llm_meta_widget 0.4.2 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 2503371d1d6229fbe6e57a807fb8d672190de589ccdb877d2a90a06feae20be0
4
- data.tar.gz: bce91755bb903ab02f886be7e9ab90733566699359e468624f3244f1202ad17a
3
+ metadata.gz: f019c93684d03e4f65ead3c8427b369af41caf2b9f4b0637f60aa2014239f080
4
+ data.tar.gz: 28bf06bc291b4bc0f745b85ae7f067f9e9c3f8adfd4902d747a4496d1b2e34da
5
5
  SHA512:
6
- metadata.gz: d833790ba0fcdb48cb291a9cde65ab45e2c07284a18a9699c755289a2ded830695e84b3125e552f770c66db8e9930e9c28410baaef2c5e0459d75ef9b93636a4
7
- data.tar.gz: 819cd9aa17be97c44b3828980b9faa6a3f79476a2802337b3e59a4149c17a45058a0355a7e5a29f39594ed8c8dfae373cb6040dde762795b6c031b9467d70128
6
+ metadata.gz: 50a79ea50cd118b4cafb1a26ff5af6713fd6b98dd1374ad4a2b2fa10fff24677fe417cee3b4cf4d900e121b3b4be3c013fcbc1aa95d4183f714e629a100432c8
7
+ data.tar.gz: 1bbed09f23f1c5c7f9b08500db77d2dd1b7ba97e1b564405b40f8cadd9fb1bac920a4d612e187119d08cd90db51301b96dc8d3d3ad22e0196adc04b5ee1621c7
data/README.md CHANGED
@@ -8,19 +8,25 @@ Client-orchestrated: the widget fetches host-side action schemas + host-publishe
8
8
 
9
9
  ## What you need first
10
10
 
11
- The widget is a browser front end. It does not talk to OpenAI, Anthropic or
12
- Ollama itself — it talks to **[llm_meta_server](https://github.com/pubannotation)**,
13
- a hub that holds the provider credentials and streams responses back over SSE.
14
- So adopting the widget means either running a hub or being given the URL of one.
11
+ The widget is a browser front end: it needs something to answer the chat. That
12
+ can be **[llm_meta_server](https://github.com/pubannotation)** — a hub that
13
+ holds provider credentials and streams responses back — or **an Ollama**,
14
+ which the widget talks to directly with no server component of ours in
15
+ between.
15
16
 
16
- Nothing else is required: no database, no migrations, no JavaScript build step,
17
- no Node at runtime. The gem ships plain ES modules that the engine serves.
17
+ Those are separate from where tools come from. A hub also registers MCP tools
18
+ (Class 1 below), and you can point at one for tools while a local Ollama
19
+ answers the chat, or use no hub at all. Page actions and your own
20
+ `.well-known/mcp.json` never involve a hub either way.
21
+
22
+ Nothing else is required: no database, no migrations, no JavaScript build
23
+ step, no Node at runtime.
18
24
 
19
25
  | Requirement | Version |
20
26
  |---|---|
21
27
  | Ruby | >= 3.2 |
22
28
  | Rails | >= 8.0 (8.1 not required) |
23
- | A reachable llm_meta_server | any |
29
+ | An llm_meta_server **or** an Ollama | any |
24
30
 
25
31
  ## Installation
26
32
 
@@ -45,7 +51,7 @@ Put this on any view — a fresh `pages/demo.html.erb` is fine:
45
51
  ```erb
46
52
  <h1>Demo</h1>
47
53
 
48
- <%= llm_meta_widget(base_url: "https://your-meta-server.example",
54
+ <%= llm_meta_widget(llm_url: "https://your-meta-server.example",
49
55
  model: "qwen3-6-35b-fast",
50
56
  greeting: "Hi — ask me anything about this page.") %>
51
57
  ```
@@ -63,8 +69,9 @@ says so explicitly.
63
69
 
64
70
  Two more things worth knowing before you go further:
65
71
 
66
- - `base_url` must be reachable from your visitors' browsers, not just from
67
- your server. `localhost` works only while you are the visitor.
72
+ - `llm_url` must be reachable from your visitors' browsers, not just from
73
+ your server. `localhost` works only while you are the visitor — which is
74
+ fine when each visitor runs their own Ollama.
68
75
  - The model name is the hub's name for it (`GET /api/llms` lists them), not
69
76
  the provider's.
70
77
 
@@ -185,7 +192,8 @@ To adjust picker behavior at the helper call site:
185
192
 
186
193
  ```erb
187
194
  <%= llm_meta_widget(
188
- base_url: "https://your-meta-server.example",
195
+ llm_url: "https://your-meta-server.example",
196
+ tool_hub_url: "https://your-meta-server.example",
189
197
  model: "qwen3-6-35b-fast", # initial selection
190
198
  enable_model_picker: true, # false → hide picker, use fixed `model:`
191
199
  enable_tool_picker: true, # false → hide picker, no Class-1 tools
@@ -197,7 +205,7 @@ To adjust picker behavior at the helper call site:
197
205
  To lock the widget to the fixed `model:` prop and disable Class-1 tools entirely (level-0 mode — Class 2 & 3 still work):
198
206
 
199
207
  ```erb
200
- <%= llm_meta_widget(base_url: "…", model: "…",
208
+ <%= llm_meta_widget(llm_url: "…", model: "…",
201
209
  enable_model_picker: false,
202
210
  enable_tool_picker: false) %>
203
211
  ```
@@ -206,7 +214,9 @@ To lock the widget to the fixed `model:` prop and disable Class-1 tools entirely
206
214
 
207
215
  | Option | Default | Purpose |
208
216
  |---|---|---|
209
- | `base_url:` | required | Meta-server URL |
217
+ | `llm_url:` | required | Who answers the chat — an llm_meta_server, or an Ollama |
218
+ | `llm_provider:` | `:llm_meta_server` | What `llm_url` speaks. `:ollama` talks to Ollama's `/api/chat` directly |
219
+ | `tool_hub_url:` | `nil` | An llm_meta_server whose registered MCP tools to offer. Absent = none (page actions and your own `.well-known` MCP are unaffected) |
210
220
  | `model:` | required | Initial model (also the fallback when picker is disabled) |
211
221
  | `api_key_uuid:` | `"ollama-local"` | Hub API-key uuid to invoke |
212
222
  | `orchestrator_path:` | `"/llm_meta_widget_assets/orchestrator.js"` | Served by the gem's engine; rarely overridden |
@@ -222,6 +232,62 @@ To lock the widget to the fixed `model:` prop and disable Class-1 tools entirely
222
232
  | `models:` | `nil` | Model-name allowlist; `nil` = all anon-available |
223
233
  | `hub_tools:` | `nil` | MCP-server-name allowlist; `nil` = all anon-public |
224
234
 
235
+ ## Surviving navigation
236
+
237
+ A page action that navigates — submitting a form, following a link — no longer
238
+ loses the conversation. The transcript, the record of which tools ran, and the
239
+ model's own context are kept in `sessionStorage`: per tab, gone when the tab
240
+ closes, never sent anywhere.
241
+
242
+ The key is the page's **path**, not its full URL, so a form submit that returns
243
+ to the same page with a new query string keeps the thread, while moving to a
244
+ different page starts a fresh one. Nothing to configure.
245
+
246
+ Two consequences worth knowing. An action of yours that navigates is now safe
247
+ to declare — before this, offering one meant offering to wipe the visitor's
248
+ chat. And if the navigation lands on an error page that does not render the
249
+ widget, the conversation is still in storage but there is no panel to show it
250
+ until the visitor returns to a page that has one.
251
+
252
+ ## Choosing what answers the chat, and whose tools to offer
253
+
254
+ Two independent settings, because they are two jobs. Neither implies the
255
+ other, so there is nothing to disable and nothing to inherit — what a widget
256
+ does is what its call site says.
257
+
258
+ ```erb
259
+ <%# a hub answers the chat; its registered tools are offered too %>
260
+ <% hub = "https://your-meta-server.example" %>
261
+ <%= llm_meta_widget(llm_url: hub, tool_hub_url: hub, model: "qwen3-6-35b-fast") %>
262
+
263
+ <%# a hub answers the chat; no hub-registered tools %>
264
+ <%= llm_meta_widget(llm_url: hub, model: "qwen3-6-35b-fast") %>
265
+
266
+ <%# the visitor's own Ollama answers; a hub still supplies its tools %>
267
+ <%= llm_meta_widget(llm_url: "http://localhost:11434", llm_provider: :ollama,
268
+ tool_hub_url: hub, model: "qwen3.8:27b") %>
269
+
270
+ <%# nothing of ours in the path at all %>
271
+ <%= llm_meta_widget(llm_url: "http://localhost:11434", llm_provider: :ollama,
272
+ model: "qwen3.8:27b") %>
273
+ ```
274
+
275
+ **The Ollama path.** The widget POSTs to `{llm_url}/api/chat` and reads
276
+ Ollama's NDJSON stream; tool schemas go over as OpenAI-shaped functions and
277
+ `tool_calls` come back, so Class 2 and Class 3 tools work exactly as they do
278
+ against a hub. The model picker lists `{llm_url}/api/tags`. Ollama must be
279
+ told to accept your page's origin — `OLLAMA_ORIGINS=https://your-site.example`
280
+ — which is the same CORS story as the hub, configured elsewhere.
281
+
282
+ What you give up without a `tool_hub_url` is Class 1 only: tools registered on
283
+ somebody's hub. That is not "no tools" — page actions and your own
284
+ `.well-known/mcp.json` are untouched, and they are the interesting ones for an
285
+ assistant embedded in your page.
286
+
287
+ **Why anyone would want the last shape:** with a local Ollama and local MCP
288
+ endpoints, nothing a visitor types leaves the machine. No credentials to hold,
289
+ no retention policy to write.
290
+
225
291
  ## Declaring resources and prompts (static-primitives extension)
226
292
 
227
293
  If you operate the MCP endpoint behind your `.well-known/mcp.json`, you can
@@ -325,6 +325,14 @@ export async function runChatLoop(opts) {
325
325
  remoteTools = [],
326
326
  hostWideTools = [],
327
327
  maxRounds = 10,
328
+ // "hub" (default) talks to llm_meta_server; "ollama" talks straight to an
329
+ // Ollama, with no server component of ours in between. The loop is the
330
+ // same either way — only the call is swapped.
331
+ provider = "hub",
332
+ // Where Class 1 (hub-registered) tools are proxied. Usually the same
333
+ // llm_meta_server that answers the chat, but not when the chat goes
334
+ // straight to an Ollama — the tools still belong to the hub.
335
+ toolHubUrl,
328
336
  signal,
329
337
  onRoundStart,
330
338
  onTextDelta,
@@ -399,7 +407,9 @@ export async function runChatLoop(opts) {
399
407
  if (signal?.aborted) throw new DOMException("aborted", "AbortError")
400
408
  onRoundStart?.(round)
401
409
 
402
- const turnResult = await singleLlmCall({
410
+ // The loop is the same whoever answers; only the call differs.
411
+ const call = provider === "ollama" ? ollamaChatCall : singleLlmCall
412
+ const turnResult = await call({
403
413
  ...singleOpts,
404
414
  messages,
405
415
  toolIds,
@@ -490,7 +500,7 @@ export async function runChatLoop(opts) {
490
500
  const args = coerceArguments(tc.arguments)
491
501
  try {
492
502
  const value = await dispatchRemoteToolCall({
493
- baseUrl: singleOpts.baseUrl,
503
+ baseUrl: toolHubUrl || singleOpts.baseUrl,
494
504
  bearerToken: singleOpts.bearerToken,
495
505
  toolId: tool.id,
496
506
  args: args,
@@ -1006,3 +1016,243 @@ export function promptButtonProps(prompt) {
1006
1016
  }
1007
1017
  }
1008
1018
 
1019
+
1020
+ // ---- direct-to-Ollama provider ------------------------------------------
1021
+ //
1022
+ // The smallest adoption: a Rails app and an Ollama, with no hub of ours in
1023
+ // between. Ollama's /api/chat already does the hard part — it takes tool
1024
+ // schemas and returns tool_calls — so what differs from the hub path is the
1025
+ // envelope: NDJSON rather than SSE, and its own request shape.
1026
+ //
1027
+ // What an adopter gives up is Class 1 (hub-registered MCP), which is the
1028
+ // hub's job by definition. Page actions and the host's own .well-known MCP
1029
+ // both work unchanged, and nothing the visitor types leaves their machine.
1030
+ //
1031
+ // Ollama must be told to accept the page's origin (OLLAMA_ORIGINS), the same
1032
+ // CORS story as the hub, configured elsewhere.
1033
+
1034
+ export async function* parseNdjsonStream(readableStream, signal) {
1035
+ const reader = readableStream.getReader()
1036
+ const decoder = new TextDecoder("utf-8")
1037
+ let buffer = ""
1038
+
1039
+ const onAbort = () => { try { reader.cancel() } catch { /* noop */ } }
1040
+ signal?.addEventListener("abort", onAbort)
1041
+
1042
+ try {
1043
+ while (true) {
1044
+ const { value, done } = await reader.read()
1045
+ if (done) break
1046
+ buffer += decoder.decode(value, { stream: true })
1047
+
1048
+ let nl
1049
+ while ((nl = buffer.indexOf("\n")) !== -1) {
1050
+ const line = buffer.slice(0, nl).trim()
1051
+ buffer = buffer.slice(nl + 1)
1052
+ if (!line) continue
1053
+ try {
1054
+ yield JSON.parse(line)
1055
+ } catch (e) {
1056
+ // A truncated or non-JSON line is not worth aborting a stream for.
1057
+ }
1058
+ }
1059
+ }
1060
+ const tail = buffer.trim()
1061
+ if (tail) { try { yield JSON.parse(tail) } catch (e) { /* ignore */ } }
1062
+ } finally {
1063
+ signal?.removeEventListener("abort", onAbort)
1064
+ try { reader.releaseLock() } catch { /* noop */ }
1065
+ }
1066
+ }
1067
+
1068
+ // Tool schemas travel as MCP-flavoured `input_schema`; Ollama wants OpenAI's
1069
+ // function shape.
1070
+ function toolsForOllama(localTools) {
1071
+ return (localTools || []).map((t) => ({
1072
+ type: "function",
1073
+ function: {
1074
+ name: t.name,
1075
+ description: t.description,
1076
+ parameters: t.input_schema || t.inputSchema || { type: "object", properties: {} }
1077
+ }
1078
+ }))
1079
+ }
1080
+
1081
+ // The hub's history carries `tool_call_id`; Ollama identifies a result by the
1082
+ // tool's name instead, and rejects unknown keys on some versions.
1083
+ function messagesForOllama(messages) {
1084
+ return (messages || []).map((m) => {
1085
+ if (m.role === "tool") {
1086
+ return { role: "tool", tool_name: m.name, content: m.content }
1087
+ }
1088
+ if (m.tool_calls) {
1089
+ return {
1090
+ role: m.role,
1091
+ content: m.content || "",
1092
+ tool_calls: m.tool_calls.map((tc) => ({
1093
+ function: { name: tc.name, arguments: coerceArguments(tc.arguments) }
1094
+ }))
1095
+ }
1096
+ }
1097
+ return { role: m.role, content: m.content }
1098
+ })
1099
+ }
1100
+
1101
+ // Same signature and return shape as singleLlmCall, so runChatLoop does not
1102
+ // care which provider answered.
1103
+ export async function ollamaChatCall({
1104
+ baseUrl,
1105
+ modelName,
1106
+ messages,
1107
+ localTools = [],
1108
+ generationSettings = {},
1109
+ onTextDelta,
1110
+ onThinkingDelta,
1111
+ onToolCall,
1112
+ onPhase,
1113
+ signal,
1114
+ }) {
1115
+ if (!baseUrl || !modelName) throw new Error("ollamaChatCall: baseUrl and modelName are required")
1116
+ if (!Array.isArray(messages) || messages.length === 0) {
1117
+ throw new Error("ollamaChatCall: messages must be a non-empty array")
1118
+ }
1119
+
1120
+ const { think, ...options } = generationSettings || {}
1121
+ const body = { model: modelName, messages: messagesForOllama(messages), stream: true }
1122
+ if (localTools && localTools.length) body.tools = toolsForOllama(localTools)
1123
+ if (think !== undefined) body.think = think
1124
+ if (Object.keys(options).length) body.options = options
1125
+
1126
+ const response = await fetch(`${baseUrl.replace(/\/$/, "")}/api/chat`, {
1127
+ method: "POST",
1128
+ headers: { "Content-Type": "application/json" },
1129
+ body: JSON.stringify(body),
1130
+ signal
1131
+ })
1132
+ if (!response.ok) {
1133
+ const text = await response.text().catch(() => "")
1134
+ throw new Error(`ollamaChatCall: HTTP ${response.status} ${response.statusText}${text ? " — " + text.slice(0, 200) : ""}`)
1135
+ }
1136
+ if (!response.body) throw new Error("ollamaChatCall: response has no body (streaming unsupported?)")
1137
+
1138
+ const toolCalls = []
1139
+ let content = ""
1140
+ let finishReason = null
1141
+ let announcedPhase = null
1142
+
1143
+ const phase = (name) => {
1144
+ if (announcedPhase === name) return
1145
+ announcedPhase = name
1146
+ onPhase?.(name)
1147
+ }
1148
+
1149
+ for await (const frame of parseNdjsonStream(response.body, signal)) {
1150
+ const message = frame.message || {}
1151
+
1152
+ if (message.thinking) {
1153
+ phase("thinking")
1154
+ onThinkingDelta?.(message.thinking)
1155
+ }
1156
+ if (message.content) {
1157
+ phase("responding")
1158
+ content += message.content
1159
+ onTextDelta?.(message.content)
1160
+ }
1161
+ for (const tc of message.tool_calls || []) {
1162
+ const call = {
1163
+ id: tc.id || `ollama-${toolCalls.length}`,
1164
+ name: tc.function?.name,
1165
+ arguments: tc.function?.arguments ?? {}
1166
+ }
1167
+ toolCalls.push(call)
1168
+ onToolCall?.(call)
1169
+ }
1170
+ if (frame.done) finishReason = frame.done_reason || "stop"
1171
+ }
1172
+
1173
+ return { content, toolCalls, finishReason }
1174
+ }
1175
+
1176
+ // Ollama's own catalogue, for the model picker when there is no hub to ask.
1177
+ export async function fetchOllamaModels({ baseUrl, signal }) {
1178
+ try {
1179
+ const response = await fetch(`${baseUrl.replace(/\/$/, "")}/api/tags`, { signal })
1180
+ if (!response.ok) return []
1181
+ const data = await response.json()
1182
+ return (data.models || []).map((m) => m.name).filter(Boolean)
1183
+ } catch (e) {
1184
+ // A picker that cannot be populated is not a reason to break the widget.
1185
+ return []
1186
+ }
1187
+ }
1188
+
1189
+ // ---- surviving navigation ------------------------------------------------
1190
+ //
1191
+ // A page action that navigates — submitting a form, following a link — used to
1192
+ // destroy the conversation, the record of what ran and the scroll position.
1193
+ // That made such actions unsafe to offer at all: PubDictionaries had to stop
1194
+ // declaring its submit action rather than let the assistant wipe the chat.
1195
+ //
1196
+ // The transcript is kept in sessionStorage: per tab, gone when the tab closes,
1197
+ // invisible to other visitors. Keyed by pathname rather than full URL, because
1198
+ // submitting a form usually returns to the same page with different query
1199
+ // parameters and that is exactly the case worth surviving.
1200
+
1201
+ export const CONVERSATION_FORMAT = 1
1202
+
1203
+ export function packConversation({ turns = [], open = false, model = null,
1204
+ maxTurns = 40, maxBytes = 200000 } = {}) {
1205
+ // Oldest turns go first when trimming: the recent ones carry the thread.
1206
+ let kept = turns.slice(-maxTurns)
1207
+ let payload = { v: CONVERSATION_FORMAT, savedAt: Date.now(), open, model, turns: kept }
1208
+ let json = JSON.stringify(payload)
1209
+
1210
+ while (json.length > maxBytes && kept.length > 1) {
1211
+ kept = kept.slice(1)
1212
+ payload = { ...payload, turns: kept }
1213
+ json = JSON.stringify(payload)
1214
+ }
1215
+ return json
1216
+ }
1217
+
1218
+ export function unpackConversation(raw) {
1219
+ if (!raw) return null
1220
+ let payload
1221
+ try {
1222
+ payload = JSON.parse(raw)
1223
+ } catch (e) {
1224
+ return null
1225
+ }
1226
+ // A payload from a different format is not worth guessing at; starting
1227
+ // fresh is better than rendering something half-understood.
1228
+ if (!payload || payload.v !== CONVERSATION_FORMAT || !Array.isArray(payload.turns)) return null
1229
+
1230
+ const turns = payload.turns.filter((t) => t && typeof t.content === "string" &&
1231
+ (t.role === "user" || t.role === "assistant"))
1232
+ return { turns, open: !!payload.open, model: payload.model || null, savedAt: payload.savedAt || 0 }
1233
+ }
1234
+
1235
+ // Storage can throw — private windows, blocked site data, quota — and none of
1236
+ // that is a reason for the widget to stop working.
1237
+ export function createConversationStore({ storage, key }) {
1238
+ const safely = (fn, fallback = null) => {
1239
+ try {
1240
+ return fn()
1241
+ } catch (e) {
1242
+ return fallback
1243
+ }
1244
+ }
1245
+
1246
+ return {
1247
+ key,
1248
+ save(state) { return safely(() => { storage.setItem(key, packConversation(state)); return true }, false) },
1249
+ load() { return safely(() => unpackConversation(storage.getItem(key))) },
1250
+ clear() { return safely(() => { storage.removeItem(key); return true }, false) }
1251
+ }
1252
+ }
1253
+
1254
+ export function conversationKeyFor(location, prefix = "llm_meta_widget") {
1255
+ // Path only: a form submit returns to the same page with a new query string,
1256
+ // and that conversation is still the same conversation.
1257
+ return `${prefix}:${location.pathname}`
1258
+ }
@@ -16,7 +16,18 @@ module LlmMetaWidget
16
16
  # (page-embedded aiActions / host-wide well-known / hub-registered).
17
17
  module WidgetHelper
18
18
  DEFAULTS = {
19
- api_key_uuid: "ollama-local",
19
+ api_key_uuid: "ollama-local", # llm_meta_server provider only
20
+ # Answering the chat and registering tools are separate jobs. Name each
21
+ # endpoint for the job it does; neither implies the other.
22
+ #
23
+ # llm_url: who answers the chat (required)
24
+ # llm_provider: what it speaks — :llm_meta_server (default) or :ollama
25
+ # tool_hub_url: an llm_meta_server whose registered MCP tools to
26
+ # offer. Absent means none — which does NOT mean "no
27
+ # tools": page actions and the host's own
28
+ # .well-known/mcp.json are unaffected.
29
+ llm_provider: "llm_meta_server",
30
+ tool_hub_url: nil,
20
31
  orchestrator_path: "/llm_meta_widget_assets/orchestrator.js",
21
32
  actions_schema_id: "ai-actions",
22
33
  state_global: "aiState",
@@ -42,9 +53,20 @@ module LlmMetaWidget
42
53
  greeting: nil
43
54
  }.freeze
44
55
 
45
- def llm_meta_widget(base_url:, model:, **overrides)
46
- locals = DEFAULTS.merge(base_url: base_url, model: model, **overrides)
47
- render partial: "llm_meta_widget/chat_panel", locals: locals
56
+ def llm_meta_widget(llm_url:, model:, **overrides)
57
+ locals = DEFAULTS.merge(llm_url: llm_url, model: model, **overrides)
58
+ provider = locals[:llm_provider].to_s
59
+ unless %w[llm_meta_server ollama].include?(provider)
60
+ raise ArgumentError,
61
+ "llm_meta_widget: llm_provider must be :llm_meta_server or :ollama, got #{provider.inspect}"
62
+ end
63
+ raise ArgumentError, "llm_meta_widget: llm_url is required" if llm_url.to_s.strip.empty?
64
+
65
+ hub = locals[:tool_hub_url]
66
+ hub = nil if hub.to_s.strip.empty?
67
+
68
+ render partial: "llm_meta_widget/chat_panel",
69
+ locals: locals.merge(llm_provider: provider, tool_hub_url: hub)
48
70
  end
49
71
  end
50
72
  end
@@ -452,7 +452,8 @@
452
452
  import { runChatLoop, fetchMcpManifest, listMcpPrompts, getMcpPrompt,
453
453
  listMcpResources, readMcpResource, promptMessagesToText,
454
454
  loadHostResource, resourceLinesForTurn, resolvePromptArguments,
455
- promptButtonProps, promptArgumentSummary } from "<%= orchestrator_path %>";
455
+ promptButtonProps, promptArgumentSummary, fetchOllamaModels,
456
+ createConversationStore, conversationKeyFor } from "<%= orchestrator_path %>";
456
457
  import { marked } from "/llm_meta_widget_assets/marked.esm.js";
457
458
 
458
459
  // Standard prose settings — GFM (tables, autolinks, strikethrough),
@@ -460,7 +461,12 @@
460
461
  // bare newlines still reads naturally.
461
462
  marked.setOptions({ gfm: true, breaks: true });
462
463
 
463
- var META_BASE = <%= base_url.to_json.html_safe %>;
464
+ // Two endpoints, two jobs: one answers the chat, the other registers the
465
+ // MCP tools on offer. They are usually the same llm_meta_server, but need
466
+ // not be — an adopter can run their own Ollama and still borrow a hub's
467
+ // tools, or use no hub at all.
468
+ var LLM_BASE = <%= llm_url.to_json.html_safe %>;
469
+ var TOOL_HUB_BASE = <%= tool_hub_url.to_json.html_safe %>;
464
470
  var API_KEY_UUID = <%= api_key_uuid.to_json.html_safe %>;
465
471
  var MODEL = <%= model.to_json.html_safe %>;
466
472
  var ACTIONS_SCHEMA_ID = <%= actions_schema_id.to_json.html_safe %>;
@@ -470,8 +476,12 @@
470
476
  var REMOTE_TOOLS_SCHEMA_ID = <%= remote_tools_schema_id.to_json.html_safe %>;
471
477
  var MAX_ROUNDS = <%= max_rounds.to_json.html_safe %>;
472
478
  var WELL_KNOWN_URLS = <%= raw(well_known_urls.nil? ? "null" : well_known_urls.to_json) %>;
479
+ var LLM_PROVIDER = <%= llm_provider.to_json.html_safe %>;
473
480
  var ENABLE_MODEL_PICKER = <%= enable_model_picker.to_json.html_safe %>;
474
- var ENABLE_TOOL_PICKER = <%= enable_tool_picker.to_json.html_safe %>;
481
+ // Class 1 tools are registered on a hub, so the picker needs one — which
482
+ // is independent of who answers the chat. Page actions and the host's own
483
+ // .well-known MCP work regardless.
484
+ var ENABLE_TOOL_PICKER = <%= enable_tool_picker.to_json.html_safe %> && <%= tool_hub_url.to_json.html_safe %> !== null;
475
485
  var MODEL_ALLOWLIST = <%= raw(models.nil? ? "null" : models.to_json) %>;
476
486
  var HUB_TOOLS_ALLOWLIST = <%= raw(hub_tools.nil? ? "null" : hub_tools.to_json) %>;
477
487
 
@@ -567,7 +577,26 @@
567
577
  ro.observe(root);
568
578
  }
569
579
 
570
- var conversation = []; // [{role, content}, ...] — grows across turns
580
+ // [{role, content, tools}] — grows across turns. `tools` is for the
581
+ // transcript only; the model is sent role and content.
582
+ var conversation = [];
583
+
584
+ // A page action that navigates used to take the conversation with it. The
585
+ // transcript lives in sessionStorage — per tab, gone when the tab closes —
586
+ // keyed by path, so submitting a form and landing back on the same page
587
+ // with a new query string keeps the thread.
588
+ var conversationStore = createConversationStore({
589
+ storage: window.sessionStorage,
590
+ key: conversationKeyFor(window.location)
591
+ });
592
+
593
+ function persistConversation() {
594
+ conversationStore.save({
595
+ turns: conversation,
596
+ open: !root.classList.contains("lmw-collapsed"),
597
+ model: MODEL
598
+ });
599
+ }
571
600
  var actionsSchemaEl = document.getElementById(ACTIONS_SCHEMA_ID);
572
601
  var localTools = actionsSchemaEl ? JSON.parse(actionsSchemaEl.textContent) : [];
573
602
  var remoteToolsEl = document.getElementById(REMOTE_TOOLS_SCHEMA_ID);
@@ -633,8 +662,27 @@
633
662
  // still works with the initial `model:` + no hub tools).
634
663
  var tasks = [];
635
664
 
636
- if (ENABLE_MODEL_PICKER && modelPicker) {
637
- tasks.push(fetch(META_BASE + "/api/llms", { headers: { "Accept": "application/json" } })
665
+ if (ENABLE_MODEL_PICKER && modelPicker && LLM_PROVIDER === "ollama") {
666
+ // No hub to ask: Ollama lists its own models.
667
+ tasks.push(fetchOllamaModels({ baseUrl: LLM_BASE }).then(function(names) {
668
+ var flat = names
669
+ .filter(function(n) { return anyAllowedByAllowlist(n, MODEL_ALLOWLIST); })
670
+ .map(function(n) { return { value: n, label: n }; });
671
+ if (!flat.some(function(m) { return m.value === MODEL; })) {
672
+ flat.unshift({ value: MODEL, label: MODEL });
673
+ }
674
+ modelPicker.innerHTML = "";
675
+ flat.forEach(function(m) {
676
+ var opt = document.createElement("option");
677
+ opt.value = m.value;
678
+ opt.textContent = m.label;
679
+ if (m.value === MODEL) opt.selected = true;
680
+ modelPicker.appendChild(opt);
681
+ });
682
+ if (flat.length > 1) modelPicker.style.display = "";
683
+ }));
684
+ } else if (ENABLE_MODEL_PICKER && modelPicker) {
685
+ tasks.push(fetch(LLM_BASE + "/api/llms", { headers: { "Accept": "application/json" } })
638
686
  .then(function(r) { return r.ok ? r.json() : { llms: [] }; })
639
687
  .then(function(payload) {
640
688
  // /api/llms is heterogeneous per family:
@@ -671,7 +719,7 @@
671
719
  }
672
720
 
673
721
  if (ENABLE_TOOL_PICKER && toolsListEl) {
674
- tasks.push(fetch(META_BASE + "/api/mcp_servers", { headers: { "Accept": "application/json" } })
722
+ tasks.push(fetch(TOOL_HUB_BASE + "/api/mcp_servers", { headers: { "Accept": "application/json" } })
675
723
  .then(function(r) { return r.ok ? r.json() : { mcp_servers: [] }; })
676
724
  .then(function(payload) {
677
725
  hubMcpServers = (payload.mcp_servers || []).filter(function(s) {
@@ -817,6 +865,10 @@
817
865
  var RESOURCE_BUDGET_BYTES = 8000;
818
866
 
819
867
 
868
+ // Before anything asynchronous: if this page was navigated away from and
869
+ // back, the visitor should see their conversation immediately.
870
+ restoreConversation();
871
+
820
872
  var wellKnownReady = (async function() {
821
873
  var urls = WELL_KNOWN_URLS === null
822
874
  ? [ window.location.origin + "/.well-known/mcp.json" ]
@@ -893,11 +945,39 @@
893
945
  promptsEl.style.display = "";
894
946
  }
895
947
 
948
+ // Put a saved transcript back on screen: the same bubbles, the same
949
+ // markdown, and the same record of which tools ran. Called once at boot,
950
+ // before the welcome block, which is only for a genuinely fresh start.
951
+ function restoreConversation() {
952
+ var saved = conversationStore.load();
953
+ if (!saved || !saved.turns.length) return false;
954
+
955
+ conversation = saved.turns;
956
+ saved.turns.forEach(function(turn) {
957
+ var body = appendTurn(turn.role, turn.role === "assistant" ? "" : turn.content);
958
+ if (turn.role === "assistant") renderMarkdownInto(body, turn.content || "");
959
+ if (!turn.tools || !turn.tools.length) return;
960
+ var row = document.createElement("div");
961
+ row.className = "lmw-tool-chips";
962
+ turn.tools.forEach(function(tool) {
963
+ var chip = document.createElement("span");
964
+ chip.className = "lmw-tool-chip" + (tool.error ? " error" : "");
965
+ chip.textContent = (tool.error ? "❌ " : "🔧 ") + tool.name;
966
+ row.appendChild(chip);
967
+ });
968
+ if (body.parentNode) body.parentNode.appendChild(row);
969
+ });
970
+ historyEl.scrollTop = historyEl.scrollHeight;
971
+ return true;
972
+ }
973
+
896
974
  // A blank panel tells a first-time visitor nothing. Open with a greeting
897
975
  // and the offers themselves — each template showing what it will take from
898
976
  // the page and what it will ask for — so the assistant is the page's way
899
977
  // in rather than a box you must already know how to talk to.
900
978
  function renderWelcome() {
979
+ // A restored transcript is not a fresh start; the greeting would push
980
+ // the conversation the visitor is mid-way through off the screen.
901
981
  if (!historyEl || historyEl.querySelector(".message")) return;
902
982
  historyEl.textContent = "";
903
983
 
@@ -1210,6 +1290,7 @@
1210
1290
  clearBtn.addEventListener("click", function() {
1211
1291
  if (currentAbort) { try { currentAbort.abort(); } catch (e) { /* noop */ } }
1212
1292
  conversation = [];
1293
+ conversationStore.clear();
1213
1294
  historyEl.innerHTML = "";
1214
1295
  renderWelcome();
1215
1296
  });
@@ -1252,11 +1333,12 @@
1252
1333
 
1253
1334
  var resourceLines = await resourceLinesForThisTurn();
1254
1335
  var messages = [{ role: "system", content: currentSystemPrompt(resourceLines) }]
1255
- .concat(conversation)
1336
+ .concat(conversation.map(function(t) { return { role: t.role, content: t.content }; }))
1256
1337
  .concat([{ role: "user", content: userText }]);
1257
1338
 
1258
1339
  var result = await runChatLoop({
1259
- baseUrl: META_BASE,
1340
+ baseUrl: LLM_BASE,
1341
+ toolHubUrl: TOOL_HUB_BASE,
1260
1342
  apiKeyUuid: API_KEY_UUID,
1261
1343
  modelName: MODEL,
1262
1344
  messages: messages,
@@ -1265,6 +1347,7 @@
1265
1347
  hostWideTools: hostWideTools,
1266
1348
  aiActions: window[ACTIONS_GLOBAL] || {},
1267
1349
  maxRounds: MAX_ROUNDS,
1350
+ provider: LLM_PROVIDER === "ollama" ? "ollama" : "hub",
1268
1351
  signal: currentAbort.signal,
1269
1352
  onToolCall: function(toolCall) { announceToolCall(toolCall); },
1270
1353
  onToolDispatched: function(outcome) { resolveToolCall(outcome); },
@@ -1301,8 +1384,17 @@
1301
1384
  historyEl.scrollTop = historyEl.scrollHeight;
1302
1385
  }
1303
1386
  });
1304
- conversation.push({ role: "user", content: userText });
1305
- conversation.push({ role: "assistant", content: result.content });
1387
+ conversation.push({ role: "user", content: userText });
1388
+ conversation.push({
1389
+ role: "assistant",
1390
+ content: result.content,
1391
+ // Kept so a restored transcript still shows what ran — losing
1392
+ // that record is what made a navigating action unbearable.
1393
+ tools: result.dispatched.map(function(d) {
1394
+ return d.error ? { name: d.toolCall.name, error: true } : { name: d.toolCall.name };
1395
+ })
1396
+ });
1397
+ persistConversation();
1306
1398
 
1307
1399
  // Anything still pending never reported an outcome (a class that
1308
1400
  // does not round-trip, or a dispatch that vanished). Reconcile
@@ -1,3 +1,3 @@
1
1
  module LlmMetaWidget
2
- VERSION = "0.4.2"
2
+ VERSION = "0.6.0"
3
3
  end
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: llm_meta_widget
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.4.2
4
+ version: 0.6.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - jdkim