llm_meta_widget 0.4.0 → 0.4.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 20aa5babd648e80015577ece5c815b37f6fe5e06d6136c91647f49831016833b
4
- data.tar.gz: 8dfc055787110255fbf32368bc3f29cf11b3be178477dac9dadf08243a48d096
3
+ metadata.gz: 2503371d1d6229fbe6e57a807fb8d672190de589ccdb877d2a90a06feae20be0
4
+ data.tar.gz: bce91755bb903ab02f886be7e9ab90733566699359e468624f3244f1202ad17a
5
5
  SHA512:
6
- metadata.gz: 2545397e740fefb805c8b8b3f9c90ca78814d71b2138f91a35afd38e1499f9bace74428faaf9160cbda4d9a5b51421a81f16f51df82c7ce88389d07acb97293d
7
- data.tar.gz: a05e51a8b8e94a0153d5c7612901cba9ff2f31beb3142e656852cfa43c87b65ecc513d42c2a6a7a9b2a7ef6b0be3f240ec067bd468afd4b320684089cddfff3b
6
+ metadata.gz: d833790ba0fcdb48cb291a9cde65ab45e2c07284a18a9699c755289a2ded830695e84b3125e552f770c66db8e9930e9c28410baaef2c5e0459d75ef9b93636a4
7
+ data.tar.gz: 819cd9aa17be97c44b3828980b9faa6a3f79476a2802337b3e59a4149c17a45058a0355a7e5a29f39594ed8c8dfae373cb6040dde762795b6c031b9467d70128
data/README.md CHANGED
@@ -6,28 +6,71 @@ Client-orchestrated: the widget fetches host-side action schemas + host-publishe
6
6
 
7
7
  **No** Devise, DB migrations, ChatManager, or PromptNavigator. Adds only `rails >= 8.0` as a runtime dep — so hosts that haven't bumped to 8.1 can adopt it without a Rails upgrade.
8
8
 
9
- ## Installation (local, path-based)
9
+ ## What you need first
10
10
 
11
- Add to the host app's Gemfile:
11
+ The widget is a browser front end. It does not talk to OpenAI, Anthropic or
12
+ Ollama itself — it talks to **[llm_meta_server](https://github.com/pubannotation)**,
13
+ a hub that holds the provider credentials and streams responses back over SSE.
14
+ So adopting the widget means either running a hub or being given the URL of one.
15
+
16
+ Nothing else is required: no database, no migrations, no JavaScript build step,
17
+ no Node at runtime. The gem ships plain ES modules that the engine serves.
18
+
19
+ | Requirement | Version |
20
+ |---|---|
21
+ | Ruby | >= 3.2 |
22
+ | Rails | >= 8.0 (8.1 not required) |
23
+ | A reachable llm_meta_server | any |
24
+
25
+ ## Installation
12
26
 
13
27
  ```ruby
14
- gem "llm_meta_widget", path: "../llm_meta/llm_meta_widget"
28
+ # Gemfile
29
+ gem "llm_meta_widget", "~> 0.4"
30
+ ```
31
+
32
+ ```
33
+ bundle install
15
34
  ```
16
35
 
17
- Then `bundle install`. The engine auto-includes the widget helper into ActionView, so nothing else to wire up.
36
+ That is the whole install. There is **no `mount` to add** — the engine
37
+ deliberately skips `isolate_namespace`, so the helper is included into
38
+ ActionView and the asset routes appear at the host's top level on their own.
39
+ There are no migrations and no generators to run.
18
40
 
19
- ## Usage
41
+ ## Minimal working example
20
42
 
21
- On any view where you want the widget:
43
+ Put this on any view — a fresh `pages/demo.html.erb` is fine:
22
44
 
23
45
  ```erb
46
+ <h1>Demo</h1>
47
+
24
48
  <%= llm_meta_widget(base_url: "https://your-meta-server.example",
25
- model: "qwen3-6-35b-fast") %>
49
+ model: "qwen3-6-35b-fast",
50
+ greeting: "Hi — ask me anything about this page.") %>
26
51
  ```
27
52
 
28
- That's the minimum. The widget renders a floating chat panel; visitors get a message-list, an input box, and (by default) a model dropdown + tool picker below the input.
53
+ Start the app, open that page, and click the round toggle at the bottom right.
54
+ **Success looks like:** a panel opens showing your greeting, you type "hello",
55
+ the role label turns into a spinning gear reading "Working…", and the reply
56
+ streams in.
57
+
58
+ If the panel opens but nothing comes back, it is almost always CORS. The hub
59
+ must list your app's origin in its `config/initializers/cors.rb` for the
60
+ `/api/*` resource, because the widget calls the hub **from the visitor's
61
+ browser**, not from your server. Open the browser console: a blocked request
62
+ says so explicitly.
63
+
64
+ Two more things worth knowing before you go further:
65
+
66
+ - `base_url` must be reachable from your visitors' browsers, not just from
67
+ your server. `localhost` works only while you are the visitor.
68
+ - The model name is the hub's name for it (`GET /api/llms` lists them), not
69
+ the provider's.
29
70
 
30
- To make the LLM able to do more than converse, declare tools using any of the **three tool classes** below.
71
+ Once that works, the widget can converse but cannot *do* anything. To let the
72
+ LLM act on your page or call your own services, declare tools using any of the
73
+ **three tool classes** below.
31
74
 
32
75
  ## Three tool classes
33
76
 
@@ -72,7 +115,17 @@ Declared inline on the same view as the widget. Runs as JavaScript in the browse
72
115
 
73
116
  **What the LLM sees.** The `ai-actions` JSON is passed to the meta-server as `local_tools` for every turn. Alongside it, every reader in `window.aiState` is invoked (each turn) and the results are JSON-serialized into a `Current page state:` block appended to the system prompt — the LLM can answer from state directly instead of tool-calling for lookups.
74
117
 
75
- **When the tool_call fires.** Fire-and-forget, dispatched **AFTER** the LLM's response text finishes streaming. The action's return value is NOT fed back to the LLM this turn — the model already produced its user-facing answer and moved on. This is a deliberate design choice (see `project_page_embedded_actions` memory): writes are separated from the conversation loop so the LLM can commit to an action without waiting for a round-trip. Use Class 1/2 instead when you need the LLM to actually see the tool's result.
118
+ **When the tool_call fires.** During the turn, as soon as the LLM emits the
119
+ call. Since 0.4.0 the action's outcome — `{ok: true, applied: "<name>"}`, or
120
+ the error message if it raised — is fed back to the LLM as a tool result, and
121
+ the turn continues. Before 0.4.0 these were fire-and-forget, which meant a
122
+ turn whose only calls were page actions ended there: any task shaped *change
123
+ the page, then do something with it* was cut off after the write. The change
124
+ also means the LLM can see that an action failed and correct itself, instead
125
+ of the failure being visible only to the visitor.
126
+
127
+ Budget for it: every page action now costs a round, so a flow that writes twice
128
+ and then calls a tool needs roughly five or six. See `max_rounds:` below.
76
129
 
77
130
  **Args**: parsed JS object matching the `input_schema`. Return value ignored. Async allowed (widget doesn't await, but browser will still execute the promise). Exceptions logged and surfaced as an `❌ <name>` chip in the message footer.
78
131
 
@@ -115,7 +168,7 @@ Declared out-of-band on the meta-server (via the hub's admin UI or `/user/:id/mc
115
168
 
116
169
  **What the LLM sees.** Only tools from servers the visitor has **enabled via the tool picker** (see "Level-1 pickers" below). Nothing is auto-selected — the visitor opts in per session.
117
170
 
118
- **When the tool_call fires.** Synchronously during the turn — widget POSTs to the hub's `/api/llm_api_keys/:uuid/models/:name/chat_streams` endpoint with `tool_ids: [...]`; the hub proxies to each MCP server and streams results back through SSE.
171
+ **When the tool_call fires.** Synchronously during the turn — widget POSTs to the hub's `/api/llm_api_keys/:uuid/models/:name/single_llm_calls` endpoint with `tool_ids: [...]`; the hub proxies to each MCP server and streams results back through SSE.
119
172
 
120
173
  **Error path.** Hub-side errors (rate limit, timeout, MCP server unavailable, upstream failure) arrive as SSE `event: error` frames with typed codes (`mcp_unavailable`, `timeout`, `rate_limit`, …); the widget surfaces them in the message bubble with a per-code prefix.
121
174
 
@@ -162,12 +215,69 @@ To lock the widget to the fixed `model:` prop and disable Class-1 tools entirely
162
215
  | `actions_global:` | `"aiActions"` | Global window object holding Class-3 implementations |
163
216
  | `remote_tools_schema_id:` | `"remote-mcp-tools"` | Optional DOM id for pre-configured Class-1 tools (bypasses picker) |
164
217
  | `well_known_urls:` | `nil` | `nil` = auto-discover same-origin; explicit array = fetch those; `[]` = disable |
165
- | `max_rounds:` | `3` | Cap on tool-call rounds per LLM turn |
218
+ | `greeting:` | `nil` | First thing a visitor sees when the panel opens, above the offered prompt templates. `nil` = a generic line |
219
+ | `max_rounds:` | `3` | Cap on tool-call rounds per LLM turn. Page actions cost a round each since 0.4.0 — raise it for multi-step flows |
166
220
  | `enable_model_picker:` | `true` | Show model dropdown (Level-1) |
167
221
  | `enable_tool_picker:` | `true` | Show tool picker (Level-1) |
168
222
  | `models:` | `nil` | Model-name allowlist; `nil` = all anon-available |
169
223
  | `hub_tools:` | `nil` | MCP-server-name allowlist; `nil` = all anon-public |
170
224
 
225
+ ## Declaring resources and prompts (static-primitives extension)
226
+
227
+ If you operate the MCP endpoint behind your `.well-known/mcp.json`, you can
228
+ also offer **prompts** (one-click starting points, shown as buttons above the
229
+ input) and **resources** (reference data the widget can put in the model's
230
+ context). The widget reads two optional fields the
231
+ `io.modelcontextprotocol/static-primitives` extension defines on a
232
+ `resources/list` entry:
233
+
234
+ ```json
235
+ {
236
+ "uri": "yourhost://catalog",
237
+ "name": "Catalog",
238
+ "mimeType": "application/json",
239
+ "_meta": {
240
+ "io.modelcontextprotocol/static-primitives": {
241
+ "sizeBytes": 1298,
242
+ "volatility": "stable",
243
+ "autoAttach": true
244
+ }
245
+ }
246
+ }
247
+ ```
248
+
249
+ - **`sizeBytes`** — the byte length of the payload `resources/read` returns.
250
+ The widget decides whether it can afford the resource *before* fetching it;
251
+ over its budget (8 KB), the bytes never cross the wire and the LLM is left
252
+ to ask for what it needs through your tools instead.
253
+ - **`volatility`** — `"stable"` (read once, re-used) or `"volatile"` (re-read
254
+ on every turn that uses it).
255
+ - **`autoAttach`** — may the widget put this in the model's context without
256
+ the visitor asking?
257
+
258
+ All three are optional. A server that declares none behaves exactly as one
259
+ that predates the extension: the resource is fetched, trimmed to budget and
260
+ carried in the model's context.
261
+
262
+ A prompt is offered as a button labelled by its `title`. Declare its arguments
263
+ as **optional** and fill them from page state where you can: an argument the
264
+ visitor must supply by hand makes the button useless to the visitor who needed
265
+ it most, since they must already know your form to press it.
266
+
267
+ ## Further reading
268
+
269
+ - **The static-primitives extension** — `io.modelcontextprotocol/static-primitives`,
270
+ a follow-on to [SEP-2127](https://github.com/modelcontextprotocol/modelcontextprotocol)
271
+ (Server Cards). The fields above are a reference implementation of its
272
+ current draft shape.
273
+ - **MCP itself** — <https://modelcontextprotocol.io>
274
+ - **A worked adopter** — PubDictionaries' annotation page
275
+ (<https://pubdictionaries.org/text_annotation>) runs this widget with all
276
+ three tool classes plus prompts and resources. Its
277
+ `app/views/annotation/text_annotation.html.erb` and `app/controllers/mcp_controller.rb`
278
+ are the fullest example available.
279
+ - **Issues and questions** — <https://github.com/jdkim/llm_meta_widget/issues>
280
+
171
281
  ## License
172
282
 
173
283
  Apache-2.0.
@@ -32,6 +32,7 @@
32
32
  // onTextDelta: (str) => {},
33
33
  // onThinkingDelta: (str) => {},
34
34
  // onToolCall: (tc) => {}, // { id, name, arguments }
35
+ // onToolDispatched: ({toolCall, value, error}) => {}, // after it ran
35
36
  // onPhase: (name) => {}, // 'thinking' | 'tool_execution' | ...
36
37
  // signal: abortController.signal
37
38
  // })
@@ -329,6 +330,11 @@ export async function runChatLoop(opts) {
329
330
  onTextDelta,
330
331
  onThinkingDelta,
331
332
  onToolCall,
333
+ // Fired as each tool finishes, so a caller can show what ran WHILE the
334
+ // turn is still going. The returned `dispatched` array says the same
335
+ // thing, but only once the whole loop ends — which on a slow model is
336
+ // minutes after the work happened, with the user watching unnamed tools.
337
+ onToolDispatched,
332
338
  onPhase,
333
339
  ...singleOpts
334
340
  } = opts
@@ -454,7 +460,8 @@ export async function runChatLoop(opts) {
454
460
  // the write. It also means a failed action is something the model can
455
461
  // see and correct, instead of a red mark only the user notices.
456
462
  const roundTripResults = []
457
- for (const { toolCall, error } of localOut.dispatched) {
463
+ for (const { toolCall, value, error } of localOut.dispatched) {
464
+ onToolDispatched?.({ toolCall, value, error })
458
465
  roundTripResults.push({
459
466
  tc: toolCall,
460
467
  result: error ? { error: String(error.message || error) } : { ok: true, applied: toolCall.name }
@@ -468,9 +475,11 @@ export async function runChatLoop(opts) {
468
475
  try {
469
476
  const value = await callMcpTool({ endpoint: tool.endpoint, name: tc.name, args, signal })
470
477
  allDispatched.push({ toolCall: tc, value })
478
+ onToolDispatched?.({ toolCall: tc, value })
471
479
  roundTripResults.push({ tc, result: value })
472
480
  } catch (error) {
473
481
  allDispatched.push({ toolCall: tc, error })
482
+ onToolDispatched?.({ toolCall: tc, error })
474
483
  roundTripResults.push({ tc, result: { error: String(error.message || error) } })
475
484
  }
476
485
  }
@@ -488,9 +497,11 @@ export async function runChatLoop(opts) {
488
497
  signal
489
498
  })
490
499
  allDispatched.push({ toolCall: tc, value })
500
+ onToolDispatched?.({ toolCall: tc, value })
491
501
  roundTripResults.push({ tc, result: value })
492
502
  } catch (error) {
493
503
  allDispatched.push({ toolCall: tc, error })
504
+ onToolDispatched?.({ toolCall: tc, error })
494
505
  // Feed the error text back to the LLM as the tool result — better
495
506
  // than dropping it (the LLM can react, apologize, retry differently).
496
507
  roundTripResults.push({ tc, result: { error: String(error.message || error) } })
@@ -1070,8 +1070,54 @@
1070
1070
  label.textContent = roleLabel("assistant");
1071
1071
  }
1072
1072
 
1073
+ // Tool chips, written as the tools run rather than assembled at the end.
1074
+ // The end-of-turn footer was invisible for as long as the turn lasted —
1075
+ // minutes on a thinking model — so a visitor watched tools execute with
1076
+ // no idea which ones.
1077
+ var liveChips = null; // the container inside the current bubble
1078
+ var pendingChips = []; // chips awaiting their outcome, oldest first
1079
+
1080
+ function chipRow() {
1081
+ if (liveChips && liveChips.parentNode) return liveChips;
1082
+ if (!activeAssistantBody || !activeAssistantBody.parentNode) return null;
1083
+ liveChips = document.createElement("div");
1084
+ liveChips.className = "lmw-tool-chips";
1085
+ // Inside the bubble but OUTSIDE .message-content: every text delta
1086
+ // re-renders that element's markdown from scratch, which silently
1087
+ // erased any chip already written into it.
1088
+ activeAssistantBody.parentNode.appendChild(liveChips);
1089
+ return liveChips;
1090
+ }
1091
+
1092
+ function announceToolCall(toolCall) {
1093
+ var row = chipRow();
1094
+ if (!row) return;
1095
+ var chip = document.createElement("span");
1096
+ chip.className = "lmw-tool-chip running";
1097
+ chip.textContent = "⏳ " + toolCall.name;
1098
+ chip.dataset.tool = toolCall.name;
1099
+ row.appendChild(chip);
1100
+ pendingChips.push(chip);
1101
+ historyEl.scrollTop = historyEl.scrollHeight;
1102
+ }
1103
+
1104
+ function resolveToolCall(outcome) {
1105
+ // Match by name, oldest first: ids are not always echoed back, and a
1106
+ // repeated name resolves in call order.
1107
+ var idx = pendingChips.findIndex(function(c) { return c.dataset.tool === outcome.toolCall.name; });
1108
+ var chip = idx === -1 ? null : pendingChips.splice(idx, 1)[0];
1109
+ if (!chip) return;
1110
+ chip.className = "lmw-tool-chip" + (outcome.error ? " error" : "");
1111
+ chip.textContent = (outcome.error ? "❌ " : "🔧 ") + outcome.toolCall.name;
1112
+ if (outcome.error) chip.title = outcome.error.message || String(outcome.error);
1113
+ }
1114
+
1073
1115
  var currentThinkingBlock = null;
1074
1116
  var currentThinkingBody = null;
1117
+ // The assistant bubble of the turn in flight. Reasoning belongs ABOVE that
1118
+ // bubble's content, the way llm_meta_chat places it — appending it to the
1119
+ // transcript instead left it below the answer it was reasoning towards.
1120
+ var activeAssistantBody = null;
1075
1121
 
1076
1122
  function ensureThinkingBlock() {
1077
1123
  if (currentThinkingBody) return currentThinkingBody;
@@ -1095,7 +1141,11 @@
1095
1141
  body.className = "message-thinking-content";
1096
1142
  details.appendChild(summary);
1097
1143
  details.appendChild(body);
1098
- historyEl.appendChild(details);
1144
+ if (activeAssistantBody && activeAssistantBody.parentNode) {
1145
+ activeAssistantBody.parentNode.insertBefore(details, activeAssistantBody);
1146
+ } else {
1147
+ historyEl.appendChild(details);
1148
+ }
1099
1149
  historyEl.scrollTop = historyEl.scrollHeight;
1100
1150
  currentThinkingBlock = details;
1101
1151
  currentThinkingBody = body;
@@ -1187,6 +1237,9 @@
1187
1237
  currentAbort = new AbortController();
1188
1238
 
1189
1239
  var assistantBody = appendTurn("assistant", "");
1240
+ activeAssistantBody = assistantBody;
1241
+ liveChips = null;
1242
+ pendingChips = [];
1190
1243
  markWorking(assistantBody.roleLabel);
1191
1244
  var assistantMarkdown = ""; // accumulate raw markdown, re-render on each delta
1192
1245
 
@@ -1213,6 +1266,8 @@
1213
1266
  aiActions: window[ACTIONS_GLOBAL] || {},
1214
1267
  maxRounds: MAX_ROUNDS,
1215
1268
  signal: currentAbort.signal,
1269
+ onToolCall: function(toolCall) { announceToolCall(toolCall); },
1270
+ onToolDispatched: function(outcome) { resolveToolCall(outcome); },
1216
1271
  onPhase: function(name) {
1217
1272
  // 'thinking' covers the long silence before the first
1218
1273
  // token; 'responding' means text is on its way.
@@ -1249,23 +1304,15 @@
1249
1304
  conversation.push({ role: "user", content: userText });
1250
1305
  conversation.push({ role: "assistant", content: result.content });
1251
1306
 
1252
- // Append a compact "tools used" footer INSIDE the assistant
1253
- // bubble so users have visual confirmation of what the LLM
1254
- // actually did — especially important when the LLM's own text
1255
- // is incomplete (e.g. "Let me first search…" with no follow-up
1256
- // synthesis, common with weaker tool-use models).
1257
- if (result.dispatched.length > 0) {
1258
- var chips = document.createElement("div");
1259
- chips.className = "lmw-tool-chips";
1260
- result.dispatched.forEach(function(d) {
1261
- var chip = document.createElement("span");
1262
- chip.className = "lmw-tool-chip" + (d.error ? " error" : "");
1263
- chip.textContent = (d.error ? "❌ " : "🔧 ") + d.toolCall.name;
1264
- if (d.error) chip.title = d.error.message;
1265
- chips.appendChild(chip);
1266
- });
1267
- assistantBody.appendChild(chips);
1268
- }
1307
+ // Anything still pending never reported an outcome (a class that
1308
+ // does not round-trip, or a dispatch that vanished). Reconcile
1309
+ // from the loop's own record so no chip is left spinning.
1310
+ result.dispatched.forEach(function(d) { resolveToolCall(d); });
1311
+ pendingChips.forEach(function(chip) {
1312
+ chip.className = "lmw-tool-chip";
1313
+ chip.textContent = "🔧 " + chip.dataset.tool;
1314
+ });
1315
+ pendingChips = [];
1269
1316
 
1270
1317
  // Skipped = LLM tried a tool that doesn't exist. Real signal.
1271
1318
  if (result.skipped.length > 0) {
@@ -1296,6 +1343,7 @@
1296
1343
  // no spinner: it says "still working" about a turn that ended.
1297
1344
  markDone(assistantBody.roleLabel);
1298
1345
  collapseThinkingBlock();
1346
+ activeAssistantBody = null;
1299
1347
  currentAbort = null;
1300
1348
  }
1301
1349
  });
@@ -1,3 +1,3 @@
1
1
  module LlmMetaWidget
2
- VERSION = "0.4.0"
2
+ VERSION = "0.4.2"
3
3
  end
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: llm_meta_widget
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.4.0
4
+ version: 0.4.2
5
5
  platform: ruby
6
6
  authors:
7
7
  - jdkim