llm_meta_widget 0.4.0 → 0.4.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/README.md +122 -12
- data/app/views/llm_meta_widget/_chat_panel.html.erb +11 -1
- data/lib/llm_meta_widget/version.rb +1 -1
- metadata +1 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 17cad71af08d5e4c49a3e7418055bd9c034880d9ab6cae2cbb3fb6cf3a0f0804
|
|
4
|
+
data.tar.gz: 24e7aa0c38400bccee9ddc832a6da10cf2788f42e343b67901d0cfc9a7748e24
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 2e55a245ad87d21f2d235fe9c211751520172df72f7197ad70b558a546c0b80e62ce63e9f00718ac0f91145d2a14bf920db808f70a550c7822d603c20a565833
|
|
7
|
+
data.tar.gz: bdad3eda95fd643db1dc11fc10c4e2c48d0d3442d3e2a1a79012f1e8bb2af769dc98217d897b6fd7b173a17f06e7ba55ab3209d50a3cd7186349dbce5920030b
|
data/README.md
CHANGED
|
@@ -6,28 +6,71 @@ Client-orchestrated: the widget fetches host-side action schemas + host-publishe
|
|
|
6
6
|
|
|
7
7
|
**No** Devise, DB migrations, ChatManager, or PromptNavigator. Adds only `rails >= 8.0` as a runtime dep — so hosts that haven't bumped to 8.1 can adopt it without a Rails upgrade.
|
|
8
8
|
|
|
9
|
-
##
|
|
9
|
+
## What you need first
|
|
10
10
|
|
|
11
|
-
|
|
11
|
+
The widget is a browser front end. It does not talk to OpenAI, Anthropic or
|
|
12
|
+
Ollama itself — it talks to **[llm_meta_server](https://github.com/pubannotation)**,
|
|
13
|
+
a hub that holds the provider credentials and streams responses back over SSE.
|
|
14
|
+
So adopting the widget means either running a hub or being given the URL of one.
|
|
15
|
+
|
|
16
|
+
Nothing else is required: no database, no migrations, no JavaScript build step,
|
|
17
|
+
no Node at runtime. The gem ships plain ES modules that the engine serves.
|
|
18
|
+
|
|
19
|
+
| Requirement | Version |
|
|
20
|
+
|---|---|
|
|
21
|
+
| Ruby | >= 3.2 |
|
|
22
|
+
| Rails | >= 8.0 (8.1 not required) |
|
|
23
|
+
| A reachable llm_meta_server | any |
|
|
24
|
+
|
|
25
|
+
## Installation
|
|
12
26
|
|
|
13
27
|
```ruby
|
|
14
|
-
|
|
28
|
+
# Gemfile
|
|
29
|
+
gem "llm_meta_widget", "~> 0.4"
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
```
|
|
33
|
+
bundle install
|
|
15
34
|
```
|
|
16
35
|
|
|
17
|
-
|
|
36
|
+
That is the whole install. There is **no `mount` to add** — the engine
|
|
37
|
+
deliberately skips `isolate_namespace`, so the helper is included into
|
|
38
|
+
ActionView and the asset routes appear at the host's top level on their own.
|
|
39
|
+
There are no migrations and no generators to run.
|
|
18
40
|
|
|
19
|
-
##
|
|
41
|
+
## Minimal working example
|
|
20
42
|
|
|
21
|
-
|
|
43
|
+
Put this on any view — a fresh `pages/demo.html.erb` is fine:
|
|
22
44
|
|
|
23
45
|
```erb
|
|
46
|
+
<h1>Demo</h1>
|
|
47
|
+
|
|
24
48
|
<%= llm_meta_widget(base_url: "https://your-meta-server.example",
|
|
25
|
-
model: "qwen3-6-35b-fast"
|
|
49
|
+
model: "qwen3-6-35b-fast",
|
|
50
|
+
greeting: "Hi — ask me anything about this page.") %>
|
|
26
51
|
```
|
|
27
52
|
|
|
28
|
-
|
|
53
|
+
Start the app, open that page, and click the round toggle at the bottom right.
|
|
54
|
+
**Success looks like:** a panel opens showing your greeting, you type "hello",
|
|
55
|
+
the role label turns into a spinning gear reading "Working…", and the reply
|
|
56
|
+
streams in.
|
|
57
|
+
|
|
58
|
+
If the panel opens but nothing comes back, it is almost always CORS. The hub
|
|
59
|
+
must list your app's origin in its `config/initializers/cors.rb` for the
|
|
60
|
+
`/api/*` resource, because the widget calls the hub **from the visitor's
|
|
61
|
+
browser**, not from your server. Open the browser console: a blocked request
|
|
62
|
+
says so explicitly.
|
|
63
|
+
|
|
64
|
+
Two more things worth knowing before you go further:
|
|
65
|
+
|
|
66
|
+
- `base_url` must be reachable from your visitors' browsers, not just from
|
|
67
|
+
your server. `localhost` works only while you are the visitor.
|
|
68
|
+
- The model name is the hub's name for it (`GET /api/llms` lists them), not
|
|
69
|
+
the provider's.
|
|
29
70
|
|
|
30
|
-
|
|
71
|
+
Once that works, the widget can converse but cannot *do* anything. To let the
|
|
72
|
+
LLM act on your page or call your own services, declare tools using any of the
|
|
73
|
+
**three tool classes** below.
|
|
31
74
|
|
|
32
75
|
## Three tool classes
|
|
33
76
|
|
|
@@ -72,7 +115,17 @@ Declared inline on the same view as the widget. Runs as JavaScript in the browse
|
|
|
72
115
|
|
|
73
116
|
**What the LLM sees.** The `ai-actions` JSON is passed to the meta-server as `local_tools` for every turn. Alongside it, every reader in `window.aiState` is invoked (each turn) and the results are JSON-serialized into a `Current page state:` block appended to the system prompt — the LLM can answer from state directly instead of tool-calling for lookups.
|
|
74
117
|
|
|
75
|
-
**When the tool_call fires.**
|
|
118
|
+
**When the tool_call fires.** During the turn, as soon as the LLM emits the
|
|
119
|
+
call. Since 0.4.0 the action's outcome — `{ok: true, applied: "<name>"}`, or
|
|
120
|
+
the error message if it raised — is fed back to the LLM as a tool result, and
|
|
121
|
+
the turn continues. Before 0.4.0 these were fire-and-forget, which meant a
|
|
122
|
+
turn whose only calls were page actions ended there: any task shaped *change
|
|
123
|
+
the page, then do something with it* was cut off after the write. The change
|
|
124
|
+
also means the LLM can see that an action failed and correct itself, instead
|
|
125
|
+
of the failure being visible only to the visitor.
|
|
126
|
+
|
|
127
|
+
Budget for it: every page action now costs a round, so a flow that writes twice
|
|
128
|
+
and then calls a tool needs roughly five or six. See `max_rounds:` below.
|
|
76
129
|
|
|
77
130
|
**Args**: parsed JS object matching the `input_schema`. Return value ignored. Async allowed (widget doesn't await, but browser will still execute the promise). Exceptions logged and surfaced as an `❌ <name>` chip in the message footer.
|
|
78
131
|
|
|
@@ -115,7 +168,7 @@ Declared out-of-band on the meta-server (via the hub's admin UI or `/user/:id/mc
|
|
|
115
168
|
|
|
116
169
|
**What the LLM sees.** Only tools from servers the visitor has **enabled via the tool picker** (see "Level-1 pickers" below). Nothing is auto-selected — the visitor opts in per session.
|
|
117
170
|
|
|
118
|
-
**When the tool_call fires.** Synchronously during the turn — widget POSTs to the hub's `/api/llm_api_keys/:uuid/models/:name/
|
|
171
|
+
**When the tool_call fires.** Synchronously during the turn — widget POSTs to the hub's `/api/llm_api_keys/:uuid/models/:name/single_llm_calls` endpoint with `tool_ids: [...]`; the hub proxies to each MCP server and streams results back through SSE.
|
|
119
172
|
|
|
120
173
|
**Error path.** Hub-side errors (rate limit, timeout, MCP server unavailable, upstream failure) arrive as SSE `event: error` frames with typed codes (`mcp_unavailable`, `timeout`, `rate_limit`, …); the widget surfaces them in the message bubble with a per-code prefix.
|
|
121
174
|
|
|
@@ -162,12 +215,69 @@ To lock the widget to the fixed `model:` prop and disable Class-1 tools entirely
|
|
|
162
215
|
| `actions_global:` | `"aiActions"` | Global window object holding Class-3 implementations |
|
|
163
216
|
| `remote_tools_schema_id:` | `"remote-mcp-tools"` | Optional DOM id for pre-configured Class-1 tools (bypasses picker) |
|
|
164
217
|
| `well_known_urls:` | `nil` | `nil` = auto-discover same-origin; explicit array = fetch those; `[]` = disable |
|
|
165
|
-
| `
|
|
218
|
+
| `greeting:` | `nil` | First thing a visitor sees when the panel opens, above the offered prompt templates. `nil` = a generic line |
|
|
219
|
+
| `max_rounds:` | `3` | Cap on tool-call rounds per LLM turn. Page actions cost a round each since 0.4.0 — raise it for multi-step flows |
|
|
166
220
|
| `enable_model_picker:` | `true` | Show model dropdown (Level-1) |
|
|
167
221
|
| `enable_tool_picker:` | `true` | Show tool picker (Level-1) |
|
|
168
222
|
| `models:` | `nil` | Model-name allowlist; `nil` = all anon-available |
|
|
169
223
|
| `hub_tools:` | `nil` | MCP-server-name allowlist; `nil` = all anon-public |
|
|
170
224
|
|
|
225
|
+
## Declaring resources and prompts (static-primitives extension)
|
|
226
|
+
|
|
227
|
+
If you operate the MCP endpoint behind your `.well-known/mcp.json`, you can
|
|
228
|
+
also offer **prompts** (one-click starting points, shown as buttons above the
|
|
229
|
+
input) and **resources** (reference data the widget can put in the model's
|
|
230
|
+
context). The widget reads two optional fields the
|
|
231
|
+
`io.modelcontextprotocol/static-primitives` extension defines on a
|
|
232
|
+
`resources/list` entry:
|
|
233
|
+
|
|
234
|
+
```json
|
|
235
|
+
{
|
|
236
|
+
"uri": "yourhost://catalog",
|
|
237
|
+
"name": "Catalog",
|
|
238
|
+
"mimeType": "application/json",
|
|
239
|
+
"_meta": {
|
|
240
|
+
"io.modelcontextprotocol/static-primitives": {
|
|
241
|
+
"sizeBytes": 1298,
|
|
242
|
+
"volatility": "stable",
|
|
243
|
+
"autoAttach": true
|
|
244
|
+
}
|
|
245
|
+
}
|
|
246
|
+
}
|
|
247
|
+
```
|
|
248
|
+
|
|
249
|
+
- **`sizeBytes`** — the byte length of the payload `resources/read` returns.
|
|
250
|
+
The widget decides whether it can afford the resource *before* fetching it;
|
|
251
|
+
over its budget (8 KB), the bytes never cross the wire and the LLM is left
|
|
252
|
+
to ask for what it needs through your tools instead.
|
|
253
|
+
- **`volatility`** — `"stable"` (read once, re-used) or `"volatile"` (re-read
|
|
254
|
+
on every turn that uses it).
|
|
255
|
+
- **`autoAttach`** — may the widget put this in the model's context without
|
|
256
|
+
the visitor asking?
|
|
257
|
+
|
|
258
|
+
All three are optional. A server that declares none behaves exactly as one
|
|
259
|
+
that predates the extension: the resource is fetched, trimmed to budget and
|
|
260
|
+
carried in the model's context.
|
|
261
|
+
|
|
262
|
+
A prompt is offered as a button labelled by its `title`. Declare its arguments
|
|
263
|
+
as **optional** and fill them from page state where you can: an argument the
|
|
264
|
+
visitor must supply by hand makes the button useless to the visitor who needed
|
|
265
|
+
it most, since they must already know your form to press it.
|
|
266
|
+
|
|
267
|
+
## Further reading
|
|
268
|
+
|
|
269
|
+
- **The static-primitives extension** — `io.modelcontextprotocol/static-primitives`,
|
|
270
|
+
a follow-on to [SEP-2127](https://github.com/modelcontextprotocol/modelcontextprotocol)
|
|
271
|
+
(Server Cards). The fields above are a reference implementation of its
|
|
272
|
+
current draft shape.
|
|
273
|
+
- **MCP itself** — <https://modelcontextprotocol.io>
|
|
274
|
+
- **A worked adopter** — PubDictionaries' annotation page
|
|
275
|
+
(<https://pubdictionaries.org/text_annotation>) runs this widget with all
|
|
276
|
+
three tool classes plus prompts and resources. Its
|
|
277
|
+
`app/views/annotation/text_annotation.html.erb` and `app/controllers/mcp_controller.rb`
|
|
278
|
+
are the fullest example available.
|
|
279
|
+
- **Issues and questions** — <https://github.com/jdkim/llm_meta_widget/issues>
|
|
280
|
+
|
|
171
281
|
## License
|
|
172
282
|
|
|
173
283
|
Apache-2.0.
|
|
@@ -1072,6 +1072,10 @@
|
|
|
1072
1072
|
|
|
1073
1073
|
var currentThinkingBlock = null;
|
|
1074
1074
|
var currentThinkingBody = null;
|
|
1075
|
+
// The assistant bubble of the turn in flight. Reasoning belongs ABOVE that
|
|
1076
|
+
// bubble's content, the way llm_meta_chat places it — appending it to the
|
|
1077
|
+
// transcript instead left it below the answer it was reasoning towards.
|
|
1078
|
+
var activeAssistantBody = null;
|
|
1075
1079
|
|
|
1076
1080
|
function ensureThinkingBlock() {
|
|
1077
1081
|
if (currentThinkingBody) return currentThinkingBody;
|
|
@@ -1095,7 +1099,11 @@
|
|
|
1095
1099
|
body.className = "message-thinking-content";
|
|
1096
1100
|
details.appendChild(summary);
|
|
1097
1101
|
details.appendChild(body);
|
|
1098
|
-
|
|
1102
|
+
if (activeAssistantBody && activeAssistantBody.parentNode) {
|
|
1103
|
+
activeAssistantBody.parentNode.insertBefore(details, activeAssistantBody);
|
|
1104
|
+
} else {
|
|
1105
|
+
historyEl.appendChild(details);
|
|
1106
|
+
}
|
|
1099
1107
|
historyEl.scrollTop = historyEl.scrollHeight;
|
|
1100
1108
|
currentThinkingBlock = details;
|
|
1101
1109
|
currentThinkingBody = body;
|
|
@@ -1187,6 +1195,7 @@
|
|
|
1187
1195
|
currentAbort = new AbortController();
|
|
1188
1196
|
|
|
1189
1197
|
var assistantBody = appendTurn("assistant", "");
|
|
1198
|
+
activeAssistantBody = assistantBody;
|
|
1190
1199
|
markWorking(assistantBody.roleLabel);
|
|
1191
1200
|
var assistantMarkdown = ""; // accumulate raw markdown, re-render on each delta
|
|
1192
1201
|
|
|
@@ -1296,6 +1305,7 @@
|
|
|
1296
1305
|
// no spinner: it says "still working" about a turn that ended.
|
|
1297
1306
|
markDone(assistantBody.roleLabel);
|
|
1298
1307
|
collapseThinkingBlock();
|
|
1308
|
+
activeAssistantBody = null;
|
|
1299
1309
|
currentAbort = null;
|
|
1300
1310
|
}
|
|
1301
1311
|
});
|