llm_meta_widget 0.4.2 → 0.6.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: f019c93684d03e4f65ead3c8427b369af41caf2b9f4b0637f60aa2014239f080
|
|
4
|
+
data.tar.gz: 28bf06bc291b4bc0f745b85ae7f067f9e9c3f8adfd4902d747a4496d1b2e34da
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 50a79ea50cd118b4cafb1a26ff5af6713fd6b98dd1374ad4a2b2fa10fff24677fe417cee3b4cf4d900e121b3b4be3c013fcbc1aa95d4183f714e629a100432c8
|
|
7
|
+
data.tar.gz: 1bbed09f23f1c5c7f9b08500db77d2dd1b7ba97e1b564405b40f8cadd9fb1bac920a4d612e187119d08cd90db51301b96dc8d3d3ad22e0196adc04b5ee1621c7
|
data/README.md
CHANGED
|
@@ -8,19 +8,25 @@ Client-orchestrated: the widget fetches host-side action schemas + host-publishe
|
|
|
8
8
|
|
|
9
9
|
## What you need first
|
|
10
10
|
|
|
11
|
-
The widget is a browser front end
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
11
|
+
The widget is a browser front end: it needs something to answer the chat. That
|
|
12
|
+
can be **[llm_meta_server](https://github.com/pubannotation)** — a hub that
|
|
13
|
+
holds provider credentials and streams responses back — or **an Ollama**,
|
|
14
|
+
which the widget talks to directly with no server component of ours in
|
|
15
|
+
between.
|
|
15
16
|
|
|
16
|
-
|
|
17
|
-
|
|
17
|
+
Those are separate from where tools come from. A hub also registers MCP tools
|
|
18
|
+
(Class 1 below), and you can point at one for tools while a local Ollama
|
|
19
|
+
answers the chat, or use no hub at all. Page actions and your own
|
|
20
|
+
`.well-known/mcp.json` never involve a hub either way.
|
|
21
|
+
|
|
22
|
+
Nothing else is required: no database, no migrations, no JavaScript build
|
|
23
|
+
step, no Node at runtime.
|
|
18
24
|
|
|
19
25
|
| Requirement | Version |
|
|
20
26
|
|---|---|
|
|
21
27
|
| Ruby | >= 3.2 |
|
|
22
28
|
| Rails | >= 8.0 (8.1 not required) |
|
|
23
|
-
|
|
|
29
|
+
| An llm_meta_server **or** an Ollama | any |
|
|
24
30
|
|
|
25
31
|
## Installation
|
|
26
32
|
|
|
@@ -45,7 +51,7 @@ Put this on any view — a fresh `pages/demo.html.erb` is fine:
|
|
|
45
51
|
```erb
|
|
46
52
|
<h1>Demo</h1>
|
|
47
53
|
|
|
48
|
-
<%= llm_meta_widget(
|
|
54
|
+
<%= llm_meta_widget(llm_url: "https://your-meta-server.example",
|
|
49
55
|
model: "qwen3-6-35b-fast",
|
|
50
56
|
greeting: "Hi — ask me anything about this page.") %>
|
|
51
57
|
```
|
|
@@ -63,8 +69,9 @@ says so explicitly.
|
|
|
63
69
|
|
|
64
70
|
Two more things worth knowing before you go further:
|
|
65
71
|
|
|
66
|
-
- `
|
|
67
|
-
your server. `localhost` works only while you are the visitor
|
|
72
|
+
- `llm_url` must be reachable from your visitors' browsers, not just from
|
|
73
|
+
your server. `localhost` works only while you are the visitor — which is
|
|
74
|
+
fine when each visitor runs their own Ollama.
|
|
68
75
|
- The model name is the hub's name for it (`GET /api/llms` lists them), not
|
|
69
76
|
the provider's.
|
|
70
77
|
|
|
@@ -185,7 +192,8 @@ To adjust picker behavior at the helper call site:
|
|
|
185
192
|
|
|
186
193
|
```erb
|
|
187
194
|
<%= llm_meta_widget(
|
|
188
|
-
|
|
195
|
+
llm_url: "https://your-meta-server.example",
|
|
196
|
+
tool_hub_url: "https://your-meta-server.example",
|
|
189
197
|
model: "qwen3-6-35b-fast", # initial selection
|
|
190
198
|
enable_model_picker: true, # false → hide picker, use fixed `model:`
|
|
191
199
|
enable_tool_picker: true, # false → hide picker, no Class-1 tools
|
|
@@ -197,7 +205,7 @@ To adjust picker behavior at the helper call site:
|
|
|
197
205
|
To lock the widget to the fixed `model:` prop and disable Class-1 tools entirely (level-0 mode — Class 2 & 3 still work):
|
|
198
206
|
|
|
199
207
|
```erb
|
|
200
|
-
<%= llm_meta_widget(
|
|
208
|
+
<%= llm_meta_widget(llm_url: "…", model: "…",
|
|
201
209
|
enable_model_picker: false,
|
|
202
210
|
enable_tool_picker: false) %>
|
|
203
211
|
```
|
|
@@ -206,7 +214,9 @@ To lock the widget to the fixed `model:` prop and disable Class-1 tools entirely
|
|
|
206
214
|
|
|
207
215
|
| Option | Default | Purpose |
|
|
208
216
|
|---|---|---|
|
|
209
|
-
| `
|
|
217
|
+
| `llm_url:` | required | Who answers the chat — an llm_meta_server, or an Ollama |
|
|
218
|
+
| `llm_provider:` | `:llm_meta_server` | What `llm_url` speaks. `:ollama` talks to Ollama's `/api/chat` directly |
|
|
219
|
+
| `tool_hub_url:` | `nil` | An llm_meta_server whose registered MCP tools to offer. Absent = none (page actions and your own `.well-known` MCP are unaffected) |
|
|
210
220
|
| `model:` | required | Initial model (also the fallback when picker is disabled) |
|
|
211
221
|
| `api_key_uuid:` | `"ollama-local"` | Hub API-key uuid to invoke |
|
|
212
222
|
| `orchestrator_path:` | `"/llm_meta_widget_assets/orchestrator.js"` | Served by the gem's engine; rarely overridden |
|
|
@@ -222,6 +232,62 @@ To lock the widget to the fixed `model:` prop and disable Class-1 tools entirely
|
|
|
222
232
|
| `models:` | `nil` | Model-name allowlist; `nil` = all anon-available |
|
|
223
233
|
| `hub_tools:` | `nil` | MCP-server-name allowlist; `nil` = all anon-public |
|
|
224
234
|
|
|
235
|
+
## Surviving navigation
|
|
236
|
+
|
|
237
|
+
A page action that navigates — submitting a form, following a link — no longer
|
|
238
|
+
loses the conversation. The transcript, the record of which tools ran, and the
|
|
239
|
+
model's own context are kept in `sessionStorage`: per tab, gone when the tab
|
|
240
|
+
closes, never sent anywhere.
|
|
241
|
+
|
|
242
|
+
The key is the page's **path**, not its full URL, so a form submit that returns
|
|
243
|
+
to the same page with a new query string keeps the thread, while moving to a
|
|
244
|
+
different page starts a fresh one. Nothing to configure.
|
|
245
|
+
|
|
246
|
+
Two consequences worth knowing. An action of yours that navigates is now safe
|
|
247
|
+
to declare — before this, offering one meant offering to wipe the visitor's
|
|
248
|
+
chat. And if the navigation lands on an error page that does not render the
|
|
249
|
+
widget, the conversation is still in storage but there is no panel to show it
|
|
250
|
+
until the visitor returns to a page that has one.
|
|
251
|
+
|
|
252
|
+
## Choosing what answers the chat, and whose tools to offer
|
|
253
|
+
|
|
254
|
+
Two independent settings, because they are two jobs. Neither implies the
|
|
255
|
+
other, so there is nothing to disable and nothing to inherit — what a widget
|
|
256
|
+
does is what its call site says.
|
|
257
|
+
|
|
258
|
+
```erb
|
|
259
|
+
<%# a hub answers the chat; its registered tools are offered too %>
|
|
260
|
+
<% hub = "https://your-meta-server.example" %>
|
|
261
|
+
<%= llm_meta_widget(llm_url: hub, tool_hub_url: hub, model: "qwen3-6-35b-fast") %>
|
|
262
|
+
|
|
263
|
+
<%# a hub answers the chat; no hub-registered tools %>
|
|
264
|
+
<%= llm_meta_widget(llm_url: hub, model: "qwen3-6-35b-fast") %>
|
|
265
|
+
|
|
266
|
+
<%# the visitor's own Ollama answers; a hub still supplies its tools %>
|
|
267
|
+
<%= llm_meta_widget(llm_url: "http://localhost:11434", llm_provider: :ollama,
|
|
268
|
+
tool_hub_url: hub, model: "qwen3.8:27b") %>
|
|
269
|
+
|
|
270
|
+
<%# nothing of ours in the path at all %>
|
|
271
|
+
<%= llm_meta_widget(llm_url: "http://localhost:11434", llm_provider: :ollama,
|
|
272
|
+
model: "qwen3.8:27b") %>
|
|
273
|
+
```
|
|
274
|
+
|
|
275
|
+
**The Ollama path.** The widget POSTs to `{llm_url}/api/chat` and reads
|
|
276
|
+
Ollama's NDJSON stream; tool schemas go over as OpenAI-shaped functions and
|
|
277
|
+
`tool_calls` come back, so Class 2 and Class 3 tools work exactly as they do
|
|
278
|
+
against a hub. The model picker lists `{llm_url}/api/tags`. Ollama must be
|
|
279
|
+
told to accept your page's origin — `OLLAMA_ORIGINS=https://your-site.example`
|
|
280
|
+
— which is the same CORS story as the hub, configured elsewhere.
|
|
281
|
+
|
|
282
|
+
What you give up without a `tool_hub_url` is Class 1 only: tools registered on
|
|
283
|
+
somebody's hub. That is not "no tools" — page actions and your own
|
|
284
|
+
`.well-known/mcp.json` are untouched, and they are the interesting ones for an
|
|
285
|
+
assistant embedded in your page.
|
|
286
|
+
|
|
287
|
+
**Why anyone would want the last shape:** with a local Ollama and local MCP
|
|
288
|
+
endpoints, nothing a visitor types leaves the machine. No credentials to hold,
|
|
289
|
+
no retention policy to write.
|
|
290
|
+
|
|
225
291
|
## Declaring resources and prompts (static-primitives extension)
|
|
226
292
|
|
|
227
293
|
If you operate the MCP endpoint behind your `.well-known/mcp.json`, you can
|
|
@@ -325,6 +325,14 @@ export async function runChatLoop(opts) {
|
|
|
325
325
|
remoteTools = [],
|
|
326
326
|
hostWideTools = [],
|
|
327
327
|
maxRounds = 10,
|
|
328
|
+
// "hub" (default) talks to llm_meta_server; "ollama" talks straight to an
|
|
329
|
+
// Ollama, with no server component of ours in between. The loop is the
|
|
330
|
+
// same either way — only the call is swapped.
|
|
331
|
+
provider = "hub",
|
|
332
|
+
// Where Class 1 (hub-registered) tools are proxied. Usually the same
|
|
333
|
+
// llm_meta_server that answers the chat, but not when the chat goes
|
|
334
|
+
// straight to an Ollama — the tools still belong to the hub.
|
|
335
|
+
toolHubUrl,
|
|
328
336
|
signal,
|
|
329
337
|
onRoundStart,
|
|
330
338
|
onTextDelta,
|
|
@@ -399,7 +407,9 @@ export async function runChatLoop(opts) {
|
|
|
399
407
|
if (signal?.aborted) throw new DOMException("aborted", "AbortError")
|
|
400
408
|
onRoundStart?.(round)
|
|
401
409
|
|
|
402
|
-
|
|
410
|
+
// The loop is the same whoever answers; only the call differs.
|
|
411
|
+
const call = provider === "ollama" ? ollamaChatCall : singleLlmCall
|
|
412
|
+
const turnResult = await call({
|
|
403
413
|
...singleOpts,
|
|
404
414
|
messages,
|
|
405
415
|
toolIds,
|
|
@@ -490,7 +500,7 @@ export async function runChatLoop(opts) {
|
|
|
490
500
|
const args = coerceArguments(tc.arguments)
|
|
491
501
|
try {
|
|
492
502
|
const value = await dispatchRemoteToolCall({
|
|
493
|
-
baseUrl: singleOpts.baseUrl,
|
|
503
|
+
baseUrl: toolHubUrl || singleOpts.baseUrl,
|
|
494
504
|
bearerToken: singleOpts.bearerToken,
|
|
495
505
|
toolId: tool.id,
|
|
496
506
|
args: args,
|
|
@@ -1006,3 +1016,243 @@ export function promptButtonProps(prompt) {
|
|
|
1006
1016
|
}
|
|
1007
1017
|
}
|
|
1008
1018
|
|
|
1019
|
+
|
|
1020
|
+
// ---- direct-to-Ollama provider ------------------------------------------
|
|
1021
|
+
//
|
|
1022
|
+
// The smallest adoption: a Rails app and an Ollama, with no hub of ours in
|
|
1023
|
+
// between. Ollama's /api/chat already does the hard part — it takes tool
|
|
1024
|
+
// schemas and returns tool_calls — so what differs from the hub path is the
|
|
1025
|
+
// envelope: NDJSON rather than SSE, and its own request shape.
|
|
1026
|
+
//
|
|
1027
|
+
// What an adopter gives up is Class 1 (hub-registered MCP), which is the
|
|
1028
|
+
// hub's job by definition. Page actions and the host's own .well-known MCP
|
|
1029
|
+
// both work unchanged, and nothing the visitor types leaves their machine.
|
|
1030
|
+
//
|
|
1031
|
+
// Ollama must be told to accept the page's origin (OLLAMA_ORIGINS), the same
|
|
1032
|
+
// CORS story as the hub, configured elsewhere.
|
|
1033
|
+
|
|
1034
|
+
export async function* parseNdjsonStream(readableStream, signal) {
|
|
1035
|
+
const reader = readableStream.getReader()
|
|
1036
|
+
const decoder = new TextDecoder("utf-8")
|
|
1037
|
+
let buffer = ""
|
|
1038
|
+
|
|
1039
|
+
const onAbort = () => { try { reader.cancel() } catch { /* noop */ } }
|
|
1040
|
+
signal?.addEventListener("abort", onAbort)
|
|
1041
|
+
|
|
1042
|
+
try {
|
|
1043
|
+
while (true) {
|
|
1044
|
+
const { value, done } = await reader.read()
|
|
1045
|
+
if (done) break
|
|
1046
|
+
buffer += decoder.decode(value, { stream: true })
|
|
1047
|
+
|
|
1048
|
+
let nl
|
|
1049
|
+
while ((nl = buffer.indexOf("\n")) !== -1) {
|
|
1050
|
+
const line = buffer.slice(0, nl).trim()
|
|
1051
|
+
buffer = buffer.slice(nl + 1)
|
|
1052
|
+
if (!line) continue
|
|
1053
|
+
try {
|
|
1054
|
+
yield JSON.parse(line)
|
|
1055
|
+
} catch (e) {
|
|
1056
|
+
// A truncated or non-JSON line is not worth aborting a stream for.
|
|
1057
|
+
}
|
|
1058
|
+
}
|
|
1059
|
+
}
|
|
1060
|
+
const tail = buffer.trim()
|
|
1061
|
+
if (tail) { try { yield JSON.parse(tail) } catch (e) { /* ignore */ } }
|
|
1062
|
+
} finally {
|
|
1063
|
+
signal?.removeEventListener("abort", onAbort)
|
|
1064
|
+
try { reader.releaseLock() } catch { /* noop */ }
|
|
1065
|
+
}
|
|
1066
|
+
}
|
|
1067
|
+
|
|
1068
|
+
// Tool schemas travel as MCP-flavoured `input_schema`; Ollama wants OpenAI's
|
|
1069
|
+
// function shape.
|
|
1070
|
+
function toolsForOllama(localTools) {
|
|
1071
|
+
return (localTools || []).map((t) => ({
|
|
1072
|
+
type: "function",
|
|
1073
|
+
function: {
|
|
1074
|
+
name: t.name,
|
|
1075
|
+
description: t.description,
|
|
1076
|
+
parameters: t.input_schema || t.inputSchema || { type: "object", properties: {} }
|
|
1077
|
+
}
|
|
1078
|
+
}))
|
|
1079
|
+
}
|
|
1080
|
+
|
|
1081
|
+
// The hub's history carries `tool_call_id`; Ollama identifies a result by the
|
|
1082
|
+
// tool's name instead, and rejects unknown keys on some versions.
|
|
1083
|
+
function messagesForOllama(messages) {
|
|
1084
|
+
return (messages || []).map((m) => {
|
|
1085
|
+
if (m.role === "tool") {
|
|
1086
|
+
return { role: "tool", tool_name: m.name, content: m.content }
|
|
1087
|
+
}
|
|
1088
|
+
if (m.tool_calls) {
|
|
1089
|
+
return {
|
|
1090
|
+
role: m.role,
|
|
1091
|
+
content: m.content || "",
|
|
1092
|
+
tool_calls: m.tool_calls.map((tc) => ({
|
|
1093
|
+
function: { name: tc.name, arguments: coerceArguments(tc.arguments) }
|
|
1094
|
+
}))
|
|
1095
|
+
}
|
|
1096
|
+
}
|
|
1097
|
+
return { role: m.role, content: m.content }
|
|
1098
|
+
})
|
|
1099
|
+
}
|
|
1100
|
+
|
|
1101
|
+
// Same signature and return shape as singleLlmCall, so runChatLoop does not
|
|
1102
|
+
// care which provider answered.
|
|
1103
|
+
export async function ollamaChatCall({
|
|
1104
|
+
baseUrl,
|
|
1105
|
+
modelName,
|
|
1106
|
+
messages,
|
|
1107
|
+
localTools = [],
|
|
1108
|
+
generationSettings = {},
|
|
1109
|
+
onTextDelta,
|
|
1110
|
+
onThinkingDelta,
|
|
1111
|
+
onToolCall,
|
|
1112
|
+
onPhase,
|
|
1113
|
+
signal,
|
|
1114
|
+
}) {
|
|
1115
|
+
if (!baseUrl || !modelName) throw new Error("ollamaChatCall: baseUrl and modelName are required")
|
|
1116
|
+
if (!Array.isArray(messages) || messages.length === 0) {
|
|
1117
|
+
throw new Error("ollamaChatCall: messages must be a non-empty array")
|
|
1118
|
+
}
|
|
1119
|
+
|
|
1120
|
+
const { think, ...options } = generationSettings || {}
|
|
1121
|
+
const body = { model: modelName, messages: messagesForOllama(messages), stream: true }
|
|
1122
|
+
if (localTools && localTools.length) body.tools = toolsForOllama(localTools)
|
|
1123
|
+
if (think !== undefined) body.think = think
|
|
1124
|
+
if (Object.keys(options).length) body.options = options
|
|
1125
|
+
|
|
1126
|
+
const response = await fetch(`${baseUrl.replace(/\/$/, "")}/api/chat`, {
|
|
1127
|
+
method: "POST",
|
|
1128
|
+
headers: { "Content-Type": "application/json" },
|
|
1129
|
+
body: JSON.stringify(body),
|
|
1130
|
+
signal
|
|
1131
|
+
})
|
|
1132
|
+
if (!response.ok) {
|
|
1133
|
+
const text = await response.text().catch(() => "")
|
|
1134
|
+
throw new Error(`ollamaChatCall: HTTP ${response.status} ${response.statusText}${text ? " — " + text.slice(0, 200) : ""}`)
|
|
1135
|
+
}
|
|
1136
|
+
if (!response.body) throw new Error("ollamaChatCall: response has no body (streaming unsupported?)")
|
|
1137
|
+
|
|
1138
|
+
const toolCalls = []
|
|
1139
|
+
let content = ""
|
|
1140
|
+
let finishReason = null
|
|
1141
|
+
let announcedPhase = null
|
|
1142
|
+
|
|
1143
|
+
const phase = (name) => {
|
|
1144
|
+
if (announcedPhase === name) return
|
|
1145
|
+
announcedPhase = name
|
|
1146
|
+
onPhase?.(name)
|
|
1147
|
+
}
|
|
1148
|
+
|
|
1149
|
+
for await (const frame of parseNdjsonStream(response.body, signal)) {
|
|
1150
|
+
const message = frame.message || {}
|
|
1151
|
+
|
|
1152
|
+
if (message.thinking) {
|
|
1153
|
+
phase("thinking")
|
|
1154
|
+
onThinkingDelta?.(message.thinking)
|
|
1155
|
+
}
|
|
1156
|
+
if (message.content) {
|
|
1157
|
+
phase("responding")
|
|
1158
|
+
content += message.content
|
|
1159
|
+
onTextDelta?.(message.content)
|
|
1160
|
+
}
|
|
1161
|
+
for (const tc of message.tool_calls || []) {
|
|
1162
|
+
const call = {
|
|
1163
|
+
id: tc.id || `ollama-${toolCalls.length}`,
|
|
1164
|
+
name: tc.function?.name,
|
|
1165
|
+
arguments: tc.function?.arguments ?? {}
|
|
1166
|
+
}
|
|
1167
|
+
toolCalls.push(call)
|
|
1168
|
+
onToolCall?.(call)
|
|
1169
|
+
}
|
|
1170
|
+
if (frame.done) finishReason = frame.done_reason || "stop"
|
|
1171
|
+
}
|
|
1172
|
+
|
|
1173
|
+
return { content, toolCalls, finishReason }
|
|
1174
|
+
}
|
|
1175
|
+
|
|
1176
|
+
// Ollama's own catalogue, for the model picker when there is no hub to ask.
|
|
1177
|
+
export async function fetchOllamaModels({ baseUrl, signal }) {
|
|
1178
|
+
try {
|
|
1179
|
+
const response = await fetch(`${baseUrl.replace(/\/$/, "")}/api/tags`, { signal })
|
|
1180
|
+
if (!response.ok) return []
|
|
1181
|
+
const data = await response.json()
|
|
1182
|
+
return (data.models || []).map((m) => m.name).filter(Boolean)
|
|
1183
|
+
} catch (e) {
|
|
1184
|
+
// A picker that cannot be populated is not a reason to break the widget.
|
|
1185
|
+
return []
|
|
1186
|
+
}
|
|
1187
|
+
}
|
|
1188
|
+
|
|
1189
|
+
// ---- surviving navigation ------------------------------------------------
|
|
1190
|
+
//
|
|
1191
|
+
// A page action that navigates — submitting a form, following a link — used to
|
|
1192
|
+
// destroy the conversation, the record of what ran and the scroll position.
|
|
1193
|
+
// That made such actions unsafe to offer at all: PubDictionaries had to stop
|
|
1194
|
+
// declaring its submit action rather than let the assistant wipe the chat.
|
|
1195
|
+
//
|
|
1196
|
+
// The transcript is kept in sessionStorage: per tab, gone when the tab closes,
|
|
1197
|
+
// invisible to other visitors. Keyed by pathname rather than full URL, because
|
|
1198
|
+
// submitting a form usually returns to the same page with different query
|
|
1199
|
+
// parameters and that is exactly the case worth surviving.
|
|
1200
|
+
|
|
1201
|
+
export const CONVERSATION_FORMAT = 1
|
|
1202
|
+
|
|
1203
|
+
export function packConversation({ turns = [], open = false, model = null,
|
|
1204
|
+
maxTurns = 40, maxBytes = 200000 } = {}) {
|
|
1205
|
+
// Oldest turns go first when trimming: the recent ones carry the thread.
|
|
1206
|
+
let kept = turns.slice(-maxTurns)
|
|
1207
|
+
let payload = { v: CONVERSATION_FORMAT, savedAt: Date.now(), open, model, turns: kept }
|
|
1208
|
+
let json = JSON.stringify(payload)
|
|
1209
|
+
|
|
1210
|
+
while (json.length > maxBytes && kept.length > 1) {
|
|
1211
|
+
kept = kept.slice(1)
|
|
1212
|
+
payload = { ...payload, turns: kept }
|
|
1213
|
+
json = JSON.stringify(payload)
|
|
1214
|
+
}
|
|
1215
|
+
return json
|
|
1216
|
+
}
|
|
1217
|
+
|
|
1218
|
+
export function unpackConversation(raw) {
|
|
1219
|
+
if (!raw) return null
|
|
1220
|
+
let payload
|
|
1221
|
+
try {
|
|
1222
|
+
payload = JSON.parse(raw)
|
|
1223
|
+
} catch (e) {
|
|
1224
|
+
return null
|
|
1225
|
+
}
|
|
1226
|
+
// A payload from a different format is not worth guessing at; starting
|
|
1227
|
+
// fresh is better than rendering something half-understood.
|
|
1228
|
+
if (!payload || payload.v !== CONVERSATION_FORMAT || !Array.isArray(payload.turns)) return null
|
|
1229
|
+
|
|
1230
|
+
const turns = payload.turns.filter((t) => t && typeof t.content === "string" &&
|
|
1231
|
+
(t.role === "user" || t.role === "assistant"))
|
|
1232
|
+
return { turns, open: !!payload.open, model: payload.model || null, savedAt: payload.savedAt || 0 }
|
|
1233
|
+
}
|
|
1234
|
+
|
|
1235
|
+
// Storage can throw — private windows, blocked site data, quota — and none of
|
|
1236
|
+
// that is a reason for the widget to stop working.
|
|
1237
|
+
export function createConversationStore({ storage, key }) {
|
|
1238
|
+
const safely = (fn, fallback = null) => {
|
|
1239
|
+
try {
|
|
1240
|
+
return fn()
|
|
1241
|
+
} catch (e) {
|
|
1242
|
+
return fallback
|
|
1243
|
+
}
|
|
1244
|
+
}
|
|
1245
|
+
|
|
1246
|
+
return {
|
|
1247
|
+
key,
|
|
1248
|
+
save(state) { return safely(() => { storage.setItem(key, packConversation(state)); return true }, false) },
|
|
1249
|
+
load() { return safely(() => unpackConversation(storage.getItem(key))) },
|
|
1250
|
+
clear() { return safely(() => { storage.removeItem(key); return true }, false) }
|
|
1251
|
+
}
|
|
1252
|
+
}
|
|
1253
|
+
|
|
1254
|
+
export function conversationKeyFor(location, prefix = "llm_meta_widget") {
|
|
1255
|
+
// Path only: a form submit returns to the same page with a new query string,
|
|
1256
|
+
// and that conversation is still the same conversation.
|
|
1257
|
+
return `${prefix}:${location.pathname}`
|
|
1258
|
+
}
|
|
@@ -16,7 +16,18 @@ module LlmMetaWidget
|
|
|
16
16
|
# (page-embedded aiActions / host-wide well-known / hub-registered).
|
|
17
17
|
module WidgetHelper
|
|
18
18
|
DEFAULTS = {
|
|
19
|
-
api_key_uuid: "ollama-local",
|
|
19
|
+
api_key_uuid: "ollama-local", # llm_meta_server provider only
|
|
20
|
+
# Answering the chat and registering tools are separate jobs. Name each
|
|
21
|
+
# endpoint for the job it does; neither implies the other.
|
|
22
|
+
#
|
|
23
|
+
# llm_url: who answers the chat (required)
|
|
24
|
+
# llm_provider: what it speaks — :llm_meta_server (default) or :ollama
|
|
25
|
+
# tool_hub_url: an llm_meta_server whose registered MCP tools to
|
|
26
|
+
# offer. Absent means none — which does NOT mean "no
|
|
27
|
+
# tools": page actions and the host's own
|
|
28
|
+
# .well-known/mcp.json are unaffected.
|
|
29
|
+
llm_provider: "llm_meta_server",
|
|
30
|
+
tool_hub_url: nil,
|
|
20
31
|
orchestrator_path: "/llm_meta_widget_assets/orchestrator.js",
|
|
21
32
|
actions_schema_id: "ai-actions",
|
|
22
33
|
state_global: "aiState",
|
|
@@ -42,9 +53,20 @@ module LlmMetaWidget
|
|
|
42
53
|
greeting: nil
|
|
43
54
|
}.freeze
|
|
44
55
|
|
|
45
|
-
def llm_meta_widget(
|
|
46
|
-
locals = DEFAULTS.merge(
|
|
47
|
-
|
|
56
|
+
def llm_meta_widget(llm_url:, model:, **overrides)
|
|
57
|
+
locals = DEFAULTS.merge(llm_url: llm_url, model: model, **overrides)
|
|
58
|
+
provider = locals[:llm_provider].to_s
|
|
59
|
+
unless %w[llm_meta_server ollama].include?(provider)
|
|
60
|
+
raise ArgumentError,
|
|
61
|
+
"llm_meta_widget: llm_provider must be :llm_meta_server or :ollama, got #{provider.inspect}"
|
|
62
|
+
end
|
|
63
|
+
raise ArgumentError, "llm_meta_widget: llm_url is required" if llm_url.to_s.strip.empty?
|
|
64
|
+
|
|
65
|
+
hub = locals[:tool_hub_url]
|
|
66
|
+
hub = nil if hub.to_s.strip.empty?
|
|
67
|
+
|
|
68
|
+
render partial: "llm_meta_widget/chat_panel",
|
|
69
|
+
locals: locals.merge(llm_provider: provider, tool_hub_url: hub)
|
|
48
70
|
end
|
|
49
71
|
end
|
|
50
72
|
end
|
|
@@ -452,7 +452,8 @@
|
|
|
452
452
|
import { runChatLoop, fetchMcpManifest, listMcpPrompts, getMcpPrompt,
|
|
453
453
|
listMcpResources, readMcpResource, promptMessagesToText,
|
|
454
454
|
loadHostResource, resourceLinesForTurn, resolvePromptArguments,
|
|
455
|
-
promptButtonProps, promptArgumentSummary
|
|
455
|
+
promptButtonProps, promptArgumentSummary, fetchOllamaModels,
|
|
456
|
+
createConversationStore, conversationKeyFor } from "<%= orchestrator_path %>";
|
|
456
457
|
import { marked } from "/llm_meta_widget_assets/marked.esm.js";
|
|
457
458
|
|
|
458
459
|
// Standard prose settings — GFM (tables, autolinks, strikethrough),
|
|
@@ -460,7 +461,12 @@
|
|
|
460
461
|
// bare newlines still reads naturally.
|
|
461
462
|
marked.setOptions({ gfm: true, breaks: true });
|
|
462
463
|
|
|
463
|
-
|
|
464
|
+
// Two endpoints, two jobs: one answers the chat, the other registers the
|
|
465
|
+
// MCP tools on offer. They are usually the same llm_meta_server, but need
|
|
466
|
+
// not be — an adopter can run their own Ollama and still borrow a hub's
|
|
467
|
+
// tools, or use no hub at all.
|
|
468
|
+
var LLM_BASE = <%= llm_url.to_json.html_safe %>;
|
|
469
|
+
var TOOL_HUB_BASE = <%= tool_hub_url.to_json.html_safe %>;
|
|
464
470
|
var API_KEY_UUID = <%= api_key_uuid.to_json.html_safe %>;
|
|
465
471
|
var MODEL = <%= model.to_json.html_safe %>;
|
|
466
472
|
var ACTIONS_SCHEMA_ID = <%= actions_schema_id.to_json.html_safe %>;
|
|
@@ -470,8 +476,12 @@
|
|
|
470
476
|
var REMOTE_TOOLS_SCHEMA_ID = <%= remote_tools_schema_id.to_json.html_safe %>;
|
|
471
477
|
var MAX_ROUNDS = <%= max_rounds.to_json.html_safe %>;
|
|
472
478
|
var WELL_KNOWN_URLS = <%= raw(well_known_urls.nil? ? "null" : well_known_urls.to_json) %>;
|
|
479
|
+
var LLM_PROVIDER = <%= llm_provider.to_json.html_safe %>;
|
|
473
480
|
var ENABLE_MODEL_PICKER = <%= enable_model_picker.to_json.html_safe %>;
|
|
474
|
-
|
|
481
|
+
// Class 1 tools are registered on a hub, so the picker needs one — which
|
|
482
|
+
// is independent of who answers the chat. Page actions and the host's own
|
|
483
|
+
// .well-known MCP work regardless.
|
|
484
|
+
var ENABLE_TOOL_PICKER = <%= enable_tool_picker.to_json.html_safe %> && <%= tool_hub_url.to_json.html_safe %> !== null;
|
|
475
485
|
var MODEL_ALLOWLIST = <%= raw(models.nil? ? "null" : models.to_json) %>;
|
|
476
486
|
var HUB_TOOLS_ALLOWLIST = <%= raw(hub_tools.nil? ? "null" : hub_tools.to_json) %>;
|
|
477
487
|
|
|
@@ -567,7 +577,26 @@
|
|
|
567
577
|
ro.observe(root);
|
|
568
578
|
}
|
|
569
579
|
|
|
570
|
-
|
|
580
|
+
// [{role, content, tools}] — grows across turns. `tools` is for the
|
|
581
|
+
// transcript only; the model is sent role and content.
|
|
582
|
+
var conversation = [];
|
|
583
|
+
|
|
584
|
+
// A page action that navigates used to take the conversation with it. The
|
|
585
|
+
// transcript lives in sessionStorage — per tab, gone when the tab closes —
|
|
586
|
+
// keyed by path, so submitting a form and landing back on the same page
|
|
587
|
+
// with a new query string keeps the thread.
|
|
588
|
+
var conversationStore = createConversationStore({
|
|
589
|
+
storage: window.sessionStorage,
|
|
590
|
+
key: conversationKeyFor(window.location)
|
|
591
|
+
});
|
|
592
|
+
|
|
593
|
+
function persistConversation() {
|
|
594
|
+
conversationStore.save({
|
|
595
|
+
turns: conversation,
|
|
596
|
+
open: !root.classList.contains("lmw-collapsed"),
|
|
597
|
+
model: MODEL
|
|
598
|
+
});
|
|
599
|
+
}
|
|
571
600
|
var actionsSchemaEl = document.getElementById(ACTIONS_SCHEMA_ID);
|
|
572
601
|
var localTools = actionsSchemaEl ? JSON.parse(actionsSchemaEl.textContent) : [];
|
|
573
602
|
var remoteToolsEl = document.getElementById(REMOTE_TOOLS_SCHEMA_ID);
|
|
@@ -633,8 +662,27 @@
|
|
|
633
662
|
// still works with the initial `model:` + no hub tools).
|
|
634
663
|
var tasks = [];
|
|
635
664
|
|
|
636
|
-
if (ENABLE_MODEL_PICKER && modelPicker) {
|
|
637
|
-
|
|
665
|
+
if (ENABLE_MODEL_PICKER && modelPicker && LLM_PROVIDER === "ollama") {
|
|
666
|
+
// No hub to ask: Ollama lists its own models.
|
|
667
|
+
tasks.push(fetchOllamaModels({ baseUrl: LLM_BASE }).then(function(names) {
|
|
668
|
+
var flat = names
|
|
669
|
+
.filter(function(n) { return anyAllowedByAllowlist(n, MODEL_ALLOWLIST); })
|
|
670
|
+
.map(function(n) { return { value: n, label: n }; });
|
|
671
|
+
if (!flat.some(function(m) { return m.value === MODEL; })) {
|
|
672
|
+
flat.unshift({ value: MODEL, label: MODEL });
|
|
673
|
+
}
|
|
674
|
+
modelPicker.innerHTML = "";
|
|
675
|
+
flat.forEach(function(m) {
|
|
676
|
+
var opt = document.createElement("option");
|
|
677
|
+
opt.value = m.value;
|
|
678
|
+
opt.textContent = m.label;
|
|
679
|
+
if (m.value === MODEL) opt.selected = true;
|
|
680
|
+
modelPicker.appendChild(opt);
|
|
681
|
+
});
|
|
682
|
+
if (flat.length > 1) modelPicker.style.display = "";
|
|
683
|
+
}));
|
|
684
|
+
} else if (ENABLE_MODEL_PICKER && modelPicker) {
|
|
685
|
+
tasks.push(fetch(LLM_BASE + "/api/llms", { headers: { "Accept": "application/json" } })
|
|
638
686
|
.then(function(r) { return r.ok ? r.json() : { llms: [] }; })
|
|
639
687
|
.then(function(payload) {
|
|
640
688
|
// /api/llms is heterogeneous per family:
|
|
@@ -671,7 +719,7 @@
|
|
|
671
719
|
}
|
|
672
720
|
|
|
673
721
|
if (ENABLE_TOOL_PICKER && toolsListEl) {
|
|
674
|
-
tasks.push(fetch(
|
|
722
|
+
tasks.push(fetch(TOOL_HUB_BASE + "/api/mcp_servers", { headers: { "Accept": "application/json" } })
|
|
675
723
|
.then(function(r) { return r.ok ? r.json() : { mcp_servers: [] }; })
|
|
676
724
|
.then(function(payload) {
|
|
677
725
|
hubMcpServers = (payload.mcp_servers || []).filter(function(s) {
|
|
@@ -817,6 +865,10 @@
|
|
|
817
865
|
var RESOURCE_BUDGET_BYTES = 8000;
|
|
818
866
|
|
|
819
867
|
|
|
868
|
+
// Before anything asynchronous: if this page was navigated away from and
|
|
869
|
+
// back, the visitor should see their conversation immediately.
|
|
870
|
+
restoreConversation();
|
|
871
|
+
|
|
820
872
|
var wellKnownReady = (async function() {
|
|
821
873
|
var urls = WELL_KNOWN_URLS === null
|
|
822
874
|
? [ window.location.origin + "/.well-known/mcp.json" ]
|
|
@@ -893,11 +945,39 @@
|
|
|
893
945
|
promptsEl.style.display = "";
|
|
894
946
|
}
|
|
895
947
|
|
|
948
|
+
// Put a saved transcript back on screen: the same bubbles, the same
|
|
949
|
+
// markdown, and the same record of which tools ran. Called once at boot,
|
|
950
|
+
// before the welcome block, which is only for a genuinely fresh start.
|
|
951
|
+
function restoreConversation() {
|
|
952
|
+
var saved = conversationStore.load();
|
|
953
|
+
if (!saved || !saved.turns.length) return false;
|
|
954
|
+
|
|
955
|
+
conversation = saved.turns;
|
|
956
|
+
saved.turns.forEach(function(turn) {
|
|
957
|
+
var body = appendTurn(turn.role, turn.role === "assistant" ? "" : turn.content);
|
|
958
|
+
if (turn.role === "assistant") renderMarkdownInto(body, turn.content || "");
|
|
959
|
+
if (!turn.tools || !turn.tools.length) return;
|
|
960
|
+
var row = document.createElement("div");
|
|
961
|
+
row.className = "lmw-tool-chips";
|
|
962
|
+
turn.tools.forEach(function(tool) {
|
|
963
|
+
var chip = document.createElement("span");
|
|
964
|
+
chip.className = "lmw-tool-chip" + (tool.error ? " error" : "");
|
|
965
|
+
chip.textContent = (tool.error ? "❌ " : "🔧 ") + tool.name;
|
|
966
|
+
row.appendChild(chip);
|
|
967
|
+
});
|
|
968
|
+
if (body.parentNode) body.parentNode.appendChild(row);
|
|
969
|
+
});
|
|
970
|
+
historyEl.scrollTop = historyEl.scrollHeight;
|
|
971
|
+
return true;
|
|
972
|
+
}
|
|
973
|
+
|
|
896
974
|
// A blank panel tells a first-time visitor nothing. Open with a greeting
|
|
897
975
|
// and the offers themselves — each template showing what it will take from
|
|
898
976
|
// the page and what it will ask for — so the assistant is the page's way
|
|
899
977
|
// in rather than a box you must already know how to talk to.
|
|
900
978
|
function renderWelcome() {
|
|
979
|
+
// A restored transcript is not a fresh start; the greeting would push
|
|
980
|
+
// the conversation the visitor is mid-way through off the screen.
|
|
901
981
|
if (!historyEl || historyEl.querySelector(".message")) return;
|
|
902
982
|
historyEl.textContent = "";
|
|
903
983
|
|
|
@@ -1210,6 +1290,7 @@
|
|
|
1210
1290
|
clearBtn.addEventListener("click", function() {
|
|
1211
1291
|
if (currentAbort) { try { currentAbort.abort(); } catch (e) { /* noop */ } }
|
|
1212
1292
|
conversation = [];
|
|
1293
|
+
conversationStore.clear();
|
|
1213
1294
|
historyEl.innerHTML = "";
|
|
1214
1295
|
renderWelcome();
|
|
1215
1296
|
});
|
|
@@ -1252,11 +1333,12 @@
|
|
|
1252
1333
|
|
|
1253
1334
|
var resourceLines = await resourceLinesForThisTurn();
|
|
1254
1335
|
var messages = [{ role: "system", content: currentSystemPrompt(resourceLines) }]
|
|
1255
|
-
.concat(conversation)
|
|
1336
|
+
.concat(conversation.map(function(t) { return { role: t.role, content: t.content }; }))
|
|
1256
1337
|
.concat([{ role: "user", content: userText }]);
|
|
1257
1338
|
|
|
1258
1339
|
var result = await runChatLoop({
|
|
1259
|
-
baseUrl:
|
|
1340
|
+
baseUrl: LLM_BASE,
|
|
1341
|
+
toolHubUrl: TOOL_HUB_BASE,
|
|
1260
1342
|
apiKeyUuid: API_KEY_UUID,
|
|
1261
1343
|
modelName: MODEL,
|
|
1262
1344
|
messages: messages,
|
|
@@ -1265,6 +1347,7 @@
|
|
|
1265
1347
|
hostWideTools: hostWideTools,
|
|
1266
1348
|
aiActions: window[ACTIONS_GLOBAL] || {},
|
|
1267
1349
|
maxRounds: MAX_ROUNDS,
|
|
1350
|
+
provider: LLM_PROVIDER === "ollama" ? "ollama" : "hub",
|
|
1268
1351
|
signal: currentAbort.signal,
|
|
1269
1352
|
onToolCall: function(toolCall) { announceToolCall(toolCall); },
|
|
1270
1353
|
onToolDispatched: function(outcome) { resolveToolCall(outcome); },
|
|
@@ -1301,8 +1384,17 @@
|
|
|
1301
1384
|
historyEl.scrollTop = historyEl.scrollHeight;
|
|
1302
1385
|
}
|
|
1303
1386
|
});
|
|
1304
|
-
conversation.push({ role: "user",
|
|
1305
|
-
conversation.push({
|
|
1387
|
+
conversation.push({ role: "user", content: userText });
|
|
1388
|
+
conversation.push({
|
|
1389
|
+
role: "assistant",
|
|
1390
|
+
content: result.content,
|
|
1391
|
+
// Kept so a restored transcript still shows what ran — losing
|
|
1392
|
+
// that record is what made a navigating action unbearable.
|
|
1393
|
+
tools: result.dispatched.map(function(d) {
|
|
1394
|
+
return d.error ? { name: d.toolCall.name, error: true } : { name: d.toolCall.name };
|
|
1395
|
+
})
|
|
1396
|
+
});
|
|
1397
|
+
persistConversation();
|
|
1306
1398
|
|
|
1307
1399
|
// Anything still pending never reported an outcome (a class that
|
|
1308
1400
|
// does not round-trip, or a dispatch that vanished). Reconcile
|