@aria-framework/ai 0.26.1 → 0.27.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +103 -0
- package/README.md +55 -14
- package/browser/ai-panels.js +138 -138
- package/error.js +5 -1
- package/index.js +11 -20
- package/lmxVerify.js +402 -402
- package/package.json +5 -3
- package/providers/anthropic.js +26 -3
- package/providers/lmx.js +14 -56
- package/providers/openai-compatible.js +38 -8
- package/reasoning.js +223 -0
- package/views/ai/job-card.ejs +182 -182
- package/views/ai/lmx-stack.ejs +426 -426
package/CHANGELOG.md
ADDED
|
@@ -0,0 +1,103 @@
|
|
|
1
|
+
# @aria-framework/ai changelog
|
|
2
|
+
|
|
3
|
+
This file starts at 0.26.1. Earlier versions are described in the git history only.
|
|
4
|
+
|
|
5
|
+
Each entry lists the **silent default changes** first: things that change what a caller's existing
|
|
6
|
+
code sends, or what it is billed for, without the caller changing anything.
|
|
7
|
+
|
|
8
|
+
## 0.27.1
|
|
9
|
+
|
|
10
|
+
Fixes from an independent review of 0.26.1 and 0.27.0.
|
|
11
|
+
|
|
12
|
+
### Silent default changes
|
|
13
|
+
|
|
14
|
+
- **A failed call is recorded under the model that ran.** `budget.record(cfg, { usage, model, ms,
|
|
15
|
+
failed: true })` now receives the model the server named in its reply (`err.model`). It used to
|
|
16
|
+
receive `cfg.model`, which is empty on lmx, so every failed lmx call was ledgered with no model. The
|
|
17
|
+
configured model is used only when the reply named none. A usage table keyed on model will see
|
|
18
|
+
failed lmx calls move from `''` to the engine's model.
|
|
19
|
+
- **Older Claude models: `'full'` never asks for more output than the model allows.** On the
|
|
20
|
+
pre-4.6 budget path, `max_tokens` was the caller's `maxTokens` plus a thinking budget of the same
|
|
21
|
+
size, about twice what was asked. That could pass the model's output limit (Haiku 4.5, Sonnet 4.5,
|
|
22
|
+
Opus 4.5, Sonnet 4 and 3.7 Sonnet allow 64000; Opus 4.1 and Opus 4 allow 32000), and the API
|
|
23
|
+
refused it. It is now capped at the limit. Near the limit the thinking budget shrinks first, down
|
|
24
|
+
to the API's minimum of 1024. For example, `maxTokens: 40000` on Haiku 4.5 sends
|
|
25
|
+
`max_tokens: 64000` with a 24000 budget. When `maxTokens` plus the budget is below the limit,
|
|
26
|
+
nothing changes.
|
|
27
|
+
- **Claude 3 models without extended thinking get no thinking.** On `'full'`, Claude 3.5 Haiku,
|
|
28
|
+
3 Haiku, 3 Opus and 3.5 Sonnet were sent `thinking: enabled`, which they reject with a 400. They now
|
|
29
|
+
keep the pre-0.27 request, and `'full'` logs a warning instead of failing the call, as it does
|
|
30
|
+
for a model the table does not know. 3.7 Sonnet still thinks.
|
|
31
|
+
- **Opus 4 and Sonnet 4 under their `-0` aliases and dated ids** (`claude-opus-4-0`,
|
|
32
|
+
`claude-opus-4-20250514`, `claude-sonnet-4-0`, `claude-sonnet-4-20250514`) are now recognised.
|
|
33
|
+
0.27.0 treated them as unknown: `'full'` sent no thinking and logged a warning. They now get the
|
|
34
|
+
budget like the other pre-4.6 models.
|
|
35
|
+
|
|
36
|
+
### Fixes
|
|
37
|
+
|
|
38
|
+
- If a budget's `record` throws `null` (or anything without a `message`), logging it no longer
|
|
39
|
+
throws a `TypeError` that replaces the caller's real error.
|
|
40
|
+
- `AiError` takes an optional `model`, set by both adapters on every error that carries `usage`.
|
|
41
|
+
On lmx, when the reply names no model, it is the model the engine reports running; the same
|
|
42
|
+
fallback now applies to a successful result's `model`.
|
|
43
|
+
- **Anthropic: a `maxTokens` that is not a positive number is refused** (`AiError` kind `refused`,
|
|
44
|
+
before any request) on every Anthropic call. A fraction is rounded down. Before this, a
|
|
45
|
+
string `"2048"` on the budget path became `max_tokens: 64000`, and a negative value sent a
|
|
46
|
+
`budget_tokens` above `max_tokens`. `0`, `NaN` and a missing value still take the 1024 default.
|
|
47
|
+
|
|
48
|
+
### Tests
|
|
49
|
+
|
|
50
|
+
- `test/anthropicRequests.js` (new) compares the request bodies with a fixture typed by hand from
|
|
51
|
+
the Messages API rules for 26 model ids, rather than with values read back from
|
|
52
|
+
`ANTHROPIC_FAMILIES`. It also checks the budget at 1024, 4096, 40000 and at each model's limit.
|
|
53
|
+
- `test/usageOnFailure.js` adds an lmx failure with an empty `cfg.model`, an Anthropic failure whose
|
|
54
|
+
reply names a different model, an empty completion with zero usage on both adapters, and a budget
|
|
55
|
+
that throws `null`.
|
|
56
|
+
|
|
57
|
+
## 0.27.0
|
|
58
|
+
|
|
59
|
+
`complete({ reasoning: 'off' | 'full' })` is honoured by every provider. Before this, only lmx
|
|
60
|
+
honoured it. Leaving it out means `'off'`.
|
|
61
|
+
|
|
62
|
+
### Silent default changes
|
|
63
|
+
|
|
64
|
+
These apply to callers that do not pass `reasoning`. To keep a model thinking, pass
|
|
65
|
+
`reasoning: 'full'`.
|
|
66
|
+
|
|
67
|
+
- **openai-compatible and lmstudio, qwen or gpt-oss models: a body flag is now sent on every call.**
|
|
68
|
+
qwen gets `chat_template_kwargs: { enable_thinking: false }` and gpt-oss gets
|
|
69
|
+
`reasoning_effort: 'low'`. Before this, these models used the server's default, which usually means
|
|
70
|
+
thinking on every call.
|
|
71
|
+
- **LM Studio (`provider: 'lmstudio'`), qwen models: `/no_think` is appended to the system prompt**
|
|
72
|
+
(or sent as the system prompt when there is none), and `reasoning_effort: 'none'` is added to the
|
|
73
|
+
body. LM Studio ignores `chat_template_kwargs`. A caller that compares or caches system prompts will
|
|
74
|
+
see the extra line. An LM Studio server configured as `openai-compatible` gets neither, and keeps
|
|
75
|
+
thinking.
|
|
76
|
+
- **Claude Opus 5.5 and Fable (and Mythos) get `output_config.effort: 'low'`.** These models cannot
|
|
77
|
+
turn thinking off, and Opus 5.5's own default effort is `medium`, so leaving `reasoning` out now
|
|
78
|
+
means less thinking than before.
|
|
79
|
+
- **Other Claude generations:** Opus 5 gets effort `'low'` (its own default is adaptive thinking at
|
|
80
|
+
`high`). Sonnet 5.5 gets `thinking: { type: 'between_tools' }` and Sonnet 5 gets
|
|
81
|
+
`thinking: { type: 'disabled' }`; before this, both thought by default. 4.6 and older are
|
|
82
|
+
unchanged unless `'full'` is passed.
|
|
83
|
+
- **`temperature` is no longer sent to Claude models that reject it** (Opus 4.7 and later, Sonnet 5
|
|
84
|
+
and later, Fable). Before this, those requests failed with a 400. It is also left out with `'full'`
|
|
85
|
+
on the pre-4.6 budget path, where extended thinking requires the default.
|
|
86
|
+
- The client no longer strips `reasoning` with a warning saying only lmx honours it.
|
|
87
|
+
|
|
88
|
+
## 0.26.1
|
|
89
|
+
|
|
90
|
+
### Silent default changes
|
|
91
|
+
|
|
92
|
+
- **A failed call that spent tokens is recorded through `budget.record`.** When the model answered
|
|
93
|
+
but the answer was unusable (it thought until the limit, or a structured reply was cut off),
|
|
94
|
+
`budget.record(cfg, { usage, model, ms, failed: true }, meta)` is called before the error is
|
|
95
|
+
thrown. A failure that spent nothing (refused, unreachable, timed out, or a zero usage total)
|
|
96
|
+
records nothing. **A caller that metered `err.usage` itself must stop, or the call is counted
|
|
97
|
+
twice.** A `record` that throws is logged and never replaces the original error.
|
|
98
|
+
|
|
99
|
+
### Fixes
|
|
100
|
+
|
|
101
|
+
- Both adapters put normalised usage (`{ prompt, completion, total }`) on every error raised after
|
|
102
|
+
the reply was read.
|
|
103
|
+
- A 200 whose JSON body is `null` is an `AiError` of kind `bad_response`. It used to be a `TypeError`.
|
package/README.md
CHANGED
|
@@ -35,7 +35,7 @@ const ai = createAiClient({
|
|
|
35
35
|
const r = await ai.complete({
|
|
36
36
|
system, messages, maxTokens: 400,
|
|
37
37
|
schema, // optional: constrained JSON
|
|
38
|
-
// reasoning: 'full', // optional,
|
|
38
|
+
// reasoning: 'full', // optional, every provider: 'off' (default) or 'full' (see Reasoning below)
|
|
39
39
|
});
|
|
40
40
|
// r.text, r.json (when schema), r.model, r.usage.total, r.ms
|
|
41
41
|
```
|
|
@@ -48,7 +48,10 @@ an admin screen and never includes a URL or key.
|
|
|
48
48
|
was unusable (it thought until the limit, or a structured reply was cut off), the error carries
|
|
49
49
|
`usage` (`{ prompt, completion, total }`) and `budget.record(cfg, { usage, model, ms, failed: true }, meta)`
|
|
50
50
|
is called before the error is thrown. A failure that spent nothing (refused, unreachable, timeout)
|
|
51
|
-
records nothing. If `record` throws, the original error still reaches the caller.
|
|
51
|
+
records nothing. If `record` throws, the original error still reaches the caller. The error also
|
|
52
|
+
carries `model`, the model the server says answered, and `record` receives that (0.27.1); it falls
|
|
53
|
+
back to the configured model only when the reply named none. On lmx the configured model is empty.
|
|
54
|
+
A caller that metered `err.usage` itself before 0.26.1 must stop, or the call is counted twice.
|
|
52
55
|
|
|
53
56
|
`providerStore`, `usageStore`, `speedStore` and `lmxStore` are optional db-worker-backed
|
|
54
57
|
stores. Each exports a `schemaFor(dialect)`, so the app's migration can be checked against it.
|
|
@@ -69,7 +72,7 @@ to write yourself:
|
|
|
69
72
|
| Engine URLs come from the document | Resolved on every call, and never stored |
|
|
70
73
|
| A status outage is not an inference outage | Keeps routing on the last good document for 3 minutes (`staleMs`), logging loudly |
|
|
71
74
|
| Pin the certificate; never disable verification | Uses an undici `Agent({ connect: { ca } })` per stack. Never `NODE_EXTRA_CA_CERTS`, never `rejectUnauthorized:false` |
|
|
72
|
-
| Choose the reasoning flag from the model | By default thinking is kept to a minimum: `qwen` → `chat_template_kwargs.enable_thinking=false`, `gpt-oss` → `reasoning_effort:'low'`, anything else gets neither (with a warning). `reasoning: 'full'` on the call (since 0.26.0) sends `enable_thinking=true` / `reasoning_effort:'high'` instead. The flag always overrides a raw `extra`; only the named option changes it. See [Reasoning](#reasoning-reasoning-full) |
|
|
75
|
+
| Choose the reasoning flag from the model | By default thinking is kept to a minimum: `qwen` → `chat_template_kwargs.enable_thinking=false`, `gpt-oss` → `reasoning_effort:'low'`, anything else gets neither (with a warning). `reasoning: 'full'` on the call (since 0.26.0) sends `enable_thinking=true` / `reasoning_effort:'high'` instead. The flag always overrides a raw `extra`; only the named option changes it. Since 0.27.0 the same rule applies to every provider. See [Reasoning](#reasoning-reasoning-off--full) |
|
|
73
76
|
| Size requests against `maxInputTokens` (per slot) | `contextTokens` is taken from the engine, not from config |
|
|
74
77
|
| Identity is `(instance, id)`, falling back to `name` | `lmx.engineId` is matched first; `rekeyPlan()` migrates rows when ids are minted or engines renamed |
|
|
75
78
|
| A `429` is about the key, not the engine | Fails fast (since 0.25.0): throws `rate_limit` with `lmxSkip: 'lmx_throttled'` and `retryAfterMs`. The client never waits; whether to wait or move on is the caller's decision |
|
|
@@ -176,10 +179,11 @@ instead:
|
|
|
176
179
|
Failover across several engines for one job is the app's job; lmx provides none. List
|
|
177
180
|
endpoints in preference order and take the first one that serves.
|
|
178
181
|
|
|
179
|
-
### Reasoning (`reasoning: 'full'`)
|
|
182
|
+
### Reasoning (`reasoning: 'off' | 'full'`)
|
|
180
183
|
|
|
181
|
-
|
|
182
|
-
|
|
184
|
+
How much a model may think is decided by the **app, per call**, and every provider honours it (since
|
|
185
|
+
0.27.0; before that only lmx did). Leave it out, or pass `reasoning: 'off'`, and the model thinks as
|
|
186
|
+
little as its family allows, so a quick answer stays quick. Pass `reasoning: 'full'` to let it think:
|
|
183
187
|
|
|
184
188
|
```js
|
|
185
189
|
const r = await ai.complete({
|
|
@@ -204,23 +208,60 @@ a `bad_response` AiError saying it ran out of room (and, when the server says so
|
|
|
204
208
|
tokens went on internal reasoning), but the work is lost either way. Use **`maxTokens` ≥ 4096** and a
|
|
205
209
|
**`timeoutMs` ≥ 120 s** (120000) on the endpoint, or the deadline cuts the answer off first.
|
|
206
210
|
|
|
207
|
-
**
|
|
211
|
+
**lmx, openai-compatible and lmstudio: what is added to the request body, per model family.** The
|
|
212
|
+
family is read from the model the engine is running (lmx) or the endpoint's model name (others):
|
|
208
213
|
|
|
209
|
-
| Family |
|
|
214
|
+
| Family | `'off'` (the default) | `reasoning: 'full'` |
|
|
210
215
|
|---|---|---|
|
|
211
216
|
| qwen | `{"chat_template_kwargs":{"enable_thinking":false}}` | `{"chat_template_kwargs":{"enable_thinking":true}}` |
|
|
212
217
|
| gpt-oss | `{"reasoning_effort":"low"}` | `{"reasoning_effort":"high"}` |
|
|
213
|
-
| anything else | nothing (warning
|
|
218
|
+
| anything else | nothing (lmx logs a warning) | nothing (a warning that `'full'` could not be honoured) |
|
|
219
|
+
|
|
220
|
+
**LM Studio ignores `chat_template_kwargs`** (measured live: Qwen3 thought just as long with
|
|
221
|
+
`enable_thinking: false`). So on **`provider: 'lmstudio'`** a qwen model gets two more switches. One is
|
|
222
|
+
`{"reasoning_effort":"none"}` for `'off'` and `{"reasoning_effort":"high"}` for `'full'`: verified live, but
|
|
223
|
+
LM Studio documents `reasoning_effort` only for gpt-oss. The other is Qwen's own documented prompt switch,
|
|
224
|
+
`/no_think` for `'off'` and `/think` for `'full'`, appended to the system prompt (or sent as one when there
|
|
225
|
+
is none): documented on LM Studio's Qwen3 model pages, but only hybrid Qwen3 reads it. Both are sent, so
|
|
226
|
+
if a future LM Studio drops one the other still holds. Configure an LM Studio server as `lmstudio`, not
|
|
227
|
+
`openai-compatible`; as `openai-compatible` it keeps thinking on every call. Plain `openai-compatible`
|
|
228
|
+
gets neither, because vLLM validates `reasoning_effort` and may refuse `'none'`.
|
|
229
|
+
|
|
230
|
+
**anthropic: what is sent, per Claude generation.** The Messages API differs by generation: Opus 5.5
|
|
231
|
+
and Fable cannot turn thinking off (effort is the only lever), Sonnet 5.5 turns it off with
|
|
232
|
+
`between_tools`, and several current models refuse `temperature`, so none is sent to them.
|
|
233
|
+
|
|
234
|
+
| Models | `'off'` (the default) | `reasoning: 'full'` | temperature |
|
|
235
|
+
|---|---|---|---|
|
|
236
|
+
| claude-fable / claude-opus-5-5 | `{"output_config":{"effort":"low"}}` | `{"output_config":{"effort":"high"}}` | not sent |
|
|
237
|
+
| claude-opus-5 | `{"output_config":{"effort":"low"}}` | `{"thinking":{"type":"adaptive"},"output_config":{"effort":"high"}}` | not sent |
|
|
238
|
+
| claude-sonnet-5-5 | `{"thinking":{"type":"between_tools"}}` | `{"thinking":{"type":"adaptive"},"output_config":{"effort":"high"}}` | not sent |
|
|
239
|
+
| claude-sonnet-5 | `{"thinking":{"type":"disabled"}}` | `{"thinking":{"type":"adaptive"},"output_config":{"effort":"high"}}` | not sent |
|
|
240
|
+
| claude-opus-4-7 / 4-8 | nothing | `{"thinking":{"type":"adaptive"},"output_config":{"effort":"high"}}` | not sent |
|
|
241
|
+
| claude-*-4-6 | nothing | `{"thinking":{"type":"adaptive"},"output_config":{"effort":"high"}}` | sent |
|
|
242
|
+
| claude-*-4-5 | nothing | thinking enabled with a budget of `maxTokens` (at least 1024), added on top of `max_tokens`, which never exceeds 64000 | sent; not sent with `'full'` |
|
|
243
|
+
| claude-opus-4-1 / claude-opus-4 | nothing | thinking enabled with a budget of `maxTokens` (at least 1024), added on top of `max_tokens`, which never exceeds 32000 | sent; not sent with `'full'` |
|
|
244
|
+
| claude-sonnet-4 | nothing | thinking enabled with a budget of `maxTokens` (at least 1024), added on top of `max_tokens`, which never exceeds 64000 | sent; not sent with `'full'` |
|
|
245
|
+
| claude-3-7-sonnet | nothing | thinking enabled with a budget of `maxTokens` (at least 1024), added on top of `max_tokens`, which never exceeds 64000 | sent; not sent with `'full'` |
|
|
246
|
+
| claude-3 without extended thinking | nothing | nothing (a warning that `'full'` could not be honoured) | sent |
|
|
247
|
+
| anything else | nothing | nothing (a warning that `'full'` could not be honoured) | sent |
|
|
248
|
+
|
|
249
|
+
`effort` is merged into any `output_config` the structured-output schema already put there.
|
|
250
|
+
|
|
251
|
+
On the budget rows `max_tokens` counts the thinking as well as the answer, so it is capped at the
|
|
252
|
+
model's output limit (0.27.1). Near the limit the thinking budget shrinks first, down to the API's
|
|
253
|
+
minimum of 1024, and only then the answer's room: `maxTokens: 40000` on Haiku 4.5 sends
|
|
254
|
+
`max_tokens: 64000` with a 24000 budget. Claude 3 models other than 3.7 Sonnet have no extended
|
|
255
|
+
thinking, so `'full'` sends nothing to them and logs a warning; it does not fail the call.
|
|
214
256
|
|
|
215
257
|
Rules:
|
|
216
258
|
|
|
217
|
-
- `'
|
|
218
|
-
for example `'high'`, `true` or `''`, is refused with `AiError` kind `refused` before any
|
|
259
|
+
- The values are `'off'` and `'full'`. Leaving the option out (or `null`) means `'off'`. Any other
|
|
260
|
+
value, for example `'high'`, `true` or `''`, is refused with `AiError` kind `refused` before any
|
|
219
261
|
request is made, whichever provider the endpoint uses.
|
|
220
262
|
- The flag is merged **after** `extra`, so a raw `extra: { chat_template_kwargs: … }` or
|
|
221
|
-
`extra: { reasoning_effort: … }` cannot change it. Only `reasoning` can.
|
|
222
|
-
|
|
223
|
-
warning, removes the option, and nothing reaches that provider's request body.
|
|
263
|
+
`extra: { reasoning_effort: … }` cannot change it. Only `reasoning` can. The option itself is never
|
|
264
|
+
copied into a request body.
|
|
224
265
|
- `REASONING_MODES` (package root) lists the accepted values for an app that validates its own
|
|
225
266
|
settings.
|
|
226
267
|
|
package/browser/ai-panels.js
CHANGED
|
@@ -1,138 +1,138 @@
|
|
|
1
|
-
/**
|
|
2
|
-
* The fold on the Inference list, and prefilling the supervisor form from a panel.
|
|
3
|
-
*
|
|
4
|
-
* WHY ANYTHING FOLDS. A stable stack is four engines, three credentials, a certificate fingerprint
|
|
5
|
-
* and a verification report — worth having, not worth reading every time you open this page to do
|
|
6
|
-
* something else. Folded, a panel is one line that still carries the whole verdict: status, the
|
|
7
|
-
* count of what depends on it, all three credentials and when it was last checked. Folding hides
|
|
8
|
-
* detail; it must never hide the reason you would have opened it.
|
|
9
|
-
*
|
|
10
|
-
* A PANEL WITH A PROBLEM IGNORES WHAT YOU REMEMBERED. `data-panel-attention` is stamped by the
|
|
11
|
-
* server on anything with a missing engine, an expiring certificate, a failed check or a nearly
|
|
12
|
-
* spent cap. Those open and stay open. Attention beats tidiness — the alternative is a page that
|
|
13
|
-
* quietly honours a fold you chose last week and hides the thing that broke yesterday.
|
|
14
|
-
*
|
|
15
|
-
* THE PREFERENCE IS PER BROWSER AND DISPOSABLE. It is a convenience about how a page looks to one
|
|
16
|
-
* person, so localStorage is the right home and losing it costs nothing. Every access is guarded:
|
|
17
|
-
* a browser set to block site data throws on read, and a settings page must not break because
|
|
18
|
-
* somebody tightened their privacy settings.
|
|
19
|
-
*/
|
|
20
|
-
(function () {
|
|
21
|
-
'use strict';
|
|
22
|
-
|
|
23
|
-
var KEY = 's101.ai.panels';
|
|
24
|
-
|
|
25
|
-
function readPrefs() {
|
|
26
|
-
try {
|
|
27
|
-
return JSON.parse(window.localStorage.getItem(KEY) || '{}') || {};
|
|
28
|
-
} catch (e) {
|
|
29
|
-
return {};
|
|
30
|
-
}
|
|
31
|
-
}
|
|
32
|
-
|
|
33
|
-
function writePref(id, open) {
|
|
34
|
-
try {
|
|
35
|
-
var prefs = readPrefs();
|
|
36
|
-
prefs[id] = !!open;
|
|
37
|
-
window.localStorage.setItem(KEY, JSON.stringify(prefs));
|
|
38
|
-
} catch (e) { /* private window, or site data blocked — the fold still works for this visit */ }
|
|
39
|
-
}
|
|
40
|
-
|
|
41
|
-
function bodyOf(panel) {
|
|
42
|
-
var btn = panel.querySelector('[data-panel-toggle]');
|
|
43
|
-
if (!btn) return null;
|
|
44
|
-
return document.getElementById(btn.getAttribute('aria-controls'));
|
|
45
|
-
}
|
|
46
|
-
|
|
47
|
-
function setOpen(panel, open, remember) {
|
|
48
|
-
var btn = panel.querySelector('[data-panel-toggle]');
|
|
49
|
-
var body = bodyOf(panel);
|
|
50
|
-
if (!btn || !body) return;
|
|
51
|
-
// `hidden`, not a style: the server renders the closed state the same way, so a panel does not
|
|
52
|
-
// flicker open on load before this script runs.
|
|
53
|
-
body.hidden = !open;
|
|
54
|
-
btn.setAttribute('aria-expanded', open ? 'true' : 'false');
|
|
55
|
-
panel.classList.toggle('panel-open', open);
|
|
56
|
-
if (remember) writePref(panel.getAttribute('data-panel'), open);
|
|
57
|
-
}
|
|
58
|
-
|
|
59
|
-
function restore() {
|
|
60
|
-
var prefs = readPrefs();
|
|
61
|
-
var panels = document.querySelectorAll('[data-panel]');
|
|
62
|
-
for (var i = 0; i < panels.length; i += 1) {
|
|
63
|
-
var panel = panels[i];
|
|
64
|
-
var id = panel.getAttribute('data-panel');
|
|
65
|
-
// The server already opened this one and means it. Do not consult the preference at all —
|
|
66
|
-
// reading it and then ignoring it is the same thing, but invites somebody to "fix" it later.
|
|
67
|
-
if (panel.hasAttribute('data-panel-attention')) {
|
|
68
|
-
setOpen(panel, true, false);
|
|
69
|
-
continue;
|
|
70
|
-
}
|
|
71
|
-
if (Object.prototype.hasOwnProperty.call(prefs, id)) setOpen(panel, !!prefs[id], false);
|
|
72
|
-
}
|
|
73
|
-
}
|
|
74
|
-
|
|
75
|
-
document.addEventListener('click', function (ev) {
|
|
76
|
-
var toggle = ev.target.closest('[data-panel-toggle]');
|
|
77
|
-
if (toggle) {
|
|
78
|
-
var panel = toggle.closest('[data-panel]');
|
|
79
|
-
if (!panel) return;
|
|
80
|
-
var body = bodyOf(panel);
|
|
81
|
-
setOpen(panel, !!(body && body.hidden), true);
|
|
82
|
-
return;
|
|
83
|
-
}
|
|
84
|
-
|
|
85
|
-
// ── EDITING A STACK HAPPENS IN THE PANEL ────────────────────────────────────────────────────
|
|
86
|
-
//
|
|
87
|
-
// It used to happen in a form at the foot of the page, and Edit scrolled you down to it — away
|
|
88
|
-
// from the stack you were reading, into a form that also served as "Add a supervisor" and said
|
|
89
|
-
// nothing about which stack it had been filled with. Nothing needs prefilling now: every panel
|
|
90
|
-
// renders its own form, from the server, with its own values already in it.
|
|
91
|
-
var open = ev.target.closest('[data-stack-edit-open]');
|
|
92
|
-
if (open) { stackMode(open.closest('[data-stack]'), true); return; }
|
|
93
|
-
|
|
94
|
-
var cancel = ev.target.closest('[data-stack-cancel]');
|
|
95
|
-
if (cancel) {
|
|
96
|
-
var panel2 = cancel.closest('[data-stack]');
|
|
97
|
-
// A NEW stack has no reading half to go back to, so cancelling it puts the blank panel away.
|
|
98
|
-
if (panel2 && panel2.hasAttribute('data-stack-new')) { panel2.hidden = true; return; }
|
|
99
|
-
// Otherwise reload rather than restore: a Cancel that left half-typed values behind, ready to
|
|
100
|
-
// be posted by the next Save, would be worse than the extra request.
|
|
101
|
-
window.location.reload();
|
|
102
|
-
return;
|
|
103
|
-
}
|
|
104
|
-
|
|
105
|
-
// "Add LMX" reveals the blank panel that is already on the page, at the top of the list.
|
|
106
|
-
var add = ev.target.closest('[data-stack-add]');
|
|
107
|
-
if (add) {
|
|
108
|
-
var blank = document.querySelector('[data-stack-new]');
|
|
109
|
-
if (!blank) return;
|
|
110
|
-
blank.hidden = false;
|
|
111
|
-
blank.scrollIntoView({ block: 'nearest' });
|
|
112
|
-
var id = blank.querySelector('[name="id"]');
|
|
113
|
-
if (id) id.focus();
|
|
114
|
-
}
|
|
115
|
-
});
|
|
116
|
-
|
|
117
|
-
/** Swap one stack panel between reading and editing. */
|
|
118
|
-
function stackMode(panel, editing) {
|
|
119
|
-
if (!panel) return;
|
|
120
|
-
var view = panel.querySelector('[data-stack-view]');
|
|
121
|
-
var form = panel.querySelector('[data-stack-edit]');
|
|
122
|
-
if (!view || !form) return;
|
|
123
|
-
view.hidden = editing;
|
|
124
|
-
form.hidden = !editing;
|
|
125
|
-
if (editing) {
|
|
126
|
-
// The id is readonly on an existing stack, so focus the first field somebody can actually
|
|
127
|
-
// change rather than one that will not accept a keystroke.
|
|
128
|
-
var first = form.querySelector('[name="label"]') || form.querySelector('[name="status_url"]');
|
|
129
|
-
if (first) first.focus();
|
|
130
|
-
}
|
|
131
|
-
}
|
|
132
|
-
|
|
133
|
-
if (document.readyState === 'loading') {
|
|
134
|
-
document.addEventListener('DOMContentLoaded', restore);
|
|
135
|
-
} else {
|
|
136
|
-
restore();
|
|
137
|
-
}
|
|
138
|
-
})();
|
|
1
|
+
/**
|
|
2
|
+
* The fold on the Inference list, and prefilling the supervisor form from a panel.
|
|
3
|
+
*
|
|
4
|
+
* WHY ANYTHING FOLDS. A stable stack is four engines, three credentials, a certificate fingerprint
|
|
5
|
+
* and a verification report — worth having, not worth reading every time you open this page to do
|
|
6
|
+
* something else. Folded, a panel is one line that still carries the whole verdict: status, the
|
|
7
|
+
* count of what depends on it, all three credentials and when it was last checked. Folding hides
|
|
8
|
+
* detail; it must never hide the reason you would have opened it.
|
|
9
|
+
*
|
|
10
|
+
* A PANEL WITH A PROBLEM IGNORES WHAT YOU REMEMBERED. `data-panel-attention` is stamped by the
|
|
11
|
+
* server on anything with a missing engine, an expiring certificate, a failed check or a nearly
|
|
12
|
+
* spent cap. Those open and stay open. Attention beats tidiness — the alternative is a page that
|
|
13
|
+
* quietly honours a fold you chose last week and hides the thing that broke yesterday.
|
|
14
|
+
*
|
|
15
|
+
* THE PREFERENCE IS PER BROWSER AND DISPOSABLE. It is a convenience about how a page looks to one
|
|
16
|
+
* person, so localStorage is the right home and losing it costs nothing. Every access is guarded:
|
|
17
|
+
* a browser set to block site data throws on read, and a settings page must not break because
|
|
18
|
+
* somebody tightened their privacy settings.
|
|
19
|
+
*/
|
|
20
|
+
(function () {
|
|
21
|
+
'use strict';
|
|
22
|
+
|
|
23
|
+
var KEY = 's101.ai.panels';
|
|
24
|
+
|
|
25
|
+
function readPrefs() {
|
|
26
|
+
try {
|
|
27
|
+
return JSON.parse(window.localStorage.getItem(KEY) || '{}') || {};
|
|
28
|
+
} catch (e) {
|
|
29
|
+
return {};
|
|
30
|
+
}
|
|
31
|
+
}
|
|
32
|
+
|
|
33
|
+
function writePref(id, open) {
|
|
34
|
+
try {
|
|
35
|
+
var prefs = readPrefs();
|
|
36
|
+
prefs[id] = !!open;
|
|
37
|
+
window.localStorage.setItem(KEY, JSON.stringify(prefs));
|
|
38
|
+
} catch (e) { /* private window, or site data blocked — the fold still works for this visit */ }
|
|
39
|
+
}
|
|
40
|
+
|
|
41
|
+
function bodyOf(panel) {
|
|
42
|
+
var btn = panel.querySelector('[data-panel-toggle]');
|
|
43
|
+
if (!btn) return null;
|
|
44
|
+
return document.getElementById(btn.getAttribute('aria-controls'));
|
|
45
|
+
}
|
|
46
|
+
|
|
47
|
+
function setOpen(panel, open, remember) {
|
|
48
|
+
var btn = panel.querySelector('[data-panel-toggle]');
|
|
49
|
+
var body = bodyOf(panel);
|
|
50
|
+
if (!btn || !body) return;
|
|
51
|
+
// `hidden`, not a style: the server renders the closed state the same way, so a panel does not
|
|
52
|
+
// flicker open on load before this script runs.
|
|
53
|
+
body.hidden = !open;
|
|
54
|
+
btn.setAttribute('aria-expanded', open ? 'true' : 'false');
|
|
55
|
+
panel.classList.toggle('panel-open', open);
|
|
56
|
+
if (remember) writePref(panel.getAttribute('data-panel'), open);
|
|
57
|
+
}
|
|
58
|
+
|
|
59
|
+
function restore() {
|
|
60
|
+
var prefs = readPrefs();
|
|
61
|
+
var panels = document.querySelectorAll('[data-panel]');
|
|
62
|
+
for (var i = 0; i < panels.length; i += 1) {
|
|
63
|
+
var panel = panels[i];
|
|
64
|
+
var id = panel.getAttribute('data-panel');
|
|
65
|
+
// The server already opened this one and means it. Do not consult the preference at all —
|
|
66
|
+
// reading it and then ignoring it is the same thing, but invites somebody to "fix" it later.
|
|
67
|
+
if (panel.hasAttribute('data-panel-attention')) {
|
|
68
|
+
setOpen(panel, true, false);
|
|
69
|
+
continue;
|
|
70
|
+
}
|
|
71
|
+
if (Object.prototype.hasOwnProperty.call(prefs, id)) setOpen(panel, !!prefs[id], false);
|
|
72
|
+
}
|
|
73
|
+
}
|
|
74
|
+
|
|
75
|
+
document.addEventListener('click', function (ev) {
|
|
76
|
+
var toggle = ev.target.closest('[data-panel-toggle]');
|
|
77
|
+
if (toggle) {
|
|
78
|
+
var panel = toggle.closest('[data-panel]');
|
|
79
|
+
if (!panel) return;
|
|
80
|
+
var body = bodyOf(panel);
|
|
81
|
+
setOpen(panel, !!(body && body.hidden), true);
|
|
82
|
+
return;
|
|
83
|
+
}
|
|
84
|
+
|
|
85
|
+
// ── EDITING A STACK HAPPENS IN THE PANEL ────────────────────────────────────────────────────
|
|
86
|
+
//
|
|
87
|
+
// It used to happen in a form at the foot of the page, and Edit scrolled you down to it — away
|
|
88
|
+
// from the stack you were reading, into a form that also served as "Add a supervisor" and said
|
|
89
|
+
// nothing about which stack it had been filled with. Nothing needs prefilling now: every panel
|
|
90
|
+
// renders its own form, from the server, with its own values already in it.
|
|
91
|
+
var open = ev.target.closest('[data-stack-edit-open]');
|
|
92
|
+
if (open) { stackMode(open.closest('[data-stack]'), true); return; }
|
|
93
|
+
|
|
94
|
+
var cancel = ev.target.closest('[data-stack-cancel]');
|
|
95
|
+
if (cancel) {
|
|
96
|
+
var panel2 = cancel.closest('[data-stack]');
|
|
97
|
+
// A NEW stack has no reading half to go back to, so cancelling it puts the blank panel away.
|
|
98
|
+
if (panel2 && panel2.hasAttribute('data-stack-new')) { panel2.hidden = true; return; }
|
|
99
|
+
// Otherwise reload rather than restore: a Cancel that left half-typed values behind, ready to
|
|
100
|
+
// be posted by the next Save, would be worse than the extra request.
|
|
101
|
+
window.location.reload();
|
|
102
|
+
return;
|
|
103
|
+
}
|
|
104
|
+
|
|
105
|
+
// "Add LMX" reveals the blank panel that is already on the page, at the top of the list.
|
|
106
|
+
var add = ev.target.closest('[data-stack-add]');
|
|
107
|
+
if (add) {
|
|
108
|
+
var blank = document.querySelector('[data-stack-new]');
|
|
109
|
+
if (!blank) return;
|
|
110
|
+
blank.hidden = false;
|
|
111
|
+
blank.scrollIntoView({ block: 'nearest' });
|
|
112
|
+
var id = blank.querySelector('[name="id"]');
|
|
113
|
+
if (id) id.focus();
|
|
114
|
+
}
|
|
115
|
+
});
|
|
116
|
+
|
|
117
|
+
/** Swap one stack panel between reading and editing. */
|
|
118
|
+
function stackMode(panel, editing) {
|
|
119
|
+
if (!panel) return;
|
|
120
|
+
var view = panel.querySelector('[data-stack-view]');
|
|
121
|
+
var form = panel.querySelector('[data-stack-edit]');
|
|
122
|
+
if (!view || !form) return;
|
|
123
|
+
view.hidden = editing;
|
|
124
|
+
form.hidden = !editing;
|
|
125
|
+
if (editing) {
|
|
126
|
+
// The id is readonly on an existing stack, so focus the first field somebody can actually
|
|
127
|
+
// change rather than one that will not accept a keystroke.
|
|
128
|
+
var first = form.querySelector('[name="label"]') || form.querySelector('[name="status_url"]');
|
|
129
|
+
if (first) first.focus();
|
|
130
|
+
}
|
|
131
|
+
}
|
|
132
|
+
|
|
133
|
+
if (document.readyState === 'loading') {
|
|
134
|
+
document.addEventListener('DOMContentLoaded', restore);
|
|
135
|
+
} else {
|
|
136
|
+
restore();
|
|
137
|
+
}
|
|
138
|
+
})();
|
package/error.js
CHANGED
|
@@ -18,7 +18,7 @@ class AiError extends Error {
|
|
|
18
18
|
/**
|
|
19
19
|
* @param {AiErrorKind} kind
|
|
20
20
|
* @param {string} message written for the admin screen, not the log
|
|
21
|
-
* @param {{cause?: Error, status?: number, usage?: object}} [opts]
|
|
21
|
+
* @param {{cause?: Error, status?: number, usage?: object, model?: string}} [opts]
|
|
22
22
|
*/
|
|
23
23
|
constructor(kind, message, opts = {}) {
|
|
24
24
|
super(message);
|
|
@@ -30,6 +30,10 @@ class AiError extends Error {
|
|
|
30
30
|
// measuring throughput should not have to treat that as zero work. Optional everywhere: most
|
|
31
31
|
// failures happen before a single token is spent.
|
|
32
32
|
if (opts.usage) this.usage = opts.usage;
|
|
33
|
+
// ...AND WHICH MODEL SPENT IT (0.27.1), as the server named it in its reply. The client meters a
|
|
34
|
+
// failed call under this, not cfg.model: on lmx cfg.model is empty (the engine decides what
|
|
35
|
+
// runs), and an alias on any provider resolves to an id other than the one configured.
|
|
36
|
+
if (opts.model) this.model = opts.model;
|
|
33
37
|
if (opts.cause) this.cause = opts.cause;
|
|
34
38
|
}
|
|
35
39
|
|
package/index.js
CHANGED
|
@@ -26,7 +26,7 @@ const facts = require('./facts');
|
|
|
26
26
|
const { AiError, fromFetchFailure, redact } = require('./error');
|
|
27
27
|
const { polish } = require('./polish');
|
|
28
28
|
const { generate } = require('./generate');
|
|
29
|
-
const { assertReasoning, REASONING_MODES } = require('./
|
|
29
|
+
const { assertReasoning, REASONING_MODES } = require('./reasoning');
|
|
30
30
|
|
|
31
31
|
const PROVIDERS = {
|
|
32
32
|
// 'lmstudio' and 'openai-compatible' are the SAME adapter with different defaults — a kindness to
|
|
@@ -52,9 +52,6 @@ const DEFAULTS = {
|
|
|
52
52
|
const RETRY_AFTER_MS = 400;
|
|
53
53
|
const RETRY_ONLY_IF_FAILED_WITHIN_MS = 5000;
|
|
54
54
|
|
|
55
|
-
/** Providers already warned that they ignore `reasoning` — once per provider per process. */
|
|
56
|
-
const warnedReasoning = new Set();
|
|
57
|
-
|
|
58
55
|
const NOOP_LOGGER = { info() {}, warn() {}, error() {} };
|
|
59
56
|
const NOOP_BUDGET = { async assertWithinBudget() {}, async record() {} };
|
|
60
57
|
|
|
@@ -106,7 +103,7 @@ function createAiClient(deps = {}) {
|
|
|
106
103
|
* Ask the configured model for something.
|
|
107
104
|
* @param {{system?:string, messages:Array, maxTokens?:number, temperature?:number,
|
|
108
105
|
* schema?:object, signal?:AbortSignal, ticketId?:*, skipBudget?:boolean,
|
|
109
|
-
* reasoning?:'full'}} opts
|
|
106
|
+
* reasoning?:'off'|'full'}} opts
|
|
110
107
|
* @param {object} [cfgOverride] the resolved config, when the caller already has it
|
|
111
108
|
*/
|
|
112
109
|
async function complete(opts, cfgOverride) {
|
|
@@ -134,19 +131,9 @@ function createAiClient(deps = {}) {
|
|
|
134
131
|
// by whichever caller forgets. `skipBudget` is for the admin's Test connection; it still records.
|
|
135
132
|
if (!opts.skipBudget) await meter.assertWithinBudget(cfg, { ticketId: opts.ticketId });
|
|
136
133
|
|
|
137
|
-
//
|
|
138
|
-
//
|
|
139
|
-
|
|
140
|
-
if (opts.reasoning != null && cfg.provider !== 'lmx') {
|
|
141
|
-
if (!warnedReasoning.has(cfg.provider)) {
|
|
142
|
-
warnedReasoning.add(cfg.provider);
|
|
143
|
-
log.warn(`AI: reasoning '${opts.reasoning}' is honoured only by lmx engines — `
|
|
144
|
-
+ `${cfg.label || cfg.provider} sends no reasoning flag and runs on its own default `
|
|
145
|
-
+ '(warned once per provider).');
|
|
146
|
-
}
|
|
147
|
-
const { reasoning: _ignored, ...rest } = opts;
|
|
148
|
-
callOpts = rest;
|
|
149
|
-
}
|
|
134
|
+
// EVERY PROVIDER HONOURS `reasoning` (0.27.0): each adapter translates it for its model family
|
|
135
|
+
// (see reasoning.js). It used to be lmx-only, warned about and stripped everywhere else.
|
|
136
|
+
const callOpts = opts;
|
|
150
137
|
|
|
151
138
|
const started = Date.now();
|
|
152
139
|
let result;
|
|
@@ -158,12 +145,16 @@ function createAiClient(deps = {}) {
|
|
|
158
145
|
// connection test on such a model run outside the ceiling. Adapters put normalised usage on
|
|
159
146
|
// the error; a failure that spent nothing (refused, unreachable) carries none and is not counted.
|
|
160
147
|
// The ledger failing must never replace the error the caller needs to see.
|
|
148
|
+
// THE MODEL THAT RAN (0.27.1), as the adapter read it off the reply — cfg.model only when the
|
|
149
|
+
// reply named none. On lmx cfg.model is '' and every failed call was ledgered under no model.
|
|
161
150
|
if (err && err.usage && err.usage.total > 0) {
|
|
162
151
|
try {
|
|
163
|
-
await meter.record(cfg,
|
|
152
|
+
await meter.record(cfg,
|
|
153
|
+
{ usage: err.usage, model: err.model || cfg.model, ms: Date.now() - started, failed: true },
|
|
164
154
|
{ ticketId: opts.ticketId });
|
|
165
155
|
} catch (recErr) {
|
|
166
|
-
|
|
156
|
+
// A budget may throw anything, null included; logging it must not throw in its turn.
|
|
157
|
+
log.warn(`AI: usage of a failed call not recorded (${String(recErr?.message ?? recErr)})`);
|
|
167
158
|
}
|
|
168
159
|
}
|
|
169
160
|
throw err;
|