@aria-framework/ai 0.26.1 → 0.27.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md ADDED
@@ -0,0 +1,103 @@
1
+ # @aria-framework/ai changelog
2
+
3
+ This file starts at 0.26.1. Earlier versions are described in the git history only.
4
+
5
+ Each entry lists the **silent default changes** first: things that change what a caller's existing
6
+ code sends, or what it is billed for, without the caller changing anything.
7
+
8
+ ## 0.27.1
9
+
10
+ Fixes from an independent review of 0.26.1 and 0.27.0.
11
+
12
+ ### Silent default changes
13
+
14
+ - **A failed call is recorded under the model that ran.** `budget.record(cfg, { usage, model, ms,
15
+ failed: true })` now receives the model the server named in its reply (`err.model`). It used to
16
+ receive `cfg.model`, which is empty on lmx, so every failed lmx call was ledgered with no model. The
17
+ configured model is used only when the reply named none. A usage table keyed on model will see
18
+ failed lmx calls move from `''` to the engine's model.
19
+ - **Older Claude models: `'full'` never asks for more output than the model allows.** On the
20
+ pre-4.6 budget path, `max_tokens` was the caller's `maxTokens` plus a thinking budget of the same
21
+ size, about twice what was asked. That could pass the model's output limit (Haiku 4.5, Sonnet 4.5,
22
+ Opus 4.5, Sonnet 4 and 3.7 Sonnet allow 64000; Opus 4.1 and Opus 4 allow 32000), and the API
23
+ refused it. It is now capped at the limit. Near the limit the thinking budget shrinks first, down
24
+ to the API's minimum of 1024. For example, `maxTokens: 40000` on Haiku 4.5 sends
25
+ `max_tokens: 64000` with a 24000 budget. When `maxTokens` plus the budget is below the limit,
26
+ nothing changes.
27
+ - **Claude 3 models without extended thinking get no thinking.** On `'full'`, Claude 3.5 Haiku,
28
+ 3 Haiku, 3 Opus and 3.5 Sonnet were sent `thinking: enabled`, which they reject with a 400. They now
29
+ keep the pre-0.27 request, and `'full'` logs a warning instead of failing the call, as it does
30
+ for a model the table does not know. 3.7 Sonnet still thinks.
31
+ - **Opus 4 and Sonnet 4 under their `-0` aliases and dated ids** (`claude-opus-4-0`,
32
+ `claude-opus-4-20250514`, `claude-sonnet-4-0`, `claude-sonnet-4-20250514`) are now recognised.
33
+ 0.27.0 treated them as unknown: `'full'` sent no thinking and logged a warning. They now get the
34
+ budget like the other pre-4.6 models.
35
+
36
+ ### Fixes
37
+
38
+ - If a budget's `record` throws `null` (or anything without a `message`), logging it no longer
39
+ throws a `TypeError` that replaces the caller's real error.
40
+ - `AiError` takes an optional `model`, set by both adapters on every error that carries `usage`.
41
+ On lmx, when the reply names no model, it is the model the engine reports running; the same
42
+ fallback now applies to a successful result's `model`.
43
+ - **Anthropic: a `maxTokens` that is not a positive number is refused** (`AiError` kind `refused`,
44
+ before any request) on every Anthropic call. A fraction is rounded down. Before this, a
45
+ string `"2048"` on the budget path became `max_tokens: 64000`, and a negative value sent a
46
+ `budget_tokens` above `max_tokens`. `0`, `NaN` and a missing value still take the 1024 default.
47
+
48
+ ### Tests
49
+
50
+ - `test/anthropicRequests.js` (new) compares the request bodies with a fixture typed by hand from
51
+ the Messages API rules for 26 model ids, rather than with values read back from
52
+ `ANTHROPIC_FAMILIES`. It also checks the budget at 1024, 4096, 40000 and at each model's limit.
53
+ - `test/usageOnFailure.js` adds an lmx failure with an empty `cfg.model`, an Anthropic failure whose
54
+ reply names a different model, an empty completion with zero usage on both adapters, and a budget
55
+ that throws `null`.
56
+
57
+ ## 0.27.0
58
+
59
+ `complete({ reasoning: 'off' | 'full' })` is honoured by every provider. Before this, only lmx
60
+ honoured it. Leaving it out means `'off'`.
61
+
62
+ ### Silent default changes
63
+
64
+ These apply to callers that do not pass `reasoning`. To keep a model thinking, pass
65
+ `reasoning: 'full'`.
66
+
67
+ - **openai-compatible and lmstudio, qwen or gpt-oss models: a body flag is now sent on every call.**
68
+ qwen gets `chat_template_kwargs: { enable_thinking: false }` and gpt-oss gets
69
+ `reasoning_effort: 'low'`. Before this, these models used the server's default, which usually means
70
+ thinking on every call.
71
+ - **LM Studio (`provider: 'lmstudio'`), qwen models: `/no_think` is appended to the system prompt**
72
+ (or sent as the system prompt when there is none), and `reasoning_effort: 'none'` is added to the
73
+ body. LM Studio ignores `chat_template_kwargs`. A caller that compares or caches system prompts will
74
+ see the extra line. An LM Studio server configured as `openai-compatible` gets neither, and keeps
75
+ thinking.
76
+ - **Claude Opus 5.5 and Fable (and Mythos) get `output_config.effort: 'low'`.** These models cannot
77
+ turn thinking off, and Opus 5.5's own default effort is `medium`, so leaving `reasoning` out now
78
+ means less thinking than before.
79
+ - **Other Claude generations:** Opus 5 gets effort `'low'` (its own default is adaptive thinking at
80
+ `high`). Sonnet 5.5 gets `thinking: { type: 'between_tools' }` and Sonnet 5 gets
81
+ `thinking: { type: 'disabled' }`; before this, both thought by default. 4.6 and older are
82
+ unchanged unless `'full'` is passed.
83
+ - **`temperature` is no longer sent to Claude models that reject it** (Opus 4.7 and later, Sonnet 5
84
+ and later, Fable). Before this, those requests failed with a 400. It is also left out with `'full'`
85
+ on the pre-4.6 budget path, where extended thinking requires the default.
86
+ - The client no longer strips `reasoning` with a warning saying only lmx honours it.
87
+
88
+ ## 0.26.1
89
+
90
+ ### Silent default changes
91
+
92
+ - **A failed call that spent tokens is recorded through `budget.record`.** When the model answered
93
+ but the answer was unusable (it thought until the limit, or a structured reply was cut off),
94
+ `budget.record(cfg, { usage, model, ms, failed: true }, meta)` is called before the error is
95
+ thrown. A failure that spent nothing (refused, unreachable, timed out, or a zero usage total)
96
+ records nothing. **A caller that metered `err.usage` itself must stop, or the call is counted
97
+ twice.** A `record` that throws is logged and never replaces the original error.
98
+
99
+ ### Fixes
100
+
101
+ - Both adapters put normalised usage (`{ prompt, completion, total }`) on every error raised after
102
+ the reply was read.
103
+ - A 200 whose JSON body is `null` is an `AiError` of kind `bad_response`. It used to be a `TypeError`.
package/README.md CHANGED
@@ -35,7 +35,7 @@ const ai = createAiClient({
35
35
  const r = await ai.complete({
36
36
  system, messages, maxTokens: 400,
37
37
  schema, // optional: constrained JSON
38
- // reasoning: 'full', // optional, lmx only: let the model think fully (see Reasoning below)
38
+ // reasoning: 'full', // optional, every provider: 'off' (default) or 'full' (see Reasoning below)
39
39
  });
40
40
  // r.text, r.json (when schema), r.model, r.usage.total, r.ms
41
41
  ```
@@ -48,7 +48,10 @@ an admin screen and never includes a URL or key.
48
48
  was unusable (it thought until the limit, or a structured reply was cut off), the error carries
49
49
  `usage` (`{ prompt, completion, total }`) and `budget.record(cfg, { usage, model, ms, failed: true }, meta)`
50
50
  is called before the error is thrown. A failure that spent nothing (refused, unreachable, timeout)
51
- records nothing. If `record` throws, the original error still reaches the caller.
51
+ records nothing. If `record` throws, the original error still reaches the caller. The error also
52
+ carries `model`, the model the server says answered, and `record` receives that (0.27.1); it falls
53
+ back to the configured model only when the reply named none. On lmx the configured model is empty.
54
+ A caller that metered `err.usage` itself before 0.26.1 must stop, or the call is counted twice.
52
55
 
53
56
  `providerStore`, `usageStore`, `speedStore` and `lmxStore` are optional db-worker-backed
54
57
  stores. Each exports a `schemaFor(dialect)`, so the app's migration can be checked against it.
@@ -69,7 +72,7 @@ to write yourself:
69
72
  | Engine URLs come from the document | Resolved on every call, and never stored |
70
73
  | A status outage is not an inference outage | Keeps routing on the last good document for 3 minutes (`staleMs`), logging loudly |
71
74
  | Pin the certificate; never disable verification | Uses an undici `Agent({ connect: { ca } })` per stack. Never `NODE_EXTRA_CA_CERTS`, never `rejectUnauthorized:false` |
72
- | Choose the reasoning flag from the model | By default thinking is kept to a minimum: `qwen` → `chat_template_kwargs.enable_thinking=false`, `gpt-oss` → `reasoning_effort:'low'`, anything else gets neither (with a warning). `reasoning: 'full'` on the call (since 0.26.0) sends `enable_thinking=true` / `reasoning_effort:'high'` instead. The flag always overrides a raw `extra`; only the named option changes it. See [Reasoning](#reasoning-reasoning-full) |
75
+ | Choose the reasoning flag from the model | By default thinking is kept to a minimum: `qwen` → `chat_template_kwargs.enable_thinking=false`, `gpt-oss` → `reasoning_effort:'low'`, anything else gets neither (with a warning). `reasoning: 'full'` on the call (since 0.26.0) sends `enable_thinking=true` / `reasoning_effort:'high'` instead. The flag always overrides a raw `extra`; only the named option changes it. Since 0.27.0 the same rule applies to every provider. See [Reasoning](#reasoning-reasoning-off--full) |
73
76
  | Size requests against `maxInputTokens` (per slot) | `contextTokens` is taken from the engine, not from config |
74
77
  | Identity is `(instance, id)`, falling back to `name` | `lmx.engineId` is matched first; `rekeyPlan()` migrates rows when ids are minted or engines renamed |
75
78
  | A `429` is about the key, not the engine | Fails fast (since 0.25.0): throws `rate_limit` with `lmxSkip: 'lmx_throttled'` and `retryAfterMs`. The client never waits; whether to wait or move on is the caller's decision |
@@ -176,10 +179,11 @@ instead:
176
179
  Failover across several engines for one job is the app's job; lmx provides none. List
177
180
  endpoints in preference order and take the first one that serves.
178
181
 
179
- ### Reasoning (`reasoning: 'full'`)
182
+ ### Reasoning (`reasoning: 'off' | 'full'`)
180
183
 
181
- By default the adapter makes the model think as little as its family allows, so a quick answer
182
- stays quick. Pass `reasoning: 'full'` on a `complete()` call to let the model think fully:
184
+ How much a model may think is decided by the **app, per call**, and every provider honours it (since
185
+ 0.27.0; before that only lmx did). Leave it out, or pass `reasoning: 'off'`, and the model thinks as
186
+ little as its family allows, so a quick answer stays quick. Pass `reasoning: 'full'` to let it think:
183
187
 
184
188
  ```js
185
189
  const r = await ai.complete({
@@ -204,23 +208,60 @@ a `bad_response` AiError saying it ran out of room (and, when the server says so
204
208
  tokens went on internal reasoning), but the work is lost either way. Use **`maxTokens` ≥ 4096** and a
205
209
  **`timeoutMs` ≥ 120 s** (120000) on the endpoint, or the deadline cuts the answer off first.
206
210
 
207
- **What is sent, per model family** (read from the model the engine is running on each call):
211
+ **lmx, openai-compatible and lmstudio: what is added to the request body, per model family.** The
212
+ family is read from the model the engine is running (lmx) or the endpoint's model name (others):
208
213
 
209
- | Family | Default (no `reasoning`) | `reasoning: 'full'` |
214
+ | Family | `'off'` (the default) | `reasoning: 'full'` |
210
215
  |---|---|---|
211
216
  | qwen | `{"chat_template_kwargs":{"enable_thinking":false}}` | `{"chat_template_kwargs":{"enable_thinking":true}}` |
212
217
  | gpt-oss | `{"reasoning_effort":"low"}` | `{"reasoning_effort":"high"}` |
213
- | anything else | nothing (warning logged) | nothing (warning logged that `'full'` could not be honoured) |
218
+ | anything else | nothing (lmx logs a warning) | nothing (a warning that `'full'` could not be honoured) |
219
+
220
+ **LM Studio ignores `chat_template_kwargs`** (measured live: Qwen3 thought just as long with
221
+ `enable_thinking: false`). So on **`provider: 'lmstudio'`** a qwen model gets two more switches. One is
222
+ `{"reasoning_effort":"none"}` for `'off'` and `{"reasoning_effort":"high"}` for `'full'`: verified live, but
223
+ LM Studio documents `reasoning_effort` only for gpt-oss. The other is Qwen's own documented prompt switch,
224
+ `/no_think` for `'off'` and `/think` for `'full'`, appended to the system prompt (or sent as one when there
225
+ is none): documented on LM Studio's Qwen3 model pages, but only hybrid Qwen3 reads it. Both are sent, so
226
+ if a future LM Studio drops one the other still holds. Configure an LM Studio server as `lmstudio`, not
227
+ `openai-compatible`; as `openai-compatible` it keeps thinking on every call. Plain `openai-compatible`
228
+ gets neither, because vLLM validates `reasoning_effort` and may refuse `'none'`.
229
+
230
+ **anthropic: what is sent, per Claude generation.** The Messages API differs by generation: Opus 5.5
231
+ and Fable cannot turn thinking off (effort is the only lever), Sonnet 5.5 turns it off with
232
+ `between_tools`, and several current models refuse `temperature`, so none is sent to them.
233
+
234
+ | Models | `'off'` (the default) | `reasoning: 'full'` | temperature |
235
+ |---|---|---|---|
236
+ | claude-fable / claude-opus-5-5 | `{"output_config":{"effort":"low"}}` | `{"output_config":{"effort":"high"}}` | not sent |
237
+ | claude-opus-5 | `{"output_config":{"effort":"low"}}` | `{"thinking":{"type":"adaptive"},"output_config":{"effort":"high"}}` | not sent |
238
+ | claude-sonnet-5-5 | `{"thinking":{"type":"between_tools"}}` | `{"thinking":{"type":"adaptive"},"output_config":{"effort":"high"}}` | not sent |
239
+ | claude-sonnet-5 | `{"thinking":{"type":"disabled"}}` | `{"thinking":{"type":"adaptive"},"output_config":{"effort":"high"}}` | not sent |
240
+ | claude-opus-4-7 / 4-8 | nothing | `{"thinking":{"type":"adaptive"},"output_config":{"effort":"high"}}` | not sent |
241
+ | claude-*-4-6 | nothing | `{"thinking":{"type":"adaptive"},"output_config":{"effort":"high"}}` | sent |
242
+ | claude-*-4-5 | nothing | thinking enabled with a budget of `maxTokens` (at least 1024), added on top of `max_tokens`, which never exceeds 64000 | sent; not sent with `'full'` |
243
+ | claude-opus-4-1 / claude-opus-4 | nothing | thinking enabled with a budget of `maxTokens` (at least 1024), added on top of `max_tokens`, which never exceeds 32000 | sent; not sent with `'full'` |
244
+ | claude-sonnet-4 | nothing | thinking enabled with a budget of `maxTokens` (at least 1024), added on top of `max_tokens`, which never exceeds 64000 | sent; not sent with `'full'` |
245
+ | claude-3-7-sonnet | nothing | thinking enabled with a budget of `maxTokens` (at least 1024), added on top of `max_tokens`, which never exceeds 64000 | sent; not sent with `'full'` |
246
+ | claude-3 without extended thinking | nothing | nothing (a warning that `'full'` could not be honoured) | sent |
247
+ | anything else | nothing | nothing (a warning that `'full'` could not be honoured) | sent |
248
+
249
+ `effort` is merged into any `output_config` the structured-output schema already put there.
250
+
251
+ On the budget rows `max_tokens` counts the thinking as well as the answer, so it is capped at the
252
+ model's output limit (0.27.1). Near the limit the thinking budget shrinks first, down to the API's
253
+ minimum of 1024, and only then the answer's room: `maxTokens: 40000` on Haiku 4.5 sends
254
+ `max_tokens: 64000` with a 24000 budget. Claude 3 models other than 3.7 Sonnet have no extended
255
+ thinking, so `'full'` sends nothing to them and logs a warning; it does not fail the call.
214
256
 
215
257
  Rules:
216
258
 
217
- - `'full'` is the only value. Leaving the option out (or `null`) means the default. Any other value,
218
- for example `'high'`, `true` or `''`, is refused with `AiError` kind `refused` before any
259
+ - The values are `'off'` and `'full'`. Leaving the option out (or `null`) means `'off'`. Any other
260
+ value, for example `'high'`, `true` or `''`, is refused with `AiError` kind `refused` before any
219
261
  request is made, whichever provider the endpoint uses.
220
262
  - The flag is merged **after** `extra`, so a raw `extra: { chat_template_kwargs: … }` or
221
- `extra: { reasoning_effort: … }` cannot change it. Only `reasoning` can.
222
- - Only `lmx` honours it. On `openai-compatible`, `lmstudio` or `anthropic` the client logs a
223
- warning, removes the option, and nothing reaches that provider's request body.
263
+ `extra: { reasoning_effort: … }` cannot change it. Only `reasoning` can. The option itself is never
264
+ copied into a request body.
224
265
  - `REASONING_MODES` (package root) lists the accepted values for an app that validates its own
225
266
  settings.
226
267
 
@@ -1,138 +1,138 @@
1
- /**
2
- * The fold on the Inference list, and prefilling the supervisor form from a panel.
3
- *
4
- * WHY ANYTHING FOLDS. A stable stack is four engines, three credentials, a certificate fingerprint
5
- * and a verification report — worth having, not worth reading every time you open this page to do
6
- * something else. Folded, a panel is one line that still carries the whole verdict: status, the
7
- * count of what depends on it, all three credentials and when it was last checked. Folding hides
8
- * detail; it must never hide the reason you would have opened it.
9
- *
10
- * A PANEL WITH A PROBLEM IGNORES WHAT YOU REMEMBERED. `data-panel-attention` is stamped by the
11
- * server on anything with a missing engine, an expiring certificate, a failed check or a nearly
12
- * spent cap. Those open and stay open. Attention beats tidiness — the alternative is a page that
13
- * quietly honours a fold you chose last week and hides the thing that broke yesterday.
14
- *
15
- * THE PREFERENCE IS PER BROWSER AND DISPOSABLE. It is a convenience about how a page looks to one
16
- * person, so localStorage is the right home and losing it costs nothing. Every access is guarded:
17
- * a browser set to block site data throws on read, and a settings page must not break because
18
- * somebody tightened their privacy settings.
19
- */
20
- (function () {
21
- 'use strict';
22
-
23
- var KEY = 's101.ai.panels';
24
-
25
- function readPrefs() {
26
- try {
27
- return JSON.parse(window.localStorage.getItem(KEY) || '{}') || {};
28
- } catch (e) {
29
- return {};
30
- }
31
- }
32
-
33
- function writePref(id, open) {
34
- try {
35
- var prefs = readPrefs();
36
- prefs[id] = !!open;
37
- window.localStorage.setItem(KEY, JSON.stringify(prefs));
38
- } catch (e) { /* private window, or site data blocked — the fold still works for this visit */ }
39
- }
40
-
41
- function bodyOf(panel) {
42
- var btn = panel.querySelector('[data-panel-toggle]');
43
- if (!btn) return null;
44
- return document.getElementById(btn.getAttribute('aria-controls'));
45
- }
46
-
47
- function setOpen(panel, open, remember) {
48
- var btn = panel.querySelector('[data-panel-toggle]');
49
- var body = bodyOf(panel);
50
- if (!btn || !body) return;
51
- // `hidden`, not a style: the server renders the closed state the same way, so a panel does not
52
- // flicker open on load before this script runs.
53
- body.hidden = !open;
54
- btn.setAttribute('aria-expanded', open ? 'true' : 'false');
55
- panel.classList.toggle('panel-open', open);
56
- if (remember) writePref(panel.getAttribute('data-panel'), open);
57
- }
58
-
59
- function restore() {
60
- var prefs = readPrefs();
61
- var panels = document.querySelectorAll('[data-panel]');
62
- for (var i = 0; i < panels.length; i += 1) {
63
- var panel = panels[i];
64
- var id = panel.getAttribute('data-panel');
65
- // The server already opened this one and means it. Do not consult the preference at all —
66
- // reading it and then ignoring it is the same thing, but invites somebody to "fix" it later.
67
- if (panel.hasAttribute('data-panel-attention')) {
68
- setOpen(panel, true, false);
69
- continue;
70
- }
71
- if (Object.prototype.hasOwnProperty.call(prefs, id)) setOpen(panel, !!prefs[id], false);
72
- }
73
- }
74
-
75
- document.addEventListener('click', function (ev) {
76
- var toggle = ev.target.closest('[data-panel-toggle]');
77
- if (toggle) {
78
- var panel = toggle.closest('[data-panel]');
79
- if (!panel) return;
80
- var body = bodyOf(panel);
81
- setOpen(panel, !!(body && body.hidden), true);
82
- return;
83
- }
84
-
85
- // ── EDITING A STACK HAPPENS IN THE PANEL ────────────────────────────────────────────────────
86
- //
87
- // It used to happen in a form at the foot of the page, and Edit scrolled you down to it — away
88
- // from the stack you were reading, into a form that also served as "Add a supervisor" and said
89
- // nothing about which stack it had been filled with. Nothing needs prefilling now: every panel
90
- // renders its own form, from the server, with its own values already in it.
91
- var open = ev.target.closest('[data-stack-edit-open]');
92
- if (open) { stackMode(open.closest('[data-stack]'), true); return; }
93
-
94
- var cancel = ev.target.closest('[data-stack-cancel]');
95
- if (cancel) {
96
- var panel2 = cancel.closest('[data-stack]');
97
- // A NEW stack has no reading half to go back to, so cancelling it puts the blank panel away.
98
- if (panel2 && panel2.hasAttribute('data-stack-new')) { panel2.hidden = true; return; }
99
- // Otherwise reload rather than restore: a Cancel that left half-typed values behind, ready to
100
- // be posted by the next Save, would be worse than the extra request.
101
- window.location.reload();
102
- return;
103
- }
104
-
105
- // "Add LMX" reveals the blank panel that is already on the page, at the top of the list.
106
- var add = ev.target.closest('[data-stack-add]');
107
- if (add) {
108
- var blank = document.querySelector('[data-stack-new]');
109
- if (!blank) return;
110
- blank.hidden = false;
111
- blank.scrollIntoView({ block: 'nearest' });
112
- var id = blank.querySelector('[name="id"]');
113
- if (id) id.focus();
114
- }
115
- });
116
-
117
- /** Swap one stack panel between reading and editing. */
118
- function stackMode(panel, editing) {
119
- if (!panel) return;
120
- var view = panel.querySelector('[data-stack-view]');
121
- var form = panel.querySelector('[data-stack-edit]');
122
- if (!view || !form) return;
123
- view.hidden = editing;
124
- form.hidden = !editing;
125
- if (editing) {
126
- // The id is readonly on an existing stack, so focus the first field somebody can actually
127
- // change rather than one that will not accept a keystroke.
128
- var first = form.querySelector('[name="label"]') || form.querySelector('[name="status_url"]');
129
- if (first) first.focus();
130
- }
131
- }
132
-
133
- if (document.readyState === 'loading') {
134
- document.addEventListener('DOMContentLoaded', restore);
135
- } else {
136
- restore();
137
- }
138
- })();
1
+ /**
2
+ * The fold on the Inference list, and prefilling the supervisor form from a panel.
3
+ *
4
+ * WHY ANYTHING FOLDS. A stable stack is four engines, three credentials, a certificate fingerprint
5
+ * and a verification report — worth having, not worth reading every time you open this page to do
6
+ * something else. Folded, a panel is one line that still carries the whole verdict: status, the
7
+ * count of what depends on it, all three credentials and when it was last checked. Folding hides
8
+ * detail; it must never hide the reason you would have opened it.
9
+ *
10
+ * A PANEL WITH A PROBLEM IGNORES WHAT YOU REMEMBERED. `data-panel-attention` is stamped by the
11
+ * server on anything with a missing engine, an expiring certificate, a failed check or a nearly
12
+ * spent cap. Those open and stay open. Attention beats tidiness — the alternative is a page that
13
+ * quietly honours a fold you chose last week and hides the thing that broke yesterday.
14
+ *
15
+ * THE PREFERENCE IS PER BROWSER AND DISPOSABLE. It is a convenience about how a page looks to one
16
+ * person, so localStorage is the right home and losing it costs nothing. Every access is guarded:
17
+ * a browser set to block site data throws on read, and a settings page must not break because
18
+ * somebody tightened their privacy settings.
19
+ */
20
+ (function () {
21
+ 'use strict';
22
+
23
+ var KEY = 's101.ai.panels';
24
+
25
+ function readPrefs() {
26
+ try {
27
+ return JSON.parse(window.localStorage.getItem(KEY) || '{}') || {};
28
+ } catch (e) {
29
+ return {};
30
+ }
31
+ }
32
+
33
+ function writePref(id, open) {
34
+ try {
35
+ var prefs = readPrefs();
36
+ prefs[id] = !!open;
37
+ window.localStorage.setItem(KEY, JSON.stringify(prefs));
38
+ } catch (e) { /* private window, or site data blocked — the fold still works for this visit */ }
39
+ }
40
+
41
+ function bodyOf(panel) {
42
+ var btn = panel.querySelector('[data-panel-toggle]');
43
+ if (!btn) return null;
44
+ return document.getElementById(btn.getAttribute('aria-controls'));
45
+ }
46
+
47
+ function setOpen(panel, open, remember) {
48
+ var btn = panel.querySelector('[data-panel-toggle]');
49
+ var body = bodyOf(panel);
50
+ if (!btn || !body) return;
51
+ // `hidden`, not a style: the server renders the closed state the same way, so a panel does not
52
+ // flicker open on load before this script runs.
53
+ body.hidden = !open;
54
+ btn.setAttribute('aria-expanded', open ? 'true' : 'false');
55
+ panel.classList.toggle('panel-open', open);
56
+ if (remember) writePref(panel.getAttribute('data-panel'), open);
57
+ }
58
+
59
+ function restore() {
60
+ var prefs = readPrefs();
61
+ var panels = document.querySelectorAll('[data-panel]');
62
+ for (var i = 0; i < panels.length; i += 1) {
63
+ var panel = panels[i];
64
+ var id = panel.getAttribute('data-panel');
65
+ // The server already opened this one and means it. Do not consult the preference at all —
66
+ // reading it and then ignoring it is the same thing, but invites somebody to "fix" it later.
67
+ if (panel.hasAttribute('data-panel-attention')) {
68
+ setOpen(panel, true, false);
69
+ continue;
70
+ }
71
+ if (Object.prototype.hasOwnProperty.call(prefs, id)) setOpen(panel, !!prefs[id], false);
72
+ }
73
+ }
74
+
75
+ document.addEventListener('click', function (ev) {
76
+ var toggle = ev.target.closest('[data-panel-toggle]');
77
+ if (toggle) {
78
+ var panel = toggle.closest('[data-panel]');
79
+ if (!panel) return;
80
+ var body = bodyOf(panel);
81
+ setOpen(panel, !!(body && body.hidden), true);
82
+ return;
83
+ }
84
+
85
+ // ── EDITING A STACK HAPPENS IN THE PANEL ────────────────────────────────────────────────────
86
+ //
87
+ // It used to happen in a form at the foot of the page, and Edit scrolled you down to it — away
88
+ // from the stack you were reading, into a form that also served as "Add a supervisor" and said
89
+ // nothing about which stack it had been filled with. Nothing needs prefilling now: every panel
90
+ // renders its own form, from the server, with its own values already in it.
91
+ var open = ev.target.closest('[data-stack-edit-open]');
92
+ if (open) { stackMode(open.closest('[data-stack]'), true); return; }
93
+
94
+ var cancel = ev.target.closest('[data-stack-cancel]');
95
+ if (cancel) {
96
+ var panel2 = cancel.closest('[data-stack]');
97
+ // A NEW stack has no reading half to go back to, so cancelling it puts the blank panel away.
98
+ if (panel2 && panel2.hasAttribute('data-stack-new')) { panel2.hidden = true; return; }
99
+ // Otherwise reload rather than restore: a Cancel that left half-typed values behind, ready to
100
+ // be posted by the next Save, would be worse than the extra request.
101
+ window.location.reload();
102
+ return;
103
+ }
104
+
105
+ // "Add LMX" reveals the blank panel that is already on the page, at the top of the list.
106
+ var add = ev.target.closest('[data-stack-add]');
107
+ if (add) {
108
+ var blank = document.querySelector('[data-stack-new]');
109
+ if (!blank) return;
110
+ blank.hidden = false;
111
+ blank.scrollIntoView({ block: 'nearest' });
112
+ var id = blank.querySelector('[name="id"]');
113
+ if (id) id.focus();
114
+ }
115
+ });
116
+
117
+ /** Swap one stack panel between reading and editing. */
118
+ function stackMode(panel, editing) {
119
+ if (!panel) return;
120
+ var view = panel.querySelector('[data-stack-view]');
121
+ var form = panel.querySelector('[data-stack-edit]');
122
+ if (!view || !form) return;
123
+ view.hidden = editing;
124
+ form.hidden = !editing;
125
+ if (editing) {
126
+ // The id is readonly on an existing stack, so focus the first field somebody can actually
127
+ // change rather than one that will not accept a keystroke.
128
+ var first = form.querySelector('[name="label"]') || form.querySelector('[name="status_url"]');
129
+ if (first) first.focus();
130
+ }
131
+ }
132
+
133
+ if (document.readyState === 'loading') {
134
+ document.addEventListener('DOMContentLoaded', restore);
135
+ } else {
136
+ restore();
137
+ }
138
+ })();
package/error.js CHANGED
@@ -18,7 +18,7 @@ class AiError extends Error {
18
18
  /**
19
19
  * @param {AiErrorKind} kind
20
20
  * @param {string} message written for the admin screen, not the log
21
- * @param {{cause?: Error, status?: number, usage?: object}} [opts]
21
+ * @param {{cause?: Error, status?: number, usage?: object, model?: string}} [opts]
22
22
  */
23
23
  constructor(kind, message, opts = {}) {
24
24
  super(message);
@@ -30,6 +30,10 @@ class AiError extends Error {
30
30
  // measuring throughput should not have to treat that as zero work. Optional everywhere: most
31
31
  // failures happen before a single token is spent.
32
32
  if (opts.usage) this.usage = opts.usage;
33
+ // ...AND WHICH MODEL SPENT IT (0.27.1), as the server named it in its reply. The client meters a
34
+ // failed call under this, not cfg.model: on lmx cfg.model is empty (the engine decides what
35
+ // runs), and an alias on any provider resolves to an id other than the one configured.
36
+ if (opts.model) this.model = opts.model;
33
37
  if (opts.cause) this.cause = opts.cause;
34
38
  }
35
39
 
package/index.js CHANGED
@@ -26,7 +26,7 @@ const facts = require('./facts');
26
26
  const { AiError, fromFetchFailure, redact } = require('./error');
27
27
  const { polish } = require('./polish');
28
28
  const { generate } = require('./generate');
29
- const { assertReasoning, REASONING_MODES } = require('./providers/lmx');
29
+ const { assertReasoning, REASONING_MODES } = require('./reasoning');
30
30
 
31
31
  const PROVIDERS = {
32
32
  // 'lmstudio' and 'openai-compatible' are the SAME adapter with different defaults — a kindness to
@@ -52,9 +52,6 @@ const DEFAULTS = {
52
52
  const RETRY_AFTER_MS = 400;
53
53
  const RETRY_ONLY_IF_FAILED_WITHIN_MS = 5000;
54
54
 
55
- /** Providers already warned that they ignore `reasoning` — once per provider per process. */
56
- const warnedReasoning = new Set();
57
-
58
55
  const NOOP_LOGGER = { info() {}, warn() {}, error() {} };
59
56
  const NOOP_BUDGET = { async assertWithinBudget() {}, async record() {} };
60
57
 
@@ -106,7 +103,7 @@ function createAiClient(deps = {}) {
106
103
  * Ask the configured model for something.
107
104
  * @param {{system?:string, messages:Array, maxTokens?:number, temperature?:number,
108
105
  * schema?:object, signal?:AbortSignal, ticketId?:*, skipBudget?:boolean,
109
- * reasoning?:'full'}} opts
106
+ * reasoning?:'off'|'full'}} opts
110
107
  * @param {object} [cfgOverride] the resolved config, when the caller already has it
111
108
  */
112
109
  async function complete(opts, cfgOverride) {
@@ -134,19 +131,9 @@ function createAiClient(deps = {}) {
134
131
  // by whichever caller forgets. `skipBudget` is for the admin's Test connection; it still records.
135
132
  if (!opts.skipBudget) await meter.assertWithinBudget(cfg, { ticketId: opts.ticketId });
136
133
 
137
- // ONLY lmx HONOURS `reasoning`. Elsewhere it is warned about and removed, so no other adapter's
138
- // request body can ever carry it.
139
- let callOpts = opts;
140
- if (opts.reasoning != null && cfg.provider !== 'lmx') {
141
- if (!warnedReasoning.has(cfg.provider)) {
142
- warnedReasoning.add(cfg.provider);
143
- log.warn(`AI: reasoning '${opts.reasoning}' is honoured only by lmx engines — `
144
- + `${cfg.label || cfg.provider} sends no reasoning flag and runs on its own default `
145
- + '(warned once per provider).');
146
- }
147
- const { reasoning: _ignored, ...rest } = opts;
148
- callOpts = rest;
149
- }
134
+ // EVERY PROVIDER HONOURS `reasoning` (0.27.0): each adapter translates it for its model family
135
+ // (see reasoning.js). It used to be lmx-only, warned about and stripped everywhere else.
136
+ const callOpts = opts;
150
137
 
151
138
  const started = Date.now();
152
139
  let result;
@@ -158,12 +145,16 @@ function createAiClient(deps = {}) {
158
145
  // connection test on such a model run outside the ceiling. Adapters put normalised usage on
159
146
  // the error; a failure that spent nothing (refused, unreachable) carries none and is not counted.
160
147
  // The ledger failing must never replace the error the caller needs to see.
148
+ // THE MODEL THAT RAN (0.27.1), as the adapter read it off the reply — cfg.model only when the
149
+ // reply named none. On lmx cfg.model is '' and every failed call was ledgered under no model.
161
150
  if (err && err.usage && err.usage.total > 0) {
162
151
  try {
163
- await meter.record(cfg, { usage: err.usage, model: cfg.model, ms: Date.now() - started, failed: true },
152
+ await meter.record(cfg,
153
+ { usage: err.usage, model: err.model || cfg.model, ms: Date.now() - started, failed: true },
164
154
  { ticketId: opts.ticketId });
165
155
  } catch (recErr) {
166
- log.warn(`AI: usage of a failed call not recorded (${recErr.message})`);
156
+ // A budget may throw anything, null included; logging it must not throw in its turn.
157
+ log.warn(`AI: usage of a failed call not recorded (${String(recErr?.message ?? recErr)})`);
167
158
  }
168
159
  }
169
160
  throw err;