@aria-framework/ai 0.26.1 → 0.27.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -35,7 +35,7 @@ const ai = createAiClient({
35
35
  const r = await ai.complete({
36
36
  system, messages, maxTokens: 400,
37
37
  schema, // optional: constrained JSON
38
- // reasoning: 'full', // optional, lmx only: let the model think fully (see Reasoning below)
38
+ // reasoning: 'full', // optional, every provider: 'off' (default) or 'full' (see Reasoning below)
39
39
  });
40
40
  // r.text, r.json (when schema), r.model, r.usage.total, r.ms
41
41
  ```
@@ -69,7 +69,7 @@ to write yourself:
69
69
  | Engine URLs come from the document | Resolved on every call, and never stored |
70
70
  | A status outage is not an inference outage | Keeps routing on the last good document for 3 minutes (`staleMs`), logging loudly |
71
71
  | Pin the certificate; never disable verification | Uses an undici `Agent({ connect: { ca } })` per stack. Never `NODE_EXTRA_CA_CERTS`, never `rejectUnauthorized:false` |
72
- | Choose the reasoning flag from the model | By default thinking is kept to a minimum: `qwen` → `chat_template_kwargs.enable_thinking=false`, `gpt-oss` → `reasoning_effort:'low'`, anything else gets neither (with a warning). `reasoning: 'full'` on the call (since 0.26.0) sends `enable_thinking=true` / `reasoning_effort:'high'` instead. The flag always overrides a raw `extra`; only the named option changes it. See [Reasoning](#reasoning-reasoning-full) |
72
+ | Choose the reasoning flag from the model | By default thinking is kept to a minimum: `qwen` → `chat_template_kwargs.enable_thinking=false`, `gpt-oss` → `reasoning_effort:'low'`, anything else gets neither (with a warning). `reasoning: 'full'` on the call (since 0.26.0) sends `enable_thinking=true` / `reasoning_effort:'high'` instead. The flag always overrides a raw `extra`; only the named option changes it. Since 0.27.0 the same rule applies to every provider. See [Reasoning](#reasoning-reasoning-off--full) |
73
73
  | Size requests against `maxInputTokens` (per slot) | `contextTokens` is taken from the engine, not from config |
74
74
  | Identity is `(instance, id)`, falling back to `name` | `lmx.engineId` is matched first; `rekeyPlan()` migrates rows when ids are minted or engines renamed |
75
75
  | A `429` is about the key, not the engine | Fails fast (since 0.25.0): throws `rate_limit` with `lmxSkip: 'lmx_throttled'` and `retryAfterMs`. The client never waits; whether to wait or move on is the caller's decision |
@@ -176,10 +176,11 @@ instead:
176
176
  Failover across several engines for one job is the app's job; lmx provides none. List
177
177
  endpoints in preference order and take the first one that serves.
178
178
 
179
- ### Reasoning (`reasoning: 'full'`)
179
+ ### Reasoning (`reasoning: 'off' | 'full'`)
180
180
 
181
- By default the adapter makes the model think as little as its family allows, so a quick answer
182
- stays quick. Pass `reasoning: 'full'` on a `complete()` call to let the model think fully:
181
+ How much a model may think is decided by the **app, per call**, and every provider honours it (since
182
+ 0.27.0; before that only lmx did). Leave it out, or pass `reasoning: 'off'`, and the model thinks as
183
+ little as its family allows, so a quick answer stays quick. Pass `reasoning: 'full'` to let it think:
183
184
 
184
185
  ```js
185
186
  const r = await ai.complete({
@@ -204,23 +205,50 @@ a `bad_response` AiError saying it ran out of room (and, when the server says so
204
205
  tokens went on internal reasoning), but the work is lost either way. Use **`maxTokens` ≥ 4096** and a
205
206
  **`timeoutMs` ≥ 120 s** (120000) on the endpoint, or the deadline cuts the answer off first.
206
207
 
207
- **What is sent, per model family** (read from the model the engine is running on each call):
208
+ **lmx, openai-compatible and lmstudio: what is added to the request body, per model family.** The
209
+ family is read from the model the engine is running (lmx) or the endpoint's model name (others):
208
210
 
209
- | Family | Default (no `reasoning`) | `reasoning: 'full'` |
211
+ | Family | `'off'` (the default) | `reasoning: 'full'` |
210
212
  |---|---|---|
211
213
  | qwen | `{"chat_template_kwargs":{"enable_thinking":false}}` | `{"chat_template_kwargs":{"enable_thinking":true}}` |
212
214
  | gpt-oss | `{"reasoning_effort":"low"}` | `{"reasoning_effort":"high"}` |
213
- | anything else | nothing (warning logged) | nothing (warning logged that `'full'` could not be honoured) |
215
+ | anything else | nothing (lmx logs a warning) | nothing (a warning that `'full'` could not be honoured) |
216
+
217
+ **LM Studio ignores `chat_template_kwargs`** (measured live: Qwen3 thought just as long with
218
+ `enable_thinking: false`). So on **`provider: 'lmstudio'`** a qwen model gets two more switches. One is
219
+ `{"reasoning_effort":"none"}` for `'off'` and `{"reasoning_effort":"high"}` for `'full'`: verified live, but
220
+ LM Studio documents `reasoning_effort` only for gpt-oss. The other is Qwen's own documented prompt switch,
221
+ `/no_think` for `'off'` and `/think` for `'full'`, appended to the system prompt (or sent as one when there
222
+ is none): documented on LM Studio's Qwen3 model pages, but only hybrid Qwen3 reads it. Both are sent, so
223
+ if a future LM Studio drops one the other still holds. Configure an LM Studio server as `lmstudio`, not
224
+ `openai-compatible`; as `openai-compatible` it keeps thinking on every call. Plain `openai-compatible`
225
+ gets neither, because vLLM validates `reasoning_effort` and may refuse `'none'`.
226
+
227
+ **anthropic: what is sent, per Claude generation.** The Messages API differs by generation: Opus 5.5
228
+ and Fable cannot turn thinking off (effort is the only lever), Sonnet 5.5 turns it off with
229
+ `between_tools`, and several current models refuse `temperature`, so none is sent to them.
230
+
231
+ | Models | `'off'` (the default) | `reasoning: 'full'` | temperature |
232
+ |---|---|---|---|
233
+ | claude-fable / claude-opus-5-5 | `{"output_config":{"effort":"low"}}` | `{"output_config":{"effort":"high"}}` | not sent |
234
+ | claude-opus-5 | `{"output_config":{"effort":"low"}}` | `{"thinking":{"type":"adaptive"},"output_config":{"effort":"high"}}` | not sent |
235
+ | claude-sonnet-5-5 | `{"thinking":{"type":"between_tools"}}` | `{"thinking":{"type":"adaptive"},"output_config":{"effort":"high"}}` | not sent |
236
+ | claude-sonnet-5 | `{"thinking":{"type":"disabled"}}` | `{"thinking":{"type":"adaptive"},"output_config":{"effort":"high"}}` | not sent |
237
+ | claude-opus-4-7 / 4-8 | nothing | `{"thinking":{"type":"adaptive"},"output_config":{"effort":"high"}}` | not sent |
238
+ | claude-*-4-6 | nothing | `{"thinking":{"type":"adaptive"},"output_config":{"effort":"high"}}` | sent |
239
+ | claude-*-4-5 and older | nothing | thinking enabled with a budget of `maxTokens` (at least 1024), added on top of `max_tokens` | sent; not sent with `'full'` |
240
+ | anything else | nothing | nothing (a warning that `'full'` could not be honoured) | sent |
241
+
242
+ `effort` is merged into any `output_config` the structured-output schema already put there.
214
243
 
215
244
  Rules:
216
245
 
217
- - `'full'` is the only value. Leaving the option out (or `null`) means the default. Any other value,
218
- for example `'high'`, `true` or `''`, is refused with `AiError` kind `refused` before any
246
+ - The values are `'off'` and `'full'`. Leaving the option out (or `null`) means `'off'`. Any other
247
+ value, for example `'high'`, `true` or `''`, is refused with `AiError` kind `refused` before any
219
248
  request is made, whichever provider the endpoint uses.
220
249
  - The flag is merged **after** `extra`, so a raw `extra: { chat_template_kwargs: … }` or
221
- `extra: { reasoning_effort: … }` cannot change it. Only `reasoning` can.
222
- - Only `lmx` honours it. On `openai-compatible`, `lmstudio` or `anthropic` the client logs a
223
- warning, removes the option, and nothing reaches that provider's request body.
250
+ `extra: { reasoning_effort: … }` cannot change it. Only `reasoning` can. The option itself is never
251
+ copied into a request body.
224
252
  - `REASONING_MODES` (package root) lists the accepted values for an app that validates its own
225
253
  settings.
226
254
 
@@ -1,138 +1,138 @@
1
- /**
2
- * The fold on the Inference list, and prefilling the supervisor form from a panel.
3
- *
4
- * WHY ANYTHING FOLDS. A stable stack is four engines, three credentials, a certificate fingerprint
5
- * and a verification report — worth having, not worth reading every time you open this page to do
6
- * something else. Folded, a panel is one line that still carries the whole verdict: status, the
7
- * count of what depends on it, all three credentials and when it was last checked. Folding hides
8
- * detail; it must never hide the reason you would have opened it.
9
- *
10
- * A PANEL WITH A PROBLEM IGNORES WHAT YOU REMEMBERED. `data-panel-attention` is stamped by the
11
- * server on anything with a missing engine, an expiring certificate, a failed check or a nearly
12
- * spent cap. Those open and stay open. Attention beats tidiness — the alternative is a page that
13
- * quietly honours a fold you chose last week and hides the thing that broke yesterday.
14
- *
15
- * THE PREFERENCE IS PER BROWSER AND DISPOSABLE. It is a convenience about how a page looks to one
16
- * person, so localStorage is the right home and losing it costs nothing. Every access is guarded:
17
- * a browser set to block site data throws on read, and a settings page must not break because
18
- * somebody tightened their privacy settings.
19
- */
20
- (function () {
21
- 'use strict';
22
-
23
- var KEY = 's101.ai.panels';
24
-
25
- function readPrefs() {
26
- try {
27
- return JSON.parse(window.localStorage.getItem(KEY) || '{}') || {};
28
- } catch (e) {
29
- return {};
30
- }
31
- }
32
-
33
- function writePref(id, open) {
34
- try {
35
- var prefs = readPrefs();
36
- prefs[id] = !!open;
37
- window.localStorage.setItem(KEY, JSON.stringify(prefs));
38
- } catch (e) { /* private window, or site data blocked — the fold still works for this visit */ }
39
- }
40
-
41
- function bodyOf(panel) {
42
- var btn = panel.querySelector('[data-panel-toggle]');
43
- if (!btn) return null;
44
- return document.getElementById(btn.getAttribute('aria-controls'));
45
- }
46
-
47
- function setOpen(panel, open, remember) {
48
- var btn = panel.querySelector('[data-panel-toggle]');
49
- var body = bodyOf(panel);
50
- if (!btn || !body) return;
51
- // `hidden`, not a style: the server renders the closed state the same way, so a panel does not
52
- // flicker open on load before this script runs.
53
- body.hidden = !open;
54
- btn.setAttribute('aria-expanded', open ? 'true' : 'false');
55
- panel.classList.toggle('panel-open', open);
56
- if (remember) writePref(panel.getAttribute('data-panel'), open);
57
- }
58
-
59
- function restore() {
60
- var prefs = readPrefs();
61
- var panels = document.querySelectorAll('[data-panel]');
62
- for (var i = 0; i < panels.length; i += 1) {
63
- var panel = panels[i];
64
- var id = panel.getAttribute('data-panel');
65
- // The server already opened this one and means it. Do not consult the preference at all —
66
- // reading it and then ignoring it is the same thing, but invites somebody to "fix" it later.
67
- if (panel.hasAttribute('data-panel-attention')) {
68
- setOpen(panel, true, false);
69
- continue;
70
- }
71
- if (Object.prototype.hasOwnProperty.call(prefs, id)) setOpen(panel, !!prefs[id], false);
72
- }
73
- }
74
-
75
- document.addEventListener('click', function (ev) {
76
- var toggle = ev.target.closest('[data-panel-toggle]');
77
- if (toggle) {
78
- var panel = toggle.closest('[data-panel]');
79
- if (!panel) return;
80
- var body = bodyOf(panel);
81
- setOpen(panel, !!(body && body.hidden), true);
82
- return;
83
- }
84
-
85
- // ── EDITING A STACK HAPPENS IN THE PANEL ────────────────────────────────────────────────────
86
- //
87
- // It used to happen in a form at the foot of the page, and Edit scrolled you down to it — away
88
- // from the stack you were reading, into a form that also served as "Add a supervisor" and said
89
- // nothing about which stack it had been filled with. Nothing needs prefilling now: every panel
90
- // renders its own form, from the server, with its own values already in it.
91
- var open = ev.target.closest('[data-stack-edit-open]');
92
- if (open) { stackMode(open.closest('[data-stack]'), true); return; }
93
-
94
- var cancel = ev.target.closest('[data-stack-cancel]');
95
- if (cancel) {
96
- var panel2 = cancel.closest('[data-stack]');
97
- // A NEW stack has no reading half to go back to, so cancelling it puts the blank panel away.
98
- if (panel2 && panel2.hasAttribute('data-stack-new')) { panel2.hidden = true; return; }
99
- // Otherwise reload rather than restore: a Cancel that left half-typed values behind, ready to
100
- // be posted by the next Save, would be worse than the extra request.
101
- window.location.reload();
102
- return;
103
- }
104
-
105
- // "Add LMX" reveals the blank panel that is already on the page, at the top of the list.
106
- var add = ev.target.closest('[data-stack-add]');
107
- if (add) {
108
- var blank = document.querySelector('[data-stack-new]');
109
- if (!blank) return;
110
- blank.hidden = false;
111
- blank.scrollIntoView({ block: 'nearest' });
112
- var id = blank.querySelector('[name="id"]');
113
- if (id) id.focus();
114
- }
115
- });
116
-
117
- /** Swap one stack panel between reading and editing. */
118
- function stackMode(panel, editing) {
119
- if (!panel) return;
120
- var view = panel.querySelector('[data-stack-view]');
121
- var form = panel.querySelector('[data-stack-edit]');
122
- if (!view || !form) return;
123
- view.hidden = editing;
124
- form.hidden = !editing;
125
- if (editing) {
126
- // The id is readonly on an existing stack, so focus the first field somebody can actually
127
- // change rather than one that will not accept a keystroke.
128
- var first = form.querySelector('[name="label"]') || form.querySelector('[name="status_url"]');
129
- if (first) first.focus();
130
- }
131
- }
132
-
133
- if (document.readyState === 'loading') {
134
- document.addEventListener('DOMContentLoaded', restore);
135
- } else {
136
- restore();
137
- }
138
- })();
1
+ /**
2
+ * The fold on the Inference list, and prefilling the supervisor form from a panel.
3
+ *
4
+ * WHY ANYTHING FOLDS. A stable stack is four engines, three credentials, a certificate fingerprint
5
+ * and a verification report — worth having, not worth reading every time you open this page to do
6
+ * something else. Folded, a panel is one line that still carries the whole verdict: status, the
7
+ * count of what depends on it, all three credentials and when it was last checked. Folding hides
8
+ * detail; it must never hide the reason you would have opened it.
9
+ *
10
+ * A PANEL WITH A PROBLEM IGNORES WHAT YOU REMEMBERED. `data-panel-attention` is stamped by the
11
+ * server on anything with a missing engine, an expiring certificate, a failed check or a nearly
12
+ * spent cap. Those open and stay open. Attention beats tidiness — the alternative is a page that
13
+ * quietly honours a fold you chose last week and hides the thing that broke yesterday.
14
+ *
15
+ * THE PREFERENCE IS PER BROWSER AND DISPOSABLE. It is a convenience about how a page looks to one
16
+ * person, so localStorage is the right home and losing it costs nothing. Every access is guarded:
17
+ * a browser set to block site data throws on read, and a settings page must not break because
18
+ * somebody tightened their privacy settings.
19
+ */
20
+ (function () {
21
+ 'use strict';
22
+
23
+ var KEY = 's101.ai.panels';
24
+
25
+ function readPrefs() {
26
+ try {
27
+ return JSON.parse(window.localStorage.getItem(KEY) || '{}') || {};
28
+ } catch (e) {
29
+ return {};
30
+ }
31
+ }
32
+
33
+ function writePref(id, open) {
34
+ try {
35
+ var prefs = readPrefs();
36
+ prefs[id] = !!open;
37
+ window.localStorage.setItem(KEY, JSON.stringify(prefs));
38
+ } catch (e) { /* private window, or site data blocked — the fold still works for this visit */ }
39
+ }
40
+
41
+ function bodyOf(panel) {
42
+ var btn = panel.querySelector('[data-panel-toggle]');
43
+ if (!btn) return null;
44
+ return document.getElementById(btn.getAttribute('aria-controls'));
45
+ }
46
+
47
+ function setOpen(panel, open, remember) {
48
+ var btn = panel.querySelector('[data-panel-toggle]');
49
+ var body = bodyOf(panel);
50
+ if (!btn || !body) return;
51
+ // `hidden`, not a style: the server renders the closed state the same way, so a panel does not
52
+ // flicker open on load before this script runs.
53
+ body.hidden = !open;
54
+ btn.setAttribute('aria-expanded', open ? 'true' : 'false');
55
+ panel.classList.toggle('panel-open', open);
56
+ if (remember) writePref(panel.getAttribute('data-panel'), open);
57
+ }
58
+
59
+ function restore() {
60
+ var prefs = readPrefs();
61
+ var panels = document.querySelectorAll('[data-panel]');
62
+ for (var i = 0; i < panels.length; i += 1) {
63
+ var panel = panels[i];
64
+ var id = panel.getAttribute('data-panel');
65
+ // The server already opened this one and means it. Do not consult the preference at all —
66
+ // reading it and then ignoring it is the same thing, but invites somebody to "fix" it later.
67
+ if (panel.hasAttribute('data-panel-attention')) {
68
+ setOpen(panel, true, false);
69
+ continue;
70
+ }
71
+ if (Object.prototype.hasOwnProperty.call(prefs, id)) setOpen(panel, !!prefs[id], false);
72
+ }
73
+ }
74
+
75
+ document.addEventListener('click', function (ev) {
76
+ var toggle = ev.target.closest('[data-panel-toggle]');
77
+ if (toggle) {
78
+ var panel = toggle.closest('[data-panel]');
79
+ if (!panel) return;
80
+ var body = bodyOf(panel);
81
+ setOpen(panel, !!(body && body.hidden), true);
82
+ return;
83
+ }
84
+
85
+ // ── EDITING A STACK HAPPENS IN THE PANEL ────────────────────────────────────────────────────
86
+ //
87
+ // It used to happen in a form at the foot of the page, and Edit scrolled you down to it — away
88
+ // from the stack you were reading, into a form that also served as "Add a supervisor" and said
89
+ // nothing about which stack it had been filled with. Nothing needs prefilling now: every panel
90
+ // renders its own form, from the server, with its own values already in it.
91
+ var open = ev.target.closest('[data-stack-edit-open]');
92
+ if (open) { stackMode(open.closest('[data-stack]'), true); return; }
93
+
94
+ var cancel = ev.target.closest('[data-stack-cancel]');
95
+ if (cancel) {
96
+ var panel2 = cancel.closest('[data-stack]');
97
+ // A NEW stack has no reading half to go back to, so cancelling it puts the blank panel away.
98
+ if (panel2 && panel2.hasAttribute('data-stack-new')) { panel2.hidden = true; return; }
99
+ // Otherwise reload rather than restore: a Cancel that left half-typed values behind, ready to
100
+ // be posted by the next Save, would be worse than the extra request.
101
+ window.location.reload();
102
+ return;
103
+ }
104
+
105
+ // "Add LMX" reveals the blank panel that is already on the page, at the top of the list.
106
+ var add = ev.target.closest('[data-stack-add]');
107
+ if (add) {
108
+ var blank = document.querySelector('[data-stack-new]');
109
+ if (!blank) return;
110
+ blank.hidden = false;
111
+ blank.scrollIntoView({ block: 'nearest' });
112
+ var id = blank.querySelector('[name="id"]');
113
+ if (id) id.focus();
114
+ }
115
+ });
116
+
117
+ /** Swap one stack panel between reading and editing. */
118
+ function stackMode(panel, editing) {
119
+ if (!panel) return;
120
+ var view = panel.querySelector('[data-stack-view]');
121
+ var form = panel.querySelector('[data-stack-edit]');
122
+ if (!view || !form) return;
123
+ view.hidden = editing;
124
+ form.hidden = !editing;
125
+ if (editing) {
126
+ // The id is readonly on an existing stack, so focus the first field somebody can actually
127
+ // change rather than one that will not accept a keystroke.
128
+ var first = form.querySelector('[name="label"]') || form.querySelector('[name="status_url"]');
129
+ if (first) first.focus();
130
+ }
131
+ }
132
+
133
+ if (document.readyState === 'loading') {
134
+ document.addEventListener('DOMContentLoaded', restore);
135
+ } else {
136
+ restore();
137
+ }
138
+ })();
package/index.js CHANGED
@@ -26,7 +26,7 @@ const facts = require('./facts');
26
26
  const { AiError, fromFetchFailure, redact } = require('./error');
27
27
  const { polish } = require('./polish');
28
28
  const { generate } = require('./generate');
29
- const { assertReasoning, REASONING_MODES } = require('./providers/lmx');
29
+ const { assertReasoning, REASONING_MODES } = require('./reasoning');
30
30
 
31
31
  const PROVIDERS = {
32
32
  // 'lmstudio' and 'openai-compatible' are the SAME adapter with different defaults — a kindness to
@@ -52,9 +52,6 @@ const DEFAULTS = {
52
52
  const RETRY_AFTER_MS = 400;
53
53
  const RETRY_ONLY_IF_FAILED_WITHIN_MS = 5000;
54
54
 
55
- /** Providers already warned that they ignore `reasoning` — once per provider per process. */
56
- const warnedReasoning = new Set();
57
-
58
55
  const NOOP_LOGGER = { info() {}, warn() {}, error() {} };
59
56
  const NOOP_BUDGET = { async assertWithinBudget() {}, async record() {} };
60
57
 
@@ -106,7 +103,7 @@ function createAiClient(deps = {}) {
106
103
  * Ask the configured model for something.
107
104
  * @param {{system?:string, messages:Array, maxTokens?:number, temperature?:number,
108
105
  * schema?:object, signal?:AbortSignal, ticketId?:*, skipBudget?:boolean,
109
- * reasoning?:'full'}} opts
106
+ * reasoning?:'off'|'full'}} opts
110
107
  * @param {object} [cfgOverride] the resolved config, when the caller already has it
111
108
  */
112
109
  async function complete(opts, cfgOverride) {
@@ -134,19 +131,9 @@ function createAiClient(deps = {}) {
134
131
  // by whichever caller forgets. `skipBudget` is for the admin's Test connection; it still records.
135
132
  if (!opts.skipBudget) await meter.assertWithinBudget(cfg, { ticketId: opts.ticketId });
136
133
 
137
- // ONLY lmx HONOURS `reasoning`. Elsewhere it is warned about and removed, so no other adapter's
138
- // request body can ever carry it.
139
- let callOpts = opts;
140
- if (opts.reasoning != null && cfg.provider !== 'lmx') {
141
- if (!warnedReasoning.has(cfg.provider)) {
142
- warnedReasoning.add(cfg.provider);
143
- log.warn(`AI: reasoning '${opts.reasoning}' is honoured only by lmx engines — `
144
- + `${cfg.label || cfg.provider} sends no reasoning flag and runs on its own default `
145
- + '(warned once per provider).');
146
- }
147
- const { reasoning: _ignored, ...rest } = opts;
148
- callOpts = rest;
149
- }
134
+ // EVERY PROVIDER HONOURS `reasoning` (0.27.0): each adapter translates it for its model family
135
+ // (see reasoning.js). It used to be lmx-only, warned about and stripped everywhere else.
136
+ const callOpts = opts;
150
137
 
151
138
  const started = Date.now();
152
139
  let result;