@aria-framework/ai 0.25.0 → 0.26.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +54 -2
- package/index.js +333 -309
- package/lmxStatus.js +1 -1
- package/package.json +2 -2
- package/providers/lmx.js +52 -12
- package/providers/openai-compatible.js +19 -3
package/README.md
CHANGED
|
@@ -32,7 +32,11 @@ const ai = createAiClient({
|
|
|
32
32
|
logger: console // optional
|
|
33
33
|
});
|
|
34
34
|
|
|
35
|
-
const r = await ai.complete({
|
|
35
|
+
const r = await ai.complete({
|
|
36
|
+
system, messages, maxTokens: 400,
|
|
37
|
+
schema, // optional: constrained JSON
|
|
38
|
+
// reasoning: 'full', // optional, lmx only: let the model think fully (see Reasoning below)
|
|
39
|
+
});
|
|
36
40
|
// r.text, r.json (when schema), r.model, r.usage.total, r.ms
|
|
37
41
|
```
|
|
38
42
|
|
|
@@ -59,7 +63,7 @@ to write yourself:
|
|
|
59
63
|
| Engine URLs come from the document | Resolved on every call, and never stored |
|
|
60
64
|
| A status outage is not an inference outage | Keeps routing on the last good document for 3 minutes (`staleMs`), logging loudly |
|
|
61
65
|
| Pin the certificate; never disable verification | Uses an undici `Agent({ connect: { ca } })` per stack. Never `NODE_EXTRA_CA_CERTS`, never `rejectUnauthorized:false` |
|
|
62
|
-
| Choose the reasoning flag from the model | `qwen` → `chat_template_kwargs.enable_thinking=false`, `gpt-oss` → `reasoning_effort:'low'`, anything else gets neither (with a warning) |
|
|
66
|
+
| Choose the reasoning flag from the model | By default thinking is kept to a minimum: `qwen` → `chat_template_kwargs.enable_thinking=false`, `gpt-oss` → `reasoning_effort:'low'`, anything else gets neither (with a warning). `reasoning: 'full'` on the call (since 0.26.0) sends `enable_thinking=true` / `reasoning_effort:'high'` instead. The flag always overrides a raw `extra`; only the named option changes it. See [Reasoning](#reasoning-reasoning-full) |
|
|
63
67
|
| Size requests against `maxInputTokens` (per slot) | `contextTokens` is taken from the engine, not from config |
|
|
64
68
|
| Identity is `(instance, id)`, falling back to `name` | `lmx.engineId` is matched first; `rekeyPlan()` migrates rows when ids are minted or engines renamed |
|
|
65
69
|
| A `429` is about the key, not the engine | Fails fast (since 0.25.0): throws `rate_limit` with `lmxSkip: 'lmx_throttled'` and `retryAfterMs`. The client never waits; whether to wait or move on is the caller's decision |
|
|
@@ -166,6 +170,54 @@ instead:
|
|
|
166
170
|
Failover across several engines for one job is the app's job; lmx provides none. List
|
|
167
171
|
endpoints in preference order and take the first one that serves.
|
|
168
172
|
|
|
173
|
+
### Reasoning (`reasoning: 'full'`)
|
|
174
|
+
|
|
175
|
+
By default the adapter makes the model think as little as its family allows, so a quick answer
|
|
176
|
+
stays quick. Pass `reasoning: 'full'` on a `complete()` call to let the model think fully:
|
|
177
|
+
|
|
178
|
+
```js
|
|
179
|
+
const r = await ai.complete({
|
|
180
|
+
system, messages, schema,
|
|
181
|
+
reasoning: 'full',
|
|
182
|
+
maxTokens: 4096, // at least 4096; see the trap below
|
|
183
|
+
}); // and give the endpoint timeoutMs: 120000 or more
|
|
184
|
+
```
|
|
185
|
+
|
|
186
|
+
**When to use it.** Use it for judgement tasks where the answer depends on weighing evidence,
|
|
187
|
+
such as alert triage. Do not use it for quick summaries, rewrites or extraction. They do not get
|
|
188
|
+
better, only slower.
|
|
189
|
+
|
|
190
|
+
**What it costs.** It is much slower, roughly 30–45 s per answer on lab hardware instead of a few
|
|
191
|
+
seconds, and it spends many more tokens, because the thinking is generated and billed like any
|
|
192
|
+
other output. Budget for both.
|
|
193
|
+
|
|
194
|
+
**The trap.** Thinking comes out of `maxTokens`. If the limit is too small, the model spends all
|
|
195
|
+
of it thinking and returns **an empty answer, with no error from the server** (`finish_reason:
|
|
196
|
+
"length"`). The adapter turns that, and a reply cut off inside an unclosed `<think>` block, into
|
|
197
|
+
a `bad_response` AiError saying it ran out of room (and, when the server says so, that the
|
|
198
|
+
tokens went on internal reasoning), but the work is lost either way. Use **`maxTokens` ≥ 4096** and a
|
|
199
|
+
**`timeoutMs` ≥ 120 s** (120000) on the endpoint, or the deadline cuts the answer off first.
|
|
200
|
+
|
|
201
|
+
**What is sent, per model family** (read from the model the engine is running on each call):
|
|
202
|
+
|
|
203
|
+
| Family | Default (no `reasoning`) | `reasoning: 'full'` |
|
|
204
|
+
|---|---|---|
|
|
205
|
+
| qwen | `{"chat_template_kwargs":{"enable_thinking":false}}` | `{"chat_template_kwargs":{"enable_thinking":true}}` |
|
|
206
|
+
| gpt-oss | `{"reasoning_effort":"low"}` | `{"reasoning_effort":"high"}` |
|
|
207
|
+
| anything else | nothing (warning logged) | nothing (warning logged that `'full'` could not be honoured) |
|
|
208
|
+
|
|
209
|
+
Rules:
|
|
210
|
+
|
|
211
|
+
- `'full'` is the only value. Leaving the option out (or `null`) means the default. Any other value,
|
|
212
|
+
for example `'high'`, `true` or `''`, is refused with `AiError` kind `refused` before any
|
|
213
|
+
request is made, whichever provider the endpoint uses.
|
|
214
|
+
- The flag is merged **after** `extra`, so a raw `extra: { chat_template_kwargs: … }` or
|
|
215
|
+
`extra: { reasoning_effort: … }` cannot change it. Only `reasoning` can.
|
|
216
|
+
- Only `lmx` honours it. On `openai-compatible`, `lmstudio` or `anthropic` the client logs a
|
|
217
|
+
warning, removes the option, and nothing reaches that provider's request body.
|
|
218
|
+
- `REASONING_MODES` (package root) lists the accepted values for an app that validates its own
|
|
219
|
+
settings.
|
|
220
|
+
|
|
169
221
|
### Embeddings
|
|
170
222
|
|
|
171
223
|
Call `client.embed(cfg, texts, { signal })` and pass your own resolved embedding config; it never
|
package/index.js
CHANGED
|
@@ -1,309 +1,333 @@
|
|
|
1
|
-
/**
|
|
2
|
-
* @aria-framework/ai — the AI seam. One `complete()`, several providers behind it, plus the
|
|
3
|
-
* writing-assist engines (polish/generate) and the fact-preservation guard.
|
|
4
|
-
*
|
|
5
|
-
* DEPENDENCY-INJECTED, DATABASE-FREE. The package knows how to talk to a model; it does NOT know
|
|
6
|
-
* where an app keeps its settings, its credentials or its token ledger. The consumer builds a
|
|
7
|
-
* client with two functions of its own:
|
|
8
|
-
*
|
|
9
|
-
* const ai = createAiClient({
|
|
10
|
-
* resolveConfig, // async () => resolved config (provider, baseUrl, model, apiKey, caps…)
|
|
11
|
-
* budget, // { assertWithinBudget(cfg, ctx), record(cfg, result, ctx) } — optional
|
|
12
|
-
* logger // { info, warn, error } — optional
|
|
13
|
-
* });
|
|
14
|
-
*
|
|
15
|
-
* The provider adapters already take an explicit config and never read a database, which is what
|
|
16
|
-
* makes the seam testable: a stub adapter and a real adapter are called identically.
|
|
17
|
-
*
|
|
18
|
-
* PROMPTS ARE CONTENT AND LIVE IN THE APP. This package carries the mechanism (how to call a model,
|
|
19
|
-
* how to enforce a token ceiling, how to check a rewrite kept its facts) and generic writing
|
|
20
|
-
* operations; the words that say "you are editing a reply to a customer" belong to the app.
|
|
21
|
-
*/
|
|
22
|
-
|
|
23
|
-
'use strict';
|
|
24
|
-
|
|
25
|
-
const facts = require('./facts');
|
|
26
|
-
const { AiError, fromFetchFailure, redact } = require('./error');
|
|
27
|
-
const { polish } = require('./polish');
|
|
28
|
-
const { generate } = require('./generate');
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
//
|
|
33
|
-
//
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
//
|
|
39
|
-
//
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
}
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
const
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
const
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
*
|
|
75
|
-
*
|
|
76
|
-
*
|
|
77
|
-
*
|
|
78
|
-
*
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
//
|
|
114
|
-
//
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
if (!cfg.
|
|
118
|
-
throw new AiError('
|
|
119
|
-
}
|
|
120
|
-
|
|
121
|
-
|
|
122
|
-
|
|
123
|
-
|
|
124
|
-
|
|
125
|
-
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
|
|
131
|
-
|
|
132
|
-
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
|
|
137
|
-
|
|
138
|
-
|
|
139
|
-
|
|
140
|
-
|
|
141
|
-
|
|
142
|
-
|
|
143
|
-
|
|
144
|
-
|
|
145
|
-
|
|
146
|
-
|
|
147
|
-
|
|
148
|
-
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
|
|
152
|
-
|
|
153
|
-
//
|
|
154
|
-
|
|
155
|
-
|
|
156
|
-
|
|
157
|
-
|
|
158
|
-
|
|
159
|
-
|
|
160
|
-
|
|
161
|
-
|
|
162
|
-
|
|
163
|
-
|
|
164
|
-
|
|
165
|
-
|
|
166
|
-
|
|
167
|
-
|
|
168
|
-
|
|
169
|
-
|
|
170
|
-
|
|
171
|
-
|
|
172
|
-
|
|
173
|
-
async function
|
|
174
|
-
|
|
175
|
-
|
|
176
|
-
|
|
177
|
-
|
|
178
|
-
|
|
179
|
-
|
|
180
|
-
|
|
181
|
-
|
|
182
|
-
|
|
183
|
-
|
|
184
|
-
|
|
185
|
-
|
|
186
|
-
|
|
187
|
-
|
|
188
|
-
|
|
189
|
-
|
|
190
|
-
|
|
191
|
-
|
|
192
|
-
}
|
|
193
|
-
|
|
194
|
-
/**
|
|
195
|
-
async function
|
|
196
|
-
|
|
197
|
-
|
|
198
|
-
|
|
199
|
-
|
|
200
|
-
|
|
201
|
-
|
|
202
|
-
|
|
203
|
-
|
|
204
|
-
|
|
205
|
-
|
|
206
|
-
|
|
207
|
-
|
|
208
|
-
|
|
209
|
-
|
|
210
|
-
}
|
|
211
|
-
|
|
212
|
-
|
|
213
|
-
|
|
214
|
-
|
|
215
|
-
|
|
216
|
-
|
|
217
|
-
|
|
218
|
-
|
|
219
|
-
|
|
220
|
-
|
|
221
|
-
|
|
222
|
-
|
|
223
|
-
|
|
224
|
-
|
|
225
|
-
*
|
|
226
|
-
*
|
|
227
|
-
*/
|
|
228
|
-
async function
|
|
229
|
-
const cfg = cfgOverride || await resolveConfig();
|
|
230
|
-
|
|
231
|
-
|
|
232
|
-
|
|
233
|
-
|
|
234
|
-
|
|
235
|
-
|
|
236
|
-
|
|
237
|
-
|
|
238
|
-
|
|
239
|
-
|
|
240
|
-
|
|
241
|
-
|
|
242
|
-
|
|
243
|
-
|
|
244
|
-
}
|
|
245
|
-
|
|
246
|
-
|
|
247
|
-
|
|
248
|
-
|
|
249
|
-
|
|
250
|
-
|
|
251
|
-
|
|
252
|
-
|
|
253
|
-
|
|
254
|
-
|
|
255
|
-
//
|
|
256
|
-
//
|
|
257
|
-
|
|
258
|
-
|
|
259
|
-
|
|
260
|
-
|
|
261
|
-
|
|
262
|
-
|
|
263
|
-
|
|
264
|
-
|
|
265
|
-
|
|
266
|
-
|
|
267
|
-
|
|
268
|
-
|
|
269
|
-
|
|
270
|
-
//
|
|
271
|
-
//
|
|
272
|
-
|
|
273
|
-
|
|
274
|
-
|
|
275
|
-
//
|
|
276
|
-
|
|
277
|
-
//
|
|
278
|
-
|
|
279
|
-
//
|
|
280
|
-
//
|
|
281
|
-
|
|
282
|
-
//
|
|
283
|
-
|
|
284
|
-
|
|
285
|
-
|
|
286
|
-
|
|
287
|
-
|
|
288
|
-
|
|
289
|
-
|
|
290
|
-
|
|
291
|
-
|
|
292
|
-
|
|
293
|
-
|
|
294
|
-
|
|
295
|
-
//
|
|
296
|
-
|
|
297
|
-
//
|
|
298
|
-
|
|
299
|
-
//
|
|
300
|
-
|
|
301
|
-
//
|
|
302
|
-
//
|
|
303
|
-
|
|
304
|
-
|
|
305
|
-
//
|
|
306
|
-
|
|
307
|
-
|
|
308
|
-
|
|
309
|
-
|
|
1
|
+
/**
|
|
2
|
+
* @aria-framework/ai — the AI seam. One `complete()`, several providers behind it, plus the
|
|
3
|
+
* writing-assist engines (polish/generate) and the fact-preservation guard.
|
|
4
|
+
*
|
|
5
|
+
* DEPENDENCY-INJECTED, DATABASE-FREE. The package knows how to talk to a model; it does NOT know
|
|
6
|
+
* where an app keeps its settings, its credentials or its token ledger. The consumer builds a
|
|
7
|
+
* client with two functions of its own:
|
|
8
|
+
*
|
|
9
|
+
* const ai = createAiClient({
|
|
10
|
+
* resolveConfig, // async () => resolved config (provider, baseUrl, model, apiKey, caps…)
|
|
11
|
+
* budget, // { assertWithinBudget(cfg, ctx), record(cfg, result, ctx) } — optional
|
|
12
|
+
* logger // { info, warn, error } — optional
|
|
13
|
+
* });
|
|
14
|
+
*
|
|
15
|
+
* The provider adapters already take an explicit config and never read a database, which is what
|
|
16
|
+
* makes the seam testable: a stub adapter and a real adapter are called identically.
|
|
17
|
+
*
|
|
18
|
+
* PROMPTS ARE CONTENT AND LIVE IN THE APP. This package carries the mechanism (how to call a model,
|
|
19
|
+
* how to enforce a token ceiling, how to check a rewrite kept its facts) and generic writing
|
|
20
|
+
* operations; the words that say "you are editing a reply to a customer" belong to the app.
|
|
21
|
+
*/
|
|
22
|
+
|
|
23
|
+
'use strict';
|
|
24
|
+
|
|
25
|
+
const facts = require('./facts');
|
|
26
|
+
const { AiError, fromFetchFailure, redact } = require('./error');
|
|
27
|
+
const { polish } = require('./polish');
|
|
28
|
+
const { generate } = require('./generate');
|
|
29
|
+
const { assertReasoning, REASONING_MODES } = require('./providers/lmx');
|
|
30
|
+
|
|
31
|
+
const PROVIDERS = {
|
|
32
|
+
// 'lmstudio' and 'openai-compatible' are the SAME adapter with different defaults — a kindness to
|
|
33
|
+
// whoever configures it: an operator running LM Studio should not have to know it speaks a shape
|
|
34
|
+
// named after somebody else.
|
|
35
|
+
lmstudio: require('./providers/openai-compatible'),
|
|
36
|
+
'openai-compatible': require('./providers/openai-compatible'),
|
|
37
|
+
anthropic: require('./providers/anthropic'),
|
|
38
|
+
// 0.14.0 — a supervised fleet rather than an address. The engine URL is discovered from the
|
|
39
|
+
// supervisor's status document per call, the reasoning flag is read from the model the engine
|
|
40
|
+
// is actually running, and the whole conversation is pinned to a self-signed certificate.
|
|
41
|
+
lmx: require('./providers/lmx')
|
|
42
|
+
};
|
|
43
|
+
|
|
44
|
+
const DEFAULTS = {
|
|
45
|
+
lmstudio: { baseUrl: 'http://localhost:1234/v1', model: 'qwen3.5-9b', label: 'LM Studio' },
|
|
46
|
+
'openai-compatible': { baseUrl: 'http://localhost:11434/v1', model: '', label: 'The model server' },
|
|
47
|
+
anthropic: { baseUrl: 'https://api.anthropic.com/v1', model: 'claude-sonnet-4-5', label: 'Claude' },
|
|
48
|
+
// No baseUrl: an lmx engine's address is never configured, only discovered.
|
|
49
|
+
lmx: { baseUrl: '', model: '', label: 'lmx engine' }
|
|
50
|
+
};
|
|
51
|
+
|
|
52
|
+
const RETRY_AFTER_MS = 400;
|
|
53
|
+
const RETRY_ONLY_IF_FAILED_WITHIN_MS = 5000;
|
|
54
|
+
|
|
55
|
+
/** Providers already warned that they ignore `reasoning` — once per provider per process. */
|
|
56
|
+
const warnedReasoning = new Set();
|
|
57
|
+
|
|
58
|
+
const NOOP_LOGGER = { info() {}, warn() {}, error() {} };
|
|
59
|
+
const NOOP_BUDGET = { async assertWithinBudget() {}, async record() {} };
|
|
60
|
+
|
|
61
|
+
/**
|
|
62
|
+
* Build an AI client bound to one app's config resolution and token budget.
|
|
63
|
+
* @param {{resolveConfig: () => Promise<object>, budget?: object, logger?: object}} deps
|
|
64
|
+
*/
|
|
65
|
+
function createAiClient(deps = {}) {
|
|
66
|
+
const resolveConfig = deps.resolveConfig;
|
|
67
|
+
if (typeof resolveConfig !== 'function') {
|
|
68
|
+
throw new Error('createAiClient: resolveConfig must be an async function returning the resolved config');
|
|
69
|
+
}
|
|
70
|
+
const log = deps.logger || NOOP_LOGGER;
|
|
71
|
+
const meter = deps.budget || NOOP_BUDGET;
|
|
72
|
+
|
|
73
|
+
/**
|
|
74
|
+
* Try once more, but only for the failure where trying again could help — a local provider that
|
|
75
|
+
* dropped the connection while loading a model. A cancelled call, a timeout, a rate limit or a slow
|
|
76
|
+
* failure is never retried (see the guards below).
|
|
77
|
+
*
|
|
78
|
+
* A 429 FAILS FAST (0.25.0). 0.23/0.24 slept its Retry-After here (up to 10 s), below every app's
|
|
79
|
+
* dispatcher: every engine on an lmx stack shares one gateway key, so each failover re-hit the same
|
|
80
|
+
* throttle and paid the wait again, and an OpenAI-compatible primary that sent Retry-After delayed a
|
|
81
|
+
* healthy backup by the full wait. Decided 2026-10-01: the client does not wait. The error carries
|
|
82
|
+
* `retryAfterMs` (and `lmxSkip: 'lmx_throttled'` from an lmx engine) for a caller that chooses to.
|
|
83
|
+
*/
|
|
84
|
+
async function withOneRetry(run, opts = {}) {
|
|
85
|
+
const startedAt = Date.now();
|
|
86
|
+
try {
|
|
87
|
+
return await run();
|
|
88
|
+
} catch (err) {
|
|
89
|
+
if (opts.signal && opts.signal.aborted) throw err;
|
|
90
|
+
if (!err || !err.retryable) throw err;
|
|
91
|
+
const elapsed = Date.now() - startedAt;
|
|
92
|
+
// A TIMEOUT is the deadline itself being reached — retrying waits the whole deadline again. A
|
|
93
|
+
// RATE LIMIT is the provider asking for less pressure. A slow `unreachable` is not the
|
|
94
|
+
// sub-second dropped-connection transient this retry exists for. None of those retry.
|
|
95
|
+
if (err.kind === 'timeout' || err.kind === 'rate_limit' || elapsed > RETRY_ONLY_IF_FAILED_WITHIN_MS) {
|
|
96
|
+
log.warn(`AI: ${err.kind} after ${elapsed}ms — not retrying (${err.message})`);
|
|
97
|
+
throw err;
|
|
98
|
+
}
|
|
99
|
+
log.warn(`AI: ${err.kind} — trying once more in ${RETRY_AFTER_MS}ms (${err.message})`);
|
|
100
|
+
await new Promise((r) => setTimeout(r, RETRY_AFTER_MS));
|
|
101
|
+
return run();
|
|
102
|
+
}
|
|
103
|
+
}
|
|
104
|
+
|
|
105
|
+
/**
|
|
106
|
+
* Ask the configured model for something.
|
|
107
|
+
* @param {{system?:string, messages:Array, maxTokens?:number, temperature?:number,
|
|
108
|
+
* schema?:object, signal?:AbortSignal, ticketId?:*, skipBudget?:boolean,
|
|
109
|
+
* reasoning?:'full'}} opts
|
|
110
|
+
* @param {object} [cfgOverride] the resolved config, when the caller already has it
|
|
111
|
+
*/
|
|
112
|
+
async function complete(opts, cfgOverride) {
|
|
113
|
+
// `reasoning` (0.26.0) is checked here for EVERY provider, first: a typo is the caller's bug
|
|
114
|
+
// whichever endpoint the route happens to pick, and must not wait for an lmx one to surface.
|
|
115
|
+
assertReasoning(opts && opts.reasoning);
|
|
116
|
+
const cfg = cfgOverride || await resolveConfig();
|
|
117
|
+
if (!cfg.enabled) {
|
|
118
|
+
throw new AiError('disabled', 'AI assistance is switched off. An administrator can enable it in Settings.');
|
|
119
|
+
}
|
|
120
|
+
// A MODEL NAME IS REQUIRED OF EVERY PROVIDER THAT HAS ONE TO CONFIGURE — which is all of them
|
|
121
|
+
// except a supervised stack. There the model is a FACT the supervisor reports about an engine,
|
|
122
|
+
// not a setting: the operator picks an engine and whatever it is running answers. Insisting on
|
|
123
|
+
// a model name would make an lmx endpoint unusable by demanding the one field the design says
|
|
124
|
+
// not to store, and would say so with a message naming nothing an operator could go and fill in.
|
|
125
|
+
if (!cfg.model && cfg.provider !== 'lmx') {
|
|
126
|
+
throw new AiError('unconfigured', 'No model name is configured.');
|
|
127
|
+
}
|
|
128
|
+
const adapter = PROVIDERS[cfg.provider];
|
|
129
|
+
if (!adapter) {
|
|
130
|
+
throw new AiError('unconfigured', `No adapter is registered for provider "${cfg.provider}".`);
|
|
131
|
+
}
|
|
132
|
+
|
|
133
|
+
// The ceiling, before the call — the only place a limit can be enforced without being bypassable
|
|
134
|
+
// by whichever caller forgets. `skipBudget` is for the admin's Test connection; it still records.
|
|
135
|
+
if (!opts.skipBudget) await meter.assertWithinBudget(cfg, { ticketId: opts.ticketId });
|
|
136
|
+
|
|
137
|
+
// ONLY lmx HONOURS `reasoning`. Elsewhere it is warned about and removed, so no other adapter's
|
|
138
|
+
// request body can ever carry it.
|
|
139
|
+
let callOpts = opts;
|
|
140
|
+
if (opts.reasoning != null && cfg.provider !== 'lmx') {
|
|
141
|
+
if (!warnedReasoning.has(cfg.provider)) {
|
|
142
|
+
warnedReasoning.add(cfg.provider);
|
|
143
|
+
log.warn(`AI: reasoning '${opts.reasoning}' is honoured only by lmx engines — `
|
|
144
|
+
+ `${cfg.label || cfg.provider} sends no reasoning flag and runs on its own default `
|
|
145
|
+
+ '(warned once per provider).');
|
|
146
|
+
}
|
|
147
|
+
const { reasoning: _ignored, ...rest } = opts;
|
|
148
|
+
callOpts = rest;
|
|
149
|
+
}
|
|
150
|
+
|
|
151
|
+
const result = await withOneRetry(() => adapter.complete(cfg, callOpts), opts);
|
|
152
|
+
|
|
153
|
+
// ...and the counter after it, AWAITED: two calls in quick succession must both be counted
|
|
154
|
+
// before the second's ceiling check reads the total, or the limit is enforced against a stale one.
|
|
155
|
+
await meter.record(cfg, result, { ticketId: opts.ticketId });
|
|
156
|
+
|
|
157
|
+
log.info(`AI: ${cfg.provider}/${result.model} ${result.usage.total} tokens in ${result.ms}ms`);
|
|
158
|
+
return result;
|
|
159
|
+
}
|
|
160
|
+
|
|
161
|
+
/**
|
|
162
|
+
* Embed one or more strings, with the same retry policy as complete() — a fast dropped connection
|
|
163
|
+
* is retried once; a 429 fails fast (0.25.0, see withOneRetry). Embeddings are idempotent, so the
|
|
164
|
+
* one resend is always safe.
|
|
165
|
+
*
|
|
166
|
+
* THE CONFIG IS THE CALLER'S. An app resolves its embedding settings separately from completion
|
|
167
|
+
* (embeddingModel, often a different endpoint), so this never falls back to resolveConfig().
|
|
168
|
+
* No budget metering - embeddings were never metered, and starting is a separate decision.
|
|
169
|
+
* @param {object} cfg the resolved embedding config (provider, embeddingModel, …)
|
|
170
|
+
* @param {string|string[]} texts
|
|
171
|
+
* @param {{signal?: AbortSignal}} [opts]
|
|
172
|
+
*/
|
|
173
|
+
async function embed(cfg, texts, opts = {}) {
|
|
174
|
+
// EXPLICITLY DISABLED is refused, as complete() refuses it (the review found embed ignored it).
|
|
175
|
+
// Strictly `false`: an embedding config the caller built without the flag is still honoured.
|
|
176
|
+
if (cfg && cfg.enabled === false) {
|
|
177
|
+
throw new AiError('disabled', 'This endpoint is switched off.');
|
|
178
|
+
}
|
|
179
|
+
const adapter = cfg && PROVIDERS[cfg.provider];
|
|
180
|
+
if (!adapter) {
|
|
181
|
+
throw new AiError('unconfigured', `No adapter is registered for provider "${cfg && cfg.provider}".`);
|
|
182
|
+
}
|
|
183
|
+
if (typeof adapter.embed !== 'function') {
|
|
184
|
+
throw new AiError('unsupported', `${cfg.label || cfg.provider} does not support embeddings.`);
|
|
185
|
+
}
|
|
186
|
+
return withOneRetry(() => adapter.embed(cfg, texts), opts);
|
|
187
|
+
}
|
|
188
|
+
|
|
189
|
+
/** Is there a provider configured at all? Callers use this to decide whether to render a control. */
|
|
190
|
+
async function isEnabled() {
|
|
191
|
+
return (await resolveConfig()).enabled;
|
|
192
|
+
}
|
|
193
|
+
|
|
194
|
+
/** A short round trip for an admin "Test connection". NEVER THROWS — it reports what is wrong. */
|
|
195
|
+
async function test(cfgOverride) {
|
|
196
|
+
const cfg = cfgOverride || await resolveConfig();
|
|
197
|
+
if (!cfg.enabled) return { ok: false, kind: 'disabled', error: 'No provider is selected.' };
|
|
198
|
+
try {
|
|
199
|
+
const r = await complete({
|
|
200
|
+
system: 'Reply with the single word: ready. Do not explain.',
|
|
201
|
+
messages: [{ role: 'user', content: 'ready?' }],
|
|
202
|
+
maxTokens: 512,
|
|
203
|
+
temperature: 0,
|
|
204
|
+
skipBudget: true // an admin diagnosing a provider must not be blocked by a full budget
|
|
205
|
+
}, cfg);
|
|
206
|
+
return {
|
|
207
|
+
ok: true, model: r.model, ms: r.ms, reply: (r.text || '').trim().slice(0, 60), usage: r.usage,
|
|
208
|
+
finishReason: r.finishReason || null, reasoned: !!r.reasonedFor
|
|
209
|
+
};
|
|
210
|
+
} catch (err) {
|
|
211
|
+
if (err instanceof AiError) return { ok: false, kind: err.kind, error: err.message, retryable: err.retryable };
|
|
212
|
+
return { ok: false, kind: 'bad_response', error: err.message };
|
|
213
|
+
}
|
|
214
|
+
}
|
|
215
|
+
|
|
216
|
+
/** What the server has loaded, for an admin model picker. Empty when it cannot say. */
|
|
217
|
+
async function listModels(cfgOverride) {
|
|
218
|
+
return (await listModelsResult(cfgOverride)).models;
|
|
219
|
+
}
|
|
220
|
+
|
|
221
|
+
/**
|
|
222
|
+
* The model list and WHY it is the length it is.
|
|
223
|
+
*
|
|
224
|
+
* Callers that only want names should use listModels(). An admin screen wants this one: an empty
|
|
225
|
+
* array on its own cannot tell "the server has nothing loaded" from "that address is not an
|
|
226
|
+
* OpenAI-compatible API root", and those need different fixes.
|
|
227
|
+
*/
|
|
228
|
+
async function listModelsResult(cfgOverride) {
|
|
229
|
+
const cfg = cfgOverride || await resolveConfig();
|
|
230
|
+
if (!cfg.enabled) {
|
|
231
|
+
return { ok: false, models: [], url: null, status: 0, error: 'No provider is selected.' };
|
|
232
|
+
}
|
|
233
|
+
const adapter = PROVIDERS[cfg.provider];
|
|
234
|
+
if (!adapter) {
|
|
235
|
+
return { ok: false, models: [], url: null, status: 0, error: `Unknown provider “${cfg.provider}”.` };
|
|
236
|
+
}
|
|
237
|
+
try {
|
|
238
|
+
if (adapter.listModelsResult) return await adapter.listModelsResult(cfg);
|
|
239
|
+
// An adapter that predates this contract still works; it simply cannot explain itself.
|
|
240
|
+
return { ok: true, models: await adapter.listModels(cfg), url: null, status: 0, error: null };
|
|
241
|
+
} catch (err) {
|
|
242
|
+
return { ok: false, models: [], url: null, status: 0, error: err.message };
|
|
243
|
+
}
|
|
244
|
+
}
|
|
245
|
+
|
|
246
|
+
/**
|
|
247
|
+
* A fixed-workload speed test for one endpoint. See benchmark.js for why the workload is fixed
|
|
248
|
+
* rather than the prompt, and why the warm-up is reported rather than discarded.
|
|
249
|
+
*/
|
|
250
|
+
async function benchmarkEndpoint(cfgOverride, benchOpts) {
|
|
251
|
+
const cfg = cfgOverride || await resolveConfig();
|
|
252
|
+
return require('./benchmark').benchmark(complete, cfg, benchOpts || {});
|
|
253
|
+
}
|
|
254
|
+
|
|
255
|
+
// The writing-assist engines are bound to this client's complete() so a caller gets config +
|
|
256
|
+
// budget + retry for free. Prompt framing is supplied per call by the app (content).
|
|
257
|
+
const boundPolish = (opts) => polish(complete, opts);
|
|
258
|
+
const boundGenerate = (opts) => generate(complete, opts);
|
|
259
|
+
|
|
260
|
+
return {
|
|
261
|
+
complete, embed, isEnabled, test, listModels, listModelsResult, withOneRetry,
|
|
262
|
+
benchmark: benchmarkEndpoint,
|
|
263
|
+
polish: boundPolish, generate: boundGenerate,
|
|
264
|
+
facts, AiError, PROVIDERS, DEFAULTS
|
|
265
|
+
};
|
|
266
|
+
}
|
|
267
|
+
|
|
268
|
+
module.exports = {
|
|
269
|
+
// The usage counter behind every ceiling. LAZY: it needs the db-worker driver contract, which
|
|
270
|
+
// is an OPTIONAL peer — a consumer using only createAiClient/polish/facts must not be made to
|
|
271
|
+
// install a database package to require this one.
|
|
272
|
+
get createUsageStore() { return require('./usageStore').createUsageStore; },
|
|
273
|
+
get createProviderStore() { return require('./providerStore').createProviderStore; },
|
|
274
|
+
// Speed history. Lazy for the same reason as the others: it needs the db-worker driver contract,
|
|
275
|
+
// which is an optional peer.
|
|
276
|
+
get createSpeedStore() { return require('./speedStore').createSpeedStore; },
|
|
277
|
+
// EXPORTED SO A CONSUMER CAN ASSERT ITS TABLE MATCHES. An app writes its own migration, which is
|
|
278
|
+
// a hand copy of this DDL — and a copy with nothing comparing it to the original is the failure
|
|
279
|
+
// mode this repo has already documented twice. providerSchemaFor and usageSchemaFor exist for the
|
|
280
|
+
// same reason; leaving this one out meant a drift would surface as an INSERT throwing at runtime.
|
|
281
|
+
get speedSchemaFor() { return require('./speedStore').schemaFor; },
|
|
282
|
+
// No database behind health, so it loads eagerly like the rest of the seam.
|
|
283
|
+
...require('./health'),
|
|
284
|
+
/**
|
|
285
|
+
* Where this package's EJS partials live, for the consumer's view-roots list.
|
|
286
|
+
*
|
|
287
|
+
* Same contract as backup/server/notify/uploads: the package knows its own layout, the app
|
|
288
|
+
* puts its own views FIRST so a local file of the same name wins.
|
|
289
|
+
*/
|
|
290
|
+
viewsDir: require('path').join(__dirname, 'views'),
|
|
291
|
+
get providerSchemaFor() { return require('./providerStore').schemaFor; },
|
|
292
|
+
// ── SUPERVISED STACKS (lmx) ─────────────────────────────────────────────────────────────────
|
|
293
|
+
// The store is LAZY for the same reason as the others: it needs the db-worker driver contract,
|
|
294
|
+
// which is an optional peer. A consumer using only createAiClient must not be made to install a
|
|
295
|
+
// database package to require this one.
|
|
296
|
+
get createLmxStore() { return require('./lmxStore').createLmxStore; },
|
|
297
|
+
// EXPORTED SO A CONSUMER CAN ASSERT ITS TABLE MATCHES — the migration is a hand copy of this, and
|
|
298
|
+
// a copy with nothing comparing it to the original is the failure this repo has documented three
|
|
299
|
+
// times now. The first app to carry the table wrote it with nothing to check against.
|
|
300
|
+
get lmxSchemaFor() { return require('./lmxStore').schemaFor; },
|
|
301
|
+
// Does this stack actually work? Four checks in the only order they can run. No database and no
|
|
302
|
+
// keystore — every credential arrives as an argument — so it loads eagerly.
|
|
303
|
+
lmxVerify: require('./lmxVerify'),
|
|
304
|
+
// Engine identity - (instance, id), falling back to name - shared by the router, the stack panel
|
|
305
|
+
// and the verifier. Apps used to deep-require providers/lmxDiscovery for it; that path still
|
|
306
|
+
// works, but this is the supported one.
|
|
307
|
+
findEngine: require('./providers/lmxDiscovery').findEngine,
|
|
308
|
+
// The counting rules a screen needs. Separate from the verifier because "what did the stack say"
|
|
309
|
+
// and "what does that mean for what I am relying on" are different questions, and only the second
|
|
310
|
+
// one needs to know which engines this app has adopted.
|
|
311
|
+
...require('./lmxStatus'),
|
|
312
|
+
get usageSchemaFor() { return require('./usageStore').schemaFor; },
|
|
313
|
+
createAiClient,
|
|
314
|
+
PROVIDERS, DEFAULTS,
|
|
315
|
+
AiError, fromFetchFailure, redact,
|
|
316
|
+
facts,
|
|
317
|
+
// THE INPUT SIDE of the same concern facts.js covers on the output side: text somebody else
|
|
318
|
+
// wrote, placed where a model can read it without being able to give orders. See
|
|
319
|
+
// untrusted.js for why the fence marker has to be generated per call.
|
|
320
|
+
untrusted: require('./untrusted'),
|
|
321
|
+
// ...and the correct ASSEMBLY of it. untrusted.js hands over `fence()` and `rule()` separately;
|
|
322
|
+
// fenced.js puts them together, because one app assembled them two different ways in two prompt
|
|
323
|
+
// builders and one of the divergences started refusing correct answers. See fenced.js for the
|
|
324
|
+
// measurements, including why a fenced document is worth about a third of the defence unless the
|
|
325
|
+
// caller also states its schema's field names.
|
|
326
|
+
fenced: require('./fenced'),
|
|
327
|
+
// Default writing-op catalogues, so an app can build its menus without re-declaring them.
|
|
328
|
+
POLISH_MODES: require('./polish').MODES,
|
|
329
|
+
POLISH_TONES: require('./polish').TONES,
|
|
330
|
+
// The values complete({ reasoning }) accepts (0.26.0), for a caller that validates its own config.
|
|
331
|
+
REASONING_MODES,
|
|
332
|
+
RETRY_AFTER_MS
|
|
333
|
+
};
|
package/lmxStatus.js
CHANGED
|
@@ -50,7 +50,7 @@ function reasoningLabel(model) {
|
|
|
50
50
|
const flag = reasoningFor(model);
|
|
51
51
|
if (!flag) return null;
|
|
52
52
|
if (flag.reasoning_effort) return `reasoning_effort: ${flag.reasoning_effort}`;
|
|
53
|
-
return
|
|
53
|
+
return `enable_thinking: ${flag.chat_template_kwargs.enable_thinking}`;
|
|
54
54
|
}
|
|
55
55
|
|
|
56
56
|
/**
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@aria-framework/ai",
|
|
3
3
|
"description": "Aria App Framework \u2014 AI module. A dependency-injected model seam (createAiClient) over several providers (LM Studio / OpenAI-compatible / Anthropic), with a fact-preservation guard, generic Polish and Generate writing engines, and a browser polish widget. Prompts and config stay in the consuming app.",
|
|
4
|
-
"version": "0.
|
|
4
|
+
"version": "0.26.0",
|
|
5
5
|
"license": "UNLICENSED",
|
|
6
6
|
"private": false,
|
|
7
7
|
"publishConfig": {
|
|
@@ -47,7 +47,7 @@
|
|
|
47
47
|
}
|
|
48
48
|
},
|
|
49
49
|
"scripts": {
|
|
50
|
-
"test": "node test/smoke.js && node test/usageStore.js && node test/providerStore.js && node test/speedStore.js && node test/health.js && node test/listModels.js && node test/benchmark.js && node test/lmxDiscovery.js && node test/lmx.js && node test/lmxVerify.js && node test/lmxStore.js && node test/lmxStatus.js && node test/jobCard.js && node test/untrusted.js && node test/packaging.js && node test/views.js && node test/fenced.js && node test/polish.js",
|
|
50
|
+
"test": "node test/smoke.js && node test/usageStore.js && node test/providerStore.js && node test/speedStore.js && node test/health.js && node test/listModels.js && node test/benchmark.js && node test/lmxDiscovery.js && node test/lmx.js && node test/lmxVerify.js && node test/lmxStore.js && node test/lmxStatus.js && node test/jobCard.js && node test/untrusted.js && node test/packaging.js && node test/views.js && node test/fenced.js && node test/polish.js && node test/reasoning.js",
|
|
51
51
|
"prepublishOnly": "node ../../test/packaging.js ai"
|
|
52
52
|
},
|
|
53
53
|
"devDependencies": {
|
package/providers/lmx.js
CHANGED
|
@@ -149,12 +149,43 @@ function skipError(name, res) {
|
|
|
149
149
|
* An UNRECOGNISED family gets NEITHER flag. Guessing would be worse than not guessing — sending a
|
|
150
150
|
* Qwen argument to a model that ignores it wastes the budget silently, which is the failure this
|
|
151
151
|
* exists to avoid.
|
|
152
|
+
*
|
|
153
|
+
* THE DEFAULT IS "THINK AS LITTLE AS THE FAMILY ALLOWS"; `reasoning: 'full'` (0.26.0) is the one
|
|
154
|
+
* named way to ask for the opposite — Qwen thinking on, gpt-oss effort high — for judgement work
|
|
155
|
+
* (triage) where a thinking-off answer measured as a different product. It is a NAMED per-call
|
|
156
|
+
* option rather than a raw `extra`, because the forced flag deliberately outranks `extra`: a stray
|
|
157
|
+
* passthrough must never be able to turn thinking on and spend a budget sized for a quick answer.
|
|
158
|
+
*/
|
|
159
|
+
const REASONING_MODES = Object.freeze(['full']);
|
|
160
|
+
|
|
161
|
+
/**
|
|
162
|
+
* Refuse a `reasoning` value this package does not define. `undefined`/`null` mean the default.
|
|
163
|
+
* Anything else — 'high', 'Full', true, '' — is a caller's mistake, and silently treating it as the
|
|
164
|
+
* default would hand a triage caller a thinking-off answer while it believed it had asked for more.
|
|
165
|
+
*/
|
|
166
|
+
function assertReasoning(value) {
|
|
167
|
+
if (value === undefined || value === null) return;
|
|
168
|
+
if (!REASONING_MODES.includes(value)) {
|
|
169
|
+
throw new AiError('refused',
|
|
170
|
+
`Unknown reasoning option ${JSON.stringify(value)}. The only value is `
|
|
171
|
+
+ `${REASONING_MODES.map((m) => `'${m}'`).join(', ')}; leave it out for the default.`);
|
|
172
|
+
}
|
|
173
|
+
}
|
|
174
|
+
|
|
175
|
+
/**
|
|
176
|
+
* The families with a known flag, in match order. A TABLE rather than a chain of ifs so the README
|
|
177
|
+
* test can enumerate it: a family added here and not documented fails that test.
|
|
152
178
|
*/
|
|
153
|
-
|
|
179
|
+
const REASONING_FAMILIES = Object.freeze([
|
|
180
|
+
{ family: 'qwen', match: /qwen/, flag: (full) => ({ chat_template_kwargs: { enable_thinking: full } }) },
|
|
181
|
+
{ family: 'gpt-oss', match: /gpt-oss/, flag: (full) => ({ reasoning_effort: full ? 'high' : 'low' }) }
|
|
182
|
+
]);
|
|
183
|
+
|
|
184
|
+
function reasoningFor(modelPath, mode) {
|
|
185
|
+
assertReasoning(mode);
|
|
154
186
|
const m = String(modelPath || '').toLowerCase();
|
|
155
|
-
|
|
156
|
-
|
|
157
|
-
return null;
|
|
187
|
+
const f = REASONING_FAMILIES.find((x) => x.match.test(m));
|
|
188
|
+
return f ? f.flag(mode === 'full') : null;
|
|
158
189
|
}
|
|
159
190
|
|
|
160
191
|
/** The config openai-compatible needs, once the address is known. */
|
|
@@ -204,23 +235,31 @@ function tagThrottle(err) {
|
|
|
204
235
|
}
|
|
205
236
|
|
|
206
237
|
async function complete(cfg, opts) {
|
|
238
|
+
// Validated BEFORE discovery, so a typo is refused without a status read or an engine call.
|
|
239
|
+
const mode = opts && opts.reasoning;
|
|
240
|
+
assertReasoning(mode);
|
|
207
241
|
const d = discoveryFor(cfg.lmx, cfg.logger);
|
|
208
242
|
const name = cfg.lmx.engine;
|
|
209
243
|
const res = await resolveWarm(d, refOf(cfg.lmx));
|
|
210
244
|
if (!res.url) throw skipError(name, res);
|
|
211
245
|
|
|
212
|
-
const reasoning = reasoningFor(res.engine.model);
|
|
246
|
+
const reasoning = reasoningFor(res.engine.model, mode);
|
|
213
247
|
if (!reasoning && cfg.logger) {
|
|
214
|
-
cfg.logger.warn(
|
|
215
|
-
`lmx:
|
|
216
|
-
|
|
217
|
-
|
|
248
|
+
cfg.logger.warn(mode === 'full'
|
|
249
|
+
? `lmx: reasoning '${mode}' was asked for, but no reasoning flag is known for model `
|
|
250
|
+
+ `"${res.engine.model}" on engine "${name}" — sending none; the model runs on its own default.`
|
|
251
|
+
: `lmx: no reasoning flag known for model "${res.engine.model}" on engine "${name}" — sending `
|
|
252
|
+
+ 'none. If it reasons before answering, the token budget may be consumed with no content '
|
|
253
|
+
+ 'returned.'
|
|
218
254
|
);
|
|
219
255
|
}
|
|
220
256
|
|
|
257
|
+
// The named option is consumed here; openai-compatible never sees it.
|
|
258
|
+
const { reasoning: _consumed, ...rest } = opts || {};
|
|
221
259
|
return openai.complete(engineConfig(cfg, res.url, res.engine), {
|
|
222
|
-
...
|
|
223
|
-
// Merged rather than replacing: a caller's own extras survive.
|
|
260
|
+
...rest,
|
|
261
|
+
// Merged rather than replacing: a caller's own extras survive. The flag is spread LAST, so a
|
|
262
|
+
// raw extra can never override it — only the named `reasoning` option changes what it says.
|
|
224
263
|
extra: { ...(opts && opts.extra), ...(reasoning || {}) }
|
|
225
264
|
}).catch((err) => { throw tagThrottle(err); });
|
|
226
265
|
}
|
|
@@ -289,5 +328,6 @@ function _resetRegistry() {
|
|
|
289
328
|
module.exports = {
|
|
290
329
|
complete, embed, listModels, listModelsResult,
|
|
291
330
|
apiRoot: openai.apiRoot,
|
|
292
|
-
reasoningFor,
|
|
331
|
+
reasoningFor, assertReasoning, REASONING_MODES, REASONING_FAMILIES,
|
|
332
|
+
discoveryFor, forgetDiscovery, _resetRegistry, SKIP_CODE
|
|
293
333
|
};
|
|
@@ -154,8 +154,17 @@ async function complete(cfg, opts) {
|
|
|
154
154
|
// that thinking back separately (`reasoning_content`) or inline, wrapped in <think> tags. Neither
|
|
155
155
|
// is the answer, and both have to be recognised — otherwise a model that reasoned and then ran
|
|
156
156
|
// out of room looks identical to a broken server.
|
|
157
|
-
|
|
158
|
-
|
|
157
|
+
let reasoning = message.reasoning_content || message.reasoning || '';
|
|
158
|
+
let text = stripThinking(message.content || '');
|
|
159
|
+
|
|
160
|
+
// A LENGTH STOP INSIDE AN UNCLOSED <think> IS NO ANSWER (0.26.0). stripThinking only removes a
|
|
161
|
+
// CLOSED block, so a model that ran out of room mid-thought used to resolve with its half-finished
|
|
162
|
+
// reasoning as the "answer" (truncated: true). Narrow on purpose: only a reply that STARTS with
|
|
163
|
+
// the tag, has no closing tag, and stopped on length — the case reasoning: 'full' makes common.
|
|
164
|
+
if (finishReason === 'length' && unclosedThinking(text)) {
|
|
165
|
+
reasoning = reasoning || text;
|
|
166
|
+
text = '';
|
|
167
|
+
}
|
|
159
168
|
|
|
160
169
|
if (!text) {
|
|
161
170
|
throw emptyCompletion({ label, finishReason, reasoning, maxTokens: body.max_tokens, usage: payload.usage });
|
|
@@ -214,6 +223,12 @@ function stripThinking(content) {
|
|
|
214
223
|
.trim();
|
|
215
224
|
}
|
|
216
225
|
|
|
226
|
+
/** A reply that opens a <think>/<thinking> block and never closes it. */
|
|
227
|
+
function unclosedThinking(text) {
|
|
228
|
+
const m = /^<(think|thinking)>/i.exec(String(text || '').trimStart());
|
|
229
|
+
return !!m && !new RegExp(`</${m[1]}>`, 'i').test(text);
|
|
230
|
+
}
|
|
231
|
+
|
|
217
232
|
/**
|
|
218
233
|
* Why nothing came back — the message an operator can act on.
|
|
219
234
|
*
|
|
@@ -227,7 +242,8 @@ function emptyCompletion({ label, finishReason, reasoning, maxTokens, usage }) {
|
|
|
227
242
|
return new AiError('bad_response',
|
|
228
243
|
`${label} ran out of room before it answered — it used all ${maxTokens} reply tokens` +
|
|
229
244
|
(reasoning ? ' on internal reasoning' : '') +
|
|
230
|
-
'. This model thinks before it replies, so raise the reply limit (1024 is a sensible floor
|
|
245
|
+
'. This model thinks before it replies, so raise the reply limit (1024 is a sensible floor; ' +
|
|
246
|
+
"4096 or more with reasoning: 'full').",
|
|
231
247
|
{ usage });
|
|
232
248
|
}
|
|
233
249
|
if (reasoning) {
|