@aria-framework/ai 0.27.0 → 0.27.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +103 -0
- package/README.md +15 -2
- package/error.js +5 -1
- package/index.js +6 -2
- package/package.json +3 -2
- package/providers/anthropic.js +9 -3
- package/providers/openai-compatible.js +13 -8
- package/reasoning.js +55 -8
package/CHANGELOG.md
ADDED
|
@@ -0,0 +1,103 @@
|
|
|
1
|
+
# @aria-framework/ai changelog
|
|
2
|
+
|
|
3
|
+
This file starts at 0.26.1. Earlier versions are described in the git history only.
|
|
4
|
+
|
|
5
|
+
Each entry lists the **silent default changes** first: things that change what a caller's existing
|
|
6
|
+
code sends, or what it is billed for, without the caller changing anything.
|
|
7
|
+
|
|
8
|
+
## 0.27.1
|
|
9
|
+
|
|
10
|
+
Fixes from an independent review of 0.26.1 and 0.27.0.
|
|
11
|
+
|
|
12
|
+
### Silent default changes
|
|
13
|
+
|
|
14
|
+
- **A failed call is recorded under the model that ran.** `budget.record(cfg, { usage, model, ms,
|
|
15
|
+
failed: true })` now receives the model the server named in its reply (`err.model`). It used to
|
|
16
|
+
receive `cfg.model`, which is empty on lmx, so every failed lmx call was ledgered with no model. The
|
|
17
|
+
configured model is used only when the reply named none. A usage table keyed on model will see
|
|
18
|
+
failed lmx calls move from `''` to the engine's model.
|
|
19
|
+
- **Older Claude models: `'full'` never asks for more output than the model allows.** On the
|
|
20
|
+
pre-4.6 budget path, `max_tokens` was the caller's `maxTokens` plus a thinking budget of the same
|
|
21
|
+
size, about twice what was asked. That could pass the model's output limit (Haiku 4.5, Sonnet 4.5,
|
|
22
|
+
Opus 4.5, Sonnet 4 and 3.7 Sonnet allow 64000; Opus 4.1 and Opus 4 allow 32000), and the API
|
|
23
|
+
refused it. It is now capped at the limit. Near the limit the thinking budget shrinks first, down
|
|
24
|
+
to the API's minimum of 1024. For example, `maxTokens: 40000` on Haiku 4.5 sends
|
|
25
|
+
`max_tokens: 64000` with a 24000 budget. When `maxTokens` plus the budget is below the limit,
|
|
26
|
+
nothing changes.
|
|
27
|
+
- **Claude 3 models without extended thinking get no thinking.** On `'full'`, Claude 3.5 Haiku,
|
|
28
|
+
3 Haiku, 3 Opus and 3.5 Sonnet were sent `thinking: enabled`, which they reject with a 400. They now
|
|
29
|
+
keep the pre-0.27 request, and `'full'` logs a warning instead of failing the call, as it does
|
|
30
|
+
for a model the table does not know. 3.7 Sonnet still thinks.
|
|
31
|
+
- **Opus 4 and Sonnet 4 under their `-0` aliases and dated ids** (`claude-opus-4-0`,
|
|
32
|
+
`claude-opus-4-20250514`, `claude-sonnet-4-0`, `claude-sonnet-4-20250514`) are now recognised.
|
|
33
|
+
0.27.0 treated them as unknown: `'full'` sent no thinking and logged a warning. They now get the
|
|
34
|
+
budget like the other pre-4.6 models.
|
|
35
|
+
|
|
36
|
+
### Fixes
|
|
37
|
+
|
|
38
|
+
- If a budget's `record` throws `null` (or anything without a `message`), logging it no longer
|
|
39
|
+
throws a `TypeError` that replaces the caller's real error.
|
|
40
|
+
- `AiError` takes an optional `model`, set by both adapters on every error that carries `usage`.
|
|
41
|
+
On lmx, when the reply names no model, it is the model the engine reports running; the same
|
|
42
|
+
fallback now applies to a successful result's `model`.
|
|
43
|
+
- **Anthropic: a `maxTokens` that is not a positive number is refused** (`AiError` kind `refused`,
|
|
44
|
+
before any request) on every Anthropic call. A fraction is rounded down. Before this, a
|
|
45
|
+
string `"2048"` on the budget path became `max_tokens: 64000`, and a negative value sent a
|
|
46
|
+
`budget_tokens` above `max_tokens`. `0`, `NaN` and a missing value still take the 1024 default.
|
|
47
|
+
|
|
48
|
+
### Tests
|
|
49
|
+
|
|
50
|
+
- `test/anthropicRequests.js` (new) compares the request bodies with a fixture typed by hand from
|
|
51
|
+
the Messages API rules for 26 model ids, rather than with values read back from
|
|
52
|
+
`ANTHROPIC_FAMILIES`. It also checks the budget at 1024, 4096, 40000 and at each model's limit.
|
|
53
|
+
- `test/usageOnFailure.js` adds an lmx failure with an empty `cfg.model`, an Anthropic failure whose
|
|
54
|
+
reply names a different model, an empty completion with zero usage on both adapters, and a budget
|
|
55
|
+
that throws `null`.
|
|
56
|
+
|
|
57
|
+
## 0.27.0
|
|
58
|
+
|
|
59
|
+
`complete({ reasoning: 'off' | 'full' })` is honoured by every provider. Before this, only lmx
|
|
60
|
+
honoured it. Leaving it out means `'off'`.
|
|
61
|
+
|
|
62
|
+
### Silent default changes
|
|
63
|
+
|
|
64
|
+
These apply to callers that do not pass `reasoning`. To keep a model thinking, pass
|
|
65
|
+
`reasoning: 'full'`.
|
|
66
|
+
|
|
67
|
+
- **openai-compatible and lmstudio, qwen or gpt-oss models: a body flag is now sent on every call.**
|
|
68
|
+
qwen gets `chat_template_kwargs: { enable_thinking: false }` and gpt-oss gets
|
|
69
|
+
`reasoning_effort: 'low'`. Before this, these models used the server's default, which usually means
|
|
70
|
+
thinking on every call.
|
|
71
|
+
- **LM Studio (`provider: 'lmstudio'`), qwen models: `/no_think` is appended to the system prompt**
|
|
72
|
+
(or sent as the system prompt when there is none), and `reasoning_effort: 'none'` is added to the
|
|
73
|
+
body. LM Studio ignores `chat_template_kwargs`. A caller that compares or caches system prompts will
|
|
74
|
+
see the extra line. An LM Studio server configured as `openai-compatible` gets neither, and keeps
|
|
75
|
+
thinking.
|
|
76
|
+
- **Claude Opus 5.5 and Fable (and Mythos) get `output_config.effort: 'low'`.** These models cannot
|
|
77
|
+
turn thinking off, and Opus 5.5's own default effort is `medium`, so leaving `reasoning` out now
|
|
78
|
+
means less thinking than before.
|
|
79
|
+
- **Other Claude generations:** Opus 5 gets effort `'low'` (its own default is adaptive thinking at
|
|
80
|
+
`high`). Sonnet 5.5 gets `thinking: { type: 'between_tools' }` and Sonnet 5 gets
|
|
81
|
+
`thinking: { type: 'disabled' }`; before this, both thought by default. 4.6 and older are
|
|
82
|
+
unchanged unless `'full'` is passed.
|
|
83
|
+
- **`temperature` is no longer sent to Claude models that reject it** (Opus 4.7 and later, Sonnet 5
|
|
84
|
+
and later, Fable). Before this, those requests failed with a 400. It is also left out with `'full'`
|
|
85
|
+
on the pre-4.6 budget path, where extended thinking requires the default.
|
|
86
|
+
- The client no longer strips `reasoning` with a warning saying only lmx honours it.
|
|
87
|
+
|
|
88
|
+
## 0.26.1
|
|
89
|
+
|
|
90
|
+
### Silent default changes
|
|
91
|
+
|
|
92
|
+
- **A failed call that spent tokens is recorded through `budget.record`.** When the model answered
|
|
93
|
+
but the answer was unusable (it thought until the limit, or a structured reply was cut off),
|
|
94
|
+
`budget.record(cfg, { usage, model, ms, failed: true }, meta)` is called before the error is
|
|
95
|
+
thrown. A failure that spent nothing (refused, unreachable, timed out, or a zero usage total)
|
|
96
|
+
records nothing. **A caller that metered `err.usage` itself must stop, or the call is counted
|
|
97
|
+
twice.** A `record` that throws is logged and never replaces the original error.
|
|
98
|
+
|
|
99
|
+
### Fixes
|
|
100
|
+
|
|
101
|
+
- Both adapters put normalised usage (`{ prompt, completion, total }`) on every error raised after
|
|
102
|
+
the reply was read.
|
|
103
|
+
- A 200 whose JSON body is `null` is an `AiError` of kind `bad_response`. It used to be a `TypeError`.
|
package/README.md
CHANGED
|
@@ -48,7 +48,10 @@ an admin screen and never includes a URL or key.
|
|
|
48
48
|
was unusable (it thought until the limit, or a structured reply was cut off), the error carries
|
|
49
49
|
`usage` (`{ prompt, completion, total }`) and `budget.record(cfg, { usage, model, ms, failed: true }, meta)`
|
|
50
50
|
is called before the error is thrown. A failure that spent nothing (refused, unreachable, timeout)
|
|
51
|
-
records nothing. If `record` throws, the original error still reaches the caller.
|
|
51
|
+
records nothing. If `record` throws, the original error still reaches the caller. The error also
|
|
52
|
+
carries `model`, the model the server says answered, and `record` receives that (0.27.1); it falls
|
|
53
|
+
back to the configured model only when the reply named none. On lmx the configured model is empty.
|
|
54
|
+
A caller that metered `err.usage` itself before 0.26.1 must stop, or the call is counted twice.
|
|
52
55
|
|
|
53
56
|
`providerStore`, `usageStore`, `speedStore` and `lmxStore` are optional db-worker-backed
|
|
54
57
|
stores. Each exports a `schemaFor(dialect)`, so the app's migration can be checked against it.
|
|
@@ -236,11 +239,21 @@ and Fable cannot turn thinking off (effort is the only lever), Sonnet 5.5 turns
|
|
|
236
239
|
| claude-sonnet-5 | `{"thinking":{"type":"disabled"}}` | `{"thinking":{"type":"adaptive"},"output_config":{"effort":"high"}}` | not sent |
|
|
237
240
|
| claude-opus-4-7 / 4-8 | nothing | `{"thinking":{"type":"adaptive"},"output_config":{"effort":"high"}}` | not sent |
|
|
238
241
|
| claude-*-4-6 | nothing | `{"thinking":{"type":"adaptive"},"output_config":{"effort":"high"}}` | sent |
|
|
239
|
-
| claude-*-4-5
|
|
242
|
+
| claude-*-4-5 | nothing | thinking enabled with a budget of `maxTokens` (at least 1024), added on top of `max_tokens`, which never exceeds 64000 | sent; not sent with `'full'` |
|
|
243
|
+
| claude-opus-4-1 / claude-opus-4 | nothing | thinking enabled with a budget of `maxTokens` (at least 1024), added on top of `max_tokens`, which never exceeds 32000 | sent; not sent with `'full'` |
|
|
244
|
+
| claude-sonnet-4 | nothing | thinking enabled with a budget of `maxTokens` (at least 1024), added on top of `max_tokens`, which never exceeds 64000 | sent; not sent with `'full'` |
|
|
245
|
+
| claude-3-7-sonnet | nothing | thinking enabled with a budget of `maxTokens` (at least 1024), added on top of `max_tokens`, which never exceeds 64000 | sent; not sent with `'full'` |
|
|
246
|
+
| claude-3 without extended thinking | nothing | nothing (a warning that `'full'` could not be honoured) | sent |
|
|
240
247
|
| anything else | nothing | nothing (a warning that `'full'` could not be honoured) | sent |
|
|
241
248
|
|
|
242
249
|
`effort` is merged into any `output_config` the structured-output schema already put there.
|
|
243
250
|
|
|
251
|
+
On the budget rows `max_tokens` counts the thinking as well as the answer, so it is capped at the
|
|
252
|
+
model's output limit (0.27.1). Near the limit the thinking budget shrinks first, down to the API's
|
|
253
|
+
minimum of 1024, and only then the answer's room: `maxTokens: 40000` on Haiku 4.5 sends
|
|
254
|
+
`max_tokens: 64000` with a 24000 budget. Claude 3 models other than 3.7 Sonnet have no extended
|
|
255
|
+
thinking, so `'full'` sends nothing to them and logs a warning; it does not fail the call.
|
|
256
|
+
|
|
244
257
|
Rules:
|
|
245
258
|
|
|
246
259
|
- The values are `'off'` and `'full'`. Leaving the option out (or `null`) means `'off'`. Any other
|
package/error.js
CHANGED
|
@@ -18,7 +18,7 @@ class AiError extends Error {
|
|
|
18
18
|
/**
|
|
19
19
|
* @param {AiErrorKind} kind
|
|
20
20
|
* @param {string} message written for the admin screen, not the log
|
|
21
|
-
* @param {{cause?: Error, status?: number, usage?: object}} [opts]
|
|
21
|
+
* @param {{cause?: Error, status?: number, usage?: object, model?: string}} [opts]
|
|
22
22
|
*/
|
|
23
23
|
constructor(kind, message, opts = {}) {
|
|
24
24
|
super(message);
|
|
@@ -30,6 +30,10 @@ class AiError extends Error {
|
|
|
30
30
|
// measuring throughput should not have to treat that as zero work. Optional everywhere: most
|
|
31
31
|
// failures happen before a single token is spent.
|
|
32
32
|
if (opts.usage) this.usage = opts.usage;
|
|
33
|
+
// ...AND WHICH MODEL SPENT IT (0.27.1), as the server named it in its reply. The client meters a
|
|
34
|
+
// failed call under this, not cfg.model: on lmx cfg.model is empty (the engine decides what
|
|
35
|
+
// runs), and an alias on any provider resolves to an id other than the one configured.
|
|
36
|
+
if (opts.model) this.model = opts.model;
|
|
33
37
|
if (opts.cause) this.cause = opts.cause;
|
|
34
38
|
}
|
|
35
39
|
|
package/index.js
CHANGED
|
@@ -145,12 +145,16 @@ function createAiClient(deps = {}) {
|
|
|
145
145
|
// connection test on such a model run outside the ceiling. Adapters put normalised usage on
|
|
146
146
|
// the error; a failure that spent nothing (refused, unreachable) carries none and is not counted.
|
|
147
147
|
// The ledger failing must never replace the error the caller needs to see.
|
|
148
|
+
// THE MODEL THAT RAN (0.27.1), as the adapter read it off the reply — cfg.model only when the
|
|
149
|
+
// reply named none. On lmx cfg.model is '' and every failed call was ledgered under no model.
|
|
148
150
|
if (err && err.usage && err.usage.total > 0) {
|
|
149
151
|
try {
|
|
150
|
-
await meter.record(cfg,
|
|
152
|
+
await meter.record(cfg,
|
|
153
|
+
{ usage: err.usage, model: err.model || cfg.model, ms: Date.now() - started, failed: true },
|
|
151
154
|
{ ticketId: opts.ticketId });
|
|
152
155
|
} catch (recErr) {
|
|
153
|
-
|
|
156
|
+
// A budget may throw anything, null included; logging it must not throw in its turn.
|
|
157
|
+
log.warn(`AI: usage of a failed call not recorded (${String(recErr?.message ?? recErr)})`);
|
|
154
158
|
}
|
|
155
159
|
}
|
|
156
160
|
throw err;
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@aria-framework/ai",
|
|
3
3
|
"description": "Aria App Framework — AI module. A dependency-injected model seam (createAiClient) over several providers (LM Studio / OpenAI-compatible / Anthropic), with a fact-preservation guard, generic Polish and Generate writing engines, and a browser polish widget. Prompts and config stay in the consuming app.",
|
|
4
|
-
"version": "0.27.
|
|
4
|
+
"version": "0.27.1",
|
|
5
5
|
"license": "UNLICENSED",
|
|
6
6
|
"private": false,
|
|
7
7
|
"publishConfig": {
|
|
@@ -9,6 +9,7 @@
|
|
|
9
9
|
},
|
|
10
10
|
"main": "index.js",
|
|
11
11
|
"files": [
|
|
12
|
+
"CHANGELOG.md",
|
|
12
13
|
"benchmark.js",
|
|
13
14
|
"browser/ai-jobs.js",
|
|
14
15
|
"browser/ai-panels.js",
|
|
@@ -48,7 +49,7 @@
|
|
|
48
49
|
}
|
|
49
50
|
},
|
|
50
51
|
"scripts": {
|
|
51
|
-
"test": "node test/smoke.js && node test/usageStore.js && node test/providerStore.js && node test/speedStore.js && node test/health.js && node test/listModels.js && node test/benchmark.js && node test/lmxDiscovery.js && node test/lmx.js && node test/lmxVerify.js && node test/lmxStore.js && node test/lmxStatus.js && node test/jobCard.js && node test/untrusted.js && node test/packaging.js && node test/views.js && node test/fenced.js && node test/polish.js && node test/reasoning.js && node test/reasoningProviders.js && node test/usageOnFailure.js",
|
|
52
|
+
"test": "node test/smoke.js && node test/usageStore.js && node test/providerStore.js && node test/speedStore.js && node test/health.js && node test/listModels.js && node test/benchmark.js && node test/lmxDiscovery.js && node test/lmx.js && node test/lmxVerify.js && node test/lmxStore.js && node test/lmxStatus.js && node test/jobCard.js && node test/untrusted.js && node test/packaging.js && node test/views.js && node test/fenced.js && node test/polish.js && node test/reasoning.js && node test/reasoningProviders.js && node test/anthropicRequests.js && node test/usageOnFailure.js",
|
|
52
53
|
"prepublishOnly": "node ../../test/packaging.js ai"
|
|
53
54
|
},
|
|
54
55
|
"devDependencies": {
|
package/providers/anthropic.js
CHANGED
|
@@ -72,6 +72,12 @@ async function complete(cfg, opts) {
|
|
|
72
72
|
// requires the default. Sending it would be a 400, so it is left out.
|
|
73
73
|
if (!r.sendTemperature) delete body.temperature;
|
|
74
74
|
body.max_tokens = r.maxTokens;
|
|
75
|
+
// A MODEL WITH NO EXTENDED THINKING (0.27.1): nothing is sent, and 'full' is said to be
|
|
76
|
+
// unhonoured — a warning, as for an unknown model, never an error.
|
|
77
|
+
if (r.cannotThink && cfg.logger) {
|
|
78
|
+
cfg.logger.warn(`${label}: reasoning 'full' was asked for, but model "${cfg.model}" has no extended `
|
|
79
|
+
+ 'thinking — sending no thinking setting; the model answers without it.');
|
|
80
|
+
}
|
|
75
81
|
} else if (opts.reasoning === 'full' && cfg.logger) {
|
|
76
82
|
cfg.logger.warn(`${label}: reasoning 'full' was asked for, but model "${cfg.model}" is not in the `
|
|
77
83
|
+ 'known Claude generations — sending no thinking setting; the model runs on its own default.');
|
|
@@ -133,7 +139,7 @@ async function complete(cfg, opts) {
|
|
|
133
139
|
const thought = (payload.content || []).some((b) => b && b.type === 'thinking');
|
|
134
140
|
throw new AiError('bad_response', thought
|
|
135
141
|
? `${label} returned only its internal reasoning. Raise the reply limit and try again.`
|
|
136
|
-
: `${label} returned no text (stop reason: ${payload.stop_reason || 'none given'}).`, { usage });
|
|
142
|
+
: `${label} returned no text (stop reason: ${payload.stop_reason || 'none given'}).`, { usage, model: payload.model });
|
|
137
143
|
}
|
|
138
144
|
|
|
139
145
|
let json = null;
|
|
@@ -145,10 +151,10 @@ async function complete(cfg, opts) {
|
|
|
145
151
|
if (payload.stop_reason === 'max_tokens') {
|
|
146
152
|
throw new AiError('bad_response',
|
|
147
153
|
`${label} ran out of room mid-answer — the structured reply was cut off before it was ` +
|
|
148
|
-
'finished. A Regenerate usually succeeds.', { cause: err, usage });
|
|
154
|
+
'finished. A Regenerate usually succeeds.', { cause: err, usage, model: payload.model });
|
|
149
155
|
}
|
|
150
156
|
throw new AiError('bad_response',
|
|
151
|
-
`${label} was asked for structured output and returned text that will not parse.`, { cause: err, usage });
|
|
157
|
+
`${label} was asked for structured output and returned text that will not parse.`, { cause: err, usage, model: payload.model });
|
|
152
158
|
}
|
|
153
159
|
}
|
|
154
160
|
|
|
@@ -164,6 +164,10 @@ async function complete(cfg, opts) {
|
|
|
164
164
|
throw new AiError('bad_response', `${label} answered with an empty body.`);
|
|
165
165
|
}
|
|
166
166
|
|
|
167
|
+
// THE MODEL THAT RAN (0.27.1): what the reply names, else — on lmx — the model the engine reports
|
|
168
|
+
// running (`reasoningModel`). Carried on errors for metering and on the result.
|
|
169
|
+
const ranModel = payload.model || cfg.reasoningModel;
|
|
170
|
+
|
|
167
171
|
if (payload.error && !payload.choices) {
|
|
168
172
|
const said = typeof payload.error === 'string'
|
|
169
173
|
? payload.error
|
|
@@ -198,7 +202,8 @@ async function complete(cfg, opts) {
|
|
|
198
202
|
}
|
|
199
203
|
|
|
200
204
|
if (!text) {
|
|
201
|
-
throw emptyCompletion({ label, finishReason, reasoning, maxTokens: body.max_tokens, usage: payload.usage
|
|
205
|
+
throw emptyCompletion({ label, finishReason, reasoning, maxTokens: body.max_tokens, usage: payload.usage,
|
|
206
|
+
model: ranModel });
|
|
202
207
|
}
|
|
203
208
|
|
|
204
209
|
// TOKENS WERE SPENT EVEN WHEN THE ANSWER IS UNUSABLE (0.26.1). Carried on every error raised after
|
|
@@ -223,17 +228,17 @@ async function complete(cfg, opts) {
|
|
|
223
228
|
throw new AiError('bad_response',
|
|
224
229
|
`${label} ran out of room mid-answer (${body.max_tokens} tokens) — the structured reply ` +
|
|
225
230
|
'was cut off before it was finished. If this repeats, the schema may be letting the ' +
|
|
226
|
-
'model ramble; a Regenerate usually succeeds.', { cause: err, usage });
|
|
231
|
+
'model ramble; a Regenerate usually succeeds.', { cause: err, usage, model: ranModel });
|
|
227
232
|
}
|
|
228
233
|
throw new AiError('bad_response',
|
|
229
|
-
`${label} was asked for structured output and returned text that will not parse.`, { cause: err, usage });
|
|
234
|
+
`${label} was asked for structured output and returned text that will not parse.`, { cause: err, usage, model: ranModel });
|
|
230
235
|
}
|
|
231
236
|
}
|
|
232
237
|
|
|
233
238
|
return {
|
|
234
239
|
text,
|
|
235
240
|
json,
|
|
236
|
-
model:
|
|
241
|
+
model: ranModel || cfg.model,
|
|
237
242
|
usage,
|
|
238
243
|
ms: Date.now() - started,
|
|
239
244
|
finishReason,
|
|
@@ -271,7 +276,7 @@ function unclosedThinking(text) {
|
|
|
271
276
|
* payload, so they are distinguished: the budget ran out, the budget went on thinking, or the server
|
|
272
277
|
* genuinely said nothing.
|
|
273
278
|
*/
|
|
274
|
-
function emptyCompletion({ label, finishReason, reasoning, maxTokens, usage: raw }) {
|
|
279
|
+
function emptyCompletion({ label, finishReason, reasoning, maxTokens, usage: raw, model }) {
|
|
275
280
|
const spent = (raw && raw.completion_tokens) || 0;
|
|
276
281
|
const usage = normaliseUsage(raw);
|
|
277
282
|
if (finishReason === 'length' || (reasoning && !spentLeftRoom(spent, maxTokens))) {
|
|
@@ -280,16 +285,16 @@ function emptyCompletion({ label, finishReason, reasoning, maxTokens, usage: raw
|
|
|
280
285
|
(reasoning ? ' on internal reasoning' : '') +
|
|
281
286
|
'. This model thinks before it replies, so raise the reply limit (1024 is a sensible floor; ' +
|
|
282
287
|
"4096 or more with reasoning: 'full').",
|
|
283
|
-
{ usage });
|
|
288
|
+
{ usage, model });
|
|
284
289
|
}
|
|
285
290
|
if (reasoning) {
|
|
286
291
|
return new AiError('bad_response',
|
|
287
292
|
`${label} returned only its internal reasoning and no answer. Raise the reply limit and try again.`,
|
|
288
|
-
{ usage });
|
|
293
|
+
{ usage, model });
|
|
289
294
|
}
|
|
290
295
|
return new AiError('bad_response',
|
|
291
296
|
`${label} returned an empty completion (finish reason: ${finishReason || 'none given'}). ` +
|
|
292
|
-
'Check the model is fully loaded on the server.', { usage });
|
|
297
|
+
'Check the model is fully loaded on the server.', { usage, model });
|
|
293
298
|
}
|
|
294
299
|
|
|
295
300
|
/** Did the completion stop well short of the ceiling? Then the ceiling was not the problem. */
|
package/reasoning.js
CHANGED
|
@@ -108,6 +108,9 @@ function reasoningFor(model, mode, provider) {
|
|
|
108
108
|
* effort → output_config.effort
|
|
109
109
|
* thinking → the `thinking` object
|
|
110
110
|
* budget → `{ type: 'enabled', budget_tokens }` sized from maxTokens (pre-4.6 models)
|
|
111
|
+
* maxOutput → the model's output limit; a budget row never sends max_tokens above it (0.27.1)
|
|
112
|
+
* cannotThink → the model has no extended thinking: 'full' sends nothing and is warned about,
|
|
113
|
+
* exactly as for an unknown model (0.27.1)
|
|
111
114
|
* sampling → false: the model rejects temperature, so none is sent
|
|
112
115
|
*
|
|
113
116
|
* `examples` are real model ids that must land on their own row — the test checks each one, which is
|
|
@@ -135,22 +138,61 @@ const ANTHROPIC_FAMILIES = Object.freeze([
|
|
|
135
138
|
{ family: 'claude-*-4-6', match: /(opus|sonnet)-4-6/,
|
|
136
139
|
examples: ['claude-opus-4-6', 'claude-sonnet-4-6'],
|
|
137
140
|
off: {}, full: { thinking: { type: 'adaptive' }, effort: 'high' }, sampling: true },
|
|
138
|
-
|
|
139
|
-
|
|
140
|
-
|
|
141
|
+
// PRE-4.6, ONE ROW PER OUTPUT LIMIT (0.27.1). 0.27.0 had a single row and doubled max_tokens on
|
|
142
|
+
// 'full' with no ceiling — 80000 asked of Haiku 4.5, which allows 64000, is a 400. Limits from the
|
|
143
|
+
// model tables (cached 2026-09-25). Dated ids (claude-opus-4-20250514) and the -0 aliases land here
|
|
144
|
+
// too; 0.27.0 matched neither, so 'full' on them sent no thinking and warned "unknown model".
|
|
145
|
+
{ family: 'claude-*-4-5', match: /-4-5/,
|
|
146
|
+
examples: ['claude-haiku-4-5', 'claude-sonnet-4-5', 'claude-opus-4-5', 'claude-sonnet-4-5-20250929'],
|
|
147
|
+
off: {}, full: { budget: true }, maxOutput: 64000, sampling: true },
|
|
148
|
+
{ family: 'claude-opus-4-1 / claude-opus-4', match: /opus-4-1|opus-4(?:-0|-\d{8})?(?![-.]?\d)/,
|
|
149
|
+
examples: ['claude-opus-4-1', 'claude-opus-4-1-20250805', 'claude-opus-4-0', 'claude-opus-4-20250514'],
|
|
150
|
+
off: {}, full: { budget: true }, maxOutput: 32000, sampling: true },
|
|
151
|
+
{ family: 'claude-sonnet-4', match: /sonnet-4(?:-0|-\d{8})?(?![-.]?\d)/,
|
|
152
|
+
examples: ['claude-sonnet-4-0', 'claude-sonnet-4-20250514'],
|
|
153
|
+
off: {}, full: { budget: true }, maxOutput: 64000, sampling: true },
|
|
154
|
+
{ family: 'claude-3-7-sonnet', match: /claude-3-7/,
|
|
155
|
+
examples: ['claude-3-7-sonnet-latest', 'claude-3-7-sonnet-20250219'],
|
|
156
|
+
off: {}, full: { budget: true }, maxOutput: 64000, sampling: true },
|
|
157
|
+
// EVERY OTHER CLAUDE 3 HAS NO EXTENDED THINKING (0.27.1). 0.27.0's /claude-3/ sent them
|
|
158
|
+
// `thinking: enabled` on 'full', which the API refuses. Now they keep the pre-0.27 request and
|
|
159
|
+
// 'full' is a warning — the same as a model this table does not know — never an error: the caller
|
|
160
|
+
// asked for more thought, not for the call to fail.
|
|
161
|
+
{ family: 'claude-3 without extended thinking', match: /claude-3/,
|
|
162
|
+
examples: ['claude-3-5-haiku-20241022', 'claude-3-5-haiku-latest', 'claude-3-haiku-20240307',
|
|
163
|
+
'claude-3-opus-20240229', 'claude-3-5-sonnet-20241022'],
|
|
164
|
+
off: {}, full: {}, cannotThink: true, sampling: true }
|
|
141
165
|
]);
|
|
142
166
|
|
|
143
167
|
/** Thinking room added on top of the answer's own maxTokens when a pre-4.6 model is asked to think. */
|
|
144
168
|
const MIN_THINKING_BUDGET = 1024;
|
|
145
169
|
|
|
170
|
+
/**
|
|
171
|
+
* maxTokens as a positive integer, or refused (0.27.1). A fraction is floored; anything that is not
|
|
172
|
+
* a finite number of at least 1 is the caller's bug. A string "2048" used to turn the budget sum into
|
|
173
|
+
* a string concatenation that Math.min read as the model's whole limit, and a negative value left
|
|
174
|
+
* budget_tokens above max_tokens — both sent without complaint.
|
|
175
|
+
*/
|
|
176
|
+
function positiveTokens(value) {
|
|
177
|
+
const n = typeof value === 'number' && Number.isFinite(value) ? Math.floor(value) : NaN;
|
|
178
|
+
if (!(n >= 1)) {
|
|
179
|
+
const shown = typeof value === 'string' ? JSON.stringify(value) : String(value);
|
|
180
|
+
throw new AiError('refused', `maxTokens must be a positive number of tokens; got ${shown}.`);
|
|
181
|
+
}
|
|
182
|
+
return n;
|
|
183
|
+
}
|
|
184
|
+
|
|
146
185
|
/**
|
|
147
186
|
* What to change on an Anthropic request body for this model and mode, or null for an unknown model
|
|
148
|
-
* (which keeps the pre-0.27 request: no thinking field, temperature as given).
|
|
187
|
+
* (which keeps the pre-0.27 request: no thinking field, temperature as given). `cannotThink` is set
|
|
188
|
+
* when 'full' was asked of a model with no extended thinking, so the adapter can say so.
|
|
149
189
|
*
|
|
150
|
-
* @returns {{family:string, thinking?:object, effort?:string, sendTemperature:boolean, maxTokens:number
|
|
190
|
+
* @returns {{family:string, thinking?:object, effort?:string, sendTemperature:boolean, maxTokens:number,
|
|
191
|
+
* cannotThink?:boolean}|null}
|
|
151
192
|
*/
|
|
152
|
-
function anthropicReasoning(model, mode,
|
|
193
|
+
function anthropicReasoning(model, mode, maxTokensIn) {
|
|
153
194
|
assertReasoning(mode);
|
|
195
|
+
const maxTokens = positiveTokens(maxTokensIn);
|
|
154
196
|
const m = String(model || '').toLowerCase();
|
|
155
197
|
const f = ANTHROPIC_FAMILIES.find((x) => x.match.test(m));
|
|
156
198
|
if (!f) return null;
|
|
@@ -158,12 +200,17 @@ function anthropicReasoning(model, mode, maxTokens) {
|
|
|
158
200
|
const out = { family: f.family, sendTemperature: f.sampling, maxTokens };
|
|
159
201
|
if (want.thinking) out.thinking = want.thinking;
|
|
160
202
|
if (want.effort) out.effort = want.effort;
|
|
203
|
+
if (isFull(mode) && f.cannotThink) out.cannotThink = true;
|
|
161
204
|
if (want.budget) {
|
|
162
205
|
// The caller's maxTokens stays the ANSWER's room; the thinking budget is added on top, because
|
|
163
206
|
// budget_tokens must be below max_tokens and a budget carved out of the answer would starve it.
|
|
164
|
-
|
|
207
|
+
// CAPPED AT THE MODEL'S OUTPUT LIMIT (0.27.1): max_tokens counts the thinking too, and a request
|
|
208
|
+
// past the limit is refused. Near the limit the thinking room shrinks first, down to the API's
|
|
209
|
+
// minimum of 1024; only then does the answer's room give way.
|
|
210
|
+
const total = Math.min(maxTokens + Math.max(MIN_THINKING_BUDGET, maxTokens), f.maxOutput);
|
|
211
|
+
const budget = Math.max(MIN_THINKING_BUDGET, total - maxTokens);
|
|
165
212
|
out.thinking = { type: 'enabled', budget_tokens: budget };
|
|
166
|
-
out.maxTokens =
|
|
213
|
+
out.maxTokens = total;
|
|
167
214
|
// Extended thinking requires the default temperature on these models.
|
|
168
215
|
out.sendTemperature = false;
|
|
169
216
|
}
|