@owlmeans/llm 0.1.18-rc.12 → 0.1.18-rc.13

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (46) hide show
  1. package/README.md +2 -2
  2. package/agent-meta/manifest.json +2 -2
  3. package/agent-meta/skills/llm/SKILL.md +189 -232
  4. package/agent-meta/skills/llm-prompt-caching/SKILL.md +48 -4
  5. package/build/errors.d.ts.map +1 -1
  6. package/build/errors.js.map +1 -1
  7. package/build/execution/service.d.ts.map +1 -1
  8. package/build/execution/service.js.map +1 -1
  9. package/build/execution/utils.d.ts.map +1 -1
  10. package/build/helpers/cache.d.ts.map +1 -1
  11. package/build/helpers/json.d.ts.map +1 -1
  12. package/build/helpers/messages.d.ts.map +1 -1
  13. package/build/helpers/retry.d.ts +14 -0
  14. package/build/helpers/retry.d.ts.map +1 -1
  15. package/build/helpers/retry.js +14 -0
  16. package/build/helpers/retry.js.map +1 -1
  17. package/build/helpers/spectate.d.ts.map +1 -1
  18. package/build/model.d.ts.map +1 -1
  19. package/build/model.js +3 -1
  20. package/build/model.js.map +1 -1
  21. package/build/plugins/anthropic.d.ts +15 -3
  22. package/build/plugins/anthropic.d.ts.map +1 -1
  23. package/build/plugins/anthropic.js +18 -3
  24. package/build/plugins/anthropic.js.map +1 -1
  25. package/build/plugins/index.d.ts.map +1 -1
  26. package/build/plugins/openai.d.ts.map +1 -1
  27. package/build/plugins/types.d.ts +7 -0
  28. package/build/plugins/types.d.ts.map +1 -1
  29. package/build/plugins/utils.d.ts.map +1 -1
  30. package/build/prompt/render.d.ts.map +1 -1
  31. package/build/prompt/service.d.ts.map +1 -1
  32. package/build/prompt/service.js.map +1 -1
  33. package/build/service.d.ts.map +1 -1
  34. package/build/service.js.map +1 -1
  35. package/build/utils/config.d.ts.map +1 -1
  36. package/build/utils/null-report.d.ts.map +1 -1
  37. package/build/utils/prompt.d.ts.map +1 -1
  38. package/build/utils/schema.d.ts.map +1 -1
  39. package/build/utils/stream.d.ts.map +1 -1
  40. package/build/utils/stream.js.map +1 -1
  41. package/package.json +6 -6
  42. package/src/helpers/retry.ts +16 -0
  43. package/src/model.ts +3 -1
  44. package/src/plugins/anthropic.ts +22 -3
  45. package/src/plugins/types.ts +8 -0
  46. package/tests/plugins.spec.ts +30 -0
package/README.md CHANGED
@@ -20,7 +20,7 @@ abstraction that resolves models from an inheritable policy.
20
20
  ## Installation
21
21
 
22
22
  ```bash
23
- bun add @owlmeans/llm @owlmeans/llm-common
23
+ bun add @owlmeans/llm@^0.1.18-rc.12 @owlmeans/llm-common@^0.1.18-rc.11
24
24
  bun add @langchain/core @langchain/openai @langchain/anthropic # peer dependencies
25
25
  ```
26
26
 
@@ -184,7 +184,7 @@ This package ships embedded agent skills under `agent-meta/`. After installing y
184
184
  your project's skill store (`.agents/skills/`):
185
185
 
186
186
  ```sh
187
- npx @owlmeans/agent-skills
187
+ npx @owlmeans/agent-skills@^0.1.18-rc.12
188
188
  ```
189
189
 
190
190
  The embedded files are version-matched to this package release. Do not edit them
@@ -1,8 +1,8 @@
1
1
  {
2
2
  "schemaVersion": 2,
3
3
  "package": "@owlmeans/llm",
4
- "version": "0.1.18-rc.12",
5
- "generatedAt": "2026-09-01T22:50:56.459Z",
4
+ "version": "0.1.18-rc.13",
5
+ "generatedAt": "2026-09-04T22:43:25.441Z",
6
6
  "canonicalRepo": "https://github.com/owlmeans/common",
7
7
  "entries": [
8
8
  {
@@ -8,48 +8,37 @@ user-invocable: false
8
8
  # @owlmeans/llm
9
9
 
10
10
  **Layer:** Core
11
- **Install:** `"@owlmeans/llm": "^0.1.18-rc.12"` in `dependencies` (plus the `@langchain/*` peers)
11
+ **Install:** `"@owlmeans/llm": "^0.1.18-rc.13"` in `dependencies` (plus the `@langchain/*` peers)
12
12
 
13
- The inference runtime. Everything provider-specific is a **plugin**; the model itself only
14
- owns the provider-independent parts (streaming discipline, retries, validation,
15
- observability). Serializable contracts live in `@owlmeans/llm-common`.
13
+ The inference runtime. Everything provider-specific is a **plugin**; the model itself only owns the
14
+ provider-independent parts (streaming discipline, retries, validation, observability). Serializable
15
+ contracts live in `@owlmeans/llm-common`. `src/helpers/` is exported — functions a **consumer** may
16
+ use alongside a model; `src/utils/` is internal and **never exported** (a spec that needs one imports
17
+ it from `../src/utils/…`). Decide the side before placing a function, and never export a `utils/`
18
+ symbol "because a test needs it".
16
19
 
17
20
  ## Key Exports
18
21
 
19
22
  | Export | Description |
20
23
  |--------|-------------|
21
- | `makeLlmModel(options, spectator)` | The four-method model: `ask` / `talk` / `invoke` / `request`. |
22
- | `makeLlmService(options, alias?)` · `appendLlmService(ctx, options, alias?)` | Model factory/registry — resolves a `ModelConfig` by alias, memoized per alias+override. |
23
- | `llmServiceApi(options, self)` | The factory half WITHOUT `createService`, to compose into your own service (role accessors, domain helpers). |
24
- | `makeExecutionService(alias?, options?)` · `appendExecutionService(ctx, alias?, options?)` | Frozen 3-level executions + policy resolution + snapshot/restore/checkpoint + advice. |
24
+ | `makeLlmModel({ model, purpose, prompt?, prompts?, files?, utility?, retries?, … }, spectator)` | The four-method model: `ask` / `talk` / `invoke(input, schema, opts)` / `request`. `model` is an already-resolved `BaseChatModel`. |
25
+ | `makeLlmService(options, alias?)` · `appendLlmService(ctx, options, alias?)` · `llmServiceApi(options, self)` | Model factory/registry — `makeLlmService({ models: () => configs }).getModel(alias, override?)` resolves a `ModelConfig` by alias, memoized per alias+override. The `…Api` half omits `createService`, to compose into your own service (role accessors, domain helpers). |
26
+ | `makeExecutionService(alias?, options?)` · `appendExecutionService(ctx, alias?, options?)` · `executionServiceApi(options, self)` | Frozen 3-level executions + policy resolution + snapshot/restore/checkpoint + advice, and the same composable half. |
25
27
  | `makePromptService(options?, alias?)` · `appendPromptService(ctx, options?, alias?)` · `promptServiceApi(options, self)` | Skill registry + the composition plugin chain. Also at `@owlmeans/llm/prompt`. |
26
28
  | `rolePlugin`, `skillsPlugin`, `contextPlugin`, `BUILT_IN_PROMPT_PLUGINS` | The built-in composition plugins. |
27
29
  | `PromptContext.claim(key)` · `PromptComposeParams.utility` | Per-composition ownership of a key; a cheap model for one plugin-side call. |
28
- | `renderSkill`, `sortSkills`, `joinChunks`, `compareAlias`, `prefixHash` | Deterministic rendering primitives — reuse them, never re-implement. |
29
- | `readCacheUsage`, `hasCacheActivity` | Normalized prompt-cache accounting from a completion. |
30
- | `executionServiceApi(options, self)` | The execution half without `createService`, for the same composition pattern. |
30
+ | `renderSkill`, `sortSkills`, `joinChunks`, `compareAlias`, `prefixHash` · `readCacheUsage`, `hasCacheActivity` | Deterministic rendering primitives — reuse them, never re-implement — and normalized prompt-cache accounting from a completion. |
31
31
  | `plugins`, `registerLlmPlugin`, `resolvePlugin`, `pluginOf`, `pluginFor` | The provider-plugin registry. Also at `@owlmeans/llm/plugins`. |
32
32
  | `anthropicPlugin`, `openAiPlugin`, `compatiblePlugin`, `openAiFamily` | Built-in providers; `openAiFamily` is the shared OpenAI-client behaviour to spread into a new plugin. |
33
- | `NO_SAMPLING_PREFIXES`, `rejectsSampling(model)` | Which Claude families reject `temperature`/`top_p`/`top_k` (4.7+ and the 5 family). Consumers pin presets against it. |
34
- | `RESPONSES_API_PREFIXES`, `usesResponsesApi(model)` | Which OpenAI families go through the Responses API and therefore reject `temperature`/`top_p` (`gpt-5*`, `codex-*`). The OpenAI counterpart of the pair above; consumers pin presets against it. |
35
- | `withRetry`, `registerFatalError`, `spectate`, `normalizeInput`, `parseJsonContent`, `coerceToSchema` | Helpers usable alongside a model. Also at `@owlmeans/llm/helpers`. |
33
+ | `NO_SAMPLING_PREFIXES` / `rejectsSampling(model)` · `RESPONSES_API_PREFIXES` / `usesResponsesApi(model)` | Which families reject which sampling parameters — see the table below. Consumers pin presets against them. |
34
+ | `withRetry`, `registerFatalError`, `isFatalError`, `spectate`, `normalizeInput`, `parseJsonContent`, `coerceToSchema` | Helpers usable alongside a model. Also at `@owlmeans/llm/helpers`. |
36
35
  | `LlmError`, `LlmModelError`, `LlmMissconfiguredError`, `LlmPluginError`, `LlmRetryExceededError` | `ResilientError` family. `LlmModelError` is the RETRYABLE one. |
37
36
  | `mergePrompt`, `mergePolicy`, `resolveRole`, `effortPatch` | Execution merge helpers; `mergePrompt` unions skills and takes the deepest role. |
38
37
  | `DEFAULT_MODEL_RETRIES`, `MODEL_STREAM_TIMEOUT_MS` (3 min idle), `FALLBACK_AFTER_ATTEMPTS`, `DEFAULT_EFFORT`, `EFFORT_TABLE`, `MAX_CACHE_BREAKPOINTS`, `MAX_SYSTEM_BREAKPOINTS`, `MIN_CACHEABLE_TOKENS`, `LLM_SERVICE`, `EXECUTION_SERVICE`, `PROMPT_SERVICE` | Tuning + aliases. |
39
38
 
40
- ## helpers/ vs utils/ — the rule this package follows
41
-
42
- - `src/helpers/` — functions a **consumer** may use alongside a model. Exported.
43
- - `src/utils/` — used only inside the library. **Never exported**; a spec that needs one
44
- imports it from `../src/utils/…`.
45
-
46
- Adding a function? Decide which side it belongs to first, then place it. Do not export a
47
- `utils/` symbol "because a test needs it".
48
-
49
39
  ## Provider differences are plugins, never `if`s
50
40
 
51
- `LlmPlugin` is the single seam. If you find yourself writing `instanceof ChatAnthropic` or
52
- `config.provider === …` in `model.ts` or `service.ts`, it belongs on the plugin instead:
41
+ `LlmPlugin` is the single seam: an `instanceof ChatAnthropic` or `config.provider === …` in `model.ts` or `service.ts` belongs on the plugin.
53
42
 
54
43
  | Plugin member | Replaces |
55
44
  |---|---|
@@ -62,198 +51,155 @@ Adding a function? Decide which side it belongs to first, then place it. Do not
62
51
  | `patchCache` | the message-prefix cache marker |
63
52
  | `cacheKey` (via `ModelConfig.cacheKey`) | provider cache-routing hints such as OpenAI's `prompt_cache_key` |
64
53
  | `isFatal` | "this error can never be retried" |
54
+ | `suppressesThinking` | whether the plugin turns reasoning off on the wire FOR THIS CONFIG, which is what drops the `/no_think` prompt directive from the prepared messages |
65
55
 
66
56
  ### A provider that removes a parameter is a plugin concern too
67
57
 
68
- `build` is not the only place a parameter reaches the wire, and `refine` is not only a retry
69
- hook — **every** call is made on the instance `refine` returns, attempt 0 included. A knob
70
- suppressed in `build` and re-applied in `refine` therefore ships on every single request, not
71
- on a retry: an unsupported `temperature` restored there 400s the model's whole family on the
72
- first call, and `isFatal` correctly refuses to retry it, so the failure is immediate and total.
73
-
74
- Both built-in plugins gate their rejected parameters in **both** hooks, through a predicate the
75
- package root exports:
58
+ `build` is not the only place a parameter reaches the wire, and `refine` is not only a retry hook —
59
+ **every** call is made on the instance `refine` returns, attempt 0 included. A knob suppressed in
60
+ `build` and re-applied in `refine` therefore ships on every single request: an unsupported
61
+ `temperature` restored there 400s the model's whole family on the first call, and `isFatal`
62
+ correctly refuses to retry it. Both built-in plugins gate their rejected parameters in **both**
63
+ hooks, through a predicate the package root exports:
76
64
 
77
65
  | Family | Rejects | Predicate |
78
66
  |---|---|---|
79
67
  | Claude 4.7+ and the 5 family | `temperature`, `top_p`, `top_k` | `NO_SAMPLING_PREFIXES` / `rejectsSampling(model)` |
80
68
  | OpenAI Responses API (`gpt-5*`, `codex-*`) | `temperature`, `top_p` | `RESPONSES_API_PREFIXES` / `usesResponsesApi(model)` |
81
69
 
82
- Models below those lines keep the deterministic `temperature: 0` default. `refine` reads the
83
- id off the ACTIVE base (`model.model ?? lc_kwargs.model`), so a same-family `fallback` rung is
84
- judged on its own id rather than the primary's, and the OpenAI check also accepts a
85
- `useResponsesApi` already on `lc_kwargs`.
86
-
87
- Export the predicate rather than keeping it private: consumers pin presets against it
88
- (`viable-agent`'s `tests/presets.test.ts` asserts no preset entry declares a parameter its
89
- model rejects), and a second hand-written copy of the list drifts exactly when a new family is
90
- added — the moment it has to be right.
70
+ Models below those lines keep the deterministic `temperature: 0` default. Each `refine` re-derives
71
+ the family from the ACTIVE base instance it is handed — Anthropic from `modelName ?? model`, OpenAI
72
+ from `model ?? lc_kwargs.model` plus a `useResponsesApi` already on `lc_kwargs` — so a same-family
73
+ `fallback` rung is judged on its own id, not the primary's. Keep the predicate exported: consumers
74
+ pin presets against it (viable-agent's `tests/presets.test.ts` asserts no preset entry declares a
75
+ parameter its model rejects), and a second hand-written copy drifts when a family is added.
91
76
 
92
- **Registration order is load-bearing.** Instance-based lookup (`pluginFor`) returns the
93
- FIRST plugin whose `owns` matches. `compatible` is registered before `openai` because both
94
- build a `ChatOpenAI`, and assuming the tool-calling hack for an unlabelled model is safe
95
- everywhere while assuming native JSON-schema support is not.
77
+ **Registration order is load-bearing.** Instance lookup (`pluginFor`) returns the FIRST plugin whose
78
+ `owns` matches, and `compatible` is registered before `openai` because both build a `ChatOpenAI`:
79
+ assuming the tool-calling hack for an unlabelled model is safe everywhere, assuming native
80
+ JSON-schema support is not.
96
81
 
97
82
  ## System prompts: a role and skills, never a hand-built message
98
83
 
99
84
  `makeLlmModel` takes `prompt` (a `PromptInput`) and a `prompts` resolver. The prompt service
100
- composes them into an ordered, cacheable system message; a caller's own leading
101
- `SystemMessage` is folded into the volatile `Context` block, so an unmigrated call site
102
- still works. **Do not build a persona as a `SystemMessage` in a helper** — declare it as
103
- `PromptPolicy.role` plus registered skills, or the knowledge duplicates and the cache
104
- prefix stops being stable.
105
-
106
- Skills accumulate down the execution chain (project → task → helper) and the deepest
107
- declared `role` wins — see `mergePrompt`. Full rules, breakpoint budget and the provider
108
- facts behind them: [[llm-prompt-caching]].
109
-
110
- Two seams exist for plugins that would otherwise collide or overspend:
111
- `PromptContext.claim(key)` grants a key to the first plugin that asks within one
112
- composition, so a static catalogue and a detector never emit the same skill twice; and
113
- `PromptComposeParams.utility` offers a cheap model for ONE side call. Both are threaded
114
- from where a prompt is composed — `makeLlmModel`'s `utility` option (beside `files`) and
115
- `AgentOptions.utility`. Neither may change a cached block: see [[llm-prompt-caching]].
85
+ composes them into an ordered, cacheable system message; a caller's own leading `SystemMessage` is
86
+ folded into the volatile `Context` block, so an unmigrated call site still works. **Do not build a
87
+ persona as a `SystemMessage` in a helper** — declare it as `PromptPolicy.role` plus registered
88
+ skills, or the knowledge duplicates and the cache prefix stops being stable. Skills accumulate down
89
+ the execution chain (project → task → helper) and the deepest declared `role` wins (`mergePrompt`).
90
+ Block order, the breakpoint budget, the provider facts behind them, and the plugin seams
91
+ `claim(key)` / `utility`: [[llm-prompt-caching]].
116
92
 
117
93
  ## Execution: policy in, model out
118
94
 
119
- ```
120
- ProjectExecution ← root: policy + purpose + models resolver
121
- └─ TaskExecution ← + resumable state (phase/cursor/completed/data)
122
- └─ HelperExecution ← + a RESOLVED model + temperatureFactory, bound to a role
123
- ```
95
+ `ProjectExecution` (policy + purpose + models resolver) → `TaskExecution` (+ resumable state:
96
+ phase/cursor/completed/data) → `HelperExecution` (+ a RESOLVED model + `temperatureFactory`, a role).
124
97
 
125
- `prompt` (role + skills) travels on `ExecutionState`, so it survives snapshot/restore;
126
- `prompts` and `files` are collaborators and never enter a snapshot.
98
+ `prompt` (role + skills) travels on `ExecutionState`, so it survives snapshot/restore; `prompts` and
99
+ `files` are collaborators and never enter a snapshot. Every method returns a NEW `Object.freeze`d
100
+ object. Resolution precedence in `model(exec, role, override)`: **roleOverride → modelOverride →
101
+ effort tier → `LlmService.getModel`**; `escalate(exec, { effort })` raises the tier once and cascades
102
+ to everything derived from it. `snapshot` excludes `state` itself — without that, every
103
+ `derive`/`escalate`/`withPurpose` on a task would nest another copy of the previous state.
127
104
 
128
- Every method returns a NEW `Object.freeze`d object. Resolution precedence in
129
- `model(exec, role, override)`: **roleOverride → modelOverride → effort tier →
130
- `LlmService.getModel`**. `escalate(exec, { effort })` raises the tier once and cascades to
131
- everything derived from it.
105
+ Extending it for a domain: declare your own `Execution`/input types, list collaborator fields in
106
+ `ExecutionServiceOptions.collaboratorKeys` (kept out of snapshots), instantiate the service generic
107
+ with your own `ExecutionShape`, and **never narrow an inherited method signature** (contravariance).
132
108
 
133
- Extending it for a domain: declare your own `Execution`/input types, list your collaborator
134
- fields in `ExecutionServiceOptions.collaboratorKeys` so they stay out of snapshots, and
135
- instantiate the service generic with your own `ExecutionShape` — **do not narrow the
136
- inherited method signatures**, which would be a contravariance error.
137
-
138
- `snapshot` excludes `state` itself; without that, every `derive`/`escalate`/`withPurpose`
139
- on a task would nest another copy of the previous state.
140
-
141
- `forHelper` also accepts `output` — an initial `maxTokens` for a helper whose one answer is
142
- genuinely large (a whole source file rather than a decision). It is a model selector, not a
143
- field the helper carries: it becomes a call override, is clamped to `maxOutput`, doubles from
144
- there under retry, and survives `temperatureFactory`.
109
+ `forHelper` also accepts `output` — an initial `maxTokens` for a helper whose one answer is genuinely
110
+ large (a whole source file, not a decision). It is a model selector, not a field the helper carries:
111
+ it becomes a call override, clamped to `maxOutput`, doubling under retry, surviving `temperatureFactory`.
145
112
 
146
113
  ### The cheap tier: `utility(exec, override?)`
147
114
 
148
- Work that is not the work — a relevance pick, a classification, a one-line judgement a
149
- plugin needs before the real call can be shaped — runs on `ExecutionService.utility`, never
150
- on the helper's own model. It resolves `policy.utilityRole ?? UTILITY_ROLE`
151
- (`@owlmeans/llm-common`, value `'utility'`) at `ExecutionEffort.Economy`, through the SAME
152
- ladder as `model()`: `roleOverrides` remap it and `modelOverrides` pin it exactly as for any
153
- other role. The economy floor is local — the execution it was asked on keeps its own tier —
154
- and `utilityRole` travels on `ModelPolicy`, so it survives `forTask`/`escalate` and a
155
- snapshot/restore round trip.
156
-
157
- A deployment that wants the tier must register a config under that alias; an unregistered
158
- alias is a misconfiguration like any other, which is why a plugin asking for one is handed a
159
- resolver that may yield `undefined` rather than the model itself.
115
+ Work that is not the work — a relevance pick, a classification, a one-line judgement a plugin needs
116
+ before the real call can be shaped — runs on `ExecutionService.utility`, never on the helper's own
117
+ model. It resolves `policy.utilityRole ?? UTILITY_ROLE` (`@owlmeans/llm-common`, value `'utility'`)
118
+ at `ExecutionEffort.Economy`, through the SAME ladder as `model()`: `roleOverrides` remap it and
119
+ `modelOverrides` pin it as for any other role. The economy floor is local — the execution it was
120
+ asked on keeps its own tier — and `utilityRole` travels on `ModelPolicy`, so it survives
121
+ `forTask`/`escalate` and a snapshot/restore round trip. `utility` returns a `BaseChatModel`, never
122
+ `undefined`: an alias with no registered config reaches `createModel` through `model()` and throws
123
+ `LlmMissconfiguredError`, like any other unregistered role. The `undefined` a prompt plugin has to
124
+ handle comes from the other end — `PromptComposeParams.utility` (and `AgentOptions.utility`) is an
125
+ OPTIONAL resolver, unset wherever no cheap tier is wired, so a plugin that cannot get one degrades
126
+ rather than fails.
160
127
 
161
128
  ### The plugin seam has two hooks, dispatched independently
162
129
 
163
- `ExecutionPlugin` carries `onCheckpoint`/`onRestore` (persist and resume an
164
- `ExecutionState`) and `advise` (answer a performer's question about the project it is
165
- working in — `ExecutionService.advise(exec, request)`, first usable answer wins, a throwing
166
- plugin is skipped, `null` when nobody answers). Advice is advisory by contract: a caller
167
- appends whatever comes back and proceeds unchanged on `null`.
168
-
169
- `checkpoint` dispatches on plugins that declare **`onCheckpoint`**, not on the plugin count —
170
- otherwise registering an advise-only plugin would silently start composing snapshots nobody
171
- consumes.
130
+ `ExecutionPlugin` carries `onCheckpoint`/`onRestore` (persist and resume an `ExecutionState`) and
131
+ `advise` (answer a performer's question about the project it works in —
132
+ `ExecutionService.advise(exec, request)`, first usable answer wins, a throwing plugin is skipped,
133
+ `null` when nobody answers).
134
+ Advice is advisory by contract: a caller appends whatever comes back and proceeds unchanged on
135
+ `null`. `checkpoint` dispatches on plugins declaring **`onCheckpoint`**, not on the plugin count, or
136
+ an advise-only plugin would silently start composing snapshots nobody consumes.
172
137
 
173
138
  ## Classify a provider error by walking `cause`, never by `instanceof` or a surface read
174
139
 
175
- Two independent layers hide the wire shape of a provider failure, and each one alone is
176
- enough to make a fatal error look retryable — which costs all eight attempts with the real
177
- message buried under the repeats.
140
+ Two independent layers hide the wire shape of a provider failure, and each alone makes a fatal error
141
+ look retryable — costing the whole retry budget with the real message buried under the repeats.
178
142
 
179
- 1. **Nested SDK copies.** `@langchain/anthropic` and `@langchain/openai` bundle their OWN
180
- copies of the provider SDKs, so an error they throw is an instance of a DIFFERENT class
181
- than the one this package imports — `e instanceof BadRequestError` silently returns
182
- `false`.
143
+ 1. **Nested SDK copies.** `@langchain/anthropic` and `@langchain/openai` bundle their OWN copies of
144
+ the provider SDKs, so `e instanceof BadRequestError` compares against a DIFFERENT class than the
145
+ one this package imports and silently returns `false`.
183
146
  2. **Langchain's own error wrappers.** A failure is re-wrapped in a typed langchain error
184
- (`ContextOverflowError` for an input past the context window, and its siblings) that
185
- carries the original **only under `cause`** and has no `status` of its own — so
186
- `e.status === 400` misses it too.
147
+ (`ContextOverflowError` for an input past the context window, and its siblings) carrying the
148
+ original **only under `cause`**, no `status` of its own — so `e.status === 400` misses it.
187
149
 
188
- Use `isBadRequest` from `plugins/utils.ts`: it walks the `cause` chain looking for
189
- `status === 400`, bounded in depth so a self-referential chain terminates. Both built-in
190
- `isFatal` implementations go through it.
191
-
192
- A context overflow is the case that makes this urgent rather than merely untidy: `refine`
193
- escalates the **output** budget on each retry, so an over-limit **input** can never improve —
194
- every attempt re-sends the identical oversized request. Consumers hold their locks for the
195
- whole loop, so a single unfixable call becomes minutes of thrash on the caller's side.
196
-
197
- The same trap applies to any cross-copy `instanceof`; it is the runtime face of the
198
- peer-dependency identity rule below.
150
+ Use `isBadRequest` from `plugins/utils.ts`: it walks the `cause` chain for `status === 400`, bounded
151
+ in depth so a self-referential chain terminates. Both built-in `isFatal` implementations go through
152
+ it. A context overflow makes this urgent: `refine` escalates the **output** budget on each retry, so
153
+ an over-limit **input** can never improve — every attempt re-sends the identical oversized request,
154
+ and consumers hold their locks for the whole loop. The same trap applies to any cross-copy
155
+ `instanceof`; it is the runtime face of the peer-dependency rule below.
199
156
 
200
157
  ## Peer-dependency rule (langchain identity)
201
158
 
202
- `@langchain/core`, `@langchain/openai` and `@langchain/anthropic` are **peer** dependencies:
203
- model instances cross the package boundary, and two installed copies of a class with
204
- protected members are nominally distinct types. In a linked-workspace checkout the consumer
205
- must pin them to a single copy — see the `bun-linked-workspaces` skill.
206
-
207
- ## Usage
208
-
209
- ```typescript
210
- import { makeLlmModel, makeLlmService } from '@owlmeans/llm'
211
- import { ModelProvider } from '@owlmeans/llm-common'
212
-
213
- const llm = makeLlmService({ models: () => configs })
214
- const model = makeLlmModel(
215
- { model: llm.getModel('analyst'), purpose: { type: 'analysis' } }, spectator
216
- )
217
- const spec = await model.invoke('Describe the app', SpecSchema, { action: 'spec' })
218
- ```
159
+ `@langchain/core`, `@langchain/openai` and `@langchain/anthropic` are **peer** dependencies: model
160
+ instances cross the package boundary, and two installed copies of a class with protected members are
161
+ nominally distinct types. A consumer must end up with exactly ONE copy of each — declare them at one
162
+ range and check no nested `node_modules` holds a second, or every model instance crossing a boundary
163
+ becomes a foreign type and `instanceof` starts lying.
219
164
 
220
165
  ## Hangs are bounded by an IDLE deadline, not a total one
221
166
 
222
- A stalled provider is aborted after `MODEL_STREAM_TIMEOUT_MS` (3 min) of SILENCE and
223
- surfaces as a retryable `LlmModelError`, so the escalator moves on. The timer re-arms on
224
- every token, so a long-but-productive generation is never cut off — which is why the value
225
- can be low. Set it for a whole deployment with `LlmServiceOptions.streamTimeout` where the
226
- application composes its context; a preset that names its own `ModelConfig.streamTimeout`
227
- keeps it.
228
-
229
- Note what this does NOT bound: a call that keeps streaming forever, and retries. A fatal
230
- error misclassified as retryable multiplies its own latency by `DEFAULT_MODEL_RETRIES` —
231
- see the `instanceof` trap above.
167
+ A stalled provider is aborted after `MODEL_STREAM_TIMEOUT_MS` (3 min) of SILENCE and surfaces as a
168
+ retryable `LlmModelError`, so the escalator moves on. The timer re-arms on every token, so a
169
+ long-but-productive generation is never cut off — which is why the value can be low. Set it for a
170
+ deployment with `LlmServiceOptions.streamTimeout` where the application composes its context; a
171
+ preset naming its own `ModelConfig.streamTimeout` keeps it. It does NOT bound a call that keeps
172
+ streaming forever, nor retries — a fatal error misclassified as retryable multiplies its own latency
173
+ by `DEFAULT_MODEL_RETRIES`.
232
174
 
233
175
  ## An outer validation loop must pass its attempt in
234
176
 
235
- A caller that validates the OUTPUT — a diff that has to apply, a file that must not come
236
- back truncated — runs its own retry loop around whole `ask` calls. Every one of those calls
237
- starts a FRESH inner loop at attempt 0, so the escalator's two rungs never move: same model,
238
- same output budget, same deterministic answer, N times. Pass `escalation: <outer attempt>`
239
- in `LlmCallOptions` and the per-call escalator starts that far up its ladder instead —
240
- `maxTokens` doubling and the `FALLBACK_AFTER_ATTEMPTS` switch to `ModelConfig.fallback` both
241
- advance. It is clamped to `retries - 1`, moves the STARTING rung only, and never changes how
242
- many attempts the call itself makes.
243
-
244
- Two things this depends on, and both are preset data rather than code: the role must
245
- actually declare a `fallback`, and that fallback must be in the same plugin `family` (a
246
- cross-family one is skipped with a warning, because switching provider mid-call flips the
247
- structured-output shape).
248
-
249
- `LlmCallOptions.fatal` is the matching lever in the other direction — a per-call resolver
250
- consulted before the global ones and the plugin's `isFatal`, for an error the caller knows
251
- no retry can fix.
177
+ A caller that validates the OUTPUT — a diff that has to apply, a file that must not come back
178
+ truncated — runs its own retry loop around whole `ask` calls. Every one of those calls starts a FRESH
179
+ inner loop at attempt 0, so the escalator's two rungs never move: same model, same output budget,
180
+ same deterministic answer, N times. Pass `escalation: <outer attempt>` in `LlmCallOptions` and the
181
+ per-call escalator starts that far up its ladder instead — `maxTokens` doubling and the
182
+ `FALLBACK_AFTER_ATTEMPTS` switch to `ModelConfig.fallback` both advance. It is clamped to
183
+ `retries - 1`, moves the STARTING rung only, and never changes how many attempts the call makes. Two
184
+ things it depends on, both preset data rather than code: the role must declare a `fallback`, and that
185
+ fallback must be in the same plugin `family` (a cross-family one is skipped with a warning, because
186
+ switching provider mid-call flips the structured-output shape). `LlmCallOptions.fatal` is the lever
187
+ in the other direction — a per-call resolver consulted before the global ones and the plugin's
188
+ `isFatal`, for an error the caller knows no retry can fix.
189
+
190
+ ### A loop ABOVE the model asks the same question with `isFatalError`
191
+
192
+ A retry loop is not the only place that decides to carry on: a fix ladder rescues a failed repair and
193
+ climbs to a stronger model, an agent runner catches a round that threw and reports "gave up". Both
194
+ are right for a model that answered badly and wrong for a budget that ran out, and a blanket `catch`
195
+ cannot tell them apart — an exhausted balance becomes more expensive calls instead of a halt.
196
+ `isFatalError(e, fatal?)` runs the same resolvers, in the same order, that `withRetry` uses, and
197
+ returns the error to abort WITH (a resolver may unwrap a carrier and hand back the real cause) or
198
+ `null` when nothing considers it terminal. Ask it rather than re-deriving the rule.
252
199
 
253
200
  ## Output caps: what the deployment wants vs what the provider allows
254
201
 
255
- Four fields, and conflating them is what turns an escalation into a fatal 400 hours into a
256
- run:
202
+ Four fields, and conflating them turns an escalation into a fatal 400 hours into a run:
257
203
 
258
204
  | Field | Means |
259
205
  |---|---|
@@ -262,83 +208,94 @@ run:
262
208
  | `maxOutput` | what the PROVIDER accepts in one request — a fact about the model |
263
209
  | `contextWindow` | total window (input + output); informational, never sent |
264
210
 
265
- `resolveOutputCap` (`utils/config.ts`) is the one place that reconciles them: the declared
266
- cap chooses the ceiling and the capability trims it, and `DEFAULT_MAX_OUTPUT_CAP` applies
267
- only when neither is stated. `createModel` additionally clamps `maxTokens` to `maxOutput`
268
- and warns about a cap above it. For an aggregated model `maxOutput` is the limit of the
269
- `inferenceProvider` actually pinned, which is often far below what the model can do
270
- elsewhere. `combinedWindow: true` marks a model whose window is shared between input and
271
- output (MiniMax M2.x, gpt-oss) — nothing enforces it at runtime; it keeps presets honest
272
- about leaving room for the prompt.
273
-
274
- **A `fallback` that changes `model` must restate `contextWindow`/`maxOutput`** (and reset
275
- `combinedWindow`): the fallback config is `{...primaryConfig, ...fallback}`, so every field
276
- the patch does not name is inherited from a different model.
277
-
278
- ### Reasoning is billed against the same budget as the answer
279
-
280
- The Claude 5-series models — the same set as `NO_SAMPLING_PREFIXES` in the Anthropic plugin —
281
- think ADAPTIVELY whether or not the request asks them to, and by default that thinking is not
282
- displayed: it streams as thinking blocks with empty text. It is spent from `max_tokens`, the same
283
- allowance as the answer. So a budget sized for the answer alone is spent entirely on reasoning and
284
- the completion arrives well-formed, `stop_reason: "max_tokens"`, carrying no text block at all.
285
-
286
- `ADAPTIVE_MIN_MAX_TOKENS` (32k) is the floor that buys room for both. It is a FLOOR, not an
287
- override — a preset asking for more keeps it — and it is clamped through `resolveOutputCap`, so it
288
- can never exceed what the provider accepts and turn a retryable empty answer into a fatal 400.
289
- Raising or removing it re-opens the failure.
290
-
291
- Two things made this expensive to find, both now fixed and worth not undoing:
292
-
293
- - **An empty completion is a null result, not a filter rejection.** `ask` tests emptiness BEFORE
294
- the caller's filter, because every shipped filter returns null only for empty input — running
295
- one first reported a provider problem as the caller's fault and, worse, skipped `reportNull`, so
296
- the `finishReason` and `outputTokens` that name the cause were never produced.
211
+ `resolveOutputCap` (`utils/config.ts`) reconciles them: the declared cap chooses the ceiling and the
212
+ capability trims it, and `DEFAULT_MAX_OUTPUT_CAP` applies only when neither is stated. `createModel`
213
+ also clamps `maxTokens` to `maxOutput` and warns about a cap above it. For an aggregated model
214
+ `maxOutput` is the limit of the `inferenceProvider` actually pinned, often far below what the model
215
+ can do elsewhere. `combinedWindow: true` marks a model whose window is shared between input and
216
+ output (MiniMax M2.x, gpt-oss) — nothing enforces it at runtime; it keeps presets honest about
217
+ leaving room for the prompt. **A `fallback` that changes `model` must restate
218
+ `contextWindow`/`maxOutput`** (and reset `combinedWindow`): the fallback config is
219
+ `{...primaryConfig, ...fallback}`, so every field the patch does not name is inherited from a
220
+ different model.
221
+
222
+ ### Reasoning is off unless a preset asks for it — and it is billed against the same budget
223
+
224
+ The models `NO_SAMPLING_PREFIXES` names think ADAPTIVELY unless the request says otherwise: an absent
225
+ `thinking` parameter means adaptive, and `@langchain/anthropic` forwards the parameter only when the
226
+ caller sets it (its own field default of `disabled` is never sent). Left on, the reasoning costs
227
+ twice — it is spent from `max_tokens`, the same allowance as the answer (a budget sized for the
228
+ answer alone goes entirely on reasoning and the completion arrives well-formed, `stop_reason:
229
+ "max_tokens"`, with no text block), and its SUMMARISED stream arrives in bursts minutes apart, which
230
+ the idle deadline reads as a dead connection and retries from scratch
231
+ (`llm:model:stream-stalled:no token for 180000ms (idle deadline)`).
232
+
233
+ **`ModelConfig.disableThinking` is the switch, and where it lands depends on the plugin.**
234
+ `makeLlmModel` appends the literal `/no_think` to every request's prepared messages whenever the flag
235
+ is set AND the plugin's `suppressesThinking(config)` does not answer `true` — the soft switch for
236
+ models with no request-level control (Qwen3). The Anthropic plugin answers `true` only for
237
+ `rejectsSampling(model)`, and for those puts `thinking: { type: 'disabled' }` on the request in
238
+ `build`, which `refine` carries through `lc_kwargs` on every attempt. Below that line
239
+ (`claude-haiku-4-5`, `claude-sonnet-4-6`) and under any plugin declaring no hook the flag injects
240
+ prompt text instead, so set it where the wire honours it. **A preset must set the flag on every
241
+ adaptive Anthropic role**; it reaches the `fallback` rung only by inheritance from the entry
242
+ (viable-agent's `presets.test.ts` pins this). Turning reasoning ON is a per-role decision.
243
+
244
+ `ADAPTIVE_MIN_MAX_TOKENS` (32k) is the output floor the Anthropic plugin's `build` applies to every
245
+ `rejectsSampling(model)` config — `disableThinking` is not consulted, so a role with reasoning turned
246
+ off is floored just the same. It is a floor, not an override (a preset asking for more keeps it) and
247
+ it is clamped through `resolveOutputCap`, so it can never exceed what the provider accepts and turn a
248
+ retryable empty answer into a fatal 400. Raising or removing it re-opens empty completions.
249
+
250
+ Two diagnostics the above depends on:
251
+
252
+ - **An empty completion is a null result, not a filter rejection.** `ask` tests emptiness BEFORE the
253
+ caller's filter — every shipped filter returns null only for empty input, so a filter run first
254
+ blames the caller for a provider problem and skips `reportNull`, losing the `finishReason` and
255
+ `outputTokens` that name the cause.
297
256
  - **Anthropic's stop reason is not `finish_reason`.** langchain puts it in
298
- `additional_kwargs.stop_reason`; `response_metadata.finish_reason` does not exist on that
299
- provider, so every Anthropic null report printed `finishReason: undefined`. `null-report.ts`
300
- reads both, and `thinkingOnly` marks a completion that was all reasoning.
257
+ `additional_kwargs.stop_reason` and `response_metadata.finish_reason` does not exist there, so
258
+ `null-report.ts` reads both; `thinkingOnly` marks a completion that was all reasoning.
301
259
 
302
- The same class is handled for OpenAI reasoning models by shrinking the reasoning cap on retry
303
- (`plugins/openai.ts`). Escalating `maxTokens` alone does not fix it for either family: the retry
304
- draws again from an unchanged distribution.
260
+ OpenAI reasoning models get it handled by shrinking the reasoning cap on retry (`plugins/openai.ts`);
261
+ escalating `maxTokens` alone fixes neither family — the retry draws from an unchanged distribution.
305
262
 
306
263
  ## Config precedence: a preset is a base, not a final word
307
264
 
308
- `createModel` layers four sources, lowest first:
309
-
310
- presetOf(base.preset) < base < presetOf(override.preset) < override
265
+ `createModel` layers four sources, lowest first: `presetOf(base.preset)` < `base` < `presetOf(override.preset)` < `override`.
311
266
 
312
- A `preset` is a BASE that its referent refines, so it sits UNDER the config naming it.
313
- Assigning it last — as this did until the layering was fixed — meant a role declaring
314
- `preset:` silently discarded both its own fields and the caller's override, which is how
315
- effort-tier token caps and `temperatureFactory`'s temperature vanished for preset-based
316
- roles. An override naming a preset (how the execution layer delivers a `modelOverrides`
317
- string pin) outranks the alias but still yields to explicit override fields. Resolution is
318
- ONE level deep: a preset meant to carry a model must name one.
267
+ A `preset` is a BASE that its referent refines, so it sits UNDER the config naming it. Assign it last
268
+ and a role declaring `preset:` silently discards both its own fields and the caller's override — that
269
+ is how effort-tier token caps and `temperatureFactory`'s temperature disappear for preset-based
270
+ roles. An override naming a preset (how the execution layer delivers a `modelOverrides` string pin)
271
+ outranks the alias but yields to explicit override fields; resolution is ONE level deep, so a preset
272
+ meant to carry a model must name one.
319
273
 
320
274
  ## Resilience already handled — do not reimplement
321
275
 
322
- Idle stream deadline · duplicate-final-chunk dedup · output-budget escalation · reasoning-cap
323
- shrink · adaptive-thinking budget floor · same-family fallback model · caller-seeded ladder position (`escalation`) · schema
324
- coercion · JSON salvage from prose · `NullCapture`
325
- diagnostics · fatal-error short-circuit · blank-content sanitization (whitespace-only text
326
- blocks are dropped before every call — a blank block, e.g. an empty file read pasted into a
327
- prompt, is otherwise a fatal Anthropic 400; blank tool results are stubbed to keep their
328
- `tool_use` pairing). Details: package `README.md`.
276
+ Idle stream deadline · duplicate-final-chunk dedup · output-budget escalation · reasoning-cap shrink ·
277
+ adaptive-thinking budget floor · same-family fallback model · caller-seeded ladder position
278
+ (`escalation`) · schema coercion · JSON salvage from prose · `NullCapture` diagnostics · fatal-error
279
+ short-circuit · blank-content sanitization (whitespace-only text blocks are dropped before every call
280
+ — a blank block, e.g. an empty file read pasted into a prompt, is otherwise a fatal Anthropic 400;
281
+ blank tool results are stubbed to keep their `tool_use` pairing). Details: package `README.md`.
329
282
 
330
283
  ## Tests
331
284
 
332
- `bun test ./tests` in the package. Offline specs always run; `tests/model.spec.ts` is gated
333
- on `OPENROUTER_SECRET` / `ANTHROPIC_SECRET` in the repo-root `.env` and self-skips otherwise.
285
+ `bun test ./tests` in the package; offline specs always run. In `tests/model.spec.ts` the anthropic live
286
+ suite is gated on `ANTHROPIC_SECRET` and self-skips with a printed reason without it, and the OpenRouter
287
+ suite is disabled unconditionally: an aggregator on a separate account serving models no deployment runs,
288
+ whose `402 requires more credits` reads as a failure of the code under test. `plugins.spec.ts` covers the
289
+ `Compatible` provider offline.
334
290
 
335
291
  ## Depends On
336
292
 
337
293
  - `@owlmeans/llm-common` · `@owlmeans/context` · `@owlmeans/error` · `@owlmeans/basic-ids` · `ajv`
294
+ - `@anthropic-ai/sdk` — runtime, for the `BadRequestError` the fatal-error rules are written around
338
295
  - peer `@langchain/core`, `@langchain/openai`, `@langchain/anthropic`
339
296
 
340
297
  ## Related
341
298
 
342
- - [[llm-common]] — the serializable contracts
343
- - [[llm-prompt-caching]] — prompt composition, block order and the cache invariants
299
+ - [[llm-common]] — the serializable contracts · [[llm-prompt-caching]] — prompt composition, block
300
+ order and the cache invariants
344
301
  - [[context]] — service registration · [[error]] — the `ResilientError` family