@owlmeans/llm 0.1.18-rc.12 → 0.1.18-rc.13
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +2 -2
- package/agent-meta/manifest.json +2 -2
- package/agent-meta/skills/llm/SKILL.md +189 -232
- package/agent-meta/skills/llm-prompt-caching/SKILL.md +48 -4
- package/build/errors.d.ts.map +1 -1
- package/build/errors.js.map +1 -1
- package/build/execution/service.d.ts.map +1 -1
- package/build/execution/service.js.map +1 -1
- package/build/execution/utils.d.ts.map +1 -1
- package/build/helpers/cache.d.ts.map +1 -1
- package/build/helpers/json.d.ts.map +1 -1
- package/build/helpers/messages.d.ts.map +1 -1
- package/build/helpers/retry.d.ts +14 -0
- package/build/helpers/retry.d.ts.map +1 -1
- package/build/helpers/retry.js +14 -0
- package/build/helpers/retry.js.map +1 -1
- package/build/helpers/spectate.d.ts.map +1 -1
- package/build/model.d.ts.map +1 -1
- package/build/model.js +3 -1
- package/build/model.js.map +1 -1
- package/build/plugins/anthropic.d.ts +15 -3
- package/build/plugins/anthropic.d.ts.map +1 -1
- package/build/plugins/anthropic.js +18 -3
- package/build/plugins/anthropic.js.map +1 -1
- package/build/plugins/index.d.ts.map +1 -1
- package/build/plugins/openai.d.ts.map +1 -1
- package/build/plugins/types.d.ts +7 -0
- package/build/plugins/types.d.ts.map +1 -1
- package/build/plugins/utils.d.ts.map +1 -1
- package/build/prompt/render.d.ts.map +1 -1
- package/build/prompt/service.d.ts.map +1 -1
- package/build/prompt/service.js.map +1 -1
- package/build/service.d.ts.map +1 -1
- package/build/service.js.map +1 -1
- package/build/utils/config.d.ts.map +1 -1
- package/build/utils/null-report.d.ts.map +1 -1
- package/build/utils/prompt.d.ts.map +1 -1
- package/build/utils/schema.d.ts.map +1 -1
- package/build/utils/stream.d.ts.map +1 -1
- package/build/utils/stream.js.map +1 -1
- package/package.json +6 -6
- package/src/helpers/retry.ts +16 -0
- package/src/model.ts +3 -1
- package/src/plugins/anthropic.ts +22 -3
- package/src/plugins/types.ts +8 -0
- package/tests/plugins.spec.ts +30 -0
package/README.md
CHANGED
|
@@ -20,7 +20,7 @@ abstraction that resolves models from an inheritable policy.
|
|
|
20
20
|
## Installation
|
|
21
21
|
|
|
22
22
|
```bash
|
|
23
|
-
bun add @owlmeans/llm @owlmeans/llm-common
|
|
23
|
+
bun add @owlmeans/llm@^0.1.18-rc.12 @owlmeans/llm-common@^0.1.18-rc.11
|
|
24
24
|
bun add @langchain/core @langchain/openai @langchain/anthropic # peer dependencies
|
|
25
25
|
```
|
|
26
26
|
|
|
@@ -184,7 +184,7 @@ This package ships embedded agent skills under `agent-meta/`. After installing y
|
|
|
184
184
|
your project's skill store (`.agents/skills/`):
|
|
185
185
|
|
|
186
186
|
```sh
|
|
187
|
-
npx @owlmeans/agent-skills
|
|
187
|
+
npx @owlmeans/agent-skills@^0.1.18-rc.12
|
|
188
188
|
```
|
|
189
189
|
|
|
190
190
|
The embedded files are version-matched to this package release. Do not edit them
|
package/agent-meta/manifest.json
CHANGED
|
@@ -1,8 +1,8 @@
|
|
|
1
1
|
{
|
|
2
2
|
"schemaVersion": 2,
|
|
3
3
|
"package": "@owlmeans/llm",
|
|
4
|
-
"version": "0.1.18-rc.
|
|
5
|
-
"generatedAt": "2026-09-
|
|
4
|
+
"version": "0.1.18-rc.13",
|
|
5
|
+
"generatedAt": "2026-09-04T22:43:25.441Z",
|
|
6
6
|
"canonicalRepo": "https://github.com/owlmeans/common",
|
|
7
7
|
"entries": [
|
|
8
8
|
{
|
|
@@ -8,48 +8,37 @@ user-invocable: false
|
|
|
8
8
|
# @owlmeans/llm
|
|
9
9
|
|
|
10
10
|
**Layer:** Core
|
|
11
|
-
**Install:** `"@owlmeans/llm": "^0.1.18-rc.
|
|
11
|
+
**Install:** `"@owlmeans/llm": "^0.1.18-rc.13"` in `dependencies` (plus the `@langchain/*` peers)
|
|
12
12
|
|
|
13
|
-
The inference runtime. Everything provider-specific is a **plugin**; the model itself only
|
|
14
|
-
|
|
15
|
-
|
|
13
|
+
The inference runtime. Everything provider-specific is a **plugin**; the model itself only owns the
|
|
14
|
+
provider-independent parts (streaming discipline, retries, validation, observability). Serializable
|
|
15
|
+
contracts live in `@owlmeans/llm-common`. `src/helpers/` is exported — functions a **consumer** may
|
|
16
|
+
use alongside a model; `src/utils/` is internal and **never exported** (a spec that needs one imports
|
|
17
|
+
it from `../src/utils/…`). Decide the side before placing a function, and never export a `utils/`
|
|
18
|
+
symbol "because a test needs it".
|
|
16
19
|
|
|
17
20
|
## Key Exports
|
|
18
21
|
|
|
19
22
|
| Export | Description |
|
|
20
23
|
|--------|-------------|
|
|
21
|
-
| `makeLlmModel(
|
|
22
|
-
| `makeLlmService(options, alias?)` · `appendLlmService(ctx, options, alias?)` | Model factory/registry — resolves a `ModelConfig` by alias, memoized per alias+override. |
|
|
23
|
-
| `
|
|
24
|
-
| `makeExecutionService(alias?, options?)` · `appendExecutionService(ctx, alias?, options?)` | Frozen 3-level executions + policy resolution + snapshot/restore/checkpoint + advice. |
|
|
24
|
+
| `makeLlmModel({ model, purpose, prompt?, prompts?, files?, utility?, retries?, … }, spectator)` | The four-method model: `ask` / `talk` / `invoke(input, schema, opts)` / `request`. `model` is an already-resolved `BaseChatModel`. |
|
|
25
|
+
| `makeLlmService(options, alias?)` · `appendLlmService(ctx, options, alias?)` · `llmServiceApi(options, self)` | Model factory/registry — `makeLlmService({ models: () => configs }).getModel(alias, override?)` resolves a `ModelConfig` by alias, memoized per alias+override. The `…Api` half omits `createService`, to compose into your own service (role accessors, domain helpers). |
|
|
26
|
+
| `makeExecutionService(alias?, options?)` · `appendExecutionService(ctx, alias?, options?)` · `executionServiceApi(options, self)` | Frozen 3-level executions + policy resolution + snapshot/restore/checkpoint + advice, and the same composable half. |
|
|
25
27
|
| `makePromptService(options?, alias?)` · `appendPromptService(ctx, options?, alias?)` · `promptServiceApi(options, self)` | Skill registry + the composition plugin chain. Also at `@owlmeans/llm/prompt`. |
|
|
26
28
|
| `rolePlugin`, `skillsPlugin`, `contextPlugin`, `BUILT_IN_PROMPT_PLUGINS` | The built-in composition plugins. |
|
|
27
29
|
| `PromptContext.claim(key)` · `PromptComposeParams.utility` | Per-composition ownership of a key; a cheap model for one plugin-side call. |
|
|
28
|
-
| `renderSkill`, `sortSkills`, `joinChunks`, `compareAlias`, `prefixHash` | Deterministic rendering primitives — reuse them, never re-implement. |
|
|
29
|
-
| `readCacheUsage`, `hasCacheActivity` | Normalized prompt-cache accounting from a completion. |
|
|
30
|
-
| `executionServiceApi(options, self)` | The execution half without `createService`, for the same composition pattern. |
|
|
30
|
+
| `renderSkill`, `sortSkills`, `joinChunks`, `compareAlias`, `prefixHash` · `readCacheUsage`, `hasCacheActivity` | Deterministic rendering primitives — reuse them, never re-implement — and normalized prompt-cache accounting from a completion. |
|
|
31
31
|
| `plugins`, `registerLlmPlugin`, `resolvePlugin`, `pluginOf`, `pluginFor` | The provider-plugin registry. Also at `@owlmeans/llm/plugins`. |
|
|
32
32
|
| `anthropicPlugin`, `openAiPlugin`, `compatiblePlugin`, `openAiFamily` | Built-in providers; `openAiFamily` is the shared OpenAI-client behaviour to spread into a new plugin. |
|
|
33
|
-
| `NO_SAMPLING_PREFIXES
|
|
34
|
-
| `
|
|
35
|
-
| `withRetry`, `registerFatalError`, `spectate`, `normalizeInput`, `parseJsonContent`, `coerceToSchema` | Helpers usable alongside a model. Also at `@owlmeans/llm/helpers`. |
|
|
33
|
+
| `NO_SAMPLING_PREFIXES` / `rejectsSampling(model)` · `RESPONSES_API_PREFIXES` / `usesResponsesApi(model)` | Which families reject which sampling parameters — see the table below. Consumers pin presets against them. |
|
|
34
|
+
| `withRetry`, `registerFatalError`, `isFatalError`, `spectate`, `normalizeInput`, `parseJsonContent`, `coerceToSchema` | Helpers usable alongside a model. Also at `@owlmeans/llm/helpers`. |
|
|
36
35
|
| `LlmError`, `LlmModelError`, `LlmMissconfiguredError`, `LlmPluginError`, `LlmRetryExceededError` | `ResilientError` family. `LlmModelError` is the RETRYABLE one. |
|
|
37
36
|
| `mergePrompt`, `mergePolicy`, `resolveRole`, `effortPatch` | Execution merge helpers; `mergePrompt` unions skills and takes the deepest role. |
|
|
38
37
|
| `DEFAULT_MODEL_RETRIES`, `MODEL_STREAM_TIMEOUT_MS` (3 min idle), `FALLBACK_AFTER_ATTEMPTS`, `DEFAULT_EFFORT`, `EFFORT_TABLE`, `MAX_CACHE_BREAKPOINTS`, `MAX_SYSTEM_BREAKPOINTS`, `MIN_CACHEABLE_TOKENS`, `LLM_SERVICE`, `EXECUTION_SERVICE`, `PROMPT_SERVICE` | Tuning + aliases. |
|
|
39
38
|
|
|
40
|
-
## helpers/ vs utils/ — the rule this package follows
|
|
41
|
-
|
|
42
|
-
- `src/helpers/` — functions a **consumer** may use alongside a model. Exported.
|
|
43
|
-
- `src/utils/` — used only inside the library. **Never exported**; a spec that needs one
|
|
44
|
-
imports it from `../src/utils/…`.
|
|
45
|
-
|
|
46
|
-
Adding a function? Decide which side it belongs to first, then place it. Do not export a
|
|
47
|
-
`utils/` symbol "because a test needs it".
|
|
48
|
-
|
|
49
39
|
## Provider differences are plugins, never `if`s
|
|
50
40
|
|
|
51
|
-
`LlmPlugin` is the single seam
|
|
52
|
-
`config.provider === …` in `model.ts` or `service.ts`, it belongs on the plugin instead:
|
|
41
|
+
`LlmPlugin` is the single seam: an `instanceof ChatAnthropic` or `config.provider === …` in `model.ts` or `service.ts` belongs on the plugin.
|
|
53
42
|
|
|
54
43
|
| Plugin member | Replaces |
|
|
55
44
|
|---|---|
|
|
@@ -62,198 +51,155 @@ Adding a function? Decide which side it belongs to first, then place it. Do not
|
|
|
62
51
|
| `patchCache` | the message-prefix cache marker |
|
|
63
52
|
| `cacheKey` (via `ModelConfig.cacheKey`) | provider cache-routing hints such as OpenAI's `prompt_cache_key` |
|
|
64
53
|
| `isFatal` | "this error can never be retried" |
|
|
54
|
+
| `suppressesThinking` | whether the plugin turns reasoning off on the wire FOR THIS CONFIG, which is what drops the `/no_think` prompt directive from the prepared messages |
|
|
65
55
|
|
|
66
56
|
### A provider that removes a parameter is a plugin concern too
|
|
67
57
|
|
|
68
|
-
`build` is not the only place a parameter reaches the wire, and `refine` is not only a retry
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
Both built-in plugins gate their rejected parameters in **both** hooks, through a predicate the
|
|
75
|
-
package root exports:
|
|
58
|
+
`build` is not the only place a parameter reaches the wire, and `refine` is not only a retry hook —
|
|
59
|
+
**every** call is made on the instance `refine` returns, attempt 0 included. A knob suppressed in
|
|
60
|
+
`build` and re-applied in `refine` therefore ships on every single request: an unsupported
|
|
61
|
+
`temperature` restored there 400s the model's whole family on the first call, and `isFatal`
|
|
62
|
+
correctly refuses to retry it. Both built-in plugins gate their rejected parameters in **both**
|
|
63
|
+
hooks, through a predicate the package root exports:
|
|
76
64
|
|
|
77
65
|
| Family | Rejects | Predicate |
|
|
78
66
|
|---|---|---|
|
|
79
67
|
| Claude 4.7+ and the 5 family | `temperature`, `top_p`, `top_k` | `NO_SAMPLING_PREFIXES` / `rejectsSampling(model)` |
|
|
80
68
|
| OpenAI Responses API (`gpt-5*`, `codex-*`) | `temperature`, `top_p` | `RESPONSES_API_PREFIXES` / `usesResponsesApi(model)` |
|
|
81
69
|
|
|
82
|
-
Models below those lines keep the deterministic `temperature: 0` default. `refine`
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
`
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
(`viable-agent`'s `tests/presets.test.ts` asserts no preset entry declares a parameter its
|
|
89
|
-
model rejects), and a second hand-written copy of the list drifts exactly when a new family is
|
|
90
|
-
added — the moment it has to be right.
|
|
70
|
+
Models below those lines keep the deterministic `temperature: 0` default. Each `refine` re-derives
|
|
71
|
+
the family from the ACTIVE base instance it is handed — Anthropic from `modelName ?? model`, OpenAI
|
|
72
|
+
from `model ?? lc_kwargs.model` plus a `useResponsesApi` already on `lc_kwargs` — so a same-family
|
|
73
|
+
`fallback` rung is judged on its own id, not the primary's. Keep the predicate exported: consumers
|
|
74
|
+
pin presets against it (viable-agent's `tests/presets.test.ts` asserts no preset entry declares a
|
|
75
|
+
parameter its model rejects), and a second hand-written copy drifts when a family is added.
|
|
91
76
|
|
|
92
|
-
**Registration order is load-bearing.** Instance
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
77
|
+
**Registration order is load-bearing.** Instance lookup (`pluginFor`) returns the FIRST plugin whose
|
|
78
|
+
`owns` matches, and `compatible` is registered before `openai` because both build a `ChatOpenAI`:
|
|
79
|
+
assuming the tool-calling hack for an unlabelled model is safe everywhere, assuming native
|
|
80
|
+
JSON-schema support is not.
|
|
96
81
|
|
|
97
82
|
## System prompts: a role and skills, never a hand-built message
|
|
98
83
|
|
|
99
84
|
`makeLlmModel` takes `prompt` (a `PromptInput`) and a `prompts` resolver. The prompt service
|
|
100
|
-
composes them into an ordered, cacheable system message; a caller's own leading
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
declared `role` wins — see `mergePrompt`. Full rules, breakpoint budget and the provider
|
|
108
|
-
facts behind them: [[llm-prompt-caching]].
|
|
109
|
-
|
|
110
|
-
Two seams exist for plugins that would otherwise collide or overspend:
|
|
111
|
-
`PromptContext.claim(key)` grants a key to the first plugin that asks within one
|
|
112
|
-
composition, so a static catalogue and a detector never emit the same skill twice; and
|
|
113
|
-
`PromptComposeParams.utility` offers a cheap model for ONE side call. Both are threaded
|
|
114
|
-
from where a prompt is composed — `makeLlmModel`'s `utility` option (beside `files`) and
|
|
115
|
-
`AgentOptions.utility`. Neither may change a cached block: see [[llm-prompt-caching]].
|
|
85
|
+
composes them into an ordered, cacheable system message; a caller's own leading `SystemMessage` is
|
|
86
|
+
folded into the volatile `Context` block, so an unmigrated call site still works. **Do not build a
|
|
87
|
+
persona as a `SystemMessage` in a helper** — declare it as `PromptPolicy.role` plus registered
|
|
88
|
+
skills, or the knowledge duplicates and the cache prefix stops being stable. Skills accumulate down
|
|
89
|
+
the execution chain (project → task → helper) and the deepest declared `role` wins (`mergePrompt`).
|
|
90
|
+
Block order, the breakpoint budget, the provider facts behind them, and the plugin seams
|
|
91
|
+
`claim(key)` / `utility`: [[llm-prompt-caching]].
|
|
116
92
|
|
|
117
93
|
## Execution: policy in, model out
|
|
118
94
|
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
└─ TaskExecution ← + resumable state (phase/cursor/completed/data)
|
|
122
|
-
└─ HelperExecution ← + a RESOLVED model + temperatureFactory, bound to a role
|
|
123
|
-
```
|
|
95
|
+
`ProjectExecution` (policy + purpose + models resolver) → `TaskExecution` (+ resumable state:
|
|
96
|
+
phase/cursor/completed/data) → `HelperExecution` (+ a RESOLVED model + `temperatureFactory`, a role).
|
|
124
97
|
|
|
125
|
-
`prompt` (role + skills) travels on `ExecutionState`, so it survives snapshot/restore;
|
|
126
|
-
`
|
|
98
|
+
`prompt` (role + skills) travels on `ExecutionState`, so it survives snapshot/restore; `prompts` and
|
|
99
|
+
`files` are collaborators and never enter a snapshot. Every method returns a NEW `Object.freeze`d
|
|
100
|
+
object. Resolution precedence in `model(exec, role, override)`: **roleOverride → modelOverride →
|
|
101
|
+
effort tier → `LlmService.getModel`**; `escalate(exec, { effort })` raises the tier once and cascades
|
|
102
|
+
to everything derived from it. `snapshot` excludes `state` itself — without that, every
|
|
103
|
+
`derive`/`escalate`/`withPurpose` on a task would nest another copy of the previous state.
|
|
127
104
|
|
|
128
|
-
|
|
129
|
-
`
|
|
130
|
-
|
|
131
|
-
everything derived from it.
|
|
105
|
+
Extending it for a domain: declare your own `Execution`/input types, list collaborator fields in
|
|
106
|
+
`ExecutionServiceOptions.collaboratorKeys` (kept out of snapshots), instantiate the service generic
|
|
107
|
+
with your own `ExecutionShape`, and **never narrow an inherited method signature** (contravariance).
|
|
132
108
|
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
inherited method signatures**, which would be a contravariance error.
|
|
137
|
-
|
|
138
|
-
`snapshot` excludes `state` itself; without that, every `derive`/`escalate`/`withPurpose`
|
|
139
|
-
on a task would nest another copy of the previous state.
|
|
140
|
-
|
|
141
|
-
`forHelper` also accepts `output` — an initial `maxTokens` for a helper whose one answer is
|
|
142
|
-
genuinely large (a whole source file rather than a decision). It is a model selector, not a
|
|
143
|
-
field the helper carries: it becomes a call override, is clamped to `maxOutput`, doubles from
|
|
144
|
-
there under retry, and survives `temperatureFactory`.
|
|
109
|
+
`forHelper` also accepts `output` — an initial `maxTokens` for a helper whose one answer is genuinely
|
|
110
|
+
large (a whole source file, not a decision). It is a model selector, not a field the helper carries:
|
|
111
|
+
it becomes a call override, clamped to `maxOutput`, doubling under retry, surviving `temperatureFactory`.
|
|
145
112
|
|
|
146
113
|
### The cheap tier: `utility(exec, override?)`
|
|
147
114
|
|
|
148
|
-
Work that is not the work — a relevance pick, a classification, a one-line judgement a
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
|
|
152
|
-
|
|
153
|
-
|
|
154
|
-
|
|
155
|
-
|
|
156
|
-
|
|
157
|
-
|
|
158
|
-
|
|
159
|
-
|
|
115
|
+
Work that is not the work — a relevance pick, a classification, a one-line judgement a plugin needs
|
|
116
|
+
before the real call can be shaped — runs on `ExecutionService.utility`, never on the helper's own
|
|
117
|
+
model. It resolves `policy.utilityRole ?? UTILITY_ROLE` (`@owlmeans/llm-common`, value `'utility'`)
|
|
118
|
+
at `ExecutionEffort.Economy`, through the SAME ladder as `model()`: `roleOverrides` remap it and
|
|
119
|
+
`modelOverrides` pin it as for any other role. The economy floor is local — the execution it was
|
|
120
|
+
asked on keeps its own tier — and `utilityRole` travels on `ModelPolicy`, so it survives
|
|
121
|
+
`forTask`/`escalate` and a snapshot/restore round trip. `utility` returns a `BaseChatModel`, never
|
|
122
|
+
`undefined`: an alias with no registered config reaches `createModel` through `model()` and throws
|
|
123
|
+
`LlmMissconfiguredError`, like any other unregistered role. The `undefined` a prompt plugin has to
|
|
124
|
+
handle comes from the other end — `PromptComposeParams.utility` (and `AgentOptions.utility`) is an
|
|
125
|
+
OPTIONAL resolver, unset wherever no cheap tier is wired, so a plugin that cannot get one degrades
|
|
126
|
+
rather than fails.
|
|
160
127
|
|
|
161
128
|
### The plugin seam has two hooks, dispatched independently
|
|
162
129
|
|
|
163
|
-
`ExecutionPlugin` carries `onCheckpoint`/`onRestore` (persist and resume an
|
|
164
|
-
`
|
|
165
|
-
|
|
166
|
-
|
|
167
|
-
appends whatever comes back and proceeds unchanged on
|
|
168
|
-
|
|
169
|
-
|
|
170
|
-
otherwise registering an advise-only plugin would silently start composing snapshots nobody
|
|
171
|
-
consumes.
|
|
130
|
+
`ExecutionPlugin` carries `onCheckpoint`/`onRestore` (persist and resume an `ExecutionState`) and
|
|
131
|
+
`advise` (answer a performer's question about the project it works in —
|
|
132
|
+
`ExecutionService.advise(exec, request)`, first usable answer wins, a throwing plugin is skipped,
|
|
133
|
+
`null` when nobody answers).
|
|
134
|
+
Advice is advisory by contract: a caller appends whatever comes back and proceeds unchanged on
|
|
135
|
+
`null`. `checkpoint` dispatches on plugins declaring **`onCheckpoint`**, not on the plugin count, or
|
|
136
|
+
an advise-only plugin would silently start composing snapshots nobody consumes.
|
|
172
137
|
|
|
173
138
|
## Classify a provider error by walking `cause`, never by `instanceof` or a surface read
|
|
174
139
|
|
|
175
|
-
Two independent layers hide the wire shape of a provider failure, and each
|
|
176
|
-
|
|
177
|
-
message buried under the repeats.
|
|
140
|
+
Two independent layers hide the wire shape of a provider failure, and each alone makes a fatal error
|
|
141
|
+
look retryable — costing the whole retry budget with the real message buried under the repeats.
|
|
178
142
|
|
|
179
|
-
1. **Nested SDK copies.** `@langchain/anthropic` and `@langchain/openai` bundle their OWN
|
|
180
|
-
|
|
181
|
-
|
|
182
|
-
`false`.
|
|
143
|
+
1. **Nested SDK copies.** `@langchain/anthropic` and `@langchain/openai` bundle their OWN copies of
|
|
144
|
+
the provider SDKs, so `e instanceof BadRequestError` compares against a DIFFERENT class than the
|
|
145
|
+
one this package imports and silently returns `false`.
|
|
183
146
|
2. **Langchain's own error wrappers.** A failure is re-wrapped in a typed langchain error
|
|
184
|
-
(`ContextOverflowError` for an input past the context window, and its siblings)
|
|
185
|
-
|
|
186
|
-
`e.status === 400` misses it too.
|
|
147
|
+
(`ContextOverflowError` for an input past the context window, and its siblings) carrying the
|
|
148
|
+
original **only under `cause`**, no `status` of its own — so `e.status === 400` misses it.
|
|
187
149
|
|
|
188
|
-
Use `isBadRequest` from `plugins/utils.ts`: it walks the `cause` chain
|
|
189
|
-
|
|
190
|
-
`
|
|
191
|
-
|
|
192
|
-
|
|
193
|
-
|
|
194
|
-
every attempt re-sends the identical oversized request. Consumers hold their locks for the
|
|
195
|
-
whole loop, so a single unfixable call becomes minutes of thrash on the caller's side.
|
|
196
|
-
|
|
197
|
-
The same trap applies to any cross-copy `instanceof`; it is the runtime face of the
|
|
198
|
-
peer-dependency identity rule below.
|
|
150
|
+
Use `isBadRequest` from `plugins/utils.ts`: it walks the `cause` chain for `status === 400`, bounded
|
|
151
|
+
in depth so a self-referential chain terminates. Both built-in `isFatal` implementations go through
|
|
152
|
+
it. A context overflow makes this urgent: `refine` escalates the **output** budget on each retry, so
|
|
153
|
+
an over-limit **input** can never improve — every attempt re-sends the identical oversized request,
|
|
154
|
+
and consumers hold their locks for the whole loop. The same trap applies to any cross-copy
|
|
155
|
+
`instanceof`; it is the runtime face of the peer-dependency rule below.
|
|
199
156
|
|
|
200
157
|
## Peer-dependency rule (langchain identity)
|
|
201
158
|
|
|
202
|
-
`@langchain/core`, `@langchain/openai` and `@langchain/anthropic` are **peer** dependencies:
|
|
203
|
-
|
|
204
|
-
|
|
205
|
-
|
|
206
|
-
|
|
207
|
-
## Usage
|
|
208
|
-
|
|
209
|
-
```typescript
|
|
210
|
-
import { makeLlmModel, makeLlmService } from '@owlmeans/llm'
|
|
211
|
-
import { ModelProvider } from '@owlmeans/llm-common'
|
|
212
|
-
|
|
213
|
-
const llm = makeLlmService({ models: () => configs })
|
|
214
|
-
const model = makeLlmModel(
|
|
215
|
-
{ model: llm.getModel('analyst'), purpose: { type: 'analysis' } }, spectator
|
|
216
|
-
)
|
|
217
|
-
const spec = await model.invoke('Describe the app', SpecSchema, { action: 'spec' })
|
|
218
|
-
```
|
|
159
|
+
`@langchain/core`, `@langchain/openai` and `@langchain/anthropic` are **peer** dependencies: model
|
|
160
|
+
instances cross the package boundary, and two installed copies of a class with protected members are
|
|
161
|
+
nominally distinct types. A consumer must end up with exactly ONE copy of each — declare them at one
|
|
162
|
+
range and check no nested `node_modules` holds a second, or every model instance crossing a boundary
|
|
163
|
+
becomes a foreign type and `instanceof` starts lying.
|
|
219
164
|
|
|
220
165
|
## Hangs are bounded by an IDLE deadline, not a total one
|
|
221
166
|
|
|
222
|
-
A stalled provider is aborted after `MODEL_STREAM_TIMEOUT_MS` (3 min) of SILENCE and
|
|
223
|
-
|
|
224
|
-
|
|
225
|
-
|
|
226
|
-
|
|
227
|
-
|
|
228
|
-
|
|
229
|
-
Note what this does NOT bound: a call that keeps streaming forever, and retries. A fatal
|
|
230
|
-
error misclassified as retryable multiplies its own latency by `DEFAULT_MODEL_RETRIES` —
|
|
231
|
-
see the `instanceof` trap above.
|
|
167
|
+
A stalled provider is aborted after `MODEL_STREAM_TIMEOUT_MS` (3 min) of SILENCE and surfaces as a
|
|
168
|
+
retryable `LlmModelError`, so the escalator moves on. The timer re-arms on every token, so a
|
|
169
|
+
long-but-productive generation is never cut off — which is why the value can be low. Set it for a
|
|
170
|
+
deployment with `LlmServiceOptions.streamTimeout` where the application composes its context; a
|
|
171
|
+
preset naming its own `ModelConfig.streamTimeout` keeps it. It does NOT bound a call that keeps
|
|
172
|
+
streaming forever, nor retries — a fatal error misclassified as retryable multiplies its own latency
|
|
173
|
+
by `DEFAULT_MODEL_RETRIES`.
|
|
232
174
|
|
|
233
175
|
## An outer validation loop must pass its attempt in
|
|
234
176
|
|
|
235
|
-
A caller that validates the OUTPUT — a diff that has to apply, a file that must not come
|
|
236
|
-
|
|
237
|
-
|
|
238
|
-
same
|
|
239
|
-
|
|
240
|
-
`
|
|
241
|
-
|
|
242
|
-
|
|
243
|
-
|
|
244
|
-
|
|
245
|
-
|
|
246
|
-
|
|
247
|
-
|
|
248
|
-
|
|
249
|
-
|
|
250
|
-
|
|
251
|
-
|
|
177
|
+
A caller that validates the OUTPUT — a diff that has to apply, a file that must not come back
|
|
178
|
+
truncated — runs its own retry loop around whole `ask` calls. Every one of those calls starts a FRESH
|
|
179
|
+
inner loop at attempt 0, so the escalator's two rungs never move: same model, same output budget,
|
|
180
|
+
same deterministic answer, N times. Pass `escalation: <outer attempt>` in `LlmCallOptions` and the
|
|
181
|
+
per-call escalator starts that far up its ladder instead — `maxTokens` doubling and the
|
|
182
|
+
`FALLBACK_AFTER_ATTEMPTS` switch to `ModelConfig.fallback` both advance. It is clamped to
|
|
183
|
+
`retries - 1`, moves the STARTING rung only, and never changes how many attempts the call makes. Two
|
|
184
|
+
things it depends on, both preset data rather than code: the role must declare a `fallback`, and that
|
|
185
|
+
fallback must be in the same plugin `family` (a cross-family one is skipped with a warning, because
|
|
186
|
+
switching provider mid-call flips the structured-output shape). `LlmCallOptions.fatal` is the lever
|
|
187
|
+
in the other direction — a per-call resolver consulted before the global ones and the plugin's
|
|
188
|
+
`isFatal`, for an error the caller knows no retry can fix.
|
|
189
|
+
|
|
190
|
+
### A loop ABOVE the model asks the same question with `isFatalError`
|
|
191
|
+
|
|
192
|
+
A retry loop is not the only place that decides to carry on: a fix ladder rescues a failed repair and
|
|
193
|
+
climbs to a stronger model, an agent runner catches a round that threw and reports "gave up". Both
|
|
194
|
+
are right for a model that answered badly and wrong for a budget that ran out, and a blanket `catch`
|
|
195
|
+
cannot tell them apart — an exhausted balance becomes more expensive calls instead of a halt.
|
|
196
|
+
`isFatalError(e, fatal?)` runs the same resolvers, in the same order, that `withRetry` uses, and
|
|
197
|
+
returns the error to abort WITH (a resolver may unwrap a carrier and hand back the real cause) or
|
|
198
|
+
`null` when nothing considers it terminal. Ask it rather than re-deriving the rule.
|
|
252
199
|
|
|
253
200
|
## Output caps: what the deployment wants vs what the provider allows
|
|
254
201
|
|
|
255
|
-
Four fields, and conflating them
|
|
256
|
-
run:
|
|
202
|
+
Four fields, and conflating them turns an escalation into a fatal 400 hours into a run:
|
|
257
203
|
|
|
258
204
|
| Field | Means |
|
|
259
205
|
|---|---|
|
|
@@ -262,83 +208,94 @@ run:
|
|
|
262
208
|
| `maxOutput` | what the PROVIDER accepts in one request — a fact about the model |
|
|
263
209
|
| `contextWindow` | total window (input + output); informational, never sent |
|
|
264
210
|
|
|
265
|
-
`resolveOutputCap` (`utils/config.ts`)
|
|
266
|
-
|
|
267
|
-
|
|
268
|
-
|
|
269
|
-
`
|
|
270
|
-
|
|
271
|
-
|
|
272
|
-
|
|
273
|
-
|
|
274
|
-
|
|
275
|
-
|
|
276
|
-
|
|
277
|
-
|
|
278
|
-
|
|
279
|
-
|
|
280
|
-
|
|
281
|
-
|
|
282
|
-
|
|
283
|
-
|
|
284
|
-
the
|
|
285
|
-
|
|
286
|
-
|
|
287
|
-
|
|
288
|
-
|
|
289
|
-
|
|
290
|
-
|
|
291
|
-
|
|
292
|
-
|
|
293
|
-
-
|
|
294
|
-
|
|
295
|
-
|
|
296
|
-
|
|
211
|
+
`resolveOutputCap` (`utils/config.ts`) reconciles them: the declared cap chooses the ceiling and the
|
|
212
|
+
capability trims it, and `DEFAULT_MAX_OUTPUT_CAP` applies only when neither is stated. `createModel`
|
|
213
|
+
also clamps `maxTokens` to `maxOutput` and warns about a cap above it. For an aggregated model
|
|
214
|
+
`maxOutput` is the limit of the `inferenceProvider` actually pinned, often far below what the model
|
|
215
|
+
can do elsewhere. `combinedWindow: true` marks a model whose window is shared between input and
|
|
216
|
+
output (MiniMax M2.x, gpt-oss) — nothing enforces it at runtime; it keeps presets honest about
|
|
217
|
+
leaving room for the prompt. **A `fallback` that changes `model` must restate
|
|
218
|
+
`contextWindow`/`maxOutput`** (and reset `combinedWindow`): the fallback config is
|
|
219
|
+
`{...primaryConfig, ...fallback}`, so every field the patch does not name is inherited from a
|
|
220
|
+
different model.
|
|
221
|
+
|
|
222
|
+
### Reasoning is off unless a preset asks for it — and it is billed against the same budget
|
|
223
|
+
|
|
224
|
+
The models `NO_SAMPLING_PREFIXES` names think ADAPTIVELY unless the request says otherwise: an absent
|
|
225
|
+
`thinking` parameter means adaptive, and `@langchain/anthropic` forwards the parameter only when the
|
|
226
|
+
caller sets it (its own field default of `disabled` is never sent). Left on, the reasoning costs
|
|
227
|
+
twice — it is spent from `max_tokens`, the same allowance as the answer (a budget sized for the
|
|
228
|
+
answer alone goes entirely on reasoning and the completion arrives well-formed, `stop_reason:
|
|
229
|
+
"max_tokens"`, with no text block), and its SUMMARISED stream arrives in bursts minutes apart, which
|
|
230
|
+
the idle deadline reads as a dead connection and retries from scratch
|
|
231
|
+
(`llm:model:stream-stalled:no token for 180000ms (idle deadline)`).
|
|
232
|
+
|
|
233
|
+
**`ModelConfig.disableThinking` is the switch, and where it lands depends on the plugin.**
|
|
234
|
+
`makeLlmModel` appends the literal `/no_think` to every request's prepared messages whenever the flag
|
|
235
|
+
is set AND the plugin's `suppressesThinking(config)` does not answer `true` — the soft switch for
|
|
236
|
+
models with no request-level control (Qwen3). The Anthropic plugin answers `true` only for
|
|
237
|
+
`rejectsSampling(model)`, and for those puts `thinking: { type: 'disabled' }` on the request in
|
|
238
|
+
`build`, which `refine` carries through `lc_kwargs` on every attempt. Below that line
|
|
239
|
+
(`claude-haiku-4-5`, `claude-sonnet-4-6`) and under any plugin declaring no hook the flag injects
|
|
240
|
+
prompt text instead, so set it where the wire honours it. **A preset must set the flag on every
|
|
241
|
+
adaptive Anthropic role**; it reaches the `fallback` rung only by inheritance from the entry
|
|
242
|
+
(viable-agent's `presets.test.ts` pins this). Turning reasoning ON is a per-role decision.
|
|
243
|
+
|
|
244
|
+
`ADAPTIVE_MIN_MAX_TOKENS` (32k) is the output floor the Anthropic plugin's `build` applies to every
|
|
245
|
+
`rejectsSampling(model)` config — `disableThinking` is not consulted, so a role with reasoning turned
|
|
246
|
+
off is floored just the same. It is a floor, not an override (a preset asking for more keeps it) and
|
|
247
|
+
it is clamped through `resolveOutputCap`, so it can never exceed what the provider accepts and turn a
|
|
248
|
+
retryable empty answer into a fatal 400. Raising or removing it re-opens empty completions.
|
|
249
|
+
|
|
250
|
+
Two diagnostics the above depends on:
|
|
251
|
+
|
|
252
|
+
- **An empty completion is a null result, not a filter rejection.** `ask` tests emptiness BEFORE the
|
|
253
|
+
caller's filter — every shipped filter returns null only for empty input, so a filter run first
|
|
254
|
+
blames the caller for a provider problem and skips `reportNull`, losing the `finishReason` and
|
|
255
|
+
`outputTokens` that name the cause.
|
|
297
256
|
- **Anthropic's stop reason is not `finish_reason`.** langchain puts it in
|
|
298
|
-
`additional_kwargs.stop_reason
|
|
299
|
-
|
|
300
|
-
reads both, and `thinkingOnly` marks a completion that was all reasoning.
|
|
257
|
+
`additional_kwargs.stop_reason` and `response_metadata.finish_reason` does not exist there, so
|
|
258
|
+
`null-report.ts` reads both; `thinkingOnly` marks a completion that was all reasoning.
|
|
301
259
|
|
|
302
|
-
|
|
303
|
-
|
|
304
|
-
draws again from an unchanged distribution.
|
|
260
|
+
OpenAI reasoning models get it handled by shrinking the reasoning cap on retry (`plugins/openai.ts`);
|
|
261
|
+
escalating `maxTokens` alone fixes neither family — the retry draws from an unchanged distribution.
|
|
305
262
|
|
|
306
263
|
## Config precedence: a preset is a base, not a final word
|
|
307
264
|
|
|
308
|
-
`createModel` layers four sources, lowest first:
|
|
309
|
-
|
|
310
|
-
presetOf(base.preset) < base < presetOf(override.preset) < override
|
|
265
|
+
`createModel` layers four sources, lowest first: `presetOf(base.preset)` < `base` < `presetOf(override.preset)` < `override`.
|
|
311
266
|
|
|
312
|
-
A `preset` is a BASE that its referent refines, so it sits UNDER the config naming it.
|
|
313
|
-
|
|
314
|
-
|
|
315
|
-
|
|
316
|
-
|
|
317
|
-
|
|
318
|
-
ONE level deep: a preset meant to carry a model must name one.
|
|
267
|
+
A `preset` is a BASE that its referent refines, so it sits UNDER the config naming it. Assign it last
|
|
268
|
+
and a role declaring `preset:` silently discards both its own fields and the caller's override — that
|
|
269
|
+
is how effort-tier token caps and `temperatureFactory`'s temperature disappear for preset-based
|
|
270
|
+
roles. An override naming a preset (how the execution layer delivers a `modelOverrides` string pin)
|
|
271
|
+
outranks the alias but yields to explicit override fields; resolution is ONE level deep, so a preset
|
|
272
|
+
meant to carry a model must name one.
|
|
319
273
|
|
|
320
274
|
## Resilience already handled — do not reimplement
|
|
321
275
|
|
|
322
|
-
Idle stream deadline · duplicate-final-chunk dedup · output-budget escalation · reasoning-cap
|
|
323
|
-
|
|
324
|
-
coercion · JSON salvage from prose · `NullCapture`
|
|
325
|
-
|
|
326
|
-
|
|
327
|
-
|
|
328
|
-
`tool_use` pairing). Details: package `README.md`.
|
|
276
|
+
Idle stream deadline · duplicate-final-chunk dedup · output-budget escalation · reasoning-cap shrink ·
|
|
277
|
+
adaptive-thinking budget floor · same-family fallback model · caller-seeded ladder position
|
|
278
|
+
(`escalation`) · schema coercion · JSON salvage from prose · `NullCapture` diagnostics · fatal-error
|
|
279
|
+
short-circuit · blank-content sanitization (whitespace-only text blocks are dropped before every call
|
|
280
|
+
— a blank block, e.g. an empty file read pasted into a prompt, is otherwise a fatal Anthropic 400;
|
|
281
|
+
blank tool results are stubbed to keep their `tool_use` pairing). Details: package `README.md`.
|
|
329
282
|
|
|
330
283
|
## Tests
|
|
331
284
|
|
|
332
|
-
`bun test ./tests` in the package
|
|
333
|
-
|
|
285
|
+
`bun test ./tests` in the package; offline specs always run. In `tests/model.spec.ts` the anthropic live
|
|
286
|
+
suite is gated on `ANTHROPIC_SECRET` and self-skips with a printed reason without it, and the OpenRouter
|
|
287
|
+
suite is disabled unconditionally: an aggregator on a separate account serving models no deployment runs,
|
|
288
|
+
whose `402 requires more credits` reads as a failure of the code under test. `plugins.spec.ts` covers the
|
|
289
|
+
`Compatible` provider offline.
|
|
334
290
|
|
|
335
291
|
## Depends On
|
|
336
292
|
|
|
337
293
|
- `@owlmeans/llm-common` · `@owlmeans/context` · `@owlmeans/error` · `@owlmeans/basic-ids` · `ajv`
|
|
294
|
+
- `@anthropic-ai/sdk` — runtime, for the `BadRequestError` the fatal-error rules are written around
|
|
338
295
|
- peer `@langchain/core`, `@langchain/openai`, `@langchain/anthropic`
|
|
339
296
|
|
|
340
297
|
## Related
|
|
341
298
|
|
|
342
|
-
- [[llm-common]] — the serializable contracts
|
|
343
|
-
|
|
299
|
+
- [[llm-common]] — the serializable contracts · [[llm-prompt-caching]] — prompt composition, block
|
|
300
|
+
order and the cache invariants
|
|
344
301
|
- [[context]] — service registration · [[error]] — the `ResilientError` family
|