@plurnk/plurnk-providers 1.3.5 → 1.3.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (84) hide show
  1. package/.env.defaults +35 -45
  2. package/README.md +44 -53
  3. package/SPEC.md +215 -371
  4. package/dist/AiSdkProvider.d.ts +78 -0
  5. package/dist/AiSdkProvider.d.ts.map +1 -0
  6. package/dist/AiSdkProvider.js +591 -0
  7. package/dist/AiSdkProvider.js.map +1 -0
  8. package/dist/OpenAICompat.d.ts +1 -2
  9. package/dist/OpenAICompat.d.ts.map +1 -1
  10. package/dist/OpenAICompat.js +39 -117
  11. package/dist/OpenAICompat.js.map +1 -1
  12. package/dist/ProviderRegistry.d.ts.map +1 -1
  13. package/dist/ProviderRegistry.js +37 -24
  14. package/dist/ProviderRegistry.js.map +1 -1
  15. package/dist/aiSdkTransport.d.ts +52 -0
  16. package/dist/aiSdkTransport.d.ts.map +1 -0
  17. package/dist/aiSdkTransport.js +294 -0
  18. package/dist/aiSdkTransport.js.map +1 -0
  19. package/dist/catalogProvider.d.ts +15 -0
  20. package/dist/catalogProvider.d.ts.map +1 -0
  21. package/dist/catalogProvider.js +103 -0
  22. package/dist/catalogProvider.js.map +1 -0
  23. package/dist/compatibleProvider.d.ts +3 -0
  24. package/dist/compatibleProvider.d.ts.map +1 -0
  25. package/dist/compatibleProvider.js +146 -0
  26. package/dist/compatibleProvider.js.map +1 -0
  27. package/dist/discover.d.ts.map +1 -1
  28. package/dist/discover.js.map +1 -1
  29. package/dist/env.d.ts +1 -0
  30. package/dist/env.d.ts.map +1 -1
  31. package/dist/env.js +13 -6
  32. package/dist/env.js.map +1 -1
  33. package/dist/index.d.ts +4 -6
  34. package/dist/index.d.ts.map +1 -1
  35. package/dist/index.js +3 -7
  36. package/dist/index.js.map +1 -1
  37. package/dist/ollama.d.ts +3 -0
  38. package/dist/ollama.d.ts.map +1 -0
  39. package/dist/ollama.js +39 -0
  40. package/dist/ollama.js.map +1 -0
  41. package/dist/openai.d.ts +2 -4
  42. package/dist/openai.d.ts.map +1 -1
  43. package/dist/openai.js +1 -2
  44. package/dist/openai.js.map +1 -1
  45. package/dist/sdkModels.d.ts +13 -0
  46. package/dist/sdkModels.d.ts.map +1 -0
  47. package/dist/sdkModels.js +153 -0
  48. package/dist/sdkModels.js.map +1 -0
  49. package/dist/standardProviders.d.ts.map +1 -1
  50. package/dist/standardProviders.js +0 -1
  51. package/dist/standardProviders.js.map +1 -1
  52. package/dist/telemetry.d.ts.map +1 -1
  53. package/dist/telemetry.js +20 -9
  54. package/dist/telemetry.js.map +1 -1
  55. package/dist/types.d.ts +3 -2
  56. package/dist/types.d.ts.map +1 -1
  57. package/package.json +18 -10
  58. package/src/{OpenAICompat.test.ts → AiSdkProvider.test.ts} +193 -152
  59. package/src/{OpenAICompat.ts → AiSdkProvider.ts} +83 -127
  60. package/src/Mock.test.ts +1 -1
  61. package/src/ProviderRegistry.test.ts +40 -27
  62. package/src/ProviderRegistry.ts +35 -24
  63. package/src/aiSdkTransport.test.ts +253 -0
  64. package/src/aiSdkTransport.ts +369 -0
  65. package/src/boundaries.test.ts +2 -2
  66. package/src/catalogProvider.test.ts +100 -0
  67. package/src/catalogProvider.ts +151 -0
  68. package/src/compatibleProvider.test.ts +44 -0
  69. package/src/compatibleProvider.ts +205 -0
  70. package/src/discover.test.ts +12 -12
  71. package/src/discover.ts +3 -6
  72. package/src/env.ts +14 -6
  73. package/src/index.ts +6 -10
  74. package/src/ollama.ts +63 -0
  75. package/src/openai.ts +2 -8
  76. package/src/sdkModels.test.ts +47 -0
  77. package/src/sdkModels.ts +194 -0
  78. package/src/telemetry.test.ts +17 -10
  79. package/src/telemetry.ts +22 -14
  80. package/src/types.ts +5 -8
  81. package/src/aiSdkAdapter.spike.test.ts +0 -242
  82. package/src/openaiStream.ts +0 -310
  83. package/src/standardProviders.test.ts +0 -939
  84. package/src/standardProviders.ts +0 -631
package/SPEC.md CHANGED
@@ -1,453 +1,297 @@
1
- # plurnk-providers — Specification
1
+ # Provider Contract
2
2
 
3
- Contract for `@plurnk/plurnk-providers-*` sibling packages. Audience: implementer of an LLM transport. Consumer: [plurnk-service](https://github.com/plurnk/plurnk-service) (SPEC.md §2).
3
+ `@plurnk/plurnk-providers` adapts model endpoints to one stable PLURNK
4
+ `Provider`. It does not maintain a parallel vendor registry or reproduce
5
+ ordinary provider protocols.
4
6
 
5
- ## §1 Manifest
7
+ ## §1 Ownership
6
8
 
7
- Each provider package's `package.json`:
9
+ The provider stack has four owners:
8
10
 
9
- ```json
10
- {
11
- "name": "@plurnk/plurnk-providers-<name>",
12
- "plurnk": { "kind": "provider", "name": "<name>" }
13
- }
14
- ```
11
+ 1. Models.dev supplies a build-time snapshot of provider package, API endpoint,
12
+ credential names, models, context windows, output limits, and USD prices.
13
+ 2. Official AI SDK providers own vendor request and response protocols.
14
+ 3. This package owns the PLURNK contract: aliases, envelopes, normalized usage
15
+ and errors, evidence, local capabilities, and first-party metadata.
16
+ 4. The operator owns secrets, machine-specific endpoints, and deliberate
17
+ metadata overrides through environment variables.
15
18
 
16
- - `kind` MUST be `"provider"`.
17
- - `name` is a vendor identifier (`openai`, `anthropic`, `ollama`).
18
-
19
- Collision on `(kind: "provider", name)` at discovery: fail-hard.
19
+ Facts MUST have one owner. Do not copy a cataloged endpoint, credential name,
20
+ model prefix, context window, price, or vendor request shape into a PLURNK
21
+ table. A missing or wrong catalog fact is fixed upstream, overridden through a
22
+ provider declaration, or left explicitly unknown.
20
23
 
21
24
  ## §2 Provider interface
22
25
 
26
+ `Provider` exposes immutable model facts and one generation operation:
27
+
23
28
  ```ts
24
29
  interface Provider {
25
- // Identity (immutable across lifetime)
26
- readonly contextWindow: number | null; // context tokens, null if unresolved.
27
- // PER SLOT under llama-server --parallel N
28
- // (the server splits --ctx-size and reports
29
- // the divided value; verified live).
30
- readonly model: string; // configured model id (the alias for a local backend)
31
- // OPTIONAL (#37): backend's SELF-REPORTED served id, from a /v1/models-shaped
32
- // probe. For a local alias `model` is the alias; this is the real served name
33
- // (the .gguf) the tokenizer seam maps. Absent when unprobed. Consumers resolve
34
- // `servedModel ?? model`.
35
- readonly servedModel?: string;
36
-
37
- // Tokenomic primitives (synchronous, pure)
38
- countTokens(text: string): number;
39
- calculateCost(usage: ProviderUsage): number; // estimated USD
40
-
41
- // OPTIONAL capability: exact tokenization served by the backend's own vocab
42
- // (llama-server /tokenize). Probe-gated — undefined means the backend can't.
43
- tokenize?(text: string): Promise<number[]>;
44
-
45
- // OPTIONAL resolved capabilities — introspectable facts for boot-time policy:
46
- // constrainsOutput (#34): a transported grammar WILL constrain this decode
47
- // (rails live) — consumers fail hard on a dark-rails boot instead of
48
- // discovering it from unconstrained emissions.
49
- // requiresMaxTokens (#43): this backend decodes UNBOUNDED absent a caller cap
50
- // (llama-server honors n_predict to the context wall, providers#10) — a
51
- // consumer MUST bring an output envelope (§13), and can refuse AT BOOT a
52
- // local alias whose envelope was never declared. Self-clamping cloud
53
- // backends never set it; undefined = no claim.
54
- readonly constrainsOutput?: boolean;
55
- readonly requiresMaxTokens?: boolean;
56
- // #507 (owner-ruled): generation-envelope reserves derived from the DETECTED
57
- // window (floor percentages; absolute per-alias pins win outright). null =
58
- // underivable -> the consumer's no-cap path. The consumer's prompt budget is
59
- // window - reasoningReserve - completionReserve - its OWN safety margin;
60
- // its generation cap is the two reserves pooled. Absent = no claim.
61
- readonly reasoningReserve?: number | null;
62
- readonly completionReserve?: number | null;
63
-
64
- // Transport. `workerId` is REQUIRED: the opaque, stable identity of the
65
- // consumer's work stream — providers may key backend affinity on it and
66
- // never interpret it. `grammar` is an optional GBNF string for
67
- // grammar-constrained sampling (§13) — attached verbatim by capable
68
- // backends, ignored by all others. `maxTokens` is the consumer's per-call
69
- // output ceiling (wire `max_tokens`); absent means the server default,
70
- // which is typically UNBOUNDED. `attributions`/`client` are optional
71
- // first-party metadata, forwarded as `Plurnk-*` headers ONLY by a provider
72
- // configured with `firstPartyMetadata` (the plurnk endpoint) and dropped by
73
- // every other — structurally unable to reach a third-party backend (§11).
74
- // `sampling` is an optional bag of standard OpenAI-compat sampling params
75
- // (temperature, top_p, top_k, min_p, penalties, stop, seed, …) merged into the
76
- // body UNDER the managed fields — model/messages/grammar/reasoning/max_tokens/
77
- // slot always win, and reserved keys are stripped (#477): transport/protocol
78
- // (stream, response_format, grammar, id_slot, logprobs), paradigm breakers
79
- // (n, the tools/functions family, modalities/audio, prediction), and the
80
- // token caps (max_tokens/max_completion_tokens -- the envelope is the managed
81
- // maxTokens, never bypassable). It carries sampling intent + platform knobs
82
- // only (§8). For a PROXY consumer forwarding its own
83
- // caller's sampling knobs (the plurnk endpoint fronting gemma/Fireworks); a
84
- // direct consumer leaves it unset.
85
- // `strikes` is the worker's CURRENT rail-strike streak at time-of-generate
86
- // (0 = clean, distinct from absent = unreported; contract plurnk-service#313).
87
- // Forwarded as `Plurnk-Strikes` ONLY under the firstPartyMetadata gate,
88
- // dropped everywhere else. Headers only — never placed in the packet.
89
- // `workspaceId`/`loop`/`turn` (#404): the turn coordinate, stamped as
90
- // `Plurnk-Workspace-Id`/`Plurnk-Loop`/`Plurnk-Turn` under the SAME gate.
91
- // 1-based; absent/0 emits no header. Headers only, never the packet.
92
- // `primaryWorkerId` (#522): the ROOT worker of this turn's lineage (the
93
- // no-parent ancestor the worker tree descends from) — a worker id, stamped as
94
- // `Plurnk-Worker-Primary` under the SAME gate. Lets a consumer classify
95
- // root-vs-descendant by equality (`primaryWorkerId == workerId` ⇒ the primary
96
- // worker) and group a worker tree by its root, for telemetry/analytics. The
97
- // provider EMITS what the consumer supplies and never invents a primary; the
98
- // consumer's contract is to supply it every turn (including the primary's own,
99
- // where it equals workerId). Absent emits no header.
100
- generate(args: { messages: ChatMessage[]; workerId: string; primaryWorkerId?: string; signal?: AbortSignal; grammar?: string; maxTokens?: number; attributions?: string[]; client?: string; strikes?: number; workspaceId?: string; loop?: number; turn?: number; sampling?: Record<string, unknown> }): Promise<ProviderResponse>;
101
- }
102
-
103
- interface ProviderResponse {
104
- assistant: {
105
- content: string; // raw model emission; consumer parses
106
- reasoning: string | null; // wire-reported reasoning content; null if absent
107
- // sealed relay reasoning (#482): items { id, subtype, encrypted:
108
- // [{data, format}] } verbatim, never decoded; `id` from the wire,
109
- // `subtype` from wire position (message-attached; plurnk is tools-in-body
110
- // so it is constant). Absent when none. agui projects REASONING_ENCRYPTED_VALUE.
111
- reasoningEncrypted?: Array<{ id: string | null; subtype: string; encrypted: Array<{ data: string; format: string | null }> }>;
112
- usage: ProviderUsage; // { prompt, completion, reasoning, cached, total }
113
- finishReason: "stop" | "length" | "tool_calls" | "content_filter" | null;
114
- model: string; // wire-reported (may differ from requested for relay providers)
115
- };
116
- assistantRaw: unknown; // verbatim wire response for forensics
117
- meta?: Record<string, unknown>; // verbatim per-turn provider metadata; absent when empty (#23)
118
- }
119
-
120
- interface ProviderUsage {
121
- prompt: number; // input tokens (cached ones included)
122
- completion: number; // visible output tokens, EXCLUDING reasoning
123
- reasoning: number; // reasoning tokens, billed as output
124
- cached: number; // subset of prompt served from cache
125
- total: number; // prompt + completion + reasoning
30
+ readonly model: string;
31
+ readonly contextWindow: number | null;
32
+ readonly servedModel?: string;
33
+ readonly constrainsOutput?: boolean;
34
+ readonly requiresMaxTokens?: boolean;
35
+ readonly reasoningReserve?: number | null;
36
+ readonly completionReserve?: number | null;
37
+
38
+ countTokens(text: string): number;
39
+ tokenize?(text: string): Promise<number[]>;
40
+ calculateCost(usage: ProviderUsage): number;
41
+ generate(args: GenerateArgs): Promise<ProviderResponse>;
126
42
  }
127
43
  ```
128
44
 
129
- Usage invariant: `total = prompt + completion + reasoning`; `cached ⊆ prompt`; `completion` excludes reasoning; **billable output = `completion + reasoning`**. Providers report reasoning THREE ways: inside `completion_tokens` (OpenAI, via `completion_tokens_details.reasoning_tokens`), only as the `total - prompt - completion` gap (Gemini), or folded into `completion_tokens` with NO itemization while shipping the reasoning as TEXT (Fireworks, #425). The framework's `normalizeUsage` (§11) collapses all three to this invariant -- re-splitting the Fireworks case by the emitted text lengths, sum-preserving so cost is unchanged -- so siblings on `OpenAICompatProvider` get it for free.
45
+ `contextWindow: null` means genuinely unknown. A consumer MUST NOT invent a
46
+ stand-in. A cataloged cloud model without a context window fails construction
47
+ unless the operator pins `PLURNK_PROVIDERS_CONTEXT_WINDOW`. A local probe
48
+ failure degrades to `null` and emits one warning because a transient probe
49
+ failure must not make a usable local endpoint unbootable.
130
50
 
131
- ### OpenAI-compatible request execution
51
+ `countTokens` is synchronous and non-negative. The common fallback is a
52
+ conservative chars/2 ruler and is announced once. `tokenize` exists only when
53
+ the endpoint exposes its real vocabulary.
132
54
 
133
- `OpenAICompatConfig.fetch?: ProviderFetch`, where `ProviderFetch` is
134
- `typeof globalThis.fetch`, selects request execution for one provider instance.
135
- When omitted, the provider uses `globalThis.fetch`. The same function MUST
136
- execute streaming completions, buffered completions, retries, and optional
137
- backend tokenization. The provider supplies its complete URL, headers, body,
138
- and effective `AbortSignal`; an injected implementation MUST preserve those Web
139
- API semantics.
55
+ `calculateCost` returns estimated USD. Models.dev rates are converted at the
56
+ provider boundary. Unknown pricing returns `0`; it is not represented as a
57
+ fabricated rate.
140
58
 
141
- Injection changes only how a shaped request executes. Request construction,
142
- response and error interpretation, retry policy, cancellation and timeout
143
- signals, usage and reasoning normalization, telemetry, and raw capture remain
144
- owned by `OpenAICompatProvider`. The function is per-instance; replacing
145
- `globalThis.fetch` is not part of the contract. A platform or vendor binding
146
- adapter belongs to its consuming integration and returns a fetch-compatible
147
- `Response`.
59
+ ### Generation
148
60
 
149
- The `@plurnk/plurnk-providers/openai` subpath is the runtime-neutral import
150
- surface for this contract. Its transitive module graph MUST NOT import provider
151
- discovery, registries, filesystem access, or environment-owned construction.
152
- The package root remains the Node daemon integration surface.
61
+ `generate` requires a non-empty, stable, opaque `workerId`. It accepts:
153
62
 
154
- ### Promises
63
+ - `messages`: system, user, and assistant text messages;
64
+ - caller cancellation through `signal`;
65
+ - optional `grammar` and `maxTokens`;
66
+ - standard `sampling` intent;
67
+ - first-party attribution, client, strike, workspace, loop, and turn metadata.
155
68
 
156
- - `assistant.content` is the **verbatim** model emission. Consumer parses via `@plurnk/plurnk-grammar` — providers MUST NOT parse. plurnk uses a **tools-in-body** design: tool invocations are expressed as plurnk DSL *inside the message content*, never via a provider's native tool-calling API, so `content` is always raw text and providers never request, parse, or translate native tool calls.[^tools]
69
+ It returns the model's raw content and reasoning, normalized usage, normalized
70
+ finish reason, model identity, opaque evidence, optional metadata, and optional
71
+ telemetry. The provider transports and observes model output; it never retries,
72
+ discards, or repairs an otherwise completed exchange because PLURNK grammar did
73
+ not accept it.
157
74
 
158
- [^tools]: A provider that wants to drive native tool-calling (OpenAI `tool_calls`, Anthropic `tool_use`) would have to normalize those emissions back into a plurnk DSL string at the boundary. No provider does this and it is out of scope for v0 — tools-in-body sidesteps the whole problem. The clause is recorded only so a future native-tools mode has a defined contract.
159
- - `assistant.usage` is authoritative and follows the invariant above. Fill `0`s when the wire response omits a breakdown.
160
- - `countTokens` is **synchronous**, returns a non-negative integer, deterministic for the same input. Without an exact tokenizer family configured it is the **chars/2 UPPER BOUND** — deliberately conservative (real agentic text measures ~2.9–3.2 chars/token on gemma/deepseek, so the former chars/4 silently UNDERcounted 20–27%; a fallback may overcount, never under) — and it is **surfaced at construction** (`process.emitWarning`, code `PLURNK_TOKENIZER_HEURISTIC`), never silent. Exact counting is the tokenizer seam's job (mimetypes family), fed by `tokenize()` where available.
161
- - `tokenize?` is an **optional async capability**: token ids in the model's real vocabulary, served by the backend itself (llama-server's native root `/tokenize`, surfaced when the §11 probe fingerprints a llama-server and `detectLlamaServer` isn't false). `tokenize === undefined` is the honest "backend can't" signal. Exact-counting consumers prefer it over any client-side tokenizer data — the local model's own vocab needs no bundled `tokenizer.json` at all.
162
- - `calculateCost` is **pure**, returns USD non-negative integer. Returns `0` for siblings with no known rates (local Ollama, generic OpenAI-compat shims).
163
- - `contextWindow` resolves to `null` when a PROBING provider (openai/llama-server) can't determine the window (consumer treats null as "no budget info"); a CLOUD provider with no window source FAILS HARD instead (#419, §11).
164
- - `generate` rejects on signal abort — does NOT resolve with partial content.
165
- - `generate` transports `grammar` verbatim when the backend supports grammar-constrained sampling, and silently ignores it otherwise (§13). The provider never chooses or modifies the grammar.
166
- - `generate` **returns for every completed exchange — bytes always present, the conformance verdict attached as an observation.** When a grammar was transported (or validated in filter mode), the returned `content` is checked against it; a non-accept verdict rides `response.telemetry` as a `grammar_unenforced` event (message + divergence `position`) and the response returns normally. The provider transports and observes; it never adjudicates — discard, retry, escalate, or feed-back is consumer policy. This is a grammar-**conformance** check against the grammar the provider already holds — *not* a plurnk-DSL parse (that stays consumer-side, below) — so it remains backend- and DSL-agnostic. `ProviderError` remains reserved for exchanges that did NOT complete (transport failure, abort, boundary violations).
167
- - **Backend affinity is the provider's internal guarantee, keyed by `workerId`.** The consumer says *which run this is*, never *which backend resource serves it* — raw resource identifiers (slot integers, connections) never cross the contract in either direction. On slot-pinning backends (llama-server `--parallel N>1`), the provider keeps each worker sticky to one slot and spreads distinct runs across slots, so each concurrent run keeps its KV-cache prefix warm (un-pinned routing is the server's similarity heuristic — slot hops re-pay full prefills). Backends without affinity semantics ignore `workerId` entirely.
75
+ Usage obeys:
168
76
 
169
- ## §3 `fromEnv(env, model, options?)` factory
170
-
171
- Default export MUST have a static `fromEnv(env, model, options?)` factory:
172
-
173
- ```ts
174
- class OpenAI {
175
- static fromEnv(env: NodeJS.ProcessEnv, model: string, options?: ProviderOptions): OpenAI | Promise<OpenAI> {
176
- // Read provider-specific env (OPENAI_BASE_URL, OPENAI_API_KEY, ...)
177
- // plus universal operator knobs (PLURNK_PROVIDERS_REASONING, PLURNK_PROVIDERS_FETCH_TIMEOUT,
178
- // PLURNK_PROVIDERS_CONTEXT_WINDOW). `options.baseUrl`, when set, is the per-alias
179
- // endpoint override (PLURNK_BASEURL_<alias>, §5) and wins over the env base URL.
180
- return new OpenAI({ /* ... */ });
181
- }
182
- constructor(config: OpenAIConfig) { /* ... */ }
183
- }
77
+ ```text
78
+ total = prompt + completion + reasoning
79
+ cached ⊆ prompt
184
80
  ```
185
81
 
186
- The consumer's instantiation path calls `mod.default.fromEnv(env, alias.model, options?)` generically (§5). `options` is optional — a factory that ignores the third arg keeps working unchanged; a self-hosted provider (`openai`, `ollama`) honors `options.baseUrl` so two aliases can target two boxes.
82
+ `completion` excludes reasoning. Known vendor finish reasons normalize to
83
+ `stop`, `length`, `tool_calls`, or `content_filter`; an unknown value becomes
84
+ `null` and emits a warning.
187
85
 
188
- `fromEnv` MAY be sync or async; return type `Provider | Promise<Provider>`. (Why a factory, where execs/mimes use a base class + constructor injection and schemes a `static manifest`: a provider often **async-probes** at construction — model catalog, context window, slot count — which a constructor can't express. The factory is the seam for that probe.)
86
+ ## §3 AI SDK boundary
189
87
 
190
- `fromEnv` MUST fail fast with a clear error if required env is missing — name the env var the operator needs to set.
88
+ Cataloged providers instantiate their Models.dev-declared AI SDK package.
89
+ Standard request shaping, streaming, retries, cancellation, timeouts, usage,
90
+ and vendor error parsing belong to the SDK.
191
91
 
192
- ## §4 Universal operator knobs
92
+ PLURNK maps its generic settings to AI SDK call settings:
193
93
 
194
- Defaults use four evidence states:
94
+ - `temperature`, `top_p`, `top_k`;
95
+ - presence and frequency penalties;
96
+ - stop sequences and seed;
97
+ - output-token ceiling;
98
+ - `off`, `adaptive`, or budget-derived reasoning intent.
195
99
 
196
- - **Measured** — a controlled behavioral experiment demonstrates the parameter's
197
- effect on the relevant endpoint and model family.
198
- - **Documented** — the provider's primary documentation defines the parameter and
199
- its semantics.
200
- - **Accepted** — a live request carrying the parameter succeeds. This proves wire
201
- compatibility only; an endpoint may accept and ignore a field.
202
- - **Unknown** — neither semantics nor compatibility are established.
100
+ Provider-specific options are permitted only where they preserve a documented
101
+ PLURNK product contract the generic SDK surface cannot express.
203
102
 
204
- Measured evidence outranks documented evidence when they conflict. A portable
205
- nonzero tuning default requires measured or documented semantics that actually
206
- generalize across its scope. Accepted evidence permits transport but never
207
- justifies a magnitude. PLURNK does not maintain a provider-characterization
208
- harness; upstream provider contracts and focused regression specimens own that
209
- knowledge.
103
+ The compatible transport is deliberately retained for:
210
104
 
211
- Each provider's `fromEnv` reads these:
105
+ - `openai` local endpoints, including llama-server and vLLM;
106
+ - `ollama`, after its native `/api/show` probe;
107
+ - the first-party `plurnk` endpoint;
108
+ - operator-declared `@ai-sdk/openai-compatible` providers.
212
109
 
213
- - **`PLURNK_PROVIDERS_REASONING`** — REQUIRED, one of `off | adaptive | on`. **`PLURNK_PROVIDERS_REASONING_BUDGET`** is a positive integer required when reasoning is `on`. The provider maps this intent to the backend's supported controls. Backends may omit, combine, or expose reasoning differently when grammar-constrained output is enabled; the provider reports the channels it actually receives rather than synthesizing a separate reasoning channel. `max_tokens` remains an output limit, not a reasoning budget. The model-facing `PLAN` operation is part of the grammar and is independent of provider reasoning controls.
110
+ It carries PLURNK-only fields and raw wire evidence without reimplementing the
111
+ SDK's ordinary transport.
214
112
 
215
- Read via `reasoningFromEnv` and **fail hard when unset**: the budget is required only when `on` and carries no floor default. Configuration lives in the operator's env over the package's `.env.defaults` floor (which declares every var and ships its default); the framework never bakes a knob default into code.
216
- - **`PLURNK_PROVIDERS_FETCH_TIMEOUT`** — service-wide ms ceiling on any single outbound request (**per attempt**, not shared across retries). Each `fromEnv` reads and passes as `AbortSignal.timeout`. Per-provider override envs are NOT part of the contract.
217
- - **`PLURNK_PROVIDERS_STREAM_IDLE_TIMEOUT`** — REQUIRED non-negative milliseconds from the shipped floor; the maximum silence between streamed response-body chunks after the response body becomes available. The clock resets on every byte chunk, `0` disables it, and expiry is a retryable `network_failure`. The portable default is `0`: slow local inference may legitimately pause for minutes, so a nonzero deadline is a measured per-alias deployment policy, not a universal provider assumption. It does not shorten time-to-response for a large prefill; `FETCH_TIMEOUT` remains the whole-attempt ceiling. A successful retry records its attempt, elapsed time, reason, and message in `ProviderResponse.meta.transportRetries`.
218
- - **`PLURNK_PROVIDERS_RETRY_ATTEMPTS`** — REQUIRED non-negative integer (read via `parseRequiredInt`). The transient-failure retry budget: **`0`** surfaces the first failure; **`N`** retries up to `N` times on a *transient* classification only (`rate_limit` / `network_failure` — 429, 5xx, timeout, connection reset), plus **`grammar_invalid`** (a 422 output reject a fresh sample may satisfy, #548). Terminal kinds (`unauthorized`, `quota_exceeded`, `invalid_response`, `model_refused`) are never retried. Backoff is exponential from a `2000ms` base (`base * 2^(attempt-1)`), unless the server sent a `Retry-After` (which wins). The caller's `signal` aborts both the in-flight request and the backoff sleep. Lives in the shared `OpenAICompatProvider` so every provider inherits it uniformly; rides on the existing `classifyProviderError` (#18).
219
- - **`PLURNK_PROVIDERS_CONTEXT_WINDOW`** -- optional positive-integer override (alias-scopable) for the model's context window. Resolution (#419): this env var -> endpoint `n_ctx` probe (probing specs only) -> `@plurnk/plurnk-models` catalog -> then a PROBING provider degrades to `null`, a CLOUD provider (no probe) FAILS HARD (an uncataloged, unpinned cloud model is a config error, not a guessable window). See §11.
220
- - **`PLURNK_PROVIDERS_REASONING_RESERVE` / `PLURNK_PROVIDERS_COMPLETION_RESERVE`** (#507, owner-ruled) -- the generation-envelope reserves, REQUIRED (floor ships `10%` / `25%`). A percentage derives from the DETECTED window (llama-server n_ctx, the plurnk.ai router, the catalog) so every advertising endpoint gets sane defaults with ZERO operator tuning; an absolute token count wins outright (per-alias-scopable -- the measured-envelope override). These MIGRATED from core's `PLURNK_SERVICE_{CONTEXT_WINDOW,REASONING,COMPLETION}` (provider quantities wearing a service prefix; core keeps only its own packing-safety margin). Surfaced as `Provider.reasoningReserve`/`completionReserve`.
221
- - **`PLURNK_PROVIDERS_TEMPERATURE` / `PLURNK_PROVIDERS_REPEAT_PENALTY` / `PLURNK_PROVIDERS_FREQUENCY_PENALTY`** -- REQUIRED sampling controls (read via `parseRequiredFloat`, values from the `.env.defaults` floor). `REPEAT_PENALTY` (canonical `1.15`) is the measured llama.cpp multiplier. `FREQUENCY_PENALTY` is the optional cloud analogue, but its portable floor is `0`: endpoint acceptance proves only that the field may ride, not that one magnitude has a portable semantic effect. Enable it per alias from provider documentation or controlled behavioral evidence. The controls are keyed per backend (§13).
222
- - **`PLURNK_PROVIDERS_SERVICE_TIER`** -- optional fixed request tier, alias-scopable. Fireworks accepts its published `auto | default | flex | priority` vocabulary and owns those values' routing semantics; when configured, the value is validated at construction and wins over per-call sampling on every request. Unset delegates to the provider default. Other standard providers reject the knob rather than silently ignoring a paid routing choice.
223
- - **`PLURNK_PROVIDERS_DRY_MULTIPLIER` / `_DRY_BASE` / `_DRY_ALLOWED_LENGTH`** -- the llama.cpp DRY loop-breaker (#567), customer-overridable per alias. The generic floor is deliberately **off** (`MULTIPLIER=0`): the turboderp/Gemma sweep proved the community-standard `0.8`/`1.75`/`2` settings worse than off for both runaway emissions and exact-identifier corruption. That same sweep measured `0.8`/`1.75`/`32` as a safe alias-specific deployment setting (zero corruption, about 6% runaways versus 19% off), so `.env.defaults` carries `BASE=1.75` and `ALLOWED_LENGTH=32` as inert override companions without pretending one model's multiplier is universal. DRY penalizes repeated sequences with a penalty escalating in run length -- the tool for a plan-restart loop a single-token `repeat_penalty` over a short window cannot see. **`PLURNK_PROVIDERS_REPEAT_LAST_N`** (optional) widens the older `repeat_penalty` window past the box's 64. These knobs are sent **only on the detected `llamacpp` path**; cloud providers parse but never emit them.
113
+ ## §4 Operator configuration
224
114
 
225
- ## §5 Alias cascade resolution
115
+ Every operational value is an environment knob documented in `.env.defaults`.
116
+ There are no hidden tuning constants. Every `PLURNK_PROVIDERS_*` knob may be
117
+ scoped to an alias by appending `_<alias>`; the scoped value wins.
226
118
 
227
- `PLURNK_MODEL_<alias>=<provider>/<model-id>` declares an alias; `PLURNK_MODEL=<alias>` selects which is active.
119
+ The universal groups are:
228
120
 
229
- ```
230
- PLURNK_MODEL_gemma=openai/macher.gguf
231
- PLURNK_MODEL_opus=openrouter/anthropic/claude-opus-latest
232
- PLURNK_MODEL=gemma
233
- ```
121
+ - reasoning activation and optional explicit budget;
122
+ - decode tuning;
123
+ - request, stream-idle, retry, and probe budgets;
124
+ - local GBNF and llama-server capability pins;
125
+ - context-window and generation-envelope overrides;
126
+ - opt-in logprob and raw-body capture.
234
127
 
235
- First path segment names the provider; rest is the model identifier (may contain `/` for tri-level providers like openrouter's `publisher/model`).
128
+ Operator secrets and machine-specific values never belong in committed
129
+ defaults.
236
130
 
237
- **Per-alias endpoint override — `PLURNK_BASEURL_<alias>`.** A provider's base URL otherwise binds **one URL per provider *name*** (its `baseUrlVar`, §11), so two `openai/…` aliases collapse onto the same `OPENAI_BASE_URL`. That makes running **N self-hosted boxes of the same kind** — the real case for the two "bring your own box" providers, `openai` (llama.cpp/vLLM/LM Studio) and `ollama` — impossible by name alone. `PLURNK_BASEURL_<alias>` attaches an endpoint to the *alias*, case-folded to match its `PLURNK_MODEL_<alias>`, and **wins over** the provider's own base-URL var:
131
+ ## §5 Resolution
238
132
 
239
- ```
240
- PLURNK_MODEL_HAZEL1=openai/qwen2.5-coder
241
- PLURNK_BASEURL_HAZEL1=http://hazel1:8080/v1 # llama.cpp box 1
242
- PLURNK_MODEL_HAZEL2=openai/qwen3-coder
243
- PLURNK_BASEURL_HAZEL2=http://hazel2:8080/v1 # llama.cpp box 2 — same provider name, different box
244
- PLURNK_MODEL_NOOK=ollama/qwen2.5-coder
245
- PLURNK_BASEURL_NOOK=http://nook:11434 # an ollama box; drives BOTH /v1/chat and /api/show
246
- ```
133
+ `PLURNK_MODEL_<alias>=<provider>/<model-id>` declares an alias.
134
+ `PLURNK_MODEL=<alias>` selects the boot alias. Model IDs may contain `/`.
135
+ `PLURNK_BASEURL_<alias>` is a per-alias endpoint override.
247
136
 
248
- Each alias instantiates against its own URL and probes its own box (openai's `n_ctx`/slots, ollama's `/api/show`). The override threads through the alias → `instantiateProvider` → both tiers; on the tier-2 (plugin) path it arrives as the third `fromEnv(env, model, { baseUrl })` argument (a factory that ignores it is unaffected). A `PLURNK_BASEURL_*` with no matching alias **fails hard** (a typo, not a silent no-op). It's accepted for *any* provider but only meaningful for the self-hosted ones; the hosted providers (`groq`, `fireworks`, …) carry a canonical endpoint you'd never multi-home.
137
+ `instantiateProvider` resolves in this order:
249
138
 
250
- This package's exported resolution surface:
139
+ 1. A Models.dev provider and model, using its declared AI SDK package.
140
+ 2. An operator provider declaration:
141
+ `PLURNK_PROVIDERS_PROVIDER_<NAME>_{NPM,BASE_URL,API_KEY_ENV}`.
142
+ 3. The local `openai`, `ollama`, or first-party `plurnk` adapter.
143
+ 4. A discovered AI SDK provider plugin.
144
+ 5. A precise unknown-provider error.
251
145
 
252
- - `parseAliasesFromEnv(env)` — extracts alias entries.
253
- - `resolveActiveAlias(env)` — `{ alias, provider, model } | null`.
254
- - `instantiateProvider(name, env, model)` — full two-tier resolution (below).
255
- - `loadActiveProvider(env)` — boot convenience: active alias → `instantiateProvider`.
256
- - `standardProviderFromEnv(name, env, model)` / `isStandardProvider(name)` — the tier-1 internals, still exported.
257
- - `discover({ cwd?, packageDirs?, env? })` — the tier-2 scan: a `Discovery` (`{ registry, skipped, attributions }`) of every installed provider package. `registry`/`skipped` are `Map<name, packageSpecifier>` partitioned by the trust gate; `attributions` is `Map<name, string | string[]>` carrying each registered provider's raw `plurnk.attribution` for author credit (#21 — surfaced verbatim; the consumer applies the `@plurnk/`-scope reservation policy). Exported so a consumer can enumerate trusted providers — and see which were declined — without instantiating them.
146
+ Earlier sources are authoritative. Installed plugins cannot shadow a cataloged
147
+ or explicitly declared name.
258
148
 
259
- **Two-tier provider resolution — owned entirely by this package.** A provider name resolves in this order:
149
+ Model IDs resolve exactly first. A unique catalog suffix is accepted to avoid
150
+ forcing a vendor-owned resource prefix into PLURNK aliases. Ambiguous suffixes
151
+ fail to resolve.
260
152
 
261
- 1. **Standard provider** (`§11`) — if `isStandardProvider(name)`, instantiated directly via `standardProviderFromEnv(name, env, model)`. No package is imported. Covers every plain OpenAI-compatible endpoint (`openai`, `groq`, `deepseek`, `mistral`, `together`, `fireworks`, `deepinfra`, `anthropic`, `bedrock`, `plurnk`, …) — including first-party Claude (Anthropic's compat endpoint, bearer auth, the `thinking` reasoning param), AWS Bedrock (its compat endpoint at the `/openai/v1` path, Bedrock-API-key bearer; the base is `BEDROCK_BASE_URL` if set, else derived as `https://bedrock-runtime.{region}.amazonaws.com/openai/v1` from the standard `AWS_REGION`/`AWS_DEFAULT_REGION`; its inference-profile model ids resolve a catalog **context window** by stripping the region and looking the model up under its publisher (#22), while cost stays unknown — bedrock marks up over the native rate), and the **`plurnk` hosted model** (a plain remote OpenAI endpoint at its `PLURNK_BASE_URL` base, `.env.defaults` → `https://plurnk.ai/v1`; reads its server-controlled context window from upstream, sends no grammar and no tuning -- `suppressTuningFloors` drops the client temperature/penalty floors so the router's per-model tuning is never overridden (#507); caller `sampling` still passes); none needs a plugin. `plurnk` authenticates with a single optional bearer (`PLURNK_API_KEY`, sent only when set), like any other standard bearer. A spec's `apiKeyVar`/`baseUrlVar` may be a **list** of accepted env-var aliases — the conventional names the wild uses for one credential/base (e.g. `deepinfra` → `DEEPINFRA_API_KEY` / `DEEPINFRA_API_TOKEN` / `DEEPINFRA_TOKEN`; `openai` base → `OPENAI_BASE_URL` / `OPENAI_API_BASE`). First set non-empty wins; a required key unset across *all* aliases fails hard naming each. This is alias resolution over operator-set values, never a fabricated default. (An entry MAY instead supply a custom `headersFromEnv` builder for auth a bearer can't express — multi-header/credential schemes — or a `baseUrlFromEnv` builder to template the base from env, as bedrock does.) The standard table is **authoritative**: a scanned package whose name duplicates a standard one is shadowed (tier 1 returns first).
262
- 2. **Discovered package** — otherwise, a **scope-agnostic `node_modules` scan** (`discover()`) maps the provider name to the installed package that declares `plurnk: { kind: "provider", name }`, which is dynamic-imported and `fromEnv(env, model, options?)`-called. Covers first-party plugins with real runtime surface (`openrouter`, `ollama`, `google`, `xai`, `cloudflare`, the planned `vertex`/`cohere`) **and any third-party provider published under its own scope** (`@acme/llm-provider-foo`) — no involvement from us, no `@plurnk` scope assumption. Two installed packages claiming the same name → fail-hard naming both. Unknown name → fail-hard. The scan runs once per process and is memoized.
153
+ Provider declarations configure facts, not credentials:
263
154
 
264
- Discovery honors the host **trust gate** `PLURNK_PLUGINS_TRUSTED_ONLY` — the same env var the execs/mimes/schemes families read (plurnk-service#229, #15). OFF (unset/empty/`0`) trusts every installed provider; ON (any value) trusts `@plurnk/*` plus a comma-separated allowlist and declines the rest. A declined package is recorded in `skipped` (never registered, never thrown), so requesting its name yields a precise *untrusted* error rather than *unknown*.
265
-
266
- The framework is **contract-only**: it does not depend on provider plugins. The daemon declares its bundled providers as ordinary direct dependencies, and operators install additional providers at the application root so Node's package resolution and the scope-agnostic scan can find them. Provider plugins declare the framework as a peer dependency using the repository's normal same-minor compatibility range; this preserves one shared contract instance without coupling a plugin release to every framework patch.
267
-
268
- ## §6 Engine → provider guarantees (consumer side)
269
-
270
- - `messages` is a complete prompt. Consumer has pre-assembled all sections. Provider does not add, reorder, or inject turns — the wire `messages` are exactly what the consumer passed. (The provider injects no `PLAN` turn — `PLAN` is part of the consumer's grammar contract, §4.)
271
- - Every `generate` carries `workerId` — the worker's stable, opaque identity. Same run → same string across its turns; distinct runs → distinct strings.
272
- - `signal` is wired to the worker's AbortController.
273
- - `generate` is single-call per turn. No parallel calls on the same instance.
274
- - `assistantRaw` is opaque to the consumer (forensics-only).
275
- - `meta` is the per-turn provider→client metadata bag: the backend's non-standard top-level response fields pass through verbatim. Monetary metadata carries an explicit decimal-string `amount` and `currency`; the provider never guesses or converts its unit. Absent when the backend reported no extras. The consumer (service) merges `meta` into its Turn metadata and filters what reaches the client; it reads `meta`, never mines `assistantRaw` (#23).
276
- - `countTokens` is cheap by contract; consumer calls frequently.
277
-
278
- ## §7 Provider → engine guarantees
279
-
280
- - **No DB access.** Provider never touches `node:sqlite` or storage layers.
281
- - **No service access.** No imports from `@plurnk/plurnk-service`.
282
- - **No grammar runtime dep.** Type-imports from `@plurnk/plurnk-grammar` are fine; invoking `PlurnkParser.parse` is consumer-side.
283
- - **Raw `content`.** Returned verbatim. Tools are expressed in-body as plurnk DSL (see §2); providers do not use native tool-calling.
284
- - **Atomic.** One `generate` call resolves with one complete `ProviderResponse`. No streaming partial resolves (v0).
285
- - **Honors `signal`.** Aborted calls reject; resources free; no orphaned connections.
286
- - **Single model.** One provider instance speaks to one model.
287
- - **Synchronous `countTokens`, pure `calculateCost`.** No I/O, no async, no state beyond cached tokenizer artifacts.
288
-
289
- ## §8 Forbidden
290
-
291
- | ❌ |
292
- |---|
293
- | Database access |
294
- | Filesystem access beyond reading provider-internal config |
295
- | Imports from `@plurnk/plurnk-service/*` |
296
- | Resolving with partial content on abort |
297
- | Mutating `messages` |
298
- | Parsing `content` into `PlurnkStatement[]` |
299
- | Streaming the resolve (atomic only; v0) |
300
- | Holding state across `generate` calls beyond connection pooling, config, and backend-affinity bookkeeping (run→resource maps, §2) |
301
- | Exposing backend resource identifiers (slot ids, connections) on the consumer surface |
302
- | Reading model output via `console.*` |
303
- | Ignoring `signal` |
304
- | Spawning subprocesses for inference |
305
-
306
- ## §9 Reference — `Mock`
307
-
308
- `./src/Mock.ts` — test-fixture provider + worked example. Queue of pre-built responses; `generate` shifts one off.
309
-
310
- ```ts
311
- import { Mock } from "@plurnk/plurnk-providers";
312
-
313
- const mock = new Mock({
314
- contextWindow: 100000,
315
- responses: [{ assistant: { content: "<<SEND[200]:hi:SEND", reasoning: null } }],
316
- });
317
- const result = await mock.generate({ messages: [] });
155
+ ```dotenv
156
+ PLURNK_PROVIDERS_PROVIDER_ACME_NPM=@ai-sdk/openai-compatible
157
+ PLURNK_PROVIDERS_PROVIDER_ACME_BASE_URL=https://api.acme.example/v1
158
+ PLURNK_PROVIDERS_PROVIDER_ACME_API_KEY_ENV=ACME_API_KEY,ACME_TOKEN
318
159
  ```
319
160
 
320
- `MockResponse.assistant.ops?: unknown[]` is a pre-parsed escape hatch consumed by plurnk-service intg tests (skips parse roundtrip); the consumer casts to `PlurnkStatement[]` on its side. It is typed `unknown[]` deliberately so **this package has no dependency on `@plurnk/plurnk-grammar` at all** — not runtime, not peer, not even a type import — so grammar releases never force a providers re-pin. Production providers don't expose `ops`.
161
+ The named secret remains in the operator environment. `${ENV_NAME}` inside a
162
+ catalog or declared endpoint is expanded at construction and fails clearly
163
+ when absent.
321
164
 
322
- ## §10 Conformance
165
+ ## §6 Provider plugins
323
166
 
324
- A sibling package satisfies the contract when:
167
+ Provider plugins are the escape hatch for a protocol binding not represented by
168
+ Models.dev, installed SDK providers, or a declaration. Most extensibility
169
+ belongs in MCP, schemes, executors, or mimetypes instead.
325
170
 
326
- 1. Default export is a class with `static fromEnv(env, model, options?)` factory.
327
- 2. Instance exposes `contextWindow: number | null` and `model: string` (non-empty).
328
- 3. Instance exposes `countTokens(text): number` and `calculateCost(usage): number`.
329
- 4. `countTokens("")` returns `0`; `countTokens("…")` returns a non-negative integer.
330
- 5. `calculateCost({prompt:0,completion:0,reasoning:0,cached:0,total:0})` returns `0` (or non-negative USD for non-free models).
331
- 6. Identity getters return stable values across reads.
332
- 7. `generate` resolves with a valid `ProviderResponse` shape.
333
- 8. `generate` invoked with a pre-aborted `signal` rejects without making a wire call.
334
- 9. `generate` invoked, then aborted mid-flight rejects within ≤5s; no connection leak.
335
- 10. `assistantRaw` is present (any value, including `null`).
336
- 11. No DB access, no imports from `@plurnk/plurnk-service`.
337
- 12. No runtime import of `@plurnk/plurnk-grammar` parser entry points.
338
- 13. `generate` invoked with `grammar` against a backend without grammar support sends no grammar-related wire fields and does not error (§13).
339
- 14. `generate` that transported a grammar ALWAYS resolves with the bytes; non-conforming output (reject or incomplete) attaches a `grammar_unenforced` `TelemetryEvent` on `response.telemetry` with the divergence position — an observation, never a throw (#24, §13). (Inherited from `OpenAICompatProvider`; bespoke siblings on a non-compat transport implement it.)
171
+ A provider plugin:
340
172
 
341
- Sibling-specific behavioral tests (wire-format compliance, model-family quirks, retry logic) live in each package's own test surface.
173
+ 1. declares `plurnk: { kind: "provider", name }` in `package.json`;
174
+ 2. may use any npm scope;
175
+ 3. default-exports an AI SDK provider with `languageModel(modelId)`;
176
+ 4. peers on compatible `ai` and `@plurnk/plurnk-providers` majors.
342
177
 
343
- ## §11 Shared OpenAI-compatible machinery
178
+ PLURNK adapts the returned language model. The plugin does not implement the
179
+ PLURNK `Provider`, read PLURNK tuning knobs, or reproduce transport policy.
344
180
 
345
- The framework ships the transport spine every OpenAI-compatible provider had been duplicating. Build a sibling *on top of these* — don't re-implement them.
181
+ Discovery is scope-agnostic and memoized per process. Duplicate names fail hard.
182
+ The common plugin trust gate applies before import. A plugin absent from
183
+ Models.dev requires an explicit context-window pin because PLURNK will not guess
184
+ model physics.
346
185
 
347
- - **`OpenAICompatProvider`** — a `Provider` implementation built by composition. Its `generate` does the universal work (merge `signal` with a `PLURNK_PROVIDERS_FETCH_TIMEOUT` deadline, stream the completion, map `usage`, normalize `finishReason` to the §2 set (translating known per-backend cap synonyms -- `max_tokens`, `MAX_TOKENS`, ... -> `length` -- so a consumer's `=== "length"` holds across backends, warning once on any unmapped value, #425), assemble the response). Per-provider deltas arrive as config:
186
+ ## §7 Local capabilities
348
187
 
349
- ```ts
350
- new OpenAICompatProvider({
351
- model, url, // fully-resolved chat-completions URL
352
- fetchTimeoutMs,
353
- headers, // fully-resolved request headers (incl. auth)
354
- contextWindow, // number | null
355
- reasoning, reasoningStyle, // {mode,budget} intent + style: "none"|"think"|"include_reasoning"|"effort"|"effort_explicit"|"template"|"anthropic"
356
- temperature, repeatPenalty, frequencyPenalty, // sampling + anti-degeneration floor; frequency_penalty guards the plain cloud path (#426)
357
- countTokens, calculateCost, // strategies; default heuristic / free
358
- grammarStyle, // "none" | "llamacpp" — optional local GBNF transport (§13)
359
- gbnfDebug, // PLURNK_PROVIDERS_GBNF_DEBUG: validate a grammar locally + throw on invalid, but DON'T send it (§13); default false
360
- streaming, // SSE transport; default true (false → one non-streamed JSON)
361
- supportsSlotPinning, slotCount, // INTERNAL slot-affinity wiring (run→id_slot); never consumer-facing
362
- topLogprobs, rawBody, // #36 opt-in data capture (PLURNK_PROVIDERS_TOP_LOGPROBS / _RAWBODY); default off
363
- servedModel, // #37 backend's self-reported served id (from the probe) → Provider.servedModel
364
- });
365
- ```
188
+ The `openai` local adapter probes `/v1/models`. A llama-server fingerprint may
189
+ also expose:
366
190
 
367
- The `openai` standard provider sets `grammarStyle: "llamacpp"`, `supportsSlotPinning`, and `slotCount` from the same llama-server fingerprint (`/v1/models` `meta` block + `/props`). The worker→slot mapping lives inside `OpenAICompatProvider`: sticky per `workerId`, round-robin across new runs, LRU-bounded.
191
+ - the actual served model and per-slot context window;
192
+ - GBNF constrained sampling;
193
+ - slot count and worker-sticky slot affinity;
194
+ - EOS marker removal;
195
+ - exact `/tokenize`;
196
+ - the requirement that the caller provide `maxTokens`.
368
197
 
369
- - **`chatCompletionStream` / `chatCompletion` / `OpenAiHttpError` / `StreamResponse`** — the shared HTTP client (`chatCompletionStream` for SSE, `chatCompletion` for the non-streamed JSON the `streaming: false` path uses). One shared copy.
370
- - **`normalizeUsage(raw, reasoningText?, contentText?)` / `calculateCostUsd(usage, rates)`** — usage normalization to the §2 invariant (handles all three reasoning-reporting conventions; the optional text args feed the Fireworks re-split, #425) and the single cost formula. Rates use the Models.dev convention of USD per million tokens; billable output is `completion + reasoning`. `OpenAICompatProvider` applies `normalizeUsage` automatically; provider instances expose `calculateCost(usage)`.
371
- - **`parseRequiredInt` / `parseOptionalInt` / `requireEnv`** — env helpers; each takes a provider `label` for error prefixing.
372
- - **`effortFromBudget(budget)`** — the shared reasoning-budget → `low|medium|high` breakpoints.
198
+ `PLURNK_PROVIDERS_LLAMA_SERVER` may force or disable detection. Probe attempts
199
+ and delay are knobs. A failed probe does not silently assert capabilities.
373
200
 
374
- A **bespoke sibling** therefore reduces to a thin class whose `fromEnv` probes whatever it needs (model catalog, pricing, context window), builds the config, and returns `new OpenAICompatProvider(config)`. A **standard provider** (§5 tier 1) needs no sibling at all — it's a frozen entry in `STANDARD_PROVIDERS` describing its key var, base-URL var, reasoning style, and tokenizer; `standardProviderFromEnv(name, env, model)` (async — returns `Promise<Provider | null>`) does the rest. The endpoint's **canonical URL ships as a floored default** in `.env.defaults` (set-if-unset, overridable in the operator's env or per-alias); it is read from the base-URL var (or a `baseUrlFromEnv` deriver) with **no in-code default**, the value living in the shipped floor, never baked into the table. Only the API **key** is required operator config (a secret with no default; fail-hard when unset).
201
+ Ollama probes `/api/show` for its model context and uses its documented
202
+ OpenAI-compatible generation endpoint through the SDK adapter.
375
203
 
376
- The `plurnk` entry alone sets **`firstPartyMetadata: true`** — it forwards the consumer's per-turn `generate()` `attributions` (which installed plugin packages dispatched) and `client` (the originating frontend, e.g. `plurnk.nvim/1.4.0`) as `Plurnk-Attribution` / `Plurnk-Client` headers, and the opaque `workerId` as `Plurnk-Worker-Id` (#26, wire-name completed #511), and the lineage root `primaryWorkerId` as `Plurnk-Worker-Primary` (#522 — root-vs-descendant classification and worker-tree grouping key, always stamped when supplied). The gate lives on the provider, not the call site, so these first-party signals are structurally incapable of reaching a third-party backend. Empty values emit no header. Endpoint response metadata follows the same general pass-through contract as every provider.
204
+ ## §8 Request authority
377
205
 
378
- **Prompt-cache affinity (`promptCacheKey`, #518).** Standard providers send the OpenAI-standard `prompt_cache_key` set to the `workerId` on every request, **default-ON** (opt out per spec). Serverless backends prompt-cache automatically but the cache is REPLICA-LOCAL; without an affinity key a worker's turns scatter across replicas and the stable prefix never hits (verified live: `cached_tokens` 0 without the key). `workerId` -- already the slot-affinity identity -- is exactly the opaque per-conversation key the cache wants, so a worker's turns pin to one replica and its stable prefix caches. Managed + reserved from caller `sampling`. It's the OpenAI-standard field and broadly accepted -- verified live on fireworks, together, deepinfra, xai, openrouter, and llama-server (6/6, all accept it, every serverless one caches). A backend that caches by a DIFFERENT mechanism opts out (`anthropic`: cache_control breakpoints); a backend later found to strict-reject the field opts out the same way.
206
+ The caller's `sampling` bag expresses sampling intent. It cannot override:
379
207
 
380
- A spec may carry a **`modelPrefix`** — a constant model-id segment the backend requires but the operator's alias shouldn't repeat. `fireworks` sets `"accounts/fireworks/models/"`, so `PLURNK_MODEL_fast=fireworks/deepseek-v4-pro` carries only the distinctive tail; `standardProviderFromEnv` prepends it idempotently to form the wire id, which is **also** the catalog key (models.dev keys fireworks-ai on the full id). A fully qualified Fireworks resource under `accounts/fireworks/` is preserved verbatim, including `routers/` and `deployments/`. Specs without a `modelPrefix` use the model string verbatim.
208
+ - model or messages;
209
+ - stream mode;
210
+ - grammar or response format;
211
+ - backend slot;
212
+ - data-capture settings;
213
+ - tool, modality, or multi-choice behavior;
214
+ - the consumer-owned output envelope;
215
+ - prompt-cache identity.
381
216
 
382
- `contextWindow` for a standard provider resolves (#419): `PLURNK_PROVIDERS_CONTEXT_WINDOW` -> endpoint `n_ctx` (for `probeNctx`-flagged specs like `openai`, queried from `GET /v1/models`: llama-server reports its loaded window at `data[].meta.n_ctx`, vLLM top-level; cloud endpoints don't) -> the `@plurnk/plurnk-models` catalog -> **then the hybrid: a PROBING provider degrades to `null`, a CLOUD provider (no probe) FAILS HARD** (uncataloged + unpinned = config error, the #417 kimi case, not a guessed window). The same probe fingerprints llama-server (the `meta` block) to enable grammar transport (§13), and reads the row's `id` as `servedModel` (#37) — the real served name behind a local alias — so it runs even when the env var pins the window. The probe is best-effort: any failure resolves to `null` context / no grammar capability (a legitimate "unknown"), never throws. For a PROBING provider, an underivable window (env, probe, and catalog ALL missed) is surfaced once via a **`PLURNK_CONTEXT_UNKNOWN`** warning naming the model and the remediation (`PLURNK_PROVIDERS_CONTEXT_WINDOW`, alias-scopable) -- null stays legitimate but never silent (a CLOUD provider throws here instead, above). Operator-facing warnings (`PLURNK_TOKENIZER_HEURISTIC`, `PLURNK_PROBE_FAILED`, `PLURNK_GRAMMAR_UNVERIFIABLE`, `PLURNK_CONTEXT_UNKNOWN`, `PLURNK_FINISH_REASON_UNKNOWN`) are deduplicated **once per process per (code, message)** (#40) — repeat constructions don't re-fire them, but a *different* provider/model's first surfacing is never suppressed.
217
+ Generic AI SDK calls accept only settings represented by the SDK's portable
218
+ surface. Compatible endpoints may carry additional sampling keys after reserved
219
+ keys are removed.
383
220
 
384
- ## §12 Telemetry — provider failures
221
+ First-party attribution, client, strike, workspace, loop, turn, and worker
222
+ headers are sent only by the `plurnk` provider. They never leak to another
223
+ backend.
385
224
 
386
- Transport failures surface as a `ProviderError` (extends `Error`, so existing catchers keep working) that carries the plurnk **TelemetryEvent** envelope via `toTelemetryEvent()`:
387
-
388
- ```ts
389
- { source: "provider:<vendor>", kind: ProviderTelemetryKind, message: string, position: null }
390
- ```
225
+ ## §9 Failures, retries, and cancellation
391
226
 
392
- - `source` is `provider:<vendor>` (schema pattern `^[a-z]+(:[a-z][a-z0-9-]*)?$`); standard providers set it from their name, siblings via the `source` config field (default `"provider"`).
393
- - `kind` ∈ `rate_limit | network_failure | model_refused | invalid_response | unauthorized | quota_exceeded | grammar_invalid | grammar_unenforced`. HTTP status maps: 401/403→`unauthorized`, 402→`quota_exceeded`, 429→`rate_limit`, ≥500→`network_failure`, a 422 whose `error.type` is `grammar_invalid`→`grammar_invalid` (transient: a stochastic output reject a fresh sample may satisfy, #548), other 4xx→`invalid_response`; timeouts/fetch errors→`network_failure`. (`model_refused` is response-level — minted consumer-side from a `content_filter` finish reason, not from a thrown error.) **`grammar_unenforced`** is response-level too, and ALWAYS an observation, never a throw (#24, §2 Promises): whether the grammar was transported or withheld (GBNF-filter mode), a completed exchange returns its bytes with a **non-fatal `TelemetryEvent`** on `ProviderResponse.telemetry` carrying the divergence `position`, so the consumer can drive discard/retry/self-correction. `ProviderError` stays reserved for exchanges that did NOT complete.
394
- - `message` is terse and factual (no guidance prose); `position` is `null` (provider failures aren't localizable into prior content).
395
- - **Caller-initiated abort is NOT telemetry** — an aborted `signal` rethrows the original abort, never a `ProviderError`.
227
+ Provider failures normalize to `ProviderError` with source, kind, status where
228
+ available, and the original cause. A caught failure is surfaced or deliberately
229
+ preserved; it is never converted into an empty model turn.
396
230
 
397
- The `TelemetryEvent` shape is mirrored **locally** (`./telemetry.ts`), structurally matching `@plurnk/plurnk-grammar`'s `TelemetryEvent.json`, so the framework keeps zero grammar dependency (§11). Consumers route provider events through the same `source`+`kind` discriminator as parse/rail events.
231
+ The AI SDK owns attempt scheduling. `PLURNK_PROVIDERS_RETRY_ATTEMPTS` is the
232
+ maximum retry count. Caller cancellation spans the operation. Total and
233
+ stream-chunk deadlines are separately configurable.
398
234
 
399
- ## §13 Grammar-constrained sampling (GBNF)
235
+ HTTP 408, 409, 429, and ordinary 5xx responses are retryable unless the endpoint
236
+ explicitly says otherwise. Endpoint control responses 520–527 are final so a
237
+ router can prevent multiplicative retries behind its own retry policy.
400
238
 
401
- GBNF is an optional aid for local llama.cpp hobbyists, not the PLURNK language
402
- contract and not a baseline cloud capability. The canonical language is parsed
403
- by `@plurnk/plurnk-grammar`'s ANTLR grammar. That package also ships a generated
404
- `plurnk.gbnf` whose language is a tested subset of the canonical grammar.
239
+ ## §10 Grammar
405
240
 
406
- - **plurnk-grammar** owns the artifact (canonical-form GBNF, `L(GBNF) ⊂ L(ANTLR)` invariant, tests).
407
- - **This layer** detects or accepts an operator pin for llama-server, transports
408
- the caller's GBNF verbatim as its top-level `grammar` field, and reports
409
- conformance. Cloud providers never receive a grammar-related field.
410
- - **The consumer** decides whether to configure a local constraint and which
411
- artifact to send. Endpoint-managed constraints are endpoint settings, not a
412
- provider capability inferred from their absence here.
241
+ GBNF is a local llama-server capability, not a generic provider expectation.
242
+ The consumer chooses whether to supply a grammar. The provider never creates or
243
+ rewrites one.
413
244
 
414
- `grammarStyle` is `"none"` or `"llamacpp"`. A llama-server fingerprint or
415
- `PLURNK_PROVIDERS_LLAMA_SERVER=1` selects `"llamacpp"`; all other providers
416
- remain `"none"`. `constrainsOutput` is true only for the former.
245
+ When transported, output is validated locally after completion. Divergence
246
+ attaches `grammar_unenforced` telemetry with its position; the bytes still
247
+ return. `PLURNK_PROVIDERS_GBNF_DEBUG` validates but withholds the grammar and
248
+ compares the unconstrained result for diagnostics.
417
249
 
418
- When a grammar is transported, the provider independently validates returned
419
- content with `@plurnk/gbnf`. A non-accept verdict attaches
420
- `grammar_unenforced` telemetry without discarding the completed response.
421
- `meta.railsAttached` and `meta.railsVerdict` record the observed local
422
- transport and verdict. If the validator cannot parse the supplied grammar, the
423
- provider emits `PLURNK_GRAMMAR_UNVERIFIABLE`.
250
+ ## §11 Evidence and metadata
424
251
 
425
- `PLURNK_PROVIDERS_GBNF_DEBUG` validates and withholds an otherwise transportable
426
- local grammar, then reports how the unconstrained output diverges. It is a
427
- development diagnostic, not a cloud compatibility mode.
252
+ `assistantRaw` is an opaque normalized transport record. Provider top-level
253
+ metadata is forwarded as an open bag without reinterpreting currencies or
254
+ vendor fields.
428
255
 
429
- Hard constraints can amplify repetition and do not replace the normal output
430
- envelope. The llama.cpp path therefore carries its configured
431
- `repeat_penalty`; ordinary cloud requests use the standard configured
432
- `frequency_penalty`. The GBNF string still arrives per call, so this package
433
- does not depend on the PLURNK grammar artifact.
256
+ Logprobs and verbatim response capture are opt-in, alias-scoped dataset features.
257
+ When disabled, the request asks for neither and the response carries neither.
258
+ When enabled, raw per-token model logprob is canonical; alternatives are
259
+ preserved when returned. Raw body/chunks preserve wire evidence the normalized
260
+ record omits.
434
261
 
435
- ## §14 Data capture — logprobs + verbatim body (#36)
262
+ Readable reasoning and encrypted reasoning are separate. Encrypted reasoning is
263
+ preserved verbatim and never decoded or synthesized.
436
264
 
437
- Two OPT-IN knobs surface the full signal of a paid turn for downstream IQ scoring and model distillation. Both are **OFF by default** and **per-alias-scopable** (`PLURNK_PROVIDERS_<KNOB>_<alias>`): the flag *is* the isolation, so a serving turn requests nothing on the wire and carries nothing on the response — only a dataset-scraping alias opts in. Universal: any provider (standard or plugin), any backend that returns the data.
265
+ ## §12 Generation envelopes
438
266
 
439
- **`PLURNK_PROVIDERS_TOP_LOGPROBS`** (non-negative int = the OpenAI `top_logprobs`; unset = off). When set, `generate` requests `logprobs:true, top_logprobs:<n>` and surfaces `response.assistant.logprobs: Array<{ token, logprob, top? }>` plus `assistant.meanLogprob`. These are **managed fields** — reserved from caller `sampling`, so the env flag is the single control (a proxy consumer can't forge them). A backend that returns no logprobs yields an absent field — **never synthesized**.
267
+ Reasoning and completion reserves are percentages of the resolved context
268
+ window or absolute token counts. Absolute pins win. The provider reports the
269
+ resolved reserves; the consumer owns prompt packing and the per-call output cap.
440
270
 
441
- **The `logprob` vs `sampling_logprob` decision (the honest-confidence call).** Fireworks returns both per token: `logprob` (raw model log-probability) and `sampling_logprob` (post-sampling-transform). The structured `logprob` we surface is the **raw** value — the sampling-transform-invariant measure of the model's native belief, the correct confidence signal AND distillation target. This was settled empirically, not by assumption: under grammar the two measured **identical to full float precision**, including an *adversarial* mask (grammar forcing a token the model assigned ~8%: `logprob` −2.5229365 == `sampling_logprob` −2.5229365). A post-mask renormalization would inflate confidence toward the constraint; the raw value stays honest. Anyone wanting `sampling_logprob` reads it from `rawBody`.
271
+ Models.dev `maxOutput` constrains a percentage-derived completion reserve but
272
+ does not override an explicit absolute operator choice. Unknown model physics
273
+ remain unknown.
442
274
 
443
- **`PLURNK_PROVIDERS_RAWBODY`** (truthy = on). When on, `response.rawBody` carries the **verbatim** backend body — the full wire JSON for a non-streamed turn (exact; grammar turns already run non-streamed), or the reassembled equivalent for a streamed one. `assistantRaw` remains the normalized **digest** (it drops `choices[]`); `rawBody` is the capture-everything record that keeps `sampling_logprob`, `token_id`, `bytes`, and any backend-specific per-token fields. Off by default so serving turns never pay the retention cost.
275
+ ## §13 Capacity pool
444
276
 
445
- ## §15 Capacity pool (`Pool`)
277
+ `Pool` fronts interchangeable `Provider` instances. It keeps workers sticky for
278
+ cache locality, selects a healthy sibling for overflow, and preserves the same
279
+ Provider contract. Whether endpoints are interchangeable is a consumer
280
+ decision, not inferred from provider names.
446
281
 
447
- `Pool` fronts **N interchangeable backends as one `Provider`** - capacity scaling, not model blend. It ships the MECHANISM (round-robin across workers, sticky within a worker, overflow to a healthy sibling); the blend/escalation DECISION (which SKU, when to switch) stays the **consumer's**, one level up, by choosing WHICH pool to call. `new Pool(backends: Provider[])` - the consumer resolves the backends (per-alias `instantiateProvider`) and composes them.
282
+ ## §14 Conformance
448
283
 
449
- **Interchangeable, or it throws.** Construction fails on mixed `model`: a heterogeneous "pool" is the consumer's per-turn selection, not this primitive. The surface is the honest aggregate - `contextWindow` is the **safe floor** (min; `null` if any backend's window is unknown, so the consumer never improvises a cap, #421) with its matching reserves; `constrainsOutput` is claimed only if EVERY backend does; `requiresMaxTokens` if ANY does; `servedModel` the common id (else absent); `countTokens`/`tokenize`/`calculateCost` delegate.
284
+ Coverage MUST prove:
450
285
 
451
- **Affinity is the point (§11, one level up).** A worker's turns stick to one backend so its stable prompt prefix keeps hitting the same KV cache; scattering a worker across backends shreds the prefix cache (#531). `worker -> backend` is the `worker -> slot` slot-affinity pattern (#11) across a fleet: round-robin assigns a NEW worker, a returning worker re-pins, the map is LRU-bounded (`N*8`).
286
+ - catalog, declaration, local, and plugin resolution;
287
+ - exact and unique-suffix model lookup;
288
+ - native SDK request mapping and normalized responses;
289
+ - compatible extension preservation;
290
+ - timeout, retry, cancellation, and final-error behavior;
291
+ - local capability probes and pins;
292
+ - usage, costs, evidence, metadata isolation, and grammar observation;
293
+ - alias scoping and fail-hard invalid configuration.
452
294
 
453
- **Overflow is availability-only.** A backend that throws `network_failure`/`rate_limit` (having already exhausted its own transient retries, §11) hands the worker to the next untried backend and **re-sticks** it there (its cache moves with it; worst case one cold prefill). Auth/quota/content failures and caller aborts **propagate** - a peer fails them the same, and failing over would only multiply the spend. Whole fleet down -> the last error is thrown.
295
+ Mock-only green tests do not establish a vendor integration. Live drills and
296
+ integration tests complement this contract; they do not replace its unit-level
297
+ proof.