@plurnk/plurnk-providers 1.3.4 → 1.3.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.env.defaults +41 -52
- package/README.md +44 -53
- package/SPEC.md +215 -354
- package/dist/AiSdkProvider.d.ts +78 -0
- package/dist/AiSdkProvider.d.ts.map +1 -0
- package/dist/AiSdkProvider.js +591 -0
- package/dist/AiSdkProvider.js.map +1 -0
- package/dist/Mock.d.ts +1 -1
- package/dist/Mock.d.ts.map +1 -1
- package/dist/Mock.js +1 -1
- package/dist/Mock.js.map +1 -1
- package/dist/OpenAICompat.d.ts +3 -5
- package/dist/OpenAICompat.d.ts.map +1 -1
- package/dist/OpenAICompat.js +44 -133
- package/dist/OpenAICompat.js.map +1 -1
- package/dist/Pool.d.ts +1 -1
- package/dist/Pool.d.ts.map +1 -1
- package/dist/Pool.js +1 -1
- package/dist/Pool.js.map +1 -1
- package/dist/ProviderRegistry.d.ts.map +1 -1
- package/dist/ProviderRegistry.js +37 -24
- package/dist/ProviderRegistry.js.map +1 -1
- package/dist/aiSdkTransport.d.ts +52 -0
- package/dist/aiSdkTransport.d.ts.map +1 -0
- package/dist/aiSdkTransport.js +294 -0
- package/dist/aiSdkTransport.js.map +1 -0
- package/dist/catalogProvider.d.ts +15 -0
- package/dist/catalogProvider.d.ts.map +1 -0
- package/dist/catalogProvider.js +103 -0
- package/dist/catalogProvider.js.map +1 -0
- package/dist/compatibleProvider.d.ts +3 -0
- package/dist/compatibleProvider.d.ts.map +1 -0
- package/dist/compatibleProvider.js +146 -0
- package/dist/compatibleProvider.js.map +1 -0
- package/dist/discover.d.ts.map +1 -1
- package/dist/discover.js.map +1 -1
- package/dist/env.d.ts +1 -0
- package/dist/env.d.ts.map +1 -1
- package/dist/env.js +13 -6
- package/dist/env.js.map +1 -1
- package/dist/index.d.ts +5 -7
- package/dist/index.d.ts.map +1 -1
- package/dist/index.js +4 -8
- package/dist/index.js.map +1 -1
- package/dist/ollama.d.ts +3 -0
- package/dist/ollama.d.ts.map +1 -0
- package/dist/ollama.js +39 -0
- package/dist/ollama.js.map +1 -0
- package/dist/openai.d.ts +2 -4
- package/dist/openai.d.ts.map +1 -1
- package/dist/openai.js +1 -2
- package/dist/openai.js.map +1 -1
- package/dist/sdkModels.d.ts +13 -0
- package/dist/sdkModels.d.ts.map +1 -0
- package/dist/sdkModels.js +153 -0
- package/dist/sdkModels.js.map +1 -0
- package/dist/standardProviders.d.ts +0 -1
- package/dist/standardProviders.d.ts.map +1 -1
- package/dist/standardProviders.js +9 -11
- package/dist/standardProviders.js.map +1 -1
- package/dist/telemetry.d.ts.map +1 -1
- package/dist/telemetry.js +20 -9
- package/dist/telemetry.js.map +1 -1
- package/dist/types.d.ts +4 -3
- package/dist/types.d.ts.map +1 -1
- package/dist/usage.d.ts +1 -1
- package/dist/usage.d.ts.map +1 -1
- package/dist/usage.js +4 -2
- package/dist/usage.js.map +1 -1
- package/package.json +19 -8
- package/src/{OpenAICompat.test.ts → AiSdkProvider.test.ts} +202 -177
- package/src/{OpenAICompat.ts → AiSdkProvider.ts} +89 -144
- package/src/Mock.test.ts +3 -3
- package/src/Mock.ts +1 -1
- package/src/Pool.test.ts +3 -3
- package/src/Pool.ts +1 -1
- package/src/ProviderRegistry.test.ts +40 -27
- package/src/ProviderRegistry.ts +35 -24
- package/src/aiSdkTransport.test.ts +253 -0
- package/src/aiSdkTransport.ts +369 -0
- package/src/boundaries.test.ts +2 -2
- package/src/catalogProvider.test.ts +100 -0
- package/src/catalogProvider.ts +151 -0
- package/src/compatibleProvider.test.ts +44 -0
- package/src/compatibleProvider.ts +205 -0
- package/src/discover.test.ts +12 -12
- package/src/discover.ts +3 -6
- package/src/env.ts +14 -6
- package/src/index.ts +7 -11
- package/src/ollama.ts +63 -0
- package/src/openai.ts +2 -8
- package/src/sdkModels.test.ts +47 -0
- package/src/sdkModels.ts +194 -0
- package/src/telemetry.test.ts +17 -10
- package/src/telemetry.ts +22 -14
- package/src/types.ts +10 -13
- package/src/usage.test.ts +8 -10
- package/src/usage.ts +7 -3
- package/src/openaiStream.ts +0 -310
- package/src/standardProviders.test.ts +0 -949
- package/src/standardProviders.ts +0 -635
package/SPEC.md
CHANGED
|
@@ -1,436 +1,297 @@
|
|
|
1
|
-
#
|
|
1
|
+
# Provider Contract
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
`@plurnk/plurnk-providers` adapts model endpoints to one stable PLURNK
|
|
4
|
+
`Provider`. It does not maintain a parallel vendor registry or reproduce
|
|
5
|
+
ordinary provider protocols.
|
|
4
6
|
|
|
5
|
-
## §1
|
|
7
|
+
## §1 Ownership
|
|
6
8
|
|
|
7
|
-
|
|
9
|
+
The provider stack has four owners:
|
|
8
10
|
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
11
|
+
1. Models.dev supplies a build-time snapshot of provider package, API endpoint,
|
|
12
|
+
credential names, models, context windows, output limits, and USD prices.
|
|
13
|
+
2. Official AI SDK providers own vendor request and response protocols.
|
|
14
|
+
3. This package owns the PLURNK contract: aliases, envelopes, normalized usage
|
|
15
|
+
and errors, evidence, local capabilities, and first-party metadata.
|
|
16
|
+
4. The operator owns secrets, machine-specific endpoints, and deliberate
|
|
17
|
+
metadata overrides through environment variables.
|
|
15
18
|
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
19
|
+
Facts MUST have one owner. Do not copy a cataloged endpoint, credential name,
|
|
20
|
+
model prefix, context window, price, or vendor request shape into a PLURNK
|
|
21
|
+
table. A missing or wrong catalog fact is fixed upstream, overridden through a
|
|
22
|
+
provider declaration, or left explicitly unknown.
|
|
20
23
|
|
|
21
24
|
## §2 Provider interface
|
|
22
25
|
|
|
26
|
+
`Provider` exposes immutable model facts and one generation operation:
|
|
27
|
+
|
|
23
28
|
```ts
|
|
24
29
|
interface Provider {
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
// Tokenomic primitives (synchronous, pure)
|
|
38
|
-
countTokens(text: string): number;
|
|
39
|
-
costFor(usage: ProviderUsage): number; // pico-USD (1e-12 USD)
|
|
40
|
-
|
|
41
|
-
// OPTIONAL capability: exact tokenization served by the backend's own vocab
|
|
42
|
-
// (llama-server /tokenize). Probe-gated — undefined means the backend can't.
|
|
43
|
-
tokenize?(text: string): Promise<number[]>;
|
|
44
|
-
|
|
45
|
-
// OPTIONAL resolved capabilities — introspectable facts for boot-time policy:
|
|
46
|
-
// constrainsOutput (#34): a transported grammar WILL constrain this decode
|
|
47
|
-
// (rails live) — consumers fail hard on a dark-rails boot instead of
|
|
48
|
-
// discovering it from unconstrained emissions.
|
|
49
|
-
// requiresMaxTokens (#43): this backend decodes UNBOUNDED absent a caller cap
|
|
50
|
-
// (llama-server honors n_predict to the context wall, providers#10) — a
|
|
51
|
-
// consumer MUST bring an output envelope (§13), and can refuse AT BOOT a
|
|
52
|
-
// local alias whose envelope was never declared. Self-clamping cloud
|
|
53
|
-
// backends never set it; undefined = no claim.
|
|
54
|
-
readonly constrainsOutput?: boolean;
|
|
55
|
-
readonly requiresMaxTokens?: boolean;
|
|
56
|
-
// #507 (owner-ruled): generation-envelope reserves derived from the DETECTED
|
|
57
|
-
// window (floor percentages; absolute per-alias pins win outright). null =
|
|
58
|
-
// underivable -> the consumer's no-cap path. The consumer's prompt budget is
|
|
59
|
-
// window - reasoningReserve - completionReserve - its OWN safety margin;
|
|
60
|
-
// its generation cap is the two reserves pooled. Absent = no claim.
|
|
61
|
-
readonly reasoningReserve?: number | null;
|
|
62
|
-
readonly completionReserve?: number | null;
|
|
63
|
-
|
|
64
|
-
// Transport. `workerId` is REQUIRED: the opaque, stable identity of the
|
|
65
|
-
// consumer's work stream — providers may key backend affinity on it and
|
|
66
|
-
// never interpret it. `grammar` is an optional GBNF string for
|
|
67
|
-
// grammar-constrained sampling (§13) — attached verbatim by capable
|
|
68
|
-
// backends, ignored by all others. `maxTokens` is the consumer's per-call
|
|
69
|
-
// output ceiling (wire `max_tokens`); absent means the server default,
|
|
70
|
-
// which is typically UNBOUNDED. `attributions`/`client` are optional
|
|
71
|
-
// first-party metadata, forwarded as `Plurnk-*` headers ONLY by a provider
|
|
72
|
-
// configured with `firstPartyMetadata` (the plurnk endpoint) and dropped by
|
|
73
|
-
// every other — structurally unable to reach a third-party backend (§11).
|
|
74
|
-
// `sampling` is an optional bag of standard OpenAI-compat sampling params
|
|
75
|
-
// (temperature, top_p, top_k, min_p, penalties, stop, seed, …) merged into the
|
|
76
|
-
// body UNDER the managed fields — model/messages/grammar/reasoning/max_tokens/
|
|
77
|
-
// slot always win, and reserved keys are stripped (#477): transport/protocol
|
|
78
|
-
// (stream, response_format, grammar, id_slot, logprobs), paradigm breakers
|
|
79
|
-
// (n, the tools/functions family, modalities/audio, prediction), and the
|
|
80
|
-
// token caps (max_tokens/max_completion_tokens -- the envelope is the managed
|
|
81
|
-
// maxTokens, never bypassable). It carries sampling intent + platform knobs
|
|
82
|
-
// only (§8). For a PROXY consumer forwarding its own
|
|
83
|
-
// caller's sampling knobs (the plurnk endpoint fronting gemma/Fireworks); a
|
|
84
|
-
// direct consumer leaves it unset.
|
|
85
|
-
// `strikes` is the worker's CURRENT rail-strike streak at time-of-generate
|
|
86
|
-
// (0 = clean, distinct from absent = unreported; contract plurnk-service#313).
|
|
87
|
-
// Forwarded as `Plurnk-Strikes` ONLY under the firstPartyMetadata gate,
|
|
88
|
-
// dropped everywhere else. Headers only — never placed in the packet.
|
|
89
|
-
// `workspaceId`/`loop`/`turn` (#404): the turn coordinate, stamped as
|
|
90
|
-
// `Plurnk-Workspace-Id`/`Plurnk-Loop`/`Plurnk-Turn` under the SAME gate.
|
|
91
|
-
// 1-based; absent/0 emits no header. Headers only, never the packet.
|
|
92
|
-
// `primaryWorkerId` (#522): the ROOT worker of this turn's lineage (the
|
|
93
|
-
// no-parent ancestor the worker tree descends from) — a worker id, stamped as
|
|
94
|
-
// `Plurnk-Worker-Primary` under the SAME gate. Lets a consumer classify
|
|
95
|
-
// root-vs-descendant by equality (`primaryWorkerId == workerId` ⇒ the primary
|
|
96
|
-
// worker) and group a worker tree by its root, for telemetry/analytics. The
|
|
97
|
-
// provider EMITS what the consumer supplies and never invents a primary; the
|
|
98
|
-
// consumer's contract is to supply it every turn (including the primary's own,
|
|
99
|
-
// where it equals workerId). Absent emits no header.
|
|
100
|
-
generate(args: { messages: ChatMessage[]; workerId: string; primaryWorkerId?: string; signal?: AbortSignal; grammar?: string; maxTokens?: number; attributions?: string[]; client?: string; strikes?: number; workspaceId?: string; loop?: number; turn?: number; sampling?: Record<string, unknown> }): Promise<ProviderResponse>;
|
|
101
|
-
}
|
|
102
|
-
|
|
103
|
-
interface ProviderResponse {
|
|
104
|
-
assistant: {
|
|
105
|
-
content: string; // raw model emission; consumer parses
|
|
106
|
-
reasoning: string | null; // wire-reported reasoning content; null if absent
|
|
107
|
-
// sealed relay reasoning (#482): items { id, subtype, encrypted:
|
|
108
|
-
// [{data, format}] } verbatim, never decoded; `id` from the wire,
|
|
109
|
-
// `subtype` from wire position (message-attached; plurnk is tools-in-body
|
|
110
|
-
// so it is constant). Absent when none. agui projects REASONING_ENCRYPTED_VALUE.
|
|
111
|
-
reasoningEncrypted?: Array<{ id: string | null; subtype: string; encrypted: Array<{ data: string; format: string | null }> }>;
|
|
112
|
-
usage: ProviderUsage; // { prompt, completion, reasoning, cached, total }
|
|
113
|
-
finishReason: "stop" | "length" | "tool_calls" | "content_filter" | null;
|
|
114
|
-
model: string; // wire-reported (may differ from requested for relay providers)
|
|
115
|
-
};
|
|
116
|
-
assistantRaw: unknown; // verbatim wire response for forensics
|
|
117
|
-
meta?: Record<string, unknown>; // per-turn provider→client bag: backend extra fields passed through + validated known keys (e.g. balancePico, pico-USD); absent when empty (#23)
|
|
118
|
-
}
|
|
119
|
-
|
|
120
|
-
interface ProviderUsage {
|
|
121
|
-
prompt: number; // input tokens (cached ones included)
|
|
122
|
-
completion: number; // visible output tokens, EXCLUDING reasoning
|
|
123
|
-
reasoning: number; // reasoning tokens, billed as output
|
|
124
|
-
cached: number; // subset of prompt served from cache
|
|
125
|
-
total: number; // prompt + completion + reasoning
|
|
30
|
+
readonly model: string;
|
|
31
|
+
readonly contextWindow: number | null;
|
|
32
|
+
readonly servedModel?: string;
|
|
33
|
+
readonly constrainsOutput?: boolean;
|
|
34
|
+
readonly requiresMaxTokens?: boolean;
|
|
35
|
+
readonly reasoningReserve?: number | null;
|
|
36
|
+
readonly completionReserve?: number | null;
|
|
37
|
+
|
|
38
|
+
countTokens(text: string): number;
|
|
39
|
+
tokenize?(text: string): Promise<number[]>;
|
|
40
|
+
calculateCost(usage: ProviderUsage): number;
|
|
41
|
+
generate(args: GenerateArgs): Promise<ProviderResponse>;
|
|
126
42
|
}
|
|
127
43
|
```
|
|
128
44
|
|
|
129
|
-
|
|
130
|
-
|
|
131
|
-
|
|
45
|
+
`contextWindow: null` means genuinely unknown. A consumer MUST NOT invent a
|
|
46
|
+
stand-in. A cataloged cloud model without a context window fails construction
|
|
47
|
+
unless the operator pins `PLURNK_PROVIDERS_CONTEXT_WINDOW`. A local probe
|
|
48
|
+
failure degrades to `null` and emits one warning because a transient probe
|
|
49
|
+
failure must not make a usable local endpoint unbootable.
|
|
132
50
|
|
|
133
|
-
`
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
execute streaming completions, buffered completions, retries, and optional
|
|
137
|
-
backend tokenization. The provider supplies its complete URL, headers, body,
|
|
138
|
-
and effective `AbortSignal`; an injected implementation MUST preserve those Web
|
|
139
|
-
API semantics.
|
|
51
|
+
`countTokens` is synchronous and non-negative. The common fallback is a
|
|
52
|
+
conservative chars/2 ruler and is announced once. `tokenize` exists only when
|
|
53
|
+
the endpoint exposes its real vocabulary.
|
|
140
54
|
|
|
141
|
-
|
|
142
|
-
|
|
143
|
-
|
|
144
|
-
owned by `OpenAICompatProvider`. The function is per-instance; replacing
|
|
145
|
-
`globalThis.fetch` is not part of the contract. A platform or vendor binding
|
|
146
|
-
adapter belongs to its consuming integration and returns a fetch-compatible
|
|
147
|
-
`Response`.
|
|
55
|
+
`calculateCost` returns estimated USD. Models.dev rates are converted at the
|
|
56
|
+
provider boundary. Unknown pricing returns `0`; it is not represented as a
|
|
57
|
+
fabricated rate.
|
|
148
58
|
|
|
149
|
-
|
|
150
|
-
surface for this contract. Its transitive module graph MUST NOT import provider
|
|
151
|
-
discovery, registries, filesystem access, or environment-owned construction.
|
|
152
|
-
The package root remains the Node daemon integration surface.
|
|
59
|
+
### Generation
|
|
153
60
|
|
|
154
|
-
|
|
61
|
+
`generate` requires a non-empty, stable, opaque `workerId`. It accepts:
|
|
155
62
|
|
|
156
|
-
- `
|
|
63
|
+
- `messages`: system, user, and assistant text messages;
|
|
64
|
+
- caller cancellation through `signal`;
|
|
65
|
+
- optional `grammar` and `maxTokens`;
|
|
66
|
+
- standard `sampling` intent;
|
|
67
|
+
- first-party attribution, client, strike, workspace, loop, and turn metadata.
|
|
157
68
|
|
|
158
|
-
|
|
159
|
-
|
|
160
|
-
|
|
161
|
-
|
|
162
|
-
|
|
163
|
-
- `contextWindow` resolves to `null` when a PROBING provider (openai/llama-server) can't determine the window (consumer treats null as "no budget info"); a CLOUD provider with no window source FAILS HARD instead (#419, §11).
|
|
164
|
-
- `generate` rejects on signal abort — does NOT resolve with partial content.
|
|
165
|
-
- `generate` transports `grammar` verbatim when the backend supports grammar-constrained sampling, and silently ignores it otherwise (§13). The provider never chooses or modifies the grammar.
|
|
166
|
-
- `generate` **returns for every completed exchange — bytes always present, the conformance verdict attached as an observation.** When a grammar was transported (or validated in filter mode), the returned `content` is checked against it; a non-accept verdict rides `response.telemetry` as a `grammar_unenforced` event (message + divergence `position`) and the response returns normally. The provider transports and observes; it never adjudicates — discard, retry, escalate, or feed-back is consumer policy. This is a grammar-**conformance** check against the grammar the provider already holds — *not* a plurnk-DSL parse (that stays consumer-side, below) — so it remains backend- and DSL-agnostic. `ProviderError` remains reserved for exchanges that did NOT complete (transport failure, abort, boundary violations).
|
|
167
|
-
- **Backend affinity is the provider's internal guarantee, keyed by `workerId`.** The consumer says *which run this is*, never *which backend resource serves it* — raw resource identifiers (slot integers, connections) never cross the contract in either direction. On slot-pinning backends (llama-server `--parallel N>1`), the provider keeps each worker sticky to one slot and spreads distinct runs across slots, so each concurrent run keeps its KV-cache prefix warm (un-pinned routing is the server's similarity heuristic — slot hops re-pay full prefills). Backends without affinity semantics ignore `workerId` entirely.
|
|
69
|
+
It returns the model's raw content and reasoning, normalized usage, normalized
|
|
70
|
+
finish reason, model identity, opaque evidence, optional metadata, and optional
|
|
71
|
+
telemetry. The provider transports and observes model output; it never retries,
|
|
72
|
+
discards, or repairs an otherwise completed exchange because PLURNK grammar did
|
|
73
|
+
not accept it.
|
|
168
74
|
|
|
169
|
-
|
|
75
|
+
Usage obeys:
|
|
170
76
|
|
|
171
|
-
|
|
172
|
-
|
|
173
|
-
|
|
174
|
-
class OpenAI {
|
|
175
|
-
static fromEnv(env: NodeJS.ProcessEnv, model: string, options?: ProviderOptions): OpenAI | Promise<OpenAI> {
|
|
176
|
-
// Read provider-specific env (OPENAI_BASE_URL, OPENAI_API_KEY, ...)
|
|
177
|
-
// plus universal operator knobs (PLURNK_PROVIDERS_REASONING, PLURNK_PROVIDERS_FETCH_TIMEOUT,
|
|
178
|
-
// PLURNK_PROVIDERS_CONTEXT_WINDOW). `options.baseUrl`, when set, is the per-alias
|
|
179
|
-
// endpoint override (PLURNK_BASEURL_<alias>, §5) and wins over the env base URL.
|
|
180
|
-
return new OpenAI({ /* ... */ });
|
|
181
|
-
}
|
|
182
|
-
constructor(config: OpenAIConfig) { /* ... */ }
|
|
183
|
-
}
|
|
77
|
+
```text
|
|
78
|
+
total = prompt + completion + reasoning
|
|
79
|
+
cached ⊆ prompt
|
|
184
80
|
```
|
|
185
81
|
|
|
186
|
-
|
|
187
|
-
|
|
188
|
-
`
|
|
189
|
-
|
|
190
|
-
`fromEnv` MUST fail fast with a clear error if required env is missing — name the env var the operator needs to set.
|
|
82
|
+
`completion` excludes reasoning. Known vendor finish reasons normalize to
|
|
83
|
+
`stop`, `length`, `tool_calls`, or `content_filter`; an unknown value becomes
|
|
84
|
+
`null` and emits a warning.
|
|
191
85
|
|
|
192
|
-
## §
|
|
86
|
+
## §3 AI SDK boundary
|
|
193
87
|
|
|
194
|
-
|
|
88
|
+
Cataloged providers instantiate their Models.dev-declared AI SDK package.
|
|
89
|
+
Standard request shaping, streaming, retries, cancellation, timeouts, usage,
|
|
90
|
+
and vendor error parsing belong to the SDK.
|
|
195
91
|
|
|
196
|
-
|
|
92
|
+
PLURNK maps its generic settings to AI SDK call settings:
|
|
197
93
|
|
|
198
|
-
|
|
199
|
-
-
|
|
200
|
-
-
|
|
201
|
-
-
|
|
202
|
-
-
|
|
203
|
-
- **`PLURNK_PROVIDERS_REASONING_RESERVE` / `PLURNK_PROVIDERS_COMPLETION_RESERVE`** (#507, owner-ruled) -- the generation-envelope reserves, REQUIRED (floor ships `10%` / `25%`). A percentage derives from the DETECTED window (llama-server n_ctx, the plurnk.ai router, the catalog) so every advertising endpoint gets sane defaults with ZERO operator tuning; an absolute token count wins outright (per-alias-scopable -- the measured-envelope override). These MIGRATED from core's `PLURNK_SERVICE_{CONTEXT_WINDOW,REASONING,COMPLETION}` (provider quantities wearing a service prefix; core keeps only its own packing-safety margin). Surfaced as `Provider.reasoningReserve`/`completionReserve`.
|
|
204
|
-
- **`PLURNK_PROVIDERS_TEMPERATURE` / `PLURNK_PROVIDERS_REPEAT_PENALTY` / `PLURNK_PROVIDERS_FREQUENCY_PENALTY`** -- REQUIRED sampling + anti-degeneration floors (read via `parseRequiredFloat`, values from the `.env.defaults` floor). `REPEAT_PENALTY` (canonical `1.15`) is the llama.cpp/Fireworks multiplier; `FREQUENCY_PENALTY` (canonical `0.4`, #426) is the OpenAI-standard guard on the plain cloud path, which has no `repeat_penalty`. Applied to EVERY request, keyed per backend (§13).
|
|
205
|
-
- **`PLURNK_PROVIDERS_SERVICE_TIER`** -- optional fixed request tier, alias-scopable. Fireworks accepts its published `auto | default | flex | priority` vocabulary and owns those values' routing semantics; when configured, the value is validated at construction and wins over per-call sampling on every request. Unset delegates to the provider default. Other standard providers reject the knob rather than silently ignoring a paid routing choice.
|
|
206
|
-
- **`PLURNK_PROVIDERS_DRY_MULTIPLIER` / `_DRY_BASE` / `_DRY_ALLOWED_LENGTH`** -- the llama.cpp DRY loop-breaker (#567), customer-overridable per alias. The generic floor is deliberately **off** (`MULTIPLIER=0`): the turboderp/Gemma sweep proved the community-standard `0.8`/`1.75`/`2` settings worse than off for both runaway emissions and exact-identifier corruption. That same sweep measured `0.8`/`1.75`/`32` as a safe alias-specific deployment setting (zero corruption, about 6% runaways versus 19% off), so `.env.defaults` carries `BASE=1.75` and `ALLOWED_LENGTH=32` as inert override companions without pretending one model's multiplier is universal. DRY penalizes repeated sequences with a penalty escalating in run length -- the tool for a plan-restart loop a single-token `repeat_penalty` over a short window cannot see. **`PLURNK_PROVIDERS_REPEAT_LAST_N`** (optional) widens the older `repeat_penalty` window past the box's 64. These knobs are sent **only on the detected `llamacpp` path**; cloud providers parse but never emit them.
|
|
94
|
+
- `temperature`, `top_p`, `top_k`;
|
|
95
|
+
- presence and frequency penalties;
|
|
96
|
+
- stop sequences and seed;
|
|
97
|
+
- output-token ceiling;
|
|
98
|
+
- `off`, `adaptive`, or budget-derived reasoning intent.
|
|
207
99
|
|
|
208
|
-
|
|
100
|
+
Provider-specific options are permitted only where they preserve a documented
|
|
101
|
+
PLURNK product contract the generic SDK surface cannot express.
|
|
209
102
|
|
|
210
|
-
|
|
103
|
+
The compatible transport is deliberately retained for:
|
|
211
104
|
|
|
212
|
-
|
|
213
|
-
|
|
214
|
-
|
|
215
|
-
|
|
216
|
-
```
|
|
217
|
-
|
|
218
|
-
First path segment names the provider; rest is the model identifier (may contain `/` for tri-level providers like openrouter's `publisher/model`).
|
|
219
|
-
|
|
220
|
-
**Per-alias endpoint override — `PLURNK_BASEURL_<alias>`.** A provider's base URL otherwise binds **one URL per provider *name*** (its `baseUrlVar`, §11), so two `openai/…` aliases collapse onto the same `OPENAI_BASE_URL`. That makes running **N self-hosted boxes of the same kind** — the real case for the two "bring your own box" providers, `openai` (llama.cpp/vLLM/LM Studio) and `ollama` — impossible by name alone. `PLURNK_BASEURL_<alias>` attaches an endpoint to the *alias*, case-folded to match its `PLURNK_MODEL_<alias>`, and **wins over** the provider's own base-URL var:
|
|
221
|
-
|
|
222
|
-
```
|
|
223
|
-
PLURNK_MODEL_HAZEL1=openai/qwen2.5-coder
|
|
224
|
-
PLURNK_BASEURL_HAZEL1=http://hazel1:8080/v1 # llama.cpp box 1
|
|
225
|
-
PLURNK_MODEL_HAZEL2=openai/qwen3-coder
|
|
226
|
-
PLURNK_BASEURL_HAZEL2=http://hazel2:8080/v1 # llama.cpp box 2 — same provider name, different box
|
|
227
|
-
PLURNK_MODEL_NOOK=ollama/qwen2.5-coder
|
|
228
|
-
PLURNK_BASEURL_NOOK=http://nook:11434 # an ollama box; drives BOTH /v1/chat and /api/show
|
|
229
|
-
```
|
|
105
|
+
- `openai` local endpoints, including llama-server and vLLM;
|
|
106
|
+
- `ollama`, after its native `/api/show` probe;
|
|
107
|
+
- the first-party `plurnk` endpoint;
|
|
108
|
+
- operator-declared `@ai-sdk/openai-compatible` providers.
|
|
230
109
|
|
|
231
|
-
|
|
110
|
+
It carries PLURNK-only fields and raw wire evidence without reimplementing the
|
|
111
|
+
SDK's ordinary transport.
|
|
232
112
|
|
|
233
|
-
|
|
113
|
+
## §4 Operator configuration
|
|
234
114
|
|
|
235
|
-
|
|
236
|
-
|
|
237
|
-
|
|
238
|
-
- `loadActiveProvider(env)` — boot convenience: active alias → `instantiateProvider`.
|
|
239
|
-
- `standardProviderFromEnv(name, env, model)` / `isStandardProvider(name)` — the tier-1 internals, still exported.
|
|
240
|
-
- `discover({ cwd?, packageDirs?, env? })` — the tier-2 scan: a `Discovery` (`{ registry, skipped, attributions }`) of every installed provider package. `registry`/`skipped` are `Map<name, packageSpecifier>` partitioned by the trust gate; `attributions` is `Map<name, string | string[]>` carrying each registered provider's raw `plurnk.attribution` for author credit (#21 — surfaced verbatim; the consumer applies the `@plurnk/`-scope reservation policy). Exported so a consumer can enumerate trusted providers — and see which were declined — without instantiating them.
|
|
115
|
+
Every operational value is an environment knob documented in `.env.defaults`.
|
|
116
|
+
There are no hidden tuning constants. Every `PLURNK_PROVIDERS_*` knob may be
|
|
117
|
+
scoped to an alias by appending `_<alias>`; the scoped value wins.
|
|
241
118
|
|
|
242
|
-
|
|
119
|
+
The universal groups are:
|
|
243
120
|
|
|
244
|
-
|
|
245
|
-
|
|
121
|
+
- reasoning activation and optional explicit budget;
|
|
122
|
+
- decode tuning;
|
|
123
|
+
- request, stream-idle, retry, and probe budgets;
|
|
124
|
+
- local GBNF and llama-server capability pins;
|
|
125
|
+
- context-window and generation-envelope overrides;
|
|
126
|
+
- opt-in logprob and raw-body capture.
|
|
246
127
|
|
|
247
|
-
|
|
128
|
+
Operator secrets and machine-specific values never belong in committed
|
|
129
|
+
defaults.
|
|
248
130
|
|
|
249
|
-
|
|
131
|
+
## §5 Resolution
|
|
250
132
|
|
|
251
|
-
|
|
133
|
+
`PLURNK_MODEL_<alias>=<provider>/<model-id>` declares an alias.
|
|
134
|
+
`PLURNK_MODEL=<alias>` selects the boot alias. Model IDs may contain `/`.
|
|
135
|
+
`PLURNK_BASEURL_<alias>` is a per-alias endpoint override.
|
|
252
136
|
|
|
253
|
-
|
|
254
|
-
- Every `generate` carries `workerId` — the worker's stable, opaque identity. Same run → same string across its turns; distinct runs → distinct strings.
|
|
255
|
-
- `signal` is wired to the worker's AbortController.
|
|
256
|
-
- `generate` is single-call per turn. No parallel calls on the same instance.
|
|
257
|
-
- `assistantRaw` is opaque to the consumer (forensics-only).
|
|
258
|
-
- `meta` is the per-turn provider→client metadata bag: the backend's **non-standard top-level response fields passed through verbatim** (every provider), PLUS **validated known keys** the framework holds a contract for — currently `balancePico` (a finite pico-USD number normalized from the plurnk endpoint's balance field, renamed off its raw key; dropped if non-numeric). Absent when the backend reported no extras. The consumer (service) merges `meta` into its Turn metadata and filters what reaches the client; it reads `meta`, never mines `assistantRaw` (#23).
|
|
259
|
-
- `countTokens` is cheap by contract; consumer calls frequently.
|
|
137
|
+
`instantiateProvider` resolves in this order:
|
|
260
138
|
|
|
261
|
-
|
|
139
|
+
1. A Models.dev provider and model, using its declared AI SDK package.
|
|
140
|
+
2. An operator provider declaration:
|
|
141
|
+
`PLURNK_PROVIDERS_PROVIDER_<NAME>_{NPM,BASE_URL,API_KEY_ENV}`.
|
|
142
|
+
3. The local `openai`, `ollama`, or first-party `plurnk` adapter.
|
|
143
|
+
4. A discovered AI SDK provider plugin.
|
|
144
|
+
5. A precise unknown-provider error.
|
|
262
145
|
|
|
263
|
-
|
|
264
|
-
|
|
265
|
-
- **No grammar runtime dep.** Type-imports from `@plurnk/plurnk-grammar` are fine; invoking `PlurnkParser.parse` is consumer-side.
|
|
266
|
-
- **Raw `content`.** Returned verbatim. Tools are expressed in-body as plurnk DSL (see §2); providers do not use native tool-calling.
|
|
267
|
-
- **Atomic.** One `generate` call resolves with one complete `ProviderResponse`. No streaming partial resolves (v0).
|
|
268
|
-
- **Honors `signal`.** Aborted calls reject; resources free; no orphaned connections.
|
|
269
|
-
- **Single model.** One provider instance speaks to one model.
|
|
270
|
-
- **Synchronous `countTokens`, pure `costFor`.** No I/O, no async, no state beyond cached tokenizer artifacts.
|
|
146
|
+
Earlier sources are authoritative. Installed plugins cannot shadow a cataloged
|
|
147
|
+
or explicitly declared name.
|
|
271
148
|
|
|
272
|
-
|
|
149
|
+
Model IDs resolve exactly first. A unique catalog suffix is accepted to avoid
|
|
150
|
+
forcing a vendor-owned resource prefix into PLURNK aliases. Ambiguous suffixes
|
|
151
|
+
fail to resolve.
|
|
273
152
|
|
|
274
|
-
|
|
275
|
-
|---|
|
|
276
|
-
| Database access |
|
|
277
|
-
| Filesystem access beyond reading provider-internal config |
|
|
278
|
-
| Imports from `@plurnk/plurnk-service/*` |
|
|
279
|
-
| Resolving with partial content on abort |
|
|
280
|
-
| Mutating `messages` |
|
|
281
|
-
| Parsing `content` into `PlurnkStatement[]` |
|
|
282
|
-
| Streaming the resolve (atomic only; v0) |
|
|
283
|
-
| Holding state across `generate` calls beyond connection pooling, config, and backend-affinity bookkeeping (run→resource maps, §2) |
|
|
284
|
-
| Exposing backend resource identifiers (slot ids, connections) on the consumer surface |
|
|
285
|
-
| Reading model output via `console.*` |
|
|
286
|
-
| Ignoring `signal` |
|
|
287
|
-
| Spawning subprocesses for inference |
|
|
153
|
+
Provider declarations configure facts, not credentials:
|
|
288
154
|
|
|
289
|
-
|
|
290
|
-
|
|
291
|
-
|
|
292
|
-
|
|
293
|
-
```ts
|
|
294
|
-
import { Mock } from "@plurnk/plurnk-providers";
|
|
295
|
-
|
|
296
|
-
const mock = new Mock({
|
|
297
|
-
contextWindow: 100000,
|
|
298
|
-
responses: [{ assistant: { content: "<<SEND[200]:hi:SEND", reasoning: null } }],
|
|
299
|
-
});
|
|
300
|
-
const result = await mock.generate({ messages: [] });
|
|
155
|
+
```dotenv
|
|
156
|
+
PLURNK_PROVIDERS_PROVIDER_ACME_NPM=@ai-sdk/openai-compatible
|
|
157
|
+
PLURNK_PROVIDERS_PROVIDER_ACME_BASE_URL=https://api.acme.example/v1
|
|
158
|
+
PLURNK_PROVIDERS_PROVIDER_ACME_API_KEY_ENV=ACME_API_KEY,ACME_TOKEN
|
|
301
159
|
```
|
|
302
160
|
|
|
303
|
-
|
|
161
|
+
The named secret remains in the operator environment. `${ENV_NAME}` inside a
|
|
162
|
+
catalog or declared endpoint is expanded at construction and fails clearly
|
|
163
|
+
when absent.
|
|
304
164
|
|
|
305
|
-
## §
|
|
165
|
+
## §6 Provider plugins
|
|
306
166
|
|
|
307
|
-
|
|
167
|
+
Provider plugins are the escape hatch for a protocol binding not represented by
|
|
168
|
+
Models.dev, installed SDK providers, or a declaration. Most extensibility
|
|
169
|
+
belongs in MCP, schemes, executors, or mimetypes instead.
|
|
308
170
|
|
|
309
|
-
|
|
310
|
-
2. Instance exposes `contextWindow: number | null` and `model: string` (non-empty).
|
|
311
|
-
3. Instance exposes `countTokens(text): number` and `costFor(usage): number`.
|
|
312
|
-
4. `countTokens("")` returns `0`; `countTokens("…")` returns a non-negative integer.
|
|
313
|
-
5. `costFor({prompt:0,completion:0,reasoning:0,cached:0,total:0})` returns `0` (or non-negative pico-USD for non-free models).
|
|
314
|
-
6. Identity getters return stable values across reads.
|
|
315
|
-
7. `generate` resolves with a valid `ProviderResponse` shape.
|
|
316
|
-
8. `generate` invoked with a pre-aborted `signal` rejects without making a wire call.
|
|
317
|
-
9. `generate` invoked, then aborted mid-flight rejects within ≤5s; no connection leak.
|
|
318
|
-
10. `assistantRaw` is present (any value, including `null`).
|
|
319
|
-
11. No DB access, no imports from `@plurnk/plurnk-service`.
|
|
320
|
-
12. No runtime import of `@plurnk/plurnk-grammar` parser entry points.
|
|
321
|
-
13. `generate` invoked with `grammar` against a backend without grammar support sends no grammar-related wire fields and does not error (§13).
|
|
322
|
-
14. `generate` that transported a grammar ALWAYS resolves with the bytes; non-conforming output (reject or incomplete) attaches a `grammar_unenforced` `TelemetryEvent` on `response.telemetry` with the divergence position — an observation, never a throw (#24, §13). (Inherited from `OpenAICompatProvider`; bespoke siblings on a non-compat transport implement it.)
|
|
171
|
+
A provider plugin:
|
|
323
172
|
|
|
324
|
-
|
|
173
|
+
1. declares `plurnk: { kind: "provider", name }` in `package.json`;
|
|
174
|
+
2. may use any npm scope;
|
|
175
|
+
3. default-exports an AI SDK provider with `languageModel(modelId)`;
|
|
176
|
+
4. peers on compatible `ai` and `@plurnk/plurnk-providers` majors.
|
|
325
177
|
|
|
326
|
-
|
|
178
|
+
PLURNK adapts the returned language model. The plugin does not implement the
|
|
179
|
+
PLURNK `Provider`, read PLURNK tuning knobs, or reproduce transport policy.
|
|
327
180
|
|
|
328
|
-
|
|
181
|
+
Discovery is scope-agnostic and memoized per process. Duplicate names fail hard.
|
|
182
|
+
The common plugin trust gate applies before import. A plugin absent from
|
|
183
|
+
Models.dev requires an explicit context-window pin because PLURNK will not guess
|
|
184
|
+
model physics.
|
|
329
185
|
|
|
330
|
-
|
|
186
|
+
## §7 Local capabilities
|
|
331
187
|
|
|
332
|
-
|
|
333
|
-
|
|
334
|
-
model, url, // fully-resolved chat-completions URL
|
|
335
|
-
fetchTimeoutMs,
|
|
336
|
-
headers, // fully-resolved request headers (incl. auth)
|
|
337
|
-
contextWindow, // number | null
|
|
338
|
-
reasoning, reasoningStyle, // {mode,budget} intent + style: "none"|"think"|"include_reasoning"|"effort"|"effort_explicit"|"template"|"anthropic"
|
|
339
|
-
temperature, repeatPenalty, frequencyPenalty, // sampling + anti-degeneration floor; frequency_penalty guards the plain cloud path (#426)
|
|
340
|
-
countTokens, costFor, // strategies; default heuristic / free
|
|
341
|
-
grammarStyle, // "none" | "llamacpp" — optional local GBNF transport (§13)
|
|
342
|
-
gbnfDebug, // PLURNK_PROVIDERS_GBNF_DEBUG: validate a grammar locally + throw on invalid, but DON'T send it (§13); default false
|
|
343
|
-
streaming, // SSE transport; default true (false → one non-streamed JSON)
|
|
344
|
-
supportsSlotPinning, slotCount, // INTERNAL slot-affinity wiring (run→id_slot); never consumer-facing
|
|
345
|
-
topLogprobs, rawBody, // #36 opt-in data capture (PLURNK_PROVIDERS_TOP_LOGPROBS / _RAWBODY); default off
|
|
346
|
-
servedModel, // #37 backend's self-reported served id (from the probe) → Provider.servedModel
|
|
347
|
-
});
|
|
348
|
-
```
|
|
188
|
+
The `openai` local adapter probes `/v1/models`. A llama-server fingerprint may
|
|
189
|
+
also expose:
|
|
349
190
|
|
|
350
|
-
|
|
191
|
+
- the actual served model and per-slot context window;
|
|
192
|
+
- GBNF constrained sampling;
|
|
193
|
+
- slot count and worker-sticky slot affinity;
|
|
194
|
+
- EOS marker removal;
|
|
195
|
+
- exact `/tokenize`;
|
|
196
|
+
- the requirement that the caller provide `maxTokens`.
|
|
351
197
|
|
|
352
|
-
|
|
353
|
-
|
|
354
|
-
- **`parseRequiredInt` / `parseOptionalInt` / `requireEnv`** — env helpers; each takes a provider `label` for error prefixing.
|
|
355
|
-
- **`effortFromBudget(budget)`** — the shared reasoning-budget → `low|medium|high` breakpoints.
|
|
198
|
+
`PLURNK_PROVIDERS_LLAMA_SERVER` may force or disable detection. Probe attempts
|
|
199
|
+
and delay are knobs. A failed probe does not silently assert capabilities.
|
|
356
200
|
|
|
357
|
-
|
|
201
|
+
Ollama probes `/api/show` for its model context and uses its documented
|
|
202
|
+
OpenAI-compatible generation endpoint through the SDK adapter.
|
|
358
203
|
|
|
359
|
-
|
|
204
|
+
## §8 Request authority
|
|
360
205
|
|
|
361
|
-
|
|
206
|
+
The caller's `sampling` bag expresses sampling intent. It cannot override:
|
|
362
207
|
|
|
363
|
-
|
|
208
|
+
- model or messages;
|
|
209
|
+
- stream mode;
|
|
210
|
+
- grammar or response format;
|
|
211
|
+
- backend slot;
|
|
212
|
+
- data-capture settings;
|
|
213
|
+
- tool, modality, or multi-choice behavior;
|
|
214
|
+
- the consumer-owned output envelope;
|
|
215
|
+
- prompt-cache identity.
|
|
364
216
|
|
|
365
|
-
|
|
217
|
+
Generic AI SDK calls accept only settings represented by the SDK's portable
|
|
218
|
+
surface. Compatible endpoints may carry additional sampling keys after reserved
|
|
219
|
+
keys are removed.
|
|
366
220
|
|
|
367
|
-
|
|
221
|
+
First-party attribution, client, strike, workspace, loop, turn, and worker
|
|
222
|
+
headers are sent only by the `plurnk` provider. They never leak to another
|
|
223
|
+
backend.
|
|
368
224
|
|
|
369
|
-
|
|
370
|
-
|
|
371
|
-
```ts
|
|
372
|
-
{ source: "provider:<vendor>", kind: ProviderTelemetryKind, message: string, position: null }
|
|
373
|
-
```
|
|
225
|
+
## §9 Failures, retries, and cancellation
|
|
374
226
|
|
|
375
|
-
|
|
376
|
-
|
|
377
|
-
|
|
378
|
-
- **Caller-initiated abort is NOT telemetry** — an aborted `signal` rethrows the original abort, never a `ProviderError`.
|
|
227
|
+
Provider failures normalize to `ProviderError` with source, kind, status where
|
|
228
|
+
available, and the original cause. A caught failure is surfaced or deliberately
|
|
229
|
+
preserved; it is never converted into an empty model turn.
|
|
379
230
|
|
|
380
|
-
The
|
|
231
|
+
The AI SDK owns attempt scheduling. `PLURNK_PROVIDERS_RETRY_ATTEMPTS` is the
|
|
232
|
+
maximum retry count. Caller cancellation spans the operation. Total and
|
|
233
|
+
stream-chunk deadlines are separately configurable.
|
|
381
234
|
|
|
382
|
-
|
|
235
|
+
HTTP 408, 409, 429, and ordinary 5xx responses are retryable unless the endpoint
|
|
236
|
+
explicitly says otherwise. Endpoint control responses 520–527 are final so a
|
|
237
|
+
router can prevent multiplicative retries behind its own retry policy.
|
|
383
238
|
|
|
384
|
-
|
|
385
|
-
contract and not a baseline cloud capability. The canonical language is parsed
|
|
386
|
-
by `@plurnk/plurnk-grammar`'s ANTLR grammar. That package also ships a generated
|
|
387
|
-
`plurnk.gbnf` whose language is a tested subset of the canonical grammar.
|
|
239
|
+
## §10 Grammar
|
|
388
240
|
|
|
389
|
-
|
|
390
|
-
|
|
391
|
-
|
|
392
|
-
conformance. Cloud providers never receive a grammar-related field.
|
|
393
|
-
- **The consumer** decides whether to configure a local constraint and which
|
|
394
|
-
artifact to send. Endpoint-managed constraints are endpoint settings, not a
|
|
395
|
-
provider capability inferred from their absence here.
|
|
241
|
+
GBNF is a local llama-server capability, not a generic provider expectation.
|
|
242
|
+
The consumer chooses whether to supply a grammar. The provider never creates or
|
|
243
|
+
rewrites one.
|
|
396
244
|
|
|
397
|
-
|
|
398
|
-
`
|
|
399
|
-
|
|
245
|
+
When transported, output is validated locally after completion. Divergence
|
|
246
|
+
attaches `grammar_unenforced` telemetry with its position; the bytes still
|
|
247
|
+
return. `PLURNK_PROVIDERS_GBNF_DEBUG` validates but withholds the grammar and
|
|
248
|
+
compares the unconstrained result for diagnostics.
|
|
400
249
|
|
|
401
|
-
|
|
402
|
-
content with `@plurnk/gbnf`. A non-accept verdict attaches
|
|
403
|
-
`grammar_unenforced` telemetry without discarding the completed response.
|
|
404
|
-
`meta.railsAttached` and `meta.railsVerdict` record the observed local
|
|
405
|
-
transport and verdict. If the validator cannot parse the supplied grammar, the
|
|
406
|
-
provider emits `PLURNK_GRAMMAR_UNVERIFIABLE`.
|
|
250
|
+
## §11 Evidence and metadata
|
|
407
251
|
|
|
408
|
-
`
|
|
409
|
-
|
|
410
|
-
|
|
252
|
+
`assistantRaw` is an opaque normalized transport record. Provider top-level
|
|
253
|
+
metadata is forwarded as an open bag without reinterpreting currencies or
|
|
254
|
+
vendor fields.
|
|
411
255
|
|
|
412
|
-
|
|
413
|
-
|
|
414
|
-
|
|
415
|
-
|
|
416
|
-
|
|
256
|
+
Logprobs and verbatim response capture are opt-in, alias-scoped dataset features.
|
|
257
|
+
When disabled, the request asks for neither and the response carries neither.
|
|
258
|
+
When enabled, raw per-token model logprob is canonical; alternatives are
|
|
259
|
+
preserved when returned. Raw body/chunks preserve wire evidence the normalized
|
|
260
|
+
record omits.
|
|
417
261
|
|
|
418
|
-
|
|
262
|
+
Readable reasoning and encrypted reasoning are separate. Encrypted reasoning is
|
|
263
|
+
preserved verbatim and never decoded or synthesized.
|
|
419
264
|
|
|
420
|
-
|
|
265
|
+
## §12 Generation envelopes
|
|
421
266
|
|
|
422
|
-
|
|
267
|
+
Reasoning and completion reserves are percentages of the resolved context
|
|
268
|
+
window or absolute token counts. Absolute pins win. The provider reports the
|
|
269
|
+
resolved reserves; the consumer owns prompt packing and the per-call output cap.
|
|
423
270
|
|
|
424
|
-
|
|
271
|
+
Models.dev `maxOutput` constrains a percentage-derived completion reserve but
|
|
272
|
+
does not override an explicit absolute operator choice. Unknown model physics
|
|
273
|
+
remain unknown.
|
|
425
274
|
|
|
426
|
-
|
|
275
|
+
## §13 Capacity pool
|
|
427
276
|
|
|
428
|
-
|
|
277
|
+
`Pool` fronts interchangeable `Provider` instances. It keeps workers sticky for
|
|
278
|
+
cache locality, selects a healthy sibling for overflow, and preserves the same
|
|
279
|
+
Provider contract. Whether endpoints are interchangeable is a consumer
|
|
280
|
+
decision, not inferred from provider names.
|
|
429
281
|
|
|
430
|
-
|
|
282
|
+
## §14 Conformance
|
|
431
283
|
|
|
432
|
-
|
|
284
|
+
Coverage MUST prove:
|
|
433
285
|
|
|
434
|
-
|
|
286
|
+
- catalog, declaration, local, and plugin resolution;
|
|
287
|
+
- exact and unique-suffix model lookup;
|
|
288
|
+
- native SDK request mapping and normalized responses;
|
|
289
|
+
- compatible extension preservation;
|
|
290
|
+
- timeout, retry, cancellation, and final-error behavior;
|
|
291
|
+
- local capability probes and pins;
|
|
292
|
+
- usage, costs, evidence, metadata isolation, and grammar observation;
|
|
293
|
+
- alias scoping and fail-hard invalid configuration.
|
|
435
294
|
|
|
436
|
-
|
|
295
|
+
Mock-only green tests do not establish a vendor integration. Live drills and
|
|
296
|
+
integration tests complement this contract; they do not replace its unit-level
|
|
297
|
+
proof.
|