@arnilo/prism 0.3.0 → 0.3.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (69) hide show
  1. package/CHANGELOG.md +27 -0
  2. package/README.md +3 -1
  3. package/dist/agent-loops.js +45 -8
  4. package/dist/agent-session/helpers.js +2 -2
  5. package/dist/cache-helpers.d.ts +11 -0
  6. package/dist/cache-helpers.js +29 -5
  7. package/dist/cli-provider-add.js +2 -1
  8. package/dist/context-budget.js +9 -6
  9. package/dist/contracts-core/agent.d.ts +2 -0
  10. package/dist/contracts-core/provider.d.ts +2 -0
  11. package/dist/event-multiplexer.js +0 -4
  12. package/dist/index.d.ts +6 -4
  13. package/dist/index.js +5 -3
  14. package/dist/input.js +19 -11
  15. package/dist/node/session-store-jsonl.js +7 -3
  16. package/dist/providers/openai-compatible.js +2 -1
  17. package/dist/providers/openai-primitives.js +2 -1
  18. package/dist/providers/schema.d.ts +7 -0
  19. package/dist/providers/schema.js +25 -0
  20. package/dist/testing/provider-conformance.d.ts +10 -0
  21. package/dist/testing/provider-conformance.js +37 -0
  22. package/dist/trim-trailing-slashes.d.ts +8 -0
  23. package/dist/trim-trailing-slashes.js +14 -0
  24. package/docs/0.1.0-readiness.md +1 -1
  25. package/docs/acp.md +1 -0
  26. package/docs/ag-ui.md +1 -0
  27. package/docs/agent-loops.md +3 -0
  28. package/docs/agent-session-runtime.md +1 -0
  29. package/docs/browser-automation.md +1 -0
  30. package/docs/database-persistence.md +1 -1
  31. package/docs/graft.md +125 -0
  32. package/docs/host-security.md +3 -1
  33. package/docs/index.md +11 -9
  34. package/docs/input-and-prompt-assembly.md +11 -6
  35. package/docs/instruction-injection.md +1 -1
  36. package/docs/mcp-tools.md +1 -0
  37. package/docs/migration.md +10 -0
  38. package/docs/node-jsonl-session-store.md +1 -1
  39. package/docs/obscura.md +175 -0
  40. package/docs/observability.md +21 -1
  41. package/docs/performance.md +58 -4
  42. package/docs/ponytail.md +1 -1
  43. package/docs/provider-caching.md +13 -11
  44. package/docs/provider-conformance.md +6 -0
  45. package/docs/provider-packages.md +1 -1
  46. package/docs/provider-primitives.md +15 -2
  47. package/docs/providers/ai-sdk.md +1 -1
  48. package/docs/providers/anthropic.md +1 -1
  49. package/docs/providers/azure.md +1 -0
  50. package/docs/providers/bedrock.md +1 -0
  51. package/docs/providers/kimi.md +2 -1
  52. package/docs/providers/openai.md +19 -7
  53. package/docs/providers/opencode-go.md +3 -1
  54. package/docs/providers/openrouter.md +4 -3
  55. package/docs/providers/vertex.md +1 -0
  56. package/docs/public-contracts.md +2 -1
  57. package/docs/rag.md +55 -8
  58. package/docs/release-and-install.md +45 -7
  59. package/docs/server.md +1 -0
  60. package/docs/supervisors.md +3 -2
  61. package/docs/system-prompts.md +1 -1
  62. package/docs/tools.md +1 -1
  63. package/docs/web-tools.md +2 -0
  64. package/docs/wiki.md +140 -0
  65. package/docs/workflows.md +4 -3
  66. package/docs/working-and-semantic-memory.md +20 -0
  67. package/package.json +12 -5
  68. package/docs/api-page-template.md +0 -32
  69. package/docs/release-0.2.7-evidence.md +0 -514
package/docs/migration.md CHANGED
@@ -1,5 +1,15 @@
1
1
  # Migration guide
2
2
 
3
+ ## 0.3.0 → 0.3.1 production RAG engine (independent patch)
4
+
5
+ Only `@arnilo/prism-rag`, `@arnilo/prism-memory`, and `@arnilo/prism-observability-opentelemetry` move to `0.3.1`. Keep every other first-party package on `^0.3.0` — those ranges already satisfy `0.3.1`.
6
+
7
+ **Required for Embedder implementers.** `Embedder.id` is now a required `readonly id: string` (stable model/deploy identity, ≤256 chars). Hosts that construct their own embedder must set it; `createHashEmbedder` defaults to `"prism-hash-embedder"` and the Alibaba embedder uses `options.model`. `retrieveContext` fails closed with `ERR_PRISM_RAG_EMBEDDER_MISMATCH` when a stored `embedderId` is missing or differs (or dimensions differ). Re-index the source after an embedder/model change. Existing 0.3.0 rows without `embedderId` also fail closed until re-indexed.
8
+
9
+ Everything else is additive and opt-in: `createPostgresVectorStore`, hybrid `lexical` retrieve, `contentHash` skip, heading metadata, generation pointers, `createRagTelemetry`, `createTeiReranker`, multi-scope retrieve (`scopes` — pass `scope` for one corpus or `scopes` for one-or-many exact corpora; `scope` stays valid, both or neither throws). Default `retrieveContext` / `replaceSource` paths without those options stay 0.3.0-compatible (vector-only, single-scope, no skip, no telemetry).
10
+
11
+ Postgres DDL is additive (`IF NOT EXISTS` columns/indexes/tables). 0.3.0 rows remain readable. Rollback = restore the 0.3.0 package versions; no down migration.
12
+
3
13
  ## 0.2.9 → 0.3.0 lockstep cut and independent package versions (additive)
4
14
 
5
15
  Release **0.3.0** is the final lockstep cut on the 0.3.x line: all 57 publishable manifests move from `0.2.9` to `0.3.0`, then internal first-party `dependencies`, `optionalDependencies`, and `peerDependencies` use `^0.3.0`. The package graph is now **Decision B**: changed packages may patch/minor independently inside `>=0.3.0 <0.4.0`; unchanged packages keep their version.
@@ -72,7 +72,7 @@ Use `createMemorySessionStore()` for tests or throwaway sessions; use the JSONL
72
72
  - Reads and writes use only the caller-provided path.
73
73
  - Errors include path/reason or line number, not file contents.
74
74
  - Do not put secrets in messages, metadata, summaries, labels, or custom entries.
75
- - Reads are linear in file size. Appends also re-read and re-parse the whole file for duplicate/parent/corruption checks before writing one line, and are serialized per store instance.
75
+ - Reads are linear in file size. Appends also re-read and re-parse the whole file for duplicate/parent/corruption checks before writing one line, and are serialized per store instance. A rejected append does not poison later appends; the rejected line is not written.
76
76
  - There is no cross-process lock or durable idempotency table; two processes writing the same file can race. Add a database or external lock if multiple processes write the same file.
77
77
  - Treat this adapter as development/single-process storage. Production multi-writer hosts should use an indexed database `SessionStore` adapter.
78
78
 
@@ -0,0 +1,175 @@
1
+ # Obscura browser engine
2
+
3
+ Optional `@arnilo/prism-obscura` support for a host-installed
4
+ [Obscura](https://github.com/h4ckf0r0day/obscura) headless browser. Obscura is never
5
+ bundled — install the binary (or use the `h4ckf0r0day/obscura` Docker image) and point
6
+ the package at it.
7
+
8
+ - Capability evidence: `docs/_evidence/phase39-obscura-capability-matrix.md`
9
+ (pinned to Obscura `0.1.0` @ `f449e6f`).
10
+
11
+ ## Install
12
+
13
+ ```bash
14
+ npm install @arnilo/prism-obscura
15
+ ```
16
+
17
+ ## Process lifecycle (`spawnObscuraProcess`)
18
+
19
+ Runs a shell-free, bounded, ownership-tracked Obscura process. Configuration is
20
+ validated fail-closed: absolute NUL-free command, bounded argv/env, minimal default
21
+ environment (`PATH`, `HOME`), and insecure flags (`--allow-private-network`,
22
+ `--allow-file-access`, non-loopback `--host`) rejected unless `allowInsecureFlags`
23
+ is set explicitly.
24
+
25
+ ```ts
26
+ import { spawnObscuraProcess } from "@arnilo/prism-obscura";
27
+
28
+ const obscura = spawnObscuraProcess({
29
+ command: "/usr/local/bin/obscura",
30
+ args: ["serve", "--host", "127.0.0.1", "--port", "9222"],
31
+ });
32
+ await obscura.waitReady(() => canConnect("ws://127.0.0.1:9222"));
33
+ await obscura.close(); // SIGTERM → SIGKILL after grace; process-group kill on POSIX
34
+ ```
35
+
36
+ Docker works through the same argv seam — no Docker SDK:
37
+
38
+ ```ts
39
+ const mcp = spawnObscuraProcess({
40
+ command: "/usr/bin/docker",
41
+ args: ["run", "--rm", "-i", "h4ckf0r0day/obscura", "mcp"],
42
+ });
43
+ ```
44
+
45
+ Guarantees: argv byte-for-byte; capped stderr capture; errors never echo argv/env;
46
+ `close()` idempotent and kills only owned resources; abort signals kill owned
47
+ processes immediately.
48
+
49
+ ## MCP tools (`createObscuraMcpTools`)
50
+
51
+ Connects to `obscura mcp` (stdio or Streamable HTTP) through
52
+ [`@arnilo/prism-mcp`](mcp-tools.md) and exposes **every advertised tool** — no static
53
+ allow-list, so future Obscura tools keep flowing through.
54
+
55
+ ```ts
56
+ import { createObscuraMcpTools } from "@arnilo/prism-obscura";
57
+
58
+ const obscura = await createObscuraMcpTools({
59
+ transport: { type: "stdio", command: "/usr/local/bin/obscura", args: ["mcp"] },
60
+ });
61
+ agent.tools = [...agent.tools, ...obscura.tools];
62
+ await obscura.close();
63
+ ```
64
+
65
+ - Tool inventory at the pinned revision: 37 `browser_*` tools; render-enabled builds
66
+ add `browser_screenshot` and `browser_pdf`.
67
+ - **`browser_search` is in-page text search**, not public-web search.
68
+ - Effects: read/diagnostic/waiter/capture tools are effect-free; navigation,
69
+ interaction, evaluation, cookie/storage writes, tabs, and any unknown future tool
70
+ are exclusive, serialized external mutations (Obscura keeps one live page).
71
+ - Naming: default `obscura_` prefix coexists with `@arnilo/prism-browser`;
72
+ `namePrefix: ""` preserves native Obscura names.
73
+ - Transports: stdio configs are validated with the same fail-closed command policy;
74
+ Streamable HTTP endpoints outside loopback require explicit `allowRemoteHttp` and
75
+ remain subject to `@arnilo/prism-mcp` origin/transport security.
76
+
77
+ ## CDP and Playwright (`connectObscuraCdp`)
78
+
79
+ Attach to a running Obscura CDP endpoint — or spawn `obscura serve` and attach once
80
+ it is ready (bounded, abortable readiness; no fixed post-start sleep) — through the
81
+ host's Playwright via `chromium.connectOverCDP`. `connect()` and browser launch are
82
+ never used; Prism never launches browsers.
83
+
84
+ ```ts
85
+ import { connectObscuraCdp } from "@arnilo/prism-obscura";
86
+ import { createBrowserTools } from "@arnilo/prism-browser";
87
+
88
+ const session = await connectObscuraCdp({
89
+ command: "/usr/local/bin/obscura",
90
+ args: ["serve", "--host", "127.0.0.1", "--port", "9222"],
91
+ });
92
+ const tools = createBrowserTools({ browser: session.browser, networkPolicy });
93
+ // ... run ...
94
+ await session.close(); // browser first, then the owned process
95
+ ```
96
+
97
+ - External mode: pass `endpoint` (`ws://`/`http://`) with no `command` — the server
98
+ stays alive after `close()`; only resources this call created are terminated.
99
+ - Endpoints are loopback-only unless `allowRemoteEndpoint` is set; credentials in the
100
+ URL are always rejected; remote plain `ws:`/`http:` is refused (no authentication —
101
+ require an authenticated `wss:`/`https:` tunnel).
102
+ - The Playwright import is an optional exact `playwright-core@1.61.0` peer; supply
103
+ `connectObscuraCdp({ playwright })` to inject a host-selected build.
104
+ - The returned browser composes with `createBrowserManager`/`createBrowserTools`:
105
+ snapshots, actions, policy, checkpoints, and artifacts are Prism-owned. Raw CDP
106
+ (screenshots, PDF, screencast) stays available through Playwright's CDP session
107
+ APIs (`browser.newBrowserCDPSession()`, `context.newCDPSession(page)`); the package
108
+ adds no CDP command allow-list.
109
+ - Concurrency limit: pages served by one Obscura worker share one V8 isolate —
110
+ CPU-bound page JavaScript can delay sibling pages. Keep `@arnilo/prism-browser`
111
+ limits authoritative; size Obscura's `--workers` for the host.
112
+ - Screenshots/PDF require a render-enabled Obscura build and still obey the browser
113
+ package's artifact/byte policy.
114
+
115
+ ## Web search, fetch, and scrape (`createObscuraWebTools`)
116
+
117
+ Bounded CLI-backed web tools built on short-lived `obscura fetch`/`obscura scrape`
118
+ child processes. Returns standard Prism `web_search`/`web_fetch` tools plus explicit
119
+ `obscura_fetch`/`obscura_scrape` (disable with `nativeTools: false`).
120
+
121
+ ```ts
122
+ import { createObscuraWebTools } from "@arnilo/prism-obscura";
123
+
124
+ const web = createObscuraWebTools({ command: "/usr/local/bin/obscura" });
125
+ agent.tools = [...agent.tools, ...web.tools];
126
+ ```
127
+
128
+ - **Search truth**: `web_search` runs public-web search through one replaceable HTML
129
+ search profile (default: DuckDuckGo HTML endpoint). The extraction JavaScript is a
130
+ constant — the query travels only URL-encoded inside the search URL, never inside
131
+ evaluated source. Supply `searchProfile` to swap engines. Obscura's native
132
+ `browser_search` is in-page text search and is **not** exposed by this package.
133
+ - **web_search / web_fetch** return the same normalized untrusted shapes as
134
+ [`@arnilo/prism-web-tools`](web-tools.md) (`provider: "obscura"`, citations,
135
+ `untrusted: true`); content is labeled untrusted external content.
136
+ - **obscura_fetch**: one URL, bounded dump mode (`html|text|links|markdown|original`),
137
+ optional CSS selector (1-256 chars). No evaluation, screenshots, output paths, file
138
+ URLs, or private-network access.
139
+ - **obscura_scrape**: batch of public URLs (deduplicated, capped) with a constant
140
+ default expression; Obscura enforces `--concurrency` itself. Custom expressions
141
+ require explicit `allowEval: true` and stay byte-capped. Input association is
142
+ preserved by index; missing rows surface as `{ url, error }` entries.
143
+ - **Bounds**: query/result/batch/concurrency/output/eval byte caps, per-run timeout,
144
+ and child-process kill on timeout/abort (`DEFAULT_OBSCURA_WEB_LIMITS`,
145
+ `HARD_OBSCURA_WEB_LIMITS`). Malformed JSON, oversized output, and nonzero exits
146
+ fail closed with redacted diagnostics; nothing is retried.
147
+ - **URL policy**: every URL is validated as a public HTTP(S) target (private/
148
+ loopback/metadata and credentialed URLs denied) before any child process starts.
149
+ - Docker-style invocations work through `argsBefore` (e.g.
150
+ `["run", "--rm", "-i", "h4ckf0r0day/obscura"]`).
151
+ - An opt-in live smoke test runs against a real installed binary with
152
+ `npm run test:live -w @arnilo/prism-obscura` plus `PRISM_LIVE_OBSCURA=1` and
153
+ `PRISM_OBSCURA_BIN=/path/to/obscura`.
154
+
155
+ ## Host conformance (one generic integration)
156
+
157
+ `scripts/obscura-host-conformance.test.mjs` runs one shared fake-Obscura fixture —
158
+ the same `ToolDefinition[]` with one read tool (`web_fetch`) and one mutating tool
159
+ (`obscura_scrape`) — through every Prism host's public API: core agent/session
160
+ execution, the Prism MCP server, the `createPrismHandler` server lifecycle, AG-UI
161
+ MCP-tool injection, ACP fronting, workflow `toolNode`/`agentNode`s, supervisor
162
+ children, and Antigravity delegated MCP exposure. It verifies host authorization
163
+ and selection deny before execution, that no host needs an Obscura-specific branch,
164
+ and that an aborted in-flight call settles and kills the owned child. Composition
165
+ walkthrough: [`examples/obscura.ts`](../examples/obscura.ts).
166
+
167
+ ## Security
168
+
169
+ Obscura's CDP and MCP HTTP endpoints have no built-in authentication. This package
170
+ binds or connects to loopback by default. For remote deployments use an
171
+ authenticating reverse proxy or network isolation
172
+ (see [host security](host-security.md)).
173
+
174
+ CLI-backed web tools land in a subsequent release
175
+ (see `plans/039-Obscura-Full-Host-Support-And-Changed-Package-Release.md`).
@@ -4,13 +4,14 @@
4
4
 
5
5
  Prism exposes provider and tool timing through stable, metadata-only `AgentEvent` variants. Hosts subscribe via `session.subscribe()` or persist events through `RunLedger`. Core helpers build `ProviderTurnMetadata` and classify HTTP failures without echoing prompts, tool arguments, or credentials.
6
6
 
7
- Optional package `@arnilo/prism-observability-opentelemetry` maps those events to OpenTelemetry spans and low-cardinality metrics. OpenTelemetry is **not** a dependency of `@arnilo/prism`.
7
+ Optional package `@arnilo/prism-observability-opentelemetry` maps those events to OpenTelemetry spans and low-cardinality metrics, and adapts `@arnilo/prism-rag`'s dependency-free telemetry seam (`createRagTelemetry()`) onto the same tracer. OpenTelemetry is **not** a dependency of `@arnilo/prism`.
8
8
 
9
9
  APIs:
10
10
 
11
11
  - `ProviderTurnMetadata`, `ToolExecutionMetadata` on `AgentEvent`
12
12
  - `createProviderTurnMetadata()`, `readProviderHttpStatus()` in `@arnilo/prism`
13
13
  - `createOpenTelemetryInstrumentation()`, `wrapOpenTelemetryApi()`, `createInMemoryTelemetry()` in `@arnilo/prism-observability-opentelemetry`
14
+ - `createRagTelemetry()` in `@arnilo/prism-observability-opentelemetry` (RAG spans/events; see span tree below)
14
15
  - `handleRunFeedback()` / `handleEvaluation()` for explicit safe post-run projection
15
16
 
16
17
  ## When to use it
@@ -96,6 +97,21 @@ OpenTelemetry mapping (when enabled):
96
97
  | `handleRunFeedback` | active-run `prism.run.feedback` event or ended-run span | `prism.run.feedback` |
97
98
  | `handleEvaluation` | active-run `gen_ai.evaluation.result` event or ended-run span | `prism.run.evaluation` (`status`) |
98
99
 
100
+ RAG span tree (`@arnilo/prism-rag` + `createRagTelemetry()`):
101
+
102
+ | Span | Parent | Notes |
103
+ | --- | --- | --- |
104
+ | `rag_request` | host/chosen parent or root | One per `retrieveContext()`; carries `rag.top_k`, `rag.scope_count`, `rag.result_count`, `rag.index_generation` (single-scope only), scope/embedder id. `chunk_retrieved` events add `rag.chunk.tenant_id` + `rag.chunk.corpus_id`. |
105
+ | `embedding.query` | `rag_request` | Embedder call for the query |
106
+ | `retrieval.vector_search` / `retrieval.lexical` | `rag_request` | Present when the leg runs (lexical only when enabled and supported) |
107
+ | `retrieval.fusion` | `rag_request` | RRF fusion of the legs; `rag.fused_candidates` count |
108
+ | `retrieval.rerank` | `rag_request` | Only when a reranker is configured |
109
+ | `prompt.assembly` | `rag_request` | Context/template rendering |
110
+ | `rag_index` | host/chosen parent or root | One per `replaceSource()`/`indexChunks()`; `rag.chunk_count`, `rag.index_generation`, `rag.embedder_id`, `rag.source_id` |
111
+ | `embedding.index` | `rag_index` | Embedder batch call; skipped when unchanged-content skip fires |
112
+
113
+ `createRagTelemetry({ tracer, meter, attributeFilter? })` adapts the dependency-free `RagTelemetry` seam to a PrismTracer. Only the fixed span names and `rag.*`-shaped attribute keys pass through; anything else (including raw chunk text) is dropped before export. `attributeFilter` can further reduce or drop attributes.
114
+
99
115
  High-cardinality identifiers (`sessionId`, `runId`, `requestId`, `toolCallId`) are **span attributes only**, never metric labels.
100
116
 
101
117
  ## Request/response example
@@ -149,6 +165,10 @@ detach();
149
165
  telemetry.handleRunFeedback({ runId: result.runId, rating: 1, hasComment: true, tagCount: 1, scorerCount: 1, evaluationCount: 1 });
150
166
  telemetry.handleEvaluation({ runId: result.runId, name: "citation", status: "scored", score: 0.9, hasReason: true });
151
167
  console.log(traceId, memory.spans.map((span) => span.name));
168
+
169
+ // RAG: attach the same tracer to retrieveContext via the dependency-free seam
170
+ const ragTelemetry = createRagTelemetry({ tracer: memory.tracer, meter: memory.meter });
171
+ const found = await retrieveContext("policy", { embedder, store, scope, telemetry: ragTelemetry }); // rag_request tree
152
172
  ```
153
173
 
154
174
  ## Extension and configuration notes
@@ -28,6 +28,47 @@ node scripts/benchmark-0.1.0.mjs --out scripts/benchmark-0.1.0.json
28
28
  PRISM_TEST_POSTGRES_URL="postgresql://…" node scripts/benchmark-0.1.0.mjs --out scripts/benchmark-0.1.0.json # adds protected legs
29
29
  ```
30
30
 
31
+ ## Multi-agent runtime concurrency (phase 35)
32
+
33
+ `node scripts/benchmark.mjs --scenario multi-agent-runtime` is network-free (mock providers, in-process memory stores, no credentials). It measures concurrent independent sessions (1/4/16/32), supervisor fan-out and saturation (32 attempted delegates vs `maxActiveChildren`), parallel workflow fan-out maps (8×20 ms items at concurrency 2, ≥1.75× vs sequential), parallel workflow agent nodes, in-run tool concurrency, and an abort storm. Each result row carries p50/p95, throughput, heap delta, queued/dropped events, peak active provider calls, completions, and abort settle. Ceilings live in `scripts/budgets.json#multiAgentRuntime` (sanity bounds, machine-dependent). Exhaustive 59-manifest classification and recorded numbers: [`docs/_evidence/phase35-ai-runtime-package-matrix.md`](./_evidence/phase35-ai-runtime-package-matrix.md). Schema/safety/invariants: `scripts/benchmark-multi-agent.test.mjs`. Fan-out row: 8×20 ms items at concurrency 2, ≥1.75× vs sequential, peak workers ≤ 2. Supervisor saturation: 32 attempted delegates vs `maxActiveChildren` 4, overflow rejected, `activeAfter` 0.
34
+
35
+ ```bash
36
+ node scripts/benchmark.mjs --scenario multi-agent-runtime --out /tmp/prism-multi-agent.json
37
+ ```
38
+
39
+ Recorded 2026-08-27, Node v24.19.0 / Linux x64, 5 warmups + 20 waves, 8 ms mock delay. 32 independent sessions p95 10.1 ms (vs 9.0 ms at n=1); supervisor cap-4 fan-out p95 9.4 ms; workflow 4 agent nodes at concurrency 2 p95 17.9 ms; 8 tools at concurrency 4 p95 17.4 ms; abort storm settled in 5.5 ms with zero leftover provider calls. Dropped events: 0 on every row. Task 6 three-run median p95 (2026-08-28, same fixture): sessions 9.9/9.9/10.6/12.0, supervisorFanOut 10.4, supervisorSaturation 10.3, workflowFanOut 84.5 (1.87×, peak workers 2), workflowAgentNodes 19.7, toolConcurrency 19.1, abortStorm 3.3 — all under `scripts/budgets.json#multiAgentRuntime` ceilings. Protected PostgreSQL (`PRISM_TEST_POSTGRES_URL`) skipped on this host; `release:gate` blocked until durable evidence exists. Memory-store router 16/32-worker reservations do not oversubscribe.
40
+
41
+ ## Large-history and streamed-delta hot paths (plan 036)
42
+
43
+ The `multi-agent-runtime` scenario also covers 10,000 context-budget history rows and
44
+ 5,000 streamed provider deltas. `applyContextBudget` measures the keep-set once,
45
+ advances a history head cursor during eviction, and slices the retained suffix once;
46
+ it never front-mutates the history array. Runtime request/response limit accounting
47
+ uses `Buffer.byteLength(JSON.stringify(value), "utf8")`, so UTF-8 byte limits do not
48
+ allocate an encoded buffer per provider event.
49
+
50
+ Run with the existing network-free fixture:
51
+
52
+ ```bash
53
+ node scripts/benchmark.mjs --scenario multi-agent-runtime
54
+ ```
55
+
56
+ Recorded 2026-08-28 on Node v24.19.0 / Linux x64, 5 warmups + 20 measured waves.
57
+ `contextBudget-10k-history` completed with zero history remaining; its p50/p95 were
58
+ 2.481/3.708 ms and peak measured heap delta was 7,748,848 bytes. `provider-5k-deltas`
59
+ processed 5,000 deltas (320,015 serialized response bytes) at 4.253/5.847 ms p50/p95
60
+ with a 4,414,944-byte peak measured heap delta. These are local comparison evidence,
61
+ not portable SLOs; ceilings are in `scripts/budgets.json#multiAgentRuntime`.
62
+
63
+ `Buffer.byteLength` counts encoded bytes rather than JavaScript string length. A
64
+ serialized provider event at the exact response-byte cap succeeds; one byte below it
65
+ fails closed, including multibyte Unicode deltas. Context-budget omission order and
66
+ newest-history preservation remain covered by the root context-budget tests.
67
+
68
+ ## Current-line root artifact diet
69
+
70
+ `npm pack --dry-run --json` on `@arnilo/prism` is gated by `scripts/budget-gate.test.mjs` against `scripts/budgets.json#root` (±5%). Repository-only history stays out of the tarball: `docs/_evidence/**`, `docs/release-*-evidence.md`, `docs/api-page-template.md`, `dist/__tests__`, and `*.map`. Every page linked from shipped `docs/index.md` must be in the pack. Recorded 2026-08-27: **923,045 packed / 3,149,665 unpacked / 375 files** (226 `dist` js+d.ts, 124 index-linked docs, 25 other). 0.1.0 freeze 713,454 / 293 stays historical.
71
+
31
72
  ## 0.1.4 tree-shake measurement (static-reachability proxy)
32
73
 
33
74
  The 0.1.4 god-module split (agents/contracts → per-concern modules behind barrels) is
@@ -89,9 +130,10 @@ file count fail above baseline × 1.05. Labels: **network-free** = runs in
89
130
  | reconnectCatchup | 8.374 | 100 | distributed events (0.0.24) | protected |
90
131
 
91
132
  Install/startup rows (same helpers as the budget gate — no duplicate
92
- measurement): startup import 41.7 ms (ceiling 250 ms); root packed 711,755
93
- bytes vs baseline 678,541 (+5%, tolerance 5%); root file count 295 vs 293
94
- (+5%). Storage-growth rows and query plans from the protected legs are in the
133
+ measurement): startup import 41.7 ms (ceiling 250 ms). The recorded 0.1.0.json
134
+ pack rows (711,755 bytes / 295 files vs freeze 678,541 / 293) stay historical.
135
+ Live root pack is the current-line diet in `scripts/budgets.json#root` (see below).
136
+ Storage-growth rows and query plans from the protected legs are in the
95
137
  recorded JSON (`storageBeforeCleanup` / `storageAfterCleanup` per leg).
96
138
 
97
139
  Conformance companions: `scripts/phase8–11-conformance.test.mjs` plus the
@@ -375,7 +417,7 @@ On default overflow, the affected subscriber receives one `event_subscriber_over
375
417
  }
376
418
  ```
377
419
 
378
- `drop_oldest` keeps the newest queued events. `drop_newest` ignores incoming events while the queue is full. These policies are live-view policies only; they do not affect `RunLedger` writes or stored session entries.
420
+ `drop_oldest` keeps the newest queued events. `drop_newest` ignores incoming events while the queue is full. These policies are live-view policies only; they do not affect `RunLedger` writes or stored session entries. Graceful `createEventMultiplexer().close()` (and abort) stop new publishes/sources and drain already-queued events within `maxQueuedEvents` before the subscriber completes. Overflow `close` still drops the backlog, emits one overflow notice, and terminates.
379
421
 
380
422
  ## Request/response example
381
423
 
@@ -707,6 +749,18 @@ Enterprise governance and connector caps (defaults / hard). Timings: `node scrip
707
749
 
708
750
  Offline behavior tests (identity propagation, policy export, router deny paths, fake CLI argv) are release gates; live tenant canaries remain operator-gated.
709
751
 
752
+ ### 0.3.x Phase 39 Obscura browser-engine envelopes (2026-08-29)
753
+
754
+ `@arnilo/prism-obscura` binary-backed legs, network-free, driven by a deterministic fake CLI: `node scripts/benchmark-obscura.mjs` (3 runs, medians vs reviewed ceilings; artifact `scripts/benchmark-obscura.json`). Startup leg probes SIG-0 liveness after spawn — a real host waits on its readiness endpoint inside the same bound.
755
+
756
+ | Leg | Median (3 runs) | Ceiling | Notes |
757
+ | --- | --- | --- | --- |
758
+ | Managed startup (`spawnObscuraProcess` + `waitReady`) | ~0.02 ms | 250 ms | fake child; machine-dependent sanity bound, catches catastrophic lifecycle regression |
759
+ | Bounded CLI `web_search` call | ~20 ms | 100 ms | one `runObscuraCli` round trip through the public tool surface |
760
+ | Group close (SIGTERM drain) | ~0.6 ms | 250 ms | idempotent group-wide close; real children exit on signal |
761
+
762
+ No new release gate: the ceilings are evidence, not gates. Concurrent-resource evidence is behavioral, not timing: the MCP bridge serializes mutations (one live page), and abort tests prove an aborted in-flight call settles and kills the owned child with zero leaked processes (`scripts/obscura-host-conformance.test.mjs` abort leg; process/web suite timeout/abort-kill tests). Packed tarball 34.4 kB / 16 files; the package installs no binary, image, or browser.
763
+
710
764
  ## Related APIs
711
765
 
712
766
  - [Agent events](agent-events.md): `SubscribeOptions` and `event_subscriber_overflow` event details.
package/docs/ponytail.md CHANGED
@@ -106,7 +106,7 @@ See `examples/caveman-ponytail.ts` for combined Caveman + Ponytail progressive d
106
106
  - Import alone registers nothing (`sideEffects: false`); no timers, watchers, network, or shell scripts.
107
107
  - Upstream hook modules load via `createRequire` from resolved root — instruction strings are not forked in Prism.
108
108
  - Mode restore scans `getEntries()` for latest `data.type === "ponytail-mode"` (OM attach pattern).
109
- - `ponytail-subagent` hook is not wired; nested-agent behavior is host responsibility.
109
+ - `ponytail-subagent` hook is not wired; nested-agent behavior is host responsibility. When hosts wire the upstream hook, `PONYTAIL_SUBAGENT_MATCHER` accepts only the documented safe subset — `"explore|general"` (any literal substring) or `"^general$"` (exact), case-insensitive, max 256 chars. No `RegExp` is compiled from the environment, so arbitrary regex (including catastrophic nested quantifiers) is never evaluated; unset/invalid patterns inject into every subagent.
110
110
  - No TUI statusline scripts; use `ponytail status` command or extension events.
111
111
  - Not included in `@arnilo/prism-code` or `@arnilo/prism-sdk` profiles — opt-in install only.
112
112
 
@@ -8,7 +8,7 @@ Provider caching documents Prism's cache intent surface:
8
8
  - Legacy aliases `cacheKey` and `cacheRetention`, still supported for backwards compatibility.
9
9
  - `PromptCacheBreakpoint` locations for reusable prompt regions.
10
10
  - `ModelCacheCapabilities` for model/provider cache support metadata.
11
- - Shared helpers: `sanitizeCacheKey`, `mapCacheRetention`, `applyCacheControl`, `cacheHitRate`, `cacheSavings`, and `cacheUsageReport`.
11
+ - Shared helpers: `sanitizeCacheKey`, `mapCacheRetention`, `applyCacheControl`, `resolveBreakpoint`, `canonicalizeJsonSchema`, `cacheHitRate`, `cacheSavings`, and `cacheUsageReport`.
12
12
 
13
13
  Cache hints are best-effort. They describe intent; providers decide whether their native API can use them. Prism does not guarantee cache hits.
14
14
 
@@ -17,7 +17,7 @@ Cache hints are best-effort. They describe intent; providers decide whether thei
17
17
  Use this page when a host or provider package needs to:
18
18
 
19
19
  - Mark stable system prompts, tools, context, or messages as cacheable.
20
- - Opt into cache-aware default input ordering so stable attachments/resources, summaries, and prior history form a reusable prefix before the current user turn.
20
+ - Use cache-aware default input ordering so stable instructions, attachments/resources, summaries, and prior history form a reusable prefix before the current user turn.
21
21
  - Carry a stable cache key across turns without putting provider-specific fields in core.
22
22
  - Read `ModelConfig.cache` to decide whether to map hints to implicit caching, key-based caching, cache-control breakpoints, provider-specific caching, or no caching.
23
23
  - Compute normalized cache diagnostics from `Usage.cacheReadTokens` / `Usage.cacheWriteTokens`, including providers that only report reads.
@@ -60,13 +60,15 @@ Cache helpers return plain data:
60
60
  | `sanitizeCacheKey(value, maxLength)` | Safe key string or `undefined`. |
61
61
  | `mapCacheRetention(retention, model)` | `"short"`, `"long"`, or `undefined`. |
62
62
  | `applyCacheControl(messages, breakpoints, options)` | New message array with `cache_control: { type: "ephemeral" }` on selected message anchors. |
63
+ | `resolveBreakpoint(messages, breakpoint)` | Message index for a `PromptCacheBreakpoint` (`-1` when unresolved); shared anchor selection for `applyCacheControl` and OpenAI explicit breakpoints. |
64
+ | `canonicalizeJsonSchema(value)` | Clone with sorted object keys and `required` names; semantic arrays stay ordered. Used by first-party tool serializers. |
63
65
  | `cacheHitRate(usage)` | Cached input ratio or `undefined`. |
64
66
  | `cacheSavings(usage, model)` | Estimated read-token savings or `undefined` without pricing. |
65
67
  | `cacheUsageReport(usage, model?)` | Normalized read/write tokens, hit rate, estimated savings, and currency when available; `undefined` when no usage is supplied. |
66
68
 
67
69
  Provider events do not change. Cache accounting stays in normalized `Usage.cacheReadTokens` and `Usage.cacheWriteTokens`.
68
70
 
69
- For stable-prefix payloads, `inputLayout: "cache_aware"` is the default on the default input builder, `assembleProviderInput()`, `AgentConfig`, and `RunOptions`; set `inputLayout: "legacy"` to restore the prior order. The default prompt builder already places context, selected skills, and tool declarations before input messages; cache-aware input ordering then places attachments/resources, summaries, prior history, and pending tool results before the current user suffix. The prefix is byte-stable only when those stable inputs are unchanged; Prism still does not guarantee provider cache hits.
71
+ For stable-prefix payloads, `inputLayout: "cache_aware"` is the default on the default input builder, `assembleProviderInput()`, `AgentConfig`, and `RunOptions`; set `inputLayout: "legacy"` to restore the prior order. The default prompt builder's cache-aware order is leading system instructions → resolved context blocks → selected/progressively disclosed skills → fallback text tool declarations → attachments/resources → summaries → prior history → pending tool results → current input. Declared tool schemas remain in `ProviderRequest.tools` and are never granted by prompt middleware. First-party tool serializers run `canonicalizeJsonSchema` so property insertion order cannot break that prefix. Changing only current input preserves the serialized message prefix before the final user suffix; changing dynamic context or loaded skills changes only from its own boundary onward, while tool schemas remain independently stable. The prefix is byte-stable only when those stable inputs are unchanged; Prism still does not guarantee provider cache hits.
70
72
 
71
73
  ## Request/response example
72
74
 
@@ -135,7 +137,7 @@ Provider request policies can set `ProviderRequestOptions.cache` or the legacy `
135
137
 
136
138
  | `ModelCacheCapabilities.kind` | Typical mapping |
137
139
  | --- | --- |
138
- | `implicit` | No request mutation; provider caches automatically. |
140
+ | `implicit` | No request mutation; provider caches automatically. Conformance (`assertNoForeignCacheFields`) proves the serialized body carries no cache wire fields. |
139
141
  | `openai_key` | Send sanitized cache key and mapped retention where supported. |
140
142
  | `cache_control` | Use `applyCacheControl()` on provider-native message anchors. |
141
143
  | `provider_specific` | Provider package uses `compat`/native options intentionally. |
@@ -145,8 +147,8 @@ Provider request policies can set `ProviderRequestOptions.cache` or the legacy `
145
147
 
146
148
  | Provider package | Cache kind | Explicit cache hints | Multi-turn reuse notes | Caveats |
147
149
  | --- | --- | --- | --- | --- |
148
- | `@arnilo/prism-provider-openai` | `openai_key` | Sends sanitized `prompt_cache_key`; `prompt_cache_retention: "24h"` only when the model declares `longRetention`. | Stable cache key + stable prefix can improve reuse. | Best-effort only; `"short"`/`"none"` omit retention. |
149
- | `@arnilo/prism-provider-anthropic` | `cache_control` | Marks only selected Anthropic message anchors; `"long"` maps to documented `ttl: "1h"`. | Keep selected anchors stable. | Best-effort; never stamp every block. |
150
+ | `@arnilo/prism-provider-openai` | `openai_key` | Sends sanitized `prompt_cache_key`; pre-5.6 models emit `prompt_cache_retention: "24h"` when `longRetention`; GPT-5.6+ models (`explicitBreakpoints`) map `cache.breakpoints`/`cache.mode: "on"` to `prompt_cache_options: { mode: "explicit" }` + `prompt_cache_breakpoint` markers (≤4 writes). | Stable cache key + stable prefix can improve reuse; keep selected anchors stable. | Best-effort only; `"short"`/`"none"` omit retention; `"30m"` TTL is the default and never emitted. |
151
+ | `@arnilo/prism-provider-anthropic` | `cache_control` | Marks only selected Anthropic message anchors; `system_prompt` breakpoints emit native `system` text blocks with the marker; `"long"` maps to documented `ttl: "1h"`. | Keep selected anchors stable. | Best-effort; never stamp every block. |
150
152
  | `@arnilo/prism-provider-google` | none | Sends no Prism cache marker. | Host/model may have upstream behavior. | Gemini cache controls are not mapped in this package. |
151
153
  | `@arnilo/prism-provider-openrouter` | `cache_control` | Top-level automatic `cache_control` when enabled without breakpoints; otherwise markers only on caller-selected `cache.breakpoints`; `"long"` may add `ttl: "1h"`. Sticky `session_id` routing. | Breakpoint-stable / automatic prefixes can be reused by upstream providers. | Best-effort only; top-level automatic may exclude some backends from routing. |
152
154
  | `@arnilo/prism-provider-opencode-go` | route-specific | Sends sanitized `x-opencode-session`; Anthropic route applies selected `cache_control` breakpoints; OpenAI route sends none. | Session id + unchanged selected anchors can help route-native caches. | Best-effort and route-dependent. |
@@ -156,7 +158,7 @@ Provider request policies can set `ProviderRequestOptions.cache` or the legacy `
156
158
  | `@arnilo/prism-provider-ai-sdk` | host-owned | No Prism cache payload; host `LanguageModelV4` owns upstream caching. | Host model/provider decides cache keys, breakpoints, and sticky routing. | Adapter maps `inputTokens.cacheRead`/`cacheWrite` from `finish.usage` only; does not invent cache fields. |
157
159
  | `@arnilo/prism-provider-alibaba` | implicit by default, optional `cache_control` | DashScope implicit prefix caching is automatic; opt-in `cache_control: {"type":"ephemeral"}` markers only on caller-selected `cache.breakpoints`, capped at 4. | Keep selected anchors and prior history stable; each cached prefix needs ≥1024 tokens and lives ~5 minutes upstream. | Best-effort and model-dependent; `cached_tokens`→read, `cache_creation_input_tokens`→write. |
158
160
  | `@arnilo/prism-provider-ollama` | `implicit` | No `cache_control`, `cacheKey`, `prompt_cache`, or `cacheRetention` payload; Ollama KV/prefix caching is automatic with no request knob. | Resend unchanged prior history for implicit KV reuse. | Best-effort only; Ollama reports no cached-token count, so `Usage.cacheReadTokens` stays `undefined`. |
159
- | `@arnilo/prism-provider-deepseek` | `implicit` | No `cache_control` / `prompt_cache_key`; tools schemas are key-sorted so the prefix stays byte-stable. | Resend unchanged history from token 0; append only the new turn. Thinking-on strips temperature/top_p/penalties so they cannot break the prefix. | Best-effort prefix units (~1024 practical min). `prompt_cache_hit_tokens` → `cacheReadTokens`. |
161
+ | `@arnilo/prism-provider-deepseek` | `implicit` | No `cache_control` / `prompt_cache_key`; tool `parameters` go through shared `canonicalizeJsonSchema`. | Resend unchanged history from token 0; append only the new turn. Thinking-on strips temperature/top_p/penalties so they cannot break the prefix. | Best-effort prefix units (~1024 practical min). `prompt_cache_hit_tokens` → `cacheReadTokens`. |
160
162
  | `@arnilo/prism-provider-xai` | `implicit` | No `prompt_cache_key`. Package-local `x-grok-conv-id` is `sanitizeCacheKey(cache.key ?? cacheKey ?? sessionId, 128)`. | Same server + unchanged message prefix. Replay `reasoning_content` on reasoning models or the prefix breaks. | Conv-id is never a credential or SuperGrok token. Omitted when `cache.mode` is `off` or `cacheRetention` is `none`. `cached_tokens` → `cacheReadTokens` (inclusive or exclusive reports kept as-is). |
161
163
  | `@arnilo/prism-provider-clinepass` | `implicit` | No `cache_control` / `prompt_cache_key`. Gateway-owned prefix cache. | Resend unchanged prior history. Stream only. | Best-effort and backend-dependent (`cline-pass/*` slugs). `cached_tokens` / `prompt_cache_hit_tokens` map when present. |
162
164
  | `@arnilo/prism-provider-azure` | none | No Prism cache mapping. | Endpoint/model-specific. | Host owns Azure cache policy. |
@@ -165,11 +167,11 @@ Provider request policies can set `ProviderRequestOptions.cache` or the legacy `
165
167
 
166
168
  Detailed first-party provider notes:
167
169
 
168
- - OpenAI Responses (`@arnilo/prism-provider-openai`): `kind: "openai_key"`. Sanitizes/clamps `prompt_cache_key` to 64 chars; `"long"` retention maps to `prompt_cache_retention: "24h"` only when the model declares `cache.longRetention`; `"short"`/`"none"` omit the field. GPT-5.6+ official docs use `prompt_cache_options` / breakpoints instead of retention — `listOpenAIModels` sets `longRetention: false` for those ids. `input_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
170
+ - OpenAI Responses (`@arnilo/prism-provider-openai`): `kind: "openai_key"`. Sanitizes/clamps `prompt_cache_key` to 64 chars; pre-GPT-5.6 models (`cache.longRetention: true`) map `"long"` retention to `prompt_cache_retention: "24h"`; GPT-5.6+ models (`cache.explicitBreakpoints: true`) map `cache.breakpoints`/`cache.mode: "on"` to `prompt_cache_options: { mode: "explicit" }` plus `prompt_cache_breakpoint: { mode: "explicit" }` markers on selected message anchors (≤4 writes; the only TTL `"30m"` is the default, so none is emitted). Resolved cache fields win over caller `extra`. `input_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`Usage.cacheWriteTokens`.
169
171
  - OpenAI-compatible Chat Completions adapter: minimal scope, sends no cache payload; see [OpenAI-compatible provider](providers/openai-compatible.md).
170
- - Anthropic (`@arnilo/prism-provider-anthropic`): `kind: "cache_control"`; selected Anthropic message anchors receive `cache_control` and eligible long retention maps to `ttl: "1h"`. Cache read/create usage maps to normalized cache read/write tokens.
172
+ - Anthropic (`@arnilo/prism-provider-anthropic`): `kind: "cache_control"`; selected Anthropic message anchors receive `cache_control` and eligible long retention maps to `ttl: "1h"`. A `system_prompt` breakpoint serializes `system` as native text blocks carrying the marker (shared `systemCacheControlField()` helper; plain joined string when unmarked). Cache read/create usage maps to normalized cache read/write tokens.
171
173
  - Google (`@arnilo/prism-provider-google`): sends no Prism cache-control payload. Do not infer cache hits or cache token counts from absent Gemini fields.
172
- - OpenRouter (`@arnilo/prism-provider-openrouter`): `kind: "cache_control"`. Sanitizes/clamps `session_id`/`X-Session-Id` to 256 chars for sticky routing; with no breakpoints emits top-level automatic `cache_control: { type: "ephemeral" }`; with breakpoints applies Anthropic-style markers only to caller-selected locations (last content block of each selected message); `"long"` retention adds `ttl: "1h"` when the model allows it. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`. Optional `listOpenRouterModels()` may populate `ModelConfig.cache`/`cost` from live pricing.
174
+ - OpenRouter (`@arnilo/prism-provider-openrouter`): `kind: "cache_control"`. Sanitizes/clamps `session_id`/`X-Session-Id` to 256 chars for sticky routing (from `cache.key` ?? legacy `cacheKey` ?? `sessionId`); with no breakpoints emits top-level automatic `cache_control: { type: "ephemeral" }`; with breakpoints applies Anthropic-style markers only to caller-selected locations (last content block of each selected message); `"long"` retention adds `ttl: "1h"` when the model allows it. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`. Optional `listOpenRouterModels()` may populate `ModelConfig.cache`/`cost` from live pricing.
173
175
  - OpenCode Go (`@arnilo/prism-provider-opencode-go`): default base `https://opencode.ai/zen/go/v1`; `x-opencode-session` from `cacheKey ?? sessionId`, sanitized to 128 chars; the Anthropic route (MiniMax/Qwen) applies `cache_control` markers only to selected breakpoints (`"long"` → `ttl: "1h"`), the OpenAI route (Grok/GLM/Kimi/MiMo/DeepSeek) sends none and preserves `reasoning_content`. OpenAI route maps `prompt_tokens_details.cached_tokens`/`cache_write_tokens`; Anthropic route maps `cache_read_input_tokens`/`cache_creation_input_tokens`. Caller-gated `listOpenCodeGoModels` against official `GET /zen/go/v1/models`.
174
176
  - Z.AI (`@arnilo/prism-provider-zai`): `kind: "implicit"`. GLM context caching is automatic; sends no explicit cache payload regardless of cache options. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`.
175
177
  - NeuralWatt (`@arnilo/prism-provider-neuralwatt`): `kind: "implicit"`. NeuralWatt prefix caching is automatic; sends no explicit cache payload regardless of cache options. `cacheRetention: "none"` disables Prism cache-control hints only (not the implicit backend prefix cache). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`; NeuralWatt does not report a cache-write token, so `Usage.cacheWriteTokens` is never fabricated (stays `undefined`). NeuralWatt's `/v1/models` catalog advertises exact `cached_input_per_million` rates for cache reads and `cached_output_per_million: null`; static curated aliases do not guess those prices.
@@ -177,7 +179,7 @@ Detailed first-party provider notes:
177
179
  - AI SDK adapter (`@arnilo/prism-provider-ai-sdk`): **host-owned**. Sends no Prism cache payload; the supplied `LanguageModelV4` and its upstream provider own request caching. Maps AI SDK v4 `finish.usage.inputTokens.cacheRead`/`cacheWrite` to `Usage.cacheReadTokens`/`cacheWriteTokens`. No `list*Models()` export.
178
180
  - Alibaba Cloud (`@arnilo/prism-provider-alibaba`): implicit by default, optional `cache_control`. DashScope implicit prefix caching is automatic (no marker); explicit opt-in `cache_control: {"type":"ephemeral"}` markers apply only to selected breakpoints when `ModelConfig.cache.kind: "cache_control"` and the caller supplies breakpoints, capped at 4 (each prefix ≥1024 tokens, ~5 minute TTL). `prompt_tokens_details.cached_tokens`/`cache_creation_input_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`. Caller-gated `listAlibabaModels` against OpenAI-compatible `GET {base}/models`.
179
181
  - Ollama (`@arnilo/prism-provider-ollama`): `kind: "implicit"`. Ollama reuses its KV/prompt cache automatically; there is no request knob and no wire marker, so Prism never emits `cache_control`. Ollama reports no cached-token count, so `Usage.cacheReadTokens` is intentionally left `undefined` (not `0`). Caller-gated `listOllamaModels` against OpenAI-compatible `GET {base}/models`.
180
- - DeepSeek (`@arnilo/prism-provider-deepseek`): `kind: "implicit"`. Official disk prefix cache is automatic (byte-identical prefix from token 0). Adapter sends no cache payload; tool `parameters` are canonicalized for stable JSON key order. `prompt_cache_hit_tokens` maps to `Usage.cacheReadTokens`. Caller-gated `listDeepSeekModels`.
182
+ - DeepSeek (`@arnilo/prism-provider-deepseek`): `kind: "implicit"`. Official disk prefix cache is automatic (byte-identical prefix from token 0). Adapter sends no cache payload; tool `parameters` use shared `canonicalizeJsonSchema` (object keys + unordered `required` only; `enum`/`prefixItems`/`examples` keep caller order). `prompt_cache_hit_tokens` maps to `Usage.cacheReadTokens`. Caller-gated `listDeepSeekModels`.
181
183
  - xAI (`@arnilo/prism-provider-xai`): `kind: "implicit"`. Automatic prefix cache. Sticky `x-grok-conv-id` is a sanitized session/cache key (128 chars), never an OAuth access token. Reasoning models must replay `reasoning_content`. `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`. Caller-gated `listXaiModels`.
182
184
  - ClinePass (`@arnilo/prism-provider-clinepass`): `kind: "implicit"`. No explicit cache payload; multi-backend gateway may report `cached_tokens` or `prompt_cache_hit_tokens`. Static `cline-pass/*` catalog only — no `listClinePassModels`.
183
185
  - Azure, Bedrock, and Vertex: their OpenAI-compatible packages intentionally emit no Prism cache fields. Endpoint/model-specific cache controls remain host-owned rather than guessed from another provider family.
@@ -12,6 +12,9 @@ Exported from `@arnilo/prism/testing/provider-conformance`:
12
12
  - `assertToolCallDeltasReconstruct(events, expected)`
13
13
  - `assertUsageAccounting(events, expected)`
14
14
  - `assertSerializedRequestCoversContent(request, body, options?)`
15
+ - `assertCanonicalToolParameters(serialized, original)`
16
+ - `assertNoForeignCacheFields(body, allowed?)`
17
+ - `assertNoFetches(calls)`
15
18
  - `assertProviderOwnedHeadersWin(captured, options)`
16
19
  - `assertNoSecretLeak(events, secrets)`
17
20
 
@@ -68,6 +71,9 @@ Helpers accept normal `AIProvider`, `ProviderRequest`, `ProviderEvent`, `Usage`,
68
71
  - `assertAbortIsObserved()` passes an already-aborted signal and expects provider generation to reject. This is the supported timeout primitive; use a host abort controller or `RunOptions.signal`.
69
72
  - `assertToolCallDeltasReconstruct()` rebuilds streamed `tool_call_delta` fragments into tool calls and validates expected id/name/arguments. Malformed JSON with id+name present yields `argumentsError` (no throw); missing id/name throws typed `incomplete_delta`. The runtime uses the same reconstruction before tool execution when a provider streams deltas.
70
73
  - `assertUsageAccounting()` finds `usage` or `done.usage` and checks selected token fields including `cacheReadTokens` and `cacheWriteTokens`. This is the provider-neutral check for normalized cache read/write token extraction; every first-party provider package exercises it against server-specific fields (`cached_tokens`, `cache_read_input_tokens`, etc.).
74
+ - `assertCanonicalToolParameters()` checks a serialized tool schema matches `canonicalizeJsonSchema(original)` so property insertion order and `required` name order cannot drift on the wire while `enum`/`prefixItems`/`examples` stay caller-ordered.
75
+ - `assertNoForeignCacheFields()` fails when a request body carries a cache wire field the route does not document (`cache_control`, `prompt_cache_*`, `cachedContent`, `cachePoint`); pass documented fields in `allowed` for routes that opt in (Alibaba/OpenRouter markers, Gemini `extra.cachedContent`). Implicit-cache providers must serialize no foreign cache fields — implicit caching works by byte-stable prefix reuse, not request payloads.
76
+ - `assertNoFetches()` fails when a provider performed network calls outside caller-gated discovery/stream; provider construction and `setup()` must be network-free.
71
77
  - `assertSerializedRequestCoversContent()` scans a serialized provider request body for primitive canaries from each Prism content block and fails if any supported block type is silently dropped. Provider-valid transcripts place assistant `tool_call` messages before matching role `tool` `tool_result` messages; runtime, cache-aware input layout, and observational-memory worker loops preserve that order before serialization.
72
78
  - `assertProviderOwnedHeadersWin()` compares captured request headers against the provider's authoritative owned header values and a caller-supplied header bag; it fails if any owned header (`authorization`, `content-type`, session/security headers) was overridden by caller headers, and also fails if a non-owned caller header was dropped. This is the provider-neutral check that caller `ProviderRequest.options.headers` cannot hijack provider credentials or sessions; every first-party provider package exercises it.
73
79
  - `assertNoSecretLeak()` stringifies all collected events and fails if any known secret string is present.
@@ -115,7 +115,7 @@ Every package remains explicit, setup-zero-fetch, and late-credential-bound. `Mo
115
115
 
116
116
  Every first-party provider package hardens prompt-cache behavior so it cannot emit invalid cache retention values or over-broad cache-control markers, and so provider-owned `authorization`/session/security headers cannot be overridden by caller `ProviderRequest.options.headers`. Cache behavior is provider-specific and best-effort: OpenAI/OpenRouter use explicit hints, NeuralWatt/Z.AI/DeepSeek/ClinePass use implicit caching, xAI adds a sanitized `x-grok-conv-id`, and OpenCode Go/Kimi are route/model-dependent. See [Provider caching](provider-caching.md#per-provider-cache-behavior) for the canonical explicit/implicit matrix.
117
117
 
118
- - **OpenAI** (`kind: openai_key`): `prompt_cache_key` is sanitized and clamped to 64 chars; `prompt_cache_retention` is emitted as `24h` only when the model declares `cache.longRetention`, and omitted for `short`/`none` (the API only accepts absent or `24h`). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
118
+ - **OpenAI** (`kind: openai_key`): `prompt_cache_key` is sanitized and clamped to 64 chars. Pre-GPT-5.6 models emit `prompt_cache_retention: "24h"` when the model declares `cache.longRetention`; GPT-5.6+ models (`cache.explicitBreakpoints`) map `cache.breakpoints` / `cache.mode: "on"` to `prompt_cache_options: { mode: "explicit" }` with `prompt_cache_breakpoint` markers on selected anchors. `input_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`Usage.cacheWriteTokens`.
119
119
  - **OpenAI-compatible core adapter**: Chat Completions sends no `prompt_cache_key`/`prompt_cache_retention`/`cache_control` fields; endpoints cache implicitly. `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
120
120
  - **OpenRouter** (`kind: cache_control`): `session_id`/`x-session-id` sanitized and clamped to 256 chars; with no breakpoints, emits top-level automatic `cache_control: { type: ephemeral }`; with breakpoints, markers applied only to caller-selected locations (not every block); `cacheRetention: long` adds `ttl: 1h` when allowed. Preserves `cached_tokens`/`cache_write_tokens`.
121
121
  - **OpenCode Go**: default base `https://opencode.ai/zen/go/v1`; `x-opencode-session` from `cacheKey ?? sessionId` sanitized to 128 chars; Anthropic route (MiniMax/Qwen) applies `cache_control` only to selected breakpoints (`long` → `ttl: 1h`); OpenAI route (Grok/GLM/Kimi/MiMo/DeepSeek) sends none and preserves `reasoning_content`. Per-route usage mapping. Caller-gated `listOpenCodeGoModels`.
@@ -8,7 +8,7 @@ Implementation is **shipped** for transport and OpenAI serialization primitives
8
8
 
9
9
  ## When to use it
10
10
 
11
- - **Provider package authors** implementing or migrating a first-party adapter should import shared primitives from `@arnilo/prism/providers/transport` and `@arnilo/prism/providers/openai` instead of copying `sse.ts`, `safeText`, `parseArgs`, or serializers.
11
+ - **Provider package authors** implementing or migrating a first-party adapter should import shared primitives from `@arnilo/prism/providers/transport`, `@arnilo/prism/providers/openai`, and `@arnilo/prism/providers/schema` instead of copying `sse.ts`, `safeText`, `parseArgs`, serializers, or JSON Schema key-sorting.
12
12
  - **Host apps** choose native structured output via `ProviderRequestOptions.structuredOutput` when the model declares support; otherwise they keep the artifact generate→validate→revise loop ([Structured output](structured-output.md)).
13
13
  - **Operators** enable observability through extended agent events and the optional OpenTelemetry adapter package ([Observability](observability.md)).
14
14
 
@@ -190,6 +190,19 @@ export function assertOpenAIChatMessage(message: unknown, path: string): asserts
190
190
 
191
191
  `src/providers/openai-compatible.ts` becomes a thin adapter over these helpers in Task 2.
192
192
 
193
+ ### `@arnilo/prism/providers/schema` — **shipped**
194
+
195
+ Deterministic JSON Schema clone for tool/function parameters. Sorts object keys and unordered `required` names. Leaves semantic arrays (`prefixItems`, `examples`, `enum`, tuple `items`) in caller order. Does not resolve `$ref`, mutate input, or enforce schema bounds.
196
+
197
+ ```ts
198
+ import { canonicalizeJsonSchema } from "@arnilo/prism/providers/schema";
199
+
200
+ canonicalizeJsonSchema({ required: ["b", "a"], properties: { b: {}, a: {} } });
201
+ // keys and required names stable; ordered schema arrays remain ordered
202
+ ```
203
+
204
+ `serializeOpenAITool` and first-party native `toTool` mappers reuse this helper so logically identical schemas stringify identically.
205
+
193
206
  ### Structured output capability (Task 4 — **shipped**)
194
207
 
195
208
  ```ts
@@ -286,7 +299,7 @@ Every migrated provider must pass this shared matrix (implemented in Task 1 test
286
299
  ## Related APIs
287
300
 
288
301
  - [Provider layer](provider-layer.md): registry, mock provider, event helpers
289
- - [Provider conformance](provider-conformance.md): stream order, abort, header ownership
302
+ - [Provider conformance](provider-conformance.md): stream order, abort, header ownership, canonical tool schemas
290
303
  - [OpenAI-compatible provider](providers/openai-compatible.md): reference adapter subpath
291
304
  - [Structured output](structured-output.md): artifact loop fallback
292
305
  - [Provider request policies](provider-request-policies.md): cache and request hooks
@@ -112,7 +112,7 @@ console.log(result.text);
112
112
 
113
113
  There is **no Prism-side model catalog** and **no `list*Models()` export** by design. Hosts supply a ready-made `LanguageModelV4` instance (typically from `@ai-sdk/openai`, `@ai-sdk/anthropic`, AI Gateway, or a custom provider) and register a matching `ModelConfig` for capabilities/limits.
114
114
 
115
- Prism setup remains network-free: `createAiSdkProvider` only wraps the supplied model and never fetches catalogs or credentials.
115
+ Prism setup remains network-free: `createAiSdkProvider` only wraps the supplied model and never fetches catalogs or credentials. Stream and abort behavior are conformance-proven: an already-aborted signal fails fast as an `error` event, a host stream ending without a `finish` part fails loudly (typed `AiSdkProviderError { code: "model_error" }`) instead of synthesizing a `done`, and unmappable stream parts fail closed.
116
116
 
117
117
  ## Prompt caching
118
118
 
@@ -41,7 +41,7 @@ Featured offline aliases: `claude-opus-4-8`, `claude-sonnet-5`, `claude-haiku-4-
41
41
  | Surface | Behavior |
42
42
  | --- | --- |
43
43
  | Stream | Prism text, thinking deltas, tool-call delta/final, usage (incl. cache read/create when present), `done`, redacted `error`. |
44
- | Cache | Featured models use `cache.kind: "cache_control"`; markers on selected breakpoints (`long` → `ttl: "1h"`). |
44
+ | Cache | Featured models use `cache.kind: "cache_control"`; markers on selected breakpoints (`long` → `ttl: "1h"`). A `system_prompt` breakpoint serializes `system` as native text blocks with the marker (plain string otherwise). |
45
45
  | Thinking | Model-family aware (`adaptive` vs `enabled`+`budget_tokens`); helpers `anthropicThinking` / `anthropicEffort` / `anthropicPreserveThinking`. |
46
46
  | Auth | `api_key` for provider id; provider-owned `content-type`, `x-api-key`, `anthropic-version` win over caller headers. No OAuth descriptor or subscription adapter is registered. |
47
47
 
@@ -64,6 +64,7 @@ Register via `createExtensionKernel().load([createAzureOpenAIProviderPackage(...
64
64
  - Endpoint host is never rewritten to public DNS.
65
65
  - Errors redact credential values via shared transport helpers.
66
66
  - No Azure SDK dependency.
67
+ - Conformance-proven (Task 6): package `setup()` performs zero fetch and zero credential resolution; an already-aborted signal fails fast; a truncated SSE stream (no `data: [DONE]`) ends in an `error` event; Azure cache policy stays host-owned, so no cache wire fields (`cache_control`, `prompt_cache_*`) are emitted even when the request carries Prism cache hints — only upstream-reported `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
67
68
 
68
69
  ## Related APIs
69
70
 
@@ -62,6 +62,7 @@ Uses Bedrock’s OpenAI-compatible runtime route (not Converse eventstream). Hos
62
62
  - No AWS SDK; package-local SigV4 only for `bedrock` service.
63
63
  - Input headers are normalized once before signing: names are lowercased and duplicate-case keys merge last-wins, so the canonical request always matches the signed header list (no duplicate-case mismatch); query parameters are canonicalized sorted by encoded key then value.
64
64
  - Private endpoint hosts are not rewritten to public DNS.
65
+ - Conformance-proven (Task 6): package `setup()` performs zero fetch and zero credential resolution; an already-aborted signal fails fast; a truncated SSE stream (no `data: [DONE]`) ends in an `error` event; native Bedrock caching (`Converse cachePoint`) is intentionally unsupported on the OpenAI-compatible route — no cache wire fields are emitted even when the request carries Prism cache hints.
65
66
  - Credential secrets are redacted from provider errors.
66
67
  - No credential prefetch at import.
67
68
 
@@ -179,7 +179,8 @@ await kernel.load([
179
179
  - When opted in, `cache_control: { type: "ephemeral" }` markers apply only to
180
180
  caller-selected `ProviderRequestOptions.cache.breakpoints` on the last content
181
181
  block of each selected message. `cacheRetention: "long"` adds `ttl: "1h"` when
182
- the model allows long retention.
182
+ the model allows long retention. A `system_prompt` breakpoint serializes
183
+ `system` as native text blocks with the marker (plain string otherwise).
183
184
  - The Moonshot Open Platform route never receives Anthropic `cache_control` fields.
184
185
  - Coding usage: `cache_read_input_tokens` → `Usage.cacheReadTokens`,
185
186
  `cache_creation_input_tokens` → `Usage.cacheWriteTokens`.