@arnilo/prism 0.2.9 → 0.3.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (77) hide show
  1. package/CHANGELOG.md +39 -0
  2. package/README.md +12 -5
  3. package/dist/agent-loops.js +45 -8
  4. package/dist/agent-session/helpers.js +2 -2
  5. package/dist/cache-helpers.d.ts +11 -0
  6. package/dist/cache-helpers.js +29 -5
  7. package/dist/cli-provider-add.js +2 -1
  8. package/dist/context-budget.js +9 -6
  9. package/dist/contracts-core/agent.d.ts +2 -0
  10. package/dist/contracts-core/provider.d.ts +2 -0
  11. package/dist/contracts-protocol.d.ts +31 -1
  12. package/dist/delegated-agent-step.d.ts +20 -0
  13. package/dist/delegated-agent-step.js +99 -0
  14. package/dist/event-multiplexer.js +0 -4
  15. package/dist/index.d.ts +6 -2
  16. package/dist/index.js +5 -2
  17. package/dist/input.js +19 -11
  18. package/dist/node/session-store-jsonl.js +7 -3
  19. package/dist/providers/openai-compatible.js +2 -1
  20. package/dist/providers/openai-primitives.js +2 -1
  21. package/dist/providers/schema.d.ts +7 -0
  22. package/dist/providers/schema.js +25 -0
  23. package/dist/testing/provider-conformance.d.ts +10 -0
  24. package/dist/testing/provider-conformance.js +37 -0
  25. package/dist/trim-trailing-slashes.d.ts +8 -0
  26. package/dist/trim-trailing-slashes.js +14 -0
  27. package/docs/0.1.0-readiness.md +8 -8
  28. package/docs/acp.md +5 -3
  29. package/docs/ag-ui.md +6 -2
  30. package/docs/agent-events.md +8 -1
  31. package/docs/agent-loops.md +3 -0
  32. package/docs/agent-session-runtime.md +1 -0
  33. package/docs/antigravity-agent.md +207 -0
  34. package/docs/browser-automation.md +1 -0
  35. package/docs/coding-agent-tools.md +32 -4
  36. package/docs/computer-use-linux.md +122 -0
  37. package/docs/database-persistence.md +1 -1
  38. package/docs/device-adapters.md +4 -3
  39. package/docs/graft.md +125 -0
  40. package/docs/host-security.md +3 -1
  41. package/docs/index.md +21 -10
  42. package/docs/input-and-prompt-assembly.md +11 -6
  43. package/docs/instruction-injection.md +1 -1
  44. package/docs/mcp-tools.md +2 -1
  45. package/docs/migration.md +25 -2
  46. package/docs/node-jsonl-session-store.md +1 -1
  47. package/docs/obscura.md +175 -0
  48. package/docs/observability.md +21 -1
  49. package/docs/performance.md +58 -4
  50. package/docs/ponytail.md +1 -1
  51. package/docs/provider-caching.md +13 -11
  52. package/docs/provider-conformance.md +6 -0
  53. package/docs/provider-packages.md +1 -1
  54. package/docs/provider-primitives.md +15 -2
  55. package/docs/providers/ai-sdk.md +1 -1
  56. package/docs/providers/anthropic.md +1 -1
  57. package/docs/providers/azure.md +1 -0
  58. package/docs/providers/bedrock.md +1 -0
  59. package/docs/providers/kimi.md +2 -1
  60. package/docs/providers/openai.md +19 -7
  61. package/docs/providers/opencode-go.md +3 -1
  62. package/docs/providers/openrouter.md +4 -3
  63. package/docs/providers/vertex.md +1 -0
  64. package/docs/public-contracts.md +2 -1
  65. package/docs/rag.md +55 -8
  66. package/docs/release-and-install.md +105 -25
  67. package/docs/server.md +1 -0
  68. package/docs/supervisors.md +3 -2
  69. package/docs/system-prompts.md +1 -1
  70. package/docs/tools.md +1 -1
  71. package/docs/web-tools.md +2 -0
  72. package/docs/wiki.md +140 -0
  73. package/docs/workflows.md +4 -3
  74. package/docs/working-and-semantic-memory.md +20 -0
  75. package/package.json +14 -5
  76. package/docs/api-page-template.md +0 -32
  77. package/docs/release-0.2.7-evidence.md +0 -514
@@ -28,6 +28,47 @@ node scripts/benchmark-0.1.0.mjs --out scripts/benchmark-0.1.0.json
28
28
  PRISM_TEST_POSTGRES_URL="postgresql://…" node scripts/benchmark-0.1.0.mjs --out scripts/benchmark-0.1.0.json # adds protected legs
29
29
  ```
30
30
 
31
+ ## Multi-agent runtime concurrency (phase 35)
32
+
33
+ `node scripts/benchmark.mjs --scenario multi-agent-runtime` is network-free (mock providers, in-process memory stores, no credentials). It measures concurrent independent sessions (1/4/16/32), supervisor fan-out and saturation (32 attempted delegates vs `maxActiveChildren`), parallel workflow fan-out maps (8×20 ms items at concurrency 2, ≥1.75× vs sequential), parallel workflow agent nodes, in-run tool concurrency, and an abort storm. Each result row carries p50/p95, throughput, heap delta, queued/dropped events, peak active provider calls, completions, and abort settle. Ceilings live in `scripts/budgets.json#multiAgentRuntime` (sanity bounds, machine-dependent). Exhaustive 59-manifest classification and recorded numbers: [`docs/_evidence/phase35-ai-runtime-package-matrix.md`](./_evidence/phase35-ai-runtime-package-matrix.md). Schema/safety/invariants: `scripts/benchmark-multi-agent.test.mjs`. Fan-out row: 8×20 ms items at concurrency 2, ≥1.75× vs sequential, peak workers ≤ 2. Supervisor saturation: 32 attempted delegates vs `maxActiveChildren` 4, overflow rejected, `activeAfter` 0.
34
+
35
+ ```bash
36
+ node scripts/benchmark.mjs --scenario multi-agent-runtime --out /tmp/prism-multi-agent.json
37
+ ```
38
+
39
+ Recorded 2026-08-27, Node v24.19.0 / Linux x64, 5 warmups + 20 waves, 8 ms mock delay. 32 independent sessions p95 10.1 ms (vs 9.0 ms at n=1); supervisor cap-4 fan-out p95 9.4 ms; workflow 4 agent nodes at concurrency 2 p95 17.9 ms; 8 tools at concurrency 4 p95 17.4 ms; abort storm settled in 5.5 ms with zero leftover provider calls. Dropped events: 0 on every row. Task 6 three-run median p95 (2026-08-28, same fixture): sessions 9.9/9.9/10.6/12.0, supervisorFanOut 10.4, supervisorSaturation 10.3, workflowFanOut 84.5 (1.87×, peak workers 2), workflowAgentNodes 19.7, toolConcurrency 19.1, abortStorm 3.3 — all under `scripts/budgets.json#multiAgentRuntime` ceilings. Protected PostgreSQL (`PRISM_TEST_POSTGRES_URL`) skipped on this host; `release:gate` blocked until durable evidence exists. Memory-store router 16/32-worker reservations do not oversubscribe.
40
+
41
+ ## Large-history and streamed-delta hot paths (plan 036)
42
+
43
+ The `multi-agent-runtime` scenario also covers 10,000 context-budget history rows and
44
+ 5,000 streamed provider deltas. `applyContextBudget` measures the keep-set once,
45
+ advances a history head cursor during eviction, and slices the retained suffix once;
46
+ it never front-mutates the history array. Runtime request/response limit accounting
47
+ uses `Buffer.byteLength(JSON.stringify(value), "utf8")`, so UTF-8 byte limits do not
48
+ allocate an encoded buffer per provider event.
49
+
50
+ Run with the existing network-free fixture:
51
+
52
+ ```bash
53
+ node scripts/benchmark.mjs --scenario multi-agent-runtime
54
+ ```
55
+
56
+ Recorded 2026-08-28 on Node v24.19.0 / Linux x64, 5 warmups + 20 measured waves.
57
+ `contextBudget-10k-history` completed with zero history remaining; its p50/p95 were
58
+ 2.481/3.708 ms and peak measured heap delta was 7,748,848 bytes. `provider-5k-deltas`
59
+ processed 5,000 deltas (320,015 serialized response bytes) at 4.253/5.847 ms p50/p95
60
+ with a 4,414,944-byte peak measured heap delta. These are local comparison evidence,
61
+ not portable SLOs; ceilings are in `scripts/budgets.json#multiAgentRuntime`.
62
+
63
+ `Buffer.byteLength` counts encoded bytes rather than JavaScript string length. A
64
+ serialized provider event at the exact response-byte cap succeeds; one byte below it
65
+ fails closed, including multibyte Unicode deltas. Context-budget omission order and
66
+ newest-history preservation remain covered by the root context-budget tests.
67
+
68
+ ## Current-line root artifact diet
69
+
70
+ `npm pack --dry-run --json` on `@arnilo/prism` is gated by `scripts/budget-gate.test.mjs` against `scripts/budgets.json#root` (±5%). Repository-only history stays out of the tarball: `docs/_evidence/**`, `docs/release-*-evidence.md`, `docs/api-page-template.md`, `dist/__tests__`, and `*.map`. Every page linked from shipped `docs/index.md` must be in the pack. Recorded 2026-08-27: **923,045 packed / 3,149,665 unpacked / 375 files** (226 `dist` js+d.ts, 124 index-linked docs, 25 other). 0.1.0 freeze 713,454 / 293 stays historical.
71
+
31
72
  ## 0.1.4 tree-shake measurement (static-reachability proxy)
32
73
 
33
74
  The 0.1.4 god-module split (agents/contracts → per-concern modules behind barrels) is
@@ -89,9 +130,10 @@ file count fail above baseline × 1.05. Labels: **network-free** = runs in
89
130
  | reconnectCatchup | 8.374 | 100 | distributed events (0.0.24) | protected |
90
131
 
91
132
  Install/startup rows (same helpers as the budget gate — no duplicate
92
- measurement): startup import 41.7 ms (ceiling 250 ms); root packed 711,755
93
- bytes vs baseline 678,541 (+5%, tolerance 5%); root file count 295 vs 293
94
- (+5%). Storage-growth rows and query plans from the protected legs are in the
133
+ measurement): startup import 41.7 ms (ceiling 250 ms). The recorded 0.1.0.json
134
+ pack rows (711,755 bytes / 295 files vs freeze 678,541 / 293) stay historical.
135
+ Live root pack is the current-line diet in `scripts/budgets.json#root` (see below).
136
+ Storage-growth rows and query plans from the protected legs are in the
95
137
  recorded JSON (`storageBeforeCleanup` / `storageAfterCleanup` per leg).
96
138
 
97
139
  Conformance companions: `scripts/phase8–11-conformance.test.mjs` plus the
@@ -375,7 +417,7 @@ On default overflow, the affected subscriber receives one `event_subscriber_over
375
417
  }
376
418
  ```
377
419
 
378
- `drop_oldest` keeps the newest queued events. `drop_newest` ignores incoming events while the queue is full. These policies are live-view policies only; they do not affect `RunLedger` writes or stored session entries.
420
+ `drop_oldest` keeps the newest queued events. `drop_newest` ignores incoming events while the queue is full. These policies are live-view policies only; they do not affect `RunLedger` writes or stored session entries. Graceful `createEventMultiplexer().close()` (and abort) stop new publishes/sources and drain already-queued events within `maxQueuedEvents` before the subscriber completes. Overflow `close` still drops the backlog, emits one overflow notice, and terminates.
379
421
 
380
422
  ## Request/response example
381
423
 
@@ -707,6 +749,18 @@ Enterprise governance and connector caps (defaults / hard). Timings: `node scrip
707
749
 
708
750
  Offline behavior tests (identity propagation, policy export, router deny paths, fake CLI argv) are release gates; live tenant canaries remain operator-gated.
709
751
 
752
+ ### 0.3.x Phase 39 Obscura browser-engine envelopes (2026-08-29)
753
+
754
+ `@arnilo/prism-obscura` binary-backed legs, network-free, driven by a deterministic fake CLI: `node scripts/benchmark-obscura.mjs` (3 runs, medians vs reviewed ceilings; artifact `scripts/benchmark-obscura.json`). Startup leg probes SIG-0 liveness after spawn — a real host waits on its readiness endpoint inside the same bound.
755
+
756
+ | Leg | Median (3 runs) | Ceiling | Notes |
757
+ | --- | --- | --- | --- |
758
+ | Managed startup (`spawnObscuraProcess` + `waitReady`) | ~0.02 ms | 250 ms | fake child; machine-dependent sanity bound, catches catastrophic lifecycle regression |
759
+ | Bounded CLI `web_search` call | ~20 ms | 100 ms | one `runObscuraCli` round trip through the public tool surface |
760
+ | Group close (SIGTERM drain) | ~0.6 ms | 250 ms | idempotent group-wide close; real children exit on signal |
761
+
762
+ No new release gate: the ceilings are evidence, not gates. Concurrent-resource evidence is behavioral, not timing: the MCP bridge serializes mutations (one live page), and abort tests prove an aborted in-flight call settles and kills the owned child with zero leaked processes (`scripts/obscura-host-conformance.test.mjs` abort leg; process/web suite timeout/abort-kill tests). Packed tarball 34.4 kB / 16 files; the package installs no binary, image, or browser.
763
+
710
764
  ## Related APIs
711
765
 
712
766
  - [Agent events](agent-events.md): `SubscribeOptions` and `event_subscriber_overflow` event details.
package/docs/ponytail.md CHANGED
@@ -106,7 +106,7 @@ See `examples/caveman-ponytail.ts` for combined Caveman + Ponytail progressive d
106
106
  - Import alone registers nothing (`sideEffects: false`); no timers, watchers, network, or shell scripts.
107
107
  - Upstream hook modules load via `createRequire` from resolved root — instruction strings are not forked in Prism.
108
108
  - Mode restore scans `getEntries()` for latest `data.type === "ponytail-mode"` (OM attach pattern).
109
- - `ponytail-subagent` hook is not wired; nested-agent behavior is host responsibility.
109
+ - `ponytail-subagent` hook is not wired; nested-agent behavior is host responsibility. When hosts wire the upstream hook, `PONYTAIL_SUBAGENT_MATCHER` accepts only the documented safe subset — `"explore|general"` (any literal substring) or `"^general$"` (exact), case-insensitive, max 256 chars. No `RegExp` is compiled from the environment, so arbitrary regex (including catastrophic nested quantifiers) is never evaluated; unset/invalid patterns inject into every subagent.
110
110
  - No TUI statusline scripts; use `ponytail status` command or extension events.
111
111
  - Not included in `@arnilo/prism-code` or `@arnilo/prism-sdk` profiles — opt-in install only.
112
112
 
@@ -8,7 +8,7 @@ Provider caching documents Prism's cache intent surface:
8
8
  - Legacy aliases `cacheKey` and `cacheRetention`, still supported for backwards compatibility.
9
9
  - `PromptCacheBreakpoint` locations for reusable prompt regions.
10
10
  - `ModelCacheCapabilities` for model/provider cache support metadata.
11
- - Shared helpers: `sanitizeCacheKey`, `mapCacheRetention`, `applyCacheControl`, `cacheHitRate`, `cacheSavings`, and `cacheUsageReport`.
11
+ - Shared helpers: `sanitizeCacheKey`, `mapCacheRetention`, `applyCacheControl`, `resolveBreakpoint`, `canonicalizeJsonSchema`, `cacheHitRate`, `cacheSavings`, and `cacheUsageReport`.
12
12
 
13
13
  Cache hints are best-effort. They describe intent; providers decide whether their native API can use them. Prism does not guarantee cache hits.
14
14
 
@@ -17,7 +17,7 @@ Cache hints are best-effort. They describe intent; providers decide whether thei
17
17
  Use this page when a host or provider package needs to:
18
18
 
19
19
  - Mark stable system prompts, tools, context, or messages as cacheable.
20
- - Opt into cache-aware default input ordering so stable attachments/resources, summaries, and prior history form a reusable prefix before the current user turn.
20
+ - Use cache-aware default input ordering so stable instructions, attachments/resources, summaries, and prior history form a reusable prefix before the current user turn.
21
21
  - Carry a stable cache key across turns without putting provider-specific fields in core.
22
22
  - Read `ModelConfig.cache` to decide whether to map hints to implicit caching, key-based caching, cache-control breakpoints, provider-specific caching, or no caching.
23
23
  - Compute normalized cache diagnostics from `Usage.cacheReadTokens` / `Usage.cacheWriteTokens`, including providers that only report reads.
@@ -60,13 +60,15 @@ Cache helpers return plain data:
60
60
  | `sanitizeCacheKey(value, maxLength)` | Safe key string or `undefined`. |
61
61
  | `mapCacheRetention(retention, model)` | `"short"`, `"long"`, or `undefined`. |
62
62
  | `applyCacheControl(messages, breakpoints, options)` | New message array with `cache_control: { type: "ephemeral" }` on selected message anchors. |
63
+ | `resolveBreakpoint(messages, breakpoint)` | Message index for a `PromptCacheBreakpoint` (`-1` when unresolved); shared anchor selection for `applyCacheControl` and OpenAI explicit breakpoints. |
64
+ | `canonicalizeJsonSchema(value)` | Clone with sorted object keys and `required` names; semantic arrays stay ordered. Used by first-party tool serializers. |
63
65
  | `cacheHitRate(usage)` | Cached input ratio or `undefined`. |
64
66
  | `cacheSavings(usage, model)` | Estimated read-token savings or `undefined` without pricing. |
65
67
  | `cacheUsageReport(usage, model?)` | Normalized read/write tokens, hit rate, estimated savings, and currency when available; `undefined` when no usage is supplied. |
66
68
 
67
69
  Provider events do not change. Cache accounting stays in normalized `Usage.cacheReadTokens` and `Usage.cacheWriteTokens`.
68
70
 
69
- For stable-prefix payloads, `inputLayout: "cache_aware"` is the default on the default input builder, `assembleProviderInput()`, `AgentConfig`, and `RunOptions`; set `inputLayout: "legacy"` to restore the prior order. The default prompt builder already places context, selected skills, and tool declarations before input messages; cache-aware input ordering then places attachments/resources, summaries, prior history, and pending tool results before the current user suffix. The prefix is byte-stable only when those stable inputs are unchanged; Prism still does not guarantee provider cache hits.
71
+ For stable-prefix payloads, `inputLayout: "cache_aware"` is the default on the default input builder, `assembleProviderInput()`, `AgentConfig`, and `RunOptions`; set `inputLayout: "legacy"` to restore the prior order. The default prompt builder's cache-aware order is leading system instructions → resolved context blocks → selected/progressively disclosed skills → fallback text tool declarations → attachments/resources → summaries → prior history → pending tool results → current input. Declared tool schemas remain in `ProviderRequest.tools` and are never granted by prompt middleware. First-party tool serializers run `canonicalizeJsonSchema` so property insertion order cannot break that prefix. Changing only current input preserves the serialized message prefix before the final user suffix; changing dynamic context or loaded skills changes only from its own boundary onward, while tool schemas remain independently stable. The prefix is byte-stable only when those stable inputs are unchanged; Prism still does not guarantee provider cache hits.
70
72
 
71
73
  ## Request/response example
72
74
 
@@ -135,7 +137,7 @@ Provider request policies can set `ProviderRequestOptions.cache` or the legacy `
135
137
 
136
138
  | `ModelCacheCapabilities.kind` | Typical mapping |
137
139
  | --- | --- |
138
- | `implicit` | No request mutation; provider caches automatically. |
140
+ | `implicit` | No request mutation; provider caches automatically. Conformance (`assertNoForeignCacheFields`) proves the serialized body carries no cache wire fields. |
139
141
  | `openai_key` | Send sanitized cache key and mapped retention where supported. |
140
142
  | `cache_control` | Use `applyCacheControl()` on provider-native message anchors. |
141
143
  | `provider_specific` | Provider package uses `compat`/native options intentionally. |
@@ -145,8 +147,8 @@ Provider request policies can set `ProviderRequestOptions.cache` or the legacy `
145
147
 
146
148
  | Provider package | Cache kind | Explicit cache hints | Multi-turn reuse notes | Caveats |
147
149
  | --- | --- | --- | --- | --- |
148
- | `@arnilo/prism-provider-openai` | `openai_key` | Sends sanitized `prompt_cache_key`; `prompt_cache_retention: "24h"` only when the model declares `longRetention`. | Stable cache key + stable prefix can improve reuse. | Best-effort only; `"short"`/`"none"` omit retention. |
149
- | `@arnilo/prism-provider-anthropic` | `cache_control` | Marks only selected Anthropic message anchors; `"long"` maps to documented `ttl: "1h"`. | Keep selected anchors stable. | Best-effort; never stamp every block. |
150
+ | `@arnilo/prism-provider-openai` | `openai_key` | Sends sanitized `prompt_cache_key`; pre-5.6 models emit `prompt_cache_retention: "24h"` when `longRetention`; GPT-5.6+ models (`explicitBreakpoints`) map `cache.breakpoints`/`cache.mode: "on"` to `prompt_cache_options: { mode: "explicit" }` + `prompt_cache_breakpoint` markers (≤4 writes). | Stable cache key + stable prefix can improve reuse; keep selected anchors stable. | Best-effort only; `"short"`/`"none"` omit retention; `"30m"` TTL is the default and never emitted. |
151
+ | `@arnilo/prism-provider-anthropic` | `cache_control` | Marks only selected Anthropic message anchors; `system_prompt` breakpoints emit native `system` text blocks with the marker; `"long"` maps to documented `ttl: "1h"`. | Keep selected anchors stable. | Best-effort; never stamp every block. |
150
152
  | `@arnilo/prism-provider-google` | none | Sends no Prism cache marker. | Host/model may have upstream behavior. | Gemini cache controls are not mapped in this package. |
151
153
  | `@arnilo/prism-provider-openrouter` | `cache_control` | Top-level automatic `cache_control` when enabled without breakpoints; otherwise markers only on caller-selected `cache.breakpoints`; `"long"` may add `ttl: "1h"`. Sticky `session_id` routing. | Breakpoint-stable / automatic prefixes can be reused by upstream providers. | Best-effort only; top-level automatic may exclude some backends from routing. |
152
154
  | `@arnilo/prism-provider-opencode-go` | route-specific | Sends sanitized `x-opencode-session`; Anthropic route applies selected `cache_control` breakpoints; OpenAI route sends none. | Session id + unchanged selected anchors can help route-native caches. | Best-effort and route-dependent. |
@@ -156,7 +158,7 @@ Provider request policies can set `ProviderRequestOptions.cache` or the legacy `
156
158
  | `@arnilo/prism-provider-ai-sdk` | host-owned | No Prism cache payload; host `LanguageModelV4` owns upstream caching. | Host model/provider decides cache keys, breakpoints, and sticky routing. | Adapter maps `inputTokens.cacheRead`/`cacheWrite` from `finish.usage` only; does not invent cache fields. |
157
159
  | `@arnilo/prism-provider-alibaba` | implicit by default, optional `cache_control` | DashScope implicit prefix caching is automatic; opt-in `cache_control: {"type":"ephemeral"}` markers only on caller-selected `cache.breakpoints`, capped at 4. | Keep selected anchors and prior history stable; each cached prefix needs ≥1024 tokens and lives ~5 minutes upstream. | Best-effort and model-dependent; `cached_tokens`→read, `cache_creation_input_tokens`→write. |
158
160
  | `@arnilo/prism-provider-ollama` | `implicit` | No `cache_control`, `cacheKey`, `prompt_cache`, or `cacheRetention` payload; Ollama KV/prefix caching is automatic with no request knob. | Resend unchanged prior history for implicit KV reuse. | Best-effort only; Ollama reports no cached-token count, so `Usage.cacheReadTokens` stays `undefined`. |
159
- | `@arnilo/prism-provider-deepseek` | `implicit` | No `cache_control` / `prompt_cache_key`; tools schemas are key-sorted so the prefix stays byte-stable. | Resend unchanged history from token 0; append only the new turn. Thinking-on strips temperature/top_p/penalties so they cannot break the prefix. | Best-effort prefix units (~1024 practical min). `prompt_cache_hit_tokens` → `cacheReadTokens`. |
161
+ | `@arnilo/prism-provider-deepseek` | `implicit` | No `cache_control` / `prompt_cache_key`; tool `parameters` go through shared `canonicalizeJsonSchema`. | Resend unchanged history from token 0; append only the new turn. Thinking-on strips temperature/top_p/penalties so they cannot break the prefix. | Best-effort prefix units (~1024 practical min). `prompt_cache_hit_tokens` → `cacheReadTokens`. |
160
162
  | `@arnilo/prism-provider-xai` | `implicit` | No `prompt_cache_key`. Package-local `x-grok-conv-id` is `sanitizeCacheKey(cache.key ?? cacheKey ?? sessionId, 128)`. | Same server + unchanged message prefix. Replay `reasoning_content` on reasoning models or the prefix breaks. | Conv-id is never a credential or SuperGrok token. Omitted when `cache.mode` is `off` or `cacheRetention` is `none`. `cached_tokens` → `cacheReadTokens` (inclusive or exclusive reports kept as-is). |
161
163
  | `@arnilo/prism-provider-clinepass` | `implicit` | No `cache_control` / `prompt_cache_key`. Gateway-owned prefix cache. | Resend unchanged prior history. Stream only. | Best-effort and backend-dependent (`cline-pass/*` slugs). `cached_tokens` / `prompt_cache_hit_tokens` map when present. |
162
164
  | `@arnilo/prism-provider-azure` | none | No Prism cache mapping. | Endpoint/model-specific. | Host owns Azure cache policy. |
@@ -165,11 +167,11 @@ Provider request policies can set `ProviderRequestOptions.cache` or the legacy `
165
167
 
166
168
  Detailed first-party provider notes:
167
169
 
168
- - OpenAI Responses (`@arnilo/prism-provider-openai`): `kind: "openai_key"`. Sanitizes/clamps `prompt_cache_key` to 64 chars; `"long"` retention maps to `prompt_cache_retention: "24h"` only when the model declares `cache.longRetention`; `"short"`/`"none"` omit the field. GPT-5.6+ official docs use `prompt_cache_options` / breakpoints instead of retention — `listOpenAIModels` sets `longRetention: false` for those ids. `input_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
170
+ - OpenAI Responses (`@arnilo/prism-provider-openai`): `kind: "openai_key"`. Sanitizes/clamps `prompt_cache_key` to 64 chars; pre-GPT-5.6 models (`cache.longRetention: true`) map `"long"` retention to `prompt_cache_retention: "24h"`; GPT-5.6+ models (`cache.explicitBreakpoints: true`) map `cache.breakpoints`/`cache.mode: "on"` to `prompt_cache_options: { mode: "explicit" }` plus `prompt_cache_breakpoint: { mode: "explicit" }` markers on selected message anchors (≤4 writes; the only TTL `"30m"` is the default, so none is emitted). Resolved cache fields win over caller `extra`. `input_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`Usage.cacheWriteTokens`.
169
171
  - OpenAI-compatible Chat Completions adapter: minimal scope, sends no cache payload; see [OpenAI-compatible provider](providers/openai-compatible.md).
170
- - Anthropic (`@arnilo/prism-provider-anthropic`): `kind: "cache_control"`; selected Anthropic message anchors receive `cache_control` and eligible long retention maps to `ttl: "1h"`. Cache read/create usage maps to normalized cache read/write tokens.
172
+ - Anthropic (`@arnilo/prism-provider-anthropic`): `kind: "cache_control"`; selected Anthropic message anchors receive `cache_control` and eligible long retention maps to `ttl: "1h"`. A `system_prompt` breakpoint serializes `system` as native text blocks carrying the marker (shared `systemCacheControlField()` helper; plain joined string when unmarked). Cache read/create usage maps to normalized cache read/write tokens.
171
173
  - Google (`@arnilo/prism-provider-google`): sends no Prism cache-control payload. Do not infer cache hits or cache token counts from absent Gemini fields.
172
- - OpenRouter (`@arnilo/prism-provider-openrouter`): `kind: "cache_control"`. Sanitizes/clamps `session_id`/`X-Session-Id` to 256 chars for sticky routing; with no breakpoints emits top-level automatic `cache_control: { type: "ephemeral" }`; with breakpoints applies Anthropic-style markers only to caller-selected locations (last content block of each selected message); `"long"` retention adds `ttl: "1h"` when the model allows it. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`. Optional `listOpenRouterModels()` may populate `ModelConfig.cache`/`cost` from live pricing.
174
+ - OpenRouter (`@arnilo/prism-provider-openrouter`): `kind: "cache_control"`. Sanitizes/clamps `session_id`/`X-Session-Id` to 256 chars for sticky routing (from `cache.key` ?? legacy `cacheKey` ?? `sessionId`); with no breakpoints emits top-level automatic `cache_control: { type: "ephemeral" }`; with breakpoints applies Anthropic-style markers only to caller-selected locations (last content block of each selected message); `"long"` retention adds `ttl: "1h"` when the model allows it. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`. Optional `listOpenRouterModels()` may populate `ModelConfig.cache`/`cost` from live pricing.
173
175
  - OpenCode Go (`@arnilo/prism-provider-opencode-go`): default base `https://opencode.ai/zen/go/v1`; `x-opencode-session` from `cacheKey ?? sessionId`, sanitized to 128 chars; the Anthropic route (MiniMax/Qwen) applies `cache_control` markers only to selected breakpoints (`"long"` → `ttl: "1h"`), the OpenAI route (Grok/GLM/Kimi/MiMo/DeepSeek) sends none and preserves `reasoning_content`. OpenAI route maps `prompt_tokens_details.cached_tokens`/`cache_write_tokens`; Anthropic route maps `cache_read_input_tokens`/`cache_creation_input_tokens`. Caller-gated `listOpenCodeGoModels` against official `GET /zen/go/v1/models`.
174
176
  - Z.AI (`@arnilo/prism-provider-zai`): `kind: "implicit"`. GLM context caching is automatic; sends no explicit cache payload regardless of cache options. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`.
175
177
  - NeuralWatt (`@arnilo/prism-provider-neuralwatt`): `kind: "implicit"`. NeuralWatt prefix caching is automatic; sends no explicit cache payload regardless of cache options. `cacheRetention: "none"` disables Prism cache-control hints only (not the implicit backend prefix cache). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`; NeuralWatt does not report a cache-write token, so `Usage.cacheWriteTokens` is never fabricated (stays `undefined`). NeuralWatt's `/v1/models` catalog advertises exact `cached_input_per_million` rates for cache reads and `cached_output_per_million: null`; static curated aliases do not guess those prices.
@@ -177,7 +179,7 @@ Detailed first-party provider notes:
177
179
  - AI SDK adapter (`@arnilo/prism-provider-ai-sdk`): **host-owned**. Sends no Prism cache payload; the supplied `LanguageModelV4` and its upstream provider own request caching. Maps AI SDK v4 `finish.usage.inputTokens.cacheRead`/`cacheWrite` to `Usage.cacheReadTokens`/`cacheWriteTokens`. No `list*Models()` export.
178
180
  - Alibaba Cloud (`@arnilo/prism-provider-alibaba`): implicit by default, optional `cache_control`. DashScope implicit prefix caching is automatic (no marker); explicit opt-in `cache_control: {"type":"ephemeral"}` markers apply only to selected breakpoints when `ModelConfig.cache.kind: "cache_control"` and the caller supplies breakpoints, capped at 4 (each prefix ≥1024 tokens, ~5 minute TTL). `prompt_tokens_details.cached_tokens`/`cache_creation_input_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`. Caller-gated `listAlibabaModels` against OpenAI-compatible `GET {base}/models`.
179
181
  - Ollama (`@arnilo/prism-provider-ollama`): `kind: "implicit"`. Ollama reuses its KV/prompt cache automatically; there is no request knob and no wire marker, so Prism never emits `cache_control`. Ollama reports no cached-token count, so `Usage.cacheReadTokens` is intentionally left `undefined` (not `0`). Caller-gated `listOllamaModels` against OpenAI-compatible `GET {base}/models`.
180
- - DeepSeek (`@arnilo/prism-provider-deepseek`): `kind: "implicit"`. Official disk prefix cache is automatic (byte-identical prefix from token 0). Adapter sends no cache payload; tool `parameters` are canonicalized for stable JSON key order. `prompt_cache_hit_tokens` maps to `Usage.cacheReadTokens`. Caller-gated `listDeepSeekModels`.
182
+ - DeepSeek (`@arnilo/prism-provider-deepseek`): `kind: "implicit"`. Official disk prefix cache is automatic (byte-identical prefix from token 0). Adapter sends no cache payload; tool `parameters` use shared `canonicalizeJsonSchema` (object keys + unordered `required` only; `enum`/`prefixItems`/`examples` keep caller order). `prompt_cache_hit_tokens` maps to `Usage.cacheReadTokens`. Caller-gated `listDeepSeekModels`.
181
183
  - xAI (`@arnilo/prism-provider-xai`): `kind: "implicit"`. Automatic prefix cache. Sticky `x-grok-conv-id` is a sanitized session/cache key (128 chars), never an OAuth access token. Reasoning models must replay `reasoning_content`. `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`. Caller-gated `listXaiModels`.
182
184
  - ClinePass (`@arnilo/prism-provider-clinepass`): `kind: "implicit"`. No explicit cache payload; multi-backend gateway may report `cached_tokens` or `prompt_cache_hit_tokens`. Static `cline-pass/*` catalog only — no `listClinePassModels`.
183
185
  - Azure, Bedrock, and Vertex: their OpenAI-compatible packages intentionally emit no Prism cache fields. Endpoint/model-specific cache controls remain host-owned rather than guessed from another provider family.
@@ -12,6 +12,9 @@ Exported from `@arnilo/prism/testing/provider-conformance`:
12
12
  - `assertToolCallDeltasReconstruct(events, expected)`
13
13
  - `assertUsageAccounting(events, expected)`
14
14
  - `assertSerializedRequestCoversContent(request, body, options?)`
15
+ - `assertCanonicalToolParameters(serialized, original)`
16
+ - `assertNoForeignCacheFields(body, allowed?)`
17
+ - `assertNoFetches(calls)`
15
18
  - `assertProviderOwnedHeadersWin(captured, options)`
16
19
  - `assertNoSecretLeak(events, secrets)`
17
20
 
@@ -68,6 +71,9 @@ Helpers accept normal `AIProvider`, `ProviderRequest`, `ProviderEvent`, `Usage`,
68
71
  - `assertAbortIsObserved()` passes an already-aborted signal and expects provider generation to reject. This is the supported timeout primitive; use a host abort controller or `RunOptions.signal`.
69
72
  - `assertToolCallDeltasReconstruct()` rebuilds streamed `tool_call_delta` fragments into tool calls and validates expected id/name/arguments. Malformed JSON with id+name present yields `argumentsError` (no throw); missing id/name throws typed `incomplete_delta`. The runtime uses the same reconstruction before tool execution when a provider streams deltas.
70
73
  - `assertUsageAccounting()` finds `usage` or `done.usage` and checks selected token fields including `cacheReadTokens` and `cacheWriteTokens`. This is the provider-neutral check for normalized cache read/write token extraction; every first-party provider package exercises it against server-specific fields (`cached_tokens`, `cache_read_input_tokens`, etc.).
74
+ - `assertCanonicalToolParameters()` checks a serialized tool schema matches `canonicalizeJsonSchema(original)` so property insertion order and `required` name order cannot drift on the wire while `enum`/`prefixItems`/`examples` stay caller-ordered.
75
+ - `assertNoForeignCacheFields()` fails when a request body carries a cache wire field the route does not document (`cache_control`, `prompt_cache_*`, `cachedContent`, `cachePoint`); pass documented fields in `allowed` for routes that opt in (Alibaba/OpenRouter markers, Gemini `extra.cachedContent`). Implicit-cache providers must serialize no foreign cache fields — implicit caching works by byte-stable prefix reuse, not request payloads.
76
+ - `assertNoFetches()` fails when a provider performed network calls outside caller-gated discovery/stream; provider construction and `setup()` must be network-free.
71
77
  - `assertSerializedRequestCoversContent()` scans a serialized provider request body for primitive canaries from each Prism content block and fails if any supported block type is silently dropped. Provider-valid transcripts place assistant `tool_call` messages before matching role `tool` `tool_result` messages; runtime, cache-aware input layout, and observational-memory worker loops preserve that order before serialization.
72
78
  - `assertProviderOwnedHeadersWin()` compares captured request headers against the provider's authoritative owned header values and a caller-supplied header bag; it fails if any owned header (`authorization`, `content-type`, session/security headers) was overridden by caller headers, and also fails if a non-owned caller header was dropped. This is the provider-neutral check that caller `ProviderRequest.options.headers` cannot hijack provider credentials or sessions; every first-party provider package exercises it.
73
79
  - `assertNoSecretLeak()` stringifies all collected events and fails if any known secret string is present.
@@ -115,7 +115,7 @@ Every package remains explicit, setup-zero-fetch, and late-credential-bound. `Mo
115
115
 
116
116
  Every first-party provider package hardens prompt-cache behavior so it cannot emit invalid cache retention values or over-broad cache-control markers, and so provider-owned `authorization`/session/security headers cannot be overridden by caller `ProviderRequest.options.headers`. Cache behavior is provider-specific and best-effort: OpenAI/OpenRouter use explicit hints, NeuralWatt/Z.AI/DeepSeek/ClinePass use implicit caching, xAI adds a sanitized `x-grok-conv-id`, and OpenCode Go/Kimi are route/model-dependent. See [Provider caching](provider-caching.md#per-provider-cache-behavior) for the canonical explicit/implicit matrix.
117
117
 
118
- - **OpenAI** (`kind: openai_key`): `prompt_cache_key` is sanitized and clamped to 64 chars; `prompt_cache_retention` is emitted as `24h` only when the model declares `cache.longRetention`, and omitted for `short`/`none` (the API only accepts absent or `24h`). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
118
+ - **OpenAI** (`kind: openai_key`): `prompt_cache_key` is sanitized and clamped to 64 chars. Pre-GPT-5.6 models emit `prompt_cache_retention: "24h"` when the model declares `cache.longRetention`; GPT-5.6+ models (`cache.explicitBreakpoints`) map `cache.breakpoints` / `cache.mode: "on"` to `prompt_cache_options: { mode: "explicit" }` with `prompt_cache_breakpoint` markers on selected anchors. `input_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`Usage.cacheWriteTokens`.
119
119
  - **OpenAI-compatible core adapter**: Chat Completions sends no `prompt_cache_key`/`prompt_cache_retention`/`cache_control` fields; endpoints cache implicitly. `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
120
120
  - **OpenRouter** (`kind: cache_control`): `session_id`/`x-session-id` sanitized and clamped to 256 chars; with no breakpoints, emits top-level automatic `cache_control: { type: ephemeral }`; with breakpoints, markers applied only to caller-selected locations (not every block); `cacheRetention: long` adds `ttl: 1h` when allowed. Preserves `cached_tokens`/`cache_write_tokens`.
121
121
  - **OpenCode Go**: default base `https://opencode.ai/zen/go/v1`; `x-opencode-session` from `cacheKey ?? sessionId` sanitized to 128 chars; Anthropic route (MiniMax/Qwen) applies `cache_control` only to selected breakpoints (`long` → `ttl: 1h`); OpenAI route (Grok/GLM/Kimi/MiMo/DeepSeek) sends none and preserves `reasoning_content`. Per-route usage mapping. Caller-gated `listOpenCodeGoModels`.
@@ -8,7 +8,7 @@ Implementation is **shipped** for transport and OpenAI serialization primitives
8
8
 
9
9
  ## When to use it
10
10
 
11
- - **Provider package authors** implementing or migrating a first-party adapter should import shared primitives from `@arnilo/prism/providers/transport` and `@arnilo/prism/providers/openai` instead of copying `sse.ts`, `safeText`, `parseArgs`, or serializers.
11
+ - **Provider package authors** implementing or migrating a first-party adapter should import shared primitives from `@arnilo/prism/providers/transport`, `@arnilo/prism/providers/openai`, and `@arnilo/prism/providers/schema` instead of copying `sse.ts`, `safeText`, `parseArgs`, serializers, or JSON Schema key-sorting.
12
12
  - **Host apps** choose native structured output via `ProviderRequestOptions.structuredOutput` when the model declares support; otherwise they keep the artifact generate→validate→revise loop ([Structured output](structured-output.md)).
13
13
  - **Operators** enable observability through extended agent events and the optional OpenTelemetry adapter package ([Observability](observability.md)).
14
14
 
@@ -190,6 +190,19 @@ export function assertOpenAIChatMessage(message: unknown, path: string): asserts
190
190
 
191
191
  `src/providers/openai-compatible.ts` becomes a thin adapter over these helpers in Task 2.
192
192
 
193
+ ### `@arnilo/prism/providers/schema` — **shipped**
194
+
195
+ Deterministic JSON Schema clone for tool/function parameters. Sorts object keys and unordered `required` names. Leaves semantic arrays (`prefixItems`, `examples`, `enum`, tuple `items`) in caller order. Does not resolve `$ref`, mutate input, or enforce schema bounds.
196
+
197
+ ```ts
198
+ import { canonicalizeJsonSchema } from "@arnilo/prism/providers/schema";
199
+
200
+ canonicalizeJsonSchema({ required: ["b", "a"], properties: { b: {}, a: {} } });
201
+ // keys and required names stable; ordered schema arrays remain ordered
202
+ ```
203
+
204
+ `serializeOpenAITool` and first-party native `toTool` mappers reuse this helper so logically identical schemas stringify identically.
205
+
193
206
  ### Structured output capability (Task 4 — **shipped**)
194
207
 
195
208
  ```ts
@@ -286,7 +299,7 @@ Every migrated provider must pass this shared matrix (implemented in Task 1 test
286
299
  ## Related APIs
287
300
 
288
301
  - [Provider layer](provider-layer.md): registry, mock provider, event helpers
289
- - [Provider conformance](provider-conformance.md): stream order, abort, header ownership
302
+ - [Provider conformance](provider-conformance.md): stream order, abort, header ownership, canonical tool schemas
290
303
  - [OpenAI-compatible provider](providers/openai-compatible.md): reference adapter subpath
291
304
  - [Structured output](structured-output.md): artifact loop fallback
292
305
  - [Provider request policies](provider-request-policies.md): cache and request hooks
@@ -112,7 +112,7 @@ console.log(result.text);
112
112
 
113
113
  There is **no Prism-side model catalog** and **no `list*Models()` export** by design. Hosts supply a ready-made `LanguageModelV4` instance (typically from `@ai-sdk/openai`, `@ai-sdk/anthropic`, AI Gateway, or a custom provider) and register a matching `ModelConfig` for capabilities/limits.
114
114
 
115
- Prism setup remains network-free: `createAiSdkProvider` only wraps the supplied model and never fetches catalogs or credentials.
115
+ Prism setup remains network-free: `createAiSdkProvider` only wraps the supplied model and never fetches catalogs or credentials. Stream and abort behavior are conformance-proven: an already-aborted signal fails fast as an `error` event, a host stream ending without a `finish` part fails loudly (typed `AiSdkProviderError { code: "model_error" }`) instead of synthesizing a `done`, and unmappable stream parts fail closed.
116
116
 
117
117
  ## Prompt caching
118
118
 
@@ -41,7 +41,7 @@ Featured offline aliases: `claude-opus-4-8`, `claude-sonnet-5`, `claude-haiku-4-
41
41
  | Surface | Behavior |
42
42
  | --- | --- |
43
43
  | Stream | Prism text, thinking deltas, tool-call delta/final, usage (incl. cache read/create when present), `done`, redacted `error`. |
44
- | Cache | Featured models use `cache.kind: "cache_control"`; markers on selected breakpoints (`long` → `ttl: "1h"`). |
44
+ | Cache | Featured models use `cache.kind: "cache_control"`; markers on selected breakpoints (`long` → `ttl: "1h"`). A `system_prompt` breakpoint serializes `system` as native text blocks with the marker (plain string otherwise). |
45
45
  | Thinking | Model-family aware (`adaptive` vs `enabled`+`budget_tokens`); helpers `anthropicThinking` / `anthropicEffort` / `anthropicPreserveThinking`. |
46
46
  | Auth | `api_key` for provider id; provider-owned `content-type`, `x-api-key`, `anthropic-version` win over caller headers. No OAuth descriptor or subscription adapter is registered. |
47
47
 
@@ -64,6 +64,7 @@ Register via `createExtensionKernel().load([createAzureOpenAIProviderPackage(...
64
64
  - Endpoint host is never rewritten to public DNS.
65
65
  - Errors redact credential values via shared transport helpers.
66
66
  - No Azure SDK dependency.
67
+ - Conformance-proven (Task 6): package `setup()` performs zero fetch and zero credential resolution; an already-aborted signal fails fast; a truncated SSE stream (no `data: [DONE]`) ends in an `error` event; Azure cache policy stays host-owned, so no cache wire fields (`cache_control`, `prompt_cache_*`) are emitted even when the request carries Prism cache hints — only upstream-reported `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
67
68
 
68
69
  ## Related APIs
69
70
 
@@ -62,6 +62,7 @@ Uses Bedrock’s OpenAI-compatible runtime route (not Converse eventstream). Hos
62
62
  - No AWS SDK; package-local SigV4 only for `bedrock` service.
63
63
  - Input headers are normalized once before signing: names are lowercased and duplicate-case keys merge last-wins, so the canonical request always matches the signed header list (no duplicate-case mismatch); query parameters are canonicalized sorted by encoded key then value.
64
64
  - Private endpoint hosts are not rewritten to public DNS.
65
+ - Conformance-proven (Task 6): package `setup()` performs zero fetch and zero credential resolution; an already-aborted signal fails fast; a truncated SSE stream (no `data: [DONE]`) ends in an `error` event; native Bedrock caching (`Converse cachePoint`) is intentionally unsupported on the OpenAI-compatible route — no cache wire fields are emitted even when the request carries Prism cache hints.
65
66
  - Credential secrets are redacted from provider errors.
66
67
  - No credential prefetch at import.
67
68
 
@@ -179,7 +179,8 @@ await kernel.load([
179
179
  - When opted in, `cache_control: { type: "ephemeral" }` markers apply only to
180
180
  caller-selected `ProviderRequestOptions.cache.breakpoints` on the last content
181
181
  block of each selected message. `cacheRetention: "long"` adds `ttl: "1h"` when
182
- the model allows long retention.
182
+ the model allows long retention. A `system_prompt` breakpoint serializes
183
+ `system` as native text blocks with the marker (plain string otherwise).
183
184
  - The Moonshot Open Platform route never receives Anthropic `cache_control` fields.
184
185
  - Coding usage: `cache_read_input_tokens` → `Usage.cacheReadTokens`,
185
186
  `cache_creation_input_tokens` → `Usage.cacheWriteTokens`.
@@ -159,14 +159,26 @@ Official: [Prompt caching](https://developers.openai.com/api/docs/guides/prompt-
159
159
  model declares `ModelConfig.cache.longRetention === true`; models without that
160
160
  metadata omit the field. Featured `gpt-5.1` declares
161
161
  `cache: { kind: "openai_key", longRetention: true, maxKeyLength: 64 }`.
162
- - GPT-5.6+ official docs prefer `prompt_cache_options` / explicit breakpoints;
163
- `listOpenAIModels` sets `longRetention: false` for those ids so Prism does not
164
- emit deprecated `prompt_cache_retention` for them. Breakpoint helpers are not
165
- shipped in this package yet — hosts may pass `prompt_cache_options` through
166
- `compat` / `extra` when needed.
162
+ - GPT-5.6+ models use current `prompt_cache_options` instead of retention:
163
+ `listOpenAIModels` / `mapOpenAIModel` set `cache.explicitBreakpoints: true` and
164
+ `longRetention: false` for those ids, so Prism never emits
165
+ `prompt_cache_retention` for them. When the host supplies
166
+ `cache.breakpoints` (or forces `cache.mode: "on"`), Prism emits
167
+ `prompt_cache_options: { mode: "explicit" }` and stamps
168
+ `prompt_cache_breakpoint: { mode: "explicit" }` on the last text block of each
169
+ selected message anchor (shared breakpoint selection with
170
+ `applyCacheControl`, capped at the official 4 cache writes per request;
171
+ `tools` breakpoints are skipped — tool definitions are not markable blocks).
172
+ The only supported TTL is `"30m"` (also the default), so no ttl field is ever
173
+ emitted. `cache.mode: "off"` suppresses explicit options and markers.
174
+ - Resolved cache fields win over caller `extra`: `prompt_cache_key`,
175
+ `prompt_cache_retention`, and `prompt_cache_options` are re-applied after the
176
+ `extra` spread, so invalid caller values cannot replace the resolved policy.
167
177
  - Cache accounting is preserved in normalized `Usage`: OpenAI
168
- `input_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`. OpenAI
169
- Responses does not report a cache-write token field on older models.
178
+ `input_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens` and
179
+ `input_tokens_details.cache_write_tokens` maps to `Usage.cacheWriteTokens`
180
+ (GPT-5.6+ report cache writes; older models omit the field, leaving
181
+ `Usage.cacheWriteTokens` undefined).
170
182
  - Provider-owned headers (`content-type`, `authorization`, `x-client-request-id`)
171
183
  are applied after caller `ProviderRequestOptions.headers` so caller config
172
184
  cannot replace credentials, content type, or the session request id; non-owned
@@ -219,7 +219,9 @@ Owned compat keys (`route`, `thinking`, `reasoning`, `reasoning_effort`,
219
219
  shared `applyCacheControl()` helper) on the last content block of each selected
220
220
  message — not to every block. Caching is enabled unless disabled
221
221
  (`cacheRetention: "none"` / `cache.mode: "off"`) and the model opts in via
222
- `ModelConfig.cache.kind: "cache_control"` (or `cache.mode: "on"`).
222
+ `ModelConfig.cache.kind: "cache_control"` (or `cache.mode: "on"`). A
223
+ `system_prompt` breakpoint serializes `system` as native text blocks with the
224
+ marker (plain string otherwise).
223
225
  - `cacheRetention: "long"` emits `cache_control: { type: "ephemeral", ttl: "1h" }`
224
226
  markers when the model allows long retention
225
227
  (`ModelConfig.cache.longRetention !== false`); otherwise the default ephemeral
@@ -136,9 +136,10 @@ await kernel.load([
136
136
  ### Cache and session behavior
137
137
 
138
138
  - `session_id` (request body) and the `X-Session-Id` header are derived from
139
- `ProviderRequestOptions.cacheKey` (falling back to `sessionId`) and sanitized
140
- + clamped to 256 characters via the shared `sanitizeCacheKey()` helper.
141
- OpenRouter uses this for provider sticky routing to maximize cache hits.
139
+ `ProviderRequestOptions.cache.key` (falling back to legacy `cacheKey`, then
140
+ `sessionId`) and sanitized + clamped to 256 characters via the shared
141
+ `sanitizeCacheKey()` helper. OpenRouter uses this for provider sticky routing
142
+ to maximize cache hits.
142
143
  - **Automatic caching** (no breakpoints): when caching is enabled for an
143
144
  explicit `cache_control` model (or `compat.openRouterCache` /
144
145
  `cache.mode: "on"`), Prism emits a top-level
@@ -60,6 +60,7 @@ const provider = createVertexProvider({
60
60
  - No Google Cloud SDK dependency in the package.
61
61
  - Custom/private endpoint hosts are preserved.
62
62
  - Tokens redacted from errors; no import-time credential prefetch — the credential is resolved exactly once per request (a rotating `CredentialValueSource` is never consumed twice; the same resolved token drives the wrapper check and the inner auth header).
63
+ - Conformance-proven (Task 6): package `setup()` performs zero fetch and zero credential resolution; an already-aborted signal fails fast; a truncated SSE stream (no `data: [DONE]`) ends in an `error` event; native Vertex cached-content lifecycle is intentionally unsupported on the OpenAI-compatible route — no cache wire fields are emitted even when the request carries Prism cache hints (use `@arnilo/prism-provider-google`'s `extra.cachedContent` on that package, or manage cache resources host-side).
63
64
  - Pair with model-router residency allow-lists on `location`.
64
65
 
65
66
  ## Related APIs
@@ -124,6 +124,7 @@ Important request shapes:
124
124
  | `ContextResolutionContext` | Context provider input: messages plus optional session/run ids, metadata, and signal. |
125
125
  | `InputAssemblyLayout` | Default input layout selector: `"cache_aware"` (default) or opt-in `"legacy"`. |
126
126
  | `DefaultInputBuildContext` | Optional default input assembly context: input layout, instructions, history, summaries, attachments, explicit resources, tool results, middleware, ids, metadata, and signal. |
127
+ | `PromptBuildRequest` | Prompt-builder input: messages, context, selected skills, active tools, model, metadata, signal, and optional `inputLayout`; the default builder uses cache-aware ordering unless `legacy` is explicit. |
127
128
  | `ResolveContextOptions` | Ordered context resolution input: selected providers, messages, ids, metadata, signal, and optional middleware. |
128
129
  | `AssembleProviderInputOptions` | Provider input assembly input: model, input, optional builders, selected context providers/skills, active tools, metadata, and signal. |
129
130
  | `PromptTemplateOptions` | Missing-variable behavior for tiny `renderPromptTemplate()` substitutions. |
@@ -148,7 +149,7 @@ Important request shapes:
148
149
  | `CheckpointStore` | Generic versioned checkpoint capability: save/load/bounded-list/delete by namespace and key, with ownership, exact-version CAS, and lease fencing. `createMemoryCheckpointStore()` is the reference implementation; it is bounded — `maxRecords` (default 10,000, evicts least-recently-saved) and `maxValueBytes` (default 1 MiB per JSON value). |
149
150
  | `LeaseStore` | Atomic acquire/renew/release/get by namespace and key, with opaque claim tokens, expiry, ownership scope, and monotonically increasing takeover fences. `createMemoryLeaseStore()` is the reference implementation. |
150
151
  | `RunFeedbackStore` | Immutable append, bounded owned query, and owned deletion for ratings/comments/tags linked to existing run/trace/evaluation IDs. `createMemoryRunFeedbackStore()` is the reference implementation. |
151
- | `EventMultiplexer<T>` | Generic bounded fan-in from async sources. `createEventMultiplexer()` owns queue limits, overflow policy, abort, source teardown, and close behavior. Single-consumer contract: a second concurrent `subscribe()` throws `EventMultiplexerError` (`ERR_PRISM_EVENT_MULTIPLEXER_SINGLE_CONSUMER`); the slot frees when the active consumer completes/is `return()`ed at a yield or the multiplexer closes. `observe` fan-in is unchanged (broadcast happens at the source). |
152
+ | `EventMultiplexer<T>` | Generic bounded fan-in from async sources. `createEventMultiplexer()` owns queue limits, overflow policy, abort, source teardown, and close behavior. Graceful `close()` stops publishes/sources and drains already-queued events before the subscriber completes; overflow `close` still emits one notice and terminates. Single-consumer contract: a second concurrent `subscribe()` throws `EventMultiplexerError` (`ERR_PRISM_EVENT_MULTIPLEXER_SINGLE_CONSUMER`); the slot frees when the active consumer completes or is `return()`ed at a yield. `observe` fan-in is unchanged (broadcast happens at the source). |
152
153
  | `PersistencePage<T>` | Cursor-paginated result page: `items`, optional `nextCursor`, optional `total`. |
153
154
  | `PersistenceQuery` | Common pagination controls: `cursor?`, `limit?`, `order?: "asc" \| "desc"`. |
154
155
  | `OwnershipScope` | Multi-tenant scope: `tenantId?`, `accountId?`, `userId?`. Included in records and queries. |
package/docs/rag.md CHANGED
@@ -2,7 +2,7 @@
2
2
 
3
3
  ## What it does
4
4
 
5
- `@arnilo/prism-rag` is an optional package for deterministic text/Markdown chunking, bounded embedding/vector indexing, atomic scoped source replacement/deletion, focused text/Markdown/HTML/PDF parsing, bounded reranking, ingestion status, attributable citations, content-trust metadata, and explicit `ContextProvider` injection. It reuses `Embedder` and `VectorStore` from `@arnilo/prism-memory`; Prism core input assembly is unchanged.
5
+ `@arnilo/prism-rag` is an optional package for deterministic text/Markdown chunking (with ATX heading-stack metadata), bounded embedding/vector indexing with embedder-identity drift guards, atomic scoped source replacement/deletion with content-hash skip and generation visibility, hybrid vector+lexical retrieval with reciprocal-rank fusion (one embed / one RRF / one rerank across one or many exact scopes), focused text/Markdown/HTML/PDF parsing, bounded reranking (host seam plus a TEI REST adapter), ingestion status, attributable citations, content-trust metadata, and explicit `ContextProvider` injection. It reuses `Embedder` and `VectorStore` from `@arnilo/prism-memory`; Prism core input assembly is unchanged.
6
6
 
7
7
  ## When to use it
8
8
 
@@ -18,7 +18,7 @@ Chunking:
18
18
  | `chunkMarkdown(markdown, options)` | Same engine, preferring heading/paragraph boundaries |
19
19
  | `sourceId` | Required stable, non-secret source identifier |
20
20
  | `size` / `overlap` | Character ceiling and repeated context |
21
- | `metadata` | JSON metadata copied to every chunk |
21
+ | `metadata` | JSON metadata copied to every chunk; Markdown chunking additionally stamps `heading` (ordered parent-first heading stack, e.g. `["Policy", "3.2 Leave"]`) unless the caller supplies one.
22
22
 
23
23
  Document lifecycle:
24
24
 
@@ -35,13 +35,18 @@ Index/retrieve:
35
35
  | Field | Required | Meaning |
36
36
  | --- | --- | --- |
37
37
  | `embedder` / `store` | yes | Phase 7 `Embedder` and `VectorStore` |
38
- | `scope` | yes | `{ tenantId, resourceId, corpusId }`; corpus maps to vector thread isolation |
38
+ | `scope` / `scopes` | one or the other | Exact `{ tenantId, resourceId, corpusId }` (corpus → vector thread). `scope` is the single-corpus path; `scopes` is 0..`HARD_RETRIEVE_SCOPE_CAP` (8) exact scopes. Empty `scopes` returns no hits and does not embed/search/rerank. Passing both or neither throws. |
39
39
  | `chunks` | indexing | `RagChunk[]` from package chunkers or compatible host parser |
40
- | `topK` / `queryCandidates` | retrieval | Returned result count and bounded pre-filter candidates |
40
+ | `topK` / `queryCandidates` | retrieval | Returned result count and bounded pre-filter candidates (`queryCandidates` is **per scope**) |
41
+ | `lexical` | no | `"fts"` \| `"bm25"` \| `"off"` (default `"off"`); enables the lexical retrieval leg when the store advertises it |
42
+ | `fusion` / `rrfK` | no | `"rrf"` fusion of vector+lexical legs (default `"rrf"` when `lexical` is on; `rrfK` default 60, hard cap 1,000) |
41
43
  | `filter` | no | Shallow JSON metadata equality filter |
42
44
  | `reranker` | no | Host-owned `Reranker` receives redacted bounded `RagHit[]` and must return the same IDs once each, in preferred order. |
43
45
  | `maxRerankBytes` / `maxRerankMs` / `rerankConcurrency` | no | Reranker caps; defaults/hard limits are 64/256 KiB, 2/10 s, and 2/8 active calls per reranker object. |
44
46
  | `statusStore` | no | `IngestionStatusStore` records per-source pending/indexed/failed/partial byte/chunk progress; use `listIngestionStatus()` for capped exact-scope pages. |
47
+ | `contentHash` | no | Host-computed document digest; stamped on records and enables unchanged-source skip in `replaceSource` (`skipIfUnchanged`, default true when present) |
48
+ | `reuseEmbeddings` | no | `ReadonlyMap<string, ReusableEmbedding>` — chunk id → `{ text, embedding }`; embeddings reused (no embed call) when texts match |
49
+ | `telemetry` / `telemetryParent` | no | `RagTelemetry` seam (e.g. `createRagTelemetry()` from `@arnilo/prism-observability-opentelemetry`); spans nest under `telemetryParent` |
45
50
  | `redactor` / `secrets` | no | Redact before embedding, persistence, reranking, and injection |
46
51
  | `signal` | no | Abort embedding, vector operations, reranking, and batch progression |
47
52
 
@@ -51,7 +56,8 @@ Index/retrieve:
51
56
  - `indexChunks()` returns `{ indexed, sourceIds }` after bounded batch upserts.
52
57
  - `replaceSource()` / `deleteSource()` return `{ sourceId, deleted, indexed }`.
53
58
  - `replaceDocument()` carries loader parser metadata into chunk metadata; the web loader preserves web-tools citation ID and `untrusted: true`.
54
- - `retrieveContext()` returns `{ query, trust, text, hits, citations, truncated }`. Every hit/citation carries `{ provenance: { sourceId, chunkId, citationId, provider, retrieval: "vector", retrievedAt }, trust: { untrusted: true, inert: true, injectionCapable: true } }`; `retrievalRank` preserves pre-rerank order. Rendered text uses `[citation-id] text` blocks.
59
+ - `retrieveContext()` returns `{ query, trust, text, hits, citations, truncated }`. Every hit/citation carries `{ provenance: { sourceId, chunkId, citationId, provider, tenantId, resourceId, corpusId, retrieval: "vector" | "lexical" | "hybrid", retrievedAt }, trust: { untrusted: true, inert: true, injectionCapable: true } }`; `retrieval` labels the leg(s) that surfaced the hit after RRF fusion, and `retrievalRank` preserves pre-rerank order. Rendered text uses `[citation-id] text` blocks.
60
+ - `replaceSource()` returns `{ sourceId, deleted, indexed, skipped? }` (skipped when the stored `contentHash` matched and no writes occurred). Records carry `embedderId` (from `Embedder.id`, the Task 2 identity contract) and `generation` (scope-level monotonically bumped index per replacement; `_rag` metadata carries `contentHash` when supplied). `store.getCurrentGeneration(scope)` / `store.setCurrentGeneration(scope, n)` let hosts read and roll back the visible generation; retrieval filters to the current generation while legacy generation-less rows stay visible.
55
61
  - `createMemoryIngestionStatusStore()` is a bounded in-memory reference adapter. `listIngestionStatus({ store, scope, limit, cursor })` returns capped status pages; hosts supply durable stores when status must survive process restart.
56
62
  - `createRagContextProvider()` returns one ordinary context provider. Empty queries/results contribute no block.
57
63
  - No events, tools, permissions, provider calls, loaders, or network requests are added.
@@ -94,7 +100,7 @@ await indexChunks({ chunks, embedder, store, scope, statusStore });
94
100
  const found = await retrieveContext("approval policy", {
95
101
  embedder,
96
102
  store,
97
- scope,
103
+ scopes: [scope], // or `scope` for one corpus
98
104
  topK: 4,
99
105
  filter: { category: "security" },
100
106
  reranker: { rerank: async ({ hits }) => [...hits].sort((a, b) => b.score - a.score) },
@@ -109,11 +115,47 @@ const agent = createAgent({
109
115
  console.log(found.text, await agent.createSession().run("How do approvals work?"));
110
116
  ```
111
117
 
118
+ Content-hash skip and hash validation:
119
+
120
+ ```ts
121
+ import { isValidContentHash } from "@arnilo/prism-rag";
122
+
123
+ const digest = "ab12..."; // host-computed SHA-256 hex of the document
124
+ if (!isValidContentHash(digest)) throw new Error("invalid digest");
125
+ await replaceSource({ sourceId: "doc", chunks, embedder, store, scope, contentHash: digest }); // unchanged → skipped, zero embeds
126
+ ```
127
+
128
+ Hybrid retrieval, TEI reranking, and telemetry:
129
+
130
+ ```ts
131
+ import { createRagTelemetry } from "@arnilo/prism-observability-opentelemetry";
132
+ import { createTeiReranker } from "@arnilo/prism-rag";
133
+
134
+ const telemetry = createRagTelemetry({ tracer, meter }); // @opentelemetry/api instruments
135
+ const org = { tenantId: "t1", resourceId: "docs", corpusId: "org" };
136
+ const user = { tenantId: "t1", resourceId: "docs", corpusId: "user" };
137
+ const session = { tenantId: "t1", resourceId: "docs", corpusId: "session" };
138
+ const found = await retrieveContext("leave balance", {
139
+ embedder,
140
+ store, // a store that advertises lexicalModes: ["fts"]
141
+ scopes: [org, user, session], // one embed, per-scope legs, one RRF, one rerank
142
+ lexical: "fts",
143
+ topK: 8,
144
+ reranker: createTeiReranker({ baseUrl: "https://tei.svc:8080" }),
145
+ telemetry, // roots a rag_request span tree; attachSession/handleAgentEvent NOT required
146
+ });
147
+ ```
148
+
112
149
  ## Extension and configuration notes
113
150
 
114
151
  - Supply any Phase 7-conforming embedder/vector store, including the in-memory reference or PostgreSQL/pgvector adapter.
115
152
  - Metadata filtering is package-local after a bounded candidate query so existing vector contracts/adapters remain unchanged. Increase `queryCandidates` only when selective filters measurably need it.
116
153
  - `Reranker` is a host seam, not a provider integration. Return each redacted candidate ID exactly once; Prism retains canonical hit/provenance/trust fields and exposes `retrievalRank` for diagnostics. Add a hosted reranker only when a host owns its credentials, quota, and retry policy.
154
+ - `createTeiReranker({ baseUrl, model?, timeoutMs?, maxResponseBytes?, ssrf?, allowLoopback?, fetch? })` (`CreateTeiRerankerOptions`) adapts a Hugging Face TEI `POST <baseUrl>/rerank` endpoint (`{query, texts, raw_scores:false}` → `{results:[{index,score}]}`) into the `Reranker` seam. It returns a permutation-only reorder of the same hit objects, so provenance/trust move untouched. Response parsing is strict — short/duplicate/out-of-range indices, non-finite scores, HTTP errors, timeouts, and oversized bodies all fail closed; the `rerankHits` caps (`maxRerankBytes`, `maxRerankMs`, `rerankConcurrency`) still apply around it. The default transport is the core DNS-pinned `pinnedFetch` (redirect-free, byte-bounded to 65,536 by default); HTTPS is required unless `allowLoopback: true` (loopback dev/test) or the host supplies `ssrf`/`fetch` for cluster networking. The adapter validates URL shape only — SSRF policy enforcement stays host-side. No credentials are ever sent; there is no SaaS default URL.
155
+ - Hybrid retrieval: pass `lexical: "fts"` (or `"bm25"` when the store supports it) to `retrieveContext()`; the two legs are fused with reciprocal-rank fusion (`fusion: "rrf"`, `rrfK` 60 default; the pure helper `fuseReciprocalRank()` returns `FusedCandidate[]` for custom orchestration). Stores advertise support via `lexicalModes?: readonly LexicalMode[]` and `tokenizeLexical()` is the shared tokenizer. Each hit's provenance `retrieval` field reports `vector`/`lexical`/`hybrid`; fusion internals expose `RetrievalLeg`.
156
+ - Multi-scope retrieve: `scopes: RagScope[]` searches each exact scope against that scope's current generation, then runs **one** RRF over the union and **one** rerank. The query is embedded once. `queryCandidates` is per scope. Duplicate scopes are dropped. `HARD_RETRIEVE_SCOPE_CAP` is 8.
157
+ - Embedder identity/drift guard: `Embedder.id` (memory contract) is stamped onto every vector record as `embedderId`. `retrieveContext()` fails closed with `ERR_PRISM_RAG_EMBEDDER_MISMATCH` when a stored record's `embedderId` or dimensions differ from the active embedder (for example after a model change) — re-index the source before retrieving. Legacy records without an `embedderId` also fail closed, naming the re-index path.
158
+ - Generations: `replaceSource()` stamps a scope-level generation (auto-incremented per replacement) on staged records and the vector store filters retrieval to the current generation. `setCurrentGeneration()` supports rollback; stores without generation tracking keep legacy behavior (everything visible).
117
159
  - `IngestionStatusStore` is optional observability storage. It is keyed by exact scope and source ID; use `listIngestionStatus()` rather than an unbounded corpus scan. The reference memory store is process-local; implement the same capped scope behavior for durable status.
118
160
  - `createRagContextProvider()` derives its query from latest user text by default; pass a fixed string or callback for host-controlled query generation.
119
161
  - `createResourceDocumentLoader({ loader })` calls one host-owned `ResourceLoader`; it scans nothing and performs no filesystem or network I/O itself. Pass the host's permission/trust context to that loader.
@@ -123,14 +165,19 @@ console.log(found.text, await agent.createSession().run("How do approvals work?"
123
165
 
124
166
  ## Security and performance notes
125
167
 
126
- - Every index/query includes exact tenant/resource/corpus scope; returned records are rechecked and malformed/foreign records fail closed.
168
+ - Every index/query includes exact tenant/resource/corpus scope; returned records are rechecked and malformed/foreign records fail closed. `retrieveContext` accepts `scope` or `scopes` (never both, never neither). Empty `scopes` is the host “no allowed corpora” path — no embed, no search, no rerank. A hit whose stored scope is not in the requested list fails closed. Generation filters stay per scope.
169
+ - Embedding identity is a privacy/consistency boundary: records from a different embedder (or dimension) never silently mingle with new ones — retrieval fails closed and names the re-index path. Generation pointers are scope-scoped: a pointer row belongs to exactly one scope, and visibility is computed inside the store (SQL), never by post-filtering in JS.
127
170
  - Source IDs become citation/storage IDs and must be stable non-secret identifiers. Text and user metadata can be redacted before external embedding and persistence.
171
+ - Heading metadata is document text only — it passes through the existing `maxMetadataBytes` cap as chunk metadata; no new content path is introduced.
172
+ - `contentHash` skip and `reuseEmbeddings` never leak embeddings: reused embeddings are keyed by chunk id within one replacement and only accepted when the stored text matches exactly.
128
173
  - Retrieved documents are untrusted inert context. Prompt-injection text cannot activate tools, skills, credentials, permissions, or extensions.
129
174
  - Remote sources must pass existing resource/media trust, SSRF, MIME, and byte policies before their decoded text reaches this package.
130
175
  - `replaceSource()` stages every bounded embedding before opening the store transaction. It requires a source-aware transactional store and fails closed rather than pretending generic upserts are atomic. `createMemoryVectorStore()` supplies the reference `getBySource()` / transaction capability; durable stores must implement equivalent exact-scope behavior.
131
176
  - `deleteSource()` rechecks every returned record's tenant/resource/corpus and source metadata before delete. Same source IDs in another corpus remain untouched.
132
177
  - Parsers enforce byte/page/time caps, abort before and after parsing, decode UTF-8 strictly, and strip HTML script/style content. Parsed and retrieved text remains untrusted inert context; it never gains tool authority.
133
- - Rerankers receive redacted input under byte/time/concurrency caps. Timeout, abort, unknown/duplicate/missing IDs, oversized input, and reranker failures fail closed; returned objects cannot overwrite Prism provenance/trust fields.
178
+ - Rerankers receive redacted input under byte/time/concurrency caps. Timeout, abort, unknown/duplicate/missing IDs, oversized input, and reranker failures fail closed; returned objects cannot overwrite Prism provenance/trust fields. The TEI adapter adds fail-closed response parsing (permutation completeness, finite scores) and honors the 65,536-byte response ceiling; SSRF/URL policy is host-side (see Extension notes).
179
+ - Telemetry is a host-owned seam: `RagTelemetry` adapter (`createRagTelemetry()`) drops anything outside a fixed span-name set and `rag.*`-shaped attribute keys, so raw chunk text never reaches the tracer unless the host's own `attributeFilter` opts it in; when the seam is absent, instrumentation costs nothing.
180
+ - Durable vector stores (PostgreSQL/pgvector path via `@arnilo/prism-memory`) run their DDL against the host's knowledge database — tables are created in a schema/table the host names (default `prism_memory.semantic_memory`), and hosts must own backup/retention of that database. See [Working and semantic memory](working-and-semantic-memory.md).
134
181
  - Ingestion failure errors are redacted before status storage. Status reads reject foreign scope entries and page-limit violations; status itself creates no permission or tool authority.
135
182
  - Filtering scans at most `queryCandidates` hits; rendering stops at top-K, UTF-8 result bytes, or estimated context-token ceiling.
136
183