@arnilo/prism 0.2.9 → 0.3.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +39 -0
- package/README.md +12 -5
- package/dist/agent-loops.js +45 -8
- package/dist/agent-session/helpers.js +2 -2
- package/dist/cache-helpers.d.ts +11 -0
- package/dist/cache-helpers.js +29 -5
- package/dist/cli-provider-add.js +2 -1
- package/dist/context-budget.js +9 -6
- package/dist/contracts-core/agent.d.ts +2 -0
- package/dist/contracts-core/provider.d.ts +2 -0
- package/dist/contracts-protocol.d.ts +31 -1
- package/dist/delegated-agent-step.d.ts +20 -0
- package/dist/delegated-agent-step.js +99 -0
- package/dist/event-multiplexer.js +0 -4
- package/dist/index.d.ts +6 -2
- package/dist/index.js +5 -2
- package/dist/input.js +19 -11
- package/dist/node/session-store-jsonl.js +7 -3
- package/dist/providers/openai-compatible.js +2 -1
- package/dist/providers/openai-primitives.js +2 -1
- package/dist/providers/schema.d.ts +7 -0
- package/dist/providers/schema.js +25 -0
- package/dist/testing/provider-conformance.d.ts +10 -0
- package/dist/testing/provider-conformance.js +37 -0
- package/dist/trim-trailing-slashes.d.ts +8 -0
- package/dist/trim-trailing-slashes.js +14 -0
- package/docs/0.1.0-readiness.md +8 -8
- package/docs/acp.md +5 -3
- package/docs/ag-ui.md +6 -2
- package/docs/agent-events.md +8 -1
- package/docs/agent-loops.md +3 -0
- package/docs/agent-session-runtime.md +1 -0
- package/docs/antigravity-agent.md +207 -0
- package/docs/browser-automation.md +1 -0
- package/docs/coding-agent-tools.md +32 -4
- package/docs/computer-use-linux.md +122 -0
- package/docs/database-persistence.md +1 -1
- package/docs/device-adapters.md +4 -3
- package/docs/graft.md +125 -0
- package/docs/host-security.md +3 -1
- package/docs/index.md +21 -10
- package/docs/input-and-prompt-assembly.md +11 -6
- package/docs/instruction-injection.md +1 -1
- package/docs/mcp-tools.md +2 -1
- package/docs/migration.md +25 -2
- package/docs/node-jsonl-session-store.md +1 -1
- package/docs/obscura.md +175 -0
- package/docs/observability.md +21 -1
- package/docs/performance.md +58 -4
- package/docs/ponytail.md +1 -1
- package/docs/provider-caching.md +13 -11
- package/docs/provider-conformance.md +6 -0
- package/docs/provider-packages.md +1 -1
- package/docs/provider-primitives.md +15 -2
- package/docs/providers/ai-sdk.md +1 -1
- package/docs/providers/anthropic.md +1 -1
- package/docs/providers/azure.md +1 -0
- package/docs/providers/bedrock.md +1 -0
- package/docs/providers/kimi.md +2 -1
- package/docs/providers/openai.md +19 -7
- package/docs/providers/opencode-go.md +3 -1
- package/docs/providers/openrouter.md +4 -3
- package/docs/providers/vertex.md +1 -0
- package/docs/public-contracts.md +2 -1
- package/docs/rag.md +55 -8
- package/docs/release-and-install.md +105 -25
- package/docs/server.md +1 -0
- package/docs/supervisors.md +3 -2
- package/docs/system-prompts.md +1 -1
- package/docs/tools.md +1 -1
- package/docs/web-tools.md +2 -0
- package/docs/wiki.md +140 -0
- package/docs/workflows.md +4 -3
- package/docs/working-and-semantic-memory.md +20 -0
- package/package.json +14 -5
- package/docs/api-page-template.md +0 -32
- package/docs/release-0.2.7-evidence.md +0 -514
package/docs/performance.md
CHANGED
|
@@ -28,6 +28,47 @@ node scripts/benchmark-0.1.0.mjs --out scripts/benchmark-0.1.0.json
|
|
|
28
28
|
PRISM_TEST_POSTGRES_URL="postgresql://…" node scripts/benchmark-0.1.0.mjs --out scripts/benchmark-0.1.0.json # adds protected legs
|
|
29
29
|
```
|
|
30
30
|
|
|
31
|
+
## Multi-agent runtime concurrency (phase 35)
|
|
32
|
+
|
|
33
|
+
`node scripts/benchmark.mjs --scenario multi-agent-runtime` is network-free (mock providers, in-process memory stores, no credentials). It measures concurrent independent sessions (1/4/16/32), supervisor fan-out and saturation (32 attempted delegates vs `maxActiveChildren`), parallel workflow fan-out maps (8×20 ms items at concurrency 2, ≥1.75× vs sequential), parallel workflow agent nodes, in-run tool concurrency, and an abort storm. Each result row carries p50/p95, throughput, heap delta, queued/dropped events, peak active provider calls, completions, and abort settle. Ceilings live in `scripts/budgets.json#multiAgentRuntime` (sanity bounds, machine-dependent). Exhaustive 59-manifest classification and recorded numbers: [`docs/_evidence/phase35-ai-runtime-package-matrix.md`](./_evidence/phase35-ai-runtime-package-matrix.md). Schema/safety/invariants: `scripts/benchmark-multi-agent.test.mjs`. Fan-out row: 8×20 ms items at concurrency 2, ≥1.75× vs sequential, peak workers ≤ 2. Supervisor saturation: 32 attempted delegates vs `maxActiveChildren` 4, overflow rejected, `activeAfter` 0.
|
|
34
|
+
|
|
35
|
+
```bash
|
|
36
|
+
node scripts/benchmark.mjs --scenario multi-agent-runtime --out /tmp/prism-multi-agent.json
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
Recorded 2026-08-27, Node v24.19.0 / Linux x64, 5 warmups + 20 waves, 8 ms mock delay. 32 independent sessions p95 10.1 ms (vs 9.0 ms at n=1); supervisor cap-4 fan-out p95 9.4 ms; workflow 4 agent nodes at concurrency 2 p95 17.9 ms; 8 tools at concurrency 4 p95 17.4 ms; abort storm settled in 5.5 ms with zero leftover provider calls. Dropped events: 0 on every row. Task 6 three-run median p95 (2026-08-28, same fixture): sessions 9.9/9.9/10.6/12.0, supervisorFanOut 10.4, supervisorSaturation 10.3, workflowFanOut 84.5 (1.87×, peak workers 2), workflowAgentNodes 19.7, toolConcurrency 19.1, abortStorm 3.3 — all under `scripts/budgets.json#multiAgentRuntime` ceilings. Protected PostgreSQL (`PRISM_TEST_POSTGRES_URL`) skipped on this host; `release:gate` blocked until durable evidence exists. Memory-store router 16/32-worker reservations do not oversubscribe.
|
|
40
|
+
|
|
41
|
+
## Large-history and streamed-delta hot paths (plan 036)
|
|
42
|
+
|
|
43
|
+
The `multi-agent-runtime` scenario also covers 10,000 context-budget history rows and
|
|
44
|
+
5,000 streamed provider deltas. `applyContextBudget` measures the keep-set once,
|
|
45
|
+
advances a history head cursor during eviction, and slices the retained suffix once;
|
|
46
|
+
it never front-mutates the history array. Runtime request/response limit accounting
|
|
47
|
+
uses `Buffer.byteLength(JSON.stringify(value), "utf8")`, so UTF-8 byte limits do not
|
|
48
|
+
allocate an encoded buffer per provider event.
|
|
49
|
+
|
|
50
|
+
Run with the existing network-free fixture:
|
|
51
|
+
|
|
52
|
+
```bash
|
|
53
|
+
node scripts/benchmark.mjs --scenario multi-agent-runtime
|
|
54
|
+
```
|
|
55
|
+
|
|
56
|
+
Recorded 2026-08-28 on Node v24.19.0 / Linux x64, 5 warmups + 20 measured waves.
|
|
57
|
+
`contextBudget-10k-history` completed with zero history remaining; its p50/p95 were
|
|
58
|
+
2.481/3.708 ms and peak measured heap delta was 7,748,848 bytes. `provider-5k-deltas`
|
|
59
|
+
processed 5,000 deltas (320,015 serialized response bytes) at 4.253/5.847 ms p50/p95
|
|
60
|
+
with a 4,414,944-byte peak measured heap delta. These are local comparison evidence,
|
|
61
|
+
not portable SLOs; ceilings are in `scripts/budgets.json#multiAgentRuntime`.
|
|
62
|
+
|
|
63
|
+
`Buffer.byteLength` counts encoded bytes rather than JavaScript string length. A
|
|
64
|
+
serialized provider event at the exact response-byte cap succeeds; one byte below it
|
|
65
|
+
fails closed, including multibyte Unicode deltas. Context-budget omission order and
|
|
66
|
+
newest-history preservation remain covered by the root context-budget tests.
|
|
67
|
+
|
|
68
|
+
## Current-line root artifact diet
|
|
69
|
+
|
|
70
|
+
`npm pack --dry-run --json` on `@arnilo/prism` is gated by `scripts/budget-gate.test.mjs` against `scripts/budgets.json#root` (±5%). Repository-only history stays out of the tarball: `docs/_evidence/**`, `docs/release-*-evidence.md`, `docs/api-page-template.md`, `dist/__tests__`, and `*.map`. Every page linked from shipped `docs/index.md` must be in the pack. Recorded 2026-08-27: **923,045 packed / 3,149,665 unpacked / 375 files** (226 `dist` js+d.ts, 124 index-linked docs, 25 other). 0.1.0 freeze 713,454 / 293 stays historical.
|
|
71
|
+
|
|
31
72
|
## 0.1.4 tree-shake measurement (static-reachability proxy)
|
|
32
73
|
|
|
33
74
|
The 0.1.4 god-module split (agents/contracts → per-concern modules behind barrels) is
|
|
@@ -89,9 +130,10 @@ file count fail above baseline × 1.05. Labels: **network-free** = runs in
|
|
|
89
130
|
| reconnectCatchup | 8.374 | 100 | distributed events (0.0.24) | protected |
|
|
90
131
|
|
|
91
132
|
Install/startup rows (same helpers as the budget gate — no duplicate
|
|
92
|
-
measurement): startup import 41.7 ms (ceiling 250 ms)
|
|
93
|
-
bytes vs
|
|
94
|
-
|
|
133
|
+
measurement): startup import 41.7 ms (ceiling 250 ms). The recorded 0.1.0.json
|
|
134
|
+
pack rows (711,755 bytes / 295 files vs freeze 678,541 / 293) stay historical.
|
|
135
|
+
Live root pack is the current-line diet in `scripts/budgets.json#root` (see below).
|
|
136
|
+
Storage-growth rows and query plans from the protected legs are in the
|
|
95
137
|
recorded JSON (`storageBeforeCleanup` / `storageAfterCleanup` per leg).
|
|
96
138
|
|
|
97
139
|
Conformance companions: `scripts/phase8–11-conformance.test.mjs` plus the
|
|
@@ -375,7 +417,7 @@ On default overflow, the affected subscriber receives one `event_subscriber_over
|
|
|
375
417
|
}
|
|
376
418
|
```
|
|
377
419
|
|
|
378
|
-
`drop_oldest` keeps the newest queued events. `drop_newest` ignores incoming events while the queue is full. These policies are live-view policies only; they do not affect `RunLedger` writes or stored session entries.
|
|
420
|
+
`drop_oldest` keeps the newest queued events. `drop_newest` ignores incoming events while the queue is full. These policies are live-view policies only; they do not affect `RunLedger` writes or stored session entries. Graceful `createEventMultiplexer().close()` (and abort) stop new publishes/sources and drain already-queued events within `maxQueuedEvents` before the subscriber completes. Overflow `close` still drops the backlog, emits one overflow notice, and terminates.
|
|
379
421
|
|
|
380
422
|
## Request/response example
|
|
381
423
|
|
|
@@ -707,6 +749,18 @@ Enterprise governance and connector caps (defaults / hard). Timings: `node scrip
|
|
|
707
749
|
|
|
708
750
|
Offline behavior tests (identity propagation, policy export, router deny paths, fake CLI argv) are release gates; live tenant canaries remain operator-gated.
|
|
709
751
|
|
|
752
|
+
### 0.3.x Phase 39 Obscura browser-engine envelopes (2026-08-29)
|
|
753
|
+
|
|
754
|
+
`@arnilo/prism-obscura` binary-backed legs, network-free, driven by a deterministic fake CLI: `node scripts/benchmark-obscura.mjs` (3 runs, medians vs reviewed ceilings; artifact `scripts/benchmark-obscura.json`). Startup leg probes SIG-0 liveness after spawn — a real host waits on its readiness endpoint inside the same bound.
|
|
755
|
+
|
|
756
|
+
| Leg | Median (3 runs) | Ceiling | Notes |
|
|
757
|
+
| --- | --- | --- | --- |
|
|
758
|
+
| Managed startup (`spawnObscuraProcess` + `waitReady`) | ~0.02 ms | 250 ms | fake child; machine-dependent sanity bound, catches catastrophic lifecycle regression |
|
|
759
|
+
| Bounded CLI `web_search` call | ~20 ms | 100 ms | one `runObscuraCli` round trip through the public tool surface |
|
|
760
|
+
| Group close (SIGTERM drain) | ~0.6 ms | 250 ms | idempotent group-wide close; real children exit on signal |
|
|
761
|
+
|
|
762
|
+
No new release gate: the ceilings are evidence, not gates. Concurrent-resource evidence is behavioral, not timing: the MCP bridge serializes mutations (one live page), and abort tests prove an aborted in-flight call settles and kills the owned child with zero leaked processes (`scripts/obscura-host-conformance.test.mjs` abort leg; process/web suite timeout/abort-kill tests). Packed tarball 34.4 kB / 16 files; the package installs no binary, image, or browser.
|
|
763
|
+
|
|
710
764
|
## Related APIs
|
|
711
765
|
|
|
712
766
|
- [Agent events](agent-events.md): `SubscribeOptions` and `event_subscriber_overflow` event details.
|
package/docs/ponytail.md
CHANGED
|
@@ -106,7 +106,7 @@ See `examples/caveman-ponytail.ts` for combined Caveman + Ponytail progressive d
|
|
|
106
106
|
- Import alone registers nothing (`sideEffects: false`); no timers, watchers, network, or shell scripts.
|
|
107
107
|
- Upstream hook modules load via `createRequire` from resolved root — instruction strings are not forked in Prism.
|
|
108
108
|
- Mode restore scans `getEntries()` for latest `data.type === "ponytail-mode"` (OM attach pattern).
|
|
109
|
-
- `ponytail-subagent` hook is not wired; nested-agent behavior is host responsibility.
|
|
109
|
+
- `ponytail-subagent` hook is not wired; nested-agent behavior is host responsibility. When hosts wire the upstream hook, `PONYTAIL_SUBAGENT_MATCHER` accepts only the documented safe subset — `"explore|general"` (any literal substring) or `"^general$"` (exact), case-insensitive, max 256 chars. No `RegExp` is compiled from the environment, so arbitrary regex (including catastrophic nested quantifiers) is never evaluated; unset/invalid patterns inject into every subagent.
|
|
110
110
|
- No TUI statusline scripts; use `ponytail status` command or extension events.
|
|
111
111
|
- Not included in `@arnilo/prism-code` or `@arnilo/prism-sdk` profiles — opt-in install only.
|
|
112
112
|
|
package/docs/provider-caching.md
CHANGED
|
@@ -8,7 +8,7 @@ Provider caching documents Prism's cache intent surface:
|
|
|
8
8
|
- Legacy aliases `cacheKey` and `cacheRetention`, still supported for backwards compatibility.
|
|
9
9
|
- `PromptCacheBreakpoint` locations for reusable prompt regions.
|
|
10
10
|
- `ModelCacheCapabilities` for model/provider cache support metadata.
|
|
11
|
-
- Shared helpers: `sanitizeCacheKey`, `mapCacheRetention`, `applyCacheControl`, `cacheHitRate`, `cacheSavings`, and `cacheUsageReport`.
|
|
11
|
+
- Shared helpers: `sanitizeCacheKey`, `mapCacheRetention`, `applyCacheControl`, `resolveBreakpoint`, `canonicalizeJsonSchema`, `cacheHitRate`, `cacheSavings`, and `cacheUsageReport`.
|
|
12
12
|
|
|
13
13
|
Cache hints are best-effort. They describe intent; providers decide whether their native API can use them. Prism does not guarantee cache hits.
|
|
14
14
|
|
|
@@ -17,7 +17,7 @@ Cache hints are best-effort. They describe intent; providers decide whether thei
|
|
|
17
17
|
Use this page when a host or provider package needs to:
|
|
18
18
|
|
|
19
19
|
- Mark stable system prompts, tools, context, or messages as cacheable.
|
|
20
|
-
-
|
|
20
|
+
- Use cache-aware default input ordering so stable instructions, attachments/resources, summaries, and prior history form a reusable prefix before the current user turn.
|
|
21
21
|
- Carry a stable cache key across turns without putting provider-specific fields in core.
|
|
22
22
|
- Read `ModelConfig.cache` to decide whether to map hints to implicit caching, key-based caching, cache-control breakpoints, provider-specific caching, or no caching.
|
|
23
23
|
- Compute normalized cache diagnostics from `Usage.cacheReadTokens` / `Usage.cacheWriteTokens`, including providers that only report reads.
|
|
@@ -60,13 +60,15 @@ Cache helpers return plain data:
|
|
|
60
60
|
| `sanitizeCacheKey(value, maxLength)` | Safe key string or `undefined`. |
|
|
61
61
|
| `mapCacheRetention(retention, model)` | `"short"`, `"long"`, or `undefined`. |
|
|
62
62
|
| `applyCacheControl(messages, breakpoints, options)` | New message array with `cache_control: { type: "ephemeral" }` on selected message anchors. |
|
|
63
|
+
| `resolveBreakpoint(messages, breakpoint)` | Message index for a `PromptCacheBreakpoint` (`-1` when unresolved); shared anchor selection for `applyCacheControl` and OpenAI explicit breakpoints. |
|
|
64
|
+
| `canonicalizeJsonSchema(value)` | Clone with sorted object keys and `required` names; semantic arrays stay ordered. Used by first-party tool serializers. |
|
|
63
65
|
| `cacheHitRate(usage)` | Cached input ratio or `undefined`. |
|
|
64
66
|
| `cacheSavings(usage, model)` | Estimated read-token savings or `undefined` without pricing. |
|
|
65
67
|
| `cacheUsageReport(usage, model?)` | Normalized read/write tokens, hit rate, estimated savings, and currency when available; `undefined` when no usage is supplied. |
|
|
66
68
|
|
|
67
69
|
Provider events do not change. Cache accounting stays in normalized `Usage.cacheReadTokens` and `Usage.cacheWriteTokens`.
|
|
68
70
|
|
|
69
|
-
For stable-prefix payloads, `inputLayout: "cache_aware"` is the default on the default input builder, `assembleProviderInput()`, `AgentConfig`, and `RunOptions`; set `inputLayout: "legacy"` to restore the prior order. The default prompt builder
|
|
71
|
+
For stable-prefix payloads, `inputLayout: "cache_aware"` is the default on the default input builder, `assembleProviderInput()`, `AgentConfig`, and `RunOptions`; set `inputLayout: "legacy"` to restore the prior order. The default prompt builder's cache-aware order is leading system instructions → resolved context blocks → selected/progressively disclosed skills → fallback text tool declarations → attachments/resources → summaries → prior history → pending tool results → current input. Declared tool schemas remain in `ProviderRequest.tools` and are never granted by prompt middleware. First-party tool serializers run `canonicalizeJsonSchema` so property insertion order cannot break that prefix. Changing only current input preserves the serialized message prefix before the final user suffix; changing dynamic context or loaded skills changes only from its own boundary onward, while tool schemas remain independently stable. The prefix is byte-stable only when those stable inputs are unchanged; Prism still does not guarantee provider cache hits.
|
|
70
72
|
|
|
71
73
|
## Request/response example
|
|
72
74
|
|
|
@@ -135,7 +137,7 @@ Provider request policies can set `ProviderRequestOptions.cache` or the legacy `
|
|
|
135
137
|
|
|
136
138
|
| `ModelCacheCapabilities.kind` | Typical mapping |
|
|
137
139
|
| --- | --- |
|
|
138
|
-
| `implicit` | No request mutation; provider caches automatically. |
|
|
140
|
+
| `implicit` | No request mutation; provider caches automatically. Conformance (`assertNoForeignCacheFields`) proves the serialized body carries no cache wire fields. |
|
|
139
141
|
| `openai_key` | Send sanitized cache key and mapped retention where supported. |
|
|
140
142
|
| `cache_control` | Use `applyCacheControl()` on provider-native message anchors. |
|
|
141
143
|
| `provider_specific` | Provider package uses `compat`/native options intentionally. |
|
|
@@ -145,8 +147,8 @@ Provider request policies can set `ProviderRequestOptions.cache` or the legacy `
|
|
|
145
147
|
|
|
146
148
|
| Provider package | Cache kind | Explicit cache hints | Multi-turn reuse notes | Caveats |
|
|
147
149
|
| --- | --- | --- | --- | --- |
|
|
148
|
-
| `@arnilo/prism-provider-openai` | `openai_key` | Sends sanitized `prompt_cache_key`; `prompt_cache_retention: "24h"`
|
|
149
|
-
| `@arnilo/prism-provider-anthropic` | `cache_control` | Marks only selected Anthropic message anchors; `"long"` maps to documented `ttl: "1h"`. | Keep selected anchors stable. | Best-effort; never stamp every block. |
|
|
150
|
+
| `@arnilo/prism-provider-openai` | `openai_key` | Sends sanitized `prompt_cache_key`; pre-5.6 models emit `prompt_cache_retention: "24h"` when `longRetention`; GPT-5.6+ models (`explicitBreakpoints`) map `cache.breakpoints`/`cache.mode: "on"` to `prompt_cache_options: { mode: "explicit" }` + `prompt_cache_breakpoint` markers (≤4 writes). | Stable cache key + stable prefix can improve reuse; keep selected anchors stable. | Best-effort only; `"short"`/`"none"` omit retention; `"30m"` TTL is the default and never emitted. |
|
|
151
|
+
| `@arnilo/prism-provider-anthropic` | `cache_control` | Marks only selected Anthropic message anchors; `system_prompt` breakpoints emit native `system` text blocks with the marker; `"long"` maps to documented `ttl: "1h"`. | Keep selected anchors stable. | Best-effort; never stamp every block. |
|
|
150
152
|
| `@arnilo/prism-provider-google` | none | Sends no Prism cache marker. | Host/model may have upstream behavior. | Gemini cache controls are not mapped in this package. |
|
|
151
153
|
| `@arnilo/prism-provider-openrouter` | `cache_control` | Top-level automatic `cache_control` when enabled without breakpoints; otherwise markers only on caller-selected `cache.breakpoints`; `"long"` may add `ttl: "1h"`. Sticky `session_id` routing. | Breakpoint-stable / automatic prefixes can be reused by upstream providers. | Best-effort only; top-level automatic may exclude some backends from routing. |
|
|
152
154
|
| `@arnilo/prism-provider-opencode-go` | route-specific | Sends sanitized `x-opencode-session`; Anthropic route applies selected `cache_control` breakpoints; OpenAI route sends none. | Session id + unchanged selected anchors can help route-native caches. | Best-effort and route-dependent. |
|
|
@@ -156,7 +158,7 @@ Provider request policies can set `ProviderRequestOptions.cache` or the legacy `
|
|
|
156
158
|
| `@arnilo/prism-provider-ai-sdk` | host-owned | No Prism cache payload; host `LanguageModelV4` owns upstream caching. | Host model/provider decides cache keys, breakpoints, and sticky routing. | Adapter maps `inputTokens.cacheRead`/`cacheWrite` from `finish.usage` only; does not invent cache fields. |
|
|
157
159
|
| `@arnilo/prism-provider-alibaba` | implicit by default, optional `cache_control` | DashScope implicit prefix caching is automatic; opt-in `cache_control: {"type":"ephemeral"}` markers only on caller-selected `cache.breakpoints`, capped at 4. | Keep selected anchors and prior history stable; each cached prefix needs ≥1024 tokens and lives ~5 minutes upstream. | Best-effort and model-dependent; `cached_tokens`→read, `cache_creation_input_tokens`→write. |
|
|
158
160
|
| `@arnilo/prism-provider-ollama` | `implicit` | No `cache_control`, `cacheKey`, `prompt_cache`, or `cacheRetention` payload; Ollama KV/prefix caching is automatic with no request knob. | Resend unchanged prior history for implicit KV reuse. | Best-effort only; Ollama reports no cached-token count, so `Usage.cacheReadTokens` stays `undefined`. |
|
|
159
|
-
| `@arnilo/prism-provider-deepseek` | `implicit` | No `cache_control` / `prompt_cache_key`;
|
|
161
|
+
| `@arnilo/prism-provider-deepseek` | `implicit` | No `cache_control` / `prompt_cache_key`; tool `parameters` go through shared `canonicalizeJsonSchema`. | Resend unchanged history from token 0; append only the new turn. Thinking-on strips temperature/top_p/penalties so they cannot break the prefix. | Best-effort prefix units (~1024 practical min). `prompt_cache_hit_tokens` → `cacheReadTokens`. |
|
|
160
162
|
| `@arnilo/prism-provider-xai` | `implicit` | No `prompt_cache_key`. Package-local `x-grok-conv-id` is `sanitizeCacheKey(cache.key ?? cacheKey ?? sessionId, 128)`. | Same server + unchanged message prefix. Replay `reasoning_content` on reasoning models or the prefix breaks. | Conv-id is never a credential or SuperGrok token. Omitted when `cache.mode` is `off` or `cacheRetention` is `none`. `cached_tokens` → `cacheReadTokens` (inclusive or exclusive reports kept as-is). |
|
|
161
163
|
| `@arnilo/prism-provider-clinepass` | `implicit` | No `cache_control` / `prompt_cache_key`. Gateway-owned prefix cache. | Resend unchanged prior history. Stream only. | Best-effort and backend-dependent (`cline-pass/*` slugs). `cached_tokens` / `prompt_cache_hit_tokens` map when present. |
|
|
162
164
|
| `@arnilo/prism-provider-azure` | none | No Prism cache mapping. | Endpoint/model-specific. | Host owns Azure cache policy. |
|
|
@@ -165,11 +167,11 @@ Provider request policies can set `ProviderRequestOptions.cache` or the legacy `
|
|
|
165
167
|
|
|
166
168
|
Detailed first-party provider notes:
|
|
167
169
|
|
|
168
|
-
- OpenAI Responses (`@arnilo/prism-provider-openai`): `kind: "openai_key"`. Sanitizes/clamps `prompt_cache_key` to 64 chars; `"long"` retention
|
|
170
|
+
- OpenAI Responses (`@arnilo/prism-provider-openai`): `kind: "openai_key"`. Sanitizes/clamps `prompt_cache_key` to 64 chars; pre-GPT-5.6 models (`cache.longRetention: true`) map `"long"` retention to `prompt_cache_retention: "24h"`; GPT-5.6+ models (`cache.explicitBreakpoints: true`) map `cache.breakpoints`/`cache.mode: "on"` to `prompt_cache_options: { mode: "explicit" }` plus `prompt_cache_breakpoint: { mode: "explicit" }` markers on selected message anchors (≤4 writes; the only TTL `"30m"` is the default, so none is emitted). Resolved cache fields win over caller `extra`. `input_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`Usage.cacheWriteTokens`.
|
|
169
171
|
- OpenAI-compatible Chat Completions adapter: minimal scope, sends no cache payload; see [OpenAI-compatible provider](providers/openai-compatible.md).
|
|
170
|
-
- Anthropic (`@arnilo/prism-provider-anthropic`): `kind: "cache_control"`; selected Anthropic message anchors receive `cache_control` and eligible long retention maps to `ttl: "1h"`. Cache read/create usage maps to normalized cache read/write tokens.
|
|
172
|
+
- Anthropic (`@arnilo/prism-provider-anthropic`): `kind: "cache_control"`; selected Anthropic message anchors receive `cache_control` and eligible long retention maps to `ttl: "1h"`. A `system_prompt` breakpoint serializes `system` as native text blocks carrying the marker (shared `systemCacheControlField()` helper; plain joined string when unmarked). Cache read/create usage maps to normalized cache read/write tokens.
|
|
171
173
|
- Google (`@arnilo/prism-provider-google`): sends no Prism cache-control payload. Do not infer cache hits or cache token counts from absent Gemini fields.
|
|
172
|
-
- OpenRouter (`@arnilo/prism-provider-openrouter`): `kind: "cache_control"`. Sanitizes/clamps `session_id`/`X-Session-Id` to 256 chars for sticky routing; with no breakpoints emits top-level automatic `cache_control: { type: "ephemeral" }`; with breakpoints applies Anthropic-style markers only to caller-selected locations (last content block of each selected message); `"long"` retention adds `ttl: "1h"` when the model allows it. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`. Optional `listOpenRouterModels()` may populate `ModelConfig.cache`/`cost` from live pricing.
|
|
174
|
+
- OpenRouter (`@arnilo/prism-provider-openrouter`): `kind: "cache_control"`. Sanitizes/clamps `session_id`/`X-Session-Id` to 256 chars for sticky routing (from `cache.key` ?? legacy `cacheKey` ?? `sessionId`); with no breakpoints emits top-level automatic `cache_control: { type: "ephemeral" }`; with breakpoints applies Anthropic-style markers only to caller-selected locations (last content block of each selected message); `"long"` retention adds `ttl: "1h"` when the model allows it. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`. Optional `listOpenRouterModels()` may populate `ModelConfig.cache`/`cost` from live pricing.
|
|
173
175
|
- OpenCode Go (`@arnilo/prism-provider-opencode-go`): default base `https://opencode.ai/zen/go/v1`; `x-opencode-session` from `cacheKey ?? sessionId`, sanitized to 128 chars; the Anthropic route (MiniMax/Qwen) applies `cache_control` markers only to selected breakpoints (`"long"` → `ttl: "1h"`), the OpenAI route (Grok/GLM/Kimi/MiMo/DeepSeek) sends none and preserves `reasoning_content`. OpenAI route maps `prompt_tokens_details.cached_tokens`/`cache_write_tokens`; Anthropic route maps `cache_read_input_tokens`/`cache_creation_input_tokens`. Caller-gated `listOpenCodeGoModels` against official `GET /zen/go/v1/models`.
|
|
174
176
|
- Z.AI (`@arnilo/prism-provider-zai`): `kind: "implicit"`. GLM context caching is automatic; sends no explicit cache payload regardless of cache options. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`.
|
|
175
177
|
- NeuralWatt (`@arnilo/prism-provider-neuralwatt`): `kind: "implicit"`. NeuralWatt prefix caching is automatic; sends no explicit cache payload regardless of cache options. `cacheRetention: "none"` disables Prism cache-control hints only (not the implicit backend prefix cache). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`; NeuralWatt does not report a cache-write token, so `Usage.cacheWriteTokens` is never fabricated (stays `undefined`). NeuralWatt's `/v1/models` catalog advertises exact `cached_input_per_million` rates for cache reads and `cached_output_per_million: null`; static curated aliases do not guess those prices.
|
|
@@ -177,7 +179,7 @@ Detailed first-party provider notes:
|
|
|
177
179
|
- AI SDK adapter (`@arnilo/prism-provider-ai-sdk`): **host-owned**. Sends no Prism cache payload; the supplied `LanguageModelV4` and its upstream provider own request caching. Maps AI SDK v4 `finish.usage.inputTokens.cacheRead`/`cacheWrite` to `Usage.cacheReadTokens`/`cacheWriteTokens`. No `list*Models()` export.
|
|
178
180
|
- Alibaba Cloud (`@arnilo/prism-provider-alibaba`): implicit by default, optional `cache_control`. DashScope implicit prefix caching is automatic (no marker); explicit opt-in `cache_control: {"type":"ephemeral"}` markers apply only to selected breakpoints when `ModelConfig.cache.kind: "cache_control"` and the caller supplies breakpoints, capped at 4 (each prefix ≥1024 tokens, ~5 minute TTL). `prompt_tokens_details.cached_tokens`/`cache_creation_input_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`. Caller-gated `listAlibabaModels` against OpenAI-compatible `GET {base}/models`.
|
|
179
181
|
- Ollama (`@arnilo/prism-provider-ollama`): `kind: "implicit"`. Ollama reuses its KV/prompt cache automatically; there is no request knob and no wire marker, so Prism never emits `cache_control`. Ollama reports no cached-token count, so `Usage.cacheReadTokens` is intentionally left `undefined` (not `0`). Caller-gated `listOllamaModels` against OpenAI-compatible `GET {base}/models`.
|
|
180
|
-
- DeepSeek (`@arnilo/prism-provider-deepseek`): `kind: "implicit"`. Official disk prefix cache is automatic (byte-identical prefix from token 0). Adapter sends no cache payload; tool `parameters`
|
|
182
|
+
- DeepSeek (`@arnilo/prism-provider-deepseek`): `kind: "implicit"`. Official disk prefix cache is automatic (byte-identical prefix from token 0). Adapter sends no cache payload; tool `parameters` use shared `canonicalizeJsonSchema` (object keys + unordered `required` only; `enum`/`prefixItems`/`examples` keep caller order). `prompt_cache_hit_tokens` maps to `Usage.cacheReadTokens`. Caller-gated `listDeepSeekModels`.
|
|
181
183
|
- xAI (`@arnilo/prism-provider-xai`): `kind: "implicit"`. Automatic prefix cache. Sticky `x-grok-conv-id` is a sanitized session/cache key (128 chars), never an OAuth access token. Reasoning models must replay `reasoning_content`. `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`. Caller-gated `listXaiModels`.
|
|
182
184
|
- ClinePass (`@arnilo/prism-provider-clinepass`): `kind: "implicit"`. No explicit cache payload; multi-backend gateway may report `cached_tokens` or `prompt_cache_hit_tokens`. Static `cline-pass/*` catalog only — no `listClinePassModels`.
|
|
183
185
|
- Azure, Bedrock, and Vertex: their OpenAI-compatible packages intentionally emit no Prism cache fields. Endpoint/model-specific cache controls remain host-owned rather than guessed from another provider family.
|
|
@@ -12,6 +12,9 @@ Exported from `@arnilo/prism/testing/provider-conformance`:
|
|
|
12
12
|
- `assertToolCallDeltasReconstruct(events, expected)`
|
|
13
13
|
- `assertUsageAccounting(events, expected)`
|
|
14
14
|
- `assertSerializedRequestCoversContent(request, body, options?)`
|
|
15
|
+
- `assertCanonicalToolParameters(serialized, original)`
|
|
16
|
+
- `assertNoForeignCacheFields(body, allowed?)`
|
|
17
|
+
- `assertNoFetches(calls)`
|
|
15
18
|
- `assertProviderOwnedHeadersWin(captured, options)`
|
|
16
19
|
- `assertNoSecretLeak(events, secrets)`
|
|
17
20
|
|
|
@@ -68,6 +71,9 @@ Helpers accept normal `AIProvider`, `ProviderRequest`, `ProviderEvent`, `Usage`,
|
|
|
68
71
|
- `assertAbortIsObserved()` passes an already-aborted signal and expects provider generation to reject. This is the supported timeout primitive; use a host abort controller or `RunOptions.signal`.
|
|
69
72
|
- `assertToolCallDeltasReconstruct()` rebuilds streamed `tool_call_delta` fragments into tool calls and validates expected id/name/arguments. Malformed JSON with id+name present yields `argumentsError` (no throw); missing id/name throws typed `incomplete_delta`. The runtime uses the same reconstruction before tool execution when a provider streams deltas.
|
|
70
73
|
- `assertUsageAccounting()` finds `usage` or `done.usage` and checks selected token fields including `cacheReadTokens` and `cacheWriteTokens`. This is the provider-neutral check for normalized cache read/write token extraction; every first-party provider package exercises it against server-specific fields (`cached_tokens`, `cache_read_input_tokens`, etc.).
|
|
74
|
+
- `assertCanonicalToolParameters()` checks a serialized tool schema matches `canonicalizeJsonSchema(original)` so property insertion order and `required` name order cannot drift on the wire while `enum`/`prefixItems`/`examples` stay caller-ordered.
|
|
75
|
+
- `assertNoForeignCacheFields()` fails when a request body carries a cache wire field the route does not document (`cache_control`, `prompt_cache_*`, `cachedContent`, `cachePoint`); pass documented fields in `allowed` for routes that opt in (Alibaba/OpenRouter markers, Gemini `extra.cachedContent`). Implicit-cache providers must serialize no foreign cache fields — implicit caching works by byte-stable prefix reuse, not request payloads.
|
|
76
|
+
- `assertNoFetches()` fails when a provider performed network calls outside caller-gated discovery/stream; provider construction and `setup()` must be network-free.
|
|
71
77
|
- `assertSerializedRequestCoversContent()` scans a serialized provider request body for primitive canaries from each Prism content block and fails if any supported block type is silently dropped. Provider-valid transcripts place assistant `tool_call` messages before matching role `tool` `tool_result` messages; runtime, cache-aware input layout, and observational-memory worker loops preserve that order before serialization.
|
|
72
78
|
- `assertProviderOwnedHeadersWin()` compares captured request headers against the provider's authoritative owned header values and a caller-supplied header bag; it fails if any owned header (`authorization`, `content-type`, session/security headers) was overridden by caller headers, and also fails if a non-owned caller header was dropped. This is the provider-neutral check that caller `ProviderRequest.options.headers` cannot hijack provider credentials or sessions; every first-party provider package exercises it.
|
|
73
79
|
- `assertNoSecretLeak()` stringifies all collected events and fails if any known secret string is present.
|
|
@@ -115,7 +115,7 @@ Every package remains explicit, setup-zero-fetch, and late-credential-bound. `Mo
|
|
|
115
115
|
|
|
116
116
|
Every first-party provider package hardens prompt-cache behavior so it cannot emit invalid cache retention values or over-broad cache-control markers, and so provider-owned `authorization`/session/security headers cannot be overridden by caller `ProviderRequest.options.headers`. Cache behavior is provider-specific and best-effort: OpenAI/OpenRouter use explicit hints, NeuralWatt/Z.AI/DeepSeek/ClinePass use implicit caching, xAI adds a sanitized `x-grok-conv-id`, and OpenCode Go/Kimi are route/model-dependent. See [Provider caching](provider-caching.md#per-provider-cache-behavior) for the canonical explicit/implicit matrix.
|
|
117
117
|
|
|
118
|
-
- **OpenAI** (`kind: openai_key`): `prompt_cache_key` is sanitized and clamped to 64 chars
|
|
118
|
+
- **OpenAI** (`kind: openai_key`): `prompt_cache_key` is sanitized and clamped to 64 chars. Pre-GPT-5.6 models emit `prompt_cache_retention: "24h"` when the model declares `cache.longRetention`; GPT-5.6+ models (`cache.explicitBreakpoints`) map `cache.breakpoints` / `cache.mode: "on"` to `prompt_cache_options: { mode: "explicit" }` with `prompt_cache_breakpoint` markers on selected anchors. `input_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`Usage.cacheWriteTokens`.
|
|
119
119
|
- **OpenAI-compatible core adapter**: Chat Completions sends no `prompt_cache_key`/`prompt_cache_retention`/`cache_control` fields; endpoints cache implicitly. `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
|
|
120
120
|
- **OpenRouter** (`kind: cache_control`): `session_id`/`x-session-id` sanitized and clamped to 256 chars; with no breakpoints, emits top-level automatic `cache_control: { type: ephemeral }`; with breakpoints, markers applied only to caller-selected locations (not every block); `cacheRetention: long` adds `ttl: 1h` when allowed. Preserves `cached_tokens`/`cache_write_tokens`.
|
|
121
121
|
- **OpenCode Go**: default base `https://opencode.ai/zen/go/v1`; `x-opencode-session` from `cacheKey ?? sessionId` sanitized to 128 chars; Anthropic route (MiniMax/Qwen) applies `cache_control` only to selected breakpoints (`long` → `ttl: 1h`); OpenAI route (Grok/GLM/Kimi/MiMo/DeepSeek) sends none and preserves `reasoning_content`. Per-route usage mapping. Caller-gated `listOpenCodeGoModels`.
|
|
@@ -8,7 +8,7 @@ Implementation is **shipped** for transport and OpenAI serialization primitives
|
|
|
8
8
|
|
|
9
9
|
## When to use it
|
|
10
10
|
|
|
11
|
-
- **Provider package authors** implementing or migrating a first-party adapter should import shared primitives from `@arnilo/prism/providers/transport
|
|
11
|
+
- **Provider package authors** implementing or migrating a first-party adapter should import shared primitives from `@arnilo/prism/providers/transport`, `@arnilo/prism/providers/openai`, and `@arnilo/prism/providers/schema` instead of copying `sse.ts`, `safeText`, `parseArgs`, serializers, or JSON Schema key-sorting.
|
|
12
12
|
- **Host apps** choose native structured output via `ProviderRequestOptions.structuredOutput` when the model declares support; otherwise they keep the artifact generate→validate→revise loop ([Structured output](structured-output.md)).
|
|
13
13
|
- **Operators** enable observability through extended agent events and the optional OpenTelemetry adapter package ([Observability](observability.md)).
|
|
14
14
|
|
|
@@ -190,6 +190,19 @@ export function assertOpenAIChatMessage(message: unknown, path: string): asserts
|
|
|
190
190
|
|
|
191
191
|
`src/providers/openai-compatible.ts` becomes a thin adapter over these helpers in Task 2.
|
|
192
192
|
|
|
193
|
+
### `@arnilo/prism/providers/schema` — **shipped**
|
|
194
|
+
|
|
195
|
+
Deterministic JSON Schema clone for tool/function parameters. Sorts object keys and unordered `required` names. Leaves semantic arrays (`prefixItems`, `examples`, `enum`, tuple `items`) in caller order. Does not resolve `$ref`, mutate input, or enforce schema bounds.
|
|
196
|
+
|
|
197
|
+
```ts
|
|
198
|
+
import { canonicalizeJsonSchema } from "@arnilo/prism/providers/schema";
|
|
199
|
+
|
|
200
|
+
canonicalizeJsonSchema({ required: ["b", "a"], properties: { b: {}, a: {} } });
|
|
201
|
+
// keys and required names stable; ordered schema arrays remain ordered
|
|
202
|
+
```
|
|
203
|
+
|
|
204
|
+
`serializeOpenAITool` and first-party native `toTool` mappers reuse this helper so logically identical schemas stringify identically.
|
|
205
|
+
|
|
193
206
|
### Structured output capability (Task 4 — **shipped**)
|
|
194
207
|
|
|
195
208
|
```ts
|
|
@@ -286,7 +299,7 @@ Every migrated provider must pass this shared matrix (implemented in Task 1 test
|
|
|
286
299
|
## Related APIs
|
|
287
300
|
|
|
288
301
|
- [Provider layer](provider-layer.md): registry, mock provider, event helpers
|
|
289
|
-
- [Provider conformance](provider-conformance.md): stream order, abort, header ownership
|
|
302
|
+
- [Provider conformance](provider-conformance.md): stream order, abort, header ownership, canonical tool schemas
|
|
290
303
|
- [OpenAI-compatible provider](providers/openai-compatible.md): reference adapter subpath
|
|
291
304
|
- [Structured output](structured-output.md): artifact loop fallback
|
|
292
305
|
- [Provider request policies](provider-request-policies.md): cache and request hooks
|
package/docs/providers/ai-sdk.md
CHANGED
|
@@ -112,7 +112,7 @@ console.log(result.text);
|
|
|
112
112
|
|
|
113
113
|
There is **no Prism-side model catalog** and **no `list*Models()` export** by design. Hosts supply a ready-made `LanguageModelV4` instance (typically from `@ai-sdk/openai`, `@ai-sdk/anthropic`, AI Gateway, or a custom provider) and register a matching `ModelConfig` for capabilities/limits.
|
|
114
114
|
|
|
115
|
-
Prism setup remains network-free: `createAiSdkProvider` only wraps the supplied model and never fetches catalogs or credentials.
|
|
115
|
+
Prism setup remains network-free: `createAiSdkProvider` only wraps the supplied model and never fetches catalogs or credentials. Stream and abort behavior are conformance-proven: an already-aborted signal fails fast as an `error` event, a host stream ending without a `finish` part fails loudly (typed `AiSdkProviderError { code: "model_error" }`) instead of synthesizing a `done`, and unmappable stream parts fail closed.
|
|
116
116
|
|
|
117
117
|
## Prompt caching
|
|
118
118
|
|
|
@@ -41,7 +41,7 @@ Featured offline aliases: `claude-opus-4-8`, `claude-sonnet-5`, `claude-haiku-4-
|
|
|
41
41
|
| Surface | Behavior |
|
|
42
42
|
| --- | --- |
|
|
43
43
|
| Stream | Prism text, thinking deltas, tool-call delta/final, usage (incl. cache read/create when present), `done`, redacted `error`. |
|
|
44
|
-
| Cache | Featured models use `cache.kind: "cache_control"`; markers on selected breakpoints (`long` → `ttl: "1h"`). |
|
|
44
|
+
| Cache | Featured models use `cache.kind: "cache_control"`; markers on selected breakpoints (`long` → `ttl: "1h"`). A `system_prompt` breakpoint serializes `system` as native text blocks with the marker (plain string otherwise). |
|
|
45
45
|
| Thinking | Model-family aware (`adaptive` vs `enabled`+`budget_tokens`); helpers `anthropicThinking` / `anthropicEffort` / `anthropicPreserveThinking`. |
|
|
46
46
|
| Auth | `api_key` for provider id; provider-owned `content-type`, `x-api-key`, `anthropic-version` win over caller headers. No OAuth descriptor or subscription adapter is registered. |
|
|
47
47
|
|
package/docs/providers/azure.md
CHANGED
|
@@ -64,6 +64,7 @@ Register via `createExtensionKernel().load([createAzureOpenAIProviderPackage(...
|
|
|
64
64
|
- Endpoint host is never rewritten to public DNS.
|
|
65
65
|
- Errors redact credential values via shared transport helpers.
|
|
66
66
|
- No Azure SDK dependency.
|
|
67
|
+
- Conformance-proven (Task 6): package `setup()` performs zero fetch and zero credential resolution; an already-aborted signal fails fast; a truncated SSE stream (no `data: [DONE]`) ends in an `error` event; Azure cache policy stays host-owned, so no cache wire fields (`cache_control`, `prompt_cache_*`) are emitted even when the request carries Prism cache hints — only upstream-reported `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
|
|
67
68
|
|
|
68
69
|
## Related APIs
|
|
69
70
|
|
|
@@ -62,6 +62,7 @@ Uses Bedrock’s OpenAI-compatible runtime route (not Converse eventstream). Hos
|
|
|
62
62
|
- No AWS SDK; package-local SigV4 only for `bedrock` service.
|
|
63
63
|
- Input headers are normalized once before signing: names are lowercased and duplicate-case keys merge last-wins, so the canonical request always matches the signed header list (no duplicate-case mismatch); query parameters are canonicalized sorted by encoded key then value.
|
|
64
64
|
- Private endpoint hosts are not rewritten to public DNS.
|
|
65
|
+
- Conformance-proven (Task 6): package `setup()` performs zero fetch and zero credential resolution; an already-aborted signal fails fast; a truncated SSE stream (no `data: [DONE]`) ends in an `error` event; native Bedrock caching (`Converse cachePoint`) is intentionally unsupported on the OpenAI-compatible route — no cache wire fields are emitted even when the request carries Prism cache hints.
|
|
65
66
|
- Credential secrets are redacted from provider errors.
|
|
66
67
|
- No credential prefetch at import.
|
|
67
68
|
|
package/docs/providers/kimi.md
CHANGED
|
@@ -179,7 +179,8 @@ await kernel.load([
|
|
|
179
179
|
- When opted in, `cache_control: { type: "ephemeral" }` markers apply only to
|
|
180
180
|
caller-selected `ProviderRequestOptions.cache.breakpoints` on the last content
|
|
181
181
|
block of each selected message. `cacheRetention: "long"` adds `ttl: "1h"` when
|
|
182
|
-
the model allows long retention.
|
|
182
|
+
the model allows long retention. A `system_prompt` breakpoint serializes
|
|
183
|
+
`system` as native text blocks with the marker (plain string otherwise).
|
|
183
184
|
- The Moonshot Open Platform route never receives Anthropic `cache_control` fields.
|
|
184
185
|
- Coding usage: `cache_read_input_tokens` → `Usage.cacheReadTokens`,
|
|
185
186
|
`cache_creation_input_tokens` → `Usage.cacheWriteTokens`.
|
package/docs/providers/openai.md
CHANGED
|
@@ -159,14 +159,26 @@ Official: [Prompt caching](https://developers.openai.com/api/docs/guides/prompt-
|
|
|
159
159
|
model declares `ModelConfig.cache.longRetention === true`; models without that
|
|
160
160
|
metadata omit the field. Featured `gpt-5.1` declares
|
|
161
161
|
`cache: { kind: "openai_key", longRetention: true, maxKeyLength: 64 }`.
|
|
162
|
-
- GPT-5.6+
|
|
163
|
-
`listOpenAIModels`
|
|
164
|
-
|
|
165
|
-
|
|
166
|
-
`
|
|
162
|
+
- GPT-5.6+ models use current `prompt_cache_options` instead of retention:
|
|
163
|
+
`listOpenAIModels` / `mapOpenAIModel` set `cache.explicitBreakpoints: true` and
|
|
164
|
+
`longRetention: false` for those ids, so Prism never emits
|
|
165
|
+
`prompt_cache_retention` for them. When the host supplies
|
|
166
|
+
`cache.breakpoints` (or forces `cache.mode: "on"`), Prism emits
|
|
167
|
+
`prompt_cache_options: { mode: "explicit" }` and stamps
|
|
168
|
+
`prompt_cache_breakpoint: { mode: "explicit" }` on the last text block of each
|
|
169
|
+
selected message anchor (shared breakpoint selection with
|
|
170
|
+
`applyCacheControl`, capped at the official 4 cache writes per request;
|
|
171
|
+
`tools` breakpoints are skipped — tool definitions are not markable blocks).
|
|
172
|
+
The only supported TTL is `"30m"` (also the default), so no ttl field is ever
|
|
173
|
+
emitted. `cache.mode: "off"` suppresses explicit options and markers.
|
|
174
|
+
- Resolved cache fields win over caller `extra`: `prompt_cache_key`,
|
|
175
|
+
`prompt_cache_retention`, and `prompt_cache_options` are re-applied after the
|
|
176
|
+
`extra` spread, so invalid caller values cannot replace the resolved policy.
|
|
167
177
|
- Cache accounting is preserved in normalized `Usage`: OpenAI
|
|
168
|
-
`input_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens
|
|
169
|
-
|
|
178
|
+
`input_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens` and
|
|
179
|
+
`input_tokens_details.cache_write_tokens` maps to `Usage.cacheWriteTokens`
|
|
180
|
+
(GPT-5.6+ report cache writes; older models omit the field, leaving
|
|
181
|
+
`Usage.cacheWriteTokens` undefined).
|
|
170
182
|
- Provider-owned headers (`content-type`, `authorization`, `x-client-request-id`)
|
|
171
183
|
are applied after caller `ProviderRequestOptions.headers` so caller config
|
|
172
184
|
cannot replace credentials, content type, or the session request id; non-owned
|
|
@@ -219,7 +219,9 @@ Owned compat keys (`route`, `thinking`, `reasoning`, `reasoning_effort`,
|
|
|
219
219
|
shared `applyCacheControl()` helper) on the last content block of each selected
|
|
220
220
|
message — not to every block. Caching is enabled unless disabled
|
|
221
221
|
(`cacheRetention: "none"` / `cache.mode: "off"`) and the model opts in via
|
|
222
|
-
`ModelConfig.cache.kind: "cache_control"` (or `cache.mode: "on"`).
|
|
222
|
+
`ModelConfig.cache.kind: "cache_control"` (or `cache.mode: "on"`). A
|
|
223
|
+
`system_prompt` breakpoint serializes `system` as native text blocks with the
|
|
224
|
+
marker (plain string otherwise).
|
|
223
225
|
- `cacheRetention: "long"` emits `cache_control: { type: "ephemeral", ttl: "1h" }`
|
|
224
226
|
markers when the model allows long retention
|
|
225
227
|
(`ModelConfig.cache.longRetention !== false`); otherwise the default ephemeral
|
|
@@ -136,9 +136,10 @@ await kernel.load([
|
|
|
136
136
|
### Cache and session behavior
|
|
137
137
|
|
|
138
138
|
- `session_id` (request body) and the `X-Session-Id` header are derived from
|
|
139
|
-
`ProviderRequestOptions.
|
|
140
|
-
+ clamped to 256 characters via the shared
|
|
141
|
-
OpenRouter uses this for provider sticky routing
|
|
139
|
+
`ProviderRequestOptions.cache.key` (falling back to legacy `cacheKey`, then
|
|
140
|
+
`sessionId`) and sanitized + clamped to 256 characters via the shared
|
|
141
|
+
`sanitizeCacheKey()` helper. OpenRouter uses this for provider sticky routing
|
|
142
|
+
to maximize cache hits.
|
|
142
143
|
- **Automatic caching** (no breakpoints): when caching is enabled for an
|
|
143
144
|
explicit `cache_control` model (or `compat.openRouterCache` /
|
|
144
145
|
`cache.mode: "on"`), Prism emits a top-level
|
package/docs/providers/vertex.md
CHANGED
|
@@ -60,6 +60,7 @@ const provider = createVertexProvider({
|
|
|
60
60
|
- No Google Cloud SDK dependency in the package.
|
|
61
61
|
- Custom/private endpoint hosts are preserved.
|
|
62
62
|
- Tokens redacted from errors; no import-time credential prefetch — the credential is resolved exactly once per request (a rotating `CredentialValueSource` is never consumed twice; the same resolved token drives the wrapper check and the inner auth header).
|
|
63
|
+
- Conformance-proven (Task 6): package `setup()` performs zero fetch and zero credential resolution; an already-aborted signal fails fast; a truncated SSE stream (no `data: [DONE]`) ends in an `error` event; native Vertex cached-content lifecycle is intentionally unsupported on the OpenAI-compatible route — no cache wire fields are emitted even when the request carries Prism cache hints (use `@arnilo/prism-provider-google`'s `extra.cachedContent` on that package, or manage cache resources host-side).
|
|
63
64
|
- Pair with model-router residency allow-lists on `location`.
|
|
64
65
|
|
|
65
66
|
## Related APIs
|
package/docs/public-contracts.md
CHANGED
|
@@ -124,6 +124,7 @@ Important request shapes:
|
|
|
124
124
|
| `ContextResolutionContext` | Context provider input: messages plus optional session/run ids, metadata, and signal. |
|
|
125
125
|
| `InputAssemblyLayout` | Default input layout selector: `"cache_aware"` (default) or opt-in `"legacy"`. |
|
|
126
126
|
| `DefaultInputBuildContext` | Optional default input assembly context: input layout, instructions, history, summaries, attachments, explicit resources, tool results, middleware, ids, metadata, and signal. |
|
|
127
|
+
| `PromptBuildRequest` | Prompt-builder input: messages, context, selected skills, active tools, model, metadata, signal, and optional `inputLayout`; the default builder uses cache-aware ordering unless `legacy` is explicit. |
|
|
127
128
|
| `ResolveContextOptions` | Ordered context resolution input: selected providers, messages, ids, metadata, signal, and optional middleware. |
|
|
128
129
|
| `AssembleProviderInputOptions` | Provider input assembly input: model, input, optional builders, selected context providers/skills, active tools, metadata, and signal. |
|
|
129
130
|
| `PromptTemplateOptions` | Missing-variable behavior for tiny `renderPromptTemplate()` substitutions. |
|
|
@@ -148,7 +149,7 @@ Important request shapes:
|
|
|
148
149
|
| `CheckpointStore` | Generic versioned checkpoint capability: save/load/bounded-list/delete by namespace and key, with ownership, exact-version CAS, and lease fencing. `createMemoryCheckpointStore()` is the reference implementation; it is bounded — `maxRecords` (default 10,000, evicts least-recently-saved) and `maxValueBytes` (default 1 MiB per JSON value). |
|
|
149
150
|
| `LeaseStore` | Atomic acquire/renew/release/get by namespace and key, with opaque claim tokens, expiry, ownership scope, and monotonically increasing takeover fences. `createMemoryLeaseStore()` is the reference implementation. |
|
|
150
151
|
| `RunFeedbackStore` | Immutable append, bounded owned query, and owned deletion for ratings/comments/tags linked to existing run/trace/evaluation IDs. `createMemoryRunFeedbackStore()` is the reference implementation. |
|
|
151
|
-
| `EventMultiplexer<T>` | Generic bounded fan-in from async sources. `createEventMultiplexer()` owns queue limits, overflow policy, abort, source teardown, and close behavior. Single-consumer contract: a second concurrent `subscribe()` throws `EventMultiplexerError` (`ERR_PRISM_EVENT_MULTIPLEXER_SINGLE_CONSUMER`); the slot frees when the active consumer completes
|
|
152
|
+
| `EventMultiplexer<T>` | Generic bounded fan-in from async sources. `createEventMultiplexer()` owns queue limits, overflow policy, abort, source teardown, and close behavior. Graceful `close()` stops publishes/sources and drains already-queued events before the subscriber completes; overflow `close` still emits one notice and terminates. Single-consumer contract: a second concurrent `subscribe()` throws `EventMultiplexerError` (`ERR_PRISM_EVENT_MULTIPLEXER_SINGLE_CONSUMER`); the slot frees when the active consumer completes or is `return()`ed at a yield. `observe` fan-in is unchanged (broadcast happens at the source). |
|
|
152
153
|
| `PersistencePage<T>` | Cursor-paginated result page: `items`, optional `nextCursor`, optional `total`. |
|
|
153
154
|
| `PersistenceQuery` | Common pagination controls: `cursor?`, `limit?`, `order?: "asc" \| "desc"`. |
|
|
154
155
|
| `OwnershipScope` | Multi-tenant scope: `tenantId?`, `accountId?`, `userId?`. Included in records and queries. |
|
package/docs/rag.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
## What it does
|
|
4
4
|
|
|
5
|
-
`@arnilo/prism-rag` is an optional package for deterministic text/Markdown chunking, bounded embedding/vector indexing, atomic scoped source replacement/deletion, focused text/Markdown/HTML/PDF parsing, bounded reranking, ingestion status, attributable citations, content-trust metadata, and explicit `ContextProvider` injection. It reuses `Embedder` and `VectorStore` from `@arnilo/prism-memory`; Prism core input assembly is unchanged.
|
|
5
|
+
`@arnilo/prism-rag` is an optional package for deterministic text/Markdown chunking (with ATX heading-stack metadata), bounded embedding/vector indexing with embedder-identity drift guards, atomic scoped source replacement/deletion with content-hash skip and generation visibility, hybrid vector+lexical retrieval with reciprocal-rank fusion (one embed / one RRF / one rerank across one or many exact scopes), focused text/Markdown/HTML/PDF parsing, bounded reranking (host seam plus a TEI REST adapter), ingestion status, attributable citations, content-trust metadata, and explicit `ContextProvider` injection. It reuses `Embedder` and `VectorStore` from `@arnilo/prism-memory`; Prism core input assembly is unchanged.
|
|
6
6
|
|
|
7
7
|
## When to use it
|
|
8
8
|
|
|
@@ -18,7 +18,7 @@ Chunking:
|
|
|
18
18
|
| `chunkMarkdown(markdown, options)` | Same engine, preferring heading/paragraph boundaries |
|
|
19
19
|
| `sourceId` | Required stable, non-secret source identifier |
|
|
20
20
|
| `size` / `overlap` | Character ceiling and repeated context |
|
|
21
|
-
| `metadata` | JSON metadata copied to every chunk
|
|
21
|
+
| `metadata` | JSON metadata copied to every chunk; Markdown chunking additionally stamps `heading` (ordered parent-first heading stack, e.g. `["Policy", "3.2 Leave"]`) unless the caller supplies one.
|
|
22
22
|
|
|
23
23
|
Document lifecycle:
|
|
24
24
|
|
|
@@ -35,13 +35,18 @@ Index/retrieve:
|
|
|
35
35
|
| Field | Required | Meaning |
|
|
36
36
|
| --- | --- | --- |
|
|
37
37
|
| `embedder` / `store` | yes | Phase 7 `Embedder` and `VectorStore` |
|
|
38
|
-
| `scope` |
|
|
38
|
+
| `scope` / `scopes` | one or the other | Exact `{ tenantId, resourceId, corpusId }` (corpus → vector thread). `scope` is the single-corpus path; `scopes` is 0..`HARD_RETRIEVE_SCOPE_CAP` (8) exact scopes. Empty `scopes` returns no hits and does not embed/search/rerank. Passing both or neither throws. |
|
|
39
39
|
| `chunks` | indexing | `RagChunk[]` from package chunkers or compatible host parser |
|
|
40
|
-
| `topK` / `queryCandidates` | retrieval | Returned result count and bounded pre-filter candidates |
|
|
40
|
+
| `topK` / `queryCandidates` | retrieval | Returned result count and bounded pre-filter candidates (`queryCandidates` is **per scope**) |
|
|
41
|
+
| `lexical` | no | `"fts"` \| `"bm25"` \| `"off"` (default `"off"`); enables the lexical retrieval leg when the store advertises it |
|
|
42
|
+
| `fusion` / `rrfK` | no | `"rrf"` fusion of vector+lexical legs (default `"rrf"` when `lexical` is on; `rrfK` default 60, hard cap 1,000) |
|
|
41
43
|
| `filter` | no | Shallow JSON metadata equality filter |
|
|
42
44
|
| `reranker` | no | Host-owned `Reranker` receives redacted bounded `RagHit[]` and must return the same IDs once each, in preferred order. |
|
|
43
45
|
| `maxRerankBytes` / `maxRerankMs` / `rerankConcurrency` | no | Reranker caps; defaults/hard limits are 64/256 KiB, 2/10 s, and 2/8 active calls per reranker object. |
|
|
44
46
|
| `statusStore` | no | `IngestionStatusStore` records per-source pending/indexed/failed/partial byte/chunk progress; use `listIngestionStatus()` for capped exact-scope pages. |
|
|
47
|
+
| `contentHash` | no | Host-computed document digest; stamped on records and enables unchanged-source skip in `replaceSource` (`skipIfUnchanged`, default true when present) |
|
|
48
|
+
| `reuseEmbeddings` | no | `ReadonlyMap<string, ReusableEmbedding>` — chunk id → `{ text, embedding }`; embeddings reused (no embed call) when texts match |
|
|
49
|
+
| `telemetry` / `telemetryParent` | no | `RagTelemetry` seam (e.g. `createRagTelemetry()` from `@arnilo/prism-observability-opentelemetry`); spans nest under `telemetryParent` |
|
|
45
50
|
| `redactor` / `secrets` | no | Redact before embedding, persistence, reranking, and injection |
|
|
46
51
|
| `signal` | no | Abort embedding, vector operations, reranking, and batch progression |
|
|
47
52
|
|
|
@@ -51,7 +56,8 @@ Index/retrieve:
|
|
|
51
56
|
- `indexChunks()` returns `{ indexed, sourceIds }` after bounded batch upserts.
|
|
52
57
|
- `replaceSource()` / `deleteSource()` return `{ sourceId, deleted, indexed }`.
|
|
53
58
|
- `replaceDocument()` carries loader parser metadata into chunk metadata; the web loader preserves web-tools citation ID and `untrusted: true`.
|
|
54
|
-
- `retrieveContext()` returns `{ query, trust, text, hits, citations, truncated }`. Every hit/citation carries `{ provenance: { sourceId, chunkId, citationId, provider, retrieval: "vector", retrievedAt }, trust: { untrusted: true, inert: true, injectionCapable: true } }`; `retrievalRank` preserves pre-rerank order. Rendered text uses `[citation-id] text` blocks.
|
|
59
|
+
- `retrieveContext()` returns `{ query, trust, text, hits, citations, truncated }`. Every hit/citation carries `{ provenance: { sourceId, chunkId, citationId, provider, tenantId, resourceId, corpusId, retrieval: "vector" | "lexical" | "hybrid", retrievedAt }, trust: { untrusted: true, inert: true, injectionCapable: true } }`; `retrieval` labels the leg(s) that surfaced the hit after RRF fusion, and `retrievalRank` preserves pre-rerank order. Rendered text uses `[citation-id] text` blocks.
|
|
60
|
+
- `replaceSource()` returns `{ sourceId, deleted, indexed, skipped? }` (skipped when the stored `contentHash` matched and no writes occurred). Records carry `embedderId` (from `Embedder.id`, the Task 2 identity contract) and `generation` (scope-level monotonically bumped index per replacement; `_rag` metadata carries `contentHash` when supplied). `store.getCurrentGeneration(scope)` / `store.setCurrentGeneration(scope, n)` let hosts read and roll back the visible generation; retrieval filters to the current generation while legacy generation-less rows stay visible.
|
|
55
61
|
- `createMemoryIngestionStatusStore()` is a bounded in-memory reference adapter. `listIngestionStatus({ store, scope, limit, cursor })` returns capped status pages; hosts supply durable stores when status must survive process restart.
|
|
56
62
|
- `createRagContextProvider()` returns one ordinary context provider. Empty queries/results contribute no block.
|
|
57
63
|
- No events, tools, permissions, provider calls, loaders, or network requests are added.
|
|
@@ -94,7 +100,7 @@ await indexChunks({ chunks, embedder, store, scope, statusStore });
|
|
|
94
100
|
const found = await retrieveContext("approval policy", {
|
|
95
101
|
embedder,
|
|
96
102
|
store,
|
|
97
|
-
scope,
|
|
103
|
+
scopes: [scope], // or `scope` for one corpus
|
|
98
104
|
topK: 4,
|
|
99
105
|
filter: { category: "security" },
|
|
100
106
|
reranker: { rerank: async ({ hits }) => [...hits].sort((a, b) => b.score - a.score) },
|
|
@@ -109,11 +115,47 @@ const agent = createAgent({
|
|
|
109
115
|
console.log(found.text, await agent.createSession().run("How do approvals work?"));
|
|
110
116
|
```
|
|
111
117
|
|
|
118
|
+
Content-hash skip and hash validation:
|
|
119
|
+
|
|
120
|
+
```ts
|
|
121
|
+
import { isValidContentHash } from "@arnilo/prism-rag";
|
|
122
|
+
|
|
123
|
+
const digest = "ab12..."; // host-computed SHA-256 hex of the document
|
|
124
|
+
if (!isValidContentHash(digest)) throw new Error("invalid digest");
|
|
125
|
+
await replaceSource({ sourceId: "doc", chunks, embedder, store, scope, contentHash: digest }); // unchanged → skipped, zero embeds
|
|
126
|
+
```
|
|
127
|
+
|
|
128
|
+
Hybrid retrieval, TEI reranking, and telemetry:
|
|
129
|
+
|
|
130
|
+
```ts
|
|
131
|
+
import { createRagTelemetry } from "@arnilo/prism-observability-opentelemetry";
|
|
132
|
+
import { createTeiReranker } from "@arnilo/prism-rag";
|
|
133
|
+
|
|
134
|
+
const telemetry = createRagTelemetry({ tracer, meter }); // @opentelemetry/api instruments
|
|
135
|
+
const org = { tenantId: "t1", resourceId: "docs", corpusId: "org" };
|
|
136
|
+
const user = { tenantId: "t1", resourceId: "docs", corpusId: "user" };
|
|
137
|
+
const session = { tenantId: "t1", resourceId: "docs", corpusId: "session" };
|
|
138
|
+
const found = await retrieveContext("leave balance", {
|
|
139
|
+
embedder,
|
|
140
|
+
store, // a store that advertises lexicalModes: ["fts"]
|
|
141
|
+
scopes: [org, user, session], // one embed, per-scope legs, one RRF, one rerank
|
|
142
|
+
lexical: "fts",
|
|
143
|
+
topK: 8,
|
|
144
|
+
reranker: createTeiReranker({ baseUrl: "https://tei.svc:8080" }),
|
|
145
|
+
telemetry, // roots a rag_request span tree; attachSession/handleAgentEvent NOT required
|
|
146
|
+
});
|
|
147
|
+
```
|
|
148
|
+
|
|
112
149
|
## Extension and configuration notes
|
|
113
150
|
|
|
114
151
|
- Supply any Phase 7-conforming embedder/vector store, including the in-memory reference or PostgreSQL/pgvector adapter.
|
|
115
152
|
- Metadata filtering is package-local after a bounded candidate query so existing vector contracts/adapters remain unchanged. Increase `queryCandidates` only when selective filters measurably need it.
|
|
116
153
|
- `Reranker` is a host seam, not a provider integration. Return each redacted candidate ID exactly once; Prism retains canonical hit/provenance/trust fields and exposes `retrievalRank` for diagnostics. Add a hosted reranker only when a host owns its credentials, quota, and retry policy.
|
|
154
|
+
- `createTeiReranker({ baseUrl, model?, timeoutMs?, maxResponseBytes?, ssrf?, allowLoopback?, fetch? })` (`CreateTeiRerankerOptions`) adapts a Hugging Face TEI `POST <baseUrl>/rerank` endpoint (`{query, texts, raw_scores:false}` → `{results:[{index,score}]}`) into the `Reranker` seam. It returns a permutation-only reorder of the same hit objects, so provenance/trust move untouched. Response parsing is strict — short/duplicate/out-of-range indices, non-finite scores, HTTP errors, timeouts, and oversized bodies all fail closed; the `rerankHits` caps (`maxRerankBytes`, `maxRerankMs`, `rerankConcurrency`) still apply around it. The default transport is the core DNS-pinned `pinnedFetch` (redirect-free, byte-bounded to 65,536 by default); HTTPS is required unless `allowLoopback: true` (loopback dev/test) or the host supplies `ssrf`/`fetch` for cluster networking. The adapter validates URL shape only — SSRF policy enforcement stays host-side. No credentials are ever sent; there is no SaaS default URL.
|
|
155
|
+
- Hybrid retrieval: pass `lexical: "fts"` (or `"bm25"` when the store supports it) to `retrieveContext()`; the two legs are fused with reciprocal-rank fusion (`fusion: "rrf"`, `rrfK` 60 default; the pure helper `fuseReciprocalRank()` returns `FusedCandidate[]` for custom orchestration). Stores advertise support via `lexicalModes?: readonly LexicalMode[]` and `tokenizeLexical()` is the shared tokenizer. Each hit's provenance `retrieval` field reports `vector`/`lexical`/`hybrid`; fusion internals expose `RetrievalLeg`.
|
|
156
|
+
- Multi-scope retrieve: `scopes: RagScope[]` searches each exact scope against that scope's current generation, then runs **one** RRF over the union and **one** rerank. The query is embedded once. `queryCandidates` is per scope. Duplicate scopes are dropped. `HARD_RETRIEVE_SCOPE_CAP` is 8.
|
|
157
|
+
- Embedder identity/drift guard: `Embedder.id` (memory contract) is stamped onto every vector record as `embedderId`. `retrieveContext()` fails closed with `ERR_PRISM_RAG_EMBEDDER_MISMATCH` when a stored record's `embedderId` or dimensions differ from the active embedder (for example after a model change) — re-index the source before retrieving. Legacy records without an `embedderId` also fail closed, naming the re-index path.
|
|
158
|
+
- Generations: `replaceSource()` stamps a scope-level generation (auto-incremented per replacement) on staged records and the vector store filters retrieval to the current generation. `setCurrentGeneration()` supports rollback; stores without generation tracking keep legacy behavior (everything visible).
|
|
117
159
|
- `IngestionStatusStore` is optional observability storage. It is keyed by exact scope and source ID; use `listIngestionStatus()` rather than an unbounded corpus scan. The reference memory store is process-local; implement the same capped scope behavior for durable status.
|
|
118
160
|
- `createRagContextProvider()` derives its query from latest user text by default; pass a fixed string or callback for host-controlled query generation.
|
|
119
161
|
- `createResourceDocumentLoader({ loader })` calls one host-owned `ResourceLoader`; it scans nothing and performs no filesystem or network I/O itself. Pass the host's permission/trust context to that loader.
|
|
@@ -123,14 +165,19 @@ console.log(found.text, await agent.createSession().run("How do approvals work?"
|
|
|
123
165
|
|
|
124
166
|
## Security and performance notes
|
|
125
167
|
|
|
126
|
-
- Every index/query includes exact tenant/resource/corpus scope; returned records are rechecked and malformed/foreign records fail closed.
|
|
168
|
+
- Every index/query includes exact tenant/resource/corpus scope; returned records are rechecked and malformed/foreign records fail closed. `retrieveContext` accepts `scope` or `scopes` (never both, never neither). Empty `scopes` is the host “no allowed corpora” path — no embed, no search, no rerank. A hit whose stored scope is not in the requested list fails closed. Generation filters stay per scope.
|
|
169
|
+
- Embedding identity is a privacy/consistency boundary: records from a different embedder (or dimension) never silently mingle with new ones — retrieval fails closed and names the re-index path. Generation pointers are scope-scoped: a pointer row belongs to exactly one scope, and visibility is computed inside the store (SQL), never by post-filtering in JS.
|
|
127
170
|
- Source IDs become citation/storage IDs and must be stable non-secret identifiers. Text and user metadata can be redacted before external embedding and persistence.
|
|
171
|
+
- Heading metadata is document text only — it passes through the existing `maxMetadataBytes` cap as chunk metadata; no new content path is introduced.
|
|
172
|
+
- `contentHash` skip and `reuseEmbeddings` never leak embeddings: reused embeddings are keyed by chunk id within one replacement and only accepted when the stored text matches exactly.
|
|
128
173
|
- Retrieved documents are untrusted inert context. Prompt-injection text cannot activate tools, skills, credentials, permissions, or extensions.
|
|
129
174
|
- Remote sources must pass existing resource/media trust, SSRF, MIME, and byte policies before their decoded text reaches this package.
|
|
130
175
|
- `replaceSource()` stages every bounded embedding before opening the store transaction. It requires a source-aware transactional store and fails closed rather than pretending generic upserts are atomic. `createMemoryVectorStore()` supplies the reference `getBySource()` / transaction capability; durable stores must implement equivalent exact-scope behavior.
|
|
131
176
|
- `deleteSource()` rechecks every returned record's tenant/resource/corpus and source metadata before delete. Same source IDs in another corpus remain untouched.
|
|
132
177
|
- Parsers enforce byte/page/time caps, abort before and after parsing, decode UTF-8 strictly, and strip HTML script/style content. Parsed and retrieved text remains untrusted inert context; it never gains tool authority.
|
|
133
|
-
- Rerankers receive redacted input under byte/time/concurrency caps. Timeout, abort, unknown/duplicate/missing IDs, oversized input, and reranker failures fail closed; returned objects cannot overwrite Prism provenance/trust fields.
|
|
178
|
+
- Rerankers receive redacted input under byte/time/concurrency caps. Timeout, abort, unknown/duplicate/missing IDs, oversized input, and reranker failures fail closed; returned objects cannot overwrite Prism provenance/trust fields. The TEI adapter adds fail-closed response parsing (permutation completeness, finite scores) and honors the 65,536-byte response ceiling; SSRF/URL policy is host-side (see Extension notes).
|
|
179
|
+
- Telemetry is a host-owned seam: `RagTelemetry` adapter (`createRagTelemetry()`) drops anything outside a fixed span-name set and `rag.*`-shaped attribute keys, so raw chunk text never reaches the tracer unless the host's own `attributeFilter` opts it in; when the seam is absent, instrumentation costs nothing.
|
|
180
|
+
- Durable vector stores (PostgreSQL/pgvector path via `@arnilo/prism-memory`) run their DDL against the host's knowledge database — tables are created in a schema/table the host names (default `prism_memory.semantic_memory`), and hosts must own backup/retention of that database. See [Working and semantic memory](working-and-semantic-memory.md).
|
|
134
181
|
- Ingestion failure errors are redacted before status storage. Status reads reject foreign scope entries and page-limit violations; status itself creates no permission or tool authority.
|
|
135
182
|
- Filtering scans at most `queryCandidates` hits; rendering stops at top-K, UTF-8 result bytes, or estimated context-token ceiling.
|
|
136
183
|
|