@arnilo/prism 0.0.1 → 0.0.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (120) hide show
  1. package/CHANGELOG.md +4 -2
  2. package/README.md +17 -7
  3. package/dist/agent-definitions.d.ts +12 -0
  4. package/dist/agent-definitions.js +131 -0
  5. package/dist/agent-loops.d.ts +14 -0
  6. package/dist/agent-loops.js +161 -0
  7. package/dist/agents.js +263 -76
  8. package/dist/cache-helpers.d.ts +28 -0
  9. package/dist/cache-helpers.js +73 -0
  10. package/dist/cli-runner.d.ts +38 -2
  11. package/dist/cli-runner.js +167 -5
  12. package/dist/compaction.js +2 -0
  13. package/dist/config.js +47 -12
  14. package/dist/contracts.d.ts +581 -6
  15. package/dist/contracts.js +41 -1
  16. package/dist/contribution-parsing.d.ts +19 -0
  17. package/dist/contribution-parsing.js +124 -0
  18. package/dist/contributions.d.ts +13 -3
  19. package/dist/contributions.js +96 -20
  20. package/dist/extensions.js +3 -0
  21. package/dist/index.d.ts +19 -9
  22. package/dist/index.js +10 -4
  23. package/dist/input.d.ts +7 -1
  24. package/dist/input.js +52 -11
  25. package/dist/instruction-injection.d.ts +28 -0
  26. package/dist/instruction-injection.js +55 -0
  27. package/dist/manifests.d.ts +1 -1
  28. package/dist/manifests.js +3 -3
  29. package/dist/models.d.ts +4 -1
  30. package/dist/models.js +5 -2
  31. package/dist/node/agent-definitions.d.ts +98 -0
  32. package/dist/node/agent-definitions.js +389 -0
  33. package/dist/node/contribution-discovery.d.ts +17 -0
  34. package/dist/node/contribution-discovery.js +163 -0
  35. package/dist/node/instruction-injectors.d.ts +32 -0
  36. package/dist/node/instruction-injectors.js +72 -0
  37. package/dist/node/session-store-jsonl.d.ts +1 -1
  38. package/dist/node/session-store-jsonl.js +42 -4
  39. package/dist/node/system-project-prompts.d.ts +30 -0
  40. package/dist/node/system-project-prompts.js +53 -0
  41. package/dist/provider-events.d.ts +3 -1
  42. package/dist/provider-events.js +34 -0
  43. package/dist/provider-request-policy.js +15 -1
  44. package/dist/providers/openai-compatible.js +1 -1
  45. package/dist/providers.d.ts +6 -2
  46. package/dist/providers.js +15 -1
  47. package/dist/redaction.d.ts +2 -1
  48. package/dist/redaction.js +3 -0
  49. package/dist/registry-options.d.ts +5 -0
  50. package/dist/registry-options.js +5 -0
  51. package/dist/rpc.d.ts +6 -2
  52. package/dist/rpc.js +71 -13
  53. package/dist/session-stores.d.ts +3 -1
  54. package/dist/session-stores.js +67 -6
  55. package/dist/skills.d.ts +4 -1
  56. package/dist/skills.js +3 -1
  57. package/dist/system-prompts.js +6 -2
  58. package/dist/testing/compaction-conformance.d.ts +17 -0
  59. package/dist/testing/compaction-conformance.js +61 -0
  60. package/dist/testing/extension-conformance.d.ts +26 -0
  61. package/dist/testing/extension-conformance.js +55 -0
  62. package/dist/testing/provider-conformance.d.ts +7 -0
  63. package/dist/testing/provider-conformance.js +18 -31
  64. package/dist/testing/session-store-conformance.d.ts +20 -0
  65. package/dist/testing/session-store-conformance.js +92 -0
  66. package/dist/testing/tool-conformance.d.ts +39 -0
  67. package/dist/testing/tool-conformance.js +79 -0
  68. package/dist/tools.d.ts +7 -2
  69. package/dist/tools.js +50 -13
  70. package/docs/agent-definitions.md +251 -0
  71. package/docs/agent-events.md +199 -0
  72. package/docs/agent-loops.md +217 -0
  73. package/docs/agent-session-runtime.md +20 -8
  74. package/docs/cli-rpc.md +39 -4
  75. package/docs/compaction-and-retry.md +2 -2
  76. package/docs/compaction-conformance.md +76 -0
  77. package/docs/compaction-llm.md +6 -3
  78. package/docs/compaction-observational-memory.md +4 -4
  79. package/docs/configuration-and-manifests.md +6 -1
  80. package/docs/context-and-skills.md +79 -6
  81. package/docs/contribution-discovery.md +149 -0
  82. package/docs/contribution-registries.md +9 -6
  83. package/docs/credentials-and-redaction.md +2 -0
  84. package/docs/customization.md +191 -0
  85. package/docs/database-persistence.md +407 -0
  86. package/docs/extension-authoring.md +193 -0
  87. package/docs/extension-conformance.md +80 -0
  88. package/docs/extensions.md +6 -0
  89. package/docs/host-security.md +141 -0
  90. package/docs/index.md +40 -19
  91. package/docs/input-and-prompt-assembly.md +19 -3
  92. package/docs/instruction-injection.md +183 -0
  93. package/docs/migration.md +201 -0
  94. package/docs/model-registry.md +122 -0
  95. package/docs/node-jsonl-session-store.md +5 -4
  96. package/docs/performance.md +127 -0
  97. package/docs/provider-caching.md +206 -0
  98. package/docs/provider-conformance.md +32 -5
  99. package/docs/provider-layer.md +51 -11
  100. package/docs/provider-packages.md +65 -5
  101. package/docs/provider-request-policies.md +113 -0
  102. package/docs/providers/kimi.md +22 -0
  103. package/docs/providers/neuralwatt.md +388 -0
  104. package/docs/providers/openai-compatible.md +1 -0
  105. package/docs/providers/openai.md +21 -0
  106. package/docs/providers/opencode-go.md +31 -3
  107. package/docs/providers/openrouter.md +29 -0
  108. package/docs/providers/zai.md +17 -0
  109. package/docs/public-contracts.md +87 -12
  110. package/docs/release-and-install.md +76 -26
  111. package/docs/runs-and-usage.md +236 -0
  112. package/docs/session-store-conformance.md +78 -0
  113. package/docs/session-stores-and-branching.md +10 -6
  114. package/docs/session-stores.md +126 -0
  115. package/docs/settings-auth-trust-security.md +18 -4
  116. package/docs/structured-output.md +247 -0
  117. package/docs/system-prompts.md +104 -2
  118. package/docs/tool-conformance.md +87 -0
  119. package/docs/tools.md +64 -8
  120. package/package.json +35 -2
@@ -0,0 +1,127 @@
1
+ # Performance limits
2
+
3
+ ## What it does
4
+
5
+ This page states Prism runtime limits that keep slow consumers and long sessions from becoming unbounded memory or latency problems.
6
+
7
+ Current surfaces:
8
+
9
+ - `SubscribeOptions` for bounded live `AgentEvent` subscriber queues.
10
+ - `SessionStore.readBranchPath(query)` for branch reads that avoid full-session scans.
11
+ - `ProductionPersistenceStore` cursor queries for entries, events, runs, tool calls, and usage.
12
+ - JSONL and memory stores documented as development/local adapters, not production multi-writer stores.
13
+
14
+ ## When to use it
15
+
16
+ Use these limits when embedding Prism in a UI, API server, job worker, or multi-tenant app that may have slow event consumers or long-lived sessions.
17
+
18
+ Do not treat Prism's live event subscribers as a durable queue. Use `RunLedger` / database persistence for replay, audit, billing, and timelines.
19
+
20
+ ## Inputs / request
21
+
22
+ ```ts
23
+ import { createAgent, type SubscribeOptions } from "@arnilo/prism";
24
+
25
+ const options: SubscribeOptions = {
26
+ maxQueuedEvents: 256,
27
+ overflow: "close",
28
+ };
29
+
30
+ const events = session.subscribe(options);
31
+ ```
32
+
33
+ `SubscribeOptions` fields:
34
+
35
+ | Field | Default | Purpose |
36
+ | --- | --- | --- |
37
+ | `maxQueuedEvents` | `1024` | Maximum events queued for one subscriber while it is not awaiting `next()`. Values below `1` are clamped to `1`. |
38
+ | `overflow` | `"close"` | Overflow policy: `"close"`, `"drop_oldest"`, or `"drop_newest"`. |
39
+
40
+ ## Outputs / response / events
41
+
42
+ On default overflow, the affected subscriber receives one `event_subscriber_overflow` event and then finishes:
43
+
44
+ ```json
45
+ {
46
+ "type": "event_subscriber_overflow",
47
+ "sessionId": "session_1",
48
+ "droppedEvents": 257,
49
+ "maxQueuedEvents": 256,
50
+ "overflow": "close"
51
+ }
52
+ ```
53
+
54
+ `drop_oldest` keeps the newest queued events. `drop_newest` ignores incoming events while the queue is full. These policies are live-view policies only; they do not affect `RunLedger` writes or stored session entries.
55
+
56
+ ## Request/response example
57
+
58
+ ```json
59
+ {
60
+ "subscribe": { "maxQueuedEvents": 256, "overflow": "close" },
61
+ "store": "database-backed SessionStore with readBranchPath",
62
+ "eventLedger": "cursor-paginated by runId and sequence"
63
+ }
64
+ ```
65
+
66
+ ## Implementation example
67
+
68
+ ```ts
69
+ import { createAgent, createMockProvider, providerDone, providerTextDelta } from "@arnilo/prism";
70
+
71
+ const agent = createAgent({
72
+ model: { provider: "mock", model: "demo" },
73
+ provider: createMockProvider([providerTextDelta("Hello"), providerDone()]),
74
+ });
75
+
76
+ const session = agent.createSession();
77
+ const reader = (async () => {
78
+ for await (const event of session.subscribe({ maxQueuedEvents: 256, overflow: "close" })) {
79
+ if (event.type === "event_subscriber_overflow") break;
80
+ render(event);
81
+ }
82
+ })();
83
+
84
+ await session.run("Hi");
85
+ await reader;
86
+
87
+ function render(_event: unknown) {}
88
+ ```
89
+
90
+ For production branch reads, implement `SessionStore.readBranchPath` instead of loading every entry:
91
+
92
+ ```ts
93
+ const store = {
94
+ async append(entry, options) { /* transaction + parent/idempotency checks */ },
95
+ async list(sessionId) { /* development fallback only */ return []; },
96
+ async readBranchPath(query) {
97
+ // Use one ancestor query / recursive CTE and return a cursor page.
98
+ return { items: [], nextCursor: undefined };
99
+ },
100
+ };
101
+ ```
102
+
103
+ ## Extension and configuration notes
104
+
105
+ - `SubscribeOptions` is per subscriber. One slow UI can be closed or dropped without affecting other subscribers, the active run, ledger writes, or session storage.
106
+ - `RunLedger` remains the durable event/timeline surface. Hosts may batch inside their ledger adapter, but Prism awaits ledger writes at safe boundaries; preserve per-run event order before acknowledging a batch.
107
+ - Database-backed stores should implement `readBranchPath` and cursor-paginated `ProductionPersistenceStore` queries. Memory and JSONL stores intentionally use full-session/file reads.
108
+ - Cursor pagination should use indexed keys, not offsets: `(run_id, sequence)` for events, `(session_id, started_at, id)` for runs, `(run_id, recorded_at, id)` for usage, and `(session_id, timestamp, id)` for entries.
109
+ - Hosts own queue sizes, page-size caps, database indexes, connection pools, transaction timeouts, retention jobs, partitioning, and multi-process coordination.
110
+
111
+ ## Security and performance notes
112
+
113
+ - Overflow events contain only counts and policy, never message text, tool arguments, prompts, provider payloads, or credentials.
114
+ - Runtime event payloads can be large (`Message`, content deltas, tool results, summaries, artifact metadata). Size queues by events and keep payload size in mind.
115
+ - Live subscriber queues are bounded by default. Durable replay belongs to host storage.
116
+ - `SessionStore.list(sessionId)` is a full-session read. It is fine for memory/JSONL development stores, but production adapters should use `readBranchPath` for provider context and branch views.
117
+ - The JSONL store rereads/parses the file for validation/list/get and serializes appends only within one process. It has no cross-process lock, pagination, migrations, tenant isolation, or retention.
118
+ - Recommended database indexes: session id, run id, parent id, branch leaf id, timestamps, tenant/account/user, event type, entry kind, `(run_id, sequence)` for event timelines, and `(run_id, recorded_at, id)` for usage. Allocate event `sequence` per run for stable timeline pagination.
119
+
120
+ ## Related APIs
121
+
122
+ - [Agent events](agent-events.md): `SubscribeOptions` and `event_subscriber_overflow` event details.
123
+ - [Agent/session runtime](agent-session-runtime.md): `session.subscribe()` and runtime event flow.
124
+ - [Session stores](session-stores.md): `SessionStore.readBranchPath` and dev-vs-production branch reads.
125
+ - [Database persistence](database-persistence.md): cursor queries, reference schema, indexes, and event sequence guidance.
126
+ - [Runs and usage ledger](runs-and-usage.md): durable event, tool-call, and usage persistence.
127
+ - [Node JSONL session store](node-jsonl-session-store.md): development-only JSONL limits.
@@ -0,0 +1,206 @@
1
+ # Provider caching
2
+
3
+ ## What it does
4
+
5
+ Provider caching documents Prism's cache intent surface:
6
+
7
+ - `ProviderRequestOptions.cache?: PromptCacheHints` for structured, provider-agnostic cache hints.
8
+ - Legacy aliases `cacheKey` and `cacheRetention`, still supported for backwards compatibility.
9
+ - `PromptCacheBreakpoint` locations for reusable prompt regions.
10
+ - `ModelCacheCapabilities` for model/provider cache support metadata.
11
+ - Shared helpers: `sanitizeCacheKey`, `mapCacheRetention`, `applyCacheControl`, `cacheHitRate`, `cacheSavings`, and `cacheUsageReport`.
12
+
13
+ Cache hints are best-effort. They describe intent; providers decide whether their native API can use them. Prism does not guarantee cache hits.
14
+
15
+ ## When to use it
16
+
17
+ Use this page when a host or provider package needs to:
18
+
19
+ - Mark stable system prompts, tools, context, or messages as cacheable.
20
+ - Opt into cache-aware default input ordering so stable attachments/resources, summaries, and prior history form a reusable prefix before the current user turn.
21
+ - Carry a stable cache key across turns without putting provider-specific fields in core.
22
+ - Read `ModelConfig.cache` to decide whether to map hints to implicit caching, key-based caching, cache-control breakpoints, provider-specific caching, or no caching.
23
+ - Compute normalized cache diagnostics from `Usage.cacheReadTokens` / `Usage.cacheWriteTokens`, including providers that only report reads.
24
+
25
+ Do not use cache keys for credentials, bearer tokens, API keys, OAuth tokens, user secrets, or raw private prompts.
26
+
27
+ ## Inputs / request
28
+
29
+ ```ts
30
+ import type {
31
+ ModelCacheCapabilities,
32
+ PromptCacheBreakpoint,
33
+ PromptCacheHints,
34
+ ProviderRequestOptions,
35
+ } from "@arnilo/prism";
36
+ ```
37
+
38
+ | Type / field | Purpose |
39
+ | --- | --- |
40
+ | `PromptCacheHints.mode?: "auto" | "on" | "off"` | Host intent. Providers may ignore unsupported modes. |
41
+ | `PromptCacheHints.key?: string` | Stable, untrusted cache key. Sanitize before sending to provider APIs. |
42
+ | `PromptCacheHints.retention?: "none" | "short" | "long"` | Desired retention. `mapCacheRetention()` downgrades unsupported long retention. |
43
+ | `PromptCacheHints.breakpoints?: readonly PromptCacheBreakpoint[]` | Stable prompt locations to mark for cache-control style providers. |
44
+ | `PromptCacheBreakpoint.location` | `system_prompt`, `tools`, `stable_context`, `last_stable_message`, `last_user_message`, or `message_id`. |
45
+ | `PromptCacheBreakpoint.messageId?` | Required when `location: "message_id"`. |
46
+ | `PromptCacheBreakpoint.ttl?` | Generic `short` / `long` hint. Provider packages map to native TTL shape. |
47
+ | `ModelConfig.cache?: ModelCacheCapabilities` | Static model/provider cache support metadata. |
48
+
49
+ `ModelCacheCapabilities.kind` values are generic: `implicit`, `openai_key`, `cache_control`, `provider_specific`, or `none`. Core never branches on provider names; provider packages read the metadata and map it to native requests.
50
+
51
+ Legacy alias note: `cacheKey` maps to `cache.key`, and `cacheRetention` maps to `cache.retention`. When both are present, structured `cache.key` / `cache.retention` is the authoritative cache intent for providers that read structured hints; legacy fields remain for older adapters.
52
+
53
+ ## Outputs / response / events
54
+
55
+ Cache helpers return plain data:
56
+
57
+ | Helper | Output |
58
+ | --- | --- |
59
+ | `sanitizeCacheKey(value, maxLength)` | Safe key string or `undefined`. |
60
+ | `mapCacheRetention(retention, model)` | `"short"`, `"long"`, or `undefined`. |
61
+ | `applyCacheControl(messages, breakpoints, options)` | New message array with `cache_control: { type: "ephemeral" }` on selected message anchors. |
62
+ | `cacheHitRate(usage)` | Cached input ratio or `undefined`. |
63
+ | `cacheSavings(usage, model)` | Estimated read-token savings or `undefined` without pricing. |
64
+ | `cacheUsageReport(usage, model?)` | Normalized read/write tokens, hit rate, estimated savings, and currency when available; `undefined` when no usage is supplied. |
65
+
66
+ Provider events do not change. Cache accounting stays in normalized `Usage.cacheReadTokens` and `Usage.cacheWriteTokens`.
67
+
68
+ For stable-prefix payloads, set `inputLayout: "cache_aware"` on the default input builder, `assembleProviderInput()`, `AgentConfig`, or `RunOptions`. The default prompt builder already places context, selected skills, and tool declarations before input messages; cache-aware input ordering then places attachments/resources, summaries, prior history, and pending tool results before the current user suffix. The prefix is byte-stable only when those stable inputs are unchanged; Prism still does not guarantee provider cache hits.
69
+
70
+ ## Request/response example
71
+
72
+ ```json
73
+ {
74
+ "providerRequest.options": {
75
+ "sessionId": "sess_123",
76
+ "cache": {
77
+ "mode": "on",
78
+ "key": "sess_123",
79
+ "retention": "long",
80
+ "breakpoints": [
81
+ { "location": "system_prompt" },
82
+ { "location": "last_user_message" }
83
+ ]
84
+ }
85
+ },
86
+ "model.cache": {
87
+ "kind": "cache_control",
88
+ "maxBreakpoints": 4,
89
+ "minCacheableTokens": 1024,
90
+ "longRetention": true
91
+ }
92
+ }
93
+ ```
94
+
95
+ ## Implementation example
96
+
97
+ ```ts
98
+ import {
99
+ applyCacheControl,
100
+ cacheHitRate,
101
+ cacheUsageReport,
102
+ mapCacheRetention,
103
+ sanitizeCacheKey,
104
+ type ModelConfig,
105
+ type PromptCacheHints,
106
+ } from "@arnilo/prism";
107
+
108
+ const model: ModelConfig = {
109
+ provider: "demo",
110
+ model: "demo-large",
111
+ cache: { kind: "cache_control", maxBreakpoints: 4, longRetention: true },
112
+ };
113
+
114
+ const hints: PromptCacheHints = {
115
+ mode: "on",
116
+ key: "workspace:agent#1",
117
+ retention: "long",
118
+ breakpoints: [{ location: "system_prompt" }, { location: "last_user_message" }],
119
+ };
120
+
121
+ const key = sanitizeCacheKey(hints.key, model.cache?.maxKeyLength ?? 128);
122
+ const retention = mapCacheRetention(hints.retention, model);
123
+ const stamped = applyCacheControl(messages, hints.breakpoints ?? [], { maxBreakpoints: model.cache?.maxBreakpoints });
124
+ const hitRate = cacheHitRate({ inputTokens: 1000, cacheReadTokens: 800 });
125
+ const report = cacheUsageReport({ inputTokens: 1000, cacheReadTokens: 800 }, model);
126
+ // { cacheReadTokens: 800, cacheWriteTokens: 0, hitRate: 0.8, ... }
127
+
128
+ await session.run("Explain this", { inputLayout: "cache_aware" });
129
+ ```
130
+
131
+ ## Extension and configuration notes
132
+
133
+ Provider request policies can set `ProviderRequestOptions.cache` or the legacy `cacheKey` / `cacheRetention` aliases. Provider packages decide how to map hints to native payloads:
134
+
135
+ | `ModelCacheCapabilities.kind` | Typical mapping |
136
+ | --- | --- |
137
+ | `implicit` | No request mutation; provider caches automatically. |
138
+ | `openai_key` | Send sanitized cache key and mapped retention where supported. |
139
+ | `cache_control` | Use `applyCacheControl()` on provider-native message anchors. |
140
+ | `provider_specific` | Provider package uses `compat`/native options intentionally. |
141
+ | `none` | Do not send cache fields. |
142
+
143
+ ### Per-provider cache behavior
144
+
145
+ | Provider package | Cache kind | Explicit cache hints | Multi-turn reuse notes | Caveats |
146
+ | --- | --- | --- | --- | --- |
147
+ | `@arnilo/prism-provider-openai` | `openai_key` | Sends sanitized `prompt_cache_key`; `prompt_cache_retention: "24h"` only when the model declares `longRetention`. | Stable cache key + stable prefix can improve reuse. | Best-effort only; `"short"`/`"none"` omit retention. |
148
+ | `@arnilo/prism-provider-openrouter` | `cache_control` | Applies `cache_control` markers only to caller-selected `cache.breakpoints`; `"long"` may add `ttl: "1h"`. | Breakpoint-stable prefixes can be reused by upstream providers. | Best-effort only; no marker is added to every block. |
149
+ | `@arnilo/prism-provider-opencode-go` | route-specific | Sends sanitized `x-opencode-session`; Anthropic route applies selected `cache_control` breakpoints; OpenAI route sends none. | Session id + unchanged selected anchors can help route-native caches. | Best-effort and route-dependent. |
150
+ | `@arnilo/prism-provider-zai` | `implicit` | No explicit cache payload; GLM context caching is automatic. | Resend unchanged prior history for implicit context-cache reuse. | Best-effort only; cache options do not force hits. |
151
+ | `@arnilo/prism-provider-kimi` | implicit by default, optional `cache_control` | Default catalog models send no `cache_control`; hosts may opt in on Anthropic `/messages` models with `ModelConfig.cache.kind: "cache_control"`. | Keep selected Anthropic anchors and prior history stable. | Best-effort and model/route-dependent. |
152
+ | `@arnilo/prism-provider-neuralwatt` | `implicit` | No `cache_control`, `cacheKey`, `prompt_cache`, or `cacheRetention` payload; NeuralWatt vLLM prefix caching is automatic. | Full prior history must be resent unchanged with only the new turn appended; `inputLayout: "cache_aware"` keeps stable prefixes first. | Best-effort only; does not promise cache hits; `cacheRetention: "none"` disables Prism hints only, not the implicit backend prefix cache. |
153
+
154
+ Detailed first-party provider notes:
155
+
156
+ - OpenAI Responses (`@arnilo/prism-provider-openai`): `kind: "openai_key"`. Sanitizes/clamps `prompt_cache_key` to 64 chars; `"long"` retention maps to `prompt_cache_retention: "24h"` only when the model declares `cache.longRetention`; `"short"`/`"none"` omit the field. `input_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`.
157
+ - OpenAI-compatible Chat Completions adapter: minimal scope, sends no cache payload; see [OpenAI-compatible provider](providers/openai-compatible.md).
158
+ - OpenRouter (`@arnilo/prism-provider-openrouter`): `kind: "cache_control"`. Sanitizes/clamps `session_id`/`X-Session-Id` to 256 chars; applies Anthropic-style `cache_control` markers only to caller-selected `cache.breakpoints` (last content block of each selected message), not every block; `"long"` retention adds `ttl: "1h"` when the model allows it. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`.
159
+ - OpenCode Go (`@arnilo/prism-provider-opencode-go`): `x-opencode-session` from `cacheKey ?? sessionId`, sanitized to 128 chars; the Anthropic route applies `cache_control` markers only to selected breakpoints (`"long"` → `ttl: "1h"`), the OpenAI route sends none. OpenAI route maps `prompt_tokens_details.cached_tokens`/`cache_write_tokens`; Anthropic route maps `cache_read_input_tokens`/`cache_creation_input_tokens`.
160
+ - Z.AI (`@arnilo/prism-provider-zai`): `kind: "implicit"`. GLM context caching is automatic; sends no explicit cache payload regardless of cache options. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`.
161
+ - NeuralWatt (`@arnilo/prism-provider-neuralwatt`): `kind: "implicit"`. NeuralWatt prefix caching is automatic; sends no explicit cache payload regardless of cache options. `cacheRetention: "none"` disables Prism cache-control hints only (not the implicit backend prefix cache). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`; NeuralWatt does not report a cache-write token, so `Usage.cacheWriteTokens` is never fabricated (stays `undefined`). NeuralWatt's `/v1/models` catalog advertises exact `cached_input_per_million` rates for cache reads and `cached_output_per_million: null`; static curated aliases do not guess those prices.
162
+ - Kimi (`@arnilo/prism-provider-kimi`): default catalog models use implicit caching (no `cache_control`); hosts opt in via `ModelConfig.cache.kind: "cache_control"` on the Anthropic `/messages` route, then `cache_control` markers apply only to selected breakpoints (`"long"` → `ttl: "1h"`); the Moonshot OpenAI route sends none. `cache_read_input_tokens`/`cache_creation_input_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`.
163
+
164
+ ### NeuralWatt cache-aware limiter
165
+
166
+ NeuralWatt (`@arnilo/prism-provider-neuralwatt`) runs a cache-aware backend rate
167
+ limiter on top of its implicit vLLM prefix cache. This shapes long-running agent
168
+ sessions differently from one-shot chat:
169
+
170
+ - **Uncached TPM counts cold prefill only.** The tokens-per-minute budget charges the
171
+ prefix that is not already cached. A request whose prefix is fully cached consumes
172
+ far less TPM than a cold request of the same total prompt length.
173
+ - **Warm-prefix requests can avoid some `503` fleet-capacity blocks.** Near fleet
174
+ capacity, requests that reuse a cached prefix are more likely to be admitted than
175
+ fully cold requests. Prefix reuse is both an availability and a latency lever.
176
+ - **Full prior history is required for multi-turn cache reuse.** The prefix cache is
177
+ keyed by request content, so each follow-up turn must resend the entire prior
178
+ transcript (system prompt + all prior turns) unchanged, with only the new turn
179
+ appended. Use `inputLayout: "cache_aware"` so Prism keeps the stable prefix first.
180
+ - Cache behavior is best-effort and **does not guarantee cache hits**. Admission and
181
+ eviction are server-side decisions and vary with fleet load. `cacheRetention:
182
+ "none"` disables Prism cache-control hints only; it does not disable the implicit
183
+ backend prefix cache.
184
+
185
+ See [NeuralWatt provider](providers/neuralwatt.md) for the package-level cache,
186
+ usage, and retry details.
187
+
188
+ ## Security and performance notes
189
+
190
+ - Cache hints are best-effort and do not guarantee cache hits.
191
+ - Cache keys are untrusted input; sanitize and truncate with `sanitizeCacheKey()` before provider I/O.
192
+ - Cache keys must never be credentials or secrets.
193
+ - Provider-owned auth/session/security headers always win over caller headers.
194
+ - Helpers are pure, network-free, and O(messages) at most. `cacheUsageReport()` is O(1).
195
+ - Cache-aware input ordering does not change resource loading: URI attachments/resources still load only through the caller-provided `ResourceLoader`.
196
+ - Cache usage reports contain only usage counts and optional pricing/currency; they do not include prompt text, cache keys, headers, credentials, or provider payloads.
197
+ - `applyCacheControl()` returns new message objects for stamped anchors and does not mutate input messages.
198
+
199
+ ## Related APIs
200
+
201
+ - [Input and prompt assembly](input-and-prompt-assembly.md): opt-in cache-aware ordering for stable provider payload prefixes.
202
+ - [Provider request policies](provider-request-policies.md): set cache hints before provider calls.
203
+ - [Model registry](model-registry.md): register `ModelConfig.cache` capability metadata.
204
+ - [Provider layer](provider-layer.md): provider/model registries and provider events.
205
+ - [Provider packages](provider-packages.md): package-owned mapping to provider-native cache APIs.
206
+ - [Public contracts](public-contracts.md): public type list for cache contracts and helpers.
@@ -12,14 +12,17 @@ Exported from `@arnilo/prism/testing/provider-conformance`:
12
12
  - `assertToolCallDeltasReconstruct(events, expected)`
13
13
  - `assertUsageAccounting(events, expected)`
14
14
  - `assertSerializedRequestCoversContent(request, body, options?)`
15
+ - `assertProviderOwnedHeadersWin(captured, options)`
15
16
  - `assertNoSecretLeak(events, secrets)`
16
17
 
17
18
  ## When to use it
18
19
 
19
- Use these helpers in provider package tests to check event order, terminal events, abort propagation, streamed tool-call deltas, usage/cache accounting, request body content preservation, and secret redaction.
20
+ Use these helpers in provider package tests to check event order, terminal events, abort propagation via `ProviderRequest.signal`, streamed tool-call deltas, usage/cache accounting, request body content preservation, protected header ownership, and secret redaction. Do not treat deprecated `ProviderRequestOptions.timeoutMs`/`maxRetries`/`maxRetryDelayMs` as conformance requirements; first-party providers use runtime abort signals and `AgentConfig.retry`/`RunOptions.retry` instead.
20
21
 
21
22
  Do not use them as a live integration runner, provider simulator, retry framework, credential loader, or test framework replacement.
22
23
 
24
+ For real network smoke tests, each first-party provider package ships an env-gated `src/__tests__/live.test.ts` that exercises the live API when `PRISM_LIVE_PROVIDER_TESTS=1` and a provider-specific API key are set. These live tests reuse the same conformance helpers (`assertProviderStreamConforms`, `assertAbortIsObserved`, `assertNoSecretLeak`) against the real provider, so offline and live assertions stay consistent. The default `npm test` never sets these gates and stays network-free; see [Release and install](release-and-install.md) for the full env-var list.
25
+
23
26
  ## Inputs / request
24
27
 
25
28
  ```ts
@@ -41,10 +44,11 @@ Helpers accept normal `AIProvider`, `ProviderRequest`, `ProviderEvent`, `Usage`,
41
44
 
42
45
  - `collectProviderEvents()` returns provider events in stream order.
43
46
  - `assertProviderStreamConforms()` returns collected events after verifying the stream ends with `done` or `error`, terminal events are last, and optional text/usage expectations match.
44
- - `assertAbortIsObserved()` passes an already-aborted signal and expects provider generation to reject.
45
- - `assertToolCallDeltasReconstruct()` rebuilds streamed `tool_call_delta` fragments into tool calls and validates expected id/name/arguments.
46
- - `assertUsageAccounting()` finds `usage` or `done.usage` and checks selected token fields including `cacheReadTokens` and `cacheWriteTokens`.
47
- - `assertSerializedRequestCoversContent()` scans a serialized provider request body for primitive canaries from each Prism content block and fails if any supported block type is silently dropped.
47
+ - `assertAbortIsObserved()` passes an already-aborted signal and expects provider generation to reject. This is the supported timeout primitive; use a host abort controller or `RunOptions.signal` rather than deprecated provider-level `timeoutMs`.
48
+ - `assertToolCallDeltasReconstruct()` rebuilds streamed `tool_call_delta` fragments into tool calls and validates expected id/name/arguments. The runtime uses the same reconstruction behavior before tool execution when a provider streams deltas.
49
+ - `assertUsageAccounting()` finds `usage` or `done.usage` and checks selected token fields including `cacheReadTokens` and `cacheWriteTokens`. This is the provider-neutral check for normalized cache read/write token extraction; every first-party provider package exercises it against server-specific fields (`cached_tokens`, `cache_read_input_tokens`, etc.).
50
+ - `assertSerializedRequestCoversContent()` scans a serialized provider request body for primitive canaries from each Prism content block and fails if any supported block type is silently dropped. Provider-valid transcripts place assistant `tool_call` messages before matching role `tool` `tool_result` messages; runtime, cache-aware input layout, and observational-memory worker loops preserve that order before serialization.
51
+ - `assertProviderOwnedHeadersWin()` compares captured request headers against the provider's authoritative owned header values and a caller-supplied header bag; it fails if any owned header (`authorization`, `content-type`, session/security headers) was overridden by caller headers, and also fails if a non-owned caller header was dropped. This is the provider-neutral check that caller `ProviderRequest.options.headers` cannot hijack provider credentials or sessions; every first-party provider package exercises it.
48
52
  - `assertNoSecretLeak()` stringifies all collected events and fails if any known secret string is present.
49
53
 
50
54
  ## Request/response example
@@ -56,6 +60,18 @@ Helpers accept normal `AIProvider`, `ProviderRequest`, `ProviderEvent`, `Usage`,
56
60
  }
57
61
  ```
58
62
 
63
+ Tool-call delta reconstruction example:
64
+
65
+ ```ts
66
+ import { assertToolCallDeltasReconstruct } from "@arnilo/prism/testing/provider-conformance";
67
+
68
+ assertToolCallDeltasReconstruct([
69
+ { type: "tool_call_delta", index: 0, id: "call_1", name: "lookup", argumentsText: "{\"q\":" },
70
+ { type: "tool_call_delta", index: 0, argumentsText: "\"prism\"}" },
71
+ { type: "done" },
72
+ ], [{ index: 0, id: "call_1", name: "lookup", arguments: { q: "prism" } }]);
73
+ ```
74
+
59
75
  Content-preservation example:
60
76
 
61
77
  ```ts
@@ -77,6 +93,17 @@ const body = JSON.parse(String(fetchInit.body));
77
93
  assertSerializedRequestCoversContent(request, body, { unsupported: ["image"] });
78
94
  ```
79
95
 
96
+ Protected header-ownership example:
97
+
98
+ ```ts
99
+ import { assertProviderOwnedHeadersWin } from "@arnilo/prism/testing/provider-conformance";
100
+
101
+ assertProviderOwnedHeadersWin(capturedHeaders, {
102
+ owned: { authorization: "Bearer provider-key", "content-type": "application/json" },
103
+ caller: { authorization: "Bearer attacker", "content-type": "text/plain", "x-caller": "kept" },
104
+ });
105
+ ```
106
+
80
107
  ## Implementation example
81
108
 
82
109
  ```ts
@@ -5,9 +5,10 @@
5
5
  The provider layer contains the small runtime pieces Prism already ships for host-owned model access:
6
6
 
7
7
  - `createProviderRegistry()` / `ProviderRegistry`: register and resolve `AIProvider` instances by id.
8
- - `createModelRegistry()` / `ModelRegistry`: register and resolve `ModelConfig` values by provider/model key.
9
- - `ModelConfig` metadata fields for display names, capabilities, limits, cost/cache pricing, opaque provider compat data, and host metadata.
10
- - Provider event helpers: create normalized `ProviderEvent` values for text, thinking, tool calls, usage, done, and errors, including optional cache read/write usage fields.
8
+ - `createProviderResolver()` / `ProviderResolver`: build a resolver that maps a `ModelConfig` to an `AIProvider` (or `undefined`), from a `ProviderRegistry` or a plain `AIProvider[]`.
9
+ - `createModelRegistry()` / `ModelRegistry`: register and resolve `ModelConfig` values by provider/model key. See [Model registry](model-registry.md).
10
+ - `ModelConfig` metadata fields for display names, capabilities, limits, cost/cache pricing, cache support metadata, opaque provider compat data, and host metadata. See [Provider caching](provider-caching.md).
11
+ - Provider event helpers: create normalized `ProviderEvent` values for text, thinking, streamed tool-call deltas, final tool calls, usage, done, and errors, including optional cache read/write usage fields.
11
12
  - `toolCallContent()`: create a `ToolCallContent` block.
12
13
  - `createMockProvider()` / `MockProviderOptions`: create a deterministic scripted `AIProvider` for tests and examples.
13
14
  - `@arnilo/prism/testing/provider-conformance`: optional network-free assertion helpers for provider adapter tests.
@@ -23,21 +24,21 @@ Use this layer when a host app, extension package, or test needs to:
23
24
  - Emit provider events without hand-writing event objects.
24
25
  - Test agent/provider flows without timers, credentials, SDKs, or network calls.
25
26
 
26
- Do not use this layer for credential storage, settings loading, tool dispatch, agent loops, package discovery, cache stores, or provider SDK configuration. Those stay host-owned or belong to provider packages.
27
+ Do not use this layer for credential storage, settings loading, tool dispatch, agent loops, package discovery, cache stores, or provider SDK configuration. Those stay host-owned or belong to provider packages. Request option hooks are covered in [Provider request policies](provider-request-policies.md).
27
28
 
28
29
  ## Inputs / request
29
30
 
30
31
  ### Provider registry
31
32
 
32
33
  ```ts
33
- createProviderRegistry(providers?: readonly AIProvider[]): ProviderRegistry
34
+ createProviderRegistry(providers?: readonly AIProvider[], options?: { duplicate?: "replace" | "error" }): ProviderRegistry
34
35
  ```
35
36
 
36
37
  `ProviderRegistry` methods:
37
38
 
38
39
  | Method | Input | Result |
39
40
  | --- | --- | --- |
40
- | `register(provider)` | `AIProvider` | Stores provider by `provider.id`. |
41
+ | `register(provider)` | `AIProvider` | Stores/replaces provider by `provider.id`; throws `Duplicate provider: <id>` when `duplicate: "error"`. |
41
42
  | `get(id)` | provider id string | Returns provider or `undefined`. |
42
43
  | `resolve(model)` | provider id string or `{ provider: string }` | Returns provider or throws `Unknown provider: <id>`. |
43
44
  | `list()` | none | Returns registered providers in insertion order. |
@@ -45,14 +46,14 @@ createProviderRegistry(providers?: readonly AIProvider[]): ProviderRegistry
45
46
  ### Model registry
46
47
 
47
48
  ```ts
48
- createModelRegistry(models?: readonly ModelConfig[]): ModelRegistry
49
+ createModelRegistry(models?: readonly ModelConfig[], options?: { duplicate?: "replace" | "error" }): ModelRegistry
49
50
  ```
50
51
 
51
52
  `ModelRegistry` methods:
52
53
 
53
54
  | Method | Input | Result |
54
55
  | --- | --- | --- |
55
- | `register(model)` | `ModelConfig` | Stores model by provider/model key, preserving inert metadata. |
56
+ | `register(model)` | `ModelConfig` | Stores/replaces model by provider/model key, preserving inert metadata; throws `Duplicate model: <provider>/<model>` when `duplicate: "error"`. |
56
57
  | `get(provider, model)` | provider id and model id | Returns model config or `undefined`. |
57
58
  | `resolve(provider, model)` | provider id and model id | Returns model config or throws `Unknown model: <provider>/<model>`. |
58
59
  | `list()` | none | Returns registered model configs in insertion order. |
@@ -71,6 +72,8 @@ providerError(error: unknown, secrets?: readonly (string | undefined)[]): Provid
71
72
  toolCallContent(id: string, name: string, args?: JsonObject): ToolCallContent
72
73
  ```
73
74
 
75
+ `tool_call_delta` fragments use the same `{ index, id?, name?, argumentsText? }` shape as live `message_delta` content. Runtime reconstructs final `ToolCallContent` before tool execution; conformance tests use the same reconstruction rules.
76
+
74
77
  ### Mock provider
75
78
 
76
79
  ```ts
@@ -84,13 +87,49 @@ createMockProvider(events?: readonly ProviderEvent[], options?: MockProviderOpti
84
87
  | `id` | `string` | Optional provider id. Defaults to `mock`. |
85
88
  | `onRequest` | `(request: ProviderRequest) => void` | Optional request observer for tests. |
86
89
 
90
+ ### Provider resolver
91
+
92
+ ```ts
93
+ export type ProviderResolver = (model: ModelConfig) => AIProvider | undefined;
94
+
95
+ createProviderResolver(source: ProviderRegistry | readonly AIProvider[]): ProviderResolver
96
+ ```
97
+
98
+ A `ProviderResolver` maps a `ModelConfig` to an `AIProvider` for the current
99
+ run. `createProviderResolver()` builds one from a `ProviderRegistry` (reuses
100
+ `ProviderRegistry.get`) or a plain `AIProvider[]` (builds an id-keyed lookup
101
+ once at construction; last duplicate id wins). A custom function is the
102
+ zero-helper path for hosts with their own provider map (lazy construction,
103
+ per-request routing).
104
+
105
+ The resolver returns `undefined` on a miss; the agent runtime fails closed with
106
+ `Unknown provider: ${model.provider}` before any provider turn (see
107
+ [Agent/session runtime](agent-session-runtime.md)).
108
+
109
+ Wire `providerSource` on `AgentConfig`, override per run with
110
+ `RunOptions.providerSource` (RunOptions wins). When `AgentConfig.provider` is
111
+ set it takes first precedence and the resolver is bypassed. The resolver is
112
+ called once per run with `options.model ?? config.model`; per-turn
113
+ re-resolution is unnecessary.
114
+
115
+ ```ts
116
+ import { createAgent, createProviderResolver, createProviderRegistry } from "@arnilo/prism";
117
+
118
+ const own = createMyProvider();
119
+ const providerSource = createProviderResolver(createProviderRegistry([own]));
120
+ // or mix first-party + own in one list:
121
+ // const providerSource = createProviderResolver([firstPartyProvider, own]);
122
+
123
+ const agent = createAgent({ model: { provider: own.id, model: "demo" }, providerSource });
124
+ ```
125
+
87
126
  ## Outputs / response / events
88
127
 
89
128
  - Registry `resolve()` returns the matching provider/model or throws before any provider `generate()` call.
90
129
  - Provider event helpers return plain `ProviderEvent` objects.
91
130
  - `providerError()` converts unknown errors to redacted `ErrorInfo` through `errorToErrorInfo()` and preserves safe string/number `code` fields for retry classification.
92
131
  - `createMockProvider()` returns an `AIProvider` whose `generate()` yields the scripted events in order and checks `request.signal?.aborted` before each event.
93
- - The agent/session runtime passes its per-run abort signal as `ProviderRequest.signal`.
132
+ - The agent/session runtime passes its per-run abort signal as `ProviderRequest.signal`. `ProviderRequestOptions.timeoutMs`, `maxRetries`, and `maxRetryDelayMs` are deprecated inert hints in first-party providers; use `RunOptions.signal`/host abort controllers for timeouts and `AgentConfig.retry`/`RunOptions.retry` for retry.
94
133
 
95
134
  ## Request/response example
96
135
 
@@ -147,6 +186,7 @@ for await (const event of resolvedProvider.generate({
147
186
  ## Extension and configuration notes
148
187
 
149
188
  - Registries are explicit objects returned by factories. Prism does not create a hidden global provider/model registry.
189
+ - Default duplicate policy is `"replace"` for compatibility. Hosts that load third-party provider/model contributions can pass `duplicate: "error"` to reject silent shadowing.
150
190
  - Extension packages can contribute `AIProvider` and `ModelConfig` values by registering them with host-owned registries.
151
191
  - Model resolution and provider resolution are separate on purpose: hosts can validate a model exists before selecting a provider.
152
192
  - Credential resolvers stay outside these registries; pass credentials directly to the provider adapter or runtime edge that needs them.
@@ -154,7 +194,7 @@ for await (const event of resolvedProvider.generate({
154
194
 
155
195
  ## Security and performance notes
156
196
 
157
- - Provider/model registries are `Map`-backed and perform O(1) lookup.
197
+ - Provider/model registries are `Map`-backed and perform O(1) lookup. Strict duplicate mode adds one O(1) `Map.has()` check during registration only.
158
198
  - Registries store providers and model metadata only. Do not store API keys, credential resolvers, headers, tokens, or secret-bearing settings in them.
159
199
  - Unknown provider/model resolution fails before provider execution or network I/O.
160
200
  - `createMockProvider()` uses scripted events only: no timers, credentials, SDKs, or network.
@@ -164,7 +204,7 @@ for await (const event of resolvedProvider.generate({
164
204
 
165
205
  ## Related APIs
166
206
 
167
- - [Agent/session runtime](agent-session-runtime.md): passes abort signals to providers, maps provider errors to session `error` events, and can retry configured transient provider-turn failures before output.
207
+ - [Agent/session runtime](agent-session-runtime.md): passes abort signals to providers, maps provider errors to session `error` events, and can retry configured transient provider-turn failures before output; this is the supported replacement for deprecated provider-level retry options.
168
208
  - [Provider packages](provider-packages.md): explicit package primitive for registering providers, models, auth descriptors, request/cache policies, and prompt contributions.
169
209
  - [Public contracts](public-contracts.md): `AIProvider`, `ProviderRequest`, `ProviderEvent`, `ModelConfig`, `Usage`, and content/tool-call contracts.
170
210
  - [Credentials and redaction](credentials-and-redaction.md): credential and redaction helpers used by provider adapters.