@arnilo/prism 0.0.4 → 0.0.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (97) hide show
  1. package/CHANGELOG.md +46 -1
  2. package/README.md +34 -10
  3. package/dist/agent-loops.d.ts +1 -0
  4. package/dist/agent-loops.js +26 -16
  5. package/dist/agents.js +147 -21
  6. package/dist/cli-init.d.ts +41 -0
  7. package/dist/cli-init.js +390 -0
  8. package/dist/cli-runner.d.ts +7 -1
  9. package/dist/cli-runner.js +13 -1
  10. package/dist/content.d.ts +19 -0
  11. package/dist/content.js +197 -69
  12. package/dist/contracts.d.ts +96 -9
  13. package/dist/contracts.js +8 -0
  14. package/dist/feedback.d.ts +48 -0
  15. package/dist/feedback.js +230 -0
  16. package/dist/ids.d.ts +2 -0
  17. package/dist/ids.js +6 -0
  18. package/dist/index.d.ts +10 -4
  19. package/dist/index.js +6 -3
  20. package/dist/providers/media.d.ts +3 -1
  21. package/dist/providers/media.js +11 -1
  22. package/dist/session-stores.js +2 -3
  23. package/dist/testing/feedback.d.ts +6 -0
  24. package/dist/testing/feedback.js +37 -0
  25. package/dist/testing/persistence-schema.d.ts +48 -10
  26. package/dist/testing/persistence-schema.js +166 -22
  27. package/dist/testing/run-ledger-conformance.js +7 -1
  28. package/dist/thinking.d.ts +42 -0
  29. package/dist/thinking.js +92 -0
  30. package/dist/tools.js +2 -3
  31. package/dist/use-case-model.d.ts +63 -0
  32. package/dist/use-case-model.js +52 -0
  33. package/docs/a2a.md +75 -0
  34. package/docs/agent-events.md +14 -21
  35. package/docs/agent-loops.md +12 -9
  36. package/docs/agent-session-runtime.md +14 -16
  37. package/docs/cli-rpc.md +35 -7
  38. package/docs/coding-agent-tools.md +35 -14
  39. package/docs/coding-security.md +7 -3
  40. package/docs/compaction-llm.md +17 -7
  41. package/docs/compaction-observational-memory.md +30 -4
  42. package/docs/context-and-skills.md +1 -0
  43. package/docs/credential-storage.md +58 -9
  44. package/docs/credentials-and-redaction.md +3 -3
  45. package/docs/database-persistence.md +17 -9
  46. package/docs/evaluations.md +122 -0
  47. package/docs/extensions.md +2 -2
  48. package/docs/host-security.md +26 -5
  49. package/docs/index.md +43 -28
  50. package/docs/mcp-tools.md +74 -13
  51. package/docs/migration.md +177 -3
  52. package/docs/multimodal-content.md +14 -6
  53. package/docs/node-filesystem-config.md +1 -0
  54. package/docs/node-jsonl-session-store.md +5 -4
  55. package/docs/observability.md +14 -6
  56. package/docs/performance.md +209 -0
  57. package/docs/postgres-persistence.md +8 -6
  58. package/docs/provider-caching.md +16 -4
  59. package/docs/provider-conformance.md +40 -1
  60. package/docs/provider-packages.md +62 -3
  61. package/docs/providers/ai-sdk.md +149 -0
  62. package/docs/providers/kimi.md +124 -61
  63. package/docs/providers/neuralwatt.md +19 -13
  64. package/docs/providers/openai.md +56 -13
  65. package/docs/providers/opencode-go.md +118 -30
  66. package/docs/providers/openrouter.md +105 -35
  67. package/docs/providers/zai.md +94 -45
  68. package/docs/public-contracts.md +6 -5
  69. package/docs/rag.md +113 -0
  70. package/docs/release-and-install.md +100 -79
  71. package/docs/review-coverage-2026-07-15.md +193 -0
  72. package/docs/review-coverage-2026-07-17-provider-validation.md +192 -0
  73. package/docs/runs-and-usage.md +42 -5
  74. package/docs/server.md +139 -0
  75. package/docs/settings-auth-trust-security.md +5 -5
  76. package/docs/sqlite-persistence.md +6 -5
  77. package/docs/structured-output.md +1 -1
  78. package/docs/supervisors.md +71 -0
  79. package/docs/thinking-and-reasoning.md +98 -0
  80. package/docs/tool-execution-primitives.md +3 -3
  81. package/docs/tools.md +15 -0
  82. package/docs/use-case-model-selection.md +109 -0
  83. package/docs/workflow-orchestration-primitives.md +20 -3
  84. package/docs/workflows.md +114 -33
  85. package/docs/working-and-semantic-memory.md +170 -0
  86. package/package.json +13 -3
  87. package/templates/init/README.md.tmpl +28 -0
  88. package/templates/init/env.example.tmpl +1 -0
  89. package/templates/init/gitignore.tmpl +11 -0
  90. package/templates/init/optional/evals-example.ts.tmpl +17 -0
  91. package/templates/init/optional/workflows-example.ts.tmpl +27 -0
  92. package/templates/init/package.json.tmpl +22 -0
  93. package/templates/init/providers.json +76 -0
  94. package/templates/init/src/agent.ts.tmpl +10 -0
  95. package/templates/init/src/index.ts.tmpl +12 -0
  96. package/templates/init/src/tests/agent.test.ts.tmpl +24 -0
  97. package/templates/init/tsconfig.json.tmpl +15 -0
@@ -1,62 +1,85 @@
1
- # ZAI provider package
1
+ # Z.AI provider package
2
2
 
3
3
  ## What it does
4
4
 
5
- `@arnilo/prism-provider-zai` provides explicit, side-effect-free setup for the ZAI GLM
6
- API-key provider using Prism's OpenAI-compatible route with the
7
- `thinkingFormat: "zai"` model-compat setting, developer-role fallback, and
8
- GLM tool-stream quirks.
5
+ `@arnilo/prism-provider-zai` provides explicit, side-effect-free setup for the Z.AI
6
+ GLM Chat Completions API (`POST /chat/completions`) with official deep-thinking,
7
+ reasoning-effort, and tool-stream request fields.
9
8
 
10
- The package registers a provider, default model metadata, and an `api_key` auth
11
- method through `createExtensionKernel().load([...])`.
9
+ The package registers a provider, featured GLM model metadata, and an `api_key`
10
+ auth method through `createExtensionKernel().load([...])`.
12
11
 
13
12
  ## When to use it
14
13
 
15
- Use it when a host app wants to run the ZAI GLM endpoint through Prism's
16
- `AgentSession` runtime with ZAI-specific thinking/tool-stream handling.
14
+ Use it when a host app wants to run Z.AI GLM models through Prism's
15
+ `AgentSession` runtime with official `thinking` / `reasoning_effort` /
16
+ `tool_stream` mapping and implicit context caching.
17
17
 
18
- Do not use it for automatic credential discovery, catalog fetches, or
18
+ Do not use it for automatic credential discovery, setup-time catalog fetches, or
19
19
  real-network tests.
20
20
 
21
21
  ## Inputs / request
22
22
 
23
23
  ```ts
24
- import { createZaiProviderPackage } from "@arnilo/prism-provider-zai";
24
+ import {
25
+ createZaiProviderPackage,
26
+ defineZaiModel,
27
+ listZaiModels,
28
+ } from "@arnilo/prism-provider-zai";
25
29
 
26
30
  createZaiProviderPackage(options: ZaiProviderPackageOptions): ProviderPackage
27
- defineZaiModel(config: ZaiModelConfig): ZaiModelConfig
31
+ defineZaiModel(config: ZaiModelConfig): ModelConfig
32
+ listZaiModels(options?: ListZaiModelsOptions): Promise<ModelConfig[]>
28
33
  ```
29
34
 
30
35
  | Field | Type | Purpose |
31
36
  | --- | --- | --- |
32
37
  | `apiKey` | `CredentialValueSource` | Direct/callback/resolver API-key source. |
33
38
  | `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
34
- | `baseUrl` | `string` | Overrides the ZAI base URL. |
39
+ | `baseUrl` | `string` | Overrides the Z.AI base URL (default `https://api.z.ai/api/paas/v4`). |
35
40
  | `id` | `string` | Overrides the provider id (default `zai`). |
36
- | `models` | `readonly ModelConfig[]` | Overrides `zaiModels` defaults. |
41
+ | `models` | `readonly ModelConfig[]` | Overrides featured `zaiModels` defaults. |
37
42
 
38
- `ModelConfig.compat.thinkingFormat: "zai"` enables ZAI thinking handling;
39
- `developerRoleFallback` controls developer-role fallback.
43
+ ### Thinking / reasoning compat
44
+
45
+ Official body fields (request `options.compat` wins over `model.compat`):
46
+
47
+ | Compat / body field | Wire shape | Notes |
48
+ | --- | --- | --- |
49
+ | `thinking` | `boolean` or `{ type: "enabled" \| "disabled", clear_thinking?: boolean }` | Boolean `true`/`false` maps to `{ type: "enabled" }` / `{ type: "disabled" }`. |
50
+ | `reasoning_effort` | string | GLM-5.2+: `max` (default) \| `xhigh` \| `high` \| `medium` \| `low` \| `minimal` \| `none`. |
51
+ | `tool_stream` | boolean | GLM-4.6+: stream tool-call argument deltas (`false` by API default; featured 4.6+ models opt in). |
52
+ | `clear_thinking` | boolean | Nested into `thinking.clear_thinking`. Default on the API is `true` (drop prior reasoning). Set `false` for Preserved Thinking. |
53
+ | `preserveThinking` | boolean | Prism-local: when true (or when `clear_thinking: false`), replay prior thinking blocks as assistant `reasoning_content`. |
54
+
55
+ `ProviderRequestOptions.cacheRetention: "none"` forces `thinking: { type: "disabled" }`.
56
+
57
+ Shared Prism helpers (`applyThinkingLevel` / `thinkingFamilyForModel`) map portable
58
+ levels into the `reasoning_effort` family for Z.AI; hosts can also set
59
+ `thinking.type` directly for enable/disable.
40
60
 
41
61
  ## Outputs / response / events
42
62
 
43
63
  | Surface | Behavior |
44
64
  | --- | --- |
45
- | Provider stream | Prism text, thinking (preserved only when `model.compat.preserveThinking` is true, otherwise downgraded to text), tool-call delta/final, `usage`, `done`, redacted `error`. |
46
- | Block preservation | Text, thinking, assistant `tool_call` → `tool_calls`, `tool_result` → role `tool` messages, images when `capabilities.input` includes `"image"`. |
65
+ | Provider stream | Prism text, thinking (`delta.reasoning_content`), tool-call delta/final, `usage`, `done`, redacted `error`. |
66
+ | Block preservation | Text; thinking → `reasoning_content` when Preserved Thinking is active (otherwise dropped, never flattened into text); assistant `tool_call` → `tool_calls`; `tool_result` → role `tool`; images when `capabilities.input` includes `"image"`. |
47
67
  | Auth method | `api_key` for the configured provider id, credential name `apiKey`. |
48
68
 
49
69
  Unsupported block placements or unclaimed images fail before fetch.
50
70
 
51
71
  ## Request/response example
52
72
 
53
- Example request body (OpenAI-compatible Chat Completions shape):
73
+ Example request body (official Chat Completions shape):
54
74
 
55
75
  ```json
56
76
  {
57
- "model": "glm-4.6",
77
+ "model": "glm-5.2",
58
78
  "messages": [{ "role": "user", "content": "Hello" }],
59
- "stream": true
79
+ "stream": true,
80
+ "thinking": { "type": "enabled" },
81
+ "reasoning_effort": "max",
82
+ "tool_stream": true
60
83
  }
61
84
  ```
62
85
 
@@ -64,63 +87,89 @@ Example request body (OpenAI-compatible Chat Completions shape):
64
87
 
65
88
  ```ts
66
89
  import { createExtensionKernel } from "@arnilo/prism";
67
- import { createZaiProviderPackage } from "@arnilo/prism-provider-zai";
90
+ import { createZaiProviderPackage, listZaiModels } from "@arnilo/prism-provider-zai";
68
91
 
69
92
  const kernel = createExtensionKernel();
70
93
  await kernel.load([createZaiProviderPackage({ apiKey: "fake-zai-key" })]);
94
+
95
+ // Optional caller-gated discovery (never runs during package setup):
96
+ const live = await listZaiModels({ apiKey: "fake-zai-key" });
97
+ await kernel.load([createZaiProviderPackage({ apiKey: "fake-zai-key", models: live })]);
71
98
  ```
72
99
 
73
- Override the provider id and models:
100
+ Per-turn thinking override:
74
101
 
75
102
  ```ts
76
- import { createZaiProviderPackage, defineZaiModel, zaiModels } from "@arnilo/prism-provider-zai";
77
-
78
- await kernel.load([
79
- createZaiProviderPackage({ id: "zai", apiKey: "fake", models: zaiModels }),
80
- ]);
103
+ await session.prompt("Plan the refactor", {
104
+ providerOptions: {
105
+ compat: {
106
+ thinking: { type: "enabled", clear_thinking: false },
107
+ reasoning_effort: "high",
108
+ tool_stream: true,
109
+ },
110
+ },
111
+ });
81
112
  ```
82
113
 
83
114
  ## Extension and configuration notes
84
115
 
85
- - Hosts choose base URL, provider id, model list, credential source, and `fetch`
86
- impl.
87
- - `defineZaiModel` lets apps set ZAI-specific `compat` (thinking format, developer
88
- fallback).
89
- - Package contributes models via the extension `api` and an `api_key` auth method.
116
+ - Default base URL is the official international endpoint
117
+ `https://api.z.ai/api/paas/v4`. Hosts targeting China can pass
118
+ `baseUrl: "https://open.bigmodel.cn/api/paas/v4"`. Coding Plan hosts may use
119
+ `https://api.z.ai/api/coding/paas/v4`.
120
+ - Featured `zaiModels` are offline bootstrap aliases (`glm-5.2`, `glm-5.1`,
121
+ `glm-5`, `glm-5-turbo`, `glm-4.7`, `glm-4.6`, `glm-4.5`) curated from the
122
+ official Chat Completions model enum and overview context sizes.
123
+ - `listZaiModels()` is caller-gated OpenAI-compatible `GET /models` discovery
124
+ (not a first-class docs.z.ai list page). Setup never fetches.
125
+ - `defineZaiModel` sets Z.AI-specific `compat` (`thinking`, `reasoning_effort`,
126
+ `tool_stream`, `clear_thinking`, `preserveThinking`).
90
127
 
91
128
  ### Cache behavior
92
129
 
93
130
  - Z.AI GLM models use **implicit context caching**: the server caches prompt
94
131
  prefixes automatically based on request content, with no explicit request-side
95
132
  cache payload. Catalog models declare `cache: { kind: "implicit" }`.
133
+ - Official docs: hits appear in `usage.prompt_tokens_details.cached_tokens`.
96
134
  - The provider sends no `cache_control`, `prompt_cache_key`, `prompt_cache_retention`,
97
135
  or other explicit cache-control fields regardless of `ProviderRequestOptions.cache`
98
136
  / `cacheKey` / `cacheRetention` settings — those options have no effect on the
99
- Z.AI request body. Hosts relying on cache hits should keep their stable prompt
100
- prefix byte-stable and stable inputs unchanged.
101
- - Usage accounting is preserved: `prompt_tokens_details.cached_tokens` maps to
102
- `Usage.cacheReadTokens` and `prompt_tokens_details.cache_write_tokens` maps to
103
- `Usage.cacheWriteTokens` when the server reports them.
137
+ Z.AI request body (except `cacheRetention: "none"` disabling thinking).
138
+ - Usage accounting: `prompt_tokens_details.cached_tokens` → `Usage.cacheReadTokens`
139
+ and `prompt_tokens_details.cache_write_tokens` → `Usage.cacheWriteTokens` when
140
+ the server reports them.
104
141
 
105
142
  ## Security and performance notes
106
143
 
107
- - SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport` helpers (`readSseData`, `readBoundedResponseText`).
144
+ - SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport`
145
+ helpers (`readSseData`, `readBoundedResponseText`).
108
146
  - No network calls during import, setup, build, or default tests.
109
147
  - No automatic environment, file, keychain, or shell credential lookup.
110
148
  - API keys are resolved per request from caller-supplied values or resolvers and
111
- redacted from errors.
149
+ redacted from errors (including discovery failures).
112
150
  - Caller-supplied `ProviderRequest.options.headers` can add non-owned headers,
113
151
  but provider-owned headers (`content-type`, `authorization`) are applied last
114
152
  and cannot be overridden by caller headers.
115
- - Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus fake-safe
116
- provider-specific env names; default tests are network-free.
153
+ - Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus `ZAI_API_KEY`;
154
+ default tests are network-free.
117
155
 
118
156
  ## Related APIs
119
157
 
120
158
  - [Provider packages](../provider-packages.md): `defineProviderPackage`,
121
- `ModelConfig`/`compat`, thinking formats.
159
+ caller-gated discovery, per-turn thinking.
160
+ - [Thinking and reasoning](../thinking-and-reasoning.md): portable
161
+ `applyThinkingLevel` → Z.AI `reasoning_effort` / `thinking.type`.
122
162
  - [Credentials and redaction](../credentials-and-redaction.md):
123
163
  `resolveCredentialValue`, `redactSecrets`.
124
- - [OpenAI-compatible provider](openai-compatible.md): underlying Chat Completions
125
- adapter.
164
+ - [Provider caching](../provider-caching.md): implicit GLM context caching.
126
165
  - [Provider conformance](../provider-conformance.md): network-free adapter tests.
166
+
167
+ ## Official evidence
168
+
169
+ - [Deep Thinking](https://docs.z.ai/guides/capabilities/thinking)
170
+ - [Thinking Mode](https://docs.z.ai/guides/capabilities/thinking-mode) (Preserved Thinking / `clear_thinking`)
171
+ - [Tool Streaming](https://docs.z.ai/guides/capabilities/stream-tool)
172
+ - [Context Caching](https://docs.z.ai/guides/capabilities/cache)
173
+ - [Chat Completion](https://docs.z.ai/api-reference/llm/chat-completion)
174
+ - [Migrate to GLM-5.2](https://docs.z.ai/guides/overview/migrate-to-glm-new)
175
+ - [Models overview](https://docs.z.ai/guides/overview/overview)
@@ -15,7 +15,7 @@ Current contract groups:
15
15
  - Extensions/middleware: `ExtensionLifecycleEventName`, `ExtensionEvent`, `Extension`, `ExtensionAPI`, `MiddlewareHookName`, `Middleware`, `MiddlewareNext`, `MiddlewareRegistry`
16
16
  - Configuration/manifests: `ConfigProvider`, `ConfigLayer`, `ConfigLoadContext`, `PrismManifest`, `ManifestContributionDeclaration`, `ManifestResourceDeclaration`, `ManifestContributionKind`
17
17
  - Stores/resources/settings/credentials/compaction/retry/cache helpers: `SessionEntry`, `SessionStore`, `StoreFactory`, `Resource`, `ResourceLoader`, `ResourceLoadContext`, `SettingsProvider`, `CredentialRequest`, `Credential`, `CredentialResolver`, `CompactionStrategy`, `CompactionContext`, `CompactionResult`, `CompactionOptions`, `CompactionMiddlewarePayload`, `CompactionEntryData`, `DefaultCompactionStrategyOptions`, `RetryPolicy`, `RetryContext`, `RetryDecision`, `RetryOptions`, `RetryMiddlewarePayload`, `DefaultRetryPolicyOptions`, `CacheUsageReport`, `sanitizeCacheKey`, `mapCacheRetention`, `applyCacheControl`, `cacheHitRate`, `cacheSavings`, `cacheUsageReport`
18
- - Production persistence (adapter-facing): `ProductionPersistenceStore`, `CheckpointStore`, `CheckpointKey`, `CheckpointSaveInput`, `CheckpointRecord`, `CheckpointQuery`, `LeaseStore`, `LeaseKey`, `LeaseAcquireInput`, `LeaseClaimInput`, `LeaseRecord`, `PersistencePage`, `PersistenceQuery`, `OwnershipScope`, `SessionRecord`, `SessionQuery`, `BranchRecord`, `BranchQuery`, `SessionEntryQuery`, `RunRecord`, `RunQuery`, `AgentEventRecord`, `AgentEventQuery`, `ToolCallRecord`, `ToolCallQuery`, `UsageRecord`, `UsageQuery`, `AgentDefinitionRecord`, `AgentDefinitionQuery`, `RetentionPolicy`, `RetentionPolicyQuery`, `MigrationRecord`, `MigrationQuery`
18
+ - Production persistence (adapter-facing): `ProductionPersistenceStore`, `CheckpointStore`, `CheckpointKey`, `CheckpointSaveInput`, `CheckpointRecord`, `CheckpointQuery`, `LeaseStore`, `LeaseKey`, `LeaseAcquireInput`, `LeaseClaimInput`, `LeaseRecord`, `PersistencePage`, `PersistenceQuery`, `OwnershipScope`, `SessionRecord`, `SessionQuery`, `BranchRecord`, `BranchQuery`, `SessionEntryQuery`, `RunRecord`, `RunQuery`, `RunFeedbackRecord`, `RunFeedbackStore`, `RunFeedbackQuery`, `AgentEventRecord`, `AgentEventQuery`, `ToolCallRecord`, `ToolCallQuery`, `UsageRecord`, `UsageQuery`, `AgentDefinitionRecord`, `AgentDefinitionQuery`, `RetentionPolicy`, `RetentionPolicyQuery`, `MigrationRecord`, `MigrationQuery`
19
19
 
20
20
  ## When to use it
21
21
 
@@ -142,9 +142,10 @@ Important request shapes:
142
142
  | `SystemPromptContribution` | Explicit caller-selected prompt layer with source, mode, text, and metadata. |
143
143
  | `ConfigLayer` | Named JSON config layer consumed by `mergeConfigLayers()`. |
144
144
  | `PrismManifest` | Data-only package manifest with config defaults, contribution declarations, and resource declarations. |
145
- | `ProductionPersistenceStore` | Adapter-facing interface for durable, paginated, multi-tenant storage plus optional `checkpoints?: CheckpointStore` and `leases?: LeaseStore`. No SQL/ORM/host file storage/network dependency. |
145
+ | `ProductionPersistenceStore` | Adapter-facing interface for durable, paginated, multi-tenant storage plus optional `checkpoints?: CheckpointStore`, `leases?: LeaseStore`, and `feedback?: RunFeedbackStore`. No SQL/ORM/host file storage/network dependency. |
146
146
  | `CheckpointStore` | Generic versioned checkpoint capability: save/load/bounded-list/delete by namespace and key, with ownership, exact-version CAS, and lease fencing. `createMemoryCheckpointStore()` is the reference implementation. |
147
147
  | `LeaseStore` | Atomic acquire/renew/release/get by namespace and key, with opaque claim tokens, expiry, ownership scope, and monotonically increasing takeover fences. `createMemoryLeaseStore()` is the reference implementation. |
148
+ | `RunFeedbackStore` | Immutable append, bounded owned query, and owned deletion for ratings/comments/tags linked to existing run/trace/evaluation IDs. `createMemoryRunFeedbackStore()` is the reference implementation. |
148
149
  | `EventMultiplexer<T>` | Generic bounded fan-in from async sources. `createEventMultiplexer()` owns queue limits, overflow policy, abort, source teardown, and close behavior. |
149
150
  | `PersistencePage<T>` | Cursor-paginated result page: `items`, optional `nextCursor`, optional `total`. |
150
151
  | `PersistenceQuery` | Common pagination controls: `cursor?`, `limit?`, `order?: "asc" \| "desc"`. |
@@ -155,7 +156,7 @@ Important request shapes:
155
156
  | `RunRecord` / `RunQuery` | Stored run and filters: session, branch, status, timestamps, ownership. |
156
157
  | `AgentEventRecord` / `AgentEventQuery` | Event ledger row with `redacted` flag and filters by type, session, run, entry, timestamp, ownership. |
157
158
  | `ToolCallRecord` / `ToolCallQuery` | Tool-call row with `redacted` flag and filters by name, status, session, run, entry, timestamps, ownership. |
158
- | `UsageRecord` / `UsageQuery` | Usage row and filters: session, run, entry, recorded-at range, ownership. |
159
+ | `UsageRecord` / `UsageQuery` | Usage row and filters: `provider_turn`/`run_total` scope, turn/attempt, session, run, entry, recorded-at range, ownership. |
159
160
  | `CacheUsageReport` | Numeric cache diagnostics from normalized `Usage`: read/write tokens, hit rate, estimated savings, and optional currency. |
160
161
  | `AgentDefinitionRecord` / `AgentDefinitionQuery` | Versioned agent-definition snapshot and filters. Does not store credentials or provider instances. |
161
162
  | `RetentionPolicy` / `RetentionPolicyQuery` | Retention policy and filters: age, entry count, byte limits, archive store, applied kinds. |
@@ -412,8 +413,8 @@ void credentials;
412
413
  - Contracts are host-owned and package-friendly. External packages can implement `AIProvider`, `ToolDefinition`, `CommandDefinition`, `AgentDefinition`, `InputBuilder`, `PromptBuilder`, `Middleware`, `ContextProvider`, `Skill`, `Extension`, config providers, data-only manifests, compaction strategies, store factories, resource loaders, settings providers, and credential resolvers.
413
414
  - `ExtensionAPI` is implemented by the extension kernel. It exposes explicit registries, ordered middleware registration, ordered event subscription/emission, and registration methods for Phase 2 contribution categories.
414
415
  - `AgentConfig.provider` can hold a direct provider instance for simple host wiring. Hosts that need config-driven selection should use `ModelConfig.provider` with explicit `createProviderRegistry()` / `createModelRegistry()` objects; Prism does not create a hidden global provider registry.
415
- - `SettingsProvider` and `CredentialResolver` are explicit dependencies. Prism must not hide global settings or credentials behind these contracts, and `CredentialResolver` should be passed only to the edge that needs a credential. `AgentConfig.settings` and `AgentConfig.credentials` are host-owned metadata; the session runtime does not call `settings.get()` or `credentials.resolve()`.
416
- - `AgentConfig.extensions` is host-owned metadata; the session runtime does not load extensions or call `Extension.setup()`. Use `createExtensionKernel().load(...)` before creating an agent, then pass selected contributions into `AgentConfig`.
416
+ - `SettingsProvider` and `CredentialResolver` are explicit dependencies. Prism must not hide global settings or credentials behind these contracts, and `CredentialResolver` should be passed only to the edge that needs a credential. These seams are host-owned outside `AgentConfig`; the session runtime does not call `settings.get()` or `credentials.resolve()`.
417
+ - Extension loading is host-owned outside `AgentConfig`; the session runtime does not load extensions or call `Extension.setup()`. Use `createExtensionKernel().load(...)` before creating an agent, then pass selected contributions into `AgentConfig`.
417
418
  - `PrismManifest` is data-only. It can describe contribution modules/resources and config defaults, but parsing it does not import modules, execute package code, or mutate registries.
418
419
  - Resource helper functions decode resources from a caller-provided `ResourceLoader`; Prism does not include host file storage, network, package, or URI router loaders.
419
420
  - `createDefaultInputBuilder()` is a small default implementation of `InputBuilder`. It is replaceable and only loads explicit URI resources through a caller-provided `ResourceLoader`.
package/docs/rag.md ADDED
@@ -0,0 +1,113 @@
1
+ # Retrieval-augmented generation (RAG)
2
+
3
+ ## What it does
4
+
5
+ `@arnilo/prism-rag` is an optional package for deterministic plain-text/Markdown chunking, bounded embedding/vector indexing, filtered semantic retrieval, stable citations, and explicit `ContextProvider` injection. It reuses `Embedder` and `VectorStore` from `@arnilo/prism-memory`; Prism core input assembly is unchanged.
6
+
7
+ ## When to use it
8
+
9
+ Use it when a host already owns trusted document text and needs small retrieval primitives without a document framework. Do not use it for PDF/HTML/LaTeX parsing, semantic chunking, metadata extraction agents, reranker pipelines, GraphRAG, crawling, URL fetching, or filesystem discovery.
10
+
11
+ ## Inputs / request
12
+
13
+ Chunking:
14
+
15
+ | API/field | Meaning |
16
+ | --- | --- |
17
+ | `chunkText(text, options)` | Character-bounded plain-text chunks |
18
+ | `chunkMarkdown(markdown, options)` | Same engine, preferring heading/paragraph boundaries |
19
+ | `sourceId` | Required stable, non-secret source identifier |
20
+ | `size` / `overlap` | Character ceiling and repeated context |
21
+ | `metadata` | JSON metadata copied to every chunk |
22
+
23
+ Index/retrieve:
24
+
25
+ | Field | Required | Meaning |
26
+ | --- | --- | --- |
27
+ | `embedder` / `store` | yes | Phase 7 `Embedder` and `VectorStore` |
28
+ | `scope` | yes | `{ tenantId, resourceId, corpusId }`; corpus maps to vector thread isolation |
29
+ | `chunks` | indexing | `RagChunk[]` from package chunkers or compatible host parser |
30
+ | `topK` / `queryCandidates` | retrieval | Returned result count and bounded pre-filter candidates |
31
+ | `filter` | no | Shallow JSON metadata equality filter |
32
+ | `redactor` / `secrets` | no | Redact before embedding, persistence, and injection |
33
+ | `signal` | no | Abort embedding, vector operations, and batch progression |
34
+
35
+ ## Outputs / response / events
36
+
37
+ - `chunkText()` / `chunkMarkdown()` return frozen `RagChunk[]` with `sourceId`, zero-based index, offsets, and stable IDs such as `guide#0001`.
38
+ - `indexChunks()` returns `{ indexed, sourceIds }` after bounded batch upserts.
39
+ - `retrieveContext()` returns `{ query, text, hits, citations, truncated }`. Rendered text uses `[citation-id] text` blocks.
40
+ - `createRagContextProvider()` returns one ordinary context provider. Empty queries/results contribute no block.
41
+ - No events, tools, permissions, provider calls, loaders, or network requests are added.
42
+
43
+ Default hard ceilings include 1,000/16,384 chunk characters, 100/4,096 overlap, 1,048,576/8,388,608 document characters, 2,048/8,192 chunks, 32/128 embed batch, top-K 5/32, candidates 20/128, result 64/512 KiB, and context 2,000/8,000 estimated tokens.
44
+
45
+ ## Request/response example
46
+
47
+ ```json
48
+ {
49
+ "scope": { "tenantId": "t1", "resourceId": "docs", "corpusId": "handbook" },
50
+ "query": "How do approvals work?",
51
+ "topK": 1,
52
+ "result": {
53
+ "text": "[security-guide#0001] Recheck policy before side effects.",
54
+ "citations": [{ "id": "security-guide#0001", "sourceId": "security-guide" }]
55
+ }
56
+ }
57
+ ```
58
+
59
+ ## Implementation example
60
+
61
+ ```ts
62
+ import { createAgent, createMockProvider, providerDone, providerTextDelta } from "@arnilo/prism";
63
+ import { createHashEmbedder, createMemoryVectorStore } from "@arnilo/prism-memory";
64
+ import { chunkMarkdown, createRagContextProvider, indexChunks, retrieveContext } from "@arnilo/prism-rag";
65
+
66
+ const embedder = createHashEmbedder(); // deterministic demo/test helper, not production semantic quality
67
+ const store = createMemoryVectorStore();
68
+ const scope = { tenantId: "t1", resourceId: "docs", corpusId: "handbook" };
69
+ const chunks = chunkMarkdown("# Approval\n\nRecheck current policy before side effects.", {
70
+ sourceId: "security-guide",
71
+ metadata: { category: "security" },
72
+ });
73
+ await indexChunks({ chunks, embedder, store, scope });
74
+
75
+ const found = await retrieveContext("approval policy", {
76
+ embedder,
77
+ store,
78
+ scope,
79
+ topK: 4,
80
+ filter: { category: "security" },
81
+ });
82
+
83
+ const agent = createAgent({
84
+ model: { provider: "mock", model: "demo" },
85
+ provider: createMockProvider([providerTextDelta("Policy checked."), providerDone()]),
86
+ context: [createRagContextProvider({ embedder, store, scope })],
87
+ });
88
+ console.log(found.text, await agent.createSession().run("How do approvals work?"));
89
+ ```
90
+
91
+ ## Extension and configuration notes
92
+
93
+ - Supply any Phase 7-conforming embedder/vector store, including the in-memory reference or PostgreSQL/pgvector adapter.
94
+ - Metadata filtering is package-local after a bounded candidate query so existing vector contracts/adapters remain unchanged. Increase `queryCandidates` only when selective filters measurably need it.
95
+ - `createRagContextProvider()` derives its query from latest user text by default; pass a fixed string or callback for host-controlled query generation.
96
+ - Load source text separately with a host-owned `ResourceLoader`. The RAG package intentionally accepts text, not URLs or filesystem paths.
97
+ - Package is available directly or through `@arnilo/prism-all`; installation does not create an embedder, vector store, or context provider.
98
+
99
+ ## Security and performance notes
100
+
101
+ - Every index/query includes exact tenant/resource/corpus scope; returned records are rechecked and malformed/foreign records fail closed.
102
+ - Source IDs become citation/storage IDs and must be stable non-secret identifiers. Text and user metadata can be redacted before external embedding and persistence.
103
+ - Retrieved documents are untrusted inert context. Prompt-injection text cannot activate tools, skills, credentials, permissions, or extensions.
104
+ - Remote sources must pass existing resource/media trust, SSRF, MIME, and byte policies before their decoded text reaches this package.
105
+ - Indexing is bounded per batch and checks abort between embed/upsert operations. A failure can leave completed batches persisted; retry is idempotent for the same stable source/chunk IDs.
106
+ - Filtering scans at most `queryCandidates` hits; rendering stops at top-K, UTF-8 result bytes, or estimated context-token ceiling.
107
+
108
+ ## Related APIs
109
+
110
+ - [Working and semantic memory](working-and-semantic-memory.md): shared `Embedder`/`VectorStore` contracts and adapters.
111
+ - [Context and skills](context-and-skills.md): explicit `ContextProvider` injection and inert context semantics.
112
+ - [Resource loading](resource-loading.md): host-owned trusted source loading.
113
+ - [Multimodal content](multimodal-content.md): remote media SSRF/MIME/byte policies before text extraction.