@arnilo/prism 0.0.10 → 0.0.12
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +27 -0
- package/dist/agent-loops.d.ts +1 -0
- package/dist/agent-loops.js +37 -2
- package/dist/agent-run-lifecycle.d.ts +5 -1
- package/dist/agent-run-lifecycle.js +17 -1
- package/dist/agents.d.ts +3 -1
- package/dist/agents.js +202 -24
- package/dist/context-budget.d.ts +63 -0
- package/dist/context-budget.js +235 -0
- package/dist/contracts.d.ts +109 -0
- package/dist/contracts.js +77 -0
- package/dist/index.d.ts +8 -6
- package/dist/index.js +5 -4
- package/dist/input.d.ts +3 -0
- package/dist/input.js +71 -28
- package/dist/node/session-store-jsonl.js +4 -1
- package/dist/rpc.js +13 -2
- package/dist/session-stores.d.ts +7 -2
- package/dist/session-stores.js +174 -4
- package/dist/structured-output.d.ts +5 -1
- package/dist/structured-output.js +18 -0
- package/dist/testing/persistence-schema.d.ts +1 -1
- package/dist/testing/persistence-schema.js +8 -2
- package/dist/testing/session-store-conformance.d.ts +6 -0
- package/dist/testing/session-store-conformance.js +36 -1
- package/docs/a2a.md +1 -0
- package/docs/ag-ui.md +123 -0
- package/docs/agent-events.md +2 -1
- package/docs/agent-loops.md +8 -1
- package/docs/agent-session-runtime.md +10 -2
- package/docs/cli-rpc.md +2 -1
- package/docs/coding-agent-tools.md +70 -1
- package/docs/compaction-and-retry.md +2 -1
- package/docs/compaction-llm.md +20 -1
- package/docs/credential-storage.md +2 -1
- package/docs/credentials-and-redaction.md +8 -0
- package/docs/evaluations.md +1 -1
- package/docs/host-security.md +2 -0
- package/docs/index.md +20 -17
- package/docs/input-and-prompt-assembly.md +4 -1
- package/docs/migration.md +32 -0
- package/docs/node-jsonl-session-store.md +1 -1
- package/docs/performance.md +36 -0
- package/docs/postgres-persistence.md +3 -3
- package/docs/provider-packages.md +11 -1
- package/docs/providers/anthropic.md +93 -0
- package/docs/providers/google.md +88 -0
- package/docs/providers/openai.md +2 -2
- package/docs/public-contracts.md +6 -0
- package/docs/release-and-install.md +232 -65
- package/docs/review-coverage-2026-07-22-phase-6.md +209 -0
- package/docs/review-coverage-2026-07-22-phase-7.md +173 -0
- package/docs/runs-and-usage.md +1 -0
- package/docs/server.md +1 -0
- package/docs/session-store-conformance.md +2 -0
- package/docs/session-stores.md +40 -1
- package/docs/sqlite-persistence.md +3 -3
- package/docs/structured-output.md +7 -1
- package/docs/workflows.md +2 -0
- package/package.json +3 -2
package/docs/ag-ui.md
ADDED
|
@@ -0,0 +1,123 @@
|
|
|
1
|
+
# Frontend interoperability (AG-UI and ACP)
|
|
2
|
+
|
|
3
|
+
## What it does
|
|
4
|
+
|
|
5
|
+
`@arnilo/prism-ag-ui` is an optional, framework-free protocol adapter over Prism's existing redacted `AgentEvent`, session, durable-run, and persistence seams.
|
|
6
|
+
|
|
7
|
+
- Root export maps Prism events to AG-UI `@ag-ui/core` **0.0.57** events and offers `createAgUiHandler()` (`Request` → SSE `Response`) plus `createPersistenceAgUiReplay()`.
|
|
8
|
+
- `@arnilo/prism-ag-ui/acp` uses stable `@agentclientprotocol/sdk` **1.3.0** root exports for `createAcpEventMapper()` and `createPrismAcpAgent()`.
|
|
9
|
+
- Core remains protocol-free. `resumeAgentRunStream()` / `AgentRunLifecycle.resumeStream()` are generic durable-resume streams shared by adapters.
|
|
10
|
+
|
|
11
|
+
This is not an app TUI, desktop shell, conversation database, terminal/filesystem bridge, A2A implementation, or frontend tool registry.
|
|
12
|
+
|
|
13
|
+
## When to use it
|
|
14
|
+
|
|
15
|
+
Use AG-UI when a host already authenticates users, owns sessions and durable run correlation, and needs a bounded Web endpoint for a browser/TUI/desktop client. Use ACP when an editor client already supplies an ACP transport and needs text, safe tool status, usage, and approval updates from a Prism session.
|
|
16
|
+
|
|
17
|
+
Use [A2A interoperability](a2a.md) for remote agent-to-agent JSON-RPC/HTTPS tasks. AG-UI/ACP are frontend/client protocol adapters; neither replaces A2A task lifecycle or storage.
|
|
18
|
+
|
|
19
|
+
## Inputs / request
|
|
20
|
+
|
|
21
|
+
Install the optional package beside the core runtime (it becomes publishable with the 0.0.12 release graph):
|
|
22
|
+
|
|
23
|
+
```bash
|
|
24
|
+
npm install @arnilo/prism @arnilo/prism-ag-ui
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
`createAgUiHandler()` takes host-owned callbacks:
|
|
28
|
+
|
|
29
|
+
| Input | Purpose |
|
|
30
|
+
| --- | --- |
|
|
31
|
+
| `authorize` | Rebinds untrusted AG-UI thread/run selectors to host ownership on every request. `false` returns 403. |
|
|
32
|
+
| `sessionFactory` | Returns an authorized Prism `AgentSession`; client input never selects tools or capabilities. |
|
|
33
|
+
| `lifecycle` + `resolveRun` | Optional durable status/resume path. Required only for a resumed interruption. |
|
|
34
|
+
| `replay` | Optional `createPersistenceAgUiReplay(store, options)` adapter for ownership-scoped durable pages. |
|
|
35
|
+
| `projection` | Explicit safe tool args/results, paths, or state projection. Omit it for default deny. |
|
|
36
|
+
| `redactor`, `limits` | Host redaction and narrowing-only finite caps. |
|
|
37
|
+
|
|
38
|
+
The handler accepts only `POST` JSON validated with AG-UI `RunAgentInputSchema`. IDs are bounded URL-safe values; it uses only the last text user message. Frontend tools and non-empty frontend state are rejected before authorization or session lookup. Start a run with no `resume` and no `?cursor=`; resume has exactly one entry; replay supplies `?cursor=`.
|
|
39
|
+
|
|
40
|
+
## Outputs / response / events
|
|
41
|
+
|
|
42
|
+
The handler returns `text/event-stream`, one `data: <AG-UI event>\n\n` frame per output. Mapper lifecycle is ordered: Prism `agent_started`/assistant text/tool events map to `RUN_STARTED`, `TEXT_MESSAGE_*`, and `TOOL_CALL_*`; terminal success maps to `RUN_FINISHED`; runtime errors map to `RUN_ERROR`. Active AG-UI message/tool sequences close before an error, interruption, or finish.
|
|
43
|
+
|
|
44
|
+
A Prism durable `agent_suspended` returns `RUN_FINISHED` with interrupt id `${runId}:${version}` and a strict `{ decision: "approve" | "deny" }` schema. A client must address that exact current id. `cancelled` means deny; a resolved resume payload must contain only that decision. The adapter checks host authorization, selected run, suspended status, and checkpoint version, then calls `AgentRunLifecycle.resumeStream()` once. Claimed/dispatched tools are never replayed.
|
|
45
|
+
|
|
46
|
+
`createPersistenceAgUiReplay()` queries only the host-resolved run with ownership and ascending bounded pagination. Every record must already be redacted. Events carry `prismEventId` for at-least-once page-boundary de-duplication; a nonterminal final page may attach a filtered live subscriber. Terminal pages never create a session or rerun a provider/tool.
|
|
47
|
+
|
|
48
|
+
ACP maps assistant text to `agent_message_chunk`, safe tool lifecycle to `tool_call`/`tool_call_update`, provider usage to `usage_update`, and durable suspension to `session/request_permission`. Only `allow_once` approves; reject, cancellation, unknown outcomes, and request failure deny. It advertises only close-session capability—no terminal, filesystem, MCP, editor state, location, diff, or raw input/output capability.
|
|
49
|
+
|
|
50
|
+
## Request/response example
|
|
51
|
+
|
|
52
|
+
```json
|
|
53
|
+
{
|
|
54
|
+
"threadId": "thread-1",
|
|
55
|
+
"runId": "run-1",
|
|
56
|
+
"messages": [{ "id": "message-1", "role": "user", "content": "Summarize this" }],
|
|
57
|
+
"tools": [],
|
|
58
|
+
"state": {}
|
|
59
|
+
}
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
A suspended response includes this resumable interrupt shape:
|
|
63
|
+
|
|
64
|
+
```json
|
|
65
|
+
{
|
|
66
|
+
"type": "RUN_FINISHED",
|
|
67
|
+
"threadId": "thread-1",
|
|
68
|
+
"runId": "run-1",
|
|
69
|
+
"outcome": {
|
|
70
|
+
"type": "interrupt",
|
|
71
|
+
"interrupts": [{ "id": "run-1:4", "responseSchema": { "required": ["decision"] } }]
|
|
72
|
+
}
|
|
73
|
+
}
|
|
74
|
+
```
|
|
75
|
+
|
|
76
|
+
Resume the same host thread/run with `resume: [{ "interruptId": "run-1:4", "status": "resolved", "payload": { "decision": "approve" } }]`. Do not send a copied session transcript, tool definitions, or mutable application state.
|
|
77
|
+
|
|
78
|
+
## Implementation example
|
|
79
|
+
|
|
80
|
+
```ts
|
|
81
|
+
import { createAgent, createMockProvider, providerDone, providerTextDelta } from "@arnilo/prism";
|
|
82
|
+
import { createAgUiHandler } from "@arnilo/prism-ag-ui";
|
|
83
|
+
|
|
84
|
+
const agent = createAgent({
|
|
85
|
+
model: { provider: "mock", model: "offline" },
|
|
86
|
+
provider: createMockProvider([providerTextDelta("ready"), providerDone()]),
|
|
87
|
+
});
|
|
88
|
+
|
|
89
|
+
const handle = createAgUiHandler({
|
|
90
|
+
authorize: ({ request }) => request.headers.get("authorization") === "Bearer host-checked"
|
|
91
|
+
? { ownership: { userId: "user-1" } }
|
|
92
|
+
: false,
|
|
93
|
+
sessionFactory: () => agent.createSession({ id: "host-owned-thread" }),
|
|
94
|
+
projection: { toolArguments: () => undefined, toolResult: () => undefined },
|
|
95
|
+
});
|
|
96
|
+
|
|
97
|
+
const response = await handle(request); // adapt this Web Response in host framework
|
|
98
|
+
```
|
|
99
|
+
|
|
100
|
+
See runnable network-free [`examples/ag-ui-server.ts`](../examples/ag-ui-server.ts). For ACP, construct `createPrismAcpAgent({ authorize, sessionFactory, lifecycle })` and connect the returned stable SDK agent through the host's ACP transport.
|
|
101
|
+
|
|
102
|
+
## Extension and configuration notes
|
|
103
|
+
|
|
104
|
+
All identity, authorization, session/thread mapping, durable checkpoint lookup, persistence selection, replay cursor persistence, transport adaptation, and optional projection are host-owned. The adapter owns no listener, database, background reconnect loop, credential resolver, or UI state.
|
|
105
|
+
|
|
106
|
+
`AgUiProjection` is an allow-list. Without a callback, raw tool arguments/results/progress, paths, arbitrary state, raw Prism events, ACP locations/diffs/terminals/raw I/O, and frontend-supplied tools remain absent. Use a projector that returns a redacted display value, not a host filesystem path or tool payload.
|
|
107
|
+
|
|
108
|
+
## Security and performance notes
|
|
109
|
+
|
|
110
|
+
Authorize every start, replay, resume, ACP new/prompt/cancel/close request. Treat thread IDs, run IDs, cursors, client messages, resume payloads, and protocol output as untrusted. Persist run ↔ protocol correlation before exposing an interrupt. Keep `SecretRedactor` active for streaming and ledger writes.
|
|
111
|
+
|
|
112
|
+
Defaults / hard caps: request 64 KiB / 1 MiB; input 128 / 1024 messages and 64 KiB / 1 MiB text; projected event 64 KiB / 1 MiB; error 8 KiB / 64 KiB; cursor 4 / 16 KiB; replay page 100 / 500 records; queue 128 / 4096 events; stream 10,000 / 100,000 events and 10 / 64 MiB; wall time 120 seconds / 30 minutes. Overflow yields a bounded error/closed stream, not an unbounded queue. Reconnect is at-least-once, so clients de-duplicate stable event/message/tool IDs.
|
|
113
|
+
|
|
114
|
+
Benchmark command/result placeholder: Task 8 adds `node scripts/benchmark-0.0.12.mjs` for mapper throughput, replay/handler latency, queue/heap, bytes, and coding-compaction preparation. No 0.0.12 timing result is claimed before that gate.
|
|
115
|
+
|
|
116
|
+
## Related APIs
|
|
117
|
+
|
|
118
|
+
- [Agent/session runtime](agent-session-runtime.md): `session.stream()`, `resumeAgentRunStream()`, and durable lifecycle.
|
|
119
|
+
- [Agent events](agent-events.md): normalized source events and ledger redaction.
|
|
120
|
+
- [Runs and usage ledger](runs-and-usage.md): durable `AgentEventRecord` query source.
|
|
121
|
+
- [Web-standard server handler](server.md): generic Prism HTTP API, separate from AG-UI.
|
|
122
|
+
- [A2A interoperability](a2a.md): remote agent-to-agent tasks, not frontend protocol mapping.
|
|
123
|
+
- [Host security guide](host-security.md): authorization, ownership, redaction, and credential boundaries.
|
package/docs/agent-events.md
CHANGED
|
@@ -191,7 +191,7 @@ for await (const event of session.stream("draft", { loop: { strategy: "generate-
|
|
|
191
191
|
|
|
192
192
|
- All events flow through `redactAgentEvent(event, activeRedactor)` before subscribers observe them. Configure `AgentConfig.redactor` / `RunOptions.redactor` via `createSecretRedactor([...knownSecretStrings])` so secret values are redacted in `message` content, `errors[].message`, `metadata`, and artifact `result`/`failure` payloads.
|
|
193
193
|
- The artifact variants are emitted only by `generateValidateReviseLoop`. `singleShotLoop` (the default when no `AgentConfig.loop` / `RunOptions.loop` is set) emits zero artifact events. See [Agent loops](agent-loops.md).
|
|
194
|
-
- Subscribers are in-process; the broadcaster is in-memory and live-only. Multiple `subscribe()` calls receive the same stream.
|
|
194
|
+
- Subscribers are in-process; the broadcaster is in-memory and live-only. Multiple `subscribe()` calls receive the same stream. `resumeAgentRunStream()` and `AgentRunLifecycle.resumeStream()` subscribe before resumed execution and yield only their selected durable `runId`; approval emits the normal `agent_started` then `agent_resumed` envelope, denial emits only `agent_denied`.
|
|
195
195
|
- `session.subscribe(options)` accepts `maxQueuedEvents` (default `1024`, minimum `1`) and `overflow` (default `"close"`). The `close` policy clears queued payload events, queues one `event_subscriber_overflow` notice for that subscriber, then closes it. `drop_oldest` keeps the newest queued events; `drop_newest` ignores new events while full.
|
|
196
196
|
- The union is additive: new variants are appended without renumbering; subscribers should handle unknown `event.type` gracefully.
|
|
197
197
|
|
|
@@ -212,3 +212,4 @@ for await (const event of session.stream("draft", { loop: { strategy: "generate-
|
|
|
212
212
|
- [Observability](observability.md): `ProviderTurnMetadata`; optional adapter builds one parented GenAI span tree from metadata-only lifecycle events and ignores message/progress deltas.
|
|
213
213
|
- [Tools](tools.md): `tool_execution_*` variants.
|
|
214
214
|
- [Compaction and retry policies](compaction-and-retry.md): `compaction_*` and `retry_scheduled` variants.
|
|
215
|
+
- [Frontend interoperability (AG-UI and ACP)](ag-ui.md): optional redacted mapping of this stream; durable replay is ledger-backed and at-least-once, never a live-subscriber substitute.
|
package/docs/agent-loops.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
## What it does
|
|
4
4
|
|
|
5
|
-
Agent loops make the agent's per-run turn-control flow a replaceable strategy without forking the runtime. The runtime owns provider calls, retry, abort, store appends, redaction, and event emission; a loop only orchestrates those shared primitives through a `LoopContext`. The default `singleShotLoop` is the former inline turn loop extracted verbatim — assemble → generate → append assistant message → optional tool dispatch → next turn. `generateValidateReviseLoop` is the first alternative loop: generate → parse → validate → revise up to a budget.
|
|
5
|
+
Agent loops make the agent's per-run turn-control flow a replaceable strategy without forking the runtime. The runtime owns provider calls, retry, abort, store appends, redaction, and event emission; a loop only orchestrates those shared primitives through a `LoopContext`. The default `singleShotLoop` is the former inline turn loop extracted verbatim — optional drain pending steers → assemble → generate → append assistant message → optional tool dispatch → next turn (continue while steers remain even if the provider returned no tool calls). `generateValidateReviseLoop` is the first alternative loop: generate → parse → validate → revise up to a budget.
|
|
6
6
|
|
|
7
7
|
Loops are opt-in. When no `loop` is configured, the runtime runs `singleShotLoop` and behavior is bit-for-bit with the pre-loop runtime.
|
|
8
8
|
|
|
@@ -58,6 +58,7 @@ await session.run(input, {
|
|
|
58
58
|
repairer: hostRepairer, // optional; default stringifies validation.errors[].message
|
|
59
59
|
maxRevisions: 3, // optional; default 3
|
|
60
60
|
toolCalls: "bounded", // optional; default "disabled"; uses limits.maxToolRounds
|
|
61
|
+
structuredOutputTiming: "final-turn-only", // optional; default "every-turn"
|
|
61
62
|
},
|
|
62
63
|
});
|
|
63
64
|
|
|
@@ -82,6 +83,10 @@ type AgentLoopOptions =
|
|
|
82
83
|
readonly maxRevisions?: number;
|
|
83
84
|
/** Default "disabled". "bounded" dispatches sequentially up to limits.maxToolRounds. */
|
|
84
85
|
readonly toolCalls?: "disabled" | "bounded";
|
|
86
|
+
readonly structuredOutput?: StructuredOutputOptions;
|
|
87
|
+
readonly structuredOutputMode?: "native" | "artifact-loop";
|
|
88
|
+
/** Default "every-turn". "final-turn-only" omits schema while tools may run. */
|
|
89
|
+
readonly structuredOutputTiming?: "every-turn" | "final-turn-only";
|
|
85
90
|
};
|
|
86
91
|
```
|
|
87
92
|
|
|
@@ -96,6 +101,8 @@ Host callback contracts (all generic over host `T`):
|
|
|
96
101
|
| `ArtifactContext` | `{ sessionId, runId, turn, signal, metadata }` — passed to every callback. |
|
|
97
102
|
| `ArtifactParseResult<T>` | `{ ok: boolean; value?: T; error?: string }`. |
|
|
98
103
|
|
|
104
|
+
Optional steer hooks on `LoopContext` (0.0.11): `hasPendingSteers?()` / `applyPendingSteers?()`. Hosts/custom loops that omit them keep pre-steer behavior; built-in loops drain at turn start.
|
|
105
|
+
|
|
99
106
|
`LoopContext` (what the runtime builds for the loop each run):
|
|
100
107
|
|
|
101
108
|
| Field | Purpose |
|
|
@@ -19,6 +19,7 @@ The agent/session runtime adds the minimal shared SDK surface for running provid
|
|
|
19
19
|
- `session.fork(options?)`
|
|
20
20
|
- `session.clone(options?)`
|
|
21
21
|
- `resumeAgentRun(agent, ref, decision, options)`
|
|
22
|
+
- `resumeAgentRunStream(agent, ref, decision, options)` → owned durable-resume `AsyncIterable<AgentEvent>`
|
|
22
23
|
- `createAgentRunLifecycle({ checkpoints, resolveAgent })` for host-selected remote status/resume adapters
|
|
23
24
|
|
|
24
25
|
The runtime streams provider text/tool-call content into `AgentEvent` values. Complete `tool_call` events are dispatched through the active host `ToolRegistry`, then returned as tool-result messages on the next provider turn. When a store is supplied, user, assistant, tool-result, and model-change entries are appended under the current branch leaf. Abort propagation and run exclusivity use native `AbortController`.
|
|
@@ -56,6 +57,8 @@ string | Message | readonly Message[]
|
|
|
56
57
|
|
|
57
58
|
`session.stream(input, options?)` subscribes first, starts exactly one run, yields only that run's events, and terminates when the run succeeds, fails, or aborts. Early consumer return aborts the owned run and releases the session. `SubscribeOptions.maxQueuedEvents` / `overflow` may be passed alongside `RunOptions`.
|
|
58
59
|
|
|
60
|
+
`resumeAgentRunStream(agent, ref, resume, options?)` does the same for one existing suspended durable run. It validates checkpoint ownership, revision/fingerprint, and `expectedVersion`, then subscribes before emitting `agent_started` / `agent_resumed` and resumed message/tool/terminal events. `AgentRunResumeStreamOptions` combines existing resume options with `signal`, `maxQueuedEvents`, and `overflow`; early return aborts only resumed execution. It does not replay a claimed/dispatched tool, poll a ledger, or retain a worker. `createAgentRunLifecycle().resumeStream(ref, resume, request?)` adds the same behavior after host agent-capability resolution.
|
|
61
|
+
|
|
59
62
|
`session.subscribe(options?)` remains available for hosts that want a long-lived subscriber across runs. Subscribe before `run()` to observe that run's events. The consumer loop and `session.run()` must run concurrently (e.g. start the `for await` consumer, then `await Promise.all([consumer, session.run("Hi")])`): events are only emitted during a live run, so awaiting the subscribe loop before calling `run()` deadlocks. Prefer `session.stream()` when you only need one run's events. `SubscribeOptions.maxQueuedEvents` defaults to `1024` (minimum `1`) and caps events queued while the consumer is not awaiting `next()`. `SubscribeOptions.overflow` defaults to `"close"`; it clears queued payload events, delivers one `event_subscriber_overflow` notice to that subscriber, then closes it. `"drop_oldest"` keeps newest events; `"drop_newest"` ignores new events while full.
|
|
60
63
|
|
|
61
64
|
For a text-only provider turn, the runtime emits:
|
|
@@ -78,7 +81,11 @@ Provider `thinking`/`reasoning` content emitted during a turn is preserved as `t
|
|
|
78
81
|
|
|
79
82
|
Missing providers fail closed: `run()` emits `error` and rejects before calling any provider. Provider `error` events emit session `error` and reject unless configured retry handles a transient provider-turn failure before output. Unknown tools fail closed through the tool harness and do not execute. Tool exceptions emit `tool_execution_error`, return an error `ToolResult`, and may still continue to the next provider turn.
|
|
80
83
|
|
|
81
|
-
Only one `run()` may be active per session. Concurrent `run()` calls emit `error` and reject immediately; Prism does not queue
|
|
84
|
+
Only one `run()` may be active per session. Concurrent `run()` / `prompt` / `followUp` calls emit `error` and reject immediately; Prism does not queue second prompts. Manual `compact()` also rejects while a run is active.
|
|
85
|
+
|
|
86
|
+
### Mid-run steer (0.0.11)
|
|
87
|
+
|
|
88
|
+
`session.steer(input, options?)` enqueues user text into the **same** active run (fail closed when no run). Default: inject at the next turn boundary (after tool rounds / before next provider assemble). `options.softInterrupt: true` aborts only the current provider stream, then continues the same `runId` with steered text. Pending queue caps: **8** messages / **64 KiB** UTF-8 total (`DEFAULT_MAX_PENDING_STEERS` / `DEFAULT_MAX_PENDING_STEER_BYTES`); overflow throws. Steered messages pass input guardrails + normal session append/redaction. Loops drain via optional `LoopContext.hasPendingSteers` / `applyPendingSteers`.
|
|
82
89
|
|
|
83
90
|
`session.abort(reason)` aborts the active run. The abort signal is passed to input assembly, provider requests, and tool execution; if a tool/provider path aborts after a tool call, Prism does not start another provider turn.
|
|
84
91
|
|
|
@@ -180,7 +187,7 @@ if (result.status === "suspended") {
|
|
|
180
187
|
}
|
|
181
188
|
```
|
|
182
189
|
|
|
183
|
-
Resume requires exact checkpoint ownership, version, agent fingerprint, and revision. Prism CAS-claims approval before work, rechecks normal guardrail/permission/validation/limit paths, and marks a pending tool dispatched before its side effect. `createAgentRunLifecycle()` wraps the same core path for server/MCP hosts: adapters pass only authorized ownership, status returns only `{ state, version }`, and `resolveAgent()` supplies current agent/revision. Remote restart requires both checkpoint and session stores to be durable. A crash after that mark is ambiguous and is never replayed automatically; use host tool idempotency keyed by `runId`/`toolCallId` or resolve it manually. Checkpoints contain bounded redacted state plus session/leaf references, never provider objects, callbacks, signals, credentials, or raw secrets. Only built-in loop options are durable; custom `AgentLoopStrategy` rejects before provider work.
|
|
190
|
+
Resume requires exact checkpoint ownership, version, agent fingerprint, and revision. Prism CAS-claims approval before work, rechecks normal guardrail/permission/validation/limit paths, and marks a pending tool dispatched before its side effect. `createAgentRunLifecycle()` wraps the same core path for server/MCP hosts: adapters pass only authorized ownership, status returns only `{ state, version }`, and `resolveAgent()` supplies current agent/revision. `resumeStream()` uses that same claim path and bounded subscriber, so adapters do not poll or duplicate resume logic. Remote restart requires both checkpoint and session stores to be durable. A crash after that mark is ambiguous and is never replayed automatically; use host tool idempotency keyed by `runId`/`toolCallId` or resolve it manually. Checkpoints contain bounded redacted state plus session/leaf references, never provider objects, callbacks, signals, credentials, or raw secrets. Only built-in loop options are durable; custom `AgentLoopStrategy` rejects before provider work.
|
|
184
191
|
|
|
185
192
|
## Secure composition
|
|
186
193
|
|
|
@@ -205,6 +212,7 @@ Per-run options may narrow `limits` and append `guardrails`; they cannot replace
|
|
|
205
212
|
- [CLI/RPC](cli-rpc.md): terminal and JSONL adapters over this runtime.
|
|
206
213
|
- [Workflows](workflows.md): optional DAG orchestration that calls `AgentSession.run()` for agent nodes.
|
|
207
214
|
- [A2A interoperability](a2a.md): direct text exposure calls `AgentSession.run()`; durable/rich/reconnect behavior uses host `A2ATaskLifecycle` over existing checkpoints/persistence, never an in-memory runtime cache.
|
|
215
|
+
- [Frontend interoperability (AG-UI and ACP)](ag-ui.md): optional adapters use `session.stream()` and `resumeAgentRunStream()` / `AgentRunLifecycle.resumeStream()`; protocol/UI state remains outside core.
|
|
208
216
|
|
|
209
217
|
`AgentConfig.loop` and `RunOptions.loop` select a replaceable per-run control loop (`singleShotLoop` default, or `generate-validate-revise` with host callbacks); see [Agent loops](agent-loops.md). `RunOptions.loop` wins over `AgentConfig.loop`. Built-in loops emit the same normal turn/message envelope around provider turns, and both add the first run input to live history once after the first provider turn so later turns see the same transcript shape.
|
|
210
218
|
|
package/docs/cli-rpc.md
CHANGED
|
@@ -106,7 +106,7 @@ Branch-aware session commands return live handle details:
|
|
|
106
106
|
|
|
107
107
|
`sessionId` identifies the durable session. `leafId` is the selected branch tip. `handleId` is the RPC map key used by `switchSession`; forks that share the same `sessionId` get stable ids like `session-1#2` so the parent handle is not overwritten.
|
|
108
108
|
|
|
109
|
-
Invalid CLI flags return exit code `2`. Invalid JSON, missing ids, unknown RPC commands,
|
|
109
|
+
Invalid CLI flags return exit code `2`. Invalid JSON, missing ids, unknown RPC commands, unknown command contributions, and runtime failures return `ok: false` response envelopes without executing unknown tools or commands. `steer` with no active run (or overflow) returns `ok: false`.
|
|
110
110
|
|
|
111
111
|
## Request/response example
|
|
112
112
|
|
|
@@ -128,6 +128,7 @@ Invalid CLI flags return exit code `2`. Invalid JSON, missing ids, unknown RPC c
|
|
|
128
128
|
- `state`, `messages`, `setModel`, `switchSession`, `forkSession`, `cloneSession`, `checkout`, and registered `command` requests are processed immediately.
|
|
129
129
|
- `compact` is fail-closed: if the current session has an active run, it returns `ok: false` because the session rejects compaction during a run.
|
|
130
130
|
- A second `prompt` or `followUp` for the same session while it already has an active run returns `ok: false` immediately instead of blocking the input loop.
|
|
131
|
+
- `steer` enqueues mid-run user text for the active session (`params.input`, optional `params.softInterrupt`). Fails closed when no active run or when the pending steer queue overflows (8 messages / 64 KiB). Soft interrupt aborts the current provider stream only; the run continues.
|
|
131
132
|
|
|
132
133
|
Events streamed during a run keep the original prompt request id, even when an `abort` with a different request id cancels the run. The completion or error response for the prompt also uses the original prompt request id.
|
|
133
134
|
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
## What it does
|
|
4
4
|
|
|
5
|
-
`@arnilo/prism-coding-agent` is an optional first-party package that provides host shell/filesystem/repository tools as Prism `ToolDefinition` objects. It ships six default coding tools — `shell`, `read`, `write`, `edit`, `repo_list`, `repo_search` — plus
|
|
5
|
+
`@arnilo/prism-coding-agent` is an optional first-party package that provides host shell/filesystem/repository tools as Prism `ToolDefinition` objects. It ships six default coding tools — `shell`, `read`, `write`, `edit`, `repo_list`, `repo_search` — plus opt-in structured Git/check set (`createGitTools`), opt-in `createAskUserDecisionTool({ ask })`, and bounded coding-plan/checkpoint helpers. The tools are **inert** until a host imports them and registers them into a `ToolRegistry`. Hosts may register any subset, omit aggregators entirely, or mix first-party tools with host-owned `ToolDefinition`s. Behavior for shell/read/write/edit is a behavioral port of the pi coding agent's tools, adapted to Prism's `ToolDefinition` / `ToolResult` contracts (no `@earendil-works/*` or `typebox` dependencies; only `diff` plus the Node standard library). List/search/Git are native Prism tools with no glob/ripgrep/Git-library dependency.
|
|
6
6
|
|
|
7
7
|
| Export | Purpose |
|
|
8
8
|
| --- | --- |
|
|
@@ -17,6 +17,7 @@
|
|
|
17
17
|
| `createAllTools(cwd, options?)` | Identical to `createCodingTools` (Git tools remain opt-in via `createGitTools`). |
|
|
18
18
|
| `createGitTools(cwd, options?)` | Opt-in Git tools (`git_status`/`git_diff`/`git_branch`/`git_worktree`/`git_apply`/`git_commit`/`git_pr_handoff`) plus optional `coding_check`. |
|
|
19
19
|
| `createCodingCheckTool(cwd, options)` | Named host-declared checks; model selects only a name. |
|
|
20
|
+
| `createAskUserDecisionTool(options)` | Opt-in user decision tool (`ask_user_decision`); host supplies `ask` callback. Not in default aggregators. |
|
|
20
21
|
| `createLocalRepositoryOperations(limits?)` | Default streaming Node filesystem backend for list/search. |
|
|
21
22
|
| `createGitOperations(options)` | Typed Git operations backend (argument arrays, safe config, finite output). |
|
|
22
23
|
| `buildCodingCheckpointMetadata` / `validateCodingCheckpointMetadata` / `assertCodingResumeAllowed` | Bounded durable coding-task metadata for workflow `state.coding` (no second runtime). |
|
|
@@ -231,6 +232,72 @@ const gitTools = createGitTools(workspaceRoot, {
|
|
|
231
232
|
});
|
|
232
233
|
```
|
|
233
234
|
|
|
235
|
+
### Ask-user decision (`createAskUserDecisionTool`)
|
|
236
|
+
|
|
237
|
+
Opt-in `ask_user_decision` for ambiguous, high-impact direction choices. Model must pass a question plus 2+ options, each with **exactly 3 pros and 3 cons**. Host supplies `ask` (blocks until the user picks). Not in `createCodingTools` / `createAllTools` / `createReadOnlyTools`.
|
|
238
|
+
|
|
239
|
+
| Mode | How |
|
|
240
|
+
| --- | --- |
|
|
241
|
+
| Single (default) | `selectionMode: "single"` → host returns `{ selectedId }` (or length-1 `selectedIds`) |
|
|
242
|
+
| Multi | `selectionMode: "multiple"` → `{ selectedIds: [...] }` (non-empty, known ids) |
|
|
243
|
+
| Free-text | `allowCustom: true` → host may return `{ customText }` **XOR** selection (never both) |
|
|
244
|
+
| Blocking tool | `createAskUserDecisionTool({ ask })` — in-process UI callback |
|
|
245
|
+
| Durable workflow | `suspendAskUserDecision(request)` + `createAskUserDecisionResumeValidator()` / `validateAskUserDecisionResume` on `resumeWorkflow` |
|
|
246
|
+
| Agent durable adapter | `validateAskUserDecisionAgentResume({ request, answer })` — same validation; **no** new `AgentRunInterruption` kinds in 0.0.11 |
|
|
247
|
+
|
|
248
|
+
Custom-text caps match question defaults (2 KiB / hard 8 KiB). Options default max 6 (hard 16).
|
|
249
|
+
|
|
250
|
+
```ts
|
|
251
|
+
import { createToolRegistry } from "@arnilo/prism";
|
|
252
|
+
import {
|
|
253
|
+
createAskUserDecisionTool,
|
|
254
|
+
createCodingTools,
|
|
255
|
+
suspendAskUserDecision,
|
|
256
|
+
createAskUserDecisionResumeValidator,
|
|
257
|
+
} from "@arnilo/prism-coding-agent";
|
|
258
|
+
|
|
259
|
+
const tools = createToolRegistry([
|
|
260
|
+
...createCodingTools(workspaceRoot),
|
|
261
|
+
createAskUserDecisionTool({
|
|
262
|
+
ask: async ({ question, options, selectionMode, allowCustom }) =>
|
|
263
|
+
ui.ask({ question, options, selectionMode, allowCustom }),
|
|
264
|
+
}),
|
|
265
|
+
]);
|
|
266
|
+
|
|
267
|
+
// Workflow node:
|
|
268
|
+
return suspendAskUserDecision({
|
|
269
|
+
question: "Ship sqlite or postgres?",
|
|
270
|
+
options: [/* ≥2 with 3 pros + 3 cons each */],
|
|
271
|
+
selectionMode: "single",
|
|
272
|
+
allowCustom: false,
|
|
273
|
+
});
|
|
274
|
+
// resumeWorkflow(..., { validateResume: createAskUserDecisionResumeValidator() })
|
|
275
|
+
```
|
|
276
|
+
|
|
277
|
+
### Goal → verify helper (`runCodingGoalVerify`)
|
|
278
|
+
|
|
279
|
+
Thin composition over existing plan Markdown, named checks, workflow `suspend`/`resumeWorkflow`, and bounded PR handoff. **No Goal table / second runtime.** Peer `@arnilo/prism-workflows`. Example: `examples/coding-goal-verify.ts`.
|
|
280
|
+
|
|
281
|
+
```ts
|
|
282
|
+
import { runCodingGoalVerify } from "@arnilo/prism-coding-agent";
|
|
283
|
+
|
|
284
|
+
const result = await runCodingGoalVerify({
|
|
285
|
+
goal: "Fix the flake",
|
|
286
|
+
cwd: process.cwd(),
|
|
287
|
+
taskId: "flake-1",
|
|
288
|
+
baseBranch: "main",
|
|
289
|
+
branch: "fix/flake",
|
|
290
|
+
checkNames: ["test"],
|
|
291
|
+
checkDefinitions: { test: { file: "/usr/bin/npm", args: ["test"] } },
|
|
292
|
+
runCheck: hostRunCheck,
|
|
293
|
+
buildHandoff: hostBuildHandoff,
|
|
294
|
+
approval: { validateResume: hostValidate },
|
|
295
|
+
checkpoints,
|
|
296
|
+
ownership,
|
|
297
|
+
redactor,
|
|
298
|
+
});
|
|
299
|
+
```
|
|
300
|
+
|
|
234
301
|
### Durable coding plans and checkpoints
|
|
235
302
|
|
|
236
303
|
There is no `CodingRun`, todo database, or second approval engine. Persist executable plan/todos as ordinary workspace Markdown (for example `plans/<task>.md`) and store only bounded metadata under workflow `state.coding`:
|
|
@@ -312,6 +379,7 @@ const remoteWrite = createWriteTool("/repo", {
|
|
|
312
379
|
|
|
313
380
|
## Extension and configuration notes
|
|
314
381
|
|
|
382
|
+
- **Long coding sessions.** Use `createCodingCompactionStrategy()` from optional `@arnilo/prism-compaction-llm` when history needs a bounded coding handoff. It is selected explicitly through normal `session.compact()` / agent compaction configuration, preserves raw session entries, and prioritizes file paths, patch intent, checks, plan/todo state, blockers, and verification steps. It does not read files, retain full diffs, or create a second coding runtime.
|
|
315
383
|
- **Pluggable operation backends.** Every tool accepts an `operations` seam. Custom `ReadOperations` must implement bounded `readText` plus `statFile`; custom `EditOperations` must implement `statFile`; read/write methods receive caps/signals. `BashOperations` must stream through `onData` and honor `signal`/`timeout`. Custom `RepositoryOperations` must honor depth/entry/file/match/scan/time caps and abort. A hostile custom backend can still violate its host-owned contract, so isolate it separately.
|
|
316
384
|
- **Per-tool options.** `ShellToolOptions` adds `timeout` and `maxTotalOutputBytes`; `ReadToolOptions` adds `maxScanBytes`; `WriteToolOptions` adds `maxInputBytes`; `EditToolOptions` adds `maxFileBytes`, `maxInputBytes`, and `maxEdits`; list/search accept `repository` limits and shared aggregator `ToolsOptions.repository`.
|
|
317
385
|
- **Aggregator options.** `ToolsOptions` (`{ executionPolicy?, shell?, read?, write?, edit?, list?, search?, repository? }`) threads each sub-object to the matching tool. `createCodingTools()`, `createAllTools()`, and `createReadOnlyTools()` apply the shared policy unless that tool has an explicit per-tool override. Read-only membership is deliberately `read` + `repo_list` + `repo_search` (0.0.9 behavior change).
|
|
@@ -358,3 +426,4 @@ Every configurable value is a positive safe integer (context may be zero); Prism
|
|
|
358
426
|
- [Public contracts](public-contracts.md): `ToolDefinition`, `ToolResult`, `ToolExecutionContext`, `ContentBlock`, and `JsonObject` shapes.
|
|
359
427
|
- [Host security guide](host-security.md): fail-closed checklist for permission policies, tool validation, and trust boundaries that must gate these tools.
|
|
360
428
|
- [Tool conformance](tool-conformance.md): assertions for the tool-dispatch blocked-reason matrix these tools participate in.
|
|
429
|
+
- [LLM compaction package](compaction-llm.md): optional `createCodingCompactionStrategy()` retains bounded paths, patch intent, checks, plan/todo state, blockers, and next verification—not complete diffs or raw command output.
|
|
@@ -149,7 +149,7 @@ Retry policies are ordinary `RetryPolicy` implementations and can be registered
|
|
|
149
149
|
|
|
150
150
|
Compaction strategies are ordinary `CompactionStrategy` implementations. Extensions can register strategies through the existing compaction strategy contribution registry, but registration is inert until a host explicitly selects and passes a strategy to runtime code. Extensions can also register `compaction` middleware; the runtime calls it only when the agent/session has that middleware registry configured.
|
|
151
151
|
|
|
152
|
-
The default strategy does not call a provider. Hosts that need model-generated summaries can use the optional [`@arnilo/prism-compaction-llm` package](compaction-llm.md); its `maxOutputTokens`/`maxSummaryTokens` budget is passed through `model.parameters.maxTokens` and first-party providers serialize that to provider output-token fields. Hosts that need prepared source-backed memory without a compaction-time model call can use [`@arnilo/prism-compaction-observational-memory`](compaction-observational-memory.md).
|
|
152
|
+
The default strategy does not call a provider. Hosts that need model-generated summaries can use the optional [`@arnilo/prism-compaction-llm` package](compaction-llm.md); its `maxOutputTokens`/`maxSummaryTokens` budget is passed through `model.parameters.maxTokens` and first-party providers serialize that to provider output-token fields. Coding sessions can select that package's `createCodingCompactionStrategy()` preset for paths, patch intent, checks, plans/todos, blockers, and next verification steps; it remains an ordinary `CompactionStrategy` and does not retain complete diffs or add a coding runtime. Hosts that need prepared source-backed memory without a compaction-time model call can use [`@arnilo/prism-compaction-observational-memory`](compaction-observational-memory.md).
|
|
153
153
|
|
|
154
154
|
## Security and performance notes
|
|
155
155
|
|
|
@@ -173,5 +173,6 @@ The default strategy does not call a provider. Hosts that need model-generated s
|
|
|
173
173
|
- [Configuration and manifests](configuration-and-manifests.md): `compactionStrategy` and `retryPolicy` manifest contribution kinds.
|
|
174
174
|
- [Provider layer](provider-layer.md): safe provider error codes used by retry classification.
|
|
175
175
|
- [Credentials and redaction](credentials-and-redaction.md): exact secret redaction helper used by default compaction and retry error handling.
|
|
176
|
+
- [LLM compaction package](compaction-llm.md): `createCodingCompactionStrategy()` is the thin coding-focused preset; see `examples/coding-compaction.ts` for a network-free mock.
|
|
176
177
|
|
|
177
178
|
Runtime redaction composes with compaction and retry secret lists: configured redactors apply at session serialization boundaries, while compaction/retry `secrets` still redact their local summaries and errors.
|
package/docs/compaction-llm.md
CHANGED
|
@@ -12,6 +12,7 @@ Key exports:
|
|
|
12
12
|
| Export | Purpose |
|
|
13
13
|
| --- | --- |
|
|
14
14
|
| `createLlmCompactionStrategy(options)` | Returns a provider-backed `CompactionStrategy`. |
|
|
15
|
+
| `createCodingCompactionStrategy(options)` | Fixed `coding` preset over the LLM strategy: prioritizes file paths, patch intent, commands/checks, plan/todos, blockers, and next verification while retaining normal limits and raw history. |
|
|
15
16
|
| `createLlmCompactionExtension(options)` | Registers the strategy into an explicit extension kernel compaction registry. |
|
|
16
17
|
| `prepareLlmCompaction(context, options?)` | Splits branch entries into summary input, kept suffix, optional split-turn prefix, file details, and compaction data. |
|
|
17
18
|
| `findLlmCompactionCutPoint(entries, options?)` | Finds the last entry covered by a summary using approximate token budgets. |
|
|
@@ -76,6 +77,23 @@ const strategy = createLlmCompactionStrategy({
|
|
|
76
77
|
await session.compact({ strategy, secrets: [apiKey] });
|
|
77
78
|
```
|
|
78
79
|
|
|
80
|
+
Coding-session example:
|
|
81
|
+
|
|
82
|
+
```ts
|
|
83
|
+
import { createCodingCompactionStrategy } from "@arnilo/prism-compaction-llm";
|
|
84
|
+
|
|
85
|
+
const strategy = createCodingCompactionStrategy({
|
|
86
|
+
provider: summaryProvider,
|
|
87
|
+
summaryModel: { provider: "openai", model: "gpt-4.1-mini" },
|
|
88
|
+
keepRecentTokens: 20_000,
|
|
89
|
+
maxSummaryTokens: 800,
|
|
90
|
+
customInstructions: "Keep migration blockers prominent.",
|
|
91
|
+
});
|
|
92
|
+
await session.compact({ strategy });
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
The preset always uses strategy name `coding` and enables existing read/modified-file retention. It adds no provider call, parser, worker, filesystem access, or complete-diff retention beyond `createLlmCompactionStrategy()`.
|
|
96
|
+
|
|
79
97
|
Credential factory example:
|
|
80
98
|
|
|
81
99
|
```ts
|
|
@@ -108,7 +126,7 @@ Preparation is O(n) over branch entries and uses only arrays, strings, and JSON
|
|
|
108
126
|
|
|
109
127
|
Provider deltas are redacted while retained and stop at `maxSummaryTokens * 4` UTF-16 code units without splitting a surrogate pair. Provider iteration is closed/aborted on overflow. A derived finite event ceiling also stops endless empty/non-text deltas. Final history/turn/file composition receives the same cap. Provider error events, generator throws, provider-factory failures, and policy failures expose only bounded redacted detail; host abort remains authoritative.
|
|
110
128
|
|
|
111
|
-
The strategy makes only the needed provider call(s): one history summary plus one split-turn prefix summary when needed.
|
|
129
|
+
The strategy makes only the needed provider call(s): one history summary plus one split-turn prefix summary when needed. The coding preset makes the same calls and uses the same bounded file-operation preparation. Neither discovers credentials, reads files, starts background jobs, or adds provider SDK dependencies. Redaction is exact-string only; pass every known secret that may appear in history, paths, instructions, or provider output.
|
|
112
130
|
|
|
113
131
|
## Related APIs
|
|
114
132
|
|
|
@@ -119,3 +137,4 @@ The strategy makes only the needed provider call(s): one history summary plus on
|
|
|
119
137
|
- [Agent/session runtime](agent-session-runtime.md): `AgentSession.compact()` and opt-in auto-compaction.
|
|
120
138
|
- [Provider layer](provider-layer.md): mock providers and provider request contracts.
|
|
121
139
|
- [Credentials and redaction](credentials-and-redaction.md): exact known-secret redaction behavior.
|
|
140
|
+
- [Frontend interoperability (AG-UI and ACP)](ag-ui.md): separate optional frontend transport; coding compaction adds no UI protocol dependency.
|
|
@@ -205,7 +205,7 @@ const providers = createOpenAIProviderPackage({ apiKey });
|
|
|
205
205
|
- Use distinct `namespace` or vault paths per tenant/environment.
|
|
206
206
|
- Keychain `list()` / `listOAuth()` are intentionally unsupported — enumerate credentials through host configuration instead of scanning the OS store.
|
|
207
207
|
- Combine with `createExplicitCredentialResolver()` so runtime overrides still win over stored values.
|
|
208
|
-
- Wire `createOAuthCredentialStoreAdapter(store)` into `refreshOAuthCredential()`
|
|
208
|
+
- Wire `createOAuthCredentialStoreAdapter(store)` into `refreshOAuthCredential()` only for an OAuth flow explicitly selected by the host and authorized by that provider. In 0.0.12 that means OpenAI Codex; Anthropic and Google packages accept API keys only. Never import or migrate Claude Code/Gemini CLI credential files, setup tokens, browser sessions, or CLI OAuth rows into this store.
|
|
209
209
|
|
|
210
210
|
## Security and performance notes
|
|
211
211
|
|
|
@@ -216,6 +216,7 @@ const providers = createOpenAIProviderPackage({ apiKey });
|
|
|
216
216
|
- Keychain operations use `@napi-rs/keyring`'s abort-aware `AsyncEntry`, so native work runs outside the JavaScript event loop. A main-loop timer aborts and rejects at `timeoutMs`; native cancellation remains OS/backend-dependent and may briefly retain one libuv worker after rejection.
|
|
217
217
|
- Keychain payloads are bytes rather than password strings and are zeroed after parse/write. Unknown native errors are mapped to sanitized typed errors; no native message or secret value is echoed.
|
|
218
218
|
- Never log passphrases, derived keys, or decrypted credential payloads.
|
|
219
|
+
- Storage is not OAuth eligibility. A durable store may persist credentials for a provider only after the host selects a provider-authorized flow; it must not be used to piggyback on a vendor CLI or consumer subscription.
|
|
219
220
|
- Live keychain tests are opt-in (`PRISM_TEST_KEYCHAIN=1`); default `npm test` stays offline.
|
|
220
221
|
|
|
221
222
|
## MCP authentication boundary
|
|
@@ -102,6 +102,14 @@ console.log(error.message);
|
|
|
102
102
|
- Keep resolved credential values local to the request path. Do not put them in registries, model configs, messages, provider events, agent events, session entries, compaction summaries, or logs.
|
|
103
103
|
- Future settings/config loaders may provide credential resolver instances, but core helpers remain storage-free.
|
|
104
104
|
|
|
105
|
+
### Subscription OAuth eligibility
|
|
106
|
+
|
|
107
|
+
In 0.0.12, OpenAI Codex is Prism's only first-party subscription OAuth flow. It is explicit and host-invoked through `createOpenAICodexOAuthProvider()`; hosts own login UI and may use `createOAuthCredentialStoreAdapter()` for deliberately selected durable storage.
|
|
108
|
+
|
|
109
|
+
Anthropic and Google provider packages are API-key-only. Do not scrape or import Claude Code/Gemini CLI credential files, setup tokens, environment values, or browser sessions, and do not route a user's Claude.ai/Gemini subscription through Prism. Anthropic states that developers building products must use Claude Console API keys or a supported cloud provider and may not offer Claude.ai login or route Free/Pro/Max credentials ([legal and compliance](https://docs.anthropic.com/en/docs/claude-code/legal-and-compliance)). Gemini CLI states that third-party software using its OAuth to access backend services violates applicable terms; its FAQ names Vertex AI or Google AI Studio API keys as the supported third-party path ([terms](https://github.com/google-gemini/gemini-cli/blob/main/docs/resources/tos-privacy.md), [FAQ](https://github.com/google-gemini/gemini-cli/blob/main/docs/resources/faq.md)).
|
|
110
|
+
|
|
111
|
+
A future provider-local OAuth adapter needs published permission for third-party products, documented authorize/token/refresh endpoints and scopes, PKCE/state where required, abort/expiry/bounded-response/redaction/store-round-trip fixtures, and legal review before registration. Until then, absence is intentional.
|
|
112
|
+
|
|
105
113
|
## Security and performance notes
|
|
106
114
|
|
|
107
115
|
- Redaction only removes exact known secret values passed to the helper. It is not a general-purpose secret detector.
|
package/docs/evaluations.md
CHANGED
|
@@ -152,5 +152,5 @@ Fixtures reuse `@arnilo/prism-evals` (`defineDataset` / `defineScorer` / `scoreR
|
|
|
152
152
|
- [Runs and usage ledger](runs-and-usage.md): run/session identity for score linkage
|
|
153
153
|
- [Observability](observability.md): use `onTraceReference` or bounded `traceId(runId)` to supply `ScoreRunOptions.traceId`; evaluation telemetry emits no reason/explanation content
|
|
154
154
|
- [Coding agent tools](coding-agent-tools.md) / [Browser automation](browser-automation.md) / [Workflows](workflows.md): network-free coding-task composition at `examples/durable-coding-workflow.ts`; adversarial coding/browser eval example at `examples/coding-browser-evaluation.ts`
|
|
155
|
-
- [Performance limits](performance.md): `scripts/benchmark-0.0.10.mjs` workspace-mode evidence and `scripts/benchmark-0.0.9.mjs` coding/browser evidence fields
|
|
155
|
+
- [Performance limits](performance.md): `scripts/benchmark-0.0.11.mjs` search/budget evidence, `scripts/benchmark-0.0.10.mjs` workspace-mode evidence, and `scripts/benchmark-0.0.9.mjs` coding/browser evidence fields
|
|
156
156
|
- [Release and install](release-and-install.md): optional package install and protected sandbox-browser workflow
|
package/docs/host-security.md
CHANGED
|
@@ -117,6 +117,7 @@ Wire those values where they matter: provider adapters receive the resolved cred
|
|
|
117
117
|
- Resolve credentials at the provider/request edge, as late as possible. Do not put resolved credentials in configs, manifests, registries, prompts, messages, events, session entries, run ledgers, idempotency keys, cache keys, or logs.
|
|
118
118
|
- Use `createExplicitCredentialResolver()` to document source order such as runtime override → stored credential → caller-supplied env object → fallback.
|
|
119
119
|
- Use `createEnvCredentialResolver()` only with an object the host passes in. Prism does not read `process.env` for credentials.
|
|
120
|
+
- Do not treat a local vendor CLI credential file, setup token, browser session, or consumer subscription as a Prism credential source. In 0.0.12 only OpenAI Codex has a first-party subscription OAuth flow; Anthropic and Google providers remain API-key-only under their published third-party restrictions. See [Credentials and redaction](credentials-and-redaction.md#subscription-oauth-eligibility).
|
|
120
121
|
- Use `createPathTrustPolicy()` for workspace/resource roots and fail closed on symlink escapes.
|
|
121
122
|
- Use `createContributionRegistries({ duplicate: "error" })` and prefixed names for third-party packages to prevent silent shadowing.
|
|
122
123
|
- Extension contributions are inert until selected. Loading an extension package runs its `setup(api)` code, so hosts should load only trusted packages or isolate untrusted code outside Prism.
|
|
@@ -190,6 +191,7 @@ PostgreSQL TLS/network policy, MCP endpoint trust/credentials and egress policy
|
|
|
190
191
|
## Related APIs
|
|
191
192
|
|
|
192
193
|
- [Web-standard server handler](server.md): remote agent/workflow route, ownership, limits, abort, and deployment boundary.
|
|
194
|
+
- [Frontend interoperability (AG-UI and ACP)](ag-ui.md): authorize every protocol selector/operation; default-deny tool/state/path projection; exact interrupt/version resume; redacted, ownership-scoped replay.
|
|
193
195
|
- [Supervisor delegation](supervisors.md): local child permission/memory/budget boundary.
|
|
194
196
|
- [A2A interoperability](a2a.md): remote card/auth/origin/signature boundary.
|
|
195
197
|
- [Settings, auth, trust, and security controls](settings-auth-trust-security.md): low-level helpers and boundary hardening table.
|