@arnilo/prism 0.0.7 → 0.0.10
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +50 -2
- package/README.md +3 -1
- package/dist/agent-loops.js +14 -8
- package/dist/agents.js +37 -3
- package/dist/contracts.d.ts +17 -0
- package/dist/index.d.ts +4 -2
- package/dist/index.js +3 -2
- package/dist/provider-events.d.ts +2 -0
- package/dist/provider-events.js +21 -13
- package/dist/providers/openai-compatible.js +8 -5
- package/dist/providers/transport.d.ts +10 -1
- package/dist/providers/transport.js +24 -8
- package/dist/run-ledger.d.ts +21 -0
- package/dist/run-ledger.js +115 -0
- package/dist/tools.js +2 -0
- package/docs/a2a.md +61 -42
- package/docs/agent-events.md +5 -4
- package/docs/agent-loops.md +2 -2
- package/docs/agent-session-runtime.md +1 -0
- package/docs/browser-automation.md +124 -0
- package/docs/coding-agent-tools.md +111 -14
- package/docs/coding-security.md +84 -11
- package/docs/credential-storage.md +9 -0
- package/docs/database-persistence.md +1 -1
- package/docs/evaluations.md +38 -4
- package/docs/guardrails.md +3 -2
- package/docs/host-security.md +29 -4
- package/docs/index.md +20 -15
- package/docs/mcp-tools.md +29 -5
- package/docs/migration.md +104 -0
- package/docs/observability.md +26 -14
- package/docs/performance.md +54 -0
- package/docs/postgres-persistence.md +1 -0
- package/docs/provider-conformance.md +1 -1
- package/docs/provider-primitives.md +7 -1
- package/docs/providers/kimi.md +16 -2
- package/docs/providers/opencode-go.md +43 -2
- package/docs/release-and-install.md +117 -62
- package/docs/resource-loading.md +4 -0
- package/docs/review-coverage-2026-07-19-phase-3.md +174 -0
- package/docs/review-coverage-2026-07-20-phase-4.md +175 -0
- package/docs/review-coverage-2026-07-21-phase-5.md +172 -0
- package/docs/run-ledger-conformance.md +1 -0
- package/docs/runs-and-usage.md +17 -2
- package/docs/sqlite-persistence.md +1 -0
- package/docs/structured-output.md +2 -2
- package/docs/supervisors.md +2 -2
- package/docs/tools.md +5 -1
- package/docs/web-tools.md +78 -0
- package/docs/workflows.md +2 -0
- package/package.json +6 -4
package/docs/migration.md
CHANGED
|
@@ -7,6 +7,110 @@ Prism 0.0.6 preserves documented 0.0.3 agent construction except for two intenti
|
|
|
7
7
|
1. **`session.run()` / `session.prompt()` return `AgentRunResult`** and `session.stream()` starts one owned run after subscribing. Callers that ignored the previous `Promise<void>` keep working; failed/aborted runs reject with `AgentRunError` (`.result` attached).
|
|
8
8
|
2. **`AgentConfig.extensions` / `settings` / `credentials` are removed.** Wire extensions through `createExtensionKernel()`, read settings in the host, and pass credential resolvers to the provider edge.
|
|
9
9
|
|
|
10
|
+
## 0.0.9 / 0.0.96 → 0.0.10 coding workspace modes (breaking composition)
|
|
11
|
+
|
|
12
|
+
`@arnilo/prism-coding-security` composition now requires explicit `workspaceMode: "host" | "sandbox"`. Missing mode throws at construction. The `0.0.9` default that wired sandbox shell while keeping read/write/edit/list/search on the host cwd is **superseded** and fail-closed.
|
|
13
|
+
|
|
14
|
+
| Before (0.0.9) | After (0.0.10) |
|
|
15
|
+
| --- | --- |
|
|
16
|
+
| `createSandboxCodingTools(cwd, { sandbox })` — shell in sandbox, FS on host | Must pass `workspaceMode`. Prefer `createSandboxCodingComposition(...)`. |
|
|
17
|
+
| Silent split-brain treated as normal | Throws unless `allowMixedWorkspaceWiring: true` (warnings; `containmentClaim: false`). |
|
|
18
|
+
| No containment metadata | `composition.containmentClaim` / `warnings` / optional `treeIdentity`. Host mode never claims containment. |
|
|
19
|
+
|
|
20
|
+
```ts
|
|
21
|
+
// Contained: one disposable tree
|
|
22
|
+
const { tools, composition } = createSandboxCodingComposition(sourceRoot, {
|
|
23
|
+
workspaceMode: "sandbox",
|
|
24
|
+
sandbox, // DisposableSandbox auto-wires FS backends
|
|
25
|
+
});
|
|
26
|
+
|
|
27
|
+
// Explicit host (non-contained)
|
|
28
|
+
createSandboxCodingTools(cwd, { workspaceMode: "host" });
|
|
29
|
+
|
|
30
|
+
// Escape hatch (documented split; no containment claim)
|
|
31
|
+
createSandboxCodingTools(cwd, {
|
|
32
|
+
workspaceMode: "sandbox",
|
|
33
|
+
sandbox,
|
|
34
|
+
allowMixedWorkspaceWiring: true,
|
|
35
|
+
});
|
|
36
|
+
|
|
37
|
+
// Same-tree Git
|
|
38
|
+
createGitTools(composition.workspaceRoot, {
|
|
39
|
+
execFile: sandbox.execFile.bind(sandbox),
|
|
40
|
+
commitIdentity: { name: "bot", email: "bot@example.com" },
|
|
41
|
+
});
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
Docker defaults unchanged: digest-pinned image, non-root user, network none, absolute Docker CLI, no host-env inheritance. Unified mode adds no unbounded sync; caps stay in `sandbox-limits.ts` / coding-agent limits. Benchmark evidence: `scripts/benchmark-0.0.10.mjs`.
|
|
45
|
+
|
|
46
|
+
## 0.0.8 → 0.0.9 release overview
|
|
47
|
+
|
|
48
|
+
All 32 first-party manifests and exact internal ranges move together to `0.0.9`; mixed first-party versions are unsupported. Core remains dependency-free at runtime and existing low-level agent/session APIs remain compatible. New coding sandbox, repository/Git, durable coding-plan, and browser surfaces are opt-in. `@arnilo/prism-browser` is included by `@arnilo/prism-all` but not by `@arnilo/prism-code` — install it explicitly when interactive browser automation is required. Office execution remains outside Prism packaging (host-selected skills/instructions only). No tag or publication is automatic from this migration.
|
|
49
|
+
|
|
50
|
+
### Malformed streamed tool-call arguments (recoverable)
|
|
51
|
+
|
|
52
|
+
Malformed streamed tool-call JSON (id+name present) no longer terminates the run as `ProviderTransportError("invalid_json_arguments")`. First-party providers emit a tool call carrying `argumentsError`; dispatch blocks with `tool_execution_blocked` / `invalid_arguments` (`error.code: "invalid_json_arguments"`), never calls `execute()`, and the model can self-correct within existing turn/tool-round budgets. Prefer `toolCallFromArgumentsText` / `tryParseJsonObjectArguments` in custom providers.
|
|
53
|
+
|
|
54
|
+
### Incomplete tool-call deltas (typed failure)
|
|
55
|
+
|
|
56
|
+
Tool-call deltas missing `id` and/or `name` at stream end no longer throw a bare `Error("Incomplete tool call delta...")`. Core reconstruction and the openai-compatible finalizer surface `ProviderTransportError` / `ErrorInfo.code: "incomplete_delta"`, fail the provider turn (no tool execution), and keep OpenCode Go / Kimi dangling fail-closed behavior. Distinguish from Defect 1a: missing identity fails the turn; present identity with bad JSON recovers via failed tool results.
|
|
57
|
+
|
|
58
|
+
### Empty call-free artifact candidates (parse_error)
|
|
59
|
+
|
|
60
|
+
`generateValidateReviseLoop` treats empty/whitespace-only call-free assistant text (including thinking-only/reasoning-only turns) as `parse_error` before the host parser/identity default. Session runs succeed only after `artifact_finished`; terminal `artifact_failed` fails the run (`AgentRunError`, typically `error.code: "parse_error"`).
|
|
61
|
+
|
|
62
|
+
## 0.0.9 coding-security Docker sandbox (additive)
|
|
63
|
+
|
|
64
|
+
`@arnilo/prism-coding-security` adds `createDockerSandbox()` / `DisposableSandbox` while preserving `SandboxAdapter.exec` and `createSandboxBashOperations()`. Hosts opt in with an absolute Docker executable and digest-pinned image; default network is none, host env is never inherited, and workspace export is an explicit bounded host callback. Existing approval-policy callers need no changes.
|
|
65
|
+
|
|
66
|
+
## 0.0.9 coding-agent repository list/search (additive behavior change)
|
|
67
|
+
|
|
68
|
+
`@arnilo/prism-coding-agent` adds native `repo_list` / `repo_search` tools. `createCodingTools()` / `createAllTools()` now return six tools. **`createReadOnlyTools()` deliberately expands from `[read]` to `[read, repo_list, repo_search]`** — update hosts that asserted the previous read-only membership. Prefer `createSandboxCodingComposition(cwd, { workspaceMode, sandbox, repository })` (or the tools-only wrappers) from `@arnilo/prism-coding-security`. Pass required `workspaceMode`; sandbox mode keeps shell and FS/list/search on one disposable tree. The 0.0.9 split (sandbox shell + host FS) is superseded — see **0.0.9 / 0.0.96 → 0.0.10 coding workspace modes** above.
|
|
69
|
+
|
|
70
|
+
Opt-in structured Git/check tools are available via `createGitTools(cwd, { commitIdentity, checks? })` and are **not** added to `createCodingTools()`/`createAllTools()`. Commits require an explicit host `commitIdentity`; PR handoff returns bounded metadata/artifacts only and never pushes.
|
|
71
|
+
|
|
72
|
+
Durable coding-task composition uses existing workflows plus coding-agent helpers (`writeCodingPlanFile`, `buildCodingCheckpointMetadata`, `assertCodingResumeAllowed`). Plan/todos remain workspace Markdown; checkpoint state keeps only references/hashes/summaries/fingerprints under `state.coding`. No `CodingRun` or todo database is introduced. See `examples/durable-coding-workflow.ts`.
|
|
73
|
+
|
|
74
|
+
## 0.0.9 browser automation (additive)
|
|
75
|
+
|
|
76
|
+
Install `@arnilo/prism-browser` explicitly (or through `@arnilo/prism-all`) for interactive browser tools. Hosts supply a pinned Playwright `Browser` (`playwright-core@1.61.0` optional peer); package import launches and downloads nothing. `createBrowserTools()` returns exactly `browser_open`, `browser_snapshot`, `browser_act`, and `browser_close` (all `exclusive: true`). Network policy defaults to require contained-proxy attestation; configure `uploads`/`downloads` for file transfer; `browser_act` adds `upload`/`screenshot`/`download_release`. Use `createBrowserManager().closeRun(runId)` / `close()` on terminal/abort. Align with a disposable sandbox via `createSharedSandboxBrowserOptions()` and `assertBrowserSandboxNetwork()`. CSS/XPath/evaluate/CDP/persistent profiles remain unsupported.
|
|
77
|
+
|
|
78
|
+
## 0.0.7 → 0.0.8 release overview
|
|
79
|
+
|
|
80
|
+
All 31 first-party manifests and exact internal ranges move together to `0.0.8`; mixed first-party versions are unsupported. Core remains dependency-free at runtime and existing low-level agent/session APIs remain compatible. New telemetry, evaluation, MCP, A2A, ledger batching, and web research surfaces are opt-in. Release CI now requires CodeQL, dependency/license/SBOM/secret checks, packed-artifact attestations, PostgreSQL integration, and protected live-canary prerequisites; no tag or publication is automatic from this migration.
|
|
81
|
+
|
|
82
|
+
## 0.0.7 → 0.0.8 evaluations and ledger operation
|
|
83
|
+
|
|
84
|
+
`@arnilo/prism-evals` adds owner-scoped trace resolution, optional host model judges, deterministic pairwise reports, and `assertEvaluationThreshold()` without changing stored evaluation schemas. Hosts select all judge/provider credentials and should version rubrics. Core adds optional `createBatchedRunLedger()`; direct ledgers remain write-through. Choose `flush_on_terminal` only after accepting bounded pre-flush crash loss, and call `dispose()` during shutdown. Runtime snapshot caching is session/leaf-local and requires no persistence migration.
|
|
85
|
+
|
|
86
|
+
## 0.0.7 → 0.0.8 web research tools
|
|
87
|
+
|
|
88
|
+
Install `@arnilo/prism-web-tools` explicitly (or through `@arnilo/prism-all`) to add web capability; core and existing profiles remain inert. Select Brave or Exa at construction, provide Firecrawl separately for Markdown/schema extraction, and register returned `web_search`/`web_fetch`/`web_extract` tools through normal permission/trust/validation dispatch. Provider selection, credentials, target DNS policy, and extraction schema are host-only. All returned content is marked untrusted; no browser or vendor SDK is added.
|
|
89
|
+
|
|
90
|
+
## 0.0.7 → 0.0.8 A2A durable tasks
|
|
91
|
+
|
|
92
|
+
Existing text `createA2AHandler({ exposure })`, `client.send()`, and `client.stream()` remain compatible. Add host `tasks` to enable `GetTask`/`ListTasks`/`CancelTask`/`SubscribeToTask`, rich parts, interrupted states, and replay cursors; no task store or migration is created. Add host `push` for push-config CRUD and matching card capability. Raw/data/URL parts are disabled until selected in `parts`; URL/push endpoints additionally require host URL policy and are never fetched by part parsing. Push delivery/retries/idempotency remain host-owned.
|
|
93
|
+
|
|
94
|
+
## 0.0.7 → 0.0.8 MCP capabilities and sessions
|
|
95
|
+
|
|
96
|
+
`@arnilo/prism-mcp` now pins official SDK 1.29.0. Existing `connectMcpTools()` and stateless web handlers remain compatible. Use `connectMcpCapabilities()` for bounded resources/prompts and explicit roots/sampling/elicitation callbacks. Server resources/prompts must be selected explicitly and authorize every operation. Stateful Streamable HTTP additionally requires `sessionIdGenerator`, exact `allowedOrigins`, and host `resolveIdentity`; omission preserves stateless mode. `Last-Event-ID` replay is not enabled. Missing capability calls fail with `ERR_PRISM_MCP_UNSUPPORTED_CAPABILITY`.
|
|
97
|
+
|
|
98
|
+
## 0.0.7 → 0.0.8 OpenTelemetry adapter
|
|
99
|
+
|
|
100
|
+
The optional observability package now emits OTel GenAI names and units instead of independent `prism.agent.run` / `prism.provider.turn` / `prism.tool.execute` spans and millisecond metrics. Update dashboards to `invoke_agent prism`, `chat {model}`, `execute_tool {tool}`, `gen_ai.*.duration` (seconds), and `gen_ai.client.token.usage`. Pass `{ context, trace }` as third `wrapOpenTelemetryApi()` argument for native parent context, and use `onTraceReference` or `traceId(runId)` for evaluation linkage. Core APIs and persistence schemas are unchanged.
|
|
101
|
+
|
|
102
|
+
## 0.0.7 → 0.0.8 Kimi provider alignment
|
|
103
|
+
|
|
104
|
+
`@arnilo/prism-provider-kimi` now matches the official contracts: featured Coding `k3` defaults to `reasoning_effort: "high"` (Open Platform `kimi-k3` keeps `"max"`); featured context windows use the official `262_144` for 256K-class models; the featured Moonshot catalog adds `kimi-k2.7-code-highspeed`, `kimi-k2.6`, and `kimi-k2.5` (K2.5 intentionally without Preserved Thinking). Provider-owned compat keys (`route`, `preserveThinking`, `preserve_thinking`) are stripped before the opaque compat spread and no longer leak into request bodies. The Coding route additionally sends provider-owned `x-api-key` and `anthropic-version: 2023-06-01` headers per the official third-party setup. Streams emit `done` only on protocol completion evidence (`message_stop` on the Coding route, `[DONE]` + `finish_reason` on the Moonshot route); truncated streams now surface as run failures.
|
|
105
|
+
|
|
106
|
+
## 0.0.7 → 0.0.8 artifact-loop parse failures
|
|
107
|
+
|
|
108
|
+
`generateValidateReviseLoop` no longer returns silently on artifact parse failure. A parser returning `{ ok: false }` (or no `value`) now consumes revision budget exactly like a validation failure: the repairer receives `value: undefined` plus a synthetic failure (`metadata.reason: "parse_error"`), and exhaustion ends with terminal `artifact_failed`. Host repairers must already tolerate `value: undefined` per the `ArtifactRepairer` contract; runs that previously ended after one silent parse failure now spend up to `maxRevisions` repair turns first.
|
|
109
|
+
|
|
110
|
+
## 0.0.7 → 0.0.8 OpenCode Go provider fixes
|
|
111
|
+
|
|
112
|
+
`@arnilo/prism-provider-opencode-go` no longer infers `structuredOutput: "json_schema"` from OpenAI-compatible routing alone. Only verified models (`mimo-v2.5`, `mimo-v2.5-pro`) advertise it; other OpenAI-route models (for example `deepseek-v4-pro`) now use the artifact-loop parsing/validation path, and requests that still pass `options.structuredOutput` for an unverified model fail before dispatch with `unsupported_model`. Hosts with their own verification evidence can set the capability explicitly through `defineOpenCodeGoModel({ capabilities })`. The Anthropic route additionally sends provider-owned `x-api-key` and `anthropic-version: 2023-06-01` headers alongside Bearer, fixing HTTP 401 on MiniMax/Qwen models; caller headers cannot override them. Streams now emit `done` only on protocol completion evidence (`[DONE]` plus a terminal `finish_reason` on the OpenAI route, `message_stop` on the Anthropic route) with no dangling tool-call accumulators; truncated connections and incomplete tool calls terminate with an `error` event, so hosts may see previously silent truncations surface as run failures.
|
|
113
|
+
|
|
10
114
|
## 0.0.6 → 0.0.7 secure run lifecycle
|
|
11
115
|
|
|
12
116
|
`createAgent()` remains backward-compatible. Version 0.0.7 adds opt-in typed `Guardrails` (`input`, provider `output`, `toolInput`, `toolOutput`) and narrowing-only `RunLimits`. Output guardrails and configured output-token/total-token/cost limits buffer provider output before exposure; blocked content is neither emitted nor persisted. A breach emits one redacted `run_limit_exceeded` event and rejects with `AgentRunError.result.limit`.
|
package/docs/observability.md
CHANGED
|
@@ -52,8 +52,17 @@ OpenTelemetry adapter:
|
|
|
52
52
|
import { trace, metrics } from "@opentelemetry/api";
|
|
53
53
|
import { createOpenTelemetryInstrumentation, wrapOpenTelemetryApi } from "@arnilo/prism-observability-opentelemetry";
|
|
54
54
|
|
|
55
|
-
const { tracer, meter } = wrapOpenTelemetryApi(
|
|
56
|
-
|
|
55
|
+
const { tracer, meter } = wrapOpenTelemetryApi(
|
|
56
|
+
trace.getTracer("app"),
|
|
57
|
+
metrics.getMeter("app"),
|
|
58
|
+
{ context, trace },
|
|
59
|
+
);
|
|
60
|
+
const telemetry = createOpenTelemetryInstrumentation({
|
|
61
|
+
tracer,
|
|
62
|
+
meter,
|
|
63
|
+
onTraceReference: ({ runId, traceId }) => saveRunTrace(runId, traceId),
|
|
64
|
+
onExporterError: console.error,
|
|
65
|
+
});
|
|
57
66
|
|
|
58
67
|
const detach = telemetry.attachSession(session);
|
|
59
68
|
// or: for await (const event of session.subscribe()) telemetry.handleAgentEvent(event);
|
|
@@ -79,13 +88,13 @@ OpenTelemetry mapping (when enabled):
|
|
|
79
88
|
|
|
80
89
|
| Agent event | Span | Metric labels |
|
|
81
90
|
| --- | --- | --- |
|
|
82
|
-
| `agent_started` /
|
|
83
|
-
| `provider_turn_*` | `
|
|
84
|
-
| `tool_execution_*`
|
|
85
|
-
| `
|
|
86
|
-
| `
|
|
87
|
-
| `handleRunFeedback` | active-run `prism.run.feedback` event or ended-run span | `prism.run.feedback`
|
|
88
|
-
| `handleEvaluation` | active-run `
|
|
91
|
+
| `agent_started` / terminal event | `invoke_agent prism` (`INTERNAL`) | `gen_ai.invoke_agent.duration` |
|
|
92
|
+
| `provider_turn_*` | `chat {model}` (`CLIENT`) | `gen_ai.client.operation.duration`, `gen_ai.client.token.usage` |
|
|
93
|
+
| `tool_execution_*` | `execute_tool {tool}` (`INTERNAL`) when started | `gen_ai.execute_tool.duration` |
|
|
94
|
+
| `guardrail_decision` | `prism.guardrail.evaluate` child (`INTERNAL`) | none |
|
|
95
|
+
| `handleDelegation()` | `prism.agent.delegate` child (`INTERNAL`) | none |
|
|
96
|
+
| `handleRunFeedback` | active-run `prism.run.feedback` event or ended-run span | `prism.run.feedback` |
|
|
97
|
+
| `handleEvaluation` | active-run `gen_ai.evaluation.result` event or ended-run span | `prism.run.evaluation` (`status`) |
|
|
89
98
|
|
|
90
99
|
High-cardinality identifiers (`sessionId`, `runId`, `requestId`, `toolCallId`) are **span attributes only**, never metric labels.
|
|
91
100
|
|
|
@@ -135,11 +144,11 @@ const session = createAgent({
|
|
|
135
144
|
|
|
136
145
|
const detach = telemetry.attachSession(session);
|
|
137
146
|
const result = await session.run("hello");
|
|
147
|
+
const traceId = telemetry.traceId(result.runId); // or persist onTraceReference immediately
|
|
138
148
|
detach();
|
|
139
149
|
telemetry.handleRunFeedback({ runId: result.runId, rating: 1, hasComment: true, tagCount: 1, scorerCount: 1, evaluationCount: 1 });
|
|
140
|
-
telemetry.handleEvaluation({ runId: result.runId, status: "scored", score: 0.9, hasReason: true });
|
|
141
|
-
|
|
142
|
-
console.log(memory.spans.map((span) => span.name));
|
|
150
|
+
telemetry.handleEvaluation({ runId: result.runId, name: "citation", status: "scored", score: 0.9, hasReason: true });
|
|
151
|
+
console.log(traceId, memory.spans.map((span) => span.name));
|
|
143
152
|
```
|
|
144
153
|
|
|
145
154
|
## Extension and configuration notes
|
|
@@ -149,14 +158,17 @@ console.log(memory.spans.map((span) => span.name));
|
|
|
149
158
|
- NeuralWatt `neuralwatt:telemetry` provider events remain package-local; hosts may forward numeric cost/energy into custom metrics.
|
|
150
159
|
- `@arnilo/prism-observability-opentelemetry` is optional and included through `@arnilo/prism-sdk` and `@arnilo/prism-all`; instrumentation remains disabled until a host configures it.
|
|
151
160
|
- Exporter failures are isolated: instrumentation catches tracer/meter errors and invokes `onExporterError` without affecting the run, feedback persistence, or evaluation scoring.
|
|
152
|
-
-
|
|
161
|
+
- Trace grading uses `createPersistenceTraceResolver()` with explicit session/run/ownership and finite pages/bytes. Judge reasons remain evaluation data; `gen_ai.evaluation.result` receives only name, finite score, controlled status, and reason-presence.
|
|
162
|
+
- Run spans parent provider, tool, guardrail, and explicit delegation spans. Pass `{ context, trace }` to `wrapOpenTelemetryApi()` for native parent context creation; `parentContext` can attach the run to host ambient/remote context.
|
|
163
|
+
- `onTraceReference` receives `{ runId, traceId }` when a run starts. `traceId(runId)` keeps only the newest 1,024 mappings by default (`maxTraceReferences`, hard cap 10,000); durable linkage remains host-owned.
|
|
164
|
+
- Run `error`, suspension, denial, and detach close every attributable span. Repeated terminal events are idempotent and cannot end a span twice.
|
|
153
165
|
- Disabled instrumentation performs no per-delta span work (`enabled: false` or missing tracer/meter).
|
|
154
166
|
|
|
155
167
|
## Security and performance notes
|
|
156
168
|
|
|
157
169
|
- Default events are metadata-only — no prompts, streamed deltas, tool arguments, or credentials.
|
|
158
170
|
- Opt-in content in other event types (`message_delta`, tool `result`) is still subject to `redactAgentEvent`.
|
|
159
|
-
- Metric labels stay low-cardinality (`
|
|
171
|
+
- Metric labels stay low-cardinality (`gen_ai.operation.name`, `gen_ai.provider.name`, token type, controlled outcome/status, feedback rating bucket/link presence); never use session/run/request/call IDs, model output, comments, tag values, scorer/evaluation IDs, or arbitrary metadata as labels. Token usage is recorded once at provider operation scope.
|
|
160
172
|
- Target overhead when enabled is under 5% excluding exporter I/O; disabled hooks allocate no spans.
|
|
161
173
|
- Provider transport limits and redaction order are documented in [Provider primitives](provider-primitives.md).
|
|
162
174
|
|
package/docs/performance.md
CHANGED
|
@@ -1,9 +1,63 @@
|
|
|
1
1
|
# Performance limits
|
|
2
2
|
|
|
3
|
+
Evaluation defaults are finite: 100 trace rows × 20 pages and 4 MiB aggregate trace data; one model-judge attempt with 30-second/16-KiB bounds; 8 comparison candidates, 1-MiB candidate results, 10,000 dataset items, and 4-MiB serialized reports. Hard caps are exported by `@arnilo/prism-evals`; overflow fails rather than truncating grading evidence.
|
|
4
|
+
|
|
3
5
|
## What it does
|
|
4
6
|
|
|
5
7
|
This page states Prism runtime limits that keep slow consumers and long sessions from becoming unbounded memory or latency problems.
|
|
6
8
|
|
|
9
|
+
## Release 0.0.10 reproducible workspace-mode evidence
|
|
10
|
+
|
|
11
|
+
Run `node scripts/benchmark-0.0.10.mjs`; `PRISM_BENCH_ITERATIONS` accepts 10–100,000 (default 100). Schema/bounds test: `node --test scripts/benchmark-0.0.10.test.mjs`. Default mode is network-free: host-composition write/read/list plus sandbox-fake composition write/read/list/search (in-memory `DisposableSandbox`). Emits environment, scenario mode, throughput, p50/p95 latency, heap, disk bytes, process counts, zero external cost, backpressure, and resource-limit signals. Optional `PRISM_BENCH_DOCKER=1` (with `PRISM_TEST_DOCKER_*`) appends real local Docker composition rows. Unified workspace mode reuses existing sandbox/repo hard caps and adds no unbounded host↔container sync. These are evidence fields, not CI timing gates.
|
|
12
|
+
|
|
13
|
+
## Release 0.0.9 reproducible coding/browser evidence
|
|
14
|
+
|
|
15
|
+
Run `node scripts/benchmark-0.0.9.mjs`; `PRISM_BENCH_ITERATIONS` accepts 10–100,000 (default 100). Schema/bounds test: `node --test scripts/benchmark-0.0.9.test.mjs`. Default mode is network-free fake/in-process only and emits environment, scenario mode, throughput, p50/p95 latency, heap, disk bytes, process counts, zero external cost, backpressure, and resource-limit signals for repository list/search, Git status, and browser open/snapshot/action/close. Optional `PRISM_BENCH_DOCKER=1` (with `PRISM_TEST_DOCKER_*`) and `PRISM_BENCH_PLAYWRIGHT=1` append real local Docker / protected Playwright rows. These are evidence fields, not CI timing gates.
|
|
16
|
+
|
|
17
|
+
2026-07-21 baseline: Node v24.18.0, Linux x64, 100 iterations/scenario, network=false, credentials=false, docker=false, playwright=false.
|
|
18
|
+
|
|
19
|
+
| Scenario | mode | ops/s | p95 ms | heap bytes | disk bytes | processes | cost USD | backpressure | resource limits |
|
|
20
|
+
| --- | --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
|
|
21
|
+
| repo-list | fake-in-process | 1,343 | 1.11 | 14,747,560 | 0 | 1 | 0 | 0 | 0 |
|
|
22
|
+
| repo-search | fake-in-process | 380 | 3.72 | 16,093,176 | 0 | 1 | 0 | 0 | 0 |
|
|
23
|
+
| git-status | fake-in-process | 479 | 2.62 | 13,821,688 | 0 | 1 | 0 | 0 | 0 |
|
|
24
|
+
| browser-open-snapshot-action-close | fake-in-process | 17,141 | 0.11 | 18,590,488 | 0 | 1 | 0 | 0 | 0 |
|
|
25
|
+
|
|
26
|
+
Rows exercise shipped repository/Git helpers and fake Playwright APIs only. Real Docker sandbox and Playwright browser timings remain explicit protected-gate evidence (`PRISM_TEST_DOCKER_SANDBOX=1`, `PRISM_LIVE_PLAYWRIGHT=1` / `PRISM_BENCH_DOCKER=1` / `PRISM_BENCH_PLAYWRIGHT=1`) because this release-candidate host did not enable those gates for the dated baseline. No live claim is inferred from skipped gates.
|
|
27
|
+
|
|
28
|
+
## Release 0.0.8 reproducible synthetic evidence
|
|
29
|
+
|
|
30
|
+
Run `node scripts/benchmark-0.0.8.mjs`; `PRISM_BENCH_ITERATIONS` accepts 10–100,000. Script uses no network/credentials and emits environment, throughput, p50/p95 latency, heap, synthetic disk bytes, zero external cost, and backpressure signals. These are evidence fields, not CI timing gates.
|
|
31
|
+
|
|
32
|
+
2026-07-20 baseline: Node v24.18.0, Linux x64, 1,000 operations/scenario.
|
|
33
|
+
|
|
34
|
+
| Scenario | ops/s | p95 ms | heap bytes | disk bytes | cost USD | backpressure |
|
|
35
|
+
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
|
|
36
|
+
| provider envelope | 675,430 | 0.0026 | 10,974,296 | 0 | 0 | 0 |
|
|
37
|
+
| actual `createBatchedRunLedger` enqueue/flush | 514,493 | 0.0023 | 12,601,840 | 0 | 0 | 7 |
|
|
38
|
+
| one-entry snapshot-cache hit | 4,582,216 | 0.0001 | 10,072,656 | 0 | 0 | 0 |
|
|
39
|
+
| actual in-memory OTel agent span start/end | 295,372 | 0.0049 | 10,357,048 | 0 | 0 | 0 |
|
|
40
|
+
| PostgreSQL-ledger-shaped file workload | 494,403 | 0.0009 | 11,584,760 | 54,890 | 0 | 7 |
|
|
41
|
+
| MCP envelope | 1,243,254 | 0.0008 | 12,915,288 | 0 | 0 | 0 |
|
|
42
|
+
| A2A envelope | 961,492 | 0.0007 | 10,425,024 | 0 | 0 | 0 |
|
|
43
|
+
| web-tools envelope | 1,388,694 | 0.0007 | 11,734,208 | 0 | 0 | 0 |
|
|
44
|
+
|
|
45
|
+
Ledger and OTel rows exercise shipped implementations; cache row isolates the runtime's one-entry lookup shape. Provider/PostgreSQL/MCP/A2A/web rows remain local serialization/file envelopes and prove repeatability/schema/backpressure instrumentation only—not external latency, throughput, or billing. Real PostgreSQL correctness runs in protected CI; provider/MCP/A2A/web timings and costs remain explicit protected live-canary/release-host evidence because this release-candidate host has no credentials/endpoints. No live claim is inferred from skipped gates.
|
|
46
|
+
|
|
47
|
+
Security automation is isolated from `npm test`: CodeQL/supply-chain jobs have 10-minute backstops, dependency review and live workflow have 5-minute job backstops, live probe step has 3 minutes, SBOM is capped at 16 MiB/10,000 packages, packed release/security tarballs at 128 MiB aggregate, secret scan at 100,000 files/16 MiB each, and retained security/canary reports expire after 7 days. Live canaries issue four probes plus at most one MCP cleanup, cap responses at 64 KiB and requests at 15 seconds (30 seconds hard), and never enter `sdk:ready`.
|
|
48
|
+
|
|
49
|
+
Web tools default/hard ceilings are query 4/16 KiB, results 10/20, URLs 5/20, request 256 KiB/1 MiB, response/aggregate 2/16 MiB, Markdown 1/8 MiB, extraction 256 KiB/1 MiB, schema 64/256 KiB, concurrency 4/16, retries 2/4, polling 20/100, and wall time 60 seconds/30 minutes. Bounds charge before request, retention, retry, or polling; overflow fails rather than truncating citation/extraction evidence.
|
|
50
|
+
|
|
51
|
+
Docker sandbox defaults/hard caps from `@arnilo/prism-coding-security`: startup 30 s/120 s; wall 20 min/30 min; idle 5 min/15 min; CPUs 2/8; memory 2 GiB/16 GiB (swap equal to memory); PIDs 256/1,024; FDs 1,024/8,192; workspace/tmp/download tmpfs 1 GiB/8 GiB, 256 MiB/2 GiB, 64 MiB/512 MiB; commands 100/256 with concurrent execs 1/8; env 64/256 names and 64 KiB/256 KiB values; export 50,000/250,000 entries and 256 MiB/2 GiB bytes with 16/64 retained artifacts; stop grace 5 s/30 s and cleanup 30 s/120 s. Caps validate before `docker create`/exec/export; overflow aborts and cleans the recorded container. Output still streams into the coding-agent `OutputAccumulator` ceilings (64 MiB/1 GiB).
|
|
52
|
+
|
|
53
|
+
Repository list/search defaults/hard caps from `@arnilo/prism-coding-agent`: depth 32/128; entries/files 10,000/100,000; page/results 1,000/10,000; search scan 64 MiB/1 GiB aggregate and 8 MiB/64 MiB per file; matches 1,000/10,000; pattern 512 B/4 KiB; line 50 KiB/1 MiB; context 5/20; wall 30 s/300 s; concurrency config 8/32. Walks stream via `opendir`/`lstat`, never follow symlink escapes, and stop immediately on aggregate limits or abort.
|
|
54
|
+
|
|
55
|
+
Structured Git/check/handoff defaults/hard caps: paths 1,000/10,000; refs 1 KiB/4 KiB; commit message 64 KiB/256 KiB; inline Git output 4 MiB/64 MiB; diff lines 10,000/100,000; changed files 1,000/10,000; patch input 16 MiB/64 MiB; worktrees 4/16; named checks 8/32 names, concurrency 1/4, timeout 10 min/60 min, diagnostic lines 2,000/100,000, output 4 MiB/64 MiB; PR handoff JSON 256 KiB/1 MiB with 100/1,000 commits. Git tools use typed argument arrays (never shell), disable hooks/credential prompts/external diff by default, and emit host-owned PR handoff data only — no push/network/PR client.
|
|
56
|
+
|
|
57
|
+
Durable coding plan/checkpoint defaults/hard caps: plan Markdown 256 KiB/1 MiB; todos 1,000/10,000 with 512 B/4 KiB text; checkpoint metadata 64 KiB/512 KiB; artifact references 16/64 at 256 MiB/2 GiB each; check summaries 1 KiB/8 KiB. Checkpoints store URI/hash/summaries/fingerprints only; resume revalidates workspace root, base branch, plan hash, and tool/policy/image fingerprints before import.
|
|
58
|
+
|
|
59
|
+
Browser automation defaults/hard caps from `@arnilo/prism-browser`: pages 4/16; actions 100/256; queued actions 16/64; snapshot refs 2,000/10,000; depth 30/100; snapshot bytes 256 KiB/2 MiB; navigation 30 s/120 s; action 10 s/60 s; wait 30 s/120 s; run wall 20 min/30 min; popups 4/16; dialogs 16/64; listeners 64/256; action input 64 KiB/256 KiB; close grace 5 s/30 s; network requests 1,000/10,000 with 10/32 redirects per request and 8/32 WebSockets; screenshots 16/64 with 16/64 megapixels and 10 MiB/32 MiB encoded; uploads 8/32 files, 16 MiB/64 MiB each, 64 MiB/256 MiB aggregate; downloads 8/32 files, 32 MiB/256 MiB each, 64 MiB/512 MiB aggregate. Caps charge before context/page/action/queue/snapshot/network/artifact retention. Host supplies Playwright and egress proxy attestation; package import launches nothing.
|
|
60
|
+
|
|
7
61
|
Current surfaces:
|
|
8
62
|
|
|
9
63
|
- `SubscribeOptions` for bounded live `AgentEvent` subscriber queues.
|
|
@@ -125,6 +125,7 @@ PRISM_TEST_POSTGRES_URL="$DATABASE_URL" npm run test:postgres --workspace @arnil
|
|
|
125
125
|
- **Identifier validation.** Configurable `schema` is validated and double-quoted; table names are fixed constants in adapter SQL.
|
|
126
126
|
- **TLS and credentials.** Configure via `pg` `Pool` / `PoolConfig`; the adapter does not read environment variables unless the host passes them into `connectionString` or `poolConfig`.
|
|
127
127
|
- **Redaction upstream.** Event and tool-call payloads may contain secrets; redact before ledger writes. The adapter does not scan or rewrite row contents.
|
|
128
|
+
- **Optional batching.** PostgreSQL remains write-through by default. Hosts may wrap its ledger with core `createBatchedRunLedger()`; the bounded FIFO retains a failed head record and propagates flush failure instead of silently acknowledging it.
|
|
128
129
|
- **Bounded pool.** Adapter-owned pools default to `max: 10`. Hosts with heavy concurrency should supply their own pool sizing.
|
|
129
130
|
- **Indexed operations.** Append, parent validation, idempotency dedup, branch reads, and pagination use the indexes documented in [Database persistence](database-persistence.md); normal paths avoid sequential scans.
|
|
130
131
|
- **Migration locking.** `pg_advisory_xact_lock` prevents concurrent migration races when multiple processes open the adapter at once. Startup catalog reads are bounded metadata queries, not application-row scans.
|
|
@@ -45,7 +45,7 @@ Helpers accept normal `AIProvider`, `ProviderRequest`, `ProviderEvent`, `Usage`,
|
|
|
45
45
|
- `collectProviderEvents()` returns provider events in stream order.
|
|
46
46
|
- `assertProviderStreamConforms()` returns collected events after verifying the stream ends with `done` or `error`, terminal events are last, and optional text/usage expectations match.
|
|
47
47
|
- `assertAbortIsObserved()` passes an already-aborted signal and expects provider generation to reject. This is the supported timeout primitive; use a host abort controller or `RunOptions.signal` rather than deprecated provider-level `timeoutMs`.
|
|
48
|
-
- `assertToolCallDeltasReconstruct()` rebuilds streamed `tool_call_delta` fragments into tool calls and validates expected id/name/arguments. The runtime uses the same reconstruction
|
|
48
|
+
- `assertToolCallDeltasReconstruct()` rebuilds streamed `tool_call_delta` fragments into tool calls and validates expected id/name/arguments. Malformed JSON with id+name present yields `argumentsError` (no throw); missing id/name throws typed `incomplete_delta`. The runtime uses the same reconstruction before tool execution when a provider streams deltas.
|
|
49
49
|
- `assertUsageAccounting()` finds `usage` or `done.usage` and checks selected token fields including `cacheReadTokens` and `cacheWriteTokens`. This is the provider-neutral check for normalized cache read/write token extraction; every first-party provider package exercises it against server-specific fields (`cached_tokens`, `cache_read_input_tokens`, etc.).
|
|
50
50
|
- `assertSerializedRequestCoversContent()` scans a serialized provider request body for primitive canaries from each Prism content block and fails if any supported block type is silently dropped. Provider-valid transcripts place assistant `tool_call` messages before matching role `tool` `tool_result` messages; runtime, cache-aware input layout, and observational-memory worker loops preserve that order before serialization.
|
|
51
51
|
- `assertProviderOwnedHeadersWin()` compares captured request headers against the provider's authoritative owned header values and a caller-supplied header bag; it fails if any owned header (`authorization`, `content-type`, session/security headers) was overridden by caller headers, and also fails if a non-owned caller header was dropped. This is the provider-neutral check that caller `ProviderRequest.options.headers` cannot hijack provider credentials or sessions; every first-party provider package exercises it.
|
|
@@ -110,7 +110,7 @@ export interface SseEvent {
|
|
|
110
110
|
}
|
|
111
111
|
|
|
112
112
|
export class ProviderTransportError extends Error {
|
|
113
|
-
readonly code: "sse_buffer_overflow" | "sse_event_overflow" | "response_body_overflow" | "aborted";
|
|
113
|
+
readonly code: "sse_buffer_overflow" | "sse_event_overflow" | "response_body_overflow" | "aborted" | "invalid_json_arguments" | "incomplete_delta";
|
|
114
114
|
readonly limitBytes?: number;
|
|
115
115
|
}
|
|
116
116
|
|
|
@@ -131,6 +131,12 @@ export function parseJsonObjectArguments(
|
|
|
131
131
|
text: string,
|
|
132
132
|
options?: { toolName?: string; maxBytes?: number },
|
|
133
133
|
): JsonObject;
|
|
134
|
+
|
|
135
|
+
/** Non-throwing variant for recoverable tool-call recovery; prefer with \`toolCallFromArgumentsText\`. */
|
|
136
|
+
export function tryParseJsonObjectArguments(
|
|
137
|
+
text: string,
|
|
138
|
+
options?: { toolName?: string; maxBytes?: number },
|
|
139
|
+
): { ok: true; value: JsonObject } | { ok: false; error: ProviderTransportError };
|
|
134
140
|
```
|
|
135
141
|
|
|
136
142
|
**Performance:** Single pass over chunks; retained memory is `O(min(buffer, maxBufferBytes))`, not `O(stream)`. No full-stream accumulation.
|
package/docs/providers/kimi.md
CHANGED
|
@@ -70,7 +70,7 @@ Unsupported block placements or unclaimed images fail before fetch.
|
|
|
70
70
|
| --- | --- | --- |
|
|
71
71
|
| Base URL | `https://api.kimi.com/coding` | `https://api.moonshot.ai/v1` (or `.cn`) |
|
|
72
72
|
| Wire API | Anthropic `/messages` | OpenAI `/chat/completions` |
|
|
73
|
-
| Featured ids | `kimi-for-coding`, `kimi-for-coding-highspeed`, `k3` | `kimi-k2.7-code`, `kimi-k3` (+ discovery) |
|
|
73
|
+
| Featured ids | `kimi-for-coding`, `kimi-for-coding-highspeed`, `k3` | `kimi-k2.7-code`, `kimi-k2.7-code-highspeed`, `kimi-k2.6`, `kimi-k2.5`, `kimi-k3` (+ discovery) |
|
|
74
74
|
| Discovery | No public list API — curated featured aliases | Official `GET /v1/models` via `listKimiModels()` |
|
|
75
75
|
| Cache | Implicit by default; opt-in Anthropic `cache_control` | Implicit only — never emits Anthropic `cache_control` |
|
|
76
76
|
| Thinking | Block replay + body `thinking` / `reasoning_effort` | `reasoning_content` replay + body `thinking` / `reasoning_effort` |
|
|
@@ -86,12 +86,26 @@ Official fields (Open Platform docs; Coding docs for `k3` effort mapping):
|
|
|
86
86
|
|
|
87
87
|
| Model family | Official control | Prism `compat` |
|
|
88
88
|
| --- | --- | --- |
|
|
89
|
-
| K3 / Coding `k3` | top-level `reasoning_effort
|
|
89
|
+
| K3 / Coding `k3` | top-level `reasoning_effort`: `"low"`/`"high"`/`"max"` (Open Platform default `"max"`; Kimi Code default `"high"`) | `compat.reasoning_effort` — use Task 4 family `reasoning_effort` |
|
|
90
90
|
| K2.7-code / Coding | thinking always on; Preserved Thinking always on | omit `thinking` by default; `preserveThinking: true` for replay; do not send `disabled` |
|
|
91
91
|
| K2.6 / K2.5 | `thinking.type` enabled/disabled; K2.6 optional `keep: "all"` | `compat.thinking` — Task 4 family `thinking_type` |
|
|
92
92
|
|
|
93
93
|
Per-turn `ProviderRequestOptions.compat` wins over `ModelConfig.compat`. Helpers:
|
|
94
94
|
`kimiThinking`, `kimiReasoningEffort`, `kimiPreserveThinking`.
|
|
95
|
+
`stripKimiThinkingCompat` removes provider-owned routing/serialization keys
|
|
96
|
+
(`route`, `preserveThinking`, `preserve_thinking`, thinking/effort keys) before
|
|
97
|
+
the opaque compat spread, so they never leak into wire bodies.
|
|
98
|
+
|
|
99
|
+
Featured context windows follow the official docs exactly (`262_144` for the
|
|
100
|
+
256K-class models, `1_048_576` for K3). Both stream parsers emit `done` only on
|
|
101
|
+
protocol completion evidence — Coding route: `message_stop` with all `tool_use`
|
|
102
|
+
blocks closed; Moonshot route: `[DONE]` plus a terminal `finish_reason` with no
|
|
103
|
+
dangling tool calls. Truncated streams terminate with an `error` event instead.
|
|
104
|
+
|
|
105
|
+
The Coding route authenticates with provider-owned `authorization: Bearer`,
|
|
106
|
+
`x-api-key`, and `anthropic-version: 2023-06-01` headers (official third-party
|
|
107
|
+
setup uses `ANTHROPIC_API_KEY` semantics); caller-supplied headers cannot
|
|
108
|
+
override them.
|
|
95
109
|
|
|
96
110
|
## Request/response example
|
|
97
111
|
|
|
@@ -57,6 +57,7 @@ createOpenCodeGoProviderPackage(options: OpenCodeGoProviderPackageOptions): Prov
|
|
|
57
57
|
| Surface | Behavior |
|
|
58
58
|
| --- | --- |
|
|
59
59
|
| Provider stream | Prism text, thinking, tool-call delta/final, `usage`, `done`, redacted `error`. |
|
|
60
|
+
| Stream completion | `done` is emitted only on completion evidence — OpenAI route: `[DONE]` marker plus a terminal `finish_reason`; Anthropic route: `message_stop`. Truncated or incomplete streams (including dangling tool-call blocks) end with a terminal `error` instead, so partial output never surfaces as `succeeded`. |
|
|
60
61
|
| OpenAI thinking | `delta.reasoning_content` → thinking deltas; replay via `reasoning_content` when `preserveThinking`. |
|
|
61
62
|
| Anthropic thinking | `thinking_delta` → thinking deltas; replay via Anthropic thinking blocks when `preserveThinking`. |
|
|
62
63
|
| Session/cache | `x-opencode-session` + route-specific cache markers. |
|
|
@@ -72,6 +73,12 @@ createOpenCodeGoProviderPackage(options: OpenCodeGoProviderPackageOptions): Prov
|
|
|
72
73
|
}
|
|
73
74
|
```
|
|
74
75
|
|
|
76
|
+
The Anthropic route (`POST /messages`) additionally sends provider-owned
|
|
77
|
+
`x-api-key: <resolved-key>` and `anthropic-version: 2023-06-01` headers;
|
|
78
|
+
Bearer-only authentication returns HTTP 401 on that route. These headers are
|
|
79
|
+
applied after caller headers and cannot be overridden. The OpenAI route never
|
|
80
|
+
receives Anthropic-only headers.
|
|
81
|
+
|
|
75
82
|
OpenAI-route body (thinking passthrough + preserved reasoning):
|
|
76
83
|
|
|
77
84
|
```json
|
|
@@ -131,6 +138,39 @@ table; Pi secondary metadata is used only for context/output limits when docs om
|
|
|
131
138
|
| `grok-4.5`, `glm-5.2`, `glm-5.1`, `kimi-k3`, `kimi-k2.7-code`, `kimi-k2.6`, `mimo-v2.5`, `mimo-v2.5-pro`, `deepseek-v4-pro`, `deepseek-v4-flash` | `openai` | `implicit` |
|
|
132
139
|
| `minimax-m3`, `minimax-m2.7`, `minimax-m2.5`, `qwen3.7-max`, `qwen3.7-plus`, `qwen3.6-plus` | `anthropic` | `cache_control` |
|
|
133
140
|
|
|
141
|
+
### Structured output capability
|
|
142
|
+
|
|
143
|
+
`capabilities.structuredOutput: "json_schema"` is advertised only for models
|
|
144
|
+
verified against the live gateway to accept JSON Schema `response_format` —
|
|
145
|
+
OpenAI-compatible routing alone never implies support (unverified models such
|
|
146
|
+
as `deepseek-v4-pro` reject it upstream with HTTP 400). Verified models:
|
|
147
|
+
`mimo-v2.5`, `mimo-v2.5-pro`. All other models leave the capability undefined
|
|
148
|
+
and use Prism's artifact-loop parsing/validation path without
|
|
149
|
+
`response_format`; hosts with their own verification evidence can set the
|
|
150
|
+
capability explicitly via `defineOpenCodeGoModel({ capabilities })`.
|
|
151
|
+
|
|
152
|
+
Extending the verified set requires per-model live evidence. Run the
|
|
153
|
+
credential-gated probe:
|
|
154
|
+
|
|
155
|
+
```sh
|
|
156
|
+
PRISM_LIVE_PROVIDER_TESTS=1 OPENCODE_API_KEY=... \
|
|
157
|
+
npm run test --workspace=@arnilo/prism-provider-opencode-go
|
|
158
|
+
```
|
|
159
|
+
|
|
160
|
+
`live_json_schema_structured_output_succeeds_<model>` must pass for a model
|
|
161
|
+
before it joins the set; `live_json_schema_structured_output_rejected_deepseek_v4_pro`
|
|
162
|
+
documents the current boundary (upstream DeepSeek supports only
|
|
163
|
+
`response_format` `"text" | "json_object"`, not `json_schema`) and fails loudly
|
|
164
|
+
if the gateway ever starts accepting `json_schema`, which is the signal to
|
|
165
|
+
extend the set.
|
|
166
|
+
|
|
167
|
+
Discovery is capability-aware: when a `/models` entry carries
|
|
168
|
+
`capabilities.structured_output` (`"json_schema"` or `true`),
|
|
169
|
+
`listOpenCodeGoModels()` honors it as gateway-authoritative — including an
|
|
170
|
+
explicit `false`, which overrides the static verified set. Today's sparse
|
|
171
|
+
payload carries no capability fields, so discovery falls back to the static
|
|
172
|
+
verified set with no behavior change.
|
|
173
|
+
|
|
134
174
|
## Model discovery
|
|
135
175
|
|
|
136
176
|
Official list endpoint (sparse OpenAI-compatible shape):
|
|
@@ -199,8 +239,9 @@ Owned compat keys (`route`, `thinking`, `reasoning`, `reasoning_effort`,
|
|
|
199
239
|
- API keys are resolved per request from caller-supplied values or resolvers and
|
|
200
240
|
redacted from errors (including discovery failures).
|
|
201
241
|
- Caller-supplied `ProviderRequest.options.headers` can add non-owned headers, but
|
|
202
|
-
provider-owned headers (`content-type`, `x-opencode-session`, `authorization
|
|
203
|
-
|
|
242
|
+
provider-owned headers (`content-type`, `x-opencode-session`, `authorization`,
|
|
243
|
+
and on the Anthropic route `x-api-key`/`anthropic-version`) are applied last
|
|
244
|
+
and cannot be overridden by caller headers.
|
|
204
245
|
- Live tests stay opt-in behind `PRISM_LIVE_PROVIDER_TESTS=1` plus `OPENCODE_API_KEY`;
|
|
205
246
|
default tests are network-free.
|
|
206
247
|
|