@arnilo/prism 0.0.13 → 0.0.14
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +12 -0
- package/README.md +8 -2
- package/dist/artifacts.d.ts +78 -0
- package/dist/artifacts.js +24 -0
- package/dist/contracts.d.ts +10 -0
- package/dist/contracts.js +8 -0
- package/dist/conversations.d.ts +50 -0
- package/dist/conversations.js +97 -0
- package/dist/credentials.d.ts +14 -0
- package/dist/credentials.js +9 -0
- package/dist/devices.d.ts +94 -0
- package/dist/devices.js +138 -0
- package/dist/index.d.ts +10 -4
- package/dist/index.js +6 -3
- package/dist/providers/openai-primitives.js +5 -2
- package/docs/ag-ui.md +5 -0
- package/docs/browser-automation.md +3 -0
- package/docs/conversations.md +135 -0
- package/docs/credential-storage.md +28 -1
- package/docs/credentials-and-redaction.md +2 -0
- package/docs/database-persistence.md +5 -1
- package/docs/device-adapters.md +97 -0
- package/docs/host-security.md +4 -2
- package/docs/index.md +16 -12
- package/docs/migration.md +21 -0
- package/docs/performance.md +2 -0
- package/docs/policy-and-audit.md +1 -0
- package/docs/provider-caching.md +4 -0
- package/docs/provider-packages.md +4 -1
- package/docs/providers/alibaba.md +179 -0
- package/docs/providers/ollama.md +166 -0
- package/docs/release-and-install.md +75 -5
- package/docs/review-coverage-2026-07-25-phase-9.md +256 -0
- package/docs/server.md +4 -0
- package/docs/work-artifacts-and-review.md +100 -0
- package/docs/work-connectors.md +5 -1
- package/docs/work-tools.md +3 -0
- package/docs/workflows.md +4 -0
- package/docs/working-and-semantic-memory.md +20 -5
- package/package.json +1 -1
- package/templates/init/providers.json +22 -0
package/docs/index.md
CHANGED
|
@@ -26,13 +26,15 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
|
|
|
26
26
|
- [Compaction and retry policies](compaction-and-retry.md): summarize branch history and retry transient provider failures with host-replaceable policies.
|
|
27
27
|
- [LLM compaction package](compaction-llm.md): optional provider-backed strategy with finite summary/reserve/error caps, bounded redacted streaming retention, mandatory finite post-policy `model.parameters.maxTokens`, and `createCodingCompactionStrategy()` for coding handoff focus.
|
|
28
28
|
- [Observational memory compaction package](compaction-observational-memory.md): optional source-backed memory with owned append callback, finite turn/call/argument/result/transcript/error worker limits, redacted provider-valid transcripts, fast compaction, recall, and status/view commands; worker model falls back to host-supplied `sessionModel`.
|
|
29
|
-
- [Working and semantic memory](working-and-semantic-memory.md): optional `@arnilo/prism-memory` working-memory store, semantic recall, finite Embedder/VectorStore contracts, in-memory adapters,
|
|
29
|
+
- [Working and semantic memory](working-and-semantic-memory.md): optional `@arnilo/prism-memory` working-memory store, semantic recall, finite Embedder/VectorStore contracts, in-memory adapters, PostgreSQL/pgvector path, and consent/source/visibility lifecycle (grant/correct/forget/retention) enforced at injection.
|
|
30
30
|
- [Session stores](session-stores.md): `SessionStore` contract, `SessionAppendOptions`, `SessionAppendConflictError`, branch handles, `readBranchPath`, optional bounded `searchSessions` / `SessionIndex` (memory linear|unsupported), and dev-vs-production branch reads — start here for session persistence.
|
|
31
|
+
- [Conversations](conversations.md): durable user-scoped conversation threads (create/list/continue/branch/archive/export/delete) on session + event-ledger seams, thread-bound reconnectable replay, frozen caps, and legal-hold-aware deletion.
|
|
32
|
+
- [Work artifacts and review](work-artifacts-and-review.md): durable artifact co-work review — authorized attach (MIME/hash/version, producer run, citations, preview metadata), revision compare, approve/reject with last-validated recovery, and authorized expiring delivery links; records persist as versioned checkpoints, never file bodies.
|
|
31
33
|
- [Session stores and branching](session-stores-and-branching.md): detailed branch semantics and helper reference (kept for compatibility; links back to the canonical atomic append / branch-handle sections).
|
|
32
34
|
- [Database persistence](database-persistence.md): production persistence contracts, shared checksummed migration/full-shape catalog primitives (`@arnilo/prism/testing/persistence-schema`), conditional append, indexes, `readBranchPath`, reference relational schema, retention/legal-hold/quota lifecycle (`lifecycle`), and NoSQL mapping.
|
|
33
35
|
- [SQLite persistence](sqlite-persistence.md): optional `better-sqlite3` adapter with session/run storage, checkpoints/leases, feedback, FTS `searchSessions` (migration-v4), and transactionally verified/backfilled migration metadata.
|
|
34
36
|
- [PostgreSQL persistence](postgres-persistence.md): optional pooled `pg` adapter with session/run/checkpoint/lease/feedback storage, FTS `searchSessions` (migration-v4), advisory-locked checksummed/full-shape migrations, and opt-in live conformance.
|
|
35
|
-
- [Migration guide](migration.md): **0.0.13** enterprise identity/policy/router/work connectors, cloud providers, server deployment seams, persistence schema v5
|
|
37
|
+
- [Migration guide](migration.md): **0.0.14** personal/work-agent conversations, memory consent/lifecycle, artifact co-work review, AG-UI co-work events, scoped M365/GWS OAuth connectors, browser checkpoints, device contracts, and Alibaba/Ollama providers; **0.0.13** enterprise identity/policy/router/work connectors, cloud providers, server deployment seams, and persistence schema v5; plus 0.0.12 AG-UI/ACP, 0.0.11 coding-harness fundamentals, 0.0.10 workspace modes, and 0.0.9 coding/browser surfaces.
|
|
36
38
|
- [Node JSONL session store](node-jsonl-session-store.md): development-only JSONL file adapter for single-process Node hosts; no cross-process safety; `searchSessions` throws `SessionSearchUnsupportedError`.
|
|
37
39
|
- [Persistence, credentials, and multimodality primitives](persistence-credentials-multimodality-primitives.md): Plan 056 inventory — session/run-ledger/persistence contracts, credential/OAuth seams, content/resource/model capabilities, package dependency matrix, conformance matrix, and threat model for production adapters.
|
|
38
40
|
|
|
@@ -45,7 +47,7 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
|
|
|
45
47
|
- [Use-case model selection](use-case-model-selection.md): bind `{ model?, provider?, thinkingLevel? }` for observational memory, LLM compaction, and other non-session LLM jobs with explicit session-model fallback via `resolveUseCaseModel`.
|
|
46
48
|
- [Provider request policies](provider-request-policies.md): chain `ProviderRequestPolicy` hooks, use `createSessionCachePolicy`, and merge legacy/structured cache options safely.
|
|
47
49
|
- [Provider packages](provider-packages.md): define explicit provider packages, model metadata, auth descriptors, request/cache policies, provider-owned header precedence, and the 0.0.12 provider-authorized OAuth matrix without package discovery or provider-specific core behavior; includes a first-party cache behavior summary and the **caller-gated on-demand model discovery** contract (`list*Models`, setup zero-fetch).
|
|
48
|
-
- Phase 12 package workspaces: [`@arnilo/prism-provider-openai`](providers/openai.md), [`@arnilo/prism-provider-anthropic`](providers/anthropic.md) (native Messages, `cache_control`, thinking, caller-gated `listAnthropicModels`), [`@arnilo/prism-provider-google`](providers/google.md) (native Gemini `generateContent` SSE, caller-gated `listGoogleModels`), [`@arnilo/prism-provider-opencode-go`](providers/opencode-go.md) (official Go open models, dual-route Anthropic/OpenAI, caller-gated `listOpenCodeGoModels`, `reasoning_content`/thinking preserve), [`@arnilo/prism-provider-openrouter`](providers/openrouter.md) (app-controlled catalog, caller-gated `listOpenRouterModels`, `reasoning` merge/preserve, `cache_control` + sticky `session_id`), [`@arnilo/prism-provider-zai`](providers/zai.md) (official `thinking`/`reasoning_effort`/`tool_stream`, implicit cache, caller-gated `listZaiModels`), [`@arnilo/prism-provider-kimi`](providers/kimi.md), and [`@arnilo/prism-provider-neuralwatt`](providers/neuralwatt.md) with implicit vLLM prefix caching, reasoning controls (`reasoning_effort`/`thinking_token_budget`/`enable_thinking`/`preserve_thinking`/`clear_thinking`), reasoning preservation, OpenAI-style tool-call loop, quota, telemetry, and retry classification helpers.
|
|
50
|
+
- Phase 12 package workspaces: [`@arnilo/prism-provider-openai`](providers/openai.md), [`@arnilo/prism-provider-anthropic`](providers/anthropic.md) (native Messages, `cache_control`, thinking, caller-gated `listAnthropicModels`), [`@arnilo/prism-provider-google`](providers/google.md) (native Gemini `generateContent` SSE, caller-gated `listGoogleModels`), [`@arnilo/prism-provider-opencode-go`](providers/opencode-go.md) (official Go open models, dual-route Anthropic/OpenAI, caller-gated `listOpenCodeGoModels`, `reasoning_content`/thinking preserve), [`@arnilo/prism-provider-openrouter`](providers/openrouter.md) (app-controlled catalog, caller-gated `listOpenRouterModels`, `reasoning` merge/preserve, `cache_control` + sticky `session_id`), [`@arnilo/prism-provider-zai`](providers/zai.md) (official `thinking`/`reasoning_effort`/`tool_stream`, implicit cache, caller-gated `listZaiModels`), [`@arnilo/prism-provider-kimi`](providers/kimi.md), [`@arnilo/prism-provider-alibaba`](providers/alibaba.md) (Alibaba Cloud Model Studio / DashScope + Coding Plan, OpenAI-compatible, caller-gated `listAlibabaModels`, implicit + explicit `cache_control` caching, Qwen `enable_thinking`), [`@arnilo/prism-provider-ollama`](providers/ollama.md) (Ollama Cloud + local, OpenAI-compatible, caller-gated `listOllamaModels`, implicit-only caching, `reasoning_effort`), and [`@arnilo/prism-provider-neuralwatt`](providers/neuralwatt.md) with implicit vLLM prefix caching, reasoning controls (`reasoning_effort`/`thinking_token_budget`/`enable_thinking`/`preserve_thinking`/`clear_thinking`), reasoning preservation, OpenAI-style tool-call loop, quota, telemetry, and retry classification helpers.
|
|
49
51
|
- Phase 8 enterprise cloud (workload identity; separate from consumer Anthropic/Google): [`@arnilo/prism-provider-azure`](providers/azure.md) (Entra / Foundry), [`@arnilo/prism-provider-bedrock`](providers/bedrock.md) (IAM/IRSA + region/PrivateLink), [`@arnilo/prism-provider-vertex`](providers/vertex.md) (ADC / Vertex OpenAPI).
|
|
50
52
|
- Optional AI SDK adapter: [`@arnilo/prism-provider-ai-sdk`](providers/ai-sdk.md) maps host-owned `LanguageModelV4` models onto Prism `AIProvider` streams (specification v4; no Prism catalog; maps `finish.usage` cache read/write tokens; reasoning is host-model-owned).
|
|
51
53
|
- [OpenAI-compatible provider](providers/openai-compatible.md): optional provider subpath using native or injected `fetch` for Chat Completions streaming (`chatCompletionsUrl` / `authStyle` overrides for enterprise adapters).
|
|
@@ -65,9 +67,10 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
|
|
|
65
67
|
- [Tool validator JSON Schema package](../packages/tool-validator-json-schema/README.md): optional `@arnilo/prism-tool-validator-json-schema` adapter for `tool.parameters`.
|
|
66
68
|
- [MCP client bridge and server exposure](mcp-tools.md): SDK-1.29.0 bounded tools/resources/prompts, host-owned roots/sampling/elicitation, exact-origin DNS-pinned client transport, and principal-bound opt-in Streamable HTTP sessions.
|
|
67
69
|
- [Web search, fetch, and extraction](web-tools.md): optional host-selected Brave/Exa discovery and Firecrawl Markdown/schema tools with native fetch, stable citations, late credentials, finite limits, and explicit untrusted-content boundaries.
|
|
68
|
-
- [Work tools](work-tools.md): optional `@arnilo/prism-work-tools` identity-scoped M365 + GWS connectors (hard-coded CLI argv, draft-then-approve, idempotency, shared result shapes).
|
|
69
|
-
- [Work connectors](work-connectors.md): connector principles, capability gates, and out-of-scope boundaries for Microsoft 365 / Google Workspace.
|
|
70
|
-
- [Browser automation](browser-automation.md): optional `@arnilo/prism-browser` with host-supplied Playwright contexts, AI-mode snapshots/refs, ordered `browser_open`/`browser_snapshot`/`browser_act`/`browser_close`, egress/side-effect/upload/download/screenshot policy,
|
|
70
|
+
- [Work tools](work-tools.md): optional `@arnilo/prism-work-tools` identity-scoped M365 + GWS connectors (hard-coded CLI argv, draft-then-approve, idempotency, shared result shapes); 0.0.14 adds a late-bound per-identity `tokenProvider` (env-only, fail-closed).
|
|
71
|
+
- [Work connectors](work-connectors.md): connector principles, capability gates, scoped OAuth establishment (0.0.14), and out-of-scope boundaries (Slack/Teams channels not shipped) for Microsoft 365 / Google Workspace.
|
|
72
|
+
- [Browser automation](browser-automation.md): optional `@arnilo/prism-browser` with host-supplied Playwright contexts, AI-mode snapshots/refs, ordered `browser_open`/`browser_snapshot`/`browser_act`/`browser_close`, egress/side-effect/upload/download/screenshot policy, finite page/action/snapshot/network/artifact caps, and 0.0.14 verified-state checkpoints with reload/verify-before-side-effect.
|
|
73
|
+
- [Device adapters](device-adapters.md): deny-by-default realtime voice / desktop-control contract + conformance (0.0.14); no vendor package — admission fails closed without explicit consent+sandbox+approval, stream bounds, shared `RunLimits`, redacted telemetry.
|
|
71
74
|
- [Coding agent tools](coding-agent-tools.md): optional `shell`, `read`, `write`, `edit`, `repo_list`, and `repo_search` definitions plus opt-in `createGitTools()` / `coding_check`, opt-in `createAskUserDecisionTool` (single/multi/free-text + durable suspend glue), and `runCodingGoalVerify`; durable plan/todo Markdown helpers with workflow `state.coding` checkpoint metadata; streamed text pages, bounded repository list/search, finite Git/check/plan/ask caps, bounded image/edit reads and write/edit payloads, finite shell wall/total-output limits, secure host-owned spill cleanup, pluggable bounded operation contracts, per-path mutation serialization, and optional `ExecutionPolicy`. Limits do not sandbox host access—gate with permission/trust policy and `@arnilo/prism-coding-security`.
|
|
72
75
|
- [Coding execution approval and sandboxing](coding-security.md): path/command approval, identity-scoped caching, shell-turn exclusivity, required `workspaceMode` (`host`/`sandbox`) with fail-closed mixed wiring, `createSandboxCodingComposition()` containment metadata, and the disposable Docker/OCI sandbox reference with bounded workspace import/export.
|
|
73
76
|
|
|
@@ -89,19 +92,19 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
|
|
|
89
92
|
## Multi-agent and interoperability
|
|
90
93
|
- [Supervisor delegation](supervisors.md): optional explicit child allow-list, derived memory scopes, narrowing-only permissions, lifecycle hooks, nested delegation, cancellation, finite budgets, host-projected delegation telemetry, and separate A2A durable adapter boundary.
|
|
91
94
|
- [A2A interoperability](a2a.md): A2A 1.0 JSON-RPC/HTTPS cards plus host-owned durable task get/list/cancel/subscribe, bounded rich parts/replay, principal-scoped push configs, and exact-origin verified client.
|
|
92
|
-
- [Frontend interoperability (AG-UI and ACP)](ag-ui.md): optional `@arnilo/prism-ag-ui` AG-UI mapper/authorized Web handler/replay and stable ACP sibling over shared redacted event and durable-approval seams; no TUI, editor, filesystem, or A2A runtime.
|
|
95
|
+
- [Frontend interoperability (AG-UI and ACP)](ag-ui.md): optional `@arnilo/prism-ag-ui` AG-UI mapper/authorized Web handler/replay and stable ACP sibling over shared redacted event and durable-approval seams; 0.0.14 adds reconnectable co-work events (artifact progress/approval/download-link, connector drafts, redacted browser snapshots); no TUI, editor, filesystem, or A2A runtime.
|
|
93
96
|
|
|
94
97
|
## CLI/RPC
|
|
95
98
|
- [CLI/RPC](cli-rpc.md): Run print/json modes and LF-delimited RPC over the public AgentSession runtime, including mid-run `steer`, branch-handle results, fixed `forkSession`, and `checkout`. `prism init` scaffolds a tiny TypeScript project with one selected provider and an offline mock test.
|
|
96
|
-
- [Workflows](workflows.md): optional `@arnilo/prism-workflows` typed bounded DAG orchestration — explicit recursive definition revisions, exact-owner cancellation/active identity, finite hard limits, durable human suspend/resume, schedules/background execution, nested workflows, replay, coordination, events, and optional RPC/Web bindings. Compose coding plans/checkpoints via workspace Markdown + `state.coding` without a second runtime. Interactive TUI (C-012) deferred.
|
|
99
|
+
- [Workflows](workflows.md): optional `@arnilo/prism-workflows` typed bounded DAG orchestration — explicit recursive definition revisions, exact-owner cancellation/active identity, finite hard limits, durable human suspend/resume, schedules/background execution, revocable proactive schedule capability tokens, nested workflows, replay, coordination, events, and optional RPC/Web bindings. Compose coding plans/checkpoints via workspace Markdown + `state.coding` without a second runtime. Interactive TUI (C-012) deferred.
|
|
97
100
|
- [Workflow orchestration primitives](workflow-orchestration-primitives.md): architecture inventory — workflow adapters consume core `CheckpointStore`, `LeaseStore`, and bounded `EventMultiplexer`; run control and optional RPC commands stay package-local.
|
|
98
101
|
- [Workflow/TUI scope](workflow-tui-primitives.md): records why 0.0.5 ships workflow APIs/RPC control but no interactive terminal UI.
|
|
99
102
|
|
|
100
103
|
## Security and credentials
|
|
101
104
|
- [Host security guide](host-security.md): fail-closed checklist for supply-chain/attestation/canary isolation, bounded credentials, AG-UI/ACP/A2A/web remote boundaries, untrusted external content, settings, redaction, trust roots, workflow ownership, coding I/O, permissions, persistence, extensions, and tool validation.
|
|
102
105
|
- [Security/auth/trust](settings-auth-trust-security.md): settings providers, credential helpers, trust/permission policies, redaction controls, host-owned settings/credentials wiring outside `AgentConfig`, and security-boundary hardening summary.
|
|
103
|
-
- [Credentials and redaction](credentials-and-redaction.md): compose explicit credential resolver order, use caller-supplied env objects/OAuth refresh helpers, resolve credentials only at the provider edge, redact known secret values, and follow the provider-authorized subscription OAuth matrix.
|
|
104
|
-
- [Credential storage](credential-storage.md): optional `@arnilo/prism-credentials-node` adapter with strict bounded AES-GCM envelopes, async finite scrypt, restrictive Unix files, abort-aware bounded system-keychain calls,
|
|
106
|
+
- [Credentials and redaction](credentials-and-redaction.md): compose explicit credential resolver order, use caller-supplied env objects/OAuth refresh + revoke helpers, resolve credentials only at the provider edge, redact known secret values, and follow the provider-authorized subscription OAuth matrix.
|
|
107
|
+
- [Credential storage](credential-storage.md): optional `@arnilo/prism-credentials-node` adapter with strict bounded AES-GCM envelopes, async finite scrypt, restrictive Unix files, abort-aware bounded system-keychain calls, optional host-KMS wrap (`encryptWithHostKms`), and 0.0.14 Microsoft 365 / Google Workspace OAuth providers (PKCE/device-code, least-privilege scope bundles, per-identity work-token bridge).
|
|
105
108
|
|
|
106
109
|
## Testing and examples
|
|
107
110
|
- Provider test doubles: `createMockProvider()` and provider event helpers are documented on the canonical Provider layer page above.
|
|
@@ -111,10 +114,11 @@ Prism is a TypeScript/Node.js agent harness. Host apps and extension packages ow
|
|
|
111
114
|
- [Compaction conformance](compaction-conformance.md): assert any `CompactionStrategy` returns a non-empty redacted summary and observes abort from `@arnilo/prism/testing/compaction-conformance`.
|
|
112
115
|
- [Tool conformance](tool-conformance.md): assert the tool-dispatch blocked-reason matrix (unknown/denied/invalid/permission/validator) and success path from `@arnilo/prism/testing/tool-conformance`.
|
|
113
116
|
- [Extension conformance](extension-conformance.md): assert an `Extension` setup runs, contributions stay inert, and setup errors are redacted or rethrown from `@arnilo/prism/testing/extension-conformance`.
|
|
114
|
-
- `examples/`: compile-checked typed examples and runnable mock demos (SDK basics, provider registration, auth, tools, [`examples/ag-ui-server.ts`](../examples/ag-ui-server.ts), [`examples/enterprise-identity.ts`](../examples/enterprise-identity.ts), [`examples/enterprise-policy-audit.ts`](../examples/enterprise-policy-audit.ts), [`examples/enterprise-work-connectors.ts`](../examples/enterprise-work-connectors.ts), [`examples/server-deployment-seams.ts`](../examples/server-deployment-seams.ts), cache-aware prompt assembly, NeuralWatt agent run ([`examples/neuralwatt-agent-run.ts`](../examples/neuralwatt-agent-run.ts)), [`examples/coding-compaction.ts`](../examples/coding-compaction.ts), stores/branching, structured-output/artifact-loop, CLI, RPC, workflow orchestration).
|
|
117
|
+
- `examples/`: compile-checked typed examples and runnable mock demos (SDK basics, provider registration, auth, tools, [`examples/ag-ui-server.ts`](../examples/ag-ui-server.ts), [`examples/enterprise-identity.ts`](../examples/enterprise-identity.ts), [`examples/enterprise-policy-audit.ts`](../examples/enterprise-policy-audit.ts), [`examples/enterprise-work-connectors.ts`](../examples/enterprise-work-connectors.ts), [`examples/conversation-durable-replay.ts`](../examples/conversation-durable-replay.ts), [`examples/artifact-review-delivery.ts`](../examples/artifact-review-delivery.ts), [`examples/server-deployment-seams.ts`](../examples/server-deployment-seams.ts), cache-aware prompt assembly, NeuralWatt agent run ([`examples/neuralwatt-agent-run.ts`](../examples/neuralwatt-agent-run.ts)), [`examples/coding-compaction.ts`](../examples/coding-compaction.ts), stores/branching, structured-output/artifact-loop, CLI, RPC, workflow orchestration).
|
|
115
118
|
|
|
116
119
|
## Release and install
|
|
117
|
-
- [Release and install](release-and-install.md): current **
|
|
120
|
+
- [Release and install](release-and-install.md): current **43**-package graph (Phase 9 conversations/artifacts/co-work events, scoped OAuth connectors, browser checkpoints, device contracts, and `@arnilo/prism-provider-alibaba`/`@arnilo/prism-provider-ollama` ship at 0.0.14), install/tarball rules, pinned CodeQL/dependency/SBOM/license/secret/attestation gates, deterministic resumable publication, offline tests, protected live canaries, and sandbox-browser Docker/Playwright gates.
|
|
121
|
+
- [Review coverage (2026-07-25 Phase 9)](review-coverage-2026-07-25-phase-9.md): Plan 077 evidence freeze — conversation service, memory consent/lifecycle, artifact co-work review, AG-UI co-work events, scoped M365/GWS OAuth, browser checkpoint composition, and deny-by-default device contracts for 0.0.14 (41 → 43 manifests; only the two provider packages are new).
|
|
118
122
|
- [Review coverage (2026-07-23 Phase 8)](review-coverage-2026-07-23-phase-8.md): Plan 076 evidence freeze — enterprise identity/policy/router packages, Azure/Bedrock/Vertex adapters, server deployment seams, persistence lifecycle hooks, and M365/GWS work-connector bounds for 0.0.13.
|
|
119
123
|
- [Review coverage (2026-07-22 Phase 7)](review-coverage-2026-07-22-phase-7.md): Plan 075 evidence freeze — AG-UI/ACP package boundary, streamed durable resume, bounded replay/projection, coding compaction preset, and provider-authorized OAuth policy for 0.0.12.
|
|
120
124
|
- [Review coverage (2026-07-22 Phase 6)](review-coverage-2026-07-22-phase-6.md): Plan 074 evidence freeze — SessionIndex/search, contextBudget, native Anthropic/Google packages, goal→verify, steer, ask_user_decision (multi/free-text/suspend), finite limits, threats, and 0.0.11 release gates.
|
package/docs/migration.md
CHANGED
|
@@ -7,6 +7,27 @@ Prism 0.0.6 preserves documented 0.0.3 agent construction except for two intenti
|
|
|
7
7
|
1. **`session.run()` / `session.prompt()` return `AgentRunResult`** and `session.stream()` starts one owned run after subscribing. Callers that ignored the previous `Promise<void>` keep working; failed/aborted runs reject with `AgentRunError` (`.result` attached).
|
|
8
8
|
2. **`AgentConfig.extensions` / `settings` / `credentials` are removed.** Wire extensions through `createExtensionKernel()`, read settings in the host, and pass credential resolvers to the provider edge.
|
|
9
9
|
|
|
10
|
+
## 0.0.13 → 0.0.14 personal/work-agent conversations, co-work review, and channel/device gates (additive, pre-release)
|
|
11
|
+
|
|
12
|
+
Release **0.0.14** is strictly additive: every surface extends a shipped package and reuses the AG-UI adapter shipped in 0.0.12. The only new packages are two optional provider adapters (41 → 43 manifests): `@arnilo/prism-provider-alibaba` and `@arnilo/prism-provider-ollama`, both enrolled via the `@arnilo/prism-providers` family. No permission broadening — channel/device/co-work features cannot widen consent, memory, network, file, browser, connector, or tool permissions (roadmap gate 8). See [Phase 9 evidence](review-coverage-2026-07-25-phase-9.md).
|
|
13
|
+
|
|
14
|
+
| Surface | Before (0.0.13) | After (0.0.14) |
|
|
15
|
+
| --- | --- | --- |
|
|
16
|
+
| Conversations | n/a | `@arnilo/prism-server` `createConversationService` / `createConversationHandler`: durable user-scoped threads, reconnectable redacted replay, branch/archive caps |
|
|
17
|
+
| Memory consent/lifecycle | Scope only | `consent { source, scope, visible }` on records; `recall()` injection filter; `setConsent` / `correct` / `forget` / `applyRetention` |
|
|
18
|
+
| Artifacts / review | n/a | `createArtifactService` / `createArtifactHandler` over the existing checkpoint store: revisions, approve/reject, `lastValidated`, expiring authorized delivery links |
|
|
19
|
+
| AG-UI co-work events | Run events only | `mapCoWork()` (+ ACP parity) for artifact progress/approval/download-link, connector drafts, redacted browser snapshots |
|
|
20
|
+
| OAuth connectors | Codex only | `createMicrosoft365OAuthProvider` / `createGoogleWorkspaceOAuthProvider` (PKCE/device-code), least-privilege scope bundles, `revokeOAuthCredential`, per-identity `createOAuthWorkTokenProvider` |
|
|
21
|
+
| Browser composition | Run policy only | `createBrowserCheckpointLedger`: verified-state checkpoints + reload/verify-before-side-effect |
|
|
22
|
+
| Device adapters | n/a | Core `DeviceAdapter` contract + deny-by-default `resolveDevicePolicy` / `assertDeviceAdmit` + conformance (no vendor package) |
|
|
23
|
+
| Providers | 9 HTTP adapters in `@arnilo/prism-providers` | Optional `@arnilo/prism-provider-alibaba` (Model Studio / DashScope + Coding Plan, dynamic `listAlibabaModels`, explicit + implicit cache) and `@arnilo/prism-provider-ollama` (cloud/local, dynamic `listOllamaModels`, implicit-only cache); both join the `@arnilo/prism-providers` family (11 adapters) |
|
|
24
|
+
|
|
25
|
+
**Identity requirement:** every new conversation/artifact/memory/connector/browser/device surface starts from a host-verified `AgentIdentity` (0.0.13 `IdentityVerifier`); ownership is rechecked on resume and at schedule fire time. Caller-asserted identity fails closed.
|
|
26
|
+
|
|
27
|
+
**Deferred to 0.0.15 / 0.1.x (demand-gated):** Slack/Teams chat-channel packages, realtime-voice and desktop-control vendor packages (contract + conformance only in 0.0.14), Studio/control plane, local Office runtime, a second memory/event runtime, and memory production conformance canaries. PostgreSQL/pgvector memory and M365/GWS OAuth / Playwright / keychain live canaries remain explicit operator gates.
|
|
28
|
+
|
|
29
|
+
Benchmark placeholder: `node scripts/benchmark-0.0.14.mjs` (release Task 12). Caps documented in [Performance limits](performance.md).
|
|
30
|
+
|
|
10
31
|
## 0.0.12 → 0.0.13 enterprise identity, policy, routing, and work connectors (additive, pre-release)
|
|
11
32
|
|
|
12
33
|
Release **0.0.13** adds host-verified `Principal` / `AgentIdentity` on runs, tools, server/MCP/A2A/workflow seams. Hosts must supply an `IdentityVerifier` (`verify()` → `AgentIdentity` with `verified: true`); caller-asserted identity without host verification fails closed. See [Agent identity](agent-identity.md).
|
package/docs/performance.md
CHANGED
|
@@ -94,6 +94,8 @@ Durable coding plan/checkpoint defaults/hard caps: plan Markdown 256 KiB/1 MiB;
|
|
|
94
94
|
|
|
95
95
|
Browser automation defaults/hard caps from `@arnilo/prism-browser`: pages 4/16; actions 100/256; queued actions 16/64; snapshot refs 2,000/10,000; depth 30/100; snapshot bytes 256 KiB/2 MiB; navigation 30 s/120 s; action 10 s/60 s; wait 30 s/120 s; run wall 20 min/30 min; popups 4/16; dialogs 16/64; listeners 64/256; action input 64 KiB/256 KiB; close grace 5 s/30 s; network requests 1,000/10,000 with 10/32 redirects per request and 8/32 WebSockets; screenshots 16/64 with 16/64 megapixels and 10 MiB/32 MiB encoded; uploads 8/32 files, 16 MiB/64 MiB each, 64 MiB/256 MiB aggregate; downloads 8/32 files, 32 MiB/256 MiB each, 64 MiB/512 MiB aggregate. Caps charge before context/page/action/queue/snapshot/network/artifact retention. Host supplies Playwright and egress proxy attestation; package import launches nothing.
|
|
96
96
|
|
|
97
|
+
0.0.14 co-work defaults/hard caps (frozen in [Phase 9 evidence](review-coverage-2026-07-25-phase-9.md)): conversation thread list pages 50/200, active branches per thread 16/64, replay/export page 100/500 events; artifact revisions per artifact 32/128, artifacts per thread 64/256, metadata record 8/64 KiB, preview 16/64 KiB, citations 32/128 (2/8 KiB each), delivery-link TTL 5 min/24 h, delivery token 4/16 KiB, compare exactly 2 revisions; memory retention batch 500/5000; proactive capability TTL 24 h/31 d, capability token record 16 KiB; browser checkpoint URL 8 KiB/16 KiB, domain-state hash 256 B/1 KiB, host-data ref 2 KiB/8 KiB, 16/64 checkpoints per run; device stream chunk 1 MiB/8 MiB, concurrent device sessions per identity 1/4 (device wall/turns/tool calls consume shared `RunLimits`). All caps charge before persist/emit and fail closed on overflow. Benchmark placeholder: `node scripts/benchmark-0.0.14.mjs` (release Task 12) reports conversation replay, memory injection/consent, artifact revision/delivery, AG-UI co-work mapping, and connector refresh overhead against these budgets.
|
|
98
|
+
|
|
97
99
|
Current surfaces:
|
|
98
100
|
|
|
99
101
|
- `SubscribeOptions` for bounded live `AgentEvent` subscriber queues.
|
package/docs/policy-and-audit.md
CHANGED
|
@@ -123,5 +123,6 @@ Policy is optional. Hosts wire `record*` helpers or `evaluateAndAppend` at permi
|
|
|
123
123
|
- [Agent identity](agent-identity.md)
|
|
124
124
|
- [Guardrails](guardrails.md)
|
|
125
125
|
- [Runs and usage ledger](runs-and-usage.md)
|
|
126
|
+
- [Workflows](workflows.md): proactive schedule capability enable/revoke events bridge here via `onCapability`.
|
|
126
127
|
- [Host security](host-security.md)
|
|
127
128
|
- Package README: [`@arnilo/prism-policy`](../packages/policy/README.md)
|
package/docs/provider-caching.md
CHANGED
|
@@ -152,6 +152,8 @@ Provider request policies can set `ProviderRequestOptions.cache` or the legacy `
|
|
|
152
152
|
| `@arnilo/prism-provider-kimi` | implicit by default, optional `cache_control` | Default catalog models send no `cache_control`; hosts may opt in on Anthropic `/messages` models with `ModelConfig.cache.kind: "cache_control"`. | Keep selected Anthropic anchors and prior history stable. | Best-effort and model/route-dependent. |
|
|
153
153
|
| `@arnilo/prism-provider-neuralwatt` | `implicit` | No `cache_control`, `cacheKey`, `prompt_cache`, or `cacheRetention` payload; NeuralWatt vLLM prefix caching is automatic. | Full prior history must be resent unchanged with only the new turn appended; `inputLayout: "cache_aware"` keeps stable prefixes first. | Best-effort only; does not promise cache hits; `cacheRetention: "none"` disables Prism hints only, not the implicit backend prefix cache. |
|
|
154
154
|
| `@arnilo/prism-provider-ai-sdk` | host-owned | No Prism cache payload; host `LanguageModelV4` owns upstream caching. | Host model/provider decides cache keys, breakpoints, and sticky routing. | Adapter maps `inputTokens.cacheRead`/`cacheWrite` from `finish.usage` only; does not invent cache fields. |
|
|
155
|
+
| `@arnilo/prism-provider-alibaba` | implicit by default, optional `cache_control` | DashScope implicit prefix caching is automatic; opt-in `cache_control: {"type":"ephemeral"}` markers only on caller-selected `cache.breakpoints`, capped at 4. | Keep selected anchors and prior history stable; each cached prefix needs ≥1024 tokens and lives ~5 minutes upstream. | Best-effort and model-dependent; `cached_tokens`→read, `cache_creation_input_tokens`→write. |
|
|
156
|
+
| `@arnilo/prism-provider-ollama` | `implicit` | No `cache_control`, `cacheKey`, `prompt_cache`, or `cacheRetention` payload; Ollama KV/prefix caching is automatic with no request knob. | Resend unchanged prior history for implicit KV reuse. | Best-effort only; Ollama reports no cached-token count, so `Usage.cacheReadTokens` stays `undefined`. |
|
|
155
157
|
|
|
156
158
|
Detailed first-party provider notes:
|
|
157
159
|
|
|
@@ -163,6 +165,8 @@ Detailed first-party provider notes:
|
|
|
163
165
|
- NeuralWatt (`@arnilo/prism-provider-neuralwatt`): `kind: "implicit"`. NeuralWatt prefix caching is automatic; sends no explicit cache payload regardless of cache options. `cacheRetention: "none"` disables Prism cache-control hints only (not the implicit backend prefix cache). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`; NeuralWatt does not report a cache-write token, so `Usage.cacheWriteTokens` is never fabricated (stays `undefined`). NeuralWatt's `/v1/models` catalog advertises exact `cached_input_per_million` rates for cache reads and `cached_output_per_million: null`; static curated aliases do not guess those prices.
|
|
164
166
|
- Kimi (`@arnilo/prism-provider-kimi`): default catalog models use implicit caching (no `cache_control`); hosts opt in via `ModelConfig.cache.kind: "cache_control"` on the Anthropic `/messages` route, then `cache_control` markers apply only to selected breakpoints (`"long"` → `ttl: "1h"`); the Moonshot OpenAI route sends none. `cache_read_input_tokens`/`cache_creation_input_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`.
|
|
165
167
|
- AI SDK adapter (`@arnilo/prism-provider-ai-sdk`): **host-owned**. Sends no Prism cache payload; the supplied `LanguageModelV4` and its upstream provider own request caching. Maps AI SDK v4 `finish.usage.inputTokens.cacheRead`/`cacheWrite` to `Usage.cacheReadTokens`/`cacheWriteTokens`. No `list*Models()` export.
|
|
168
|
+
- Alibaba Cloud (`@arnilo/prism-provider-alibaba`): implicit by default, optional `cache_control`. DashScope implicit prefix caching is automatic (no marker); explicit opt-in `cache_control: {"type":"ephemeral"}` markers apply only to selected breakpoints when `ModelConfig.cache.kind: "cache_control"` and the caller supplies breakpoints, capped at 4 (each prefix ≥1024 tokens, ~5 minute TTL). `prompt_tokens_details.cached_tokens`/`cache_creation_input_tokens` map to `Usage.cacheReadTokens`/`cacheWriteTokens`. Caller-gated `listAlibabaModels` against OpenAI-compatible `GET {base}/models`.
|
|
169
|
+
- Ollama (`@arnilo/prism-provider-ollama`): `kind: "implicit"`. Ollama reuses its KV/prompt cache automatically; there is no request knob and no wire marker, so Prism never emits `cache_control`. Ollama reports no cached-token count, so `Usage.cacheReadTokens` is intentionally left `undefined` (not `0`). Caller-gated `listOllamaModels` against OpenAI-compatible `GET {base}/models`.
|
|
166
170
|
|
|
167
171
|
### NeuralWatt cache-aware limiter
|
|
168
172
|
|
|
@@ -93,6 +93,8 @@ Every first-party provider package hardens prompt-cache behavior so it cannot em
|
|
|
93
93
|
- **Z.AI** (`kind: implicit`): GLM context caching is automatic; no explicit cache payload sent regardless of cache options. `prompt_tokens_details.cached_tokens`/`cache_write_tokens` map to cache usage.
|
|
94
94
|
- **NeuralWatt** (`kind: implicit`): NeuralWatt prefix caching is automatic; sends no explicit cache payload regardless of cache options. `cacheRetention: "none"` disables Prism cache-control hints only (not the implicit backend prefix cache). `prompt_tokens_details.cached_tokens` maps to `Usage.cacheReadTokens`; NeuralWatt does not report a cache-write token so `Usage.cacheWriteTokens` is never fabricated.
|
|
95
95
|
- **Kimi**: default catalog models use implicit caching (no `cache_control`); hosts opt in via `ModelConfig.cache.kind: cache_control` on the Anthropic `/messages` route, then markers apply only to selected breakpoints (`long` → `ttl: 1h`); the Moonshot OpenAI route sends none. `cache_read_input_tokens`/`cache_creation_input_tokens` map to cache usage.
|
|
96
|
+
- **Alibaba Cloud** (implicit by default, optional `cache_control`): DashScope implicit prefix caching is automatic; hosts opt in via `ModelConfig.cache.kind: cache_control`, then `cache_control: {"type":"ephemeral"}` markers apply only to selected breakpoints, capped at 4. `prompt_tokens_details.cached_tokens`/`cache_creation_input_tokens` map to cache usage. Caller-gated `listAlibabaModels`.
|
|
97
|
+
- **Ollama** (`kind: implicit`): Ollama KV/prefix caching is automatic with no request knob; sends no explicit cache payload. Ollama reports no cached-token count, so `Usage.cacheReadTokens` stays `undefined`. Caller-gated `listOllamaModels`.
|
|
96
98
|
|
|
97
99
|
See [Provider caching](provider-caching.md) for the `PromptCacheHints` surface and shared helpers, and [Provider conformance](provider-conformance.md) for the `assertUsageAccounting` and `assertProviderOwnedHeadersWin` checks every first-party package exercises.
|
|
98
100
|
|
|
@@ -160,7 +162,8 @@ provider packages: an `Extension` whose `setup(api)` calls
|
|
|
160
162
|
`api.registerProvider(provider)` for each provider it owns. First-party
|
|
161
163
|
provider packages (`@arnilo/prism-provider-openai`, `@arnilo/prism-provider-openrouter`,
|
|
162
164
|
`@arnilo/prism-provider-kimi`, `@arnilo/prism-provider-zai`,
|
|
163
|
-
`@arnilo/prism-provider-opencode-go
|
|
165
|
+
`@arnilo/prism-provider-opencode-go`, `@arnilo/prism-provider-alibaba`,
|
|
166
|
+
`@arnilo/prism-provider-ollama`) are **opt-in and individually installable**;
|
|
164
167
|
`@arnilo/prism` core runs without any first-party provider package (mock-only).
|
|
165
168
|
|
|
166
169
|
A host mixes first-party packages and third-party providers in one resolver.
|
|
@@ -0,0 +1,179 @@
|
|
|
1
|
+
# Alibaba Cloud provider package
|
|
2
|
+
|
|
3
|
+
## What it does
|
|
4
|
+
|
|
5
|
+
`@arnilo/prism-provider-alibaba` is a side-effect-free adapter for Alibaba Cloud
|
|
6
|
+
Model Studio / DashScope (including the Coding Plan) over the **OpenAI-compatible**
|
|
7
|
+
`POST {base}/chat/completions` endpoint.
|
|
8
|
+
|
|
9
|
+
- **Dynamic model discovery** — `listAlibabaModels()` calls the OpenAI-compatible
|
|
10
|
+
`GET {base}/models`. No model catalog is hard-coded in the package: available
|
|
11
|
+
models vary by region, workspace, and billing plan, so discovery is the source of
|
|
12
|
+
truth. Package setup never fetches.
|
|
13
|
+
- **Context cache** — DashScope implicit prefix caching is automatic. Explicit
|
|
14
|
+
caching is opt-in via Anthropic-style `cache_control: {"type":"ephemeral"}`
|
|
15
|
+
markers (at most 4 per request). Cache hits are accounted from
|
|
16
|
+
`usage.prompt_tokens_details.cached_tokens` (read) and
|
|
17
|
+
`cache_creation_input_tokens` (write).
|
|
18
|
+
- **Qwen thinking** — `enable_thinking` passthrough toggles reasoning on Qwen models.
|
|
19
|
+
|
|
20
|
+
The API key is region/plan-scoped: it must match the base URL's billing plan
|
|
21
|
+
(pay-as-you-go regional, workspace-dedicated, or Coding Plan).
|
|
22
|
+
|
|
23
|
+
## When to use it
|
|
24
|
+
|
|
25
|
+
Use it when a host app wants Alibaba Cloud Qwen models (Model Studio / DashScope or
|
|
26
|
+
the Coding Plan) through Prism's `AgentSession` runtime with OpenAI-compatible
|
|
27
|
+
serialization, dynamic model discovery, and explicit/implicit cache accounting.
|
|
28
|
+
|
|
29
|
+
Do not use it for automatic credential discovery, setup-time catalog fetches, or
|
|
30
|
+
real-network tests (live tests stay opt-in).
|
|
31
|
+
|
|
32
|
+
## Inputs / request
|
|
33
|
+
|
|
34
|
+
```ts
|
|
35
|
+
import {
|
|
36
|
+
createAlibabaProviderPackage,
|
|
37
|
+
createAlibabaProvider,
|
|
38
|
+
listAlibabaModels,
|
|
39
|
+
defineAlibabaModel,
|
|
40
|
+
alibabaBaseUrl,
|
|
41
|
+
} from "@arnilo/prism-provider-alibaba";
|
|
42
|
+
|
|
43
|
+
createAlibabaProviderPackage(options: AlibabaProviderPackageOptions): ProviderPackage
|
|
44
|
+
createAlibabaProvider(options?: AlibabaProviderOptions): AIProvider
|
|
45
|
+
listAlibabaModels(options?: ListAlibabaModelsOptions): Promise<ModelConfig[]>
|
|
46
|
+
defineAlibabaModel(config: AlibabaModelConfig): ModelConfig
|
|
47
|
+
alibabaBaseUrl(options?: { baseUrl?: string; preset?: AlibabaBasePreset }): string
|
|
48
|
+
```
|
|
49
|
+
|
|
50
|
+
| Field | Type | Purpose |
|
|
51
|
+
| --- | --- | --- |
|
|
52
|
+
| `apiKey` | `CredentialValueSource` | DashScope API key (`DASHSCOPE_API_KEY`), region/plan-scoped. |
|
|
53
|
+
| `baseUrl` | `string` | Explicit OpenAI-compatible base URL (wins over `preset`). |
|
|
54
|
+
| `preset` | `AlibabaBasePreset` | `"singapore"` (default) / `"beijing"` / `"us"` / `"coding-plan"`. |
|
|
55
|
+
| `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
|
|
56
|
+
| `id` | `string` | Provider id (default `alibaba`). |
|
|
57
|
+
| `models` | `readonly ModelConfig[]` | Host-supplied models (from `listAlibabaModels`) to register. |
|
|
58
|
+
|
|
59
|
+
Base URLs resolved by preset:
|
|
60
|
+
|
|
61
|
+
| Preset | Base URL |
|
|
62
|
+
| --- | --- |
|
|
63
|
+
| `singapore` | `https://dashscope-intl.aliyuncs.com/compatible-mode/v1` |
|
|
64
|
+
| `beijing` | `https://dashscope.aliyuncs.com/compatible-mode/v1` |
|
|
65
|
+
| `us` | `https://dashscope-us.aliyuncs.com/compatible-mode/v1` |
|
|
66
|
+
| `coding-plan` | `https://coding-intl.dashscope.aliyuncs.com/v1` |
|
|
67
|
+
|
|
68
|
+
Workspace-dedicated endpoints
|
|
69
|
+
(`https://{workspaceId}.{region}.maas.aliyuncs.com/compatible-mode/v1`) are supplied
|
|
70
|
+
verbatim via `baseUrl`.
|
|
71
|
+
|
|
72
|
+
## Outputs / response / events
|
|
73
|
+
|
|
74
|
+
| Surface | Behavior |
|
|
75
|
+
| --- | --- |
|
|
76
|
+
| Stream | Prism text deltas, `delta.reasoning_content` → thinking deltas, tool-call delta/final, `usage`, `done`, redacted `error`. |
|
|
77
|
+
| Usage | `prompt_tokens`/`completion_tokens`/`total_tokens`; `prompt_tokens_details.cached_tokens` → `cacheReadTokens`, `cache_creation_input_tokens` → `cacheWriteTokens`. |
|
|
78
|
+
| Discovery | `listAlibabaModels()` maps `GET {base}/models` entries → `ModelConfig` (reasoning/vision inferred from id). |
|
|
79
|
+
| Auth methods | `api_key` for `alibaba`. |
|
|
80
|
+
|
|
81
|
+
The stream parser emits `done` only on completion evidence (`[DONE]` plus a terminal
|
|
82
|
+
`finish_reason` with no dangling tool calls). Truncated streams terminate with an
|
|
83
|
+
`error` event instead. Unsupported block placements or unclaimed images fail before
|
|
84
|
+
fetch.
|
|
85
|
+
|
|
86
|
+
## Request/response example
|
|
87
|
+
|
|
88
|
+
```bash
|
|
89
|
+
curl 'https://dashscope-intl.aliyuncs.com/compatible-mode/v1/chat/completions' \
|
|
90
|
+
-H "Authorization: Bearer $DASHSCOPE_API_KEY" \
|
|
91
|
+
-H 'Content-Type: application/json' \
|
|
92
|
+
-d '{
|
|
93
|
+
"model": "qwen-plus",
|
|
94
|
+
"messages": [{ "role": "user", "content": "Hello" }],
|
|
95
|
+
"stream": true,
|
|
96
|
+
"stream_options": { "include_usage": true }
|
|
97
|
+
}'
|
|
98
|
+
```
|
|
99
|
+
|
|
100
|
+
Usage in the final streamed chunk:
|
|
101
|
+
|
|
102
|
+
```json
|
|
103
|
+
{
|
|
104
|
+
"usage": {
|
|
105
|
+
"prompt_tokens": 100,
|
|
106
|
+
"completion_tokens": 5,
|
|
107
|
+
"total_tokens": 105,
|
|
108
|
+
"prompt_tokens_details": { "cached_tokens": 80, "cache_creation_input_tokens": 10 }
|
|
109
|
+
}
|
|
110
|
+
}
|
|
111
|
+
```
|
|
112
|
+
|
|
113
|
+
## Implementation example
|
|
114
|
+
|
|
115
|
+
```ts
|
|
116
|
+
import { createExtensionKernel } from "@arnilo/prism";
|
|
117
|
+
import {
|
|
118
|
+
createAlibabaProviderPackage,
|
|
119
|
+
listAlibabaModels,
|
|
120
|
+
} from "@arnilo/prism-provider-alibaba";
|
|
121
|
+
|
|
122
|
+
const kernel = createExtensionKernel();
|
|
123
|
+
|
|
124
|
+
// Caller-gated discovery — never runs during setup.
|
|
125
|
+
const models = await listAlibabaModels({ apiKey: process.env.DASHSCOPE_API_KEY });
|
|
126
|
+
|
|
127
|
+
await kernel.load([
|
|
128
|
+
createAlibabaProviderPackage({
|
|
129
|
+
apiKey: process.env.DASHSCOPE_API_KEY,
|
|
130
|
+
preset: "singapore", // or "coding-plan" with a Coding Plan key
|
|
131
|
+
models,
|
|
132
|
+
}),
|
|
133
|
+
]);
|
|
134
|
+
```
|
|
135
|
+
|
|
136
|
+
## Extension and configuration notes
|
|
137
|
+
|
|
138
|
+
- Hosts choose base URL/preset, provider id, model list, credential source, and
|
|
139
|
+
`fetch` impl. Nothing is hard-coded; register discovered models via `models:`.
|
|
140
|
+
- Qwen thinking: `compat.enable_thinking` (request wins over model default) maps to
|
|
141
|
+
the top-level `enable_thinking` wire field; omitted unless explicitly boolean.
|
|
142
|
+
- Provider-owned compat keys (`route`, `enable_thinking`, `alibaba`) are stripped
|
|
143
|
+
before the opaque `compat` spread so they never leak into wire bodies.
|
|
144
|
+
|
|
145
|
+
### Cache behavior
|
|
146
|
+
|
|
147
|
+
- **Implicit** prefix caching is automatic upstream and sends no markers.
|
|
148
|
+
- **Explicit** caching is opt-in: when `ModelConfig.cache.kind === "cache_control"`
|
|
149
|
+
(or `cache.mode === "on"`) and the caller supplies
|
|
150
|
+
`ProviderRequestOptions.cache.breakpoints`, `cache_control: {"type":"ephemeral"}`
|
|
151
|
+
markers land on the last content block of each selected message, capped at
|
|
152
|
+
`ALIBABA_MAX_CACHE_BREAKPOINTS` (4). Each cached prefix needs ≥1024 tokens and
|
|
153
|
+
lives ~5 minutes upstream.
|
|
154
|
+
- Usage accounting: `cached_tokens` → `Usage.cacheReadTokens`,
|
|
155
|
+
`cache_creation_input_tokens` → `Usage.cacheWriteTokens`.
|
|
156
|
+
|
|
157
|
+
## Security and performance notes
|
|
158
|
+
|
|
159
|
+
- SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport`
|
|
160
|
+
helpers (`readSseData`, `readBoundedResponseText`).
|
|
161
|
+
- No network calls during import, setup, build, or default tests.
|
|
162
|
+
- No automatic environment, file, keychain, or shell credential lookup.
|
|
163
|
+
- The API key is resolved per request via `resolveCredentialValue` and sent only as
|
|
164
|
+
`Authorization: Bearer`; keys are redacted from all thrown errors (including
|
|
165
|
+
discovery failures). No local filesystem paths enter request payloads.
|
|
166
|
+
- Caller-supplied `ProviderRequest.options.headers` can add non-owned headers, but
|
|
167
|
+
provider-owned headers (`content-type`, `authorization`) are applied last and
|
|
168
|
+
cannot be overridden.
|
|
169
|
+
- Model discovery is caller-gated and never invoked in the provider hot path.
|
|
170
|
+
- Live tests stay opt-in; default tests are network-free.
|
|
171
|
+
|
|
172
|
+
## Related APIs
|
|
173
|
+
|
|
174
|
+
- [Provider packages](../provider-packages.md): `defineProviderPackage`,
|
|
175
|
+
caller-gated discovery, OpenAI-compatible routes.
|
|
176
|
+
- [Provider caching](../provider-caching.md): explicit/implicit matrix.
|
|
177
|
+
- [Credentials and redaction](../credentials-and-redaction.md):
|
|
178
|
+
`resolveCredentialValue`, `redactSecrets`.
|
|
179
|
+
- [Provider conformance](../provider-conformance.md): network-free adapter tests.
|
|
@@ -0,0 +1,166 @@
|
|
|
1
|
+
# Ollama Cloud provider package
|
|
2
|
+
|
|
3
|
+
## What it does
|
|
4
|
+
|
|
5
|
+
`@arnilo/prism-provider-ollama` is a side-effect-free adapter for Ollama — both
|
|
6
|
+
**Ollama Cloud** (`https://ollama.com`) and a **local** `ollama serve`
|
|
7
|
+
(`http://localhost:11434`) — over the OpenAI-compatible
|
|
8
|
+
`POST {base}/chat/completions` endpoint.
|
|
9
|
+
|
|
10
|
+
- **Dynamic model discovery** — `listOllamaModels()` calls the OpenAI-compatible
|
|
11
|
+
`GET {base}/models`. No model catalog is hard-coded: available models vary by cloud
|
|
12
|
+
account or local pull, so discovery is the source of truth. Package setup never
|
|
13
|
+
fetches. (The native `GET {base}/api/tags` endpoint is an alternate catalog source;
|
|
14
|
+
Prism uses the OpenAI-compatible route for a uniform shape.)
|
|
15
|
+
- **Implicit cache only** — Ollama reuses its KV/prompt cache automatically. There is
|
|
16
|
+
no request knob and no cached-token count in usage, so `Usage.cacheReadTokens` is
|
|
17
|
+
intentionally left undefined (documented ceiling below).
|
|
18
|
+
- **Reasoning** — `reasoning_effort` passthrough (e.g. gpt-oss models).
|
|
19
|
+
|
|
20
|
+
Cloud auth is an ollama.com API key sent as `Authorization: Bearer`; local instances
|
|
21
|
+
are typically unauthenticated (omit the key).
|
|
22
|
+
|
|
23
|
+
## When to use it
|
|
24
|
+
|
|
25
|
+
Use it when a host app wants Ollama Cloud or local Ollama models through Prism's
|
|
26
|
+
`AgentSession` runtime with OpenAI-compatible serialization and dynamic model
|
|
27
|
+
discovery.
|
|
28
|
+
|
|
29
|
+
Do not use it for automatic credential discovery, setup-time catalog fetches, explicit
|
|
30
|
+
cache control (Ollama has none), or real-network tests (live tests stay opt-in).
|
|
31
|
+
|
|
32
|
+
## Inputs / request
|
|
33
|
+
|
|
34
|
+
```ts
|
|
35
|
+
import {
|
|
36
|
+
createOllamaProviderPackage,
|
|
37
|
+
createOllamaProvider,
|
|
38
|
+
listOllamaModels,
|
|
39
|
+
defineOllamaModel,
|
|
40
|
+
ollamaBaseUrl,
|
|
41
|
+
} from "@arnilo/prism-provider-ollama";
|
|
42
|
+
|
|
43
|
+
createOllamaProviderPackage(options: OllamaProviderPackageOptions): ProviderPackage
|
|
44
|
+
createOllamaProvider(options?: OllamaProviderOptions): AIProvider
|
|
45
|
+
listOllamaModels(options?: ListOllamaModelsOptions): Promise<ModelConfig[]>
|
|
46
|
+
defineOllamaModel(config: OllamaModelConfig): ModelConfig
|
|
47
|
+
ollamaBaseUrl(options?: { baseUrl?: string; preset?: OllamaBasePreset }): string
|
|
48
|
+
```
|
|
49
|
+
|
|
50
|
+
| Field | Type | Purpose |
|
|
51
|
+
| --- | --- | --- |
|
|
52
|
+
| `apiKey` | `CredentialValueSource` | Ollama Cloud API key; omit for unauthenticated local. |
|
|
53
|
+
| `baseUrl` | `string` | Explicit OpenAI-compatible base URL (wins over `preset`). |
|
|
54
|
+
| `preset` | `OllamaBasePreset` | `"cloud"` (default) / `"local"`. |
|
|
55
|
+
| `fetch` | `typeof fetch` | Optional fetch implementation for tests/hosts. |
|
|
56
|
+
| `id` | `string` | Provider id (default `ollama`). |
|
|
57
|
+
| `models` | `readonly ModelConfig[]` | Host-supplied models (from `listOllamaModels`) to register. |
|
|
58
|
+
|
|
59
|
+
Base URLs resolved by preset (each includes the `/v1` segment):
|
|
60
|
+
|
|
61
|
+
| Preset | Base URL |
|
|
62
|
+
| --- | --- |
|
|
63
|
+
| `cloud` | `https://ollama.com/v1` |
|
|
64
|
+
| `local` | `http://localhost:11434/v1` |
|
|
65
|
+
|
|
66
|
+
## Outputs / response / events
|
|
67
|
+
|
|
68
|
+
| Surface | Behavior |
|
|
69
|
+
| --- | --- |
|
|
70
|
+
| Stream | Prism text deltas, `delta.reasoning_content` → thinking deltas, tool-call delta/final, `usage`, `done`, redacted `error`. |
|
|
71
|
+
| Usage | `prompt_tokens` → `inputTokens`, `completion_tokens` → `outputTokens` (native `prompt_eval_count`/`eval_count` are the equivalent). `cacheReadTokens` stays undefined. |
|
|
72
|
+
| Discovery | `listOllamaModels()` maps `GET {base}/models` entries → `ModelConfig` (reasoning/vision inferred from id). |
|
|
73
|
+
| Auth methods | `api_key` for `ollama`. |
|
|
74
|
+
|
|
75
|
+
The stream parser emits `done` only on completion evidence (`[DONE]` plus a terminal
|
|
76
|
+
`finish_reason` with no dangling tool calls). Truncated streams terminate with an
|
|
77
|
+
`error` event instead. Unsupported block placements or unclaimed images fail before
|
|
78
|
+
fetch.
|
|
79
|
+
|
|
80
|
+
## Request/response example
|
|
81
|
+
|
|
82
|
+
```bash
|
|
83
|
+
curl 'https://ollama.com/v1/chat/completions' \
|
|
84
|
+
-H "Authorization: Bearer $OLLAMA_API_KEY" \
|
|
85
|
+
-H 'Content-Type: application/json' \
|
|
86
|
+
-d '{
|
|
87
|
+
"model": "gpt-oss:20b",
|
|
88
|
+
"messages": [{ "role": "user", "content": "Hello" }],
|
|
89
|
+
"stream": true,
|
|
90
|
+
"stream_options": { "include_usage": true }
|
|
91
|
+
}'
|
|
92
|
+
|
|
93
|
+
curl 'https://ollama.com/v1/models' -H "Authorization: Bearer $OLLAMA_API_KEY"
|
|
94
|
+
```
|
|
95
|
+
|
|
96
|
+
Usage in the final streamed chunk:
|
|
97
|
+
|
|
98
|
+
```json
|
|
99
|
+
{ "usage": { "prompt_tokens": 100, "completion_tokens": 5, "total_tokens": 105 } }
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
## Implementation example
|
|
103
|
+
|
|
104
|
+
```ts
|
|
105
|
+
import { createExtensionKernel } from "@arnilo/prism";
|
|
106
|
+
import {
|
|
107
|
+
createOllamaProviderPackage,
|
|
108
|
+
listOllamaModels,
|
|
109
|
+
} from "@arnilo/prism-provider-ollama";
|
|
110
|
+
|
|
111
|
+
const kernel = createExtensionKernel();
|
|
112
|
+
|
|
113
|
+
// Caller-gated discovery — never runs during setup.
|
|
114
|
+
const models = await listOllamaModels({ apiKey: process.env.OLLAMA_API_KEY });
|
|
115
|
+
|
|
116
|
+
await kernel.load([
|
|
117
|
+
createOllamaProviderPackage({
|
|
118
|
+
apiKey: process.env.OLLAMA_API_KEY, // omit for local
|
|
119
|
+
preset: "cloud", // or "local"
|
|
120
|
+
models,
|
|
121
|
+
}),
|
|
122
|
+
]);
|
|
123
|
+
```
|
|
124
|
+
|
|
125
|
+
## Extension and configuration notes
|
|
126
|
+
|
|
127
|
+
- Hosts choose base URL/preset, provider id, model list, credential source, and
|
|
128
|
+
`fetch` impl. Nothing is hard-coded; register discovered models via `models:`.
|
|
129
|
+
- Reasoning: `compat.reasoning_effort` (request wins over model default) maps to the
|
|
130
|
+
top-level `reasoning_effort` wire field; omitted unless explicitly a string.
|
|
131
|
+
- Provider-owned compat keys (`route`, `reasoning_effort`, `ollama`) are stripped
|
|
132
|
+
before the opaque `compat` spread so they never leak into wire bodies.
|
|
133
|
+
|
|
134
|
+
### Cache behavior
|
|
135
|
+
|
|
136
|
+
- **Implicit only.** Ollama reuses its KV/prompt cache automatically; there is no
|
|
137
|
+
request knob and no wire marker. Prism never emits `cache_control` for Ollama.
|
|
138
|
+
- **Documented ceiling:** Ollama exposes no cached-token count, so
|
|
139
|
+
`Usage.cacheReadTokens` is intentionally left `undefined` (not `0`). If a future
|
|
140
|
+
Ollama release reports cached tokens, map them in `mapOllamaModel`/usage handling.
|
|
141
|
+
|
|
142
|
+
## Security and performance notes
|
|
143
|
+
|
|
144
|
+
- SSE streams and HTTP error bodies use bounded `@arnilo/prism/providers/transport`
|
|
145
|
+
helpers (`readSseData`, `readBoundedResponseText`).
|
|
146
|
+
- No network calls during import, setup, build, or default tests.
|
|
147
|
+
- No automatic environment, file, keychain, or shell credential lookup.
|
|
148
|
+
- The cloud API key is resolved per request via `resolveCredentialValue` and sent only
|
|
149
|
+
as `Authorization: Bearer`; keys are redacted from all thrown errors (including
|
|
150
|
+
discovery failures). Local presets send no auth header when no key is configured.
|
|
151
|
+
No local filesystem paths enter request payloads.
|
|
152
|
+
- Caller-supplied `ProviderRequest.options.headers` can add non-owned headers, but
|
|
153
|
+
provider-owned headers (`content-type`, `authorization`) are applied last and
|
|
154
|
+
cannot be overridden.
|
|
155
|
+
- Model discovery is caller-gated and never invoked in the provider hot path.
|
|
156
|
+
- Live tests stay opt-in; default tests are network-free.
|
|
157
|
+
|
|
158
|
+
## Related APIs
|
|
159
|
+
|
|
160
|
+
- [Provider packages](../provider-packages.md): `defineProviderPackage`,
|
|
161
|
+
caller-gated discovery, OpenAI-compatible routes.
|
|
162
|
+
- [Provider caching](../provider-caching.md): explicit/implicit matrix (Ollama =
|
|
163
|
+
implicit only).
|
|
164
|
+
- [Credentials and redaction](../credentials-and-redaction.md):
|
|
165
|
+
`resolveCredentialValue`, `redactSecrets`.
|
|
166
|
+
- [Provider conformance](../provider-conformance.md): network-free adapter tests.
|