ds4-context-engine 0.1.0 → 0.1.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -16,7 +16,7 @@ bounded active context with provenance
16
16
  Pi provider
17
17
  ```
18
18
 
19
- > **Project status:** M0–M13 are implemented. The Pi adapter and standalone `ds4-context-core` package are version `0.1.0`; the adapter targets Pi `0.84.3`.
19
+ > **Project status:** M0–M13 are implemented. The Pi adapter and standalone `ds4-context-core` package are version `0.1.2`; the adapter targets Pi `0.84.3`.
20
20
 
21
21
  ## Why DS4
22
22
 
@@ -371,6 +371,7 @@ scripts package and release-readiness checks
371
371
  - [Native continuation](docs/NATIVE_CONTINUATION.md)
372
372
  - [Portable core](docs/PORTABLE_CORE.md)
373
373
  - [Storage](docs/STORAGE.md)
374
+ - [Roadmap 0.2.0](docs/ROADMAP_0.2.0.md)
374
375
  - [Release process](docs/RELEASING.md)
375
376
  - [Architecture decisions](docs/ADR/README.md)
376
377
  - [Original development plan](DS4_Context_Engine_Extension_Piano_Sviluppo.md)
@@ -379,7 +380,7 @@ scripts package and release-readiness checks
379
380
 
380
381
  The original M0–M13 roadmap is complete. `ds4-context-core` now contains the compiled Pi-independent implementation, while runtime-specific behavior remains in the Pi adapter.
381
382
 
382
- Possible later work includes additional agent-runtime adapters, semantic retrieval, richer symbol indexing, cross-session project memory, context quality metrics, learned ranking and local KV integration. These are not required by the current MVP.
383
+ The planned [0.2.0 roadmap](docs/ROADMAP_0.2.0.md) covers context-quality metrics, richer symbol indexing, hybrid semantic retrieval, cross-session project memory, optional learned ranking, a runtime adapter kit with one reference adapter, and optional local KV reuse. Sensitive or transport-specific behavior remains opt-in, and the 0.1 lexical planner stays available as the deterministic fallback.
383
384
 
384
385
  ## Contributing
385
386
 
@@ -16,7 +16,9 @@ This is independent of global `project.enabled`: a global setting cannot overrid
16
16
 
17
17
  ## Discovery and exclusions
18
18
 
19
- When Git is available, discovery combines tracked and non-ignored untracked files under `ctx.cwd`. Outside Git, DS4 walks the project tree deterministically. It never follows symbolic links and rejects paths outside the canonical project root.
19
+ When Git is available, discovery combines tracked and non-ignored untracked files under `ctx.cwd`. Outside Git, DS4 walks the project tree deterministically. Discovery stops as soon as the configured file bound is reached and also bounds visited directories. It never follows symbolic links and rejects paths outside the canonical project root.
20
+
21
+ DS4 skips project indexing when `ctx.cwd` is the filesystem root or the user's home directory. This prevents a normal Pi launch from recursively scanning an entire drive or user profile; start Pi inside the intended project directory to enable project knowledge.
20
22
 
21
23
  Default bounds:
22
24
 
@@ -0,0 +1,254 @@
1
+ # DS4 0.2.0 Roadmap
2
+
3
+ Status: **planned**. No release date is committed.
4
+
5
+ Version 0.2.0 focuses on evidence quality, safe project-wide reuse and runtime portability. It extends the released 0.1.0 architecture without changing its canonical-state or failure guarantees.
6
+
7
+ ## Goals
8
+
9
+ The release targets seven improvements:
10
+
11
+ 1. deterministic context-quality metrics;
12
+ 2. richer symbol indexing;
13
+ 3. hybrid lexical and semantic retrieval;
14
+ 4. cross-session project memory;
15
+ 5. optional learned ranking;
16
+ 6. an adapter kit plus one non-Pi reference adapter;
17
+ 7. optional local KV reuse through an explicit runtime capability.
18
+
19
+ The existing planner, lexical retrieval and Pi adapter remain the baseline. New behavior that can affect privacy, ranking or provider transport is opt-in until its release gate is satisfied.
20
+
21
+ ## Invariants
22
+
23
+ Every milestone must preserve these rules:
24
+
25
+ - Pi JSONL remains canonical for Pi conversation state and memory/pin mutations.
26
+ - Live project files remain canonical for project knowledge.
27
+ - SQLite contains only disposable projections. Embeddings, symbol graphs, quality aggregates and learned weights are reproducible from canonical sources plus versioned local evaluation inputs, or are discarded when those inputs are unavailable.
28
+ - No migration rewrites Pi JSONL or deletes raw history.
29
+ - Retrieved content remains quoted, provenance-carrying evidence rather than instructions.
30
+ - Tool call/result groups remain atomic.
31
+ - Strict compaction validation remains enabled; unsafe output falls back to Pi.
32
+ - Remote processing never receives `local-only` content. Semantic and training paths pass through the same privacy policy as prompt construction.
33
+ - Manifests and telemetry contain aggregate metadata, hashes and source identifiers, not prompts, private content or provider response IDs.
34
+ - Static deterministic ranking remains available as the fallback for every new subsystem.
35
+ - A feature-specific failure cannot prevent Pi from continuing with the 0.1 behavior. Privacy enforcement remains fail-closed before remote transport.
36
+
37
+ ## Dependency order
38
+
39
+ ```text
40
+ M14 quality baseline ───────────────┐
41
+ ├─> M18 learned ranking
42
+ M15 symbol index -> M16 semantics ──┤
43
+ M17 cross-session memory ───────────┘
44
+
45
+ M19 adapter kit -> M20 local KV capability
46
+ ```
47
+
48
+ M14 is first because adaptive ranking must be measured against a stable baseline. M15 precedes M16 so semantic chunks have structural boundaries. M18 cannot become active until replay evaluation proves it improves over deterministic ranking.
49
+
50
+ ## M14 — Context Quality Metrics
51
+
52
+ ### Deliverables
53
+
54
+ - A sanitized replay corpus with expected evidence source IDs and task descriptors.
55
+ - Deterministic metrics for evidence recall, irrelevant-token ratio, duplicate evidence, provenance coverage, current-request retention, atomic-group validity, overflow/fallback rate and planning latency.
56
+ - Per-category budget utilization and selection/drop reasons in metadata-only diagnostics.
57
+ - A comparison harness for 0.1 static ranking versus candidate 0.2 strategies.
58
+ - `/context quality` output that exposes aggregate scores and sample counts without content.
59
+
60
+ ### Storage
61
+
62
+ Quality samples store planner/profile versions, source-kind counts, token totals, decisions, outcome labels and timings. They must not store message, summary, memory, artifact or project text.
63
+
64
+ ### Acceptance
65
+
66
+ - Replaying an identical fixture produces byte-stable non-timing metric output; wall-clock measurements are reported separately.
67
+ - Deleting SQLite and replaying canonical sources reproduces the same quality aggregates.
68
+ - Metrics add no more than 10% p95 latency when enabled and effectively no regression when disabled.
69
+ - Corrupt or incomplete samples are ignored without affecting planning.
70
+
71
+ ## M15 — Rich Symbol Indexing
72
+
73
+ ### Deliverables
74
+
75
+ - A parser interface in `ds4-context-core` with deterministic regex fallback.
76
+ - Structural chunks for declarations, signatures, parent symbols, imports/references and line ranges.
77
+ - Stable symbol IDs derived from project identity, path, file hash and structural location.
78
+ - Incremental invalidation at changed-file granularity.
79
+ - Initial parser coverage for TypeScript/JavaScript and at least two additional common repository languages, selected by fixture coverage and packaging feasibility.
80
+ - Exact symbol, qualified-name and path lookup ahead of fuzzy retrieval.
81
+
82
+ The parser implementation must not introduce a mandatory native build dependency. A WASM or optional adapter may be used only if clean installation on supported Node versions remains reproducible.
83
+
84
+ ### Acceptance
85
+
86
+ - False symbol matches are lower than the 0.1 regex baseline on the versioned corpus.
87
+ - A file edit invalidates only rows derived from the previous file hash.
88
+ - Unsupported or syntactically invalid files fall back to current text chunking.
89
+ - Project trust, sensitive-file exclusion and live-hash verification remain enforced.
90
+
91
+ ## M16 — Hybrid Semantic Retrieval
92
+
93
+ ### Deliverables
94
+
95
+ - A runtime-neutral embedding port; model invocation remains outside core.
96
+ - Local embedding as the default supported mode. Remote embedding requires explicit provider/model consent and privacy filtering.
97
+ - Derived embedding rows keyed by source hash, chunking version, embedding provider/model and dimensions.
98
+ - Hybrid candidate generation using exact identifiers, FTS and vectors.
99
+ - Deterministic rank fusion, bounded candidate pools, stable tie-breaking and lexical-only fallback.
100
+ - Diagnostics for lexical/vector contribution, model identity, index freshness and fallback reason without storing query or evidence text.
101
+
102
+ Semantic retrieval does not replace exact matching. Exact paths, symbols, quoted phrases and identifiers retain priority.
103
+
104
+ ### Acceptance
105
+
106
+ - Hybrid retrieval improves evidence recall on the M14 corpus without exceeding the irrelevant-token regression threshold.
107
+ - `local-only` sources and queries never reach a remote embedding provider.
108
+ - Model/dimension changes invalidate only the affected derived vectors.
109
+ - Missing model, corrupt vector, timeout or provider failure produces a lexical result rather than a planning failure.
110
+ - No embedding call occurs in the normal turn when all required vectors and query features are already available locally.
111
+
112
+ ## M17 — Cross-Session Project Memory
113
+
114
+ ### Deliverables
115
+
116
+ - Discovery of canonical Pi session JSONL files associated with the same trusted canonical project identity.
117
+ - Materialization of existing project-scoped memory/pin mutations across those sessions into derived SQLite state.
118
+ - Source-session, source-entry, branch, classification, supersession and contradiction provenance for every selected claim.
119
+ - Incremental checkpoints per source session and deterministic rebuild after database deletion.
120
+ - Explicit diagnostics and commands to inspect contributing sessions and exclude a source session.
121
+
122
+ Version 0.2.0 does not silently extract new memories from conversation text. A durable item still originates from an explicit canonical memory/pin mutation.
123
+
124
+ ### Acceptance
125
+
126
+ - No prompt or memory copy is written into project source files or manifests.
127
+ - Only sessions for the exact trusted project identity contribute.
128
+ - Supersession and branch rules remain deterministic across session order changes.
129
+ - Missing, moved, truncated or corrupt session files degrade by excluding unverifiable claims.
130
+ - Privacy classification survives materialization and is rechecked for the active provider.
131
+
132
+ ## M18 — Learned Ranking
133
+
134
+ ### Deliverables
135
+
136
+ - A bounded feature schema based on metadata such as source kind, exact/FTS/vector scores, recency, branch relation, symbol relation, classification eligibility, token cost and prior selection outcome.
137
+ - Local training from versioned sanitized replay labels and explicit feedback recorded as classified Pi custom entries; raw text is never a training feature or persisted sample.
138
+ - Versioned, checksummed model artifacts with deterministic inference and stable tie-breaking.
139
+ - `off`, `shadow` and `active` modes. Shadow mode records aggregate comparison only and does not change context.
140
+ - Automatic fallback to the static ranker for missing, incompatible, corrupt or regressing models.
141
+
142
+ ### Promotion gate
143
+
144
+ The learned ranker may become active only when it:
145
+
146
+ - improves the primary quality score on held-out repositories;
147
+ - does not reduce exact-identifier recall;
148
+ - does not increase privacy violations, atomicity failures or overflow;
149
+ - stays within the planner latency budget;
150
+ - reproduces identical ordering for identical features, model and configuration.
151
+
152
+ If the gate is not met, 0.2.0 ships shadow mode but keeps static ranking active.
153
+
154
+ ## M19 — Runtime Adapter Kit
155
+
156
+ ### Deliverables
157
+
158
+ - A documented adapter contract for canonical history snapshots, model limits, tool atomicity, trusted project roots, completion, privacy enforcement and lifecycle shutdown.
159
+ - A conformance test kit reusable by adapter packages.
160
+ - Capability negotiation so unsupported compaction, provider continuation, embeddings or KV reuse disable only those features.
161
+ - One non-Pi reference adapter chosen after a short compatibility spike.
162
+ - Packaging rules that keep runtime SDK dependencies outside `ds4-context-core`.
163
+
164
+ ### Acceptance
165
+
166
+ - Core imports no runtime SDK and passes the existing boundary test.
167
+ - The reference adapter passes canonical-history, rebuild, privacy, fallback and lifecycle conformance tests.
168
+ - Unsupported capabilities have explicit diagnostics and safe behavior.
169
+ - Pi remains fully compatible and shows no regression when the new adapter kit is installed.
170
+
171
+ ## M20 — Local KV Capability
172
+
173
+ ### Deliverables
174
+
175
+ - An optional adapter capability for local inference runtimes that expose reusable prefix/KV state.
176
+ - Eligibility based on exact provider/model, prompt-prefix hash, tool/system options, privacy policy and model revision.
177
+ - Conservative invalidation on every prefix, model, option, privacy or runtime change.
178
+ - Aggregate hit/miss/saved-prefill diagnostics without cache handles or content.
179
+ - Full-prompt replay after any stale, rejected or unavailable KV state.
180
+
181
+ KV state is an inference optimization, not memory, retrieval evidence or canonical history. Core decides eligibility from hashes; the runtime adapter owns cache handles and transport.
182
+
183
+ ### Acceptance
184
+
185
+ - Identical eligible prefixes reuse KV state; any changed byte or option rejects reuse.
186
+ - Cache loss and runtime restart produce a transparent full replay.
187
+ - No cache handle is persisted in Pi JSONL, manifests or DS4 SQLite.
188
+ - Privacy filtering runs before prefix verification and cannot be bypassed by cached state.
189
+ - Benchmarks report prefill latency/token savings separately from context occupancy.
190
+
191
+ ## Configuration and compatibility
192
+
193
+ Configuration additions are additive and schema-validated. Names are finalized during each milestone, but the feature groups will map to:
194
+
195
+ - quality measurement;
196
+ - structural project indexing;
197
+ - semantic retrieval and embedding policy;
198
+ - cross-session project memory;
199
+ - learned-ranking mode/model;
200
+ - runtime capabilities and local KV reuse.
201
+
202
+ Semantic retrieval, cross-session memory, learned active ranking and local KV reuse default to disabled for upgrades from 0.1. Existing 0.1 configurations retain their behavior. Unknown or invalid new configuration fails safely through the existing configuration fallback path.
203
+
204
+ SQLite schema changes use forward migrations plus complete rebuild tests from canonical sources. No 0.2 feature may require a canonical-history migration.
205
+
206
+ ## Delivery sequence
207
+
208
+ ### `0.2.0-alpha.1`
209
+
210
+ - M14 quality corpus, metrics and comparison harness.
211
+ - M15 parser interface, structural chunks and fallback.
212
+
213
+ ### `0.2.0-alpha.2`
214
+
215
+ - M16 local hybrid retrieval behind an opt-in flag.
216
+ - M17 cross-session materialization behind an opt-in flag.
217
+
218
+ ### `0.2.0-beta.1`
219
+
220
+ - M18 learned ranker in shadow mode.
221
+ - Privacy, rebuild, corruption and performance hardening for M14–M18.
222
+
223
+ ### `0.2.0-beta.2`
224
+
225
+ - M19 adapter contract, conformance kit and reference adapter.
226
+ - M20 local KV capability for runtimes that support it.
227
+
228
+ ### `0.2.0-rc.1`
229
+
230
+ - Upgrade/rebuild testing from 0.1.0 state.
231
+ - Long-session dogfooding and provider-switch tests.
232
+ - Registry package smoke tests, documentation and release notes.
233
+ - Freeze config, database and adapter-contract schemas for 0.2.0.
234
+
235
+ ## Release gates
236
+
237
+ Version 0.2.0 is ready only when:
238
+
239
+ 1. all 0.1 tests and package-boundary checks still pass;
240
+ 2. new projections rebuild from canonical sources after complete SQLite deletion;
241
+ 3. lexical-only operation remains available with no embedding, rank model or KV runtime;
242
+ 4. privacy E2E tests cover local and explicitly remote embedding paths;
243
+ 5. cross-session tests cover branch, supersession, corruption and project-identity isolation;
244
+ 6. learned ranking passes its promotion gate or remains shadow-only;
245
+ 7. Pi and the reference adapter pass the shared conformance suite;
246
+ 8. feature-disabled p95 planning latency regresses by no more than 10% from the versioned 0.1 baseline;
247
+ 9. long-session tests show no canonical corruption, avoidable overflow or unbounded derived growth;
248
+ 10. `ds4-context-core` and every adapter package use matching versions and pass clean registry-consumer smoke tests;
249
+ 11. CI passes on the minimum supported Node version and the current LTS line;
250
+ 12. migration, privacy, limitations and rollback behavior are documented.
251
+
252
+ ## Deferred beyond 0.2.0
253
+
254
+ Unless required to satisfy a release gate, 0.2.0 does not include automatic memory extraction, summary consensus, a hosted DS4 backend, web UI, visual graph UI, autonomous online training, automatic context-policy tuning or multiple production-grade non-Pi adapters. These remain candidates for later releases.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "ds4-context-engine",
3
- "version": "0.1.0",
3
+ "version": "0.1.2",
4
4
  "description": "Non-destructive, provider-independent context management for Pi.",
5
5
  "type": "module",
6
6
  "license": "MIT",
@@ -50,7 +50,7 @@
50
50
  ]
51
51
  },
52
52
  "dependencies": {
53
- "ds4-context-core": "0.1.0"
53
+ "ds4-context-core": "0.1.2"
54
54
  },
55
55
  "peerDependencies": {
56
56
  "@earendil-works/pi-ai": "0.84.3",
@@ -1,6 +1,7 @@
1
1
  import { randomUUID } from "node:crypto";
2
2
  import { existsSync, realpathSync } from "node:fs";
3
- import { join, resolve } from "node:path";
3
+ import { homedir } from "node:os";
4
+ import { join, parse, resolve } from "node:path";
4
5
  import type {
5
6
  Api,
6
7
  AssistantMessage,
@@ -118,6 +119,23 @@ export type RuntimePhase = "idle" | "initializing" | "disabled" | "observer" | "
118
119
 
119
120
  const READ_ONLY_PROJECT_TOOLS = new Set(["read", "grep", "find", "ls"]);
120
121
 
122
+ function comparablePath(path: string): string {
123
+ let canonical: string;
124
+ try {
125
+ canonical = realpathSync(path);
126
+ } catch {
127
+ canonical = resolve(path);
128
+ }
129
+ return process.platform === "win32" ? canonical.toLowerCase() : canonical;
130
+ }
131
+
132
+ export function isBroadProjectRoot(projectPath: string, homePath = homedir()): boolean {
133
+ const project = comparablePath(projectPath);
134
+ const home = comparablePath(homePath);
135
+ const filesystemRoot = comparablePath(parse(project).root);
136
+ return project === home || project === filesystemRoot;
137
+ }
138
+
121
139
  export interface RuntimeDependencies {
122
140
  agentDir: string;
123
141
  configDirName: string;
@@ -1713,6 +1731,23 @@ export class Ds4ContextRuntime {
1713
1731
  );
1714
1732
  return;
1715
1733
  }
1734
+ if (isBroadProjectRoot(ctx.cwd, this.dependencies.homeDir ?? homedir())) {
1735
+ this.lastProject = {
1736
+ ...emptyProjectDiagnostics(
1737
+ "disabled",
1738
+ true,
1739
+ this.config.context.maxProjectTokens,
1740
+ this.config.project.maxResults,
1741
+ ),
1742
+ projectPath: resolve(ctx.cwd),
1743
+ fallbackReason: "Project indexing is skipped for filesystem roots and the user home directory",
1744
+ };
1745
+ this.logger.warn("project_index.skipped", {
1746
+ reason: "broad-root",
1747
+ projectPath: resolve(ctx.cwd),
1748
+ });
1749
+ return;
1750
+ }
1716
1751
  if (!this.database) return;
1717
1752
 
1718
1753
  try {
@@ -7,18 +7,15 @@ import {
7
7
  type Context,
8
8
  type Model,
9
9
  type SimpleStreamOptions,
10
+ type Tool,
10
11
  } from "@earendil-works/pi-ai";
11
- import { createGrammarToolInputProperties } from "@earendil-works/pi-ai/api/constrained-sampling";
12
- import { openAIResponsesApi } from "@earendil-works/pi-ai/api/openai-responses.lazy";
13
- import { convertResponsesMessages } from "@earendil-works/pi-ai/api/openai-responses-shared";
12
+ import { openAIResponsesApi } from "@earendil-works/pi-ai/compat";
14
13
  import {
15
14
  continuationItemHashes,
16
15
  type NativeContinuationAttempt,
17
16
  type PreparedNativeContinuation,
18
17
  } from "ds4-context-core/continuation/native-continuation";
19
18
 
20
- const OPENAI_TOOL_CALL_PROVIDERS = new Set(["openai", "openai-codex", "opencode"]);
21
-
22
19
  export type OpenAIResponsesStreamDelegate = (
23
20
  model: Model<Api>,
24
21
  context: Context,
@@ -63,32 +60,170 @@ function retryableContinuationError(event: AssistantMessageEvent): boolean {
63
60
  return /previous[ _-]?response|previous_response_id|response(?:\s+with)?\s+id[^\n]{0,80}(?:not found|expired|invalid|unknown)|conversation(?:\s+with)?\s+id[^\n]{0,80}(?:not found|expired|invalid|unknown)/iu.test(message);
64
61
  }
65
62
 
66
- function outputItemHashes(
63
+ interface ParsedTextSignature {
64
+ id: string;
65
+ phase?: "commentary" | "final_answer";
66
+ }
67
+
68
+ function parseTextSignature(signature: string | undefined): ParsedTextSignature | undefined {
69
+ if (!signature) return undefined;
70
+ if (signature.startsWith("{")) {
71
+ try {
72
+ const parsed = JSON.parse(signature) as Record<string, unknown>;
73
+ if (parsed.v === 1 && typeof parsed.id === "string") {
74
+ return parsed.phase === "commentary" || parsed.phase === "final_answer"
75
+ ? { id: parsed.id, phase: parsed.phase }
76
+ : { id: parsed.id };
77
+ }
78
+ } catch {
79
+ // Fall through to the legacy plain-string signature.
80
+ }
81
+ }
82
+ return { id: signature };
83
+ }
84
+
85
+ function shortHash(value: string): string {
86
+ let first = 0xdeadbeef;
87
+ let second = 0x41c6ce57;
88
+ for (let index = 0; index < value.length; index++) {
89
+ const code = value.charCodeAt(index);
90
+ first = Math.imul(first ^ code, 2654435761);
91
+ second = Math.imul(second ^ code, 1597334677);
92
+ }
93
+ first = Math.imul(first ^ (first >>> 16), 2246822507)
94
+ ^ Math.imul(second ^ (second >>> 13), 3266489909);
95
+ second = Math.imul(second ^ (second >>> 16), 2246822507)
96
+ ^ Math.imul(first ^ (first >>> 13), 3266489909);
97
+ return (second >>> 0).toString(36) + (first >>> 0).toString(36);
98
+ }
99
+
100
+ function sanitizeSurrogates(value: string): string {
101
+ return value.replace(
102
+ /[\uD800-\uDBFF](?![\uDC00-\uDFFF])|(?<![\uD800-\uDBFF])[\uDC00-\uDFFF]/gu,
103
+ "",
104
+ );
105
+ }
106
+
107
+ function grammarInputProperties(
108
+ tools: Tool[] | undefined,
109
+ supported: boolean,
110
+ ): ReadonlyMap<string, string> {
111
+ const properties = new Map<string, string>();
112
+ if (!supported) return properties;
113
+
114
+ for (const tool of tools ?? []) {
115
+ const config = tool.constrainedSampling;
116
+ if (!config || config.type !== "grammar") continue;
117
+ const hasDefinition = [config.variants.openai_lark, config.variants.openai_regex]
118
+ .some((value) => typeof value === "string" && value.trim().length > 0);
119
+ if (!hasDefinition) throw new Error(`Grammar tool ${tool.name} has no OpenAI definition`);
120
+
121
+ const schema = tool.parameters as {
122
+ type?: unknown;
123
+ required?: unknown;
124
+ properties?: Record<string, { type?: unknown }>;
125
+ };
126
+ if (schema.type !== "object"
127
+ || !Array.isArray(schema.required)
128
+ || schema.required.length !== 1
129
+ || typeof schema.required[0] !== "string") {
130
+ throw new Error(`Grammar tool ${tool.name} requires one string property`);
131
+ }
132
+ const inputProperty = schema.required[0];
133
+ if (schema.properties?.[inputProperty]?.type !== "string") {
134
+ throw new Error(`Grammar tool ${tool.name} requires one string property`);
135
+ }
136
+ properties.set(tool.name, inputProperty);
137
+ }
138
+ return properties;
139
+ }
140
+
141
+ function responseItems(
142
+ model: Model<Api>,
143
+ context: Context,
144
+ message: AssistantMessage,
145
+ ): unknown[] {
146
+ if (message.provider !== model.provider
147
+ || message.api !== model.api
148
+ || message.model !== model.id
149
+ || message.stopReason === "error"
150
+ || message.stopReason === "aborted") {
151
+ return [];
152
+ }
153
+
154
+ const supportsGrammar = Boolean(
155
+ model.compat
156
+ && "supportsOpenAIGrammarTools" in model.compat
157
+ && model.compat.supportsOpenAIGrammarTools === true,
158
+ );
159
+ const grammarProperties = grammarInputProperties(context.tools, supportsGrammar);
160
+ const items: unknown[] = [];
161
+ let textBlockIndex = 0;
162
+
163
+ for (const block of message.content) {
164
+ if (block.type === "thinking") {
165
+ if (block.thinkingSignature) items.push(JSON.parse(block.thinkingSignature));
166
+ continue;
167
+ }
168
+ if (block.type === "text") {
169
+ const signature = parseTextSignature(block.textSignature);
170
+ const fallbackId = textBlockIndex === 0 ? "msg_pi_0" : `msg_pi_0_${textBlockIndex}`;
171
+ textBlockIndex++;
172
+ const id = !signature?.id
173
+ ? fallbackId
174
+ : signature.id.length > 64
175
+ ? `msg_${shortHash(signature.id)}`
176
+ : signature.id;
177
+ items.push({
178
+ type: "message",
179
+ role: "assistant",
180
+ content: [{ type: "output_text", text: sanitizeSurrogates(block.text), annotations: [] }],
181
+ status: "completed",
182
+ id,
183
+ phase: signature?.phase,
184
+ });
185
+ continue;
186
+ }
187
+
188
+ const [callId, rawItemId] = block.id.split("|");
189
+ const grammarProperty = grammarProperties.get(block.name);
190
+ let itemId = rawItemId;
191
+ if (grammarProperty === undefined && !itemId?.startsWith("fc_")) itemId = undefined;
192
+ const namespace = block.namespace === undefined ? {} : { namespace: block.namespace };
193
+ if (grammarProperty !== undefined) {
194
+ const input = block.arguments[grammarProperty];
195
+ if (typeof input !== "string") {
196
+ throw new Error(`Grammar tool ${block.name} requires string input`);
197
+ }
198
+ items.push({
199
+ type: "custom_tool_call",
200
+ id: itemId,
201
+ call_id: callId,
202
+ name: block.name,
203
+ input: sanitizeSurrogates(input),
204
+ ...namespace,
205
+ });
206
+ } else {
207
+ items.push({
208
+ type: "function_call",
209
+ id: itemId,
210
+ call_id: callId,
211
+ name: block.name,
212
+ arguments: JSON.stringify(block.arguments),
213
+ ...namespace,
214
+ });
215
+ }
216
+ }
217
+ return items;
218
+ }
219
+
220
+ export function openAIResponseItemHashes(
67
221
  model: Model<Api>,
68
222
  context: Context,
69
223
  message: AssistantMessage,
70
224
  ): string[] {
71
225
  try {
72
- const supportsGrammar = Boolean(
73
- model.compat
74
- && "supportsOpenAIGrammarTools" in model.compat
75
- && model.compat.supportsOpenAIGrammarTools === true,
76
- );
77
- const grammarToolInputProperties = createGrammarToolInputProperties(
78
- context.tools,
79
- supportsGrammar,
80
- );
81
- const items = convertResponsesMessages(
82
- model,
83
- { messages: [message] },
84
- OPENAI_TOOL_CALL_PROVIDERS,
85
- {
86
- includeSystemPrompt: false,
87
- grammarToolInputProperties,
88
- },
89
- ).filter((item) => item.type !== "function_call_output"
90
- && item.type !== "custom_tool_call_output");
91
- return continuationItemHashes(items);
226
+ return continuationItemHashes(responseItems(model, context, message));
92
227
  } catch {
93
228
  return [];
94
229
  }
@@ -183,7 +318,11 @@ export function createOpenAIResponsesContinuationStream(
183
318
 
184
319
  if (event.type === "done" && attempt) {
185
320
  try {
186
- controller.complete(attempt, event.message, outputItemHashes(model, context, event.message));
321
+ controller.complete(
322
+ attempt,
323
+ event.message,
324
+ openAIResponseItemHashes(model, context, event.message),
325
+ );
187
326
  } catch {
188
327
  // Provider output remains authoritative if optimization bookkeeping fails.
189
328
  }
@@ -1,4 +1,4 @@
1
- export const EXTENSION_VERSION = "0.1.0";
1
+ export const EXTENSION_VERSION = "0.1.2";
2
2
  export const SUPPORTED_PI_VERSION = "0.84.3";
3
3
  export const OBSERVER_PLANNER_VERSION = "observer-model-aware-v1";
4
4
  export const PLANNER_VERSION = "managed-native-continuation-v1";