pi-ui-extend 1.0.38 → 1.0.39

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (40) hide show
  1. package/dist/app/extensions/extension-ui-controller.js +3 -0
  2. package/dist/app/rendering/extension-entry-renderer.js +2 -0
  3. package/dist/bundled-extensions/question/contract.d.ts +3 -0
  4. package/dist/bundled-extensions/question/contract.js +24 -0
  5. package/dist/bundled-extensions/question/desktop.d.ts +11 -0
  6. package/dist/bundled-extensions/question/desktop.js +143 -0
  7. package/dist/bundled-extensions/question/index.d.ts +1 -0
  8. package/dist/bundled-extensions/question/index.js +5 -1
  9. package/dist/bundled-extensions/question/render.js +6 -0
  10. package/dist/bundled-extensions/question/result.js +71 -9
  11. package/dist/bundled-extensions/question/tool-description.js +4 -3
  12. package/dist/bundled-extensions/question/tui.js +127 -12
  13. package/dist/bundled-extensions/question/types.d.ts +23 -2
  14. package/dist/tool-renderers/question.js +20 -1
  15. package/docs/desktop-markdown-media.md +77 -0
  16. package/docs/desktop-mvp.md +134 -0
  17. package/docs/desktop-task-manager.md +124 -0
  18. package/external/pi-tools-suite/package.json +3 -3
  19. package/external/pi-tools-suite/src/async-subagents/async-subagents.sample.jsonc +8 -0
  20. package/external/pi-tools-suite/src/async-subagents/core/agent-strategy.ts +41 -2
  21. package/external/pi-tools-suite/src/async-subagents/core/config.ts +12 -0
  22. package/external/pi-tools-suite/src/async-subagents/index.ts +6 -2
  23. package/external/pi-tools-suite/src/async-subagents/subagent-overlay.ts +1 -1
  24. package/external/pi-tools-suite/src/dcp/auto-compress.ts +1 -1
  25. package/external/pi-tools-suite/src/dcp/compression-blocks.ts +1 -51
  26. package/external/pi-tools-suite/src/dcp/debug-log.ts +6 -0
  27. package/external/pi-tools-suite/src/dcp/index.ts +27 -125
  28. package/external/pi-tools-suite/src/dcp/prompts.ts +2 -2
  29. package/external/pi-tools-suite/src/dcp/provider-tool-results.ts +2 -1
  30. package/external/pi-tools-suite/src/dcp/pruner-candidates.ts +31 -10
  31. package/external/pi-tools-suite/src/dcp/pruner-compression-blocks.ts +6 -7
  32. package/external/pi-tools-suite/src/dcp/pruner-message-ids.ts +153 -111
  33. package/external/pi-tools-suite/src/dcp/pruner-metadata.ts +28 -17
  34. package/external/pi-tools-suite/src/dcp/pruner-nudge.ts +61 -27
  35. package/external/pi-tools-suite/src/dcp/pruner.ts +21 -13
  36. package/external/pi-tools-suite/src/dcp/state.ts +59 -0
  37. package/external/pi-tools-suite/src/default-pi-tools-suite-config.ts +4 -0
  38. package/external/pi-tools-suite/src/lib/rpc-session-state.ts +34 -0
  39. package/external/pi-tools-suite/src/todo/todo.ts +8 -3
  40. package/package.json +8 -4
@@ -0,0 +1,124 @@
1
+ # Spec: Desktop Project Task Manager
2
+
3
+ ## Type
4
+
5
+ Change
6
+
7
+ ## Goal
8
+
9
+ Add a project-scoped task list to Pix Desktop in a persistent left sidebar and
10
+ allow a saved task to start work immediately in a new session tab.
11
+
12
+ ## Scope
13
+
14
+ - A collapsible, resizable left sidebar with `Tasks` and `Project` tabs.
15
+ - A compact task list with create, edit, delete, filtering, and manual status
16
+ changes.
17
+ - Task type (`bug`, `feature`, `improvement`), status (`backlog`, `todo`,
18
+ `in-progress`, `done`), and priority (`low`, `medium`, `high`, `urgent`).
19
+ - Project-local persistence in `.pi/tasks.json` with a versioned schema.
20
+ - Starting a task in a new ACP session, automatically sending a generated
21
+ prompt, marking the task in progress, and storing the new session id.
22
+ - Opening an already-linked session instead of creating a duplicate.
23
+
24
+ ## Non-goals
25
+
26
+ - Kanban, drag and drop, subtasks, dependencies, assignees, due dates,
27
+ comments, or task history.
28
+ - Automatic transition to `done` when an agent turn finishes.
29
+ - Synchronizing the project task file with the agent's session-local todo list.
30
+ - Concurrent multi-window file merging.
31
+
32
+ ## Behavior
33
+
34
+ 1. The sidebar is visible below the title/session chrome, starts on `Tasks`,
35
+ remembers its width and collapsed state locally, and preserves the main
36
+ session workspace at narrow window sizes.
37
+ 2. Task rows show a title, type, status, priority, and actions. Status and
38
+ priority are communicated by a Lucide icon, text, and color so meaning never
39
+ depends on color alone. `done` is green, `in-progress` is yellow, and idle
40
+ states are neutral; all colors support light and dark themes.
41
+ 3. Creating or editing validates a non-empty title. A description is optional.
42
+ 4. Filters affect only the visible list and do not modify persisted ordering.
43
+ 5. A missing `.pi/tasks.json` is treated as an empty version-1 document.
44
+ 6. Starting an unlinked task creates and selects a new ACP session, persists
45
+ its id and `in-progress` status, appends the generated user prompt to the
46
+ transcript, and sends it immediately.
47
+ 7. Starting a linked task loads that session. A missing/stale linked session is
48
+ reported as a recoverable error and does not silently create another one.
49
+ 8. Task completion remains manual.
50
+
51
+ ## Contracts
52
+
53
+ - Project file: `.pi/tasks.json` containing `{ version: 1, tasks: [...] }`.
54
+ - Every task has a unique id, title, type, status, priority, `createdAt`, and
55
+ `updatedAt`; description and `sessionId` may be omitted.
56
+ - Tauri exposes narrow read/write commands for this one file. Workspace paths
57
+ are validated and `.pi` symlinks may not escape the workspace.
58
+ - The webview sends a complete validated document on each mutation. Writes use
59
+ a same-directory temporary file and replacement to avoid partial JSON.
60
+
61
+ ## Invariants
62
+
63
+ - The task document version must be supported before it is displayed or saved.
64
+ - Duplicate ids, unknown enum values, empty titles, and malformed timestamps
65
+ are rejected.
66
+ - A task is linked to at most one session in this MVP.
67
+ - Starting a task is disabled while another prompt/session operation is active.
68
+ - A failed save does not leave the UI claiming a task state that was not
69
+ persisted.
70
+
71
+ ## Edge cases
72
+
73
+ - Switching workspaces discards the previous workspace's in-memory task view
74
+ and loads the new project file.
75
+ - Empty and missing files have distinct behavior: missing means no tasks;
76
+ malformed or empty JSON is an error.
77
+ - Save, session creation, and prompt failures remain visible and retryable.
78
+ - Collapsing the sidebar keeps its tab rail available; choosing a tab expands
79
+ it.
80
+
81
+ ## Related files
82
+
83
+ - `desktop/src/App.svelte`
84
+ - `desktop/src/components/`
85
+ - `desktop/src/lib/`
86
+ - `desktop/src/styles.css`
87
+ - `desktop/src-tauri/src/lib.rs`
88
+
89
+ ## Verification
90
+
91
+ - Unit tests for task parsing, validation, filtering, prompt generation, and
92
+ sidebar preference parsing.
93
+ - Rust tests for missing/read/write/malformed files and workspace escape.
94
+ - `npm --prefix desktop test`
95
+ - `npm --prefix desktop run check`
96
+ - `npm --prefix desktop run build:web`
97
+ - `cargo test --manifest-path desktop/src-tauri/Cargo.toml`
98
+
99
+ ## Risks / unknowns
100
+
101
+ - Whole-document writes assume one active desktop writer per project.
102
+ - A linked session may later be removed outside Pix Desktop; the first version
103
+ reports this rather than automatically unlinking it.
104
+ - Windows replacement semantics may require a fallback around replacing an
105
+ existing task file while retaining best-effort crash safety.
106
+
107
+ ## Evidence
108
+
109
+ - Confirmed by code: desktop session creation/loading and prompt submission are
110
+ orchestrated in `desktop/src/App.svelte` through `AcpClient`.
111
+ - Confirmed by code: Tauri already validates workspace-contained project file
112
+ reads in `desktop/src-tauri/src/lib.rs`.
113
+ - Confirmed by tests: session tab ordering and ACP request behavior have focused
114
+ Vitest coverage under `desktop/src/lib/`.
115
+ - Confirmed by docs: `DESIGN.md` specifies a compact persistent sidebar,
116
+ semantic theme roles, Lucide icons, and non-color status cues.
117
+ - Implemented: `.pi/tasks.json` now has an explicit version-1 contract and
118
+ rejects unsupported versions rather than guessing a migration.
119
+ - Verified by tests: all 82 desktop Vitest tests and all 8 Rust tests pass;
120
+ Svelte/TypeScript checks report no errors or warnings.
121
+ - Verified visually: browser QA covered clean preference state, resize/collapse,
122
+ tabs, filters, and semantic colors in light/dark themes.
123
+ - Verified natively: actual Tauri UI/backend CRUD, Run, persisted linkage,
124
+ Open session without duplication, and delete cleanup all passed.
@@ -44,9 +44,9 @@
44
44
  "vscode-languageserver-protocol": "^3.17.5"
45
45
  },
46
46
  "peerDependencies": {
47
- "@earendil-works/pi-ai": "0.84.4",
48
- "@earendil-works/pi-coding-agent": "0.84.4",
49
- "@earendil-works/pi-tui": "0.84.4",
47
+ "@earendil-works/pi-ai": "0.85.0",
48
+ "@earendil-works/pi-coding-agent": "0.85.0",
49
+ "@earendil-works/pi-tui": "0.85.0",
50
50
  "typebox": "*"
51
51
  },
52
52
  "devDependencies": {
@@ -234,6 +234,14 @@
234
234
  "description": "Use when the sub-agent should make or plan code changes for a feature, bug fix, or refactor.",
235
235
  "model": "openai-codex/gpt-5.6-sol",
236
236
  "fallbackModels": ["zai/glm-5.3"],
237
+ // Avoid Sol recursively doing routine implementation work for a Sol parent.
238
+ // Luna also escalates substantial implementation to Terra. modelByParent
239
+ // wins over preset role models, so this prevents both Luna -> Sol and
240
+ // Sol -> Sol for implement tasks under the built-in `gpt` preset.
241
+ "modelByParent": {
242
+ "openai-codex/gpt-5.6-luna*": { "model": "openai-codex/gpt-5.6-terra", "fallbackModels": ["zai/glm-5.3"] },
243
+ "openai-codex/gpt-5.6-sol*": { "model": "openai-codex/gpt-5.6-terra", "fallbackModels": ["zai/glm-5.3"] }
244
+ },
237
245
  "thinking": "high"
238
246
  },
239
247
 
@@ -1,6 +1,6 @@
1
1
  import { isGptLikeModel } from "./ultrawork-auto.js";
2
2
 
3
- export type AgentStrategyName = "parallel-first" | "deep-work";
3
+ export type AgentStrategyName = "parallel-first" | "deep-work" | "escalation-aware" | "cost-aware-orchestrator";
4
4
 
5
5
  export interface AgentStrategyOptions {
6
6
  modelRef?: string;
@@ -27,13 +27,40 @@ Default: autonomous deep worker. Build context directly, make progress, edit, an
27
27
  For broad work, keep delegation explicit and bounded: focused review/research/tests/frontend/deep tracks, plus one oracle only for high-stakes uncertainty or final plan checks. Read compact results, decide in the parent session, and report only what matters. If compressing unfinished work, preserve active objective + next step via todo/DCP rules.
28
28
  </agent_strategy>`;
29
29
 
30
+ const ESCALATION_AWARE_STRATEGY_PROMPT = `<agent_strategy name="escalation-aware">
31
+ Execution hint for Pi, not a replacement for system/developer/user instructions.
32
+
33
+ Default: self-sufficient with tiered escalation. Solve narrow, well-bounded work directly, but do not grind through a large or uncertain task at the current model tier when a stronger focused subagent is appropriate. Use todo for the plan and async subagents for escalation; read compact results first and keep the parent context lean.
34
+
35
+ For a Luna parent, prefer Terra workers for substantial multi-file research, tests, or implementation, and escalate deep root-cause analysis, architecture/security review, high-risk decisions, or repeatedly failing complex work to Sol through the deep/review roles. For a Terra parent, handle routine research/tests/implementation directly; escalate deep root-cause analysis, architecture/security review, high-risk decisions, or stubborn complex failures to Sol through deep/review. Do not escalate merely because a plan has several steps, and do not delegate a tiny known-file edit or exact lookup.
36
+
37
+ Keep user questions, plan/todo changes, integration decisions, and the final report in the parent. Independent read-only escalations may run in parallel; serialize overlapping edits unless scopes are clearly disjoint. When work is delegated, synchronize its todo lifecycle: mark it in progress, collect and verify the worker result, then complete/update it before moving on.
38
+ </agent_strategy>`;
39
+
40
+ const COST_AWARE_ORCHESTRATOR_STRATEGY_PROMPT = `<agent_strategy name="cost-aware-orchestrator">
41
+ Execution hint for Pi, not a replacement for system/developer/user instructions.
42
+
43
+ Default: cost-aware orchestration. You are an expensive parent model, so keep the parent session focused on planning, decisions, integration, verification, and the final user-facing answer. For non-trivial todo work, prefer focused async subagents for repo scanning, multi-file research, documentation, tests, frontend work, and implementation steps that would otherwise require several repository tool calls. Read compact subagent results first; inspect raw artifacts or redo work in the parent only when verification or uncertainty requires it.
44
+
45
+ Keep user questions, todo/plan changes, architecture tradeoffs, cross-worker integration decisions, high-stakes review, and the final report in the parent. Do not delegate merely to avoid one cheap exact lookup or a tiny known-file edit. Independent read-only tracks may run in parallel; serialize overlapping edits unless scopes are clearly disjoint. When a todo item is delegated, keep its lifecycle synchronized: mark it in progress, collect and verify the worker result, then complete/update it before moving on.
46
+ </agent_strategy>`;
47
+
30
48
  export function agentStrategyPrompt(options: AgentStrategyOptions = {}): string | undefined {
31
49
  const env = options.env ?? process.env;
32
50
  const override = strategyOverride(env);
33
51
  if (override === "off") return undefined;
34
52
  if (options.customPrompt && shouldSkipCustomPrompt(env)) return undefined;
35
53
 
36
- const strategy = override ?? (isGptLikeModel(options.modelRef) ? "deep-work" : "parallel-first");
54
+ const strategy = override
55
+ ?? (isExpensiveGptParent(options.modelRef)
56
+ ? "cost-aware-orchestrator"
57
+ : isEscalationAwareGptParent(options.modelRef)
58
+ ? "escalation-aware"
59
+ : isGptLikeModel(options.modelRef)
60
+ ? "deep-work"
61
+ : "parallel-first");
62
+ if (strategy === "cost-aware-orchestrator") return COST_AWARE_ORCHESTRATOR_STRATEGY_PROMPT;
63
+ if (strategy === "escalation-aware") return ESCALATION_AWARE_STRATEGY_PROMPT;
37
64
  return strategy === "deep-work" ? DEEP_WORK_STRATEGY_PROMPT : PARALLEL_FIRST_STRATEGY_PROMPT;
38
65
  }
39
66
 
@@ -49,10 +76,22 @@ function strategyOverride(env: NodeJS.ProcessEnv): AgentStrategyName | "off" | u
49
76
  if (FALSE_ENV_PATTERN.test(value)) return "off";
50
77
  if (value === "parallel-first") return "parallel-first";
51
78
  if (value === "deep-work") return "deep-work";
79
+ if (value === "escalation" || value === "escalation-aware") return "escalation-aware";
80
+ if (value === "cost-aware" || value === "cost-aware-orchestrator" || value === "orchestrator") return "cost-aware-orchestrator";
52
81
  if (TRUE_ENV_PATTERN.test(value)) return undefined;
53
82
  return undefined;
54
83
  }
55
84
 
85
+ function isExpensiveGptParent(modelRef: string | undefined): boolean {
86
+ if (!modelRef) return false;
87
+ return /(?:^|\/)gpt-5\.6-sol(?:$|[-.:])/i.test(modelRef.trim());
88
+ }
89
+
90
+ function isEscalationAwareGptParent(modelRef: string | undefined): boolean {
91
+ if (!modelRef) return false;
92
+ return /(?:^|\/)gpt-5\.6-(?:luna|terra)(?:$|[-.:])/i.test(modelRef.trim());
93
+ }
94
+
56
95
  function shouldSkipCustomPrompt(env: NodeJS.ProcessEnv): boolean {
57
96
  const raw = firstEnv(env, "PI_AGENT_STRATEGY_WITH_CUSTOM_PROMPT", "ASYNC_SUBAGENTS_AGENT_STRATEGY_WITH_CUSTOM_PROMPT");
58
97
  return raw ? !TRUE_ENV_PATTERN.test(raw.trim()) : true;
@@ -224,6 +224,10 @@ const BUILTIN_CONFIG: SubagentConfig = {
224
224
  },
225
225
  implement: {
226
226
  description: "Use when the sub-agent should make or plan code changes for a feature, bug fix, or refactor.",
227
+ modelByParent: {
228
+ "openai-codex/gpt-5.6-luna*": { model: "openai-codex/gpt-5.6-terra", fallbackModels: ["zai/glm-5.3"] },
229
+ "openai-codex/gpt-5.6-sol*": { model: "openai-codex/gpt-5.6-terra", fallbackModels: ["zai/glm-5.3"] },
230
+ },
227
231
  thinking: "high",
228
232
  },
229
233
  tests: {
@@ -234,10 +238,18 @@ const BUILTIN_CONFIG: SubagentConfig = {
234
238
  },
235
239
  review: {
236
240
  description: "Use for review/audit of existing code or changes: correctness, security, performance, maintainability, API risks, quality. Do not implement new code.",
241
+ modelByParent: {
242
+ "openai-codex/gpt-5.6-luna*": { model: "openai-codex/gpt-5.6-sol", fallbackModels: ["zai/glm-5.3"] },
243
+ "openai-codex/gpt-5.6-terra*": { model: "openai-codex/gpt-5.6-sol", fallbackModels: ["zai/glm-5.3"] },
244
+ },
237
245
  thinking: "high",
238
246
  },
239
247
  deep: {
240
248
  description: "Use for broad hard reasoning: architecture, root-cause analysis, cross-module impact, complex debugging or tradeoffs.",
249
+ modelByParent: {
250
+ "openai-codex/gpt-5.6-luna*": { model: "openai-codex/gpt-5.6-sol", fallbackModels: ["zai/glm-5.3"] },
251
+ "openai-codex/gpt-5.6-terra*": { model: "openai-codex/gpt-5.6-sol", fallbackModels: ["zai/glm-5.3"] },
252
+ },
241
253
  thinking: "high",
242
254
  },
243
255
  oracle: {
@@ -32,6 +32,7 @@ import { registerSubagentsTool } from "./tools/subagents.js";
32
32
  import type { LiveAgent, SubagentsLiveStateEvent } from "./types.js";
33
33
  import type { AgentState } from "./core/types.js";
34
34
  import { publishStartupSection } from "../startup-section.js";
35
+ import { publishRpcSessionState } from "../lib/rpc-session-state.js";
35
36
 
36
37
  function isTerminalAgentStatus(status: AgentState["status"]): boolean {
37
38
  return status === "done" || status === "failed" || status === "stopped";
@@ -83,8 +84,8 @@ function createLiveStatePayload(
83
84
  }
84
85
 
85
86
  function agentMatchesSession(agent: LiveAgent, sessionFile: string | undefined): boolean {
86
- if (!sessionFile || !agent.parentSession) return true;
87
- return pathsEqual(sessionFile, agent.parentSession);
87
+ if (!sessionFile) return true;
88
+ return agent.parentSession !== undefined && pathsEqual(sessionFile, agent.parentSession);
88
89
  }
89
90
 
90
91
  function isStaleExtensionContextError(error: unknown): boolean {
@@ -100,6 +101,7 @@ export default function (pi: ExtensionAPI) {
100
101
  const subagentOverlay = new SubagentOverlay(liveAgents);
101
102
  let sawAutoUltraworkCandidate = false;
102
103
  let currentSessionFile: string | undefined;
104
+ let currentSessionStateContext: Parameters<typeof publishRpcSessionState>[0];
103
105
  let completionWatchTimer: ReturnType<typeof setInterval> | undefined;
104
106
  publishSubagentPresetsStartupSection();
105
107
 
@@ -109,6 +111,7 @@ export default function (pi: ExtensionAPI) {
109
111
  const liveState = createLiveStatePayload(liveAgents, currentSessionFile);
110
112
  pi.events?.emit?.(SUBAGENTS_LIVE_COUNT_EVENT, { count: liveState.count });
111
113
  pi.events?.emit?.(SUBAGENTS_LIVE_STATE_EVENT, liveState);
114
+ publishRpcSessionState(currentSessionStateContext, SUBAGENTS_LIVE_STATE_EVENT, liveState);
112
115
  updateCompletionWatcher();
113
116
  } catch (error) {
114
117
  ignoreStaleExtensionContextError(error);
@@ -168,6 +171,7 @@ export default function (pi: ExtensionAPI) {
168
171
  try {
169
172
  sawAutoUltraworkCandidate = false;
170
173
  currentSessionFile = sessionFileFromContext(ctx);
174
+ currentSessionStateContext = ctx;
171
175
  subagentOverlay.restoreRunningAgents(ctx.cwd, currentSessionFile);
172
176
  refreshSubagentOverlay();
173
177
  } catch (error) {
@@ -23,7 +23,7 @@ export class SubagentOverlay {
23
23
  if (liveRun.has(agent.id)) continue;
24
24
  const agentDir = path.join(runDir, agent.id);
25
25
  const agentParentSession = readParentSessionLink(agentDir);
26
- if (parentSession && agentParentSession && !pathsEqual(parentSession, agentParentSession)) continue;
26
+ if (parentSession && (!agentParentSession || !pathsEqual(parentSession, agentParentSession))) continue;
27
27
  liveRun.set(agent.id, { runDir, agentId: agent.id, parentSession: agentParentSession, completed: Promise.resolve() });
28
28
  }
29
29
  if (liveRun.size === 0) this.liveAgents.delete(runDir);
@@ -116,7 +116,7 @@ export function buildProgrammaticSummary(
116
116
  return lines.join("\n")
117
117
  }
118
118
 
119
- const SUMMARIZER_SYSTEM_PROMPT = `You summarize a slice of a coding agent's conversation so it can replace the raw messages in context. Produce a dense, continuation-focused summary: preserve user intent, decisions made, files/symbols changed or inspected, exact errors still actionable, verification status, and next steps. Do not infer, invent, or add facts absent from the source; preserve uncertainty instead of filling gaps. Drop full logs, repeated output, and incidental detail. Be concise (roughly 4-10 bullets). Output ONLY the summary text, no preamble.`
119
+ const SUMMARIZER_SYSTEM_PROMPT = `You summarize a slice of a coding agent's conversation so it can replace the raw messages in context. Produce a dense, continuation-focused summary: preserve user intent, decisions made, files/symbols changed or inspected, exact errors still actionable, verification status, and next steps. Preserve exact identifiers and explicit continuity markers verbatim, including uppercase labels before colons; never paraphrase or omit those labels. Do not infer, invent, or add facts absent from the source; preserve uncertainty instead of filling gaps. Drop full logs, repeated output, and incidental detail without quoting or naming the discarded log lines or their markers. Be concise (roughly 4-10 bullets). Output ONLY the summary text, no preamble.`
120
120
 
121
121
  /** Outcome of one summarizer-model attempt, surfaced in DCP debug logs. */
122
122
  export interface ModelSummaryAttempt {
@@ -388,55 +388,6 @@ export function resolveIdToBoundary(
388
388
  const ts = state.messageIdSnapshot.get(id)
389
389
  if (ts !== undefined) return { timestamp: ts }
390
390
 
391
- // ── Stale mNNN fallback: direction-only clamp ────────────────────────
392
- // When compression or pruning removes messages between context passes,
393
- // positional mNNN IDs shift (e.g. an end boundary m145 becomes m123
394
- // after 22 messages are removed). Clamp to the closest valid ID, but
395
- // ONLY in a direction that preserves the range's semantics:
396
- // - start boundary: clamp upward to the first available ID at or after
397
- // the requested number. The start must never move backwards into
398
- // older content; if no such ID exists, the requested start is gone
399
- // and we cannot safely compress — throw.
400
- // - end boundary: clamp downward to the last available ID at or before
401
- // the requested number. The end must never move forwards into newer
402
- // content; if no such ID exists, throw.
403
- // The previous implementation fell back to the highest available ID in
404
- // both "no match" cases, which could clamp e.g. m010..m010 over a
405
- // snapshot of m001..m003 to a single-message block over m003 — silently
406
- // compressing the wrong content.
407
- const mMatch = id.match(/^m(\d+)$/i)
408
- if (mMatch && state.messageIdSnapshot.size > 0) {
409
- const requestedNum = parseInt(mMatch[1]!, 10)
410
- const allIds = sortIds([...state.messageIdSnapshot.keys()])
411
- const allNums = allIds
412
- .map((mid) => {
413
- const n = mid.match(/^m(\d+)$/i)
414
- return n ? { id: mid, num: parseInt(n[1]!, 10) } : null
415
- })
416
- .filter((entry): entry is { id: string; num: number } => entry !== null)
417
- .sort((a, b) => a.num - b.num)
418
-
419
- if (allNums.length > 0) {
420
- let clamped: { id: string; num: number } | undefined
421
- if (field === "startTimestamp") {
422
- clamped = allNums.find((entry) => entry.num >= requestedNum)
423
- } else {
424
- for (let i = allNums.length - 1; i >= 0; i--) {
425
- if (allNums[i]!.num <= requestedNum) {
426
- clamped = allNums[i]
427
- break
428
- }
429
- }
430
- }
431
- if (clamped) {
432
- const clampedMeta = state.messageMetaSnapshot.get(clamped.id)
433
- if (clampedMeta) return resolveMetaBoundary(clampedMeta, field, state)
434
- const clampedTs = state.messageIdSnapshot.get(clamped.id)
435
- if (clampedTs !== undefined) return { timestamp: clampedTs }
436
- }
437
- }
438
- }
439
-
440
391
  throw unknownCompressionIdError(id, state)
441
392
  }
442
393
 
@@ -445,8 +396,7 @@ export function resolveIdToBoundary(
445
396
  * synthetic placeholder for an active compression block (`meta.blockId` is
446
397
  * set), resolve to that block's stored boundary so the caller rolls the
447
398
  * block up instead of nesting a new block on top of the placeholder. This
448
- * matters for both exact-match and clamped mNNN IDs: a model-visible mNNN
449
- * may itself represent a previously compressed section.
399
+ * A model-visible mNNN may itself represent a previously compressed section.
450
400
  */
451
401
  function resolveMetaBoundary(
452
402
  meta: MessageIdMeta,
@@ -127,7 +127,13 @@ export function summarizeDcpState(state: DcpState): Record<string, unknown> {
127
127
  nextBlockId: state.nextBlockId,
128
128
  },
129
129
  inactiveBlocksTail: inactiveBlocks,
130
+ persistentMessageIds: state.messageIdsByStableId.size,
131
+ nextMessageId: state.nextMessageId,
130
132
  prunedTools: state.prunedToolIds.size,
133
+ automaticPruneCheckpoint: {
134
+ turn: state.lastAutomaticPruneTurn,
135
+ blockId: state.lastAutomaticPruneBlockId,
136
+ },
131
137
  providerSeenTools: state.providerSeenToolIds.size,
132
138
  consecutiveEmergencyPasses: state.consecutiveIgnoredStrongNudges,
133
139
  nudgeAnchors: state.nudgeAnchors.map((anchor) => ({
@@ -45,13 +45,7 @@ import {
45
45
  resolveContextThresholds,
46
46
  estimateTokens,
47
47
  } from "./pruner.js"
48
- import {
49
- stripStaleDcpMetadataFromAssistantMessage,
50
- stripStaleDcpMetadataFromMessage,
51
- } from "./pruner-metadata.js"
52
- import {
53
- buildMessageIdControlText,
54
- } from "./pruner-message-ids.js"
48
+ import { stripStaleDcpMetadataFromMessage } from "./pruner-metadata.js"
55
49
  import { summarizeDcpState, writeDcpDebugLog } from "./debug-log.js"
56
50
  import type { DcpNudgeType } from "./pruner-types.js"
57
51
  import { registerCompressTool } from "./compress-tool.js"
@@ -119,104 +113,6 @@ function isDcpControlPlaneMessage(message: any): boolean {
119
113
  return message?.role === "custom" && DCP_CONTROL_PLANE_CUSTOM_TYPES.has(message.customType)
120
114
  }
121
115
 
122
- const DCP_PROVIDER_CONTROL_HEADER = "DCP message ID control data (do not quote or output):"
123
-
124
- function appendTextToContent(content: unknown, text: string): unknown {
125
- if (typeof content === "string") return `${content}\n\n${text}`
126
- if (Array.isArray(content)) {
127
- const textType = content.some((part: any) => part?.type === "input_text") ? "input_text" : "text"
128
- return [...content, { type: textType, text }]
129
- }
130
- return text
131
- }
132
-
133
- function appendDcpControlToMessages(messages: unknown, text: string): unknown {
134
- if (!Array.isArray(messages)) return messages
135
- const block = `${DCP_PROVIDER_CONTROL_HEADER}\n${text}`
136
-
137
- // Keep DCP's volatile ID map at the tail of the provider context instead of
138
- // mutating the stable system/developer prefix. Prefix cache reuse is much
139
- // better when only the newest transcript item changes on each request.
140
- let targetIndex = -1
141
- for (let index = messages.length - 1; index >= 0; index--) {
142
- const message = messages[index] as any
143
- if (!message || typeof message !== "object") continue
144
- // Responses items such as reasoning/function_call_output do not accept a
145
- // `content` field. Keep a function output at the tail by appending to its
146
- // valid `output` string; otherwise scan back to an actual message item.
147
- if (typeof message.type === "string" && message.type !== "message" && message.role === undefined) {
148
- if (message.type === "function_call_output" && typeof message.output === "string") {
149
- return messages.map((candidate: any, candidateIndex) => candidateIndex === index
150
- ? { ...candidate, output: `${candidate.output}\n\n${block}` }
151
- : candidate)
152
- }
153
- continue
154
- }
155
- if (message.role === "system" || message.role === "developer") continue
156
- targetIndex = index
157
- break
158
- }
159
-
160
- if (targetIndex < 0) {
161
- targetIndex = messages.findIndex((message: any) =>
162
- message?.role === "system" || message?.role === "developer"
163
- )
164
- }
165
-
166
- if (targetIndex >= 0) {
167
- return messages.map((message: any, index) => index === targetIndex
168
- ? { ...message, content: appendTextToContent(message.content, block) }
169
- : message)
170
- }
171
- return [{ role: "system", content: block }, ...messages]
172
- }
173
-
174
- function appendDcpControlToAnthropicSystem(system: unknown, text: string): unknown {
175
- const block = `${DCP_PROVIDER_CONTROL_HEADER}\n${text}`
176
- if (typeof system === "string") return `${system}\n\n${block}`
177
- if (Array.isArray(system)) return [...system, { type: "text", text: block }]
178
- if (system === undefined || system === null) return [{ type: "text", text: block }]
179
- return system
180
- }
181
-
182
- function appendDcpControlToGoogleSystemInstruction(systemInstruction: unknown, text: string): unknown {
183
- const block = `${DCP_PROVIDER_CONTROL_HEADER}\n${text}`
184
- if (typeof systemInstruction === "string") return `${systemInstruction}\n\n${block}`
185
- if (systemInstruction === undefined || systemInstruction === null) return block
186
- return systemInstruction
187
- }
188
-
189
- function appendDcpControlToProviderPayload(payload: unknown, text: string): unknown {
190
- if (Array.isArray(payload)) return appendDcpControlToMessages(payload, text)
191
- if (!payload || typeof payload !== "object") return payload
192
- const record = payload as Record<string, unknown>
193
-
194
- if (Array.isArray(record.input)) {
195
- return { ...record, input: appendDcpControlToMessages(record.input, text) }
196
- }
197
-
198
- if (Array.isArray(record.messages)) {
199
- return { ...record, messages: appendDcpControlToMessages(record.messages, text) }
200
- }
201
-
202
- if ("system" in record) {
203
- return { ...record, system: appendDcpControlToAnthropicSystem(record.system, text) }
204
- }
205
-
206
- if (record.config && typeof record.config === "object") {
207
- const config = record.config as Record<string, unknown>
208
- return {
209
- ...record,
210
- config: {
211
- ...config,
212
- systemInstruction: appendDcpControlToGoogleSystemInstruction(config.systemInstruction, text),
213
- },
214
- }
215
- }
216
-
217
- return payload
218
- }
219
-
220
116
  // ---------------------------------------------------------------------------
221
117
  // Module export
222
118
  // ---------------------------------------------------------------------------
@@ -344,15 +240,6 @@ export default async function dcpModule(pi: ExtensionAPI): Promise<void> {
344
240
  }
345
241
  })
346
242
 
347
- // ── 7b. message_end: never persist provider-echoed DCP control markers ─────
348
- pi.on("message_end", async (event, ctx) => {
349
- const effectiveConfig = configForContext(ctx)
350
- if (!effectiveConfig.enabled || event.message?.role !== "assistant") return undefined
351
-
352
- const sanitized = stripStaleDcpMetadataFromAssistantMessage(event.message)
353
- return { message: sanitized }
354
- })
355
-
356
243
  // ── 8. tool_call: record input args for dedup / purge fingerprinting ───────
357
244
  pi.on("tool_call", async (event, _ctx) => {
358
245
  if (!state.toolCalls.has(event.toolCallId)) {
@@ -418,7 +305,7 @@ export default async function dcpModule(pi: ExtensionAPI): Promise<void> {
418
305
  inputMessages: event.messages.length,
419
306
  filteredMessages: contextMessages.length,
420
307
  outputMessages: messages.length,
421
- messageIdControl: "provider-payload",
308
+ messageIdControl: "distributed-carriers",
422
309
  state: summarizeDcpState(state),
423
310
  ...details,
424
311
  }, ctx)
@@ -454,7 +341,20 @@ export default async function dcpModule(pi: ExtensionAPI): Promise<void> {
454
341
  await saveDcpState(ctx, state)
455
342
  }
456
343
  }
344
+ const prunedToolCountBeforeCheckpoint = state.prunedToolIds.size
457
345
  let prunedMessages = applyPruning(contextMessages, state, effectiveConfig)
346
+ const automaticPrunesCommitted = state.prunedToolIds.size - prunedToolCountBeforeCheckpoint
347
+ if (automaticPrunesCommitted > 0) {
348
+ const clearedAnchors = clearDcpNudgeAnchors(state)
349
+ await saveDcpState(ctx, state)
350
+ writeDcpDebugLog(effectiveConfig, "prune.tool_checkpoint", {
351
+ committed: automaticPrunesCommitted,
352
+ clearedAnchors,
353
+ turn: state.currentTurn,
354
+ blockId: state.lastAutomaticPruneBlockId,
355
+ state: summarizeDcpState(state),
356
+ }, ctx)
357
+ }
458
358
  let candidate = null as ReturnType<typeof detectCompressionCandidate>
459
359
  let messageCandidates = [] as ReturnType<typeof detectMessageCompressionCandidates>
460
360
  let emergencySelection = null as ReturnType<typeof analyzeEmergencyCurrentTurn> | null
@@ -746,6 +646,7 @@ export default async function dcpModule(pi: ExtensionAPI): Promise<void> {
746
646
  )
747
647
  if (emergencyPruneResult.prunedToolCallIds.length > 0) {
748
648
  prunedMessages = applyPruning(contextMessages, state, effectiveConfig)
649
+ const clearedAnchors = clearDcpNudgeAnchors(state)
749
650
  state.consecutiveIgnoredStrongNudges = 0
750
651
  emergencySelection = analyzeEmergencyCurrentTurn(prunedMessages, state, effectiveConfig)
751
652
  messageCandidates = emergencyCurrentTurnMessageCandidates(emergencySelection, effectiveConfig)
@@ -758,6 +659,7 @@ export default async function dcpModule(pi: ExtensionAPI): Promise<void> {
758
659
  targetContextPercent,
759
660
  targetRecoveryTokens,
760
661
  prunedOutputs: emergencyPruneResult.prunedToolCallIds.length,
662
+ clearedAnchors,
761
663
  estimatedTokensRecovered: emergencyPruneResult.estimatedTokensRecovered,
762
664
  estimatedContextPercentAfter: Math.max(
763
665
  0,
@@ -786,7 +688,7 @@ export default async function dcpModule(pi: ExtensionAPI): Promise<void> {
786
688
  prunedMessages,
787
689
  state,
788
690
  nudgeType,
789
- { contextPercent },
691
+ { contextPercent, renderedReminder: nudgeText },
790
692
  )
791
693
  if (anchorResult.anchor) {
792
694
  if (anchorResult.updated) {
@@ -836,10 +738,10 @@ export default async function dcpModule(pi: ExtensionAPI): Promise<void> {
836
738
  anchor.type === "context-strong" || anchor.type === "context-soft",
837
739
  )
838
740
  }
839
- applyAnchoredNudges(prunedMessages, state, (anchor) =>
741
+ const nudgeApplication = applyAnchoredNudges(prunedMessages, state, (anchor) =>
840
742
  appendConcreteNudgeGuidance(baseNudgeText(anchor.type), candidate, messageCandidates, state),
841
743
  )
842
- if (state.nudgeAnchors.length !== anchorsBeforeFinalization) {
744
+ if (state.nudgeAnchors.length !== anchorsBeforeFinalization || nudgeApplication.stateChanged) {
843
745
  await saveDcpState(ctx, state)
844
746
  }
845
747
 
@@ -851,7 +753,7 @@ export default async function dcpModule(pi: ExtensionAPI): Promise<void> {
851
753
  })
852
754
  })
853
755
 
854
- // ── 10b. before_provider_request: inject DCP IDs outside transcript ────────
756
+ // ── 10b. before_provider_request: observe accepted tool-result payloads ────
855
757
  pi.on("before_provider_request", async (event, ctx) => {
856
758
  const effectiveConfig = configForContext(ctx)
857
759
  pendingProviderToolIds.clear()
@@ -868,16 +770,16 @@ export default async function dcpModule(pi: ExtensionAPI): Promise<void> {
868
770
  }
869
771
  }
870
772
 
871
- const controlText = buildMessageIdControlText(state)
872
- if (!controlText) return undefined
873
-
874
- const payload = appendDcpControlToProviderPayload(event.payload, controlText)
875
773
  writeDcpDebugLog(effectiveConfig, "provider_payload.message_ids", {
876
- injected: payload !== event.payload,
774
+ injected: false,
775
+ delivery: "distributed-carriers",
877
776
  pendingToolResults: pendingProviderToolIds.size,
878
777
  state: summarizeDcpState(state),
879
778
  }, ctx)
880
- return payload === event.payload ? undefined : payload
779
+ // IDs are already attached to deterministic user/tool-result context
780
+ // carriers. Replacing the payload here would move metadata from the old
781
+ // tail to the new one and break strict append-only Responses continuation.
782
+ return undefined
881
783
  })
882
784
 
883
785
  // Promote pending IDs only after the provider accepted the exact payload.
@@ -118,9 +118,9 @@ You specify boundaries by ID using the injected metadata IDs present in the conv
118
118
  - \`mNNN\` IDs identify raw messages (3 digits, zero-padded, e.g. \`m001\`, \`m042\`)
119
119
  - \`bN\` IDs identify previously compressed blocks
120
120
 
121
- Current raw message IDs are provided in hidden DCP control metadata at the end of the model context.
121
+ Raw message IDs are provided in hidden DCP control metadata attached to stable user/tool-result carriers. IDs may be sparse; determine range order from where their carriers appear in context, not from the numeric suffix.
122
122
  Some message-compression candidate hints include a low/medium/high priority; prefer high-priority stale message IDs for message-mode compression when a full range would be too broad.
123
- The ID reference line appears at the end of the message it belongs to — it identifies the message above it, not the one below it.
123
+ Each carrier's metadata labels that carrier and any immediately preceding assistant message(s).
124
124
  Treat these reference lines as boundary metadata only, not as tool result content.
125
125
 
126
126
  Rules: