pi-ui-extend 1.0.40 → 1.0.41

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (110) hide show
  1. package/dist/app/commands/command-registry.js +2 -2
  2. package/dist/app/commands/command-session-actions.d.ts +0 -1
  3. package/dist/app/commands/command-session-actions.js +22 -13
  4. package/dist/app/icons.d.ts +14 -0
  5. package/dist/app/icons.js +33 -0
  6. package/dist/app/rendering/conversation-tool-renderer.js +2 -2
  7. package/dist/app/rendering/dcp-stats.d.ts +6 -1
  8. package/dist/app/rendering/dcp-stats.js +214 -46
  9. package/dist/app/rendering/editor-panels.js +8 -5
  10. package/dist/app/session/lazy-session-manager.js +12 -1
  11. package/dist/app/session/tabs-controller.d.ts +2 -5
  12. package/dist/app/session/tabs-controller.js +12 -21
  13. package/dist/app/subagents/subagents-model.d.ts +14 -1
  14. package/dist/app/subagents/subagents-model.js +34 -15
  15. package/dist/app/types.d.ts +2 -0
  16. package/dist/bundled-extensions/session-title/config.js +1 -1
  17. package/dist/markdown-format.js +27 -9
  18. package/dist/schemas/pi-tools-suite-schema.d.ts +29 -16
  19. package/dist/schemas/pi-tools-suite-schema.js +46 -31
  20. package/external/pi-tools-suite/README.md +188 -55
  21. package/external/pi-tools-suite/docs/browser-qa-subagent.md +31 -21
  22. package/external/pi-tools-suite/docs/context-gateway-p00-adr.md +216 -0
  23. package/external/pi-tools-suite/docs/context-gateway-p01n-gate-review.md +122 -0
  24. package/external/pi-tools-suite/docs/context-gateway-p01n-measurement.md +133 -0
  25. package/external/pi-tools-suite/docs/context-gateway-p01r-ra-evidence.md +111 -0
  26. package/external/pi-tools-suite/docs/context-gateway-p01r-rb-evidence.md +100 -0
  27. package/external/pi-tools-suite/docs/context-gateway-p01r-rc-evidence.md +69 -0
  28. package/external/pi-tools-suite/docs/context-gateway-p01r-rd-evidence.md +100 -0
  29. package/external/pi-tools-suite/docs/context-gateway-p01r-re-evidence.md +74 -0
  30. package/external/pi-tools-suite/docs/context-gateway-p01r-rf-evidence.md +153 -0
  31. package/external/pi-tools-suite/docs/context-gateway-p01r-rg-evidence.md +235 -0
  32. package/external/pi-tools-suite/docs/subagent-model-pools.md +109 -0
  33. package/external/pi-tools-suite/package.json +3 -0
  34. package/external/pi-tools-suite/src/async-subagents/{private-skills → agents}/browser-qa/scripts/browser-qa-runner.mjs +82 -1
  35. package/external/pi-tools-suite/src/async-subagents/{private-skills/browser-qa/SKILL.md → agents/browser-qa.md} +261 -12
  36. package/external/pi-tools-suite/src/async-subagents/agents/implement.md +20 -0
  37. package/external/pi-tools-suite/src/async-subagents/agents/oracle.md +16 -0
  38. package/external/pi-tools-suite/src/async-subagents/agents/research.md +18 -0
  39. package/external/pi-tools-suite/src/async-subagents/agents/verify.md +18 -0
  40. package/external/pi-tools-suite/src/async-subagents/async-subagents.sample.jsonc +27 -243
  41. package/external/pi-tools-suite/src/async-subagents/commands.ts +6 -2
  42. package/external/pi-tools-suite/src/async-subagents/core/agent-catalog.ts +41 -0
  43. package/external/pi-tools-suite/src/async-subagents/core/agent-strategy.ts +13 -93
  44. package/external/pi-tools-suite/src/async-subagents/core/agents-dir.ts +494 -0
  45. package/external/pi-tools-suite/src/async-subagents/core/browser-qa.ts +9 -0
  46. package/external/pi-tools-suite/src/async-subagents/core/config.ts +200 -143
  47. package/external/pi-tools-suite/src/async-subagents/core/model-fallback.ts +1 -1
  48. package/external/pi-tools-suite/src/async-subagents/core/model-selection.ts +54 -0
  49. package/external/pi-tools-suite/src/async-subagents/core/prompt.ts +7 -6
  50. package/external/pi-tools-suite/src/async-subagents/core/routing.ts +52 -45
  51. package/external/pi-tools-suite/src/async-subagents/core/spawn.ts +12 -4
  52. package/external/pi-tools-suite/src/async-subagents/index.ts +11 -1
  53. package/external/pi-tools-suite/src/async-subagents/lib.ts +6 -2
  54. package/external/pi-tools-suite/src/async-subagents/tools/spawn.ts +46 -18
  55. package/external/pi-tools-suite/src/async-subagents/tools/subagents.ts +3 -2
  56. package/external/pi-tools-suite/src/async-subagents/types.ts +2 -0
  57. package/external/pi-tools-suite/src/config.ts +1 -1
  58. package/external/pi-tools-suite/src/context-gateway/accounting.ts +151 -0
  59. package/external/pi-tools-suite/src/context-gateway/config.ts +111 -0
  60. package/external/pi-tools-suite/src/context-gateway/index.ts +160 -0
  61. package/external/pi-tools-suite/src/context-gateway/metadata-normalization.ts +88 -0
  62. package/external/pi-tools-suite/src/context-gateway/storeless-capabilities.ts +89 -0
  63. package/external/pi-tools-suite/src/context-gateway/telemetry.ts +429 -0
  64. package/external/pi-tools-suite/src/context-gateway/test-output-parser.ts +326 -0
  65. package/external/pi-tools-suite/src/context-gateway/types.ts +152 -0
  66. package/external/pi-tools-suite/src/dcp/auto-compress-budget.ts +106 -0
  67. package/external/pi-tools-suite/src/dcp/auto-compress.ts +810 -106
  68. package/external/pi-tools-suite/src/dcp/commands.ts +64 -139
  69. package/external/pi-tools-suite/src/dcp/compress-tool.ts +369 -35
  70. package/external/pi-tools-suite/src/dcp/compression-blocks.ts +510 -64
  71. package/external/pi-tools-suite/src/dcp/compression-preview.ts +113 -0
  72. package/external/pi-tools-suite/src/dcp/compression-progress.ts +70 -0
  73. package/external/pi-tools-suite/src/dcp/config.ts +36 -61
  74. package/external/pi-tools-suite/src/dcp/conversation-index.ts +421 -0
  75. package/external/pi-tools-suite/src/dcp/debug-log.ts +7 -5
  76. package/external/pi-tools-suite/src/dcp/index.ts +617 -203
  77. package/external/pi-tools-suite/src/dcp/journal.ts +566 -0
  78. package/external/pi-tools-suite/src/dcp/progress-controller.ts +244 -0
  79. package/external/pi-tools-suite/src/dcp/prompts.ts +10 -7
  80. package/external/pi-tools-suite/src/dcp/provider-tool-results.ts +189 -0
  81. package/external/pi-tools-suite/src/dcp/pruner-candidates.ts +298 -78
  82. package/external/pi-tools-suite/src/dcp/pruner-compression-blocks.ts +173 -281
  83. package/external/pi-tools-suite/src/dcp/pruner-emergency.ts +2 -4
  84. package/external/pi-tools-suite/src/dcp/pruner-message-ids.ts +17 -5
  85. package/external/pi-tools-suite/src/dcp/pruner-metadata.ts +11 -1
  86. package/external/pi-tools-suite/src/dcp/pruner-nudge.ts +30 -82
  87. package/external/pi-tools-suite/src/dcp/pruner-tools.ts +22 -133
  88. package/external/pi-tools-suite/src/dcp/pruner.ts +18 -33
  89. package/external/pi-tools-suite/src/dcp/recovery.ts +129 -0
  90. package/external/pi-tools-suite/src/dcp/shadow-plan.ts +127 -0
  91. package/external/pi-tools-suite/src/dcp/state-transaction.ts +102 -0
  92. package/external/pi-tools-suite/src/dcp/state.ts +158 -580
  93. package/external/pi-tools-suite/src/dcp/ui.ts +1 -0
  94. package/external/pi-tools-suite/src/default-pi-tools-suite-config.ts +32 -214
  95. package/external/pi-tools-suite/src/index.ts +9 -0
  96. package/external/pi-tools-suite/src/model-tools/index.ts +76 -42
  97. package/external/pi-tools-suite/src/repo-discovery/index.ts +84 -18
  98. package/external/pi-tools-suite/src/repo-discovery/native-compact.ts +458 -0
  99. package/external/pi-tools-suite/src/session-recovery/index.ts +189 -43
  100. package/external/pi-tools-suite/src/tool-descriptions.ts +39 -35
  101. package/external/pi-tools-suite/src/truncation-metadata-normalizer/index.ts +17 -0
  102. package/package.json +3 -2
  103. package/schemas/pi-tools-suite.json +159 -78
  104. package/external/pi-tools-suite/src/async-subagents/private-skills/browser-qa/references/auth-scaffold-spec.md +0 -78
  105. package/external/pi-tools-suite/src/async-subagents/private-skills/browser-qa/references/qa-design.md +0 -223
  106. package/external/pi-tools-suite/src/dcp/state-persistence.ts +0 -195
  107. /package/external/pi-tools-suite/src/async-subagents/{private-skills/browser-qa/references → agents/browser-qa/examples}/qa-auth.example.jsonc +0 -0
  108. /package/external/pi-tools-suite/src/async-subagents/{private-skills/browser-qa/references → agents/browser-qa/examples}/qa-flow.example.jsonc +0 -0
  109. /package/external/pi-tools-suite/src/async-subagents/{private-skills → agents}/browser-qa/vendor/fflate.LICENSE +0 -0
  110. /package/external/pi-tools-suite/src/async-subagents/{private-skills → agents}/browser-qa/vendor/fflate.mjs +0 -0
@@ -0,0 +1,100 @@
1
+ # Context Gateway P01-R / R-D evidence
2
+
3
+ <!-- markdownlint-disable MD013 -->
4
+
5
+ > Date: 7 September 2026.
6
+ > Repository HEAD during deterministic gate: `daa1b06` with a dirty tested tree.
7
+ > Scope: pure test/build parsing plus passive observe classification. Production tool-result delivery remains passthrough; no archive, reference, fetch, execution or LLM summary is introduced.
8
+
9
+ ## Supported parser scope
10
+
11
+ R-D adds `src/context-gateway/test-output-parser.ts` as a pure in-memory parser for a deliberately small format set:
12
+
13
+ - Bun test terminal summaries (`pass`, `fail`, `Ran ... tests across ... files`) and exact failure/warning lines;
14
+ - TAP v13-style terminal plan/count summaries and exact `not ok` / diagnostic lines;
15
+ - bounded TypeScript compiler diagnostics plus an explicit terminal `Found N error(s)` summary.
16
+
17
+ The parser returns `recognised`, `partial`, or `unrecognised`. Host execution outcome remains authoritative; text does not turn an errored shell call into success. A failed test summary without an exact parsed diagnostic is not considered complete.
18
+
19
+ The parser's ANSI and bare-CR processing creates only an internal parsing view. It does not mutate the original tool result or claim source byte coordinates.
20
+
21
+ ## Conservative failure rules
22
+
23
+ The parser/delivery planner intentionally falls back to passthrough when any of these conditions holds:
24
+
25
+ - SDK/upstream truncation is already reported;
26
+ - timeout/abort is visible;
27
+ - terminal summary conflicts with the host outcome (for example PASS summary followed by command exit 1);
28
+ - multiple recognised formats appear in one output;
29
+ - TypeScript pretty/source/caret lines remain unparsed;
30
+ - a failed summary has no captured exact failure diagnostic;
31
+ - the terminal summary is followed by significant unrecognised output;
32
+ - parser input exceeds the bounded scan budget (default 1 Mi-character); head/tail resemblance is not promoted to completeness;
33
+ - the originating shell command is compound or its scope is unknown;
34
+ - a prospective complete compact representation would exceed the configured byte budget.
35
+
36
+ Compound-command detection is transient and conservative. Observe stores only `simple / compound / unknown`, never the shell command text. False positives merely preserve passthrough; they never rewrite or block execution.
37
+
38
+ ## Prospective compact delivery is all-or-passthrough
39
+
40
+ `planProspectiveTestOutputDelivery` is a decision helper, not a production shaper. For a complete recognised **simple** command it can build a candidate containing:
41
+
42
+ 1. exact host outcome;
43
+ 2. recognised format;
44
+ 3. terminal summary counts;
45
+ 4. every parsed mandatory error/warning line.
46
+
47
+ The candidate is accepted only when the entire representation fits the budget. Diagnostics are never sliced merely to hit a target. `partial` and `unrecognised` inputs have no generic head/tail summary path.
48
+
49
+ No runtime adapter consumes this candidate in R-D. The current Context Gateway `tool_result` handler still returns `undefined`, so delivered content stays byte-equivalent. Connecting a compact delivery requires a later explicit adapter decision with a recovery/lifetime mechanism appropriate to the omitted data. Native truncation alone is not that permission.
50
+
51
+ ## Observe telemetry
52
+
53
+ Shell output is parsed transiently only while Context Gateway is in `observe` mode. The telemetry snapshot stores:
54
+
55
+ - parser version;
56
+ - classification and format counts;
57
+ - scan-limited count;
58
+ - terminal-summary / diagnostic / warning counts for the last observation;
59
+ - safe command-scope enum;
60
+ - prospective `compact-candidate / passthrough` decision and allowlisted reason.
61
+
62
+ It does **not** store command args, diagnostic strings, filenames from diagnostics, raw body, snippets, or hidden-fact fingerprints. Tests compare the input event before/after telemetry and prove it is not mutated.
63
+
64
+ ## Corpus cases
65
+
66
+ The deterministic corpus covers:
67
+
68
+ - complete Bun success;
69
+ - a Bun error in the middle of output plus a warning and later passing test;
70
+ - PASS summary followed by host exit 1;
71
+ - timeout;
72
+ - upstream-truncated output;
73
+ - ANSI and bare-CR overwrite;
74
+ - TAP failure with exact diagnostics;
75
+ - compact TypeScript diagnostics and a pretty/unparsed partial case;
76
+ - mixed/nested recognised formats;
77
+ - unknown custom build output;
78
+ - compound and unknown command scope;
79
+ - an intentionally huge output exceeding the parser scan bound;
80
+ - prospective output that fits the byte budget and the same mandatory diagnostics under a too-small budget.
81
+
82
+ These are format contracts, not claims that every Bun/TAP/TypeScript version is supported. Unknown variants remain passthrough.
83
+
84
+ ## Deterministic gate
85
+
86
+ ```text
87
+ bun test \
88
+ test/context-gateway/test-output-parser.test.ts \
89
+ test/context-gateway/observe.test.ts \
90
+ test/evals/harness.test.ts
91
+ ```
92
+
93
+ Result: **28 pass, 0 fail, 152 assertions**.
94
+
95
+ Additional checks:
96
+
97
+ - `npm run typecheck` — pass.
98
+ - `git diff --check` — pass.
99
+
100
+ No live model call, manual sync, user config change, result replacement or new full-output archive was performed for R-D.
@@ -0,0 +1,74 @@
1
+ # Context Gateway P01-R / R-E evidence
2
+
3
+ <!-- markdownlint-disable MD013 -->
4
+
5
+ > Date: 7 September 2026.
6
+ > Repository HEAD during deterministic gate: `daa1b06` with a dirty tested tree.
7
+ > Scope: mutation/LSP outcome integrity and storeless capability inventory for web/docs, structured JSON, subagents and visual results. No archival adapter, fetcher, grant or lifetime extension was added.
8
+
9
+ ## Mutation and LSP
10
+
11
+ R-E keeps mutations as native passthrough. An actual `applyPatch` run in a temporary fixture proves that its returned `changedFiles`/summary contain only the file changed by that invocation while an unrelated dirty file remains untouched and absent from the result. `getEventPaths` prefers `details.changedFiles` over broader input/patch-shaped paths, so downstream LSP refresh does not derive a whole-workspace diff.
12
+
13
+ The pre-approved patch string is not rewritten by Context Gateway. Errored/cancelled mutation results are no-ops for LSP enrichment and storeless metadata cleanup, preserving their original result object/outcome. Existing LSP integration tests continue to cover successful post-edit diagnostics, changed-file deduplication, deleted-file behavior and aliases. R-B SDK-chain tests separately prove that enrichment/protocol fields survive later storeless cleanup/security stages.
14
+
15
+ No new “partial mutation” inference is introduced. If a producer reports partial/error/cancelled state, Gateway preserves that producer outcome rather than constructing a synthetic success or retry.
16
+
17
+ ## Capability matrix
18
+
19
+ `src/context-gateway/storeless-capabilities.ts` records the current non-archival decisions. It is metadata only; importing it registers no tools/readers/adapters.
20
+
21
+ | Surface | Status | Storeless strategy | Lifetime claim |
22
+ | --- | --- | --- | --- |
23
+ | test/build | limited | pure-parser candidate in observe only | current result |
24
+ | mutation/LSP | supported | native passthrough | current result |
25
+ | web/document | limited | native passthrough | current result |
26
+ | structured JSON | limited | native passthrough | current result |
27
+ | subagent result | limited | native passthrough | producer-managed |
28
+ | visual/image | supported | native passthrough | visual message |
29
+ | direct browser | unsupported | none | none |
30
+ | direct MCP | unsupported | none | none |
31
+
32
+ The matrix is a claim boundary, not feature discovery by name. Unsupported paths do not become supported merely because an output resembles JSON/HTML or contains a filesystem path.
33
+
34
+ ## Web/document path
35
+
36
+ Existing web tool tests cover actual producer-owned HTTP behavior: request metadata, content/link metadata, Ollama→Tavily fallback, auth/API errors, invalid JSON, cancellation and timeout. Storeless Gateway adds no fetch callback and does not reinterpret headings/code or producer pagination/truncation metadata. A bounded R-E fixture confirms that `web_fetch` content remains byte-equivalent through observation.
37
+
38
+ There is no storeless historical page reader. Producer truncation/pagination remains producer-owned, which is why web/document is `limited` rather than an archive-capable adapter.
39
+
40
+ ## Structured JSON
41
+
42
+ No generic JSON parser/field selector is registered. A fixture containing the exact literal `9007199254740993123456789` remains opaque text and is not parsed into a JavaScript number, rounded, ranked or field-selected. Upstream pagination metadata is left on the producer result. A dedicated typed adapter requires a future concrete use case and precision contract.
43
+
44
+ ## Subagents
45
+
46
+ The existing async-subagent result tool already emits a bounded summary plus paths to producer artifacts and explicitly avoids inlining raw logs. Its tests cover missing/running/completed states, structured result generation, artifact paths, cleanup and session lifecycle.
47
+
48
+ P01-R does not read internal subagent history, does not turn an artifact path into an authorization grant and does not extend artifact lifetime. A completed path therefore remains `producer-managed`; it is not promised to survive producer cleanup/resume/export.
49
+
50
+ ## Visual and unsupported direct paths
51
+
52
+ Image parts remain image parts and are accounted separately from text. No textual placeholder replaces them as a context optimization.
53
+
54
+ P00 did not prove a direct parent Context Gateway boundary for browser DOM/network/console or a current direct MCP execution adapter. Both remain explicitly `unsupported`; R-E does not add placeholder adapters to make the matrix look complete.
55
+
56
+ ## Deterministic gate
57
+
58
+ ```text
59
+ bun test \
60
+ test/context-gateway/storeless-format-contracts.test.ts \
61
+ test/lsp.test.ts \
62
+ test/web-search.test.ts \
63
+ test/async-subagents/tools.test.ts \
64
+ test/context-gateway/sdk-pipeline.test.ts
65
+ ```
66
+
67
+ Result: **107 pass, 0 fail, 689 assertions**.
68
+
69
+ Additional checks:
70
+
71
+ - `npm run typecheck` — pass.
72
+ - `git diff --check` — pass.
73
+
74
+ No live model call, manual sync, user-config change or result-shaping adapter was added for R-E.
@@ -0,0 +1,153 @@
1
+ # Context Gateway P01-R / R-F evidence
2
+
3
+ <!-- markdownlint-disable MD013 -->
4
+
5
+ > Date: 7 September 2026.
6
+ > Repository HEAD during deterministic gate: `daa1b06` with a dirty tested tree.
7
+ > Scope: storeless lifecycle and current-result UI/replay only. No new persistent artifact service, native-temp reader endpoint, fork grant, export bundle, DCP lifecycle implementation or manual sync was added.
8
+
9
+ ## Runtime / transient lifecycle
10
+
11
+ Context Gateway observe state is runtime-local. The current implementation tracks
12
+ in-flight tool call IDs separately from aggregate telemetry and refuses an
13
+ `off ↔ observe` mode change while any tool call is still in flight. At a safe
14
+ boundary the mode change clears transient call bindings. `session_start`,
15
+ `session_tree` and `session_shutdown` reset transient and aggregate observation
16
+ state.
17
+
18
+ Deterministic contracts additionally prove:
19
+
20
+ - a mode change requested during an observed call is rejected until that result
21
+ completes; the result is attributed to the old effective observe epoch once;
22
+ - lifecycle reset removes pending bindings and counters rather than carrying them
23
+ into a resumed/switched branch;
24
+ - two Read/shell calls may complete in reverse order without crossing tool class,
25
+ parser command-scope or call identity;
26
+ - an unbound result increments the explicit `unboundResults` counter but its
27
+ bytes/class/parser facts are not aggregated as current-session evidence;
28
+ - extension reload during an in-flight SDK tool remains a separate P00 limitation
29
+ and is not hidden by telemetry; root tab/session generation guards keep a late
30
+ result owned by its origin tab.
31
+
32
+ No suppression ledger is introduced, and no persisted JSONL message is rewritten
33
+ when mode or session lifecycle changes.
34
+
35
+ ## TUI / current result rendering
36
+
37
+ R-B already proves collapsed/expanded renderer equivalence for the exact
38
+ Read/Bash/ast_grep metadata-normalization surfaces, including image preservation.
39
+ The root renderer/session gate additionally covers:
40
+
41
+ - shell running/nonzero/signal rendering;
42
+ - LSP diagnostic severity rendering and mutation tool blocks;
43
+ - truncated collapsed previews and expanded full current-result bodies;
44
+ - persisted history tail/replay, unresolved/completed historical tool calls and
45
+ cancellation while older history is being prepended;
46
+ - stale runtime/session event rejection and late tool-result binding to the
47
+ original inactive tab.
48
+
49
+ These are current/persisted tool-result views; they are not a new route to the
50
+ producer's full native temp output.
51
+
52
+ ## ACP / Desktop lazy current-result hydration
53
+
54
+ Pix already has `pix/session/history`, `pix/session/tool_result` and deferred image
55
+ hydration for Desktop. The lazy tool-result route is explicitly scoped by the ACP
56
+ `sessionId` and the deferred result map created for that session. For a persisted
57
+ result, the backend stores a trusted `sessionPath + byte offset + byte length`
58
+ reference and materializes exactly that persisted JSONL message line on demand.
59
+
60
+ The ACP contract now proves that requesting a deferred result:
61
+
62
+ - does not replay or append a `session/update`;
63
+ - returns the persisted current tool-result content/details only on explicit
64
+ hydration;
65
+ - cannot be performed with another session's ID even when the tool call ID is
66
+ known;
67
+ - frees the duplicate backend deferred copy after successful hydration.
68
+
69
+ Desktop transcript logic keeps deferred results lightweight until expansion,
70
+ hydrates them into the local transcript, and separately hydrates deferred image
71
+ bodies. This is a human/UI operation; it does not create a new model tool result
72
+ or provider input.
73
+
74
+ ## Native temp output: deliberate limitation
75
+
76
+ There is **no existing authenticated host route that reads the contents of a
77
+ `Read`/Bash/ast_grep native temp/full-output handle for TUI/ACP/Desktop**.
78
+ `pix/session/tool_result` is not such a route: it reads a persisted JSONL result,
79
+ not the producer's `fullOutputPath` contents.
80
+
81
+ Therefore R-F adds no “open full native source” button or ACP endpoint. The UI may
82
+ show the current result/status/diagnostics already available through the normal
83
+ message path, but P01-R does not widen filesystem access merely because a result
84
+ contains a temp path. R-C's lifetime limitations remain in force: temp paths can
85
+ be absent, mutable, expired, or unavailable on error/timeout/abort and have no
86
+ resume/fork/export guarantee.
87
+
88
+ ## Deterministic gates
89
+
90
+ Context Gateway lifecycle/result gate:
91
+
92
+ ```text
93
+ bun test \
94
+ test/context-gateway/observe.test.ts \
95
+ test/context-gateway/sdk-pipeline.test.ts \
96
+ test/context-gateway/metadata-normalization.test.ts
97
+ ```
98
+
99
+ Result: **49 pass, 0 fail, 337 assertions**. Suite typecheck and `git diff --check`
100
+ also pass for this gate.
101
+
102
+ Root TUI/session gate:
103
+
104
+ ```text
105
+ node --import tsx --test \
106
+ tests/conversation-tool-renderer.test.ts \
107
+ tests/tool-block-renderer.test.ts \
108
+ tests/conversation-shell-renderer.test.ts \
109
+ tests/session-history.test.ts \
110
+ tests/session-lifecycle-controller.test.ts \
111
+ tests/tabs-controller.test.ts
112
+ ```
113
+
114
+ Result: **98 pass, 0 fail**.
115
+
116
+ ACP lazy current-result gate:
117
+
118
+ ```text
119
+ node --import tsx --test \
120
+ --test-name-pattern='desktop lazy session/load omits tool bodies and retrieves them on demand|desktop lazy session/load reuses an already-live tab runtime|deferred' \
121
+ test/agent.test.ts test/session-replay.test.ts
122
+ npm run typecheck
123
+ ```
124
+
125
+ Result: **4 pass, 0 fail**; ACP typecheck passes. The targeted lazy-result test
126
+ also asserts no hydration `session/update` and denial for another session ID.
127
+
128
+ An additional ACP boundary suite covers persisted lazy materialization and replay:
129
+
130
+ ```text
131
+ node --import tsx --test \
132
+ test/session-replay.test.ts \
133
+ test/session-history-file.test.ts \
134
+ test/desktop-commands.test.ts
135
+ ```
136
+
137
+ Result: **6 pass, 0 fail**. It additionally proves that the Desktop tool-result
138
+ request surface is `sessionId + toolCallId` scoped, ignores arbitrary path-like
139
+ extras, legacy persisted metadata is not retroactively normalized during replay,
140
+ and deferred image/tool bodies are hydrated only from their existing persisted
141
+ session references.
142
+
143
+ Desktop gate:
144
+
145
+ ```text
146
+ npm run check
147
+ npm test -- src/lib/transcript.test.ts src/lib/acp-client.test.ts
148
+ ```
149
+
150
+ Result: Desktop check has **0 errors** and two pre-existing Svelte accessibility
151
+ warnings in `WorkspaceSidebar.svelte`; targeted tests are **28 pass, 0 fail**.
152
+
153
+ No live model call or manual suite sync was used for R-F.
@@ -0,0 +1,235 @@
1
+ # Context Gateway P01-R / R-G storeless release evidence
2
+
3
+ <!-- markdownlint-disable MD013 -->
4
+
5
+ > Scope frozen before the R-G measurement gate on 7 September 2026.
6
+ > Repository HEAD: `daa1b06`; tested source tree is dirty and earlier live run identities must not be inferred from HEAD alone.
7
+ > This document does not authorise Context Gateway enforce mode or P02 durable storage.
8
+
9
+ ## Scope fixed before measurement
10
+
11
+ The candidate storeless release is deliberately small:
12
+
13
+ | Surface | Decision for R-G | Runtime effect |
14
+ | --- | --- | --- |
15
+ | Context Gateway observe | Keep available, `off` by default | Aggregate measurement/classification only; no result replacement or archive. |
16
+ | `repo_*` Native Compact | Accept as explicit opt-in profile | Existing bounded native flags/cursors/validation; no store. |
17
+ | `truncation-metadata-normalizer` | Accept as separate explicit opt-in module | Removes only proven duplicate truncation metadata for measured Read/shell/ast_grep shapes. |
18
+ | test/build parser | Keep observe-only | Safe classification/metrics only; no result replacement. |
19
+ | mutation/LSP | Native passthrough | Existing outcome/diagnostics behavior only. |
20
+ | web/document | Native passthrough / limited | No second fetch, no archive. |
21
+ | structured JSON | Native passthrough / limited | No generic parse/rewrite. |
22
+ | subagent result | Native passthrough / producer-managed lifetime | No parent grant or lifetime extension. |
23
+ | images | Native passthrough | No text replacement. |
24
+ | direct browser / MCP | Unsupported | Not in the claimed denominator. |
25
+
26
+ Provider-specific equality claims are limited to the installed OpenAI-completions
27
+ serialization path already exercised by the deterministic harness. UI claims are
28
+ limited to the current TUI result renderers and the existing ACP/Desktop lazy
29
+ persisted-result flow. Native temp-output contents do not get a new UI/ACP reader.
30
+
31
+ ## Metrics and release criteria fixed before measurement
32
+
33
+ R-G evaluates the following dimensions independently:
34
+
35
+ - task/fact correctness and execution outcome;
36
+ - initial delivered result bytes and all continuation/recovery calls;
37
+ - JSONL/details bytes for metadata-only cleanup;
38
+ - actual checked provider payload equality/usage rather than converting metadata
39
+ bytes into token estimates;
40
+ - tool-call/refusal/retry counts and elapsed time for model-driven evidence that
41
+ already exists;
42
+ - parser/capability CPU work through explicit scan/output bounds rather than a
43
+ claim based on one timing sample;
44
+ - native temp/current-file/index lifetime limitations and unavailable recovery;
45
+ - unsupported paths and failures remain in the report rather than being removed
46
+ from the denominator.
47
+
48
+ There is no aggregate batch-budget claim. P00 proved stable per-call identity and
49
+ source-order delivery, but not a host API that supplies a durable whole-batch
50
+ budget before execution. Batch cost is therefore the sum of actual per-result
51
+ traffic/calls in measurements.
52
+
53
+ ## Measurement results
54
+
55
+ The measurements below were run after the scope and criteria above were written.
56
+
57
+ ### Native Compact deterministic paired corpus
58
+
59
+ `PI_CONTEXT_GATEWAY_BENCHMARK_REPORT=1 bun test test/context-gateway/benchmark.test.ts`
60
+ compares the historical baseline behavior with the current Native Compact profile
61
+ over search, structure, AST, explain and dependency scenarios. All eight critical
62
+ facts are recovered in both arms.
63
+
64
+ | Metric | Baseline | Native Compact | Delta |
65
+ | --- | ---: | ---: | ---: |
66
+ | Delivered repo bytes including continuations | 71,025 | 27,519 | **-61.25%** |
67
+ | Tool calls | 6 | 8 | +2 |
68
+ | Continuation calls | 1 | 3 | +2 |
69
+ | Refusals | 0 | 0 | 0 |
70
+ | Full overrides | 0 | 0 | 0 |
71
+ | Critical facts | 8/8 | 8/8 | equal |
72
+
73
+ This is the expected trade-off for narrow native paging: substantially fewer
74
+ delivered bytes, but potentially more calls. It is not evidence of lower total
75
+ model cost by itself.
76
+
77
+ ### Historical model-driven paired evidence
78
+
79
+ The retained `zai/glm-5.3` paired report
80
+ `test/evals/artifacts/p01n-2026-09-07T17-30-21-479Z/p01n-paired-report.json`
81
+ contains three Prompt Compact / Native Compact pairs; all **6 arm-runs passed**.
82
+ It predates R-A v2 source/corpus identity, so R-G uses it only as historical
83
+ model-behavior/cost evidence rather than pretending it certifies the current dirty
84
+ tree byte-for-byte.
85
+
86
+ | Aggregate metric | Prompt Compact | Native Compact | Delta |
87
+ | --- | ---: | ---: | ---: |
88
+ | Repo-result bytes | 1,286 | 983 | **-23.6%** |
89
+ | All tool-result bytes | 9,737 | 9,149 | **-6.0%** |
90
+ | Tool calls | 14 | 12 | **-14.3%** |
91
+ | Parent tokens | 179,855 | 207,611 | **+15.4%** |
92
+ | Parent cost | $0.0682 | $0.0923 | **+35.4%** |
93
+ | Elapsed | 88.682s | 96.219s | **+8.5%** |
94
+ | Native refusals / full overrides / refusal retries | 0 / 0 / 0 | 0 / 0 / 0 | equal |
95
+
96
+ The live evidence therefore argues **against** making Native Compact the default:
97
+ repo traffic improved, but total token/cost/latency did not consistently improve.
98
+ It remains a useful explicit opt-in for workloads where repo-result volume is the
99
+ binding constraint.
100
+
101
+ ### Metadata normalizer
102
+
103
+ The normalizer changes only `details.truncation.content` on proven Read/shell/
104
+ `ast_grep` SDK shapes. A deterministic 24,007-byte duplicate fixture measured:
105
+
106
+ | Metric | Raw | Normalized | Delta |
107
+ | --- | ---: | ---: | ---: |
108
+ | Serialized tool-result JSON | 48,396 B | 24,376 B | **-24,020 B (-49.6%)** |
109
+ | Ten identical persisted results | — | — | **240,200 B avoided** |
110
+
111
+ The actual SDK fixtures in R-B also remove more than 40 KiB of duplicated metadata
112
+ from large Read/Bash results while preserving visible content and structural
113
+ truncation fields.
114
+
115
+ This is deliberately reported as JSONL/metadata savings only. The installed
116
+ OpenAI-completions serialization contract is byte-equivalent before/after cleanup
117
+ because tool-result `details` are not serialized to that provider payload. R-G
118
+ therefore claims **no token, cache or model-quality saving** from the normalizer.
119
+
120
+ An additional 52,000-byte duplicate sample, matching the approximate size of the
121
+ live observe residual details, measured `104,408 → 52,395 B` for the serialized
122
+ tool-result object and `52,279 → 266 B` for `details`: **52,013 B** removed while
123
+ visible content remained `52,035 B`. Seven reference-machine rounds of 10,000
124
+ normalizer calls on that sample measured approximately `3.5–4.3 µs/call`. These
125
+ timings are diagnostics, not a portable release threshold.
126
+
127
+ ### Residual and negative corpus
128
+
129
+ The live observe-only residual run shows the measured large classes remain real:
130
+ Read/Bash/ast_grep each delivered roughly 52 KiB of result content and roughly
131
+ another 52 KiB of details at the checked boundary. The Bash task assertion failed
132
+ because the model made an extra call; the large Bash observation itself remained
133
+ valid. That failure stays visible rather than being removed from the denominator.
134
+
135
+ R-A–R-F deterministic contracts cover the rest of the offline release corpus:
136
+
137
+ - small/exact/control results and current-file Read continuation;
138
+ - Unicode/CRLF/long-line and changed file/index behavior;
139
+ - successful Bash/ast_grep native recovery plus missing/replaced temp output;
140
+ - timeout/abort/nonzero and unavailable structured recovery handles;
141
+ - parser middle failures, PASS→exit-1 conflict, ANSI/CR, mixed/compound commands,
142
+ unknown formats and a bounded 1 Mi-character parser scan;
143
+ - mutation/LSP outcomes, web errors/cancellation, opaque large-integer JSON,
144
+ subagent producer-managed artifacts, image passthrough and unsupported browser/MCP;
145
+ - late/reloaded/session-switched results, parallel calls and UI/history replay.
146
+
147
+ The test/build parser is bounded but remains observe-only. Its prospective compact
148
+ candidate is all-or-passthrough and must fit the configured result budget; no
149
+ runtime result delivery, provider traffic, calls or token usage are changed by it.
150
+ On a 4,550,063-character synthetic input the 1 Mi-character scan bound produced
151
+ `scan-limited / partial`; seven rounds of 100 parses measured about
152
+ `16.2–16.9 ms/call` on the reference machine. This is CPU/telemetry evidence only,
153
+ not a shaping or provider-saving claim.
154
+
155
+ ## Live-run decision
156
+
157
+ R-G does **not** request another live model run for the selected release scope:
158
+
159
+ - the metadata normalizer is provider-invisible on the checked serializer and has
160
+ deterministic JSONL/provider equality contracts;
161
+ - the test/build parser does not replace results;
162
+ - mutation/LSP, web/JSON/subagent/image surfaces remain native passthrough;
163
+ - Native Compact already has retained paired model-driven evidence, while all
164
+ post-run changes relevant to this scope are deterministic policy/validation
165
+ hardening rather than a new model-visible adapter.
166
+
167
+ The R-A v2 recovery runner remains available for a future claim that specifically
168
+ depends on exact model-driven native-handle recovery. Such a run still requires
169
+ separate authorization and must create a new v2-identity report. R-G does not
170
+ reinterpret the old v1 recovery reports as v2 passes.
171
+
172
+ ## Release decision
173
+
174
+ Accept the verified storeless scope with conservative defaults:
175
+
176
+ 1. **Context Gateway stays `off` by default.** Observe remains an explicit passive
177
+ measurement mode; `enforce` remains unavailable.
178
+ 2. **Native Compact is accepted only as explicit opt-in.** Default
179
+ `repoDiscovery.profile` remains `baseline` because live total-cost evidence is
180
+ mixed despite lower repo-result traffic.
181
+ 3. **`truncation-metadata-normalizer` is accepted as a separate explicit opt-in.**
182
+ It remains disabled by default and claims JSONL/metadata reduction, not token
183
+ savings.
184
+ 4. **Test/build parsing stays observe-only.** No production compact-result adapter
185
+ is enabled without a proven recovery/lifetime contract for omitted data.
186
+ 5. **Mutation/LSP, web/document, JSON, subagent and visual paths remain their R-E
187
+ passthrough/limited decisions.** Browser/MCP direct adapters remain unsupported.
188
+ 6. **P02 durable store remains deferred.** R-G found no task requiring new durable
189
+ snapshot identity strongly enough to justify store/readers/quotas now.
190
+
191
+ Rollback is configuration-only for future calls: disable the normalizer module,
192
+ return `repoDiscovery.profile` to `baseline`, and/or set Context Gateway mode to
193
+ `off`. Rollback does not delete session history, native temp files, user data or
194
+ DCP state and does not require a manual sync operation.
195
+
196
+ ## Deterministic R-G gate
197
+
198
+ The final gate reuses the independently proven R-A–R-F contracts rather than one
199
+ monolithic test process where unrelated LSP/AgentSession suites can interfere with
200
+ shared process/temp state. The required groups are:
201
+
202
+ - Context Gateway/storeless contracts, Native Compact, recovery validator/corpus,
203
+ eval harness and coverage registry;
204
+ - ACP lazy history/replay and typecheck;
205
+ - root tab ownership/session UI contracts;
206
+ - suite/root `git diff --check` and source typechecks.
207
+
208
+ No new live model call, manual suite sync, durable store, reader tool or DCP change
209
+ is part of the R-G release gate.
210
+
211
+ The final external-suite P01-R gate is:
212
+
213
+ ```text
214
+ bun test \
215
+ test/context-gateway \
216
+ test/repo-native-compact.test.ts \
217
+ test/repo-discovery.test.ts \
218
+ test/config.test.ts \
219
+ test/evals/extension-contracts.test.ts \
220
+ test/evals/harness.test.ts \
221
+ test/evals/recovery-corpus.test.ts \
222
+ test/evals/recovery-run-identity.test.ts \
223
+ test/evals/recovery-report.test.ts \
224
+ test/evals/recovery-validation.test.ts
225
+ ```
226
+
227
+ Result: **144 pass, 0 fail, 1011 assertions**; suite typecheck and
228
+ `git diff --check` pass. Root SDK pin remains `0.85.1`; generated schemas are
229
+ unchanged when checked through `node --import tsx` (the sandbox blocks the normal
230
+ `tsx` CLI IPC socket), and root `tsc --noEmit` passes.
231
+
232
+ R-F's separate host/UI gates remain green: root TUI/session **98/98**, ACP lazy
233
+ current-result **4/4** plus an additional **6/6** persisted-history/request-boundary
234
+ suite and typecheck, and Desktop transcript/client **28/28** with Desktop check at
235
+ zero errors (two pre-existing accessibility warnings).
@@ -0,0 +1,109 @@
1
+ # Sub-agent model pools
2
+
3
+ Sub-agents primarily reduce the cost of bounded work and keep intermediate
4
+ source, searches and logs out of the parent context. The parent owns planning,
5
+ integration, decisions and the final answer. Actual savings depend on worker
6
+ quality, retries and how much work the parent repeats; the configuration is
7
+ not a price oracle.
8
+
9
+ ## Five execution modes
10
+
11
+ - `research`: read-only evidence gathering, searches and independent diff review.
12
+ - `implement`: bounded code, documentation, test and frontend changes.
13
+ - `verify`: run checks and interpret logs, without fixing source or tests.
14
+ - `browser-qa`: isolated browser workflow with assertions and visual artifacts.
15
+ - `oracle`: a deliberate strong second opinion, not automatic worker escalation.
16
+
17
+ Task-specific discipline belongs in the brief or `promptAppend`. A new project
18
+ agent is warranted when it adds a durable contract, capabilities or resources,
19
+ not merely a professional title. `verify` has a behavioral no-edit contract;
20
+ shell access is not a read-only filesystem sandbox.
21
+
22
+ ## Agent priority, preset availability
23
+
24
+ Each Markdown profile owns its ordered `models` list. This is one candidate
25
+ chain for initial selection and subsequent quota fallbacks:
26
+
27
+ ```yaml
28
+ ---
29
+ description: Make bounded implementation changes.
30
+ models:
31
+ - zai/glm-5.3-flash
32
+ - openai-codex/gpt-5.6-terra
33
+ - openai-codex/gpt-5.6-luna
34
+ thinking: medium
35
+ ---
36
+ ```
37
+
38
+ A preset contains a set of available models, not a per-agent matrix:
39
+
40
+ ```jsonc
41
+ {
42
+ "asyncSubagents": {
43
+ "presets": {
44
+ "gpt": {
45
+ "description": "Models available for this session",
46
+ "models": [
47
+ "openai-codex/gpt-5.6-luna",
48
+ "openai-codex/gpt-5.6-terra",
49
+ "openai-codex/gpt-5.6-sol"
50
+ ]
51
+ }
52
+ }
53
+ }
54
+ }
55
+ ```
56
+
57
+ The example agent selects Terra, not Luna: the agent's order wins. Sol is
58
+ available in the pool but absent from this worker's chain, so it cannot become
59
+ an automatic implementation fallback. The oracle can declare Sol in its own
60
+ chain. Model references must be exact `provider/model` values, not wildcards.
61
+
62
+ The resolver intersects the agent chain with the selected pool. Runtime
63
+ selection then skips unregistered, unauthenticated or session-exhausted models.
64
+ Tasks with images and browser QA require confirmed image support. The first
65
+ eligible candidate runs; only the remaining eligible candidates are passed to
66
+ quota fallback. An empty intersection or unavailable chain rejects the batch
67
+ before any children or run state are created. Model selection makes no LLM
68
+ completion request; the optional role router is a separate operation.
69
+
70
+ Oracle prefers another provider when possible, but still respects the pool.
71
+ It never substitutes an ordinary cheap candidate merely to avoid a selection
72
+ error. A single-provider pool cannot promise cross-provider independence.
73
+
74
+ Explicit task `model`, CLI `--model`, and `FORCE_CURRENT_MODEL` remain deliberate
75
+ overrides: they bypass the pool and do not add automatic fallback candidates.
76
+ The parent should not use these to evade the configured budget. The pool is
77
+ a selection policy, not a security boundary against explicit overrides.
78
+
79
+ ## Selection and compatibility
80
+
81
+ Use `/subagent-preset <name>`, `AGENTS_PRESET=<name>` or
82
+ `/subagent-preset session <name>`. Clearing the preset uses agent priorities
83
+ without a pool filter. The shipped names remain compatible with saved choices:
84
+ `cheap` is the GLM pool, `gpt` the GPT pool, and `deep` the mixed pool. The last
85
+ name no longer means that ordinary workers should escalate to flagship models.
86
+
87
+ Old role names are not implicit aliases. `quick`, `scan`, `review`, `deep`,
88
+ `docs`, `frontend`, and `tests` work only when explicitly defined as ordinary
89
+ custom/project types. This keeps the effective catalog and accepted names exact.
90
+
91
+ Legacy `model` plus `fallbackModels` and `modelByParent` still load. New profile
92
+ `models` replaces inherited legacy selection fields; an explicit legacy model
93
+ override can still replace an inherited new list. Empty `models` means no
94
+ candidates, not permission to inherit the parent model. Model-less project
95
+ specialists must declare candidates or receive an explicit model override.
96
+
97
+ Legacy preset role matrices remain readable. A preset with `models` uses only
98
+ the pool contract, dropping stale legacy model/thinking/type overrides. Switching
99
+ a higher-priority config layer back to a legacy preset removes the inherited
100
+ pool. Configuration loading never rewrites user files; review old overrides
101
+ when migrating, since explicitly saved profiles can retain expensive models.
102
+
103
+ ## Compact handoff
104
+
105
+ Give workers a scope, acceptance criteria and the evidence needed to start.
106
+ Read compact results first and inspect raw artifacts selectively. One noisy
107
+ sequential investigation can justify a worker; a command whose exit status is
108
+ sufficient usually only needs a saved log, not another LLM. Independent review
109
+ uses a fresh `research` invocation, not a separate built-in persona.