pi-ui-extend 1.0.40 → 1.0.41
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/app/commands/command-registry.js +2 -2
- package/dist/app/commands/command-session-actions.d.ts +0 -1
- package/dist/app/commands/command-session-actions.js +22 -13
- package/dist/app/icons.d.ts +14 -0
- package/dist/app/icons.js +33 -0
- package/dist/app/rendering/conversation-tool-renderer.js +2 -2
- package/dist/app/rendering/dcp-stats.d.ts +6 -1
- package/dist/app/rendering/dcp-stats.js +214 -46
- package/dist/app/rendering/editor-panels.js +8 -5
- package/dist/app/session/lazy-session-manager.js +12 -1
- package/dist/app/session/tabs-controller.d.ts +2 -5
- package/dist/app/session/tabs-controller.js +12 -21
- package/dist/app/subagents/subagents-model.d.ts +14 -1
- package/dist/app/subagents/subagents-model.js +34 -15
- package/dist/app/types.d.ts +2 -0
- package/dist/bundled-extensions/session-title/config.js +1 -1
- package/dist/markdown-format.js +27 -9
- package/dist/schemas/pi-tools-suite-schema.d.ts +29 -16
- package/dist/schemas/pi-tools-suite-schema.js +46 -31
- package/external/pi-tools-suite/README.md +188 -55
- package/external/pi-tools-suite/docs/browser-qa-subagent.md +31 -21
- package/external/pi-tools-suite/docs/context-gateway-p00-adr.md +216 -0
- package/external/pi-tools-suite/docs/context-gateway-p01n-gate-review.md +122 -0
- package/external/pi-tools-suite/docs/context-gateway-p01n-measurement.md +133 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-ra-evidence.md +111 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-rb-evidence.md +100 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-rc-evidence.md +69 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-rd-evidence.md +100 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-re-evidence.md +74 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-rf-evidence.md +153 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-rg-evidence.md +235 -0
- package/external/pi-tools-suite/docs/subagent-model-pools.md +109 -0
- package/external/pi-tools-suite/package.json +3 -0
- package/external/pi-tools-suite/src/async-subagents/{private-skills → agents}/browser-qa/scripts/browser-qa-runner.mjs +82 -1
- package/external/pi-tools-suite/src/async-subagents/{private-skills/browser-qa/SKILL.md → agents/browser-qa.md} +261 -12
- package/external/pi-tools-suite/src/async-subagents/agents/implement.md +20 -0
- package/external/pi-tools-suite/src/async-subagents/agents/oracle.md +16 -0
- package/external/pi-tools-suite/src/async-subagents/agents/research.md +18 -0
- package/external/pi-tools-suite/src/async-subagents/agents/verify.md +18 -0
- package/external/pi-tools-suite/src/async-subagents/async-subagents.sample.jsonc +27 -243
- package/external/pi-tools-suite/src/async-subagents/commands.ts +6 -2
- package/external/pi-tools-suite/src/async-subagents/core/agent-catalog.ts +41 -0
- package/external/pi-tools-suite/src/async-subagents/core/agent-strategy.ts +13 -93
- package/external/pi-tools-suite/src/async-subagents/core/agents-dir.ts +494 -0
- package/external/pi-tools-suite/src/async-subagents/core/browser-qa.ts +9 -0
- package/external/pi-tools-suite/src/async-subagents/core/config.ts +200 -143
- package/external/pi-tools-suite/src/async-subagents/core/model-fallback.ts +1 -1
- package/external/pi-tools-suite/src/async-subagents/core/model-selection.ts +54 -0
- package/external/pi-tools-suite/src/async-subagents/core/prompt.ts +7 -6
- package/external/pi-tools-suite/src/async-subagents/core/routing.ts +52 -45
- package/external/pi-tools-suite/src/async-subagents/core/spawn.ts +12 -4
- package/external/pi-tools-suite/src/async-subagents/index.ts +11 -1
- package/external/pi-tools-suite/src/async-subagents/lib.ts +6 -2
- package/external/pi-tools-suite/src/async-subagents/tools/spawn.ts +46 -18
- package/external/pi-tools-suite/src/async-subagents/tools/subagents.ts +3 -2
- package/external/pi-tools-suite/src/async-subagents/types.ts +2 -0
- package/external/pi-tools-suite/src/config.ts +1 -1
- package/external/pi-tools-suite/src/context-gateway/accounting.ts +151 -0
- package/external/pi-tools-suite/src/context-gateway/config.ts +111 -0
- package/external/pi-tools-suite/src/context-gateway/index.ts +160 -0
- package/external/pi-tools-suite/src/context-gateway/metadata-normalization.ts +88 -0
- package/external/pi-tools-suite/src/context-gateway/storeless-capabilities.ts +89 -0
- package/external/pi-tools-suite/src/context-gateway/telemetry.ts +429 -0
- package/external/pi-tools-suite/src/context-gateway/test-output-parser.ts +326 -0
- package/external/pi-tools-suite/src/context-gateway/types.ts +152 -0
- package/external/pi-tools-suite/src/dcp/auto-compress-budget.ts +106 -0
- package/external/pi-tools-suite/src/dcp/auto-compress.ts +810 -106
- package/external/pi-tools-suite/src/dcp/commands.ts +64 -139
- package/external/pi-tools-suite/src/dcp/compress-tool.ts +369 -35
- package/external/pi-tools-suite/src/dcp/compression-blocks.ts +510 -64
- package/external/pi-tools-suite/src/dcp/compression-preview.ts +113 -0
- package/external/pi-tools-suite/src/dcp/compression-progress.ts +70 -0
- package/external/pi-tools-suite/src/dcp/config.ts +36 -61
- package/external/pi-tools-suite/src/dcp/conversation-index.ts +421 -0
- package/external/pi-tools-suite/src/dcp/debug-log.ts +7 -5
- package/external/pi-tools-suite/src/dcp/index.ts +617 -203
- package/external/pi-tools-suite/src/dcp/journal.ts +566 -0
- package/external/pi-tools-suite/src/dcp/progress-controller.ts +244 -0
- package/external/pi-tools-suite/src/dcp/prompts.ts +10 -7
- package/external/pi-tools-suite/src/dcp/provider-tool-results.ts +189 -0
- package/external/pi-tools-suite/src/dcp/pruner-candidates.ts +298 -78
- package/external/pi-tools-suite/src/dcp/pruner-compression-blocks.ts +173 -281
- package/external/pi-tools-suite/src/dcp/pruner-emergency.ts +2 -4
- package/external/pi-tools-suite/src/dcp/pruner-message-ids.ts +17 -5
- package/external/pi-tools-suite/src/dcp/pruner-metadata.ts +11 -1
- package/external/pi-tools-suite/src/dcp/pruner-nudge.ts +30 -82
- package/external/pi-tools-suite/src/dcp/pruner-tools.ts +22 -133
- package/external/pi-tools-suite/src/dcp/pruner.ts +18 -33
- package/external/pi-tools-suite/src/dcp/recovery.ts +129 -0
- package/external/pi-tools-suite/src/dcp/shadow-plan.ts +127 -0
- package/external/pi-tools-suite/src/dcp/state-transaction.ts +102 -0
- package/external/pi-tools-suite/src/dcp/state.ts +158 -580
- package/external/pi-tools-suite/src/dcp/ui.ts +1 -0
- package/external/pi-tools-suite/src/default-pi-tools-suite-config.ts +32 -214
- package/external/pi-tools-suite/src/index.ts +9 -0
- package/external/pi-tools-suite/src/model-tools/index.ts +76 -42
- package/external/pi-tools-suite/src/repo-discovery/index.ts +84 -18
- package/external/pi-tools-suite/src/repo-discovery/native-compact.ts +458 -0
- package/external/pi-tools-suite/src/session-recovery/index.ts +189 -43
- package/external/pi-tools-suite/src/tool-descriptions.ts +39 -35
- package/external/pi-tools-suite/src/truncation-metadata-normalizer/index.ts +17 -0
- package/package.json +3 -2
- package/schemas/pi-tools-suite.json +159 -78
- package/external/pi-tools-suite/src/async-subagents/private-skills/browser-qa/references/auth-scaffold-spec.md +0 -78
- package/external/pi-tools-suite/src/async-subagents/private-skills/browser-qa/references/qa-design.md +0 -223
- package/external/pi-tools-suite/src/dcp/state-persistence.ts +0 -195
- /package/external/pi-tools-suite/src/async-subagents/{private-skills/browser-qa/references → agents/browser-qa/examples}/qa-auth.example.jsonc +0 -0
- /package/external/pi-tools-suite/src/async-subagents/{private-skills/browser-qa/references → agents/browser-qa/examples}/qa-flow.example.jsonc +0 -0
- /package/external/pi-tools-suite/src/async-subagents/{private-skills → agents}/browser-qa/vendor/fflate.LICENSE +0 -0
- /package/external/pi-tools-suite/src/async-subagents/{private-skills → agents}/browser-qa/vendor/fflate.mjs +0 -0
|
@@ -0,0 +1,100 @@
|
|
|
1
|
+
# Context Gateway P01-R / R-D evidence
|
|
2
|
+
|
|
3
|
+
<!-- markdownlint-disable MD013 -->
|
|
4
|
+
|
|
5
|
+
> Date: 7 September 2026.
|
|
6
|
+
> Repository HEAD during deterministic gate: `daa1b06` with a dirty tested tree.
|
|
7
|
+
> Scope: pure test/build parsing plus passive observe classification. Production tool-result delivery remains passthrough; no archive, reference, fetch, execution or LLM summary is introduced.
|
|
8
|
+
|
|
9
|
+
## Supported parser scope
|
|
10
|
+
|
|
11
|
+
R-D adds `src/context-gateway/test-output-parser.ts` as a pure in-memory parser for a deliberately small format set:
|
|
12
|
+
|
|
13
|
+
- Bun test terminal summaries (`pass`, `fail`, `Ran ... tests across ... files`) and exact failure/warning lines;
|
|
14
|
+
- TAP v13-style terminal plan/count summaries and exact `not ok` / diagnostic lines;
|
|
15
|
+
- bounded TypeScript compiler diagnostics plus an explicit terminal `Found N error(s)` summary.
|
|
16
|
+
|
|
17
|
+
The parser returns `recognised`, `partial`, or `unrecognised`. Host execution outcome remains authoritative; text does not turn an errored shell call into success. A failed test summary without an exact parsed diagnostic is not considered complete.
|
|
18
|
+
|
|
19
|
+
The parser's ANSI and bare-CR processing creates only an internal parsing view. It does not mutate the original tool result or claim source byte coordinates.
|
|
20
|
+
|
|
21
|
+
## Conservative failure rules
|
|
22
|
+
|
|
23
|
+
The parser/delivery planner intentionally falls back to passthrough when any of these conditions holds:
|
|
24
|
+
|
|
25
|
+
- SDK/upstream truncation is already reported;
|
|
26
|
+
- timeout/abort is visible;
|
|
27
|
+
- terminal summary conflicts with the host outcome (for example PASS summary followed by command exit 1);
|
|
28
|
+
- multiple recognised formats appear in one output;
|
|
29
|
+
- TypeScript pretty/source/caret lines remain unparsed;
|
|
30
|
+
- a failed summary has no captured exact failure diagnostic;
|
|
31
|
+
- the terminal summary is followed by significant unrecognised output;
|
|
32
|
+
- parser input exceeds the bounded scan budget (default 1 Mi-character); head/tail resemblance is not promoted to completeness;
|
|
33
|
+
- the originating shell command is compound or its scope is unknown;
|
|
34
|
+
- a prospective complete compact representation would exceed the configured byte budget.
|
|
35
|
+
|
|
36
|
+
Compound-command detection is transient and conservative. Observe stores only `simple / compound / unknown`, never the shell command text. False positives merely preserve passthrough; they never rewrite or block execution.
|
|
37
|
+
|
|
38
|
+
## Prospective compact delivery is all-or-passthrough
|
|
39
|
+
|
|
40
|
+
`planProspectiveTestOutputDelivery` is a decision helper, not a production shaper. For a complete recognised **simple** command it can build a candidate containing:
|
|
41
|
+
|
|
42
|
+
1. exact host outcome;
|
|
43
|
+
2. recognised format;
|
|
44
|
+
3. terminal summary counts;
|
|
45
|
+
4. every parsed mandatory error/warning line.
|
|
46
|
+
|
|
47
|
+
The candidate is accepted only when the entire representation fits the budget. Diagnostics are never sliced merely to hit a target. `partial` and `unrecognised` inputs have no generic head/tail summary path.
|
|
48
|
+
|
|
49
|
+
No runtime adapter consumes this candidate in R-D. The current Context Gateway `tool_result` handler still returns `undefined`, so delivered content stays byte-equivalent. Connecting a compact delivery requires a later explicit adapter decision with a recovery/lifetime mechanism appropriate to the omitted data. Native truncation alone is not that permission.
|
|
50
|
+
|
|
51
|
+
## Observe telemetry
|
|
52
|
+
|
|
53
|
+
Shell output is parsed transiently only while Context Gateway is in `observe` mode. The telemetry snapshot stores:
|
|
54
|
+
|
|
55
|
+
- parser version;
|
|
56
|
+
- classification and format counts;
|
|
57
|
+
- scan-limited count;
|
|
58
|
+
- terminal-summary / diagnostic / warning counts for the last observation;
|
|
59
|
+
- safe command-scope enum;
|
|
60
|
+
- prospective `compact-candidate / passthrough` decision and allowlisted reason.
|
|
61
|
+
|
|
62
|
+
It does **not** store command args, diagnostic strings, filenames from diagnostics, raw body, snippets, or hidden-fact fingerprints. Tests compare the input event before/after telemetry and prove it is not mutated.
|
|
63
|
+
|
|
64
|
+
## Corpus cases
|
|
65
|
+
|
|
66
|
+
The deterministic corpus covers:
|
|
67
|
+
|
|
68
|
+
- complete Bun success;
|
|
69
|
+
- a Bun error in the middle of output plus a warning and later passing test;
|
|
70
|
+
- PASS summary followed by host exit 1;
|
|
71
|
+
- timeout;
|
|
72
|
+
- upstream-truncated output;
|
|
73
|
+
- ANSI and bare-CR overwrite;
|
|
74
|
+
- TAP failure with exact diagnostics;
|
|
75
|
+
- compact TypeScript diagnostics and a pretty/unparsed partial case;
|
|
76
|
+
- mixed/nested recognised formats;
|
|
77
|
+
- unknown custom build output;
|
|
78
|
+
- compound and unknown command scope;
|
|
79
|
+
- an intentionally huge output exceeding the parser scan bound;
|
|
80
|
+
- prospective output that fits the byte budget and the same mandatory diagnostics under a too-small budget.
|
|
81
|
+
|
|
82
|
+
These are format contracts, not claims that every Bun/TAP/TypeScript version is supported. Unknown variants remain passthrough.
|
|
83
|
+
|
|
84
|
+
## Deterministic gate
|
|
85
|
+
|
|
86
|
+
```text
|
|
87
|
+
bun test \
|
|
88
|
+
test/context-gateway/test-output-parser.test.ts \
|
|
89
|
+
test/context-gateway/observe.test.ts \
|
|
90
|
+
test/evals/harness.test.ts
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
Result: **28 pass, 0 fail, 152 assertions**.
|
|
94
|
+
|
|
95
|
+
Additional checks:
|
|
96
|
+
|
|
97
|
+
- `npm run typecheck` — pass.
|
|
98
|
+
- `git diff --check` — pass.
|
|
99
|
+
|
|
100
|
+
No live model call, manual sync, user config change, result replacement or new full-output archive was performed for R-D.
|
|
@@ -0,0 +1,74 @@
|
|
|
1
|
+
# Context Gateway P01-R / R-E evidence
|
|
2
|
+
|
|
3
|
+
<!-- markdownlint-disable MD013 -->
|
|
4
|
+
|
|
5
|
+
> Date: 7 September 2026.
|
|
6
|
+
> Repository HEAD during deterministic gate: `daa1b06` with a dirty tested tree.
|
|
7
|
+
> Scope: mutation/LSP outcome integrity and storeless capability inventory for web/docs, structured JSON, subagents and visual results. No archival adapter, fetcher, grant or lifetime extension was added.
|
|
8
|
+
|
|
9
|
+
## Mutation and LSP
|
|
10
|
+
|
|
11
|
+
R-E keeps mutations as native passthrough. An actual `applyPatch` run in a temporary fixture proves that its returned `changedFiles`/summary contain only the file changed by that invocation while an unrelated dirty file remains untouched and absent from the result. `getEventPaths` prefers `details.changedFiles` over broader input/patch-shaped paths, so downstream LSP refresh does not derive a whole-workspace diff.
|
|
12
|
+
|
|
13
|
+
The pre-approved patch string is not rewritten by Context Gateway. Errored/cancelled mutation results are no-ops for LSP enrichment and storeless metadata cleanup, preserving their original result object/outcome. Existing LSP integration tests continue to cover successful post-edit diagnostics, changed-file deduplication, deleted-file behavior and aliases. R-B SDK-chain tests separately prove that enrichment/protocol fields survive later storeless cleanup/security stages.
|
|
14
|
+
|
|
15
|
+
No new “partial mutation” inference is introduced. If a producer reports partial/error/cancelled state, Gateway preserves that producer outcome rather than constructing a synthetic success or retry.
|
|
16
|
+
|
|
17
|
+
## Capability matrix
|
|
18
|
+
|
|
19
|
+
`src/context-gateway/storeless-capabilities.ts` records the current non-archival decisions. It is metadata only; importing it registers no tools/readers/adapters.
|
|
20
|
+
|
|
21
|
+
| Surface | Status | Storeless strategy | Lifetime claim |
|
|
22
|
+
| --- | --- | --- | --- |
|
|
23
|
+
| test/build | limited | pure-parser candidate in observe only | current result |
|
|
24
|
+
| mutation/LSP | supported | native passthrough | current result |
|
|
25
|
+
| web/document | limited | native passthrough | current result |
|
|
26
|
+
| structured JSON | limited | native passthrough | current result |
|
|
27
|
+
| subagent result | limited | native passthrough | producer-managed |
|
|
28
|
+
| visual/image | supported | native passthrough | visual message |
|
|
29
|
+
| direct browser | unsupported | none | none |
|
|
30
|
+
| direct MCP | unsupported | none | none |
|
|
31
|
+
|
|
32
|
+
The matrix is a claim boundary, not feature discovery by name. Unsupported paths do not become supported merely because an output resembles JSON/HTML or contains a filesystem path.
|
|
33
|
+
|
|
34
|
+
## Web/document path
|
|
35
|
+
|
|
36
|
+
Existing web tool tests cover actual producer-owned HTTP behavior: request metadata, content/link metadata, Ollama→Tavily fallback, auth/API errors, invalid JSON, cancellation and timeout. Storeless Gateway adds no fetch callback and does not reinterpret headings/code or producer pagination/truncation metadata. A bounded R-E fixture confirms that `web_fetch` content remains byte-equivalent through observation.
|
|
37
|
+
|
|
38
|
+
There is no storeless historical page reader. Producer truncation/pagination remains producer-owned, which is why web/document is `limited` rather than an archive-capable adapter.
|
|
39
|
+
|
|
40
|
+
## Structured JSON
|
|
41
|
+
|
|
42
|
+
No generic JSON parser/field selector is registered. A fixture containing the exact literal `9007199254740993123456789` remains opaque text and is not parsed into a JavaScript number, rounded, ranked or field-selected. Upstream pagination metadata is left on the producer result. A dedicated typed adapter requires a future concrete use case and precision contract.
|
|
43
|
+
|
|
44
|
+
## Subagents
|
|
45
|
+
|
|
46
|
+
The existing async-subagent result tool already emits a bounded summary plus paths to producer artifacts and explicitly avoids inlining raw logs. Its tests cover missing/running/completed states, structured result generation, artifact paths, cleanup and session lifecycle.
|
|
47
|
+
|
|
48
|
+
P01-R does not read internal subagent history, does not turn an artifact path into an authorization grant and does not extend artifact lifetime. A completed path therefore remains `producer-managed`; it is not promised to survive producer cleanup/resume/export.
|
|
49
|
+
|
|
50
|
+
## Visual and unsupported direct paths
|
|
51
|
+
|
|
52
|
+
Image parts remain image parts and are accounted separately from text. No textual placeholder replaces them as a context optimization.
|
|
53
|
+
|
|
54
|
+
P00 did not prove a direct parent Context Gateway boundary for browser DOM/network/console or a current direct MCP execution adapter. Both remain explicitly `unsupported`; R-E does not add placeholder adapters to make the matrix look complete.
|
|
55
|
+
|
|
56
|
+
## Deterministic gate
|
|
57
|
+
|
|
58
|
+
```text
|
|
59
|
+
bun test \
|
|
60
|
+
test/context-gateway/storeless-format-contracts.test.ts \
|
|
61
|
+
test/lsp.test.ts \
|
|
62
|
+
test/web-search.test.ts \
|
|
63
|
+
test/async-subagents/tools.test.ts \
|
|
64
|
+
test/context-gateway/sdk-pipeline.test.ts
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
Result: **107 pass, 0 fail, 689 assertions**.
|
|
68
|
+
|
|
69
|
+
Additional checks:
|
|
70
|
+
|
|
71
|
+
- `npm run typecheck` — pass.
|
|
72
|
+
- `git diff --check` — pass.
|
|
73
|
+
|
|
74
|
+
No live model call, manual sync, user-config change or result-shaping adapter was added for R-E.
|
|
@@ -0,0 +1,153 @@
|
|
|
1
|
+
# Context Gateway P01-R / R-F evidence
|
|
2
|
+
|
|
3
|
+
<!-- markdownlint-disable MD013 -->
|
|
4
|
+
|
|
5
|
+
> Date: 7 September 2026.
|
|
6
|
+
> Repository HEAD during deterministic gate: `daa1b06` with a dirty tested tree.
|
|
7
|
+
> Scope: storeless lifecycle and current-result UI/replay only. No new persistent artifact service, native-temp reader endpoint, fork grant, export bundle, DCP lifecycle implementation or manual sync was added.
|
|
8
|
+
|
|
9
|
+
## Runtime / transient lifecycle
|
|
10
|
+
|
|
11
|
+
Context Gateway observe state is runtime-local. The current implementation tracks
|
|
12
|
+
in-flight tool call IDs separately from aggregate telemetry and refuses an
|
|
13
|
+
`off ↔ observe` mode change while any tool call is still in flight. At a safe
|
|
14
|
+
boundary the mode change clears transient call bindings. `session_start`,
|
|
15
|
+
`session_tree` and `session_shutdown` reset transient and aggregate observation
|
|
16
|
+
state.
|
|
17
|
+
|
|
18
|
+
Deterministic contracts additionally prove:
|
|
19
|
+
|
|
20
|
+
- a mode change requested during an observed call is rejected until that result
|
|
21
|
+
completes; the result is attributed to the old effective observe epoch once;
|
|
22
|
+
- lifecycle reset removes pending bindings and counters rather than carrying them
|
|
23
|
+
into a resumed/switched branch;
|
|
24
|
+
- two Read/shell calls may complete in reverse order without crossing tool class,
|
|
25
|
+
parser command-scope or call identity;
|
|
26
|
+
- an unbound result increments the explicit `unboundResults` counter but its
|
|
27
|
+
bytes/class/parser facts are not aggregated as current-session evidence;
|
|
28
|
+
- extension reload during an in-flight SDK tool remains a separate P00 limitation
|
|
29
|
+
and is not hidden by telemetry; root tab/session generation guards keep a late
|
|
30
|
+
result owned by its origin tab.
|
|
31
|
+
|
|
32
|
+
No suppression ledger is introduced, and no persisted JSONL message is rewritten
|
|
33
|
+
when mode or session lifecycle changes.
|
|
34
|
+
|
|
35
|
+
## TUI / current result rendering
|
|
36
|
+
|
|
37
|
+
R-B already proves collapsed/expanded renderer equivalence for the exact
|
|
38
|
+
Read/Bash/ast_grep metadata-normalization surfaces, including image preservation.
|
|
39
|
+
The root renderer/session gate additionally covers:
|
|
40
|
+
|
|
41
|
+
- shell running/nonzero/signal rendering;
|
|
42
|
+
- LSP diagnostic severity rendering and mutation tool blocks;
|
|
43
|
+
- truncated collapsed previews and expanded full current-result bodies;
|
|
44
|
+
- persisted history tail/replay, unresolved/completed historical tool calls and
|
|
45
|
+
cancellation while older history is being prepended;
|
|
46
|
+
- stale runtime/session event rejection and late tool-result binding to the
|
|
47
|
+
original inactive tab.
|
|
48
|
+
|
|
49
|
+
These are current/persisted tool-result views; they are not a new route to the
|
|
50
|
+
producer's full native temp output.
|
|
51
|
+
|
|
52
|
+
## ACP / Desktop lazy current-result hydration
|
|
53
|
+
|
|
54
|
+
Pix already has `pix/session/history`, `pix/session/tool_result` and deferred image
|
|
55
|
+
hydration for Desktop. The lazy tool-result route is explicitly scoped by the ACP
|
|
56
|
+
`sessionId` and the deferred result map created for that session. For a persisted
|
|
57
|
+
result, the backend stores a trusted `sessionPath + byte offset + byte length`
|
|
58
|
+
reference and materializes exactly that persisted JSONL message line on demand.
|
|
59
|
+
|
|
60
|
+
The ACP contract now proves that requesting a deferred result:
|
|
61
|
+
|
|
62
|
+
- does not replay or append a `session/update`;
|
|
63
|
+
- returns the persisted current tool-result content/details only on explicit
|
|
64
|
+
hydration;
|
|
65
|
+
- cannot be performed with another session's ID even when the tool call ID is
|
|
66
|
+
known;
|
|
67
|
+
- frees the duplicate backend deferred copy after successful hydration.
|
|
68
|
+
|
|
69
|
+
Desktop transcript logic keeps deferred results lightweight until expansion,
|
|
70
|
+
hydrates them into the local transcript, and separately hydrates deferred image
|
|
71
|
+
bodies. This is a human/UI operation; it does not create a new model tool result
|
|
72
|
+
or provider input.
|
|
73
|
+
|
|
74
|
+
## Native temp output: deliberate limitation
|
|
75
|
+
|
|
76
|
+
There is **no existing authenticated host route that reads the contents of a
|
|
77
|
+
`Read`/Bash/ast_grep native temp/full-output handle for TUI/ACP/Desktop**.
|
|
78
|
+
`pix/session/tool_result` is not such a route: it reads a persisted JSONL result,
|
|
79
|
+
not the producer's `fullOutputPath` contents.
|
|
80
|
+
|
|
81
|
+
Therefore R-F adds no “open full native source” button or ACP endpoint. The UI may
|
|
82
|
+
show the current result/status/diagnostics already available through the normal
|
|
83
|
+
message path, but P01-R does not widen filesystem access merely because a result
|
|
84
|
+
contains a temp path. R-C's lifetime limitations remain in force: temp paths can
|
|
85
|
+
be absent, mutable, expired, or unavailable on error/timeout/abort and have no
|
|
86
|
+
resume/fork/export guarantee.
|
|
87
|
+
|
|
88
|
+
## Deterministic gates
|
|
89
|
+
|
|
90
|
+
Context Gateway lifecycle/result gate:
|
|
91
|
+
|
|
92
|
+
```text
|
|
93
|
+
bun test \
|
|
94
|
+
test/context-gateway/observe.test.ts \
|
|
95
|
+
test/context-gateway/sdk-pipeline.test.ts \
|
|
96
|
+
test/context-gateway/metadata-normalization.test.ts
|
|
97
|
+
```
|
|
98
|
+
|
|
99
|
+
Result: **49 pass, 0 fail, 337 assertions**. Suite typecheck and `git diff --check`
|
|
100
|
+
also pass for this gate.
|
|
101
|
+
|
|
102
|
+
Root TUI/session gate:
|
|
103
|
+
|
|
104
|
+
```text
|
|
105
|
+
node --import tsx --test \
|
|
106
|
+
tests/conversation-tool-renderer.test.ts \
|
|
107
|
+
tests/tool-block-renderer.test.ts \
|
|
108
|
+
tests/conversation-shell-renderer.test.ts \
|
|
109
|
+
tests/session-history.test.ts \
|
|
110
|
+
tests/session-lifecycle-controller.test.ts \
|
|
111
|
+
tests/tabs-controller.test.ts
|
|
112
|
+
```
|
|
113
|
+
|
|
114
|
+
Result: **98 pass, 0 fail**.
|
|
115
|
+
|
|
116
|
+
ACP lazy current-result gate:
|
|
117
|
+
|
|
118
|
+
```text
|
|
119
|
+
node --import tsx --test \
|
|
120
|
+
--test-name-pattern='desktop lazy session/load omits tool bodies and retrieves them on demand|desktop lazy session/load reuses an already-live tab runtime|deferred' \
|
|
121
|
+
test/agent.test.ts test/session-replay.test.ts
|
|
122
|
+
npm run typecheck
|
|
123
|
+
```
|
|
124
|
+
|
|
125
|
+
Result: **4 pass, 0 fail**; ACP typecheck passes. The targeted lazy-result test
|
|
126
|
+
also asserts no hydration `session/update` and denial for another session ID.
|
|
127
|
+
|
|
128
|
+
An additional ACP boundary suite covers persisted lazy materialization and replay:
|
|
129
|
+
|
|
130
|
+
```text
|
|
131
|
+
node --import tsx --test \
|
|
132
|
+
test/session-replay.test.ts \
|
|
133
|
+
test/session-history-file.test.ts \
|
|
134
|
+
test/desktop-commands.test.ts
|
|
135
|
+
```
|
|
136
|
+
|
|
137
|
+
Result: **6 pass, 0 fail**. It additionally proves that the Desktop tool-result
|
|
138
|
+
request surface is `sessionId + toolCallId` scoped, ignores arbitrary path-like
|
|
139
|
+
extras, legacy persisted metadata is not retroactively normalized during replay,
|
|
140
|
+
and deferred image/tool bodies are hydrated only from their existing persisted
|
|
141
|
+
session references.
|
|
142
|
+
|
|
143
|
+
Desktop gate:
|
|
144
|
+
|
|
145
|
+
```text
|
|
146
|
+
npm run check
|
|
147
|
+
npm test -- src/lib/transcript.test.ts src/lib/acp-client.test.ts
|
|
148
|
+
```
|
|
149
|
+
|
|
150
|
+
Result: Desktop check has **0 errors** and two pre-existing Svelte accessibility
|
|
151
|
+
warnings in `WorkspaceSidebar.svelte`; targeted tests are **28 pass, 0 fail**.
|
|
152
|
+
|
|
153
|
+
No live model call or manual suite sync was used for R-F.
|
|
@@ -0,0 +1,235 @@
|
|
|
1
|
+
# Context Gateway P01-R / R-G storeless release evidence
|
|
2
|
+
|
|
3
|
+
<!-- markdownlint-disable MD013 -->
|
|
4
|
+
|
|
5
|
+
> Scope frozen before the R-G measurement gate on 7 September 2026.
|
|
6
|
+
> Repository HEAD: `daa1b06`; tested source tree is dirty and earlier live run identities must not be inferred from HEAD alone.
|
|
7
|
+
> This document does not authorise Context Gateway enforce mode or P02 durable storage.
|
|
8
|
+
|
|
9
|
+
## Scope fixed before measurement
|
|
10
|
+
|
|
11
|
+
The candidate storeless release is deliberately small:
|
|
12
|
+
|
|
13
|
+
| Surface | Decision for R-G | Runtime effect |
|
|
14
|
+
| --- | --- | --- |
|
|
15
|
+
| Context Gateway observe | Keep available, `off` by default | Aggregate measurement/classification only; no result replacement or archive. |
|
|
16
|
+
| `repo_*` Native Compact | Accept as explicit opt-in profile | Existing bounded native flags/cursors/validation; no store. |
|
|
17
|
+
| `truncation-metadata-normalizer` | Accept as separate explicit opt-in module | Removes only proven duplicate truncation metadata for measured Read/shell/ast_grep shapes. |
|
|
18
|
+
| test/build parser | Keep observe-only | Safe classification/metrics only; no result replacement. |
|
|
19
|
+
| mutation/LSP | Native passthrough | Existing outcome/diagnostics behavior only. |
|
|
20
|
+
| web/document | Native passthrough / limited | No second fetch, no archive. |
|
|
21
|
+
| structured JSON | Native passthrough / limited | No generic parse/rewrite. |
|
|
22
|
+
| subagent result | Native passthrough / producer-managed lifetime | No parent grant or lifetime extension. |
|
|
23
|
+
| images | Native passthrough | No text replacement. |
|
|
24
|
+
| direct browser / MCP | Unsupported | Not in the claimed denominator. |
|
|
25
|
+
|
|
26
|
+
Provider-specific equality claims are limited to the installed OpenAI-completions
|
|
27
|
+
serialization path already exercised by the deterministic harness. UI claims are
|
|
28
|
+
limited to the current TUI result renderers and the existing ACP/Desktop lazy
|
|
29
|
+
persisted-result flow. Native temp-output contents do not get a new UI/ACP reader.
|
|
30
|
+
|
|
31
|
+
## Metrics and release criteria fixed before measurement
|
|
32
|
+
|
|
33
|
+
R-G evaluates the following dimensions independently:
|
|
34
|
+
|
|
35
|
+
- task/fact correctness and execution outcome;
|
|
36
|
+
- initial delivered result bytes and all continuation/recovery calls;
|
|
37
|
+
- JSONL/details bytes for metadata-only cleanup;
|
|
38
|
+
- actual checked provider payload equality/usage rather than converting metadata
|
|
39
|
+
bytes into token estimates;
|
|
40
|
+
- tool-call/refusal/retry counts and elapsed time for model-driven evidence that
|
|
41
|
+
already exists;
|
|
42
|
+
- parser/capability CPU work through explicit scan/output bounds rather than a
|
|
43
|
+
claim based on one timing sample;
|
|
44
|
+
- native temp/current-file/index lifetime limitations and unavailable recovery;
|
|
45
|
+
- unsupported paths and failures remain in the report rather than being removed
|
|
46
|
+
from the denominator.
|
|
47
|
+
|
|
48
|
+
There is no aggregate batch-budget claim. P00 proved stable per-call identity and
|
|
49
|
+
source-order delivery, but not a host API that supplies a durable whole-batch
|
|
50
|
+
budget before execution. Batch cost is therefore the sum of actual per-result
|
|
51
|
+
traffic/calls in measurements.
|
|
52
|
+
|
|
53
|
+
## Measurement results
|
|
54
|
+
|
|
55
|
+
The measurements below were run after the scope and criteria above were written.
|
|
56
|
+
|
|
57
|
+
### Native Compact deterministic paired corpus
|
|
58
|
+
|
|
59
|
+
`PI_CONTEXT_GATEWAY_BENCHMARK_REPORT=1 bun test test/context-gateway/benchmark.test.ts`
|
|
60
|
+
compares the historical baseline behavior with the current Native Compact profile
|
|
61
|
+
over search, structure, AST, explain and dependency scenarios. All eight critical
|
|
62
|
+
facts are recovered in both arms.
|
|
63
|
+
|
|
64
|
+
| Metric | Baseline | Native Compact | Delta |
|
|
65
|
+
| --- | ---: | ---: | ---: |
|
|
66
|
+
| Delivered repo bytes including continuations | 71,025 | 27,519 | **-61.25%** |
|
|
67
|
+
| Tool calls | 6 | 8 | +2 |
|
|
68
|
+
| Continuation calls | 1 | 3 | +2 |
|
|
69
|
+
| Refusals | 0 | 0 | 0 |
|
|
70
|
+
| Full overrides | 0 | 0 | 0 |
|
|
71
|
+
| Critical facts | 8/8 | 8/8 | equal |
|
|
72
|
+
|
|
73
|
+
This is the expected trade-off for narrow native paging: substantially fewer
|
|
74
|
+
delivered bytes, but potentially more calls. It is not evidence of lower total
|
|
75
|
+
model cost by itself.
|
|
76
|
+
|
|
77
|
+
### Historical model-driven paired evidence
|
|
78
|
+
|
|
79
|
+
The retained `zai/glm-5.3` paired report
|
|
80
|
+
`test/evals/artifacts/p01n-2026-09-07T17-30-21-479Z/p01n-paired-report.json`
|
|
81
|
+
contains three Prompt Compact / Native Compact pairs; all **6 arm-runs passed**.
|
|
82
|
+
It predates R-A v2 source/corpus identity, so R-G uses it only as historical
|
|
83
|
+
model-behavior/cost evidence rather than pretending it certifies the current dirty
|
|
84
|
+
tree byte-for-byte.
|
|
85
|
+
|
|
86
|
+
| Aggregate metric | Prompt Compact | Native Compact | Delta |
|
|
87
|
+
| --- | ---: | ---: | ---: |
|
|
88
|
+
| Repo-result bytes | 1,286 | 983 | **-23.6%** |
|
|
89
|
+
| All tool-result bytes | 9,737 | 9,149 | **-6.0%** |
|
|
90
|
+
| Tool calls | 14 | 12 | **-14.3%** |
|
|
91
|
+
| Parent tokens | 179,855 | 207,611 | **+15.4%** |
|
|
92
|
+
| Parent cost | $0.0682 | $0.0923 | **+35.4%** |
|
|
93
|
+
| Elapsed | 88.682s | 96.219s | **+8.5%** |
|
|
94
|
+
| Native refusals / full overrides / refusal retries | 0 / 0 / 0 | 0 / 0 / 0 | equal |
|
|
95
|
+
|
|
96
|
+
The live evidence therefore argues **against** making Native Compact the default:
|
|
97
|
+
repo traffic improved, but total token/cost/latency did not consistently improve.
|
|
98
|
+
It remains a useful explicit opt-in for workloads where repo-result volume is the
|
|
99
|
+
binding constraint.
|
|
100
|
+
|
|
101
|
+
### Metadata normalizer
|
|
102
|
+
|
|
103
|
+
The normalizer changes only `details.truncation.content` on proven Read/shell/
|
|
104
|
+
`ast_grep` SDK shapes. A deterministic 24,007-byte duplicate fixture measured:
|
|
105
|
+
|
|
106
|
+
| Metric | Raw | Normalized | Delta |
|
|
107
|
+
| --- | ---: | ---: | ---: |
|
|
108
|
+
| Serialized tool-result JSON | 48,396 B | 24,376 B | **-24,020 B (-49.6%)** |
|
|
109
|
+
| Ten identical persisted results | — | — | **240,200 B avoided** |
|
|
110
|
+
|
|
111
|
+
The actual SDK fixtures in R-B also remove more than 40 KiB of duplicated metadata
|
|
112
|
+
from large Read/Bash results while preserving visible content and structural
|
|
113
|
+
truncation fields.
|
|
114
|
+
|
|
115
|
+
This is deliberately reported as JSONL/metadata savings only. The installed
|
|
116
|
+
OpenAI-completions serialization contract is byte-equivalent before/after cleanup
|
|
117
|
+
because tool-result `details` are not serialized to that provider payload. R-G
|
|
118
|
+
therefore claims **no token, cache or model-quality saving** from the normalizer.
|
|
119
|
+
|
|
120
|
+
An additional 52,000-byte duplicate sample, matching the approximate size of the
|
|
121
|
+
live observe residual details, measured `104,408 → 52,395 B` for the serialized
|
|
122
|
+
tool-result object and `52,279 → 266 B` for `details`: **52,013 B** removed while
|
|
123
|
+
visible content remained `52,035 B`. Seven reference-machine rounds of 10,000
|
|
124
|
+
normalizer calls on that sample measured approximately `3.5–4.3 µs/call`. These
|
|
125
|
+
timings are diagnostics, not a portable release threshold.
|
|
126
|
+
|
|
127
|
+
### Residual and negative corpus
|
|
128
|
+
|
|
129
|
+
The live observe-only residual run shows the measured large classes remain real:
|
|
130
|
+
Read/Bash/ast_grep each delivered roughly 52 KiB of result content and roughly
|
|
131
|
+
another 52 KiB of details at the checked boundary. The Bash task assertion failed
|
|
132
|
+
because the model made an extra call; the large Bash observation itself remained
|
|
133
|
+
valid. That failure stays visible rather than being removed from the denominator.
|
|
134
|
+
|
|
135
|
+
R-A–R-F deterministic contracts cover the rest of the offline release corpus:
|
|
136
|
+
|
|
137
|
+
- small/exact/control results and current-file Read continuation;
|
|
138
|
+
- Unicode/CRLF/long-line and changed file/index behavior;
|
|
139
|
+
- successful Bash/ast_grep native recovery plus missing/replaced temp output;
|
|
140
|
+
- timeout/abort/nonzero and unavailable structured recovery handles;
|
|
141
|
+
- parser middle failures, PASS→exit-1 conflict, ANSI/CR, mixed/compound commands,
|
|
142
|
+
unknown formats and a bounded 1 Mi-character parser scan;
|
|
143
|
+
- mutation/LSP outcomes, web errors/cancellation, opaque large-integer JSON,
|
|
144
|
+
subagent producer-managed artifacts, image passthrough and unsupported browser/MCP;
|
|
145
|
+
- late/reloaded/session-switched results, parallel calls and UI/history replay.
|
|
146
|
+
|
|
147
|
+
The test/build parser is bounded but remains observe-only. Its prospective compact
|
|
148
|
+
candidate is all-or-passthrough and must fit the configured result budget; no
|
|
149
|
+
runtime result delivery, provider traffic, calls or token usage are changed by it.
|
|
150
|
+
On a 4,550,063-character synthetic input the 1 Mi-character scan bound produced
|
|
151
|
+
`scan-limited / partial`; seven rounds of 100 parses measured about
|
|
152
|
+
`16.2–16.9 ms/call` on the reference machine. This is CPU/telemetry evidence only,
|
|
153
|
+
not a shaping or provider-saving claim.
|
|
154
|
+
|
|
155
|
+
## Live-run decision
|
|
156
|
+
|
|
157
|
+
R-G does **not** request another live model run for the selected release scope:
|
|
158
|
+
|
|
159
|
+
- the metadata normalizer is provider-invisible on the checked serializer and has
|
|
160
|
+
deterministic JSONL/provider equality contracts;
|
|
161
|
+
- the test/build parser does not replace results;
|
|
162
|
+
- mutation/LSP, web/JSON/subagent/image surfaces remain native passthrough;
|
|
163
|
+
- Native Compact already has retained paired model-driven evidence, while all
|
|
164
|
+
post-run changes relevant to this scope are deterministic policy/validation
|
|
165
|
+
hardening rather than a new model-visible adapter.
|
|
166
|
+
|
|
167
|
+
The R-A v2 recovery runner remains available for a future claim that specifically
|
|
168
|
+
depends on exact model-driven native-handle recovery. Such a run still requires
|
|
169
|
+
separate authorization and must create a new v2-identity report. R-G does not
|
|
170
|
+
reinterpret the old v1 recovery reports as v2 passes.
|
|
171
|
+
|
|
172
|
+
## Release decision
|
|
173
|
+
|
|
174
|
+
Accept the verified storeless scope with conservative defaults:
|
|
175
|
+
|
|
176
|
+
1. **Context Gateway stays `off` by default.** Observe remains an explicit passive
|
|
177
|
+
measurement mode; `enforce` remains unavailable.
|
|
178
|
+
2. **Native Compact is accepted only as explicit opt-in.** Default
|
|
179
|
+
`repoDiscovery.profile` remains `baseline` because live total-cost evidence is
|
|
180
|
+
mixed despite lower repo-result traffic.
|
|
181
|
+
3. **`truncation-metadata-normalizer` is accepted as a separate explicit opt-in.**
|
|
182
|
+
It remains disabled by default and claims JSONL/metadata reduction, not token
|
|
183
|
+
savings.
|
|
184
|
+
4. **Test/build parsing stays observe-only.** No production compact-result adapter
|
|
185
|
+
is enabled without a proven recovery/lifetime contract for omitted data.
|
|
186
|
+
5. **Mutation/LSP, web/document, JSON, subagent and visual paths remain their R-E
|
|
187
|
+
passthrough/limited decisions.** Browser/MCP direct adapters remain unsupported.
|
|
188
|
+
6. **P02 durable store remains deferred.** R-G found no task requiring new durable
|
|
189
|
+
snapshot identity strongly enough to justify store/readers/quotas now.
|
|
190
|
+
|
|
191
|
+
Rollback is configuration-only for future calls: disable the normalizer module,
|
|
192
|
+
return `repoDiscovery.profile` to `baseline`, and/or set Context Gateway mode to
|
|
193
|
+
`off`. Rollback does not delete session history, native temp files, user data or
|
|
194
|
+
DCP state and does not require a manual sync operation.
|
|
195
|
+
|
|
196
|
+
## Deterministic R-G gate
|
|
197
|
+
|
|
198
|
+
The final gate reuses the independently proven R-A–R-F contracts rather than one
|
|
199
|
+
monolithic test process where unrelated LSP/AgentSession suites can interfere with
|
|
200
|
+
shared process/temp state. The required groups are:
|
|
201
|
+
|
|
202
|
+
- Context Gateway/storeless contracts, Native Compact, recovery validator/corpus,
|
|
203
|
+
eval harness and coverage registry;
|
|
204
|
+
- ACP lazy history/replay and typecheck;
|
|
205
|
+
- root tab ownership/session UI contracts;
|
|
206
|
+
- suite/root `git diff --check` and source typechecks.
|
|
207
|
+
|
|
208
|
+
No new live model call, manual suite sync, durable store, reader tool or DCP change
|
|
209
|
+
is part of the R-G release gate.
|
|
210
|
+
|
|
211
|
+
The final external-suite P01-R gate is:
|
|
212
|
+
|
|
213
|
+
```text
|
|
214
|
+
bun test \
|
|
215
|
+
test/context-gateway \
|
|
216
|
+
test/repo-native-compact.test.ts \
|
|
217
|
+
test/repo-discovery.test.ts \
|
|
218
|
+
test/config.test.ts \
|
|
219
|
+
test/evals/extension-contracts.test.ts \
|
|
220
|
+
test/evals/harness.test.ts \
|
|
221
|
+
test/evals/recovery-corpus.test.ts \
|
|
222
|
+
test/evals/recovery-run-identity.test.ts \
|
|
223
|
+
test/evals/recovery-report.test.ts \
|
|
224
|
+
test/evals/recovery-validation.test.ts
|
|
225
|
+
```
|
|
226
|
+
|
|
227
|
+
Result: **144 pass, 0 fail, 1011 assertions**; suite typecheck and
|
|
228
|
+
`git diff --check` pass. Root SDK pin remains `0.85.1`; generated schemas are
|
|
229
|
+
unchanged when checked through `node --import tsx` (the sandbox blocks the normal
|
|
230
|
+
`tsx` CLI IPC socket), and root `tsc --noEmit` passes.
|
|
231
|
+
|
|
232
|
+
R-F's separate host/UI gates remain green: root TUI/session **98/98**, ACP lazy
|
|
233
|
+
current-result **4/4** plus an additional **6/6** persisted-history/request-boundary
|
|
234
|
+
suite and typecheck, and Desktop transcript/client **28/28** with Desktop check at
|
|
235
|
+
zero errors (two pre-existing accessibility warnings).
|
|
@@ -0,0 +1,109 @@
|
|
|
1
|
+
# Sub-agent model pools
|
|
2
|
+
|
|
3
|
+
Sub-agents primarily reduce the cost of bounded work and keep intermediate
|
|
4
|
+
source, searches and logs out of the parent context. The parent owns planning,
|
|
5
|
+
integration, decisions and the final answer. Actual savings depend on worker
|
|
6
|
+
quality, retries and how much work the parent repeats; the configuration is
|
|
7
|
+
not a price oracle.
|
|
8
|
+
|
|
9
|
+
## Five execution modes
|
|
10
|
+
|
|
11
|
+
- `research`: read-only evidence gathering, searches and independent diff review.
|
|
12
|
+
- `implement`: bounded code, documentation, test and frontend changes.
|
|
13
|
+
- `verify`: run checks and interpret logs, without fixing source or tests.
|
|
14
|
+
- `browser-qa`: isolated browser workflow with assertions and visual artifacts.
|
|
15
|
+
- `oracle`: a deliberate strong second opinion, not automatic worker escalation.
|
|
16
|
+
|
|
17
|
+
Task-specific discipline belongs in the brief or `promptAppend`. A new project
|
|
18
|
+
agent is warranted when it adds a durable contract, capabilities or resources,
|
|
19
|
+
not merely a professional title. `verify` has a behavioral no-edit contract;
|
|
20
|
+
shell access is not a read-only filesystem sandbox.
|
|
21
|
+
|
|
22
|
+
## Agent priority, preset availability
|
|
23
|
+
|
|
24
|
+
Each Markdown profile owns its ordered `models` list. This is one candidate
|
|
25
|
+
chain for initial selection and subsequent quota fallbacks:
|
|
26
|
+
|
|
27
|
+
```yaml
|
|
28
|
+
---
|
|
29
|
+
description: Make bounded implementation changes.
|
|
30
|
+
models:
|
|
31
|
+
- zai/glm-5.3-flash
|
|
32
|
+
- openai-codex/gpt-5.6-terra
|
|
33
|
+
- openai-codex/gpt-5.6-luna
|
|
34
|
+
thinking: medium
|
|
35
|
+
---
|
|
36
|
+
```
|
|
37
|
+
|
|
38
|
+
A preset contains a set of available models, not a per-agent matrix:
|
|
39
|
+
|
|
40
|
+
```jsonc
|
|
41
|
+
{
|
|
42
|
+
"asyncSubagents": {
|
|
43
|
+
"presets": {
|
|
44
|
+
"gpt": {
|
|
45
|
+
"description": "Models available for this session",
|
|
46
|
+
"models": [
|
|
47
|
+
"openai-codex/gpt-5.6-luna",
|
|
48
|
+
"openai-codex/gpt-5.6-terra",
|
|
49
|
+
"openai-codex/gpt-5.6-sol"
|
|
50
|
+
]
|
|
51
|
+
}
|
|
52
|
+
}
|
|
53
|
+
}
|
|
54
|
+
}
|
|
55
|
+
```
|
|
56
|
+
|
|
57
|
+
The example agent selects Terra, not Luna: the agent's order wins. Sol is
|
|
58
|
+
available in the pool but absent from this worker's chain, so it cannot become
|
|
59
|
+
an automatic implementation fallback. The oracle can declare Sol in its own
|
|
60
|
+
chain. Model references must be exact `provider/model` values, not wildcards.
|
|
61
|
+
|
|
62
|
+
The resolver intersects the agent chain with the selected pool. Runtime
|
|
63
|
+
selection then skips unregistered, unauthenticated or session-exhausted models.
|
|
64
|
+
Tasks with images and browser QA require confirmed image support. The first
|
|
65
|
+
eligible candidate runs; only the remaining eligible candidates are passed to
|
|
66
|
+
quota fallback. An empty intersection or unavailable chain rejects the batch
|
|
67
|
+
before any children or run state are created. Model selection makes no LLM
|
|
68
|
+
completion request; the optional role router is a separate operation.
|
|
69
|
+
|
|
70
|
+
Oracle prefers another provider when possible, but still respects the pool.
|
|
71
|
+
It never substitutes an ordinary cheap candidate merely to avoid a selection
|
|
72
|
+
error. A single-provider pool cannot promise cross-provider independence.
|
|
73
|
+
|
|
74
|
+
Explicit task `model`, CLI `--model`, and `FORCE_CURRENT_MODEL` remain deliberate
|
|
75
|
+
overrides: they bypass the pool and do not add automatic fallback candidates.
|
|
76
|
+
The parent should not use these to evade the configured budget. The pool is
|
|
77
|
+
a selection policy, not a security boundary against explicit overrides.
|
|
78
|
+
|
|
79
|
+
## Selection and compatibility
|
|
80
|
+
|
|
81
|
+
Use `/subagent-preset <name>`, `AGENTS_PRESET=<name>` or
|
|
82
|
+
`/subagent-preset session <name>`. Clearing the preset uses agent priorities
|
|
83
|
+
without a pool filter. The shipped names remain compatible with saved choices:
|
|
84
|
+
`cheap` is the GLM pool, `gpt` the GPT pool, and `deep` the mixed pool. The last
|
|
85
|
+
name no longer means that ordinary workers should escalate to flagship models.
|
|
86
|
+
|
|
87
|
+
Old role names are not implicit aliases. `quick`, `scan`, `review`, `deep`,
|
|
88
|
+
`docs`, `frontend`, and `tests` work only when explicitly defined as ordinary
|
|
89
|
+
custom/project types. This keeps the effective catalog and accepted names exact.
|
|
90
|
+
|
|
91
|
+
Legacy `model` plus `fallbackModels` and `modelByParent` still load. New profile
|
|
92
|
+
`models` replaces inherited legacy selection fields; an explicit legacy model
|
|
93
|
+
override can still replace an inherited new list. Empty `models` means no
|
|
94
|
+
candidates, not permission to inherit the parent model. Model-less project
|
|
95
|
+
specialists must declare candidates or receive an explicit model override.
|
|
96
|
+
|
|
97
|
+
Legacy preset role matrices remain readable. A preset with `models` uses only
|
|
98
|
+
the pool contract, dropping stale legacy model/thinking/type overrides. Switching
|
|
99
|
+
a higher-priority config layer back to a legacy preset removes the inherited
|
|
100
|
+
pool. Configuration loading never rewrites user files; review old overrides
|
|
101
|
+
when migrating, since explicitly saved profiles can retain expensive models.
|
|
102
|
+
|
|
103
|
+
## Compact handoff
|
|
104
|
+
|
|
105
|
+
Give workers a scope, acceptance criteria and the evidence needed to start.
|
|
106
|
+
Read compact results first and inspect raw artifacts selectively. One noisy
|
|
107
|
+
sequential investigation can justify a worker; a command whose exit status is
|
|
108
|
+
sufficient usually only needs a saved log, not another LLM. Independent review
|
|
109
|
+
uses a fresh `research` invocation, not a separate built-in persona.
|