pi-ui-extend 1.0.39 → 1.0.41
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +1 -1
- package/dist/app/commands/command-registry.js +2 -2
- package/dist/app/commands/command-session-actions.d.ts +0 -1
- package/dist/app/commands/command-session-actions.js +22 -13
- package/dist/app/icons.d.ts +14 -0
- package/dist/app/icons.js +33 -0
- package/dist/app/rendering/conversation-tool-renderer.js +2 -2
- package/dist/app/rendering/dcp-stats.d.ts +6 -1
- package/dist/app/rendering/dcp-stats.js +214 -46
- package/dist/app/rendering/editor-panels.js +8 -5
- package/dist/app/session/lazy-session-manager.js +12 -1
- package/dist/app/session/tabs-controller.d.ts +2 -5
- package/dist/app/session/tabs-controller.js +12 -21
- package/dist/app/subagents/subagents-model.d.ts +14 -1
- package/dist/app/subagents/subagents-model.js +34 -15
- package/dist/app/types.d.ts +2 -0
- package/dist/bundled-extensions/session-title/config.js +1 -1
- package/dist/markdown-format.js +27 -9
- package/dist/schemas/pi-tools-suite-schema.d.ts +29 -16
- package/dist/schemas/pi-tools-suite-schema.js +46 -31
- package/external/pi-tools-suite/README.md +392 -52
- package/external/pi-tools-suite/docs/browser-qa-subagent.md +31 -21
- package/external/pi-tools-suite/docs/context-gateway-p00-adr.md +216 -0
- package/external/pi-tools-suite/docs/context-gateway-p01n-gate-review.md +122 -0
- package/external/pi-tools-suite/docs/context-gateway-p01n-measurement.md +133 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-ra-evidence.md +111 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-rb-evidence.md +100 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-rc-evidence.md +69 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-rd-evidence.md +100 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-re-evidence.md +74 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-rf-evidence.md +153 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-rg-evidence.md +235 -0
- package/external/pi-tools-suite/docs/evals.md +684 -0
- package/external/pi-tools-suite/docs/subagent-model-pools.md +109 -0
- package/external/pi-tools-suite/package.json +10 -3
- package/external/pi-tools-suite/src/async-subagents/{private-skills → agents}/browser-qa/scripts/browser-qa-runner.mjs +82 -1
- package/external/pi-tools-suite/src/async-subagents/{private-skills/browser-qa/SKILL.md → agents/browser-qa.md} +261 -12
- package/external/pi-tools-suite/src/async-subagents/agents/implement.md +20 -0
- package/external/pi-tools-suite/src/async-subagents/agents/oracle.md +16 -0
- package/external/pi-tools-suite/src/async-subagents/agents/research.md +18 -0
- package/external/pi-tools-suite/src/async-subagents/agents/verify.md +18 -0
- package/external/pi-tools-suite/src/async-subagents/async-subagents.sample.jsonc +27 -243
- package/external/pi-tools-suite/src/async-subagents/commands.ts +6 -2
- package/external/pi-tools-suite/src/async-subagents/core/agent-catalog.ts +41 -0
- package/external/pi-tools-suite/src/async-subagents/core/agent-strategy.ts +13 -93
- package/external/pi-tools-suite/src/async-subagents/core/agents-dir.ts +494 -0
- package/external/pi-tools-suite/src/async-subagents/core/browser-qa.ts +9 -0
- package/external/pi-tools-suite/src/async-subagents/core/config.ts +200 -143
- package/external/pi-tools-suite/src/async-subagents/core/model-fallback.ts +1 -1
- package/external/pi-tools-suite/src/async-subagents/core/model-selection.ts +54 -0
- package/external/pi-tools-suite/src/async-subagents/core/prompt.ts +7 -6
- package/external/pi-tools-suite/src/async-subagents/core/routing.ts +52 -45
- package/external/pi-tools-suite/src/async-subagents/core/spawn.ts +12 -4
- package/external/pi-tools-suite/src/async-subagents/index.ts +11 -1
- package/external/pi-tools-suite/src/async-subagents/lib.ts +6 -2
- package/external/pi-tools-suite/src/async-subagents/tools/spawn.ts +46 -18
- package/external/pi-tools-suite/src/async-subagents/tools/subagents.ts +3 -2
- package/external/pi-tools-suite/src/async-subagents/types.ts +2 -0
- package/external/pi-tools-suite/src/coding-discipline/index.ts +41 -142
- package/external/pi-tools-suite/src/config.ts +1 -22
- package/external/pi-tools-suite/src/context-gateway/accounting.ts +151 -0
- package/external/pi-tools-suite/src/context-gateway/config.ts +111 -0
- package/external/pi-tools-suite/src/context-gateway/index.ts +160 -0
- package/external/pi-tools-suite/src/context-gateway/metadata-normalization.ts +88 -0
- package/external/pi-tools-suite/src/context-gateway/storeless-capabilities.ts +89 -0
- package/external/pi-tools-suite/src/context-gateway/telemetry.ts +429 -0
- package/external/pi-tools-suite/src/context-gateway/test-output-parser.ts +326 -0
- package/external/pi-tools-suite/src/context-gateway/types.ts +152 -0
- package/external/pi-tools-suite/src/dcp/auto-compress-budget.ts +106 -0
- package/external/pi-tools-suite/src/dcp/auto-compress.ts +810 -106
- package/external/pi-tools-suite/src/dcp/commands.ts +64 -139
- package/external/pi-tools-suite/src/dcp/compress-tool.ts +369 -35
- package/external/pi-tools-suite/src/dcp/compression-blocks.ts +510 -64
- package/external/pi-tools-suite/src/dcp/compression-preview.ts +113 -0
- package/external/pi-tools-suite/src/dcp/compression-progress.ts +70 -0
- package/external/pi-tools-suite/src/dcp/config.ts +36 -61
- package/external/pi-tools-suite/src/dcp/conversation-index.ts +421 -0
- package/external/pi-tools-suite/src/dcp/debug-log.ts +7 -5
- package/external/pi-tools-suite/src/dcp/index.ts +617 -203
- package/external/pi-tools-suite/src/dcp/journal.ts +566 -0
- package/external/pi-tools-suite/src/dcp/progress-controller.ts +244 -0
- package/external/pi-tools-suite/src/dcp/prompts.ts +10 -7
- package/external/pi-tools-suite/src/dcp/provider-tool-results.ts +189 -0
- package/external/pi-tools-suite/src/dcp/pruner-candidates.ts +298 -78
- package/external/pi-tools-suite/src/dcp/pruner-compression-blocks.ts +173 -281
- package/external/pi-tools-suite/src/dcp/pruner-emergency.ts +2 -4
- package/external/pi-tools-suite/src/dcp/pruner-message-ids.ts +17 -5
- package/external/pi-tools-suite/src/dcp/pruner-metadata.ts +11 -1
- package/external/pi-tools-suite/src/dcp/pruner-nudge.ts +30 -82
- package/external/pi-tools-suite/src/dcp/pruner-tools.ts +22 -133
- package/external/pi-tools-suite/src/dcp/pruner.ts +18 -33
- package/external/pi-tools-suite/src/dcp/recovery.ts +129 -0
- package/external/pi-tools-suite/src/dcp/shadow-plan.ts +127 -0
- package/external/pi-tools-suite/src/dcp/state-transaction.ts +102 -0
- package/external/pi-tools-suite/src/dcp/state.ts +158 -580
- package/external/pi-tools-suite/src/dcp/ui.ts +1 -0
- package/external/pi-tools-suite/src/default-pi-tools-suite-config.ts +55 -220
- package/external/pi-tools-suite/src/index.ts +9 -0
- package/external/pi-tools-suite/src/model-tools/index.ts +76 -42
- package/external/pi-tools-suite/src/repo-discovery/index.ts +84 -18
- package/external/pi-tools-suite/src/repo-discovery/native-compact.ts +458 -0
- package/external/pi-tools-suite/src/session-recovery/index.ts +189 -43
- package/external/pi-tools-suite/src/tool-descriptions.ts +43 -38
- package/external/pi-tools-suite/src/truncation-metadata-normalizer/index.ts +17 -0
- package/package.json +6 -6
- package/schemas/pi-tools-suite.json +159 -78
- package/external/pi-tools-suite/src/async-subagents/private-skills/browser-qa/references/auth-scaffold-spec.md +0 -78
- package/external/pi-tools-suite/src/async-subagents/private-skills/browser-qa/references/qa-design.md +0 -223
- package/external/pi-tools-suite/src/dcp/state-persistence.ts +0 -195
- /package/external/pi-tools-suite/src/async-subagents/{private-skills/browser-qa/references → agents/browser-qa/examples}/qa-auth.example.jsonc +0 -0
- /package/external/pi-tools-suite/src/async-subagents/{private-skills/browser-qa/references → agents/browser-qa/examples}/qa-flow.example.jsonc +0 -0
- /package/external/pi-tools-suite/src/async-subagents/{private-skills → agents}/browser-qa/vendor/fflate.LICENSE +0 -0
- /package/external/pi-tools-suite/src/async-subagents/{private-skills → agents}/browser-qa/vendor/fflate.mjs +0 -0
|
@@ -0,0 +1,100 @@
|
|
|
1
|
+
# Context Gateway P01-R / R-B evidence
|
|
2
|
+
|
|
3
|
+
<!-- markdownlint-disable MD013 -->
|
|
4
|
+
|
|
5
|
+
> Date: 7 September 2026.
|
|
6
|
+
> Repository HEAD during deterministic gate: `daa1b06` with an explicitly dirty tested source tree.
|
|
7
|
+
> Installed Pi SDK: `@earendil-works/pi-coding-agent` `0.85.1`.
|
|
8
|
+
> Scope: storeless result-pipeline ordering and optional truncation-metadata cleanup. No durable store, Gateway enforce mode, DCP redesign or live model run was introduced here.
|
|
9
|
+
|
|
10
|
+
## Result
|
|
11
|
+
|
|
12
|
+
R-B is complete for the currently claimed storeless combinations. The optional
|
|
13
|
+
`truncation-metadata-normalizer` remains disabled by default and returns only a
|
|
14
|
+
`details` patch. Context Gateway `off`/`observe` semantics remain unchanged.
|
|
15
|
+
|
|
16
|
+
The verified result order is:
|
|
17
|
+
|
|
18
|
+
1. LSP/comment result enrichment;
|
|
19
|
+
2. passive Context Gateway observe;
|
|
20
|
+
3. optional truncation metadata normalization;
|
|
21
|
+
4. downstream result observers/modifiers;
|
|
22
|
+
5. opt-in credential-firewall session-hygiene redaction.
|
|
23
|
+
|
|
24
|
+
For provider hooks the credential firewall runs before the final
|
|
25
|
+
`codex-reasoning-fix` sanitizer, which remains last in `MODULES`.
|
|
26
|
+
|
|
27
|
+
No suite-local result coordinator was added. The ADR now treats a coordinator as
|
|
28
|
+
conditional future enforce/store work only if an actual enabled-handler conflict
|
|
29
|
+
cannot be expressed safely through the tested event-specific order.
|
|
30
|
+
|
|
31
|
+
## Normalizer scope and invariants
|
|
32
|
+
|
|
33
|
+
The normalizer is restricted to the measured `Read`/shell/`ast_grep` names and a
|
|
34
|
+
complete SDK truncation shape. It removes only `details.truncation.content` when
|
|
35
|
+
that string is a prefix of the actually delivered text. It preserves the rest of
|
|
36
|
+
the truncation fields and any native `fullOutputPath`.
|
|
37
|
+
|
|
38
|
+
Deterministic contracts cover:
|
|
39
|
+
|
|
40
|
+
- actual installed SDK `Read` and `Bash` truncation results;
|
|
41
|
+
- the real suite `ast_grep` truncation result and its native full-output handle;
|
|
42
|
+
- `bash` / `shell` / `shell_command` aliases;
|
|
43
|
+
- Unicode duplicate text and multipart text + image content;
|
|
44
|
+
- idempotence and no mutation of the original content/details objects;
|
|
45
|
+
- no-op for unknown tool names, malformed shapes and non-matching metadata;
|
|
46
|
+
- collapsed/expanded rendering equivalence for installed Bash, installed Read and suite `ast_grep` renderers;
|
|
47
|
+
- persisted JSONL removal of the duplicate metadata copy while structural fields remain;
|
|
48
|
+
- an exactly equal checked OpenAI-completions provider payload before/after metadata normalization.
|
|
49
|
+
|
|
50
|
+
The generic `tool_result` event cannot prove that a separately loaded replacement
|
|
51
|
+
using the same measured tool name and exact SDK-looking shape is the original SDK
|
|
52
|
+
definition. The ADR therefore marks same-name replacement provenance as limited;
|
|
53
|
+
name+shape is not promoted to a strict-enforce trust primitive.
|
|
54
|
+
|
|
55
|
+
## Result/security composition
|
|
56
|
+
|
|
57
|
+
An executable installed `ExtensionRunner` contract uses an enrichment stage,
|
|
58
|
+
Context Gateway observe, the normalizer, a DCP-free fake downstream observer and
|
|
59
|
+
the real credential firewall. It proves that:
|
|
60
|
+
|
|
61
|
+
- tool call ID observed downstream is unchanged;
|
|
62
|
+
- `content`, image parts, `isError`, usage and structural completeness metadata survive normalization;
|
|
63
|
+
- enrichment diagnostics survive until the later firewall;
|
|
64
|
+
- the downstream observer sees normalized metadata before firewall redaction;
|
|
65
|
+
- with session hygiene enabled, synthetic secrets are removed from final result content/details without restoring the deleted duplicate;
|
|
66
|
+
- with session hygiene disabled, visible content is not redacted while metadata-only normalization still occurs;
|
|
67
|
+
- the input result object itself is not mutated.
|
|
68
|
+
|
|
69
|
+
A headless AgentSession contract separately proves that the normalized/redacted
|
|
70
|
+
result, rather than the earlier secret-bearing result, is what reaches JSONL and
|
|
71
|
+
the next model context. The provider-hook contract proves that credential
|
|
72
|
+
redaction followed by the final Codex sanitizer neither restores the secret nor
|
|
73
|
+
restores rejected reasoning/prompt-cache fields.
|
|
74
|
+
|
|
75
|
+
Tool-result `details` bytes are therefore reported as JSONL/metadata overhead,
|
|
76
|
+
not as provider-token savings. The checked OpenAI-completions serializer already
|
|
77
|
+
omits tool-result details.
|
|
78
|
+
|
|
79
|
+
## Deterministic gate
|
|
80
|
+
|
|
81
|
+
```text
|
|
82
|
+
bun test test/context-gateway \
|
|
83
|
+
test/evals/recovery-corpus.test.ts \
|
|
84
|
+
test/evals/recovery-run-identity.test.ts \
|
|
85
|
+
test/evals/recovery-report.test.ts \
|
|
86
|
+
test/evals/recovery-validation.test.ts \
|
|
87
|
+
test/evals/harness.test.ts \
|
|
88
|
+
test/config.test.ts \
|
|
89
|
+
test/evals/extension-contracts.test.ts
|
|
90
|
+
```
|
|
91
|
+
|
|
92
|
+
Result: **88 pass, 0 fail, 667 assertions**.
|
|
93
|
+
|
|
94
|
+
Additional checks:
|
|
95
|
+
|
|
96
|
+
- `npm run typecheck` — pass.
|
|
97
|
+
- `git diff --check` — pass.
|
|
98
|
+
|
|
99
|
+
No manual suite sync was run; the user-owned watcher remains the synchronization
|
|
100
|
+
mechanism and is not evidence of which bytes a future live run loads.
|
|
@@ -0,0 +1,69 @@
|
|
|
1
|
+
# Context Gateway P01-R / R-C evidence
|
|
2
|
+
|
|
3
|
+
<!-- markdownlint-disable MD013 -->
|
|
4
|
+
|
|
5
|
+
> Date: 7 September 2026.
|
|
6
|
+
> Repository HEAD during deterministic gate: `daa1b06` with a dirty tested tree.
|
|
7
|
+
> Installed Pi SDK: `@earendil-works/pi-coding-agent` `0.85.1`.
|
|
8
|
+
> Scope: native paging/current-view semantics and temporary full-output lifetime. No durable snapshot, artifact reader, shell rewrite or live model call was added.
|
|
9
|
+
|
|
10
|
+
## Read is a current decoded text view, not a snapshot reader
|
|
11
|
+
|
|
12
|
+
Installed SDK `Read` uses 1-based line `offset`/`limit` over `Buffer.toString("utf-8").split("\n")`.
|
|
13
|
+
R-C contracts verify:
|
|
14
|
+
|
|
15
|
+
- CRLF and Unicode survive relative to that decoded SDK text view;
|
|
16
|
+
- explicit `limit` produces an actionable next line offset;
|
|
17
|
+
- empty files and EOF/out-of-range offsets are distinguishable;
|
|
18
|
+
- a very large limit does not create a separate byte pagination protocol;
|
|
19
|
+
- a single line larger than the SDK byte limit is an honest **limited** case: the result points at a shell fallback and does not fabricate `Use offset=` recovery;
|
|
20
|
+
- the Read schema exposes no `byteOffset` or generic cursor;
|
|
21
|
+
- re-executing the same path after the source changes returns the new file view. Same path/offset is therefore not an immutable historical snapshot.
|
|
22
|
+
|
|
23
|
+
This project does not claim byte-for-byte source-file coordinates beyond the SDK's decoded text semantics.
|
|
24
|
+
|
|
25
|
+
## Repo Native Compact cursors are current-index cursors
|
|
26
|
+
|
|
27
|
+
The checked native cursor surface exists only on `repo_structure` and `repo_ast`.
|
|
28
|
+
`repo_search` has no cursor flag in the current wrapper policy and is not described as if it did.
|
|
29
|
+
|
|
30
|
+
R-C tightens integer native flags to `Number.isSafeInteger`: negative or unsafe integer cursors are refused before `idx`. Re-running the same `repo_structure --cursor 20` executes `idx` again; a deterministic fake backend returning version 1 and then version 2 produces two different results for the same cursor. The cursor is therefore a continuation token for current backend/index state, not a snapshot identifier.
|
|
31
|
+
|
|
32
|
+
No command execution is suppressed because args/cursor match an earlier call.
|
|
33
|
+
|
|
34
|
+
## Successful temp output versus error paths
|
|
35
|
+
|
|
36
|
+
For a successful large built-in Bash result, the installed SDK returns a structured `details.fullOutputPath`. The file contains the complete emitted output and a normal Read can recover an omitted head fact.
|
|
37
|
+
|
|
38
|
+
The same path has no durability guarantee. The test replaces its contents and a later Read sees the replacement; after deletion a later Read fails. A native temp path is thus an ephemeral current file handle, not historical snapshot identity.
|
|
39
|
+
|
|
40
|
+
Bash timeout, abort and non-zero exit preserve their visible status and captured visible output but reject the tool execution. The current exception bridge does not expose a structured `fullOutputPath` result. SDK-formatted error text may itself mention its generated temp path when the partial output was truncated; R-C does **not** parse such text into a trusted capability, because look-alike paths in arbitrary text are not authorization/provenance.
|
|
41
|
+
|
|
42
|
+
The suite `ast_grep` path behaves similarly but is suite-owned: successful truncated output returns a structured full-output handle containing the complete combined output. Cancelled/killed output returns an explicit cancelled result without a full-output handle, and a real tool error throws instead of publishing a successful artifact capability.
|
|
43
|
+
|
|
44
|
+
## Broad-output decision
|
|
45
|
+
|
|
46
|
+
R-C does not introduce another generic cap for built-in Read or rewrite user shell commands. Current evidence already shows successful native recovery for ordinary line-based Read and successful Bash/ast-grep temp-output cases, while long single-line Read and error-path handles remain explicitly limited.
|
|
47
|
+
|
|
48
|
+
For repo tools the existing opt-in Native Compact profile remains the only new scope/budget enforcement: narrow native defaults, explicit same-call bounded `outputMode=full`, and current-index cursors where the backend exposes them. Search without a native cursor is not given a synthetic one.
|
|
49
|
+
|
|
50
|
+
A stricter broad-read/search delivery policy requires a demonstrated task where its required facts remain recoverable under the claimed mechanism. Until then small/exact/instruction reads and shell execution stay unchanged.
|
|
51
|
+
|
|
52
|
+
## Deterministic gate
|
|
53
|
+
|
|
54
|
+
```text
|
|
55
|
+
bun test \
|
|
56
|
+
test/context-gateway/native-recovery-contracts.test.ts \
|
|
57
|
+
test/repo-native-compact.test.ts \
|
|
58
|
+
test/context-gateway/capture-contracts.test.ts \
|
|
59
|
+
test/context-gateway/sdk-pipeline.test.ts
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
Result: **46 pass, 0 fail, 352 assertions**.
|
|
63
|
+
|
|
64
|
+
Additional checks:
|
|
65
|
+
|
|
66
|
+
- `npm run typecheck` — pass.
|
|
67
|
+
- `git diff --check` — pass.
|
|
68
|
+
|
|
69
|
+
Session-history recovery remains plan 32 work. R-C creates neither `artifact_read` nor a catalog and makes no resume/export guarantee for native temp files or current-index cursors.
|
|
@@ -0,0 +1,100 @@
|
|
|
1
|
+
# Context Gateway P01-R / R-D evidence
|
|
2
|
+
|
|
3
|
+
<!-- markdownlint-disable MD013 -->
|
|
4
|
+
|
|
5
|
+
> Date: 7 September 2026.
|
|
6
|
+
> Repository HEAD during deterministic gate: `daa1b06` with a dirty tested tree.
|
|
7
|
+
> Scope: pure test/build parsing plus passive observe classification. Production tool-result delivery remains passthrough; no archive, reference, fetch, execution or LLM summary is introduced.
|
|
8
|
+
|
|
9
|
+
## Supported parser scope
|
|
10
|
+
|
|
11
|
+
R-D adds `src/context-gateway/test-output-parser.ts` as a pure in-memory parser for a deliberately small format set:
|
|
12
|
+
|
|
13
|
+
- Bun test terminal summaries (`pass`, `fail`, `Ran ... tests across ... files`) and exact failure/warning lines;
|
|
14
|
+
- TAP v13-style terminal plan/count summaries and exact `not ok` / diagnostic lines;
|
|
15
|
+
- bounded TypeScript compiler diagnostics plus an explicit terminal `Found N error(s)` summary.
|
|
16
|
+
|
|
17
|
+
The parser returns `recognised`, `partial`, or `unrecognised`. Host execution outcome remains authoritative; text does not turn an errored shell call into success. A failed test summary without an exact parsed diagnostic is not considered complete.
|
|
18
|
+
|
|
19
|
+
The parser's ANSI and bare-CR processing creates only an internal parsing view. It does not mutate the original tool result or claim source byte coordinates.
|
|
20
|
+
|
|
21
|
+
## Conservative failure rules
|
|
22
|
+
|
|
23
|
+
The parser/delivery planner intentionally falls back to passthrough when any of these conditions holds:
|
|
24
|
+
|
|
25
|
+
- SDK/upstream truncation is already reported;
|
|
26
|
+
- timeout/abort is visible;
|
|
27
|
+
- terminal summary conflicts with the host outcome (for example PASS summary followed by command exit 1);
|
|
28
|
+
- multiple recognised formats appear in one output;
|
|
29
|
+
- TypeScript pretty/source/caret lines remain unparsed;
|
|
30
|
+
- a failed summary has no captured exact failure diagnostic;
|
|
31
|
+
- the terminal summary is followed by significant unrecognised output;
|
|
32
|
+
- parser input exceeds the bounded scan budget (default 1 Mi-character); head/tail resemblance is not promoted to completeness;
|
|
33
|
+
- the originating shell command is compound or its scope is unknown;
|
|
34
|
+
- a prospective complete compact representation would exceed the configured byte budget.
|
|
35
|
+
|
|
36
|
+
Compound-command detection is transient and conservative. Observe stores only `simple / compound / unknown`, never the shell command text. False positives merely preserve passthrough; they never rewrite or block execution.
|
|
37
|
+
|
|
38
|
+
## Prospective compact delivery is all-or-passthrough
|
|
39
|
+
|
|
40
|
+
`planProspectiveTestOutputDelivery` is a decision helper, not a production shaper. For a complete recognised **simple** command it can build a candidate containing:
|
|
41
|
+
|
|
42
|
+
1. exact host outcome;
|
|
43
|
+
2. recognised format;
|
|
44
|
+
3. terminal summary counts;
|
|
45
|
+
4. every parsed mandatory error/warning line.
|
|
46
|
+
|
|
47
|
+
The candidate is accepted only when the entire representation fits the budget. Diagnostics are never sliced merely to hit a target. `partial` and `unrecognised` inputs have no generic head/tail summary path.
|
|
48
|
+
|
|
49
|
+
No runtime adapter consumes this candidate in R-D. The current Context Gateway `tool_result` handler still returns `undefined`, so delivered content stays byte-equivalent. Connecting a compact delivery requires a later explicit adapter decision with a recovery/lifetime mechanism appropriate to the omitted data. Native truncation alone is not that permission.
|
|
50
|
+
|
|
51
|
+
## Observe telemetry
|
|
52
|
+
|
|
53
|
+
Shell output is parsed transiently only while Context Gateway is in `observe` mode. The telemetry snapshot stores:
|
|
54
|
+
|
|
55
|
+
- parser version;
|
|
56
|
+
- classification and format counts;
|
|
57
|
+
- scan-limited count;
|
|
58
|
+
- terminal-summary / diagnostic / warning counts for the last observation;
|
|
59
|
+
- safe command-scope enum;
|
|
60
|
+
- prospective `compact-candidate / passthrough` decision and allowlisted reason.
|
|
61
|
+
|
|
62
|
+
It does **not** store command args, diagnostic strings, filenames from diagnostics, raw body, snippets, or hidden-fact fingerprints. Tests compare the input event before/after telemetry and prove it is not mutated.
|
|
63
|
+
|
|
64
|
+
## Corpus cases
|
|
65
|
+
|
|
66
|
+
The deterministic corpus covers:
|
|
67
|
+
|
|
68
|
+
- complete Bun success;
|
|
69
|
+
- a Bun error in the middle of output plus a warning and later passing test;
|
|
70
|
+
- PASS summary followed by host exit 1;
|
|
71
|
+
- timeout;
|
|
72
|
+
- upstream-truncated output;
|
|
73
|
+
- ANSI and bare-CR overwrite;
|
|
74
|
+
- TAP failure with exact diagnostics;
|
|
75
|
+
- compact TypeScript diagnostics and a pretty/unparsed partial case;
|
|
76
|
+
- mixed/nested recognised formats;
|
|
77
|
+
- unknown custom build output;
|
|
78
|
+
- compound and unknown command scope;
|
|
79
|
+
- an intentionally huge output exceeding the parser scan bound;
|
|
80
|
+
- prospective output that fits the byte budget and the same mandatory diagnostics under a too-small budget.
|
|
81
|
+
|
|
82
|
+
These are format contracts, not claims that every Bun/TAP/TypeScript version is supported. Unknown variants remain passthrough.
|
|
83
|
+
|
|
84
|
+
## Deterministic gate
|
|
85
|
+
|
|
86
|
+
```text
|
|
87
|
+
bun test \
|
|
88
|
+
test/context-gateway/test-output-parser.test.ts \
|
|
89
|
+
test/context-gateway/observe.test.ts \
|
|
90
|
+
test/evals/harness.test.ts
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
Result: **28 pass, 0 fail, 152 assertions**.
|
|
94
|
+
|
|
95
|
+
Additional checks:
|
|
96
|
+
|
|
97
|
+
- `npm run typecheck` — pass.
|
|
98
|
+
- `git diff --check` — pass.
|
|
99
|
+
|
|
100
|
+
No live model call, manual sync, user config change, result replacement or new full-output archive was performed for R-D.
|
|
@@ -0,0 +1,74 @@
|
|
|
1
|
+
# Context Gateway P01-R / R-E evidence
|
|
2
|
+
|
|
3
|
+
<!-- markdownlint-disable MD013 -->
|
|
4
|
+
|
|
5
|
+
> Date: 7 September 2026.
|
|
6
|
+
> Repository HEAD during deterministic gate: `daa1b06` with a dirty tested tree.
|
|
7
|
+
> Scope: mutation/LSP outcome integrity and storeless capability inventory for web/docs, structured JSON, subagents and visual results. No archival adapter, fetcher, grant or lifetime extension was added.
|
|
8
|
+
|
|
9
|
+
## Mutation and LSP
|
|
10
|
+
|
|
11
|
+
R-E keeps mutations as native passthrough. An actual `applyPatch` run in a temporary fixture proves that its returned `changedFiles`/summary contain only the file changed by that invocation while an unrelated dirty file remains untouched and absent from the result. `getEventPaths` prefers `details.changedFiles` over broader input/patch-shaped paths, so downstream LSP refresh does not derive a whole-workspace diff.
|
|
12
|
+
|
|
13
|
+
The pre-approved patch string is not rewritten by Context Gateway. Errored/cancelled mutation results are no-ops for LSP enrichment and storeless metadata cleanup, preserving their original result object/outcome. Existing LSP integration tests continue to cover successful post-edit diagnostics, changed-file deduplication, deleted-file behavior and aliases. R-B SDK-chain tests separately prove that enrichment/protocol fields survive later storeless cleanup/security stages.
|
|
14
|
+
|
|
15
|
+
No new “partial mutation” inference is introduced. If a producer reports partial/error/cancelled state, Gateway preserves that producer outcome rather than constructing a synthetic success or retry.
|
|
16
|
+
|
|
17
|
+
## Capability matrix
|
|
18
|
+
|
|
19
|
+
`src/context-gateway/storeless-capabilities.ts` records the current non-archival decisions. It is metadata only; importing it registers no tools/readers/adapters.
|
|
20
|
+
|
|
21
|
+
| Surface | Status | Storeless strategy | Lifetime claim |
|
|
22
|
+
| --- | --- | --- | --- |
|
|
23
|
+
| test/build | limited | pure-parser candidate in observe only | current result |
|
|
24
|
+
| mutation/LSP | supported | native passthrough | current result |
|
|
25
|
+
| web/document | limited | native passthrough | current result |
|
|
26
|
+
| structured JSON | limited | native passthrough | current result |
|
|
27
|
+
| subagent result | limited | native passthrough | producer-managed |
|
|
28
|
+
| visual/image | supported | native passthrough | visual message |
|
|
29
|
+
| direct browser | unsupported | none | none |
|
|
30
|
+
| direct MCP | unsupported | none | none |
|
|
31
|
+
|
|
32
|
+
The matrix is a claim boundary, not feature discovery by name. Unsupported paths do not become supported merely because an output resembles JSON/HTML or contains a filesystem path.
|
|
33
|
+
|
|
34
|
+
## Web/document path
|
|
35
|
+
|
|
36
|
+
Existing web tool tests cover actual producer-owned HTTP behavior: request metadata, content/link metadata, Ollama→Tavily fallback, auth/API errors, invalid JSON, cancellation and timeout. Storeless Gateway adds no fetch callback and does not reinterpret headings/code or producer pagination/truncation metadata. A bounded R-E fixture confirms that `web_fetch` content remains byte-equivalent through observation.
|
|
37
|
+
|
|
38
|
+
There is no storeless historical page reader. Producer truncation/pagination remains producer-owned, which is why web/document is `limited` rather than an archive-capable adapter.
|
|
39
|
+
|
|
40
|
+
## Structured JSON
|
|
41
|
+
|
|
42
|
+
No generic JSON parser/field selector is registered. A fixture containing the exact literal `9007199254740993123456789` remains opaque text and is not parsed into a JavaScript number, rounded, ranked or field-selected. Upstream pagination metadata is left on the producer result. A dedicated typed adapter requires a future concrete use case and precision contract.
|
|
43
|
+
|
|
44
|
+
## Subagents
|
|
45
|
+
|
|
46
|
+
The existing async-subagent result tool already emits a bounded summary plus paths to producer artifacts and explicitly avoids inlining raw logs. Its tests cover missing/running/completed states, structured result generation, artifact paths, cleanup and session lifecycle.
|
|
47
|
+
|
|
48
|
+
P01-R does not read internal subagent history, does not turn an artifact path into an authorization grant and does not extend artifact lifetime. A completed path therefore remains `producer-managed`; it is not promised to survive producer cleanup/resume/export.
|
|
49
|
+
|
|
50
|
+
## Visual and unsupported direct paths
|
|
51
|
+
|
|
52
|
+
Image parts remain image parts and are accounted separately from text. No textual placeholder replaces them as a context optimization.
|
|
53
|
+
|
|
54
|
+
P00 did not prove a direct parent Context Gateway boundary for browser DOM/network/console or a current direct MCP execution adapter. Both remain explicitly `unsupported`; R-E does not add placeholder adapters to make the matrix look complete.
|
|
55
|
+
|
|
56
|
+
## Deterministic gate
|
|
57
|
+
|
|
58
|
+
```text
|
|
59
|
+
bun test \
|
|
60
|
+
test/context-gateway/storeless-format-contracts.test.ts \
|
|
61
|
+
test/lsp.test.ts \
|
|
62
|
+
test/web-search.test.ts \
|
|
63
|
+
test/async-subagents/tools.test.ts \
|
|
64
|
+
test/context-gateway/sdk-pipeline.test.ts
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
Result: **107 pass, 0 fail, 689 assertions**.
|
|
68
|
+
|
|
69
|
+
Additional checks:
|
|
70
|
+
|
|
71
|
+
- `npm run typecheck` — pass.
|
|
72
|
+
- `git diff --check` — pass.
|
|
73
|
+
|
|
74
|
+
No live model call, manual sync, user-config change or result-shaping adapter was added for R-E.
|
|
@@ -0,0 +1,153 @@
|
|
|
1
|
+
# Context Gateway P01-R / R-F evidence
|
|
2
|
+
|
|
3
|
+
<!-- markdownlint-disable MD013 -->
|
|
4
|
+
|
|
5
|
+
> Date: 7 September 2026.
|
|
6
|
+
> Repository HEAD during deterministic gate: `daa1b06` with a dirty tested tree.
|
|
7
|
+
> Scope: storeless lifecycle and current-result UI/replay only. No new persistent artifact service, native-temp reader endpoint, fork grant, export bundle, DCP lifecycle implementation or manual sync was added.
|
|
8
|
+
|
|
9
|
+
## Runtime / transient lifecycle
|
|
10
|
+
|
|
11
|
+
Context Gateway observe state is runtime-local. The current implementation tracks
|
|
12
|
+
in-flight tool call IDs separately from aggregate telemetry and refuses an
|
|
13
|
+
`off ↔ observe` mode change while any tool call is still in flight. At a safe
|
|
14
|
+
boundary the mode change clears transient call bindings. `session_start`,
|
|
15
|
+
`session_tree` and `session_shutdown` reset transient and aggregate observation
|
|
16
|
+
state.
|
|
17
|
+
|
|
18
|
+
Deterministic contracts additionally prove:
|
|
19
|
+
|
|
20
|
+
- a mode change requested during an observed call is rejected until that result
|
|
21
|
+
completes; the result is attributed to the old effective observe epoch once;
|
|
22
|
+
- lifecycle reset removes pending bindings and counters rather than carrying them
|
|
23
|
+
into a resumed/switched branch;
|
|
24
|
+
- two Read/shell calls may complete in reverse order without crossing tool class,
|
|
25
|
+
parser command-scope or call identity;
|
|
26
|
+
- an unbound result increments the explicit `unboundResults` counter but its
|
|
27
|
+
bytes/class/parser facts are not aggregated as current-session evidence;
|
|
28
|
+
- extension reload during an in-flight SDK tool remains a separate P00 limitation
|
|
29
|
+
and is not hidden by telemetry; root tab/session generation guards keep a late
|
|
30
|
+
result owned by its origin tab.
|
|
31
|
+
|
|
32
|
+
No suppression ledger is introduced, and no persisted JSONL message is rewritten
|
|
33
|
+
when mode or session lifecycle changes.
|
|
34
|
+
|
|
35
|
+
## TUI / current result rendering
|
|
36
|
+
|
|
37
|
+
R-B already proves collapsed/expanded renderer equivalence for the exact
|
|
38
|
+
Read/Bash/ast_grep metadata-normalization surfaces, including image preservation.
|
|
39
|
+
The root renderer/session gate additionally covers:
|
|
40
|
+
|
|
41
|
+
- shell running/nonzero/signal rendering;
|
|
42
|
+
- LSP diagnostic severity rendering and mutation tool blocks;
|
|
43
|
+
- truncated collapsed previews and expanded full current-result bodies;
|
|
44
|
+
- persisted history tail/replay, unresolved/completed historical tool calls and
|
|
45
|
+
cancellation while older history is being prepended;
|
|
46
|
+
- stale runtime/session event rejection and late tool-result binding to the
|
|
47
|
+
original inactive tab.
|
|
48
|
+
|
|
49
|
+
These are current/persisted tool-result views; they are not a new route to the
|
|
50
|
+
producer's full native temp output.
|
|
51
|
+
|
|
52
|
+
## ACP / Desktop lazy current-result hydration
|
|
53
|
+
|
|
54
|
+
Pix already has `pix/session/history`, `pix/session/tool_result` and deferred image
|
|
55
|
+
hydration for Desktop. The lazy tool-result route is explicitly scoped by the ACP
|
|
56
|
+
`sessionId` and the deferred result map created for that session. For a persisted
|
|
57
|
+
result, the backend stores a trusted `sessionPath + byte offset + byte length`
|
|
58
|
+
reference and materializes exactly that persisted JSONL message line on demand.
|
|
59
|
+
|
|
60
|
+
The ACP contract now proves that requesting a deferred result:
|
|
61
|
+
|
|
62
|
+
- does not replay or append a `session/update`;
|
|
63
|
+
- returns the persisted current tool-result content/details only on explicit
|
|
64
|
+
hydration;
|
|
65
|
+
- cannot be performed with another session's ID even when the tool call ID is
|
|
66
|
+
known;
|
|
67
|
+
- frees the duplicate backend deferred copy after successful hydration.
|
|
68
|
+
|
|
69
|
+
Desktop transcript logic keeps deferred results lightweight until expansion,
|
|
70
|
+
hydrates them into the local transcript, and separately hydrates deferred image
|
|
71
|
+
bodies. This is a human/UI operation; it does not create a new model tool result
|
|
72
|
+
or provider input.
|
|
73
|
+
|
|
74
|
+
## Native temp output: deliberate limitation
|
|
75
|
+
|
|
76
|
+
There is **no existing authenticated host route that reads the contents of a
|
|
77
|
+
`Read`/Bash/ast_grep native temp/full-output handle for TUI/ACP/Desktop**.
|
|
78
|
+
`pix/session/tool_result` is not such a route: it reads a persisted JSONL result,
|
|
79
|
+
not the producer's `fullOutputPath` contents.
|
|
80
|
+
|
|
81
|
+
Therefore R-F adds no “open full native source” button or ACP endpoint. The UI may
|
|
82
|
+
show the current result/status/diagnostics already available through the normal
|
|
83
|
+
message path, but P01-R does not widen filesystem access merely because a result
|
|
84
|
+
contains a temp path. R-C's lifetime limitations remain in force: temp paths can
|
|
85
|
+
be absent, mutable, expired, or unavailable on error/timeout/abort and have no
|
|
86
|
+
resume/fork/export guarantee.
|
|
87
|
+
|
|
88
|
+
## Deterministic gates
|
|
89
|
+
|
|
90
|
+
Context Gateway lifecycle/result gate:
|
|
91
|
+
|
|
92
|
+
```text
|
|
93
|
+
bun test \
|
|
94
|
+
test/context-gateway/observe.test.ts \
|
|
95
|
+
test/context-gateway/sdk-pipeline.test.ts \
|
|
96
|
+
test/context-gateway/metadata-normalization.test.ts
|
|
97
|
+
```
|
|
98
|
+
|
|
99
|
+
Result: **49 pass, 0 fail, 337 assertions**. Suite typecheck and `git diff --check`
|
|
100
|
+
also pass for this gate.
|
|
101
|
+
|
|
102
|
+
Root TUI/session gate:
|
|
103
|
+
|
|
104
|
+
```text
|
|
105
|
+
node --import tsx --test \
|
|
106
|
+
tests/conversation-tool-renderer.test.ts \
|
|
107
|
+
tests/tool-block-renderer.test.ts \
|
|
108
|
+
tests/conversation-shell-renderer.test.ts \
|
|
109
|
+
tests/session-history.test.ts \
|
|
110
|
+
tests/session-lifecycle-controller.test.ts \
|
|
111
|
+
tests/tabs-controller.test.ts
|
|
112
|
+
```
|
|
113
|
+
|
|
114
|
+
Result: **98 pass, 0 fail**.
|
|
115
|
+
|
|
116
|
+
ACP lazy current-result gate:
|
|
117
|
+
|
|
118
|
+
```text
|
|
119
|
+
node --import tsx --test \
|
|
120
|
+
--test-name-pattern='desktop lazy session/load omits tool bodies and retrieves them on demand|desktop lazy session/load reuses an already-live tab runtime|deferred' \
|
|
121
|
+
test/agent.test.ts test/session-replay.test.ts
|
|
122
|
+
npm run typecheck
|
|
123
|
+
```
|
|
124
|
+
|
|
125
|
+
Result: **4 pass, 0 fail**; ACP typecheck passes. The targeted lazy-result test
|
|
126
|
+
also asserts no hydration `session/update` and denial for another session ID.
|
|
127
|
+
|
|
128
|
+
An additional ACP boundary suite covers persisted lazy materialization and replay:
|
|
129
|
+
|
|
130
|
+
```text
|
|
131
|
+
node --import tsx --test \
|
|
132
|
+
test/session-replay.test.ts \
|
|
133
|
+
test/session-history-file.test.ts \
|
|
134
|
+
test/desktop-commands.test.ts
|
|
135
|
+
```
|
|
136
|
+
|
|
137
|
+
Result: **6 pass, 0 fail**. It additionally proves that the Desktop tool-result
|
|
138
|
+
request surface is `sessionId + toolCallId` scoped, ignores arbitrary path-like
|
|
139
|
+
extras, legacy persisted metadata is not retroactively normalized during replay,
|
|
140
|
+
and deferred image/tool bodies are hydrated only from their existing persisted
|
|
141
|
+
session references.
|
|
142
|
+
|
|
143
|
+
Desktop gate:
|
|
144
|
+
|
|
145
|
+
```text
|
|
146
|
+
npm run check
|
|
147
|
+
npm test -- src/lib/transcript.test.ts src/lib/acp-client.test.ts
|
|
148
|
+
```
|
|
149
|
+
|
|
150
|
+
Result: Desktop check has **0 errors** and two pre-existing Svelte accessibility
|
|
151
|
+
warnings in `WorkspaceSidebar.svelte`; targeted tests are **28 pass, 0 fail**.
|
|
152
|
+
|
|
153
|
+
No live model call or manual suite sync was used for R-F.
|