pi-ui-extend 1.0.40 → 1.0.41

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (110) hide show
  1. package/dist/app/commands/command-registry.js +2 -2
  2. package/dist/app/commands/command-session-actions.d.ts +0 -1
  3. package/dist/app/commands/command-session-actions.js +22 -13
  4. package/dist/app/icons.d.ts +14 -0
  5. package/dist/app/icons.js +33 -0
  6. package/dist/app/rendering/conversation-tool-renderer.js +2 -2
  7. package/dist/app/rendering/dcp-stats.d.ts +6 -1
  8. package/dist/app/rendering/dcp-stats.js +214 -46
  9. package/dist/app/rendering/editor-panels.js +8 -5
  10. package/dist/app/session/lazy-session-manager.js +12 -1
  11. package/dist/app/session/tabs-controller.d.ts +2 -5
  12. package/dist/app/session/tabs-controller.js +12 -21
  13. package/dist/app/subagents/subagents-model.d.ts +14 -1
  14. package/dist/app/subagents/subagents-model.js +34 -15
  15. package/dist/app/types.d.ts +2 -0
  16. package/dist/bundled-extensions/session-title/config.js +1 -1
  17. package/dist/markdown-format.js +27 -9
  18. package/dist/schemas/pi-tools-suite-schema.d.ts +29 -16
  19. package/dist/schemas/pi-tools-suite-schema.js +46 -31
  20. package/external/pi-tools-suite/README.md +188 -55
  21. package/external/pi-tools-suite/docs/browser-qa-subagent.md +31 -21
  22. package/external/pi-tools-suite/docs/context-gateway-p00-adr.md +216 -0
  23. package/external/pi-tools-suite/docs/context-gateway-p01n-gate-review.md +122 -0
  24. package/external/pi-tools-suite/docs/context-gateway-p01n-measurement.md +133 -0
  25. package/external/pi-tools-suite/docs/context-gateway-p01r-ra-evidence.md +111 -0
  26. package/external/pi-tools-suite/docs/context-gateway-p01r-rb-evidence.md +100 -0
  27. package/external/pi-tools-suite/docs/context-gateway-p01r-rc-evidence.md +69 -0
  28. package/external/pi-tools-suite/docs/context-gateway-p01r-rd-evidence.md +100 -0
  29. package/external/pi-tools-suite/docs/context-gateway-p01r-re-evidence.md +74 -0
  30. package/external/pi-tools-suite/docs/context-gateway-p01r-rf-evidence.md +153 -0
  31. package/external/pi-tools-suite/docs/context-gateway-p01r-rg-evidence.md +235 -0
  32. package/external/pi-tools-suite/docs/subagent-model-pools.md +109 -0
  33. package/external/pi-tools-suite/package.json +3 -0
  34. package/external/pi-tools-suite/src/async-subagents/{private-skills → agents}/browser-qa/scripts/browser-qa-runner.mjs +82 -1
  35. package/external/pi-tools-suite/src/async-subagents/{private-skills/browser-qa/SKILL.md → agents/browser-qa.md} +261 -12
  36. package/external/pi-tools-suite/src/async-subagents/agents/implement.md +20 -0
  37. package/external/pi-tools-suite/src/async-subagents/agents/oracle.md +16 -0
  38. package/external/pi-tools-suite/src/async-subagents/agents/research.md +18 -0
  39. package/external/pi-tools-suite/src/async-subagents/agents/verify.md +18 -0
  40. package/external/pi-tools-suite/src/async-subagents/async-subagents.sample.jsonc +27 -243
  41. package/external/pi-tools-suite/src/async-subagents/commands.ts +6 -2
  42. package/external/pi-tools-suite/src/async-subagents/core/agent-catalog.ts +41 -0
  43. package/external/pi-tools-suite/src/async-subagents/core/agent-strategy.ts +13 -93
  44. package/external/pi-tools-suite/src/async-subagents/core/agents-dir.ts +494 -0
  45. package/external/pi-tools-suite/src/async-subagents/core/browser-qa.ts +9 -0
  46. package/external/pi-tools-suite/src/async-subagents/core/config.ts +200 -143
  47. package/external/pi-tools-suite/src/async-subagents/core/model-fallback.ts +1 -1
  48. package/external/pi-tools-suite/src/async-subagents/core/model-selection.ts +54 -0
  49. package/external/pi-tools-suite/src/async-subagents/core/prompt.ts +7 -6
  50. package/external/pi-tools-suite/src/async-subagents/core/routing.ts +52 -45
  51. package/external/pi-tools-suite/src/async-subagents/core/spawn.ts +12 -4
  52. package/external/pi-tools-suite/src/async-subagents/index.ts +11 -1
  53. package/external/pi-tools-suite/src/async-subagents/lib.ts +6 -2
  54. package/external/pi-tools-suite/src/async-subagents/tools/spawn.ts +46 -18
  55. package/external/pi-tools-suite/src/async-subagents/tools/subagents.ts +3 -2
  56. package/external/pi-tools-suite/src/async-subagents/types.ts +2 -0
  57. package/external/pi-tools-suite/src/config.ts +1 -1
  58. package/external/pi-tools-suite/src/context-gateway/accounting.ts +151 -0
  59. package/external/pi-tools-suite/src/context-gateway/config.ts +111 -0
  60. package/external/pi-tools-suite/src/context-gateway/index.ts +160 -0
  61. package/external/pi-tools-suite/src/context-gateway/metadata-normalization.ts +88 -0
  62. package/external/pi-tools-suite/src/context-gateway/storeless-capabilities.ts +89 -0
  63. package/external/pi-tools-suite/src/context-gateway/telemetry.ts +429 -0
  64. package/external/pi-tools-suite/src/context-gateway/test-output-parser.ts +326 -0
  65. package/external/pi-tools-suite/src/context-gateway/types.ts +152 -0
  66. package/external/pi-tools-suite/src/dcp/auto-compress-budget.ts +106 -0
  67. package/external/pi-tools-suite/src/dcp/auto-compress.ts +810 -106
  68. package/external/pi-tools-suite/src/dcp/commands.ts +64 -139
  69. package/external/pi-tools-suite/src/dcp/compress-tool.ts +369 -35
  70. package/external/pi-tools-suite/src/dcp/compression-blocks.ts +510 -64
  71. package/external/pi-tools-suite/src/dcp/compression-preview.ts +113 -0
  72. package/external/pi-tools-suite/src/dcp/compression-progress.ts +70 -0
  73. package/external/pi-tools-suite/src/dcp/config.ts +36 -61
  74. package/external/pi-tools-suite/src/dcp/conversation-index.ts +421 -0
  75. package/external/pi-tools-suite/src/dcp/debug-log.ts +7 -5
  76. package/external/pi-tools-suite/src/dcp/index.ts +617 -203
  77. package/external/pi-tools-suite/src/dcp/journal.ts +566 -0
  78. package/external/pi-tools-suite/src/dcp/progress-controller.ts +244 -0
  79. package/external/pi-tools-suite/src/dcp/prompts.ts +10 -7
  80. package/external/pi-tools-suite/src/dcp/provider-tool-results.ts +189 -0
  81. package/external/pi-tools-suite/src/dcp/pruner-candidates.ts +298 -78
  82. package/external/pi-tools-suite/src/dcp/pruner-compression-blocks.ts +173 -281
  83. package/external/pi-tools-suite/src/dcp/pruner-emergency.ts +2 -4
  84. package/external/pi-tools-suite/src/dcp/pruner-message-ids.ts +17 -5
  85. package/external/pi-tools-suite/src/dcp/pruner-metadata.ts +11 -1
  86. package/external/pi-tools-suite/src/dcp/pruner-nudge.ts +30 -82
  87. package/external/pi-tools-suite/src/dcp/pruner-tools.ts +22 -133
  88. package/external/pi-tools-suite/src/dcp/pruner.ts +18 -33
  89. package/external/pi-tools-suite/src/dcp/recovery.ts +129 -0
  90. package/external/pi-tools-suite/src/dcp/shadow-plan.ts +127 -0
  91. package/external/pi-tools-suite/src/dcp/state-transaction.ts +102 -0
  92. package/external/pi-tools-suite/src/dcp/state.ts +158 -580
  93. package/external/pi-tools-suite/src/dcp/ui.ts +1 -0
  94. package/external/pi-tools-suite/src/default-pi-tools-suite-config.ts +32 -214
  95. package/external/pi-tools-suite/src/index.ts +9 -0
  96. package/external/pi-tools-suite/src/model-tools/index.ts +76 -42
  97. package/external/pi-tools-suite/src/repo-discovery/index.ts +84 -18
  98. package/external/pi-tools-suite/src/repo-discovery/native-compact.ts +458 -0
  99. package/external/pi-tools-suite/src/session-recovery/index.ts +189 -43
  100. package/external/pi-tools-suite/src/tool-descriptions.ts +39 -35
  101. package/external/pi-tools-suite/src/truncation-metadata-normalizer/index.ts +17 -0
  102. package/package.json +3 -2
  103. package/schemas/pi-tools-suite.json +159 -78
  104. package/external/pi-tools-suite/src/async-subagents/private-skills/browser-qa/references/auth-scaffold-spec.md +0 -78
  105. package/external/pi-tools-suite/src/async-subagents/private-skills/browser-qa/references/qa-design.md +0 -223
  106. package/external/pi-tools-suite/src/dcp/state-persistence.ts +0 -195
  107. /package/external/pi-tools-suite/src/async-subagents/{private-skills/browser-qa/references → agents/browser-qa/examples}/qa-auth.example.jsonc +0 -0
  108. /package/external/pi-tools-suite/src/async-subagents/{private-skills/browser-qa/references → agents/browser-qa/examples}/qa-flow.example.jsonc +0 -0
  109. /package/external/pi-tools-suite/src/async-subagents/{private-skills → agents}/browser-qa/vendor/fflate.LICENSE +0 -0
  110. /package/external/pi-tools-suite/src/async-subagents/{private-skills → agents}/browser-qa/vendor/fflate.mjs +0 -0
@@ -100,7 +100,13 @@ DCP settings are stored only under `dcp` in the user shared config file `~/.conf
100
100
  "nudgeFrequency": 1,
101
101
  "iterationNudgeThreshold": 6,
102
102
  "nudgeForce": "strong",
103
- "protectedTools": ["compress", "write", "edit", "subagents"]
103
+ "protectedTools": ["compress", "write", "edit", "subagents"],
104
+ "autoCompress": {
105
+ "enabled": false,
106
+ "patience": 2,
107
+ "summarizerModel": [],
108
+ "timeoutMs": 20000
109
+ }
104
110
  },
105
111
  "strategies": {
106
112
  "emergencyCurrentTurnPruning": {
@@ -170,7 +176,11 @@ DCP settings are stored only under `dcp` in the user shared config file `~/.conf
170
176
 
171
177
  `minContextPercent` / `maxContextPercent` accept legacy fractions (`0.25`), percent strings (`"25%"`), or absolute token counts when Pi knows the current model context window. `minContextLimit` / `maxContextLimit` and `modelMinContextLimits` / `modelMaxContextLimits` are explicit absolute-or-percent aliases. `modelOverrides` and the `modelMin*` / `modelMax*` maps support exact model keys plus `*` / `?` wildcard patterns; matching is applied from generic to specific so exact bare-model matches override bare wildcards, and exact `provider/model` matches override provider wildcards. Array fields are union-merged, so model-specific `protectedTools` extend the defaults instead of replacing them. If `compress.protectUserMessages` is enabled, range compression appends selected user messages verbatim instead of rejecting the range; individual message compression still skips protected raw user messages. Protected tool outputs are copied into summaries for tools protected by name or `protectedFilePatterns`; protected `subagents` result reads also try to include the saved `result.md` artifact when available.
172
178
 
173
- `strategies.emergencyCurrentTurnPruning` is the default-enabled lossy safety floor for a single unfinished turn that has no normal compression candidate. DCP first emits emergency reminders and offers only safe old same-turn tool-result candidates. After `patience` ignored reminders, or at the model-independent `hardContextPercent`, it replaces eligible oldest result bodies until the estimated provider context reaches `targetContextPercent` or a margin below the model emergency threshold. User messages, configured/protected data, the newest `keepRecentToolPairs`, and results not present in an accepted provider request are never selected. The raw session transcript is unchanged. Setting `enabled` to `false` disables same-turn candidates and lossy pruning, but keeps the non-destructive emergency reminder.
179
+ `compress.autoCompress.enabled` is `false` by default. When explicitly enabled, its `patience` counts completed correlated main-provider opportunities, not repeated context transforms; `summarizerModel: []` uses the bounded extractive fallback without a model call, while configured models share the single `timeoutMs` summarizer deadline. If that deadline expires, DCP still has a bounded finalization grace to commit the extractive fallback. Auto commit is accepted only when the full projected replacement has positive gain and meets the current budget-recovery target.
180
+
181
+ `strategies.emergencyCurrentTurnPruning` is the default-enabled lossy safety floor for a single unfinished turn that has no normal compression candidate. DCP first emits emergency reminders and offers only safe old same-turn tool-result candidates. After `patience` completed ignored opportunities, or at the model-independent `hardContextPercent`, it replaces eligible oldest result bodies until the estimated provider context reaches `targetContextPercent` or a margin below the model emergency threshold. User messages, configured/protected data, the newest `keepRecentToolPairs`, and results without **completed** provider evidence are never selected. HTTP 2xx alone is not evidence; DCP promotes eligibility only after an unambiguously correlated successful finalized assistant response, and ambiguous retries/interleaving fail closed. The raw session transcript is unchanged. Setting `enabled` to `false` disables same-turn candidates and lossy pruning, but keeps the non-destructive emergency reminder.
182
+
183
+ DCP sidecars are session-private versioned generation envelopes under `<sessionDir>/dcp-state/`. Saves use private atomic files, retain a last-valid `.prev` generation, validate payload/block-graph integrity on load, quarantine corrupt primaries, and use a cross-process exclusive lock rather than silent last-writer-wins. Live paused sessions are not deleted merely for age. Protected subagent artifacts are optional bounded recovery input: reads are async, rooted at the session cwd, reject symlink escapes and oversized files, and never silently truncate a required protected artifact.
174
184
 
175
185
  Set `dcp.debug: true` to write a JSONL debug log of DCP context/prune/compress events to `~/.pi/agent/dcp-debug.jsonl` (override the path with `PI_DCP_DEBUG_LOG`, or enable without config via `PI_DCP_DEBUG=1`); off by default. The log is size-limited and rotated: once it reaches `dcp.debugLog.maxBytes` (default `5242880` = 5 MB) it is renamed to `.1`, older backups shift down (`.1`→`.2`, …) and the oldest beyond `dcp.debugLog.maxBackups` (default `3`, minimum `1`) is dropped; override either with `PI_DCP_DEBUG_MAX_BYTES` / `PI_DCP_DEBUG_MAX_BACKUPS`.
176
186
 
@@ -427,21 +437,120 @@ Notes:
427
437
 
428
438
  ## Async sub-agents
429
439
 
430
- Sub-agent model routing normally follows task overrides, subagent type config, then `ASYNC_SUBAGENTS_MODEL` / `PI_SUBAGENTS_MODEL` fallbacks. Set `ASYNC_SUBAGENTS_FORCE_CURRENT_MODEL=1` (or `PI_SUBAGENTS_FORCE_CURRENT_MODEL=1`) to ignore task/config/env model choices and launch every sub-agent with the current parent session model. When this flag is enabled, any `--model` entries in sub-agent extra args are stripped so they cannot override the current model.
440
+ Model selection uses the ordered candidates from each agent's Markdown file,
441
+ filtered by the selected preset's available models and runtime capabilities.
442
+ Explicit task/CLI model overrides bypass the pool. Setting
443
+ `ASYNC_SUBAGENTS_FORCE_CURRENT_MODEL=1` (or
444
+ `PI_SUBAGENTS_FORCE_CURRENT_MODEL=1`) deliberately selects the parent model and
445
+ strips conflicting model arguments; this is not the economical default.
446
+
447
+ The five built-in modes are `research` (read-only evidence and independent
448
+ review), `implement` (bounded code, docs, tests, or UI changes), `verify`
449
+ (run checks and diagnose logs without fixing files), `browser-qa` (trusted
450
+ browser workflow), and `oracle` (deliberate strong second opinion).
451
+ Ordinary workers use economical model candidates; no built-in parent-tier
452
+ rule promotes them to a flagship. Oracle is the exception, not an automatic
453
+ retry for difficult work. Task-specific discipline belongs in the brief.
454
+
455
+ Delegate when a suitable lower-cost worker can handle bounded work or noisy
456
+ intermediate evidence should stay outside the parent context. One sequential
457
+ task can qualify. Keep decisions and integration in the parent; read compact
458
+ results and verify selectively rather than repeating the worker's investigation.
459
+ Do trivial reads/edits directly. Redirect a noisy command to a log without an
460
+ extra LLM when no interpretation is needed. `verify`'s no-edit instruction is
461
+ a behavioral contract, not a read-only filesystem sandbox for its shell.
462
+
463
+ Run `/ultrawork` or `/ulw` for orchestration, `/hyperplan` to pressure-test a
464
+ plan, or set `ULTRAWORK=1` to apply the orchestration prompt to normal inputs.
465
+ `ULTRAWORK_AUTO=1` classifies only the first normal input on non-GPT parents;
466
+ GPT-like parents skip that automatic transform, not ordinary delegation.
467
+
468
+ See [Model pools and migration](docs/subagent-model-pools.md) for the selection
469
+ contract, configuration examples, override rules and legacy compatibility.
470
+
471
+ ### Parent-first role selection
472
+
473
+ The parent normally selects an explicit `subagentType` from the effective
474
+ system-prompt catalog, preferring a matching project-local specialist. Valid
475
+ explicit types bypass the LLM router entirely; presets, model selection, tools,
476
+ skills, and role instructions are still applied by the normal config resolver.
477
+ Model/thinking overrides are not substitutes for selecting a role.
478
+
479
+ The router remains enabled as a fallback for omitted types: use it when the role
480
+ is unclear or the user explicitly requests automatic routing. Only omitted
481
+ tasks are classified, in one batch; the parent's explicit choices are preserved.
482
+ Real-browser QA still requires explicit `subagentType: "browser-qa"`.
483
+
484
+ Unknown explicit types and failed/incomplete automatic routing reject the
485
+ **entire spawn batch before run state or child processes are created**. The tool
486
+ returns an error with affected task IDs and available types; the parent should
487
+ correct the roles and resubmit the whole batch. Provider error responses are
488
+ failures too, not successful routes. Configured fallback router models may be
489
+ tried, but missing routes are never silently replaced with `quick`/`defaultType`.
490
+
491
+ With `routing.enabled: false`, every spawn task must supply a valid explicit
492
+ type. `defaultType` remains a preference for genuinely ambiguous LLM choices
493
+ and a legacy config-resolver default, not a spawn error fallback. Existing
494
+ callers relying on an implicit default must now choose a type explicitly.
495
+
496
+ ### Project-local agents (`.pi/agents/*.md`)
497
+
498
+ A project can ship sub-agent roles as individual Markdown files in
499
+ `<project>/.pi/agents/`. The first such directory found walking up from the
500
+ session cwd is used; each top-level `*.md` file becomes a `subagentType` named
501
+ after the file. Parent and router see the short `description`; only the child
502
+ receives the Markdown body. A project's ordered `models` are filtered through
503
+ the same active preset pool as built-in agents.
504
+
505
+ ```markdown
506
+ ---
507
+ description: Use for reviewing this repo's diff — knows the house rules.
508
+ icon: eye
509
+ models:
510
+ - zai/glm-5-turbo
511
+ - openai-codex/gpt-5.6-luna
512
+ thinking: low
513
+ tools: read, grep
514
+ retry:
515
+ maxRetries: 1
516
+ backoffMs: 2000
517
+ ---
518
+
519
+ You are this project's staff reviewer. Apply the repo rules from
520
+ AGENTS.md before approving anything; cite file paths first.
521
+ ```
431
522
 
432
- For an oh-my-openagent-style workflow, run `/ultrawork` or `/ulw` to ask the parent agent to split broad work into configured async-subagents roles (`quick`, `scan`, `research`, `docs`, `frontend`, `browser-qa`, `implement`, `tests`, `review`, `deep`, `oracle`). Set `ULTRAWORK=1` before launching Pi to apply that compact routing prompt to normal non-slash user inputs automatically. Set `ULTRAWORK_AUTO=1` to ask the lightweight router model to classify only the first normal user input on non-GPT parent models: clear broad/parallel work is transformed into ultrawork, vague potentially-complex work gets a soft delegation hint, and narrow work is left unchanged. GPT-like parent models skip only this automatic transform; they can still use `/ultrawork` and `subagents` normally. `frontend` is for UI/UX, styling, layout, responsive behavior, and visual component polish; `browser-qa` reproduces browser bugs and proves fixes with deterministic assertions plus screenshot/video/trace evidence; `review` covers security/performance/audit tracks; `implement` covers refactors; `deep` covers debugging/root-cause; `oracle` is for sparse cross-provider second opinions on high-stakes uncertainty. Run `/hyperplan` to pressure-test a plan before implementation.
523
+ - Frontmatter keys: `name` (must match the filename), `description`, `icon`, `models`, `thinking`, `tools`, `isolatedSkills`, `extraArgs`, `promptAppend`, `promptOverride`, `retry`, `maxResultBytes`, `timeoutMs`. Legacy `model`, `fallbackModels`, and `modelByParent` still load. Unknown keys are rejected with an error naming the file.
524
+ - Array fields accept block lists (`- item`), inline arrays (`[a, b]`), or comma-separated strings (`tools: read, grep, bash`). The frontmatter YAML subset is intentionally small: scalars, quoted strings, numbers, comments, lists, and nested maps for `modelByParent`/`retry`. Tabs, block scalars (`|`/`>`), anchors/aliases, and flow maps are hard errors naming file and line.
525
+ - The markdown body becomes `promptAppend`: it is appended after the standard generated prompt (parent objective + task + output format), so the agent still receives its task in the usual structure. Use frontmatter `promptOverride` for full prompt replacement.
526
+ - Precedence: project agent fields override same-named types from user/project JSONC config (field-level; other fields are kept), which in turn override built-ins. Setting `ASYNC_SUBAGENTS_CONFIG` / `PI_SUBAGENTS_CONFIG` disables the directory (explicit config = full control).
527
+ - Files without frontmatter are skipped (a `README.md` there is fine). Definition loading is uncached: edits apply on the next config read/spawn without a restart, and the effective system-prompt catalog is rebuilt at parent-agent start.
528
+ - Bundled roles use the same format internally under `src/async-subagents/agents/*.md`; built-in and project-local profiles therefore share one parser and normalization path instead of maintaining a second role-description schema in TypeScript.
529
+ - `icon` names an agent glyph for UIs that render sub-agent widgets (pix TUI panel, Pix Desktop subagents panel): `agent` (neutral default), `search`, `code`, `flask`, `globe`, `sparkles`, `brain`, `wrench`, `terminal`, `bug`, `book`, `eye`, `zap`, `rocket`. The value is passed through opaquely; unknown names render as the neutral agent icon, and status stays color-coded next to it.
433
530
 
434
531
  ### Private browser QA and project auth
435
532
 
436
533
  The built-in `browser-qa` role runs on `zai/glm-5.3-flash`, with
437
- `openai-codex/gpt-5.6-luna` as its fallback. Its browser
438
- workflow is an explicit private skill under `src/async-subagents/private-skills/`,
439
- outside normal Pi skill discovery. The role's first-class `isolatedSkills` setting launches the child with
440
- `--no-skills` plus one self-contained private workflow. It bundles the relevant
441
- scenario-design, locator, waiting, assertion, evidence, and cleanup guidance next
442
- to its trusted runner, so browser QA does not depend on a separately installed
443
- skill or CLI. The private workflow remains mandatory when configuration appends
444
- other isolated skills; parent and ordinary sub-agent sessions do not discover it.
534
+ `openai-codex/gpt-5.6-luna` as its fallback. Its complete workflow and detailed
535
+ scenario-design guidance live in the Markdown body of
536
+ `src/async-subagents/agents/browser-qa.md`. The normal profile loader appends
537
+ that body to the QA child's task prompt; the parent and LLM router receive only
538
+ the short `description`. There is no additional QA skill to discover or read.
539
+
540
+ Executable resources live under `src/async-subagents/agents/browser-qa/`.
541
+ The launcher supplies the installed runner's absolute path in
542
+ `PI_BROWSER_QA_RUNNER`; the child invokes `node "$PI_BROWSER_QA_RUNNER"` from
543
+ the delegated project's cwd. This non-secret path is set only for QA children.
544
+ QA always launches with `--no-skills`, even when `isolatedSkills` is empty,
545
+ and skill flags in `extraArgs` cannot bypass that isolation. Explicitly
546
+ configured `isolatedSkills` remain supported as optional additions; no built-in
547
+ QA `--skill` is injected. Other roles retain their normal discovery behavior.
548
+
549
+ Model/thinking/tool-only profile overrides inherit the Markdown workflow.
550
+ An explicit profile `promptAppend` replaces the inherited body under the usual
551
+ field-level merge rules; custom QA instructions must preserve the runner-only,
552
+ credential, target, and evidence contracts. Runner-enforced isolation and
553
+ credential handling remain in code, not in the prompt.
445
554
 
446
555
  Public browser QA does not require an auth profile or `.pi/qa_auth.jsonc`: run it
447
556
  with an explicit base URL, whose exact origin becomes the fail-closed allowlist.
@@ -498,8 +607,8 @@ creating a template. Only an explicit authenticated request may create the
498
607
  private template. Missing, rejected, or expired selected auth returns
499
608
  `QA_AUTH_UPDATE_REQUIRED`, naming only the profile/file/reason needed for the
500
609
  parent to ask the user for an update and rerun. See
501
- `src/async-subagents/private-skills/browser-qa/references/qa-auth.example.jsonc`
502
- for complete profile shapes and `references/qa-flow.example.jsonc` beside it for
610
+ `src/async-subagents/agents/browser-qa/examples/qa-auth.example.jsonc`
611
+ for complete profile shapes and `examples/qa-flow.example.jsonc` beside it for
503
612
  the declarative, non-executable QA action/assertion format.
504
613
 
505
614
  Browser QA videos automatically visualize pointer interactions. Clicks and
@@ -514,68 +623,92 @@ Async-subagents also injects a lightweight oh-my-openagent-style system-prompt s
514
623
 
515
624
  For blind-model screenshot/image inspection, use the main-session `coding-discipline` lookup tool; the bundled default uses vision-capable `zai/glm-5.3-flash`. Async-subagents still supports `imagePaths` on tasks when a broader delegated track genuinely needs images, but it no longer ships a dedicated `vision` role. Dynamic provider capabilities can be missing or stale after switching models, so blind parent models can still be configured explicitly with case-insensitive `*` masks under `asyncSubagents.vision.blindModelPatterns` in `~/.config/pi/pi-tools-suite.jsonc`; do not include `zai/glm-5.3-flash` because it accepts image input. This keeps guidance honest, not a sub-agent role.
516
625
 
517
- When a task omits `subagentType`, async-subagents asks a lightweight router model to choose one configured type for each task from the task text/scope and the `types.<name>.description` metadata. Explicit task `subagentType` still wins. Keep type descriptions short, literal, and distinct because they are inserted into the router prompt for a small model. Router settings live under `asyncSubagents.routing` (`enabled`, `model`, `maxTaskChars`, `maxTokens`, `maxRetries`, `timeoutMs`, `debug`); the default router model is `zai/glm-4.5-air`. If the router is disabled, unavailable, aborted, or returns invalid JSON, omitted types fall back to `defaultType`.
518
-
519
- Define optional `presets` under `asyncSubagents` in `~/.config/pi/pi-tools-suite.jsonc`, `$PI_CONFIG_DIR/pi-tools-suite.jsonc`, or project `.pi/pi-tools-suite.jsonc`, then use `/subagent-preset` or `/subagent-preset-config` to pick one persistent active preset for future spawns across all sessions. Set `AGENTS_PRESET=<name>` before launching Pi to override the saved preset for only the current process/session without changing the saved selection. If Pi is already running, use `/subagent-preset session <name>` for the same process-only override, and `/subagent-preset session-clear` to remove that runtime override. The TUI only selects presets already present in config; it does not edit JSON. If no `asyncSubagents` section exists, run `/subagent-preset init` to insert the bundled sample from `src/async-subagents/async-subagents.sample.jsonc` into the shared config (or to copy a standalone override file when `ASYNC_SUBAGENTS_CONFIG` / `PI_SUBAGENTS_CONFIG` is set). Existing config sections/files are never overwritten. Presets select an agent/model configuration: they can provide global fallback `model`/`thinking`/`extraArgs` and per-role overrides under `asyncSubagents.presets.<name>.types.<subagentType>`. They can also provide ordered `fallbackModels` globally or per-role; when a sub-agent fails with quota/rate-limit errors such as 429, async-subagents immediately tries the next fallback model and remembers the exhausted provider for the current Pi process/session, so later spawns skip that provider until Pi exits. This is intended for provider-level fallback chains such as `antigravity/* → openai-codex/* → zai/*` or `openai-codex/* → zai/*`; omit fallbacks for effectively unlimited providers. Antigravity account rotation has priority over preset fallback: async-subagents only falls back after Antigravity reports that all configured accounts are exhausted for that model. Explicit task model overrides and force-current-model disable preset fallback for that task. The active preset name is stored separately in `~/.pi/agent/subagent-preset-selection.json`.
626
+ When `subagentType` is omitted, the lightweight role router classifies the task
627
+ using the descriptions. Explicit types bypass it. Unknown types or failed
628
+ routing reject the batch, never substitute `defaultType`. Choosing a worker
629
+ model from its candidate list does not involve an LLM call.
630
+
631
+ ### Presets are available-model pools
632
+
633
+ Each agent declares an ordered `models` list in Markdown. A preset declares
634
+ which model references may be used, not another role/model/thinking matrix.
635
+ Selection preserves agent order, intersects it with `preset.models`, checks
636
+ runtime registration/auth availability, and takes the first usable candidate.
637
+ Pool order does not change preference and pool-only models are never appended.
638
+ Without a preset, the full agent list is eligible. Candidate order expresses
639
+ the configured budget preference; runtime does not infer current API prices.
640
+
641
+ Image-bearing tasks and `browser-qa` require confirmed image support; configured
642
+ blind-model masks override runtime image metadata. Remaining eligible models
643
+ form the quota fallback chain, so neither quota history nor image fallback can
644
+ escape the pool. Antigravity account rotation still happens before provider
645
+ fallback. No match, no usable model, or an explicitly empty list rejects the
646
+ whole batch before run directories or child processes are created. A new custom
647
+ agent must supply candidates instead of silently inheriting the parent model.
648
+
649
+ Oracle uses its separate strong-model list and prefers another provider when
650
+ available, but also respects the pool. A same-provider choice is allowed when
651
+ the pool offers no alternative; cross-provider independence is not guaranteed.
652
+ Explicit task/CLI model overrides and `FORCE_CURRENT_MODEL` remain deliberate
653
+ escape hatches and disable automatic model fallback for that task. They do not
654
+ bypass the image-capability check.
655
+
656
+ Define pools in the shared or project `pi-tools-suite.jsonc`. Select a saved
657
+ pool with `/subagent-preset`; use `AGENTS_PRESET=<name>` or
658
+ `/subagent-preset session <name>` for a process-only override and
659
+ `/subagent-preset session-clear` to remove it. The saved selection lives in
660
+ `~/.pi/agent/subagent-preset-selection.json`. `/subagent-preset init` inserts the
661
+ sample only when config is missing. The shipped pools are `cheap` (GLM), `gpt`,
662
+ and `deep` (the retained legacy name for the mixed pool, not worker escalation).
663
+ Initial user config and the sample share one source; descriptions and worker
664
+ model order exist only in the agent files. Existing user files are not rewritten.
520
665
 
521
666
  Example shared async-subagents config section:
522
667
 
523
668
  ```jsonc
524
669
  {
525
670
  "asyncSubagents": {
526
- "defaultType": "quick",
671
+ "defaultType": "research",
527
672
  "routing": {
528
673
  "enabled": true,
529
- "model": "zai/glm-4.5-air",
674
+ "model": "zai/glm-5-turbo",
530
675
  "timeoutMs": 12000
531
676
  },
532
677
  "presets": {
533
678
  "cheap": {
534
- "description": "Use GLM by role, including GLM-5.3 Flash for multimodal work.",
535
- "types": {
536
- "quick": { "model": "zai/glm-5.3", "thinking": "off" },
537
- "frontend": { "model": "zai/glm-5.3-flash", "thinking": "medium" },
538
- "browser-qa": { "model": "zai/glm-5.3-flash", "fallbackModels": ["openai-codex/gpt-5.6-luna"], "thinking": "low" },
539
- "review": { "model": "zai/glm-5.3", "thinking": "high" }
540
- }
679
+ "description": "GLM workers with a strong oracle candidate.",
680
+ "models": ["zai/glm-5-turbo", "zai/glm-5.3-flash", "zai/glm-5.3"]
541
681
  }
542
682
  },
543
683
  "types": {
544
- "frontend": {
545
- "description": "Use for frontend UI/UX visual work: styling, layout, typography, animation, responsive states, component polish, accessibility. Avoid backend/business logic unless needed for UI behavior.",
546
- "thinking": "medium"
547
- },
548
- "review": {
549
- "description": "Use for review/audit of existing code or changes: correctness, security, performance, maintainability, API risks, quality. Do not implement new code.",
550
- "thinking": "high"
684
+ "research": {
685
+ "models": ["zai/glm-5-turbo", "openai-codex/gpt-5.6-luna"],
686
+ "thinking": "low"
551
687
  }
552
688
  }
553
689
  }
554
690
  }
555
691
  ```
556
692
 
557
- ### Parent-model-aware model selection (`modelByParent`)
558
-
559
- Any type profile can carry `modelByParent`: a map from glob model refs (matched against the **current parent model**, e.g. `"zai/*"`) to a model for that role. The first matching key wins. Values may be a model string or `{ "model": "...", "fallbackModels": [...] }`. It is resolved after an explicit task `model` / `forcedModel`, but **before** the preset/static profile `model`, so a role can always pick a model based on who the parent is — independent of the active preset.
560
-
561
- The canonical use case is an **`oracle`** role that consults a flagship model from a *different* provider than the parent for a second opinion:
562
-
563
- ```jsonc
564
- "oracle": {
565
- "description": "Cross-provider second opinion: consult a flagship from a different provider than the parent to pressure-test a hard decision. Read-only; advise, do not edit.",
566
- "model": "openai-codex/gpt-5.6-sol",
567
- "fallbackModels": ["zai/glm-5.3"],
568
- "thinking": "max",
569
- "modelByParent": {
570
- "zai/*": { "model": "openai-codex/gpt-5.6-sol", "fallbackModels": ["zai/glm-5.3"] },
571
- "openai-codex/*": "zai/glm-5.3",
572
- "antigravity/*": { "model": "zai/glm-5.3", "fallbackModels": ["openai-codex/gpt-5.6-sol"] },
573
- "anthropic/*": { "model": "openai-codex/gpt-5.6-sol", "fallbackModels": ["zai/glm-5.3"] }
574
- }
575
- }
576
- ```
577
-
578
- With this config a GLM parent (`zai/*`) spawns the oracle on `gpt-5.6-sol`, a GPT parent (`openai-codex/*`) spawns it on `glm-5.3`, and so on — automatically, at spawn time, with no `task.model` needed. The parent model ref is read from the spawn context (`ctx.model`) and passed into resolution. Pattern matching is case-insensitive `*` glob (same engine as `vision.blindModelPatterns`). When no key matches (or no parent model is known), the role falls back to its static `model` + `fallbackModels`. An explicit `task.model` or `ASYNC_SUBAGENTS_FORCE_CURRENT_MODEL=1` still overrides the match.
693
+ ### Legacy configuration compatibility
694
+
695
+ Old built-in role names are no longer implicit aliases. `quick`, `scan`,
696
+ `review`, `deep`, `docs`, `frontend`, and `tests` are valid only when explicitly
697
+ defined as ordinary custom/project types. Old preset per-role keys likewise
698
+ apply only when a type with that exact name exists.
699
+
700
+ Legacy `model` plus `fallbackModels` remains readable. `models` is a complete
701
+ replacement list: it clears inherited legacy model/fallback/parent routing.
702
+ A later old-format model override still replaces the primary candidate, and a
703
+ later `fallbackModels` replaces the remaining candidates; `[]` disables them.
704
+ Old `modelByParent` configs remain supported, but ordinary roles give legacy
705
+ preset models precedence. New built-ins contain no parent-tier escalation maps.
706
+
707
+ When a preset specifies `models`, it is exclusively a pool; inherited legacy
708
+ `model`, `types`, thinking, arguments and timeout overrides do not run. A later
709
+ explicit old-format preset selector can still replace a pool for compatibility.
710
+ Runtime retry structures and the separate role router continue to use the
711
+ term `fallbackModels` for actual fallback-only lists, not agent candidates.
579
712
 
580
713
  Sub-agents run with `--no-session` by default to avoid writing duplicate Pi session JSONL files for fire-and-forget background work. Set `ASYNC_SUBAGENTS_ENABLE_SESSIONS=1` to restore persisted per-agent sessions under each agent's `sessions/` directory; this also registers the session-navigation slash commands (`/sub-open`, `/sub-back`, `/sub-where`) needed for switching and deeper post-mortem navigation.
581
714
 
@@ -4,25 +4,32 @@
4
4
 
5
5
  Provide a cheap, fast `browser-qa` async-subagent that reproduces browser bugs
6
6
  and proves fixes with deterministic assertions plus screenshot, video, and trace
7
- evidence. The role uses `zai/glm-5.3-flash`, falling back to
8
- `openai-codex/gpt-5.6-luna`.
9
-
10
- ## Private skill isolation
11
-
12
- - The browser QA skill lives under `src/async-subagents/private-skills/`, outside
13
- Pi's normal skill discovery roots.
7
+ evidence. Its ranked `models` list prefers `zai/glm-5.3-flash`, then
8
+ `openai-codex/gpt-5.6-luna`, filtered by the active preset's model pool and
9
+ confirmed runtime image support.
10
+
11
+ ## Inline agent workflow and skill isolation
12
+
13
+ - All operating instructions, flow contracts, scenario-design guidance, and
14
+ auth-scaffolding rules live in `src/async-subagents/agents/browser-qa.md`.
15
+ Its body becomes the QA child's `promptAppend` through the shared agent
16
+ loader. Parent and router catalogs include only its short `description`.
17
+ - Runner code, vendor dependencies/licenses, and optional JSONC examples live
18
+ under `src/async-subagents/agents/browser-qa/`; none is a discoverable skill.
14
19
  - Sub-agent processes disable normal extension discovery, then always load the
15
- suite's model-tools and Antigravity provider extensions explicitly, regardless
16
- of the selected model. This keeps the process isolated without making any
17
- Antigravity-backed role unavailable.
20
+ suite's model-tools extension. They load the Antigravity provider extension
21
+ only when an Antigravity model is explicitly selected.
18
22
  - A type profile may declare `isolatedSkills`. Spawning that profile adds
19
23
  `--no-skills` followed by one explicit `--skill` per configured path.
20
- - The `browser-qa` profile always loads one self-contained private workflow.
21
- Relevant browser-test design guidance is bundled beside its trusted runner;
22
- no separately discovered skill or browser CLI is required. Configuration may
23
- append isolated skills but cannot remove the mandatory private workflow.
24
- - Other sub-agent profiles and the parent session must not discover the private
25
- skill automatically.
24
+ - `browser-qa` always disables normal skill discovery and filters skill flags
25
+ out of `extraArgs`, even without configured skills. It no longer injects a
26
+ mandatory QA skill. Explicitly configured skills are optional additions.
27
+ - The launcher sets `PI_BROWSER_QA_RUNNER` to the absolute installed runner path,
28
+ replacing inherited values, and strips it from ordinary child environments.
29
+ QA invokes `node "$PI_BROWSER_QA_RUNNER"` without guessing paths from cwd.
30
+ - Model-only profile overrides inherit the workflow. An explicit profile
31
+ `promptAppend` replaces the body like any other agent profile; it is not an
32
+ immutable security boundary. Runtime protections remain in the runner.
26
33
 
27
34
  ## Authentication contract
28
35
 
@@ -98,7 +105,7 @@ evidence. The role uses `zai/glm-5.3-flash`, falling back to
98
105
 
99
106
  ## Reliability and shutdown contract
100
107
 
101
- - The built-in `browser-qa` profile has a 120-second wall-clock budget unless
108
+ - The built-in `browser-qa` profile has a 300-second wall-clock budget unless
102
109
  the caller explicitly supplies a task or spawn timeout. This bounds model
103
110
  stalls as well as browser work.
104
111
  - The trusted runner has its own bounded lifecycle. Browser launch, context
@@ -124,11 +131,14 @@ evidence. The role uses `zai/glm-5.3-flash`, falling back to
124
131
 
125
132
  ## Acceptance criteria
126
133
 
127
- 1. `browser-qa` resolves to the intended model/fallback and its self-contained
128
- private workflow, and its isolated child process can register the configured
134
+ 1. `browser-qa` resolves to the intended model/fallback and its inline Markdown
135
+ workflow, and its isolated child process can register the configured
129
136
  model provider.
130
- 2. Spawn args contain `--no-skills` and the mandatory private skill for this
131
- profile; ordinary profiles retain existing skill discovery behavior.
137
+ 2. Default QA spawn args contain `--no-skills` but no `--skill`. The child
138
+ receives the full workflow in its initial prompt and can invoke the bundled
139
+ runner through `PI_BROWSER_QA_RUNNER` from an unrelated project directory.
140
+ Optional configured skills still load; ordinary profiles retain existing
141
+ skill discovery behavior and do not receive QA-only environment paths.
132
142
  3. Auth profile listing and all error output are redacted; model-authored input
133
143
  cannot execute code in the credential-bearing process.
134
144
  4. Runner tests cover public execution without an auth file, explicit profile
@@ -0,0 +1,216 @@
1
+ # Context Gateway P00 ADR: tool-result stage ordering
2
+
3
+ <!-- markdownlint-disable MD013 -->
4
+
5
+ > Status: accepted implementation decision for future Context Gateway work.
6
+ > This ADR does not enable Gateway, change user configuration, or certify rollout.
7
+ > Evidence baseline: repository `daa1b06`, installed Pi SDK `0.85.1`, 7 September 2026.
8
+
9
+ ## Context
10
+
11
+ The installed extension runner chains `tool_result` handlers in registration
12
+ order. The suite currently has four independent `tool_result` registrations:
13
+ LSP, comment-checker, DCP, and the opt-in credential firewall. Their module load
14
+ order is LSP → comment-checker → DCP → credential firewall.
15
+
16
+ That order is unsuitable for Gateway enforce mode:
17
+
18
+ - mutation diagnostics must be present before Gateway renders a bounded result;
19
+ - session-hygiene redaction, when enabled, must happen before durable capture;
20
+ - Gateway must shape before DCP records the delivered result;
21
+ - changing the global `MODULES` order would also change unrelated lifecycle and
22
+ provider hooks, including the intentionally late provider firewall.
23
+
24
+ P00 tests also prove that throwing from a result handler is fail-open, and that
25
+ an in-flight extension tool can cross an extension reload: the old wrapper then
26
+ fails on its stale runner while the new runner handles the resulting error.
27
+ Therefore current active-runner state is not an origin binding.
28
+
29
+ ## Decision
30
+
31
+ For the current **P01-R storeless scope**, keep the existing independent
32
+ `tool_result` chain. A suite-local coordinator is **not** introduced merely to
33
+ make the architecture look uniform. It becomes conditional future work only if
34
+ a store-backed/enforce stage is authorised and a concrete conflict between
35
+ enabled suite-owned result handlers cannot be expressed safely through the
36
+ verified event-specific order.
37
+
38
+ The verified storeless order is:
39
+
40
+ 1. LSP and comment-checker result enrichers.
41
+ 2. Context Gateway `observe` (passive; no result patch).
42
+ 3. Optional `truncation-metadata-normalizer` (details-only duplicate cleanup).
43
+ 4. Other downstream result observers/modifiers; the current legacy DCP happens
44
+ to be in this position but P01-R does not depend on DCP state or persistence.
45
+ 5. Opt-in credential-firewall session-hygiene redaction.
46
+
47
+ For provider hooks, credential-firewall redaction runs before
48
+ `codex-reasoning-fix`, and `codex-reasoning-fix` remains the final
49
+ `before_provider_request` sanitizer.
50
+
51
+ This order is covered both by the `MODULES` contract and by an executable
52
+ `ExtensionRunner` chain with a DCP-free fake downstream observer. The optional
53
+ normalizer removes only redundant metadata; the later firewall still owns secret
54
+ redaction. With firewall session hygiene disabled, the normalizer does not redact
55
+ or otherwise alter visible content.
56
+
57
+ If a future enforce/store stage is authorised, the previously proposed
58
+ coordinator remains the candidate design: modules participating in it must not
59
+ also register an independent `tool_result` handler, while their non-result hooks
60
+ remain registered normally. That future decision must be revalidated against the
61
+ then-current DCP/session-recovery implementation rather than copied mechanically
62
+ from the legacy chain.
63
+
64
+ The **future enforce coordinator candidate** order is:
65
+
66
+ 1. **Enrich** — LSP mutation diagnostics, then comment-checker output.
67
+ 2. **Session-hygiene sanitize** — the credential-firewall tool-result transform
68
+ when that existing opt-in policy is enabled.
69
+ 3. **Gateway sanitize/capture/render** — apply Gateway's archive-specific
70
+ permitted-snapshot sanitizer, consume any trusted pre-truncation capture
71
+ handle, publish if policy requires it, and return passthrough/exact/compact/
72
+ degraded delivery.
73
+ 4. **DCP observe** — record only the result that will be delivered after the
74
+ preceding stages.
75
+
76
+ The credential firewall's `before_provider_request` hook stays in its existing
77
+ late provider position. `codex-reasoning-fix` remains the last provider-payload
78
+ sanitizer. The coordinator is event-specific; it is not a replacement for
79
+ global module ordering.
80
+
81
+ ### P01-R storeless ordering before any coordinator exists
82
+
83
+ The accepted coordinator above is conditional future **store-backed enforce**
84
+ work. P01-R does not add it merely to clean redundant SDK metadata. The checked
85
+ storeless path keeps independent handlers and the current suite registration
86
+ order:
87
+
88
+ 1. suite-owned result enrichers such as LSP/comment-checker;
89
+ 2. Context Gateway `observe`, which records the original result boundary and
90
+ does not transform it;
91
+ 3. optional `truncation-metadata-normalizer`, which may remove only a proven
92
+ duplicate `details.truncation.content` while preserving the delivered body,
93
+ outcome and every other details field;
94
+ 4. downstream result observers. The R-B contract uses a fake observer with no
95
+ DCP imports so this boundary is not coupled to the current or future DCP
96
+ implementation;
97
+ 5. optional credential-firewall `tool_result` hygiene, which remains the
98
+ security transform for delivered result content/details when that module is
99
+ enabled;
100
+ 6. provider hooks later run in their normal order: credential firewall before
101
+ the final `codex-reasoning-fix` payload sanitizer.
102
+
103
+ This ordering intentionally lets `observe` measure the pre-cleanup metadata
104
+ boundary. The normalizer is not a security boundary and does not restore or
105
+ create source bytes. If firewall hygiene is enabled after it, secrets still
106
+ present in delivered content/remaining details are redacted before JSONL/next
107
+ provider context. When hygiene is disabled, the normalizer must not silently
108
+ pretend those bytes were redacted.
109
+
110
+ The normalizer's tool-name/SDK-shape allowlist is a compatibility guard, not
111
+ trusted provenance. `tool_result` carries no authenticated extension owner, so
112
+ a replacement tool can reuse `Read`/shell/`ast_grep` names. Exact duplicate
113
+ cleanup can remain semantics-preserving while all stronger capability/lifetime
114
+ claims for such a replacement stay **limited**. A future capture/store path
115
+ cannot use this name/shape test as permission to open a path or publish a
116
+ snapshot.
117
+
118
+ R-B does **not** introduce a coordinator because no conflict requiring one is
119
+ present in the storeless path. Existing tests prove independent handler
120
+ composition, exception fail-open behavior and the fact that a later extension
121
+ can reinsert raw content. Therefore hard enforce remains unsupported; if a
122
+ future store-backed stage needs different security ordering, it must implement
123
+ the event-specific coordinator without leaving duplicate independent result
124
+ registrations active.
125
+
126
+ ## P01-R same-name replacement limitation
127
+
128
+ The metadata normalizer can conservatively reject unknown tool names and
129
+ malformed/non-matching truncation metadata, but the generic `tool_result` event
130
+ does not carry a cryptographic or host-owned identity proving which tool
131
+ definition produced a result. A separately loaded replacement that deliberately
132
+ uses a measured name such as `read` plus the exact SDK truncation shape is
133
+ therefore **not independently provenance-certified** by name+shape alone. P01-R
134
+ support claims are limited to the verified suite/SDK tool combinations. Strict
135
+ enforce must not generalise this heuristic to arbitrary replacements without a
136
+ host-owned tool-definition identity seam.
137
+
138
+ ## Failure contract
139
+
140
+ Gateway stage failure must not rely on `throw`. A Gateway failure returns an
141
+ explicit bounded degraded result that preserves the real execution outcome and
142
+ states that archival/retrieval is unavailable. A successful mutation must never
143
+ be reported as "not executed" merely because publication failed.
144
+
145
+ Enrichment or optional session-hygiene failures keep their existing semantics
146
+ until their coordinator adapters are implemented and tested. DCP observation
147
+ must never be allowed to restore a raw source after Gateway shaping.
148
+
149
+ ## Origin binding and reload
150
+
151
+ `toolCallId` is necessary but not sufficient. Gateway execution identity must
152
+ also bind the originating session/workspace and attempt/runtime epoch before an
153
+ await boundary. The active tab or current extension runner after execution is
154
+ not authoritative.
155
+
156
+ The installed SDK currently makes extension reload during an in-flight custom
157
+ tool a limited path: `wrapRegisteredTool()` touches the old runner after the
158
+ tool returns and can turn the original completion into a stale-runner error.
159
+ Gateway strict-enforce support for such cross-reload executions is therefore
160
+ **not claimed**. Wrapper-level capture may preserve permitted bytes as an
161
+ orphaned source, but it must not fabricate a successful delivered result.
162
+
163
+ At the app layer, tab ownership already uses runtime/session/generation guards.
164
+ Gateway bindings should use equivalent host-owned identity rather than the
165
+ currently active tab.
166
+
167
+ ## Capture implications
168
+
169
+ - Built-in `read`: generic `tool_result` is after truncation and there is no
170
+ full-output handle. Treat as limited unless a supported execution wrapper is
171
+ introduced; exact native paging remains the primary path.
172
+ - Built-in `bash`: generic result is bounded, but a trusted `fullOutputPath`
173
+ exists on successful truncated output. Timeout/abort throw paths preserve a
174
+ text prefix/status but lose structured details.
175
+ - `repo_*`: add any future capture seam inside the suite wrapper before
176
+ `truncateOutput`; the generic result hook cannot recover omitted bytes.
177
+ - `ast_grep`: the suite wrapper can capture before truncation and already
178
+ exposes a full-output temp path when truncated.
179
+ - Unknown/custom tools remain limited unless their concrete execution path is
180
+ separately proven.
181
+
182
+ ## Unsupported integrations in the current evidence
183
+
184
+ The current Pix ACP `session/new` implementation consumes `cwd` and `_meta` and
185
+ does not plumb the protocol `mcpServers` field into a local MCP execution path.
186
+ MCP result capture is therefore unsupported, not implicitly covered by the
187
+ generic hook.
188
+
189
+ Browser QA runs in a child Pi process launched with `--no-extensions` plus a
190
+ restricted extension set. The parent Gateway cannot observe Playwright DOM,
191
+ network, or console bytes as parent `tool_result` events. Future integration may
192
+ reuse completed subagent artifacts; it is not a direct browser adapter.
193
+
194
+ ## Storage/platform boundary
195
+
196
+ P00 chooses no durable publication primitive. Current tests run on macOS and do
197
+ not certify directory rename/fsync/no-clobber or cross-process quota behavior on
198
+ Linux or Windows. Until P02 fault/platform tests exist, Gateway must not describe
199
+ its storage as crash-durable across the supported platform matrix.
200
+
201
+ ## Third-party extensions
202
+
203
+ The suite coordinator controls only suite-owned stages. A separately loaded
204
+ third-party extension may register a later `tool_result` modifier and reinsert
205
+ large/raw data. Strict enforce is unsupported for an unverified extension
206
+ combination until a final session/provider-boundary test proves that the raw
207
+ sentinel cannot reappear.
208
+
209
+ ## Evidence
210
+
211
+ - `test/context-gateway/sdk-pipeline.test.ts`
212
+ - `test/context-gateway/capture-contracts.test.ts`
213
+ - `test/context-gateway/lifecycle-contracts.test.ts`
214
+ - `test/context-gateway/provider-serialization.test.ts`
215
+ - `test/context-gateway/p00-capabilities.md`
216
+ - root `tests/tabs-controller.test.ts`, late origin-tab result contract