pi-ui-extend 1.0.40 → 1.0.41
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/app/commands/command-registry.js +2 -2
- package/dist/app/commands/command-session-actions.d.ts +0 -1
- package/dist/app/commands/command-session-actions.js +22 -13
- package/dist/app/icons.d.ts +14 -0
- package/dist/app/icons.js +33 -0
- package/dist/app/rendering/conversation-tool-renderer.js +2 -2
- package/dist/app/rendering/dcp-stats.d.ts +6 -1
- package/dist/app/rendering/dcp-stats.js +214 -46
- package/dist/app/rendering/editor-panels.js +8 -5
- package/dist/app/session/lazy-session-manager.js +12 -1
- package/dist/app/session/tabs-controller.d.ts +2 -5
- package/dist/app/session/tabs-controller.js +12 -21
- package/dist/app/subagents/subagents-model.d.ts +14 -1
- package/dist/app/subagents/subagents-model.js +34 -15
- package/dist/app/types.d.ts +2 -0
- package/dist/bundled-extensions/session-title/config.js +1 -1
- package/dist/markdown-format.js +27 -9
- package/dist/schemas/pi-tools-suite-schema.d.ts +29 -16
- package/dist/schemas/pi-tools-suite-schema.js +46 -31
- package/external/pi-tools-suite/README.md +188 -55
- package/external/pi-tools-suite/docs/browser-qa-subagent.md +31 -21
- package/external/pi-tools-suite/docs/context-gateway-p00-adr.md +216 -0
- package/external/pi-tools-suite/docs/context-gateway-p01n-gate-review.md +122 -0
- package/external/pi-tools-suite/docs/context-gateway-p01n-measurement.md +133 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-ra-evidence.md +111 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-rb-evidence.md +100 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-rc-evidence.md +69 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-rd-evidence.md +100 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-re-evidence.md +74 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-rf-evidence.md +153 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-rg-evidence.md +235 -0
- package/external/pi-tools-suite/docs/subagent-model-pools.md +109 -0
- package/external/pi-tools-suite/package.json +3 -0
- package/external/pi-tools-suite/src/async-subagents/{private-skills → agents}/browser-qa/scripts/browser-qa-runner.mjs +82 -1
- package/external/pi-tools-suite/src/async-subagents/{private-skills/browser-qa/SKILL.md → agents/browser-qa.md} +261 -12
- package/external/pi-tools-suite/src/async-subagents/agents/implement.md +20 -0
- package/external/pi-tools-suite/src/async-subagents/agents/oracle.md +16 -0
- package/external/pi-tools-suite/src/async-subagents/agents/research.md +18 -0
- package/external/pi-tools-suite/src/async-subagents/agents/verify.md +18 -0
- package/external/pi-tools-suite/src/async-subagents/async-subagents.sample.jsonc +27 -243
- package/external/pi-tools-suite/src/async-subagents/commands.ts +6 -2
- package/external/pi-tools-suite/src/async-subagents/core/agent-catalog.ts +41 -0
- package/external/pi-tools-suite/src/async-subagents/core/agent-strategy.ts +13 -93
- package/external/pi-tools-suite/src/async-subagents/core/agents-dir.ts +494 -0
- package/external/pi-tools-suite/src/async-subagents/core/browser-qa.ts +9 -0
- package/external/pi-tools-suite/src/async-subagents/core/config.ts +200 -143
- package/external/pi-tools-suite/src/async-subagents/core/model-fallback.ts +1 -1
- package/external/pi-tools-suite/src/async-subagents/core/model-selection.ts +54 -0
- package/external/pi-tools-suite/src/async-subagents/core/prompt.ts +7 -6
- package/external/pi-tools-suite/src/async-subagents/core/routing.ts +52 -45
- package/external/pi-tools-suite/src/async-subagents/core/spawn.ts +12 -4
- package/external/pi-tools-suite/src/async-subagents/index.ts +11 -1
- package/external/pi-tools-suite/src/async-subagents/lib.ts +6 -2
- package/external/pi-tools-suite/src/async-subagents/tools/spawn.ts +46 -18
- package/external/pi-tools-suite/src/async-subagents/tools/subagents.ts +3 -2
- package/external/pi-tools-suite/src/async-subagents/types.ts +2 -0
- package/external/pi-tools-suite/src/config.ts +1 -1
- package/external/pi-tools-suite/src/context-gateway/accounting.ts +151 -0
- package/external/pi-tools-suite/src/context-gateway/config.ts +111 -0
- package/external/pi-tools-suite/src/context-gateway/index.ts +160 -0
- package/external/pi-tools-suite/src/context-gateway/metadata-normalization.ts +88 -0
- package/external/pi-tools-suite/src/context-gateway/storeless-capabilities.ts +89 -0
- package/external/pi-tools-suite/src/context-gateway/telemetry.ts +429 -0
- package/external/pi-tools-suite/src/context-gateway/test-output-parser.ts +326 -0
- package/external/pi-tools-suite/src/context-gateway/types.ts +152 -0
- package/external/pi-tools-suite/src/dcp/auto-compress-budget.ts +106 -0
- package/external/pi-tools-suite/src/dcp/auto-compress.ts +810 -106
- package/external/pi-tools-suite/src/dcp/commands.ts +64 -139
- package/external/pi-tools-suite/src/dcp/compress-tool.ts +369 -35
- package/external/pi-tools-suite/src/dcp/compression-blocks.ts +510 -64
- package/external/pi-tools-suite/src/dcp/compression-preview.ts +113 -0
- package/external/pi-tools-suite/src/dcp/compression-progress.ts +70 -0
- package/external/pi-tools-suite/src/dcp/config.ts +36 -61
- package/external/pi-tools-suite/src/dcp/conversation-index.ts +421 -0
- package/external/pi-tools-suite/src/dcp/debug-log.ts +7 -5
- package/external/pi-tools-suite/src/dcp/index.ts +617 -203
- package/external/pi-tools-suite/src/dcp/journal.ts +566 -0
- package/external/pi-tools-suite/src/dcp/progress-controller.ts +244 -0
- package/external/pi-tools-suite/src/dcp/prompts.ts +10 -7
- package/external/pi-tools-suite/src/dcp/provider-tool-results.ts +189 -0
- package/external/pi-tools-suite/src/dcp/pruner-candidates.ts +298 -78
- package/external/pi-tools-suite/src/dcp/pruner-compression-blocks.ts +173 -281
- package/external/pi-tools-suite/src/dcp/pruner-emergency.ts +2 -4
- package/external/pi-tools-suite/src/dcp/pruner-message-ids.ts +17 -5
- package/external/pi-tools-suite/src/dcp/pruner-metadata.ts +11 -1
- package/external/pi-tools-suite/src/dcp/pruner-nudge.ts +30 -82
- package/external/pi-tools-suite/src/dcp/pruner-tools.ts +22 -133
- package/external/pi-tools-suite/src/dcp/pruner.ts +18 -33
- package/external/pi-tools-suite/src/dcp/recovery.ts +129 -0
- package/external/pi-tools-suite/src/dcp/shadow-plan.ts +127 -0
- package/external/pi-tools-suite/src/dcp/state-transaction.ts +102 -0
- package/external/pi-tools-suite/src/dcp/state.ts +158 -580
- package/external/pi-tools-suite/src/dcp/ui.ts +1 -0
- package/external/pi-tools-suite/src/default-pi-tools-suite-config.ts +32 -214
- package/external/pi-tools-suite/src/index.ts +9 -0
- package/external/pi-tools-suite/src/model-tools/index.ts +76 -42
- package/external/pi-tools-suite/src/repo-discovery/index.ts +84 -18
- package/external/pi-tools-suite/src/repo-discovery/native-compact.ts +458 -0
- package/external/pi-tools-suite/src/session-recovery/index.ts +189 -43
- package/external/pi-tools-suite/src/tool-descriptions.ts +39 -35
- package/external/pi-tools-suite/src/truncation-metadata-normalizer/index.ts +17 -0
- package/package.json +3 -2
- package/schemas/pi-tools-suite.json +159 -78
- package/external/pi-tools-suite/src/async-subagents/private-skills/browser-qa/references/auth-scaffold-spec.md +0 -78
- package/external/pi-tools-suite/src/async-subagents/private-skills/browser-qa/references/qa-design.md +0 -223
- package/external/pi-tools-suite/src/dcp/state-persistence.ts +0 -195
- /package/external/pi-tools-suite/src/async-subagents/{private-skills/browser-qa/references → agents/browser-qa/examples}/qa-auth.example.jsonc +0 -0
- /package/external/pi-tools-suite/src/async-subagents/{private-skills/browser-qa/references → agents/browser-qa/examples}/qa-flow.example.jsonc +0 -0
- /package/external/pi-tools-suite/src/async-subagents/{private-skills → agents}/browser-qa/vendor/fflate.LICENSE +0 -0
- /package/external/pi-tools-suite/src/async-subagents/{private-skills → agents}/browser-qa/vendor/fflate.mjs +0 -0
|
@@ -100,7 +100,13 @@ DCP settings are stored only under `dcp` in the user shared config file `~/.conf
|
|
|
100
100
|
"nudgeFrequency": 1,
|
|
101
101
|
"iterationNudgeThreshold": 6,
|
|
102
102
|
"nudgeForce": "strong",
|
|
103
|
-
"protectedTools": ["compress", "write", "edit", "subagents"]
|
|
103
|
+
"protectedTools": ["compress", "write", "edit", "subagents"],
|
|
104
|
+
"autoCompress": {
|
|
105
|
+
"enabled": false,
|
|
106
|
+
"patience": 2,
|
|
107
|
+
"summarizerModel": [],
|
|
108
|
+
"timeoutMs": 20000
|
|
109
|
+
}
|
|
104
110
|
},
|
|
105
111
|
"strategies": {
|
|
106
112
|
"emergencyCurrentTurnPruning": {
|
|
@@ -170,7 +176,11 @@ DCP settings are stored only under `dcp` in the user shared config file `~/.conf
|
|
|
170
176
|
|
|
171
177
|
`minContextPercent` / `maxContextPercent` accept legacy fractions (`0.25`), percent strings (`"25%"`), or absolute token counts when Pi knows the current model context window. `minContextLimit` / `maxContextLimit` and `modelMinContextLimits` / `modelMaxContextLimits` are explicit absolute-or-percent aliases. `modelOverrides` and the `modelMin*` / `modelMax*` maps support exact model keys plus `*` / `?` wildcard patterns; matching is applied from generic to specific so exact bare-model matches override bare wildcards, and exact `provider/model` matches override provider wildcards. Array fields are union-merged, so model-specific `protectedTools` extend the defaults instead of replacing them. If `compress.protectUserMessages` is enabled, range compression appends selected user messages verbatim instead of rejecting the range; individual message compression still skips protected raw user messages. Protected tool outputs are copied into summaries for tools protected by name or `protectedFilePatterns`; protected `subagents` result reads also try to include the saved `result.md` artifact when available.
|
|
172
178
|
|
|
173
|
-
`
|
|
179
|
+
`compress.autoCompress.enabled` is `false` by default. When explicitly enabled, its `patience` counts completed correlated main-provider opportunities, not repeated context transforms; `summarizerModel: []` uses the bounded extractive fallback without a model call, while configured models share the single `timeoutMs` summarizer deadline. If that deadline expires, DCP still has a bounded finalization grace to commit the extractive fallback. Auto commit is accepted only when the full projected replacement has positive gain and meets the current budget-recovery target.
|
|
180
|
+
|
|
181
|
+
`strategies.emergencyCurrentTurnPruning` is the default-enabled lossy safety floor for a single unfinished turn that has no normal compression candidate. DCP first emits emergency reminders and offers only safe old same-turn tool-result candidates. After `patience` completed ignored opportunities, or at the model-independent `hardContextPercent`, it replaces eligible oldest result bodies until the estimated provider context reaches `targetContextPercent` or a margin below the model emergency threshold. User messages, configured/protected data, the newest `keepRecentToolPairs`, and results without **completed** provider evidence are never selected. HTTP 2xx alone is not evidence; DCP promotes eligibility only after an unambiguously correlated successful finalized assistant response, and ambiguous retries/interleaving fail closed. The raw session transcript is unchanged. Setting `enabled` to `false` disables same-turn candidates and lossy pruning, but keeps the non-destructive emergency reminder.
|
|
182
|
+
|
|
183
|
+
DCP sidecars are session-private versioned generation envelopes under `<sessionDir>/dcp-state/`. Saves use private atomic files, retain a last-valid `.prev` generation, validate payload/block-graph integrity on load, quarantine corrupt primaries, and use a cross-process exclusive lock rather than silent last-writer-wins. Live paused sessions are not deleted merely for age. Protected subagent artifacts are optional bounded recovery input: reads are async, rooted at the session cwd, reject symlink escapes and oversized files, and never silently truncate a required protected artifact.
|
|
174
184
|
|
|
175
185
|
Set `dcp.debug: true` to write a JSONL debug log of DCP context/prune/compress events to `~/.pi/agent/dcp-debug.jsonl` (override the path with `PI_DCP_DEBUG_LOG`, or enable without config via `PI_DCP_DEBUG=1`); off by default. The log is size-limited and rotated: once it reaches `dcp.debugLog.maxBytes` (default `5242880` = 5 MB) it is renamed to `.1`, older backups shift down (`.1`→`.2`, …) and the oldest beyond `dcp.debugLog.maxBackups` (default `3`, minimum `1`) is dropped; override either with `PI_DCP_DEBUG_MAX_BYTES` / `PI_DCP_DEBUG_MAX_BACKUPS`.
|
|
176
186
|
|
|
@@ -427,21 +437,120 @@ Notes:
|
|
|
427
437
|
|
|
428
438
|
## Async sub-agents
|
|
429
439
|
|
|
430
|
-
|
|
440
|
+
Model selection uses the ordered candidates from each agent's Markdown file,
|
|
441
|
+
filtered by the selected preset's available models and runtime capabilities.
|
|
442
|
+
Explicit task/CLI model overrides bypass the pool. Setting
|
|
443
|
+
`ASYNC_SUBAGENTS_FORCE_CURRENT_MODEL=1` (or
|
|
444
|
+
`PI_SUBAGENTS_FORCE_CURRENT_MODEL=1`) deliberately selects the parent model and
|
|
445
|
+
strips conflicting model arguments; this is not the economical default.
|
|
446
|
+
|
|
447
|
+
The five built-in modes are `research` (read-only evidence and independent
|
|
448
|
+
review), `implement` (bounded code, docs, tests, or UI changes), `verify`
|
|
449
|
+
(run checks and diagnose logs without fixing files), `browser-qa` (trusted
|
|
450
|
+
browser workflow), and `oracle` (deliberate strong second opinion).
|
|
451
|
+
Ordinary workers use economical model candidates; no built-in parent-tier
|
|
452
|
+
rule promotes them to a flagship. Oracle is the exception, not an automatic
|
|
453
|
+
retry for difficult work. Task-specific discipline belongs in the brief.
|
|
454
|
+
|
|
455
|
+
Delegate when a suitable lower-cost worker can handle bounded work or noisy
|
|
456
|
+
intermediate evidence should stay outside the parent context. One sequential
|
|
457
|
+
task can qualify. Keep decisions and integration in the parent; read compact
|
|
458
|
+
results and verify selectively rather than repeating the worker's investigation.
|
|
459
|
+
Do trivial reads/edits directly. Redirect a noisy command to a log without an
|
|
460
|
+
extra LLM when no interpretation is needed. `verify`'s no-edit instruction is
|
|
461
|
+
a behavioral contract, not a read-only filesystem sandbox for its shell.
|
|
462
|
+
|
|
463
|
+
Run `/ultrawork` or `/ulw` for orchestration, `/hyperplan` to pressure-test a
|
|
464
|
+
plan, or set `ULTRAWORK=1` to apply the orchestration prompt to normal inputs.
|
|
465
|
+
`ULTRAWORK_AUTO=1` classifies only the first normal input on non-GPT parents;
|
|
466
|
+
GPT-like parents skip that automatic transform, not ordinary delegation.
|
|
467
|
+
|
|
468
|
+
See [Model pools and migration](docs/subagent-model-pools.md) for the selection
|
|
469
|
+
contract, configuration examples, override rules and legacy compatibility.
|
|
470
|
+
|
|
471
|
+
### Parent-first role selection
|
|
472
|
+
|
|
473
|
+
The parent normally selects an explicit `subagentType` from the effective
|
|
474
|
+
system-prompt catalog, preferring a matching project-local specialist. Valid
|
|
475
|
+
explicit types bypass the LLM router entirely; presets, model selection, tools,
|
|
476
|
+
skills, and role instructions are still applied by the normal config resolver.
|
|
477
|
+
Model/thinking overrides are not substitutes for selecting a role.
|
|
478
|
+
|
|
479
|
+
The router remains enabled as a fallback for omitted types: use it when the role
|
|
480
|
+
is unclear or the user explicitly requests automatic routing. Only omitted
|
|
481
|
+
tasks are classified, in one batch; the parent's explicit choices are preserved.
|
|
482
|
+
Real-browser QA still requires explicit `subagentType: "browser-qa"`.
|
|
483
|
+
|
|
484
|
+
Unknown explicit types and failed/incomplete automatic routing reject the
|
|
485
|
+
**entire spawn batch before run state or child processes are created**. The tool
|
|
486
|
+
returns an error with affected task IDs and available types; the parent should
|
|
487
|
+
correct the roles and resubmit the whole batch. Provider error responses are
|
|
488
|
+
failures too, not successful routes. Configured fallback router models may be
|
|
489
|
+
tried, but missing routes are never silently replaced with `quick`/`defaultType`.
|
|
490
|
+
|
|
491
|
+
With `routing.enabled: false`, every spawn task must supply a valid explicit
|
|
492
|
+
type. `defaultType` remains a preference for genuinely ambiguous LLM choices
|
|
493
|
+
and a legacy config-resolver default, not a spawn error fallback. Existing
|
|
494
|
+
callers relying on an implicit default must now choose a type explicitly.
|
|
495
|
+
|
|
496
|
+
### Project-local agents (`.pi/agents/*.md`)
|
|
497
|
+
|
|
498
|
+
A project can ship sub-agent roles as individual Markdown files in
|
|
499
|
+
`<project>/.pi/agents/`. The first such directory found walking up from the
|
|
500
|
+
session cwd is used; each top-level `*.md` file becomes a `subagentType` named
|
|
501
|
+
after the file. Parent and router see the short `description`; only the child
|
|
502
|
+
receives the Markdown body. A project's ordered `models` are filtered through
|
|
503
|
+
the same active preset pool as built-in agents.
|
|
504
|
+
|
|
505
|
+
```markdown
|
|
506
|
+
---
|
|
507
|
+
description: Use for reviewing this repo's diff — knows the house rules.
|
|
508
|
+
icon: eye
|
|
509
|
+
models:
|
|
510
|
+
- zai/glm-5-turbo
|
|
511
|
+
- openai-codex/gpt-5.6-luna
|
|
512
|
+
thinking: low
|
|
513
|
+
tools: read, grep
|
|
514
|
+
retry:
|
|
515
|
+
maxRetries: 1
|
|
516
|
+
backoffMs: 2000
|
|
517
|
+
---
|
|
518
|
+
|
|
519
|
+
You are this project's staff reviewer. Apply the repo rules from
|
|
520
|
+
AGENTS.md before approving anything; cite file paths first.
|
|
521
|
+
```
|
|
431
522
|
|
|
432
|
-
|
|
523
|
+
- Frontmatter keys: `name` (must match the filename), `description`, `icon`, `models`, `thinking`, `tools`, `isolatedSkills`, `extraArgs`, `promptAppend`, `promptOverride`, `retry`, `maxResultBytes`, `timeoutMs`. Legacy `model`, `fallbackModels`, and `modelByParent` still load. Unknown keys are rejected with an error naming the file.
|
|
524
|
+
- Array fields accept block lists (`- item`), inline arrays (`[a, b]`), or comma-separated strings (`tools: read, grep, bash`). The frontmatter YAML subset is intentionally small: scalars, quoted strings, numbers, comments, lists, and nested maps for `modelByParent`/`retry`. Tabs, block scalars (`|`/`>`), anchors/aliases, and flow maps are hard errors naming file and line.
|
|
525
|
+
- The markdown body becomes `promptAppend`: it is appended after the standard generated prompt (parent objective + task + output format), so the agent still receives its task in the usual structure. Use frontmatter `promptOverride` for full prompt replacement.
|
|
526
|
+
- Precedence: project agent fields override same-named types from user/project JSONC config (field-level; other fields are kept), which in turn override built-ins. Setting `ASYNC_SUBAGENTS_CONFIG` / `PI_SUBAGENTS_CONFIG` disables the directory (explicit config = full control).
|
|
527
|
+
- Files without frontmatter are skipped (a `README.md` there is fine). Definition loading is uncached: edits apply on the next config read/spawn without a restart, and the effective system-prompt catalog is rebuilt at parent-agent start.
|
|
528
|
+
- Bundled roles use the same format internally under `src/async-subagents/agents/*.md`; built-in and project-local profiles therefore share one parser and normalization path instead of maintaining a second role-description schema in TypeScript.
|
|
529
|
+
- `icon` names an agent glyph for UIs that render sub-agent widgets (pix TUI panel, Pix Desktop subagents panel): `agent` (neutral default), `search`, `code`, `flask`, `globe`, `sparkles`, `brain`, `wrench`, `terminal`, `bug`, `book`, `eye`, `zap`, `rocket`. The value is passed through opaquely; unknown names render as the neutral agent icon, and status stays color-coded next to it.
|
|
433
530
|
|
|
434
531
|
### Private browser QA and project auth
|
|
435
532
|
|
|
436
533
|
The built-in `browser-qa` role runs on `zai/glm-5.3-flash`, with
|
|
437
|
-
`openai-codex/gpt-5.6-luna` as its fallback. Its
|
|
438
|
-
|
|
439
|
-
|
|
440
|
-
|
|
441
|
-
|
|
442
|
-
|
|
443
|
-
|
|
444
|
-
|
|
534
|
+
`openai-codex/gpt-5.6-luna` as its fallback. Its complete workflow and detailed
|
|
535
|
+
scenario-design guidance live in the Markdown body of
|
|
536
|
+
`src/async-subagents/agents/browser-qa.md`. The normal profile loader appends
|
|
537
|
+
that body to the QA child's task prompt; the parent and LLM router receive only
|
|
538
|
+
the short `description`. There is no additional QA skill to discover or read.
|
|
539
|
+
|
|
540
|
+
Executable resources live under `src/async-subagents/agents/browser-qa/`.
|
|
541
|
+
The launcher supplies the installed runner's absolute path in
|
|
542
|
+
`PI_BROWSER_QA_RUNNER`; the child invokes `node "$PI_BROWSER_QA_RUNNER"` from
|
|
543
|
+
the delegated project's cwd. This non-secret path is set only for QA children.
|
|
544
|
+
QA always launches with `--no-skills`, even when `isolatedSkills` is empty,
|
|
545
|
+
and skill flags in `extraArgs` cannot bypass that isolation. Explicitly
|
|
546
|
+
configured `isolatedSkills` remain supported as optional additions; no built-in
|
|
547
|
+
QA `--skill` is injected. Other roles retain their normal discovery behavior.
|
|
548
|
+
|
|
549
|
+
Model/thinking/tool-only profile overrides inherit the Markdown workflow.
|
|
550
|
+
An explicit profile `promptAppend` replaces the inherited body under the usual
|
|
551
|
+
field-level merge rules; custom QA instructions must preserve the runner-only,
|
|
552
|
+
credential, target, and evidence contracts. Runner-enforced isolation and
|
|
553
|
+
credential handling remain in code, not in the prompt.
|
|
445
554
|
|
|
446
555
|
Public browser QA does not require an auth profile or `.pi/qa_auth.jsonc`: run it
|
|
447
556
|
with an explicit base URL, whose exact origin becomes the fail-closed allowlist.
|
|
@@ -498,8 +607,8 @@ creating a template. Only an explicit authenticated request may create the
|
|
|
498
607
|
private template. Missing, rejected, or expired selected auth returns
|
|
499
608
|
`QA_AUTH_UPDATE_REQUIRED`, naming only the profile/file/reason needed for the
|
|
500
609
|
parent to ask the user for an update and rerun. See
|
|
501
|
-
`src/async-subagents/
|
|
502
|
-
for complete profile shapes and `
|
|
610
|
+
`src/async-subagents/agents/browser-qa/examples/qa-auth.example.jsonc`
|
|
611
|
+
for complete profile shapes and `examples/qa-flow.example.jsonc` beside it for
|
|
503
612
|
the declarative, non-executable QA action/assertion format.
|
|
504
613
|
|
|
505
614
|
Browser QA videos automatically visualize pointer interactions. Clicks and
|
|
@@ -514,68 +623,92 @@ Async-subagents also injects a lightweight oh-my-openagent-style system-prompt s
|
|
|
514
623
|
|
|
515
624
|
For blind-model screenshot/image inspection, use the main-session `coding-discipline` lookup tool; the bundled default uses vision-capable `zai/glm-5.3-flash`. Async-subagents still supports `imagePaths` on tasks when a broader delegated track genuinely needs images, but it no longer ships a dedicated `vision` role. Dynamic provider capabilities can be missing or stale after switching models, so blind parent models can still be configured explicitly with case-insensitive `*` masks under `asyncSubagents.vision.blindModelPatterns` in `~/.config/pi/pi-tools-suite.jsonc`; do not include `zai/glm-5.3-flash` because it accepts image input. This keeps guidance honest, not a sub-agent role.
|
|
516
625
|
|
|
517
|
-
When
|
|
518
|
-
|
|
519
|
-
|
|
626
|
+
When `subagentType` is omitted, the lightweight role router classifies the task
|
|
627
|
+
using the descriptions. Explicit types bypass it. Unknown types or failed
|
|
628
|
+
routing reject the batch, never substitute `defaultType`. Choosing a worker
|
|
629
|
+
model from its candidate list does not involve an LLM call.
|
|
630
|
+
|
|
631
|
+
### Presets are available-model pools
|
|
632
|
+
|
|
633
|
+
Each agent declares an ordered `models` list in Markdown. A preset declares
|
|
634
|
+
which model references may be used, not another role/model/thinking matrix.
|
|
635
|
+
Selection preserves agent order, intersects it with `preset.models`, checks
|
|
636
|
+
runtime registration/auth availability, and takes the first usable candidate.
|
|
637
|
+
Pool order does not change preference and pool-only models are never appended.
|
|
638
|
+
Without a preset, the full agent list is eligible. Candidate order expresses
|
|
639
|
+
the configured budget preference; runtime does not infer current API prices.
|
|
640
|
+
|
|
641
|
+
Image-bearing tasks and `browser-qa` require confirmed image support; configured
|
|
642
|
+
blind-model masks override runtime image metadata. Remaining eligible models
|
|
643
|
+
form the quota fallback chain, so neither quota history nor image fallback can
|
|
644
|
+
escape the pool. Antigravity account rotation still happens before provider
|
|
645
|
+
fallback. No match, no usable model, or an explicitly empty list rejects the
|
|
646
|
+
whole batch before run directories or child processes are created. A new custom
|
|
647
|
+
agent must supply candidates instead of silently inheriting the parent model.
|
|
648
|
+
|
|
649
|
+
Oracle uses its separate strong-model list and prefers another provider when
|
|
650
|
+
available, but also respects the pool. A same-provider choice is allowed when
|
|
651
|
+
the pool offers no alternative; cross-provider independence is not guaranteed.
|
|
652
|
+
Explicit task/CLI model overrides and `FORCE_CURRENT_MODEL` remain deliberate
|
|
653
|
+
escape hatches and disable automatic model fallback for that task. They do not
|
|
654
|
+
bypass the image-capability check.
|
|
655
|
+
|
|
656
|
+
Define pools in the shared or project `pi-tools-suite.jsonc`. Select a saved
|
|
657
|
+
pool with `/subagent-preset`; use `AGENTS_PRESET=<name>` or
|
|
658
|
+
`/subagent-preset session <name>` for a process-only override and
|
|
659
|
+
`/subagent-preset session-clear` to remove it. The saved selection lives in
|
|
660
|
+
`~/.pi/agent/subagent-preset-selection.json`. `/subagent-preset init` inserts the
|
|
661
|
+
sample only when config is missing. The shipped pools are `cheap` (GLM), `gpt`,
|
|
662
|
+
and `deep` (the retained legacy name for the mixed pool, not worker escalation).
|
|
663
|
+
Initial user config and the sample share one source; descriptions and worker
|
|
664
|
+
model order exist only in the agent files. Existing user files are not rewritten.
|
|
520
665
|
|
|
521
666
|
Example shared async-subagents config section:
|
|
522
667
|
|
|
523
668
|
```jsonc
|
|
524
669
|
{
|
|
525
670
|
"asyncSubagents": {
|
|
526
|
-
"defaultType": "
|
|
671
|
+
"defaultType": "research",
|
|
527
672
|
"routing": {
|
|
528
673
|
"enabled": true,
|
|
529
|
-
"model": "zai/glm-
|
|
674
|
+
"model": "zai/glm-5-turbo",
|
|
530
675
|
"timeoutMs": 12000
|
|
531
676
|
},
|
|
532
677
|
"presets": {
|
|
533
678
|
"cheap": {
|
|
534
|
-
"description": "
|
|
535
|
-
"
|
|
536
|
-
"quick": { "model": "zai/glm-5.3", "thinking": "off" },
|
|
537
|
-
"frontend": { "model": "zai/glm-5.3-flash", "thinking": "medium" },
|
|
538
|
-
"browser-qa": { "model": "zai/glm-5.3-flash", "fallbackModels": ["openai-codex/gpt-5.6-luna"], "thinking": "low" },
|
|
539
|
-
"review": { "model": "zai/glm-5.3", "thinking": "high" }
|
|
540
|
-
}
|
|
679
|
+
"description": "GLM workers with a strong oracle candidate.",
|
|
680
|
+
"models": ["zai/glm-5-turbo", "zai/glm-5.3-flash", "zai/glm-5.3"]
|
|
541
681
|
}
|
|
542
682
|
},
|
|
543
683
|
"types": {
|
|
544
|
-
"
|
|
545
|
-
"
|
|
546
|
-
"thinking": "
|
|
547
|
-
},
|
|
548
|
-
"review": {
|
|
549
|
-
"description": "Use for review/audit of existing code or changes: correctness, security, performance, maintainability, API risks, quality. Do not implement new code.",
|
|
550
|
-
"thinking": "high"
|
|
684
|
+
"research": {
|
|
685
|
+
"models": ["zai/glm-5-turbo", "openai-codex/gpt-5.6-luna"],
|
|
686
|
+
"thinking": "low"
|
|
551
687
|
}
|
|
552
688
|
}
|
|
553
689
|
}
|
|
554
690
|
}
|
|
555
691
|
```
|
|
556
692
|
|
|
557
|
-
###
|
|
558
|
-
|
|
559
|
-
|
|
560
|
-
|
|
561
|
-
|
|
562
|
-
|
|
563
|
-
|
|
564
|
-
|
|
565
|
-
|
|
566
|
-
|
|
567
|
-
|
|
568
|
-
|
|
569
|
-
|
|
570
|
-
|
|
571
|
-
|
|
572
|
-
|
|
573
|
-
|
|
574
|
-
|
|
575
|
-
|
|
576
|
-
```
|
|
577
|
-
|
|
578
|
-
With this config a GLM parent (`zai/*`) spawns the oracle on `gpt-5.6-sol`, a GPT parent (`openai-codex/*`) spawns it on `glm-5.3`, and so on — automatically, at spawn time, with no `task.model` needed. The parent model ref is read from the spawn context (`ctx.model`) and passed into resolution. Pattern matching is case-insensitive `*` glob (same engine as `vision.blindModelPatterns`). When no key matches (or no parent model is known), the role falls back to its static `model` + `fallbackModels`. An explicit `task.model` or `ASYNC_SUBAGENTS_FORCE_CURRENT_MODEL=1` still overrides the match.
|
|
693
|
+
### Legacy configuration compatibility
|
|
694
|
+
|
|
695
|
+
Old built-in role names are no longer implicit aliases. `quick`, `scan`,
|
|
696
|
+
`review`, `deep`, `docs`, `frontend`, and `tests` are valid only when explicitly
|
|
697
|
+
defined as ordinary custom/project types. Old preset per-role keys likewise
|
|
698
|
+
apply only when a type with that exact name exists.
|
|
699
|
+
|
|
700
|
+
Legacy `model` plus `fallbackModels` remains readable. `models` is a complete
|
|
701
|
+
replacement list: it clears inherited legacy model/fallback/parent routing.
|
|
702
|
+
A later old-format model override still replaces the primary candidate, and a
|
|
703
|
+
later `fallbackModels` replaces the remaining candidates; `[]` disables them.
|
|
704
|
+
Old `modelByParent` configs remain supported, but ordinary roles give legacy
|
|
705
|
+
preset models precedence. New built-ins contain no parent-tier escalation maps.
|
|
706
|
+
|
|
707
|
+
When a preset specifies `models`, it is exclusively a pool; inherited legacy
|
|
708
|
+
`model`, `types`, thinking, arguments and timeout overrides do not run. A later
|
|
709
|
+
explicit old-format preset selector can still replace a pool for compatibility.
|
|
710
|
+
Runtime retry structures and the separate role router continue to use the
|
|
711
|
+
term `fallbackModels` for actual fallback-only lists, not agent candidates.
|
|
579
712
|
|
|
580
713
|
Sub-agents run with `--no-session` by default to avoid writing duplicate Pi session JSONL files for fire-and-forget background work. Set `ASYNC_SUBAGENTS_ENABLE_SESSIONS=1` to restore persisted per-agent sessions under each agent's `sessions/` directory; this also registers the session-navigation slash commands (`/sub-open`, `/sub-back`, `/sub-where`) needed for switching and deeper post-mortem navigation.
|
|
581
714
|
|
|
@@ -4,25 +4,32 @@
|
|
|
4
4
|
|
|
5
5
|
Provide a cheap, fast `browser-qa` async-subagent that reproduces browser bugs
|
|
6
6
|
and proves fixes with deterministic assertions plus screenshot, video, and trace
|
|
7
|
-
evidence.
|
|
8
|
-
`openai-codex/gpt-5.6-luna
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
7
|
+
evidence. Its ranked `models` list prefers `zai/glm-5.3-flash`, then
|
|
8
|
+
`openai-codex/gpt-5.6-luna`, filtered by the active preset's model pool and
|
|
9
|
+
confirmed runtime image support.
|
|
10
|
+
|
|
11
|
+
## Inline agent workflow and skill isolation
|
|
12
|
+
|
|
13
|
+
- All operating instructions, flow contracts, scenario-design guidance, and
|
|
14
|
+
auth-scaffolding rules live in `src/async-subagents/agents/browser-qa.md`.
|
|
15
|
+
Its body becomes the QA child's `promptAppend` through the shared agent
|
|
16
|
+
loader. Parent and router catalogs include only its short `description`.
|
|
17
|
+
- Runner code, vendor dependencies/licenses, and optional JSONC examples live
|
|
18
|
+
under `src/async-subagents/agents/browser-qa/`; none is a discoverable skill.
|
|
14
19
|
- Sub-agent processes disable normal extension discovery, then always load the
|
|
15
|
-
suite's model-tools
|
|
16
|
-
|
|
17
|
-
Antigravity-backed role unavailable.
|
|
20
|
+
suite's model-tools extension. They load the Antigravity provider extension
|
|
21
|
+
only when an Antigravity model is explicitly selected.
|
|
18
22
|
- A type profile may declare `isolatedSkills`. Spawning that profile adds
|
|
19
23
|
`--no-skills` followed by one explicit `--skill` per configured path.
|
|
20
|
-
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
24
|
+
- `browser-qa` always disables normal skill discovery and filters skill flags
|
|
25
|
+
out of `extraArgs`, even without configured skills. It no longer injects a
|
|
26
|
+
mandatory QA skill. Explicitly configured skills are optional additions.
|
|
27
|
+
- The launcher sets `PI_BROWSER_QA_RUNNER` to the absolute installed runner path,
|
|
28
|
+
replacing inherited values, and strips it from ordinary child environments.
|
|
29
|
+
QA invokes `node "$PI_BROWSER_QA_RUNNER"` without guessing paths from cwd.
|
|
30
|
+
- Model-only profile overrides inherit the workflow. An explicit profile
|
|
31
|
+
`promptAppend` replaces the body like any other agent profile; it is not an
|
|
32
|
+
immutable security boundary. Runtime protections remain in the runner.
|
|
26
33
|
|
|
27
34
|
## Authentication contract
|
|
28
35
|
|
|
@@ -98,7 +105,7 @@ evidence. The role uses `zai/glm-5.3-flash`, falling back to
|
|
|
98
105
|
|
|
99
106
|
## Reliability and shutdown contract
|
|
100
107
|
|
|
101
|
-
- The built-in `browser-qa` profile has a
|
|
108
|
+
- The built-in `browser-qa` profile has a 300-second wall-clock budget unless
|
|
102
109
|
the caller explicitly supplies a task or spawn timeout. This bounds model
|
|
103
110
|
stalls as well as browser work.
|
|
104
111
|
- The trusted runner has its own bounded lifecycle. Browser launch, context
|
|
@@ -124,11 +131,14 @@ evidence. The role uses `zai/glm-5.3-flash`, falling back to
|
|
|
124
131
|
|
|
125
132
|
## Acceptance criteria
|
|
126
133
|
|
|
127
|
-
1. `browser-qa` resolves to the intended model/fallback and its
|
|
128
|
-
|
|
134
|
+
1. `browser-qa` resolves to the intended model/fallback and its inline Markdown
|
|
135
|
+
workflow, and its isolated child process can register the configured
|
|
129
136
|
model provider.
|
|
130
|
-
2.
|
|
131
|
-
|
|
137
|
+
2. Default QA spawn args contain `--no-skills` but no `--skill`. The child
|
|
138
|
+
receives the full workflow in its initial prompt and can invoke the bundled
|
|
139
|
+
runner through `PI_BROWSER_QA_RUNNER` from an unrelated project directory.
|
|
140
|
+
Optional configured skills still load; ordinary profiles retain existing
|
|
141
|
+
skill discovery behavior and do not receive QA-only environment paths.
|
|
132
142
|
3. Auth profile listing and all error output are redacted; model-authored input
|
|
133
143
|
cannot execute code in the credential-bearing process.
|
|
134
144
|
4. Runner tests cover public execution without an auth file, explicit profile
|
|
@@ -0,0 +1,216 @@
|
|
|
1
|
+
# Context Gateway P00 ADR: tool-result stage ordering
|
|
2
|
+
|
|
3
|
+
<!-- markdownlint-disable MD013 -->
|
|
4
|
+
|
|
5
|
+
> Status: accepted implementation decision for future Context Gateway work.
|
|
6
|
+
> This ADR does not enable Gateway, change user configuration, or certify rollout.
|
|
7
|
+
> Evidence baseline: repository `daa1b06`, installed Pi SDK `0.85.1`, 7 September 2026.
|
|
8
|
+
|
|
9
|
+
## Context
|
|
10
|
+
|
|
11
|
+
The installed extension runner chains `tool_result` handlers in registration
|
|
12
|
+
order. The suite currently has four independent `tool_result` registrations:
|
|
13
|
+
LSP, comment-checker, DCP, and the opt-in credential firewall. Their module load
|
|
14
|
+
order is LSP → comment-checker → DCP → credential firewall.
|
|
15
|
+
|
|
16
|
+
That order is unsuitable for Gateway enforce mode:
|
|
17
|
+
|
|
18
|
+
- mutation diagnostics must be present before Gateway renders a bounded result;
|
|
19
|
+
- session-hygiene redaction, when enabled, must happen before durable capture;
|
|
20
|
+
- Gateway must shape before DCP records the delivered result;
|
|
21
|
+
- changing the global `MODULES` order would also change unrelated lifecycle and
|
|
22
|
+
provider hooks, including the intentionally late provider firewall.
|
|
23
|
+
|
|
24
|
+
P00 tests also prove that throwing from a result handler is fail-open, and that
|
|
25
|
+
an in-flight extension tool can cross an extension reload: the old wrapper then
|
|
26
|
+
fails on its stale runner while the new runner handles the resulting error.
|
|
27
|
+
Therefore current active-runner state is not an origin binding.
|
|
28
|
+
|
|
29
|
+
## Decision
|
|
30
|
+
|
|
31
|
+
For the current **P01-R storeless scope**, keep the existing independent
|
|
32
|
+
`tool_result` chain. A suite-local coordinator is **not** introduced merely to
|
|
33
|
+
make the architecture look uniform. It becomes conditional future work only if
|
|
34
|
+
a store-backed/enforce stage is authorised and a concrete conflict between
|
|
35
|
+
enabled suite-owned result handlers cannot be expressed safely through the
|
|
36
|
+
verified event-specific order.
|
|
37
|
+
|
|
38
|
+
The verified storeless order is:
|
|
39
|
+
|
|
40
|
+
1. LSP and comment-checker result enrichers.
|
|
41
|
+
2. Context Gateway `observe` (passive; no result patch).
|
|
42
|
+
3. Optional `truncation-metadata-normalizer` (details-only duplicate cleanup).
|
|
43
|
+
4. Other downstream result observers/modifiers; the current legacy DCP happens
|
|
44
|
+
to be in this position but P01-R does not depend on DCP state or persistence.
|
|
45
|
+
5. Opt-in credential-firewall session-hygiene redaction.
|
|
46
|
+
|
|
47
|
+
For provider hooks, credential-firewall redaction runs before
|
|
48
|
+
`codex-reasoning-fix`, and `codex-reasoning-fix` remains the final
|
|
49
|
+
`before_provider_request` sanitizer.
|
|
50
|
+
|
|
51
|
+
This order is covered both by the `MODULES` contract and by an executable
|
|
52
|
+
`ExtensionRunner` chain with a DCP-free fake downstream observer. The optional
|
|
53
|
+
normalizer removes only redundant metadata; the later firewall still owns secret
|
|
54
|
+
redaction. With firewall session hygiene disabled, the normalizer does not redact
|
|
55
|
+
or otherwise alter visible content.
|
|
56
|
+
|
|
57
|
+
If a future enforce/store stage is authorised, the previously proposed
|
|
58
|
+
coordinator remains the candidate design: modules participating in it must not
|
|
59
|
+
also register an independent `tool_result` handler, while their non-result hooks
|
|
60
|
+
remain registered normally. That future decision must be revalidated against the
|
|
61
|
+
then-current DCP/session-recovery implementation rather than copied mechanically
|
|
62
|
+
from the legacy chain.
|
|
63
|
+
|
|
64
|
+
The **future enforce coordinator candidate** order is:
|
|
65
|
+
|
|
66
|
+
1. **Enrich** — LSP mutation diagnostics, then comment-checker output.
|
|
67
|
+
2. **Session-hygiene sanitize** — the credential-firewall tool-result transform
|
|
68
|
+
when that existing opt-in policy is enabled.
|
|
69
|
+
3. **Gateway sanitize/capture/render** — apply Gateway's archive-specific
|
|
70
|
+
permitted-snapshot sanitizer, consume any trusted pre-truncation capture
|
|
71
|
+
handle, publish if policy requires it, and return passthrough/exact/compact/
|
|
72
|
+
degraded delivery.
|
|
73
|
+
4. **DCP observe** — record only the result that will be delivered after the
|
|
74
|
+
preceding stages.
|
|
75
|
+
|
|
76
|
+
The credential firewall's `before_provider_request` hook stays in its existing
|
|
77
|
+
late provider position. `codex-reasoning-fix` remains the last provider-payload
|
|
78
|
+
sanitizer. The coordinator is event-specific; it is not a replacement for
|
|
79
|
+
global module ordering.
|
|
80
|
+
|
|
81
|
+
### P01-R storeless ordering before any coordinator exists
|
|
82
|
+
|
|
83
|
+
The accepted coordinator above is conditional future **store-backed enforce**
|
|
84
|
+
work. P01-R does not add it merely to clean redundant SDK metadata. The checked
|
|
85
|
+
storeless path keeps independent handlers and the current suite registration
|
|
86
|
+
order:
|
|
87
|
+
|
|
88
|
+
1. suite-owned result enrichers such as LSP/comment-checker;
|
|
89
|
+
2. Context Gateway `observe`, which records the original result boundary and
|
|
90
|
+
does not transform it;
|
|
91
|
+
3. optional `truncation-metadata-normalizer`, which may remove only a proven
|
|
92
|
+
duplicate `details.truncation.content` while preserving the delivered body,
|
|
93
|
+
outcome and every other details field;
|
|
94
|
+
4. downstream result observers. The R-B contract uses a fake observer with no
|
|
95
|
+
DCP imports so this boundary is not coupled to the current or future DCP
|
|
96
|
+
implementation;
|
|
97
|
+
5. optional credential-firewall `tool_result` hygiene, which remains the
|
|
98
|
+
security transform for delivered result content/details when that module is
|
|
99
|
+
enabled;
|
|
100
|
+
6. provider hooks later run in their normal order: credential firewall before
|
|
101
|
+
the final `codex-reasoning-fix` payload sanitizer.
|
|
102
|
+
|
|
103
|
+
This ordering intentionally lets `observe` measure the pre-cleanup metadata
|
|
104
|
+
boundary. The normalizer is not a security boundary and does not restore or
|
|
105
|
+
create source bytes. If firewall hygiene is enabled after it, secrets still
|
|
106
|
+
present in delivered content/remaining details are redacted before JSONL/next
|
|
107
|
+
provider context. When hygiene is disabled, the normalizer must not silently
|
|
108
|
+
pretend those bytes were redacted.
|
|
109
|
+
|
|
110
|
+
The normalizer's tool-name/SDK-shape allowlist is a compatibility guard, not
|
|
111
|
+
trusted provenance. `tool_result` carries no authenticated extension owner, so
|
|
112
|
+
a replacement tool can reuse `Read`/shell/`ast_grep` names. Exact duplicate
|
|
113
|
+
cleanup can remain semantics-preserving while all stronger capability/lifetime
|
|
114
|
+
claims for such a replacement stay **limited**. A future capture/store path
|
|
115
|
+
cannot use this name/shape test as permission to open a path or publish a
|
|
116
|
+
snapshot.
|
|
117
|
+
|
|
118
|
+
R-B does **not** introduce a coordinator because no conflict requiring one is
|
|
119
|
+
present in the storeless path. Existing tests prove independent handler
|
|
120
|
+
composition, exception fail-open behavior and the fact that a later extension
|
|
121
|
+
can reinsert raw content. Therefore hard enforce remains unsupported; if a
|
|
122
|
+
future store-backed stage needs different security ordering, it must implement
|
|
123
|
+
the event-specific coordinator without leaving duplicate independent result
|
|
124
|
+
registrations active.
|
|
125
|
+
|
|
126
|
+
## P01-R same-name replacement limitation
|
|
127
|
+
|
|
128
|
+
The metadata normalizer can conservatively reject unknown tool names and
|
|
129
|
+
malformed/non-matching truncation metadata, but the generic `tool_result` event
|
|
130
|
+
does not carry a cryptographic or host-owned identity proving which tool
|
|
131
|
+
definition produced a result. A separately loaded replacement that deliberately
|
|
132
|
+
uses a measured name such as `read` plus the exact SDK truncation shape is
|
|
133
|
+
therefore **not independently provenance-certified** by name+shape alone. P01-R
|
|
134
|
+
support claims are limited to the verified suite/SDK tool combinations. Strict
|
|
135
|
+
enforce must not generalise this heuristic to arbitrary replacements without a
|
|
136
|
+
host-owned tool-definition identity seam.
|
|
137
|
+
|
|
138
|
+
## Failure contract
|
|
139
|
+
|
|
140
|
+
Gateway stage failure must not rely on `throw`. A Gateway failure returns an
|
|
141
|
+
explicit bounded degraded result that preserves the real execution outcome and
|
|
142
|
+
states that archival/retrieval is unavailable. A successful mutation must never
|
|
143
|
+
be reported as "not executed" merely because publication failed.
|
|
144
|
+
|
|
145
|
+
Enrichment or optional session-hygiene failures keep their existing semantics
|
|
146
|
+
until their coordinator adapters are implemented and tested. DCP observation
|
|
147
|
+
must never be allowed to restore a raw source after Gateway shaping.
|
|
148
|
+
|
|
149
|
+
## Origin binding and reload
|
|
150
|
+
|
|
151
|
+
`toolCallId` is necessary but not sufficient. Gateway execution identity must
|
|
152
|
+
also bind the originating session/workspace and attempt/runtime epoch before an
|
|
153
|
+
await boundary. The active tab or current extension runner after execution is
|
|
154
|
+
not authoritative.
|
|
155
|
+
|
|
156
|
+
The installed SDK currently makes extension reload during an in-flight custom
|
|
157
|
+
tool a limited path: `wrapRegisteredTool()` touches the old runner after the
|
|
158
|
+
tool returns and can turn the original completion into a stale-runner error.
|
|
159
|
+
Gateway strict-enforce support for such cross-reload executions is therefore
|
|
160
|
+
**not claimed**. Wrapper-level capture may preserve permitted bytes as an
|
|
161
|
+
orphaned source, but it must not fabricate a successful delivered result.
|
|
162
|
+
|
|
163
|
+
At the app layer, tab ownership already uses runtime/session/generation guards.
|
|
164
|
+
Gateway bindings should use equivalent host-owned identity rather than the
|
|
165
|
+
currently active tab.
|
|
166
|
+
|
|
167
|
+
## Capture implications
|
|
168
|
+
|
|
169
|
+
- Built-in `read`: generic `tool_result` is after truncation and there is no
|
|
170
|
+
full-output handle. Treat as limited unless a supported execution wrapper is
|
|
171
|
+
introduced; exact native paging remains the primary path.
|
|
172
|
+
- Built-in `bash`: generic result is bounded, but a trusted `fullOutputPath`
|
|
173
|
+
exists on successful truncated output. Timeout/abort throw paths preserve a
|
|
174
|
+
text prefix/status but lose structured details.
|
|
175
|
+
- `repo_*`: add any future capture seam inside the suite wrapper before
|
|
176
|
+
`truncateOutput`; the generic result hook cannot recover omitted bytes.
|
|
177
|
+
- `ast_grep`: the suite wrapper can capture before truncation and already
|
|
178
|
+
exposes a full-output temp path when truncated.
|
|
179
|
+
- Unknown/custom tools remain limited unless their concrete execution path is
|
|
180
|
+
separately proven.
|
|
181
|
+
|
|
182
|
+
## Unsupported integrations in the current evidence
|
|
183
|
+
|
|
184
|
+
The current Pix ACP `session/new` implementation consumes `cwd` and `_meta` and
|
|
185
|
+
does not plumb the protocol `mcpServers` field into a local MCP execution path.
|
|
186
|
+
MCP result capture is therefore unsupported, not implicitly covered by the
|
|
187
|
+
generic hook.
|
|
188
|
+
|
|
189
|
+
Browser QA runs in a child Pi process launched with `--no-extensions` plus a
|
|
190
|
+
restricted extension set. The parent Gateway cannot observe Playwright DOM,
|
|
191
|
+
network, or console bytes as parent `tool_result` events. Future integration may
|
|
192
|
+
reuse completed subagent artifacts; it is not a direct browser adapter.
|
|
193
|
+
|
|
194
|
+
## Storage/platform boundary
|
|
195
|
+
|
|
196
|
+
P00 chooses no durable publication primitive. Current tests run on macOS and do
|
|
197
|
+
not certify directory rename/fsync/no-clobber or cross-process quota behavior on
|
|
198
|
+
Linux or Windows. Until P02 fault/platform tests exist, Gateway must not describe
|
|
199
|
+
its storage as crash-durable across the supported platform matrix.
|
|
200
|
+
|
|
201
|
+
## Third-party extensions
|
|
202
|
+
|
|
203
|
+
The suite coordinator controls only suite-owned stages. A separately loaded
|
|
204
|
+
third-party extension may register a later `tool_result` modifier and reinsert
|
|
205
|
+
large/raw data. Strict enforce is unsupported for an unverified extension
|
|
206
|
+
combination until a final session/provider-boundary test proves that the raw
|
|
207
|
+
sentinel cannot reappear.
|
|
208
|
+
|
|
209
|
+
## Evidence
|
|
210
|
+
|
|
211
|
+
- `test/context-gateway/sdk-pipeline.test.ts`
|
|
212
|
+
- `test/context-gateway/capture-contracts.test.ts`
|
|
213
|
+
- `test/context-gateway/lifecycle-contracts.test.ts`
|
|
214
|
+
- `test/context-gateway/provider-serialization.test.ts`
|
|
215
|
+
- `test/context-gateway/p00-capabilities.md`
|
|
216
|
+
- root `tests/tabs-controller.test.ts`, late origin-tab result contract
|