pi-ui-extend 1.0.39 → 1.0.41

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (113) hide show
  1. package/README.md +1 -1
  2. package/dist/app/commands/command-registry.js +2 -2
  3. package/dist/app/commands/command-session-actions.d.ts +0 -1
  4. package/dist/app/commands/command-session-actions.js +22 -13
  5. package/dist/app/icons.d.ts +14 -0
  6. package/dist/app/icons.js +33 -0
  7. package/dist/app/rendering/conversation-tool-renderer.js +2 -2
  8. package/dist/app/rendering/dcp-stats.d.ts +6 -1
  9. package/dist/app/rendering/dcp-stats.js +214 -46
  10. package/dist/app/rendering/editor-panels.js +8 -5
  11. package/dist/app/session/lazy-session-manager.js +12 -1
  12. package/dist/app/session/tabs-controller.d.ts +2 -5
  13. package/dist/app/session/tabs-controller.js +12 -21
  14. package/dist/app/subagents/subagents-model.d.ts +14 -1
  15. package/dist/app/subagents/subagents-model.js +34 -15
  16. package/dist/app/types.d.ts +2 -0
  17. package/dist/bundled-extensions/session-title/config.js +1 -1
  18. package/dist/markdown-format.js +27 -9
  19. package/dist/schemas/pi-tools-suite-schema.d.ts +29 -16
  20. package/dist/schemas/pi-tools-suite-schema.js +46 -31
  21. package/external/pi-tools-suite/README.md +392 -52
  22. package/external/pi-tools-suite/docs/browser-qa-subagent.md +31 -21
  23. package/external/pi-tools-suite/docs/context-gateway-p00-adr.md +216 -0
  24. package/external/pi-tools-suite/docs/context-gateway-p01n-gate-review.md +122 -0
  25. package/external/pi-tools-suite/docs/context-gateway-p01n-measurement.md +133 -0
  26. package/external/pi-tools-suite/docs/context-gateway-p01r-ra-evidence.md +111 -0
  27. package/external/pi-tools-suite/docs/context-gateway-p01r-rb-evidence.md +100 -0
  28. package/external/pi-tools-suite/docs/context-gateway-p01r-rc-evidence.md +69 -0
  29. package/external/pi-tools-suite/docs/context-gateway-p01r-rd-evidence.md +100 -0
  30. package/external/pi-tools-suite/docs/context-gateway-p01r-re-evidence.md +74 -0
  31. package/external/pi-tools-suite/docs/context-gateway-p01r-rf-evidence.md +153 -0
  32. package/external/pi-tools-suite/docs/context-gateway-p01r-rg-evidence.md +235 -0
  33. package/external/pi-tools-suite/docs/evals.md +684 -0
  34. package/external/pi-tools-suite/docs/subagent-model-pools.md +109 -0
  35. package/external/pi-tools-suite/package.json +10 -3
  36. package/external/pi-tools-suite/src/async-subagents/{private-skills → agents}/browser-qa/scripts/browser-qa-runner.mjs +82 -1
  37. package/external/pi-tools-suite/src/async-subagents/{private-skills/browser-qa/SKILL.md → agents/browser-qa.md} +261 -12
  38. package/external/pi-tools-suite/src/async-subagents/agents/implement.md +20 -0
  39. package/external/pi-tools-suite/src/async-subagents/agents/oracle.md +16 -0
  40. package/external/pi-tools-suite/src/async-subagents/agents/research.md +18 -0
  41. package/external/pi-tools-suite/src/async-subagents/agents/verify.md +18 -0
  42. package/external/pi-tools-suite/src/async-subagents/async-subagents.sample.jsonc +27 -243
  43. package/external/pi-tools-suite/src/async-subagents/commands.ts +6 -2
  44. package/external/pi-tools-suite/src/async-subagents/core/agent-catalog.ts +41 -0
  45. package/external/pi-tools-suite/src/async-subagents/core/agent-strategy.ts +13 -93
  46. package/external/pi-tools-suite/src/async-subagents/core/agents-dir.ts +494 -0
  47. package/external/pi-tools-suite/src/async-subagents/core/browser-qa.ts +9 -0
  48. package/external/pi-tools-suite/src/async-subagents/core/config.ts +200 -143
  49. package/external/pi-tools-suite/src/async-subagents/core/model-fallback.ts +1 -1
  50. package/external/pi-tools-suite/src/async-subagents/core/model-selection.ts +54 -0
  51. package/external/pi-tools-suite/src/async-subagents/core/prompt.ts +7 -6
  52. package/external/pi-tools-suite/src/async-subagents/core/routing.ts +52 -45
  53. package/external/pi-tools-suite/src/async-subagents/core/spawn.ts +12 -4
  54. package/external/pi-tools-suite/src/async-subagents/index.ts +11 -1
  55. package/external/pi-tools-suite/src/async-subagents/lib.ts +6 -2
  56. package/external/pi-tools-suite/src/async-subagents/tools/spawn.ts +46 -18
  57. package/external/pi-tools-suite/src/async-subagents/tools/subagents.ts +3 -2
  58. package/external/pi-tools-suite/src/async-subagents/types.ts +2 -0
  59. package/external/pi-tools-suite/src/coding-discipline/index.ts +41 -142
  60. package/external/pi-tools-suite/src/config.ts +1 -22
  61. package/external/pi-tools-suite/src/context-gateway/accounting.ts +151 -0
  62. package/external/pi-tools-suite/src/context-gateway/config.ts +111 -0
  63. package/external/pi-tools-suite/src/context-gateway/index.ts +160 -0
  64. package/external/pi-tools-suite/src/context-gateway/metadata-normalization.ts +88 -0
  65. package/external/pi-tools-suite/src/context-gateway/storeless-capabilities.ts +89 -0
  66. package/external/pi-tools-suite/src/context-gateway/telemetry.ts +429 -0
  67. package/external/pi-tools-suite/src/context-gateway/test-output-parser.ts +326 -0
  68. package/external/pi-tools-suite/src/context-gateway/types.ts +152 -0
  69. package/external/pi-tools-suite/src/dcp/auto-compress-budget.ts +106 -0
  70. package/external/pi-tools-suite/src/dcp/auto-compress.ts +810 -106
  71. package/external/pi-tools-suite/src/dcp/commands.ts +64 -139
  72. package/external/pi-tools-suite/src/dcp/compress-tool.ts +369 -35
  73. package/external/pi-tools-suite/src/dcp/compression-blocks.ts +510 -64
  74. package/external/pi-tools-suite/src/dcp/compression-preview.ts +113 -0
  75. package/external/pi-tools-suite/src/dcp/compression-progress.ts +70 -0
  76. package/external/pi-tools-suite/src/dcp/config.ts +36 -61
  77. package/external/pi-tools-suite/src/dcp/conversation-index.ts +421 -0
  78. package/external/pi-tools-suite/src/dcp/debug-log.ts +7 -5
  79. package/external/pi-tools-suite/src/dcp/index.ts +617 -203
  80. package/external/pi-tools-suite/src/dcp/journal.ts +566 -0
  81. package/external/pi-tools-suite/src/dcp/progress-controller.ts +244 -0
  82. package/external/pi-tools-suite/src/dcp/prompts.ts +10 -7
  83. package/external/pi-tools-suite/src/dcp/provider-tool-results.ts +189 -0
  84. package/external/pi-tools-suite/src/dcp/pruner-candidates.ts +298 -78
  85. package/external/pi-tools-suite/src/dcp/pruner-compression-blocks.ts +173 -281
  86. package/external/pi-tools-suite/src/dcp/pruner-emergency.ts +2 -4
  87. package/external/pi-tools-suite/src/dcp/pruner-message-ids.ts +17 -5
  88. package/external/pi-tools-suite/src/dcp/pruner-metadata.ts +11 -1
  89. package/external/pi-tools-suite/src/dcp/pruner-nudge.ts +30 -82
  90. package/external/pi-tools-suite/src/dcp/pruner-tools.ts +22 -133
  91. package/external/pi-tools-suite/src/dcp/pruner.ts +18 -33
  92. package/external/pi-tools-suite/src/dcp/recovery.ts +129 -0
  93. package/external/pi-tools-suite/src/dcp/shadow-plan.ts +127 -0
  94. package/external/pi-tools-suite/src/dcp/state-transaction.ts +102 -0
  95. package/external/pi-tools-suite/src/dcp/state.ts +158 -580
  96. package/external/pi-tools-suite/src/dcp/ui.ts +1 -0
  97. package/external/pi-tools-suite/src/default-pi-tools-suite-config.ts +55 -220
  98. package/external/pi-tools-suite/src/index.ts +9 -0
  99. package/external/pi-tools-suite/src/model-tools/index.ts +76 -42
  100. package/external/pi-tools-suite/src/repo-discovery/index.ts +84 -18
  101. package/external/pi-tools-suite/src/repo-discovery/native-compact.ts +458 -0
  102. package/external/pi-tools-suite/src/session-recovery/index.ts +189 -43
  103. package/external/pi-tools-suite/src/tool-descriptions.ts +43 -38
  104. package/external/pi-tools-suite/src/truncation-metadata-normalizer/index.ts +17 -0
  105. package/package.json +6 -6
  106. package/schemas/pi-tools-suite.json +159 -78
  107. package/external/pi-tools-suite/src/async-subagents/private-skills/browser-qa/references/auth-scaffold-spec.md +0 -78
  108. package/external/pi-tools-suite/src/async-subagents/private-skills/browser-qa/references/qa-design.md +0 -223
  109. package/external/pi-tools-suite/src/dcp/state-persistence.ts +0 -195
  110. /package/external/pi-tools-suite/src/async-subagents/{private-skills/browser-qa/references → agents/browser-qa/examples}/qa-auth.example.jsonc +0 -0
  111. /package/external/pi-tools-suite/src/async-subagents/{private-skills/browser-qa/references → agents/browser-qa/examples}/qa-flow.example.jsonc +0 -0
  112. /package/external/pi-tools-suite/src/async-subagents/{private-skills → agents}/browser-qa/vendor/fflate.LICENSE +0 -0
  113. /package/external/pi-tools-suite/src/async-subagents/{private-skills → agents}/browser-qa/vendor/fflate.mjs +0 -0
@@ -100,7 +100,13 @@ DCP settings are stored only under `dcp` in the user shared config file `~/.conf
100
100
  "nudgeFrequency": 1,
101
101
  "iterationNudgeThreshold": 6,
102
102
  "nudgeForce": "strong",
103
- "protectedTools": ["compress", "write", "edit", "subagents"]
103
+ "protectedTools": ["compress", "write", "edit", "subagents"],
104
+ "autoCompress": {
105
+ "enabled": false,
106
+ "patience": 2,
107
+ "summarizerModel": [],
108
+ "timeoutMs": 20000
109
+ }
104
110
  },
105
111
  "strategies": {
106
112
  "emergencyCurrentTurnPruning": {
@@ -170,7 +176,11 @@ DCP settings are stored only under `dcp` in the user shared config file `~/.conf
170
176
 
171
177
  `minContextPercent` / `maxContextPercent` accept legacy fractions (`0.25`), percent strings (`"25%"`), or absolute token counts when Pi knows the current model context window. `minContextLimit` / `maxContextLimit` and `modelMinContextLimits` / `modelMaxContextLimits` are explicit absolute-or-percent aliases. `modelOverrides` and the `modelMin*` / `modelMax*` maps support exact model keys plus `*` / `?` wildcard patterns; matching is applied from generic to specific so exact bare-model matches override bare wildcards, and exact `provider/model` matches override provider wildcards. Array fields are union-merged, so model-specific `protectedTools` extend the defaults instead of replacing them. If `compress.protectUserMessages` is enabled, range compression appends selected user messages verbatim instead of rejecting the range; individual message compression still skips protected raw user messages. Protected tool outputs are copied into summaries for tools protected by name or `protectedFilePatterns`; protected `subagents` result reads also try to include the saved `result.md` artifact when available.
172
178
 
173
- `strategies.emergencyCurrentTurnPruning` is the default-enabled lossy safety floor for a single unfinished turn that has no normal compression candidate. DCP first emits emergency reminders and offers only safe old same-turn tool-result candidates. After `patience` ignored reminders, or at the model-independent `hardContextPercent`, it replaces eligible oldest result bodies until the estimated provider context reaches `targetContextPercent` or a margin below the model emergency threshold. User messages, configured/protected data, the newest `keepRecentToolPairs`, and results not present in an accepted provider request are never selected. The raw session transcript is unchanged. Setting `enabled` to `false` disables same-turn candidates and lossy pruning, but keeps the non-destructive emergency reminder.
179
+ `compress.autoCompress.enabled` is `false` by default. When explicitly enabled, its `patience` counts completed correlated main-provider opportunities, not repeated context transforms; `summarizerModel: []` uses the bounded extractive fallback without a model call, while configured models share the single `timeoutMs` summarizer deadline. If that deadline expires, DCP still has a bounded finalization grace to commit the extractive fallback. Auto commit is accepted only when the full projected replacement has positive gain and meets the current budget-recovery target.
180
+
181
+ `strategies.emergencyCurrentTurnPruning` is the default-enabled lossy safety floor for a single unfinished turn that has no normal compression candidate. DCP first emits emergency reminders and offers only safe old same-turn tool-result candidates. After `patience` completed ignored opportunities, or at the model-independent `hardContextPercent`, it replaces eligible oldest result bodies until the estimated provider context reaches `targetContextPercent` or a margin below the model emergency threshold. User messages, configured/protected data, the newest `keepRecentToolPairs`, and results without **completed** provider evidence are never selected. HTTP 2xx alone is not evidence; DCP promotes eligibility only after an unambiguously correlated successful finalized assistant response, and ambiguous retries/interleaving fail closed. The raw session transcript is unchanged. Setting `enabled` to `false` disables same-turn candidates and lossy pruning, but keeps the non-destructive emergency reminder.
182
+
183
+ DCP sidecars are session-private versioned generation envelopes under `<sessionDir>/dcp-state/`. Saves use private atomic files, retain a last-valid `.prev` generation, validate payload/block-graph integrity on load, quarantine corrupt primaries, and use a cross-process exclusive lock rather than silent last-writer-wins. Live paused sessions are not deleted merely for age. Protected subagent artifacts are optional bounded recovery input: reads are async, rooted at the session cwd, reject symlink escapes and oversized files, and never silently truncate a required protected artifact.
174
184
 
175
185
  Set `dcp.debug: true` to write a JSONL debug log of DCP context/prune/compress events to `~/.pi/agent/dcp-debug.jsonl` (override the path with `PI_DCP_DEBUG_LOG`, or enable without config via `PI_DCP_DEBUG=1`); off by default. The log is size-limited and rotated: once it reaches `dcp.debugLog.maxBytes` (default `5242880` = 5 MB) it is renamed to `.1`, older backups shift down (`.1`→`.2`, …) and the oldest beyond `dcp.debugLog.maxBackups` (default `3`, minimum `1`) is dropped; override either with `PI_DCP_DEBUG_MAX_BYTES` / `PI_DCP_DEBUG_MAX_BACKUPS`.
176
186
 
@@ -188,9 +198,28 @@ Install the language servers used by the bundled example config. The commands be
188
198
  # TypeScript / JavaScript
189
199
  npm install -g typescript typescript-language-server
190
200
 
201
+ # Svelte
202
+ npm install -g svelte-language-server
203
+
204
+ # Vue
205
+ npm install -g @vue/language-server
206
+
191
207
  # Python
192
208
  python3 -m pip install --user python-lsp-server
193
209
 
210
+ # Go
211
+ go install golang.org/x/tools/gopls@latest
212
+
213
+ # C / C++ (clangd)
214
+ brew install llvm
215
+ # or use your distro's clangd package: apt install clangd, dnf install clang-tools-extra, ...
216
+
217
+ # Lua
218
+ brew install lua-language-server
219
+
220
+ # Bash
221
+ npm install -g bash-language-server
222
+
194
223
  # C# / Unity
195
224
  dotnet tool install -g Microsoft.CodeAnalysis.LanguageServer
196
225
 
@@ -252,23 +281,276 @@ Minimal shared config shape:
252
281
 
253
282
  Project-local overrides can be added in `.pi/pi-tools-suite.jsonc`; pi-tools-suite asks for trust before using project-local LSP binaries.
254
283
 
284
+ ### Popular language-server examples
285
+
286
+ Copy the entries you need into `lsp.servers` of the shared config. Values mirror the commented templates shipped in the generated config file; where they disagree, the generated file is authoritative. Both diagnostics modes are on by default: servers that support pull diagnostics (`textDocument/diagnostic`) are queried directly, and push diagnostics (`publishDiagnostics`) are awaited for every server — `pullDiagnostics: false` / `waitForPublishDiagnostics: false` disable either side explicitly. Servers start lazily: one spawns only after a mutating tool (Edit/Write/ast-grep/apply_patch) touches a file matching `include`, and diagnostics land in that tool's result.
287
+
288
+ ```jsonc
289
+ {
290
+ "lsp": {
291
+ "servers": [
292
+ // Svelte (verified with svelte-language-server): compiler + embedded TS/JS diagnostics
293
+ {
294
+ "id": "svelte",
295
+ "include": ["**/*.svelte"],
296
+ "exclude": ["**/node_modules/**"],
297
+ "rootMarkers": ["svelte.config.js", "package.json"],
298
+ "bin": "svelteserver",
299
+ "args": ["--stdio"],
300
+ "startupTimeoutMs": 30000,
301
+ "diagnosticsWaitMs": 8000,
302
+ "languageIdByExtension": { ".svelte": "svelte" }
303
+ },
304
+ // Vue (Volar)
305
+ {
306
+ "id": "vue",
307
+ "include": ["**/*.vue"],
308
+ "exclude": ["**/node_modules/**"],
309
+ "rootMarkers": ["package.json"],
310
+ "bin": "vue-language-server",
311
+ "args": ["--stdio"],
312
+ "startupTimeoutMs": 30000,
313
+ "diagnosticsWaitMs": 8000,
314
+ "languageIdByExtension": { ".vue": "vue" }
315
+ },
316
+ // Python (python-lsp-server)
317
+ {
318
+ "id": "python",
319
+ "include": ["**/*.py", "**/*.pyi"],
320
+ "exclude": ["**/.git/**", "**/node_modules/**", "**/__pycache__/**", "**/.venv/**", "**/venv/**", "**/.tox/**", "**/.mypy_cache/**", "**/.ruff_cache/**"],
321
+ "rootMarkers": ["pyproject.toml", "setup.py", "setup.cfg", "requirements.txt", "Pipfile", "poetry.lock", ".git"],
322
+ "bin": "pylsp",
323
+ "args": [],
324
+ "languageIdByExtension": { ".py": "python", ".pyi": "python" }
325
+ },
326
+ // Go (gopls, push diagnostics)
327
+ {
328
+ "id": "go",
329
+ "include": ["**/*.go"],
330
+ "exclude": ["**/.git/**", "**/vendor/**"],
331
+ "rootMarkers": ["go.mod", ".git"],
332
+ "bin": "gopls",
333
+ "args": [],
334
+ "startupTimeoutMs": 20000,
335
+ "diagnosticsWaitMs": 8000,
336
+ "languageIdByExtension": { ".go": "go" }
337
+ },
338
+ // Rust (rust-analyzer, push diagnostics)
339
+ {
340
+ "id": "rust",
341
+ "include": ["**/*.rs"],
342
+ "exclude": ["**/.git/**", "**/node_modules/**", "**/target/**"],
343
+ "rootMarkers": ["Cargo.toml", "rust-project.json", ".git"],
344
+ "bin": "rust-analyzer",
345
+ "args": [],
346
+ "startupTimeoutMs": 20000,
347
+ "diagnosticsWaitMs": 20000,
348
+ "pullDiagnostics": false,
349
+ "waitForPublishDiagnostics": true,
350
+ "languageIdByExtension": { ".rs": "rust" }
351
+ },
352
+ // C / C++ (clangd, push diagnostics)
353
+ {
354
+ "id": "clangd",
355
+ "include": ["**/*.c", "**/*.cc", "**/*.cpp", "**/*.cxx", "**/*.h", "**/*.hh", "**/*.hpp"],
356
+ "exclude": ["**/.git/**", "**/node_modules/**"],
357
+ "rootMarkers": ["compile_commands.json", "CMakeLists.txt", "Makefile", ".clang-format", ".git"],
358
+ "bin": "clangd",
359
+ "args": [],
360
+ "startupTimeoutMs": 20000,
361
+ "diagnosticsWaitMs": 8000,
362
+ "languageIdByExtension": {
363
+ ".c": "c", ".cc": "cpp", ".cpp": "cpp", ".cxx": "cpp",
364
+ ".h": "c", ".hh": "cpp", ".hpp": "cpp"
365
+ }
366
+ },
367
+ // C# / Unity (Roslyn)
368
+ {
369
+ "id": "csharp",
370
+ "include": ["**/*.cs", "**/*.csx"],
371
+ "exclude": ["**/.git/**", "**/node_modules/**", "**/bin/**", "**/obj/**", "**/.vs/**", "**/Library/**", "**/Temp/**", "**/Logs/**"],
372
+ "rootMarkers": ["*.sln", "*.csproj", "global.json", "Directory.Build.props", "Directory.Packages.props", "Packages/manifest.json", "ProjectSettings/ProjectVersion.txt", ".git"],
373
+ "bin": "~/.dotnet/tools/roslyn-language-server",
374
+ "args": ["--stdio", "--autoLoadProjects", "--logLevel", "Error"],
375
+ "startupTimeoutMs": 30000,
376
+ "diagnosticsWaitMs": 15000,
377
+ "languageIdByExtension": { ".cs": "csharp", ".csx": "csharp" }
378
+ },
379
+ // Ruby
380
+ {
381
+ "id": "ruby",
382
+ "include": ["**/*.rb", "**/*.rake", "**/Gemfile", "**/Rakefile", "**/*.gemspec"],
383
+ "exclude": ["**/.git/**", "**/node_modules/**", "**/vendor/bundle/**", "**/.bundle/**", "**/tmp/**", "**/log/**"],
384
+ "rootMarkers": ["Gemfile.lock", "*.gemspec", "Rakefile", ".ruby-version", ".git"],
385
+ "bin": "ruby-lsp",
386
+ "args": [],
387
+ "startupTimeoutMs": 60000,
388
+ "diagnosticsWaitMs": 10000,
389
+ "languageIdByExtension": { ".rb": "ruby", ".rake": "ruby", ".gemspec": "ruby" }
390
+ },
391
+ // Lua (lua-language-server)
392
+ {
393
+ "id": "lua",
394
+ "include": ["**/*.lua"],
395
+ "exclude": ["**/.git/**", "**/node_modules/**"],
396
+ "rootMarkers": [".luarc.json", ".git"],
397
+ "bin": "lua-language-server",
398
+ "args": [],
399
+ "startupTimeoutMs": 30000,
400
+ "diagnosticsWaitMs": 6000,
401
+ "languageIdByExtension": { ".lua": "lua" }
402
+ },
403
+ // Bash
404
+ {
405
+ "id": "bash",
406
+ "include": ["**/*.sh", "**/*.bash"],
407
+ "exclude": ["**/.git/**", "**/node_modules/**"],
408
+ "rootMarkers": [".git"],
409
+ "bin": "bash-language-server",
410
+ "args": ["start"],
411
+ "startupTimeoutMs": 15000,
412
+ "diagnosticsWaitMs": 5000,
413
+ "languageIdByExtension": { ".sh": "shellscript", ".bash": "shellscript" }
414
+ },
415
+ // Markdown (link validation settings ship in the generated config template)
416
+ {
417
+ "id": "markdown",
418
+ "include": ["**/*.md", "**/*.markdown", "**/*.mdown", "**/*.mkd", "**/*.mmd"],
419
+ "exclude": ["**/.git/**", "**/node_modules/**"],
420
+ "rootMarkers": [".git", "package.json", "README.md"],
421
+ "bin": "vscode-markdown-language-server",
422
+ "args": ["--stdio"],
423
+ "startupTimeoutMs": 15000,
424
+ "diagnosticsWaitMs": 5000,
425
+ "languageIdByExtension": { ".md": "markdown", ".markdown": "markdown", ".mdown": "markdown", ".mkd": "markdown", ".mmd": "markdown" }
426
+ }
427
+ ]
428
+ }
429
+ }
430
+ ```
431
+
432
+ Notes:
433
+
434
+ - Svelte resolves `svelte` and `typescript` from the workspace `node_modules`, so project-local versions win; `.svelte.js`/`.svelte.ts` runes modules are not covered because their extensions collide with the TypeScript server.
435
+ - Vue requires `@vue/language-server` v2+ (the `vue-language-server` binary) plus the workspace's `vue` package for template type-checking.
436
+ - The full commented templates (including GDScript via a headless Godot wrapper and the complete Markdown link-validation `settings`) are written to the shared config file on first run.
437
+
255
438
  ## Async sub-agents
256
439
 
257
- Sub-agent model routing normally follows task overrides, subagent type config, then `ASYNC_SUBAGENTS_MODEL` / `PI_SUBAGENTS_MODEL` fallbacks. Set `ASYNC_SUBAGENTS_FORCE_CURRENT_MODEL=1` (or `PI_SUBAGENTS_FORCE_CURRENT_MODEL=1`) to ignore task/config/env model choices and launch every sub-agent with the current parent session model. When this flag is enabled, any `--model` entries in sub-agent extra args are stripped so they cannot override the current model.
440
+ Model selection uses the ordered candidates from each agent's Markdown file,
441
+ filtered by the selected preset's available models and runtime capabilities.
442
+ Explicit task/CLI model overrides bypass the pool. Setting
443
+ `ASYNC_SUBAGENTS_FORCE_CURRENT_MODEL=1` (or
444
+ `PI_SUBAGENTS_FORCE_CURRENT_MODEL=1`) deliberately selects the parent model and
445
+ strips conflicting model arguments; this is not the economical default.
446
+
447
+ The five built-in modes are `research` (read-only evidence and independent
448
+ review), `implement` (bounded code, docs, tests, or UI changes), `verify`
449
+ (run checks and diagnose logs without fixing files), `browser-qa` (trusted
450
+ browser workflow), and `oracle` (deliberate strong second opinion).
451
+ Ordinary workers use economical model candidates; no built-in parent-tier
452
+ rule promotes them to a flagship. Oracle is the exception, not an automatic
453
+ retry for difficult work. Task-specific discipline belongs in the brief.
454
+
455
+ Delegate when a suitable lower-cost worker can handle bounded work or noisy
456
+ intermediate evidence should stay outside the parent context. One sequential
457
+ task can qualify. Keep decisions and integration in the parent; read compact
458
+ results and verify selectively rather than repeating the worker's investigation.
459
+ Do trivial reads/edits directly. Redirect a noisy command to a log without an
460
+ extra LLM when no interpretation is needed. `verify`'s no-edit instruction is
461
+ a behavioral contract, not a read-only filesystem sandbox for its shell.
462
+
463
+ Run `/ultrawork` or `/ulw` for orchestration, `/hyperplan` to pressure-test a
464
+ plan, or set `ULTRAWORK=1` to apply the orchestration prompt to normal inputs.
465
+ `ULTRAWORK_AUTO=1` classifies only the first normal input on non-GPT parents;
466
+ GPT-like parents skip that automatic transform, not ordinary delegation.
467
+
468
+ See [Model pools and migration](docs/subagent-model-pools.md) for the selection
469
+ contract, configuration examples, override rules and legacy compatibility.
470
+
471
+ ### Parent-first role selection
472
+
473
+ The parent normally selects an explicit `subagentType` from the effective
474
+ system-prompt catalog, preferring a matching project-local specialist. Valid
475
+ explicit types bypass the LLM router entirely; presets, model selection, tools,
476
+ skills, and role instructions are still applied by the normal config resolver.
477
+ Model/thinking overrides are not substitutes for selecting a role.
478
+
479
+ The router remains enabled as a fallback for omitted types: use it when the role
480
+ is unclear or the user explicitly requests automatic routing. Only omitted
481
+ tasks are classified, in one batch; the parent's explicit choices are preserved.
482
+ Real-browser QA still requires explicit `subagentType: "browser-qa"`.
483
+
484
+ Unknown explicit types and failed/incomplete automatic routing reject the
485
+ **entire spawn batch before run state or child processes are created**. The tool
486
+ returns an error with affected task IDs and available types; the parent should
487
+ correct the roles and resubmit the whole batch. Provider error responses are
488
+ failures too, not successful routes. Configured fallback router models may be
489
+ tried, but missing routes are never silently replaced with `quick`/`defaultType`.
490
+
491
+ With `routing.enabled: false`, every spawn task must supply a valid explicit
492
+ type. `defaultType` remains a preference for genuinely ambiguous LLM choices
493
+ and a legacy config-resolver default, not a spawn error fallback. Existing
494
+ callers relying on an implicit default must now choose a type explicitly.
495
+
496
+ ### Project-local agents (`.pi/agents/*.md`)
497
+
498
+ A project can ship sub-agent roles as individual Markdown files in
499
+ `<project>/.pi/agents/`. The first such directory found walking up from the
500
+ session cwd is used; each top-level `*.md` file becomes a `subagentType` named
501
+ after the file. Parent and router see the short `description`; only the child
502
+ receives the Markdown body. A project's ordered `models` are filtered through
503
+ the same active preset pool as built-in agents.
504
+
505
+ ```markdown
506
+ ---
507
+ description: Use for reviewing this repo's diff — knows the house rules.
508
+ icon: eye
509
+ models:
510
+ - zai/glm-5-turbo
511
+ - openai-codex/gpt-5.6-luna
512
+ thinking: low
513
+ tools: read, grep
514
+ retry:
515
+ maxRetries: 1
516
+ backoffMs: 2000
517
+ ---
518
+
519
+ You are this project's staff reviewer. Apply the repo rules from
520
+ AGENTS.md before approving anything; cite file paths first.
521
+ ```
258
522
 
259
- For an oh-my-openagent-style workflow, run `/ultrawork` or `/ulw` to ask the parent agent to split broad work into configured async-subagents roles (`quick`, `scan`, `research`, `docs`, `frontend`, `browser-qa`, `implement`, `tests`, `review`, `deep`, `oracle`). Set `ULTRAWORK=1` before launching Pi to apply that compact routing prompt to normal non-slash user inputs automatically. Set `ULTRAWORK_AUTO=1` to ask the lightweight router model to classify only the first normal user input on non-GPT parent models: clear broad/parallel work is transformed into ultrawork, vague potentially-complex work gets a soft delegation hint, and narrow work is left unchanged. GPT-like parent models skip only this automatic transform; they can still use `/ultrawork` and `subagents` normally. `frontend` is for UI/UX, styling, layout, responsive behavior, and visual component polish; `browser-qa` reproduces browser bugs and proves fixes with deterministic assertions plus screenshot/video/trace evidence; `review` covers security/performance/audit tracks; `implement` covers refactors; `deep` covers debugging/root-cause; `oracle` is for sparse cross-provider second opinions on high-stakes uncertainty. Run `/hyperplan` to pressure-test a plan before implementation.
523
+ - Frontmatter keys: `name` (must match the filename), `description`, `icon`, `models`, `thinking`, `tools`, `isolatedSkills`, `extraArgs`, `promptAppend`, `promptOverride`, `retry`, `maxResultBytes`, `timeoutMs`. Legacy `model`, `fallbackModels`, and `modelByParent` still load. Unknown keys are rejected with an error naming the file.
524
+ - Array fields accept block lists (`- item`), inline arrays (`[a, b]`), or comma-separated strings (`tools: read, grep, bash`). The frontmatter YAML subset is intentionally small: scalars, quoted strings, numbers, comments, lists, and nested maps for `modelByParent`/`retry`. Tabs, block scalars (`|`/`>`), anchors/aliases, and flow maps are hard errors naming file and line.
525
+ - The markdown body becomes `promptAppend`: it is appended after the standard generated prompt (parent objective + task + output format), so the agent still receives its task in the usual structure. Use frontmatter `promptOverride` for full prompt replacement.
526
+ - Precedence: project agent fields override same-named types from user/project JSONC config (field-level; other fields are kept), which in turn override built-ins. Setting `ASYNC_SUBAGENTS_CONFIG` / `PI_SUBAGENTS_CONFIG` disables the directory (explicit config = full control).
527
+ - Files without frontmatter are skipped (a `README.md` there is fine). Definition loading is uncached: edits apply on the next config read/spawn without a restart, and the effective system-prompt catalog is rebuilt at parent-agent start.
528
+ - Bundled roles use the same format internally under `src/async-subagents/agents/*.md`; built-in and project-local profiles therefore share one parser and normalization path instead of maintaining a second role-description schema in TypeScript.
529
+ - `icon` names an agent glyph for UIs that render sub-agent widgets (pix TUI panel, Pix Desktop subagents panel): `agent` (neutral default), `search`, `code`, `flask`, `globe`, `sparkles`, `brain`, `wrench`, `terminal`, `bug`, `book`, `eye`, `zap`, `rocket`. The value is passed through opaquely; unknown names render as the neutral agent icon, and status stays color-coded next to it.
260
530
 
261
531
  ### Private browser QA and project auth
262
532
 
263
533
  The built-in `browser-qa` role runs on `zai/glm-5.3-flash`, with
264
- `openai-codex/gpt-5.6-luna` as its fallback. Its browser
265
- workflow is an explicit private skill under `src/async-subagents/private-skills/`,
266
- outside normal Pi skill discovery. The role's first-class `isolatedSkills` setting launches the child with
267
- `--no-skills` plus one self-contained private workflow. It bundles the relevant
268
- scenario-design, locator, waiting, assertion, evidence, and cleanup guidance next
269
- to its trusted runner, so browser QA does not depend on a separately installed
270
- skill or CLI. The private workflow remains mandatory when configuration appends
271
- other isolated skills; parent and ordinary sub-agent sessions do not discover it.
534
+ `openai-codex/gpt-5.6-luna` as its fallback. Its complete workflow and detailed
535
+ scenario-design guidance live in the Markdown body of
536
+ `src/async-subagents/agents/browser-qa.md`. The normal profile loader appends
537
+ that body to the QA child's task prompt; the parent and LLM router receive only
538
+ the short `description`. There is no additional QA skill to discover or read.
539
+
540
+ Executable resources live under `src/async-subagents/agents/browser-qa/`.
541
+ The launcher supplies the installed runner's absolute path in
542
+ `PI_BROWSER_QA_RUNNER`; the child invokes `node "$PI_BROWSER_QA_RUNNER"` from
543
+ the delegated project's cwd. This non-secret path is set only for QA children.
544
+ QA always launches with `--no-skills`, even when `isolatedSkills` is empty,
545
+ and skill flags in `extraArgs` cannot bypass that isolation. Explicitly
546
+ configured `isolatedSkills` remain supported as optional additions; no built-in
547
+ QA `--skill` is injected. Other roles retain their normal discovery behavior.
548
+
549
+ Model/thinking/tool-only profile overrides inherit the Markdown workflow.
550
+ An explicit profile `promptAppend` replaces the inherited body under the usual
551
+ field-level merge rules; custom QA instructions must preserve the runner-only,
552
+ credential, target, and evidence contracts. Runner-enforced isolation and
553
+ credential handling remain in code, not in the prompt.
272
554
 
273
555
  Public browser QA does not require an auth profile or `.pi/qa_auth.jsonc`: run it
274
556
  with an explicit base URL, whose exact origin becomes the fail-closed allowlist.
@@ -325,8 +607,8 @@ creating a template. Only an explicit authenticated request may create the
325
607
  private template. Missing, rejected, or expired selected auth returns
326
608
  `QA_AUTH_UPDATE_REQUIRED`, naming only the profile/file/reason needed for the
327
609
  parent to ask the user for an update and rerun. See
328
- `src/async-subagents/private-skills/browser-qa/references/qa-auth.example.jsonc`
329
- for complete profile shapes and `references/qa-flow.example.jsonc` beside it for
610
+ `src/async-subagents/agents/browser-qa/examples/qa-auth.example.jsonc`
611
+ for complete profile shapes and `examples/qa-flow.example.jsonc` beside it for
330
612
  the declarative, non-executable QA action/assertion format.
331
613
 
332
614
  Browser QA videos automatically visualize pointer interactions. Clicks and
@@ -341,68 +623,92 @@ Async-subagents also injects a lightweight oh-my-openagent-style system-prompt s
341
623
 
342
624
  For blind-model screenshot/image inspection, use the main-session `coding-discipline` lookup tool; the bundled default uses vision-capable `zai/glm-5.3-flash`. Async-subagents still supports `imagePaths` on tasks when a broader delegated track genuinely needs images, but it no longer ships a dedicated `vision` role. Dynamic provider capabilities can be missing or stale after switching models, so blind parent models can still be configured explicitly with case-insensitive `*` masks under `asyncSubagents.vision.blindModelPatterns` in `~/.config/pi/pi-tools-suite.jsonc`; do not include `zai/glm-5.3-flash` because it accepts image input. This keeps guidance honest, not a sub-agent role.
343
625
 
344
- When a task omits `subagentType`, async-subagents asks a lightweight router model to choose one configured type for each task from the task text/scope and the `types.<name>.description` metadata. Explicit task `subagentType` still wins. Keep type descriptions short, literal, and distinct because they are inserted into the router prompt for a small model. Router settings live under `asyncSubagents.routing` (`enabled`, `model`, `maxTaskChars`, `maxTokens`, `maxRetries`, `timeoutMs`, `debug`); the default router model is `zai/glm-4.5-air`. If the router is disabled, unavailable, aborted, or returns invalid JSON, omitted types fall back to `defaultType`.
345
-
346
- Define optional `presets` under `asyncSubagents` in `~/.config/pi/pi-tools-suite.jsonc`, `$PI_CONFIG_DIR/pi-tools-suite.jsonc`, or project `.pi/pi-tools-suite.jsonc`, then use `/subagent-preset` or `/subagent-preset-config` to pick one persistent active preset for future spawns across all sessions. Set `AGENTS_PRESET=<name>` before launching Pi to override the saved preset for only the current process/session without changing the saved selection. If Pi is already running, use `/subagent-preset session <name>` for the same process-only override, and `/subagent-preset session-clear` to remove that runtime override. The TUI only selects presets already present in config; it does not edit JSON. If no `asyncSubagents` section exists, run `/subagent-preset init` to insert the bundled sample from `src/async-subagents/async-subagents.sample.jsonc` into the shared config (or to copy a standalone override file when `ASYNC_SUBAGENTS_CONFIG` / `PI_SUBAGENTS_CONFIG` is set). Existing config sections/files are never overwritten. Presets select an agent/model configuration: they can provide global fallback `model`/`thinking`/`extraArgs` and per-role overrides under `asyncSubagents.presets.<name>.types.<subagentType>`. They can also provide ordered `fallbackModels` globally or per-role; when a sub-agent fails with quota/rate-limit errors such as 429, async-subagents immediately tries the next fallback model and remembers the exhausted provider for the current Pi process/session, so later spawns skip that provider until Pi exits. This is intended for provider-level fallback chains such as `antigravity/* → openai-codex/* → zai/*` or `openai-codex/* → zai/*`; omit fallbacks for effectively unlimited providers. Antigravity account rotation has priority over preset fallback: async-subagents only falls back after Antigravity reports that all configured accounts are exhausted for that model. Explicit task model overrides and force-current-model disable preset fallback for that task. The active preset name is stored separately in `~/.pi/agent/subagent-preset-selection.json`.
626
+ When `subagentType` is omitted, the lightweight role router classifies the task
627
+ using the descriptions. Explicit types bypass it. Unknown types or failed
628
+ routing reject the batch, never substitute `defaultType`. Choosing a worker
629
+ model from its candidate list does not involve an LLM call.
630
+
631
+ ### Presets are available-model pools
632
+
633
+ Each agent declares an ordered `models` list in Markdown. A preset declares
634
+ which model references may be used, not another role/model/thinking matrix.
635
+ Selection preserves agent order, intersects it with `preset.models`, checks
636
+ runtime registration/auth availability, and takes the first usable candidate.
637
+ Pool order does not change preference and pool-only models are never appended.
638
+ Without a preset, the full agent list is eligible. Candidate order expresses
639
+ the configured budget preference; runtime does not infer current API prices.
640
+
641
+ Image-bearing tasks and `browser-qa` require confirmed image support; configured
642
+ blind-model masks override runtime image metadata. Remaining eligible models
643
+ form the quota fallback chain, so neither quota history nor image fallback can
644
+ escape the pool. Antigravity account rotation still happens before provider
645
+ fallback. No match, no usable model, or an explicitly empty list rejects the
646
+ whole batch before run directories or child processes are created. A new custom
647
+ agent must supply candidates instead of silently inheriting the parent model.
648
+
649
+ Oracle uses its separate strong-model list and prefers another provider when
650
+ available, but also respects the pool. A same-provider choice is allowed when
651
+ the pool offers no alternative; cross-provider independence is not guaranteed.
652
+ Explicit task/CLI model overrides and `FORCE_CURRENT_MODEL` remain deliberate
653
+ escape hatches and disable automatic model fallback for that task. They do not
654
+ bypass the image-capability check.
655
+
656
+ Define pools in the shared or project `pi-tools-suite.jsonc`. Select a saved
657
+ pool with `/subagent-preset`; use `AGENTS_PRESET=<name>` or
658
+ `/subagent-preset session <name>` for a process-only override and
659
+ `/subagent-preset session-clear` to remove it. The saved selection lives in
660
+ `~/.pi/agent/subagent-preset-selection.json`. `/subagent-preset init` inserts the
661
+ sample only when config is missing. The shipped pools are `cheap` (GLM), `gpt`,
662
+ and `deep` (the retained legacy name for the mixed pool, not worker escalation).
663
+ Initial user config and the sample share one source; descriptions and worker
664
+ model order exist only in the agent files. Existing user files are not rewritten.
347
665
 
348
666
  Example shared async-subagents config section:
349
667
 
350
668
  ```jsonc
351
669
  {
352
670
  "asyncSubagents": {
353
- "defaultType": "quick",
671
+ "defaultType": "research",
354
672
  "routing": {
355
673
  "enabled": true,
356
- "model": "zai/glm-4.5-air",
674
+ "model": "zai/glm-5-turbo",
357
675
  "timeoutMs": 12000
358
676
  },
359
677
  "presets": {
360
678
  "cheap": {
361
- "description": "Use GLM by role, including GLM-5.3 Flash for multimodal work.",
362
- "types": {
363
- "quick": { "model": "zai/glm-5.3", "thinking": "off" },
364
- "frontend": { "model": "zai/glm-5.3-flash", "thinking": "medium" },
365
- "browser-qa": { "model": "zai/glm-5.3-flash", "fallbackModels": ["openai-codex/gpt-5.6-luna"], "thinking": "low" },
366
- "review": { "model": "zai/glm-5.3", "thinking": "high" }
367
- }
679
+ "description": "GLM workers with a strong oracle candidate.",
680
+ "models": ["zai/glm-5-turbo", "zai/glm-5.3-flash", "zai/glm-5.3"]
368
681
  }
369
682
  },
370
683
  "types": {
371
- "frontend": {
372
- "description": "Use for frontend UI/UX visual work: styling, layout, typography, animation, responsive states, component polish, accessibility. Avoid backend/business logic unless needed for UI behavior.",
373
- "thinking": "medium"
374
- },
375
- "review": {
376
- "description": "Use for review/audit of existing code or changes: correctness, security, performance, maintainability, API risks, quality. Do not implement new code.",
377
- "thinking": "high"
684
+ "research": {
685
+ "models": ["zai/glm-5-turbo", "openai-codex/gpt-5.6-luna"],
686
+ "thinking": "low"
378
687
  }
379
688
  }
380
689
  }
381
690
  }
382
691
  ```
383
692
 
384
- ### Parent-model-aware model selection (`modelByParent`)
693
+ ### Legacy configuration compatibility
385
694
 
386
- Any type profile can carry `modelByParent`: a map from glob model refs (matched against the **current parent model**, e.g. `"zai/*"`) to a model for that role. The first matching key wins. Values may be a model string or `{ "model": "...", "fallbackModels": [...] }`. It is resolved after an explicit task `model` / `forcedModel`, but **before** the preset/static profile `model`, so a role can always pick a model based on who the parent is — independent of the active preset.
695
+ Old built-in role names are no longer implicit aliases. `quick`, `scan`,
696
+ `review`, `deep`, `docs`, `frontend`, and `tests` are valid only when explicitly
697
+ defined as ordinary custom/project types. Old preset per-role keys likewise
698
+ apply only when a type with that exact name exists.
387
699
 
388
- The canonical use case is an **`oracle`** role that consults a flagship model from a *different* provider than the parent for a second opinion:
700
+ Legacy `model` plus `fallbackModels` remains readable. `models` is a complete
701
+ replacement list: it clears inherited legacy model/fallback/parent routing.
702
+ A later old-format model override still replaces the primary candidate, and a
703
+ later `fallbackModels` replaces the remaining candidates; `[]` disables them.
704
+ Old `modelByParent` configs remain supported, but ordinary roles give legacy
705
+ preset models precedence. New built-ins contain no parent-tier escalation maps.
389
706
 
390
- ```jsonc
391
- "oracle": {
392
- "description": "Cross-provider second opinion: consult a flagship from a different provider than the parent to pressure-test a hard decision. Read-only; advise, do not edit.",
393
- "model": "openai-codex/gpt-5.6-sol",
394
- "fallbackModels": ["zai/glm-5.3"],
395
- "thinking": "max",
396
- "modelByParent": {
397
- "zai/*": { "model": "openai-codex/gpt-5.6-sol", "fallbackModels": ["zai/glm-5.3"] },
398
- "openai-codex/*": "zai/glm-5.3",
399
- "antigravity/*": { "model": "zai/glm-5.3", "fallbackModels": ["openai-codex/gpt-5.6-sol"] },
400
- "anthropic/*": { "model": "openai-codex/gpt-5.6-sol", "fallbackModels": ["zai/glm-5.3"] }
401
- }
402
- }
403
- ```
404
-
405
- With this config a GLM parent (`zai/*`) spawns the oracle on `gpt-5.6-sol`, a GPT parent (`openai-codex/*`) spawns it on `glm-5.3`, and so on — automatically, at spawn time, with no `task.model` needed. The parent model ref is read from the spawn context (`ctx.model`) and passed into resolution. Pattern matching is case-insensitive `*` glob (same engine as `vision.blindModelPatterns`). When no key matches (or no parent model is known), the role falls back to its static `model` + `fallbackModels`. An explicit `task.model` or `ASYNC_SUBAGENTS_FORCE_CURRENT_MODEL=1` still overrides the match.
707
+ When a preset specifies `models`, it is exclusively a pool; inherited legacy
708
+ `model`, `types`, thinking, arguments and timeout overrides do not run. A later
709
+ explicit old-format preset selector can still replace a pool for compatibility.
710
+ Runtime retry structures and the separate role router continue to use the
711
+ term `fallbackModels` for actual fallback-only lists, not agent candidates.
406
712
 
407
713
  Sub-agents run with `--no-session` by default to avoid writing duplicate Pi session JSONL files for fire-and-forget background work. Set `ASYNC_SUBAGENTS_ENABLE_SESSIONS=1` to restore persisted per-agent sessions under each agent's `sessions/` directory; this also registers the session-navigation slash commands (`/sub-open`, `/sub-back`, `/sub-where`) needed for switching and deeper post-mortem navigation.
408
714
 
@@ -520,6 +826,40 @@ npm run test:prompt-evals:dcp
520
826
 
521
827
  The default live model is `zai/glm-5-turbo`. Override it for the whole suite with `PI_TOOLS_SUITE_E2E_MODEL=provider/model`, or use the existing component variables such as `TOOL_SELECTION_E2E_MODEL`, `ASYNC_SUBAGENTS_MODEL`, `ASYNC_SUBAGENTS_ROUTING_E2E_MODEL`, and `DCP_SUMMARY_E2E_MODEL`. The normal deterministic coverage remains `npm test`; run prompt evals after changing tool descriptions, routing/classifier prompts, DCP summary prompts, or the default evaluation model.
522
828
 
829
+ ### Unified eval harness
830
+
831
+ `test/evals/` adds a shared deterministic + live-model eval layer. The coverage
832
+ gate requires every registered extension and model-facing tool to have a
833
+ deterministic contract. The initial live corpus contains 20 cases across tool
834
+ selection, coding quality, orchestration/escalation, and negative overuse
835
+ controls. Coding-quality fixtures use executable behavioral checks rather than
836
+ an LLM judge, while reports compare parent/worker tokens, provider-reported cost,
837
+ tool calls, changed files, and elapsed time.
838
+
839
+ ```bash
840
+ # Deterministic coverage/contract gate only
841
+ npm run test:evals:contracts
842
+
843
+ # Live matrix as Bun tests. Models are comma/semicolon separated.
844
+ PI_TOOLS_SUITE_EVAL_MODELS='zai/glm-5.3,openai-codex/gpt-5.6-luna,openai-codex/gpt-5.6-terra,openai-codex/gpt-5.6-sol' \
845
+ npm run test:evals:live
846
+
847
+ # Produce JSON + Markdown comparison artifacts.
848
+ PI_TOOLS_SUITE_EVAL_MODELS='zai/glm-5.3,openai-codex/gpt-5.6-terra,openai-codex/gpt-5.6-sol' \
849
+ npm run evals:report
850
+
851
+ # Focus the report runner when iterating
852
+ PI_TOOLS_SUITE_EVAL_MODELS='zai/glm-5.3' \
853
+ PI_TOOLS_SUITE_EVAL_CATEGORIES='coding-quality,negative' \
854
+ npm run evals:report
855
+ ```
856
+
857
+ Live evals are opt-in. The deterministic coverage registry is part of normal
858
+ tests, so adding an extension or tool without eval coverage fails the gate. See
859
+ [`docs/evals.md`](docs/evals.md) for the architecture, complete 20-case catalog,
860
+ fixtures, assertions, metrics, model matrix, report format, environment
861
+ variables, CI recommendations, and the procedure for adding new evals.
862
+
523
863
  Supporting docs and historical standalone README content are kept in `docs/`; third-party license texts are kept in `licenses/`.
524
864
 
525
865
  ## SDK pin
@@ -4,25 +4,32 @@
4
4
 
5
5
  Provide a cheap, fast `browser-qa` async-subagent that reproduces browser bugs
6
6
  and proves fixes with deterministic assertions plus screenshot, video, and trace
7
- evidence. The role uses `zai/glm-5.3-flash`, falling back to
8
- `openai-codex/gpt-5.6-luna`.
9
-
10
- ## Private skill isolation
11
-
12
- - The browser QA skill lives under `src/async-subagents/private-skills/`, outside
13
- Pi's normal skill discovery roots.
7
+ evidence. Its ranked `models` list prefers `zai/glm-5.3-flash`, then
8
+ `openai-codex/gpt-5.6-luna`, filtered by the active preset's model pool and
9
+ confirmed runtime image support.
10
+
11
+ ## Inline agent workflow and skill isolation
12
+
13
+ - All operating instructions, flow contracts, scenario-design guidance, and
14
+ auth-scaffolding rules live in `src/async-subagents/agents/browser-qa.md`.
15
+ Its body becomes the QA child's `promptAppend` through the shared agent
16
+ loader. Parent and router catalogs include only its short `description`.
17
+ - Runner code, vendor dependencies/licenses, and optional JSONC examples live
18
+ under `src/async-subagents/agents/browser-qa/`; none is a discoverable skill.
14
19
  - Sub-agent processes disable normal extension discovery, then always load the
15
- suite's model-tools and Antigravity provider extensions explicitly, regardless
16
- of the selected model. This keeps the process isolated without making any
17
- Antigravity-backed role unavailable.
20
+ suite's model-tools extension. They load the Antigravity provider extension
21
+ only when an Antigravity model is explicitly selected.
18
22
  - A type profile may declare `isolatedSkills`. Spawning that profile adds
19
23
  `--no-skills` followed by one explicit `--skill` per configured path.
20
- - The `browser-qa` profile always loads one self-contained private workflow.
21
- Relevant browser-test design guidance is bundled beside its trusted runner;
22
- no separately discovered skill or browser CLI is required. Configuration may
23
- append isolated skills but cannot remove the mandatory private workflow.
24
- - Other sub-agent profiles and the parent session must not discover the private
25
- skill automatically.
24
+ - `browser-qa` always disables normal skill discovery and filters skill flags
25
+ out of `extraArgs`, even without configured skills. It no longer injects a
26
+ mandatory QA skill. Explicitly configured skills are optional additions.
27
+ - The launcher sets `PI_BROWSER_QA_RUNNER` to the absolute installed runner path,
28
+ replacing inherited values, and strips it from ordinary child environments.
29
+ QA invokes `node "$PI_BROWSER_QA_RUNNER"` without guessing paths from cwd.
30
+ - Model-only profile overrides inherit the workflow. An explicit profile
31
+ `promptAppend` replaces the body like any other agent profile; it is not an
32
+ immutable security boundary. Runtime protections remain in the runner.
26
33
 
27
34
  ## Authentication contract
28
35
 
@@ -98,7 +105,7 @@ evidence. The role uses `zai/glm-5.3-flash`, falling back to
98
105
 
99
106
  ## Reliability and shutdown contract
100
107
 
101
- - The built-in `browser-qa` profile has a 120-second wall-clock budget unless
108
+ - The built-in `browser-qa` profile has a 300-second wall-clock budget unless
102
109
  the caller explicitly supplies a task or spawn timeout. This bounds model
103
110
  stalls as well as browser work.
104
111
  - The trusted runner has its own bounded lifecycle. Browser launch, context
@@ -124,11 +131,14 @@ evidence. The role uses `zai/glm-5.3-flash`, falling back to
124
131
 
125
132
  ## Acceptance criteria
126
133
 
127
- 1. `browser-qa` resolves to the intended model/fallback and its self-contained
128
- private workflow, and its isolated child process can register the configured
134
+ 1. `browser-qa` resolves to the intended model/fallback and its inline Markdown
135
+ workflow, and its isolated child process can register the configured
129
136
  model provider.
130
- 2. Spawn args contain `--no-skills` and the mandatory private skill for this
131
- profile; ordinary profiles retain existing skill discovery behavior.
137
+ 2. Default QA spawn args contain `--no-skills` but no `--skill`. The child
138
+ receives the full workflow in its initial prompt and can invoke the bundled
139
+ runner through `PI_BROWSER_QA_RUNNER` from an unrelated project directory.
140
+ Optional configured skills still load; ordinary profiles retain existing
141
+ skill discovery behavior and do not receive QA-only environment paths.
132
142
  3. Auth profile listing and all error output are redacted; model-authored input
133
143
  cannot execute code in the credential-bearing process.
134
144
  4. Runner tests cover public execution without an auth file, explicit profile