pi-ui-extend 1.0.39 → 1.0.41
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +1 -1
- package/dist/app/commands/command-registry.js +2 -2
- package/dist/app/commands/command-session-actions.d.ts +0 -1
- package/dist/app/commands/command-session-actions.js +22 -13
- package/dist/app/icons.d.ts +14 -0
- package/dist/app/icons.js +33 -0
- package/dist/app/rendering/conversation-tool-renderer.js +2 -2
- package/dist/app/rendering/dcp-stats.d.ts +6 -1
- package/dist/app/rendering/dcp-stats.js +214 -46
- package/dist/app/rendering/editor-panels.js +8 -5
- package/dist/app/session/lazy-session-manager.js +12 -1
- package/dist/app/session/tabs-controller.d.ts +2 -5
- package/dist/app/session/tabs-controller.js +12 -21
- package/dist/app/subagents/subagents-model.d.ts +14 -1
- package/dist/app/subagents/subagents-model.js +34 -15
- package/dist/app/types.d.ts +2 -0
- package/dist/bundled-extensions/session-title/config.js +1 -1
- package/dist/markdown-format.js +27 -9
- package/dist/schemas/pi-tools-suite-schema.d.ts +29 -16
- package/dist/schemas/pi-tools-suite-schema.js +46 -31
- package/external/pi-tools-suite/README.md +392 -52
- package/external/pi-tools-suite/docs/browser-qa-subagent.md +31 -21
- package/external/pi-tools-suite/docs/context-gateway-p00-adr.md +216 -0
- package/external/pi-tools-suite/docs/context-gateway-p01n-gate-review.md +122 -0
- package/external/pi-tools-suite/docs/context-gateway-p01n-measurement.md +133 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-ra-evidence.md +111 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-rb-evidence.md +100 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-rc-evidence.md +69 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-rd-evidence.md +100 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-re-evidence.md +74 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-rf-evidence.md +153 -0
- package/external/pi-tools-suite/docs/context-gateway-p01r-rg-evidence.md +235 -0
- package/external/pi-tools-suite/docs/evals.md +684 -0
- package/external/pi-tools-suite/docs/subagent-model-pools.md +109 -0
- package/external/pi-tools-suite/package.json +10 -3
- package/external/pi-tools-suite/src/async-subagents/{private-skills → agents}/browser-qa/scripts/browser-qa-runner.mjs +82 -1
- package/external/pi-tools-suite/src/async-subagents/{private-skills/browser-qa/SKILL.md → agents/browser-qa.md} +261 -12
- package/external/pi-tools-suite/src/async-subagents/agents/implement.md +20 -0
- package/external/pi-tools-suite/src/async-subagents/agents/oracle.md +16 -0
- package/external/pi-tools-suite/src/async-subagents/agents/research.md +18 -0
- package/external/pi-tools-suite/src/async-subagents/agents/verify.md +18 -0
- package/external/pi-tools-suite/src/async-subagents/async-subagents.sample.jsonc +27 -243
- package/external/pi-tools-suite/src/async-subagents/commands.ts +6 -2
- package/external/pi-tools-suite/src/async-subagents/core/agent-catalog.ts +41 -0
- package/external/pi-tools-suite/src/async-subagents/core/agent-strategy.ts +13 -93
- package/external/pi-tools-suite/src/async-subagents/core/agents-dir.ts +494 -0
- package/external/pi-tools-suite/src/async-subagents/core/browser-qa.ts +9 -0
- package/external/pi-tools-suite/src/async-subagents/core/config.ts +200 -143
- package/external/pi-tools-suite/src/async-subagents/core/model-fallback.ts +1 -1
- package/external/pi-tools-suite/src/async-subagents/core/model-selection.ts +54 -0
- package/external/pi-tools-suite/src/async-subagents/core/prompt.ts +7 -6
- package/external/pi-tools-suite/src/async-subagents/core/routing.ts +52 -45
- package/external/pi-tools-suite/src/async-subagents/core/spawn.ts +12 -4
- package/external/pi-tools-suite/src/async-subagents/index.ts +11 -1
- package/external/pi-tools-suite/src/async-subagents/lib.ts +6 -2
- package/external/pi-tools-suite/src/async-subagents/tools/spawn.ts +46 -18
- package/external/pi-tools-suite/src/async-subagents/tools/subagents.ts +3 -2
- package/external/pi-tools-suite/src/async-subagents/types.ts +2 -0
- package/external/pi-tools-suite/src/coding-discipline/index.ts +41 -142
- package/external/pi-tools-suite/src/config.ts +1 -22
- package/external/pi-tools-suite/src/context-gateway/accounting.ts +151 -0
- package/external/pi-tools-suite/src/context-gateway/config.ts +111 -0
- package/external/pi-tools-suite/src/context-gateway/index.ts +160 -0
- package/external/pi-tools-suite/src/context-gateway/metadata-normalization.ts +88 -0
- package/external/pi-tools-suite/src/context-gateway/storeless-capabilities.ts +89 -0
- package/external/pi-tools-suite/src/context-gateway/telemetry.ts +429 -0
- package/external/pi-tools-suite/src/context-gateway/test-output-parser.ts +326 -0
- package/external/pi-tools-suite/src/context-gateway/types.ts +152 -0
- package/external/pi-tools-suite/src/dcp/auto-compress-budget.ts +106 -0
- package/external/pi-tools-suite/src/dcp/auto-compress.ts +810 -106
- package/external/pi-tools-suite/src/dcp/commands.ts +64 -139
- package/external/pi-tools-suite/src/dcp/compress-tool.ts +369 -35
- package/external/pi-tools-suite/src/dcp/compression-blocks.ts +510 -64
- package/external/pi-tools-suite/src/dcp/compression-preview.ts +113 -0
- package/external/pi-tools-suite/src/dcp/compression-progress.ts +70 -0
- package/external/pi-tools-suite/src/dcp/config.ts +36 -61
- package/external/pi-tools-suite/src/dcp/conversation-index.ts +421 -0
- package/external/pi-tools-suite/src/dcp/debug-log.ts +7 -5
- package/external/pi-tools-suite/src/dcp/index.ts +617 -203
- package/external/pi-tools-suite/src/dcp/journal.ts +566 -0
- package/external/pi-tools-suite/src/dcp/progress-controller.ts +244 -0
- package/external/pi-tools-suite/src/dcp/prompts.ts +10 -7
- package/external/pi-tools-suite/src/dcp/provider-tool-results.ts +189 -0
- package/external/pi-tools-suite/src/dcp/pruner-candidates.ts +298 -78
- package/external/pi-tools-suite/src/dcp/pruner-compression-blocks.ts +173 -281
- package/external/pi-tools-suite/src/dcp/pruner-emergency.ts +2 -4
- package/external/pi-tools-suite/src/dcp/pruner-message-ids.ts +17 -5
- package/external/pi-tools-suite/src/dcp/pruner-metadata.ts +11 -1
- package/external/pi-tools-suite/src/dcp/pruner-nudge.ts +30 -82
- package/external/pi-tools-suite/src/dcp/pruner-tools.ts +22 -133
- package/external/pi-tools-suite/src/dcp/pruner.ts +18 -33
- package/external/pi-tools-suite/src/dcp/recovery.ts +129 -0
- package/external/pi-tools-suite/src/dcp/shadow-plan.ts +127 -0
- package/external/pi-tools-suite/src/dcp/state-transaction.ts +102 -0
- package/external/pi-tools-suite/src/dcp/state.ts +158 -580
- package/external/pi-tools-suite/src/dcp/ui.ts +1 -0
- package/external/pi-tools-suite/src/default-pi-tools-suite-config.ts +55 -220
- package/external/pi-tools-suite/src/index.ts +9 -0
- package/external/pi-tools-suite/src/model-tools/index.ts +76 -42
- package/external/pi-tools-suite/src/repo-discovery/index.ts +84 -18
- package/external/pi-tools-suite/src/repo-discovery/native-compact.ts +458 -0
- package/external/pi-tools-suite/src/session-recovery/index.ts +189 -43
- package/external/pi-tools-suite/src/tool-descriptions.ts +43 -38
- package/external/pi-tools-suite/src/truncation-metadata-normalizer/index.ts +17 -0
- package/package.json +6 -6
- package/schemas/pi-tools-suite.json +159 -78
- package/external/pi-tools-suite/src/async-subagents/private-skills/browser-qa/references/auth-scaffold-spec.md +0 -78
- package/external/pi-tools-suite/src/async-subagents/private-skills/browser-qa/references/qa-design.md +0 -223
- package/external/pi-tools-suite/src/dcp/state-persistence.ts +0 -195
- /package/external/pi-tools-suite/src/async-subagents/{private-skills/browser-qa/references → agents/browser-qa/examples}/qa-auth.example.jsonc +0 -0
- /package/external/pi-tools-suite/src/async-subagents/{private-skills/browser-qa/references → agents/browser-qa/examples}/qa-flow.example.jsonc +0 -0
- /package/external/pi-tools-suite/src/async-subagents/{private-skills → agents}/browser-qa/vendor/fflate.LICENSE +0 -0
- /package/external/pi-tools-suite/src/async-subagents/{private-skills → agents}/browser-qa/vendor/fflate.mjs +0 -0
|
@@ -100,7 +100,13 @@ DCP settings are stored only under `dcp` in the user shared config file `~/.conf
|
|
|
100
100
|
"nudgeFrequency": 1,
|
|
101
101
|
"iterationNudgeThreshold": 6,
|
|
102
102
|
"nudgeForce": "strong",
|
|
103
|
-
"protectedTools": ["compress", "write", "edit", "subagents"]
|
|
103
|
+
"protectedTools": ["compress", "write", "edit", "subagents"],
|
|
104
|
+
"autoCompress": {
|
|
105
|
+
"enabled": false,
|
|
106
|
+
"patience": 2,
|
|
107
|
+
"summarizerModel": [],
|
|
108
|
+
"timeoutMs": 20000
|
|
109
|
+
}
|
|
104
110
|
},
|
|
105
111
|
"strategies": {
|
|
106
112
|
"emergencyCurrentTurnPruning": {
|
|
@@ -170,7 +176,11 @@ DCP settings are stored only under `dcp` in the user shared config file `~/.conf
|
|
|
170
176
|
|
|
171
177
|
`minContextPercent` / `maxContextPercent` accept legacy fractions (`0.25`), percent strings (`"25%"`), or absolute token counts when Pi knows the current model context window. `minContextLimit` / `maxContextLimit` and `modelMinContextLimits` / `modelMaxContextLimits` are explicit absolute-or-percent aliases. `modelOverrides` and the `modelMin*` / `modelMax*` maps support exact model keys plus `*` / `?` wildcard patterns; matching is applied from generic to specific so exact bare-model matches override bare wildcards, and exact `provider/model` matches override provider wildcards. Array fields are union-merged, so model-specific `protectedTools` extend the defaults instead of replacing them. If `compress.protectUserMessages` is enabled, range compression appends selected user messages verbatim instead of rejecting the range; individual message compression still skips protected raw user messages. Protected tool outputs are copied into summaries for tools protected by name or `protectedFilePatterns`; protected `subagents` result reads also try to include the saved `result.md` artifact when available.
|
|
172
178
|
|
|
173
|
-
`
|
|
179
|
+
`compress.autoCompress.enabled` is `false` by default. When explicitly enabled, its `patience` counts completed correlated main-provider opportunities, not repeated context transforms; `summarizerModel: []` uses the bounded extractive fallback without a model call, while configured models share the single `timeoutMs` summarizer deadline. If that deadline expires, DCP still has a bounded finalization grace to commit the extractive fallback. Auto commit is accepted only when the full projected replacement has positive gain and meets the current budget-recovery target.
|
|
180
|
+
|
|
181
|
+
`strategies.emergencyCurrentTurnPruning` is the default-enabled lossy safety floor for a single unfinished turn that has no normal compression candidate. DCP first emits emergency reminders and offers only safe old same-turn tool-result candidates. After `patience` completed ignored opportunities, or at the model-independent `hardContextPercent`, it replaces eligible oldest result bodies until the estimated provider context reaches `targetContextPercent` or a margin below the model emergency threshold. User messages, configured/protected data, the newest `keepRecentToolPairs`, and results without **completed** provider evidence are never selected. HTTP 2xx alone is not evidence; DCP promotes eligibility only after an unambiguously correlated successful finalized assistant response, and ambiguous retries/interleaving fail closed. The raw session transcript is unchanged. Setting `enabled` to `false` disables same-turn candidates and lossy pruning, but keeps the non-destructive emergency reminder.
|
|
182
|
+
|
|
183
|
+
DCP sidecars are session-private versioned generation envelopes under `<sessionDir>/dcp-state/`. Saves use private atomic files, retain a last-valid `.prev` generation, validate payload/block-graph integrity on load, quarantine corrupt primaries, and use a cross-process exclusive lock rather than silent last-writer-wins. Live paused sessions are not deleted merely for age. Protected subagent artifacts are optional bounded recovery input: reads are async, rooted at the session cwd, reject symlink escapes and oversized files, and never silently truncate a required protected artifact.
|
|
174
184
|
|
|
175
185
|
Set `dcp.debug: true` to write a JSONL debug log of DCP context/prune/compress events to `~/.pi/agent/dcp-debug.jsonl` (override the path with `PI_DCP_DEBUG_LOG`, or enable without config via `PI_DCP_DEBUG=1`); off by default. The log is size-limited and rotated: once it reaches `dcp.debugLog.maxBytes` (default `5242880` = 5 MB) it is renamed to `.1`, older backups shift down (`.1`→`.2`, …) and the oldest beyond `dcp.debugLog.maxBackups` (default `3`, minimum `1`) is dropped; override either with `PI_DCP_DEBUG_MAX_BYTES` / `PI_DCP_DEBUG_MAX_BACKUPS`.
|
|
176
186
|
|
|
@@ -188,9 +198,28 @@ Install the language servers used by the bundled example config. The commands be
|
|
|
188
198
|
# TypeScript / JavaScript
|
|
189
199
|
npm install -g typescript typescript-language-server
|
|
190
200
|
|
|
201
|
+
# Svelte
|
|
202
|
+
npm install -g svelte-language-server
|
|
203
|
+
|
|
204
|
+
# Vue
|
|
205
|
+
npm install -g @vue/language-server
|
|
206
|
+
|
|
191
207
|
# Python
|
|
192
208
|
python3 -m pip install --user python-lsp-server
|
|
193
209
|
|
|
210
|
+
# Go
|
|
211
|
+
go install golang.org/x/tools/gopls@latest
|
|
212
|
+
|
|
213
|
+
# C / C++ (clangd)
|
|
214
|
+
brew install llvm
|
|
215
|
+
# or use your distro's clangd package: apt install clangd, dnf install clang-tools-extra, ...
|
|
216
|
+
|
|
217
|
+
# Lua
|
|
218
|
+
brew install lua-language-server
|
|
219
|
+
|
|
220
|
+
# Bash
|
|
221
|
+
npm install -g bash-language-server
|
|
222
|
+
|
|
194
223
|
# C# / Unity
|
|
195
224
|
dotnet tool install -g Microsoft.CodeAnalysis.LanguageServer
|
|
196
225
|
|
|
@@ -252,23 +281,276 @@ Minimal shared config shape:
|
|
|
252
281
|
|
|
253
282
|
Project-local overrides can be added in `.pi/pi-tools-suite.jsonc`; pi-tools-suite asks for trust before using project-local LSP binaries.
|
|
254
283
|
|
|
284
|
+
### Popular language-server examples
|
|
285
|
+
|
|
286
|
+
Copy the entries you need into `lsp.servers` of the shared config. Values mirror the commented templates shipped in the generated config file; where they disagree, the generated file is authoritative. Both diagnostics modes are on by default: servers that support pull diagnostics (`textDocument/diagnostic`) are queried directly, and push diagnostics (`publishDiagnostics`) are awaited for every server — `pullDiagnostics: false` / `waitForPublishDiagnostics: false` disable either side explicitly. Servers start lazily: one spawns only after a mutating tool (Edit/Write/ast-grep/apply_patch) touches a file matching `include`, and diagnostics land in that tool's result.
|
|
287
|
+
|
|
288
|
+
```jsonc
|
|
289
|
+
{
|
|
290
|
+
"lsp": {
|
|
291
|
+
"servers": [
|
|
292
|
+
// Svelte (verified with svelte-language-server): compiler + embedded TS/JS diagnostics
|
|
293
|
+
{
|
|
294
|
+
"id": "svelte",
|
|
295
|
+
"include": ["**/*.svelte"],
|
|
296
|
+
"exclude": ["**/node_modules/**"],
|
|
297
|
+
"rootMarkers": ["svelte.config.js", "package.json"],
|
|
298
|
+
"bin": "svelteserver",
|
|
299
|
+
"args": ["--stdio"],
|
|
300
|
+
"startupTimeoutMs": 30000,
|
|
301
|
+
"diagnosticsWaitMs": 8000,
|
|
302
|
+
"languageIdByExtension": { ".svelte": "svelte" }
|
|
303
|
+
},
|
|
304
|
+
// Vue (Volar)
|
|
305
|
+
{
|
|
306
|
+
"id": "vue",
|
|
307
|
+
"include": ["**/*.vue"],
|
|
308
|
+
"exclude": ["**/node_modules/**"],
|
|
309
|
+
"rootMarkers": ["package.json"],
|
|
310
|
+
"bin": "vue-language-server",
|
|
311
|
+
"args": ["--stdio"],
|
|
312
|
+
"startupTimeoutMs": 30000,
|
|
313
|
+
"diagnosticsWaitMs": 8000,
|
|
314
|
+
"languageIdByExtension": { ".vue": "vue" }
|
|
315
|
+
},
|
|
316
|
+
// Python (python-lsp-server)
|
|
317
|
+
{
|
|
318
|
+
"id": "python",
|
|
319
|
+
"include": ["**/*.py", "**/*.pyi"],
|
|
320
|
+
"exclude": ["**/.git/**", "**/node_modules/**", "**/__pycache__/**", "**/.venv/**", "**/venv/**", "**/.tox/**", "**/.mypy_cache/**", "**/.ruff_cache/**"],
|
|
321
|
+
"rootMarkers": ["pyproject.toml", "setup.py", "setup.cfg", "requirements.txt", "Pipfile", "poetry.lock", ".git"],
|
|
322
|
+
"bin": "pylsp",
|
|
323
|
+
"args": [],
|
|
324
|
+
"languageIdByExtension": { ".py": "python", ".pyi": "python" }
|
|
325
|
+
},
|
|
326
|
+
// Go (gopls, push diagnostics)
|
|
327
|
+
{
|
|
328
|
+
"id": "go",
|
|
329
|
+
"include": ["**/*.go"],
|
|
330
|
+
"exclude": ["**/.git/**", "**/vendor/**"],
|
|
331
|
+
"rootMarkers": ["go.mod", ".git"],
|
|
332
|
+
"bin": "gopls",
|
|
333
|
+
"args": [],
|
|
334
|
+
"startupTimeoutMs": 20000,
|
|
335
|
+
"diagnosticsWaitMs": 8000,
|
|
336
|
+
"languageIdByExtension": { ".go": "go" }
|
|
337
|
+
},
|
|
338
|
+
// Rust (rust-analyzer, push diagnostics)
|
|
339
|
+
{
|
|
340
|
+
"id": "rust",
|
|
341
|
+
"include": ["**/*.rs"],
|
|
342
|
+
"exclude": ["**/.git/**", "**/node_modules/**", "**/target/**"],
|
|
343
|
+
"rootMarkers": ["Cargo.toml", "rust-project.json", ".git"],
|
|
344
|
+
"bin": "rust-analyzer",
|
|
345
|
+
"args": [],
|
|
346
|
+
"startupTimeoutMs": 20000,
|
|
347
|
+
"diagnosticsWaitMs": 20000,
|
|
348
|
+
"pullDiagnostics": false,
|
|
349
|
+
"waitForPublishDiagnostics": true,
|
|
350
|
+
"languageIdByExtension": { ".rs": "rust" }
|
|
351
|
+
},
|
|
352
|
+
// C / C++ (clangd, push diagnostics)
|
|
353
|
+
{
|
|
354
|
+
"id": "clangd",
|
|
355
|
+
"include": ["**/*.c", "**/*.cc", "**/*.cpp", "**/*.cxx", "**/*.h", "**/*.hh", "**/*.hpp"],
|
|
356
|
+
"exclude": ["**/.git/**", "**/node_modules/**"],
|
|
357
|
+
"rootMarkers": ["compile_commands.json", "CMakeLists.txt", "Makefile", ".clang-format", ".git"],
|
|
358
|
+
"bin": "clangd",
|
|
359
|
+
"args": [],
|
|
360
|
+
"startupTimeoutMs": 20000,
|
|
361
|
+
"diagnosticsWaitMs": 8000,
|
|
362
|
+
"languageIdByExtension": {
|
|
363
|
+
".c": "c", ".cc": "cpp", ".cpp": "cpp", ".cxx": "cpp",
|
|
364
|
+
".h": "c", ".hh": "cpp", ".hpp": "cpp"
|
|
365
|
+
}
|
|
366
|
+
},
|
|
367
|
+
// C# / Unity (Roslyn)
|
|
368
|
+
{
|
|
369
|
+
"id": "csharp",
|
|
370
|
+
"include": ["**/*.cs", "**/*.csx"],
|
|
371
|
+
"exclude": ["**/.git/**", "**/node_modules/**", "**/bin/**", "**/obj/**", "**/.vs/**", "**/Library/**", "**/Temp/**", "**/Logs/**"],
|
|
372
|
+
"rootMarkers": ["*.sln", "*.csproj", "global.json", "Directory.Build.props", "Directory.Packages.props", "Packages/manifest.json", "ProjectSettings/ProjectVersion.txt", ".git"],
|
|
373
|
+
"bin": "~/.dotnet/tools/roslyn-language-server",
|
|
374
|
+
"args": ["--stdio", "--autoLoadProjects", "--logLevel", "Error"],
|
|
375
|
+
"startupTimeoutMs": 30000,
|
|
376
|
+
"diagnosticsWaitMs": 15000,
|
|
377
|
+
"languageIdByExtension": { ".cs": "csharp", ".csx": "csharp" }
|
|
378
|
+
},
|
|
379
|
+
// Ruby
|
|
380
|
+
{
|
|
381
|
+
"id": "ruby",
|
|
382
|
+
"include": ["**/*.rb", "**/*.rake", "**/Gemfile", "**/Rakefile", "**/*.gemspec"],
|
|
383
|
+
"exclude": ["**/.git/**", "**/node_modules/**", "**/vendor/bundle/**", "**/.bundle/**", "**/tmp/**", "**/log/**"],
|
|
384
|
+
"rootMarkers": ["Gemfile.lock", "*.gemspec", "Rakefile", ".ruby-version", ".git"],
|
|
385
|
+
"bin": "ruby-lsp",
|
|
386
|
+
"args": [],
|
|
387
|
+
"startupTimeoutMs": 60000,
|
|
388
|
+
"diagnosticsWaitMs": 10000,
|
|
389
|
+
"languageIdByExtension": { ".rb": "ruby", ".rake": "ruby", ".gemspec": "ruby" }
|
|
390
|
+
},
|
|
391
|
+
// Lua (lua-language-server)
|
|
392
|
+
{
|
|
393
|
+
"id": "lua",
|
|
394
|
+
"include": ["**/*.lua"],
|
|
395
|
+
"exclude": ["**/.git/**", "**/node_modules/**"],
|
|
396
|
+
"rootMarkers": [".luarc.json", ".git"],
|
|
397
|
+
"bin": "lua-language-server",
|
|
398
|
+
"args": [],
|
|
399
|
+
"startupTimeoutMs": 30000,
|
|
400
|
+
"diagnosticsWaitMs": 6000,
|
|
401
|
+
"languageIdByExtension": { ".lua": "lua" }
|
|
402
|
+
},
|
|
403
|
+
// Bash
|
|
404
|
+
{
|
|
405
|
+
"id": "bash",
|
|
406
|
+
"include": ["**/*.sh", "**/*.bash"],
|
|
407
|
+
"exclude": ["**/.git/**", "**/node_modules/**"],
|
|
408
|
+
"rootMarkers": [".git"],
|
|
409
|
+
"bin": "bash-language-server",
|
|
410
|
+
"args": ["start"],
|
|
411
|
+
"startupTimeoutMs": 15000,
|
|
412
|
+
"diagnosticsWaitMs": 5000,
|
|
413
|
+
"languageIdByExtension": { ".sh": "shellscript", ".bash": "shellscript" }
|
|
414
|
+
},
|
|
415
|
+
// Markdown (link validation settings ship in the generated config template)
|
|
416
|
+
{
|
|
417
|
+
"id": "markdown",
|
|
418
|
+
"include": ["**/*.md", "**/*.markdown", "**/*.mdown", "**/*.mkd", "**/*.mmd"],
|
|
419
|
+
"exclude": ["**/.git/**", "**/node_modules/**"],
|
|
420
|
+
"rootMarkers": [".git", "package.json", "README.md"],
|
|
421
|
+
"bin": "vscode-markdown-language-server",
|
|
422
|
+
"args": ["--stdio"],
|
|
423
|
+
"startupTimeoutMs": 15000,
|
|
424
|
+
"diagnosticsWaitMs": 5000,
|
|
425
|
+
"languageIdByExtension": { ".md": "markdown", ".markdown": "markdown", ".mdown": "markdown", ".mkd": "markdown", ".mmd": "markdown" }
|
|
426
|
+
}
|
|
427
|
+
]
|
|
428
|
+
}
|
|
429
|
+
}
|
|
430
|
+
```
|
|
431
|
+
|
|
432
|
+
Notes:
|
|
433
|
+
|
|
434
|
+
- Svelte resolves `svelte` and `typescript` from the workspace `node_modules`, so project-local versions win; `.svelte.js`/`.svelte.ts` runes modules are not covered because their extensions collide with the TypeScript server.
|
|
435
|
+
- Vue requires `@vue/language-server` v2+ (the `vue-language-server` binary) plus the workspace's `vue` package for template type-checking.
|
|
436
|
+
- The full commented templates (including GDScript via a headless Godot wrapper and the complete Markdown link-validation `settings`) are written to the shared config file on first run.
|
|
437
|
+
|
|
255
438
|
## Async sub-agents
|
|
256
439
|
|
|
257
|
-
|
|
440
|
+
Model selection uses the ordered candidates from each agent's Markdown file,
|
|
441
|
+
filtered by the selected preset's available models and runtime capabilities.
|
|
442
|
+
Explicit task/CLI model overrides bypass the pool. Setting
|
|
443
|
+
`ASYNC_SUBAGENTS_FORCE_CURRENT_MODEL=1` (or
|
|
444
|
+
`PI_SUBAGENTS_FORCE_CURRENT_MODEL=1`) deliberately selects the parent model and
|
|
445
|
+
strips conflicting model arguments; this is not the economical default.
|
|
446
|
+
|
|
447
|
+
The five built-in modes are `research` (read-only evidence and independent
|
|
448
|
+
review), `implement` (bounded code, docs, tests, or UI changes), `verify`
|
|
449
|
+
(run checks and diagnose logs without fixing files), `browser-qa` (trusted
|
|
450
|
+
browser workflow), and `oracle` (deliberate strong second opinion).
|
|
451
|
+
Ordinary workers use economical model candidates; no built-in parent-tier
|
|
452
|
+
rule promotes them to a flagship. Oracle is the exception, not an automatic
|
|
453
|
+
retry for difficult work. Task-specific discipline belongs in the brief.
|
|
454
|
+
|
|
455
|
+
Delegate when a suitable lower-cost worker can handle bounded work or noisy
|
|
456
|
+
intermediate evidence should stay outside the parent context. One sequential
|
|
457
|
+
task can qualify. Keep decisions and integration in the parent; read compact
|
|
458
|
+
results and verify selectively rather than repeating the worker's investigation.
|
|
459
|
+
Do trivial reads/edits directly. Redirect a noisy command to a log without an
|
|
460
|
+
extra LLM when no interpretation is needed. `verify`'s no-edit instruction is
|
|
461
|
+
a behavioral contract, not a read-only filesystem sandbox for its shell.
|
|
462
|
+
|
|
463
|
+
Run `/ultrawork` or `/ulw` for orchestration, `/hyperplan` to pressure-test a
|
|
464
|
+
plan, or set `ULTRAWORK=1` to apply the orchestration prompt to normal inputs.
|
|
465
|
+
`ULTRAWORK_AUTO=1` classifies only the first normal input on non-GPT parents;
|
|
466
|
+
GPT-like parents skip that automatic transform, not ordinary delegation.
|
|
467
|
+
|
|
468
|
+
See [Model pools and migration](docs/subagent-model-pools.md) for the selection
|
|
469
|
+
contract, configuration examples, override rules and legacy compatibility.
|
|
470
|
+
|
|
471
|
+
### Parent-first role selection
|
|
472
|
+
|
|
473
|
+
The parent normally selects an explicit `subagentType` from the effective
|
|
474
|
+
system-prompt catalog, preferring a matching project-local specialist. Valid
|
|
475
|
+
explicit types bypass the LLM router entirely; presets, model selection, tools,
|
|
476
|
+
skills, and role instructions are still applied by the normal config resolver.
|
|
477
|
+
Model/thinking overrides are not substitutes for selecting a role.
|
|
478
|
+
|
|
479
|
+
The router remains enabled as a fallback for omitted types: use it when the role
|
|
480
|
+
is unclear or the user explicitly requests automatic routing. Only omitted
|
|
481
|
+
tasks are classified, in one batch; the parent's explicit choices are preserved.
|
|
482
|
+
Real-browser QA still requires explicit `subagentType: "browser-qa"`.
|
|
483
|
+
|
|
484
|
+
Unknown explicit types and failed/incomplete automatic routing reject the
|
|
485
|
+
**entire spawn batch before run state or child processes are created**. The tool
|
|
486
|
+
returns an error with affected task IDs and available types; the parent should
|
|
487
|
+
correct the roles and resubmit the whole batch. Provider error responses are
|
|
488
|
+
failures too, not successful routes. Configured fallback router models may be
|
|
489
|
+
tried, but missing routes are never silently replaced with `quick`/`defaultType`.
|
|
490
|
+
|
|
491
|
+
With `routing.enabled: false`, every spawn task must supply a valid explicit
|
|
492
|
+
type. `defaultType` remains a preference for genuinely ambiguous LLM choices
|
|
493
|
+
and a legacy config-resolver default, not a spawn error fallback. Existing
|
|
494
|
+
callers relying on an implicit default must now choose a type explicitly.
|
|
495
|
+
|
|
496
|
+
### Project-local agents (`.pi/agents/*.md`)
|
|
497
|
+
|
|
498
|
+
A project can ship sub-agent roles as individual Markdown files in
|
|
499
|
+
`<project>/.pi/agents/`. The first such directory found walking up from the
|
|
500
|
+
session cwd is used; each top-level `*.md` file becomes a `subagentType` named
|
|
501
|
+
after the file. Parent and router see the short `description`; only the child
|
|
502
|
+
receives the Markdown body. A project's ordered `models` are filtered through
|
|
503
|
+
the same active preset pool as built-in agents.
|
|
504
|
+
|
|
505
|
+
```markdown
|
|
506
|
+
---
|
|
507
|
+
description: Use for reviewing this repo's diff — knows the house rules.
|
|
508
|
+
icon: eye
|
|
509
|
+
models:
|
|
510
|
+
- zai/glm-5-turbo
|
|
511
|
+
- openai-codex/gpt-5.6-luna
|
|
512
|
+
thinking: low
|
|
513
|
+
tools: read, grep
|
|
514
|
+
retry:
|
|
515
|
+
maxRetries: 1
|
|
516
|
+
backoffMs: 2000
|
|
517
|
+
---
|
|
518
|
+
|
|
519
|
+
You are this project's staff reviewer. Apply the repo rules from
|
|
520
|
+
AGENTS.md before approving anything; cite file paths first.
|
|
521
|
+
```
|
|
258
522
|
|
|
259
|
-
|
|
523
|
+
- Frontmatter keys: `name` (must match the filename), `description`, `icon`, `models`, `thinking`, `tools`, `isolatedSkills`, `extraArgs`, `promptAppend`, `promptOverride`, `retry`, `maxResultBytes`, `timeoutMs`. Legacy `model`, `fallbackModels`, and `modelByParent` still load. Unknown keys are rejected with an error naming the file.
|
|
524
|
+
- Array fields accept block lists (`- item`), inline arrays (`[a, b]`), or comma-separated strings (`tools: read, grep, bash`). The frontmatter YAML subset is intentionally small: scalars, quoted strings, numbers, comments, lists, and nested maps for `modelByParent`/`retry`. Tabs, block scalars (`|`/`>`), anchors/aliases, and flow maps are hard errors naming file and line.
|
|
525
|
+
- The markdown body becomes `promptAppend`: it is appended after the standard generated prompt (parent objective + task + output format), so the agent still receives its task in the usual structure. Use frontmatter `promptOverride` for full prompt replacement.
|
|
526
|
+
- Precedence: project agent fields override same-named types from user/project JSONC config (field-level; other fields are kept), which in turn override built-ins. Setting `ASYNC_SUBAGENTS_CONFIG` / `PI_SUBAGENTS_CONFIG` disables the directory (explicit config = full control).
|
|
527
|
+
- Files without frontmatter are skipped (a `README.md` there is fine). Definition loading is uncached: edits apply on the next config read/spawn without a restart, and the effective system-prompt catalog is rebuilt at parent-agent start.
|
|
528
|
+
- Bundled roles use the same format internally under `src/async-subagents/agents/*.md`; built-in and project-local profiles therefore share one parser and normalization path instead of maintaining a second role-description schema in TypeScript.
|
|
529
|
+
- `icon` names an agent glyph for UIs that render sub-agent widgets (pix TUI panel, Pix Desktop subagents panel): `agent` (neutral default), `search`, `code`, `flask`, `globe`, `sparkles`, `brain`, `wrench`, `terminal`, `bug`, `book`, `eye`, `zap`, `rocket`. The value is passed through opaquely; unknown names render as the neutral agent icon, and status stays color-coded next to it.
|
|
260
530
|
|
|
261
531
|
### Private browser QA and project auth
|
|
262
532
|
|
|
263
533
|
The built-in `browser-qa` role runs on `zai/glm-5.3-flash`, with
|
|
264
|
-
`openai-codex/gpt-5.6-luna` as its fallback. Its
|
|
265
|
-
|
|
266
|
-
|
|
267
|
-
|
|
268
|
-
|
|
269
|
-
|
|
270
|
-
|
|
271
|
-
|
|
534
|
+
`openai-codex/gpt-5.6-luna` as its fallback. Its complete workflow and detailed
|
|
535
|
+
scenario-design guidance live in the Markdown body of
|
|
536
|
+
`src/async-subagents/agents/browser-qa.md`. The normal profile loader appends
|
|
537
|
+
that body to the QA child's task prompt; the parent and LLM router receive only
|
|
538
|
+
the short `description`. There is no additional QA skill to discover or read.
|
|
539
|
+
|
|
540
|
+
Executable resources live under `src/async-subagents/agents/browser-qa/`.
|
|
541
|
+
The launcher supplies the installed runner's absolute path in
|
|
542
|
+
`PI_BROWSER_QA_RUNNER`; the child invokes `node "$PI_BROWSER_QA_RUNNER"` from
|
|
543
|
+
the delegated project's cwd. This non-secret path is set only for QA children.
|
|
544
|
+
QA always launches with `--no-skills`, even when `isolatedSkills` is empty,
|
|
545
|
+
and skill flags in `extraArgs` cannot bypass that isolation. Explicitly
|
|
546
|
+
configured `isolatedSkills` remain supported as optional additions; no built-in
|
|
547
|
+
QA `--skill` is injected. Other roles retain their normal discovery behavior.
|
|
548
|
+
|
|
549
|
+
Model/thinking/tool-only profile overrides inherit the Markdown workflow.
|
|
550
|
+
An explicit profile `promptAppend` replaces the inherited body under the usual
|
|
551
|
+
field-level merge rules; custom QA instructions must preserve the runner-only,
|
|
552
|
+
credential, target, and evidence contracts. Runner-enforced isolation and
|
|
553
|
+
credential handling remain in code, not in the prompt.
|
|
272
554
|
|
|
273
555
|
Public browser QA does not require an auth profile or `.pi/qa_auth.jsonc`: run it
|
|
274
556
|
with an explicit base URL, whose exact origin becomes the fail-closed allowlist.
|
|
@@ -325,8 +607,8 @@ creating a template. Only an explicit authenticated request may create the
|
|
|
325
607
|
private template. Missing, rejected, or expired selected auth returns
|
|
326
608
|
`QA_AUTH_UPDATE_REQUIRED`, naming only the profile/file/reason needed for the
|
|
327
609
|
parent to ask the user for an update and rerun. See
|
|
328
|
-
`src/async-subagents/
|
|
329
|
-
for complete profile shapes and `
|
|
610
|
+
`src/async-subagents/agents/browser-qa/examples/qa-auth.example.jsonc`
|
|
611
|
+
for complete profile shapes and `examples/qa-flow.example.jsonc` beside it for
|
|
330
612
|
the declarative, non-executable QA action/assertion format.
|
|
331
613
|
|
|
332
614
|
Browser QA videos automatically visualize pointer interactions. Clicks and
|
|
@@ -341,68 +623,92 @@ Async-subagents also injects a lightweight oh-my-openagent-style system-prompt s
|
|
|
341
623
|
|
|
342
624
|
For blind-model screenshot/image inspection, use the main-session `coding-discipline` lookup tool; the bundled default uses vision-capable `zai/glm-5.3-flash`. Async-subagents still supports `imagePaths` on tasks when a broader delegated track genuinely needs images, but it no longer ships a dedicated `vision` role. Dynamic provider capabilities can be missing or stale after switching models, so blind parent models can still be configured explicitly with case-insensitive `*` masks under `asyncSubagents.vision.blindModelPatterns` in `~/.config/pi/pi-tools-suite.jsonc`; do not include `zai/glm-5.3-flash` because it accepts image input. This keeps guidance honest, not a sub-agent role.
|
|
343
625
|
|
|
344
|
-
When
|
|
345
|
-
|
|
346
|
-
|
|
626
|
+
When `subagentType` is omitted, the lightweight role router classifies the task
|
|
627
|
+
using the descriptions. Explicit types bypass it. Unknown types or failed
|
|
628
|
+
routing reject the batch, never substitute `defaultType`. Choosing a worker
|
|
629
|
+
model from its candidate list does not involve an LLM call.
|
|
630
|
+
|
|
631
|
+
### Presets are available-model pools
|
|
632
|
+
|
|
633
|
+
Each agent declares an ordered `models` list in Markdown. A preset declares
|
|
634
|
+
which model references may be used, not another role/model/thinking matrix.
|
|
635
|
+
Selection preserves agent order, intersects it with `preset.models`, checks
|
|
636
|
+
runtime registration/auth availability, and takes the first usable candidate.
|
|
637
|
+
Pool order does not change preference and pool-only models are never appended.
|
|
638
|
+
Without a preset, the full agent list is eligible. Candidate order expresses
|
|
639
|
+
the configured budget preference; runtime does not infer current API prices.
|
|
640
|
+
|
|
641
|
+
Image-bearing tasks and `browser-qa` require confirmed image support; configured
|
|
642
|
+
blind-model masks override runtime image metadata. Remaining eligible models
|
|
643
|
+
form the quota fallback chain, so neither quota history nor image fallback can
|
|
644
|
+
escape the pool. Antigravity account rotation still happens before provider
|
|
645
|
+
fallback. No match, no usable model, or an explicitly empty list rejects the
|
|
646
|
+
whole batch before run directories or child processes are created. A new custom
|
|
647
|
+
agent must supply candidates instead of silently inheriting the parent model.
|
|
648
|
+
|
|
649
|
+
Oracle uses its separate strong-model list and prefers another provider when
|
|
650
|
+
available, but also respects the pool. A same-provider choice is allowed when
|
|
651
|
+
the pool offers no alternative; cross-provider independence is not guaranteed.
|
|
652
|
+
Explicit task/CLI model overrides and `FORCE_CURRENT_MODEL` remain deliberate
|
|
653
|
+
escape hatches and disable automatic model fallback for that task. They do not
|
|
654
|
+
bypass the image-capability check.
|
|
655
|
+
|
|
656
|
+
Define pools in the shared or project `pi-tools-suite.jsonc`. Select a saved
|
|
657
|
+
pool with `/subagent-preset`; use `AGENTS_PRESET=<name>` or
|
|
658
|
+
`/subagent-preset session <name>` for a process-only override and
|
|
659
|
+
`/subagent-preset session-clear` to remove it. The saved selection lives in
|
|
660
|
+
`~/.pi/agent/subagent-preset-selection.json`. `/subagent-preset init` inserts the
|
|
661
|
+
sample only when config is missing. The shipped pools are `cheap` (GLM), `gpt`,
|
|
662
|
+
and `deep` (the retained legacy name for the mixed pool, not worker escalation).
|
|
663
|
+
Initial user config and the sample share one source; descriptions and worker
|
|
664
|
+
model order exist only in the agent files. Existing user files are not rewritten.
|
|
347
665
|
|
|
348
666
|
Example shared async-subagents config section:
|
|
349
667
|
|
|
350
668
|
```jsonc
|
|
351
669
|
{
|
|
352
670
|
"asyncSubagents": {
|
|
353
|
-
"defaultType": "
|
|
671
|
+
"defaultType": "research",
|
|
354
672
|
"routing": {
|
|
355
673
|
"enabled": true,
|
|
356
|
-
"model": "zai/glm-
|
|
674
|
+
"model": "zai/glm-5-turbo",
|
|
357
675
|
"timeoutMs": 12000
|
|
358
676
|
},
|
|
359
677
|
"presets": {
|
|
360
678
|
"cheap": {
|
|
361
|
-
"description": "
|
|
362
|
-
"
|
|
363
|
-
"quick": { "model": "zai/glm-5.3", "thinking": "off" },
|
|
364
|
-
"frontend": { "model": "zai/glm-5.3-flash", "thinking": "medium" },
|
|
365
|
-
"browser-qa": { "model": "zai/glm-5.3-flash", "fallbackModels": ["openai-codex/gpt-5.6-luna"], "thinking": "low" },
|
|
366
|
-
"review": { "model": "zai/glm-5.3", "thinking": "high" }
|
|
367
|
-
}
|
|
679
|
+
"description": "GLM workers with a strong oracle candidate.",
|
|
680
|
+
"models": ["zai/glm-5-turbo", "zai/glm-5.3-flash", "zai/glm-5.3"]
|
|
368
681
|
}
|
|
369
682
|
},
|
|
370
683
|
"types": {
|
|
371
|
-
"
|
|
372
|
-
"
|
|
373
|
-
"thinking": "
|
|
374
|
-
},
|
|
375
|
-
"review": {
|
|
376
|
-
"description": "Use for review/audit of existing code or changes: correctness, security, performance, maintainability, API risks, quality. Do not implement new code.",
|
|
377
|
-
"thinking": "high"
|
|
684
|
+
"research": {
|
|
685
|
+
"models": ["zai/glm-5-turbo", "openai-codex/gpt-5.6-luna"],
|
|
686
|
+
"thinking": "low"
|
|
378
687
|
}
|
|
379
688
|
}
|
|
380
689
|
}
|
|
381
690
|
}
|
|
382
691
|
```
|
|
383
692
|
|
|
384
|
-
###
|
|
693
|
+
### Legacy configuration compatibility
|
|
385
694
|
|
|
386
|
-
|
|
695
|
+
Old built-in role names are no longer implicit aliases. `quick`, `scan`,
|
|
696
|
+
`review`, `deep`, `docs`, `frontend`, and `tests` are valid only when explicitly
|
|
697
|
+
defined as ordinary custom/project types. Old preset per-role keys likewise
|
|
698
|
+
apply only when a type with that exact name exists.
|
|
387
699
|
|
|
388
|
-
|
|
700
|
+
Legacy `model` plus `fallbackModels` remains readable. `models` is a complete
|
|
701
|
+
replacement list: it clears inherited legacy model/fallback/parent routing.
|
|
702
|
+
A later old-format model override still replaces the primary candidate, and a
|
|
703
|
+
later `fallbackModels` replaces the remaining candidates; `[]` disables them.
|
|
704
|
+
Old `modelByParent` configs remain supported, but ordinary roles give legacy
|
|
705
|
+
preset models precedence. New built-ins contain no parent-tier escalation maps.
|
|
389
706
|
|
|
390
|
-
|
|
391
|
-
|
|
392
|
-
|
|
393
|
-
|
|
394
|
-
|
|
395
|
-
"thinking": "max",
|
|
396
|
-
"modelByParent": {
|
|
397
|
-
"zai/*": { "model": "openai-codex/gpt-5.6-sol", "fallbackModels": ["zai/glm-5.3"] },
|
|
398
|
-
"openai-codex/*": "zai/glm-5.3",
|
|
399
|
-
"antigravity/*": { "model": "zai/glm-5.3", "fallbackModels": ["openai-codex/gpt-5.6-sol"] },
|
|
400
|
-
"anthropic/*": { "model": "openai-codex/gpt-5.6-sol", "fallbackModels": ["zai/glm-5.3"] }
|
|
401
|
-
}
|
|
402
|
-
}
|
|
403
|
-
```
|
|
404
|
-
|
|
405
|
-
With this config a GLM parent (`zai/*`) spawns the oracle on `gpt-5.6-sol`, a GPT parent (`openai-codex/*`) spawns it on `glm-5.3`, and so on — automatically, at spawn time, with no `task.model` needed. The parent model ref is read from the spawn context (`ctx.model`) and passed into resolution. Pattern matching is case-insensitive `*` glob (same engine as `vision.blindModelPatterns`). When no key matches (or no parent model is known), the role falls back to its static `model` + `fallbackModels`. An explicit `task.model` or `ASYNC_SUBAGENTS_FORCE_CURRENT_MODEL=1` still overrides the match.
|
|
707
|
+
When a preset specifies `models`, it is exclusively a pool; inherited legacy
|
|
708
|
+
`model`, `types`, thinking, arguments and timeout overrides do not run. A later
|
|
709
|
+
explicit old-format preset selector can still replace a pool for compatibility.
|
|
710
|
+
Runtime retry structures and the separate role router continue to use the
|
|
711
|
+
term `fallbackModels` for actual fallback-only lists, not agent candidates.
|
|
406
712
|
|
|
407
713
|
Sub-agents run with `--no-session` by default to avoid writing duplicate Pi session JSONL files for fire-and-forget background work. Set `ASYNC_SUBAGENTS_ENABLE_SESSIONS=1` to restore persisted per-agent sessions under each agent's `sessions/` directory; this also registers the session-navigation slash commands (`/sub-open`, `/sub-back`, `/sub-where`) needed for switching and deeper post-mortem navigation.
|
|
408
714
|
|
|
@@ -520,6 +826,40 @@ npm run test:prompt-evals:dcp
|
|
|
520
826
|
|
|
521
827
|
The default live model is `zai/glm-5-turbo`. Override it for the whole suite with `PI_TOOLS_SUITE_E2E_MODEL=provider/model`, or use the existing component variables such as `TOOL_SELECTION_E2E_MODEL`, `ASYNC_SUBAGENTS_MODEL`, `ASYNC_SUBAGENTS_ROUTING_E2E_MODEL`, and `DCP_SUMMARY_E2E_MODEL`. The normal deterministic coverage remains `npm test`; run prompt evals after changing tool descriptions, routing/classifier prompts, DCP summary prompts, or the default evaluation model.
|
|
522
828
|
|
|
829
|
+
### Unified eval harness
|
|
830
|
+
|
|
831
|
+
`test/evals/` adds a shared deterministic + live-model eval layer. The coverage
|
|
832
|
+
gate requires every registered extension and model-facing tool to have a
|
|
833
|
+
deterministic contract. The initial live corpus contains 20 cases across tool
|
|
834
|
+
selection, coding quality, orchestration/escalation, and negative overuse
|
|
835
|
+
controls. Coding-quality fixtures use executable behavioral checks rather than
|
|
836
|
+
an LLM judge, while reports compare parent/worker tokens, provider-reported cost,
|
|
837
|
+
tool calls, changed files, and elapsed time.
|
|
838
|
+
|
|
839
|
+
```bash
|
|
840
|
+
# Deterministic coverage/contract gate only
|
|
841
|
+
npm run test:evals:contracts
|
|
842
|
+
|
|
843
|
+
# Live matrix as Bun tests. Models are comma/semicolon separated.
|
|
844
|
+
PI_TOOLS_SUITE_EVAL_MODELS='zai/glm-5.3,openai-codex/gpt-5.6-luna,openai-codex/gpt-5.6-terra,openai-codex/gpt-5.6-sol' \
|
|
845
|
+
npm run test:evals:live
|
|
846
|
+
|
|
847
|
+
# Produce JSON + Markdown comparison artifacts.
|
|
848
|
+
PI_TOOLS_SUITE_EVAL_MODELS='zai/glm-5.3,openai-codex/gpt-5.6-terra,openai-codex/gpt-5.6-sol' \
|
|
849
|
+
npm run evals:report
|
|
850
|
+
|
|
851
|
+
# Focus the report runner when iterating
|
|
852
|
+
PI_TOOLS_SUITE_EVAL_MODELS='zai/glm-5.3' \
|
|
853
|
+
PI_TOOLS_SUITE_EVAL_CATEGORIES='coding-quality,negative' \
|
|
854
|
+
npm run evals:report
|
|
855
|
+
```
|
|
856
|
+
|
|
857
|
+
Live evals are opt-in. The deterministic coverage registry is part of normal
|
|
858
|
+
tests, so adding an extension or tool without eval coverage fails the gate. See
|
|
859
|
+
[`docs/evals.md`](docs/evals.md) for the architecture, complete 20-case catalog,
|
|
860
|
+
fixtures, assertions, metrics, model matrix, report format, environment
|
|
861
|
+
variables, CI recommendations, and the procedure for adding new evals.
|
|
862
|
+
|
|
523
863
|
Supporting docs and historical standalone README content are kept in `docs/`; third-party license texts are kept in `licenses/`.
|
|
524
864
|
|
|
525
865
|
## SDK pin
|
|
@@ -4,25 +4,32 @@
|
|
|
4
4
|
|
|
5
5
|
Provide a cheap, fast `browser-qa` async-subagent that reproduces browser bugs
|
|
6
6
|
and proves fixes with deterministic assertions plus screenshot, video, and trace
|
|
7
|
-
evidence.
|
|
8
|
-
`openai-codex/gpt-5.6-luna
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
7
|
+
evidence. Its ranked `models` list prefers `zai/glm-5.3-flash`, then
|
|
8
|
+
`openai-codex/gpt-5.6-luna`, filtered by the active preset's model pool and
|
|
9
|
+
confirmed runtime image support.
|
|
10
|
+
|
|
11
|
+
## Inline agent workflow and skill isolation
|
|
12
|
+
|
|
13
|
+
- All operating instructions, flow contracts, scenario-design guidance, and
|
|
14
|
+
auth-scaffolding rules live in `src/async-subagents/agents/browser-qa.md`.
|
|
15
|
+
Its body becomes the QA child's `promptAppend` through the shared agent
|
|
16
|
+
loader. Parent and router catalogs include only its short `description`.
|
|
17
|
+
- Runner code, vendor dependencies/licenses, and optional JSONC examples live
|
|
18
|
+
under `src/async-subagents/agents/browser-qa/`; none is a discoverable skill.
|
|
14
19
|
- Sub-agent processes disable normal extension discovery, then always load the
|
|
15
|
-
suite's model-tools
|
|
16
|
-
|
|
17
|
-
Antigravity-backed role unavailable.
|
|
20
|
+
suite's model-tools extension. They load the Antigravity provider extension
|
|
21
|
+
only when an Antigravity model is explicitly selected.
|
|
18
22
|
- A type profile may declare `isolatedSkills`. Spawning that profile adds
|
|
19
23
|
`--no-skills` followed by one explicit `--skill` per configured path.
|
|
20
|
-
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
24
|
+
- `browser-qa` always disables normal skill discovery and filters skill flags
|
|
25
|
+
out of `extraArgs`, even without configured skills. It no longer injects a
|
|
26
|
+
mandatory QA skill. Explicitly configured skills are optional additions.
|
|
27
|
+
- The launcher sets `PI_BROWSER_QA_RUNNER` to the absolute installed runner path,
|
|
28
|
+
replacing inherited values, and strips it from ordinary child environments.
|
|
29
|
+
QA invokes `node "$PI_BROWSER_QA_RUNNER"` without guessing paths from cwd.
|
|
30
|
+
- Model-only profile overrides inherit the workflow. An explicit profile
|
|
31
|
+
`promptAppend` replaces the body like any other agent profile; it is not an
|
|
32
|
+
immutable security boundary. Runtime protections remain in the runner.
|
|
26
33
|
|
|
27
34
|
## Authentication contract
|
|
28
35
|
|
|
@@ -98,7 +105,7 @@ evidence. The role uses `zai/glm-5.3-flash`, falling back to
|
|
|
98
105
|
|
|
99
106
|
## Reliability and shutdown contract
|
|
100
107
|
|
|
101
|
-
- The built-in `browser-qa` profile has a
|
|
108
|
+
- The built-in `browser-qa` profile has a 300-second wall-clock budget unless
|
|
102
109
|
the caller explicitly supplies a task or spawn timeout. This bounds model
|
|
103
110
|
stalls as well as browser work.
|
|
104
111
|
- The trusted runner has its own bounded lifecycle. Browser launch, context
|
|
@@ -124,11 +131,14 @@ evidence. The role uses `zai/glm-5.3-flash`, falling back to
|
|
|
124
131
|
|
|
125
132
|
## Acceptance criteria
|
|
126
133
|
|
|
127
|
-
1. `browser-qa` resolves to the intended model/fallback and its
|
|
128
|
-
|
|
134
|
+
1. `browser-qa` resolves to the intended model/fallback and its inline Markdown
|
|
135
|
+
workflow, and its isolated child process can register the configured
|
|
129
136
|
model provider.
|
|
130
|
-
2.
|
|
131
|
-
|
|
137
|
+
2. Default QA spawn args contain `--no-skills` but no `--skill`. The child
|
|
138
|
+
receives the full workflow in its initial prompt and can invoke the bundled
|
|
139
|
+
runner through `PI_BROWSER_QA_RUNNER` from an unrelated project directory.
|
|
140
|
+
Optional configured skills still load; ordinary profiles retain existing
|
|
141
|
+
skill discovery behavior and do not receive QA-only environment paths.
|
|
132
142
|
3. Auth profile listing and all error output are redacted; model-authored input
|
|
133
143
|
cannot execute code in the credential-bearing process.
|
|
134
144
|
4. Runner tests cover public execution without an auth file, explicit profile
|