pi-ui-extend 1.0.37 → 1.0.39

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (44) hide show
  1. package/README.md +1 -0
  2. package/dist/app/extensions/extension-ui-controller.js +3 -0
  3. package/dist/app/rendering/extension-entry-renderer.js +2 -0
  4. package/dist/bundled-extensions/question/contract.d.ts +3 -0
  5. package/dist/bundled-extensions/question/contract.js +24 -0
  6. package/dist/bundled-extensions/question/desktop.d.ts +11 -0
  7. package/dist/bundled-extensions/question/desktop.js +143 -0
  8. package/dist/bundled-extensions/question/index.d.ts +1 -0
  9. package/dist/bundled-extensions/question/index.js +5 -1
  10. package/dist/bundled-extensions/question/render.js +6 -0
  11. package/dist/bundled-extensions/question/result.js +71 -9
  12. package/dist/bundled-extensions/question/tool-description.js +4 -3
  13. package/dist/bundled-extensions/question/tui.js +127 -12
  14. package/dist/bundled-extensions/question/types.d.ts +23 -2
  15. package/dist/tool-renderers/question.js +20 -1
  16. package/docs/desktop-markdown-media.md +77 -0
  17. package/docs/desktop-mvp.md +134 -0
  18. package/docs/desktop-task-manager.md +124 -0
  19. package/external/pi-tools-suite/package.json +3 -3
  20. package/external/pi-tools-suite/src/async-subagents/async-subagents.sample.jsonc +9 -1
  21. package/external/pi-tools-suite/src/async-subagents/core/agent-strategy.ts +41 -2
  22. package/external/pi-tools-suite/src/async-subagents/core/config.ts +13 -1
  23. package/external/pi-tools-suite/src/async-subagents/index.ts +6 -2
  24. package/external/pi-tools-suite/src/async-subagents/private-skills/browser-qa/SKILL.md +65 -22
  25. package/external/pi-tools-suite/src/async-subagents/private-skills/browser-qa/references/qa-design.md +23 -8
  26. package/external/pi-tools-suite/src/async-subagents/private-skills/browser-qa/scripts/browser-qa-runner.mjs +180 -41
  27. package/external/pi-tools-suite/src/async-subagents/subagent-overlay.ts +1 -1
  28. package/external/pi-tools-suite/src/dcp/auto-compress.ts +1 -1
  29. package/external/pi-tools-suite/src/dcp/compression-blocks.ts +1 -51
  30. package/external/pi-tools-suite/src/dcp/debug-log.ts +6 -0
  31. package/external/pi-tools-suite/src/dcp/index.ts +27 -125
  32. package/external/pi-tools-suite/src/dcp/prompts.ts +2 -2
  33. package/external/pi-tools-suite/src/dcp/provider-tool-results.ts +2 -1
  34. package/external/pi-tools-suite/src/dcp/pruner-candidates.ts +31 -10
  35. package/external/pi-tools-suite/src/dcp/pruner-compression-blocks.ts +6 -7
  36. package/external/pi-tools-suite/src/dcp/pruner-message-ids.ts +153 -111
  37. package/external/pi-tools-suite/src/dcp/pruner-metadata.ts +28 -17
  38. package/external/pi-tools-suite/src/dcp/pruner-nudge.ts +61 -27
  39. package/external/pi-tools-suite/src/dcp/pruner.ts +21 -13
  40. package/external/pi-tools-suite/src/dcp/state.ts +59 -0
  41. package/external/pi-tools-suite/src/default-pi-tools-suite-config.ts +4 -0
  42. package/external/pi-tools-suite/src/lib/rpc-session-state.ts +34 -0
  43. package/external/pi-tools-suite/src/todo/todo.ts +8 -3
  44. package/package.json +10 -4
@@ -0,0 +1,124 @@
1
+ # Spec: Desktop Project Task Manager
2
+
3
+ ## Type
4
+
5
+ Change
6
+
7
+ ## Goal
8
+
9
+ Add a project-scoped task list to Pix Desktop in a persistent left sidebar and
10
+ allow a saved task to start work immediately in a new session tab.
11
+
12
+ ## Scope
13
+
14
+ - A collapsible, resizable left sidebar with `Tasks` and `Project` tabs.
15
+ - A compact task list with create, edit, delete, filtering, and manual status
16
+ changes.
17
+ - Task type (`bug`, `feature`, `improvement`), status (`backlog`, `todo`,
18
+ `in-progress`, `done`), and priority (`low`, `medium`, `high`, `urgent`).
19
+ - Project-local persistence in `.pi/tasks.json` with a versioned schema.
20
+ - Starting a task in a new ACP session, automatically sending a generated
21
+ prompt, marking the task in progress, and storing the new session id.
22
+ - Opening an already-linked session instead of creating a duplicate.
23
+
24
+ ## Non-goals
25
+
26
+ - Kanban, drag and drop, subtasks, dependencies, assignees, due dates,
27
+ comments, or task history.
28
+ - Automatic transition to `done` when an agent turn finishes.
29
+ - Synchronizing the project task file with the agent's session-local todo list.
30
+ - Concurrent multi-window file merging.
31
+
32
+ ## Behavior
33
+
34
+ 1. The sidebar is visible below the title/session chrome, starts on `Tasks`,
35
+ remembers its width and collapsed state locally, and preserves the main
36
+ session workspace at narrow window sizes.
37
+ 2. Task rows show a title, type, status, priority, and actions. Status and
38
+ priority are communicated by a Lucide icon, text, and color so meaning never
39
+ depends on color alone. `done` is green, `in-progress` is yellow, and idle
40
+ states are neutral; all colors support light and dark themes.
41
+ 3. Creating or editing validates a non-empty title. A description is optional.
42
+ 4. Filters affect only the visible list and do not modify persisted ordering.
43
+ 5. A missing `.pi/tasks.json` is treated as an empty version-1 document.
44
+ 6. Starting an unlinked task creates and selects a new ACP session, persists
45
+ its id and `in-progress` status, appends the generated user prompt to the
46
+ transcript, and sends it immediately.
47
+ 7. Starting a linked task loads that session. A missing/stale linked session is
48
+ reported as a recoverable error and does not silently create another one.
49
+ 8. Task completion remains manual.
50
+
51
+ ## Contracts
52
+
53
+ - Project file: `.pi/tasks.json` containing `{ version: 1, tasks: [...] }`.
54
+ - Every task has a unique id, title, type, status, priority, `createdAt`, and
55
+ `updatedAt`; description and `sessionId` may be omitted.
56
+ - Tauri exposes narrow read/write commands for this one file. Workspace paths
57
+ are validated and `.pi` symlinks may not escape the workspace.
58
+ - The webview sends a complete validated document on each mutation. Writes use
59
+ a same-directory temporary file and replacement to avoid partial JSON.
60
+
61
+ ## Invariants
62
+
63
+ - The task document version must be supported before it is displayed or saved.
64
+ - Duplicate ids, unknown enum values, empty titles, and malformed timestamps
65
+ are rejected.
66
+ - A task is linked to at most one session in this MVP.
67
+ - Starting a task is disabled while another prompt/session operation is active.
68
+ - A failed save does not leave the UI claiming a task state that was not
69
+ persisted.
70
+
71
+ ## Edge cases
72
+
73
+ - Switching workspaces discards the previous workspace's in-memory task view
74
+ and loads the new project file.
75
+ - Empty and missing files have distinct behavior: missing means no tasks;
76
+ malformed or empty JSON is an error.
77
+ - Save, session creation, and prompt failures remain visible and retryable.
78
+ - Collapsing the sidebar keeps its tab rail available; choosing a tab expands
79
+ it.
80
+
81
+ ## Related files
82
+
83
+ - `desktop/src/App.svelte`
84
+ - `desktop/src/components/`
85
+ - `desktop/src/lib/`
86
+ - `desktop/src/styles.css`
87
+ - `desktop/src-tauri/src/lib.rs`
88
+
89
+ ## Verification
90
+
91
+ - Unit tests for task parsing, validation, filtering, prompt generation, and
92
+ sidebar preference parsing.
93
+ - Rust tests for missing/read/write/malformed files and workspace escape.
94
+ - `npm --prefix desktop test`
95
+ - `npm --prefix desktop run check`
96
+ - `npm --prefix desktop run build:web`
97
+ - `cargo test --manifest-path desktop/src-tauri/Cargo.toml`
98
+
99
+ ## Risks / unknowns
100
+
101
+ - Whole-document writes assume one active desktop writer per project.
102
+ - A linked session may later be removed outside Pix Desktop; the first version
103
+ reports this rather than automatically unlinking it.
104
+ - Windows replacement semantics may require a fallback around replacing an
105
+ existing task file while retaining best-effort crash safety.
106
+
107
+ ## Evidence
108
+
109
+ - Confirmed by code: desktop session creation/loading and prompt submission are
110
+ orchestrated in `desktop/src/App.svelte` through `AcpClient`.
111
+ - Confirmed by code: Tauri already validates workspace-contained project file
112
+ reads in `desktop/src-tauri/src/lib.rs`.
113
+ - Confirmed by tests: session tab ordering and ACP request behavior have focused
114
+ Vitest coverage under `desktop/src/lib/`.
115
+ - Confirmed by docs: `DESIGN.md` specifies a compact persistent sidebar,
116
+ semantic theme roles, Lucide icons, and non-color status cues.
117
+ - Implemented: `.pi/tasks.json` now has an explicit version-1 contract and
118
+ rejects unsupported versions rather than guessing a migration.
119
+ - Verified by tests: all 82 desktop Vitest tests and all 8 Rust tests pass;
120
+ Svelte/TypeScript checks report no errors or warnings.
121
+ - Verified visually: browser QA covered clean preference state, resize/collapse,
122
+ tabs, filters, and semantic colors in light/dark themes.
123
+ - Verified natively: actual Tauri UI/backend CRUD, Run, persisted linkage,
124
+ Open session without duplication, and delete cleanup all passed.
@@ -44,9 +44,9 @@
44
44
  "vscode-languageserver-protocol": "^3.17.5"
45
45
  },
46
46
  "peerDependencies": {
47
- "@earendil-works/pi-ai": "0.84.4",
48
- "@earendil-works/pi-coding-agent": "0.84.4",
49
- "@earendil-works/pi-tui": "0.84.4",
47
+ "@earendil-works/pi-ai": "0.85.0",
48
+ "@earendil-works/pi-coding-agent": "0.85.0",
49
+ "@earendil-works/pi-tui": "0.85.0",
50
50
  "typebox": "*"
51
51
  },
52
52
  "devDependencies": {
@@ -210,7 +210,7 @@
210
210
  "model": "zai/glm-5.3-flash",
211
211
  "fallbackModels": ["openai-codex/gpt-5.6-luna"],
212
212
  "thinking": "low",
213
- "timeoutMs": 120000,
213
+ "timeoutMs": 300000,
214
214
  "tools": ["read", "grep", "bash"]
215
215
  },
216
216
 
@@ -234,6 +234,14 @@
234
234
  "description": "Use when the sub-agent should make or plan code changes for a feature, bug fix, or refactor.",
235
235
  "model": "openai-codex/gpt-5.6-sol",
236
236
  "fallbackModels": ["zai/glm-5.3"],
237
+ // Avoid Sol recursively doing routine implementation work for a Sol parent.
238
+ // Luna also escalates substantial implementation to Terra. modelByParent
239
+ // wins over preset role models, so this prevents both Luna -> Sol and
240
+ // Sol -> Sol for implement tasks under the built-in `gpt` preset.
241
+ "modelByParent": {
242
+ "openai-codex/gpt-5.6-luna*": { "model": "openai-codex/gpt-5.6-terra", "fallbackModels": ["zai/glm-5.3"] },
243
+ "openai-codex/gpt-5.6-sol*": { "model": "openai-codex/gpt-5.6-terra", "fallbackModels": ["zai/glm-5.3"] }
244
+ },
237
245
  "thinking": "high"
238
246
  },
239
247
 
@@ -1,6 +1,6 @@
1
1
  import { isGptLikeModel } from "./ultrawork-auto.js";
2
2
 
3
- export type AgentStrategyName = "parallel-first" | "deep-work";
3
+ export type AgentStrategyName = "parallel-first" | "deep-work" | "escalation-aware" | "cost-aware-orchestrator";
4
4
 
5
5
  export interface AgentStrategyOptions {
6
6
  modelRef?: string;
@@ -27,13 +27,40 @@ Default: autonomous deep worker. Build context directly, make progress, edit, an
27
27
  For broad work, keep delegation explicit and bounded: focused review/research/tests/frontend/deep tracks, plus one oracle only for high-stakes uncertainty or final plan checks. Read compact results, decide in the parent session, and report only what matters. If compressing unfinished work, preserve active objective + next step via todo/DCP rules.
28
28
  </agent_strategy>`;
29
29
 
30
+ const ESCALATION_AWARE_STRATEGY_PROMPT = `<agent_strategy name="escalation-aware">
31
+ Execution hint for Pi, not a replacement for system/developer/user instructions.
32
+
33
+ Default: self-sufficient with tiered escalation. Solve narrow, well-bounded work directly, but do not grind through a large or uncertain task at the current model tier when a stronger focused subagent is appropriate. Use todo for the plan and async subagents for escalation; read compact results first and keep the parent context lean.
34
+
35
+ For a Luna parent, prefer Terra workers for substantial multi-file research, tests, or implementation, and escalate deep root-cause analysis, architecture/security review, high-risk decisions, or repeatedly failing complex work to Sol through the deep/review roles. For a Terra parent, handle routine research/tests/implementation directly; escalate deep root-cause analysis, architecture/security review, high-risk decisions, or stubborn complex failures to Sol through deep/review. Do not escalate merely because a plan has several steps, and do not delegate a tiny known-file edit or exact lookup.
36
+
37
+ Keep user questions, plan/todo changes, integration decisions, and the final report in the parent. Independent read-only escalations may run in parallel; serialize overlapping edits unless scopes are clearly disjoint. When work is delegated, synchronize its todo lifecycle: mark it in progress, collect and verify the worker result, then complete/update it before moving on.
38
+ </agent_strategy>`;
39
+
40
+ const COST_AWARE_ORCHESTRATOR_STRATEGY_PROMPT = `<agent_strategy name="cost-aware-orchestrator">
41
+ Execution hint for Pi, not a replacement for system/developer/user instructions.
42
+
43
+ Default: cost-aware orchestration. You are an expensive parent model, so keep the parent session focused on planning, decisions, integration, verification, and the final user-facing answer. For non-trivial todo work, prefer focused async subagents for repo scanning, multi-file research, documentation, tests, frontend work, and implementation steps that would otherwise require several repository tool calls. Read compact subagent results first; inspect raw artifacts or redo work in the parent only when verification or uncertainty requires it.
44
+
45
+ Keep user questions, todo/plan changes, architecture tradeoffs, cross-worker integration decisions, high-stakes review, and the final report in the parent. Do not delegate merely to avoid one cheap exact lookup or a tiny known-file edit. Independent read-only tracks may run in parallel; serialize overlapping edits unless scopes are clearly disjoint. When a todo item is delegated, keep its lifecycle synchronized: mark it in progress, collect and verify the worker result, then complete/update it before moving on.
46
+ </agent_strategy>`;
47
+
30
48
  export function agentStrategyPrompt(options: AgentStrategyOptions = {}): string | undefined {
31
49
  const env = options.env ?? process.env;
32
50
  const override = strategyOverride(env);
33
51
  if (override === "off") return undefined;
34
52
  if (options.customPrompt && shouldSkipCustomPrompt(env)) return undefined;
35
53
 
36
- const strategy = override ?? (isGptLikeModel(options.modelRef) ? "deep-work" : "parallel-first");
54
+ const strategy = override
55
+ ?? (isExpensiveGptParent(options.modelRef)
56
+ ? "cost-aware-orchestrator"
57
+ : isEscalationAwareGptParent(options.modelRef)
58
+ ? "escalation-aware"
59
+ : isGptLikeModel(options.modelRef)
60
+ ? "deep-work"
61
+ : "parallel-first");
62
+ if (strategy === "cost-aware-orchestrator") return COST_AWARE_ORCHESTRATOR_STRATEGY_PROMPT;
63
+ if (strategy === "escalation-aware") return ESCALATION_AWARE_STRATEGY_PROMPT;
37
64
  return strategy === "deep-work" ? DEEP_WORK_STRATEGY_PROMPT : PARALLEL_FIRST_STRATEGY_PROMPT;
38
65
  }
39
66
 
@@ -49,10 +76,22 @@ function strategyOverride(env: NodeJS.ProcessEnv): AgentStrategyName | "off" | u
49
76
  if (FALSE_ENV_PATTERN.test(value)) return "off";
50
77
  if (value === "parallel-first") return "parallel-first";
51
78
  if (value === "deep-work") return "deep-work";
79
+ if (value === "escalation" || value === "escalation-aware") return "escalation-aware";
80
+ if (value === "cost-aware" || value === "cost-aware-orchestrator" || value === "orchestrator") return "cost-aware-orchestrator";
52
81
  if (TRUE_ENV_PATTERN.test(value)) return undefined;
53
82
  return undefined;
54
83
  }
55
84
 
85
+ function isExpensiveGptParent(modelRef: string | undefined): boolean {
86
+ if (!modelRef) return false;
87
+ return /(?:^|\/)gpt-5\.6-sol(?:$|[-.:])/i.test(modelRef.trim());
88
+ }
89
+
90
+ function isEscalationAwareGptParent(modelRef: string | undefined): boolean {
91
+ if (!modelRef) return false;
92
+ return /(?:^|\/)gpt-5\.6-(?:luna|terra)(?:$|[-.:])/i.test(modelRef.trim());
93
+ }
94
+
56
95
  function shouldSkipCustomPrompt(env: NodeJS.ProcessEnv): boolean {
57
96
  const raw = firstEnv(env, "PI_AGENT_STRATEGY_WITH_CUSTOM_PROMPT", "ASYNC_SUBAGENTS_AGENT_STRATEGY_WITH_CUSTOM_PROMPT");
58
97
  return raw ? !TRUE_ENV_PATTERN.test(raw.trim()) : true;
@@ -218,12 +218,16 @@ const BUILTIN_CONFIG: SubagentConfig = {
218
218
  model: "zai/glm-5.3-flash",
219
219
  fallbackModels: ["openai-codex/gpt-5.6-luna"],
220
220
  thinking: "low",
221
- timeoutMs: 120_000,
221
+ timeoutMs: 300_000,
222
222
  tools: ["read", "grep", "bash"],
223
223
  isolatedSkills: [getBrowserQaSkillPath()],
224
224
  },
225
225
  implement: {
226
226
  description: "Use when the sub-agent should make or plan code changes for a feature, bug fix, or refactor.",
227
+ modelByParent: {
228
+ "openai-codex/gpt-5.6-luna*": { model: "openai-codex/gpt-5.6-terra", fallbackModels: ["zai/glm-5.3"] },
229
+ "openai-codex/gpt-5.6-sol*": { model: "openai-codex/gpt-5.6-terra", fallbackModels: ["zai/glm-5.3"] },
230
+ },
227
231
  thinking: "high",
228
232
  },
229
233
  tests: {
@@ -234,10 +238,18 @@ const BUILTIN_CONFIG: SubagentConfig = {
234
238
  },
235
239
  review: {
236
240
  description: "Use for review/audit of existing code or changes: correctness, security, performance, maintainability, API risks, quality. Do not implement new code.",
241
+ modelByParent: {
242
+ "openai-codex/gpt-5.6-luna*": { model: "openai-codex/gpt-5.6-sol", fallbackModels: ["zai/glm-5.3"] },
243
+ "openai-codex/gpt-5.6-terra*": { model: "openai-codex/gpt-5.6-sol", fallbackModels: ["zai/glm-5.3"] },
244
+ },
237
245
  thinking: "high",
238
246
  },
239
247
  deep: {
240
248
  description: "Use for broad hard reasoning: architecture, root-cause analysis, cross-module impact, complex debugging or tradeoffs.",
249
+ modelByParent: {
250
+ "openai-codex/gpt-5.6-luna*": { model: "openai-codex/gpt-5.6-sol", fallbackModels: ["zai/glm-5.3"] },
251
+ "openai-codex/gpt-5.6-terra*": { model: "openai-codex/gpt-5.6-sol", fallbackModels: ["zai/glm-5.3"] },
252
+ },
241
253
  thinking: "high",
242
254
  },
243
255
  oracle: {
@@ -32,6 +32,7 @@ import { registerSubagentsTool } from "./tools/subagents.js";
32
32
  import type { LiveAgent, SubagentsLiveStateEvent } from "./types.js";
33
33
  import type { AgentState } from "./core/types.js";
34
34
  import { publishStartupSection } from "../startup-section.js";
35
+ import { publishRpcSessionState } from "../lib/rpc-session-state.js";
35
36
 
36
37
  function isTerminalAgentStatus(status: AgentState["status"]): boolean {
37
38
  return status === "done" || status === "failed" || status === "stopped";
@@ -83,8 +84,8 @@ function createLiveStatePayload(
83
84
  }
84
85
 
85
86
  function agentMatchesSession(agent: LiveAgent, sessionFile: string | undefined): boolean {
86
- if (!sessionFile || !agent.parentSession) return true;
87
- return pathsEqual(sessionFile, agent.parentSession);
87
+ if (!sessionFile) return true;
88
+ return agent.parentSession !== undefined && pathsEqual(sessionFile, agent.parentSession);
88
89
  }
89
90
 
90
91
  function isStaleExtensionContextError(error: unknown): boolean {
@@ -100,6 +101,7 @@ export default function (pi: ExtensionAPI) {
100
101
  const subagentOverlay = new SubagentOverlay(liveAgents);
101
102
  let sawAutoUltraworkCandidate = false;
102
103
  let currentSessionFile: string | undefined;
104
+ let currentSessionStateContext: Parameters<typeof publishRpcSessionState>[0];
103
105
  let completionWatchTimer: ReturnType<typeof setInterval> | undefined;
104
106
  publishSubagentPresetsStartupSection();
105
107
 
@@ -109,6 +111,7 @@ export default function (pi: ExtensionAPI) {
109
111
  const liveState = createLiveStatePayload(liveAgents, currentSessionFile);
110
112
  pi.events?.emit?.(SUBAGENTS_LIVE_COUNT_EVENT, { count: liveState.count });
111
113
  pi.events?.emit?.(SUBAGENTS_LIVE_STATE_EVENT, liveState);
114
+ publishRpcSessionState(currentSessionStateContext, SUBAGENTS_LIVE_STATE_EVENT, liveState);
112
115
  updateCompletionWatcher();
113
116
  } catch (error) {
114
117
  ignoreStaleExtensionContextError(error);
@@ -168,6 +171,7 @@ export default function (pi: ExtensionAPI) {
168
171
  try {
169
172
  sawAutoUltraworkCandidate = false;
170
173
  currentSessionFile = sessionFileFromContext(ctx);
174
+ currentSessionStateContext = ctx;
171
175
  subagentOverlay.restoreRunningAgents(ctx.cwd, currentSessionFile);
172
176
  refreshSubagentOverlay();
173
177
  } catch (error) {
@@ -24,12 +24,18 @@ the requested target in the browser.
24
24
 
25
25
  ## Workflow
26
26
 
27
- Treat target discovery as a 30-second preflight and invoke the runner within 45
28
- seconds of starting. Use at most one runner invocation unless the task explicitly
29
- requests multiple auth profiles. If you cannot identify a reachable target and
30
- a supported deterministic assertion inside that preflight, return a structured
31
- `BLOCKED` result immediately. Do not consume the launcher budget on further
32
- source reading, server polling, capability probing, or retries.
27
+ Treat target discovery as a 30-second preflight and begin browser execution within
28
+ 45 seconds of starting. For a verification task, use at most one actual `run`
29
+ browser execution per profile. Metadata/preparation commands such as `profiles`
30
+ and form-auth `scaffold` do not count as browser verification runs. For an
31
+ explicit exploratory/manual-QA task, use at most three `run` rounds, give each a
32
+ unique `--run-id`, bound each with `--runner-timeout-ms 60000`, and let each round
33
+ test one concrete hypothesis learned from the prior evidence. Stop earlier once
34
+ the requested behavior is explained or no new supported hypothesis remains. If
35
+ you cannot identify a reachable target and a supported deterministic assertion
36
+ inside the preflight, return a structured `BLOCKED` result immediately. Do not
37
+ consume the launcher budget on open-ended source reading, server polling,
38
+ capability probing, or retries.
33
39
 
34
40
  1. Resolve `scripts/browser-qa-runner.mjs` relative to this skill.
35
41
  2. Use the launcher-provided `$PI_SUBAGENT_AGENT_DIR/browser-qa/` workspace.
@@ -56,15 +62,27 @@ source reading, server polling, capability probing, or retries.
56
62
  operations documented below; it does not accept expressions or scripts.
57
63
  6. Run public QA with
58
64
  `node <runner> run --base-url <url> --flow <flow.jsonc>`. The URL's exact
59
- origin becomes the fail-closed allowlist. Only for authenticated QA, add
60
- `--profile <id>`; the selected profile then owns the URL and allowlist.
65
+ origin becomes the fail-closed allowlist. When the real app requires known
66
+ API/CDN origins, add repeatable `--allow-origin <exact-origin>` flags. Each
67
+ value must be an exact `http(s)` origin with no path, credentials, wildcard,
68
+ or inferred sibling domain; undeclared origins remain blocked. Only for
69
+ authenticated QA, add `--profile <id>`; the selected profile then owns the
70
+ URL and allowlist.
61
71
  Profile id, URL, and flow path are non-secret; never pass credentials as
62
72
  arguments or environment variables.
63
73
  7. Report deterministic assertions and every artifact returned by the runner.
64
- For each screenshot, video, trace, or retained download, emit a separate clickable Markdown
65
- link using its `uri` and also show its absolute `path`. Do this for failed
66
- runs too whenever `artifacts` is present; never report only `evidenceDir`.
67
- Visual inspection supplements assertions; it does not replace them.
74
+ For each screenshot, video, trace, or retained download, emit a separate
75
+ clickable Markdown link using its `uri` and also show its absolute `path`.
76
+ Do this for failed runs too whenever `artifacts` is present; never report only
77
+ `evidenceDir`.
78
+ When screenshots are present, inspect at least one representative meaningful
79
+ PNG directly with the `read` tool on its absolute path before claiming visual
80
+ QA. Inspect additional screenshots when they represent materially different
81
+ states or popups. Record `visualInspection: inspected` plus the inspected
82
+ paths and concrete findings. If image reading is unavailable in the active
83
+ model, report `visualInspection: unavailable` and do not claim a visual pass;
84
+ deterministic assertions may still be reported separately. Visual inspection
85
+ supplements assertions; it does not replace them.
68
86
 
69
87
  ## Scenario design
70
88
 
@@ -84,11 +102,14 @@ source reading, server polling, capability probing, or retries.
84
102
  animation/debounce. Set flow `timeoutMs` only as high as the target
85
103
  legitimately needs.
86
104
  - The runner waits after every visible interaction until the document is ready,
87
- requests started by the interaction have finished, and common visible busy
105
+ observes requests started by the interaction for a bounded readiness window,
106
+ and waits for common visible busy
88
107
  markers (including `aria-busy`, progress bars, loading/spinner/skeleton test
89
- ids and classes) disappear. It then keeps the stable state on video for 500
90
- ms. A busy page that does not settle within `timeoutMs` fails instead of
91
- continuing against a skeleton. For an app-specific loader not covered by
108
+ ids and classes) to disappear. EventSource/WebSocket traffic is excluded and
109
+ a long-poll/background request cannot pin readiness for the full flow timeout;
110
+ visible busy UI can. It then keeps the stable state on video for 500 ms. A
111
+ busy page that does not settle within `timeoutMs` fails instead of continuing
112
+ against a skeleton. For an app-specific loader not covered by
92
113
  those conventions, add an explicit `waitFor`/`assertHidden` for that loader
93
114
  and assert the loaded content before interacting with it.
94
115
  - Recorded pointer actions are annotated automatically: clicks and double-clicks
@@ -101,7 +122,8 @@ source reading, server polling, capability probing, or retries.
101
122
  pointer input, and clears before the runner's post-action stable interval
102
123
  completes.
103
124
  - Place `authRejectedIf` immediately after navigation or any transition that
104
- may reveal expired authentication.
125
+ may reveal expired authentication. It must declare `urlIncludes` or a locator;
126
+ an empty rejection check is invalid.
105
127
  - Never weaken an assertion merely to make a failing run pass. If the observed
106
128
  product behavior differs from the expectation, preserve the failure evidence
107
129
  and report the mismatch.
@@ -120,7 +142,7 @@ Supported actions:
120
142
  `uncheck`, `selectOption`, `wheel`, `evaluate`, `dragTo`, `uploadFiles`,
121
143
  `openPopup`, `download`
122
144
  - assertions: `assertVisible`, `assertHidden`, `assertEnabled`,
123
- `assertDisabled`, `assertChecked`, `assertUnchecked`, `assertText`,
145
+ `assertDisabled`, `assertChecked`, `assertUnchecked`, `assertText`, `assertTextContent`,
124
146
  `assertValue`, `assertAttribute`, `assertCount`, `assertURL`,
125
147
  `assertDOMMetric`
126
148
  - evidence/auth: `screenshot`, `authRejectedIf`
@@ -128,6 +150,9 @@ Supported actions:
128
150
  Locators accept one of `testId`, `role` (plus optional `name`), `label`,
129
151
  `placeholder`, `text`, or `css`; add `exact: true` where useful. String
130
152
  assertions require exactly one of `equals` or `includes`.
153
+ `assertText` matches visible, user-facing `innerText` and therefore fails for a
154
+ hidden locator. Use `assertTextContent` only when raw DOM text, including hidden
155
+ content, is intentionally the oracle.
131
156
  `assertAttribute` additionally requires a bounded `attribute` name and is
132
157
  useful for `aria-*`, `data-*`, `href`, and similar observable state. All
133
158
  assertions retry until `timeoutMs` and report generic failures without exposing
@@ -208,7 +233,11 @@ oracle. Raw JavaScript remains intentionally unsupported.
208
233
  ```jsonc
209
234
  {
210
235
  "viewport": { "width": 844, "height": 847 },
211
- "environment": { "locale": "en-GB", "timezoneId": "Europe/London", "colorScheme": "dark" },
236
+ "environment": {
237
+ "locale": "en-GB",
238
+ "timezoneId": "Europe/London",
239
+ "colorScheme": "dark"
240
+ },
212
241
  "steps": [
213
242
  { "action": "goto", "path": "/settings" },
214
243
  { "action": "authRejectedIf", "urlIncludes": "/login" },
@@ -216,7 +245,11 @@ oracle. Raw JavaScript remains intentionally unsupported.
216
245
  "action": "assertVisible",
217
246
  "locator": { "role": "heading", "name": "Settings", "exact": true }
218
247
  },
219
- { "action": "wheel", "locator": { "css": ".settings-panel" }, "deltaY": 500 },
248
+ {
249
+ "action": "wheel",
250
+ "locator": { "css": ".settings-panel" },
251
+ "deltaY": 500
252
+ },
220
253
  {
221
254
  "action": "assertDOMMetric",
222
255
  "locator": { "css": ".settings-panel" },
@@ -226,9 +259,17 @@ oracle. Raw JavaScript remains intentionally unsupported.
226
259
  {
227
260
  "action": "click",
228
261
  "locator": { "testId": "save-settings" },
229
- "expectResponse": { "path": "/api/settings", "method": "PUT", "status": 200 }
262
+ "expectResponse": {
263
+ "path": "/api/settings",
264
+ "method": "PUT",
265
+ "status": 200
266
+ }
267
+ },
268
+ {
269
+ "action": "assertText",
270
+ "locator": { "testId": "toast" },
271
+ "includes": "Saved"
230
272
  },
231
- { "action": "assertText", "locator": { "testId": "toast" }, "includes": "Saved" },
232
273
  { "action": "screenshot", "name": "settings-saved" }
233
274
  ]
234
275
  }
@@ -301,6 +342,8 @@ After any runner invocation that actually performed browser testing, include
301
342
  all non-empty `artifacts.screenshots`, `artifacts.videos`, and
302
343
  `artifacts.traces`, and `artifacts.downloads` groups in the final response.
303
344
  These links are mandatory so the user can open the evidence directly.
345
+ Also include `visualInspection` with `inspected` or `unavailable`; never infer a
346
+ visual pass from a successful runner status alone.
304
347
 
305
348
  See `references/qa-auth.example.jsonc`, `references/qa-flow.example.jsonc`,
306
349
  `references/qa-design.md`, and `references/auth-scaffold-spec.md`.
@@ -19,6 +19,12 @@ proof.
19
19
  When verifying a fix, prefer a focused regression flow over a broad tour of the
20
20
  application. If multiple independent states matter, assert each one explicitly.
21
21
 
22
+ For explicit exploratory/manual QA, keep exploration bounded rather than turning
23
+ it into an open-ended crawl. Run at most three minimal rounds. Each round should
24
+ start from one concrete hypothesis, produce a deterministic observation plus a
25
+ meaningful screenshot, and use that evidence to decide whether another round is
26
+ justified. Verification tasks remain one browser run per profile.
27
+
22
28
  ## Choose resilient locators
23
29
 
24
30
  Prefer locators that match how users and accessibility APIs identify controls:
@@ -42,13 +48,14 @@ by an assertion is enough. Use `waitFor` only when the next operation depends on
42
48
  a distinct attached/detached/visible/hidden transition.
43
49
 
44
50
  After navigation and visible interactions, the runner also waits for DOM
45
- readiness, completion of requests started by that action, disappearance of
46
- common visible `aria-busy`/progress/loading/spinner/skeleton markers, and a
47
- 500 ms stable interval. If those signals remain busy through the flow timeout,
48
- the run fails rather than interacting with a loading shell. This is a safe
49
- baseline, not an application-specific oracle: explicitly wait for a custom
50
- loader to become hidden and assert the loaded content when the application uses
51
- different readiness semantics.
51
+ readiness, tracks requests causally started by that action through a bounded
52
+ readiness window, waits for common visible
53
+ `aria-busy`/progress/loading/spinner/skeleton markers, and keeps a 500 ms stable
54
+ interval. EventSource/WebSocket traffic is ignored, and a long poll is not
55
+ allowed to pin the entire flow timeout. A visible busy indicator may still hold
56
+ readiness until `timeoutMs`. This is a safe baseline, not an application-specific
57
+ oracle: explicitly wait for a custom loader to become hidden and assert the
58
+ loaded content when the application uses different readiness semantics.
52
59
 
53
60
  `waitForTimeout` is bounded to five seconds and should be exceptional—for a
54
61
  known animation, debounce, or externally scheduled transition with no
@@ -58,7 +65,9 @@ behavior. If a normal operation legitimately needs more time, adjust the flow's
58
65
 
59
66
  All assertion actions retry until that timeout. This makes an action followed
60
67
  directly by `assertText`, `assertVisible`, `assertURL`, or another assertion
61
- safe for asynchronously rendered outcomes. Use `assertAttribute` for
68
+ safe for asynchronously rendered outcomes. `assertText` requires the locator to
69
+ be visible and matches rendered `innerText`; use `assertTextContent` only when
70
+ hidden/raw DOM text is deliberately part of the oracle. Use `assertAttribute` for
62
71
  observable state such as `aria-expanded`, `aria-invalid`, or `data-state`
63
72
  instead of reading DOM state through executable JavaScript.
64
73
 
@@ -177,6 +186,12 @@ they explain the chronology without becoming screenshot or assertion oracles.
177
186
  Assertions determine pass/fail; evidence explains it. Preserve and link every
178
187
  artifact group returned on both passed and failed runs.
179
188
 
189
+ For visual QA, do not stop at artifact generation. Open at least one meaningful
190
+ PNG with the model's image-capable `read` path and inspect layout, clipping,
191
+ overlap, state styling, and other visual defects relevant to the scenario. If
192
+ the active model cannot read images, explicitly report visual inspection as
193
+ unavailable rather than treating deterministic assertions as a visual pass.
194
+
180
195
  ## Diagnose failures without weakening the test
181
196
 
182
197
  Classify the first failing step: