pi-agent-browser-native 0.2.71 → 0.2.74

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (52) hide show
  1. package/CHANGELOG.md +46 -0
  2. package/README.md +14 -12
  3. package/dist/extensions/agent-browser/index.js +104 -16
  4. package/dist/extensions/agent-browser/lib/argv-grammar.js +124 -0
  5. package/dist/extensions/agent-browser/lib/command-taxonomy.js +12 -1
  6. package/dist/extensions/agent-browser/lib/electron/cdp.js +2 -2
  7. package/dist/extensions/agent-browser/lib/electron/launch.js +48 -12
  8. package/dist/extensions/agent-browser/lib/input-modes/params.js +96 -98
  9. package/dist/extensions/agent-browser/lib/launch-scoped-flags.js +88 -2
  10. package/dist/extensions/agent-browser/lib/managed-session-capabilities.js +22 -0
  11. package/dist/extensions/agent-browser/lib/managed-session-policy-lock.js +432 -0
  12. package/dist/extensions/agent-browser/lib/managed-session-restore.js +367 -0
  13. package/dist/extensions/agent-browser/lib/managed-session-snapshots.js +367 -0
  14. package/dist/extensions/agent-browser/lib/managed-session-state-policy.js +589 -0
  15. package/dist/extensions/agent-browser/lib/managed-session-storage.js +299 -0
  16. package/dist/extensions/agent-browser/lib/orchestration/batch-stdin.js +35 -0
  17. package/dist/extensions/agent-browser/lib/orchestration/browser-run/artifact-paths.js +9 -2
  18. package/dist/extensions/agent-browser/lib/orchestration/browser-run/diagnostics.js +40 -22
  19. package/dist/extensions/agent-browser/lib/orchestration/browser-run/final-result.js +15 -6
  20. package/dist/extensions/agent-browser/lib/orchestration/browser-run/index.js +54 -33
  21. package/dist/extensions/agent-browser/lib/orchestration/browser-run/managed-session-daemon-policy.js +182 -0
  22. package/dist/extensions/agent-browser/lib/orchestration/browser-run/prepare/direct-anchor-download.js +1 -1
  23. package/dist/extensions/agent-browser/lib/orchestration/browser-run/prepare/network-page-filter.js +1 -1
  24. package/dist/extensions/agent-browser/lib/orchestration/browser-run/prepare/scroll-shims.js +1 -1
  25. package/dist/extensions/agent-browser/lib/orchestration/browser-run/prepare/snapshot-filter.js +1 -1
  26. package/dist/extensions/agent-browser/lib/orchestration/browser-run/prepare.js +625 -429
  27. package/dist/extensions/agent-browser/lib/orchestration/browser-run/process-output.js +136 -56
  28. package/dist/extensions/agent-browser/lib/orchestration/browser-run/session-state.js +28 -40
  29. package/dist/extensions/agent-browser/lib/orchestration/electron-host/index.js +102 -19
  30. package/dist/extensions/agent-browser/lib/orchestration/output-file.js +13 -1
  31. package/dist/extensions/agent-browser/lib/playbook.js +10 -9
  32. package/dist/extensions/agent-browser/lib/process-identity.js +82 -0
  33. package/dist/extensions/agent-browser/lib/process.js +270 -34
  34. package/dist/extensions/agent-browser/lib/results/artifact-manifest.js +5 -3
  35. package/dist/extensions/agent-browser/lib/results/categories.js +21 -2
  36. package/dist/extensions/agent-browser/lib/results/presentation/common.js +2 -1
  37. package/dist/extensions/agent-browser/lib/results/presentation/diagnostics.js +80 -12
  38. package/dist/extensions/agent-browser/lib/results/presentation/managed-list-filter.js +42 -0
  39. package/dist/extensions/agent-browser/lib/results/recovery-actions.js +7 -0
  40. package/dist/extensions/agent-browser/lib/runtime.js +85 -85
  41. package/dist/extensions/agent-browser/lib/session-page-state.js +48 -17
  42. package/dist/extensions/agent-browser/lib/temp.js +13 -25
  43. package/docs/ARCHITECTURE.md +9 -8
  44. package/docs/COMMAND_REFERENCE.md +97 -32
  45. package/docs/ELECTRON.md +10 -10
  46. package/docs/RELEASE.md +3 -2
  47. package/docs/SUPPORT_MATRIX.md +22 -19
  48. package/docs/TOOL_CONTRACT.md +31 -28
  49. package/docs/platform-smoke.md +5 -5
  50. package/package.json +1 -1
  51. package/platform-smoke.config.mjs +3 -1
  52. package/scripts/agent-browser-capability-baseline.mjs +45 -3
@@ -18,13 +18,31 @@ This project intentionally blocks normal `agent-browser` bash usage in most agen
18
18
 
19
19
  <!-- agent-browser-capability-baseline:start upstream-baseline -->
20
20
  <!-- Generated from scripts/agent-browser-capability-baseline.mjs. Run `npm run docs -- command-reference write` to update. Do not edit manually. -->
21
- This reference is baselined to the locally installed `agent-browser 0.32.2` command/help surface, audited against vercel-labs/agent-browser@6ede7a9470ac4b681cabf838af8668b9aa99e957. Upstream `agent-browser` remains the source of truth for command semantics; this file is the local fallback for Pi agent sessions where direct binary help is blocked or discouraged.
21
+ This reference is baselined to the locally installed `agent-browser 0.33.2` command/help surface, audited against vercel-labs/agent-browser@93cdda5709e8861c0c26b0b955d8d746e9fda0d7. Upstream `agent-browser` remains the source of truth for command semantics; this file is the local fallback for Pi agent sessions where direct binary help is blocked or discouraged.
22
22
 
23
23
  The lightweight drift check is `npm run verify -- command-reference`. Run it whenever the installed upstream `agent-browser` version changes or this reference is edited.
24
24
 
25
25
  Use `npm run benchmark:agent-browser` or `npm run verify -- benchmark` before and after agent-facing workflow abstractions to measure task success, tool calls, model-visible output size, stale-ref behavior, artifact success, failure-category coverage, and elapsed-time estimates.
26
26
  <!-- agent-browser-capability-baseline:end upstream-baseline -->
27
27
 
28
+ ### Upstream 0.33.2 rebaseline
29
+
30
+ The 0.33.1–0.33.2 releases harden daemon lifecycle and live streaming without new core page commands:
31
+
32
+ - 0.33.1 ships a default daemon idle timeout of 1 hour (`AGENT_BROWSER_IDLE_TIMEOUT_MS`, default `3600000`; `0` disables). Sessions without a restore key discard transient cookies/tabs on idle shutdown. Headed, Safari/iOS WebDriver, and user-attached browsers are exempt from that default. Tab recovery also revives Memory Saver-discarded tabs on connect/switch/close and reports recovery fields such as `revived` / `dialogBlocked` / `activeTabRevived`.
33
+ - 0.33.2 makes stream frame delivery latest-wins, prioritizes input over frame writes, adds per-client `maxFps` / ack pacing, and adds `AGENT_BROWSER_STREAM_QUALITY`, `AGENT_BROWSER_STREAM_MAX_WIDTH`, and `AGENT_BROWSER_STREAM_MAX_HEIGHT` for screencast bandwidth control.
34
+ - This wrapper keeps its managed-session idle override (default 15 minutes via `PI_AGENT_BROWSER_IMPLICIT_SESSION_IDLE_TIMEOUT_MS` / `AGENT_BROWSER_IDLE_TIMEOUT_MS`) and enables Git-checkout-generation-stable `AGENT_BROWSER_RESTORE` for wrapper-owned managed sessions so SSO cookies survive browser relaunches. Caller-owned `--session` names do not get that inject, and foreign `piab-*` names are rejected unless this extension instance owns the exact namespace/session. Managed daemon inspection uses its own bounded timeout rather than a caller's short `PI_AGENT_BROWSER_PROCESS_TIMEOUT_MS`. Local browser navigation into `.agent-browser` state storage is blocked, including encoded `file:` paths, local-directory file ref interactions, batches, and later capture from a tracked protected target. Set `PI_AGENT_BROWSER_MANAGED_SESSION_RESTORE=0` to disable restore injection. No Eve-specific Pi runtime is added.
35
+
36
+ ### Upstream 0.33.0 rebaseline
37
+
38
+ The 0.32.3–0.33.0 releases add HAR body capture, fix semantic locators, and ship accessibility audits:
39
+
40
+ - 0.32.3 embeds text response bodies in HAR captures by default and adds `network har start --content <mode>` with `text` (default), `all`, and `none`. It also ships the `derive-client` skill for recording traffic and generating a standalone API client from HAR data.
41
+ - 0.32.4 makes `find role` match implicit ARIA roles and browser-computed accessible names (for example `find role heading text --name` against `<h2>`, lists, and banners), keeps case-insensitive substring name matching by default, preserves locator detail in element-not-found errors (`Names seen: …`), and aligns advertised `find` actions to `click, fill, check, hover, text`.
42
+ - 0.33.0 adds `a11y [url]` axe-core accessibility audits (`--tags`, `--selector`, structured JSON) with an embedded offline engine, plus a native fix that revives discarded tabs on tab switch instead of hanging the daemon.
43
+
44
+ This wrapper classifies 0.32.4+ locator-detail misses as `selector-not-found`, treats HAR `--content` and a11y `--tags` as command-scoped value flags, renders compact `a11y` summaries, and documents the skill/HAR/a11y surfaces. No Eve-specific Pi runtime is added.
45
+
28
46
  ### Upstream 0.32.2 rebaseline
29
47
 
30
48
  The 0.32.1 and 0.32.2 releases only update the separate `@agent-browser/eve` integration; the CLI/help/schema surface is unchanged:
@@ -48,7 +66,7 @@ The 0.32.0 rebaseline hardens domain containment and fixes completed-page waits
48
66
 
49
67
  The 0.31.2 rebaseline adds a WebGPU launch preset and periodic restore-state autosaves:
50
68
 
51
- - `--webgpu` (also `AGENT_BROWSER_WEBGPU` or `"webgpu": true` in `agent-browser.json`) enables the upstream platform preset. It uses Metal on macOS, D3D on Windows, and SwiftShader software Vulkan on Linux. The native Pi wrapper passes it through as a launch-scoped optional boolean, so use `sessionMode: "fresh"` after an implicit session exists; `--webgpu false` explicitly disables a config/environment default.
69
+ - `--webgpu` (also `AGENT_BROWSER_WEBGPU`; standalone upstream additionally accepts `"webgpu": true` in `agent-browser.json`) enables the upstream platform preset. Browser-backed native calls reject upstream config files, so use the flag or environment form through this tool. It uses Metal on macOS, D3D on Windows, and SwiftShader software Vulkan on Linux. The native Pi wrapper passes it through as a launch-scoped optional boolean, so use `sessionMode: "fresh"` after an implicit session exists; `--webgpu false` explicitly disables a config/environment default.
52
70
  - WebGPU requires a local browser launch. Upstream rejects enabled WebGPU with `--cdp`, `--auto-connect`, or `-p` / `--provider`. Use `doctor --webgpu` to pixel-check rendering and screenshot capture; use `doctor --webgpu --headed` when validating the headed capture path.
53
71
  - Headless WebGPU screenshots work on macOS. Upstream documents black headless WebGPU canvas captures on Windows and Linux even when in-page rendering succeeds; Windows needs a logged-in headed desktop, while Linux can use `--headed` with automatic Xvfb unless `AGENT_BROWSER_NO_XVFB=1`. Linux software rendering also needs `libvulkan1` and `mesa-vulkan-drivers`.
54
72
  - Restore-enabled sessions now save periodically while the browser remains open, including idle page-driven cookie/storage changes. `AGENT_BROWSER_AUTOSAVE_INTERVAL_MS` defaults to `30000`; `0` disables periodic saves but keeps save-on-close. The existing `--restore-save` policy still controls whether automatic saves are allowed.
@@ -62,7 +80,7 @@ The 0.31.1 rebaseline is a React bugfix release: `react tree`, `react inspect <i
62
80
 
63
81
  The 0.31.0 rebaseline adds restore workflow and namespace/session lifecycle surfaces: `--restore [name]`, `--restore-save <policy>`, restore check flags, `--namespace <name>`, `session id`, and `session info`. The wrapper parses those globals, keeps `--namespace` before `--session`, carries namespace context through managed-session probes and state, and keeps `session id` / `session info` sessionless. Use `agent_browser` with `args: ["session", "id", "--scope", "worktree", "--prefix", "my-skill"]` to derive reusable session ids from inside Pi; use `--restore=<key>` when passing an explicit key that could be confused with a command word.
64
82
 
65
- Runtime probes retain the 0.30.1 `wait --url` fix: `wait --url "**/dashboard"` succeeds after a `pushstate /dashboard`, so `job.assertUrl` delegates exact and glob patterns to upstream `wait --url`. Two old caveats still stand: `find ... uncheck` and `wait <selector> --state hidden|detached` remain advertised by help but fail at runtime. Keep the wrapper's direct `uncheck` passthrough and `wait --fn` disappearance guidance.
83
+ Runtime probes retain the 0.30.1 `wait --url` fix: `wait --url "**/dashboard"` succeeds after a `pushstate /dashboard`, so `job.assertUrl` delegates exact and glob patterns to upstream `wait --url`. Upstream 0.32.4 aligns advertised `find` actions with the dispatcher: `click, fill, check, hover, text` only—use top-level `uncheck <selector-or-ref>` (and `type` / `focus`) instead of `find ... uncheck|type|focus`. One older caveat still stands: `wait <selector> --state hidden|detached` remains advertised by some help paths but fails at runtime, so keep `wait --fn` disappearance guidance.
66
84
 
67
85
  ### Upstream 0.29.1 rebaseline
68
86
 
@@ -148,7 +166,7 @@ Tool parameters (use exactly one of `args`, `semanticAction`, `job`, `qa`, `sour
148
166
 
149
167
  ### Debug, diff, stream, dashboard, and chat families
150
168
 
151
- Upstream also exposes non-core families (`network`, `diff`, `trace` / `profiler` / `record`, `console` / `errors` / `highlight` / `inspect` / `clipboard`, `stream`, `dashboard`, `chat`, and related subcommands). The wrapper still owns argv planning, `--json`, managed sessions where applicable, artifact metadata, and model-facing presentation: structured results are compacted and scrubbed in `extensions/agent-browser/lib/results/presentation.ts`, and echoed argv uses the same `redactInvocationArgs` rules as core commands (see [`TOOL_CONTRACT.md`](TOOL_CONTRACT.md#details) for the field contract). Deterministic fake-upstream coverage for representative JSON shapes and redaction lives in `test/agent-browser.extension-validation.test.ts` under `agentBrowserExtension passes through non-core network debug diff stream dashboard and chat families`.
169
+ Upstream also exposes non-core families (`network`, `diff`, `trace` / `profiler` / `record`, `console` / `errors` / `a11y` / `highlight` / `inspect` / `clipboard`, `stream`, `dashboard`, `chat`, and related subcommands). The wrapper still owns argv planning, `--json`, managed sessions where applicable, artifact metadata, and model-facing presentation: structured results are compacted and scrubbed in `extensions/agent-browser/lib/results/presentation.ts`, and echoed argv uses the same `redactInvocationArgs` rules as core commands (see [`TOOL_CONTRACT.md`](TOOL_CONTRACT.md#details) for the field contract). Deterministic fake-upstream coverage for representative JSON shapes and redaction lives in `test/agent-browser.extension-validation.test.ts` under `agentBrowserExtension passes through non-core network debug diff stream dashboard and chat families`.
152
170
 
153
171
  ## Recommended workflow
154
172
 
@@ -183,7 +201,9 @@ Run `{ "args": ["doctor", "--webgpu"] }` before trusting a black or blank WebGPU
183
201
 
184
202
  Treat headed success as browser-context success, not proof that a window is visible on the user's display. Remote shells, containers, virtual framebuffers, or upstream/provider-owned browser hosts can still put the visible window somewhere the user cannot see. If a user reports no window, gather evidence with `screenshot`, `tab list`, `get url`, or `snapshot -i`; then relaunch with the right display/profile/provider setup rather than assuming the user missed it.
185
203
 
186
- For local fixtures, remember that `localhost` and `127.0.0.1` are resolved from the browser host, which may differ from the shell that started a temporary HTTP server. `net::ERR_EMPTY_RESPONSE` on `http://localhost:<port>` usually means the browser could not reach that server, not that the page itself rendered blank; the wrapper appends a local fixture hint for common loopback navigation failures. Prefer a host-reachable address when your environment provides one; otherwise use `file://` only for static fixtures and note its limits. `file://` does not provide HTTP headers and may change MIME/CORS/storage/debugger behavior. If `eval --stdin` on a `file://` page returns `null` for even simple DOM expressions, first make sure the JavaScript is in the native tool `stdin` field rather than trailing after `--stdin` in `args`; then treat the result as inconclusive and verify with `snapshot -i`, `get text` on current refs, or screenshots until the fixture can run over reachable HTTP.
204
+ For local fixtures, remember that `localhost` and `127.0.0.1` are resolved from the browser host, which may differ from the shell that started a temporary HTTP server. `net::ERR_EMPTY_RESPONSE` on `http://localhost:<port>` usually means the browser could not reach that server, not that the page rendered blank; the wrapper appends a local fixture hint for common loopback failures. Prefer an environment-specific host-reachable HTTP(S) address. Do not switch to `file://`: the native wrapper blocks content-returning local-URL calls plus follow-up inspection, scripting, interaction, and Electron probes on local file pages to protect authenticated `.agent-browser` state. Protected artifact destinations and top-level `outputPath` also fail before browser spawn or directory creation.
205
+
206
+ For a caller-owned explicit `--session`, content-bearing reads and interactions first run a session-scoped `get url`; missing or stale transcript page state is not trusted. If the probe fails or resolves to a protected local target, the requested content command does not run. Calls to the same effective canonical namespace/session are serialized inside one extension instance; explicit namespace argv overrides inherited `AGENT_BROWSER_NAMESPACE`, including an explicit empty default across preparation helpers, that live probe, any semantic-action snapshot, and the requested command; different caller-owned sessions can still overlap. This does not coordinate direct `agent-browser` calls or another Pi process. Windows drive-relative forms such as `C:.agent-browser\\state\\...` are paths, not URL schemes. Nested `batch` steps are rejected. Raw batch command strings mirror upstream's ASCII-space tokenizer, including its single/double-quote and backslash handling; other Unicode whitespace remains part of a token. When later batch content depends on navigation, use exact `batch --bail` or split the calls: a non-bail batch is rejected if any failed transition could leave a local or unverified page active. Non-bail diagnostics remain available when every retained target is already verified safe.
187
207
 
188
208
  Temporary HTTP servers and their port/process lifecycle stay outside the native tool. Extension maintainers running real-upstream contract tests can reuse `startAgentBrowserContractFixtureServer()` in [`test/helpers/agent-browser-harness.ts`](https://github.com/fitchmultz/pi-agent-browser-native/blob/main/test/helpers/agent-browser-harness.ts) instead of ad-hoc `python3 -m http.server` processes.
189
209
 
@@ -242,7 +262,7 @@ Examples:
242
262
  { "args": ["snapshot", "-i"] }
243
263
  ```
244
264
 
245
- The optional native `semanticAction` object is only a thin schema for common locator-based actions, direct selector/ref click/check/fill, and native dropdown selection; it compiles locator actions to existing upstream `find` commands, direct selector/ref actions to `click` / `check` / `fill`, compiles `action: "select"` to upstream `select <selector> <value...>`, and reports the compiled argv in `details.compiledSemanticAction` (see [`TOOL_CONTRACT.md`](TOOL_CONTRACT.md#semanticaction) for the full field rules). For `locator: "role"`, pass either `value: "button"` or `role: "button"`; if both are present they must match. It is a top-level alternative to `args`, `job`, `qa`, `sourceLookup`, `networkSourceLookup`, and `electron`, not a nested shape inside `batch` stdin arrays. Add `session` inside `semanticAction` when the shorthand should target a named upstream browser session; the compiled argv prepends `--session <name>` before `find`, direct selector/ref commands, or `select`, and fallback candidate actions preserve that prefix. For active sessions, role/name click/check/fill shorthands may resolve through the current `snapshot -i` refs before execution so hidden duplicate matches do not steal the action; fill only resolves when there is one exact editable current ref match. Inspect `details.effectiveArgs` when you need the exact executed argv. `semanticAction` does not expose `uncheck` while upstream `find ... uncheck` is not runtime-supported; use raw `uncheck <selector-or-ref>` after choosing a stable selector or current snapshot ref. `select` shorthand intentionally requires a stable selector or current `@ref` plus `value`/`values`; upstream `find` does not expose a verified `select` action, so role/name/label dropdown resolution stays a snapshot/selector decision instead of hidden wrapper magic. If a raw `find` or semantic action misses with `selector-not-found`, the wrapper may take one fresh snapshot and append `Current snapshot ref fallback` when that snapshot has exact visible role/name matches for the failed target. Non-fill matches can include direct `try-current-visible-ref*` next actions. Semantic click misses may also include `Agent-browser candidate fallbacks`; `details.nextActions` first recommends a fresh `snapshot -i` and may include bounded role/name retries such as `button`/`link` for a missed `text` click, each as a `try-*-candidate` entry carrying redacted `find role …` argv.
265
+ The optional native `semanticAction` object is only a thin schema for common locator-based actions, direct selector/ref click/check/fill, and native dropdown selection; it compiles locator actions to existing upstream `find` commands, direct selector/ref actions to `click` / `check` / `fill`, compiles `action: "select"` to upstream `select <selector> <value...>`, and reports the compiled argv in `details.compiledSemanticAction` (see [`TOOL_CONTRACT.md`](TOOL_CONTRACT.md#semanticaction) for the full field rules). For `locator: "role"`, pass either `value: "button"` or `role: "button"`; if both are present they must match. It is a top-level alternative to `args`, `job`, `qa`, `sourceLookup`, `networkSourceLookup`, and `electron`, not a nested shape inside `batch` stdin arrays. Add `session` inside `semanticAction` when the shorthand should target a named upstream browser session; the compiled argv prepends `--session <name>` before `find`, direct selector/ref commands, or `select`, and fallback candidate actions preserve that prefix. For active sessions, role/name click/check/fill shorthands may resolve through the current `snapshot -i` refs before execution so hidden duplicate matches do not steal the action; fill only resolves when there is one exact editable current ref match. Inspect `details.effectiveArgs` when you need the exact executed argv. `semanticAction` does not expose `uncheck` because upstream `find` actions are only `click, fill, check, hover, text`; use raw `uncheck <selector-or-ref>` after choosing a stable selector or current snapshot ref. `select` shorthand intentionally requires a stable selector or current `@ref` plus `value`/`values`; upstream `find` does not expose a verified `select` action, so role/name/label dropdown resolution stays a snapshot/selector decision instead of hidden wrapper magic. If a raw `find` or semantic action misses with `selector-not-found`, the wrapper may take one fresh snapshot and append `Current snapshot ref fallback` when that snapshot has exact visible role/name matches for the failed target. Non-fill matches can include direct `try-current-visible-ref*` next actions. Semantic click misses may also include `Agent-browser candidate fallbacks`; `details.nextActions` first recommends a fresh `snapshot -i` and may include bounded role/name retries such as `button`/`link` for a missed `text` click, each as a `try-*-candidate` entry carrying redacted `find role …` argv.
246
266
 
247
267
  For desktop, contenteditable, or host-controlled rich inputs, treat a semantic `fill` miss or mismatch differently. Active-session role/name fills can execute through one exact current editable `combobox`, `searchbox`, or `textbox` ref before upstream `find` runs. If a later selector miss still finds an exact current editable ref (`searchbox` or `textbox`), `details.richInputRecovery` and visible `Rich input recovery` describe the candidate and append `focus-current-editable-ref*` / `click-current-editable-ref*` next actions. Those actions deliberately do **not** copy the fill text and never press `Enter` or submit. Direct `fill @ref <text>` on contenteditable refs may also append/prepend instead of replacing; when the latest snapshot proves the target is contenteditable, the wrapper verifies `get text` after a successful fill and appends `details.fillVerification` plus `inspect-after-fill-verification` / `verify-filled-value` if the visible text does not match. Use the safe ladder instead: refresh refs, choose the current editable `@ref`, focus or click it, then send the intended text with `keyboard inserttext` or `keyboard type` in a separate call. Do not auto-submit unless the user flow explicitly calls for it.
248
268
 
@@ -287,7 +307,7 @@ On tabbed or hidden-DOM pages, `get text <selector>` reads the upstream-selected
287
307
 
288
308
  Use `batch --bail` when later steps should stop after the first failed command.
289
309
 
290
- For short constrained flows, use top-level `job` instead of hand-writing `batch` stdin. Supported job steps are `open`, `click`, `fill`, `type`, `select`, `wait`, `assertText`, `assertUrl`, `waitForDownload`, `snapshot`, and `screenshot`. `open` can include `loadState: "domcontentloaded" | "load" | "networkidle"` to insert a `wait --load …` row immediately after navigation before the next click/read step. `click` and `fill` accept either a stable `selector` or the same semantic locator fields as top-level `semanticAction` (`locator`, plus `role`/`name` or `value` as appropriate) and compile locator steps to upstream `find` argv. `type` focuses an optional selector, sends text through upstream keyboard typing, can insert `wait` rows via `delayMs` for human-paced input, and can append a final `press` key such as `Enter`; delayed typing is capped at 200 characters per step, and generated per-character rows are compacted in model-visible batch text while remaining available in `details.batchSteps`. `select` requires `selector` plus `value` or `values`, and compiles to upstream `select <selector> <value...>`. By default the wrapper compiles steps to upstream `batch --bail` so a failed setup/fill/assertion step stops later mutating clicks; set `failFast: false` only when you explicitly need continue-after-error diagnostics. The wrapper records `details.compiledJob.steps[]` plus `details.compiledJob.failFast`. There is still no separate first-class catalog of reusable named browser recipes above `job`, the `qa` preset, and raw `batch`; see [`ARCHITECTURE.md`](ARCHITECTURE.md#no-reusable-recipe-layer-yet) for the closed `RQ-0068` decision and revisit bar.
310
+ For short constrained flows, use top-level `job` instead of hand-writing `batch` stdin. Supported job steps are `open`, `click`, `fill`, `type`, `select`, `wait`, `assertText`, `assertUrl`, `waitForDownload`, `snapshot`, and `screenshot`. `open` can include `loadState: "domcontentloaded" | "load" | "networkidle"` to insert a `wait --load …` row immediately after navigation before the next click/read step. `click` and `fill` accept either a stable `selector` or the same semantic locator fields as top-level `semanticAction` (`locator`, plus `role`/`name` or `value` as appropriate) and compile locator steps to upstream `find` argv. `type` focuses an optional selector, sends text through upstream keyboard typing, can insert `wait` rows via `delayMs` for human-paced input, and can append a final `press` key such as `Enter`; delayed typing is capped at 200 characters per step, and generated per-character rows are compacted in model-visible batch text while remaining available in `details.batchSteps`. `select` requires `selector` plus `value` or `values`, and compiles to upstream `select <selector> <value...>`. By default the wrapper compiles steps to upstream `batch --bail` so a failed setup/fill/assertion step stops later mutating clicks; set `failFast: false` only when you explicitly need continue-after-error diagnostics and those later steps remain safe if an earlier navigation fails; otherwise keep fail-fast or split navigation from content. The wrapper records `details.compiledJob.steps[]` plus `details.compiledJob.failFast`. There is still no separate first-class catalog of reusable named browser recipes above `job`, the `qa` preset, and raw `batch`; see [`ARCHITECTURE.md`](ARCHITECTURE.md#no-reusable-recipe-layer-yet) for the closed `RQ-0068` decision and revisit bar.
291
311
 
292
312
  **Job navigation is explicit.** A `click` step (or other navigation-prone interaction) does not prove the next page loaded. The wrapper does not auto-insert `assertUrl` or `assertText` after clicks inside `job`; add those steps yourself with the exact URL, a `*` / `**` glob-style URL pattern, or on-page text you expect, especially after forms, checkout, tabs, or submit buttons, before screenshots or later steps. Exact and glob-style `assertUrl` values compile to `wait --url` unchanged, including query strings and literal `?`; upstream `agent-browser 0.31.1` matches `*` / `**` patterns against the full active URL. Do not put a whole dynamic checkout into one long job: split around login, sorting/cart mutations, checkout navigation, and final evidence capture so refs and app state can be rechecked between phases.
293
313
 
@@ -374,7 +394,7 @@ Typical lifecycle:
374
394
  { "electron": { "action": "cleanup", "launchId": "electron-…" } }
375
395
  ```
376
396
 
377
- `electron.status` and `electron.cleanup` take either `launchId`, **`all: true`** (literal boolean) to walk every wrapper-tracked launch in one call, or neither when exactly one active launch exists—never both `launchId` and `all`. They can target the current branch-visible launch plus still-owned off-branch launch records by `launchId`; default no-arg calls are intentionally ambiguous when more than one active launch is owned. `/reload` preserves the current branch-visible active Electron launch and its isolated temp `userDataDir` for continuity, and cleans off-branch owned Electron launches; if cleanup is partial and skips or fails profile removal, the generic temp sweep preserves that `userDataDir` across reload, quit, later temp cleanup, process exit, and stale temp-root pruning after restart. For `electron.launch`, `timeoutMs` bounds host CDP readiness with a **15s** default and **120s** cap in `extensions/agent-browser/lib/electron/launch.ts`. Optional `timeoutMs` on **`status`** applies to managed-session `get title` / `get url` reads (localhost CDP probes stay on a short fixed fetch budget). On **`cleanup`**, it caps upstream `close` **and** host teardown (process exit, debug-port idle check, isolated profile removal); when omitted it follows the implicit session close default (**5s** unless `PI_AGENT_BROWSER_IMPLICIT_SESSION_CLOSE_TIMEOUT_MS` overrides). A successful managed-session close step retires that wrapper-managed session even when host process/profile cleanup remains partial. On **`probe`**, it bounds each underlying upstream read subprocess—omit it to use the normal tool subprocess default, or raise it on slow desktops.
397
+ `electron.status` and `electron.cleanup` take either `launchId`, **`all: true`** (literal boolean) to walk every wrapper-tracked launch in one call, or neither when exactly one active launch exists—never both `launchId` and `all`. They can target the current branch-visible launch plus still-owned off-branch launch records by `launchId`; default no-arg calls are intentionally ambiguous when more than one active launch is owned. `/reload` preserves the current branch-visible active Electron launch and its isolated temp `userDataDir` for continuity, and cleans off-branch owned Electron launches; if cleanup is partial and skips or fails profile removal, the generic temp sweep preserves that `userDataDir` across reload, quit, later temp cleanup, process exit, and stale temp-root pruning after restart. For `electron.launch`, `timeoutMs` bounds host CDP readiness with a **15s** default and **120s** cap in `extensions/agent-browser/lib/electron/launch.ts`. Optional `timeoutMs` on **`status`** applies to managed-session `get url`, then `get title` reads (localhost CDP probes stay on a short fixed fetch budget). On **`cleanup`**, it caps upstream `close` **and** host teardown (process exit, debug-port idle check, isolated profile removal); when omitted it follows the implicit session close default (**5s** unless `PI_AGENT_BROWSER_IMPLICIT_SESSION_CLOSE_TIMEOUT_MS` overrides). A successful managed-session close step retires that wrapper-managed session even when host process/profile cleanup remains partial. On **`probe`**, it bounds each underlying upstream read subprocess—omit it to use the normal tool subprocess default, or raise it on slow desktops.
378
398
 
379
399
  `launch.handoff` defaults to `"snapshot"`, which attaches through upstream `connect`, lists targets, and captures a current `snapshot -i` in one call. Snapshot handoff retries briefly when the first Electron snapshot has no refs; if it still reports no refs, run `snapshot -i` once more before assuming the app is blank. Use `handoff: "tabs"` as the safer diagnostic starting point when you only need target discovery and do not want to snapshot app content yet, or `handoff: "connect"` when you want to attach first and run your own follow-up commands. `targetType` defaults to `"page"`; use `"webview"` or `"any"` for apps that expose useful webviews. When a matching CDP target exposes a WebSocket URL, launch connects to that target; otherwise it falls back to the browser port.
380
400
 
@@ -389,9 +409,9 @@ Manual path for externally launched apps: if you started the Electron app yourse
389
409
  { "args": ["snapshot", "-i"] }
390
410
  ```
391
411
 
392
- A successful raw `connect` means the debug endpoint accepted the session, not that the app has an active ready page. Prefer `details.nextActions` when present: `list-connected-session-tabs` runs the session-scoped tab inspection. After that read-only list, select or confirm the stable `t<N>` target and run `snapshot -i` explicitly before trusting refs. If a `snapshot -i` says `No active page`, the wrapper clears any prior refs for that session; follow `list-tabs-after-no-active-page`, select the stable `t<N>` surface, then use a condition wait or retry `snapshot -i` before trusting refs.
412
+ A successful raw `connect` means the debug endpoint accepted the session, not that the app has an active ready page. Prefer `details.nextActions` when present: `verify-connected-session-url` performs the only page read allowed while the attached target is unverified, and `list-connected-session-tabs` runs session-scoped tab inspection. A verified HTTP(S)/app target clears the guard; otherwise navigate explicitly to a safe URL. After the read-only tab list, select or confirm the stable `t<N>` target, verify it with `get url`, and run `snapshot -i` explicitly before trusting refs. If a `snapshot -i` says `No active page`, the wrapper clears any prior refs for that session; follow `list-tabs-after-no-active-page`, select the stable `t<N>` surface, then use a condition wait or retry `snapshot -i` before trusting refs.
393
413
 
394
- For current-session smoke checks after either path, use `qa.attached`; for compact state instead of separate title/url/focus/tab/snapshot calls, use `electron.probe`. `electron.probe.timeoutMs` bounds each underlying read subprocess; `electron.probe.launchId` ties the probe to a wrapper launch and can surface session or target mismatch guidance before you trust page refs. For VS Code-style quick inputs, treat a successful `fill` as tentative: the wrapper may append `details.fillVerification` if `get value` still reads empty or different, and Electron `@e…` mutations can append `refresh-electron-refs-after-rerender` because same-URL UI rerenders commonly churn refs.
414
+ For current-session smoke checks after either path, use `qa.attached`; for compact state instead of separate title/url/focus/tab/snapshot calls, use `electron.probe`. `electron.probe.timeoutMs` bounds each underlying read subprocess; `electron.probe.launchId` ties the probe to a wrapper launch and can surface session or target mismatch guidance before you trust page refs. Electron status target reads and probe reads use the same daemon-policy lock and owned restore decision as ordinary managed commands. A probe reads and validates the live URL before title, focus, tab, or snapshot helpers. Electron `launch` snapshot/tabs handoff likewise validates the URL before tab/snapshot reads; handoff failure or cancellation closes the new managed session and host process/profile. Current-managed probe results persist their top-level namespace, tab target, and ref snapshot so Pi reload/branch replay restores the same page identity. For VS Code-style quick inputs, treat a successful `fill` as tentative: the wrapper may append `details.fillVerification` if `get value` still reads empty or different, and Electron `@e…` mutations can append `refresh-electron-refs-after-rerender` because same-URL UI rerenders commonly churn refs.
395
415
 
396
416
  For local app debugging, top-level `sourceLookup` can gather candidate component/file locations for a visible element from selector DOM hints, React DevTools inspection, and a bounded workspace component-name search rooted at the Pi session working directory (`maxWorkspaceFiles` defaults to 2000 and cannot exceed 5000; the scan records at most ten `workspace-search` candidates). With a `selector`, the wrapper runs `is visible` and, unless `includeDomHints` is `false`, `get html` so DOM data attributes and embedded source-like paths can become `dom-attribute` candidates. It reports evidence and confidence in `details.sourceLookup` instead of claiming a guaranteed source file. React hints require a session opened with `--enable react-devtools`. The `details.sourceLookup.status` field reads `unsupported` only when no candidates were collected **and** a `react` batch step failed (inspect errors, missing renderer, and similar); it reads `no-candidates` when the batch succeeded but nothing matched. If selector or workspace hints still yield candidates, `status` remains `candidates-found` even when React inspection failed. Unlike `qa`, the wrapper does not downgrade a **fully successful** upstream batch to `isError` solely because those statuses appear—though failed batch steps still produce normal tool errors. For wrapper-tracked packaged Electron sessions with no candidates, `details.sourceLookup.workspaceRoot` and optional `details.sourceLookup.electronContext` explain that the scan only covered the Pi tool cwd; installed app resources or `app.asar` bundles are outside that scan and are not unpacked. Those results may add `snapshot-electron-session`, `probe-electron-launch`, and `list-electron-tabs` next actions so you can inspect the live packaged app before deciding whether to change the workspace or app bundle.
397
417
 
@@ -540,7 +560,7 @@ Operational notes:
540
560
  - Visible page content from real authenticated profiles is still model-visible and may persist in transcripts or saved artifacts. The wrapper redacts credential-like cookie/storage/auth data, not the ordinary page text you asked it to read.
541
561
  - `stdin` is accepted only for `batch`, `eval --stdin`, and `auth save --password-stdin`; other stdin-bearing calls are rejected before launch.
542
562
  - `auth list/show/save/login/delete` summaries avoid expanding profile secrets. Prefer `auth save --password-stdin` over `--password <value>`.
543
- - `session list` and `tab list` are formatted as compact field lists so generated names, labels, active markers, page titles, and URLs are visible without relying on raw JSON.
563
+ - `session list` and `tab list` are formatted as compact field lists so caller-owned names, labels, active markers, page titles, and URLs are visible without relying on raw JSON. Wrapper-managed `piab-*` session rows are removed from `session list`.
544
564
  - `state save <path>` is a verified file-artifact workflow; the wrapper creates missing parent directories before invoking upstream, then inspect `details.artifactVerification` before relying on the file. `state load <path>` is not treated as a newly saved artifact.
545
565
  - `cookies get` can expose real authenticated-profile cookies; prefer task-specific page actions and only inspect cookies when the user needs cookie data.
546
566
  - `storage local|session` summaries redact sensitive keys and likely secret values but may keep benign primitive local QA values visible, for example `theme: dark`; still avoid broad storage dumps unless necessary.
@@ -562,7 +582,7 @@ Session note: `skills list`, `skills get …`, and `skills path …` are **state
562
582
  | `skills list` | List available CLI-bundled skills. |
563
583
  | `skills get core` | Print the core usage guide. |
564
584
  | `skills get core --full` | Print the full version-matched core command reference and templates. |
565
- | `skills get <name>` | Load a specialized skill such as `electron` or `slack`. Common specialized calls include `skills get electron`, `skills get slack`, `skills get dogfood`, `skills get vercel-sandbox`, and `skills get agentcore`. |
585
+ | `skills get <name>` | Load a specialized skill such as `electron` or `slack`. Common specialized calls include `skills get electron`, `skills get slack`, `skills get dogfood`, `skills get vercel-sandbox`, `skills get agentcore`, and `skills get derive-client` (HAR-to-API-client workflow). |
566
586
  | `skills get <name> --full` | Include a skill's supplementary references/templates when present. |
567
587
  | `skills get --all` | Print all visible bundled skills for broad audit/debug work. |
568
588
  | `skills path [name]` | Print a skill directory path. |
@@ -631,7 +651,7 @@ Comboboxes vary by app. For native `<select>` controls, prefer raw `select <sele
631
651
  | `state save <path>` | Save cookies, local storage, and session storage to a state file. |
632
652
  | `state load <path>` | Load cookies and storage from a state file. |
633
653
  | `state list` | List saved state files. |
634
- | `state show <filename>` | Show saved-state metadata without dumping secrets. |
654
+ | `state show <filename>` | Show saved-state metadata without dumping cookie or storage values. |
635
655
  | `state rename <old-name> <new-name>` | Rename a saved state file. |
636
656
  | `state clear [session-name] [--all]` | Clear saved states for one name or all names; `state clear -a` is the upstream short alias for clearing all names. |
637
657
  | `session id --scope worktree --prefix <name>` | Generate a stable session id for agent/worktree-scoped browser state. |
@@ -663,10 +683,10 @@ These calls return plain text and stay stateless: the extension does not inject
663
683
  | `get text/html/value/count <selector>` | Read matched elements; use `get text body` for whole-page text. |
664
684
  | `get attr <selector> <name>`, `get box <selector>`, `get styles <selector>` | Read an attribute, bounding box, or computed styles from matched elements. |
665
685
  | `is <what> <selector>` | Check `visible`, `enabled`, or `checked`. |
666
- | `find <locator> <value> <action> [text]` | Locator types include `role`, `text`, `label`, `placeholder`, `alt`, `title`, and `testid`; selector helpers include `find first <sel>`, `find last <sel>`, and `find nth <n> <sel>`. Role/text filters include `find role <role> --name <name>` and `find ... --exact`. |
667
- | `mouse <action> [args]` | `move <x> <y>`, `down [btn]`, `up [btn]`, `wheel <dy> [dx]`. |
686
+ | `find <locator> <value> <action> [text]` | Locator types include `role`, `text`, `label`, `placeholder`, `alt`, `title`, and `testid`; selector helpers include `find first <sel>`, `find last <sel>`, and `find nth <n> <sel>`. Role/text filters include `find role <role> --name <name>` and `find ... --exact`. Actions are `click, fill, check, hover, text` only. Prefer `find role` for semantic elements: implicit roles work (`find role heading text --name` for `<h2>`, list/banner landmarks, and similar). Default name matching is a case-insensitive substring; `--exact` makes the accessible name case-sensitive. On misses, upstream 0.32.4+ keeps locator detail such as `Names seen: …` or `No element found: getByRole(...)` instead of a generic flatten. |
687
+ | `mouse <action> [args]` | `move <x> <y>`, `down [btn]`, `up [btn]`, `wheel <dy> [dx]`. Local directory `file:` pages reject interaction commands that could navigate through a ref or scripted event into protected `.agent-browser` state storage. |
668
688
  | `set <setting> [value]` | `viewport <w> <h>`, `device <name>`, `geo <lat> <lng>`, `offline [on|off]`, `headers <json>`, `credentials <user> <pass>`, and `set media <features>` (`dark`, `light`, and/or `reduced-motion`). |
669
- | `network <action>` | `network route <url> [--abort|--body <json>] [--resource-type <csv>]`, `network unroute [url]`, `network requests [--clear] [--filter <pattern>] [--type <csv>] [--method <method>] [--status <code|range>]`, `network request <requestId>`, `network har start`, and `network har stop [path]`. `--resource-type` filters intercepted requests by CDP resource type, such as `script`, `image`, `font`, `xhr`, or `fetch`; request listing filters accept resource types (`xhr,fetch`), methods (`POST`), and statuses (`2xx`, `400-499`). |
689
+ | `network <action>` | `network route <url> [--abort|--body <json>] [--resource-type <csv>]`, `network unroute [url]`, `network requests [--clear] [--filter <pattern>] [--type <csv>] [--method <method>] [--status <code|range>]`, `network request <requestId>`, `network har start`, `network har start --content text` (default; embeds text bodies), `network har start --content all`, `network har start --content none`, and `network har stop [path]`. `--resource-type` filters intercepted requests by CDP resource type, such as `script`, `image`, `font`, `xhr`, or `fetch`; request listing filters accept resource types (`xhr,fetch`), methods (`POST`), and statuses (`2xx`, `400-499`). HAR files can include auth headers and bodies—do not share them unredacted. For turning a recording into a reusable API client, load `skills get derive-client`. |
670
690
  | `cookies [get|set|clear]` | Manage cookies. Full set form: `cookies set <name> <value> --url <url> --domain <domain> --path <path> --httpOnly --secure --sameSite <Strict|Lax|None> --expires <timestamp>`; also supports `cookies set --curl <file>` for JSON, cURL, or bare Cookie-header bulk imports. |
671
691
  | `storage <local|session>` | Manage web storage. |
672
692
 
@@ -742,6 +762,7 @@ Current upstream still does not parse `wait <selector> --state hidden` / `wait <
742
762
  | `react renders start` | Start recording React render activity. |
743
763
  | `react renders stop [--json]` | Stop render recording and print mount/re-render counts and changed details. |
744
764
  | `react suspense [--only-dynamic] [--json]` | Classify Suspense boundaries with grouped root-cause recommendations. |
765
+ | `a11y [url]` | Run an embedded axe-core accessibility audit on the current page, or navigate to `url` first. Options: `a11y --tags wcag2a,wcag2aa`, `a11y --selector "#main"`. CDP browsers only (not Safari/iOS WebDriver). Model-facing text summarizes violation/incomplete counts and top rules; full node targets stay in `details.data`. |
745
766
  | `vitals [url] [--json]` | Report Core Web Vitals: LCP, CLS, TTFB, FCP, INP, plus React hydration timing when available. `web-vitals [url] [--json]` is the upstream alias. |
746
767
  | `pushstate <url>` | Perform SPA client-side navigation; detects Next.js router pushes and falls back to history navigation events. |
747
768
  | `removeinitscript <id>` | Remove an init script registered through upstream init-script mechanisms. |
@@ -785,7 +806,7 @@ Long-running or lifecycle commands should be explicitly paired with cleanup call
785
806
  | `mcp` | Start a local MCP stdio server for external MCP clients; bare native-tool calls are rejected before spawn. |
786
807
  | `profiles` | List available Chrome profiles. |
787
808
 
788
- When these commands are invoked through the native `agent_browser` tool, structured diagnostic/status outputs are rendered as compact summaries. Local inspection/setup calls (`auth save/list/show/delete/remove`, `dashboard start/stop`, `device list`, `doctor`, `install`, `upgrade`, `profiles`, `session id`, `session info`, `session list`, `plugin add/list/show/run`, `state list/show/rename`, `state clean --older-than <days>`, `state clear --all`, `state clear -a`, and `state clear <session-name>`) are sessionless unless you explicitly pass `--session`; bare `mcp` server calls are blocked except help. Context-dependent calls such as root `session`, untargeted `state clear`, `auth login`, `chat`, and `state save/load` keep normal session behavior. List-like outputs such as sessions, Chrome profiles, auth profiles, network requests, console messages, and page errors include counts and key fields; large outputs are previewed with a `Full output path:` spill file instead of dumping the entire payload into context. For `network requests`, the wrapper shows a failed-request summary split into actionable versus benign low-impact rows, then status, method, URL, resource/mime type, request id, and, when the installed upstream output includes body-like fields, bounded redacted payload, response, and failure/error snippets. Safe request IDs also produce `details.nextActions` for exact request details, actionable failed-request source lookup candidates, filtered request lists, or starting HAR capture before a repro. If the same session has active wrapper-observed network routes, failed/pending/CORS-looking matched request rows add `details.networkRouteDiagnostics` and executable route-mock next actions before the generic request actions. `data:image` artifact rows are omitted from compact request previews but remain in raw `details.data.requests`. `network request <requestId>` can expose upstream full-detail body fields such as response bodies using the same bounded model-facing preview; its request URL stays diagnostic-only and does not overwrite `details.sessionTabTarget` for later ref guards. Clipboard failures that mention `NotAllowedError` or permission denial are usually browser/OS capability limits, not proof that a read, paste, or page mutation happened; prefer page-native reads (`snapshot -i`, `get text`, `eval --stdin`) or direct typing (`keyboard inserttext` / `keyboard type`) when the workflow allows it, and retry true clipboard flows only from an allowed profile/session on a normal `http(s)` page. Header, cookie, auth, token, and other secret-like fields are not expanded in model-facing text or `details.data`; low-risk primitive storage values may remain visible, while command echoes still redact `--body`, `--headers`, `--password`, proxy credentials, auth-bearing URLs, `clipboard write` text, cookie/storage set values, and bearer/basic credential text in positional arguments. Use upstream HAR or full raw details only when complete data is required.
809
+ When these commands are invoked through the native `agent_browser` tool, structured diagnostic/status outputs are rendered as compact summaries. As a checkout-auth isolation boundary, `session list` omits wrapper-managed `piab-*` live-session rows, `state list` omits wrapper-managed `piab-r2-*` rows and legacy `piab-r-*` rows, and explicit `piab-*` session targets are rejected unless this extension instance owns that exact namespace/session; foreign managed `--restore`, `--state`, `state show`, and `state load` references fail before spawn. Broad `state clear` / `state clean` and managed save/rename targets are also blocked; targeted caller-owned state names and paths remain available. Local inspection/setup calls (`auth save/list/show/delete/remove`, `dashboard start/stop`, `device list`, `doctor`, `install`, `upgrade`, `profiles`, `session id`, `session info`, `session list`, `plugin add/list/show/run`, `state list/show/rename`, and targeted `state clear <caller-owned-name>`) are sessionless unless you explicitly pass `--session`; bare `mcp` server calls are blocked except help. Context-dependent calls such as root `session`, untargeted `state clear`, `auth login`, `chat`, and `state save/load` keep normal session behavior. List-like outputs such as sessions, Chrome profiles, auth profiles, network requests, console messages, and page errors include counts and key fields; large outputs are previewed with a `Full output path:` spill file instead of dumping the entire payload into context. For `network requests`, the wrapper shows a failed-request summary split into actionable versus benign low-impact rows, then status, method, URL, resource/mime type, request id, and, when the installed upstream output includes body-like fields, bounded redacted payload, response, and failure/error snippets. Safe request IDs also produce `details.nextActions` for exact request details, actionable failed-request source lookup candidates, filtered request lists, or starting HAR capture before a repro. If the same session has active wrapper-observed network routes, failed/pending/CORS-looking matched request rows add `details.networkRouteDiagnostics` and executable route-mock next actions before the generic request actions. `data:image` artifact rows are omitted from compact request previews but remain in raw `details.data.requests`. `network request <requestId>` can expose upstream full-detail body fields such as response bodies using the same bounded model-facing preview; its request URL stays diagnostic-only and does not overwrite `details.sessionTabTarget` for later ref guards. Clipboard failures that mention `NotAllowedError` or permission denial are usually browser/OS capability limits, not proof that a read, paste, or page mutation happened; prefer page-native reads (`snapshot -i`, `get text`, `eval --stdin`) or direct typing (`keyboard inserttext` / `keyboard type`) when the workflow allows it, and retry true clipboard flows only from an allowed profile/session on a normal `http(s)` page. Header, cookie, auth, token, and other secret-like fields are not expanded in model-facing text or `details.data`; low-risk primitive storage values may remain visible, while command echoes still redact `--body`, `--headers`, `--password`, proxy credentials, auth-bearing URLs, `clipboard write` text, cookie/storage set values, and bearer/basic credential text in positional arguments. Use upstream HAR or full raw details only when complete data is required.
789
810
 
790
811
  ## Optional package config and companion web search
791
812
 
@@ -862,13 +883,13 @@ Browser default config is conservative: it adds agent guidance for signed-in/acc
862
883
 
863
884
  - `--profile <name|path>`: reuse Chrome profile login state by directory name from `profiles`, or use a persistent custom profile/profile-directory path when upstream accepts it. Environment: `AGENT_BROWSER_PROFILE`.
864
885
  - `--session <name>`: use an isolated session. Environment: `AGENT_BROWSER_SESSION`.
865
- - `--restore [name]`: auto-save/restore cookies and local storage; bare `--restore` uses `--session` as the key. Environment: `AGENT_BROWSER_RESTORE`.
886
+ - `--restore [name]`: auto-save/restore cookies, local storage, and session storage; bare `--restore` uses `--session` as the key. Environment: `AGENT_BROWSER_RESTORE`. Wrapper-owned managed sessions set a Git-checkout-generation-stable restore key automatically unless disabled with `PI_AGENT_BROWSER_MANAGED_SESSION_RESTORE=0` or sticky-disabled after an incompatible launch; a successfully spawned suppressed identity reports `details.managedSessionRestoreDisabled`, meaning later plain follow-ups do not inject the wrapper key. Same-policy follow-ups may keep using that live daemon; any later call that requests incompatible policy first inspects the daemon and blocks if it retains a restore key or cannot be inspected. Any upstream config discovered while planning or browser mutation flag/env disables automatic managed restore without reading caller-selected config content. Subprocesses that receive the wrapper restore key and wrapper-owned closes pin a process-private empty `AGENT_BROWSER_CONFIG` (`0400` on POSIX) in the marked secure-temp lifecycle, so config created after planning cannot alter the browser that receives restored auth. Raw batch argv and batch stdin containing nested `connect`/`batch` also disable restore so an attached browser cannot receive wrapper-managed auth state. A user-private immutable ticket-claim lock serializes daemon policy inspection through the receiving spawn and bridges the pre-update v2 lock path; every lock winner re-inspects the live daemon, and abandoned v2 locks fail closed for manual repair. POSIX process identity probes use absolute `/bin/ps` then `/usr/bin/ps` paths. Before incompatible reuse the wrapper inspects the actual daemon and blocks if it retains any restore key, cannot be inspected, or reports restore-disabled policy without current-process provenance. Same-process `session_tree` changes keep that provenance, while reload/restart/resume intentionally do not restore it from transcript rows; close a still-live blocked daemon first, use a fresh wrapper session, or choose a distinct explicit session. Wrapper-owned subprocesses pin their canonical namespace, including the empty default namespace. Upstream restore files under `~/.agent-browser/` are plaintext unless a valid 64-character hex `AGENT_BROWSER_ENCRYPTION_KEY` is set; automatic restore requires a durable Git checkout generation plus an absolute home root, and on POSIX the wrapper canonicalizes and pins `HOME`, requires trusted non-writable owner ancestry plus stable device/inode/birth-time metadata for the checkout and Git-admin directories, enforces mode `0700` without silently repairing unsafe existing directories, and rejects symlinks/non-directories along the exact restore `sessions` path and `.tmp` write area, while Windows requires that key because POSIX mode checks cannot verify profile ACLs. Wrapper-owned close discards caller config/restore globals and preserves the live daemon's existing restore key rather than injecting one derived from a replacement checkout and records a returned old-generation snapshot against that observed wrapper key. A failed fresh command is followed by an exact-identity daemon probe; live or uninspectable starts remain owned for shutdown cleanup. After a wrapper-owned close succeeds, the wrapper persists its returned state path as an atomic record in a lockless convergent per-key ownership directory (`0700`, with `0600` records, on POSIX), retains the two newest proven snapshots for the exact restore key across Pi restarts, self-heals malformed regular records, removes additional proven snapshots older than 30 days, expires stale ownership-proven snapshots and empty manifests from other restore-key generations after 30 days only when a private lineage record proves the same canonical checkout path, and caps young close churn at 256 records per restore key without invoking namespace-wide `state clean`; unrecorded matching files and the current checkout key are untouched. Restore capabilities and key-bearing paths are redacted from output/transcripts; malformed oversized upstream output is discarded instead of persisted as a parse-failure spill. Checkout/storage/managed-session/state-access policy is revalidated after async setup immediately before spawn. Native Windows command-first adaptation moves only valid leading globals, rewrites valued `--restore <name>` as `--restore=<name>`, consumes only exact lowercase boolean literals, and leaves invalid or command-scoped leading input unchanged.
866
887
  - `--restore-save <policy>` (`auto`, `always`, or `never`): restore auto-save policy. Environment: `AGENT_BROWSER_RESTORE_SAVE`. Restore-enabled sessions also save periodically while open; `AGENT_BROWSER_AUTOSAVE_INTERVAL_MS` sets the minimum interval in milliseconds (`30000` by default, `0` disables periodic saves but not save-on-close).
867
888
  - `--restore-check-url <glob>`, `--restore-check-text <txt>`, `--restore-check-fn <js>`: validate restored state before auto-save. Environments: `AGENT_BROWSER_RESTORE_CHECK_URL`, `AGENT_BROWSER_RESTORE_CHECK_TEXT`, `AGENT_BROWSER_RESTORE_CHECK_FN`.
868
- - `--namespace <name>`: isolate daemon sockets and restore-state directories. Environment: `AGENT_BROWSER_NAMESPACE`.
889
+ - `--namespace <name>`: isolate daemon sockets and restore-state directories. Environment: `AGENT_BROWSER_NAMESPACE`. Upstream and the wrapper canonicalize namespace identity to a lowercase sanitized component (for example, `Team Name` becomes `team-name`).
869
890
  - `--session-name <name>`: legacy alias for restore persistence key. Environment: `AGENT_BROWSER_SESSION_NAME`.
870
891
  - `--state <path>`: load saved auth state from JSON. Environment: `AGENT_BROWSER_STATE`.
871
- - `--auto-connect`: connect to a running Chrome to reuse auth state. Environment: `AGENT_BROWSER_AUTO_CONNECT`.
892
+ - `--auto-connect`: connect to a running Chrome to reuse auth state. Environment: `AGENT_BROWSER_AUTO_CONNECT`. Optional booleans use separated tokens (`--auto-connect false`); upstream 0.33.2 does not recognize `--auto-connect=false`, so that token cannot disable an earlier bare `--auto-connect`.
872
893
  - `--headers <json>`: apply HTTP headers scoped to the opened URL's origin.
873
894
  - `--init-script <path>`: register a script before first navigation; repeatable. Environment: `AGENT_BROWSER_INIT_SCRIPTS`.
874
895
  - `--enable <feature>`: enable built-in init scripts such as `react-devtools`; repeatable or comma-separated. Environment: `AGENT_BROWSER_ENABLE`.
@@ -882,7 +903,8 @@ Browser default config is conservative: it adds agent guidance for signed-in/acc
882
903
  - `--proxy <server>`: proxy server URL. Environments: `AGENT_BROWSER_PROXY`, `HTTP_PROXY`, `HTTPS_PROXY`, `ALL_PROXY`.
883
904
  - `--proxy-bypass <hosts>`: proxy bypass hosts. Environments: `AGENT_BROWSER_PROXY_BYPASS`, `NO_PROXY`.
884
905
  - `--ignore-https-errors`: ignore HTTPS certificate errors. Environment: `AGENT_BROWSER_IGNORE_HTTPS_ERRORS`.
885
- - `--allow-file-access`: allow `file://` URLs to access local files. Environment: `AGENT_BROWSER_ALLOW_FILE_ACCESS`.
906
+ - `--allow-file-access`: upstream capability, but enabled argv/`AGENT_BROWSER_ALLOW_FILE_ACCESS` forms and file-access-enabling `--args` / `AGENT_BROWSER_ARGS` Chrome switches are rejected by this native wrapper. Every spawn adds canonical `--args "" --allow-file-access false` defaults so project/user config cannot re-enable file access; an explicit validated safe CLI `--args` value remains usable. Every spawn removes all caller occurrences before adding one canonical separated `--allow-file-access false`, so unsupported equals forms cannot preserve an earlier enabled flag and config cannot re-enable local filesystem access. Unknown top-level or batch tab/attachment/script/state-load transitions remain blocked for page inspection until `get url` or explicit safe navigation establishes the target. `tab list` and non-content `tab <id>` selection remain available while unknown, but selection stays unverified until `get url`; post-transition summaries (including after arbitrary `eval`) read the live URL before title and stop if the target is a local file page.
907
+ - `--hide-scrollbars <bool>`: explicitly show or hide native scrollbars in headless Chromium screenshots.
886
908
  - `--headed`: ask upstream to show the browser window. Environment: `AGENT_BROWSER_HEADED`. Use it on the first launch, normally with `sessionMode: "fresh"` when changing an existing managed session; verify visibility with screenshot/tab evidence because the wrapper cannot yet prove the OS window is visible to the user.
887
909
  - `--webgpu`: enable upstream's platform-specific WebGPU launch preset. Environment: `AGENT_BROWSER_WEBGPU`; config: `"webgpu": true`. Use it on a fresh local launch. It is incompatible while enabled with `--cdp`, `--auto-connect`, and provider launches. `AGENT_BROWSER_NO_XVFB=1` disables upstream's automatic Xvfb for displayless headed Linux sessions.
888
910
  - `--cdp <port>`: connect through Chrome DevTools Protocol.
@@ -917,20 +939,21 @@ Browser default config is conservative: it adds agent guidance for signed-in/acc
917
939
 
918
940
  ### Config precedence
919
941
 
920
- `agent-browser` looks for `agent-browser.json` in these locations, from lowest to highest priority:
942
+ Standalone `agent-browser` looks for `agent-browser.json` in these locations, from lowest to highest priority:
921
943
 
922
944
  1. `~/.agent-browser/config.json` for user defaults.
923
945
  2. `./agent-browser.json` for project overrides.
924
946
  3. Environment variables, including `AGENT_BROWSER_CONFIG`.
925
947
  4. CLI flags.
926
948
 
927
- Use `--config <path>` to load a specific config file. Boolean flags accept optional `true` or `false` values, such as `--headed false` or `--webgpu false`, to override config. Browser extensions from user and project configs are merged rather than replaced.
949
+ Use separated `--config <path>` to load a specific config file in standalone upstream; upstream 0.33.2 does not recognize `--config=<path>` as the global config selector. The native wrapper rejects discovered, environment-selected, or explicit upstream config for browser-backed calls without reading it, then pins a process-private empty config for every accepted browser-backed spawn so a file created after planning cannot change the browser. Sessionless local/setup commands keep upstream config behavior. This policy is separate from the Pi-scoped package config under `.pi/config/pi-agent-browser-native/`; pass safe browser settings through native `args`/environment or that package's advisory browser guidance. Boolean flags accept optional `true` or `false` values, such as `--headed false` or `--webgpu false`, to override config. Browser extensions from user and project configs are merged rather than replaced.
928
950
 
929
- Other useful environment variables include `AGENT_BROWSER_DEFAULT_TIMEOUT`, `AGENT_BROWSER_AUTOSAVE_INTERVAL_MS`, `AGENT_BROWSER_STREAM_PORT`, `AGENT_BROWSER_IDLE_TIMEOUT_MS`, `AGENT_BROWSER_ENCRYPTION_KEY`, `AGENT_BROWSER_STATE_EXPIRE_DAYS`, `AGENT_BROWSER_IOS_DEVICE`, `AGENT_BROWSER_IOS_UDID`, `AI_GATEWAY_URL`, `AI_GATEWAY_API_KEY`, provider credential names, and AWS credential names when using AgentCore. The upstream child receives the parent environment plus wrapper overrides such as the managed socket directory and clamped default operation timeout (`buildAgentBrowserProcessEnv` in `extensions/agent-browser/lib/process.ts`). Model-facing output still redacts recognized secret values.
951
+ Other useful environment variables include `AGENT_BROWSER_DEFAULT_TIMEOUT`, `AGENT_BROWSER_AUTOSAVE_INTERVAL_MS`, `AGENT_BROWSER_STREAM_PORT`, `AGENT_BROWSER_STREAM_QUALITY`, `AGENT_BROWSER_STREAM_MAX_WIDTH`, `AGENT_BROWSER_STREAM_MAX_HEIGHT`, `AGENT_BROWSER_IDLE_TIMEOUT_MS`, `AGENT_BROWSER_ENCRYPTION_KEY`, `AGENT_BROWSER_STATE_EXPIRE_DAYS`, `AGENT_BROWSER_IOS_DEVICE`, `AGENT_BROWSER_IOS_UDID`, `AI_GATEWAY_URL`, `AI_GATEWAY_API_KEY`, provider credential names, and AWS credential names when using AgentCore. The upstream child receives the parent environment plus wrapper overrides such as the managed socket directory, clamped default operation timeout, canonical owned-session namespace (including empty default), and Git-checkout-generation-stable `AGENT_BROWSER_RESTORE` for wrapper-owned managed sessions (`buildAgentBrowserProcessEnv` in `extensions/agent-browser/lib/process.ts`, ownership carried by the wrapper's typed process options and call-scoped managed-session context). Model-facing output still redacts recognized secret values.
930
952
 
931
953
  ## Wrapper-specific behavior worth knowing
932
954
 
933
955
  - The extension may keep following one implicit managed session across later tool calls.
956
+ - Protected `.agent-browser` paths are rejected equally in CLI operands (including dash-prefixed values), raw Chrome args, and path-bearing environment mirrors, including `AGENT_BROWSER_STATE`, `AGENT_BROWSER_PROFILE`, `AGENT_BROWSER_CONFIG`, `AGENT_BROWSER_EXECUTABLE_PATH`, `AGENT_BROWSER_EXTENSIONS`, `AGENT_BROWSER_INIT_SCRIPTS`, `AGENT_BROWSER_ACTION_POLICY`, download/screenshot directories, `AGENT_BROWSER_SKILLS_DIR`, and the wrapper socket directory.
934
957
  - If launch-scoped flags like `--profile`, `--executable-path`, `--webgpu`, `--restore`, `--restore-save`, restore check flags, `--namespace`, `--session-name`, `--cdp`, `--state`, `--auto-connect`, `--init-script`, `--enable`, `--provider` / `-p`, or provider device flags like `--device` would be ignored because that implicit session is already active, retry with `sessionMode: "fresh"`.
935
958
  - If a `sessionMode: "fresh"` call fails (including upstream failure, timeout, missing binary, or **`qa`** reclassification after a nominally successful batch), read `details.managedSessionOutcome` before assuming where the next default call will go: `preserved` means the prior managed session remains current, while `abandoned` means no managed session became current. When the failure reason is not the fresh launch itself—for example `failureCategory: "qa-failure"`—`status`/`summary` may still describe the managed-session transition while `succeeded` on this object matches the final tool outcome.
936
959
  <!-- agent-browser-playbook:start wrapper-tab-recovery -->
@@ -940,7 +963,7 @@ Other useful environment variables include `AGENT_BROWSER_DEFAULT_TIMEOUT`, `AGE
940
963
  - For sessions with observed tab-drift risk, after a successful command on a known target tab, agent_browser also best-effort restores that intended tab if a restored/background tab steals focus after the command completes. Routine same-session commands skip this post-command tab-list probe.
941
964
  - If a known session target unexpectedly reports about:blank, agent_browser best-effort re-selects the prior intended target when it still exists; if recovery fails, it records the observed about:blank target and reports exact recovery guidance instead of treating the prior page as active.
942
965
  <!-- agent-browser-playbook:end wrapper-tab-recovery -->
943
- - Wrapper-spawned commands clamp `AGENT_BROWSER_DEFAULT_TIMEOUT` to the upstream documented 25-second default and use a 35-second child-process watchdog (`PI_AGENT_BROWSER_PROCESS_TIMEOUT_MS` overrides the default 35s budget; top-level `timeoutMs` overrides it per browser CLI call). Explicit `wait <ms>` or `wait --timeout <ms>` calls can exceed that default; when top-level `timeoutMs` is omitted, the wrapper derives a subprocess watchdog from the requested wait duration plus a small grace window. Dialog commands are additionally bounded to 5 seconds (`PI_AGENT_BROWSER_DIALOG_PROCESS_TIMEOUT_MS`), and click/tap/find refs or tokens plus `eval --stdin` snippets that look like alert/confirm/prompt/dialog triggers are bounded to 8 seconds (`PI_AGENT_BROWSER_DIALOG_TRIGGER_PROCESS_TIMEOUT_MS`). When any watchdog fires, `details.timeoutPartialProgress` may include a planned step list with per-step status (including `generatedFrom` labels for wrapper-inserted rows such as `open.loadState`) and a `retry-timeout-step` next action only when the first incomplete step is read-only or idempotent, or `inspect-current-page-after-timeout` when the session is still inspectable but the incomplete step may be mutating and should not be blindly retried. It also includes current page title/URL from best-effort session `get url` / `get title` (or a planned URL inferred from the step list when the session cannot answer), an `openedButPostOpenTimedOut` classification only when a live page URL was recovered before a later step hung, and declared artifact paths such as `screenshot`, `pdf`, `download`, or `wait --download` outputs with existence/state checks; the same evidence is appended under `Timeout partial progress` in visible text with URL/path redaction.
966
+ - Wrapper-spawned commands clamp `AGENT_BROWSER_DEFAULT_TIMEOUT` to the upstream documented 25-second default and use a 35-second child-process watchdog (`PI_AGENT_BROWSER_PROCESS_TIMEOUT_MS` overrides the default 35s budget; top-level `timeoutMs` overrides it per browser CLI call). Explicit `wait <ms>` or `wait --timeout <ms>` calls can exceed that default; when top-level `timeoutMs` is omitted, the wrapper derives a subprocess watchdog from the requested wait duration plus a small grace window. Dialog commands are additionally bounded to 5 seconds (`PI_AGENT_BROWSER_DIALOG_PROCESS_TIMEOUT_MS`), and click/tap/find refs or tokens plus `eval --stdin` snippets that look like alert/confirm/prompt/dialog triggers are bounded to 8 seconds (`PI_AGENT_BROWSER_DIALOG_TRIGGER_PROCESS_TIMEOUT_MS`). When any watchdog fires, `details.timeoutPartialProgress` may include a planned step list with per-step status (including `generatedFrom` labels for wrapper-inserted rows such as `open.loadState`) and a `retry-timeout-step` next action only when the first incomplete step is read-only or idempotent, or `inspect-current-page-after-timeout` when the session is still inspectable but the incomplete step may be mutating and should not be blindly retried. It also includes current page URL from best-effort session `get url`, followed by `get title` only for a verified non-file URL (or a planned URL inferred from the step list when the session cannot answer), an `openedButPostOpenTimedOut` classification only when a live page URL was recovered before a later step hung, and declared artifact paths such as `screenshot`, `pdf`, `download`, or `wait --download` outputs with existence/state checks; the same evidence is appended under `Timeout partial progress` in visible text with URL/path redaction.
944
967
  - Oversized snapshots and oversized generic outputs may be compacted in tool content, with the full raw output written to a spill file path shown directly in the tool result. Recent artifact metadata is bounded by `PI_AGENT_BROWSER_SESSION_ARTIFACT_MANIFEST_MAX_ENTRIES` (default 100); persisted spill files are separately bounded by `PI_AGENT_BROWSER_SESSION_ARTIFACT_MAX_BYTES` (default 32 MiB).
945
968
  - The wrapper keeps `--help` and `--version` stateless so they do not consume the implicit managed-session slot.
946
969
 
@@ -949,19 +972,23 @@ Other useful environment variables include `AGENT_BROWSER_DEFAULT_TIMEOUT`, `AGE
949
972
  <!-- agent-browser-capability-baseline:start capability-token-baseline -->
950
973
  <!-- Generated from scripts/agent-browser-capability-baseline.mjs. Run `npm run docs -- command-reference write` to update. Do not edit manually. -->
951
974
  <details>
952
- <summary>Generated verifier capability baseline for agent-browser 0.32.2</summary>
975
+ <summary>Generated verifier capability baseline for agent-browser 0.33.2</summary>
953
976
 
954
977
  This generated block is review data for maintainers. The human-authored reference sections above remain the readable command guide.
955
978
 
956
979
  #### Source evidence
957
980
  - repository: `vercel-labs/agent-browser`
958
- - upstream HEAD: `6ede7a9470ac4b681cabf838af8668b9aa99e957`
959
- - upstream package version: `0.32.2`
981
+ - upstream HEAD: `93cdda5709e8861c0c26b0b955d8d746e9fda0d7`
982
+ - upstream package version: `0.33.2`
960
983
  - inspected: `agent-browser --version`
961
984
  - inspected: `agent-browser --help`
962
985
  - inspected: `selected agent-browser <command> --help output`
986
+ - inspected: `agent-browser a11y --help`
963
987
  - inspected: `agent-browser mcp --help`
964
988
  - inspected: `agent-browser plugin --help`
989
+ - inspected: `agent-browser skills list`
990
+ - inspected: `agent-browser skills get core --full`
991
+ - inspected: `agent-browser skills get derive-client --full`
965
992
  - inspected: `README.md`
966
993
  - inspected: `CHANGELOG.md`
967
994
  - inspected: `agent-browser.schema.json`
@@ -970,8 +997,17 @@ This generated block is review data for maintainers. The human-authored referenc
970
997
  - inspected: `cli/src/read.rs`
971
998
  - inspected: `cli/src/doctor/webgpu.rs`
972
999
  - inspected: `cli/src/native/actions.rs`
1000
+ - inspected: `cli/src/native/a11y/mod.rs`
1001
+ - inspected: `cli/src/native/browser.rs`
973
1002
  - inspected: `cli/src/native/daemon.rs`
1003
+ - inspected: `cli/src/output.rs`
974
1004
  - inspected: `docs/src/app/webgpu/page.mdx`
1005
+ - inspected: `docs/src/app/network/page.mdx`
1006
+ - inspected: `docs/src/app/selectors/page.mdx`
1007
+ - inspected: `docs/src/app/skills/page.mdx`
1008
+ - inspected: `docs/src/app/commands/page.mdx`
1009
+ - inspected: `skill-data/derive-client/SKILL.md`
1010
+ - inspected: `skill-data/core/SKILL.md`
975
1011
  - inspected: `packages/@agent-browser/eve/README.md`
976
1012
  - inspected: `packages/@agent-browser/eve/package.json`
977
1013
  - inspected: `packages/@agent-browser/eve/test/extension.test.mjs`
@@ -1027,6 +1063,7 @@ This generated block is review data for maintainers. The human-authored referenc
1027
1063
  - record help: `agent-browser record --help`
1028
1064
  - console help: `agent-browser console --help`
1029
1065
  - errors help: `agent-browser errors --help`
1066
+ - a11y help: `agent-browser a11y --help`
1030
1067
  - clipboard help: `agent-browser clipboard --help`
1031
1068
  - tap help: `agent-browser tap --help`
1032
1069
  - swipe help: `agent-browser swipe --help`
@@ -1038,12 +1075,12 @@ This generated block is review data for maintainers. The human-authored referenc
1038
1075
  - plugin help: `agent-browser plugin --help`
1039
1076
 
1040
1077
  #### Inventory sections
1041
- - Built-in skills: 15 human-doc token(s), 15 upstream token(s)
1042
- - Core page, element, navigation, and extraction commands: 81 human-doc token(s), 82 upstream token(s)
1078
+ - Built-in skills: 16 human-doc token(s), 18 upstream token(s)
1079
+ - Core page, element, navigation, and extraction commands: 82 human-doc token(s), 84 upstream token(s)
1043
1080
  - Sessions, state, tabs, frames, dialogs, and windows: 24 human-doc token(s), 20 upstream token(s)
1044
- - Network, storage, artifacts, diagnostics, and performance: 43 human-doc token(s), 53 upstream token(s)
1081
+ - Network, storage, artifacts, diagnostics, and performance: 49 human-doc token(s), 60 upstream token(s)
1045
1082
  - Batch, auth, confirmations, setup, dashboard, devices, and AI commands: 33 human-doc token(s), 37 upstream token(s)
1046
- - Global flags, config, providers, policy, and environment: 138 human-doc token(s), 106 upstream token(s)
1083
+ - Global flags, config, providers, policy, and environment: 142 human-doc token(s), 110 upstream token(s)
1047
1084
 
1048
1085
  #### Human-authored doc tokens required
1049
1086
  ##### Built-in skills
@@ -1058,6 +1095,7 @@ This generated block is review data for maintainers. The human-authored referenc
1058
1095
  - `skills get dogfood`
1059
1096
  - `skills get vercel-sandbox`
1060
1097
  - `skills get agentcore`
1098
+ - `skills get derive-client`
1061
1099
  - `@agent-browser/sandbox`
1062
1100
  - `installSystemDependencies: false`
1063
1101
  - `skills path [name]`
@@ -1139,6 +1177,7 @@ This generated block is review data for maintainers. The human-authored referenc
1139
1177
  - `find last <sel>`
1140
1178
  - `find nth <n> <sel>`
1141
1179
  - `find role <role> --name <name>`
1180
+ - `find role heading text --name`
1142
1181
  - `find ... --exact`
1143
1182
  - `mouse <action> [args]`
1144
1183
  - `set <setting> [value]`
@@ -1179,6 +1218,9 @@ This generated block is review data for maintainers. The human-authored referenc
1179
1218
  - `network requests [--clear] [--filter <pattern>] [--type <csv>] [--method <method>] [--status <code|range>]`
1180
1219
  - `network request <requestId>`
1181
1220
  - `network har start`
1221
+ - `network har start --content all`
1222
+ - `network har start --content none`
1223
+ - `network har start --content text`
1182
1224
  - `network har stop [path]`
1183
1225
  - `cookies [get|set|clear]`
1184
1226
  - `cookies set <name> <value> --url <url> --domain <domain> --path <path> --httpOnly --secure --sameSite <Strict|Lax|None> --expires <timestamp>`
@@ -1215,6 +1257,9 @@ This generated block is review data for maintainers. The human-authored referenc
1215
1257
  - `react suspense [--only-dynamic] [--json]`
1216
1258
  - `vitals [url] [--json]`
1217
1259
  - `web-vitals [url] [--json]`
1260
+ - `a11y [url]`
1261
+ - `a11y --tags wcag2a,wcag2aa`
1262
+ - `a11y --selector "#main"`
1218
1263
  - `removeinitscript <id>`
1219
1264
 
1220
1265
  ##### Batch, auth, confirmations, setup, dashboard, devices, and AI commands
@@ -1300,6 +1345,7 @@ This generated block is review data for maintainers. The human-authored referenc
1300
1345
  - `AGENT_BROWSER_IGNORE_HTTPS_ERRORS`
1301
1346
  - `--allow-file-access`
1302
1347
  - `AGENT_BROWSER_ALLOW_FILE_ACCESS`
1348
+ - `--hide-scrollbars <bool>`
1303
1349
  - `--headed`
1304
1350
  - `AGENT_BROWSER_HEADED`
1305
1351
  - `--webgpu`
@@ -1361,6 +1407,9 @@ This generated block is review data for maintainers. The human-authored referenc
1361
1407
  - `--idle-timeout <ms>`
1362
1408
  - `AGENT_BROWSER_AUTOSAVE_INTERVAL_MS`
1363
1409
  - `AGENT_BROWSER_STREAM_PORT`
1410
+ - `AGENT_BROWSER_STREAM_QUALITY`
1411
+ - `AGENT_BROWSER_STREAM_MAX_WIDTH`
1412
+ - `AGENT_BROWSER_STREAM_MAX_HEIGHT`
1364
1413
  - `AGENT_BROWSER_IDLE_TIMEOUT_MS`
1365
1414
  - `AGENT_BROWSER_ENCRYPTION_KEY`
1366
1415
  - `AGENT_BROWSER_STATE_EXPIRE_DAYS`
@@ -1404,11 +1453,14 @@ This generated block is review data for maintainers. The human-authored referenc
1404
1453
  - skills list: `dogfood`
1405
1454
  - skills list: `vercel-sandbox`
1406
1455
  - skills list: `agentcore`
1456
+ - skills list: `derive-client`
1407
1457
  - vercel sandbox skill full: `@agent-browser/sandbox`
1408
1458
  - vercel sandbox skill full: `installSystemDependencies: false`
1409
1459
  - core skill full: `agent-browser frame @e3`
1410
1460
  - core skill full: `agent-browser dialog accept`
1411
1461
  - core skill full: `agent-browser --session "$SESSION" --restore open https://app.example.com`
1462
+ - core skill full: `network har start --content all`
1463
+ - core skill full: `implicit roles work`
1412
1464
 
1413
1465
  ##### Core page, element, navigation, and extraction commands
1414
1466
  - open help: `open [url]`
@@ -1482,6 +1534,8 @@ This generated block is review data for maintainers. The human-authored referenc
1482
1534
  - find help: `nth <index> <selector>`
1483
1535
  - find help: `--name <name>`
1484
1536
  - find help: `--exact`
1537
+ - find help: `case-insensitive`
1538
+ - find help: `click, fill, check, hover, text`
1485
1539
  - root help: `Mouse: agent-browser mouse <action> [args]`
1486
1540
  - root help: `Browser Settings: agent-browser set <setting> [value]`
1487
1541
  - set help: `media [dark|light]`
@@ -1521,6 +1575,8 @@ This generated block is review data for maintainers. The human-authored referenc
1521
1575
  - root help: `--resource-type <csv>`
1522
1576
  - network help: `unroute [url]`
1523
1577
  - network help: `network har start`
1578
+ - network help: `network har start --content all`
1579
+ - network help: `--content <mode>`
1524
1580
  - network help: `network har stop ./capture.har`
1525
1581
  - root help: `cookies [get|set|clear]`
1526
1582
  - root help: `cookies set --curl <file>`
@@ -1550,6 +1606,11 @@ This generated block is review data for maintainers. The human-authored referenc
1550
1606
  - root help: `react renders stop [--json]`
1551
1607
  - root help: `react suspense [--only-dynamic] [--json]`
1552
1608
  - root help: `vitals [url] [--json]`
1609
+ - root help: `a11y [url] [--tags <t1,t2>] [--selector <css>] [--json]`
1610
+ - a11y help: `a11y [url]`
1611
+ - a11y help: `--tags <tag1,tag2>`
1612
+ - a11y help: `-s, --selector <css>`
1613
+ - core skill full: `agent-browser a11y`
1553
1614
  - root help: `removeinitscript <id>`
1554
1615
  - network help: `requests [options]`
1555
1616
  - network help: `--type <types>`
@@ -1657,6 +1718,7 @@ This generated block is review data for maintainers. The human-authored referenc
1657
1718
  - root help: `AGENT_BROWSER_IGNORE_HTTPS_ERRORS`
1658
1719
  - root help: `--allow-file-access`
1659
1720
  - root help: `AGENT_BROWSER_ALLOW_FILE_ACCESS`
1721
+ - root help: `--hide-scrollbars <bool>`
1660
1722
  - root help: `--headed`
1661
1723
  - root help: `AGENT_BROWSER_HEADED`
1662
1724
  - root help: `--webgpu`
@@ -1711,6 +1773,9 @@ This generated block is review data for maintainers. The human-authored referenc
1711
1773
  - root help: `AGENT_BROWSER_DEFAULT_TIMEOUT`
1712
1774
  - root help: `AGENT_BROWSER_AUTOSAVE_INTERVAL_MS`
1713
1775
  - root help: `AGENT_BROWSER_STREAM_PORT`
1776
+ - root help: `AGENT_BROWSER_STREAM_QUALITY`
1777
+ - root help: `AGENT_BROWSER_STREAM_MAX_WIDTH`
1778
+ - root help: `AGENT_BROWSER_STREAM_MAX_HEIGHT`
1714
1779
  - root help: `AGENT_BROWSER_IDLE_TIMEOUT_MS`
1715
1780
  - root help: `AGENT_BROWSER_ENCRYPTION_KEY`
1716
1781
  - root help: `AGENT_BROWSER_STATE_EXPIRE_DAYS`
package/docs/ELECTRON.md CHANGED
@@ -95,7 +95,7 @@ Then attach and choose a ready target before using refs:
95
95
  { "qa": { "attached": true, "expectedText": "Channels" } }
96
96
  ```
97
97
 
98
- A successful `connect` means the CDP endpoint accepted the session; it does **not** prove the app has an active rendered page yet. Prefer `details.nextActions` when present: `list-connected-session-tabs` inspects the attached session targets. After that read-only list, select or confirm a stable `t<N>` target and run `snapshot -i` explicitly before trusting refs. If the first `snapshot -i` says `No active page`, follow `list-tabs-after-no-active-page`. If it returns no useful refs without that error, manually run `tab list`, select a stable `t<N>` id for the app surface, then retry a condition wait or `snapshot -i` on that selected target.
98
+ A successful `connect` means the CDP endpoint accepted the session; it does **not** prove the app has an active rendered page yet. Prefer `details.nextActions` when present: `verify-connected-session-url` performs the only page read allowed while the attached target is unverified, and `list-connected-session-tabs` inspects attached targets. A verified non-file target clears the guard; otherwise navigate explicitly to a safe URL. After the read-only list, select or confirm a stable `t<N>` target, verify it with `get url`, and run `snapshot -i` explicitly before trusting refs. If the first `snapshot -i` says `No active page`, follow `list-tabs-after-no-active-page`. If it returns no useful refs without that error, manually run `tab list`, select a stable `t<N>` id for the app surface, then retry a condition wait or `snapshot -i` on that selected target.
99
99
 
100
100
  If the app is already running without a debug port, ask before relaunching it — relaunching may lose unsaved state and Electron's single-instance behavior will silently drop a second invocation's `--remote-debugging-port` flag.
101
101
 
@@ -147,15 +147,15 @@ Handoff selection (`handoff` field):
147
147
 
148
148
  | Value | Behavior | When to use |
149
149
  |---|---|---|
150
- | `"snapshot"` (default) | Attach, list targets, capture `snapshot -i` in one call | You need interactive refs immediately for clicks/fills |
151
- | `"tabs"` | Attach and list targets only | Safer diagnostic start when you only need target discovery |
150
+ | `"snapshot"` (default) | Attach, verify `get url`, list targets, capture `snapshot -i` in one call | You need interactive refs immediately for clicks/fills |
151
+ | `"tabs"` | Attach, verify `get url`, and list targets only | Safer diagnostic start when you only need target discovery |
152
152
  | `"connect"` | Attach and stop | You will run your own follow-up commands |
153
153
 
154
154
  `targetType` defaults to `"page"`; use `"webview"` or `"any"` for apps whose useful UI is exposed as a webview target.
155
155
 
156
- Optional `timeoutMs` on `electron.launch` bounds host-side CDP readiness (waiting for `DevToolsActivePort` and attach). When omitted, the default is **15 seconds** with a hard maximum of **120 seconds**, matching `ELECTRON_LAUNCH_DEFAULT_TIMEOUT_MS` and `ELECTRON_LAUNCH_MAX_TIMEOUT_MS` in `extensions/agent-browser/lib/electron/launch.ts`.
156
+ Optional `timeoutMs` on `electron.launch` bounds host-side CDP readiness (waiting for `DevToolsActivePort` and attach). When omitted, the default is **15 seconds** with a hard maximum of **120 seconds**, matching `ELECTRON_LAUNCH_DEFAULT_TIMEOUT_MS` and `ELECTRON_LAUNCH_MAX_TIMEOUT_MS` in `extensions/agent-browser/lib/electron/launch.ts`. Pi cancellation is separate: an already-cancelled call never launches the app, while cancellation during readiness polling or URL/tab/snapshot handoff closes the managed session, stops the tracked process, removes its isolated profile, and returns `failureCategory: "aborted"` without waiting for the launch timeout.
157
157
 
158
- Wrapper-owned launches **always** use an isolated temp profile and an OS-chosen port. `--user-data-dir`, `--remote-debugging-port`, `--remote-debugging-address`, `--remote-debugging-pipe`, and bare `--` in `appArgs` are rejected. There is no caller-supplied port and no way to make `electron.launch` reuse the app's normal signed-in profile or attach to an already-running app — by design. Use the manual path described above when those are the actual requirements.
158
+ Wrapper-owned launches **always** use an isolated temp profile and an OS-chosen port. If wrapper validation, managed-session policy, or the post-attach live-URL handoff guard fails after the host app starts, the wrapper immediately stops that process and removes the isolated profile; it retains a partial tracked record only when cleanup itself cannot finish. `--user-data-dir`, `--remote-debugging-port`, `--remote-debugging-address`, `--remote-debugging-pipe`, and bare `--` in `appArgs` are rejected. There is no caller-supplied port and no way to make `electron.launch` reuse the app's normal signed-in profile or attach to an already-running app — by design. Use the manual path described above when those are the actual requirements.
159
159
 
160
160
  ### `electron.status` — liveness and targets
161
161
 
@@ -167,18 +167,18 @@ Read-only inspection of one or more tracked launches. Without `launchId` or `all
167
167
  { "electron": { "action": "status", "all": true } }
168
168
  ```
169
169
 
170
- Reports `cleanupState`, debug-port and PID liveness, and bounded CDP target metadata under `details.electron.statuses`. Mismatch fields surface when the current managed session or tab no longer matches a live wrapper launch target — typically the cue to follow `reattach-electron-launch` before trusting old refs.
170
+ Reports `cleanupState`, debug-port and PID liveness, and bounded CDP target metadata under `details.electron.statuses`. Its managed-session title/URL reads hold the normal daemon-policy lock and owned restore context. Mismatch fields surface when the current managed session or tab no longer matches a live wrapper launch target — typically the cue to follow `reattach-electron-launch` before trusting old refs.
171
171
 
172
172
  ### `electron.probe` — compact state read
173
173
 
174
- `probe` collapses what would otherwise be separate `get title` / `get url` / focused-element `eval` / `tab list` / `snapshot -i` calls into one bounded result. Use it instead of chaining those reads when you just need a quick "where are we?" check.
174
+ `probe` collapses what would otherwise be separate `get url` / `get title` / focused-element `eval` / `tab list` / `snapshot -i` calls into one bounded result. Use it instead of chaining those reads when you just need a quick "where are we?" check. The wrapper holds the managed-session daemon-policy lock for the probe and runs every underlying read with the session's owned restore decision, so probing cannot restart the daemon under a different restore key.
175
175
 
176
176
  ```json
177
177
  { "electron": { "action": "probe" } }
178
178
  { "electron": { "action": "probe", "launchId": "electron-…", "timeoutMs": 5000 } }
179
179
  ```
180
180
 
181
- Output appears under `details.electron.probe`: `title`, `url`, `focusedElement`, `activeTab`, `tabs`, compact `snapshot` metadata (`refCount`, `refIds`, optional text preview and omission counts), and `errors`. When `launchId` is given, the probe is tied to that tracked launch and will surface mismatch guidance if the wrapper sees a session or target drift; visible output also includes debug-port/pid liveness so a stale `about:blank` against a dead launch is unmistakable.
181
+ Output appears under `details.electron.probe`: `title`, `url`, `focusedElement`, `activeTab`, `tabs`, compact `snapshot` metadata (`refCount`, `refIds`, optional text preview and omission counts), and `errors`. If every underlying read fails, the tool fails with `failureCategory: "upstream-error"`; it does not report a successful empty partial probe. Probes reject a persisted local/unverified target before helper reads, then verify the live URL again before title, eval, tab, or snapshot helpers so external target drift cannot expose local-page content. A current-managed probe also persists top-level `details.namespace`, `sessionTabTarget`, and `refSnapshot` so Pi reload/branch replay keeps the same namespaced page identity; unverified transitions persist as `details.sessionTabTargetUnknown: true` until a safe explicit navigation establishes a trustworthy target. When `launchId` is given, the probe is tied to that tracked launch and will surface mismatch guidance if the wrapper sees a session or target drift; visible output also includes debug-port/pid liveness so a stale `about:blank` against a dead launch is unmistakable.
182
182
 
183
183
  `timeoutMs` bounds each underlying read subprocess. Use it for dense desktop apps when the default budget is too short, or to fail fast when you suspect the app process is wedged.
184
184
 
@@ -209,9 +209,9 @@ On Pi `quit`, active wrapper-owned Electron launches are best-effort cleaned. On
209
209
  | Action | What `timeoutMs` covers when set | Typical default when omitted |
210
210
  | --- | --- | --- |
211
211
  | `launch` | Host-side wait for `DevToolsActivePort` and CDP readiness | **15 s**, hard-capped at **120 s** (`normalizeTimeoutMs` in `extensions/agent-browser/lib/electron/launch.ts`) |
212
- | `status` | Optional managed-session `get title` / `get url` reads used for mismatch diagnostics | Normal tool subprocess budget from `runAgentBrowserProcess` / `AGENT_BROWSER_DEFAULT_TIMEOUT`; localhost CDP HTTP probes keep a short fixed budget (`ELECTRON_STATUS_FETCH_TIMEOUT_MS` in `extensions/agent-browser/lib/electron/cleanup.ts`) |
212
+ | `status` | Optional managed-session `get url`, then `get title` reads used for mismatch diagnostics | Normal tool subprocess budget from `runAgentBrowserProcess` / `AGENT_BROWSER_DEFAULT_TIMEOUT`; localhost CDP HTTP probes keep a short fixed budget (`ELECTRON_STATUS_FETCH_TIMEOUT_MS` in `extensions/agent-browser/lib/electron/cleanup.ts`) |
213
213
  | `cleanup` | One combined budget for managed-session `close`, tracked process exit, debug-port verification, and temp profile removal | `PI_AGENT_BROWSER_IMPLICIT_SESSION_CLOSE_TIMEOUT_MS` when set, else **5000 ms** (`getImplicitSessionCloseTimeoutMs` in `extensions/agent-browser/lib/runtime.ts`, passed through `cleanupTrackedElectronHostLaunches` in `extensions/agent-browser/lib/orchestration/electron-host/index.ts`) |
214
- | `probe` | **Each** upstream read in the probe chain (`get title`, `get url`, focused `eval --stdin`, `tab list`, `snapshot -i`) | Same default as other tool calls (typically **28 s** per subprocess unless `AGENT_BROWSER_DEFAULT_TIMEOUT` / `PI_AGENT_BROWSER_PROCESS_TIMEOUT_MS` overrides `runAgentBrowserProcess` in `extensions/agent-browser/lib/process.ts`) |
214
+ | `probe` | **Each** upstream read in the probe chain (`get url`, then `get title`, focused `eval --stdin`, `tab list`, `snapshot -i`) | Same default as other tool calls (typically **28 s** per subprocess unless `AGENT_BROWSER_DEFAULT_TIMEOUT` / `PI_AGENT_BROWSER_PROCESS_TIMEOUT_MS` overrides `runAgentBrowserProcess` in `extensions/agent-browser/lib/process.ts`) |
215
215
 
216
216
  ## `qa.attached` — current-session smoke check
217
217