pi-chrome 0.15.51 → 0.15.54

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -2,6 +2,32 @@
2
2
 
3
3
  All notable user-facing changes to `pi-chrome`.
4
4
 
5
+ ## 0.15.54 — 2026-09-26
6
+
7
+ - **Hidden tabs report instead of silently dropping input.** Chrome ignores trusted input to hidden pages. A page is hidden when it is an inactive tab, its window is minimized, its window is fully covered by another window, or (on macOS) its window became a hidden window tab behind a full-screen window. In a hidden page, mouse presses and keys were dropped while the tool reported success, and mouse moves could hang until a timeout (the source of `Detached while handling command` in background runs). Input tools now check page visibility first. `chrome_type`, `chrome_key`, `chrome_hover`, `chrome_scroll`, `chrome_tap` and `chrome_drag` fail fast with the reason and the options (`background:false`, or `/chrome background off`). `chrome_click` and `chrome_fill` go straight to their existing DOM fallback and report the reason, or reject with `domFallback:false`. Hard background mode is unchanged: nothing is activated or focused.
8
+ - **No hidden automation windows in full screen.** While your Chrome window is full screen, a new automation window would become a hidden macOS window tab or land on another Space. The automation target is now an inactive tab in your window instead.
9
+ - **Targeted `chrome_type` appends.** With a uid/selector, `chrome_type` clicked a random point inside the field, which left the caret mid-text and spliced input into existing content. After the focus click it now moves the caret to the end with Chrome's `moveToEndOfDocument` editing command. Untargeted typing still uses the current caret; `replace:true` is unchanged.
10
+ - **Upload buttons, labels, wrappers, and iframes.** `chrome_upload_file` accepts a label for a file input, or a wrapper/dropzone containing exactly one, and searches same-origin iframes for selectors. For upload buttons that open the native picker (no persistent `<input type=file>`), it intercepts the file chooser (`Page.setInterceptFileChooserDialog`) so no OS dialog appears. The trigger click uses trusted Chrome input on visible pages. On hidden pages it uses `element.click()` with a CDP user gesture, unless `domFallback:false`. It rejects multiple paths for single-file inputs. Adapted from kkunkunya's fork.
11
+ - **No duplicate `change` on upload.** Chrome's `DOM.setFileInputFiles` already fires `input`/`change`, and pi-chrome dispatched a second pair, which could make apps upload twice. It now dispatches only when Chrome did not, and reports `events: native|dispatched`.
12
+ - **Bounded screenshot folder.** Default-path `chrome_screenshot` captures prune `.pi/chrome-screenshots` at capture time. Captures older than 7 days are removed, but the newest 20 are always kept. Only files matching the tool's own naming (including full-page tiles and manifests) are eligible, and explicit `path:` captures are never pruned. `retentionDays` overrides the window; `0` disables pruning. Adapted from ardhiqii's fork.
13
+ - **`input.debug` without a target.** It no longer fails with `Value must be at least 0`; it describes this session's automation tab if one exists.
14
+ - **Validation.** New `hidden-input` unit suite. Live-checked on Chrome/macOS: hidden-tab errors and DOM fallbacks, caret-end typing in inputs and contenteditables, all upload shapes in visible and hidden tabs, single `change` events, and pruning through a real Pi session.
15
+
16
+ ## 0.15.53 — 2026-09-26
17
+
18
+ Changes adapted from community forks (ardhiqii, kkunkunya, nihar-oracle, steimerbyte).
19
+
20
+ - **Connector recovers from stalled connections.** The extension's `/next` long poll now gives up after 45s (the bridge holds it for up to 25s). Before this, a half-open socket (Pi process died, machine slept, network changed) could leave the extension waiting forever until someone reloaded it by hand. Result posts time out, retry on network/5xx failures, and are never posted twice for the same command.
21
+ - **Page actions work on a fresh automation tab.** Chrome refuses `chrome.scripting` on a top-level `about:blank`, which is where the dedicated automation tab starts. As a result, `chrome_snapshot`, `chrome_inspect`, console/network listing, uid click targeting, and the `/chrome doctor` page probe failed on that tab with `Cannot access contents of url "about:blank"`. When Chrome refuses scripting for a permission reason, these calls now fall back to CDP `Runtime.evaluate`. Ordinary pages still use `chrome.scripting` first.
22
+ - **Chrome tools in subagent sessions.** A second load of the same pi-chrome install in one process (for example, pi-subagents) was skipped as a duplicate, so subagents had no `chrome_*` tools. It now loads as a client of the shared bridge. Subagents share the parent's authorization and get their own automation tab and tab group. Two different install roots are still treated as duplicates. `chrome_launch` in a client session now reports the shared connection state instead of always saying "waiting for extension".
23
+ - **New `chrome_cdp` and `chrome_cdp_targets` tools.** `chrome_cdp` runs one raw Chrome DevTools Protocol method on a tab and validates the method/params first. `timeoutMs` can extend the deadline up to 120s, and screenshot/binary or oversized results are summarized instead of returned in full. Background mode blocks `Page.bringToFront` and `Target.activateTarget`. `chrome_cdp_targets` lists CDP targets on a tab, such as password-manager overlays or DevTools, to help diagnose `Detached while handling command`. It never attaches the debugger or creates a tab.
24
+ - **`chrome_type` shows what it typed.** Results include the field value before and after, and `insertedAt` (`empty`, `caret-end`, `caret-middle`, `replaced-selection`, `replaced-all`). The text warns when input was spliced into existing content, when Enter may have submitted the spliced value, and when the field did not change at all (keystrokes did not reach it). Password/OTP/card-like fields report only lengths. New `replace: true` selects all with Chrome's platform-neutral `selectAll` editing command and deletes before typing. Tool descriptions now state that `chrome_type` inserts at the caret and `chrome_fill` replaces.
25
+ - **`includeSnapshot` waits for navigations.** If a click/type/fill/key action starts a navigation, the included snapshot waits up to 5s for the new page and reports `navigation {from, to, settled, waitedMs}`. Before, it could describe the page being replaced. Actions without navigation do not wait. A page that was already loading is reported but not waited on.
26
+ - **CDP timeouts report as timeouts.** A timed-out CDP command now fails with `CDP <method> timed out after Nms`. Before, the cleanup detach surfaced as `Detached while handling command`, and the command was re-sent once, so a slow command could run twice.
27
+ - **`chrome_tab new` reports the loaded tab.** It waits up to 5s for the URL to load and returns `loadStatus` instead of Chrome's initial empty `loading` tab.
28
+ - **Tool activation.** `chrome_find`, `chrome_inspect`, and the new CDP tools are now activated by `/chrome authorize` and removed from the active tool set by `/chrome revoke` or grant expiry, like the other `chrome_*` tools.
29
+ - **Validation.** New unit suites: `bridge-resilience`, `cdp-passthrough`, `type-evidence`.
30
+
5
31
  ## 0.15.51 — 2026-09-10
6
32
 
7
33
  - **Fewer Chrome commands.** Removed `/chrome status`; use bare `/chrome` for the quick connection, authorization, and background dashboard plus controls. The dashboard remains lightweight and does not run page probes.
package/README.md CHANGED
@@ -16,7 +16,7 @@ Try prompts like these after setup:
16
16
  | **Understand an existing page** | “Find my open dashboard tab and summarize what's on the page. Don't change anything.” |
17
17
  | **Create evidence for a PR** | “On my local app, capture the empty, loading, and populated states of this feature for my PR.” |
18
18
 
19
- Pi gets tools to inspect pages, click, type, fill forms, scroll, upload files, capture screenshots, and inspect captured console logs and `fetch`/`XMLHttpRequest` responses. You describe the task; Pi handles the agent loop.
19
+ Pi gets tools to inspect pages, click, type, fill forms, scroll, upload files, capture screenshots, and inspect captured console logs and `fetch`/`XMLHttpRequest` responses. A raw Chrome DevTools Protocol tool (`chrome_cdp`) covers anything else, such as device emulation, cookies, PDFs, or the accessibility tree. You describe the task; Pi handles the agent loop.
20
20
 
21
21
  **Best fit:** interactive workflows in the Chrome profile you already use. For deterministic CI tests, consider a test framework such as Playwright; for fleets of isolated browsers, consider a hosted browser service. See [more workflows](./docs/EXAMPLES.md) and [browser-tool comparisons](./docs/COMPARISON.md).
22
22
 
@@ -85,7 +85,7 @@ Run `/chrome revoke` when finished. Use `/chrome authorize` again whenever you w
85
85
 
86
86
  This is browser automation, not full OS control. Native Chrome/OS dialogs, password-manager prompts, passkeys/security keys/biometrics, CAPTCHA challenges, cross-origin iframe DOM access, rich multitouch/stylus gestures, and arbitrary desktop apps are outside its reliable tool surface. Some workflows need human assistance.
87
87
 
88
- If page inspection or evaluation is blocked, use screenshots and coordinate input where possible. Background pages can throttle rendering or reject focus-gated actions. See the [FAQ](./docs/FAQ.md) for details.
88
+ If page inspection or evaluation is blocked, use screenshots and coordinate input where possible. Background pages can throttle rendering or reject focus-gated actions. Chrome ignores real input to hidden pages (inactive tabs, covered or minimized windows, windows behind a macOS full-screen window); input tools report this instead of pretending to succeed. See the [FAQ](./docs/FAQ.md) for details.
89
89
 
90
90
  ## Commands
91
91
 
package/SECURITY.md CHANGED
@@ -27,6 +27,7 @@ The Chrome extension under `extensions/chrome-profile-bridge/browser-extension/`
27
27
  - Loopback bridge only. No remote port. No telemetry.
28
28
  - Chrome real input layer for interactive controls.
29
29
  - Chrome control locked by default; `/chrome authorize` unlocks current Pi session after terminal confirmation, `/chrome revoke` locks it again.
30
+ - `chrome_cdp` sends raw Chrome DevTools Protocol commands to a tab. It is locked behind `/chrome authorize` like every other tool, but it is not filtered against a safe list. Background mode blocks only its explicit focus methods (`Page.bringToFront`, `Target.activateTarget`).
30
31
  - Hard background mode is on by default: tools cannot override it to explicitly focus windows or activate tabs. `/chrome background off` allows foreground/watch mode. This is not a security sandbox: trusted input, page scripts, native prompts, and Chrome/OS behavior can still affect focus.
31
32
 
32
33
  ## Custom ports
package/docs/FAQ.md CHANGED
@@ -32,6 +32,8 @@ Chrome control is also locked per Pi session until you run `/chrome authorize`;
32
32
 
33
33
  Yes. The first session opens the local bridge; later sessions detect it and pipe their commands through the same bridge. Each Pi session must be authorized with `/chrome authorize` before its chrome_* tools work. Each session also owns its **own** dedicated automation window (ownership is keyed by session id inside the one extension), so concurrent sessions never navigate into or close each other's tabs.
34
34
 
35
+ Subagent sessions that run inside an authorized Pi process (for example with pi-subagents) also get the chrome_* tools. They share the parent's `/chrome authorize` grant and bridge, but each subagent gets its own automation tab and tab group, which are cleaned up when the subagent session ends.
36
+
35
37
  ## Does pi-chrome navigate my current tab?
36
38
 
37
39
  No. The first chrome_* action that has no explicit target opens a **dedicated automation window** that pi-chrome owns (falling back to a dedicated tab only if a separate window can't be created), and reuses it for the rest of the session. Your existing tabs and windows are never reused or overwritten. Pass `targetId`/`urlIncludes`/`titleIncludes` to deliberately act on a tab you already have open.
@@ -50,7 +52,7 @@ pi-chrome ships as an unpacked extension so the source and broad browser permiss
50
52
 
51
53
  ## What's the install footprint?
52
54
 
53
- - Pi side: one extension that registers 19 tools and a few slash commands.
55
+ - Pi side: one extension that registers 23 tools and a few slash commands.
54
56
  - Chrome side: one unpacked extension, ~2000 LOC of plain JavaScript, no dependencies.
55
57
 
56
58
  ## Can I script it without Pi?
@@ -73,7 +75,7 @@ If the page did not change, take a fresh snapshot or screenshot and check for ov
73
75
 
74
76
  ## How do I attach a file to a React file input?
75
77
 
76
- `chrome_upload_file` — uses Chrome DevTools file-input control and fires `input` + `change` events. It does **not** open the native file picker. Works with React/Vue/Angular controlled inputs.
78
+ `chrome_upload_file` — uses Chrome DevTools file-input control; the page gets one `input` + `change`. Target the file input, a label or wrapper containing one, or an upload button: for buttons, the native file chooser is intercepted, so no OS dialog appears. Selectors also search same-origin iframes. Works with React/Vue/Angular controlled inputs.
77
79
 
78
80
  ## Can it record videos?
79
81
 
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "manifest_version": 3,
3
3
  "name": "Pi Chrome Connector",
4
- "version": "0.15.51",
4
+ "version": "0.15.54",
5
5
  "description": "Lets Pi control tabs in Chrome via a local connector at 127.0.0.1.",
6
6
  "permissions": [
7
7
  "tabs",