@brunosps00/dev-workflow 0.0.5 → 0.0.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (164) hide show
  1. package/bin/dev-workflow.js +6 -4
  2. package/lib/init.js +28 -11
  3. package/package.json +1 -1
  4. package/scaffold/pt-br/commands/dw-analyze-project.md +3 -3
  5. package/scaffold/pt-br/commands/dw-bugfix.md +6 -6
  6. package/scaffold/pt-br/commands/dw-code-review.md +2 -2
  7. package/scaffold/pt-br/commands/dw-create-tasks.md +4 -4
  8. package/scaffold/pt-br/commands/dw-generate-pr.md +3 -3
  9. package/scaffold/pt-br/commands/dw-help.md +50 -50
  10. package/scaffold/pt-br/commands/dw-review-implementation.md +3 -3
  11. package/scaffold/pt-br/commands/dw-run-plan.md +8 -8
  12. package/scaffold/pt-br/commands/dw-run-task.md +3 -3
  13. package/scaffold/pt-br/templates/tasks-template.md +2 -2
  14. package/scaffold/skills/agent-browser/SKILL.md +750 -0
  15. package/scaffold/skills/agent-browser/references/authentication.md +303 -0
  16. package/scaffold/skills/agent-browser/references/commands.md +295 -0
  17. package/scaffold/skills/agent-browser/references/profiling.md +120 -0
  18. package/scaffold/skills/agent-browser/references/proxy-support.md +194 -0
  19. package/scaffold/skills/agent-browser/references/session-management.md +193 -0
  20. package/scaffold/skills/agent-browser/references/snapshot-refs.md +219 -0
  21. package/scaffold/skills/agent-browser/references/video-recording.md +173 -0
  22. package/scaffold/skills/agent-browser/templates/authenticated-session.sh +105 -0
  23. package/scaffold/skills/agent-browser/templates/capture-workflow.sh +69 -0
  24. package/scaffold/skills/agent-browser/templates/form-automation.sh +62 -0
  25. package/scaffold/skills/humanizer/README.md +143 -0
  26. package/scaffold/skills/humanizer/SKILL.md +488 -0
  27. package/scaffold/skills/humanizer/WARP.md +53 -0
  28. package/scaffold/skills/remotion-best-practices/SKILL.md +61 -0
  29. package/scaffold/skills/remotion-best-practices/rules/3d.md +86 -0
  30. package/scaffold/skills/remotion-best-practices/rules/animations.md +27 -0
  31. package/scaffold/skills/remotion-best-practices/rules/assets/charts-bar-chart.tsx +173 -0
  32. package/scaffold/skills/remotion-best-practices/rules/assets/text-animations-typewriter.tsx +100 -0
  33. package/scaffold/skills/remotion-best-practices/rules/assets/text-animations-word-highlight.tsx +103 -0
  34. package/scaffold/skills/remotion-best-practices/rules/assets.md +78 -0
  35. package/scaffold/skills/remotion-best-practices/rules/audio-visualization.md +198 -0
  36. package/scaffold/skills/remotion-best-practices/rules/audio.md +169 -0
  37. package/scaffold/skills/remotion-best-practices/rules/calculate-metadata.md +134 -0
  38. package/scaffold/skills/remotion-best-practices/rules/can-decode.md +75 -0
  39. package/scaffold/skills/remotion-best-practices/rules/charts.md +120 -0
  40. package/scaffold/skills/remotion-best-practices/rules/compositions.md +154 -0
  41. package/scaffold/skills/remotion-best-practices/rules/display-captions.md +184 -0
  42. package/scaffold/skills/remotion-best-practices/rules/extract-frames.md +229 -0
  43. package/scaffold/skills/remotion-best-practices/rules/ffmpeg.md +38 -0
  44. package/scaffold/skills/remotion-best-practices/rules/fonts.md +152 -0
  45. package/scaffold/skills/remotion-best-practices/rules/get-audio-duration.md +58 -0
  46. package/scaffold/skills/remotion-best-practices/rules/get-video-dimensions.md +68 -0
  47. package/scaffold/skills/remotion-best-practices/rules/get-video-duration.md +60 -0
  48. package/scaffold/skills/remotion-best-practices/rules/gifs.md +141 -0
  49. package/scaffold/skills/remotion-best-practices/rules/images.md +134 -0
  50. package/scaffold/skills/remotion-best-practices/rules/import-srt-captions.md +69 -0
  51. package/scaffold/skills/remotion-best-practices/rules/light-leaks.md +73 -0
  52. package/scaffold/skills/remotion-best-practices/rules/lottie.md +70 -0
  53. package/scaffold/skills/remotion-best-practices/rules/maps.md +412 -0
  54. package/scaffold/skills/remotion-best-practices/rules/measuring-dom-nodes.md +34 -0
  55. package/scaffold/skills/remotion-best-practices/rules/measuring-text.md +140 -0
  56. package/scaffold/skills/remotion-best-practices/rules/parameters.md +109 -0
  57. package/scaffold/skills/remotion-best-practices/rules/sequencing.md +118 -0
  58. package/scaffold/skills/remotion-best-practices/rules/sfx.md +26 -0
  59. package/scaffold/skills/remotion-best-practices/rules/subtitles.md +36 -0
  60. package/scaffold/skills/remotion-best-practices/rules/tailwind.md +11 -0
  61. package/scaffold/skills/remotion-best-practices/rules/text-animations.md +20 -0
  62. package/scaffold/skills/remotion-best-practices/rules/timing.md +179 -0
  63. package/scaffold/skills/remotion-best-practices/rules/transcribe-captions.md +70 -0
  64. package/scaffold/skills/remotion-best-practices/rules/transitions.md +197 -0
  65. package/scaffold/skills/remotion-best-practices/rules/transparent-videos.md +106 -0
  66. package/scaffold/skills/remotion-best-practices/rules/trimming.md +51 -0
  67. package/scaffold/skills/remotion-best-practices/rules/videos.md +171 -0
  68. package/scaffold/skills/remotion-best-practices/rules/voiceover.md +99 -0
  69. package/scaffold/skills/security-review/LICENSE +22 -0
  70. package/scaffold/skills/security-review/SKILL.md +312 -0
  71. package/scaffold/skills/security-review/infrastructure/docker.md +432 -0
  72. package/scaffold/skills/security-review/languages/javascript.md +388 -0
  73. package/scaffold/skills/security-review/languages/python.md +363 -0
  74. package/scaffold/skills/security-review/references/api-security.md +519 -0
  75. package/scaffold/skills/security-review/references/authentication.md +353 -0
  76. package/scaffold/skills/security-review/references/authorization.md +372 -0
  77. package/scaffold/skills/security-review/references/business-logic.md +443 -0
  78. package/scaffold/skills/security-review/references/cryptography.md +329 -0
  79. package/scaffold/skills/security-review/references/csrf.md +398 -0
  80. package/scaffold/skills/security-review/references/data-protection.md +378 -0
  81. package/scaffold/skills/security-review/references/deserialization.md +410 -0
  82. package/scaffold/skills/security-review/references/error-handling.md +436 -0
  83. package/scaffold/skills/security-review/references/file-security.md +457 -0
  84. package/scaffold/skills/security-review/references/injection.md +259 -0
  85. package/scaffold/skills/security-review/references/logging.md +433 -0
  86. package/scaffold/skills/security-review/references/misconfiguration.md +435 -0
  87. package/scaffold/skills/security-review/references/modern-threats.md +475 -0
  88. package/scaffold/skills/security-review/references/ssrf.md +415 -0
  89. package/scaffold/skills/security-review/references/supply-chain.md +405 -0
  90. package/scaffold/skills/security-review/references/xss.md +336 -0
  91. package/scaffold/skills/vercel-react-best-practices/AGENTS.md +3648 -0
  92. package/scaffold/skills/vercel-react-best-practices/README.md +123 -0
  93. package/scaffold/skills/vercel-react-best-practices/SKILL.md +146 -0
  94. package/scaffold/skills/vercel-react-best-practices/rules/_sections.md +46 -0
  95. package/scaffold/skills/vercel-react-best-practices/rules/_template.md +28 -0
  96. package/scaffold/skills/vercel-react-best-practices/rules/advanced-event-handler-refs.md +55 -0
  97. package/scaffold/skills/vercel-react-best-practices/rules/advanced-init-once.md +42 -0
  98. package/scaffold/skills/vercel-react-best-practices/rules/advanced-use-latest.md +39 -0
  99. package/scaffold/skills/vercel-react-best-practices/rules/async-api-routes.md +38 -0
  100. package/scaffold/skills/vercel-react-best-practices/rules/async-cheap-condition-before-await.md +37 -0
  101. package/scaffold/skills/vercel-react-best-practices/rules/async-defer-await.md +82 -0
  102. package/scaffold/skills/vercel-react-best-practices/rules/async-dependencies.md +51 -0
  103. package/scaffold/skills/vercel-react-best-practices/rules/async-parallel.md +28 -0
  104. package/scaffold/skills/vercel-react-best-practices/rules/async-suspense-boundaries.md +99 -0
  105. package/scaffold/skills/vercel-react-best-practices/rules/bundle-barrel-imports.md +60 -0
  106. package/scaffold/skills/vercel-react-best-practices/rules/bundle-conditional.md +31 -0
  107. package/scaffold/skills/vercel-react-best-practices/rules/bundle-defer-third-party.md +49 -0
  108. package/scaffold/skills/vercel-react-best-practices/rules/bundle-dynamic-imports.md +35 -0
  109. package/scaffold/skills/vercel-react-best-practices/rules/bundle-preload.md +50 -0
  110. package/scaffold/skills/vercel-react-best-practices/rules/client-event-listeners.md +74 -0
  111. package/scaffold/skills/vercel-react-best-practices/rules/client-localstorage-schema.md +71 -0
  112. package/scaffold/skills/vercel-react-best-practices/rules/client-passive-event-listeners.md +48 -0
  113. package/scaffold/skills/vercel-react-best-practices/rules/client-swr-dedup.md +56 -0
  114. package/scaffold/skills/vercel-react-best-practices/rules/js-batch-dom-css.md +107 -0
  115. package/scaffold/skills/vercel-react-best-practices/rules/js-cache-function-results.md +80 -0
  116. package/scaffold/skills/vercel-react-best-practices/rules/js-cache-property-access.md +28 -0
  117. package/scaffold/skills/vercel-react-best-practices/rules/js-cache-storage.md +70 -0
  118. package/scaffold/skills/vercel-react-best-practices/rules/js-combine-iterations.md +32 -0
  119. package/scaffold/skills/vercel-react-best-practices/rules/js-early-exit.md +50 -0
  120. package/scaffold/skills/vercel-react-best-practices/rules/js-flatmap-filter.md +60 -0
  121. package/scaffold/skills/vercel-react-best-practices/rules/js-hoist-regexp.md +45 -0
  122. package/scaffold/skills/vercel-react-best-practices/rules/js-index-maps.md +37 -0
  123. package/scaffold/skills/vercel-react-best-practices/rules/js-length-check-first.md +49 -0
  124. package/scaffold/skills/vercel-react-best-practices/rules/js-min-max-loop.md +82 -0
  125. package/scaffold/skills/vercel-react-best-practices/rules/js-request-idle-callback.md +105 -0
  126. package/scaffold/skills/vercel-react-best-practices/rules/js-set-map-lookups.md +24 -0
  127. package/scaffold/skills/vercel-react-best-practices/rules/js-tosorted-immutable.md +57 -0
  128. package/scaffold/skills/vercel-react-best-practices/rules/rendering-activity.md +26 -0
  129. package/scaffold/skills/vercel-react-best-practices/rules/rendering-animate-svg-wrapper.md +47 -0
  130. package/scaffold/skills/vercel-react-best-practices/rules/rendering-conditional-render.md +40 -0
  131. package/scaffold/skills/vercel-react-best-practices/rules/rendering-content-visibility.md +38 -0
  132. package/scaffold/skills/vercel-react-best-practices/rules/rendering-hoist-jsx.md +46 -0
  133. package/scaffold/skills/vercel-react-best-practices/rules/rendering-hydration-no-flicker.md +82 -0
  134. package/scaffold/skills/vercel-react-best-practices/rules/rendering-hydration-suppress-warning.md +30 -0
  135. package/scaffold/skills/vercel-react-best-practices/rules/rendering-resource-hints.md +85 -0
  136. package/scaffold/skills/vercel-react-best-practices/rules/rendering-script-defer-async.md +68 -0
  137. package/scaffold/skills/vercel-react-best-practices/rules/rendering-svg-precision.md +28 -0
  138. package/scaffold/skills/vercel-react-best-practices/rules/rendering-usetransition-loading.md +75 -0
  139. package/scaffold/skills/vercel-react-best-practices/rules/rerender-defer-reads.md +39 -0
  140. package/scaffold/skills/vercel-react-best-practices/rules/rerender-dependencies.md +45 -0
  141. package/scaffold/skills/vercel-react-best-practices/rules/rerender-derived-state-no-effect.md +40 -0
  142. package/scaffold/skills/vercel-react-best-practices/rules/rerender-derived-state.md +29 -0
  143. package/scaffold/skills/vercel-react-best-practices/rules/rerender-functional-setstate.md +74 -0
  144. package/scaffold/skills/vercel-react-best-practices/rules/rerender-lazy-state-init.md +58 -0
  145. package/scaffold/skills/vercel-react-best-practices/rules/rerender-memo-with-default-value.md +38 -0
  146. package/scaffold/skills/vercel-react-best-practices/rules/rerender-memo.md +44 -0
  147. package/scaffold/skills/vercel-react-best-practices/rules/rerender-move-effect-to-event.md +45 -0
  148. package/scaffold/skills/vercel-react-best-practices/rules/rerender-no-inline-components.md +82 -0
  149. package/scaffold/skills/vercel-react-best-practices/rules/rerender-simple-expression-in-memo.md +35 -0
  150. package/scaffold/skills/vercel-react-best-practices/rules/rerender-split-combined-hooks.md +64 -0
  151. package/scaffold/skills/vercel-react-best-practices/rules/rerender-transitions.md +40 -0
  152. package/scaffold/skills/vercel-react-best-practices/rules/rerender-use-deferred-value.md +59 -0
  153. package/scaffold/skills/vercel-react-best-practices/rules/rerender-use-ref-transient-values.md +73 -0
  154. package/scaffold/skills/vercel-react-best-practices/rules/server-after-nonblocking.md +73 -0
  155. package/scaffold/skills/vercel-react-best-practices/rules/server-auth-actions.md +96 -0
  156. package/scaffold/skills/vercel-react-best-practices/rules/server-cache-lru.md +41 -0
  157. package/scaffold/skills/vercel-react-best-practices/rules/server-cache-react.md +76 -0
  158. package/scaffold/skills/vercel-react-best-practices/rules/server-dedup-props.md +65 -0
  159. package/scaffold/skills/vercel-react-best-practices/rules/server-hoist-static-io.md +149 -0
  160. package/scaffold/skills/vercel-react-best-practices/rules/server-parallel-fetching.md +83 -0
  161. package/scaffold/skills/vercel-react-best-practices/rules/server-parallel-nested-fetching.md +34 -0
  162. package/scaffold/skills/vercel-react-best-practices/rules/server-serialization.md +38 -0
  163. package/scaffold/skills/webapp-testing/SKILL.md +133 -0
  164. package/scaffold/skills/webapp-testing/assets/test-helper.js +56 -0
@@ -0,0 +1,750 @@
1
+ ---
2
+ name: agent-browser
3
+ description: Browser automation CLI for AI agents. Use when the user needs to interact with websites, including navigating pages, filling forms, clicking buttons, taking screenshots, extracting data, testing web apps, or automating any browser task. Triggers include requests to "open a website", "fill out a form", "click a button", "take a screenshot", "scrape data from a page", "test this web app", "login to a site", "automate browser actions", or any task requiring programmatic web interaction.
4
+ allowed-tools: Bash(npx agent-browser:*), Bash(agent-browser:*)
5
+ ---
6
+
7
+ # Browser Automation with agent-browser
8
+
9
+ The CLI uses Chrome/Chromium via CDP directly. Install via `npm i -g agent-browser`, `brew install agent-browser`, or `cargo install agent-browser`. Run `agent-browser install` to download Chrome. Existing Chrome, Brave, Playwright, and Puppeteer installations are detected automatically. Run `agent-browser upgrade` to update to the latest version.
10
+
11
+ ## Core Workflow
12
+
13
+ Every browser automation follows this pattern:
14
+
15
+ 1. **Navigate**: `agent-browser open <url>`
16
+ 2. **Snapshot**: `agent-browser snapshot -i` (get element refs like `@e1`, `@e2`)
17
+ 3. **Interact**: Use refs to click, fill, select
18
+ 4. **Re-snapshot**: After navigation or DOM changes, get fresh refs
19
+
20
+ ```bash
21
+ agent-browser open https://example.com/form
22
+ agent-browser snapshot -i
23
+ # Output: @e1 [input type="email"], @e2 [input type="password"], @e3 [button] "Submit"
24
+
25
+ agent-browser fill @e1 "user@example.com"
26
+ agent-browser fill @e2 "password123"
27
+ agent-browser click @e3
28
+ agent-browser wait --load networkidle
29
+ agent-browser snapshot -i # Check result
30
+ ```
31
+
32
+ ## Command Chaining
33
+
34
+ Commands can be chained with `&&` in a single shell invocation. The browser persists between commands via a background daemon, so chaining is safe and more efficient than separate calls.
35
+
36
+ ```bash
37
+ # Chain open + wait + snapshot in one call
38
+ agent-browser open https://example.com && agent-browser wait --load networkidle && agent-browser snapshot -i
39
+
40
+ # Chain multiple interactions
41
+ agent-browser fill @e1 "user@example.com" && agent-browser fill @e2 "password123" && agent-browser click @e3
42
+
43
+ # Navigate and capture
44
+ agent-browser open https://example.com && agent-browser wait --load networkidle && agent-browser screenshot page.png
45
+ ```
46
+
47
+ **When to chain:** Use `&&` when you don't need to read the output of an intermediate command before proceeding (e.g., open + wait + screenshot). Run commands separately when you need to parse the output first (e.g., snapshot to discover refs, then interact using those refs).
48
+
49
+ ## Handling Authentication
50
+
51
+ When automating a site that requires login, choose the approach that fits:
52
+
53
+ **Option 1: Import auth from the user's browser (fastest for one-off tasks)**
54
+
55
+ ```bash
56
+ # Connect to the user's running Chrome (they're already logged in)
57
+ agent-browser --auto-connect state save ./auth.json
58
+ # Use that auth state
59
+ agent-browser --state ./auth.json open https://app.example.com/dashboard
60
+ ```
61
+
62
+ State files contain session tokens in plaintext -- add to `.gitignore` and delete when no longer needed. Set `AGENT_BROWSER_ENCRYPTION_KEY` for encryption at rest.
63
+
64
+ **Option 2: Persistent profile (simplest for recurring tasks)**
65
+
66
+ ```bash
67
+ # First run: login manually or via automation
68
+ agent-browser --profile ~/.myapp open https://app.example.com/login
69
+ # ... fill credentials, submit ...
70
+
71
+ # All future runs: already authenticated
72
+ agent-browser --profile ~/.myapp open https://app.example.com/dashboard
73
+ ```
74
+
75
+ **Option 3: Session name (auto-save/restore cookies + localStorage)**
76
+
77
+ ```bash
78
+ agent-browser --session-name myapp open https://app.example.com/login
79
+ # ... login flow ...
80
+ agent-browser close # State auto-saved
81
+
82
+ # Next time: state auto-restored
83
+ agent-browser --session-name myapp open https://app.example.com/dashboard
84
+ ```
85
+
86
+ **Option 4: Auth vault (credentials stored encrypted, login by name)**
87
+
88
+ ```bash
89
+ echo "$PASSWORD" | agent-browser auth save myapp --url https://app.example.com/login --username user --password-stdin
90
+ agent-browser auth login myapp
91
+ ```
92
+
93
+ `auth login` navigates with `load` and then waits for login form selectors to appear before filling/clicking, which is more reliable on delayed SPA login screens.
94
+
95
+ **Option 5: State file (manual save/load)**
96
+
97
+ ```bash
98
+ # After logging in:
99
+ agent-browser state save ./auth.json
100
+ # In a future session:
101
+ agent-browser state load ./auth.json
102
+ agent-browser open https://app.example.com/dashboard
103
+ ```
104
+
105
+ See [references/authentication.md](references/authentication.md) for OAuth, 2FA, cookie-based auth, and token refresh patterns.
106
+
107
+ ## Essential Commands
108
+
109
+ ```bash
110
+ # Navigation
111
+ agent-browser open <url> # Navigate (aliases: goto, navigate)
112
+ agent-browser close # Close browser
113
+ agent-browser close --all # Close all active sessions
114
+
115
+ # Snapshot
116
+ agent-browser snapshot -i # Interactive elements with refs (recommended)
117
+ agent-browser snapshot -s "#selector" # Scope to CSS selector
118
+
119
+ # Interaction (use @refs from snapshot)
120
+ agent-browser click @e1 # Click element
121
+ agent-browser click @e1 --new-tab # Click and open in new tab
122
+ agent-browser fill @e2 "text" # Clear and type text
123
+ agent-browser type @e2 "text" # Type without clearing
124
+ agent-browser select @e1 "option" # Select dropdown option
125
+ agent-browser check @e1 # Check checkbox
126
+ agent-browser press Enter # Press key
127
+ agent-browser keyboard type "text" # Type at current focus (no selector)
128
+ agent-browser keyboard inserttext "text" # Insert without key events
129
+ agent-browser scroll down 500 # Scroll page
130
+ agent-browser scroll down 500 --selector "div.content" # Scroll within a specific container
131
+
132
+ # Get information
133
+ agent-browser get text @e1 # Get element text
134
+ agent-browser get url # Get current URL
135
+ agent-browser get title # Get page title
136
+ agent-browser get cdp-url # Get CDP WebSocket URL
137
+
138
+ # Wait
139
+ agent-browser wait @e1 # Wait for element
140
+ agent-browser wait --load networkidle # Wait for network idle
141
+ agent-browser wait --url "**/page" # Wait for URL pattern
142
+ agent-browser wait 2000 # Wait milliseconds
143
+ agent-browser wait --text "Welcome" # Wait for text to appear (substring match)
144
+ agent-browser wait --fn "!document.body.innerText.includes('Loading...')" # Wait for text to disappear
145
+ agent-browser wait "#spinner" --state hidden # Wait for element to disappear
146
+
147
+ # Downloads
148
+ agent-browser download @e1 ./file.pdf # Click element to trigger download
149
+ agent-browser wait --download ./output.zip # Wait for any download to complete
150
+ agent-browser --download-path ./downloads open <url> # Set default download directory
151
+
152
+ # Network
153
+ agent-browser network requests # Inspect tracked requests
154
+ agent-browser network requests --type xhr,fetch # Filter by resource type
155
+ agent-browser network requests --method POST # Filter by HTTP method
156
+ agent-browser network requests --status 2xx # Filter by status (200, 2xx, 400-499)
157
+ agent-browser network request <requestId> # View full request/response detail
158
+ agent-browser network route "**/api/*" --abort # Block matching requests
159
+ agent-browser network har start # Start HAR recording
160
+ agent-browser network har stop ./capture.har # Stop and save HAR file
161
+
162
+ # Viewport & Device Emulation
163
+ agent-browser set viewport 1920 1080 # Set viewport size (default: 1280x720)
164
+ agent-browser set viewport 1920 1080 2 # 2x retina (same CSS size, higher res screenshots)
165
+ agent-browser set device "iPhone 14" # Emulate device (viewport + user agent)
166
+
167
+ # Capture
168
+ agent-browser screenshot # Screenshot to temp dir
169
+ agent-browser screenshot --full # Full page screenshot
170
+ agent-browser screenshot --annotate # Annotated screenshot with numbered element labels
171
+ agent-browser screenshot --screenshot-dir ./shots # Save to custom directory
172
+ agent-browser screenshot --screenshot-format jpeg --screenshot-quality 80
173
+ agent-browser pdf output.pdf # Save as PDF
174
+
175
+ # Live preview / streaming
176
+ agent-browser stream enable # Start runtime WebSocket streaming on an auto-selected port
177
+ agent-browser stream enable --port 9223 # Bind a specific localhost port
178
+ agent-browser stream status # Inspect enabled state, port, connection, and screencasting
179
+ agent-browser stream disable # Stop runtime streaming and remove the .stream metadata file
180
+
181
+ # Clipboard
182
+ agent-browser clipboard read # Read text from clipboard
183
+ agent-browser clipboard write "Hello, World!" # Write text to clipboard
184
+ agent-browser clipboard copy # Copy current selection
185
+ agent-browser clipboard paste # Paste from clipboard
186
+
187
+ # Dialogs (alert, confirm, prompt, beforeunload)
188
+ # By default, alert and beforeunload dialogs are auto-accepted so they never block the agent.
189
+ # confirm and prompt dialogs still require explicit handling.
190
+ # Use --no-auto-dialog (or AGENT_BROWSER_NO_AUTO_DIALOG=1) to disable automatic handling.
191
+ agent-browser dialog accept # Accept dialog
192
+ agent-browser dialog accept "my input" # Accept prompt dialog with text
193
+ agent-browser dialog dismiss # Dismiss/cancel dialog
194
+ agent-browser dialog status # Check if a dialog is currently open
195
+
196
+ # Diff (compare page states)
197
+ agent-browser diff snapshot # Compare current vs last snapshot
198
+ agent-browser diff snapshot --baseline before.txt # Compare current vs saved file
199
+ agent-browser diff screenshot --baseline before.png # Visual pixel diff
200
+ agent-browser diff url <url1> <url2> # Compare two pages
201
+ agent-browser diff url <url1> <url2> --wait-until networkidle # Custom wait strategy
202
+ agent-browser diff url <url1> <url2> --selector "#main" # Scope to element
203
+ ```
204
+
205
+ ## Streaming
206
+
207
+ Every session automatically starts a WebSocket stream server on an OS-assigned port. Use `agent-browser stream status` to see the bound port and connection state. Use `stream disable` to tear it down, and `stream enable --port <port>` to re-enable on a specific port.
208
+
209
+ ## Batch Execution
210
+
211
+ Execute multiple commands in a single invocation by piping a JSON array of string arrays to `batch`. This avoids per-command process startup overhead when running multi-step workflows.
212
+
213
+ ```bash
214
+ echo '[
215
+ ["open", "https://example.com"],
216
+ ["snapshot", "-i"],
217
+ ["click", "@e1"],
218
+ ["screenshot", "result.png"]
219
+ ]' | agent-browser batch --json
220
+
221
+ # Stop on first error
222
+ agent-browser batch --bail < commands.json
223
+ ```
224
+
225
+ Use `batch` when you have a known sequence of commands that don't depend on intermediate output. Use separate commands or `&&` chaining when you need to parse output between steps (e.g., snapshot to discover refs, then interact).
226
+
227
+ ## Common Patterns
228
+
229
+ ### Form Submission
230
+
231
+ ```bash
232
+ agent-browser open https://example.com/signup
233
+ agent-browser snapshot -i
234
+ agent-browser fill @e1 "Jane Doe"
235
+ agent-browser fill @e2 "jane@example.com"
236
+ agent-browser select @e3 "California"
237
+ agent-browser check @e4
238
+ agent-browser click @e5
239
+ agent-browser wait --load networkidle
240
+ ```
241
+
242
+ ### Authentication with Auth Vault (Recommended)
243
+
244
+ ```bash
245
+ # Save credentials once (encrypted with AGENT_BROWSER_ENCRYPTION_KEY)
246
+ # Recommended: pipe password via stdin to avoid shell history exposure
247
+ echo "pass" | agent-browser auth save github --url https://github.com/login --username user --password-stdin
248
+
249
+ # Login using saved profile (LLM never sees password)
250
+ agent-browser auth login github
251
+
252
+ # List/show/delete profiles
253
+ agent-browser auth list
254
+ agent-browser auth show github
255
+ agent-browser auth delete github
256
+ ```
257
+
258
+ `auth login` waits for username/password/submit selectors before interacting, with a timeout tied to the default action timeout.
259
+
260
+ ### Authentication with State Persistence
261
+
262
+ ```bash
263
+ # Login once and save state
264
+ agent-browser open https://app.example.com/login
265
+ agent-browser snapshot -i
266
+ agent-browser fill @e1 "$USERNAME"
267
+ agent-browser fill @e2 "$PASSWORD"
268
+ agent-browser click @e3
269
+ agent-browser wait --url "**/dashboard"
270
+ agent-browser state save auth.json
271
+
272
+ # Reuse in future sessions
273
+ agent-browser state load auth.json
274
+ agent-browser open https://app.example.com/dashboard
275
+ ```
276
+
277
+ ### Session Persistence
278
+
279
+ ```bash
280
+ # Auto-save/restore cookies and localStorage across browser restarts
281
+ agent-browser --session-name myapp open https://app.example.com/login
282
+ # ... login flow ...
283
+ agent-browser close # State auto-saved to ~/.agent-browser/sessions/
284
+
285
+ # Next time, state is auto-loaded
286
+ agent-browser --session-name myapp open https://app.example.com/dashboard
287
+
288
+ # Encrypt state at rest
289
+ export AGENT_BROWSER_ENCRYPTION_KEY=$(openssl rand -hex 32)
290
+ agent-browser --session-name secure open https://app.example.com
291
+
292
+ # Manage saved states
293
+ agent-browser state list
294
+ agent-browser state show myapp-default.json
295
+ agent-browser state clear myapp
296
+ agent-browser state clean --older-than 7
297
+ ```
298
+
299
+ ### Working with Iframes
300
+
301
+ Iframe content is automatically inlined in snapshots. Refs inside iframes carry frame context, so you can interact with them directly.
302
+
303
+ ```bash
304
+ agent-browser open https://example.com/checkout
305
+ agent-browser snapshot -i
306
+ # @e1 [heading] "Checkout"
307
+ # @e2 [Iframe] "payment-frame"
308
+ # @e3 [input] "Card number"
309
+ # @e4 [input] "Expiry"
310
+ # @e5 [button] "Pay"
311
+
312
+ # Interact directly — no frame switch needed
313
+ agent-browser fill @e3 "4111111111111111"
314
+ agent-browser fill @e4 "12/28"
315
+ agent-browser click @e5
316
+
317
+ # To scope a snapshot to one iframe:
318
+ agent-browser frame @e2
319
+ agent-browser snapshot -i # Only iframe content
320
+ agent-browser frame main # Return to main frame
321
+ ```
322
+
323
+ ### Data Extraction
324
+
325
+ ```bash
326
+ agent-browser open https://example.com/products
327
+ agent-browser snapshot -i
328
+ agent-browser get text @e5 # Get specific element text
329
+ agent-browser get text body > page.txt # Get all page text
330
+
331
+ # JSON output for parsing
332
+ agent-browser snapshot -i --json
333
+ agent-browser get text @e1 --json
334
+ ```
335
+
336
+ ### Parallel Sessions
337
+
338
+ ```bash
339
+ agent-browser --session site1 open https://site-a.com
340
+ agent-browser --session site2 open https://site-b.com
341
+
342
+ agent-browser --session site1 snapshot -i
343
+ agent-browser --session site2 snapshot -i
344
+
345
+ agent-browser session list
346
+ ```
347
+
348
+ ### Connect to Existing Chrome
349
+
350
+ ```bash
351
+ # Auto-discover running Chrome with remote debugging enabled
352
+ agent-browser --auto-connect open https://example.com
353
+ agent-browser --auto-connect snapshot
354
+
355
+ # Or with explicit CDP port
356
+ agent-browser --cdp 9222 snapshot
357
+ ```
358
+
359
+ Auto-connect discovers Chrome via `DevToolsActivePort`, common debugging ports (9222, 9229), and falls back to a direct WebSocket connection if HTTP-based CDP discovery fails.
360
+
361
+ ### Color Scheme (Dark Mode)
362
+
363
+ ```bash
364
+ # Persistent dark mode via flag (applies to all pages and new tabs)
365
+ agent-browser --color-scheme dark open https://example.com
366
+
367
+ # Or via environment variable
368
+ AGENT_BROWSER_COLOR_SCHEME=dark agent-browser open https://example.com
369
+
370
+ # Or set during session (persists for subsequent commands)
371
+ agent-browser set media dark
372
+ ```
373
+
374
+ ### Viewport & Responsive Testing
375
+
376
+ ```bash
377
+ # Set a custom viewport size (default is 1280x720)
378
+ agent-browser set viewport 1920 1080
379
+ agent-browser screenshot desktop.png
380
+
381
+ # Test mobile-width layout
382
+ agent-browser set viewport 375 812
383
+ agent-browser screenshot mobile.png
384
+
385
+ # Retina/HiDPI: same CSS layout at 2x pixel density
386
+ # Screenshots stay at logical viewport size, but content renders at higher DPI
387
+ agent-browser set viewport 1920 1080 2
388
+ agent-browser screenshot retina.png
389
+
390
+ # Device emulation (sets viewport + user agent in one step)
391
+ agent-browser set device "iPhone 14"
392
+ agent-browser screenshot device.png
393
+ ```
394
+
395
+ The `scale` parameter (3rd argument) sets `window.devicePixelRatio` without changing CSS layout. Use it when testing retina rendering or capturing higher-resolution screenshots.
396
+
397
+ ### Visual Browser (Debugging)
398
+
399
+ ```bash
400
+ agent-browser --headed open https://example.com
401
+ agent-browser highlight @e1 # Highlight element
402
+ agent-browser inspect # Open Chrome DevTools for the active page
403
+ agent-browser record start demo.webm # Record session
404
+ agent-browser profiler start # Start Chrome DevTools profiling
405
+ agent-browser profiler stop trace.json # Stop and save profile (path optional)
406
+ ```
407
+
408
+ Use `AGENT_BROWSER_HEADED=1` to enable headed mode via environment variable. Browser extensions work in both headed and headless mode.
409
+
410
+ ### Local Files (PDFs, HTML)
411
+
412
+ ```bash
413
+ # Open local files with file:// URLs
414
+ agent-browser --allow-file-access open file:///path/to/document.pdf
415
+ agent-browser --allow-file-access open file:///path/to/page.html
416
+ agent-browser screenshot output.png
417
+ ```
418
+
419
+ ### iOS Simulator (Mobile Safari)
420
+
421
+ ```bash
422
+ # List available iOS simulators
423
+ agent-browser device list
424
+
425
+ # Launch Safari on a specific device
426
+ agent-browser -p ios --device "iPhone 16 Pro" open https://example.com
427
+
428
+ # Same workflow as desktop - snapshot, interact, re-snapshot
429
+ agent-browser -p ios snapshot -i
430
+ agent-browser -p ios tap @e1 # Tap (alias for click)
431
+ agent-browser -p ios fill @e2 "text"
432
+ agent-browser -p ios swipe up # Mobile-specific gesture
433
+
434
+ # Take screenshot
435
+ agent-browser -p ios screenshot mobile.png
436
+
437
+ # Close session (shuts down simulator)
438
+ agent-browser -p ios close
439
+ ```
440
+
441
+ **Requirements:** macOS with Xcode, Appium (`npm install -g appium && appium driver install xcuitest`)
442
+
443
+ **Real devices:** Works with physical iOS devices if pre-configured. Use `--device "<UDID>"` where UDID is from `xcrun xctrace list devices`.
444
+
445
+ ## Security
446
+
447
+ All security features are opt-in. By default, agent-browser imposes no restrictions on navigation, actions, or output.
448
+
449
+ ### Content Boundaries (Recommended for AI Agents)
450
+
451
+ Enable `--content-boundaries` to wrap page-sourced output in markers that help LLMs distinguish tool output from untrusted page content:
452
+
453
+ ```bash
454
+ export AGENT_BROWSER_CONTENT_BOUNDARIES=1
455
+ agent-browser snapshot
456
+ # Output:
457
+ # --- AGENT_BROWSER_PAGE_CONTENT nonce=<hex> origin=https://example.com ---
458
+ # [accessibility tree]
459
+ # --- END_AGENT_BROWSER_PAGE_CONTENT nonce=<hex> ---
460
+ ```
461
+
462
+ ### Domain Allowlist
463
+
464
+ Restrict navigation to trusted domains. Wildcards like `*.example.com` also match the bare domain `example.com`. Sub-resource requests, WebSocket, and EventSource connections to non-allowed domains are also blocked. Include CDN domains your target pages depend on:
465
+
466
+ ```bash
467
+ export AGENT_BROWSER_ALLOWED_DOMAINS="example.com,*.example.com"
468
+ agent-browser open https://example.com # OK
469
+ agent-browser open https://malicious.com # Blocked
470
+ ```
471
+
472
+ ### Action Policy
473
+
474
+ Use a policy file to gate destructive actions:
475
+
476
+ ```bash
477
+ export AGENT_BROWSER_ACTION_POLICY=./policy.json
478
+ ```
479
+
480
+ Example `policy.json`:
481
+
482
+ ```json
483
+ { "default": "deny", "allow": ["navigate", "snapshot", "click", "scroll", "wait", "get"] }
484
+ ```
485
+
486
+ Auth vault operations (`auth login`, etc.) bypass action policy but domain allowlist still applies.
487
+
488
+ ### Output Limits
489
+
490
+ Prevent context flooding from large pages:
491
+
492
+ ```bash
493
+ export AGENT_BROWSER_MAX_OUTPUT=50000
494
+ ```
495
+
496
+ ## Diffing (Verifying Changes)
497
+
498
+ Use `diff snapshot` after performing an action to verify it had the intended effect. This compares the current accessibility tree against the last snapshot taken in the session.
499
+
500
+ ```bash
501
+ # Typical workflow: snapshot -> action -> diff
502
+ agent-browser snapshot -i # Take baseline snapshot
503
+ agent-browser click @e2 # Perform action
504
+ agent-browser diff snapshot # See what changed (auto-compares to last snapshot)
505
+ ```
506
+
507
+ For visual regression testing or monitoring:
508
+
509
+ ```bash
510
+ # Save a baseline screenshot, then compare later
511
+ agent-browser screenshot baseline.png
512
+ # ... time passes or changes are made ...
513
+ agent-browser diff screenshot --baseline baseline.png
514
+
515
+ # Compare staging vs production
516
+ agent-browser diff url https://staging.example.com https://prod.example.com --screenshot
517
+ ```
518
+
519
+ `diff snapshot` output uses `+` for additions and `-` for removals, similar to git diff. `diff screenshot` produces a diff image with changed pixels highlighted in red, plus a mismatch percentage.
520
+
521
+ ## Timeouts and Slow Pages
522
+
523
+ The default timeout is 25 seconds. This can be overridden with the `AGENT_BROWSER_DEFAULT_TIMEOUT` environment variable (value in milliseconds). For slow websites or large pages, use explicit waits instead of relying on the default timeout:
524
+
525
+ ```bash
526
+ # Wait for network activity to settle (best for slow pages)
527
+ agent-browser wait --load networkidle
528
+
529
+ # Wait for a specific element to appear
530
+ agent-browser wait "#content"
531
+ agent-browser wait @e1
532
+
533
+ # Wait for a specific URL pattern (useful after redirects)
534
+ agent-browser wait --url "**/dashboard"
535
+
536
+ # Wait for a JavaScript condition
537
+ agent-browser wait --fn "document.readyState === 'complete'"
538
+
539
+ # Wait a fixed duration (milliseconds) as a last resort
540
+ agent-browser wait 5000
541
+ ```
542
+
543
+ When dealing with consistently slow websites, use `wait --load networkidle` after `open` to ensure the page is fully loaded before taking a snapshot. If a specific element is slow to render, wait for it directly with `wait <selector>` or `wait @ref`.
544
+
545
+ ## JavaScript Dialogs (alert / confirm / prompt)
546
+
547
+ When a page opens a JavaScript dialog (`alert()`, `confirm()`, or `prompt()`), it blocks all other browser commands (snapshot, screenshot, click, etc.) until the dialog is dismissed. If commands start timing out unexpectedly, check for a pending dialog:
548
+
549
+ ```bash
550
+ # Check if a dialog is blocking
551
+ agent-browser dialog status
552
+
553
+ # Accept the dialog (dismiss the alert / click OK)
554
+ agent-browser dialog accept
555
+
556
+ # Accept a prompt dialog with input text
557
+ agent-browser dialog accept "my input"
558
+
559
+ # Dismiss the dialog (click Cancel)
560
+ agent-browser dialog dismiss
561
+ ```
562
+
563
+ When a dialog is pending, all command responses include a `warning` field indicating the dialog type and message. In `--json` mode this appears as a `"warning"` key in the response object.
564
+
565
+ ## Session Management and Cleanup
566
+
567
+ When running multiple agents or automations concurrently, always use named sessions to avoid conflicts:
568
+
569
+ ```bash
570
+ # Each agent gets its own isolated session
571
+ agent-browser --session agent1 open site-a.com
572
+ agent-browser --session agent2 open site-b.com
573
+
574
+ # Check active sessions
575
+ agent-browser session list
576
+ ```
577
+
578
+ Always close your browser session when done to avoid leaked processes:
579
+
580
+ ```bash
581
+ agent-browser close # Close default session
582
+ agent-browser --session agent1 close # Close specific session
583
+ agent-browser close --all # Close all active sessions
584
+ ```
585
+
586
+ If a previous session was not closed properly, the daemon may still be running. Use `agent-browser close` to clean it up, or `agent-browser close --all` to shut down every session at once.
587
+
588
+ To auto-shutdown the daemon after a period of inactivity (useful for ephemeral/CI environments):
589
+
590
+ ```bash
591
+ AGENT_BROWSER_IDLE_TIMEOUT_MS=60000 agent-browser open example.com
592
+ ```
593
+
594
+ ## Ref Lifecycle (Important)
595
+
596
+ Refs (`@e1`, `@e2`, etc.) are invalidated when the page changes. Always re-snapshot after:
597
+
598
+ - Clicking links or buttons that navigate
599
+ - Form submissions
600
+ - Dynamic content loading (dropdowns, modals)
601
+
602
+ ```bash
603
+ agent-browser click @e5 # Navigates to new page
604
+ agent-browser snapshot -i # MUST re-snapshot
605
+ agent-browser click @e1 # Use new refs
606
+ ```
607
+
608
+ ## Annotated Screenshots (Vision Mode)
609
+
610
+ Use `--annotate` to take a screenshot with numbered labels overlaid on interactive elements. Each label `[N]` maps to ref `@eN`. This also caches refs, so you can interact with elements immediately without a separate snapshot.
611
+
612
+ ```bash
613
+ agent-browser screenshot --annotate
614
+ # Output includes the image path and a legend:
615
+ # [1] @e1 button "Submit"
616
+ # [2] @e2 link "Home"
617
+ # [3] @e3 textbox "Email"
618
+ agent-browser click @e2 # Click using ref from annotated screenshot
619
+ ```
620
+
621
+ Use annotated screenshots when:
622
+
623
+ - The page has unlabeled icon buttons or visual-only elements
624
+ - You need to verify visual layout or styling
625
+ - Canvas or chart elements are present (invisible to text snapshots)
626
+ - You need spatial reasoning about element positions
627
+
628
+ ## Semantic Locators (Alternative to Refs)
629
+
630
+ When refs are unavailable or unreliable, use semantic locators:
631
+
632
+ ```bash
633
+ agent-browser find text "Sign In" click
634
+ agent-browser find label "Email" fill "user@test.com"
635
+ agent-browser find role button click --name "Submit"
636
+ agent-browser find placeholder "Search" type "query"
637
+ agent-browser find testid "submit-btn" click
638
+ ```
639
+
640
+ ## JavaScript Evaluation (eval)
641
+
642
+ Use `eval` to run JavaScript in the browser context. **Shell quoting can corrupt complex expressions** -- use `--stdin` or `-b` to avoid issues.
643
+
644
+ ```bash
645
+ # Simple expressions work with regular quoting
646
+ agent-browser eval 'document.title'
647
+ agent-browser eval 'document.querySelectorAll("img").length'
648
+
649
+ # Complex JS: use --stdin with heredoc (RECOMMENDED)
650
+ agent-browser eval --stdin <<'EVALEOF'
651
+ JSON.stringify(
652
+ Array.from(document.querySelectorAll("img"))
653
+ .filter(i => !i.alt)
654
+ .map(i => ({ src: i.src.split("/").pop(), width: i.width }))
655
+ )
656
+ EVALEOF
657
+
658
+ # Alternative: base64 encoding (avoids all shell escaping issues)
659
+ agent-browser eval -b "$(echo -n 'Array.from(document.querySelectorAll("a")).map(a => a.href)' | base64)"
660
+ ```
661
+
662
+ **Why this matters:** When the shell processes your command, inner double quotes, `!` characters (history expansion), backticks, and `$()` can all corrupt the JavaScript before it reaches agent-browser. The `--stdin` and `-b` flags bypass shell interpretation entirely.
663
+
664
+ **Rules of thumb:**
665
+
666
+ - Single-line, no nested quotes -> regular `eval 'expression'` with single quotes is fine
667
+ - Nested quotes, arrow functions, template literals, or multiline -> use `eval --stdin <<'EVALEOF'`
668
+ - Programmatic/generated scripts -> use `eval -b` with base64
669
+
670
+ ## Configuration File
671
+
672
+ Create `agent-browser.json` in the project root for persistent settings:
673
+
674
+ ```json
675
+ {
676
+ "headed": true,
677
+ "proxy": "http://localhost:8080",
678
+ "profile": "./browser-data"
679
+ }
680
+ ```
681
+
682
+ Priority (lowest to highest): `~/.agent-browser/config.json` < `./agent-browser.json` < env vars < CLI flags. Use `--config <path>` or `AGENT_BROWSER_CONFIG` env var for a custom config file (exits with error if missing/invalid). All CLI options map to camelCase keys (e.g., `--executable-path` -> `"executablePath"`). Boolean flags accept `true`/`false` values (e.g., `--headed false` overrides config). Extensions from user and project configs are merged, not replaced.
683
+
684
+ ## Deep-Dive Documentation
685
+
686
+ | Reference | When to Use |
687
+ | -------------------------------------------------------------------- | --------------------------------------------------------- |
688
+ | [references/commands.md](references/commands.md) | Full command reference with all options |
689
+ | [references/snapshot-refs.md](references/snapshot-refs.md) | Ref lifecycle, invalidation rules, troubleshooting |
690
+ | [references/session-management.md](references/session-management.md) | Parallel sessions, state persistence, concurrent scraping |
691
+ | [references/authentication.md](references/authentication.md) | Login flows, OAuth, 2FA handling, state reuse |
692
+ | [references/video-recording.md](references/video-recording.md) | Recording workflows for debugging and documentation |
693
+ | [references/profiling.md](references/profiling.md) | Chrome DevTools profiling for performance analysis |
694
+ | [references/proxy-support.md](references/proxy-support.md) | Proxy configuration, geo-testing, rotating proxies |
695
+
696
+ ## Browser Engine Selection
697
+
698
+ Use `--engine` to choose a local browser engine. The default is `chrome`.
699
+
700
+ ```bash
701
+ # Use Lightpanda (fast headless browser, requires separate install)
702
+ agent-browser --engine lightpanda open example.com
703
+
704
+ # Via environment variable
705
+ export AGENT_BROWSER_ENGINE=lightpanda
706
+ agent-browser open example.com
707
+
708
+ # With custom binary path
709
+ agent-browser --engine lightpanda --executable-path /path/to/lightpanda open example.com
710
+ ```
711
+
712
+ Supported engines:
713
+ - `chrome` (default) -- Chrome/Chromium via CDP
714
+ - `lightpanda` -- Lightpanda headless browser via CDP (10x faster, 10x less memory than Chrome)
715
+
716
+ Lightpanda does not support `--extension`, `--profile`, `--state`, or `--allow-file-access`. Install Lightpanda from https://lightpanda.io/docs/open-source/installation.
717
+
718
+ ## Observability Dashboard
719
+
720
+ The dashboard is a standalone background server that shows live browser viewports, command activity, and console output for all sessions.
721
+
722
+ ```bash
723
+ # Install the dashboard once
724
+ agent-browser dashboard install
725
+
726
+ # Start the dashboard server (background, port 4848)
727
+ agent-browser dashboard start
728
+
729
+ # All sessions are automatically visible in the dashboard
730
+ agent-browser open example.com
731
+
732
+ # Stop the dashboard
733
+ agent-browser dashboard stop
734
+ ```
735
+
736
+ The dashboard runs independently of browser sessions on port 4848 (configurable with `--port`). All sessions automatically stream to the dashboard. Sessions can also be created from the dashboard UI with local engines or cloud providers.
737
+
738
+ ## Ready-to-Use Templates
739
+
740
+ | Template | Description |
741
+ | ------------------------------------------------------------------------ | ----------------------------------- |
742
+ | [templates/form-automation.sh](templates/form-automation.sh) | Form filling with validation |
743
+ | [templates/authenticated-session.sh](templates/authenticated-session.sh) | Login once, reuse state |
744
+ | [templates/capture-workflow.sh](templates/capture-workflow.sh) | Content extraction with screenshots |
745
+
746
+ ```bash
747
+ ./templates/form-automation.sh https://example.com/form
748
+ ./templates/authenticated-session.sh https://app.example.com/login
749
+ ./templates/capture-workflow.sh https://example.com ./output
750
+ ```