scenescout 3.23.0 โ†’ 3.23.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -2,7 +2,7 @@
2
2
 
3
3
  # ๐Ÿ”ญ SceneScout
4
4
 
5
- **Exploratory UI testing, driven by the AI agent you already use.**
5
+ **Your coding agent, turned into an exploratory QA tester for any web app.**
6
6
 
7
7
  Works with Claude Code ยท Cursor ยท VS Code (Copilot) ยท Codex CLI ยท Gemini CLI ยท Copilot CLI ยท Windsurf ยท any [MCP](https://modelcontextprotocol.io) client
8
8
 
@@ -10,841 +10,248 @@ Works with Claude Code ยท Cursor ยท VS Code (Copilot) ยท Codex CLI ยท Gemini CLI
10
10
  [![npm](https://img.shields.io/npm/v/scenescout.svg)](https://www.npmjs.com/package/scenescout)
11
11
  [![license: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)
12
12
  ![node >= 20](https://img.shields.io/badge/node-%E2%89%A5%2020-339933?logo=node.js&logoColor=white)
13
- ![MCP server](https://img.shields.io/badge/MCP-server-8A2BE2)
14
13
 
15
- [๐Ÿ“– Guide](docs/guide/Home.md) ยท [๐Ÿ‘€ See it work](#-see-it-work) ยท [โœจ Why](#-why-its-different) ยท [๐ŸŽฏ Two ways to use it](#-two-ways-to-use-it) ยท [๐Ÿš€ Quickstart](#-quickstart) ยท [๐Ÿงฐ Toolbox](#-the-toolbox) ยท [๐Ÿ”Œ Other clients](#-other-mcp-clients) ยท [๐Ÿ”’ Safety](#-safety-model) ยท [๐Ÿฉบ Troubleshooting](#-troubleshooting)
14
+ [๐Ÿ“– Guide](docs/guide/Home.md) ยท [๐Ÿ› What it catches](#-what-it-catches) ยท [๐ŸŽฌ QA with evidence](#-automated-qa-with-evidence) ยท [๐Ÿš€ Get started](#-get-started) ยท [๐Ÿšฆ CI](#-in-ci) ยท [๐Ÿ”’ Safety](#-safe-by-default) ยท [๐Ÿ“š Docs](#-documentation)
16
15
 
17
16
  </div>
18
17
 
19
- SceneScout is an [MCP](https://modelcontextprotocol.io) server that hands an agent a *structured view* of a running web app โ€” every element, its geometry, and a set of always-on correctness oracles โ€” and lets the agent explore it like a curious user. Your coding agent is the brain; SceneScout is the hands, eyes, and memory. Any MCP client can drive it, and the testing method comes with the server, so the agent knows how to use the tools wherever it runs.
18
+ Scripted end-to-end tests answer one question: *does this exact flow still work?* They say nothing about the rest of the app. SceneScout lets the agent you already use explore a running web app like a curious, thorough tester. It clicks, fills forms, switches roles and calls the API behind a hidden button, then writes a report of what is **broken** and what could be **better**, with a picture and evidence for every line.
20
19
 
20
+ Try it on any app you are allowed to test. No account, no API key, no setup:
21
+
22
+ ```bash
23
+ npx -y scenescout http://localhost:3000
21
24
  ```
22
- โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ” MCP (stdio) โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
23
- โ”‚ Your coding agent โ”‚ โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ถ โ”‚ SceneScout engine โ”‚
24
- โ”‚ (intent, judgment, โ”‚ โ—€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ โ”‚ Playwright ยท oracles ยท memory โ”‚
25
- โ”‚ your subscription) โ”‚ tool results โ”‚ findings ยท report โ€” no LLM โ”‚
26
- โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
25
+
26
+ ```mermaid
27
+ flowchart LR
28
+ A["๐Ÿง  Your coding agent<br/>Claude Code, Cursor, Copilotโ€ฆ"] -- MCP --> S["๐Ÿ”ญ SceneScout<br/>browser ยท checks ยท memory"]
29
+ S --> App["๐ŸŒ Your running app"]
30
+ S --> R["๐Ÿ“‹ Report<br/>plain words + pictures"]
31
+ S --> E["๐ŸŽฌ Evidence<br/>replay, videos, test report"]
27
32
  ```
28
33
 
29
- Scripted E2E suites answer one question โ€” *"does this exact flow still work?"* โ€” and say nothing about the 95% of the app they don't touch. SceneScout covers both gaps: it finds what's **broken** (crashes, dead ends, permission leaks) *and* reports how the product could be **better** (confusing flows, weak hierarchy, design-system drift), with concrete measurements.
34
+ ## ๐Ÿ‘ฅ Who it's for
30
35
 
36
+ | ๐Ÿ› ๏ธ **Developers** | ๐Ÿง‘โ€๐Ÿ’ผ **QA, product and non-technical teams** |
37
+ |---|---|
38
+ | Findings with the request that failed (`GET /api/orders?status=archived โ†’ 500`), the steps, and a Playwright regression-test skeleton | Each problem in plain words first: what was done, what was expected, what happened, and a picture of the page |
39
+ | Next to the source, the file behind the bug and a likely fix | Visible problems shown, not described: a covered button, a broken image, text too faint to read |
40
+ | A deterministic gate for pull requests, with SARIF for code scanning | Journeys recorded step by step, as frames and video, so a pass is something you can watch |
41
+ | Runs from your editor's agent, or unattended in CI | A test report laid out by **your** template, with expected and actual results, deviations and blank sign-off rows |
31
42
 
32
- ## ๐Ÿ‘€ See it work
43
+ ## ๐Ÿ› What it catches
33
44
 
34
- This is a real run against the small demo app bundled in this repository. The app has bugs planted in it on purpose, and two of them are visible on its dashboard:
45
+ A real run against the small demo app in this repository, which has bugs planted on purpose. The two you can see are outlined:
35
46
 
36
47
  <p align="center"><img src="examples/screenshots/dashboard-annotated.png" alt="The demo app's dashboard with two defects outlined in red: 1, a yellow badge covering the All orders button; 2, the weekly chart image failing to load" width="760" /></p>
37
48
 
38
- **The broken chart is the demo app's bug, not this page's** โ€” it is one of the twelve findings SceneScout filed, next to the badge sitting on a button. The red callouts were added for this README; the [unmarked screenshots](examples/screenshots/) are the ones the engine took.
39
-
40
- An excerpt of the report it wrote โ€” [read the whole thing](examples/report.md):
49
+ It filed twelve findings. A few from [the report](examples/report.md):
41
50
 
42
51
  > **๐Ÿ”ด [HIGH] A double-click on Create order creates two orders**
43
- > Evidence: `2ร— click fired the same state-changing request 2ร— (POST /api/orders)`
44
- > The submit button stays enabled while the request is in flight, and the endpoint accepts the repeat.
45
- >
46
- > **๐Ÿ”ด [HIGH] Filtering orders by Archived fails, and the page shows an empty table instead of an error**
47
- > Evidence: `GET /api/orders?status=archived โ†’ HTTP 500`
52
+ > `2ร— click fired the same state-changing request 2ร— (POST /api/orders)`
48
53
  >
49
54
  > **๐Ÿ”ด [HIGH] A clerk can approve an order by calling the endpoint the page hides from them**
50
- > Evidence: `POST /api/orders/1037/approve 200 as clerk; POST /api/orders/1038/reject 403 as clerk` โ€” the button was hidden, the server did not agree.
51
- >
52
- > **๐ŸŸ  [MEDIUM] The "New: bulk import" badge sits on top of the All orders button** *(callout 1)*
53
- > Evidence: `"All orders" overlaps "New: bulk import" (81%)` โ€” measured from layout boxes, no screenshot needed.
55
+ > `POST /api/orders/1037/approve 200 as clerk`: the button was hidden, the server did not agree.
54
56
  >
55
- > **๐ŸŸก [LOW] The dashboard chart image is missing** *(callout 2)*
56
- > Evidence: `GET /img/weekly-chart.png โ†’ HTTP 404`
57
+ > **๐Ÿ”ด [HIGH] Filtering orders by Archived fails, and the page shows an empty table instead of an error**
58
+ > `GET /api/orders?status=archived โ†’ HTTP 500`
57
59
  >
58
- > **Gap ledger โ€” what was NOT tested:** 9 of 12 known routes visited this run and never design-audited ยท single-role run, so permission boundaries are untested
59
-
60
- Every finding comes with a repro trace and a Playwright regression-test skeleton. To try it yourself, clone this repository, run `npm run demo:serve`, then `/scenescout --url http://127.0.0.1:4173` โ€” see [demo-app/](demo-app/). Its README lists every seeded defect and which oracle catches it.
61
-
62
- ---
63
-
64
- ## โœจ Why it's different
65
-
66
- - ๐Ÿง  **Your agent is the brain โ€” no API key.** The engine contains no LLM. Exploration runs on the agent and subscription you already have (Claude Code, Cursor, Copilot, Codex, Gemini CLI and others); SceneScout just gives it deterministic tools and the method for using them.
67
- - ๐Ÿ“ **Structured scene, not pixels.** The agent reads element lists *with layout geometry*, not screenshots. Overlap and off-screen bugs are computed from boxes โ€” deterministic, no vision guessing. Images that failed to load are read from the DOM too. (Screenshots exist only for pixel-native residue like a canvas or a rendering glitch.)
68
- - ๐Ÿ›ก๏ธ **Read-only by default, enforced on the wire.** Destructive actions are blocked at the network layer, not by asking the model nicely. Opt into writes only against disposable data.
69
- - โœ… **Completion is a contract, not a vibe.** The engine knows the app's routes and *refuses* to file an "extensive" report while any known route is unvisited, unexercised, or un-audited. "Explored a bit and stopped" is structurally impossible.
70
- - ๐Ÿงญ **It remembers.** UI states are fingerprinted and stored in the project's `.scenescout/`. Run N+1 skips what run N already covered, and every run starts smarter than the last.
71
-
72
- ---
73
-
74
- ## ๐ŸŽฏ Two ways to use it
75
-
76
- SceneScout needs only a URL. Give it the source code as well and it gets noticeably better.
77
-
78
- | | ๐Ÿ  **Next to the codebase** *(recommended)* | ๐ŸŒ **Against a remote URL** |
79
- |---|---|---|
80
- | **You run it from** | the app's repository | any folder โ€” an empty `qa/` directory is fine |
81
- | **It plays the role of** | a developer-tester who can read the code | a black-box QA tester, like a person with a browser |
82
- | **How it finds pages** | ๐Ÿ“‚ reads routes from the source **and** follows links: file-based routing (Next.js, SvelteKit, Nuxt) and router configuration written in code (React Router, Vue Router, Angular). Routes built at runtime are not seen | ๐Ÿ”— follows same-origin links only โ€” pages nothing links to, or on another subdomain, stay unknown |
83
- | **"Did we cover everything?"** | checked against the routes found in source *plus* discovered links โ€” an unvisited one blocks the report | checked against the pages it managed to discover |
84
- | **Setup it figures out** | framework, dev command, saved Playwright logins (`playwright/.auth/`), whether the app uses `data-testid` | none โ€” you pass the URL, and the path to a login state if the app needs one |
85
- | **What a finding looks like** | the symptom, **plus** the file behind it and a suggested fix | the symptom, a repro trace, and a regression-test skeleton |
86
- | **Typical target** | `localhost` while you build | staging, a preview deploy, a client's site |
87
-
88
- **Why the codebase helps.** The agent driving SceneScout is a coding agent, which can already read your repository. With the source at hand it knows the app's static routes before opening the browser, so coverage is measured against the real app instead of whatever happened to be linked. It can also check a suspicion against the code before reporting it: "there is no way to export this table" is a much stronger finding once the agent has confirmed no export handler exists. And when something breaks it can open the component or handler responsible and tell you *where* and *how* to fix it โ€” "the save button does nothing" becomes "`OrderForm` swallows the rejected promise in `onSubmit`; surface the error and re-enable the button".
89
-
90
- **Why it still works without it.** Everything SceneScout *observes* comes from the running page โ€” elements, layout geometry, console and network errors, design-audit scores, task-ease measurements โ€” and none of that needs source code. Point it at a URL you are allowed to test and it behaves like a thorough QA tester: it explores, reproduces, and files findings with evidence.
91
-
92
- ```
93
- # next to the code โ€” run inside the app's repository
94
- /scenescout --url http://localhost:3000
95
-
96
- # remote โ€” run from any folder; memory and the report are kept there
97
- /scenescout --url https://staging.example.com --role ./auth/qa.json
98
- ```
60
+ > **๐ŸŸ  [MEDIUM] The "New: bulk import" badge sits on top of the All orders button** *(callout 1)*
61
+ > `"All orders" overlaps "New: bulk import" (81%)`, measured from layout boxes, no screenshot needed.
99
62
 
100
- > [!IMPORTANT]
101
- > Only test sites you own or are authorized to test. A remote environment is more likely to hold real data, so for a remote URL with no source the skill attaches in **`observe`** mode: nothing but `GET` requests leaves the page. The default **read-only** mode blocks `PUT`/`PATCH`/`DELETE` and destructive-looking requests, but an ordinary form submission (a plain `POST`: contact form, comment, order, signup) still reaches the server and can create a record. Say so when that is acceptable on your target. See the [safety model](#-safety-model).
63
+ What it looks for, on every page, after every action:
102
64
 
103
- ---
65
+ - ๐Ÿ“ **Broken layout, from geometry:** controls that overlap, sit off-screen, hide under a sticky bar or can never be scrolled into view. The CSS bugs a person spots at a glance, found without one.
66
+ - ๐Ÿงจ **Real breakage:** console errors, crashes, failed requests, 4xx and 5xx responses, broken images, dead-end pages.
67
+ - ๐Ÿ”“ **Permission leaks:** it calls the app's own API as each role, so "the button is hidden" becomes "the server refuses it", or doesn't.
68
+ - ๐Ÿคฅ **Pages that lie:** "Saved!" after the server refused the save, or an empty table after the request failed.
69
+ - ๐Ÿ‘† **Impatient users:** a double-click that sends the same order twice.
70
+ - โ™ฟ **Accessibility and craft:** contrast, focus, labels, target sizes, spacing and type, with a 0 to 100 score per page.
71
+ - ๐Ÿงญ **Friction:** how many steps a task takes, and where a user had to go back.
72
+ - ๐Ÿ’‰ **Security smells:** typed markup that comes back as an element, and tokens posted to any window.
104
73
 
105
- ## ๐Ÿš€ Quickstart
74
+ [Everything it checks](docs/guide/What-it-checks.md).
106
75
 
107
- ### โšก A first look, nothing to set up
76
+ ## ๐ŸŽฌ Automated QA with evidence
108
77
 
109
- Node 20 or newer and the address of an app you are allowed to test:
78
+ Save the journeys that matter, the happy paths and the ones that must fail politely, and `scenescout check` replays them on every pull request with no model involved. Recorded, each run leaves proof you can watch and hand to someone who never opens a terminal.
110
79
 
111
- ```bash
112
- npx -y scenescout http://localhost:3000
80
+ ```mermaid
81
+ flowchart LR
82
+ PR["๐Ÿ”€ Pull request"] --> C["๐Ÿšฆ scenescout check<br/>--record --video --template"]
83
+ F["๐Ÿ“ Saved journeys<br/>.scenescout/flows/*.json"] --> C
84
+ C --> G["โœ… / โŒ A gate on the pull request"]
85
+ C --> RP["๐Ÿ–ผ๏ธ replay.html<br/>every step, with frames"]
86
+ C --> V["๐ŸŽž๏ธ A video per journey"]
87
+ C --> TR["๐Ÿ“„ Test report<br/>your template, SHA-256 manifest,<br/>blank sign-off rows"]
113
88
  ```
114
89
 
115
- It needs no model, no API key and no MCP client. It downloads the headless Chromium build if the machine has none (once, about 200 MB) and changes nothing else: no skill, no MCP registration, nothing on your PATH. Then it opens up to 20 pages in `observe` mode, where nothing but reads leaves the page (signing in and refreshing a token apart), starting none after 3 minutes. It measures each one the way [`scenescout check`](docs/guide/Ways-to-use-it.md#scenescout-check-a-gate-in-ci) does, writes `scenescout-report/report.md` in the current folder and prints the three issues to look at first:
90
+ A journey is a few lines of JSON, written by hand or kept from a flow your agent just walked. This is [the demo's happy path](examples/flows/01-place-an-order.json):
116
91
 
92
+ ```json
93
+ {
94
+ "name": "place an order",
95
+ "id": "TC-01",
96
+ "requirements": ["REQ-ORD-1", "REQ-ORD-2"],
97
+ "steps": [
98
+ { "action": "navigate", "target": "/orders-new.html", "expected": "The new-order form is shown" },
99
+ { "action": "type", "target": "testid=new-order-customer", "value": "Harbour Bakery", "expected": "The customer is filled in" },
100
+ { "action": "type", "target": "testid=new-order-items", "value": "4", "replace": true, "expected": "The item count is 4" },
101
+ { "action": "click", "target": "testid=new-order-submit", "expected": "The order is created" },
102
+ { "action": "expect-request", "request": "POST /api/orders", "status": "2xx" },
103
+ { "action": "expect-element", "target": "testid=new-order-created-link", "state": "visible", "expected": "A link to the new order is shown" }
104
+ ]
105
+ }
117
106
  ```
118
- Look at these first:
119
- 1. [medium] Request failed with a client error: GET /img/weekly-chart.png โ†’ HTTP 404 (on /)
120
- 2. [medium] Dead end: /reports-scheduled.html: 0 controls (on /reports-scheduled.html)
121
- 3. [medium] Control covered by pinned chrome: button "Save notes" is COVERED by pinned chrome [order-stickybar] at this scroll position โ€” a click aimed at it lands on that element instead (on /order.html?id=1042)
122
-
123
- 12 pages looked at in 12 s in observe mode: 0 high ยท 6 medium ยท 2 low ยท 5 worth a look, never counted.
124
- Report: scenescout-report/report.md
125
- ```
126
-
127
- That is the [demo app](demo-app/). It exits 0 whatever it finds (a look, not a gate), and 2 when the address cannot be reached or the report cannot be written. After the address, `--max-routes` and `--max-minutes` raise the limits, `--mode read-only` lets a plain POST through, and `--out` names another folder. A `scenescout-report/` holding files a first look did not write is left alone, and no report it did not write is ever replaced. A first look only opens pages. To have your agent click, fill forms, compare roles and remember what it learned, set SceneScout up as below.
128
-
129
- ### ๐Ÿ“ฆ Prerequisites
130
107
 
131
- | | |
132
- |---|---|
133
- | **Node** | โ‰ฅ 20 |
134
- | **An MCP client** | Claude Code, Cursor, VS Code with Copilot, Codex CLI, Gemini CLI, GitHub Copilot CLI, Windsurf, or [any other](#-other-mcp-clients) |
135
- | **A web app to test** | SceneScout tests a *live* app: start yours locally first (e.g. `npm run dev`, `make dev-up`), or have the URL of a deployed one you're allowed to test |
136
-
137
- ### 1๏ธโƒฃ Install
138
-
139
- It is on npm. Nothing to clone:
108
+ Against the [demo app](demo-app/), with the three journeys in [examples/flows](examples/flows/) (one creates an order, so the check is allowed to send it):
140
109
 
141
110
  ```bash
142
- npx -y scenescout install # Claude Code: skill + server + Chromium (one-time download)
143
- npx -y scenescout install --client cursor # or: vscode, codex, gemini, copilot, windsurf (comma-separated for several)
111
+ npx scenescout check http://127.0.0.1:4173 --flows examples/flows --mode read-only --flow-writes allow \
112
+ --record --video --template examples/report-template.json
144
113
  ```
145
114
 
146
- Either way it downloads the browser and registers the server with the client you named. Claude Code also gets the method as a skill; every other client receives the same method from the server. [What each client gets](#-other-mcp-clients).
115
+ **The journey, as it ran** (one of the videos it filmed):
147
116
 
148
- **Prefer a Claude Code plugin?** The skill and the server arrive together:
117
+ <p align="center"><img src="examples/screenshots/journey-place-order.gif" alt="A recorded journey on the demo app: the new-order form is filled with a customer and four items, Create order is clicked, and the confirmation with a link to the new order appears" width="760" /></p>
149
118
 
150
- ```
151
- /plugin marketplace add brunoboto96/SceneScout
152
- /plugin install scenescout@scenescout-marketplace
153
- ```
119
+ **The replay page:** each journey with a pass or fail badge. A journey that broke opens at the step that broke, with the page as it was after every step:
154
120
 
155
- Then start a new chat to use SceneScout; the test browser downloads on first use (to have it ready beforehand: `npx -y scenescout install --browser-only`). The command becomes `/scenescout:scenescout`. A plugin's skill comes from this repository and its server from the latest npm release, so right after a release lands here the two can differ for a short while; `/plugin marketplace update scenescout-marketplace` brings the skill up to date.
121
+ <p align="center"><img src="examples/screenshots/check-replay.png" alt="The check's replay page: verdict FAILED, the app's address, start and end times, 24 frames and 3 videos; two journeys passed and one, see archived orders, failed at step 3" width="820" /></p>
156
122
 
157
- An optional second plugin, `scenescout-mod@scenescout-marketplace`, adds a run pane (`/scenescout-pane`) and a setting for the model lane agents run on, in the Claude Code CLI and the desktop Code tab. It is a mod: unsandboxed JavaScript that runs inside Claude Code, so it is opt-in. [The Claude Code mod](docs/guide/Ways-to-use-it.md#the-claude-code-mod).
123
+ <p align="center"><img src="examples/screenshots/check-replay-failed.png" alt="The failed journey opened: step 1 navigates to the orders page and passes, step 2 selects Archived and passes with an empty table, step 3 expected GET /api/orders to answer 200 and it answered 500, highlighted in red" width="760" /></p>
158
124
 
159
- **Using Claude Desktop?** Install the extension: download `scenescout-X.Y.Z.mcpb` from the [latest release](https://github.com/brunoboto96/SceneScout/releases/latest) and open it (or Settings > Extensions > Advanced settings > Install Extension). It works as soon as it is installed, with no terminal step: the test browser downloads on first use. Start a new chat and ask *"Use SceneScout to test http://localhost:3000"*. [More in the guide](docs/guide/Start-here.md#as-a-claude-desktop-extension).
160
-
161
- **A client that is not in that list?** [Add the server to its config by hand](#-other-mcp-clients); the test browser downloads on first use.
162
-
163
- <details>
164
- <summary>What <code>install</code> actually does</summary>
165
-
166
- 1. puts the `/scenescout` skill into `~/.claude/skills/` (or `$CLAUDE_CONFIG_DIR/skills/`) โ€” a `scenescout` folder it didn't create is moved aside to a `.backup-โ€ฆ` copy, never deleted,
167
- 2. downloads the browser SceneScout drives (skipped if you already have it). By default that is Chromium, as two builds: the full browser for headed runs and the headless shell every other run uses. [Choose something else](#-choosing-browsers) with `--browsers`,
168
- 3. registers the MCP server with Claude Code at user scope. Run through `npx`, the launcher is `npx -y scenescout serve`, with the absolute path of `npx` where one sits beside node, so it works under nvm/fnm. From a clone or a global install it is the absolute node path plus that install's `dist/mcp-server.js`,
169
- 4. puts the `scenescout` command on your PATH, so `scenescout status`, `scenescout watch` and `scenescout doctor` work from any terminal. Run through `npx`, that is `npm install -g` of the version you just ran; from a clone it is `npm link`, so the command always runs what you last built. If npm refuses (a system-wide node usually needs `sudo` for this), the step prints the command to run by hand and the rest of the setup still counts as done: `npx -y scenescout <command>` works without it.
170
-
171
- Re-run it any time: after moving the folder or switching node versions it refreshes the stored paths. It exits non-zero if a step the tool depends on failed, so it is safe to chain. Opt out of a step with `--no-register`, `--skip-browser` or `--no-command`.
172
-
173
- If `claude` isn't on the PATH of the shell you ran it from, it prints the registration command instead of running it:
174
-
175
- ```bash
176
- claude mcp add --scope user scenescout -- npx -y scenescout serve
177
- ```
125
+ **The test report**, laid out by [a template](examples/report-template.json) you write once: test IDs, the requirements each covers, expected and actual results, a screenshot per step, every deviation listed again for the reviewer, and the SHA-256 of each piece of evidence. SceneScout signs nothing; the sign-off rows are for your people.
178
126
 
179
- </details>
127
+ <p align="center"><img src="examples/screenshots/test-report-results.png" alt="The test report's results table: test TC-01, place an order, covering REQ-ORD-1 and REQ-ORD-2; each step with its expected result, actual result As expected, result Pass, and a screenshot of the page after the step" width="820" /></p>
180
128
 
181
- ### 2๏ธโƒฃ Check it
129
+ <p align="center"><img src="examples/screenshots/test-report-deviations.png" alt="The test report's deviations: TC-03, see archived orders, step 3 expected GET /api/orders to answer 200, actual GET /api/orders answered 500, result Fail" width="820" /></p>
182
130
 
183
- ```bash
184
- npx -y scenescout doctor --engine # any client: node + build + browser
185
- npx -y scenescout doctor # Claude Code: the above, plus the skill and the registration
186
- ```
131
+ Everything is one self-contained HTML page per report, with no scripts, nothing loaded from the network, and a layout that prints. [Recording a check](docs/ci.md#recording-a-check) and [a test report from a template](docs/ci.md#a-test-report-from-a-template) have the details, including what to keep out of the pictures.
187
132
 
188
- Every line should be a โœ“. Anything that isn't prints the exact command that fixes it. Then **start a fresh session** in your client so it picks up the new tools.
133
+ ## ๐Ÿ“บ Watch it work
189
134
 
190
- ### 3๏ธโƒฃ Run it
135
+ Each exploratory run opens a live view on your machine, with one card per agent: what it is doing, the page it is on, and a feed of every action. Three agents are testing the demo app in parallel here:
191
136
 
192
- No app handy? Clone this repository and run `npm run demo:serve`: the [demo app](demo-app/) starts on `http://127.0.0.1:4173`.
137
+ <p align="center"><img src="examples/screenshots/live-view.png" alt="The live view during a run of three parallel agents against the demo app: one card per session, each with its role and objective, the task it is on, the tool it is running, the page it is on, a live thumbnail, and a feed of the actions it just took, tinted one colour per task" width="880" /></p>
193
138
 
194
- Open your agent inside the project you want to test (or, for a [remote URL](#-two-ways-to-use-it), any folder) and ask:
139
+ You can read the report while the agents are still working, and scrub back through any session's timeline. Beside `report.md`, every run writes `report.html`, the whole run as one self-contained page; ask for a recorded run and it also keeps a frame after every action.
195
140
 
196
- ```
197
- Use SceneScout to test http://localhost:3000 at medium level
198
- ```
141
+ ## โœจ Why it's different
199
142
 
200
- In Claude Code the skill gives you a command with flags for the same thing:
143
+ - ๐Ÿง  **Your agent is the brain, so no extra API key.** The engine contains no model. It gives the agent you already pay for deterministic tools and the testing method to use them.
144
+ - ๐Ÿ“ **It reads structure, not pixels.** The agent sees every element with its role, state and layout box, so overlap and broken images are measured, not guessed from a screenshot.
145
+ - ๐Ÿ›ก๏ธ **Safety is enforced on the network, not requested in a prompt.** Nothing existing is changed unless you allow it, and a blocked write never reaches your server.
146
+ - โœ… **"Done" is a contract.** The report lists everything not tested, and at the `extensive` level refuses to finish while any known page is unvisited.
147
+ - ๐Ÿง  **It remembers.** Each run starts from what the last one learned, and re-tests the bugs earlier runs left open.
148
+ - ๐Ÿ“Š **It is measured, not asserted.** Every change to how it explores is scored against an app with planted bugs and a held-out app it is never tuned on ([the log](docs/benchmark.md)):
201
149
 
202
- ```
203
- /scenescout --level medium --url http://localhost:3000 --role qa
150
+ ```mermaid
151
+ xychart-beta
152
+ title "Planted defects found per run, demo app (13 planted)"
153
+ x-axis "Run" ["0", "1", "2", "3", "4", "5", "6", "7", "8", "9", "10", "11", "12"]
154
+ y-axis "Found" 0 --> 13
155
+ bar [11, 12, 9, 10, 11, 11, 10, 10, 11, 13, 11, 11, 12]
204
156
  ```
205
157
 
206
- The agent scans the project (if there is one), attaches read-only, explores, and writes findings to `.scenescout/report.md`. That's it.
158
+ <sub>Each bar is one run of eight parallel agents, re-scored against today's answer key. One run is noisy: read the trend across runs, not a single bar.</sub>
207
159
 
208
- **Common flags** โ€” `--level minimal|medium|extensive` ยท `--url <app>` ยท `--role <name\|path>` (who to explore as: a login saved with `scenescout login`, a storage state found by the scan, or a path to a Playwright storage-state JSON) ยท `--focus <text>` (a ticket or a sentence to check) ยท `--observe` / `--read-only` / `--safe-write` / `--allow-destructive`.
160
+ ## ๐Ÿš€ Get started
209
161
 
210
- **No flags at all** (`/scenescout` on its own) and the agent asks four plain questions instead: the address, whether and how you sign in, what to check (tickets or a description), and whether the site holds real data. Where your client can show a form, `scout_intake` asks them as one form, never for a password; otherwise the agent asks in chat. Real data, or not being sure, means nothing but `GET` requests leave the page; you are never asked to pick a mode. Any flag skips the questions. See [Plain questions instead of flags](docs/guide/Ways-to-use-it.md#plain-questions-instead-of-flags).
211
-
212
- ### ๐Ÿ”‘ Signing in as a role
213
-
214
- For an app behind SSO or MFA, sign in once yourself and let every session reuse it:
162
+ **1. Install** for your agent (Node 20 or newer):
215
163
 
216
164
  ```bash
217
- scenescout login http://localhost:3000 --role admin
165
+ npx -y scenescout install # Claude Code: skill, MCP server and the test browser
166
+ npx -y scenescout install --client cursor # or vscode, codex, gemini, copilot, windsurf
167
+ npx -y scenescout doctor # every line should be a โœ“
218
168
  ```
219
169
 
220
- A browser window opens at the URL. Sign in however the app asks: once you are back on the app with a new session, the window saves it as `.scenescout/auth/admin.json` in the project and closes by itself. A trip through a single sign-on provider and back is followed, not taken for the end. Pressing **Enter** in the terminal saves at once, and `--save enter` makes Enter the only way, as before. Closing the window or pressing Ctrl+C saves nothing. From a conversation, the agent opens the same window with `scout_login`, so no terminal is needed. The file is readable by your account only, `.scenescout/` keeps itself out of git, and the command prints where it saved, how many cookies, origins and databases it holds, never what they are, and how long it will last: read from each cookie's expiry and the `exp` of any JWT in a cookie or in localStorage (the payload is decoded for that one claim, never verified, never printed). The profile keeps cookies, localStorage, IndexedDB and sessionStorage, so an app whose sign-in library keeps its token in sessionStorage or IndexedDB still comes back signed in; sessionStorage is put back only on the origin it came from, once per tab, so a lane that signs out stays signed out. A login saved by an earlier version has no sessionStorage or IndexedDB: record it again if the app keeps its token there. `--project <dir>` saves into another project; `--browser firefox|webkit` records in another browser.
221
-
222
- Then `/scenescout --role admin`, or `scout_attach { role: "admin" }` from any agent. Every session attached with the same role gets its own browser built from that one login, so parallel lanes can all run as `admin`. A role with no saved login is refused with the command to run. `role` and `storageStatePath` are alternatives: pass one.
223
-
224
- Sessions of one role share one saved login, so they share its refresh token too. An app that rotates refresh tokens and treats a second use of a spent one as theft would revoke the whole token family, and sign every session of that role out, the moment two of them refreshed with the same token. SceneScout stops that for a session attached by role. When the page is about to send a refresh token from the role's profile, the session first takes a lock beside the profile (`.scenescout/auth/<role>.json.lock`, owner-only, taken over if its holder has not touched it in 30 seconds). Holding the lock, it re-reads the profile: if another session has rotated the token in the meantime, it loads that profile into its own browser and sends the current token in place of the spent one. Once the page has stored the rotated token, the session writes its state back over the profile and releases the lock. Sessions in separate processes share the lock through the file. A refresh token is recognised by name (a cookie, a storage key, or a field inside a JSON storage value whose name contains `refresh`) and is never printed or logged. A token that is only in a cookie makes a request a refresh only when the request is plausibly the refresh call, a POST to a path named for one or an endpoint seen to rotate the cookie, so a refresh cookie scoped to `/` never holds back the app's other requests; scripts, stylesheets, images and fonts are never brokered, and a broker that cannot do its job lets the request through unchanged. See [the guide](docs/guide/Signing-in.md#sessions-that-share-a-refresh-token). An app whose sign-in renews through the identity provider's own session cookie needs none of this, since no refresh token is shared. `SCENESCOUT_REFRESH_BROKER=off` turns the broker off.
225
-
226
- In CI, where nobody can type, `--script` signs in headless as a test user from `SCENESCOUT_LOGIN_USERNAME`, `SCENESCOUT_LOGIN_PASSWORD` and, for a one-time code, `SCENESCOUT_LOGIN_TOTP_SECRET` or a fixed code the test environment accepts in `SCENESCOUT_LOGIN_OTP_CODE` (with no password for a passwordless sign-in), and saves the same profile. No credential value is ever printed. See [signing in from CI](docs/ci.md#signing-in-from-ci) for the options and the rules: a test tenant's user, never production or a real person's account.
227
-
228
- Before a parallel run, `scout_lane_brief` checks that the planner's saved login will outlast it: `runMinutes` (default 60) plus `expiryMarginMinutes` (default 10). It refuses only when it is sure, meaning every credential in the profile has a date, none was set for another host, and the last of them ends before the run does, and then names the `scenescout login` command to run again. A profile holds cookies other than the sign-in (analytics, preferences), so the first one to expire is reported as a warning rather than a reason to refuse, and a profile with undated credentials in it (a session cookie, or a refresh token with no expiry) is a warning that its lifetime is unknown.
229
-
230
- ---
231
-
232
- ## ๐Ÿ“บ Watching a run live
233
-
234
- When a session attaches, the engine starts a small live view and hands the agent its address on a `Live view:` line, which the agent passes on to you. On a local desktop it also opens that page in your default browser as the session attaches, and opens `report.html` when `scout_report` writes it, whether or not the browser window is shown. Nothing opens in CI, over SSH, or on Linux with no display. `SCENESCOUT_OPEN` (`live`, `report`, `both` or `none`) in the server's environment chooses otherwise, and `scout_attach {open}` wins over it. `scenescout ci` opens nothing unless `SCENESCOUT_OPEN` is set. From a terminal, `scenescout watch` opens the same page. There is one card per session:
235
-
236
- <p align="center"><img src="examples/screenshots/live-view.png" alt="The live view during a run of three parallel agents against the demo app: one card per session, each with its role and objective, the task it is on, the tool it is running, the page it is on, a live thumbnail, and a feed of the actions it just took, tinted one colour per task" width="880" /></p>
237
-
238
- - **What it is doing:** the tool it is running and for how long, the page it is on, and a thumbnail of that page. This works for headless runs too, which have no window to look at.
239
- - **What it just did:** a rolling feed of its actions, each with its target and how it turned out, with failures in red. It is the same trail a finding's repro trace uses. The engine never sees the agent's reasoning, so this is what the session *did*, not what it thought.
240
- - **Stuck, not slow:** a call still running past its own tool's watchdog budget turns the card red, so a wedged session is visible without asking. A crawl legitimately runs for minutes; it is judged against the crawl's budget, not a click's.
241
- - **Live stream:** switch it on for one card, or for all of them. Click a thumbnail for a close-up.
242
- - **The report, as it stands:** the Report button in the top bar shows the same document `scout_report` writes at the end, rendered from the run's current state, so findings can be read while the agents are still working.
243
- - **What it is for:** the close-up puts the feed beside the session's brief โ€” the objective it was given when it attached (`scout_attach {objective}`), and underneath it the task it is on right now (`scout_task`), which the engine requires before any tool will act. Each task tints its own block of actions, so a change of task is a change of colour; point at a block and the brief names the task those actions served.
244
- - **Scrub it back:** under the page is a tick per action, coloured by task. Click one to see the frame from that moment, and `Back to live` to return. On a run that was not recorded the ticks still read the trail; they just have no picture behind them.
245
-
246
- <p align="center"><img src="examples/screenshots/live-view-closeup.png" alt="A close-up of one session: a frame from a step picked out of the timeline, the timeline itself as a tick per action coloured by task, the feed of the session's actions in the same colours, and beside it the objective and the task it is on" width="880" /></p>
247
-
248
- <p align="center"><img src="examples/screenshots/live-view-report.png" alt="The report opened from the live view's top bar while the run is still going: summary table, gap ledger, and the findings filed so far, each with an accordion of the screenshots taken around it" width="880" /></p>
249
-
250
- The view is served on `127.0.0.1` only, behind a token that changes every time the engine starts. It answers `GET` and nothing else, so a viewer can watch a run but not act in it, and no frame it shows is written to disk ([ADR 7](docs/adr/0007-the-live-view-is-local-read-only-and-leaves-nothing-behind.md)) unless the run was recorded, which is asked for and off by default ([ADR 8](docs/adr/0008-a-recorded-run-is-evidence-and-must-be-asked-for.md)). A stream runs only while someone is watching it. `SCENESCOUT_LIVE=off` keeps the port closed.
251
-
252
- **Try it with parallel agents.** The demo app has three roles and several separate areas, so a run can be split between agents. Start it with `npm run demo:serve`, then ask your agent to explore it with several agents in parallel, one role and one area each. The pictures above come from a run of three. Two things keep a parallel run efficient:
253
-
254
- - **Each agent opens its own session when it starts, and the planner closes it once it has folded that agent's report.** An agent waiting for its turn then holds no browser. Opening every session up front leaves browsers idling while the machine runs out of memory for the agents that are working. Closing before the fold loses the lane's decisions, which have nowhere to be kept.
255
- - **Slow it down to follow along.** `scout_attach {paceMs}` (or `scout_session {paceMs}` mid-run) sets a floor between actions, for when you want to watch a flow rather than let it run as fast as the page allows.
256
- - **Run about as many agents at once as your machine has cores, less two.** Each one drives a real browser.
257
-
258
- ---
259
-
260
- ## ๐ŸŽฌ Recording a run, and reading it back
170
+ <details>
171
+ <summary>Other ways to install: a Claude Code plugin, the Claude Desktop extension, or by hand</summary>
261
172
 
262
- A report says what happened. For QA work that is not always enough โ€” the point
263
- is often to *show* what was checked, not to assert it.
173
+ - **Claude Code plugin:** `/plugin marketplace add brunoboto96/SceneScout`, then `/plugin install scenescout@scenescout-marketplace`. The command becomes `/scenescout:scenescout`.
174
+ - **Claude Desktop:** download `scenescout-X.Y.Z.mcpb` from the [latest release](https://github.com/brunoboto96/SceneScout/releases/latest) and open it. No terminal needed.
175
+ - **Any other MCP client:** add a stdio server whose command is `npx -y scenescout serve`. [Each client's config](docs/guide/Start-here.md#a-client-that-install-does-not-know).
264
176
 
265
- Every finding already carries a picture: the element it is about, with a margin,
266
- or the page as it was. It is kept in `.scenescout/recordings/`, shown under the
267
- finding in `report.html`, and returned with the `scout_finding` result, so a
268
- chat client shows the evidence the moment it is filed. Pictures are bounded in
269
- size and in how many reach the conversation, and a CI job keeps them on file
270
- only; `SCENESCOUT_EVIDENCE` and `scout_attach {evidence}` change that
271
- ([configuration reference](docs/guide/Configuration-reference.md#environment-variables)).
177
+ </details>
272
178
 
273
- Ask for a recorded run and the engine also keeps a frame of the page after every action:
179
+ **2. Start your app**, then a fresh session of your agent, and ask:
274
180
 
181
+ ```text
182
+ Use SceneScout to test http://localhost:3000
275
183
  ```
276
- Use SceneScout to test http://localhost:3000, record the run
277
- ```
278
-
279
- or, on the tool directly, `scout_attach {record: true}`. `SCENESCOUT_RECORD=on` in
280
- the server's environment records every run. A CI gate records too:
281
- `scenescout check --record` writes `replay.html`, every journey step by step with
282
- its frames ([recording a check](docs/ci.md#recording-a-check)), and
283
- `--template <file.json>` also writes it up as a test report laid out as a template
284
- says, with expected and actual results, deviations, blank sign-off rows and a
285
- SHA-256 manifest of the evidence ([test reports](docs/ci.md#a-test-report-from-a-template)).
286
-
287
- Then `scout_report` writes two files side by side in `.scenescout/`:
288
- `report.md` as always, and `report.html` โ€” the whole run as one self-contained
289
- page. It opens from the file system with nothing running, needs no network, and
290
- holds:
291
-
292
- - **The report**, rendered from the same Markdown: the plain-language view
293
- first (each problem's steps, what was expected, what happened and its
294
- picture), with each problem's technical detail one click away.
295
- - **The screenshots around each finding**, in an accordion under it, from the
296
- session that filed it.
297
- - **Every session's trail**, in the blocks its tasks made, each step with the
298
- page as it was at that moment.
299
-
300
- <p align="center"><img src="examples/screenshots/run-page.png" alt="The saved copy of a run, opened from the file system with nothing running: a finding with its evidence accordion open, showing the four screenshots taken around it with the action and time under each" width="880" /></p>
301
184
 
302
- The live view serves the same document at `run` while the engine is still up,
303
- and sends you there when the run ends โ€” so the address survives a refresh
304
- instead of a panel over a dead board.
185
+ In Claude Code there is a command too: `/scenescout --url http://localhost:3000 --level medium`. On its own, `/scenescout` asks you four plain questions instead: where the app is, how you sign in, what to check, and whether it holds real data.
305
186
 
306
- **What it costs.** Frames are pictures of the app under test, inside the tested
307
- project's folder, and the secret redaction that protects everything else the
308
- engine writes cannot read a picture. That is why it is off unless asked for,
309
- capped per session, and written only under `.scenescout/`, which ignores itself
310
- so `git add -A` in the tested project cannot pick the frames up. The reasoning is in
311
- [ADR 8](docs/adr/0008-a-recorded-run-is-evidence-and-must-be-asked-for.md).
187
+ **3. Read the report** in `.scenescout/report.md`. It opens in plain words (each problem, its steps, what was expected and what happened), with the technical detail one click away. [A complete example](examples/report.md).
312
188
 
313
- ---
189
+ **No app handy?** Clone this repository and run `npm run demo:serve`: the [demo app](demo-app/) starts on `http://127.0.0.1:4173`, and its README lists every planted bug.
314
190
 
315
- ## ๐Ÿ”„ How a run works
191
+ **Behind a sign-in?** Sign in once yourself, SSO and MFA included, and every session reuses it:
316
192
 
317
- One curiosity loop, repeated โ€” breadth first, then judgment where it matters:
318
-
319
- ```
320
- scan โ”€โ”€โ–ถ attach โ”€โ”€โ–ถ crawl โ”€โ”€โ–ถ investigate โ”€โ”€โ–ถ measure โ”€โ”€โ–ถ report
321
- โ”‚ โ”‚ โ”‚ โ”‚ โ”‚ โ”‚
322
- routes browser every route reproduce & journeys + gap-checked
323
- & auth (r/o) in ONE call file findings design audit markdown
193
+ ```bash
194
+ npx -y scenescout login http://localhost:3000 --role admin # then: /scenescout --role admin
324
195
  ```
325
196
 
326
- 1. **Scan** the project โ€” framework, routes, auth states.
327
- 2. **Attach** a browser (read-only unless you said otherwise).
328
- 3. **Crawl** every known route in a *single* call โ€” per-route HTTP status, element counts, oracle violations, dead ends.
329
- 4. **Investigate** what the crawl flagged: navigate, snapshot, reproduce, file a structured finding.
330
- 5. **Measure** task ease (`scout_journey`) and design quality (`scout_design_audit`) on representative pages.
331
- 6. **Report** โ€” the engine checks the gap ledger and writes `.scenescout/report.md`.
332
-
333
- Snapshots are cheap: re-snapshotting a route returns only *what changed*, with stable refs (measured on a 130-element page: 10.7 kB โ†’ 0.7 kB).
334
-
335
- ---
197
+ [Signing in](docs/guide/Signing-in.md) covers roles, expiry and scripted sign-in for CI.
336
198
 
337
- ## ๐Ÿงฐ The toolbox
199
+ ## ๐Ÿšฆ In CI
338
200
 
339
- 35 deterministic tools. The agent picks; you rarely call these by hand.
340
-
341
- | Phase | Tools | What they do |
201
+ | | What it does | Needs a model? |
342
202
  |---|---|---|
343
- | **Set up** | `scout_playbook` `scout_intake` `scout_scan` `scout_attach` `scout_session` | Hand the testing method to an agent that has no skill loaded; ask the start-of-run questions as one form where the client can show one; discover routes; launch a browser in a write-mode; keep several authenticated roles alive at once |
344
- | **Explore** | `scout_crawl` `scout_coverage` | Sweep every route in one call; ask what's still untested |
345
- | **Look** | `scout_snapshot` `scout_hover` `scout_screenshot` `scout_capture` | Read the structured scene (diffed); reveal tooltips/hover cards; capture pixels only when needed; save one element as a PNG to show someone |
346
- | **Ask the server** | `scout_request` `scout_network` | Call the app's own API as this session, with the UI bypassed โ€” the check that turns a hidden button into a proven refusal; list the requests the page itself made since it loaded |
347
- | **Act** | `scout_click` `scout_type` `scout_select` `scout_upload` `scout_press` `scout_scroll` `scout_navigate` `scout_back` `scout_run_plan` | Drive the UI like a user; `scout_run_plan` batches a whole mechanical sequence into one call |
348
- | **Assess** | `scout_design_audit` `scout_journey` | Score a page's craft/a11y/consistency; measure how hard a task is to complete |
349
- | **Record** | `scout_note` `scout_finding` `scout_resolve` `scout_report` | Curate durable notes; file deduped findings; mark fixes; write the report, and on a recorded run the whole run as one page |
350
- | **Answer tickets** | `scout_tickets` `scout_criterion` | Read the acceptance criteria in pasted or uploaded tickets; record each criterion as passed, failed (with the findings that show it) or not tested (and why), with a confidence |
351
- | **Re-test** | `scout_verify` | List the findings earlier runs left open, worst route first, and record whether each is gone, still present, or changed |
352
- | **Split the work** | `scout_lane_brief` `scout_lane_report` | Divide the app between parallel agents by whole module, each with its own landing route and rules; fold what each hands back as one typed JSON object, and name any defect it judged but never filed |
353
- | **Close** | `scout_close` | Tear down one session or all |
354
-
355
- A few that punch above their weight:
356
-
357
- - **`scout_crawl`** โ€” the entire breadth pass in one tool call. No visiting routes one-by-one.
358
- - **`scout_run_plan`** โ€” up to 20 actions (fill form โ†’ submit โ†’ check) with semantic targets (`testid=โ€ฆ`, `text=โ€ฆ`), aborting at the first anomaly.
359
- - **`scout_journey`** โ€” wraps one goal and reports interaction count, screens seen, and **backtracks**; an abandoned journey is a finding no passing E2E suite can produce.
360
- - **`scout_upload`** โ€” generates a *valid* in-memory fixture (real PDF/PNG, kind inferred from `accept`) so file-upload flows stop being a blind spot.
361
- - **`scout_click {clicks: 2}`** โ€” the impatient-user probe: states whether a double-click fired the same state-changing request twice (the classic double-submit bug).
362
- - **`scout_request`** โ€” calls the app's own API as the session, so "the button is hidden" becomes "the server refuses it" (or doesn't).
363
-
364
- Beyond crashes and HTTP errors, two oracles catch a page **contradicting the server**: `refused_empty` (a list request was refused and the page shows its empty state with no error) and `false_success` (a save was refused and the page says it worked). A third, `dom_injection`, reports a typed markup value coming back as an element on any page any session opens. A fourth, `postmessage_token`, reports a page calling `postMessage` with targetOrigin `"*"` on a message that carries a token (a JWT, a `Bearer` value, or an opaque value under a key such as `access_token`): the report names where in the message it was, its shape and its first four characters, never the token.
365
-
366
- ---
367
-
368
- ## ๐Ÿ“Š Test levels
369
-
370
- Each level is a contract. `scout_report` enforces what the engine can see for itself โ€” routes visited, pages audited, and at `extensive` an empty gap ledger (every route exercised and audited, every filled form submitted, a completed journey, two roles) โ€” and the report's gap ledger discloses the rest.
371
-
372
- | Level | What it guarantees | Rough size |
373
- |---|---|---|
374
- | `minimal` | Every route visited, โ‰ฅ1 design audit, key journeys as plans, crawl problems triaged. Remaining gaps **disclosed**. | ~40 actions |
375
- | `medium` *(default)* | minimal + design audits across several routes (enforced) + every element class exercised + every form submitted valid **and** invalid (asked of the agent) | ~150 actions |
376
- | `extensive` | medium + fuzzing, back/refresh/deep-link resilience, keyboard-only pass, a journey per module, โ‰ฅ2 roles compared, anonymous auth-surface walk. **Refuses to finalize while any gap in the ledger remains.** | budget-capped |
377
-
378
- That refusal *is* the guarantee: an extensive report can only exist when nothing the engine can measure was left untested. Fuzzing, the keyboard pass and the auth-surface walk are the agent's to do; the engine cannot see whether they were done well.
379
-
380
- ---
381
-
382
- ## ๐Ÿ”’ Safety model
203
+ | [`scenescout check`](docs/guide/Ways-to-use-it.md#scenescout-check-a-gate-in-ci) | A deterministic gate: measures every page, replays your saved journeys and visual baselines, fails only on what it can prove; recorded, it leaves the [evidence above](#-automated-qa-with-evidence) | No |
204
+ | [`scenescout ci`](docs/guide/Ways-to-use-it.md#scenescout-ci-an-unattended-exploratory-run) | An unattended exploratory run, driven by the Anthropic or OpenAI API. It reports and never fails the build | An API key |
205
+ | [`/scenescout qa`](docs/guide/Ways-to-use-it.md#scenescout-qa-on-a-pull-request) | A comment on a pull request that tests its preview deploy and replies with the results; `/scenescout qa check` runs your own check instead | An API key (`qa check`: none) |
206
+ | [`scenescout export`](docs/guide/Ways-to-use-it.md#filing-findings-as-issues) | Files the findings as GitHub or Jira issues, each once | No |
383
207
 
384
- - ๐Ÿ”ต **`observe`** (`--observe`) lets nothing but `GET` requests leave the page. The one exception is what a session needs in order to exist: logging in, logging out and refreshing a token. Signing up, changing or resetting a password and creating users are blocked like any other write. WebSocket frames are not inspected; the engine says so when the app opens a socket. It is what the skill picks for a remote URL with no source, where an ordinary form POST would create a real record. Forms that could not be submitted are listed in the gap ledger.
385
- - ๐ŸŸข **`read-only` by default.** Destructive-labeled elements (delete/revoke/archive/โ€ฆ) **and** all `PUT/PATCH/DELETE` + destructive `POST`s are blocked at the network layer โ€” see [`src/engine/policy.ts`](src/engine/policy.ts). Non-destructive `POST`s are allowed, because submitting forms is how a tester finds validation bugs โ€” so read-only means *nothing existing is changed or removed*, not *nothing is ever created*.
386
- - ๐ŸŸก **`safe-write`** (`--safe-write`) lets the agent create data and edit/delete **only what it created** this run โ€” never pre-existing records.
387
- - ๐Ÿ”ด **`destructive`** (`--allow-destructive`) allows everything, and only ever when *you* confirm the environment is disposable. The skill will never choose this itself.
388
- - ๐Ÿ“‚ Findings, memory, and reports live in a `.scenescout/` folder in the project. A client with no project folder, such as a desktop chat, gets one folder per tested site under `Documents/SceneScout/<host>/` by default, and the attach says where; `SCENESCOUT_PROJECTS_DIR` moves it. The folder ignores itself in git, so a stray `git add -A` never commits test data.
389
-
390
- A `๐Ÿ›ก WRITE-POLICY blocked` notice is the safety net doing its job, not an app bug. The server never sees a blocked request, but a page's own `fetch` or XHR is answered with a `403` in its place rather than dropped, so the page's handling of a refusal really runs: a page that then claims success is reported as a `false_success` ([ADR 9](docs/adr/0009-a-refused-write-is-answered-not-dropped.md)).
391
-
392
- A control is judged by its own label: a dropdown by the option picked, a row by its own text rather than the buttons inside it, and "Discard changes" on an unsent form is allowed. When a page asks to confirm leaving unsent input, the result says so; `observe` and `read-only` stay unless the call passes `leave: true`. See the [safety model](docs/guide/Safety-model.md).
393
-
394
- ---
395
-
396
- ## ๐Ÿ“‹ What you get
397
-
398
- `.scenescout/report.md` and `report.html` open **In plain words**: a short summary, then each problem this run found, worst first, with its impact, the steps that led to it, what was expected, what happened and a picture when there is one. The technical detail (id, category, evidence, route) stays one click away.
399
-
400
- `.scenescout/report.md` โ€” a deduplicated, worst-first report with:
401
-
402
- - ๐Ÿ› **Findings** with repro traces and generated Playwright regression-test skeletons.
403
- - ๐Ÿ”Ž **Worth a look** โ€” observations that are defects only under a convention of your project the run cannot see (a spacing scale, link styling in navigation, test ids on every control), each naming that convention. Listed below the findings and not counted as defects ([ADR 13](docs/adr/0013-a-convention-is-the-projects-to-decide.md)).
404
- - ๐Ÿ’ฏ **Page scores** (0โ€“100: a11y ยท craft ยท consistency ยท task-clarity), ranked worst-first, with stale scores from old runs marked as such.
405
- - ๐Ÿ‘ฅ **A role capability matrix** โ€” what each role could and couldn't reach.
406
- - ๐Ÿงพ **A gap ledger** โ€” everything *not* done, so the report is honest about its own coverage.
407
- - โฑ๏ธ **How the run was paced** โ€” how closely each session kept working, and apart from that, how long finished lanes held their browsers waiting to be collected, so neither hides the other.
408
- - ๐ŸŽฏ **How well the lanes judged** โ€” on a parallel run, whether the confidence each lane stated matched what the project went on to file, beside what later re-tests found ([ADR 10](docs/adr/0010-a-confidence-is-checked-not-trusted.md)).
409
-
410
- `.scenescout/report.html` โ€” the same report as one self-contained page, with every session's trail beside it. Each finding shows a picture of the element it is about, or of the page as it was. A [recorded run](#-recording-a-run-and-reading-it-back) also shows the screenshots around each finding.
411
-
412
- ๐Ÿ‘€ Watch a run live: `npx scenescout watch`, or `npx scenescout status <project-path>` for the same information as text.
413
-
414
- ---
415
-
416
- ## ๐Ÿšฆ In CI: a deterministic check
417
-
418
- An exploratory run is driven by a model, so two runs never find exactly the same things. That suits a report, but not a gate. `scenescout check` is the part that needs no model. It visits the start URL, the project's scanned routes and every same-origin link it finds, and measures each page:
419
-
420
- - HTTP and page errors
421
- - layout geometry (covered, clipped and overlapping controls, blocking overlays)
422
- - broken images
423
- - controls with no name, and fields whose only label is a placeholder
424
- - contrast and focus
425
- - pages with no way out
426
-
427
- Some of what it measures is a defect only under a convention the check cannot see: paddings off a 4px grid, and links styled like body text. Those are listed under **Worth a look**, each with the convention that would make it a defect. They are never counted and never fail the gate, at any `--fail-on`; SARIF reports them at level `note` ([ADR 13](docs/adr/0013-a-convention-is-the-projects-to-decide.md)).
428
-
429
- ```bash
430
- npx scenescout check http://127.0.0.1:3000 --fail-on high
431
- ```
432
-
433
- | Exit code | Meaning |
434
- |---|---|
435
- | 0 | Passed the gate |
436
- | 1 | Failed it: something at the `--fail-on` severity or worse |
437
- | 2 | Could not run, or not all of it: a bad argument, an app that never answered, only the sign-in page reached, a saved flow that is not valid, or a flow step the write policy refused |
438
-
439
- It writes `report.md`, `check.sarif` (for code-scanning dashboards) and `check.json` to `.scenescout/check/`, and on GitHub Actions it also puts the report on the job's summary page.
440
-
441
- On GitHub Actions, this repository is also an action that installs everything and keeps the results:
208
+ A gate on every pull request that keeps its evidence, as a GitHub Action:
442
209
 
443
210
  ```yaml
444
- - uses: brunoboto96/SceneScout@v3.10.0
211
+ - uses: brunoboto96/SceneScout@v3
445
212
  with:
446
213
  url: http://127.0.0.1:3000
214
+ record: on
215
+ video: on
216
+ template: tests/report-template.json
447
217
  ```
448
218
 
449
- [docs/ci.md](docs/ci.md) has a complete workflow (start the app, wait for it, check it), the action's inputs and outputs, code-scanning upload, and the same check on GitLab CI, CircleCI or any shell.
450
-
451
- With the default settings its saved flows send no HTTP write (they replay under observe's rule), and its crawl runs under `--mode observe` or `read-only`; `--flow-writes allow` lets flows write as `--mode` allows. By default it fails only on facts that mean a page is broken: a page that did not load, an uncaught exception, a 5xx, a failure shown as success. Other options:
452
-
453
- - `--fail-on medium` or `low` makes the gate stricter.
454
- - `--ignore <rule>` drops a rule, worth-a-look rules included.
455
- - `--paths /a,/b` checks only those pages.
456
- - `--storage-state <file>` checks while signed in.
457
- - `--flows <dir>` or `off` chooses which saved flows to replay; `--retest off` skips re-testing open findings.
458
- - `--flow-writes never|allow` (default `never`): `never` replays flows under observe's rule whatever `--mode` says; `allow` replays them under `--mode`, so in `read-only` a flow's form submissions are sent to the target on every run.
459
- - `--on-refused-step report|stop` (default `report`): `report` marks a flow whose step was refused "could not run", keeps every other verdict and exits 2; `stop` exits 2 at that step with no results.
460
- - `--gate-retests never|high|all` (default `high`): which still-reproducing re-tested findings fail the gate.
461
- - `--baseline off|compare|update` (default `off`), with `--baselines <dir>` and `--baseline-threshold <percent>` (default `0.1`, so small anti-aliasing noise between machines passes): visual baselines, below.
462
- - `--record` (or `SCENESCOUT_RECORD=on`; the action's `record: on`) keeps a frame after each route visit and each saved-flow step and writes `replay.html` beside the report: each role, each journey with a pass or fail badge, each step with its caption, result and frame, the first failing step highlighted. `--video` (the action's `video: on`) adds a WebM of each journey, played on that page beside its steps. Both off by default. [docs/ci.md](docs/ci.md#recording-a-check) covers size, privacy and publishing it.
463
-
464
- The defaults are what an unconfigured check does, for a first try or an AI agent running it unattended: its flows send no HTTP write and it never silently hides a result. Each setting is a choice for the project; the report and `check.json` print the values a check ran with.
465
-
466
- `scenescout check --help` lists every option. Why the defaults are what they are: [ADR 11](docs/adr/0011-a-gate-is-deterministic-and-fails-only-on-what-it-can-prove.md).
467
-
468
- It also replays the flows saved in `.scenescout/flows/*.json`, with no model: the steps `scout_run_plan` takes (navigate, click, type, select, press, and upload, which attaches a generated file) plus `expect-text`, `expect-element` (a control is visible, hidden, enabled, disabled, checked or unchecked), `expect-url` and `expect-request`, and `repeat` to run actions until a condition holds (page through a document until Continue is enabled). A flow can name the `role` it runs as, signed in with a profile `scenescout login --role` saved, so a journey that passes between people (one submits, another approves) is a sequence of flows. Values can come from the environment, `${env:NAME}`, so a code or a password stays in the CI's secret store and is masked in everything the check writes. A flow whose step breaks fails the gate, naming the flow and the step. And it re-tests the open findings earlier runs left in the project's memory that a page load can reproduce, reporting each as still reproducing or possibly fixed; by default a finding filed high that still reproduces fails the gate. [docs/ci.md](docs/ci.md#saved-flows) has the flow format; [ADR 12](docs/adr/0012-a-check-replays-saved-flows-and-reports-re-tests.md) says why it works this way.
469
-
470
- With `--baseline compare` it also holds pages and elements to approved pictures: list them in a `targets.json`, take the baselines once with `--baseline update`, and a later check that finds one changed past `--baseline-threshold` fails the gate with the share of pixels changed and a diff picture beside the report. Baselines are kept per browser, in a git-ignored folder unless `--baselines` names one the project commits, and only `--baseline update` ever writes one. [The guide](docs/guide/Ways-to-use-it.md#visual-baselines) has the details; [ADR 19](docs/adr/0019-a-visual-baseline-changes-only-when-asked.md) says why.
471
-
472
- Beyond those flows it explores nothing and fills no forms. That is the exploratory run's job, and its findings belong in a report, not a gate.
473
-
474
- ## ๐Ÿค– In CI: an unattended exploratory run
475
-
476
- `scenescout ci` runs the exploratory side in a CI job, with no person and no coding agent. A model reached through its API drives the same `scout_*` tools by the same method, and the run ends in the ordinary report:
477
-
478
- ```bash
479
- export OPENAI_API_KEY=โ€ฆ # or ANTHROPIC_API_KEY; read from the environment only
480
- npx scenescout ci http://127.0.0.1:3000
481
- ```
482
-
483
- - **It reports and never gates.** Exit 0 when the run ran, whatever it found; exit 2 when it could not run (no key, a key the API refused, an app that never answered). Two runs of the same app find different things, so a finding is something to read, never a reason to fail a build. `scenescout check` is the gate.
484
- - **Providers:** the Anthropic Messages API (default model `claude-sonnet-5`) or the OpenAI Responses API (default `gpt-6-luna`), chosen by which key is set; with both set, `--provider` decides. `--model` and `--effort` (default `low`) override; `--base-url` points at another endpoint that implements the same API.
485
- - **Caps:** at most 80 model turns, 3,000,000 tokens and 20 minutes (`--max-turns`, `--max-tokens`, `--max-minutes`), shared by two model loops that each explore their own part of the app (`--lanes`, default 2). The first cap reached ends the exploration; the report is still written, and says which cap ended it. On the benchmark's demo app a run at these defaults cost about $0.03 on `gpt-6-luna` ([docs/benchmark.md](docs/benchmark.md#choosing-the-defaults-issue-419)).
486
- - **Mode:** `read-only` by default; `--mode observe` sends no form at all, `--mode safe-write` lets the run create records and change only the ones it created. `--mode destructive` runs only with `--allow-destructive` as well.
487
- - **Duplicates:** when the dedup rule keeps a filed finding apart, the run's model is asked at its lowest effort whether it is one already open on the same page, and merges it on a "same", keeping the filing's title, category, severity and evidence under that finding. The two findings' titles, categories and evidence, and the page's path, are sent; `--dedup rule` turns it off ([ADR 17](docs/adr/0017-a-model-judges-only-the-merges-the-rule-misses.md)).
488
- - **Starting from an earlier run (opt-in):** every run that writes its report leaves a record in `ci.json` and the project's memory: the routes it worked on, its steps, and what it left on each route. `--from-run <ci.json or project directory>` continues where that run left off (routes it never worked on first, then the ones it left work on, with exactly which forms, options and controls to take first), and `--from-run-mode replay` follows its routes and steps in order. `scout_lane_brief {fromRun, fromRunMode}` does the same for a parallel run. On the demo app a chain of continued runs found 8 defects against 7 for fresh runs at the default budget, and 4, 3 and 4 against 4 at a budget cut to stand in for a larger app, within the noise, so it stays off unless asked for ([the measurement](docs/benchmark.md#starting-from-an-earlier-run-issue-418)).
489
- - **Output**, in `.scenescout/ci/` (or `--out`): `report.md` and `report.html` (the report an agent's run writes), `summary.md` (also appended to the GitHub job summary), `ci.json` and `ci.sarif`, with a usage line: turns, tokens, time and an estimated cost where the model's price is known (`--price-in`, `--price-out` give one for any model).
219
+ [docs/ci.md](docs/ci.md) has complete workflows, every option, and the same check on GitLab CI, CircleCI or any shell.
490
220
 
491
- There is a GitHub Action for it (`uses: brunoboto96/SceneScout/ci@โ€ฆ`). [docs/ci.md](docs/ci.md#an-unattended-exploratory-run) has the workflow and every option; [ADR 14](docs/adr/0014-an-unattended-run-reports-and-never-gates.md) says why it works this way.
221
+ ## ๐Ÿ”’ Safe by default
492
222
 
493
- On a pull request, an allowed account can comment `/scenescout qa` to run it against that pull request's deployed preview and get the results as a reply. The job that holds the key checks out nothing and runs SceneScout from an exact release tag, so the pull request's code never runs beside the key. A project without previews can comment `/scenescout qa check` instead: it dispatches the project's own recorded `scenescout check` workflow on the pull request's branch and replies with the verdict, per journey, with no model key involved. [docs/ci.md](docs/ci.md#a-qa-review-from-a-pull-request-comment) has the workflows to copy and what a project configures; [ADR 15](docs/adr/0015-a-qa-comment-tests-a-preview-and-never-runs-the-pull-requests-code.md) says why.
494
-
495
- ## ๐Ÿ“ฎ Filing findings as issues
496
-
497
- `scenescout export` turns the project's open findings into GitHub or Jira issues, where the team already works. It reads `.scenescout/memory.json`, so it runs after an interactive run, `scenescout ci` or anything else that wrote findings:
498
-
499
- ```bash
500
- export GH_TOKEN=โ€ฆ # or GITHUB_TOKEN; for Jira, JIRA_EMAIL and JIRA_API_TOKEN. Read from the environment only
501
- npx scenescout export --to github --repo owner/app # a dry run: lists what it would file
502
- npx scenescout export --to github --repo owner/app --yes # files it
503
- npx scenescout export --to jira --jira-url https://your-site.atlassian.net --jira-project QA --yes
504
- ```
505
-
506
- - **Each finding once.** Every issue carries the `scenescout` label and a marker with the finding's id. Before filing, the export reads the labelled issues, open or closed, and skips every finding already filed, naming its issue, so a second export of the same run files only what the first left over the cap. A closed won't-fix is not filed again; `--refile-closed` files a finding again when its issue is closed.
507
- - **Jira issues are kept up to date and linked to the ticket.** A later export updates an open Jira issue instead of filing another, leaving text someone edited in Jira as they wrote it. The finding's picture is attached, and a finding that fails a ticket's acceptance criterion is linked to that ticket (`--jira-link-type`, default `Relates`).
508
- - **A dry run unless `--yes`**, and at most `--max-issues` (default 20) per export; the next export files the rest. `--min-severity`, `--only <ids>` and `--include-worth-a-look` choose what goes.
509
- - **Inert issues.** Titles, descriptions, steps and evidence come from the run and the app's pages, so no `@mention`, link, `#123` reference, HTML or Markdown in them does anything.
510
- - **Severity** becomes a label on GitHub and a priority in Jira; `--severity-map` renames them or turns them off. **Screenshots** from a recorded run are attached in Jira; GitHub's API takes no uploads, so a GitHub issue names the frames in the run's `.scenescout/` folder.
511
- - **Credentials** are never printed. Every request has a timeout, a short rate limit is waited out, failed reads are retried with backoff, and redirects are refused. Jira's search can take a little while to show a new issue, so leave a few minutes between two exports to the same Jira project.
512
-
513
- [The guide](docs/guide/Ways-to-use-it.md#filing-findings-as-issues) has the details and a GitHub Actions step.
514
-
515
- ---
516
-
517
- ## ๐Ÿฉบ Troubleshooting
518
-
519
- Run `npx -y scenescout doctor` first โ€” it checks every setup item below (everything but the last row, which is about your app) and prints the fix.
520
-
521
- | Symptom | Cause and fix |
223
+ | Mode | What may leave the page |
522
224
  |---|---|
523
- | `/scenescout` isn't a known command | The skill isn't linked, or the session predates it. `npx -y scenescout install`, then start a **fresh** Claude Code session. |
524
- | The `scout_*` tools don't appear | The MCP server isn't registered, or points at an old path. `npx -y scenescout install` re-registers it; `claude mcp list` should show `scenescout` as connected. |
525
- | *"Executable not found in $PATH"* | The server was registered with a bare `node`. `npx -y scenescout install` registers an absolute path. |
526
- | Installed as a plugin, and the tools fail with *"Executable not found in $PATH: npx"* | A plugin starts the server with a bare `npx`, which Claude Code can only find if it was launched from an environment that has Node on its `PATH`. Under nvm or fnm that means starting Claude Code from a terminal, not from a dock or launcher. Or use `npx -y scenescout install` instead, which registers the absolute path of `npx`. |
527
- | *"โ€ฆ build has not been downloaded yet"* on attach | The attach downloads a missing browser itself, once, except in CI or with `SCENESCOUT_BROWSER_DOWNLOAD=off`; there, or when that download failed, it names the command to run. Run the command the message names, for example `npx -y scenescout install --browser-only --browsers firefox`. On Linux, system libraries may be missing too: `npx playwright install --with-deps chromium`. |
528
- | Tools broke after moving the folder or changing node version | The registration stores absolute paths. `npx -y scenescout install` refreshes them. |
529
- | Attach fails or every route lands on the login page | Your app isn't running at `--url`, or the `--role` session has expired. For a saved login, run `scenescout login <url> --role <name>` again; for a storage-state file, regenerate it the way your project's Playwright setup does. |
530
-
531
- ### โฌ†๏ธ Upgrading from an older version
532
-
533
- - **Tools are now `scout_*`.** Up to v0.23 they were prefixed `ft_`. The rename happened before the first npm release, with no aliases, so an agent's context carries one tool list rather than two. Re-run `npx -y scenescout install` so the installed skill matches the server.
534
- - **Earlier names.** This tool was previously called SceneCraft (and, before that, frontend-tester). `scenescout install` cleans up after both: it removes the old skill link and the old `scenecraft` MCP registration when they point at this install, and the first attach in a project moves its `.scenecraft/` memory folder to `.scenescout/` so earlier coverage and findings carry over.
535
-
536
- ### ๐Ÿงน Uninstall
537
-
538
- ```bash
539
- # Claude Code
540
- claude mcp remove --scope user scenescout
541
- rm -rf ~/.claude/skills/scenescout
542
- # Codex / Gemini / Copilot CLI
543
- codex mcp remove scenescout # likewise: gemini mcp remove โ€ฆ, copilot mcp remove โ€ฆ
544
- ```
545
-
546
- For Cursor, Windsurf and VS Code, delete the `scenescout` entry from the client's MCP server list.
225
+ | ๐Ÿ”ต `observe` | Reads only. The default for a first look, for `scenescout check`, and for an agent's run on a remote site with no source |
226
+ | ๐ŸŸข `read-only` | Reads and ordinary form posts; nothing existing is changed or deleted. The default for an agent's run on a local app |
227
+ | ๐ŸŸก `safe-write` | Creates records, and edits or deletes only the ones it created |
228
+ | ๐Ÿ”ด `destructive` | Everything. Only when you say the data is disposable; the agent never picks it |
547
229
 
548
- Nothing else is installed: `npx` runs the package from npm's cache. Per-project memory lives in each tested project's `.scenescout/` folder; delete it there if you want it gone.
230
+ The policy sits on the network, so a blocked request never reaches your server. The page gets a refusal instead, which is how SceneScout catches a page that claims success anyway. A saved journey that creates something, like the order above, runs only when the check is told `--flow-writes allow`, against a test environment. Findings, memory and reports stay in a `.scenescout/` folder that keeps itself out of git.
549
231
 
550
- ---
551
-
552
- ## ๐ŸŒ Choosing browsers
553
-
554
- `install` downloads Chromium and nothing else unless you ask. `--browsers` takes one name, a comma-separated list, or `all`:
555
-
556
- | `--browsers` | What is downloaded | About, on disk |
557
- |---|---|---|
558
- | `chromium` *(default)* | the full browser and the headless shell | 550 MB |
559
- | `chromium-headless-shell` | the headless shell only: every run works except `headed` | 200 MB |
560
- | `firefox` | Firefox | 270 MB |
561
- | `webkit` | WebKit, the engine behind Safari | 290 MB |
562
- | `all` | Chromium, Firefox and WebKit | 1.1 GB |
563
-
564
- ```bash
565
- npx -y scenescout install --browsers chromium-headless-shell # the smallest working setup
566
- npx -y scenescout install --browser-only --browsers firefox,webkit # add two more later
567
- ```
568
-
569
- Sizes vary by platform. The builds go to Playwright's shared cache, so a build another tool already fetched is not downloaded again.
570
-
571
- The first attach that needs a build which is not on disk downloads it itself, once, and says so ("Getting the test browser ready"). It does not in CI unless `SCENESCOUT_BROWSER_DOWNLOAD=on` is set, and never with `SCENESCOUT_BROWSER_DOWNLOAD=off`, for a machine where nothing may be downloaded.
572
-
573
- To drive another browser, pass `browser` when attaching (`scout_attach { browser: "firefox" }`), or set `SCENESCOUT_BROWSER=webkit` in the server's environment to change the default. `scenescout doctor` checks the browser named by that variable in the shell it runs from, so check another one with `SCENESCOUT_BROWSER=webkit scenescout doctor`. Two things differ outside Chromium:
574
-
575
- - **Service workers are not allowed to register** in Firefox and WebKit. The write policy works by intercepting requests, and only Chromium lets a request issued by a service worker be intercepted. An app that depends on its worker may behave differently there.
576
- - **A Firefox or WebKit left behind by a crash is not cleaned up** on the next start the way a leftover Chromium is.
577
-
578
- In every browser, pages are not given shared workers unless the mode is `destructive`: a request a shared worker sends cannot be intercepted anywhere, so the app is made to do that work on the page, where the policy sees it.
579
-
580
- ### Time limits
581
-
582
- An action on the page (a click, typing, a hover, a pick from a list) may take 5 s, and a page 20 s to load (15 s for a page the crawl opens). On a loaded machine these can run out while the app is fine; the timeout then says which limit ran out and how to raise it. Raise them per session with `scout_attach { actionTimeoutMs: 15000, navTimeoutMs: 60000 }`, for every session with `SCENESCOUT_ACTION_TIMEOUT_MS` and `SCENESCOUT_NAV_TIMEOUT_MS` in the server's environment, or on `scenescout check` and `scenescout ci` with `--action-timeout-ms` and `--nav-timeout-ms`. An option wins over the variable, and the variable over the default. The action limit takes 1000 to 120000 ms and the page-load limit 1000 to 300000 ms; anything else refuses the attach with a sentence naming the value to fix. Saving a login profile with `scenescout login` honours the two variables as well, and otherwise keeps its own longer waits (30 s for the page, 10 s for a field or the submit).
583
-
584
- ## ๐Ÿ”Œ Other MCP clients
232
+ > [!IMPORTANT]
233
+ > Only test sites you own or are allowed to test. [Safety model](docs/guide/Safety-model.md).
585
234
 
586
- The engine is a plain MCP server over stdio, so any client can drive it, and the testing method reaches the agent through the server itself (see the end of this section). `install` can register it for you:
235
+ ## ๐Ÿ“š Documentation
587
236
 
588
- ```bash
589
- npx -y scenescout install --client cursor # one client
590
- npx -y scenescout install --client vscode,codex # several; add claude-code to keep that one too
591
- ```
592
-
593
- | `--client` | How it is registered |
237
+ | | |
594
238
  |---|---|
595
- | `claude-code` *(default)* | `claude mcp add`, plus the skill |
596
- | `cursor` | adds an entry to `~/.cursor/mcp.json`, keeping the others |
597
- | `vscode` | VS Code's own `code --add-mcp`. A `code` command that belongs to another editor is not used |
598
- | `codex` | `codex mcp add` |
599
- | `gemini` | `gemini mcp add --scope user` |
600
- | `copilot` | `copilot mcp add` (GitHub Copilot CLI) |
601
- | `windsurf` | adds an entry to `~/.codeium/windsurf/mcp_config.json`, keeping the others |
602
-
603
- A config file that is not valid JSON is left untouched, and the entry to add by hand is printed instead; a config that is a link into a dotfiles repository is written through the link. When a client that is registered through its own command is not installed, `install` says so and prints the command to run later. Cursor and Windsurf are files, so their entry is written whether or not the editor is installed yet. On Windows, a client installed through npm is a `.cmd` shim that `install` cannot start; it prints the command for you to run instead. Then restart the client and ask its agent: *"Use SceneScout to test http://localhost:3000"*.
604
-
605
- What has been checked: registering through each command above was run against Codex CLI, Gemini CLI, GitHub Copilot CLI and VS Code, and Cursor's command line agent read the entry `install` wrote, connected and listed the tools. The Windsurf path follows its documentation. A full test session has been run in Claude Code, with and without the skill. If a client behaves differently for you, a correction is welcome (say which client version you checked).
606
-
607
- To register by hand instead, the server entry is always the same command, `npx -y scenescout serve`:
608
-
609
- <details>
610
- <summary><strong>Cursor</strong> โ€” <code>~/.cursor/mcp.json</code> (or <code>.cursor/mcp.json</code> in a project)</summary>
611
-
612
- ```json
613
- {
614
- "mcpServers": {
615
- "scenescout": { "command": "npx", "args": ["-y", "scenescout", "serve"] }
616
- }
617
- }
618
- ```
619
-
620
- </details>
621
-
622
- <details>
623
- <summary><strong>VS Code</strong> (GitHub Copilot agent mode) โ€” <code>.vscode/mcp.json</code></summary>
239
+ | [Start here](docs/guide/Start-here.md) | A first look, install, a first run, reading the report, the live view |
240
+ | [Ways to use it](docs/guide/Ways-to-use-it.md) | Interactive runs, parallel agents, CI, pull-request QA, filing issues |
241
+ | [What it checks](docs/guide/What-it-checks.md) | Every check and every `scout_*` tool |
242
+ | [Signing in](docs/guide/Signing-in.md) ยท [Safety model](docs/guide/Safety-model.md) | Roles and saved logins; what each mode refuses and why |
243
+ | [Recipes](docs/guide/Recipes.md) | Setups for seven kinds of project |
244
+ | [Configuration reference](docs/guide/Configuration-reference.md) | Every option, environment variable and action input |
245
+ | [Troubleshooting](docs/guide/Troubleshooting.md) | Symptoms and fixes, upgrading and uninstalling |
246
+ | [Running it in CI](docs/ci.md) | Saved journeys, recording, test reports from templates, every workflow |
247
+ | [How it works](docs/how-it-works.md) | Diagrams of a run, an action, the write policy, lanes |
248
+ | [Benchmark](docs/benchmark.md) ยท [Validation](docs/validation.md) | How runs are scored against answer keys, and runs on public apps |
249
+ | [Design decisions](docs/adr/README.md) | Why the rules are what they are |
624
250
 
625
- ```json
626
- {
627
- "servers": {
628
- "scenescout": { "type": "stdio", "command": "npx", "args": ["-y", "scenescout", "serve"] }
629
- }
630
- }
631
- ```
251
+ ## ๐Ÿ”ง Contributing
632
252
 
633
- </details>
634
-
635
- <details>
636
- <summary><strong>Codex CLI</strong> โ€” <code>~/.codex/config.toml</code></summary>
637
-
638
- ```toml
639
- [mcp_servers.scenescout]
640
- command = "npx"
641
- args = ["-y", "scenescout", "serve"]
642
- ```
253
+ Start with [VISION.md](VISION.md) (what is in scope) and [CONTRIBUTING.md](CONTRIBUTING.md) (local setup and how changes land). [AGENTS.md](AGENTS.md) holds the house rules for people and coding agents alike.
643
254
 
644
- </details>
645
-
646
- <details>
647
- <summary><strong>Gemini CLI</strong> โ€” <code>~/.gemini/settings.json</code> (or <code>.gemini/settings.json</code> in a project)</summary>
648
-
649
- ```json
650
- {
651
- "mcpServers": {
652
- "scenescout": { "command": "npx", "args": ["-y", "scenescout", "serve"] }
653
- }
654
- }
655
- ```
255
+ Found a way past the write policy, or another security problem? Report it privately: [SECURITY.md](SECURITY.md).
656
256
 
657
- </details>
658
-
659
- <details>
660
- <summary><strong>Windsurf</strong> โ€” <code>~/.codeium/windsurf/mcp_config.json</code></summary>
661
-
662
- ```json
663
- {
664
- "mcpServers": {
665
- "scenescout": { "command": "npx", "args": ["-y", "scenescout", "serve"] }
666
- }
667
- }
668
- ```
669
-
670
- </details>
671
-
672
- <details>
673
- <summary><strong>Cline</strong> โ€” MCP Servers โ†’ Configure โ†’ Configure MCP Servers (or <code>~/.cline/mcp.json</code> for the CLI)</summary>
674
-
675
- ```json
676
- {
677
- "mcpServers": {
678
- "scenescout": { "command": "npx", "args": ["-y", "scenescout", "serve"], "disabled": false, "autoApprove": [] }
679
- }
680
- }
681
- ```
682
-
683
- </details>
684
-
685
- <details>
686
- <summary><strong>Zed</strong> โ€” <code>settings.json</code> (command palette: <code>zed: open settings file</code>)</summary>
687
-
688
- ```json
689
- {
690
- "context_servers": {
691
- "scenescout": { "command": "npx", "args": ["-y", "scenescout", "serve"], "env": {} }
692
- }
693
- }
694
- ```
695
-
696
- </details>
697
-
698
- <details>
699
- <summary><strong>Anything else</strong></summary>
700
-
701
- Most clients accept the same `mcpServers` JSON shape shown for Cursor.
702
-
703
- </details>
704
-
705
- **The method travels with the server.** The tools are only hands and eyes; [`skills/scenescout/SKILL.md`](skills/scenescout/SKILL.md) is the method: what to look at first, when to stop, what counts as a finding. Claude Code loads it as a skill. Every other client gets the same text from the server, with nothing to copy:
706
-
707
- - the server's instructions tell the agent to call `scout_playbook` before its first attach, and that tool returns the method,
708
- - clients that list server prompts as commands also get an `explore` prompt, which loads the method and takes an optional URL, level and focus.
709
-
710
- So in any client, a first message like *"Use SceneScout to test http://localhost:3000"* is enough. If an agent starts clicking without having called `scout_playbook`, tell it to call that first; how closely a model follows server instructions varies by client.
711
-
712
- The CLI is also useful on its own:
713
-
714
- ```bash
715
- npx -y scenescout scan <path> # project discovery: framework, routes, saved logins
716
- npx -y scenescout status <path> # what every session of a running engine is doing right now
717
- npx -y scenescout watch <path> # the same, live in your browser, with each session's page
718
- npx -y scenescout login <url> --role admin # sign in once in a visible browser; sessions attach with role: "admin"
719
- ```
720
-
721
- ---
722
-
723
- ## ๐Ÿ“ Project layout
724
-
725
- ```
726
- src/
727
- mcp-server.ts the 29 tools + per-session dispatch
728
- scan.ts project discovery (framework, routes, auth)
729
- cli.ts scan ยท serve ยท install ยท doctor ยท check ยท ci ยท login ยท export ยท status ยท watch
730
- check-run.ts drives a check: attach, crawl every route, collect what was measured
731
- ci-run.ts drives a CI run: the MCP server as a child, the model's API, the agent loop
732
- login-run.ts drives `scenescout login` and scout_login: a visible browser that saves the role's profile once the sign-in is seen to finish (or on Enter); or --script, headless from the environment
733
- export-run.ts drives `scenescout export`: reads the findings, asks GitHub or Jira what is filed, files the rest
734
- installer.ts setup logic (skill link, MCP registration, diagnostics)
735
- engine/
736
- browser.ts the engine class: attach, snapshot, actions, crawl, plans
737
- probes.ts in-page scroll + overlay + focus probes (needs a browser too)
738
- fingerprint.ts route + element-set identity (state hashing)
739
- oracles.ts console/page/network/HTTP error detection
740
- injection.ts the DOM-injection oracle's rules (what to watch for, how to find it)
741
- claims.ts when the page contradicts the server (refused_empty, false_success)
742
- request.ts what a replayed API call may be and where it may go
743
- brief.ts splitting the app between parallel lanes
744
- lane.ts the typed report a lane hands back
745
- calibration.ts whether a lane's confidence held up; what it judged and never filed
746
- pace.ts how a run spent its time
747
- bench.ts scoring a run against the demo app's answer key
748
- policy.ts the write-policy safety net
749
- ownership.ts safe-write: which records were created in this process?
750
- uploads.ts disk uploads, fenced to the project by real path
751
- journey.ts task-ease measurement from the action log
752
- design.ts the design audit + page scoring
753
- memory.ts cross-run storage + finding dedup
754
- profiles.ts saved sign-ins: role names, where a profile lives, owner-only files, attach by role, sessionStorage restore
755
- refresh.ts the refresh broker: which values are a role's refresh tokens, the lock beside the profile, swapping a spent token
756
- scripted-login.ts a CI sign-in: env and flags, TOTP (RFC 6238) or a fixed code, which field is which, redaction
757
- signed-in.ts when a person's sign-in in the window has finished: back on the app, a new session, past any SSO round trip
758
- expiry.ts how long a saved sign-in lasts: cookie dates and JWT exp, checked before lanes start
759
- report.ts the gap ledger + report generation
760
- check.ts the check's rules, gate, report and SARIF
761
- baseline.ts visual baselines: targets.json, where each picture is kept, when one is met
762
- sarif.ts which repository file a SARIF result points at, so code scanning keeps it
763
- ci.ts a CI run's options, provider choice, caps, key redaction, tools and files
764
- export.ts which findings an export files, the inert issue it writes, the marker that dedups it
765
- provider.ts the Anthropic and OpenAI message shapes, and retries
766
- replay.ts the run as one page: steps, tasks, frames under each finding
767
- โ€ฆ collector ยท dispatch ยท fixtures ยท authloss ยท reaper
768
- scripts/ the test suites (smoke/ holds the real-browser ones)
769
- test-app/ fixtures for the real-browser smoke tests
770
- skills/scenescout/ the testing method (SKILL.md): a skill in Claude Code, served by the server everywhere else
771
- docs/how-it-works.md what happens at each stage, in diagrams
772
- docs/benchmark.md measuring whether a change made runs better
773
- docs/validation.md scorecards from runs against public open-source apps
774
- docs/adr/ why it's built this way
775
- ```
776
-
777
- > Design principle: logic that *doesn't* need Playwright lives outside `browser.ts`, so it can be unit-tested without launching a browser. That's why `fingerprint`, `policy`, `memory`, `report`, etc. are their own modules.
778
-
779
- ---
780
-
781
- ## ๐Ÿง  Design decisions
782
-
783
- **[How it works, stage by stage](docs/how-it-works.md)** โ€” diagrams of the run lifecycle, what happens inside one action, the write policy on the wire, how a violation becomes a finding, how a parallel run is split and folded, how roles hand work to each other, where a run's time goes, and how a lane's confidence is checked afterwards.
784
-
785
- **[Measuring whether a change helped](docs/benchmark.md)** โ€” the demo app's answer key, the scorecard (recall, precision, judged-not-filed, severity, calibration), and the results log of every run, including what did not help.
786
-
787
- **[Validation on public open-source apps](docs/validation.md)** โ€” runs against three well-known open-source web apps, with a scorecard for each: issues by severity, how many were real and how many were false positives, and the engine problems the runs exposed.
788
-
789
- The load-bearing choices are recorded as ADRs โ€” read the relevant one before changing a rule it covers:
790
-
791
- - [1 ยท Completion is an enforced contract, not a claim](docs/adr/0001-completion-is-a-contract-not-a-vibe.md)
792
- - [2 ยท The write policy is enforced on the wire, not in the prompt](docs/adr/0002-enforce-the-write-policy-at-the-network-layer.md)
793
- - [3 ยท A gap-ledger entry must be actionable, and suppression must be visible](docs/adr/0003-a-noisy-ledger-is-a-broken-ledger.md)
794
- - [4 ยท Findings dedup on machine signals, and a merge must never lose a finding](docs/adr/0004-dedup-on-machine-signals-not-prose.md)
795
- - [5 ยท Testable logic lives outside `browser.ts`](docs/adr/0005-keep-testable-logic-out-of-the-browser-module.md)
796
- - [6 ยท Nothing in this repo names or is tuned for a tested app](docs/adr/0006-stay-project-agnostic.md)
797
- - [7 ยท The live view is local, read-only, and leaves nothing behind](docs/adr/0007-the-live-view-is-local-read-only-and-leaves-nothing-behind.md)
798
- - [8 ยท Recording is opt-in, and a recorded run is one self-contained page](docs/adr/0008-a-recorded-run-is-evidence-and-must-be-asked-for.md)
799
- - [9 ยท A refused write is answered, not dropped](docs/adr/0009-a-refused-write-is-answered-not-dropped.md)
800
- - [10 ยท A lane's confidence is checked, not trusted](docs/adr/0010-a-confidence-is-checked-not-trusted.md)
801
- - [11 ยท A gate is deterministic, and fails only on what it can prove](docs/adr/0011-a-gate-is-deterministic-and-fails-only-on-what-it-can-prove.md)
802
- - [12 ยท A check replays saved flows and re-tests open findings, within settings whose defaults do the least harm](docs/adr/0012-a-check-replays-saved-flows-and-reports-re-tests.md)
803
- - [13 ยท What depends on a project's convention is the project's to decide](docs/adr/0013-a-convention-is-the-projects-to-decide.md)
804
- - [21 ยท Update the docs in the same pull request, unless the change has no user-facing surface](docs/adr/0021-update-the-docs-in-the-pull-request.md)
805
-
806
- ---
807
-
808
- ## ๐Ÿ”ง Development
809
-
810
- Working on SceneScout itself is the only reason to clone it:
811
-
812
- ```bash
813
- git clone https://github.com/brunoboto96/SceneScout.git scenescout && cd scenescout
814
- npm install # installs dependencies and builds
815
- npm run setup # same as `scenescout install`, but registers THIS checkout (the skill is linked, so edits are live)
816
- npm test # build + 24 suites: 22 pure-logic suites (scan, oracle, policy, โ€ฆ bench, hygiene),
817
- # then smoke and mcp-check (the server over stdio), both with real browsers
818
- npm run bench -- --all # re-score every archived benchmark run against the current answer key
819
- npm run demo # regenerate examples/ from the demo app
820
- ```
821
-
822
- Contributing? Start with [VISION.md](VISION.md) (what is in scope) and [CONTRIBUTING.md](CONTRIBUTING.md) (how changes land), then see [AGENTS.md](AGENTS.md) for the house rules โ€” chiefly: bug fixes need a regression test at the cheapest layer that can fail, keep the repo project-agnostic (ADR 6), and `npm test` must pass.
823
-
824
- ## ๐Ÿ” Security
825
-
826
- Found a way past the write policy, or another security problem? Please report it privately โ€” see [SECURITY.md](SECURITY.md).
827
-
828
- ## ๐Ÿ“„ License
829
-
830
- [MIT](LICENSE).
831
-
832
- <details>
833
- <summary><strong>Full capability list</strong> โ€” every behavior, for the curious</summary>
834
-
835
- - **Structured render-state, not pixels.** Element lists with geometry; screenshots reserved for pixel-native residue (canvas, rendering glitches). Images that failed to load are reported from the DOM, including ones whose URL answered 200 with something that is not an image.
836
- - **Diff snapshots with stable refs.** Re-snapshots return only what changed (10.7 kB โ†’ 0.7 kB on a 130-element page); old refs stay valid.
837
- - **Geometry oracles.** Overlap and off-screen defects computed from layout boxes.
838
- - **Oracles after every action.** Console errors, page errors, failed requests, HTTP 4xx/5xx drained into every tool result โ€” and DOM injection: a markup-shaped value the agent typed that later renders as an element on any page (stored or reflected XSS).
839
- - **Multi-role, genuinely concurrent.** Commands to *different* sessions run in parallel; safe-write ownership is shared, so role A can create what role B approves. The report renders a role capability matrix.
840
- - **Task ease, not just correctness.** `scout_journey` measures interaction cost, distinct screens, path, and backtracks.
841
- - **Design audit with page scores.** Two tiers (โš  measurable defects / โ†’ craft suggestions incl. AI-slop tells), per-page 0โ€“100 score persisted per route, plus an automatic overlay/modal probe on every snapshot. Shared shell scored once, separately.
842
- - **Scrolls like a user โ€” and notices when it can't.** Reports `SCROLL LOCKED` for a leaked modal scroll-lock, finds the real inner scroll pane on app-shell layouts, and flags `UNREACHABLE` controls clipped inside `overflow:hidden`.
843
- - **Uploads like a user.** Answers a styled file-chooser or sets a hidden input directly, with a valid in-memory fixture; `filePath` is fenced to the project under test; files violating `accept` are flagged at selection.
844
- - **Auth via Playwright storage states.** Expired tokens caught at attach; repeated login-bounces raise `SESSION AUTH LOST` (a session attached by role first re-attaches once from its role's latest saved profile and carries on); a bounced route is recorded as *not* covered โ€” a dead session can't certify routes it never reached.
845
- - **A trustworthy gap ledger.** Entries must be actionable (a search box or wizard sub-step isn't "form filled but never submitted"); API/download URLs never enter the route contract.
846
- - **Honest reporting.** Shared chrome counted once, stale scores marked, role matrix compares only roles that actually attempted a route.
847
- - **Cross-run written knowledge.** `scout_note` curates `.scenescout/ASSUMPTIONS.md` โ€” app model, personas, constraints, risks โ€” in prose.
848
- - **Daemon-grade robustness.** Per-tool watchdogs, orphaned-browser reaping, bounded teardown, live status via `scenescout status <project>`, and a live view of every session's page: the agent gives you its address when it attaches, or run `scenescout watch <project>` (loopback only, read-only, nothing written to disk: [ADR 7](docs/adr/0007-the-live-view-is-local-read-only-and-leaves-nothing-behind.md)).
849
-
850
- </details>
257
+ [MIT](LICENSE) licensed.