scenescout 3.23.0 โ 3.23.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +12 -0
- package/README.md +154 -747
- package/dist/engine/check-replay.js +1 -1
- package/dist/engine/check-report.js +3 -2
- package/dist/engine/dispatch.js +70 -0
- package/dist/engine/memory.js +12 -0
- package/dist/mcp-server.js +45 -6
- package/package.json +2 -2
package/README.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
# ๐ญ SceneScout
|
|
4
4
|
|
|
5
|
-
**
|
|
5
|
+
**Your coding agent, turned into an exploratory QA tester for any web app.**
|
|
6
6
|
|
|
7
7
|
Works with Claude Code ยท Cursor ยท VS Code (Copilot) ยท Codex CLI ยท Gemini CLI ยท Copilot CLI ยท Windsurf ยท any [MCP](https://modelcontextprotocol.io) client
|
|
8
8
|
|
|
@@ -10,841 +10,248 @@ Works with Claude Code ยท Cursor ยท VS Code (Copilot) ยท Codex CLI ยท Gemini CLI
|
|
|
10
10
|
[](https://www.npmjs.com/package/scenescout)
|
|
11
11
|
[](LICENSE)
|
|
12
12
|

|
|
13
|
-

|
|
14
13
|
|
|
15
|
-
[๐ Guide](docs/guide/Home.md) ยท [
|
|
14
|
+
[๐ Guide](docs/guide/Home.md) ยท [๐ What it catches](#-what-it-catches) ยท [๐ฌ QA with evidence](#-automated-qa-with-evidence) ยท [๐ Get started](#-get-started) ยท [๐ฆ CI](#-in-ci) ยท [๐ Safety](#-safe-by-default) ยท [๐ Docs](#-documentation)
|
|
16
15
|
|
|
17
16
|
</div>
|
|
18
17
|
|
|
19
|
-
|
|
18
|
+
Scripted end-to-end tests answer one question: *does this exact flow still work?* They say nothing about the rest of the app. SceneScout lets the agent you already use explore a running web app like a curious, thorough tester. It clicks, fills forms, switches roles and calls the API behind a hidden button, then writes a report of what is **broken** and what could be **better**, with a picture and evidence for every line.
|
|
20
19
|
|
|
20
|
+
Try it on any app you are allowed to test. No account, no API key, no setup:
|
|
21
|
+
|
|
22
|
+
```bash
|
|
23
|
+
npx -y scenescout http://localhost:3000
|
|
21
24
|
```
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
25
|
+
|
|
26
|
+
```mermaid
|
|
27
|
+
flowchart LR
|
|
28
|
+
A["๐ง Your coding agent<br/>Claude Code, Cursor, Copilotโฆ"] -- MCP --> S["๐ญ SceneScout<br/>browser ยท checks ยท memory"]
|
|
29
|
+
S --> App["๐ Your running app"]
|
|
30
|
+
S --> R["๐ Report<br/>plain words + pictures"]
|
|
31
|
+
S --> E["๐ฌ Evidence<br/>replay, videos, test report"]
|
|
27
32
|
```
|
|
28
33
|
|
|
29
|
-
|
|
34
|
+
## ๐ฅ Who it's for
|
|
30
35
|
|
|
36
|
+
| ๐ ๏ธ **Developers** | ๐งโ๐ผ **QA, product and non-technical teams** |
|
|
37
|
+
|---|---|
|
|
38
|
+
| Findings with the request that failed (`GET /api/orders?status=archived โ 500`), the steps, and a Playwright regression-test skeleton | Each problem in plain words first: what was done, what was expected, what happened, and a picture of the page |
|
|
39
|
+
| Next to the source, the file behind the bug and a likely fix | Visible problems shown, not described: a covered button, a broken image, text too faint to read |
|
|
40
|
+
| A deterministic gate for pull requests, with SARIF for code scanning | Journeys recorded step by step, as frames and video, so a pass is something you can watch |
|
|
41
|
+
| Runs from your editor's agent, or unattended in CI | A test report laid out by **your** template, with expected and actual results, deviations and blank sign-off rows |
|
|
31
42
|
|
|
32
|
-
##
|
|
43
|
+
## ๐ What it catches
|
|
33
44
|
|
|
34
|
-
|
|
45
|
+
A real run against the small demo app in this repository, which has bugs planted on purpose. The two you can see are outlined:
|
|
35
46
|
|
|
36
47
|
<p align="center"><img src="examples/screenshots/dashboard-annotated.png" alt="The demo app's dashboard with two defects outlined in red: 1, a yellow badge covering the All orders button; 2, the weekly chart image failing to load" width="760" /></p>
|
|
37
48
|
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
An excerpt of the report it wrote โ [read the whole thing](examples/report.md):
|
|
49
|
+
It filed twelve findings. A few from [the report](examples/report.md):
|
|
41
50
|
|
|
42
51
|
> **๐ด [HIGH] A double-click on Create order creates two orders**
|
|
43
|
-
>
|
|
44
|
-
> The submit button stays enabled while the request is in flight, and the endpoint accepts the repeat.
|
|
45
|
-
>
|
|
46
|
-
> **๐ด [HIGH] Filtering orders by Archived fails, and the page shows an empty table instead of an error**
|
|
47
|
-
> Evidence: `GET /api/orders?status=archived โ HTTP 500`
|
|
52
|
+
> `2ร click fired the same state-changing request 2ร (POST /api/orders)`
|
|
48
53
|
>
|
|
49
54
|
> **๐ด [HIGH] A clerk can approve an order by calling the endpoint the page hides from them**
|
|
50
|
-
>
|
|
51
|
-
>
|
|
52
|
-
> **๐ [MEDIUM] The "New: bulk import" badge sits on top of the All orders button** *(callout 1)*
|
|
53
|
-
> Evidence: `"All orders" overlaps "New: bulk import" (81%)` โ measured from layout boxes, no screenshot needed.
|
|
55
|
+
> `POST /api/orders/1037/approve 200 as clerk`: the button was hidden, the server did not agree.
|
|
54
56
|
>
|
|
55
|
-
>
|
|
56
|
-
>
|
|
57
|
+
> **๐ด [HIGH] Filtering orders by Archived fails, and the page shows an empty table instead of an error**
|
|
58
|
+
> `GET /api/orders?status=archived โ HTTP 500`
|
|
57
59
|
>
|
|
58
|
-
>
|
|
59
|
-
|
|
60
|
-
Every finding comes with a repro trace and a Playwright regression-test skeleton. To try it yourself, clone this repository, run `npm run demo:serve`, then `/scenescout --url http://127.0.0.1:4173` โ see [demo-app/](demo-app/). Its README lists every seeded defect and which oracle catches it.
|
|
61
|
-
|
|
62
|
-
---
|
|
63
|
-
|
|
64
|
-
## โจ Why it's different
|
|
65
|
-
|
|
66
|
-
- ๐ง **Your agent is the brain โ no API key.** The engine contains no LLM. Exploration runs on the agent and subscription you already have (Claude Code, Cursor, Copilot, Codex, Gemini CLI and others); SceneScout just gives it deterministic tools and the method for using them.
|
|
67
|
-
- ๐ **Structured scene, not pixels.** The agent reads element lists *with layout geometry*, not screenshots. Overlap and off-screen bugs are computed from boxes โ deterministic, no vision guessing. Images that failed to load are read from the DOM too. (Screenshots exist only for pixel-native residue like a canvas or a rendering glitch.)
|
|
68
|
-
- ๐ก๏ธ **Read-only by default, enforced on the wire.** Destructive actions are blocked at the network layer, not by asking the model nicely. Opt into writes only against disposable data.
|
|
69
|
-
- โ
**Completion is a contract, not a vibe.** The engine knows the app's routes and *refuses* to file an "extensive" report while any known route is unvisited, unexercised, or un-audited. "Explored a bit and stopped" is structurally impossible.
|
|
70
|
-
- ๐งญ **It remembers.** UI states are fingerprinted and stored in the project's `.scenescout/`. Run N+1 skips what run N already covered, and every run starts smarter than the last.
|
|
71
|
-
|
|
72
|
-
---
|
|
73
|
-
|
|
74
|
-
## ๐ฏ Two ways to use it
|
|
75
|
-
|
|
76
|
-
SceneScout needs only a URL. Give it the source code as well and it gets noticeably better.
|
|
77
|
-
|
|
78
|
-
| | ๐ **Next to the codebase** *(recommended)* | ๐ **Against a remote URL** |
|
|
79
|
-
|---|---|---|
|
|
80
|
-
| **You run it from** | the app's repository | any folder โ an empty `qa/` directory is fine |
|
|
81
|
-
| **It plays the role of** | a developer-tester who can read the code | a black-box QA tester, like a person with a browser |
|
|
82
|
-
| **How it finds pages** | ๐ reads routes from the source **and** follows links: file-based routing (Next.js, SvelteKit, Nuxt) and router configuration written in code (React Router, Vue Router, Angular). Routes built at runtime are not seen | ๐ follows same-origin links only โ pages nothing links to, or on another subdomain, stay unknown |
|
|
83
|
-
| **"Did we cover everything?"** | checked against the routes found in source *plus* discovered links โ an unvisited one blocks the report | checked against the pages it managed to discover |
|
|
84
|
-
| **Setup it figures out** | framework, dev command, saved Playwright logins (`playwright/.auth/`), whether the app uses `data-testid` | none โ you pass the URL, and the path to a login state if the app needs one |
|
|
85
|
-
| **What a finding looks like** | the symptom, **plus** the file behind it and a suggested fix | the symptom, a repro trace, and a regression-test skeleton |
|
|
86
|
-
| **Typical target** | `localhost` while you build | staging, a preview deploy, a client's site |
|
|
87
|
-
|
|
88
|
-
**Why the codebase helps.** The agent driving SceneScout is a coding agent, which can already read your repository. With the source at hand it knows the app's static routes before opening the browser, so coverage is measured against the real app instead of whatever happened to be linked. It can also check a suspicion against the code before reporting it: "there is no way to export this table" is a much stronger finding once the agent has confirmed no export handler exists. And when something breaks it can open the component or handler responsible and tell you *where* and *how* to fix it โ "the save button does nothing" becomes "`OrderForm` swallows the rejected promise in `onSubmit`; surface the error and re-enable the button".
|
|
89
|
-
|
|
90
|
-
**Why it still works without it.** Everything SceneScout *observes* comes from the running page โ elements, layout geometry, console and network errors, design-audit scores, task-ease measurements โ and none of that needs source code. Point it at a URL you are allowed to test and it behaves like a thorough QA tester: it explores, reproduces, and files findings with evidence.
|
|
91
|
-
|
|
92
|
-
```
|
|
93
|
-
# next to the code โ run inside the app's repository
|
|
94
|
-
/scenescout --url http://localhost:3000
|
|
95
|
-
|
|
96
|
-
# remote โ run from any folder; memory and the report are kept there
|
|
97
|
-
/scenescout --url https://staging.example.com --role ./auth/qa.json
|
|
98
|
-
```
|
|
60
|
+
> **๐ [MEDIUM] The "New: bulk import" badge sits on top of the All orders button** *(callout 1)*
|
|
61
|
+
> `"All orders" overlaps "New: bulk import" (81%)`, measured from layout boxes, no screenshot needed.
|
|
99
62
|
|
|
100
|
-
|
|
101
|
-
> Only test sites you own or are authorized to test. A remote environment is more likely to hold real data, so for a remote URL with no source the skill attaches in **`observe`** mode: nothing but `GET` requests leaves the page. The default **read-only** mode blocks `PUT`/`PATCH`/`DELETE` and destructive-looking requests, but an ordinary form submission (a plain `POST`: contact form, comment, order, signup) still reaches the server and can create a record. Say so when that is acceptable on your target. See the [safety model](#-safety-model).
|
|
63
|
+
What it looks for, on every page, after every action:
|
|
102
64
|
|
|
103
|
-
|
|
65
|
+
- ๐ **Broken layout, from geometry:** controls that overlap, sit off-screen, hide under a sticky bar or can never be scrolled into view. The CSS bugs a person spots at a glance, found without one.
|
|
66
|
+
- ๐งจ **Real breakage:** console errors, crashes, failed requests, 4xx and 5xx responses, broken images, dead-end pages.
|
|
67
|
+
- ๐ **Permission leaks:** it calls the app's own API as each role, so "the button is hidden" becomes "the server refuses it", or doesn't.
|
|
68
|
+
- ๐คฅ **Pages that lie:** "Saved!" after the server refused the save, or an empty table after the request failed.
|
|
69
|
+
- ๐ **Impatient users:** a double-click that sends the same order twice.
|
|
70
|
+
- โฟ **Accessibility and craft:** contrast, focus, labels, target sizes, spacing and type, with a 0 to 100 score per page.
|
|
71
|
+
- ๐งญ **Friction:** how many steps a task takes, and where a user had to go back.
|
|
72
|
+
- ๐ **Security smells:** typed markup that comes back as an element, and tokens posted to any window.
|
|
104
73
|
|
|
105
|
-
|
|
74
|
+
[Everything it checks](docs/guide/What-it-checks.md).
|
|
106
75
|
|
|
107
|
-
|
|
76
|
+
## ๐ฌ Automated QA with evidence
|
|
108
77
|
|
|
109
|
-
|
|
78
|
+
Save the journeys that matter, the happy paths and the ones that must fail politely, and `scenescout check` replays them on every pull request with no model involved. Recorded, each run leaves proof you can watch and hand to someone who never opens a terminal.
|
|
110
79
|
|
|
111
|
-
```
|
|
112
|
-
|
|
80
|
+
```mermaid
|
|
81
|
+
flowchart LR
|
|
82
|
+
PR["๐ Pull request"] --> C["๐ฆ scenescout check<br/>--record --video --template"]
|
|
83
|
+
F["๐ Saved journeys<br/>.scenescout/flows/*.json"] --> C
|
|
84
|
+
C --> G["โ
/ โ A gate on the pull request"]
|
|
85
|
+
C --> RP["๐ผ๏ธ replay.html<br/>every step, with frames"]
|
|
86
|
+
C --> V["๐๏ธ A video per journey"]
|
|
87
|
+
C --> TR["๐ Test report<br/>your template, SHA-256 manifest,<br/>blank sign-off rows"]
|
|
113
88
|
```
|
|
114
89
|
|
|
115
|
-
|
|
90
|
+
A journey is a few lines of JSON, written by hand or kept from a flow your agent just walked. This is [the demo's happy path](examples/flows/01-place-an-order.json):
|
|
116
91
|
|
|
92
|
+
```json
|
|
93
|
+
{
|
|
94
|
+
"name": "place an order",
|
|
95
|
+
"id": "TC-01",
|
|
96
|
+
"requirements": ["REQ-ORD-1", "REQ-ORD-2"],
|
|
97
|
+
"steps": [
|
|
98
|
+
{ "action": "navigate", "target": "/orders-new.html", "expected": "The new-order form is shown" },
|
|
99
|
+
{ "action": "type", "target": "testid=new-order-customer", "value": "Harbour Bakery", "expected": "The customer is filled in" },
|
|
100
|
+
{ "action": "type", "target": "testid=new-order-items", "value": "4", "replace": true, "expected": "The item count is 4" },
|
|
101
|
+
{ "action": "click", "target": "testid=new-order-submit", "expected": "The order is created" },
|
|
102
|
+
{ "action": "expect-request", "request": "POST /api/orders", "status": "2xx" },
|
|
103
|
+
{ "action": "expect-element", "target": "testid=new-order-created-link", "state": "visible", "expected": "A link to the new order is shown" }
|
|
104
|
+
]
|
|
105
|
+
}
|
|
117
106
|
```
|
|
118
|
-
Look at these first:
|
|
119
|
-
1. [medium] Request failed with a client error: GET /img/weekly-chart.png โ HTTP 404 (on /)
|
|
120
|
-
2. [medium] Dead end: /reports-scheduled.html: 0 controls (on /reports-scheduled.html)
|
|
121
|
-
3. [medium] Control covered by pinned chrome: button "Save notes" is COVERED by pinned chrome [order-stickybar] at this scroll position โ a click aimed at it lands on that element instead (on /order.html?id=1042)
|
|
122
|
-
|
|
123
|
-
12 pages looked at in 12 s in observe mode: 0 high ยท 6 medium ยท 2 low ยท 5 worth a look, never counted.
|
|
124
|
-
Report: scenescout-report/report.md
|
|
125
|
-
```
|
|
126
|
-
|
|
127
|
-
That is the [demo app](demo-app/). It exits 0 whatever it finds (a look, not a gate), and 2 when the address cannot be reached or the report cannot be written. After the address, `--max-routes` and `--max-minutes` raise the limits, `--mode read-only` lets a plain POST through, and `--out` names another folder. A `scenescout-report/` holding files a first look did not write is left alone, and no report it did not write is ever replaced. A first look only opens pages. To have your agent click, fill forms, compare roles and remember what it learned, set SceneScout up as below.
|
|
128
|
-
|
|
129
|
-
### ๐ฆ Prerequisites
|
|
130
107
|
|
|
131
|
-
|
|
132
|
-
|---|---|
|
|
133
|
-
| **Node** | โฅ 20 |
|
|
134
|
-
| **An MCP client** | Claude Code, Cursor, VS Code with Copilot, Codex CLI, Gemini CLI, GitHub Copilot CLI, Windsurf, or [any other](#-other-mcp-clients) |
|
|
135
|
-
| **A web app to test** | SceneScout tests a *live* app: start yours locally first (e.g. `npm run dev`, `make dev-up`), or have the URL of a deployed one you're allowed to test |
|
|
136
|
-
|
|
137
|
-
### 1๏ธโฃ Install
|
|
138
|
-
|
|
139
|
-
It is on npm. Nothing to clone:
|
|
108
|
+
Against the [demo app](demo-app/), with the three journeys in [examples/flows](examples/flows/) (one creates an order, so the check is allowed to send it):
|
|
140
109
|
|
|
141
110
|
```bash
|
|
142
|
-
npx
|
|
143
|
-
|
|
111
|
+
npx scenescout check http://127.0.0.1:4173 --flows examples/flows --mode read-only --flow-writes allow \
|
|
112
|
+
--record --video --template examples/report-template.json
|
|
144
113
|
```
|
|
145
114
|
|
|
146
|
-
|
|
115
|
+
**The journey, as it ran** (one of the videos it filmed):
|
|
147
116
|
|
|
148
|
-
|
|
117
|
+
<p align="center"><img src="examples/screenshots/journey-place-order.gif" alt="A recorded journey on the demo app: the new-order form is filled with a customer and four items, Create order is clicked, and the confirmation with a link to the new order appears" width="760" /></p>
|
|
149
118
|
|
|
150
|
-
|
|
151
|
-
/plugin marketplace add brunoboto96/SceneScout
|
|
152
|
-
/plugin install scenescout@scenescout-marketplace
|
|
153
|
-
```
|
|
119
|
+
**The replay page:** each journey with a pass or fail badge. A journey that broke opens at the step that broke, with the page as it was after every step:
|
|
154
120
|
|
|
155
|
-
|
|
121
|
+
<p align="center"><img src="examples/screenshots/check-replay.png" alt="The check's replay page: verdict FAILED, the app's address, start and end times, 24 frames and 3 videos; two journeys passed and one, see archived orders, failed at step 3" width="820" /></p>
|
|
156
122
|
|
|
157
|
-
|
|
123
|
+
<p align="center"><img src="examples/screenshots/check-replay-failed.png" alt="The failed journey opened: step 1 navigates to the orders page and passes, step 2 selects Archived and passes with an empty table, step 3 expected GET /api/orders to answer 200 and it answered 500, highlighted in red" width="760" /></p>
|
|
158
124
|
|
|
159
|
-
**
|
|
160
|
-
|
|
161
|
-
**A client that is not in that list?** [Add the server to its config by hand](#-other-mcp-clients); the test browser downloads on first use.
|
|
162
|
-
|
|
163
|
-
<details>
|
|
164
|
-
<summary>What <code>install</code> actually does</summary>
|
|
165
|
-
|
|
166
|
-
1. puts the `/scenescout` skill into `~/.claude/skills/` (or `$CLAUDE_CONFIG_DIR/skills/`) โ a `scenescout` folder it didn't create is moved aside to a `.backup-โฆ` copy, never deleted,
|
|
167
|
-
2. downloads the browser SceneScout drives (skipped if you already have it). By default that is Chromium, as two builds: the full browser for headed runs and the headless shell every other run uses. [Choose something else](#-choosing-browsers) with `--browsers`,
|
|
168
|
-
3. registers the MCP server with Claude Code at user scope. Run through `npx`, the launcher is `npx -y scenescout serve`, with the absolute path of `npx` where one sits beside node, so it works under nvm/fnm. From a clone or a global install it is the absolute node path plus that install's `dist/mcp-server.js`,
|
|
169
|
-
4. puts the `scenescout` command on your PATH, so `scenescout status`, `scenescout watch` and `scenescout doctor` work from any terminal. Run through `npx`, that is `npm install -g` of the version you just ran; from a clone it is `npm link`, so the command always runs what you last built. If npm refuses (a system-wide node usually needs `sudo` for this), the step prints the command to run by hand and the rest of the setup still counts as done: `npx -y scenescout <command>` works without it.
|
|
170
|
-
|
|
171
|
-
Re-run it any time: after moving the folder or switching node versions it refreshes the stored paths. It exits non-zero if a step the tool depends on failed, so it is safe to chain. Opt out of a step with `--no-register`, `--skip-browser` or `--no-command`.
|
|
172
|
-
|
|
173
|
-
If `claude` isn't on the PATH of the shell you ran it from, it prints the registration command instead of running it:
|
|
174
|
-
|
|
175
|
-
```bash
|
|
176
|
-
claude mcp add --scope user scenescout -- npx -y scenescout serve
|
|
177
|
-
```
|
|
125
|
+
**The test report**, laid out by [a template](examples/report-template.json) you write once: test IDs, the requirements each covers, expected and actual results, a screenshot per step, every deviation listed again for the reviewer, and the SHA-256 of each piece of evidence. SceneScout signs nothing; the sign-off rows are for your people.
|
|
178
126
|
|
|
179
|
-
|
|
127
|
+
<p align="center"><img src="examples/screenshots/test-report-results.png" alt="The test report's results table: test TC-01, place an order, covering REQ-ORD-1 and REQ-ORD-2; each step with its expected result, actual result As expected, result Pass, and a screenshot of the page after the step" width="820" /></p>
|
|
180
128
|
|
|
181
|
-
|
|
129
|
+
<p align="center"><img src="examples/screenshots/test-report-deviations.png" alt="The test report's deviations: TC-03, see archived orders, step 3 expected GET /api/orders to answer 200, actual GET /api/orders answered 500, result Fail" width="820" /></p>
|
|
182
130
|
|
|
183
|
-
|
|
184
|
-
npx -y scenescout doctor --engine # any client: node + build + browser
|
|
185
|
-
npx -y scenescout doctor # Claude Code: the above, plus the skill and the registration
|
|
186
|
-
```
|
|
131
|
+
Everything is one self-contained HTML page per report, with no scripts, nothing loaded from the network, and a layout that prints. [Recording a check](docs/ci.md#recording-a-check) and [a test report from a template](docs/ci.md#a-test-report-from-a-template) have the details, including what to keep out of the pictures.
|
|
187
132
|
|
|
188
|
-
|
|
133
|
+
## ๐บ Watch it work
|
|
189
134
|
|
|
190
|
-
|
|
135
|
+
Each exploratory run opens a live view on your machine, with one card per agent: what it is doing, the page it is on, and a feed of every action. Three agents are testing the demo app in parallel here:
|
|
191
136
|
|
|
192
|
-
|
|
137
|
+
<p align="center"><img src="examples/screenshots/live-view.png" alt="The live view during a run of three parallel agents against the demo app: one card per session, each with its role and objective, the task it is on, the tool it is running, the page it is on, a live thumbnail, and a feed of the actions it just took, tinted one colour per task" width="880" /></p>
|
|
193
138
|
|
|
194
|
-
|
|
139
|
+
You can read the report while the agents are still working, and scrub back through any session's timeline. Beside `report.md`, every run writes `report.html`, the whole run as one self-contained page; ask for a recorded run and it also keeps a frame after every action.
|
|
195
140
|
|
|
196
|
-
|
|
197
|
-
Use SceneScout to test http://localhost:3000 at medium level
|
|
198
|
-
```
|
|
141
|
+
## โจ Why it's different
|
|
199
142
|
|
|
200
|
-
|
|
143
|
+
- ๐ง **Your agent is the brain, so no extra API key.** The engine contains no model. It gives the agent you already pay for deterministic tools and the testing method to use them.
|
|
144
|
+
- ๐ **It reads structure, not pixels.** The agent sees every element with its role, state and layout box, so overlap and broken images are measured, not guessed from a screenshot.
|
|
145
|
+
- ๐ก๏ธ **Safety is enforced on the network, not requested in a prompt.** Nothing existing is changed unless you allow it, and a blocked write never reaches your server.
|
|
146
|
+
- โ
**"Done" is a contract.** The report lists everything not tested, and at the `extensive` level refuses to finish while any known page is unvisited.
|
|
147
|
+
- ๐ง **It remembers.** Each run starts from what the last one learned, and re-tests the bugs earlier runs left open.
|
|
148
|
+
- ๐ **It is measured, not asserted.** Every change to how it explores is scored against an app with planted bugs and a held-out app it is never tuned on ([the log](docs/benchmark.md)):
|
|
201
149
|
|
|
202
|
-
```
|
|
203
|
-
|
|
150
|
+
```mermaid
|
|
151
|
+
xychart-beta
|
|
152
|
+
title "Planted defects found per run, demo app (13 planted)"
|
|
153
|
+
x-axis "Run" ["0", "1", "2", "3", "4", "5", "6", "7", "8", "9", "10", "11", "12"]
|
|
154
|
+
y-axis "Found" 0 --> 13
|
|
155
|
+
bar [11, 12, 9, 10, 11, 11, 10, 10, 11, 13, 11, 11, 12]
|
|
204
156
|
```
|
|
205
157
|
|
|
206
|
-
|
|
158
|
+
<sub>Each bar is one run of eight parallel agents, re-scored against today's answer key. One run is noisy: read the trend across runs, not a single bar.</sub>
|
|
207
159
|
|
|
208
|
-
|
|
160
|
+
## ๐ Get started
|
|
209
161
|
|
|
210
|
-
**
|
|
211
|
-
|
|
212
|
-
### ๐ Signing in as a role
|
|
213
|
-
|
|
214
|
-
For an app behind SSO or MFA, sign in once yourself and let every session reuse it:
|
|
162
|
+
**1. Install** for your agent (Node 20 or newer):
|
|
215
163
|
|
|
216
164
|
```bash
|
|
217
|
-
scenescout
|
|
165
|
+
npx -y scenescout install # Claude Code: skill, MCP server and the test browser
|
|
166
|
+
npx -y scenescout install --client cursor # or vscode, codex, gemini, copilot, windsurf
|
|
167
|
+
npx -y scenescout doctor # every line should be a โ
|
|
218
168
|
```
|
|
219
169
|
|
|
220
|
-
|
|
221
|
-
|
|
222
|
-
Then `/scenescout --role admin`, or `scout_attach { role: "admin" }` from any agent. Every session attached with the same role gets its own browser built from that one login, so parallel lanes can all run as `admin`. A role with no saved login is refused with the command to run. `role` and `storageStatePath` are alternatives: pass one.
|
|
223
|
-
|
|
224
|
-
Sessions of one role share one saved login, so they share its refresh token too. An app that rotates refresh tokens and treats a second use of a spent one as theft would revoke the whole token family, and sign every session of that role out, the moment two of them refreshed with the same token. SceneScout stops that for a session attached by role. When the page is about to send a refresh token from the role's profile, the session first takes a lock beside the profile (`.scenescout/auth/<role>.json.lock`, owner-only, taken over if its holder has not touched it in 30 seconds). Holding the lock, it re-reads the profile: if another session has rotated the token in the meantime, it loads that profile into its own browser and sends the current token in place of the spent one. Once the page has stored the rotated token, the session writes its state back over the profile and releases the lock. Sessions in separate processes share the lock through the file. A refresh token is recognised by name (a cookie, a storage key, or a field inside a JSON storage value whose name contains `refresh`) and is never printed or logged. A token that is only in a cookie makes a request a refresh only when the request is plausibly the refresh call, a POST to a path named for one or an endpoint seen to rotate the cookie, so a refresh cookie scoped to `/` never holds back the app's other requests; scripts, stylesheets, images and fonts are never brokered, and a broker that cannot do its job lets the request through unchanged. See [the guide](docs/guide/Signing-in.md#sessions-that-share-a-refresh-token). An app whose sign-in renews through the identity provider's own session cookie needs none of this, since no refresh token is shared. `SCENESCOUT_REFRESH_BROKER=off` turns the broker off.
|
|
225
|
-
|
|
226
|
-
In CI, where nobody can type, `--script` signs in headless as a test user from `SCENESCOUT_LOGIN_USERNAME`, `SCENESCOUT_LOGIN_PASSWORD` and, for a one-time code, `SCENESCOUT_LOGIN_TOTP_SECRET` or a fixed code the test environment accepts in `SCENESCOUT_LOGIN_OTP_CODE` (with no password for a passwordless sign-in), and saves the same profile. No credential value is ever printed. See [signing in from CI](docs/ci.md#signing-in-from-ci) for the options and the rules: a test tenant's user, never production or a real person's account.
|
|
227
|
-
|
|
228
|
-
Before a parallel run, `scout_lane_brief` checks that the planner's saved login will outlast it: `runMinutes` (default 60) plus `expiryMarginMinutes` (default 10). It refuses only when it is sure, meaning every credential in the profile has a date, none was set for another host, and the last of them ends before the run does, and then names the `scenescout login` command to run again. A profile holds cookies other than the sign-in (analytics, preferences), so the first one to expire is reported as a warning rather than a reason to refuse, and a profile with undated credentials in it (a session cookie, or a refresh token with no expiry) is a warning that its lifetime is unknown.
|
|
229
|
-
|
|
230
|
-
---
|
|
231
|
-
|
|
232
|
-
## ๐บ Watching a run live
|
|
233
|
-
|
|
234
|
-
When a session attaches, the engine starts a small live view and hands the agent its address on a `Live view:` line, which the agent passes on to you. On a local desktop it also opens that page in your default browser as the session attaches, and opens `report.html` when `scout_report` writes it, whether or not the browser window is shown. Nothing opens in CI, over SSH, or on Linux with no display. `SCENESCOUT_OPEN` (`live`, `report`, `both` or `none`) in the server's environment chooses otherwise, and `scout_attach {open}` wins over it. `scenescout ci` opens nothing unless `SCENESCOUT_OPEN` is set. From a terminal, `scenescout watch` opens the same page. There is one card per session:
|
|
235
|
-
|
|
236
|
-
<p align="center"><img src="examples/screenshots/live-view.png" alt="The live view during a run of three parallel agents against the demo app: one card per session, each with its role and objective, the task it is on, the tool it is running, the page it is on, a live thumbnail, and a feed of the actions it just took, tinted one colour per task" width="880" /></p>
|
|
237
|
-
|
|
238
|
-
- **What it is doing:** the tool it is running and for how long, the page it is on, and a thumbnail of that page. This works for headless runs too, which have no window to look at.
|
|
239
|
-
- **What it just did:** a rolling feed of its actions, each with its target and how it turned out, with failures in red. It is the same trail a finding's repro trace uses. The engine never sees the agent's reasoning, so this is what the session *did*, not what it thought.
|
|
240
|
-
- **Stuck, not slow:** a call still running past its own tool's watchdog budget turns the card red, so a wedged session is visible without asking. A crawl legitimately runs for minutes; it is judged against the crawl's budget, not a click's.
|
|
241
|
-
- **Live stream:** switch it on for one card, or for all of them. Click a thumbnail for a close-up.
|
|
242
|
-
- **The report, as it stands:** the Report button in the top bar shows the same document `scout_report` writes at the end, rendered from the run's current state, so findings can be read while the agents are still working.
|
|
243
|
-
- **What it is for:** the close-up puts the feed beside the session's brief โ the objective it was given when it attached (`scout_attach {objective}`), and underneath it the task it is on right now (`scout_task`), which the engine requires before any tool will act. Each task tints its own block of actions, so a change of task is a change of colour; point at a block and the brief names the task those actions served.
|
|
244
|
-
- **Scrub it back:** under the page is a tick per action, coloured by task. Click one to see the frame from that moment, and `Back to live` to return. On a run that was not recorded the ticks still read the trail; they just have no picture behind them.
|
|
245
|
-
|
|
246
|
-
<p align="center"><img src="examples/screenshots/live-view-closeup.png" alt="A close-up of one session: a frame from a step picked out of the timeline, the timeline itself as a tick per action coloured by task, the feed of the session's actions in the same colours, and beside it the objective and the task it is on" width="880" /></p>
|
|
247
|
-
|
|
248
|
-
<p align="center"><img src="examples/screenshots/live-view-report.png" alt="The report opened from the live view's top bar while the run is still going: summary table, gap ledger, and the findings filed so far, each with an accordion of the screenshots taken around it" width="880" /></p>
|
|
249
|
-
|
|
250
|
-
The view is served on `127.0.0.1` only, behind a token that changes every time the engine starts. It answers `GET` and nothing else, so a viewer can watch a run but not act in it, and no frame it shows is written to disk ([ADR 7](docs/adr/0007-the-live-view-is-local-read-only-and-leaves-nothing-behind.md)) unless the run was recorded, which is asked for and off by default ([ADR 8](docs/adr/0008-a-recorded-run-is-evidence-and-must-be-asked-for.md)). A stream runs only while someone is watching it. `SCENESCOUT_LIVE=off` keeps the port closed.
|
|
251
|
-
|
|
252
|
-
**Try it with parallel agents.** The demo app has three roles and several separate areas, so a run can be split between agents. Start it with `npm run demo:serve`, then ask your agent to explore it with several agents in parallel, one role and one area each. The pictures above come from a run of three. Two things keep a parallel run efficient:
|
|
253
|
-
|
|
254
|
-
- **Each agent opens its own session when it starts, and the planner closes it once it has folded that agent's report.** An agent waiting for its turn then holds no browser. Opening every session up front leaves browsers idling while the machine runs out of memory for the agents that are working. Closing before the fold loses the lane's decisions, which have nowhere to be kept.
|
|
255
|
-
- **Slow it down to follow along.** `scout_attach {paceMs}` (or `scout_session {paceMs}` mid-run) sets a floor between actions, for when you want to watch a flow rather than let it run as fast as the page allows.
|
|
256
|
-
- **Run about as many agents at once as your machine has cores, less two.** Each one drives a real browser.
|
|
257
|
-
|
|
258
|
-
---
|
|
259
|
-
|
|
260
|
-
## ๐ฌ Recording a run, and reading it back
|
|
170
|
+
<details>
|
|
171
|
+
<summary>Other ways to install: a Claude Code plugin, the Claude Desktop extension, or by hand</summary>
|
|
261
172
|
|
|
262
|
-
|
|
263
|
-
|
|
173
|
+
- **Claude Code plugin:** `/plugin marketplace add brunoboto96/SceneScout`, then `/plugin install scenescout@scenescout-marketplace`. The command becomes `/scenescout:scenescout`.
|
|
174
|
+
- **Claude Desktop:** download `scenescout-X.Y.Z.mcpb` from the [latest release](https://github.com/brunoboto96/SceneScout/releases/latest) and open it. No terminal needed.
|
|
175
|
+
- **Any other MCP client:** add a stdio server whose command is `npx -y scenescout serve`. [Each client's config](docs/guide/Start-here.md#a-client-that-install-does-not-know).
|
|
264
176
|
|
|
265
|
-
|
|
266
|
-
or the page as it was. It is kept in `.scenescout/recordings/`, shown under the
|
|
267
|
-
finding in `report.html`, and returned with the `scout_finding` result, so a
|
|
268
|
-
chat client shows the evidence the moment it is filed. Pictures are bounded in
|
|
269
|
-
size and in how many reach the conversation, and a CI job keeps them on file
|
|
270
|
-
only; `SCENESCOUT_EVIDENCE` and `scout_attach {evidence}` change that
|
|
271
|
-
([configuration reference](docs/guide/Configuration-reference.md#environment-variables)).
|
|
177
|
+
</details>
|
|
272
178
|
|
|
273
|
-
|
|
179
|
+
**2. Start your app**, then a fresh session of your agent, and ask:
|
|
274
180
|
|
|
181
|
+
```text
|
|
182
|
+
Use SceneScout to test http://localhost:3000
|
|
275
183
|
```
|
|
276
|
-
Use SceneScout to test http://localhost:3000, record the run
|
|
277
|
-
```
|
|
278
|
-
|
|
279
|
-
or, on the tool directly, `scout_attach {record: true}`. `SCENESCOUT_RECORD=on` in
|
|
280
|
-
the server's environment records every run. A CI gate records too:
|
|
281
|
-
`scenescout check --record` writes `replay.html`, every journey step by step with
|
|
282
|
-
its frames ([recording a check](docs/ci.md#recording-a-check)), and
|
|
283
|
-
`--template <file.json>` also writes it up as a test report laid out as a template
|
|
284
|
-
says, with expected and actual results, deviations, blank sign-off rows and a
|
|
285
|
-
SHA-256 manifest of the evidence ([test reports](docs/ci.md#a-test-report-from-a-template)).
|
|
286
|
-
|
|
287
|
-
Then `scout_report` writes two files side by side in `.scenescout/`:
|
|
288
|
-
`report.md` as always, and `report.html` โ the whole run as one self-contained
|
|
289
|
-
page. It opens from the file system with nothing running, needs no network, and
|
|
290
|
-
holds:
|
|
291
|
-
|
|
292
|
-
- **The report**, rendered from the same Markdown: the plain-language view
|
|
293
|
-
first (each problem's steps, what was expected, what happened and its
|
|
294
|
-
picture), with each problem's technical detail one click away.
|
|
295
|
-
- **The screenshots around each finding**, in an accordion under it, from the
|
|
296
|
-
session that filed it.
|
|
297
|
-
- **Every session's trail**, in the blocks its tasks made, each step with the
|
|
298
|
-
page as it was at that moment.
|
|
299
|
-
|
|
300
|
-
<p align="center"><img src="examples/screenshots/run-page.png" alt="The saved copy of a run, opened from the file system with nothing running: a finding with its evidence accordion open, showing the four screenshots taken around it with the action and time under each" width="880" /></p>
|
|
301
184
|
|
|
302
|
-
|
|
303
|
-
and sends you there when the run ends โ so the address survives a refresh
|
|
304
|
-
instead of a panel over a dead board.
|
|
185
|
+
In Claude Code there is a command too: `/scenescout --url http://localhost:3000 --level medium`. On its own, `/scenescout` asks you four plain questions instead: where the app is, how you sign in, what to check, and whether it holds real data.
|
|
305
186
|
|
|
306
|
-
**
|
|
307
|
-
project's folder, and the secret redaction that protects everything else the
|
|
308
|
-
engine writes cannot read a picture. That is why it is off unless asked for,
|
|
309
|
-
capped per session, and written only under `.scenescout/`, which ignores itself
|
|
310
|
-
so `git add -A` in the tested project cannot pick the frames up. The reasoning is in
|
|
311
|
-
[ADR 8](docs/adr/0008-a-recorded-run-is-evidence-and-must-be-asked-for.md).
|
|
187
|
+
**3. Read the report** in `.scenescout/report.md`. It opens in plain words (each problem, its steps, what was expected and what happened), with the technical detail one click away. [A complete example](examples/report.md).
|
|
312
188
|
|
|
313
|
-
|
|
189
|
+
**No app handy?** Clone this repository and run `npm run demo:serve`: the [demo app](demo-app/) starts on `http://127.0.0.1:4173`, and its README lists every planted bug.
|
|
314
190
|
|
|
315
|
-
|
|
191
|
+
**Behind a sign-in?** Sign in once yourself, SSO and MFA included, and every session reuses it:
|
|
316
192
|
|
|
317
|
-
|
|
318
|
-
|
|
319
|
-
```
|
|
320
|
-
scan โโโถ attach โโโถ crawl โโโถ investigate โโโถ measure โโโถ report
|
|
321
|
-
โ โ โ โ โ โ
|
|
322
|
-
routes browser every route reproduce & journeys + gap-checked
|
|
323
|
-
& auth (r/o) in ONE call file findings design audit markdown
|
|
193
|
+
```bash
|
|
194
|
+
npx -y scenescout login http://localhost:3000 --role admin # then: /scenescout --role admin
|
|
324
195
|
```
|
|
325
196
|
|
|
326
|
-
|
|
327
|
-
2. **Attach** a browser (read-only unless you said otherwise).
|
|
328
|
-
3. **Crawl** every known route in a *single* call โ per-route HTTP status, element counts, oracle violations, dead ends.
|
|
329
|
-
4. **Investigate** what the crawl flagged: navigate, snapshot, reproduce, file a structured finding.
|
|
330
|
-
5. **Measure** task ease (`scout_journey`) and design quality (`scout_design_audit`) on representative pages.
|
|
331
|
-
6. **Report** โ the engine checks the gap ledger and writes `.scenescout/report.md`.
|
|
332
|
-
|
|
333
|
-
Snapshots are cheap: re-snapshotting a route returns only *what changed*, with stable refs (measured on a 130-element page: 10.7 kB โ 0.7 kB).
|
|
334
|
-
|
|
335
|
-
---
|
|
197
|
+
[Signing in](docs/guide/Signing-in.md) covers roles, expiry and scripted sign-in for CI.
|
|
336
198
|
|
|
337
|
-
##
|
|
199
|
+
## ๐ฆ In CI
|
|
338
200
|
|
|
339
|
-
|
|
340
|
-
|
|
341
|
-
| Phase | Tools | What they do |
|
|
201
|
+
| | What it does | Needs a model? |
|
|
342
202
|
|---|---|---|
|
|
343
|
-
|
|
|
344
|
-
|
|
|
345
|
-
|
|
|
346
|
-
|
|
|
347
|
-
| **Act** | `scout_click` `scout_type` `scout_select` `scout_upload` `scout_press` `scout_scroll` `scout_navigate` `scout_back` `scout_run_plan` | Drive the UI like a user; `scout_run_plan` batches a whole mechanical sequence into one call |
|
|
348
|
-
| **Assess** | `scout_design_audit` `scout_journey` | Score a page's craft/a11y/consistency; measure how hard a task is to complete |
|
|
349
|
-
| **Record** | `scout_note` `scout_finding` `scout_resolve` `scout_report` | Curate durable notes; file deduped findings; mark fixes; write the report, and on a recorded run the whole run as one page |
|
|
350
|
-
| **Answer tickets** | `scout_tickets` `scout_criterion` | Read the acceptance criteria in pasted or uploaded tickets; record each criterion as passed, failed (with the findings that show it) or not tested (and why), with a confidence |
|
|
351
|
-
| **Re-test** | `scout_verify` | List the findings earlier runs left open, worst route first, and record whether each is gone, still present, or changed |
|
|
352
|
-
| **Split the work** | `scout_lane_brief` `scout_lane_report` | Divide the app between parallel agents by whole module, each with its own landing route and rules; fold what each hands back as one typed JSON object, and name any defect it judged but never filed |
|
|
353
|
-
| **Close** | `scout_close` | Tear down one session or all |
|
|
354
|
-
|
|
355
|
-
A few that punch above their weight:
|
|
356
|
-
|
|
357
|
-
- **`scout_crawl`** โ the entire breadth pass in one tool call. No visiting routes one-by-one.
|
|
358
|
-
- **`scout_run_plan`** โ up to 20 actions (fill form โ submit โ check) with semantic targets (`testid=โฆ`, `text=โฆ`), aborting at the first anomaly.
|
|
359
|
-
- **`scout_journey`** โ wraps one goal and reports interaction count, screens seen, and **backtracks**; an abandoned journey is a finding no passing E2E suite can produce.
|
|
360
|
-
- **`scout_upload`** โ generates a *valid* in-memory fixture (real PDF/PNG, kind inferred from `accept`) so file-upload flows stop being a blind spot.
|
|
361
|
-
- **`scout_click {clicks: 2}`** โ the impatient-user probe: states whether a double-click fired the same state-changing request twice (the classic double-submit bug).
|
|
362
|
-
- **`scout_request`** โ calls the app's own API as the session, so "the button is hidden" becomes "the server refuses it" (or doesn't).
|
|
363
|
-
|
|
364
|
-
Beyond crashes and HTTP errors, two oracles catch a page **contradicting the server**: `refused_empty` (a list request was refused and the page shows its empty state with no error) and `false_success` (a save was refused and the page says it worked). A third, `dom_injection`, reports a typed markup value coming back as an element on any page any session opens. A fourth, `postmessage_token`, reports a page calling `postMessage` with targetOrigin `"*"` on a message that carries a token (a JWT, a `Bearer` value, or an opaque value under a key such as `access_token`): the report names where in the message it was, its shape and its first four characters, never the token.
|
|
365
|
-
|
|
366
|
-
---
|
|
367
|
-
|
|
368
|
-
## ๐ Test levels
|
|
369
|
-
|
|
370
|
-
Each level is a contract. `scout_report` enforces what the engine can see for itself โ routes visited, pages audited, and at `extensive` an empty gap ledger (every route exercised and audited, every filled form submitted, a completed journey, two roles) โ and the report's gap ledger discloses the rest.
|
|
371
|
-
|
|
372
|
-
| Level | What it guarantees | Rough size |
|
|
373
|
-
|---|---|---|
|
|
374
|
-
| `minimal` | Every route visited, โฅ1 design audit, key journeys as plans, crawl problems triaged. Remaining gaps **disclosed**. | ~40 actions |
|
|
375
|
-
| `medium` *(default)* | minimal + design audits across several routes (enforced) + every element class exercised + every form submitted valid **and** invalid (asked of the agent) | ~150 actions |
|
|
376
|
-
| `extensive` | medium + fuzzing, back/refresh/deep-link resilience, keyboard-only pass, a journey per module, โฅ2 roles compared, anonymous auth-surface walk. **Refuses to finalize while any gap in the ledger remains.** | budget-capped |
|
|
377
|
-
|
|
378
|
-
That refusal *is* the guarantee: an extensive report can only exist when nothing the engine can measure was left untested. Fuzzing, the keyboard pass and the auth-surface walk are the agent's to do; the engine cannot see whether they were done well.
|
|
379
|
-
|
|
380
|
-
---
|
|
381
|
-
|
|
382
|
-
## ๐ Safety model
|
|
203
|
+
| [`scenescout check`](docs/guide/Ways-to-use-it.md#scenescout-check-a-gate-in-ci) | A deterministic gate: measures every page, replays your saved journeys and visual baselines, fails only on what it can prove; recorded, it leaves the [evidence above](#-automated-qa-with-evidence) | No |
|
|
204
|
+
| [`scenescout ci`](docs/guide/Ways-to-use-it.md#scenescout-ci-an-unattended-exploratory-run) | An unattended exploratory run, driven by the Anthropic or OpenAI API. It reports and never fails the build | An API key |
|
|
205
|
+
| [`/scenescout qa`](docs/guide/Ways-to-use-it.md#scenescout-qa-on-a-pull-request) | A comment on a pull request that tests its preview deploy and replies with the results; `/scenescout qa check` runs your own check instead | An API key (`qa check`: none) |
|
|
206
|
+
| [`scenescout export`](docs/guide/Ways-to-use-it.md#filing-findings-as-issues) | Files the findings as GitHub or Jira issues, each once | No |
|
|
383
207
|
|
|
384
|
-
|
|
385
|
-
- ๐ข **`read-only` by default.** Destructive-labeled elements (delete/revoke/archive/โฆ) **and** all `PUT/PATCH/DELETE` + destructive `POST`s are blocked at the network layer โ see [`src/engine/policy.ts`](src/engine/policy.ts). Non-destructive `POST`s are allowed, because submitting forms is how a tester finds validation bugs โ so read-only means *nothing existing is changed or removed*, not *nothing is ever created*.
|
|
386
|
-
- ๐ก **`safe-write`** (`--safe-write`) lets the agent create data and edit/delete **only what it created** this run โ never pre-existing records.
|
|
387
|
-
- ๐ด **`destructive`** (`--allow-destructive`) allows everything, and only ever when *you* confirm the environment is disposable. The skill will never choose this itself.
|
|
388
|
-
- ๐ Findings, memory, and reports live in a `.scenescout/` folder in the project. A client with no project folder, such as a desktop chat, gets one folder per tested site under `Documents/SceneScout/<host>/` by default, and the attach says where; `SCENESCOUT_PROJECTS_DIR` moves it. The folder ignores itself in git, so a stray `git add -A` never commits test data.
|
|
389
|
-
|
|
390
|
-
A `๐ก WRITE-POLICY blocked` notice is the safety net doing its job, not an app bug. The server never sees a blocked request, but a page's own `fetch` or XHR is answered with a `403` in its place rather than dropped, so the page's handling of a refusal really runs: a page that then claims success is reported as a `false_success` ([ADR 9](docs/adr/0009-a-refused-write-is-answered-not-dropped.md)).
|
|
391
|
-
|
|
392
|
-
A control is judged by its own label: a dropdown by the option picked, a row by its own text rather than the buttons inside it, and "Discard changes" on an unsent form is allowed. When a page asks to confirm leaving unsent input, the result says so; `observe` and `read-only` stay unless the call passes `leave: true`. See the [safety model](docs/guide/Safety-model.md).
|
|
393
|
-
|
|
394
|
-
---
|
|
395
|
-
|
|
396
|
-
## ๐ What you get
|
|
397
|
-
|
|
398
|
-
`.scenescout/report.md` and `report.html` open **In plain words**: a short summary, then each problem this run found, worst first, with its impact, the steps that led to it, what was expected, what happened and a picture when there is one. The technical detail (id, category, evidence, route) stays one click away.
|
|
399
|
-
|
|
400
|
-
`.scenescout/report.md` โ a deduplicated, worst-first report with:
|
|
401
|
-
|
|
402
|
-
- ๐ **Findings** with repro traces and generated Playwright regression-test skeletons.
|
|
403
|
-
- ๐ **Worth a look** โ observations that are defects only under a convention of your project the run cannot see (a spacing scale, link styling in navigation, test ids on every control), each naming that convention. Listed below the findings and not counted as defects ([ADR 13](docs/adr/0013-a-convention-is-the-projects-to-decide.md)).
|
|
404
|
-
- ๐ฏ **Page scores** (0โ100: a11y ยท craft ยท consistency ยท task-clarity), ranked worst-first, with stale scores from old runs marked as such.
|
|
405
|
-
- ๐ฅ **A role capability matrix** โ what each role could and couldn't reach.
|
|
406
|
-
- ๐งพ **A gap ledger** โ everything *not* done, so the report is honest about its own coverage.
|
|
407
|
-
- โฑ๏ธ **How the run was paced** โ how closely each session kept working, and apart from that, how long finished lanes held their browsers waiting to be collected, so neither hides the other.
|
|
408
|
-
- ๐ฏ **How well the lanes judged** โ on a parallel run, whether the confidence each lane stated matched what the project went on to file, beside what later re-tests found ([ADR 10](docs/adr/0010-a-confidence-is-checked-not-trusted.md)).
|
|
409
|
-
|
|
410
|
-
`.scenescout/report.html` โ the same report as one self-contained page, with every session's trail beside it. Each finding shows a picture of the element it is about, or of the page as it was. A [recorded run](#-recording-a-run-and-reading-it-back) also shows the screenshots around each finding.
|
|
411
|
-
|
|
412
|
-
๐ Watch a run live: `npx scenescout watch`, or `npx scenescout status <project-path>` for the same information as text.
|
|
413
|
-
|
|
414
|
-
---
|
|
415
|
-
|
|
416
|
-
## ๐ฆ In CI: a deterministic check
|
|
417
|
-
|
|
418
|
-
An exploratory run is driven by a model, so two runs never find exactly the same things. That suits a report, but not a gate. `scenescout check` is the part that needs no model. It visits the start URL, the project's scanned routes and every same-origin link it finds, and measures each page:
|
|
419
|
-
|
|
420
|
-
- HTTP and page errors
|
|
421
|
-
- layout geometry (covered, clipped and overlapping controls, blocking overlays)
|
|
422
|
-
- broken images
|
|
423
|
-
- controls with no name, and fields whose only label is a placeholder
|
|
424
|
-
- contrast and focus
|
|
425
|
-
- pages with no way out
|
|
426
|
-
|
|
427
|
-
Some of what it measures is a defect only under a convention the check cannot see: paddings off a 4px grid, and links styled like body text. Those are listed under **Worth a look**, each with the convention that would make it a defect. They are never counted and never fail the gate, at any `--fail-on`; SARIF reports them at level `note` ([ADR 13](docs/adr/0013-a-convention-is-the-projects-to-decide.md)).
|
|
428
|
-
|
|
429
|
-
```bash
|
|
430
|
-
npx scenescout check http://127.0.0.1:3000 --fail-on high
|
|
431
|
-
```
|
|
432
|
-
|
|
433
|
-
| Exit code | Meaning |
|
|
434
|
-
|---|---|
|
|
435
|
-
| 0 | Passed the gate |
|
|
436
|
-
| 1 | Failed it: something at the `--fail-on` severity or worse |
|
|
437
|
-
| 2 | Could not run, or not all of it: a bad argument, an app that never answered, only the sign-in page reached, a saved flow that is not valid, or a flow step the write policy refused |
|
|
438
|
-
|
|
439
|
-
It writes `report.md`, `check.sarif` (for code-scanning dashboards) and `check.json` to `.scenescout/check/`, and on GitHub Actions it also puts the report on the job's summary page.
|
|
440
|
-
|
|
441
|
-
On GitHub Actions, this repository is also an action that installs everything and keeps the results:
|
|
208
|
+
A gate on every pull request that keeps its evidence, as a GitHub Action:
|
|
442
209
|
|
|
443
210
|
```yaml
|
|
444
|
-
- uses: brunoboto96/SceneScout@v3
|
|
211
|
+
- uses: brunoboto96/SceneScout@v3
|
|
445
212
|
with:
|
|
446
213
|
url: http://127.0.0.1:3000
|
|
214
|
+
record: on
|
|
215
|
+
video: on
|
|
216
|
+
template: tests/report-template.json
|
|
447
217
|
```
|
|
448
218
|
|
|
449
|
-
[docs/ci.md](docs/ci.md) has
|
|
450
|
-
|
|
451
|
-
With the default settings its saved flows send no HTTP write (they replay under observe's rule), and its crawl runs under `--mode observe` or `read-only`; `--flow-writes allow` lets flows write as `--mode` allows. By default it fails only on facts that mean a page is broken: a page that did not load, an uncaught exception, a 5xx, a failure shown as success. Other options:
|
|
452
|
-
|
|
453
|
-
- `--fail-on medium` or `low` makes the gate stricter.
|
|
454
|
-
- `--ignore <rule>` drops a rule, worth-a-look rules included.
|
|
455
|
-
- `--paths /a,/b` checks only those pages.
|
|
456
|
-
- `--storage-state <file>` checks while signed in.
|
|
457
|
-
- `--flows <dir>` or `off` chooses which saved flows to replay; `--retest off` skips re-testing open findings.
|
|
458
|
-
- `--flow-writes never|allow` (default `never`): `never` replays flows under observe's rule whatever `--mode` says; `allow` replays them under `--mode`, so in `read-only` a flow's form submissions are sent to the target on every run.
|
|
459
|
-
- `--on-refused-step report|stop` (default `report`): `report` marks a flow whose step was refused "could not run", keeps every other verdict and exits 2; `stop` exits 2 at that step with no results.
|
|
460
|
-
- `--gate-retests never|high|all` (default `high`): which still-reproducing re-tested findings fail the gate.
|
|
461
|
-
- `--baseline off|compare|update` (default `off`), with `--baselines <dir>` and `--baseline-threshold <percent>` (default `0.1`, so small anti-aliasing noise between machines passes): visual baselines, below.
|
|
462
|
-
- `--record` (or `SCENESCOUT_RECORD=on`; the action's `record: on`) keeps a frame after each route visit and each saved-flow step and writes `replay.html` beside the report: each role, each journey with a pass or fail badge, each step with its caption, result and frame, the first failing step highlighted. `--video` (the action's `video: on`) adds a WebM of each journey, played on that page beside its steps. Both off by default. [docs/ci.md](docs/ci.md#recording-a-check) covers size, privacy and publishing it.
|
|
463
|
-
|
|
464
|
-
The defaults are what an unconfigured check does, for a first try or an AI agent running it unattended: its flows send no HTTP write and it never silently hides a result. Each setting is a choice for the project; the report and `check.json` print the values a check ran with.
|
|
465
|
-
|
|
466
|
-
`scenescout check --help` lists every option. Why the defaults are what they are: [ADR 11](docs/adr/0011-a-gate-is-deterministic-and-fails-only-on-what-it-can-prove.md).
|
|
467
|
-
|
|
468
|
-
It also replays the flows saved in `.scenescout/flows/*.json`, with no model: the steps `scout_run_plan` takes (navigate, click, type, select, press, and upload, which attaches a generated file) plus `expect-text`, `expect-element` (a control is visible, hidden, enabled, disabled, checked or unchecked), `expect-url` and `expect-request`, and `repeat` to run actions until a condition holds (page through a document until Continue is enabled). A flow can name the `role` it runs as, signed in with a profile `scenescout login --role` saved, so a journey that passes between people (one submits, another approves) is a sequence of flows. Values can come from the environment, `${env:NAME}`, so a code or a password stays in the CI's secret store and is masked in everything the check writes. A flow whose step breaks fails the gate, naming the flow and the step. And it re-tests the open findings earlier runs left in the project's memory that a page load can reproduce, reporting each as still reproducing or possibly fixed; by default a finding filed high that still reproduces fails the gate. [docs/ci.md](docs/ci.md#saved-flows) has the flow format; [ADR 12](docs/adr/0012-a-check-replays-saved-flows-and-reports-re-tests.md) says why it works this way.
|
|
469
|
-
|
|
470
|
-
With `--baseline compare` it also holds pages and elements to approved pictures: list them in a `targets.json`, take the baselines once with `--baseline update`, and a later check that finds one changed past `--baseline-threshold` fails the gate with the share of pixels changed and a diff picture beside the report. Baselines are kept per browser, in a git-ignored folder unless `--baselines` names one the project commits, and only `--baseline update` ever writes one. [The guide](docs/guide/Ways-to-use-it.md#visual-baselines) has the details; [ADR 19](docs/adr/0019-a-visual-baseline-changes-only-when-asked.md) says why.
|
|
471
|
-
|
|
472
|
-
Beyond those flows it explores nothing and fills no forms. That is the exploratory run's job, and its findings belong in a report, not a gate.
|
|
473
|
-
|
|
474
|
-
## ๐ค In CI: an unattended exploratory run
|
|
475
|
-
|
|
476
|
-
`scenescout ci` runs the exploratory side in a CI job, with no person and no coding agent. A model reached through its API drives the same `scout_*` tools by the same method, and the run ends in the ordinary report:
|
|
477
|
-
|
|
478
|
-
```bash
|
|
479
|
-
export OPENAI_API_KEY=โฆ # or ANTHROPIC_API_KEY; read from the environment only
|
|
480
|
-
npx scenescout ci http://127.0.0.1:3000
|
|
481
|
-
```
|
|
482
|
-
|
|
483
|
-
- **It reports and never gates.** Exit 0 when the run ran, whatever it found; exit 2 when it could not run (no key, a key the API refused, an app that never answered). Two runs of the same app find different things, so a finding is something to read, never a reason to fail a build. `scenescout check` is the gate.
|
|
484
|
-
- **Providers:** the Anthropic Messages API (default model `claude-sonnet-5`) or the OpenAI Responses API (default `gpt-6-luna`), chosen by which key is set; with both set, `--provider` decides. `--model` and `--effort` (default `low`) override; `--base-url` points at another endpoint that implements the same API.
|
|
485
|
-
- **Caps:** at most 80 model turns, 3,000,000 tokens and 20 minutes (`--max-turns`, `--max-tokens`, `--max-minutes`), shared by two model loops that each explore their own part of the app (`--lanes`, default 2). The first cap reached ends the exploration; the report is still written, and says which cap ended it. On the benchmark's demo app a run at these defaults cost about $0.03 on `gpt-6-luna` ([docs/benchmark.md](docs/benchmark.md#choosing-the-defaults-issue-419)).
|
|
486
|
-
- **Mode:** `read-only` by default; `--mode observe` sends no form at all, `--mode safe-write` lets the run create records and change only the ones it created. `--mode destructive` runs only with `--allow-destructive` as well.
|
|
487
|
-
- **Duplicates:** when the dedup rule keeps a filed finding apart, the run's model is asked at its lowest effort whether it is one already open on the same page, and merges it on a "same", keeping the filing's title, category, severity and evidence under that finding. The two findings' titles, categories and evidence, and the page's path, are sent; `--dedup rule` turns it off ([ADR 17](docs/adr/0017-a-model-judges-only-the-merges-the-rule-misses.md)).
|
|
488
|
-
- **Starting from an earlier run (opt-in):** every run that writes its report leaves a record in `ci.json` and the project's memory: the routes it worked on, its steps, and what it left on each route. `--from-run <ci.json or project directory>` continues where that run left off (routes it never worked on first, then the ones it left work on, with exactly which forms, options and controls to take first), and `--from-run-mode replay` follows its routes and steps in order. `scout_lane_brief {fromRun, fromRunMode}` does the same for a parallel run. On the demo app a chain of continued runs found 8 defects against 7 for fresh runs at the default budget, and 4, 3 and 4 against 4 at a budget cut to stand in for a larger app, within the noise, so it stays off unless asked for ([the measurement](docs/benchmark.md#starting-from-an-earlier-run-issue-418)).
|
|
489
|
-
- **Output**, in `.scenescout/ci/` (or `--out`): `report.md` and `report.html` (the report an agent's run writes), `summary.md` (also appended to the GitHub job summary), `ci.json` and `ci.sarif`, with a usage line: turns, tokens, time and an estimated cost where the model's price is known (`--price-in`, `--price-out` give one for any model).
|
|
219
|
+
[docs/ci.md](docs/ci.md) has complete workflows, every option, and the same check on GitLab CI, CircleCI or any shell.
|
|
490
220
|
|
|
491
|
-
|
|
221
|
+
## ๐ Safe by default
|
|
492
222
|
|
|
493
|
-
|
|
494
|
-
|
|
495
|
-
## ๐ฎ Filing findings as issues
|
|
496
|
-
|
|
497
|
-
`scenescout export` turns the project's open findings into GitHub or Jira issues, where the team already works. It reads `.scenescout/memory.json`, so it runs after an interactive run, `scenescout ci` or anything else that wrote findings:
|
|
498
|
-
|
|
499
|
-
```bash
|
|
500
|
-
export GH_TOKEN=โฆ # or GITHUB_TOKEN; for Jira, JIRA_EMAIL and JIRA_API_TOKEN. Read from the environment only
|
|
501
|
-
npx scenescout export --to github --repo owner/app # a dry run: lists what it would file
|
|
502
|
-
npx scenescout export --to github --repo owner/app --yes # files it
|
|
503
|
-
npx scenescout export --to jira --jira-url https://your-site.atlassian.net --jira-project QA --yes
|
|
504
|
-
```
|
|
505
|
-
|
|
506
|
-
- **Each finding once.** Every issue carries the `scenescout` label and a marker with the finding's id. Before filing, the export reads the labelled issues, open or closed, and skips every finding already filed, naming its issue, so a second export of the same run files only what the first left over the cap. A closed won't-fix is not filed again; `--refile-closed` files a finding again when its issue is closed.
|
|
507
|
-
- **Jira issues are kept up to date and linked to the ticket.** A later export updates an open Jira issue instead of filing another, leaving text someone edited in Jira as they wrote it. The finding's picture is attached, and a finding that fails a ticket's acceptance criterion is linked to that ticket (`--jira-link-type`, default `Relates`).
|
|
508
|
-
- **A dry run unless `--yes`**, and at most `--max-issues` (default 20) per export; the next export files the rest. `--min-severity`, `--only <ids>` and `--include-worth-a-look` choose what goes.
|
|
509
|
-
- **Inert issues.** Titles, descriptions, steps and evidence come from the run and the app's pages, so no `@mention`, link, `#123` reference, HTML or Markdown in them does anything.
|
|
510
|
-
- **Severity** becomes a label on GitHub and a priority in Jira; `--severity-map` renames them or turns them off. **Screenshots** from a recorded run are attached in Jira; GitHub's API takes no uploads, so a GitHub issue names the frames in the run's `.scenescout/` folder.
|
|
511
|
-
- **Credentials** are never printed. Every request has a timeout, a short rate limit is waited out, failed reads are retried with backoff, and redirects are refused. Jira's search can take a little while to show a new issue, so leave a few minutes between two exports to the same Jira project.
|
|
512
|
-
|
|
513
|
-
[The guide](docs/guide/Ways-to-use-it.md#filing-findings-as-issues) has the details and a GitHub Actions step.
|
|
514
|
-
|
|
515
|
-
---
|
|
516
|
-
|
|
517
|
-
## ๐ฉบ Troubleshooting
|
|
518
|
-
|
|
519
|
-
Run `npx -y scenescout doctor` first โ it checks every setup item below (everything but the last row, which is about your app) and prints the fix.
|
|
520
|
-
|
|
521
|
-
| Symptom | Cause and fix |
|
|
223
|
+
| Mode | What may leave the page |
|
|
522
224
|
|---|---|
|
|
523
|
-
|
|
|
524
|
-
|
|
|
525
|
-
|
|
|
526
|
-
|
|
|
527
|
-
| *"โฆ build has not been downloaded yet"* on attach | The attach downloads a missing browser itself, once, except in CI or with `SCENESCOUT_BROWSER_DOWNLOAD=off`; there, or when that download failed, it names the command to run. Run the command the message names, for example `npx -y scenescout install --browser-only --browsers firefox`. On Linux, system libraries may be missing too: `npx playwright install --with-deps chromium`. |
|
|
528
|
-
| Tools broke after moving the folder or changing node version | The registration stores absolute paths. `npx -y scenescout install` refreshes them. |
|
|
529
|
-
| Attach fails or every route lands on the login page | Your app isn't running at `--url`, or the `--role` session has expired. For a saved login, run `scenescout login <url> --role <name>` again; for a storage-state file, regenerate it the way your project's Playwright setup does. |
|
|
530
|
-
|
|
531
|
-
### โฌ๏ธ Upgrading from an older version
|
|
532
|
-
|
|
533
|
-
- **Tools are now `scout_*`.** Up to v0.23 they were prefixed `ft_`. The rename happened before the first npm release, with no aliases, so an agent's context carries one tool list rather than two. Re-run `npx -y scenescout install` so the installed skill matches the server.
|
|
534
|
-
- **Earlier names.** This tool was previously called SceneCraft (and, before that, frontend-tester). `scenescout install` cleans up after both: it removes the old skill link and the old `scenecraft` MCP registration when they point at this install, and the first attach in a project moves its `.scenecraft/` memory folder to `.scenescout/` so earlier coverage and findings carry over.
|
|
535
|
-
|
|
536
|
-
### ๐งน Uninstall
|
|
537
|
-
|
|
538
|
-
```bash
|
|
539
|
-
# Claude Code
|
|
540
|
-
claude mcp remove --scope user scenescout
|
|
541
|
-
rm -rf ~/.claude/skills/scenescout
|
|
542
|
-
# Codex / Gemini / Copilot CLI
|
|
543
|
-
codex mcp remove scenescout # likewise: gemini mcp remove โฆ, copilot mcp remove โฆ
|
|
544
|
-
```
|
|
545
|
-
|
|
546
|
-
For Cursor, Windsurf and VS Code, delete the `scenescout` entry from the client's MCP server list.
|
|
225
|
+
| ๐ต `observe` | Reads only. The default for a first look, for `scenescout check`, and for an agent's run on a remote site with no source |
|
|
226
|
+
| ๐ข `read-only` | Reads and ordinary form posts; nothing existing is changed or deleted. The default for an agent's run on a local app |
|
|
227
|
+
| ๐ก `safe-write` | Creates records, and edits or deletes only the ones it created |
|
|
228
|
+
| ๐ด `destructive` | Everything. Only when you say the data is disposable; the agent never picks it |
|
|
547
229
|
|
|
548
|
-
|
|
230
|
+
The policy sits on the network, so a blocked request never reaches your server. The page gets a refusal instead, which is how SceneScout catches a page that claims success anyway. A saved journey that creates something, like the order above, runs only when the check is told `--flow-writes allow`, against a test environment. Findings, memory and reports stay in a `.scenescout/` folder that keeps itself out of git.
|
|
549
231
|
|
|
550
|
-
|
|
551
|
-
|
|
552
|
-
## ๐ Choosing browsers
|
|
553
|
-
|
|
554
|
-
`install` downloads Chromium and nothing else unless you ask. `--browsers` takes one name, a comma-separated list, or `all`:
|
|
555
|
-
|
|
556
|
-
| `--browsers` | What is downloaded | About, on disk |
|
|
557
|
-
|---|---|---|
|
|
558
|
-
| `chromium` *(default)* | the full browser and the headless shell | 550 MB |
|
|
559
|
-
| `chromium-headless-shell` | the headless shell only: every run works except `headed` | 200 MB |
|
|
560
|
-
| `firefox` | Firefox | 270 MB |
|
|
561
|
-
| `webkit` | WebKit, the engine behind Safari | 290 MB |
|
|
562
|
-
| `all` | Chromium, Firefox and WebKit | 1.1 GB |
|
|
563
|
-
|
|
564
|
-
```bash
|
|
565
|
-
npx -y scenescout install --browsers chromium-headless-shell # the smallest working setup
|
|
566
|
-
npx -y scenescout install --browser-only --browsers firefox,webkit # add two more later
|
|
567
|
-
```
|
|
568
|
-
|
|
569
|
-
Sizes vary by platform. The builds go to Playwright's shared cache, so a build another tool already fetched is not downloaded again.
|
|
570
|
-
|
|
571
|
-
The first attach that needs a build which is not on disk downloads it itself, once, and says so ("Getting the test browser ready"). It does not in CI unless `SCENESCOUT_BROWSER_DOWNLOAD=on` is set, and never with `SCENESCOUT_BROWSER_DOWNLOAD=off`, for a machine where nothing may be downloaded.
|
|
572
|
-
|
|
573
|
-
To drive another browser, pass `browser` when attaching (`scout_attach { browser: "firefox" }`), or set `SCENESCOUT_BROWSER=webkit` in the server's environment to change the default. `scenescout doctor` checks the browser named by that variable in the shell it runs from, so check another one with `SCENESCOUT_BROWSER=webkit scenescout doctor`. Two things differ outside Chromium:
|
|
574
|
-
|
|
575
|
-
- **Service workers are not allowed to register** in Firefox and WebKit. The write policy works by intercepting requests, and only Chromium lets a request issued by a service worker be intercepted. An app that depends on its worker may behave differently there.
|
|
576
|
-
- **A Firefox or WebKit left behind by a crash is not cleaned up** on the next start the way a leftover Chromium is.
|
|
577
|
-
|
|
578
|
-
In every browser, pages are not given shared workers unless the mode is `destructive`: a request a shared worker sends cannot be intercepted anywhere, so the app is made to do that work on the page, where the policy sees it.
|
|
579
|
-
|
|
580
|
-
### Time limits
|
|
581
|
-
|
|
582
|
-
An action on the page (a click, typing, a hover, a pick from a list) may take 5 s, and a page 20 s to load (15 s for a page the crawl opens). On a loaded machine these can run out while the app is fine; the timeout then says which limit ran out and how to raise it. Raise them per session with `scout_attach { actionTimeoutMs: 15000, navTimeoutMs: 60000 }`, for every session with `SCENESCOUT_ACTION_TIMEOUT_MS` and `SCENESCOUT_NAV_TIMEOUT_MS` in the server's environment, or on `scenescout check` and `scenescout ci` with `--action-timeout-ms` and `--nav-timeout-ms`. An option wins over the variable, and the variable over the default. The action limit takes 1000 to 120000 ms and the page-load limit 1000 to 300000 ms; anything else refuses the attach with a sentence naming the value to fix. Saving a login profile with `scenescout login` honours the two variables as well, and otherwise keeps its own longer waits (30 s for the page, 10 s for a field or the submit).
|
|
583
|
-
|
|
584
|
-
## ๐ Other MCP clients
|
|
232
|
+
> [!IMPORTANT]
|
|
233
|
+
> Only test sites you own or are allowed to test. [Safety model](docs/guide/Safety-model.md).
|
|
585
234
|
|
|
586
|
-
|
|
235
|
+
## ๐ Documentation
|
|
587
236
|
|
|
588
|
-
|
|
589
|
-
npx -y scenescout install --client cursor # one client
|
|
590
|
-
npx -y scenescout install --client vscode,codex # several; add claude-code to keep that one too
|
|
591
|
-
```
|
|
592
|
-
|
|
593
|
-
| `--client` | How it is registered |
|
|
237
|
+
| | |
|
|
594
238
|
|---|---|
|
|
595
|
-
|
|
|
596
|
-
|
|
|
597
|
-
|
|
|
598
|
-
|
|
|
599
|
-
|
|
|
600
|
-
|
|
|
601
|
-
|
|
|
602
|
-
|
|
603
|
-
|
|
604
|
-
|
|
605
|
-
|
|
606
|
-
|
|
607
|
-
To register by hand instead, the server entry is always the same command, `npx -y scenescout serve`:
|
|
608
|
-
|
|
609
|
-
<details>
|
|
610
|
-
<summary><strong>Cursor</strong> โ <code>~/.cursor/mcp.json</code> (or <code>.cursor/mcp.json</code> in a project)</summary>
|
|
611
|
-
|
|
612
|
-
```json
|
|
613
|
-
{
|
|
614
|
-
"mcpServers": {
|
|
615
|
-
"scenescout": { "command": "npx", "args": ["-y", "scenescout", "serve"] }
|
|
616
|
-
}
|
|
617
|
-
}
|
|
618
|
-
```
|
|
619
|
-
|
|
620
|
-
</details>
|
|
621
|
-
|
|
622
|
-
<details>
|
|
623
|
-
<summary><strong>VS Code</strong> (GitHub Copilot agent mode) โ <code>.vscode/mcp.json</code></summary>
|
|
239
|
+
| [Start here](docs/guide/Start-here.md) | A first look, install, a first run, reading the report, the live view |
|
|
240
|
+
| [Ways to use it](docs/guide/Ways-to-use-it.md) | Interactive runs, parallel agents, CI, pull-request QA, filing issues |
|
|
241
|
+
| [What it checks](docs/guide/What-it-checks.md) | Every check and every `scout_*` tool |
|
|
242
|
+
| [Signing in](docs/guide/Signing-in.md) ยท [Safety model](docs/guide/Safety-model.md) | Roles and saved logins; what each mode refuses and why |
|
|
243
|
+
| [Recipes](docs/guide/Recipes.md) | Setups for seven kinds of project |
|
|
244
|
+
| [Configuration reference](docs/guide/Configuration-reference.md) | Every option, environment variable and action input |
|
|
245
|
+
| [Troubleshooting](docs/guide/Troubleshooting.md) | Symptoms and fixes, upgrading and uninstalling |
|
|
246
|
+
| [Running it in CI](docs/ci.md) | Saved journeys, recording, test reports from templates, every workflow |
|
|
247
|
+
| [How it works](docs/how-it-works.md) | Diagrams of a run, an action, the write policy, lanes |
|
|
248
|
+
| [Benchmark](docs/benchmark.md) ยท [Validation](docs/validation.md) | How runs are scored against answer keys, and runs on public apps |
|
|
249
|
+
| [Design decisions](docs/adr/README.md) | Why the rules are what they are |
|
|
624
250
|
|
|
625
|
-
|
|
626
|
-
{
|
|
627
|
-
"servers": {
|
|
628
|
-
"scenescout": { "type": "stdio", "command": "npx", "args": ["-y", "scenescout", "serve"] }
|
|
629
|
-
}
|
|
630
|
-
}
|
|
631
|
-
```
|
|
251
|
+
## ๐ง Contributing
|
|
632
252
|
|
|
633
|
-
|
|
634
|
-
|
|
635
|
-
<details>
|
|
636
|
-
<summary><strong>Codex CLI</strong> โ <code>~/.codex/config.toml</code></summary>
|
|
637
|
-
|
|
638
|
-
```toml
|
|
639
|
-
[mcp_servers.scenescout]
|
|
640
|
-
command = "npx"
|
|
641
|
-
args = ["-y", "scenescout", "serve"]
|
|
642
|
-
```
|
|
253
|
+
Start with [VISION.md](VISION.md) (what is in scope) and [CONTRIBUTING.md](CONTRIBUTING.md) (local setup and how changes land). [AGENTS.md](AGENTS.md) holds the house rules for people and coding agents alike.
|
|
643
254
|
|
|
644
|
-
|
|
645
|
-
|
|
646
|
-
<details>
|
|
647
|
-
<summary><strong>Gemini CLI</strong> โ <code>~/.gemini/settings.json</code> (or <code>.gemini/settings.json</code> in a project)</summary>
|
|
648
|
-
|
|
649
|
-
```json
|
|
650
|
-
{
|
|
651
|
-
"mcpServers": {
|
|
652
|
-
"scenescout": { "command": "npx", "args": ["-y", "scenescout", "serve"] }
|
|
653
|
-
}
|
|
654
|
-
}
|
|
655
|
-
```
|
|
255
|
+
Found a way past the write policy, or another security problem? Report it privately: [SECURITY.md](SECURITY.md).
|
|
656
256
|
|
|
657
|
-
|
|
658
|
-
|
|
659
|
-
<details>
|
|
660
|
-
<summary><strong>Windsurf</strong> โ <code>~/.codeium/windsurf/mcp_config.json</code></summary>
|
|
661
|
-
|
|
662
|
-
```json
|
|
663
|
-
{
|
|
664
|
-
"mcpServers": {
|
|
665
|
-
"scenescout": { "command": "npx", "args": ["-y", "scenescout", "serve"] }
|
|
666
|
-
}
|
|
667
|
-
}
|
|
668
|
-
```
|
|
669
|
-
|
|
670
|
-
</details>
|
|
671
|
-
|
|
672
|
-
<details>
|
|
673
|
-
<summary><strong>Cline</strong> โ MCP Servers โ Configure โ Configure MCP Servers (or <code>~/.cline/mcp.json</code> for the CLI)</summary>
|
|
674
|
-
|
|
675
|
-
```json
|
|
676
|
-
{
|
|
677
|
-
"mcpServers": {
|
|
678
|
-
"scenescout": { "command": "npx", "args": ["-y", "scenescout", "serve"], "disabled": false, "autoApprove": [] }
|
|
679
|
-
}
|
|
680
|
-
}
|
|
681
|
-
```
|
|
682
|
-
|
|
683
|
-
</details>
|
|
684
|
-
|
|
685
|
-
<details>
|
|
686
|
-
<summary><strong>Zed</strong> โ <code>settings.json</code> (command palette: <code>zed: open settings file</code>)</summary>
|
|
687
|
-
|
|
688
|
-
```json
|
|
689
|
-
{
|
|
690
|
-
"context_servers": {
|
|
691
|
-
"scenescout": { "command": "npx", "args": ["-y", "scenescout", "serve"], "env": {} }
|
|
692
|
-
}
|
|
693
|
-
}
|
|
694
|
-
```
|
|
695
|
-
|
|
696
|
-
</details>
|
|
697
|
-
|
|
698
|
-
<details>
|
|
699
|
-
<summary><strong>Anything else</strong></summary>
|
|
700
|
-
|
|
701
|
-
Most clients accept the same `mcpServers` JSON shape shown for Cursor.
|
|
702
|
-
|
|
703
|
-
</details>
|
|
704
|
-
|
|
705
|
-
**The method travels with the server.** The tools are only hands and eyes; [`skills/scenescout/SKILL.md`](skills/scenescout/SKILL.md) is the method: what to look at first, when to stop, what counts as a finding. Claude Code loads it as a skill. Every other client gets the same text from the server, with nothing to copy:
|
|
706
|
-
|
|
707
|
-
- the server's instructions tell the agent to call `scout_playbook` before its first attach, and that tool returns the method,
|
|
708
|
-
- clients that list server prompts as commands also get an `explore` prompt, which loads the method and takes an optional URL, level and focus.
|
|
709
|
-
|
|
710
|
-
So in any client, a first message like *"Use SceneScout to test http://localhost:3000"* is enough. If an agent starts clicking without having called `scout_playbook`, tell it to call that first; how closely a model follows server instructions varies by client.
|
|
711
|
-
|
|
712
|
-
The CLI is also useful on its own:
|
|
713
|
-
|
|
714
|
-
```bash
|
|
715
|
-
npx -y scenescout scan <path> # project discovery: framework, routes, saved logins
|
|
716
|
-
npx -y scenescout status <path> # what every session of a running engine is doing right now
|
|
717
|
-
npx -y scenescout watch <path> # the same, live in your browser, with each session's page
|
|
718
|
-
npx -y scenescout login <url> --role admin # sign in once in a visible browser; sessions attach with role: "admin"
|
|
719
|
-
```
|
|
720
|
-
|
|
721
|
-
---
|
|
722
|
-
|
|
723
|
-
## ๐ Project layout
|
|
724
|
-
|
|
725
|
-
```
|
|
726
|
-
src/
|
|
727
|
-
mcp-server.ts the 29 tools + per-session dispatch
|
|
728
|
-
scan.ts project discovery (framework, routes, auth)
|
|
729
|
-
cli.ts scan ยท serve ยท install ยท doctor ยท check ยท ci ยท login ยท export ยท status ยท watch
|
|
730
|
-
check-run.ts drives a check: attach, crawl every route, collect what was measured
|
|
731
|
-
ci-run.ts drives a CI run: the MCP server as a child, the model's API, the agent loop
|
|
732
|
-
login-run.ts drives `scenescout login` and scout_login: a visible browser that saves the role's profile once the sign-in is seen to finish (or on Enter); or --script, headless from the environment
|
|
733
|
-
export-run.ts drives `scenescout export`: reads the findings, asks GitHub or Jira what is filed, files the rest
|
|
734
|
-
installer.ts setup logic (skill link, MCP registration, diagnostics)
|
|
735
|
-
engine/
|
|
736
|
-
browser.ts the engine class: attach, snapshot, actions, crawl, plans
|
|
737
|
-
probes.ts in-page scroll + overlay + focus probes (needs a browser too)
|
|
738
|
-
fingerprint.ts route + element-set identity (state hashing)
|
|
739
|
-
oracles.ts console/page/network/HTTP error detection
|
|
740
|
-
injection.ts the DOM-injection oracle's rules (what to watch for, how to find it)
|
|
741
|
-
claims.ts when the page contradicts the server (refused_empty, false_success)
|
|
742
|
-
request.ts what a replayed API call may be and where it may go
|
|
743
|
-
brief.ts splitting the app between parallel lanes
|
|
744
|
-
lane.ts the typed report a lane hands back
|
|
745
|
-
calibration.ts whether a lane's confidence held up; what it judged and never filed
|
|
746
|
-
pace.ts how a run spent its time
|
|
747
|
-
bench.ts scoring a run against the demo app's answer key
|
|
748
|
-
policy.ts the write-policy safety net
|
|
749
|
-
ownership.ts safe-write: which records were created in this process?
|
|
750
|
-
uploads.ts disk uploads, fenced to the project by real path
|
|
751
|
-
journey.ts task-ease measurement from the action log
|
|
752
|
-
design.ts the design audit + page scoring
|
|
753
|
-
memory.ts cross-run storage + finding dedup
|
|
754
|
-
profiles.ts saved sign-ins: role names, where a profile lives, owner-only files, attach by role, sessionStorage restore
|
|
755
|
-
refresh.ts the refresh broker: which values are a role's refresh tokens, the lock beside the profile, swapping a spent token
|
|
756
|
-
scripted-login.ts a CI sign-in: env and flags, TOTP (RFC 6238) or a fixed code, which field is which, redaction
|
|
757
|
-
signed-in.ts when a person's sign-in in the window has finished: back on the app, a new session, past any SSO round trip
|
|
758
|
-
expiry.ts how long a saved sign-in lasts: cookie dates and JWT exp, checked before lanes start
|
|
759
|
-
report.ts the gap ledger + report generation
|
|
760
|
-
check.ts the check's rules, gate, report and SARIF
|
|
761
|
-
baseline.ts visual baselines: targets.json, where each picture is kept, when one is met
|
|
762
|
-
sarif.ts which repository file a SARIF result points at, so code scanning keeps it
|
|
763
|
-
ci.ts a CI run's options, provider choice, caps, key redaction, tools and files
|
|
764
|
-
export.ts which findings an export files, the inert issue it writes, the marker that dedups it
|
|
765
|
-
provider.ts the Anthropic and OpenAI message shapes, and retries
|
|
766
|
-
replay.ts the run as one page: steps, tasks, frames under each finding
|
|
767
|
-
โฆ collector ยท dispatch ยท fixtures ยท authloss ยท reaper
|
|
768
|
-
scripts/ the test suites (smoke/ holds the real-browser ones)
|
|
769
|
-
test-app/ fixtures for the real-browser smoke tests
|
|
770
|
-
skills/scenescout/ the testing method (SKILL.md): a skill in Claude Code, served by the server everywhere else
|
|
771
|
-
docs/how-it-works.md what happens at each stage, in diagrams
|
|
772
|
-
docs/benchmark.md measuring whether a change made runs better
|
|
773
|
-
docs/validation.md scorecards from runs against public open-source apps
|
|
774
|
-
docs/adr/ why it's built this way
|
|
775
|
-
```
|
|
776
|
-
|
|
777
|
-
> Design principle: logic that *doesn't* need Playwright lives outside `browser.ts`, so it can be unit-tested without launching a browser. That's why `fingerprint`, `policy`, `memory`, `report`, etc. are their own modules.
|
|
778
|
-
|
|
779
|
-
---
|
|
780
|
-
|
|
781
|
-
## ๐ง Design decisions
|
|
782
|
-
|
|
783
|
-
**[How it works, stage by stage](docs/how-it-works.md)** โ diagrams of the run lifecycle, what happens inside one action, the write policy on the wire, how a violation becomes a finding, how a parallel run is split and folded, how roles hand work to each other, where a run's time goes, and how a lane's confidence is checked afterwards.
|
|
784
|
-
|
|
785
|
-
**[Measuring whether a change helped](docs/benchmark.md)** โ the demo app's answer key, the scorecard (recall, precision, judged-not-filed, severity, calibration), and the results log of every run, including what did not help.
|
|
786
|
-
|
|
787
|
-
**[Validation on public open-source apps](docs/validation.md)** โ runs against three well-known open-source web apps, with a scorecard for each: issues by severity, how many were real and how many were false positives, and the engine problems the runs exposed.
|
|
788
|
-
|
|
789
|
-
The load-bearing choices are recorded as ADRs โ read the relevant one before changing a rule it covers:
|
|
790
|
-
|
|
791
|
-
- [1 ยท Completion is an enforced contract, not a claim](docs/adr/0001-completion-is-a-contract-not-a-vibe.md)
|
|
792
|
-
- [2 ยท The write policy is enforced on the wire, not in the prompt](docs/adr/0002-enforce-the-write-policy-at-the-network-layer.md)
|
|
793
|
-
- [3 ยท A gap-ledger entry must be actionable, and suppression must be visible](docs/adr/0003-a-noisy-ledger-is-a-broken-ledger.md)
|
|
794
|
-
- [4 ยท Findings dedup on machine signals, and a merge must never lose a finding](docs/adr/0004-dedup-on-machine-signals-not-prose.md)
|
|
795
|
-
- [5 ยท Testable logic lives outside `browser.ts`](docs/adr/0005-keep-testable-logic-out-of-the-browser-module.md)
|
|
796
|
-
- [6 ยท Nothing in this repo names or is tuned for a tested app](docs/adr/0006-stay-project-agnostic.md)
|
|
797
|
-
- [7 ยท The live view is local, read-only, and leaves nothing behind](docs/adr/0007-the-live-view-is-local-read-only-and-leaves-nothing-behind.md)
|
|
798
|
-
- [8 ยท Recording is opt-in, and a recorded run is one self-contained page](docs/adr/0008-a-recorded-run-is-evidence-and-must-be-asked-for.md)
|
|
799
|
-
- [9 ยท A refused write is answered, not dropped](docs/adr/0009-a-refused-write-is-answered-not-dropped.md)
|
|
800
|
-
- [10 ยท A lane's confidence is checked, not trusted](docs/adr/0010-a-confidence-is-checked-not-trusted.md)
|
|
801
|
-
- [11 ยท A gate is deterministic, and fails only on what it can prove](docs/adr/0011-a-gate-is-deterministic-and-fails-only-on-what-it-can-prove.md)
|
|
802
|
-
- [12 ยท A check replays saved flows and re-tests open findings, within settings whose defaults do the least harm](docs/adr/0012-a-check-replays-saved-flows-and-reports-re-tests.md)
|
|
803
|
-
- [13 ยท What depends on a project's convention is the project's to decide](docs/adr/0013-a-convention-is-the-projects-to-decide.md)
|
|
804
|
-
- [21 ยท Update the docs in the same pull request, unless the change has no user-facing surface](docs/adr/0021-update-the-docs-in-the-pull-request.md)
|
|
805
|
-
|
|
806
|
-
---
|
|
807
|
-
|
|
808
|
-
## ๐ง Development
|
|
809
|
-
|
|
810
|
-
Working on SceneScout itself is the only reason to clone it:
|
|
811
|
-
|
|
812
|
-
```bash
|
|
813
|
-
git clone https://github.com/brunoboto96/SceneScout.git scenescout && cd scenescout
|
|
814
|
-
npm install # installs dependencies and builds
|
|
815
|
-
npm run setup # same as `scenescout install`, but registers THIS checkout (the skill is linked, so edits are live)
|
|
816
|
-
npm test # build + 24 suites: 22 pure-logic suites (scan, oracle, policy, โฆ bench, hygiene),
|
|
817
|
-
# then smoke and mcp-check (the server over stdio), both with real browsers
|
|
818
|
-
npm run bench -- --all # re-score every archived benchmark run against the current answer key
|
|
819
|
-
npm run demo # regenerate examples/ from the demo app
|
|
820
|
-
```
|
|
821
|
-
|
|
822
|
-
Contributing? Start with [VISION.md](VISION.md) (what is in scope) and [CONTRIBUTING.md](CONTRIBUTING.md) (how changes land), then see [AGENTS.md](AGENTS.md) for the house rules โ chiefly: bug fixes need a regression test at the cheapest layer that can fail, keep the repo project-agnostic (ADR 6), and `npm test` must pass.
|
|
823
|
-
|
|
824
|
-
## ๐ Security
|
|
825
|
-
|
|
826
|
-
Found a way past the write policy, or another security problem? Please report it privately โ see [SECURITY.md](SECURITY.md).
|
|
827
|
-
|
|
828
|
-
## ๐ License
|
|
829
|
-
|
|
830
|
-
[MIT](LICENSE).
|
|
831
|
-
|
|
832
|
-
<details>
|
|
833
|
-
<summary><strong>Full capability list</strong> โ every behavior, for the curious</summary>
|
|
834
|
-
|
|
835
|
-
- **Structured render-state, not pixels.** Element lists with geometry; screenshots reserved for pixel-native residue (canvas, rendering glitches). Images that failed to load are reported from the DOM, including ones whose URL answered 200 with something that is not an image.
|
|
836
|
-
- **Diff snapshots with stable refs.** Re-snapshots return only what changed (10.7 kB โ 0.7 kB on a 130-element page); old refs stay valid.
|
|
837
|
-
- **Geometry oracles.** Overlap and off-screen defects computed from layout boxes.
|
|
838
|
-
- **Oracles after every action.** Console errors, page errors, failed requests, HTTP 4xx/5xx drained into every tool result โ and DOM injection: a markup-shaped value the agent typed that later renders as an element on any page (stored or reflected XSS).
|
|
839
|
-
- **Multi-role, genuinely concurrent.** Commands to *different* sessions run in parallel; safe-write ownership is shared, so role A can create what role B approves. The report renders a role capability matrix.
|
|
840
|
-
- **Task ease, not just correctness.** `scout_journey` measures interaction cost, distinct screens, path, and backtracks.
|
|
841
|
-
- **Design audit with page scores.** Two tiers (โ measurable defects / โ craft suggestions incl. AI-slop tells), per-page 0โ100 score persisted per route, plus an automatic overlay/modal probe on every snapshot. Shared shell scored once, separately.
|
|
842
|
-
- **Scrolls like a user โ and notices when it can't.** Reports `SCROLL LOCKED` for a leaked modal scroll-lock, finds the real inner scroll pane on app-shell layouts, and flags `UNREACHABLE` controls clipped inside `overflow:hidden`.
|
|
843
|
-
- **Uploads like a user.** Answers a styled file-chooser or sets a hidden input directly, with a valid in-memory fixture; `filePath` is fenced to the project under test; files violating `accept` are flagged at selection.
|
|
844
|
-
- **Auth via Playwright storage states.** Expired tokens caught at attach; repeated login-bounces raise `SESSION AUTH LOST` (a session attached by role first re-attaches once from its role's latest saved profile and carries on); a bounced route is recorded as *not* covered โ a dead session can't certify routes it never reached.
|
|
845
|
-
- **A trustworthy gap ledger.** Entries must be actionable (a search box or wizard sub-step isn't "form filled but never submitted"); API/download URLs never enter the route contract.
|
|
846
|
-
- **Honest reporting.** Shared chrome counted once, stale scores marked, role matrix compares only roles that actually attempted a route.
|
|
847
|
-
- **Cross-run written knowledge.** `scout_note` curates `.scenescout/ASSUMPTIONS.md` โ app model, personas, constraints, risks โ in prose.
|
|
848
|
-
- **Daemon-grade robustness.** Per-tool watchdogs, orphaned-browser reaping, bounded teardown, live status via `scenescout status <project>`, and a live view of every session's page: the agent gives you its address when it attaches, or run `scenescout watch <project>` (loopback only, read-only, nothing written to disk: [ADR 7](docs/adr/0007-the-live-view-is-local-read-only-and-leaves-nothing-behind.md)).
|
|
849
|
-
|
|
850
|
-
</details>
|
|
257
|
+
[MIT](LICENSE) licensed.
|