scenescout 3.22.0 โ 3.23.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +12 -0
- package/README.md +88 -769
- package/dist/ci-run.js +259 -36
- package/dist/cli.js +1 -1
- package/dist/engine/brief.js +35 -3
- package/dist/engine/ci-lanes.js +28 -12
- package/dist/engine/ci.js +39 -1
- package/dist/engine/dispatch.js +70 -0
- package/dist/engine/from-run.js +649 -0
- package/dist/engine/memory.js +135 -0
- package/dist/engine/report.js +26 -0
- package/dist/mcp-server.js +108 -11
- package/package.json +2 -2
- package/skills/scenescout/SKILL.md +2 -2
package/README.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
# ๐ญ SceneScout
|
|
4
4
|
|
|
5
|
-
**
|
|
5
|
+
**Your coding agent, turned into an exploratory QA tester for any web app.**
|
|
6
6
|
|
|
7
7
|
Works with Claude Code ยท Cursor ยท VS Code (Copilot) ยท Codex CLI ยท Gemini CLI ยท Copilot CLI ยท Windsurf ยท any [MCP](https://modelcontextprotocol.io) client
|
|
8
8
|
|
|
@@ -10,840 +10,159 @@ Works with Claude Code ยท Cursor ยท VS Code (Copilot) ยท Codex CLI ยท Gemini CLI
|
|
|
10
10
|
[](https://www.npmjs.com/package/scenescout)
|
|
11
11
|
[](LICENSE)
|
|
12
12
|

|
|
13
|
-

|
|
14
13
|
|
|
15
|
-
[๐ Guide](docs/guide/Home.md) ยท [
|
|
14
|
+
[๐ Guide](docs/guide/Home.md) ยท [๐ What it catches](#-what-it-catches) ยท [๐ Get started](#-get-started) ยท [๐ฆ CI](#-in-ci) ยท [๐ Safety](#-safe-by-default) ยท [๐ Docs](#-documentation)
|
|
16
15
|
|
|
17
16
|
</div>
|
|
18
17
|
|
|
19
|
-
|
|
18
|
+
Scripted end-to-end tests answer one question: *does this exact flow still work?* They say nothing about the rest of the app. SceneScout lets the agent you already use explore a running web app like a curious, thorough tester. It clicks, fills forms, switches roles and calls the API behind a hidden button, then writes a report of what is **broken** and what could be **better**, with evidence for every line.
|
|
20
19
|
|
|
21
|
-
|
|
22
|
-
โโโโโโโโโโโโโโโโโโโโโโโ MCP (stdio) โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
|
|
23
|
-
โ Your coding agent โ โโโโโโโโโโโโโโโถ โ SceneScout engine โ
|
|
24
|
-
โ (intent, judgment, โ โโโโโโโโโโโโโโโ โ Playwright ยท oracles ยท memory โ
|
|
25
|
-
โ your subscription) โ tool results โ findings ยท report โ no LLM โ
|
|
26
|
-
โโโโโโโโโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
|
|
27
|
-
```
|
|
28
|
-
|
|
29
|
-
Scripted E2E suites answer one question โ *"does this exact flow still work?"* โ and say nothing about the 95% of the app they don't touch. SceneScout covers both gaps: it finds what's **broken** (crashes, dead ends, permission leaks) *and* reports how the product could be **better** (confusing flows, weak hierarchy, design-system drift), with concrete measurements.
|
|
20
|
+
Try it on any app you are allowed to test. No account, no API key, no setup:
|
|
30
21
|
|
|
22
|
+
```bash
|
|
23
|
+
npx -y scenescout http://localhost:3000
|
|
24
|
+
```
|
|
31
25
|
|
|
32
|
-
##
|
|
26
|
+
## ๐ What it catches
|
|
33
27
|
|
|
34
|
-
|
|
28
|
+
A real run against the small demo app in this repository, which has bugs planted on purpose:
|
|
35
29
|
|
|
36
30
|
<p align="center"><img src="examples/screenshots/dashboard-annotated.png" alt="The demo app's dashboard with two defects outlined in red: 1, a yellow badge covering the All orders button; 2, the weekly chart image failing to load" width="760" /></p>
|
|
37
31
|
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
An excerpt of the report it wrote โ [read the whole thing](examples/report.md):
|
|
32
|
+
It filed twelve findings. A few from [the report](examples/report.md):
|
|
41
33
|
|
|
42
34
|
> **๐ด [HIGH] A double-click on Create order creates two orders**
|
|
43
|
-
>
|
|
44
|
-
> The submit button stays enabled while the request is in flight, and the endpoint accepts the repeat.
|
|
45
|
-
>
|
|
46
|
-
> **๐ด [HIGH] Filtering orders by Archived fails, and the page shows an empty table instead of an error**
|
|
47
|
-
> Evidence: `GET /api/orders?status=archived โ HTTP 500`
|
|
35
|
+
> `2ร click fired the same state-changing request 2ร (POST /api/orders)`
|
|
48
36
|
>
|
|
49
37
|
> **๐ด [HIGH] A clerk can approve an order by calling the endpoint the page hides from them**
|
|
50
|
-
>
|
|
38
|
+
> `POST /api/orders/1037/approve 200 as clerk`: the button was hidden, the server did not agree.
|
|
51
39
|
>
|
|
52
|
-
>
|
|
53
|
-
>
|
|
54
|
-
>
|
|
55
|
-
> **๐ก [LOW] The dashboard chart image is missing** *(callout 2)*
|
|
56
|
-
> Evidence: `GET /img/weekly-chart.png โ HTTP 404`
|
|
40
|
+
> **๐ด [HIGH] Filtering orders by Archived fails, and the page shows an empty table instead of an error**
|
|
41
|
+
> `GET /api/orders?status=archived โ HTTP 500`
|
|
57
42
|
>
|
|
58
|
-
>
|
|
59
|
-
|
|
60
|
-
Every finding comes with a repro trace and a Playwright regression-test skeleton. To try it yourself, clone this repository, run `npm run demo:serve`, then `/scenescout --url http://127.0.0.1:4173` โ see [demo-app/](demo-app/). Its README lists every seeded defect and which oracle catches it.
|
|
61
|
-
|
|
62
|
-
---
|
|
63
|
-
|
|
64
|
-
## โจ Why it's different
|
|
65
|
-
|
|
66
|
-
- ๐ง **Your agent is the brain โ no API key.** The engine contains no LLM. Exploration runs on the agent and subscription you already have (Claude Code, Cursor, Copilot, Codex, Gemini CLI and others); SceneScout just gives it deterministic tools and the method for using them.
|
|
67
|
-
- ๐ **Structured scene, not pixels.** The agent reads element lists *with layout geometry*, not screenshots. Overlap and off-screen bugs are computed from boxes โ deterministic, no vision guessing. Images that failed to load are read from the DOM too. (Screenshots exist only for pixel-native residue like a canvas or a rendering glitch.)
|
|
68
|
-
- ๐ก๏ธ **Read-only by default, enforced on the wire.** Destructive actions are blocked at the network layer, not by asking the model nicely. Opt into writes only against disposable data.
|
|
69
|
-
- โ
**Completion is a contract, not a vibe.** The engine knows the app's routes and *refuses* to file an "extensive" report while any known route is unvisited, unexercised, or un-audited. "Explored a bit and stopped" is structurally impossible.
|
|
70
|
-
- ๐งญ **It remembers.** UI states are fingerprinted and stored in the project's `.scenescout/`. Run N+1 skips what run N already covered, and every run starts smarter than the last.
|
|
71
|
-
|
|
72
|
-
---
|
|
73
|
-
|
|
74
|
-
## ๐ฏ Two ways to use it
|
|
75
|
-
|
|
76
|
-
SceneScout needs only a URL. Give it the source code as well and it gets noticeably better.
|
|
77
|
-
|
|
78
|
-
| | ๐ **Next to the codebase** *(recommended)* | ๐ **Against a remote URL** |
|
|
79
|
-
|---|---|---|
|
|
80
|
-
| **You run it from** | the app's repository | any folder โ an empty `qa/` directory is fine |
|
|
81
|
-
| **It plays the role of** | a developer-tester who can read the code | a black-box QA tester, like a person with a browser |
|
|
82
|
-
| **How it finds pages** | ๐ reads routes from the source **and** follows links: file-based routing (Next.js, SvelteKit, Nuxt) and router configuration written in code (React Router, Vue Router, Angular). Routes built at runtime are not seen | ๐ follows same-origin links only โ pages nothing links to, or on another subdomain, stay unknown |
|
|
83
|
-
| **"Did we cover everything?"** | checked against the routes found in source *plus* discovered links โ an unvisited one blocks the report | checked against the pages it managed to discover |
|
|
84
|
-
| **Setup it figures out** | framework, dev command, saved Playwright logins (`playwright/.auth/`), whether the app uses `data-testid` | none โ you pass the URL, and the path to a login state if the app needs one |
|
|
85
|
-
| **What a finding looks like** | the symptom, **plus** the file behind it and a suggested fix | the symptom, a repro trace, and a regression-test skeleton |
|
|
86
|
-
| **Typical target** | `localhost` while you build | staging, a preview deploy, a client's site |
|
|
87
|
-
|
|
88
|
-
**Why the codebase helps.** The agent driving SceneScout is a coding agent, which can already read your repository. With the source at hand it knows the app's static routes before opening the browser, so coverage is measured against the real app instead of whatever happened to be linked. It can also check a suspicion against the code before reporting it: "there is no way to export this table" is a much stronger finding once the agent has confirmed no export handler exists. And when something breaks it can open the component or handler responsible and tell you *where* and *how* to fix it โ "the save button does nothing" becomes "`OrderForm` swallows the rejected promise in `onSubmit`; surface the error and re-enable the button".
|
|
89
|
-
|
|
90
|
-
**Why it still works without it.** Everything SceneScout *observes* comes from the running page โ elements, layout geometry, console and network errors, design-audit scores, task-ease measurements โ and none of that needs source code. Point it at a URL you are allowed to test and it behaves like a thorough QA tester: it explores, reproduces, and files findings with evidence.
|
|
91
|
-
|
|
92
|
-
```
|
|
93
|
-
# next to the code โ run inside the app's repository
|
|
94
|
-
/scenescout --url http://localhost:3000
|
|
95
|
-
|
|
96
|
-
# remote โ run from any folder; memory and the report are kept there
|
|
97
|
-
/scenescout --url https://staging.example.com --role ./auth/qa.json
|
|
98
|
-
```
|
|
99
|
-
|
|
100
|
-
> [!IMPORTANT]
|
|
101
|
-
> Only test sites you own or are authorized to test. A remote environment is more likely to hold real data, so for a remote URL with no source the skill attaches in **`observe`** mode: nothing but `GET` requests leaves the page. The default **read-only** mode blocks `PUT`/`PATCH`/`DELETE` and destructive-looking requests, but an ordinary form submission (a plain `POST`: contact form, comment, order, signup) still reaches the server and can create a record. Say so when that is acceptable on your target. See the [safety model](#-safety-model).
|
|
43
|
+
> **๐ [MEDIUM] The "New: bulk import" badge sits on top of the All orders button** *(callout 1)*
|
|
44
|
+
> `"All orders" overlaps "New: bulk import" (81%)`, measured from layout boxes, no screenshot needed.
|
|
102
45
|
|
|
103
|
-
|
|
46
|
+
What it looks for, on every page, after every action:
|
|
104
47
|
|
|
105
|
-
|
|
48
|
+
- ๐งจ **Real breakage:** console errors, crashes, failed requests, 4xx and 5xx responses, broken images, dead-end pages.
|
|
49
|
+
- ๐ **Permission leaks:** it calls the app's own API as each role, so "the button is hidden" becomes "the server refuses it", or doesn't.
|
|
50
|
+
- ๐คฅ **Pages that lie:** "Saved!" after the server refused the save, or an empty table after the request failed.
|
|
51
|
+
- ๐ **Impatient users:** a double-click that sends the same order twice.
|
|
52
|
+
- ๐ **Broken layout, from geometry:** controls that overlap, sit off-screen, hide under a sticky bar or can never be scrolled into view.
|
|
53
|
+
- โฟ **Accessibility and craft:** contrast, focus, labels, target sizes, spacing and type, with a 0 to 100 score per page.
|
|
54
|
+
- ๐งญ **Friction:** how many steps a task takes, and where a user had to go back.
|
|
55
|
+
- ๐ **Security smells:** typed markup that comes back as an element, and tokens posted to any window.
|
|
106
56
|
|
|
107
|
-
|
|
57
|
+
Every finding comes with the evidence, the steps to reproduce it, a picture, and a Playwright regression-test skeleton. Next to the source code, it also names the file behind the bug and a likely fix. [Everything it checks](docs/guide/What-it-checks.md).
|
|
108
58
|
|
|
109
|
-
|
|
59
|
+
## ๐บ Watch it work
|
|
110
60
|
|
|
111
|
-
|
|
112
|
-
npx -y scenescout http://localhost:3000
|
|
113
|
-
```
|
|
61
|
+
Each run opens a live view on your machine, with one card per agent: what it is doing, the page it is on, and a feed of every action. Three agents are testing the demo app in parallel here:
|
|
114
62
|
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
```
|
|
118
|
-
Look at these first:
|
|
119
|
-
1. [medium] Request failed with a client error: GET /img/weekly-chart.png โ HTTP 404 (on /)
|
|
120
|
-
2. [medium] Dead end: /reports-scheduled.html: 0 controls (on /reports-scheduled.html)
|
|
121
|
-
3. [medium] Control covered by pinned chrome: button "Save notes" is COVERED by pinned chrome [order-stickybar] at this scroll position โ a click aimed at it lands on that element instead (on /order.html?id=1042)
|
|
122
|
-
|
|
123
|
-
12 pages looked at in 12 s in observe mode: 0 high ยท 6 medium ยท 2 low ยท 5 worth a look, never counted.
|
|
124
|
-
Report: scenescout-report/report.md
|
|
125
|
-
```
|
|
63
|
+
<p align="center"><img src="examples/screenshots/live-view.png" alt="The live view during a run of three parallel agents against the demo app: one card per session, each with its role and objective, the task it is on, the tool it is running, the page it is on, a live thumbnail, and a feed of the actions it just took, tinted one colour per task" width="880" /></p>
|
|
126
64
|
|
|
127
|
-
|
|
65
|
+
You can read the report while the agents are still working, and scrub back through any session's timeline. Beside `report.md`, every run writes `report.html`, the whole run as one self-contained page; ask for a recorded run and it also keeps a frame after every action.
|
|
128
66
|
|
|
129
|
-
|
|
67
|
+
## โจ Why it's different
|
|
130
68
|
|
|
131
|
-
|
|
132
|
-
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
|
|
69
|
+
- ๐ง **Your agent is the brain, so no extra API key.** The engine contains no model. It gives the agent you already pay for deterministic tools and the testing method to use them.
|
|
70
|
+
- ๐ **It reads structure, not pixels.** The agent sees every element with its role, state and layout box, so overlap and broken images are measured, not guessed from a screenshot.
|
|
71
|
+
- ๐ก๏ธ **Safety is enforced on the network, not requested in a prompt.** Nothing existing is changed unless you allow it, and a blocked write never reaches your server.
|
|
72
|
+
- โ
**"Done" is a contract.** The report lists everything not tested, and at the `extensive` level refuses to finish while any known page is unvisited.
|
|
73
|
+
- ๐ง **It remembers.** Each run starts from what the last one learned, and re-tests the bugs earlier runs left open.
|
|
136
74
|
|
|
137
|
-
|
|
75
|
+
## ๐ Get started
|
|
138
76
|
|
|
139
|
-
|
|
77
|
+
**1. Install** for your agent (Node 20 or newer):
|
|
140
78
|
|
|
141
79
|
```bash
|
|
142
|
-
npx -y scenescout install # Claude Code: skill
|
|
143
|
-
npx -y scenescout install --client cursor # or
|
|
80
|
+
npx -y scenescout install # Claude Code: skill, MCP server and the test browser
|
|
81
|
+
npx -y scenescout install --client cursor # or vscode, codex, gemini, copilot, windsurf
|
|
82
|
+
npx -y scenescout doctor # every line should be a โ
|
|
144
83
|
```
|
|
145
84
|
|
|
146
|
-
Either way it downloads the browser and registers the server with the client you named. Claude Code also gets the method as a skill; every other client receives the same method from the server. [What each client gets](#-other-mcp-clients).
|
|
147
|
-
|
|
148
|
-
**Prefer a Claude Code plugin?** The skill and the server arrive together:
|
|
149
|
-
|
|
150
|
-
```
|
|
151
|
-
/plugin marketplace add brunoboto96/SceneScout
|
|
152
|
-
/plugin install scenescout@scenescout-marketplace
|
|
153
|
-
```
|
|
154
|
-
|
|
155
|
-
Then start a new chat to use SceneScout; the test browser downloads on first use (to have it ready beforehand: `npx -y scenescout install --browser-only`). The command becomes `/scenescout:scenescout`. A plugin's skill comes from this repository and its server from the latest npm release, so right after a release lands here the two can differ for a short while; `/plugin marketplace update scenescout-marketplace` brings the skill up to date.
|
|
156
|
-
|
|
157
|
-
An optional second plugin, `scenescout-mod@scenescout-marketplace`, adds a run pane (`/scenescout-pane`) and a setting for the model lane agents run on, in the Claude Code CLI and the desktop Code tab. It is a mod: unsandboxed JavaScript that runs inside Claude Code, so it is opt-in. [The Claude Code mod](docs/guide/Ways-to-use-it.md#the-claude-code-mod).
|
|
158
|
-
|
|
159
|
-
**Using Claude Desktop?** Install the extension: download `scenescout-X.Y.Z.mcpb` from the [latest release](https://github.com/brunoboto96/SceneScout/releases/latest) and open it (or Settings > Extensions > Advanced settings > Install Extension). It works as soon as it is installed, with no terminal step: the test browser downloads on first use. Start a new chat and ask *"Use SceneScout to test http://localhost:3000"*. [More in the guide](docs/guide/Start-here.md#as-a-claude-desktop-extension).
|
|
160
|
-
|
|
161
|
-
**A client that is not in that list?** [Add the server to its config by hand](#-other-mcp-clients); the test browser downloads on first use.
|
|
162
|
-
|
|
163
85
|
<details>
|
|
164
|
-
<summary>
|
|
86
|
+
<summary>Other ways to install: a Claude Code plugin, the Claude Desktop extension, or by hand</summary>
|
|
165
87
|
|
|
166
|
-
|
|
167
|
-
|
|
168
|
-
|
|
169
|
-
4. puts the `scenescout` command on your PATH, so `scenescout status`, `scenescout watch` and `scenescout doctor` work from any terminal. Run through `npx`, that is `npm install -g` of the version you just ran; from a clone it is `npm link`, so the command always runs what you last built. If npm refuses (a system-wide node usually needs `sudo` for this), the step prints the command to run by hand and the rest of the setup still counts as done: `npx -y scenescout <command>` works without it.
|
|
170
|
-
|
|
171
|
-
Re-run it any time: after moving the folder or switching node versions it refreshes the stored paths. It exits non-zero if a step the tool depends on failed, so it is safe to chain. Opt out of a step with `--no-register`, `--skip-browser` or `--no-command`.
|
|
172
|
-
|
|
173
|
-
If `claude` isn't on the PATH of the shell you ran it from, it prints the registration command instead of running it:
|
|
174
|
-
|
|
175
|
-
```bash
|
|
176
|
-
claude mcp add --scope user scenescout -- npx -y scenescout serve
|
|
177
|
-
```
|
|
88
|
+
- **Claude Code plugin:** `/plugin marketplace add brunoboto96/SceneScout`, then `/plugin install scenescout@scenescout-marketplace`. The command becomes `/scenescout:scenescout`.
|
|
89
|
+
- **Claude Desktop:** download `scenescout-X.Y.Z.mcpb` from the [latest release](https://github.com/brunoboto96/SceneScout/releases/latest) and open it. No terminal needed.
|
|
90
|
+
- **Any other MCP client:** add a stdio server whose command is `npx -y scenescout serve`. [Each client's config](docs/guide/Start-here.md#a-client-that-install-does-not-know).
|
|
178
91
|
|
|
179
92
|
</details>
|
|
180
93
|
|
|
181
|
-
|
|
182
|
-
|
|
183
|
-
```bash
|
|
184
|
-
npx -y scenescout doctor --engine # any client: node + build + browser
|
|
185
|
-
npx -y scenescout doctor # Claude Code: the above, plus the skill and the registration
|
|
186
|
-
```
|
|
187
|
-
|
|
188
|
-
Every line should be a โ. Anything that isn't prints the exact command that fixes it. Then **start a fresh session** in your client so it picks up the new tools.
|
|
189
|
-
|
|
190
|
-
### 3๏ธโฃ Run it
|
|
191
|
-
|
|
192
|
-
No app handy? Clone this repository and run `npm run demo:serve`: the [demo app](demo-app/) starts on `http://127.0.0.1:4173`.
|
|
193
|
-
|
|
194
|
-
Open your agent inside the project you want to test (or, for a [remote URL](#-two-ways-to-use-it), any folder) and ask:
|
|
94
|
+
**2. Start your app**, then a fresh session of your agent, and ask:
|
|
195
95
|
|
|
96
|
+
```text
|
|
97
|
+
Use SceneScout to test http://localhost:3000
|
|
196
98
|
```
|
|
197
|
-
Use SceneScout to test http://localhost:3000 at medium level
|
|
198
|
-
```
|
|
199
|
-
|
|
200
|
-
In Claude Code the skill gives you a command with flags for the same thing:
|
|
201
99
|
|
|
202
|
-
|
|
203
|
-
/scenescout --level medium --url http://localhost:3000 --role qa
|
|
204
|
-
```
|
|
100
|
+
In Claude Code there is a command too: `/scenescout --url http://localhost:3000 --level medium`. On its own, `/scenescout` asks you four plain questions instead: where the app is, how you sign in, what to check, and whether it holds real data.
|
|
205
101
|
|
|
206
|
-
|
|
102
|
+
**3. Read the report** in `.scenescout/report.md`. It opens in plain words (each problem, its steps, what was expected and what happened), with the technical detail one click away. [A complete example](examples/report.md).
|
|
207
103
|
|
|
208
|
-
**
|
|
104
|
+
**No app handy?** Clone this repository and run `npm run demo:serve`: the [demo app](demo-app/) starts on `http://127.0.0.1:4173`, and its README lists every planted bug.
|
|
209
105
|
|
|
210
|
-
**
|
|
211
|
-
|
|
212
|
-
### ๐ Signing in as a role
|
|
213
|
-
|
|
214
|
-
For an app behind SSO or MFA, sign in once yourself and let every session reuse it:
|
|
106
|
+
**Behind a sign-in?** Sign in once yourself, SSO and MFA included, and every session reuses it:
|
|
215
107
|
|
|
216
108
|
```bash
|
|
217
|
-
scenescout login http://localhost:3000 --role admin
|
|
218
|
-
```
|
|
219
|
-
|
|
220
|
-
A browser window opens at the URL. Sign in however the app asks: once you are back on the app with a new session, the window saves it as `.scenescout/auth/admin.json` in the project and closes by itself. A trip through a single sign-on provider and back is followed, not taken for the end. Pressing **Enter** in the terminal saves at once, and `--save enter` makes Enter the only way, as before. Closing the window or pressing Ctrl+C saves nothing. From a conversation, the agent opens the same window with `scout_login`, so no terminal is needed. The file is readable by your account only, `.scenescout/` keeps itself out of git, and the command prints where it saved, how many cookies, origins and databases it holds, never what they are, and how long it will last: read from each cookie's expiry and the `exp` of any JWT in a cookie or in localStorage (the payload is decoded for that one claim, never verified, never printed). The profile keeps cookies, localStorage, IndexedDB and sessionStorage, so an app whose sign-in library keeps its token in sessionStorage or IndexedDB still comes back signed in; sessionStorage is put back only on the origin it came from, once per tab, so a lane that signs out stays signed out. A login saved by an earlier version has no sessionStorage or IndexedDB: record it again if the app keeps its token there. `--project <dir>` saves into another project; `--browser firefox|webkit` records in another browser.
|
|
221
|
-
|
|
222
|
-
Then `/scenescout --role admin`, or `scout_attach { role: "admin" }` from any agent. Every session attached with the same role gets its own browser built from that one login, so parallel lanes can all run as `admin`. A role with no saved login is refused with the command to run. `role` and `storageStatePath` are alternatives: pass one.
|
|
223
|
-
|
|
224
|
-
Sessions of one role share one saved login, so they share its refresh token too. An app that rotates refresh tokens and treats a second use of a spent one as theft would revoke the whole token family, and sign every session of that role out, the moment two of them refreshed with the same token. SceneScout stops that for a session attached by role. When the page is about to send a refresh token from the role's profile, the session first takes a lock beside the profile (`.scenescout/auth/<role>.json.lock`, owner-only, taken over if its holder has not touched it in 30 seconds). Holding the lock, it re-reads the profile: if another session has rotated the token in the meantime, it loads that profile into its own browser and sends the current token in place of the spent one. Once the page has stored the rotated token, the session writes its state back over the profile and releases the lock. Sessions in separate processes share the lock through the file. A refresh token is recognised by name (a cookie, a storage key, or a field inside a JSON storage value whose name contains `refresh`) and is never printed or logged. A token that is only in a cookie makes a request a refresh only when the request is plausibly the refresh call, a POST to a path named for one or an endpoint seen to rotate the cookie, so a refresh cookie scoped to `/` never holds back the app's other requests; scripts, stylesheets, images and fonts are never brokered, and a broker that cannot do its job lets the request through unchanged. See [the guide](docs/guide/Signing-in.md#sessions-that-share-a-refresh-token). An app whose sign-in renews through the identity provider's own session cookie needs none of this, since no refresh token is shared. `SCENESCOUT_REFRESH_BROKER=off` turns the broker off.
|
|
225
|
-
|
|
226
|
-
In CI, where nobody can type, `--script` signs in headless as a test user from `SCENESCOUT_LOGIN_USERNAME`, `SCENESCOUT_LOGIN_PASSWORD` and, for a one-time code, `SCENESCOUT_LOGIN_TOTP_SECRET` or a fixed code the test environment accepts in `SCENESCOUT_LOGIN_OTP_CODE` (with no password for a passwordless sign-in), and saves the same profile. No credential value is ever printed. See [signing in from CI](docs/ci.md#signing-in-from-ci) for the options and the rules: a test tenant's user, never production or a real person's account.
|
|
227
|
-
|
|
228
|
-
Before a parallel run, `scout_lane_brief` checks that the planner's saved login will outlast it: `runMinutes` (default 60) plus `expiryMarginMinutes` (default 10). It refuses only when it is sure, meaning every credential in the profile has a date, none was set for another host, and the last of them ends before the run does, and then names the `scenescout login` command to run again. A profile holds cookies other than the sign-in (analytics, preferences), so the first one to expire is reported as a warning rather than a reason to refuse, and a profile with undated credentials in it (a session cookie, or a refresh token with no expiry) is a warning that its lifetime is unknown.
|
|
229
|
-
|
|
230
|
-
---
|
|
231
|
-
|
|
232
|
-
## ๐บ Watching a run live
|
|
233
|
-
|
|
234
|
-
When a session attaches, the engine starts a small live view and hands the agent its address on a `Live view:` line, which the agent passes on to you. On a local desktop it also opens that page in your default browser as the session attaches, and opens `report.html` when `scout_report` writes it, whether or not the browser window is shown. Nothing opens in CI, over SSH, or on Linux with no display. `SCENESCOUT_OPEN` (`live`, `report`, `both` or `none`) in the server's environment chooses otherwise, and `scout_attach {open}` wins over it. `scenescout ci` opens nothing unless `SCENESCOUT_OPEN` is set. From a terminal, `scenescout watch` opens the same page. There is one card per session:
|
|
235
|
-
|
|
236
|
-
<p align="center"><img src="examples/screenshots/live-view.png" alt="The live view during a run of three parallel agents against the demo app: one card per session, each with its role and objective, the task it is on, the tool it is running, the page it is on, a live thumbnail, and a feed of the actions it just took, tinted one colour per task" width="880" /></p>
|
|
237
|
-
|
|
238
|
-
- **What it is doing:** the tool it is running and for how long, the page it is on, and a thumbnail of that page. This works for headless runs too, which have no window to look at.
|
|
239
|
-
- **What it just did:** a rolling feed of its actions, each with its target and how it turned out, with failures in red. It is the same trail a finding's repro trace uses. The engine never sees the agent's reasoning, so this is what the session *did*, not what it thought.
|
|
240
|
-
- **Stuck, not slow:** a call still running past its own tool's watchdog budget turns the card red, so a wedged session is visible without asking. A crawl legitimately runs for minutes; it is judged against the crawl's budget, not a click's.
|
|
241
|
-
- **Live stream:** switch it on for one card, or for all of them. Click a thumbnail for a close-up.
|
|
242
|
-
- **The report, as it stands:** the Report button in the top bar shows the same document `scout_report` writes at the end, rendered from the run's current state, so findings can be read while the agents are still working.
|
|
243
|
-
- **What it is for:** the close-up puts the feed beside the session's brief โ the objective it was given when it attached (`scout_attach {objective}`), and underneath it the task it is on right now (`scout_task`), which the engine requires before any tool will act. Each task tints its own block of actions, so a change of task is a change of colour; point at a block and the brief names the task those actions served.
|
|
244
|
-
- **Scrub it back:** under the page is a tick per action, coloured by task. Click one to see the frame from that moment, and `Back to live` to return. On a run that was not recorded the ticks still read the trail; they just have no picture behind them.
|
|
245
|
-
|
|
246
|
-
<p align="center"><img src="examples/screenshots/live-view-closeup.png" alt="A close-up of one session: a frame from a step picked out of the timeline, the timeline itself as a tick per action coloured by task, the feed of the session's actions in the same colours, and beside it the objective and the task it is on" width="880" /></p>
|
|
247
|
-
|
|
248
|
-
<p align="center"><img src="examples/screenshots/live-view-report.png" alt="The report opened from the live view's top bar while the run is still going: summary table, gap ledger, and the findings filed so far, each with an accordion of the screenshots taken around it" width="880" /></p>
|
|
249
|
-
|
|
250
|
-
The view is served on `127.0.0.1` only, behind a token that changes every time the engine starts. It answers `GET` and nothing else, so a viewer can watch a run but not act in it, and no frame it shows is written to disk ([ADR 7](docs/adr/0007-the-live-view-is-local-read-only-and-leaves-nothing-behind.md)) unless the run was recorded, which is asked for and off by default ([ADR 8](docs/adr/0008-a-recorded-run-is-evidence-and-must-be-asked-for.md)). A stream runs only while someone is watching it. `SCENESCOUT_LIVE=off` keeps the port closed.
|
|
251
|
-
|
|
252
|
-
**Try it with parallel agents.** The demo app has three roles and several separate areas, so a run can be split between agents. Start it with `npm run demo:serve`, then ask your agent to explore it with several agents in parallel, one role and one area each. The pictures above come from a run of three. Two things keep a parallel run efficient:
|
|
253
|
-
|
|
254
|
-
- **Each agent opens its own session when it starts, and the planner closes it once it has folded that agent's report.** An agent waiting for its turn then holds no browser. Opening every session up front leaves browsers idling while the machine runs out of memory for the agents that are working. Closing before the fold loses the lane's decisions, which have nowhere to be kept.
|
|
255
|
-
- **Slow it down to follow along.** `scout_attach {paceMs}` (or `scout_session {paceMs}` mid-run) sets a floor between actions, for when you want to watch a flow rather than let it run as fast as the page allows.
|
|
256
|
-
- **Run about as many agents at once as your machine has cores, less two.** Each one drives a real browser.
|
|
257
|
-
|
|
258
|
-
---
|
|
259
|
-
|
|
260
|
-
## ๐ฌ Recording a run, and reading it back
|
|
261
|
-
|
|
262
|
-
A report says what happened. For QA work that is not always enough โ the point
|
|
263
|
-
is often to *show* what was checked, not to assert it.
|
|
264
|
-
|
|
265
|
-
Every finding already carries a picture: the element it is about, with a margin,
|
|
266
|
-
or the page as it was. It is kept in `.scenescout/recordings/`, shown under the
|
|
267
|
-
finding in `report.html`, and returned with the `scout_finding` result, so a
|
|
268
|
-
chat client shows the evidence the moment it is filed. Pictures are bounded in
|
|
269
|
-
size and in how many reach the conversation, and a CI job keeps them on file
|
|
270
|
-
only; `SCENESCOUT_EVIDENCE` and `scout_attach {evidence}` change that
|
|
271
|
-
([configuration reference](docs/guide/Configuration-reference.md#environment-variables)).
|
|
272
|
-
|
|
273
|
-
Ask for a recorded run and the engine also keeps a frame of the page after every action:
|
|
274
|
-
|
|
275
|
-
```
|
|
276
|
-
Use SceneScout to test http://localhost:3000, record the run
|
|
277
|
-
```
|
|
278
|
-
|
|
279
|
-
or, on the tool directly, `scout_attach {record: true}`. `SCENESCOUT_RECORD=on` in
|
|
280
|
-
the server's environment records every run. A CI gate records too:
|
|
281
|
-
`scenescout check --record` writes `replay.html`, every journey step by step with
|
|
282
|
-
its frames ([recording a check](docs/ci.md#recording-a-check)), and
|
|
283
|
-
`--template <file.json>` also writes it up as a test report laid out as a template
|
|
284
|
-
says, with expected and actual results, deviations, blank sign-off rows and a
|
|
285
|
-
SHA-256 manifest of the evidence ([test reports](docs/ci.md#a-test-report-from-a-template)).
|
|
286
|
-
|
|
287
|
-
Then `scout_report` writes two files side by side in `.scenescout/`:
|
|
288
|
-
`report.md` as always, and `report.html` โ the whole run as one self-contained
|
|
289
|
-
page. It opens from the file system with nothing running, needs no network, and
|
|
290
|
-
holds:
|
|
291
|
-
|
|
292
|
-
- **The report**, rendered from the same Markdown: the plain-language view
|
|
293
|
-
first (each problem's steps, what was expected, what happened and its
|
|
294
|
-
picture), with each problem's technical detail one click away.
|
|
295
|
-
- **The screenshots around each finding**, in an accordion under it, from the
|
|
296
|
-
session that filed it.
|
|
297
|
-
- **Every session's trail**, in the blocks its tasks made, each step with the
|
|
298
|
-
page as it was at that moment.
|
|
299
|
-
|
|
300
|
-
<p align="center"><img src="examples/screenshots/run-page.png" alt="The saved copy of a run, opened from the file system with nothing running: a finding with its evidence accordion open, showing the four screenshots taken around it with the action and time under each" width="880" /></p>
|
|
301
|
-
|
|
302
|
-
The live view serves the same document at `run` while the engine is still up,
|
|
303
|
-
and sends you there when the run ends โ so the address survives a refresh
|
|
304
|
-
instead of a panel over a dead board.
|
|
305
|
-
|
|
306
|
-
**What it costs.** Frames are pictures of the app under test, inside the tested
|
|
307
|
-
project's folder, and the secret redaction that protects everything else the
|
|
308
|
-
engine writes cannot read a picture. That is why it is off unless asked for,
|
|
309
|
-
capped per session, and written only under `.scenescout/`, which ignores itself
|
|
310
|
-
so `git add -A` in the tested project cannot pick the frames up. The reasoning is in
|
|
311
|
-
[ADR 8](docs/adr/0008-a-recorded-run-is-evidence-and-must-be-asked-for.md).
|
|
312
|
-
|
|
313
|
-
---
|
|
314
|
-
|
|
315
|
-
## ๐ How a run works
|
|
316
|
-
|
|
317
|
-
One curiosity loop, repeated โ breadth first, then judgment where it matters:
|
|
318
|
-
|
|
319
|
-
```
|
|
320
|
-
scan โโโถ attach โโโถ crawl โโโถ investigate โโโถ measure โโโถ report
|
|
321
|
-
โ โ โ โ โ โ
|
|
322
|
-
routes browser every route reproduce & journeys + gap-checked
|
|
323
|
-
& auth (r/o) in ONE call file findings design audit markdown
|
|
109
|
+
npx -y scenescout login http://localhost:3000 --role admin # then: /scenescout --role admin
|
|
324
110
|
```
|
|
325
111
|
|
|
326
|
-
|
|
327
|
-
2. **Attach** a browser (read-only unless you said otherwise).
|
|
328
|
-
3. **Crawl** every known route in a *single* call โ per-route HTTP status, element counts, oracle violations, dead ends.
|
|
329
|
-
4. **Investigate** what the crawl flagged: navigate, snapshot, reproduce, file a structured finding.
|
|
330
|
-
5. **Measure** task ease (`scout_journey`) and design quality (`scout_design_audit`) on representative pages.
|
|
331
|
-
6. **Report** โ the engine checks the gap ledger and writes `.scenescout/report.md`.
|
|
112
|
+
[Signing in](docs/guide/Signing-in.md) covers roles, expiry and scripted sign-in for CI.
|
|
332
113
|
|
|
333
|
-
|
|
114
|
+
## ๐ฆ In CI
|
|
334
115
|
|
|
335
|
-
|
|
336
|
-
|
|
337
|
-
## ๐งฐ The toolbox
|
|
338
|
-
|
|
339
|
-
35 deterministic tools. The agent picks; you rarely call these by hand.
|
|
340
|
-
|
|
341
|
-
| Phase | Tools | What they do |
|
|
116
|
+
| | What it does | Needs a model? |
|
|
342
117
|
|---|---|---|
|
|
343
|
-
|
|
|
344
|
-
|
|
|
345
|
-
|
|
|
346
|
-
|
|
|
347
|
-
| **Act** | `scout_click` `scout_type` `scout_select` `scout_upload` `scout_press` `scout_scroll` `scout_navigate` `scout_back` `scout_run_plan` | Drive the UI like a user; `scout_run_plan` batches a whole mechanical sequence into one call |
|
|
348
|
-
| **Assess** | `scout_design_audit` `scout_journey` | Score a page's craft/a11y/consistency; measure how hard a task is to complete |
|
|
349
|
-
| **Record** | `scout_note` `scout_finding` `scout_resolve` `scout_report` | Curate durable notes; file deduped findings; mark fixes; write the report, and on a recorded run the whole run as one page |
|
|
350
|
-
| **Answer tickets** | `scout_tickets` `scout_criterion` | Read the acceptance criteria in pasted or uploaded tickets; record each criterion as passed, failed (with the findings that show it) or not tested (and why), with a confidence |
|
|
351
|
-
| **Re-test** | `scout_verify` | List the findings earlier runs left open, worst route first, and record whether each is gone, still present, or changed |
|
|
352
|
-
| **Split the work** | `scout_lane_brief` `scout_lane_report` | Divide the app between parallel agents by whole module, each with its own landing route and rules; fold what each hands back as one typed JSON object, and name any defect it judged but never filed |
|
|
353
|
-
| **Close** | `scout_close` | Tear down one session or all |
|
|
354
|
-
|
|
355
|
-
A few that punch above their weight:
|
|
356
|
-
|
|
357
|
-
- **`scout_crawl`** โ the entire breadth pass in one tool call. No visiting routes one-by-one.
|
|
358
|
-
- **`scout_run_plan`** โ up to 20 actions (fill form โ submit โ check) with semantic targets (`testid=โฆ`, `text=โฆ`), aborting at the first anomaly.
|
|
359
|
-
- **`scout_journey`** โ wraps one goal and reports interaction count, screens seen, and **backtracks**; an abandoned journey is a finding no passing E2E suite can produce.
|
|
360
|
-
- **`scout_upload`** โ generates a *valid* in-memory fixture (real PDF/PNG, kind inferred from `accept`) so file-upload flows stop being a blind spot.
|
|
361
|
-
- **`scout_click {clicks: 2}`** โ the impatient-user probe: states whether a double-click fired the same state-changing request twice (the classic double-submit bug).
|
|
362
|
-
- **`scout_request`** โ calls the app's own API as the session, so "the button is hidden" becomes "the server refuses it" (or doesn't).
|
|
363
|
-
|
|
364
|
-
Beyond crashes and HTTP errors, two oracles catch a page **contradicting the server**: `refused_empty` (a list request was refused and the page shows its empty state with no error) and `false_success` (a save was refused and the page says it worked). A third, `dom_injection`, reports a typed markup value coming back as an element on any page any session opens. A fourth, `postmessage_token`, reports a page calling `postMessage` with targetOrigin `"*"` on a message that carries a token (a JWT, a `Bearer` value, or an opaque value under a key such as `access_token`): the report names where in the message it was, its shape and its first four characters, never the token.
|
|
365
|
-
|
|
366
|
-
---
|
|
367
|
-
|
|
368
|
-
## ๐ Test levels
|
|
369
|
-
|
|
370
|
-
Each level is a contract. `scout_report` enforces what the engine can see for itself โ routes visited, pages audited, and at `extensive` an empty gap ledger (every route exercised and audited, every filled form submitted, a completed journey, two roles) โ and the report's gap ledger discloses the rest.
|
|
371
|
-
|
|
372
|
-
| Level | What it guarantees | Rough size |
|
|
373
|
-
|---|---|---|
|
|
374
|
-
| `minimal` | Every route visited, โฅ1 design audit, key journeys as plans, crawl problems triaged. Remaining gaps **disclosed**. | ~40 actions |
|
|
375
|
-
| `medium` *(default)* | minimal + design audits across several routes (enforced) + every element class exercised + every form submitted valid **and** invalid (asked of the agent) | ~150 actions |
|
|
376
|
-
| `extensive` | medium + fuzzing, back/refresh/deep-link resilience, keyboard-only pass, a journey per module, โฅ2 roles compared, anonymous auth-surface walk. **Refuses to finalize while any gap in the ledger remains.** | budget-capped |
|
|
377
|
-
|
|
378
|
-
That refusal *is* the guarantee: an extensive report can only exist when nothing the engine can measure was left untested. Fuzzing, the keyboard pass and the auth-surface walk are the agent's to do; the engine cannot see whether they were done well.
|
|
379
|
-
|
|
380
|
-
---
|
|
381
|
-
|
|
382
|
-
## ๐ Safety model
|
|
383
|
-
|
|
384
|
-
- ๐ต **`observe`** (`--observe`) lets nothing but `GET` requests leave the page. The one exception is what a session needs in order to exist: logging in, logging out and refreshing a token. Signing up, changing or resetting a password and creating users are blocked like any other write. WebSocket frames are not inspected; the engine says so when the app opens a socket. It is what the skill picks for a remote URL with no source, where an ordinary form POST would create a real record. Forms that could not be submitted are listed in the gap ledger.
|
|
385
|
-
- ๐ข **`read-only` by default.** Destructive-labeled elements (delete/revoke/archive/โฆ) **and** all `PUT/PATCH/DELETE` + destructive `POST`s are blocked at the network layer โ see [`src/engine/policy.ts`](src/engine/policy.ts). Non-destructive `POST`s are allowed, because submitting forms is how a tester finds validation bugs โ so read-only means *nothing existing is changed or removed*, not *nothing is ever created*.
|
|
386
|
-
- ๐ก **`safe-write`** (`--safe-write`) lets the agent create data and edit/delete **only what it created** this run โ never pre-existing records.
|
|
387
|
-
- ๐ด **`destructive`** (`--allow-destructive`) allows everything, and only ever when *you* confirm the environment is disposable. The skill will never choose this itself.
|
|
388
|
-
- ๐ Findings, memory, and reports live in a `.scenescout/` folder in the project. A client with no project folder, such as a desktop chat, gets one folder per tested site under `Documents/SceneScout/<host>/` by default, and the attach says where; `SCENESCOUT_PROJECTS_DIR` moves it. The folder ignores itself in git, so a stray `git add -A` never commits test data.
|
|
389
|
-
|
|
390
|
-
A `๐ก WRITE-POLICY blocked` notice is the safety net doing its job, not an app bug. The server never sees a blocked request, but a page's own `fetch` or XHR is answered with a `403` in its place rather than dropped, so the page's handling of a refusal really runs: a page that then claims success is reported as a `false_success` ([ADR 9](docs/adr/0009-a-refused-write-is-answered-not-dropped.md)).
|
|
391
|
-
|
|
392
|
-
A control is judged by its own label: a dropdown by the option picked, a row by its own text rather than the buttons inside it, and "Discard changes" on an unsent form is allowed. When a page asks to confirm leaving unsent input, the result says so; `observe` and `read-only` stay unless the call passes `leave: true`. See the [safety model](docs/guide/Safety-model.md).
|
|
393
|
-
|
|
394
|
-
---
|
|
395
|
-
|
|
396
|
-
## ๐ What you get
|
|
397
|
-
|
|
398
|
-
`.scenescout/report.md` and `report.html` open **In plain words**: a short summary, then each problem this run found, worst first, with its impact, the steps that led to it, what was expected, what happened and a picture when there is one. The technical detail (id, category, evidence, route) stays one click away.
|
|
399
|
-
|
|
400
|
-
`.scenescout/report.md` โ a deduplicated, worst-first report with:
|
|
401
|
-
|
|
402
|
-
- ๐ **Findings** with repro traces and generated Playwright regression-test skeletons.
|
|
403
|
-
- ๐ **Worth a look** โ observations that are defects only under a convention of your project the run cannot see (a spacing scale, link styling in navigation, test ids on every control), each naming that convention. Listed below the findings and not counted as defects ([ADR 13](docs/adr/0013-a-convention-is-the-projects-to-decide.md)).
|
|
404
|
-
- ๐ฏ **Page scores** (0โ100: a11y ยท craft ยท consistency ยท task-clarity), ranked worst-first, with stale scores from old runs marked as such.
|
|
405
|
-
- ๐ฅ **A role capability matrix** โ what each role could and couldn't reach.
|
|
406
|
-
- ๐งพ **A gap ledger** โ everything *not* done, so the report is honest about its own coverage.
|
|
407
|
-
- โฑ๏ธ **How the run was paced** โ how closely each session kept working, and apart from that, how long finished lanes held their browsers waiting to be collected, so neither hides the other.
|
|
408
|
-
- ๐ฏ **How well the lanes judged** โ on a parallel run, whether the confidence each lane stated matched what the project went on to file, beside what later re-tests found ([ADR 10](docs/adr/0010-a-confidence-is-checked-not-trusted.md)).
|
|
409
|
-
|
|
410
|
-
`.scenescout/report.html` โ the same report as one self-contained page, with every session's trail beside it. Each finding shows a picture of the element it is about, or of the page as it was. A [recorded run](#-recording-a-run-and-reading-it-back) also shows the screenshots around each finding.
|
|
411
|
-
|
|
412
|
-
๐ Watch a run live: `npx scenescout watch`, or `npx scenescout status <project-path>` for the same information as text.
|
|
413
|
-
|
|
414
|
-
---
|
|
415
|
-
|
|
416
|
-
## ๐ฆ In CI: a deterministic check
|
|
417
|
-
|
|
418
|
-
An exploratory run is driven by a model, so two runs never find exactly the same things. That suits a report, but not a gate. `scenescout check` is the part that needs no model. It visits the start URL, the project's scanned routes and every same-origin link it finds, and measures each page:
|
|
419
|
-
|
|
420
|
-
- HTTP and page errors
|
|
421
|
-
- layout geometry (covered, clipped and overlapping controls, blocking overlays)
|
|
422
|
-
- broken images
|
|
423
|
-
- controls with no name, and fields whose only label is a placeholder
|
|
424
|
-
- contrast and focus
|
|
425
|
-
- pages with no way out
|
|
118
|
+
| [`scenescout check`](docs/guide/Ways-to-use-it.md#scenescout-check-a-gate-in-ci) | A deterministic gate: measures every page, replays your saved flows and visual baselines, fails only on what it can prove | No |
|
|
119
|
+
| [`scenescout ci`](docs/guide/Ways-to-use-it.md#scenescout-ci-an-unattended-exploratory-run) | An unattended exploratory run, driven by the Anthropic or OpenAI API. It reports and never fails the build | An API key |
|
|
120
|
+
| [`/scenescout qa`](docs/guide/Ways-to-use-it.md#scenescout-qa-on-a-pull-request) | A comment on a pull request that tests its preview deploy and replies with the results; `/scenescout qa check` runs your own check instead | An API key (`qa check`: none) |
|
|
121
|
+
| [`scenescout export`](docs/guide/Ways-to-use-it.md#filing-findings-as-issues) | Files the findings as GitHub or Jira issues, each once | No |
|
|
426
122
|
|
|
427
|
-
|
|
428
|
-
|
|
429
|
-
```bash
|
|
430
|
-
npx scenescout check http://127.0.0.1:3000 --fail-on high
|
|
431
|
-
```
|
|
432
|
-
|
|
433
|
-
| Exit code | Meaning |
|
|
434
|
-
|---|---|
|
|
435
|
-
| 0 | Passed the gate |
|
|
436
|
-
| 1 | Failed it: something at the `--fail-on` severity or worse |
|
|
437
|
-
| 2 | Could not run, or not all of it: a bad argument, an app that never answered, only the sign-in page reached, a saved flow that is not valid, or a flow step the write policy refused |
|
|
438
|
-
|
|
439
|
-
It writes `report.md`, `check.sarif` (for code-scanning dashboards) and `check.json` to `.scenescout/check/`, and on GitHub Actions it also puts the report on the job's summary page.
|
|
440
|
-
|
|
441
|
-
On GitHub Actions, this repository is also an action that installs everything and keeps the results:
|
|
123
|
+
A gate on every pull request, as a GitHub Action:
|
|
442
124
|
|
|
443
125
|
```yaml
|
|
444
|
-
- uses: brunoboto96/SceneScout@v3
|
|
126
|
+
- uses: brunoboto96/SceneScout@v3
|
|
445
127
|
with:
|
|
446
128
|
url: http://127.0.0.1:3000
|
|
447
129
|
```
|
|
448
130
|
|
|
449
|
-
[docs/ci.md](docs/ci.md) has
|
|
450
|
-
|
|
451
|
-
With the default settings its saved flows send no HTTP write (they replay under observe's rule), and its crawl runs under `--mode observe` or `read-only`; `--flow-writes allow` lets flows write as `--mode` allows. By default it fails only on facts that mean a page is broken: a page that did not load, an uncaught exception, a 5xx, a failure shown as success. Other options:
|
|
452
|
-
|
|
453
|
-
- `--fail-on medium` or `low` makes the gate stricter.
|
|
454
|
-
- `--ignore <rule>` drops a rule, worth-a-look rules included.
|
|
455
|
-
- `--paths /a,/b` checks only those pages.
|
|
456
|
-
- `--storage-state <file>` checks while signed in.
|
|
457
|
-
- `--flows <dir>` or `off` chooses which saved flows to replay; `--retest off` skips re-testing open findings.
|
|
458
|
-
- `--flow-writes never|allow` (default `never`): `never` replays flows under observe's rule whatever `--mode` says; `allow` replays them under `--mode`, so in `read-only` a flow's form submissions are sent to the target on every run.
|
|
459
|
-
- `--on-refused-step report|stop` (default `report`): `report` marks a flow whose step was refused "could not run", keeps every other verdict and exits 2; `stop` exits 2 at that step with no results.
|
|
460
|
-
- `--gate-retests never|high|all` (default `high`): which still-reproducing re-tested findings fail the gate.
|
|
461
|
-
- `--baseline off|compare|update` (default `off`), with `--baselines <dir>` and `--baseline-threshold <percent>` (default `0.1`, so small anti-aliasing noise between machines passes): visual baselines, below.
|
|
462
|
-
- `--record` (or `SCENESCOUT_RECORD=on`; the action's `record: on`) keeps a frame after each route visit and each saved-flow step and writes `replay.html` beside the report: each role, each journey with a pass or fail badge, each step with its caption, result and frame, the first failing step highlighted. `--video` (the action's `video: on`) adds a WebM of each journey, played on that page beside its steps. Both off by default. [docs/ci.md](docs/ci.md#recording-a-check) covers size, privacy and publishing it.
|
|
463
|
-
|
|
464
|
-
The defaults are what an unconfigured check does, for a first try or an AI agent running it unattended: its flows send no HTTP write and it never silently hides a result. Each setting is a choice for the project; the report and `check.json` print the values a check ran with.
|
|
465
|
-
|
|
466
|
-
`scenescout check --help` lists every option. Why the defaults are what they are: [ADR 11](docs/adr/0011-a-gate-is-deterministic-and-fails-only-on-what-it-can-prove.md).
|
|
467
|
-
|
|
468
|
-
It also replays the flows saved in `.scenescout/flows/*.json`, with no model: the steps `scout_run_plan` takes (navigate, click, type, select, press, and upload, which attaches a generated file) plus `expect-text`, `expect-element` (a control is visible, hidden, enabled, disabled, checked or unchecked), `expect-url` and `expect-request`, and `repeat` to run actions until a condition holds (page through a document until Continue is enabled). A flow can name the `role` it runs as, signed in with a profile `scenescout login --role` saved, so a journey that passes between people (one submits, another approves) is a sequence of flows. Values can come from the environment, `${env:NAME}`, so a code or a password stays in the CI's secret store and is masked in everything the check writes. A flow whose step breaks fails the gate, naming the flow and the step. And it re-tests the open findings earlier runs left in the project's memory that a page load can reproduce, reporting each as still reproducing or possibly fixed; by default a finding filed high that still reproduces fails the gate. [docs/ci.md](docs/ci.md#saved-flows) has the flow format; [ADR 12](docs/adr/0012-a-check-replays-saved-flows-and-reports-re-tests.md) says why it works this way.
|
|
469
|
-
|
|
470
|
-
With `--baseline compare` it also holds pages and elements to approved pictures: list them in a `targets.json`, take the baselines once with `--baseline update`, and a later check that finds one changed past `--baseline-threshold` fails the gate with the share of pixels changed and a diff picture beside the report. Baselines are kept per browser, in a git-ignored folder unless `--baselines` names one the project commits, and only `--baseline update` ever writes one. [The guide](docs/guide/Ways-to-use-it.md#visual-baselines) has the details; [ADR 19](docs/adr/0019-a-visual-baseline-changes-only-when-asked.md) says why.
|
|
471
|
-
|
|
472
|
-
Beyond those flows it explores nothing and fills no forms. That is the exploratory run's job, and its findings belong in a report, not a gate.
|
|
473
|
-
|
|
474
|
-
## ๐ค In CI: an unattended exploratory run
|
|
475
|
-
|
|
476
|
-
`scenescout ci` runs the exploratory side in a CI job, with no person and no coding agent. A model reached through its API drives the same `scout_*` tools by the same method, and the run ends in the ordinary report:
|
|
477
|
-
|
|
478
|
-
```bash
|
|
479
|
-
export OPENAI_API_KEY=โฆ # or ANTHROPIC_API_KEY; read from the environment only
|
|
480
|
-
npx scenescout ci http://127.0.0.1:3000
|
|
481
|
-
```
|
|
482
|
-
|
|
483
|
-
- **It reports and never gates.** Exit 0 when the run ran, whatever it found; exit 2 when it could not run (no key, a key the API refused, an app that never answered). Two runs of the same app find different things, so a finding is something to read, never a reason to fail a build. `scenescout check` is the gate.
|
|
484
|
-
- **Providers:** the Anthropic Messages API (default model `claude-sonnet-5`) or the OpenAI Responses API (default `gpt-6-luna`), chosen by which key is set; with both set, `--provider` decides. `--model` and `--effort` (default `low`) override; `--base-url` points at another endpoint that implements the same API.
|
|
485
|
-
- **Caps:** at most 80 model turns, 3,000,000 tokens and 20 minutes (`--max-turns`, `--max-tokens`, `--max-minutes`), shared by two model loops that each explore their own part of the app (`--lanes`, default 2). The first cap reached ends the exploration; the report is still written, and says which cap ended it. On the benchmark's demo app a run at these defaults cost about $0.03 on `gpt-6-luna` ([docs/benchmark.md](docs/benchmark.md#choosing-the-defaults-issue-419)).
|
|
486
|
-
- **Mode:** `read-only` by default; `--mode observe` sends no form at all, `--mode safe-write` lets the run create records and change only the ones it created. `--mode destructive` runs only with `--allow-destructive` as well.
|
|
487
|
-
- **Duplicates:** when the dedup rule keeps a filed finding apart, the run's model is asked at its lowest effort whether it is one already open on the same page, and merges it on a "same", keeping the filing's title, category, severity and evidence under that finding. The two findings' titles, categories and evidence, and the page's path, are sent; `--dedup rule` turns it off ([ADR 17](docs/adr/0017-a-model-judges-only-the-merges-the-rule-misses.md)).
|
|
488
|
-
- **Output**, in `.scenescout/ci/` (or `--out`): `report.md` and `report.html` (the report an agent's run writes), `summary.md` (also appended to the GitHub job summary), `ci.json` and `ci.sarif`, with a usage line: turns, tokens, time and an estimated cost where the model's price is known (`--price-in`, `--price-out` give one for any model).
|
|
489
|
-
|
|
490
|
-
There is a GitHub Action for it (`uses: brunoboto96/SceneScout/ci@โฆ`). [docs/ci.md](docs/ci.md#an-unattended-exploratory-run) has the workflow and every option; [ADR 14](docs/adr/0014-an-unattended-run-reports-and-never-gates.md) says why it works this way.
|
|
491
|
-
|
|
492
|
-
On a pull request, an allowed account can comment `/scenescout qa` to run it against that pull request's deployed preview and get the results as a reply. The job that holds the key checks out nothing and runs SceneScout from an exact release tag, so the pull request's code never runs beside the key. A project without previews can comment `/scenescout qa check` instead: it dispatches the project's own recorded `scenescout check` workflow on the pull request's branch and replies with the verdict, per journey, with no model key involved. [docs/ci.md](docs/ci.md#a-qa-review-from-a-pull-request-comment) has the workflows to copy and what a project configures; [ADR 15](docs/adr/0015-a-qa-comment-tests-a-preview-and-never-runs-the-pull-requests-code.md) says why.
|
|
131
|
+
[docs/ci.md](docs/ci.md) has complete workflows, every option, and the same check on GitLab CI, CircleCI or any shell.
|
|
493
132
|
|
|
494
|
-
##
|
|
495
|
-
|
|
496
|
-
`scenescout export` turns the project's open findings into GitHub or Jira issues, where the team already works. It reads `.scenescout/memory.json`, so it runs after an interactive run, `scenescout ci` or anything else that wrote findings:
|
|
497
|
-
|
|
498
|
-
```bash
|
|
499
|
-
export GH_TOKEN=โฆ # or GITHUB_TOKEN; for Jira, JIRA_EMAIL and JIRA_API_TOKEN. Read from the environment only
|
|
500
|
-
npx scenescout export --to github --repo owner/app # a dry run: lists what it would file
|
|
501
|
-
npx scenescout export --to github --repo owner/app --yes # files it
|
|
502
|
-
npx scenescout export --to jira --jira-url https://your-site.atlassian.net --jira-project QA --yes
|
|
503
|
-
```
|
|
133
|
+
## ๐ Safe by default
|
|
504
134
|
|
|
505
|
-
|
|
506
|
-
- **Jira issues are kept up to date and linked to the ticket.** A later export updates an open Jira issue instead of filing another, leaving text someone edited in Jira as they wrote it. The finding's picture is attached, and a finding that fails a ticket's acceptance criterion is linked to that ticket (`--jira-link-type`, default `Relates`).
|
|
507
|
-
- **A dry run unless `--yes`**, and at most `--max-issues` (default 20) per export; the next export files the rest. `--min-severity`, `--only <ids>` and `--include-worth-a-look` choose what goes.
|
|
508
|
-
- **Inert issues.** Titles, descriptions, steps and evidence come from the run and the app's pages, so no `@mention`, link, `#123` reference, HTML or Markdown in them does anything.
|
|
509
|
-
- **Severity** becomes a label on GitHub and a priority in Jira; `--severity-map` renames them or turns them off. **Screenshots** from a recorded run are attached in Jira; GitHub's API takes no uploads, so a GitHub issue names the frames in the run's `.scenescout/` folder.
|
|
510
|
-
- **Credentials** are never printed. Every request has a timeout, a short rate limit is waited out, failed reads are retried with backoff, and redirects are refused. Jira's search can take a little while to show a new issue, so leave a few minutes between two exports to the same Jira project.
|
|
511
|
-
|
|
512
|
-
[The guide](docs/guide/Ways-to-use-it.md#filing-findings-as-issues) has the details and a GitHub Actions step.
|
|
513
|
-
|
|
514
|
-
---
|
|
515
|
-
|
|
516
|
-
## ๐ฉบ Troubleshooting
|
|
517
|
-
|
|
518
|
-
Run `npx -y scenescout doctor` first โ it checks every setup item below (everything but the last row, which is about your app) and prints the fix.
|
|
519
|
-
|
|
520
|
-
| Symptom | Cause and fix |
|
|
135
|
+
| Mode | What may leave the page |
|
|
521
136
|
|---|---|
|
|
522
|
-
|
|
|
523
|
-
|
|
|
524
|
-
|
|
|
525
|
-
|
|
|
526
|
-
| *"โฆ build has not been downloaded yet"* on attach | The attach downloads a missing browser itself, once, except in CI or with `SCENESCOUT_BROWSER_DOWNLOAD=off`; there, or when that download failed, it names the command to run. Run the command the message names, for example `npx -y scenescout install --browser-only --browsers firefox`. On Linux, system libraries may be missing too: `npx playwright install --with-deps chromium`. |
|
|
527
|
-
| Tools broke after moving the folder or changing node version | The registration stores absolute paths. `npx -y scenescout install` refreshes them. |
|
|
528
|
-
| Attach fails or every route lands on the login page | Your app isn't running at `--url`, or the `--role` session has expired. For a saved login, run `scenescout login <url> --role <name>` again; for a storage-state file, regenerate it the way your project's Playwright setup does. |
|
|
529
|
-
|
|
530
|
-
### โฌ๏ธ Upgrading from an older version
|
|
137
|
+
| ๐ต `observe` | Reads only. The default for a first look, for `scenescout check`, and for an agent's run on a remote site with no source |
|
|
138
|
+
| ๐ข `read-only` | Reads and ordinary form posts; nothing existing is changed or deleted. The default for an agent's run on a local app |
|
|
139
|
+
| ๐ก `safe-write` | Creates records, and edits or deletes only the ones it created |
|
|
140
|
+
| ๐ด `destructive` | Everything. Only when you say the data is disposable; the agent never picks it |
|
|
531
141
|
|
|
532
|
-
|
|
533
|
-
- **Earlier names.** This tool was previously called SceneCraft (and, before that, frontend-tester). `scenescout install` cleans up after both: it removes the old skill link and the old `scenecraft` MCP registration when they point at this install, and the first attach in a project moves its `.scenecraft/` memory folder to `.scenescout/` so earlier coverage and findings carry over.
|
|
142
|
+
The policy sits on the network, so a blocked request never reaches your server. The page gets a refusal instead, which is how SceneScout catches a page that claims success anyway. Findings, memory and reports stay in a `.scenescout/` folder that keeps itself out of git.
|
|
534
143
|
|
|
535
|
-
|
|
536
|
-
|
|
537
|
-
```bash
|
|
538
|
-
# Claude Code
|
|
539
|
-
claude mcp remove --scope user scenescout
|
|
540
|
-
rm -rf ~/.claude/skills/scenescout
|
|
541
|
-
# Codex / Gemini / Copilot CLI
|
|
542
|
-
codex mcp remove scenescout # likewise: gemini mcp remove โฆ, copilot mcp remove โฆ
|
|
543
|
-
```
|
|
544
|
-
|
|
545
|
-
For Cursor, Windsurf and VS Code, delete the `scenescout` entry from the client's MCP server list.
|
|
546
|
-
|
|
547
|
-
Nothing else is installed: `npx` runs the package from npm's cache. Per-project memory lives in each tested project's `.scenescout/` folder; delete it there if you want it gone.
|
|
548
|
-
|
|
549
|
-
---
|
|
550
|
-
|
|
551
|
-
## ๐ Choosing browsers
|
|
552
|
-
|
|
553
|
-
`install` downloads Chromium and nothing else unless you ask. `--browsers` takes one name, a comma-separated list, or `all`:
|
|
554
|
-
|
|
555
|
-
| `--browsers` | What is downloaded | About, on disk |
|
|
556
|
-
|---|---|---|
|
|
557
|
-
| `chromium` *(default)* | the full browser and the headless shell | 550 MB |
|
|
558
|
-
| `chromium-headless-shell` | the headless shell only: every run works except `headed` | 200 MB |
|
|
559
|
-
| `firefox` | Firefox | 270 MB |
|
|
560
|
-
| `webkit` | WebKit, the engine behind Safari | 290 MB |
|
|
561
|
-
| `all` | Chromium, Firefox and WebKit | 1.1 GB |
|
|
562
|
-
|
|
563
|
-
```bash
|
|
564
|
-
npx -y scenescout install --browsers chromium-headless-shell # the smallest working setup
|
|
565
|
-
npx -y scenescout install --browser-only --browsers firefox,webkit # add two more later
|
|
566
|
-
```
|
|
567
|
-
|
|
568
|
-
Sizes vary by platform. The builds go to Playwright's shared cache, so a build another tool already fetched is not downloaded again.
|
|
569
|
-
|
|
570
|
-
The first attach that needs a build which is not on disk downloads it itself, once, and says so ("Getting the test browser ready"). It does not in CI unless `SCENESCOUT_BROWSER_DOWNLOAD=on` is set, and never with `SCENESCOUT_BROWSER_DOWNLOAD=off`, for a machine where nothing may be downloaded.
|
|
571
|
-
|
|
572
|
-
To drive another browser, pass `browser` when attaching (`scout_attach { browser: "firefox" }`), or set `SCENESCOUT_BROWSER=webkit` in the server's environment to change the default. `scenescout doctor` checks the browser named by that variable in the shell it runs from, so check another one with `SCENESCOUT_BROWSER=webkit scenescout doctor`. Two things differ outside Chromium:
|
|
573
|
-
|
|
574
|
-
- **Service workers are not allowed to register** in Firefox and WebKit. The write policy works by intercepting requests, and only Chromium lets a request issued by a service worker be intercepted. An app that depends on its worker may behave differently there.
|
|
575
|
-
- **A Firefox or WebKit left behind by a crash is not cleaned up** on the next start the way a leftover Chromium is.
|
|
576
|
-
|
|
577
|
-
In every browser, pages are not given shared workers unless the mode is `destructive`: a request a shared worker sends cannot be intercepted anywhere, so the app is made to do that work on the page, where the policy sees it.
|
|
578
|
-
|
|
579
|
-
### Time limits
|
|
580
|
-
|
|
581
|
-
An action on the page (a click, typing, a hover, a pick from a list) may take 5 s, and a page 20 s to load (15 s for a page the crawl opens). On a loaded machine these can run out while the app is fine; the timeout then says which limit ran out and how to raise it. Raise them per session with `scout_attach { actionTimeoutMs: 15000, navTimeoutMs: 60000 }`, for every session with `SCENESCOUT_ACTION_TIMEOUT_MS` and `SCENESCOUT_NAV_TIMEOUT_MS` in the server's environment, or on `scenescout check` and `scenescout ci` with `--action-timeout-ms` and `--nav-timeout-ms`. An option wins over the variable, and the variable over the default. The action limit takes 1000 to 120000 ms and the page-load limit 1000 to 300000 ms; anything else refuses the attach with a sentence naming the value to fix. Saving a login profile with `scenescout login` honours the two variables as well, and otherwise keeps its own longer waits (30 s for the page, 10 s for a field or the submit).
|
|
582
|
-
|
|
583
|
-
## ๐ Other MCP clients
|
|
584
|
-
|
|
585
|
-
The engine is a plain MCP server over stdio, so any client can drive it, and the testing method reaches the agent through the server itself (see the end of this section). `install` can register it for you:
|
|
144
|
+
> [!IMPORTANT]
|
|
145
|
+
> Only test sites you own or are allowed to test. [Safety model](docs/guide/Safety-model.md).
|
|
586
146
|
|
|
587
|
-
|
|
588
|
-
npx -y scenescout install --client cursor # one client
|
|
589
|
-
npx -y scenescout install --client vscode,codex # several; add claude-code to keep that one too
|
|
590
|
-
```
|
|
147
|
+
## ๐ Documentation
|
|
591
148
|
|
|
592
|
-
|
|
|
149
|
+
| | |
|
|
593
150
|
|---|---|
|
|
594
|
-
|
|
|
595
|
-
|
|
|
596
|
-
|
|
|
597
|
-
|
|
|
598
|
-
|
|
|
599
|
-
|
|
|
600
|
-
|
|
|
601
|
-
|
|
602
|
-
|
|
151
|
+
| [Start here](docs/guide/Start-here.md) | A first look, install, a first run, reading the report, the live view |
|
|
152
|
+
| [Ways to use it](docs/guide/Ways-to-use-it.md) | Interactive runs, parallel agents, CI, pull-request QA, filing issues |
|
|
153
|
+
| [What it checks](docs/guide/What-it-checks.md) | Every check and every `scout_*` tool |
|
|
154
|
+
| [Signing in](docs/guide/Signing-in.md) ยท [Safety model](docs/guide/Safety-model.md) | Roles and saved logins; what each mode refuses and why |
|
|
155
|
+
| [Recipes](docs/guide/Recipes.md) | Setups for seven kinds of project |
|
|
156
|
+
| [Configuration reference](docs/guide/Configuration-reference.md) | Every option, environment variable and action input |
|
|
157
|
+
| [Troubleshooting](docs/guide/Troubleshooting.md) | Symptoms and fixes, upgrading and uninstalling |
|
|
158
|
+
| [How it works](docs/how-it-works.md) | Diagrams of a run, an action, the write policy, lanes |
|
|
159
|
+
| [Benchmark](docs/benchmark.md) ยท [Validation](docs/validation.md) | How runs are scored against answer keys, and runs on public apps |
|
|
160
|
+
| [Design decisions](docs/adr/README.md) | Why the rules are what they are |
|
|
603
161
|
|
|
604
|
-
|
|
162
|
+
## ๐ง Contributing
|
|
605
163
|
|
|
606
|
-
|
|
164
|
+
Start with [VISION.md](VISION.md) (what is in scope) and [CONTRIBUTING.md](CONTRIBUTING.md) (local setup and how changes land). [AGENTS.md](AGENTS.md) holds the house rules for people and coding agents alike.
|
|
607
165
|
|
|
608
|
-
|
|
609
|
-
<summary><strong>Cursor</strong> โ <code>~/.cursor/mcp.json</code> (or <code>.cursor/mcp.json</code> in a project)</summary>
|
|
610
|
-
|
|
611
|
-
```json
|
|
612
|
-
{
|
|
613
|
-
"mcpServers": {
|
|
614
|
-
"scenescout": { "command": "npx", "args": ["-y", "scenescout", "serve"] }
|
|
615
|
-
}
|
|
616
|
-
}
|
|
617
|
-
```
|
|
618
|
-
|
|
619
|
-
</details>
|
|
166
|
+
Found a way past the write policy, or another security problem? Report it privately: [SECURITY.md](SECURITY.md).
|
|
620
167
|
|
|
621
|
-
|
|
622
|
-
<summary><strong>VS Code</strong> (GitHub Copilot agent mode) โ <code>.vscode/mcp.json</code></summary>
|
|
623
|
-
|
|
624
|
-
```json
|
|
625
|
-
{
|
|
626
|
-
"servers": {
|
|
627
|
-
"scenescout": { "type": "stdio", "command": "npx", "args": ["-y", "scenescout", "serve"] }
|
|
628
|
-
}
|
|
629
|
-
}
|
|
630
|
-
```
|
|
631
|
-
|
|
632
|
-
</details>
|
|
633
|
-
|
|
634
|
-
<details>
|
|
635
|
-
<summary><strong>Codex CLI</strong> โ <code>~/.codex/config.toml</code></summary>
|
|
636
|
-
|
|
637
|
-
```toml
|
|
638
|
-
[mcp_servers.scenescout]
|
|
639
|
-
command = "npx"
|
|
640
|
-
args = ["-y", "scenescout", "serve"]
|
|
641
|
-
```
|
|
642
|
-
|
|
643
|
-
</details>
|
|
644
|
-
|
|
645
|
-
<details>
|
|
646
|
-
<summary><strong>Gemini CLI</strong> โ <code>~/.gemini/settings.json</code> (or <code>.gemini/settings.json</code> in a project)</summary>
|
|
647
|
-
|
|
648
|
-
```json
|
|
649
|
-
{
|
|
650
|
-
"mcpServers": {
|
|
651
|
-
"scenescout": { "command": "npx", "args": ["-y", "scenescout", "serve"] }
|
|
652
|
-
}
|
|
653
|
-
}
|
|
654
|
-
```
|
|
655
|
-
|
|
656
|
-
</details>
|
|
657
|
-
|
|
658
|
-
<details>
|
|
659
|
-
<summary><strong>Windsurf</strong> โ <code>~/.codeium/windsurf/mcp_config.json</code></summary>
|
|
660
|
-
|
|
661
|
-
```json
|
|
662
|
-
{
|
|
663
|
-
"mcpServers": {
|
|
664
|
-
"scenescout": { "command": "npx", "args": ["-y", "scenescout", "serve"] }
|
|
665
|
-
}
|
|
666
|
-
}
|
|
667
|
-
```
|
|
668
|
-
|
|
669
|
-
</details>
|
|
670
|
-
|
|
671
|
-
<details>
|
|
672
|
-
<summary><strong>Cline</strong> โ MCP Servers โ Configure โ Configure MCP Servers (or <code>~/.cline/mcp.json</code> for the CLI)</summary>
|
|
673
|
-
|
|
674
|
-
```json
|
|
675
|
-
{
|
|
676
|
-
"mcpServers": {
|
|
677
|
-
"scenescout": { "command": "npx", "args": ["-y", "scenescout", "serve"], "disabled": false, "autoApprove": [] }
|
|
678
|
-
}
|
|
679
|
-
}
|
|
680
|
-
```
|
|
681
|
-
|
|
682
|
-
</details>
|
|
683
|
-
|
|
684
|
-
<details>
|
|
685
|
-
<summary><strong>Zed</strong> โ <code>settings.json</code> (command palette: <code>zed: open settings file</code>)</summary>
|
|
686
|
-
|
|
687
|
-
```json
|
|
688
|
-
{
|
|
689
|
-
"context_servers": {
|
|
690
|
-
"scenescout": { "command": "npx", "args": ["-y", "scenescout", "serve"], "env": {} }
|
|
691
|
-
}
|
|
692
|
-
}
|
|
693
|
-
```
|
|
694
|
-
|
|
695
|
-
</details>
|
|
696
|
-
|
|
697
|
-
<details>
|
|
698
|
-
<summary><strong>Anything else</strong></summary>
|
|
699
|
-
|
|
700
|
-
Most clients accept the same `mcpServers` JSON shape shown for Cursor.
|
|
701
|
-
|
|
702
|
-
</details>
|
|
703
|
-
|
|
704
|
-
**The method travels with the server.** The tools are only hands and eyes; [`skills/scenescout/SKILL.md`](skills/scenescout/SKILL.md) is the method: what to look at first, when to stop, what counts as a finding. Claude Code loads it as a skill. Every other client gets the same text from the server, with nothing to copy:
|
|
705
|
-
|
|
706
|
-
- the server's instructions tell the agent to call `scout_playbook` before its first attach, and that tool returns the method,
|
|
707
|
-
- clients that list server prompts as commands also get an `explore` prompt, which loads the method and takes an optional URL, level and focus.
|
|
708
|
-
|
|
709
|
-
So in any client, a first message like *"Use SceneScout to test http://localhost:3000"* is enough. If an agent starts clicking without having called `scout_playbook`, tell it to call that first; how closely a model follows server instructions varies by client.
|
|
710
|
-
|
|
711
|
-
The CLI is also useful on its own:
|
|
712
|
-
|
|
713
|
-
```bash
|
|
714
|
-
npx -y scenescout scan <path> # project discovery: framework, routes, saved logins
|
|
715
|
-
npx -y scenescout status <path> # what every session of a running engine is doing right now
|
|
716
|
-
npx -y scenescout watch <path> # the same, live in your browser, with each session's page
|
|
717
|
-
npx -y scenescout login <url> --role admin # sign in once in a visible browser; sessions attach with role: "admin"
|
|
718
|
-
```
|
|
719
|
-
|
|
720
|
-
---
|
|
721
|
-
|
|
722
|
-
## ๐ Project layout
|
|
723
|
-
|
|
724
|
-
```
|
|
725
|
-
src/
|
|
726
|
-
mcp-server.ts the 29 tools + per-session dispatch
|
|
727
|
-
scan.ts project discovery (framework, routes, auth)
|
|
728
|
-
cli.ts scan ยท serve ยท install ยท doctor ยท check ยท ci ยท login ยท export ยท status ยท watch
|
|
729
|
-
check-run.ts drives a check: attach, crawl every route, collect what was measured
|
|
730
|
-
ci-run.ts drives a CI run: the MCP server as a child, the model's API, the agent loop
|
|
731
|
-
login-run.ts drives `scenescout login` and scout_login: a visible browser that saves the role's profile once the sign-in is seen to finish (or on Enter); or --script, headless from the environment
|
|
732
|
-
export-run.ts drives `scenescout export`: reads the findings, asks GitHub or Jira what is filed, files the rest
|
|
733
|
-
installer.ts setup logic (skill link, MCP registration, diagnostics)
|
|
734
|
-
engine/
|
|
735
|
-
browser.ts the engine class: attach, snapshot, actions, crawl, plans
|
|
736
|
-
probes.ts in-page scroll + overlay + focus probes (needs a browser too)
|
|
737
|
-
fingerprint.ts route + element-set identity (state hashing)
|
|
738
|
-
oracles.ts console/page/network/HTTP error detection
|
|
739
|
-
injection.ts the DOM-injection oracle's rules (what to watch for, how to find it)
|
|
740
|
-
claims.ts when the page contradicts the server (refused_empty, false_success)
|
|
741
|
-
request.ts what a replayed API call may be and where it may go
|
|
742
|
-
brief.ts splitting the app between parallel lanes
|
|
743
|
-
lane.ts the typed report a lane hands back
|
|
744
|
-
calibration.ts whether a lane's confidence held up; what it judged and never filed
|
|
745
|
-
pace.ts how a run spent its time
|
|
746
|
-
bench.ts scoring a run against the demo app's answer key
|
|
747
|
-
policy.ts the write-policy safety net
|
|
748
|
-
ownership.ts safe-write: which records were created in this process?
|
|
749
|
-
uploads.ts disk uploads, fenced to the project by real path
|
|
750
|
-
journey.ts task-ease measurement from the action log
|
|
751
|
-
design.ts the design audit + page scoring
|
|
752
|
-
memory.ts cross-run storage + finding dedup
|
|
753
|
-
profiles.ts saved sign-ins: role names, where a profile lives, owner-only files, attach by role, sessionStorage restore
|
|
754
|
-
refresh.ts the refresh broker: which values are a role's refresh tokens, the lock beside the profile, swapping a spent token
|
|
755
|
-
scripted-login.ts a CI sign-in: env and flags, TOTP (RFC 6238) or a fixed code, which field is which, redaction
|
|
756
|
-
signed-in.ts when a person's sign-in in the window has finished: back on the app, a new session, past any SSO round trip
|
|
757
|
-
expiry.ts how long a saved sign-in lasts: cookie dates and JWT exp, checked before lanes start
|
|
758
|
-
report.ts the gap ledger + report generation
|
|
759
|
-
check.ts the check's rules, gate, report and SARIF
|
|
760
|
-
baseline.ts visual baselines: targets.json, where each picture is kept, when one is met
|
|
761
|
-
sarif.ts which repository file a SARIF result points at, so code scanning keeps it
|
|
762
|
-
ci.ts a CI run's options, provider choice, caps, key redaction, tools and files
|
|
763
|
-
export.ts which findings an export files, the inert issue it writes, the marker that dedups it
|
|
764
|
-
provider.ts the Anthropic and OpenAI message shapes, and retries
|
|
765
|
-
replay.ts the run as one page: steps, tasks, frames under each finding
|
|
766
|
-
โฆ collector ยท dispatch ยท fixtures ยท authloss ยท reaper
|
|
767
|
-
scripts/ the test suites (smoke/ holds the real-browser ones)
|
|
768
|
-
test-app/ fixtures for the real-browser smoke tests
|
|
769
|
-
skills/scenescout/ the testing method (SKILL.md): a skill in Claude Code, served by the server everywhere else
|
|
770
|
-
docs/how-it-works.md what happens at each stage, in diagrams
|
|
771
|
-
docs/benchmark.md measuring whether a change made runs better
|
|
772
|
-
docs/validation.md scorecards from runs against public open-source apps
|
|
773
|
-
docs/adr/ why it's built this way
|
|
774
|
-
```
|
|
775
|
-
|
|
776
|
-
> Design principle: logic that *doesn't* need Playwright lives outside `browser.ts`, so it can be unit-tested without launching a browser. That's why `fingerprint`, `policy`, `memory`, `report`, etc. are their own modules.
|
|
777
|
-
|
|
778
|
-
---
|
|
779
|
-
|
|
780
|
-
## ๐ง Design decisions
|
|
781
|
-
|
|
782
|
-
**[How it works, stage by stage](docs/how-it-works.md)** โ diagrams of the run lifecycle, what happens inside one action, the write policy on the wire, how a violation becomes a finding, how a parallel run is split and folded, how roles hand work to each other, where a run's time goes, and how a lane's confidence is checked afterwards.
|
|
783
|
-
|
|
784
|
-
**[Measuring whether a change helped](docs/benchmark.md)** โ the demo app's answer key, the scorecard (recall, precision, judged-not-filed, severity, calibration), and the results log of every run, including what did not help.
|
|
785
|
-
|
|
786
|
-
**[Validation on public open-source apps](docs/validation.md)** โ runs against three well-known open-source web apps, with a scorecard for each: issues by severity, how many were real and how many were false positives, and the engine problems the runs exposed.
|
|
787
|
-
|
|
788
|
-
The load-bearing choices are recorded as ADRs โ read the relevant one before changing a rule it covers:
|
|
789
|
-
|
|
790
|
-
- [1 ยท Completion is an enforced contract, not a claim](docs/adr/0001-completion-is-a-contract-not-a-vibe.md)
|
|
791
|
-
- [2 ยท The write policy is enforced on the wire, not in the prompt](docs/adr/0002-enforce-the-write-policy-at-the-network-layer.md)
|
|
792
|
-
- [3 ยท A gap-ledger entry must be actionable, and suppression must be visible](docs/adr/0003-a-noisy-ledger-is-a-broken-ledger.md)
|
|
793
|
-
- [4 ยท Findings dedup on machine signals, and a merge must never lose a finding](docs/adr/0004-dedup-on-machine-signals-not-prose.md)
|
|
794
|
-
- [5 ยท Testable logic lives outside `browser.ts`](docs/adr/0005-keep-testable-logic-out-of-the-browser-module.md)
|
|
795
|
-
- [6 ยท Nothing in this repo names or is tuned for a tested app](docs/adr/0006-stay-project-agnostic.md)
|
|
796
|
-
- [7 ยท The live view is local, read-only, and leaves nothing behind](docs/adr/0007-the-live-view-is-local-read-only-and-leaves-nothing-behind.md)
|
|
797
|
-
- [8 ยท Recording is opt-in, and a recorded run is one self-contained page](docs/adr/0008-a-recorded-run-is-evidence-and-must-be-asked-for.md)
|
|
798
|
-
- [9 ยท A refused write is answered, not dropped](docs/adr/0009-a-refused-write-is-answered-not-dropped.md)
|
|
799
|
-
- [10 ยท A lane's confidence is checked, not trusted](docs/adr/0010-a-confidence-is-checked-not-trusted.md)
|
|
800
|
-
- [11 ยท A gate is deterministic, and fails only on what it can prove](docs/adr/0011-a-gate-is-deterministic-and-fails-only-on-what-it-can-prove.md)
|
|
801
|
-
- [12 ยท A check replays saved flows and re-tests open findings, within settings whose defaults do the least harm](docs/adr/0012-a-check-replays-saved-flows-and-reports-re-tests.md)
|
|
802
|
-
- [13 ยท What depends on a project's convention is the project's to decide](docs/adr/0013-a-convention-is-the-projects-to-decide.md)
|
|
803
|
-
- [21 ยท Update the docs in the same pull request, unless the change has no user-facing surface](docs/adr/0021-update-the-docs-in-the-pull-request.md)
|
|
804
|
-
|
|
805
|
-
---
|
|
806
|
-
|
|
807
|
-
## ๐ง Development
|
|
808
|
-
|
|
809
|
-
Working on SceneScout itself is the only reason to clone it:
|
|
810
|
-
|
|
811
|
-
```bash
|
|
812
|
-
git clone https://github.com/brunoboto96/SceneScout.git scenescout && cd scenescout
|
|
813
|
-
npm install # installs dependencies and builds
|
|
814
|
-
npm run setup # same as `scenescout install`, but registers THIS checkout (the skill is linked, so edits are live)
|
|
815
|
-
npm test # build + 24 suites: 22 pure-logic suites (scan, oracle, policy, โฆ bench, hygiene),
|
|
816
|
-
# then smoke and mcp-check (the server over stdio), both with real browsers
|
|
817
|
-
npm run bench -- --all # re-score every archived benchmark run against the current answer key
|
|
818
|
-
npm run demo # regenerate examples/ from the demo app
|
|
819
|
-
```
|
|
820
|
-
|
|
821
|
-
Contributing? Start with [VISION.md](VISION.md) (what is in scope) and [CONTRIBUTING.md](CONTRIBUTING.md) (how changes land), then see [AGENTS.md](AGENTS.md) for the house rules โ chiefly: bug fixes need a regression test at the cheapest layer that can fail, keep the repo project-agnostic (ADR 6), and `npm test` must pass.
|
|
822
|
-
|
|
823
|
-
## ๐ Security
|
|
824
|
-
|
|
825
|
-
Found a way past the write policy, or another security problem? Please report it privately โ see [SECURITY.md](SECURITY.md).
|
|
826
|
-
|
|
827
|
-
## ๐ License
|
|
828
|
-
|
|
829
|
-
[MIT](LICENSE).
|
|
830
|
-
|
|
831
|
-
<details>
|
|
832
|
-
<summary><strong>Full capability list</strong> โ every behavior, for the curious</summary>
|
|
833
|
-
|
|
834
|
-
- **Structured render-state, not pixels.** Element lists with geometry; screenshots reserved for pixel-native residue (canvas, rendering glitches). Images that failed to load are reported from the DOM, including ones whose URL answered 200 with something that is not an image.
|
|
835
|
-
- **Diff snapshots with stable refs.** Re-snapshots return only what changed (10.7 kB โ 0.7 kB on a 130-element page); old refs stay valid.
|
|
836
|
-
- **Geometry oracles.** Overlap and off-screen defects computed from layout boxes.
|
|
837
|
-
- **Oracles after every action.** Console errors, page errors, failed requests, HTTP 4xx/5xx drained into every tool result โ and DOM injection: a markup-shaped value the agent typed that later renders as an element on any page (stored or reflected XSS).
|
|
838
|
-
- **Multi-role, genuinely concurrent.** Commands to *different* sessions run in parallel; safe-write ownership is shared, so role A can create what role B approves. The report renders a role capability matrix.
|
|
839
|
-
- **Task ease, not just correctness.** `scout_journey` measures interaction cost, distinct screens, path, and backtracks.
|
|
840
|
-
- **Design audit with page scores.** Two tiers (โ measurable defects / โ craft suggestions incl. AI-slop tells), per-page 0โ100 score persisted per route, plus an automatic overlay/modal probe on every snapshot. Shared shell scored once, separately.
|
|
841
|
-
- **Scrolls like a user โ and notices when it can't.** Reports `SCROLL LOCKED` for a leaked modal scroll-lock, finds the real inner scroll pane on app-shell layouts, and flags `UNREACHABLE` controls clipped inside `overflow:hidden`.
|
|
842
|
-
- **Uploads like a user.** Answers a styled file-chooser or sets a hidden input directly, with a valid in-memory fixture; `filePath` is fenced to the project under test; files violating `accept` are flagged at selection.
|
|
843
|
-
- **Auth via Playwright storage states.** Expired tokens caught at attach; repeated login-bounces raise `SESSION AUTH LOST` (a session attached by role first re-attaches once from its role's latest saved profile and carries on); a bounced route is recorded as *not* covered โ a dead session can't certify routes it never reached.
|
|
844
|
-
- **A trustworthy gap ledger.** Entries must be actionable (a search box or wizard sub-step isn't "form filled but never submitted"); API/download URLs never enter the route contract.
|
|
845
|
-
- **Honest reporting.** Shared chrome counted once, stale scores marked, role matrix compares only roles that actually attempted a route.
|
|
846
|
-
- **Cross-run written knowledge.** `scout_note` curates `.scenescout/ASSUMPTIONS.md` โ app model, personas, constraints, risks โ in prose.
|
|
847
|
-
- **Daemon-grade robustness.** Per-tool watchdogs, orphaned-browser reaping, bounded teardown, live status via `scenescout status <project>`, and a live view of every session's page: the agent gives you its address when it attaches, or run `scenescout watch <project>` (loopback only, read-only, nothing written to disk: [ADR 7](docs/adr/0007-the-live-view-is-local-read-only-and-leaves-nothing-behind.md)).
|
|
848
|
-
|
|
849
|
-
</details>
|
|
168
|
+
[MIT](LICENSE) licensed.
|