scenescout 3.23.1 → 3.23.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +13 -0
- package/README.md +98 -9
- package/dist/engine/check-replay.js +1 -1
- package/dist/engine/check-report.js +3 -2
- package/dist/engine/live-page.js +8 -5
- package/dist/engine/replay.js +34 -0
- package/dist/engine/report.js +3 -2
- package/dist/mcp-server.js +29 -9
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,18 @@
|
|
|
1
1
|
# scenescout
|
|
2
2
|
|
|
3
|
+
## 3.23.3
|
|
4
|
+
|
|
5
|
+
### Patch Changes
|
|
6
|
+
|
|
7
|
+
- 43012d7: The live board's report panel keeps showing a finding's pictures when sessions in two project folders are attached at once and the session that filed the finding has closed. The report now names the project it describes, and the board loads its pictures from that folder rather than from whichever project is attached first.
|
|
8
|
+
- 1cda780: The live view's run page keeps showing its pictures after a session attaches to another project folder. Its picture addresses now name the project the page was written for, so a page already open in a browser loads them from that folder rather than from whichever project attached last.
|
|
9
|
+
|
|
10
|
+
## 3.23.2
|
|
11
|
+
|
|
12
|
+
### Patch Changes
|
|
13
|
+
|
|
14
|
+
- 63329f7: On a recorded check's `replay.html`, each step's frame now fills its box, so the note shown when a frame is missing no longer shows beside a frame that loaded. In a test report laid out from a template, table cells wrap at word boundaries, so short columns such as Result no longer break a word across lines; hashes, paths and frame links still wrap anywhere.
|
|
15
|
+
|
|
3
16
|
## 3.23.1
|
|
4
17
|
|
|
5
18
|
### Patch Changes
|
package/README.md
CHANGED
|
@@ -11,11 +11,11 @@ Works with Claude Code · Cursor · VS Code (Copilot) · Codex CLI · Gemini CLI
|
|
|
11
11
|
[](LICENSE)
|
|
12
12
|

|
|
13
13
|
|
|
14
|
-
[📖 Guide](docs/guide/Home.md) · [🐛 What it catches](#-what-it-catches) · [🚀 Get started](#-get-started) · [🚦 CI](#-in-ci) · [🔒 Safety](#-safe-by-default) · [📚 Docs](#-documentation)
|
|
14
|
+
[📖 Guide](docs/guide/Home.md) · [🐛 What it catches](#-what-it-catches) · [🎬 QA with evidence](#-automated-qa-with-evidence) · [🚀 Get started](#-get-started) · [🚦 CI](#-in-ci) · [🔒 Safety](#-safe-by-default) · [📚 Docs](#-documentation)
|
|
15
15
|
|
|
16
16
|
</div>
|
|
17
17
|
|
|
18
|
-
Scripted end-to-end tests answer one question: *does this exact flow still work?* They say nothing about the rest of the app. SceneScout lets the agent you already use explore a running web app like a curious, thorough tester. It clicks, fills forms, switches roles and calls the API behind a hidden button, then writes a report of what is **broken** and what could be **better**, with evidence for every line.
|
|
18
|
+
Scripted end-to-end tests answer one question: *does this exact flow still work?* They say nothing about the rest of the app. SceneScout lets the agent you already use explore a running web app like a curious, thorough tester. It clicks, fills forms, switches roles and calls the API behind a hidden button, then writes a report of what is **broken** and what could be **better**, with a picture and evidence for every line.
|
|
19
19
|
|
|
20
20
|
Try it on any app you are allowed to test. No account, no API key, no setup:
|
|
21
21
|
|
|
@@ -23,9 +23,26 @@ Try it on any app you are allowed to test. No account, no API key, no setup:
|
|
|
23
23
|
npx -y scenescout http://localhost:3000
|
|
24
24
|
```
|
|
25
25
|
|
|
26
|
+
```mermaid
|
|
27
|
+
flowchart LR
|
|
28
|
+
A["🧠 Your coding agent<br/>Claude Code, Cursor, Copilot…"] -- MCP --> S["🔭 SceneScout<br/>browser · checks · memory"]
|
|
29
|
+
S --> App["🌐 Your running app"]
|
|
30
|
+
S --> R["📋 Report<br/>plain words + pictures"]
|
|
31
|
+
S --> E["🎬 Evidence<br/>replay, videos, test report"]
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
## 👥 Who it's for
|
|
35
|
+
|
|
36
|
+
| 🛠️ **Developers** | 🧑💼 **QA, product and non-technical teams** |
|
|
37
|
+
|---|---|
|
|
38
|
+
| Findings with the request that failed (`GET /api/orders?status=archived → 500`), the steps, and a Playwright regression-test skeleton | Each problem in plain words first: what was done, what was expected, what happened, and a picture of the page |
|
|
39
|
+
| Next to the source, the file behind the bug and a likely fix | Visible problems shown, not described: a covered button, a broken image, text too faint to read |
|
|
40
|
+
| A deterministic gate for pull requests, with SARIF for code scanning | Journeys recorded step by step, as frames and video, so a pass is something you can watch |
|
|
41
|
+
| Runs from your editor's agent, or unattended in CI | A test report laid out by **your** template, with expected and actual results, deviations and blank sign-off rows |
|
|
42
|
+
|
|
26
43
|
## 🐛 What it catches
|
|
27
44
|
|
|
28
|
-
A real run against the small demo app in this repository, which has bugs planted on purpose:
|
|
45
|
+
A real run against the small demo app in this repository, which has bugs planted on purpose. The two you can see are outlined:
|
|
29
46
|
|
|
30
47
|
<p align="center"><img src="examples/screenshots/dashboard-annotated.png" alt="The demo app's dashboard with two defects outlined in red: 1, a yellow badge covering the All orders button; 2, the weekly chart image failing to load" width="760" /></p>
|
|
31
48
|
|
|
@@ -45,20 +62,77 @@ It filed twelve findings. A few from [the report](examples/report.md):
|
|
|
45
62
|
|
|
46
63
|
What it looks for, on every page, after every action:
|
|
47
64
|
|
|
65
|
+
- 📐 **Broken layout, from geometry:** controls that overlap, sit off-screen, hide under a sticky bar or can never be scrolled into view. The CSS bugs a person spots at a glance, found without one.
|
|
48
66
|
- 🧨 **Real breakage:** console errors, crashes, failed requests, 4xx and 5xx responses, broken images, dead-end pages.
|
|
49
67
|
- 🔓 **Permission leaks:** it calls the app's own API as each role, so "the button is hidden" becomes "the server refuses it", or doesn't.
|
|
50
68
|
- 🤥 **Pages that lie:** "Saved!" after the server refused the save, or an empty table after the request failed.
|
|
51
69
|
- 👆 **Impatient users:** a double-click that sends the same order twice.
|
|
52
|
-
- 📐 **Broken layout, from geometry:** controls that overlap, sit off-screen, hide under a sticky bar or can never be scrolled into view.
|
|
53
70
|
- ♿ **Accessibility and craft:** contrast, focus, labels, target sizes, spacing and type, with a 0 to 100 score per page.
|
|
54
71
|
- 🧭 **Friction:** how many steps a task takes, and where a user had to go back.
|
|
55
72
|
- 💉 **Security smells:** typed markup that comes back as an element, and tokens posted to any window.
|
|
56
73
|
|
|
57
|
-
|
|
74
|
+
[Everything it checks](docs/guide/What-it-checks.md).
|
|
75
|
+
|
|
76
|
+
## 🎬 Automated QA with evidence
|
|
77
|
+
|
|
78
|
+
Save the journeys that matter, the happy paths and the ones that must fail politely, and `scenescout check` replays them on every pull request with no model involved. Recorded, each run leaves proof you can watch and hand to someone who never opens a terminal.
|
|
79
|
+
|
|
80
|
+
```mermaid
|
|
81
|
+
flowchart LR
|
|
82
|
+
PR["🔀 Pull request"] --> C["🚦 scenescout check<br/>--record --video --template"]
|
|
83
|
+
F["📝 Saved journeys<br/>.scenescout/flows/*.json"] --> C
|
|
84
|
+
C --> G["✅ / ❌ A gate on the pull request"]
|
|
85
|
+
C --> RP["🖼️ replay.html<br/>every step, with frames"]
|
|
86
|
+
C --> V["🎞️ A video per journey"]
|
|
87
|
+
C --> TR["📄 Test report<br/>your template, SHA-256 manifest,<br/>blank sign-off rows"]
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
A journey is a few lines of JSON, written by hand or kept from a flow your agent just walked. This is [the demo's happy path](examples/flows/01-place-an-order.json):
|
|
91
|
+
|
|
92
|
+
```json
|
|
93
|
+
{
|
|
94
|
+
"name": "place an order",
|
|
95
|
+
"id": "TC-01",
|
|
96
|
+
"requirements": ["REQ-ORD-1", "REQ-ORD-2"],
|
|
97
|
+
"steps": [
|
|
98
|
+
{ "action": "navigate", "target": "/orders-new.html", "expected": "The new-order form is shown" },
|
|
99
|
+
{ "action": "type", "target": "testid=new-order-customer", "value": "Harbour Bakery", "expected": "The customer is filled in" },
|
|
100
|
+
{ "action": "type", "target": "testid=new-order-items", "value": "4", "replace": true, "expected": "The item count is 4" },
|
|
101
|
+
{ "action": "click", "target": "testid=new-order-submit", "expected": "The order is created" },
|
|
102
|
+
{ "action": "expect-request", "request": "POST /api/orders", "status": "2xx" },
|
|
103
|
+
{ "action": "expect-element", "target": "testid=new-order-created-link", "state": "visible", "expected": "A link to the new order is shown" }
|
|
104
|
+
]
|
|
105
|
+
}
|
|
106
|
+
```
|
|
107
|
+
|
|
108
|
+
Against the [demo app](demo-app/), with the three journeys in [examples/flows](examples/flows/) (one creates an order, so the check is allowed to send it):
|
|
109
|
+
|
|
110
|
+
```bash
|
|
111
|
+
npx scenescout check http://127.0.0.1:4173 --flows examples/flows --mode read-only --flow-writes allow \
|
|
112
|
+
--record --video --template examples/report-template.json
|
|
113
|
+
```
|
|
114
|
+
|
|
115
|
+
**The journey, as it ran** (one of the videos it filmed):
|
|
116
|
+
|
|
117
|
+
<p align="center"><img src="examples/screenshots/journey-place-order.gif" alt="A recorded journey on the demo app: the new-order form is filled with a customer and four items, Create order is clicked, and the confirmation with a link to the new order appears" width="760" /></p>
|
|
118
|
+
|
|
119
|
+
**The replay page:** each journey with a pass or fail badge. A journey that broke opens at the step that broke, with the page as it was after every step:
|
|
120
|
+
|
|
121
|
+
<p align="center"><img src="examples/screenshots/check-replay.png" alt="The check's replay page: verdict FAILED, the app's address, start and end times, 24 frames and 3 videos; two journeys passed and one, see archived orders, failed at step 3" width="820" /></p>
|
|
122
|
+
|
|
123
|
+
<p align="center"><img src="examples/screenshots/check-replay-failed.png" alt="The failed journey opened: step 1 navigates to the orders page and passes, step 2 selects Archived and passes with an empty table, step 3 expected GET /api/orders to answer 200 and it answered 500, highlighted in red" width="760" /></p>
|
|
124
|
+
|
|
125
|
+
**The test report**, laid out by [a template](examples/report-template.json) you write once: test IDs, the requirements each covers, expected and actual results, a screenshot per step, every deviation listed again for the reviewer, and the SHA-256 of each piece of evidence. SceneScout signs nothing; the sign-off rows are for your people.
|
|
126
|
+
|
|
127
|
+
<p align="center"><img src="examples/screenshots/test-report-results.png" alt="The test report's results table: test TC-01, place an order, covering REQ-ORD-1 and REQ-ORD-2; each step with its expected result, actual result As expected, result Pass, and a screenshot of the page after the step" width="820" /></p>
|
|
128
|
+
|
|
129
|
+
<p align="center"><img src="examples/screenshots/test-report-deviations.png" alt="The test report's deviations: TC-03, see archived orders, step 3 expected GET /api/orders to answer 200, actual GET /api/orders answered 500, result Fail" width="820" /></p>
|
|
130
|
+
|
|
131
|
+
Everything is one self-contained HTML page per report, with no scripts, nothing loaded from the network, and a layout that prints. [Recording a check](docs/ci.md#recording-a-check) and [a test report from a template](docs/ci.md#a-test-report-from-a-template) have the details, including what to keep out of the pictures.
|
|
58
132
|
|
|
59
133
|
## 📺 Watch it work
|
|
60
134
|
|
|
61
|
-
Each run opens a live view on your machine, with one card per agent: what it is doing, the page it is on, and a feed of every action. Three agents are testing the demo app in parallel here:
|
|
135
|
+
Each exploratory run opens a live view on your machine, with one card per agent: what it is doing, the page it is on, and a feed of every action. Three agents are testing the demo app in parallel here:
|
|
62
136
|
|
|
63
137
|
<p align="center"><img src="examples/screenshots/live-view.png" alt="The live view during a run of three parallel agents against the demo app: one card per session, each with its role and objective, the task it is on, the tool it is running, the page it is on, a live thumbnail, and a feed of the actions it just took, tinted one colour per task" width="880" /></p>
|
|
64
138
|
|
|
@@ -71,6 +145,17 @@ You can read the report while the agents are still working, and scrub back throu
|
|
|
71
145
|
- 🛡️ **Safety is enforced on the network, not requested in a prompt.** Nothing existing is changed unless you allow it, and a blocked write never reaches your server.
|
|
72
146
|
- ✅ **"Done" is a contract.** The report lists everything not tested, and at the `extensive` level refuses to finish while any known page is unvisited.
|
|
73
147
|
- 🧠 **It remembers.** Each run starts from what the last one learned, and re-tests the bugs earlier runs left open.
|
|
148
|
+
- 📊 **It is measured, not asserted.** Every change to how it explores is scored against an app with planted bugs and a held-out app it is never tuned on ([the log](docs/benchmark.md)):
|
|
149
|
+
|
|
150
|
+
```mermaid
|
|
151
|
+
xychart-beta
|
|
152
|
+
title "Planted defects found per run, demo app (13 planted)"
|
|
153
|
+
x-axis "Run" ["0", "1", "2", "3", "4", "5", "6", "7", "8", "9", "10", "11", "12"]
|
|
154
|
+
y-axis "Found" 0 --> 13
|
|
155
|
+
bar [11, 12, 9, 10, 11, 11, 10, 10, 11, 13, 11, 11, 12]
|
|
156
|
+
```
|
|
157
|
+
|
|
158
|
+
<sub>Each bar is one run of eight parallel agents, re-scored against today's answer key. One run is noisy: read the trend across runs, not a single bar.</sub>
|
|
74
159
|
|
|
75
160
|
## 🚀 Get started
|
|
76
161
|
|
|
@@ -115,17 +200,20 @@ npx -y scenescout login http://localhost:3000 --role admin # then: /scenescou
|
|
|
115
200
|
|
|
116
201
|
| | What it does | Needs a model? |
|
|
117
202
|
|---|---|---|
|
|
118
|
-
| [`scenescout check`](docs/guide/Ways-to-use-it.md#scenescout-check-a-gate-in-ci) | A deterministic gate: measures every page, replays your saved
|
|
203
|
+
| [`scenescout check`](docs/guide/Ways-to-use-it.md#scenescout-check-a-gate-in-ci) | A deterministic gate: measures every page, replays your saved journeys and visual baselines, fails only on what it can prove; recorded, it leaves the [evidence above](#-automated-qa-with-evidence) | No |
|
|
119
204
|
| [`scenescout ci`](docs/guide/Ways-to-use-it.md#scenescout-ci-an-unattended-exploratory-run) | An unattended exploratory run, driven by the Anthropic or OpenAI API. It reports and never fails the build | An API key |
|
|
120
205
|
| [`/scenescout qa`](docs/guide/Ways-to-use-it.md#scenescout-qa-on-a-pull-request) | A comment on a pull request that tests its preview deploy and replies with the results; `/scenescout qa check` runs your own check instead | An API key (`qa check`: none) |
|
|
121
206
|
| [`scenescout export`](docs/guide/Ways-to-use-it.md#filing-findings-as-issues) | Files the findings as GitHub or Jira issues, each once | No |
|
|
122
207
|
|
|
123
|
-
A gate on every pull request, as a GitHub Action:
|
|
208
|
+
A gate on every pull request that keeps its evidence, as a GitHub Action:
|
|
124
209
|
|
|
125
210
|
```yaml
|
|
126
211
|
- uses: brunoboto96/SceneScout@v3
|
|
127
212
|
with:
|
|
128
213
|
url: http://127.0.0.1:3000
|
|
214
|
+
record: on
|
|
215
|
+
video: on
|
|
216
|
+
template: tests/report-template.json
|
|
129
217
|
```
|
|
130
218
|
|
|
131
219
|
[docs/ci.md](docs/ci.md) has complete workflows, every option, and the same check on GitLab CI, CircleCI or any shell.
|
|
@@ -139,7 +227,7 @@ A gate on every pull request, as a GitHub Action:
|
|
|
139
227
|
| 🟡 `safe-write` | Creates records, and edits or deletes only the ones it created |
|
|
140
228
|
| 🔴 `destructive` | Everything. Only when you say the data is disposable; the agent never picks it |
|
|
141
229
|
|
|
142
|
-
The policy sits on the network, so a blocked request never reaches your server. The page gets a refusal instead, which is how SceneScout catches a page that claims success anyway. Findings, memory and reports stay in a `.scenescout/` folder that keeps itself out of git.
|
|
230
|
+
The policy sits on the network, so a blocked request never reaches your server. The page gets a refusal instead, which is how SceneScout catches a page that claims success anyway. A saved journey that creates something, like the order above, runs only when the check is told `--flow-writes allow`, against a test environment. Findings, memory and reports stay in a `.scenescout/` folder that keeps itself out of git.
|
|
143
231
|
|
|
144
232
|
> [!IMPORTANT]
|
|
145
233
|
> Only test sites you own or are allowed to test. [Safety model](docs/guide/Safety-model.md).
|
|
@@ -155,6 +243,7 @@ The policy sits on the network, so a blocked request never reaches your server.
|
|
|
155
243
|
| [Recipes](docs/guide/Recipes.md) | Setups for seven kinds of project |
|
|
156
244
|
| [Configuration reference](docs/guide/Configuration-reference.md) | Every option, environment variable and action input |
|
|
157
245
|
| [Troubleshooting](docs/guide/Troubleshooting.md) | Symptoms and fixes, upgrading and uninstalling |
|
|
246
|
+
| [Running it in CI](docs/ci.md) | Saved journeys, recording, test reports from templates, every workflow |
|
|
158
247
|
| [How it works](docs/how-it-works.md) | Diagrams of a run, an action, the write policy, lanes |
|
|
159
248
|
| [Benchmark](docs/benchmark.md) · [Validation](docs/validation.md) | How runs are scored against answer keys, and runs on public apps |
|
|
160
249
|
| [Design decisions](docs/adr/README.md) | Why the rules are what they are |
|
|
@@ -285,7 +285,7 @@ ol.steps { list-style:none; margin:10px 0 0; padding:0; }
|
|
|
285
285
|
.why { margin:4px 0; font-size:13px; overflow-wrap:anywhere; }
|
|
286
286
|
a.frame { display:grid; margin:6px 0 2px; max-width:min(100%,720px); color:var(--muted); font-size:12px; }
|
|
287
287
|
a.frame > * { grid-area:1 / 1; }
|
|
288
|
-
a.frame img { position:relative;
|
|
288
|
+
a.frame img { position:relative; width:100%; max-height:400px; object-fit:cover; object-position:top; border:1px solid var(--line); border-radius:6px; display:block; background:var(--panel); }
|
|
289
289
|
.gone-note { align-self:end; padding:2.6em 12px 12px; border:1px dashed var(--line); border-radius:6px; }
|
|
290
290
|
.noframe, .none { color:var(--muted); font-size:12px; font-style:italic; margin:4px 0; }
|
|
291
291
|
.note { color:var(--muted); font-size:13px; }
|
|
@@ -431,7 +431,7 @@ header.title h1 { font-size:22px; margin:0 0 4px; }
|
|
|
431
431
|
header.title .subtitle { margin:0 0 8px; color:var(--muted); }
|
|
432
432
|
h2 { font-size:17px; margin:28px 0 8px; border-bottom:1px solid var(--line); padding-bottom:4px; }
|
|
433
433
|
table { border-collapse:collapse; width:100%; margin:8px 0; }
|
|
434
|
-
th, td { border:1px solid var(--line); padding:5px 8px; text-align:left; vertical-align:top; overflow-wrap:
|
|
434
|
+
th, td { border:1px solid var(--line); padding:5px 8px; text-align:left; vertical-align:top; overflow-wrap:break-word; }
|
|
435
435
|
thead th, table.facts th { background:var(--head); }
|
|
436
436
|
table.facts { width:auto; min-width:50%; }
|
|
437
437
|
tr.test-head th { background:var(--head); font-weight:400; }
|
|
@@ -444,7 +444,8 @@ td.r-passed { color:var(--pass); font-weight:700; } td.r-failed, td.r-refused {
|
|
|
444
444
|
tr.failed td, tr.refused td { background:var(--fail-bg); }
|
|
445
445
|
a.frame { display:block; color:var(--muted); font-size:11px; }
|
|
446
446
|
a.frame img { display:block; max-width:220px; max-height:150px; object-fit:cover; object-position:top; border:1px solid var(--line); }
|
|
447
|
-
td.hash, td.path { font:12px/1.5 ui-monospace,SFMono-Regular,Menlo,Consolas,monospace; }
|
|
447
|
+
td.hash, td.path { font:12px/1.5 ui-monospace,SFMono-Regular,Menlo,Consolas,monospace; overflow-wrap:anywhere; }
|
|
448
|
+
a.frame { overflow-wrap:anywhere; }
|
|
448
449
|
table.sign td.blank { height:2.6em; min-width:8em; }
|
|
449
450
|
.none { color:var(--muted); font-style:italic; }
|
|
450
451
|
@media print { main { max-width:none; padding:0; } a { color:inherit; text-decoration:none; } tbody, tr { break-inside:avoid; } }
|
package/dist/engine/live-page.js
CHANGED
|
@@ -297,6 +297,8 @@ export const LIVE_PAGE = `<!doctype html>
|
|
|
297
297
|
var reportKey = null;
|
|
298
298
|
// Per-finding frames from api/report: empty unless the run was recorded.
|
|
299
299
|
var reportEvidence = [];
|
|
300
|
+
// Where the report's pictures are served from: the project it was rendered from (api/report names it).
|
|
301
|
+
var reportFramePrefix = 'record/';
|
|
300
302
|
var savedACopy = false;
|
|
301
303
|
// A result that reads as a failure is shown in red.
|
|
302
304
|
var BAD_RESULT = /error|fail|refus|block|violation|abandoned/i;
|
|
@@ -509,13 +511,13 @@ export const LIVE_PAGE = `<!doctype html>
|
|
|
509
511
|
picture = el('figure', 'picture');
|
|
510
512
|
picture.setAttribute('data-testid', 'live-report-picture-' + id);
|
|
511
513
|
var link = el('a');
|
|
512
|
-
link.href =
|
|
514
|
+
link.href = reportFramePrefix + found.picture.file;
|
|
513
515
|
link.target = '_blank';
|
|
514
516
|
link.rel = 'noreferrer';
|
|
515
517
|
link.setAttribute('data-testid', 'live-report-picture-open');
|
|
516
518
|
var pic = el('img');
|
|
517
519
|
pic.loading = 'lazy';
|
|
518
|
-
pic.src =
|
|
520
|
+
pic.src = reportFramePrefix + found.picture.file;
|
|
519
521
|
pic.alt = 'What the finding is about, when it was filed';
|
|
520
522
|
pic.addEventListener('error', function () { picture.classList.add('gone'); });
|
|
521
523
|
link.appendChild(pic);
|
|
@@ -533,7 +535,7 @@ export const LIVE_PAGE = `<!doctype html>
|
|
|
533
535
|
var fig = el('figure');
|
|
534
536
|
var img = el('img');
|
|
535
537
|
img.loading = 'lazy';
|
|
536
|
-
img.src =
|
|
538
|
+
img.src = reportFramePrefix + f.frame;
|
|
537
539
|
img.alt = f.action + ' ' + f.detail;
|
|
538
540
|
// A recording deleted since the run leaves a broken icon under a finding
|
|
539
541
|
// that still counts it as evidence. Say which it is.
|
|
@@ -607,7 +609,7 @@ export const LIVE_PAGE = `<!doctype html>
|
|
|
607
609
|
var fig = el('figure', 'picture');
|
|
608
610
|
var pic = el('img');
|
|
609
611
|
pic.loading = 'lazy';
|
|
610
|
-
pic.src =
|
|
612
|
+
pic.src = reportFramePrefix + m[2];
|
|
611
613
|
pic.alt = m[1];
|
|
612
614
|
fig.appendChild(pic);
|
|
613
615
|
container.appendChild(fig);
|
|
@@ -680,12 +682,13 @@ export const LIVE_PAGE = `<!doctype html>
|
|
|
680
682
|
}
|
|
681
683
|
reportMarkdown = d.markdown;
|
|
682
684
|
reportEvidence = d.evidence || [];
|
|
685
|
+
reportFramePrefix = d.framePrefix || 'record/';
|
|
683
686
|
meta.textContent = (finished ? 'as the run left it at ' : 'as the run stands at ') + clock(d.at) + ' · ' + whereItIs();
|
|
684
687
|
// The report is re-read every few seconds while the run goes on.
|
|
685
688
|
// Re-rendering an unchanged document threw away what the reader was
|
|
686
689
|
// doing with it: an opened accordion shut itself, and the page jumped
|
|
687
690
|
// back to the top, every five seconds.
|
|
688
|
-
var key = d.markdown + '\\u0000' + JSON.stringify(reportEvidence);
|
|
691
|
+
var key = d.markdown + '\\u0000' + JSON.stringify(reportEvidence) + '\\u0000' + reportFramePrefix;
|
|
689
692
|
if (key === reportKey) return;
|
|
690
693
|
var open = {};
|
|
691
694
|
var was = doc.querySelectorAll('details[open][data-testid]');
|
package/dist/engine/replay.js
CHANGED
|
@@ -18,6 +18,7 @@
|
|
|
18
18
|
* Everything here is pure string work — no browser, no filesystem — so the
|
|
19
19
|
* escaping and the grouping are table-tested.
|
|
20
20
|
*/
|
|
21
|
+
import crypto from "node:crypto";
|
|
21
22
|
import path from "node:path";
|
|
22
23
|
import { isSafeRelativePath } from "./plain.js";
|
|
23
24
|
/** Most frames one recorded session keeps. A long run is thousands of actions, and a project folder is not a video store. */
|
|
@@ -59,6 +60,39 @@ export function resolveFrame(root, relPath) {
|
|
|
59
60
|
const file = path.resolve(base, inside);
|
|
60
61
|
return file.startsWith(base + path.sep) ? file : null;
|
|
61
62
|
}
|
|
63
|
+
/**
|
|
64
|
+
* A short name for a project folder, used in the live view's frame addresses so
|
|
65
|
+
* a page says which project its pictures belong to. One engine can serve runs
|
|
66
|
+
* in several folders in turn, and a page already open in a browser loads its
|
|
67
|
+
* images later, after another folder may have attached. The name is derived
|
|
68
|
+
* from the folder's resolved path and reveals nothing of it.
|
|
69
|
+
*/
|
|
70
|
+
export function projectKey(dir) {
|
|
71
|
+
return "p-" + crypto.createHash("sha256").update(path.resolve(dir)).digest("hex").slice(0, 12);
|
|
72
|
+
}
|
|
73
|
+
/** A frame address split into the project it names, when it names one, and the path inside that project's recordings. */
|
|
74
|
+
export function splitProjectFrame(relPath) {
|
|
75
|
+
const m = /^(p-[0-9a-f]{12})\/(.+)$/.exec(relPath);
|
|
76
|
+
return m ? { key: m[1], rest: m[2] } : { rest: relPath };
|
|
77
|
+
}
|
|
78
|
+
/**
|
|
79
|
+
* The project folders this process has rendered a run page or a board report
|
|
80
|
+
* for, by key. Only a folder one was rendered from can be named, so a viewer
|
|
81
|
+
* cannot reach any other folder by guessing a key.
|
|
82
|
+
*/
|
|
83
|
+
export class ProjectDirs {
|
|
84
|
+
byKey = new Map();
|
|
85
|
+
/** Remember `dir` and return its key. */
|
|
86
|
+
remember(dir) {
|
|
87
|
+
const key = projectKey(dir);
|
|
88
|
+
this.byKey.set(key, dir);
|
|
89
|
+
return key;
|
|
90
|
+
}
|
|
91
|
+
/** The folder a key names, or undefined when nothing was rendered from one with that key. */
|
|
92
|
+
dir(key) {
|
|
93
|
+
return this.byKey.get(key);
|
|
94
|
+
}
|
|
95
|
+
}
|
|
62
96
|
/** Text from the app under test reaches this document, so nothing is interpolated unescaped. */
|
|
63
97
|
export function escapeHtml(text) {
|
|
64
98
|
return text.replace(/[&<>"']/g, (c) => ({ "&": "&", "<": "<", ">": ">", '"': """, "'": "'" })[c]);
|
package/dist/engine/report.js
CHANGED
|
@@ -4,7 +4,7 @@ import { baseRoute, reachedRoutes, routeIdentity, isEmbedKey, isWorthALook, judg
|
|
|
4
4
|
import { sayVerification } from "./verify.js";
|
|
5
5
|
import { describePicture, readFindingPicture } from "./capture.js";
|
|
6
6
|
import { feedForSession } from "./live.js";
|
|
7
|
-
import { buildReplayHtml, evidenceFor } from "./replay.js";
|
|
7
|
+
import { buildReplayHtml, evidenceFor, projectKey } from "./replay.js";
|
|
8
8
|
import { calibrate, formatCalibration } from "./calibration.js";
|
|
9
9
|
import { formatPace, measurePace } from "./pace.js";
|
|
10
10
|
import { formatNeverSubmittedEmpty } from "./forms.js";
|
|
@@ -221,7 +221,8 @@ export function replayDocument(memory, markdown, version = "") {
|
|
|
221
221
|
project: projectName(memory.dir),
|
|
222
222
|
at: new Date().toISOString(),
|
|
223
223
|
version,
|
|
224
|
-
|
|
224
|
+
// Keyed by project, so the page's pictures still load after a session attaches to another folder (replay.ts).
|
|
225
|
+
framePrefix: `record/${projectKey(memory.dir)}/`,
|
|
225
226
|
savedAt: memory.dir,
|
|
226
227
|
});
|
|
227
228
|
}
|
package/dist/mcp-server.js
CHANGED
|
@@ -70,7 +70,7 @@ import { CRITERION_VERDICTS, findCriterion, formatCriteriaForLanes, formatReadin
|
|
|
70
70
|
import { ACTION_TIMEOUT_ENV, DEFAULT_ACTION_TIMEOUT_MS, DEFAULT_CRAWL_NAV_TIMEOUT_MS, DEFAULT_NAV_TIMEOUT_MS, LIMIT_BOUNDS, NAV_TIMEOUT_ENV, watchdogFor, } from "./engine/limits.js";
|
|
71
71
|
import { MAX_READ_POSTS, READ_POSTS_ENV } from "./engine/policy.js";
|
|
72
72
|
import { chooseProjectFolder, PROJECTS_DIR_ENV, workspaceFromRoots } from "./engine/project-folder.js";
|
|
73
|
-
import { RECORD_MAX_FRAMES, resolveFrame } from "./engine/replay.js";
|
|
73
|
+
import { ProjectDirs, RECORD_MAX_FRAMES, resolveFrame, splitProjectFrame } from "./engine/replay.js";
|
|
74
74
|
import { describePace, normalizePace } from "./engine/settle.js";
|
|
75
75
|
import { needsTask, taskRefusal, TASK_MAX } from "./engine/task.js";
|
|
76
76
|
import { formElicitationSupported, INTAKE_FORM_MESSAGE, INTAKE_TOOL, intakeFormSchema, intakeFromForm, intakeReplyText } from "./intake.js";
|
|
@@ -149,6 +149,8 @@ let liveError = null;
|
|
|
149
149
|
* found — which is the moment somebody most wants to read it.
|
|
150
150
|
*/
|
|
151
151
|
let lastRun = null;
|
|
152
|
+
/** Every project folder a run page or the board's report has been rendered from, so their pictures resolve against it later (#457, #463). */
|
|
153
|
+
const pageProjects = new ProjectDirs();
|
|
152
154
|
/** Render the report for a session that is about to close, so the live view keeps it. */
|
|
153
155
|
function keepReport(eng) {
|
|
154
156
|
if (!eng.memory)
|
|
@@ -167,6 +169,7 @@ function keepReport(eng) {
|
|
|
167
169
|
}
|
|
168
170
|
try {
|
|
169
171
|
lastRun.evidence = reportEvidence(eng.memory);
|
|
172
|
+
pageProjects.remember(eng.memory.dir);
|
|
170
173
|
lastRun.replay = replayDocument(eng.memory, lastRun.markdown, PKG_VERSION);
|
|
171
174
|
}
|
|
172
175
|
catch (err) {
|
|
@@ -185,6 +188,10 @@ function reportFile(dir) {
|
|
|
185
188
|
return { path: file, written: false };
|
|
186
189
|
}
|
|
187
190
|
}
|
|
191
|
+
/** The address a report's frames start with: keyed by project, and the folder is remembered so the key resolves (#463). */
|
|
192
|
+
function framePrefixFor(dir) {
|
|
193
|
+
return `record/${pageProjects.remember(dir)}/`;
|
|
194
|
+
}
|
|
188
195
|
const liveProvider = {
|
|
189
196
|
snapshot: () => ({
|
|
190
197
|
pid: process.pid,
|
|
@@ -204,9 +211,14 @@ const liveProvider = {
|
|
|
204
211
|
report: () => {
|
|
205
212
|
const eng = (lastWriter && engines.get(lastWriter.session)) ?? engines.values().next().value;
|
|
206
213
|
if (!eng?.memory)
|
|
207
|
-
return lastRun ? { markdown: lastRun.markdown, at: lastRun.at, evidence: lastRun.evidence } : null;
|
|
214
|
+
return lastRun ? { markdown: lastRun.markdown, at: lastRun.at, evidence: lastRun.evidence, framePrefix: framePrefixFor(lastRun.dir) } : null;
|
|
208
215
|
const { markdown } = generateReport(eng.memory, eng.oracleLog.all, reportExtras(eng), { write: false });
|
|
209
|
-
return {
|
|
216
|
+
return {
|
|
217
|
+
markdown: redactSecrets(markdown),
|
|
218
|
+
at: new Date().toISOString(),
|
|
219
|
+
evidence: reportEvidence(eng.memory),
|
|
220
|
+
framePrefix: framePrefixFor(eng.memory.dir),
|
|
221
|
+
};
|
|
210
222
|
},
|
|
211
223
|
/**
|
|
212
224
|
* The whole run as one page, served at its own address. The live view sends
|
|
@@ -219,6 +231,7 @@ const liveProvider = {
|
|
|
219
231
|
return lastRun?.replay || null;
|
|
220
232
|
try {
|
|
221
233
|
const { markdown } = generateReport(eng.memory, eng.oracleLog.all, reportExtras(eng), { write: false });
|
|
234
|
+
pageProjects.remember(eng.memory.dir);
|
|
222
235
|
return replayDocument(eng.memory, redactSecrets(markdown), PKG_VERSION);
|
|
223
236
|
}
|
|
224
237
|
catch (err) {
|
|
@@ -233,12 +246,19 @@ const liveProvider = {
|
|
|
233
246
|
* comes from a viewer, so it is resolved and then required to still be
|
|
234
247
|
* inside that directory: nothing else in the project is reachable this way.
|
|
235
248
|
*/
|
|
236
|
-
frame: async (
|
|
237
|
-
//
|
|
238
|
-
//
|
|
239
|
-
//
|
|
240
|
-
|
|
241
|
-
const
|
|
249
|
+
frame: async (address) => {
|
|
250
|
+
// A run page names the project it was rendered from, so its pictures load
|
|
251
|
+
// after a session attaches to another folder. The live board's own frames
|
|
252
|
+
// carry no project: they belong to the session their path names, and
|
|
253
|
+
// falling back keeps a finished run's frames reachable.
|
|
254
|
+
const { key, rest: relPath } = splitProjectFrame(address);
|
|
255
|
+
let dir;
|
|
256
|
+
if (key)
|
|
257
|
+
dir = pageProjects.dir(key);
|
|
258
|
+
else {
|
|
259
|
+
const named = relPath.replace(/^recordings[\\/]/, "").split("/")[0];
|
|
260
|
+
dir = engines.get(named)?.memory?.dir ?? engines.values().next().value?.memory?.dir ?? lastRun?.dir;
|
|
261
|
+
}
|
|
242
262
|
if (!dir)
|
|
243
263
|
return null;
|
|
244
264
|
const file = resolveFrame(path.join(dir, "recordings"), relPath);
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "scenescout",
|
|
3
|
-
"version": "3.23.
|
|
3
|
+
"version": "3.23.3",
|
|
4
4
|
"description": "SceneScout — exploratory UI testing for AI coding agents. An MCP server that gives any agent (Claude Code, Cursor, VS Code Copilot, Codex, Gemini CLI and others) a structured view of a running web app, always-on oracles, a network-level write policy, memory across runs and a gap-checked report.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"author": "brunoboto96",
|