scenescout 3.15.0 → 3.17.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +87 -0
- package/README.md +70 -18
- package/dist/browsers.js +28 -0
- package/dist/check-run.js +191 -14
- package/dist/ci-run.js +268 -52
- package/dist/cli.js +107 -47
- package/dist/commands.js +3 -2
- package/dist/engine/baseline.js +377 -0
- package/dist/engine/brief.js +16 -7
- package/dist/engine/browser.js +1147 -286
- package/dist/engine/calibration.js +61 -30
- package/dist/engine/capture.js +164 -0
- package/dist/engine/check.js +244 -42
- package/dist/engine/ci-lanes.js +215 -0
- package/dist/engine/ci.js +136 -18
- package/dist/engine/claims.js +159 -3
- package/dist/engine/collector.js +561 -30
- package/dist/engine/crawl.js +49 -0
- package/dist/engine/design.js +281 -38
- package/dist/engine/export.js +877 -0
- package/dist/engine/fingerprint.js +92 -4
- package/dist/engine/flow.js +18 -6
- package/dist/engine/forms.js +181 -18
- package/dist/engine/journey.js +29 -1
- package/dist/engine/lane.js +13 -3
- package/dist/engine/launch.js +45 -6
- package/dist/engine/limits.js +7 -0
- package/dist/engine/live-page.js +49 -2
- package/dist/engine/live.js +4 -1
- package/dist/engine/memory.js +501 -47
- package/dist/engine/open.js +118 -0
- package/dist/engine/oracles.js +41 -1
- package/dist/engine/plain.js +268 -0
- package/dist/engine/png.js +127 -0
- package/dist/engine/policy.js +379 -9
- package/dist/engine/probes.js +3 -2
- package/dist/engine/profiles.js +45 -9
- package/dist/engine/project-folder.js +191 -0
- package/dist/engine/refresh.js +68 -3
- package/dist/engine/replay.js +63 -10
- package/dist/engine/report.js +241 -40
- package/dist/engine/request.js +317 -23
- package/dist/engine/sarif.js +120 -0
- package/dist/engine/settle.js +67 -0
- package/dist/engine/signed-in.js +256 -0
- package/dist/engine/status-pane-page.js +441 -0
- package/dist/engine/status-pane.js +128 -0
- package/dist/engine/tickets.js +671 -0
- package/dist/engine/unload.js +3 -2
- package/dist/export-run.js +633 -0
- package/dist/first-run.js +5 -0
- package/dist/installer.js +378 -8
- package/dist/intake.js +104 -0
- package/dist/login-run.js +250 -36
- package/dist/mcp-server.js +660 -65
- package/dist/playbook.js +5 -0
- package/dist/prompts.js +106 -0
- package/package.json +8 -5
- package/skills/scenescout/SKILL.md +49 -16
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,92 @@
|
|
|
1
1
|
# scenescout
|
|
2
2
|
|
|
3
|
+
## 3.17.0
|
|
4
|
+
|
|
5
|
+
### Minor Changes
|
|
6
|
+
|
|
7
|
+
- b0709f0: `scenescout check` accepts `--ignore-path`: a path drops every rule filed on that route, and `rule:/path` drops that one rule there. A page meant to answer HTTP 500 can stay off the gate while the same status on another path still fails `--fail-on high`. The GitHub Action takes the same input.
|
|
8
|
+
- 547d006: Add the `live` and `login` MCP prompts. `live` returns the current session's loopback live-view URL. `login` takes a role and tells the agent to call `scout_login`, with no password argument.
|
|
9
|
+
- 7ba8712: Add `scout_status`, a run-status pane built as an MCP App (`io.modelcontextprotocol/ui`, specification 2026-01-26). In a client that renders MCP Apps it shows each session's objective and task, open findings by severity, coverage and a button for the live view, and refreshes itself every 2.5 seconds through the app-only `scout_status_poll` tool. Every other client gets the same as text, starting with the live view's loopback address. The pane's page loads nothing from outside and shows only what the live view already shows.
|
|
10
|
+
|
|
11
|
+
### Patch Changes
|
|
12
|
+
|
|
13
|
+
- 745a38e: On Windows, `scenescout install --client` starts an npm-installed client (a `.cmd` or `.bat` shim) through `cmd.exe`, so registration runs the client's own command instead of stopping and printing it to run by hand.
|
|
14
|
+
|
|
15
|
+
## 3.16.0
|
|
16
|
+
|
|
17
|
+
### Minor Changes
|
|
18
|
+
|
|
19
|
+
- 8dacac8: `scout_select` matches the value against the dropdown's options before picking: an exact value or label, either ignoring case, then a label it starts with ("Low" picks "Low — minor impact"). A value that matches no option, or several, is refused at once with the options listed, instead of waiting out the action limit. Typing into a date or time field puts an obvious near miss in the field's format (a plain date into a date-and-time field is entered at T00:00) and refuses anything else naming the format the field takes, instead of failing with "Malformed value". A click forced past something on top of its target names that element, and says when it came after a write-policy block. `scout_run_plan` takes `onViolation: "continue"` for a sweep of independent steps: a new error status is listed on its step's line and the plan goes on, while a failed step, a policy refusal or any other violation still stops it.
|
|
20
|
+
- 11fa895: A run can answer the tickets it was given. `scout_tickets` reads acceptance criteria from pasted text or a ticket file (`.md`, `.txt`, `.feature`): Given/When/Then scenarios, checklists, numbered or `AC1:` criteria, and lists under an "Acceptance criteria" heading; a ticket with none of these is reported as having no recognisable criteria rather than guessed at. `scout_criterion` records each criterion as passed, failed (linked to the findings that show it) or not tested (`no-access`, `observe-blocked` or `out-of-scope`), with the agent's confidence. The report answers each ticket at the top of the plain section, with a failing criterion's pictures, and lists every verdict with its confidence in a new "Acceptance criteria" section of the technical report. A parallel run's lane briefs list the criteria, and the plain questions' "what to check" reads tickets this way.
|
|
21
|
+
- 9ebbce6: Add visual regression baselines to `scenescout check`. List pages and elements in a `targets.json` in the baselines folder (`page`, or an element named the way a saved flow names its target), take their pictures with `--baseline update`, and `--baseline compare` takes them again and compares pixel by pixel: a picture where more pixels changed than `--baseline-threshold` allows (default 0.1%, so anti-aliasing noise between machines and browser builds does not fail the gate; `0` counts every changed pixel), whose size changed, or that could not be taken is a high `visual-change` issue that fails the default gate, with the share of pixels changed and the baseline, the picture now and a diff with the changed pixels in red written under `visual/` beside the report. A baseline that cannot be used (half there, unreadable, or taken with other settings) is filed the same way; a target with no baseline yet is listed, counted beside the verdict, and never fails. Pictures are taken in a fixed 1280×900 window at one pixel per CSS pixel, after a fresh load, with reduced motion requested, fonts loaded, animations stopped and the caret hidden; baselines are kept per browser with a JSON of the settings they were taken with. They live in `.scenescout/baselines/`, which git ignores, unless `--baselines <dir>` names a folder the project commits, and only `--baseline update` writes one, rewriting those compare would not accept and any taken on another operating system. The GitHub Action takes the same three inputs (`baseline`, `baselines`, `baseline-threshold`) and keeps the pictures in its artifact.
|
|
22
|
+
- 819fa87: `scenescout ci --lanes <n>` (and the ci action's `lanes` input) splits an unattended run between up to 8 model loops that explore at once, each in its own browser session and its own modules of the app. A crawl plans the split with no model call, the lanes share the run's turn, token and time caps rather than getting them each, and their findings fold into one report; the summary and `ci.json` list each lane. The default stays one loop: at the default caps, four lanes found no more than one loop on the benchmark's demo app, so raise `--max-turns` and `--max-tokens` with `--lanes`; `--lanes 4 --max-turns 160 --max-tokens 6000000` found the most, at about 2.3 times one loop's cost.
|
|
23
|
+
- 5a83d6e: New tool `scout_network` lists the fetch and XHR requests the current page made since it loaded, with method, path, status and time. Each request shows the route it was sent from, so requests after a client-side route change can be placed. A request that failed, one still pending and one that never ran can be told apart. The tool is read-only, query-string credentials are redacted, the list is bounded, and `scout_request`'s own calls are marked.
|
|
24
|
+
|
|
25
|
+
`scout_request` takes `select` (one value of a JSON body by dotted path, such as `stats.open` or `items.0.name`) and `offset`/`limit` (a window of characters), so a field past the 2000-character cut can be read. Without them the output is the same as before, except that the truncation line now names the options.
|
|
26
|
+
|
|
27
|
+
`scout_coverage` shows a session its own work when other sessions share the project: the routes it reached this run and the forms it saw. `scope: "project"` shows every session's coverage and tags each route and form with the sessions that saw it.
|
|
28
|
+
|
|
29
|
+
Smaller fixes:
|
|
30
|
+
|
|
31
|
+
- A form whose date or time field the app pre-fills, for example with the current time, counts as submitted empty when that value is left as the app set it and every other field is blank. A pre-filled text field still counts as filled, so saving an edit form unchanged is not taken for the empty submit.
|
|
32
|
+
- "Seen in N runs" counts runs, not filings. When the same session files a finding again in the same run, its newer convention and detail replace the earlier ones, and the result says which fields changed or which were kept.
|
|
33
|
+
- A path crawled by name that answers as a page joins the route list. A crawled path that ends on another route is marked `REDIRECTED → <route>`.
|
|
34
|
+
- A journey's time is active time. Gaps over 30 seconds between steps are left out, and the result says how many were left out.
|
|
35
|
+
- 86e9537: Finding dedup no longer merges two findings that both carry evidence on a quoted literal that one names in its title and the other only mentions in its detail: a control's label quoted in passing names where two defects were found, not one defect. The literal must be in the other finding's title or evidence; with no evidence on one side the detail still counts. A finding merged from another page now records that page (`seenOn`, at most 20), and the report prints it as "also seen on" beside the finding's route; the merge note returned to the lane says so. A lane report's decision can name the finding it was filed as (`finding`, the id `scout_finding` returned), and the fold's check for judged defects nobody filed counts it as filed when the project holds that defect. The check's text match also reads underscores inside kebab-case ids and treats ids in paths (`/things/5,/1`) as the route's (`/things/:id`).
|
|
36
|
+
- 8f1cb0f: `scout_attach` and `scout_login` no longer need a `projectPath`; left out, both choose the same folder for a site, so a sign-in saved by `scout_login` is found by the attach after it. Given, it still always wins. Left out, a client that offers a workspace folder gets that folder, and otherwise each tested site gets its own folder, `Documents/SceneScout/<host>/` by default (`localhost-3000` for `http://localhost:3000`), created on first use and named in the attach's result so the person knows where the report is. `SCENESCOUT_PROJECTS_DIR` moves that folder, or `off` makes `projectPath` required again. The default is refused when it would sit inside a git repository below the home folder; a home folder that is itself a repository, as with dotfiles, does not count.
|
|
37
|
+
- 561bc4e: SceneScout installs in Claude Desktop as a desktop extension. Each release now carries `scenescout-X.Y.Z.mcpb`, a bundle in the MCPB manifest format (manifest version 0.3) holding the engine and its dependencies; opening it installs SceneScout with no new chat or restart needed for the install itself. `scenescout doctor` recognises the extension: it checks that the extension's server is in place, checks the Chromium build the extension launches when that differs from its own, with the command that downloads it, and it no longer asks a Claude Desktop-only user to set up the Claude Code skill and registration. After a plugin install, `scenescout install --browser-only` now ends with "Start a new chat to use SceneScout."
|
|
38
|
+
- fd610f0: The first `scout_attach` on a machine without the browser build it needs downloads that build itself, once, saying "Getting the test browser ready" while it does, then carries on with the attach. A Claude Desktop extension or a plugin install needs no terminal step before the first test. CI keeps its explicit install step: there the attach names the command instead, unless `SCENESCOUT_BROWSER_DOWNLOAD=on`. `SCENESCOUT_BROWSER_DOWNLOAD=off` turns the download off for a machine where nothing may be fetched. A failed download names the command to run by hand. `scenescout doctor` no longer fails a desktop extension whose browser is not downloaded yet when the extension is at least as new as `doctor`, since it downloads on first use; it still names the command to have it ready beforehand.
|
|
39
|
+
- d7af6e7: Add `scenescout export --to github|jira`, which files the project's open findings as GitHub or Jira issues, each once: every issue carries the `scenescout` label and a marker with the finding's id, and an export skips every finding that already has an issue, open or closed (`--refile-closed` files one again when its issue is closed), so a second export of the same run files only what the first left over the cap. It is a dry run that lists what it would file unless `--yes` is given, and files at most `--max-issues` (default 20) per export. Issue text is rendered inert (no mention, link, cross-reference or markup from a finding does anything); severity becomes a label on GitHub and a priority in Jira (`--severity-map`); a recorded run's frames are attached in Jira and named in a GitHub issue. Credentials come from `GH_TOKEN` or `GITHUB_TOKEN`, or `JIRA_EMAIL` and `JIRA_API_TOKEN`, only, and are never printed; requests time out, wait out short rate limits, retry failed reads with backoff, never re-send a create the tracker may have carried out without first looking for its marker, and refuse redirects.
|
|
40
|
+
- 362e52d: Every finding is filed with a picture of what it is about: the element `scout_finding {ref}` names, with a margin, or the page as it was. The picture is kept in `.scenescout/recordings/`, shown under the finding in `report.html` and in the live view's report, named in `report.md`, and returned in the `scout_finding` result as image content so a chat client shows it as the finding is filed. Pictures are bounded by `SCENESCOUT_EVIDENCE_MAX_PX` (default 800 pixels on the longer side) and `SCENESCOUT_EVIDENCE_MAX_KB` (default 200), and a session returns at most `SCENESCOUT_EVIDENCE_INLINE` (default 10) in its results. `scout_attach {evidence}` or `SCENESCOUT_EVIDENCE` chooses `inline`, `file` or `off`; a CI job, and `scenescout ci`, default to `file`. `SCENESCOUT_RECORD=on` makes every session record a frame after each action without each attach asking.
|
|
41
|
+
- ffb9a38: `scenescout export --to jira` keeps a filed issue up to date and links it to the ticket it fails. A later export rewrites an open issue's summary and description when the finding has changed, unless someone has edited them in Jira since, and adds the picture, frames and ticket links it lacks, rather than leaving it as first filed (`--jira-update off` only lists it). The finding's picture is attached first, and a finding that fails a ticket's acceptance criterion is linked to that ticket (`--jira-link-type`, default `Relates`, or `JIRA_LINK_TYPE`; `none` links nothing). GitHub issues name the picture and list the failed criteria.
|
|
42
|
+
- 267e24a: `scenescout login` no longer needs Enter in a terminal: the window saves the role's profile and closes by itself once the person is signed in, meaning back on the app, with no password or code field on the page, past any return from a single sign-on provider, and holding a session cookie or storage entry it did not hold when the window opened. A round trip through an identity provider on another site, a sign-in popup still open on one, and the app's own page before it has exchanged the provider's code are never taken for the end. Enter still saves at once; `--save enter` makes it the only way, as before, and `--success-url` names the signed-in address instead. The new `scout_login { url, role, projectPath }` tool opens the same window from a conversation and returns once the sign-in is saved, or after `waitSeconds` (default 120) with the window still open, so the agent can call it again; a window nobody finishes closes after 15 minutes, saving nothing. A role with no saved login now names `scout_login` beside the command when an attach is refused.
|
|
43
|
+
- df40cfe: The engine can open the live view in the default browser when a session attaches, and report.html when `scout_report` writes it. The new `open` setting (`scout_attach {open}` or `SCENESCOUT_OPEN`: `live`, `report`, `both` or `none`) defaults to both on a local desktop session, headed or headless, and to none in CI, over SSH, or on Linux with no display. `scenescout ci` opens nothing unless `SCENESCOUT_OPEN` is set. The live view keeps its loopback-and-token rules.
|
|
44
|
+
- 180d502: The report now opens in plain words: a short summary, then each problem with its impact (blocks users, annoying or cosmetic), numbered steps, what was expected, what happened and its picture, with the technical detail folded beneath each one. `report.md` puts this section before the technical report, and `report.html` opens on it. `scout_report {report}` chooses the parts: `both` (default), `qa` for the plain section alone, or `dev` for the technical report alone.
|
|
45
|
+
- 8046682: Start a run with plain questions instead of settings. `/scenescout` with no flags, or the `explore` prompt with no arguments, now asks for the address, whether and how to sign in, what to check (tickets or a description) and whether the site holds real data, and chooses the URL, the sign-in, the objective and the write mode from the answers: real data, or an unsure answer, means observe. Any flag skips the questions. The skill gains `--focus <text>` and `--read-only`.
|
|
46
|
+
- 4d4443d: Observe mode can be told which POST endpoints only read, and the label check no longer refuses controls for the record text they show.
|
|
47
|
+
|
|
48
|
+
- `scout_attach` takes `readPosts` (`["POST /api/search"]`), and `SCENESCOUT_READ_POSTS` sets the same for `check`, `ci` and a first look. Observe then lets those POSTs out, so a search or query page that loads its data through POST can be tested. Nothing is named by default. A named endpoint is still refused when its path or body looks destructive or its body is a GraphQL mutation, and each one let out is logged. The gap ledger names the pages where observe refused a script's POST, with the endpoint.
|
|
49
|
+
- A row, card, heading or panel is judged by its test id and the control at its centre, not by the record text it shows. Its own text counts only for a clickable element whose text is a short command, and a heading's never does.
|
|
50
|
+
- Removing a filter chip ("Remove Status: Open filter") is allowed; "Remove member" is still refused.
|
|
51
|
+
- "Sign off" is refused only as a command, at the start of a label or joined to another verb ("Save and sign off"). "Manager sign-off", "Final sign-off recorded" and "Confirm sign off" are allowed.
|
|
52
|
+
- 1a98cb4: Snapshot refs last longer and re-snapshots cost less. Refs from the last snapshot keep working after a search or filter that rewrites only the query string, and a control a re-render replaced is found again by its unique test id (the action says it was re-bound); a route change still refuses them, and the next diff now says when refs were dropped instead of calling them stable. A route you come back to is shown as a diff against its own last snapshot, and another tab of the same screen as a diff against that screen's last tab. Repeated rows are matched by their text or link, so a filtered list reads as the rows that went rather than as the first row relabeled. A truncated snapshot says what it cut, by role and test-id family, and keeps pager and "Load more" controls in the list. An element listed without a control role that a click or Tab still reaches is marked `[clickable]` or `[focusable]`.
|
|
53
|
+
|
|
54
|
+
### Patch Changes
|
|
55
|
+
|
|
56
|
+
- c6a75e2: `scout_request` replays the Authorization header the page last sent to the app's own origin on any request, reads included. It used to replay the header of the page's last write, which went stale once the app rotated its access token and then only read, so a replay got 401 while the page's own calls got 200. A header sent to another origin is never replayed, nor is one a `scout_request` call chose for itself. When a replay still gets 401 while the page's latest authorised call succeeded, the result says the replayed credential may be stale.
|
|
57
|
+
|
|
58
|
+
The refresh broker's write-back keeps the profile's IndexedDB. It used to save the page's storage state without it, so an app keeping part of its sign-in in IndexedDB lost it from the role's profile at the first brokered refresh.
|
|
59
|
+
|
|
60
|
+
An action's result now says when the refresh broker acted during it: the token refreshed and stored, another session's rotated token loaded and sent, an endpoint learned, or a refresh that could not be brokered. Counts only, no token values.
|
|
61
|
+
- fe6ca3c: `false_success` no longer pairs a write sent as a clicked link loads the next page (the old page's save as it is left, the new page's beacon as it loads) with that page's static text, and it now reports, at medium, a refused write whose control or counter shows the change as kept with no error. Errors an HTTP client raises over the write policy's stand-in 403 ("Request failed with status code 403", "403 Forbidden") are attributed to the policy, an alert that appeared after a block is marked `(after a write-policy block)` in the snapshot, and the silent-submit note no longer fires when the click opened a dialog or client-side validation answered. Error-monitor tunnel and client-error endpoints count as infrastructure writes.
|
|
62
|
+
- 4cb1441: The gap ledger's three route lines now count over the known routes and name that set (`3 of 54 known route(s) …`), treat a tab or section of a page as part of that page, and judge "nothing exercised" and "never design-audited" over the routes this run reached rather than every earlier run's. `scout_coverage` in a session's scope lists only the controls on the pages that session saw, not another role's on the same route, and labels the route figure as the project's. A wrapper flagged as not a control in any state of a route no longer counts in coverage because an older state left it unflagged. Record ids shaped like codes (`WID-2025-001`, `A1B2C3`) collapse to `:id` like numeric ones, and a dropdown's option already selected when the page loaded is no longer listed as never chosen.
|
|
63
|
+
- 313dfab: Make the design audit agree with the page it measures:
|
|
64
|
+
|
|
65
|
+
- A filter panel is no longer judged as a form. With no `<form>` on the page, the form-burden lines ("NONE marked required", "no obvious submit") count only the fields a user types into, and fields in a panel that names itself a filter or facet (its test id, id, label or legend) are left out wherever they are. Six text fields with no `<form>` and no submit are still flagged.
|
|
66
|
+
- The shadow census counts the layers a box-shadow draws. Empty ring layers (no offset, blur or spread) and transparent layers are dropped, so utility-CSS rings are no longer counted as elevations, and each example shows the layer that sets it apart instead of a truncated value.
|
|
67
|
+
- Tinted grays (a slate, a warm stone) are counted as grays: any colour with no hue family is one. Grays within a few units of each other count as one step of the scale.
|
|
68
|
+
- Elements are named in audit lines by their accessible name, computed as the snapshot computes it, so an icon button with `aria-label="Dismiss"` reads as "Dismiss" rather than "(no text)". The focus-indicator line names tab stops the same way.
|
|
69
|
+
- A new NAMES section lists controls with no accessible name and fields labelled only by their placeholder, and both now lower the a11y subscore (a link or button whose only content is an image with alt text, or an svg with a title or aria-label, is named by it and not listed), which before measured contrast, focus visibility and target size only. Page scores on pages with such controls go down.
|
|
70
|
+
- c6a75e2: A click that submits a native form to a refresh endpoint the refresh broker handles returns once the navigation commits, with the broker's line in its own result. In Chromium it used to wait out the action limit, because loading another session's rotated profile opened a page that Chromium did not finish while the form's navigation was held. While it holds a navigation, the broker now loads only the profile's cookies; the rest of the profile is loaded at the next refresh a script sends, and the form's write-back saves the page's cookies while keeping the profile's storage.
|
|
71
|
+
- 664c24d: A request still loading when its page is left, its frame is removed or its frame navigates away is no longer counted as in flight, so the next wait for the page to go quiet no longer runs to its 2 s cap.
|
|
72
|
+
- c6dc360: Stop the layout checks reporting intended layering as defects, while the defective form of each layout is still reported:
|
|
73
|
+
|
|
74
|
+
- A skip link parked off the page until it takes focus (a link to a place in the page, or any control a `:focus` rule moves) is no longer "outside the reachable page area", and once focused it is not reported as covering the header under it. A link parked off the page with nothing to bring it back, or a hash-route link, still is.
|
|
75
|
+
- A control inside a list that scrolls within an overflow-hidden card is no longer UNREACHABLE, nor is a slide of a viewer that a control naming it in `aria-controls`, or a next, previous or numbered control beside a row of slides, reveals. The same content with no scroller and no pager still is, and so is a column a card clips beside a "Next page" that pages rows.
|
|
76
|
+
- The overlap check skips a decorative overlay with `pointer-events: none` and a clear button or icon lying in the padding a text field reserves for it. A button over the field's text, an overlay that takes clicks, or a disabled control drawn over another still overlaps.
|
|
77
|
+
- The small-target rule measures a native input together with the label that wraps or touches it, and a visually hidden input as the label or drop zone that operates it, and applies the WCAG 2.5.8 spacing exception: a small target with nothing else inside its 24px circle passes.
|
|
78
|
+
- 327e0c6: A path given to `scout_navigate`, `scout_crawl`, a plan or a replayed flow now resolves against the attached origin, as `scout_request` already did, so a session attached on a page below the root (`/things`) goes to `/widgets` for `/widgets` instead of `/things/widgets`, and a crawl given a full same-origin URL visits it as given. A crawled route whose main area holds only an alert or a loading placeholder is flagged `ERROR-VIEW` or `STILL-LOADING`, listed under problem routes, and no longer joins the route contract.
|
|
79
|
+
- 28f9b79: The write policy's notices and label check are easier to work with, and refuse as much as before.
|
|
80
|
+
|
|
81
|
+
- A blocked endpoint is named and explained once per session; later blocks of it are counted on one line, so a page that beacons on every load no longer buries each tool result.
|
|
82
|
+
- A control is judged by its own label. A dropdown is judged by the option picked, so a filter that offers "Delete" can be set to "Create". A row or panel is judged by its own text, not by the buttons inside it, plus any control covering its centre. A pick whose label cannot be read is refused.
|
|
83
|
+
- "Discard changes" and similar labels, which drop only unsent input, are no longer refused. "Discard draft" or "Discard record" still is, and `discard` in a request path is now treated as destructive on the network.
|
|
84
|
+
- When a page asks to confirm leaving unsent input, the result says so by name instead of failing with `ERR_ABORTED`. `scout_navigate`, `scout_click` and `scout_back` take `leave`: observe and read-only stay unless it is `true`, and other modes leave unless it is `false`. Every native dialog the page opens is reported in the action's result.
|
|
85
|
+
- f0c8bf3: `check.sarif` and `ci.sarif` results now point at a repository file, so GitHub code scanning keeps them instead of dropping every one. An issue a saved flow raised points at that flow's file; every other result points at the anchor: the new `--sarif-file-anchor` option (and `sarif-file-anchor` action input) when given, else the running workflow's file from `GITHUB_WORKFLOW_REF`, else `package.json`, else `README.md`, whichever exists first; a missing option or workflow file is named in a warning, and when none exists the SARIF is still written with a warning that code scanning will drop its results. The page each result was seen on moves to the message, a logical location and `properties`; fingerprints are unchanged, so existing alerts keep their identity. In `check.json`, an issue a saved flow raised names that flow's file in `flow`.
|
|
86
|
+
- e698580: The snapshot names a link or button whose only content is an image by that image: an `<img>`'s alt text, or an `<svg>`'s or `role="img"` element's aria-label or `<title>`, skipping anything under `aria-hidden="true"`. `<a href="/"><img alt="Home"></a>` is now listed as `link "Home"` rather than `link ""`, so the crawl, `scenescout check` and the a11y counts no longer report it as unnamed, and the read-only policy now judges an image button by the name it is announced with. Text still comes first: a control with both text and an image is named by its text, as before. The design audit reads this same name, so the audit and the snapshot use one rule.
|
|
87
|
+
|
|
88
|
+
The name is half of a control's coverage key, so such a control gets a new key (`button "Search"` instead of `button ""`), and two unnamed image buttons that shared one key now have one each. Memory written before this change carries over: each snapshot records the key a control had under the earlier rule, and coverage reads older states through it, so what was exercised stays exercised and the earlier key is not left as a gap. On a route this run does not reach, the earlier keys stay as they are until a snapshot there lists the controls under their new names.
|
|
89
|
+
|
|
3
90
|
## 3.15.0
|
|
4
91
|
|
|
5
92
|
### Minor Changes
|
package/README.md
CHANGED
|
@@ -55,7 +55,7 @@ An excerpt of the report it wrote — [read the whole thing](examples/report.md)
|
|
|
55
55
|
> **🟡 [LOW] The dashboard chart image is missing** *(callout 2)*
|
|
56
56
|
> Evidence: `GET /img/weekly-chart.png → HTTP 404`
|
|
57
57
|
>
|
|
58
|
-
> **Gap ledger — what was NOT tested:** 9
|
|
58
|
+
> **Gap ledger — what was NOT tested:** 9 of 12 known routes visited this run and never design-audited · single-role run, so permission boundaries are untested
|
|
59
59
|
|
|
60
60
|
Every finding comes with a repro trace and a Playwright regression-test skeleton. To try it yourself, clone this repository, run `npm run demo:serve`, then `/scenescout --url http://127.0.0.1:4173` — see [demo-app/](demo-app/). Its README lists every seeded defect and which oracle catches it.
|
|
61
61
|
|
|
@@ -152,9 +152,11 @@ Either way it downloads the browser and registers the server with the client you
|
|
|
152
152
|
/plugin install scenescout@scenescout-marketplace
|
|
153
153
|
```
|
|
154
154
|
|
|
155
|
-
Then
|
|
155
|
+
Then start a new chat to use SceneScout; the test browser downloads on first use (to have it ready beforehand: `npx -y scenescout install --browser-only`). The command becomes `/scenescout:scenescout`. A plugin's skill comes from this repository and its server from the latest npm release, so right after a release lands here the two can differ for a short while; `/plugin marketplace update scenescout-marketplace` brings the skill up to date.
|
|
156
156
|
|
|
157
|
-
**
|
|
157
|
+
**Using Claude Desktop?** Install the extension: download `scenescout-X.Y.Z.mcpb` from the [latest release](https://github.com/brunoboto96/SceneScout/releases/latest) and open it (or Settings > Extensions > Advanced settings > Install Extension). It works as soon as it is installed, with no terminal step: the test browser downloads on first use. Start a new chat and ask *"Use SceneScout to test http://localhost:3000"*. [More in the guide](docs/guide/Start-here.md#as-a-claude-desktop-extension).
|
|
158
|
+
|
|
159
|
+
**A client that is not in that list?** [Add the server to its config by hand](#-other-mcp-clients); the test browser downloads on first use.
|
|
158
160
|
|
|
159
161
|
<details>
|
|
160
162
|
<summary>What <code>install</code> actually does</summary>
|
|
@@ -201,7 +203,9 @@ In Claude Code the skill gives you a command with flags for the same thing:
|
|
|
201
203
|
|
|
202
204
|
The agent scans the project (if there is one), attaches read-only, explores, and writes findings to `.scenescout/report.md`. That's it.
|
|
203
205
|
|
|
204
|
-
**Common flags** — `--level minimal|medium|extensive` · `--url <app>` · `--role <name\|path>` (who to explore as: a login saved with `scenescout login`, a storage state found by the scan, or a path to a Playwright storage-state JSON) · `--observe` / `--safe-write` / `--allow-destructive`.
|
|
206
|
+
**Common flags** — `--level minimal|medium|extensive` · `--url <app>` · `--role <name\|path>` (who to explore as: a login saved with `scenescout login`, a storage state found by the scan, or a path to a Playwright storage-state JSON) · `--focus <text>` (a ticket or a sentence to check) · `--observe` / `--read-only` / `--safe-write` / `--allow-destructive`.
|
|
207
|
+
|
|
208
|
+
**No flags at all** (`/scenescout` on its own) and the agent asks four plain questions instead: the address, whether and how you sign in, what to check (tickets or a description), and whether the site holds real data. Real data, or not being sure, means nothing but `GET` requests leave the page; you are never asked to pick a mode. Any flag skips the questions. See [Plain questions instead of flags](docs/guide/Ways-to-use-it.md#plain-questions-instead-of-flags).
|
|
205
209
|
|
|
206
210
|
### 🔑 Signing in as a role
|
|
207
211
|
|
|
@@ -211,7 +215,7 @@ For an app behind SSO or MFA, sign in once yourself and let every session reuse
|
|
|
211
215
|
scenescout login http://localhost:3000 --role admin
|
|
212
216
|
```
|
|
213
217
|
|
|
214
|
-
A browser window opens at the URL. Sign in however the app asks
|
|
218
|
+
A browser window opens at the URL. Sign in however the app asks: once you are back on the app with a new session, the window saves it as `.scenescout/auth/admin.json` in the project and closes by itself. A trip through a single sign-on provider and back is followed, not taken for the end. Pressing **Enter** in the terminal saves at once, and `--save enter` makes Enter the only way, as before. Closing the window or pressing Ctrl+C saves nothing. From a conversation, the agent opens the same window with `scout_login`, so no terminal is needed. The file is readable by your account only, `.scenescout/` keeps itself out of git, and the command prints where it saved, how many cookies, origins and databases it holds, never what they are, and how long it will last: read from each cookie's expiry and the `exp` of any JWT in a cookie or in localStorage (the payload is decoded for that one claim, never verified, never printed). The profile keeps cookies, localStorage, IndexedDB and sessionStorage, so an app whose sign-in library keeps its token in sessionStorage or IndexedDB still comes back signed in; sessionStorage is put back only on the origin it came from, once per tab, so a lane that signs out stays signed out. A login saved by an earlier version has no sessionStorage or IndexedDB: record it again if the app keeps its token there. `--project <dir>` saves into another project; `--browser firefox|webkit` records in another browser.
|
|
215
219
|
|
|
216
220
|
Then `/scenescout --role admin`, or `scout_attach { role: "admin" }` from any agent. Every session attached with the same role gets its own browser built from that one login, so parallel lanes can all run as `admin`. A role with no saved login is refused with the command to run. `role` and `storageStatePath` are alternatives: pass one.
|
|
217
221
|
|
|
@@ -225,7 +229,7 @@ Before a parallel run, `scout_lane_brief` checks that the planner's saved login
|
|
|
225
229
|
|
|
226
230
|
## 📺 Watching a run live
|
|
227
231
|
|
|
228
|
-
When a session attaches, the engine starts a small live view and hands the agent its address on a `Live view:` line, which the agent passes on to you. From a terminal, `scenescout watch` opens the same page. There is one card per session:
|
|
232
|
+
When a session attaches, the engine starts a small live view and hands the agent its address on a `Live view:` line, which the agent passes on to you. On a local desktop it also opens that page in your default browser as the session attaches, and opens `report.html` when `scout_report` writes it, whether or not the browser window is shown. Nothing opens in CI, over SSH, or on Linux with no display. `SCENESCOUT_OPEN` (`live`, `report`, `both` or `none`) in the server's environment chooses otherwise, and `scout_attach {open}` wins over it. `scenescout ci` opens nothing unless `SCENESCOUT_OPEN` is set. From a terminal, `scenescout watch` opens the same page. There is one card per session:
|
|
229
233
|
|
|
230
234
|
<p align="center"><img src="examples/screenshots/live-view.png" alt="The live view during a run of three parallel agents against the demo app: one card per session, each with its role and objective, the task it is on, the tool it is running, the page it is on, a live thumbnail, and a feed of the actions it just took, tinted one colour per task" width="880" /></p>
|
|
231
235
|
|
|
@@ -254,21 +258,33 @@ The view is served on `127.0.0.1` only, behind a token that changes every time t
|
|
|
254
258
|
## 🎬 Recording a run, and reading it back
|
|
255
259
|
|
|
256
260
|
A report says what happened. For QA work that is not always enough — the point
|
|
257
|
-
is often to *show* what was checked, not to assert it.
|
|
258
|
-
|
|
261
|
+
is often to *show* what was checked, not to assert it.
|
|
262
|
+
|
|
263
|
+
Every finding already carries a picture: the element it is about, with a margin,
|
|
264
|
+
or the page as it was. It is kept in `.scenescout/recordings/`, shown under the
|
|
265
|
+
finding in `report.html`, and returned with the `scout_finding` result, so a
|
|
266
|
+
chat client shows the evidence the moment it is filed. Pictures are bounded in
|
|
267
|
+
size and in how many reach the conversation, and a CI job keeps them on file
|
|
268
|
+
only; `SCENESCOUT_EVIDENCE` and `scout_attach {evidence}` change that
|
|
269
|
+
([configuration reference](docs/guide/Configuration-reference.md#environment-variables)).
|
|
270
|
+
|
|
271
|
+
Ask for a recorded run and the engine also keeps a frame of the page after every action:
|
|
259
272
|
|
|
260
273
|
```
|
|
261
274
|
Use SceneScout to test http://localhost:3000, record the run
|
|
262
275
|
```
|
|
263
276
|
|
|
264
|
-
or, on the tool directly, `scout_attach {record: true}`.
|
|
277
|
+
or, on the tool directly, `scout_attach {record: true}`. `SCENESCOUT_RECORD=on` in
|
|
278
|
+
the server's environment records every run.
|
|
265
279
|
|
|
266
280
|
Then `scout_report` writes two files side by side in `.scenescout/`:
|
|
267
281
|
`report.md` as always, and `report.html` — the whole run as one self-contained
|
|
268
282
|
page. It opens from the file system with nothing running, needs no network, and
|
|
269
283
|
holds:
|
|
270
284
|
|
|
271
|
-
- **The report**, rendered from the same Markdown
|
|
285
|
+
- **The report**, rendered from the same Markdown: the plain-language view
|
|
286
|
+
first (each problem's steps, what was expected, what happened and its
|
|
287
|
+
picture), with each problem's technical detail one click away.
|
|
272
288
|
- **The screenshots around each finding**, in an accordion under it, from the
|
|
273
289
|
session that filed it.
|
|
274
290
|
- **Every session's trail**, in the blocks its tasks made, each step with the
|
|
@@ -313,17 +329,18 @@ Snapshots are cheap: re-snapshotting a route returns only *what changed*, with s
|
|
|
313
329
|
|
|
314
330
|
## 🧰 The toolbox
|
|
315
331
|
|
|
316
|
-
|
|
332
|
+
34 deterministic tools. The agent picks; you rarely call these by hand.
|
|
317
333
|
|
|
318
334
|
| Phase | Tools | What they do |
|
|
319
335
|
|---|---|---|
|
|
320
336
|
| **Set up** | `scout_playbook` `scout_scan` `scout_attach` `scout_session` | Hand the testing method to an agent that has no skill loaded; discover routes; launch a browser in a write-mode; keep several authenticated roles alive at once |
|
|
321
337
|
| **Explore** | `scout_crawl` `scout_coverage` | Sweep every route in one call; ask what's still untested |
|
|
322
338
|
| **Look** | `scout_snapshot` `scout_hover` `scout_screenshot` `scout_capture` | Read the structured scene (diffed); reveal tooltips/hover cards; capture pixels only when needed; save one element as a PNG to show someone |
|
|
323
|
-
| **Ask the server** | `scout_request` | Call the app's own API as this session, with the UI bypassed — the check that turns a hidden button into a proven refusal |
|
|
339
|
+
| **Ask the server** | `scout_request` `scout_network` | Call the app's own API as this session, with the UI bypassed — the check that turns a hidden button into a proven refusal; list the requests the page itself made since it loaded |
|
|
324
340
|
| **Act** | `scout_click` `scout_type` `scout_select` `scout_upload` `scout_press` `scout_scroll` `scout_navigate` `scout_back` `scout_run_plan` | Drive the UI like a user; `scout_run_plan` batches a whole mechanical sequence into one call |
|
|
325
341
|
| **Assess** | `scout_design_audit` `scout_journey` | Score a page's craft/a11y/consistency; measure how hard a task is to complete |
|
|
326
342
|
| **Record** | `scout_note` `scout_finding` `scout_resolve` `scout_report` | Curate durable notes; file deduped findings; mark fixes; write the report, and on a recorded run the whole run as one page |
|
|
343
|
+
| **Answer tickets** | `scout_tickets` `scout_criterion` | Read the acceptance criteria in pasted or uploaded tickets; record each criterion as passed, failed (with the findings that show it) or not tested (and why), with a confidence |
|
|
327
344
|
| **Re-test** | `scout_verify` | List the findings earlier runs left open, worst route first, and record whether each is gone, still present, or changed |
|
|
328
345
|
| **Split the work** | `scout_lane_brief` `scout_lane_report` | Divide the app between parallel agents by whole module, each with its own landing route and rules; fold what each hands back as one typed JSON object, and name any defect it judged but never filed |
|
|
329
346
|
| **Close** | `scout_close` | Tear down one session or all |
|
|
@@ -361,14 +378,18 @@ That refusal *is* the guarantee: an extensive report can only exist when nothing
|
|
|
361
378
|
- 🟢 **`read-only` by default.** Destructive-labeled elements (delete/revoke/archive/…) **and** all `PUT/PATCH/DELETE` + destructive `POST`s are blocked at the network layer — see [`src/engine/policy.ts`](src/engine/policy.ts). Non-destructive `POST`s are allowed, because submitting forms is how a tester finds validation bugs — so read-only means *nothing existing is changed or removed*, not *nothing is ever created*.
|
|
362
379
|
- 🟡 **`safe-write`** (`--safe-write`) lets the agent create data and edit/delete **only what it created** this run — never pre-existing records.
|
|
363
380
|
- 🔴 **`destructive`** (`--allow-destructive`) allows everything, and only ever when *you* confirm the environment is disposable. The skill will never choose this itself.
|
|
364
|
-
- 📂 Findings, memory, and reports live in a `.scenescout/` folder where
|
|
381
|
+
- 📂 Findings, memory, and reports live in a `.scenescout/` folder in the project. A client with no project folder, such as a desktop chat, gets one folder per tested site under `Documents/SceneScout/<host>/` by default, and the attach says where; `SCENESCOUT_PROJECTS_DIR` moves it. The folder ignores itself in git, so a stray `git add -A` never commits test data.
|
|
365
382
|
|
|
366
383
|
A `🛡 WRITE-POLICY blocked` notice is the safety net doing its job, not an app bug. The server never sees a blocked request, but a page's own `fetch` or XHR is answered with a `403` in its place rather than dropped, so the page's handling of a refusal really runs: a page that then claims success is reported as a `false_success` ([ADR 9](docs/adr/0009-a-refused-write-is-answered-not-dropped.md)).
|
|
367
384
|
|
|
385
|
+
A control is judged by its own label: a dropdown by the option picked, a row by its own text rather than the buttons inside it, and "Discard changes" on an unsent form is allowed. When a page asks to confirm leaving unsent input, the result says so; `observe` and `read-only` stay unless the call passes `leave: true`. See the [safety model](docs/guide/Safety-model.md).
|
|
386
|
+
|
|
368
387
|
---
|
|
369
388
|
|
|
370
389
|
## 📋 What you get
|
|
371
390
|
|
|
391
|
+
`.scenescout/report.md` and `report.html` open **In plain words**: a short summary, then each problem this run found, worst first, with its impact, the steps that led to it, what was expected, what happened and a picture when there is one. The technical detail (id, category, evidence, route) stays one click away.
|
|
392
|
+
|
|
372
393
|
`.scenescout/report.md` — a deduplicated, worst-first report with:
|
|
373
394
|
|
|
374
395
|
- 🐛 **Findings** with repro traces and generated Playwright regression-test skeletons.
|
|
@@ -379,9 +400,9 @@ A `🛡 WRITE-POLICY blocked` notice is the safety net doing its job, not an app
|
|
|
379
400
|
- ⏱️ **How the run was paced** — how closely each session kept working, and apart from that, how long finished lanes held their browsers waiting to be collected, so neither hides the other.
|
|
380
401
|
- 🎯 **How well the lanes judged** — on a parallel run, whether the confidence each lane stated matched what the project went on to file, beside what later re-tests found ([ADR 10](docs/adr/0010-a-confidence-is-checked-not-trusted.md)).
|
|
381
402
|
|
|
382
|
-
`.scenescout/report.html` — the same report as one self-contained page, with every session's trail beside it
|
|
403
|
+
`.scenescout/report.html` — the same report as one self-contained page, with every session's trail beside it. Each finding shows a picture of the element it is about, or of the page as it was. A [recorded run](#-recording-a-run-and-reading-it-back) also shows the screenshots around each finding.
|
|
383
404
|
|
|
384
|
-
👀 Watch a run live: `
|
|
405
|
+
👀 Watch a run live: `npx scenescout watch`, or `npx scenescout status <project-path>` for the same information as text.
|
|
385
406
|
|
|
386
407
|
---
|
|
387
408
|
|
|
@@ -430,6 +451,7 @@ With the default settings its saved flows send no HTTP write (they replay under
|
|
|
430
451
|
- `--flow-writes never|allow` (default `never`): `never` replays flows under observe's rule whatever `--mode` says; `allow` replays them under `--mode`, so in `read-only` a flow's form submissions are sent to the target on every run.
|
|
431
452
|
- `--on-refused-step report|stop` (default `report`): `report` marks a flow whose step was refused "could not run", keeps every other verdict and exits 2; `stop` exits 2 at that step with no results.
|
|
432
453
|
- `--gate-retests never|high|all` (default `high`): which still-reproducing re-tested findings fail the gate.
|
|
454
|
+
- `--baseline off|compare|update` (default `off`), with `--baselines <dir>` and `--baseline-threshold <percent>` (default `0.1`, so small anti-aliasing noise between machines passes): visual baselines, below.
|
|
433
455
|
|
|
434
456
|
The defaults are what an unconfigured check does, for a first try or an AI agent running it unattended: its flows send no HTTP write and it never silently hides a result. Each setting is a choice for the project; the report and `check.json` print the values a check ran with.
|
|
435
457
|
|
|
@@ -437,6 +459,8 @@ The defaults are what an unconfigured check does, for a first try or an AI agent
|
|
|
437
459
|
|
|
438
460
|
It also replays the flows saved in `.scenescout/flows/*.json`, with no model: the steps `scout_run_plan` takes (navigate, click, type, select, press) plus `expect-text`, `expect-url` and `expect-request`. A flow whose step breaks fails the gate, naming the flow and the step. And it re-tests the open findings earlier runs left in the project's memory that a page load can reproduce, reporting each as still reproducing or possibly fixed; by default a finding filed high that still reproduces fails the gate. [docs/ci.md](docs/ci.md#saved-flows) has the flow format; [ADR 12](docs/adr/0012-a-check-replays-saved-flows-and-reports-re-tests.md) says why it works this way.
|
|
439
461
|
|
|
462
|
+
With `--baseline compare` it also holds pages and elements to approved pictures: list them in a `targets.json`, take the baselines once with `--baseline update`, and a later check that finds one changed past `--baseline-threshold` fails the gate with the share of pixels changed and a diff picture beside the report. Baselines are kept per browser, in a git-ignored folder unless `--baselines` names one the project commits, and only `--baseline update` ever writes one. [The guide](docs/guide/Ways-to-use-it.md#visual-baselines) has the details; [ADR 19](docs/adr/0019-a-visual-baseline-changes-only-when-asked.md) says why.
|
|
463
|
+
|
|
440
464
|
Beyond those flows it explores nothing and fills no forms. That is the exploratory run's job, and its findings belong in a report, not a gate.
|
|
441
465
|
|
|
442
466
|
## 🤖 In CI: an unattended exploratory run
|
|
@@ -459,6 +483,26 @@ There is a GitHub Action for it (`uses: brunoboto96/SceneScout/ci@…`). [docs/c
|
|
|
459
483
|
|
|
460
484
|
On a pull request, an allowed account can comment `/scenescout qa` to run it against that pull request's deployed preview and get the results as a reply. The job that holds the key checks out nothing and runs SceneScout from an exact release tag, so the pull request's code never runs beside the key. [docs/ci.md](docs/ci.md#a-qa-review-from-a-pull-request-comment) has the workflow to copy and what a project configures; [ADR 15](docs/adr/0015-a-qa-comment-tests-a-preview-and-never-runs-the-pull-requests-code.md) says why.
|
|
461
485
|
|
|
486
|
+
## 📮 Filing findings as issues
|
|
487
|
+
|
|
488
|
+
`scenescout export` turns the project's open findings into GitHub or Jira issues, where the team already works. It reads `.scenescout/memory.json`, so it runs after an interactive run, `scenescout ci` or anything else that wrote findings:
|
|
489
|
+
|
|
490
|
+
```bash
|
|
491
|
+
export GH_TOKEN=… # or GITHUB_TOKEN; for Jira, JIRA_EMAIL and JIRA_API_TOKEN. Read from the environment only
|
|
492
|
+
npx scenescout export --to github --repo owner/app # a dry run: lists what it would file
|
|
493
|
+
npx scenescout export --to github --repo owner/app --yes # files it
|
|
494
|
+
npx scenescout export --to jira --jira-url https://your-site.atlassian.net --jira-project QA --yes
|
|
495
|
+
```
|
|
496
|
+
|
|
497
|
+
- **Each finding once.** Every issue carries the `scenescout` label and a marker with the finding's id. Before filing, the export reads the labelled issues, open or closed, and skips every finding already filed, naming its issue, so a second export of the same run files only what the first left over the cap. A closed won't-fix is not filed again; `--refile-closed` files a finding again when its issue is closed.
|
|
498
|
+
- **Jira issues are kept up to date and linked to the ticket.** A later export updates an open Jira issue instead of filing another, leaving text someone edited in Jira as they wrote it. The finding's picture is attached, and a finding that fails a ticket's acceptance criterion is linked to that ticket (`--jira-link-type`, default `Relates`).
|
|
499
|
+
- **A dry run unless `--yes`**, and at most `--max-issues` (default 20) per export; the next export files the rest. `--min-severity`, `--only <ids>` and `--include-worth-a-look` choose what goes.
|
|
500
|
+
- **Inert issues.** Titles, descriptions, steps and evidence come from the run and the app's pages, so no `@mention`, link, `#123` reference, HTML or Markdown in them does anything.
|
|
501
|
+
- **Severity** becomes a label on GitHub and a priority in Jira; `--severity-map` renames them or turns them off. **Screenshots** from a recorded run are attached in Jira; GitHub's API takes no uploads, so a GitHub issue names the frames in the run's `.scenescout/` folder.
|
|
502
|
+
- **Credentials** are never printed. Every request has a timeout, a short rate limit is waited out, failed reads are retried with backoff, and redirects are refused. Jira's search can take a little while to show a new issue, so leave a few minutes between two exports to the same Jira project.
|
|
503
|
+
|
|
504
|
+
[The guide](docs/guide/Ways-to-use-it.md#filing-findings-as-issues) has the details and a GitHub Actions step.
|
|
505
|
+
|
|
462
506
|
---
|
|
463
507
|
|
|
464
508
|
## 🩺 Troubleshooting
|
|
@@ -471,7 +515,7 @@ Run `npx -y scenescout doctor` first — it checks every setup item below (every
|
|
|
471
515
|
| The `scout_*` tools don't appear | The MCP server isn't registered, or points at an old path. `npx -y scenescout install` re-registers it; `claude mcp list` should show `scenescout` as connected. |
|
|
472
516
|
| *"Executable not found in $PATH"* | The server was registered with a bare `node`. `npx -y scenescout install` registers an absolute path. |
|
|
473
517
|
| Installed as a plugin, and the tools fail with *"Executable not found in $PATH: npx"* | A plugin starts the server with a bare `npx`, which Claude Code can only find if it was launched from an environment that has Node on its `PATH`. Under nvm or fnm that means starting Claude Code from a terminal, not from a dock or launcher. Or use `npx -y scenescout install` instead, which registers the absolute path of `npx`. |
|
|
474
|
-
| *"… build has not been downloaded yet"* on attach | The browser
|
|
518
|
+
| *"… build has not been downloaded yet"* on attach | The attach downloads a missing browser itself, once, except in CI or with `SCENESCOUT_BROWSER_DOWNLOAD=off`; there, or when that download failed, it names the command to run. Run the command the message names, for example `npx -y scenescout install --browser-only --browsers firefox`. On Linux, system libraries may be missing too: `npx playwright install --with-deps chromium`. |
|
|
475
519
|
| Tools broke after moving the folder or changing node version | The registration stores absolute paths. `npx -y scenescout install` refreshes them. |
|
|
476
520
|
| Attach fails or every route lands on the login page | Your app isn't running at `--url`, or the `--role` session has expired. For a saved login, run `scenescout login <url> --role <name>` again; for a storage-state file, regenerate it the way your project's Playwright setup does. |
|
|
477
521
|
|
|
@@ -515,6 +559,8 @@ npx -y scenescout install --browser-only --browsers firefox,webkit # add two m
|
|
|
515
559
|
|
|
516
560
|
Sizes vary by platform. The builds go to Playwright's shared cache, so a build another tool already fetched is not downloaded again.
|
|
517
561
|
|
|
562
|
+
The first attach that needs a build which is not on disk downloads it itself, once, and says so ("Getting the test browser ready"). It does not in CI unless `SCENESCOUT_BROWSER_DOWNLOAD=on` is set, and never with `SCENESCOUT_BROWSER_DOWNLOAD=off`, for a machine where nothing may be downloaded.
|
|
563
|
+
|
|
518
564
|
To drive another browser, pass `browser` when attaching (`scout_attach { browser: "firefox" }`), or set `SCENESCOUT_BROWSER=webkit` in the server's environment to change the default. `scenescout doctor` checks the browser named by that variable in the shell it runs from, so check another one with `SCENESCOUT_BROWSER=webkit scenescout doctor`. Two things differ outside Chromium:
|
|
519
565
|
|
|
520
566
|
- **Service workers are not allowed to register** in Firefox and WebKit. The write policy works by intercepting requests, and only Chromium lets a request issued by a service worker be intercepted. An app that depends on its worker may behave differently there.
|
|
@@ -671,10 +717,11 @@ npx -y scenescout login <url> --role admin # sign in once in a visible browser
|
|
|
671
717
|
src/
|
|
672
718
|
mcp-server.ts the 29 tools + per-session dispatch
|
|
673
719
|
scan.ts project discovery (framework, routes, auth)
|
|
674
|
-
cli.ts scan · serve · install · doctor · check · ci · login · status · watch
|
|
720
|
+
cli.ts scan · serve · install · doctor · check · ci · login · export · status · watch
|
|
675
721
|
check-run.ts drives a check: attach, crawl every route, collect what was measured
|
|
676
722
|
ci-run.ts drives a CI run: the MCP server as a child, the model's API, the agent loop
|
|
677
|
-
login-run.ts drives `scenescout login
|
|
723
|
+
login-run.ts drives `scenescout login` and scout_login: a visible browser that saves the role's profile once the sign-in is seen to finish (or on Enter); or --script, headless from the environment
|
|
724
|
+
export-run.ts drives `scenescout export`: reads the findings, asks GitHub or Jira what is filed, files the rest
|
|
678
725
|
installer.ts setup logic (skill link, MCP registration, diagnostics)
|
|
679
726
|
engine/
|
|
680
727
|
browser.ts the engine class: attach, snapshot, actions, crawl, plans
|
|
@@ -698,10 +745,14 @@ src/
|
|
|
698
745
|
profiles.ts saved sign-ins: role names, where a profile lives, owner-only files, attach by role, sessionStorage restore
|
|
699
746
|
refresh.ts the refresh broker: which values are a role's refresh tokens, the lock beside the profile, swapping a spent token
|
|
700
747
|
scripted-login.ts a CI sign-in: env and flags, TOTP (RFC 6238) or a fixed code, which field is which, redaction
|
|
748
|
+
signed-in.ts when a person's sign-in in the window has finished: back on the app, a new session, past any SSO round trip
|
|
701
749
|
expiry.ts how long a saved sign-in lasts: cookie dates and JWT exp, checked before lanes start
|
|
702
750
|
report.ts the gap ledger + report generation
|
|
703
751
|
check.ts the check's rules, gate, report and SARIF
|
|
752
|
+
baseline.ts visual baselines: targets.json, where each picture is kept, when one is met
|
|
753
|
+
sarif.ts which repository file a SARIF result points at, so code scanning keeps it
|
|
704
754
|
ci.ts a CI run's options, provider choice, caps, key redaction, tools and files
|
|
755
|
+
export.ts which findings an export files, the inert issue it writes, the marker that dedups it
|
|
705
756
|
provider.ts the Anthropic and OpenAI message shapes, and retries
|
|
706
757
|
replay.ts the run as one page: steps, tasks, frames under each finding
|
|
707
758
|
… collector · dispatch · fixtures · authloss · reaper
|
|
@@ -738,6 +789,7 @@ The load-bearing choices are recorded as ADRs — read the relevant one before c
|
|
|
738
789
|
- [11 · A gate is deterministic, and fails only on what it can prove](docs/adr/0011-a-gate-is-deterministic-and-fails-only-on-what-it-can-prove.md)
|
|
739
790
|
- [12 · A check replays saved flows and re-tests open findings, within settings whose defaults do the least harm](docs/adr/0012-a-check-replays-saved-flows-and-reports-re-tests.md)
|
|
740
791
|
- [13 · What depends on a project's convention is the project's to decide](docs/adr/0013-a-convention-is-the-projects-to-decide.md)
|
|
792
|
+
- [21 · Update the docs in the same pull request, unless the change has no user-facing surface](docs/adr/0021-update-the-docs-in-the-pull-request.md)
|
|
741
793
|
|
|
742
794
|
---
|
|
743
795
|
|
package/dist/browsers.js
CHANGED
|
@@ -7,6 +7,7 @@
|
|
|
7
7
|
*/
|
|
8
8
|
import fs from "node:fs";
|
|
9
9
|
import path from "node:path";
|
|
10
|
+
import { isCiEnv } from "./engine/capture.js";
|
|
10
11
|
/** A browser the engine can launch. */
|
|
11
12
|
export const BROWSER_ENGINES = ["chromium", "firefox", "webkit"];
|
|
12
13
|
/**
|
|
@@ -283,3 +284,30 @@ export function defaultAttachNote(opts) {
|
|
|
283
284
|
return (`an attach drives ${opts.defaultEngine} unless told otherwise, and that build is not installed. ` +
|
|
284
285
|
`Pass browser: "${engine}" when attaching, or set ${DEFAULT_ENGINE_ENV}=${engine} in the MCP server's environment.`);
|
|
285
286
|
}
|
|
287
|
+
/** Environment variable deciding whether an attach downloads a missing browser build itself. */
|
|
288
|
+
export const BROWSER_DOWNLOAD_ENV = "SCENESCOUT_BROWSER_DOWNLOAD";
|
|
289
|
+
/** `auto` downloads except in CI; `on` downloads in CI too; `off` never downloads, for a machine where nothing may be fetched. */
|
|
290
|
+
export const BROWSER_DOWNLOAD_SETTINGS = ["auto", "on", "off"];
|
|
291
|
+
export const DEFAULT_BROWSER_DOWNLOAD = "auto";
|
|
292
|
+
/**
|
|
293
|
+
* What an attach does when the build it needs is not on disk: download it
|
|
294
|
+
* itself, refuse because the setting says never, or tell CI to keep its
|
|
295
|
+
* explicit install step (a CI job downloads only when it asks to, so a
|
|
296
|
+
* pipeline's browser is never fetched behind its back). An unknown setting
|
|
297
|
+
* refuses the attach rather than being read as a default.
|
|
298
|
+
*/
|
|
299
|
+
export function browserDownloadDecision(env) {
|
|
300
|
+
const raw = env[BROWSER_DOWNLOAD_ENV]?.trim().toLowerCase() || DEFAULT_BROWSER_DOWNLOAD;
|
|
301
|
+
if (!BROWSER_DOWNLOAD_SETTINGS.includes(raw)) {
|
|
302
|
+
throw new Error(`${BROWSER_DOWNLOAD_ENV}="${env[BROWSER_DOWNLOAD_ENV]}" is not a setting SceneScout knows. Use one of: ${BROWSER_DOWNLOAD_SETTINGS.join(", ")}.`);
|
|
303
|
+
}
|
|
304
|
+
if (raw === "off")
|
|
305
|
+
return "refuse";
|
|
306
|
+
if (raw === "on")
|
|
307
|
+
return "download";
|
|
308
|
+
return isCiEnv(env) ? "tell" : "download";
|
|
309
|
+
}
|
|
310
|
+
/** The plain-words line shown while an attach downloads the build it needs. */
|
|
311
|
+
export function attachDownloadLine(target) {
|
|
312
|
+
return `Getting the test browser ready — a one-time download of about ${APPROX_DISK_MB[target]} MB (${target}). The test carries on once it is done.`;
|
|
313
|
+
}
|