scenescout 3.17.0 → 3.19.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,42 @@
1
1
  # scenescout
2
2
 
3
+ ## 3.19.0
4
+
5
+ ### Minor Changes
6
+
7
+ - fb7a3e5: An optional Claude Code mod, `scenescout-mod`, is listed in the plugin marketplace beside `scenescout`. In the Claude Code CLI and the desktop Code tab, `/scenescout-pane` opens a pane that refreshes every 2.5 seconds from the server's `scout_status_poll` tool: each session with its task and objective, open findings by severity, coverage, and a link to the live view. Its `lane_model` setting, empty by default, sets the model SceneScout lane agents start on. The `scenescout` plugin, skill and server are unchanged and work the same without it.
8
+ - 1a2db2f: `scenescout export --from <file>` exports a `check.json` from `scenescout check` or a `ci.json` from `scenescout ci`, so the unattended runs can file GitHub or Jira issues without an interactive session's memory. Their issues get the same body, label and marker as memory findings: a check issue is keyed on its fingerprint and a ci finding keeps its memory id, so a re-export files nothing twice. The project's memory stays the default source.
9
+ - 498dbdf: A saved flow can take a `type` or `select` value from the environment with `${env:NAME}`, so a one-time code or a password stays in the CI's secret store rather than in the flow file. A variable that is not set stops the check before it starts, naming the flow and the variable, and every substituted value of four characters or more is masked as `[$NAME]` in everything the check writes and prints.
10
+ - ea8fbfa: Saved flows can assert an element's state with a new step, `expect-element`: a target and a state of `visible`, `hidden`, `enabled`, `disabled`, `checked` or `unchecked`. `hidden` also holds when nothing matches, so a flow can say a dialog closed or a receipt was never shown, and a failure says what is true instead (`testid=receipt is visible, expected hidden`).
11
+ - c07941d: Saved flows have a `repeat` step: run a few click, type, select or press steps until an `expect-text`, `expect-element` or `expect-url` step holds, at most `max` times. It is for gates a fixed script cannot walk, such as paging through a document until its Continue button is enabled. The condition is checked first, so a page already in that state runs nothing, and a failure names what never held and how many rounds were tried.
12
+ - 54257fc: A saved flow can name the `role` it runs as. `scenescout check` replays it in a browser of its own, signed in with the profile `scenescout login --role <name>` saved in the project, while flows naming no role keep the check's own session. Flows run in file-name order, so a journey that passes between people (one submits, another approves) is a sequence of flows. A role with no saved profile stops the check before it starts, with exit 2 and the command that saves one.
13
+ - d439d28: Saved flows have an `upload` step: attach a small valid file generated on the spot (`pdf`, `png`, `txt`, `csv` or `json`, or the kind the input's `accept` asks for) to a file input, the control that opens its chooser, or the page's only file input. A journey that starts from an attached file can now be saved and replayed by `scenescout check`. Submitting the upload is a write, so it reaches the server only under `--flow-writes allow`.
14
+
15
+ ### Patch Changes
16
+
17
+ - 36c6e11: `scenescout check` no longer passes when the only page it measured was a start page showing at most one control and linking nowhere, which is what an app measured before it finished drawing looks like. Without `--paths`, that check now exits 2 saying only the start page was measured, rather than going green having checked nothing.
18
+ - 87fa5bf: The contrast rule no longer reports text whose own colour is fully transparent, such as the selectable text layer a PDF viewer lays over the rendered page. That text is not painted, so its 1.00:1 ratio is not what anyone sees, and how many of those spans were measured depended on whether the viewer had finished loading. Translucent text that is painted, a faint watermark say, is still measured.
19
+ - 8857bab: `scenescout export` no longer files a finding twice when an export runs straight after another, or after a create whose result was uncertain, while the tracker's issue listing has not caught up. Each filed issue is recorded in `.scenescout/exported.json` and read back by number on the next export; GitHub's newest issues are also read without the label filter; and in Jira a finding whose create may have been made is held back for 15 minutes unless its issue is found.
20
+ - 82f775c: A saved flow step on a control that is shown but stays disabled (or, for `type`, read-only) until the action limit runs out now fails saying so first, `testid=save is visible but disabled after 5s`, followed by the limit hint. It used to report only an action timeout, which never said the control was there but disabled.
21
+ - 830b663: `scenescout check --baseline` now takes each picture again, a frame later, until two in a row are the same, so a baseline is no longer of a moment the page was still drawing: the last frame of an animation it had just stopped, or a script still filling something in after the load. A page that never holds still within the action limit keeps its last picture, and the report says so on that target's line.
22
+
23
+ ## 3.18.0
24
+
25
+ ### Minor Changes
26
+
27
+ - 1a6ce12: Ask the start-of-run questions as one form where the client can show one. A new `scout_intake` tool checks whether the client declared MCP elicitation in form mode and, if it did, asks the address, whether and how to sign in, what to check and whether the site holds real data in a single form, then returns the `scout_login` and `scout_attach` calls the answers choose. With no form support, or when the person declines or closes the form, it returns the questions for the agent to ask in chat, as before. The form never asks for a password or a code: signing in stays in the window `scout_login` opens. The skill and the `explore` prompt call `scout_intake` before setup.
28
+
29
+ ### Patch Changes
30
+
31
+ - 96bee5c: `scenescout check` and the first look no longer report a feed, a plain-text file, XML, JSON, a PDF or an image as a dead end. A route's response content type now decides whether it is a page: one served as anything but `text/html` or `application/xhtml+xml` is listed under "Not pages" in the report and as `resources` in `check.json`, is not checked against the page rules, and does not count towards the routes checked or `--max-routes`. One that answers 4xx or 5xx is still reported as the route's error.
32
+ - 952fbf3: A check or first look with no signed-in session no longer reports each route that redirects to sign-in as an `auth-redirect` issue. It lists them once as a coverage gap ("N routes need sign-in; give a role to cover them"), under "Needs sign-in" in the report and as `needsSignIn` in `check.json`. With `--storage-state`, a redirect to sign-in is still reported as a lost session.
33
+ - 6073a0d: A link or button with no text is now named by a descendant's `aria-label` (an icon element inside the link) and by an image's `title` when its alt text is empty, as the browser names it, so `unnamed-control` no longer reports those controls. Coverage recorded for them under their earlier keys carries over.
34
+ - cd0639b: Text drawn over a positioned `<img>`, `<picture>`, `<video>` or `<canvas>` is no longer reported as a contrast failure measured against the page background (often as 1.00:1). Like text over a CSS background image, it has no single background colour, so the design audit and `scenescout check` leave it unmeasured.
35
+ - 680e4f5: The focus-indicator rule no longer reports an iframe as a control with no visible focus indicator. A Tab onto a frame moves focus into the frame's document, so the design audit and `scenescout check` now follow focus into a same-origin frame and measure the control focused there, reported by its own name. A press that leaves focus inside another site's frame is skipped.
36
+ - fe16d4c: A page that locks scrolling behind a modal opened inside a full-viewport frame is no longer reported as a leaked scroll lock. The overlay probe now counts a visible iframe that covers the viewport, is pinned itself or through a fixed ancestor, and is neither faded out nor click-through as an open overlay, provided that, when the frame is same-origin, a dialog or a dimming backdrop is showing inside it. A frame left mounted after its modal closed does not justify the lock, so that page is still reported at high severity.
37
+ - 96b74aa: The design audit's `image-aspect` rule no longer reports an image that keeps its proportions through `object-fit: cover`, `contain`, `scale-down` or `none`, whether set inline or by a class. Only an image stretched to its box under `fill`, the default, is reported as distorted.
38
+ - 7fc6013: The gap ledger's list of pages whose POST observe refused now names the page that sent the POST. A page whose script posted as it loaded could be listed as the page the session came from, because the browser had not yet reported the new page; the request's Referer now decides, with the session's page as the fallback.
39
+
3
40
  ## 3.17.0
4
41
 
5
42
  ### Minor Changes
package/README.md CHANGED
@@ -154,6 +154,8 @@ Either way it downloads the browser and registers the server with the client you
154
154
 
155
155
  Then start a new chat to use SceneScout; the test browser downloads on first use (to have it ready beforehand: `npx -y scenescout install --browser-only`). The command becomes `/scenescout:scenescout`. A plugin's skill comes from this repository and its server from the latest npm release, so right after a release lands here the two can differ for a short while; `/plugin marketplace update scenescout-marketplace` brings the skill up to date.
156
156
 
157
+ An optional second plugin, `scenescout-mod@scenescout-marketplace`, adds a run pane (`/scenescout-pane`) and a setting for the model lane agents run on, in the Claude Code CLI and the desktop Code tab. It is a mod: unsandboxed JavaScript that runs inside Claude Code, so it is opt-in. [The Claude Code mod](docs/guide/Ways-to-use-it.md#the-claude-code-mod).
158
+
157
159
  **Using Claude Desktop?** Install the extension: download `scenescout-X.Y.Z.mcpb` from the [latest release](https://github.com/brunoboto96/SceneScout/releases/latest) and open it (or Settings > Extensions > Advanced settings > Install Extension). It works as soon as it is installed, with no terminal step: the test browser downloads on first use. Start a new chat and ask *"Use SceneScout to test http://localhost:3000"*. [More in the guide](docs/guide/Start-here.md#as-a-claude-desktop-extension).
158
160
 
159
161
  **A client that is not in that list?** [Add the server to its config by hand](#-other-mcp-clients); the test browser downloads on first use.
@@ -205,7 +207,7 @@ The agent scans the project (if there is one), attaches read-only, explores, and
205
207
 
206
208
  **Common flags** — `--level minimal|medium|extensive` · `--url <app>` · `--role <name\|path>` (who to explore as: a login saved with `scenescout login`, a storage state found by the scan, or a path to a Playwright storage-state JSON) · `--focus <text>` (a ticket or a sentence to check) · `--observe` / `--read-only` / `--safe-write` / `--allow-destructive`.
207
209
 
208
- **No flags at all** (`/scenescout` on its own) and the agent asks four plain questions instead: the address, whether and how you sign in, what to check (tickets or a description), and whether the site holds real data. Real data, or not being sure, means nothing but `GET` requests leave the page; you are never asked to pick a mode. Any flag skips the questions. See [Plain questions instead of flags](docs/guide/Ways-to-use-it.md#plain-questions-instead-of-flags).
210
+ **No flags at all** (`/scenescout` on its own) and the agent asks four plain questions instead: the address, whether and how you sign in, what to check (tickets or a description), and whether the site holds real data. Where your client can show a form, `scout_intake` asks them as one form, never for a password; otherwise the agent asks in chat. Real data, or not being sure, means nothing but `GET` requests leave the page; you are never asked to pick a mode. Any flag skips the questions. See [Plain questions instead of flags](docs/guide/Ways-to-use-it.md#plain-questions-instead-of-flags).
209
211
 
210
212
  ### 🔑 Signing in as a role
211
213
 
@@ -329,11 +331,11 @@ Snapshots are cheap: re-snapshotting a route returns only *what changed*, with s
329
331
 
330
332
  ## 🧰 The toolbox
331
333
 
332
- 34 deterministic tools. The agent picks; you rarely call these by hand.
334
+ 35 deterministic tools. The agent picks; you rarely call these by hand.
333
335
 
334
336
  | Phase | Tools | What they do |
335
337
  |---|---|---|
336
- | **Set up** | `scout_playbook` `scout_scan` `scout_attach` `scout_session` | Hand the testing method to an agent that has no skill loaded; discover routes; launch a browser in a write-mode; keep several authenticated roles alive at once |
338
+ | **Set up** | `scout_playbook` `scout_intake` `scout_scan` `scout_attach` `scout_session` | Hand the testing method to an agent that has no skill loaded; ask the start-of-run questions as one form where the client can show one; discover routes; launch a browser in a write-mode; keep several authenticated roles alive at once |
337
339
  | **Explore** | `scout_crawl` `scout_coverage` | Sweep every route in one call; ask what's still untested |
338
340
  | **Look** | `scout_snapshot` `scout_hover` `scout_screenshot` `scout_capture` | Read the structured scene (diffed); reveal tooltips/hover cards; capture pixels only when needed; save one element as a PNG to show someone |
339
341
  | **Ask the server** | `scout_request` `scout_network` | Call the app's own API as this session, with the UI bypassed — the check that turns a hidden button into a proven refusal; list the requests the page itself made since it loaded |
@@ -457,7 +459,7 @@ The defaults are what an unconfigured check does, for a first try or an AI agent
457
459
 
458
460
  `scenescout check --help` lists every option. Why the defaults are what they are: [ADR 11](docs/adr/0011-a-gate-is-deterministic-and-fails-only-on-what-it-can-prove.md).
459
461
 
460
- It also replays the flows saved in `.scenescout/flows/*.json`, with no model: the steps `scout_run_plan` takes (navigate, click, type, select, press) plus `expect-text`, `expect-url` and `expect-request`. A flow whose step breaks fails the gate, naming the flow and the step. And it re-tests the open findings earlier runs left in the project's memory that a page load can reproduce, reporting each as still reproducing or possibly fixed; by default a finding filed high that still reproduces fails the gate. [docs/ci.md](docs/ci.md#saved-flows) has the flow format; [ADR 12](docs/adr/0012-a-check-replays-saved-flows-and-reports-re-tests.md) says why it works this way.
462
+ It also replays the flows saved in `.scenescout/flows/*.json`, with no model: the steps `scout_run_plan` takes (navigate, click, type, select, press, and upload, which attaches a generated file) plus `expect-text`, `expect-element` (a control is visible, hidden, enabled, disabled, checked or unchecked), `expect-url` and `expect-request`, and `repeat` to run actions until a condition holds (page through a document until Continue is enabled). A flow can name the `role` it runs as, signed in with a profile `scenescout login --role` saved, so a journey that passes between people (one submits, another approves) is a sequence of flows. Values can come from the environment, `${env:NAME}`, so a code or a password stays in the CI's secret store and is masked in everything the check writes. A flow whose step breaks fails the gate, naming the flow and the step. And it re-tests the open findings earlier runs left in the project's memory that a page load can reproduce, reporting each as still reproducing or possibly fixed; by default a finding filed high that still reproduces fails the gate. [docs/ci.md](docs/ci.md#saved-flows) has the flow format; [ADR 12](docs/adr/0012-a-check-replays-saved-flows-and-reports-re-tests.md) says why it works this way.
461
463
 
462
464
  With `--baseline compare` it also holds pages and elements to approved pictures: list them in a `targets.json`, take the baselines once with `--baseline update`, and a later check that finds one changed past `--baseline-threshold` fails the gate with the share of pixels changed and a diff picture beside the report. Baselines are kept per browser, in a git-ignored folder unless `--baselines` names one the project commits, and only `--baseline update` ever writes one. [The guide](docs/guide/Ways-to-use-it.md#visual-baselines) has the details; [ADR 19](docs/adr/0019-a-visual-baseline-changes-only-when-asked.md) says why.
463
465
 
@@ -761,6 +763,7 @@ test-app/ fixtures for the real-browser smoke tests
761
763
  skills/scenescout/ the testing method (SKILL.md): a skill in Claude Code, served by the server everywhere else
762
764
  docs/how-it-works.md what happens at each stage, in diagrams
763
765
  docs/benchmark.md measuring whether a change made runs better
766
+ docs/validation.md scorecards from runs against public open-source apps
764
767
  docs/adr/ why it's built this way
765
768
  ```
766
769
 
@@ -774,6 +777,8 @@ docs/adr/ why it's built this way
774
777
 
775
778
  **[Measuring whether a change helped](docs/benchmark.md)** — the demo app's answer key, the scorecard (recall, precision, judged-not-filed, severity, calibration), and the results log of every run, including what did not help.
776
779
 
780
+ **[Validation on public open-source apps](docs/validation.md)** — runs against three well-known open-source web apps, with a scorecard for each: issues by severity, how many were real and how many were false positives, and the engine problems the runs exposed.
781
+
777
782
  The load-bearing choices are recorded as ADRs — read the relevant one before changing a rule it covers:
778
783
 
779
784
  - [1 · Completion is an enforced contract, not a claim](docs/adr/0001-completion-is-a-contract-not-a-vibe.md)
package/dist/check-run.js CHANGED
@@ -11,11 +11,13 @@ import os from "node:os";
11
11
  import path from "node:path";
12
12
  import { BASELINE_CAPTURE, baselineFiles, baselineMeta, defaultBaselinesDir, elementTarget, isVisualPicture, judgeBaseline, missingTargetsMessage, parseBaselineTargets, readStoredBaseline, TARGETS_FILE, VISUAL_DIRNAME, visualFiles, } from "./engine/baseline.js";
13
13
  import { BrowserEngine } from "./engine/browser.js";
14
- import { checkFindings, redactBaselineRun, MAX_DISCOVERY_ROUNDS, redactFlowRuns, redactRoute, redactRoutes, settingsOf, unmeasuredReason, withoutOwnResponse, } from "./engine/check.js";
15
- import { loadFlows, resolveFlowsDir } from "./engine/flow.js";
14
+ import { checkFindings, redactBaselineRun, MAX_DISCOVERY_ROUNDS, redactFlowRuns, redactRoute, redactRoutes, settingsOf, splitResources, unmeasuredReason, withoutOwnResponse, } from "./engine/check.js";
15
+ import { flowRoles, loadFlows, maskEnvValues, resolveFlowEnv, resolveFlowsDir } from "./engine/flow.js";
16
+ import { resolveAttachAuth } from "./engine/profiles.js";
16
17
  import { MemoryStore, MEMORY_DIRNAME, writeSelfIgnore } from "./engine/memory.js";
17
18
  import { decodePng, encodePng } from "./engine/png.js";
18
19
  import { firstLineOf } from "./engine/limits.js";
20
+ import { isNonPageResource } from "./engine/crawl.js";
19
21
  import { checkRetestPlan, retestResults, wellFormedFindings } from "./engine/verify.js";
20
22
  /** The baselines folder a check uses: the one --baselines names, else the project's own. */
21
23
  export function baselinesDirOf(options) {
@@ -80,7 +82,23 @@ export async function runCheck(options, log = () => { }, inputs = { flows: [], f
80
82
  // Also on an exit the finally below never reaches, such as Ctrl+C, which the browser's driver answers with process.exit.
81
83
  const removeScratch = () => fs.rmSync(scratch, { recursive: true, force: true });
82
84
  process.once("exit", removeScratch);
85
+ // A flow that names a role whose sign-in is not saved would only fail after the crawl: refuse before anything starts,
86
+ // naming the command that saves it.
87
+ for (const role of flowRoles(inputs.flows))
88
+ resolveAttachAuth({ projectDir: options.projectDir, url: options.url, role });
89
+ // ${env:NAME} in a flow's values, from the environment; an unset name stops the check before anything starts.
90
+ const envValues = new Map();
91
+ const runFlows = inputs.flows.map((flow) => {
92
+ const resolved = resolveFlowEnv(flow, process.env);
93
+ if (!resolved.ok)
94
+ throw new Error(`flow "${flow.name}" (${flow.file}) takes ${resolved.missing.map((n) => `\${env:${n}}`).join(", ")} from the environment, which is not set`);
95
+ for (const [name, value] of resolved.values)
96
+ envValues.set(name, value);
97
+ return resolved.flow;
98
+ });
83
99
  const engine = new BrowserEngine();
100
+ /** One browser per role a flow runs as, each attached once, on first use, and closed with the check's own. */
101
+ const roleEngines = new Map();
84
102
  const start = new URL(options.url);
85
103
  // From the start of the run, browser launch included: the budget is wall-clock time a person waits.
86
104
  const deadline = options.timeBudgetMs !== undefined ? Date.now() + options.timeBudgetMs : undefined;
@@ -107,6 +125,8 @@ export async function runCheck(options, log = () => { }, inputs = { flows: [], f
107
125
  if (authFailed)
108
126
  throw new Error(authFailed.replace(/ Continuing now tests a logged-out app\.$/, "").replace(/re-attach/, "run the check again"));
109
127
  const routes = [];
128
+ // A feed or a file the crawl followed is not a page, so it does not use up --max-routes.
129
+ const pagesSoFar = () => routes.filter((r) => !isNonPageResource(r)).length;
110
130
  const crawl = async (paths, limit, deadline) => {
111
131
  await engine.crawl(paths, { inspect: true, limit, deadline });
112
132
  routes.push(...engine.lastCrawlHealth);
@@ -118,15 +138,15 @@ export async function runCheck(options, log = () => { }, inputs = { flows: [], f
118
138
  // The start page first, whatever else is known: it is the one route the user named. The
119
139
  // time budget does not apply to it, so a run always measures at least the page it was given.
120
140
  await crawl([`${start.pathname}${start.search}${start.hash}`], 1);
121
- for (let round = 0; round < MAX_DISCOVERY_ROUNDS && routes.length < options.maxRoutes; round++) {
141
+ for (let round = 0; round < MAX_DISCOVERY_ROUNDS && pagesSoFar() < options.maxRoutes; round++) {
122
142
  if (engine.crawlableRoutes().length === 0 || pastDeadline())
123
143
  break;
124
- await crawl(undefined, options.maxRoutes - routes.length, deadline);
125
- log(` ${routes.length} route(s) checked`);
144
+ await crawl(undefined, options.maxRoutes - pagesSoFar(), deadline);
145
+ log(` ${pagesSoFar()} route(s) checked`);
126
146
  }
127
147
  }
128
148
  // Out of time with routes still to visit; a cap of routes reached first is --max-routes's to report.
129
- const timeLimitReached = pastDeadline() && routes.length < options.maxRoutes && engine.crawlableRoutes().length > 0;
149
+ const timeLimitReached = pastDeadline() && pagesSoFar() < options.maxRoutes && engine.crawlableRoutes().length > 0;
130
150
  // Pages of open findings the crawl did not load exactly: loaded now, so each re-test has its own measurement.
131
151
  // Kept apart from `routes`: they are measured only to re-test, never checked against the page rules, and do not count
132
152
  // towards --max-routes. With --paths the check stays on the paths it was given.
@@ -144,15 +164,45 @@ export async function runCheck(options, log = () => { }, inputs = { flows: [], f
144
164
  // reached only sign-in pages: the check has no verdict then, and an update would write sign-in pages as baselines.
145
165
  if (options.baseline !== "off" && !inputs.baselineTargets)
146
166
  throw new Error(`--baseline ${options.baseline} was given no targets: read them with readCheckInputs`);
147
- const baselines = options.baseline !== "off" && inputs.baselineTargets && unmeasuredReason(routes, !options.paths) === null
167
+ const baselines = options.baseline !== "off" && inputs.baselineTargets && unmeasuredReason(routes, !options.paths, options.paths ? [] : engine.crawlableRoutes()) === null
148
168
  ? await takeBaselines(engine, options, options.baseline, inputs.baselineTargets, log)
149
169
  : null;
150
170
  const flowRuns = [];
151
- for (const flow of inputs.flows) {
171
+ // A flow that names a role runs in that role's own browser, so the check's session (and what its crawl found) is
172
+ // left as it was.
173
+ const engineFor = async (role) => {
174
+ if (role === undefined)
175
+ return engine;
176
+ const known = roleEngines.get(role);
177
+ if (known)
178
+ return known;
179
+ const roleEngine = new BrowserEngine();
180
+ roleEngines.set(role, roleEngine);
181
+ const roleDir = path.join(scratch, `role-${role}`);
182
+ fs.mkdirSync(roleDir, { recursive: true });
183
+ const attachedAs = await roleEngine.attach({
184
+ url: start.origin,
185
+ projectDir: options.projectDir,
186
+ memoryStore: new MemoryStore(roleDir),
187
+ mode: options.mode,
188
+ role,
189
+ ...(options.browser ? { browser: options.browser } : {}),
190
+ actionTimeoutMs: options.actionTimeoutMs,
191
+ navTimeoutMs: options.navTimeoutMs,
192
+ objective: `Deterministic check: replay the saved flows that run as ${role}`,
193
+ task: `Replaying flows as ${role}`,
194
+ });
195
+ const lost = attachedAs.split("\n").find((line) => line.startsWith("⚠ AUTH FAILED"));
196
+ if (lost)
197
+ throw new Error(`the saved sign-in for role "${role}" no longer signs in: ${lost.replace(/ Continuing now tests a logged-out app\.$/, "")}`);
198
+ return roleEngine;
199
+ };
200
+ for (const flow of runFlows) {
201
+ const runner = await engineFor(flow.role);
152
202
  // --flow-writes never: observe's rule, whatever --mode lets the crawl do.
153
- const replay = await engine.replayFlow(flow.steps, options.flowWrites === "never" ? "observe" : options.mode);
154
- flowRuns.push({ name: flow.name, file: flow.file, steps: flow.steps.length, ...replay });
155
- log(` flow ${flow.name}: ${replay.outcome.status}${replay.outcome.status === "passed" ? "" : ` at step ${replay.outcome.step}`}`);
203
+ const replay = await runner.replayFlow(flow.steps, options.flowWrites === "never" ? "observe" : options.mode);
204
+ flowRuns.push({ name: flow.name, file: flow.file, steps: flow.steps.length, ...(flow.role === undefined ? {} : { role: flow.role }), ...replay });
205
+ log(` flow ${flow.name}${flow.role === undefined ? "" : ` (as ${flow.role})`}: ${replay.outcome.status}${replay.outcome.status === "passed" ? "" : ` at step ${replay.outcome.step}`}`);
156
206
  // --on-refused-step stop: nothing after a refused step runs, and the check ends without a verdict.
157
207
  if (replay.outcome.status === "refused" && options.onRefusedStep === "stop")
158
208
  break;
@@ -172,18 +222,22 @@ export async function runCheck(options, log = () => { }, inputs = { flows: [], f
172
222
  }))),
173
223
  }
174
224
  : null;
175
- const measured = redactRoutes(routes.map(withoutOwnResponse));
225
+ const { pages: measured, resources } = splitResources(redactRoutes(routes.map(withoutOwnResponse)));
176
226
  const flows = redactFlowRuns(flowRuns);
177
227
  const pictured = baselines ? redactBaselineRun(baselines) : null;
178
- const { issues, worthALook } = checkFindings(measured, start.origin, options.ignore, flows, pictured, options.ignorePaths);
179
- return {
228
+ // Signed in only when given a session: with none, a route that sends the browser to sign-in needs one, and is not a lost session.
229
+ const signedIn = options.storageStatePath !== undefined;
230
+ const { issues, worthALook, needsSignIn } = checkFindings(measured, start.origin, options.ignore, flows, pictured, options.ignorePaths, signedIn);
231
+ const result = {
180
232
  url: redactRoute(options.url),
181
233
  generatedAt: new Date().toISOString(),
182
234
  mode: options.mode,
183
235
  failOn: options.failOn,
184
236
  routes: measured,
237
+ resources,
185
238
  issues,
186
239
  worthALook,
240
+ ...(signedIn ? {} : { needsSignIn }),
187
241
  // Routes that failed to load are issues already; "not visited" is only what --max-routes (or the time budget) left out.
188
242
  unvisited: options.paths ? [] : engine.crawlableRoutes().map(redactRoute),
189
243
  ...(options.timeBudgetMs !== undefined ? { timeBudget: { ms: options.timeBudgetMs, reached: timeLimitReached } } : {}),
@@ -195,8 +249,12 @@ export async function runCheck(options, log = () => { }, inputs = { flows: [], f
195
249
  settings: settingsOf(options),
196
250
  baselines: pictured,
197
251
  };
252
+ // A value taken from the environment (a code, a password) is never written, wherever the page echoed it.
253
+ return envValues.size > 0 ? JSON.parse(maskEnvValues(JSON.stringify(result), envValues)) : result;
198
254
  }
199
255
  finally {
256
+ for (const [role, roleEngine] of roleEngines)
257
+ await roleEngine.close().catch((err) => log(`closing the browser for role ${role} failed: ${err instanceof Error ? err.message : String(err)}`));
200
258
  await engine.close().catch((err) => log(`closing the browser failed: ${err instanceof Error ? err.message : String(err)}`));
201
259
  process.off("exit", removeScratch);
202
260
  removeScratch();
@@ -308,7 +366,12 @@ async function takeBaselines(engine, options, mode, targets, log) {
308
366
  stored,
309
367
  now: { capture, platform: process.platform, image },
310
368
  });
311
- const result = { ...own, ...verdict, ...(shot.cut ? { partial: shot.cut } : {}) };
369
+ const result = {
370
+ ...own,
371
+ ...verdict,
372
+ ...(shot.cut ? { partial: shot.cut } : {}),
373
+ ...(shot.unsteady ? { unsteady: shot.unsteady } : {}),
374
+ };
312
375
  if (verdict.status === "updated") {
313
376
  fs.mkdirSync(path.dirname(at(dir, files.png)), { recursive: true });
314
377
  fs.writeFileSync(at(dir, files.png), shot.png);
package/dist/cli.js CHANGED
@@ -162,8 +162,9 @@ Usage:
162
162
  from SCENESCOUT_LOGIN_<FLAG>, e.g. SCENESCOUT_LOGIN_SUCCESS_URL;
163
163
  --timeout seconds (default 60))
164
164
  scenescout export --to github|jira
165
- File the project's open findings (from .scenescout/memory.json) as issues,
166
- each once: a finding whose marker is already on an issue is skipped. A dry run
165
+ File the project's open findings (from .scenescout/memory.json, or from a
166
+ check.json or ci.json given with --from file) as issues, each once: a finding
167
+ whose marker is already on an issue is skipped. A dry run
167
168
  that lists what it would file unless --yes is given. Credentials come from the
168
169
  environment only: GH_TOKEN or GITHUB_TOKEN; JIRA_EMAIL and JIRA_API_TOKEN.
169
170
  (--repo owner/name for GitHub (GITHUB_API_URL for GitHub Enterprise Server);
@@ -181,7 +182,8 @@ Usage:
181
182
  in Jira (default severity: high… / High, Medium, Low); --labels a,b;
182
183
  --screenshots on|off (default on: the finding's picture and the run's frames,
183
184
  attached in Jira, named on GitHub);
184
- --project dir (default: here); --dry-run; --yes)
185
+ --project dir (default: here; its .scenescout folder keeps the record of
186
+ filed issues); --from check.json|ci.json; --dry-run; --yes)
185
187
  Exit code: 0 done (findings over the cap wait for the next export), 2 could not
186
188
  export (it lists what it filed before it stopped), or a screenshot was not
187
189
  attached or a ticket not linked.
@@ -587,7 +589,7 @@ async function check(args) {
587
589
  console.error(`scenescout check: could not run: ${err instanceof Error ? err.message : String(err)}`);
588
590
  process.exit(EXIT.error);
589
591
  }
590
- const unmeasured = unmeasuredReason(result.routes, !options.paths);
592
+ const unmeasured = unmeasuredReason(result.routes, !options.paths, result.unvisited);
591
593
  if (unmeasured) {
592
594
  console.error(`scenescout check: could not measure ${options.url}: ${unmeasured}`);
593
595
  process.exit(EXIT.error);
@@ -74,6 +74,49 @@ export const STOP_ANIMATIONS_SCRIPT = `(() => {
74
74
  }
75
75
  return true;
76
76
  })()`;
77
+ /**
78
+ * Page-side source that resolves once the page has drawn a frame after this
79
+ * call (two animation frames: the first starts one, the second runs once it
80
+ * is done). A page whose frames are not running at all resolves after
81
+ * NEXT_FRAME_WAIT_MS instead; the browser side bounds the wait by the same
82
+ * time, for a page too busy to run either.
83
+ */
84
+ export const NEXT_FRAME_WAIT_MS = 1000;
85
+ export const NEXT_FRAME_SCRIPT = `new Promise((resolve) => {
86
+ requestAnimationFrame(() => requestAnimationFrame(() => resolve(true)));
87
+ setTimeout(() => resolve(false), ${NEXT_FRAME_WAIT_MS});
88
+ })`;
89
+ /**
90
+ * A picture of a page that has stopped changing: pictures are taken, a frame
91
+ * apart, until two in a row are the same, and that one is kept. Stopping CSS
92
+ * animations does not make a page still. A frame drawn before the stop took
93
+ * effect, or a script that draws for a moment after the load (a count-up, an
94
+ * entrance), each make the first picture depend on how fast the machine is,
95
+ * and a baseline taken on a fast run then fails on a slow one. No new
96
+ * picture is started once `budgetMs` has passed (the one in progress still
97
+ * finishes): a page that never holds still keeps its last picture, and
98
+ * `steady` is false so the result can say so. `take` and `nextFrame` are the
99
+ * browser's (BrowserEngine.captureForBaseline); the loop is table-tested.
100
+ */
101
+ export async function steadyPicture(take, nextFrame, budgetMs, now = Date.now) {
102
+ const until = now() + budgetMs;
103
+ let last = await take();
104
+ let takes = 1;
105
+ for (;;) {
106
+ await nextFrame();
107
+ const png = await take();
108
+ takes += 1;
109
+ if (png.equals(last))
110
+ return { png, steady: true, takes };
111
+ last = png;
112
+ if (now() >= until)
113
+ return { png, steady: false, takes };
114
+ }
115
+ }
116
+ /** What a result says when its picture never held still (steadyPicture): a comparison of it may not repeat. */
117
+ export function unsteadyNote(budgetMs) {
118
+ return `the picture was still changing after about ${budgetMs / 1000}s, so the last one taken is used: something on the page keeps moving, and a comparison of it may not repeat`;
119
+ }
77
120
  const ELEMENT_HELP = `${PAGE_ELEMENT} or ${TARGET_HELP}`;
78
121
  /** What an element names: null for the page itself, else the target a saved flow would use. Undefined when it is neither. */
79
122
  export function elementTarget(element) {