scenescout 3.20.1 → 3.21.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,18 @@
1
1
  # scenescout
2
2
 
3
+ ## 3.21.0
4
+
5
+ ### Minor Changes
6
+
7
+ - 64a719b: `scenescout ci` and its GitHub Action now default to two lanes sharing 80 model turns and 3,000,000 tokens (from one loop, 40 turns and 1,500,000 tokens). On the benchmark's demo app the new defaults found 5 to 7 of 13 planted defects in three runs, against 2 to 4 for the old ones, and 4 to 5 of 10 on the held-out app, at about $0.03 a run on `gpt-6-luna` instead of about $0.011. Without `--lanes`, a run with `--show` or with `--max-turns 1` runs as one loop. `--lanes 1 --max-turns 40 --max-tokens 1500000` restores the old behaviour.
8
+ - a4531a8: `/scenescout qa check [focus]` runs a project's own recorded check from a pull-request comment, for projects with no preview deployments. Set the repository variable `SCENESCOUT_QA_CHECK_WORKFLOW` to a workflow that runs `scenescout check --record --video` (`examples/workflows/scenescout-qa-check.yml` is one): the comment workflow dispatches it on the pull request's branch, waits for it, and replies with the verdict, each journey's result, the findings by severity, the first failing step and links to the run and its artifact. No model key is involved, and pull requests from forks are refused. `/scenescout qa`, `show` and `compare` are unchanged.
9
+
10
+ ## 3.20.2
11
+
12
+ ### Patch Changes
13
+
14
+ - f8f7ad2: A recorded `scenescout check` (`--record` or `--video`) no longer overwrites a `replay.html` in the output folder that it did not write: it stops before it starts, with exit code 2 and a message naming the file. With `--video`, a video that fails to start no longer stops every later flow from being filmed.
15
+
3
16
  ## 3.20.1
4
17
 
5
18
  ### Patch Changes
package/README.md CHANGED
@@ -479,14 +479,14 @@ npx scenescout ci http://127.0.0.1:3000
479
479
 
480
480
  - **It reports and never gates.** Exit 0 when the run ran, whatever it found; exit 2 when it could not run (no key, a key the API refused, an app that never answered). Two runs of the same app find different things, so a finding is something to read, never a reason to fail a build. `scenescout check` is the gate.
481
481
  - **Providers:** the Anthropic Messages API (default model `claude-sonnet-5`) or the OpenAI Responses API (default `gpt-6-luna`), chosen by which key is set; with both set, `--provider` decides. `--model` and `--effort` (default `low`) override; `--base-url` points at another endpoint that implements the same API.
482
- - **Caps:** at most 40 model turns, 1,500,000 tokens and 20 minutes (`--max-turns`, `--max-tokens`, `--max-minutes`). The first cap reached ends the exploration; the report is still written, and says which cap ended it.
482
+ - **Caps:** at most 80 model turns, 3,000,000 tokens and 20 minutes (`--max-turns`, `--max-tokens`, `--max-minutes`), shared by two model loops that each explore their own part of the app (`--lanes`, default 2). The first cap reached ends the exploration; the report is still written, and says which cap ended it. On the benchmark's demo app a run at these defaults cost about $0.03 on `gpt-6-luna` ([docs/benchmark.md](docs/benchmark.md#choosing-the-defaults-issue-419)).
483
483
  - **Mode:** `read-only` by default; `--mode observe` sends no form at all, `--mode safe-write` lets the run create records and change only the ones it created. `--mode destructive` runs only with `--allow-destructive` as well.
484
484
  - **Duplicates:** when the dedup rule keeps a filed finding apart, the run's model is asked at its lowest effort whether it is one already open on the same page, and merges it on a "same", keeping the filing's title, category, severity and evidence under that finding. The two findings' titles, categories and evidence, and the page's path, are sent; `--dedup rule` turns it off ([ADR 17](docs/adr/0017-a-model-judges-only-the-merges-the-rule-misses.md)).
485
485
  - **Output**, in `.scenescout/ci/` (or `--out`): `report.md` and `report.html` (the report an agent's run writes), `summary.md` (also appended to the GitHub job summary), `ci.json` and `ci.sarif`, with a usage line: turns, tokens, time and an estimated cost where the model's price is known (`--price-in`, `--price-out` give one for any model).
486
486
 
487
487
  There is a GitHub Action for it (`uses: brunoboto96/SceneScout/ci@…`). [docs/ci.md](docs/ci.md#an-unattended-exploratory-run) has the workflow and every option; [ADR 14](docs/adr/0014-an-unattended-run-reports-and-never-gates.md) says why it works this way.
488
488
 
489
- On a pull request, an allowed account can comment `/scenescout qa` to run it against that pull request's deployed preview and get the results as a reply. The job that holds the key checks out nothing and runs SceneScout from an exact release tag, so the pull request's code never runs beside the key. [docs/ci.md](docs/ci.md#a-qa-review-from-a-pull-request-comment) has the workflow to copy and what a project configures; [ADR 15](docs/adr/0015-a-qa-comment-tests-a-preview-and-never-runs-the-pull-requests-code.md) says why.
489
+ On a pull request, an allowed account can comment `/scenescout qa` to run it against that pull request's deployed preview and get the results as a reply. The job that holds the key checks out nothing and runs SceneScout from an exact release tag, so the pull request's code never runs beside the key. A project without previews can comment `/scenescout qa check` instead: it dispatches the project's own recorded `scenescout check` workflow on the pull request's branch and replies with the verdict, per journey, with no model key involved. [docs/ci.md](docs/ci.md#a-qa-review-from-a-pull-request-comment) has the workflows to copy and what a project configures; [ADR 15](docs/adr/0015-a-qa-comment-tests-a-preview-and-never-runs-the-pull-requests-code.md) says why.
490
490
 
491
491
  ## 📮 Filing findings as issues
492
492
 
package/dist/check-run.js CHANGED
@@ -512,6 +512,27 @@ export function clearReplayOutput(outDir) {
512
512
  if (fs.readdirSync(root).length === 0)
513
513
  fs.rmdirSync(root);
514
514
  }
515
+ /**
516
+ * Why a recorded check (--record or --video) must not start, or null: a
517
+ * replay.html in the output folder that a check did not write is the
518
+ * project's own, and is neither removed nor overwritten. Asked before the
519
+ * browser starts, so the run is not spent first.
520
+ */
521
+ export function replayPageConflict(outDir) {
522
+ const page = path.join(outDir, REPLAY_FILE);
523
+ let text;
524
+ try {
525
+ text = fs.readFileSync(page, "utf8");
526
+ }
527
+ catch (err) {
528
+ if (err.code === "ENOENT")
529
+ return null;
530
+ throw err;
531
+ }
532
+ if (isGeneratedReplay(text))
533
+ return null;
534
+ return `${page} is not a page SceneScout wrote, so a recorded check will not overwrite it. Move or rename it, or pass --out to write the check somewhere else`;
535
+ }
515
536
  /**
516
537
  * The replay page's model for a recorded check, with its frames copied from
517
538
  * the browsers' scratch folders to replay-frames/ beside the report: routes
package/dist/cli.js CHANGED
@@ -21,7 +21,7 @@ import { APPROX_DISK_MB, defaultAttachNote, defaultEngine, launchTarget, parseBr
21
21
  import { CLIENT_LABELS, firstMessageHint, manualFor, parseClients, registerWithClient, vscodeBinary } from "./clients.js";
22
22
  import { CLAUDE_CODE_NOT_NEEDED, CLI_NAME, desktopExtensionRoots, diagnose, doctorAllGood, findDesktopExtension, installClosing, ensureCommand, findOnUserPath, installSkill, isEphemeralRoot, launchCommand, manualRegisterCommand, planCommand, registerMcp, resolveClaudeDir, spawnRunner, } from "./installer.js";
23
23
  import { downloadBrowsers, presentBrowsers } from "./installer.js";
24
- import { baselinesDirOf, clearReplayOutput, defaultCheckDir, readCheckInputs, runCheck } from "./check-run.js";
24
+ import { baselinesDirOf, clearReplayOutput, replayPageConflict, defaultCheckDir, readCheckInputs, runCheck } from "./check-run.js";
25
25
  import { buildCheckReplayHtml, commitOf, REPLAY_FILE, replayFrames, replayVideos } from "./engine/check-replay.js";
26
26
  import { recordChoice } from "./engine/capture.js";
27
27
  import { httpClient, httpJudgeAsk, runCi } from "./ci-run.js";
@@ -122,10 +122,10 @@ Usage:
122
122
  --model id (default claude-sonnet-5 / gpt-6-luna); --effort none|low|medium|
123
123
  high|xhigh|max (default low; none is OpenAI only); --base-url https://…/v1 for
124
124
  another endpoint that implements the same API;
125
- --max-turns N (default 40); --max-tokens N (default 1500000);
125
+ --max-turns N (default 80); --max-tokens N (default 3000000);
126
126
  --max-minutes N (default 20): the run stops at the first cap reached and still
127
127
  writes the report;
128
- --lanes N (default 1, at most 8): split the app between N model loops that
128
+ --lanes N (default 2, 1 with --show; at most 8): split the app between N model loops that
129
129
  explore at once, each in its own browser, sharing those caps;
130
130
  --price-in, --price-cached-in, --price-out: US dollars per million tokens,
131
131
  over the built-in prices, for the cost estimate of any model;
@@ -583,6 +583,10 @@ async function check(args) {
583
583
  try {
584
584
  // --record, else SCENESCOUT_RECORD, else off.
585
585
  options = { ...options, record: recordChoice(options.record, process.env) };
586
+ // A replay.html the project keeps there is its own: a recorded check refuses to start rather than overwrite it.
587
+ const conflict = options.record || options.video ? replayPageConflict(outDir) : null;
588
+ if (conflict)
589
+ throw new Error(conflict);
586
590
  // A replay page an earlier run left must never be read, or uploaded, as this run's.
587
591
  clearReplayOutput(outDir);
588
592
  inputs = readCheckInputs(options);
@@ -16,6 +16,7 @@ import { extractCreatedIds, isOwnedResource, normalizeId } from "./ownership.js"
16
16
  import { formatJourney, journeyTime, measureJourney } from "./journey.js";
17
17
  import { describeStep, elementStateMatches, FLOW_AFTER_LAST_STEP_MS, isAction, matchRequest, notActionable, parseTarget, repeatFailure, splitRefusals, TARGET_HELP, urlMatches, } from "./flow.js";
18
18
  import { framePath, RECORD_MAX_FRAMES } from "./replay.js";
19
+ import { film } from "./film.js";
19
20
  import { InFlightRequests, keepWatchingUrl, normalizePace, SETTLE_TICK_MS, shouldKeepWaiting } from "./settle.js";
20
21
  import { crawledRoute, crawlLine, isFileMediaType, isNonPageResource, mainStateFlag, mediaTypeOf } from "./crawl.js";
21
22
  import { authToRemember, BODY_FETCH_MAX, buildRequestScript, formatPageRequests, formatReplay, PageRequests, replaySignature, requestHeaders, resolveMethod, resolveRequestUrl, resolveTarget, staleCredentialNote, toReplayResult, wantsView, } from "./request.js";
@@ -1689,30 +1690,11 @@ export class BrowserEngine {
1689
1690
  */
1690
1691
  async filming(videoTo, work) {
1691
1692
  const page = this.requirePage();
1692
- const why = (err) => (err instanceof Error ? err.message.split("\n")[0] : String(err));
1693
- let videoError;
1694
- let filming = false;
1695
- try {
1696
- await fs.promises.mkdir(path.dirname(videoTo), { recursive: true });
1697
- await page.screencast.start({ path: videoTo, ...(page.viewportSize() ? { size: page.viewportSize() } : {}) });
1698
- filming = true;
1699
- }
1700
- catch (err) {
1701
- videoError = `the video could not be started: ${why(err)}`;
1702
- }
1703
- let value;
1704
- try {
1705
- value = await work();
1706
- }
1707
- finally {
1708
- // The page the filming began on, even if the work moved the session to a popup.
1709
- if (filming)
1710
- await page.screencast.stop().catch((err) => (videoError = `the video could not be saved: ${why(err)}`));
1711
- }
1712
- if (videoError)
1713
- return { value, video: null, videoError };
1714
- const saved = await fs.promises.stat(videoTo).then((st) => st.size > 0, () => false);
1715
- return saved ? { value, video: videoTo } : { value, video: null, videoError: "no video was written" };
1693
+ // The page the filming began on is the one stopped, even if the work moved the session to a popup.
1694
+ return film(page.screencast, videoTo, page.viewportSize(), work, {
1695
+ prepare: (file) => fs.promises.mkdir(path.dirname(file), { recursive: true }).then(() => undefined),
1696
+ written: (file) => fs.promises.stat(file).then((st) => st.size > 0, () => false),
1697
+ });
1716
1698
  }
1717
1699
  requirePage() {
1718
1700
  if (!this.page || !this.memory) {
package/dist/engine/ci.js CHANGED
@@ -53,15 +53,19 @@ export const DEFAULT_CI_DEDUP = "judge";
53
53
  export const DEDUP_ENV = "SCENESCOUT_DEDUP";
54
54
  /** Which provider the MCP server's dedup judge uses when both keys are in its environment. */
55
55
  export const DEDUP_PROVIDER_ENV = "SCENESCOUT_DEDUP_PROVIDER";
56
- export const DEFAULT_CAPS = { turns: 40, tokens: 1_500_000, wallMs: 20 * 60_000 };
56
+ /** Why these are the defaults: docs/benchmark.md, "Choosing the defaults (issue 419)". */
57
+ export const DEFAULT_CAPS = { turns: 80, tokens: 3_000_000, wallMs: 20 * 60_000 };
57
58
  const CAP_BOUNDS = { turns: [1, 500], tokens: [1_000, 20_000_000], minutes: [1, 360] };
58
59
  /**
59
60
  * How many model loops explore at once, each in its own browser session and
60
61
  * its own part of the app (engine/ci-lanes.ts). 1 is the single loop. The most
61
62
  * is scout_lane_brief's, since the split is the same one. Why the default is
62
- * what it is: docs/benchmark.md, "Unattended runs".
63
+ * what it is: docs/benchmark.md, "Choosing the defaults (issue 419)". Without
64
+ * --lanes, a run that cannot take this many (--show, or fewer turns than
65
+ * lanes) runs as one loop rather than failing; --lanes given explicitly is
66
+ * held to those rules.
63
67
  */
64
- export const DEFAULT_LANES = 1;
68
+ export const DEFAULT_LANES = 2;
65
69
  export const MAX_CI_LANES = MAX_LANES;
66
70
  /** Every option `scenescout ci` accepts; the ci action's inputs are these names (ci-test holds them equal). */
67
71
  export const CI_OPTION_NAMES = [
@@ -183,7 +187,9 @@ export function parseCiArgs(args, cwd) {
183
187
  const turns = whole("max-turns", CAP_BOUNDS.turns, DEFAULT_CAPS.turns);
184
188
  const tokens = whole("max-tokens", CAP_BOUNDS.tokens, DEFAULT_CAPS.tokens);
185
189
  const minutes = whole("max-minutes", CAP_BOUNDS.minutes, DEFAULT_CAPS.wallMs / 60_000);
186
- const lanes = whole("lanes", [1, MAX_CI_LANES], DEFAULT_LANES);
190
+ // The default split yields to --show and to a turn cap below it; an explicit --lanes is checked against them below.
191
+ const defaultLanes = flags.has("show") || typeof turns !== "number" ? 1 : Math.min(DEFAULT_LANES, turns);
192
+ const lanes = whole("lanes", [1, MAX_CI_LANES], defaultLanes);
187
193
  for (const v of [turns, tokens, minutes, lanes])
188
194
  if (typeof v === "string")
189
195
  return { ok: false, error: v };
@@ -0,0 +1,41 @@
1
+ /**
2
+ * Filming one piece of work into a video file (`scenescout check --video`):
3
+ * start, run the work, stop, and say whether a video was saved. The browser
4
+ * hands in its page's screencast; nothing here needs Playwright, so the
5
+ * failure paths are table-tested.
6
+ */
7
+ const why = (err) => (err instanceof Error ? err.message.split("\n")[0] : String(err));
8
+ /**
9
+ * Run `work` while `screencast` films it into `videoTo`. The work's own result
10
+ * and errors are never touched by the filming: a video that cannot be started
11
+ * or saved is reported in `videoError`, and the work runs regardless.
12
+ *
13
+ * A start that fails is followed by a stop, its error ignored: a screencast
14
+ * left half-started would refuse every later start ("already started"), so
15
+ * one failure would cost the video of every flow after it.
16
+ */
17
+ export async function film(screencast, videoTo, size, work, io) {
18
+ let videoError;
19
+ let filming = false;
20
+ try {
21
+ await io.prepare(videoTo);
22
+ await screencast.start({ path: videoTo, ...(size ? { size } : {}) });
23
+ filming = true;
24
+ }
25
+ catch (err) {
26
+ videoError = `the video could not be started: ${why(err)}`;
27
+ // Whatever the failed start left running is stopped, so the next flow can start afresh; there is nothing to save.
28
+ await screencast.stop().catch(() => undefined);
29
+ }
30
+ let value;
31
+ try {
32
+ value = await work();
33
+ }
34
+ finally {
35
+ if (filming)
36
+ await screencast.stop().catch((err) => (videoError = `the video could not be saved: ${why(err)}`));
37
+ }
38
+ if (videoError)
39
+ return { value, video: null, videoError };
40
+ return (await io.written(videoTo)) ? { value, video: videoTo } : { value, video: null, videoError: "no video was written" };
41
+ }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "scenescout",
3
- "version": "3.20.1",
3
+ "version": "3.21.0",
4
4
  "description": "SceneScout — exploratory UI testing for AI coding agents. An MCP server that gives any agent (Claude Code, Cursor, VS Code Copilot, Codex, Gemini CLI and others) a structured view of a running web app, always-on oracles, a network-level write policy, memory across runs and a gap-checked report.",
5
5
  "license": "MIT",
6
6
  "author": "brunoboto96",