scenescout 3.20.1 → 3.21.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +13 -0
- package/README.md +2 -2
- package/dist/check-run.js +21 -0
- package/dist/cli.js +7 -3
- package/dist/engine/browser.js +6 -24
- package/dist/engine/ci.js +10 -4
- package/dist/engine/film.js +41 -0
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,18 @@
|
|
|
1
1
|
# scenescout
|
|
2
2
|
|
|
3
|
+
## 3.21.0
|
|
4
|
+
|
|
5
|
+
### Minor Changes
|
|
6
|
+
|
|
7
|
+
- 64a719b: `scenescout ci` and its GitHub Action now default to two lanes sharing 80 model turns and 3,000,000 tokens (from one loop, 40 turns and 1,500,000 tokens). On the benchmark's demo app the new defaults found 5 to 7 of 13 planted defects in three runs, against 2 to 4 for the old ones, and 4 to 5 of 10 on the held-out app, at about $0.03 a run on `gpt-6-luna` instead of about $0.011. Without `--lanes`, a run with `--show` or with `--max-turns 1` runs as one loop. `--lanes 1 --max-turns 40 --max-tokens 1500000` restores the old behaviour.
|
|
8
|
+
- a4531a8: `/scenescout qa check [focus]` runs a project's own recorded check from a pull-request comment, for projects with no preview deployments. Set the repository variable `SCENESCOUT_QA_CHECK_WORKFLOW` to a workflow that runs `scenescout check --record --video` (`examples/workflows/scenescout-qa-check.yml` is one): the comment workflow dispatches it on the pull request's branch, waits for it, and replies with the verdict, each journey's result, the findings by severity, the first failing step and links to the run and its artifact. No model key is involved, and pull requests from forks are refused. `/scenescout qa`, `show` and `compare` are unchanged.
|
|
9
|
+
|
|
10
|
+
## 3.20.2
|
|
11
|
+
|
|
12
|
+
### Patch Changes
|
|
13
|
+
|
|
14
|
+
- f8f7ad2: A recorded `scenescout check` (`--record` or `--video`) no longer overwrites a `replay.html` in the output folder that it did not write: it stops before it starts, with exit code 2 and a message naming the file. With `--video`, a video that fails to start no longer stops every later flow from being filmed.
|
|
15
|
+
|
|
3
16
|
## 3.20.1
|
|
4
17
|
|
|
5
18
|
### Patch Changes
|
package/README.md
CHANGED
|
@@ -479,14 +479,14 @@ npx scenescout ci http://127.0.0.1:3000
|
|
|
479
479
|
|
|
480
480
|
- **It reports and never gates.** Exit 0 when the run ran, whatever it found; exit 2 when it could not run (no key, a key the API refused, an app that never answered). Two runs of the same app find different things, so a finding is something to read, never a reason to fail a build. `scenescout check` is the gate.
|
|
481
481
|
- **Providers:** the Anthropic Messages API (default model `claude-sonnet-5`) or the OpenAI Responses API (default `gpt-6-luna`), chosen by which key is set; with both set, `--provider` decides. `--model` and `--effort` (default `low`) override; `--base-url` points at another endpoint that implements the same API.
|
|
482
|
-
- **Caps:** at most
|
|
482
|
+
- **Caps:** at most 80 model turns, 3,000,000 tokens and 20 minutes (`--max-turns`, `--max-tokens`, `--max-minutes`), shared by two model loops that each explore their own part of the app (`--lanes`, default 2). The first cap reached ends the exploration; the report is still written, and says which cap ended it. On the benchmark's demo app a run at these defaults cost about $0.03 on `gpt-6-luna` ([docs/benchmark.md](docs/benchmark.md#choosing-the-defaults-issue-419)).
|
|
483
483
|
- **Mode:** `read-only` by default; `--mode observe` sends no form at all, `--mode safe-write` lets the run create records and change only the ones it created. `--mode destructive` runs only with `--allow-destructive` as well.
|
|
484
484
|
- **Duplicates:** when the dedup rule keeps a filed finding apart, the run's model is asked at its lowest effort whether it is one already open on the same page, and merges it on a "same", keeping the filing's title, category, severity and evidence under that finding. The two findings' titles, categories and evidence, and the page's path, are sent; `--dedup rule` turns it off ([ADR 17](docs/adr/0017-a-model-judges-only-the-merges-the-rule-misses.md)).
|
|
485
485
|
- **Output**, in `.scenescout/ci/` (or `--out`): `report.md` and `report.html` (the report an agent's run writes), `summary.md` (also appended to the GitHub job summary), `ci.json` and `ci.sarif`, with a usage line: turns, tokens, time and an estimated cost where the model's price is known (`--price-in`, `--price-out` give one for any model).
|
|
486
486
|
|
|
487
487
|
There is a GitHub Action for it (`uses: brunoboto96/SceneScout/ci@…`). [docs/ci.md](docs/ci.md#an-unattended-exploratory-run) has the workflow and every option; [ADR 14](docs/adr/0014-an-unattended-run-reports-and-never-gates.md) says why it works this way.
|
|
488
488
|
|
|
489
|
-
On a pull request, an allowed account can comment `/scenescout qa` to run it against that pull request's deployed preview and get the results as a reply. The job that holds the key checks out nothing and runs SceneScout from an exact release tag, so the pull request's code never runs beside the key. [docs/ci.md](docs/ci.md#a-qa-review-from-a-pull-request-comment) has the
|
|
489
|
+
On a pull request, an allowed account can comment `/scenescout qa` to run it against that pull request's deployed preview and get the results as a reply. The job that holds the key checks out nothing and runs SceneScout from an exact release tag, so the pull request's code never runs beside the key. A project without previews can comment `/scenescout qa check` instead: it dispatches the project's own recorded `scenescout check` workflow on the pull request's branch and replies with the verdict, per journey, with no model key involved. [docs/ci.md](docs/ci.md#a-qa-review-from-a-pull-request-comment) has the workflows to copy and what a project configures; [ADR 15](docs/adr/0015-a-qa-comment-tests-a-preview-and-never-runs-the-pull-requests-code.md) says why.
|
|
490
490
|
|
|
491
491
|
## 📮 Filing findings as issues
|
|
492
492
|
|
package/dist/check-run.js
CHANGED
|
@@ -512,6 +512,27 @@ export function clearReplayOutput(outDir) {
|
|
|
512
512
|
if (fs.readdirSync(root).length === 0)
|
|
513
513
|
fs.rmdirSync(root);
|
|
514
514
|
}
|
|
515
|
+
/**
|
|
516
|
+
* Why a recorded check (--record or --video) must not start, or null: a
|
|
517
|
+
* replay.html in the output folder that a check did not write is the
|
|
518
|
+
* project's own, and is neither removed nor overwritten. Asked before the
|
|
519
|
+
* browser starts, so the run is not spent first.
|
|
520
|
+
*/
|
|
521
|
+
export function replayPageConflict(outDir) {
|
|
522
|
+
const page = path.join(outDir, REPLAY_FILE);
|
|
523
|
+
let text;
|
|
524
|
+
try {
|
|
525
|
+
text = fs.readFileSync(page, "utf8");
|
|
526
|
+
}
|
|
527
|
+
catch (err) {
|
|
528
|
+
if (err.code === "ENOENT")
|
|
529
|
+
return null;
|
|
530
|
+
throw err;
|
|
531
|
+
}
|
|
532
|
+
if (isGeneratedReplay(text))
|
|
533
|
+
return null;
|
|
534
|
+
return `${page} is not a page SceneScout wrote, so a recorded check will not overwrite it. Move or rename it, or pass --out to write the check somewhere else`;
|
|
535
|
+
}
|
|
515
536
|
/**
|
|
516
537
|
* The replay page's model for a recorded check, with its frames copied from
|
|
517
538
|
* the browsers' scratch folders to replay-frames/ beside the report: routes
|
package/dist/cli.js
CHANGED
|
@@ -21,7 +21,7 @@ import { APPROX_DISK_MB, defaultAttachNote, defaultEngine, launchTarget, parseBr
|
|
|
21
21
|
import { CLIENT_LABELS, firstMessageHint, manualFor, parseClients, registerWithClient, vscodeBinary } from "./clients.js";
|
|
22
22
|
import { CLAUDE_CODE_NOT_NEEDED, CLI_NAME, desktopExtensionRoots, diagnose, doctorAllGood, findDesktopExtension, installClosing, ensureCommand, findOnUserPath, installSkill, isEphemeralRoot, launchCommand, manualRegisterCommand, planCommand, registerMcp, resolveClaudeDir, spawnRunner, } from "./installer.js";
|
|
23
23
|
import { downloadBrowsers, presentBrowsers } from "./installer.js";
|
|
24
|
-
import { baselinesDirOf, clearReplayOutput, defaultCheckDir, readCheckInputs, runCheck } from "./check-run.js";
|
|
24
|
+
import { baselinesDirOf, clearReplayOutput, replayPageConflict, defaultCheckDir, readCheckInputs, runCheck } from "./check-run.js";
|
|
25
25
|
import { buildCheckReplayHtml, commitOf, REPLAY_FILE, replayFrames, replayVideos } from "./engine/check-replay.js";
|
|
26
26
|
import { recordChoice } from "./engine/capture.js";
|
|
27
27
|
import { httpClient, httpJudgeAsk, runCi } from "./ci-run.js";
|
|
@@ -122,10 +122,10 @@ Usage:
|
|
|
122
122
|
--model id (default claude-sonnet-5 / gpt-6-luna); --effort none|low|medium|
|
|
123
123
|
high|xhigh|max (default low; none is OpenAI only); --base-url https://…/v1 for
|
|
124
124
|
another endpoint that implements the same API;
|
|
125
|
-
--max-turns N (default
|
|
125
|
+
--max-turns N (default 80); --max-tokens N (default 3000000);
|
|
126
126
|
--max-minutes N (default 20): the run stops at the first cap reached and still
|
|
127
127
|
writes the report;
|
|
128
|
-
--lanes N (default
|
|
128
|
+
--lanes N (default 2, 1 with --show; at most 8): split the app between N model loops that
|
|
129
129
|
explore at once, each in its own browser, sharing those caps;
|
|
130
130
|
--price-in, --price-cached-in, --price-out: US dollars per million tokens,
|
|
131
131
|
over the built-in prices, for the cost estimate of any model;
|
|
@@ -583,6 +583,10 @@ async function check(args) {
|
|
|
583
583
|
try {
|
|
584
584
|
// --record, else SCENESCOUT_RECORD, else off.
|
|
585
585
|
options = { ...options, record: recordChoice(options.record, process.env) };
|
|
586
|
+
// A replay.html the project keeps there is its own: a recorded check refuses to start rather than overwrite it.
|
|
587
|
+
const conflict = options.record || options.video ? replayPageConflict(outDir) : null;
|
|
588
|
+
if (conflict)
|
|
589
|
+
throw new Error(conflict);
|
|
586
590
|
// A replay page an earlier run left must never be read, or uploaded, as this run's.
|
|
587
591
|
clearReplayOutput(outDir);
|
|
588
592
|
inputs = readCheckInputs(options);
|
package/dist/engine/browser.js
CHANGED
|
@@ -16,6 +16,7 @@ import { extractCreatedIds, isOwnedResource, normalizeId } from "./ownership.js"
|
|
|
16
16
|
import { formatJourney, journeyTime, measureJourney } from "./journey.js";
|
|
17
17
|
import { describeStep, elementStateMatches, FLOW_AFTER_LAST_STEP_MS, isAction, matchRequest, notActionable, parseTarget, repeatFailure, splitRefusals, TARGET_HELP, urlMatches, } from "./flow.js";
|
|
18
18
|
import { framePath, RECORD_MAX_FRAMES } from "./replay.js";
|
|
19
|
+
import { film } from "./film.js";
|
|
19
20
|
import { InFlightRequests, keepWatchingUrl, normalizePace, SETTLE_TICK_MS, shouldKeepWaiting } from "./settle.js";
|
|
20
21
|
import { crawledRoute, crawlLine, isFileMediaType, isNonPageResource, mainStateFlag, mediaTypeOf } from "./crawl.js";
|
|
21
22
|
import { authToRemember, BODY_FETCH_MAX, buildRequestScript, formatPageRequests, formatReplay, PageRequests, replaySignature, requestHeaders, resolveMethod, resolveRequestUrl, resolveTarget, staleCredentialNote, toReplayResult, wantsView, } from "./request.js";
|
|
@@ -1689,30 +1690,11 @@ export class BrowserEngine {
|
|
|
1689
1690
|
*/
|
|
1690
1691
|
async filming(videoTo, work) {
|
|
1691
1692
|
const page = this.requirePage();
|
|
1692
|
-
|
|
1693
|
-
|
|
1694
|
-
|
|
1695
|
-
|
|
1696
|
-
|
|
1697
|
-
await page.screencast.start({ path: videoTo, ...(page.viewportSize() ? { size: page.viewportSize() } : {}) });
|
|
1698
|
-
filming = true;
|
|
1699
|
-
}
|
|
1700
|
-
catch (err) {
|
|
1701
|
-
videoError = `the video could not be started: ${why(err)}`;
|
|
1702
|
-
}
|
|
1703
|
-
let value;
|
|
1704
|
-
try {
|
|
1705
|
-
value = await work();
|
|
1706
|
-
}
|
|
1707
|
-
finally {
|
|
1708
|
-
// The page the filming began on, even if the work moved the session to a popup.
|
|
1709
|
-
if (filming)
|
|
1710
|
-
await page.screencast.stop().catch((err) => (videoError = `the video could not be saved: ${why(err)}`));
|
|
1711
|
-
}
|
|
1712
|
-
if (videoError)
|
|
1713
|
-
return { value, video: null, videoError };
|
|
1714
|
-
const saved = await fs.promises.stat(videoTo).then((st) => st.size > 0, () => false);
|
|
1715
|
-
return saved ? { value, video: videoTo } : { value, video: null, videoError: "no video was written" };
|
|
1693
|
+
// The page the filming began on is the one stopped, even if the work moved the session to a popup.
|
|
1694
|
+
return film(page.screencast, videoTo, page.viewportSize(), work, {
|
|
1695
|
+
prepare: (file) => fs.promises.mkdir(path.dirname(file), { recursive: true }).then(() => undefined),
|
|
1696
|
+
written: (file) => fs.promises.stat(file).then((st) => st.size > 0, () => false),
|
|
1697
|
+
});
|
|
1716
1698
|
}
|
|
1717
1699
|
requirePage() {
|
|
1718
1700
|
if (!this.page || !this.memory) {
|
package/dist/engine/ci.js
CHANGED
|
@@ -53,15 +53,19 @@ export const DEFAULT_CI_DEDUP = "judge";
|
|
|
53
53
|
export const DEDUP_ENV = "SCENESCOUT_DEDUP";
|
|
54
54
|
/** Which provider the MCP server's dedup judge uses when both keys are in its environment. */
|
|
55
55
|
export const DEDUP_PROVIDER_ENV = "SCENESCOUT_DEDUP_PROVIDER";
|
|
56
|
-
|
|
56
|
+
/** Why these are the defaults: docs/benchmark.md, "Choosing the defaults (issue 419)". */
|
|
57
|
+
export const DEFAULT_CAPS = { turns: 80, tokens: 3_000_000, wallMs: 20 * 60_000 };
|
|
57
58
|
const CAP_BOUNDS = { turns: [1, 500], tokens: [1_000, 20_000_000], minutes: [1, 360] };
|
|
58
59
|
/**
|
|
59
60
|
* How many model loops explore at once, each in its own browser session and
|
|
60
61
|
* its own part of the app (engine/ci-lanes.ts). 1 is the single loop. The most
|
|
61
62
|
* is scout_lane_brief's, since the split is the same one. Why the default is
|
|
62
|
-
* what it is: docs/benchmark.md, "
|
|
63
|
+
* what it is: docs/benchmark.md, "Choosing the defaults (issue 419)". Without
|
|
64
|
+
* --lanes, a run that cannot take this many (--show, or fewer turns than
|
|
65
|
+
* lanes) runs as one loop rather than failing; --lanes given explicitly is
|
|
66
|
+
* held to those rules.
|
|
63
67
|
*/
|
|
64
|
-
export const DEFAULT_LANES =
|
|
68
|
+
export const DEFAULT_LANES = 2;
|
|
65
69
|
export const MAX_CI_LANES = MAX_LANES;
|
|
66
70
|
/** Every option `scenescout ci` accepts; the ci action's inputs are these names (ci-test holds them equal). */
|
|
67
71
|
export const CI_OPTION_NAMES = [
|
|
@@ -183,7 +187,9 @@ export function parseCiArgs(args, cwd) {
|
|
|
183
187
|
const turns = whole("max-turns", CAP_BOUNDS.turns, DEFAULT_CAPS.turns);
|
|
184
188
|
const tokens = whole("max-tokens", CAP_BOUNDS.tokens, DEFAULT_CAPS.tokens);
|
|
185
189
|
const minutes = whole("max-minutes", CAP_BOUNDS.minutes, DEFAULT_CAPS.wallMs / 60_000);
|
|
186
|
-
|
|
190
|
+
// The default split yields to --show and to a turn cap below it; an explicit --lanes is checked against them below.
|
|
191
|
+
const defaultLanes = flags.has("show") || typeof turns !== "number" ? 1 : Math.min(DEFAULT_LANES, turns);
|
|
192
|
+
const lanes = whole("lanes", [1, MAX_CI_LANES], defaultLanes);
|
|
187
193
|
for (const v of [turns, tokens, minutes, lanes])
|
|
188
194
|
if (typeof v === "string")
|
|
189
195
|
return { ok: false, error: v };
|
|
@@ -0,0 +1,41 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* Filming one piece of work into a video file (`scenescout check --video`):
|
|
3
|
+
* start, run the work, stop, and say whether a video was saved. The browser
|
|
4
|
+
* hands in its page's screencast; nothing here needs Playwright, so the
|
|
5
|
+
* failure paths are table-tested.
|
|
6
|
+
*/
|
|
7
|
+
const why = (err) => (err instanceof Error ? err.message.split("\n")[0] : String(err));
|
|
8
|
+
/**
|
|
9
|
+
* Run `work` while `screencast` films it into `videoTo`. The work's own result
|
|
10
|
+
* and errors are never touched by the filming: a video that cannot be started
|
|
11
|
+
* or saved is reported in `videoError`, and the work runs regardless.
|
|
12
|
+
*
|
|
13
|
+
* A start that fails is followed by a stop, its error ignored: a screencast
|
|
14
|
+
* left half-started would refuse every later start ("already started"), so
|
|
15
|
+
* one failure would cost the video of every flow after it.
|
|
16
|
+
*/
|
|
17
|
+
export async function film(screencast, videoTo, size, work, io) {
|
|
18
|
+
let videoError;
|
|
19
|
+
let filming = false;
|
|
20
|
+
try {
|
|
21
|
+
await io.prepare(videoTo);
|
|
22
|
+
await screencast.start({ path: videoTo, ...(size ? { size } : {}) });
|
|
23
|
+
filming = true;
|
|
24
|
+
}
|
|
25
|
+
catch (err) {
|
|
26
|
+
videoError = `the video could not be started: ${why(err)}`;
|
|
27
|
+
// Whatever the failed start left running is stopped, so the next flow can start afresh; there is nothing to save.
|
|
28
|
+
await screencast.stop().catch(() => undefined);
|
|
29
|
+
}
|
|
30
|
+
let value;
|
|
31
|
+
try {
|
|
32
|
+
value = await work();
|
|
33
|
+
}
|
|
34
|
+
finally {
|
|
35
|
+
if (filming)
|
|
36
|
+
await screencast.stop().catch((err) => (videoError = `the video could not be saved: ${why(err)}`));
|
|
37
|
+
}
|
|
38
|
+
if (videoError)
|
|
39
|
+
return { value, video: null, videoError };
|
|
40
|
+
return (await io.written(videoTo)) ? { value, video: videoTo } : { value, video: null, videoError: "no video was written" };
|
|
41
|
+
}
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "scenescout",
|
|
3
|
-
"version": "3.
|
|
3
|
+
"version": "3.21.0",
|
|
4
4
|
"description": "SceneScout — exploratory UI testing for AI coding agents. An MCP server that gives any agent (Claude Code, Cursor, VS Code Copilot, Codex, Gemini CLI and others) a structured view of a running web app, always-on oracles, a network-level write policy, memory across runs and a gap-checked report.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"author": "brunoboto96",
|