scenescout 3.20.2 → 3.21.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +13 -0
- package/README.md +2 -2
- package/dist/cli.js +2 -2
- package/dist/engine/ci.js +10 -4
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,18 @@
|
|
|
1
1
|
# scenescout
|
|
2
2
|
|
|
3
|
+
## 3.21.1
|
|
4
|
+
|
|
5
|
+
### Patch Changes
|
|
6
|
+
|
|
7
|
+
- 828d177: `/scenescout qa check` follow-ups. Behaviour change: a comment beginning `/scenescout qa check …` used to be a preview run with the focus "check …"; it now always means the project's own check, so give a preview run's focus in other words. The reply no longer calls a run "Passed" when a focus matched no journeys. It says why GitHub refused a dispatch: a deleted branch, a branch whose copy of the workflow lacks the trigger or the file, or undeclared inputs. A cancelled run gets its own reply. The dispatch names the branch as `refs/heads/<branch>`, so a tag of the same name is never run. The artifact is downloaded as its zip and never extracted. Only `check.json` is read from it, capped at 10 MB, and an artifact over 1 GB is not downloaded. The example check workflow runs one check per pull request.
|
|
8
|
+
|
|
9
|
+
## 3.21.0
|
|
10
|
+
|
|
11
|
+
### Minor Changes
|
|
12
|
+
|
|
13
|
+
- 64a719b: `scenescout ci` and its GitHub Action now default to two lanes sharing 80 model turns and 3,000,000 tokens (from one loop, 40 turns and 1,500,000 tokens). On the benchmark's demo app the new defaults found 5 to 7 of 13 planted defects in three runs, against 2 to 4 for the old ones, and 4 to 5 of 10 on the held-out app, at about $0.03 a run on `gpt-6-luna` instead of about $0.011. Without `--lanes`, a run with `--show` or with `--max-turns 1` runs as one loop. `--lanes 1 --max-turns 40 --max-tokens 1500000` restores the old behaviour.
|
|
14
|
+
- a4531a8: `/scenescout qa check [focus]` runs a project's own recorded check from a pull-request comment, for projects with no preview deployments. Set the repository variable `SCENESCOUT_QA_CHECK_WORKFLOW` to a workflow that runs `scenescout check --record --video` (`examples/workflows/scenescout-qa-check.yml` is one): the comment workflow dispatches it on the pull request's branch, waits for it, and replies with the verdict, each journey's result, the findings by severity, the first failing step and links to the run and its artifact. No model key is involved, and pull requests from forks are refused. `/scenescout qa`, `show` and `compare` are unchanged.
|
|
15
|
+
|
|
3
16
|
## 3.20.2
|
|
4
17
|
|
|
5
18
|
### Patch Changes
|
package/README.md
CHANGED
|
@@ -479,14 +479,14 @@ npx scenescout ci http://127.0.0.1:3000
|
|
|
479
479
|
|
|
480
480
|
- **It reports and never gates.** Exit 0 when the run ran, whatever it found; exit 2 when it could not run (no key, a key the API refused, an app that never answered). Two runs of the same app find different things, so a finding is something to read, never a reason to fail a build. `scenescout check` is the gate.
|
|
481
481
|
- **Providers:** the Anthropic Messages API (default model `claude-sonnet-5`) or the OpenAI Responses API (default `gpt-6-luna`), chosen by which key is set; with both set, `--provider` decides. `--model` and `--effort` (default `low`) override; `--base-url` points at another endpoint that implements the same API.
|
|
482
|
-
- **Caps:** at most
|
|
482
|
+
- **Caps:** at most 80 model turns, 3,000,000 tokens and 20 minutes (`--max-turns`, `--max-tokens`, `--max-minutes`), shared by two model loops that each explore their own part of the app (`--lanes`, default 2). The first cap reached ends the exploration; the report is still written, and says which cap ended it. On the benchmark's demo app a run at these defaults cost about $0.03 on `gpt-6-luna` ([docs/benchmark.md](docs/benchmark.md#choosing-the-defaults-issue-419)).
|
|
483
483
|
- **Mode:** `read-only` by default; `--mode observe` sends no form at all, `--mode safe-write` lets the run create records and change only the ones it created. `--mode destructive` runs only with `--allow-destructive` as well.
|
|
484
484
|
- **Duplicates:** when the dedup rule keeps a filed finding apart, the run's model is asked at its lowest effort whether it is one already open on the same page, and merges it on a "same", keeping the filing's title, category, severity and evidence under that finding. The two findings' titles, categories and evidence, and the page's path, are sent; `--dedup rule` turns it off ([ADR 17](docs/adr/0017-a-model-judges-only-the-merges-the-rule-misses.md)).
|
|
485
485
|
- **Output**, in `.scenescout/ci/` (or `--out`): `report.md` and `report.html` (the report an agent's run writes), `summary.md` (also appended to the GitHub job summary), `ci.json` and `ci.sarif`, with a usage line: turns, tokens, time and an estimated cost where the model's price is known (`--price-in`, `--price-out` give one for any model).
|
|
486
486
|
|
|
487
487
|
There is a GitHub Action for it (`uses: brunoboto96/SceneScout/ci@…`). [docs/ci.md](docs/ci.md#an-unattended-exploratory-run) has the workflow and every option; [ADR 14](docs/adr/0014-an-unattended-run-reports-and-never-gates.md) says why it works this way.
|
|
488
488
|
|
|
489
|
-
On a pull request, an allowed account can comment `/scenescout qa` to run it against that pull request's deployed preview and get the results as a reply. The job that holds the key checks out nothing and runs SceneScout from an exact release tag, so the pull request's code never runs beside the key. [docs/ci.md](docs/ci.md#a-qa-review-from-a-pull-request-comment) has the
|
|
489
|
+
On a pull request, an allowed account can comment `/scenescout qa` to run it against that pull request's deployed preview and get the results as a reply. The job that holds the key checks out nothing and runs SceneScout from an exact release tag, so the pull request's code never runs beside the key. A project without previews can comment `/scenescout qa check` instead: it dispatches the project's own recorded `scenescout check` workflow on the pull request's branch and replies with the verdict, per journey, with no model key involved. [docs/ci.md](docs/ci.md#a-qa-review-from-a-pull-request-comment) has the workflows to copy and what a project configures; [ADR 15](docs/adr/0015-a-qa-comment-tests-a-preview-and-never-runs-the-pull-requests-code.md) says why.
|
|
490
490
|
|
|
491
491
|
## 📮 Filing findings as issues
|
|
492
492
|
|
package/dist/cli.js
CHANGED
|
@@ -122,10 +122,10 @@ Usage:
|
|
|
122
122
|
--model id (default claude-sonnet-5 / gpt-6-luna); --effort none|low|medium|
|
|
123
123
|
high|xhigh|max (default low; none is OpenAI only); --base-url https://…/v1 for
|
|
124
124
|
another endpoint that implements the same API;
|
|
125
|
-
--max-turns N (default
|
|
125
|
+
--max-turns N (default 80); --max-tokens N (default 3000000);
|
|
126
126
|
--max-minutes N (default 20): the run stops at the first cap reached and still
|
|
127
127
|
writes the report;
|
|
128
|
-
--lanes N (default
|
|
128
|
+
--lanes N (default 2, 1 with --show; at most 8): split the app between N model loops that
|
|
129
129
|
explore at once, each in its own browser, sharing those caps;
|
|
130
130
|
--price-in, --price-cached-in, --price-out: US dollars per million tokens,
|
|
131
131
|
over the built-in prices, for the cost estimate of any model;
|
package/dist/engine/ci.js
CHANGED
|
@@ -53,15 +53,19 @@ export const DEFAULT_CI_DEDUP = "judge";
|
|
|
53
53
|
export const DEDUP_ENV = "SCENESCOUT_DEDUP";
|
|
54
54
|
/** Which provider the MCP server's dedup judge uses when both keys are in its environment. */
|
|
55
55
|
export const DEDUP_PROVIDER_ENV = "SCENESCOUT_DEDUP_PROVIDER";
|
|
56
|
-
|
|
56
|
+
/** Why these are the defaults: docs/benchmark.md, "Choosing the defaults (issue 419)". */
|
|
57
|
+
export const DEFAULT_CAPS = { turns: 80, tokens: 3_000_000, wallMs: 20 * 60_000 };
|
|
57
58
|
const CAP_BOUNDS = { turns: [1, 500], tokens: [1_000, 20_000_000], minutes: [1, 360] };
|
|
58
59
|
/**
|
|
59
60
|
* How many model loops explore at once, each in its own browser session and
|
|
60
61
|
* its own part of the app (engine/ci-lanes.ts). 1 is the single loop. The most
|
|
61
62
|
* is scout_lane_brief's, since the split is the same one. Why the default is
|
|
62
|
-
* what it is: docs/benchmark.md, "
|
|
63
|
+
* what it is: docs/benchmark.md, "Choosing the defaults (issue 419)". Without
|
|
64
|
+
* --lanes, a run that cannot take this many (--show, or fewer turns than
|
|
65
|
+
* lanes) runs as one loop rather than failing; --lanes given explicitly is
|
|
66
|
+
* held to those rules.
|
|
63
67
|
*/
|
|
64
|
-
export const DEFAULT_LANES =
|
|
68
|
+
export const DEFAULT_LANES = 2;
|
|
65
69
|
export const MAX_CI_LANES = MAX_LANES;
|
|
66
70
|
/** Every option `scenescout ci` accepts; the ci action's inputs are these names (ci-test holds them equal). */
|
|
67
71
|
export const CI_OPTION_NAMES = [
|
|
@@ -183,7 +187,9 @@ export function parseCiArgs(args, cwd) {
|
|
|
183
187
|
const turns = whole("max-turns", CAP_BOUNDS.turns, DEFAULT_CAPS.turns);
|
|
184
188
|
const tokens = whole("max-tokens", CAP_BOUNDS.tokens, DEFAULT_CAPS.tokens);
|
|
185
189
|
const minutes = whole("max-minutes", CAP_BOUNDS.minutes, DEFAULT_CAPS.wallMs / 60_000);
|
|
186
|
-
|
|
190
|
+
// The default split yields to --show and to a turn cap below it; an explicit --lanes is checked against them below.
|
|
191
|
+
const defaultLanes = flags.has("show") || typeof turns !== "number" ? 1 : Math.min(DEFAULT_LANES, turns);
|
|
192
|
+
const lanes = whole("lanes", [1, MAX_CI_LANES], defaultLanes);
|
|
187
193
|
for (const v of [turns, tokens, minutes, lanes])
|
|
188
194
|
if (typeof v === "string")
|
|
189
195
|
return { ok: false, error: v };
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "scenescout",
|
|
3
|
-
"version": "3.
|
|
3
|
+
"version": "3.21.1",
|
|
4
4
|
"description": "SceneScout — exploratory UI testing for AI coding agents. An MCP server that gives any agent (Claude Code, Cursor, VS Code Copilot, Codex, Gemini CLI and others) a structured view of a running web app, always-on oracles, a network-level write policy, memory across runs and a gap-checked report.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"author": "brunoboto96",
|