scenescout 3.20.2 → 3.21.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,18 @@
1
1
  # scenescout
2
2
 
3
+ ## 3.21.1
4
+
5
+ ### Patch Changes
6
+
7
+ - 828d177: `/scenescout qa check` follow-ups. Behaviour change: a comment beginning `/scenescout qa check …` used to be a preview run with the focus "check …"; it now always means the project's own check, so give a preview run's focus in other words. The reply no longer calls a run "Passed" when a focus matched no journeys. It says why GitHub refused a dispatch: a deleted branch, a branch whose copy of the workflow lacks the trigger or the file, or undeclared inputs. A cancelled run gets its own reply. The dispatch names the branch as `refs/heads/<branch>`, so a tag of the same name is never run. The artifact is downloaded as its zip and never extracted. Only `check.json` is read from it, capped at 10 MB, and an artifact over 1 GB is not downloaded. The example check workflow runs one check per pull request.
8
+
9
+ ## 3.21.0
10
+
11
+ ### Minor Changes
12
+
13
+ - 64a719b: `scenescout ci` and its GitHub Action now default to two lanes sharing 80 model turns and 3,000,000 tokens (from one loop, 40 turns and 1,500,000 tokens). On the benchmark's demo app the new defaults found 5 to 7 of 13 planted defects in three runs, against 2 to 4 for the old ones, and 4 to 5 of 10 on the held-out app, at about $0.03 a run on `gpt-6-luna` instead of about $0.011. Without `--lanes`, a run with `--show` or with `--max-turns 1` runs as one loop. `--lanes 1 --max-turns 40 --max-tokens 1500000` restores the old behaviour.
14
+ - a4531a8: `/scenescout qa check [focus]` runs a project's own recorded check from a pull-request comment, for projects with no preview deployments. Set the repository variable `SCENESCOUT_QA_CHECK_WORKFLOW` to a workflow that runs `scenescout check --record --video` (`examples/workflows/scenescout-qa-check.yml` is one): the comment workflow dispatches it on the pull request's branch, waits for it, and replies with the verdict, each journey's result, the findings by severity, the first failing step and links to the run and its artifact. No model key is involved, and pull requests from forks are refused. `/scenescout qa`, `show` and `compare` are unchanged.
15
+
3
16
  ## 3.20.2
4
17
 
5
18
  ### Patch Changes
package/README.md CHANGED
@@ -479,14 +479,14 @@ npx scenescout ci http://127.0.0.1:3000
479
479
 
480
480
  - **It reports and never gates.** Exit 0 when the run ran, whatever it found; exit 2 when it could not run (no key, a key the API refused, an app that never answered). Two runs of the same app find different things, so a finding is something to read, never a reason to fail a build. `scenescout check` is the gate.
481
481
  - **Providers:** the Anthropic Messages API (default model `claude-sonnet-5`) or the OpenAI Responses API (default `gpt-6-luna`), chosen by which key is set; with both set, `--provider` decides. `--model` and `--effort` (default `low`) override; `--base-url` points at another endpoint that implements the same API.
482
- - **Caps:** at most 40 model turns, 1,500,000 tokens and 20 minutes (`--max-turns`, `--max-tokens`, `--max-minutes`). The first cap reached ends the exploration; the report is still written, and says which cap ended it.
482
+ - **Caps:** at most 80 model turns, 3,000,000 tokens and 20 minutes (`--max-turns`, `--max-tokens`, `--max-minutes`), shared by two model loops that each explore their own part of the app (`--lanes`, default 2). The first cap reached ends the exploration; the report is still written, and says which cap ended it. On the benchmark's demo app a run at these defaults cost about $0.03 on `gpt-6-luna` ([docs/benchmark.md](docs/benchmark.md#choosing-the-defaults-issue-419)).
483
483
  - **Mode:** `read-only` by default; `--mode observe` sends no form at all, `--mode safe-write` lets the run create records and change only the ones it created. `--mode destructive` runs only with `--allow-destructive` as well.
484
484
  - **Duplicates:** when the dedup rule keeps a filed finding apart, the run's model is asked at its lowest effort whether it is one already open on the same page, and merges it on a "same", keeping the filing's title, category, severity and evidence under that finding. The two findings' titles, categories and evidence, and the page's path, are sent; `--dedup rule` turns it off ([ADR 17](docs/adr/0017-a-model-judges-only-the-merges-the-rule-misses.md)).
485
485
  - **Output**, in `.scenescout/ci/` (or `--out`): `report.md` and `report.html` (the report an agent's run writes), `summary.md` (also appended to the GitHub job summary), `ci.json` and `ci.sarif`, with a usage line: turns, tokens, time and an estimated cost where the model's price is known (`--price-in`, `--price-out` give one for any model).
486
486
 
487
487
  There is a GitHub Action for it (`uses: brunoboto96/SceneScout/ci@…`). [docs/ci.md](docs/ci.md#an-unattended-exploratory-run) has the workflow and every option; [ADR 14](docs/adr/0014-an-unattended-run-reports-and-never-gates.md) says why it works this way.
488
488
 
489
- On a pull request, an allowed account can comment `/scenescout qa` to run it against that pull request's deployed preview and get the results as a reply. The job that holds the key checks out nothing and runs SceneScout from an exact release tag, so the pull request's code never runs beside the key. [docs/ci.md](docs/ci.md#a-qa-review-from-a-pull-request-comment) has the workflow to copy and what a project configures; [ADR 15](docs/adr/0015-a-qa-comment-tests-a-preview-and-never-runs-the-pull-requests-code.md) says why.
489
+ On a pull request, an allowed account can comment `/scenescout qa` to run it against that pull request's deployed preview and get the results as a reply. The job that holds the key checks out nothing and runs SceneScout from an exact release tag, so the pull request's code never runs beside the key. A project without previews can comment `/scenescout qa check` instead: it dispatches the project's own recorded `scenescout check` workflow on the pull request's branch and replies with the verdict, per journey, with no model key involved. [docs/ci.md](docs/ci.md#a-qa-review-from-a-pull-request-comment) has the workflows to copy and what a project configures; [ADR 15](docs/adr/0015-a-qa-comment-tests-a-preview-and-never-runs-the-pull-requests-code.md) says why.
490
490
 
491
491
  ## 📮 Filing findings as issues
492
492
 
package/dist/cli.js CHANGED
@@ -122,10 +122,10 @@ Usage:
122
122
  --model id (default claude-sonnet-5 / gpt-6-luna); --effort none|low|medium|
123
123
  high|xhigh|max (default low; none is OpenAI only); --base-url https://…/v1 for
124
124
  another endpoint that implements the same API;
125
- --max-turns N (default 40); --max-tokens N (default 1500000);
125
+ --max-turns N (default 80); --max-tokens N (default 3000000);
126
126
  --max-minutes N (default 20): the run stops at the first cap reached and still
127
127
  writes the report;
128
- --lanes N (default 1, at most 8): split the app between N model loops that
128
+ --lanes N (default 2, 1 with --show; at most 8): split the app between N model loops that
129
129
  explore at once, each in its own browser, sharing those caps;
130
130
  --price-in, --price-cached-in, --price-out: US dollars per million tokens,
131
131
  over the built-in prices, for the cost estimate of any model;
package/dist/engine/ci.js CHANGED
@@ -53,15 +53,19 @@ export const DEFAULT_CI_DEDUP = "judge";
53
53
  export const DEDUP_ENV = "SCENESCOUT_DEDUP";
54
54
  /** Which provider the MCP server's dedup judge uses when both keys are in its environment. */
55
55
  export const DEDUP_PROVIDER_ENV = "SCENESCOUT_DEDUP_PROVIDER";
56
- export const DEFAULT_CAPS = { turns: 40, tokens: 1_500_000, wallMs: 20 * 60_000 };
56
+ /** Why these are the defaults: docs/benchmark.md, "Choosing the defaults (issue 419)". */
57
+ export const DEFAULT_CAPS = { turns: 80, tokens: 3_000_000, wallMs: 20 * 60_000 };
57
58
  const CAP_BOUNDS = { turns: [1, 500], tokens: [1_000, 20_000_000], minutes: [1, 360] };
58
59
  /**
59
60
  * How many model loops explore at once, each in its own browser session and
60
61
  * its own part of the app (engine/ci-lanes.ts). 1 is the single loop. The most
61
62
  * is scout_lane_brief's, since the split is the same one. Why the default is
62
- * what it is: docs/benchmark.md, "Unattended runs".
63
+ * what it is: docs/benchmark.md, "Choosing the defaults (issue 419)". Without
64
+ * --lanes, a run that cannot take this many (--show, or fewer turns than
65
+ * lanes) runs as one loop rather than failing; --lanes given explicitly is
66
+ * held to those rules.
63
67
  */
64
- export const DEFAULT_LANES = 1;
68
+ export const DEFAULT_LANES = 2;
65
69
  export const MAX_CI_LANES = MAX_LANES;
66
70
  /** Every option `scenescout ci` accepts; the ci action's inputs are these names (ci-test holds them equal). */
67
71
  export const CI_OPTION_NAMES = [
@@ -183,7 +187,9 @@ export function parseCiArgs(args, cwd) {
183
187
  const turns = whole("max-turns", CAP_BOUNDS.turns, DEFAULT_CAPS.turns);
184
188
  const tokens = whole("max-tokens", CAP_BOUNDS.tokens, DEFAULT_CAPS.tokens);
185
189
  const minutes = whole("max-minutes", CAP_BOUNDS.minutes, DEFAULT_CAPS.wallMs / 60_000);
186
- const lanes = whole("lanes", [1, MAX_CI_LANES], DEFAULT_LANES);
190
+ // The default split yields to --show and to a turn cap below it; an explicit --lanes is checked against them below.
191
+ const defaultLanes = flags.has("show") || typeof turns !== "number" ? 1 : Math.min(DEFAULT_LANES, turns);
192
+ const lanes = whole("lanes", [1, MAX_CI_LANES], defaultLanes);
187
193
  for (const v of [turns, tokens, minutes, lanes])
188
194
  if (typeof v === "string")
189
195
  return { ok: false, error: v };
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "scenescout",
3
- "version": "3.20.2",
3
+ "version": "3.21.1",
4
4
  "description": "SceneScout — exploratory UI testing for AI coding agents. An MCP server that gives any agent (Claude Code, Cursor, VS Code Copilot, Codex, Gemini CLI and others) a structured view of a running web app, always-on oracles, a network-level write policy, memory across runs and a gap-checked report.",
5
5
  "license": "MIT",
6
6
  "author": "brunoboto96",