scenescout 3.14.0 → 3.15.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +25 -0
- package/README.md +27 -4
- package/dist/check-run.js +22 -11
- package/dist/ci-run.js +113 -17
- package/dist/cli.js +150 -75
- package/dist/commands.js +141 -0
- package/dist/engine/browser.js +546 -133
- package/dist/engine/check.js +77 -30
- package/dist/engine/ci.js +97 -0
- package/dist/engine/claims.js +81 -15
- package/dist/engine/collector.js +307 -65
- package/dist/engine/dedup.js +329 -9
- package/dist/engine/design.js +81 -12
- package/dist/engine/forms.js +17 -0
- package/dist/engine/memory.js +180 -36
- package/dist/engine/oracles.js +126 -1
- package/dist/engine/provider.js +4 -3
- package/dist/engine/refresh.js +270 -22
- package/dist/engine/report.js +8 -1
- package/dist/engine/scripted-login.js +342 -62
- package/dist/first-run.js +577 -0
- package/dist/login-run.js +136 -35
- package/dist/mcp-server.js +75 -16
- package/package.json +3 -2
- package/skills/scenescout/SKILL.md +3 -3
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,30 @@
|
|
|
1
1
|
# scenescout
|
|
2
2
|
|
|
3
|
+
## 3.15.0
|
|
4
|
+
|
|
5
|
+
### Minor Changes
|
|
6
|
+
|
|
7
|
+
- 3c488f3: Finding dedup can ask a model. When the dedup rule keeps a newly filed finding apart from everything recorded, a model judge is asked whether it is the same defect as one of the open findings on the same page (the three most alike, at most), and a "same" merges it; the merged filing's title, category, severity and evidence are kept under the finding and shown in the report with the judge's probability. Any failure, unsure answer or slow call leaves the rule's decision and is logged once per kind; three failed calls in a row switch the judge off for the run.
|
|
8
|
+
|
|
9
|
+
`scenescout ci` judges by default with the run's model at the lowest effort its API takes, sending each asked pair's titles, categories and evidence, and the page's path; the calls count in the run's usage, and `--dedup rule` (or the action's `dedup` input) turns it off. The MCP server judges only when `SCENESCOUT_DEDUP=judge` or `scout_attach {dedup: "judge"}` asks for it, with `ANTHROPIC_API_KEY` or `OPENAI_API_KEY` in its environment (`SCENESCOUT_DEDUP_PROVIDER` picks one when both are set).
|
|
10
|
+
- 47d5293: Add `scenescout <url>`, a first look at any app with nothing set up: `npx -y scenescout http://localhost:3000`. It runs the deterministic check's crawl and measurements in observe mode, where nothing but reads leaves the page (`--mode read-only` lets a plain POST through), visiting up to 20 pages and starting none after 3 minutes (`--max-routes`, `--max-minutes`), with no model and no API key. When the headless Chromium build is missing it downloads that and nothing else: no skill, no MCP registration, nothing on the PATH. It writes `report.md` and `check.json` to `scenescout-report/` in the current folder (a temporary folder when that cannot be written, or `--out`), and only into a folder that is new, empty or an earlier first look's: any other `scenescout-report/` is left as it is and the run exits 2, pointing at `--out`, and a `report.md` or `check.json` a first look did not write is never replaced; and the summary and the report open with the three issues to look at first: the highest severity, then the most pages affected, with one failure seen several ways (a missing image is a failed request, a broken image and a console line) taking one of the three places. It exits 0 once it has looked, whatever it found, and 2 when it could not run, such as when the address cannot be reached, or could not write the report. The guide's Start here page and the README's quickstart now lead with it.
|
|
11
|
+
- 400add2: `scenescout login --script` signs in to passwordless apps and takes a fixed one-time code. `SCENESCOUT_LOGIN_OTP_CODE` is a code the test environment accepts, typed where the form asks for one, as an alternative to `SCENESCOUT_LOGIN_TOTP_SECRET` (setting both is refused at startup); it is redacted from all output like the other credentials. `SCENESCOUT_LOGIN_PASSWORD` may be left unset when a code is given: the run fills the username, presses the button that sends the code ("Send code", "Continue"), waits through the moment with no field on screen, and types the code; a password field that appears anyway stops the run naming the variable.
|
|
12
|
+
|
|
13
|
+
Finding and submitting the code: a code split into one box per character is typed one character per box; a numeric field sized for a code is the code field once the username has gone; an email field whose label mentions a code stays the username. After typing, the run reads the page until it is clear how to go on, so a button enabled only once the form is complete is waited for, a page that takes the code itself is not submitted again, and a field the typing revealed is filled first. A button that signs in or verifies is pressed rather than one that sends a code where a page has both, and buttons that resend a code or change the address are never pressed. With no success URL or selector, a page with no sign-in field now counts as signed in only once a password or a code has gone, so the pause while a code is sent is not taken for a sign-in.
|
|
14
|
+
- 3471e73: Snapshots say more about the page. Alert and status regions are listed by their text whether or not they carry a test id, and the diff reports a region that says something new. Controls show their state (`pressed`, `selected`, `checked`, `expanded`, `current`) and the diff reports it moving; the per-control `done` marker is now `exercised`. A `main:` line summarises the main area's heading and static text, or says `main: EMPTY`, and crawl lines carry the same count. Accessible names follow the accessible-name computation's order: wrapping labels and `title` name a control, a button-like input is named by its value, and a select is never named by its options. Elements listed only for a test id (wrappers, headings, decorative badges) are no longer counted as unnamed controls or in the coverage denominator, and nothing inside an `aria-hidden` subtree is counted as unnamed. Controls held outside the visible width of a horizontally scrolling container are reported once per container, under the new worth-a-look check rule `scrolled-out-controls`.
|
|
15
|
+
|
|
16
|
+
### Patch Changes
|
|
17
|
+
|
|
18
|
+
- f6997dc: The design audit's task-efficiency lines count only what they describe. Competing actions are buttons and links painted as buttons; text fields and breadcrumb links no longer count. Form burden counts the fields of a form; row-selection checkboxes, selects that edit a table row in place and search boxes no longer count, and when the page has a `<form>`, fields outside it no longer count either. The app shell's landmarks (navigation, banner, sidebar and footer regions outside the main content) are kept out of a page's score from the first audit, so the first pages of a run are no longer scored with the shell in them while later ones are scored without it. A shell built without landmarks is still recognised only once the shared-chrome census has seen it on several routes.
|
|
19
|
+
- c9edcac: Fewer false oracle violations. A `scout_request` probe the server refuses is no longer reported as an `http_error` or `console_error` of the page visited next. `false_success` now pairs a refused write only with a success message the action put on screen: a status badge or heading already there, a column header and a write the page sent in the background or to its own telemetry no longer count; a success message beside an action whose other writes went through is reported as a partial `false_success` at medium ("partial: N of M writes from this action were refused"); and an announced "was refused", "rejected" or "could not" counts as the page admitting the refusal. The silent-submit note matches its words ("sign", "post", "save") as whole words and waits for a client-side route change before calling a click silent. A page error raised by a link click that left the URL unchanged and opened a confirmation, or whose message says a route change was cancelled, is reported at medium with a note instead of high.
|
|
20
|
+
- a5f9c33: The refresh broker no longer holds back an app's requests when its refresh token is a cookie scoped to "/". A cookie alone no longer makes a request a refresh: only a POST, PUT or PATCH to a path named for a refresh, or an endpoint the broker has seen rotate the cookie (learned for every session of the role), goes through the lock. Scripts, stylesheets, images and fonts are never brokered, and when the broker cannot take the lock or read the profile the request goes out as the page sent it instead of being dropped. A refresh-named storage value that is an address is no longer taken for a token. On Windows, a lock file still being deleted as it changes hands (EPERM, EBUSY or EACCES) is waited for like a held lock instead of failing.
|
|
21
|
+
|
|
22
|
+
## 3.14.1
|
|
23
|
+
|
|
24
|
+
### Patch Changes
|
|
25
|
+
|
|
26
|
+
- 5358a81: `--help` and `-h` now print the usage and exit 0 on every subcommand before it does anything; `scenescout install --help` used to run a real install. `install`, `doctor`, `scan`, `status` and `watch` now refuse a flag or argument they do not know instead of ignoring it.
|
|
27
|
+
|
|
3
28
|
## 3.14.0
|
|
4
29
|
|
|
5
30
|
### Minor Changes
|
package/README.md
CHANGED
|
@@ -12,7 +12,7 @@ Works with Claude Code · Cursor · VS Code (Copilot) · Codex CLI · Gemini CLI
|
|
|
12
12
|

|
|
13
13
|

|
|
14
14
|
|
|
15
|
-
[👀 See it work](#-see-it-work) · [✨ Why](#-why-its-different) · [🎯 Two ways to use it](#-two-ways-to-use-it) · [🚀 Quickstart](#-quickstart) · [🧰 Toolbox](#-the-toolbox) · [🔌 Other clients](#-other-mcp-clients) · [🔒 Safety](#-safety-model) · [🩺 Troubleshooting](#-troubleshooting)
|
|
15
|
+
[📖 Guide](docs/guide/Home.md) · [👀 See it work](#-see-it-work) · [✨ Why](#-why-its-different) · [🎯 Two ways to use it](#-two-ways-to-use-it) · [🚀 Quickstart](#-quickstart) · [🧰 Toolbox](#-the-toolbox) · [🔌 Other clients](#-other-mcp-clients) · [🔒 Safety](#-safety-model) · [🩺 Troubleshooting](#-troubleshooting)
|
|
16
16
|
|
|
17
17
|
</div>
|
|
18
18
|
|
|
@@ -104,6 +104,28 @@ SceneScout needs only a URL. Give it the source code as well and it gets noticea
|
|
|
104
104
|
|
|
105
105
|
## 🚀 Quickstart
|
|
106
106
|
|
|
107
|
+
### ⚡ A first look, nothing to set up
|
|
108
|
+
|
|
109
|
+
Node 20 or newer and the address of an app you are allowed to test:
|
|
110
|
+
|
|
111
|
+
```bash
|
|
112
|
+
npx -y scenescout http://localhost:3000
|
|
113
|
+
```
|
|
114
|
+
|
|
115
|
+
It needs no model, no API key and no MCP client. It downloads the headless Chromium build if the machine has none (once, about 200 MB) and changes nothing else: no skill, no MCP registration, nothing on your PATH. Then it opens up to 20 pages in `observe` mode, where nothing but reads leaves the page (signing in and refreshing a token apart), starting none after 3 minutes. It measures each one the way [`scenescout check`](docs/guide/Ways-to-use-it.md#scenescout-check-a-gate-in-ci) does, writes `scenescout-report/report.md` in the current folder and prints the three issues to look at first:
|
|
116
|
+
|
|
117
|
+
```
|
|
118
|
+
Look at these first:
|
|
119
|
+
1. [medium] Request failed with a client error: GET /img/weekly-chart.png → HTTP 404 (on /)
|
|
120
|
+
2. [medium] Dead end: /reports-scheduled.html: 0 controls (on /reports-scheduled.html)
|
|
121
|
+
3. [medium] Control covered by pinned chrome: button "Save notes" is COVERED by pinned chrome [order-stickybar] at this scroll position — a click aimed at it lands on that element instead (on /order.html?id=1042)
|
|
122
|
+
|
|
123
|
+
12 pages looked at in 12 s in observe mode: 0 high · 6 medium · 2 low · 5 worth a look, never counted.
|
|
124
|
+
Report: scenescout-report/report.md
|
|
125
|
+
```
|
|
126
|
+
|
|
127
|
+
That is the [demo app](demo-app/). It exits 0 whatever it finds (a look, not a gate), and 2 when the address cannot be reached or the report cannot be written. After the address, `--max-routes` and `--max-minutes` raise the limits, `--mode read-only` lets a plain POST through, and `--out` names another folder. A `scenescout-report/` holding files a first look did not write is left alone, and no report it did not write is ever replaced. A first look only opens pages. To have your agent click, fill forms, compare roles and remember what it learned, set SceneScout up as below.
|
|
128
|
+
|
|
107
129
|
### 📦 Prerequisites
|
|
108
130
|
|
|
109
131
|
| | |
|
|
@@ -193,9 +215,9 @@ A browser window opens at the URL. Sign in however the app asks, then come back
|
|
|
193
215
|
|
|
194
216
|
Then `/scenescout --role admin`, or `scout_attach { role: "admin" }` from any agent. Every session attached with the same role gets its own browser built from that one login, so parallel lanes can all run as `admin`. A role with no saved login is refused with the command to run. `role` and `storageStatePath` are alternatives: pass one.
|
|
195
217
|
|
|
196
|
-
Sessions of one role share one saved login, so they share its refresh token too. An app that rotates refresh tokens and treats a second use of a spent one as theft would revoke the whole token family, and sign every session of that role out, the moment two of them refreshed with the same token. SceneScout stops that for a session attached by role. When the page is about to send a refresh token from the role's profile, the session first takes a lock beside the profile (`.scenescout/auth/<role>.json.lock`, owner-only, taken over if its holder has not touched it in 30 seconds). Holding the lock, it re-reads the profile: if another session has rotated the token in the meantime, it loads that profile into its own browser and sends the current token in place of the spent one. Once the page has stored the rotated token, the session writes its state back over the profile and releases the lock. Sessions in separate processes share the lock through the file. A refresh token is recognised by name (a cookie, a storage key, or a field inside a JSON storage value whose name contains `refresh`) and is never printed or logged. An app whose sign-in renews through the identity provider's own session cookie needs none of this, since no refresh token is shared. `SCENESCOUT_REFRESH_BROKER=off` turns the broker off.
|
|
218
|
+
Sessions of one role share one saved login, so they share its refresh token too. An app that rotates refresh tokens and treats a second use of a spent one as theft would revoke the whole token family, and sign every session of that role out, the moment two of them refreshed with the same token. SceneScout stops that for a session attached by role. When the page is about to send a refresh token from the role's profile, the session first takes a lock beside the profile (`.scenescout/auth/<role>.json.lock`, owner-only, taken over if its holder has not touched it in 30 seconds). Holding the lock, it re-reads the profile: if another session has rotated the token in the meantime, it loads that profile into its own browser and sends the current token in place of the spent one. Once the page has stored the rotated token, the session writes its state back over the profile and releases the lock. Sessions in separate processes share the lock through the file. A refresh token is recognised by name (a cookie, a storage key, or a field inside a JSON storage value whose name contains `refresh`) and is never printed or logged. A token that is only in a cookie makes a request a refresh only when the request is plausibly the refresh call, a POST to a path named for one or an endpoint seen to rotate the cookie, so a refresh cookie scoped to `/` never holds back the app's other requests; scripts, stylesheets, images and fonts are never brokered, and a broker that cannot do its job lets the request through unchanged. See [the guide](docs/guide/Signing-in.md#sessions-that-share-a-refresh-token). An app whose sign-in renews through the identity provider's own session cookie needs none of this, since no refresh token is shared. `SCENESCOUT_REFRESH_BROKER=off` turns the broker off.
|
|
197
219
|
|
|
198
|
-
In CI, where nobody can type, `--script` signs in headless as a test user from `SCENESCOUT_LOGIN_USERNAME`, `SCENESCOUT_LOGIN_PASSWORD` and, for a one-time code, `SCENESCOUT_LOGIN_TOTP_SECRET
|
|
220
|
+
In CI, where nobody can type, `--script` signs in headless as a test user from `SCENESCOUT_LOGIN_USERNAME`, `SCENESCOUT_LOGIN_PASSWORD` and, for a one-time code, `SCENESCOUT_LOGIN_TOTP_SECRET` or a fixed code the test environment accepts in `SCENESCOUT_LOGIN_OTP_CODE` (with no password for a passwordless sign-in), and saves the same profile. No credential value is ever printed. See [signing in from CI](docs/ci.md#signing-in-from-ci) for the options and the rules: a test tenant's user, never production or a real person's account.
|
|
199
221
|
|
|
200
222
|
Before a parallel run, `scout_lane_brief` checks that the planner's saved login will outlast it: `runMinutes` (default 60) plus `expiryMarginMinutes` (default 10). It refuses only when it is sure, meaning every credential in the profile has a date, none was set for another host, and the last of them ends before the run does, and then names the `scenescout login` command to run again. A profile holds cookies other than the sign-in (analytics, preferences), so the first one to expire is reported as a warning rather than a reason to refuse, and a profile with undated credentials in it (a session cookie, or a refresh token with no expiry) is a warning that its lifetime is unknown.
|
|
201
223
|
|
|
@@ -430,6 +452,7 @@ npx scenescout ci http://127.0.0.1:3000
|
|
|
430
452
|
- **Providers:** the Anthropic Messages API (default model `claude-sonnet-5`) or the OpenAI Responses API (default `gpt-6-luna`), chosen by which key is set; with both set, `--provider` decides. `--model` and `--effort` (default `low`) override; `--base-url` points at another endpoint that implements the same API.
|
|
431
453
|
- **Caps:** at most 40 model turns, 1,500,000 tokens and 20 minutes (`--max-turns`, `--max-tokens`, `--max-minutes`). The first cap reached ends the exploration; the report is still written, and says which cap ended it.
|
|
432
454
|
- **Mode:** `read-only` by default; `--mode observe` sends no form at all, `--mode safe-write` lets the run create records and change only the ones it created. `--mode destructive` runs only with `--allow-destructive` as well.
|
|
455
|
+
- **Duplicates:** when the dedup rule keeps a filed finding apart, the run's model is asked at its lowest effort whether it is one already open on the same page, and merges it on a "same", keeping the filing's title, category, severity and evidence under that finding. The two findings' titles, categories and evidence, and the page's path, are sent; `--dedup rule` turns it off ([ADR 17](docs/adr/0017-a-model-judges-only-the-merges-the-rule-misses.md)).
|
|
433
456
|
- **Output**, in `.scenescout/ci/` (or `--out`): `report.md` and `report.html` (the report an agent's run writes), `summary.md` (also appended to the GitHub job summary), `ci.json` and `ci.sarif`, with a usage line: turns, tokens, time and an estimated cost where the model's price is known (`--price-in`, `--price-out` give one for any model).
|
|
434
457
|
|
|
435
458
|
There is a GitHub Action for it (`uses: brunoboto96/SceneScout/ci@…`). [docs/ci.md](docs/ci.md#an-unattended-exploratory-run) has the workflow and every option; [ADR 14](docs/adr/0014-an-unattended-run-reports-and-never-gates.md) says why it works this way.
|
|
@@ -674,7 +697,7 @@ src/
|
|
|
674
697
|
memory.ts cross-run storage + finding dedup
|
|
675
698
|
profiles.ts saved sign-ins: role names, where a profile lives, owner-only files, attach by role, sessionStorage restore
|
|
676
699
|
refresh.ts the refresh broker: which values are a role's refresh tokens, the lock beside the profile, swapping a spent token
|
|
677
|
-
scripted-login.ts a CI sign-in: env and flags, TOTP (RFC 6238), which field is which, redaction
|
|
700
|
+
scripted-login.ts a CI sign-in: env and flags, TOTP (RFC 6238) or a fixed code, which field is which, redaction
|
|
678
701
|
expiry.ts how long a saved sign-in lasts: cookie dates and JWT exp, checked before lanes start
|
|
679
702
|
report.ts the gap ledger + report generation
|
|
680
703
|
check.ts the check's rules, gate, report and SARIF
|
package/dist/check-run.js
CHANGED
|
@@ -9,12 +9,10 @@ import fs from "node:fs";
|
|
|
9
9
|
import os from "node:os";
|
|
10
10
|
import path from "node:path";
|
|
11
11
|
import { BrowserEngine } from "./engine/browser.js";
|
|
12
|
-
import { checkFindings, redactFlowRuns, redactRoute, redactRoutes, settingsOf, withoutOwnResponse, } from "./engine/check.js";
|
|
12
|
+
import { checkFindings, MAX_DISCOVERY_ROUNDS, redactFlowRuns, redactRoute, redactRoutes, settingsOf, withoutOwnResponse, } from "./engine/check.js";
|
|
13
13
|
import { loadFlows, resolveFlowsDir } from "./engine/flow.js";
|
|
14
14
|
import { MemoryStore, MEMORY_DIRNAME } from "./engine/memory.js";
|
|
15
15
|
import { checkRetestPlan, retestResults, wellFormedFindings } from "./engine/verify.js";
|
|
16
|
-
/** Link discovery rounds: each crawl reveals the routes its pages link to. Past a few, a site is paginating rather than revealing. */
|
|
17
|
-
const MAX_ROUNDS = 6;
|
|
18
16
|
/**
|
|
19
17
|
* Read the flows and the findings, before any browser starts. Throws with a
|
|
20
18
|
* sentence naming the file and the field on anything it cannot read: a flow
|
|
@@ -43,11 +41,19 @@ export function readCheckInputs(options) {
|
|
|
43
41
|
return { flows, skippedFlows, findings: wellFormedFindings(findings ?? []) };
|
|
44
42
|
}
|
|
45
43
|
export async function runCheck(options, log = () => { }, inputs = { flows: [], findings: null }) {
|
|
44
|
+
if (options.paths && options.timeBudgetMs !== undefined)
|
|
45
|
+
throw new Error("a time budget applies to route discovery, not to a list of paths");
|
|
46
46
|
// A throwaway memory: a check is one run, and a store shared with earlier
|
|
47
47
|
// exploratory runs would count their visits as this check's and skip those routes.
|
|
48
48
|
const scratch = fs.mkdtempSync(path.join(os.tmpdir(), "scenescout-check-"));
|
|
49
|
+
// Also on an exit the finally below never reaches, such as Ctrl+C, which the browser's driver answers with process.exit.
|
|
50
|
+
const removeScratch = () => fs.rmSync(scratch, { recursive: true, force: true });
|
|
51
|
+
process.once("exit", removeScratch);
|
|
49
52
|
const engine = new BrowserEngine();
|
|
50
53
|
const start = new URL(options.url);
|
|
54
|
+
// From the start of the run, browser launch included: the budget is wall-clock time a person waits.
|
|
55
|
+
const deadline = options.timeBudgetMs !== undefined ? Date.now() + options.timeBudgetMs : undefined;
|
|
56
|
+
const pastDeadline = () => deadline !== undefined && Date.now() >= deadline;
|
|
51
57
|
try {
|
|
52
58
|
// Attached at the origin: the engine joins every crawled path onto the URL
|
|
53
59
|
// it attached to, so attaching to a start page with a path would double it.
|
|
@@ -68,23 +74,26 @@ export async function runCheck(options, log = () => { }, inputs = { flows: [], f
|
|
|
68
74
|
if (authFailed)
|
|
69
75
|
throw new Error(authFailed.replace(/ Continuing now tests a logged-out app\.$/, "").replace(/re-attach/, "run the check again"));
|
|
70
76
|
const routes = [];
|
|
71
|
-
const crawl = async (paths, limit) => {
|
|
72
|
-
await engine.crawl(paths, { inspect: true, limit });
|
|
77
|
+
const crawl = async (paths, limit, deadline) => {
|
|
78
|
+
await engine.crawl(paths, { inspect: true, limit, deadline });
|
|
73
79
|
routes.push(...engine.lastCrawlHealth);
|
|
74
80
|
};
|
|
75
81
|
if (options.paths) {
|
|
76
82
|
await crawl(options.paths.slice(0, options.maxRoutes), options.maxRoutes);
|
|
77
83
|
}
|
|
78
84
|
else {
|
|
79
|
-
// The start page first, whatever else is known: it is the one route the user named.
|
|
85
|
+
// The start page first, whatever else is known: it is the one route the user named. The
|
|
86
|
+
// time budget does not apply to it, so a run always measures at least the page it was given.
|
|
80
87
|
await crawl([`${start.pathname}${start.search}${start.hash}`], 1);
|
|
81
|
-
for (let round = 0; round <
|
|
82
|
-
if (engine.crawlableRoutes().length === 0)
|
|
88
|
+
for (let round = 0; round < MAX_DISCOVERY_ROUNDS && routes.length < options.maxRoutes; round++) {
|
|
89
|
+
if (engine.crawlableRoutes().length === 0 || pastDeadline())
|
|
83
90
|
break;
|
|
84
|
-
await crawl(undefined, options.maxRoutes - routes.length);
|
|
91
|
+
await crawl(undefined, options.maxRoutes - routes.length, deadline);
|
|
85
92
|
log(` ${routes.length} route(s) checked`);
|
|
86
93
|
}
|
|
87
94
|
}
|
|
95
|
+
// Out of time with routes still to visit; a cap of routes reached first is --max-routes's to report.
|
|
96
|
+
const timeLimitReached = pastDeadline() && routes.length < options.maxRoutes && engine.crawlableRoutes().length > 0;
|
|
88
97
|
// Pages of open findings the crawl did not load exactly: loaded now, so each re-test has its own measurement.
|
|
89
98
|
// Kept apart from `routes`: they are measured only to re-test, never checked against the page rules, and do not count
|
|
90
99
|
// towards --max-routes. With --paths the check stays on the paths it was given.
|
|
@@ -134,8 +143,9 @@ export async function runCheck(options, log = () => { }, inputs = { flows: [], f
|
|
|
134
143
|
routes: measured,
|
|
135
144
|
issues,
|
|
136
145
|
worthALook,
|
|
137
|
-
// Routes that failed to load are issues already; "not visited" is only what --max-routes left out.
|
|
146
|
+
// Routes that failed to load are issues already; "not visited" is only what --max-routes (or the time budget) left out.
|
|
138
147
|
unvisited: options.paths ? [] : engine.crawlableRoutes().map(redactRoute),
|
|
148
|
+
...(options.timeBudgetMs !== undefined ? { timeBudget: { ms: options.timeBudgetMs, reached: timeLimitReached } } : {}),
|
|
139
149
|
ignored: options.ignore,
|
|
140
150
|
flows,
|
|
141
151
|
skippedFlows: inputs.skippedFlows ?? [],
|
|
@@ -145,7 +155,8 @@ export async function runCheck(options, log = () => { }, inputs = { flows: [], f
|
|
|
145
155
|
}
|
|
146
156
|
finally {
|
|
147
157
|
await engine.close().catch((err) => log(`closing the browser failed: ${err instanceof Error ? err.message : String(err)}`));
|
|
148
|
-
|
|
158
|
+
process.off("exit", removeScratch);
|
|
159
|
+
removeScratch();
|
|
149
160
|
}
|
|
150
161
|
}
|
|
151
162
|
/** Where a check writes when not told: beside the project's other SceneScout output, which is already ignored by git. */
|
package/dist/ci-run.js
CHANGED
|
@@ -17,6 +17,8 @@ import path from "node:path";
|
|
|
17
17
|
import { fileURLToPath } from "node:url";
|
|
18
18
|
import { Client } from "@modelcontextprotocol/sdk/client/index.js";
|
|
19
19
|
import { StdioClientTransport } from "@modelcontextprotocol/sdk/client/stdio.js";
|
|
20
|
+
import { CreateMessageRequestSchema, ErrorCode, McpError } from "@modelcontextprotocol/sdk/types.js";
|
|
21
|
+
import { DEDUP_JUDGE_CAPABILITY, durationText, JUDGE_CALL_MS, JUDGE_MAX_OUTPUT_TOKENS, JUDGE_SYSTEM, JUDGE_TOOL, judgeKickoffOf, samplingResultOf, } from "./engine/dedup.js";
|
|
20
22
|
import { CAPTURE_MARGIN, parseCaptureResult, rebaseUrl, SHOT_FILES, SHOTS_DIRNAME } from "./engine/capture.js";
|
|
21
23
|
import { addUsage, wallLeftMs, guardToolArgs, capReached, CAPTURE_TOOLS, childEnv, ciCaptureKickoff, ciCaptureSystemPrompt, CI_DIRNAME, ciExitCode, ciKickoff, ciSarif, ciSummaryJson, ciSummaryMarkdown, ciSystemPrompt, ciToolArgs, ciTools, describeStop, findingsThisRun, NO_USAGE, readFindings, redactKeys, toolResultText, usageLine, } from "./engine/ci.js";
|
|
22
24
|
import { resolveTimeLimits } from "./engine/limits.js";
|
|
@@ -135,22 +137,88 @@ export class HttpModelClient {
|
|
|
135
137
|
this.conversation.addResults(results);
|
|
136
138
|
}
|
|
137
139
|
}
|
|
140
|
+
/** One conversation with the provider `resolved` names. */
|
|
141
|
+
function conversationFor(resolved, o, kickoff) {
|
|
142
|
+
return resolved.provider === "anthropic" ? new AnthropicConversation(o, kickoff) : new OpenAIConversation(o, kickoff);
|
|
143
|
+
}
|
|
138
144
|
export function httpClient(resolved, key, system, tools, kickoff) {
|
|
139
145
|
const o = { baseUrl: resolved.baseUrl, model: resolved.model, effort: resolved.effort, system, tools };
|
|
140
|
-
|
|
141
|
-
return new HttpModelClient(conversation, key);
|
|
146
|
+
return new HttpModelClient(conversationFor(resolved, o, kickoff), key);
|
|
142
147
|
}
|
|
143
|
-
|
|
144
|
-
|
|
145
|
-
|
|
146
|
-
|
|
147
|
-
|
|
148
|
-
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
|
|
152
|
-
|
|
153
|
-
|
|
148
|
+
/**
|
|
149
|
+
* The dedup judge's model call over HTTP: one short conversation per pair, at
|
|
150
|
+
* `resolved`'s effort, with its output capped and every attempt and retry
|
|
151
|
+
* inside `callMs`. Used by `scenescout ci` (the run's provider and model at
|
|
152
|
+
* the judge's effort) and by the MCP server when a key is in its environment.
|
|
153
|
+
*/
|
|
154
|
+
export function httpJudgeAsk(resolved, key, deps = {}) {
|
|
155
|
+
return async (system, tools, kickoff, limitMs) => {
|
|
156
|
+
const ms = Math.min(deps.callMs ?? JUDGE_CALL_MS, limitMs ?? Infinity);
|
|
157
|
+
const o = { baseUrl: resolved.baseUrl, model: resolved.model, effort: resolved.effort, system, tools, maxOutputTokens: JUDGE_MAX_OUTPUT_TOKENS };
|
|
158
|
+
try {
|
|
159
|
+
return await new HttpModelClient(conversationFor(resolved, o, kickoff), key, { fetch: deps.fetch, sleep: deps.sleep, attemptMs: ms }).next(ms);
|
|
160
|
+
}
|
|
161
|
+
catch (err) {
|
|
162
|
+
// The "time cap" here is the judge call's own limit, not the run's: say which.
|
|
163
|
+
if (err instanceof OutOfTime)
|
|
164
|
+
throw new Error(err.message.replace(/^the time cap was reached/, `no answer within ${durationText(ms)}`));
|
|
165
|
+
throw err;
|
|
166
|
+
}
|
|
167
|
+
};
|
|
168
|
+
}
|
|
169
|
+
/**
|
|
170
|
+
* The client's answer to the server's dedup judge (a sampling request): the
|
|
171
|
+
* question goes to the run's model through `ask`, under this side's own
|
|
172
|
+
* JUDGE_SYSTEM and JUDGE_TOOL, and the answer goes back as the tool call. Its
|
|
173
|
+
* tokens are added to the run's usage, so the caps count them, and no call
|
|
174
|
+
* runs past the time cap: one asked after it is refused, and one asked before
|
|
175
|
+
* it gets only the time left. A request that is not the judge's question is
|
|
176
|
+
* refused; a failed call is reported to the server, whose judge then leaves
|
|
177
|
+
* the rule's decision and logs it.
|
|
178
|
+
*/
|
|
179
|
+
export function judgeHandler(o) {
|
|
180
|
+
const now = o.now ?? Date.now;
|
|
181
|
+
return async (request) => {
|
|
182
|
+
o.calls.calls += 1;
|
|
183
|
+
const question = judgeKickoffOf(request.params);
|
|
184
|
+
if (!question.ok) {
|
|
185
|
+
o.calls.failed += 1;
|
|
186
|
+
throw new McpError(ErrorCode.InvalidRequest, `this client answers only the dedup judge's question: ${question.error}`);
|
|
187
|
+
}
|
|
188
|
+
const left = wallLeftMs(o.spend, o.caps, now());
|
|
189
|
+
if (left <= 0) {
|
|
190
|
+
o.calls.failed += 1;
|
|
191
|
+
throw new McpError(ErrorCode.InternalError, "the run's time cap was reached, so the judge was not asked");
|
|
192
|
+
}
|
|
193
|
+
const started = now();
|
|
194
|
+
try {
|
|
195
|
+
const turn = await o.ask(JUDGE_SYSTEM, [JUDGE_TOOL], question.kickoff, Math.min(JUDGE_CALL_MS, left));
|
|
196
|
+
o.spend.usage = addUsage(o.spend.usage, turn.usage);
|
|
197
|
+
o.calls.usage = addUsage(o.calls.usage, turn.usage);
|
|
198
|
+
return samplingResultOf(turn, o.model);
|
|
199
|
+
}
|
|
200
|
+
catch (err) {
|
|
201
|
+
o.calls.failed += 1;
|
|
202
|
+
throw new McpError(ErrorCode.InternalError, redactKeys(err instanceof Error ? err.message : String(err), o.secrets ?? []));
|
|
203
|
+
}
|
|
204
|
+
finally {
|
|
205
|
+
o.calls.ms += now() - started;
|
|
206
|
+
}
|
|
207
|
+
};
|
|
208
|
+
}
|
|
209
|
+
/**
|
|
210
|
+
* The run's MCP client. With a judge handler it declares sampling with tools
|
|
211
|
+
* and DEDUP_JUDGE_CAPABILITY, and answers the server's dedup judge with it;
|
|
212
|
+
* without one it declares nothing, and the server dedups by the rule.
|
|
213
|
+
*/
|
|
214
|
+
export function ciClient(judge) {
|
|
215
|
+
const client = new Client({ name: "scenescout-ci", version: "1" }, judge ? { capabilities: { sampling: { tools: {} }, experimental: { [DEDUP_JUDGE_CAPABILITY]: {} } } } : undefined);
|
|
216
|
+
if (judge)
|
|
217
|
+
client.setRequestHandler(CreateMessageRequestSchema, judge);
|
|
218
|
+
return client;
|
|
219
|
+
}
|
|
220
|
+
/** A connected client as the loop's tool host. */
|
|
221
|
+
export function toolHost(client) {
|
|
154
222
|
return {
|
|
155
223
|
tools: async () => (await client.listTools()).tools,
|
|
156
224
|
call: async (name, args, timeoutMs) => {
|
|
@@ -162,6 +230,19 @@ async function startServer(log) {
|
|
|
162
230
|
},
|
|
163
231
|
};
|
|
164
232
|
}
|
|
233
|
+
async function startServer(log, judge) {
|
|
234
|
+
if (!fs.existsSync(serverPath))
|
|
235
|
+
throw new Error(`${serverPath} is missing: run \`npm run build\` first`);
|
|
236
|
+
const transport = new StdioClientTransport({ command: process.execPath, args: [serverPath], env: childEnv(process.env), stderr: "pipe" });
|
|
237
|
+
transport.stderr?.on("data", (chunk) => {
|
|
238
|
+
for (const line of chunk.toString("utf8").split(/\r?\n/))
|
|
239
|
+
if (line.trim())
|
|
240
|
+
log(` [server] ${line}`);
|
|
241
|
+
});
|
|
242
|
+
const client = ciClient(judge);
|
|
243
|
+
await client.connect(transport);
|
|
244
|
+
return toolHost(client);
|
|
245
|
+
}
|
|
165
246
|
function readMemoryFindings(projectDir) {
|
|
166
247
|
try {
|
|
167
248
|
return readFindings(JSON.parse(fs.readFileSync(path.join(projectDir, MEMORY_DIRNAME, "memory.json"), "utf8")));
|
|
@@ -182,7 +263,7 @@ function readMemoryFindings(projectDir) {
|
|
|
182
263
|
*/
|
|
183
264
|
export async function agentLoop(o) {
|
|
184
265
|
const now = o.now ?? Date.now;
|
|
185
|
-
const spend = { turns: 0, usage: { ...NO_USAGE }, startedAt: o.startedAt ?? now() };
|
|
266
|
+
const spend = o.spend ?? { turns: 0, usage: { ...NO_USAGE }, startedAt: o.startedAt ?? now() };
|
|
186
267
|
const allowed = new Set(o.tools.map((t) => t.name));
|
|
187
268
|
for (;;) {
|
|
188
269
|
const cap = capReached(spend, o.caps, now());
|
|
@@ -350,7 +431,15 @@ export async function runCi(options, resolved, deps) {
|
|
|
350
431
|
const before = readMemoryFindings(options.projectDir);
|
|
351
432
|
// Pictures an earlier run left in the same output are not this run's: they must never be uploaded as its.
|
|
352
433
|
fs.rmSync(path.join(outDir, SHOTS_DIRNAME), { recursive: true, force: true });
|
|
353
|
-
|
|
434
|
+
// One spend for the run: the loop adds its turns, and the dedup judge's calls add their tokens, which the caps count.
|
|
435
|
+
const spend = { turns: 0, usage: { ...NO_USAGE }, startedAt };
|
|
436
|
+
// A run asked to show an element files no findings, so it has nothing to deduplicate.
|
|
437
|
+
const wantsJudge = options.dedup === "judge" && !options.show;
|
|
438
|
+
if (wantsJudge && !deps.judge)
|
|
439
|
+
log("No model was given for the dedup judge; the rule decides duplicates.");
|
|
440
|
+
const judgeAsk = wantsJudge ? deps.judge : undefined;
|
|
441
|
+
const judgeCalls = { calls: 0, failed: 0, usage: { ...NO_USAGE }, ms: 0 };
|
|
442
|
+
let outcome = { stop: "could-not-start", spend };
|
|
354
443
|
let contractMet = false;
|
|
355
444
|
let reportWritten = false;
|
|
356
445
|
let capture;
|
|
@@ -361,11 +450,13 @@ export async function runCi(options, resolved, deps) {
|
|
|
361
450
|
try {
|
|
362
451
|
// The page-load limit may be longer than the usual attach budget; the attach gets that limit and a minute to launch.
|
|
363
452
|
const attachMs = Math.max(ATTACH_MS, resolveTimeLimits(options, process.env).navMs + 60_000);
|
|
364
|
-
host = await startServer(log);
|
|
453
|
+
host = await startServer(log, judgeAsk ? judgeHandler({ ask: judgeAsk, model: resolved.model, spend, caps: options.caps, calls: judgeCalls, secrets, now }) : undefined);
|
|
365
454
|
const attached = await host.call("scout_attach", {
|
|
366
455
|
url: options.url,
|
|
367
456
|
projectPath: options.projectDir,
|
|
368
457
|
mode: options.mode,
|
|
458
|
+
// Named either way, so a SCENESCOUT_DEDUP in the job's environment never decides for the option.
|
|
459
|
+
dedup: judgeAsk ? "judge" : "rule",
|
|
369
460
|
objective: `CI run: explore at level ${options.level}${options.focus ? `, focusing on ${options.focus}` : ""}`.slice(0, 300),
|
|
370
461
|
task: "Starting the CI run",
|
|
371
462
|
...(options.storageStatePath ? { storageStatePath: options.storageStatePath } : {}),
|
|
@@ -390,7 +481,7 @@ export async function runCi(options, resolved, deps) {
|
|
|
390
481
|
caps: options.caps,
|
|
391
482
|
log,
|
|
392
483
|
now,
|
|
393
|
-
|
|
484
|
+
spend,
|
|
394
485
|
projectDir: options.projectDir,
|
|
395
486
|
onResult: (name, _args, r) => {
|
|
396
487
|
if (name === "scout_capture" && !r.isError)
|
|
@@ -449,6 +540,11 @@ export async function runCi(options, resolved, deps) {
|
|
|
449
540
|
endedAt,
|
|
450
541
|
findings: findingsThisRun(before, readMemoryFindings(options.projectDir)),
|
|
451
542
|
...(capture ? { capture } : {}),
|
|
543
|
+
...(options.show
|
|
544
|
+
? {}
|
|
545
|
+
: {
|
|
546
|
+
dedup: judgeAsk ? { by: "judge", ...(deps.judgeEffort ? { effort: deps.judgeEffort } : {}), ...judgeCalls } : { by: "rule" },
|
|
547
|
+
}),
|
|
452
548
|
};
|
|
453
549
|
const written = [];
|
|
454
550
|
try {
|