scenescout 3.14.0 → 3.15.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,30 @@
1
1
  # scenescout
2
2
 
3
+ ## 3.15.0
4
+
5
+ ### Minor Changes
6
+
7
+ - 3c488f3: Finding dedup can ask a model. When the dedup rule keeps a newly filed finding apart from everything recorded, a model judge is asked whether it is the same defect as one of the open findings on the same page (the three most alike, at most), and a "same" merges it; the merged filing's title, category, severity and evidence are kept under the finding and shown in the report with the judge's probability. Any failure, unsure answer or slow call leaves the rule's decision and is logged once per kind; three failed calls in a row switch the judge off for the run.
8
+
9
+ `scenescout ci` judges by default with the run's model at the lowest effort its API takes, sending each asked pair's titles, categories and evidence, and the page's path; the calls count in the run's usage, and `--dedup rule` (or the action's `dedup` input) turns it off. The MCP server judges only when `SCENESCOUT_DEDUP=judge` or `scout_attach {dedup: "judge"}` asks for it, with `ANTHROPIC_API_KEY` or `OPENAI_API_KEY` in its environment (`SCENESCOUT_DEDUP_PROVIDER` picks one when both are set).
10
+ - 47d5293: Add `scenescout <url>`, a first look at any app with nothing set up: `npx -y scenescout http://localhost:3000`. It runs the deterministic check's crawl and measurements in observe mode, where nothing but reads leaves the page (`--mode read-only` lets a plain POST through), visiting up to 20 pages and starting none after 3 minutes (`--max-routes`, `--max-minutes`), with no model and no API key. When the headless Chromium build is missing it downloads that and nothing else: no skill, no MCP registration, nothing on the PATH. It writes `report.md` and `check.json` to `scenescout-report/` in the current folder (a temporary folder when that cannot be written, or `--out`), and only into a folder that is new, empty or an earlier first look's: any other `scenescout-report/` is left as it is and the run exits 2, pointing at `--out`, and a `report.md` or `check.json` a first look did not write is never replaced; and the summary and the report open with the three issues to look at first: the highest severity, then the most pages affected, with one failure seen several ways (a missing image is a failed request, a broken image and a console line) taking one of the three places. It exits 0 once it has looked, whatever it found, and 2 when it could not run, such as when the address cannot be reached, or could not write the report. The guide's Start here page and the README's quickstart now lead with it.
11
+ - 400add2: `scenescout login --script` signs in to passwordless apps and takes a fixed one-time code. `SCENESCOUT_LOGIN_OTP_CODE` is a code the test environment accepts, typed where the form asks for one, as an alternative to `SCENESCOUT_LOGIN_TOTP_SECRET` (setting both is refused at startup); it is redacted from all output like the other credentials. `SCENESCOUT_LOGIN_PASSWORD` may be left unset when a code is given: the run fills the username, presses the button that sends the code ("Send code", "Continue"), waits through the moment with no field on screen, and types the code; a password field that appears anyway stops the run naming the variable.
12
+
13
+ Finding and submitting the code: a code split into one box per character is typed one character per box; a numeric field sized for a code is the code field once the username has gone; an email field whose label mentions a code stays the username. After typing, the run reads the page until it is clear how to go on, so a button enabled only once the form is complete is waited for, a page that takes the code itself is not submitted again, and a field the typing revealed is filled first. A button that signs in or verifies is pressed rather than one that sends a code where a page has both, and buttons that resend a code or change the address are never pressed. With no success URL or selector, a page with no sign-in field now counts as signed in only once a password or a code has gone, so the pause while a code is sent is not taken for a sign-in.
14
+ - 3471e73: Snapshots say more about the page. Alert and status regions are listed by their text whether or not they carry a test id, and the diff reports a region that says something new. Controls show their state (`pressed`, `selected`, `checked`, `expanded`, `current`) and the diff reports it moving; the per-control `done` marker is now `exercised`. A `main:` line summarises the main area's heading and static text, or says `main: EMPTY`, and crawl lines carry the same count. Accessible names follow the accessible-name computation's order: wrapping labels and `title` name a control, a button-like input is named by its value, and a select is never named by its options. Elements listed only for a test id (wrappers, headings, decorative badges) are no longer counted as unnamed controls or in the coverage denominator, and nothing inside an `aria-hidden` subtree is counted as unnamed. Controls held outside the visible width of a horizontally scrolling container are reported once per container, under the new worth-a-look check rule `scrolled-out-controls`.
15
+
16
+ ### Patch Changes
17
+
18
+ - f6997dc: The design audit's task-efficiency lines count only what they describe. Competing actions are buttons and links painted as buttons; text fields and breadcrumb links no longer count. Form burden counts the fields of a form; row-selection checkboxes, selects that edit a table row in place and search boxes no longer count, and when the page has a `<form>`, fields outside it no longer count either. The app shell's landmarks (navigation, banner, sidebar and footer regions outside the main content) are kept out of a page's score from the first audit, so the first pages of a run are no longer scored with the shell in them while later ones are scored without it. A shell built without landmarks is still recognised only once the shared-chrome census has seen it on several routes.
19
+ - c9edcac: Fewer false oracle violations. A `scout_request` probe the server refuses is no longer reported as an `http_error` or `console_error` of the page visited next. `false_success` now pairs a refused write only with a success message the action put on screen: a status badge or heading already there, a column header and a write the page sent in the background or to its own telemetry no longer count; a success message beside an action whose other writes went through is reported as a partial `false_success` at medium ("partial: N of M writes from this action were refused"); and an announced "was refused", "rejected" or "could not" counts as the page admitting the refusal. The silent-submit note matches its words ("sign", "post", "save") as whole words and waits for a client-side route change before calling a click silent. A page error raised by a link click that left the URL unchanged and opened a confirmation, or whose message says a route change was cancelled, is reported at medium with a note instead of high.
20
+ - a5f9c33: The refresh broker no longer holds back an app's requests when its refresh token is a cookie scoped to "/". A cookie alone no longer makes a request a refresh: only a POST, PUT or PATCH to a path named for a refresh, or an endpoint the broker has seen rotate the cookie (learned for every session of the role), goes through the lock. Scripts, stylesheets, images and fonts are never brokered, and when the broker cannot take the lock or read the profile the request goes out as the page sent it instead of being dropped. A refresh-named storage value that is an address is no longer taken for a token. On Windows, a lock file still being deleted as it changes hands (EPERM, EBUSY or EACCES) is waited for like a held lock instead of failing.
21
+
22
+ ## 3.14.1
23
+
24
+ ### Patch Changes
25
+
26
+ - 5358a81: `--help` and `-h` now print the usage and exit 0 on every subcommand before it does anything; `scenescout install --help` used to run a real install. `install`, `doctor`, `scan`, `status` and `watch` now refuse a flag or argument they do not know instead of ignoring it.
27
+
3
28
  ## 3.14.0
4
29
 
5
30
  ### Minor Changes
package/README.md CHANGED
@@ -12,7 +12,7 @@ Works with Claude Code · Cursor · VS Code (Copilot) · Codex CLI · Gemini CLI
12
12
  ![node >= 20](https://img.shields.io/badge/node-%E2%89%A5%2020-339933?logo=node.js&logoColor=white)
13
13
  ![MCP server](https://img.shields.io/badge/MCP-server-8A2BE2)
14
14
 
15
- [👀 See it work](#-see-it-work) · [✨ Why](#-why-its-different) · [🎯 Two ways to use it](#-two-ways-to-use-it) · [🚀 Quickstart](#-quickstart) · [🧰 Toolbox](#-the-toolbox) · [🔌 Other clients](#-other-mcp-clients) · [🔒 Safety](#-safety-model) · [🩺 Troubleshooting](#-troubleshooting)
15
+ [📖 Guide](docs/guide/Home.md) · [👀 See it work](#-see-it-work) · [✨ Why](#-why-its-different) · [🎯 Two ways to use it](#-two-ways-to-use-it) · [🚀 Quickstart](#-quickstart) · [🧰 Toolbox](#-the-toolbox) · [🔌 Other clients](#-other-mcp-clients) · [🔒 Safety](#-safety-model) · [🩺 Troubleshooting](#-troubleshooting)
16
16
 
17
17
  </div>
18
18
 
@@ -104,6 +104,28 @@ SceneScout needs only a URL. Give it the source code as well and it gets noticea
104
104
 
105
105
  ## 🚀 Quickstart
106
106
 
107
+ ### ⚡ A first look, nothing to set up
108
+
109
+ Node 20 or newer and the address of an app you are allowed to test:
110
+
111
+ ```bash
112
+ npx -y scenescout http://localhost:3000
113
+ ```
114
+
115
+ It needs no model, no API key and no MCP client. It downloads the headless Chromium build if the machine has none (once, about 200 MB) and changes nothing else: no skill, no MCP registration, nothing on your PATH. Then it opens up to 20 pages in `observe` mode, where nothing but reads leaves the page (signing in and refreshing a token apart), starting none after 3 minutes. It measures each one the way [`scenescout check`](docs/guide/Ways-to-use-it.md#scenescout-check-a-gate-in-ci) does, writes `scenescout-report/report.md` in the current folder and prints the three issues to look at first:
116
+
117
+ ```
118
+ Look at these first:
119
+ 1. [medium] Request failed with a client error: GET /img/weekly-chart.png → HTTP 404 (on /)
120
+ 2. [medium] Dead end: /reports-scheduled.html: 0 controls (on /reports-scheduled.html)
121
+ 3. [medium] Control covered by pinned chrome: button "Save notes" is COVERED by pinned chrome [order-stickybar] at this scroll position — a click aimed at it lands on that element instead (on /order.html?id=1042)
122
+
123
+ 12 pages looked at in 12 s in observe mode: 0 high · 6 medium · 2 low · 5 worth a look, never counted.
124
+ Report: scenescout-report/report.md
125
+ ```
126
+
127
+ That is the [demo app](demo-app/). It exits 0 whatever it finds (a look, not a gate), and 2 when the address cannot be reached or the report cannot be written. After the address, `--max-routes` and `--max-minutes` raise the limits, `--mode read-only` lets a plain POST through, and `--out` names another folder. A `scenescout-report/` holding files a first look did not write is left alone, and no report it did not write is ever replaced. A first look only opens pages. To have your agent click, fill forms, compare roles and remember what it learned, set SceneScout up as below.
128
+
107
129
  ### 📦 Prerequisites
108
130
 
109
131
  | | |
@@ -193,9 +215,9 @@ A browser window opens at the URL. Sign in however the app asks, then come back
193
215
 
194
216
  Then `/scenescout --role admin`, or `scout_attach { role: "admin" }` from any agent. Every session attached with the same role gets its own browser built from that one login, so parallel lanes can all run as `admin`. A role with no saved login is refused with the command to run. `role` and `storageStatePath` are alternatives: pass one.
195
217
 
196
- Sessions of one role share one saved login, so they share its refresh token too. An app that rotates refresh tokens and treats a second use of a spent one as theft would revoke the whole token family, and sign every session of that role out, the moment two of them refreshed with the same token. SceneScout stops that for a session attached by role. When the page is about to send a refresh token from the role's profile, the session first takes a lock beside the profile (`.scenescout/auth/<role>.json.lock`, owner-only, taken over if its holder has not touched it in 30 seconds). Holding the lock, it re-reads the profile: if another session has rotated the token in the meantime, it loads that profile into its own browser and sends the current token in place of the spent one. Once the page has stored the rotated token, the session writes its state back over the profile and releases the lock. Sessions in separate processes share the lock through the file. A refresh token is recognised by name (a cookie, a storage key, or a field inside a JSON storage value whose name contains `refresh`) and is never printed or logged. An app whose sign-in renews through the identity provider's own session cookie needs none of this, since no refresh token is shared. `SCENESCOUT_REFRESH_BROKER=off` turns the broker off.
218
+ Sessions of one role share one saved login, so they share its refresh token too. An app that rotates refresh tokens and treats a second use of a spent one as theft would revoke the whole token family, and sign every session of that role out, the moment two of them refreshed with the same token. SceneScout stops that for a session attached by role. When the page is about to send a refresh token from the role's profile, the session first takes a lock beside the profile (`.scenescout/auth/<role>.json.lock`, owner-only, taken over if its holder has not touched it in 30 seconds). Holding the lock, it re-reads the profile: if another session has rotated the token in the meantime, it loads that profile into its own browser and sends the current token in place of the spent one. Once the page has stored the rotated token, the session writes its state back over the profile and releases the lock. Sessions in separate processes share the lock through the file. A refresh token is recognised by name (a cookie, a storage key, or a field inside a JSON storage value whose name contains `refresh`) and is never printed or logged. A token that is only in a cookie makes a request a refresh only when the request is plausibly the refresh call, a POST to a path named for one or an endpoint seen to rotate the cookie, so a refresh cookie scoped to `/` never holds back the app's other requests; scripts, stylesheets, images and fonts are never brokered, and a broker that cannot do its job lets the request through unchanged. See [the guide](docs/guide/Signing-in.md#sessions-that-share-a-refresh-token). An app whose sign-in renews through the identity provider's own session cookie needs none of this, since no refresh token is shared. `SCENESCOUT_REFRESH_BROKER=off` turns the broker off.
197
219
 
198
- In CI, where nobody can type, `--script` signs in headless as a test user from `SCENESCOUT_LOGIN_USERNAME`, `SCENESCOUT_LOGIN_PASSWORD` and, for a one-time code, `SCENESCOUT_LOGIN_TOTP_SECRET`, and saves the same profile. No credential value is ever printed. See [signing in from CI](docs/ci.md#signing-in-from-ci) for the options and the rules: a test tenant's user, never production or a real person's account.
220
+ In CI, where nobody can type, `--script` signs in headless as a test user from `SCENESCOUT_LOGIN_USERNAME`, `SCENESCOUT_LOGIN_PASSWORD` and, for a one-time code, `SCENESCOUT_LOGIN_TOTP_SECRET` or a fixed code the test environment accepts in `SCENESCOUT_LOGIN_OTP_CODE` (with no password for a passwordless sign-in), and saves the same profile. No credential value is ever printed. See [signing in from CI](docs/ci.md#signing-in-from-ci) for the options and the rules: a test tenant's user, never production or a real person's account.
199
221
 
200
222
  Before a parallel run, `scout_lane_brief` checks that the planner's saved login will outlast it: `runMinutes` (default 60) plus `expiryMarginMinutes` (default 10). It refuses only when it is sure, meaning every credential in the profile has a date, none was set for another host, and the last of them ends before the run does, and then names the `scenescout login` command to run again. A profile holds cookies other than the sign-in (analytics, preferences), so the first one to expire is reported as a warning rather than a reason to refuse, and a profile with undated credentials in it (a session cookie, or a refresh token with no expiry) is a warning that its lifetime is unknown.
201
223
 
@@ -430,6 +452,7 @@ npx scenescout ci http://127.0.0.1:3000
430
452
  - **Providers:** the Anthropic Messages API (default model `claude-sonnet-5`) or the OpenAI Responses API (default `gpt-6-luna`), chosen by which key is set; with both set, `--provider` decides. `--model` and `--effort` (default `low`) override; `--base-url` points at another endpoint that implements the same API.
431
453
  - **Caps:** at most 40 model turns, 1,500,000 tokens and 20 minutes (`--max-turns`, `--max-tokens`, `--max-minutes`). The first cap reached ends the exploration; the report is still written, and says which cap ended it.
432
454
  - **Mode:** `read-only` by default; `--mode observe` sends no form at all, `--mode safe-write` lets the run create records and change only the ones it created. `--mode destructive` runs only with `--allow-destructive` as well.
455
+ - **Duplicates:** when the dedup rule keeps a filed finding apart, the run's model is asked at its lowest effort whether it is one already open on the same page, and merges it on a "same", keeping the filing's title, category, severity and evidence under that finding. The two findings' titles, categories and evidence, and the page's path, are sent; `--dedup rule` turns it off ([ADR 17](docs/adr/0017-a-model-judges-only-the-merges-the-rule-misses.md)).
433
456
  - **Output**, in `.scenescout/ci/` (or `--out`): `report.md` and `report.html` (the report an agent's run writes), `summary.md` (also appended to the GitHub job summary), `ci.json` and `ci.sarif`, with a usage line: turns, tokens, time and an estimated cost where the model's price is known (`--price-in`, `--price-out` give one for any model).
434
457
 
435
458
  There is a GitHub Action for it (`uses: brunoboto96/SceneScout/ci@…`). [docs/ci.md](docs/ci.md#an-unattended-exploratory-run) has the workflow and every option; [ADR 14](docs/adr/0014-an-unattended-run-reports-and-never-gates.md) says why it works this way.
@@ -674,7 +697,7 @@ src/
674
697
  memory.ts cross-run storage + finding dedup
675
698
  profiles.ts saved sign-ins: role names, where a profile lives, owner-only files, attach by role, sessionStorage restore
676
699
  refresh.ts the refresh broker: which values are a role's refresh tokens, the lock beside the profile, swapping a spent token
677
- scripted-login.ts a CI sign-in: env and flags, TOTP (RFC 6238), which field is which, redaction
700
+ scripted-login.ts a CI sign-in: env and flags, TOTP (RFC 6238) or a fixed code, which field is which, redaction
678
701
  expiry.ts how long a saved sign-in lasts: cookie dates and JWT exp, checked before lanes start
679
702
  report.ts the gap ledger + report generation
680
703
  check.ts the check's rules, gate, report and SARIF
package/dist/check-run.js CHANGED
@@ -9,12 +9,10 @@ import fs from "node:fs";
9
9
  import os from "node:os";
10
10
  import path from "node:path";
11
11
  import { BrowserEngine } from "./engine/browser.js";
12
- import { checkFindings, redactFlowRuns, redactRoute, redactRoutes, settingsOf, withoutOwnResponse, } from "./engine/check.js";
12
+ import { checkFindings, MAX_DISCOVERY_ROUNDS, redactFlowRuns, redactRoute, redactRoutes, settingsOf, withoutOwnResponse, } from "./engine/check.js";
13
13
  import { loadFlows, resolveFlowsDir } from "./engine/flow.js";
14
14
  import { MemoryStore, MEMORY_DIRNAME } from "./engine/memory.js";
15
15
  import { checkRetestPlan, retestResults, wellFormedFindings } from "./engine/verify.js";
16
- /** Link discovery rounds: each crawl reveals the routes its pages link to. Past a few, a site is paginating rather than revealing. */
17
- const MAX_ROUNDS = 6;
18
16
  /**
19
17
  * Read the flows and the findings, before any browser starts. Throws with a
20
18
  * sentence naming the file and the field on anything it cannot read: a flow
@@ -43,11 +41,19 @@ export function readCheckInputs(options) {
43
41
  return { flows, skippedFlows, findings: wellFormedFindings(findings ?? []) };
44
42
  }
45
43
  export async function runCheck(options, log = () => { }, inputs = { flows: [], findings: null }) {
44
+ if (options.paths && options.timeBudgetMs !== undefined)
45
+ throw new Error("a time budget applies to route discovery, not to a list of paths");
46
46
  // A throwaway memory: a check is one run, and a store shared with earlier
47
47
  // exploratory runs would count their visits as this check's and skip those routes.
48
48
  const scratch = fs.mkdtempSync(path.join(os.tmpdir(), "scenescout-check-"));
49
+ // Also on an exit the finally below never reaches, such as Ctrl+C, which the browser's driver answers with process.exit.
50
+ const removeScratch = () => fs.rmSync(scratch, { recursive: true, force: true });
51
+ process.once("exit", removeScratch);
49
52
  const engine = new BrowserEngine();
50
53
  const start = new URL(options.url);
54
+ // From the start of the run, browser launch included: the budget is wall-clock time a person waits.
55
+ const deadline = options.timeBudgetMs !== undefined ? Date.now() + options.timeBudgetMs : undefined;
56
+ const pastDeadline = () => deadline !== undefined && Date.now() >= deadline;
51
57
  try {
52
58
  // Attached at the origin: the engine joins every crawled path onto the URL
53
59
  // it attached to, so attaching to a start page with a path would double it.
@@ -68,23 +74,26 @@ export async function runCheck(options, log = () => { }, inputs = { flows: [], f
68
74
  if (authFailed)
69
75
  throw new Error(authFailed.replace(/ Continuing now tests a logged-out app\.$/, "").replace(/re-attach/, "run the check again"));
70
76
  const routes = [];
71
- const crawl = async (paths, limit) => {
72
- await engine.crawl(paths, { inspect: true, limit });
77
+ const crawl = async (paths, limit, deadline) => {
78
+ await engine.crawl(paths, { inspect: true, limit, deadline });
73
79
  routes.push(...engine.lastCrawlHealth);
74
80
  };
75
81
  if (options.paths) {
76
82
  await crawl(options.paths.slice(0, options.maxRoutes), options.maxRoutes);
77
83
  }
78
84
  else {
79
- // The start page first, whatever else is known: it is the one route the user named.
85
+ // The start page first, whatever else is known: it is the one route the user named. The
86
+ // time budget does not apply to it, so a run always measures at least the page it was given.
80
87
  await crawl([`${start.pathname}${start.search}${start.hash}`], 1);
81
- for (let round = 0; round < MAX_ROUNDS && routes.length < options.maxRoutes; round++) {
82
- if (engine.crawlableRoutes().length === 0)
88
+ for (let round = 0; round < MAX_DISCOVERY_ROUNDS && routes.length < options.maxRoutes; round++) {
89
+ if (engine.crawlableRoutes().length === 0 || pastDeadline())
83
90
  break;
84
- await crawl(undefined, options.maxRoutes - routes.length);
91
+ await crawl(undefined, options.maxRoutes - routes.length, deadline);
85
92
  log(` ${routes.length} route(s) checked`);
86
93
  }
87
94
  }
95
+ // Out of time with routes still to visit; a cap of routes reached first is --max-routes's to report.
96
+ const timeLimitReached = pastDeadline() && routes.length < options.maxRoutes && engine.crawlableRoutes().length > 0;
88
97
  // Pages of open findings the crawl did not load exactly: loaded now, so each re-test has its own measurement.
89
98
  // Kept apart from `routes`: they are measured only to re-test, never checked against the page rules, and do not count
90
99
  // towards --max-routes. With --paths the check stays on the paths it was given.
@@ -134,8 +143,9 @@ export async function runCheck(options, log = () => { }, inputs = { flows: [], f
134
143
  routes: measured,
135
144
  issues,
136
145
  worthALook,
137
- // Routes that failed to load are issues already; "not visited" is only what --max-routes left out.
146
+ // Routes that failed to load are issues already; "not visited" is only what --max-routes (or the time budget) left out.
138
147
  unvisited: options.paths ? [] : engine.crawlableRoutes().map(redactRoute),
148
+ ...(options.timeBudgetMs !== undefined ? { timeBudget: { ms: options.timeBudgetMs, reached: timeLimitReached } } : {}),
139
149
  ignored: options.ignore,
140
150
  flows,
141
151
  skippedFlows: inputs.skippedFlows ?? [],
@@ -145,7 +155,8 @@ export async function runCheck(options, log = () => { }, inputs = { flows: [], f
145
155
  }
146
156
  finally {
147
157
  await engine.close().catch((err) => log(`closing the browser failed: ${err instanceof Error ? err.message : String(err)}`));
148
- fs.rmSync(scratch, { recursive: true, force: true });
158
+ process.off("exit", removeScratch);
159
+ removeScratch();
149
160
  }
150
161
  }
151
162
  /** Where a check writes when not told: beside the project's other SceneScout output, which is already ignored by git. */
package/dist/ci-run.js CHANGED
@@ -17,6 +17,8 @@ import path from "node:path";
17
17
  import { fileURLToPath } from "node:url";
18
18
  import { Client } from "@modelcontextprotocol/sdk/client/index.js";
19
19
  import { StdioClientTransport } from "@modelcontextprotocol/sdk/client/stdio.js";
20
+ import { CreateMessageRequestSchema, ErrorCode, McpError } from "@modelcontextprotocol/sdk/types.js";
21
+ import { DEDUP_JUDGE_CAPABILITY, durationText, JUDGE_CALL_MS, JUDGE_MAX_OUTPUT_TOKENS, JUDGE_SYSTEM, JUDGE_TOOL, judgeKickoffOf, samplingResultOf, } from "./engine/dedup.js";
20
22
  import { CAPTURE_MARGIN, parseCaptureResult, rebaseUrl, SHOT_FILES, SHOTS_DIRNAME } from "./engine/capture.js";
21
23
  import { addUsage, wallLeftMs, guardToolArgs, capReached, CAPTURE_TOOLS, childEnv, ciCaptureKickoff, ciCaptureSystemPrompt, CI_DIRNAME, ciExitCode, ciKickoff, ciSarif, ciSummaryJson, ciSummaryMarkdown, ciSystemPrompt, ciToolArgs, ciTools, describeStop, findingsThisRun, NO_USAGE, readFindings, redactKeys, toolResultText, usageLine, } from "./engine/ci.js";
22
24
  import { resolveTimeLimits } from "./engine/limits.js";
@@ -135,22 +137,88 @@ export class HttpModelClient {
135
137
  this.conversation.addResults(results);
136
138
  }
137
139
  }
140
+ /** One conversation with the provider `resolved` names. */
141
+ function conversationFor(resolved, o, kickoff) {
142
+ return resolved.provider === "anthropic" ? new AnthropicConversation(o, kickoff) : new OpenAIConversation(o, kickoff);
143
+ }
138
144
  export function httpClient(resolved, key, system, tools, kickoff) {
139
145
  const o = { baseUrl: resolved.baseUrl, model: resolved.model, effort: resolved.effort, system, tools };
140
- const conversation = resolved.provider === "anthropic" ? new AnthropicConversation(o, kickoff) : new OpenAIConversation(o, kickoff);
141
- return new HttpModelClient(conversation, key);
146
+ return new HttpModelClient(conversationFor(resolved, o, kickoff), key);
142
147
  }
143
- async function startServer(log) {
144
- if (!fs.existsSync(serverPath))
145
- throw new Error(`${serverPath} is missing: run \`npm run build\` first`);
146
- const transport = new StdioClientTransport({ command: process.execPath, args: [serverPath], env: childEnv(process.env), stderr: "pipe" });
147
- transport.stderr?.on("data", (chunk) => {
148
- for (const line of chunk.toString("utf8").split(/\r?\n/))
149
- if (line.trim())
150
- log(` [server] ${line}`);
151
- });
152
- const client = new Client({ name: "scenescout-ci", version: "1" });
153
- await client.connect(transport);
148
+ /**
149
+ * The dedup judge's model call over HTTP: one short conversation per pair, at
150
+ * `resolved`'s effort, with its output capped and every attempt and retry
151
+ * inside `callMs`. Used by `scenescout ci` (the run's provider and model at
152
+ * the judge's effort) and by the MCP server when a key is in its environment.
153
+ */
154
+ export function httpJudgeAsk(resolved, key, deps = {}) {
155
+ return async (system, tools, kickoff, limitMs) => {
156
+ const ms = Math.min(deps.callMs ?? JUDGE_CALL_MS, limitMs ?? Infinity);
157
+ const o = { baseUrl: resolved.baseUrl, model: resolved.model, effort: resolved.effort, system, tools, maxOutputTokens: JUDGE_MAX_OUTPUT_TOKENS };
158
+ try {
159
+ return await new HttpModelClient(conversationFor(resolved, o, kickoff), key, { fetch: deps.fetch, sleep: deps.sleep, attemptMs: ms }).next(ms);
160
+ }
161
+ catch (err) {
162
+ // The "time cap" here is the judge call's own limit, not the run's: say which.
163
+ if (err instanceof OutOfTime)
164
+ throw new Error(err.message.replace(/^the time cap was reached/, `no answer within ${durationText(ms)}`));
165
+ throw err;
166
+ }
167
+ };
168
+ }
169
+ /**
170
+ * The client's answer to the server's dedup judge (a sampling request): the
171
+ * question goes to the run's model through `ask`, under this side's own
172
+ * JUDGE_SYSTEM and JUDGE_TOOL, and the answer goes back as the tool call. Its
173
+ * tokens are added to the run's usage, so the caps count them, and no call
174
+ * runs past the time cap: one asked after it is refused, and one asked before
175
+ * it gets only the time left. A request that is not the judge's question is
176
+ * refused; a failed call is reported to the server, whose judge then leaves
177
+ * the rule's decision and logs it.
178
+ */
179
+ export function judgeHandler(o) {
180
+ const now = o.now ?? Date.now;
181
+ return async (request) => {
182
+ o.calls.calls += 1;
183
+ const question = judgeKickoffOf(request.params);
184
+ if (!question.ok) {
185
+ o.calls.failed += 1;
186
+ throw new McpError(ErrorCode.InvalidRequest, `this client answers only the dedup judge's question: ${question.error}`);
187
+ }
188
+ const left = wallLeftMs(o.spend, o.caps, now());
189
+ if (left <= 0) {
190
+ o.calls.failed += 1;
191
+ throw new McpError(ErrorCode.InternalError, "the run's time cap was reached, so the judge was not asked");
192
+ }
193
+ const started = now();
194
+ try {
195
+ const turn = await o.ask(JUDGE_SYSTEM, [JUDGE_TOOL], question.kickoff, Math.min(JUDGE_CALL_MS, left));
196
+ o.spend.usage = addUsage(o.spend.usage, turn.usage);
197
+ o.calls.usage = addUsage(o.calls.usage, turn.usage);
198
+ return samplingResultOf(turn, o.model);
199
+ }
200
+ catch (err) {
201
+ o.calls.failed += 1;
202
+ throw new McpError(ErrorCode.InternalError, redactKeys(err instanceof Error ? err.message : String(err), o.secrets ?? []));
203
+ }
204
+ finally {
205
+ o.calls.ms += now() - started;
206
+ }
207
+ };
208
+ }
209
+ /**
210
+ * The run's MCP client. With a judge handler it declares sampling with tools
211
+ * and DEDUP_JUDGE_CAPABILITY, and answers the server's dedup judge with it;
212
+ * without one it declares nothing, and the server dedups by the rule.
213
+ */
214
+ export function ciClient(judge) {
215
+ const client = new Client({ name: "scenescout-ci", version: "1" }, judge ? { capabilities: { sampling: { tools: {} }, experimental: { [DEDUP_JUDGE_CAPABILITY]: {} } } } : undefined);
216
+ if (judge)
217
+ client.setRequestHandler(CreateMessageRequestSchema, judge);
218
+ return client;
219
+ }
220
+ /** A connected client as the loop's tool host. */
221
+ export function toolHost(client) {
154
222
  return {
155
223
  tools: async () => (await client.listTools()).tools,
156
224
  call: async (name, args, timeoutMs) => {
@@ -162,6 +230,19 @@ async function startServer(log) {
162
230
  },
163
231
  };
164
232
  }
233
+ async function startServer(log, judge) {
234
+ if (!fs.existsSync(serverPath))
235
+ throw new Error(`${serverPath} is missing: run \`npm run build\` first`);
236
+ const transport = new StdioClientTransport({ command: process.execPath, args: [serverPath], env: childEnv(process.env), stderr: "pipe" });
237
+ transport.stderr?.on("data", (chunk) => {
238
+ for (const line of chunk.toString("utf8").split(/\r?\n/))
239
+ if (line.trim())
240
+ log(` [server] ${line}`);
241
+ });
242
+ const client = ciClient(judge);
243
+ await client.connect(transport);
244
+ return toolHost(client);
245
+ }
165
246
  function readMemoryFindings(projectDir) {
166
247
  try {
167
248
  return readFindings(JSON.parse(fs.readFileSync(path.join(projectDir, MEMORY_DIRNAME, "memory.json"), "utf8")));
@@ -182,7 +263,7 @@ function readMemoryFindings(projectDir) {
182
263
  */
183
264
  export async function agentLoop(o) {
184
265
  const now = o.now ?? Date.now;
185
- const spend = { turns: 0, usage: { ...NO_USAGE }, startedAt: o.startedAt ?? now() };
266
+ const spend = o.spend ?? { turns: 0, usage: { ...NO_USAGE }, startedAt: o.startedAt ?? now() };
186
267
  const allowed = new Set(o.tools.map((t) => t.name));
187
268
  for (;;) {
188
269
  const cap = capReached(spend, o.caps, now());
@@ -350,7 +431,15 @@ export async function runCi(options, resolved, deps) {
350
431
  const before = readMemoryFindings(options.projectDir);
351
432
  // Pictures an earlier run left in the same output are not this run's: they must never be uploaded as its.
352
433
  fs.rmSync(path.join(outDir, SHOTS_DIRNAME), { recursive: true, force: true });
353
- let outcome = { stop: "could-not-start", spend: { turns: 0, usage: { ...NO_USAGE }, startedAt } };
434
+ // One spend for the run: the loop adds its turns, and the dedup judge's calls add their tokens, which the caps count.
435
+ const spend = { turns: 0, usage: { ...NO_USAGE }, startedAt };
436
+ // A run asked to show an element files no findings, so it has nothing to deduplicate.
437
+ const wantsJudge = options.dedup === "judge" && !options.show;
438
+ if (wantsJudge && !deps.judge)
439
+ log("No model was given for the dedup judge; the rule decides duplicates.");
440
+ const judgeAsk = wantsJudge ? deps.judge : undefined;
441
+ const judgeCalls = { calls: 0, failed: 0, usage: { ...NO_USAGE }, ms: 0 };
442
+ let outcome = { stop: "could-not-start", spend };
354
443
  let contractMet = false;
355
444
  let reportWritten = false;
356
445
  let capture;
@@ -361,11 +450,13 @@ export async function runCi(options, resolved, deps) {
361
450
  try {
362
451
  // The page-load limit may be longer than the usual attach budget; the attach gets that limit and a minute to launch.
363
452
  const attachMs = Math.max(ATTACH_MS, resolveTimeLimits(options, process.env).navMs + 60_000);
364
- host = await startServer(log);
453
+ host = await startServer(log, judgeAsk ? judgeHandler({ ask: judgeAsk, model: resolved.model, spend, caps: options.caps, calls: judgeCalls, secrets, now }) : undefined);
365
454
  const attached = await host.call("scout_attach", {
366
455
  url: options.url,
367
456
  projectPath: options.projectDir,
368
457
  mode: options.mode,
458
+ // Named either way, so a SCENESCOUT_DEDUP in the job's environment never decides for the option.
459
+ dedup: judgeAsk ? "judge" : "rule",
369
460
  objective: `CI run: explore at level ${options.level}${options.focus ? `, focusing on ${options.focus}` : ""}`.slice(0, 300),
370
461
  task: "Starting the CI run",
371
462
  ...(options.storageStatePath ? { storageStatePath: options.storageStatePath } : {}),
@@ -390,7 +481,7 @@ export async function runCi(options, resolved, deps) {
390
481
  caps: options.caps,
391
482
  log,
392
483
  now,
393
- startedAt,
484
+ spend,
394
485
  projectDir: options.projectDir,
395
486
  onResult: (name, _args, r) => {
396
487
  if (name === "scout_capture" && !r.isError)
@@ -449,6 +540,11 @@ export async function runCi(options, resolved, deps) {
449
540
  endedAt,
450
541
  findings: findingsThisRun(before, readMemoryFindings(options.projectDir)),
451
542
  ...(capture ? { capture } : {}),
543
+ ...(options.show
544
+ ? {}
545
+ : {
546
+ dedup: judgeAsk ? { by: "judge", ...(deps.judgeEffort ? { effort: deps.judgeEffort } : {}), ...judgeCalls } : { by: "rule" },
547
+ }),
452
548
  };
453
549
  const written = [];
454
550
  try {