scenescout 3.21.3 → 3.23.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,17 @@
1
1
  # scenescout
2
2
 
3
+ ## 3.23.0
4
+
5
+ ### Minor Changes
6
+
7
+ - ed24973: Start a run from an earlier run's record, opt-in. Every run that writes its report now keeps a record of what it worked on and left (routes in order, each session's steps, and on each route the controls never exercised, forms never submitted and options never chosen), in the project's memory and in `ci.json` under `record`. `scenescout ci --from-run <ci.json or project directory>` (or `SCENESCOUT_FROM_RUN`) with `--from-run-mode continue` (the default) takes first the routes that run never worked on, then the ones it left work on, told exactly what to do first on each, then the rest, reaching a page the earlier run got to by acting on another page through the same steps; a run takes on only its budget's worth of pages (`--from-run-turns-per-page`, default 7 turns a page) and the next continued run takes the pages after them; `--from-run-mode replay` follows its routes and steps in order, one lane per session it had. `scout_lane_brief` takes `fromRun` and `fromRunMode`, the ci action takes `from-run` and `from-run-mode`, and the run it started from is named in the report, `summary.md` and `ci.json` (`fromRun`). Without it, runs behave as before.
8
+
9
+ ## 3.22.0
10
+
11
+ ### Minor Changes
12
+
13
+ - 072b68b: `scenescout check --record --template <file.json>` (the action's `template` input) also writes the recorded run up as a test report laid out by the template: each saved flow a test and each step a row with its expected and actual result and its frame, failed steps as deviations, a SHA-256 manifest of the evidence and blank sign-off rows. Flows gain optional `id`, `requirements` and a step's `expected` for traceability; an example template is in `examples/report-template.json`.
14
+
3
15
  ## 3.21.3
4
16
 
5
17
  ### Patch Changes
package/README.md CHANGED
@@ -279,7 +279,10 @@ Use SceneScout to test http://localhost:3000, record the run
279
279
  or, on the tool directly, `scout_attach {record: true}`. `SCENESCOUT_RECORD=on` in
280
280
  the server's environment records every run. A CI gate records too:
281
281
  `scenescout check --record` writes `replay.html`, every journey step by step with
282
- its frames ([recording a check](docs/ci.md#recording-a-check)).
282
+ its frames ([recording a check](docs/ci.md#recording-a-check)), and
283
+ `--template <file.json>` also writes it up as a test report laid out as a template
284
+ says, with expected and actual results, deviations, blank sign-off rows and a
285
+ SHA-256 manifest of the evidence ([test reports](docs/ci.md#a-test-report-from-a-template)).
283
286
 
284
287
  Then `scout_report` writes two files side by side in `.scenescout/`:
285
288
  `report.md` as always, and `report.html` — the whole run as one self-contained
@@ -482,6 +485,7 @@ npx scenescout ci http://127.0.0.1:3000
482
485
  - **Caps:** at most 80 model turns, 3,000,000 tokens and 20 minutes (`--max-turns`, `--max-tokens`, `--max-minutes`), shared by two model loops that each explore their own part of the app (`--lanes`, default 2). The first cap reached ends the exploration; the report is still written, and says which cap ended it. On the benchmark's demo app a run at these defaults cost about $0.03 on `gpt-6-luna` ([docs/benchmark.md](docs/benchmark.md#choosing-the-defaults-issue-419)).
483
486
  - **Mode:** `read-only` by default; `--mode observe` sends no form at all, `--mode safe-write` lets the run create records and change only the ones it created. `--mode destructive` runs only with `--allow-destructive` as well.
484
487
  - **Duplicates:** when the dedup rule keeps a filed finding apart, the run's model is asked at its lowest effort whether it is one already open on the same page, and merges it on a "same", keeping the filing's title, category, severity and evidence under that finding. The two findings' titles, categories and evidence, and the page's path, are sent; `--dedup rule` turns it off ([ADR 17](docs/adr/0017-a-model-judges-only-the-merges-the-rule-misses.md)).
488
+ - **Starting from an earlier run (opt-in):** every run that writes its report leaves a record in `ci.json` and the project's memory: the routes it worked on, its steps, and what it left on each route. `--from-run <ci.json or project directory>` continues where that run left off (routes it never worked on first, then the ones it left work on, with exactly which forms, options and controls to take first), and `--from-run-mode replay` follows its routes and steps in order. `scout_lane_brief {fromRun, fromRunMode}` does the same for a parallel run. On the demo app a chain of continued runs found 8 defects against 7 for fresh runs at the default budget, and 4, 3 and 4 against 4 at a budget cut to stand in for a larger app, within the noise, so it stays off unless asked for ([the measurement](docs/benchmark.md#starting-from-an-earlier-run-issue-418)).
485
489
  - **Output**, in `.scenescout/ci/` (or `--out`): `report.md` and `report.html` (the report an agent's run writes), `summary.md` (also appended to the GitHub job summary), `ci.json` and `ci.sarif`, with a usage line: turns, tokens, time and an estimated cost where the model's price is known (`--price-in`, `--price-out` give one for any model).
486
490
 
487
491
  There is a GitHub Action for it (`uses: brunoboto96/SceneScout/ci@…`). [docs/ci.md](docs/ci.md#an-unattended-exploratory-run) has the workflow and every option; [ADR 14](docs/adr/0014-an-unattended-run-reports-and-never-gates.md) says why it works this way.
package/dist/check-run.js CHANGED
@@ -20,6 +20,7 @@ import { firstLineOf } from "./engine/limits.js";
20
20
  import { isNonPageResource } from "./engine/crawl.js";
21
21
  import { checkRetestPlan, retestResults, wellFormedFindings } from "./engine/verify.js";
22
22
  import { capFrames, isReplayFrameFile, journeyOf, journeyVideoPath, isGeneratedReplay, isJourneyVideoFile, redactReplay, REPLAY_VIDEOS_DIRNAME, replayVideos, REPLAY_FILE, REPLAY_FRAMES_DIRNAME, replayFramePath, replayFrames, replaySessionKey, visitOf, } from "./engine/check-replay.js";
23
+ import { buildTemplateReport, evidenceFiles, evidenceIndex, isGeneratedReport, readReportTemplate, recordedReportFile, reportFileOf, templateRecordError, } from "./engine/check-report.js";
23
24
  /** The baselines folder a check uses: the one --baselines names, else the project's own. */
24
25
  export function baselinesDirOf(options) {
25
26
  return options.baselinesDir ?? defaultBaselinesDir(options.projectDir);
@@ -470,7 +471,34 @@ async function takeBaselines(engine, options, mode, targets, log) {
470
471
  * check, recorded or not, so a page from an earlier run is never read, or
471
472
  * uploaded, as this one's; whatever else the folder holds stays.
472
473
  */
473
- export function clearReplayOutput(outDir) {
474
+ export function clearReplayOutput(outDir, reportFile) {
475
+ // The test report the earlier run's check.json says it wrote (--template), and the one this run's template names:
476
+ // each only when it still carries the mark. A copy renamed or kept under any other name is the project's own.
477
+ const reports = new Set();
478
+ try {
479
+ const earlier = recordedReportFile(JSON.parse(fs.readFileSync(path.join(outDir, "check.json"), "utf8")));
480
+ if (earlier)
481
+ reports.add(earlier);
482
+ }
483
+ catch {
484
+ // No earlier check.json, or one that cannot be read: it names no report to remove.
485
+ }
486
+ if (reportFile)
487
+ reports.add(reportFile);
488
+ for (const name of reports) {
489
+ const at = path.join(outDir, name);
490
+ let text;
491
+ try {
492
+ text = fs.readFileSync(at, "utf8");
493
+ }
494
+ catch (err) {
495
+ if (err.code === "ENOENT")
496
+ continue;
497
+ throw err;
498
+ }
499
+ if (isGeneratedReport(text))
500
+ fs.rmSync(at);
501
+ }
474
502
  // Only a page a check wrote: a file of the same name the project keeps there is its own.
475
503
  const page = path.join(outDir, REPLAY_FILE);
476
504
  let text = null;
@@ -512,6 +540,39 @@ export function clearReplayOutput(outDir) {
512
540
  if (fs.readdirSync(root).length === 0)
513
541
  fs.rmdirSync(root);
514
542
  }
543
+ /**
544
+ * The report template a check was given (--template), read and validated, or
545
+ * null for none. Throws, before the browser starts, when the check is not
546
+ * recorded, when the template is not valid, or when the file it would write is
547
+ * one the project keeps there and a check did not write.
548
+ */
549
+ export function prepareTemplateReport(options, outDir) {
550
+ if (options.template === undefined)
551
+ return null;
552
+ const needsRecord = templateRecordError(options);
553
+ if (needsRecord)
554
+ throw new Error(needsRecord);
555
+ const template = readReportTemplate(options.template);
556
+ const file = path.join(outDir, reportFileOf(template));
557
+ let text = null;
558
+ try {
559
+ text = fs.readFileSync(file, "utf8");
560
+ }
561
+ catch (err) {
562
+ if (err.code !== "ENOENT")
563
+ throw err;
564
+ }
565
+ if (text !== null && !isGeneratedReport(text))
566
+ throw new Error(`${file} is not a report SceneScout wrote, so a check will not overwrite it. Move or rename it, name another file in the template's "file", or pass --out to write the check somewhere else`);
567
+ return template;
568
+ }
569
+ /** Write the template's report beside replay.html, after check.json and the replay page, so the manifest hashes the files as written. Returns its file name. */
570
+ export function writeTemplateReport(outDir, result, template, meta) {
571
+ const file = reportFileOf(template);
572
+ const evidence = evidenceIndex(outDir, evidenceFiles(result));
573
+ fs.writeFileSync(path.join(outDir, file), buildTemplateReport(result, template, evidence, meta));
574
+ return file;
575
+ }
515
576
  /**
516
577
  * Why a recorded check (--record or --video) must not start, or null: a
517
578
  * replay.html in the output folder that a check did not write is the
package/dist/ci-run.js CHANGED
@@ -22,9 +22,10 @@ import { CreateMessageRequestSchema, ErrorCode, McpError } from "@modelcontextpr
22
22
  import { DEDUP_JUDGE_CAPABILITY, durationText, JUDGE_CALL_MS, JUDGE_MAX_OUTPUT_TOKENS, JUDGE_SYSTEM, JUDGE_TOOL, judgeKickoffOf, samplingResultOf, } from "./engine/dedup.js";
23
23
  import { CAPTURE_MARGIN, parseCaptureResult, rebaseUrl, SHOT_FILES, SHOTS_DIRNAME } from "./engine/capture.js";
24
24
  import { addUsage, attachFailure, budgetSpend, wallLeftMs, guardToolArgs, CAPTURE_TOOLS, childEnv, ciCaptureKickoff, ciCaptureSystemPrompt, CI_DIRNAME, ciExitCode, ciKickoff, ciSarif, ciSummaryJson, ciSummaryMarkdown, ciSystemPrompt, ciToolArgs, ciTools, describeStop, findingsThisRun, newBudget, NO_USAGE, readFindings, redactKeys, settleTurn, takeTurn, toolResultText, usageLine, } from "./engine/ci.js";
25
- import { ciLaneKickoff, ciLaneSystemPrompt, crawlFoundNothing, crawlNotes, LANE_TOOLS, mergeLaneStops, PLAN_CRAWL_ROUNDS, planCiLanes, PLANNER_SESSION, } from "./engine/ci-lanes.js";
25
+ import { ciLaneKickoff, ciLaneSystemPrompt, crawlFoundNothing, crawlNotes, LANE_TOOLS, mergeLaneStops, PLAN_CRAWL_ROUNDS, planCiLanes, plannedRoutes, planReplayLanes, PLANNER_SESSION, } from "./engine/ci-lanes.js";
26
26
  import { resolveTimeLimits } from "./engine/limits.js";
27
- import { MEMORY_DIRNAME, writeSelfIgnore } from "./engine/memory.js";
27
+ import { loadRunRecord, MEMORY_DIRNAME, noteFromRunOnDisk, readRunRecordsOnDisk, writeSelfIgnore } from "./engine/memory.js";
28
+ import { carryForward, continueLines, continueExhausted, continuePlan, DEFAULT_TURNS_PER_PAGE, pageCap, takenPages, fromRunLine, isPattern, pathLine, prefixOutcome, prefixPlan, replayLines, replayPlan, } from "./engine/from-run.js";
28
29
  import { sarifFilesFor } from "./engine/sarif.js";
29
30
  import { decodePng, diffImages, encodePng } from "./engine/png.js";
30
31
  import { AnthropicConversation, backoffMs, errorMessage, MalformedReply, OpenAIConversation, retryable, retryAfterMs, } from "./engine/provider.js";
@@ -434,43 +435,119 @@ function takesSession(t) {
434
435
  }
435
436
  const messageOf = (err) => (err instanceof Error ? err.message : String(err));
436
437
  /**
437
- * A run split into lanes (--lanes): plan, then run every lane at once, then
438
- * fold what they did. The plan is a snapshot and a crawl from the planner's
439
- * session (no model call: only their time counts), the crawl repeated while it
440
- * finds routes, and the split brief.ts makes of them. Each lane attaches its
441
- * own session on the run's target URL, as the planner did (the engine resolves
442
- * every path against the URL a session attached with), opens its first route,
443
- * runs the agent loop there with its own conversation, draws its turns from
444
- * the run's one budget, and is closed when it ends. Their findings are already
445
- * one: every session files into the project's one memory, whose dedup folds a
446
- * defect two lanes filed. `outcome` is absent when there was nothing to split,
447
- * and the run explores in one loop instead.
438
+ * Take the earlier run's path to a page (from-run.ts prefixPlan) in one
439
+ * scout_run_plan call, which the write policy governs like any other step.
440
+ * Nothing happens when the record never reached the page or opened it by its
441
+ * address. When a step cannot be repeated (the record cannot say it, or the
442
+ * page changed), the session navigates to the page directly instead, if it is
443
+ * not a pattern. Returns the note for the record and the line the model is told.
448
444
  */
449
- async function exploreInLanes(o) {
450
- const { host, options, log, now, budget } = o;
451
- const timeLeft = () => wallLeftMs(budgetSpend(budget), options.caps, now());
452
- // ── plan ──
445
+ async function reachByPath(o) {
446
+ if (!o.route)
447
+ return { lines: [] };
448
+ const plan = prefixPlan(o.record, o.route);
449
+ if (!plan || ("steps" in plan && plan.direct))
450
+ return { lines: [] };
451
+ let why;
452
+ let steps = 0;
453
+ if ("cannot" in plan)
454
+ why = `its path cannot be repeated: ${plan.cannot}`;
455
+ else {
456
+ steps = plan.steps.length;
457
+ const r = await o.host
458
+ .call("scout_run_plan", { session: o.session, steps: plan.steps, task: `Taking the earlier run's path to ${o.route}` }, Math.min(240_000, o.timeLeft()))
459
+ .catch((err) => ({ text: `ERROR: ${messageOf(err)}`, isError: true }));
460
+ const out = prefixOutcome(r.text.replace(/^\[session [^\]\n]*\]\n/, ""), steps, o.route);
461
+ if (out.ok) {
462
+ o.say(`reached ${o.route} by the earlier run's path (${steps} step(s)).`);
463
+ return {
464
+ note: { session: o.session, target: o.route, steps, outcome: "replayed" },
465
+ lines: [`Your browser reached ${o.route} the way the earlier run did: ${pathLine(plan.steps)}. Start there.`],
466
+ };
467
+ }
468
+ why = out.why;
469
+ }
470
+ o.say(`could not take the earlier run's path to ${o.route} (${why}); navigating to it instead.`);
471
+ const direct = isPattern(o.route)
472
+ ? undefined
473
+ : await o.host
474
+ .call("scout_navigate", { session: o.session, target: o.route, task: `Opening ${o.route}` }, Math.min(120_000, o.timeLeft()))
475
+ .catch((err) => ({ text: `ERROR: ${messageOf(err)}`, isError: true }));
476
+ const opened = !!direct && !direct.isError && !/^(ERROR|REFUSED):/.test(direct.text.replace(/^\[session [^\]\n]*\]\n/, ""));
477
+ return {
478
+ note: { session: o.session, target: o.route, steps, outcome: opened ? "navigated" : "unreached", why: why.slice(0, 300) },
479
+ lines: [
480
+ opened
481
+ ? `The earlier run's path to ${o.route} could not be repeated (${why}), so your browser was sent there directly.`
482
+ : `The earlier run's path to ${o.route} could not be repeated (${why}); reach it from the page you are on.`,
483
+ ],
484
+ };
485
+ }
486
+ /** The planning crawl's notes with these routes added, each with no note: routes a record knew that the crawl did not list. */
487
+ function withRoutes(notes, routes) {
488
+ const out = new Map(notes);
489
+ for (const r of routes)
490
+ if (!out.has(r))
491
+ out.set(r, []);
492
+ return out;
493
+ }
494
+ /**
495
+ * The record this run's report left in the project's memory (from-run.ts), for
496
+ * ci.json: the newest written since the run started. A continued run's carries
497
+ * the record it continued, so a chain of runs each given the last one's
498
+ * ci.json accumulates what the chain covered. Never fatal: without one, ci.json
499
+ * has no record, and the log says why.
500
+ */
501
+ function thisRunsRecord(o) {
502
+ let own;
503
+ try {
504
+ own = readRunRecordsOnDisk(o.projectDir)
505
+ .filter((r) => r.at >= o.since)
506
+ .at(-1);
507
+ }
508
+ catch (err) {
509
+ o.log(`The run's record could not be read from the project's memory, so ci.json holds none: ${messageOf(err)}`);
510
+ return undefined;
511
+ }
512
+ if (!own) {
513
+ o.log("The report left no run record in the project's memory, so ci.json holds none.");
514
+ return undefined;
515
+ }
516
+ const withPrefixes = {
517
+ ...own,
518
+ ...(o.prefixes && o.prefixes.length > 0 ? { prefixes: [...o.prefixes] } : {}),
519
+ ...(o.continuedFresh ? { continuedFresh: true } : {}),
520
+ ...(o.assigned && o.assigned.length > 0 ? { assigned: [...new Set(o.assigned)] } : {}),
521
+ };
522
+ return o.continued ? carryForward(o.continued, withPrefixes) : withPrefixes;
523
+ }
524
+ /**
525
+ * The planning crawl a run split into lanes, or a continued run, starts with: a
526
+ * snapshot from the planner's session (attaching harvests no links; a snapshot
527
+ * of the page it landed on does), then a crawl repeated while it finds routes,
528
+ * up to PLAN_CRAWL_ROUNDS. No model call: only its time counts.
529
+ */
530
+ async function planningCrawl(o) {
453
531
  // What went wrong while planning, so a plan left with nothing to split says why rather than blaming the app.
454
532
  let planningFailed;
455
533
  /** One planner call: its text, or undefined after saying why it failed. */
456
534
  const plannerCall = async (tool, maxMs, what) => {
457
535
  try {
458
- const r = await host.call(tool, { session: PLANNER_SESSION }, Math.min(maxMs, timeLeft()));
536
+ const r = await o.host.call(tool, { session: PLANNER_SESSION }, Math.min(maxMs, o.timeLeft()));
459
537
  if (r.isError)
460
538
  throw new Error(r.text.replace(/^ERROR:\s*/, ""));
461
539
  return r.text;
462
540
  }
463
541
  catch (err) {
464
542
  planningFailed = `the planning ${what} failed: ${messageOf(err).slice(0, 300)}`;
465
- log(`Lanes: ${planningFailed}.`);
543
+ o.log(`${o.label}: ${planningFailed}.`);
466
544
  return undefined;
467
545
  }
468
546
  };
469
- // Attaching harvests no links; a snapshot of the page it landed on does, so the first crawl has routes to visit.
470
- if (timeLeft() > 0)
547
+ if (o.timeLeft() > 0)
471
548
  await plannerCall("scout_snapshot", 120_000, "snapshot");
472
549
  const notes = new Map();
473
- for (let round = 0; round < PLAN_CRAWL_ROUNDS && timeLeft() > 0; round += 1) {
550
+ for (let round = 0; round < PLAN_CRAWL_ROUNDS && o.timeLeft() > 0; round += 1) {
474
551
  const text = await plannerCall("scout_crawl", 600_000, "crawl");
475
552
  if (text === undefined)
476
553
  break;
@@ -481,14 +558,24 @@ async function exploreInLanes(o) {
481
558
  if (crawlFoundNothing(text) || notes.size === known)
482
559
  break;
483
560
  }
484
- const plan = planCiLanes({
485
- target: options.url,
486
- notes,
487
- count: options.lanes,
488
- focus: options.focus,
489
- mode: options.mode,
490
- ...(planningFailed ? { planningFailed } : {}),
491
- });
561
+ return { notes, ...(planningFailed ? { planningFailed } : {}) };
562
+ }
563
+ /**
564
+ * A run split into lanes (--lanes): plan, then run every lane at once, then
565
+ * fold what they did. The plan is a snapshot and a crawl from the planner's
566
+ * session (no model call: only their time counts), the crawl repeated while it
567
+ * finds routes, and the split brief.ts makes of them. Each lane attaches its
568
+ * own session on the run's target URL, as the planner did (the engine resolves
569
+ * every path against the URL a session attached with), opens its first route,
570
+ * runs the agent loop there with its own conversation, draws its turns from
571
+ * the run's one budget, and is closed when it ends. Their findings are already
572
+ * one: every session files into the project's one memory, whose dedup folds a
573
+ * defect two lanes filed. `outcome` is absent when there was nothing to split,
574
+ * and the run explores in one loop instead.
575
+ */
576
+ async function exploreInLanes(o) {
577
+ const { host, options, log, now, budget, plan } = o;
578
+ const timeLeft = () => wallLeftMs(budgetSpend(budget), options.caps, now());
492
579
  if (plan.oneLoop) {
493
580
  log(`Lanes: ${plan.oneLoop}. Exploring in one loop.`);
494
581
  return { lanes: { asked: options.lanes, sessions: [], oneLoop: plan.oneLoop } };
@@ -539,7 +626,9 @@ async function exploreInLanes(o) {
539
626
  on = lane.url;
540
627
  }
541
628
  say(`attached on ${new URL(on).pathname}, owning ${lane.modules.join(", ")}.`);
542
- return await exploreLane(lane, on, say, { ...result, attached: true });
629
+ // A continued run: the earlier run's path to this lane's first page, when it got there by acting on another.
630
+ const reached = o.reach && timeLeft() > 0 ? await o.reach(lane.session, lane.routes[0], say) : [];
631
+ return await exploreLane(lane, on, say, { ...result, attached: true }, reached);
543
632
  }
544
633
  catch (err) {
545
634
  // Not a cap and not the model's API: the lane itself broke. Reported as the lane's, and the run's (mergeLaneStops), never dropped.
@@ -548,7 +637,7 @@ async function exploreInLanes(o) {
548
637
  return { ...result, attached, stop: "could-not-start", stopDetail: why };
549
638
  }
550
639
  };
551
- const exploreLane = async (lane, on, say, result) => {
640
+ const exploreLane = async (lane, on, say, result, reached = []) => {
552
641
  const outcome = await agentLoop({
553
642
  client: o.makeClient(system, tools, ciLaneKickoff({
554
643
  lane: { ...lane, url: on },
@@ -559,6 +648,7 @@ async function exploreInLanes(o) {
559
648
  level: options.level,
560
649
  focus: options.focus,
561
650
  caps: options.caps,
651
+ ...(o.laneLines || reached.length > 0 ? { fromRun: [...(o.laneLines?.(lane) ?? []), ...reached] } : {}),
562
652
  })),
563
653
  host,
564
654
  tools,
@@ -625,6 +715,11 @@ export async function runCi(options, resolved, deps) {
625
715
  let reportWritten = false;
626
716
  let capture;
627
717
  let lanes;
718
+ let runFromRun;
719
+ let runRecord;
720
+ let earlier;
721
+ // The real clock, as the server's: this run's record is the one its report wrote after this.
722
+ const startedIso = new Date().toISOString();
628
723
  let host = null;
629
724
  // Set when the exploration ends (or never starts): the report and the close share FINISH_MS from then.
630
725
  let finishBy = 0;
@@ -636,6 +731,17 @@ export async function runCi(options, resolved, deps) {
636
731
  const judge = judgeAsk
637
732
  ? judgeHandler({ ask: judgeAsk, model: resolved.model, spend: budget, caps: options.caps, calls: judgeCalls, secrets, now })
638
733
  : undefined;
734
+ // The earlier run's record, read before anything starts: a run asked to continue or replay one must not quietly start fresh.
735
+ if (options.fromRun && !options.show) {
736
+ try {
737
+ earlier = loadRunRecord(options.fromRun.path);
738
+ }
739
+ catch (err) {
740
+ throw new Error(`--from-run: ${messageOf(err)}`);
741
+ }
742
+ if (options.fromRun.mode === "replay" && replayPlan(earlier).length === 0)
743
+ throw new Error(`--from-run: the run recorded in ${options.fromRun.path} took no steps, so there is nothing to replay`);
744
+ }
639
745
  host = await (deps.startHost ?? startServer)(log, judge);
640
746
  // What every session of the run attaches with: the planner's here, and each lane's when the run is split.
641
747
  const attachArgs = (a) => ({
@@ -666,10 +772,67 @@ export async function runCi(options, resolved, deps) {
666
772
  const listed = await host.tools();
667
773
  const toolHost = host;
668
774
  let captured = null;
669
- const oneLoop = () => {
775
+ // Named as it was given: a resolved path would put a local directory in the report and ci.json.
776
+ const from = options.fromRun ? (options.fromRun.given ?? options.fromRun.path) : "";
777
+ const replay = earlier && options.fromRun?.mode === "replay" ? replayPlan(earlier) : undefined;
778
+ const continuing = earlier && options.fromRun?.mode === "continue" ? earlier : undefined;
779
+ // A run split into lanes, or a continued one, crawls first: the lanes are split, and a continued run's routes ordered, from what it finds.
780
+ // A replay does not: its routes and steps are the record's.
781
+ const timeLeft = () => wallLeftMs(budgetSpend(budget), options.caps, now());
782
+ const plans = !options.show && !replay && (options.lanes > 1 || !!continuing);
783
+ const planned = plans ? await planningCrawl({ host, log, timeLeft, label: options.lanes > 1 ? "Lanes" : "From run" }) : undefined;
784
+ const planItems = continuing && planned ? continuePlan(continuing, plannedRoutes(options.url, planned.notes)) : undefined;
785
+ // A record with no work left anywhere: explore as a fresh run (the stable split, the landing page, no path), and say so.
786
+ const continuedFresh = !!planItems && continueExhausted(planItems);
787
+ const items = continuedFresh ? undefined : planItems;
788
+ // How many pages a continued run takes on, from its budget (from-run.ts pageCap): each lane's share of the turns, or the run's.
789
+ const perPage = options.fromRun?.turnsPerPage ?? DEFAULT_TURNS_PER_PAGE;
790
+ const cap = items ? pageCap(options.caps.turns, perPage) : undefined;
791
+ const laneCap = items ? pageCap(Math.floor(options.caps.turns / Math.max(1, options.lanes)), perPage) : undefined;
792
+ // The pages this run was given, kept in its record so the next continued run takes the next ones.
793
+ const assigned = [];
794
+ if (options.fromRun && earlier && !options.show) {
795
+ runFromRun = { mode: options.fromRun.mode, source: from, runId: earlier.runId, recordAt: earlier.at };
796
+ log(`From run: ${fromRunLine(runFromRun)}.`);
797
+ if (items)
798
+ log(`From run: ${items.filter((i) => i.tier === 1).length} route(s) it never worked on, ${items.filter((i) => i.tier === 2).length} with work left, ${items.filter((i) => i.tier === 3).length} worked through.`);
799
+ if (continuedFresh)
800
+ log("From run: the earlier run left no recorded work on any route, so this run explores as a fresh one.");
801
+ // Noted in the project's memory for the report, which the server writes. Never fatal: the run goes on.
802
+ try {
803
+ noteFromRunOnDisk(options.projectDir, { at: new Date().toISOString(), ...runFromRun });
804
+ }
805
+ catch (err) {
806
+ log(`From run: the project's memory could not be told, so the report will not name the earlier run: ${messageOf(err)}`);
807
+ }
808
+ }
809
+ // How a continued run's lanes (or its one loop) reached their first page by the earlier run's path, for the record.
810
+ const prefixes = [];
811
+ const reach = items
812
+ ? async (session, route, say) => {
813
+ const r = await reachByPath({ host: toolHost, session, record: continuing, route, timeLeft, say });
814
+ if (r.note)
815
+ prefixes.push(r.note);
816
+ return r.lines;
817
+ }
818
+ : undefined;
819
+ const oneLoop = async () => {
670
820
  const tools = ciTools(listed, options.show ? CAPTURE_TOOLS : undefined);
671
821
  const system = options.show ? ciCaptureSystemPrompt() : ciSystemPrompt(loadPlaybook(packageRoot), options);
672
- const kickoff = options.show ? ciCaptureKickoff({ url: options.url, show: options.show }) : ciKickoff(options);
822
+ const reached = reach && items && !options.show ? await reach(PLANNER_SESSION, items[0]?.route, (l) => log(`From run: ${l}`)) : [];
823
+ if (items && cap !== undefined) {
824
+ const mine = takenPages(items, cap);
825
+ assigned.push(...mine);
826
+ log(`From run: takes on ${mine.join(", ")} (${cap} page(s) at ${perPage} turns a page).`);
827
+ }
828
+ const fromRun = planItems
829
+ ? [...continueLines(planItems, { from, ...(cap !== undefined ? { cap } : {}) }), ...reached]
830
+ : replay && replay.length > 0
831
+ ? replayLines(replay[0], { from })
832
+ : undefined;
833
+ const kickoff = options.show
834
+ ? ciCaptureKickoff({ url: options.url, show: options.show })
835
+ : ciKickoff({ ...options, ...(fromRun && fromRun.length > 0 ? { fromRunLines: fromRun } : {}) });
673
836
  return agentLoop({
674
837
  client: deps.makeClient(system, tools, kickoff),
675
838
  host: toolHost,
@@ -685,8 +848,56 @@ export async function runCi(options, resolved, deps) {
685
848
  },
686
849
  });
687
850
  };
688
- if (options.lanes > 1 && !options.show) {
689
- const split = await exploreInLanes({ host, listed, options, makeClient: deps.makeClient, log, now, budget, attachArgs, attachMs });
851
+ // A replay has as many lanes as the recorded run had sessions; otherwise the planning crawl is split as asked.
852
+ const lanePlan = options.show
853
+ ? undefined
854
+ : replay
855
+ ? planReplayLanes({ target: options.url, lanes: replay, ...(options.focus ? { focus: options.focus } : {}) })
856
+ : options.lanes > 1 && planned
857
+ ? planCiLanes({
858
+ target: options.url,
859
+ // A continued run splits every route it orders, the record's own included, so none is left out of every lane.
860
+ notes: items
861
+ ? withRoutes(planned.notes, items.map((i) => i.route))
862
+ : planned.notes,
863
+ count: options.lanes,
864
+ focus: options.focus,
865
+ mode: options.mode,
866
+ ...(planned.planningFailed ? { planningFailed: planned.planningFailed } : {}),
867
+ ...(items ? { order: items.map((i) => i.route) } : {}),
868
+ })
869
+ : undefined;
870
+ if (replay && replay.length > 1 && replay.length !== options.lanes)
871
+ log(`From run: replaying the recorded run's ${replay.length} sessions as ${replay.length} lanes.`);
872
+ if (lanePlan && (lanePlan.lanes.length > 0 || options.lanes > 1)) {
873
+ // A replayed lane is told its own session's steps: the plan keeps the recorded sessions' order.
874
+ const replayed = new Map(replay ? lanePlan.lanes.map((l, i) => [l.session, replay[i]]) : []);
875
+ const laneLines = planItems
876
+ ? (lane) => {
877
+ if (items && laneCap !== undefined) {
878
+ const mine = takenPages(items, laneCap, lane.routes);
879
+ assigned.push(...mine);
880
+ log(`From run: lane ${lane.session} takes on ${mine.join(", ") || "no page with work left"} (${laneCap} page(s) at ${perPage} turns a page).`);
881
+ }
882
+ return continueLines(planItems, { from, only: lane.routes, ...(laneCap !== undefined ? { cap: laneCap } : {}) });
883
+ }
884
+ : replay
885
+ ? (lane) => replayLines(replayed.get(lane.session), { from })
886
+ : undefined;
887
+ const split = await exploreInLanes({
888
+ host,
889
+ listed,
890
+ options,
891
+ makeClient: deps.makeClient,
892
+ log,
893
+ now,
894
+ budget,
895
+ attachArgs,
896
+ attachMs,
897
+ plan: lanePlan,
898
+ ...(laneLines ? { laneLines } : {}),
899
+ ...(reach ? { reach } : {}),
900
+ });
690
901
  lanes = split.lanes;
691
902
  outcome = split.outcome ?? (await oneLoop());
692
903
  }
@@ -709,6 +920,16 @@ export async function runCi(options, resolved, deps) {
709
920
  reportWritten = !report.isError && !/^ERROR:/.test(report.text) && fs.existsSync(path.join(options.projectDir, MEMORY_DIRNAME, "report.md"));
710
921
  if (!reportWritten)
711
922
  log(`The report could not be generated: ${report.text.slice(0, 400)}`);
923
+ else
924
+ runRecord = thisRunsRecord({
925
+ projectDir: options.projectDir,
926
+ since: startedIso,
927
+ continued: earlier && options.fromRun?.mode === "continue" ? earlier : undefined,
928
+ prefixes,
929
+ continuedFresh,
930
+ assigned,
931
+ log,
932
+ });
712
933
  }
713
934
  }
714
935
  }
@@ -748,6 +969,8 @@ export async function runCi(options, resolved, deps) {
748
969
  findings: findingsThisRun(before, readMemoryFindings(options.projectDir)),
749
970
  ...(capture ? { capture } : {}),
750
971
  ...(lanes ? { lanes } : {}),
972
+ ...(runFromRun ? { fromRun: runFromRun } : {}),
973
+ ...(runRecord ? { record: runRecord } : {}),
751
974
  ...(options.show
752
975
  ? {}
753
976
  : {
package/dist/cli.js CHANGED
@@ -21,7 +21,8 @@ import { APPROX_DISK_MB, defaultAttachNote, defaultEngine, launchTarget, parseBr
21
21
  import { CLIENT_LABELS, firstMessageHint, manualFor, parseClients, registerWithClient, vscodeBinary } from "./clients.js";
22
22
  import { CLAUDE_CODE_NOT_NEEDED, CLI_NAME, desktopExtensionRoots, diagnose, doctorAllGood, findDesktopExtension, installClosing, ensureCommand, findOnUserPath, installSkill, isEphemeralRoot, launchCommand, manualRegisterCommand, planCommand, registerMcp, resolveClaudeDir, spawnRunner, } from "./installer.js";
23
23
  import { downloadBrowsers, presentBrowsers } from "./installer.js";
24
- import { baselinesDirOf, clearReplayOutput, replayPageConflict, defaultCheckDir, readCheckInputs, runCheck } from "./check-run.js";
24
+ import { baselinesDirOf, clearReplayOutput, replayPageConflict, defaultCheckDir, prepareTemplateReport, readCheckInputs, runCheck, writeTemplateReport, } from "./check-run.js";
25
+ import { reportFileOf } from "./engine/check-report.js";
25
26
  import { buildCheckReplayHtml, commitOf, REPLAY_FILE, replayFrames, replayVideos } from "./engine/check-replay.js";
26
27
  import { recordChoice } from "./engine/capture.js";
27
28
  import { httpClient, httpJudgeAsk, runCi } from "./ci-run.js";
@@ -112,7 +113,10 @@ Usage:
112
113
  and write replay.html beside the report, role → journey → step (default:
113
114
  SCENESCOUT_RECORD, else off; the frames go in replay-frames/);
114
115
  --video [on|off]: record a WebM of each saved flow, and only of the flows,
115
- into replay-videos/, played on replay.html beside its steps (default off))
116
+ into replay-videos/, played on replay.html beside its steps (default off);
117
+ --template file.json: on a recorded check, also write a test report laid out
118
+ as the template says (report.html unless it names a file), with the frames,
119
+ deviations and a SHA-256 manifest of the evidence; needs --record)
116
120
  Exit code: 0 passed, 1 failed the gate, 2 could not run.
117
121
  scenescout ci <url> An exploratory run with no person present: a model reached through its API
118
122
  drives the tools by the SceneScout method and the run ends in the report.
@@ -580,15 +584,18 @@ async function check(args) {
580
584
  let options = parsed.options;
581
585
  const outDir = options.outDir ?? defaultCheckDir(options.projectDir);
582
586
  let inputs;
587
+ let template = null;
583
588
  try {
584
589
  // --record, else SCENESCOUT_RECORD, else off.
585
590
  options = { ...options, record: recordChoice(options.record, process.env) };
591
+ // --template: read and checked before the browser starts, so a bad template never costs a run.
592
+ template = prepareTemplateReport(options, outDir);
586
593
  // A replay.html the project keeps there is its own: a recorded check refuses to start rather than overwrite it.
587
594
  const conflict = options.record || options.video ? replayPageConflict(outDir) : null;
588
595
  if (conflict)
589
596
  throw new Error(conflict);
590
597
  // A replay page an earlier run left must never be read, or uploaded, as this run's.
591
- clearReplayOutput(outDir);
598
+ clearReplayOutput(outDir, template ? reportFileOf(template) : undefined);
592
599
  inputs = readCheckInputs(options);
593
600
  }
594
601
  catch (err) {
@@ -615,7 +622,11 @@ async function check(args) {
615
622
  console.error(`scenescout check: could not run a saved flow: ${refused}`);
616
623
  process.exit(EXIT.error);
617
624
  }
625
+ // Recorded in check.json, so the next check removes exactly this file and the action finds it.
626
+ if (template && result.replay)
627
+ result = { ...result, testReport: reportFileOf(template) };
618
628
  const markdown = formatCheck(result);
629
+ let reportFile = null;
619
630
  try {
620
631
  const version = packageVersion();
621
632
  if (!options.outDir)
@@ -645,6 +656,8 @@ async function check(args) {
645
656
  couldNotRun,
646
657
  });
647
658
  fs.writeFileSync(path.join(outDir, REPLAY_FILE), html);
659
+ if (template)
660
+ reportFile = writeTemplateReport(outDir, result, template, { version, commit: commitOf(process.env) });
648
661
  }
649
662
  // On GitHub Actions the verdict also goes on the run's summary page.
650
663
  if (process.env.GITHUB_STEP_SUMMARY)
@@ -659,6 +672,8 @@ async function check(args) {
659
672
  console.log(`Wrote report.md, check.sarif and check.json to ${outDir}`);
660
673
  if (result.replay)
661
674
  console.log(`Wrote ${REPLAY_FILE} to ${outDir}, with ${replayFrames(result.replay).length} frame(s) and ${replayVideos(result.replay).length} journey video(s) beside it: open it in a browser to see each step`);
675
+ if (reportFile)
676
+ console.log(`Wrote ${reportFile} to ${outDir}: the test report the template lays out, with a SHA-256 manifest of its evidence`);
662
677
  const pictured = result.baselines?.results.filter((r) => r.files).length ?? 0;
663
678
  if (pictured > 0)
664
679
  console.log(`Wrote the pictures of ${pictured} changed baseline(s) under ${path.join(outDir, VISUAL_DIRNAME)}`);
@@ -762,7 +777,7 @@ async function ci(args) {
762
777
  console.error(redactKeys(`scenescout ci: ${message}`, secrets));
763
778
  process.exit(EXIT_CI.couldNotRun);
764
779
  };
765
- const parsed = parseCiArgs(args, process.cwd());
780
+ const parsed = parseCiArgs(args, process.cwd(), process.env);
766
781
  if (!parsed.ok)
767
782
  return fail(parsed.error);
768
783
  const options = parsed.options;