scenescout 3.21.3 → 3.23.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +12 -0
- package/README.md +5 -1
- package/dist/check-run.js +62 -1
- package/dist/ci-run.js +259 -36
- package/dist/cli.js +19 -4
- package/dist/engine/brief.js +35 -3
- package/dist/engine/check-replay.js +12 -2
- package/dist/engine/check-report.js +503 -0
- package/dist/engine/check.js +9 -0
- package/dist/engine/ci-lanes.js +28 -12
- package/dist/engine/ci.js +39 -1
- package/dist/engine/flow.js +28 -8
- package/dist/engine/from-run.js +649 -0
- package/dist/engine/memory.js +123 -0
- package/dist/engine/report.js +26 -0
- package/dist/mcp-server.js +63 -5
- package/package.json +3 -2
- package/skills/scenescout/SKILL.md +2 -2
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,17 @@
|
|
|
1
1
|
# scenescout
|
|
2
2
|
|
|
3
|
+
## 3.23.0
|
|
4
|
+
|
|
5
|
+
### Minor Changes
|
|
6
|
+
|
|
7
|
+
- ed24973: Start a run from an earlier run's record, opt-in. Every run that writes its report now keeps a record of what it worked on and left (routes in order, each session's steps, and on each route the controls never exercised, forms never submitted and options never chosen), in the project's memory and in `ci.json` under `record`. `scenescout ci --from-run <ci.json or project directory>` (or `SCENESCOUT_FROM_RUN`) with `--from-run-mode continue` (the default) takes first the routes that run never worked on, then the ones it left work on, told exactly what to do first on each, then the rest, reaching a page the earlier run got to by acting on another page through the same steps; a run takes on only its budget's worth of pages (`--from-run-turns-per-page`, default 7 turns a page) and the next continued run takes the pages after them; `--from-run-mode replay` follows its routes and steps in order, one lane per session it had. `scout_lane_brief` takes `fromRun` and `fromRunMode`, the ci action takes `from-run` and `from-run-mode`, and the run it started from is named in the report, `summary.md` and `ci.json` (`fromRun`). Without it, runs behave as before.
|
|
8
|
+
|
|
9
|
+
## 3.22.0
|
|
10
|
+
|
|
11
|
+
### Minor Changes
|
|
12
|
+
|
|
13
|
+
- 072b68b: `scenescout check --record --template <file.json>` (the action's `template` input) also writes the recorded run up as a test report laid out by the template: each saved flow a test and each step a row with its expected and actual result and its frame, failed steps as deviations, a SHA-256 manifest of the evidence and blank sign-off rows. Flows gain optional `id`, `requirements` and a step's `expected` for traceability; an example template is in `examples/report-template.json`.
|
|
14
|
+
|
|
3
15
|
## 3.21.3
|
|
4
16
|
|
|
5
17
|
### Patch Changes
|
package/README.md
CHANGED
|
@@ -279,7 +279,10 @@ Use SceneScout to test http://localhost:3000, record the run
|
|
|
279
279
|
or, on the tool directly, `scout_attach {record: true}`. `SCENESCOUT_RECORD=on` in
|
|
280
280
|
the server's environment records every run. A CI gate records too:
|
|
281
281
|
`scenescout check --record` writes `replay.html`, every journey step by step with
|
|
282
|
-
its frames ([recording a check](docs/ci.md#recording-a-check))
|
|
282
|
+
its frames ([recording a check](docs/ci.md#recording-a-check)), and
|
|
283
|
+
`--template <file.json>` also writes it up as a test report laid out as a template
|
|
284
|
+
says, with expected and actual results, deviations, blank sign-off rows and a
|
|
285
|
+
SHA-256 manifest of the evidence ([test reports](docs/ci.md#a-test-report-from-a-template)).
|
|
283
286
|
|
|
284
287
|
Then `scout_report` writes two files side by side in `.scenescout/`:
|
|
285
288
|
`report.md` as always, and `report.html` — the whole run as one self-contained
|
|
@@ -482,6 +485,7 @@ npx scenescout ci http://127.0.0.1:3000
|
|
|
482
485
|
- **Caps:** at most 80 model turns, 3,000,000 tokens and 20 minutes (`--max-turns`, `--max-tokens`, `--max-minutes`), shared by two model loops that each explore their own part of the app (`--lanes`, default 2). The first cap reached ends the exploration; the report is still written, and says which cap ended it. On the benchmark's demo app a run at these defaults cost about $0.03 on `gpt-6-luna` ([docs/benchmark.md](docs/benchmark.md#choosing-the-defaults-issue-419)).
|
|
483
486
|
- **Mode:** `read-only` by default; `--mode observe` sends no form at all, `--mode safe-write` lets the run create records and change only the ones it created. `--mode destructive` runs only with `--allow-destructive` as well.
|
|
484
487
|
- **Duplicates:** when the dedup rule keeps a filed finding apart, the run's model is asked at its lowest effort whether it is one already open on the same page, and merges it on a "same", keeping the filing's title, category, severity and evidence under that finding. The two findings' titles, categories and evidence, and the page's path, are sent; `--dedup rule` turns it off ([ADR 17](docs/adr/0017-a-model-judges-only-the-merges-the-rule-misses.md)).
|
|
488
|
+
- **Starting from an earlier run (opt-in):** every run that writes its report leaves a record in `ci.json` and the project's memory: the routes it worked on, its steps, and what it left on each route. `--from-run <ci.json or project directory>` continues where that run left off (routes it never worked on first, then the ones it left work on, with exactly which forms, options and controls to take first), and `--from-run-mode replay` follows its routes and steps in order. `scout_lane_brief {fromRun, fromRunMode}` does the same for a parallel run. On the demo app a chain of continued runs found 8 defects against 7 for fresh runs at the default budget, and 4, 3 and 4 against 4 at a budget cut to stand in for a larger app, within the noise, so it stays off unless asked for ([the measurement](docs/benchmark.md#starting-from-an-earlier-run-issue-418)).
|
|
485
489
|
- **Output**, in `.scenescout/ci/` (or `--out`): `report.md` and `report.html` (the report an agent's run writes), `summary.md` (also appended to the GitHub job summary), `ci.json` and `ci.sarif`, with a usage line: turns, tokens, time and an estimated cost where the model's price is known (`--price-in`, `--price-out` give one for any model).
|
|
486
490
|
|
|
487
491
|
There is a GitHub Action for it (`uses: brunoboto96/SceneScout/ci@…`). [docs/ci.md](docs/ci.md#an-unattended-exploratory-run) has the workflow and every option; [ADR 14](docs/adr/0014-an-unattended-run-reports-and-never-gates.md) says why it works this way.
|
package/dist/check-run.js
CHANGED
|
@@ -20,6 +20,7 @@ import { firstLineOf } from "./engine/limits.js";
|
|
|
20
20
|
import { isNonPageResource } from "./engine/crawl.js";
|
|
21
21
|
import { checkRetestPlan, retestResults, wellFormedFindings } from "./engine/verify.js";
|
|
22
22
|
import { capFrames, isReplayFrameFile, journeyOf, journeyVideoPath, isGeneratedReplay, isJourneyVideoFile, redactReplay, REPLAY_VIDEOS_DIRNAME, replayVideos, REPLAY_FILE, REPLAY_FRAMES_DIRNAME, replayFramePath, replayFrames, replaySessionKey, visitOf, } from "./engine/check-replay.js";
|
|
23
|
+
import { buildTemplateReport, evidenceFiles, evidenceIndex, isGeneratedReport, readReportTemplate, recordedReportFile, reportFileOf, templateRecordError, } from "./engine/check-report.js";
|
|
23
24
|
/** The baselines folder a check uses: the one --baselines names, else the project's own. */
|
|
24
25
|
export function baselinesDirOf(options) {
|
|
25
26
|
return options.baselinesDir ?? defaultBaselinesDir(options.projectDir);
|
|
@@ -470,7 +471,34 @@ async function takeBaselines(engine, options, mode, targets, log) {
|
|
|
470
471
|
* check, recorded or not, so a page from an earlier run is never read, or
|
|
471
472
|
* uploaded, as this one's; whatever else the folder holds stays.
|
|
472
473
|
*/
|
|
473
|
-
export function clearReplayOutput(outDir) {
|
|
474
|
+
export function clearReplayOutput(outDir, reportFile) {
|
|
475
|
+
// The test report the earlier run's check.json says it wrote (--template), and the one this run's template names:
|
|
476
|
+
// each only when it still carries the mark. A copy renamed or kept under any other name is the project's own.
|
|
477
|
+
const reports = new Set();
|
|
478
|
+
try {
|
|
479
|
+
const earlier = recordedReportFile(JSON.parse(fs.readFileSync(path.join(outDir, "check.json"), "utf8")));
|
|
480
|
+
if (earlier)
|
|
481
|
+
reports.add(earlier);
|
|
482
|
+
}
|
|
483
|
+
catch {
|
|
484
|
+
// No earlier check.json, or one that cannot be read: it names no report to remove.
|
|
485
|
+
}
|
|
486
|
+
if (reportFile)
|
|
487
|
+
reports.add(reportFile);
|
|
488
|
+
for (const name of reports) {
|
|
489
|
+
const at = path.join(outDir, name);
|
|
490
|
+
let text;
|
|
491
|
+
try {
|
|
492
|
+
text = fs.readFileSync(at, "utf8");
|
|
493
|
+
}
|
|
494
|
+
catch (err) {
|
|
495
|
+
if (err.code === "ENOENT")
|
|
496
|
+
continue;
|
|
497
|
+
throw err;
|
|
498
|
+
}
|
|
499
|
+
if (isGeneratedReport(text))
|
|
500
|
+
fs.rmSync(at);
|
|
501
|
+
}
|
|
474
502
|
// Only a page a check wrote: a file of the same name the project keeps there is its own.
|
|
475
503
|
const page = path.join(outDir, REPLAY_FILE);
|
|
476
504
|
let text = null;
|
|
@@ -512,6 +540,39 @@ export function clearReplayOutput(outDir) {
|
|
|
512
540
|
if (fs.readdirSync(root).length === 0)
|
|
513
541
|
fs.rmdirSync(root);
|
|
514
542
|
}
|
|
543
|
+
/**
|
|
544
|
+
* The report template a check was given (--template), read and validated, or
|
|
545
|
+
* null for none. Throws, before the browser starts, when the check is not
|
|
546
|
+
* recorded, when the template is not valid, or when the file it would write is
|
|
547
|
+
* one the project keeps there and a check did not write.
|
|
548
|
+
*/
|
|
549
|
+
export function prepareTemplateReport(options, outDir) {
|
|
550
|
+
if (options.template === undefined)
|
|
551
|
+
return null;
|
|
552
|
+
const needsRecord = templateRecordError(options);
|
|
553
|
+
if (needsRecord)
|
|
554
|
+
throw new Error(needsRecord);
|
|
555
|
+
const template = readReportTemplate(options.template);
|
|
556
|
+
const file = path.join(outDir, reportFileOf(template));
|
|
557
|
+
let text = null;
|
|
558
|
+
try {
|
|
559
|
+
text = fs.readFileSync(file, "utf8");
|
|
560
|
+
}
|
|
561
|
+
catch (err) {
|
|
562
|
+
if (err.code !== "ENOENT")
|
|
563
|
+
throw err;
|
|
564
|
+
}
|
|
565
|
+
if (text !== null && !isGeneratedReport(text))
|
|
566
|
+
throw new Error(`${file} is not a report SceneScout wrote, so a check will not overwrite it. Move or rename it, name another file in the template's "file", or pass --out to write the check somewhere else`);
|
|
567
|
+
return template;
|
|
568
|
+
}
|
|
569
|
+
/** Write the template's report beside replay.html, after check.json and the replay page, so the manifest hashes the files as written. Returns its file name. */
|
|
570
|
+
export function writeTemplateReport(outDir, result, template, meta) {
|
|
571
|
+
const file = reportFileOf(template);
|
|
572
|
+
const evidence = evidenceIndex(outDir, evidenceFiles(result));
|
|
573
|
+
fs.writeFileSync(path.join(outDir, file), buildTemplateReport(result, template, evidence, meta));
|
|
574
|
+
return file;
|
|
575
|
+
}
|
|
515
576
|
/**
|
|
516
577
|
* Why a recorded check (--record or --video) must not start, or null: a
|
|
517
578
|
* replay.html in the output folder that a check did not write is the
|
package/dist/ci-run.js
CHANGED
|
@@ -22,9 +22,10 @@ import { CreateMessageRequestSchema, ErrorCode, McpError } from "@modelcontextpr
|
|
|
22
22
|
import { DEDUP_JUDGE_CAPABILITY, durationText, JUDGE_CALL_MS, JUDGE_MAX_OUTPUT_TOKENS, JUDGE_SYSTEM, JUDGE_TOOL, judgeKickoffOf, samplingResultOf, } from "./engine/dedup.js";
|
|
23
23
|
import { CAPTURE_MARGIN, parseCaptureResult, rebaseUrl, SHOT_FILES, SHOTS_DIRNAME } from "./engine/capture.js";
|
|
24
24
|
import { addUsage, attachFailure, budgetSpend, wallLeftMs, guardToolArgs, CAPTURE_TOOLS, childEnv, ciCaptureKickoff, ciCaptureSystemPrompt, CI_DIRNAME, ciExitCode, ciKickoff, ciSarif, ciSummaryJson, ciSummaryMarkdown, ciSystemPrompt, ciToolArgs, ciTools, describeStop, findingsThisRun, newBudget, NO_USAGE, readFindings, redactKeys, settleTurn, takeTurn, toolResultText, usageLine, } from "./engine/ci.js";
|
|
25
|
-
import { ciLaneKickoff, ciLaneSystemPrompt, crawlFoundNothing, crawlNotes, LANE_TOOLS, mergeLaneStops, PLAN_CRAWL_ROUNDS, planCiLanes, PLANNER_SESSION, } from "./engine/ci-lanes.js";
|
|
25
|
+
import { ciLaneKickoff, ciLaneSystemPrompt, crawlFoundNothing, crawlNotes, LANE_TOOLS, mergeLaneStops, PLAN_CRAWL_ROUNDS, planCiLanes, plannedRoutes, planReplayLanes, PLANNER_SESSION, } from "./engine/ci-lanes.js";
|
|
26
26
|
import { resolveTimeLimits } from "./engine/limits.js";
|
|
27
|
-
import { MEMORY_DIRNAME, writeSelfIgnore } from "./engine/memory.js";
|
|
27
|
+
import { loadRunRecord, MEMORY_DIRNAME, noteFromRunOnDisk, readRunRecordsOnDisk, writeSelfIgnore } from "./engine/memory.js";
|
|
28
|
+
import { carryForward, continueLines, continueExhausted, continuePlan, DEFAULT_TURNS_PER_PAGE, pageCap, takenPages, fromRunLine, isPattern, pathLine, prefixOutcome, prefixPlan, replayLines, replayPlan, } from "./engine/from-run.js";
|
|
28
29
|
import { sarifFilesFor } from "./engine/sarif.js";
|
|
29
30
|
import { decodePng, diffImages, encodePng } from "./engine/png.js";
|
|
30
31
|
import { AnthropicConversation, backoffMs, errorMessage, MalformedReply, OpenAIConversation, retryable, retryAfterMs, } from "./engine/provider.js";
|
|
@@ -434,43 +435,119 @@ function takesSession(t) {
|
|
|
434
435
|
}
|
|
435
436
|
const messageOf = (err) => (err instanceof Error ? err.message : String(err));
|
|
436
437
|
/**
|
|
437
|
-
*
|
|
438
|
-
*
|
|
439
|
-
*
|
|
440
|
-
*
|
|
441
|
-
*
|
|
442
|
-
*
|
|
443
|
-
* runs the agent loop there with its own conversation, draws its turns from
|
|
444
|
-
* the run's one budget, and is closed when it ends. Their findings are already
|
|
445
|
-
* one: every session files into the project's one memory, whose dedup folds a
|
|
446
|
-
* defect two lanes filed. `outcome` is absent when there was nothing to split,
|
|
447
|
-
* and the run explores in one loop instead.
|
|
438
|
+
* Take the earlier run's path to a page (from-run.ts prefixPlan) in one
|
|
439
|
+
* scout_run_plan call, which the write policy governs like any other step.
|
|
440
|
+
* Nothing happens when the record never reached the page or opened it by its
|
|
441
|
+
* address. When a step cannot be repeated (the record cannot say it, or the
|
|
442
|
+
* page changed), the session navigates to the page directly instead, if it is
|
|
443
|
+
* not a pattern. Returns the note for the record and the line the model is told.
|
|
448
444
|
*/
|
|
449
|
-
async function
|
|
450
|
-
|
|
451
|
-
|
|
452
|
-
|
|
445
|
+
async function reachByPath(o) {
|
|
446
|
+
if (!o.route)
|
|
447
|
+
return { lines: [] };
|
|
448
|
+
const plan = prefixPlan(o.record, o.route);
|
|
449
|
+
if (!plan || ("steps" in plan && plan.direct))
|
|
450
|
+
return { lines: [] };
|
|
451
|
+
let why;
|
|
452
|
+
let steps = 0;
|
|
453
|
+
if ("cannot" in plan)
|
|
454
|
+
why = `its path cannot be repeated: ${plan.cannot}`;
|
|
455
|
+
else {
|
|
456
|
+
steps = plan.steps.length;
|
|
457
|
+
const r = await o.host
|
|
458
|
+
.call("scout_run_plan", { session: o.session, steps: plan.steps, task: `Taking the earlier run's path to ${o.route}` }, Math.min(240_000, o.timeLeft()))
|
|
459
|
+
.catch((err) => ({ text: `ERROR: ${messageOf(err)}`, isError: true }));
|
|
460
|
+
const out = prefixOutcome(r.text.replace(/^\[session [^\]\n]*\]\n/, ""), steps, o.route);
|
|
461
|
+
if (out.ok) {
|
|
462
|
+
o.say(`reached ${o.route} by the earlier run's path (${steps} step(s)).`);
|
|
463
|
+
return {
|
|
464
|
+
note: { session: o.session, target: o.route, steps, outcome: "replayed" },
|
|
465
|
+
lines: [`Your browser reached ${o.route} the way the earlier run did: ${pathLine(plan.steps)}. Start there.`],
|
|
466
|
+
};
|
|
467
|
+
}
|
|
468
|
+
why = out.why;
|
|
469
|
+
}
|
|
470
|
+
o.say(`could not take the earlier run's path to ${o.route} (${why}); navigating to it instead.`);
|
|
471
|
+
const direct = isPattern(o.route)
|
|
472
|
+
? undefined
|
|
473
|
+
: await o.host
|
|
474
|
+
.call("scout_navigate", { session: o.session, target: o.route, task: `Opening ${o.route}` }, Math.min(120_000, o.timeLeft()))
|
|
475
|
+
.catch((err) => ({ text: `ERROR: ${messageOf(err)}`, isError: true }));
|
|
476
|
+
const opened = !!direct && !direct.isError && !/^(ERROR|REFUSED):/.test(direct.text.replace(/^\[session [^\]\n]*\]\n/, ""));
|
|
477
|
+
return {
|
|
478
|
+
note: { session: o.session, target: o.route, steps, outcome: opened ? "navigated" : "unreached", why: why.slice(0, 300) },
|
|
479
|
+
lines: [
|
|
480
|
+
opened
|
|
481
|
+
? `The earlier run's path to ${o.route} could not be repeated (${why}), so your browser was sent there directly.`
|
|
482
|
+
: `The earlier run's path to ${o.route} could not be repeated (${why}); reach it from the page you are on.`,
|
|
483
|
+
],
|
|
484
|
+
};
|
|
485
|
+
}
|
|
486
|
+
/** The planning crawl's notes with these routes added, each with no note: routes a record knew that the crawl did not list. */
|
|
487
|
+
function withRoutes(notes, routes) {
|
|
488
|
+
const out = new Map(notes);
|
|
489
|
+
for (const r of routes)
|
|
490
|
+
if (!out.has(r))
|
|
491
|
+
out.set(r, []);
|
|
492
|
+
return out;
|
|
493
|
+
}
|
|
494
|
+
/**
|
|
495
|
+
* The record this run's report left in the project's memory (from-run.ts), for
|
|
496
|
+
* ci.json: the newest written since the run started. A continued run's carries
|
|
497
|
+
* the record it continued, so a chain of runs each given the last one's
|
|
498
|
+
* ci.json accumulates what the chain covered. Never fatal: without one, ci.json
|
|
499
|
+
* has no record, and the log says why.
|
|
500
|
+
*/
|
|
501
|
+
function thisRunsRecord(o) {
|
|
502
|
+
let own;
|
|
503
|
+
try {
|
|
504
|
+
own = readRunRecordsOnDisk(o.projectDir)
|
|
505
|
+
.filter((r) => r.at >= o.since)
|
|
506
|
+
.at(-1);
|
|
507
|
+
}
|
|
508
|
+
catch (err) {
|
|
509
|
+
o.log(`The run's record could not be read from the project's memory, so ci.json holds none: ${messageOf(err)}`);
|
|
510
|
+
return undefined;
|
|
511
|
+
}
|
|
512
|
+
if (!own) {
|
|
513
|
+
o.log("The report left no run record in the project's memory, so ci.json holds none.");
|
|
514
|
+
return undefined;
|
|
515
|
+
}
|
|
516
|
+
const withPrefixes = {
|
|
517
|
+
...own,
|
|
518
|
+
...(o.prefixes && o.prefixes.length > 0 ? { prefixes: [...o.prefixes] } : {}),
|
|
519
|
+
...(o.continuedFresh ? { continuedFresh: true } : {}),
|
|
520
|
+
...(o.assigned && o.assigned.length > 0 ? { assigned: [...new Set(o.assigned)] } : {}),
|
|
521
|
+
};
|
|
522
|
+
return o.continued ? carryForward(o.continued, withPrefixes) : withPrefixes;
|
|
523
|
+
}
|
|
524
|
+
/**
|
|
525
|
+
* The planning crawl a run split into lanes, or a continued run, starts with: a
|
|
526
|
+
* snapshot from the planner's session (attaching harvests no links; a snapshot
|
|
527
|
+
* of the page it landed on does), then a crawl repeated while it finds routes,
|
|
528
|
+
* up to PLAN_CRAWL_ROUNDS. No model call: only its time counts.
|
|
529
|
+
*/
|
|
530
|
+
async function planningCrawl(o) {
|
|
453
531
|
// What went wrong while planning, so a plan left with nothing to split says why rather than blaming the app.
|
|
454
532
|
let planningFailed;
|
|
455
533
|
/** One planner call: its text, or undefined after saying why it failed. */
|
|
456
534
|
const plannerCall = async (tool, maxMs, what) => {
|
|
457
535
|
try {
|
|
458
|
-
const r = await host.call(tool, { session: PLANNER_SESSION }, Math.min(maxMs, timeLeft()));
|
|
536
|
+
const r = await o.host.call(tool, { session: PLANNER_SESSION }, Math.min(maxMs, o.timeLeft()));
|
|
459
537
|
if (r.isError)
|
|
460
538
|
throw new Error(r.text.replace(/^ERROR:\s*/, ""));
|
|
461
539
|
return r.text;
|
|
462
540
|
}
|
|
463
541
|
catch (err) {
|
|
464
542
|
planningFailed = `the planning ${what} failed: ${messageOf(err).slice(0, 300)}`;
|
|
465
|
-
log(
|
|
543
|
+
o.log(`${o.label}: ${planningFailed}.`);
|
|
466
544
|
return undefined;
|
|
467
545
|
}
|
|
468
546
|
};
|
|
469
|
-
|
|
470
|
-
if (timeLeft() > 0)
|
|
547
|
+
if (o.timeLeft() > 0)
|
|
471
548
|
await plannerCall("scout_snapshot", 120_000, "snapshot");
|
|
472
549
|
const notes = new Map();
|
|
473
|
-
for (let round = 0; round < PLAN_CRAWL_ROUNDS && timeLeft() > 0; round += 1) {
|
|
550
|
+
for (let round = 0; round < PLAN_CRAWL_ROUNDS && o.timeLeft() > 0; round += 1) {
|
|
474
551
|
const text = await plannerCall("scout_crawl", 600_000, "crawl");
|
|
475
552
|
if (text === undefined)
|
|
476
553
|
break;
|
|
@@ -481,14 +558,24 @@ async function exploreInLanes(o) {
|
|
|
481
558
|
if (crawlFoundNothing(text) || notes.size === known)
|
|
482
559
|
break;
|
|
483
560
|
}
|
|
484
|
-
|
|
485
|
-
|
|
486
|
-
|
|
487
|
-
|
|
488
|
-
|
|
489
|
-
|
|
490
|
-
|
|
491
|
-
|
|
561
|
+
return { notes, ...(planningFailed ? { planningFailed } : {}) };
|
|
562
|
+
}
|
|
563
|
+
/**
|
|
564
|
+
* A run split into lanes (--lanes): plan, then run every lane at once, then
|
|
565
|
+
* fold what they did. The plan is a snapshot and a crawl from the planner's
|
|
566
|
+
* session (no model call: only their time counts), the crawl repeated while it
|
|
567
|
+
* finds routes, and the split brief.ts makes of them. Each lane attaches its
|
|
568
|
+
* own session on the run's target URL, as the planner did (the engine resolves
|
|
569
|
+
* every path against the URL a session attached with), opens its first route,
|
|
570
|
+
* runs the agent loop there with its own conversation, draws its turns from
|
|
571
|
+
* the run's one budget, and is closed when it ends. Their findings are already
|
|
572
|
+
* one: every session files into the project's one memory, whose dedup folds a
|
|
573
|
+
* defect two lanes filed. `outcome` is absent when there was nothing to split,
|
|
574
|
+
* and the run explores in one loop instead.
|
|
575
|
+
*/
|
|
576
|
+
async function exploreInLanes(o) {
|
|
577
|
+
const { host, options, log, now, budget, plan } = o;
|
|
578
|
+
const timeLeft = () => wallLeftMs(budgetSpend(budget), options.caps, now());
|
|
492
579
|
if (plan.oneLoop) {
|
|
493
580
|
log(`Lanes: ${plan.oneLoop}. Exploring in one loop.`);
|
|
494
581
|
return { lanes: { asked: options.lanes, sessions: [], oneLoop: plan.oneLoop } };
|
|
@@ -539,7 +626,9 @@ async function exploreInLanes(o) {
|
|
|
539
626
|
on = lane.url;
|
|
540
627
|
}
|
|
541
628
|
say(`attached on ${new URL(on).pathname}, owning ${lane.modules.join(", ")}.`);
|
|
542
|
-
|
|
629
|
+
// A continued run: the earlier run's path to this lane's first page, when it got there by acting on another.
|
|
630
|
+
const reached = o.reach && timeLeft() > 0 ? await o.reach(lane.session, lane.routes[0], say) : [];
|
|
631
|
+
return await exploreLane(lane, on, say, { ...result, attached: true }, reached);
|
|
543
632
|
}
|
|
544
633
|
catch (err) {
|
|
545
634
|
// Not a cap and not the model's API: the lane itself broke. Reported as the lane's, and the run's (mergeLaneStops), never dropped.
|
|
@@ -548,7 +637,7 @@ async function exploreInLanes(o) {
|
|
|
548
637
|
return { ...result, attached, stop: "could-not-start", stopDetail: why };
|
|
549
638
|
}
|
|
550
639
|
};
|
|
551
|
-
const exploreLane = async (lane, on, say, result) => {
|
|
640
|
+
const exploreLane = async (lane, on, say, result, reached = []) => {
|
|
552
641
|
const outcome = await agentLoop({
|
|
553
642
|
client: o.makeClient(system, tools, ciLaneKickoff({
|
|
554
643
|
lane: { ...lane, url: on },
|
|
@@ -559,6 +648,7 @@ async function exploreInLanes(o) {
|
|
|
559
648
|
level: options.level,
|
|
560
649
|
focus: options.focus,
|
|
561
650
|
caps: options.caps,
|
|
651
|
+
...(o.laneLines || reached.length > 0 ? { fromRun: [...(o.laneLines?.(lane) ?? []), ...reached] } : {}),
|
|
562
652
|
})),
|
|
563
653
|
host,
|
|
564
654
|
tools,
|
|
@@ -625,6 +715,11 @@ export async function runCi(options, resolved, deps) {
|
|
|
625
715
|
let reportWritten = false;
|
|
626
716
|
let capture;
|
|
627
717
|
let lanes;
|
|
718
|
+
let runFromRun;
|
|
719
|
+
let runRecord;
|
|
720
|
+
let earlier;
|
|
721
|
+
// The real clock, as the server's: this run's record is the one its report wrote after this.
|
|
722
|
+
const startedIso = new Date().toISOString();
|
|
628
723
|
let host = null;
|
|
629
724
|
// Set when the exploration ends (or never starts): the report and the close share FINISH_MS from then.
|
|
630
725
|
let finishBy = 0;
|
|
@@ -636,6 +731,17 @@ export async function runCi(options, resolved, deps) {
|
|
|
636
731
|
const judge = judgeAsk
|
|
637
732
|
? judgeHandler({ ask: judgeAsk, model: resolved.model, spend: budget, caps: options.caps, calls: judgeCalls, secrets, now })
|
|
638
733
|
: undefined;
|
|
734
|
+
// The earlier run's record, read before anything starts: a run asked to continue or replay one must not quietly start fresh.
|
|
735
|
+
if (options.fromRun && !options.show) {
|
|
736
|
+
try {
|
|
737
|
+
earlier = loadRunRecord(options.fromRun.path);
|
|
738
|
+
}
|
|
739
|
+
catch (err) {
|
|
740
|
+
throw new Error(`--from-run: ${messageOf(err)}`);
|
|
741
|
+
}
|
|
742
|
+
if (options.fromRun.mode === "replay" && replayPlan(earlier).length === 0)
|
|
743
|
+
throw new Error(`--from-run: the run recorded in ${options.fromRun.path} took no steps, so there is nothing to replay`);
|
|
744
|
+
}
|
|
639
745
|
host = await (deps.startHost ?? startServer)(log, judge);
|
|
640
746
|
// What every session of the run attaches with: the planner's here, and each lane's when the run is split.
|
|
641
747
|
const attachArgs = (a) => ({
|
|
@@ -666,10 +772,67 @@ export async function runCi(options, resolved, deps) {
|
|
|
666
772
|
const listed = await host.tools();
|
|
667
773
|
const toolHost = host;
|
|
668
774
|
let captured = null;
|
|
669
|
-
|
|
775
|
+
// Named as it was given: a resolved path would put a local directory in the report and ci.json.
|
|
776
|
+
const from = options.fromRun ? (options.fromRun.given ?? options.fromRun.path) : "";
|
|
777
|
+
const replay = earlier && options.fromRun?.mode === "replay" ? replayPlan(earlier) : undefined;
|
|
778
|
+
const continuing = earlier && options.fromRun?.mode === "continue" ? earlier : undefined;
|
|
779
|
+
// A run split into lanes, or a continued one, crawls first: the lanes are split, and a continued run's routes ordered, from what it finds.
|
|
780
|
+
// A replay does not: its routes and steps are the record's.
|
|
781
|
+
const timeLeft = () => wallLeftMs(budgetSpend(budget), options.caps, now());
|
|
782
|
+
const plans = !options.show && !replay && (options.lanes > 1 || !!continuing);
|
|
783
|
+
const planned = plans ? await planningCrawl({ host, log, timeLeft, label: options.lanes > 1 ? "Lanes" : "From run" }) : undefined;
|
|
784
|
+
const planItems = continuing && planned ? continuePlan(continuing, plannedRoutes(options.url, planned.notes)) : undefined;
|
|
785
|
+
// A record with no work left anywhere: explore as a fresh run (the stable split, the landing page, no path), and say so.
|
|
786
|
+
const continuedFresh = !!planItems && continueExhausted(planItems);
|
|
787
|
+
const items = continuedFresh ? undefined : planItems;
|
|
788
|
+
// How many pages a continued run takes on, from its budget (from-run.ts pageCap): each lane's share of the turns, or the run's.
|
|
789
|
+
const perPage = options.fromRun?.turnsPerPage ?? DEFAULT_TURNS_PER_PAGE;
|
|
790
|
+
const cap = items ? pageCap(options.caps.turns, perPage) : undefined;
|
|
791
|
+
const laneCap = items ? pageCap(Math.floor(options.caps.turns / Math.max(1, options.lanes)), perPage) : undefined;
|
|
792
|
+
// The pages this run was given, kept in its record so the next continued run takes the next ones.
|
|
793
|
+
const assigned = [];
|
|
794
|
+
if (options.fromRun && earlier && !options.show) {
|
|
795
|
+
runFromRun = { mode: options.fromRun.mode, source: from, runId: earlier.runId, recordAt: earlier.at };
|
|
796
|
+
log(`From run: ${fromRunLine(runFromRun)}.`);
|
|
797
|
+
if (items)
|
|
798
|
+
log(`From run: ${items.filter((i) => i.tier === 1).length} route(s) it never worked on, ${items.filter((i) => i.tier === 2).length} with work left, ${items.filter((i) => i.tier === 3).length} worked through.`);
|
|
799
|
+
if (continuedFresh)
|
|
800
|
+
log("From run: the earlier run left no recorded work on any route, so this run explores as a fresh one.");
|
|
801
|
+
// Noted in the project's memory for the report, which the server writes. Never fatal: the run goes on.
|
|
802
|
+
try {
|
|
803
|
+
noteFromRunOnDisk(options.projectDir, { at: new Date().toISOString(), ...runFromRun });
|
|
804
|
+
}
|
|
805
|
+
catch (err) {
|
|
806
|
+
log(`From run: the project's memory could not be told, so the report will not name the earlier run: ${messageOf(err)}`);
|
|
807
|
+
}
|
|
808
|
+
}
|
|
809
|
+
// How a continued run's lanes (or its one loop) reached their first page by the earlier run's path, for the record.
|
|
810
|
+
const prefixes = [];
|
|
811
|
+
const reach = items
|
|
812
|
+
? async (session, route, say) => {
|
|
813
|
+
const r = await reachByPath({ host: toolHost, session, record: continuing, route, timeLeft, say });
|
|
814
|
+
if (r.note)
|
|
815
|
+
prefixes.push(r.note);
|
|
816
|
+
return r.lines;
|
|
817
|
+
}
|
|
818
|
+
: undefined;
|
|
819
|
+
const oneLoop = async () => {
|
|
670
820
|
const tools = ciTools(listed, options.show ? CAPTURE_TOOLS : undefined);
|
|
671
821
|
const system = options.show ? ciCaptureSystemPrompt() : ciSystemPrompt(loadPlaybook(packageRoot), options);
|
|
672
|
-
const
|
|
822
|
+
const reached = reach && items && !options.show ? await reach(PLANNER_SESSION, items[0]?.route, (l) => log(`From run: ${l}`)) : [];
|
|
823
|
+
if (items && cap !== undefined) {
|
|
824
|
+
const mine = takenPages(items, cap);
|
|
825
|
+
assigned.push(...mine);
|
|
826
|
+
log(`From run: takes on ${mine.join(", ")} (${cap} page(s) at ${perPage} turns a page).`);
|
|
827
|
+
}
|
|
828
|
+
const fromRun = planItems
|
|
829
|
+
? [...continueLines(planItems, { from, ...(cap !== undefined ? { cap } : {}) }), ...reached]
|
|
830
|
+
: replay && replay.length > 0
|
|
831
|
+
? replayLines(replay[0], { from })
|
|
832
|
+
: undefined;
|
|
833
|
+
const kickoff = options.show
|
|
834
|
+
? ciCaptureKickoff({ url: options.url, show: options.show })
|
|
835
|
+
: ciKickoff({ ...options, ...(fromRun && fromRun.length > 0 ? { fromRunLines: fromRun } : {}) });
|
|
673
836
|
return agentLoop({
|
|
674
837
|
client: deps.makeClient(system, tools, kickoff),
|
|
675
838
|
host: toolHost,
|
|
@@ -685,8 +848,56 @@ export async function runCi(options, resolved, deps) {
|
|
|
685
848
|
},
|
|
686
849
|
});
|
|
687
850
|
};
|
|
688
|
-
|
|
689
|
-
|
|
851
|
+
// A replay has as many lanes as the recorded run had sessions; otherwise the planning crawl is split as asked.
|
|
852
|
+
const lanePlan = options.show
|
|
853
|
+
? undefined
|
|
854
|
+
: replay
|
|
855
|
+
? planReplayLanes({ target: options.url, lanes: replay, ...(options.focus ? { focus: options.focus } : {}) })
|
|
856
|
+
: options.lanes > 1 && planned
|
|
857
|
+
? planCiLanes({
|
|
858
|
+
target: options.url,
|
|
859
|
+
// A continued run splits every route it orders, the record's own included, so none is left out of every lane.
|
|
860
|
+
notes: items
|
|
861
|
+
? withRoutes(planned.notes, items.map((i) => i.route))
|
|
862
|
+
: planned.notes,
|
|
863
|
+
count: options.lanes,
|
|
864
|
+
focus: options.focus,
|
|
865
|
+
mode: options.mode,
|
|
866
|
+
...(planned.planningFailed ? { planningFailed: planned.planningFailed } : {}),
|
|
867
|
+
...(items ? { order: items.map((i) => i.route) } : {}),
|
|
868
|
+
})
|
|
869
|
+
: undefined;
|
|
870
|
+
if (replay && replay.length > 1 && replay.length !== options.lanes)
|
|
871
|
+
log(`From run: replaying the recorded run's ${replay.length} sessions as ${replay.length} lanes.`);
|
|
872
|
+
if (lanePlan && (lanePlan.lanes.length > 0 || options.lanes > 1)) {
|
|
873
|
+
// A replayed lane is told its own session's steps: the plan keeps the recorded sessions' order.
|
|
874
|
+
const replayed = new Map(replay ? lanePlan.lanes.map((l, i) => [l.session, replay[i]]) : []);
|
|
875
|
+
const laneLines = planItems
|
|
876
|
+
? (lane) => {
|
|
877
|
+
if (items && laneCap !== undefined) {
|
|
878
|
+
const mine = takenPages(items, laneCap, lane.routes);
|
|
879
|
+
assigned.push(...mine);
|
|
880
|
+
log(`From run: lane ${lane.session} takes on ${mine.join(", ") || "no page with work left"} (${laneCap} page(s) at ${perPage} turns a page).`);
|
|
881
|
+
}
|
|
882
|
+
return continueLines(planItems, { from, only: lane.routes, ...(laneCap !== undefined ? { cap: laneCap } : {}) });
|
|
883
|
+
}
|
|
884
|
+
: replay
|
|
885
|
+
? (lane) => replayLines(replayed.get(lane.session), { from })
|
|
886
|
+
: undefined;
|
|
887
|
+
const split = await exploreInLanes({
|
|
888
|
+
host,
|
|
889
|
+
listed,
|
|
890
|
+
options,
|
|
891
|
+
makeClient: deps.makeClient,
|
|
892
|
+
log,
|
|
893
|
+
now,
|
|
894
|
+
budget,
|
|
895
|
+
attachArgs,
|
|
896
|
+
attachMs,
|
|
897
|
+
plan: lanePlan,
|
|
898
|
+
...(laneLines ? { laneLines } : {}),
|
|
899
|
+
...(reach ? { reach } : {}),
|
|
900
|
+
});
|
|
690
901
|
lanes = split.lanes;
|
|
691
902
|
outcome = split.outcome ?? (await oneLoop());
|
|
692
903
|
}
|
|
@@ -709,6 +920,16 @@ export async function runCi(options, resolved, deps) {
|
|
|
709
920
|
reportWritten = !report.isError && !/^ERROR:/.test(report.text) && fs.existsSync(path.join(options.projectDir, MEMORY_DIRNAME, "report.md"));
|
|
710
921
|
if (!reportWritten)
|
|
711
922
|
log(`The report could not be generated: ${report.text.slice(0, 400)}`);
|
|
923
|
+
else
|
|
924
|
+
runRecord = thisRunsRecord({
|
|
925
|
+
projectDir: options.projectDir,
|
|
926
|
+
since: startedIso,
|
|
927
|
+
continued: earlier && options.fromRun?.mode === "continue" ? earlier : undefined,
|
|
928
|
+
prefixes,
|
|
929
|
+
continuedFresh,
|
|
930
|
+
assigned,
|
|
931
|
+
log,
|
|
932
|
+
});
|
|
712
933
|
}
|
|
713
934
|
}
|
|
714
935
|
}
|
|
@@ -748,6 +969,8 @@ export async function runCi(options, resolved, deps) {
|
|
|
748
969
|
findings: findingsThisRun(before, readMemoryFindings(options.projectDir)),
|
|
749
970
|
...(capture ? { capture } : {}),
|
|
750
971
|
...(lanes ? { lanes } : {}),
|
|
972
|
+
...(runFromRun ? { fromRun: runFromRun } : {}),
|
|
973
|
+
...(runRecord ? { record: runRecord } : {}),
|
|
751
974
|
...(options.show
|
|
752
975
|
? {}
|
|
753
976
|
: {
|
package/dist/cli.js
CHANGED
|
@@ -21,7 +21,8 @@ import { APPROX_DISK_MB, defaultAttachNote, defaultEngine, launchTarget, parseBr
|
|
|
21
21
|
import { CLIENT_LABELS, firstMessageHint, manualFor, parseClients, registerWithClient, vscodeBinary } from "./clients.js";
|
|
22
22
|
import { CLAUDE_CODE_NOT_NEEDED, CLI_NAME, desktopExtensionRoots, diagnose, doctorAllGood, findDesktopExtension, installClosing, ensureCommand, findOnUserPath, installSkill, isEphemeralRoot, launchCommand, manualRegisterCommand, planCommand, registerMcp, resolveClaudeDir, spawnRunner, } from "./installer.js";
|
|
23
23
|
import { downloadBrowsers, presentBrowsers } from "./installer.js";
|
|
24
|
-
import { baselinesDirOf, clearReplayOutput, replayPageConflict, defaultCheckDir, readCheckInputs, runCheck } from "./check-run.js";
|
|
24
|
+
import { baselinesDirOf, clearReplayOutput, replayPageConflict, defaultCheckDir, prepareTemplateReport, readCheckInputs, runCheck, writeTemplateReport, } from "./check-run.js";
|
|
25
|
+
import { reportFileOf } from "./engine/check-report.js";
|
|
25
26
|
import { buildCheckReplayHtml, commitOf, REPLAY_FILE, replayFrames, replayVideos } from "./engine/check-replay.js";
|
|
26
27
|
import { recordChoice } from "./engine/capture.js";
|
|
27
28
|
import { httpClient, httpJudgeAsk, runCi } from "./ci-run.js";
|
|
@@ -112,7 +113,10 @@ Usage:
|
|
|
112
113
|
and write replay.html beside the report, role → journey → step (default:
|
|
113
114
|
SCENESCOUT_RECORD, else off; the frames go in replay-frames/);
|
|
114
115
|
--video [on|off]: record a WebM of each saved flow, and only of the flows,
|
|
115
|
-
into replay-videos/, played on replay.html beside its steps (default off)
|
|
116
|
+
into replay-videos/, played on replay.html beside its steps (default off);
|
|
117
|
+
--template file.json: on a recorded check, also write a test report laid out
|
|
118
|
+
as the template says (report.html unless it names a file), with the frames,
|
|
119
|
+
deviations and a SHA-256 manifest of the evidence; needs --record)
|
|
116
120
|
Exit code: 0 passed, 1 failed the gate, 2 could not run.
|
|
117
121
|
scenescout ci <url> An exploratory run with no person present: a model reached through its API
|
|
118
122
|
drives the tools by the SceneScout method and the run ends in the report.
|
|
@@ -580,15 +584,18 @@ async function check(args) {
|
|
|
580
584
|
let options = parsed.options;
|
|
581
585
|
const outDir = options.outDir ?? defaultCheckDir(options.projectDir);
|
|
582
586
|
let inputs;
|
|
587
|
+
let template = null;
|
|
583
588
|
try {
|
|
584
589
|
// --record, else SCENESCOUT_RECORD, else off.
|
|
585
590
|
options = { ...options, record: recordChoice(options.record, process.env) };
|
|
591
|
+
// --template: read and checked before the browser starts, so a bad template never costs a run.
|
|
592
|
+
template = prepareTemplateReport(options, outDir);
|
|
586
593
|
// A replay.html the project keeps there is its own: a recorded check refuses to start rather than overwrite it.
|
|
587
594
|
const conflict = options.record || options.video ? replayPageConflict(outDir) : null;
|
|
588
595
|
if (conflict)
|
|
589
596
|
throw new Error(conflict);
|
|
590
597
|
// A replay page an earlier run left must never be read, or uploaded, as this run's.
|
|
591
|
-
clearReplayOutput(outDir);
|
|
598
|
+
clearReplayOutput(outDir, template ? reportFileOf(template) : undefined);
|
|
592
599
|
inputs = readCheckInputs(options);
|
|
593
600
|
}
|
|
594
601
|
catch (err) {
|
|
@@ -615,7 +622,11 @@ async function check(args) {
|
|
|
615
622
|
console.error(`scenescout check: could not run a saved flow: ${refused}`);
|
|
616
623
|
process.exit(EXIT.error);
|
|
617
624
|
}
|
|
625
|
+
// Recorded in check.json, so the next check removes exactly this file and the action finds it.
|
|
626
|
+
if (template && result.replay)
|
|
627
|
+
result = { ...result, testReport: reportFileOf(template) };
|
|
618
628
|
const markdown = formatCheck(result);
|
|
629
|
+
let reportFile = null;
|
|
619
630
|
try {
|
|
620
631
|
const version = packageVersion();
|
|
621
632
|
if (!options.outDir)
|
|
@@ -645,6 +656,8 @@ async function check(args) {
|
|
|
645
656
|
couldNotRun,
|
|
646
657
|
});
|
|
647
658
|
fs.writeFileSync(path.join(outDir, REPLAY_FILE), html);
|
|
659
|
+
if (template)
|
|
660
|
+
reportFile = writeTemplateReport(outDir, result, template, { version, commit: commitOf(process.env) });
|
|
648
661
|
}
|
|
649
662
|
// On GitHub Actions the verdict also goes on the run's summary page.
|
|
650
663
|
if (process.env.GITHUB_STEP_SUMMARY)
|
|
@@ -659,6 +672,8 @@ async function check(args) {
|
|
|
659
672
|
console.log(`Wrote report.md, check.sarif and check.json to ${outDir}`);
|
|
660
673
|
if (result.replay)
|
|
661
674
|
console.log(`Wrote ${REPLAY_FILE} to ${outDir}, with ${replayFrames(result.replay).length} frame(s) and ${replayVideos(result.replay).length} journey video(s) beside it: open it in a browser to see each step`);
|
|
675
|
+
if (reportFile)
|
|
676
|
+
console.log(`Wrote ${reportFile} to ${outDir}: the test report the template lays out, with a SHA-256 manifest of its evidence`);
|
|
662
677
|
const pictured = result.baselines?.results.filter((r) => r.files).length ?? 0;
|
|
663
678
|
if (pictured > 0)
|
|
664
679
|
console.log(`Wrote the pictures of ${pictured} changed baseline(s) under ${path.join(outDir, VISUAL_DIRNAME)}`);
|
|
@@ -762,7 +777,7 @@ async function ci(args) {
|
|
|
762
777
|
console.error(redactKeys(`scenescout ci: ${message}`, secrets));
|
|
763
778
|
process.exit(EXIT_CI.couldNotRun);
|
|
764
779
|
};
|
|
765
|
-
const parsed = parseCiArgs(args, process.cwd());
|
|
780
|
+
const parsed = parseCiArgs(args, process.cwd(), process.env);
|
|
766
781
|
if (!parsed.ok)
|
|
767
782
|
return fail(parsed.error);
|
|
768
783
|
const options = parsed.options;
|