scenescout 3.2.0 → 3.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,96 @@
1
1
  # scenescout
2
2
 
3
+ ## 3.4.0
4
+
5
+ ### Minor Changes
6
+
7
+ - 04f0f22: Measure whether a lane's confidence means anything, and make one lane rubric serve a whole wave.
8
+
9
+ Every lane has been told its confidence must be calibrated — "0.5 means a coin flip, 0.95 means you would bet on it" — and the number was then averaged into one line and discarded. Nothing was stored, so nothing could ever be checked, and a confidence nobody checks is decoration.
10
+
11
+ Lane decisions are now kept, and the report carries a calibration section: how often a decision at a stated confidence matched a finding the project holds, bucketed, with an expected calibration error. It says plainly what the number is not — agreement between the lanes and the bar the project applies, over its whole history, not evidence that the app is broken — and names the two ways a lane is counted wrong through no fault of its own. A decision naming no failing endpoint cannot be looked up at all, so it is excluded and disclosed rather than scored as a miss. Below eight checkable decisions the figure is withheld, but the section says so rather than vanishing. Where findings have since been re-tested through `scout_verify`, those verdicts are reported beside it, because they *are* evidence about the app.
12
+
13
+ Separately, the lane name used to sit in the second sentence of the instruction every lane receives, so two lanes' prompts diverged almost immediately and shared no prefix. The rubric is now identical for every lane in a wave and the name is the last thing said, which makes it one cacheable prefix instead of one per lane.
14
+
15
+ ## 3.3.0
16
+
17
+ ### Minor Changes
18
+
19
+ - 0778653: Report the page contradicting the server: a refused list shown as an empty state, and a refused save shown as a success.
20
+
21
+ Two of the most expensive bugs a web app ships were invisible to every oracle that watches one side of the wire. A list request is refused with a 403 and the page renders its empty state, so the user is told they have nothing when the truth is that nothing could be loaded — which is how a permission regression reaches production without anyone noticing. A save is refused and the page says "Saved", so the user walks away believing their work is stored.
22
+
23
+ Neither is a crash. The HTTP oracle already saw the refusal and reported it as a medium, indistinguishable from the dozens of expected 401s an auth probe produces; the defect is not the refusal but the page contradicting it. Both now raise a high-severity `refused_empty` or `false_success` violation on the action that caused them, naming the endpoint and quoting what the user was shown instead.
24
+
25
+ The rules pair an exact half with a fuzzy one — a status code either is an error or is not, and the page half is never enough alone — so a page that is refused and says so raises nothing. Against the demo app, whose only 4xx is a missing image, they are silent.
26
+ - 22cd8ca: Add `scout_lane_brief`, and remember how a login state is regenerated.
27
+
28
+ Dividing an app between parallel lanes by hand fails in two ways a finished run cannot tell apart from success. Lanes overlap, so two browsers audit the same register while a third module is never opened — and route coverage reads complete either way, because both lanes visiting a route makes it covered. And lanes launch underspecified: in one real four-session run the first two sessions acted with no task set, so the person watching the live view saw browsers clicking through their app with nothing to say why.
29
+
30
+ `scout_lane_brief {lanes, goal}` computes the split instead. Routes are grouped into whole modules by their first path segment, so a lane that owns everything under one module carries state between its own steps rather than re-learning the app on every route, and modules are dealt out so the lanes come out within a route or two of each other. It returns each lane's session name, the `objective` to attach with, the routes it owns, and the two rules a hand-written brief keeps dropping. The same routes always produce the same split, so a lane that has to be re-run is handed the same brief.
31
+
32
+ Separately, `scout_note` gains a `setup` section for how to get an app testable at all, and the `⚠ AUTH FAILED` message now quotes back whatever an earlier run recorded there. A storage state expires on a timer nobody remembers, and "regenerate it" is advice the reader already had; the command that worked last time is the part worth keeping.
33
+ - 8f06c37: The report is a worklist again.
34
+
35
+ A project that has been tested for a while accumulates findings, and the report printed every one of them in full. On one real project that was 737 findings, 1.75 MB, of which 423 were open but unverified by that run and 314 were already fixed — and the eleven findings the run had actually just made were buried in the middle of it. A document nobody opens is not a report.
36
+
37
+ Findings from this run still print in full. Findings from earlier runs, and resolved ones, are now an index: one row each with the id, severity, how long ago it was last seen, how many runs have seen it, and the title. The same project's report becomes 113 KB, 94% smaller, with nothing lost — every id is there, and `scout_report {history: "full"}` prints all of it exactly as before, which is what to use when handing the document to someone who cannot read the project's memory.
38
+
39
+ Age is on every row because it is what decides whether an unverified finding is worth re-testing: one nobody has re-confirmed in four months is a different proposition from one seen last week.
40
+ - 93f9937: Add `scout_verify`: re-test what earlier runs left open, and record what each re-test found.
41
+
42
+ The report has always carried two kinds of finding and been honest that they are not the same thing — what this run saw, and what some earlier run saw. The second kind was labelled historical and unverified, which is accurate and almost useless: a reader cannot tell a bug fixed three weeks ago from one still costing users money today, and neither can the next run. Closing that by hand meant copying each finding's route and evidence out of the report, re-walking them one at a time, and calling `scout_resolve` on the ones that were gone. One project's history held over three hundred.
43
+
44
+ `scout_verify` called bare returns the open findings in the order to re-test them — worst route first, grouped so a route is walked once rather than once per finding — each with the evidence that identifies it and the steps that produced it. `scout_verify {ids}` narrows it, and names any that are not open rather than quietly shortening the list.
45
+
46
+ After re-testing one, `scout_verify {id, verdict, note}` records it: `gone` resolves it, `present` stamps it confirmed so the report dates the confirmation instead of calling it unverified, and `changed` keeps it open and says the behaviour differs. The history index gains a "Re-tested" column, so a reader can see at a glance which of it is still believed.
47
+ - b706095: Make the live board answer, at a glance, which session is stuck and which is in trouble.
48
+
49
+ On a board of eleven cards the questions actually being asked are "which one has been on the same thing for ten minutes", "which one is having trouble", and "where is the one on the orders register". The board could answer none of them, and two of the three answers were already in the status payload on every poll and reached nobody: `taskSince` was rendered only inside the close-up, and a step whose result went wrong was only ever a red word in a six-line feed somebody had to read.
50
+
51
+ Each card now carries a line under its task: how long the session has been on it, and how many of its recent steps went wrong — counted with the same rule the feed colours red, so a card and the feed beneath it cannot disagree. The line is absent on a session that has stated no task and had no trouble.
52
+
53
+ The header gains a filter over everything a card shows — name, role, objective, task, page, tool. It hides cards and nothing else: a filtered-out session is still running, still streaming and still counted in the header, and a filter that matches nothing says so rather than showing a blank page that reads as every session having gone.
54
+
55
+ The close-up's timeline can be walked from the keyboard: arrows step, Home and End jump to the ends, and Space returns to what the session is showing now. Scrubbing a long run by clicking 16-pixel ticks was the thing a mouse was worst at, and the run worth examining is always the one with hundreds of steps.
56
+ - 2e5a808: What a run shows about itself.
57
+
58
+ A four-session validation pass exposed several things the engine knew but never said. All of them are fixed here.
59
+
60
+ **Recording covers the breadth pass.** A crawl now keeps a frame per route it visits, and a snapshot keeps one too. The run that prompted this kept 9 frames out of 67 actions, none of them from the 30 routes a crawl had just swept — the evidence artifact was missing exactly where the coverage happened.
61
+
62
+ **A session says what it is doing from the moment it appears.** `scout_attach` takes a `task`, and puts up a placeholder when none is given, so a fresh card no longer reads "Nothing stated yet" while the session works. The placeholder is display only: it does not satisfy the requirement that an agent state its task before a tool acts.
63
+
64
+ **Several engines on one project no longer erase each other.** Each writes `status.<pid>.json` and its own token file, and `scenescout watch` lists every live engine with its address instead of finding only whichever attached last. The shared `status.json` is still written for older readers.
65
+
66
+ **The report says how the run was paced** — actions, span, median gap, longest gap, idle share and frames per session — and warns about a session that has held a browser with nothing to do for over five minutes. Idle share is labelled as time the browser waited for the agent, because it is not a measure of the engine.
67
+
68
+ **Memory stops growing without limit.** A route keeps its most recent states, capped, so a history that had reached 6,075 states and 36 MB — parsed and re-serialised on every save — is trimmed on open. States a finding points at are never dropped, and coverage is unchanged because it is asked per route.
69
+
70
+ **`scout_scan` says which saved logins have expired**, rather than leaving it to be discovered by attaching and landing on a login page.
71
+ - cfdc508: `scout_request` — call the app's own API as the session, with the UI bypassed.
72
+
73
+ A refusal shown by hiding or disabling a button is not a refusal. Confirming that the server refuses the same action is the most valuable check a permission pass makes, and until now it could only be done outside the tool, in a shell with curl and a hand-extracted token. None of that evidence reached the report: a whole validation run's permission matrices lived in shell history and went with it.
74
+
75
+ The request is made by the page, not beside it, which matters twice. It goes through the same interception the write policy is enforced on, so a safe-write session cannot reach past the policy by calling an endpoint instead of clicking it — the browser suite proves a replayed `DELETE` on a record the session did not create is refused exactly as a click would be. And it carries the session's own credentials, because it is the same origin with the same cookies. Bearer schemes work by replaying whatever `Authorization` header the app itself last sent, so nothing in the engine knows what a token looks like or where an app keeps one.
76
+
77
+ The result leads with the signature a finding should quote (`GET /api/admin/users 403`), then the timing, then the headers that decide whether two responses are genuinely identical — content-type, location, www-authenticate, retry-after, cache-control — then the body. Every call is recorded in the run's trail.
78
+
79
+ Paths are fenced to the attached origin, as navigation is: a session talks to its own app, and another host needs another session.
80
+ - b0f75ec: Wait for the requests an action fired, rather than a fixed sleep, and let a session ask to be slowed down.
81
+
82
+ Every action used to be followed by a flat 400 ms sleep. Measured against the demo app that was 54% of a snapshot's wall time, and a run of two hundred actions spent over a minute asleep — while any page slower than 400 ms was still read before it had finished changing. The engine already intercepts every request, so it now waits on what is actually in flight, with a quiet window after the last one starts and after the action itself, and the old constant survives as a ceiling instead of a floor. On the demo app a navigate costs 149 ms rather than 430, and a snapshot 446 rather than 740.
83
+
84
+ The same rule carries the opposite need. `scout_attach {paceMs}` and `scout_session {paceMs}` set a floor between actions so a person watching can follow along — useful when taking notes beside a run or demonstrating a flow. Unset, a session runs as fast as its page allows; `scout_session {paceMs}` with no `name` changes every attached session at once.
85
+
86
+ ### Patch Changes
87
+
88
+ - aa77244: Catch a client-side auth guard that redirects after the page has gone quiet.
89
+
90
+ Settling on the requests an action fired is faster than a fixed sleep and more patient with a slow page, but it cannot wait for something that has not been scheduled. A client-side auth guard issues no request until its timer fires, so the page goes quiet, the URL is read, and the gated route is recorded as reached — the bounce invisible, and a dead session along with it. The removed 400 ms sleep had been covering this by accident, and this project's own CI began failing intermittently on a 40 ms guard that a loaded runner delayed past the quiet window.
91
+
92
+ Where a bounce verdict is made — attach judging a storage state, navigate judging coverage — the URL is now watched until it has held still rather than read once. It is a window rather than a guarantee: a guard slower than it still lands after the verdict, and is caught on the next action. Only a session that was given credentials pays for it on every navigation, so an anonymous crawl keeps its full speed: navigate 179 ms and 470 ms per route, unchanged.
93
+
3
94
  ## 3.2.0
4
95
 
5
96
  ### Minor Changes
package/README.md CHANGED
@@ -272,13 +272,14 @@ Snapshots are cheap: re-snapshotting a route returns only *what changed*, with s
272
272
 
273
273
  ## 🧰 The toolbox
274
274
 
275
- 25 deterministic tools. The agent picks; you rarely call these by hand.
275
+ 26 deterministic tools. The agent picks; you rarely call these by hand.
276
276
 
277
277
  | Phase | Tools | What they do |
278
278
  |---|---|---|
279
279
  | **Set up** | `scout_playbook` `scout_scan` `scout_attach` `scout_session` | Hand the testing method to an agent that has no skill loaded; discover routes; launch a browser in a write-mode; keep several authenticated roles alive at once |
280
280
  | **Explore** | `scout_crawl` `scout_coverage` | Sweep every route in one call; ask what's still untested |
281
281
  | **Look** | `scout_snapshot` `scout_hover` `scout_screenshot` | Read the structured scene (diffed); reveal tooltips/hover cards; capture pixels only when needed |
282
+ | **Ask the server** | `scout_request` | Call the app's own API as this session, with the UI bypassed — the check that turns a hidden button into a proven refusal |
282
283
  | **Act** | `scout_click` `scout_type` `scout_select` `scout_upload` `scout_press` `scout_scroll` `scout_navigate` `scout_back` `scout_run_plan` | Drive the UI like a user; `scout_run_plan` batches a whole mechanical sequence into one call |
283
284
  | **Assess** | `scout_design_audit` `scout_journey` | Score a page's craft/a11y/consistency; measure how hard a task is to complete |
284
285
  | **Record** | `scout_note` `scout_finding` `scout_resolve` `scout_report` | Curate durable notes; file deduped findings; mark fixes; write the report, and on a recorded run the whole run as one page |
@@ -557,7 +558,7 @@ src/
557
558
  report.ts the gap ledger + report generation
558
559
  replay.ts the run as one page: steps, tasks, frames under each finding
559
560
  … collector · dispatch · fixtures · authloss · reaper
560
- scripts/ the 14 test suites (smoke/ holds the real-browser ones)
561
+ scripts/ the 17 test suites (smoke/ holds the real-browser ones)
561
562
  test-app/ fixtures for the real-browser smoke tests
562
563
  skills/scenescout/ the testing method (SKILL.md): a skill in Claude Code, served by the server everywhere else
563
564
  docs/adr/ why it's built this way
package/dist/cli.js CHANGED
@@ -17,7 +17,7 @@ import { APPROX_DISK_MB, BROWSER_ENGINES, browserPresence, defaultAttachNote, de
17
17
  import { CLIENT_LABELS, firstMessageHint, manualFor, parseClients, registerWithClient, vscodeBinary } from "./clients.js";
18
18
  import { CLI_NAME, diagnose, ensureCommand, findOnUserPath, installSkill, isEphemeralRoot, launchCommand, manualRegisterCommand, planCommand, registerMcp, resolveClaudeDir, spawnRunner, } from "./installer.js";
19
19
  import { LEGACY_MEMORY_DIRNAME, MEMORY_DIRNAME } from "./engine/memory.js";
20
- import { formatStatus, localClock, LIVE_TOKEN_FILE, watchTarget } from "./engine/live.js";
20
+ import { formatStatus, liveEngines, liveTokenFileName, localClock, LIVE_TOKEN_FILE, pidAlive, watchTarget, wholeSessions, } from "./engine/live.js";
21
21
  import { formatScan, scanProject } from "./scan.js";
22
22
  const here = path.dirname(fileURLToPath(import.meta.url));
23
23
  const packageRoot = path.resolve(here, "..");
@@ -70,20 +70,6 @@ function readStatusFile(dir) {
70
70
  return "unreadable";
71
71
  }
72
72
  }
73
- function pidAlive(pid) {
74
- if (!pid)
75
- return false;
76
- try {
77
- process.kill(pid, 0);
78
- return true;
79
- }
80
- catch (err) {
81
- // EPERM means the process EXISTS but belongs to another user — only
82
- // ESRCH actually means "no such process". Treating both as dead reported
83
- // a live engine as stale.
84
- return err?.code === "EPERM";
85
- }
86
- }
87
73
  /** Realtime observability: read the status file + recent action log the running engine maintains. */
88
74
  function status(projectPath) {
89
75
  const dir = statusDir(projectPath);
@@ -136,15 +122,45 @@ function status(projectPath) {
136
122
  /** Open the engine's live view. The engine serves it; this only finds the address and hands it to a browser. */
137
123
  function watch(projectPath, open) {
138
124
  const dir = statusDir(projectPath);
139
- const st = readStatusFile(dir);
140
- let token = null;
141
- try {
142
- token = fs.readFileSync(path.join(dir, LIVE_TOKEN_FILE), "utf8");
143
- }
144
- catch {
145
- // watchTarget explains a missing token in context.
125
+ // Several engines can be attached to one project at once — one per client,
126
+ // say. Each writes its own status and token, so every live board is
127
+ // reachable instead of only whichever attached last.
128
+ const engines = liveEngines(dir, pidAlive);
129
+ const tokenFor = (pid) => {
130
+ for (const name of [liveTokenFileName(pid), LIVE_TOKEN_FILE]) {
131
+ try {
132
+ return fs.readFileSync(path.join(dir, name), "utf8");
133
+ }
134
+ catch {
135
+ // Try the shared name next; watchTarget explains a missing token.
136
+ }
137
+ }
138
+ return null;
139
+ };
140
+ if (engines.length > 1) {
141
+ console.log(`${engines.length} engines are attached to this project:`);
142
+ let shown = 0;
143
+ for (const { pid, status } of engines) {
144
+ const one = watchTarget({ status, alive: true, token: tokenFor(pid) });
145
+ const sessions = wholeSessions(status.detail);
146
+ const who = sessions.length > 0 ? sessions.map((x) => x.session).join(", ") : (status.session ?? "no session");
147
+ console.log(`\n pid ${pid} — ${who}`);
148
+ console.log("problem" in one ? ` ${one.problem}` : ` ${one.url}`);
149
+ if (!("problem" in one))
150
+ shown += 1;
151
+ }
152
+ console.log("\nEach address holds its own access token: treat them like passwords.");
153
+ if (shown === 0)
154
+ process.exitCode = 1;
155
+ return;
146
156
  }
147
- const target = watchTarget({ status: st, alive: st !== null && st !== "unreadable" && pidAlive(st.pid), token });
157
+ const only = engines[0];
158
+ const st = only ? only.status : readStatusFile(dir);
159
+ const target = watchTarget({
160
+ status: st,
161
+ alive: only ? true : st !== null && st !== "unreadable" && pidAlive(st.pid),
162
+ token: tokenFor(only?.pid ?? (typeof st === "object" && st !== null ? (st.pid ?? 0) : 0)),
163
+ });
148
164
  if ("problem" in target) {
149
165
  console.log(target.problem);
150
166
  process.exitCode = 1;
@@ -0,0 +1,134 @@
1
+ /**
2
+ * Splitting an app between parallel lanes.
3
+ *
4
+ * A parallel run has one planning agent and several lanes, each driving its
5
+ * own browser. `lane.ts` is how a lane hands its answers BACK. This is the
6
+ * other half: what the planner hands each lane in the first place.
7
+ *
8
+ * Done by hand it goes wrong in two ways, both seen in a real four-session
9
+ * run. Lanes overlap, so two browsers audit the same register and the third
10
+ * module is never opened at all — the run's coverage looks fine because every
11
+ * route was visited, and the gap ledger only catches it at report time. And
12
+ * lanes launch underspecified: the first two sessions in that run acted with
13
+ * no task set, so the person watching the live view saw two browsers clicking
14
+ * through their app with nothing to say why.
15
+ *
16
+ * So the split is computed, not improvised. Routes are grouped into modules by
17
+ * their first path segment, because a lane that owns "everything under
18
+ * /orders" can carry state between its own steps, while a lane handed nine
19
+ * unrelated routes re-learns the app nine times. Modules are then dealt to
20
+ * lanes worst-first by size, which keeps the lanes within a route or two of
21
+ * each other without ever splitting a module across two browsers.
22
+ *
23
+ * Pure, so the split and the wording are table-tested.
24
+ */
25
+ import { normalizePath } from "./fingerprint.js";
26
+ /** Most lanes worth running at once. Past this the planner spends longer reading reports than the lanes spend testing. */
27
+ export const MAX_LANES = 8;
28
+ /** A lane name is a session name: short, and it shows in the live view. */
29
+ export const LANE_NAME_MAX = 40;
30
+ /**
31
+ * The module a route belongs to: its first path segment, or "/" for the root.
32
+ *
33
+ * The query string is dropped, unlike everywhere else in the engine, where a
34
+ * `?tab=` screen is deliberately its own state. Two tabs of one register are
35
+ * one module: handing them to different lanes would mean two browsers learning
36
+ * the same screen, which is the opposite of what the split is for.
37
+ */
38
+ export function moduleOf(route) {
39
+ const path = normalizePath(route).split("?")[0];
40
+ const first = path.split("/").filter(Boolean)[0];
41
+ return first ? `/${first}` : "/";
42
+ }
43
+ /**
44
+ * Deal modules to lanes so each lane owns whole modules and the lanes come out
45
+ * close to even.
46
+ *
47
+ * Largest module first into the lane with fewest routes so far: the standard
48
+ * greedy split, which is within a fraction of optimal for this shape and, more
49
+ * to the point, is stable — the same routes produce the same split every time,
50
+ * so a re-run of a lane can be given the same brief.
51
+ */
52
+ export function splitRoutes(routes, laneCount) {
53
+ const lanes = Math.max(1, Math.min(Math.floor(laneCount) || 1, MAX_LANES));
54
+ const byModule = new Map();
55
+ for (const route of routes) {
56
+ const key = moduleOf(route);
57
+ const list = byModule.get(key);
58
+ if (list)
59
+ list.push(route);
60
+ else
61
+ byModule.set(key, [route]);
62
+ }
63
+ // Biggest first, then by name, and each module's own routes sorted: nothing
64
+ // about the split may depend on the order the routes were discovered in, or
65
+ // a lane that has to be re-run cannot be handed the same brief.
66
+ const modules = [...byModule.entries()]
67
+ .map(([name, list]) => [name, [...list].sort((a, b) => a.localeCompare(b))])
68
+ .sort((a, b) => b[1].length - a[1].length || a[0].localeCompare(b[0]));
69
+ const out = Array.from({ length: lanes }, () => ({ modules: [], routes: [] }));
70
+ for (const [name, list] of modules) {
71
+ let smallest = 0;
72
+ for (let i = 1; i < out.length; i += 1)
73
+ if (out[i].routes.length < out[smallest].routes.length)
74
+ smallest = i;
75
+ out[smallest].modules.push(name);
76
+ out[smallest].routes.push(...list);
77
+ }
78
+ // A lane with nothing to do is a browser held open for no reason.
79
+ return out.filter((lane) => lane.routes.length > 0);
80
+ }
81
+ /** A lane's name, derived from what it owns so the live view reads as the app rather than as "lane-3". */
82
+ export function laneName(modules, index) {
83
+ const first = modules[0]?.replace(/^\//, "") ?? "";
84
+ const base = first ? first.replace(/[^a-zA-Z0-9-]+/g, "-").replace(/^-|-$/g, "") : `lane-${index + 1}`;
85
+ const name = modules.length > 1 ? `${base}+${modules.length - 1}` : base || `lane-${index + 1}`;
86
+ return name.slice(0, LANE_NAME_MAX);
87
+ }
88
+ /** The lanes to run, each with the objective to attach with. */
89
+ export function planLanes(routes, laneCount, opts = {}) {
90
+ return splitRoutes(routes, laneCount).map((lane, i) => ({
91
+ lane: laneName(lane.modules, i),
92
+ objective: laneObjective(lane.modules, opts.goal),
93
+ modules: lane.modules,
94
+ routes: lane.routes,
95
+ }));
96
+ }
97
+ function laneObjective(modules, goal) {
98
+ const owned = modules.length === 1 ? modules[0] : `${modules.slice(0, -1).join(", ")} and ${modules[modules.length - 1]}`;
99
+ const what = `Own ${owned}`;
100
+ return goal ? `${what} — ${goal}` : what;
101
+ }
102
+ /**
103
+ * The briefing the planner gives the lanes.
104
+ *
105
+ * It states the two rules a hand-written brief keeps dropping: a lane touches
106
+ * only its own routes, and a lane says what it is doing before it does it.
107
+ * Both are cheap to write here and expensive to discover missing at report
108
+ * time.
109
+ */
110
+ export function formatBriefs(briefs, opts = {}) {
111
+ if (briefs.length === 0)
112
+ return "Nothing to split — no routes are known yet. Crawl first, then ask again.";
113
+ const mode = opts.mode ?? "read-only";
114
+ const lines = [
115
+ `LANE PLAN — ${briefs.length} lane(s) over ${briefs.reduce((n, b) => n + b.routes.length, 0)} route(s).`,
116
+ ``,
117
+ `Give each lane its own agent. Every lane attaches with its own session name, so the browsers run genuinely in parallel:`,
118
+ ` scout_attach { session: "<lane>", url, projectPath, mode: "${mode}"${opts.role ? `, storageStatePath: "<${opts.role}>"` : ""}, objective: "<objective>" }`,
119
+ ``,
120
+ `Rules to pass on, both of which a hand-written brief tends to drop:`,
121
+ ` · A lane works ITS routes only. Two lanes auditing the same register while a third module is never opened is the failure this plan exists to prevent — and route coverage will look complete either way.`,
122
+ ` · Every acting tool takes a \`task\`. A lane that acts with none is refused, and the person watching the live view would otherwise see a browser clicking through their app with nothing to say why.`,
123
+ ` · A lane reports back with scout_lane_report, not prose. Call scout_lane_report with no reply to get the instruction to put in its prompt.`,
124
+ ``,
125
+ ];
126
+ for (const b of briefs) {
127
+ lines.push(`── ${b.lane} ──`);
128
+ lines.push(`objective: ${b.objective}`);
129
+ lines.push(`owns: ${b.modules.join(", ")} (${b.routes.length} route(s))`);
130
+ lines.push(`routes: ${b.routes.slice(0, 20).join(", ")}${b.routes.length > 20 ? ` … and ${b.routes.length - 20} more` : ""}`);
131
+ lines.push(``);
132
+ }
133
+ return lines.join("\n");
134
+ }