scenescout 3.3.0 → 3.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,17 @@
1
1
  # scenescout
2
2
 
3
+ ## 3.4.0
4
+
5
+ ### Minor Changes
6
+
7
+ - 04f0f22: Measure whether a lane's confidence means anything, and make one lane rubric serve a whole wave.
8
+
9
+ Every lane has been told its confidence must be calibrated — "0.5 means a coin flip, 0.95 means you would bet on it" — and the number was then averaged into one line and discarded. Nothing was stored, so nothing could ever be checked, and a confidence nobody checks is decoration.
10
+
11
+ Lane decisions are now kept, and the report carries a calibration section: how often a decision at a stated confidence matched a finding the project holds, bucketed, with an expected calibration error. It says plainly what the number is not — agreement between the lanes and the bar the project applies, over its whole history, not evidence that the app is broken — and names the two ways a lane is counted wrong through no fault of its own. A decision naming no failing endpoint cannot be looked up at all, so it is excluded and disclosed rather than scored as a miss. Below eight checkable decisions the figure is withheld, but the section says so rather than vanishing. Where findings have since been re-tested through `scout_verify`, those verdicts are reported beside it, because they *are* evidence about the app.
12
+
13
+ Separately, the lane name used to sit in the second sentence of the instruction every lane receives, so two lanes' prompts diverged almost immediately and shared no prefix. The rubric is now identical for every lane in a wave and the name is the last thing said, which makes it one cacheable prefix instead of one per lane.
14
+
3
15
  ## 3.3.0
4
16
 
5
17
  ### Minor Changes
@@ -0,0 +1,217 @@
1
+ /**
2
+ * Whether a lane's confidence means anything.
3
+ *
4
+ * `lane.ts` tells every lane its confidence must be calibrated — "0.5 means a
5
+ * coin flip, 0.95 means you would bet on it" — and then the number was
6
+ * averaged into one line and thrown away. Nothing was stored, so nothing could
7
+ * ever be checked, and a confidence nobody checks is decoration. Decision-only
8
+ * models report an expected calibration error for exactly this reason; asking
9
+ * for calibration without measuring it is the half of the idea that costs
10
+ * nothing and buys nothing.
11
+ *
12
+ * WHAT THIS MEASURES, precisely, because the number is easy to over-read: a
13
+ * lane says "defect" about an observation and attaches a machine signature.
14
+ * The run either went on to file a finding carrying that signature or it did
15
+ * not. That is agreement between the lane and the bar the run actually
16
+ * applied — NOT ground truth about the app. A lane can be perfectly calibrated
17
+ * against a planner that files the wrong things. Where a finding has since
18
+ * been re-tested through `scout_verify`, the verdict is reported beside it,
19
+ * and that IS evidence about the app.
20
+ *
21
+ * Pure, so every rule here is table-tested.
22
+ */
23
+ import { failingSignatures } from "./memory.js";
24
+ /**
25
+ * Upper edge of each confidence bucket. Five is enough to see a shape and few
26
+ * enough that each holds a usable count on a run of a few dozen decisions;
27
+ * twenty buckets of one decision each measure nothing.
28
+ */
29
+ export const BUCKET_EDGES = [0.2, 0.4, 0.6, 0.8, 1.0];
30
+ /**
31
+ * Which bucket a confidence falls in. The top bucket is closed so 1.0 has
32
+ * somewhere to go.
33
+ *
34
+ * A value that is not a number lands in the LOWEST bucket, not the highest.
35
+ * The schema refuses those, but stored history is read back without one, and
36
+ * falling through the comparisons put a NaN in the 0.8–1.0 bucket — a
37
+ * confidence nobody stated, reported as near-certainty.
38
+ */
39
+ export function bucketOf(confidence) {
40
+ if (!Number.isFinite(confidence))
41
+ return 0;
42
+ const c = Math.min(1, Math.max(0, confidence));
43
+ for (let i = 0; i < BUCKET_EDGES.length; i += 1) {
44
+ if (c <= BUCKET_EDGES[i])
45
+ return i;
46
+ }
47
+ return BUCKET_EDGES.length - 1;
48
+ }
49
+ export function bucketLabel(i) {
50
+ const low = i === 0 ? 0 : BUCKET_EDGES[i - 1];
51
+ return `${low.toFixed(1)}–${BUCKET_EDGES[i].toFixed(1)}`;
52
+ }
53
+ /**
54
+ * The keys a piece of evidence can be joined on, or nothing when it cannot be
55
+ * joined at all.
56
+ *
57
+ * Findings MERGE: the store treats `GET /api/things/3 500` and
58
+ * `GET /api/things/7 500` as one bug, because the path normalises to
59
+ * `/api/things/:id`. Joining on the literal string reported five lane
60
+ * decisions about one endpoint as one filed and four dropped — the lanes
61
+ * understated fivefold, and the number read as a finding about them rather
62
+ * than a bug in this join. So the join uses the store's own rule.
63
+ *
64
+ * Evidence carrying no failure signature returns NOTHING rather than falling
65
+ * back to its literal text. That fallback looked harmless and was the single
66
+ * worst thing here: `500 on GET /api/r0` — ordinary English, and whichever
67
+ * agent wrote the finding chose the word order — produced a key that matched
68
+ * nothing, and a lane that was right about a bug that WAS filed published an
69
+ * expected calibration error of 0.90. An unjoinable decision is now counted
70
+ * and disclosed as unjoinable, not scored as a lane being wrong.
71
+ */
72
+ export function joinKeys(evidence) {
73
+ return failingSignatures(evidence);
74
+ }
75
+ /** A confidence as stored may be anything; the schema guards the wire, not the file. */
76
+ function usableConfidence(value) {
77
+ return typeof value === "number" && Number.isFinite(value) && value >= 0 && value <= 1 ? value : null;
78
+ }
79
+ /**
80
+ * Join lane decisions to what the run filed.
81
+ *
82
+ * Only a "defect" carrying a signature has a checkable outcome. An "unsure" is
83
+ * a request to look closer rather than a claim, and a decision with no
84
+ * signature cannot be joined to anything — counting either would measure the
85
+ * lane's willingness to attach evidence, not its judgement.
86
+ */
87
+ export function calibrate(decisions, findings) {
88
+ const filedKeys = new Map();
89
+ for (const f of findings) {
90
+ if (!f.evidence)
91
+ continue;
92
+ // A finding verified as still present is the most informative match, so it
93
+ // wins a key two findings share; otherwise first write wins and the result
94
+ // does not depend on the order findings came back in.
95
+ for (const key of joinKeys(f.evidence)) {
96
+ const prev = filedKeys.get(key);
97
+ if (!prev || (f.verdict && f.verifiedAt && !(prev.verdict && prev.verifiedAt)))
98
+ filedKeys.set(key, f);
99
+ }
100
+ }
101
+ const claims = decisions.filter((d) => d.verdict === "defect" && d.evidence !== null);
102
+ const checkable = claims.filter((d) => usableConfidence(d.confidence) !== null && joinKeys(d.evidence).size > 0);
103
+ const unjoinable = claims.length - checkable.length;
104
+ if (checkable.length === 0)
105
+ return unjoinable > 0 ? { checkable: 0, filed: 0, stated: 0, buckets: [], ece: 0, verified: { present: 0, gone: 0, changed: 0 }, unjoinable } : null;
106
+ const buckets = BUCKET_EDGES.map(() => ({ n: 0, conf: 0, hits: 0 }));
107
+ const verified = { present: 0, gone: 0, changed: 0 };
108
+ const matched = new Map();
109
+ let filed = 0;
110
+ let stated = 0;
111
+ for (const d of checkable) {
112
+ let match;
113
+ for (const key of joinKeys(d.evidence)) {
114
+ match = filedKeys.get(key);
115
+ if (match)
116
+ break;
117
+ }
118
+ const hit = match !== undefined;
119
+ const confidence = usableConfidence(d.confidence);
120
+ const b = buckets[bucketOf(confidence)];
121
+ b.n += 1;
122
+ b.conf += confidence;
123
+ if (hit)
124
+ b.hits += 1;
125
+ if (hit)
126
+ filed += 1;
127
+ stated += confidence;
128
+ // By finding, not by decision: several decisions can match one finding,
129
+ // and this sentence counts findings.
130
+ if (match)
131
+ matched.set(match.id, match);
132
+ }
133
+ for (const f of matched.values()) {
134
+ // A verdict read from disk is not guaranteed to be one of the three.
135
+ if (f.verifiedAt && f.verdict && f.verdict in verified)
136
+ verified[f.verdict] += 1;
137
+ }
138
+ const n = checkable.length;
139
+ let ece = 0;
140
+ const out = [];
141
+ buckets.forEach((b, i) => {
142
+ if (b.n === 0)
143
+ return;
144
+ const meanConf = b.conf / b.n;
145
+ const rate = b.hits / b.n;
146
+ ece += (b.n / n) * Math.abs(meanConf - rate);
147
+ out.push({ label: bucketLabel(i), decisions: b.n, stated: meanConf, filed: rate });
148
+ });
149
+ return { checkable: n, filed, stated: stated / n, buckets: out, ece, verified, unjoinable };
150
+ }
151
+ const pct = (x) => `${Math.round(x * 100)}%`;
152
+ /**
153
+ * The report section. It leads with what the number is not, because "ECE 0.24"
154
+ * above a table of percentages reads as a verdict on the app rather than on
155
+ * the lanes that judged it.
156
+ */
157
+ export function formatCalibration(c) {
158
+ if (!c)
159
+ return [];
160
+ if (c.checkable < MIN_FOR_A_VERDICT) {
161
+ // Suppress the NUMBER, not the fact. A run that recorded seven decisions
162
+ // and a run that recorded none look identical when the section simply
163
+ // vanishes, and the reader concludes the feature is broken — which is this
164
+ // project's own definition of a silent path.
165
+ const had = c.checkable + c.unjoinable;
166
+ if (had === 0)
167
+ return [];
168
+ return [
169
+ `## How well the lanes judged`,
170
+ ``,
171
+ `Not enough to say yet: ${c.checkable} lane decision(s) could be checked${c.unjoinable > 0 ? ` (and ${c.unjoinable} could not be looked up at all)` : ""}, ` +
172
+ `and ${MIN_FOR_A_VERDICT} are needed before a calibration figure survives one of them changing.`,
173
+ ``,
174
+ ];
175
+ }
176
+ const lines = [
177
+ `## How well the lanes judged`,
178
+ ``,
179
+ `${c.checkable} lane decision(s) called a defect on an endpoint that failed; ${c.filed} of them match a finding this project holds (${pct(c.filed / c.checkable)}), ` +
180
+ `against a mean stated confidence of ${c.stated.toFixed(2)}.`,
181
+ ``,
182
+ `**What this is not.** It covers the project's whole history, not this run, and it measures agreement between the lanes and the bar this project applies — not whether the app is broken. ` +
183
+ `A lane can be perfectly calibrated against a planner that files the wrong things. Two things also count against a lane through no fault of its own: a defect filed with no machine signature, ` +
184
+ `and a finding whose signature was dropped when the store merged it into another. Read the figure as an ordering, not a measurement.`,
185
+ ``,
186
+ `| Stated confidence | Decisions | Said | Filed |`,
187
+ `|---|---:|---:|---:|`,
188
+ ];
189
+ for (const b of c.buckets) {
190
+ lines.push(`| ${b.label} | ${b.decisions} | ${b.stated.toFixed(2)} | ${pct(b.filed)} |`);
191
+ }
192
+ lines.push(``, `Expected calibration error **${c.ece.toFixed(2)}** — the gap between what the lanes said and what happened, weighted by bucket. ${readEce(c.ece)}`);
193
+ if (c.unjoinable > 0) {
194
+ lines.push(``, `${c.unjoinable} further decision(s) called a defect but named no failing endpoint, so nothing could be looked up for them. They are excluded above rather than counted as wrong.`);
195
+ }
196
+ const seen = c.verified.present + c.verified.gone + c.verified.changed;
197
+ if (seen > 0) {
198
+ lines.push(``, `${seen} of the findings those decisions matched ${seen === 1 ? "has" : "have"} since been re-tested with \`scout_verify\`: ${c.verified.present} still present, ${c.verified.gone} gone, ${c.verified.changed} changed. ` +
199
+ `Those verdicts are evidence about the app, unlike the table above. Findings no lane decision matched are not counted here, so this will not agree with the re-tested column elsewhere in the report.`);
200
+ }
201
+ lines.push(``);
202
+ return lines;
203
+ }
204
+ /** Below this, the number swings on a single decision and is worse than no number. */
205
+ export const MIN_FOR_A_VERDICT = 8;
206
+ /**
207
+ * What the figure suggests, hedged on purpose. The join loses decisions it
208
+ * cannot key and findings whose signature a merge discarded, and both push the
209
+ * error upward — so a confident "these numbers are wrong" can itself be wrong.
210
+ */
211
+ function readEce(ece) {
212
+ if (ece <= 0.05)
213
+ return "The lanes' numbers track what happened closely.";
214
+ if (ece <= 0.15)
215
+ return "Close enough to compare lanes by.";
216
+ return "Far enough apart that a lane's confidence is worth reading as an ordering rather than a rate — after checking the exclusions above, which push this upward.";
217
+ }
@@ -138,14 +138,28 @@ export function parseLaneReport(text, expectedLane) {
138
138
  */
139
139
  export function laneReportInstruction(lane) {
140
140
  return [
141
- "Reply with ONE JSON object and nothing else — no prose before or after it, no explanation, no headings. A fenced ```json block is fine.",
142
- `Shape: {"lane":${JSON.stringify(lane)},"status":<${quoteAll(LANE_STATUSES)}>,"decisions":[…],"routes":[…],"blocked_by":<string or null>}.`,
143
- `Each decision: {"observation":<a short id for what was observed, unique in the report, at most ${LANE_OBSERVATION_MAX} characters>,"verdict":<${quoteAll(LANE_VERDICTS)}>,"severity":<${quoteAll(LANE_SEVERITIES)} or null>,"category":<${quoteAll(LANE_CATEGORIES)} or null>,"confidence":<0..1>,"evidence":<machine signature such as "GET /api/things 500", or null>}.`,
144
- `A "defect" must carry a severity and a category. "evidence" is a signature, not a sentence: at most ${LANE_EVIDENCE_MAX} characters. "confidence" is how sure you are of the verdict, calibrated: 0.5 means a coin flip, 0.95 means you would bet on it.`,
145
- `"routes" lists the routes you covered, each at most ${LANE_ROUTE_MAX} characters. "blocked_by" is one line of at most ${LANE_BLOCKED_BY_MAX} characters saying what stopped you: required when the status is "blocked", allowed with "partial", null with "complete"; the detail belongs in a finding. The lane name is at most ${LANE_NAME_MAX} characters. At most ${LANE_MAX_ITEMS} decisions and ${LANE_MAX_ITEMS} routes. Unknown keys are refused.`,
146
- "The object IS your final report: whatever hands it back must hand back the object verbatim, not a summary of it.",
141
+ ...LANE_RUBRIC,
142
+ // The ONLY lane-specific sentence, and it comes last on purpose. Every
143
+ // lane in a wave is given the same rubric, so keeping it byte-identical up
144
+ // to here makes it one shared prompt prefix: the cache hits from the
145
+ // second lane onward instead of diverging at the first sentence, which is
146
+ // what putting the name in the shape line used to do.
147
+ `Your lane name is ${JSON.stringify(lane)}; put exactly that in "lane".`,
147
148
  ].join(" ");
148
149
  }
150
+ /**
151
+ * Everything every lane is told, identical for all of them. Built once from
152
+ * the same constants the parser enforces, because a limit a lane is not told
153
+ * refuses good replies.
154
+ */
155
+ const LANE_RUBRIC = [
156
+ "Reply with ONE JSON object and nothing else — no prose before or after it, no explanation, no headings. A fenced ```json block is fine.",
157
+ `Shape: {"lane":<your lane name>,"status":<${quoteAll(LANE_STATUSES)}>,"decisions":[…],"routes":[…],"blocked_by":<string or null>}.`,
158
+ `Each decision: {"observation":<a short id for what was observed, unique in the report, at most ${LANE_OBSERVATION_MAX} characters>,"verdict":<${quoteAll(LANE_VERDICTS)}>,"severity":<${quoteAll(LANE_SEVERITIES)} or null>,"category":<${quoteAll(LANE_CATEGORIES)} or null>,"confidence":<0..1>,"evidence":<machine signature such as "GET /api/things 500", or null>}.`,
159
+ `A "defect" must carry a severity and a category. "evidence" is a signature, not a sentence: at most ${LANE_EVIDENCE_MAX} characters. "confidence" is how sure you are of the verdict, calibrated: 0.5 means a coin flip, 0.95 means you would bet on it.`,
160
+ `"routes" lists the routes you covered, each at most ${LANE_ROUTE_MAX} characters. "blocked_by" is one line of at most ${LANE_BLOCKED_BY_MAX} characters saying what stopped you: required when the status is "blocked", allowed with "partial", null with "complete"; the detail belongs in a finding. The lane name is at most ${LANE_NAME_MAX} characters. At most ${LANE_MAX_ITEMS} decisions and ${LANE_MAX_ITEMS} routes. Unknown keys are refused.`,
161
+ "The object IS your final report: whatever hands it back must hand back the object verbatim, not a summary of it.",
162
+ ];
149
163
  function quoteAll(values) {
150
164
  return values.map((v) => `"${v}"`).join("|");
151
165
  }
@@ -250,8 +250,39 @@ export function mergeMemory(mine, theirs) {
250
250
  }
251
251
  if (Object.keys(out.roleAccess).length === 0)
252
252
  delete out.roleAccess;
253
+ // Lanes commonly run in SEPARATE processes — that is the point of a lane —
254
+ // so without this the spread at the top would keep one process's decisions
255
+ // and drop every other lane's, which is the exact bug this merge exists to
256
+ // prevent for findings. Keyed so merging the same foreign document twice is
257
+ // idempotent.
258
+ const byDecision = new Map();
259
+ for (const d of theirs.laneDecisions ?? [])
260
+ byDecision.set(decisionKey(d), d);
261
+ for (const d of mine.laneDecisions ?? [])
262
+ byDecision.set(decisionKey(d), d);
263
+ out.laneDecisions = [...byDecision.values()].sort((a, b) => a.at.localeCompare(b.at)).slice(-MAX_LANE_DECISIONS);
264
+ if (out.laneDecisions.length === 0)
265
+ delete out.laneDecisions;
253
266
  return out;
254
267
  }
268
+ /**
269
+ * What makes two stored decisions the same judgement.
270
+ *
271
+ * Deliberately NOT the timestamp. `at` is stamped when the planner folds the
272
+ * reply, not when the lane judged, so relaying one reply twice — which the
273
+ * protocol invites, since a refused report is asked for again — wrote every
274
+ * decision a second time and doubled the lane's weight in the calibration.
275
+ * Identity is what was decided, so re-folding the same reply is a no-op.
276
+ */
277
+ function decisionKey(d) {
278
+ return [d.lane, d.observation, d.verdict, d.severity ?? "", d.category ?? "", d.confidence, d.evidence ?? ""].join("|");
279
+ }
280
+ /**
281
+ * Most lane decisions kept. Calibration wants a few dozen; a long-lived
282
+ * project would otherwise accumulate every decision ever made and re-serialise
283
+ * them on each save, which is what made an old history slow to open.
284
+ */
285
+ export const MAX_LANE_DECISIONS = 1000;
255
286
  const MAX_DISCOVERED_ROUTES = 300;
256
287
  /** Shared finding-similarity helpers (used by live dedup and retro-merge). */
257
288
  function findingTokens(s) {
@@ -295,7 +326,22 @@ function findingLiterals(...texts) {
295
326
  * (`/api/users/7` and `/api/users/9` are one endpoint) so a per-record repro
296
327
  * does not read as a per-record bug.
297
328
  */
298
- function endpointSignatures(evidence) {
329
+ /**
330
+ * The signatures that identify a BUG: only those carrying a failure status.
331
+ *
332
+ * A bare `POST /api/orders` says which endpoint was involved, not what went
333
+ * wrong — a double submit and an accepted negative quantity both name it, and
334
+ * are two bugs. Exported so calibration applies the store's own rule instead
335
+ * of a second copy of this regex, which is what it had.
336
+ */
337
+ export function failingSignatures(evidence) {
338
+ return new Set([...endpointSignatures(evidence)].filter((sig) => /\s[45]\d{2}$/.test(sig)));
339
+ }
340
+ /**
341
+ * The endpoint signatures a piece of evidence names: method + normalised path,
342
+ * with the failure status when one follows.
343
+ */
344
+ export function endpointSignatures(evidence) {
299
345
  const out = new Set();
300
346
  if (!evidence)
301
347
  return out;
@@ -338,11 +384,10 @@ function sharesEndpointSignature(a, b) {
338
384
  // `POST /api/orders` (the endpoint answered 2xx, or no status was named)
339
385
  // says which endpoint was involved, not what went wrong: a double submit
340
386
  // and an accepted negative quantity both name it, and are two bugs.
341
- const failing = (evidence) => new Set([...endpointSignatures(evidence)].filter((sig) => /\s[45]\d{2}$/.test(sig)));
342
- const aSigs = failing(a.evidence);
387
+ const aSigs = failingSignatures(a.evidence);
343
388
  if (aSigs.size === 0)
344
389
  return false;
345
- for (const sig of failing(b.evidence))
390
+ for (const sig of failingSignatures(b.evidence))
346
391
  if (aSigs.has(sig))
347
392
  return true;
348
393
  return false;
@@ -782,6 +827,45 @@ export class MemoryStore {
782
827
  fs.writeFileSync(this.assumptionsPath, content);
783
828
  return true;
784
829
  }
830
+ /** What the lanes decided, oldest first. Empty on a project that has never run one. */
831
+ get laneDecisions() {
832
+ return this.data.laneDecisions ?? [];
833
+ }
834
+ /**
835
+ * Record what a lane decided. Called once per accepted lane report, so the
836
+ * confidence it stated can be checked later against what the run filed.
837
+ * Free text from the lane is redacted like every other stored string: an
838
+ * observation is written by a model reading the app under test.
839
+ */
840
+ addLaneDecisions(lane, decisions) {
841
+ if (decisions.length === 0)
842
+ return 0;
843
+ const list = this.data.laneDecisions ?? [];
844
+ const seen = new Set(list.map(decisionKey));
845
+ let added = 0;
846
+ for (const d of decisions) {
847
+ const record = {
848
+ ...d,
849
+ lane,
850
+ observation: redactSecrets(d.observation).slice(0, 200),
851
+ evidence: d.evidence === null ? null : redactSecrets(d.evidence).slice(0, 200),
852
+ };
853
+ if (seen.has(decisionKey(record)))
854
+ continue;
855
+ seen.add(decisionKey(record));
856
+ list.push(record);
857
+ added += 1;
858
+ }
859
+ this.data.laneDecisions = list.slice(-MAX_LANE_DECISIONS);
860
+ if (added > 0)
861
+ this.flush();
862
+ // What survives the cap. Appends go to the tail and the cap keeps the
863
+ // tail, so all of `added` survives unless the call itself exceeds the cap.
864
+ // Measuring it as growth instead looked right and was not: `list` aliases
865
+ // the stored array, so the "before" length was read after the appends and
866
+ // every call after the first reported nothing kept — while storing fine.
867
+ return Math.min(added, MAX_LANE_DECISIONS);
868
+ }
785
869
  /** Mark a finding resolved; returns it or null. */
786
870
  resolveFinding(id) {
787
871
  const f = this.data.findings.find((x) => x.id === id);
@@ -4,6 +4,7 @@ import { SHARED_CHROME_ROUTE } from "./memory.js";
4
4
  import { sayVerification } from "./verify.js";
5
5
  import { feedForSession } from "./live.js";
6
6
  import { buildReplayHtml, evidenceFor } from "./replay.js";
7
+ import { calibrate, formatCalibration } from "./calibration.js";
7
8
  import { formatPace, measurePace } from "./pace.js";
8
9
  function playwrightSkeleton(f) {
9
10
  const routeClass = f.state.split("#")[0].split("?")[0];
@@ -616,6 +617,9 @@ export function generateReport(memory, oracleLog, extras, opts = {}) {
616
617
  if (extras?.pace !== false) {
617
618
  lines.push(...formatPace(measurePace(memory.actionLog, Date.now(), extras?.attachedSessions ?? [])));
618
619
  }
620
+ // Only on a run that used lanes, and only once enough of them have been
621
+ // judged for the number to mean anything; formatCalibration decides both.
622
+ lines.push(...formatCalibration(calibrate(memory.laneDecisions, memory.findings)));
619
623
  const markdown = lines.join("\n");
620
624
  const outPath = path.join(memory.dir, "report.md");
621
625
  const htmlPath = path.join(memory.dir, "report.html");
@@ -528,10 +528,35 @@ server.registerTool("scout_lane_report", {
528
528
  if (reply === undefined)
529
529
  return { content: [{ type: "text", text: laneReportInstruction(lane) }] };
530
530
  const parsed = parseLaneReport(reply, lane);
531
- const out = parsed.ok
532
- ? `Lane report accepted — ${summarizeLaneReport(parsed.report)}`
533
- : `Lane report REFUSED: ${parsed.reason}. Ask the lane once for the corrected object; do not re-judge its prose.`;
534
- return { content: [{ type: "text", text: out }] };
531
+ if (!parsed.ok) {
532
+ return {
533
+ content: [
534
+ { type: "text", text: `Lane report REFUSED: ${parsed.reason}. Ask the lane once for the corrected object; do not re-judge its prose.` },
535
+ ],
536
+ };
537
+ }
538
+ // Keep what the lane decided, so the confidence it stated can be checked
539
+ // against what the run goes on to file. Best-effort: a lane report is
540
+ // still accepted if this project has no memory open yet, because the
541
+ // planner's fold must not depend on where the report was written.
542
+ // ONLY the lane's own session. Falling back to any engine with memory
543
+ // open put one project's decisions into another project's store whenever
544
+ // two sessions were attached to different apps — and the lane having
545
+ // already closed makes that the ordinary case, not an edge one.
546
+ const owner = engines.get(lane)?.memory;
547
+ const at = new Date().toISOString();
548
+ const kept = owner
549
+ ? owner.addLaneDecisions(lane, parsed.report.decisions.map((d) => ({ ...d, lane, at })))
550
+ : 0;
551
+ // Say when nothing was kept. Every lane closing its session before the
552
+ // planner folds its report is the order the method describes, and it
553
+ // leaves no memory to write to — reporting a bare "accepted" while the
554
+ // skill promises the decisions are kept is the kind of silence that
555
+ // makes a later calibration section look wrong rather than absent.
556
+ const note = kept > 0
557
+ ? ` (${kept} decision(s) kept for calibration)`
558
+ : ` (nothing kept for calibration — session ${JSON.stringify(lane)} is not attached here, so there is no project memory to write to. Fold a lane report before closing that lane's session.)`;
559
+ return { content: [{ type: "text", text: `Lane report accepted — ${summarizeLaneReport(parsed.report)}${note}` }] };
535
560
  }
536
561
  catch (err) {
537
562
  return errorText(err);
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "scenescout",
3
- "version": "3.3.0",
3
+ "version": "3.4.0",
4
4
  "description": "SceneScout — exploratory UI testing for AI coding agents. An MCP server that gives any agent (Claude Code, Cursor, VS Code Copilot, Codex, Gemini CLI and others) a structured view of a running web app, always-on oracles, a network-level write policy, memory across runs and a gap-checked report.",
5
5
  "license": "MIT",
6
6
  "author": "brunoboto96",
@@ -71,7 +71,7 @@
71
71
  "mcp-check": "npm run build && npm run mcp-check:run",
72
72
  "mcp-check:run": "tsx scripts/mcp-check.ts",
73
73
  "test": "npm run build && npm run test:unit && npm run smoke:run && npm run mcp-check:run",
74
- "test:unit": "npm run scan-test && npm run oracle-test && npm run policy-test && npm run fixture-test && npm run dispatch-test && npm run design-test && npm run contract-test && npm run pace-test && npm run request-test && npm run settle-test && npm run claims-test && npm run verify-test && npm run brief-test && npm run lane-test && npm run memory-test && npm run install-test && npm run live-test && npm run demo-test && npm run hygiene-test",
74
+ "test:unit": "npm run scan-test && npm run oracle-test && npm run policy-test && npm run fixture-test && npm run dispatch-test && npm run design-test && npm run contract-test && npm run pace-test && npm run request-test && npm run settle-test && npm run claims-test && npm run verify-test && npm run brief-test && npm run lane-test && npm run calibration-test && npm run memory-test && npm run install-test && npm run live-test && npm run demo-test && npm run hygiene-test",
75
75
  "scan-test": "tsx scripts/scan-test.ts",
76
76
  "oracle-test": "tsx --test scripts/oracle-test.ts",
77
77
  "policy-test": "tsx --test scripts/policy-test.ts",
@@ -90,7 +90,8 @@
90
90
  "settle-test": "tsx --test scripts/settle-test.ts",
91
91
  "claims-test": "tsx --test scripts/claims-test.ts",
92
92
  "verify-test": "tsx --test scripts/verify-test.ts",
93
- "brief-test": "tsx --test scripts/brief-test.ts"
93
+ "brief-test": "tsx --test scripts/brief-test.ts",
94
+ "calibration-test": "tsx --test scripts/calibration-test.ts"
94
95
  },
95
96
  "dependencies": {
96
97
  "@modelcontextprotocol/sdk": "^1.12.0",
@@ -74,7 +74,7 @@ Some flows need a TEAM — a document one role submits and another approves, a r
74
74
  - `scout_close {all: true}` at the end of a multi-role run; `scout_close {session}` to drop one role early.
75
75
  - **Several agents in parallel** (subagents or a workflow, each driving its own session): each agent attaches its session when it STARTS and closes it by name when it FINISHES. Never open sessions ahead for agents that have not started, and never hand an open session from one agent to the next: an agent waiting for its turn should hold no browser. Give each session an `objective` when you attach it (`scout_attach {session, objective:"Approve and reject orders as a manager"}`) and wrap each goal in `scout_journey`: the live view shows that objective and the task it is on beside the session's feed, which is how the person watching knows what every agent is for. Keep that goal TRUE: one journey per goal, one goal per thing you are checking ("Save a settings change as the auditor", not "Check every page"), ended the moment it is decided and the next one started before you move on. A journey that outlives its goal shows the viewer an objective the session left behind minutes ago. Run roughly as many agents at once as the machine has cores, less two, since each drives a real browser; beyond that they only queue. Exploring one area is well within a mid-tier model, so run these agents on one (Sonnet or its equivalent in your client) unless the user names a model; keep the larger model for the agent that plans the split and writes the report. No agent may call `scout_close {all: true}` while others run — only the last step, once every agent has finished.
76
76
  - **Let the engine divide the app: `scout_lane_brief {lanes, goal}`.** Call it after the first crawl, when route knowledge is complete. It splits the known routes into whole modules — everything under `/orders` goes to one lane, so that lane carries state between its own steps instead of re-learning the app on every route — deals the modules out so the lanes come out within a route or two of each other, and returns each lane's session name, the `objective` to attach it with, and the routes it owns. Pass `routes` to split a subset instead. Hand each brief to its agent verbatim. Dividing by hand fails in two ways that a finished run cannot tell apart from success: two lanes audit the same register while a third module is never opened, and route coverage reads complete either way; and lanes launch without an objective, so the person watching sees browsers clicking through their app with nothing to say why.
77
- - **Lanes report in typed decisions, not prose.** Each lane files its findings with `scout_finding` as it goes, so the report already has them; what the lane hands BACK to the planner is one JSON object, the lane report, and nothing else. Get the paragraph to put in a lane's prompt from `scout_lane_report {lane}`: it names the shape (`status`, one decision per observation with `verdict`, `severity`, `category`, a calibrated `confidence` and a machine-signature `evidence`, the `routes` covered, `blocked_by`), every value's closed set (the categories are `scout_finding`'s) and every length limit, because a limit a lane is not told refuses good replies. When the lane hands back, pass its text to `scout_lane_report {lane, reply}`: it returns the one-line fold (defects, highs, unsure, mean confidence, routes, what blocked it) or the reason the reply was refused. Relay a refusal to the lane once and ask for the corrected object; never re-judge its prose yourself, and tell every lane that the object IS its final report, since a lane that writes the object and then hands back a summary of it has failed. The planner then folds lanes by counting, not by reading: a lane's reply is a few hundred tokens whatever it found, and a value the schema refuses is caught at the boundary instead of becoming a severity like "Low-Medium" in the report. Where the client lets you set a lane's reasoning effort, use **medium**: in this project's benchmark it judged as well as high in three-quarters of the time and under half of xhigh, low inflated severities, and max took three minutes per lane for no better agreement.
77
+ - **Lanes report in typed decisions, not prose.** Each lane files its findings with `scout_finding` as it goes, so the report already has them; what the lane hands BACK to the planner is one JSON object, the lane report, and nothing else. Get the paragraph to put in a lane's prompt from `scout_lane_report {lane}`: it names the shape (`status`, one decision per observation with `verdict`, `severity`, `category`, a calibrated `confidence` and a machine-signature `evidence`, the `routes` covered, `blocked_by`), every value's closed set (the categories are `scout_finding`'s) and every length limit, because a limit a lane is not told refuses good replies. When the lane hands back, pass its text to `scout_lane_report {lane, reply}`: it returns the one-line fold (defects, highs, unsure, mean confidence, routes, what blocked it) or the reason the reply was refused, and keeps each decision so the confidence the lane stated can be checked against what the run went on to file — the report's calibration section is built from that, and it only appears once enough decisions exist to mean something. **So the confidence is worth stating honestly**: a lane that writes 0.95 on everything makes the section say so. Relay a refusal to the lane once and ask for the corrected object; never re-judge its prose yourself, and tell every lane that the object IS its final report, since a lane that writes the object and then hands back a summary of it has failed. The planner then folds lanes by counting, not by reading: a lane's reply is a few hundred tokens whatever it found, and a value the schema refuses is caught at the boundary instead of becoming a severity like "Low-Medium" in the report. Where the client lets you set a lane's reasoning effort, use **medium**: in this project's benchmark it judged as well as high in three-quarters of the time and under half of xhigh, low inflated severities, and max took three minutes per lane for no better agreement.
78
78
 
79
79
  ## The impatient-user pass (extensive)
80
80