scenescout 3.3.0 → 3.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +12 -0
- package/dist/engine/calibration.js +217 -0
- package/dist/engine/lane.js +20 -6
- package/dist/engine/memory.js +88 -4
- package/dist/engine/report.js +4 -0
- package/dist/mcp-server.js +29 -4
- package/package.json +4 -3
- package/skills/scenescout/SKILL.md +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,17 @@
|
|
|
1
1
|
# scenescout
|
|
2
2
|
|
|
3
|
+
## 3.4.0
|
|
4
|
+
|
|
5
|
+
### Minor Changes
|
|
6
|
+
|
|
7
|
+
- 04f0f22: Measure whether a lane's confidence means anything, and make one lane rubric serve a whole wave.
|
|
8
|
+
|
|
9
|
+
Every lane has been told its confidence must be calibrated — "0.5 means a coin flip, 0.95 means you would bet on it" — and the number was then averaged into one line and discarded. Nothing was stored, so nothing could ever be checked, and a confidence nobody checks is decoration.
|
|
10
|
+
|
|
11
|
+
Lane decisions are now kept, and the report carries a calibration section: how often a decision at a stated confidence matched a finding the project holds, bucketed, with an expected calibration error. It says plainly what the number is not — agreement between the lanes and the bar the project applies, over its whole history, not evidence that the app is broken — and names the two ways a lane is counted wrong through no fault of its own. A decision naming no failing endpoint cannot be looked up at all, so it is excluded and disclosed rather than scored as a miss. Below eight checkable decisions the figure is withheld, but the section says so rather than vanishing. Where findings have since been re-tested through `scout_verify`, those verdicts are reported beside it, because they *are* evidence about the app.
|
|
12
|
+
|
|
13
|
+
Separately, the lane name used to sit in the second sentence of the instruction every lane receives, so two lanes' prompts diverged almost immediately and shared no prefix. The rubric is now identical for every lane in a wave and the name is the last thing said, which makes it one cacheable prefix instead of one per lane.
|
|
14
|
+
|
|
3
15
|
## 3.3.0
|
|
4
16
|
|
|
5
17
|
### Minor Changes
|
|
@@ -0,0 +1,217 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* Whether a lane's confidence means anything.
|
|
3
|
+
*
|
|
4
|
+
* `lane.ts` tells every lane its confidence must be calibrated — "0.5 means a
|
|
5
|
+
* coin flip, 0.95 means you would bet on it" — and then the number was
|
|
6
|
+
* averaged into one line and thrown away. Nothing was stored, so nothing could
|
|
7
|
+
* ever be checked, and a confidence nobody checks is decoration. Decision-only
|
|
8
|
+
* models report an expected calibration error for exactly this reason; asking
|
|
9
|
+
* for calibration without measuring it is the half of the idea that costs
|
|
10
|
+
* nothing and buys nothing.
|
|
11
|
+
*
|
|
12
|
+
* WHAT THIS MEASURES, precisely, because the number is easy to over-read: a
|
|
13
|
+
* lane says "defect" about an observation and attaches a machine signature.
|
|
14
|
+
* The run either went on to file a finding carrying that signature or it did
|
|
15
|
+
* not. That is agreement between the lane and the bar the run actually
|
|
16
|
+
* applied — NOT ground truth about the app. A lane can be perfectly calibrated
|
|
17
|
+
* against a planner that files the wrong things. Where a finding has since
|
|
18
|
+
* been re-tested through `scout_verify`, the verdict is reported beside it,
|
|
19
|
+
* and that IS evidence about the app.
|
|
20
|
+
*
|
|
21
|
+
* Pure, so every rule here is table-tested.
|
|
22
|
+
*/
|
|
23
|
+
import { failingSignatures } from "./memory.js";
|
|
24
|
+
/**
|
|
25
|
+
* Upper edge of each confidence bucket. Five is enough to see a shape and few
|
|
26
|
+
* enough that each holds a usable count on a run of a few dozen decisions;
|
|
27
|
+
* twenty buckets of one decision each measure nothing.
|
|
28
|
+
*/
|
|
29
|
+
export const BUCKET_EDGES = [0.2, 0.4, 0.6, 0.8, 1.0];
|
|
30
|
+
/**
|
|
31
|
+
* Which bucket a confidence falls in. The top bucket is closed so 1.0 has
|
|
32
|
+
* somewhere to go.
|
|
33
|
+
*
|
|
34
|
+
* A value that is not a number lands in the LOWEST bucket, not the highest.
|
|
35
|
+
* The schema refuses those, but stored history is read back without one, and
|
|
36
|
+
* falling through the comparisons put a NaN in the 0.8–1.0 bucket — a
|
|
37
|
+
* confidence nobody stated, reported as near-certainty.
|
|
38
|
+
*/
|
|
39
|
+
export function bucketOf(confidence) {
|
|
40
|
+
if (!Number.isFinite(confidence))
|
|
41
|
+
return 0;
|
|
42
|
+
const c = Math.min(1, Math.max(0, confidence));
|
|
43
|
+
for (let i = 0; i < BUCKET_EDGES.length; i += 1) {
|
|
44
|
+
if (c <= BUCKET_EDGES[i])
|
|
45
|
+
return i;
|
|
46
|
+
}
|
|
47
|
+
return BUCKET_EDGES.length - 1;
|
|
48
|
+
}
|
|
49
|
+
export function bucketLabel(i) {
|
|
50
|
+
const low = i === 0 ? 0 : BUCKET_EDGES[i - 1];
|
|
51
|
+
return `${low.toFixed(1)}–${BUCKET_EDGES[i].toFixed(1)}`;
|
|
52
|
+
}
|
|
53
|
+
/**
|
|
54
|
+
* The keys a piece of evidence can be joined on, or nothing when it cannot be
|
|
55
|
+
* joined at all.
|
|
56
|
+
*
|
|
57
|
+
* Findings MERGE: the store treats `GET /api/things/3 500` and
|
|
58
|
+
* `GET /api/things/7 500` as one bug, because the path normalises to
|
|
59
|
+
* `/api/things/:id`. Joining on the literal string reported five lane
|
|
60
|
+
* decisions about one endpoint as one filed and four dropped — the lanes
|
|
61
|
+
* understated fivefold, and the number read as a finding about them rather
|
|
62
|
+
* than a bug in this join. So the join uses the store's own rule.
|
|
63
|
+
*
|
|
64
|
+
* Evidence carrying no failure signature returns NOTHING rather than falling
|
|
65
|
+
* back to its literal text. That fallback looked harmless and was the single
|
|
66
|
+
* worst thing here: `500 on GET /api/r0` — ordinary English, and whichever
|
|
67
|
+
* agent wrote the finding chose the word order — produced a key that matched
|
|
68
|
+
* nothing, and a lane that was right about a bug that WAS filed published an
|
|
69
|
+
* expected calibration error of 0.90. An unjoinable decision is now counted
|
|
70
|
+
* and disclosed as unjoinable, not scored as a lane being wrong.
|
|
71
|
+
*/
|
|
72
|
+
export function joinKeys(evidence) {
|
|
73
|
+
return failingSignatures(evidence);
|
|
74
|
+
}
|
|
75
|
+
/** A confidence as stored may be anything; the schema guards the wire, not the file. */
|
|
76
|
+
function usableConfidence(value) {
|
|
77
|
+
return typeof value === "number" && Number.isFinite(value) && value >= 0 && value <= 1 ? value : null;
|
|
78
|
+
}
|
|
79
|
+
/**
|
|
80
|
+
* Join lane decisions to what the run filed.
|
|
81
|
+
*
|
|
82
|
+
* Only a "defect" carrying a signature has a checkable outcome. An "unsure" is
|
|
83
|
+
* a request to look closer rather than a claim, and a decision with no
|
|
84
|
+
* signature cannot be joined to anything — counting either would measure the
|
|
85
|
+
* lane's willingness to attach evidence, not its judgement.
|
|
86
|
+
*/
|
|
87
|
+
export function calibrate(decisions, findings) {
|
|
88
|
+
const filedKeys = new Map();
|
|
89
|
+
for (const f of findings) {
|
|
90
|
+
if (!f.evidence)
|
|
91
|
+
continue;
|
|
92
|
+
// A finding verified as still present is the most informative match, so it
|
|
93
|
+
// wins a key two findings share; otherwise first write wins and the result
|
|
94
|
+
// does not depend on the order findings came back in.
|
|
95
|
+
for (const key of joinKeys(f.evidence)) {
|
|
96
|
+
const prev = filedKeys.get(key);
|
|
97
|
+
if (!prev || (f.verdict && f.verifiedAt && !(prev.verdict && prev.verifiedAt)))
|
|
98
|
+
filedKeys.set(key, f);
|
|
99
|
+
}
|
|
100
|
+
}
|
|
101
|
+
const claims = decisions.filter((d) => d.verdict === "defect" && d.evidence !== null);
|
|
102
|
+
const checkable = claims.filter((d) => usableConfidence(d.confidence) !== null && joinKeys(d.evidence).size > 0);
|
|
103
|
+
const unjoinable = claims.length - checkable.length;
|
|
104
|
+
if (checkable.length === 0)
|
|
105
|
+
return unjoinable > 0 ? { checkable: 0, filed: 0, stated: 0, buckets: [], ece: 0, verified: { present: 0, gone: 0, changed: 0 }, unjoinable } : null;
|
|
106
|
+
const buckets = BUCKET_EDGES.map(() => ({ n: 0, conf: 0, hits: 0 }));
|
|
107
|
+
const verified = { present: 0, gone: 0, changed: 0 };
|
|
108
|
+
const matched = new Map();
|
|
109
|
+
let filed = 0;
|
|
110
|
+
let stated = 0;
|
|
111
|
+
for (const d of checkable) {
|
|
112
|
+
let match;
|
|
113
|
+
for (const key of joinKeys(d.evidence)) {
|
|
114
|
+
match = filedKeys.get(key);
|
|
115
|
+
if (match)
|
|
116
|
+
break;
|
|
117
|
+
}
|
|
118
|
+
const hit = match !== undefined;
|
|
119
|
+
const confidence = usableConfidence(d.confidence);
|
|
120
|
+
const b = buckets[bucketOf(confidence)];
|
|
121
|
+
b.n += 1;
|
|
122
|
+
b.conf += confidence;
|
|
123
|
+
if (hit)
|
|
124
|
+
b.hits += 1;
|
|
125
|
+
if (hit)
|
|
126
|
+
filed += 1;
|
|
127
|
+
stated += confidence;
|
|
128
|
+
// By finding, not by decision: several decisions can match one finding,
|
|
129
|
+
// and this sentence counts findings.
|
|
130
|
+
if (match)
|
|
131
|
+
matched.set(match.id, match);
|
|
132
|
+
}
|
|
133
|
+
for (const f of matched.values()) {
|
|
134
|
+
// A verdict read from disk is not guaranteed to be one of the three.
|
|
135
|
+
if (f.verifiedAt && f.verdict && f.verdict in verified)
|
|
136
|
+
verified[f.verdict] += 1;
|
|
137
|
+
}
|
|
138
|
+
const n = checkable.length;
|
|
139
|
+
let ece = 0;
|
|
140
|
+
const out = [];
|
|
141
|
+
buckets.forEach((b, i) => {
|
|
142
|
+
if (b.n === 0)
|
|
143
|
+
return;
|
|
144
|
+
const meanConf = b.conf / b.n;
|
|
145
|
+
const rate = b.hits / b.n;
|
|
146
|
+
ece += (b.n / n) * Math.abs(meanConf - rate);
|
|
147
|
+
out.push({ label: bucketLabel(i), decisions: b.n, stated: meanConf, filed: rate });
|
|
148
|
+
});
|
|
149
|
+
return { checkable: n, filed, stated: stated / n, buckets: out, ece, verified, unjoinable };
|
|
150
|
+
}
|
|
151
|
+
const pct = (x) => `${Math.round(x * 100)}%`;
|
|
152
|
+
/**
|
|
153
|
+
* The report section. It leads with what the number is not, because "ECE 0.24"
|
|
154
|
+
* above a table of percentages reads as a verdict on the app rather than on
|
|
155
|
+
* the lanes that judged it.
|
|
156
|
+
*/
|
|
157
|
+
export function formatCalibration(c) {
|
|
158
|
+
if (!c)
|
|
159
|
+
return [];
|
|
160
|
+
if (c.checkable < MIN_FOR_A_VERDICT) {
|
|
161
|
+
// Suppress the NUMBER, not the fact. A run that recorded seven decisions
|
|
162
|
+
// and a run that recorded none look identical when the section simply
|
|
163
|
+
// vanishes, and the reader concludes the feature is broken — which is this
|
|
164
|
+
// project's own definition of a silent path.
|
|
165
|
+
const had = c.checkable + c.unjoinable;
|
|
166
|
+
if (had === 0)
|
|
167
|
+
return [];
|
|
168
|
+
return [
|
|
169
|
+
`## How well the lanes judged`,
|
|
170
|
+
``,
|
|
171
|
+
`Not enough to say yet: ${c.checkable} lane decision(s) could be checked${c.unjoinable > 0 ? ` (and ${c.unjoinable} could not be looked up at all)` : ""}, ` +
|
|
172
|
+
`and ${MIN_FOR_A_VERDICT} are needed before a calibration figure survives one of them changing.`,
|
|
173
|
+
``,
|
|
174
|
+
];
|
|
175
|
+
}
|
|
176
|
+
const lines = [
|
|
177
|
+
`## How well the lanes judged`,
|
|
178
|
+
``,
|
|
179
|
+
`${c.checkable} lane decision(s) called a defect on an endpoint that failed; ${c.filed} of them match a finding this project holds (${pct(c.filed / c.checkable)}), ` +
|
|
180
|
+
`against a mean stated confidence of ${c.stated.toFixed(2)}.`,
|
|
181
|
+
``,
|
|
182
|
+
`**What this is not.** It covers the project's whole history, not this run, and it measures agreement between the lanes and the bar this project applies — not whether the app is broken. ` +
|
|
183
|
+
`A lane can be perfectly calibrated against a planner that files the wrong things. Two things also count against a lane through no fault of its own: a defect filed with no machine signature, ` +
|
|
184
|
+
`and a finding whose signature was dropped when the store merged it into another. Read the figure as an ordering, not a measurement.`,
|
|
185
|
+
``,
|
|
186
|
+
`| Stated confidence | Decisions | Said | Filed |`,
|
|
187
|
+
`|---|---:|---:|---:|`,
|
|
188
|
+
];
|
|
189
|
+
for (const b of c.buckets) {
|
|
190
|
+
lines.push(`| ${b.label} | ${b.decisions} | ${b.stated.toFixed(2)} | ${pct(b.filed)} |`);
|
|
191
|
+
}
|
|
192
|
+
lines.push(``, `Expected calibration error **${c.ece.toFixed(2)}** — the gap between what the lanes said and what happened, weighted by bucket. ${readEce(c.ece)}`);
|
|
193
|
+
if (c.unjoinable > 0) {
|
|
194
|
+
lines.push(``, `${c.unjoinable} further decision(s) called a defect but named no failing endpoint, so nothing could be looked up for them. They are excluded above rather than counted as wrong.`);
|
|
195
|
+
}
|
|
196
|
+
const seen = c.verified.present + c.verified.gone + c.verified.changed;
|
|
197
|
+
if (seen > 0) {
|
|
198
|
+
lines.push(``, `${seen} of the findings those decisions matched ${seen === 1 ? "has" : "have"} since been re-tested with \`scout_verify\`: ${c.verified.present} still present, ${c.verified.gone} gone, ${c.verified.changed} changed. ` +
|
|
199
|
+
`Those verdicts are evidence about the app, unlike the table above. Findings no lane decision matched are not counted here, so this will not agree with the re-tested column elsewhere in the report.`);
|
|
200
|
+
}
|
|
201
|
+
lines.push(``);
|
|
202
|
+
return lines;
|
|
203
|
+
}
|
|
204
|
+
/** Below this, the number swings on a single decision and is worse than no number. */
|
|
205
|
+
export const MIN_FOR_A_VERDICT = 8;
|
|
206
|
+
/**
|
|
207
|
+
* What the figure suggests, hedged on purpose. The join loses decisions it
|
|
208
|
+
* cannot key and findings whose signature a merge discarded, and both push the
|
|
209
|
+
* error upward — so a confident "these numbers are wrong" can itself be wrong.
|
|
210
|
+
*/
|
|
211
|
+
function readEce(ece) {
|
|
212
|
+
if (ece <= 0.05)
|
|
213
|
+
return "The lanes' numbers track what happened closely.";
|
|
214
|
+
if (ece <= 0.15)
|
|
215
|
+
return "Close enough to compare lanes by.";
|
|
216
|
+
return "Far enough apart that a lane's confidence is worth reading as an ordering rather than a rate — after checking the exclusions above, which push this upward.";
|
|
217
|
+
}
|
package/dist/engine/lane.js
CHANGED
|
@@ -138,14 +138,28 @@ export function parseLaneReport(text, expectedLane) {
|
|
|
138
138
|
*/
|
|
139
139
|
export function laneReportInstruction(lane) {
|
|
140
140
|
return [
|
|
141
|
-
|
|
142
|
-
|
|
143
|
-
|
|
144
|
-
|
|
145
|
-
|
|
146
|
-
|
|
141
|
+
...LANE_RUBRIC,
|
|
142
|
+
// The ONLY lane-specific sentence, and it comes last on purpose. Every
|
|
143
|
+
// lane in a wave is given the same rubric, so keeping it byte-identical up
|
|
144
|
+
// to here makes it one shared prompt prefix: the cache hits from the
|
|
145
|
+
// second lane onward instead of diverging at the first sentence, which is
|
|
146
|
+
// what putting the name in the shape line used to do.
|
|
147
|
+
`Your lane name is ${JSON.stringify(lane)}; put exactly that in "lane".`,
|
|
147
148
|
].join(" ");
|
|
148
149
|
}
|
|
150
|
+
/**
|
|
151
|
+
* Everything every lane is told, identical for all of them. Built once from
|
|
152
|
+
* the same constants the parser enforces, because a limit a lane is not told
|
|
153
|
+
* refuses good replies.
|
|
154
|
+
*/
|
|
155
|
+
const LANE_RUBRIC = [
|
|
156
|
+
"Reply with ONE JSON object and nothing else — no prose before or after it, no explanation, no headings. A fenced ```json block is fine.",
|
|
157
|
+
`Shape: {"lane":<your lane name>,"status":<${quoteAll(LANE_STATUSES)}>,"decisions":[…],"routes":[…],"blocked_by":<string or null>}.`,
|
|
158
|
+
`Each decision: {"observation":<a short id for what was observed, unique in the report, at most ${LANE_OBSERVATION_MAX} characters>,"verdict":<${quoteAll(LANE_VERDICTS)}>,"severity":<${quoteAll(LANE_SEVERITIES)} or null>,"category":<${quoteAll(LANE_CATEGORIES)} or null>,"confidence":<0..1>,"evidence":<machine signature such as "GET /api/things 500", or null>}.`,
|
|
159
|
+
`A "defect" must carry a severity and a category. "evidence" is a signature, not a sentence: at most ${LANE_EVIDENCE_MAX} characters. "confidence" is how sure you are of the verdict, calibrated: 0.5 means a coin flip, 0.95 means you would bet on it.`,
|
|
160
|
+
`"routes" lists the routes you covered, each at most ${LANE_ROUTE_MAX} characters. "blocked_by" is one line of at most ${LANE_BLOCKED_BY_MAX} characters saying what stopped you: required when the status is "blocked", allowed with "partial", null with "complete"; the detail belongs in a finding. The lane name is at most ${LANE_NAME_MAX} characters. At most ${LANE_MAX_ITEMS} decisions and ${LANE_MAX_ITEMS} routes. Unknown keys are refused.`,
|
|
161
|
+
"The object IS your final report: whatever hands it back must hand back the object verbatim, not a summary of it.",
|
|
162
|
+
];
|
|
149
163
|
function quoteAll(values) {
|
|
150
164
|
return values.map((v) => `"${v}"`).join("|");
|
|
151
165
|
}
|
package/dist/engine/memory.js
CHANGED
|
@@ -250,8 +250,39 @@ export function mergeMemory(mine, theirs) {
|
|
|
250
250
|
}
|
|
251
251
|
if (Object.keys(out.roleAccess).length === 0)
|
|
252
252
|
delete out.roleAccess;
|
|
253
|
+
// Lanes commonly run in SEPARATE processes — that is the point of a lane —
|
|
254
|
+
// so without this the spread at the top would keep one process's decisions
|
|
255
|
+
// and drop every other lane's, which is the exact bug this merge exists to
|
|
256
|
+
// prevent for findings. Keyed so merging the same foreign document twice is
|
|
257
|
+
// idempotent.
|
|
258
|
+
const byDecision = new Map();
|
|
259
|
+
for (const d of theirs.laneDecisions ?? [])
|
|
260
|
+
byDecision.set(decisionKey(d), d);
|
|
261
|
+
for (const d of mine.laneDecisions ?? [])
|
|
262
|
+
byDecision.set(decisionKey(d), d);
|
|
263
|
+
out.laneDecisions = [...byDecision.values()].sort((a, b) => a.at.localeCompare(b.at)).slice(-MAX_LANE_DECISIONS);
|
|
264
|
+
if (out.laneDecisions.length === 0)
|
|
265
|
+
delete out.laneDecisions;
|
|
253
266
|
return out;
|
|
254
267
|
}
|
|
268
|
+
/**
|
|
269
|
+
* What makes two stored decisions the same judgement.
|
|
270
|
+
*
|
|
271
|
+
* Deliberately NOT the timestamp. `at` is stamped when the planner folds the
|
|
272
|
+
* reply, not when the lane judged, so relaying one reply twice — which the
|
|
273
|
+
* protocol invites, since a refused report is asked for again — wrote every
|
|
274
|
+
* decision a second time and doubled the lane's weight in the calibration.
|
|
275
|
+
* Identity is what was decided, so re-folding the same reply is a no-op.
|
|
276
|
+
*/
|
|
277
|
+
function decisionKey(d) {
|
|
278
|
+
return [d.lane, d.observation, d.verdict, d.severity ?? "", d.category ?? "", d.confidence, d.evidence ?? ""].join("|");
|
|
279
|
+
}
|
|
280
|
+
/**
|
|
281
|
+
* Most lane decisions kept. Calibration wants a few dozen; a long-lived
|
|
282
|
+
* project would otherwise accumulate every decision ever made and re-serialise
|
|
283
|
+
* them on each save, which is what made an old history slow to open.
|
|
284
|
+
*/
|
|
285
|
+
export const MAX_LANE_DECISIONS = 1000;
|
|
255
286
|
const MAX_DISCOVERED_ROUTES = 300;
|
|
256
287
|
/** Shared finding-similarity helpers (used by live dedup and retro-merge). */
|
|
257
288
|
function findingTokens(s) {
|
|
@@ -295,7 +326,22 @@ function findingLiterals(...texts) {
|
|
|
295
326
|
* (`/api/users/7` and `/api/users/9` are one endpoint) so a per-record repro
|
|
296
327
|
* does not read as a per-record bug.
|
|
297
328
|
*/
|
|
298
|
-
|
|
329
|
+
/**
|
|
330
|
+
* The signatures that identify a BUG: only those carrying a failure status.
|
|
331
|
+
*
|
|
332
|
+
* A bare `POST /api/orders` says which endpoint was involved, not what went
|
|
333
|
+
* wrong — a double submit and an accepted negative quantity both name it, and
|
|
334
|
+
* are two bugs. Exported so calibration applies the store's own rule instead
|
|
335
|
+
* of a second copy of this regex, which is what it had.
|
|
336
|
+
*/
|
|
337
|
+
export function failingSignatures(evidence) {
|
|
338
|
+
return new Set([...endpointSignatures(evidence)].filter((sig) => /\s[45]\d{2}$/.test(sig)));
|
|
339
|
+
}
|
|
340
|
+
/**
|
|
341
|
+
* The endpoint signatures a piece of evidence names: method + normalised path,
|
|
342
|
+
* with the failure status when one follows.
|
|
343
|
+
*/
|
|
344
|
+
export function endpointSignatures(evidence) {
|
|
299
345
|
const out = new Set();
|
|
300
346
|
if (!evidence)
|
|
301
347
|
return out;
|
|
@@ -338,11 +384,10 @@ function sharesEndpointSignature(a, b) {
|
|
|
338
384
|
// `POST /api/orders` (the endpoint answered 2xx, or no status was named)
|
|
339
385
|
// says which endpoint was involved, not what went wrong: a double submit
|
|
340
386
|
// and an accepted negative quantity both name it, and are two bugs.
|
|
341
|
-
const
|
|
342
|
-
const aSigs = failing(a.evidence);
|
|
387
|
+
const aSigs = failingSignatures(a.evidence);
|
|
343
388
|
if (aSigs.size === 0)
|
|
344
389
|
return false;
|
|
345
|
-
for (const sig of
|
|
390
|
+
for (const sig of failingSignatures(b.evidence))
|
|
346
391
|
if (aSigs.has(sig))
|
|
347
392
|
return true;
|
|
348
393
|
return false;
|
|
@@ -782,6 +827,45 @@ export class MemoryStore {
|
|
|
782
827
|
fs.writeFileSync(this.assumptionsPath, content);
|
|
783
828
|
return true;
|
|
784
829
|
}
|
|
830
|
+
/** What the lanes decided, oldest first. Empty on a project that has never run one. */
|
|
831
|
+
get laneDecisions() {
|
|
832
|
+
return this.data.laneDecisions ?? [];
|
|
833
|
+
}
|
|
834
|
+
/**
|
|
835
|
+
* Record what a lane decided. Called once per accepted lane report, so the
|
|
836
|
+
* confidence it stated can be checked later against what the run filed.
|
|
837
|
+
* Free text from the lane is redacted like every other stored string: an
|
|
838
|
+
* observation is written by a model reading the app under test.
|
|
839
|
+
*/
|
|
840
|
+
addLaneDecisions(lane, decisions) {
|
|
841
|
+
if (decisions.length === 0)
|
|
842
|
+
return 0;
|
|
843
|
+
const list = this.data.laneDecisions ?? [];
|
|
844
|
+
const seen = new Set(list.map(decisionKey));
|
|
845
|
+
let added = 0;
|
|
846
|
+
for (const d of decisions) {
|
|
847
|
+
const record = {
|
|
848
|
+
...d,
|
|
849
|
+
lane,
|
|
850
|
+
observation: redactSecrets(d.observation).slice(0, 200),
|
|
851
|
+
evidence: d.evidence === null ? null : redactSecrets(d.evidence).slice(0, 200),
|
|
852
|
+
};
|
|
853
|
+
if (seen.has(decisionKey(record)))
|
|
854
|
+
continue;
|
|
855
|
+
seen.add(decisionKey(record));
|
|
856
|
+
list.push(record);
|
|
857
|
+
added += 1;
|
|
858
|
+
}
|
|
859
|
+
this.data.laneDecisions = list.slice(-MAX_LANE_DECISIONS);
|
|
860
|
+
if (added > 0)
|
|
861
|
+
this.flush();
|
|
862
|
+
// What survives the cap. Appends go to the tail and the cap keeps the
|
|
863
|
+
// tail, so all of `added` survives unless the call itself exceeds the cap.
|
|
864
|
+
// Measuring it as growth instead looked right and was not: `list` aliases
|
|
865
|
+
// the stored array, so the "before" length was read after the appends and
|
|
866
|
+
// every call after the first reported nothing kept — while storing fine.
|
|
867
|
+
return Math.min(added, MAX_LANE_DECISIONS);
|
|
868
|
+
}
|
|
785
869
|
/** Mark a finding resolved; returns it or null. */
|
|
786
870
|
resolveFinding(id) {
|
|
787
871
|
const f = this.data.findings.find((x) => x.id === id);
|
package/dist/engine/report.js
CHANGED
|
@@ -4,6 +4,7 @@ import { SHARED_CHROME_ROUTE } from "./memory.js";
|
|
|
4
4
|
import { sayVerification } from "./verify.js";
|
|
5
5
|
import { feedForSession } from "./live.js";
|
|
6
6
|
import { buildReplayHtml, evidenceFor } from "./replay.js";
|
|
7
|
+
import { calibrate, formatCalibration } from "./calibration.js";
|
|
7
8
|
import { formatPace, measurePace } from "./pace.js";
|
|
8
9
|
function playwrightSkeleton(f) {
|
|
9
10
|
const routeClass = f.state.split("#")[0].split("?")[0];
|
|
@@ -616,6 +617,9 @@ export function generateReport(memory, oracleLog, extras, opts = {}) {
|
|
|
616
617
|
if (extras?.pace !== false) {
|
|
617
618
|
lines.push(...formatPace(measurePace(memory.actionLog, Date.now(), extras?.attachedSessions ?? [])));
|
|
618
619
|
}
|
|
620
|
+
// Only on a run that used lanes, and only once enough of them have been
|
|
621
|
+
// judged for the number to mean anything; formatCalibration decides both.
|
|
622
|
+
lines.push(...formatCalibration(calibrate(memory.laneDecisions, memory.findings)));
|
|
619
623
|
const markdown = lines.join("\n");
|
|
620
624
|
const outPath = path.join(memory.dir, "report.md");
|
|
621
625
|
const htmlPath = path.join(memory.dir, "report.html");
|
package/dist/mcp-server.js
CHANGED
|
@@ -528,10 +528,35 @@ server.registerTool("scout_lane_report", {
|
|
|
528
528
|
if (reply === undefined)
|
|
529
529
|
return { content: [{ type: "text", text: laneReportInstruction(lane) }] };
|
|
530
530
|
const parsed = parseLaneReport(reply, lane);
|
|
531
|
-
|
|
532
|
-
|
|
533
|
-
|
|
534
|
-
|
|
531
|
+
if (!parsed.ok) {
|
|
532
|
+
return {
|
|
533
|
+
content: [
|
|
534
|
+
{ type: "text", text: `Lane report REFUSED: ${parsed.reason}. Ask the lane once for the corrected object; do not re-judge its prose.` },
|
|
535
|
+
],
|
|
536
|
+
};
|
|
537
|
+
}
|
|
538
|
+
// Keep what the lane decided, so the confidence it stated can be checked
|
|
539
|
+
// against what the run goes on to file. Best-effort: a lane report is
|
|
540
|
+
// still accepted if this project has no memory open yet, because the
|
|
541
|
+
// planner's fold must not depend on where the report was written.
|
|
542
|
+
// ONLY the lane's own session. Falling back to any engine with memory
|
|
543
|
+
// open put one project's decisions into another project's store whenever
|
|
544
|
+
// two sessions were attached to different apps — and the lane having
|
|
545
|
+
// already closed makes that the ordinary case, not an edge one.
|
|
546
|
+
const owner = engines.get(lane)?.memory;
|
|
547
|
+
const at = new Date().toISOString();
|
|
548
|
+
const kept = owner
|
|
549
|
+
? owner.addLaneDecisions(lane, parsed.report.decisions.map((d) => ({ ...d, lane, at })))
|
|
550
|
+
: 0;
|
|
551
|
+
// Say when nothing was kept. Every lane closing its session before the
|
|
552
|
+
// planner folds its report is the order the method describes, and it
|
|
553
|
+
// leaves no memory to write to — reporting a bare "accepted" while the
|
|
554
|
+
// skill promises the decisions are kept is the kind of silence that
|
|
555
|
+
// makes a later calibration section look wrong rather than absent.
|
|
556
|
+
const note = kept > 0
|
|
557
|
+
? ` (${kept} decision(s) kept for calibration)`
|
|
558
|
+
: ` (nothing kept for calibration — session ${JSON.stringify(lane)} is not attached here, so there is no project memory to write to. Fold a lane report before closing that lane's session.)`;
|
|
559
|
+
return { content: [{ type: "text", text: `Lane report accepted — ${summarizeLaneReport(parsed.report)}${note}` }] };
|
|
535
560
|
}
|
|
536
561
|
catch (err) {
|
|
537
562
|
return errorText(err);
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "scenescout",
|
|
3
|
-
"version": "3.
|
|
3
|
+
"version": "3.4.0",
|
|
4
4
|
"description": "SceneScout — exploratory UI testing for AI coding agents. An MCP server that gives any agent (Claude Code, Cursor, VS Code Copilot, Codex, Gemini CLI and others) a structured view of a running web app, always-on oracles, a network-level write policy, memory across runs and a gap-checked report.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"author": "brunoboto96",
|
|
@@ -71,7 +71,7 @@
|
|
|
71
71
|
"mcp-check": "npm run build && npm run mcp-check:run",
|
|
72
72
|
"mcp-check:run": "tsx scripts/mcp-check.ts",
|
|
73
73
|
"test": "npm run build && npm run test:unit && npm run smoke:run && npm run mcp-check:run",
|
|
74
|
-
"test:unit": "npm run scan-test && npm run oracle-test && npm run policy-test && npm run fixture-test && npm run dispatch-test && npm run design-test && npm run contract-test && npm run pace-test && npm run request-test && npm run settle-test && npm run claims-test && npm run verify-test && npm run brief-test && npm run lane-test && npm run memory-test && npm run install-test && npm run live-test && npm run demo-test && npm run hygiene-test",
|
|
74
|
+
"test:unit": "npm run scan-test && npm run oracle-test && npm run policy-test && npm run fixture-test && npm run dispatch-test && npm run design-test && npm run contract-test && npm run pace-test && npm run request-test && npm run settle-test && npm run claims-test && npm run verify-test && npm run brief-test && npm run lane-test && npm run calibration-test && npm run memory-test && npm run install-test && npm run live-test && npm run demo-test && npm run hygiene-test",
|
|
75
75
|
"scan-test": "tsx scripts/scan-test.ts",
|
|
76
76
|
"oracle-test": "tsx --test scripts/oracle-test.ts",
|
|
77
77
|
"policy-test": "tsx --test scripts/policy-test.ts",
|
|
@@ -90,7 +90,8 @@
|
|
|
90
90
|
"settle-test": "tsx --test scripts/settle-test.ts",
|
|
91
91
|
"claims-test": "tsx --test scripts/claims-test.ts",
|
|
92
92
|
"verify-test": "tsx --test scripts/verify-test.ts",
|
|
93
|
-
"brief-test": "tsx --test scripts/brief-test.ts"
|
|
93
|
+
"brief-test": "tsx --test scripts/brief-test.ts",
|
|
94
|
+
"calibration-test": "tsx --test scripts/calibration-test.ts"
|
|
94
95
|
},
|
|
95
96
|
"dependencies": {
|
|
96
97
|
"@modelcontextprotocol/sdk": "^1.12.0",
|
|
@@ -74,7 +74,7 @@ Some flows need a TEAM — a document one role submits and another approves, a r
|
|
|
74
74
|
- `scout_close {all: true}` at the end of a multi-role run; `scout_close {session}` to drop one role early.
|
|
75
75
|
- **Several agents in parallel** (subagents or a workflow, each driving its own session): each agent attaches its session when it STARTS and closes it by name when it FINISHES. Never open sessions ahead for agents that have not started, and never hand an open session from one agent to the next: an agent waiting for its turn should hold no browser. Give each session an `objective` when you attach it (`scout_attach {session, objective:"Approve and reject orders as a manager"}`) and wrap each goal in `scout_journey`: the live view shows that objective and the task it is on beside the session's feed, which is how the person watching knows what every agent is for. Keep that goal TRUE: one journey per goal, one goal per thing you are checking ("Save a settings change as the auditor", not "Check every page"), ended the moment it is decided and the next one started before you move on. A journey that outlives its goal shows the viewer an objective the session left behind minutes ago. Run roughly as many agents at once as the machine has cores, less two, since each drives a real browser; beyond that they only queue. Exploring one area is well within a mid-tier model, so run these agents on one (Sonnet or its equivalent in your client) unless the user names a model; keep the larger model for the agent that plans the split and writes the report. No agent may call `scout_close {all: true}` while others run — only the last step, once every agent has finished.
|
|
76
76
|
- **Let the engine divide the app: `scout_lane_brief {lanes, goal}`.** Call it after the first crawl, when route knowledge is complete. It splits the known routes into whole modules — everything under `/orders` goes to one lane, so that lane carries state between its own steps instead of re-learning the app on every route — deals the modules out so the lanes come out within a route or two of each other, and returns each lane's session name, the `objective` to attach it with, and the routes it owns. Pass `routes` to split a subset instead. Hand each brief to its agent verbatim. Dividing by hand fails in two ways that a finished run cannot tell apart from success: two lanes audit the same register while a third module is never opened, and route coverage reads complete either way; and lanes launch without an objective, so the person watching sees browsers clicking through their app with nothing to say why.
|
|
77
|
-
- **Lanes report in typed decisions, not prose.** Each lane files its findings with `scout_finding` as it goes, so the report already has them; what the lane hands BACK to the planner is one JSON object, the lane report, and nothing else. Get the paragraph to put in a lane's prompt from `scout_lane_report {lane}`: it names the shape (`status`, one decision per observation with `verdict`, `severity`, `category`, a calibrated `confidence` and a machine-signature `evidence`, the `routes` covered, `blocked_by`), every value's closed set (the categories are `scout_finding`'s) and every length limit, because a limit a lane is not told refuses good replies. When the lane hands back, pass its text to `scout_lane_report {lane, reply}`: it returns the one-line fold (defects, highs, unsure, mean confidence, routes, what blocked it) or the reason the reply was refused. Relay a refusal to the lane once and ask for the corrected object; never re-judge its prose yourself, and tell every lane that the object IS its final report, since a lane that writes the object and then hands back a summary of it has failed. The planner then folds lanes by counting, not by reading: a lane's reply is a few hundred tokens whatever it found, and a value the schema refuses is caught at the boundary instead of becoming a severity like "Low-Medium" in the report. Where the client lets you set a lane's reasoning effort, use **medium**: in this project's benchmark it judged as well as high in three-quarters of the time and under half of xhigh, low inflated severities, and max took three minutes per lane for no better agreement.
|
|
77
|
+
- **Lanes report in typed decisions, not prose.** Each lane files its findings with `scout_finding` as it goes, so the report already has them; what the lane hands BACK to the planner is one JSON object, the lane report, and nothing else. Get the paragraph to put in a lane's prompt from `scout_lane_report {lane}`: it names the shape (`status`, one decision per observation with `verdict`, `severity`, `category`, a calibrated `confidence` and a machine-signature `evidence`, the `routes` covered, `blocked_by`), every value's closed set (the categories are `scout_finding`'s) and every length limit, because a limit a lane is not told refuses good replies. When the lane hands back, pass its text to `scout_lane_report {lane, reply}`: it returns the one-line fold (defects, highs, unsure, mean confidence, routes, what blocked it) or the reason the reply was refused, and keeps each decision so the confidence the lane stated can be checked against what the run went on to file — the report's calibration section is built from that, and it only appears once enough decisions exist to mean something. **So the confidence is worth stating honestly**: a lane that writes 0.95 on everything makes the section say so. Relay a refusal to the lane once and ask for the corrected object; never re-judge its prose yourself, and tell every lane that the object IS its final report, since a lane that writes the object and then hands back a summary of it has failed. The planner then folds lanes by counting, not by reading: a lane's reply is a few hundred tokens whatever it found, and a value the schema refuses is caught at the boundary instead of becoming a severity like "Low-Medium" in the report. Where the client lets you set a lane's reasoning effort, use **medium**: in this project's benchmark it judged as well as high in three-quarters of the time and under half of xhigh, low inflated severities, and max took three minutes per lane for no better agreement.
|
|
78
78
|
|
|
79
79
|
## The impatient-user pass (extensive)
|
|
80
80
|
|