scenescout 3.11.1 โ 3.12.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +14 -0
- package/README.md +5 -1
- package/dist/check-run.js +4 -2
- package/dist/engine/bench.js +147 -14
- package/dist/engine/calibration.js +23 -6
- package/dist/engine/check.js +86 -14
- package/dist/engine/design.js +15 -1
- package/dist/engine/fingerprint.js +57 -0
- package/dist/engine/lane.js +57 -3
- package/dist/engine/memory.js +97 -5
- package/dist/engine/report.js +37 -2
- package/dist/engine/verify.js +4 -3
- package/dist/mcp-server.js +23 -9
- package/package.json +1 -1
- package/skills/scenescout/SKILL.md +4 -3
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,19 @@
|
|
|
1
1
|
# scenescout
|
|
2
2
|
|
|
3
|
+
## 3.12.0
|
|
4
|
+
|
|
5
|
+
### Minor Changes
|
|
6
|
+
|
|
7
|
+
- 967e98c: Add a "worth a look" tier for observations that are defects only under a convention of the project the run cannot see, such as a spacing scale, link styling in navigation, or test ids on every control. SceneScout reports each one with the convention that would decide it, and never counts it as a defect.
|
|
8
|
+
|
|
9
|
+
- Lane reports accept the verdict `worth_a_look`, which must name its `convention`. A `convention` on any other verdict is ignored, and the fold says so. Lane calibration and the benchmark's key calibration leave it unscored and count it under its own reason.
|
|
10
|
+
- `scout_finding` takes an optional `convention`, which files the finding in this tier. The report lists these findings under "Worth a look", below the findings, as "a defect only if your project uses โฆ", and leaves them out of every defect total. Filing the same thing again as a defect promotes it, at the defect's severity, and the reply says so. The benchmark sets such findings aside from recall and precision, and the fold lists a lane's judged defect as unfiled when it was filed only as worth a look.
|
|
11
|
+
- `scenescout check` reports two new rules, `off-grid-spacing` and `indistinct-link`, in this tier. They have no severity and never fail the gate at any `--fail-on`. SARIF reports them at level `note`. The report and the job summary list them in their own section, `check.json` puts them under `worthALook`, separate from `issues`, and the GitHub Action publishes their number as the `worth-a-look` output. Every existing rule is unchanged.
|
|
12
|
+
|
|
13
|
+
### Patch Changes
|
|
14
|
+
|
|
15
|
+
- c40892f: `scout_lane_report` keeps the routes each lane's accepted report lists in the project's memory (`laneRoutes` in `.scenescout/memory.json`), with query strings and fragments removed so no token in an address is stored, and `npm run bench -- --archive` carries them into the run archive. The benchmark then decides whether a lane's "not a defect" is a remark about another lane's page by where the matched defect is, not by how the verdict is worded: a defect none of whose pages the lane covered is set aside as another lane's and listed on the scorecard, and one on the lane's own page, or one every lane can reach, is scored. Answer-key entries can name further pages with `alsoOn` and mark a defect every lane can reach with `everyPage`. Archives made before routes were kept, and lanes with a route that names no page, are scored by wording, as before.
|
|
16
|
+
|
|
3
17
|
## 3.11.1
|
|
4
18
|
|
|
5
19
|
### Patch Changes
|
package/README.md
CHANGED
|
@@ -332,6 +332,7 @@ A `๐ก WRITE-POLICY blocked` notice is the safety net doing its job, not an app
|
|
|
332
332
|
`.scenescout/report.md` โ a deduplicated, worst-first report with:
|
|
333
333
|
|
|
334
334
|
- ๐ **Findings** with repro traces and generated Playwright regression-test skeletons.
|
|
335
|
+
- ๐ **Worth a look** โ observations that are defects only under a convention of your project the run cannot see (a spacing scale, link styling in navigation, test ids on every control), each naming that convention. Listed below the findings and not counted as defects ([ADR 13](docs/adr/0013-a-convention-is-the-projects-to-decide.md)).
|
|
335
336
|
- ๐ฏ **Page scores** (0โ100: a11y ยท craft ยท consistency ยท task-clarity), ranked worst-first, with stale scores from old runs marked as such.
|
|
336
337
|
- ๐ฅ **A role capability matrix** โ what each role could and couldn't reach.
|
|
337
338
|
- ๐งพ **A gap ledger** โ everything *not* done, so the report is honest about its own coverage.
|
|
@@ -355,6 +356,8 @@ An exploratory run is driven by a model, so two runs never find exactly the same
|
|
|
355
356
|
- contrast and focus
|
|
356
357
|
- pages with no way out
|
|
357
358
|
|
|
359
|
+
Some of what it measures is a defect only under a convention the check cannot see: paddings off a 4px grid, and links styled like body text. Those are listed under **Worth a look**, each with the convention that would make it a defect. They are never counted and never fail the gate, at any `--fail-on`; SARIF reports them at level `note` ([ADR 13](docs/adr/0013-a-convention-is-the-projects-to-decide.md)).
|
|
360
|
+
|
|
358
361
|
```bash
|
|
359
362
|
npx scenescout check http://127.0.0.1:3000 --fail-on high
|
|
360
363
|
```
|
|
@@ -380,7 +383,7 @@ On GitHub Actions, this repository is also an action that installs everything an
|
|
|
380
383
|
With the default settings its saved flows send no HTTP write (they replay under observe's rule), and its crawl runs under `--mode observe` or `read-only`; `--flow-writes allow` lets flows write as `--mode` allows. By default it fails only on facts that mean a page is broken: a page that did not load, an uncaught exception, a 5xx, a failure shown as success. Other options:
|
|
381
384
|
|
|
382
385
|
- `--fail-on medium` or `low` makes the gate stricter.
|
|
383
|
-
- `--ignore <rule>` drops a rule.
|
|
386
|
+
- `--ignore <rule>` drops a rule, worth-a-look rules included.
|
|
384
387
|
- `--paths /a,/b` checks only those pages.
|
|
385
388
|
- `--storage-state <file>` checks while signed in.
|
|
386
389
|
- `--flows <dir>` or `off` chooses which saved flows to replay; `--retest off` skips re-testing open findings.
|
|
@@ -661,6 +664,7 @@ The load-bearing choices are recorded as ADRs โ read the relevant one before c
|
|
|
661
664
|
- [10 ยท A lane's confidence is checked, not trusted](docs/adr/0010-a-confidence-is-checked-not-trusted.md)
|
|
662
665
|
- [11 ยท A gate is deterministic, and fails only on what it can prove](docs/adr/0011-a-gate-is-deterministic-and-fails-only-on-what-it-can-prove.md)
|
|
663
666
|
- [12 ยท A check replays saved flows and re-tests open findings, within settings whose defaults do the least harm](docs/adr/0012-a-check-replays-saved-flows-and-reports-re-tests.md)
|
|
667
|
+
- [13 ยท What depends on a project's convention is the project's to decide](docs/adr/0013-a-convention-is-the-projects-to-decide.md)
|
|
664
668
|
|
|
665
669
|
---
|
|
666
670
|
|
package/dist/check-run.js
CHANGED
|
@@ -9,7 +9,7 @@ import fs from "node:fs";
|
|
|
9
9
|
import os from "node:os";
|
|
10
10
|
import path from "node:path";
|
|
11
11
|
import { BrowserEngine } from "./engine/browser.js";
|
|
12
|
-
import {
|
|
12
|
+
import { checkFindings, redactFlowRuns, redactRoute, redactRoutes, settingsOf, withoutOwnResponse, } from "./engine/check.js";
|
|
13
13
|
import { loadFlows, resolveFlowsDir } from "./engine/flow.js";
|
|
14
14
|
import { MemoryStore, MEMORY_DIRNAME } from "./engine/memory.js";
|
|
15
15
|
import { checkRetestPlan, retestResults, wellFormedFindings } from "./engine/verify.js";
|
|
@@ -123,13 +123,15 @@ export async function runCheck(options, log = () => { }, inputs = { flows: [], f
|
|
|
123
123
|
: null;
|
|
124
124
|
const measured = redactRoutes(routes.map(withoutOwnResponse));
|
|
125
125
|
const flows = redactFlowRuns(flowRuns);
|
|
126
|
+
const { issues, worthALook } = checkFindings(measured, start.origin, options.ignore, flows);
|
|
126
127
|
return {
|
|
127
128
|
url: redactRoute(options.url),
|
|
128
129
|
generatedAt: new Date().toISOString(),
|
|
129
130
|
mode: options.mode,
|
|
130
131
|
failOn: options.failOn,
|
|
131
132
|
routes: measured,
|
|
132
|
-
issues
|
|
133
|
+
issues,
|
|
134
|
+
worthALook,
|
|
133
135
|
// Routes that failed to load are issues already; "not visited" is only what --max-routes left out.
|
|
134
136
|
unvisited: options.paths ? [] : engine.crawlableRoutes().map(redactRoute),
|
|
135
137
|
ignored: options.ignore,
|
package/dist/engine/bench.js
CHANGED
|
@@ -31,6 +31,9 @@
|
|
|
31
31
|
import { createHash } from "node:crypto";
|
|
32
32
|
import { z } from "zod";
|
|
33
33
|
import { BUCKET_EDGES, bucketLabel, bucketOf } from "./calibration.js";
|
|
34
|
+
import { laneRoutePaths, stripRouteQuery } from "./fingerprint.js";
|
|
35
|
+
import { isWorthALook } from "./memory.js";
|
|
36
|
+
export { laneRoutePaths };
|
|
34
37
|
export const LEVELS = ["minimal", "medium", "extensive"];
|
|
35
38
|
const SEVERITIES = ["high", "medium", "low"];
|
|
36
39
|
const Pattern = z.string().refine((p) => {
|
|
@@ -42,10 +45,24 @@ const Pattern = z.string().refine((p) => {
|
|
|
42
45
|
return false;
|
|
43
46
|
}
|
|
44
47
|
}, { message: "is not a valid regular expression" });
|
|
48
|
+
/**
|
|
49
|
+
* Other pages the same defect shows on, beyond `route`. A lane that owns any
|
|
50
|
+
* of them owns the defect, so its "not a defect" is a verdict, not a remark
|
|
51
|
+
* about another lane's page.
|
|
52
|
+
*/
|
|
53
|
+
const AlsoOn = z.array(z.string().startsWith("/")).min(1).optional();
|
|
54
|
+
/**
|
|
55
|
+
* Every lane can reach the defect: it is in chrome every page carries (a
|
|
56
|
+
* shared nav, header or footer), or in an endpoint any lane can call. No lane
|
|
57
|
+
* can call it another lane's, so a verdict on it is always scored.
|
|
58
|
+
*/
|
|
59
|
+
const EveryPage = z.boolean().optional();
|
|
45
60
|
const Entry = z
|
|
46
61
|
.object({
|
|
47
62
|
id: z.string().min(1),
|
|
48
63
|
route: z.string().startsWith("/"),
|
|
64
|
+
alsoOn: AlsoOn,
|
|
65
|
+
everyPage: EveryPage,
|
|
49
66
|
title: z.string().min(1),
|
|
50
67
|
/** The category a finding for it should carry. Reported, not enforced: two categories can both be defensible. */
|
|
51
68
|
category: z.string().min(1),
|
|
@@ -83,6 +100,8 @@ const Contextual = z
|
|
|
83
100
|
.object({
|
|
84
101
|
id: z.string().min(1),
|
|
85
102
|
route: z.string().startsWith("/"),
|
|
103
|
+
alsoOn: AlsoOn,
|
|
104
|
+
everyPage: EveryPage,
|
|
86
105
|
title: z.string().min(1),
|
|
87
106
|
category: z.string().min(1),
|
|
88
107
|
/** The convention that decides it: where it would be a defect, and why the run cannot tell whether it holds here. */
|
|
@@ -130,6 +149,14 @@ export function parseKey(raw) {
|
|
|
130
149
|
const dup = ids.find((id, i) => ids.indexOf(id) !== i);
|
|
131
150
|
if (dup)
|
|
132
151
|
throw new Error(`The answer key uses the id ${JSON.stringify(dup)} twice.`);
|
|
152
|
+
for (const e of [...parsed.data.defects, ...parsed.data.alsoReal, ...parsed.data.contextual]) {
|
|
153
|
+
if (e.alsoOn && e.everyPage)
|
|
154
|
+
throw new Error(`The entry ${JSON.stringify(e.id)} has both alsoOn and everyPage; an entry on every page names no other pages.`);
|
|
155
|
+
const pages = [e.route, ...(e.alsoOn ?? [])];
|
|
156
|
+
const dup = pages.find((p, i) => pages.indexOf(p) !== i);
|
|
157
|
+
if (dup)
|
|
158
|
+
throw new Error(`The entry ${JSON.stringify(e.id)} names the page ${JSON.stringify(dup)} twice in route and alsoOn.`);
|
|
159
|
+
}
|
|
133
160
|
const real = new Set([...parsed.data.defects, ...parsed.data.alsoReal].map((e) => e.id));
|
|
134
161
|
for (const nd of parsed.data.nonDefects) {
|
|
135
162
|
const unknown = nd.overrides.find((id) => !real.has(id));
|
|
@@ -241,14 +268,14 @@ export function lintKey(key) {
|
|
|
241
268
|
*
|
|
242
269
|
* "defect" on a planted or also-real defect is right, and on a known
|
|
243
270
|
* non-defect is wrong. "not_a_defect" is the reverse. An "unsure" verdict, a
|
|
244
|
-
* dismissal as another lane's, a decision the key does not
|
|
245
|
-
* names ambiguously, and one about a contextual entry are not scored โ the same rule the in-product calibration
|
|
271
|
+
* "worth_a_look", a dismissal as another lane's, a decision the key does not
|
|
272
|
+
* name, one it names ambiguously, and one about a contextual entry are not scored โ the same rule the in-product calibration
|
|
246
273
|
* follows, for the same reason.
|
|
247
274
|
*/
|
|
248
|
-
export function judgeDecision(d, key) {
|
|
249
|
-
if (d.verdict === "unsure")
|
|
275
|
+
export function judgeDecision(d, key, laneRoutes) {
|
|
276
|
+
if (d.verdict === "unsure" || d.verdict === "worth_a_look")
|
|
250
277
|
return null;
|
|
251
|
-
if (
|
|
278
|
+
if (isOwnershipRemark(d, key, laneRoutes))
|
|
252
279
|
return null;
|
|
253
280
|
const m = classify(decisionText(d), key);
|
|
254
281
|
if (!m || m.kind === "ambiguous" || m.kind === "contextual")
|
|
@@ -258,13 +285,79 @@ export function judgeDecision(d, key) {
|
|
|
258
285
|
}
|
|
259
286
|
/** How a lane says "not mine": the wording lanes actually used, none of it about a particular app. */
|
|
260
287
|
const SCOPE_DISMISSAL_RE = /out.of.lane.scope|out-of-scope|out of (my|this) (lane|scope)|belongs?.to.{0,20}\b(lane|route|page)\b|belongs-to-|(handled|owned) by (the |another )?[\w-]* ?lane|lane owns it|not (in |on |from |part of )?(my|this) (lane|assigned routes|routes?|pages?)\b|outside (my|this) (lane|routes?|pages?)|not part of my (assigned )?routes/i;
|
|
261
|
-
/**
|
|
288
|
+
/**
|
|
289
|
+
* A not-a-defect verdict whose stated reason is that the thing is another
|
|
290
|
+
* lane's. The WORDING rule: all a run archived before lane routes were kept
|
|
291
|
+
* can be judged by, and still what decides for a lane whose routes are not
|
|
292
|
+
* known. It cannot tell a lane's own page from another's, which is why it is
|
|
293
|
+
* not widened (see `isOwnershipRemark`).
|
|
294
|
+
*/
|
|
262
295
|
export function isScopeDismissal(d) {
|
|
263
296
|
return d.verdict === "not_a_defect" && SCOPE_DISMISSAL_RE.test(decisionText(d));
|
|
264
297
|
}
|
|
265
|
-
|
|
298
|
+
/** The pages a key entry is on, in the same form as `laneRoutePaths`. */
|
|
299
|
+
function entryPaths(e) {
|
|
300
|
+
return [e.route, ...(e.alsoOn ?? [])].flatMap(laneRoutePaths);
|
|
301
|
+
}
|
|
302
|
+
/**
|
|
303
|
+
* A lane's routes as a set of paths, or undefined when they are not known
|
|
304
|
+
* well enough to decide by: none were recorded, or ANY of them names no path
|
|
305
|
+
* ("the orders area"). A route that could not be read may be the very page a
|
|
306
|
+
* verdict is about, and deciding without it would set that verdict aside.
|
|
307
|
+
*/
|
|
308
|
+
function lanePaths(lane, laneRoutes) {
|
|
309
|
+
const routes = laneRoutes && Object.hasOwn(laneRoutes, lane) ? laneRoutes[lane] : undefined;
|
|
310
|
+
if (!routes || routes.length === 0)
|
|
311
|
+
return undefined;
|
|
312
|
+
const paths = new Set();
|
|
313
|
+
for (const r of routes) {
|
|
314
|
+
const named = laneRoutePaths(r);
|
|
315
|
+
if (named.length === 0)
|
|
316
|
+
return undefined;
|
|
317
|
+
for (const p of named)
|
|
318
|
+
paths.add(p);
|
|
319
|
+
}
|
|
320
|
+
return paths;
|
|
321
|
+
}
|
|
322
|
+
/**
|
|
323
|
+
* A not-a-defect that is a remark about who owns the thing, not a verdict on
|
|
324
|
+
* whether it is broken.
|
|
325
|
+
*
|
|
326
|
+
* Where the lane's routes are known, decided by WHERE the thing is: the
|
|
327
|
+
* verdict matches a planted or also-real defect whose pages are all outside
|
|
328
|
+
* the lane's routes. The wording no longer matters either way โ "raised on /
|
|
329
|
+
* before navigating" from the orders lane is a remark about the dashboard's
|
|
330
|
+
* defect, and "handled by another lane" from the lane that owns the page is
|
|
331
|
+
* its own wrong verdict. A defect in chrome every page carries (`everyPage`)
|
|
332
|
+
* is every lane's, so it is never another lane's. A verdict that matches a
|
|
333
|
+
* known non-defect, a contextual entry, several entries or none is not about
|
|
334
|
+
* a defect with a page to own, and is judged as it would be otherwise.
|
|
335
|
+
*
|
|
336
|
+
* Where they are not known โ every archive made before routes were kept, and
|
|
337
|
+
* any lane that reported no route โ the wording rule decides, unchanged, so
|
|
338
|
+
* those runs score as they always have.
|
|
339
|
+
*/
|
|
340
|
+
export function ownershipRemark(d, key, laneRoutes) {
|
|
341
|
+
if (d.verdict !== "not_a_defect")
|
|
342
|
+
return null;
|
|
343
|
+
const own = lanePaths(d.lane, laneRoutes);
|
|
344
|
+
if (!own)
|
|
345
|
+
return isScopeDismissal(d) ? { by: "wording" } : null;
|
|
346
|
+
const m = classify(decisionText(d), key);
|
|
347
|
+
if (!m || (m.kind !== "defect" && m.kind !== "alsoReal"))
|
|
348
|
+
return null;
|
|
349
|
+
if (m.entry.everyPage)
|
|
350
|
+
return null;
|
|
351
|
+
return entryPaths(m.entry).some((p) => own.has(p)) ? null : { by: "routes", id: m.entry.id };
|
|
352
|
+
}
|
|
353
|
+
export function isOwnershipRemark(d, key, laneRoutes) {
|
|
354
|
+
return ownershipRemark(d, key, laneRoutes) !== null;
|
|
355
|
+
}
|
|
356
|
+
export function calibrateAgainstKey(decisions, key, laneRoutes) {
|
|
266
357
|
if (decisions.length === 0)
|
|
267
358
|
return null;
|
|
359
|
+
const lanes = [...new Set(decisions.map((d) => d.lane))];
|
|
360
|
+
const known = lanes.filter((l) => lanePaths(l, laneRoutes) !== undefined).length;
|
|
268
361
|
const buckets = BUCKET_EDGES.map(() => ({ n: 0, conf: 0, right: 0 }));
|
|
269
362
|
const out = {
|
|
270
363
|
judged: 0,
|
|
@@ -274,7 +367,10 @@ export function calibrateAgainstKey(decisions, key) {
|
|
|
274
367
|
badConfidence: 0,
|
|
275
368
|
unsure: 0,
|
|
276
369
|
outOfScope: 0,
|
|
370
|
+
ownership: known === 0 ? "wording" : known === lanes.length ? "routes" : "mixed",
|
|
371
|
+
setAsideByRoute: [],
|
|
277
372
|
contextual: 0,
|
|
373
|
+
worthALook: 0,
|
|
278
374
|
buckets: [],
|
|
279
375
|
ece: 0,
|
|
280
376
|
brier: 0,
|
|
@@ -284,8 +380,15 @@ export function calibrateAgainstKey(decisions, key) {
|
|
|
284
380
|
out.unsure += 1;
|
|
285
381
|
continue;
|
|
286
382
|
}
|
|
287
|
-
if (
|
|
383
|
+
if (d.verdict === "worth_a_look") {
|
|
384
|
+
out.worthALook += 1;
|
|
385
|
+
continue;
|
|
386
|
+
}
|
|
387
|
+
const remark = ownershipRemark(d, key, laneRoutes);
|
|
388
|
+
if (remark) {
|
|
288
389
|
out.outOfScope += 1;
|
|
390
|
+
if (remark.by === "routes")
|
|
391
|
+
out.setAsideByRoute.push({ lane: d.lane, id: remark.id, confidence: d.confidence });
|
|
289
392
|
continue;
|
|
290
393
|
}
|
|
291
394
|
const c = d.confidence;
|
|
@@ -334,14 +437,19 @@ export function calibrateAgainstKey(decisions, key) {
|
|
|
334
437
|
* across runs, and scoring an accumulated one credits a run with what an
|
|
335
438
|
* earlier run found. Point the scorer at a fresh project directory per run.
|
|
336
439
|
*/
|
|
337
|
-
export function score(key, findings, decisions, level = "medium") {
|
|
440
|
+
export function score(key, findings, decisions, level = "medium", laneRoutes) {
|
|
338
441
|
const hitsByDefect = new Map();
|
|
339
442
|
const falsePositives = [];
|
|
340
443
|
const unknown = [];
|
|
341
444
|
const ambiguous = [];
|
|
342
445
|
const contextual = [];
|
|
446
|
+
const worthALook = [];
|
|
343
447
|
let correct = 0;
|
|
344
448
|
for (const f of findings) {
|
|
449
|
+
if (isWorthALook(f)) {
|
|
450
|
+
worthALook.push({ title: f.title, convention: f.convention ?? "" });
|
|
451
|
+
continue;
|
|
452
|
+
}
|
|
345
453
|
const m = classify(findingText(f), key);
|
|
346
454
|
if (!m) {
|
|
347
455
|
unknown.push({ title: f.title, severity: f.severity, evidence: f.evidence ?? "" });
|
|
@@ -410,11 +518,12 @@ export function score(key, findings, decisions, level = "medium") {
|
|
|
410
518
|
correct,
|
|
411
519
|
falsePositives,
|
|
412
520
|
contextual,
|
|
521
|
+
worthALook,
|
|
413
522
|
unknown,
|
|
414
523
|
ambiguous,
|
|
415
524
|
duplicates,
|
|
416
525
|
severity,
|
|
417
|
-
calibration: calibrateAgainstKey(decisions, key),
|
|
526
|
+
calibration: calibrateAgainstKey(decisions, key, laneRoutes),
|
|
418
527
|
};
|
|
419
528
|
}
|
|
420
529
|
const pct = (n, d) => (d === 0 ? "โ" : `${Math.round((n / d) * 100)}%`);
|
|
@@ -424,13 +533,18 @@ const pct = (n, d) => (d === 0 ? "โ" : `${Math.round((n / d) * 100)}%`);
|
|
|
424
533
|
* false claims nobody has judged yet. The bounds can: the lower one counts
|
|
425
534
|
* every open finding as wrong, the upper one as right. Findings set aside as
|
|
426
535
|
* contextual are in neither: they are not claims the run could get right.
|
|
536
|
+
* Nor are findings the run filed as worth a look: it did not claim them.
|
|
427
537
|
*/
|
|
428
538
|
export function precisionBounds(c) {
|
|
429
539
|
const labelled = c.correct + c.falsePositives.length;
|
|
430
540
|
const open = c.unknown.length + c.ambiguous.length;
|
|
431
|
-
const scored = c.findings - c
|
|
541
|
+
const scored = c.findings - setAside(c);
|
|
432
542
|
return { labelled: `${c.correct}/${labelled} (${pct(c.correct, labelled)})`, low: pct(c.correct, scored), high: pct(c.correct + open, scored) };
|
|
433
543
|
}
|
|
544
|
+
/** Findings in neither precision count: contextual ones, and the run's own worth-a-looks. */
|
|
545
|
+
function setAside(c) {
|
|
546
|
+
return c.contextual.length + (c.worthALook?.length ?? 0);
|
|
547
|
+
}
|
|
434
548
|
/** The scorecard as a person reads it. Leads with the two numbers, then says what each is made of. */
|
|
435
549
|
export function formatScorecard(c) {
|
|
436
550
|
const p = precisionBounds(c);
|
|
@@ -441,11 +555,13 @@ export function formatScorecard(c) {
|
|
|
441
555
|
`Recall ${c.found.length}/${c.expected} (${pct(c.found.length, c.expected)}) of the planted defects expected at this level`,
|
|
442
556
|
`Precision ${p.labelled} of the findings the key can label` +
|
|
443
557
|
(open
|
|
444
|
-
? ` โ ${open} of ${c.findings - c
|
|
558
|
+
? ` โ ${open} of ${c.findings - setAside(c)} unlabelled, so between ${p.low} and ${p.high} of ${setAside(c) ? "the findings not set aside" : "all findings"}`
|
|
445
559
|
: ""),
|
|
446
560
|
];
|
|
447
561
|
if (c.contextual.length)
|
|
448
562
|
lines.push(`Set aside ${c.contextual.length} of ${c.findings} finding(s): defects only under a convention the run cannot see, so in neither count`);
|
|
563
|
+
if (c.worthALook.length)
|
|
564
|
+
lines.push(`Set aside ${c.worthALook.length} of ${c.findings} finding(s) the run filed as worth a look: not claimed as defects, so in neither recall nor precision`);
|
|
449
565
|
lines.push(``);
|
|
450
566
|
if (c.missed.length)
|
|
451
567
|
lines.push(`Missed: ${c.missed.join(", ")}`);
|
|
@@ -463,6 +579,11 @@ export function formatScorecard(c) {
|
|
|
463
579
|
for (const x of c.contextual)
|
|
464
580
|
lines.push(` ${x.title} (${x.id})`);
|
|
465
581
|
}
|
|
582
|
+
if (c.worthALook.length) {
|
|
583
|
+
lines.push(``, `Filed as worth a look (${c.worthALook.length}):`);
|
|
584
|
+
for (const x of c.worthALook)
|
|
585
|
+
lines.push(` ${x.title}${x.convention ? ` โ a defect only if the project uses ${x.convention}` : ""}`);
|
|
586
|
+
}
|
|
466
587
|
if (c.ambiguous.length) {
|
|
467
588
|
lines.push(``, `Ambiguous (${c.ambiguous.length}) โ the key claims each twice; sharpen it:`);
|
|
468
589
|
for (const a of c.ambiguous)
|
|
@@ -490,8 +611,10 @@ export function formatScorecard(c) {
|
|
|
490
611
|
k.ambiguous && `${k.ambiguous} the key names ambiguously`,
|
|
491
612
|
k.badConfidence && `${k.badConfidence} with an unusable confidence`,
|
|
492
613
|
k.unsure && `${k.unsure} unsure`,
|
|
493
|
-
k.outOfScope &&
|
|
614
|
+
k.outOfScope &&
|
|
615
|
+
`${k.outOfScope} dismissed as another lane's (${k.ownership === "routes" ? "by the lanes' routes" : k.ownership === "wording" ? "by wording: no lane routes archived" : "by the lanes' routes where archived, else by wording"})`,
|
|
494
616
|
k.contextual && `${k.contextual} about things that are defects only under a convention the run cannot see`,
|
|
617
|
+
k.worthALook && `${k.worthALook} the lane marked worth a look`,
|
|
495
618
|
].filter(Boolean);
|
|
496
619
|
lines.push(``, k.judged === 0
|
|
497
620
|
? `Lane calibration against the key: nothing the key could judge` + (skipped.length ? ` (${skipped.join(", ")})` : "")
|
|
@@ -499,6 +622,11 @@ export function formatScorecard(c) {
|
|
|
499
622
|
(skipped.length ? ` โ not scored: ${skipped.join(", ")}` : ""));
|
|
500
623
|
for (const b of k.buckets)
|
|
501
624
|
lines.push(` stated ${b.label}: ${b.decisions} decision(s), said ${b.stated.toFixed(2)}, right ${Math.round(b.correct * 100)}%`);
|
|
625
|
+
if (k.setAsideByRoute.length) {
|
|
626
|
+
lines.push(` Set aside as another lane's, by the lanes' routes (${k.setAsideByRoute.length}):`);
|
|
627
|
+
for (const r of k.setAsideByRoute)
|
|
628
|
+
lines.push(` ${r.lane} on ${r.id}, stated ${Number.isFinite(r.confidence) ? r.confidence.toFixed(2) : "?"}`);
|
|
629
|
+
}
|
|
502
630
|
}
|
|
503
631
|
return lines.join("\n");
|
|
504
632
|
}
|
|
@@ -543,17 +671,22 @@ const LOCAL_PATH_RE = /(?:\/Users\/[^/\s"']+|\/home\/[^/\s"']+|\/private\/tmp|\/
|
|
|
543
671
|
export function sanitize(text) {
|
|
544
672
|
return text.replace(LOCAL_PATH_RE, "<path>");
|
|
545
673
|
}
|
|
546
|
-
export function toArchive(run, date, note, findings, decisions, app) {
|
|
674
|
+
export function toArchive(run, date, note, findings, decisions, app, laneRoutes = {}) {
|
|
675
|
+
// No query string or fragment reaches an archive: a lane copies routes from
|
|
676
|
+
// the address bar, and an address can carry a token.
|
|
677
|
+
const routes = Object.fromEntries(Object.entries(laneRoutes).map(([lane, list]) => [lane, list.map((r) => sanitize(stripRouteQuery(r)))]));
|
|
547
678
|
return {
|
|
548
679
|
app,
|
|
549
680
|
run,
|
|
550
681
|
date,
|
|
551
682
|
note,
|
|
683
|
+
...(Object.keys(routes).length ? { laneRoutes: routes } : {}),
|
|
552
684
|
findings: findings.map((f) => ({
|
|
553
685
|
title: sanitize(f.title),
|
|
554
686
|
severity: f.severity,
|
|
555
687
|
...(f.category ? { category: f.category } : {}),
|
|
556
688
|
...(f.evidence ? { evidence: sanitize(f.evidence) } : {}),
|
|
689
|
+
...(isWorthALook(f) ? { tier: "worth_a_look", ...(f.convention ? { convention: sanitize(f.convention) } : {}) } : {}),
|
|
557
690
|
})),
|
|
558
691
|
decisions: decisions.map((d) => ({ ...d, observation: sanitize(d.observation), evidence: d.evidence === null ? null : sanitize(d.evidence) })),
|
|
559
692
|
};
|
|
@@ -21,7 +21,7 @@
|
|
|
21
21
|
* Pure, so every rule here is table-tested.
|
|
22
22
|
*/
|
|
23
23
|
import { requestsPair, statedRequests, templatedPathsMatch } from "./lane.js";
|
|
24
|
-
import { failingSignatures } from "./memory.js";
|
|
24
|
+
import { failingSignatures, isWorthALook } from "./memory.js";
|
|
25
25
|
/**
|
|
26
26
|
* Upper edge of each confidence bucket. Five is enough to see a shape and few
|
|
27
27
|
* enough that each holds a usable count on a run of a few dozen decisions;
|
|
@@ -87,7 +87,8 @@ function usableConfidence(value) {
|
|
|
87
87
|
*/
|
|
88
88
|
export function calibrate(decisions, findings) {
|
|
89
89
|
const filedKeys = new Map();
|
|
90
|
-
|
|
90
|
+
// Filed as worth a look is not filed as a defect: a lane that called it a defect was not agreed with.
|
|
91
|
+
for (const f of findings.filter((x) => !isWorthALook(x))) {
|
|
91
92
|
if (!f.evidence)
|
|
92
93
|
continue;
|
|
93
94
|
// A finding verified as still present is the most informative match, so it
|
|
@@ -102,8 +103,11 @@ export function calibrate(decisions, findings) {
|
|
|
102
103
|
const claims = decisions.filter((d) => d.verdict === "defect" && d.evidence !== null);
|
|
103
104
|
const checkable = claims.filter((d) => usableConfidence(d.confidence) !== null && joinKeys(d.evidence).size > 0);
|
|
104
105
|
const unjoinable = claims.length - checkable.length;
|
|
106
|
+
const worthALook = decisions.filter((d) => d.verdict === "worth_a_look").length;
|
|
105
107
|
if (checkable.length === 0)
|
|
106
|
-
return unjoinable > 0
|
|
108
|
+
return unjoinable > 0 || worthALook > 0
|
|
109
|
+
? { checkable: 0, filed: 0, stated: 0, buckets: [], ece: 0, verified: { present: 0, gone: 0, changed: 0 }, unjoinable, worthALook }
|
|
110
|
+
: null;
|
|
107
111
|
const buckets = BUCKET_EDGES.map(() => ({ n: 0, conf: 0, hits: 0 }));
|
|
108
112
|
const verified = { present: 0, gone: 0, changed: 0 };
|
|
109
113
|
const matched = new Map();
|
|
@@ -147,7 +151,7 @@ export function calibrate(decisions, findings) {
|
|
|
147
151
|
ece += (b.n / n) * Math.abs(meanConf - rate);
|
|
148
152
|
out.push({ label: bucketLabel(i), decisions: b.n, stated: meanConf, filed: rate });
|
|
149
153
|
});
|
|
150
|
-
return { checkable: n, filed, stated: stated / n, buckets: out, ece, verified, unjoinable };
|
|
154
|
+
return { checkable: n, filed, stated: stated / n, buckets: out, ece, verified, unjoinable, worthALook };
|
|
151
155
|
}
|
|
152
156
|
const pct = (x) => `${Math.round(x * 100)}%`;
|
|
153
157
|
/**
|
|
@@ -163,7 +167,7 @@ export function formatCalibration(c) {
|
|
|
163
167
|
// and a run that recorded none look identical when the section simply
|
|
164
168
|
// vanishes, and the reader concludes the feature is broken โ which is this
|
|
165
169
|
// project's own definition of a silent path.
|
|
166
|
-
const had = c.checkable + c.unjoinable;
|
|
170
|
+
const had = c.checkable + c.unjoinable + c.worthALook;
|
|
167
171
|
if (had === 0)
|
|
168
172
|
return [];
|
|
169
173
|
return [
|
|
@@ -171,6 +175,7 @@ export function formatCalibration(c) {
|
|
|
171
175
|
``,
|
|
172
176
|
`Not enough to say yet: ${c.checkable} lane decision(s) could be checked${c.unjoinable > 0 ? ` (and ${c.unjoinable} could not be looked up at all)` : ""}, ` +
|
|
173
177
|
`and ${MIN_FOR_A_VERDICT} are needed before a calibration figure survives one of them changing.`,
|
|
178
|
+
...worthALookNote(c),
|
|
174
179
|
``,
|
|
175
180
|
];
|
|
176
181
|
}
|
|
@@ -194,6 +199,7 @@ export function formatCalibration(c) {
|
|
|
194
199
|
if (c.unjoinable > 0) {
|
|
195
200
|
lines.push(``, `${c.unjoinable} further decision(s) called a defect but named no failing endpoint, so nothing could be looked up for them. They are excluded above rather than counted as wrong.`);
|
|
196
201
|
}
|
|
202
|
+
lines.push(...worthALookNote(c));
|
|
197
203
|
const seen = c.verified.present + c.verified.gone + c.verified.changed;
|
|
198
204
|
if (seen > 0) {
|
|
199
205
|
lines.push(``, `${seen} of the findings those decisions matched ${seen === 1 ? "has" : "have"} since been re-tested with \`scout_verify\`: ${c.verified.present} still present, ${c.verified.gone} gone, ${c.verified.changed} changed. ` +
|
|
@@ -202,6 +208,16 @@ export function formatCalibration(c) {
|
|
|
202
208
|
lines.push(``);
|
|
203
209
|
return lines;
|
|
204
210
|
}
|
|
211
|
+
/** The line saying how many "worth a look" decisions were left out, and why; nothing when there were none. */
|
|
212
|
+
function worthALookNote(c) {
|
|
213
|
+
if (c.worthALook === 0)
|
|
214
|
+
return [];
|
|
215
|
+
return [
|
|
216
|
+
``,
|
|
217
|
+
`${c.worthALook} decision(s) were marked worth a look: real, and a defect only under a convention of the project the run cannot see. ` +
|
|
218
|
+
`Whether one was filed says nothing about whether the lane was right, so they are not scored.`,
|
|
219
|
+
];
|
|
220
|
+
}
|
|
205
221
|
/** Below this, the number swings on a single decision and is worse than no number. */
|
|
206
222
|
export const MIN_FOR_A_VERDICT = 8;
|
|
207
223
|
/**
|
|
@@ -259,7 +275,8 @@ export const MAX_UNFILED_NAMED = 10;
|
|
|
259
275
|
export function unfiledDefects(decisions, findings) {
|
|
260
276
|
const keys = [];
|
|
261
277
|
const filed = [];
|
|
262
|
-
|
|
278
|
+
// A defect filed only as worth a look is not in the report's findings, so it is still unfiled.
|
|
279
|
+
for (const f of findings.filter((x) => !isWorthALook(x))) {
|
|
263
280
|
const text = `${f.evidence ?? ""} ${f.title}`;
|
|
264
281
|
const requests = statedRequests(f.evidence ?? "");
|
|
265
282
|
// The store's signatures, and the finding's own failing requests as
|