scenescout 3.11.0 → 3.12.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +22 -0
- package/README.md +5 -1
- package/dist/browsers.js +52 -8
- package/dist/check-run.js +4 -2
- package/dist/engine/bench.js +147 -14
- package/dist/engine/browser.js +174 -7
- package/dist/engine/calibration.js +122 -20
- package/dist/engine/check.js +86 -14
- package/dist/engine/design.js +15 -1
- package/dist/engine/fingerprint.js +69 -7
- package/dist/engine/lane.js +138 -3
- package/dist/engine/memory.js +135 -10
- package/dist/engine/report.js +37 -2
- package/dist/engine/unload.js +201 -0
- package/dist/engine/verify.js +4 -3
- package/dist/mcp-server.js +23 -9
- package/package.json +1 -1
- package/skills/scenescout/SKILL.md +4 -3
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,27 @@
|
|
|
1
1
|
# scenescout
|
|
2
2
|
|
|
3
|
+
## 3.12.0
|
|
4
|
+
|
|
5
|
+
### Minor Changes
|
|
6
|
+
|
|
7
|
+
- 967e98c: Add a "worth a look" tier for observations that are defects only under a convention of the project the run cannot see, such as a spacing scale, link styling in navigation, or test ids on every control. SceneScout reports each one with the convention that would decide it, and never counts it as a defect.
|
|
8
|
+
|
|
9
|
+
- Lane reports accept the verdict `worth_a_look`, which must name its `convention`. A `convention` on any other verdict is ignored, and the fold says so. Lane calibration and the benchmark's key calibration leave it unscored and count it under its own reason.
|
|
10
|
+
- `scout_finding` takes an optional `convention`, which files the finding in this tier. The report lists these findings under "Worth a look", below the findings, as "a defect only if your project uses …", and leaves them out of every defect total. Filing the same thing again as a defect promotes it, at the defect's severity, and the reply says so. The benchmark sets such findings aside from recall and precision, and the fold lists a lane's judged defect as unfiled when it was filed only as worth a look.
|
|
11
|
+
- `scenescout check` reports two new rules, `off-grid-spacing` and `indistinct-link`, in this tier. They have no severity and never fail the gate at any `--fail-on`. SARIF reports them at level `note`. The report and the job summary list them in their own section, `check.json` puts them under `worthALook`, separate from `issues`, and the GitHub Action publishes their number as the `worth-a-look` output. Every existing rule is unchanged.
|
|
12
|
+
|
|
13
|
+
### Patch Changes
|
|
14
|
+
|
|
15
|
+
- c40892f: `scout_lane_report` keeps the routes each lane's accepted report lists in the project's memory (`laneRoutes` in `.scenescout/memory.json`), with query strings and fragments removed so no token in an address is stored, and `npm run bench -- --archive` carries them into the run archive. The benchmark then decides whether a lane's "not a defect" is a remark about another lane's page by where the matched defect is, not by how the verdict is worded: a defect none of whose pages the lane covered is set aside as another lane's and listed on the scorecard, and one on the lane's own page, or one every lane can reach, is scored. Answer-key entries can name further pages with `alsoOn` and mark a defect every lane can reach with `everyPage`. Archives made before routes were kept, and lanes with a route that names no page, are scored by wording, as before.
|
|
16
|
+
|
|
17
|
+
## 3.11.1
|
|
18
|
+
|
|
19
|
+
### Patch Changes
|
|
20
|
+
|
|
21
|
+
- 847c149: Two different defects on one element now stay two findings. Finding dedup compares what is wrong as well as where: identical evidence, a shared quoted label or a reworded title merges two findings only when they are one kind of defect (the same category, or neighbouring categories of one family such as `page-error` and `console-error`). `visual`, `ux-polish`, `a11y` and `missing-testid` are each their own kind, so a link clipped out of view and the same link styled like body text are no longer merged. When `scout_finding` merges a filing it now names the category of the finding it joined and which categories the filing could have merged with.
|
|
22
|
+
- aca7a28: When a lane report is folded, a judged defect on a failing request now counts as filed when a finding names that request with its path written as a template: `GET /api/things/{id} 500` (or `:id`, or `*`) covers a lane's `GET /api/things/7 500`. A template stands only for one id segment (a number, a UUID or a long hex id), never for a word such as `me` or `export`. A defect naming several failing requests counts as filed only when every one of them is. On a request that did not fail, the lane's evidence, once the paths are aligned, must be the finding's, a restatement of part of it, or the finding's evidence followed only by the status the call should have returned, as in `… 200 as role=viewer (expected 403)`.
|
|
23
|
+
- 2b8ac9a: The write policy now judges writes a page sends as it is being left. In Chromium, a `navigator.sendBeacon` or `fetch(..., { keepalive: true })` sent on `pagehide`, `visibilitychange` or `unload` was never intercepted and reached the server in every mode, observe included; the engine now catches it at the browser level and applies the same rules, so under observe (and under read-only, for a destructive write) it is refused and reported with the other refused writes, and a mode that allows it still sends it. Such a write is recognised by its headers, since the page that sent it is gone: one out of the app whose Origin is another site's, or `null`, is treated as an embed's and refused outside destructive, and the exception for a captcha on the app's sign-in page is decided on the Referer, the page it was sent from. Browsers send only the origin as the Referer of a request to another site by default, so that exception seldom applies and such a captcha write is refused; a sign-in the tester completes while the page is open is not affected. In Chromium, a write a 307 or 308 redirect carries on to a new address is now judged there as well. In every browser, the pages the engine closes (at the end of a session, on a re-attach, and a popup of another site) are left for `about:blank` first, so what they send on the way out meets the policy too: in Firefox and WebKit such a write had also gone out unseen.
|
|
24
|
+
|
|
3
25
|
## 3.11.0
|
|
4
26
|
|
|
5
27
|
### Minor Changes
|
package/README.md
CHANGED
|
@@ -332,6 +332,7 @@ A `🛡 WRITE-POLICY blocked` notice is the safety net doing its job, not an app
|
|
|
332
332
|
`.scenescout/report.md` — a deduplicated, worst-first report with:
|
|
333
333
|
|
|
334
334
|
- 🐛 **Findings** with repro traces and generated Playwright regression-test skeletons.
|
|
335
|
+
- 🔎 **Worth a look** — observations that are defects only under a convention of your project the run cannot see (a spacing scale, link styling in navigation, test ids on every control), each naming that convention. Listed below the findings and not counted as defects ([ADR 13](docs/adr/0013-a-convention-is-the-projects-to-decide.md)).
|
|
335
336
|
- 💯 **Page scores** (0–100: a11y · craft · consistency · task-clarity), ranked worst-first, with stale scores from old runs marked as such.
|
|
336
337
|
- 👥 **A role capability matrix** — what each role could and couldn't reach.
|
|
337
338
|
- 🧾 **A gap ledger** — everything *not* done, so the report is honest about its own coverage.
|
|
@@ -355,6 +356,8 @@ An exploratory run is driven by a model, so two runs never find exactly the same
|
|
|
355
356
|
- contrast and focus
|
|
356
357
|
- pages with no way out
|
|
357
358
|
|
|
359
|
+
Some of what it measures is a defect only under a convention the check cannot see: paddings off a 4px grid, and links styled like body text. Those are listed under **Worth a look**, each with the convention that would make it a defect. They are never counted and never fail the gate, at any `--fail-on`; SARIF reports them at level `note` ([ADR 13](docs/adr/0013-a-convention-is-the-projects-to-decide.md)).
|
|
360
|
+
|
|
358
361
|
```bash
|
|
359
362
|
npx scenescout check http://127.0.0.1:3000 --fail-on high
|
|
360
363
|
```
|
|
@@ -380,7 +383,7 @@ On GitHub Actions, this repository is also an action that installs everything an
|
|
|
380
383
|
With the default settings its saved flows send no HTTP write (they replay under observe's rule), and its crawl runs under `--mode observe` or `read-only`; `--flow-writes allow` lets flows write as `--mode` allows. By default it fails only on facts that mean a page is broken: a page that did not load, an uncaught exception, a 5xx, a failure shown as success. Other options:
|
|
381
384
|
|
|
382
385
|
- `--fail-on medium` or `low` makes the gate stricter.
|
|
383
|
-
- `--ignore <rule>` drops a rule.
|
|
386
|
+
- `--ignore <rule>` drops a rule, worth-a-look rules included.
|
|
384
387
|
- `--paths /a,/b` checks only those pages.
|
|
385
388
|
- `--storage-state <file>` checks while signed in.
|
|
386
389
|
- `--flows <dir>` or `off` chooses which saved flows to replay; `--retest off` skips re-testing open findings.
|
|
@@ -661,6 +664,7 @@ The load-bearing choices are recorded as ADRs — read the relevant one before c
|
|
|
661
664
|
- [10 · A lane's confidence is checked, not trusted](docs/adr/0010-a-confidence-is-checked-not-trusted.md)
|
|
662
665
|
- [11 · A gate is deterministic, and fails only on what it can prove](docs/adr/0011-a-gate-is-deterministic-and-fails-only-on-what-it-can-prove.md)
|
|
663
666
|
- [12 · A check replays saved flows and re-tests open findings, within settings whose defaults do the least harm](docs/adr/0012-a-check-replays-saved-flows-and-reports-re-tests.md)
|
|
667
|
+
- [13 · What depends on a project's convention is the project's to decide](docs/adr/0013-a-convention-is-the-projects-to-decide.md)
|
|
664
668
|
|
|
665
669
|
---
|
|
666
670
|
|
package/dist/browsers.js
CHANGED
|
@@ -182,14 +182,58 @@ export function beaconResourceType(engine) {
|
|
|
182
182
|
return engine === "chromium" ? "ping" : "beacon";
|
|
183
183
|
}
|
|
184
184
|
/**
|
|
185
|
-
*
|
|
186
|
-
* keepalive fetch
|
|
187
|
-
*
|
|
188
|
-
*
|
|
189
|
-
*
|
|
190
|
-
*
|
|
191
|
-
|
|
192
|
-
|
|
185
|
+
* Where the write policy meets a write a page sends as it is being left (a
|
|
186
|
+
* sendBeacon or keepalive fetch on `pagehide`, `visibilitychange` or
|
|
187
|
+
* `unload`, or one still leaving as a page is closed).
|
|
188
|
+
*
|
|
189
|
+
* - `route`: Firefox and WebKit hand it to the context's route handler like
|
|
190
|
+
* any other request.
|
|
191
|
+
* - `browser-fetch`: Chromium never routes it (it arrives with no network id,
|
|
192
|
+
* or after the page's session has gone, and the driver sends it on), so on
|
|
193
|
+
* its own it reached the server in every mode. The engine also pauses
|
|
194
|
+
* requests at the browser level with the DevTools Fetch domain and judges
|
|
195
|
+
* there, by the same rules, every write the route handler did not
|
|
196
|
+
* (unload.ts). Only in modes where the policy can refuse: destructive
|
|
197
|
+
* installs neither.
|
|
198
|
+
*
|
|
199
|
+
* Either way the write is judged and a refusal is reported like any other
|
|
200
|
+
* (ADR 2); the unload smoke suite asserts it per engine.
|
|
201
|
+
*/
|
|
202
|
+
export function unloadWriteInterception(engine) {
|
|
203
|
+
return engine === "chromium" ? "browser-fetch" : "route";
|
|
204
|
+
}
|
|
205
|
+
/**
|
|
206
|
+
* Whether a write a page sends as it is being left may be lost after the
|
|
207
|
+
* write policy lets it through. WebKit cancels such a request when the route
|
|
208
|
+
* handler answers after the page has gone, so a mode that allows the write
|
|
209
|
+
* does not promise it arrives; nothing is sent that the policy refused. Firefox
|
|
210
|
+
* sends it on, and so does Chromium's browser-level interception. In
|
|
211
|
+
* destructive mode nothing is intercepted and every engine sends it.
|
|
212
|
+
*/
|
|
213
|
+
export function allowedUnloadWritesMayBeLost(engine) {
|
|
214
|
+
return engine === "webkit";
|
|
215
|
+
}
|
|
216
|
+
/**
|
|
217
|
+
* Whether the writes another site's frame sends on `pagehide` may never be
|
|
218
|
+
* issued when the top page is left. Firefox can tear such a frame down before
|
|
219
|
+
* its requests reach the route handler, so there is nothing to refuse or to
|
|
220
|
+
* report, and nothing is sent. It depends on timing: seen in CI and under load
|
|
221
|
+
* locally, not on an idle machine. Chromium's browser-level interception and
|
|
222
|
+
* WebKit's route handler met them every time in the same runs. The writes the
|
|
223
|
+
* top page itself sends are not affected.
|
|
224
|
+
*/
|
|
225
|
+
export function frameUnloadWritesMayGoUnissued(engine) {
|
|
226
|
+
return engine === "firefox";
|
|
227
|
+
}
|
|
228
|
+
/**
|
|
229
|
+
* Whether a write carried on by a redirect (a 307 or 308 keeps the method and
|
|
230
|
+
* the body) is judged at its new address. The route handler sees only the
|
|
231
|
+
* first request of a redirect in every engine. In Chromium the browser-level
|
|
232
|
+
* interception (`unloadWriteInterception`) sees each hop and judges it like any
|
|
233
|
+
* write; Firefox and WebKit send the later hops on unjudged, a known limit of
|
|
234
|
+
* the policy there (ADR 2).
|
|
235
|
+
*/
|
|
236
|
+
export function writeRedirectHopsJudged(engine) {
|
|
193
237
|
return engine === "chromium";
|
|
194
238
|
}
|
|
195
239
|
/**
|
package/dist/check-run.js
CHANGED
|
@@ -9,7 +9,7 @@ import fs from "node:fs";
|
|
|
9
9
|
import os from "node:os";
|
|
10
10
|
import path from "node:path";
|
|
11
11
|
import { BrowserEngine } from "./engine/browser.js";
|
|
12
|
-
import {
|
|
12
|
+
import { checkFindings, redactFlowRuns, redactRoute, redactRoutes, settingsOf, withoutOwnResponse, } from "./engine/check.js";
|
|
13
13
|
import { loadFlows, resolveFlowsDir } from "./engine/flow.js";
|
|
14
14
|
import { MemoryStore, MEMORY_DIRNAME } from "./engine/memory.js";
|
|
15
15
|
import { checkRetestPlan, retestResults, wellFormedFindings } from "./engine/verify.js";
|
|
@@ -123,13 +123,15 @@ export async function runCheck(options, log = () => { }, inputs = { flows: [], f
|
|
|
123
123
|
: null;
|
|
124
124
|
const measured = redactRoutes(routes.map(withoutOwnResponse));
|
|
125
125
|
const flows = redactFlowRuns(flowRuns);
|
|
126
|
+
const { issues, worthALook } = checkFindings(measured, start.origin, options.ignore, flows);
|
|
126
127
|
return {
|
|
127
128
|
url: redactRoute(options.url),
|
|
128
129
|
generatedAt: new Date().toISOString(),
|
|
129
130
|
mode: options.mode,
|
|
130
131
|
failOn: options.failOn,
|
|
131
132
|
routes: measured,
|
|
132
|
-
issues
|
|
133
|
+
issues,
|
|
134
|
+
worthALook,
|
|
133
135
|
// Routes that failed to load are issues already; "not visited" is only what --max-routes left out.
|
|
134
136
|
unvisited: options.paths ? [] : engine.crawlableRoutes().map(redactRoute),
|
|
135
137
|
ignored: options.ignore,
|
package/dist/engine/bench.js
CHANGED
|
@@ -31,6 +31,9 @@
|
|
|
31
31
|
import { createHash } from "node:crypto";
|
|
32
32
|
import { z } from "zod";
|
|
33
33
|
import { BUCKET_EDGES, bucketLabel, bucketOf } from "./calibration.js";
|
|
34
|
+
import { laneRoutePaths, stripRouteQuery } from "./fingerprint.js";
|
|
35
|
+
import { isWorthALook } from "./memory.js";
|
|
36
|
+
export { laneRoutePaths };
|
|
34
37
|
export const LEVELS = ["minimal", "medium", "extensive"];
|
|
35
38
|
const SEVERITIES = ["high", "medium", "low"];
|
|
36
39
|
const Pattern = z.string().refine((p) => {
|
|
@@ -42,10 +45,24 @@ const Pattern = z.string().refine((p) => {
|
|
|
42
45
|
return false;
|
|
43
46
|
}
|
|
44
47
|
}, { message: "is not a valid regular expression" });
|
|
48
|
+
/**
|
|
49
|
+
* Other pages the same defect shows on, beyond `route`. A lane that owns any
|
|
50
|
+
* of them owns the defect, so its "not a defect" is a verdict, not a remark
|
|
51
|
+
* about another lane's page.
|
|
52
|
+
*/
|
|
53
|
+
const AlsoOn = z.array(z.string().startsWith("/")).min(1).optional();
|
|
54
|
+
/**
|
|
55
|
+
* Every lane can reach the defect: it is in chrome every page carries (a
|
|
56
|
+
* shared nav, header or footer), or in an endpoint any lane can call. No lane
|
|
57
|
+
* can call it another lane's, so a verdict on it is always scored.
|
|
58
|
+
*/
|
|
59
|
+
const EveryPage = z.boolean().optional();
|
|
45
60
|
const Entry = z
|
|
46
61
|
.object({
|
|
47
62
|
id: z.string().min(1),
|
|
48
63
|
route: z.string().startsWith("/"),
|
|
64
|
+
alsoOn: AlsoOn,
|
|
65
|
+
everyPage: EveryPage,
|
|
49
66
|
title: z.string().min(1),
|
|
50
67
|
/** The category a finding for it should carry. Reported, not enforced: two categories can both be defensible. */
|
|
51
68
|
category: z.string().min(1),
|
|
@@ -83,6 +100,8 @@ const Contextual = z
|
|
|
83
100
|
.object({
|
|
84
101
|
id: z.string().min(1),
|
|
85
102
|
route: z.string().startsWith("/"),
|
|
103
|
+
alsoOn: AlsoOn,
|
|
104
|
+
everyPage: EveryPage,
|
|
86
105
|
title: z.string().min(1),
|
|
87
106
|
category: z.string().min(1),
|
|
88
107
|
/** The convention that decides it: where it would be a defect, and why the run cannot tell whether it holds here. */
|
|
@@ -130,6 +149,14 @@ export function parseKey(raw) {
|
|
|
130
149
|
const dup = ids.find((id, i) => ids.indexOf(id) !== i);
|
|
131
150
|
if (dup)
|
|
132
151
|
throw new Error(`The answer key uses the id ${JSON.stringify(dup)} twice.`);
|
|
152
|
+
for (const e of [...parsed.data.defects, ...parsed.data.alsoReal, ...parsed.data.contextual]) {
|
|
153
|
+
if (e.alsoOn && e.everyPage)
|
|
154
|
+
throw new Error(`The entry ${JSON.stringify(e.id)} has both alsoOn and everyPage; an entry on every page names no other pages.`);
|
|
155
|
+
const pages = [e.route, ...(e.alsoOn ?? [])];
|
|
156
|
+
const dup = pages.find((p, i) => pages.indexOf(p) !== i);
|
|
157
|
+
if (dup)
|
|
158
|
+
throw new Error(`The entry ${JSON.stringify(e.id)} names the page ${JSON.stringify(dup)} twice in route and alsoOn.`);
|
|
159
|
+
}
|
|
133
160
|
const real = new Set([...parsed.data.defects, ...parsed.data.alsoReal].map((e) => e.id));
|
|
134
161
|
for (const nd of parsed.data.nonDefects) {
|
|
135
162
|
const unknown = nd.overrides.find((id) => !real.has(id));
|
|
@@ -241,14 +268,14 @@ export function lintKey(key) {
|
|
|
241
268
|
*
|
|
242
269
|
* "defect" on a planted or also-real defect is right, and on a known
|
|
243
270
|
* non-defect is wrong. "not_a_defect" is the reverse. An "unsure" verdict, a
|
|
244
|
-
* dismissal as another lane's, a decision the key does not
|
|
245
|
-
* names ambiguously, and one about a contextual entry are not scored — the same rule the in-product calibration
|
|
271
|
+
* "worth_a_look", a dismissal as another lane's, a decision the key does not
|
|
272
|
+
* name, one it names ambiguously, and one about a contextual entry are not scored — the same rule the in-product calibration
|
|
246
273
|
* follows, for the same reason.
|
|
247
274
|
*/
|
|
248
|
-
export function judgeDecision(d, key) {
|
|
249
|
-
if (d.verdict === "unsure")
|
|
275
|
+
export function judgeDecision(d, key, laneRoutes) {
|
|
276
|
+
if (d.verdict === "unsure" || d.verdict === "worth_a_look")
|
|
250
277
|
return null;
|
|
251
|
-
if (
|
|
278
|
+
if (isOwnershipRemark(d, key, laneRoutes))
|
|
252
279
|
return null;
|
|
253
280
|
const m = classify(decisionText(d), key);
|
|
254
281
|
if (!m || m.kind === "ambiguous" || m.kind === "contextual")
|
|
@@ -258,13 +285,79 @@ export function judgeDecision(d, key) {
|
|
|
258
285
|
}
|
|
259
286
|
/** How a lane says "not mine": the wording lanes actually used, none of it about a particular app. */
|
|
260
287
|
const SCOPE_DISMISSAL_RE = /out.of.lane.scope|out-of-scope|out of (my|this) (lane|scope)|belongs?.to.{0,20}\b(lane|route|page)\b|belongs-to-|(handled|owned) by (the |another )?[\w-]* ?lane|lane owns it|not (in |on |from |part of )?(my|this) (lane|assigned routes|routes?|pages?)\b|outside (my|this) (lane|routes?|pages?)|not part of my (assigned )?routes/i;
|
|
261
|
-
/**
|
|
288
|
+
/**
|
|
289
|
+
* A not-a-defect verdict whose stated reason is that the thing is another
|
|
290
|
+
* lane's. The WORDING rule: all a run archived before lane routes were kept
|
|
291
|
+
* can be judged by, and still what decides for a lane whose routes are not
|
|
292
|
+
* known. It cannot tell a lane's own page from another's, which is why it is
|
|
293
|
+
* not widened (see `isOwnershipRemark`).
|
|
294
|
+
*/
|
|
262
295
|
export function isScopeDismissal(d) {
|
|
263
296
|
return d.verdict === "not_a_defect" && SCOPE_DISMISSAL_RE.test(decisionText(d));
|
|
264
297
|
}
|
|
265
|
-
|
|
298
|
+
/** The pages a key entry is on, in the same form as `laneRoutePaths`. */
|
|
299
|
+
function entryPaths(e) {
|
|
300
|
+
return [e.route, ...(e.alsoOn ?? [])].flatMap(laneRoutePaths);
|
|
301
|
+
}
|
|
302
|
+
/**
|
|
303
|
+
* A lane's routes as a set of paths, or undefined when they are not known
|
|
304
|
+
* well enough to decide by: none were recorded, or ANY of them names no path
|
|
305
|
+
* ("the orders area"). A route that could not be read may be the very page a
|
|
306
|
+
* verdict is about, and deciding without it would set that verdict aside.
|
|
307
|
+
*/
|
|
308
|
+
function lanePaths(lane, laneRoutes) {
|
|
309
|
+
const routes = laneRoutes && Object.hasOwn(laneRoutes, lane) ? laneRoutes[lane] : undefined;
|
|
310
|
+
if (!routes || routes.length === 0)
|
|
311
|
+
return undefined;
|
|
312
|
+
const paths = new Set();
|
|
313
|
+
for (const r of routes) {
|
|
314
|
+
const named = laneRoutePaths(r);
|
|
315
|
+
if (named.length === 0)
|
|
316
|
+
return undefined;
|
|
317
|
+
for (const p of named)
|
|
318
|
+
paths.add(p);
|
|
319
|
+
}
|
|
320
|
+
return paths;
|
|
321
|
+
}
|
|
322
|
+
/**
|
|
323
|
+
* A not-a-defect that is a remark about who owns the thing, not a verdict on
|
|
324
|
+
* whether it is broken.
|
|
325
|
+
*
|
|
326
|
+
* Where the lane's routes are known, decided by WHERE the thing is: the
|
|
327
|
+
* verdict matches a planted or also-real defect whose pages are all outside
|
|
328
|
+
* the lane's routes. The wording no longer matters either way — "raised on /
|
|
329
|
+
* before navigating" from the orders lane is a remark about the dashboard's
|
|
330
|
+
* defect, and "handled by another lane" from the lane that owns the page is
|
|
331
|
+
* its own wrong verdict. A defect in chrome every page carries (`everyPage`)
|
|
332
|
+
* is every lane's, so it is never another lane's. A verdict that matches a
|
|
333
|
+
* known non-defect, a contextual entry, several entries or none is not about
|
|
334
|
+
* a defect with a page to own, and is judged as it would be otherwise.
|
|
335
|
+
*
|
|
336
|
+
* Where they are not known — every archive made before routes were kept, and
|
|
337
|
+
* any lane that reported no route — the wording rule decides, unchanged, so
|
|
338
|
+
* those runs score as they always have.
|
|
339
|
+
*/
|
|
340
|
+
export function ownershipRemark(d, key, laneRoutes) {
|
|
341
|
+
if (d.verdict !== "not_a_defect")
|
|
342
|
+
return null;
|
|
343
|
+
const own = lanePaths(d.lane, laneRoutes);
|
|
344
|
+
if (!own)
|
|
345
|
+
return isScopeDismissal(d) ? { by: "wording" } : null;
|
|
346
|
+
const m = classify(decisionText(d), key);
|
|
347
|
+
if (!m || (m.kind !== "defect" && m.kind !== "alsoReal"))
|
|
348
|
+
return null;
|
|
349
|
+
if (m.entry.everyPage)
|
|
350
|
+
return null;
|
|
351
|
+
return entryPaths(m.entry).some((p) => own.has(p)) ? null : { by: "routes", id: m.entry.id };
|
|
352
|
+
}
|
|
353
|
+
export function isOwnershipRemark(d, key, laneRoutes) {
|
|
354
|
+
return ownershipRemark(d, key, laneRoutes) !== null;
|
|
355
|
+
}
|
|
356
|
+
export function calibrateAgainstKey(decisions, key, laneRoutes) {
|
|
266
357
|
if (decisions.length === 0)
|
|
267
358
|
return null;
|
|
359
|
+
const lanes = [...new Set(decisions.map((d) => d.lane))];
|
|
360
|
+
const known = lanes.filter((l) => lanePaths(l, laneRoutes) !== undefined).length;
|
|
268
361
|
const buckets = BUCKET_EDGES.map(() => ({ n: 0, conf: 0, right: 0 }));
|
|
269
362
|
const out = {
|
|
270
363
|
judged: 0,
|
|
@@ -274,7 +367,10 @@ export function calibrateAgainstKey(decisions, key) {
|
|
|
274
367
|
badConfidence: 0,
|
|
275
368
|
unsure: 0,
|
|
276
369
|
outOfScope: 0,
|
|
370
|
+
ownership: known === 0 ? "wording" : known === lanes.length ? "routes" : "mixed",
|
|
371
|
+
setAsideByRoute: [],
|
|
277
372
|
contextual: 0,
|
|
373
|
+
worthALook: 0,
|
|
278
374
|
buckets: [],
|
|
279
375
|
ece: 0,
|
|
280
376
|
brier: 0,
|
|
@@ -284,8 +380,15 @@ export function calibrateAgainstKey(decisions, key) {
|
|
|
284
380
|
out.unsure += 1;
|
|
285
381
|
continue;
|
|
286
382
|
}
|
|
287
|
-
if (
|
|
383
|
+
if (d.verdict === "worth_a_look") {
|
|
384
|
+
out.worthALook += 1;
|
|
385
|
+
continue;
|
|
386
|
+
}
|
|
387
|
+
const remark = ownershipRemark(d, key, laneRoutes);
|
|
388
|
+
if (remark) {
|
|
288
389
|
out.outOfScope += 1;
|
|
390
|
+
if (remark.by === "routes")
|
|
391
|
+
out.setAsideByRoute.push({ lane: d.lane, id: remark.id, confidence: d.confidence });
|
|
289
392
|
continue;
|
|
290
393
|
}
|
|
291
394
|
const c = d.confidence;
|
|
@@ -334,14 +437,19 @@ export function calibrateAgainstKey(decisions, key) {
|
|
|
334
437
|
* across runs, and scoring an accumulated one credits a run with what an
|
|
335
438
|
* earlier run found. Point the scorer at a fresh project directory per run.
|
|
336
439
|
*/
|
|
337
|
-
export function score(key, findings, decisions, level = "medium") {
|
|
440
|
+
export function score(key, findings, decisions, level = "medium", laneRoutes) {
|
|
338
441
|
const hitsByDefect = new Map();
|
|
339
442
|
const falsePositives = [];
|
|
340
443
|
const unknown = [];
|
|
341
444
|
const ambiguous = [];
|
|
342
445
|
const contextual = [];
|
|
446
|
+
const worthALook = [];
|
|
343
447
|
let correct = 0;
|
|
344
448
|
for (const f of findings) {
|
|
449
|
+
if (isWorthALook(f)) {
|
|
450
|
+
worthALook.push({ title: f.title, convention: f.convention ?? "" });
|
|
451
|
+
continue;
|
|
452
|
+
}
|
|
345
453
|
const m = classify(findingText(f), key);
|
|
346
454
|
if (!m) {
|
|
347
455
|
unknown.push({ title: f.title, severity: f.severity, evidence: f.evidence ?? "" });
|
|
@@ -410,11 +518,12 @@ export function score(key, findings, decisions, level = "medium") {
|
|
|
410
518
|
correct,
|
|
411
519
|
falsePositives,
|
|
412
520
|
contextual,
|
|
521
|
+
worthALook,
|
|
413
522
|
unknown,
|
|
414
523
|
ambiguous,
|
|
415
524
|
duplicates,
|
|
416
525
|
severity,
|
|
417
|
-
calibration: calibrateAgainstKey(decisions, key),
|
|
526
|
+
calibration: calibrateAgainstKey(decisions, key, laneRoutes),
|
|
418
527
|
};
|
|
419
528
|
}
|
|
420
529
|
const pct = (n, d) => (d === 0 ? "—" : `${Math.round((n / d) * 100)}%`);
|
|
@@ -424,13 +533,18 @@ const pct = (n, d) => (d === 0 ? "—" : `${Math.round((n / d) * 100)}%`);
|
|
|
424
533
|
* false claims nobody has judged yet. The bounds can: the lower one counts
|
|
425
534
|
* every open finding as wrong, the upper one as right. Findings set aside as
|
|
426
535
|
* contextual are in neither: they are not claims the run could get right.
|
|
536
|
+
* Nor are findings the run filed as worth a look: it did not claim them.
|
|
427
537
|
*/
|
|
428
538
|
export function precisionBounds(c) {
|
|
429
539
|
const labelled = c.correct + c.falsePositives.length;
|
|
430
540
|
const open = c.unknown.length + c.ambiguous.length;
|
|
431
|
-
const scored = c.findings - c
|
|
541
|
+
const scored = c.findings - setAside(c);
|
|
432
542
|
return { labelled: `${c.correct}/${labelled} (${pct(c.correct, labelled)})`, low: pct(c.correct, scored), high: pct(c.correct + open, scored) };
|
|
433
543
|
}
|
|
544
|
+
/** Findings in neither precision count: contextual ones, and the run's own worth-a-looks. */
|
|
545
|
+
function setAside(c) {
|
|
546
|
+
return c.contextual.length + (c.worthALook?.length ?? 0);
|
|
547
|
+
}
|
|
434
548
|
/** The scorecard as a person reads it. Leads with the two numbers, then says what each is made of. */
|
|
435
549
|
export function formatScorecard(c) {
|
|
436
550
|
const p = precisionBounds(c);
|
|
@@ -441,11 +555,13 @@ export function formatScorecard(c) {
|
|
|
441
555
|
`Recall ${c.found.length}/${c.expected} (${pct(c.found.length, c.expected)}) of the planted defects expected at this level`,
|
|
442
556
|
`Precision ${p.labelled} of the findings the key can label` +
|
|
443
557
|
(open
|
|
444
|
-
? ` — ${open} of ${c.findings - c
|
|
558
|
+
? ` — ${open} of ${c.findings - setAside(c)} unlabelled, so between ${p.low} and ${p.high} of ${setAside(c) ? "the findings not set aside" : "all findings"}`
|
|
445
559
|
: ""),
|
|
446
560
|
];
|
|
447
561
|
if (c.contextual.length)
|
|
448
562
|
lines.push(`Set aside ${c.contextual.length} of ${c.findings} finding(s): defects only under a convention the run cannot see, so in neither count`);
|
|
563
|
+
if (c.worthALook.length)
|
|
564
|
+
lines.push(`Set aside ${c.worthALook.length} of ${c.findings} finding(s) the run filed as worth a look: not claimed as defects, so in neither recall nor precision`);
|
|
449
565
|
lines.push(``);
|
|
450
566
|
if (c.missed.length)
|
|
451
567
|
lines.push(`Missed: ${c.missed.join(", ")}`);
|
|
@@ -463,6 +579,11 @@ export function formatScorecard(c) {
|
|
|
463
579
|
for (const x of c.contextual)
|
|
464
580
|
lines.push(` ${x.title} (${x.id})`);
|
|
465
581
|
}
|
|
582
|
+
if (c.worthALook.length) {
|
|
583
|
+
lines.push(``, `Filed as worth a look (${c.worthALook.length}):`);
|
|
584
|
+
for (const x of c.worthALook)
|
|
585
|
+
lines.push(` ${x.title}${x.convention ? ` — a defect only if the project uses ${x.convention}` : ""}`);
|
|
586
|
+
}
|
|
466
587
|
if (c.ambiguous.length) {
|
|
467
588
|
lines.push(``, `Ambiguous (${c.ambiguous.length}) — the key claims each twice; sharpen it:`);
|
|
468
589
|
for (const a of c.ambiguous)
|
|
@@ -490,8 +611,10 @@ export function formatScorecard(c) {
|
|
|
490
611
|
k.ambiguous && `${k.ambiguous} the key names ambiguously`,
|
|
491
612
|
k.badConfidence && `${k.badConfidence} with an unusable confidence`,
|
|
492
613
|
k.unsure && `${k.unsure} unsure`,
|
|
493
|
-
k.outOfScope &&
|
|
614
|
+
k.outOfScope &&
|
|
615
|
+
`${k.outOfScope} dismissed as another lane's (${k.ownership === "routes" ? "by the lanes' routes" : k.ownership === "wording" ? "by wording: no lane routes archived" : "by the lanes' routes where archived, else by wording"})`,
|
|
494
616
|
k.contextual && `${k.contextual} about things that are defects only under a convention the run cannot see`,
|
|
617
|
+
k.worthALook && `${k.worthALook} the lane marked worth a look`,
|
|
495
618
|
].filter(Boolean);
|
|
496
619
|
lines.push(``, k.judged === 0
|
|
497
620
|
? `Lane calibration against the key: nothing the key could judge` + (skipped.length ? ` (${skipped.join(", ")})` : "")
|
|
@@ -499,6 +622,11 @@ export function formatScorecard(c) {
|
|
|
499
622
|
(skipped.length ? ` — not scored: ${skipped.join(", ")}` : ""));
|
|
500
623
|
for (const b of k.buckets)
|
|
501
624
|
lines.push(` stated ${b.label}: ${b.decisions} decision(s), said ${b.stated.toFixed(2)}, right ${Math.round(b.correct * 100)}%`);
|
|
625
|
+
if (k.setAsideByRoute.length) {
|
|
626
|
+
lines.push(` Set aside as another lane's, by the lanes' routes (${k.setAsideByRoute.length}):`);
|
|
627
|
+
for (const r of k.setAsideByRoute)
|
|
628
|
+
lines.push(` ${r.lane} on ${r.id}, stated ${Number.isFinite(r.confidence) ? r.confidence.toFixed(2) : "?"}`);
|
|
629
|
+
}
|
|
502
630
|
}
|
|
503
631
|
return lines.join("\n");
|
|
504
632
|
}
|
|
@@ -543,17 +671,22 @@ const LOCAL_PATH_RE = /(?:\/Users\/[^/\s"']+|\/home\/[^/\s"']+|\/private\/tmp|\/
|
|
|
543
671
|
export function sanitize(text) {
|
|
544
672
|
return text.replace(LOCAL_PATH_RE, "<path>");
|
|
545
673
|
}
|
|
546
|
-
export function toArchive(run, date, note, findings, decisions, app) {
|
|
674
|
+
export function toArchive(run, date, note, findings, decisions, app, laneRoutes = {}) {
|
|
675
|
+
// No query string or fragment reaches an archive: a lane copies routes from
|
|
676
|
+
// the address bar, and an address can carry a token.
|
|
677
|
+
const routes = Object.fromEntries(Object.entries(laneRoutes).map(([lane, list]) => [lane, list.map((r) => sanitize(stripRouteQuery(r)))]));
|
|
547
678
|
return {
|
|
548
679
|
app,
|
|
549
680
|
run,
|
|
550
681
|
date,
|
|
551
682
|
note,
|
|
683
|
+
...(Object.keys(routes).length ? { laneRoutes: routes } : {}),
|
|
552
684
|
findings: findings.map((f) => ({
|
|
553
685
|
title: sanitize(f.title),
|
|
554
686
|
severity: f.severity,
|
|
555
687
|
...(f.category ? { category: f.category } : {}),
|
|
556
688
|
...(f.evidence ? { evidence: sanitize(f.evidence) } : {}),
|
|
689
|
+
...(isWorthALook(f) ? { tier: "worth_a_look", ...(f.convention ? { convention: sanitize(f.convention) } : {}) } : {}),
|
|
557
690
|
})),
|
|
558
691
|
decisions: decisions.map((d) => ({ ...d, observation: sanitize(d.observation), evidence: d.evidence === null ? null : sanitize(d.evidence) })),
|
|
559
692
|
};
|