scenescout 3.3.0 → 3.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +27 -0
- package/README.md +38 -12
- package/dist/engine/bench.js +485 -0
- package/dist/engine/brief.js +20 -5
- package/dist/engine/browser.js +84 -27
- package/dist/engine/calibration.js +274 -0
- package/dist/engine/claims.js +10 -3
- package/dist/engine/collector.js +17 -0
- package/dist/engine/lane.js +45 -16
- package/dist/engine/memory.js +117 -4
- package/dist/engine/oracles.js +15 -8
- package/dist/engine/pace.js +57 -11
- package/dist/engine/policy.js +40 -0
- package/dist/engine/report.js +5 -1
- package/dist/engine/request.js +14 -1
- package/dist/mcp-server.js +67 -9
- package/package.json +6 -3
- package/skills/scenescout/SKILL.md +6 -6
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,32 @@
|
|
|
1
1
|
# scenescout
|
|
2
2
|
|
|
3
|
+
## 3.5.0
|
|
4
|
+
|
|
5
|
+
### Minor Changes
|
|
6
|
+
|
|
7
|
+
- 4fe8889: Parallel lanes now share what a run knows, and the planner is told what a lane left undone.
|
|
8
|
+
|
|
9
|
+
- A markup value typed in one session is watched for in every session on the same project, so a stored injection is caught when another lane opens the list that renders it.
|
|
10
|
+
- The report gate counts design audits from every session in the run, so a planner whose lanes audited the pages is no longer refused.
|
|
11
|
+
- Folding a lane report lists each judged defect that no finding matches yet, so it can be filed before the lane's session closes.
|
|
12
|
+
- A lane report wrapped in prose around one fenced JSON block is accepted, with the prose discarded unread.
|
|
13
|
+
- `scout_lane_brief` gives each lane a landing route of its own and the rules a measured run found worth stating.
|
|
14
|
+
- The pace section shows how long each session held a browser idle before its first action and after its last, including sessions that attached and never acted.
|
|
15
|
+
- An empty live region (`role="status"` and similar) is no longer shown or counted as an unnamed control.
|
|
16
|
+
- 6ec8fee: The write policy now answers a page's blocked `fetch` or XHR write with a `403` in the server's place instead of dropping it. The server is still never contacted, but the page's handling of a refusal really runs, so a page that reports a refused save or delete as a success is caught as a `false_success` (and says the refusal was the policy's stand-in). Blocked navigations are still dropped. The stand-in 403 is not reported as an HTTP error of the app.
|
|
17
|
+
|
|
18
|
+
## 3.4.0
|
|
19
|
+
|
|
20
|
+
### Minor Changes
|
|
21
|
+
|
|
22
|
+
- 04f0f22: Measure whether a lane's confidence means anything, and make one lane rubric serve a whole wave.
|
|
23
|
+
|
|
24
|
+
Every lane has been told its confidence must be calibrated — "0.5 means a coin flip, 0.95 means you would bet on it" — and the number was then averaged into one line and discarded. Nothing was stored, so nothing could ever be checked, and a confidence nobody checks is decoration.
|
|
25
|
+
|
|
26
|
+
Lane decisions are now kept, and the report carries a calibration section: how often a decision at a stated confidence matched a finding the project holds, bucketed, with an expected calibration error. It says plainly what the number is not — agreement between the lanes and the bar the project applies, over its whole history, not evidence that the app is broken — and names the two ways a lane is counted wrong through no fault of its own. A decision naming no failing endpoint cannot be looked up at all, so it is excluded and disclosed rather than scored as a miss. Below eight checkable decisions the figure is withheld, but the section says so rather than vanishing. Where findings have since been re-tested through `scout_verify`, those verdicts are reported beside it, because they *are* evidence about the app.
|
|
27
|
+
|
|
28
|
+
Separately, the lane name used to sit in the second sentence of the instruction every lane receives, so two lanes' prompts diverged almost immediately and shared no prefix. The rubric is now identical for every lane in a wave and the name is the last thing said, which makes it one cacheable prefix instead of one per lane.
|
|
29
|
+
|
|
3
30
|
## 3.3.0
|
|
4
31
|
|
|
5
32
|
### Minor Changes
|
package/README.md
CHANGED
|
@@ -205,7 +205,8 @@ The view is served on `127.0.0.1` only, behind a token that changes every time t
|
|
|
205
205
|
|
|
206
206
|
**Try it with parallel agents.** The demo app has three roles and several separate areas, so a run can be split between agents. Start it with `npm run demo:serve`, then ask your agent to explore it with several agents in parallel, one role and one area each. The pictures above come from a run of three. Two things keep a parallel run efficient:
|
|
207
207
|
|
|
208
|
-
- **Each agent opens its own session when it starts and closes it
|
|
208
|
+
- **Each agent opens its own session when it starts, and the planner closes it once it has folded that agent's report.** An agent waiting for its turn then holds no browser. Opening every session up front leaves browsers idling while the machine runs out of memory for the agents that are working. Closing before the fold loses the lane's decisions, which have nowhere to be kept.
|
|
209
|
+
- **Slow it down to follow along.** `scout_attach {paceMs}` (or `scout_session {paceMs}` mid-run) sets a floor between actions, for when you want to watch a flow rather than let it run as fast as the page allows.
|
|
209
210
|
- **Run about as many agents at once as your machine has cores, less two.** Each one drives a real browser.
|
|
210
211
|
|
|
211
212
|
---
|
|
@@ -272,7 +273,7 @@ Snapshots are cheap: re-snapshotting a route returns only *what changed*, with s
|
|
|
272
273
|
|
|
273
274
|
## 🧰 The toolbox
|
|
274
275
|
|
|
275
|
-
|
|
276
|
+
29 deterministic tools. The agent picks; you rarely call these by hand.
|
|
276
277
|
|
|
277
278
|
| Phase | Tools | What they do |
|
|
278
279
|
|---|---|---|
|
|
@@ -283,6 +284,8 @@ Snapshots are cheap: re-snapshotting a route returns only *what changed*, with s
|
|
|
283
284
|
| **Act** | `scout_click` `scout_type` `scout_select` `scout_upload` `scout_press` `scout_scroll` `scout_navigate` `scout_back` `scout_run_plan` | Drive the UI like a user; `scout_run_plan` batches a whole mechanical sequence into one call |
|
|
284
285
|
| **Assess** | `scout_design_audit` `scout_journey` | Score a page's craft/a11y/consistency; measure how hard a task is to complete |
|
|
285
286
|
| **Record** | `scout_note` `scout_finding` `scout_resolve` `scout_report` | Curate durable notes; file deduped findings; mark fixes; write the report, and on a recorded run the whole run as one page |
|
|
287
|
+
| **Re-test** | `scout_verify` | List the findings earlier runs left open, worst route first, and record whether each is gone, still present, or changed |
|
|
288
|
+
| **Split the work** | `scout_lane_brief` `scout_lane_report` | Divide the app between parallel agents by whole module, each with its own landing route and rules; fold what each hands back as one typed JSON object, and name any defect it judged but never filed |
|
|
286
289
|
| **Close** | `scout_close` | Tear down one session or all |
|
|
287
290
|
|
|
288
291
|
A few that punch above their weight:
|
|
@@ -292,20 +295,23 @@ A few that punch above their weight:
|
|
|
292
295
|
- **`scout_journey`** — wraps one goal and reports interaction count, screens seen, and **backtracks**; an abandoned journey is a finding no passing E2E suite can produce.
|
|
293
296
|
- **`scout_upload`** — generates a *valid* in-memory fixture (real PDF/PNG, kind inferred from `accept`) so file-upload flows stop being a blind spot.
|
|
294
297
|
- **`scout_click {clicks: 2}`** — the impatient-user probe: states whether a double-click fired the same state-changing request twice (the classic double-submit bug).
|
|
298
|
+
- **`scout_request`** — calls the app's own API as the session, so "the button is hidden" becomes "the server refuses it" (or doesn't).
|
|
299
|
+
|
|
300
|
+
Beyond crashes and HTTP errors, two oracles catch a page **contradicting the server**: `refused_empty` (a list request was refused and the page shows its empty state with no error) and `false_success` (a save was refused and the page says it worked). A third, `dom_injection`, reports a typed markup value coming back as an element on any page any session opens.
|
|
295
301
|
|
|
296
302
|
---
|
|
297
303
|
|
|
298
304
|
## 📊 Test levels
|
|
299
305
|
|
|
300
|
-
Each level is
|
|
306
|
+
Each level is a contract. `scout_report` enforces what the engine can see for itself — routes visited, pages audited, and at `extensive` an empty gap ledger (every route exercised and audited, every filled form submitted, a completed journey, two roles) — and the report's gap ledger discloses the rest.
|
|
301
307
|
|
|
302
308
|
| Level | What it guarantees | Rough size |
|
|
303
309
|
|---|---|---|
|
|
304
310
|
| `minimal` | Every route visited, ≥1 design audit, key journeys as plans, crawl problems triaged. Remaining gaps **disclosed**. | ~40 actions |
|
|
305
|
-
| `medium` *(default)* | minimal + design audits across several routes + every element class exercised + every form submitted valid **and** invalid | ~150 actions |
|
|
306
|
-
| `extensive` | medium + fuzzing, back/refresh/deep-link resilience, keyboard-only pass, a journey per module, ≥2 roles compared, anonymous auth-surface walk. **Refuses to finalize while any gap remains.** | budget-capped |
|
|
311
|
+
| `medium` *(default)* | minimal + design audits across several routes (enforced) + every element class exercised + every form submitted valid **and** invalid (asked of the agent) | ~150 actions |
|
|
312
|
+
| `extensive` | medium + fuzzing, back/refresh/deep-link resilience, keyboard-only pass, a journey per module, ≥2 roles compared, anonymous auth-surface walk. **Refuses to finalize while any gap in the ledger remains.** | budget-capped |
|
|
307
313
|
|
|
308
|
-
That refusal *is* the guarantee: an extensive report can only exist when nothing
|
|
314
|
+
That refusal *is* the guarantee: an extensive report can only exist when nothing the engine can measure was left untested. Fuzzing, the keyboard pass and the auth-surface walk are the agent's to do; the engine cannot see whether they were done well.
|
|
309
315
|
|
|
310
316
|
---
|
|
311
317
|
|
|
@@ -317,7 +323,7 @@ That refusal *is* the guarantee: an extensive report can only exist when nothing
|
|
|
317
323
|
- 🔴 **`destructive`** (`--allow-destructive`) allows everything, and only ever when *you* confirm the environment is disposable. The skill will never choose this itself.
|
|
318
324
|
- 📂 Findings, memory, and reports live in a `.scenescout/` folder where you ran it. It ignores itself in git, so a stray `git add -A` never commits test data.
|
|
319
325
|
|
|
320
|
-
A `🛡 WRITE-POLICY blocked` notice is the safety net doing its job, not an app bug.
|
|
326
|
+
A `🛡 WRITE-POLICY blocked` notice is the safety net doing its job, not an app bug. The server never sees a blocked request, but a page's own `fetch` or XHR is answered with a `403` in its place rather than dropped, so the page's handling of a refusal really runs: a page that then claims success is reported as a `false_success` ([ADR 9](docs/adr/0009-a-refused-write-is-answered-not-dropped.md)).
|
|
321
327
|
|
|
322
328
|
---
|
|
323
329
|
|
|
@@ -329,6 +335,8 @@ A `🛡 WRITE-POLICY blocked` notice is the safety net doing its job, not an app
|
|
|
329
335
|
- 💯 **Page scores** (0–100: a11y · craft · consistency · task-clarity), ranked worst-first, with stale scores from old runs marked as such.
|
|
330
336
|
- 👥 **A role capability matrix** — what each role could and couldn't reach.
|
|
331
337
|
- 🧾 **A gap ledger** — everything *not* done, so the report is honest about its own coverage.
|
|
338
|
+
- ⏱️ **How the run was paced** — actions, median gap, idle share and held-idle time per session, so a browser held open for nothing is visible.
|
|
339
|
+
- 🎯 **How well the lanes judged** — on a parallel run, whether the confidence each lane stated matched what the project went on to file, beside what later re-tests found ([ADR 10](docs/adr/0010-a-confidence-is-checked-not-trusted.md)).
|
|
332
340
|
|
|
333
341
|
`.scenescout/report.html` — the same report as one self-contained page, with every session's trail beside it, and on a [recorded run](#-recording-a-run-and-reading-it-back) the screenshots under each finding.
|
|
334
342
|
|
|
@@ -539,7 +547,7 @@ npx -y scenescout watch <path> # the same, live in your browser, with each
|
|
|
539
547
|
|
|
540
548
|
```
|
|
541
549
|
src/
|
|
542
|
-
mcp-server.ts the
|
|
550
|
+
mcp-server.ts the 29 tools + per-session dispatch
|
|
543
551
|
scan.ts project discovery (framework, routes, auth)
|
|
544
552
|
cli.ts scan · serve · install · doctor · status
|
|
545
553
|
installer.ts setup logic (skill link, MCP registration, diagnostics)
|
|
@@ -549,8 +557,15 @@ src/
|
|
|
549
557
|
fingerprint.ts route + element-set identity (state hashing)
|
|
550
558
|
oracles.ts console/page/network/HTTP error detection
|
|
551
559
|
injection.ts the DOM-injection oracle's rules (what to watch for, how to find it)
|
|
560
|
+
claims.ts when the page contradicts the server (refused_empty, false_success)
|
|
561
|
+
request.ts what a replayed API call may be and where it may go
|
|
562
|
+
brief.ts splitting the app between parallel lanes
|
|
563
|
+
lane.ts the typed report a lane hands back
|
|
564
|
+
calibration.ts whether a lane's confidence held up; what it judged and never filed
|
|
565
|
+
pace.ts how a run spent its time
|
|
566
|
+
bench.ts scoring a run against the demo app's answer key
|
|
552
567
|
policy.ts the write-policy safety net
|
|
553
|
-
ownership.ts safe-write: which records
|
|
568
|
+
ownership.ts safe-write: which records were created in this process?
|
|
554
569
|
uploads.ts disk uploads, fenced to the project by real path
|
|
555
570
|
journey.ts task-ease measurement from the action log
|
|
556
571
|
design.ts the design audit + page scoring
|
|
@@ -558,9 +573,11 @@ src/
|
|
|
558
573
|
report.ts the gap ledger + report generation
|
|
559
574
|
replay.ts the run as one page: steps, tasks, frames under each finding
|
|
560
575
|
… collector · dispatch · fixtures · authloss · reaper
|
|
561
|
-
scripts/ the
|
|
576
|
+
scripts/ the 23 test suites (smoke/ holds the real-browser ones)
|
|
562
577
|
test-app/ fixtures for the real-browser smoke tests
|
|
563
578
|
skills/scenescout/ the testing method (SKILL.md): a skill in Claude Code, served by the server everywhere else
|
|
579
|
+
docs/how-it-works.md what happens at each stage, in diagrams
|
|
580
|
+
docs/benchmark.md measuring whether a change made runs better
|
|
564
581
|
docs/adr/ why it's built this way
|
|
565
582
|
```
|
|
566
583
|
|
|
@@ -570,6 +587,10 @@ docs/adr/ why it's built this way
|
|
|
570
587
|
|
|
571
588
|
## 🧠 Design decisions
|
|
572
589
|
|
|
590
|
+
**[How it works, stage by stage](docs/how-it-works.md)** — diagrams of the run lifecycle, what happens inside one action, the write policy on the wire, how a violation becomes a finding, how a parallel run is split and folded, how roles hand work to each other, where a run's time goes, and how a lane's confidence is checked afterwards.
|
|
591
|
+
|
|
592
|
+
**[Measuring whether a change helped](docs/benchmark.md)** — the demo app's answer key, the scorecard (recall, precision, judged-not-filed, severity, calibration), and the results log of every run, including what did not help.
|
|
593
|
+
|
|
573
594
|
The load-bearing choices are recorded as ADRs — read the relevant one before changing a rule it covers:
|
|
574
595
|
|
|
575
596
|
- [1 · Completion is an enforced contract, not a claim](docs/adr/0001-completion-is-a-contract-not-a-vibe.md)
|
|
@@ -578,6 +599,10 @@ The load-bearing choices are recorded as ADRs — read the relevant one before c
|
|
|
578
599
|
- [4 · Findings dedup on machine signals, and a merge must never lose a finding](docs/adr/0004-dedup-on-machine-signals-not-prose.md)
|
|
579
600
|
- [5 · Testable logic lives outside `browser.ts`](docs/adr/0005-keep-testable-logic-out-of-the-browser-module.md)
|
|
580
601
|
- [6 · Nothing in this repo names or is tuned for a tested app](docs/adr/0006-stay-project-agnostic.md)
|
|
602
|
+
- [7 · The live view is local, read-only, and leaves nothing behind](docs/adr/0007-the-live-view-is-local-read-only-and-leaves-nothing-behind.md)
|
|
603
|
+
- [8 · Recording is opt-in, and a recorded run is one self-contained page](docs/adr/0008-a-recorded-run-is-evidence-and-must-be-asked-for.md)
|
|
604
|
+
- [9 · A refused write is answered, not dropped](docs/adr/0009-a-refused-write-is-answered-not-dropped.md)
|
|
605
|
+
- [10 · A lane's confidence is checked, not trusted](docs/adr/0010-a-confidence-is-checked-not-trusted.md)
|
|
581
606
|
|
|
582
607
|
---
|
|
583
608
|
|
|
@@ -589,8 +614,9 @@ Working on SceneScout itself is the only reason to clone it:
|
|
|
589
614
|
git clone https://github.com/brunoboto96/SceneScout.git scenescout && cd scenescout
|
|
590
615
|
npm install # installs dependencies and builds
|
|
591
616
|
npm run setup # same as `scenescout install`, but registers THIS checkout (the skill is linked, so edits are live)
|
|
592
|
-
npm test # build +
|
|
593
|
-
#
|
|
617
|
+
npm test # build + 23 suites: 21 pure-logic suites (scan, oracle, policy, … bench, hygiene),
|
|
618
|
+
# then smoke and mcp-check (the server over stdio), both with real browsers
|
|
619
|
+
npm run bench -- --all # re-score every archived benchmark run against the current answer key
|
|
594
620
|
npm run demo # regenerate examples/ from the demo app
|
|
595
621
|
```
|
|
596
622
|
|
|
@@ -0,0 +1,485 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* Scoring a run against an answer key.
|
|
3
|
+
*
|
|
4
|
+
* Every change to this engine and to its skill has been made on judgement:
|
|
5
|
+
* something looked wrong in a run, it was fixed, and the next run looked
|
|
6
|
+
* better — or looked different, which is not the same thing. There was no
|
|
7
|
+
* number that could go up or down, so there was no way to tell a fix from a
|
|
8
|
+
* coincidence, or to notice a change that made things worse.
|
|
9
|
+
*
|
|
10
|
+
* An app with KNOWN defects makes a number possible. Given what the app
|
|
11
|
+
* actually contains and what a run reported, this computes recall (what it
|
|
12
|
+
* found of what was there), precision (how much of what it reported was
|
|
13
|
+
* real), how often it judged a defect and then never filed it, whether the
|
|
14
|
+
* severities it chose agree with the key, and whether the confidence its lanes
|
|
15
|
+
* stated was worth what they said — measured against the key, which is the
|
|
16
|
+
* one ground truth a run cannot talk itself into.
|
|
17
|
+
*
|
|
18
|
+
* Every future change is accepted or rejected by the number this produces, so
|
|
19
|
+
* the failure that matters most here is a PLAUSIBLE WRONG number. Three rules
|
|
20
|
+
* follow from that. A piece of text two key entries both claim is reported as
|
|
21
|
+
* ambiguous and scored as neither. The key is validated when it is read, so a
|
|
22
|
+
* mistyped level cannot silently drop a defect from recall. And the key is
|
|
23
|
+
* hashed into every scorecard, so two runs scored against different keys are
|
|
24
|
+
* never compared by accident.
|
|
25
|
+
*
|
|
26
|
+
* The key is data, not code: everything specific to one app lives in its key
|
|
27
|
+
* file, so this module stays as generic as the rest of the engine.
|
|
28
|
+
*
|
|
29
|
+
* Pure, so every rule is table-tested.
|
|
30
|
+
*/
|
|
31
|
+
import { createHash } from "node:crypto";
|
|
32
|
+
import { z } from "zod";
|
|
33
|
+
import { BUCKET_EDGES, bucketLabel, bucketOf } from "./calibration.js";
|
|
34
|
+
export const LEVELS = ["minimal", "medium", "extensive"];
|
|
35
|
+
const SEVERITIES = ["high", "medium", "low"];
|
|
36
|
+
const Pattern = z.string().refine((p) => {
|
|
37
|
+
try {
|
|
38
|
+
new RegExp(p, "i");
|
|
39
|
+
return true;
|
|
40
|
+
}
|
|
41
|
+
catch {
|
|
42
|
+
return false;
|
|
43
|
+
}
|
|
44
|
+
}, { message: "is not a valid regular expression" });
|
|
45
|
+
const Entry = z
|
|
46
|
+
.object({
|
|
47
|
+
id: z.string().min(1),
|
|
48
|
+
route: z.string().startsWith("/"),
|
|
49
|
+
title: z.string().min(1),
|
|
50
|
+
/** The category a finding for it should carry. Reported, not enforced: two categories can both be defensible. */
|
|
51
|
+
category: z.string().min(1),
|
|
52
|
+
/** The severity the key's author would give it. A judgement, and labelled as one wherever it is reported. */
|
|
53
|
+
severity: z.enum(SEVERITIES),
|
|
54
|
+
/** The lowest level whose contract is expected to find it. A defect only a fuzzing pass can reach is not a miss at `medium`. */
|
|
55
|
+
level: z.enum(LEVELS),
|
|
56
|
+
/** Case-insensitive regular expressions tried against a finding's evidence and title. */
|
|
57
|
+
match: z.array(Pattern).min(1),
|
|
58
|
+
/** Phrasings that MUST classify to this entry. The title is always one of them. */
|
|
59
|
+
examples: z.array(z.string()).default([]),
|
|
60
|
+
/** Near misses that must NOT classify to this entry — the other half of each contrastive pair. */
|
|
61
|
+
counterExamples: z.array(z.string()).default([]),
|
|
62
|
+
})
|
|
63
|
+
.strict();
|
|
64
|
+
const NonDefect = z
|
|
65
|
+
.object({
|
|
66
|
+
id: z.string().min(1),
|
|
67
|
+
title: z.string().min(1),
|
|
68
|
+
/** Why it is not a defect, for the reader of a scorecard. */
|
|
69
|
+
why: z.string().min(1),
|
|
70
|
+
match: z.array(Pattern).min(1),
|
|
71
|
+
/**
|
|
72
|
+
* The defects this non-defect is a narrowing of: text that matches both is
|
|
73
|
+
* the false claim, not the defect. Only these. A non-defect that happens to
|
|
74
|
+
* overlap any other entry makes the text ambiguous, because a phrasing
|
|
75
|
+
* nobody anticipated is likelier to be the real defect than the false claim.
|
|
76
|
+
*/
|
|
77
|
+
overrides: z.array(z.string()).default([]),
|
|
78
|
+
examples: z.array(z.string()).default([]),
|
|
79
|
+
counterExamples: z.array(z.string()).default([]),
|
|
80
|
+
})
|
|
81
|
+
.strict();
|
|
82
|
+
const Key = z
|
|
83
|
+
.object({
|
|
84
|
+
app: z.string().min(1),
|
|
85
|
+
_comment: z.string().optional(),
|
|
86
|
+
/** The defects the app was built to contain. Recall is measured against these. */
|
|
87
|
+
defects: z.array(Entry).min(1),
|
|
88
|
+
/**
|
|
89
|
+
* Real problems a run found that nobody planted. They count as correct for
|
|
90
|
+
* precision but not toward recall, so the number the key was designed around
|
|
91
|
+
* stays comparable from one run to the next.
|
|
92
|
+
*/
|
|
93
|
+
alsoReal: z.array(Entry).default([]),
|
|
94
|
+
nonDefects: z.array(NonDefect).default([]),
|
|
95
|
+
})
|
|
96
|
+
.strict();
|
|
97
|
+
/**
|
|
98
|
+
* Read and validate a key. A key that fails here fails loudly: a mistyped
|
|
99
|
+
* `"level": "Minimal"` used to drop the defect out of recall with no error,
|
|
100
|
+
* and the run scored as if the app had one defect fewer.
|
|
101
|
+
*/
|
|
102
|
+
export function parseKey(raw) {
|
|
103
|
+
const parsed = Key.safeParse(raw);
|
|
104
|
+
if (!parsed.success) {
|
|
105
|
+
const issue = parsed.error.issues[0];
|
|
106
|
+
throw new Error(`The answer key is not valid at ${issue.path.join(".") || "(root)"}: ${issue.message}`);
|
|
107
|
+
}
|
|
108
|
+
const ids = [...parsed.data.defects, ...parsed.data.alsoReal, ...parsed.data.nonDefects].map((e) => e.id);
|
|
109
|
+
const dup = ids.find((id, i) => ids.indexOf(id) !== i);
|
|
110
|
+
if (dup)
|
|
111
|
+
throw new Error(`The answer key uses the id ${JSON.stringify(dup)} twice.`);
|
|
112
|
+
const real = new Set([...parsed.data.defects, ...parsed.data.alsoReal].map((e) => e.id));
|
|
113
|
+
for (const nd of parsed.data.nonDefects) {
|
|
114
|
+
const unknown = nd.overrides.find((id) => !real.has(id));
|
|
115
|
+
if (unknown)
|
|
116
|
+
throw new Error(`The non-defect ${JSON.stringify(nd.id)} overrides ${JSON.stringify(unknown)}, which is not a defect in the key.`);
|
|
117
|
+
}
|
|
118
|
+
return parsed.data;
|
|
119
|
+
}
|
|
120
|
+
/** A short, stable fingerprint of a key, so a scorecard says which key produced it. */
|
|
121
|
+
export function keyHash(key) {
|
|
122
|
+
const canonical = JSON.stringify({ defects: key.defects, alsoReal: key.alsoReal, nonDefects: key.nonDefects });
|
|
123
|
+
return createHash("sha256").update(canonical).digest("hex").slice(0, 10);
|
|
124
|
+
}
|
|
125
|
+
const RANK = { minimal: 0, medium: 1, extensive: 2 };
|
|
126
|
+
const SEVERITY_RANK = { low: 0, medium: 1, high: 2 };
|
|
127
|
+
function matches(patterns, text) {
|
|
128
|
+
return patterns.some((p) => new RegExp(p, "i").test(text));
|
|
129
|
+
}
|
|
130
|
+
/**
|
|
131
|
+
* What a piece of reported text is, according to the key.
|
|
132
|
+
*
|
|
133
|
+
* A known non-defect can be a deliberate narrowing of a real defect — "the
|
|
134
|
+
* filter is stuck after the Archived 500" names the Archived failure but claims
|
|
135
|
+
* something false about it — so a non-defect wins over the defects it lists in
|
|
136
|
+
* `overrides`, and over nothing else. Anything else that more than one entry
|
|
137
|
+
* claims is AMBIGUOUS, reported, and scored as neither: crediting the first
|
|
138
|
+
* match would quietly count a finding about two defects as one, and letting a
|
|
139
|
+
* non-defect win everywhere turned realistic rewordings of real defects into
|
|
140
|
+
* false positives with a confident explanation beside them.
|
|
141
|
+
*/
|
|
142
|
+
export function classify(text, key) {
|
|
143
|
+
const nd = key.nonDefects.filter((e) => matches(e.match, text));
|
|
144
|
+
const real = [
|
|
145
|
+
...key.defects.filter((e) => matches(e.match, text)).map((entry) => ({ kind: "defect", entry })),
|
|
146
|
+
...key.alsoReal.filter((e) => matches(e.match, text)).map((entry) => ({ kind: "alsoReal", entry })),
|
|
147
|
+
];
|
|
148
|
+
if (nd.length > 1)
|
|
149
|
+
return { kind: "ambiguous", ids: [...nd, ...real.map((m) => m.entry)].map((e) => e.id) };
|
|
150
|
+
if (nd.length === 1) {
|
|
151
|
+
const rest = real.filter((m) => !nd[0].overrides.includes(m.entry.id));
|
|
152
|
+
if (rest.length === 0)
|
|
153
|
+
return { kind: "nonDefect", entry: nd[0] };
|
|
154
|
+
return { kind: "ambiguous", ids: [nd[0].id, ...rest.map((m) => m.entry.id)] };
|
|
155
|
+
}
|
|
156
|
+
if (real.length === 1)
|
|
157
|
+
return real[0];
|
|
158
|
+
if (real.length > 1)
|
|
159
|
+
return { kind: "ambiguous", ids: real.map((m) => m.entry.id) };
|
|
160
|
+
return null;
|
|
161
|
+
}
|
|
162
|
+
/** The text a finding is classified on. Evidence leads, because it is the machine signature; the title backs it up. */
|
|
163
|
+
export function findingText(f) {
|
|
164
|
+
return `${f.evidence ?? ""}\n${f.title}`;
|
|
165
|
+
}
|
|
166
|
+
/** The text a lane decision is classified on. */
|
|
167
|
+
export function decisionText(d) {
|
|
168
|
+
return `${d.evidence ?? ""}\n${d.observation}`;
|
|
169
|
+
}
|
|
170
|
+
/**
|
|
171
|
+
* A key that disagrees with its own examples is wrong before any run is
|
|
172
|
+
* scored. Returns every title or example — non-defects' included — that
|
|
173
|
+
* classifies anywhere but its own entry, and every counter-example that
|
|
174
|
+
* classifies TO its entry.
|
|
175
|
+
*/
|
|
176
|
+
export function lintKey(key) {
|
|
177
|
+
const problems = [];
|
|
178
|
+
const describe = (got) => got === null ? "nothing" : got.kind === "ambiguous" ? `ambiguous (${got.ids.join(", ")})` : got.entry.id;
|
|
179
|
+
const check = (id, kind, text, want) => {
|
|
180
|
+
const got = classify(text, key);
|
|
181
|
+
const hit = got !== null && got.kind !== "ambiguous" && got.kind === kind && got.entry.id === id;
|
|
182
|
+
if (want && !hit)
|
|
183
|
+
problems.push(`${id}: ${JSON.stringify(text)} classified as ${describe(got)}`);
|
|
184
|
+
if (!want && hit)
|
|
185
|
+
problems.push(`${id}: counter-example ${JSON.stringify(text)} classified TO it`);
|
|
186
|
+
};
|
|
187
|
+
for (const [kind, list] of [
|
|
188
|
+
["defect", key.defects],
|
|
189
|
+
["alsoReal", key.alsoReal],
|
|
190
|
+
]) {
|
|
191
|
+
for (const e of list) {
|
|
192
|
+
for (const t of [e.title, ...e.examples])
|
|
193
|
+
check(e.id, kind, t, true);
|
|
194
|
+
for (const t of e.counterExamples)
|
|
195
|
+
check(e.id, kind, t, false);
|
|
196
|
+
}
|
|
197
|
+
}
|
|
198
|
+
for (const e of key.nonDefects) {
|
|
199
|
+
for (const t of [e.title, ...e.examples])
|
|
200
|
+
check(e.id, "nonDefect", t, true);
|
|
201
|
+
for (const t of e.counterExamples)
|
|
202
|
+
check(e.id, "nonDefect", t, false);
|
|
203
|
+
}
|
|
204
|
+
return problems;
|
|
205
|
+
}
|
|
206
|
+
/**
|
|
207
|
+
* Whether a lane's verdict was right, according to the key.
|
|
208
|
+
*
|
|
209
|
+
* "defect" on a planted or also-real defect is right, and on a known
|
|
210
|
+
* non-defect is wrong. "not_a_defect" is the reverse. An "unsure" verdict, a
|
|
211
|
+
* dismissal as another lane's, a decision the key does not name, and one it
|
|
212
|
+
* names ambiguously are not scored — the same rule the in-product calibration
|
|
213
|
+
* follows, for the same reason.
|
|
214
|
+
*/
|
|
215
|
+
export function judgeDecision(d, key) {
|
|
216
|
+
if (d.verdict === "unsure")
|
|
217
|
+
return null;
|
|
218
|
+
if (isScopeDismissal(d))
|
|
219
|
+
return null;
|
|
220
|
+
const m = classify(decisionText(d), key);
|
|
221
|
+
if (!m || m.kind === "ambiguous")
|
|
222
|
+
return null;
|
|
223
|
+
const isDefect = m.kind !== "nonDefect";
|
|
224
|
+
return d.verdict === "defect" ? isDefect : !isDefect;
|
|
225
|
+
}
|
|
226
|
+
/** How a lane says "not mine": the wording lanes actually used, none of it about a particular app. */
|
|
227
|
+
const SCOPE_DISMISSAL_RE = /out.of.lane.scope|out-of-scope|out of (my|this) (lane|scope)|belongs?.to.{0,20}\b(lane|route|page)\b|belongs-to-|(handled|owned) by (the |another )?[\w-]* ?lane|lane owns it|not (in |on |from |part of )?(my|this) (lane|assigned routes|routes?|pages?)\b|outside (my|this) (lane|routes?|pages?)|not part of my (assigned )?routes/i;
|
|
228
|
+
/** A not-a-defect verdict whose stated reason is that the thing is another lane's. */
|
|
229
|
+
export function isScopeDismissal(d) {
|
|
230
|
+
return d.verdict === "not_a_defect" && SCOPE_DISMISSAL_RE.test(decisionText(d));
|
|
231
|
+
}
|
|
232
|
+
export function calibrateAgainstKey(decisions, key) {
|
|
233
|
+
if (decisions.length === 0)
|
|
234
|
+
return null;
|
|
235
|
+
const buckets = BUCKET_EDGES.map(() => ({ n: 0, conf: 0, right: 0 }));
|
|
236
|
+
const out = { judged: 0, correct: 0, notInKey: 0, ambiguous: 0, badConfidence: 0, unsure: 0, outOfScope: 0, buckets: [], ece: 0, brier: 0 };
|
|
237
|
+
for (const d of decisions) {
|
|
238
|
+
if (d.verdict === "unsure") {
|
|
239
|
+
out.unsure += 1;
|
|
240
|
+
continue;
|
|
241
|
+
}
|
|
242
|
+
if (isScopeDismissal(d)) {
|
|
243
|
+
out.outOfScope += 1;
|
|
244
|
+
continue;
|
|
245
|
+
}
|
|
246
|
+
const c = d.confidence;
|
|
247
|
+
if (typeof c !== "number" || !Number.isFinite(c) || c < 0 || c > 1) {
|
|
248
|
+
out.badConfidence += 1;
|
|
249
|
+
continue;
|
|
250
|
+
}
|
|
251
|
+
const m = classify(decisionText(d), key);
|
|
252
|
+
if (!m) {
|
|
253
|
+
out.notInKey += 1;
|
|
254
|
+
continue;
|
|
255
|
+
}
|
|
256
|
+
if (m.kind === "ambiguous") {
|
|
257
|
+
out.ambiguous += 1;
|
|
258
|
+
continue;
|
|
259
|
+
}
|
|
260
|
+
const right = d.verdict === "defect" ? m.kind !== "nonDefect" : m.kind === "nonDefect";
|
|
261
|
+
out.judged += 1;
|
|
262
|
+
if (right)
|
|
263
|
+
out.correct += 1;
|
|
264
|
+
out.brier += (c - (right ? 1 : 0)) ** 2;
|
|
265
|
+
const b = buckets[bucketOf(c)];
|
|
266
|
+
b.n += 1;
|
|
267
|
+
b.conf += c;
|
|
268
|
+
if (right)
|
|
269
|
+
b.right += 1;
|
|
270
|
+
}
|
|
271
|
+
if (out.judged > 0)
|
|
272
|
+
out.brier /= out.judged;
|
|
273
|
+
buckets.forEach((b, i) => {
|
|
274
|
+
if (b.n === 0)
|
|
275
|
+
return;
|
|
276
|
+
out.ece += (b.n / out.judged) * Math.abs(b.conf / b.n - b.right / b.n);
|
|
277
|
+
out.buckets.push({ label: bucketLabel(i), decisions: b.n, stated: b.conf / b.n, correct: b.right / b.n });
|
|
278
|
+
});
|
|
279
|
+
return out;
|
|
280
|
+
}
|
|
281
|
+
/**
|
|
282
|
+
* Score one run.
|
|
283
|
+
*
|
|
284
|
+
* `findings` should be the run's own: a project memory accumulates findings
|
|
285
|
+
* across runs, and scoring an accumulated one credits a run with what an
|
|
286
|
+
* earlier run found. Point the scorer at a fresh project directory per run.
|
|
287
|
+
*/
|
|
288
|
+
export function score(key, findings, decisions, level = "medium") {
|
|
289
|
+
const hitsByDefect = new Map();
|
|
290
|
+
const falsePositives = [];
|
|
291
|
+
const unknown = [];
|
|
292
|
+
const ambiguous = [];
|
|
293
|
+
let correct = 0;
|
|
294
|
+
for (const f of findings) {
|
|
295
|
+
const m = classify(findingText(f), key);
|
|
296
|
+
if (!m) {
|
|
297
|
+
unknown.push({ title: f.title, severity: f.severity, evidence: f.evidence ?? "" });
|
|
298
|
+
continue;
|
|
299
|
+
}
|
|
300
|
+
if (m.kind === "ambiguous") {
|
|
301
|
+
ambiguous.push({ title: f.title, ids: m.ids });
|
|
302
|
+
continue;
|
|
303
|
+
}
|
|
304
|
+
if (m.kind === "nonDefect") {
|
|
305
|
+
falsePositives.push({ id: m.entry.id, title: f.title, why: m.entry.why, severity: f.severity });
|
|
306
|
+
continue;
|
|
307
|
+
}
|
|
308
|
+
correct += 1;
|
|
309
|
+
if (m.kind === "defect") {
|
|
310
|
+
const list = hitsByDefect.get(m.entry.id) ?? [];
|
|
311
|
+
list.push(f);
|
|
312
|
+
hitsByDefect.set(m.entry.id, list);
|
|
313
|
+
}
|
|
314
|
+
}
|
|
315
|
+
const expectedDefects = key.defects.filter((d) => RANK[d.level] <= RANK[level]);
|
|
316
|
+
const found = expectedDefects.filter((d) => hitsByDefect.has(d.id)).map((d) => d.id);
|
|
317
|
+
const missed = expectedDefects.filter((d) => !hitsByDefect.has(d.id)).map((d) => d.id);
|
|
318
|
+
const beyondLevel = key.defects.filter((d) => RANK[d.level] > RANK[level] && hitsByDefect.has(d.id)).map((d) => d.id);
|
|
319
|
+
// A defect a lane called out in its report, which no finding covers.
|
|
320
|
+
const judged = new Set();
|
|
321
|
+
for (const d of decisions) {
|
|
322
|
+
if (d.verdict !== "defect")
|
|
323
|
+
continue;
|
|
324
|
+
const m = classify(decisionText(d), key);
|
|
325
|
+
if (m?.kind === "defect" && !hitsByDefect.has(m.entry.id))
|
|
326
|
+
judged.add(m.entry.id);
|
|
327
|
+
}
|
|
328
|
+
const severity = { agree: 0, higher: 0, lower: 0, detail: [] };
|
|
329
|
+
let duplicates = 0;
|
|
330
|
+
for (const d of key.defects) {
|
|
331
|
+
const hits = hitsByDefect.get(d.id);
|
|
332
|
+
if (!hits)
|
|
333
|
+
continue;
|
|
334
|
+
duplicates += hits.length - 1;
|
|
335
|
+
// The most severe finding for it is the one a reader acts on.
|
|
336
|
+
const filed = hits.map((h) => h.severity).sort((a, b) => (SEVERITY_RANK[b] ?? 0) - (SEVERITY_RANK[a] ?? 0))[0];
|
|
337
|
+
const diff = (SEVERITY_RANK[filed] ?? 0) - SEVERITY_RANK[d.severity];
|
|
338
|
+
if (diff === 0)
|
|
339
|
+
severity.agree += 1;
|
|
340
|
+
else if (diff > 0)
|
|
341
|
+
severity.higher += 1;
|
|
342
|
+
else
|
|
343
|
+
severity.lower += 1;
|
|
344
|
+
severity.detail.push({ id: d.id, filed, key: d.severity });
|
|
345
|
+
}
|
|
346
|
+
return {
|
|
347
|
+
app: key.app,
|
|
348
|
+
key: keyHash(key),
|
|
349
|
+
level,
|
|
350
|
+
expected: expectedDefects.length,
|
|
351
|
+
found,
|
|
352
|
+
missed,
|
|
353
|
+
judgedNotFiled: [...judged],
|
|
354
|
+
beyondLevel,
|
|
355
|
+
findings: findings.length,
|
|
356
|
+
correct,
|
|
357
|
+
falsePositives,
|
|
358
|
+
unknown,
|
|
359
|
+
ambiguous,
|
|
360
|
+
duplicates,
|
|
361
|
+
severity,
|
|
362
|
+
calibration: calibrateAgainstKey(decisions, key),
|
|
363
|
+
};
|
|
364
|
+
}
|
|
365
|
+
const pct = (n, d) => (d === 0 ? "—" : `${Math.round((n / d) * 100)}%`);
|
|
366
|
+
/**
|
|
367
|
+
* Precision, with its bounds. The labelled ratio leaves unlabelled findings out
|
|
368
|
+
* of the denominator, so on its own it cannot move when a change adds five new
|
|
369
|
+
* false claims nobody has judged yet. The bounds can: the lower one counts
|
|
370
|
+
* every open finding as wrong, the upper one as right.
|
|
371
|
+
*/
|
|
372
|
+
export function precisionBounds(c) {
|
|
373
|
+
const labelled = c.correct + c.falsePositives.length;
|
|
374
|
+
const open = c.unknown.length + c.ambiguous.length;
|
|
375
|
+
return { labelled: `${c.correct}/${labelled} (${pct(c.correct, labelled)})`, low: pct(c.correct, c.findings), high: pct(c.correct + open, c.findings) };
|
|
376
|
+
}
|
|
377
|
+
/** The scorecard as a person reads it. Leads with the two numbers, then says what each is made of. */
|
|
378
|
+
export function formatScorecard(c) {
|
|
379
|
+
const p = precisionBounds(c);
|
|
380
|
+
const open = c.unknown.length + c.ambiguous.length;
|
|
381
|
+
const lines = [
|
|
382
|
+
`SCORECARD — ${c.app}, level ${c.level}, key ${c.key}`,
|
|
383
|
+
``,
|
|
384
|
+
`Recall ${c.found.length}/${c.expected} (${pct(c.found.length, c.expected)}) of the planted defects expected at this level`,
|
|
385
|
+
`Precision ${p.labelled} of the findings the key can label` +
|
|
386
|
+
(open ? ` — ${open} of ${c.findings} unlabelled, so between ${p.low} and ${p.high} of all findings` : ""),
|
|
387
|
+
``,
|
|
388
|
+
];
|
|
389
|
+
if (c.missed.length)
|
|
390
|
+
lines.push(`Missed: ${c.missed.join(", ")}`);
|
|
391
|
+
if (c.judgedNotFiled.length)
|
|
392
|
+
lines.push(`Judged a defect in a lane report, never filed: ${c.judgedNotFiled.join(", ")}`);
|
|
393
|
+
if (c.beyondLevel.length)
|
|
394
|
+
lines.push(`Found above this level's contract: ${c.beyondLevel.join(", ")}`);
|
|
395
|
+
if (c.falsePositives.length) {
|
|
396
|
+
lines.push(``, `False positives (${c.falsePositives.length}):`);
|
|
397
|
+
for (const fp of c.falsePositives)
|
|
398
|
+
lines.push(` [${fp.severity}] ${fp.title} — ${fp.why}`);
|
|
399
|
+
}
|
|
400
|
+
if (c.ambiguous.length) {
|
|
401
|
+
lines.push(``, `Ambiguous (${c.ambiguous.length}) — the key claims each twice; sharpen it:`);
|
|
402
|
+
for (const a of c.ambiguous)
|
|
403
|
+
lines.push(` ${a.title} (${a.ids.join(" / ")})`);
|
|
404
|
+
}
|
|
405
|
+
if (c.unknown.length) {
|
|
406
|
+
lines.push(``, `Unlabelled (${c.unknown.length}) — add to the key once a person has judged them:`);
|
|
407
|
+
for (const u of c.unknown)
|
|
408
|
+
lines.push(` [${u.severity}] ${u.title}`);
|
|
409
|
+
}
|
|
410
|
+
const s = c.severity;
|
|
411
|
+
lines.push(``, `Severity vs the key's judgement: ${s.agree} agree, ${s.higher} filed higher, ${s.lower} filed lower` +
|
|
412
|
+
(s.detail.some((d) => d.filed !== d.key)
|
|
413
|
+
? ` (${s.detail
|
|
414
|
+
.filter((d) => d.filed !== d.key)
|
|
415
|
+
.map((d) => `${d.id} ${d.filed}→${d.key}`)
|
|
416
|
+
.join(", ")})`
|
|
417
|
+
: ""));
|
|
418
|
+
if (c.duplicates)
|
|
419
|
+
lines.push(`Duplicates: at least ${c.duplicates} extra finding(s) for a defect that already had one (after the store's own merge)`);
|
|
420
|
+
const k = c.calibration;
|
|
421
|
+
if (k) {
|
|
422
|
+
const skipped = [
|
|
423
|
+
k.notInKey && `${k.notInKey} about things the key does not name`,
|
|
424
|
+
k.ambiguous && `${k.ambiguous} the key names ambiguously`,
|
|
425
|
+
k.badConfidence && `${k.badConfidence} with an unusable confidence`,
|
|
426
|
+
k.unsure && `${k.unsure} unsure`,
|
|
427
|
+
k.outOfScope && `${k.outOfScope} dismissed as another lane's`,
|
|
428
|
+
].filter(Boolean);
|
|
429
|
+
lines.push(``, k.judged === 0
|
|
430
|
+
? `Lane calibration against the key: nothing the key could judge` + (skipped.length ? ` (${skipped.join(", ")})` : "")
|
|
431
|
+
: `Lane calibration against the key: ${k.correct}/${k.judged} verdicts right (${pct(k.correct, k.judged)}), expected calibration error ${k.ece.toFixed(2)}, Brier ${k.brier.toFixed(3)}` +
|
|
432
|
+
(skipped.length ? ` — not scored: ${skipped.join(", ")}` : ""));
|
|
433
|
+
for (const b of k.buckets)
|
|
434
|
+
lines.push(` stated ${b.label}: ${b.decisions} decision(s), said ${b.stated.toFixed(2)}, right ${Math.round(b.correct * 100)}%`);
|
|
435
|
+
}
|
|
436
|
+
return lines.join("\n");
|
|
437
|
+
}
|
|
438
|
+
/**
|
|
439
|
+
* The date a run is archived under. A run's decisions carry the time each was
|
|
440
|
+
* made, so the date defaults to the last of them, and a date given by hand
|
|
441
|
+
* may not be later: `--all` sorts by it, and a run dated after it happened
|
|
442
|
+
* reads as the newer of two in the results log. A run with no decisions has
|
|
443
|
+
* nothing to check against and must be given one.
|
|
444
|
+
*/
|
|
445
|
+
export function runDate(decisions, given) {
|
|
446
|
+
if (given !== undefined && !/^\d{4}-\d{2}-\d{2}$/.test(given))
|
|
447
|
+
throw new Error(`The date must be YYYY-MM-DD, not ${JSON.stringify(given)}.`);
|
|
448
|
+
const last = decisions
|
|
449
|
+
.map((d) => d.at)
|
|
450
|
+
.filter((at) => typeof at === "string" && /^\d{4}-\d{2}-\d{2}/.test(at))
|
|
451
|
+
.sort()
|
|
452
|
+
.at(-1)
|
|
453
|
+
?.slice(0, 10);
|
|
454
|
+
if (!last) {
|
|
455
|
+
if (given === undefined)
|
|
456
|
+
throw new Error("The run has no timestamped decisions to date it by; give it a date.");
|
|
457
|
+
return given;
|
|
458
|
+
}
|
|
459
|
+
if (given !== undefined && given > last)
|
|
460
|
+
throw new Error(`The run's last decision was made on ${last}; it cannot be dated ${given}.`);
|
|
461
|
+
return given ?? last;
|
|
462
|
+
}
|
|
463
|
+
/**
|
|
464
|
+
* Paths into a machine's own directories, which a finding's evidence can pick
|
|
465
|
+
* up from a stack trace or an upload. An archive is committed, and a home
|
|
466
|
+
* directory names a person.
|
|
467
|
+
*/
|
|
468
|
+
const LOCAL_PATH_RE = /(?:\/Users\/[^/\s"']+|\/home\/[^/\s"']+|\/private\/tmp|\/tmp|[A-Za-z]:\\Users\\[^\\\s"']+)[^\s"']*/g;
|
|
469
|
+
export function sanitize(text) {
|
|
470
|
+
return text.replace(LOCAL_PATH_RE, "<path>");
|
|
471
|
+
}
|
|
472
|
+
export function toArchive(run, date, note, findings, decisions) {
|
|
473
|
+
return {
|
|
474
|
+
run,
|
|
475
|
+
date,
|
|
476
|
+
note,
|
|
477
|
+
findings: findings.map((f) => ({
|
|
478
|
+
title: sanitize(f.title),
|
|
479
|
+
severity: f.severity,
|
|
480
|
+
...(f.category ? { category: f.category } : {}),
|
|
481
|
+
...(f.evidence ? { evidence: sanitize(f.evidence) } : {}),
|
|
482
|
+
})),
|
|
483
|
+
decisions: decisions.map((d) => ({ ...d, observation: sanitize(d.observation), evidence: d.evidence === null ? null : sanitize(d.evidence) })),
|
|
484
|
+
};
|
|
485
|
+
}
|