dsh-vet 0.3.0 → 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -34,12 +34,32 @@ npx dsh-vet --json <specifier> # dsh-vet/v1 report on stdout
34
34
  npx dsh-vet --strict <specifier> # exit 1 on findings >= high (confidence >= medium)
35
35
  npx dsh-vet --rules dep.install-scripts <specifier>
36
36
  npx dsh-vet validate <report.json> # check a report against the contract
37
+ npx dsh-vet diff base.report.json head.report.json --json
37
38
  ```
38
39
 
39
40
  Any completed report exits `0` — grades describe findings, they do not gate.
40
41
  Scanner failures exit non-zero. The scanner runs locally, reads the npm
41
42
  registry for dependency metadata only, and never transmits audited code.
42
43
 
44
+ Every report records what the scan actually covered (`x-dsh-vet` scan
45
+ context: rule profile, coverage, content digests, stable finding
46
+ identities); human output and badges show coverage alongside the grade,
47
+ and a missing context reads as `unknown`, never complete.
48
+
49
+ ### Comparing two reports (release review)
50
+
51
+ `dsh-vet diff` compares two reports produced by the same scanner version
52
+ and rule profile over complete scans, and reports what changed:
53
+ [`dsh-vet/diff/v1`](docs/report-diff-v1.md) — added/removed/changed
54
+ findings by stable identity, behavior-observation deltas, both sides of
55
+ every transition. Exit `0` for a comparable result regardless of risk
56
+ changes, `1` when the pair cannot be trusted to describe the same
57
+ subject under the same checks (with explicit reasons), `2` on usage or
58
+ invalid reports. Two local-directory scans additionally need
59
+ `--subject <label>`. An incomparable pair is a prompt to rescan both
60
+ artifacts with the same configuration — never a claim that nothing
61
+ changed.
62
+
43
63
  Shipped rules (each with a public rationale under
44
64
  [`docs/rules/`](docs/rules)):
45
65
 
@@ -64,13 +84,19 @@ from your repo, so its value is auditable through git history and no badge
64
84
  service is involved:
65
85
 
66
86
  ```yaml
67
- - uses: rogerdigital/dsh-vet/action@v0.2.0
87
+ - uses: rogerdigital/dsh-vet/action@v0.4.0
68
88
  with:
69
89
  specifier: '.'
70
90
  commit-report: true
71
91
  ```
72
92
 
73
93
  Every run uploads the full report as an artifact; PRs get a single
94
+ comment, edited in place. Set `baseline-report` to a report scanned from
95
+ the merge base and PRs additionally show what changed since it — see
96
+ [comparing releases](action/README.md#comparing-releases) for the
97
+ trusted-baseline recipe, and the
98
+ [pilot record](docs/release-risk-pilot.md) for how the comparison
99
+ behaved on real release pairs.
74
100
  edited-in-place findings comment. Badge snippet and all inputs:
75
101
  [`action/README.md`](action/README.md). The `dsh-vet badge <report.json>`
76
102
  command renders the shields endpoint JSON if you wire CI yourself.
@@ -48,9 +48,17 @@ grade from the findings, so an emitter cannot assert a grade its evidence
48
48
  does not support:
49
49
 
50
50
  ```sh
51
- npx dsh-vet validate report.json && echo trustworthy-shape
51
+ npx dsh-vet validate report.json && echo structurally-conformant
52
52
  ```
53
53
 
54
+ `validate` proves structural conformance only. It does not prove
55
+ completeness (that the emitter ran every check and omitted nothing),
56
+ artifact identity (that the report describes the plugin version you are
57
+ serving), or provenance (who produced the report). For those guarantees,
58
+ fetch reports through channels you control — the Action's report branch in
59
+ the author's own repo, or a scan you run yourself — and treat reports from
60
+ unverified emitters as unreviewed claims.
61
+
54
62
  ## Consumer rules (from the spec)
55
63
 
56
64
  - **Ignore unknown fields** — emitters may add `x-`-prefixed extras; tolerate
@@ -101,13 +101,43 @@ initially broke vendor-prefixed check ids — do not repeat that.
101
101
 
102
102
  | Grade | Condition (over findings with confidence ≥ `medium`) |
103
103
  |---|---|
104
- | `A` | none, or only `info`/`low`-severity findings |
104
+ | `A` | no graded findings — the report is empty, or every finding is `info`-severity or `low`-confidence |
105
105
  | `B` | worst graded finding is `low` |
106
106
  | `C` | worst graded finding is `medium` |
107
107
  | `D` | worst graded finding is `high` |
108
108
  | `F` | at least one `critical` |
109
109
  | `X` | scan incomplete or errored — never presented as the plugin's grade |
110
110
 
111
+ ## What validation proves — and what it does not
112
+
113
+ `dsh-vet validate` checks **structural conformance** with this contract:
114
+ field types and enums, rule-id shape, evidence presence, the deterministic
115
+ sort, and — the load-bearing check — that `summary` is exactly what the
116
+ `findings` derive. A conformant report cannot assert a grade its evidence
117
+ does not support.
118
+
119
+ It does not, and cannot, verify:
120
+
121
+ - **Emitter honesty** — whether the emitter omitted findings or fabricated
122
+ evidence. Structural consistency is not evidence that the audit was
123
+ thorough.
124
+ - **Completeness** — v1 records nothing about which files or rules a scan
125
+ covered. A subset scan's report is structurally indistinguishable from a
126
+ full one.
127
+ - **Artifact identity** — nothing ties a report to the artifact you are
128
+ holding. `target.resolved.integrity` is recorded by the emitter, not
129
+ checked against your copy.
130
+ - **Provenance** — who actually produced the report, and whether it changed
131
+ in transit.
132
+
133
+ Those guarantees come from delivery channels, not from the report's shape:
134
+ see [emitters.md](emitters.md) for how emitters are verified and
135
+ [adopt-marketplace.md](adopt-marketplace.md) for consumer guidance. The
136
+ optional scan-context extension ([scan-context-v1.md](scan-context-v1.md))
137
+ records coverage and content identity; the reference scanner begins
138
+ emitting it in a later release, and until a report carries it, absent
139
+ coverage data means **unknown**, never complete.
140
+
111
141
  ## CLI recommendations (non-normative)
112
142
 
113
143
  Emitters that ship a CLI should exit `0` whenever a report was produced —
package/docs/emitters.md CHANGED
@@ -17,7 +17,9 @@ honesty.
17
17
  `dsh-vet validate` (`npx dsh-vet validate <report.json>`). Either path
18
18
  guarantees the derived summary, the deterministic sort, and well-formed
19
19
  rule ids; an emitter that hand-assembles reports and skips validation is
20
- not verified.
20
+ not verified. Validation proves structure, not honesty — it cannot
21
+ detect omitted findings, which is why the remaining items on this list
22
+ exist.
21
23
  2. **Determinism.** Two runs over the same artifact with the same emitter
22
24
  version produce identical reports, `scanner.ranAt` aside.
23
25
  3. **Conservative severity.** Findings follow the severity ladder in the
@@ -1,12 +1,15 @@
1
1
  # Outreach drafts
2
2
 
3
- Ready-to-send drafts for the v0.3 announcement round. Each file names its
4
- destination. None are sent yet — send from the maintainer's account, then
5
- record the thread link in ROADMAP.md v0.3 so feedback has a traceable home.
3
+ Ready-to-send drafts. Each file names its destination. The v0.3
4
+ announcement round went live on 2026-09-02 (thread links in ROADMAP.md);
5
+ the release-risk round below is pending. Send from the maintainer's
6
+ account, then record the thread link so feedback has a traceable home.
6
7
 
7
8
  | File | Destination | Depends on |
8
9
  |---|---|---|
9
- | [deepseek-harness-1115-reply.md](deepseek-harness-1115-reply.md) | reply in [deepseek-harness#1115](https://github.com/deepseek-ai/deepseek-harness/discussions/1115) | npm 0.3.0 published |
10
- | [show-your-plugins-post.md](show-your-plugins-post.md) | community "Show Your Plugins" thread | 0.3.0 published |
11
- | [awesome-list-pr.md](awesome-list-pr.md) | PR against a dsh/agent awesome-list | — |
12
- | [dsh-plugin-audit-collab.md](dsh-plugin-audit-collab.md) | issue/DM to dsh-plugin-audit maintainers | — |
10
+ | [deepseek-harness-1115-reply.md](deepseek-harness-1115-reply.md) | reply in [deepseek-harness#1115](https://github.com/deepseek-ai/deepseek-harness/discussions/1115) | npm 0.3.0 published — **sent 2026-09-02** |
11
+ | [show-your-plugins-post.md](show-your-plugins-post.md) | community "Show Your Plugins" thread | 0.3.0 published — **sent 2026-09-02** |
12
+ | [awesome-list-pr.md](awesome-list-pr.md) | PR against a dsh/agent awesome-list | — **merged 2026-09-03** |
13
+ | [dsh-plugin-audit-collab.md](dsh-plugin-audit-collab.md) | issue/DM to dsh-plugin-audit maintainers | — **sent 2026-09-02** |
14
+ | [release-risk-1115-followup.md](release-risk-1115-followup.md) | follow-up in [deepseek-harness#1115](https://github.com/deepseek-ai/deepseek-harness/discussions/1115) | next npm release ships `dsh-vet diff` |
15
+ | [release-risk-pilot-ask.md](release-risk-pilot-ask.md) | issue/DM per target plugin author (template with pre-run results) | — |
@@ -0,0 +1,37 @@
1
+ <!-- Destination: follow-up comment in deepseek-ai/deepseek-harness discussion #1115
2
+ (marketplace standards / review mechanisms), under our 2026-09-02 reply.
3
+ Depends on: an npm release that ships `dsh-vet diff` (next version) —
4
+ send from the maintainer's account, then record the thread link in ROADMAP. -->
5
+
6
+ Following up with the piece the contract was missing: **what changed since
7
+ the last release**.
8
+
9
+ "Should I install this plugin?" got a machine-readable answer with
10
+ `dsh-vet/v1`. "Should I *upgrade*?" now has one too:
11
+
12
+ - Every report records **what the scan actually covered** — rule profile,
13
+ coverage, content digests, stable finding identities — as an optional
14
+ `x-dsh-vet` extension ([spec](https://github.com/rogerdigital/dsh-vet/blob/main/docs/scan-context-v1.md)).
15
+ A grade no longer silently means "whatever subset happened to run":
16
+ partial coverage says so, in the CLI, the PR comment, and the badge.
17
+ - `dsh-vet diff` compares two reports from the same scanner and profile
18
+ and reports added / removed / changed risk by **stable identity** —
19
+ line moves and reformatting are not risk additions, a second endpoint
20
+ is ([diff spec](https://github.com/rogerdigital/dsh-vet/blob/main/docs/report-diff-v1.md)).
21
+ Two reports that can't be honestly compared say so with explicit
22
+ reasons instead of a fake "no changes".
23
+ - The GitHub Action takes an optional `baseline-report` scanned from the
24
+ merge base, so PRs show exactly what risk-relevant behavior the PR
25
+ introduces — with a recipe that keeps the baseline out of the PR's
26
+ control ([action docs](https://github.com/rogerdigital/dsh-vet/blob/main/action/README.md#comparing-releases)).
27
+
28
+ We piloted it on real releases —
29
+ [left-pad and chalk, every reported addition reconciled against the
30
+ code](https://github.com/rogerdigital/dsh-vet/blob/main/docs/release-risk-pilot.md).
31
+ The chalk case is the interesting one: 5.x vendored its dependencies and
32
+ the diff flagged exactly the two vendored files unreachable through the
33
+ `#`-imports map — nothing else moved.
34
+
35
+ If you maintain a plugin and want the same read on your last two
36
+ releases, reply here or open an issue — we'll run the comparison and
37
+ post the result for you to check against what you actually changed.
@@ -0,0 +1,71 @@
1
+ <!-- Destination: issue (or DM) to the author of each target plugin, one
2
+ thread per plugin. The template below is filled with real pre-run
3
+ results — send as-is; the recipient installs nothing and answers from
4
+ their own knowledge of the release. Send from the maintainer's account,
5
+ then record each thread link in docs/release-risk-pilot.md. -->
6
+
7
+ # Release-risk pilot ask (template)
8
+
9
+ Subject: `Did your last release change what your plugin can do? (2-minute check)`
10
+
11
+ Hi — I maintain [`dsh-vet`](https://github.com/rogerdigital/dsh-vet), the
12
+ static audit tool for DSH plugins. It just learned to **compare two
13
+ releases of the same plugin and report exactly what risk-relevant
14
+ behavior changed** — new endpoints, new capabilities, new install
15
+ scripts, changed confidence — while ignoring line moves and reformatting.
16
+
17
+ I ran it on your last two releases so you don't have to install anything.
18
+ The results are below; three questions at the end.
19
+
20
+ ---
21
+
22
+ ## Filled example — dsh-doctor 0.4.2 → 0.4.3
23
+
24
+ Both scans: same scanner, full rule profile, complete coverage. Grades
25
+ stayed **D → D**. What moved:
26
+
27
+ - **New network client call** in `lib/client.js` — the destination is
28
+ computed at runtime, so the scanner can't see where it sends
29
+ (`perm.network-client`, subject `runtime-target`).
30
+ - **Subprocess spawns went from 1 to 2** in `lib/index.js`
31
+ (`child_process.spawn`).
32
+
33
+ Nothing else changed (13 identities unchanged).
34
+
35
+ ## Filled example — dsh-searxng 0.2.1 → 0.3.0
36
+
37
+ Grades stayed **A → A**. One transition: **write/delete calls with
38
+ runtime-computed targets went from 1 to 17** in `lib/cli.mjs`
39
+ (`perm.undeclared-fs-write`, dynamic variant). These are low-severity /
40
+ low-confidence — they never affect the grade — but the count jump is the
41
+ kind of thing worth a look during a minor bump.
42
+
43
+ ## Filled example — dsh-wechat 0.9.5 → 0.9.6
44
+
45
+ **Zero risk-relevant changes.** All 7 audited behaviors identical, grade
46
+ C → C. If that matches your intent for the release, the diff just saved
47
+ you the re-review.
48
+
49
+ ---
50
+
51
+ ## The three questions
52
+
53
+ 1. Does the reported delta match what you intended to change in that
54
+ release? Anything you'd flag that the scanner missed or got wrong?
55
+ 2. Any entries that felt like noise (didn't help you understand the
56
+ release)?
57
+ 3. If this ran automatically on your PRs against the base revision —
58
+ [one Action input](https://github.com/rogerdigital/dsh-vet/blob/main/action/README.md#comparing-releases) —
59
+ would you read it before merging?
60
+
61
+ Answers in any form are useful; I'll record anonymized takeaways in the
62
+ [pilot record](https://github.com/rogerdigital/dsh-vet/blob/main/docs/release-risk-pilot.md).
63
+ This is the last feedback loop before deciding whether to build CI
64
+ gating on top (opt-in "fail on new risk"), so negative answers are as
65
+ valuable as positive ones.
66
+
67
+ <!-- Per-send checklist:
68
+ - re-run the pair with the latest main build the day of sending
69
+ - paste the actual result block for THIS author's package
70
+ - keep grades/identities verbatim from the diff, don't paraphrase
71
+ - one package per thread; don't batch multiple authors -->
@@ -0,0 +1,109 @@
1
+ # Release-risk review pilot (M1)
2
+
3
+ Status: **internal pilot complete; external feedback pending.** This page
4
+ records how the first milestone was exercised and what the comparison
5
+ showed on real release pairs. It is not an adoption claim — no external
6
+ maintainer has used the workflow yet, and the M2 entry gate (repeated
7
+ pilot use showing review friction) is not met.
8
+
9
+ ## What was exercised
10
+
11
+ | Exercise | Result |
12
+ |---|---|
13
+ | Controlled pair (committed fixtures, one added host) | Comparable; exactly the added-host identities and observation; grades equal |
14
+ | Built-output smoke (`node bin/dsh-vet.mjs`, section 8 of the plan) | Help, scan ×2, validate ×2, diff, badge all pass; same-profile comparison of two runs reports zero risk changes despite distinct `ranAt`; badge unqualified for complete coverage |
15
+ | `left-pad@1.2.0` → `1.3.0` | Comparable, zero added/removed/changed — reconciled below |
16
+ | `chalk@4.1.2` → `5.3.0` | Comparable, two added `perm.unreachable-files` identities — reconciled below |
17
+
18
+ All pairs were scanned with the same scanner build and full rule profile;
19
+ both sides validated with `dsh-vet validate` before comparison.
20
+
21
+ ## Reconciliation: left-pad 1.2.0 → 1.3.0
22
+
23
+ Both versions: identical audited behavior. The five unchanged identities
24
+ cover the benchmark/test files that ship unreachably (`perf/*`, `test.js`)
25
+ and the hex table in `perf/perf.js` (`obf.encoded-payload`, matched by
26
+ content hash) — all present in both versions, none moved. **Zero
27
+ additions to reconcile; the release changed padding internals without
28
+ touching any audited behavior.** Effort: under a minute; the diff
29
+ answered "anything new to review?" with a definite no.
30
+
31
+ ## Reconciliation: chalk 4.1.2 → 5.3.0
32
+
33
+ Reported additions, each verified against the shipped tarballs:
34
+
35
+ 1. `perm.unreachable-files` · `source/vendor/supports-color/index.js`
36
+ 2. `perm.unreachable-files` · `source/vendor/supports-color/browser.js`
37
+
38
+ Code evidence:
39
+
40
+ - 4.1.2 has no `source/vendor/` directory at all; 5.3.0 vendors its
41
+ dependencies there.
42
+ - 5.3.0's `source/index.js` imports `#supports-color`, which resolves
43
+ through the package.json `imports` map to exactly those two files
44
+ (`node` → `index.js`, `default` → `browser.js`).
45
+ - The reference analyzer resolves relative imports but not `#`-prefixed
46
+ subpath imports, so neither file is statically reachable from the
47
+ declared entry — the finding text ("no static import path from any
48
+ entry point reaches this file") is accurate for this analyzer.
49
+ - The sibling vendored `ansi-styles` **is** imported relatively
50
+ (`'./vendor/ansi-styles/index.js'`), is reachable, and is correctly
51
+ **not** reported — the delta is per-subject, not per-directory.
52
+
53
+ Both grades stay `A` (unreachable files are informational); the delta is
54
+ visible precisely because observations and identities are reported even
55
+ when grades do not move.
56
+
57
+ Honest limitation observed: `#`-subpath `imports` maps are not resolved,
58
+ so code reached only through them reads as unreachable. That is a
59
+ documented analyzer gap to close in a later revision — visible here
60
+ because the delta surfaces it, which is the point of the workflow.
61
+
62
+ ## Pilot observations
63
+
64
+ - The two real pairs took minutes to reconcile; every reported addition
65
+ mapped to a code fact, no formatting-driven false additions appeared
66
+ (line movement is excluded from identity by design and it held).
67
+ - The interesting noise class is informational unreachable-file churn on
68
+ refactors that vendor or move code — visible, cheap to dismiss, and
69
+ exactly what M2's exception mechanism would absorb if it repeats.
70
+ - One pilot round is not enough to justify policy work. The M2 entry
71
+ gate stays closed until a maintainer uses the comparison on real
72
+ releases repeatedly and reports the friction.
73
+
74
+ ## External feedback (pending)
75
+
76
+ No external maintainer has used the comparison yet — this section stays
77
+ empty until one does, and a friendly reply is not adoption. The ask and
78
+ its per-plugin pre-run results live in
79
+ [`docs/outreach/release-risk-pilot-ask.md`](outreach/release-risk-pilot-ask.md);
80
+ record each thread here as it happens.
81
+
82
+ Intake questions (mirroring the plan's §10 usefulness criteria):
83
+
84
+ 1. Does the reported delta match what the author intended to change?
85
+ 2. Anything the scanner missed or got wrong?
86
+ 3. Any entries that felt like noise?
87
+ 4. Would they read it on their PRs — and did they run it on a *second*
88
+ release? (The repeated-use answer is the M2 entry gate.)
89
+
90
+ | Date | Plugin | Thread | Verdict (used / not / second release) |
91
+ |---|---|---|---|
92
+ | — | — | — | — |
93
+
94
+ ## Reproduction
95
+
96
+ ```sh
97
+ pnpm build
98
+ node bin/dsh-vet.mjs --json left-pad@1.2.0 > /tmp/lp-base.json
99
+ node bin/dsh-vet.mjs --json left-pad@1.3.0 > /tmp/lp-head.json
100
+ node bin/dsh-vet.mjs validate /tmp/lp-base.json /tmp/lp-head.json
101
+ node bin/dsh-vet.mjs diff --json /tmp/lp-base.json /tmp/lp-head.json
102
+
103
+ node bin/dsh-vet.mjs --json chalk@4.1.2 > /tmp/ch-base.json
104
+ node bin/dsh-vet.mjs --json chalk@5.3.0 > /tmp/ch-head.json
105
+ node bin/dsh-vet.mjs diff --json /tmp/ch-base.json /tmp/ch-head.json
106
+ ```
107
+
108
+ The controlled pair lives in `test/fixtures/release-risk/` and is
109
+ asserted by `test/golden.test.ts` on every run.
@@ -0,0 +1,80 @@
1
+ # The `dsh-vet/diff/v1` comparison result
2
+
3
+ Status: **defined; the reference CLI exposes it from the next release.**
4
+ This document is normative for consumers of `compareReports()` and the
5
+ `dsh-vet diff` output. See [dsh-vet-v1.md](dsh-vet-v1.md) for the base
6
+ report and [scan-context-v1.md](scan-context-v1.md) for the extension the
7
+ comparator reads.
8
+
9
+ `compareReports(base, head)` is pure data in, pure data out: it never
10
+ fetches packages, never mutates either report, and never recalculates a
11
+ report under new rules. It compares what two validated reports actually
12
+ recorded.
13
+
14
+ ## Comparability
15
+
16
+ A pair is **comparable** only when every gate passes:
17
+
18
+ | Gate | Reason code when it fails |
19
+ |---|---|
20
+ | Both are structurally valid `dsh-vet/v1` reports | `report-invalid-base` / `report-invalid-head` |
21
+ | Both carry a validated version-1 `x-dsh-vet` context | `context-absent-*`, `context-unsupported-*`, `context-invalid-*` |
22
+ | Both carry finding identities | `identities-absent-base` / `identities-absent-head` |
23
+ | Same scanner name and version | `scanner-mismatch` |
24
+ | Same profile digest (analyzer, catalog, rules, options) | `profile-mismatch` |
25
+ | Same package identity, or an explicit shared subject label | `subject-mismatch` / `subject-unlabeled` |
26
+ | Neither side graded `X` | `grade-x-base` / `grade-x-head` |
27
+ | Complete coverage on both sides | `coverage-partial-base` / `coverage-partial-head` |
28
+
29
+ An incomparable result carries the reason codes and both report
30
+ references, and **no change collections** — an absent diff must never
31
+ read as "nothing changed". There is no force-comparison switch in M1;
32
+ when practical, rescan both artifacts with the same scanner and profile
33
+ instead. Two local-directory scans without package identity compare only
34
+ under an explicit `subjectLabel` option; identity is never inferred from
35
+ temporary absolute paths. Different analysis-input or archive digests
36
+ are expected and never block comparison.
37
+
38
+ ## Result shape
39
+
40
+ ```jsonc
41
+ {
42
+ "schema": "dsh-vet/diff/v1",
43
+ "base": { "scanner": { /* name, version, ranAt */ }, "grade": "C", "subject": { /* … */ } },
44
+ "head": { "scanner": { /* … */ }, "grade": "C", "subject": { /* … */ } },
45
+ "comparability": "comparable",
46
+ "reasons": [],
47
+ "findings": {
48
+ "added": [ /* identities present only in head */ ],
49
+ "removed": [ /* identities present only in base */ ],
50
+ "changed": [ /* matched identities that moved; both sides recorded */ ],
51
+ "unchanged": [ /* matched and identical */ ]
52
+ },
53
+ "observations": {
54
+ "added": [], "removed": [], "countChanged": [], "unchanged": []
55
+ }
56
+ }
57
+ ```
58
+
59
+ Identity entries and observations are the extension's own objects
60
+ (`rule`/`variant`/`file`/`subject`/`finding`/`count?` and
61
+ `kind`/`file`/`subject`/`count?`). Every collection is sorted and
62
+ deterministic: two comparisons of the same pair produce byte-identical
63
+ results.
64
+
65
+ A **changed** finding records both sides' severity, confidence, and
66
+ count. Severity and confidence transitions are classified separately
67
+ from additions on purpose: moving from `low` to `medium` confidence can
68
+ introduce a graded risk at unchanged severity.
69
+
70
+ ## Reading the result honestly
71
+
72
+ - `added` means newly observed under the same checks — not proof of new
73
+ malice, and not a grade change by itself.
74
+ - `removed` means no longer observed under comparable checks — never a
75
+ claim that anything was fixed.
76
+ - Observation deltas are visible even when grades are identical or the
77
+ observations are informational; an unchanged grade never hides new
78
+ behavior.
79
+ - The comparator reports line-position changes as nothing: identities
80
+ exclude line numbers, so a line insertion is not a risk addition.