dsh-vet 0.3.0 → 0.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +27 -1
- package/docs/adopt-marketplace.md +9 -1
- package/docs/dsh-vet-v1.md +31 -1
- package/docs/emitters.md +3 -1
- package/docs/outreach/README.md +10 -7
- package/docs/outreach/release-risk-1115-followup.md +37 -0
- package/docs/outreach/release-risk-pilot-ask.md +71 -0
- package/docs/release-risk-pilot.md +109 -0
- package/docs/report-diff-v1.md +80 -0
- package/docs/scan-context-v1.md +218 -0
- package/docs/superpowers/plans/2026-09-10-release-risk-development-plan.md +397 -0
- package/lib/index.d.mts +330 -38
- package/lib/index.mjs +1167 -51
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -34,12 +34,32 @@ npx dsh-vet --json <specifier> # dsh-vet/v1 report on stdout
|
|
|
34
34
|
npx dsh-vet --strict <specifier> # exit 1 on findings >= high (confidence >= medium)
|
|
35
35
|
npx dsh-vet --rules dep.install-scripts <specifier>
|
|
36
36
|
npx dsh-vet validate <report.json> # check a report against the contract
|
|
37
|
+
npx dsh-vet diff base.report.json head.report.json --json
|
|
37
38
|
```
|
|
38
39
|
|
|
39
40
|
Any completed report exits `0` — grades describe findings, they do not gate.
|
|
40
41
|
Scanner failures exit non-zero. The scanner runs locally, reads the npm
|
|
41
42
|
registry for dependency metadata only, and never transmits audited code.
|
|
42
43
|
|
|
44
|
+
Every report records what the scan actually covered (`x-dsh-vet` scan
|
|
45
|
+
context: rule profile, coverage, content digests, stable finding
|
|
46
|
+
identities); human output and badges show coverage alongside the grade,
|
|
47
|
+
and a missing context reads as `unknown`, never complete.
|
|
48
|
+
|
|
49
|
+
### Comparing two reports (release review)
|
|
50
|
+
|
|
51
|
+
`dsh-vet diff` compares two reports produced by the same scanner version
|
|
52
|
+
and rule profile over complete scans, and reports what changed:
|
|
53
|
+
[`dsh-vet/diff/v1`](docs/report-diff-v1.md) — added/removed/changed
|
|
54
|
+
findings by stable identity, behavior-observation deltas, both sides of
|
|
55
|
+
every transition. Exit `0` for a comparable result regardless of risk
|
|
56
|
+
changes, `1` when the pair cannot be trusted to describe the same
|
|
57
|
+
subject under the same checks (with explicit reasons), `2` on usage or
|
|
58
|
+
invalid reports. Two local-directory scans additionally need
|
|
59
|
+
`--subject <label>`. An incomparable pair is a prompt to rescan both
|
|
60
|
+
artifacts with the same configuration — never a claim that nothing
|
|
61
|
+
changed.
|
|
62
|
+
|
|
43
63
|
Shipped rules (each with a public rationale under
|
|
44
64
|
[`docs/rules/`](docs/rules)):
|
|
45
65
|
|
|
@@ -64,13 +84,19 @@ from your repo, so its value is auditable through git history and no badge
|
|
|
64
84
|
service is involved:
|
|
65
85
|
|
|
66
86
|
```yaml
|
|
67
|
-
- uses: rogerdigital/dsh-vet/action@v0.
|
|
87
|
+
- uses: rogerdigital/dsh-vet/action@v0.4.0
|
|
68
88
|
with:
|
|
69
89
|
specifier: '.'
|
|
70
90
|
commit-report: true
|
|
71
91
|
```
|
|
72
92
|
|
|
73
93
|
Every run uploads the full report as an artifact; PRs get a single
|
|
94
|
+
comment, edited in place. Set `baseline-report` to a report scanned from
|
|
95
|
+
the merge base and PRs additionally show what changed since it — see
|
|
96
|
+
[comparing releases](action/README.md#comparing-releases) for the
|
|
97
|
+
trusted-baseline recipe, and the
|
|
98
|
+
[pilot record](docs/release-risk-pilot.md) for how the comparison
|
|
99
|
+
behaved on real release pairs.
|
|
74
100
|
edited-in-place findings comment. Badge snippet and all inputs:
|
|
75
101
|
[`action/README.md`](action/README.md). The `dsh-vet badge <report.json>`
|
|
76
102
|
command renders the shields endpoint JSON if you wire CI yourself.
|
|
@@ -48,9 +48,17 @@ grade from the findings, so an emitter cannot assert a grade its evidence
|
|
|
48
48
|
does not support:
|
|
49
49
|
|
|
50
50
|
```sh
|
|
51
|
-
npx dsh-vet validate report.json && echo
|
|
51
|
+
npx dsh-vet validate report.json && echo structurally-conformant
|
|
52
52
|
```
|
|
53
53
|
|
|
54
|
+
`validate` proves structural conformance only. It does not prove
|
|
55
|
+
completeness (that the emitter ran every check and omitted nothing),
|
|
56
|
+
artifact identity (that the report describes the plugin version you are
|
|
57
|
+
serving), or provenance (who produced the report). For those guarantees,
|
|
58
|
+
fetch reports through channels you control — the Action's report branch in
|
|
59
|
+
the author's own repo, or a scan you run yourself — and treat reports from
|
|
60
|
+
unverified emitters as unreviewed claims.
|
|
61
|
+
|
|
54
62
|
## Consumer rules (from the spec)
|
|
55
63
|
|
|
56
64
|
- **Ignore unknown fields** — emitters may add `x-`-prefixed extras; tolerate
|
package/docs/dsh-vet-v1.md
CHANGED
|
@@ -101,13 +101,43 @@ initially broke vendor-prefixed check ids — do not repeat that.
|
|
|
101
101
|
|
|
102
102
|
| Grade | Condition (over findings with confidence ≥ `medium`) |
|
|
103
103
|
|---|---|
|
|
104
|
-
| `A` |
|
|
104
|
+
| `A` | no graded findings — the report is empty, or every finding is `info`-severity or `low`-confidence |
|
|
105
105
|
| `B` | worst graded finding is `low` |
|
|
106
106
|
| `C` | worst graded finding is `medium` |
|
|
107
107
|
| `D` | worst graded finding is `high` |
|
|
108
108
|
| `F` | at least one `critical` |
|
|
109
109
|
| `X` | scan incomplete or errored — never presented as the plugin's grade |
|
|
110
110
|
|
|
111
|
+
## What validation proves — and what it does not
|
|
112
|
+
|
|
113
|
+
`dsh-vet validate` checks **structural conformance** with this contract:
|
|
114
|
+
field types and enums, rule-id shape, evidence presence, the deterministic
|
|
115
|
+
sort, and — the load-bearing check — that `summary` is exactly what the
|
|
116
|
+
`findings` derive. A conformant report cannot assert a grade its evidence
|
|
117
|
+
does not support.
|
|
118
|
+
|
|
119
|
+
It does not, and cannot, verify:
|
|
120
|
+
|
|
121
|
+
- **Emitter honesty** — whether the emitter omitted findings or fabricated
|
|
122
|
+
evidence. Structural consistency is not evidence that the audit was
|
|
123
|
+
thorough.
|
|
124
|
+
- **Completeness** — v1 records nothing about which files or rules a scan
|
|
125
|
+
covered. A subset scan's report is structurally indistinguishable from a
|
|
126
|
+
full one.
|
|
127
|
+
- **Artifact identity** — nothing ties a report to the artifact you are
|
|
128
|
+
holding. `target.resolved.integrity` is recorded by the emitter, not
|
|
129
|
+
checked against your copy.
|
|
130
|
+
- **Provenance** — who actually produced the report, and whether it changed
|
|
131
|
+
in transit.
|
|
132
|
+
|
|
133
|
+
Those guarantees come from delivery channels, not from the report's shape:
|
|
134
|
+
see [emitters.md](emitters.md) for how emitters are verified and
|
|
135
|
+
[adopt-marketplace.md](adopt-marketplace.md) for consumer guidance. The
|
|
136
|
+
optional scan-context extension ([scan-context-v1.md](scan-context-v1.md))
|
|
137
|
+
records coverage and content identity; the reference scanner begins
|
|
138
|
+
emitting it in a later release, and until a report carries it, absent
|
|
139
|
+
coverage data means **unknown**, never complete.
|
|
140
|
+
|
|
111
141
|
## CLI recommendations (non-normative)
|
|
112
142
|
|
|
113
143
|
Emitters that ship a CLI should exit `0` whenever a report was produced —
|
package/docs/emitters.md
CHANGED
|
@@ -17,7 +17,9 @@ honesty.
|
|
|
17
17
|
`dsh-vet validate` (`npx dsh-vet validate <report.json>`). Either path
|
|
18
18
|
guarantees the derived summary, the deterministic sort, and well-formed
|
|
19
19
|
rule ids; an emitter that hand-assembles reports and skips validation is
|
|
20
|
-
not verified.
|
|
20
|
+
not verified. Validation proves structure, not honesty — it cannot
|
|
21
|
+
detect omitted findings, which is why the remaining items on this list
|
|
22
|
+
exist.
|
|
21
23
|
2. **Determinism.** Two runs over the same artifact with the same emitter
|
|
22
24
|
version produce identical reports, `scanner.ranAt` aside.
|
|
23
25
|
3. **Conservative severity.** Findings follow the severity ladder in the
|
package/docs/outreach/README.md
CHANGED
|
@@ -1,12 +1,15 @@
|
|
|
1
1
|
# Outreach drafts
|
|
2
2
|
|
|
3
|
-
Ready-to-send drafts
|
|
4
|
-
|
|
5
|
-
|
|
3
|
+
Ready-to-send drafts. Each file names its destination. The v0.3
|
|
4
|
+
announcement round went live on 2026-09-02 (thread links in ROADMAP.md);
|
|
5
|
+
the release-risk round below is pending. Send from the maintainer's
|
|
6
|
+
account, then record the thread link so feedback has a traceable home.
|
|
6
7
|
|
|
7
8
|
| File | Destination | Depends on |
|
|
8
9
|
|---|---|---|
|
|
9
|
-
| [deepseek-harness-1115-reply.md](deepseek-harness-1115-reply.md) | reply in [deepseek-harness#1115](https://github.com/deepseek-ai/deepseek-harness/discussions/1115) | npm 0.3.0 published |
|
|
10
|
-
| [show-your-plugins-post.md](show-your-plugins-post.md) | community "Show Your Plugins" thread | 0.3.0 published |
|
|
11
|
-
| [awesome-list-pr.md](awesome-list-pr.md) | PR against a dsh/agent awesome-list | — |
|
|
12
|
-
| [dsh-plugin-audit-collab.md](dsh-plugin-audit-collab.md) | issue/DM to dsh-plugin-audit maintainers | — |
|
|
10
|
+
| [deepseek-harness-1115-reply.md](deepseek-harness-1115-reply.md) | reply in [deepseek-harness#1115](https://github.com/deepseek-ai/deepseek-harness/discussions/1115) | npm 0.3.0 published — **sent 2026-09-02** |
|
|
11
|
+
| [show-your-plugins-post.md](show-your-plugins-post.md) | community "Show Your Plugins" thread | 0.3.0 published — **sent 2026-09-02** |
|
|
12
|
+
| [awesome-list-pr.md](awesome-list-pr.md) | PR against a dsh/agent awesome-list | — **merged 2026-09-03** |
|
|
13
|
+
| [dsh-plugin-audit-collab.md](dsh-plugin-audit-collab.md) | issue/DM to dsh-plugin-audit maintainers | — **sent 2026-09-02** |
|
|
14
|
+
| [release-risk-1115-followup.md](release-risk-1115-followup.md) | follow-up in [deepseek-harness#1115](https://github.com/deepseek-ai/deepseek-harness/discussions/1115) | next npm release ships `dsh-vet diff` |
|
|
15
|
+
| [release-risk-pilot-ask.md](release-risk-pilot-ask.md) | issue/DM per target plugin author (template with pre-run results) | — |
|
|
@@ -0,0 +1,37 @@
|
|
|
1
|
+
<!-- Destination: follow-up comment in deepseek-ai/deepseek-harness discussion #1115
|
|
2
|
+
(marketplace standards / review mechanisms), under our 2026-09-02 reply.
|
|
3
|
+
Depends on: an npm release that ships `dsh-vet diff` (next version) —
|
|
4
|
+
send from the maintainer's account, then record the thread link in ROADMAP. -->
|
|
5
|
+
|
|
6
|
+
Following up with the piece the contract was missing: **what changed since
|
|
7
|
+
the last release**.
|
|
8
|
+
|
|
9
|
+
"Should I install this plugin?" got a machine-readable answer with
|
|
10
|
+
`dsh-vet/v1`. "Should I *upgrade*?" now has one too:
|
|
11
|
+
|
|
12
|
+
- Every report records **what the scan actually covered** — rule profile,
|
|
13
|
+
coverage, content digests, stable finding identities — as an optional
|
|
14
|
+
`x-dsh-vet` extension ([spec](https://github.com/rogerdigital/dsh-vet/blob/main/docs/scan-context-v1.md)).
|
|
15
|
+
A grade no longer silently means "whatever subset happened to run":
|
|
16
|
+
partial coverage says so, in the CLI, the PR comment, and the badge.
|
|
17
|
+
- `dsh-vet diff` compares two reports from the same scanner and profile
|
|
18
|
+
and reports added / removed / changed risk by **stable identity** —
|
|
19
|
+
line moves and reformatting are not risk additions, a second endpoint
|
|
20
|
+
is ([diff spec](https://github.com/rogerdigital/dsh-vet/blob/main/docs/report-diff-v1.md)).
|
|
21
|
+
Two reports that can't be honestly compared say so with explicit
|
|
22
|
+
reasons instead of a fake "no changes".
|
|
23
|
+
- The GitHub Action takes an optional `baseline-report` scanned from the
|
|
24
|
+
merge base, so PRs show exactly what risk-relevant behavior the PR
|
|
25
|
+
introduces — with a recipe that keeps the baseline out of the PR's
|
|
26
|
+
control ([action docs](https://github.com/rogerdigital/dsh-vet/blob/main/action/README.md#comparing-releases)).
|
|
27
|
+
|
|
28
|
+
We piloted it on real releases —
|
|
29
|
+
[left-pad and chalk, every reported addition reconciled against the
|
|
30
|
+
code](https://github.com/rogerdigital/dsh-vet/blob/main/docs/release-risk-pilot.md).
|
|
31
|
+
The chalk case is the interesting one: 5.x vendored its dependencies and
|
|
32
|
+
the diff flagged exactly the two vendored files unreachable through the
|
|
33
|
+
`#`-imports map — nothing else moved.
|
|
34
|
+
|
|
35
|
+
If you maintain a plugin and want the same read on your last two
|
|
36
|
+
releases, reply here or open an issue — we'll run the comparison and
|
|
37
|
+
post the result for you to check against what you actually changed.
|
|
@@ -0,0 +1,71 @@
|
|
|
1
|
+
<!-- Destination: issue (or DM) to the author of each target plugin, one
|
|
2
|
+
thread per plugin. The template below is filled with real pre-run
|
|
3
|
+
results — send as-is; the recipient installs nothing and answers from
|
|
4
|
+
their own knowledge of the release. Send from the maintainer's account,
|
|
5
|
+
then record each thread link in docs/release-risk-pilot.md. -->
|
|
6
|
+
|
|
7
|
+
# Release-risk pilot ask (template)
|
|
8
|
+
|
|
9
|
+
Subject: `Did your last release change what your plugin can do? (2-minute check)`
|
|
10
|
+
|
|
11
|
+
Hi — I maintain [`dsh-vet`](https://github.com/rogerdigital/dsh-vet), the
|
|
12
|
+
static audit tool for DSH plugins. It just learned to **compare two
|
|
13
|
+
releases of the same plugin and report exactly what risk-relevant
|
|
14
|
+
behavior changed** — new endpoints, new capabilities, new install
|
|
15
|
+
scripts, changed confidence — while ignoring line moves and reformatting.
|
|
16
|
+
|
|
17
|
+
I ran it on your last two releases so you don't have to install anything.
|
|
18
|
+
The results are below; three questions at the end.
|
|
19
|
+
|
|
20
|
+
---
|
|
21
|
+
|
|
22
|
+
## Filled example — dsh-doctor 0.4.2 → 0.4.3
|
|
23
|
+
|
|
24
|
+
Both scans: same scanner, full rule profile, complete coverage. Grades
|
|
25
|
+
stayed **D → D**. What moved:
|
|
26
|
+
|
|
27
|
+
- **New network client call** in `lib/client.js` — the destination is
|
|
28
|
+
computed at runtime, so the scanner can't see where it sends
|
|
29
|
+
(`perm.network-client`, subject `runtime-target`).
|
|
30
|
+
- **Subprocess spawns went from 1 to 2** in `lib/index.js`
|
|
31
|
+
(`child_process.spawn`).
|
|
32
|
+
|
|
33
|
+
Nothing else changed (13 identities unchanged).
|
|
34
|
+
|
|
35
|
+
## Filled example — dsh-searxng 0.2.1 → 0.3.0
|
|
36
|
+
|
|
37
|
+
Grades stayed **A → A**. One transition: **write/delete calls with
|
|
38
|
+
runtime-computed targets went from 1 to 17** in `lib/cli.mjs`
|
|
39
|
+
(`perm.undeclared-fs-write`, dynamic variant). These are low-severity /
|
|
40
|
+
low-confidence — they never affect the grade — but the count jump is the
|
|
41
|
+
kind of thing worth a look during a minor bump.
|
|
42
|
+
|
|
43
|
+
## Filled example — dsh-wechat 0.9.5 → 0.9.6
|
|
44
|
+
|
|
45
|
+
**Zero risk-relevant changes.** All 7 audited behaviors identical, grade
|
|
46
|
+
C → C. If that matches your intent for the release, the diff just saved
|
|
47
|
+
you the re-review.
|
|
48
|
+
|
|
49
|
+
---
|
|
50
|
+
|
|
51
|
+
## The three questions
|
|
52
|
+
|
|
53
|
+
1. Does the reported delta match what you intended to change in that
|
|
54
|
+
release? Anything you'd flag that the scanner missed or got wrong?
|
|
55
|
+
2. Any entries that felt like noise (didn't help you understand the
|
|
56
|
+
release)?
|
|
57
|
+
3. If this ran automatically on your PRs against the base revision —
|
|
58
|
+
[one Action input](https://github.com/rogerdigital/dsh-vet/blob/main/action/README.md#comparing-releases) —
|
|
59
|
+
would you read it before merging?
|
|
60
|
+
|
|
61
|
+
Answers in any form are useful; I'll record anonymized takeaways in the
|
|
62
|
+
[pilot record](https://github.com/rogerdigital/dsh-vet/blob/main/docs/release-risk-pilot.md).
|
|
63
|
+
This is the last feedback loop before deciding whether to build CI
|
|
64
|
+
gating on top (opt-in "fail on new risk"), so negative answers are as
|
|
65
|
+
valuable as positive ones.
|
|
66
|
+
|
|
67
|
+
<!-- Per-send checklist:
|
|
68
|
+
- re-run the pair with the latest main build the day of sending
|
|
69
|
+
- paste the actual result block for THIS author's package
|
|
70
|
+
- keep grades/identities verbatim from the diff, don't paraphrase
|
|
71
|
+
- one package per thread; don't batch multiple authors -->
|
|
@@ -0,0 +1,109 @@
|
|
|
1
|
+
# Release-risk review pilot (M1)
|
|
2
|
+
|
|
3
|
+
Status: **internal pilot complete; external feedback pending.** This page
|
|
4
|
+
records how the first milestone was exercised and what the comparison
|
|
5
|
+
showed on real release pairs. It is not an adoption claim — no external
|
|
6
|
+
maintainer has used the workflow yet, and the M2 entry gate (repeated
|
|
7
|
+
pilot use showing review friction) is not met.
|
|
8
|
+
|
|
9
|
+
## What was exercised
|
|
10
|
+
|
|
11
|
+
| Exercise | Result |
|
|
12
|
+
|---|---|
|
|
13
|
+
| Controlled pair (committed fixtures, one added host) | Comparable; exactly the added-host identities and observation; grades equal |
|
|
14
|
+
| Built-output smoke (`node bin/dsh-vet.mjs`, section 8 of the plan) | Help, scan ×2, validate ×2, diff, badge all pass; same-profile comparison of two runs reports zero risk changes despite distinct `ranAt`; badge unqualified for complete coverage |
|
|
15
|
+
| `left-pad@1.2.0` → `1.3.0` | Comparable, zero added/removed/changed — reconciled below |
|
|
16
|
+
| `chalk@4.1.2` → `5.3.0` | Comparable, two added `perm.unreachable-files` identities — reconciled below |
|
|
17
|
+
|
|
18
|
+
All pairs were scanned with the same scanner build and full rule profile;
|
|
19
|
+
both sides validated with `dsh-vet validate` before comparison.
|
|
20
|
+
|
|
21
|
+
## Reconciliation: left-pad 1.2.0 → 1.3.0
|
|
22
|
+
|
|
23
|
+
Both versions: identical audited behavior. The five unchanged identities
|
|
24
|
+
cover the benchmark/test files that ship unreachably (`perf/*`, `test.js`)
|
|
25
|
+
and the hex table in `perf/perf.js` (`obf.encoded-payload`, matched by
|
|
26
|
+
content hash) — all present in both versions, none moved. **Zero
|
|
27
|
+
additions to reconcile; the release changed padding internals without
|
|
28
|
+
touching any audited behavior.** Effort: under a minute; the diff
|
|
29
|
+
answered "anything new to review?" with a definite no.
|
|
30
|
+
|
|
31
|
+
## Reconciliation: chalk 4.1.2 → 5.3.0
|
|
32
|
+
|
|
33
|
+
Reported additions, each verified against the shipped tarballs:
|
|
34
|
+
|
|
35
|
+
1. `perm.unreachable-files` · `source/vendor/supports-color/index.js`
|
|
36
|
+
2. `perm.unreachable-files` · `source/vendor/supports-color/browser.js`
|
|
37
|
+
|
|
38
|
+
Code evidence:
|
|
39
|
+
|
|
40
|
+
- 4.1.2 has no `source/vendor/` directory at all; 5.3.0 vendors its
|
|
41
|
+
dependencies there.
|
|
42
|
+
- 5.3.0's `source/index.js` imports `#supports-color`, which resolves
|
|
43
|
+
through the package.json `imports` map to exactly those two files
|
|
44
|
+
(`node` → `index.js`, `default` → `browser.js`).
|
|
45
|
+
- The reference analyzer resolves relative imports but not `#`-prefixed
|
|
46
|
+
subpath imports, so neither file is statically reachable from the
|
|
47
|
+
declared entry — the finding text ("no static import path from any
|
|
48
|
+
entry point reaches this file") is accurate for this analyzer.
|
|
49
|
+
- The sibling vendored `ansi-styles` **is** imported relatively
|
|
50
|
+
(`'./vendor/ansi-styles/index.js'`), is reachable, and is correctly
|
|
51
|
+
**not** reported — the delta is per-subject, not per-directory.
|
|
52
|
+
|
|
53
|
+
Both grades stay `A` (unreachable files are informational); the delta is
|
|
54
|
+
visible precisely because observations and identities are reported even
|
|
55
|
+
when grades do not move.
|
|
56
|
+
|
|
57
|
+
Honest limitation observed: `#`-subpath `imports` maps are not resolved,
|
|
58
|
+
so code reached only through them reads as unreachable. That is a
|
|
59
|
+
documented analyzer gap to close in a later revision — visible here
|
|
60
|
+
because the delta surfaces it, which is the point of the workflow.
|
|
61
|
+
|
|
62
|
+
## Pilot observations
|
|
63
|
+
|
|
64
|
+
- The two real pairs took minutes to reconcile; every reported addition
|
|
65
|
+
mapped to a code fact, no formatting-driven false additions appeared
|
|
66
|
+
(line movement is excluded from identity by design and it held).
|
|
67
|
+
- The interesting noise class is informational unreachable-file churn on
|
|
68
|
+
refactors that vendor or move code — visible, cheap to dismiss, and
|
|
69
|
+
exactly what M2's exception mechanism would absorb if it repeats.
|
|
70
|
+
- One pilot round is not enough to justify policy work. The M2 entry
|
|
71
|
+
gate stays closed until a maintainer uses the comparison on real
|
|
72
|
+
releases repeatedly and reports the friction.
|
|
73
|
+
|
|
74
|
+
## External feedback (pending)
|
|
75
|
+
|
|
76
|
+
No external maintainer has used the comparison yet — this section stays
|
|
77
|
+
empty until one does, and a friendly reply is not adoption. The ask and
|
|
78
|
+
its per-plugin pre-run results live in
|
|
79
|
+
[`docs/outreach/release-risk-pilot-ask.md`](outreach/release-risk-pilot-ask.md);
|
|
80
|
+
record each thread here as it happens.
|
|
81
|
+
|
|
82
|
+
Intake questions (mirroring the plan's §10 usefulness criteria):
|
|
83
|
+
|
|
84
|
+
1. Does the reported delta match what the author intended to change?
|
|
85
|
+
2. Anything the scanner missed or got wrong?
|
|
86
|
+
3. Any entries that felt like noise?
|
|
87
|
+
4. Would they read it on their PRs — and did they run it on a *second*
|
|
88
|
+
release? (The repeated-use answer is the M2 entry gate.)
|
|
89
|
+
|
|
90
|
+
| Date | Plugin | Thread | Verdict (used / not / second release) |
|
|
91
|
+
|---|---|---|---|
|
|
92
|
+
| — | — | — | — |
|
|
93
|
+
|
|
94
|
+
## Reproduction
|
|
95
|
+
|
|
96
|
+
```sh
|
|
97
|
+
pnpm build
|
|
98
|
+
node bin/dsh-vet.mjs --json left-pad@1.2.0 > /tmp/lp-base.json
|
|
99
|
+
node bin/dsh-vet.mjs --json left-pad@1.3.0 > /tmp/lp-head.json
|
|
100
|
+
node bin/dsh-vet.mjs validate /tmp/lp-base.json /tmp/lp-head.json
|
|
101
|
+
node bin/dsh-vet.mjs diff --json /tmp/lp-base.json /tmp/lp-head.json
|
|
102
|
+
|
|
103
|
+
node bin/dsh-vet.mjs --json chalk@4.1.2 > /tmp/ch-base.json
|
|
104
|
+
node bin/dsh-vet.mjs --json chalk@5.3.0 > /tmp/ch-head.json
|
|
105
|
+
node bin/dsh-vet.mjs diff --json /tmp/ch-base.json /tmp/ch-head.json
|
|
106
|
+
```
|
|
107
|
+
|
|
108
|
+
The controlled pair lives in `test/fixtures/release-risk/` and is
|
|
109
|
+
asserted by `test/golden.test.ts` on every run.
|
|
@@ -0,0 +1,80 @@
|
|
|
1
|
+
# The `dsh-vet/diff/v1` comparison result
|
|
2
|
+
|
|
3
|
+
Status: **defined; the reference CLI exposes it from the next release.**
|
|
4
|
+
This document is normative for consumers of `compareReports()` and the
|
|
5
|
+
`dsh-vet diff` output. See [dsh-vet-v1.md](dsh-vet-v1.md) for the base
|
|
6
|
+
report and [scan-context-v1.md](scan-context-v1.md) for the extension the
|
|
7
|
+
comparator reads.
|
|
8
|
+
|
|
9
|
+
`compareReports(base, head)` is pure data in, pure data out: it never
|
|
10
|
+
fetches packages, never mutates either report, and never recalculates a
|
|
11
|
+
report under new rules. It compares what two validated reports actually
|
|
12
|
+
recorded.
|
|
13
|
+
|
|
14
|
+
## Comparability
|
|
15
|
+
|
|
16
|
+
A pair is **comparable** only when every gate passes:
|
|
17
|
+
|
|
18
|
+
| Gate | Reason code when it fails |
|
|
19
|
+
|---|---|
|
|
20
|
+
| Both are structurally valid `dsh-vet/v1` reports | `report-invalid-base` / `report-invalid-head` |
|
|
21
|
+
| Both carry a validated version-1 `x-dsh-vet` context | `context-absent-*`, `context-unsupported-*`, `context-invalid-*` |
|
|
22
|
+
| Both carry finding identities | `identities-absent-base` / `identities-absent-head` |
|
|
23
|
+
| Same scanner name and version | `scanner-mismatch` |
|
|
24
|
+
| Same profile digest (analyzer, catalog, rules, options) | `profile-mismatch` |
|
|
25
|
+
| Same package identity, or an explicit shared subject label | `subject-mismatch` / `subject-unlabeled` |
|
|
26
|
+
| Neither side graded `X` | `grade-x-base` / `grade-x-head` |
|
|
27
|
+
| Complete coverage on both sides | `coverage-partial-base` / `coverage-partial-head` |
|
|
28
|
+
|
|
29
|
+
An incomparable result carries the reason codes and both report
|
|
30
|
+
references, and **no change collections** — an absent diff must never
|
|
31
|
+
read as "nothing changed". There is no force-comparison switch in M1;
|
|
32
|
+
when practical, rescan both artifacts with the same scanner and profile
|
|
33
|
+
instead. Two local-directory scans without package identity compare only
|
|
34
|
+
under an explicit `subjectLabel` option; identity is never inferred from
|
|
35
|
+
temporary absolute paths. Different analysis-input or archive digests
|
|
36
|
+
are expected and never block comparison.
|
|
37
|
+
|
|
38
|
+
## Result shape
|
|
39
|
+
|
|
40
|
+
```jsonc
|
|
41
|
+
{
|
|
42
|
+
"schema": "dsh-vet/diff/v1",
|
|
43
|
+
"base": { "scanner": { /* name, version, ranAt */ }, "grade": "C", "subject": { /* … */ } },
|
|
44
|
+
"head": { "scanner": { /* … */ }, "grade": "C", "subject": { /* … */ } },
|
|
45
|
+
"comparability": "comparable",
|
|
46
|
+
"reasons": [],
|
|
47
|
+
"findings": {
|
|
48
|
+
"added": [ /* identities present only in head */ ],
|
|
49
|
+
"removed": [ /* identities present only in base */ ],
|
|
50
|
+
"changed": [ /* matched identities that moved; both sides recorded */ ],
|
|
51
|
+
"unchanged": [ /* matched and identical */ ]
|
|
52
|
+
},
|
|
53
|
+
"observations": {
|
|
54
|
+
"added": [], "removed": [], "countChanged": [], "unchanged": []
|
|
55
|
+
}
|
|
56
|
+
}
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
Identity entries and observations are the extension's own objects
|
|
60
|
+
(`rule`/`variant`/`file`/`subject`/`finding`/`count?` and
|
|
61
|
+
`kind`/`file`/`subject`/`count?`). Every collection is sorted and
|
|
62
|
+
deterministic: two comparisons of the same pair produce byte-identical
|
|
63
|
+
results.
|
|
64
|
+
|
|
65
|
+
A **changed** finding records both sides' severity, confidence, and
|
|
66
|
+
count. Severity and confidence transitions are classified separately
|
|
67
|
+
from additions on purpose: moving from `low` to `medium` confidence can
|
|
68
|
+
introduce a graded risk at unchanged severity.
|
|
69
|
+
|
|
70
|
+
## Reading the result honestly
|
|
71
|
+
|
|
72
|
+
- `added` means newly observed under the same checks — not proof of new
|
|
73
|
+
malice, and not a grade change by itself.
|
|
74
|
+
- `removed` means no longer observed under comparable checks — never a
|
|
75
|
+
claim that anything was fixed.
|
|
76
|
+
- Observation deltas are visible even when grades are identical or the
|
|
77
|
+
observations are informational; an unchanged grade never hides new
|
|
78
|
+
behavior.
|
|
79
|
+
- The comparator reports line-position changes as nothing: identities
|
|
80
|
+
exclude line numbers, so a line insertion is not a risk addition.
|