control-arm 1.0.0 → 1.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +44 -0
- package/DESIGN.md +37 -21
- package/README.md +205 -24
- package/bin/ca.mjs +53 -6
- package/package.json +11 -4
- package/src/assertions.mjs +37 -1
- package/src/children.mjs +55 -0
- package/src/cli-spec.mjs +50 -0
- package/src/contract.mjs +193 -0
- package/src/identity.mjs +18 -5
- package/src/markdown-report.mjs +5 -1
- package/src/report.mjs +14 -1
- package/src/runner-json.mjs +47 -21
- package/src/runner.mjs +21 -8
- package/src/sample-warning.mjs +69 -0
- package/src/select-runner.mjs +7 -0
- package/src/tap.mjs +0 -10
- package/src/verify.mjs +397 -9
- package/src/worktree.mjs +11 -1
package/CHANGELOG.md
ADDED
|
@@ -0,0 +1,44 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
All notable changes to this project are documented in this file.
|
|
4
|
+
|
|
5
|
+
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this
|
|
6
|
+
project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
7
|
+
|
|
8
|
+
## [1.2.0] - 2026-10-06
|
|
9
|
+
|
|
10
|
+
> **Note on the missing 1.1.0.** `package.json` was bumped to `1.1.0` on 2026-09-25 but that
|
|
11
|
+
> version was never tagged in git, never published to npm, and never released on GitHub —
|
|
12
|
+
> npm's `latest` stayed on `1.0.0` until now. The `1.1.0` work (vitest and jest runners, the
|
|
13
|
+
> findings-contract v1 envelope, the child-process reaper, and the fixes that shipped with
|
|
14
|
+
> them) is folded into this release.
|
|
15
|
+
|
|
16
|
+
### Fixed
|
|
17
|
+
|
|
18
|
+
- The module-identity proof silently did nothing on Node < 20.6: `import.meta.resolve` is
|
|
19
|
+
undefined there, and the per-specifier error was swallowed as "genuinely missing", so every
|
|
20
|
+
verdict was "proven" without resolving a single import. The probe now reports when it
|
|
21
|
+
cannot resolve, and the verdict is withheld instead of trusted.
|
|
22
|
+
- The findings-contract `version` field was hardcoded to `1.1.0`; it is now read from
|
|
23
|
+
`package.json` so it cannot drift.
|
|
24
|
+
|
|
25
|
+
### Changed
|
|
26
|
+
|
|
27
|
+
- Minimum Node version raised to `20.6.0` (was `18`).
|
|
28
|
+
|
|
29
|
+
### Added
|
|
30
|
+
|
|
31
|
+
- `scripts/precision.mjs`: measure the precision (and recall, when the sample is labelled)
|
|
32
|
+
of the `BLIND` verdict against a hand-labelled corpus.
|
|
33
|
+
- Unit coverage for the vitest/jest report parser (`parseReport`) against captured runner
|
|
34
|
+
output.
|
|
35
|
+
|
|
36
|
+
### Removed
|
|
37
|
+
|
|
38
|
+
- Dead code: the tautological `nodeTest.matches` predicate, the unused `RUNNERS` export,
|
|
39
|
+
`isFileLevelFailure`, and a leftover destructure in `commitInfo`.
|
|
40
|
+
|
|
41
|
+
## [1.0.0] - 2026-09-25
|
|
42
|
+
|
|
43
|
+
- Initial public release: `ca doctor` / `ca verify` / `ca audit` over `node:test`, published
|
|
44
|
+
to npm and the GitHub Marketplace (`mekanhan/control-arm@v1`).
|
package/DESIGN.md
CHANGED
|
@@ -109,11 +109,19 @@ weighting: a codebase the author did not write, did not pick for a flattering re
|
|
|
109
109
|
could not tune against. 190 commits matched the subject filter; 99 ship both a test and a
|
|
110
110
|
source change; 40 were drawn from those.
|
|
111
111
|
|
|
112
|
-
**One usability trap found while running it.** `audit` defaults to a
|
|
113
|
-
|
|
114
|
-
instead of 40** and printed a confident, meaningless 100%.
|
|
115
|
-
|
|
116
|
-
|
|
112
|
+
**One usability trap found while running it.** `audit` defaults to a **twelve-month**
|
|
113
|
+
window, and dayjs ships few `fix:` commits in a year that also touch a test. The first
|
|
114
|
+
dayjs run judged **5 commits instead of 40** and printed a confident, meaningless 100%.
|
|
115
|
+
Nothing was broken: the pool was simply smaller than the request.
|
|
116
|
+
|
|
117
|
+
What exposed it was that two different `--n` values and two different seeds gave
|
|
118
|
+
**byte-identical output**. A rate that does not move when you change the sample size is
|
|
119
|
+
not a rate.
|
|
120
|
+
|
|
121
|
+
`audit` now says so when the draw comes up short, naming whichever cause applies — the
|
|
122
|
+
window, or the test-plus-source pre-filter — and never blaming a `--since` the caller
|
|
123
|
+
passed themselves (`src/sample-warning.mjs`, `SAMPLE-001..006`). Pass `--since` explicitly
|
|
124
|
+
on any repo with more than a year of history.
|
|
117
125
|
|
|
118
126
|
`INCONCLUSIVE` is 29% and 40% respectively. That is the honest denominator, not a rounding
|
|
119
127
|
error: a fix that adds an export its test imports cannot be replayed against the parent,
|
|
@@ -122,22 +130,30 @@ either way.
|
|
|
122
130
|
|
|
123
131
|
## How often is this tool wrong?
|
|
124
132
|
|
|
125
|
-
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
|
|
131
|
-
|
|
132
|
-
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
|
|
137
|
-
|
|
138
|
-
|
|
139
|
-
|
|
140
|
-
|
|
133
|
+
**[README.md → "How often is the TOOL right?"](README.md#does-it-work)** holds the audit:
|
|
134
|
+
13 raw `BLIND` on the 300-commit run, 5 genuinely open after arm C, a 62% false-positive
|
|
135
|
+
rate on the raw number.
|
|
136
|
+
|
|
137
|
+
It lives there and not here because it is the first thing a skeptic should see, and because
|
|
138
|
+
a table kept in two places drifts in one of them. What belongs here is the design
|
|
139
|
+
consequence:
|
|
140
|
+
|
|
141
|
+
**Arm C is not a nicety, it is the reason the raw signal is publishable at all.** Eight of
|
|
142
|
+
those thirteen were not defects in anybody's tests — three gaps had been closed by a later
|
|
143
|
+
commit, two named test files no longer exist, two were `fix(test):` commits where the bug
|
|
144
|
+
*was* the test, and one was a build failure no unit test could catch. None of that is
|
|
145
|
+
visible from the two arms alone; all of it needs a third look at HEAD.
|
|
146
|
+
|
|
147
|
+
That is also why the corpus filter drops `fix(test):` by default, and why
|
|
148
|
+
`--include-test-fixes` exists for anyone who wants them back.
|
|
149
|
+
|
|
150
|
+
**Measuring precision properly.** "5 for 5, too small to quote" is a count, not a rate.
|
|
151
|
+
To turn it into a rate — precision, and recall when the whole sample is labelled — label
|
|
152
|
+
each judged commit `gap`/`not-gap` and run `scripts/precision.mjs`. It reads `ca audit
|
|
153
|
+
--out` plus the labels, and prints the confusion matrix for both the raw `BLIND` signal and
|
|
154
|
+
the post-arm-C "still open" signal. Until that labelled corpus exists, the defensible claim
|
|
155
|
+
is the one the README already makes: the raw signal is over-sensitive, the filters remove
|
|
156
|
+
the eight known false positives, and the residual precision is unmeasured.
|
|
141
157
|
|
|
142
158
|
## Every bug found in this tool so far
|
|
143
159
|
|
package/README.md
CHANGED
|
@@ -40,7 +40,7 @@ because it runs the new test against the old code.
|
|
|
40
40
|
npm i -g control-arm # or run it from a clone: node bin/ca.mjs
|
|
41
41
|
```
|
|
42
42
|
|
|
43
|
-
Node
|
|
43
|
+
Node 20.6+. No other dependencies.
|
|
44
44
|
|
|
45
45
|
## Use it
|
|
46
46
|
|
|
@@ -87,9 +87,82 @@ It comments on the PR with a line per test. It does **not** fail the build by de
|
|
|
87
87
|
PR can legitimately ship only regression guards, and a gate that fires on those gets
|
|
88
88
|
switched off within a week.
|
|
89
89
|
|
|
90
|
+
**On a feature PR it refuses to claim anything.** This matters, because the check is
|
|
91
|
+
meaningless there and would otherwise read as proof. On a feature the base is code where
|
|
92
|
+
the thing does not exist yet, so essentially *any* test touching it fails —
|
|
93
|
+
`expected 0 to be greater than 0` is a real assertion failure that says nothing about
|
|
94
|
+
whether the test is well aimed. So the comment changes its own headline:
|
|
95
|
+
|
|
96
|
+
> ### `control-arm` — new code — these tests cannot be judged this way
|
|
97
|
+
>
|
|
98
|
+
> **3 of 4 tests fail without this change** — but this is a feature, so the base does not
|
|
99
|
+
> have it at all.
|
|
100
|
+
>
|
|
101
|
+
> That is expected, and it is weak evidence: on new code almost any test that reads the new
|
|
102
|
+
> thing fails on the base, whether or not it is well aimed. A test asserting `1 === 1` in
|
|
103
|
+
> the same file would NOT show up here; one that merely reads the new field would.
|
|
104
|
+
|
|
105
|
+
The classification is `fix`/`feat` from the conventional-commit prefix, plus a supporting
|
|
106
|
+
hint — a commit that only *added* source lines and deleted none looks like new code
|
|
107
|
+
regardless of its prefix. Both were measured on a real corpus before being trusted: the
|
|
108
|
+
prefix is used consistently (1,213 `fix` / 988 `feat`), while additions-only fires on 9 of
|
|
109
|
+
28 features but also on 1 of 35 fixes — **specific, not sensitive**, so it qualifies the
|
|
110
|
+
claim and never decides the verdict on its own. A `CAUGHT` on a feature is still `CAUGHT`;
|
|
111
|
+
it just does not get to say it would have caught a bug, because there was no bug.
|
|
112
|
+
|
|
113
|
+
## All flags
|
|
114
|
+
|
|
115
|
+
```bash
|
|
116
|
+
ca doctor --repo . # can this repo be measured?
|
|
117
|
+
ca verify <sha> --repo . --runs 3 --timeout 120000 --work .ca-work
|
|
118
|
+
ca verify <sha> --against origin/main --pr-comment --fail-on-blind
|
|
119
|
+
ca audit --n 100 --since '6 years' --grep '^fix' --seed 1 --out r.csv --html r.html
|
|
120
|
+
ca issues --apply --include-test-fixes
|
|
121
|
+
```
|
|
122
|
+
|
|
123
|
+
| flag | |
|
|
124
|
+
|---|---|
|
|
125
|
+
| `--repo <path>` | the repository to measure (default: cwd) |
|
|
126
|
+
| `--against <ref>` | compare against a base ref — turns `verify` into a PR check |
|
|
127
|
+
| `--runs <n>` | repeat each case N times to expose flakes |
|
|
128
|
+
| `--timeout <ms>` | per-run timeout |
|
|
129
|
+
| `--work <path>` | where the throwaway worktrees go |
|
|
130
|
+
| `--keep` | leave them behind for inspection |
|
|
131
|
+
| `--n <n>` · `--since <when>` · `--grep <re>` · `--seed <n>` | which commits `audit` samples |
|
|
132
|
+
| `--include-test-fixes` | do not skip `fix(test):` commits, where the bug WAS the test |
|
|
133
|
+
| `--out <file>` · `--html <file>` | write the audit as CSV or HTML |
|
|
134
|
+
| `--pr-comment` | print the PR comment markdown and nothing else |
|
|
135
|
+
| `--apply` | `issues` actually files them, rather than printing |
|
|
136
|
+
| **`--json`** | **findings-contract v1 envelope on stdout** (see below) |
|
|
137
|
+
| `--json <path>` | the older behaviour: write the audit summary to a file |
|
|
138
|
+
| **`--fail-on-blind`** | **exit 1 when a commit is BLIND** |
|
|
139
|
+
|
|
140
|
+
### The findings contract
|
|
141
|
+
|
|
142
|
+
Bare `--json` emits one [findings-contract v1](https://github.com/mekanhan/findings-contract)
|
|
143
|
+
object on stdout and nothing else — progress goes to stderr, so it pipes. It works on
|
|
144
|
+
`doctor`, `verify` and `audit`.
|
|
145
|
+
|
|
146
|
+
The mapping is not one-to-one, and the interesting part is what is **not** a finding:
|
|
147
|
+
|
|
148
|
+
| | |
|
|
149
|
+
|---|---|
|
|
150
|
+
| commit `CAUGHT` | no finding. Nothing is wrong |
|
|
151
|
+
| commit `BLIND` | a finding, severity `warn` |
|
|
152
|
+
| commit `INCONCLUSIVE` | **not** a finding — it goes to `skipped`, with its reason |
|
|
153
|
+
| case `NON-DISCRIMINATING` | never a finding. A fact about a guard, not a fault |
|
|
154
|
+
| case `FLAKY` | a finding. Runs disagreed, which is a real defect |
|
|
155
|
+
|
|
156
|
+
**Exit codes:** `0` ran cleanly · `1` found blockers · `2` the tool itself failed.
|
|
157
|
+
|
|
158
|
+
`ca verify` used to exit `1` on `BLIND`. It no longer does unless you pass
|
|
159
|
+
`--fail-on-blind`, because `BLIND` is a `warn` — a test that cannot catch a bug breaks
|
|
160
|
+
nothing today — and putting a warning into an exit code makes every caller guess. The
|
|
161
|
+
GitHub Action is unaffected: it reads the verdict from stdout.
|
|
162
|
+
|
|
90
163
|
## What you get
|
|
91
164
|
|
|
92
|
-
> ### `control-arm` —
|
|
165
|
+
> ### `control-arm` — 1 test here fails without this change
|
|
93
166
|
>
|
|
94
167
|
> **1 of 5 tests here genuinely catch this change.** Run against the code as it was before
|
|
95
168
|
> this PR, they fail; with the change, they pass.
|
|
@@ -121,41 +194,149 @@ measurement that cannot tell zero from failure is not a measurement.
|
|
|
121
194
|
|
|
122
195
|
## Which of your tests it looks at
|
|
123
196
|
|
|
124
|
-
|
|
197
|
+
**Unit and integration tests** — anything that runs in-process, without a deployed
|
|
198
|
+
application. That is the whole scope.
|
|
125
199
|
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
|
|
|
131
|
-
|
|
132
|
-
|
|
|
133
|
-
|
|
|
134
|
-
|
|
|
200
|
+
**Every row carries what proves it.** A row without evidence beside it is an aspiration, and
|
|
201
|
+
that is where this table was wrong before: vitest and jest were listed as covered because the
|
|
202
|
+
code has adapters for them, not because either had ever been run.
|
|
203
|
+
|
|
204
|
+
| runner | proven by |
|
|
205
|
+
|---|---|
|
|
206
|
+
| `node:test` | fixtures `01`–`12`, built and run by `npm test` on every push |
|
|
207
|
+
| vitest | observed on a real commit, 2026-09-25 — **no fixture yet** (#21) |
|
|
208
|
+
| jest | observed on a real commit, 2026-09-25 — **no fixture yet** (#21) |
|
|
209
|
+
|
|
210
|
+
A database-backed test is judged like any other, and needs its database present. The README
|
|
211
|
+
used to claim it reports `SKIPPED` without one; that has never actually been observed, so the
|
|
212
|
+
claim is withdrawn until it is (#21).
|
|
213
|
+
|
|
214
|
+
**"fixtures"** means anyone can re-run it: `npm test` builds throwaway repos where the right
|
|
215
|
+
answer is known by construction, and the suite fails if this tool cannot tell
|
|
216
|
+
`02-blind-direction` from `01-caught-value`. **"observed on a real commit"** means it worked
|
|
217
|
+
once, on a repository you cannot see — weaker, and marked weaker.
|
|
218
|
+
|
|
219
|
+
### Out of scope, permanently
|
|
220
|
+
|
|
221
|
+
Browser e2e (Playwright), device e2e (WDIO/Appium) and load tests (k6).
|
|
222
|
+
|
|
223
|
+
Arm B has to run your test against the **old code**, and for an e2e test the old code is a
|
|
224
|
+
*running application* — built and served at the parent commit, with its database and its
|
|
225
|
+
services. That is minutes per commit at best, and frequently the parent will not boot at all.
|
|
226
|
+
Load tests measure speed rather than correctness, so the question this tool asks does not
|
|
227
|
+
apply to them.
|
|
228
|
+
|
|
229
|
+
Point it at one of those and it will name the runner and decline, rather than attempting it
|
|
230
|
+
with the wrong one. That is a courtesy, not a roadmap — none of these is planned.
|
|
135
231
|
|
|
136
232
|
## Does it work?
|
|
137
233
|
|
|
138
|
-
|
|
234
|
+
**Two different questions live here, and mixing them flatters the tool.** How good are the
|
|
235
|
+
tests it measured, and how often is the tool itself right? Only the second is about
|
|
236
|
+
`control-arm`.
|
|
237
|
+
|
|
238
|
+
### How often is the TOOL right?
|
|
239
|
+
|
|
240
|
+
This is the number to judge it on. On the 300-commit run it raised **13 `BLIND` verdicts**
|
|
241
|
+
— "not one test in this commit would have caught the bug". Each was then re-run through
|
|
242
|
+
arm C and checked by hand:
|
|
243
|
+
|
|
244
|
+
| | |
|
|
245
|
+
|---|---|
|
|
246
|
+
| 13 | raw `BLIND` |
|
|
247
|
+
| −3 | `↻ REPAIRED SINCE` — the gap was real, and closed after that commit |
|
|
248
|
+
| −2 | `? cannot tell` — the test file no longer exists at HEAD |
|
|
249
|
+
| −3 | category errors: two `fix(test):` (the bug WAS the test) and one build failure no unit test can catch |
|
|
250
|
+
| **5** | genuinely open — **2.3% of answerable commits** |
|
|
251
|
+
|
|
252
|
+
**A 62% false-positive rate on the raw number.** Publish it and you hand someone thirteen
|
|
253
|
+
tickets, eight of which waste their afternoon. Arm C and the corpus filter exist entirely
|
|
254
|
+
because of that, and they run by default.
|
|
255
|
+
|
|
256
|
+
All five survivors held up under hand inspection. **That is 5 for 5 out of 13 candidates —
|
|
257
|
+
far too small a sample to quote as a precision rate**, and it is stated as a count for that
|
|
258
|
+
reason. The honest summary is: the raw signal is badly over-sensitive, the filters remove
|
|
259
|
+
eight of eight known false positives, and what precision remains after them has not been
|
|
260
|
+
measured on a sample large enough to have a rate.
|
|
261
|
+
|
|
262
|
+
`BLIND` is the only verdict that accuses anybody, so every ambiguity resolves away from it.
|
|
263
|
+
|
|
264
|
+
### How good were the TESTS it measured?
|
|
265
|
+
|
|
266
|
+
These say nothing about the tool's accuracy — they are a property of the repositories.
|
|
267
|
+
Counts, not just percentages, because a percentage over 25 commits invites arithmetic
|
|
268
|
+
nobody should have to do themselves.
|
|
139
269
|
|
|
140
270
|
| | a private monorepo | nodejs/undici | iamkun/dayjs |
|
|
141
271
|
|---|---|---|---|
|
|
142
272
|
| fix commits sampled | 300 | 25 | 40 |
|
|
143
273
|
| answerable | 214 | 15 | 25 |
|
|
144
|
-
| **CAUGHT** | **93.9%** | **80.0%** | **92.0%** |
|
|
145
|
-
| BLIND | 6.1% | 20.0% | 8.0% |
|
|
274
|
+
| **CAUGHT** | **201 — 93.9%** | **12 — 80.0%** | **23 — 92.0%** |
|
|
275
|
+
| BLIND | 13 — 6.1% | **3** — 20.0% | **2** — 8.0% |
|
|
276
|
+
| INCONCLUSIVE | 86 | 10 | 9 |
|
|
146
277
|
| runtime | 7.1 s/commit | 24.9 s/commit | 1.4 s/commit |
|
|
147
278
|
|
|
148
|
-
**The
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
|
|
279
|
+
**The small columns are small.** dayjs's 8% is **two commits** and undici's 20% is **three**.
|
|
280
|
+
Those are proof-of-concept numbers; do not read a difference between 80% and 92% as a
|
|
281
|
+
difference between the repositories.
|
|
282
|
+
|
|
283
|
+
**The dayjs column is the one to weigh** anyway: a codebase the author did not write, did
|
|
284
|
+
not choose for a flattering result, and could not tune against — 190 fix commits matched
|
|
285
|
+
the filter, 99 ship both a test and a source change, 40 drawn at random with the seed
|
|
286
|
+
recorded. It lands within a point of the private monorepo it was built on.
|
|
287
|
+
|
|
288
|
+
### Why is so much unanswerable?
|
|
289
|
+
|
|
290
|
+
29% and 40% is a lot to exclude, so here is where it goes. `INCONCLUSIVE` is never folded
|
|
291
|
+
into the other columns — a measurement that cannot tell zero from failure is not a
|
|
292
|
+
measurement — and there are exactly four ways to earn it:
|
|
293
|
+
|
|
294
|
+
| cause | what happened |
|
|
295
|
+
|---|---|
|
|
296
|
+
| **the test will not load on the parent** | the dominant one. The fix *added* an export, a module, a fixture; the test imports it; on the parent that import throws before a single assertion runs. `SyntaxError: does not provide an export named …` is not a test failing, it is a test never starting |
|
|
297
|
+
| **the test is not green on the fix either** | no before/after to compare. Usually an environment gap — a database, a browser, a missing `.env` |
|
|
298
|
+
| **module identity unprovable** | the run could not be shown to have loaded the worktree's code rather than the real repo's. Refused rather than guessed |
|
|
299
|
+
| **a skipped case is present** | the skipped one might have been the discriminating one, so the commit cannot be called `BLIND` |
|
|
300
|
+
|
|
301
|
+
The breakdown *between* these four has not been counted per repository, so no split is
|
|
302
|
+
quoted here.
|
|
303
|
+
|
|
304
|
+
## Working on this repo
|
|
305
|
+
|
|
306
|
+
```bash
|
|
307
|
+
git config core.hooksPath .githooks # once per clone; worktrees inherit it
|
|
308
|
+
```
|
|
309
|
+
|
|
310
|
+
`pre-push` refuses three things CI can only tell you about after the fact: a direct push to
|
|
311
|
+
`main` (everything lands through a PR), a branch **named** like a default that is not this
|
|
312
|
+
repo's default, and a push from a base that has already moved.
|
|
313
|
+
|
|
314
|
+
The middle one is not hypothetical. A local `master` once sat ten commits behind `main`
|
|
315
|
+
while `main` moved on through four PRs; pushing it created a parallel remote branch, and the
|
|
316
|
+
Node 20/22/24 matrix came back green against the wrong base. The only tell was a line of
|
|
317
|
+
push output — `* [new branch] master -> master` — on a repo that is anything but new.
|
|
318
|
+
|
|
319
|
+
If a push prints none of that hook's output, it is not armed. Check `core.hooksPath` before
|
|
320
|
+
trusting anything it did not say.
|
|
321
|
+
|
|
322
|
+
## Prior art
|
|
323
|
+
|
|
324
|
+
The fail-before / pass-after check is not new, and it is worth saying who got there first:
|
|
325
|
+
|
|
326
|
+
- **[Defects4J](https://github.com/rjust/defects4j)** records, for each reproducible bug,
|
|
327
|
+
the **trigger tests** that fail on the buggy version and pass on the fixed one.
|
|
328
|
+
- **[SWE-bench](https://www.swebench.com/)** builds every task around **`FAIL_TO_PASS`**
|
|
329
|
+
tests, with `PASS_TO_PASS` as the regression guard — the same two arms.
|
|
152
330
|
|
|
153
|
-
|
|
154
|
-
|
|
331
|
+
Both use the property to **construct benchmarks**: they start from a known bug and keep the
|
|
332
|
+
tests that prove it. `control-arm` runs the same check in the other direction — on ordinary
|
|
333
|
+
commits, in CI, where the answer is not known in advance and a `BLIND` result is news
|
|
334
|
+
rather than a data-cleaning step.
|
|
155
335
|
|
|
156
|
-
|
|
157
|
-
|
|
158
|
-
|
|
336
|
+
The consequence of that difference is most of this repository: a benchmark can discard
|
|
337
|
+
anything ambiguous, because it only needs *some* clean examples. A CI check cannot discard
|
|
338
|
+
the commit in front of it, so it has to be able to say `INCONCLUSIVE` out loud, and it has
|
|
339
|
+
to be right when it says `BLIND` about somebody's work.
|
|
159
340
|
|
|
160
341
|
## Scope, stated plainly
|
|
161
342
|
|
package/bin/ca.mjs
CHANGED
|
@@ -16,11 +16,28 @@ import { renderVerify, renderAudit, MARK } from '../src/report.mjs';
|
|
|
16
16
|
import { renderHtml, issueBody } from '../src/html-report.mjs';
|
|
17
17
|
import { prComment } from '../src/markdown-report.mjs';
|
|
18
18
|
import { analyseCase, extractCase } from '../src/assertions.mjs';
|
|
19
|
+
import { COMMANDS, COMMAND_NAMES } from '../src/cli-spec.mjs';
|
|
20
|
+
import { sampleWarning, relativeWindowWarning } from '../src/sample-warning.mjs';
|
|
21
|
+
import { verifyEnvelope, auditEnvelope, doctorEnvelope } from '../src/contract.mjs';
|
|
19
22
|
|
|
20
23
|
const argv = process.argv.slice(2);
|
|
21
24
|
const cmd = argv[0];
|
|
22
25
|
const flag = (n, d = null) => { const i = argv.indexOf(`--${n}`); return i === -1 ? d : argv[i + 1]; };
|
|
23
26
|
const has = n => argv.includes(`--${n}`);
|
|
27
|
+
// `--json` predates the contract and took a PATH. Bare `--json` — no value, or the next
|
|
28
|
+
// token is another flag — now means "contract envelope on stdout", which is what C-001
|
|
29
|
+
// asks for. `--json <path>` is unchanged, so nothing that worked stops working.
|
|
30
|
+
const bareJson = (() => {
|
|
31
|
+
const i = argv.indexOf('--json');
|
|
32
|
+
if (i === -1) return false;
|
|
33
|
+
const next = argv[i + 1];
|
|
34
|
+
return next === undefined || next.startsWith('-');
|
|
35
|
+
})();
|
|
36
|
+
/** C-001: one object on stdout, nothing else. Then C-007 decides the code. */
|
|
37
|
+
const emit = (env, gate = false) => {
|
|
38
|
+
process.stdout.write(JSON.stringify(env, null, 2) + '\n');
|
|
39
|
+
process.exit(gate && env.summary.blocker ? 1 : 0);
|
|
40
|
+
};
|
|
24
41
|
const repo = path.resolve(flag('repo', process.cwd()));
|
|
25
42
|
const workDir = path.resolve(flag('work', path.join(repo, '.ca-work')));
|
|
26
43
|
|
|
@@ -59,6 +76,8 @@ async function doctor() {
|
|
|
59
76
|
checks.push({ ok: true, name: 'workspaces', detail: `${[].concat(pkg.workspaces).join(', ')} — workspace links will be re-pointed into the worktree (this is the trap that produces false BLIND)` });
|
|
60
77
|
}
|
|
61
78
|
|
|
79
|
+
if (bareJson) emit(doctorEnvelope(checks, { repo }), true);
|
|
80
|
+
|
|
62
81
|
const w = 22;
|
|
63
82
|
console.log(`\n ca doctor — ${repo}\n`);
|
|
64
83
|
for (const c of checks) console.log(` ${c.ok ? MARK.ok : MARK.no} ${c.name.padEnd(w)} ${c.detail}`);
|
|
@@ -74,10 +93,18 @@ async function verify() {
|
|
|
74
93
|
const r = await verifyCommit({ repo, workDir, sha, against: flag('against'), runs, timeoutMs: Number(flag('timeout', 120_000)),
|
|
75
94
|
onStep: s => process.stderr.write(`\r … ${s} `) });
|
|
76
95
|
process.stderr.write('\r' + ' '.repeat(40) + '\r');
|
|
96
|
+
if (bareJson) { if (!has('keep')) await removeWorktrees(repo, workDir); emit(verifyEnvelope(r, { repo, sha })); }
|
|
77
97
|
if (has('pr-comment')) console.log(prComment(r, { repoName: flag('repo', '.') }));
|
|
78
98
|
else console.log(renderVerify(r));
|
|
79
99
|
if (!has('keep')) await removeWorktrees(repo, workDir);
|
|
80
|
-
|
|
100
|
+
|
|
101
|
+
// C-007. This used to exit 1 on BLIND unconditionally, which made a `warn` look like
|
|
102
|
+
// a blocker and put this tool's own opinion into an exit code every caller has to
|
|
103
|
+
// interpret. Gating is opt-in now, under the same name the Action already uses.
|
|
104
|
+
// BEHAVIOUR CHANGE: `ca verify` on a BLIND commit exits 0 unless --fail-on-blind.
|
|
105
|
+
// The Action is unaffected — it reads the verdict from stdout and already wraps the
|
|
106
|
+
// call in `|| true`.
|
|
107
|
+
process.exit(has('fail-on-blind') && r.verdict === BLIND ? 1 : 0);
|
|
81
108
|
}
|
|
82
109
|
|
|
83
110
|
async function audit() {
|
|
@@ -115,7 +142,17 @@ async function audit() {
|
|
|
115
142
|
const pool = [...eligible];
|
|
116
143
|
for (let i = pool.length - 1; i > 0; i--) { const j = Math.floor(rand() * (i + 1)); [pool[i], pool[j]] = [pool[j], pool[i]]; }
|
|
117
144
|
const sample = pool.slice(0, n);
|
|
118
|
-
process.stderr.write(` drawing ${sample.length} at random (seed ${seed})\n
|
|
145
|
+
process.stderr.write(` drawing ${sample.length} at random (seed ${seed})\n`);
|
|
146
|
+
|
|
147
|
+
// A draw that came up short is the difference between a measurement and a number.
|
|
148
|
+
const moving = relativeWindowWarning(since, argv.includes('--seed'));
|
|
149
|
+
if (moving) process.stderr.write(`\n${moving}\n`);
|
|
150
|
+
|
|
151
|
+
const short = sampleWarning({
|
|
152
|
+
requested: n, matched: candidates.length, eligible: eligible.length,
|
|
153
|
+
drawn: sample.length, since, sinceWasExplicit: argv.includes('--since'),
|
|
154
|
+
});
|
|
155
|
+
process.stderr.write(short ? `\n${short}\n\n` : '\n');
|
|
119
156
|
|
|
120
157
|
const results = [];
|
|
121
158
|
const t0 = Date.now();
|
|
@@ -130,6 +167,8 @@ async function audit() {
|
|
|
130
167
|
}
|
|
131
168
|
process.stderr.write('\r' + ' '.repeat(120) + '\r');
|
|
132
169
|
|
|
170
|
+
if (bareJson) emit(auditEnvelope(results, { repo, since, n: sample.length, seed }));
|
|
171
|
+
|
|
133
172
|
console.log(renderAudit(results, { since, n: sample.length, eligible: eligible.length, matched: candidates.length, seed, seconds: (Date.now() - t0) / 1000 }));
|
|
134
173
|
|
|
135
174
|
const htmlOut = flag('html');
|
|
@@ -180,6 +219,7 @@ async function audit() {
|
|
|
180
219
|
caught_pct: answerable ? Number(((results.filter(r => r.verdict === CAUGHT).length / answerable) * 100).toFixed(1)) : null,
|
|
181
220
|
still_open: byStatus('open'),
|
|
182
221
|
repaired_since: byStatus('repaired'),
|
|
222
|
+
likely_repaired_elsewhere: byStatus('repaired-elsewhere'),
|
|
183
223
|
cannot_tell: results.filter(r => r.verdict === BLIND && (!r.stillOpen || r.stillOpen.status === 'unknown'))
|
|
184
224
|
.map(r => ({ sha: r.sha, short: r.short, subject: r.subject })),
|
|
185
225
|
}, null, 2));
|
|
@@ -229,14 +269,21 @@ async function issues() {
|
|
|
229
269
|
if (!has('keep')) await removeWorktrees(repo, workDir);
|
|
230
270
|
}
|
|
231
271
|
|
|
232
|
-
|
|
272
|
+
// Built from the spec, not written out again here. `test/readme.test.mjs` checks the
|
|
273
|
+
// docs against COMMANDS, so a command that exists only in this file would make that
|
|
274
|
+
// check reject a command that genuinely works.
|
|
275
|
+
const impl = { doctor, verify, audit, issues };
|
|
276
|
+
const missing = COMMAND_NAMES.filter(n => !impl[n]);
|
|
277
|
+
if (missing.length) throw new Error(`cli-spec names commands with no implementation: ${missing}`);
|
|
278
|
+
const table = Object.fromEntries(COMMAND_NAMES.map(n => [n, impl[n]]));
|
|
279
|
+
|
|
233
280
|
if (!table[cmd]) {
|
|
234
281
|
console.log(`
|
|
235
282
|
ca — does a test actually fail on the code it was written to catch?
|
|
236
283
|
|
|
237
|
-
ca doctor
|
|
238
|
-
ca verify <commit> [--runs 3]
|
|
239
|
-
ca audit --n 100 [--since '6
|
|
284
|
+
ca doctor ${COMMANDS.doctor}
|
|
285
|
+
ca verify <commit> [--runs 3] ${COMMANDS.verify}
|
|
286
|
+
ca audit --n 100 [--since '6 years'] [--grep '^fix'] [--seed 1] [--out r.csv]
|
|
240
287
|
|
|
241
288
|
common: --repo <path> --timeout <ms> --keep (leave worktrees for inspection)
|
|
242
289
|
`);
|
package/package.json
CHANGED
|
@@ -1,17 +1,22 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "control-arm",
|
|
3
|
-
"version": "1.
|
|
3
|
+
"version": "1.2.0",
|
|
4
4
|
"description": "Does a test actually fail on the code it was written to catch?",
|
|
5
5
|
"type": "module",
|
|
6
|
-
"bin": {
|
|
6
|
+
"bin": {
|
|
7
|
+
"ca": "./bin/ca.mjs"
|
|
8
|
+
},
|
|
7
9
|
"files": [
|
|
8
10
|
"bin",
|
|
9
11
|
"src",
|
|
10
12
|
"README.md",
|
|
13
|
+
"CHANGELOG.md",
|
|
11
14
|
"DESIGN.md",
|
|
12
15
|
"LICENSE"
|
|
13
16
|
],
|
|
14
|
-
"engines": {
|
|
17
|
+
"engines": {
|
|
18
|
+
"node": ">=20.6.0"
|
|
19
|
+
},
|
|
15
20
|
"keywords": [
|
|
16
21
|
"testing",
|
|
17
22
|
"test-quality",
|
|
@@ -26,7 +31,9 @@
|
|
|
26
31
|
"url": "git+https://github.com/mekanhan/control-arm.git"
|
|
27
32
|
},
|
|
28
33
|
"homepage": "https://github.com/mekanhan/control-arm#readme",
|
|
29
|
-
"bugs": {
|
|
34
|
+
"bugs": {
|
|
35
|
+
"url": "https://github.com/mekanhan/control-arm/issues"
|
|
36
|
+
},
|
|
30
37
|
"scripts": {
|
|
31
38
|
"test": "node --test test/*.test.mjs",
|
|
32
39
|
"fixtures": "node fixtures/build.mjs"
|
package/src/assertions.mjs
CHANGED
|
@@ -245,6 +245,39 @@ function titleMatches(lit, caseName) {
|
|
|
245
245
|
try { return new RegExp('^' + pattern + '$').test(caseName); } catch { return false; }
|
|
246
246
|
}
|
|
247
247
|
|
|
248
|
+
/**
|
|
249
|
+
* The opening brace of the CASE BODY — skipping an options object if one is there.
|
|
250
|
+
*
|
|
251
|
+
* `node:test`, vitest and jest all accept a middle argument:
|
|
252
|
+
*
|
|
253
|
+
* test('name', { skip: SKIP }, () => { …assertions… })
|
|
254
|
+
* test('name', { timeout: 5000 }, async () => { … })
|
|
255
|
+
*
|
|
256
|
+
* Taking the first `{` after the title grabs `{ skip: SKIP }` and analyses THAT as the
|
|
257
|
+
* body. It holds no assertions, so every such case was reported as
|
|
258
|
+
* "the case body contains no assertion at all" — about tests that are full of them.
|
|
259
|
+
*
|
|
260
|
+
* FOUND IN THE WILD, on five tests holding eleven assertions between them. The VERDICT
|
|
261
|
+
* was not affected — that comes from red/green across the two arms — but the REASON was,
|
|
262
|
+
* and a wrong reason attached to a BLIND verdict sends someone to look at the wrong
|
|
263
|
+
* thing. That is the same failure this tool exists to prevent, one level up.
|
|
264
|
+
*/
|
|
265
|
+
function bodyBraceAfter(source, from) {
|
|
266
|
+
let i = from;
|
|
267
|
+
// Walk past `,` and whitespace, stepping over any balanced `{...}` met before the
|
|
268
|
+
// callback. An options object is the only thing that can legally appear there.
|
|
269
|
+
for (let guard = 0; guard < 4; guard++) {
|
|
270
|
+
while (i < source.length && /[\s,]/.test(source[i])) i++;
|
|
271
|
+
if (source[i] !== '{') break;
|
|
272
|
+
const end = matchBrace(source, i);
|
|
273
|
+
if (end === -1) return -1;
|
|
274
|
+
i = end + 1;
|
|
275
|
+
}
|
|
276
|
+
// Now at the callback — `() => {`, `async () => {`, `function () {`. Its body is the
|
|
277
|
+
// next brace, with no options object left to confuse it with.
|
|
278
|
+
return source.indexOf('{', i);
|
|
279
|
+
}
|
|
280
|
+
|
|
248
281
|
function extractExact(source, caseName) {
|
|
249
282
|
CALL.lastIndex = 0;
|
|
250
283
|
let m;
|
|
@@ -253,7 +286,10 @@ function extractExact(source, caseName) {
|
|
|
253
286
|
const lit = readLiteral(source, qi);
|
|
254
287
|
if (!lit) continue;
|
|
255
288
|
if (!titleMatches(lit, caseName)) continue;
|
|
256
|
-
|
|
289
|
+
// +1 because readLiteral returns the index OF the closing quote, not after it.
|
|
290
|
+
// Starting on the quote made the options-object skip a no-op, which is how this
|
|
291
|
+
// fix silently did nothing on its first attempt.
|
|
292
|
+
const open = bodyBraceAfter(source, lit.end + 1);
|
|
257
293
|
if (open === -1) continue;
|
|
258
294
|
const close = matchBrace(source, open);
|
|
259
295
|
// A failed brace scan on ONE site must not abandon the search — a later site may
|
package/src/children.mjs
ADDED
|
@@ -0,0 +1,55 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* Reap the test processes when this process goes away.
|
|
3
|
+
*
|
|
4
|
+
* A machine running these audits was found carrying ten stray `node --test` processes,
|
|
5
|
+
* seven of them TWO DAYS old, each holding a worktree and file descriptors open. They sat
|
|
6
|
+
* at 0% CPU, which is why nothing noticed: the tool looked idle rather than leaky.
|
|
7
|
+
*
|
|
8
|
+
* The cause is NOT the timeout — that path already killed what it spawned. It is the audit
|
|
9
|
+
* being killed itself. Observed directly:
|
|
10
|
+
*
|
|
11
|
+
* before 1242399 ppid 1242397 node --test hangs.test.mjs
|
|
12
|
+
* (kill the parent)
|
|
13
|
+
* after 1242399 ppid 2099 node --test hangs.test.mjs <- survived
|
|
14
|
+
*
|
|
15
|
+
* An audit runs for the better part of an hour, so it gets interrupted often — and each
|
|
16
|
+
* interruption stranded whatever was mid-run. `detached: true` alone makes this WORSE,
|
|
17
|
+
* because a detached child is meant to outlive its parent. So the spawn stays detached (to
|
|
18
|
+
* get a killable process GROUP for runners that fork workers) and every live group is
|
|
19
|
+
* registered here, to be swept when this process ends however it ends.
|
|
20
|
+
*/
|
|
21
|
+
const live = new Set();
|
|
22
|
+
|
|
23
|
+
/** Kill a process group, tolerating one that has already gone. */
|
|
24
|
+
export function killGroup(pid) {
|
|
25
|
+
try { process.kill(-pid, 'SIGKILL'); } catch { /* already gone, or never grouped */ }
|
|
26
|
+
}
|
|
27
|
+
|
|
28
|
+
export function register(pid) { live.add(pid); }
|
|
29
|
+
export function unregister(pid) { live.delete(pid); }
|
|
30
|
+
export function liveCount() { return live.size; }
|
|
31
|
+
|
|
32
|
+
export function reapAll() {
|
|
33
|
+
for (const pid of live) killGroup(pid);
|
|
34
|
+
live.clear();
|
|
35
|
+
}
|
|
36
|
+
|
|
37
|
+
let armed = false;
|
|
38
|
+
/**
|
|
39
|
+
* Idempotent, and deliberately not installed at import time: importing a module should not
|
|
40
|
+
* change how the process handles signals. The runners call this the first time they spawn.
|
|
41
|
+
*/
|
|
42
|
+
export function armReaper() {
|
|
43
|
+
if (armed) return;
|
|
44
|
+
armed = true;
|
|
45
|
+
process.on('exit', reapAll);
|
|
46
|
+
for (const sig of ['SIGINT', 'SIGTERM', 'SIGHUP']) {
|
|
47
|
+
process.on(sig, () => {
|
|
48
|
+
reapAll();
|
|
49
|
+
// Re-raise with the handler removed, so the exit code still says "signalled"
|
|
50
|
+
// rather than pretending this was a clean exit.
|
|
51
|
+
process.removeAllListeners(sig);
|
|
52
|
+
process.kill(process.pid, sig);
|
|
53
|
+
});
|
|
54
|
+
}
|
|
55
|
+
}
|