control-arm 1.0.0 → 1.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md ADDED
@@ -0,0 +1,44 @@
1
+ # Changelog
2
+
3
+ All notable changes to this project are documented in this file.
4
+
5
+ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this
6
+ project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
7
+
8
+ ## [1.2.0] - 2026-10-06
9
+
10
+ > **Note on the missing 1.1.0.** `package.json` was bumped to `1.1.0` on 2026-09-25 but that
11
+ > version was never tagged in git, never published to npm, and never released on GitHub —
12
+ > npm's `latest` stayed on `1.0.0` until now. The `1.1.0` work (vitest and jest runners, the
13
+ > findings-contract v1 envelope, the child-process reaper, and the fixes that shipped with
14
+ > them) is folded into this release.
15
+
16
+ ### Fixed
17
+
18
+ - The module-identity proof silently did nothing on Node < 20.6: `import.meta.resolve` is
19
+ undefined there, and the per-specifier error was swallowed as "genuinely missing", so every
20
+ verdict was "proven" without resolving a single import. The probe now reports when it
21
+ cannot resolve, and the verdict is withheld instead of trusted.
22
+ - The findings-contract `version` field was hardcoded to `1.1.0`; it is now read from
23
+ `package.json` so it cannot drift.
24
+
25
+ ### Changed
26
+
27
+ - Minimum Node version raised to `20.6.0` (was `18`).
28
+
29
+ ### Added
30
+
31
+ - `scripts/precision.mjs`: measure the precision (and recall, when the sample is labelled)
32
+ of the `BLIND` verdict against a hand-labelled corpus.
33
+ - Unit coverage for the vitest/jest report parser (`parseReport`) against captured runner
34
+ output.
35
+
36
+ ### Removed
37
+
38
+ - Dead code: the tautological `nodeTest.matches` predicate, the unused `RUNNERS` export,
39
+ `isFileLevelFailure`, and a leftover destructure in `commitInfo`.
40
+
41
+ ## [1.0.0] - 2026-09-25
42
+
43
+ - Initial public release: `ca doctor` / `ca verify` / `ca audit` over `node:test`, published
44
+ to npm and the GitHub Marketplace (`mekanhan/control-arm@v1`).
package/DESIGN.md CHANGED
@@ -109,11 +109,19 @@ weighting: a codebase the author did not write, did not pick for a flattering re
109
109
  could not tune against. 190 commits matched the subject filter; 99 ship both a test and a
110
110
  source change; 40 were drawn from those.
111
111
 
112
- **One usability trap found while running it.** `audit` defaults to a six-month window. On
113
- a mature repo that silently shrinks the sample — the first dayjs run judged **5 commits
114
- instead of 40** and printed a confident, meaningless 100%. The window *is* echoed in the
115
- header line, so it is visible rather than hidden, but a `--n 40` that quietly returns 5
116
- deserves a louder signal. Pass `--since` explicitly on any repo older than six months.
112
+ **One usability trap found while running it.** `audit` defaults to a **twelve-month**
113
+ window, and dayjs ships few `fix:` commits in a year that also touch a test. The first
114
+ dayjs run judged **5 commits instead of 40** and printed a confident, meaningless 100%.
115
+ Nothing was broken: the pool was simply smaller than the request.
116
+
117
+ What exposed it was that two different `--n` values and two different seeds gave
118
+ **byte-identical output**. A rate that does not move when you change the sample size is
119
+ not a rate.
120
+
121
+ `audit` now says so when the draw comes up short, naming whichever cause applies — the
122
+ window, or the test-plus-source pre-filter — and never blaming a `--since` the caller
123
+ passed themselves (`src/sample-warning.mjs`, `SAMPLE-001..006`). Pass `--since` explicitly
124
+ on any repo with more than a year of history.
117
125
 
118
126
  `INCONCLUSIVE` is 29% and 40% respectively. That is the honest denominator, not a rounding
119
127
  error: a fix that adds an export its test imports cannot be replayed against the parent,
@@ -122,22 +130,30 @@ either way.
122
130
 
123
131
  ## How often is this tool wrong?
124
132
 
125
- Ask any measuring instrument this. Here is the answer for this one, on the 300-commit run.
126
-
127
- The raw headline was **13 `BLIND` commits**. After running arm C on each and checking them
128
- by hand, **5 were real**:
129
-
130
- | | |
131
- |---|---|
132
- | 13 | raw `BLIND` |
133
- | −3 | `↻ REPAIRED SINCE` — the gap was closed after that commit |
134
- | −2 | `? cannot tell` — the test file no longer exists at HEAD |
135
- | −3 | category errors: two `fix(test):` (the bug WAS the test) and one build failure no unit test can catch |
136
- | **5** | genuinely open, **2.3% of answerable commits** |
137
-
138
- **A 62% false-positive rate on the raw number.** Arm C and the corpus filter exist because
139
- of it. Publish the raw count and you hand someone thirteen tickets, eight of which waste
140
- their afternoon.
133
+ **[README.md → "How often is the TOOL right?"](README.md#does-it-work)** holds the audit:
134
+ 13 raw `BLIND` on the 300-commit run, 5 genuinely open after arm C, a 62% false-positive
135
+ rate on the raw number.
136
+
137
+ It lives there and not here because it is the first thing a skeptic should see, and because
138
+ a table kept in two places drifts in one of them. What belongs here is the design
139
+ consequence:
140
+
141
+ **Arm C is not a nicety, it is the reason the raw signal is publishable at all.** Eight of
142
+ those thirteen were not defects in anybody's tests — three gaps had been closed by a later
143
+ commit, two named test files no longer exist, two were `fix(test):` commits where the bug
144
+ *was* the test, and one was a build failure no unit test could catch. None of that is
145
+ visible from the two arms alone; all of it needs a third look at HEAD.
146
+
147
+ That is also why the corpus filter drops `fix(test):` by default, and why
148
+ `--include-test-fixes` exists for anyone who wants them back.
149
+
150
+ **Measuring precision properly.** "5 for 5, too small to quote" is a count, not a rate.
151
+ To turn it into a rate — precision, and recall when the whole sample is labelled — label
152
+ each judged commit `gap`/`not-gap` and run `scripts/precision.mjs`. It reads `ca audit
153
+ --out` plus the labels, and prints the confusion matrix for both the raw `BLIND` signal and
154
+ the post-arm-C "still open" signal. Until that labelled corpus exists, the defensible claim
155
+ is the one the README already makes: the raw signal is over-sensitive, the filters remove
156
+ the eight known false positives, and the residual precision is unmeasured.
141
157
 
142
158
  ## Every bug found in this tool so far
143
159
 
package/README.md CHANGED
@@ -40,7 +40,7 @@ because it runs the new test against the old code.
40
40
  npm i -g control-arm # or run it from a clone: node bin/ca.mjs
41
41
  ```
42
42
 
43
- Node 18+. No other dependencies.
43
+ Node 20.6+. No other dependencies.
44
44
 
45
45
  ## Use it
46
46
 
@@ -87,9 +87,82 @@ It comments on the PR with a line per test. It does **not** fail the build by de
87
87
  PR can legitimately ship only regression guards, and a gate that fires on those gets
88
88
  switched off within a week.
89
89
 
90
+ **On a feature PR it refuses to claim anything.** This matters, because the check is
91
+ meaningless there and would otherwise read as proof. On a feature the base is code where
92
+ the thing does not exist yet, so essentially *any* test touching it fails —
93
+ `expected 0 to be greater than 0` is a real assertion failure that says nothing about
94
+ whether the test is well aimed. So the comment changes its own headline:
95
+
96
+ > ### `control-arm` — new code — these tests cannot be judged this way
97
+ >
98
+ > **3 of 4 tests fail without this change** — but this is a feature, so the base does not
99
+ > have it at all.
100
+ >
101
+ > That is expected, and it is weak evidence: on new code almost any test that reads the new
102
+ > thing fails on the base, whether or not it is well aimed. A test asserting `1 === 1` in
103
+ > the same file would NOT show up here; one that merely reads the new field would.
104
+
105
+ The classification is `fix`/`feat` from the conventional-commit prefix, plus a supporting
106
+ hint — a commit that only *added* source lines and deleted none looks like new code
107
+ regardless of its prefix. Both were measured on a real corpus before being trusted: the
108
+ prefix is used consistently (1,213 `fix` / 988 `feat`), while additions-only fires on 9 of
109
+ 28 features but also on 1 of 35 fixes — **specific, not sensitive**, so it qualifies the
110
+ claim and never decides the verdict on its own. A `CAUGHT` on a feature is still `CAUGHT`;
111
+ it just does not get to say it would have caught a bug, because there was no bug.
112
+
113
+ ## All flags
114
+
115
+ ```bash
116
+ ca doctor --repo . # can this repo be measured?
117
+ ca verify <sha> --repo . --runs 3 --timeout 120000 --work .ca-work
118
+ ca verify <sha> --against origin/main --pr-comment --fail-on-blind
119
+ ca audit --n 100 --since '6 years' --grep '^fix' --seed 1 --out r.csv --html r.html
120
+ ca issues --apply --include-test-fixes
121
+ ```
122
+
123
+ | flag | |
124
+ |---|---|
125
+ | `--repo <path>` | the repository to measure (default: cwd) |
126
+ | `--against <ref>` | compare against a base ref — turns `verify` into a PR check |
127
+ | `--runs <n>` | repeat each case N times to expose flakes |
128
+ | `--timeout <ms>` | per-run timeout |
129
+ | `--work <path>` | where the throwaway worktrees go |
130
+ | `--keep` | leave them behind for inspection |
131
+ | `--n <n>` · `--since <when>` · `--grep <re>` · `--seed <n>` | which commits `audit` samples |
132
+ | `--include-test-fixes` | do not skip `fix(test):` commits, where the bug WAS the test |
133
+ | `--out <file>` · `--html <file>` | write the audit as CSV or HTML |
134
+ | `--pr-comment` | print the PR comment markdown and nothing else |
135
+ | `--apply` | `issues` actually files them, rather than printing |
136
+ | **`--json`** | **findings-contract v1 envelope on stdout** (see below) |
137
+ | `--json <path>` | the older behaviour: write the audit summary to a file |
138
+ | **`--fail-on-blind`** | **exit 1 when a commit is BLIND** |
139
+
140
+ ### The findings contract
141
+
142
+ Bare `--json` emits one [findings-contract v1](https://github.com/mekanhan/findings-contract)
143
+ object on stdout and nothing else — progress goes to stderr, so it pipes. It works on
144
+ `doctor`, `verify` and `audit`.
145
+
146
+ The mapping is not one-to-one, and the interesting part is what is **not** a finding:
147
+
148
+ | | |
149
+ |---|---|
150
+ | commit `CAUGHT` | no finding. Nothing is wrong |
151
+ | commit `BLIND` | a finding, severity `warn` |
152
+ | commit `INCONCLUSIVE` | **not** a finding — it goes to `skipped`, with its reason |
153
+ | case `NON-DISCRIMINATING` | never a finding. A fact about a guard, not a fault |
154
+ | case `FLAKY` | a finding. Runs disagreed, which is a real defect |
155
+
156
+ **Exit codes:** `0` ran cleanly · `1` found blockers · `2` the tool itself failed.
157
+
158
+ `ca verify` used to exit `1` on `BLIND`. It no longer does unless you pass
159
+ `--fail-on-blind`, because `BLIND` is a `warn` — a test that cannot catch a bug breaks
160
+ nothing today — and putting a warning into an exit code makes every caller guess. The
161
+ GitHub Action is unaffected: it reads the verdict from stdout.
162
+
90
163
  ## What you get
91
164
 
92
- > ### `control-arm` — this branch is proven
165
+ > ### `control-arm` — 1 test here fails without this change
93
166
  >
94
167
  > **1 of 5 tests here genuinely catch this change.** Run against the code as it was before
95
168
  > this PR, they fail; with the change, they pass.
@@ -121,41 +194,149 @@ measurement that cannot tell zero from failure is not a measurement.
121
194
 
122
195
  ## Which of your tests it looks at
123
196
 
124
- It does not care what *kind* of test it is, only whether the runner can execute the file.
197
+ **Unit and integration tests** — anything that runs in-process, without a deployed
198
+ application. That is the whole scope.
125
199
 
126
- | kind | covered | why |
127
- |---|---|---|
128
- | unit tests | yes | the easy case |
129
- | backend / server logic | yes | same runner, same rewind |
130
- | database-backed tests | yes | needs a live DB, else they report `SKIPPED` |
131
- | component tests (React / RN) | yes | via vitest and jest |
132
- | browser e2e (Playwright) | **no** | no Playwright runner yet |
133
- | performance / load | **no** | they measure speed, not correctness |
134
- | manual QA | **no** | nothing to execute |
200
+ **Every row carries what proves it.** A row without evidence beside it is an aspiration, and
201
+ that is where this table was wrong before: vitest and jest were listed as covered because the
202
+ code has adapters for them, not because either had ever been run.
203
+
204
+ | runner | proven by |
205
+ |---|---|
206
+ | `node:test` | fixtures `01`–`12`, built and run by `npm test` on every push |
207
+ | vitest | observed on a real commit, 2026-09-25 — **no fixture yet** (#21) |
208
+ | jest | observed on a real commit, 2026-09-25 — **no fixture yet** (#21) |
209
+
210
+ A database-backed test is judged like any other, and needs its database present. The README
211
+ used to claim it reports `SKIPPED` without one; that has never actually been observed, so the
212
+ claim is withdrawn until it is (#21).
213
+
214
+ **"fixtures"** means anyone can re-run it: `npm test` builds throwaway repos where the right
215
+ answer is known by construction, and the suite fails if this tool cannot tell
216
+ `02-blind-direction` from `01-caught-value`. **"observed on a real commit"** means it worked
217
+ once, on a repository you cannot see — weaker, and marked weaker.
218
+
219
+ ### Out of scope, permanently
220
+
221
+ Browser e2e (Playwright), device e2e (WDIO/Appium) and load tests (k6).
222
+
223
+ Arm B has to run your test against the **old code**, and for an e2e test the old code is a
224
+ *running application* — built and served at the parent commit, with its database and its
225
+ services. That is minutes per commit at best, and frequently the parent will not boot at all.
226
+ Load tests measure speed rather than correctness, so the question this tool asks does not
227
+ apply to them.
228
+
229
+ Point it at one of those and it will name the runner and decline, rather than attempting it
230
+ with the wrong one. That is a courtesy, not a roadmap — none of these is planned.
135
231
 
136
232
  ## Does it work?
137
233
 
138
- Three repositories, random draws, seeds recorded so the samples are reproducible.
234
+ **Two different questions live here, and mixing them flatters the tool.** How good are the
235
+ tests it measured, and how often is the tool itself right? Only the second is about
236
+ `control-arm`.
237
+
238
+ ### How often is the TOOL right?
239
+
240
+ This is the number to judge it on. On the 300-commit run it raised **13 `BLIND` verdicts**
241
+ — "not one test in this commit would have caught the bug". Each was then re-run through
242
+ arm C and checked by hand:
243
+
244
+ | | |
245
+ |---|---|
246
+ | 13 | raw `BLIND` |
247
+ | −3 | `↻ REPAIRED SINCE` — the gap was real, and closed after that commit |
248
+ | −2 | `? cannot tell` — the test file no longer exists at HEAD |
249
+ | −3 | category errors: two `fix(test):` (the bug WAS the test) and one build failure no unit test can catch |
250
+ | **5** | genuinely open — **2.3% of answerable commits** |
251
+
252
+ **A 62% false-positive rate on the raw number.** Publish it and you hand someone thirteen
253
+ tickets, eight of which waste their afternoon. Arm C and the corpus filter exist entirely
254
+ because of that, and they run by default.
255
+
256
+ All five survivors held up under hand inspection. **That is 5 for 5 out of 13 candidates —
257
+ far too small a sample to quote as a precision rate**, and it is stated as a count for that
258
+ reason. The honest summary is: the raw signal is badly over-sensitive, the filters remove
259
+ eight of eight known false positives, and what precision remains after them has not been
260
+ measured on a sample large enough to have a rate.
261
+
262
+ `BLIND` is the only verdict that accuses anybody, so every ambiguity resolves away from it.
263
+
264
+ ### How good were the TESTS it measured?
265
+
266
+ These say nothing about the tool's accuracy — they are a property of the repositories.
267
+ Counts, not just percentages, because a percentage over 25 commits invites arithmetic
268
+ nobody should have to do themselves.
139
269
 
140
270
  | | a private monorepo | nodejs/undici | iamkun/dayjs |
141
271
  |---|---|---|---|
142
272
  | fix commits sampled | 300 | 25 | 40 |
143
273
  | answerable | 214 | 15 | 25 |
144
- | **CAUGHT** | **93.9%** | **80.0%** | **92.0%** |
145
- | BLIND | 6.1% | 20.0% | 8.0% |
274
+ | **CAUGHT** | **201 — 93.9%** | **12 — 80.0%** | **23 — 92.0%** |
275
+ | BLIND | 13 — 6.1% | **3** — 20.0% | **2** — 8.0% |
276
+ | INCONCLUSIVE | 86 | 10 | 9 |
146
277
  | runtime | 7.1 s/commit | 24.9 s/commit | 1.4 s/commit |
147
278
 
148
- **The dayjs column is the one to weigh.** It is a codebase the author did not write, did
149
- not choose for a flattering result, and could not tune the tool against — 190 fix commits
150
- matched the filter, 99 ship both a test and a source change, 40 drawn at random with the
151
- seed recorded. It lands within a point of the private monorepo it was built on.
279
+ **The small columns are small.** dayjs's 8% is **two commits** and undici's 20% is **three**.
280
+ Those are proof-of-concept numbers; do not read a difference between 80% and 92% as a
281
+ difference between the repositories.
282
+
283
+ **The dayjs column is the one to weigh** anyway: a codebase the author did not write, did
284
+ not choose for a flattering result, and could not tune against — 190 fix commits matched
285
+ the filter, 99 ship both a test and a source change, 40 drawn at random with the seed
286
+ recorded. It lands within a point of the private monorepo it was built on.
287
+
288
+ ### Why is so much unanswerable?
289
+
290
+ 29% and 40% is a lot to exclude, so here is where it goes. `INCONCLUSIVE` is never folded
291
+ into the other columns — a measurement that cannot tell zero from failure is not a
292
+ measurement — and there are exactly four ways to earn it:
293
+
294
+ | cause | what happened |
295
+ |---|---|
296
+ | **the test will not load on the parent** | the dominant one. The fix *added* an export, a module, a fixture; the test imports it; on the parent that import throws before a single assertion runs. `SyntaxError: does not provide an export named …` is not a test failing, it is a test never starting |
297
+ | **the test is not green on the fix either** | no before/after to compare. Usually an environment gap — a database, a browser, a missing `.env` |
298
+ | **module identity unprovable** | the run could not be shown to have loaded the worktree's code rather than the real repo's. Refused rather than guessed |
299
+ | **a skipped case is present** | the skipped one might have been the discriminating one, so the commit cannot be called `BLIND` |
300
+
301
+ The breakdown *between* these four has not been counted per repository, so no split is
302
+ quoted here.
303
+
304
+ ## Working on this repo
305
+
306
+ ```bash
307
+ git config core.hooksPath .githooks # once per clone; worktrees inherit it
308
+ ```
309
+
310
+ `pre-push` refuses three things CI can only tell you about after the fact: a direct push to
311
+ `main` (everything lands through a PR), a branch **named** like a default that is not this
312
+ repo's default, and a push from a base that has already moved.
313
+
314
+ The middle one is not hypothetical. A local `master` once sat ten commits behind `main`
315
+ while `main` moved on through four PRs; pushing it created a parallel remote branch, and the
316
+ Node 20/22/24 matrix came back green against the wrong base. The only tell was a line of
317
+ push output — `* [new branch] master -> master` — on a repo that is anything but new.
318
+
319
+ If a push prints none of that hook's output, it is not armed. Check `core.hooksPath` before
320
+ trusting anything it did not say.
321
+
322
+ ## Prior art
323
+
324
+ The fail-before / pass-after check is not new, and it is worth saying who got there first:
325
+
326
+ - **[Defects4J](https://github.com/rjust/defects4j)** records, for each reproducible bug,
327
+ the **trigger tests** that fail on the buggy version and pass on the fixed one.
328
+ - **[SWE-bench](https://www.swebench.com/)** builds every task around **`FAIL_TO_PASS`**
329
+ tests, with `PASS_TO_PASS` as the regression guard — the same two arms.
152
330
 
153
- Of the 13 raw `BLIND` verdicts in that 300-commit run, **5 were genuinely open** — a 62%
154
- false-positive rate on the raw number, which is why the filters exist.
331
+ Both use the property to **construct benchmarks**: they start from a known bug and keep the
332
+ tests that prove it. `control-arm` runs the same check in the other direction — on ordinary
333
+ commits, in CI, where the answer is not known in advance and a `BLIND` result is news
334
+ rather than a data-cleaning step.
155
335
 
156
- **[DESIGN.md](DESIGN.md) has the rest**: how the rewind works, the three traps that
157
- separate an instrument from a random number generator, the full accuracy audit, and every
158
- bug found in this tool so far — none of which were found by its own test suite.
336
+ The consequence of that difference is most of this repository: a benchmark can discard
337
+ anything ambiguous, because it only needs *some* clean examples. A CI check cannot discard
338
+ the commit in front of it, so it has to be able to say `INCONCLUSIVE` out loud, and it has
339
+ to be right when it says `BLIND` about somebody's work.
159
340
 
160
341
  ## Scope, stated plainly
161
342
 
package/bin/ca.mjs CHANGED
@@ -16,11 +16,28 @@ import { renderVerify, renderAudit, MARK } from '../src/report.mjs';
16
16
  import { renderHtml, issueBody } from '../src/html-report.mjs';
17
17
  import { prComment } from '../src/markdown-report.mjs';
18
18
  import { analyseCase, extractCase } from '../src/assertions.mjs';
19
+ import { COMMANDS, COMMAND_NAMES } from '../src/cli-spec.mjs';
20
+ import { sampleWarning, relativeWindowWarning } from '../src/sample-warning.mjs';
21
+ import { verifyEnvelope, auditEnvelope, doctorEnvelope } from '../src/contract.mjs';
19
22
 
20
23
  const argv = process.argv.slice(2);
21
24
  const cmd = argv[0];
22
25
  const flag = (n, d = null) => { const i = argv.indexOf(`--${n}`); return i === -1 ? d : argv[i + 1]; };
23
26
  const has = n => argv.includes(`--${n}`);
27
+ // `--json` predates the contract and took a PATH. Bare `--json` — no value, or the next
28
+ // token is another flag — now means "contract envelope on stdout", which is what C-001
29
+ // asks for. `--json <path>` is unchanged, so nothing that worked stops working.
30
+ const bareJson = (() => {
31
+ const i = argv.indexOf('--json');
32
+ if (i === -1) return false;
33
+ const next = argv[i + 1];
34
+ return next === undefined || next.startsWith('-');
35
+ })();
36
+ /** C-001: one object on stdout, nothing else. Then C-007 decides the code. */
37
+ const emit = (env, gate = false) => {
38
+ process.stdout.write(JSON.stringify(env, null, 2) + '\n');
39
+ process.exit(gate && env.summary.blocker ? 1 : 0);
40
+ };
24
41
  const repo = path.resolve(flag('repo', process.cwd()));
25
42
  const workDir = path.resolve(flag('work', path.join(repo, '.ca-work')));
26
43
 
@@ -59,6 +76,8 @@ async function doctor() {
59
76
  checks.push({ ok: true, name: 'workspaces', detail: `${[].concat(pkg.workspaces).join(', ')} — workspace links will be re-pointed into the worktree (this is the trap that produces false BLIND)` });
60
77
  }
61
78
 
79
+ if (bareJson) emit(doctorEnvelope(checks, { repo }), true);
80
+
62
81
  const w = 22;
63
82
  console.log(`\n ca doctor — ${repo}\n`);
64
83
  for (const c of checks) console.log(` ${c.ok ? MARK.ok : MARK.no} ${c.name.padEnd(w)} ${c.detail}`);
@@ -74,10 +93,18 @@ async function verify() {
74
93
  const r = await verifyCommit({ repo, workDir, sha, against: flag('against'), runs, timeoutMs: Number(flag('timeout', 120_000)),
75
94
  onStep: s => process.stderr.write(`\r … ${s} `) });
76
95
  process.stderr.write('\r' + ' '.repeat(40) + '\r');
96
+ if (bareJson) { if (!has('keep')) await removeWorktrees(repo, workDir); emit(verifyEnvelope(r, { repo, sha })); }
77
97
  if (has('pr-comment')) console.log(prComment(r, { repoName: flag('repo', '.') }));
78
98
  else console.log(renderVerify(r));
79
99
  if (!has('keep')) await removeWorktrees(repo, workDir);
80
- process.exit(r.verdict === BLIND ? 1 : 0);
100
+
101
+ // C-007. This used to exit 1 on BLIND unconditionally, which made a `warn` look like
102
+ // a blocker and put this tool's own opinion into an exit code every caller has to
103
+ // interpret. Gating is opt-in now, under the same name the Action already uses.
104
+ // BEHAVIOUR CHANGE: `ca verify` on a BLIND commit exits 0 unless --fail-on-blind.
105
+ // The Action is unaffected — it reads the verdict from stdout and already wraps the
106
+ // call in `|| true`.
107
+ process.exit(has('fail-on-blind') && r.verdict === BLIND ? 1 : 0);
81
108
  }
82
109
 
83
110
  async function audit() {
@@ -115,7 +142,17 @@ async function audit() {
115
142
  const pool = [...eligible];
116
143
  for (let i = pool.length - 1; i > 0; i--) { const j = Math.floor(rand() * (i + 1)); [pool[i], pool[j]] = [pool[j], pool[i]]; }
117
144
  const sample = pool.slice(0, n);
118
- process.stderr.write(` drawing ${sample.length} at random (seed ${seed})\n\n`);
145
+ process.stderr.write(` drawing ${sample.length} at random (seed ${seed})\n`);
146
+
147
+ // A draw that came up short is the difference between a measurement and a number.
148
+ const moving = relativeWindowWarning(since, argv.includes('--seed'));
149
+ if (moving) process.stderr.write(`\n${moving}\n`);
150
+
151
+ const short = sampleWarning({
152
+ requested: n, matched: candidates.length, eligible: eligible.length,
153
+ drawn: sample.length, since, sinceWasExplicit: argv.includes('--since'),
154
+ });
155
+ process.stderr.write(short ? `\n${short}\n\n` : '\n');
119
156
 
120
157
  const results = [];
121
158
  const t0 = Date.now();
@@ -130,6 +167,8 @@ async function audit() {
130
167
  }
131
168
  process.stderr.write('\r' + ' '.repeat(120) + '\r');
132
169
 
170
+ if (bareJson) emit(auditEnvelope(results, { repo, since, n: sample.length, seed }));
171
+
133
172
  console.log(renderAudit(results, { since, n: sample.length, eligible: eligible.length, matched: candidates.length, seed, seconds: (Date.now() - t0) / 1000 }));
134
173
 
135
174
  const htmlOut = flag('html');
@@ -180,6 +219,7 @@ async function audit() {
180
219
  caught_pct: answerable ? Number(((results.filter(r => r.verdict === CAUGHT).length / answerable) * 100).toFixed(1)) : null,
181
220
  still_open: byStatus('open'),
182
221
  repaired_since: byStatus('repaired'),
222
+ likely_repaired_elsewhere: byStatus('repaired-elsewhere'),
183
223
  cannot_tell: results.filter(r => r.verdict === BLIND && (!r.stillOpen || r.stillOpen.status === 'unknown'))
184
224
  .map(r => ({ sha: r.sha, short: r.short, subject: r.subject })),
185
225
  }, null, 2));
@@ -229,14 +269,21 @@ async function issues() {
229
269
  if (!has('keep')) await removeWorktrees(repo, workDir);
230
270
  }
231
271
 
232
- const table = { doctor, verify, audit, issues };
272
+ // Built from the spec, not written out again here. `test/readme.test.mjs` checks the
273
+ // docs against COMMANDS, so a command that exists only in this file would make that
274
+ // check reject a command that genuinely works.
275
+ const impl = { doctor, verify, audit, issues };
276
+ const missing = COMMAND_NAMES.filter(n => !impl[n]);
277
+ if (missing.length) throw new Error(`cli-spec names commands with no implementation: ${missing}`);
278
+ const table = Object.fromEntries(COMMAND_NAMES.map(n => [n, impl[n]]));
279
+
233
280
  if (!table[cmd]) {
234
281
  console.log(`
235
282
  ca — does a test actually fail on the code it was written to catch?
236
283
 
237
- ca doctor can this repo be measured?
238
- ca verify <commit> [--runs 3] one commit, per-case verdicts
239
- ca audit --n 100 [--since '6 months'] [--grep '^fix'] [--seed 1] [--out r.csv]
284
+ ca doctor ${COMMANDS.doctor}
285
+ ca verify <commit> [--runs 3] ${COMMANDS.verify}
286
+ ca audit --n 100 [--since '6 years'] [--grep '^fix'] [--seed 1] [--out r.csv]
240
287
 
241
288
  common: --repo <path> --timeout <ms> --keep (leave worktrees for inspection)
242
289
  `);
package/package.json CHANGED
@@ -1,17 +1,22 @@
1
1
  {
2
2
  "name": "control-arm",
3
- "version": "1.0.0",
3
+ "version": "1.2.0",
4
4
  "description": "Does a test actually fail on the code it was written to catch?",
5
5
  "type": "module",
6
- "bin": { "ca": "./bin/ca.mjs" },
6
+ "bin": {
7
+ "ca": "./bin/ca.mjs"
8
+ },
7
9
  "files": [
8
10
  "bin",
9
11
  "src",
10
12
  "README.md",
13
+ "CHANGELOG.md",
11
14
  "DESIGN.md",
12
15
  "LICENSE"
13
16
  ],
14
- "engines": { "node": ">=18" },
17
+ "engines": {
18
+ "node": ">=20.6.0"
19
+ },
15
20
  "keywords": [
16
21
  "testing",
17
22
  "test-quality",
@@ -26,7 +31,9 @@
26
31
  "url": "git+https://github.com/mekanhan/control-arm.git"
27
32
  },
28
33
  "homepage": "https://github.com/mekanhan/control-arm#readme",
29
- "bugs": { "url": "https://github.com/mekanhan/control-arm/issues" },
34
+ "bugs": {
35
+ "url": "https://github.com/mekanhan/control-arm/issues"
36
+ },
30
37
  "scripts": {
31
38
  "test": "node --test test/*.test.mjs",
32
39
  "fixtures": "node fixtures/build.mjs"
@@ -245,6 +245,39 @@ function titleMatches(lit, caseName) {
245
245
  try { return new RegExp('^' + pattern + '$').test(caseName); } catch { return false; }
246
246
  }
247
247
 
248
+ /**
249
+ * The opening brace of the CASE BODY — skipping an options object if one is there.
250
+ *
251
+ * `node:test`, vitest and jest all accept a middle argument:
252
+ *
253
+ * test('name', { skip: SKIP }, () => { …assertions… })
254
+ * test('name', { timeout: 5000 }, async () => { … })
255
+ *
256
+ * Taking the first `{` after the title grabs `{ skip: SKIP }` and analyses THAT as the
257
+ * body. It holds no assertions, so every such case was reported as
258
+ * "the case body contains no assertion at all" — about tests that are full of them.
259
+ *
260
+ * FOUND IN THE WILD, on five tests holding eleven assertions between them. The VERDICT
261
+ * was not affected — that comes from red/green across the two arms — but the REASON was,
262
+ * and a wrong reason attached to a BLIND verdict sends someone to look at the wrong
263
+ * thing. That is the same failure this tool exists to prevent, one level up.
264
+ */
265
+ function bodyBraceAfter(source, from) {
266
+ let i = from;
267
+ // Walk past `,` and whitespace, stepping over any balanced `{...}` met before the
268
+ // callback. An options object is the only thing that can legally appear there.
269
+ for (let guard = 0; guard < 4; guard++) {
270
+ while (i < source.length && /[\s,]/.test(source[i])) i++;
271
+ if (source[i] !== '{') break;
272
+ const end = matchBrace(source, i);
273
+ if (end === -1) return -1;
274
+ i = end + 1;
275
+ }
276
+ // Now at the callback — `() => {`, `async () => {`, `function () {`. Its body is the
277
+ // next brace, with no options object left to confuse it with.
278
+ return source.indexOf('{', i);
279
+ }
280
+
248
281
  function extractExact(source, caseName) {
249
282
  CALL.lastIndex = 0;
250
283
  let m;
@@ -253,7 +286,10 @@ function extractExact(source, caseName) {
253
286
  const lit = readLiteral(source, qi);
254
287
  if (!lit) continue;
255
288
  if (!titleMatches(lit, caseName)) continue;
256
- const open = source.indexOf('{', lit.end);
289
+ // +1 because readLiteral returns the index OF the closing quote, not after it.
290
+ // Starting on the quote made the options-object skip a no-op, which is how this
291
+ // fix silently did nothing on its first attempt.
292
+ const open = bodyBraceAfter(source, lit.end + 1);
257
293
  if (open === -1) continue;
258
294
  const close = matchBrace(source, open);
259
295
  // A failed brace scan on ONE site must not abandon the search — a later site may
@@ -0,0 +1,55 @@
1
+ /**
2
+ * Reap the test processes when this process goes away.
3
+ *
4
+ * A machine running these audits was found carrying ten stray `node --test` processes,
5
+ * seven of them TWO DAYS old, each holding a worktree and file descriptors open. They sat
6
+ * at 0% CPU, which is why nothing noticed: the tool looked idle rather than leaky.
7
+ *
8
+ * The cause is NOT the timeout — that path already killed what it spawned. It is the audit
9
+ * being killed itself. Observed directly:
10
+ *
11
+ * before 1242399 ppid 1242397 node --test hangs.test.mjs
12
+ * (kill the parent)
13
+ * after 1242399 ppid 2099 node --test hangs.test.mjs <- survived
14
+ *
15
+ * An audit runs for the better part of an hour, so it gets interrupted often — and each
16
+ * interruption stranded whatever was mid-run. `detached: true` alone makes this WORSE,
17
+ * because a detached child is meant to outlive its parent. So the spawn stays detached (to
18
+ * get a killable process GROUP for runners that fork workers) and every live group is
19
+ * registered here, to be swept when this process ends however it ends.
20
+ */
21
+ const live = new Set();
22
+
23
+ /** Kill a process group, tolerating one that has already gone. */
24
+ export function killGroup(pid) {
25
+ try { process.kill(-pid, 'SIGKILL'); } catch { /* already gone, or never grouped */ }
26
+ }
27
+
28
+ export function register(pid) { live.add(pid); }
29
+ export function unregister(pid) { live.delete(pid); }
30
+ export function liveCount() { return live.size; }
31
+
32
+ export function reapAll() {
33
+ for (const pid of live) killGroup(pid);
34
+ live.clear();
35
+ }
36
+
37
+ let armed = false;
38
+ /**
39
+ * Idempotent, and deliberately not installed at import time: importing a module should not
40
+ * change how the process handles signals. The runners call this the first time they spawn.
41
+ */
42
+ export function armReaper() {
43
+ if (armed) return;
44
+ armed = true;
45
+ process.on('exit', reapAll);
46
+ for (const sig of ['SIGINT', 'SIGTERM', 'SIGHUP']) {
47
+ process.on(sig, () => {
48
+ reapAll();
49
+ // Re-raise with the handler removed, so the exit code still says "signalled"
50
+ // rather than pretending this was a clean exit.
51
+ process.removeAllListeners(sig);
52
+ process.kill(process.pid, sig);
53
+ });
54
+ }
55
+ }