control-arm 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,82 @@
1
+ # The tool that asks whether your tests can fail, running its own.
2
+ #
3
+ # It shipped for a day without this. That is the defect it exists to find, one level up:
4
+ # a suite that is never executed by anything but its author's terminal is not a gate, and
5
+ # "46/46 green" meant "green on one laptop, when I remembered".
6
+ #
7
+ # GitHub-hosted runners on purpose: control-arm has no dependencies and must work on a
8
+ # stock box. Pinning it to a self-hosted fleet would hide exactly the assumptions —
9
+ # a preinstalled binary, a warm cache, a particular git version — that break for the
10
+ # first stranger who clones it.
11
+ name: CI
12
+
13
+ on:
14
+ push:
15
+ branches: [master, main]
16
+ pull_request:
17
+ workflow_dispatch:
18
+
19
+ concurrency:
20
+ group: ci-${{ github.ref }}
21
+ cancel-in-progress: true
22
+
23
+ jobs:
24
+ test:
25
+ name: node ${{ matrix.node }}
26
+ runs-on: ubuntu-latest
27
+ strategy:
28
+ fail-fast: false
29
+ matrix:
30
+ # 20 is the oldest LTS with a stable node:test reporter API, which the TAP parser
31
+ # depends on. 24 is what it is developed against. A break in either is worth knowing.
32
+ node: ['20', '22', '24']
33
+ steps:
34
+ - uses: actions/checkout@v4
35
+ with:
36
+ # FULL HISTORY, not the default shallow clone. The dogfood step verifies the tool
37
+ # against one of its OWN past commits, and `git show 1845c1d3` on a depth-1
38
+ # checkout fails with "unknown revision". Caught by this workflow's first run,
39
+ # which is the argument for having it.
40
+ fetch-depth: 0
41
+ - uses: actions/setup-node@v4
42
+ with:
43
+ node-version: ${{ matrix.node }}
44
+
45
+ - name: Unit + fixtures
46
+ run: node --test --test-concurrency=1 test/*.test.mjs
47
+
48
+ # DOGFOOD. The fixtures prove the verdicts on repos built to have known answers.
49
+ # This proves the whole two-arm machinery works on a REAL history with real commits,
50
+ # worktrees and module resolution — which is where every bug so far has come from.
51
+ - name: Judge its own history
52
+ run: |
53
+ set -euo pipefail
54
+ git config --global user.email ci@control-arm
55
+ git config --global user.name ci
56
+
57
+ OUT=$(node bin/ca.mjs verify 1845c1d3 --repo "$GITHUB_WORKSPACE" --timeout 120000)
58
+ echo "$OUT"
59
+
60
+ # That commit added the rule "a SKIPPED case blocks BLIND", with a test written
61
+ # for it. If the tool cannot still see that test discriminate, the tool is broken
62
+ # — regardless of what its unit suite says.
63
+ # Match the VERDICT WORD, not the whole line. This grep was 'VERDICT CAUGHT'
64
+ # and broke the moment a glyph was added between them — a cosmetic change that
65
+ # failed a correctness gate, which trains people to edit the gate rather than
66
+ # believe it. Anchor on what the check is actually about.
67
+ echo "$OUT" | grep -qE '^ *VERDICT .*\bCAUGHT\b' \
68
+ || { echo "::error::control-arm no longer judges its own fix correctly"; exit 1; }
69
+ # And it must be the STRONG claim: 1845c1d3 is a repair, so a "weak evidence"
70
+ # qualifier here would mean the new-code heuristic has started misfiring on fixes.
71
+ echo "$OUT" | grep -q 'weak evidence' \
72
+ && { echo "::error::a repair was labelled new code — the kind heuristic is wrong"; exit 1; }
73
+ echo "$OUT" | grep -q 'module identity verified' \
74
+ || { echo "::error::module identity was not proven — a verdict here is not trustworthy"; exit 1; }
75
+
76
+ - name: Determinism
77
+ run: |
78
+ set -euo pipefail
79
+ A=$(node bin/ca.mjs verify 1845c1d3 --repo "$GITHUB_WORKSPACE" --work /tmp/d1 --timeout 120000 | tail -14)
80
+ B=$(node bin/ca.mjs verify 1845c1d3 --repo "$GITHUB_WORKSPACE" --work /tmp/d2 --timeout 120000 | tail -14)
81
+ [ "$A" = "$B" ] || { echo "::error::two runs disagreed — the tool is not deterministic"; exit 1; }
82
+ echo "two independent runs are byte-identical"
package/DESIGN.md ADDED
@@ -0,0 +1,161 @@
1
+ # How it works, and why you should believe it
2
+
3
+ [README](README.md) covers what `ca` is and how to run it. This is the part a skeptic
4
+ needs: the mechanism, the traps that separate an instrument from a random number
5
+ generator, and an honest account of how often it is wrong.
6
+
7
+ ## The mechanism
8
+
9
+ You do **not** run the parent commit's tests — the parent does not have this test.
10
+ You take the new test and ask it about the old code.
11
+
12
+ ```
13
+ commit C ("fix: …")
14
+ │
15
+ ├── [worktree @ C] run the new test on the FIXED code ──► must be GREEN
16
+ │ └ not green? INCONCLUSIVE: no before/after to compare
17
+ │
18
+ ├── [worktree @ C~1] transplant ONLY the test file onto the BROKEN code
19
+ │ └ prove module identity, then run
20
+ │ ┌───────────────┬──────────────┬──────────────────┐
21
+ │ ▼ ▼ ▼ ▼
22
+ │ AssertionError passed SyntaxError/ timeout/crash
23
+ │ │ │ import failure │
24
+ │ ▼ ▼ ▼ ▼
25
+ │ CAUGHT NON-DISCRIM. INCONCLUSIVE INCONCLUSIVE
26
+ │
27
+ └── N runs disagree? ──► FLAKY
28
+ ```
29
+
30
+ ## Three things that decide whether this is an instrument or a random number generator
31
+
32
+ ### 1. Exit code is not the signal
33
+
34
+ A test that "fails" at the parent with `SyntaxError: does not provide an export named
35
+ 'TITLE_SEPARATOR'` — because the fix *added* that export — exits non-zero and has told you
36
+ nothing. It never ran. Read the exit code only and your headline number is inflated
37
+ garbage.
38
+
39
+ The discriminator is **assertion-failure vs error-failure**, which every test runner
40
+ reports and almost nothing reads.
41
+
42
+ ### 2. Prove the control arm loaded the control arm's code
43
+
44
+ On its first real run this tool reported `4 pass / 0 fail` on the parent — a `BLIND`
45
+ verdict on one of the most carefully tested commits in the target repo. The harness had
46
+ symlinked the repo's whole `node_modules` into the worktree, and workspace packages inside
47
+ it are symlinks *relative to the real repo*:
48
+
49
+ ```
50
+ node_modules/@acme/core -> ../../packages/core
51
+ resolved: file:///…/the-repo/packages/core/src/parser.js ← today's code
52
+ ```
53
+
54
+ Arm B never loaded the code under test. Every language has this trap somewhere — venvs,
55
+ module caches, classpaths. So `ca` asks the runtime where each of the test's imports
56
+ actually resolved, and **refuses to answer** if any landed outside the worktree.
57
+
58
+ ### 3. A false `BLIND` is the worst output this tool can produce
59
+
60
+ `CAUGHT` is good news. `INCONCLUSIVE` is an honest shrug. **`BLIND` accuses an engineer of
61
+ writing a test that cannot fail.** Get it wrong once and nobody trusts the tool again.
62
+
63
+ So **every ambiguity resolves away from `BLIND`**, and the verdict is withheld rather than
64
+ guessed.
65
+
66
+ That is also why the case level and the commit level use different words. The first real
67
+ run printed this:
68
+
69
+ ```
70
+ ✗ BLIND TITLE-022 CONTROL ARM: the whitespace-only patterns really do split one phrase
71
+ ✗ BLIND TITLE-022: the prefix test survives the wider separators
72
+ ✗ BLIND TITLE-022: the brands the separator change must not touch
73
+ ```
74
+
75
+ All three are **regression guards**. Staying green on both arms is their entire job.
76
+ Red/green cannot distinguish "meant to catch this bug and failed" from "meant to stay
77
+ green" — that is *intent*, and no tool can read it. So a case states the fact
78
+ (`NON-DISCRIMINATING`) and only a **commit** gets judged `BLIND`, when not one of its cases
79
+ discriminated.
80
+
81
+ ## Why you should believe a verdict
82
+
83
+ `npm test` runs eight fixture repos where the right answer is known by construction —
84
+ including three different shapes of blind test and two different shapes of unrunnable test.
85
+ **If `ca` cannot tell `02-blind-direction` from `01-caught-value`, it does not ship.**
86
+
87
+ The verdict engine is pure and has no I/O, so it is tested exhaustively. It also keeps a
88
+ control arm executing the naive exit-code implementation and asserting it still gets the
89
+ answers wrong.
90
+
91
+ That is the standard this tool asks of other people's tests, so it is the standard its own
92
+ suite is held to.
93
+
94
+ ## Measured results
95
+
96
+ Three repositories, random draws, seeds recorded so the samples are reproducible.
97
+
98
+ | | a private monorepo | nodejs/undici | iamkun/dayjs |
99
+ |---|---|---|---|
100
+ | fix commits sampled | 300 | 25 | 40 |
101
+ | answerable | 214 | 15 | 25 |
102
+ | **CAUGHT** | **201 — 93.9%** | **12 — 80.0%** | **23 — 92.0%** |
103
+ | BLIND | 13 — 6.1% | 3 — 20.0% | 2 — 8.0% |
104
+ | INCONCLUSIVE | 86 | 10 | 9 |
105
+ | runtime | 7.1 s/commit | 24.9 s/commit | 1.4 s/commit |
106
+
107
+ The dayjs run (`--n 40 --since '6 years' --grep '^fix' --seed 11`) is the one worth
108
+ weighting: a codebase the author did not write, did not pick for a flattering result, and
109
+ could not tune against. 190 commits matched the subject filter; 99 ship both a test and a
110
+ source change; 40 were drawn from those.
111
+
112
+ **One usability trap found while running it.** `audit` defaults to a six-month window. On
113
+ a mature repo that silently shrinks the sample — the first dayjs run judged **5 commits
114
+ instead of 40** and printed a confident, meaningless 100%. The window *is* echoed in the
115
+ header line, so it is visible rather than hidden, but a `--n 40` that quietly returns 5
116
+ deserves a louder signal. Pass `--since` explicitly on any repo older than six months.
117
+
118
+ `INCONCLUSIVE` is 29% and 40% respectively. That is the honest denominator, not a rounding
119
+ error: a fix that adds an export its test imports cannot be replayed against the parent,
120
+ because the test will not load there. Those are excluded from the ratio rather than assumed
121
+ either way.
122
+
123
+ ## How often is this tool wrong?
124
+
125
+ Ask any measuring instrument this. Here is the answer for this one, on the 300-commit run.
126
+
127
+ The raw headline was **13 `BLIND` commits**. After running arm C on each and checking them
128
+ by hand, **5 were real**:
129
+
130
+ | | |
131
+ |---|---|
132
+ | 13 | raw `BLIND` |
133
+ | −3 | `↻ REPAIRED SINCE` — the gap was closed after that commit |
134
+ | −2 | `? cannot tell` — the test file no longer exists at HEAD |
135
+ | −3 | category errors: two `fix(test):` (the bug WAS the test) and one build failure no unit test can catch |
136
+ | **5** | genuinely open, **2.3% of answerable commits** |
137
+
138
+ **A 62% false-positive rate on the raw number.** Arm C and the corpus filter exist because
139
+ of it. Publish the raw count and you hand someone thirteen tickets, eight of which waste
140
+ their afternoon.
141
+
142
+ ## Every bug found in this tool so far
143
+
144
+ None were found by its own test suite. All were found by pointing it at real work.
145
+
146
+ | bug | found by |
147
+ |---|---|
148
+ | a slash in a test NAME silently dropped the case | random hand-audit of `BLIND` verdicts |
149
+ | skipped cases counted toward `BLIND` | hand-audit of the `BLIND` list |
150
+ | config and CI YAML not treated as source | a live PR it was asked to check |
151
+ | commit-vs-parent is wrong for a PR branch | a reviewer's suggested command |
152
+ | `.env` absent from worktrees → 955 phantom skips | getting a real database running |
153
+ | a historical `BLIND` is not an open defect | shipping a redundant ticket to a colleague |
154
+ | test detection required BOTH a `test/` dir AND a `.test.` suffix | first run on a repo the author did not write |
155
+
156
+ That last one returned **an entirely empty audit** on undici — 284 commits matched, zero
157
+ judged — because undici names its tests `test/client-request.js`. Not a wrong answer: no
158
+ answer.
159
+
160
+ **It is the argument for running this on somebody else's code before believing any number
161
+ it prints.**
package/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Mekan Hanov
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
package/README.md ADDED
@@ -0,0 +1,164 @@
1
+ # control-arm (`ca`)
2
+
3
+ **Does a test actually fail on the code it was written to catch?**
4
+
5
+ When you fix a bug you usually add a test so it can't come back. Everyone does this.
6
+ **Nobody checks whether that test would actually have caught the bug.**
7
+
8
+ This checks. It takes the test you wrote, rewinds your code to just before the fix, and
9
+ runs the test there. Red means it works. Green means it was never capable of catching
10
+ anything — it just sits in your suite looking like protection.
11
+
12
+ It does not write tests, and there is no AI in it. It runs your tests and reports what
13
+ happened.
14
+
15
+ ## Two tests for the same bug
16
+
17
+ ```js
18
+ // A
19
+ assert.equal(priority('HIGH-PRIORITY'), 1);
20
+
21
+ // B
22
+ const p = priority('HIGH-PRIORITY');
23
+ assert.ok(p); // any number but 0 is truthy
24
+ assert.ok(SLA_HOURS[p] > 0); // 24 > 0 is true
25
+ ```
26
+
27
+ The bug: `priority()` matched `HIGH PRIORITY` with a space but not `HIGH-PRIORITY` with a
28
+ hyphen, so an urgent ticket came back as normal.
29
+
30
+ On the fixed code both pass. On the broken code **A fails and B still passes** — the wrong
31
+ answer, `2`, is also truthy and also has a positive SLA. In CI they are identical: two
32
+ green checkmarks.
33
+
34
+ **That is the gap.** Nothing in a normal pipeline can tell those two apart. `ca` can,
35
+ because it runs the new test against the old code.
36
+
37
+ ## Install
38
+
39
+ ```bash
40
+ npm i -g control-arm # or run it from a clone: node bin/ca.mjs
41
+ ```
42
+
43
+ Node 18+. No other dependencies.
44
+
45
+ ## Use it
46
+
47
+ ```bash
48
+ # Can this repo be measured at all?
49
+ ca probe
50
+
51
+ # Judge one commit — a fix, with the test that shipped alongside it
52
+ ca verify <sha>
53
+
54
+ # Judge a branch against where it will merge
55
+ ca verify HEAD --against origin/main
56
+ ```
57
+
58
+ ### As a PR check
59
+
60
+ ```yaml
61
+ # .github/workflows/control-arm.yml
62
+ name: control-arm
63
+ on: pull_request
64
+
65
+ jobs:
66
+ verify:
67
+ runs-on: ubuntu-latest
68
+ steps:
69
+ - uses: actions/checkout@v4
70
+ with: { fetch-depth: 0 } # the merge base needs real history
71
+ - uses: actions/setup-node@v4
72
+ with: { node-version: 22 }
73
+ - run: npm ci
74
+ - uses: mekanhan/control-arm@v1
75
+ with:
76
+ base: main
77
+ comment: 'true'
78
+ ```
79
+
80
+ It comments on the PR with a line per test. It does **not** fail the build by default: a
81
+ PR can legitimately ship only regression guards, and a gate that fires on those gets
82
+ switched off within a week.
83
+
84
+ ## What you get
85
+
86
+ > ### `control-arm` — this branch is proven
87
+ >
88
+ > **1 of 5 tests here genuinely catch this change.** Run against the code as it was before
89
+ > this PR, they fail; with the change, they pass.
90
+ >
91
+ > | | test | what it did on the code WITHOUT this change |
92
+ > |---|---|---|
93
+ > | ✅ | `RESET-002: does not render the next UTC midnight` | **failed** — `expect(received).toBeNull()` |
94
+ > | ⚪️ | `RESET-003: the pre-fix expression disagrees` | _passed too — so it is not what catches this bug_ |
95
+ > | ⚠️ | `RESET-001: renders the time the server sent` | _could not run there — crashed rather than disagreeing_ |
96
+
97
+ **One `CAUGHT` and a pile of `NON-DISCRIMINATING` is the healthy result**, not a problem.
98
+ One test was aimed at the bug; the rest guard other things. The only worrying outcome is a
99
+ commit where *nothing* caught it.
100
+
101
+ ## Verdicts
102
+
103
+ | level | verdict | meaning |
104
+ |---|---|---|
105
+ | case | `CAUGHT` | ran on the broken code and disagreed with it |
106
+ | case | `NON-DISCRIMINATING` | ran and was fine with it — a fact, not a fault |
107
+ | case | `INCONCLUSIVE` | never ran, or its identity could not be proven |
108
+ | case | `FLAKY` | repeated runs disagreed |
109
+ | commit | `CAUGHT` | at least one case discriminates |
110
+ | commit | `BLIND` | every case ran, **not one** discriminated |
111
+ | commit | `INCONCLUSIVE` | excluded from the ratio rather than assumed |
112
+
113
+ `INCONCLUSIVE` is reported in its own column and never folded into the others. A
114
+ measurement that cannot tell zero from failure is not a measurement.
115
+
116
+ ## Which of your tests it looks at
117
+
118
+ It does not care what *kind* of test it is, only whether the runner can execute the file.
119
+
120
+ | kind | covered | why |
121
+ |---|---|---|
122
+ | unit tests | yes | the easy case |
123
+ | backend / server logic | yes | same runner, same rewind |
124
+ | database-backed tests | yes | needs a live DB, else they report `SKIPPED` |
125
+ | component tests (React / RN) | yes | via vitest and jest |
126
+ | browser e2e (Playwright) | **no** | no Playwright runner yet |
127
+ | performance / load | **no** | they measure speed, not correctness |
128
+ | manual QA | **no** | nothing to execute |
129
+
130
+ ## Does it work?
131
+
132
+ Three repositories, random draws, seeds recorded so the samples are reproducible.
133
+
134
+ | | a private monorepo | nodejs/undici | iamkun/dayjs |
135
+ |---|---|---|---|
136
+ | fix commits sampled | 300 | 25 | 40 |
137
+ | answerable | 214 | 15 | 25 |
138
+ | **CAUGHT** | **93.9%** | **80.0%** | **92.0%** |
139
+ | BLIND | 6.1% | 20.0% | 8.0% |
140
+ | runtime | 7.1 s/commit | 24.9 s/commit | 1.4 s/commit |
141
+
142
+ **The dayjs column is the one to weigh.** It is a codebase the author did not write, did
143
+ not choose for a flattering result, and could not tune the tool against — 190 fix commits
144
+ matched the filter, 99 ship both a test and a source change, 40 drawn at random with the
145
+ seed recorded. It lands within a point of the private monorepo it was built on.
146
+
147
+ Of the 13 raw `BLIND` verdicts in that 300-commit run, **5 were genuinely open** — a 62%
148
+ false-positive rate on the raw number, which is why the filters exist.
149
+
150
+ **[DESIGN.md](DESIGN.md) has the rest**: how the rewind works, the three traps that
151
+ separate an instrument from a random number generator, the full accuracy audit, and every
152
+ bug found in this tool so far — none of which were found by its own test suite.
153
+
154
+ ## Scope, stated plainly
155
+
156
+ - **A test can only be judged against the bug it was written for.** No fix commit, no
157
+ broken code, no verdict. The measurable universe is fix commits that shipped a test. For
158
+ tests with no such pairing, the instrument is mutation testing, not this.
159
+ - One runner today (`node:test`, plus vitest and jest via adapters). The runner seam is
160
+ isolated in `src/runner.mjs`; the decision logic never sees raw runner output.
161
+ - `ca` never touches your working tree — no `stash`, no `checkout` in your clone.
162
+ Everything happens in detached worktrees under `.ca-work/`.
163
+
164
+ MIT.
package/action.yml ADDED
@@ -0,0 +1,95 @@
1
+ # A composite action, not a Docker one: the tool is plain Node with no dependencies, so a
2
+ # container would add a minute of build time to a check that otherwise takes seconds.
3
+ name: 'control-arm'
4
+ description: 'Prove the tests in this PR fail without it'
5
+ inputs:
6
+ base:
7
+ description: >-
8
+ Branch this PR targets. The comparison uses the MERGE BASE with it, never its tip.
9
+ Defaults to `main`; set it explicitly if your project integrates somewhere else
10
+ (`develop`, `trunk`, a release branch).
11
+ required: false
12
+ default: 'main'
13
+ comment:
14
+ description: 'Post the result as a PR comment. Needs pull-requests:write and GH_TOKEN.'
15
+ required: false
16
+ default: 'true'
17
+ fail-on-blind:
18
+ description: >-
19
+ Fail the check when NO case in the PR discriminates. Default false: a PR can
20
+ legitimately ship only regression guards, and a gate that fires on those gets
21
+ switched off within a week.
22
+ required: false
23
+ default: 'false'
24
+ timeout-ms:
25
+ required: false
26
+ default: '180000'
27
+ outputs:
28
+ verdict:
29
+ description: 'CAUGHT | BLIND | INCONCLUSIVE | FLAKY | SKIPPED'
30
+ value: ${{ steps.run.outputs.verdict }}
31
+ runs:
32
+ using: composite
33
+ steps:
34
+ - id: run
35
+ shell: bash
36
+ env:
37
+ CA_BASE: ${{ inputs.base }}
38
+ CA_TIMEOUT: ${{ inputs.timeout-ms }}
39
+ run: |
40
+ set -uo pipefail
41
+ # The merge base needs both histories. A shallow checkout has neither.
42
+ git -C "$GITHUB_WORKSPACE" fetch --no-tags --depth=200 origin "$CA_BASE" 2>/dev/null || true
43
+ TIP="$(git -C "$GITHUB_WORKSPACE" rev-parse HEAD)"
44
+
45
+ OUT=$(node "${{ github.action_path }}/bin/ca.mjs" verify "$TIP" \
46
+ --against "origin/$CA_BASE" \
47
+ --repo "$GITHUB_WORKSPACE" \
48
+ --timeout "$CA_TIMEOUT" 2>&1) || true
49
+ echo "$OUT"
50
+
51
+ # `|| true` is load-bearing. GitHub runs bash steps with `set -e` by default, and
52
+ # `set -uo pipefail` above does not unset it. A grep that finds nothing exits 1,
53
+ # pipefail propagates it, and the step dies BEFORE the :-INCONCLUSIVE fallback can
54
+ # run. That is exactly what happens on a PR the tool declines — a workflow-only
55
+ # change with no test — so the one case designed to be a non-event failed the check.
56
+ VERDICT=$(printf '%s' "$OUT" | grep -oE 'VERDICT +[A-Z]+' | awk '{print $2}' | head -1 || true)
57
+ VERDICT="${VERDICT:-INCONCLUSIVE}"
58
+ echo "verdict=$VERDICT" >> "$GITHUB_OUTPUT"
59
+
60
+ # `--pr-comment` returns EMPTY for a commit the tool declined to judge, and the
61
+ # comment step skips on an empty file. The decision is the tool's, not a shell's:
62
+ # scraping stdout for a VERDICT line a declined run never prints is what put
63
+ # "SKIPPED · no case fails without the change" onto a PR that simply has no tests.
64
+ MD=$(node "${{ github.action_path }}/bin/ca.mjs" verify "$TIP" \
65
+ --against "origin/$CA_BASE" --pr-comment \
66
+ --repo "$GITHUB_WORKSPACE" --timeout "$CA_TIMEOUT" 2>/dev/null) || true
67
+ printf '%s' "$MD" > "$RUNNER_TEMP/ca-comment.md"
68
+ [ -s "$RUNNER_TEMP/ca-comment.md" ] || echo "declined — nothing to post: $(printf '%s' "$OUT" | tail -1)"
69
+
70
+ - shell: bash
71
+ if: inputs.comment == 'true' && github.event_name == 'pull_request'
72
+ env:
73
+ GITHUB_TOKEN: ${{ github.token }}
74
+ run: |
75
+ # node, not `gh`. The CLI is not installed on every runner — a self-hosted
76
+ # bare-metal box failed here with `gh: command not found` after every other step
77
+ # had passed, so the tool ran, reached the right verdict, and could not say so.
78
+ node "${{ github.action_path }}/scripts/gh-api.mjs" upsert-pr-comment \
79
+ "${{ github.repository }}" \
80
+ "${{ github.event.pull_request.number }}" \
81
+ "$RUNNER_TEMP/ca-comment.md" \
82
+ '### `control-arm`'
83
+
84
+ - shell: bash
85
+ if: inputs.fail-on-blind == 'true' && steps.run.outputs.verdict == 'BLIND'
86
+ run: |
87
+ echo "::error::No test in this PR fails without the change."
88
+ exit 1
89
+
90
+ # Required by the GitHub Marketplace listing. `rewind` is not decoration: it is
91
+ # literally what this does — wind the code back to before the fix, then run the
92
+ # new test against it.
93
+ branding:
94
+ icon: 'rewind'
95
+ color: 'orange'