mjolnir-qa 1.0.8 → 2.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,25 +1,27 @@
1
1
  <div align="center">
2
2
 
3
- <img src="assets/readme/logo.png" alt="Mjölnir — Verification Trust Engine" width="800" />
3
+ <img src="assets/readme/hero.svg" alt="Mjölnir. Tests tell you what passed. Mjölnir tells you what you can trust." width="100%" />
4
4
 
5
- ### Tests tell you what passed. Mjölnir tells you what you can trust.
5
+ <br />
6
6
 
7
- **Mjölnir is a Verification Trust Engine.** Test frameworks verify your
8
- software. Mjölnir verifies the system that does the verifying — the test
9
- suite, the run artifacts and the CI pipeline — and reports a worthiness
10
- score with the evidence behind every deduction.
7
+ Mjölnir finds tests that cannot fail and pipelines that cannot go red,<br />
8
+ then scores how far the result can be trusted, with the evidence for every point.
11
9
 
12
- [![npm](https://img.shields.io/npm/v/mjolnir-qa.svg?style=flat-square&color=C19A34&labelColor=0A1119)](https://www.npmjs.com/package/mjolnir-qa)
13
- [![downloads](https://img.shields.io/npm/dm/mjolnir-qa.svg?style=flat-square&color=C19A34&labelColor=0A1119)](https://www.npmjs.com/package/mjolnir-qa)
10
+ <br />
11
+
12
+ [![npm](https://img.shields.io/npm/v/mjolnir-qa.svg?style=flat-square&color=1F6F7C&labelColor=0A1119)](https://www.npmjs.com/package/mjolnir-qa)
13
+ [![downloads](https://img.shields.io/npm/dm/mjolnir-qa.svg?style=flat-square&color=1F6F7C&labelColor=0A1119)](https://www.npmjs.com/package/mjolnir-qa)
14
14
  [![ci](https://img.shields.io/github/actions/workflow/status/Sergey-Bar/Mjolnir/ci.yml?branch=main&style=flat-square&label=ci&labelColor=0A1119)](https://github.com/Sergey-Bar/Mjolnir/actions/workflows/ci.yml)
15
- [![license](https://img.shields.io/badge/license-MIT-C19A34.svg?style=flat-square&labelColor=0A1119)](LICENSE)
16
- [![node](https://img.shields.io/badge/node-%E2%89%A5%2022.18-37ABBD.svg?style=flat-square&labelColor=0A1119)](https://nodejs.org)
15
+ [![coverage](https://img.shields.io/codecov/c/github/Sergey-Bar/Mjolnir?style=flat-square&color=1F6F7C&labelColor=0A1119&label=coverage)](https://codecov.io/gh/Sergey-Bar/Mjolnir)
16
+ [![OpenSSF Scorecard](https://api.scorecard.dev/projects/github.com/Sergey-Bar/Mjolnir/badge)](https://scorecard.dev/viewer/?uri=github.com/Sergey-Bar/Mjolnir)
17
+ [![license](https://img.shields.io/badge/license-MIT-1F6F7C.svg?style=flat-square&labelColor=0A1119)](LICENSE)
18
+ [![node](https://img.shields.io/badge/node-%E2%89%A5%2022.18-1F6F7C.svg?style=flat-square&labelColor=0A1119)](https://nodejs.org)
17
19
 
18
20
  ```bash
19
21
  npx mjolnir-qa@latest
20
22
  ```
21
23
 
22
- [See it work](#see-it-work) · [Quickstart](#quickstart) · [What it finds](#what-mjölnir-finds) · [Score](#the-worthiness-score) · [Evidence model](#the-evidence-model) · [Forensics](#runtime-forensics) · [CI](#ci-integrity) · [Agents](#ai-agents) · [Security](#trust-and-security) · [Limits](#what-mjölnir-cannot-tell-you) · [Docs](#documentation)
24
+ [See it work](#see-it-work) · [Quickstart](#quickstart) · [What it finds](#what-mjölnir-finds) · [Score](#the-worthiness-score) · [Evidence](#the-evidence-model) · [Forensics](#runtime-forensics) · [CI](#ci-integrity) · [Agents](#ai-agents) · [Security](#trust-and-security) · [Limits](#what-mjölnir-cannot-tell-you) · [Docs](#documentation)
23
25
 
24
26
  <details>
25
27
  <summary>Read this in another language — 22 translations</summary>
@@ -33,110 +35,107 @@ it; `npm run docs:translations` reports how far.
33
35
 
34
36
  </div>
35
37
 
36
- ---
38
+ <br />
37
39
 
38
- ## The problem
40
+ ## A green check is a claim, not a proof
39
41
 
40
- A green pipeline is a claim, not a proof. The same checkmark is printed
41
- whether a suite genuinely verified your product or merely failed to
42
- contradict it. Every one of these ships green:
42
+ A green check means the pipeline did not fail. It does not mean the tests
43
+ ran, or that they could have failed. Every one of these ships green:
43
44
 
44
45
  - a committed `.only` that ran 3 tests instead of 900
45
- - a `continue-on-error: true` on the job that was supposed to gate
46
- - a `|| true` after the test command
47
- - a test that asserts nothing, or whose body is empty
46
+ - `continue-on-error: true` on the job that was supposed to gate
47
+ - `|| true` after the test command
48
+ - a test that asserts nothing, or has an empty body
48
49
  - a retry wrapper that turns a real failure into a lucky pass
49
- - a report the workflow uploads but never actually generated
50
- - a hard sleep holding a race condition together until the day it doesn't
51
-
52
- None of these are exotic, and none of them turn the pipeline red. They look
53
- intentional to a reviewer — which is exactly why they survive.
54
-
55
- ## The Mjölnir Principle
50
+ - a report the workflow uploads but never generated
51
+ - a hard sleep holding a race condition together
56
52
 
57
- > ### No evidence. No proof.
53
+ None of them turns the pipeline red, and each one looks deliberate in
54
+ review. That is why they survive. Here is Mjölnir reading a real one:
58
55
 
59
- Mjölnir would rather say _unknown_ than manufacture confidence. Where a
60
- conventional tool rounds silence up to "fine", it stops and names the gap:
56
+ <p align="center">
57
+ <img src="assets/readme/scan.svg" alt="The demo repository's CI workflow, read line by line. Mjölnir flags each finding at the line it reported, with its rule, what is wrong, its evidence level and its measured false-positive rate." width="800" />
58
+ </p>
61
59
 
62
- | Situation | What Mjölnir reports |
63
- | ------------------------------------ | ------------------------------------------------------- |
64
- | No test declarations found | Score `null` — **UNKNOWN**, never a fabricated 100 |
65
- | No baseline / no comparable revision | **UNKNOWN**, with the reason named — never an assumed 0 |
66
- | Scan truncated (budget, unreadable) | **PARTIAL**, exit `2` — never presented as clean |
60
+ <sub>Every finding the demo scan reported for this workflow, at the line it
61
+ reported. Generated by `npm run docs:readme-brand` from
62
+ [`demo-report.json`](assets/readme/demo-report.json) and drift-locked in CI.</sub>
67
63
 
68
- Unknown is a valid answer, and this is the reason: a tool that says
69
- "verified" when it does not know is the same failure mode as a CI gate
70
- that says green when it never ran.
64
+ **Strict mode.** The most aggressive detections — `.only`, `continue-on-error`,
65
+ empty tests, retry abuse — live in the quarantine tier. They run only
66
+ under `--strict` and are capped to `info` severity: they flag, they never
67
+ gate. The default scan (`npx mjolnir-qa@latest` without `--strict`)
68
+ covers core and extended rules only. Add `--strict` when you want the
69
+ advisory layer too.
71
70
 
72
- ## How it works
71
+ Mjölnir reads the suite, the CI workflows and, when you have one, the
72
+ report of a real run. It does not run your tests, install your
73
+ dependencies, or execute the code it scans. And when it has no evidence,
74
+ it says so rather than inventing confidence:
73
75
 
74
- Mjölnir sits between your test system and your release decision. It reads
75
- the suite, the CI workflows and — when you point it at one — the artifacts
76
- of a real run.
76
+ | Situation | What Mjölnir reports |
77
+ | ---------------------------------------------- | -------------------------------------------------------- |
78
+ | No test declarations found | Score `null`, shown as **UNKNOWN**. Never a made-up 100. |
79
+ | No baseline or comparable revision | **UNKNOWN**, with the reason named. Never an assumed 0. |
80
+ | Scan cut short (time budget, unreadable files) | **PARTIAL**, exit `2`. Never presented as clean. |
77
81
 
78
82
  <p align="center">
79
- <img src="assets/readme/architecture.svg" alt="Mjölnir reads the test suite and the CI pipeline statically, and reads the Playwright JSON and JUnit XML artifacts of a real run. It discovers, analyzes, correlates and measures across four evidence streams — test quality, CI integrity, runtime forensics and selector health — stamping each finding E0 observation, E1 pattern evidence or E2 deterministic proof, weighted none, half and full. A trust ladder L0 to L5 shows the top three rungs require a real run. Out come findings, a worthiness score of 75 out of 100 labelled NEEDS WORK, and a CI gate on the frozen exit codes 0, 1, 2, 10 and 20. An agent loop runs scan, evidence, handoff, AI agent, re-scan, proof." width="1600" />
83
+ <img src="assets/readme/how-it-works.svg" alt="How Mjölnir works. It reads the test suite and the CI pipeline statically, and the report of a real run when there is one. It weighs every finding by its evidence level and its trust level, where only a real run can reach L3 to L5, and produces findings, a worthiness score and a CI gate on frozen exit codes. In the agent loop, AI writes the fix and Mjölnir re-scans to prove it." width="880" />
80
84
  </p>
81
85
 
82
- It does not run your tests, install your dependencies, or execute the code
83
- it scans. Static analysis reads source text; forensics reads report files
84
- that already exist on disk.
86
+ <sub>Composed for this page and shown at 1:1. Generated by
87
+ `npm run docs:readme-brand` and drift-locked in CI; the score, counts and
88
+ rule ID come from [`script.demo.json`](assets/video/script.demo.json),
89
+ [`demo-report.json`](assets/readme/demo-report.json) and the rule
90
+ registry, never typed by hand. The same picture as a poster:
91
+ [`architecture.svg`](assets/readme/architecture.svg).</sub>
85
92
 
86
- <sub>Generated by `npm run docs:architecture` and drift-locked in CI; the
87
- score, counts and rule ID are read from
88
- [`script.demo.json`](assets/video/script.demo.json), not written by hand.
89
- Open [`architecture.svg`](assets/readme/architecture.svg) on its own for
90
- the full-resolution version.</sub>
91
-
92
- ---
93
+ <br />
93
94
 
94
95
  ## See it work
95
96
 
96
- One false-green CI gate — caught, fixed with the tool's own printed fix,
97
- and re-proved by a second scan.
97
+ A real scan of [`examples/demo-repo`](examples/demo-repo), a small
98
+ Playwright suite with a CI workflow. This is where its points went:
98
99
 
99
100
  <p align="center">
100
- <img src="assets/readme/flow.svg" alt="npx mjolnir-qa@latest. A large grey 75 labelled NEEDS WORK, an arrow marked ONE FIX above and RE-SCANNED below, then a large lit 90 labelled WORTHY. Beneath: set -o pipefail, &amp;&amp; not a semicolon, no continue-on-error — the fix the report printed, closing QA-CI-009 and QA-CI-001. Then 27 findings to 23, 4 errors to 1. Finally: 90, not 100 — the suite's other problems are still real." width="900" />
101
+ <img src="assets/readme/terminal-hero.svg" alt="Mjölnir's deduction breakdown: WORTHINESS 75/100 NEEDS WORK, the score by category, the per-severity deduction box, and a FIX THIS FIRST list" width="520" />
101
102
  </p>
102
103
 
104
+ <sub>Generated by `npm run docs:hero` from a real scan and drift-locked in
105
+ CI. The full `--verbose` report of the same scan is
106
+ [`demo.svg`](assets/readme/demo.svg) (`npm run docs:demo`).</sub>
107
+
103
108
  <details>
104
- <summary><strong>Prefer to watch it?</strong> The same run, as a 42-second recording</summary>
109
+ <summary><strong>Watch it</strong> — a scan, the fix it prints, and the re-scan that proves it</summary>
110
+
111
+ <br />
105
112
 
106
- <!-- Plays inline on github.com only: <video> is rendered for GitHub's own
107
- user-content CDN, never for a repo-relative path. The link below is
108
- the fallback for every other renderer (npm, mirrors, offline clones). -->
109
113
  <p align="center">
110
- <video
111
- src="https://github.com/user-attachments/assets/0e1af1e4-1e27-4c1c-9ec4-2717d194df05"
112
- poster="https://raw.githubusercontent.com/Sergey-Bar/Mjolnir/main/assets/video/mjolnir-demo-poster.png"
113
- controls
114
- muted
115
- playsinline
116
- width="900"></video>
114
+ <a href="assets/video/mjolnir-demo.mp4">
115
+ <img src="assets/video/mjolnir-demo-poster.png" alt="A frame of the demo recording: npx mjolnir-qa@latest scanning the demo repository in a terminal window" width="900" />
116
+ </a>
117
117
  </p>
118
118
 
119
- <sub>Found, fixed, re-proved, then handed to an agent. If the player above
120
- doesn't load, the file is
121
- [`assets/video/mjolnir-demo.mp4`](assets/video/mjolnir-demo.mp4). Rendered
122
- by `npm run docs:video`. The full `--verbose` report of the same scan is
123
- [`demo.svg`](assets/readme/demo.svg) (`npm run docs:demo`).</sub>
119
+ <sub>Rendered frame by frame from a real scan by `npm run docs:video`;
120
+ never screen-recorded. Select the frame to open
121
+ [`mjolnir-demo.mp4`](assets/video/mjolnir-demo.mp4).</sub>
124
122
 
125
123
  </details>
126
124
 
127
- <sub>Every number above is read from
128
- [`script.demo.json`](assets/video/script.demo.json) — the same values
129
- [`video-script.spec.ts`](tests/contract/video-script.spec.ts) checks
130
- against real CLI output — and the diff quotes the two committed workflows
131
- verbatim. Regenerate with `npm run docs:flow`; drift-locked in CI.</sub>
132
-
133
125
  ### One finding, up close
134
126
 
135
- `mjolnir explain QA-CI-001` prints a rule's whole trust record — including
127
+ Every finding answers four questions: where it is, how sure Mjölnir is,
128
+ how often the rule is wrong, and how to fix it.
129
+
130
+ <p align="center">
131
+ <img src="assets/readme/finding-anatomy.svg" alt="The first finding of the demo scan, exactly as the terminal prints it, with its four parts marked: where, how sure, how often the rule is wrong, and the fix." width="100%" />
132
+ </p>
133
+
134
+ `mjolnir explain QA-CI-001` prints a rule's whole trust record, including
136
135
  its measured false-positive rate and the tier that rate earned it:
137
136
 
138
137
  ```text
139
- ▚ QA-CI-001 — continue-on-error masks a failing verification gate
138
+ ▍ QA-CI-001 — continue-on-error masks a failing verification gate
140
139
 
141
140
  Severity: error
142
141
  Confidence: high
@@ -179,10 +178,9 @@ HOW TO VERIFY THE FIX
179
178
  Docs: mjolnir rules --md (full catalog, this rule included)
180
179
  ```
181
180
 
182
- That is the unit of value: not a style nit, but a place where CI is
183
- reporting a pass it did not earn.
181
+ That is the unit of value: one place where CI reports a pass it did not earn.
184
182
 
185
- ---
183
+ <br />
186
184
 
187
185
  ## Quickstart
188
186
 
@@ -190,84 +188,89 @@ reporting a pass it did not earn.
190
188
  npx mjolnir-qa@latest
191
189
  ```
192
190
 
193
- That is the whole product: it scans the current directory and prints the
194
- Trust Report — what the scan found, how much you can trust it, why, and
195
- what to do next — then exits `0` if nothing at or above the gate was
196
- found. **In CI, use the changed-scope form** — it attributes findings to
197
- what your branch introduced, so a legacy suite does not drown a first PR:
191
+ It scans the current directory and prints the Trust Report: what it found,
192
+ how far you can trust it, why, and what to do next. It exits `0` when
193
+ nothing at or above the gate was found.
194
+
195
+ In CI, scan only what the branch introduced, so a legacy suite does not
196
+ drown your first pull request:
198
197
 
199
198
  ```bash
200
199
  npx mjolnir-qa@latest --scope changed
201
200
  ```
202
201
 
203
- `mjolnir ci install` writes that as a GitHub Actions workflow — the
204
- [action](https://github.com/Sergey-Bar/Mjolnir#readme) (Marketplace-grade,
205
- pinned to the `v1` major tag) by default, or plain `npx` with
206
- `--no-action`. Advisory by default, never blocking until you say so.
202
+ `mjolnir ci install` writes that as a GitHub Actions workflow, using the
203
+ [action](https://github.com/Sergey-Bar/Mjolnir#readme) pinned to the `v1`
204
+ major tag (or plain `npx` with `--no-action`). It stays advisory until you
205
+ decide it should block.
207
206
 
208
207
  | Command | What it does |
209
208
  | ----------------------------------- | ------------------------------------------------ |
210
- | `mjolnir` | Trust Report — verdict, confidence, next action |
211
- | `mjolnir --scope changed` | Only what your branch introduced — the CI form |
209
+ | `mjolnir` | Trust Report: verdict, confidence, next action |
210
+ | `mjolnir --scope changed` | Only what your branch introduced (the CI form) |
212
211
  | `mjolnir ci install` | Generate the advisory PR workflow (action-based) |
213
- | `mjolnir explain QA-CI-001` | What / why / fix + measured FP rate for one rule |
214
- | `mjolnir why src/a.spec.ts:42` | Why this exact line was flagged — never a gate |
212
+ | `mjolnir explain QA-CI-001` | What, why and fix, plus the measured FP rate |
213
+ | `mjolnir why src/a.spec.ts:42` | Why this exact line was flagged. Never gates. |
215
214
  | `mjolnir forensics ./test-results/` | Runtime evidence from a real run |
216
215
  | `mjolnir trust-report` | Self-contained Trust Artifact (md + json) |
217
216
  | `mjolnir handoff` | Remediation plan for a coding agent |
218
- | `mjolnir --json` / `--format sarif` | Machine-readable / GitHub Code Scanning |
217
+ | `mjolnir --json` / `--format sarif` | Machine-readable output, GitHub Code Scanning |
219
218
  | `mjolnir --format codequality` | GitLab Code Quality report (MR widget artifact) |
220
219
  | `mjolnir --strict` | Also run quarantine-tier rules (higher FP risk) |
221
220
 
222
221
  <details>
223
- <summary><strong>Everything else</strong> — flake triage, reporting, governance</summary>
224
-
225
- | Command | What it does |
226
- | ----------------------------------- | -------------------------------------------------------- |
227
- | `mjolnir --classic` | The pre-Trust-Report score banner render |
228
- | `mjolnir explain verdict` | Why the saved scan's verdict is what it is |
229
- | `mjolnir triage ./test-results/` | Guided triage workflow — every row ends in a next action |
230
- | `mjolnir pw-report ./test-results/` | Playwright run summary — retries / flakes / slowest |
231
- | `mjolnir doctor:playwright` | Playwright-only deep scan + Selector Health Score |
232
- | `mjolnir fix --dry-run` / `fix` | Safe auto-fixes, each re-scanned to prove it landed |
233
- | `mjolnir baseline` / `diff` | Snapshot findings, then report only new/worsened |
234
- | `mjolnir impact --since <ref>` | What a commit introduced vs resolved |
235
- | `mjolnir summary` | CI annotations + step summary from a saved report |
236
- | `mjolnir pr-comment` | A scoped PR comment, as Markdown |
237
- | `mjolnir debt` | Test-debt register with a cost model |
238
- | `mjolnir handover` | New-QA onboarding map of the suite |
239
- | `mjolnir init` | Detect frameworks + setup checklist (never overwrites) |
240
- | `mjolnir suppressions` | List suppressed findings — governance transparency |
241
- | `mjolnir rules --unmeasured` | The rules running on assumption, not measurement |
242
- | `mjolnir rules --md` | Full rule catalog (JSON or Markdown) |
243
- | `mjolnir doctor` | Self-audit of Mjölnir's own rule base |
244
- | `mjolnir create-rule <ID>` | Scaffold a new rule + fixtures |
245
- | `mjolnir stats` | Local all-time counters of fixes seen |
246
- | `mjolnir badge` | shields.io endpoint JSON + snippet |
247
- | `mjolnir --cache` | Incremental re-scans via a local verdict cache |
248
- | `mjolnir --format mermaid` | Test-architecture diagram for a PR comment |
222
+ <summary><strong>Every other command</strong> — flake triage, reporting, governance</summary>
223
+
224
+ <br />
225
+
226
+ | Command | What it does |
227
+ | ----------------------------------- | ---------------------------------------------------- |
228
+ | `mjolnir --classic` | The pre-Trust-Report score banner render |
229
+ | `mjolnir explain verdict` | Why the saved scan's verdict is what it is |
230
+ | `mjolnir triage ./test-results/` | Guided triage. Every row ends in a next action. |
231
+ | `mjolnir pw-report ./test-results/` | Playwright run summary: retries, flakes, slowest |
232
+ | `mjolnir doctor:playwright` | Playwright-only deep scan plus Selector Health Score |
233
+ | `mjolnir fix --dry-run` / `fix` | Safe auto-fixes, each re-scanned to prove it landed |
234
+ | `mjolnir baseline` / `diff` | Snapshot findings, then report only new or worse |
235
+ | `mjolnir impact --since <ref>` | What a commit introduced and resolved |
236
+ | `mjolnir summary` | CI annotations and a step summary from a report |
237
+ | `mjolnir pr-comment` | A scoped PR comment, as Markdown |
238
+ | `mjolnir debt` | Test-debt register with a cost model |
239
+ | `mjolnir handover` | Onboarding map of the suite for a new QA engineer |
240
+ | `mjolnir init` | Detect frameworks, print a setup checklist |
241
+ | `mjolnir suppressions` | List suppressed findings, for governance |
242
+ | `mjolnir rules --unmeasured` | The rules running on assumption, not measurement |
243
+ | `mjolnir rules --md` | Full rule catalog (JSON or Markdown) |
244
+ | `mjolnir doctor` | Self-audit of Mjölnir's own rule base |
245
+ | `mjolnir create-rule <ID>` | Scaffold a new rule and its fixtures |
246
+ | `mjolnir stats` | Local all-time counters of fixes seen |
247
+ | `mjolnir badge` | shields.io endpoint JSON and snippet |
248
+ | `mjolnir --cache` | Incremental re-scans via a local verdict cache |
249
+ | `mjolnir --format mermaid` | Test-architecture diagram for a PR comment |
249
250
 
250
251
  `mjolnir help <command>` prints usage, examples and the next step for any
251
252
  of them.
252
253
 
253
254
  </details>
254
255
 
255
- Requires **Node.js ≥ 22.18**. Runs on Windows, macOS and Linux. Install
256
- globally with `npm i -g mjolnir-qa` if you prefer it over `npx`.
257
- (Why ≥ 22.18? The build toolchain sets the floor — tsdown targets it and
258
- the release pipeline smoke-tests against it; the runtime dependencies
259
- have no such requirement.)
256
+ Requires **Node.js ≥ 22.18** on Windows, macOS or Linux. Prefer a global
257
+ install? `npm i -g mjolnir-qa`. The floor comes from the build toolchain
258
+ (tsdown targets it and the release pipeline smoke-tests against it); the
259
+ runtime dependencies need no more than that.
260
260
 
261
- ---
261
+ <br />
262
262
 
263
263
  ## What Mjölnir finds
264
264
 
265
- **<!-- census:total-rules -->79 rules<!-- /census:total-rules -->** in four families — **test hygiene**, **test quality**,
266
- **Playwright**, **CI integrity** — over TypeScript/JavaScript, Python,
267
- Java, C# and GitHub Actions YAML, covering Playwright in all four bindings
268
- plus pytest, JUnit, TestNG, NUnit, xUnit, MSTest, Jest, Vitest and Mocha,
269
- with starter coverage for Cypress and Selenium. Ten of them, so the shape
270
- is clear:
265
+ <p align="center">
266
+ <img src="assets/readme/stack.svg" alt="Works with your stack: the languages, test frameworks and CI systems its rules cover, from the rule registry." width="100%" />
267
+ </p>
268
+
269
+ **<!-- census:total-rules -->79 rules<!-- /census:total-rules -->** in four families — test hygiene, test quality, Playwright and
270
+ CI integrity — across TypeScript and JavaScript, Python, Java, C# and
271
+ GitHub Actions YAML. They cover Playwright in all four bindings, plus
272
+ pytest, JUnit, TestNG, NUnit, xUnit, MSTest, Jest, Vitest and Mocha, with
273
+ starter coverage for Cypress and Selenium. Nine of them, to show the shape:
271
274
 
272
275
  | ID | Rule | Severity | Tier |
273
276
  | ------------ | ----------------------------------------------------------------- | -------- | ---------- |
@@ -281,242 +284,240 @@ is clear:
281
284
  | QA-PY-002 | Skipped test (`skip`, non-strict `xfail`) | warning | core |
282
285
  | QA-CS-103 | Test method with no assertions | error | core |
283
286
 
284
- The full catalog is generated from the registry, never hand-maintained:
287
+ The full catalog is generated from the registry, never maintained by hand:
285
288
  `mjolnir rules --md`, [`docs/rules/`](docs/rules/), or the
286
289
  [what-it-checks guide](https://sergey-bar.github.io/Mjolnir/guide/what-it-checks).
287
290
 
288
291
  <details>
289
- <summary><strong>Every rule named in this README, in one table</strong> — the rest live in <code>mjolnir rules --md</code></summary>
290
-
291
- > `quarantine` rules run only under `--strict` and never gate (capped to
292
- > info); the severity shown is the authored severity.
293
-
294
- | ID | Family | Rule | Severity | Tier |
295
- | ------------ | ---------- | ------------------------------------------------------------------- | ----------------------------- | ------------------------- |
296
- | QA-TEST-001 | Hygiene | Focused test committed (`.only`, `fit`) | error | quarantine |
297
- | QA-TEST-002 | Hygiene | Skipped test — escalates to `error` without a tracked justification | warning | quarantine |
298
- | QA-TEST-003 | Hygiene | Test with no assertions | error | quarantine |
299
- | QA-TEST-004 | Hygiene | Hard sleep (`waitForTimeout`, `sleep()`, `delay()`) | warning | extended |
300
- | QA-TEST-006 | Hygiene | Retry abuse hiding flakiness | warning | quarantine |
301
- | QA-TEST-010 | Hygiene | Empty test body | error | quarantine |
302
- | QA-TQUAL-002 | Quality | Tautological assertion | error | quarantine |
303
- | QA-TQUAL-009 | Quality | Unawaited promise assertion | error | quarantine |
304
- | QA-TQUAL-011 | Quality | Commented-out tests | warning | extended |
305
- | QA-PW-002 | Playwright | Unawaited locator assertion | error | core |
306
- | QA-PW-003 | Playwright | `page.pause()` / `test.only()` committed | error | core |
307
- | QA-PW-004 | Playwright | Brittle CSS/XPath selectors | warning | quarantine |
308
- | QA-PW-123 | Playwright | Hardcoded environment URLs | warning | quarantine |
309
- | QA-PW-140 | Playwright | Screenshot without `maxDiffPixelRatio` | warning | core |
310
- | QA-CI-001 | CI | `continue-on-error` masks a failing gate | error | quarantine |
311
- | QA-CI-002 | CI | ` | | true` swallows exit codes | error | extended |
312
- | QA-CI-005 | CI | Report consumed but never generated | error | quarantine |
313
- | QA-CI-007 | CI | Retry wrappers around tests | warning | extended |
314
- | QA-CI-008 | CI | Always-success step masks failures | error | quarantine |
315
- | QA-CI-009 | CI | Exit code not propagated (` | `without pipefail,`;` chains) | error | extended |
316
- | QA-CI-010 | CI | Tests skipped where they must block | error | quarantine |
317
- | QA-PY-002 | Python | Skipped test (`skip`, non-strict `xfail`) | warning | core |
318
- | QA-PY-003 | Python | Test function with no assertions | error | quarantine |
319
- | QA-PY-005 | Python | `time.sleep()` in tests | warning | extended |
320
- | QA-PY-012 | Python | Tautological assertion | error | quarantine |
321
- | QA-JV-101 | Java | Disabled test (`@Disabled`) | warning | core |
322
- | QA-JV-102 | Java | Hard sleep (`Thread.sleep()`) | warning | extended |
323
- | QA-JV-103 | Java | Test method with no assertions | error | extended |
324
- | QA-JV-105 | Java | Playwright `waitForTimeout()` hard sleep | warning | core |
325
- | QA-JV-106 | Java | Brittle selector instead of role locator | warning | quarantine |
326
- | QA-CS-101 | C# | Skipped test (`[Ignore]`, `[Fact(Skip=)]`) | warning | core |
327
- | QA-CS-102 | C# | Hard sleep (`Thread.Sleep` / `Task.Delay`) | warning | core |
328
- | QA-CS-103 | C# | Test method with no assertions | error | core |
329
- | QA-CS-105 | C# | `WaitForTimeoutAsync()` hard sleep | warning | extended |
330
- | QA-CS-106 | C# | Brittle selector instead of role locator | warning | quarantine |
292
+ <summary><strong>Every rule named in this README</strong>, in one table</summary>
293
+
294
+ <br />
295
+
296
+ > `quarantine` rules run only under `--strict` and never gate (they are
297
+ > capped to info). The severity shown is the authored severity.
298
+
299
+ | ID | Family | Rule | Severity | Tier |
300
+ | ------------ | ---------- | ------------------------------------------------------------ | -------- | ---------- |
301
+ | QA-TEST-001 | Hygiene | Focused test committed (`.only`, `fit`) | error | quarantine |
302
+ | QA-TEST-002 | Hygiene | Skipped test. Escalates to `error` without a tracked reason. | warning | quarantine |
303
+ | QA-TEST-003 | Hygiene | Test with no assertions | error | quarantine |
304
+ | QA-TEST-004 | Hygiene | Hard sleep (`waitForTimeout`, `sleep()`, `delay()`) | warning | extended |
305
+ | QA-TEST-006 | Hygiene | Retry abuse hiding flakiness | warning | quarantine |
306
+ | QA-TEST-010 | Hygiene | Empty test body | error | quarantine |
307
+ | QA-TQUAL-002 | Quality | Tautological assertion | error | quarantine |
308
+ | QA-TQUAL-009 | Quality | Unawaited promise assertion | error | quarantine |
309
+ | QA-TQUAL-011 | Quality | Commented-out tests | warning | extended |
310
+ | QA-PW-002 | Playwright | Unawaited locator assertion | error | core |
311
+ | QA-PW-003 | Playwright | `page.pause()` / `test.only()` committed | error | core |
312
+ | QA-PW-004 | Playwright | Brittle CSS/XPath selectors | warning | quarantine |
313
+ | QA-PW-123 | Playwright | Hardcoded environment URLs | warning | quarantine |
314
+ | QA-PW-140 | Playwright | Screenshot without `maxDiffPixelRatio` | warning | core |
315
+ | QA-CI-001 | CI | `continue-on-error` masks a failing gate | error | quarantine |
316
+ | QA-CI-002 | CI | `\|\| true` swallows exit codes | error | extended |
317
+ | QA-CI-005 | CI | Report consumed but never generated | error | quarantine |
318
+ | QA-CI-007 | CI | Retry wrappers around tests | warning | extended |
319
+ | QA-CI-008 | CI | Always-success step masks failures | error | quarantine |
320
+ | QA-CI-009 | CI | Exit code not propagated (`\|` without pipefail, `;` chains) | error | extended |
321
+ | QA-CI-010 | CI | Tests skipped where they must block | error | quarantine |
322
+ | QA-PY-002 | Python | Skipped test (`skip`, non-strict `xfail`) | warning | core |
323
+ | QA-PY-003 | Python | Test function with no assertions | error | quarantine |
324
+ | QA-PY-005 | Python | `time.sleep()` in tests | warning | extended |
325
+ | QA-PY-012 | Python | Tautological assertion | error | quarantine |
326
+ | QA-JV-101 | Java | Disabled test (`@Disabled`) | warning | core |
327
+ | QA-JV-102 | Java | Hard sleep (`Thread.sleep()`) | warning | extended |
328
+ | QA-JV-103 | Java | Test method with no assertions | error | extended |
329
+ | QA-JV-105 | Java | Playwright `waitForTimeout()` hard sleep | warning | core |
330
+ | QA-JV-106 | Java | Brittle selector instead of role locator | warning | quarantine |
331
+ | QA-CS-101 | C# | Skipped test (`[Ignore]`, `[Fact(Skip=)]`) | warning | core |
332
+ | QA-CS-102 | C# | Hard sleep (`Thread.Sleep` / `Task.Delay`) | warning | core |
333
+ | QA-CS-103 | C# | Test method with no assertions | error | core |
334
+ | QA-CS-105 | C# | `WaitForTimeoutAsync()` hard sleep | warning | extended |
335
+ | QA-CS-106 | C# | Brittle selector instead of role locator | warning | quarantine |
331
336
 
332
337
  Python also ships QA-PY-001…012 (pytest hygiene) and QA-PY-101…108
333
- (Playwright-Python); Cypress and Selenium have starter sets of three
338
+ (Playwright for Python). Cypress and Selenium have starter sets of three
334
339
  rules each.
335
340
 
336
341
  </details>
337
342
 
338
- Every rule ships with a must-fire **and** a must-not-fire fixture; a rule
339
- that fires on its own negative fixture cannot ship. That is the
340
- false-positive firewall, and `mjolnir doctor` enforces it in this
341
- repository's own CI.
343
+ Every rule ships with a must-fire **and** a must-not-fire fixture, and a
344
+ rule that fires on its own negative fixture cannot ship. That is the
345
+ false-positive firewall; `mjolnir doctor` enforces it in this repository's
346
+ own CI.
342
347
 
343
348
  ### Selector Health Score
344
349
 
345
350
  `mjolnir doctor:playwright` grades every locator by how it finds an
346
- element — the way a user identifies it (role, label, text), an explicit
347
- contract (`data-testid`), or a structural accident (CSS chains, XPath) —
348
- and scores the file 0–100:
351
+ element: the way a user would (role, label, text), an explicit contract
352
+ (`data-testid`), or a structural accident (CSS chains, XPath). Each file
353
+ gets a score from 0 to 100:
349
354
 
350
355
  ```text
351
- ▚ SELECTOR HEALTH
356
+ ▍ SELECTOR HEALTH
352
357
 
353
358
  e2e/login.spec.ts
354
359
  [█████████████░░░░░░░] 65 / 100
355
- role/text: 1 · testid: 0 · css-chains: 1 ⚠ · xpath: 0
360
+ role/text: 1 · testid: 0 · plain-css: 0 · css-chains: 1 ⚠ · xpath: 0
356
361
 
357
362
  e2e/checkout.spec.ts
358
363
  [█████████████████░░░] 86 / 100
359
- role/text: 3 · testid: 1 · css-chains: 1 ⚠ · xpath: 0
364
+ role/text: 3 · testid: 1 · plain-css: 0 · css-chains: 1 ⚠ · xpath: 0
360
365
  ```
361
366
 
362
- This is **resilience, not correctness**.
367
+ This measures **resilience, not correctness**.
363
368
  `.btn.btn-primary > div:nth-child(2)` passes today and keeps passing until
364
- someone touches the markup. A low score never claims the test is broken —
365
- only that its future depends on markup nobody promised to keep.
369
+ someone touches the markup. A low score never claims the test is broken,
370
+ only that it depends on markup nobody promised to keep.
366
371
 
367
- ---
372
+ <br />
368
373
 
369
374
  ## The Worthiness Score
370
375
 
371
- <table>
372
- <tr>
373
- <td width="50%" align="center" valign="bottom">
374
- <img src="assets/readme/score-gauge.svg" alt="The hammer sweeping every score from 0 to 100 — cracked below 50 (UNWORTHY), strained 50-79 (NEEDS WORK), charged 80-99 (WORTHY), forged at 100 (FORGED) — then holding on FORGED before it loops" width="355" height="430" />
375
- </td>
376
- <td width="50%" align="center" valign="bottom">
377
- <img src="assets/readme/terminal-hero.svg" alt="Mjölnir's deduction breakdown — WORTHINESS 75/100 NEEDS WORK, a diagnostics-by-category bar chart, the per-severity deduction box, and a FIX THIS FIRST list" width="337" height="430" />
378
- </td>
379
- </tr>
380
- <tr>
381
- <td align="center"><strong>What the score means</strong></td>
382
- <td align="center"><strong>Where the points went</strong></td>
383
- </tr>
384
- </table>
385
-
386
- <sub>Left: every score 0–100 through the real `deriveScoreState`. Right: a
387
- real strict scan of `examples/demo-repo`. Both generated
388
- (`npm run docs:gauge` · `npm run docs:hero`) and drift-locked
389
- ([gauge](tests/contract/score-gauge-asset-reproducibility.spec.ts) ·
390
- [breakdown](tests/contract/hero-asset-reproducibility.spec.ts)).</sub>
391
-
392
- | Score | Verdict |
393
- | --------- | ---------------------------------------- |
394
- | `0 – 49` | **UNWORTHY** |
395
- | `50 – 79` | **NEEDS WORK** |
396
- | `80 – 99` | **WORTHY** |
397
- | `100` | **FORGED** |
398
- | `null` | **UNKNOWN** — no test declarations found |
399
-
400
- **How it is computed.** Severity sets a base deduction — `error −8`,
401
- `warning −3`, `info −1` — which the evidence level then discounts: E2 pays
376
+ <p align="center">
377
+ <img src="assets/readme/score-gauge.svg" alt="The worthiness scale from 0 to 100, with a marker sweeping every score: UNWORTHY below 50, NEEDS WORK from 50 to 79, WORTHY from 80 to 99, FORGED at 100" width="720" />
378
+ </p>
379
+
380
+ <sub>Every score from 0 to 100, placed by the real `deriveScoreState`.
381
+ Generated by `npm run docs:gauge` and drift-locked in CI.</sub>
382
+
383
+ | Score | Verdict |
384
+ | --------- | --------------------------------------- |
385
+ | `0 – 49` | **UNWORTHY** |
386
+ | `50 – 79` | **NEEDS WORK** |
387
+ | `80 – 99` | **WORTHY** |
388
+ | `100` | **FORGED** |
389
+ | `null` | **UNKNOWN**: no test declarations found |
390
+
391
+ **How it is computed.** Severity sets a base deduction (`error −8`,
392
+ `warning −3`, `info −1`) and the evidence level discounts it: E2 pays in
402
393
  full, E1 half (rounded down), E0 nothing. The total is normalized by suite
403
- exposure (deductions per test declaration, not per file), and the terminal
404
- prints the same discounted numbers the score used. No hidden second model:
405
- [docs/SCORING.md](docs/SCORING.md) ·
394
+ exposure, meaning deductions per test declaration rather than per file.
395
+ The terminal prints the same discounted numbers the score used; there is
396
+ no hidden second model. Details: [docs/SCORING.md](docs/SCORING.md) and the
406
397
  [scoring guide](https://sergey-bar.github.io/Mjolnir/guide/scoring).
407
398
 
408
- **What 100 does not mean.** Not that the software is correct, the suite
409
- adequate, or the product free of defects. Exactly one thing: **none of
410
- Mjölnir's evaluated rules produced a deduction under this scan and this
411
- evidence model.**
399
+ **What 100 does not mean.** It does not mean the software is correct, the
400
+ suite is adequate, or the product is free of defects. It means one thing:
401
+ **none of Mjölnir's evaluated rules produced a deduction under this scan
402
+ and this evidence model.**
412
403
 
413
- ---
404
+ <br />
414
405
 
415
406
  ## The evidence model
416
407
 
417
- Every finding carries the strength of the evidence behind it. This is the
418
- difference between a tool that reports patterns and a tool you can gate a
419
- release on.
408
+ Every finding carries two labels: how sure Mjölnir is, and how far the
409
+ finding was checked. That is the difference between a tool that reports
410
+ patterns and a tool you can gate a release on.
420
411
 
421
- ```text
422
- STATIC SIGNAL → EVIDENCE LEVEL → RUNTIME CORROBORATION → TRUST DECISION
423
- ```
412
+ **How sure — the evidence level.**
424
413
 
425
- | Level | Name | Means | Deduction |
426
- | ------ | ------------------- | --------------------------------------------------------- | --------- |
427
- | **E2** | Deterministic proof | The defect is structurally present in the code as written | Full |
428
- | **E1** | Pattern evidence | A pattern strongly associated with the defect was matched | Half |
429
- | **E0** | Observation | Worth knowing; not a claim that anything is wrong | Zero |
414
+ | Level | Name | Means | Deduction |
415
+ | ------ | ------------------- | -------------------------------------------------- | --------- |
416
+ | **E2** | Deterministic proof | The defect is present in the code as written | Full |
417
+ | **E1** | Pattern evidence | A pattern strongly tied to the defect matched | Half |
418
+ | **E0** | Observation | Worth knowing. Not a claim that anything is wrong. | Zero |
430
419
 
431
- Confidence in a detection is not strength of proof: a rule can be certain
432
- it matched what it looked for and still be looking at a heuristic. So E1
433
- findings are positioned to be read and judged, never applied blindly — and
420
+ Confidence in a detection is not the strength of the proof. A rule can be
421
+ certain it matched what it looked for and still be looking at a heuristic.
422
+ E1 findings are there to be read and judged, never applied blindly, and
434
423
  that boundary is stamped on the finding in the terminal, the JSON and the
435
424
  agent handoff.
436
425
 
437
- Runtime evidence raises the ceiling. Given a real run report, a finding
438
- climbs a six-rung **trust ladder** from `L0` (observation) to `L5` (the run
439
- verdict corroborates the defect class); the top three rungs structurally
440
- require runtime evidence, so a static-only finding can never claim them.
441
- Rung by rung: [docs/TERMINOLOGY.md](docs/TERMINOLOGY.md).
426
+ **How far it was checked — the trust level.** Most findings come from
427
+ reading your code. Give Mjölnir the report of a real test run and it can
428
+ confirm that the code actually ran.
429
+
430
+ <p align="center">
431
+ <img src="assets/readme/trust-ladder.svg" alt="The trust ladder from L0 to L5. L0 to L2 come from reading the code; L3 to L5 need a real run report, marked by a break in the ladder." width="100%" />
432
+ </p>
433
+
434
+ | Level | In plain words | What it takes |
435
+ | ------ | ---------------------- | ---------------------------------------------- |
436
+ | **L0** | Noted | Reading the code |
437
+ | **L1** | Looks like the problem | Reading the code: a pattern matched |
438
+ | **L2** | Proven in the code | Reading the code: the defect is structural |
439
+ | **L3** | The file ran | A run report shows the finding's file executed |
440
+ | **L4** | The test ran | A run report shows the finding's test executed |
441
+ | **L5** | The run agrees | The run's own result confirms the defect class |
442
+
443
+ A static scan stops at L2. Only a real run report (Playwright JSON, Jest or
444
+ Vitest JSON, JUnit XML) can lift a finding to L3 or above, so a finding
445
+ that was never seen running can never claim it was. Definitions:
446
+ [docs/TERMINOLOGY.md](docs/TERMINOLOGY.md).
442
447
 
443
448
  ### How much of this is measured
444
449
 
445
450
  **<!-- census:measured-of-total -->74 of 79<!-- /census:measured-of-total --> rules carry a false-positive rate measured against real OSS code**
446
- (≥ 10 hand-classified findings each — [docs/FP-AUDIT.md](docs/FP-AUDIT.md)).
447
- The other <!-- census:unmeasured -->5<!-- /census:unmeasured --> ship on the author's estimate and say so, per rule, in
448
- `mjolnir explain`; `mjolnir rules --unmeasured` lists them, and every scan
449
- footer reports how many of the rules that actually _fired_ are measured.
451
+ (at least 10 hand-classified findings each; see
452
+ [docs/FP-AUDIT.md](docs/FP-AUDIT.md)). The other <!-- census:unmeasured -->5<!-- /census:unmeasured --> ship on the author's estimate and say so, rule by rule,
453
+ in `mjolnir explain`. `mjolnir rules --unmeasured` lists them, and every
454
+ scan footer reports how many of the rules that actually _fired_ are
455
+ measured.
450
456
 
451
- The rate is published even when unflattering: QA-PW-141 audits at 43% and
452
- is quarantined for it. **Mjölnir measures its own uncertainty** — that is
453
- the product, not a caveat.
457
+ Rates stay public when they are bad. QA-TEST-001 (a committed `.only`)
458
+ audits poorly on real repositories and sits in quarantine for it. The live
459
+ number for every rule, QA-PW-141 included, is in the audit.
454
460
 
455
461
  ### Trust tiers
456
462
 
457
- Tiers follow measured false-positive behavior, not opinion:
463
+ Tiers follow the measured false-positive rate, not opinion:
458
464
 
459
- | Tier | Measured FP | Behavior |
460
- | -------------- | ----------- | --------------------------------------------- |
461
- | **core** | ≤ 10% | Default report, gates |
462
- | **extended** | ≤ 30% | Default report, lower confidence |
463
- | **quarantine** | > 30% | `--strict` only, capped to info — never gates |
464
- | _unmeasured_ | n < 10 | Cannot be promoted to core until measured |
465
+ | Tier | Measured FP | Behavior |
466
+ | -------------- | ---------------------------- | -------------------------------------------- |
467
+ | **core** | ≤ 10% | Default report, gates |
468
+ | **extended** | ≤ 30% | Default report, lower confidence |
469
+ | **quarantine** | > 30% or explicitly declared | `--strict` only, capped to info, never gates |
470
+ | _unmeasured_ | n < 10 | Cannot be promoted to core until measured |
465
471
 
466
- Promotion and demotion rules, plus per-language maturity:
467
- [rule lifecycle](https://sergey-bar.github.io/Mjolnir/reference/rule-lifecycle).
472
+ FP bands can only demote a tier — they never promote a rule out of
473
+ `quarantine` if it was explicitly declared there. An explicitly
474
+ quarantined rule stays quarantine regardless of its measured FP rate.
468
475
 
469
- ---
476
+ Promotion, demotion and per-language maturity:
477
+ [rule lifecycle](https://sergey-bar.github.io/Mjolnir/reference/rule-lifecycle).
470
478
 
471
- ## Why this is not a linter
479
+ ### Why this is not a linter
472
480
 
473
481
  Linters tell you whether code follows rules. Mjölnir tells you whether your
474
482
  verification can be trusted.
475
483
 
476
484
  | | Linters (ESLint, SonarQube) | Coverage tools | AI code review | **Mjölnir** |
477
485
  | -------------------------------------------------------- | :-------------------------: | :------------: | :------------: | :--------------: |
478
- | Scores the **verification system**, not the product code | ❌ | ❌ | ❌ | ✅ |
479
- | CI workflow integrity (`continue-on-error`, `\|\| true`) | ❌ | ❌ | only the diff | ✅ |
480
- | Grades Playwright locator resilience (Selector Health) | ❌ | ❌ | ❌ | ✅ |
481
- | Reads real run data for `TRUE-FLAKE` verdicts | ❌ | ❌ | ❌ | ✅ |
482
- | Publishes a measured false-positive rate per rule | ❌ | ❌ | ❌ | ✅ |
483
- | Flags tests with no assertions | ✅\* | ❌ | sometimes | ✅ |
484
- | Catches hard sleeps (`waitForTimeout`, `time.sleep`) | ✅\* | ❌ | sometimes | ✅ |
485
- | Deterministic (same input → same output) | ✅ | ✅ | ❌ | ✅ |
486
+ | Scores the **verification system**, not the product code | No | No | No | Yes |
487
+ | CI workflow integrity (`continue-on-error`, `\|\| true`) | No | No | only the diff | Yes |
488
+ | Grades Playwright locator resilience (Selector Health) | No | No | No | Yes |
489
+ | Reads real run data for `TRUE-FLAKE` verdicts | No | No | No | Yes |
490
+ | Publishes a measured false-positive rate per rule | No | No | No | Yes |
491
+ | Flags tests with no assertions | Yes\* | No | sometimes | Yes |
492
+ | Catches hard sleeps (`waitForTimeout`, `time.sleep`) | Yes\* | No | sometimes | Yes |
493
+ | Deterministic (same input, same output) | Yes | Yes | No | Yes |
486
494
  | Cost per scan | free | free | tokens | **zero** (local) |
487
495
 
488
- <sub>\*Covered by `eslint-plugin-jest` / `eslint-plugin-playwright`
496
+ <sub>\*Covered by `eslint-plugin-jest` and `eslint-plugin-playwright`
489
497
  (`expect-expect`, `no-wait-for-timeout`) and by SonarQube's own assertion
490
- rules. Columns describe default behavior aimed at test-suite verification;
491
- plugins, paid tiers and custom rules change some answers. A positioning
492
- summary, not a benchmark.</sub>
498
+ rules. The columns describe default behavior for test-suite verification;
499
+ plugins, paid tiers and custom rules change some answers. This is a
500
+ positioning summary, not a benchmark.</sub>
493
501
 
494
- **Use AI review too.** It catches nuance, intent and design flaws no regex
495
- can find. Mjölnir catches what AI overlooks because it looks intentional —
496
- a committed `.only`, a swallowed exit code, a `continue-on-error` on a test
497
- job. Those are not defects that need reasoning; they are facts that need
498
- scanning.
502
+ Use AI review too. It catches nuance, intent and design flaws no pattern
503
+ can find. Mjölnir catches what AI review overlooks because it looks
504
+ intentional: a committed `.only`, a swallowed exit code, a
505
+ `continue-on-error` on a test job. Those need scanning, not reasoning.
499
506
 
500
- ---
507
+ <br />
501
508
 
502
509
  ## Runtime forensics
503
510
 
504
- Static analysis reasons about code that was never run. Forensics reads what
505
- actually happened — Playwright JSON, Jest JSON, Vitest JSON, and JUnit XML
506
- from any runner:
507
-
508
- ```text
509
- Static analysis → what the code appears to do
510
- Runtime evidence → what the run actually did
511
- both → a finding that can climb the trust ladder
512
- ```
511
+ Static analysis reasons about code that never ran. Forensics reads what
512
+ actually happened: Playwright JSON, Jest JSON, Vitest JSON, and JUnit XML
513
+ from any runner.
513
514
 
514
515
  ```bash
515
516
  mjolnir forensics ./test-results/
516
517
  ```
517
518
 
518
519
  ```text
519
- ▚ FLAKINESS LEADERBOARD
520
+ ▍ FLAKINESS LEADERBOARD
520
521
 
521
522
  3 tests · 1 failed · 1 flaky · 1 retried
522
523
 
@@ -526,29 +527,29 @@ FAILING declines an expired card (e2e/checkout.spec.ts)
526
527
  ████░░░░░░░░░░░░░░░░ 1.1s · 1 attempt
527
528
  ```
528
529
 
529
- `TRUE-FLAKE` is not "this test retried". It is precise: the test **failed
530
- at least one attempt and then finished green** — a lucky pass, flagged
531
- regardless of the final checkmark. `mjolnir triage` turns that history into
532
- a quarantine proposal; `mjolnir pw-report` summarizes a run.
530
+ `TRUE-FLAKE` does not mean the test retried. It means the test **failed at
531
+ least one attempt and then finished green**: a lucky pass, flagged whatever
532
+ the final checkmark says. `mjolnir triage` turns that history into a
533
+ quarantine proposal, and `mjolnir pw-report` summarizes a run. The same run
534
+ reports are what lift findings to trust levels L3 and above.
533
535
 
534
- ---
536
+ <br />
535
537
 
536
538
  ## CI integrity
537
539
 
538
- A test can pass while the pipeline around it is incapable of failing.
539
- Mjölnir reads the workflows too — `continue-on-error`, `|| true`,
540
- unpropagated exit codes, always-success steps, reports consumed but never
541
- generated, and gates skipped on the very events that should block. Each
542
- finding names the job, the step and the line, and carries its own evidence
543
- level; none of them is a claim about CI in general.
540
+ A test can pass while the pipeline around it cannot fail. Mjölnir reads the
541
+ workflows too: `continue-on-error`, `|| true`, exit codes that never
542
+ propagate, always-success steps, reports consumed but never generated, and
543
+ gates skipped on the events that should block. Each finding names the job,
544
+ the step and the line, and carries its own evidence level.
544
545
 
545
- One command generates the PR workflow — advisory by default:
546
+ Generate the PR workflow, advisory by default:
546
547
 
547
548
  ```bash
548
549
  mjolnir ci install
549
550
  ```
550
551
 
551
- Prefer the Marketplace action over a generated workflow? It is one line:
552
+ Or add the Marketplace action to a workflow you already have:
552
553
 
553
554
  ```yaml
554
555
  - uses: Sergey-Bar/Mjolnir@v1
@@ -557,23 +558,24 @@ Prefer the Marketplace action over a generated workflow? It is one line:
557
558
  fail-on: error
558
559
  ```
559
560
 
560
- Pin `@v1` to follow the major line or an exact tag (`@v0.5.32`) for a
561
- reproducible gate — [docs/DISTRIBUTION-KIT.md](docs/DISTRIBUTION-KIT.md)
562
- covers Marketplace, Smithery and the MCP registries.
561
+ Pin `@v1` to follow the major line, or an exact tag (`@v0.5.32`) for a
562
+ reproducible gate. [docs/DISTRIBUTION-KIT.md](docs/DISTRIBUTION-KIT.md)
563
+ covers the Marketplace, Smithery and the MCP registries.
563
564
 
564
- Or wire it into GitHub Code Scanning natively via SARIF:
565
+ To put findings in GitHub Code Scanning, upload SARIF (requires
566
+ `security-events: write` at workflow or job scope):
565
567
 
566
568
  ```yaml
567
569
  - run: npx mjolnir-qa@latest --format sarif > mjolnir.sarif
570
+ continue-on-error: true
568
571
  - uses: github/codeql-action/upload-sarif@v3
572
+ if: ${{ !cancelled() }}
569
573
  with:
570
574
  sarif_file: mjolnir.sarif
571
575
  ```
572
576
 
573
- On GitLab, `--format codequality` emits the Code Quality report the MR
574
- widget and diff annotations consume
575
- ([docs/GITLAB-CI.md](docs/GITLAB-CI.md)).
576
-
577
+ On GitLab, `--format codequality` writes the Code Quality report that the
578
+ MR widget and diff annotations read ([docs/GITLAB-CI.md](docs/GITLAB-CI.md)).
577
579
  Editor and pipeline setup: [docs/SARIF-INTEGRATION.md](docs/SARIF-INTEGRATION.md).
578
580
 
579
581
  ### Changed-scope attribution
@@ -582,20 +584,20 @@ Editor and pipeline setup: [docs/SARIF-INTEGRATION.md](docs/SARIF-INTEGRATION.md
582
584
  npx mjolnir-qa@latest --scope changed
583
585
  ```
584
586
 
585
- Findings are attributed to the lines your branch added, against the
586
- **merge-base**. The scope is the same file set a full scan discovers —
587
- TS/JS specs and adapter configs, `test_*.py`, `*Test.java`, `*Tests.cs`,
588
- `.github/workflows/*.yml` — plus uncommitted and untracked working-tree
589
- changes, so it works before you commit. The base resolves
590
- `main → master → origin/main → origin/master → origin/HEAD`; override with
591
- `--base <ref>`.
587
+ Findings are attributed to the lines your branch added, measured against
588
+ the **merge-base**. The scope is the same file set a full scan discovers
589
+ (TS/JS specs and adapter configs, `test_*.py`, `*Test.java`, `*Tests.cs`,
590
+ `.github/workflows/*.yml`), plus uncommitted and untracked changes, so it
591
+ works before you commit. The base resolves
592
+ `main → master → origin/main → origin/master → origin/HEAD`; override it
593
+ with `--base <ref>`.
592
594
 
593
- When the merge-base cannot be resolved — shallow clone, detached HEAD,
594
- non-git target — findings fall back to full-file attribution **and the
595
- report says that it did.** A silent fallback would be the same class of
596
- defect this tool exists to catch.
595
+ When the merge-base cannot be resolved (a shallow clone, a detached HEAD, a
596
+ target outside git), findings fall back to whole-file attribution **and the
597
+ report says so.** A silent fallback would be the same kind of defect this
598
+ tool exists to catch.
597
599
 
598
- ---
600
+ <br />
599
601
 
600
602
  ## AI agents
601
603
 
@@ -605,14 +607,14 @@ Findings are only worth something if something acts on them.
605
607
  SCAN → EVIDENCE → HANDOFF → AGENT → RE-SCAN → PROOF
606
608
  ```
607
609
 
608
- **AI writes the fix. Mjölnir verifies the fix.** Proof comes from the
610
+ **AI writes the fix. Mjölnir verifies it.** The proof comes from the
609
611
  re-scan, never from the agent's own report of success.
610
612
 
611
613
  | Command | What the agent gets |
612
614
  | ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- |
613
- | `mjolnir mcp` | An [MCP](https://modelcontextprotocol.io) server over stdio — `scan`, `explain` and `diff` become callable tools. |
615
+ | `mjolnir mcp` | An [MCP](https://modelcontextprotocol.io) server over stdio. `scan`, `explain` and `diff` become callable tools. |
614
616
  | `mjolnir handoff` | A saved `--json` report becomes a deterministic Markdown plan: what was detected, the evidence boundary per finding, what must **not** change, how to verify. |
615
- | `mjolnir install` | Writes into the agent surfaces your repo already has — `.claude/`, `.cursor/`, `.kilo/`, `AGENTS.md` — so it re-scans before claiming it is done. |
617
+ | `mjolnir install` | Writes into the agent surfaces your repo already has (`.claude/`, `.cursor/`, `.kilo/`, `AGENTS.md`) so the agent re-scans before it claims it is done. |
616
618
 
617
619
  Add it to a client that ships its own CLI:
618
620
 
@@ -631,130 +633,121 @@ Or to any client that takes an `mcpServers` block:
631
633
  ```
632
634
 
633
635
  **The guardrail matters more than the convenience.** Every finding in a
634
- handoff carries its boundary: **E2** — _deterministic, check the location
635
- and apply the fix_; **E1** — _REQUIRES CONFIRMATION, the observation alone
636
- does not prove the defect_. An agent that fixes E1 blindly, suppresses a
637
- rule, or edits a rule to raise the score is doing the exact thing this tool
638
- exists to catch — so the artifact says so, in the prompt, next to the
639
- finding.
636
+ handoff carries its boundary. **E2** says _deterministic: check the location
637
+ and apply the fix_. **E1** says _REQUIRES CONFIRMATION: the observation
638
+ alone does not prove the defect_. An agent that fixes E1 blindly, suppresses
639
+ a rule, or edits a rule to raise the score is doing exactly what this tool
640
+ exists to catch, so the handoff says so in the prompt, next to the finding.
640
641
 
641
- ---
642
+ <br />
642
643
 
643
644
  ## Trust and security
644
645
 
645
- **Local-first, zero telemetry.** No network-capable API — `fetch`, `http`,
646
- `https`, `net`, `dns`, `dgram`, WebSocket — exists anywhere in `src/`, and
646
+ **Local-first, zero telemetry.** No network-capable API (`fetch`, `http`,
647
+ `https`, `net`, `dns`, `dgram`, WebSocket) exists anywhere in `src/`, and
647
648
  [`privacy-network-isolation.spec.ts`](tests/contract/privacy-network-isolation.spec.ts)
648
- fails the build if one appears (it also bars `eval` and `new Function`).
649
+ fails the build if one appears. It also bars `eval` and `new Function`.
649
650
  Scanning untrusted code never executes it: static analysis reads source
650
651
  text, and forensics parses report files that already exist on disk.
651
652
 
652
- Two caveats worth stating: `npx` itself fetches the package before
653
- anything runs, and the guarantee covers `src/` — not third-party plugins.
654
-
655
- **Plugins are not sandboxed, and this will not be dressed up.** JS plugins
656
- (`mjolnir-rules/*.mjs`, or npm packages under `"plugins"`) run with full
657
- Node privileges — the same trust model as ESLint or Vitest plugins. So
658
- loading them is opt-in **per scan**: without `--enable-plugins` (or
659
- `MJOLNIR_ENABLE_PLUGINS=1`) the sources are never loaded, and a stderr
660
- notice lists what was skipped. JSON rule manifests execute no code by
661
- design, and core rule-ID prefixes are reserved so a plugin cannot
662
- impersonate one. Vulnerabilities: [SECURITY.md](SECURITY.md).
663
-
664
- ### We run it on ourselves
665
-
666
- A verification trust engine has no standing unless it is itself verifiable.
667
- Every CI run scans this repository **with the build that same run
668
- produced**, and the gate fails on any error-severity finding — but also on
669
- a **partial** scan or a **crashed rule**, because a truncated self-scan
670
- that reports nothing is precisely the false green this project exists to
653
+ Two caveats: `npx` itself fetches the package before anything runs, and the
654
+ guarantee covers `src/`, not third-party plugins.
655
+
656
+ **Plugins are not sandboxed.** JS plugins (`mjolnir-rules/*.mjs`, or npm
657
+ packages listed under `"plugins"`) run with full Node privileges, the same
658
+ trust model as ESLint or Vitest plugins. Loading them is opt-in **per
659
+ scan**: without `--enable-plugins` (or `MJOLNIR_ENABLE_PLUGINS=1`) their
660
+ sources are never loaded, and a notice on stderr lists what was skipped.
661
+ JSON rule manifests execute no code, and core rule-ID prefixes are reserved
662
+ so a plugin cannot impersonate one. Report vulnerabilities through
663
+ [SECURITY.md](SECURITY.md).
664
+
665
+ **It runs on itself.** A verification trust engine has no standing unless
666
+ it is itself verifiable. Every CI run scans this repository with the build
667
+ that same run produced. The gate fails on any error-severity finding, and
668
+ also on a **partial** scan or a **crashed rule**, because a truncated
669
+ self-scan that reports nothing is the false green this project exists to
671
670
  catch. `mjolnir doctor` re-audits the rule base in the same run (fixture
672
- firewall, tier honesty, the core-tier cap), where an INCONCLUSIVE check
673
- fails exactly like a failing one. Both reports are uploaded as build
674
- artifacts.
675
-
676
- ---
677
-
678
- ## What Mjölnir cannot tell you
679
-
680
- - **It does not run your tests.** A clean scan is not a passing suite.
681
- - **It cannot tell you an assertion is _wrong_.** `expect(total).toBe(41)`
682
- looks perfectly healthy. Mjölnir finds tests that _cannot fail_ and
683
- pipelines that _cannot go red_ — not tests that check the wrong thing.
684
- - **It does not prove business correctness.** Nothing here says your
685
- product does what the requirement asked for.
686
- - **A 100 is not proof of a good suite.** Whether your suite covers your
687
- actual risk is a different question, and this tool does not answer it.
688
- - **<!-- census:unmeasured-of-total -->5 of 79<!-- /census:unmeasured-of-total --> rules ship on an estimate**, not a measured rate — disclosed
689
- per rule, not buried here.
690
- - **E1 is not E2.** Heuristic findings are worth reading, not worth
691
- applying blindly.
692
- - **An empty repo scores `null`, never 100.**
671
+ firewall, tier honesty, the core-tier cap), and an INCONCLUSIVE check fails
672
+ exactly like a failing one. Both reports are uploaded as build artifacts.
693
673
 
694
- ---
674
+ ### Exit codes and the machine contract
695
675
 
696
- ## Exit codes and the machine contract
697
-
698
- Frozen surfaces — safe to build CI logic on:
676
+ Frozen, so you can build CI logic on them:
699
677
 
700
678
  | Exit code | Meaning |
701
679
  | --------- | --------------------------------------------------------------- |
702
- | `0` | Clean — no findings at or above the gate |
680
+ | `0` | Clean: no findings at or above the gate |
703
681
  | `1` | Findings at or above the gate |
704
- | `2` | Partial scan (time budget hit, unreadable files) — never blocks |
682
+ | `2` | Partial scan (time budget hit, unreadable files). Never blocks. |
705
683
  | `10` | Usage error (bad flag, missing target) |
706
684
  | `20` | Internal error |
707
685
 
708
686
  `2` is deliberately distinct from `0`: a scan that did not finish has not
709
- found nothing — it has not finished looking.
687
+ found nothing. It has not finished looking.
710
688
 
711
- Everything a machine consumes — MCP tool results, `--json`, SARIF 2.1 —
712
- comes off one canonical result under a versioned, **additive-only** schema
713
- (`schemaVersion: 1`, `contractVersion: 1`), so no consumer reconstructs
714
- semantics from rendered text: [the machine contract](docs/machine-contract.md).
689
+ Everything a machine consumes (MCP tool results, `--json`, SARIF 2.1) comes
690
+ from one canonical result under a versioned, **additive-only** schema
691
+ (`schemaVersion: 1`, `contractVersion: 1`), so no consumer has to rebuild
692
+ meaning from rendered text. See [the machine contract](docs/machine-contract.md).
715
693
  Rule IDs (`QA-<FAMILY>-NNN`) are immutable once shipped and never reused.
716
694
 
717
- ---
695
+ <br />
696
+
697
+ ## What Mjölnir cannot tell you
698
+
699
+ - **It does not run your tests.** A clean scan is not a passing suite.
700
+ - **It cannot tell you an assertion is _wrong_.** `expect(total).toBe(41)`
701
+ looks healthy. Mjölnir finds tests that _cannot fail_ and pipelines that
702
+ _cannot go red_, not tests that check the wrong thing.
703
+ - **It does not prove business correctness.** Nothing here says your
704
+ product does what the requirement asked for.
705
+ - **A 100 is not proof of a good suite.** Whether your suite covers your
706
+ real risk is a different question, and this tool does not answer it.
707
+ - **<!-- census:unmeasured-of-total -->5 of 79<!-- /census:unmeasured-of-total --> rules ship on an estimate**, not a measured rate. Each one
708
+ says so on its own finding.
709
+ - **E1 is not E2.** Heuristic findings are worth reading, not worth applying
710
+ blindly.
711
+ - **An empty repo scores `null`, never 100.**
712
+ - **A file named `*.spec.ts` with no test declarations does not count as coverage.** A repo whose only spec files contain imports or types (zero `it`/`test` calls) scores `null`, not 100.
713
+
714
+ <br />
718
715
 
719
716
  ## Documentation
720
717
 
718
+ The full docs site is at <https://sergey-bar.github.io/Mjolnir/>.
719
+
721
720
  | Document | What's in it |
722
721
  | ------------------------------------------------------ | ------------------------------------------------- |
723
- | [docs/SCORING.md](docs/SCORING.md) | Score normalization + evidence weighting |
724
- | [docs/TERMINOLOGY.md](docs/TERMINOLOGY.md) | Canonical vocabulary — one word per concept |
725
- | [docs/FP-AUDIT.md](docs/FP-AUDIT.md) | Measured false-positive rates + method |
722
+ | [docs/SCORING.md](docs/SCORING.md) | Score normalization and evidence weighting |
723
+ | [docs/TERMINOLOGY.md](docs/TERMINOLOGY.md) | Canonical vocabulary: one word per concept |
724
+ | [docs/FP-AUDIT.md](docs/FP-AUDIT.md) | Measured false-positive rates and the method |
726
725
  | [docs/RULE-LIFECYCLE.md](docs/RULE-LIFECYCLE.md) | Rule states, tiers, suppression, deprecation |
727
726
  | [docs/VERSIONING.md](docs/VERSIONING.md) | Semver policy, frozen surfaces, deprecation cycle |
728
727
  | [docs/machine-contract.md](docs/machine-contract.md) | The canonical machine-readable result |
729
- | [docs/SARIF-INTEGRATION.md](docs/SARIF-INTEGRATION.md) | SARIF output + editor/CI setup |
728
+ | [docs/SARIF-INTEGRATION.md](docs/SARIF-INTEGRATION.md) | SARIF output and editor or CI setup |
730
729
  | [docs/GITLAB-CI.md](docs/GITLAB-CI.md) | GitLab: Code Quality report, MR recipe, gate |
731
730
  | [docs/rules/](docs/rules/) | Generated per-rule catalog |
732
- | [CONTRIBUTING.md](CONTRIBUTING.md) | Dev setup + contribution workflow |
731
+ | [CONTRIBUTING.md](CONTRIBUTING.md) | Dev setup and contribution workflow |
733
732
  | [SUPPORT.md](SUPPORT.md) | Where to ask, report and get help |
734
733
  | [SECURITY.md](SECURITY.md) | Vulnerability reporting |
735
734
  | [CHANGELOG.md](CHANGELOG.md) | Release history |
736
735
 
737
- Full docs site: <https://sergey-bar.github.io/Mjolnir/>.
738
-
739
- ---
736
+ ### Status
740
737
 
741
- ## Status
742
-
743
- **v0.5.x · open beta.** The JSON schema and the exit codes are frozen
744
- contracts. TypeScript and Python have the broadest measured coverage; Java
745
- and C# are newer — read them through the
738
+ **Version 1.** The JSON schema and the exit codes are frozen contracts.
739
+ TypeScript and Python have the broadest measured coverage. Java and C# are
740
+ newer; read them through the
746
741
  [maturity table](https://sergey-bar.github.io/Mjolnir/reference/rule-lifecycle).
747
- Honest scope, no invented dates:
742
+ What comes next, with no invented dates:
748
743
  [the public roadmap](https://sergey-bar.github.io/Mjolnir/reference/roadmap).
749
744
 
750
- ---
751
-
752
- ## Contributing
745
+ ### Contributing
753
746
 
754
747
  New rules are the easiest first contribution. One command scaffolds the
755
- rule plus its must-fire **and** must-not-fire fixtures — and the generated
756
- rule intentionally fails its own fixtures until real detection is
757
- implemented, because a stub that ships is a rule nobody measured:
748
+ rule with its must-fire **and** must-not-fire fixtures. The generated rule
749
+ fails its own fixtures on purpose until real detection is written, because
750
+ a stub that ships is a rule nobody measured:
758
751
 
759
752
  ```bash
760
753
  mjolnir create-rule QA-PW-140 --title "Screenshot without diff bound"
@@ -763,26 +756,23 @@ mjolnir create-rule QA-PW-140 --title "Screenshot without diff bound"
763
756
  Dev setup, the standing-gate commands, and the anti-creep and
764
757
  fixture-firewall laws are in [CONTRIBUTING.md](CONTRIBUTING.md).
765
758
 
766
- ---
767
-
768
- ## The Mjölnir Standard
769
-
770
- Don't ask whether the tests passed.
771
-
772
- Ask whether the evidence proves they deserve to be trusted.
759
+ <br />
773
760
 
774
761
  <div align="center">
775
762
 
776
- ---
777
-
778
- **Stop shipping tests you can't trust.**
763
+ <img src="assets/readme/closing.svg" alt="Run it on your repo." width="100%" />
779
764
 
780
765
  ```bash
781
766
  npx mjolnir-qa@latest
782
767
  ```
783
768
 
784
- **Star ⭐ · Watch 👀 · Contribute 🤝**
769
+ [Read the guide](https://sergey-bar.github.io/Mjolnir/guide/getting-started) · [Docs site](https://sergey-bar.github.io/Mjolnir/) · [npm](https://www.npmjs.com/package/mjolnir-qa)
770
+
771
+ <br />
772
+
773
+ Don't ask whether the tests passed.<br />
774
+ Ask whether the evidence proves they deserve to be trusted.
785
775
 
786
- Built by [Sergey Bar](https://www.linkedin.com/in/sergeybar/)
776
+ <sub>Built by [Sergey Bar](https://www.linkedin.com/in/sergeybar/) · MIT licensed</sub>
787
777
 
788
778
  </div>