mjolnir-qa 0.5.0 → 0.5.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -4,19 +4,15 @@
4
4
 
5
5
  ### Your tests are lying to you. We prove it.
6
6
 
7
- **Verification Trust Engine for QA.** Audits test suites and CI pipelines,
8
- reports a worthiness score, and shows exactly where trust breaks.
7
+ **Verification Trust Engine for QA.** Mjölnir audits test suites and CI
8
+ pipelines, reports a worthiness score, and shows exactly where trust breaks.
9
9
 
10
- <sub>"We prove it" is literal for the deterministic findings — a committed `.only`,
11
- a `continue-on-error` on a test job. The pattern-based rules are labelled
12
- heuristic, and **15 of 91 carry a false-positive rate measured against real
13
- OSS code**; `mjolnir rules --unmeasured` names the rest. See
14
- [docs/FP-AUDIT.md](docs/FP-AUDIT.md).</sub>
10
+ [![npm](https://img.shields.io/npm/v/mjolnir-qa.svg?style=flat-square&color=C9A227&labelColor=0B0F17)](https://www.npmjs.com/package/mjolnir-qa)
11
+ [![ci](https://img.shields.io/github/actions/workflow/status/Sergey-Bar/Mjolnir/ci.yml?branch=main&style=flat-square&label=ci&labelColor=0B0F17)](https://github.com/Sergey-Bar/Mjolnir/actions/workflows/ci.yml)
12
+ [![license](https://img.shields.io/badge/license-MIT-C9A227.svg?style=flat-square&labelColor=0B0F17)](LICENSE)
13
+ [![node](https://img.shields.io/badge/node-%E2%89%A5%2022.18-2E8C7F.svg?style=flat-square&labelColor=0B0F17)](https://nodejs.org)
15
14
 
16
- [![npm](https://img.shields.io/npm/v/mjolnir-qa.svg?style=flat-square&color=B45309&labelColor=0D0D0D)](https://www.npmjs.com/package/mjolnir-qa)
17
- [![ci](https://img.shields.io/github/actions/workflow/status/Sergey-Bar/Mjolnir/ci.yml?branch=main&style=flat-square&label=ci&labelColor=0D0D0D)](https://github.com/Sergey-Bar/Mjolnir/actions/workflows/ci.yml)
18
- [![license](https://img.shields.io/badge/license-MIT-B45309.svg?style=flat-square&labelColor=0D0D0D)](LICENSE)
19
- [![node](https://img.shields.io/badge/node-%E2%89%A5%2022.18-146A8C.svg?style=flat-square&labelColor=0D0D0D)](https://nodejs.org)
15
+ English | [简体中文](README.zh.md) | [繁體中文](README.zht.md) | [한국어](README.ko.md) | [Deutsch](README.de.md) | [Español](README.es.md) | [Français](README.fr.md) | [Italiano](README.it.md) | [Dansk](README.da.md) | [日本語](README.ja.md) | [Polski](README.pl.md) | [Русский](README.ru.md) | [Norsk](README.no.md) | [Português (Brasil)](README.br.md) | [ไทย](README.th.md) | [Türkçe](README.tr.md) | [Українська](README.uk.md) | [বাংলা](README.bn.md) | [Ελληνικά](README.gr.md) | [Tiếng Việt](README.vi.md) | [עברית](README.he.md) | [العربية](README.ar.md) | [Bosanski](README.bs.md)
20
16
 
21
17
  ```bash
22
18
  npx mjolnir-qa@latest
@@ -24,111 +20,93 @@ npx mjolnir-qa@latest
24
20
 
25
21
  **Are your tests worthy of trust?**
26
22
 
27
- [Quickstart](#-quickstart) · [Rules](#-the-rules) · [Selector Health](#-selector-health-score) · [Runtime Evidence](#-runtime-evidence) · [CI Integration](#-ci-integration) · [Contributing](#-contributing)
23
+ [See it work](#-see-it-work) · [Quickstart](#-quickstart) · [What it checks](#-what-mjölnir-checks) · [Scoring](#how-the-score-works) · [CI](#-ci-integration) · [Configuration](#configuration) · [Docs](#-documentation)
28
24
 
29
25
  </div>
30
26
 
31
27
  ---
32
28
 
33
- ## 🔨 What is Mjölnir?
29
+ ## 🎬 See it work
34
30
 
35
- | | |
36
- | --- | ----------------------------------------------------------------------------------------------------------------- |
37
- | ⚖️ | **Worthiness Score** — one number, transparent deduction table, no black box |
38
- | 🎭 | **Selector Health Score** — grades your Playwright locators, not just your pass rate |
39
- | 🔬 | **Runtime forensics** — reads real Playwright/JUnit run data to catch `TRUE-FLAKE`, not just static guesses |
40
- | 🚨 | **CI-integrity rules** — catches `continue-on-error`, `\|\| true`, and other false-green tricks |
41
- | 🐍 | **All four Playwright bindings** — TypeScript, Python, Java, C#/.NET — plus pytest, JUnit/TestNG and CI workflows |
42
- | 🔒 | **Local-first** — zero network calls while scanning, zero telemetry, runs in seconds |
43
-
44
- ---
31
+ <p align="center">
32
+ <img src="assets/readme/demo.svg" alt="Mjölnir's full --verbose report on a demo repo: WORTHINESS 75/100 NEEDS WORK, a diagnostics-by-category breakdown, a FIX THIS FIRST list, and every finding with its rule ID and line number across CI, Playwright, test-hygiene and Python rules" width="900" />
33
+ </p>
45
34
 
46
- ## ⚡ Mjölnir is not another linter
35
+ <sub>The complete `npx mjolnir-qa ./examples/demo-repo --verbose` output,
36
+ rendered from the actual reporter — nothing trimmed. Regenerated by
37
+ `npm run docs:demo`;
38
+ [`tests/demo-asset-reproducibility.spec.ts`](tests/demo-asset-reproducibility.spec.ts)
39
+ fails CI if it drifts from what the tool prints.</sub>
47
40
 
48
- Linters tell you whether code follows rules.
49
- Mjölnir tells you whether your verification can be trusted.
41
+ **What just happened:**
50
42
 
51
- | | ESLint / SonarQube | Coverage tools | Manual review | **Mjölnir** |
52
- | -------------------------------------------------------- | :----------------: | :------------: | :-----------: | :---------: |
53
- | CI workflow integrity (`continue-on-error`, `\|\| true`) | ❌ | ❌ | rarely | ✅ |
54
- | Cross-language (TS, Python, Java, C#) from one tool | ❌ | ❌ | ❌ | ✅ |
55
- | Grades Playwright locator resilience (Selector Health) | ❌ | ❌ | rarely | ✅ |
56
- | Flags tests with no real assertions | ✅ (plugin)\* | ❌ | sometimes | ✅ |
57
- | Catches hard sleeps (`waitForTimeout`, `time.sleep`) | ✅ (plugin)\* | ❌ | sometimes | ✅ |
58
- | Runs in seconds, zero network calls while scanning | ✅ | ✅ | — | ✅ |
43
+ 1. Mjölnir discovered the Playwright specs, its config, the CI workflow and
44
+ a Python test file — four languages/formats, one pass.
45
+ 2. It found evidence that weakens trust in the suite — a `continue-on-error`
46
+ masking a job, a `|| true` swallowing an exit code, hard sleeps, a
47
+ brittle selector, hardcoded staging URLs, a `networkidle` wait.
48
+ 3. It turned each into a concrete finding with a rule ID, a location and a
49
+ fix — and a single score you can gate a PR on.
59
50
 
60
- \*`eslint-plugin-jest` (`expect-expect`) and `eslint-plugin-playwright` (`expect-expect`, `no-wait-for-timeout`) cover these for their respective frameworks.
51
+ ### One finding, up close
61
52
 
62
- **Runtime Analysis** — a separate category from static linting:
53
+ Run `mjolnir explain QA-CI-001` on the first finding above and you get:
63
54
 
64
- | | Playwright retry reporter | Allure / ReportPortal | **Mjölnir forensics** |
65
- | --------------------------------------------- | :-----------------------: | :-------------------: | :-------------------: |
66
- | Reads real run data for `TRUE-FLAKE` verdicts | partial\* | partial (tag) | ✅ |
67
- | Flaky-triage report from execution history | ❌ | ✅ | ✅ |
68
- | Integrates with static worthiness score | ❌ | ❌ | ✅ |
55
+ ```text
56
+ ▚▞ QA-CI-001 — continue-on-error masks a failing verification gate
69
57
 
70
- \*Playwright tracks retries internally but does not produce a standalone flakiness report with verdict labels.
58
+ Severity: error
59
+ Confidence: high
60
+ Evidence: E2
61
+ Measured FP: not yet measured — this rule ships on assumption (see docs/FP-AUDIT.md)
71
62
 
72
- ---
63
+ WHAT WAS FOUND (real detector output, not a mockup)
64
+ Job `security-scan` runs a verification gate under `continue-on-error: true`.
73
65
 
74
- ## 🤖 Why not just use AI code review?
75
-
76
- AI reviewers are powerful — but they're expensive, intermittent, and reactive.
77
- Mjölnir is cheap, continuous, and structural. They complement each other, but
78
- only one of them catches the `continue-on-error` that's been silently hiding
79
- failures since the repo was created.
66
+ WHY IT MATTERS
67
+ This job can fail every day and CI will still show green. The checkmark
68
+ on this workflow cannot be trusted.
80
69
 
81
- | | AI code review (Copilot, etc.) | **Mjölnir** |
82
- | ------------------------------------- | :----------------------------: | :-------------------------: |
83
- | Cost per scan | Tokens (scales with diff size) | **Zero** (local, installed) |
84
- | Sees the whole suite + all CI configs | Only the PR diff you show it | **Everything, every time** |
85
- | Deterministic (same input → same out) | ❌ (non-deterministic) | **✅** |
86
- | Catches patterns dormant for months | Only if it's in the context | **✅** (scans all files) |
87
- | Remembers findings between runs | ❌ (no memory across sessions) | **✅** (baseline + diff) |
88
- | Runs without human triggering | Needs a PR or prompt | **✅** (CI hook, 3 seconds) |
89
- | Attention budget | Limited by context window | **Unlimited** (every file) |
70
+ HOW TO FIX
71
+ Remove continue-on-error, or scope it to individual non-blocking steps only.
72
+ ```
90
73
 
91
- **The honest answer:** Use both. AI catches nuance, intent, and design flaws
92
- that no regex can find. Mjölnir catches the structural patterns that AI
93
- overlooks because they look "intentional" — a committed `.only`, a swallowed
94
- exit code, a `continue-on-error` on a test job. Those aren't bugs that need
95
- reasoning; they're facts that need scanning. And scanning is what a
96
- deterministic tool does better, faster, and for free.
74
+ That is the unit of value: not a style nit, but a place where your CI is
75
+ telling you something passed when it didn't.
97
76
 
98
77
  ---
99
78
 
100
79
  ## ⚡ Quickstart
101
80
 
102
- **The product is one command in CI:**
81
+ Run it against a repo for a full report and a worthiness score:
103
82
 
104
83
  ```bash
105
- npx mjolnir-qa@latest --scope changed
84
+ npx mjolnir-qa@latest
106
85
  ```
107
86
 
108
- It scans the test files, Playwright configs and CI workflows touched in the
109
- branch, and exits non-zero when it finds new problems. Drop that in a PR check
110
- (`mjolnir ci install` writes the workflow) and you're done. Everything below is
111
- optional.
112
-
113
- Run it with no flags for a full-repo report and a WORTHINESS score:
87
+ **In CI, the product is one command.** It scans only what the branch
88
+ touched and exits non-zero on new problems:
114
89
 
115
90
  ```bash
116
- npx mjolnir-qa@latest
91
+ npx mjolnir-qa@latest --scope changed
117
92
  ```
118
93
 
119
- ### Everyday
94
+ Drop that in a PR check — `mjolnir ci install` writes the workflow —
95
+ and you're done. Everything else is optional.
120
96
 
121
97
  | Command | What it does |
122
98
  | ----------------------------------- | ------------------------------------------------ |
99
+ | `mjolnir` | Full-repo scan + worthiness score |
123
100
  | `mjolnir --scope changed` | Only what your branch introduced — the CI form |
124
- | `mjolnir` | Full-repo scan + WORTHINESS score |
125
- | `mjolnir --json` / `--format sarif` | Machine-readable / GitHub Code Scanning |
126
101
  | `mjolnir ci install` | Generate the advisory PR workflow |
127
- | `mjolnir explain <RULE-ID>` | What/why/fix + measured FP rate for one rule |
102
+ | `mjolnir explain QA-CI-001` | What / why / fix + measured FP rate for one rule |
128
103
  | `mjolnir rules --unmeasured` | The rules running on assumption, not measurement |
104
+ | `mjolnir --json` / `--format sarif` | Machine-readable / GitHub Code Scanning |
129
105
  | `mjolnir --strict` | Also run quarantine-tier rules (higher FP risk) |
106
+ | `mjolnir --cache` | Incremental re-scans via a local verdict cache |
130
107
 
131
- ### When something's flaky
108
+ <details>
109
+ <summary><strong>When something's flaky</strong></summary>
132
110
 
133
111
  | Command | What it does |
134
112
  | ----------------------------------- | --------------------------------------------------- |
@@ -137,77 +115,59 @@ npx mjolnir-qa@latest
137
115
  | `mjolnir pw-report ./test-results/` | Playwright run summary — retries / flakes / slowest |
138
116
  | `mjolnir doctor:playwright` | Playwright-only deep scan + Selector Health Score |
139
117
 
118
+ </details>
119
+
140
120
  <details>
141
121
  <summary><strong>Occasional / reporting</strong></summary>
142
122
 
143
- | Command | What it does |
144
- | -------------------------------- | ------------------------------------------------ |
145
- | `mjolnir fix --dry-run` / `fix` | Safe auto-fixes with proof |
146
- | `mjolnir baseline` / `diff` | Snapshot findings, then report only new/worsened |
147
- | `mjolnir impact [--since <ref>]` | What changed since a prior commit |
148
- | `mjolnir debt` | Test-debt register with a cost model |
149
- | `mjolnir handover` | New-QA onboarding map of the suite |
150
- | `mjolnir stats` | Local all-time counters of fixes seen |
151
- | `mjolnir badge` | shields.io endpoint JSON + snippet |
152
- | `mjolnir rules` / `rules --md` | Full rule catalog (JSON or Markdown) |
153
- | `mjolnir doctor` | Self-audit of Mjölnir's own rule base |
154
- | `mjolnir create-rule <ID>` | Scaffold a new rule + fixtures |
155
- | `mjolnir --format mermaid` | Test-architecture diagram for a PR comment |
123
+ | Command | What it does |
124
+ | ------------------------------- | ------------------------------------------------ |
125
+ | `mjolnir fix --dry-run` / `fix` | Safe auto-fixes with proof |
126
+ | `mjolnir baseline` / `diff` | Snapshot findings, then report only new/worsened |
127
+ | `mjolnir impact --since <ref>` | What changed since a prior commit |
128
+ | `mjolnir debt` | Test-debt register with a cost model |
129
+ | `mjolnir handover` | New-QA onboarding map of the suite |
130
+ | `mjolnir stats` | Local all-time counters of fixes seen |
131
+ | `mjolnir badge` | shields.io endpoint JSON + snippet |
132
+ | `mjolnir rules --md` | Full rule catalog (JSON or Markdown) |
133
+ | `mjolnir doctor` | Self-audit of Mjölnir's own rule base |
134
+ | `mjolnir create-rule <ID>` | Scaffold a new rule + fixtures |
135
+ | `mjolnir --format mermaid` | Test-architecture diagram for a PR comment |
156
136
 
157
137
  </details>
158
138
 
159
- ---
160
-
161
- ## ⚖️ Worthiness Score
162
-
163
- <p align="center">
164
- <img src="assets/readme/terminal-hero.svg" alt="Mjölnir terminal output — WORTHINESS 67/100 NEEDS WORK, a diagnostics-by-category breakdown, and a FIX THIS FIRST list" width="760" />
165
- </p>
166
-
167
- <sub>Real output of `npx mjolnir-qa ./examples/demo-repo` in this repo. Regenerated by `npm run docs:hero`; [`tests/hero-asset-reproducibility.spec.ts`](tests/hero-asset-reproducibility.spec.ts) fails CI if it drifts from what the reporter actually prints.</sub>
168
-
169
- The score is transparent: error −8, warning −3, info −1, then normalized by
170
- suite exposure (deductions per test declaration). Evidence-weighted deductions
171
- mean weak signals cost less. The terminal shows the same discounted numbers the
172
- score uses — no black box.
173
-
174
- **Verdicts:**
139
+ Install globally instead of `npx` if you prefer: `npm i -g mjolnir-qa`.
140
+ Requires Node.js ≥ 22.18. Works on Windows, macOS, and Linux.
175
141
 
176
- | Score | Verdict |
177
- | ------- | ---------------- |
178
- | ≥ 80 | ✓ **WORTHY** |
179
- | 50 – 79 | ⚠ **NEEDS WORK** |
180
- | < 50 | ✖ **UNWORTHY** |
181
-
182
- **Evidence Levels:**
142
+ ---
183
143
 
184
- Every finding carries an evidence level that determines its weight in the score:
144
+ ## 👥 Who is this for?
185
145
 
186
- | Level | Meaning | Score impact | Example |
187
- | ----- | -------------------- | ---------------- | ---------------------------------------------------------------------------------- |
188
- | E2 | Deterministic defect | Full deduction | `.only` committed — structurally provable; downgraded to E1 when confidence is low |
189
- | E1 | Heuristic pattern | Half deduction | Regex-matched `sleep()` — strong signal, not proof |
190
- | E0 | Observation | Zero (info only) | Reported but never gates CI or deducts |
146
+ - **QA / SDET** owning an e2e or integration suite who need evidence the
147
+ suite actually deserves the green checkmark it produces.
148
+ - **Platform / DevEx** teams responsible for CI integrity and release
149
+ gates — the people who care that a `continue-on-error` never silently
150
+ turns a red pipeline green.
151
+ - **OSS maintainers** who want a cheap, always-on verification gate that
152
+ runs locally and in CI with zero network calls.
191
153
 
192
- Most rules are **E1** (heuristic). The tagline "we prove it" refers to this
193
- evidence-level system — deterministic findings (E2) are structural proof;
194
- heuristic findings (E1) are correctly-positioned warnings, not formal proofs.
154
+ ---
195
155
 
196
- **How much of this is measured.** 15 of 91 rules carry a false-positive rate
197
- measured against real OSS code (≥ 10 hand-classified findings each; see
198
- [docs/FP-AUDIT.md](docs/FP-AUDIT.md)). The other 76 ship on the author's
199
- estimate. Every scan footer tells you how many of the rules that _fired_ are
200
- measured; `mjolnir rules --unmeasured` lists the ones that aren't; every rule's
201
- `mjolnir explain` page states its status. We publish the rate even when it's
202
- ugly — QA-JV-103 audits at 50% and is quarantined for it. Growing that 15
203
- is the project's main open work.
156
+ ## 🔨 What Mjölnir checks
204
157
 
205
- ---
158
+ | | |
159
+ | --- | ----------------------------------------------------------------------------------------------------------------- |
160
+ | ⚖️ | **Worthiness Score** — one number, transparent deduction table, no black box |
161
+ | 🎭 | **Selector Health Score** — grades your Playwright locators, not just your pass rate |
162
+ | 🔬 | **Runtime forensics** — reads real Playwright/JUnit run data to catch `TRUE-FLAKE`, not just static guesses |
163
+ | 🚨 | **CI-integrity rules** — catches `continue-on-error`, `\|\| true`, and other false-green tricks |
164
+ | 🐍 | **All four Playwright bindings** — TypeScript, Python, Java, C#/.NET — plus pytest, JUnit/TestNG and CI workflows |
165
+ | 🔒 | **Local-first** — zero network calls while scanning, zero telemetry, runs in seconds |
206
166
 
207
- ## 📋 The Rules
167
+ ### The rules
208
168
 
209
169
  Every rule ships with must-fire **and** must-not-fire fixtures. A rule that
210
- fires on its own negative fixture cannot ship. That's the false-positive
170
+ fires on its own negative fixture cannot ship — that's the false-positive
211
171
  firewall.
212
172
 
213
173
  <details>
@@ -230,7 +190,7 @@ firewall.
230
190
 
231
191
  | ID | Rule | Severity |
232
192
  | ------------ | --------------------------- | -------- |
233
- | QA-TQUAL-001 | Mock-only verification | warning |
193
+ | QA-TQUAL-001 | Mock-only verification | info |
234
194
  | QA-TQUAL-002 | Tautological assertion | error |
235
195
  | QA-TQUAL-009 | Unawaited promise assertion | error |
236
196
  | QA-TQUAL-011 | Commented-out tests | warning |
@@ -245,9 +205,9 @@ firewall.
245
205
  | QA-PW-002 | Unawaited locator assertion | error |
246
206
  | QA-PW-003 | `page.pause()` / `test.only()` committed | error |
247
207
  | QA-PW-004 | Brittle CSS/XPath selectors | warning |
248
- | QA-PW-005 | Business logic inside `page.evaluate()` | warning |
249
- | QA-PW-114 | Legacy element handles (`page.$`) | warning |
250
- | QA-PW-118 | `networkidle` waits (flaky by design) | warning |
208
+ | QA-PW-005 | Business logic inside `page.evaluate()` | info |
209
+ | QA-PW-114 | Legacy element handles (`page.$`) | info |
210
+ | QA-PW-118 | `networkidle` waits (flaky by design) | info |
251
211
  | QA-PW-123 | Hardcoded environment URLs | warning |
252
212
 
253
213
  </details>
@@ -275,8 +235,8 @@ firewall.
275
235
  | QA-PY-002 | Skipped test (`skip`, non-strict `xfail`) | warning |
276
236
  | QA-PY-003 | Test function with no assertions | error |
277
237
  | QA-PY-005 | `time.sleep()` in tests | warning |
278
- | QA-PY-006 | Empty test body (`pass`) | error |
279
- | QA-PY-010 | Random/time dependence without freeze | warning |
238
+ | QA-PY-006 | Empty test body (`pass`) | info |
239
+ | QA-PY-010 | Random/time dependence without freeze | info |
280
240
  | QA-PY-012 | Tautological assertion | error |
281
241
 
282
242
  20 Python rules total (QA-PY-001…012 pytest hygiene + QA-PY-101…108 Playwright-Python).
@@ -293,8 +253,8 @@ firewall.
293
253
  | QA-JV-103 | Test method with no assertions | error |
294
254
  | QA-JV-105 | Playwright `waitForTimeout()` hard sleep | warning |
295
255
  | QA-JV-106 | Brittle selector instead of role locator | warning |
296
- | QA-JV-108 | Hardcoded environment URL in test | warning |
297
- | QA-JV-111 | Blanket `page.route("**")` mock | warning |
256
+ | QA-JV-108 | Hardcoded environment URL in test | info |
257
+ | QA-JV-111 | Blanket `page.route("**")` mock | info |
298
258
 
299
259
  </details>
300
260
 
@@ -308,25 +268,92 @@ firewall.
308
268
  | QA-CS-103 | Test method with no assertions | error |
309
269
  | QA-CS-105 | `WaitForTimeoutAsync()` hard sleep | warning |
310
270
  | QA-CS-106 | Brittle selector instead of role locator | warning |
311
- | QA-CS-108 | Hardcoded environment URL in test | warning |
312
- | QA-CS-111 | Blanket `page.RouteAsync("**")` mock | warning |
271
+ | QA-CS-108 | Hardcoded environment URL in test | info |
272
+ | QA-CS-111 | Blanket `page.RouteAsync("**")` mock | info |
313
273
 
314
274
  </details>
315
275
 
316
- **Rule tiers.** Every rule is `core`, `extended`, or `quarantine`, assigned from
317
- its **measured** false-positive rate ([docs/FP-AUDIT.md](docs/FP-AUDIT.md)):
318
- `core` ≤ 10 % FP, `extended` ≤ 30 %, `quarantine` above that or not yet measured.
319
- The default scan runs core + extended; `--strict` adds quarantine. The Java and
320
- C# families are newer and mostly `extended`/`quarantine` today — they ship,
321
- they're documented, and they stay out of the headline number until a real
322
- consumer suite (not a binding-library's own tests) has been audited.
323
-
324
- > The full live catalog — every rule with tier, confidence, false-positive risk,
325
- > and autofix availability — is generated from the registry:
276
+ > The full live catalog — every rule with tier, confidence, false-positive
277
+ > risk, and autofix availability — is generated from the registry:
326
278
  >
327
279
  > ```bash
328
280
  > mjolnir rules --md
329
281
  > ```
282
+ >
283
+ > Per-rule pages live under [`docs/rules/`](docs/rules/).
284
+
285
+ ### How much of this is measured
286
+
287
+ **74 of 99 rules carry a false-positive rate measured against real OSS code** (≥ 10 hand-classified findings each; see
288
+ [docs/FP-AUDIT.md](docs/FP-AUDIT.md)). The other 19 ship on the author's
289
+ estimate. Every scan footer tells you how many of the rules that _fired_
290
+ are measured; `mjolnir rules --unmeasured` lists the ones that aren't;
291
+ every rule's `mjolnir explain` page states its status. We publish the rate
292
+ even when it's ugly — QA-CS-103 audits at 95% and is quarantined for it.
293
+ Growing that 78 is the project's continuing work.
294
+
295
+ ### Rule tiers and language maturity
296
+
297
+ Every rule is `core`, `extended`, or `quarantine`, assigned from its
298
+ **measured** false-positive rate:
299
+
300
+ | Tier | Meaning | Default scan | `--strict` |
301
+ | ------------ | ---------------------------------------- | :----------: | :--------: |
302
+ | `core` | ≤ 10 % measured FP | ✅ | ✅ |
303
+ | `extended` | ≤ 30 % measured FP | ✅ | ✅ |
304
+ | `quarantine` | above 30 %, or not yet measured (n < 10) | ❌ | ✅ |
305
+
306
+ | Language | Adapter | Coverage today |
307
+ | --------------- | ------------ | -------------------------------------------------- |
308
+ | TypeScript / JS | compiler AST | broadest, most measured — mostly `core`/`extended` |
309
+ | Python / pytest | regex layer | broad, corpus-audited — mostly `core`/`extended` |
310
+ | Java | regex layer | newer — mostly `extended`/`quarantine` |
311
+ | C# / .NET | regex layer | newer — mostly `extended`/`quarantine` |
312
+
313
+ TypeScript and Python have the broadest measured coverage. Java and C# ship,
314
+ are documented, and stay out of the headline number until a real consumer
315
+ suite (not a binding library's own tests) has been audited.
316
+
317
+ ---
318
+
319
+ ## How the score works
320
+
321
+ <p align="center">
322
+ <img src="assets/readme/terminal-hero.svg" alt="Mjölnir terminal output — WORTHINESS 75/100 NEEDS WORK, a diagnostics-by-category breakdown, and a FIX THIS FIRST list" width="820" />
323
+ </p>
324
+
325
+ <sub>Regenerated by `npm run docs:hero`;
326
+ [`tests/hero-asset-reproducibility.spec.ts`](tests/hero-asset-reproducibility.spec.ts)
327
+ fails CI if it drifts from what the reporter actually prints.</sub>
328
+
329
+ The score is transparent: **error −8, warning −3, info −1**, then normalized
330
+ by suite exposure (deductions per test declaration). Evidence-weighted
331
+ deductions mean weak signals cost less. The terminal shows the same
332
+ discounted numbers the score uses — no black box. Full method:
333
+ [docs/SCORING.md](docs/SCORING.md).
334
+
335
+ **Verdicts**
336
+
337
+ | Score | Verdict |
338
+ | ------- | ---------------- |
339
+ | ≥ 80 | ✓ **WORTHY** |
340
+ | 50 – 79 | ⚠ **NEEDS WORK** |
341
+ | < 50 | ✖ **UNWORTHY** |
342
+
343
+ **Evidence levels** — every finding carries one; it sets the finding's
344
+ weight in the score:
345
+
346
+ | Level | Meaning | Score impact | Example |
347
+ | ----- | -------------------- | ---------------- | -------------------------------------------------- |
348
+ | E2 | Deterministic defect | Full deduction | `.only` committed — structurally provable |
349
+ | E1 | Heuristic pattern | Half deduction | Regex-matched `sleep()` — strong signal, not proof |
350
+ | E0 | Observation | Zero (info only) | Reported but never gates CI or deducts |
351
+
352
+ Most rules are **E1**. The tagline "we prove it" refers to this system:
353
+ E2 findings are structural proof; E1 findings are correctly-positioned
354
+ warnings, not formal proofs.
355
+
356
+ An empty repo scores `null`, never a fake 100 — see [Trust model](#trust-model).
330
357
 
331
358
  ---
332
359
 
@@ -347,7 +374,7 @@ regressed.
347
374
 
348
375
  ---
349
376
 
350
- ## 🔬 Runtime Evidence
377
+ ## 🔬 Runtime evidence
351
378
 
352
379
  Static flakiness detection is guessing. Mjölnir reads **real execution
353
380
  data** — Playwright JSON reports and JUnit XML from any runner:
@@ -372,7 +399,61 @@ test. It gets flagged `TRUE-FLAKE` regardless of the final green checkmark.
372
399
 
373
400
  ---
374
401
 
375
- ## 🤖 CI Integration
402
+ ## ⚡ Mjölnir is not another linter
403
+
404
+ Linters tell you whether code follows rules. Mjölnir tells you whether your
405
+ verification can be trusted.
406
+
407
+ | | ESLint / SonarQube | Coverage tools | Manual review | **Mjölnir** |
408
+ | -------------------------------------------------------- | :----------------: | :------------: | :-----------: | :---------: |
409
+ | CI workflow integrity (`continue-on-error`, `\|\| true`) | ❌ | ❌ | rarely | ✅ |
410
+ | Cross-language (TS, Python, Java, C#) from one tool | ❌ | ❌ | ❌ | ✅ |
411
+ | Grades Playwright locator resilience (Selector Health) | ❌ | ❌ | rarely | ✅ |
412
+ | Flags tests with no real assertions | ✅ (plugin)\* | ❌ | sometimes | ✅ |
413
+ | Catches hard sleeps (`waitForTimeout`, `time.sleep`) | ✅ (plugin)\* | ❌ | sometimes | ✅ |
414
+ | Runs in seconds, zero network calls while scanning | ✅ | ✅ | — | ✅ |
415
+
416
+ \*`eslint-plugin-jest` (`expect-expect`) and `eslint-plugin-playwright`
417
+ (`expect-expect`, `no-wait-for-timeout`) cover these for their respective
418
+ frameworks.
419
+
420
+ **Runtime analysis** is a separate category from static linting:
421
+
422
+ | | Playwright retry reporter | Allure / ReportPortal | **Mjölnir forensics** |
423
+ | --------------------------------------------- | :-----------------------: | :-------------------: | :-------------------: |
424
+ | Reads real run data for `TRUE-FLAKE` verdicts | partial\* | partial (tag) | ✅ |
425
+ | Flaky-triage report from execution history | ❌ | ✅ | ✅ |
426
+ | Integrates with static worthiness score | ❌ | ❌ | ✅ |
427
+
428
+ \*Playwright tracks retries internally but does not produce a standalone
429
+ flakiness report with verdict labels.
430
+
431
+ ---
432
+
433
+ ## 🤖 Why not just use AI code review?
434
+
435
+ Different problem, different layer. AI review can spot a suspicious test
436
+ change in a diff; it does not prove the verification system as a whole is
437
+ trustworthy — and it only sees the diff you show it.
438
+
439
+ | | AI code review (Copilot, etc.) | **Mjölnir** |
440
+ | ------------------------------------- | :----------------------------: | :-------------------------: |
441
+ | Cost per scan | Tokens (scales with diff size) | **Zero** (local, installed) |
442
+ | Sees the whole suite + all CI configs | Only the PR diff you show it | **Everything, every time** |
443
+ | Deterministic (same input → same out) | ❌ (non-deterministic) | **✅** |
444
+ | Catches patterns dormant for months | Only if it's in the context | **✅** (scans all files) |
445
+ | Remembers findings between runs | ❌ (no memory across sessions) | **✅** (baseline + diff) |
446
+ | Runs without human triggering | Needs a PR or prompt | **✅** (CI hook, 3 seconds) |
447
+
448
+ **Use both.** AI catches nuance, intent, and design flaws no regex can
449
+ find. Mjölnir catches the structural patterns AI overlooks because they
450
+ look "intentional" — a committed `.only`, a swallowed exit code, a
451
+ `continue-on-error` on a test job. Those aren't bugs that need reasoning;
452
+ they're facts that need scanning.
453
+
454
+ ---
455
+
456
+ ## 🤖 CI integration
376
457
 
377
458
  One command generates a PR workflow — advisory by default, never blocking:
378
459
 
@@ -389,54 +470,108 @@ Or wire it into GitHub Code Scanning natively via SARIF:
389
470
  sarif_file: mjolnir.sarif
390
471
  ```
391
472
 
473
+ Editor and pipeline setup for SARIF: [docs/SARIF-INTEGRATION.md](docs/SARIF-INTEGRATION.md).
474
+
475
+ ### Changed-scope coverage
476
+
477
+ `--scope changed` attributes findings to lines added in your branch vs the
478
+ merge-base with `main`. It covers test files (`*.spec.*`, `*.test.*`) plus
479
+ GitHub workflow files and Playwright configs in the diff. When the
480
+ merge-base can't be resolved — shallow clone, detached HEAD, non-git
481
+ target, different default branch — it degrades honestly: findings fall
482
+ back to full-file attribution and the report says so. Override the base
483
+ ref with `--base <ref>`.
484
+
392
485
  ---
393
486
 
394
- ## 🛡️ Trust Model
487
+ ## Configuration
488
+
489
+ Mjölnir is zero-config. An optional `mjolnir.config.json` (or
490
+ `.mjolnir.json`) at the repo root tunes severity, gating and scope — it
491
+ never changes detection semantics.
492
+
493
+ | Key | Type | Effect |
494
+ | ------------------- | ------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
495
+ | `exclude` | `string[]` | Extra ignore globs (gitignore subset), on top of the built-in defaults |
496
+ | `gate` | `"advisory" \| "error" \| "warning"` | Which severities exit non-zero (default `error`; `advisory` never blocks) |
497
+ | `severityOverrides` | `{ "<RULE-ID>": severity }` | Re-rank a rule's findings for your repo |
498
+ | `ignore` | `IgnoreEntry[]` | Suppress findings — **`reason` is required**; entries expire after 90 days (an explicit `expires` date, or the config file's last-modified time for entries without one) |
499
+ | `plugins` | `string[]` | Third-party rule packages (see [Trust model](#trust-model)) |
500
+
501
+ ```json
502
+ {
503
+ "gate": "error",
504
+ "exclude": ["legacy/**"],
505
+ "severityOverrides": { "QA-PW-118": "warning" },
506
+ "ignore": [
507
+ {
508
+ "ruleId": "QA-TEST-004",
509
+ "files": ["e2e/legacy-login.spec.ts"],
510
+ "reason": "Third-party widget needs a settle delay; tracked in JIRA-4821",
511
+ "expires": "2026-12-31"
512
+ }
513
+ ]
514
+ }
515
+ ```
395
516
 
396
- - **Local-first** — zero network calls during scanning. Ever.
397
- - **No false proof** — we'd rather say "unknown" than "verified". An empty
398
- repo gets `score: null`, never a fake 100.
399
- - **Transparent scoring** — public deduction constants: error −8, warning −3,
400
- info −1. No black box. Verdict thresholds: **≥80 WORTHY**, **50–79 NEEDS
401
- WORK**, **<50 UNWORTHY**.
402
- - **Evidence-weighted deductions** — every finding carries an evidence level:
403
- **E2** (deterministic defect, full deduction), **E1** (heuristic pattern,
404
- half deduction), **E0** (observation — reported but costs nothing and never
405
- gates CI).
406
- - **Partial honesty** — if analysis was cut short, the output says so.
407
- Never "complete" when it isn't.
408
- - **QA-native language** — findings speak your vocabulary:
409
- `FALSE-GREEN`, `FLAKY-RISK`, `BLOCKS-RELEASE`, `HYGIENE`.
410
- - **FP firewall** — detection runs on a comment/string-free view of the code
411
- (TypeScript rules use the compiler AST): a pattern inside a prose comment
412
- or a doc-example string is documentation, not a finding.
413
- - **Measured, not asserted** — 15 of 91 rules carry a false-positive rate from
414
- real OSS code; the scan footer, `mjolnir rules --unmeasured`, and every
415
- `mjolnir explain` page tell you which rules are which.
416
- [docs/FP-AUDIT.md](docs/FP-AUDIT.md) has the numbers, ugly ones included.
517
+ - **`.mjolnirignore`** — a plain gitignore-style file for path exclusions,
518
+ same dialect as `exclude`. Use it for machine-wide noise; use `exclude`
519
+ when the list belongs in version control alongside the rest of the config.
520
+ - **CLI overrides** — `--strict` (include quarantine rules), `--width <cols>`
521
+ and `--ascii` / `--no-ascii` (terminal rendering), `--tone blunt`
522
+ (blunter messages), `--max-duration <sec>` (bounded partial scan).
523
+ - Rule suppression and deprecation lifecycle: [docs/RULE-LIFECYCLE.md](docs/RULE-LIFECYCLE.md).
524
+
525
+ `ignore` entries also power the standalone `mjolnir suppressions` command,
526
+ which lists what's currently suppressed and when each entry expires.
417
527
 
418
528
  ---
419
529
 
420
- ## 🔌 Plugin Trust Model
530
+ ## 📐 Exit codes & contracts
421
531
 
422
- Plugins are npm packages you declare in `mjolnir.config.json`. There is **no
423
- sandbox**: plugin code executes with full Node privileges against the scanned
424
- tree — the same trust model as ESLint or Vitest plugins. Only install plugins
425
- you'd trust as devDependencies. Core rule-ID prefixes (`QA-TEST`, `QA-TQUAL`,
426
- `QA-PW`, `QA-CI`, `QA-PY`, `QA-ENV`, `QA-JV`, `QA-CS`) are reserved and
427
- rejected from plugins to prevent spoofing.
532
+ Frozen — safe to build CI logic on:
533
+
534
+ | Exit code | Meaning |
535
+ | --------- | --------------------------------------------------------------- |
536
+ | `0` | Clean — no findings at or above the gate |
537
+ | `1` | Findings at or above the gate |
538
+ | `2` | Partial scan (time budget hit, unreadable files) — never blocks |
539
+ | `10` | Usage error (bad flag, missing target) |
540
+ | `20` | Internal error |
541
+
542
+ The JSON/SARIF report is `schemaVersion: 1`. Rule IDs (`QA-<FAMILY>-NNN`)
543
+ are immutable once shipped and never reused.
428
544
 
429
545
  ---
430
546
 
431
- ## 🔍 Changed-Scope Coverage
547
+ ## Trust model
432
548
 
433
- `--scope changed` attributes findings to lines added in your branch vs the
434
- merge-base with `main` (the base branch is `main`; a repo that uses `master`
435
- or a shallow clone with no merge-base degrades — see below). It covers test
436
- files (`*.spec.*`, `*.test.*`) plus GitHub workflow files and Playwright
437
- configs changed in the diff. When the merge-base can't be resolved — shallow
438
- clone, detached HEAD, non-git target, different default branch — it degrades
439
- honestly: findings fall back to full-file attribution and the report says so.
549
+ - **Local-first** — zero network calls during scanning. Ever. Zero telemetry.
550
+ - **No false proof** — we'd rather say "unknown" than "verified". An empty
551
+ repo gets `score: null`, never a fake 100.
552
+ - **Partial honesty** — if analysis was cut short, the output says so.
553
+ Never "complete" when it isn't.
554
+ - **FP firewall** — detection runs on a comment/string-free view of the code
555
+ (TypeScript rules use the compiler AST): a pattern inside a prose comment
556
+ or a doc-example string is documentation, not a finding.
557
+ - **Measured, not asserted** — only rules with a false-positive rate from
558
+ real OSS code ship in the headline tiers (see
559
+ [How much of this is measured](#how-much-of-this-is-measured)); the scan
560
+ footer and `mjolnir rules --unmeasured` tell you which is which.
561
+ - **Plugin trust** — plugins are npm packages declared under `"plugins"`.
562
+ There is **no sandbox**: plugin code runs with full Node privileges, the
563
+ same trust model as ESLint or Vitest plugins. Core rule-ID prefixes are
564
+ reserved and rejected from plugins to prevent spoofing.
565
+ - **Workspace-local external rules** (folder-based, zero network) — a
566
+ `mjolnir-rules/` directory next to the scan target loads custom rules:
567
+ JSON files declare regex patterns (no code executed), `.mjs`/`.js`
568
+ modules export `rules` (full-Node trust, same as plugins). External
569
+ rules carry the same trust metadata as core; they can never ship in
570
+ the core tier (core requires a measured FP rate from the corpus
571
+ sidecar — a declared `tier: "core"` is clamped to `extended`), obey
572
+ tier caps, and are drift-checked: `mjolnir rules --md --external`
573
+ renders the catalog from the loaded files (provenance `external`),
574
+ and the matrix generator accepts `--external <root>`.
440
575
 
441
576
  ---
442
577
 
@@ -450,15 +585,16 @@ mjolnir/
450
585
  ├── src/
451
586
  │ ├── engine/ # LanguageAdapter interface + rule runner
452
587
  │ ├── adapters/ # typescript · python · java · csharp · github-actions
453
- │ ├── rules/ # deterministic rules across 8 families
588
+ │ ├── rules/ # rules across 8 families + the measured-FP table
454
589
  │ ├── playwright/ # Selector Health Score engine
455
- │ ├── discovery/ # workspace, frameworks, safe-YAML parser
590
+ │ ├── discovery/ # workspace, frameworks, ignore resolution
456
591
  │ ├── scope/ # git merge-base changed-scope engine
457
- │ ├── scorer/ # transparent deduction table
592
+ │ ├── scorer/ # transparent deduction table + prioritization
458
593
  │ ├── reporter/ # terminal · JSON · SARIF 2.1 · Mermaid
459
594
  │ ├── forensics/ # run-data ingestion · flake verdicts · triage
460
- │ ├── commands/ # badge · baseline+diff · create-rule · debt · doctor · explain · fix · handover · impact · init · pr-comment · pw-report · rules · stats
461
- │ └── integrations/ # CI workflow generator
595
+ │ ├── config/ # mjolnir.config.json + suppressions
596
+ │ ├── plugins/ # third-party rule loading (no sandbox)
597
+ │ └── commands/ # every subcommand
462
598
  └── tests/
463
599
  ├── fixtures/ # must-fire / must-not-fire per rule
464
600
  └── golden/ # frozen score regression locks
@@ -466,60 +602,60 @@ mjolnir/
466
602
 
467
603
  </details>
468
604
 
469
- Multi-language by design: adding an ecosystem (Ruby/RSpec, Go, PHPUnit…) =
470
- one adapter + its rules. Playwright itself has exactly four official
471
- language bindings — TypeScript, Python, Java, .NET — and all four have an
472
- adapter here.
473
- TypeScript/Playwright uses the compiler AST (ts-morph); Python, Java, and
474
- C#/.NET run on a shared comment/string-masked regex layer. A tree-sitter
475
- WASM AST layer for Java and C# exists (`src/engine/tree-sitter-ast.ts`) and
476
- is the next precision step — it is not yet wired into the synchronous scan
477
- pipeline. TypeScript/Playwright and Python are the most battle-tested by
478
- measured false-positive rate; Java and C# are newer (see **Rule tiers** above).
605
+ - **Rules are pure functions** — `(SourceFileContext) → Finding[]`, no I/O,
606
+ no globals. Adding an ecosystem = one adapter + its rules.
607
+ - **TypeScript/Playwright uses the compiler AST** (ts-morph). Python, Java
608
+ and C# run on a shared comment/string-masked regex layer.
609
+ - A tree-sitter WASM AST layer for Java and C# exists and is the next
610
+ precision step — it is not yet wired into the synchronous scan pipeline.
479
611
 
480
612
  ---
481
613
 
482
- ## 📦 Install
614
+ ## 📚 Documentation
615
+
616
+ | Document | What's in it |
617
+ | ------------------------------------------------------ | ------------------------------------------------- |
618
+ | [docs/SCORING.md](docs/SCORING.md) | Score normalization + evidence weighting |
619
+ | [docs/FP-AUDIT.md](docs/FP-AUDIT.md) | Measured false-positive rates + method |
620
+ | [docs/RULE-LIFECYCLE.md](docs/RULE-LIFECYCLE.md) | Rule states, suppression, deprecation |
621
+ | [docs/VERSIONING.md](docs/VERSIONING.md) | Semver policy, frozen surfaces, deprecation cycle |
622
+ | [docs/SARIF-INTEGRATION.md](docs/SARIF-INTEGRATION.md) | SARIF output + editor/CI setup |
623
+ | [docs/rules/](docs/rules/) | Generated per-rule catalog |
624
+ | [CONTRIBUTING.md](CONTRIBUTING.md) | Dev setup + contribution workflow |
625
+ | [SUPPORT.md](SUPPORT.md) | Where to ask, report and get help |
626
+ | [CHANGELOG.md](CHANGELOG.md) | Release history |
627
+ | [SECURITY.md](SECURITY.md) | Vulnerability reporting |
483
628
 
484
- ```bash
485
- # Run directly (recommended):
486
- npx mjolnir-qa@latest
629
+ ---
487
630
 
488
- # Or install globally:
489
- npm i -g mjolnir-qa
490
- ```
631
+ ## 📈 Status
491
632
 
492
- Requires Node.js ≥ 22.18 (required by the tsdown bundler for native ESM support). Works on Windows, macOS, and Linux.
633
+ **v0.5.x · open beta.** The JSON schema and exit codes are frozen contracts.
634
+ TypeScript and Python have the broadest measured coverage; Java and C# are
635
+ newer — read them through the [maturity table](#rule-tiers-and-language-maturity).
636
+ Honest scope, no invented dates: the [public roadmap](https://sergey-bar.github.io/Mjolnir/reference/roadmap).
493
637
 
494
638
  ---
495
639
 
496
640
  ## 🤝 Contributing
497
641
 
498
- New rules are the easiest first contribution — one command scaffolds
499
- everything (anti-creep law enforced by the fixture harness):
642
+ New rules are the easiest first contribution — one command scaffolds the
643
+ rule plus its must-fire **and** must-not-fire fixtures (the generated rule
644
+ intentionally fails its fixtures until you implement real detection — a stub
645
+ cannot ship):
500
646
 
501
647
  ```bash
502
648
  mjolnir create-rule QA-PW-140 --title "Screenshot without diff bound"
503
649
  ```
504
650
 
505
- That generates the rule file plus must-fire AND must-not-fire fixture
506
- skeletons. The generated rule intentionally fails its fixtures until you
507
- implement real detection — a stub cannot ship.
508
-
509
- Manual path:
510
-
511
- 1. Fork + clone, `npm install`
512
- 2. Copy any rule folder under `src/rules/` as a template
513
- 3. Add must-fire AND must-not-fire fixtures under `tests/fixtures/<YOUR-ID>/`
514
- 4. `npx vitest run` — both directions green? Ship the PR.
515
-
516
- See [SECURITY.md](SECURITY.md) for vulnerability reporting.
651
+ Full dev setup, the standing-gate commands, and the anti-creep / fixture
652
+ firewall laws are in [CONTRIBUTING.md](CONTRIBUTING.md).
517
653
 
518
654
  ---
519
655
 
520
656
  <div align="center">
521
657
 
522
- **Are your tests worthy of trust?**
658
+ **Stop shipping tests you can't trust.**
523
659
 
524
660
  ```bash
525
661
  npx mjolnir-qa@latest