mjolnir-qa 0.4.0 → 0.5.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,23 +1,18 @@
1
1
  <div align="center">
2
2
 
3
- <pre>
4
- ╔═══════════╗
5
- ║ ║
6
- ╠═══════════╣ M J Ö L N I R
7
- ║ ║ ║
8
- ╚═════╩═════╝ VERIFICATION TRUST ENGINE
9
- ║
10
- ║
11
- </pre>
3
+ <img src="assets/readme/logo.png" alt="Mjölnir — Verification Trust Engine" width="800" />
12
4
 
13
5
  ### Your tests are lying to you. We prove it.
14
6
 
15
- **Verification Trust Engine for QA.** Audits test suites and CI pipelines,
16
- reports a worthiness score, and shows exactly where trust breaks.
7
+ **Verification Trust Engine for QA.** Mjölnir audits test suites and CI
8
+ pipelines, reports a worthiness score, and shows exactly where trust breaks.
17
9
 
18
- [![license](https://img.shields.io/badge/license-MIT-B45309.svg?style=flat-square&labelColor=0D0D0D)](LICENSE)
19
- [![node](https://img.shields.io/badge/node-%E2%89%A5%2022.18-146A8C.svg?style=flat-square&labelColor=0D0D0D)](https://nodejs.org)
20
- [![status](https://img.shields.io/badge/status-●_ONLINE-15803D.svg?style=flat-square&labelColor=0D0D0D)](#-quickstart)
10
+ [![npm](https://img.shields.io/npm/v/mjolnir-qa.svg?style=flat-square&color=C9A227&labelColor=0B0F17)](https://www.npmjs.com/package/mjolnir-qa)
11
+ [![ci](https://img.shields.io/github/actions/workflow/status/Sergey-Bar/Mjolnir/ci.yml?branch=main&style=flat-square&label=ci&labelColor=0B0F17)](https://github.com/Sergey-Bar/Mjolnir/actions/workflows/ci.yml)
12
+ [![license](https://img.shields.io/badge/license-MIT-C9A227.svg?style=flat-square&labelColor=0B0F17)](LICENSE)
13
+ [![node](https://img.shields.io/badge/node-%E2%89%A5%2022.18-2E8C7F.svg?style=flat-square&labelColor=0B0F17)](https://nodejs.org)
14
+
15
+ English | [简体中文](README.zh.md) | [繁體中文](README.zht.md) | [한국어](README.ko.md) | [Deutsch](README.de.md) | [Español](README.es.md) | [Français](README.fr.md) | [Italiano](README.it.md) | [Dansk](README.da.md) | [日本語](README.ja.md) | [Polski](README.pl.md) | [Русский](README.ru.md) | [Norsk](README.no.md) | [Português (Brasil)](README.br.md) | [ไทย](README.th.md) | [Türkçe](README.tr.md) | [Українська](README.uk.md) | [বাংলা](README.bn.md) | [Ελληνικά](README.gr.md) | [Tiếng Việt](README.vi.md) | [עברית](README.he.md) | [العربية](README.ar.md) | [Bosanski](README.bs.md)
21
16
 
22
17
  ```bash
23
18
  npx mjolnir-qa@latest
@@ -25,104 +20,154 @@ npx mjolnir-qa@latest
25
20
 
26
21
  **Are your tests worthy of trust?**
27
22
 
28
- [Quickstart](#-quickstart) · [Rules](#-the-rules) · [Selector Health](#-selector-health-score) · [Runtime Evidence](#-runtime-evidence) · [CI Integration](#-ci-integration) · [Contributing](#-contributing)
23
+ [See it work](#-see-it-work) · [Quickstart](#-quickstart) · [What it checks](#-what-mjölnir-checks) · [Scoring](#how-the-score-works) · [CI](#-ci-integration) · [Configuration](#configuration) · [Docs](#-documentation)
29
24
 
30
25
  </div>
31
26
 
32
27
  ---
33
28
 
34
- ## 🔨 What is Mjölnir?
29
+ ## 🎬 See it work
35
30
 
36
- | | |
37
- | --- | -------------------------------------------------------------------------------------------------------------- |
38
- | ⚖️ | **Worthiness Score** — one number, transparent deduction table, no black box |
39
- | 🎭 | **Selector Health Score** — grades your Playwright locators, not just your pass rate |
40
- | 🔬 | **Runtime forensics** — reads real Playwright/JUnit run data to catch `TRUE-FLAKE`, not just static guesses |
41
- | 🚨 | **CI-integrity rules** — catches `continue-on-error`, `\|\| true`, and other false-green tricks |
42
- | 🐍 | **Multi-language** — TypeScript/Playwright, Python/pytest, Java, and C#/.NET today, one adapter away from more |
43
- | 🔒 | **Local-first** — zero network calls while scanning, zero telemetry, runs in seconds |
31
+ <p align="center">
32
+ <img src="assets/readme/demo.svg" alt="Mjölnir's full --verbose report on a demo repo: WORTHINESS 75/100 NEEDS WORK, a diagnostics-by-category breakdown, a FIX THIS FIRST list, and every finding with its rule ID and line number across CI, Playwright, test-hygiene and Python rules" width="900" />
33
+ </p>
44
34
 
45
- ---
35
+ <sub>The complete `npx mjolnir-qa ./examples/demo-repo --verbose` output,
36
+ rendered from the actual reporter — nothing trimmed. Regenerated by
37
+ `npm run docs:demo`;
38
+ [`tests/demo-asset-reproducibility.spec.ts`](tests/demo-asset-reproducibility.spec.ts)
39
+ fails CI if it drifts from what the tool prints.</sub>
46
40
 
47
- ## ⚡ Mjölnir is not another linter
41
+ **What just happened:**
42
+
43
+ 1. Mjölnir discovered the Playwright specs, its config, the CI workflow and
44
+ a Python test file — four languages/formats, one pass.
45
+ 2. It found evidence that weakens trust in the suite — a `continue-on-error`
46
+ masking a job, a `|| true` swallowing an exit code, hard sleeps, a
47
+ brittle selector, hardcoded staging URLs, a `networkidle` wait.
48
+ 3. It turned each into a concrete finding with a rule ID, a location and a
49
+ fix — and a single score you can gate a PR on.
50
+
51
+ ### One finding, up close
52
+
53
+ Run `mjolnir explain QA-CI-001` on the first finding above and you get:
54
+
55
+ ```text
56
+ ▚▞ QA-CI-001 — continue-on-error masks a failing verification gate
48
57
 
49
- Linters tell you whether code follows rules.
50
- Mjölnir tells you whether your verification can be trusted.
58
+ Severity: error
59
+ Confidence: high
60
+ Evidence: E2
61
+ Measured FP: not yet measured — this rule ships on assumption (see docs/FP-AUDIT.md)
51
62
 
52
- | | ESLint / SonarQube | Coverage tools | Manual review | **Mjölnir** |
53
- | ---------------------------------------------------------- | :----------------: | :------------: | :-----------: | :---------: |
54
- | Catches syntax & style bugs | ✅ | ❌ | ✅ (slow) | — |
55
- | Flags tests with no real assertions | ❌ | ❌ | sometimes | ✅ |
56
- | Catches CI false-greens (`\|\| true`, `continue-on-error`) | ❌ | ❌ | rarely | ✅ |
57
- | Reads **real** run data for `TRUE-FLAKE` verdicts | ❌ | ❌ | ❌ | ✅ |
58
- | Grades Playwright locator resilience | ❌ | ❌ | rarely | ✅ |
59
- | Runs in seconds, zero network calls | ✅ | ✅ | — | ✅ |
63
+ WHAT WAS FOUND (real detector output, not a mockup)
64
+ Job `security-scan` runs a verification gate under `continue-on-error: true`.
65
+
66
+ WHY IT MATTERS
67
+ This job can fail every day and CI will still show green. The checkmark
68
+ on this workflow cannot be trusted.
69
+
70
+ HOW TO FIX
71
+ Remove continue-on-error, or scope it to individual non-blocking steps only.
72
+ ```
73
+
74
+ That is the unit of value: not a style nit, but a place where your CI is
75
+ telling you something passed when it didn't.
60
76
 
61
77
  ---
62
78
 
63
79
  ## ⚡ Quickstart
64
80
 
81
+ Run it against a repo for a full report and a worthiness score:
82
+
65
83
  ```bash
66
84
  npx mjolnir-qa@latest
67
85
  ```
68
86
 
69
- That's it. Zero configuration — Mjölnir detects your frameworks, finds your
70
- tests, and reports.
71
-
72
- | Command | What it does |
73
- | ------------------------------------------------------ | ------------------------------------------------------------------- |
74
- | `npx mjolnir-qa@latest --scope changed` | Only findings introduced by **your** changes — perfect for PRs |
75
- | `npx mjolnir-qa@latest --json` | Machine-readable output |
76
- | `npx mjolnir-qa@latest --format sarif > mjolnir.sarif` | GitHub Code Scanning integration |
77
- | `npx mjolnir-qa@latest --format mermaid` | Test-architecture diagram — paste into a GitHub comment or a slide |
78
- | `mjolnir doctor:playwright` | Playwright-only deep scan + Selector Health Score |
79
- | `mjolnir forensics ./test-results/` | Runtime evidence — retries, true flakes, `FLAKY.md` artifact |
80
- | `mjolnir triage ./test-results/` | The flaky-triage meeting, in 10 minutes instead of 45 |
81
- | `mjolnir fix --dry-run` / `fix` | Safe auto-fixes with proof (dry-run first) |
82
- | `mjolnir debt` | Test debt register — presentable to management |
83
- | `mjolnir handover` | New-QA-onboarding map of the suite |
84
- | `mjolnir pw-report ./test-results/` | Playwright run summary — retries / flakes / slowest |
85
- | `mjolnir badge` | Evidentiary badge (shields.io endpoint JSON) |
86
- | `mjolnir doctor` | Self-audit — prove Mjölnir's own rule base is worthy |
87
- | `mjolnir rules` / `rules --md` | Rule catalog with trust metadata (JSON or markdown) |
88
- | `mjolnir explain <RULE-ID>` | What/why/fix for one rule, with a real example from its own fixture |
89
- | `mjolnir impact [--since <ref>]` | What changed since a prior commit — fixes and new debt |
87
+ **In CI, the product is one command.** It scans only what the branch
88
+ touched and exits non-zero on new problems:
90
89
 
91
- ---
90
+ ```bash
91
+ npx mjolnir-qa@latest --scope changed
92
+ ```
92
93
 
93
- ## ⚖️ Worthiness Score
94
+ Drop that in a PR check — `mjolnir ci install` writes the workflow —
95
+ and you're done. Everything else is optional.
94
96
 
95
- ```text
96
- 🔨 MJÖLNIR
97
+ | Command | What it does |
98
+ | ----------------------------------- | ------------------------------------------------ |
99
+ | `mjolnir` | Full-repo scan + worthiness score |
100
+ | `mjolnir --scope changed` | Only what your branch introduced — the CI form |
101
+ | `mjolnir ci install` | Generate the advisory PR workflow |
102
+ | `mjolnir explain QA-CI-001` | What / why / fix + measured FP rate for one rule |
103
+ | `mjolnir rules --unmeasured` | The rules running on assumption, not measurement |
104
+ | `mjolnir --json` / `--format sarif` | Machine-readable / GitHub Code Scanning |
105
+ | `mjolnir --strict` | Also run quarantine-tier rules (higher FP risk) |
106
+ | `mjolnir --cache` | Incremental re-scans via a local verdict cache |
97
107
 
98
- WORTHINESS 80/100 — WORTHY
99
- ████████████████████████░░░░░░
108
+ <details>
109
+ <summary><strong>When something's flaky</strong></summary>
100
110
 
101
- DETECTED [playwright]
111
+ | Command | What it does |
112
+ | ----------------------------------- | --------------------------------------------------- |
113
+ | `mjolnir forensics ./test-results/` | Real run data → `TRUE-FLAKE` verdicts, `FLAKY.md` |
114
+ | `mjolnir triage ./test-results/` | Quarantine proposal from execution history |
115
+ | `mjolnir pw-report ./test-results/` | Playwright run summary — retries / flakes / slowest |
116
+ | `mjolnir doctor:playwright` | Playwright-only deep scan + Selector Health Score |
102
117
 
103
- ▚ FIX THIS FIRST
104
- +8 pts QA-CI-001 · .github/workflows/ci.yml:48
105
- +3 pts QA-TEST-004 · e2e/checkout.spec.ts:6
106
- ```
118
+ </details>
107
119
 
108
- The score is transparent: error −8, warning −3, info −1. Evidence-weighted
109
- deductions mean weak signals cost less. The terminal shows the same discounted
110
- numbers the score uses — no black box.
120
+ <details>
121
+ <summary><strong>Occasional / reporting</strong></summary>
122
+
123
+ | Command | What it does |
124
+ | ------------------------------- | ------------------------------------------------ |
125
+ | `mjolnir fix --dry-run` / `fix` | Safe auto-fixes with proof |
126
+ | `mjolnir baseline` / `diff` | Snapshot findings, then report only new/worsened |
127
+ | `mjolnir impact --since <ref>` | What changed since a prior commit |
128
+ | `mjolnir debt` | Test-debt register with a cost model |
129
+ | `mjolnir handover` | New-QA onboarding map of the suite |
130
+ | `mjolnir stats` | Local all-time counters of fixes seen |
131
+ | `mjolnir badge` | shields.io endpoint JSON + snippet |
132
+ | `mjolnir rules --md` | Full rule catalog (JSON or Markdown) |
133
+ | `mjolnir doctor` | Self-audit of Mjölnir's own rule base |
134
+ | `mjolnir create-rule <ID>` | Scaffold a new rule + fixtures |
135
+ | `mjolnir --format mermaid` | Test-architecture diagram for a PR comment |
111
136
 
112
- **Verdicts:**
137
+ </details>
113
138
 
114
- | Score | Verdict |
115
- | ------- | ---------------- |
116
- | ≥ 80 | ✓ **WORTHY** |
117
- | 50 – 79 | ⚠ **NEEDS WORK** |
118
- | < 50 | ✖ **UNWORTHY** |
139
+ Install globally instead of `npx` if you prefer: `npm i -g mjolnir-qa`.
140
+ Requires Node.js ≥ 22.18. Works on Windows, macOS, and Linux.
119
141
 
120
142
  ---
121
143
 
122
- ## 📋 The Rules
144
+ ## 👥 Who is this for?
145
+
146
+ - **QA / SDET** owning an e2e or integration suite who need evidence the
147
+ suite actually deserves the green checkmark it produces.
148
+ - **Platform / DevEx** teams responsible for CI integrity and release
149
+ gates — the people who care that a `continue-on-error` never silently
150
+ turns a red pipeline green.
151
+ - **OSS maintainers** who want a cheap, always-on verification gate that
152
+ runs locally and in CI with zero network calls.
153
+
154
+ ---
155
+
156
+ ## 🔨 What Mjölnir checks
157
+
158
+ | | |
159
+ | --- | ----------------------------------------------------------------------------------------------------------------- |
160
+ | ⚖️ | **Worthiness Score** — one number, transparent deduction table, no black box |
161
+ | 🎭 | **Selector Health Score** — grades your Playwright locators, not just your pass rate |
162
+ | 🔬 | **Runtime forensics** — reads real Playwright/JUnit run data to catch `TRUE-FLAKE`, not just static guesses |
163
+ | 🚨 | **CI-integrity rules** — catches `continue-on-error`, `\|\| true`, and other false-green tricks |
164
+ | 🐍 | **All four Playwright bindings** — TypeScript, Python, Java, C#/.NET — plus pytest, JUnit/TestNG and CI workflows |
165
+ | 🔒 | **Local-first** — zero network calls while scanning, zero telemetry, runs in seconds |
166
+
167
+ ### The rules
123
168
 
124
169
  Every rule ships with must-fire **and** must-not-fire fixtures. A rule that
125
- fires on its own negative fixture cannot ship. That's the false-positive
170
+ fires on its own negative fixture cannot ship — that's the false-positive
126
171
  firewall.
127
172
 
128
173
  <details>
@@ -145,7 +190,7 @@ firewall.
145
190
 
146
191
  | ID | Rule | Severity |
147
192
  | ------------ | --------------------------- | -------- |
148
- | QA-TQUAL-001 | Mock-only verification | warning |
193
+ | QA-TQUAL-001 | Mock-only verification | info |
149
194
  | QA-TQUAL-002 | Tautological assertion | error |
150
195
  | QA-TQUAL-009 | Unawaited promise assertion | error |
151
196
  | QA-TQUAL-011 | Commented-out tests | warning |
@@ -160,9 +205,9 @@ firewall.
160
205
  | QA-PW-002 | Unawaited locator assertion | error |
161
206
  | QA-PW-003 | `page.pause()` / `test.only()` committed | error |
162
207
  | QA-PW-004 | Brittle CSS/XPath selectors | warning |
163
- | QA-PW-005 | Business logic inside `page.evaluate()` | warning |
164
- | QA-PW-114 | Legacy element handles (`page.$`) | warning |
165
- | QA-PW-118 | `networkidle` waits (flaky by design) | warning |
208
+ | QA-PW-005 | Business logic inside `page.evaluate()` | info |
209
+ | QA-PW-114 | Legacy element handles (`page.$`) | info |
210
+ | QA-PW-118 | `networkidle` waits (flaky by design) | info |
166
211
  | QA-PW-123 | Hardcoded environment URLs | warning |
167
212
 
168
213
  </details>
@@ -190,18 +235,125 @@ firewall.
190
235
  | QA-PY-002 | Skipped test (`skip`, non-strict `xfail`) | warning |
191
236
  | QA-PY-003 | Test function with no assertions | error |
192
237
  | QA-PY-005 | `time.sleep()` in tests | warning |
193
- | QA-PY-006 | Empty test body (`pass`) | error |
194
- | QA-PY-010 | Random/time dependence without freeze | warning |
238
+ | QA-PY-006 | Empty test body (`pass`) | info |
239
+ | QA-PY-010 | Random/time dependence without freeze | info |
195
240
  | QA-PY-012 | Tautological assertion | error |
196
241
 
242
+ 20 Python rules total (QA-PY-001…012 pytest hygiene + QA-PY-101…108 Playwright-Python).
243
+
244
+ </details>
245
+
246
+ <details>
247
+ <summary><strong>Java / JUnit · TestNG ☕</strong></summary>
248
+
249
+ | ID | Rule | Severity |
250
+ | --------- | ---------------------------------------- | -------- |
251
+ | QA-JV-101 | Disabled test (`@Disabled`) | warning |
252
+ | QA-JV-102 | Hard sleep (`Thread.sleep()`) | warning |
253
+ | QA-JV-103 | Test method with no assertions | error |
254
+ | QA-JV-105 | Playwright `waitForTimeout()` hard sleep | warning |
255
+ | QA-JV-106 | Brittle selector instead of role locator | warning |
256
+ | QA-JV-108 | Hardcoded environment URL in test | info |
257
+ | QA-JV-111 | Blanket `page.route("**")` mock | info |
258
+
259
+ </details>
260
+
261
+ <details>
262
+ <summary><strong>C# / .NET — NUnit · xUnit · MSTest 🟣</strong></summary>
263
+
264
+ | ID | Rule | Severity |
265
+ | --------- | ------------------------------------------ | -------- |
266
+ | QA-CS-101 | Skipped test (`[Ignore]`, `[Fact(Skip=)]`) | warning |
267
+ | QA-CS-102 | Hard sleep (`Thread.Sleep` / `Task.Delay`) | warning |
268
+ | QA-CS-103 | Test method with no assertions | error |
269
+ | QA-CS-105 | `WaitForTimeoutAsync()` hard sleep | warning |
270
+ | QA-CS-106 | Brittle selector instead of role locator | warning |
271
+ | QA-CS-108 | Hardcoded environment URL in test | info |
272
+ | QA-CS-111 | Blanket `page.RouteAsync("**")` mock | info |
273
+
197
274
  </details>
198
275
 
199
- > The full live catalog — every rule with confidence, false-positive risk,
200
- > and autofix availability — is generated from the registry:
276
+ > The full live catalog — every rule with tier, confidence, false-positive
277
+ > risk, and autofix availability — is generated from the registry:
201
278
  >
202
279
  > ```bash
203
280
  > mjolnir rules --md
204
281
  > ```
282
+ >
283
+ > Per-rule pages live under [`docs/rules/`](docs/rules/).
284
+
285
+ ### How much of this is measured
286
+
287
+ **74 of 99 rules carry a false-positive rate measured against real OSS code** (≥ 10 hand-classified findings each; see
288
+ [docs/FP-AUDIT.md](docs/FP-AUDIT.md)). The other 19 ship on the author's
289
+ estimate. Every scan footer tells you how many of the rules that _fired_
290
+ are measured; `mjolnir rules --unmeasured` lists the ones that aren't;
291
+ every rule's `mjolnir explain` page states its status. We publish the rate
292
+ even when it's ugly — QA-CS-103 audits at 95% and is quarantined for it.
293
+ Growing that 78 is the project's continuing work.
294
+
295
+ ### Rule tiers and language maturity
296
+
297
+ Every rule is `core`, `extended`, or `quarantine`, assigned from its
298
+ **measured** false-positive rate:
299
+
300
+ | Tier | Meaning | Default scan | `--strict` |
301
+ | ------------ | ---------------------------------------- | :----------: | :--------: |
302
+ | `core` | ≤ 10 % measured FP | ✅ | ✅ |
303
+ | `extended` | ≤ 30 % measured FP | ✅ | ✅ |
304
+ | `quarantine` | above 30 %, or not yet measured (n < 10) | ❌ | ✅ |
305
+
306
+ | Language | Adapter | Coverage today |
307
+ | --------------- | ------------ | -------------------------------------------------- |
308
+ | TypeScript / JS | compiler AST | broadest, most measured — mostly `core`/`extended` |
309
+ | Python / pytest | regex layer | broad, corpus-audited — mostly `core`/`extended` |
310
+ | Java | regex layer | newer — mostly `extended`/`quarantine` |
311
+ | C# / .NET | regex layer | newer — mostly `extended`/`quarantine` |
312
+
313
+ TypeScript and Python have the broadest measured coverage. Java and C# ship,
314
+ are documented, and stay out of the headline number until a real consumer
315
+ suite (not a binding library's own tests) has been audited.
316
+
317
+ ---
318
+
319
+ ## How the score works
320
+
321
+ <p align="center">
322
+ <img src="assets/readme/terminal-hero.svg" alt="Mjölnir terminal output — WORTHINESS 75/100 NEEDS WORK, a diagnostics-by-category breakdown, and a FIX THIS FIRST list" width="820" />
323
+ </p>
324
+
325
+ <sub>Regenerated by `npm run docs:hero`;
326
+ [`tests/hero-asset-reproducibility.spec.ts`](tests/hero-asset-reproducibility.spec.ts)
327
+ fails CI if it drifts from what the reporter actually prints.</sub>
328
+
329
+ The score is transparent: **error −8, warning −3, info −1**, then normalized
330
+ by suite exposure (deductions per test declaration). Evidence-weighted
331
+ deductions mean weak signals cost less. The terminal shows the same
332
+ discounted numbers the score uses — no black box. Full method:
333
+ [docs/SCORING.md](docs/SCORING.md).
334
+
335
+ **Verdicts**
336
+
337
+ | Score | Verdict |
338
+ | ------- | ---------------- |
339
+ | ≥ 80 | ✓ **WORTHY** |
340
+ | 50 – 79 | ⚠ **NEEDS WORK** |
341
+ | < 50 | ✖ **UNWORTHY** |
342
+
343
+ **Evidence levels** — every finding carries one; it sets the finding's
344
+ weight in the score:
345
+
346
+ | Level | Meaning | Score impact | Example |
347
+ | ----- | -------------------- | ---------------- | -------------------------------------------------- |
348
+ | E2 | Deterministic defect | Full deduction | `.only` committed — structurally provable |
349
+ | E1 | Heuristic pattern | Half deduction | Regex-matched `sleep()` — strong signal, not proof |
350
+ | E0 | Observation | Zero (info only) | Reported but never gates CI or deducts |
351
+
352
+ Most rules are **E1**. The tagline "we prove it" refers to this system:
353
+ E2 findings are structural proof; E1 findings are correctly-positioned
354
+ warnings, not formal proofs.
355
+
356
+ An empty repo scores `null`, never a fake 100 — see [Trust model](#trust-model).
205
357
 
206
358
  ---
207
359
 
@@ -222,7 +374,7 @@ regressed.
222
374
 
223
375
  ---
224
376
 
225
- ## 🔬 Runtime Evidence
377
+ ## 🔬 Runtime evidence
226
378
 
227
379
  Static flakiness detection is guessing. Mjölnir reads **real execution
228
380
  data** — Playwright JSON reports and JUnit XML from any runner:
@@ -247,7 +399,61 @@ test. It gets flagged `TRUE-FLAKE` regardless of the final green checkmark.
247
399
 
248
400
  ---
249
401
 
250
- ## 🤖 CI Integration
402
+ ## ⚡ Mjölnir is not another linter
403
+
404
+ Linters tell you whether code follows rules. Mjölnir tells you whether your
405
+ verification can be trusted.
406
+
407
+ | | ESLint / SonarQube | Coverage tools | Manual review | **Mjölnir** |
408
+ | -------------------------------------------------------- | :----------------: | :------------: | :-----------: | :---------: |
409
+ | CI workflow integrity (`continue-on-error`, `\|\| true`) | ❌ | ❌ | rarely | ✅ |
410
+ | Cross-language (TS, Python, Java, C#) from one tool | ❌ | ❌ | ❌ | ✅ |
411
+ | Grades Playwright locator resilience (Selector Health) | ❌ | ❌ | rarely | ✅ |
412
+ | Flags tests with no real assertions | ✅ (plugin)\* | ❌ | sometimes | ✅ |
413
+ | Catches hard sleeps (`waitForTimeout`, `time.sleep`) | ✅ (plugin)\* | ❌ | sometimes | ✅ |
414
+ | Runs in seconds, zero network calls while scanning | ✅ | ✅ | — | ✅ |
415
+
416
+ \*`eslint-plugin-jest` (`expect-expect`) and `eslint-plugin-playwright`
417
+ (`expect-expect`, `no-wait-for-timeout`) cover these for their respective
418
+ frameworks.
419
+
420
+ **Runtime analysis** is a separate category from static linting:
421
+
422
+ | | Playwright retry reporter | Allure / ReportPortal | **Mjölnir forensics** |
423
+ | --------------------------------------------- | :-----------------------: | :-------------------: | :-------------------: |
424
+ | Reads real run data for `TRUE-FLAKE` verdicts | partial\* | partial (tag) | ✅ |
425
+ | Flaky-triage report from execution history | ❌ | ✅ | ✅ |
426
+ | Integrates with static worthiness score | ❌ | ❌ | ✅ |
427
+
428
+ \*Playwright tracks retries internally but does not produce a standalone
429
+ flakiness report with verdict labels.
430
+
431
+ ---
432
+
433
+ ## 🤖 Why not just use AI code review?
434
+
435
+ Different problem, different layer. AI review can spot a suspicious test
436
+ change in a diff; it does not prove the verification system as a whole is
437
+ trustworthy — and it only sees the diff you show it.
438
+
439
+ | | AI code review (Copilot, etc.) | **Mjölnir** |
440
+ | ------------------------------------- | :----------------------------: | :-------------------------: |
441
+ | Cost per scan | Tokens (scales with diff size) | **Zero** (local, installed) |
442
+ | Sees the whole suite + all CI configs | Only the PR diff you show it | **Everything, every time** |
443
+ | Deterministic (same input → same out) | ❌ (non-deterministic) | **✅** |
444
+ | Catches patterns dormant for months | Only if it's in the context | **✅** (scans all files) |
445
+ | Remembers findings between runs | ❌ (no memory across sessions) | **✅** (baseline + diff) |
446
+ | Runs without human triggering | Needs a PR or prompt | **✅** (CI hook, 3 seconds) |
447
+
448
+ **Use both.** AI catches nuance, intent, and design flaws no regex can
449
+ find. Mjölnir catches the structural patterns AI overlooks because they
450
+ look "intentional" — a committed `.only`, a swallowed exit code, a
451
+ `continue-on-error` on a test job. Those aren't bugs that need reasoning;
452
+ they're facts that need scanning.
453
+
454
+ ---
455
+
456
+ ## 🤖 CI integration
251
457
 
252
458
  One command generates a PR workflow — advisory by default, never blocking:
253
459
 
@@ -264,48 +470,108 @@ Or wire it into GitHub Code Scanning natively via SARIF:
264
470
  sarif_file: mjolnir.sarif
265
471
  ```
266
472
 
473
+ Editor and pipeline setup for SARIF: [docs/SARIF-INTEGRATION.md](docs/SARIF-INTEGRATION.md).
474
+
475
+ ### Changed-scope coverage
476
+
477
+ `--scope changed` attributes findings to lines added in your branch vs the
478
+ merge-base with `main`. It covers test files (`*.spec.*`, `*.test.*`) plus
479
+ GitHub workflow files and Playwright configs in the diff. When the
480
+ merge-base can't be resolved — shallow clone, detached HEAD, non-git
481
+ target, different default branch — it degrades honestly: findings fall
482
+ back to full-file attribution and the report says so. Override the base
483
+ ref with `--base <ref>`.
484
+
267
485
  ---
268
486
 
269
- ## 🛡️ Trust Model
487
+ ## Configuration
488
+
489
+ Mjölnir is zero-config. An optional `mjolnir.config.json` (or
490
+ `.mjolnir.json`) at the repo root tunes severity, gating and scope — it
491
+ never changes detection semantics.
492
+
493
+ | Key | Type | Effect |
494
+ | ------------------- | ------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
495
+ | `exclude` | `string[]` | Extra ignore globs (gitignore subset), on top of the built-in defaults |
496
+ | `gate` | `"advisory" \| "error" \| "warning"` | Which severities exit non-zero (default `error`; `advisory` never blocks) |
497
+ | `severityOverrides` | `{ "<RULE-ID>": severity }` | Re-rank a rule's findings for your repo |
498
+ | `ignore` | `IgnoreEntry[]` | Suppress findings — **`reason` is required**; entries expire after 90 days (an explicit `expires` date, or the config file's last-modified time for entries without one) |
499
+ | `plugins` | `string[]` | Third-party rule packages (see [Trust model](#trust-model)) |
500
+
501
+ ```json
502
+ {
503
+ "gate": "error",
504
+ "exclude": ["legacy/**"],
505
+ "severityOverrides": { "QA-PW-118": "warning" },
506
+ "ignore": [
507
+ {
508
+ "ruleId": "QA-TEST-004",
509
+ "files": ["e2e/legacy-login.spec.ts"],
510
+ "reason": "Third-party widget needs a settle delay; tracked in JIRA-4821",
511
+ "expires": "2026-12-31"
512
+ }
513
+ ]
514
+ }
515
+ ```
270
516
 
271
- - **Local-first** — zero network calls during scanning. Ever.
272
- - **No false proof** — we'd rather say "unknown" than "verified". An empty
273
- repo gets `score: null`, never a fake 100.
274
- - **Transparent scoring** — public deduction constants: error −8, warning −3,
275
- info −1. No black box. Verdict thresholds: **≥80 WORTHY**, **50–79 NEEDS
276
- WORK**, **<50 UNWORTHY**.
277
- - **Evidence-weighted deductions** — every finding carries an evidence level:
278
- **E2** (deterministic defect, full deduction), **E1** (heuristic pattern,
279
- half deduction), **E0** (observation — reported but costs nothing and never
280
- gates CI).
281
- - **Partial honesty** — if analysis was cut short, the output says so.
282
- Never "complete" when it isn't.
283
- - **QA-native language** — findings speak your vocabulary:
284
- `FALSE-GREEN`, `FLAKY-RISK`, `BLOCKS-RELEASE`, `HYGIENE`.
285
- - **FP firewall** — detection runs on a comment/string-free view of the code
286
- (TypeScript rules use the compiler AST): a pattern inside a prose comment
287
- or a doc-example string is documentation, not a finding.
517
+ - **`.mjolnirignore`** — a plain gitignore-style file for path exclusions,
518
+ same dialect as `exclude`. Use it for machine-wide noise; use `exclude`
519
+ when the list belongs in version control alongside the rest of the config.
520
+ - **CLI overrides** — `--strict` (include quarantine rules), `--width <cols>`
521
+ and `--ascii` / `--no-ascii` (terminal rendering), `--tone blunt`
522
+ (blunter messages), `--max-duration <sec>` (bounded partial scan).
523
+ - Rule suppression and deprecation lifecycle: [docs/RULE-LIFECYCLE.md](docs/RULE-LIFECYCLE.md).
524
+
525
+ `ignore` entries also power the standalone `mjolnir suppressions` command,
526
+ which lists what's currently suppressed and when each entry expires.
288
527
 
289
528
  ---
290
529
 
291
- ## 🔌 Plugin Trust Model
530
+ ## 📐 Exit codes & contracts
531
+
532
+ Frozen — safe to build CI logic on:
533
+
534
+ | Exit code | Meaning |
535
+ | --------- | --------------------------------------------------------------- |
536
+ | `0` | Clean — no findings at or above the gate |
537
+ | `1` | Findings at or above the gate |
538
+ | `2` | Partial scan (time budget hit, unreadable files) — never blocks |
539
+ | `10` | Usage error (bad flag, missing target) |
540
+ | `20` | Internal error |
292
541
 
293
- Plugins are npm packages you declare in `mjolnir.config.json`. There is **no
294
- sandbox**: plugin code executes with full Node privileges against the scanned
295
- tree — the same trust model as ESLint or Vitest plugins. Only install plugins
296
- you'd trust as devDependencies. Core rule-ID prefixes (`QA-TEST`, `QA-TQUAL`,
297
- `QA-PW`, `QA-CI`, `QA-PY`, `QA-ENV`, `QA-JV`, `QA-CS`) are reserved and
298
- rejected from plugins to prevent spoofing.
542
+ The JSON/SARIF report is `schemaVersion: 1`. Rule IDs (`QA-<FAMILY>-NNN`)
543
+ are immutable once shipped and never reused.
299
544
 
300
545
  ---
301
546
 
302
- ## 🔍 Changed-Scope Coverage
547
+ ## Trust model
303
548
 
304
- `--scope changed` attributes findings to lines added in your branch vs the
305
- merge-base with `main`. It covers test files (`*.spec.*`, `*.test.*`) plus
306
- GitHub workflow files and Playwright configs changed in the diff. On shallow
307
- clones or non-git targets it degrades honestly: findings fall back to
308
- full-file attribution and the report says so.
549
+ - **Local-first** — zero network calls during scanning. Ever. Zero telemetry.
550
+ - **No false proof** — we'd rather say "unknown" than "verified". An empty
551
+ repo gets `score: null`, never a fake 100.
552
+ - **Partial honesty** — if analysis was cut short, the output says so.
553
+ Never "complete" when it isn't.
554
+ - **FP firewall** — detection runs on a comment/string-free view of the code
555
+ (TypeScript rules use the compiler AST): a pattern inside a prose comment
556
+ or a doc-example string is documentation, not a finding.
557
+ - **Measured, not asserted** — only rules with a false-positive rate from
558
+ real OSS code ship in the headline tiers (see
559
+ [How much of this is measured](#how-much-of-this-is-measured)); the scan
560
+ footer and `mjolnir rules --unmeasured` tell you which is which.
561
+ - **Plugin trust** — plugins are npm packages declared under `"plugins"`.
562
+ There is **no sandbox**: plugin code runs with full Node privileges, the
563
+ same trust model as ESLint or Vitest plugins. Core rule-ID prefixes are
564
+ reserved and rejected from plugins to prevent spoofing.
565
+ - **Workspace-local external rules** (folder-based, zero network) — a
566
+ `mjolnir-rules/` directory next to the scan target loads custom rules:
567
+ JSON files declare regex patterns (no code executed), `.mjs`/`.js`
568
+ modules export `rules` (full-Node trust, same as plugins). External
569
+ rules carry the same trust metadata as core; they can never ship in
570
+ the core tier (core requires a measured FP rate from the corpus
571
+ sidecar — a declared `tier: "core"` is clamped to `extended`), obey
572
+ tier caps, and are drift-checked: `mjolnir rules --md --external`
573
+ renders the catalog from the loaded files (provenance `external`),
574
+ and the matrix generator accepts `--external <root>`.
309
575
 
310
576
  ---
311
577
 
@@ -319,15 +585,16 @@ mjolnir/
319
585
  ├── src/
320
586
  │ ├── engine/ # LanguageAdapter interface + rule runner
321
587
  │ ├── adapters/ # typescript · python · java · csharp · github-actions
322
- │ ├── rules/ # deterministic rules across 8 families
588
+ │ ├── rules/ # rules across 8 families + the measured-FP table
323
589
  │ ├── playwright/ # Selector Health Score engine
324
- │ ├── discovery/ # workspace, frameworks, safe-YAML parser
590
+ │ ├── discovery/ # workspace, frameworks, ignore resolution
325
591
  │ ├── scope/ # git merge-base changed-scope engine
326
- │ ├── scorer/ # transparent deduction table
592
+ │ ├── scorer/ # transparent deduction table + prioritization
327
593
  │ ├── reporter/ # terminal · JSON · SARIF 2.1 · Mermaid
328
594
  │ ├── forensics/ # run-data ingestion · flake verdicts · triage
329
- │ ├── commands/ # fix · badge · debt · handover · init · create-rule · doctor · rules-catalog · explain · impact · baseline · diff · pr-comment · stats
330
- │ └── integrations/ # CI workflow generator
595
+ │ ├── config/ # mjolnir.config.json + suppressions
596
+ │ ├── plugins/ # third-party rule loading (no sandbox)
597
+ │ └── commands/ # every subcommand
331
598
  └── tests/
332
599
  ├── fixtures/ # must-fire / must-not-fire per rule
333
600
  └── golden/ # frozen score regression locks
@@ -335,54 +602,60 @@ mjolnir/
335
602
 
336
603
  </details>
337
604
 
338
- Multi-language by design: adding a language = one adapter + its rules.
339
- TypeScript/Playwright and Python are the most mature; Java and C#/.NET
340
- adapters ship a regex-based core rule family with tree-sitter WASM AST
341
- precision as the next step.
605
+ - **Rules are pure functions** — `(SourceFileContext) → Finding[]`, no I/O,
606
+ no globals. Adding an ecosystem = one adapter + its rules.
607
+ - **TypeScript/Playwright uses the compiler AST** (ts-morph). Python, Java
608
+ and C# run on a shared comment/string-masked regex layer.
609
+ - A tree-sitter WASM AST layer for Java and C# exists and is the next
610
+ precision step — it is not yet wired into the synchronous scan pipeline.
342
611
 
343
612
  ---
344
613
 
345
- ## 📦 Install
614
+ ## 📚 Documentation
615
+
616
+ | Document | What's in it |
617
+ | ------------------------------------------------------ | ------------------------------------------------- |
618
+ | [docs/SCORING.md](docs/SCORING.md) | Score normalization + evidence weighting |
619
+ | [docs/FP-AUDIT.md](docs/FP-AUDIT.md) | Measured false-positive rates + method |
620
+ | [docs/RULE-LIFECYCLE.md](docs/RULE-LIFECYCLE.md) | Rule states, suppression, deprecation |
621
+ | [docs/VERSIONING.md](docs/VERSIONING.md) | Semver policy, frozen surfaces, deprecation cycle |
622
+ | [docs/SARIF-INTEGRATION.md](docs/SARIF-INTEGRATION.md) | SARIF output + editor/CI setup |
623
+ | [docs/rules/](docs/rules/) | Generated per-rule catalog |
624
+ | [CONTRIBUTING.md](CONTRIBUTING.md) | Dev setup + contribution workflow |
625
+ | [SUPPORT.md](SUPPORT.md) | Where to ask, report and get help |
626
+ | [CHANGELOG.md](CHANGELOG.md) | Release history |
627
+ | [SECURITY.md](SECURITY.md) | Vulnerability reporting |
346
628
 
347
- ```bash
348
- # Run directly (recommended):
349
- npx mjolnir-qa@latest
629
+ ---
350
630
 
351
- # Or install globally:
352
- npm i -g mjolnir-qa
353
- ```
631
+ ## 📈 Status
354
632
 
355
- Requires Node.js ≥ 22.18. Works on Windows, macOS, and Linux.
633
+ **v0.5.x · open beta.** The JSON schema and exit codes are frozen contracts.
634
+ TypeScript and Python have the broadest measured coverage; Java and C# are
635
+ newer — read them through the [maturity table](#rule-tiers-and-language-maturity).
636
+ Honest scope, no invented dates: the [public roadmap](https://sergey-bar.github.io/Mjolnir/reference/roadmap).
356
637
 
357
638
  ---
358
639
 
359
640
  ## 🤝 Contributing
360
641
 
361
- New rules are the easiest first contribution — one command scaffolds
362
- everything (anti-creep law enforced by the fixture harness):
642
+ New rules are the easiest first contribution — one command scaffolds the
643
+ rule plus its must-fire **and** must-not-fire fixtures (the generated rule
644
+ intentionally fails its fixtures until you implement real detection — a stub
645
+ cannot ship):
363
646
 
364
647
  ```bash
365
648
  mjolnir create-rule QA-PW-140 --title "Screenshot without diff bound"
366
649
  ```
367
650
 
368
- That generates the rule file plus must-fire AND must-not-fire fixture
369
- skeletons. The generated rule intentionally fails its fixtures until you
370
- implement real detection — a stub cannot ship.
371
-
372
- Manual path:
373
-
374
- 1. Fork + clone, `npm install`
375
- 2. Copy any rule folder under `src/rules/` as a template
376
- 3. Add must-fire AND must-not-fire fixtures under `tests/fixtures/<YOUR-ID>/`
377
- 4. `npx vitest run` — both directions green? Ship the PR.
378
-
379
- See [SECURITY.md](SECURITY.md) for vulnerability reporting.
651
+ Full dev setup, the standing-gate commands, and the anti-creep / fixture
652
+ firewall laws are in [CONTRIBUTING.md](CONTRIBUTING.md).
380
653
 
381
654
  ---
382
655
 
383
656
  <div align="center">
384
657
 
385
- **Are your tests worthy of trust?**
658
+ **Stop shipping tests you can't trust.**
386
659
 
387
660
  ```bash
388
661
  npx mjolnir-qa@latest