mjolnir-qa 0.5.12 → 0.5.14

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -12,15 +12,23 @@ pipelines, reports a worthiness score, and shows exactly where trust breaks.
12
12
  [![license](https://img.shields.io/badge/license-MIT-C9A227.svg?style=flat-square&labelColor=0B0F17)](LICENSE)
13
13
  [![node](https://img.shields.io/badge/node-%E2%89%A5%2022.18-2E8C7F.svg?style=flat-square&labelColor=0B0F17)](https://nodejs.org)
14
14
 
15
- English | [简体中文](README.zh.md) | [繁體中文](README.zht.md) | [한국어](README.ko.md) | [Deutsch](README.de.md) | [Español](README.es.md) | [Français](README.fr.md) | [Italiano](README.it.md) | [Dansk](README.da.md) | [日本語](README.ja.md) | [Polski](README.pl.md) | [Русский](README.ru.md) | [Norsk](README.no.md) | [Português (Brasil)](README.br.md) | [ไทย](README.th.md) | [Türkçe](README.tr.md) | [Українська](README.uk.md) | [বাংলা](README.bn.md) | [Ελληνικά](README.gr.md) | [Tiếng Việt](README.vi.md) | [עברית](README.he.md) | [العربية](README.ar.md) | [Bosanski](README.bs.md)
16
-
17
15
  ```bash
18
16
  npx mjolnir-qa@latest
19
17
  ```
20
18
 
21
19
  **Are your tests worthy of trust?**
22
20
 
23
- [See it work](#-see-it-work) · [Quickstart](#-quickstart) · [What it checks](#-what-mjölnir-checks) · [Scoring](#how-the-score-works) · [CI](#-ci-integration) · [Configuration](#configuration) · [Docs](#-documentation)
21
+ [See it work](#-see-it-work) · [Quickstart](#-quickstart) · [Who it's for](#-who-is-this-for) · [Why not a linter](#-mjölnir-is-not-another-linter) · [What it verifies](#-what-mjölnir-verifies) · [Scoring](#-how-the-score-works) · [Runtime evidence](#-runtime-evidence) · [CI](#-ci-integration) · [Agents & MCP](#-works-with-your-agent) · [Docs](#-documentation) · [Contributing](#-contributing)
22
+
23
+ <details>
24
+ <summary>Read this in another language — 22 translations</summary>
25
+
26
+ English | [简体中文](README.zh.md) | [繁體中文](README.zht.md) | [한국어](README.ko.md) | [Deutsch](README.de.md) | [Español](README.es.md) | [Français](README.fr.md) | [Italiano](README.it.md) | [Dansk](README.da.md) | [日本語](README.ja.md) | [Polski](README.pl.md) | [Русский](README.ru.md) | [Norsk](README.no.md) | [Português (Brasil)](README.br.md) | [ไทย](README.th.md) | [Türkçe](README.tr.md) | [Українська](README.uk.md) | [বাংলা](README.bn.md) | [Ελληνικά](README.gr.md) | [Tiếng Việt](README.vi.md) | [עברית](README.he.md) | [العربية](README.ar.md) | [Bosanski](README.bs.md)
27
+
28
+ English is canonical. Translations are machine-assisted and may lag behind
29
+ it; `npm run docs:translations` reports how far.
30
+
31
+ </details>
24
32
 
25
33
  </div>
26
34
 
@@ -28,15 +36,48 @@ npx mjolnir-qa@latest
28
36
 
29
37
  ## 🎬 See it work
30
38
 
39
+ <!-- Plays inline on github.com only: <video> is rendered for GitHub's own
40
+ user-content CDN, never for a repo-relative path. The <a> below is the
41
+ fallback for every other renderer (npm, mirrors, offline clones). -->
31
42
  <p align="center">
32
- <img src="assets/readme/demo.svg" alt="Mjölnir's full --verbose report on a demo repo: WORTHINESS 75/100 NEEDS WORK, a diagnostics-by-category breakdown, a FIX THIS FIRST list, and every finding with its rule ID and line number across CI, Playwright, test-hygiene and Python rules" width="900" />
43
+ <video
44
+ src="https://github.com/user-attachments/assets/0e1af1e4-1e27-4c1c-9ec4-2717d194df05"
45
+ poster="https://raw.githubusercontent.com/Sergey-Bar/Mjolnir/main/assets/video/mjolnir-demo-poster.png"
46
+ controls
47
+ muted
48
+ playsinline
49
+ width="900"></video>
50
+ </p>
51
+
52
+ <p align="center">
53
+ <strong>42 seconds</strong> — one false-green CI gate: found, fixed,
54
+ re-proved, then handed to an agent.
55
+ <a href="assets/video/mjolnir-demo.mp4">Download the 1440p MP4</a> if the
56
+ player above doesn't load.
57
+ </p>
58
+
59
+ <sub>Every frame is real CLI output — the 75 → 90 score change is a real
60
+ re-scan after applying the fix the tool itself printed, never a mockup.
61
+ Rendered by `npm run docs:video` from
62
+ [`assets/video/script.demo.json`](assets/video/script.demo.json);
63
+ [`tests/contract/video-script.spec.ts`](tests/contract/video-script.spec.ts)
64
+ fails CI if that script stops matching what the CLI prints, or if the
65
+ findings the video shows as fixed turn out to still be there.</sub>
66
+
67
+ <details>
68
+ <summary><strong>Prefer it inline?</strong> The full <code>--verbose</code> report, as an animated SVG</summary>
69
+
70
+ <p align="center">
71
+ <img src="assets/readme/demo.svg" alt="Mjölnir's full --verbose report on a demo repo: WORTHINESS 75/100 NEEDS WORK, a diagnostics-by-category breakdown, a FIX THIS FIRST list, and every finding with its rule ID and line number" width="900" />
33
72
  </p>
34
73
 
35
74
  <sub>The complete `npx mjolnir-qa ./examples/demo-repo --verbose` output,
36
75
  rendered from the actual reporter — nothing trimmed. Regenerated by
37
76
  `npm run docs:demo`;
38
- [`tests/demo-asset-reproducibility.spec.ts`](tests/demo-asset-reproducibility.spec.ts)
39
- fails CI if it drifts from what the tool prints.</sub>
77
+ [`tests/contract/demo-asset-reproducibility.spec.ts`](tests/contract/demo-asset-reproducibility.spec.ts)
78
+ fails CI if it drifts.</sub>
79
+
80
+ </details>
40
81
 
41
82
  **What just happened:**
42
83
 
@@ -47,28 +88,50 @@ fails CI if it drifts from what the tool prints.</sub>
47
88
  brittle selector, hardcoded staging URLs, a `networkidle` wait.
48
89
  3. It turned each into a concrete finding with a rule ID, a location and a
49
90
  fix — and a single score you can gate a PR on.
91
+ 4. `mjolnir handoff` turned the findings into a remediation plan — evidence,
92
+ constraints and a copy-pastable prompt per finding — that Claude Code
93
+ (or any other agent) can work through, with the tool's own verification
94
+ discipline built in.
95
+
96
+ An 89-second tour of `explain` and `forensics` is built from
97
+ [`script.tour.json`](assets/video/script.tour.json) by `npm run docs:video`
98
+ — not committed as an MP4 (~16MB), so check
99
+ [Releases](../../releases) or render it yourself.
50
100
 
51
101
  ### One finding, up close
52
102
 
53
103
  Run `mjolnir explain QA-CI-001` on the first finding above and you get:
54
104
 
55
105
  ```text
56
- ▚ QA-CI-001 — continue-on-error masks a failing verification gate
106
+ ▚ QA-CI-001 — continue-on-error masks a failing verification gate
57
107
 
58
108
  Severity: error
59
109
  Confidence: high
110
+ Tier: quarantine
60
111
  Evidence: E2
61
- Measured FP: not yet measured — this rule ships on assumption (see docs/FP-AUDIT.md)
112
+ QA impact: False-green risk (FALSE-GREEN)
113
+ Measured FP: 11% (19 hand-classified corpus verdicts)
114
+ FP risk: low (author estimate)
115
+ Languages: yaml
116
+ Frameworks: github-actions
62
117
 
63
118
  WHAT WAS FOUND (real detector output, not a mockup)
64
119
  Job `security-scan` runs a verification gate under `continue-on-error: true`.
65
120
 
66
121
  WHY IT MATTERS
67
- This job can fail every day and CI will still show green. The checkmark
68
- on this workflow cannot be trusted.
122
+ This job can fail every day and CI will still show green. The checkmark on
123
+ this workflow cannot be trusted.
69
124
 
70
125
  HOW TO FIX
71
126
  Remove continue-on-error, or scope it to individual non-blocking steps only.
127
+
128
+ Example from this rule's own must-fire fixture: QA-CI-001/must-fire/masked.yml
129
+
130
+ HOW TO VERIFY THE FIX
131
+ Re-run `mjolnir` on the changed file(s) — this finding should no longer
132
+ appear. `mjolnir --scope changed` scopes the check to just what you touched.
133
+
134
+ Docs: mjolnir rules --md (full catalog, this rule included)
72
135
  ```
73
136
 
74
137
  That is the unit of value: not a style nit, but a place where your CI is
@@ -100,6 +163,7 @@ and you're done. Everything else is optional.
100
163
  | `mjolnir --scope changed` | Only what your branch introduced — the CI form |
101
164
  | `mjolnir ci install` | Generate the advisory PR workflow |
102
165
  | `mjolnir explain QA-CI-001` | What / why / fix + measured FP rate for one rule |
166
+ | `mjolnir why src/a.spec.ts:42` | Why this exact line was flagged — never a gate |
103
167
  | `mjolnir rules --unmeasured` | The rules running on assumption, not measurement |
104
168
  | `mjolnir --json` / `--format sarif` | Machine-readable / GitHub Code Scanning |
105
169
  | `mjolnir --strict` | Also run quarantine-tier rules (higher FP risk) |
@@ -120,19 +184,23 @@ and you're done. Everything else is optional.
120
184
  <details>
121
185
  <summary><strong>Occasional / reporting</strong></summary>
122
186
 
123
- | Command | What it does |
124
- | ------------------------------- | ------------------------------------------------ |
125
- | `mjolnir fix --dry-run` / `fix` | Safe auto-fixes with proof |
126
- | `mjolnir baseline` / `diff` | Snapshot findings, then report only new/worsened |
127
- | `mjolnir impact --since <ref>` | What changed since a prior commit |
128
- | `mjolnir debt` | Test-debt register with a cost model |
129
- | `mjolnir handover` | New-QA onboarding map of the suite |
130
- | `mjolnir stats` | Local all-time counters of fixes seen |
131
- | `mjolnir badge` | shields.io endpoint JSON + snippet |
132
- | `mjolnir rules --md` | Full rule catalog (JSON or Markdown) |
133
- | `mjolnir doctor` | Self-audit of Mjölnir's own rule base |
134
- | `mjolnir create-rule <ID>` | Scaffold a new rule + fixtures |
135
- | `mjolnir --format mermaid` | Test-architecture diagram for a PR comment |
187
+ | Command | What it does |
188
+ | ------------------------------- | ------------------------------------------------------ |
189
+ | `mjolnir fix --dry-run` / `fix` | Safe auto-fixes with proof |
190
+ | `mjolnir baseline` / `diff` | Snapshot findings, then report only new/worsened |
191
+ | `mjolnir impact --since <ref>` | What changed since a prior commit |
192
+ | `mjolnir summary` | CI annotations + step summary from a saved report |
193
+ | `mjolnir pr-comment` | A scoped PR comment, as Markdown |
194
+ | `mjolnir debt` | Test-debt register with a cost model |
195
+ | `mjolnir handover` | New-QA onboarding map of the suite |
196
+ | `mjolnir init` | Detect frameworks + setup checklist (never overwrites) |
197
+ | `mjolnir suppressions` | List suppressed findings — governance transparency |
198
+ | `mjolnir stats` | Local all-time counters of fixes seen |
199
+ | `mjolnir badge` | shields.io endpoint JSON + snippet |
200
+ | `mjolnir rules --md` | Full rule catalog (JSON or Markdown) |
201
+ | `mjolnir doctor` | Self-audit of Mjölnir's own rule base |
202
+ | `mjolnir create-rule <ID>` | Scaffold a new rule + fixtures |
203
+ | `mjolnir --format mermaid` | Test-architecture diagram for a PR comment |
136
204
 
137
205
  </details>
138
206
 
@@ -153,7 +221,35 @@ Requires Node.js ≥ 22.18. Works on Windows, macOS, and Linux.
153
221
 
154
222
  ---
155
223
 
156
- ## 🔨 What Mjölnir checks
224
+ ## ⚡ Mjölnir is not another linter
225
+
226
+ Linters tell you whether code follows rules. Mjölnir tells you whether your
227
+ verification can be trusted.
228
+
229
+ | | ESLint / SonarQube | Coverage tools | AI code review | **Mjölnir** |
230
+ | -------------------------------------------------------- | :----------------: | :------------: | :------------: | :--------------: |
231
+ | CI workflow integrity (`continue-on-error`, `\|\| true`) | ❌ | ❌ | only the diff | ✅ |
232
+ | Cross-language (TS, Python, Java, C#) from one tool | ❌ | ❌ | ❌ | ✅ |
233
+ | Grades Playwright locator resilience (Selector Health) | ❌ | ❌ | ❌ | ✅ |
234
+ | Flags tests with no real assertions | ✅ (plugin)\* | ❌ | sometimes | ✅ |
235
+ | Catches hard sleeps (`waitForTimeout`, `time.sleep`) | ✅ (plugin)\* | ❌ | sometimes | ✅ |
236
+ | Reads real run data for `TRUE-FLAKE` verdicts | ❌ | ❌ | ❌ | ✅ |
237
+ | Deterministic (same input → same output) | ✅ | ✅ | ❌ | ✅ |
238
+ | Cost per scan | free | free | tokens | **zero** (local) |
239
+
240
+ \*`eslint-plugin-jest` (`expect-expect`) and `eslint-plugin-playwright`
241
+ (`expect-expect`, `no-wait-for-timeout`) cover these for their respective
242
+ frameworks.
243
+
244
+ **Use AI review too.** It catches nuance, intent, and design flaws no regex
245
+ can find. Mjölnir catches the structural patterns AI overlooks because they
246
+ look "intentional" — a committed `.only`, a swallowed exit code, a
247
+ `continue-on-error` on a test job. Those aren't bugs that need reasoning;
248
+ they're facts that need scanning.
249
+
250
+ ---
251
+
252
+ ## 🔨 What Mjölnir verifies
157
253
 
158
254
  | | |
159
255
  | --- | ----------------------------------------------------------------------------------------------------------------- |
@@ -164,12 +260,15 @@ Requires Node.js ≥ 22.18. Works on Windows, macOS, and Linux.
164
260
  | 🐍 | **All four Playwright bindings** — TypeScript, Python, Java, C#/.NET — plus pytest, JUnit/TestNG and CI workflows |
165
261
  | 🔒 | **Local-first** — zero network calls while scanning, zero telemetry, runs in seconds |
166
262
 
167
- ### The rules
168
-
169
263
  Every rule ships with must-fire **and** must-not-fire fixtures. A rule that
170
264
  fires on its own negative fixture cannot ship — that's the false-positive
171
265
  firewall.
172
266
 
267
+ **The rule catalog.** Every family is collapsed below; the generated
268
+ full catalog lives in [`docs/rules/`](docs/rules/),
269
+ [what it checks](https://sergey-bar.github.io/Mjolnir/guide/what-it-checks),
270
+ or `mjolnir rules --md`.
271
+
173
272
  <details>
174
273
  <summary><strong>Test Hygiene</strong></summary>
175
274
 
@@ -282,6 +381,29 @@ firewall.
282
381
  >
283
382
  > Per-rule pages live under [`docs/rules/`](docs/rules/).
284
383
 
384
+ ### Selector Health Score
385
+
386
+ A pass rate says the suite went green today. It says nothing about whether
387
+ it will still go green after a redesign. `mjolnir doctor:playwright` grades
388
+ every locator in a spec file by how it finds an element — the way a user
389
+ identifies it (role, label, text), an explicit contract (`data-testid`), or
390
+ a structural accident (CSS chains, XPath) — and scores the file 0–100:
391
+
392
+ ```text
393
+ ▚ SELECTOR HEALTH
394
+
395
+ e2e/login.spec.ts
396
+ [█████████████░░░░░░░] 65 / 100
397
+ role/text: 1 · testid: 0 · css-chains: 1 ⚠ · xpath: 0
398
+
399
+ e2e/checkout.spec.ts
400
+ [█████████████████░░░] 86 / 100
401
+ role/text: 3 · testid: 1 · css-chains: 1 ⚠ · xpath: 0
402
+ ```
403
+
404
+ `.btn.btn-primary > div:nth-child(2)` is not a test failure today — it's a
405
+ test failure scheduled for whenever someone touches the markup.
406
+
285
407
  ### How much of this is measured
286
408
 
287
409
  **78 of 99 rules carry a false-positive rate measured against real OSS code** (≥ 10 hand-classified findings each; see
@@ -295,83 +417,85 @@ Growing that number is the project's continuing work.
295
417
  ### Rule tiers and language maturity
296
418
 
297
419
  Every rule is `core`, `extended`, or `quarantine`, assigned from its
298
- **measured** false-positive rate:
299
-
300
- | Tier | Meaning | Default scan | `--strict` |
301
- | ------------ | ---------------------------------------- | :----------: | :--------: |
302
- | `core` | ≤ 10 % measured FP | ✅ | ✅ |
303
- | `extended` | ≤ 30 % measured FP | ✅ | ✅ |
304
- | `quarantine` | above 30 %, or not yet measured (n < 10) | ❌ | ✅ |
305
-
306
- | Language | Adapter | Coverage today |
307
- | --------------- | ------------ | -------------------------------------------------- |
308
- | TypeScript / JS | compiler AST | broadest, most measured — mostly `core`/`extended` |
309
- | Python / pytest | regex layer | broad, corpus-audited — mostly `core`/`extended` |
310
- | Java | regex layer | newer — mostly `extended`/`quarantine` |
311
- | C# / .NET | regex layer | newer — mostly `extended`/`quarantine` |
312
-
313
- TypeScript and Python have the broadest measured coverage. Java and C# ship,
314
- are documented, and stay out of the headline number until a real consumer
315
- suite (not a binding library's own tests) has been audited.
420
+ **measured** false-positive rate — quarantine rules only run under
421
+ `--strict`. Tiers, language maturity and the promotion/demotion rules:
422
+ [rule lifecycle](https://sergey-bar.github.io/Mjolnir/reference/rule-lifecycle).
423
+
424
+ ### What it can't tell you
425
+
426
+ A tool that claims your tests are lying should be blunt about its own
427
+ limits:
428
+
429
+ - **It does not run your tests.** It reads your code and, when you point
430
+ it at one, a run report. A clean scan is not a passing suite.
431
+ - **It cannot tell you an assertion is _wrong_.** `expect(total).toBe(41)`
432
+ is a perfectly healthy-looking test. Mjölnir finds tests that can't fail
433
+ and pipelines that can't go red — not tests that check the wrong thing.
434
+ - **A 100 is not proof of a good suite.** It means none of these 99 rules
435
+ fired. Coverage of your actual risk is a different question, and this
436
+ tool does not pretend to answer it.
437
+ - **21 of 99 rules ship on an estimate**, not a measured rate — and they
438
+ say so, per rule, in `mjolnir explain`.
439
+ - **E1 findings are heuristics.** They are positioned to be worth reading,
440
+ not to be applied blindly; the evidence level is attached to every
441
+ finding precisely so you can tell the difference.
442
+ - **An empty repo scores `null`, never 100.** "Unknown" is a verdict here.
443
+
444
+ ### We run it on ourselves
445
+
446
+ Every CI run scans this repository with the build produced by that same
447
+ run, and the gate fails on **any** error-severity finding — but also on a
448
+ _partial_ scan or a crashed rule, because a truncated self-scan that
449
+ reports nothing is exactly the false green this project exists to catch.
450
+ The result is uploaded as a build artifact on every run.
316
451
 
317
452
  ---
318
453
 
319
- ## How the score works
454
+ ## 📊 How the score works
320
455
 
321
- <p align="center">
322
- <img src="assets/readme/terminal-hero.svg" alt="Mjölnir terminal output — WORTHINESS 75/100 NEEDS WORK, a diagnostics-by-category breakdown, and a FIX THIS FIRST list" width="820" />
323
- </p>
456
+ <table>
457
+ <tr>
458
+ <td width="46%" valign="top" align="center">
324
459
 
325
- <sub>Regenerated by `npm run docs:hero`;
326
- [`tests/hero-asset-reproducibility.spec.ts`](tests/hero-asset-reproducibility.spec.ts)
327
- fails CI if it drifts from what the reporter actually prints.</sub>
460
+ <img src="assets/readme/score-gauge.svg" alt="The hammer sweeping every score from 0 to 100 — cracked below 50 (UNWORTHY), strained 50-79 (NEEDS WORK), charged 80-99 (WORTHY), forged at 100 (FORGED) — then holding on FORGED before it loops" width="270" />
328
461
 
329
- The score is transparent: **error −8, warning −3, info −1**, then normalized
330
- by suite exposure (deductions per test declaration). Evidence-weighted
331
- deductions mean weak signals cost less. The terminal shows the same
332
- discounted numbers the score uses — no black box. Full method:
333
- [docs/SCORING.md](docs/SCORING.md).
334
-
335
- **Verdicts**
462
+ **What the score means**
336
463
 
337
- | Score | Verdict |
338
- | ------- | ---------------- |
339
- | 100 | ⚡ **FORGED** |
340
- | ≥ 80 | ✓ **WORTHY** |
341
- | 50 – 79 | ⚠ **NEEDS WORK** |
342
- | < 50 | ✖ **UNWORTHY** |
464
+ </td>
465
+ <td width="54%" valign="top" align="center">
343
466
 
344
- **Evidence levels** — every finding carries one; it sets the finding's
345
- weight in the score:
467
+ <img src="assets/readme/terminal-hero.svg" alt="Mjölnir's deduction breakdown — WORTHINESS 75/100 NEEDS WORK, a diagnostics-by-category bar chart, the per-severity deduction box, and a FIX THIS FIRST list" width="330" />
346
468
 
347
- | Level | Meaning | Score impact | Example |
348
- | ----- | -------------------- | ---------------- | -------------------------------------------------- |
349
- | E2 | Deterministic defect | Full deduction | `.only` committed — structurally provable |
350
- | E1 | Heuristic pattern | Half deduction | Regex-matched `sleep()` — strong signal, not proof |
351
- | E0 | Observation | Zero (info only) | Reported but never gates CI or deducts |
469
+ **Where the points went**
352
470
 
353
- Most rules are **E1**. The tagline "we prove it" refers to this system:
354
- E2 findings are structural proof; E1 findings are correctly-positioned
355
- warnings, not formal proofs.
471
+ </td>
472
+ </tr>
473
+ </table>
356
474
 
357
- An empty repo scores `null`, never a fake 100 — see [Trust model](#trust-model).
475
+ <sub>Left: every score 0–100 through the real `deriveScoreState`/`renderHammer`
476
+ (`npm run docs:gauge`) — the hammer has four real states, so it snaps at the
477
+ exact thresholds rather than faking an in-between shape. Right: a real strict
478
+ scan of `examples/demo-repo` (`npm run docs:hero`), excerpted to the deduction
479
+ block. Both are drift-locked
480
+ ([gauge](tests/contract/score-gauge-asset-reproducibility.spec.ts) ·
481
+ [breakdown](tests/contract/hero-asset-reproducibility.spec.ts)).</sub>
358
482
 
359
- ---
360
-
361
- ## 🎭 Selector Health Score
362
-
363
- The headline metric for Playwright suites — how resilient your locators are:
364
-
365
- ```text
366
- ▚ SELECTOR HEALTH — e2e/checkout.spec.ts
483
+ The score is transparent: **error −8, warning −3, info −1**, then normalized
484
+ by suite exposure (deductions per test declaration). Evidence-weighted
485
+ deductions mean weak signals cost less. The terminal shows the same
486
+ discounted numbers the score uses — no black box.
367
487
 
368
- [█████████████████░░░] 83 / 100
369
- role/text: 2 · testid: 1 · css-chains: 1 ⚠ · xpath: 0
370
- ```
488
+ Each band also carries an evidence level: **100/FORGED** is **E2**
489
+ (deterministic defect, full deduction), **≥ 80/WORTHY** and **50–79/NEEDS
490
+ WORK** findings are mostly **E1** (heuristic pattern, half deduction), and
491
+ **E0** (observation) findings cost nothing — info only. The tagline "we
492
+ prove it" refers to this system: E2 findings are structural proof; E1
493
+ findings are correctly-positioned warnings, not formal proofs.
371
494
 
372
- Role-based locators score full credit. CSS class chains and XPath tank the
373
- score — they break on any DOM refactor without telling you which behavior
374
- regressed.
495
+ **No false proof.** We'd rather say "unknown" than "verified" — an empty
496
+ repo scores `null`, never a fake 100. Full method:
497
+ [docs/SCORING.md](docs/SCORING.md) ·
498
+ [scoring guide](https://sergey-bar.github.io/Mjolnir/guide/scoring).
375
499
 
376
500
  ---
377
501
 
@@ -385,7 +509,7 @@ mjolnir forensics ./test-results/
385
509
  ```
386
510
 
387
511
  ```text
388
- ▚ FLAKINESS LEADERBOARD
512
+ ▚ FLAKINESS LEADERBOARD
389
513
 
390
514
  3 tests · 1 failed · 1 flaky · 1 retried
391
515
 
@@ -400,60 +524,6 @@ test. It gets flagged `TRUE-FLAKE` regardless of the final green checkmark.
400
524
 
401
525
  ---
402
526
 
403
- ## ⚡ Mjölnir is not another linter
404
-
405
- Linters tell you whether code follows rules. Mjölnir tells you whether your
406
- verification can be trusted.
407
-
408
- | | ESLint / SonarQube | Coverage tools | Manual review | **Mjölnir** |
409
- | -------------------------------------------------------- | :----------------: | :------------: | :-----------: | :---------: |
410
- | CI workflow integrity (`continue-on-error`, `\|\| true`) | ❌ | ❌ | rarely | ✅ |
411
- | Cross-language (TS, Python, Java, C#) from one tool | ❌ | ❌ | ❌ | ✅ |
412
- | Grades Playwright locator resilience (Selector Health) | ❌ | ❌ | rarely | ✅ |
413
- | Flags tests with no real assertions | ✅ (plugin)\* | ❌ | sometimes | ✅ |
414
- | Catches hard sleeps (`waitForTimeout`, `time.sleep`) | ✅ (plugin)\* | ❌ | sometimes | ✅ |
415
- | Runs in seconds, zero network calls while scanning | ✅ | ✅ | — | ✅ |
416
-
417
- \*`eslint-plugin-jest` (`expect-expect`) and `eslint-plugin-playwright`
418
- (`expect-expect`, `no-wait-for-timeout`) cover these for their respective
419
- frameworks.
420
-
421
- **Runtime analysis** is a separate category from static linting:
422
-
423
- | | Playwright retry reporter | Allure / ReportPortal | **Mjölnir forensics** |
424
- | --------------------------------------------- | :-----------------------: | :-------------------: | :-------------------: |
425
- | Reads real run data for `TRUE-FLAKE` verdicts | partial\* | partial (tag) | ✅ |
426
- | Flaky-triage report from execution history | ❌ | ✅ | ✅ |
427
- | Integrates with static worthiness score | ❌ | ❌ | ✅ |
428
-
429
- \*Playwright tracks retries internally but does not produce a standalone
430
- flakiness report with verdict labels.
431
-
432
- ---
433
-
434
- ## 🤖 Why not just use AI code review?
435
-
436
- Different problem, different layer. AI review can spot a suspicious test
437
- change in a diff; it does not prove the verification system as a whole is
438
- trustworthy — and it only sees the diff you show it.
439
-
440
- | | AI code review (Copilot, etc.) | **Mjölnir** |
441
- | ------------------------------------- | :----------------------------: | :-------------------------------: |
442
- | Cost per scan | Tokens (scales with diff size) | **Zero** (local, installed) |
443
- | Sees the whole suite + all CI configs | Only the PR diff you show it | **Everything, every time** |
444
- | Deterministic (same input → same out) | ❌ (non-deterministic) | **✅** |
445
- | Catches patterns dormant for months | Only if it's in the context | **✅** (scans all files) |
446
- | Remembers findings between runs | ❌ (no memory across sessions) | **✅** (baseline + diff) |
447
- | Runs without human triggering | Needs a PR or prompt | **✅** (CI hook, runs in seconds) |
448
-
449
- **Use both.** AI catches nuance, intent, and design flaws no regex can
450
- find. Mjölnir catches the structural patterns AI overlooks because they
451
- look "intentional" — a committed `.only`, a swallowed exit code, a
452
- `continue-on-error` on a test job. Those aren't bugs that need reasoning;
453
- they're facts that need scanning.
454
-
455
- ---
456
-
457
527
  ## 🤖 CI integration
458
528
 
459
529
  One command generates a PR workflow — advisory by default, never blocking:
@@ -485,46 +555,46 @@ ref with `--base <ref>`.
485
555
 
486
556
  ---
487
557
 
488
- ## Configuration
558
+ ## 🧠 Works with your agent
489
559
 
490
- Mjölnir is zero-config. An optional `mjolnir.config.json` (or
491
- `.mjolnir.json`) at the repo root tunes severity, gating and scope — it
492
- never changes detection semantics.
560
+ Findings are only worth something if something acts on them. Mjölnir hands
561
+ its evidence to a coding agent three ways — none of them "paste the
562
+ terminal output and hope".
493
563
 
494
- | Key | Type | Effect |
495
- | ------------------- | ------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
496
- | `exclude` | `string[]` | Extra ignore globs (gitignore subset), on top of the built-in defaults |
497
- | `gate` | `"advisory" \| "error" \| "warning"` | Which severities exit non-zero (default `error`; `advisory` never blocks) |
498
- | `severityOverrides` | `{ "<RULE-ID>": severity }` | Re-rank a rule's findings for your repo |
499
- | `ignore` | `IgnoreEntry[]` | Suppress findings — **`reason` is required**; entries expire after 90 days (an explicit `expires` date, or the config file's last-modified time for entries without one) |
500
- | `plugins` | `string[]` | Third-party rule packages (see [Trust model](#trust-model)) |
564
+ | Command | What the agent gets |
565
+ | ----------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
566
+ | `mjolnir mcp` | Runs as an [MCP](https://modelcontextprotocol.io) server over stdio — `scan`, `explain` and `diff` become tools Claude Code, Claude Desktop or any MCP client can call. |
567
+ | `mjolnir handoff` | A saved `--json` report becomes a deterministic Markdown remediation plan: what was detected, the evidence boundary per finding, what must **not** change, and how to verify. |
568
+ | `mjolnir install` | Writes the agent instruction surfaces your repo already has (e.g. `.claude/commands/`), so the agent knows to re-scan before it claims it's done. |
569
+
570
+ Add the MCP server to Claude Code:
571
+
572
+ ```bash
573
+ claude mcp add mjolnir -- npx -y mjolnir-qa@latest mcp
574
+ ```
575
+
576
+ Or to any client that takes an `mcpServers` block:
501
577
 
502
578
  ```json
503
579
  {
504
- "gate": "error",
505
- "exclude": ["legacy/**"],
506
- "severityOverrides": { "QA-PW-118": "warning" },
507
- "ignore": [
508
- {
509
- "ruleId": "QA-TEST-004",
510
- "files": ["e2e/legacy-login.spec.ts"],
511
- "reason": "Third-party widget needs a settle delay; tracked in JIRA-4821",
512
- "expires": "2026-12-31"
513
- }
514
- ]
580
+ "mcpServers": {
581
+ "mjolnir": { "command": "npx", "args": ["-y", "mjolnir-qa@latest", "mcp"] }
582
+ }
515
583
  }
516
584
  ```
517
585
 
518
- - **`.mjolnirignore`** — a plain gitignore-style file for path exclusions,
519
- same dialect as `exclude`. Use it for machine-wide noise; use `exclude`
520
- when the list belongs in version control alongside the rest of the config.
521
- - **CLI overrides** — `--strict` (include quarantine rules), `--width <cols>`
522
- and `--ascii` / `--no-ascii` (terminal rendering), `--tone blunt`
523
- (blunter messages), `--max-duration <sec>` (bounded partial scan).
524
- - Rule suppression and deprecation lifecycle: [docs/RULE-LIFECYCLE.md](docs/RULE-LIFECYCLE.md).
586
+ Everything a machine consumes — MCP tool results, `--json`, SARIF — comes
587
+ off one canonical result under a versioned, additive-only schema, so a
588
+ consumer never has to reconstruct semantics for itself:
589
+ [the machine contract](docs/machine-contract.md) (`contractVersion: 1`).
525
590
 
526
- `ignore` entries also power the standalone `mjolnir suppressions` command,
527
- which lists what's currently suppressed and when each entry expires.
591
+ **The guardrail matters more than the convenience.** Every finding in a
592
+ handoff carries its evidence boundary: **E2** says _deterministic — check
593
+ the location and apply the fix_; **E1** says _REQUIRES CONFIRMATION — the
594
+ observation alone does not prove the defect_. An agent that "fixes" E1
595
+ findings blindly, or suppresses a rule to get a green scan, is doing the
596
+ exact thing this tool exists to catch — so the artifact says so, in the
597
+ prompt, next to the finding.
528
598
 
529
599
  ---
530
600
 
@@ -545,77 +615,25 @@ are immutable once shipped and never reused.
545
615
 
546
616
  ---
547
617
 
548
- ## Trust model
549
-
550
- - **Local-first** — zero network calls during scanning. Ever. Zero telemetry.
551
- - **No false proof** — we'd rather say "unknown" than "verified". An empty
552
- repo gets `score: null`, never a fake 100.
553
- - **Partial honesty** — if analysis was cut short, the output says so.
554
- Never "complete" when it isn't.
555
- - **FP firewall** — detection runs on a comment/string-free view of the code
556
- (TypeScript rules use the compiler AST): a pattern inside a prose comment
557
- or a doc-example string is documentation, not a finding.
558
- - **Measured, not asserted** — only rules with a false-positive rate from
559
- real OSS code ship in the headline tiers (see
560
- [How much of this is measured](#how-much-of-this-is-measured)); the scan
561
- footer and `mjolnir rules --unmeasured` tell you which is which.
562
- - **Plugin trust & execution gate** — plugins are npm packages declared
563
- under `"plugins"`; JS modules live in `mjolnir-rules/*.mjs`. There is
564
- **no sandbox**: plugin code runs with full Node privileges, the same
565
- trust model as ESLint or Vitest plugins. Because of that, code execution
566
- is **opt-in at every scan**: pass `--enable-plugins` (or set
567
- `MJOLNIR_ENABLE_PLUGINS=1`) or the sources are NOT loaded — a loud
568
- stderr notice lists exactly what was skipped. Scanning untrusted code
569
- never executes it. JSON rule manifests (`mjolnir-rules/*.json`) are
570
- unaffected: they declare regex patterns and execute no code by design.
571
- Core rule-ID prefixes are reserved and rejected from plugins and
572
- external rules to prevent spoofing.
573
- - **Workspace-local external rules** (folder-based, zero network) — a
574
- `mjolnir-rules/` directory next to the scan target loads custom rules:
575
- JSON files declare regex patterns (no code executed), `.mjs`/`.js`
576
- modules export `rules` (full-Node trust, same as plugins). External
577
- rules carry the same trust metadata as core; they can never ship in
578
- the core tier (core requires a measured FP rate from the corpus
579
- sidecar — a declared `tier: "core"` is clamped to `extended`), obey
580
- tier caps, and are drift-checked: `mjolnir rules --md --external`
581
- renders the catalog from the loaded files (provenance `external`),
582
- and the matrix generator accepts `--external <root>`.
583
-
584
- ---
618
+ ## 🔒 Trust model
585
619
 
586
- ## 🏗️ Architecture
620
+ **Local-first, zero telemetry, no false proof.** Scanning untrusted code
621
+ never executes it.
587
622
 
588
- <details>
589
- <summary>Expand tree</summary>
590
-
591
- ```
592
- mjolnir/
593
- ├── src/
594
- │ ├── engine/ # LanguageAdapter interface + rule runner
595
- │ ├── adapters/ # typescript · python · java · csharp · github-actions
596
- │ ├── rules/ # rules across 8 families + the measured-FP table
597
- │ ├── playwright/ # Selector Health Score engine
598
- │ ├── discovery/ # workspace, frameworks, ignore resolution
599
- │ ├── scope/ # git merge-base changed-scope engine
600
- │ ├── scorer/ # transparent deduction table + prioritization
601
- │ ├── reporter/ # terminal · JSON · SARIF 2.1 · Mermaid
602
- │ ├── forensics/ # run-data ingestion · flake verdicts · triage
603
- │ ├── config/ # mjolnir.config.json + suppressions
604
- │ ├── plugins/ # third-party rule loading (no sandbox)
605
- │ └── commands/ # every subcommand
606
- └── tests/
607
- ├── fixtures/ # must-fire / must-not-fire per rule
608
- └── golden/ # frozen score regression locks
609
- ```
610
-
611
- </details>
623
+ The one part that changes how you invoke the tool: **plugins are not
624
+ sandboxed.** JS plugins (`mjolnir-rules/*.mjs`, or npm packages under
625
+ `"plugins"`) run with full Node privileges — the same trust model as
626
+ ESLint or Vitest plugins — so loading them is opt-in **per scan**: without
627
+ `--enable-plugins` (or `MJOLNIR_ENABLE_PLUGINS=1`) the sources are not
628
+ loaded, and a stderr notice lists exactly what was skipped. JSON rule
629
+ manifests declare regex patterns and execute no code by design. Core
630
+ rule-ID prefixes are reserved, so a plugin cannot impersonate one.
612
631
 
613
- - **Rules are pure functions** — `(SourceFileContext) → Finding[]`, no I/O,
614
- no globals. Adding an ecosystem = one adapter + its rules.
615
- - **TypeScript/Playwright uses the compiler AST** (ts-morph). Python, Java
616
- and C# run on a shared comment/string-masked regex layer.
617
- - A tree-sitter WASM AST layer for Java and C# exists and is the next
618
- precision step — it is not yet wired into the synchronous scan pipeline.
632
+ Scoring math, rule lifecycle, architecture and the tree-sitter roadmap:
633
+ [docs/SCORING.md](docs/SCORING.md) ·
634
+ [docs/RULE-LIFECYCLE.md](docs/RULE-LIFECYCLE.md) ·
635
+ [CONTRIBUTING.md](CONTRIBUTING.md) ·
636
+ [docs site](https://sergey-bar.github.io/Mjolnir/).
619
637
 
620
638
  ---
621
639
 
@@ -641,7 +659,7 @@ mjolnir/
641
659
 
642
660
  **v0.5.x · open beta.** The JSON schema and exit codes are frozen contracts.
643
661
  TypeScript and Python have the broadest measured coverage; Java and C# are
644
- newer — read them through the [maturity table](#rule-tiers-and-language-maturity).
662
+ newer — read them through the [maturity table](https://sergey-bar.github.io/Mjolnir/reference/rule-lifecycle).
645
663
  Honest scope, no invented dates: the [public roadmap](https://sergey-bar.github.io/Mjolnir/reference/roadmap).
646
664
 
647
665
  ---