mjolnir-qa 0.5.13 → 0.5.15

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -74,6 +74,25 @@ scan`. Unknown categories are a usage error (exit 10).
74
74
  - help registry gained `why`, `handoff`, `install` and the new flags;
75
75
  site/reference/cli.md documents the handoff trust model.
76
76
 
77
+ ## [0.5.15] — 2026-09-07
78
+
79
+ ### Changes since 0.5.14
80
+
81
+ - Close-out: complete the 21-rule fixture program (blueprint §19) (#50)
82
+
83
+ ## [0.5.14] — 2026-09-07
84
+
85
+ ### Changes since 0.5.13
86
+
87
+ - Merge pull request #47 from Sergey-Bar/claude/readme-demo-video-4rzxij
88
+ - Cover the mjolnir mcp dispatch branch — CI's 100% ratchet caught it
89
+ - README: define Selector Health, state the limits, lock the samples
90
+ - Expose the MCP server, and document the agent surface in the README
91
+ - See it work: embed the real demo video inline
92
+ - Rework See it work and the score section: real video, fixed-size cards
93
+ - Replace the See it work poster+MP4 and shorten the score hero image
94
+ - Replace the score/verdict table with an animated hammer sweep
95
+
77
96
  ## [0.5.13] — 2026-09-07
78
97
 
79
98
  ### Changes since 0.5.12
package/README.md CHANGED
@@ -18,7 +18,7 @@ npx mjolnir-qa@latest
18
18
 
19
19
  **Are your tests worthy of trust?**
20
20
 
21
- [See it work](#-see-it-work) · [Quickstart](#-quickstart) · [Who it's for](#-who-is-this-for) · [Why not a linter](#-mjölnir-is-not-another-linter) · [What it verifies](#-what-mjölnir-verifies) · [Scoring](#-how-the-score-works) · [Runtime evidence](#-runtime-evidence) · [CI](#-ci-integration) · [Exit codes](#-exit-codes--contracts) · [Docs](#-documentation) · [Contributing](#-contributing)
21
+ [See it work](#-see-it-work) · [Quickstart](#-quickstart) · [Who it's for](#-who-is-this-for) · [Why not a linter](#-mjölnir-is-not-another-linter) · [What it verifies](#-what-mjölnir-verifies) · [Scoring](#-how-the-score-works) · [Runtime evidence](#-runtime-evidence) · [CI](#-ci-integration) · [Agents & MCP](#-works-with-your-agent) · [Docs](#-documentation) · [Contributing](#-contributing)
22
22
 
23
23
  <details>
24
24
  <summary>Read this in another language — 22 translations</summary>
@@ -36,19 +36,28 @@ it; `npm run docs:translations` reports how far.
36
36
 
37
37
  ## 🎬 See it work
38
38
 
39
+ <!-- Plays inline on github.com only: <video> is rendered for GitHub's own
40
+ user-content CDN, never for a repo-relative path. The <a> below is the
41
+ fallback for every other renderer (npm, mirrors, offline clones). -->
39
42
  <p align="center">
40
- <a href="assets/video/mjolnir-demo.mp4">
41
- <img src="assets/video/mjolnir-demo-poster.png" alt="Mjölnir scanning a demo repo: the hammer instrument at [STRAINED] and WORTHINESS 75/100 NEEDS WORK" width="900" />
42
- </a>
43
+ <video
44
+ src="https://github.com/user-attachments/assets/0e1af1e4-1e27-4c1c-9ec4-2717d194df05"
45
+ poster="https://raw.githubusercontent.com/Sergey-Bar/Mjolnir/main/assets/video/mjolnir-demo-poster.png"
46
+ controls
47
+ muted
48
+ playsinline
49
+ width="900"></video>
43
50
  </p>
44
51
 
45
52
  <p align="center">
46
- <strong><a href="assets/video/mjolnir-demo.mp4">▶ Watch the 42-second demo</a></strong> —
47
- one false-green CI gate: found, fixed, and re-proved.
53
+ <strong>42 seconds</strong> — one false-green CI gate: found, fixed,
54
+ re-proved, then handed to an agent.
55
+ <a href="assets/video/mjolnir-demo.mp4">Download the 1440p MP4</a> if the
56
+ player above doesn't load.
48
57
  </p>
49
58
 
50
- <sub>Every frame is real CLI output. The 75 → 90 score change is a real
51
- re-scan after applying the fix the tool itself printed — never a mockup.
59
+ <sub>Every frame is real CLI output — the 75 → 90 score change is a real
60
+ re-scan after applying the fix the tool itself printed, never a mockup.
52
61
  Rendered by `npm run docs:video` from
53
62
  [`assets/video/script.demo.json`](assets/video/script.demo.json);
54
63
  [`tests/contract/video-script.spec.ts`](tests/contract/video-script.spec.ts)
@@ -79,14 +88,15 @@ fails CI if it drifts.</sub>
79
88
  brittle selector, hardcoded staging URLs, a `networkidle` wait.
80
89
  3. It turned each into a concrete finding with a rule ID, a location and a
81
90
  fix — and a single score you can gate a PR on.
91
+ 4. `mjolnir handoff` turned the findings into a remediation plan — evidence,
92
+ constraints and a copy-pastable prompt per finding — that Claude Code
93
+ (or any other agent) can work through, with the tool's own verification
94
+ discipline built in.
82
95
 
83
- There's also an 89-second tour covering `explain`, `forensics`, and the
84
- rest of the walkthrough below — same pipeline, same guarantee (every
85
- frame is real CLI output). It's built as
86
- [`assets/video/script.tour.json`](assets/video/script.tour.json) but not
87
- committed as an MP4 (it's ~16MB; every clone shouldn't pay for a video
88
- most readers won't open) — run `npm run docs:video` to render it, or
89
- check the repo's [Releases](../../releases) for a published copy.
96
+ An 89-second tour of `explain` and `forensics` is built from
97
+ [`script.tour.json`](assets/video/script.tour.json) by `npm run docs:video`
98
+ — not committed as an MP4 (~16MB), so check
99
+ [Releases](../../releases) or render it yourself.
90
100
 
91
101
  ### One finding, up close
92
102
 
@@ -153,6 +163,7 @@ and you're done. Everything else is optional.
153
163
  | `mjolnir --scope changed` | Only what your branch introduced — the CI form |
154
164
  | `mjolnir ci install` | Generate the advisory PR workflow |
155
165
  | `mjolnir explain QA-CI-001` | What / why / fix + measured FP rate for one rule |
166
+ | `mjolnir why src/a.spec.ts:42` | Why this exact line was flagged — never a gate |
156
167
  | `mjolnir rules --unmeasured` | The rules running on assumption, not measurement |
157
168
  | `mjolnir --json` / `--format sarif` | Machine-readable / GitHub Code Scanning |
158
169
  | `mjolnir --strict` | Also run quarantine-tier rules (higher FP risk) |
@@ -173,19 +184,23 @@ and you're done. Everything else is optional.
173
184
  <details>
174
185
  <summary><strong>Occasional / reporting</strong></summary>
175
186
 
176
- | Command | What it does |
177
- | ------------------------------- | ------------------------------------------------ |
178
- | `mjolnir fix --dry-run` / `fix` | Safe auto-fixes with proof |
179
- | `mjolnir baseline` / `diff` | Snapshot findings, then report only new/worsened |
180
- | `mjolnir impact --since <ref>` | What changed since a prior commit |
181
- | `mjolnir debt` | Test-debt register with a cost model |
182
- | `mjolnir handover` | New-QA onboarding map of the suite |
183
- | `mjolnir stats` | Local all-time counters of fixes seen |
184
- | `mjolnir badge` | shields.io endpoint JSON + snippet |
185
- | `mjolnir rules --md` | Full rule catalog (JSON or Markdown) |
186
- | `mjolnir doctor` | Self-audit of Mjölnir's own rule base |
187
- | `mjolnir create-rule <ID>` | Scaffold a new rule + fixtures |
188
- | `mjolnir --format mermaid` | Test-architecture diagram for a PR comment |
187
+ | Command | What it does |
188
+ | ------------------------------- | ------------------------------------------------------ |
189
+ | `mjolnir fix --dry-run` / `fix` | Safe auto-fixes with proof |
190
+ | `mjolnir baseline` / `diff` | Snapshot findings, then report only new/worsened |
191
+ | `mjolnir impact --since <ref>` | What changed since a prior commit |
192
+ | `mjolnir summary` | CI annotations + step summary from a saved report |
193
+ | `mjolnir pr-comment` | A scoped PR comment, as Markdown |
194
+ | `mjolnir debt` | Test-debt register with a cost model |
195
+ | `mjolnir handover` | New-QA onboarding map of the suite |
196
+ | `mjolnir init` | Detect frameworks + setup checklist (never overwrites) |
197
+ | `mjolnir suppressions` | List suppressed findings — governance transparency |
198
+ | `mjolnir stats` | Local all-time counters of fixes seen |
199
+ | `mjolnir badge` | shields.io endpoint JSON + snippet |
200
+ | `mjolnir rules --md` | Full rule catalog (JSON or Markdown) |
201
+ | `mjolnir doctor` | Self-audit of Mjölnir's own rule base |
202
+ | `mjolnir create-rule <ID>` | Scaffold a new rule + fixtures |
203
+ | `mjolnir --format mermaid` | Test-architecture diagram for a PR comment |
189
204
 
190
205
  </details>
191
206
 
@@ -366,6 +381,29 @@ or `mjolnir rules --md`.
366
381
  >
367
382
  > Per-rule pages live under [`docs/rules/`](docs/rules/).
368
383
 
384
+ ### Selector Health Score
385
+
386
+ A pass rate says the suite went green today. It says nothing about whether
387
+ it will still go green after a redesign. `mjolnir doctor:playwright` grades
388
+ every locator in a spec file by how it finds an element — the way a user
389
+ identifies it (role, label, text), an explicit contract (`data-testid`), or
390
+ a structural accident (CSS chains, XPath) — and scores the file 0–100:
391
+
392
+ ```text
393
+ ▚ SELECTOR HEALTH
394
+
395
+ e2e/login.spec.ts
396
+ [█████████████░░░░░░░] 65 / 100
397
+ role/text: 1 · testid: 0 · css-chains: 1 ⚠ · xpath: 0
398
+
399
+ e2e/checkout.spec.ts
400
+ [█████████████████░░░] 86 / 100
401
+ role/text: 3 · testid: 1 · css-chains: 1 ⚠ · xpath: 0
402
+ ```
403
+
404
+ `.btn.btn-primary > div:nth-child(2)` is not a test failure today — it's a
405
+ test failure scheduled for whenever someone touches the markup.
406
+
369
407
  ### How much of this is measured
370
408
 
371
409
  **78 of 99 rules carry a false-positive rate measured against real OSS code** (≥ 10 hand-classified findings each; see
@@ -383,33 +421,76 @@ Every rule is `core`, `extended`, or `quarantine`, assigned from its
383
421
  `--strict`. Tiers, language maturity and the promotion/demotion rules:
384
422
  [rule lifecycle](https://sergey-bar.github.io/Mjolnir/reference/rule-lifecycle).
385
423
 
424
+ ### What it can't tell you
425
+
426
+ A tool that claims your tests are lying should be blunt about its own
427
+ limits:
428
+
429
+ - **It does not run your tests.** It reads your code and, when you point
430
+ it at one, a run report. A clean scan is not a passing suite.
431
+ - **It cannot tell you an assertion is _wrong_.** `expect(total).toBe(41)`
432
+ is a perfectly healthy-looking test. Mjölnir finds tests that can't fail
433
+ and pipelines that can't go red — not tests that check the wrong thing.
434
+ - **A 100 is not proof of a good suite.** It means none of these 99 rules
435
+ fired. Coverage of your actual risk is a different question, and this
436
+ tool does not pretend to answer it.
437
+ - **21 of 99 rules ship on an estimate**, not a measured rate — and they
438
+ say so, per rule, in `mjolnir explain`.
439
+ - **E1 findings are heuristics.** They are positioned to be worth reading,
440
+ not to be applied blindly; the evidence level is attached to every
441
+ finding precisely so you can tell the difference.
442
+ - **An empty repo scores `null`, never 100.** "Unknown" is a verdict here.
443
+
444
+ ### We run it on ourselves
445
+
446
+ Every CI run scans this repository with the build produced by that same
447
+ run, and the gate fails on **any** error-severity finding — but also on a
448
+ _partial_ scan or a crashed rule, because a truncated self-scan that
449
+ reports nothing is exactly the false green this project exists to catch.
450
+ The result is uploaded as a build artifact on every run.
451
+
386
452
  ---
387
453
 
388
454
  ## 📊 How the score works
389
455
 
390
- <p align="center">
391
- <img src="assets/readme/terminal-hero.svg" alt="Mjölnir terminal output — WORTHINESS 75/100 NEEDS WORK, a diagnostics-by-category breakdown, and a FIX THIS FIRST list" width="820" />
392
- </p>
456
+ <table>
457
+ <tr>
458
+ <td width="46%" valign="top" align="center">
459
+
460
+ <img src="assets/readme/score-gauge.svg" alt="The hammer sweeping every score from 0 to 100 — cracked below 50 (UNWORTHY), strained 50-79 (NEEDS WORK), charged 80-99 (WORTHY), forged at 100 (FORGED) — then holding on FORGED before it loops" width="270" />
461
+
462
+ **What the score means**
463
+
464
+ </td>
465
+ <td width="54%" valign="top" align="center">
393
466
 
394
- <sub>Regenerated by `npm run docs:hero`;
395
- [`tests/contract/hero-asset-reproducibility.spec.ts`](tests/contract/hero-asset-reproducibility.spec.ts)
396
- fails CI if it drifts from what the reporter actually prints.</sub>
467
+ <img src="assets/readme/terminal-hero.svg" alt="Mjölnir's deduction breakdown — WORTHINESS 75/100 NEEDS WORK, a diagnostics-by-category bar chart, the per-severity deduction box, and a FIX THIS FIRST list" width="330" />
468
+
469
+ **Where the points went**
470
+
471
+ </td>
472
+ </tr>
473
+ </table>
474
+
475
+ <sub>Left: every score 0–100 through the real `deriveScoreState`/`renderHammer`
476
+ (`npm run docs:gauge`) — the hammer has four real states, so it snaps at the
477
+ exact thresholds rather than faking an in-between shape. Right: a real strict
478
+ scan of `examples/demo-repo` (`npm run docs:hero`), excerpted to the deduction
479
+ block. Both are drift-locked
480
+ ([gauge](tests/contract/score-gauge-asset-reproducibility.spec.ts) ·
481
+ [breakdown](tests/contract/hero-asset-reproducibility.spec.ts)).</sub>
397
482
 
398
483
  The score is transparent: **error −8, warning −3, info −1**, then normalized
399
484
  by suite exposure (deductions per test declaration). Evidence-weighted
400
485
  deductions mean weak signals cost less. The terminal shows the same
401
486
  discounted numbers the score uses — no black box.
402
487
 
403
- | Score | Verdict | | Level | Evidence | Score impact |
404
- | ------- | ---------------- | --- | ----- | -------------------- | ---------------- |
405
- | 100 | ⚡ **FORGED** | | E2 | Deterministic defect | Full deduction |
406
- | ≥ 80 | ✓ **WORTHY** | | E1 | Heuristic pattern | Half deduction |
407
- | 50 – 79 | ⚠ **NEEDS WORK** | | E0 | Observation | Zero (info only) |
408
- | < 50 | ✖ **UNWORTHY** | | | | |
409
-
410
- Most rules are **E1**. The tagline "we prove it" refers to this system:
411
- E2 findings are structural proof; E1 findings are correctly-positioned
412
- warnings, not formal proofs.
488
+ Each band also carries an evidence level: **100/FORGED** is **E2**
489
+ (deterministic defect, full deduction), **≥ 80/WORTHY** and **50–79/NEEDS
490
+ WORK** findings are mostly **E1** (heuristic pattern, half deduction), and
491
+ **E0** (observation) findings cost nothing — info only. The tagline "we
492
+ prove it" refers to this system: E2 findings are structural proof; E1
493
+ findings are correctly-positioned warnings, not formal proofs.
413
494
 
414
495
  **No false proof.** We'd rather say "unknown" than "verified" — an empty
415
496
  repo scores `null`, never a fake 100. Full method:
@@ -428,7 +509,7 @@ mjolnir forensics ./test-results/
428
509
  ```
429
510
 
430
511
  ```text
431
- ▚ FLAKINESS LEADERBOARD
512
+ ▚ FLAKINESS LEADERBOARD
432
513
 
433
514
  3 tests · 1 failed · 1 flaky · 1 retried
434
515
 
@@ -474,6 +555,49 @@ ref with `--base <ref>`.
474
555
 
475
556
  ---
476
557
 
558
+ ## 🧠 Works with your agent
559
+
560
+ Findings are only worth something if something acts on them. Mjölnir hands
561
+ its evidence to a coding agent three ways — none of them "paste the
562
+ terminal output and hope".
563
+
564
+ | Command | What the agent gets |
565
+ | ----------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
566
+ | `mjolnir mcp` | Runs as an [MCP](https://modelcontextprotocol.io) server over stdio — `scan`, `explain` and `diff` become tools Claude Code, Claude Desktop or any MCP client can call. |
567
+ | `mjolnir handoff` | A saved `--json` report becomes a deterministic Markdown remediation plan: what was detected, the evidence boundary per finding, what must **not** change, and how to verify. |
568
+ | `mjolnir install` | Writes the agent instruction surfaces your repo already has (e.g. `.claude/commands/`), so the agent knows to re-scan before it claims it's done. |
569
+
570
+ Add the MCP server to Claude Code:
571
+
572
+ ```bash
573
+ claude mcp add mjolnir -- npx -y mjolnir-qa@latest mcp
574
+ ```
575
+
576
+ Or to any client that takes an `mcpServers` block:
577
+
578
+ ```json
579
+ {
580
+ "mcpServers": {
581
+ "mjolnir": { "command": "npx", "args": ["-y", "mjolnir-qa@latest", "mcp"] }
582
+ }
583
+ }
584
+ ```
585
+
586
+ Everything a machine consumes — MCP tool results, `--json`, SARIF — comes
587
+ off one canonical result under a versioned, additive-only schema, so a
588
+ consumer never has to reconstruct semantics for itself:
589
+ [the machine contract](docs/machine-contract.md) (`contractVersion: 1`).
590
+
591
+ **The guardrail matters more than the convenience.** Every finding in a
592
+ handoff carries its evidence boundary: **E2** says _deterministic — check
593
+ the location and apply the fix_; **E1** says _REQUIRES CONFIRMATION — the
594
+ observation alone does not prove the defect_. An agent that "fixes" E1
595
+ findings blindly, or suppresses a rule to get a green scan, is doing the
596
+ exact thing this tool exists to catch — so the artifact says so, in the
597
+ prompt, next to the finding.
598
+
599
+ ---
600
+
477
601
  ## 📐 Exit codes & contracts
478
602
 
479
603
  Frozen — safe to build CI logic on:
@@ -491,27 +615,24 @@ are immutable once shipped and never reused.
491
615
 
492
616
  ---
493
617
 
494
- ## Trust model
495
-
496
- **Local-first, zero telemetry, no false proof** — full detail in
497
- [docs/SCORING.md](docs/SCORING.md) and
498
- [docs/RULE-LIFECYCLE.md](docs/RULE-LIFECYCLE.md). The one piece worth
499
- stating here because it changes how you invoke the tool:
500
-
501
- - **Plugin trust & execution gate** — plugins are npm packages declared
502
- under `"plugins"`; JS modules live in `mjolnir-rules/*.mjs`. There is
503
- **no sandbox**: plugin code runs with full Node privileges, the same
504
- trust model as ESLint or Vitest plugins. Because of that, code execution
505
- is **opt-in at every scan**: pass `--enable-plugins` (or set
506
- `MJOLNIR_ENABLE_PLUGINS=1`) or the sources are NOT loaded — a loud
507
- stderr notice lists exactly what was skipped. Scanning untrusted code
508
- never executes it. JSON rule manifests (`mjolnir-rules/*.json`) are
509
- unaffected: they declare regex patterns and execute no code by design.
510
- Core rule-ID prefixes are reserved and rejected from plugins and
511
- external rules to prevent spoofing.
512
-
513
- Architecture, the full rule catalog, and the tree-sitter roadmap live in
514
- [CONTRIBUTING.md](CONTRIBUTING.md) and the
618
+ ## 🔒 Trust model
619
+
620
+ **Local-first, zero telemetry, no false proof.** Scanning untrusted code
621
+ never executes it.
622
+
623
+ The one part that changes how you invoke the tool: **plugins are not
624
+ sandboxed.** JS plugins (`mjolnir-rules/*.mjs`, or npm packages under
625
+ `"plugins"`) run with full Node privileges — the same trust model as
626
+ ESLint or Vitest plugins — so loading them is opt-in **per scan**: without
627
+ `--enable-plugins` (or `MJOLNIR_ENABLE_PLUGINS=1`) the sources are not
628
+ loaded, and a stderr notice lists exactly what was skipped. JSON rule
629
+ manifests declare regex patterns and execute no code by design. Core
630
+ rule-ID prefixes are reserved, so a plugin cannot impersonate one.
631
+
632
+ Scoring math, rule lifecycle, architecture and the tree-sitter roadmap:
633
+ [docs/SCORING.md](docs/SCORING.md) ·
634
+ [docs/RULE-LIFECYCLE.md](docs/RULE-LIFECYCLE.md) ·
635
+ [CONTRIBUTING.md](CONTRIBUTING.md) ·
515
636
  [docs site](https://sergey-bar.github.io/Mjolnir/).
516
637
 
517
638
  ---
package/dist/cli.d.mts CHANGED
@@ -699,7 +699,7 @@ declare const runScan: typeof runScan$1, buildUniversalRules: typeof buildUniver
699
699
  * `scripts/sync-sarif-version.cjs` on release and guarded by
700
700
  * `tests/version-consistency.spec.ts` locally.
701
701
  */
702
- declare const CLI_VERSION = "0.5.13";
702
+ declare const CLI_VERSION = "0.5.15";
703
703
  /** A usage-error detail: the offending token, when one exists. */
704
704
  interface UsageErrorDetail {
705
705
  /** The unknown flag or rejected value (e.g. `--nope`, `loud`). */