mjolnir-qa 0.5.12 → 0.5.13
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +26 -0
- package/README.md +122 -225
- package/dist/cli.d.mts +1 -1
- package/dist/cli.mjs +81 -37
- package/dist/mcp/stdio.mjs +81 -37
- package/package.json +6 -1
package/CHANGELOG.md
CHANGED
|
@@ -74,6 +74,32 @@ scan`. Unknown categories are a usage error (exit 10).
|
|
|
74
74
|
- help registry gained `why`, `handoff`, `install` and the new flags;
|
|
75
75
|
site/reference/cli.md documents the handoff trust model.
|
|
76
76
|
|
|
77
|
+
## [0.5.13] — 2026-09-07
|
|
78
|
+
|
|
79
|
+
### Changes since 0.5.12
|
|
80
|
+
|
|
81
|
+
- Merge pull request #33 from Sergey-Bar/claude/readme-demo-video-4rzxij
|
|
82
|
+
- fix(video): resolve ffmpeg/ffprobe without a shell
|
|
83
|
+
- Resync generated assets after merging origin/main (v0.5.12)
|
|
84
|
+
- Merge remote-tracking branch 'origin/main' into claude/readme-demo-video-4rzxij
|
|
85
|
+
- merge: catch up to main again (v0.5.6) — a second PR landed underneath this one
|
|
86
|
+
- Merge remote-tracking branch 'origin/main' into claude/readme-demo-video-4rzxij
|
|
87
|
+
- fix: explain's fixture path was OS-native, breaking Windows CI
|
|
88
|
+
- merge: bring in main's UX overhaul, reconcile the reporter conflicts
|
|
89
|
+
- Merge remote-tracking branch 'origin/main' into claude/readme-demo-video-4rzxij
|
|
90
|
+
- video: adopt react.doctor's terminal palette and font
|
|
91
|
+
- video: fix invisible command text, and guard the whole class
|
|
92
|
+
- video: present the terminal as a window, not a maximised screenshot
|
|
93
|
+
- video: re-render both demos against the fixed reporter
|
|
94
|
+
- report: one hammer, and output that fits the terminal it prints to
|
|
95
|
+
- video: add the manual render workflow and document the pipeline
|
|
96
|
+
- docs: restructure the README around the demo, and lead with the video
|
|
97
|
+
- video: add the media-format contract, and ship the hero MP4
|
|
98
|
+
- video: render the committed scripts to 1440p H.264, frame by frame
|
|
99
|
+
- video: capture the demo scripts from real scans, and contract them
|
|
100
|
+
- video: vendor the render font stack, gated by a cmap-exact glyph probe
|
|
101
|
+
- docs: fix the stale numbers in the honesty section, and guard them
|
|
102
|
+
|
|
77
103
|
## [0.5.12] — 2026-09-07
|
|
78
104
|
|
|
79
105
|
### Changes since 0.5.11
|
package/README.md
CHANGED
|
@@ -12,15 +12,23 @@ pipelines, reports a worthiness score, and shows exactly where trust breaks.
|
|
|
12
12
|
[](LICENSE)
|
|
13
13
|
[](https://nodejs.org)
|
|
14
14
|
|
|
15
|
-
English | [简体中文](README.zh.md) | [繁體中文](README.zht.md) | [한국어](README.ko.md) | [Deutsch](README.de.md) | [Español](README.es.md) | [Français](README.fr.md) | [Italiano](README.it.md) | [Dansk](README.da.md) | [日本語](README.ja.md) | [Polski](README.pl.md) | [Русский](README.ru.md) | [Norsk](README.no.md) | [Português (Brasil)](README.br.md) | [ไทย](README.th.md) | [Türkçe](README.tr.md) | [Українська](README.uk.md) | [বাংলা](README.bn.md) | [Ελληνικά](README.gr.md) | [Tiếng Việt](README.vi.md) | [עברית](README.he.md) | [العربية](README.ar.md) | [Bosanski](README.bs.md)
|
|
16
|
-
|
|
17
15
|
```bash
|
|
18
16
|
npx mjolnir-qa@latest
|
|
19
17
|
```
|
|
20
18
|
|
|
21
19
|
**Are your tests worthy of trust?**
|
|
22
20
|
|
|
23
|
-
[See it work](#-see-it-work) · [Quickstart](#-quickstart) · [What it
|
|
21
|
+
[See it work](#-see-it-work) · [Quickstart](#-quickstart) · [Who it's for](#-who-is-this-for) · [Why not a linter](#-mjölnir-is-not-another-linter) · [What it verifies](#-what-mjölnir-verifies) · [Scoring](#-how-the-score-works) · [Runtime evidence](#-runtime-evidence) · [CI](#-ci-integration) · [Exit codes](#-exit-codes--contracts) · [Docs](#-documentation) · [Contributing](#-contributing)
|
|
22
|
+
|
|
23
|
+
<details>
|
|
24
|
+
<summary>Read this in another language — 22 translations</summary>
|
|
25
|
+
|
|
26
|
+
English | [简体中文](README.zh.md) | [繁體中文](README.zht.md) | [한국어](README.ko.md) | [Deutsch](README.de.md) | [Español](README.es.md) | [Français](README.fr.md) | [Italiano](README.it.md) | [Dansk](README.da.md) | [日本語](README.ja.md) | [Polski](README.pl.md) | [Русский](README.ru.md) | [Norsk](README.no.md) | [Português (Brasil)](README.br.md) | [ไทย](README.th.md) | [Türkçe](README.tr.md) | [Українська](README.uk.md) | [বাংলা](README.bn.md) | [Ελληνικά](README.gr.md) | [Tiếng Việt](README.vi.md) | [עברית](README.he.md) | [العربية](README.ar.md) | [Bosanski](README.bs.md)
|
|
27
|
+
|
|
28
|
+
English is canonical. Translations are machine-assisted and may lag behind
|
|
29
|
+
it; `npm run docs:translations` reports how far.
|
|
30
|
+
|
|
31
|
+
</details>
|
|
24
32
|
|
|
25
33
|
</div>
|
|
26
34
|
|
|
@@ -29,14 +37,38 @@ npx mjolnir-qa@latest
|
|
|
29
37
|
## 🎬 See it work
|
|
30
38
|
|
|
31
39
|
<p align="center">
|
|
32
|
-
<
|
|
40
|
+
<a href="assets/video/mjolnir-demo.mp4">
|
|
41
|
+
<img src="assets/video/mjolnir-demo-poster.png" alt="Mjölnir scanning a demo repo: the hammer instrument at [STRAINED] and WORTHINESS 75/100 NEEDS WORK" width="900" />
|
|
42
|
+
</a>
|
|
43
|
+
</p>
|
|
44
|
+
|
|
45
|
+
<p align="center">
|
|
46
|
+
<strong><a href="assets/video/mjolnir-demo.mp4">▶ Watch the 42-second demo</a></strong> —
|
|
47
|
+
one false-green CI gate: found, fixed, and re-proved.
|
|
48
|
+
</p>
|
|
49
|
+
|
|
50
|
+
<sub>Every frame is real CLI output. The 75 → 90 score change is a real
|
|
51
|
+
re-scan after applying the fix the tool itself printed — never a mockup.
|
|
52
|
+
Rendered by `npm run docs:video` from
|
|
53
|
+
[`assets/video/script.demo.json`](assets/video/script.demo.json);
|
|
54
|
+
[`tests/contract/video-script.spec.ts`](tests/contract/video-script.spec.ts)
|
|
55
|
+
fails CI if that script stops matching what the CLI prints, or if the
|
|
56
|
+
findings the video shows as fixed turn out to still be there.</sub>
|
|
57
|
+
|
|
58
|
+
<details>
|
|
59
|
+
<summary><strong>Prefer it inline?</strong> The full <code>--verbose</code> report, as an animated SVG</summary>
|
|
60
|
+
|
|
61
|
+
<p align="center">
|
|
62
|
+
<img src="assets/readme/demo.svg" alt="Mjölnir's full --verbose report on a demo repo: WORTHINESS 75/100 NEEDS WORK, a diagnostics-by-category breakdown, a FIX THIS FIRST list, and every finding with its rule ID and line number" width="900" />
|
|
33
63
|
</p>
|
|
34
64
|
|
|
35
65
|
<sub>The complete `npx mjolnir-qa ./examples/demo-repo --verbose` output,
|
|
36
66
|
rendered from the actual reporter — nothing trimmed. Regenerated by
|
|
37
67
|
`npm run docs:demo`;
|
|
38
|
-
[`tests/demo-asset-reproducibility.spec.ts`](tests/demo-asset-reproducibility.spec.ts)
|
|
39
|
-
fails CI if it drifts
|
|
68
|
+
[`tests/contract/demo-asset-reproducibility.spec.ts`](tests/contract/demo-asset-reproducibility.spec.ts)
|
|
69
|
+
fails CI if it drifts.</sub>
|
|
70
|
+
|
|
71
|
+
</details>
|
|
40
72
|
|
|
41
73
|
**What just happened:**
|
|
42
74
|
|
|
@@ -48,27 +80,48 @@ fails CI if it drifts from what the tool prints.</sub>
|
|
|
48
80
|
3. It turned each into a concrete finding with a rule ID, a location and a
|
|
49
81
|
fix — and a single score you can gate a PR on.
|
|
50
82
|
|
|
83
|
+
There's also an 89-second tour covering `explain`, `forensics`, and the
|
|
84
|
+
rest of the walkthrough below — same pipeline, same guarantee (every
|
|
85
|
+
frame is real CLI output). It's built as
|
|
86
|
+
[`assets/video/script.tour.json`](assets/video/script.tour.json) but not
|
|
87
|
+
committed as an MP4 (it's ~16MB; every clone shouldn't pay for a video
|
|
88
|
+
most readers won't open) — run `npm run docs:video` to render it, or
|
|
89
|
+
check the repo's [Releases](../../releases) for a published copy.
|
|
90
|
+
|
|
51
91
|
### One finding, up close
|
|
52
92
|
|
|
53
93
|
Run `mjolnir explain QA-CI-001` on the first finding above and you get:
|
|
54
94
|
|
|
55
95
|
```text
|
|
56
|
-
▚ QA-CI-001 — continue-on-error masks a failing verification gate
|
|
96
|
+
▚ QA-CI-001 — continue-on-error masks a failing verification gate
|
|
57
97
|
|
|
58
98
|
Severity: error
|
|
59
99
|
Confidence: high
|
|
100
|
+
Tier: quarantine
|
|
60
101
|
Evidence: E2
|
|
61
|
-
|
|
102
|
+
QA impact: False-green risk (FALSE-GREEN)
|
|
103
|
+
Measured FP: 11% (19 hand-classified corpus verdicts)
|
|
104
|
+
FP risk: low (author estimate)
|
|
105
|
+
Languages: yaml
|
|
106
|
+
Frameworks: github-actions
|
|
62
107
|
|
|
63
108
|
WHAT WAS FOUND (real detector output, not a mockup)
|
|
64
109
|
Job `security-scan` runs a verification gate under `continue-on-error: true`.
|
|
65
110
|
|
|
66
111
|
WHY IT MATTERS
|
|
67
|
-
This job can fail every day and CI will still show green. The checkmark
|
|
68
|
-
|
|
112
|
+
This job can fail every day and CI will still show green. The checkmark on
|
|
113
|
+
this workflow cannot be trusted.
|
|
69
114
|
|
|
70
115
|
HOW TO FIX
|
|
71
116
|
Remove continue-on-error, or scope it to individual non-blocking steps only.
|
|
117
|
+
|
|
118
|
+
Example from this rule's own must-fire fixture: QA-CI-001/must-fire/masked.yml
|
|
119
|
+
|
|
120
|
+
HOW TO VERIFY THE FIX
|
|
121
|
+
Re-run `mjolnir` on the changed file(s) — this finding should no longer
|
|
122
|
+
appear. `mjolnir --scope changed` scopes the check to just what you touched.
|
|
123
|
+
|
|
124
|
+
Docs: mjolnir rules --md (full catalog, this rule included)
|
|
72
125
|
```
|
|
73
126
|
|
|
74
127
|
That is the unit of value: not a style nit, but a place where your CI is
|
|
@@ -153,7 +206,35 @@ Requires Node.js ≥ 22.18. Works on Windows, macOS, and Linux.
|
|
|
153
206
|
|
|
154
207
|
---
|
|
155
208
|
|
|
156
|
-
##
|
|
209
|
+
## ⚡ Mjölnir is not another linter
|
|
210
|
+
|
|
211
|
+
Linters tell you whether code follows rules. Mjölnir tells you whether your
|
|
212
|
+
verification can be trusted.
|
|
213
|
+
|
|
214
|
+
| | ESLint / SonarQube | Coverage tools | AI code review | **Mjölnir** |
|
|
215
|
+
| -------------------------------------------------------- | :----------------: | :------------: | :------------: | :--------------: |
|
|
216
|
+
| CI workflow integrity (`continue-on-error`, `\|\| true`) | ❌ | ❌ | only the diff | ✅ |
|
|
217
|
+
| Cross-language (TS, Python, Java, C#) from one tool | ❌ | ❌ | ❌ | ✅ |
|
|
218
|
+
| Grades Playwright locator resilience (Selector Health) | ❌ | ❌ | ❌ | ✅ |
|
|
219
|
+
| Flags tests with no real assertions | ✅ (plugin)\* | ❌ | sometimes | ✅ |
|
|
220
|
+
| Catches hard sleeps (`waitForTimeout`, `time.sleep`) | ✅ (plugin)\* | ❌ | sometimes | ✅ |
|
|
221
|
+
| Reads real run data for `TRUE-FLAKE` verdicts | ❌ | ❌ | ❌ | ✅ |
|
|
222
|
+
| Deterministic (same input → same output) | ✅ | ✅ | ❌ | ✅ |
|
|
223
|
+
| Cost per scan | free | free | tokens | **zero** (local) |
|
|
224
|
+
|
|
225
|
+
\*`eslint-plugin-jest` (`expect-expect`) and `eslint-plugin-playwright`
|
|
226
|
+
(`expect-expect`, `no-wait-for-timeout`) cover these for their respective
|
|
227
|
+
frameworks.
|
|
228
|
+
|
|
229
|
+
**Use AI review too.** It catches nuance, intent, and design flaws no regex
|
|
230
|
+
can find. Mjölnir catches the structural patterns AI overlooks because they
|
|
231
|
+
look "intentional" — a committed `.only`, a swallowed exit code, a
|
|
232
|
+
`continue-on-error` on a test job. Those aren't bugs that need reasoning;
|
|
233
|
+
they're facts that need scanning.
|
|
234
|
+
|
|
235
|
+
---
|
|
236
|
+
|
|
237
|
+
## 🔨 What Mjölnir verifies
|
|
157
238
|
|
|
158
239
|
| | |
|
|
159
240
|
| --- | ----------------------------------------------------------------------------------------------------------------- |
|
|
@@ -164,12 +245,15 @@ Requires Node.js ≥ 22.18. Works on Windows, macOS, and Linux.
|
|
|
164
245
|
| 🐍 | **All four Playwright bindings** — TypeScript, Python, Java, C#/.NET — plus pytest, JUnit/TestNG and CI workflows |
|
|
165
246
|
| 🔒 | **Local-first** — zero network calls while scanning, zero telemetry, runs in seconds |
|
|
166
247
|
|
|
167
|
-
### The rules
|
|
168
|
-
|
|
169
248
|
Every rule ships with must-fire **and** must-not-fire fixtures. A rule that
|
|
170
249
|
fires on its own negative fixture cannot ship — that's the false-positive
|
|
171
250
|
firewall.
|
|
172
251
|
|
|
252
|
+
**The rule catalog.** Every family is collapsed below; the generated
|
|
253
|
+
full catalog lives in [`docs/rules/`](docs/rules/),
|
|
254
|
+
[what it checks](https://sergey-bar.github.io/Mjolnir/guide/what-it-checks),
|
|
255
|
+
or `mjolnir rules --md`.
|
|
256
|
+
|
|
173
257
|
<details>
|
|
174
258
|
<summary><strong>Test Hygiene</strong></summary>
|
|
175
259
|
|
|
@@ -295,83 +379,42 @@ Growing that number is the project's continuing work.
|
|
|
295
379
|
### Rule tiers and language maturity
|
|
296
380
|
|
|
297
381
|
Every rule is `core`, `extended`, or `quarantine`, assigned from its
|
|
298
|
-
**measured** false-positive rate
|
|
299
|
-
|
|
300
|
-
|
|
301
|
-
| ------------ | ---------------------------------------- | :----------: | :--------: |
|
|
302
|
-
| `core` | ≤ 10 % measured FP | ✅ | ✅ |
|
|
303
|
-
| `extended` | ≤ 30 % measured FP | ✅ | ✅ |
|
|
304
|
-
| `quarantine` | above 30 %, or not yet measured (n < 10) | ❌ | ✅ |
|
|
305
|
-
|
|
306
|
-
| Language | Adapter | Coverage today |
|
|
307
|
-
| --------------- | ------------ | -------------------------------------------------- |
|
|
308
|
-
| TypeScript / JS | compiler AST | broadest, most measured — mostly `core`/`extended` |
|
|
309
|
-
| Python / pytest | regex layer | broad, corpus-audited — mostly `core`/`extended` |
|
|
310
|
-
| Java | regex layer | newer — mostly `extended`/`quarantine` |
|
|
311
|
-
| C# / .NET | regex layer | newer — mostly `extended`/`quarantine` |
|
|
312
|
-
|
|
313
|
-
TypeScript and Python have the broadest measured coverage. Java and C# ship,
|
|
314
|
-
are documented, and stay out of the headline number until a real consumer
|
|
315
|
-
suite (not a binding library's own tests) has been audited.
|
|
382
|
+
**measured** false-positive rate — quarantine rules only run under
|
|
383
|
+
`--strict`. Tiers, language maturity and the promotion/demotion rules:
|
|
384
|
+
[rule lifecycle](https://sergey-bar.github.io/Mjolnir/reference/rule-lifecycle).
|
|
316
385
|
|
|
317
386
|
---
|
|
318
387
|
|
|
319
|
-
## How the score works
|
|
388
|
+
## 📊 How the score works
|
|
320
389
|
|
|
321
390
|
<p align="center">
|
|
322
391
|
<img src="assets/readme/terminal-hero.svg" alt="Mjölnir terminal output — WORTHINESS 75/100 NEEDS WORK, a diagnostics-by-category breakdown, and a FIX THIS FIRST list" width="820" />
|
|
323
392
|
</p>
|
|
324
393
|
|
|
325
394
|
<sub>Regenerated by `npm run docs:hero`;
|
|
326
|
-
[`tests/hero-asset-reproducibility.spec.ts`](tests/hero-asset-reproducibility.spec.ts)
|
|
395
|
+
[`tests/contract/hero-asset-reproducibility.spec.ts`](tests/contract/hero-asset-reproducibility.spec.ts)
|
|
327
396
|
fails CI if it drifts from what the reporter actually prints.</sub>
|
|
328
397
|
|
|
329
398
|
The score is transparent: **error −8, warning −3, info −1**, then normalized
|
|
330
399
|
by suite exposure (deductions per test declaration). Evidence-weighted
|
|
331
400
|
deductions mean weak signals cost less. The terminal shows the same
|
|
332
|
-
discounted numbers the score uses — no black box.
|
|
333
|
-
[docs/SCORING.md](docs/SCORING.md).
|
|
334
|
-
|
|
335
|
-
**Verdicts**
|
|
401
|
+
discounted numbers the score uses — no black box.
|
|
336
402
|
|
|
337
|
-
| Score | Verdict |
|
|
338
|
-
| ------- | ---------------- |
|
|
339
|
-
| 100 | ⚡ **FORGED** |
|
|
340
|
-
| ≥ 80 | ✓ **WORTHY** |
|
|
341
|
-
| 50 – 79 | ⚠ **NEEDS WORK** |
|
|
342
|
-
| < 50 | ✖ **UNWORTHY** |
|
|
343
|
-
|
|
344
|
-
**Evidence levels** — every finding carries one; it sets the finding's
|
|
345
|
-
weight in the score:
|
|
346
|
-
|
|
347
|
-
| Level | Meaning | Score impact | Example |
|
|
348
|
-
| ----- | -------------------- | ---------------- | -------------------------------------------------- |
|
|
349
|
-
| E2 | Deterministic defect | Full deduction | `.only` committed — structurally provable |
|
|
350
|
-
| E1 | Heuristic pattern | Half deduction | Regex-matched `sleep()` — strong signal, not proof |
|
|
351
|
-
| E0 | Observation | Zero (info only) | Reported but never gates CI or deducts |
|
|
403
|
+
| Score | Verdict | | Level | Evidence | Score impact |
|
|
404
|
+
| ------- | ---------------- | --- | ----- | -------------------- | ---------------- |
|
|
405
|
+
| 100 | ⚡ **FORGED** | | E2 | Deterministic defect | Full deduction |
|
|
406
|
+
| ≥ 80 | ✓ **WORTHY** | | E1 | Heuristic pattern | Half deduction |
|
|
407
|
+
| 50 – 79 | ⚠ **NEEDS WORK** | | E0 | Observation | Zero (info only) |
|
|
408
|
+
| < 50 | ✖ **UNWORTHY** | | | | |
|
|
352
409
|
|
|
353
410
|
Most rules are **E1**. The tagline "we prove it" refers to this system:
|
|
354
411
|
E2 findings are structural proof; E1 findings are correctly-positioned
|
|
355
412
|
warnings, not formal proofs.
|
|
356
413
|
|
|
357
|
-
|
|
358
|
-
|
|
359
|
-
|
|
360
|
-
|
|
361
|
-
## 🎭 Selector Health Score
|
|
362
|
-
|
|
363
|
-
The headline metric for Playwright suites — how resilient your locators are:
|
|
364
|
-
|
|
365
|
-
```text
|
|
366
|
-
▚ SELECTOR HEALTH — e2e/checkout.spec.ts
|
|
367
|
-
|
|
368
|
-
[█████████████████░░░] 83 / 100
|
|
369
|
-
role/text: 2 · testid: 1 · css-chains: 1 ⚠ · xpath: 0
|
|
370
|
-
```
|
|
371
|
-
|
|
372
|
-
Role-based locators score full credit. CSS class chains and XPath tank the
|
|
373
|
-
score — they break on any DOM refactor without telling you which behavior
|
|
374
|
-
regressed.
|
|
414
|
+
**No false proof.** We'd rather say "unknown" than "verified" — an empty
|
|
415
|
+
repo scores `null`, never a fake 100. Full method:
|
|
416
|
+
[docs/SCORING.md](docs/SCORING.md) ·
|
|
417
|
+
[scoring guide](https://sergey-bar.github.io/Mjolnir/guide/scoring).
|
|
375
418
|
|
|
376
419
|
---
|
|
377
420
|
|
|
@@ -400,60 +443,6 @@ test. It gets flagged `TRUE-FLAKE` regardless of the final green checkmark.
|
|
|
400
443
|
|
|
401
444
|
---
|
|
402
445
|
|
|
403
|
-
## ⚡ Mjölnir is not another linter
|
|
404
|
-
|
|
405
|
-
Linters tell you whether code follows rules. Mjölnir tells you whether your
|
|
406
|
-
verification can be trusted.
|
|
407
|
-
|
|
408
|
-
| | ESLint / SonarQube | Coverage tools | Manual review | **Mjölnir** |
|
|
409
|
-
| -------------------------------------------------------- | :----------------: | :------------: | :-----------: | :---------: |
|
|
410
|
-
| CI workflow integrity (`continue-on-error`, `\|\| true`) | ❌ | ❌ | rarely | ✅ |
|
|
411
|
-
| Cross-language (TS, Python, Java, C#) from one tool | ❌ | ❌ | ❌ | ✅ |
|
|
412
|
-
| Grades Playwright locator resilience (Selector Health) | ❌ | ❌ | rarely | ✅ |
|
|
413
|
-
| Flags tests with no real assertions | ✅ (plugin)\* | ❌ | sometimes | ✅ |
|
|
414
|
-
| Catches hard sleeps (`waitForTimeout`, `time.sleep`) | ✅ (plugin)\* | ❌ | sometimes | ✅ |
|
|
415
|
-
| Runs in seconds, zero network calls while scanning | ✅ | ✅ | — | ✅ |
|
|
416
|
-
|
|
417
|
-
\*`eslint-plugin-jest` (`expect-expect`) and `eslint-plugin-playwright`
|
|
418
|
-
(`expect-expect`, `no-wait-for-timeout`) cover these for their respective
|
|
419
|
-
frameworks.
|
|
420
|
-
|
|
421
|
-
**Runtime analysis** is a separate category from static linting:
|
|
422
|
-
|
|
423
|
-
| | Playwright retry reporter | Allure / ReportPortal | **Mjölnir forensics** |
|
|
424
|
-
| --------------------------------------------- | :-----------------------: | :-------------------: | :-------------------: |
|
|
425
|
-
| Reads real run data for `TRUE-FLAKE` verdicts | partial\* | partial (tag) | ✅ |
|
|
426
|
-
| Flaky-triage report from execution history | ❌ | ✅ | ✅ |
|
|
427
|
-
| Integrates with static worthiness score | ❌ | ❌ | ✅ |
|
|
428
|
-
|
|
429
|
-
\*Playwright tracks retries internally but does not produce a standalone
|
|
430
|
-
flakiness report with verdict labels.
|
|
431
|
-
|
|
432
|
-
---
|
|
433
|
-
|
|
434
|
-
## 🤖 Why not just use AI code review?
|
|
435
|
-
|
|
436
|
-
Different problem, different layer. AI review can spot a suspicious test
|
|
437
|
-
change in a diff; it does not prove the verification system as a whole is
|
|
438
|
-
trustworthy — and it only sees the diff you show it.
|
|
439
|
-
|
|
440
|
-
| | AI code review (Copilot, etc.) | **Mjölnir** |
|
|
441
|
-
| ------------------------------------- | :----------------------------: | :-------------------------------: |
|
|
442
|
-
| Cost per scan | Tokens (scales with diff size) | **Zero** (local, installed) |
|
|
443
|
-
| Sees the whole suite + all CI configs | Only the PR diff you show it | **Everything, every time** |
|
|
444
|
-
| Deterministic (same input → same out) | ❌ (non-deterministic) | **✅** |
|
|
445
|
-
| Catches patterns dormant for months | Only if it's in the context | **✅** (scans all files) |
|
|
446
|
-
| Remembers findings between runs | ❌ (no memory across sessions) | **✅** (baseline + diff) |
|
|
447
|
-
| Runs without human triggering | Needs a PR or prompt | **✅** (CI hook, runs in seconds) |
|
|
448
|
-
|
|
449
|
-
**Use both.** AI catches nuance, intent, and design flaws no regex can
|
|
450
|
-
find. Mjölnir catches the structural patterns AI overlooks because they
|
|
451
|
-
look "intentional" — a committed `.only`, a swallowed exit code, a
|
|
452
|
-
`continue-on-error` on a test job. Those aren't bugs that need reasoning;
|
|
453
|
-
they're facts that need scanning.
|
|
454
|
-
|
|
455
|
-
---
|
|
456
|
-
|
|
457
446
|
## 🤖 CI integration
|
|
458
447
|
|
|
459
448
|
One command generates a PR workflow — advisory by default, never blocking:
|
|
@@ -485,49 +474,6 @@ ref with `--base <ref>`.
|
|
|
485
474
|
|
|
486
475
|
---
|
|
487
476
|
|
|
488
|
-
## Configuration
|
|
489
|
-
|
|
490
|
-
Mjölnir is zero-config. An optional `mjolnir.config.json` (or
|
|
491
|
-
`.mjolnir.json`) at the repo root tunes severity, gating and scope — it
|
|
492
|
-
never changes detection semantics.
|
|
493
|
-
|
|
494
|
-
| Key | Type | Effect |
|
|
495
|
-
| ------------------- | ------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
|
496
|
-
| `exclude` | `string[]` | Extra ignore globs (gitignore subset), on top of the built-in defaults |
|
|
497
|
-
| `gate` | `"advisory" \| "error" \| "warning"` | Which severities exit non-zero (default `error`; `advisory` never blocks) |
|
|
498
|
-
| `severityOverrides` | `{ "<RULE-ID>": severity }` | Re-rank a rule's findings for your repo |
|
|
499
|
-
| `ignore` | `IgnoreEntry[]` | Suppress findings — **`reason` is required**; entries expire after 90 days (an explicit `expires` date, or the config file's last-modified time for entries without one) |
|
|
500
|
-
| `plugins` | `string[]` | Third-party rule packages (see [Trust model](#trust-model)) |
|
|
501
|
-
|
|
502
|
-
```json
|
|
503
|
-
{
|
|
504
|
-
"gate": "error",
|
|
505
|
-
"exclude": ["legacy/**"],
|
|
506
|
-
"severityOverrides": { "QA-PW-118": "warning" },
|
|
507
|
-
"ignore": [
|
|
508
|
-
{
|
|
509
|
-
"ruleId": "QA-TEST-004",
|
|
510
|
-
"files": ["e2e/legacy-login.spec.ts"],
|
|
511
|
-
"reason": "Third-party widget needs a settle delay; tracked in JIRA-4821",
|
|
512
|
-
"expires": "2026-12-31"
|
|
513
|
-
}
|
|
514
|
-
]
|
|
515
|
-
}
|
|
516
|
-
```
|
|
517
|
-
|
|
518
|
-
- **`.mjolnirignore`** — a plain gitignore-style file for path exclusions,
|
|
519
|
-
same dialect as `exclude`. Use it for machine-wide noise; use `exclude`
|
|
520
|
-
when the list belongs in version control alongside the rest of the config.
|
|
521
|
-
- **CLI overrides** — `--strict` (include quarantine rules), `--width <cols>`
|
|
522
|
-
and `--ascii` / `--no-ascii` (terminal rendering), `--tone blunt`
|
|
523
|
-
(blunter messages), `--max-duration <sec>` (bounded partial scan).
|
|
524
|
-
- Rule suppression and deprecation lifecycle: [docs/RULE-LIFECYCLE.md](docs/RULE-LIFECYCLE.md).
|
|
525
|
-
|
|
526
|
-
`ignore` entries also power the standalone `mjolnir suppressions` command,
|
|
527
|
-
which lists what's currently suppressed and when each entry expires.
|
|
528
|
-
|
|
529
|
-
---
|
|
530
|
-
|
|
531
477
|
## 📐 Exit codes & contracts
|
|
532
478
|
|
|
533
479
|
Frozen — safe to build CI logic on:
|
|
@@ -547,18 +493,11 @@ are immutable once shipped and never reused.
|
|
|
547
493
|
|
|
548
494
|
## Trust model
|
|
549
495
|
|
|
550
|
-
|
|
551
|
-
|
|
552
|
-
|
|
553
|
-
|
|
554
|
-
|
|
555
|
-
- **FP firewall** — detection runs on a comment/string-free view of the code
|
|
556
|
-
(TypeScript rules use the compiler AST): a pattern inside a prose comment
|
|
557
|
-
or a doc-example string is documentation, not a finding.
|
|
558
|
-
- **Measured, not asserted** — only rules with a false-positive rate from
|
|
559
|
-
real OSS code ship in the headline tiers (see
|
|
560
|
-
[How much of this is measured](#how-much-of-this-is-measured)); the scan
|
|
561
|
-
footer and `mjolnir rules --unmeasured` tell you which is which.
|
|
496
|
+
**Local-first, zero telemetry, no false proof** — full detail in
|
|
497
|
+
[docs/SCORING.md](docs/SCORING.md) and
|
|
498
|
+
[docs/RULE-LIFECYCLE.md](docs/RULE-LIFECYCLE.md). The one piece worth
|
|
499
|
+
stating here because it changes how you invoke the tool:
|
|
500
|
+
|
|
562
501
|
- **Plugin trust & execution gate** — plugins are npm packages declared
|
|
563
502
|
under `"plugins"`; JS modules live in `mjolnir-rules/*.mjs`. There is
|
|
564
503
|
**no sandbox**: plugin code runs with full Node privileges, the same
|
|
@@ -570,52 +509,10 @@ are immutable once shipped and never reused.
|
|
|
570
509
|
unaffected: they declare regex patterns and execute no code by design.
|
|
571
510
|
Core rule-ID prefixes are reserved and rejected from plugins and
|
|
572
511
|
external rules to prevent spoofing.
|
|
573
|
-
- **Workspace-local external rules** (folder-based, zero network) — a
|
|
574
|
-
`mjolnir-rules/` directory next to the scan target loads custom rules:
|
|
575
|
-
JSON files declare regex patterns (no code executed), `.mjs`/`.js`
|
|
576
|
-
modules export `rules` (full-Node trust, same as plugins). External
|
|
577
|
-
rules carry the same trust metadata as core; they can never ship in
|
|
578
|
-
the core tier (core requires a measured FP rate from the corpus
|
|
579
|
-
sidecar — a declared `tier: "core"` is clamped to `extended`), obey
|
|
580
|
-
tier caps, and are drift-checked: `mjolnir rules --md --external`
|
|
581
|
-
renders the catalog from the loaded files (provenance `external`),
|
|
582
|
-
and the matrix generator accepts `--external <root>`.
|
|
583
|
-
|
|
584
|
-
---
|
|
585
|
-
|
|
586
|
-
## 🏗️ Architecture
|
|
587
|
-
|
|
588
|
-
<details>
|
|
589
|
-
<summary>Expand tree</summary>
|
|
590
|
-
|
|
591
|
-
```
|
|
592
|
-
mjolnir/
|
|
593
|
-
├── src/
|
|
594
|
-
│ ├── engine/ # LanguageAdapter interface + rule runner
|
|
595
|
-
│ ├── adapters/ # typescript · python · java · csharp · github-actions
|
|
596
|
-
│ ├── rules/ # rules across 8 families + the measured-FP table
|
|
597
|
-
│ ├── playwright/ # Selector Health Score engine
|
|
598
|
-
│ ├── discovery/ # workspace, frameworks, ignore resolution
|
|
599
|
-
│ ├── scope/ # git merge-base changed-scope engine
|
|
600
|
-
│ ├── scorer/ # transparent deduction table + prioritization
|
|
601
|
-
│ ├── reporter/ # terminal · JSON · SARIF 2.1 · Mermaid
|
|
602
|
-
│ ├── forensics/ # run-data ingestion · flake verdicts · triage
|
|
603
|
-
│ ├── config/ # mjolnir.config.json + suppressions
|
|
604
|
-
│ ├── plugins/ # third-party rule loading (no sandbox)
|
|
605
|
-
│ └── commands/ # every subcommand
|
|
606
|
-
└── tests/
|
|
607
|
-
├── fixtures/ # must-fire / must-not-fire per rule
|
|
608
|
-
└── golden/ # frozen score regression locks
|
|
609
|
-
```
|
|
610
|
-
|
|
611
|
-
</details>
|
|
612
512
|
|
|
613
|
-
|
|
614
|
-
|
|
615
|
-
|
|
616
|
-
and C# run on a shared comment/string-masked regex layer.
|
|
617
|
-
- A tree-sitter WASM AST layer for Java and C# exists and is the next
|
|
618
|
-
precision step — it is not yet wired into the synchronous scan pipeline.
|
|
513
|
+
Architecture, the full rule catalog, and the tree-sitter roadmap live in
|
|
514
|
+
[CONTRIBUTING.md](CONTRIBUTING.md) and the
|
|
515
|
+
[docs site](https://sergey-bar.github.io/Mjolnir/).
|
|
619
516
|
|
|
620
517
|
---
|
|
621
518
|
|
|
@@ -641,7 +538,7 @@ mjolnir/
|
|
|
641
538
|
|
|
642
539
|
**v0.5.x · open beta.** The JSON schema and exit codes are frozen contracts.
|
|
643
540
|
TypeScript and Python have the broadest measured coverage; Java and C# are
|
|
644
|
-
newer — read them through the [maturity table](
|
|
541
|
+
newer — read them through the [maturity table](https://sergey-bar.github.io/Mjolnir/reference/rule-lifecycle).
|
|
645
542
|
Honest scope, no invented dates: the [public roadmap](https://sergey-bar.github.io/Mjolnir/reference/roadmap).
|
|
646
543
|
|
|
647
544
|
---
|
package/dist/cli.d.mts
CHANGED
|
@@ -699,7 +699,7 @@ declare const runScan: typeof runScan$1, buildUniversalRules: typeof buildUniver
|
|
|
699
699
|
* `scripts/sync-sarif-version.cjs` on release and guarded by
|
|
700
700
|
* `tests/version-consistency.spec.ts` locally.
|
|
701
701
|
*/
|
|
702
|
-
declare const CLI_VERSION = "0.5.
|
|
702
|
+
declare const CLI_VERSION = "0.5.13";
|
|
703
703
|
/** A usage-error detail: the offending token, when one exists. */
|
|
704
704
|
interface UsageErrorDetail {
|
|
705
705
|
/** The unknown flag or rejected value (e.g. `--nope`, `loud`). */
|
package/dist/cli.mjs
CHANGED
|
@@ -12480,26 +12480,26 @@ function topFixes(findings, n = 3) {
|
|
|
12480
12480
|
}
|
|
12481
12481
|
//#endregion
|
|
12482
12482
|
//#region src/reporter/art.ts
|
|
12483
|
-
/**
|
|
12483
|
+
/**
|
|
12484
|
+
* The report's wordmark. A LOCKUP, not a picture of a hammer.
|
|
12485
|
+
*
|
|
12486
|
+
* This used to be eight rows of box-drawing that spelled out a hammer —
|
|
12487
|
+
* directly above the score instrument, which is also a hammer. Two
|
|
12488
|
+
* hammers, ten lines apart, and the smaller static one came first, so
|
|
12489
|
+
* appendScoreSection's own claim that "the hammer is the instrument: the
|
|
12490
|
+
* first thing the eye lands on" was false in every report the tool has
|
|
12491
|
+
* ever printed.
|
|
12492
|
+
*
|
|
12493
|
+
* There is one hammer now, and it is the one that means something: the
|
|
12494
|
+
* state-resolved instrument in HAMMER_STATES. The wordmark stays out of
|
|
12495
|
+
* its way.
|
|
12496
|
+
*/
|
|
12484
12497
|
const LOGO = `
|
|
12485
|
-
|
|
12486
|
-
║ ║
|
|
12487
|
-
╠═══════════╣ M J Ö L N I R
|
|
12488
|
-
║ ║ ║
|
|
12489
|
-
╚═════╩═════╝ VERIFICATION TRUST ENGINE
|
|
12490
|
-
║
|
|
12491
|
-
║
|
|
12498
|
+
M J Ö L N I R · VERIFICATION TRUST ENGINE
|
|
12492
12499
|
`;
|
|
12493
|
-
/** Plain-ASCII fallback
|
|
12494
|
-
* block-drawing LOGO above renders as mangled "?" glyphs. */
|
|
12500
|
+
/** Plain-ASCII fallback for cmd.exe/legacy consoles. */
|
|
12495
12501
|
const LOGO_ASCII = `
|
|
12496
|
-
|
|
12497
|
-
| |
|
|
12498
|
-
+-----------+ M J O L N I R
|
|
12499
|
-
| | |
|
|
12500
|
-
+-----+-----+ VERIFICATION TRUST ENGINE
|
|
12501
|
-
|
|
|
12502
|
-
|
|
|
12502
|
+
M J O L N I R - VERIFICATION TRUST ENGINE
|
|
12503
12503
|
`;
|
|
12504
12504
|
const TROPHY = String.raw`
|
|
12505
12505
|
___________
|
|
@@ -13000,7 +13000,12 @@ const CARD_GUTTER = " ";
|
|
|
13000
13000
|
function pushCard(lines, card, ui) {
|
|
13001
13001
|
const { p, width } = ui;
|
|
13002
13002
|
const contentWidth = Math.max(20, width - 2 - 4 - CARD_LABEL_PAD);
|
|
13003
|
-
|
|
13003
|
+
const header = ` ${severityIcon(card.severity, ui)} ${p.bold(card.loc)}`;
|
|
13004
|
+
if (measure(`${header} ${card.evidence}`) <= width) lines.push(`${header} ${p.dim(card.evidence)}`);
|
|
13005
|
+
else {
|
|
13006
|
+
lines.push(header);
|
|
13007
|
+
lines.push(`${CARD_GUTTER}${p.dim(card.evidence)}`);
|
|
13008
|
+
}
|
|
13004
13009
|
const fields = [
|
|
13005
13010
|
{
|
|
13006
13011
|
label: "Finding",
|
|
@@ -13040,7 +13045,7 @@ function pushCard(lines, card, ui) {
|
|
|
13040
13045
|
* overflow line; --verbose shows everything.
|
|
13041
13046
|
*/
|
|
13042
13047
|
function appendFindings(lines, result, counts, verbose, ui, tone) {
|
|
13043
|
-
const { p } = ui;
|
|
13048
|
+
const { p, width } = ui;
|
|
13044
13049
|
if (counts.total === 0) return;
|
|
13045
13050
|
if (result.findings.length === 0) {
|
|
13046
13051
|
lines.push(ui.p.dim(" filtered view: no findings in the selected category"));
|
|
@@ -13089,9 +13094,21 @@ function appendFindings(lines, result, counts, verbose, ui, tone) {
|
|
|
13089
13094
|
hiddenRules.add(unit.ruleId);
|
|
13090
13095
|
continue;
|
|
13091
13096
|
}
|
|
13092
|
-
|
|
13093
|
-
|
|
13094
|
-
|
|
13097
|
+
const groupHead = ` ${severityIcon(maxSeverity(unit.findings), ui)} ${p.bold(sanitizeData(unit.ruleId))} ${p.dim(`× ${n} — same fix applies`)}`;
|
|
13098
|
+
const groupEvidence = evidenceTag$1(first);
|
|
13099
|
+
if (measure(`${groupHead} ${groupEvidence}`) <= width) lines.push(`${groupHead} ${p.dim(groupEvidence)}`);
|
|
13100
|
+
else {
|
|
13101
|
+
lines.push(groupHead);
|
|
13102
|
+
lines.push(`${CARD_GUTTER}${p.dim(groupEvidence)}`);
|
|
13103
|
+
}
|
|
13104
|
+
const groupContentWidth = Math.max(20, width - 2 - 4 - CARD_LABEL_PAD);
|
|
13105
|
+
wrapLines(sanitizeData(first.fix), groupContentWidth).forEach((seg, i) => {
|
|
13106
|
+
const label = i === 0 ? p.accent("Fix".padEnd(CARD_LABEL_PAD)) : " ".repeat(CARD_LABEL_PAD);
|
|
13107
|
+
lines.push(`${CARD_GUTTER}${label}${p.dim(seg)}`);
|
|
13108
|
+
});
|
|
13109
|
+
for (const f of unit.findings) wrapLines(`· ${sanitizeData(f.file)}:${f.line} — ${sanitizeData(f.message)}`, groupContentWidth).forEach((seg, i) => {
|
|
13110
|
+
lines.push(`${CARD_GUTTER}${" ".repeat(CARD_LABEL_PAD)}${p.dim(i === 0 ? seg : ` ${seg}`)}`);
|
|
13111
|
+
});
|
|
13095
13112
|
lines.push("");
|
|
13096
13113
|
shown++;
|
|
13097
13114
|
continue;
|
|
@@ -13129,22 +13146,34 @@ function appendForgedBlock(lines, p, ascii) {
|
|
|
13129
13146
|
lines.push(p.forged(TROPHY));
|
|
13130
13147
|
lines.push("");
|
|
13131
13148
|
}
|
|
13149
|
+
/**
|
|
13150
|
+
* Pushes dimmed prose that respects the terminal width.
|
|
13151
|
+
*
|
|
13152
|
+
* The honesty footer used to be pushed as single unbroken strings — the
|
|
13153
|
+
* rule-coverage line alone is ~147 columns, so it overflowed every
|
|
13154
|
+
* default 80- or 100-column terminal and ignored `--width` entirely.
|
|
13155
|
+
* `wrapText` is the same helper the finding cards already use.
|
|
13156
|
+
*/
|
|
13157
|
+
function pushWrapped(lines, p, text, width) {
|
|
13158
|
+
const indent = " ";
|
|
13159
|
+
for (const line of wrapText(text, Math.max(20, width - 2))) lines.push(p.dim(`${indent}${line}`));
|
|
13160
|
+
}
|
|
13132
13161
|
function appendFooter(lines, result, ui) {
|
|
13133
|
-
const { p } = ui;
|
|
13162
|
+
const { p, width } = ui;
|
|
13134
13163
|
lines.push(...buildFooter({
|
|
13135
13164
|
ui,
|
|
13136
13165
|
complete: result.analysisStatus.discovery !== "partial",
|
|
13137
13166
|
durationMs: result.analysisStatus.durationMs
|
|
13138
13167
|
}));
|
|
13139
13168
|
const advisory = result.findings.filter((f) => (f.evidenceLevel ?? deriveEvidenceLevel(f.findingType, f.confidence)) === "E0").length;
|
|
13140
|
-
if (advisory > 0) lines
|
|
13169
|
+
if (advisory > 0) pushWrapped(lines, p, `${advisory} advisory finding${advisory === 1 ? "" : "s"} (E0 — observation only, no score impact)`, width);
|
|
13141
13170
|
if (result.findings.length > 0) {
|
|
13142
13171
|
const firedRuleIds = new Set(result.findings.map((f) => f.ruleId));
|
|
13143
13172
|
const measuredHere = [...firedRuleIds].filter((id) => MEASURED_FP[id] !== void 0).length;
|
|
13144
|
-
lines
|
|
13173
|
+
pushWrapped(lines, p, `Rule coverage: ${measuredHere}/${firedRuleIds.size} rules that fired here have a measured false-positive rate; the rest are heuristics. \`mjolnir rules --unmeasured\` lists them.`, width);
|
|
13145
13174
|
const verified = result.findings.filter((f) => f.runtimeCorroboration !== void 0).length;
|
|
13146
|
-
if (verified > 0) lines
|
|
13147
|
-
else lines
|
|
13175
|
+
if (verified > 0) pushWrapped(lines, p, `Runtime evidence: ${verified}/${result.findings.length} findings corroborated by a real run report (trust L3–L5); the rest are static-only.`, width);
|
|
13176
|
+
else pushWrapped(lines, p, `Runtime evidence: not available — no run report (mjolnir.report.json / test-results) next to the scan target; all findings are static-only (L0–L2).`, width);
|
|
13148
13177
|
}
|
|
13149
13178
|
const profile = result.agenticProfile;
|
|
13150
13179
|
if (profile && (profile.generatedMarkedFiles > 0 || profile.codegenLikeFiles > 0)) {
|
|
@@ -13224,7 +13253,7 @@ function renderSarif(result, repoRootUri) {
|
|
|
13224
13253
|
tool: { driver: {
|
|
13225
13254
|
name: "Mjölnir",
|
|
13226
13255
|
informationUri: "https://github.com/Sergey-Bar/Mjolnir",
|
|
13227
|
-
version: "0.5.
|
|
13256
|
+
version: "0.5.13",
|
|
13228
13257
|
rules: [...rules.values()].map((r) => {
|
|
13229
13258
|
const meta = RULES.find((x) => x.id === r.id);
|
|
13230
13259
|
return {
|
|
@@ -17456,14 +17485,29 @@ function explainRule(ruleId, fixturesRoot) {
|
|
|
17456
17485
|
ok: true,
|
|
17457
17486
|
rule,
|
|
17458
17487
|
exampleFinding: example,
|
|
17459
|
-
exampleFixturePath: fixturePath
|
|
17488
|
+
exampleFixturePath: fixturePath,
|
|
17489
|
+
exampleFixtureRelPath: relative(fixturesRoot, fixturePath).replaceAll("\\", "/")
|
|
17460
17490
|
};
|
|
17461
17491
|
}
|
|
17462
|
-
|
|
17492
|
+
/**
|
|
17493
|
+
* Default column budget when no width is supplied.
|
|
17494
|
+
*
|
|
17495
|
+
* `explain`'s prose used to be pushed as unbroken strings — the
|
|
17496
|
+
* "HOW TO VERIFY THE FIX" paragraph alone is 150 columns — so every
|
|
17497
|
+
* explanation overflowed a default terminal. Renderers here take a width
|
|
17498
|
+
* rather than reading process.stdout, so output stays a pure function of
|
|
17499
|
+
* its arguments (same rule the reporter's palette follows).
|
|
17500
|
+
*/
|
|
17501
|
+
const DEFAULT_EXPLAIN_WIDTH = 80;
|
|
17502
|
+
function renderExplain(result, width = DEFAULT_EXPLAIN_WIDTH) {
|
|
17463
17503
|
if (!result.ok || !result.rule) return `explain failed: ${result.error ?? "unknown error"}`;
|
|
17464
17504
|
const r = result.rule;
|
|
17465
17505
|
const evidenceLevel = r.evidenceLevel ?? deriveEvidenceLevel(r.findingType, r.confidence);
|
|
17466
17506
|
const lines = [];
|
|
17507
|
+
/** Pushes prose indented two columns, wrapped to the budget. */
|
|
17508
|
+
const pushBody = (text) => {
|
|
17509
|
+
for (const seg of wrapText(text, Math.max(20, width - 2))) lines.push(` ${seg}`);
|
|
17510
|
+
};
|
|
17467
17511
|
lines.push(sectionHeader(`${r.id} — ${r.title}`, ui$1));
|
|
17468
17512
|
lines.push("");
|
|
17469
17513
|
lines.push(`Severity: ${r.severity}`);
|
|
@@ -17480,19 +17524,19 @@ function renderExplain(result) {
|
|
|
17480
17524
|
if (result.exampleFinding) {
|
|
17481
17525
|
const f = result.exampleFinding;
|
|
17482
17526
|
lines.push("WHAT WAS FOUND (real detector output, not a mockup)");
|
|
17483
|
-
|
|
17527
|
+
pushBody(f.message);
|
|
17484
17528
|
lines.push("");
|
|
17485
17529
|
lines.push("WHY IT MATTERS");
|
|
17486
|
-
|
|
17530
|
+
pushBody(f.why);
|
|
17487
17531
|
lines.push("");
|
|
17488
17532
|
lines.push("HOW TO FIX");
|
|
17489
|
-
|
|
17533
|
+
pushBody(f.fix);
|
|
17490
17534
|
lines.push("");
|
|
17491
|
-
|
|
17492
|
-
} else
|
|
17535
|
+
pushBody(`Example from this rule's own must-fire fixture: ${result.exampleFixtureRelPath ?? result.exampleFixturePath ?? "(unknown path)"}`);
|
|
17536
|
+
} else for (const seg of wrapText("No example available — run this command from a mjolnir checkout (or pass --fixtures-root) so the fixture that proves this rule works can be shown as a real example.", width)) lines.push(seg);
|
|
17493
17537
|
lines.push("");
|
|
17494
17538
|
lines.push("HOW TO VERIFY THE FIX");
|
|
17495
|
-
|
|
17539
|
+
pushBody("Re-run `mjolnir` on the changed file(s) — this finding should no longer appear. `mjolnir --scope changed` scopes the check to just what you touched.");
|
|
17496
17540
|
lines.push("");
|
|
17497
17541
|
lines.push(`Docs: mjolnir rules --md (full catalog, this rule included)`);
|
|
17498
17542
|
return lines.join("\n");
|
|
@@ -17649,7 +17693,7 @@ const { runScan, buildUniversalRules, fallbackWorkspace, pathMatchesGlob, isVali
|
|
|
17649
17693
|
* `scripts/sync-sarif-version.cjs` on release and guarded by
|
|
17650
17694
|
* `tests/version-consistency.spec.ts` locally.
|
|
17651
17695
|
*/
|
|
17652
|
-
const CLI_VERSION = "0.5.
|
|
17696
|
+
const CLI_VERSION = "0.5.13";
|
|
17653
17697
|
function parseArgs(argv, onError) {
|
|
17654
17698
|
const args = {
|
|
17655
17699
|
target: ".",
|
package/dist/mcp/stdio.mjs
CHANGED
|
@@ -12480,26 +12480,26 @@ function topFixes(findings, n = 3) {
|
|
|
12480
12480
|
}
|
|
12481
12481
|
//#endregion
|
|
12482
12482
|
//#region src/reporter/art.ts
|
|
12483
|
-
/**
|
|
12483
|
+
/**
|
|
12484
|
+
* The report's wordmark. A LOCKUP, not a picture of a hammer.
|
|
12485
|
+
*
|
|
12486
|
+
* This used to be eight rows of box-drawing that spelled out a hammer —
|
|
12487
|
+
* directly above the score instrument, which is also a hammer. Two
|
|
12488
|
+
* hammers, ten lines apart, and the smaller static one came first, so
|
|
12489
|
+
* appendScoreSection's own claim that "the hammer is the instrument: the
|
|
12490
|
+
* first thing the eye lands on" was false in every report the tool has
|
|
12491
|
+
* ever printed.
|
|
12492
|
+
*
|
|
12493
|
+
* There is one hammer now, and it is the one that means something: the
|
|
12494
|
+
* state-resolved instrument in HAMMER_STATES. The wordmark stays out of
|
|
12495
|
+
* its way.
|
|
12496
|
+
*/
|
|
12484
12497
|
const LOGO = `
|
|
12485
|
-
|
|
12486
|
-
║ ║
|
|
12487
|
-
╠═══════════╣ M J Ö L N I R
|
|
12488
|
-
║ ║ ║
|
|
12489
|
-
╚═════╩═════╝ VERIFICATION TRUST ENGINE
|
|
12490
|
-
║
|
|
12491
|
-
║
|
|
12498
|
+
M J Ö L N I R · VERIFICATION TRUST ENGINE
|
|
12492
12499
|
`;
|
|
12493
|
-
/** Plain-ASCII fallback
|
|
12494
|
-
* block-drawing LOGO above renders as mangled "?" glyphs. */
|
|
12500
|
+
/** Plain-ASCII fallback for cmd.exe/legacy consoles. */
|
|
12495
12501
|
const LOGO_ASCII = `
|
|
12496
|
-
|
|
12497
|
-
| |
|
|
12498
|
-
+-----------+ M J O L N I R
|
|
12499
|
-
| | |
|
|
12500
|
-
+-----+-----+ VERIFICATION TRUST ENGINE
|
|
12501
|
-
|
|
|
12502
|
-
|
|
|
12502
|
+
M J O L N I R - VERIFICATION TRUST ENGINE
|
|
12503
12503
|
`;
|
|
12504
12504
|
const TROPHY = String.raw`
|
|
12505
12505
|
___________
|
|
@@ -13000,7 +13000,12 @@ const CARD_GUTTER = " ";
|
|
|
13000
13000
|
function pushCard(lines, card, ui) {
|
|
13001
13001
|
const { p, width } = ui;
|
|
13002
13002
|
const contentWidth = Math.max(20, width - 2 - 4 - CARD_LABEL_PAD);
|
|
13003
|
-
|
|
13003
|
+
const header = ` ${severityIcon(card.severity, ui)} ${p.bold(card.loc)}`;
|
|
13004
|
+
if (measure(`${header} ${card.evidence}`) <= width) lines.push(`${header} ${p.dim(card.evidence)}`);
|
|
13005
|
+
else {
|
|
13006
|
+
lines.push(header);
|
|
13007
|
+
lines.push(`${CARD_GUTTER}${p.dim(card.evidence)}`);
|
|
13008
|
+
}
|
|
13004
13009
|
const fields = [
|
|
13005
13010
|
{
|
|
13006
13011
|
label: "Finding",
|
|
@@ -13040,7 +13045,7 @@ function pushCard(lines, card, ui) {
|
|
|
13040
13045
|
* overflow line; --verbose shows everything.
|
|
13041
13046
|
*/
|
|
13042
13047
|
function appendFindings(lines, result, counts, verbose, ui, tone) {
|
|
13043
|
-
const { p } = ui;
|
|
13048
|
+
const { p, width } = ui;
|
|
13044
13049
|
if (counts.total === 0) return;
|
|
13045
13050
|
if (result.findings.length === 0) {
|
|
13046
13051
|
lines.push(ui.p.dim(" filtered view: no findings in the selected category"));
|
|
@@ -13089,9 +13094,21 @@ function appendFindings(lines, result, counts, verbose, ui, tone) {
|
|
|
13089
13094
|
hiddenRules.add(unit.ruleId);
|
|
13090
13095
|
continue;
|
|
13091
13096
|
}
|
|
13092
|
-
|
|
13093
|
-
|
|
13094
|
-
|
|
13097
|
+
const groupHead = ` ${severityIcon(maxSeverity(unit.findings), ui)} ${p.bold(sanitizeData(unit.ruleId))} ${p.dim(`× ${n} — same fix applies`)}`;
|
|
13098
|
+
const groupEvidence = evidenceTag$1(first);
|
|
13099
|
+
if (measure(`${groupHead} ${groupEvidence}`) <= width) lines.push(`${groupHead} ${p.dim(groupEvidence)}`);
|
|
13100
|
+
else {
|
|
13101
|
+
lines.push(groupHead);
|
|
13102
|
+
lines.push(`${CARD_GUTTER}${p.dim(groupEvidence)}`);
|
|
13103
|
+
}
|
|
13104
|
+
const groupContentWidth = Math.max(20, width - 2 - 4 - CARD_LABEL_PAD);
|
|
13105
|
+
wrapLines(sanitizeData(first.fix), groupContentWidth).forEach((seg, i) => {
|
|
13106
|
+
const label = i === 0 ? p.accent("Fix".padEnd(CARD_LABEL_PAD)) : " ".repeat(CARD_LABEL_PAD);
|
|
13107
|
+
lines.push(`${CARD_GUTTER}${label}${p.dim(seg)}`);
|
|
13108
|
+
});
|
|
13109
|
+
for (const f of unit.findings) wrapLines(`· ${sanitizeData(f.file)}:${f.line} — ${sanitizeData(f.message)}`, groupContentWidth).forEach((seg, i) => {
|
|
13110
|
+
lines.push(`${CARD_GUTTER}${" ".repeat(CARD_LABEL_PAD)}${p.dim(i === 0 ? seg : ` ${seg}`)}`);
|
|
13111
|
+
});
|
|
13095
13112
|
lines.push("");
|
|
13096
13113
|
shown++;
|
|
13097
13114
|
continue;
|
|
@@ -13129,22 +13146,34 @@ function appendForgedBlock(lines, p, ascii) {
|
|
|
13129
13146
|
lines.push(p.forged(TROPHY));
|
|
13130
13147
|
lines.push("");
|
|
13131
13148
|
}
|
|
13149
|
+
/**
|
|
13150
|
+
* Pushes dimmed prose that respects the terminal width.
|
|
13151
|
+
*
|
|
13152
|
+
* The honesty footer used to be pushed as single unbroken strings — the
|
|
13153
|
+
* rule-coverage line alone is ~147 columns, so it overflowed every
|
|
13154
|
+
* default 80- or 100-column terminal and ignored `--width` entirely.
|
|
13155
|
+
* `wrapText` is the same helper the finding cards already use.
|
|
13156
|
+
*/
|
|
13157
|
+
function pushWrapped(lines, p, text, width) {
|
|
13158
|
+
const indent = " ";
|
|
13159
|
+
for (const line of wrapText(text, Math.max(20, width - 2))) lines.push(p.dim(`${indent}${line}`));
|
|
13160
|
+
}
|
|
13132
13161
|
function appendFooter(lines, result, ui) {
|
|
13133
|
-
const { p } = ui;
|
|
13162
|
+
const { p, width } = ui;
|
|
13134
13163
|
lines.push(...buildFooter({
|
|
13135
13164
|
ui,
|
|
13136
13165
|
complete: result.analysisStatus.discovery !== "partial",
|
|
13137
13166
|
durationMs: result.analysisStatus.durationMs
|
|
13138
13167
|
}));
|
|
13139
13168
|
const advisory = result.findings.filter((f) => (f.evidenceLevel ?? deriveEvidenceLevel(f.findingType, f.confidence)) === "E0").length;
|
|
13140
|
-
if (advisory > 0) lines
|
|
13169
|
+
if (advisory > 0) pushWrapped(lines, p, `${advisory} advisory finding${advisory === 1 ? "" : "s"} (E0 — observation only, no score impact)`, width);
|
|
13141
13170
|
if (result.findings.length > 0) {
|
|
13142
13171
|
const firedRuleIds = new Set(result.findings.map((f) => f.ruleId));
|
|
13143
13172
|
const measuredHere = [...firedRuleIds].filter((id) => MEASURED_FP[id] !== void 0).length;
|
|
13144
|
-
lines
|
|
13173
|
+
pushWrapped(lines, p, `Rule coverage: ${measuredHere}/${firedRuleIds.size} rules that fired here have a measured false-positive rate; the rest are heuristics. \`mjolnir rules --unmeasured\` lists them.`, width);
|
|
13145
13174
|
const verified = result.findings.filter((f) => f.runtimeCorroboration !== void 0).length;
|
|
13146
|
-
if (verified > 0) lines
|
|
13147
|
-
else lines
|
|
13175
|
+
if (verified > 0) pushWrapped(lines, p, `Runtime evidence: ${verified}/${result.findings.length} findings corroborated by a real run report (trust L3–L5); the rest are static-only.`, width);
|
|
13176
|
+
else pushWrapped(lines, p, `Runtime evidence: not available — no run report (mjolnir.report.json / test-results) next to the scan target; all findings are static-only (L0–L2).`, width);
|
|
13148
13177
|
}
|
|
13149
13178
|
const profile = result.agenticProfile;
|
|
13150
13179
|
if (profile && (profile.generatedMarkedFiles > 0 || profile.codegenLikeFiles > 0)) {
|
|
@@ -13224,7 +13253,7 @@ function renderSarif(result, repoRootUri) {
|
|
|
13224
13253
|
tool: { driver: {
|
|
13225
13254
|
name: "Mjölnir",
|
|
13226
13255
|
informationUri: "https://github.com/Sergey-Bar/Mjolnir",
|
|
13227
|
-
version: "0.5.
|
|
13256
|
+
version: "0.5.13",
|
|
13228
13257
|
rules: [...rules.values()].map((r) => {
|
|
13229
13258
|
const meta = RULES.find((x) => x.id === r.id);
|
|
13230
13259
|
return {
|
|
@@ -17456,14 +17485,29 @@ function explainRule(ruleId, fixturesRoot) {
|
|
|
17456
17485
|
ok: true,
|
|
17457
17486
|
rule,
|
|
17458
17487
|
exampleFinding: example,
|
|
17459
|
-
exampleFixturePath: fixturePath
|
|
17488
|
+
exampleFixturePath: fixturePath,
|
|
17489
|
+
exampleFixtureRelPath: relative(fixturesRoot, fixturePath).replaceAll("\\", "/")
|
|
17460
17490
|
};
|
|
17461
17491
|
}
|
|
17462
|
-
|
|
17492
|
+
/**
|
|
17493
|
+
* Default column budget when no width is supplied.
|
|
17494
|
+
*
|
|
17495
|
+
* `explain`'s prose used to be pushed as unbroken strings — the
|
|
17496
|
+
* "HOW TO VERIFY THE FIX" paragraph alone is 150 columns — so every
|
|
17497
|
+
* explanation overflowed a default terminal. Renderers here take a width
|
|
17498
|
+
* rather than reading process.stdout, so output stays a pure function of
|
|
17499
|
+
* its arguments (same rule the reporter's palette follows).
|
|
17500
|
+
*/
|
|
17501
|
+
const DEFAULT_EXPLAIN_WIDTH = 80;
|
|
17502
|
+
function renderExplain(result, width = DEFAULT_EXPLAIN_WIDTH) {
|
|
17463
17503
|
if (!result.ok || !result.rule) return `explain failed: ${result.error ?? "unknown error"}`;
|
|
17464
17504
|
const r = result.rule;
|
|
17465
17505
|
const evidenceLevel = r.evidenceLevel ?? deriveEvidenceLevel(r.findingType, r.confidence);
|
|
17466
17506
|
const lines = [];
|
|
17507
|
+
/** Pushes prose indented two columns, wrapped to the budget. */
|
|
17508
|
+
const pushBody = (text) => {
|
|
17509
|
+
for (const seg of wrapText(text, Math.max(20, width - 2))) lines.push(` ${seg}`);
|
|
17510
|
+
};
|
|
17467
17511
|
lines.push(sectionHeader(`${r.id} — ${r.title}`, ui$1));
|
|
17468
17512
|
lines.push("");
|
|
17469
17513
|
lines.push(`Severity: ${r.severity}`);
|
|
@@ -17480,19 +17524,19 @@ function renderExplain(result) {
|
|
|
17480
17524
|
if (result.exampleFinding) {
|
|
17481
17525
|
const f = result.exampleFinding;
|
|
17482
17526
|
lines.push("WHAT WAS FOUND (real detector output, not a mockup)");
|
|
17483
|
-
|
|
17527
|
+
pushBody(f.message);
|
|
17484
17528
|
lines.push("");
|
|
17485
17529
|
lines.push("WHY IT MATTERS");
|
|
17486
|
-
|
|
17530
|
+
pushBody(f.why);
|
|
17487
17531
|
lines.push("");
|
|
17488
17532
|
lines.push("HOW TO FIX");
|
|
17489
|
-
|
|
17533
|
+
pushBody(f.fix);
|
|
17490
17534
|
lines.push("");
|
|
17491
|
-
|
|
17492
|
-
} else
|
|
17535
|
+
pushBody(`Example from this rule's own must-fire fixture: ${result.exampleFixtureRelPath ?? result.exampleFixturePath ?? "(unknown path)"}`);
|
|
17536
|
+
} else for (const seg of wrapText("No example available — run this command from a mjolnir checkout (or pass --fixtures-root) so the fixture that proves this rule works can be shown as a real example.", width)) lines.push(seg);
|
|
17493
17537
|
lines.push("");
|
|
17494
17538
|
lines.push("HOW TO VERIFY THE FIX");
|
|
17495
|
-
|
|
17539
|
+
pushBody("Re-run `mjolnir` on the changed file(s) — this finding should no longer appear. `mjolnir --scope changed` scopes the check to just what you touched.");
|
|
17496
17540
|
lines.push("");
|
|
17497
17541
|
lines.push(`Docs: mjolnir rules --md (full catalog, this rule included)`);
|
|
17498
17542
|
return lines.join("\n");
|
|
@@ -17649,7 +17693,7 @@ const { runScan, buildUniversalRules, fallbackWorkspace, pathMatchesGlob, isVali
|
|
|
17649
17693
|
* `scripts/sync-sarif-version.cjs` on release and guarded by
|
|
17650
17694
|
* `tests/version-consistency.spec.ts` locally.
|
|
17651
17695
|
*/
|
|
17652
|
-
const CLI_VERSION = "0.5.
|
|
17696
|
+
const CLI_VERSION = "0.5.13";
|
|
17653
17697
|
function parseArgs(argv, onError) {
|
|
17654
17698
|
const args = {
|
|
17655
17699
|
target: ".",
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "mjolnir-qa",
|
|
3
|
-
"version": "0.5.
|
|
3
|
+
"version": "0.5.13",
|
|
4
4
|
"description": "Mjölnir — the Verification Trust Engine for QA. Audits test suites and CI pipelines, reports a worthiness score and prioritized findings.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"engines": {
|
|
@@ -44,6 +44,10 @@
|
|
|
44
44
|
"site:audit": "node site/scripts/audit-live.mjs",
|
|
45
45
|
"site:lighthouse": "node site/scripts/lighthouse-run.mjs",
|
|
46
46
|
"docs:demo": "tsx scripts/generate-readme-demo.ts",
|
|
47
|
+
"docs:video:capture": "tsx scripts/video/write-scripts.ts",
|
|
48
|
+
"docs:video:glyphs": "tsx scripts/video/check-glyphs.ts",
|
|
49
|
+
"docs:video:render": "tsx scripts/video/render.ts",
|
|
50
|
+
"docs:video": "npm run docs:video:capture && npm run docs:video:render",
|
|
47
51
|
"docs:translations": "node scripts/check-readme-translations.mjs",
|
|
48
52
|
"self-scan": "node dist/cli.mjs .",
|
|
49
53
|
"prepare": "husky",
|
|
@@ -91,6 +95,7 @@
|
|
|
91
95
|
"husky": "^9.1.7",
|
|
92
96
|
"lighthouse": "^13.4.1",
|
|
93
97
|
"lint-staged": "^17.3.0",
|
|
98
|
+
"playwright-core": "^1.63.0",
|
|
94
99
|
"prettier": "^3.3.0",
|
|
95
100
|
"tsdown": "^0.22.14",
|
|
96
101
|
"tsx": "^4.16.0",
|