mjolnir-qa 0.4.0 → 0.5.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +1199 -28
- package/README.ar.md +667 -0
- package/README.bn.md +678 -0
- package/README.br.md +706 -0
- package/README.bs.md +690 -0
- package/README.da.md +698 -0
- package/README.de.md +711 -0
- package/README.es.md +709 -0
- package/README.fr.md +714 -0
- package/README.gr.md +705 -0
- package/README.he.md +664 -0
- package/README.it.md +709 -0
- package/README.ja.md +692 -0
- package/README.ko.md +681 -0
- package/README.md +425 -152
- package/README.no.md +696 -0
- package/README.pl.md +699 -0
- package/README.ru.md +701 -0
- package/README.th.md +672 -0
- package/README.tr.md +699 -0
- package/README.uk.md +693 -0
- package/README.vi.md +680 -0
- package/README.zh.md +652 -0
- package/README.zht.md +652 -0
- package/dist/cli.d.mts +469 -22
- package/dist/cli.mjs +9780 -5230
- package/package.json +25 -7
package/README.md
CHANGED
|
@@ -1,23 +1,18 @@
|
|
|
1
1
|
<div align="center">
|
|
2
2
|
|
|
3
|
-
<
|
|
4
|
-
╔═══════════╗
|
|
5
|
-
║ ║
|
|
6
|
-
╠═══════════╣ M J Ö L N I R
|
|
7
|
-
║ ║ ║
|
|
8
|
-
╚═════╩═════╝ VERIFICATION TRUST ENGINE
|
|
9
|
-
║
|
|
10
|
-
║
|
|
11
|
-
</pre>
|
|
3
|
+
<img src="assets/readme/logo.png" alt="Mjölnir — Verification Trust Engine" width="800" />
|
|
12
4
|
|
|
13
5
|
### Your tests are lying to you. We prove it.
|
|
14
6
|
|
|
15
|
-
**Verification Trust Engine for QA.**
|
|
16
|
-
reports a worthiness score, and shows exactly where trust breaks.
|
|
7
|
+
**Verification Trust Engine for QA.** Mjölnir audits test suites and CI
|
|
8
|
+
pipelines, reports a worthiness score, and shows exactly where trust breaks.
|
|
17
9
|
|
|
18
|
-
[](https://www.npmjs.com/package/mjolnir-qa)
|
|
11
|
+
[](https://github.com/Sergey-Bar/Mjolnir/actions/workflows/ci.yml)
|
|
12
|
+
[](LICENSE)
|
|
13
|
+
[](https://nodejs.org)
|
|
14
|
+
|
|
15
|
+
English | [简体中文](README.zh.md) | [繁體中文](README.zht.md) | [한국어](README.ko.md) | [Deutsch](README.de.md) | [Español](README.es.md) | [Français](README.fr.md) | [Italiano](README.it.md) | [Dansk](README.da.md) | [日本語](README.ja.md) | [Polski](README.pl.md) | [Русский](README.ru.md) | [Norsk](README.no.md) | [Português (Brasil)](README.br.md) | [ไทย](README.th.md) | [Türkçe](README.tr.md) | [Українська](README.uk.md) | [বাংলা](README.bn.md) | [Ελληνικά](README.gr.md) | [Tiếng Việt](README.vi.md) | [עברית](README.he.md) | [العربية](README.ar.md) | [Bosanski](README.bs.md)
|
|
21
16
|
|
|
22
17
|
```bash
|
|
23
18
|
npx mjolnir-qa@latest
|
|
@@ -25,104 +20,154 @@ npx mjolnir-qa@latest
|
|
|
25
20
|
|
|
26
21
|
**Are your tests worthy of trust?**
|
|
27
22
|
|
|
28
|
-
[
|
|
23
|
+
[See it work](#-see-it-work) · [Quickstart](#-quickstart) · [What it checks](#-what-mjölnir-checks) · [Scoring](#how-the-score-works) · [CI](#-ci-integration) · [Configuration](#configuration) · [Docs](#-documentation)
|
|
29
24
|
|
|
30
25
|
</div>
|
|
31
26
|
|
|
32
27
|
---
|
|
33
28
|
|
|
34
|
-
##
|
|
29
|
+
## 🎬 See it work
|
|
35
30
|
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
| 🎭 | **Selector Health Score** — grades your Playwright locators, not just your pass rate |
|
|
40
|
-
| 🔬 | **Runtime forensics** — reads real Playwright/JUnit run data to catch `TRUE-FLAKE`, not just static guesses |
|
|
41
|
-
| 🚨 | **CI-integrity rules** — catches `continue-on-error`, `\|\| true`, and other false-green tricks |
|
|
42
|
-
| 🐍 | **Multi-language** — TypeScript/Playwright, Python/pytest, Java, and C#/.NET today, one adapter away from more |
|
|
43
|
-
| 🔒 | **Local-first** — zero network calls while scanning, zero telemetry, runs in seconds |
|
|
31
|
+
<p align="center">
|
|
32
|
+
<img src="assets/readme/demo.svg" alt="Mjölnir's full --verbose report on a demo repo: WORTHINESS 75/100 NEEDS WORK, a diagnostics-by-category breakdown, a FIX THIS FIRST list, and every finding with its rule ID and line number across CI, Playwright, test-hygiene and Python rules" width="900" />
|
|
33
|
+
</p>
|
|
44
34
|
|
|
45
|
-
|
|
35
|
+
<sub>The complete `npx mjolnir-qa ./examples/demo-repo --verbose` output,
|
|
36
|
+
rendered from the actual reporter — nothing trimmed. Regenerated by
|
|
37
|
+
`npm run docs:demo`;
|
|
38
|
+
[`tests/demo-asset-reproducibility.spec.ts`](tests/demo-asset-reproducibility.spec.ts)
|
|
39
|
+
fails CI if it drifts from what the tool prints.</sub>
|
|
46
40
|
|
|
47
|
-
|
|
41
|
+
**What just happened:**
|
|
42
|
+
|
|
43
|
+
1. Mjölnir discovered the Playwright specs, its config, the CI workflow and
|
|
44
|
+
a Python test file — four languages/formats, one pass.
|
|
45
|
+
2. It found evidence that weakens trust in the suite — a `continue-on-error`
|
|
46
|
+
masking a job, a `|| true` swallowing an exit code, hard sleeps, a
|
|
47
|
+
brittle selector, hardcoded staging URLs, a `networkidle` wait.
|
|
48
|
+
3. It turned each into a concrete finding with a rule ID, a location and a
|
|
49
|
+
fix — and a single score you can gate a PR on.
|
|
50
|
+
|
|
51
|
+
### One finding, up close
|
|
52
|
+
|
|
53
|
+
Run `mjolnir explain QA-CI-001` on the first finding above and you get:
|
|
54
|
+
|
|
55
|
+
```text
|
|
56
|
+
▚▞ QA-CI-001 — continue-on-error masks a failing verification gate
|
|
48
57
|
|
|
49
|
-
|
|
50
|
-
|
|
58
|
+
Severity: error
|
|
59
|
+
Confidence: high
|
|
60
|
+
Evidence: E2
|
|
61
|
+
Measured FP: not yet measured — this rule ships on assumption (see docs/FP-AUDIT.md)
|
|
51
62
|
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
63
|
+
WHAT WAS FOUND (real detector output, not a mockup)
|
|
64
|
+
Job `security-scan` runs a verification gate under `continue-on-error: true`.
|
|
65
|
+
|
|
66
|
+
WHY IT MATTERS
|
|
67
|
+
This job can fail every day and CI will still show green. The checkmark
|
|
68
|
+
on this workflow cannot be trusted.
|
|
69
|
+
|
|
70
|
+
HOW TO FIX
|
|
71
|
+
Remove continue-on-error, or scope it to individual non-blocking steps only.
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
That is the unit of value: not a style nit, but a place where your CI is
|
|
75
|
+
telling you something passed when it didn't.
|
|
60
76
|
|
|
61
77
|
---
|
|
62
78
|
|
|
63
79
|
## ⚡ Quickstart
|
|
64
80
|
|
|
81
|
+
Run it against a repo for a full report and a worthiness score:
|
|
82
|
+
|
|
65
83
|
```bash
|
|
66
84
|
npx mjolnir-qa@latest
|
|
67
85
|
```
|
|
68
86
|
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
| Command | What it does |
|
|
73
|
-
| ------------------------------------------------------ | ------------------------------------------------------------------- |
|
|
74
|
-
| `npx mjolnir-qa@latest --scope changed` | Only findings introduced by **your** changes — perfect for PRs |
|
|
75
|
-
| `npx mjolnir-qa@latest --json` | Machine-readable output |
|
|
76
|
-
| `npx mjolnir-qa@latest --format sarif > mjolnir.sarif` | GitHub Code Scanning integration |
|
|
77
|
-
| `npx mjolnir-qa@latest --format mermaid` | Test-architecture diagram — paste into a GitHub comment or a slide |
|
|
78
|
-
| `mjolnir doctor:playwright` | Playwright-only deep scan + Selector Health Score |
|
|
79
|
-
| `mjolnir forensics ./test-results/` | Runtime evidence — retries, true flakes, `FLAKY.md` artifact |
|
|
80
|
-
| `mjolnir triage ./test-results/` | The flaky-triage meeting, in 10 minutes instead of 45 |
|
|
81
|
-
| `mjolnir fix --dry-run` / `fix` | Safe auto-fixes with proof (dry-run first) |
|
|
82
|
-
| `mjolnir debt` | Test debt register — presentable to management |
|
|
83
|
-
| `mjolnir handover` | New-QA-onboarding map of the suite |
|
|
84
|
-
| `mjolnir pw-report ./test-results/` | Playwright run summary — retries / flakes / slowest |
|
|
85
|
-
| `mjolnir badge` | Evidentiary badge (shields.io endpoint JSON) |
|
|
86
|
-
| `mjolnir doctor` | Self-audit — prove Mjölnir's own rule base is worthy |
|
|
87
|
-
| `mjolnir rules` / `rules --md` | Rule catalog with trust metadata (JSON or markdown) |
|
|
88
|
-
| `mjolnir explain <RULE-ID>` | What/why/fix for one rule, with a real example from its own fixture |
|
|
89
|
-
| `mjolnir impact [--since <ref>]` | What changed since a prior commit — fixes and new debt |
|
|
87
|
+
**In CI, the product is one command.** It scans only what the branch
|
|
88
|
+
touched and exits non-zero on new problems:
|
|
90
89
|
|
|
91
|
-
|
|
90
|
+
```bash
|
|
91
|
+
npx mjolnir-qa@latest --scope changed
|
|
92
|
+
```
|
|
92
93
|
|
|
93
|
-
|
|
94
|
+
Drop that in a PR check — `mjolnir ci install` writes the workflow —
|
|
95
|
+
and you're done. Everything else is optional.
|
|
94
96
|
|
|
95
|
-
|
|
96
|
-
|
|
97
|
+
| Command | What it does |
|
|
98
|
+
| ----------------------------------- | ------------------------------------------------ |
|
|
99
|
+
| `mjolnir` | Full-repo scan + worthiness score |
|
|
100
|
+
| `mjolnir --scope changed` | Only what your branch introduced — the CI form |
|
|
101
|
+
| `mjolnir ci install` | Generate the advisory PR workflow |
|
|
102
|
+
| `mjolnir explain QA-CI-001` | What / why / fix + measured FP rate for one rule |
|
|
103
|
+
| `mjolnir rules --unmeasured` | The rules running on assumption, not measurement |
|
|
104
|
+
| `mjolnir --json` / `--format sarif` | Machine-readable / GitHub Code Scanning |
|
|
105
|
+
| `mjolnir --strict` | Also run quarantine-tier rules (higher FP risk) |
|
|
106
|
+
| `mjolnir --cache` | Incremental re-scans via a local verdict cache |
|
|
97
107
|
|
|
98
|
-
|
|
99
|
-
|
|
108
|
+
<details>
|
|
109
|
+
<summary><strong>When something's flaky</strong></summary>
|
|
100
110
|
|
|
101
|
-
|
|
111
|
+
| Command | What it does |
|
|
112
|
+
| ----------------------------------- | --------------------------------------------------- |
|
|
113
|
+
| `mjolnir forensics ./test-results/` | Real run data → `TRUE-FLAKE` verdicts, `FLAKY.md` |
|
|
114
|
+
| `mjolnir triage ./test-results/` | Quarantine proposal from execution history |
|
|
115
|
+
| `mjolnir pw-report ./test-results/` | Playwright run summary — retries / flakes / slowest |
|
|
116
|
+
| `mjolnir doctor:playwright` | Playwright-only deep scan + Selector Health Score |
|
|
102
117
|
|
|
103
|
-
|
|
104
|
-
+8 pts QA-CI-001 · .github/workflows/ci.yml:48
|
|
105
|
-
+3 pts QA-TEST-004 · e2e/checkout.spec.ts:6
|
|
106
|
-
```
|
|
118
|
+
</details>
|
|
107
119
|
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
|
|
120
|
+
<details>
|
|
121
|
+
<summary><strong>Occasional / reporting</strong></summary>
|
|
122
|
+
|
|
123
|
+
| Command | What it does |
|
|
124
|
+
| ------------------------------- | ------------------------------------------------ |
|
|
125
|
+
| `mjolnir fix --dry-run` / `fix` | Safe auto-fixes with proof |
|
|
126
|
+
| `mjolnir baseline` / `diff` | Snapshot findings, then report only new/worsened |
|
|
127
|
+
| `mjolnir impact --since <ref>` | What changed since a prior commit |
|
|
128
|
+
| `mjolnir debt` | Test-debt register with a cost model |
|
|
129
|
+
| `mjolnir handover` | New-QA onboarding map of the suite |
|
|
130
|
+
| `mjolnir stats` | Local all-time counters of fixes seen |
|
|
131
|
+
| `mjolnir badge` | shields.io endpoint JSON + snippet |
|
|
132
|
+
| `mjolnir rules --md` | Full rule catalog (JSON or Markdown) |
|
|
133
|
+
| `mjolnir doctor` | Self-audit of Mjölnir's own rule base |
|
|
134
|
+
| `mjolnir create-rule <ID>` | Scaffold a new rule + fixtures |
|
|
135
|
+
| `mjolnir --format mermaid` | Test-architecture diagram for a PR comment |
|
|
111
136
|
|
|
112
|
-
|
|
137
|
+
</details>
|
|
113
138
|
|
|
114
|
-
|
|
115
|
-
|
|
116
|
-
| ≥ 80 | ✓ **WORTHY** |
|
|
117
|
-
| 50 – 79 | ⚠ **NEEDS WORK** |
|
|
118
|
-
| < 50 | ✖ **UNWORTHY** |
|
|
139
|
+
Install globally instead of `npx` if you prefer: `npm i -g mjolnir-qa`.
|
|
140
|
+
Requires Node.js ≥ 22.18. Works on Windows, macOS, and Linux.
|
|
119
141
|
|
|
120
142
|
---
|
|
121
143
|
|
|
122
|
-
##
|
|
144
|
+
## 👥 Who is this for?
|
|
145
|
+
|
|
146
|
+
- **QA / SDET** owning an e2e or integration suite who need evidence the
|
|
147
|
+
suite actually deserves the green checkmark it produces.
|
|
148
|
+
- **Platform / DevEx** teams responsible for CI integrity and release
|
|
149
|
+
gates — the people who care that a `continue-on-error` never silently
|
|
150
|
+
turns a red pipeline green.
|
|
151
|
+
- **OSS maintainers** who want a cheap, always-on verification gate that
|
|
152
|
+
runs locally and in CI with zero network calls.
|
|
153
|
+
|
|
154
|
+
---
|
|
155
|
+
|
|
156
|
+
## 🔨 What Mjölnir checks
|
|
157
|
+
|
|
158
|
+
| | |
|
|
159
|
+
| --- | ----------------------------------------------------------------------------------------------------------------- |
|
|
160
|
+
| ⚖️ | **Worthiness Score** — one number, transparent deduction table, no black box |
|
|
161
|
+
| 🎭 | **Selector Health Score** — grades your Playwright locators, not just your pass rate |
|
|
162
|
+
| 🔬 | **Runtime forensics** — reads real Playwright/JUnit run data to catch `TRUE-FLAKE`, not just static guesses |
|
|
163
|
+
| 🚨 | **CI-integrity rules** — catches `continue-on-error`, `\|\| true`, and other false-green tricks |
|
|
164
|
+
| 🐍 | **All four Playwright bindings** — TypeScript, Python, Java, C#/.NET — plus pytest, JUnit/TestNG and CI workflows |
|
|
165
|
+
| 🔒 | **Local-first** — zero network calls while scanning, zero telemetry, runs in seconds |
|
|
166
|
+
|
|
167
|
+
### The rules
|
|
123
168
|
|
|
124
169
|
Every rule ships with must-fire **and** must-not-fire fixtures. A rule that
|
|
125
|
-
fires on its own negative fixture cannot ship
|
|
170
|
+
fires on its own negative fixture cannot ship — that's the false-positive
|
|
126
171
|
firewall.
|
|
127
172
|
|
|
128
173
|
<details>
|
|
@@ -145,7 +190,7 @@ firewall.
|
|
|
145
190
|
|
|
146
191
|
| ID | Rule | Severity |
|
|
147
192
|
| ------------ | --------------------------- | -------- |
|
|
148
|
-
| QA-TQUAL-001 | Mock-only verification |
|
|
193
|
+
| QA-TQUAL-001 | Mock-only verification | info |
|
|
149
194
|
| QA-TQUAL-002 | Tautological assertion | error |
|
|
150
195
|
| QA-TQUAL-009 | Unawaited promise assertion | error |
|
|
151
196
|
| QA-TQUAL-011 | Commented-out tests | warning |
|
|
@@ -160,9 +205,9 @@ firewall.
|
|
|
160
205
|
| QA-PW-002 | Unawaited locator assertion | error |
|
|
161
206
|
| QA-PW-003 | `page.pause()` / `test.only()` committed | error |
|
|
162
207
|
| QA-PW-004 | Brittle CSS/XPath selectors | warning |
|
|
163
|
-
| QA-PW-005 | Business logic inside `page.evaluate()` |
|
|
164
|
-
| QA-PW-114 | Legacy element handles (`page.$`) |
|
|
165
|
-
| QA-PW-118 | `networkidle` waits (flaky by design) |
|
|
208
|
+
| QA-PW-005 | Business logic inside `page.evaluate()` | info |
|
|
209
|
+
| QA-PW-114 | Legacy element handles (`page.$`) | info |
|
|
210
|
+
| QA-PW-118 | `networkidle` waits (flaky by design) | info |
|
|
166
211
|
| QA-PW-123 | Hardcoded environment URLs | warning |
|
|
167
212
|
|
|
168
213
|
</details>
|
|
@@ -190,18 +235,125 @@ firewall.
|
|
|
190
235
|
| QA-PY-002 | Skipped test (`skip`, non-strict `xfail`) | warning |
|
|
191
236
|
| QA-PY-003 | Test function with no assertions | error |
|
|
192
237
|
| QA-PY-005 | `time.sleep()` in tests | warning |
|
|
193
|
-
| QA-PY-006 | Empty test body (`pass`) |
|
|
194
|
-
| QA-PY-010 | Random/time dependence without freeze |
|
|
238
|
+
| QA-PY-006 | Empty test body (`pass`) | info |
|
|
239
|
+
| QA-PY-010 | Random/time dependence without freeze | info |
|
|
195
240
|
| QA-PY-012 | Tautological assertion | error |
|
|
196
241
|
|
|
242
|
+
20 Python rules total (QA-PY-001…012 pytest hygiene + QA-PY-101…108 Playwright-Python).
|
|
243
|
+
|
|
244
|
+
</details>
|
|
245
|
+
|
|
246
|
+
<details>
|
|
247
|
+
<summary><strong>Java / JUnit · TestNG ☕</strong></summary>
|
|
248
|
+
|
|
249
|
+
| ID | Rule | Severity |
|
|
250
|
+
| --------- | ---------------------------------------- | -------- |
|
|
251
|
+
| QA-JV-101 | Disabled test (`@Disabled`) | warning |
|
|
252
|
+
| QA-JV-102 | Hard sleep (`Thread.sleep()`) | warning |
|
|
253
|
+
| QA-JV-103 | Test method with no assertions | error |
|
|
254
|
+
| QA-JV-105 | Playwright `waitForTimeout()` hard sleep | warning |
|
|
255
|
+
| QA-JV-106 | Brittle selector instead of role locator | warning |
|
|
256
|
+
| QA-JV-108 | Hardcoded environment URL in test | info |
|
|
257
|
+
| QA-JV-111 | Blanket `page.route("**")` mock | info |
|
|
258
|
+
|
|
259
|
+
</details>
|
|
260
|
+
|
|
261
|
+
<details>
|
|
262
|
+
<summary><strong>C# / .NET — NUnit · xUnit · MSTest 🟣</strong></summary>
|
|
263
|
+
|
|
264
|
+
| ID | Rule | Severity |
|
|
265
|
+
| --------- | ------------------------------------------ | -------- |
|
|
266
|
+
| QA-CS-101 | Skipped test (`[Ignore]`, `[Fact(Skip=)]`) | warning |
|
|
267
|
+
| QA-CS-102 | Hard sleep (`Thread.Sleep` / `Task.Delay`) | warning |
|
|
268
|
+
| QA-CS-103 | Test method with no assertions | error |
|
|
269
|
+
| QA-CS-105 | `WaitForTimeoutAsync()` hard sleep | warning |
|
|
270
|
+
| QA-CS-106 | Brittle selector instead of role locator | warning |
|
|
271
|
+
| QA-CS-108 | Hardcoded environment URL in test | info |
|
|
272
|
+
| QA-CS-111 | Blanket `page.RouteAsync("**")` mock | info |
|
|
273
|
+
|
|
197
274
|
</details>
|
|
198
275
|
|
|
199
|
-
> The full live catalog — every rule with confidence, false-positive
|
|
200
|
-
> and autofix availability — is generated from the registry:
|
|
276
|
+
> The full live catalog — every rule with tier, confidence, false-positive
|
|
277
|
+
> risk, and autofix availability — is generated from the registry:
|
|
201
278
|
>
|
|
202
279
|
> ```bash
|
|
203
280
|
> mjolnir rules --md
|
|
204
281
|
> ```
|
|
282
|
+
>
|
|
283
|
+
> Per-rule pages live under [`docs/rules/`](docs/rules/).
|
|
284
|
+
|
|
285
|
+
### How much of this is measured
|
|
286
|
+
|
|
287
|
+
**74 of 99 rules carry a false-positive rate measured against real OSS code** (≥ 10 hand-classified findings each; see
|
|
288
|
+
[docs/FP-AUDIT.md](docs/FP-AUDIT.md)). The other 19 ship on the author's
|
|
289
|
+
estimate. Every scan footer tells you how many of the rules that _fired_
|
|
290
|
+
are measured; `mjolnir rules --unmeasured` lists the ones that aren't;
|
|
291
|
+
every rule's `mjolnir explain` page states its status. We publish the rate
|
|
292
|
+
even when it's ugly — QA-CS-103 audits at 95% and is quarantined for it.
|
|
293
|
+
Growing that 78 is the project's continuing work.
|
|
294
|
+
|
|
295
|
+
### Rule tiers and language maturity
|
|
296
|
+
|
|
297
|
+
Every rule is `core`, `extended`, or `quarantine`, assigned from its
|
|
298
|
+
**measured** false-positive rate:
|
|
299
|
+
|
|
300
|
+
| Tier | Meaning | Default scan | `--strict` |
|
|
301
|
+
| ------------ | ---------------------------------------- | :----------: | :--------: |
|
|
302
|
+
| `core` | ≤ 10 % measured FP | ✅ | ✅ |
|
|
303
|
+
| `extended` | ≤ 30 % measured FP | ✅ | ✅ |
|
|
304
|
+
| `quarantine` | above 30 %, or not yet measured (n < 10) | ❌ | ✅ |
|
|
305
|
+
|
|
306
|
+
| Language | Adapter | Coverage today |
|
|
307
|
+
| --------------- | ------------ | -------------------------------------------------- |
|
|
308
|
+
| TypeScript / JS | compiler AST | broadest, most measured — mostly `core`/`extended` |
|
|
309
|
+
| Python / pytest | regex layer | broad, corpus-audited — mostly `core`/`extended` |
|
|
310
|
+
| Java | regex layer | newer — mostly `extended`/`quarantine` |
|
|
311
|
+
| C# / .NET | regex layer | newer — mostly `extended`/`quarantine` |
|
|
312
|
+
|
|
313
|
+
TypeScript and Python have the broadest measured coverage. Java and C# ship,
|
|
314
|
+
are documented, and stay out of the headline number until a real consumer
|
|
315
|
+
suite (not a binding library's own tests) has been audited.
|
|
316
|
+
|
|
317
|
+
---
|
|
318
|
+
|
|
319
|
+
## How the score works
|
|
320
|
+
|
|
321
|
+
<p align="center">
|
|
322
|
+
<img src="assets/readme/terminal-hero.svg" alt="Mjölnir terminal output — WORTHINESS 75/100 NEEDS WORK, a diagnostics-by-category breakdown, and a FIX THIS FIRST list" width="820" />
|
|
323
|
+
</p>
|
|
324
|
+
|
|
325
|
+
<sub>Regenerated by `npm run docs:hero`;
|
|
326
|
+
[`tests/hero-asset-reproducibility.spec.ts`](tests/hero-asset-reproducibility.spec.ts)
|
|
327
|
+
fails CI if it drifts from what the reporter actually prints.</sub>
|
|
328
|
+
|
|
329
|
+
The score is transparent: **error −8, warning −3, info −1**, then normalized
|
|
330
|
+
by suite exposure (deductions per test declaration). Evidence-weighted
|
|
331
|
+
deductions mean weak signals cost less. The terminal shows the same
|
|
332
|
+
discounted numbers the score uses — no black box. Full method:
|
|
333
|
+
[docs/SCORING.md](docs/SCORING.md).
|
|
334
|
+
|
|
335
|
+
**Verdicts**
|
|
336
|
+
|
|
337
|
+
| Score | Verdict |
|
|
338
|
+
| ------- | ---------------- |
|
|
339
|
+
| ≥ 80 | ✓ **WORTHY** |
|
|
340
|
+
| 50 – 79 | ⚠ **NEEDS WORK** |
|
|
341
|
+
| < 50 | ✖ **UNWORTHY** |
|
|
342
|
+
|
|
343
|
+
**Evidence levels** — every finding carries one; it sets the finding's
|
|
344
|
+
weight in the score:
|
|
345
|
+
|
|
346
|
+
| Level | Meaning | Score impact | Example |
|
|
347
|
+
| ----- | -------------------- | ---------------- | -------------------------------------------------- |
|
|
348
|
+
| E2 | Deterministic defect | Full deduction | `.only` committed — structurally provable |
|
|
349
|
+
| E1 | Heuristic pattern | Half deduction | Regex-matched `sleep()` — strong signal, not proof |
|
|
350
|
+
| E0 | Observation | Zero (info only) | Reported but never gates CI or deducts |
|
|
351
|
+
|
|
352
|
+
Most rules are **E1**. The tagline "we prove it" refers to this system:
|
|
353
|
+
E2 findings are structural proof; E1 findings are correctly-positioned
|
|
354
|
+
warnings, not formal proofs.
|
|
355
|
+
|
|
356
|
+
An empty repo scores `null`, never a fake 100 — see [Trust model](#trust-model).
|
|
205
357
|
|
|
206
358
|
---
|
|
207
359
|
|
|
@@ -222,7 +374,7 @@ regressed.
|
|
|
222
374
|
|
|
223
375
|
---
|
|
224
376
|
|
|
225
|
-
## 🔬 Runtime
|
|
377
|
+
## 🔬 Runtime evidence
|
|
226
378
|
|
|
227
379
|
Static flakiness detection is guessing. Mjölnir reads **real execution
|
|
228
380
|
data** — Playwright JSON reports and JUnit XML from any runner:
|
|
@@ -247,7 +399,61 @@ test. It gets flagged `TRUE-FLAKE` regardless of the final green checkmark.
|
|
|
247
399
|
|
|
248
400
|
---
|
|
249
401
|
|
|
250
|
-
##
|
|
402
|
+
## ⚡ Mjölnir is not another linter
|
|
403
|
+
|
|
404
|
+
Linters tell you whether code follows rules. Mjölnir tells you whether your
|
|
405
|
+
verification can be trusted.
|
|
406
|
+
|
|
407
|
+
| | ESLint / SonarQube | Coverage tools | Manual review | **Mjölnir** |
|
|
408
|
+
| -------------------------------------------------------- | :----------------: | :------------: | :-----------: | :---------: |
|
|
409
|
+
| CI workflow integrity (`continue-on-error`, `\|\| true`) | ❌ | ❌ | rarely | ✅ |
|
|
410
|
+
| Cross-language (TS, Python, Java, C#) from one tool | ❌ | ❌ | ❌ | ✅ |
|
|
411
|
+
| Grades Playwright locator resilience (Selector Health) | ❌ | ❌ | rarely | ✅ |
|
|
412
|
+
| Flags tests with no real assertions | ✅ (plugin)\* | ❌ | sometimes | ✅ |
|
|
413
|
+
| Catches hard sleeps (`waitForTimeout`, `time.sleep`) | ✅ (plugin)\* | ❌ | sometimes | ✅ |
|
|
414
|
+
| Runs in seconds, zero network calls while scanning | ✅ | ✅ | — | ✅ |
|
|
415
|
+
|
|
416
|
+
\*`eslint-plugin-jest` (`expect-expect`) and `eslint-plugin-playwright`
|
|
417
|
+
(`expect-expect`, `no-wait-for-timeout`) cover these for their respective
|
|
418
|
+
frameworks.
|
|
419
|
+
|
|
420
|
+
**Runtime analysis** is a separate category from static linting:
|
|
421
|
+
|
|
422
|
+
| | Playwright retry reporter | Allure / ReportPortal | **Mjölnir forensics** |
|
|
423
|
+
| --------------------------------------------- | :-----------------------: | :-------------------: | :-------------------: |
|
|
424
|
+
| Reads real run data for `TRUE-FLAKE` verdicts | partial\* | partial (tag) | ✅ |
|
|
425
|
+
| Flaky-triage report from execution history | ❌ | ✅ | ✅ |
|
|
426
|
+
| Integrates with static worthiness score | ❌ | ❌ | ✅ |
|
|
427
|
+
|
|
428
|
+
\*Playwright tracks retries internally but does not produce a standalone
|
|
429
|
+
flakiness report with verdict labels.
|
|
430
|
+
|
|
431
|
+
---
|
|
432
|
+
|
|
433
|
+
## 🤖 Why not just use AI code review?
|
|
434
|
+
|
|
435
|
+
Different problem, different layer. AI review can spot a suspicious test
|
|
436
|
+
change in a diff; it does not prove the verification system as a whole is
|
|
437
|
+
trustworthy — and it only sees the diff you show it.
|
|
438
|
+
|
|
439
|
+
| | AI code review (Copilot, etc.) | **Mjölnir** |
|
|
440
|
+
| ------------------------------------- | :----------------------------: | :-------------------------: |
|
|
441
|
+
| Cost per scan | Tokens (scales with diff size) | **Zero** (local, installed) |
|
|
442
|
+
| Sees the whole suite + all CI configs | Only the PR diff you show it | **Everything, every time** |
|
|
443
|
+
| Deterministic (same input → same out) | ❌ (non-deterministic) | **✅** |
|
|
444
|
+
| Catches patterns dormant for months | Only if it's in the context | **✅** (scans all files) |
|
|
445
|
+
| Remembers findings between runs | ❌ (no memory across sessions) | **✅** (baseline + diff) |
|
|
446
|
+
| Runs without human triggering | Needs a PR or prompt | **✅** (CI hook, 3 seconds) |
|
|
447
|
+
|
|
448
|
+
**Use both.** AI catches nuance, intent, and design flaws no regex can
|
|
449
|
+
find. Mjölnir catches the structural patterns AI overlooks because they
|
|
450
|
+
look "intentional" — a committed `.only`, a swallowed exit code, a
|
|
451
|
+
`continue-on-error` on a test job. Those aren't bugs that need reasoning;
|
|
452
|
+
they're facts that need scanning.
|
|
453
|
+
|
|
454
|
+
---
|
|
455
|
+
|
|
456
|
+
## 🤖 CI integration
|
|
251
457
|
|
|
252
458
|
One command generates a PR workflow — advisory by default, never blocking:
|
|
253
459
|
|
|
@@ -264,48 +470,108 @@ Or wire it into GitHub Code Scanning natively via SARIF:
|
|
|
264
470
|
sarif_file: mjolnir.sarif
|
|
265
471
|
```
|
|
266
472
|
|
|
473
|
+
Editor and pipeline setup for SARIF: [docs/SARIF-INTEGRATION.md](docs/SARIF-INTEGRATION.md).
|
|
474
|
+
|
|
475
|
+
### Changed-scope coverage
|
|
476
|
+
|
|
477
|
+
`--scope changed` attributes findings to lines added in your branch vs the
|
|
478
|
+
merge-base with `main`. It covers test files (`*.spec.*`, `*.test.*`) plus
|
|
479
|
+
GitHub workflow files and Playwright configs in the diff. When the
|
|
480
|
+
merge-base can't be resolved — shallow clone, detached HEAD, non-git
|
|
481
|
+
target, different default branch — it degrades honestly: findings fall
|
|
482
|
+
back to full-file attribution and the report says so. Override the base
|
|
483
|
+
ref with `--base <ref>`.
|
|
484
|
+
|
|
267
485
|
---
|
|
268
486
|
|
|
269
|
-
##
|
|
487
|
+
## Configuration
|
|
488
|
+
|
|
489
|
+
Mjölnir is zero-config. An optional `mjolnir.config.json` (or
|
|
490
|
+
`.mjolnir.json`) at the repo root tunes severity, gating and scope — it
|
|
491
|
+
never changes detection semantics.
|
|
492
|
+
|
|
493
|
+
| Key | Type | Effect |
|
|
494
|
+
| ------------------- | ------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
|
495
|
+
| `exclude` | `string[]` | Extra ignore globs (gitignore subset), on top of the built-in defaults |
|
|
496
|
+
| `gate` | `"advisory" \| "error" \| "warning"` | Which severities exit non-zero (default `error`; `advisory` never blocks) |
|
|
497
|
+
| `severityOverrides` | `{ "<RULE-ID>": severity }` | Re-rank a rule's findings for your repo |
|
|
498
|
+
| `ignore` | `IgnoreEntry[]` | Suppress findings — **`reason` is required**; entries expire after 90 days (an explicit `expires` date, or the config file's last-modified time for entries without one) |
|
|
499
|
+
| `plugins` | `string[]` | Third-party rule packages (see [Trust model](#trust-model)) |
|
|
500
|
+
|
|
501
|
+
```json
|
|
502
|
+
{
|
|
503
|
+
"gate": "error",
|
|
504
|
+
"exclude": ["legacy/**"],
|
|
505
|
+
"severityOverrides": { "QA-PW-118": "warning" },
|
|
506
|
+
"ignore": [
|
|
507
|
+
{
|
|
508
|
+
"ruleId": "QA-TEST-004",
|
|
509
|
+
"files": ["e2e/legacy-login.spec.ts"],
|
|
510
|
+
"reason": "Third-party widget needs a settle delay; tracked in JIRA-4821",
|
|
511
|
+
"expires": "2026-12-31"
|
|
512
|
+
}
|
|
513
|
+
]
|
|
514
|
+
}
|
|
515
|
+
```
|
|
270
516
|
|
|
271
|
-
-
|
|
272
|
-
|
|
273
|
-
|
|
274
|
-
- **
|
|
275
|
-
|
|
276
|
-
|
|
277
|
-
-
|
|
278
|
-
|
|
279
|
-
|
|
280
|
-
|
|
281
|
-
- **Partial honesty** — if analysis was cut short, the output says so.
|
|
282
|
-
Never "complete" when it isn't.
|
|
283
|
-
- **QA-native language** — findings speak your vocabulary:
|
|
284
|
-
`FALSE-GREEN`, `FLAKY-RISK`, `BLOCKS-RELEASE`, `HYGIENE`.
|
|
285
|
-
- **FP firewall** — detection runs on a comment/string-free view of the code
|
|
286
|
-
(TypeScript rules use the compiler AST): a pattern inside a prose comment
|
|
287
|
-
or a doc-example string is documentation, not a finding.
|
|
517
|
+
- **`.mjolnirignore`** — a plain gitignore-style file for path exclusions,
|
|
518
|
+
same dialect as `exclude`. Use it for machine-wide noise; use `exclude`
|
|
519
|
+
when the list belongs in version control alongside the rest of the config.
|
|
520
|
+
- **CLI overrides** — `--strict` (include quarantine rules), `--width <cols>`
|
|
521
|
+
and `--ascii` / `--no-ascii` (terminal rendering), `--tone blunt`
|
|
522
|
+
(blunter messages), `--max-duration <sec>` (bounded partial scan).
|
|
523
|
+
- Rule suppression and deprecation lifecycle: [docs/RULE-LIFECYCLE.md](docs/RULE-LIFECYCLE.md).
|
|
524
|
+
|
|
525
|
+
`ignore` entries also power the standalone `mjolnir suppressions` command,
|
|
526
|
+
which lists what's currently suppressed and when each entry expires.
|
|
288
527
|
|
|
289
528
|
---
|
|
290
529
|
|
|
291
|
-
##
|
|
530
|
+
## 📐 Exit codes & contracts
|
|
531
|
+
|
|
532
|
+
Frozen — safe to build CI logic on:
|
|
533
|
+
|
|
534
|
+
| Exit code | Meaning |
|
|
535
|
+
| --------- | --------------------------------------------------------------- |
|
|
536
|
+
| `0` | Clean — no findings at or above the gate |
|
|
537
|
+
| `1` | Findings at or above the gate |
|
|
538
|
+
| `2` | Partial scan (time budget hit, unreadable files) — never blocks |
|
|
539
|
+
| `10` | Usage error (bad flag, missing target) |
|
|
540
|
+
| `20` | Internal error |
|
|
292
541
|
|
|
293
|
-
|
|
294
|
-
|
|
295
|
-
tree — the same trust model as ESLint or Vitest plugins. Only install plugins
|
|
296
|
-
you'd trust as devDependencies. Core rule-ID prefixes (`QA-TEST`, `QA-TQUAL`,
|
|
297
|
-
`QA-PW`, `QA-CI`, `QA-PY`, `QA-ENV`, `QA-JV`, `QA-CS`) are reserved and
|
|
298
|
-
rejected from plugins to prevent spoofing.
|
|
542
|
+
The JSON/SARIF report is `schemaVersion: 1`. Rule IDs (`QA-<FAMILY>-NNN`)
|
|
543
|
+
are immutable once shipped and never reused.
|
|
299
544
|
|
|
300
545
|
---
|
|
301
546
|
|
|
302
|
-
##
|
|
547
|
+
## Trust model
|
|
303
548
|
|
|
304
|
-
|
|
305
|
-
|
|
306
|
-
|
|
307
|
-
|
|
308
|
-
|
|
549
|
+
- **Local-first** — zero network calls during scanning. Ever. Zero telemetry.
|
|
550
|
+
- **No false proof** — we'd rather say "unknown" than "verified". An empty
|
|
551
|
+
repo gets `score: null`, never a fake 100.
|
|
552
|
+
- **Partial honesty** — if analysis was cut short, the output says so.
|
|
553
|
+
Never "complete" when it isn't.
|
|
554
|
+
- **FP firewall** — detection runs on a comment/string-free view of the code
|
|
555
|
+
(TypeScript rules use the compiler AST): a pattern inside a prose comment
|
|
556
|
+
or a doc-example string is documentation, not a finding.
|
|
557
|
+
- **Measured, not asserted** — only rules with a false-positive rate from
|
|
558
|
+
real OSS code ship in the headline tiers (see
|
|
559
|
+
[How much of this is measured](#how-much-of-this-is-measured)); the scan
|
|
560
|
+
footer and `mjolnir rules --unmeasured` tell you which is which.
|
|
561
|
+
- **Plugin trust** — plugins are npm packages declared under `"plugins"`.
|
|
562
|
+
There is **no sandbox**: plugin code runs with full Node privileges, the
|
|
563
|
+
same trust model as ESLint or Vitest plugins. Core rule-ID prefixes are
|
|
564
|
+
reserved and rejected from plugins to prevent spoofing.
|
|
565
|
+
- **Workspace-local external rules** (folder-based, zero network) — a
|
|
566
|
+
`mjolnir-rules/` directory next to the scan target loads custom rules:
|
|
567
|
+
JSON files declare regex patterns (no code executed), `.mjs`/`.js`
|
|
568
|
+
modules export `rules` (full-Node trust, same as plugins). External
|
|
569
|
+
rules carry the same trust metadata as core; they can never ship in
|
|
570
|
+
the core tier (core requires a measured FP rate from the corpus
|
|
571
|
+
sidecar — a declared `tier: "core"` is clamped to `extended`), obey
|
|
572
|
+
tier caps, and are drift-checked: `mjolnir rules --md --external`
|
|
573
|
+
renders the catalog from the loaded files (provenance `external`),
|
|
574
|
+
and the matrix generator accepts `--external <root>`.
|
|
309
575
|
|
|
310
576
|
---
|
|
311
577
|
|
|
@@ -319,15 +585,16 @@ mjolnir/
|
|
|
319
585
|
├── src/
|
|
320
586
|
│ ├── engine/ # LanguageAdapter interface + rule runner
|
|
321
587
|
│ ├── adapters/ # typescript · python · java · csharp · github-actions
|
|
322
|
-
│ ├── rules/ #
|
|
588
|
+
│ ├── rules/ # rules across 8 families + the measured-FP table
|
|
323
589
|
│ ├── playwright/ # Selector Health Score engine
|
|
324
|
-
│ ├── discovery/ # workspace, frameworks,
|
|
590
|
+
│ ├── discovery/ # workspace, frameworks, ignore resolution
|
|
325
591
|
│ ├── scope/ # git merge-base changed-scope engine
|
|
326
|
-
│ ├── scorer/ # transparent deduction table
|
|
592
|
+
│ ├── scorer/ # transparent deduction table + prioritization
|
|
327
593
|
│ ├── reporter/ # terminal · JSON · SARIF 2.1 · Mermaid
|
|
328
594
|
│ ├── forensics/ # run-data ingestion · flake verdicts · triage
|
|
329
|
-
│ ├──
|
|
330
|
-
│
|
|
595
|
+
│ ├── config/ # mjolnir.config.json + suppressions
|
|
596
|
+
│ ├── plugins/ # third-party rule loading (no sandbox)
|
|
597
|
+
│ └── commands/ # every subcommand
|
|
331
598
|
└── tests/
|
|
332
599
|
├── fixtures/ # must-fire / must-not-fire per rule
|
|
333
600
|
└── golden/ # frozen score regression locks
|
|
@@ -335,54 +602,60 @@ mjolnir/
|
|
|
335
602
|
|
|
336
603
|
</details>
|
|
337
604
|
|
|
338
|
-
|
|
339
|
-
|
|
340
|
-
|
|
341
|
-
|
|
605
|
+
- **Rules are pure functions** — `(SourceFileContext) → Finding[]`, no I/O,
|
|
606
|
+
no globals. Adding an ecosystem = one adapter + its rules.
|
|
607
|
+
- **TypeScript/Playwright uses the compiler AST** (ts-morph). Python, Java
|
|
608
|
+
and C# run on a shared comment/string-masked regex layer.
|
|
609
|
+
- A tree-sitter WASM AST layer for Java and C# exists and is the next
|
|
610
|
+
precision step — it is not yet wired into the synchronous scan pipeline.
|
|
342
611
|
|
|
343
612
|
---
|
|
344
613
|
|
|
345
|
-
##
|
|
614
|
+
## 📚 Documentation
|
|
615
|
+
|
|
616
|
+
| Document | What's in it |
|
|
617
|
+
| ------------------------------------------------------ | ------------------------------------------------- |
|
|
618
|
+
| [docs/SCORING.md](docs/SCORING.md) | Score normalization + evidence weighting |
|
|
619
|
+
| [docs/FP-AUDIT.md](docs/FP-AUDIT.md) | Measured false-positive rates + method |
|
|
620
|
+
| [docs/RULE-LIFECYCLE.md](docs/RULE-LIFECYCLE.md) | Rule states, suppression, deprecation |
|
|
621
|
+
| [docs/VERSIONING.md](docs/VERSIONING.md) | Semver policy, frozen surfaces, deprecation cycle |
|
|
622
|
+
| [docs/SARIF-INTEGRATION.md](docs/SARIF-INTEGRATION.md) | SARIF output + editor/CI setup |
|
|
623
|
+
| [docs/rules/](docs/rules/) | Generated per-rule catalog |
|
|
624
|
+
| [CONTRIBUTING.md](CONTRIBUTING.md) | Dev setup + contribution workflow |
|
|
625
|
+
| [SUPPORT.md](SUPPORT.md) | Where to ask, report and get help |
|
|
626
|
+
| [CHANGELOG.md](CHANGELOG.md) | Release history |
|
|
627
|
+
| [SECURITY.md](SECURITY.md) | Vulnerability reporting |
|
|
346
628
|
|
|
347
|
-
|
|
348
|
-
# Run directly (recommended):
|
|
349
|
-
npx mjolnir-qa@latest
|
|
629
|
+
---
|
|
350
630
|
|
|
351
|
-
|
|
352
|
-
npm i -g mjolnir-qa
|
|
353
|
-
```
|
|
631
|
+
## 📈 Status
|
|
354
632
|
|
|
355
|
-
|
|
633
|
+
**v0.5.x · open beta.** The JSON schema and exit codes are frozen contracts.
|
|
634
|
+
TypeScript and Python have the broadest measured coverage; Java and C# are
|
|
635
|
+
newer — read them through the [maturity table](#rule-tiers-and-language-maturity).
|
|
636
|
+
Honest scope, no invented dates: the [public roadmap](https://sergey-bar.github.io/Mjolnir/reference/roadmap).
|
|
356
637
|
|
|
357
638
|
---
|
|
358
639
|
|
|
359
640
|
## 🤝 Contributing
|
|
360
641
|
|
|
361
|
-
New rules are the easiest first contribution — one command scaffolds
|
|
362
|
-
|
|
642
|
+
New rules are the easiest first contribution — one command scaffolds the
|
|
643
|
+
rule plus its must-fire **and** must-not-fire fixtures (the generated rule
|
|
644
|
+
intentionally fails its fixtures until you implement real detection — a stub
|
|
645
|
+
cannot ship):
|
|
363
646
|
|
|
364
647
|
```bash
|
|
365
648
|
mjolnir create-rule QA-PW-140 --title "Screenshot without diff bound"
|
|
366
649
|
```
|
|
367
650
|
|
|
368
|
-
|
|
369
|
-
|
|
370
|
-
implement real detection — a stub cannot ship.
|
|
371
|
-
|
|
372
|
-
Manual path:
|
|
373
|
-
|
|
374
|
-
1. Fork + clone, `npm install`
|
|
375
|
-
2. Copy any rule folder under `src/rules/` as a template
|
|
376
|
-
3. Add must-fire AND must-not-fire fixtures under `tests/fixtures/<YOUR-ID>/`
|
|
377
|
-
4. `npx vitest run` — both directions green? Ship the PR.
|
|
378
|
-
|
|
379
|
-
See [SECURITY.md](SECURITY.md) for vulnerability reporting.
|
|
651
|
+
Full dev setup, the standing-gate commands, and the anti-creep / fixture
|
|
652
|
+
firewall laws are in [CONTRIBUTING.md](CONTRIBUTING.md).
|
|
380
653
|
|
|
381
654
|
---
|
|
382
655
|
|
|
383
656
|
<div align="center">
|
|
384
657
|
|
|
385
|
-
**
|
|
658
|
+
**Stop shipping tests you can't trust.**
|
|
386
659
|
|
|
387
660
|
```bash
|
|
388
661
|
npx mjolnir-qa@latest
|