mjolnir-qa 1.0.5 → 1.0.9

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -9,6 +9,482 @@ Rule behavior changes (new rules, FP-rate changes against the corpus,
9
9
  severity changes) are first-class entries here — rule IDs are immutable
10
10
  once shipped, so this file is the record of what changed between versions.
11
11
 
12
+ ## [1.0.9] — 2026-09-12
13
+
14
+ ### Changes since 1
15
+
16
+ - chore: add testTimeout: 30_000 to vitest.config.ts
17
+ - test: MR-7A release-verification machinery contract (#84)
18
+ - feat: SC-8 determinism verifier + SC-11 control-state record (MR-8.C/D) (#83)
19
+ - feat: pack-audit gate wired before publish (SC-6, MR-8.B) (#82)
20
+ - test: SC-3/SC-4/SC-7 supply-chain hygiene gates (MR-8.A) (#81)
21
+ - docs: PR template + CONTRIBUTING targeted-slice ladder (MR-4, GC-2) (#80)
22
+ - feat: docs:regen aggregate — one idempotent command for every generated surface (MR-5) (#79)
23
+ - test(site): negative proof for the D-2 emitted-HTML link gate (MR-6) (#78)
24
+ - fix(release): recognize the '(Merged PR #N)' squash subject so PR labels drive the bump (#77)
25
+
26
+ ## [1.0.8] — 2026-09-11
27
+
28
+ ### Changes since 1.0.7
29
+
30
+ - chore: sync smithery.yaml in the release cut step (Merged PR #76)
31
+
32
+ ## [1.0.7] — 2026-09-11
33
+
34
+ ### Changes since 1.0.6
35
+
36
+ - docs: 1.0.6 CHANGELOG section lead-ins (Merged PR #75)
37
+
38
+ ## [1.0.6] — 2026-09-11
39
+
40
+ ### R10 2.0 preparation: breaking-set inventory + boundary-law guards (remediation/remote-first WI-25)
41
+
42
+ Preparation-only increment: the 2.0 breaking-set proposal sheet, its migration draft, and the boundary-law guard tests — nothing breaking ships in this release.
43
+
44
+ ### Added
45
+
46
+ - **2.0 breaking-set inventory** (`docs/2.0-BREAKING-SET.md`, WI-25): the
47
+ proposal sheet per the strategic blueprint's §28/§18 — two justified
48
+ breaking candidates (BS-1 default suppression expiry with the explicit
49
+ never-expire opt-out; BS-2 retirement completion into `RETIRED_RULE_IDS`),
50
+ each carrying its benefit>cost justification, its migration pointer, and a
51
+ PROPOSED decision line awaiting owner ratification, plus the locked
52
+ NOT-breaking list (`schemaVersion 1` additive extension, exit codes,
53
+ additive verbs, Node matrix, frozen surfaces). **Nothing is implemented in
54
+ this release** — nothing enters 2.0 "because large", and no frozen surface
55
+ breaks without evidence that `schemaVersion 1` cannot represent the
56
+ behavior.
57
+ - **Migration guide draft** (`docs/MIGRATION-2.0-DRAFT.md`): the working
58
+ draft of the 2.0.0 guide (publication law: CHANGELOG + site with the
59
+ release itself) covering BS-1 (init config check → explicit `expires` /
60
+ `expires: false`; no silent retroactive expiry) and BS-2 (retired-rule
61
+ list, suppression cleanup, §15-lifecycle-honest disappearance causes).
62
+ - **Boundary-law + non-goal guards** (`tests/contract/boundary-law.spec.ts`,
63
+ blueprint §9.1/§24/§36): the canonical layers (engine, forensics,
64
+ adapters, rules) never import upward into commands/ or the transports;
65
+ the MCP transport imports no detection machinery and rides the canonical
66
+ machine contract (pipeline → contract → runtime evidence → agent
67
+ transport is never reversed); the zero-network contract holds; the
68
+ dependency list carries no telemetry/cloud/hosted-backend package; src/
69
+ reads no telemetry configuration; and the breaking-set discipline is
70
+ drift-locked (every entry carries a Decision, nothing is implemented).
71
+
72
+ ### Changed
73
+
74
+ - **Final capability matrix update** (`src/capabilities.ts`,
75
+ `docs/PLAYWRIGHT-CAPABILITIES.md`): the `to agents` column now carries the
76
+ foundational Agent Skill's evidence pointer (`src/commands/install-agents.ts`,
77
+ shipped with R8/WI-22) on every row — a `no → yes` flip in the same change
78
+ set that shipped its evidence, per the claim law.
79
+
80
+ ### Fixed
81
+
82
+ - **MCP stdio bundle no longer prints the terminal Trust Report onto the
83
+ JSON-RPC stream** (`src/mcp/server.ts`): the standalone entry
84
+ (`node dist/mcp/stdio.mjs`, `npm run mcp`) dragged the CLI module in via
85
+ `import { runScan, CLI_VERSION } from "../cli.js"`, and cli.ts's entry
86
+ tail fired inside the bundle (`import.meta.url === argv[1]`), emitting the
87
+ full terminal Trust Report before/between JSON-RPC frames — a fatal
88
+ protocol violation for any MCP client. The transport now imports the
89
+ canonical homes directly (`engine/scan-pipeline.js`, `engine/version.js`),
90
+ the bundle contains no CLI entry tail, and the boundary-law guard bans the
91
+ `../cli.js` import from the MCP layer permanently. Found by the R10
92
+ bug-hunt smoke against the real stdio transport.
93
+ - **Stale dist can no longer mask new code in spawned-binary tests**
94
+ (`tests/e2e/global-setup.ts`): the EXISTS-ONLY guard skipped the build
95
+ whenever a bundle was present, so the spawned stdio binary kept answering
96
+ from a pre-R8 catalog ("unknown tool: triage") while the suite stayed
97
+ green. The setup now rebuilds whenever any `src/**/*.ts` is newer than the
98
+ bundle (the same freshness discipline as the generated-docs drift gates).
99
+ - **Freshness diagnosis names each drift class** (`checkArtifactFreshness`):
100
+ a revision bump (`rule@old -> new`), a rule retired since the render
101
+ (`rule@rev (retired)`), and a rule added since the render
102
+ (`rule@rev (new)`) are distinct remediations — a flat list hid which one
103
+ happened.
104
+ - **`trust-report --from` reports the complete artifact set** it writes
105
+ (md + html + json), not just the MD path.
106
+ - **The `provenance = bound` system-invariant item is now WIRED** (plan
107
+ §5.2 activation; `src/commands/release-trust.ts`): the release-trust
108
+ invariant previously hardcoded `provenance: UNSUPPORTED` even after the
109
+ machinery it waited for shipped. It is now PROVEN exactly when the
110
+ machine-anchored identity chain is proven — scope-integrity (runIdentity +
111
+ evidence graph, R4c) AND artifact-integrity (artifact scanId binding, R9)
112
+ both PASS — and stays UNSUPPORTED (recorded, non-blocking) otherwise. The
113
+ contract doc's activation sentence and the drift-lock are updated
114
+ accordingly; the shipped verdict is unchanged (PASS 12/12, provenance
115
+ bound).
116
+
117
+ ### R9 Trust Artifact integrity + HTML completion (remediation/remote-first WI-23+24)
118
+
119
+ Trust Artifacts gain machine-anchored identity and a deterministic HTML surface; stale, wrong-run, revision-drifted, and unbound artifacts are now detectable.
120
+
121
+ ### Added
122
+
123
+ - **Artifact integrity binding** (`src/commands/trust-report.ts`, R9): every
124
+ Trust Artifact (md · json · html) now embeds its IDENTITY — the machine
125
+ anchor (`scanId` from runIdentity), the bound commit when resolvable
126
+ (offline git read; null, never fabricated), the fired rule(rev) inventory
127
+ (deduped, sorted, undeclared revisions omitted — never a fabricated rev),
128
+ and the evidence inventory (totals, runtime-corroborated count, per-level
129
+ counts). Consumers detect **stale artifacts** (a scanId from another run),
130
+ **mismatched revisions** (rule-set drift, named per rule), and **unbound
131
+ artifacts** (pre-R9 producers) via `checkArtifactFreshness` — an unbound or
132
+ stale artifact is RECORDED, never assumed current.
133
+ - **HTML Trust Artifact** (WI-23 completion, §18): deterministic,
134
+ self-contained `mjolnir-trust-report.html` — inline CSS only, zero external
135
+ resources, hostile interpolations escaped, byte-identical regen (same
136
+ ScanResult + label + commit → same bytes), the same five-question structure
137
+ as the MD. The command writes all three formats; `--from` gains an optional
138
+ `--commit <sha>` so the Action binds the artifact to the executing run's
139
+ HEAD.
140
+ - **Artifact Integrity dimension wired** (`check:artifact-integrity` in
141
+ `src/commands/release-trust.ts`, R9 surface): the structural evaluation
142
+ asserts the identity binding, the three-format output, the freshness
143
+ detection, and the byte-regen/hostile-safety contract locks. The
144
+ release-trust contract's documented-unwired list is now EMPTY — all 12
145
+ canonical dimensions are wired and machine-evaluated.
146
+
147
+ ### R8 MCP runtime-evidence tools + Agent Safety (remediation/remote-first WI-21+22)
148
+
149
+ The MCP transport learns the runtime-evidence tools, and every installed agent surface inherits the safety contract.
150
+
151
+ ### Added
152
+
153
+ - **MCP runtime-evidence tools** (`src/mcp/server.ts`, WI-21): `forensics`,
154
+ `triage`, `pw-report` join the tool catalog as 1:1 mappings onto the SAME
155
+ engine functions the CLI verbs call — no MCP-only semantics. Parity is
156
+ drift-locked table-driven (`tests/mcp/parity.spec.ts`): for every new tool ×
157
+ every fixture class (Playwright JSON · JUnit XML · hostile corrupt report ·
158
+ no-reports directory) the MCP result deep-equals the canonical CLI
159
+ derivation, hostile inputs degrade to zero records on BOTH surfaces, and the
160
+ hostile parameter matrix (missing / empty / non-string / nonexistent path)
161
+ yields INVALID_PARAMS naming the target, never a crash. A crashing tool
162
+ never kills the server (`tests/mcp/crash-containment.spec.ts`): the failure
163
+ lands in the transport's existing catch as a structured INTERNAL error and
164
+ the server keeps answering. One scan in flight; zero network; the plugin
165
+ gate applies unchanged.
166
+ - **Agent Safety dimension wired** (`check:agent-safety` in
167
+ `src/commands/release-trust.ts`, R8 surface): the structural evaluation
168
+ asserts the §17 safety wording on every installed skill surface, that the
169
+ MCP tool surface never opens the plugin trust gate, and that the agent edge
170
+ case (`fg-agent-unsafe-action`) stays registered in the False-Green Attack
171
+ Corpus. The release-trust contract's documented-unwired list shrinks to
172
+ artifact-integrity only (ships R9).
173
+
174
+ ### Changed
175
+
176
+ - **Agent brief inherits the Constitution** (`src/commands/install-agents.ts`,
177
+ WI-22): every installed instruction surface (.claude/, .cursor/, .kilo/,
178
+ AGENTS.md) now carries the non-negotiable agent-safety contract — NEVER
179
+ declare trustworthiness without evidence · AGENT CLAIM ≠ VERIFICATION ·
180
+ NEVER manufacture, edit, or synthesize evidence · NEVER convert INCONCLUSIVE
181
+ to pass · NEVER suppress findings or weaken rules to get green — plus the
182
+ loop preconditions (FIX requires a proven actionable defect; RESCAN requires
183
+ changed-scope identification; PROOF requires fresh post-fix execution
184
+ evidence). Drift-locked by `tests/contract/agent-skill-surface.spec.ts`
185
+ (frozen surfaces only; safety wording asserted).
186
+
187
+ ### R7 Playwright capability matrix (remediation/remote-first WI-20)
188
+
189
+ The Playwright capability matrix becomes a product surface with its own drift lock.
190
+
191
+ ### Added
192
+
193
+ - **Playwright Capability Matrix** (`docs/PLAYWRIGHT-CAPABILITIES.md`,
194
+ `src/capabilities.ts`, WI-20): the product-depth surface — 12 Playwright
195
+ capabilities × 8 depth columns (detect · explain · produce evidence ·
196
+ correlate runtime · trust verdict · CLI · MCP · agents), every cell
197
+ explicitly classed (zero UNCLASSIFIED), every `yes` backed by a resolvable
198
+ evidence pointer (registered rule ID or in-repo artifact) with FAIL-CLOSED
199
+ validation: the generator refuses to render a claim on a dangling pointer.
200
+ Generated (`npm run docs:capabilities-playwright`) and drift-locked
201
+ (tests/contract/playwright-capabilities.spec.ts). The `to agents` column is
202
+ uniformly **no** until R8 ships the Agent Skill — stated, not implied.
203
+ Claims never exceed proven capability; rule counts stay out of the claim
204
+ surface entirely.
205
+
206
+ ### R6 forensic taxonomy + Selector Health v2 (remediation/remote-first WI-18+19)
207
+
208
+ Forensic verdicts gain the semantic taxonomy, and Selector Health v2 replaces the locator heuristic.
209
+
210
+ ### Added
211
+
212
+ - **Forensic verdict taxonomy** (`src/forensics/classify.ts`, WI-18): the
213
+ canonical §6 verdict set (likely-real-defect · environmental-failure ·
214
+ infrastructure-failure · flaky · retry-dependent · unstable-construction ·
215
+ **inconclusive default**) applied by a deterministic minimum-signal table —
216
+ a single weak signal can never classify confidently; conflicting signal
217
+ families force INCONCLUSIVE with an explicit `contradictory` evidence-state.
218
+ Every `TestVerdict` now carries a machine-visible `forensic` classification
219
+ (attempts + captured error text; sources without error text mark
220
+ `unsupported`, never a guess). Contradiction reconciliation
221
+ (`corroborates | contradicts | insufficient`) implements Contract H: the
222
+ runtime can corroborate but never silently weakens a static claim — a
223
+ contradiction renders the PAIR inconclusive while the claim stands.
224
+ - **Selector Health v2** (`correlateSelectorHealth`, WI-19): runtime
225
+ correlation + concrete safe next actions; **no correlation ⇒ no claim** —
226
+ absent or merely-green runtime evidence yields no health claim in either
227
+ direction; the v1 static score is secondary and never altered here.
228
+
229
+ ### Changed
230
+
231
+ - `TestRecord` gains an optional `errors` text surface (the trace ingester
232
+ populates it); `TestVerdict` gains the additive `forensic` field.
233
+
234
+ ### R5 trace ingester (remediation/remote-first WI-17)
235
+
236
+ Trace forensics: bounded ingestion of Playwright trace.zip artifacts into the evidence core.
237
+
238
+ ### Added
239
+
240
+ - **Playwright trace ingester** (`src/forensics/trace.ts`, WI-17): deterministic,
241
+ offline, bounded, version-aware ingestion of per-test traces — `trace.zip`
242
+ (a bounded, dependency-free ZIP reader: EOCD scan, central-directory
243
+ enumeration, stored/deflate members via `node:zlib` with a decompressed-output
244
+ cap) or raw `.trace`/`.ndjson` NDJSON streams. Action pairs become
245
+ Evidence-Core `TestRecord`s (start/end pairing, durations, per-action
246
+ errors; timeout errors render `timedOut`). `runForensics` recognizes trace
247
+ artifacts in both file and directory modes; the report source union gains
248
+ `playwright-trace` additively (`contractVersion 1` unchanged).
249
+ - False-Green corpus cases for the trace surface: corrupt stream, truncated
250
+ stream, event-count overflow, unsupported version marker, zip without
251
+ `trace.trace` — all degrade to the zero-record exit-2 state, never a green
252
+ empty suite; plus positive controls (real stored zip + valid stream ingest
253
+ with paired durations) proving the rejections are precision, not blindness.
254
+
255
+ ### Changed
256
+
257
+ - `ForensicsReport.source` + `RuntimeCorroboration.source` widened additively
258
+ with `"playwright-trace"`.
259
+
260
+ ### R4c Evidence Graph + Scope Integrity + Exit-Code proofs (remediation/remote-first)
261
+
262
+ Every verdict now carries a machine-anchored evidence graph, scope-integrity accounting, and exit-code proofs.
263
+
264
+ ### Added
265
+
266
+ - **Run Identity** (`src/engine/run-identity.ts`): the deterministic anchor —
267
+ `scanId = sha256(input snapshot fingerprint + rulesDigest + config
268
+ fingerprint + engine version)`; set-identity semantics (input order does not
269
+ matter); every scan report carries `runIdentity` + `evidenceGraph` — the
270
+ chain-law links VERDICT ← EVIDENCE ← EXECUTION ← SCOPE ← SOURCE ← RULE(rev)
271
+ ← FIXTURE ← REPRODUCTION, each `ref` present only when its identity input
272
+ exists (no fabrication). The engine-version literal moved to the leaf module
273
+ `src/engine/version.ts` (cli.ts re-exports it as CLI_VERSION;
274
+ sync-sarif-version.cjs + version-consistency spec follow).
275
+ - **Scope Integrity** (`ScanResult.scopeIntegrity`, additive): discovered /
276
+ analyzed / ignored / unrecognized / parseFailed / truncated counts +
277
+ `scopeVerdict` — PROVEN only when analyzed ≡ claimed scope; else PARTIAL
278
+ with named reasons. The terminal reporter renders the scope block and the
279
+ "repository verified" phrasing is forbidden output unless PROVEN. Walk-level
280
+ accounting: matcher exclusions (`onIgnored`) and unclaimed files
281
+ (`onUnrecognized`) are counted at the shared walk; parse failures are
282
+ counted at the rule stage.
283
+ - **Exit-code decision proofs** (tests/blast-radius/scope-and-exit.spec.ts):
284
+ the frozen decision points exercised in both directions — trigger present →
285
+ frozen code, trigger absent → a different code — plus the closed frozen set
286
+ {0,1,2,10,20}.
287
+ - Machine-contract doc regenerated with the three additive blocks
288
+ (`contractVersion 1` unchanged — additive within the schema).
289
+
290
+ ### Changed
291
+
292
+ - Discovery accounting: the shared walk counts matcher-excluded files and
293
+ unclaimed files (ScanContext gains optional `onIgnored`/`onUnrecognized`;
294
+ all shared-walk adapters pass them through).
295
+
296
+ ### R4b False-Green Attack Corpus (remediation/remote-first)
297
+
298
+ The False-Green Attack Corpus: hostile failure classes with mutation-based detection proofs.
299
+
300
+ ### Added
301
+
302
+ - **tests/false-green/** — the adversarial corpus (plan §6, P0): 20 cases
303
+ across the plan's seven hostile classes (execution · parser · adapter ·
304
+ evidence · rule · mcp · agent failures), each declaring the seven
305
+ owner-required fields (INPUT / EXPECTED EXECUTION / EVIDENCE / VERDICT /
306
+ EXIT CODE / REPORT FIELDS / RELEASE IMPACT) and executed against real
307
+ surfaces with specific field bindings:
308
+ - execution: empty suite (score null + no-tests-found recorded), deadline
309
+ truncation, and the partial+findings never-blocks invariant (audit C5);
310
+ - parsers (through the real `runForensics` entry): corrupt JSON, truncated
311
+ Playwright report, malformed JUnit, unsupported schema → zero records →
312
+ exit-2 state — PARSER FAILURE ≠ CLEAN; duplicate retry-storm records stay
313
+ visible;
314
+ - adapters: scalar-jobs workflow fabricates nothing; broken YAML is SKIPPED
315
+ with accounting;
316
+ - rules: a throwing local plugin rule (QA-ACME-666) → `rulesCrashed ≥ 1`
317
+ with the scan completing — RULE CRASH ≠ CLEAN;
318
+ - evidence: missing/corrupt baseline → hasBaseline=false (exit 2); stale
319
+ baseline resolutions stay scoped to their capture; the foreign
320
+ baselineCommit is recorded (binding gate ships R4c);
321
+ - MCP: unknown tool / invalid params answer JSON-RPC errors, never success;
322
+ - agent: codegen and generated-header provenance classification — AGENT
323
+ CLAIM ≠ VERIFICATION.
324
+ - **Mutation / assertion-strength protocol** (tests/false-green/mutation-
325
+ protocol.spec.ts): for every wired case and every report-field binding, the
326
+ false-green twin of the honest report (failure→success, partial→complete,
327
+ unknown→clean, crashed-rule→clean…) is injected and the case's assertion
328
+ must FAIL on it — a decorative assertion fails CI. Parser input twins flip
329
+ the hostile input to its benign form and require the observed verdict to
330
+ flip with it.
331
+ - **Generated, drift-locked index** (npm run false-green:index + index.spec.ts):
332
+ one row per case with all seven declarations, the mutation inventory, and
333
+ the UNSURFACED rows (MCP transport internals / agent-action policy → R8;
334
+ artifact binding → R9) — recorded per Constitution §5, never silently
335
+ dropped. All seven plan classes present.
336
+
337
+ ### R4a Trust Constitution + Release Trust Verdict (remediation/remote-first)
338
+
339
+ The Trust Constitution and the two-layer release-trust verdict algebra.
340
+
341
+ ### Added
342
+
343
+ - **docs/TRUST-CONSTITUTION.md** — canonical law: CERTIFICATION-POLICY A1–A4
344
+ adopted as §1; the 18 PASS-forbidden conditions (verbatim); the closed status
345
+ algebra (PROVEN evidence-state → PASS/FAILED derivation, terminality rule,
346
+ record shape); the core law (`PASS = conclusion backed by sufficient
347
+ evidence`); per-dimension applicability (UNSUPPORTED surfaces are recorded,
348
+ non-blocking, and drift-locked); publication honesty.
349
+ - **docs/RELEASE-TRUST-CONTRACT.md** — the canonical 12 dimensions (fixed set,
350
+ fixed order, governance-locked): Engine/Evidence/Rule Integrity, Failure
351
+ Containment, Corpus Integrity, Contract Compatibility, Determinism, Scope
352
+ Integrity (ships R4c), Reproducibility, Zero-Network Compliance, Agent Safety
353
+ (R8), Artifact Integrity (R9).
354
+ - New verb **`mjolnir release-trust`** emitting `mjolnir.release-trust@1` —
355
+ byte-deterministic (frozen key order, no timestamps, zero absolute paths),
356
+ per-dimension `evidence` + `determination` via the status algebra, verdict =
357
+ contract satisfaction (never a PROVEN count) with the binding system
358
+ invariant. Exit contract: 0 PASS · 1 non-PASS · 2 blocked context · 10 usage ·
359
+ 20 internal. Drift-locked by tests/contract/release-trust-contract.spec.ts
360
+ (canonical set/order, binding resolution, derivation table + terminality,
361
+ byte-stability, path-freedom).
362
+
363
+ ### Changed
364
+
365
+ - **release.yml**: the Release Trust Verdict gate is wired RELEASE-BLOCKING
366
+ pre-publish (Tests → Certification → CHANGELOG Gate → … → release-trust gate
367
+ → publish), running the BUILT binary; the verdict block + machine contract
368
+ ship with the GitHub Release (publication honesty — a missing proof renders
369
+ UNPROVEN, never omitted). No waiver path.
370
+
371
+ ### R4 blast radius audit (remediation/remote-first R4)
372
+
373
+ The blast-radius audit: a machine-verified surface manifest with its own drift lock.
374
+
375
+ ### Added
376
+
377
+ - **docs/BLAST-RADIUS-AUDIT.md** — the machine-verified surface manifest
378
+ (`npm run docs:blast-radius`): src inventory with per-area LOC, the internal
379
+ import fan-in ranking (change-blast candidates), the external dependency
380
+ allowlist, and the shipped surface (adapters, rules census, CLI flags, report
381
+ formats, frozen exit codes).
382
+ - **tests/contract/blast-radius.spec.ts** — the machine-TESTABLE boundary
383
+ contract: the committed manifest must equal a fresh render; every external
384
+ import in src/ must belong to the allowlist (`yaml`, `ts-morph`,
385
+ `web-tree-sitter`, `tree-sitter-wasms`; node builtins are platform
386
+ contracts); every CLI flag parsed must appear in the manifest; every
387
+ `process.exit(N)` in src/ must be inside the frozen set (0/1/2/10/20).
388
+
389
+ ### P6 quarantine remediation (remediation/remote-first R3)
390
+
391
+ Quarantine remediation: measured verdicts recorded, the quarantine ledger reconciled, and three rules restored to the live set.
392
+
393
+ ### Added
394
+
395
+ - **docs/QUARANTINE-REMEDIATION.md** — the ledger-first quarantine view, generated
396
+ from the live registry (`npm run docs:quarantine-ledger`) and drift-locked
397
+ (tests/contract/quarantine-ledger.spec.ts): one row per live quarantine rule
398
+ with failure-mode class, disposition, and re-measure gate; historical section
399
+ records the governed retirements.
400
+ - Python tree-sitter parse stage: `parsePythonAst` wired into the python
401
+ adapter's async `parseAst` hook (the §10 parse-or-fallback contract), with
402
+ `src/engine/python-ast.ts` structural queries — the first real python AST
403
+ substrate (the Sprint-8 "unwired" caveat is closed and re-pinned honestly).
404
+
405
+ ### Changed
406
+
407
+ - **QA-PY-007** (detectorRevision 4, AST rework): fires only on ≥2-statement
408
+ with-blocks or broad root exception types — the adjudicated FP core
409
+ (single-statement/specific-type) suppressed. Corpus: pytest-dev 167 → 11,
410
+ pallets-click 16 → 1 live findings.
411
+ - **QA-TQUAL-009** (detectorRevision 2, AST rework): skips Cypress command
412
+ chains (`cy.`-rooted — the driver awaits them) and deliberate `void`
413
+ discards. Corpus: cypress-realworld-app 10 → 0.
414
+ - **QA-PW-147** (detectorRevision 2, final attempt): AST arm fires only on real
415
+ test/it declarations — code-as-data (`test('test')` inside lint-rule test
416
+ strings) can never fire. Corpus: eslint-plugin repo 32 → 0.
417
+ - **QA-ENV-001** (detectorRevision 4, final attempt): OS-path sub-pattern
418
+ dropped (20/20 adjudicated FP — deliberate path fixtures, same undecidability
419
+ as the wave-2 host drop); locale/local-time families kept. Corpus: grafana
420
+ 7 → 4.
421
+ - Measurement: orphaned verdicts (findings the reworks suppressed) archived to
422
+ `tests/corpus/verdicts/archive/` per the established prune flow; the three
423
+ fully-reworked rules fall below the n ≥ 10 threshold and ship UNMEASURED
424
+ until owner re-adjudication (measured census 77 → 74 of 79; the
425
+ certification floor test documents the P6 invalidations).
426
+
427
+ ### P3c Jenkins (remediation/remote-first R2)
428
+
429
+ Jenkins support: a bounded Jenkinsfile scanner and the QA-CI Jenkins arms (retry masking, catchError rescue, silent swallow).
430
+
431
+ ### Added
432
+
433
+ - Jenkinsfile detection: the root `Jenkinsfile` (declarative and scripted
434
+ pipelines) is now discovered and scanned as a TEXT-target kind — a bounded,
435
+ string-aware Groovy block scanner (`sh` segments, `catchError` blocks,
436
+ `try`/`catch` pairs); no new language grammar (master-plan P3c wording).
437
+ - New rule **QA-CI-014** "try/catch swallows a verification-stage failure" —
438
+ a `try` running a gate whose `catch` neither rethrows, calls `error(...)`,
439
+ marks `currentBuild.result`, nor downgrades via `unstable()`. BORN
440
+ QUARANTINE (§15.5): opt-in via `--strict` until corpus-measured.
441
+
442
+ ### Changed
443
+
444
+ - **QA-CI-002** (detectorRevision 4): Jenkinsfile routing — the lexical
445
+ `|| true` scan now reaches `sh` strings.
446
+ - **QA-CI-008** (detectorRevision 4): Jenkinsfile arms —
447
+ `catchError(buildResult: 'SUCCESS')` wrapping a gate, and `unstable()` used
448
+ as a rescue for a failed verification stage (master-plan P3c shapes;
449
+ `buildResult: 'UNSTABLE'` is a visible downgrade and never fires).
450
+ - **QA-CI-009** (detectorRevision 3): Jenkinsfile arm — `sh` running a
451
+ verification gate with `returnStatus: true` discards the exit code.
452
+ - Measurement: sidecar + `MEASURED_FP` re-recorded for QA-CI-002/008/009
453
+ (corpus re-run: no corpus repo carries a root Jenkinsfile, so the
454
+ classified verdict evidence carries over unchanged).
455
+
456
+ ### P3b Azure DevOps (remediation/remote-first R1)
457
+
458
+ Azure DevOps support: guarded azure-pipelines.yml parsing, the QA-CI Azure arms, and the adapter's honest accounting.
459
+
460
+ ### Added
461
+
462
+ - Azure DevOps pipeline detection: `azure-pipelines.yml` at the repo root is now
463
+ discovered and scanned (`azure-pipelines` adapter, safe-YAML machinery shared
464
+ with the GitHub Actions parser — alias-bomb guard, depth cap, prototype-safe
465
+ keys; docs/AZURE-DEVOPS.md).
466
+ - New rule **QA-CI-013** "Verification gate conditioned so it can never fail the
467
+ pipeline" — `condition: failed()` rescue, `condition: false`, `enabled: false`
468
+ on Azure verification gates. BORN QUARANTINE (§15.5): opt-in via `--strict`
469
+ until corpus-measured; never silent-core.
470
+ - docs/AZURE-DEVOPS.md — platform recipe with the frozen exit-code contract.
471
+
472
+ ### Changed
473
+
474
+ - **QA-CI-001** (detectorRevision 3): Azure DevOps arm — `continueOnError: true`
475
+ on a verification step or a gate-bearing job (same mechanism, framework-tagged
476
+ `azure-pipelines`).
477
+ - **QA-CI-002** (detectorRevision 3): Azure routing — the lexical `|| true` scan
478
+ now reaches `bash:`/`pwsh:` script blocks in azure-pipelines.yml.
479
+ - **QA-CI-007** (detectorRevision 3): Azure DevOps arm — `retryCountOnTaskFailure`
480
+ on verification tasks.
481
+ - **QA-CI-008** (detectorRevision 3): Azure DevOps arm — verification gate jobs
482
+ conditioned `always()` / `succeededOrFailed()` (master-plan P3b shape).
483
+ - Measurement: sidecar + `MEASURED_FP` re-recorded at detectorRevision 3 for
484
+ QA-CI-001/002/007/008; corpus re-run showed zero QA-CI count drift (no corpus
485
+ repo carries a discoverable azure-pipelines.yml), so the existing classified
486
+ verdicts remain the measurement evidence.
487
+
12
488
  ## [1.0.5] — 2026-09-10
13
489
 
14
490
  ### Changes since 1
package/README.md CHANGED
@@ -146,7 +146,7 @@ QA impact: False-green risk (FALSE-GREEN)
146
146
  Measured FP: 11% (19 hand-classified corpus verdicts)
147
147
  FP risk: low (author estimate)
148
148
  Languages: yaml
149
- Frameworks: github-actions
149
+ Frameworks: github-actions, azure-pipelines
150
150
 
151
151
  WHAT WAS FOUND (real detector output, not a mockup)
152
152
  Job `security-scan` runs a verification gate under `continue-on-error: true`.
@@ -262,7 +262,7 @@ have no such requirement.)
262
262
 
263
263
  ## What Mjölnir finds
264
264
 
265
- **<!-- census:total-rules -->77 rules<!-- /census:total-rules -->** in four families — **test hygiene**, **test quality**,
265
+ **<!-- census:total-rules -->79 rules<!-- /census:total-rules -->** in four families — **test hygiene**, **test quality**,
266
266
  **Playwright**, **CI integrity** — over TypeScript/JavaScript, Python,
267
267
  Java, C# and GitHub Actions YAML, covering Playwright in all four bindings
268
268
  plus pytest, JUnit, TestNG, NUnit, xUnit, MSTest, Jest, Vitest and Mocha,
@@ -442,9 +442,9 @@ Rung by rung: [docs/TERMINOLOGY.md](docs/TERMINOLOGY.md).
442
442
 
443
443
  ### How much of this is measured
444
444
 
445
- **<!-- census:measured-of-total -->77 of 77<!-- /census:measured-of-total --> rules carry a false-positive rate measured against real OSS code**
445
+ **<!-- census:measured-of-total -->74 of 79<!-- /census:measured-of-total --> rules carry a false-positive rate measured against real OSS code**
446
446
  (≥ 10 hand-classified findings each — [docs/FP-AUDIT.md](docs/FP-AUDIT.md)).
447
- The other <!-- census:unmeasured -->0<!-- /census:unmeasured --> ship on the author's estimate and say so, per rule, in
447
+ The other <!-- census:unmeasured -->5<!-- /census:unmeasured --> ship on the author's estimate and say so, per rule, in
448
448
  `mjolnir explain`; `mjolnir rules --unmeasured` lists them, and every scan
449
449
  footer reports how many of the rules that actually _fired_ are measured.
450
450
 
@@ -685,7 +685,7 @@ artifacts.
685
685
  product does what the requirement asked for.
686
686
  - **A 100 is not proof of a good suite.** Whether your suite covers your
687
687
  actual risk is a different question, and this tool does not answer it.
688
- - **<!-- census:unmeasured-of-total -->0 of 77<!-- /census:unmeasured-of-total --> rules ship on an estimate**, not a measured rate — disclosed
688
+ - **<!-- census:unmeasured-of-total -->5 of 79<!-- /census:unmeasured-of-total --> rules ship on an estimate**, not a measured rate — disclosed
689
689
  per rule, not buried here.
690
690
  - **E1 is not E2.** Heuristic findings are worth reading, not worth
691
691
  applying blindly.
package/dist/cli.d.mts CHANGED
@@ -80,7 +80,7 @@ interface RuntimeCorroboration {
80
80
  * constrains TRUE-FLAKE derivation, which lives in the analysis, not
81
81
  * in the provenance label.
82
82
  */
83
- source: "playwright-json" | "junit-xml" | "jest-json" | "vitest-json";
83
+ source: "playwright-json" | "junit-xml" | "jest-json" | "vitest-json" | "playwright-trace";
84
84
  /** Number of tests executed in the finding's file (any level). */
85
85
  testsExecuted: number;
86
86
  /**
@@ -304,6 +304,62 @@ interface ScanResult {
304
304
  */
305
305
  rulesCrashed?: number;
306
306
  };
307
+ /**
308
+ * Scope Integrity block (product-gap master plan §7, R4c): the
309
+ * claimed-vs-analyzed accounting. `scopeVerdict` is PROVEN only when
310
+ * every discovered file was analyzed — no matcher exclusions, no
311
+ * unrecognized files, no parse failures, no truncation. Additive
312
+ * within schemaVersion 1.
313
+ */
314
+ scopeIntegrity?: {
315
+ /** Files discovery claimed for adapters. */
316
+ discovered: number;
317
+ /** Files that reached (and survived) the rule stage. */
318
+ analyzed: number;
319
+ /** Files excluded by the ignore matcher (counted at the walk). */
320
+ ignored: number;
321
+ /** Files the walk saw but no adapter claims. */
322
+ unrecognized: number;
323
+ /** Discovered files whose parse/analysis threw (counted, never fatal). */
324
+ parseFailed: number;
325
+ /** Named truncation events (deadline, file caps). */
326
+ truncated: number;
327
+ /** PROVEN only when analyzed ≡ claimed scope; else PARTIAL + reasons. */
328
+ scopeVerdict: "PROVEN" | "PARTIAL";
329
+ /** The named scope reasons, present only when PARTIAL. */
330
+ reasons?: string[];
331
+ };
332
+ /**
333
+ * Run Identity (R4c): the deterministic anchor binding verdict ←
334
+ * evidence ← execution ← scope ← source ← rule(rev). Present when the
335
+ * execution was machine-anchored; never fabricated.
336
+ */
337
+ runIdentity?: {
338
+ scanId: string;
339
+ inputFingerprint: string;
340
+ rulesDigest: string;
341
+ configFingerprint: string;
342
+ engineVersion: string;
343
+ };
344
+ /**
345
+ * Evidence Graph (R4c): the chain-law links (VERDICT ← EVIDENCE ←
346
+ * EXECUTION ← SCOPE ← SOURCE ← RULE(rev) ← FIXTURE ← REPRODUCTION).
347
+ * A link's `ref` is present only when its identity input exists — the
348
+ * CHAIN is always emitted so unbound links stay visible.
349
+ */
350
+ evidenceGraph?: {
351
+ chain: Array<{
352
+ link: string;
353
+ ref?: string;
354
+ }>;
355
+ runId?: {
356
+ scanId: string;
357
+ inputFingerprint: string;
358
+ rulesDigest: string;
359
+ configFingerprint: string;
360
+ engineVersion: string;
361
+ };
362
+ };
307
363
  /**
308
364
  * Local incremental cache report (Beta-to-Stable plan, M5.2). Present
309
365
  * only when the scan ran with `--cache`; additive within
@@ -848,6 +904,17 @@ declare function isValidFindingRecord$1(f: unknown): f is Omit<Finding, "ruleId"
848
904
  */
849
905
  declare function runScan$1(args: CliArgs, hooks?: ScanHooks): Promise<ScanResult>;
850
906
  //#endregion
907
+ //#region src/engine/version.d.ts
908
+ /**
909
+ * Engine version, as a leaf module: `runIdentity` (R4c) needs the engine
910
+ * version inside scan-pipeline WITHOUT importing cli.ts (a cycle — cli
911
+ * imports the pipeline). The literal follows the same discipline as
912
+ * CLI_VERSION and SARIF's driver.version: kept in sync on release by
913
+ * scripts/sync-sarif-version.cjs and guarded by the version-consistency
914
+ * spec. cli.ts re-exports this as CLI_VERSION.
915
+ */
916
+ declare const ENGINE_VERSION = "1.0.9";
917
+ //#endregion
851
918
  //#region src/cli-io.d.ts
852
919
  /**
853
920
  * Shared process IO sinks (certification-audit Phase 5, G6): the variadic
@@ -870,17 +937,6 @@ declare function runDoctorCommand(argv: string[], io?: {
870
937
  //#endregion
871
938
  //#region src/cli.d.ts
872
939
  declare const runScan: typeof runScan$1, buildUniversalRules: typeof buildUniversalRules$1, fallbackWorkspace: typeof fallbackWorkspace$1, pathMatchesGlob: typeof pathMatchesGlob$1, isValidFindingRecord: typeof isValidFindingRecord$1, discoverRuntimeReport: typeof discoverRuntimeReport$1, KNOWN_RULE_IDS: ReadonlySet<string>, OVERLAP_META_BY_RULE_ID: ReadonlyMap<string, OverlapMeta>, EVIDENCE_OVERRIDES: ReadonlyMap<string, string>, SUITE_INVALIDATING_RULE_IDS: ReadonlySet<string>;
873
- /**
874
- * Tool version for `mjolnir --version`.
875
- *
876
- * A literal, not a package.json read: the shipped artifact is a single
877
- * bundled `dist/cli.mjs`, so resolving package.json at runtime depends on
878
- * where the file happens to sit after install. This follows the same
879
- * discipline as SARIF's `driver.version` — kept in sync by
880
- * `scripts/sync-sarif-version.cjs` on release and guarded by
881
- * `tests/version-consistency.spec.ts` locally.
882
- */
883
- declare const CLI_VERSION = "1.0.5";
884
940
  /** A usage-error detail: the offending token, when one exists. */
885
941
  interface UsageErrorDetail {
886
942
  /** The unknown flag or rejected value (e.g. `--nope`, `loud`). */
@@ -1050,4 +1106,4 @@ declare function runHelpCommand(argv: string[], io?: {
1050
1106
  }): number;
1051
1107
  declare function isEntryPoint(): boolean;
1052
1108
  //#endregion
1053
- export { CLI_VERSION, type CliArgs, EVIDENCE_OVERRIDES, KNOWN_RULE_IDS, OVERLAP_META_BY_RULE_ID, type Output, SUITE_INVALIDATING_RULE_IDS, type ScanHooks, UsageErrorDetail, buildUniversalRules, discoverRuntimeReport, err, exitForFindings, fallbackWorkspace, internalErrorMessage, isEntryPoint, isValidFindingRecord, levenshtein, main, nearestFlags, out, parseArgs, pathMatchesGlob, runBadgeCommand, runBaselineCommand, runCiInstall, runCreateRuleCommand, runDebtCommand, runDiffCommand, runDoctorCommand, runDoctorPlaywright, runExplainCommand, runFixCommand, runForensicsCommand, runHandoverCommand, runHelpCommand, runImpactCommand, runInitCommand, runMutationCommand, runPrCommentCommand, runPwReportCommand, runRulesCommand, runScan, runScanCommand, runStatsCommand, runSuppressions, runTriageCommand, runVerifyCommand, usageErrorMessage };
1109
+ export { ENGINE_VERSION as CLI_VERSION, type CliArgs, EVIDENCE_OVERRIDES, KNOWN_RULE_IDS, OVERLAP_META_BY_RULE_ID, type Output, SUITE_INVALIDATING_RULE_IDS, type ScanHooks, UsageErrorDetail, buildUniversalRules, discoverRuntimeReport, err, exitForFindings, fallbackWorkspace, internalErrorMessage, isEntryPoint, isValidFindingRecord, levenshtein, main, nearestFlags, out, parseArgs, pathMatchesGlob, runBadgeCommand, runBaselineCommand, runCiInstall, runCreateRuleCommand, runDebtCommand, runDiffCommand, runDoctorCommand, runDoctorPlaywright, runExplainCommand, runFixCommand, runForensicsCommand, runHandoverCommand, runHelpCommand, runImpactCommand, runInitCommand, runMutationCommand, runPrCommentCommand, runPwReportCommand, runRulesCommand, runScan, runScanCommand, runStatsCommand, runSuppressions, runTriageCommand, runVerifyCommand, usageErrorMessage };