@mmerterden/multi-agent-pipeline 17.3.0 → 17.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -14,6 +14,73 @@ Internal file-layout changes that don't affect the slash-command surface are sti
14
14
 
15
15
  ---
16
16
 
17
+ ## [17.4.0] - 2026-09-15
18
+
19
+ Phase 4 now answers two questions it could not answer before: does the evidence a
20
+ review cites actually exist, and which files did the reviewer read.
21
+
22
+ ### Added
23
+
24
+ - `verify-citations.mjs`: resolves every `file:line` a run claims against git
25
+ rather than against the writer. Reads reviewer findings, triage
26
+ accepted/deferred/rejected, analysis `repoEvidence` buckets and markdown prose,
27
+ checks the path with `git cat-file` and the line against the blob's real
28
+ length. Exit 0/1/2/64.
29
+ - Phase 4 runs it with `--worktree`: a fix made during a review round is
30
+ uncommitted, so pinning to HEAD would call every citation to it invented.
31
+ - The analysis pre-dispatch gate runs it pinned to a commit, because a
32
+ document cites repos it only reads.
33
+ - Until now the only check was a regex asking whether a row LOOKED like a
34
+ citation, which `Foo.swift:9999` satisfies in a repo with no Foo.swift.
35
+ - `review-file-filter.mjs` and `schemas/review-file-exclusions.json`: Phase 4
36
+ Step 1.85 decides what the reviewers read before the diff cap decides what
37
+ fits. The cap truncates the largest files first, so a regenerated lockfile was
38
+ not merely wasted budget - it was what survived while real code was cut. 45
39
+ generic patterns (generated trees, lockfiles, recorded snapshots, vendored
40
+ source, build output, binary assets); every exclusion carries its reason and
41
+ the glob that matched. Fails open: an unreadable pattern list reviews
42
+ everything and says so.
43
+ - `fileCoverage[]` on `reviewer-output.schema.json` (v1.3.0), enforced by
44
+ `validate-reviewer.mjs --coverage`. A reviewer that opened one file of ten and
45
+ one that read all ten returned the identical empty `findings[]`; now each file
46
+ in the set gets a `reviewed` or `skipped` verdict, and `skipped` owes a reason.
47
+ - `graph-mermaid.mjs`: draws the blast radius `graph-affected.mjs` already
48
+ measures, as a mermaid flowchart in the PR body's impact section. GitHub
49
+ renders it natively, so the diagram costs no renderer and no dependency. Over
50
+ the node ceiling it reports what was left out rather than quietly drawing a
51
+ small radius for a large change.
52
+ - `validate-analysis-doc.mjs` check 2d-2: a sequence-diagram participant that
53
+ appears nowhere else in the document, and a named service that appears in no
54
+ diagram, are both reported.
55
+ - Shared modules `glob-match.mjs` and `git-path.mjs`. Two glob matchers would
56
+ duplicate meaning, not code: the day they disagree, a file is excluded from the
57
+ review and still counted in the conformance denominator.
58
+
59
+ ### Fixed
60
+
61
+ - `verify-citations` read `wiki.example.com:8443` out of any URL with a port and
62
+ called it an absolute path. An analysis document is made of Confluence, Jira
63
+ and Figma links, so the gate would have failed every real document with the
64
+ wrong reason. URLs are stripped first, and an extension must start with a
65
+ letter so `v1.2:30` is not a path either.
66
+ - `review-file-filter` dropped every non-ASCII filename. With `core.quotePath` on
67
+ git prints a Turkish name as `"G\303\266r..."`, which matches no glob and
68
+ resolves to no file - so it did not fail, it vanished from the denominator.
69
+ - `smoke-own-punctuation.sh` and `smoke-personal-data.sh` used `git grep` and
70
+ `git ls-files` without `--untracked` / `--others`, so a brand-new file was
71
+ invisible until committed. The run that would have caught a violation was
72
+ always the run after the one that introduced it.
73
+ - `analysis/locked.md` heading said 36 decisions over a list of 37, with nothing
74
+ counting.
75
+
76
+ ### Changed
77
+
78
+ - `phase-4-review.md` was compressed to stay under its ceiling rather than
79
+ raising it. Two passages moved to the refs that own them:
80
+ `features/design-conformance.md` gained "Why this runs as a gate", and
81
+ `cross-cli-contract.md` gained "Panel diversity per host" with the per-host
82
+ reviewer table it had been describing in prose.
83
+
17
84
  ## [17.3.0] - 2026-09-14
18
85
 
19
86
  A three-pass review of 17.2.0 found two things it had declared finished. Both
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@mmerterden/multi-agent-pipeline",
3
- "version": "17.3.0",
3
+ "version": "17.4.0",
4
4
  "description": "8-phase AI development pipeline with full orchestration on Claude Code, Copilot CLI and Codex CLI. Analysis, planning, TDD, CLI-aware parallel review with consensus surfacing + Fable triage, default-FAIL evidence gates, secret + intent guards, per-phase cost ledger, persistent learnings memory, wiki generation, commit automation. Token-preserving uninstall.",
5
5
  "type": "module",
6
6
  "main": "index.js",
@@ -25,13 +25,15 @@ Return ONLY a JSON object:
25
25
  {
26
26
  "findings": [{"severity": "blocking|important|suggestion", "file": "...", "line": N, "issue": "...", "fix": "...", "ruleId": "SEC-01", "criteriaSource": "ios-coding-standard"}],
27
27
  "conformance": [{"ruleId": "SEC-01", "verdict": "conformant|violated|not-applicable", "file": "Sources/X.swift", "line": 42, "reason": "..."}],
28
+ "fileCoverage": [{"path": "Sources/X.swift", "verdict": "reviewed|skipped", "reason": "..."}],
28
29
  "approved": true|false
29
30
  }
30
31
  ```
31
32
 
32
33
  `ruleId` + `criteriaSource` are required on a finding that comes from a cited
33
34
  rule and omitted otherwise. `conformance` is required whenever a `${CRITERIA}`
34
- block was supplied - see below.
35
+ block was supplied, and `fileCoverage` whenever a `${REVIEW_FILES}` block was -
36
+ see below.
35
37
 
36
38
  ## Severity Classification
37
39
 
@@ -90,6 +92,38 @@ obvious breach; `judgement` rules are yours, and each needs the measurement its
90
92
  When no `${CRITERIA}` block is supplied, omit `conformance` entirely and review
91
93
  on the focus areas above.
92
94
 
95
+ ## Review file set (required when supplied)
96
+
97
+ The orchestrator may pass a `${REVIEW_FILES}` block: the changed files you are
98
+ accountable for, computed before you were dispatched. Generated output,
99
+ lockfiles, recorded snapshots, vendored source and binaries are already removed,
100
+ so what is left is human-authored code and the list is not negotiable.
101
+
102
+ ```
103
+ ${REVIEW_FILES}
104
+ reviewed:
105
+ <repo-relative path>
106
+ ...
107
+ excluded (not yours, listed so you can see the decision):
108
+ <repo-relative path> - <reason>
109
+ ```
110
+
111
+ Return one `fileCoverage` row **per path under `reviewed`**, and no rows for
112
+ paths outside it. Rules:
113
+
114
+ - `reviewed` means you read the file's diff in full. There is no `partial`: what
115
+ you read but did not understand belongs in a `findings` entry, not in a
116
+ hedged verdict.
117
+ - `skipped` requires a `reason` naming why you could not read it - the diff was
118
+ truncated, the content was unreadable. "Not relevant" is not a skip reason; a
119
+ file you read and found nothing in is `reviewed`.
120
+ - A path that is neither read nor explicitly skipped fails the stage, for the
121
+ same reason an unanswered rule ID does: a reviewer that opened one file of ten
122
+ and one that read all ten return identical empty `findings` arrays, and this
123
+ list is what tells them apart.
124
+
125
+ When no `${REVIEW_FILES}` block is supplied, omit `fileCoverage` entirely.
126
+
93
127
  ## Priority Files (advisory)
94
128
 
95
129
  When the orchestrator passes a `${PRIORITY_FILES}` block, treat it as a heuristic
@@ -1,16 +1,16 @@
1
- # Locked decisions (36)
1
+ # Locked decisions (37)
2
2
 
3
- > The 36 Locked decisions of the analysis flow. Loaded by `/multi-agent:analysis`, by `/multi-agent:analysis-resolve` (which inherits them) and by pipeline Phase 1 when it runs the analysis engine. Numbering is canonical: cite as `Locked <n> (<short label>)`.
3
+ > The 37 Locked decisions of the analysis flow. Loaded by `/multi-agent:analysis`, by `/multi-agent:analysis-resolve` (which inherits them) and by pipeline Phase 1 when it runs the analysis engine. Numbering is canonical: cite as `Locked <n> (<short label>)`.
4
4
 
5
5
  ### Index by category (v9.1.0+)
6
6
 
7
- Browse-friendly grouping of the 36 Locked decisions. Numbering stays canonical (matches the list below); the index is read-only navigation.
7
+ Browse-friendly grouping of the 37 Locked decisions. Numbering stays canonical (matches the list below); the index is read-only navigation.
8
8
 
9
9
  | Category | Decisions | Concern |
10
10
  |---|---|---|
11
11
  | **A. Governance** | 1, 5, 6, 7, 10, 26, 27, 32, 36 | Run-level process rules: one feature per run, default output, auto-commit ban, punctuation policy, output picker timing, Pass B preview, evidence digest cache, analysis profile, document reviewed before publish |
12
12
  | **B. Citation and Evidence** | 3, 4, 8, 11, 24, 30, 34 | Every fact in the doc traces back to a source: citation discipline, forward-looking spec, standards binding, repo-evidence reuse-first, Pass B footnote mandatory, analysis self-contained (pipeline-wide), references built from the evidence record |
13
- | **C. Output Format and Structure** | 2, 9, 13, 14, 16, 17, 20, 21, 25, 33, 35 | How the document is laid out: section omission rule, per-platform output split, Gherkin user stories, Goals + Non-Goals paired, Files-to-Add tag, API response variants exhaustive, localization mode (ownership-aware), References at the bottom, Lite mode, corporate backbone always renders, stack-optional render |
13
+ | **C. Output Format and Structure** | 2, 9, 13, 14, 16, 17, 20, 21, 25, 33, 35, 37 | How the document is laid out: section omission rule, per-platform output split, Gherkin user stories, Goals + Non-Goals paired, Files-to-Add tag, API response variants exhaustive, localization mode (ownership-aware), References at the bottom, Lite mode, corporate backbone always renders, stack-optional render, redesign records v1 before planning v2 |
14
14
  | **D. Design Source and Pipeline Architecture** | 12, 22, 23 | Where design comes from and how the pipeline renders: Figma 3-tier access (BLOCKING), platform-agnostic template + Pass B render, convention extraction (Phase 1c) |
15
15
  | **E. UI, Variant, and Test Coverage** | 15, 18, 19, 28, 29, 31 | UI artefact rules: SVG default for new assets, screenshots embedded, all Figma variants drilled, SwiftUI Preview block (iOS), variant usage explicit, business-rule to acceptance-criterion to test traceability |
16
16
 
@@ -124,10 +124,11 @@ Result: `state.analysisSpec.outputs.requested[]`.
124
124
  for f in /tmp/analysis-<feature-slug>-<ts>/*.md; do
125
125
  node "$HOME/.claude/scripts/validate-analysis-doc.mjs" "$f" || GATE_FAILED=1
126
126
  node "$HOME/.claude/scripts/build-references.mjs" <state.json> --check "$f" || GATE_FAILED=1
127
+ node "$HOME/.claude/scripts/verify-citations.mjs" "$f" --repo "$REPO_ROOT" || GATE_FAILED=1
127
128
  done
128
129
  ```
129
130
 
130
- `validate-analysis-doc.mjs` enforces the mechanically-checkable Locked decisions on the emitted markdown itself (front-matter completeness, never-omitted sections per Locked 2, humanizer punctuation per Locked 7, Full-mode business-rule traceability per Locked 31, and in the corporate profile the backbone presence and `EKLENECEK`-to-Section-20 pairing per Locked 33). `build-references.mjs --check` runs the References coverage gate (Locked 34): a source the run consumed but did not list, or a listed row with no evidence behind it, blocks dispatch. Any ERROR blocks dispatch: fix the draft and re-validate. Warnings are advisory (run with `--strict` to treat them as blocking). This turns the "fails the dispatch gate" prose into a real, model-independent check.
131
+ `validate-analysis-doc.mjs` enforces the mechanically-checkable Locked decisions on the emitted markdown itself (front-matter completeness, never-omitted sections per Locked 2, humanizer punctuation per Locked 7, Full-mode business-rule traceability per Locked 31, and in the corporate profile the backbone presence and `EKLENECEK`-to-Section-20 pairing per Locked 33). `build-references.mjs --check` runs the References coverage gate (Locked 34): a source the run consumed but did not list, or a listed row with no evidence behind it, blocks dispatch. `verify-citations.mjs` resolves each claimed `file:line` at HEAD with `git cat-file`, so `Foo.swift:9999` in a repo with no Foo.swift blocks dispatch (Locked 3, Locked 37). Any ERROR blocks dispatch: fix the draft and re-validate. Warnings are advisory (run with `--strict` to treat them as blocking). This turns the "fails the dispatch gate" prose into a real, model-independent check.
131
132
 
132
133
  Iterate `state.analysisSpec.outputs.requested`. For each target:
133
134
 
@@ -77,6 +77,32 @@ Across stacks the same shape produces, for example: `LoginView.swift - ...` (i
77
77
  <none, or which service, contract or channel>
78
78
  ```
79
79
 
80
+ Part 3 is the one part of this body that is MEASURED rather than recalled. The
81
+ code graph already answers it, so draw the answer instead of re-typing it:
82
+
83
+ ```bash
84
+ node "$HOME/.claude/scripts/graph-mermaid.mjs" "<changed symbol[,symbol]>"
85
+ ```
86
+
87
+ Append the fenced block it prints under part 3, above the prose. GitHub renders
88
+ mermaid natively in pull requests, so this costs no renderer and no plugin. The
89
+ prose stays: the diagram says which symbols the change reaches, the sentence says
90
+ which screens and flows a tester must open, and neither answers the other.
91
+
92
+ Exit 1 means the repo has no graph yet (`/multi-agent:graph` builds it) or the
93
+ symbol is not in it. That is a gap with a reason, not a failure: write the prose
94
+ alone and say the graph was unavailable. Never hand-draw the diagram - a drawn
95
+ blast radius nobody measured is worse than none, because a diagram is read as
96
+ fact.
97
+
98
+ The commit line the script prints stays with it. A graph built before the change
99
+ draws the radius of an older tree, and the reader has no other way to notice.
100
+
101
+ This is a GitHub-only section. `channels/jira.md` has no mermaid handling at all:
102
+ a fence there converts to a literal `{code:mermaid}` block, so the Jira impact
103
+ section keeps its prose. Confluence renders it through the `ac:name="mermaid"`
104
+ macro (`md2confluence-v3.py`) when the space has the plugin.
105
+
80
106
  When the change deliberately fixes part of a wider problem, a closing **Risk and remaining scope** paragraph names what is still open and why it was left - a reviewer who can see the rest of the pattern in the repo will ask otherwise, and the honest answer is cheaper written down than defended in a thread.
81
107
 
82
108
  **`test_scenarios`** - the same titled-scenario shape the Jira adapter uses, so the tester reads one list on both surfaces, with symbols allowed here:
@@ -214,6 +214,28 @@ on skill directories would demand exactly the layout that breaks it.
214
214
 
215
215
  Future changes that break an item in the "stay identical" list must update **both** files in the same commit. `smoke-cross-cli-behavior.sh` enforces the identity-preserving axis (input parsing, routing, output shape); structural differences are left to manual review because enforcing them would require forcing the files to the same shape, which we intentionally don't want.
216
216
 
217
+ ### Panel diversity per host
218
+
219
+ Phase 4 runs three reviewers everywhere, but the diversity those three buy is not the
220
+ same on every host. Copilot CLI gets cross-VENDOR disagreement for free: GPT-5.4 sits
221
+ beside two Claude models. Claude Code and Codex each run a one-vendor panel - three
222
+ Anthropic models on one, three OpenAI models on the other - so the same three-way
223
+ agreement is weaker evidence there, and Phase 4 says so in the triage note on a
224
+ borderline finding.
225
+
226
+ Where the budget goes instead, when vendor diversity is unavailable:
227
+
228
+ | Host | Reviewer 1 | Reviewer 2 | Reviewer 3 |
229
+ |---|---|---|---|
230
+ | Copilot CLI | Fable/Opus, security + architecture | GPT-5.4, edge cases (cross-vendor) | Sonnet, quality |
231
+ | Claude Code | Fable, security + architecture | Opus, edge cases | Sonnet, quality |
232
+ | Codex | `xhigh`, security + architecture | a different family member, edge cases | `medium`, quality |
233
+
234
+ On Codex the axis is reasoning effort as much as model identity, because the family
235
+ members available there are closer to each other than Fable and Sonnet are. That is a
236
+ weaker axis, not an equivalent one, and treating it as equivalent is the error this
237
+ section exists to prevent.
238
+
217
239
  ## 3. Frontmatter Transform Rules (Claude ↔ Copilot)
218
240
 
219
241
  Each file has a different frontmatter schema. The sync flow transforms between them:
@@ -1,5 +1,11 @@
1
1
  ## Code Graph (Phase 1 Step 2.6 + Phase 7 Step 3)
2
2
 
3
+ <!-- toc -->
4
+ - [Phase 1 Step 2.6 - query before dispatching Explore](#phase-1-step-26---query-before-dispatching-explore)
5
+ - [Phase 7 Step 3 - refresh after the branch changed code](#phase-7-step-3---refresh-after-the-branch-changed-code)
6
+ - [The graph is drawable, and one place already asks for it](#the-graph-is-drawable-and-one-place-already-asks-for-it)
7
+ <!-- /toc -->
8
+
3
9
  A deterministic, LLM-free map of what a repo declares and what refers to what,
4
10
  written to `~/.claude/knowledge/<project>/code-graph.json`. Gated by
5
11
  `prefs.global.codeGraph.enabled` (default `false`); with it off, Phase 1 and
@@ -67,3 +73,37 @@ The rebuild costs no API tokens, so it runs every task rather than on a stalenes
67
73
  heuristic. A non-zero validator exit keeps the previous graph and logs
68
74
  `knowledge.graph_invalid`; it never fails the run - a stale graph is a degraded
69
75
  Phase 1, not a broken deliverable.
76
+
77
+ ### The graph is drawable, and one place already asks for it
78
+
79
+ The PR body's Impact Analysis, part 3, asks which symbols and files a change
80
+ reaches. That is `graph-affected.mjs`'s question, and until now the answer was
81
+ re-typed as prose by a model while the measurement sat on disk unread.
82
+
83
+ ```bash
84
+ node $HOME/.claude/scripts/graph-mermaid.mjs "<symbol[,symbol]>" [--depth N] [--max-nodes N]
85
+ ```
86
+
87
+ It emits a fenced `flowchart` and nothing else - no renderer, no plugin, no
88
+ dependency, because mermaid is text and GitHub renders it natively in pull
89
+ requests, issues and markdown files. Traversal is not reimplemented: `findByName`
90
+ and `affected` are imported from `graph-affected.mjs`, so the diagram and the
91
+ text report cannot disagree about what is affected.
92
+
93
+ Three properties that are enforced rather than promised
94
+ (`smoke-graph-mermaid.sh`):
95
+
96
+ - Every drawn node and edge resolves back into `code-graph.json`, with the edge
97
+ kind it claims. A diagram is read as fact and checked less than prose, so an
98
+ invented edge is the expensive failure.
99
+ - Over `--max-nodes` the leftover count is printed inside the diagram, not
100
+ dropped. A small picture of a large blast radius reads as reassurance.
101
+ - The graph's `baseCommit` is printed beside it. A graph built before the change
102
+ draws an older tree, and nothing else in the PR would reveal that.
103
+
104
+ Exit 1 with a reason on stderr means no graph or no such symbol. The caller
105
+ records the gap and writes the prose alone; it never hand-draws a replacement.
106
+
107
+ Jira is not a target: its renderer turns the fence into a literal
108
+ `{code:mermaid}` block. Confluence renders it through the `ac:name="mermaid"`
109
+ macro when the space carries the plugin (`channels/confluence.md`).
@@ -1,6 +1,7 @@
1
1
  # Design conformance - the component walk, and why a glance is not a pass
2
2
 
3
3
  <!-- toc -->
4
+ - [0. Why this runs as a gate](#0-why-this-runs-as-a-gate)
4
5
  - [1. Enumerate first, then fill every cell](#1-enumerate-first-then-fill-every-cell)
5
6
  - [2. Measure, never read the token](#2-measure-never-read-the-token)
6
7
  - [3. Adaptive per-component convergence](#3-adaptive-per-component-convergence)
@@ -24,6 +25,19 @@ Consumers: `/multi-agent:design-check` (the runner), Phase 4 review when a UI
24
25
  diff is under review, and `features/visual-evidence.md` when a capture has to
25
26
  prove a fix. Gate: `smoke-design-conformance.sh`.
26
27
 
28
+ ## 0. Why this runs as a gate
29
+
30
+ `design-check` existed as a command for a while with no phase invoking it, so the only
31
+ thing standing between a build and visual drift was the user opening the app and
32
+ looking. On one run that produced 16pt padding where the frame said `Spacing/12`, and a
33
+ full sheet rebuild afterwards.
34
+
35
+ The reason it cannot be advice is structural, not historical: a reviewer reading a diff
36
+ cannot see spacing. Every other Phase 4 check reads text and reasons about text; this
37
+ one is the only thing in the pipeline that compares a rendered result against the
38
+ design it was drawn from. Left optional, it is the check that gets skipped on exactly
39
+ the runs that are in a hurry, which are the runs that produce drift.
40
+
27
41
  ## 1. Enumerate first, then fill every cell
28
42
 
29
43
  Do **not** walk by finding, and do not walk by headline. Build the inventory
@@ -0,0 +1,132 @@
1
+ # The review file set - what was read, and what was not, on the record
2
+
3
+ <!-- toc -->
4
+ - [Why this exists](#why-this-exists)
5
+ - [The set is fixed before the reviewer sees the diff](#the-set-is-fixed-before-the-reviewer-sees-the-diff)
6
+ - [Every exclusion names the pattern that produced it](#every-exclusion-names-the-pattern-that-produced-it)
7
+ - [The reviewer answers for each file](#the-reviewer-answers-for-each-file)
8
+ - [Invocation](#invocation)
9
+ - [What this is not](#what-this-is-not)
10
+ <!-- /toc -->
11
+
12
+ > The denominator for Phase 4's other axis. Loaded on demand by
13
+ > `/multi-agent:review` and by pipeline Phase 4.
14
+
15
+ ## Why this exists
16
+
17
+ Phase 4 had a size cap and no exclusion list. When the diff exceeds the phase
18
+ token allowance the cap truncates the LARGEST files first, so a regenerated
19
+ lockfile or a snapshot dump is not merely wasted budget: it is the thing that
20
+ survives while real code is cut.
21
+
22
+ The second half is worse and quieter. A reviewer that opened one file of ten and
23
+ a reviewer that read all ten and found nothing return the identical
24
+ `{"findings": [], "approved": true}`. Nothing in the pipeline could tell them
25
+ apart, so "no findings" has been carrying two meanings at once.
26
+
27
+ Both halves are one fix: decide what is worth reading BEFORE the cap decides
28
+ what fits, and make the reviewer answer for each file it was given.
29
+
30
+ ## The set is fixed before the reviewer sees the diff
31
+
32
+ The same rule `selectedRules[]` follows, for the same reason. After a model has
33
+ seen the diff, "I did not open that one" and "there was nothing there" become
34
+ the same sentence, and whichever one is cheaper to say is the one that gets
35
+ said. So the set is computed from `git diff --name-only`, written to
36
+ `.pipeline/review-files.json`, and never recomputed inside the round.
37
+
38
+ ```bash
39
+ git -C "$WORKTREE" diff --name-only "$BASE_BRANCH"...HEAD \
40
+ | node $HOME/.claude/scripts/review-file-filter.mjs \
41
+ > "$WORKTREE/.pipeline/review-files.json"
42
+ ```
43
+
44
+ Report shape:
45
+
46
+ ```json
47
+ {
48
+ "reviewed": ["src/App.swift"],
49
+ "excluded": [
50
+ {
51
+ "path": "package-lock.json",
52
+ "reason": "lockfile - resolved by the package manager, not written by hand",
53
+ "pattern": "**/package-lock.json"
54
+ }
55
+ ],
56
+ "total": 2,
57
+ "patternsSource": ".../schemas/review-file-exclusions.json",
58
+ "patternCount": 45
59
+ }
60
+ ```
61
+
62
+ `reviewed` is what goes into the diff cap and into the reviewer prompt, as a
63
+ `${REVIEW_FILES}` block beside `${CRITERIA}` in the shared cache prefix. It has to
64
+ be in the prompt: a reviewer asked to account for a set it was never shown can
65
+ only guess, and `fileCoverage` would then fail on every single dispatch - a gate
66
+ that always fires is a gate that gets switched off. `excluded` goes into the run
67
+ report, never into silence.
68
+
69
+ ## Every exclusion names the pattern that produced it
70
+
71
+ A file that disappears between the diff and the review is indistinguishable from
72
+ a file nobody found anything in - which is the exact confusion this whole
73
+ feature exists to remove, so reintroducing it in the filter would be
74
+ self-defeating. Each excluded row carries both the human reason and the glob
75
+ that matched, so a reviewer, a PR reader or a future maintainer can dispute the
76
+ call rather than discover it.
77
+
78
+ The pattern list is data, in `schemas/review-file-exclusions.json`, and it is
79
+ generic: generated trees, lockfiles, recorded snapshots, vendored source, build
80
+ output, binary assets. No stack, project or company name appears in it. A
81
+ pattern with no reason invalidates the whole list rather than being defaulted -
82
+ the default would be exactly the sentence the caller is supposed to print.
83
+
84
+ **It fails open, on purpose.** An unreadable or malformed pattern file yields
85
+ every file reviewed, the reason on stderr, and exit 2. Failing closed would
86
+ review nothing and report a clean run.
87
+
88
+ ## The reviewer answers for each file
89
+
90
+ `reviewer-output.schema.json` (v1.3.0) carries `fileCoverage[]`: one row per
91
+ path in `reviewed`, `{path, verdict: reviewed|skipped, reason}`. There is
92
+ deliberately no `partial` - a file read in part is read, and what was not
93
+ understood belongs in a finding.
94
+
95
+ `validate-reviewer.mjs --coverage <report>` enforces it, catching the same three
96
+ failures the conformance checklist catches on the rule axis:
97
+
98
+ | Failure | Why it matters |
99
+ |---|---|
100
+ | a file in the set with no row | silently unread, and the empty `findings[]` reads as clean |
101
+ | a row for a path outside the set | an answer about something the reviewer was not given, the same shape as a hallucinated rule ID |
102
+ | `skipped` with no reason | a drop with no cause is indistinguishable from a read |
103
+
104
+ `skipped` is legitimate and expected: a file past the diff cap, a file whose
105
+ content the host truncated. What it may not be is unexplained. "Not relevant" is
106
+ a review decision and belongs in a verdict of `reviewed`, not a skip.
107
+
108
+ An empty `reviewed` set (a diff that is entirely lockfiles) demands no checklist
109
+ at all. Requiring an empty array there would fail honest output, and the
110
+ filter's `excluded[]` is what carries that information onward.
111
+
112
+ ## Invocation
113
+
114
+ ```bash
115
+ node $HOME/.claude/scripts/validate-reviewer.mjs "$REVIEWER_FILE" \
116
+ --criteria "$WORKTREE/.pipeline/criteria-manifest.json" \
117
+ --coverage "$WORKTREE/.pipeline/review-files.json"
118
+ ```
119
+
120
+ Without `--coverage` the field stays optional, so every existing caller keeps
121
+ working unchanged. With it, the checklist is enforced and exit 1 takes the same
122
+ single self-correction rework the rest of the validator gate takes.
123
+
124
+ ## What this is not
125
+
126
+ It is not a relevance filter. Nothing here decides that a file is uninteresting;
127
+ it decides that a file is not human-authored source, which is a mechanical
128
+ question with a mechanical answer. The moment a pattern starts encoding "we
129
+ probably do not care about this directory", the list has become a way to hide
130
+ work, and the near-miss assertions in `smoke-review-file-filter.sh`
131
+ (`CodeGenerator.swift`, `generated-report.md`, `distribution/`, `buildSrc/`) are
132
+ what fail when it does.
@@ -240,8 +240,7 @@ $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.criteria \
240
240
  Output conforms to `$HOME/.claude/schemas/criteria-manifest.schema.json`. The four parts that matter downstream:
241
241
 
242
242
  - **`selectedRules[]` is the denominator.** Rule IDs from every registry whose declared `scope` matches the diff, persisted BEFORE the reviewers run so the set cannot be renegotiated after one has seen the diff. That is what makes "applied completely" answerable rather than "looks fine": a reviewer finding nothing must still return a verdict per ID. Discovery is declared via `standards-registry:` frontmatter, so no stack-specific skill is named here.
243
- - **`coverage.declaredGaps[]` + `droppedReasons`.** An uncovered language, and every out-of-scope rule, is reported with a reason. "No rule applied", "every rule passed" and "the rule set was narrowed" must never render the same.
244
- - **`ledger`.** `state.telemetry.skillCalls[]` corroborates only; `ledger.source` defaults to `derived` and coverage is never computed from self-report. A declared skill the resolver cannot bind is flagged.
243
+ - **`coverage.declaredGaps[]`, `droppedReasons`, `ledger`.** Every gap, dropped rule and unbindable declared skill is reported with a reason; coverage is never computed from self-report. Rationale: the contract above.
245
244
  - **`findings[]`.** Reviewer-shaped, merging at Step 3.0 alongside test-integrity: expired / unexplained / unknown-ID exception markers, plus any registry whose delegated linter is not wired here (those rules are unverified, so reporting no violations reports that nothing was measured).
246
245
 
247
246
  **Any non-zero exit halts** (1 = setup error incl. a bad `--skills-root`; 2 = no root resolved, unparseable registry, or a declared path escaping its skill dir): continuing would drop a rule set from the denominator, and since the reviewer validator skips the checklist when zero rules were selected, the run would report clean over criteria never loaded. A coverage gap halts only under `prefs.global.skillConformance.blockOnCoverageGap`. No opt-out for the stage or the exception-expiry check, on the same grounds as Step 1.76.
@@ -266,13 +265,25 @@ Visual-fidelity mismatches against the captured screenshot are BLOCKING findings
266
265
 
267
266
  When `state.figmaAccess.tier === 3` (user-attached screenshot, no Code Connect snippet), the reviewer additionally sets `findings[i].severity = "blocking"` and `findings[i].tag = "review_blocking_tier3"` on every UI atom that lacks a confirmed canonical-component mapping. The triage step preserves these findings unless the user has explicitly cleared the open question.
268
267
 
268
+ #### Step 1.85 - Review file set (the denominator, required)
269
+
270
+ Decides what the reviewers read before the cap decides what fits: the cap truncates the largest files first, so a lockfile survives while real code is cut. Zero LLM. Full contract: `$HOME/.claude/multi-agent-refs/features/review-file-set.md`.
271
+
272
+ ```bash
273
+ git -C "$WORKTREE" diff --name-only "$BASE_BRANCH"...HEAD \
274
+ | node $HOME/.claude/scripts/review-file-filter.mjs \
275
+ > "$WORKTREE/.pipeline/review-files.json"
276
+ ```
277
+
278
+ `reviewed[]` is the denominator, fixed here for the same reason `selectedRules[]` is. `excluded[]` reaches the run report with its reason and the glob that matched. Exit 2 means the pattern list is unreadable and everything is reviewed: continue, logging `review.file_filter_failed`.
279
+
269
280
  #### Step 1.9 - Context economy (cache prefix + diff cap)
270
281
 
271
282
  Phase 4 sends the same diff to every reviewer and then to triage, so the diff is the dominant token cost. Two measures keep it bounded:
272
283
 
273
- **Shared cache prefix.** Build the reviewer and triage prompts so the large invariant context - the full diff, the `${CRITERIA}` block from Step 1.78, the Phase 1 analysis summary, the Phase 2 plan - is a byte-identical leading block across all dispatches in this iteration. Only the per-reviewer focus + skill line varies, and it goes AFTER the shared block. `${CRITERIA}` goes in the prefix, identical for every reviewer: subsetting it per reviewer would invalidate the prefix for the whole panel and re-bill the largest block in the phase. Per-reviewer emphasis stays a one-line pointer in the suffix. When the host supports prompt caching, the 2nd/3rd reviewer and the triage call then read that prefix at the discounted cache-read rate instead of re-billing it as fresh input. Forward the host-reported cache-read count as `tokens_cached` per the Token telemetry contract so the saving lands in the cost ledger. The `<scope-self-check>` block and, from iteration 2, the `<previous-round-findings>` block (Step 2.1) close the shared block, after the plan and before the per-reviewer suffix.
284
+ **Shared cache prefix.** Build the reviewer and triage prompts so the large invariant context - the full diff, the `${CRITERIA}` block from Step 1.78, the `${REVIEW_FILES}` list from Step 1.85, the Phase 1 analysis summary, the Phase 2 plan - is a byte-identical leading block across all dispatches in this iteration. Only the per-reviewer focus + skill line varies, and it goes AFTER the shared block. `${CRITERIA}` goes in the prefix, identical for every reviewer: subsetting it per reviewer would invalidate the prefix for the whole panel and re-bill the largest block in the phase. Per-reviewer emphasis stays a one-line pointer in the suffix. When the host supports prompt caching, the 2nd/3rd reviewer and the triage call then read that prefix at the discounted cache-read rate instead of re-billing it as fresh input. Forward the host-reported cache-read count as `tokens_cached` per the Token telemetry contract so the saving lands in the cost ledger. The `<scope-self-check>` block and, from iteration 2, the `<previous-round-findings>` block (Step 2.1) close the shared block, after the plan and before the per-reviewer suffix.
274
285
 
275
- **Single-repo diff cap.** If the diff exceeds the Phase 4 token allowance (`token-budget.json`), truncate the largest files and append a footer `[truncated - full diff in file://$WORKTREE/.review-diff.txt]`, writing the full diff to that path. Reviewers and triage receive the same capped view + the marker so they can flag "review the full diff manually." Log `review.diff_truncated bytes_dropped=<N>`. (Multi-repo already caps the combined diff at 80% of budget; this is the single-repo equivalent.)
286
+ **Single-repo diff cap.** Applied to the Step 1.85 `reviewed[]` set only. If it exceeds the Phase 4 token allowance (`token-budget.json`), truncate the largest files and append a footer `[truncated - full diff in file://$WORKTREE/.review-diff.txt]`, writing the full diff to that path. Reviewers and triage receive the same capped view + the marker so they can flag "review the full diff manually." Log `review.diff_truncated bytes_dropped=<N>`. (Multi-repo already caps the combined diff at 80% of budget; this is the single-repo equivalent.)
276
287
 
277
288
  #### Step 2 - Parallel AI Review (CLI-aware reviewer set)
278
289
 
@@ -291,9 +302,8 @@ Reviewer count per host: **Claude Code 3, Copilot CLI 3, Codex CLI 3** - **2**
291
302
 
292
303
  #### Codex CLI - two constraints that fail silently
293
304
 
294
- Both were measured against Codex 0.145, not inferred, and both produce a review that
295
- looks like it ran. The full statement lives in the managed block at `~/.codex/AGENTS.md`
296
- (always loaded on that host, so it is not restated here):
305
+ Both measured against Codex 0.145, both produce a review that looks like it ran. Full
306
+ statement: the managed block in `~/.codex/AGENTS.md`, always loaded on that host.
297
307
 
298
308
  1. **`fork_turns: "none"` on every `spawn_agent` that sets `model` or
299
309
  `reasoning_effort`** - a full-history fork discards the override and collapses the
@@ -304,15 +314,10 @@ looks like it ran. The full statement lives in the managed block at `~/.codex/AG
304
314
  Sub-agent delegation itself is authorized by that same managed block; without it Phase 4
305
315
  degrades to a single in-thread review.
306
316
 
307
- **Single-vendor caveat.** Every Codex reviewer is an OpenAI model, and every Claude
308
- Code reviewer is an Anthropic model, so the cross-vendor disagreement that Copilot CLI
309
- gets for free (GPT-5.4 beside two Claude models) is absent on both. The diversity budget
310
- shifts to model generation, reasoning effort and persona focus: on Codex Reviewer 1 runs
311
- `xhigh` on security and architecture, Reviewer 2 runs a different model family member
312
- on edge cases, Reviewer 3 runs `medium` on quality; on Claude Code the three slots are
313
- three different Claude tiers. Treat consensus among a single-vendor panel as weaker
314
- evidence than the same consensus on Copilot CLI, and say so in the triage note when all
315
- three agree on a borderline finding.
317
+ **Single-vendor caveat.** Claude Code and Codex both run a one-vendor panel, so their
318
+ consensus is weaker evidence than Copilot CLI's; say so in the triage note on a
319
+ borderline finding. Where the diversity budget goes instead:
320
+ `cross-cli-contract.md`, "Panel diversity per host".
316
321
 
317
322
  Each reviewer inherits the `code-reviewer` agent's focus areas (Security, Architecture, Quality, Performance) and output contract. The orchestrator overrides only the model and the stack-specific skill per-reviewer - no prompt duplication.
318
323
 
@@ -331,7 +336,7 @@ Each reviewer inherits the `code-reviewer` agent's focus areas (Security, Archit
331
336
 
332
337
  ##### 2.1 Previous-round findings (iteration >= 2) and 2.2 scope self-check (every iteration)
333
338
 
334
- A reviewer has no memory of the round before, so it rediscovers last round's findings in new words. From iteration 2, render the previous round's accepted blocking/important findings (`.pipeline/triage-round-$((ITERATION-1)).json`, max 40) into a `<previous-round-findings>` block at the end of the shared prefix: a still-present issue is reported with the SAME fingerprint and the current line, a fixed one is omitted, anything new leaves `fingerprint` unset. Every iteration also renders `.pipeline/scope-check.json` (Phase 3 Step 3.7) plus `scope-check-gate.mjs --advisory` output as `<scope-self-check>`: file reasons, unjustified files, and `notDone[]` (never re-raised as findings); a missing record logs `review.scope_check=missing`. Block text and recipes: `$HOME/.claude/multi-agent-refs/features/review-delta.md`.
339
+ From iteration 2, render the previous round's accepted blocking/important findings (`.pipeline/triage-round-$((ITERATION-1)).json`, max 40) into a `<previous-round-findings>` block at the end of the shared prefix. Every iteration also renders `.pipeline/scope-check.json` (Phase 3 Step 3.7) plus `scope-check-gate.mjs --advisory` output as `<scope-self-check>`: file reasons, unjustified files, and `notDone[]` (never re-raised as findings); a missing record logs `review.scope_check=missing`. Block text and recipes: `$HOME/.claude/multi-agent-refs/features/review-delta.md`.
335
340
 
336
341
  #### Step 2.8 - Visual conformance gate (component / screen work only)
337
342
 
@@ -350,11 +355,7 @@ disk with `Code Connect: Not published` in Figma means the binding does not exis
350
355
  anyone but the author. Assert the publish step ran; an unpublished binding is a
351
356
  blocking finding.
352
357
 
353
- Why this is a gate and not advice: `design-check` existed as a command for a while
354
- with **no phase invoking it**, so the only thing standing between a build and visual
355
- drift was the user opening the app and looking. On one run that produced 16pt padding
356
- where the frame said `Spacing/12`, and a full sheet rebuild afterwards. A reviewer
357
- reading a diff cannot see spacing; something has to compare against the design.
358
+ Why this is a gate and not advice: `design-conformance.md`, "Why this runs as a gate".
358
359
 
359
360
  Skip only when the diff has no UI change. Record the outcome in
360
361
  `consensus.visualConformance` so Phase 7 reports whether it ran.
@@ -374,10 +375,11 @@ Step 2 produces N reviewer-output objects (one per dispatched reviewer), each co
374
375
  ```json
375
376
  {"findings":[{"severity":"blocking|important|suggestion","file":"...","line":N,"issue":"...","fix":"...","ruleId":"SEC-01","criteriaSource":"ios-coding-standard"}],
376
377
  "conformance":[{"ruleId":"SEC-01","verdict":"conformant|violated|not-applicable","file":"...","line":N,"reason":"..."}],
378
+ "fileCoverage":[{"path":"src/App.swift","verdict":"reviewed|skipped","reason":"..."}],
377
379
  "approved":true|false}
378
380
  ```
379
381
 
380
- `ruleId` + `criteriaSource` appear on a finding that cites a rule from `${CRITERIA}`. `conformance` is required whenever Step 1.78 selected at least one rule, with exactly one row per selected ID and none outside the set.
382
+ `ruleId` + `criteriaSource` appear on a finding that cites a rule from `${CRITERIA}`. `conformance` is required whenever Step 1.78 selected at least one rule, and `fileCoverage` whenever Step 1.85 left at least one file in `reviewed[]`: one row per ID, one row per path, none outside either set.
381
383
 
382
384
  **Required: validator gate (deterministic) - run immediately after each reviewer returns, before merging findings.** Persist each reviewer's output and validate the file - the validator's exit code decides, not the LLM turn:
383
385
 
@@ -386,9 +388,15 @@ REVIEWER_FILE="$WORKTREE/.pipeline/reviewer-$N.json"
386
388
  printf '%s' "$REVIEWER_JSON" > "$REVIEWER_FILE"
387
389
  node $HOME/.claude/scripts/validate-reviewer.mjs "$REVIEWER_FILE" \
388
390
  --criteria "$WORKTREE/.pipeline/criteria-manifest.json" \
391
+ --coverage "$WORKTREE/.pipeline/review-files.json" \
392
+ && node $HOME/.claude/scripts/verify-citations.mjs "$REVIEWER_FILE" --repo "$WORKTREE" --worktree \
389
393
  && node $HOME/.claude/scripts/finding-fingerprint.mjs annotate --in-place "$REVIEWER_FILE"
390
394
  ```
391
395
 
396
+ `verify-citations.mjs` resolves each finding's `file:line` against the checkout,
397
+ not a commit: a round's fix is uncommitted, and HEAD would call it invented.
398
+ Exit 1 takes the single rework below; exit 2 means not a repository.
399
+
392
400
  Progress line: ` → checking validator validate-reviewer ({reviewer})`
393
401
 
394
402
  `finding-fingerprint.mjs` stamps each finding with its cross-round id once the validator passes; an echoed one is kept, and anonymization leaves it intact.