@mmerterden/multi-agent-pipeline 18.0.0 → 19.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (234) hide show
  1. package/CHANGELOG.md +287 -0
  2. package/README.md +36 -20
  3. package/README.tr.md +14 -16
  4. package/docs/adr/0002-instruction-driven-flag.md +1 -0
  5. package/docs/adr/0005-lazy-phase-docs.md +11 -1
  6. package/docs/adr/0008-installer-modularization-and-secret-leak-defense.md +1 -0
  7. package/docs/adr/0010-own-code-graph.md +1 -0
  8. package/docs/adr/0014-six-phase-consolidation.md +134 -0
  9. package/docs/adr/README.md +2 -1
  10. package/docs/architecture.md +37 -38
  11. package/docs/best-practices.md +1 -1
  12. package/docs/ecosystem.md +46 -27
  13. package/docs/engineering.md +1 -1
  14. package/docs/facts.json +61 -0
  15. package/docs/features.md +55 -54
  16. package/docs/performance.md +5 -5
  17. package/docs/recovery-guide.md +17 -17
  18. package/docs/token-budget-history.md +3 -1
  19. package/index.js +2 -2
  20. package/install/_codex-agents.mjs +1 -1
  21. package/install/templates/claude-hooks.json +1 -1
  22. package/install/templates/codex-instructions.md +1 -1
  23. package/install/templates/copilot-instructions.md +28 -28
  24. package/manifest.json +234 -216
  25. package/package.json +2 -2
  26. package/pipeline/agents/dev-critic.md +7 -7
  27. package/pipeline/commands/figma-to-swiftui.md +1 -1
  28. package/pipeline/commands/multi-agent/SKILL.md +9 -9
  29. package/pipeline/commands/multi-agent/analysis/SKILL.md +15 -15
  30. package/pipeline/commands/multi-agent/analysis-resolve/SKILL.md +2 -2
  31. package/pipeline/commands/multi-agent/autopilot/SKILL.md +7 -7
  32. package/pipeline/commands/multi-agent/channels/SKILL.md +15 -15
  33. package/pipeline/commands/multi-agent/diff-explain/SKILL.md +6 -6
  34. package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +1 -1
  35. package/pipeline/commands/multi-agent/graph/SKILL.md +1 -1
  36. package/pipeline/commands/multi-agent/help/SKILL.md +62 -62
  37. package/pipeline/commands/multi-agent/language/SKILL.md +2 -2
  38. package/pipeline/commands/multi-agent/local/SKILL.md +11 -11
  39. package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +13 -13
  40. package/pipeline/commands/multi-agent/log/SKILL.md +2 -2
  41. package/pipeline/commands/multi-agent/manual-test/SKILL.md +9 -9
  42. package/pipeline/commands/multi-agent/model/SKILL.md +69 -0
  43. package/pipeline/commands/multi-agent/refactor/SKILL.md +3 -3
  44. package/pipeline/commands/multi-agent/resume/SKILL.md +4 -4
  45. package/pipeline/commands/multi-agent/resume-local/SKILL.md +19 -17
  46. package/pipeline/commands/multi-agent/review/SKILL.md +2 -2
  47. package/pipeline/commands/multi-agent/review-analysis/SKILL.md +1 -1
  48. package/pipeline/commands/multi-agent/route-off/SKILL.md +36 -0
  49. package/pipeline/commands/multi-agent/route-on/SKILL.md +74 -0
  50. package/pipeline/commands/multi-agent/route-status/SKILL.md +56 -0
  51. package/pipeline/commands/multi-agent/setup/SKILL.md +2 -2
  52. package/pipeline/commands/multi-agent/status/SKILL.md +5 -5
  53. package/pipeline/commands/multi-agent/steer/SKILL.md +2 -2
  54. package/pipeline/commands/multi-agent/sync/SKILL.md +12 -13
  55. package/pipeline/commands/multi-agent/test/SKILL.md +1 -1
  56. package/pipeline/lib/credential-inventory.sh +1 -1
  57. package/pipeline/lib/fetch-fortify.sh +1 -1
  58. package/pipeline/lib/model-dispatch.sh +140 -0
  59. package/pipeline/lib/model-rung.sh +142 -0
  60. package/pipeline/lib/outbound-gate.mjs +14 -0
  61. package/pipeline/lib/phase-schema.mjs +88 -0
  62. package/pipeline/lib/plan-todos.sh +5 -5
  63. package/pipeline/lib/route-state.sh +161 -0
  64. package/pipeline/lib/run-paths.sh +2 -2
  65. package/pipeline/multi-agent-refs/_account-picker.md +1 -1
  66. package/pipeline/multi-agent-refs/_dev-context.md +6 -6
  67. package/pipeline/multi-agent-refs/_input-parser.md +1 -1
  68. package/pipeline/multi-agent-refs/analysis/evidence.md +2 -11
  69. package/pipeline/multi-agent-refs/analysis/intake.md +7 -7
  70. package/pipeline/multi-agent-refs/analysis/locked.md +48 -22
  71. package/pipeline/multi-agent-refs/analysis/redesign.md +1 -1
  72. package/pipeline/multi-agent-refs/analysis/render.md +10 -10
  73. package/pipeline/multi-agent-refs/analysis/resolve.md +1 -1
  74. package/pipeline/multi-agent-refs/analysis/review.md +2 -2
  75. package/pipeline/multi-agent-refs/analysis/synthesis.md +13 -7
  76. package/pipeline/multi-agent-refs/analysis-template-corporate.md +9 -9
  77. package/pipeline/multi-agent-refs/analysis-template.md +19 -19
  78. package/pipeline/multi-agent-refs/android-guide.md +1 -1
  79. package/pipeline/multi-agent-refs/audit-guide.md +13 -13
  80. package/pipeline/multi-agent-refs/channels/issue-comment.md +2 -2
  81. package/pipeline/multi-agent-refs/channels/jira.md +3 -3
  82. package/pipeline/multi-agent-refs/channels/pr.md +4 -4
  83. package/pipeline/multi-agent-refs/channels/wiki.md +1 -1
  84. package/pipeline/multi-agent-refs/component-dispatch.md +8 -8
  85. package/pipeline/multi-agent-refs/conventions-defaults.md +2 -2
  86. package/pipeline/multi-agent-refs/cross-cli-contract.md +31 -6
  87. package/pipeline/multi-agent-refs/features/analysis-jira.md +1 -1
  88. package/pipeline/multi-agent-refs/features/autopilot-circuit-breaker.md +4 -4
  89. package/pipeline/multi-agent-refs/features/code-graph.md +5 -5
  90. package/pipeline/multi-agent-refs/features/design-conformance.md +1 -1
  91. package/pipeline/multi-agent-refs/features/dev-critic.md +3 -3
  92. package/pipeline/multi-agent-refs/features/doctor.md +3 -3
  93. package/pipeline/multi-agent-refs/features/external-context-injection.md +3 -3
  94. package/pipeline/multi-agent-refs/features/maturity-followup.md +3 -3
  95. package/pipeline/multi-agent-refs/features/model-fallback.md +41 -5
  96. package/pipeline/multi-agent-refs/features/plan-todos.md +1 -1
  97. package/pipeline/multi-agent-refs/features/repo-map.md +1 -1
  98. package/pipeline/multi-agent-refs/features/review-delta.md +3 -3
  99. package/pipeline/multi-agent-refs/features/review-multi-repo.md +2 -2
  100. package/pipeline/multi-agent-refs/features/scope-check.md +4 -4
  101. package/pipeline/multi-agent-refs/features/skill-conformance.md +2 -2
  102. package/pipeline/multi-agent-refs/features/stack-skill-routing.md +1 -1
  103. package/pipeline/multi-agent-refs/features/url-enrichment.md +1 -1
  104. package/pipeline/multi-agent-refs/features/verify-by-test.md +4 -4
  105. package/pipeline/multi-agent-refs/features/visual-evidence.md +19 -19
  106. package/pipeline/multi-agent-refs/features/worktree-finalize.md +6 -6
  107. package/pipeline/multi-agent-refs/issue-jira-triad.md +10 -10
  108. package/pipeline/multi-agent-refs/knowledge.md +11 -11
  109. package/pipeline/multi-agent-refs/multi-repo-integration-build.md +13 -13
  110. package/pipeline/multi-agent-refs/payload-contracts.md +8 -8
  111. package/pipeline/multi-agent-refs/phases/log-format.md +10 -10
  112. package/pipeline/multi-agent-refs/phases/modes.md +30 -30
  113. package/pipeline/multi-agent-refs/phases/operations.md +8 -8
  114. package/pipeline/multi-agent-refs/phases/phase-0-init.md +24 -24
  115. package/pipeline/multi-agent-refs/phases/phase-1-plan.md +599 -0
  116. package/pipeline/multi-agent-refs/phases/{phase-3-dev.md → phase-2-dev.md} +129 -49
  117. package/pipeline/multi-agent-refs/phases/{phase-4-review.md → phase-3-review.md} +225 -107
  118. package/pipeline/multi-agent-refs/phases/{phase-6-commit.md → phase-4-commit.md} +23 -23
  119. package/pipeline/multi-agent-refs/phases/{phase-7-report.md → phase-5-report.md} +29 -29
  120. package/pipeline/multi-agent-refs/phases.md +44 -48
  121. package/pipeline/multi-agent-refs/picker-contract.md +1 -1
  122. package/pipeline/multi-agent-refs/progress-contract.md +6 -6
  123. package/pipeline/multi-agent-refs/readiness-review.md +1 -1
  124. package/pipeline/multi-agent-refs/rules.md +7 -7
  125. package/pipeline/multi-agent-refs/swiftui-guide.md +2 -2
  126. package/pipeline/multi-agent-refs/tracker-contract.md +31 -32
  127. package/pipeline/multi-agent-refs/wiki-capture.md +14 -14
  128. package/pipeline/preferences-template.json +9 -1
  129. package/pipeline/rules/figma-pipeline.md +8 -8
  130. package/pipeline/rules/outside-the-pipeline.md +1 -1
  131. package/pipeline/schemas/agent-state.schema.json +50 -50
  132. package/pipeline/schemas/analysis-output.schema.json +3 -3
  133. package/pipeline/schemas/analysis-spec.schema.json +2 -2
  134. package/pipeline/schemas/autopilot-config.schema.json +1 -1
  135. package/pipeline/schemas/code-graph.schema.json +1 -1
  136. package/pipeline/schemas/criteria-manifest.schema.json +1 -1
  137. package/pipeline/schemas/dev-critic-output.schema.json +1 -1
  138. package/pipeline/schemas/diff-risk.schema.json +1 -1
  139. package/pipeline/schemas/figma-project-config.schema.json +1 -1
  140. package/pipeline/schemas/migrations/prefs-2.4.0-to-2.5.0.mjs +2 -2
  141. package/pipeline/schemas/migrations/prefs-2.6.0-to-2.7.0.mjs +31 -0
  142. package/pipeline/schemas/migrations/state-2.1.0-to-2.2.0.mjs +129 -0
  143. package/pipeline/schemas/phases.json +105 -0
  144. package/pipeline/schemas/plan-todos.schema.json +5 -5
  145. package/pipeline/schemas/planning-output.schema.json +1 -1
  146. package/pipeline/schemas/prefs.schema.json +102 -58
  147. package/pipeline/schemas/reviewer-output.schema.json +3 -3
  148. package/pipeline/schemas/route-config.schema.json +74 -0
  149. package/pipeline/schemas/scope-check.schema.json +1 -1
  150. package/pipeline/schemas/secret-patterns.json +124 -0
  151. package/pipeline/schemas/test-gap.schema.json +1 -1
  152. package/pipeline/schemas/token-budget.json +12 -18
  153. package/pipeline/schemas/triage-output.schema.json +6 -6
  154. package/pipeline/scripts/README.md +3 -3
  155. package/pipeline/scripts/_code-graph.mjs +2 -2
  156. package/pipeline/scripts/_run-paths.mjs +2 -2
  157. package/pipeline/scripts/_smoke-root.sh +1 -1
  158. package/pipeline/scripts/aggregate-metrics.mjs +1 -1
  159. package/pipeline/scripts/build-references.mjs +2 -2
  160. package/pipeline/scripts/bulk-read.sh +10 -1
  161. package/pipeline/scripts/capture-flush.sh +8 -8
  162. package/pipeline/scripts/capture-resume.sh +3 -3
  163. package/pipeline/scripts/classify-plan-safety.mjs +1 -1
  164. package/pipeline/scripts/cost-table.json +8 -1
  165. package/pipeline/scripts/diff-explain.mjs +1 -1
  166. package/pipeline/scripts/doctor.mjs +3 -3
  167. package/pipeline/scripts/gc-abandoned.sh +3 -3
  168. package/pipeline/scripts/gc-tmp.sh +1 -1
  169. package/pipeline/scripts/gc-worktrees.sh +1 -1
  170. package/pipeline/scripts/gen-facts.mjs +280 -0
  171. package/pipeline/scripts/gen-mode-dispatch.mjs +32 -37
  172. package/pipeline/scripts/gen-ref-toc.mjs +1 -1
  173. package/pipeline/scripts/graph-report.mjs +1 -1
  174. package/pipeline/scripts/jira-attach.sh +1 -1
  175. package/pipeline/scripts/learn-from-transcripts.mjs +1 -1
  176. package/pipeline/scripts/learning-curve.mjs +2 -2
  177. package/pipeline/scripts/log-metric.sh +17 -4
  178. package/pipeline/scripts/memory-save.sh +1 -1
  179. package/pipeline/scripts/migrate-prefs.mjs +22 -5
  180. package/pipeline/scripts/phase-banner.sh +20 -20
  181. package/pipeline/scripts/phase-tracker.sh +12 -12
  182. package/pipeline/scripts/plan-coverage-gate.mjs +2 -2
  183. package/pipeline/scripts/pre-commit-check.sh +30 -1
  184. package/pipeline/scripts/render-agent-log-cost.sh +1 -1
  185. package/pipeline/scripts/render-work-summary.sh +3 -3
  186. package/pipeline/scripts/review-file-filter.mjs +1 -1
  187. package/pipeline/scripts/run-aggregator.mjs +13 -6
  188. package/pipeline/scripts/run-metrics.mjs +1 -1
  189. package/pipeline/scripts/runs-index.mjs +11 -1
  190. package/pipeline/scripts/scan-skills.sh +26 -0
  191. package/pipeline/scripts/smoke-cross-cli-behavior.sh +6 -6
  192. package/pipeline/scripts/smoke-schema-validation.sh +26 -7
  193. package/pipeline/scripts/token-budget-report.mjs +13 -2
  194. package/pipeline/scripts/triage-memory.mjs +2 -2
  195. package/pipeline/scripts/validate-analysis-doc.mjs +274 -43
  196. package/pipeline/scripts/validate-planning.mjs +1 -1
  197. package/pipeline/scripts/validate-reviewer.mjs +1 -1
  198. package/pipeline/scripts/validate-state.mjs +45 -5
  199. package/pipeline/scripts/validate-triage.mjs +3 -3
  200. package/pipeline/scripts/verify-citations.mjs +1 -1
  201. package/pipeline/scripts/worktree-finalize.sh +5 -5
  202. package/pipeline/scripts/write-state.mjs +32 -0
  203. package/pipeline/skills/.skill-manifest.json +38 -22
  204. package/pipeline/skills/.skills-index.json +49 -5
  205. package/pipeline/skills/shared/README.md +10 -6
  206. package/pipeline/skills/shared/core/apple-archive-compliance/SKILL.md +8 -8
  207. package/pipeline/skills/shared/core/google-play-compliance/SKILL.md +8 -8
  208. package/pipeline/skills/shared/core/multi-agent/SKILL.md +81 -82
  209. package/pipeline/skills/shared/core/multi-agent-autopilot/SKILL.md +3 -3
  210. package/pipeline/skills/shared/core/multi-agent-channels/SKILL.md +14 -14
  211. package/pipeline/skills/shared/core/multi-agent-diff-explain/SKILL.md +5 -5
  212. package/pipeline/skills/shared/core/multi-agent-graph/SKILL.md +1 -1
  213. package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +25 -23
  214. package/pipeline/skills/shared/core/multi-agent-language/SKILL.md +2 -2
  215. package/pipeline/skills/shared/core/multi-agent-local/SKILL.md +2 -2
  216. package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +8 -8
  217. package/pipeline/skills/shared/core/multi-agent-manual-test/SKILL.md +6 -6
  218. package/pipeline/skills/shared/core/multi-agent-model/SKILL.md +71 -0
  219. package/pipeline/skills/shared/core/multi-agent-refactor/SKILL.md +3 -3
  220. package/pipeline/skills/shared/core/multi-agent-resume/SKILL.md +1 -1
  221. package/pipeline/skills/shared/core/multi-agent-resume-local/SKILL.md +7 -7
  222. package/pipeline/skills/shared/core/multi-agent-route-off/SKILL.md +39 -0
  223. package/pipeline/skills/shared/core/multi-agent-route-on/SKILL.md +76 -0
  224. package/pipeline/skills/shared/core/multi-agent-route-status/SKILL.md +59 -0
  225. package/pipeline/skills/shared/core/multi-agent-setup/SKILL.md +1 -1
  226. package/pipeline/skills/shared/core/multi-agent-status/SKILL.md +5 -5
  227. package/pipeline/skills/shared/core/multi-agent-steer/SKILL.md +2 -2
  228. package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +6 -5
  229. package/pipeline/skills/shared/external/NOTICE-swift-ios-skills.md +1 -1
  230. package/pipeline/skills/shared/external/signal-community/SKILL.md +8 -1
  231. package/pipeline/skills/skills-index.md +8 -4
  232. package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +0 -263
  233. package/pipeline/multi-agent-refs/phases/phase-2-planning.md +0 -344
  234. package/pipeline/multi-agent-refs/phases/phase-5-test.md +0 -182
@@ -1,87 +1,15 @@
1
- ### Phase 4: Review (deterministic gates + parallel + triage)
1
+ ### Phase 3: Review (parallel + triage + user test)
2
2
 
3
- > **TLDR** - Three-stage review. Stage 1: deterministic gates (build + lint + test + secret scan) that MUST pass. Stage 2: 3 reviewers in parallel per host, second slot **CLI-aware** (Claude Code Fable + Opus + Sonnet; Copilot CLI GPT-5.4 + Opus + Sonnet). Stage 3: Fable triage (Opus on Copilot CLI) filters raw findings for false positives and out-of-scope items. Only triage-accepted blocking items loop back to Phase 3.
3
+ > **TLDR** - Two review stages over evidence Phase 2 already produced: the deterministic gates (build + lint + test + secret scan) are the Phase 2 exit gate now, and this phase inherits `.build.log` and `.test.log` rather than building again. Stage 1: 3 reviewers in parallel per host, second slot **CLI-aware** (Claude Code Fable + Opus + Sonnet; Copilot CLI GPT-5.4 + Opus + Sonnet). Stage 2: Fable triage (Opus on Copilot CLI) filters raw findings for false positives and out-of-scope items. The user test step closes the phase. Only triage-accepted blocking items loop back to Phase 2.
4
4
 
5
5
  <!-- progress-contract: applied -->
6
6
  Progress emission per `$HOME/.claude/multi-agent-refs/progress-contract.md` - lines for each gate, each reviewer dispatch + finish, triage start, triage verdict, fix dispatch.
7
7
 
8
8
  #### Step 0 - Analysis mode branch
9
9
 
10
- `state.mode === "analysis"` has no diff: the artefact is the document. Replace Steps 1-3 with (1) `validate-analysis-doc.mjs <file> --strict` per platform, (2) the same CLI-aware reviewer set reading the document against one question - **could an implementer build the right thing from this alone?** a finding is anything that would force a guess - and (3) `$HOME/.claude/multi-agent-refs/analysis/resolve.md` over the Section 20 rows still `Acik / Open`; deferred rows report `review_blocking`. Then go to Phase 6.
11
-
12
- Log: `Phase 4: Document review - validator:{pass/fail} | {N} findings | {M} resolved, {K} deferred`
13
-
14
- #### Step 1 - Deterministic Gates (run BEFORE AI review)
15
-
16
- If any gate fails → fix first, don't waste AI tokens reviewing broken code.
17
-
18
- ```bash
19
- # Gate 1: Build (xcodebuild/gradle assemble/tsc/py compile - stack-dependent; Xcode uses the build queue lock, see Phase 3) - tee output to a log
20
- <build-command> 2>&1 | tee "$WORKTREE/.build.log"
21
- # Gate 2: Lint (swiftlint/ktlint/ruff/eslint - stack-dependent)
22
- # Gate 3: Tests pass (xcodebuild/gradle/pytest/the resolved node command) - tee output to a log
23
- <test-command> 2>&1 | tee "$WORKTREE/.test.log"
24
- # Gate 4: Secrets - run the scanner against the staged diff
25
- bash $HOME/.claude/scripts/pre-commit-check.sh
26
- ```
27
-
28
- **Default-FAIL evidence gate (required before recording any pass):** a green exit code is not enough - the captured log must actually show success. Before marking build/test passed, run the evidence gate against the tee'd log; it fails CLOSED when the log is missing, empty, or shows failure markers:
29
-
30
- ```bash
31
- node $HOME/.claude/scripts/evidence-gate.mjs --claim build --status passed --evidence "$WORKTREE/.build.log" || BUILD_PASS=false
32
- node $HOME/.claude/scripts/evidence-gate.mjs --claim test --status passed --evidence "$WORKTREE/.test.log" || TEST_PASS=false
33
- ```
34
-
35
- This prevents a false "it built" claim with no log behind it. On exit 1, treat the gate as failed (do NOT proceed to AI review) and surface the gate's `reason`.
36
-
37
- **Inherited failures (when `state.baseline.tests` exists).** Phase 0 Step 7.6 recorded whether the suite was already red, so Gate 3 blocks on what this work broke, not what it walked into:
38
-
39
- | baseline status | Gate 3 |
40
- |---|---|
41
- | `green` | unchanged; every failure is this run's |
42
- | `red` + `failing[]` | subtract those ids. Nothing left -> pass, logged `test:pass (inherited {N})`. A NEW failure still blocks. |
43
- | `red`, empty `failing[]` | do NOT pass and do NOT silently block: report `test:inherited-red (not attributable, <logPath>)` and ask. Inventing a set here masks regressions. |
44
- | `unknown` / absent | unchanged from today |
45
-
46
- The subtraction never widens: match on identifier only, and when identifiers cannot be compared fall to the `not attributable` row.
47
-
48
- **Gate results:**
49
-
50
- - All pass (including the evidence gate) -> proceed to AI review
51
- - Any fail -> fix immediately, re-run gates (no AI review until clean)
52
- Log: "Phase 4: Gates - build:{pass/fail} lint:{pass/fail} test:{pass/fail/inherited-red} secrets:{clean/found} evidence:{ok/unverified}"
53
-
54
- ##### Gate 5 - Fortify SSC findings (runs when `state.contextLinks[]` contains a `fortify` entry, or when `prefs.global.fortify.alwaysCheck === true`)
55
-
56
- If the task description referenced a Fortify version, or named a bare issue instance id, `~/.claude/lib/fetch-fortify.sh` already populated `state.fortifyFinding` in Phase 0 (`alwaysCheck` needs `prefs.global.fortify.versionIds` to know what to scan). Phase 4 reuses that payload and applies the deterministic gate:
57
-
58
- ```bash
59
- gate=$(jq -r '.fortifyFinding.gateOutcome // empty' "$STATE_FILE")
60
- if [ -z "$gate" ]; then
61
- # No fortify entry in contextLinks; gate is N/A.
62
- echo "→ fortify gate: n/a (no Fortify URL referenced)"
63
- elif [ "$(jq -r '.fortifyFinding.gateOutcome.blocking' "$STATE_FILE")" = "true" ]; then
64
- reason=$(jq -r '.fortifyFinding.gateOutcome.reason' "$STATE_FILE")
65
- critical=$(jq -r '.fortifyFinding.severityCounts.Critical' "$STATE_FILE")
66
- echo "→ fortify gate: BLOCKED ($reason, critical=$critical)"
67
- # Treat as a deterministic gate failure - fix the critical findings before AI review.
68
- exit 1
69
- else
70
- high=$(jq -r '.fortifyFinding.severityCounts.High // 0' "$STATE_FILE")
71
- echo "→ fortify gate: pass (high=$high warnings carry into the channel summary)"
72
- fi
73
- ```
74
-
75
- Gate semantics:
76
-
77
- | `gateOutcome.reason` | Phase 4 action |
78
- |---|---|
79
- | `critical-findings` | **Blocking** - Phase 3 re-dispatches with the finding list; pipeline does not advance until Critical count is 0. |
80
- | `high-findings-warning` | Pass-with-warning - High findings appear in the channels summary (`## Security Scan (Fortify)` section in PR + Jira); they don't block the merge. |
81
- | `clean` | Pass silently - section omitted from the channels summary. |
82
-
83
- When `state.fortifyFinding.status === "skipped"` (token missing, VPN unreachable, or host mismatch) the gate logs `→ fortify gate: skipped (<reason>)` and never blocks - the user already saw the structured Save Flow signal at Phase 0 and chose to proceed.
10
+ `state.mode === "analysis"` has no diff: the artefact is the document. Replace Steps 1-3 with (1) `validate-analysis-doc.mjs <file> --strict` per platform, (2) the same CLI-aware reviewer set reading the document against one question - **could an implementer build the right thing from this alone?** a finding is anything that would force a guess - and (3) `$HOME/.claude/multi-agent-refs/analysis/resolve.md` over the Section 20 rows still `Acik / Open`; deferred rows report `review_blocking`. Then go to Phase 4.
84
11
 
12
+ Log: `Phase 3: Document review - validator:{pass/fail} | {N} findings | {M} resolved, {K} deferred`
85
13
  #### Step 1.45 - Test plan coverage cross-check (Phase 1 modes with a test plan)
86
14
 
87
15
  Every `state.dev.testPlan[]` row must map to a real test in the diff, matched on the planned name. Missing -> `important` ("planned test <name> (BR-<id>) has no implementation"); present but asserting something else -> `blocking`, plan and code disagree about the rule. This is what makes "analysis quality is output quality" measurable. Skipped when `docStatus === "not-applicable"` or the mode has no Phase 1.
@@ -115,15 +43,15 @@ If changes include UI files (iOS: `*View.swift`, `*Screen.swift`, `*Cell.swift`;
115
43
  - Privacy: API usage without purpose string / permission rationale → **blocking**
116
44
  - Deprecated API usage flagged by latest SDK → **important**
117
45
 
118
- Device-level audit runs in Phase 5 (`audit-guide.md`).
46
+ Device-level audit runs in Phase 3 (`audit-guide.md`).
119
47
 
120
48
  #### Step 1.6 - Repo Map Injection (advisory, opt-in)
121
49
 
122
50
  **Gated by `prefs.global.repoMap.enabled`** (default: `false`). Same pattern as Phase 1 Step 2.5 - runs `$HOME/.claude/scripts/repo-map.mjs` and injects the result as `${REPO_MAP}` into each reviewer's prompt context. Reviewers treat it the same way the priority files block is treated: advisory hint, not gospel.
123
51
 
124
- Phase 4 reuses the cached map from Phase 1 when both phases run in the same task (avoid recomputing the same scan twice). The orchestrator caches under `state.repoMap.<sha>` keyed by `git rev-parse HEAD`; cache miss → regenerate. When `enabled=false`, both phases skip the script entirely and no `${REPO_MAP}` placeholder appears in prompts (no empty-string artefact).
52
+ Phase 3 reuses the cached map from Phase 1 when both phases run in the same task (avoid recomputing the same scan twice). The orchestrator caches under `state.repoMap.<sha>` keyed by `git rev-parse HEAD`; cache miss → regenerate. When `enabled=false`, both phases skip the script entirely and no `${REPO_MAP}` placeholder appears in prompts (no empty-string artefact).
125
53
 
126
- Cost ledger: `phase-4.repo_map_emitted bytes=N budget=B cache_hit=true|false` - Phase 7 cost rollup distinguishes a cache hit (free) from a regeneration (~150-300ms wall, 0 LLM cost either way).
54
+ Cost ledger: `phase-4.repo_map_emitted bytes=N budget=B cache_hit=true|false` - Phase 5 cost rollup distinguishes a cache hit (free) from a regeneration (~150-300ms wall, 0 LLM cost either way).
127
55
 
128
56
  #### Step 1.8 - Platform parity cross-check (advisory, read-only)
129
57
 
@@ -152,7 +80,7 @@ echo "$RISK_FULL" | node $HOME/.claude/scripts/validate-diff-risk.mjs - >/dev/nu
152
80
  RISK_JSON=$([ -n "$RISK_FULL" ] && jq -c '.files |= (sort_by(-.score) | .[:5])' <<< "$RISK_FULL" || echo "")
153
81
  ```
154
82
 
155
- Persist the totals as `state.diffRisk` (Phase 6 `risk` section, Phase 7, `run-metrics.mjs` read them):
83
+ Persist the totals as `state.diffRisk` (Phase 4 `risk` section, Phase 5, `run-metrics.mjs` read them):
156
84
 
157
85
  ```bash
158
86
  [ -n "$RISK_FULL" ] && jq -c '{diffRisk: (.totals + {signals: ([.files[].signals[]?.name] | unique)})}' <<< "$RISK_FULL" | node $HOME/.claude/scripts/write-state.mjs "$STATE_FILE"
@@ -252,14 +180,14 @@ Output conforms to `$HOME/.claude/schemas/criteria-manifest.schema.json`. The fo
252
180
  | Absent input | Substitute |
253
181
  |---|---|
254
182
  | `detectedStack` (Phase 1 Step 2) | the language census in `criteria-manifest.json` → `languages`, from the diff's file extensions |
255
- | Phase 1 analysis summary + Phase 2 plan (triage scope, cache prefix) | the task description plus the inline task list Phase 3 generated for itself |
183
+ | Phase 1 analysis summary + Phase 1 plan (triage scope, cache prefix) | the task description plus the inline task list Phase 2 generated for itself |
256
184
  | `state.evidence.figma[]` (this step), Step 2.8 visual conformance | nothing. Record each as `not-applicable (no Phase 1 evidence)` in `consensus.visualConformance`; never silently omit |
257
185
 
258
186
  When `state.evidence.figma[]` is non-empty, the reviewer subagents MUST receive the captured screenshot URLs / paths and the canonical-component name (from `state.evidence.figma[i].screenshotUrl` and `state.evidence.figma[i].codeConnectSnippets[0].componentName` when present) so they can compare visual fidelity. Pass them inline in each reviewer prompt under a `## Figma evidence` block, one row per frame.
259
187
 
260
188
  Visual-fidelity mismatches against the captured screenshot are BLOCKING findings, not nits:
261
189
 
262
- - Canonical component usage: the Code Connect-mapped component is used verbatim - a sound-alike substitute, a forked copy, or ad-hoc inline UI where a mapping exists is blocking (Phase 3 "Design fidelity contract")
190
+ - Canonical component usage: the Code Connect-mapped component is used verbatim - a sound-alike substitute, a forked copy, or ad-hoc inline UI where a mapping exists is blocking (Phase 2 "Design fidelity contract")
263
191
  - Inter-component spacing: gaps, paddings, and alignment BETWEEN components match the design's measured values mapped to spacing tokens - invented numeric values are blocking
264
192
  - Everything else: the 30-group catalog in `$HOME/.claude/multi-agent-refs/features/design-conformance.md`. Two of its rules carry into review: a height or inset is **measured**, never read from the token under test; an unrun check is a fail-to-verify, not a pass.
265
193
 
@@ -279,11 +207,11 @@ git -C "$WORKTREE" diff --name-only "$BASE_BRANCH"...HEAD \
279
207
 
280
208
  #### Step 1.9 - Context economy (cache prefix + diff cap)
281
209
 
282
- Phase 4 sends the same diff to every reviewer and then to triage, so the diff is the dominant token cost. Two measures keep it bounded:
210
+ Phase 3 sends the same diff to every reviewer and then to triage, so the diff is the dominant token cost. Two measures keep it bounded:
283
211
 
284
- **Shared cache prefix.** Build the reviewer and triage prompts so the large invariant context - the full diff, the `${CRITERIA}` block from Step 1.78, the `${REVIEW_FILES}` list from Step 1.85, the Phase 1 analysis summary, the Phase 2 plan - is a byte-identical leading block across all dispatches in this iteration. Only the per-reviewer focus + skill line varies, and it goes AFTER the shared block. `${CRITERIA}` goes in the prefix, identical for every reviewer: subsetting it per reviewer would invalidate the prefix for the whole panel and re-bill the largest block in the phase. Per-reviewer emphasis stays a one-line pointer in the suffix. When the host supports prompt caching, the 2nd/3rd reviewer and the triage call then read that prefix at the discounted cache-read rate instead of re-billing it as fresh input. Forward the host-reported cache-read count as `tokens_cached` per the Token telemetry contract so the saving lands in the cost ledger. The `<scope-self-check>` block and, from iteration 2, the `<previous-round-findings>` block (Step 2.1) close the shared block, after the plan and before the per-reviewer suffix.
212
+ **Shared cache prefix.** Build the reviewer and triage prompts so the large invariant context - the full diff, the `${CRITERIA}` block from Step 1.78, the `${REVIEW_FILES}` list from Step 1.85, the Phase 1 analysis summary, the Phase 1 plan - is a byte-identical leading block across all dispatches in this iteration. Only the per-reviewer focus + skill line varies, and it goes AFTER the shared block. `${CRITERIA}` goes in the prefix, identical for every reviewer: subsetting it per reviewer would invalidate the prefix for the whole panel and re-bill the largest block in the phase. Per-reviewer emphasis stays a one-line pointer in the suffix. When the host supports prompt caching, the 2nd/3rd reviewer and the triage call then read that prefix at the discounted cache-read rate instead of re-billing it as fresh input. Forward the host-reported cache-read count as `tokens_cached` per the Token telemetry contract so the saving lands in the cost ledger. The `<scope-self-check>` block and, from iteration 2, the `<previous-round-findings>` block (Step 2.1) close the shared block, after the plan and before the per-reviewer suffix.
285
213
 
286
- **Single-repo diff cap.** Applied to the Step 1.85 `reviewed[]` set only. If it exceeds the Phase 4 token allowance (`token-budget.json`), truncate the largest files and append a footer `[truncated - full diff in file://$WORKTREE/.review-diff.txt]`, writing the full diff to that path. Reviewers and triage receive the same capped view + the marker so they can flag "review the full diff manually." Log `review.diff_truncated bytes_dropped=<N>`. (Multi-repo already caps the combined diff at 80% of budget; this is the single-repo equivalent.)
214
+ **Single-repo diff cap.** Applied to the Step 1.85 `reviewed[]` set only. If it exceeds the Phase 3 token allowance (`token-budget.json`), truncate the largest files and append a footer `[truncated - full diff in file://$WORKTREE/.review-diff.txt]`, writing the full diff to that path. Reviewers and triage receive the same capped view + the marker so they can flag "review the full diff manually." Log `review.diff_truncated bytes_dropped=<N>`. (Multi-repo already caps the combined diff at 80% of budget; this is the single-repo equivalent.)
287
215
 
288
216
  #### Step 2 - Parallel AI Review (CLI-aware reviewer set)
289
217
 
@@ -311,7 +239,7 @@ statement: the managed block in `~/.codex/AGENTS.md`, always loaded on that host
311
239
  2. **Three concurrent children is the ceiling** - 4 slots including the orchestrator,
312
240
  so the reviewer count on Codex is capability-derived, not a preference.
313
241
 
314
- Sub-agent delegation itself is authorized by that same managed block; without it Phase 4
242
+ Sub-agent delegation itself is authorized by that same managed block; without it Phase 3
315
243
  degrades to a single in-thread review.
316
244
 
317
245
  **Single-vendor caveat.** Claude Code and Codex both run a one-vendor panel, so their
@@ -336,7 +264,7 @@ Each reviewer inherits the `code-reviewer` agent's focus areas (Security, Archit
336
264
 
337
265
  ##### 2.1 Previous-round findings (iteration >= 2) and 2.2 scope self-check (every iteration)
338
266
 
339
- From iteration 2, render the previous round's accepted blocking/important findings (`.pipeline/triage-round-$((ITERATION-1)).json`, max 40) into a `<previous-round-findings>` block at the end of the shared prefix. Every iteration also renders `.pipeline/scope-check.json` (Phase 3 Step 3.7) plus `scope-check-gate.mjs --advisory` output as `<scope-self-check>`: file reasons, unjustified files, and `notDone[]` (never re-raised as findings); a missing record logs `review.scope_check=missing`. Block text and recipes: `$HOME/.claude/multi-agent-refs/features/review-delta.md`.
267
+ From iteration 2, render the previous round's accepted blocking/important findings (`.pipeline/triage-round-$((ITERATION-1)).json`, max 40) into a `<previous-round-findings>` block at the end of the shared prefix. Every iteration also renders `.pipeline/scope-check.json` (Phase 2 Step 3.7) plus `scope-check-gate.mjs --advisory` output as `<scope-self-check>`: file reasons, unjustified files, and `notDone[]` (never re-raised as findings); a missing record logs `review.scope_check=missing`. Block text and recipes: `$HOME/.claude/multi-agent-refs/features/review-delta.md`.
340
268
 
341
269
  #### Step 2.8 - Visual conformance gate (component / screen work only)
342
270
 
@@ -358,7 +286,7 @@ blocking finding.
358
286
  Why this is a gate and not advice: `design-conformance.md`, "Why this runs as a gate".
359
287
 
360
288
  Skip only when the diff has no UI change. Record the outcome in
361
- `consensus.visualConformance` so Phase 7 reports whether it ran.
289
+ `consensus.visualConformance` so Phase 5 reports whether it ran.
362
290
 
363
291
  **iOS/Swift - interaction & convention checks (conditional).** Step 1.78 resolves these. Where no registry covers the change, reviewers fall back to the analysis doc (Section 14 Code Connect mapping) and, when `ai-ios-toolkit` is enabled, that plugin's navigation / overlay / bottom-sheet + accessibility conventions.
364
292
 
@@ -368,7 +296,7 @@ Skip only when the diff has no UI change. Record the outcome in
368
296
 
369
297
  #### Output contract - reviewer step
370
298
 
371
- Step 2 produces N reviewer-output objects (one per dispatched reviewer), each conforming to `$HOME/.claude/schemas/reviewer-output.schema.json`. They are persisted to `state.reviewIterations[<iteration>].reviewers[]` and consumed by Step 3 (Fable triage) - never by Phase 6 directly. The triage step (below) is the producer of the only review artifact Phase 6 reads, conforming to `$HOME/.claude/schemas/triage-output.schema.json`.
299
+ Step 2 produces N reviewer-output objects (one per dispatched reviewer), each conforming to `$HOME/.claude/schemas/reviewer-output.schema.json`. They are persisted to `state.reviewIterations[<iteration>].reviewers[]` and consumed by Step 3 (Fable triage) - never by Phase 4 directly. The triage step (below) is the producer of the only review artifact Phase 4 reads, conforming to `$HOME/.claude/schemas/triage-output.schema.json`.
372
300
 
373
301
  **Subagent return format** - each reviewer returns JSON conforming to `$HOME/.claude/schemas/reviewer-output.schema.json`:
374
302
 
@@ -417,7 +345,7 @@ Exit 0 = valid. Exit 2 = contradiction (approved=true with blocking findings) -
417
345
  - Max one round. Results replace the original outputs.
418
346
  4. Proceed to Step 3 triage with the round-2 outputs.
419
347
 
420
- **Parity contract:** every host (three reviewers each: Claude Code, Copilot CLI, Codex CLI) runs the round identically. Telemetry emits `review_round_count={1|2}` per reviewer for Phase 7 rollup.
348
+ **Parity contract:** every host (three reviewers each: Claude Code, Copilot CLI, Codex CLI) runs the round identically. Telemetry emits `review_round_count={1|2}` per reviewer for Phase 5 rollup.
421
349
 
422
350
  **Cost ceiling:** rebuttal round consumes ~1× the original Step 2 token budget. Smoke + budget tests treat this as opt-in so the default cost stays the same.
423
351
 
@@ -427,9 +355,9 @@ Exit 0 = valid. Exit 2 = contradiction (approved=true with blocking findings) -
427
355
 
428
356
  **CRITICAL**: Reviewer findings are **raw signals**, not commands. Never auto-loop on every "blocking" tag - reviewers hallucinate, misread scope, or repeat each other. Run Fable triage (Opus on Copilot CLI) to evaluate merged findings against task scope.
429
357
 
430
- Optional: when `ai-analyst-toolkit` is enabled and a finding blames a third-party library rather than this diff, ask `evidence-github` whether it is already open upstream. A confirmed one is `deferred` with its `GitHub:<owner>/<repo>#<n>` citation, not `accepted` and handed to Phase 3 to fix code that is not ours.
358
+ Optional: when `ai-analyst-toolkit` is enabled and a finding blames a third-party library rather than this diff, ask `evidence-github` whether it is already open upstream. A confirmed one is `deferred` with its `GitHub:<owner>/<repo>#<n>` citation, not `accepted` and handed to Phase 2 to fix code that is not ours.
431
359
 
432
- Opt-in empirical layer: when `prefs.global.verifyByTest.enabled` is `true`, accepted blocking findings additionally go through Step 3.7 (verify-by-test), which tries to reproduce each one with a minimal failing test before the Phase 3 rework loop fires. Full wiring: `$HOME/.claude/multi-agent-refs/features/verify-by-test.md`.
360
+ Opt-in empirical layer: when `prefs.global.verifyByTest.enabled` is `true`, accepted blocking findings additionally go through Step 3.7 (verify-by-test), which tries to reproduce each one with a minimal failing test before the Phase 2 rework loop fires. Full wiring: `$HOME/.claude/multi-agent-refs/features/verify-by-test.md`.
433
361
 
434
362
  ##### 3.0 Anonymize the reviewer findings, then merge the deterministic ones
435
363
 
@@ -441,7 +369,7 @@ ANON=$(jq -n --argjson r "$REVIEWERS_JSON" --arg t "$TASK_ID" --argjson i "$ITER
441
369
  | node "$HOME/.claude/scripts/anonymize-findings.mjs" --map "/tmp/review-$TASK_ID-$ITERATION-map.json")
442
370
  ```
443
371
 
444
- `$REVIEWERS_JSON` is `state.reviewIterations[i].reviewers`. Findings come back with `foundBy: "Source A|B|C"` and every identity key removed. Persist the map to `state.reviewIterations[i].anonymizationMap` for Phase 7 per-reviewer telemetry, and **never put the map in a prompt**.
372
+ `$REVIEWERS_JSON` is `state.reviewIterations[i].reviewers`. Findings come back with `foundBy: "Source A|B|C"` and every identity key removed. Persist the map to `state.reviewIterations[i].anonymizationMap` for Phase 5 per-reviewer telemetry, and **never put the map in a prompt**.
445
373
 
446
374
  Then append the Step 1.76 test-integrity findings, so they are adjudicated rather than never seen:
447
375
 
@@ -454,14 +382,14 @@ Deterministic findings keep `tag: test_integrity` and carry no `foundBy`: a revi
454
382
 
455
383
  ##### 3.1 Short-circuit: no findings
456
384
 
457
- If **merged** findings `length === 0`, **skip triage**: write empty result `{"accepted": [], "deferred": [], "rejected": [], "approved": true}`, log, proceed to Phase 5. Note this is the merged count from 3.0: a run with zero reviewer findings but a non-empty test-integrity set must NOT short-circuit.
385
+ If **merged** findings `length === 0`, **skip triage**: write empty result `{"accepted": [], "deferred": [], "rejected": [], "approved": true}`, log, proceed to Phase 3. Note this is the merged count from 3.0: a run with zero reviewer findings but a non-empty test-integrity set must NOT short-circuit.
458
386
 
459
387
  ##### 3.2 Launch triage agent
460
388
 
461
389
  Launch **1 Agent** (subagent_type: `general-purpose`, model: `fable` on Claude Code / `opus` on Copilot CLI) with:
462
390
 
463
391
  - The anonymized merged findings from 3.0 (`Source A/B/C` labels; no model name anywhere in the prompt)
464
- - Task scope (Phase 1 analysis summary + Phase 2 plan)
392
+ - Task scope (Phase 1 analysis summary + Phase 1 plan)
465
393
  - Full diff being reviewed
466
394
  - **Prior-art context (advisory)** - per raw finding, `triage-memory.mjs query --top <prefs.global.priorArtEnrichment.topN>` (default 3). Pass `--top`: without it the script falls back to `memoryRecall.maxResults`, a different concern, and `topN` silently does nothing. Off when `priorArtEnrichment.enabled = false`.
467
395
 
@@ -527,7 +455,7 @@ a reason. Preserve `ruleId` and `criteriaSource` on every finding you keep.
527
455
 
528
456
  #### Output contract - triage step
529
457
 
530
- Step 3 produces a single triage-output object conforming to `$HOME/.claude/schemas/triage-output.schema.json` and persists it to `state.reviewIterations[<iteration>].triage`. This is the **only** Phase 4 artifact Phase 6 reads. Phase 6 commits MUST cite only `accepted` findings that were resolved; `deferred` items get linked in the PR description as follow-up work; `rejected` items never appear in any user-facing output.
458
+ Step 3 produces a single triage-output object conforming to `$HOME/.claude/schemas/triage-output.schema.json` and persists it to `state.reviewIterations[<iteration>].triage`. This is the **only** Phase 3 artifact Phase 4 reads. Phase 4 commits MUST cite only `accepted` findings that were resolved; `deferred` items get linked in the PR description as follow-up work; `rejected` items never appear in any user-facing output.
531
459
 
532
460
  Return ONLY valid JSON conforming to $HOME/.claude/schemas/triage-output.schema.json:
533
461
  {
@@ -555,7 +483,7 @@ node $HOME/.claude/scripts/validate-triage.mjs "$TRIAGE_FILE" \
555
483
 
556
484
  Progress line: ` → checking validator validate-triage`
557
485
 
558
- One file per round; `triage-output.json` is the latest copy every downstream reader expects (Phase 7, finalize, work-summary, diff-explain). Step 3.7 rewrites `$TRIAGE_FILE`; repeat the `cp` after it.
486
+ One file per round; `triage-output.json` is the latest copy every downstream reader expects (Phase 5, finalize, work-summary, diff-explain). Step 3.7 rewrites `$TRIAGE_FILE`; repeat the `cp` after it.
559
487
 
560
488
  | Exit | Meaning | Action |
561
489
  | ----- | ---------------------------- | ------------------------------------------------- |
@@ -564,7 +492,7 @@ One file per round; `triage-output.json` is the latest copy every downstream rea
564
492
  | **2** | Over-rejection guard tripped | Pause for human confirm (autopilot: log + accept) |
565
493
  | **3** | Contradiction auto-corrected | Proceed with `result.corrected` |
566
494
 
567
- Capture stdout into `state.reviewIterations[-1].validatorResult` for Phase 7 audit.
495
+ Capture stdout into `state.reviewIterations[-1].validatorResult` for Phase 5 audit.
568
496
 
569
497
  ##### 3.3 Edge case handling
570
498
 
@@ -578,7 +506,7 @@ Failure fallback (timeout >120s, or agent crash before any JSON is produced): re
578
506
 
579
507
  ##### 3.4 Telemetry emission (required)
580
508
 
581
- Emit metrics per review pass for Phase 7 cost rollup:
509
+ Emit metrics per review pass for Phase 5 cost rollup:
582
510
 
583
511
  One `review.reviewer_call` per dispatched reviewer, one `review.triage_call`, one `review.completed` to close the pass:
584
512
 
@@ -621,13 +549,13 @@ After the triage verdict is computed, populate `triage.consensus`:
621
549
  - `unverified` -> all reviewers approved BUT the diff touches a judgment-heavy surface (security, auth, concurrency, money, data migration). Agreement here may be correlated; do NOT treat it as a confirmed pass. Surface it.
622
550
  3. `disagreements[]` is populated for `split` and is also used to carry `unverified` notes (e.g. "both approved a keychain change - agreement unverified, confirm manually").
623
551
 
624
- **Surfacing (Step 4 + Phase 7):** When `verdict` is `split` or `unverified`, the disagreements are shown to the user at the Step 4 checkpoint (interactive modes) and always written to the Phase 7 report. Autopilot does not block on `unverified` (it logs `review.consensus=unverified` and proceeds), matching the maturity-check model - but the report records it so a human review can catch it. This is additive: omitting `consensus` is valid and means "not computed."
552
+ **Surfacing (Step 4 + Phase 5):** When `verdict` is `split` or `unverified`, the disagreements are shown to the user at the Step 4 checkpoint (interactive modes) and always written to the Phase 5 report. Autopilot does not block on `unverified` (it logs `review.consensus=unverified` and proceeds), matching the maturity-check model - but the report records it so a human review can catch it. This is additive: omitting `consensus` is valid and means "not computed."
625
553
 
626
554
  ##### 3.7 Verify-by-test (opt-in, empirical validation of blocking findings)
627
555
 
628
556
  A triage verdict is judgment; a failing repro test is proof. Runs only when `prefs.global.verifyByTest.enabled` is `true` AND `accepted` contains a `blocking` finding; otherwise skip silently. **Full contract (verdict table, cleanup invariant, prompts): `$HOME/.claude/multi-agent-refs/features/verify-by-test.md` - read it before executing this step.**
629
557
 
630
- Compressed flow: dispatch ONE verifier agent (model `verifyByTest.model`, default `sonnet`) for up to `maxFindings` (default 3) accepted blocking findings. Per finding it writes ONE minimal repro test and runs ONLY that test (Phase 3 single-test invocation, build lock, log tee'd to `$WORKTREE/.pipeline/verify-<i>.test.log`). Outcomes: test FAILS as predicted -> `confirmed`, finding stays blocking and the test is KEPT in `redTests[]` as the Phase 3 rework RED test; test PASSES on every one of `verifyByTest.repeatCount` runs (default 3) -> `not-reproduced` ONLY if `evidence-gate.mjs --claim test --status passed` exits 0 on each log; a run that disagrees with the others -> `inconclusive` with `flaky: passed k/N`, finding moves to `deferred[]`, test deleted; compile error / timeout / not unit-testable -> `inconclusive`, judgment verdict stands. Stamp findings with `verification` (schema v3.2.0), persist `state.reviewIterations[-1].verifyByTest = {attempted, confirmed, downgraded, inconclusive, redTests[]}`, recompute `approved`, re-run `validate-triage.mjs` under the 3.2.1 gate. Whole step bounded by `stepTimeoutSec` (default 600); on breach or crash remaining findings keep judgment verdicts - never blocks. Telemetry per 3.4: `review.verify_by_test attempted= confirmed= downgraded= inconclusive= duration_ms=`.
558
+ Compressed flow: dispatch ONE verifier agent (model `verifyByTest.model`, default `sonnet`) for up to `maxFindings` (default 3) accepted blocking findings. Per finding it writes ONE minimal repro test and runs ONLY that test (Phase 2 single-test invocation, build lock, log tee'd to `$WORKTREE/.pipeline/verify-<i>.test.log`). Outcomes: test FAILS as predicted -> `confirmed`, finding stays blocking and the test is KEPT in `redTests[]` as the Phase 2 rework RED test; test PASSES on every one of `verifyByTest.repeatCount` runs (default 3) -> `not-reproduced` ONLY if `evidence-gate.mjs --claim test --status passed` exits 0 on each log; a run that disagrees with the others -> `inconclusive` with `flaky: passed k/N`, finding moves to `deferred[]`, test deleted; compile error / timeout / not unit-testable -> `inconclusive`, judgment verdict stands. Stamp findings with `verification` (schema v3.2.0), persist `state.reviewIterations[-1].verifyByTest = {attempted, confirmed, downgraded, inconclusive, redTests[]}`, recompute `approved`, re-run `validate-triage.mjs` under the 3.2.1 gate. Whole step bounded by `stepTimeoutSec` (default 600); on breach or crash remaining findings keep judgment verdicts - never blocks. Telemetry per 3.4: `review.verify_by_test attempted= confirmed= downgraded= inconclusive= duration_ms=`.
631
559
 
632
560
  ##### 3.8 Cross-round delta + circuit-breaker trigger 2 (iteration >= 2)
633
561
 
@@ -646,7 +574,7 @@ Merge the JSON into `state.reviewIterations[-1].delta` via `write-state.mjs`; em
646
574
  | **3** | Autopilot with `prefs.global.autopilotCircuitBreaker.enabled` (default true): write `state.circuitBreaker = {tripped: true, trigger: 2, detail, checkpoint: {phase: 4, step: "3.8", iteration: N}, trippedAt, counters}`, then the `operations.md` halt protocol with `haltReason="4:circuit-breaker:identical-finding"`. Interactive: show `delta.stillPresent`, ask `Continue rework` / `Escalate to me` / `Accept as deferred`. |
647
575
  | **1** | Log `review.delta_skipped reason=invalid`, continue; the delta never blocks on its own failure. |
648
576
 
649
- `plateau` is logged, not acted on; trigger 3 (the rework cap) is recorded by the Phase 3 re-entry.
577
+ `plateau` is logged, not acted on; trigger 3 (the rework cap) is recorded by the Phase 2 re-entry.
650
578
 
651
579
  #### Step 4 - Consensus + Action (triage-driven)
652
580
 
@@ -654,15 +582,15 @@ If `triage.consensus.verdict` is `split` or `unverified`, surface `consensus.dis
654
582
 
655
583
  Act **only on triage.accepted**:
656
584
 
657
- - **accepted.blocking** → back to Phase 3 (max 3 iterations, with reflection prompt citing only accepted items). The reflection prompt names findings by `fingerprint` and quotes `delta.stillPresent` first, marked `STILL PRESENT after round N-1's fix`. When Step 3.7 ran and `state.reviewIterations[-1].verifyByTest.redTests[]` is non-empty, the reflection prompt cites each red test: "a failing repro test already exists at <testRef>; make it green; do not delete or weaken it."
585
+ - **accepted.blocking** → back to Phase 2 (max 3 iterations, with reflection prompt citing only accepted items). The reflection prompt names findings by `fingerprint` and quotes `delta.stillPresent` first, marked `STILL PRESENT after round N-1's fix`. When Step 3.7 ran and `state.reviewIterations[-1].verifyByTest.redTests[]` is non-empty, the reflection prompt cites each red test: "a failing repro test already exists at <testRef>; make it green; do not delete or weaken it."
658
586
  - **accepted.important** → fix and re-review
659
587
  - **accepted.suggestion** → apply if reasonable
660
- - **deferred** → append to Phase 7 report as "follow-up items" (do not block)
588
+ - **deferred** → append to Phase 5 report as "follow-up items" (do not block)
661
589
  - **rejected** → log reasons for audit; do not touch
662
590
 
663
591
  ##### Lesson memory loop (required, end of each fix/rework round)
664
592
 
665
- At the end of every fix/rework round (each Phase 3 re-entry that resolved accepted findings, including the final one), append ONE one-line root-cause lesson per resolved blocking/important finding to the existing learnings ledger (`learnings-ledger.mjs` - the store Phase 1 and triage already replay; never invent a parallel store):
593
+ At the end of every fix/rework round (each Phase 2 re-entry that resolved accepted findings, including the final one), append ONE one-line root-cause lesson per resolved blocking/important finding to the existing learnings ledger (`learnings-ledger.mjs` - the store Phase 1 and triage already replay; never invent a parallel store):
666
594
 
667
595
  ```bash
668
596
  node $HOME/.claude/scripts/learnings-ledger.mjs add --kind fact \
@@ -675,7 +603,7 @@ Statement shape: the durable rule/root cause, not the symptom - "force-unwrapp
675
603
 
676
604
  Progress line: ` → writing lesson to learnings ledger ({N} entries)`
677
605
 
678
- Log: "Phase 4: Review - raw={N1+N2+N3} accepted={Na} deferred={Nd} rejected={Nr} approved={bool} consensus={verdict}"
606
+ Log: "Phase 3: Review - raw={N1+N2+N3} accepted={Na} deferred={Nd} rejected={Nr} approved={bool} consensus={verdict}"
679
607
 
680
608
  ---
681
609
 
@@ -689,7 +617,197 @@ Log: "Phase 4: Review - raw={N1+N2+N3} accepted={Na} deferred={Nd} rejected={N
689
617
  bash $HOME/.claude/scripts/phase-tracker.sh tokens 4 <input_count> <output_count> [cached_count]
690
618
  ```
691
619
 
692
- The optional 4th `cached_count` is the prompt-cache-read token count when the host reports it (Anthropic `cache_read_input_tokens`); it defaults to 0 and is priced at the cheaper `cacheReadPerMtok` rate in the Phase 7 cost ledger. The tracker accumulates the totals additively, so multiple calls in the same phase compound. The render output then shows live cost on the active phase tile (e.g. `Phase 4 Dev 2m 14s · 12.4k tok`). This satisfies the contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` and the `smoke-tracker-tokens-invocation.sh` enforcement gate. Skipping this call is the #1 cause of "I can't see how much it cost" complaints.
620
+ The optional 4th `cached_count` is the prompt-cache-read token count when the host reports it (Anthropic `cache_read_input_tokens`); it defaults to 0 and is priced at the cheaper `cacheReadPerMtok` rate in the Phase 5 cost ledger. The tracker accumulates the totals additively, so multiple calls in the same phase compound. The render output then shows live cost on the active phase tile (e.g. `Phase 2 Dev 2m 14s · 12.4k tok`). This satisfies the contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` and the `smoke-tracker-tokens-invocation.sh` enforcement gate. Skipping this call is the #1 cause of "I can't see how much it cost" complaints.
693
621
 
694
622
  Contract and rationale: `progress-contract.md` -> Token telemetry forwarding.
695
623
 
624
+ ---
625
+
626
+ ## User test (was Phase 3 until v19.0.0)
627
+
628
+ Optional test gate, now the tail of Review rather than a phase of its own: it
629
+ judges work that already exists, which is what Review does. Needs an interactive
630
+ prompt AND a worktree checkout, so it is skipped where either is missing. If
631
+ issues are found, the run returns to Phase 1 Dev.
632
+
633
+
634
+ > **TLDR** - Optional test gate. Offers to boot the simulator/emulator (UI Bug Hunter) or hand off to the user for manual QA. Needs an interactive prompt AND a worktree checkout, so it is in the phase set of `/multi-agent` alone and dropped by every `autopilot` or `--local` entry. Depth does not affect it: a Short run still reaches Phase 3. If issues found, loops back to Phase 2.
635
+
636
+ <!-- progress-contract: applied -->
637
+ Progress emission per `$HOME/.claude/multi-agent-refs/progress-contract.md` - lines for local-test prompt render, user-answer capture, repo checkout (if selected).
638
+
639
+ #### Step 0 - Test Gap Report (advisory)
640
+
641
+ `state.testPolicy: none` → skip the gap scan (the gap IS the recorded policy) and run only pre-existing test targets; none → recorded no-op. Otherwise, before the local-checkout prompt, run the static test-gap detector. Heuristic, deterministic, no LLM, sub-second. The report ends up in `agent-log.md` under "Test Scenarios" and surfaces public symbols added in this branch that have no paired test.
642
+
643
+ ```bash
644
+ STACK=$(jq -r '.analysis.stack.primary // "unknown"' "$STATE_FILE")
645
+ case "$STACK" in
646
+ ios|swift) SCAN_STACK=ios ;;
647
+ android|kotlin) SCAN_STACK=android ;;
648
+ python) SCAN_STACK=python ;;
649
+ node|typescript|js) SCAN_STACK=node ;;
650
+ *) SCAN_STACK="" ;;
651
+ esac
652
+ if [ -n "$SCAN_STACK" ] && [ "${prefs_testGap_enabled:-true}" = "true" ]; then
653
+ GAP_FLAGS=""
654
+ [ "${prefs_testGap_scanTree:-false}" = "true" ] && GAP_FLAGS="$GAP_FLAGS --scan-tree"
655
+ [ "${prefs_testGap_promoteSeverity:-false}" = "true" ] && GAP_FLAGS="$GAP_FLAGS --severity-promote"
656
+ GAP_JSON=$(node $HOME/.claude/scripts/test-gap-scan.mjs \
657
+ --base "$BASE_BRANCH" --head HEAD --stack "$SCAN_STACK" $GAP_FLAGS 2>/dev/null)
658
+ echo "$GAP_JSON" | node $HOME/.claude/scripts/validate-test-gap.mjs - >/dev/null 2>&1 || GAP_JSON=""
659
+ fi
660
+ ```
661
+
662
+ **What the report contains** (per `$HOME/.claude/schemas/test-gap.schema.json`):
663
+
664
+ | Field | Meaning |
665
+ |---|---|
666
+ | `gaps[].sourcePath` | source file with the unprotected symbol |
667
+ | `gaps[].symbol` | symbol name |
668
+ | `gaps[].kind` | rule id (e.g. `public_func`, `composable_fun`, `named_export`) |
669
+ | `gaps[].severity` | `blocking` / `important` / `suggestion` (see severity table below) |
670
+ | `gaps[].expectedTestPaths` | likely test paths the user should land at, in priority order |
671
+ | `gaps[].hint` | stack-specific testing reminder (e.g. swiftui-qa.md 3-layer) |
672
+
673
+ **Severity defaults**:
674
+
675
+ | Symbol kind | Severity |
676
+ |---|---|
677
+ | `composable_fun`, `view_struct`, `config_struct`, `interface`, `objc_export`, `public_proto` | important |
678
+ | Other public API additions (`public_func`, `open_fun`, `named_export`, `default_export`, ...) | suggestion |
679
+
680
+ **Gating** (opt-in): if `prefs.testGap.blockingThreshold` is set and `gapBySeverity.important + gapBySeverity.blocking` exceeds it, Phase 3 surfaces the report as a Phase 3 rework finding and loops back. **Default off** - gaps render as advisory under "Test Gap Report" only.
681
+
682
+ **Telemetry**:
683
+
684
+ ```bash
685
+ LOG_METRIC_FORWARD_TO_TRACKER=0 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 5 test_gap.scanned \
686
+ stack=$SCAN_STACK \
687
+ sources=$(jq '.totals.sourcesScanned' <<< "$GAP_JSON") \
688
+ gaps=$(jq '.totals.gapCount' <<< "$GAP_JSON")
689
+ ```
690
+
691
+ (No tracker forwarding - the scanner has no token cost.)
692
+
693
+ **Figma reference panel (when `state.evidence.figma[]` is non-empty).** Before the local-checkout prompt, print a single block listing each captured frame so the user has a side-by-side reference during manual test:
694
+
695
+ ```
696
+ Figma evidence (tier=<n>):
697
+ <fileKey>:<nodeId> <canonicalComponentName>
698
+ screenshot: <screenshotUrl or local path>
699
+ <fileKey>:<nodeId> <canonicalComponentName>
700
+ screenshot: <screenshotUrl or local path>
701
+ ```
702
+
703
+ Tier 1 / Tier 2 records print `screenshotUrl` from the captured evidence (Tier 2 URLs expire after 30 days, re-fetch on the spot if needed). Tier 3 records print the local path to the user-attached screenshot. The block is informational; it never blocks the prompt.
704
+
705
+ 1. Ask with a native `AskUserQuestion` picker (never a typed y/N prompt). The options MUST make the local-checkout side effect explicit - testing removes the worktree and checks the branch out into the main repo:
706
+ - `question`: "Check out locally to test now?" (rendered in `outputLanguage`)
707
+ - `header`: "Test" (English, <=12 chars)
708
+ - `options`:
709
+ - `{ label: "Test now", description: "Removes the worktree and checks the branch out into the main repo for Xcode / manual test" }`
710
+ - `{ label: "Skip", description: "Stay in the worktree and go to Phase 4" }`
711
+ - **Skip** → set `state.phases["3"].status = "skipped"` (so Phase 4 can offer the local-checkout prompt) → Phase 4
712
+ - **Test now** → set `state.phases["3"].status = "in_progress"` → continue:
713
+ 2. **Commit changes in worktree BEFORE removing** (WIP commit to preserve work):
714
+ ```
715
+ git -C {worktree-path} add -A
716
+ git -C {worktree-path} commit -m "WIP: {jiraId} - changes for user test"
717
+ ```
718
+ 3. Remove worktree, checkout branch in main repo:
719
+ ```
720
+ git worktree remove .worktrees/{jiraId} --force
721
+ git checkout {branch-name}
722
+ ```
723
+ Branch now has the WIP commit - all changes are preserved.
724
+ 4. Show test instructions:
725
+ ```
726
+ Switched to branch: {branch-name}
727
+ To test: Xcode -> build -> run -> manual test
728
+ "ok" -> proceeds to Phase 4 (WIP commit will be replaced via git reset HEAD~1 + proper commit)
729
+ "fix: ..." -> worktree is recreated, returns to Phase 2
730
+ ```
731
+ 5. Mark the tracker as waiting, then wait for the user response:
732
+ ```bash
733
+ bash $HOME/.claude/scripts/phase-tracker.sh now 3 "awaiting local test (user)"
734
+ bash $HOME/.claude/scripts/phase-tracker.sh render
735
+ ```
736
+ The waiting state persists in `tracker-state.json` across the handoff; `/multi-agent:resume-local` and `/multi-agent:manual-test` CONTINUE this state file and never re-init it (`$HOME/.claude/multi-agent-refs/tracker-contract.md` "Continuation runs").
737
+
738
+ **"ok" is a structured result, not a word.** Before "ok" is accepted, the run writes `$WORKTREE/.pipeline/manual-test.json`: one entry per acceptance criterion, the criteria taken from the analysis doc test plan (Section 15 / 20), the plan tasks, and the user's own words in the reply. Every criterion records what was seen; a criterion that was not tried says so with a reason.
739
+ ```json
740
+ {"criteria":[{"spec":"<quote>","source":"analysis 15.2 | plan task 3 | user","observed":"<what was seen>","verdict":"pass|fail|not-tested","reason":"<required when not-tested>","screenshot":"<path or null>"}],"verdict":"passed|failed"}
741
+ ```
742
+ Then gate it:
743
+ ```bash
744
+ node $HOME/.claude/scripts/evidence-gate.mjs --claim manual --status passed --evidence "$WORKTREE/.pipeline/manual-test.json"
745
+ ```
746
+ Add `--require-screenshot` when `state.visualEvidence.required` is true: a passing criterion then has to name a file that exists, because a `screenshot` key pointing nowhere is not evidence.
747
+
748
+ Exit 1 means the "ok" is not accepted: tell the user which criterion is missing evidence (a `fail` verdict, or `not-tested` without a reason) and wait for the next reply. Exit 0 marks Phase 3 completed with `Result "local test passed (user)"`. The "fix: ..." path below is unchanged.
749
+ 6. If fix needed:
750
+ - Branch already has WIP commit (from step 2) - changes are safe
751
+ - **Heal stale admin state first** (same contract as Phase 0 - step 3's
752
+ `worktree remove` or an interrupted run can leave a stale entry, so a bare
753
+ re-add fails with `already exists`/`already registered`):
754
+ ```bash
755
+ git -C "$PROJECT_ROOT" worktree prune 2>/dev/null || true
756
+ if git -C "$PROJECT_ROOT" worktree list --porcelain | grep -qF "{worktree-path}"; then
757
+ git -C "$PROJECT_ROOT" worktree unlock "{worktree-path}" 2>/dev/null || true
758
+ fi
759
+ ```
760
+ Phase 0's repo residue guard is already in `.git/info/exclude` - no re-add.
761
+ - Recreate worktree from branch: `git -C $PROJECT_ROOT worktree add {worktree-path} {branch}`
762
+ - Re-set git identity: `git -C {worktree-path} config user.name/email` (from state)
763
+ - Go back to Phase 2
764
+ 7. Log: "Phase 3: Review (user test) - {result}"
765
+
766
+ **CRITICAL**: Never remove a worktree with uncommitted changes. Always WIP commit first.
767
+
768
+ #### Automated Device Checks (on-demand)
769
+
770
+ Before or during user testing, run device-level audits via Bash if user requests. See `audit-guide.md` for commands.
771
+
772
+ | Check | When | Command |
773
+ | ------------------- | ----------------- | ----------------------------------------- |
774
+ | UI flow video | `state.visualEvidence.required` AND Phase 2 recorded none | `capture-evidence.sh video start` -> drive the flow -> `video stop` -> `fit`. See below |
775
+ | Accessibility audit | UI changes | `mcp__multi-agent-toolkit__{ios,android}_accessibility_audit` |
776
+ | Biometric test | Auth flow changes | ios: `mcp__multi-agent-toolkit__ios_biometric` (android: manual) |
777
+ | Launch time | Perf-sensitive changes | ios: app-launch instrument · android: `mcp__multi-agent-toolkit__android_launch_time` |
778
+ | Visual test | Any UI changes | `/multi-agent test` (sim-test, both platforms) |
779
+ | Snapshot regression | Component / pixel-stable UI changes | ios: `mcp__multi-agent-toolkit__ios_visual_diff` · android: `mcp__multi-agent-toolkit__android_screenshot` + compare |
780
+ | Store screenshots | `taskType === screenshot` | ios: `ios_status_bar({preset: "clean"})` · android: `android_screenshot` |
781
+
782
+ ##### UI flow video, when Phase 2 produced none
783
+
784
+ Phase 2 Step 3.55 is the primary host and runs in every mode. Phase 3 is the richer one where it exists: the device is up and a person is watching, so the flow is one somebody confirmed. It adds, never replaces.
785
+
786
+ Run only when `visualEvidence.required` and `visualEvidence.video.file` is empty: re-check the device with `$HOME/.claude/scripts/probe-evidence-capability.sh --only device`, then `$HOME/.claude/scripts/capture-evidence.sh video start` -> drive the flow (`run-ui-tests.sh run`, or the Phase 3 scenarios by hand) -> `video stop` -> `fit`. The cap is `visualEvidence.maxVideoSeconds`, read through `capture-evidence.sh limits` so one value serves both. Update `videoTier` and `videoTierReason` with what ran.
787
+
788
+ When the intake answer was `unit` there is no recording here either: overriding it in a phase the user may not be watching makes the question decorative.
789
+
790
+ Results included in Phase 5 report. MCP tools preferred when available - concise structured output, lower token cost.
791
+
792
+ **Snapshot regression flow (optional):** when the task changes a stable component, capture a screenshot before the change (baseline) and after (current), then call `ios_visual_diff({baseline, current, max_diff_pct: 1.0})`. Threshold can be relaxed for animated / non-deterministic regions - keep `max_diff_pct ≤ 1.0` for static layouts.
793
+
794
+ #### Security Audit (store-readiness)
795
+
796
+ When the task touches authentication, keychain, network, or is scheduled for an imminent release, launch the `security-auditor` subagent to run an OWASP Mobile Top 10 pass plus App Store / Play Store compliance checks:
797
+
798
+ ```
799
+ Agent(subagent_type: "security-auditor", prompt: "<diff + context>")
800
+ ```
801
+
802
+ Returns severity-tagged findings (Critical / High / Medium). Critical items block Phase 4 just like Phase 3 blockers; High items are logged and surfaced in Phase 5 report. Skipped by default - opt-in for release branches or on explicit `/multi-agent "<task>" --audit` flag.
803
+
804
+ #### Telemetry - token forwarding
805
+
806
+ When the security-auditor or any other Phase 3 sub-agent runs, forward its token totals so Phase 5's Cost Breakdown captures Phase 3:
807
+
808
+ ```bash
809
+ LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 5 audit.completed \
810
+ model=opus tokens_in=$IN tokens_out=$OUT duration_ms=$DUR
811
+ ```
812
+
813
+ If Phase 3 is purely user-driven (no sub-agent ran), no token forwarding is required and the cost block stays empty for this phase. Best-effort. See `$HOME/.claude/multi-agent-refs/progress-contract.md#token-telemetry-forwarding`.