@mmerterden/multi-agent-pipeline 13.5.1 → 14.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (118) hide show
  1. package/CHANGELOG.md +232 -0
  2. package/README.md +3 -3
  3. package/docs/features.md +1 -1
  4. package/install/_common.mjs +73 -0
  5. package/install/_mcp-register.mjs +70 -31
  6. package/install/_plugin-skills.mjs +73 -14
  7. package/install/claude.mjs +28 -4
  8. package/install/codex.mjs +33 -2
  9. package/install/copilot.mjs +145 -9
  10. package/install/index.mjs +10 -6
  11. package/install/templates/copilot-instructions.md +1 -1
  12. package/package.json +1 -1
  13. package/pipeline/agents/code-reviewer.md +58 -1
  14. package/pipeline/commands/multi-agent/SKILL.md +7 -5
  15. package/pipeline/commands/multi-agent/analysis/SKILL.md +7 -7
  16. package/pipeline/commands/multi-agent/analysis-resolve/SKILL.md +1 -1
  17. package/pipeline/commands/multi-agent/build-optimize/SKILL.md +7 -7
  18. package/pipeline/commands/multi-agent/channels/SKILL.md +5 -5
  19. package/pipeline/commands/multi-agent/dev/SKILL.md +23 -18
  20. package/pipeline/commands/multi-agent/dev-autopilot/SKILL.md +19 -13
  21. package/pipeline/commands/multi-agent/dev-local/SKILL.md +14 -12
  22. package/pipeline/commands/multi-agent/dev-local-autopilot/SKILL.md +17 -12
  23. package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +1 -1
  24. package/pipeline/commands/multi-agent/help/SKILL.md +4 -4
  25. package/pipeline/commands/multi-agent/ios-coding-standard/SKILL.md +2 -2
  26. package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +4 -4
  27. package/pipeline/commands/multi-agent/resume/SKILL.md +1 -1
  28. package/pipeline/commands/multi-agent/review/SKILL.md +5 -5
  29. package/pipeline/commands/multi-agent/scan/SKILL.md +1 -1
  30. package/pipeline/commands/multi-agent/search/SKILL.md +1 -1
  31. package/pipeline/commands/multi-agent/setup/SKILL.md +6 -6
  32. package/pipeline/commands/multi-agent/{finish → ship}/SKILL.md +12 -12
  33. package/pipeline/commands/multi-agent/testflight-validation/SKILL.md +1 -1
  34. package/pipeline/commands/multi-agent/update/SKILL.md +5 -2
  35. package/pipeline/commands/sim-test.md +2 -2
  36. package/pipeline/lib/credential-store-resolver.sh +16 -0
  37. package/pipeline/lib/credential-store.sh +47 -4
  38. package/pipeline/lib/fetch-figma-annotations.sh +26 -28
  39. package/pipeline/lib/figma-screenshot.sh +28 -39
  40. package/pipeline/lib/figma-token.sh +63 -0
  41. package/pipeline/multi-agent-refs/analysis-template.md +1 -1
  42. package/pipeline/multi-agent-refs/android-guide.md +1 -1
  43. package/pipeline/multi-agent-refs/channels/issue-comment.md +1 -1
  44. package/pipeline/multi-agent-refs/component-dispatch.md +2 -2
  45. package/pipeline/multi-agent-refs/cross-cli-contract.md +4 -4
  46. package/pipeline/multi-agent-refs/features/dev-critic.md +2 -2
  47. package/pipeline/multi-agent-refs/features/model-fallback.md +35 -2
  48. package/pipeline/multi-agent-refs/features/plan-todos.md +1 -1
  49. package/pipeline/multi-agent-refs/features/repo-map.md +1 -1
  50. package/pipeline/multi-agent-refs/features/review-multi-repo.md +3 -3
  51. package/pipeline/multi-agent-refs/features/shadow-git.md +1 -1
  52. package/pipeline/multi-agent-refs/features/skill-conformance.md +116 -0
  53. package/pipeline/multi-agent-refs/features/verify-by-test.md +1 -1
  54. package/pipeline/multi-agent-refs/generate-issue.md +1 -1
  55. package/pipeline/multi-agent-refs/multi-repo-integration-build.md +1 -1
  56. package/pipeline/multi-agent-refs/phases/log-format.md +4 -4
  57. package/pipeline/multi-agent-refs/phases/modes.md +7 -7
  58. package/pipeline/multi-agent-refs/phases/phase-0-init.md +13 -11
  59. package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +17 -15
  60. package/pipeline/multi-agent-refs/phases/phase-2-planning.md +7 -7
  61. package/pipeline/multi-agent-refs/phases/phase-3-dev.md +28 -13
  62. package/pipeline/multi-agent-refs/phases/phase-4-review.md +90 -58
  63. package/pipeline/multi-agent-refs/phases/phase-5-test.md +7 -7
  64. package/pipeline/multi-agent-refs/phases/phase-6-commit.md +8 -8
  65. package/pipeline/multi-agent-refs/phases/phase-7-report.md +8 -8
  66. package/pipeline/multi-agent-refs/phases.md +13 -13
  67. package/pipeline/multi-agent-refs/progress-contract.md +2 -2
  68. package/pipeline/multi-agent-refs/rules.md +7 -5
  69. package/pipeline/multi-agent-refs/swiftui-guide.md +1 -1
  70. package/pipeline/multi-agent-refs/tracker-contract.md +16 -15
  71. package/pipeline/preferences-template.json +7 -1
  72. package/pipeline/rules/figma-pipeline.md +2 -2
  73. package/pipeline/schemas/agent-state.schema.json +333 -79
  74. package/pipeline/schemas/criteria-manifest.schema.json +228 -0
  75. package/pipeline/schemas/migrations/prefs-2.4.0-to-2.5.0.mjs +64 -0
  76. package/pipeline/schemas/prefs.schema.json +118 -262
  77. package/pipeline/schemas/reviewer-output.schema.json +48 -3
  78. package/pipeline/schemas/token-budget.json +34 -10
  79. package/pipeline/schemas/triage-output.schema.json +112 -27
  80. package/pipeline/scripts/cost-table.json +7 -4
  81. package/pipeline/scripts/gc-worktrees.sh +1 -1
  82. package/pipeline/scripts/gen-mode-dispatch.mjs +6 -6
  83. package/pipeline/scripts/match-skills.mjs +37 -4
  84. package/pipeline/scripts/migrate-prefs.mjs +88 -17
  85. package/pipeline/scripts/phase-tracker.sh +14 -3
  86. package/pipeline/scripts/pre-commit-check.sh +49 -2
  87. package/pipeline/scripts/skill-conformance.mjs +960 -0
  88. package/pipeline/scripts/smoke-schema-validation.sh +17 -4
  89. package/pipeline/scripts/uninstall.mjs +35 -9
  90. package/pipeline/scripts/validate-reviewer.mjs +108 -1
  91. package/pipeline/skills/.skill-manifest.json +1 -1
  92. package/pipeline/skills/.skills-index.json +36 -9
  93. package/pipeline/skills/shared/README.md +15 -12
  94. package/pipeline/skills/shared/core/apple-archive-compliance/SKILL.md +1 -0
  95. package/pipeline/skills/shared/core/apple-archive-compliance/references/rules.yml +167 -0
  96. package/pipeline/skills/shared/core/google-play-compliance/SKILL.md +1 -0
  97. package/pipeline/skills/shared/core/google-play-compliance/references/rules.yml +184 -0
  98. package/pipeline/skills/shared/core/multi-agent/SKILL.md +10 -10
  99. package/pipeline/skills/shared/core/multi-agent-analysis/SKILL.md +4 -4
  100. package/pipeline/skills/shared/core/multi-agent-analysis-resolve/SKILL.md +3 -3
  101. package/pipeline/skills/shared/core/multi-agent-build-optimize/SKILL.md +2 -2
  102. package/pipeline/skills/shared/core/multi-agent-create-jira/SKILL.md +1 -1
  103. package/pipeline/skills/shared/core/multi-agent-dev/SKILL.md +6 -5
  104. package/pipeline/skills/shared/core/multi-agent-dev-autopilot/SKILL.md +7 -6
  105. package/pipeline/skills/shared/core/multi-agent-dev-local/SKILL.md +4 -3
  106. package/pipeline/skills/shared/core/multi-agent-dev-local-autopilot/SKILL.md +2 -1
  107. package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +2 -2
  108. package/pipeline/skills/shared/core/multi-agent-ios-coding-standard/SKILL.md +1 -1
  109. package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +4 -4
  110. package/pipeline/skills/shared/core/multi-agent-review/SKILL.md +5 -5
  111. package/pipeline/skills/shared/core/multi-agent-scan/SKILL.md +1 -1
  112. package/pipeline/skills/shared/core/multi-agent-search/SKILL.md +1 -1
  113. package/pipeline/skills/shared/core/{multi-agent-finish → multi-agent-ship}/SKILL.md +8 -8
  114. package/pipeline/skills/shared/external/ios-coding-standard/SKILL.md +44 -5
  115. package/pipeline/skills/shared/external/ios-coding-standard/modules/_TEMPLATE.yml +82 -0
  116. package/pipeline/skills/shared/external/ios-coding-standard/references/STANDARD.md +169 -10
  117. package/pipeline/skills/shared/external/ios-coding-standard/references/rules.yml +335 -16
  118. package/pipeline/skills/skills-index.md +11 -8
@@ -16,14 +16,14 @@ xcodebuild ... 2>&1 | tee "$WORKTREE/.build.log"
16
16
  # Gate 3: Tests pass (xcodebuild test/gradle test/pytest/npm test) - tee output to a log
17
17
  <test-command> 2>&1 | tee "$WORKTREE/.test.log"
18
18
  # Gate 4: Secrets - run the scanner against the staged diff
19
- bash pipeline/scripts/pre-commit-check.sh
19
+ bash $HOME/.claude/scripts/pre-commit-check.sh
20
20
  ```
21
21
 
22
22
  **Default-FAIL evidence gate (required before recording any pass):** a green exit code is not enough - the captured log must actually show success. Before marking build/test passed, run the evidence gate against the tee'd log; it fails CLOSED when the log is missing, empty, or shows failure markers:
23
23
 
24
24
  ```bash
25
- node pipeline/scripts/evidence-gate.mjs --claim build --status passed --evidence "$WORKTREE/.build.log" || BUILD_PASS=false
26
- node pipeline/scripts/evidence-gate.mjs --claim test --status passed --evidence "$WORKTREE/.test.log" || TEST_PASS=false
25
+ node $HOME/.claude/scripts/evidence-gate.mjs --claim build --status passed --evidence "$WORKTREE/.build.log" || BUILD_PASS=false
26
+ node $HOME/.claude/scripts/evidence-gate.mjs --claim test --status passed --evidence "$WORKTREE/.test.log" || TEST_PASS=false
27
27
  ```
28
28
 
29
29
  This prevents a false "it built" claim with no log behind it. On exit 1, treat the gate as failed (do NOT proceed to AI review) and surface the gate's `reason`.
@@ -81,13 +81,7 @@ If changes include UI files (iOS: `*View.swift`, `*Screen.swift`, `*Cell.swift`;
81
81
  - Missing safe area / keyboard avoidance → **important**
82
82
  - Hardcoded colors instead of system/semantic colors → **suggestion**
83
83
 
84
- **iOS - SwiftUI interaction & accessibility conventions** (skills: `figma-navigation`, `figma-overlays`, `figma-bottom-sheets`, plus the enriched `figma-to-swiftui` accessibility rules). Gated to changed SwiftUI files; each is native-SwiftUI-first unless the project's `figma-config` `ui.*` declares a custom system, in which case check against that system instead:
85
- - A reusable component that **routes itself** (hardcoded `NavigationLink(destination: ConcreteScreen())`, calls a router) instead of emitting a typed `Output` → **important** (breaks reuse).
86
- - A reusable component that **presents its own app-level overlay** (toast/alert/modal) instead of emitting an intent for the caller → **important** (double-fires when composed).
87
- - `AnyView` routed through an overlay center / data-modal path (rich content should be a local `.sheet/.modal { }`) → **important**.
88
- - Loading HUD not **ref-counted** / missing `end` on an error path / no min-show (flash) → **important**.
89
- - A pinned bottom CTA built as a presented sheet instead of `.safeAreaInset(edge:.bottom)`, or detents that aren't ascending / use magic-number heights → **suggestion**.
90
- - **Accessibility over-annotation** (redundant `.isButton` on `Button` / `.isStaticText` on `Text`, `.accessibilityLabel` duplicating visible text, a per-enum-variant a11y key for non-interactive visual state, hints on self-explanatory actions) → **suggestion**; **missing** decorative-image hiding or interactive-state `.accessibilityValue` → **important**.
84
+ **iOS - SwiftUI interaction & accessibility conventions.** Gated to changed SwiftUI files. These are rule-registry territory as of v14.0.0, not a list transcribed here: Step 1.78 resolves them from whichever registry declares SwiftUI scope, so the criteria and their severities live in one place instead of drifting between this doc and the skill. Reviewers receive the resolved rule IDs. Native-SwiftUI-first unless the project's `figma-config` `ui.*` declares a custom system, in which case check against that system. Reference skills, when no registry covers the change: `figma-navigation`, `figma-overlays`, `figma-bottom-sheets`, `figma-to-swiftui`.
91
85
 
92
86
  **Android - Material Design compliance** (skills: `compose-components`, `android-architecture`):
93
87
  - Non-Material3 component when M3 equivalent exists → **suggestion**
@@ -102,7 +96,7 @@ Device-level audit runs in Phase 5 (`audit-guide.md`).
102
96
 
103
97
  #### Step 1.6 - Repo Map Injection (advisory, opt-in)
104
98
 
105
- **Gated by `prefs.global.repoMap.enabled`** (default: `false`). Same pattern as Phase 1 Step 2.5 - runs `pipeline/scripts/repo-map.mjs` and injects the result as `${REPO_MAP}` into each reviewer's prompt context. Reviewers treat it the same way the priority files block is treated: advisory hint, not gospel.
99
+ **Gated by `prefs.global.repoMap.enabled`** (default: `false`). Same pattern as Phase 1 Step 2.5 - runs `$HOME/.claude/scripts/repo-map.mjs` and injects the result as `${REPO_MAP}` into each reviewer's prompt context. Reviewers treat it the same way the priority files block is treated: advisory hint, not gospel.
106
100
 
107
101
  Phase 4 reuses the cached map from Phase 1 when both phases run in the same task (avoid recomputing the same scan twice). The orchestrator caches under `state.repoMap.<sha>` keyed by `git rev-parse HEAD`; cache miss → regenerate. When `enabled=false`, both phases skip the script entirely and no `${REPO_MAP}` placeholder appears in prompts (no empty-string artefact).
108
102
 
@@ -115,16 +109,16 @@ Before dispatching reviewers, run the deterministic diff risk scorer and inject
115
109
  Score the diff **once, in full** (no `--top`): Steps 1.76/1.77 need every scored file (a shrinking test file ranked 20th; whether ANY file is high-stakes). The top-N hint is derived from that report, not a second git walk.
116
110
 
117
111
  ```bash
118
- RISK_FULL=$(node pipeline/scripts/diff-risk-score.mjs \
112
+ RISK_FULL=$(node $HOME/.claude/scripts/diff-risk-score.mjs \
119
113
  --base "$BASE_BRANCH" --head HEAD \
120
114
  --task-id "$TASK_ID" 2>/dev/null)
121
- echo "$RISK_FULL" | node pipeline/scripts/validate-diff-risk.mjs - >/dev/null 2>&1 || RISK_FULL=""
115
+ echo "$RISK_FULL" | node $HOME/.claude/scripts/validate-diff-risk.mjs - >/dev/null 2>&1 || RISK_FULL=""
122
116
 
123
117
  # Priority hint for the reviewer prompts: top 5 of the full report.
124
118
  RISK_JSON=$([ -n "$RISK_FULL" ] && jq -c '.files |= (sort_by(-.score) | .[:5])' <<< "$RISK_FULL" || echo "")
125
119
  ```
126
120
 
127
- **Signals & weights** (see `pipeline/schemas/diff-risk.schema.json`):
121
+ **Signals & weights** (see `$HOME/.claude/schemas/diff-risk.schema.json`):
128
122
 
129
123
  | Signal | Weight | Triggers when |
130
124
  |---|---|---|
@@ -142,13 +136,13 @@ RISK_JSON=$([ -n "$RISK_FULL" ] && jq -c '.files |= (sort_by(-.score) | .[:5])'
142
136
  **Gate behavior**: this step is **never blocking**. If risk scoring fails (git error, parse error, validator rejection), continue with no priority hint - reviewers receive the full diff in their default order. Failures are logged via metrics:
143
137
 
144
138
  ```bash
145
- [ -z "$RISK_JSON" ] && pipeline/scripts/log-metric.sh "$TASK_ID" 4 review.diff_risk_skipped reason=$REASON
139
+ [ -z "$RISK_JSON" ] && $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.diff_risk_skipped reason=$REASON
146
140
  ```
147
141
 
148
142
  On success, emit a single summary metric:
149
143
 
150
144
  ```bash
151
- pipeline/scripts/log-metric.sh "$TASK_ID" 4 review.diff_risk \
145
+ $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.diff_risk \
152
146
  top_files=$(jq '.files | length' <<< "$RISK_JSON") \
153
147
  max_score=$(jq '.totals.max_score' <<< "$RISK_JSON") \
154
148
  loc_added=$(jq '.totals.loc_added' <<< "$RISK_JSON")
@@ -161,9 +155,9 @@ pipeline/scripts/log-metric.sh "$TASK_ID" 4 review.diff_risk \
161
155
  Step 1.75 uses `test_lines_removed` as an advisory hint only - too weak for what it detects: a suite made green by deleting tests instead of fixing code. This turns the signal into blocking findings triage must adjudicate. Pure function of the full report, no git, no LLM.
162
156
 
163
157
  ```bash
164
- TEST_INTEGRITY_JSON=$(printf '%s' "$RISK_FULL" | node pipeline/scripts/test-integrity-gate.mjs 2>/dev/null || echo "")
158
+ TEST_INTEGRITY_JSON=$(printf '%s' "$RISK_FULL" | node $HOME/.claude/scripts/test-integrity-gate.mjs 2>/dev/null || echo "")
165
159
  TI_COUNT=$(jq -r '.count // 0' <<< "${TEST_INTEGRITY_JSON:-{\}}" 2>/dev/null || echo 0)
166
- [ "$TI_COUNT" -gt 0 ] && pipeline/scripts/log-metric.sh "$TASK_ID" 4 review.test_integrity findings="$TI_COUNT"
160
+ [ "$TI_COUNT" -gt 0 ] && $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.test_integrity findings="$TI_COUNT"
167
161
  ```
168
162
 
169
163
  `findings[]` are reviewer-shaped (`test_integrity`, `blocking`), so they merge into the reviewer findings at Step 3.0 and need no triage-prompt or `validate-triage.mjs` change. Triage keeps each blocking unless the removal is justified per the immutable-test rule (spec changed AND commit body names the test) → `deferred[]`.
@@ -175,18 +169,51 @@ The gate never blocks the phase; it *emits* blocking findings. Empty or unreadab
175
169
  On a trivial diff every reviewer agrees and the extra models plus triage are paid for nothing. Reviewer count comes from the same report, no LLM.
176
170
 
177
171
  ```bash
178
- SCOPE_JSON=$(printf '%s' "$RISK_FULL" | node pipeline/scripts/review-scope.mjs 2>/dev/null \
172
+ SCOPE_JSON=$(printf '%s' "$RISK_FULL" | node $HOME/.claude/scripts/review-scope.mjs 2>/dev/null \
179
173
  || echo '{"scope":"full","reason":"no risk report - failing safe"}')
180
174
  REVIEW_SCOPE=$(jq -r '.scope // "full"' <<< "$SCOPE_JSON")
181
- pipeline/scripts/log-metric.sh "$TASK_ID" 4 review.scope scope="$REVIEW_SCOPE"
175
+ $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.scope scope="$REVIEW_SCOPE"
182
176
  ```
183
177
 
184
178
  `single` (Reviewer 1 only) requires **all** of: churn <= 20 lines, `totals.max_score` < 3.0, and no `security_path` / `migration` / `public_api` / `no_test_change` / `test_lines_removed` on any file. Anything else → `full`.
185
179
 
186
180
  Fails safe in one direction only: empty report, parse failure or validator rejection all resolve to `full`, because skipping a reviewer trades coverage for cost. `consensus.reviewerCount` records what actually ran, so a single-reviewer run never reads as cross-model agreement. Opt-out: `prefs.global.reviewScopeGate = false` forces `full`.
187
181
 
182
+ #### Step 1.78 - Criteria resolution (skill conformance, required)
183
+
184
+ Resolves WHAT the changed code was supposed to honour, before any reviewer sees it. Zero LLM. Full contract: [`$HOME/.claude/multi-agent-refs/features/skill-conformance.md`]($HOME/.claude/multi-agent-refs/features/skill-conformance.md).
185
+
186
+ ```bash
187
+ node $HOME/.claude/scripts/skill-conformance.mjs \
188
+ --diff "$WORKTREE/.review-diff.txt" --state "$WORKTREE/agent-state.json" \
189
+ --repo "$WORKTREE" --out "$WORKTREE/.pipeline/criteria-manifest.json"
190
+ CRIT_RC=$?
191
+ CRITERIA=$(cat "$WORKTREE/.pipeline/criteria-manifest.json" 2>/dev/null || echo '{}')
192
+ $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.criteria \
193
+ rules="$(jq -r '.selectedRuleCount // 0' <<< "$CRITERIA")" \
194
+ ledger="$(jq -r '.ledger.source // "derived"' <<< "$CRITERIA")"
195
+ [ "$CRIT_RC" != "0" ] && HALT "criteria could not be resolved (rc=$CRIT_RC) - see resolutionFailure / unparseableRegistries"
196
+ ```
197
+
198
+ Output conforms to `$HOME/.claude/schemas/criteria-manifest.schema.json`. The four parts that matter downstream:
199
+
200
+ - **`selectedRules[]` is the denominator.** Rule IDs from every registry whose declared `scope` matches the diff, persisted BEFORE the reviewers run so the set cannot be renegotiated after one has seen the diff. That is what makes "applied completely" answerable rather than "looks fine": a reviewer finding nothing must still return a verdict per ID. Discovery is declared via `standards-registry:` frontmatter, so no stack-specific skill is named here.
201
+ - **`coverage.declaredGaps[]` + `droppedReasons`.** An uncovered language, and every out-of-scope rule, is reported with a reason. "No rule applied", "every rule passed" and "the rule set was narrowed" must never render the same.
202
+ - **`ledger`.** `state.telemetry.skillCalls[]` corroborates only; `ledger.source` defaults to `derived` and coverage is never computed from self-report. A declared skill the resolver cannot bind is flagged.
203
+ - **`findings[]`.** Reviewer-shaped, merging at Step 3.0 alongside test-integrity: expired / unexplained / unknown-ID exception markers, plus any registry whose delegated linter is not wired here (those rules are unverified, so reporting no violations reports that nothing was measured).
204
+
205
+ **Any non-zero exit halts** (1 = setup error incl. a bad `--skills-root`; 2 = no root resolved, unparseable registry, or a declared path escaping its skill dir): continuing would drop a rule set from the denominator, and since the reviewer validator skips the checklist when zero rules were selected, the run would report clean over criteria never loaded. A coverage gap halts only under `prefs.global.skillConformance.blockOnCoverageGap`. No opt-out for the stage or the exception-expiry check, on the same grounds as Step 1.76.
206
+
188
207
  #### Step 1.8 - Figma visual-fidelity context (when task carries a Figma reference)
189
208
 
209
+ **Dev-mode inputs.** Phases 1 and 2 never run in the `--dev` family, so three inputs this phase was written around are absent. Substitute them and RECORD the substitution - a step that could not run and a step that passed must not read the same, or the completeness claim cannot be checked:
210
+
211
+ | Absent input | Substitute |
212
+ |---|---|
213
+ | `detectedStack` (Phase 1 Step 2) | the language census in `criteria-manifest.json` → `languages`, from the diff's file extensions |
214
+ | Phase 1 analysis summary + Phase 2 plan (triage scope, cache prefix) | the task description plus the inline task list Phase 3 generated for itself |
215
+ | `state.evidence.figma[]` (this step), Step 2.8 visual conformance | nothing. Record each as `not-applicable (no Phase 1 evidence)` in `consensus.visualConformance`; never silently omit |
216
+
190
217
  When `state.evidence.figma[]` is non-empty, the reviewer subagents MUST receive the captured screenshot URLs / paths and the canonical-component name (from `state.evidence.figma[i].screenshotUrl` and `state.evidence.figma[i].codeConnectSnippets[0].componentName` when present) so they can compare visual fidelity. Pass them inline in each reviewer prompt under a `## Figma evidence` block, one row per frame.
191
218
 
192
219
  Visual-fidelity mismatches against the captured screenshot are BLOCKING findings, not nits:
@@ -208,7 +235,7 @@ When `state.figmaAccess.tier === 3` (user-attached screenshot, no Code Connect s
208
235
 
209
236
  Phase 4 sends the same diff to every reviewer and then to triage, so the diff is the dominant token cost. Two measures keep it bounded:
210
237
 
211
- **Shared cache prefix.** Build the reviewer and triage prompts so the large invariant context - the full diff, the Phase 1 analysis summary, the Phase 2 plan - is a byte-identical leading block across all dispatches in this iteration. Only the per-reviewer focus + skill line varies, and it goes AFTER the shared block. When the host supports prompt caching, the 2nd/3rd reviewer and the triage call then read that prefix at the discounted cache-read rate instead of re-billing it as fresh input. Forward the host-reported cache-read count as `tokens_cached` per the Token telemetry contract so the saving lands in the cost ledger.
238
+ **Shared cache prefix.** Build the reviewer and triage prompts so the large invariant context - the full diff, the `${CRITERIA}` block from Step 1.78, the Phase 1 analysis summary, the Phase 2 plan - is a byte-identical leading block across all dispatches in this iteration. Only the per-reviewer focus + skill line varies, and it goes AFTER the shared block. `${CRITERIA}` goes in the prefix, identical for every reviewer: subsetting it per reviewer would invalidate the prefix for the whole panel and re-bill the largest block in the phase. Per-reviewer emphasis stays a one-line pointer in the suffix. When the host supports prompt caching, the 2nd/3rd reviewer and the triage call then read that prefix at the discounted cache-read rate instead of re-billing it as fresh input. Forward the host-reported cache-read count as `tokens_cached` per the Token telemetry contract so the saving lands in the cost ledger.
212
239
 
213
240
  **Single-repo diff cap.** If the diff exceeds the Phase 4 token allowance (`token-budget.json`), truncate the largest files and append a footer `[truncated - full diff in file://$WORKTREE/.review-diff.txt]`, writing the full diff to that path. Reviewers and triage receive the same capped view + the marker so they can flag "review the full diff manually." Log `review.diff_truncated bytes_dropped=<N>`. (Multi-repo already caps the combined diff at 80% of budget; this is the single-repo equivalent.)
214
241
 
@@ -220,32 +247,27 @@ Launch Agent instances **in parallel** using the shared `code-reviewer` subagent
220
247
 
221
248
  | Reviewer | subagent_type | Claude Code | Copilot CLI | Codex CLI | Focus | Skills Referenced |
222
249
  | ---------- | --------------- | --- | --- | --- | --- | --- |
223
- | Reviewer 1 | `code-reviewer` | `claude-fable-5` | `claude-opus-4-8` | `gpt-5.6` @ `xhigh` | Deep security + architecture | `api-security-best-practices`, `architecture` |
250
+ | Reviewer 1 | `code-reviewer` | `claude-fable-5` | `claude-opus-5` | `gpt-5.6` @ `xhigh` | Deep security + architecture | `api-security-best-practices`, `architecture` |
224
251
  | Reviewer 2 | `code-reviewer` | (not dispatched) | `gpt-5.4` | `gpt-5.4` @ `high` | Edge cases, different perspective | cross-model diversity |
225
- | Reviewer 3 | `code-reviewer` | `claude-sonnet-4-6` | `claude-sonnet-4-6` | `gpt-5.6` @ `medium` | Quality + correctness + naming | `clean-code`, stack-specific skill |
226
- | Triage | triage persona | `claude-fable-5` | `claude-opus-4-8` | `gpt-5.6` @ `max` | Filter false positives + out-of-scope | - |
252
+ | Reviewer 3 | `code-reviewer` | `claude-sonnet-5` | `claude-sonnet-5` | `gpt-5.6` @ `medium` | Quality + correctness + naming | `clean-code`, stack-specific skill |
253
+ | Triage | triage persona | `claude-fable-5` | `claude-opus-5` | `gpt-5.6` @ `max` | Filter false positives + out-of-scope | - |
227
254
 
228
255
  Reviewer count per host: **Claude Code 2, Copilot CLI 3, Codex CLI 3**.
229
256
 
230
257
  #### Codex CLI - two constraints that fail silently
231
258
 
232
- Both were measured against Codex 0.145, not inferred. Getting either wrong produces
233
- a review that looks like it ran.
259
+ Both were measured against Codex 0.145, not inferred, and both produce a review that
260
+ looks like it ran. The full statement lives in the managed block at `~/.codex/AGENTS.md`
261
+ (always loaded on that host, so it is not restated here):
234
262
 
235
- 1. **Every `spawn_agent` that sets `model` or `reasoning_effort` MUST pass
236
- `fork_turns: "none"`** (or a positive integer). Codex states that a full-history
237
- fork "inherits the parent model and reasoning effort and does not accept
238
- overrides" - so omitting `fork_turns` silently collapses the whole panel onto
239
- the orchestrator's model, with three reviewers reporting from one perspective
240
- and no error anywhere.
241
- 2. **Three concurrent children is the ceiling.** Codex advertises 4 concurrency
242
- slots *including the orchestrator*, so a 3-reviewer panel saturates it and a 4th
243
- child queues instead of running in parallel. The reviewer count on Codex is
244
- capability-derived, not a preference.
263
+ 1. **`fork_turns: "none"` on every `spawn_agent` that sets `model` or
264
+ `reasoning_effort`** - a full-history fork discards the override and collapses the
265
+ panel onto one model, silently.
266
+ 2. **Three concurrent children is the ceiling** - 4 slots including the orchestrator,
267
+ so the reviewer count on Codex is capability-derived, not a preference.
245
268
 
246
- Codex also refuses to spawn sub-agents at all unless instructions ask for it
247
- explicitly; the pipeline's `~/.codex/AGENTS.md` managed block carries that
248
- authorization. Without it Phase 4 degrades to a single in-thread review.
269
+ Sub-agent delegation itself is authorized by that same managed block; without it Phase 4
270
+ degrades to a single in-thread review.
249
271
 
250
272
  **Single-vendor caveat.** Every Codex reviewer is an OpenAI model, so the
251
273
  cross-vendor disagreement that Claude Code and Copilot CLI get for free is absent.
@@ -257,7 +279,7 @@ in the triage note when all three agree on a borderline finding.
257
279
 
258
280
  Each reviewer inherits the `code-reviewer` agent's focus areas (Security, Architecture, Quality, Performance) and output contract. The orchestrator overrides only the model and the stack-specific skill per-reviewer - no prompt duplication.
259
281
 
260
- **Model override wiring:** `code-reviewer.md` declares `preferredModel: fable`, so Reviewer 1 uses the persona default (Fable 5). Reviewer 2 (Copilot-only, `gpt-5.4`) and Reviewer 3 (`claude-sonnet-4-6`) set `PHASE_MODEL_OVERRIDE=<model>` before dispatch - the orchestrator exports `CLAUDE_CODE_SUBAGENT_MODEL` on Claude Code, or passes `--model` on Copilot CLI. Full precedence rule: `skills/shared/core/multi-agent/SKILL.md#agent-dispatch--per-persona-model-routing-v610`. Fable dispatches are subject to the fallback contract (`$HOME/.claude/multi-agent-refs/features/model-fallback.md`): dispatch-error retry walks `fable -> opus -> sonnet` and budget-ceiling downgrade.
282
+ **Model override wiring:** `code-reviewer.md` declares `preferredModel: fable`, so Reviewer 1 uses the persona default (Fable 5). Reviewer 2 (Copilot-only, `gpt-5.4`) and Reviewer 3 (`claude-sonnet-5`) set `PHASE_MODEL_OVERRIDE=<model>` before dispatch - the orchestrator exports `CLAUDE_CODE_SUBAGENT_MODEL` on Claude Code, or passes `--model` on Copilot CLI. Full precedence rule: `skills/shared/core/multi-agent/SKILL.md#agent-dispatch--per-persona-model-routing-v610`. Fable dispatches are subject to the fallback contract (`$HOME/.claude/multi-agent-refs/features/model-fallback.md`): dispatch-error retry walks `fable -> opus -> sonnet` and budget-ceiling downgrade.
261
283
 
262
284
  **Stack-specific skills loaded per reviewer** (from Phase 1 `detectedStack`). On Claude Code, Reviewer 2 (GPT-5.4) is not dispatched - its skill column is ignored. On Copilot CLI all three columns are used.
263
285
 
@@ -298,33 +320,38 @@ reading a diff cannot see spacing; something has to compare against the design.
298
320
  Skip only when the diff has no UI change. Record the outcome in
299
321
  `consensus.visualConformance` so Phase 7 reports whether it ran.
300
322
 
301
- **iOS/Swift - interaction & convention checks (conditional).** When the diff touches SwiftUI UI files (`*View.swift`, `*Screen.swift`, `*Configuration.swift`, `*+Modifiers.swift`), the iOS reviewers additionally apply the component/interaction conventions documented in the analysis doc (Section 14 Code Connect mapping) and, where the repo has the marketplace component toolkit enabled (`ai-ios-engineering-toolkit`), that plugin's navigation / overlay / bottom-sheet conventions (interaction: emit-intent vs self-route/self-present; native-SwiftUI-first vs the project's `ui.*` custom system) plus the component accessibility rules (minimalism). These back the Step 1.5 iOS convention checks. Generic across SwiftUI projects - not tied to any one app. Omit when the diff has no SwiftUI UI changes (keeps the reviewer prompt lean).
323
+ **iOS/Swift - interaction & convention checks (conditional).** Step 1.78 resolves these. Where no registry covers the change, reviewers fall back to the analysis doc (Section 14 Code Connect mapping) and, when `ai-ios-engineering-toolkit` is enabled, that plugin's navigation / overlay / bottom-sheet + accessibility conventions.
302
324
 
303
- **Module review guides (conditional, all stacks).** A module in the repo may carry its own CLAUDE guide - a convention/checklist file living somewhere in the module's directory tree that the host CLI never auto-loads. When a changed file's module has such a guide, the review must consult it. Discovery is deterministic, from the diff's changed paths: for each changed file, walk its directory chain up to the repo root and collect `CLAUDE.md`, `*-CLAUDE.md`, and `AGENTS.md` files (root-level ones excluded - the host CLI already loads those). Dedupe, cap at 5 (log any dropped). Inject into every reviewer prompt with the directive: read each guide before reviewing and apply its rules/checklist to the changed files under its directory - a guide governs only its own subtree, and guide violations are findings triaged by severity like any other. No guides found → no-op, prompt stays lean. Same contract as `/multi-agent:review` Step 2b.
325
+ **Module review guides (conditional, all stacks).** Step 1.78 resolves them into `criteria-manifest.json` → `moduleGuides`. Inject with the directive: read each guide, apply its rules to the changed files under its directory - a guide governs only its own subtree, and its violations are findings triaged like any other. Same contract as `/multi-agent:review` Step 2b.
304
326
 
305
327
  **Dispatch timeout (required, mirrors triage 3.3).** Reviewers run in parallel and triage waits on all of them, so one stalled reviewer hangs the phase. Bound each reviewer dispatch by `REVIEWER_TIMEOUT_SECONDS` (default 180). If a reviewer has not returned by the budget: log `review.reviewer_timeout reviewer=<name>`, treat that reviewer as absent, and proceed to triage with the reviewers that did return. The merged-findings count and `consensus.reviewerCount` reflect only the reviewers that returned. If **zero** reviewers return, retry Reviewer 1 once; on a second total failure HALT with `ERR: no reviewer returned within ${REVIEWER_TIMEOUT_SECONDS}s; resume with /multi-agent:resume #N.`. The Step 2.5 rebuttal round uses the same per-dispatch timeout. Never block indefinitely on a slow or dead reviewer dispatch.
306
328
 
307
329
  #### Output contract - reviewer step
308
330
 
309
- Step 2 produces N reviewer-output objects (one per dispatched reviewer), each conforming to `pipeline/schemas/reviewer-output.schema.json`. They are persisted to `state.reviewIterations[<iteration>].reviewers[]` and consumed by Step 3 (Fable triage) - never by Phase 6 directly. The triage step (below) is the producer of the only review artifact Phase 6 reads, conforming to `pipeline/schemas/triage-output.schema.json`.
331
+ Step 2 produces N reviewer-output objects (one per dispatched reviewer), each conforming to `$HOME/.claude/schemas/reviewer-output.schema.json`. They are persisted to `state.reviewIterations[<iteration>].reviewers[]` and consumed by Step 3 (Fable triage) - never by Phase 6 directly. The triage step (below) is the producer of the only review artifact Phase 6 reads, conforming to `$HOME/.claude/schemas/triage-output.schema.json`.
310
332
 
311
- **Subagent return format** - each reviewer returns JSON conforming to `pipeline/schemas/reviewer-output.schema.json`:
333
+ **Subagent return format** - each reviewer returns JSON conforming to `$HOME/.claude/schemas/reviewer-output.schema.json`:
312
334
 
313
335
  ```json
314
- {"findings":[{"severity":"blocking|important|suggestion","file":"...","line":N,"issue":"...","fix":"..."}],"approved":true|false}
336
+ {"findings":[{"severity":"blocking|important|suggestion","file":"...","line":N,"issue":"...","fix":"...","ruleId":"SEC-01","criteriaSource":"ios-coding-standard"}],
337
+ "conformance":[{"ruleId":"SEC-01","verdict":"conformant|violated|not-applicable","file":"...","line":N,"reason":"..."}],
338
+ "approved":true|false}
315
339
  ```
316
340
 
341
+ `ruleId` + `criteriaSource` appear on a finding that cites a rule from `${CRITERIA}`. `conformance` is required whenever Step 1.78 selected at least one rule, with exactly one row per selected ID and none outside the set.
342
+
317
343
  **Required: validator gate (deterministic) - run immediately after each reviewer returns, before merging findings.** Persist each reviewer's output and validate the file - the validator's exit code decides, not the LLM turn:
318
344
 
319
345
  ```bash
320
346
  REVIEWER_FILE="$WORKTREE/.pipeline/reviewer-$N.json"
321
347
  printf '%s' "$REVIEWER_JSON" > "$REVIEWER_FILE"
322
- node $HOME/.claude/scripts/validate-reviewer.mjs "$REVIEWER_FILE"
348
+ node $HOME/.claude/scripts/validate-reviewer.mjs "$REVIEWER_FILE" \
349
+ --criteria "$WORKTREE/.pipeline/criteria-manifest.json"
323
350
  ```
324
351
 
325
352
  Progress line: ` → checking validator validate-reviewer ({reviewer})`
326
353
 
327
- Exit 0 = valid. Exit 2 = contradiction (approved=true with blocking findings) - flip `approved` to `false`, continue. Exit 1 = malformed; gate protocol (fails CLOSED, same handling as the evidence gate): emit the validator stderr + `errors[]` verbatim, attempt ONE self-correction rework (re-invoke that reviewer with the errors quoted, overwrite the file), re-run the validator. If it fails again -> HALT the phase (no merge, no triage). Recovery hint: `ERR: reviewer output failed validate-reviewer.mjs twice. Inspect $REVIEWER_FILE against pipeline/schemas/reviewer-output.schema.json, then resume with /multi-agent:resume #N.`
354
+ Exit 0 = valid. Exit 2 = contradiction (approved=true with blocking findings) - flip `approved` to `false`, continue. With `--criteria`, exit 1 also covers the conformance checklist: a selected rule ID with no verdict, a verdict for an ID that was never selected, a `conformant` row with no file evidence, or a `violated` row with no matching finding. Those are the four ways a review can look complete without being complete, and the validator is what makes the checklist more than decoration - it is hand-written and does not apply `additionalProperties`, so an unchecked array would otherwise pass. Exit 1 = malformed; gate protocol (fails CLOSED, same handling as the evidence gate): emit the validator stderr + `errors[]` verbatim, attempt ONE self-correction rework (re-invoke that reviewer with the errors quoted, overwrite the file), re-run the validator. If it fails again -> HALT the phase (no merge, no triage). Recovery hint: `ERR: reviewer output failed validate-reviewer.mjs twice. Inspect $REVIEWER_FILE against $HOME/.claude/schemas/reviewer-output.schema.json, then resume with /multi-agent:resume #N.`
328
355
 
329
356
  #### Step 2.5 - Disagreement-round loop (opt-in)
330
357
 
@@ -379,7 +406,7 @@ PRIOR_ART="["
379
406
  for finding in $(jq -c '.findings[]' <<< "$MERGED_FINDINGS"); do
380
407
  issue=$(jq -r '.issue' <<< "$finding")
381
408
  file=$(jq -r '.file' <<< "$finding")
382
- hits=$(node pipeline/scripts/triage-memory.mjs query \
409
+ hits=$(node $HOME/.claude/scripts/triage-memory.mjs query \
383
410
  --issue "$issue" --file-glob "$(dirname "$file")/*" --top 3 2>/dev/null \
384
411
  | jq -c '.hits // []')
385
412
  PRIOR_ART="$PRIOR_ART$hits,"
@@ -394,7 +421,7 @@ The triage prompt MUST include a hedge: *"prior-art entries are context, not com
394
421
  **Rejected-preference brief (on by default via `prefs.global.learningsLedger.injectIntoTriage`).** Inject the durable rejected-preference list so triage does not re-accept a suggestion the team already rejected on this repo:
395
422
 
396
423
  ```bash
397
- node pipeline/scripts/learnings-ledger.mjs brief --max 20 2>/dev/null
424
+ node $HOME/.claude/scripts/learnings-ledger.mjs brief --max 20 2>/dev/null
398
425
  ```
399
426
 
400
427
  Exit 2 (empty ledger) skips silently. The `## Rejected review preferences` section names patterns the team chose not to act on; triage should lean toward `rejected` for a finding that restates one - but the same hedge applies (context, not command; a genuinely new instance can still be accepted).
@@ -410,11 +437,16 @@ For each finding, decide:
410
437
  - DEFERRED: real issue but out of current task scope → log for later, do not block
411
438
  - REJECTED: false positive, duplicate, style-only nitpick, or already correct
412
439
 
440
+ A finding carrying a `ruleId` cites a written rule from a registry the project
441
+ adopted, so it is not a matter of taste: it can be DEFERRED or REJECTED as out of
442
+ scope for THIS task, but "I would have written it differently" is not available as
443
+ a reason. Preserve `ruleId` and `criteriaSource` on every finding you keep.
444
+
413
445
  #### Output contract - triage step
414
446
 
415
- Step 3 produces a single triage-output object conforming to `pipeline/schemas/triage-output.schema.json` and persists it to `state.reviewIterations[<iteration>].triage`. This is the **only** Phase 4 artifact Phase 6 reads. Phase 6 commits MUST cite only `accepted` findings that were resolved; `deferred` items get linked in the PR description as follow-up work; `rejected` items never appear in any user-facing output.
447
+ Step 3 produces a single triage-output object conforming to `$HOME/.claude/schemas/triage-output.schema.json` and persists it to `state.reviewIterations[<iteration>].triage`. This is the **only** Phase 4 artifact Phase 6 reads. Phase 6 commits MUST cite only `accepted` findings that were resolved; `deferred` items get linked in the PR description as follow-up work; `rejected` items never appear in any user-facing output.
416
448
 
417
- Return ONLY valid JSON conforming to pipeline/schemas/triage-output.schema.json:
449
+ Return ONLY valid JSON conforming to $HOME/.claude/schemas/triage-output.schema.json:
418
450
  {
419
451
  "accepted": [{ "severity": "blocking|important|suggestion", "file": "...", "line": N, "issue": "...", "fix": "...", "reviewer": "fable|opus|sonnet|gpt" }],
420
452
  "deferred": [{ "finding": {...}, "reason": "..." }],
@@ -440,7 +472,7 @@ Progress line: ` → checking validator validate-triage`
440
472
  | Exit | Meaning | Action |
441
473
  | ----- | ---------------------------- | ------------------------------------------------- |
442
474
  | **0** | Valid and clean | Act on triage output as-is |
443
- | **1** | Invalid structure | Gate protocol (strict, fails CLOSED): (a) emit the validator's stderr and `errors[]` JSON into the log verbatim; (b) attempt ONE self-correction rework - re-prompt triage with the errors quoted, overwrite `$TRIAGE_FILE`; (c) re-run the validator. Second failure -> HALT the phase with the recovery hint: `ERR: triage output failed validate-triage.mjs twice. Inspect $TRIAGE_FILE against pipeline/schemas/triage-output.schema.json, fix or regenerate it, then resume with /multi-agent:resume #N.` |
475
+ | **1** | Invalid structure | Gate protocol (strict, fails CLOSED): (a) emit the validator's stderr and `errors[]` JSON into the log verbatim; (b) attempt ONE self-correction rework - re-prompt triage with the errors quoted, overwrite `$TRIAGE_FILE`; (c) re-run the validator. Second failure -> HALT the phase with the recovery hint: `ERR: triage output failed validate-triage.mjs twice. Inspect $TRIAGE_FILE against $HOME/.claude/schemas/triage-output.schema.json, fix or regenerate it, then resume with /multi-agent:resume #N.` |
444
476
  | **2** | Over-rejection guard tripped | Pause for human confirm (autopilot: log + accept) |
445
477
  | **3** | Contradiction auto-corrected | Proceed with `result.corrected` |
446
478
 
@@ -461,13 +493,13 @@ Failure fallback (timeout >120s, or agent crash before any JSON is produced): re
461
493
  Emit metrics per review pass for Phase 7 cost rollup:
462
494
 
463
495
  ```bash
464
- LOG_METRIC_FORWARD_TO_TRACKER=1 pipeline/scripts/log-metric.sh "$TASK_ID" 4 review.reviewer_call model=fable duration_ms=$R1_DURATION tokens_in=$R1_IN tokens_out=$R1_OUT # model=opus on Copilot CLI
496
+ LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.reviewer_call model=fable duration_ms=$R1_DURATION tokens_in=$R1_IN tokens_out=$R1_OUT # model=opus on Copilot CLI
465
497
  # GPT-5.4 metric emitted only on Copilot CLI (skip on Claude Code):
466
498
  [ "${CLI_HOST:-claude}" = "copilot" ] && \
467
- LOG_METRIC_FORWARD_TO_TRACKER=1 pipeline/scripts/log-metric.sh "$TASK_ID" 4 review.reviewer_call model=gpt-5.4 duration_ms=$GPT_DURATION tokens_in=$GPT_IN tokens_out=$GPT_OUT
468
- LOG_METRIC_FORWARD_TO_TRACKER=1 pipeline/scripts/log-metric.sh "$TASK_ID" 4 review.reviewer_call model=sonnet duration_ms=$SONNET_DURATION tokens_in=$SONNET_IN tokens_out=$SONNET_OUT
469
- LOG_METRIC_FORWARD_TO_TRACKER=1 pipeline/scripts/log-metric.sh "$TASK_ID" 4 review.triage_call model=fable duration_ms=$TRIAGE_DURATION tokens_in=$TRIAGE_IN tokens_out=$TRIAGE_OUT
470
- pipeline/scripts/log-metric.sh "$TASK_ID" 4 review.completed raw_count=$RAW accepted=$ACC deferred=$DEF rejected=$REJ approved=$APPROVED duration_ms=$DURATION
499
+ LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.reviewer_call model=gpt-5.4 duration_ms=$GPT_DURATION tokens_in=$GPT_IN tokens_out=$GPT_OUT
500
+ LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.reviewer_call model=sonnet duration_ms=$SONNET_DURATION tokens_in=$SONNET_IN tokens_out=$SONNET_OUT
501
+ LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.triage_call model=fable duration_ms=$TRIAGE_DURATION tokens_in=$TRIAGE_IN tokens_out=$TRIAGE_OUT
502
+ $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.completed raw_count=$RAW accepted=$ACC deferred=$DEF rejected=$REJ approved=$APPROVED duration_ms=$DURATION
471
503
  ```
472
504
 
473
505
  `LOG_METRIC_FORWARD_TO_TRACKER=1` mirrors `tokens_in`/`tokens_out`/`model` into `phase-tracker.sh` (see `$HOME/.claude/multi-agent-refs/progress-contract.md#token-telemetry-forwarding-v83`). On non-zero validator exit (1/2/3), also emit `triage.edge_case` with cause. Omit `tokens_in`/`tokens_out` if unavailable. Best-effort - never fails the pipeline.
@@ -515,7 +547,7 @@ Act **only on triage.accepted**:
515
547
  At the end of every fix/rework round (each Phase 3 re-entry that resolved accepted findings, including the final one), append ONE one-line root-cause lesson per resolved blocking/important finding to the existing learnings ledger (`learnings-ledger.mjs` - the store Phase 1 and triage already replay; never invent a parallel store):
516
548
 
517
549
  ```bash
518
- node pipeline/scripts/learnings-ledger.mjs add --kind fact \
550
+ node $HOME/.claude/scripts/learnings-ledger.mjs add --kind fact \
519
551
  --statement "<one-line rule/what-to-do so this class of finding does not recur (<=140 chars)>" \
520
552
  --diagnosis "<the causal WHY the failure happened, one line>" \
521
553
  --scope "<file-or-area glob from the finding>" --task "$TASK_ID"
@@ -1,6 +1,6 @@
1
1
  ### Phase 5: Test
2
2
 
3
- > **TLDR** - Optional test gate. Offers to boot the simulator/emulator (UI Bug Hunter) or hand off to the user for manual QA. Skipped in `autopilot` and `--dev` modes. If issues found, loops back to Phase 3.
3
+ > **TLDR** - Optional test gate. Offers to boot the simulator/emulator (UI Bug Hunter) or hand off to the user for manual QA. Needs an interactive prompt AND a worktree checkout, so it runs in `--dev` and the full pipeline but is dropped by every `autopilot` or `--local` variant. If issues found, loops back to Phase 3.
4
4
 
5
5
  <!-- progress-contract: applied -->
6
6
  Progress emission per `$HOME/.claude/multi-agent-refs/progress-contract.md` - lines for local-test prompt render, user-answer capture, repo checkout (if selected).
@@ -19,13 +19,13 @@ case "$STACK" in
19
19
  *) SCAN_STACK="" ;;
20
20
  esac
21
21
  if [ -n "$SCAN_STACK" ] && [ "${prefs_testGap_enabled:-true}" = "true" ]; then
22
- GAP_JSON=$(node pipeline/scripts/test-gap-scan.mjs \
22
+ GAP_JSON=$(node $HOME/.claude/scripts/test-gap-scan.mjs \
23
23
  --base "$BASE_BRANCH" --head HEAD --stack "$SCAN_STACK" 2>/dev/null)
24
- echo "$GAP_JSON" | node pipeline/scripts/validate-test-gap.mjs - >/dev/null 2>&1 || GAP_JSON=""
24
+ echo "$GAP_JSON" | node $HOME/.claude/scripts/validate-test-gap.mjs - >/dev/null 2>&1 || GAP_JSON=""
25
25
  fi
26
26
  ```
27
27
 
28
- **What the report contains** (per `pipeline/schemas/test-gap.schema.json`):
28
+ **What the report contains** (per `$HOME/.claude/schemas/test-gap.schema.json`):
29
29
 
30
30
  | Field | Meaning |
31
31
  |---|---|
@@ -49,7 +49,7 @@ fi
49
49
  **Telemetry**:
50
50
 
51
51
  ```bash
52
- LOG_METRIC_FORWARD_TO_TRACKER=0 pipeline/scripts/log-metric.sh "$TASK_ID" 5 test_gap.scanned \
52
+ LOG_METRIC_FORWARD_TO_TRACKER=0 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 5 test_gap.scanned \
53
53
  stack=$SCAN_STACK \
54
54
  sources=$(jq '.totals.sourcesScanned' <<< "$GAP_JSON") \
55
55
  gaps=$(jq '.totals.gapCount' <<< "$GAP_JSON")
@@ -100,7 +100,7 @@ Tier 1 / Tier 2 records print `screenshotUrl` from the captured evidence (Tier 2
100
100
  bash $HOME/.claude/scripts/phase-tracker.sh now 5 "awaiting local test (user)"
101
101
  bash $HOME/.claude/scripts/phase-tracker.sh render
102
102
  ```
103
- The waiting state persists in `tracker-state.json` across the handoff; `/multi-agent:finish` and `/multi-agent:manual-test` CONTINUE this state file and never re-init it (`$HOME/.claude/multi-agent-refs/tracker-contract.md` "Continuation runs").
103
+ The waiting state persists in `tracker-state.json` across the handoff; `/multi-agent:ship` and `/multi-agent:manual-test` CONTINUE this state file and never re-init it (`$HOME/.claude/multi-agent-refs/tracker-contract.md` "Continuation runs").
104
104
  6. If fix needed:
105
105
  - Branch already has WIP commit (from step 2) - changes are safe
106
106
  - **Heal stale admin state first** (same contract as Phase 0 - step 3's
@@ -152,7 +152,7 @@ Returns severity-tagged findings (Critical / High / Medium). Critical items bloc
152
152
  When the security-auditor or any other Phase 5 sub-agent runs, forward its token totals so Phase 7's Cost Breakdown captures Phase 5:
153
153
 
154
154
  ```bash
155
- LOG_METRIC_FORWARD_TO_TRACKER=1 pipeline/scripts/log-metric.sh "$TASK_ID" 5 audit.completed \
155
+ LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 5 audit.completed \
156
156
  model=opus tokens_in=$IN tokens_out=$OUT duration_ms=$DUR
157
157
  ```
158
158
 
@@ -15,9 +15,9 @@ Progress emission per `$HOME/.claude/multi-agent-refs/progress-contract.md` -
15
15
 
16
16
  #### Input contract
17
17
 
18
- Phase 6 consumes the latest Phase 4 triage output object conforming to `pipeline/schemas/triage-output.schema.json` (`accepted`, `deferred`, `rejected` buckets) plus the working-tree mutations from Phase 3. The commit body cites only `accepted` findings that were resolved; `deferred` items must be linked in the PR description as follow-up work, never silently dropped. If the triage output reports any unresolved blocking accepted finding, Phase 6 refuses to commit - the pipeline returns to Phase 3 for rework.
18
+ Phase 6 consumes the latest Phase 4 triage output object conforming to `$HOME/.claude/schemas/triage-output.schema.json` (`accepted`, `deferred`, `rejected` buckets) plus the working-tree mutations from Phase 3. The commit body cites only `accepted` findings that were resolved; `deferred` items must be linked in the PR description as follow-up work, never silently dropped. If the triage output reports any unresolved blocking accepted finding, Phase 6 refuses to commit - the pipeline returns to Phase 3 for rework.
19
19
 
20
- The "Build passes" checklist item below is evidence-gated, not self-asserted: Phase 6 trusts the `buildStatus.ok` that Phase 3 / Phase 4 recorded through `pipeline/scripts/evidence-gate.mjs` (a pass without a substantiating build/test log is treated as unverified and blocks the commit). The secret scanner (`pre-commit-check.sh`) runs on the staged diff as the final pre-commit gate.
20
+ The "Build passes" checklist item below is evidence-gated, not self-asserted: Phase 6 trusts the `buildStatus.ok` that Phase 3 / Phase 4 recorded through `$HOME/.claude/scripts/evidence-gate.mjs` (a pass without a substantiating build/test log is treated as unverified and blocks the commit). The secret scanner (`pre-commit-check.sh`) runs on the staged diff as the final pre-commit gate.
21
21
 
22
22
  #### Step 0 - Multi-Repo Integration Build
23
23
 
@@ -322,19 +322,19 @@ A task ending with one or more `pushStatus === "skipped"` repos does NOT bump `r
322
322
 
323
323
  ```bash
324
324
  for proj in $(jq -r '.projects[].name' "$STATE_FILE"); do
325
- pipeline/scripts/log-metric.sh "$TASK_ID" 6 commit.created repo=$proj sha=$SHA
326
- pipeline/scripts/log-metric.sh "$TASK_ID" 6 push.attempted repo=$proj attempts=$N status=$STATUS
327
- pipeline/scripts/log-metric.sh "$TASK_ID" 6 pr.opened repo=$proj url=$URL number=$N
325
+ $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 6 commit.created repo=$proj sha=$SHA
326
+ $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 6 push.attempted repo=$proj attempts=$N status=$STATUS
327
+ $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 6 pr.opened repo=$proj url=$URL number=$N
328
328
  done
329
- pipeline/scripts/log-metric.sh "$TASK_ID" 6 multi_repo.completed repos=$REPOS skipped=$SKIPPED
329
+ $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 6 multi_repo.completed repos=$REPOS skipped=$SKIPPED
330
330
  ```
331
331
 
332
332
  **Token forwarding:** the commit-message and PR-body generators run on a model. Forward those calls into the tracker so Phase 7's Cost Breakdown captures Phase 6:
333
333
 
334
334
  ```bash
335
- LOG_METRIC_FORWARD_TO_TRACKER=1 pipeline/scripts/log-metric.sh "$TASK_ID" 6 commit.message_generated \
335
+ LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 6 commit.message_generated \
336
336
  model=<sonnet|opus> tokens_in=$IN tokens_out=$OUT duration_ms=$DUR
337
- LOG_METRIC_FORWARD_TO_TRACKER=1 pipeline/scripts/log-metric.sh "$TASK_ID" 6 pr.body_generated \
337
+ LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 6 pr.body_generated \
338
338
  model=<sonnet|opus> tokens_in=$IN tokens_out=$OUT duration_ms=$DUR
339
339
  ```
340
340
 
@@ -147,7 +147,7 @@ Write `agent-log.md` with ALL sections:
147
147
 
148
148
  ## Plan Todo Rollup (when `prefs.global.planTodos.enabled`)
149
149
 
150
- Rendered via `pipeline/lib/plan-todos.sh list "$TASK_ID"`. Inline the resulting markdown checklist verbatim - one line per todo with status marker ([x] / [~] / [/] / [!]), ID, task text, and notes when present. Status summary one-liner ("4/6 done, 1 skipped, 1 failed") goes at the top from `plan-todos.sh status`.
150
+ Rendered via `$HOME/.claude/lib/plan-todos.sh list "$TASK_ID"`. Inline the resulting markdown checklist verbatim - one line per todo with status marker ([x] / [~] / [/] / [!]), ID, task text, and notes when present. Status summary one-liner ("4/6 done, 1 skipped, 1 failed") goes at the top from `plan-todos.sh status`.
151
151
 
152
152
  Skipped sections: when `planTodos.enabled` is false or no `plan.todos[]` was emitted, the Plan Todo block is omitted (no empty header). Phase 3 task-by-task progress still appears in the Timeline section either way.
153
153
 
@@ -187,9 +187,9 @@ Skipped sections: when `planTodos.enabled` is false or no `plan.todos[]` was emi
187
187
  **Telemetry emission** (mandatory): forward the phase's own LLM spend (humanizer + report compose calls) to the tracker, then emit the final event:
188
188
 
189
189
  ```bash
190
- LOG_METRIC_FORWARD_TO_TRACKER=1 pipeline/scripts/log-metric.sh "$TASK_ID" 7 report.compose \
190
+ LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 7 report.compose \
191
191
  model=$REPORT_MODEL tokens_in=$R_IN tokens_out=$R_OUT duration_ms=$R_DUR
192
- pipeline/scripts/log-metric.sh "$TASK_ID" 7 task.completed \
192
+ $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 7 task.completed \
193
193
  phases=$PHASE_COUNT review_cycles=$CYCLES lang=$PROMPT_LANG \
194
194
  channels_pr=$PR_STATUS channels_jira=$JIRA_STATUS \
195
195
  channels_confluence=$CONF_STATUS channels_wiki=$WIKI_STATUS \
@@ -203,7 +203,7 @@ pipeline/scripts/log-metric.sh "$TASK_ID" 7 task.completed \
203
203
  **Cost Breakdown emission (mandatory):** as part of agent-log compose, append the per-task cost block produced by `render-agent-log-cost.sh`. Best-effort - exit 2 (no tracker data) silently skipped:
204
204
 
205
205
  ```bash
206
- COST_BLOCK=$(bash pipeline/scripts/render-agent-log-cost.sh "$TASK_ID" 2>/dev/null) && \
206
+ COST_BLOCK=$(bash $HOME/.claude/scripts/render-agent-log-cost.sh "$TASK_ID" 2>/dev/null) && \
207
207
  printf '\n%s\n' "$COST_BLOCK" >> "$AGENT_LOG"
208
208
  ```
209
209
 
@@ -214,7 +214,7 @@ This is independent of the channels-side `reportContent.costSummary` (which gate
214
214
  ```bash
215
215
  TRIAGE_PATH="$WORKTREE/triage-output.json"
216
216
  if [ -f "$TRIAGE_PATH" ]; then
217
- node pipeline/scripts/triage-memory.mjs ingest \
217
+ node $HOME/.claude/scripts/triage-memory.mjs ingest \
218
218
  --triage "$TRIAGE_PATH" \
219
219
  --task-id "$TASK_ID" \
220
220
  --task-title "$TASK_TITLE" \
@@ -228,11 +228,11 @@ Best-effort. The corpus is JSONL at `~/.claude/memory/multi-agent/<repo-slug>/tr
228
228
 
229
229
  ```bash
230
230
  if [ -f "$TRIAGE_PATH" ]; then
231
- node pipeline/scripts/learnings-ledger.mjs from-triage \
231
+ node $HOME/.claude/scripts/learnings-ledger.mjs from-triage \
232
232
  --triage "$TRIAGE_PATH" --task "$TASK_ID" >/dev/null 2>&1 || true
233
233
  fi
234
234
  # Optionally capture a durable architectural fact the analysis established:
235
- # node pipeline/scripts/learnings-ledger.mjs add --kind fact \
235
+ # node $HOME/.claude/scripts/learnings-ledger.mjs add --kind fact \
236
236
  # --statement "<one-line fact>" --scope "<path-glob>" --task "$TASK_ID" >/dev/null 2>&1 || true
237
237
  ```
238
238
 
@@ -286,7 +286,7 @@ for entry in $(jq -c '.[]' memory-synthesis.json); do
286
286
  body=$(jq -r '.body' <<< "$entry")
287
287
  slug=$(jq -r '.slug' <<< "$entry")
288
288
  type=$(jq -r '.type' <<< "$entry")
289
- printf '%s\n' "$body" | bash pipeline/scripts/memory-save.sh "$PROJECT_ROOT" "$type" "$slug" -
289
+ printf '%s\n' "$body" | bash $HOME/.claude/scripts/memory-save.sh "$PROJECT_ROOT" "$type" "$slug" -
290
290
  done
291
291
  ```
292
292
 
@@ -20,7 +20,7 @@
20
20
 
21
21
  ```
22
22
  Normal: 0-Init -> 1-Analysis -> 2-Planning -> 3-Dev -> 4-Review -> 5-Test -> 6-Commit -> 7-Report
23
- --dev: 0-Init ------------------------------------------> 3-Dev -----------> 6-Commit -> 7-Report
23
+ --dev: 0-Init ------------------------------------------> 3-Dev -> 4-Review -> 5-Test -> 6-Commit -> 7-Report
24
24
  --local: Same as above but no worktree - works directly on local branch
25
25
  ```
26
26
 
@@ -42,9 +42,9 @@ Two channels run in parallel at every phase boundary. Both are required in their
42
42
  Phase 0 MUST initialize the tracker and register all 8 phases:
43
43
 
44
44
  ```bash
45
- pipeline/scripts/phase-tracker.sh init "$TASK_ID"
45
+ $HOME/.claude/scripts/phase-tracker.sh init "$TASK_ID"
46
46
  for p in 0:Init 1:Analysis 2:Planning 3:Dev 4:Review 5:Test 6:Commit 7:Report; do
47
- pipeline/scripts/phase-tracker.sh add "${p%%:*}" "${p#*:}"
47
+ $HOME/.claude/scripts/phase-tracker.sh add "${p%%:*}" "${p#*:}"
48
48
  done
49
49
  ```
50
50
 
@@ -55,20 +55,20 @@ This produces an initial card stack printed by both CLIs.
55
55
  As each phase enters/exits:
56
56
 
57
57
  ```bash
58
- pipeline/scripts/phase-tracker.sh update <N> in_progress # phase starts
58
+ $HOME/.claude/scripts/phase-tracker.sh update <N> in_progress # phase starts
59
59
  # ...do the work...
60
- pipeline/scripts/phase-tracker.sh update <N> completed # phase ends OK
60
+ $HOME/.claude/scripts/phase-tracker.sh update <N> completed # phase ends OK
61
61
  # or:
62
- pipeline/scripts/phase-tracker.sh update <N> failed # phase failed
63
- pipeline/scripts/phase-tracker.sh update <N> skipped # --dev mode etc.
62
+ $HOME/.claude/scripts/phase-tracker.sh update <N> failed # phase failed
63
+ $HOME/.claude/scripts/phase-tracker.sh update <N> skipped # --dev mode etc.
64
64
  ```
65
65
 
66
66
  For sub-phase progress (e.g. Phase 1's parallel Explore agents, Phase 4's reviewer dispatch + triage + validator gate, Phase 6's commit + push + PR sub-steps):
67
67
 
68
68
  ```bash
69
- pipeline/scripts/phase-tracker.sh sub <N> 1 "<sub name>" pending # register
70
- pipeline/scripts/phase-tracker.sh sub <N> 1 "<sub name>" in_progress # advance
71
- pipeline/scripts/phase-tracker.sh sub <N> 1 "<sub name>" completed # done
69
+ $HOME/.claude/scripts/phase-tracker.sh sub <N> 1 "<sub name>" pending # register
70
+ $HOME/.claude/scripts/phase-tracker.sh sub <N> 1 "<sub name>" in_progress # advance
71
+ $HOME/.claude/scripts/phase-tracker.sh sub <N> 1 "<sub name>" completed # done
72
72
  ```
73
73
 
74
74
  State persists at `$HOME/.claude/logs/multi-agent/<task_id>/tracker-state.json` - atomic writes mean concurrent worktrees don't corrupt each other.
@@ -78,9 +78,9 @@ State persists at `$HOME/.claude/logs/multi-agent/<task_id>/tracker-state.json`
78
78
  Call alongside tracker update for extra emphasis:
79
79
 
80
80
  ```bash
81
- pipeline/scripts/phase-banner.sh start <N> "<name>" "<one-line detail>"
82
- pipeline/scripts/phase-banner.sh end <N> done "<name>" "<short result>"
83
- pipeline/scripts/phase-banner.sh sub <N> 2 "<sub>" "<detail>"
81
+ $HOME/.claude/scripts/phase-banner.sh start <N> "<name>" "<one-line detail>"
82
+ $HOME/.claude/scripts/phase-banner.sh end <N> done "<name>" "<short result>"
83
+ $HOME/.claude/scripts/phase-banner.sh sub <N> 2 "<sub>" "<detail>"
84
84
  ```
85
85
 
86
86
  Banner status enum on `end`: `done` | `failed` | `skipped`. Anything else exits 64.
@@ -128,7 +128,7 @@ Overhead: ~40 bytes / line + one millisecond clock read. Negligible vs. any of t
128
128
  Every phase that dispatches a billable LLM agent MUST forward the call's token totals to the tracker so the agent-log Cost Breakdown stays complete. The single canonical call shape is:
129
129
 
130
130
  ```bash
131
- LOG_METRIC_FORWARD_TO_TRACKER=1 pipeline/scripts/log-metric.sh "$TASK_ID" <phase-id> <event> \
131
+ LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" <phase-id> <event> \
132
132
  model=<fable|opus|sonnet|haiku|gpt-5.4|gpt-5.6|gpt-5.6-terra> \
133
133
  tokens_in=$IN tokens_out=$OUT duration_ms=$DUR
134
134
  ```
@@ -142,7 +142,7 @@ This contract is enforced by `smoke-agent-log-cost.sh` (forwarder unit test) and
142
142
  When `prefs.global.costBudget.enabled` is true, run the budget check once at the END of every phase, right after that phase's token totals have been forwarded to the tracker:
143
143
 
144
144
  ```bash
145
- pipeline/scripts/cost-budget-check.mjs --task-id "$TASK_ID" \
145
+ $HOME/.claude/scripts/cost-budget-check.mjs --task-id "$TASK_ID" \
146
146
  --prefs "$HOME/.claude/multi-agent-preferences.json"
147
147
  ```
148
148