@mmerterden/multi-agent-pipeline 16.18.0 → 16.20.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (238) hide show
  1. package/CHANGELOG.md +68 -0
  2. package/README.md +5 -5
  3. package/README.tr.md +2 -2
  4. package/docs/FIGMA_PIPELINE.md +1 -1
  5. package/docs/adr/0001-three-model-triage.md +4 -2
  6. package/docs/architecture.md +2 -2
  7. package/docs/ecosystem.md +13 -11
  8. package/docs/features.md +2 -2
  9. package/index.js +1 -1
  10. package/install/_codex-agents.mjs +2 -2
  11. package/install/_common.mjs +25 -1
  12. package/install/_dev-only-files.mjs +3 -2
  13. package/install/_mcp-register.mjs +4 -3
  14. package/install/_plugin-skills.mjs +1 -3
  15. package/install/copilot.mjs +18 -9
  16. package/install/index.mjs +2 -4
  17. package/install/templates/copilot-instructions.md +7 -7
  18. package/package.json +4 -3
  19. package/pipeline/agents/android-architect.md +1 -0
  20. package/pipeline/agents/backend-architect.md +1 -0
  21. package/pipeline/agents/code-reviewer.md +1 -0
  22. package/pipeline/agents/dev-critic.md +2 -1
  23. package/pipeline/agents/explorer.md +1 -0
  24. package/pipeline/agents/ios-architect.md +1 -0
  25. package/pipeline/agents/security-auditor.md +1 -0
  26. package/pipeline/agents/task-clarifier.md +1 -0
  27. package/pipeline/claude-md-template.md +2 -2
  28. package/pipeline/commands/multi-agent/SKILL.md +5 -5
  29. package/pipeline/commands/multi-agent/complaint-analysis/SKILL.md +1 -1
  30. package/pipeline/commands/multi-agent/design-check/SKILL.md +1 -1
  31. package/pipeline/commands/multi-agent/help/SKILL.md +2 -2
  32. package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +1 -1
  33. package/pipeline/commands/multi-agent/manual-test/SKILL.md +5 -1
  34. package/pipeline/commands/multi-agent/resume/SKILL.md +1 -0
  35. package/pipeline/commands/multi-agent/review/SKILL.md +27 -14
  36. package/pipeline/commands/multi-agent/review-analysis/SKILL.md +1 -1
  37. package/pipeline/commands/multi-agent/store-ready/SKILL.md +24 -2
  38. package/pipeline/commands/multi-agent/sync/SKILL.md +2 -2
  39. package/pipeline/commands/sim-test.md +64 -20
  40. package/pipeline/lib/credential-inventory.sh +15 -2
  41. package/pipeline/lib/credential-store-resolver.sh +14 -4
  42. package/pipeline/lib/credential-store.sh +8 -2
  43. package/pipeline/lib/extract-conventions.sh +1 -14
  44. package/pipeline/lib/fetch-confluence.sh +12 -4
  45. package/pipeline/lib/fetch-crashlytics.sh +11 -8
  46. package/pipeline/lib/fetch-document.sh +1 -1
  47. package/pipeline/lib/fetch-figma-annotations.sh +7 -5
  48. package/pipeline/lib/fetch-fortify.sh +5 -3
  49. package/pipeline/lib/fetch-graylog.sh +5 -3
  50. package/pipeline/lib/figma-mcp-refresh.sh +1 -1
  51. package/pipeline/lib/figma-screenshot.sh +27 -24
  52. package/pipeline/lib/figma-token.sh +8 -4
  53. package/pipeline/lib/issue-fetcher.sh +0 -1
  54. package/pipeline/lib/jira-publish.sh +7 -5
  55. package/pipeline/lib/md2confluence-v3.py +13 -7
  56. package/pipeline/lib/multi-repo-pipeline.sh +18 -8
  57. package/pipeline/lib/plan-todos.sh +11 -0
  58. package/pipeline/lib/post-pr-review.sh +9 -2
  59. package/pipeline/lib/repo-cache.sh +18 -10
  60. package/pipeline/lib/review-watch.sh +60 -14
  61. package/pipeline/lib/shadow-git.sh +8 -4
  62. package/pipeline/lib/vercel-deploy.sh +2 -2
  63. package/pipeline/multi-agent-refs/_dev-context.md +5 -2
  64. package/pipeline/multi-agent-refs/analysis/locked.md +4 -4
  65. package/pipeline/multi-agent-refs/analysis/render.md +1 -1
  66. package/pipeline/multi-agent-refs/channels/pr.md +22 -4
  67. package/pipeline/multi-agent-refs/cross-cli-contract.md +1 -1
  68. package/pipeline/multi-agent-refs/features/autopilot-circuit-breaker.md +15 -2
  69. package/pipeline/multi-agent-refs/features/review-delta.md +89 -0
  70. package/pipeline/multi-agent-refs/features/scope-check.md +41 -0
  71. package/pipeline/multi-agent-refs/features/verify-by-test.md +6 -5
  72. package/pipeline/multi-agent-refs/features/worktree-finalize.md +1 -1
  73. package/pipeline/multi-agent-refs/outside-the-pipeline.md +6 -6
  74. package/pipeline/multi-agent-refs/payload-contracts.md +1 -1
  75. package/pipeline/multi-agent-refs/phases/log-format.md +1 -1
  76. package/pipeline/multi-agent-refs/phases/modes.md +1 -1
  77. package/pipeline/multi-agent-refs/phases/phase-0-init.md +4 -2
  78. package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +3 -3
  79. package/pipeline/multi-agent-refs/phases/phase-2-planning.md +5 -5
  80. package/pipeline/multi-agent-refs/phases/phase-3-dev.md +17 -2
  81. package/pipeline/multi-agent-refs/phases/phase-4-review.md +67 -24
  82. package/pipeline/multi-agent-refs/phases/phase-5-test.md +10 -0
  83. package/pipeline/multi-agent-refs/phases/phase-6-commit.md +2 -0
  84. package/pipeline/multi-agent-refs/phases/phase-7-report.md +4 -2
  85. package/pipeline/multi-agent-refs/rules.md +2 -2
  86. package/pipeline/multi-agent-refs/tracker-contract.md +1 -1
  87. package/pipeline/rules/figma-pipeline.md +18 -72
  88. package/pipeline/rules/outside-the-pipeline.md +4 -3
  89. package/pipeline/schemas/agent-state.schema.json +130 -1
  90. package/pipeline/schemas/dev-critic-output.schema.json +5 -0
  91. package/pipeline/schemas/prefs.schema.json +47 -0
  92. package/pipeline/schemas/reviewer-output.schema.json +7 -2
  93. package/pipeline/schemas/scope-check.schema.json +55 -0
  94. package/pipeline/schemas/token-budget.json +3 -3
  95. package/pipeline/schemas/triage-output.schema.json +12 -2
  96. package/pipeline/scripts/README.md +3 -2
  97. package/pipeline/scripts/_fingerprint.mjs +173 -0
  98. package/pipeline/scripts/_stack-routing.mjs +1 -1
  99. package/pipeline/scripts/agent-guard.py +102 -21
  100. package/pipeline/scripts/anonymize-findings.mjs +7 -6
  101. package/pipeline/scripts/build-skills-index.mjs +14 -3
  102. package/pipeline/scripts/cost-budget-check.mjs +5 -3
  103. package/pipeline/scripts/cost-lib.sh +0 -15
  104. package/pipeline/scripts/diff-explain.mjs +22 -12
  105. package/pipeline/scripts/evidence-gate.mjs +73 -5
  106. package/pipeline/scripts/finding-fingerprint.mjs +101 -0
  107. package/pipeline/scripts/gc-refs.sh +6 -2
  108. package/pipeline/scripts/gc-tmp.sh +1 -1
  109. package/pipeline/scripts/gc-worktrees.sh +1 -1
  110. package/pipeline/scripts/gen-mode-dispatch.mjs +3 -3
  111. package/pipeline/scripts/github-ssh-setup.sh +7 -2
  112. package/pipeline/scripts/graph-build.mjs +2 -2
  113. package/pipeline/scripts/jira-wiki-escape.mjs +2 -1
  114. package/pipeline/scripts/keychain.py +12 -11
  115. package/pipeline/scripts/learning-curve.mjs +1 -1
  116. package/pipeline/scripts/migrate-prefs.mjs +1 -1
  117. package/pipeline/scripts/output-quality-check.sh +3 -1
  118. package/pipeline/scripts/phase-tracker.sh +1 -1
  119. package/pipeline/scripts/phase0-exit-gate.mjs +2 -1
  120. package/pipeline/scripts/plan-coverage-gate.mjs +2 -1
  121. package/pipeline/scripts/pre-commit-check.sh +23 -13
  122. package/pipeline/scripts/prune-logs.sh +1 -1
  123. package/pipeline/scripts/render-agent-log-cost.sh +3 -1
  124. package/pipeline/scripts/render-cost-summary.sh +4 -2
  125. package/pipeline/scripts/render-work-summary.sh +5 -3
  126. package/pipeline/scripts/repo-map.mjs +3 -2
  127. package/pipeline/scripts/review-delta.mjs +217 -0
  128. package/pipeline/scripts/run-metrics.mjs +20 -0
  129. package/pipeline/scripts/scan-skills.sh +6 -2
  130. package/pipeline/scripts/scope-check-gate.mjs +90 -0
  131. package/pipeline/scripts/search-logs.sh +8 -6
  132. package/pipeline/scripts/sign-skills.sh +3 -1
  133. package/pipeline/scripts/smoke-cross-cli-behavior.sh +12 -5
  134. package/pipeline/scripts/triage-memory.mjs +25 -4
  135. package/pipeline/scripts/uninstall.mjs +20 -12
  136. package/pipeline/scripts/update-check.sh +2 -2
  137. package/pipeline/scripts/update-issue-progress.sh +6 -5
  138. package/pipeline/scripts/validate-analysis-doc.mjs +6 -6
  139. package/pipeline/scripts/validate-reviewer.mjs +6 -0
  140. package/pipeline/scripts/validate-triage.mjs +20 -0
  141. package/pipeline/scripts/verify-skills.sh +3 -1
  142. package/pipeline/scripts/worktree-finalize.sh +22 -10
  143. package/pipeline/skills/.skill-manifest.json +81 -57
  144. package/pipeline/skills/.skills-index.json +19 -19
  145. package/pipeline/skills/shared/README.md +8 -8
  146. package/pipeline/skills/shared/core/multi-agent/SKILL.md +7 -7
  147. package/pipeline/skills/shared/core/multi-agent-design-check/SKILL.md +1 -1
  148. package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +2 -2
  149. package/pipeline/skills/shared/core/multi-agent-local/SKILL.md +1 -1
  150. package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +2 -2
  151. package/pipeline/skills/shared/core/multi-agent-review/SKILL.md +23 -14
  152. package/pipeline/skills/shared/core/multi-agent-store-ready/SKILL.md +5 -0
  153. package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +2 -2
  154. package/pipeline/skills/shared/external/NOTICE-dimillian-skills.md +56 -0
  155. package/pipeline/skills/shared/external/accessibility-compliance-accessibility-audit/SKILL.md +0 -4
  156. package/pipeline/skills/shared/external/api-patterns/SKILL.md +12 -24
  157. package/pipeline/skills/shared/external/app-store-changelog/references/release-notes-guidelines.md +34 -0
  158. package/pipeline/skills/shared/external/app-store-changelog/scripts/collect_release_changes.sh +33 -0
  159. package/pipeline/skills/shared/external/architecture/SKILL.md +7 -9
  160. package/pipeline/skills/shared/external/debugging-strategies/SKILL.md +0 -4
  161. package/pipeline/skills/shared/external/fastapi-pro/SKILL.md +0 -1
  162. package/pipeline/skills/shared/external/github-actions-templates/SKILL.md +0 -14
  163. package/pipeline/skills/shared/external/hig-components-content/SKILL.md +13 -13
  164. package/pipeline/skills/shared/external/hig-components-layout/SKILL.md +16 -16
  165. package/pipeline/skills/shared/external/hig-components-status/SKILL.md +6 -6
  166. package/pipeline/skills/shared/external/hig-components-system/SKILL.md +13 -13
  167. package/pipeline/skills/shared/external/hig-foundations/SKILL.md +23 -23
  168. package/pipeline/skills/shared/external/hig-inputs/SKILL.md +18 -18
  169. package/pipeline/skills/shared/external/hig-patterns/SKILL.md +30 -30
  170. package/pipeline/skills/shared/external/hig-platforms/SKILL.md +11 -11
  171. package/pipeline/skills/shared/external/hig-technologies/SKILL.md +33 -33
  172. package/pipeline/skills/shared/external/ios-coding-standard/references/STANDARD.md +52 -52
  173. package/pipeline/skills/shared/external/ios-coding-standard/references/lint-local.sh +1 -1
  174. package/pipeline/skills/shared/external/ios-coding-standard/references/rules.yml +11 -11
  175. package/pipeline/skills/shared/external/ios-developer/SKILL.md +0 -1
  176. package/pipeline/skills/shared/external/ios-module-structure/SKILL.md +7 -3
  177. package/pipeline/skills/shared/external/localization-reuse-map/SKILL.md +9 -15
  178. package/pipeline/skills/shared/external/macos-spm-app-packaging/SKILL.md +0 -5
  179. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/bootstrap/Package.swift +17 -0
  180. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/bootstrap/Sources/MyApp/Resources/.keep +0 -0
  181. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/bootstrap/Sources/MyApp/main.swift +11 -0
  182. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/bootstrap/version.env +2 -0
  183. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/build_icon.sh +49 -0
  184. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/compile_and_run.sh +63 -0
  185. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/launch.sh +28 -0
  186. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/make_appcast.sh +82 -0
  187. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/package_app.sh +206 -0
  188. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/setup_dev_signing.sh +52 -0
  189. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/sign-and-notarize.sh +52 -0
  190. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/version.env +2 -0
  191. package/pipeline/skills/shared/external/macos-spm-app-packaging/references/packaging.md +17 -0
  192. package/pipeline/skills/shared/external/macos-spm-app-packaging/references/release.md +32 -0
  193. package/pipeline/skills/shared/external/macos-spm-app-packaging/references/scaffold.md +79 -0
  194. package/pipeline/skills/shared/external/monorepo-architect/SKILL.md +0 -1
  195. package/pipeline/skills/shared/external/nodejs-backend-patterns/SKILL.md +0 -4
  196. package/pipeline/skills/shared/external/swift-concurrency-expert/references/approachable-concurrency.md +63 -0
  197. package/pipeline/skills/shared/external/swift-concurrency-expert/references/swift-6-2-concurrency.md +272 -0
  198. package/pipeline/skills/shared/external/swift-concurrency-expert/references/swiftui-concurrency-tour-wwdc.md +33 -0
  199. package/pipeline/skills/shared/external/swiftui-performance-audit/references/code-smells.md +150 -0
  200. package/pipeline/skills/shared/external/swiftui-performance-audit/references/demystify-swiftui-performance-wwdc23.md +46 -0
  201. package/pipeline/skills/shared/external/swiftui-performance-audit/references/optimizing-swiftui-performance-instruments.md +29 -0
  202. package/pipeline/skills/shared/external/swiftui-performance-audit/references/profiling-intake.md +44 -0
  203. package/pipeline/skills/shared/external/swiftui-performance-audit/references/report-template.md +47 -0
  204. package/pipeline/skills/shared/external/swiftui-performance-audit/references/understanding-hangs-in-your-app.md +33 -0
  205. package/pipeline/skills/shared/external/swiftui-performance-audit/references/understanding-improving-swiftui-performance.md +52 -0
  206. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/app-wiring.md +201 -0
  207. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/async-state.md +96 -0
  208. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/components-index.md +46 -0
  209. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/controls.md +57 -0
  210. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/deeplinks.md +66 -0
  211. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/focus.md +90 -0
  212. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/form.md +97 -0
  213. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/grids.md +71 -0
  214. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/haptics.md +71 -0
  215. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/input-toolbar.md +51 -0
  216. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/lightweight-clients.md +93 -0
  217. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/list.md +86 -0
  218. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/loading-placeholders.md +38 -0
  219. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/macos-settings.md +71 -0
  220. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/matched-transitions.md +59 -0
  221. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/media.md +73 -0
  222. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/menu-bar.md +101 -0
  223. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/navigationstack.md +159 -0
  224. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/overlay.md +45 -0
  225. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/performance.md +62 -0
  226. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/previews.md +48 -0
  227. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/scroll-reveal.md +133 -0
  228. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/scrollview.md +87 -0
  229. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/searchable.md +71 -0
  230. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/sheets.md +155 -0
  231. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/split-views.md +72 -0
  232. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/tabview.md +114 -0
  233. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/theming.md +71 -0
  234. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/title-menus.md +93 -0
  235. package/pipeline/skills/shared/external/swiftui-ui-patterns/references/top-bar.md +49 -0
  236. package/pipeline/skills/shared/external/swiftui-view-refactor/references/mv-patterns.md +161 -0
  237. package/pipeline/skills/skills-index.md +8 -8
  238. package/pipeline/skills/shared/external/help-skills/SKILL.md +0 -166
@@ -163,10 +163,12 @@ do_list() {
163
163
  require_task_id "$task_id"
164
164
  local sd; sd=$(shadow_dir "$task_id")
165
165
  [ ! -d "$sd/.git" ] && { err "shadow not initialized for $task_id"; exit 1; }
166
- # Use the embedded `iso:` line from the commit message body so the listing
166
+ # Use the embedded `iso:` trailer from the commit message body so the listing
167
167
  # carries the snapshot timestamp without relying on committer date (which is
168
- # influenced by git env vars).
169
- sg "$sd" "$(pwd)" log --pretty=format:'%h %s' --no-decorate
168
+ # influenced by git env vars); committer date is the fallback for a snapshot
169
+ # written without one.
170
+ sg "$sd" "$(pwd)" log --pretty=format:'%h%x1f%(trailers:key=iso,valueonly,separator=%x20)%x1f%cI%x1f%s' --no-decorate \
171
+ | awk -F'\037' '{ d = $2; sub(/ .*/, "", d); if (d == "") d = $3; printf "%s %s %s\n", $1, d, $4 }'
170
172
  }
171
173
 
172
174
  do_restore() {
@@ -213,7 +215,9 @@ do_prune() {
213
215
  local older=""
214
216
  while [ "$#" -gt 0 ]; do
215
217
  case "$1" in
216
- --older-than-days) older="$2"; shift 2 ;;
218
+ --older-than-days)
219
+ [ "$#" -ge 2 ] || { err "--older-than-days needs a value"; exit 2; }
220
+ older="$2"; shift 2 ;;
217
221
  *) err "unknown flag: $1"; exit 2 ;;
218
222
  esac
219
223
  done
@@ -153,10 +153,10 @@ cmd_deploy() {
153
153
  if [ "$prod" = true ]; then
154
154
  cmd_args+=(--prod)
155
155
  fi
156
- cmd_args+=("${extra_args[@]}")
156
+ cmd_args+=(${extra_args[@]+"${extra_args[@]}"})
157
157
 
158
158
  # Refuse to run if the user is trying to pass --token=... via extra args.
159
- for a in "${extra_args[@]}"; do
159
+ for a in ${extra_args[@]+"${extra_args[@]}"}; do
160
160
  case "$a" in
161
161
  --token=*|--token)
162
162
  echo "ERROR: refused to deploy with --token argv. Pass the token via VERCEL_TOKEN env var instead." >&2
@@ -53,8 +53,11 @@ Selects extra repos the pipeline may touch beyond the primary repo(s) - typica
53
53
  entry, plus any selected `extras[]` that is not being given a worktree. The
54
54
  phases that consume them run hours later and have no access to the picker's
55
55
  return value, so a result that is not written here is a result nothing can
56
- read - Phase 4's platform-parity cross-check reads `state.siblings[]` and
57
- nothing else.
56
+ read. Phase 4's platform-parity cross-check reads `state.siblings[]` as the
57
+ fourth of its four counterpart sources, after `--with`,
58
+ `prefs.projects[<slug>].counterpartRoots[]` and the primary checkout's sibling
59
+ directories (`platform-parity.md`), so a submodule or a hand-picked repo that
60
+ is not written here is a candidate the check can never see.
58
61
 
59
62
  Resolve each entry's `stack` from its local checkout, with the marker table
60
63
  in `phases/phase-1-analysis.md` Step 2 (`.xcodeproj` / `Package.swift` →
@@ -1,14 +1,14 @@
1
- # Locked decisions (35)
1
+ # Locked decisions (36)
2
2
 
3
- > The 35 Locked decisions of the analysis flow. Loaded by `/multi-agent:analysis`, by `/multi-agent:analysis-resolve` (which inherits them) and by pipeline Phase 1 when it runs the analysis engine. Numbering is canonical: cite as `Locked <n> (<short label>)`.
3
+ > The 36 Locked decisions of the analysis flow. Loaded by `/multi-agent:analysis`, by `/multi-agent:analysis-resolve` (which inherits them) and by pipeline Phase 1 when it runs the analysis engine. Numbering is canonical: cite as `Locked <n> (<short label>)`.
4
4
 
5
5
  ### Index by category (v9.1.0+)
6
6
 
7
- Browse-friendly grouping of the 35 Locked decisions. Numbering stays canonical (matches the list below); the index is read-only navigation.
7
+ Browse-friendly grouping of the 36 Locked decisions. Numbering stays canonical (matches the list below); the index is read-only navigation.
8
8
 
9
9
  | Category | Decisions | Concern |
10
10
  |---|---|---|
11
- | **A. Governance** | 1, 5, 6, 7, 10, 26, 27, 32 | Run-level process rules: one feature per run, default output, auto-commit ban, punctuation policy, output picker timing, Pass B preview, evidence digest cache, analysis profile |
11
+ | **A. Governance** | 1, 5, 6, 7, 10, 26, 27, 32, 36 | Run-level process rules: one feature per run, default output, auto-commit ban, punctuation policy, output picker timing, Pass B preview, evidence digest cache, analysis profile, document reviewed before publish |
12
12
  | **B. Citation and Evidence** | 3, 4, 8, 11, 24, 30, 34 | Every fact in the doc traces back to a source: citation discipline, forward-looking spec, standards binding, repo-evidence reuse-first, Pass B footnote mandatory, analysis self-contained (pipeline-wide), references built from the evidence record |
13
13
  | **C. Output Format and Structure** | 2, 9, 13, 14, 16, 17, 20, 21, 25, 33, 35 | How the document is laid out: section omission rule, per-platform output split, Gherkin user stories, Goals + Non-Goals paired, Files-to-Add tag, API response variants exhaustive, localization mode (ownership-aware), References at the bottom, Lite mode, corporate backbone always renders, stack-optional render |
14
14
  | **D. Design Source and Pipeline Architecture** | 12, 22, 23 | Where design comes from and how the pipeline renders: Figma 3-tier access (BLOCKING), platform-agnostic template + Pass B render, convention extraction (Phase 1c) |
@@ -12,7 +12,7 @@
12
12
 
13
13
  **When `platforms[]` is empty** (Locked 35), the channels come from the evidence instead of from repo stack tags, so this loop runs once per derived channel (`mobile`, `web`, or one channel-agnostic pass). What drops is the development layer (corporate Part C, global Sections 13, 14, 15) and the Pass B projection are skipped, Section 20 carries a row recording that they await a repo selection, and the front-matter `platform` key reads `none`. Everything that does not need a target repository still renders in full.
14
14
 
15
- 3. **Humanizer pass (MANDATORY: actually invoke the `ai-common-toolkit:humanizer` skill on the rendered markdown - the punctuation grep alone does NOT satisfy this step)** (`technical-explanatory` tone for the scratch buffer; per-channel re-humanize happens in Phase 4 when actually emitting):
15
+ 3. **Humanizer pass (required: actually invoke the `ai-common-toolkit:humanizer` skill on the rendered markdown - the punctuation grep alone does NOT satisfy this step)** (`technical-explanatory` tone for the scratch buffer; per-channel re-humanize happens in Phase 4 when actually emitting):
16
16
  ```
17
17
  ai-common-toolkit:humanizer skill input:
18
18
  language: <tr|en>
@@ -14,14 +14,15 @@ The PR description targets code reviewers - it stays technical. Every adapter
14
14
  | 2 | `changes` | `## Değişiklikler` | `## Changes` | always |
15
15
  | 3 | `architecture` | `## Mimari Kararlar` | `## Architecture Decisions` | when a non-trivial design choice was made |
16
16
  | 4 | `verification` | `## Doğrulama` | `## Verification` | always |
17
- | 5 | `dependencies` | `## Bağımlılıklar` | `## Dependencies` | when deps added/removed/bumped |
18
- | 6 | `related` | `## İlgili` | `## Related` | always (Jira/issue ref; never `Closes/Fixes`) |
17
+ | 5 | `risk` | `## Risk ve Güvenlik` | `## Risk and Security` | when `state.diffRisk.signals` carries a high-stakes signal (`security_path`, `migration`, `public_api`, `no_test_change`, `test_lines_removed`) |
18
+ | 6 | `dependencies` | `## Bağımlılıklar` | `## Dependencies` | when deps added/removed/bumped |
19
+ | 7 | `related` | `## İlgili` | `## Related` | always (Jira/issue ref; never `Closes/Fixes`) |
19
20
 
20
21
  ### Section content rules
21
22
 
22
23
  **`summary`** - 1-3 sentences in `outputLanguage`. The "why" of the change. Past tense, no marketing voice. Code identifiers stay verbatim.
23
24
 
24
- **`changes`** - bullet list, one item per logically distinct change. Each bullet starts with the touched component and ends with a one-line "what". Use the stack's native file extensions / module paths - the example below shows the **shape**, not a stack lock-in:
25
+ **`changes`** - bullet list, one item per logically distinct change. Each bullet starts with the touched component and ends with a one-line "what". The source is `$WORKTREE/.pipeline/scope-check.json` `files[].reason` (Phase 3 Step 3.7): a file the dev could not justify there is a file this list cannot describe either, so the bullet quotes the gate output instead of inventing a reason. Use the stack's native file extensions / module paths - the example below shows the **shape**, not a stack lock-in:
25
26
 
26
27
  ```markdown
27
28
  ## Changes
@@ -50,6 +51,17 @@ Skeleton (the adapter fills the body with the actual stack-appropriate lines at
50
51
 
51
52
  Multi-repo PRs (one PR per repo) emit verification commands for that repo's stack only - never mix iOS + Android commands into a single PR body.
52
53
 
54
+ **`risk`** - only when `state.diffRisk.signals` (Phase 4 Step 1.75) contains a high-stakes signal. Four fixed lines, each answered, never left as a placeholder; the source is Phase 1 `touchedAreas` plus the signals themselves, and when a signal is present the absence of this section is a Phase 6 Step 3 blocker:
55
+
56
+ ```markdown
57
+ ## Risk and Security
58
+
59
+ - Auth flow touched: yes | no
60
+ - Secret handling changed: yes | no
61
+ - Data migration: yes | no
62
+ - Rollback: feature flag <name> | git revert <sha> | none, and why
63
+ ```
64
+
53
65
  **`dependencies`** - only when `Package.swift` / `Podfile` / `build.gradle` / `package.json` changed. Each entry: `package@old → new - reason`.
54
66
 
55
67
  **`related`** - flat list, plain text. Examples:
@@ -61,15 +73,21 @@ Multi-repo PRs (one PR per repo) emit verification commands for that repo's stac
61
73
  - Issue: #123
62
74
  - Confluence: <page-url> (if work referenced a spec)
63
75
  - Figma: <design-url> (if work referenced a design)
76
+
77
+ Follow-ups not done in this PR:
78
+ - <scope-check.json notDone[].what> - <why>
79
+ - <deferred triage finding> - <triage reason>
64
80
  ```
65
81
 
82
+ The follow-up list is present only when `scope-check.json` `notDone[]` or the final triage `deferred[]` is non-empty; the two sources merge into one list.
83
+
66
84
  Never use `Closes #N`, `Fixes #N`, `Resolves PROJ-X`. Issues require 4-approval close, the auto-close keywords break that contract.
67
85
 
68
86
  ### Assembly order (per run)
69
87
 
70
88
  ```
71
89
  1. Read agent-state.json (taskId, contextLinks, identity, language).
72
- 2. Build section bodies in markdown - summary first, then in the table order, skipping conditional sections that don't apply.
90
+ 2. Build section bodies in markdown - summary first, then in the table order, skipping conditional sections that don't apply. Section order is fixed: `summary` → `changes` → `architecture` (cond.) → `verification` → `risk` (cond.) → `dependencies` (cond.) → `related`.
73
91
  3. Run the assembled body through the `humanizer` skill.
74
92
  4. Apply Multi-repo cross-links (## Related PRs prepend when projects.length > 1).
75
93
  5. Dispatch per the Behaviour-by-remote table.
@@ -207,7 +207,7 @@ Future changes that break an item in the "stay identical" list must update **bot
207
207
 
208
208
  Each file has a different frontmatter schema. The sync flow transforms between them:
209
209
 
210
- ### 3.1 Claude Code - `commands/multi-agent/{cmd}.md`
210
+ ### 3.1 Claude Code - `commands/multi-agent/{cmd}/SKILL.md`
211
211
 
212
212
  ```yaml
213
213
  ---
@@ -2,7 +2,7 @@
2
2
 
3
3
  **Pattern**: autopilot runs with zero interaction, which is exactly when a silent failure loop is most expensive - an agent can burn a budget re-attempting the same broken fix, or thrash between two phases, with nobody watching. A circuit-breaker converts "keep going no matter what" into "keep going until a defined unsafe condition, then halt and hand back to the user." This is the sanctioned autopilot pause (same class as the Phase 7 channels pause): the run stops, records why, and waits for an explicit `resume`.
4
4
 
5
- **Gated by `prefs.global.autopilotCircuitBreaker`** (default: enabled; thresholds tunable). Halting is always safe, so the breaker itself defaults on. Disable per-run only with an explicit override. Complements, does not replace, the existing autopilot safety rules (build-fail max 3 retries, Phase 4 blocking-finding rework, destructive-op confirmations).
5
+ **Gated by `prefs.global.autopilotCircuitBreaker`** (`enabled` default true, `identicalFindingCycles` default 2, `maxReworkCycles` default 3; `schemas/prefs.schema.json`). Halting is always safe, so the breaker itself defaults on. Disable per-run only with an explicit override. Complements, does not replace, the existing autopilot safety rules (build-fail max 3 retries, Phase 4 blocking-finding rework, destructive-op confirmations).
6
6
 
7
7
  ## Trip conditions
8
8
 
@@ -18,12 +18,25 @@ Any one trips the breaker. All are evaluated from `agent-state.json` + telemetry
18
18
 
19
19
  Trigger 2 is the key addition over the plain build-retry cap: a build can "fail differently" three times (legitimate iteration) or "fail identically" twice (stuck). Only the identical-failure case is a stall; the retry cap catches the rest.
20
20
 
21
+ ## Wiring status
22
+
23
+ | Trigger | Evaluated by | Status |
24
+ |---|---|---|
25
+ | 2, finding half | `review-delta.mjs` exit 3 at Phase 4 Step 3.8: a blocking/important finding whose `fingerprint` (finding-fingerprint.mjs) stays in the accepted set for `identicalFindingCycles` consecutive rounds | **code** (v16.20.0) |
26
+ | 3 | Phase 3 re-entry item 6: the `retryCount === 3` hard-kill records the trip | **code** (v16.20.0) |
27
+ | 2, build-error half | needs a build-log signature normaliser | documented behaviour, no script yet |
28
+ | 1 | needs checkpoint-to-checkpoint artifact diffing | documented behaviour, no script yet |
29
+ | 4 | belongs to `cost-budget-check.mjs` | documented behaviour, no script yet |
30
+ | 5 | Phase 6 push | documented behaviour, no script yet |
31
+
32
+ State shape: `state.circuitBreaker = {tripped, trigger, detail, checkpoint: {phase, step, iteration}, trippedAt, counters: {identicalFindingCycles, reworkCycles}}` (`schemas/agent-state.schema.json`). The per-round classification the finding half reads lives in `state.reviewIterations[i].delta` (`new`, `stillPresent`, `resolved`, `downgraded`, `recurrence`, `plateau`). `smoke-autopilot-circuit-breaker.sh` asserts the schema fields, the scripts and the phase wiring, not only this prose.
33
+
21
34
  ## Action on trip
22
35
 
23
36
  1. Set `agent-state.json.circuitBreaker = {tripped: true, trigger: <#>, detail, checkpoint}` and flip `autopilot` handling to paused (the run does not continue unattended).
24
37
  2. Emit one actionable line per the progress contract: what tripped, the evidence (error signature / cycle count / spend vs ceiling), and the single next action (`resume #N` after a fix, or `kill #N`).
25
38
  3. Never auto-resolve the underlying cause - no force-anything, no conflict auto-merge, no budget self-raise. The breaker hands control back; it does not paper over the problem.
26
- 4. `resume #N` clears the tripped flag and continues from the recorded checkpoint. If the same trigger fires again immediately, the breaker re-trips (no silent bypass).
39
+ 4. `resume #N` clears `circuitBreaker.tripped` (keeping `counters`) and continues from the recorded checkpoint. If the same trigger fires again immediately, the breaker re-trips (no silent bypass).
27
40
 
28
41
  ## Why this is the right autopilot exception
29
42
 
@@ -0,0 +1,89 @@
1
+ # Feature: Cross-round review delta (Phase 4 Steps 2.1, 2.2, 3.8)
2
+
3
+ **Pattern**: reviewers re-read the whole diff every round with no memory of the round before, so they rediscover last round's findings in new words and nothing can tell "still broken" from "new". A finding therefore needs an identity that survives the fix: `finding-fingerprint.mjs` computes `F:xxxxxxxx` from the file and either the cited `ruleId` or the normalised issue text (lowercase, quotes stripped, paths reduced to basenames, digit runs collapsed). The line, the severity, the fix text and the reviewer take no part, because all of them change between rounds without the finding changing. `review-delta.mjs` then compares round N with round N-1 and reports `stillPresent`, `resolved`, `downgraded` (re-reported but no longer accepted by triage) and `new`, plus `recurrence`: how many consecutive rework cycles each survivor has lasted. That count is the autopilot circuit-breaker's trigger 2.
4
+
5
+ Gated by `prefs.global.autopilotCircuitBreaker` (`enabled` default true, `identicalFindingCycles` default 2). Computed by scripts, never by the model: an older triage JSON gets its fingerprints on the fly, and a reviewer that echoes one is respected but not relied on.
6
+
7
+ ## Files per round
8
+
9
+ Phase 4 Step 3.2.1 writes `$WORKTREE/.pipeline/triage-round-<N>.json` (N = `state.reviewIterations | length`), validates it, annotates it and copies it to `$WORKTREE/triage-output.json`, the name Phase 7, `worktree-finalize.sh`, `render-work-summary.sh` and `diff-explain.mjs` read. A copy rather than a symlink because the salvage is `cp -R`, and `.pipeline/` is on the salvage list, so every round survives into `artifactsPath`. Step 3.7 rewrites the round file; the copy is repeated after it.
10
+
11
+ ## Step 2.1 block: previous-round findings (iteration >= 2)
12
+
13
+ ```bash
14
+ ITERATION=$(jq '.reviewIterations | length' "$STATE_FILE")
15
+ PREV_ROUND="$WORKTREE/.pipeline/triage-round-$((ITERATION-1)).json"
16
+ PREV_BLOCK=""
17
+ [ "$ITERATION" -ge 2 ] && [ -f "$PREV_ROUND" ] && PREV_BLOCK=$(node $HOME/.claude/scripts/finding-fingerprint.mjs annotate "$PREV_ROUND" 2>/dev/null \
18
+ | jq -r '[.accepted[] | select(.severity=="blocking" or .severity=="important")] | sort_by(.severity != "blocking") | .[:40][] | "- \(.fingerprint) [\(.severity)] \(.file): \(.issue)"')
19
+ ```
20
+
21
+ Rendered at the end of the shared prefix (Step 1.9), identical for every reviewer and for the triage call of that iteration:
22
+
23
+ ```
24
+ <previous-round-findings>
25
+ Each entry below was accepted last round and sent for rework.
26
+ - If the issue is still present, report it again with the SAME fingerprint value and the current line.
27
+ - If it is fixed, omit it. Omission is how you report resolution; never emit a "resolved" finding.
28
+ - Any finding not listed here is new: leave fingerprint unset.
29
+ {PREV_BLOCK}
30
+ </previous-round-findings>
31
+ ```
32
+
33
+ The cap of 40 entries drops `important` before `blocking` so the prefix stays bounded. The block never enters the repo-stable prefix `prompt-assembly.md` describes: the diff it follows is already per-run. After each reviewer's validator gate, `finding-fingerprint.mjs annotate --in-place` fills in any fingerprint the reviewer left unset; `anonymize-findings.mjs` strips only identity keys, so the fingerprint reaches triage, and the triage prompt tells the model to preserve it verbatim.
34
+
35
+ ## Step 2.2 block: scope self-check (every iteration)
36
+
37
+ Phase 3 Step 3.7 wrote `$WORKTREE/.pipeline/scope-check.json` (contract: `features/scope-check.md`). Render it so reviewers judge the diff against the dev's stated scope and do not re-propose what was rejected:
38
+
39
+ ```bash
40
+ SCOPE_JSON="$WORKTREE/.pipeline/scope-check.json"
41
+ SCOPE_GATE=$(git -C "$WORKTREE" diff --name-only "origin/$BASE_BRANCH"...HEAD \
42
+ | node $HOME/.claude/scripts/scope-check-gate.mjs --check "$SCOPE_JSON" --diff-files - --advisory 2>/dev/null)
43
+ ```
44
+
45
+ ```
46
+ <scope-self-check>
47
+ Files and the reason the dev gave for touching each:
48
+ {jq -r '.files[] | "- \(.path): \(.reason)"' "$SCOPE_JSON"}
49
+ Files in the diff with no stated reason (flag as scope drift if the change is not obviously required):
50
+ {jq -r '.unjustified[]' <<< "$SCOPE_GATE"}
51
+ Deliberately not done (do not raise these as findings; they are known):
52
+ {jq -r '.notDone[] | "- \(.what) (\(.why))"' "$SCOPE_JSON"}
53
+ </scope-self-check>
54
+ ```
55
+
56
+ A missing record renders the block with `no scope-check.json written` and a `review.scope_check=missing` metric; the review proceeds, and the absence is itself information for the reviewer.
57
+
58
+ ## Step 3.8: delta + trigger 2
59
+
60
+ ```bash
61
+ TRIP=$(jq -r '.global.autopilotCircuitBreaker.identicalFindingCycles // 2' "$PREFS_FILE")
62
+ DELTA_JSON=$(node $HOME/.claude/scripts/review-delta.mjs --rounds-dir "$WORKTREE/.pipeline" --iteration "$ITERATION" --trip-cycles "$TRIP"); DELTA_RC=$?
63
+ jq -c --argjson d "$DELTA_JSON" --argjson i "$((ITERATION-1))" \
64
+ '{reviewIterations: (.reviewIterations | .[$i] += {delta: ($d + {computedAt: (now | todate)})})}' "$STATE_FILE" \
65
+ | node $HOME/.claude/scripts/write-state.mjs "$STATE_FILE"
66
+ $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.delta iteration=$ITERATION \
67
+ new=$(jq '.counts.new // 0' <<< "$DELTA_JSON") still_present=$(jq '.counts.stillPresent // 0' <<< "$DELTA_JSON") \
68
+ resolved=$(jq '.counts.resolved // 0' <<< "$DELTA_JSON") plateau=$(jq '.plateau // false' <<< "$DELTA_JSON") tripped=$([ "$DELTA_RC" -eq 3 ] && echo true || echo false)
69
+ ```
70
+
71
+ | Exit | Meaning | Action |
72
+ |---|---|---|
73
+ | 0 | Progress, or nothing to compare (iteration 1, missing previous round) | Continue to Step 4. |
74
+ | 3 | A blocking/important finding survived `identicalFindingCycles` consecutive rework cycles | **Autopilot** with `enabled`: trip the breaker. `state.circuitBreaker = {tripped: true, trigger: 2, detail: "finding <fingerprint> (<file>: <issue>) survived <n> consecutive rework cycles", checkpoint: {phase: 4, step: "3.8", iteration: N}, trippedAt, counters: {identicalFindingCycles: <max recurrence>, reworkCycles: N-1}}`, then the halt-visibility protocol from `phases/operations.md` (`status=paused`, `haltReason="4:circuit-breaker:identical-finding"`, tracker meta, the `>&2 HALT` line, the usage report). **Interactive modes**: do not trip; show `delta.stillPresent` and ask (picker-contract) `Continue rework` / `Escalate to me` / `Accept as deferred`, the last moving those findings to `deferred[]` with reason `circuit-breaker: accepted by user after <n> cycles`. |
75
+ | 1 | Unreadable round file | Log `review.delta_skipped reason=invalid` and continue; the delta is advisory and never blocks on its own failure. |
76
+
77
+ `plateau` (the still-present set unchanged from the previous delta) is logged, not acted on: at the default threshold it coincides with trigger 2 and with the Phase 3 `retryCount` hard-kill. The rework-storm cap itself is trigger 3, recorded by the Phase 3 re-entry. The Phase 3 reflection prompt names each accepted finding by `fingerprint` and quotes `delta.stillPresent` first, marked `STILL PRESENT after round N-1's fix`.
78
+
79
+ ## Why fingerprints are computed, not stored
80
+
81
+ Same argument `shortId()` in `_retrieval.mjs` makes for corpus rows: derived from the finding's own content, so existing triage files get ids without a migration and a re-annotated file keeps the ids it had. The known trade-off is over-merging: two findings in one file that differ only by a number share a fingerprint. The delta output prints `file: issue` beside every id and the trip needs two consecutive recurrences, so a merge is visible and cannot halt a run by itself. Under-merging (a reviewer that rewrites rather than echoes) only ever suppresses a trip; it never causes one.
82
+
83
+ ## Telemetry
84
+
85
+ `review.delta` once per iteration >= 2: `iteration`, `new`, `still_present`, `resolved`, `plateau`, `tripped`. `run-metrics.mjs` reports `reviewDelta.stillPresentFinal`, `resolvedTotal` and `tripped`; Phase 7 renders the three counts per round.
86
+
87
+ ## Reference
88
+
89
+ Scripts: `$HOME/.claude/scripts/_fingerprint.mjs`, `finding-fingerprint.mjs`, `review-delta.mjs`. Schemas: `reviewer-output` 1.2.0, `triage-output` 3.4.0, `dev-critic-output` (optional `fingerprint`), `agent-state` (`reviewIterations[].delta`, `circuitBreaker`). Breaker: `features/autopilot-circuit-breaker.md`. Smokes: `smoke-review-delta.sh`, `smoke-autopilot-circuit-breaker.sh`.
@@ -0,0 +1,41 @@
1
+ # Feature: Scope self-check (Phase 3 Step 3.7)
2
+
3
+ **Pattern**: Phase 4 reconstructs everything from the diff. The one thing it cannot reconstruct is why each file was touched and what was left out on purpose, so Dev states both before the handoff, and a deterministic gate checks the statement against the real diff. The record also carries the code-simplifier rationales (Step 3.6) that used to be discarded, and it feeds two later consumers: the `<scope-self-check>` block in the Phase 4 reviewer prefix and the PR body in Phase 6.
4
+
5
+ ## The record
6
+
7
+ `$WORKTREE/.pipeline/scope-check.json`, schema `schemas/scope-check.schema.json`:
8
+
9
+ ```json
10
+ {
11
+ "version": "1.0.0",
12
+ "taskId": "<taskId>",
13
+ "files": [{ "path": "<repo-relative path>", "reason": "<why this task needs this exact file>" }],
14
+ "notDone": [{ "what": "<change considered and not made>", "why": "out-of-scope | follow-up | rejected-abstraction" }],
15
+ "simplifier": { "applied": 0, "skipped": 0, "rationales": ["<Step 3.6 rationale strings>"] }
16
+ }
17
+ ```
18
+
19
+ Rules:
20
+
21
+ - One entry per file in the diff, no entry for a file outside it. A reason names the task requirement, never "while I was here".
22
+ - `notDone` carries every hypothetical the dev chose not to defend against and every abstraction it considered and rejected, so review does not re-propose them.
23
+ - Component tasks (`taskType === "component"`) write the record too; the plugin skill's file table is the `files[]` source.
24
+
25
+ ## The gate
26
+
27
+ ```bash
28
+ git -C "$WORKTREE" diff --name-only "origin/$BASE_BRANCH"...HEAD \
29
+ | node $HOME/.claude/scripts/scope-check-gate.mjs --check "$WORKTREE/.pipeline/scope-check.json" --diff-files -
30
+ ```
31
+
32
+ Output `{ok, justified, unjustified[], unlisted[], notDone}`. Exit 1 lists `unjustified[]` (in the diff, no reason or an empty one); `unlisted[]` (in the record, not in the diff) is reported and never fatal. Complete the record once and re-run. A second exit 1 does not block: log `dev.scope_check=incomplete unjustified=<n>` and hand off; Phase 4 receives the gate output inside `<scope-self-check>`, so reviewers see which files arrived without a stated reason. `--advisory` (used by Phase 4) always exits 0. Progress line: ` → checking scope self-check ({n} files, {m} not done)`.
33
+
34
+ ## Consumers
35
+
36
+ - Phase 4 Step 2.2 renders file reasons, unjustified files and `notDone[]` into the shared reviewer prefix (`features/review-delta.md`).
37
+ - Phase 6 Step 3 builds the PR `## Changes` bullets from `files[].reason` and lists `notDone[]` under `## Related` as "Follow-ups not done in this PR", merged with the final triage `deferred[]` (`channels/pr.md`).
38
+
39
+ ## Reference
40
+
41
+ Script: `$HOME/.claude/scripts/scope-check-gate.mjs`. Schema: `$HOME/.claude/schemas/scope-check.schema.json`. Smoke: `smoke-scope-check.sh`. Unit: `test/scope-check-gate.test.mjs`.
@@ -5,7 +5,7 @@
5
5
  **Gated by `prefs.global.verifyByTest.enabled`** (default: `false`). When enabled, after triage 3.6 and before Step 4, IF the validated triage output contains at least one `accepted` blocking finding:
6
6
 
7
7
  1. Dispatch ONE verifier sub-agent for the iteration (model: `verifyByTest.model`, default `sonnet`) - never one dispatch per finding. Input: up to `verifyByTest.maxFindings` (default 3) accepted blocking findings, the diff hunks for their files, and the Phase 1 test conventions.
8
- 2. Per finding, the verifier writes ONE minimal repro test asserting the correct behavior the finding claims is broken, then runs ONLY that test via the Phase 3 single-test invocation (`xcodebuild test -only-testing:`, `pytest {file}::{name}`, `npm test -- --testPathPattern=`, `./gradlew test --tests`) under `acquire_build_lock`/`release_build_lock`, log tee'd to `$WORKTREE/.pipeline/verify-<i>.test.log`.
8
+ 2. Per finding, the verifier writes ONE minimal repro test asserting the correct behavior the finding claims is broken, then runs ONLY that test via the Phase 3 single-test invocation (`xcodebuild test -only-testing:`, `pytest {file}::{name}`, `npm test -- --testPathPattern=`, `./gradlew test --tests`) under `acquire_build_lock`/`release_build_lock`. One green run is not trusted: the same single-test invocation runs `verifyByTest.repeatCount` times (default 3; value 1 disables the repeat) as a shell loop, each run's log tee'd to `$WORKTREE/.pipeline/verify-<i>-<k>.test.log` for `k` in `1..repeatCount`. ALL runs must pass, and each log must clear `evidence-gate.mjs --claim test --status passed`, before the outcome counts as "passes". A first run that fails as predicted ends the loop early (`confirmed` needs no repeat).
9
9
  3. Stamp each processed finding with a `verification` object (triage-output schema v3.2.0) and re-run `validate-triage.mjs` on the mutated triage file under the standard 3.2.1 gate protocol.
10
10
  4. Findings beyond `maxFindings` keep their judgment-only verdict (log `verify_by_test=cap-exceeded`).
11
11
  5. The whole step is bounded by `verifyByTest.stepTimeoutSec` (default 600); on breach or verifier crash, remaining findings keep judgment-only verdicts and the pipeline proceeds. Never blocks.
@@ -15,7 +15,8 @@
15
15
  | Repro test outcome | `verification.result` | Action |
16
16
  |---|---|---|
17
17
  | Fails as the finding predicts | `confirmed` | Finding stays accepted blocking. Repro test KEPT in the worktree, recorded in `redTests[]`. |
18
- | Passes (not reproducible) | `not-reproduced` | Downgrade gated by `evidence-gate.mjs --claim test --status passed` on the test log (exit 0 required). Finding moves `accepted[]` -> `deferred[]` with reason `verify-by-test: not reproduced - repro test <testRef> passed`; repro test file deleted. Evidence-gate failure -> treat as `inconclusive`. |
18
+ | Passes (not reproducible) | `not-reproduced` | Only when ALL `verifyByTest.repeatCount` runs pass and every `verify-<i>-<k>.test.log` clears `evidence-gate.mjs --claim test --status passed` (exit 0 required on each). Finding moves `accepted[]` -> `deferred[]` with reason `verify-by-test: not reproduced - repro test <testRef> passed`; repro test file deleted. Evidence-gate failure on any run -> treat as `inconclusive`. |
19
+ | Passes on some runs, fails on others | `inconclusive` | Flake, not proof either way. Finding stays accepted blocking, `verification.note = "flaky: passed k/N"`, repro test file deleted, telemetry line gains `flaky=<count>`. |
19
20
  | Compile error / timeout / not unit-testable | `inconclusive` | Finding stays accepted blocking (judgment stands). Partial test deleted, cause in `verification.note`. |
20
21
 
21
22
  Downgrades go to `deferred`, never `rejected`: triage judged the issue real, and deferred items surface in the Phase 7 report for a human eye.
@@ -30,12 +31,12 @@ After Step 3.7, the only uncommitted verifier artifacts are the confirmed repro
30
31
 
31
32
  ## Telemetry
32
33
 
33
- One `review.verify_by_test` metric per iteration: `attempted`, `confirmed`, `downgraded`, `inconclusive`, `duration_ms` (plus `tokens_in/out` when available), forwarded to the tracker like all Phase 4 metrics. Timeout emits `triage=verify-by-test-timeout`.
34
+ One `review.verify_by_test` metric per iteration: `attempted`, `confirmed`, `downgraded`, `inconclusive`, `flaky` (findings whose repeat runs disagreed), `duration_ms` (plus `tokens_in/out` when available), forwarded to the tracker like all Phase 4 metrics. Timeout emits `triage=verify-by-test-timeout`.
34
35
 
35
36
  ## Off by default reason
36
37
 
37
- Adds one Sonnet call plus up to `maxFindings` single-test runs (and build-lock contention on Xcode projects) per review iteration that has accepted blockers. On clean runs it never fires, but on noisy-reviewer repos it can add minutes per iteration. Flip on for security-critical work, release branches, or repos where reviewer false-positive rate is high.
38
+ Adds one Sonnet call plus up to `maxFindings` single-test runs (and build-lock contention on Xcode projects) per review iteration that has accepted blockers. Each passing repro test is re-run `verifyByTest.repeatCount` times (default 3), so a downgrade costs up to `repeatCount` single-test runs rather than one. On clean runs it never fires, but on noisy-reviewer repos it can add minutes per iteration. Flip on for security-critical work, release branches, or repos where reviewer false-positive rate is high.
38
39
 
39
40
  ## Reference
40
41
 
41
- Wiring: `$HOME/.claude/multi-agent-refs/phases/phase-4-review.md` Step 3.7. Schema: `$HOME/.claude/schemas/triage-output.schema.json` v3.2.0 (`$defs.verification`). Evidence gate: `$HOME/.claude/scripts/evidence-gate.mjs`. Prefs: `prefs.global.verifyByTest` in `$HOME/.claude/schemas/prefs.schema.json`.
42
+ Wiring: `$HOME/.claude/multi-agent-refs/phases/phase-4-review.md` Step 3.7. Schema: `$HOME/.claude/schemas/triage-output.schema.json` v3.2.0 (`$defs.verification`). Evidence gate: `$HOME/.claude/scripts/evidence-gate.mjs`. Prefs: `prefs.global.verifyByTest` (including `repeatCount`) in `$HOME/.claude/schemas/prefs.schema.json`.
@@ -49,7 +49,7 @@ This is why the removal is safe:
49
49
 
50
50
  | Consumer | Reads | Without salvage |
51
51
  |---|---|---|
52
- | Phase 7 triage-memory ingest | `triage-output.json` | `[ -f ]`-guarded, so it degrades **silently**: the triage corpus and learnings ledger stop being fed and no error appears |
52
+ | Phase 7 triage-memory ingest | `triage-output.json` | `[ -f ]`-guarded, so it degrades **silently**: the triage corpus and learnings ledger stop being fed and no error appears. Since v16.20.0 Phase 4 writes `triage-output.json` itself (Step 3.2.1, the latest copy of `.pipeline/triage-round-<N>.json`), so the salvage is a second copy, not the only bridge |
53
53
  | Phase 7 learnings-ledger distill | same file | same silent degradation |
54
54
  | `render-work-summary.sh` | `agent-state.json`, `phase-tracker.json` (falls back to `logs/multi-agent/<task>/tracker-state.json`, which survives removal) | loses the salvaged copies but keeps the tracker via the logs fallback |
55
55
  | `:resume` | `agent-state.json` | cannot continue a Phase 7 pause |
@@ -75,13 +75,13 @@ toolkit. Continue without one rather than guessing which might fit.
75
75
 
76
76
  ## multi-agent-toolkit MCP
77
77
 
78
- | Family | Reach for it when |
78
+ | Tool | Reach for it when |
79
79
  |---|---|
80
- | `ui-inspect` | you need to know what is actually on screen, not what the code implies |
81
- | `crash-logs` | a crash happened on a device or simulator |
82
- | `design-check` | a built screen has to be compared against its design |
83
- | `ios-app-store-audit` | a package is heading for review |
84
- | `ios-testflight` | validating a build before upload |
80
+ | `ios_get_ui_tree` / `android_get_ui_tree` | you need to know what is actually on screen, not what the code implies |
81
+ | `ios_list_crashes` / `android_list_crashes` | a crash happened on a device or simulator |
82
+ | `design_*` (`design_visual_compare`, `design_ui_geometry`, `design_report`, ...) | a built screen has to be compared against its design |
83
+ | `ios_app_store_audit` | a package is heading for review |
84
+ | `ios_testflight_validate` | validating a build before upload |
85
85
 
86
86
  Registered at user scope by the installer, and preserved by uninstall - the tools are
87
87
  useful with no pipeline at all. If it is not registered the tools simply are not
@@ -10,7 +10,7 @@ description: "Canonical required-reading list for outward-facing payloads (PR bo
10
10
 
11
11
  | Read | Before | Governs |
12
12
  |---|---|---|
13
- | [`channels/pr.md`]($HOME/.claude/multi-agent-refs/channels/pr.md) | assembling the PR body | fixed section set (`summary` → `changes` → `architecture` cond. → `verification` → `dependencies` cond. → `related`), Markdown-only rule, reviewer-preserving Bitbucket PUT payload |
13
+ | [`channels/pr.md`]($HOME/.claude/multi-agent-refs/channels/pr.md) | assembling the PR body | fixed section set (`summary` → `changes` → `architecture` cond. → `verification` → `risk` cond. → `dependencies` cond. → `related`), Markdown-only rule, reviewer-preserving Bitbucket PUT payload |
14
14
  | [`phases/phase-6-commit.md`]($HOME/.claude/multi-agent-refs/phases/phase-6-commit.md) | committing | commit convention, default-reviewer fetch, draft/ready prompt, push-must-succeed loop |
15
15
  | [`rules.md`]($HOME/.claude/multi-agent-refs/rules.md) "External System Outputs" | any REST payload | real newlines, no HTML entities, no hand-rolled JSON, markup dialect per surface |
16
16
 
@@ -55,7 +55,7 @@ Phase 7: Report █░░░░░░░░░░░░░░░ 2s
55
55
  (emit from the last iteration's `triage.consensus` block, schema v3.1.0. One line stating `verdict` + `reviewerCount`; when `verdict` is `split` or `unverified`, list each `disagreements[]` entry so a human can confirm what the reviewers did not independently agree on. Omit the section when triage produced no consensus block.)
56
56
 
57
57
  ```
58
- Verdict: unverified (2 reviewers)
58
+ Verdict: unverified (3 reviewers)
59
59
  - Auth/KeychainStore.swift:40 - token persisted without access-control flag
60
60
  (both approved a keychain change - agreement unverified, confirm manually)
61
61
  ```
@@ -89,7 +89,7 @@ Phase 0: Init -> Phase 3: Dev (self-contained) -> Phase 4: Review -> Phase 5: Te
89
89
  | Phase 1 (Analysis) | Parallel Explore agents + analysis document | **SKIP** - tile flips to `skipped` at Step 7.5 | **SKIP** |
90
90
  | Phase 2 (Planning) | TaskCreate + architecture review + **Plan Approval Gate** | **SKIP** (no plan means no plan gate) | **SKIP** |
91
91
  | Phase 3 (Dev) | Follows the Phase 2 plan, TDD cycle (Sonnet) | **Self-contained** (Opus): agent scans relevant files, implements with TDD, builds | Same, on the local branch |
92
- | Phase 4 (Review) | Parallel review + Fable triage (Claude: 2-model / Copilot: 3-model) | **Same** - gates, parallel review, triage; blocking findings return to Phase 3 (cap 3) | **Same**, on the local branch diff |
92
+ | Phase 4 (Review) | Parallel review + Fable triage (3 reviewers on every host: Claude Code Fable + Opus + Sonnet, Copilot GPT-5.4 + Opus + Sonnet) | **Same** - gates, parallel review, triage; blocking findings return to Phase 3 (cap 3) | **Same**, on the local branch diff |
93
93
  | Phase 5 (User Test) | Interactive prompt | **Interactive prompt** | Not in the set - no worktree to check out |
94
94
  | Phase 6 (Commit) | Commit + PR | Same - still asks | Same |
95
95
  | Phase 7 (Report) | Full report + channels multi-select | Simplified - no analysis section, review section IS present, channels menu still pauses | Same |
@@ -142,7 +142,7 @@ Results are cached in `global.serviceStatus` (existing TTL contract, default 300
142
142
  **Figma MCP expired / rejected - CRITICAL path.** The `figma_mcp` OAuth token (`figu_`) expires on a schedule (90 days), unlike the PATs. Because a dead MCP token silently degrades every Figma consumer to Tier 2/3 and has cost rebuild rounds, an expired `figma_mcp` is never deferred to mid-run:
143
143
 
144
144
  1. **Silent renewal first**: run `~/.claude/lib/figma-mcp-refresh.sh` (refresh grant via `<key>_Refresh` + the `.figma-oauth.json` client credentials next to `prefs.global.tokenScripts.figma_mcp`). Exit 0 → re-probe, log `→ figma mcp token renewed silently`, continue. No question asked.
145
- 2. **Renewal impossible/rejected** (exit 1/2) → AskUserQuestion at init (`question`/`description` in `outputLanguage`): "Figma MCP token expired - update it now?"
145
+ 2. **Renewal impossible/rejected** (exit 1/2) → ask once at init through the native picker (`picker-contract.md`; `question`/`description` in `outputLanguage`): "Figma MCP credential expired - update it now?" The answer routes to a script or the clipboard Save Flow; the value itself is never typed in chat.
146
146
  - **Regenerate now (script)** - shown only when `prefs.global.tokenScripts.figma_mcp` is set: run that script (browser OAuth flow), then re-probe and continue.
147
147
  - **Save a new token** - Token Save Flow from `setup.md` (clipboard path).
148
148
  - **Continue degraded** - proceed on Tier 2 (REST PAT) for this run; log the downgrade.
@@ -269,7 +269,7 @@ Scan `$HOME` (maxdepth 2) for project markers (`.xcodeproj`, `Package.swift`, `b
269
269
  ```
270
270
  6. Sort: `develop*` first, then `release/*`, then `main`/`master`. Surface through the
271
271
  **native picker** per `picker-contract.md` (`AskUserQuestion` on Claude Code,
272
- `ask_choice.sh` on Copilot CLI) - `question` + `description` in `outputLanguage`,
272
+ `ask-choice.sh` on Copilot CLI) - `question` + `description` in `outputLanguage`,
273
273
  `label` = the branch name verbatim (a proper noun, never translated) with the recent
274
274
  branch first and marked `(Recommended)`. The ASCII sketch
275
275
  below is what the options carry, not a menu to print:
@@ -425,6 +425,8 @@ git -C $PROJECT_ROOT config user.email "{identity.email}"
425
425
 
426
426
  **If normal mode** (worktree - default): 2. Worktree path: Jira → `.worktrees/{jiraId}/`, GitHub → `.worktrees/GH{issueNo}/`, free-text → `.worktrees/task-{shortId}/` 3. **Heal stale admin state first** (see "Worktree stale-lock heal" below) and **apply the residue guard** (see "Worktree residue guard" below), then `git -C $PROJECT_ROOT worktree add {path} -b {branch} origin/{baseBranch}` (if exists: enter, pull) 4. Set identity: `git -C {worktree-path} config user.name/email` 5. Create log dir + `agent-log.md` + `agent-state.json` at `$HOME/.claude/logs/multi-agent/{project}/{task-id}/`, never inside the worktree:
427
427
 
428
+ **Worktree location convention (cited by every other command):** always `{projectRoot}/.worktrees/{taskId}`, inside the repo, never under `$HOME`. `{taskId}` is the directory name from the rule above (`DC-<shortId>` for `/multi-agent:design-check`). `.worktrees` is fixed, not a preference: no `worktreeBasePath` key exists, and `gc-worktrees.sh`, `purge.sh`, the cost renderers and `usage-report.mjs` resolve `<repo>/.worktrees/` by name. Multi-repo tasks get one worktree per repo (the loop below); `--local` creates none and `worktreePath` is `$PROJECT_ROOT`.
429
+
428
430
  **Worktree stale-lock heal (required before every `worktree add`):** a run killed mid-`worktree add` (OOM, SIGTERM, disk full) leaves a locked or broken admin entry under `.git/worktrees/{id}/`, so the retry fails with `fatal: '<path>' already exists`. Always run the heal first - it is a no-op on a clean repo:
429
431
 
430
432
  ```bash
@@ -1,6 +1,6 @@
1
- ### Phase 1: Analysis (Fable)
1
+ ### Phase 1: Analysis (Sonnet)
2
2
 
3
- > **TLDR** - Fable-driven codebase exploration (Opus when the fallback ladder engages). Detects if the issue is already fixed (git blame, closed PRs), then launches parallel Explore sub-agents to map the affected code paths. Outputs: impact analysis, stack detection (auto-selects platform guide), relevant files, risk areas. Feeds Phase 2 planning.
3
+ > **TLDR** - Sonnet-driven codebase exploration: the `explorer` persona declares `preferredModel: sonnet` (haiku when the fallback ladder engages). Detects if the issue is already fixed (git blame, closed PRs), then launches parallel Explore sub-agents to map the affected code paths. Outputs: impact analysis, stack detection (auto-selects platform guide), relevant files, risk areas. Feeds Phase 2 planning.
4
4
 
5
5
  <!-- progress-contract: applied -->
6
6
  Progress emission per `$HOME/.claude/multi-agent-refs/progress-contract.md` - lines for each Explore dispatch, each finish, analyst synthesis start, `analysis.json` write.
@@ -217,7 +217,7 @@ Forward the explorer call's token totals into the tracker so Phase 7's Cost Brea
217
217
 
218
218
  ```bash
219
219
  LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 1 analysis.completed \
220
- model=opus tokens_in=$IN tokens_out=$OUT duration_ms=$DUR
220
+ model=sonnet tokens_in=$IN tokens_out=$OUT duration_ms=$DUR
221
221
  ```
222
222
 
223
223
  Best-effort. See `$HOME/.claude/multi-agent-refs/progress-contract.md#token-telemetry-forwarding` for the canonical contract.
@@ -215,7 +215,7 @@ Otherwise (normal mode), run the gate. The gate has **two modes** that chain: Cl
215
215
 
216
216
  ##### 5a - Clarification Mode (conditional, max 2 rounds)
217
217
 
218
- Trigger if the plan Opus produced in Step 1-4 carries ANY ambiguity signal from Phase 1 analysis:
218
+ Trigger if the plan Fable produced in Step 1-4 carries ANY ambiguity signal from Phase 1 analysis:
219
219
 
220
220
  | Signal | Check |
221
221
  |---|---|
@@ -300,12 +300,12 @@ Handle the selection:
300
300
  - Set `state.status = "paused"`, log `🧠 Phase 2: Plan aborted by user`, stop
301
301
  - **Other (free-text edit request)**:
302
302
  - Treat the typed text as an edit request. Append to `state.phases["2"].planEditRequests`, bump `planIterations`
303
- - Pass the edit request + current plan to Opus; Opus revises and returns a new plan (same schema, same validator)
303
+ - Pass the edit request + current plan to the planning model (Fable; Opus when the fallback ladder engages); it revises and returns a new plan (same schema, same validator)
304
304
  - Re-render the plan (5b), loop
305
305
 
306
306
  No hard cap on edit iterations - the user controls exit via the Approve / Cancel options. Between iterations, keep only the **latest plan** as canonical; previous renders are in the log for audit but do not re-enter the validator.
307
307
 
308
- **Validator**: every revised plan goes through `node $HOME/.claude/scripts/validate-planning.mjs -` before re-render. If validation fails after an edit, log `⚠️ Phase 2: Plan validator failed after edit request #N - retrying Opus once` and retry once; on second failure surface the validator error to the user and go back to approval prompt with the pre-edit plan.
308
+ **Validator**: every revised plan goes through `node $HOME/.claude/scripts/validate-planning.mjs -` before re-render. If validation fails after an edit, log `⚠️ Phase 2: Plan validator failed after edit request #N - retrying the planning model once` and retry once; on second failure surface the validator error to the user and go back to approval prompt with the pre-edit plan.
309
309
 
310
310
  #### Step 6 - Mode-specific short-circuit (reference)
311
311
 
@@ -321,11 +321,11 @@ The pipeline shapes interact with the gate as follows. This table is the source
321
321
 
322
322
  #### Telemetry - token forwarding
323
323
 
324
- After plan generation (and after each edit-loop iteration), forward Opus call totals so Phase 7's Cost Breakdown captures Phase 2:
324
+ After plan generation (and after each edit-loop iteration), forward the planning model's call totals so Phase 7's Cost Breakdown captures Phase 2 (`model=` names the rung that actually ran: `fable`, or `opus` after a fallback step):
325
325
 
326
326
  ```bash
327
327
  LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 2 plan.generated \
328
- model=opus tokens_in=$IN tokens_out=$OUT duration_ms=$DUR iteration=$N
328
+ model=fable tokens_in=$IN tokens_out=$OUT duration_ms=$DUR iteration=$N
329
329
  ```
330
330
 
331
331
  Best-effort. See `$HOME/.claude/multi-agent-refs/progress-contract.md#token-telemetry-forwarding`.
@@ -84,7 +84,8 @@ When re-entering from Phase 4:
84
84
  3. `accepted.suggestion` items are applied opportunistically (no TDD loop required) unless the user asked for suggestions to be treated strictly.
85
85
  4. `deferred` items are NOT actioned in this re-entry - they surface in Phase 7's "Follow-up items" section.
86
86
  5. `rejected` items are never touched. Log their IDs + triage reasons for audit only.
87
- 6. After rework, increment `state.phases["3"].retryCount`; hard-kill at `retryCount === 3` and escalate to the user. Do not loop indefinitely.
87
+ 6. After rework, increment `state.phases["3"].retryCount`; hard-kill at `retryCount === 3` and escalate to the user. Do not loop indefinitely. When `prefs.global.autopilotCircuitBreaker.enabled` is true (default), record the hard-kill before escalating as `state.circuitBreaker = {tripped: true, trigger: 3, detail: "rework cycles exhausted", checkpoint: {phase: 3, step: "re-entry", iteration: <n>}, trippedAt, counters: {reworkCycles: <n>}}`: the rework-storm trigger; `autopilotCircuitBreaker.maxReworkCycles` (default 3) is bounded above by this cap.
88
+ 7. When `state.reviewIterations[-1].delta` exists, `stillPresent[]` findings come first in the task list, quoted `STILL PRESENT after round N-1's fix`; repeating the previous attempt is the loop the circuit-breaker stops.
88
89
 
89
90
  If the latest iteration has `triage.approved === true` AND `accepted === []`, Phase 3 was entered by mistake - log the anomaly and return to Phase 5.
90
91
 
@@ -155,6 +156,7 @@ For each task (respecting dependency order):
155
156
  **REFACTOR (if needed):**
156
157
  - Only if duplication or naming is poor
157
158
  - Re-run tests after refactor → still GREEN
159
+ - **Stability (required):** run every test added or changed in this diff `prefs.global.testStability.repeatCount` times (default 3; 1 disables) with the single-test invocation above. Disagreeing outcomes are not GREEN: log `test.flake_signal file=<f> passed=<k> of=<N>` and fix the test or the code first. A pass only on retry is a flake signal, not a pass. Same rule on the Phase 4 rework re-entry.
158
160
 
159
161
  **Target resolution** (auto-detect once per project, cache in `agent-state.json`; ios resolves scheme + simulator, android resolves module + variant, backend/web need none):
160
162
  ```bash
@@ -243,7 +245,7 @@ After the build/test green step and BEFORE Phase 4 handoff, run one diff-shrink
243
245
  - **Over-abstraction** - single-call-site protocols/wrappers/helpers from this diff
244
246
 
245
247
  Progress line: ` → dispatching code-simplifier diff-shrink`
246
- 2. **Return contract** - a list of shrink edits `{ "file", "lines", "edit", "rationale", "risk": "safe|unsafe" }`; the subagent never writes files.
248
+ 2. **Return contract** - a list of shrink edits `{ "file", "lines", "edit", "rationale", "risk": "safe|unsafe" }`; the subagent never writes files. Keep the `rationale` strings: Step 3.7 writes them into `scope-check.json`.
247
249
  3. **Apply safe edits only** - skip `unsafe` (logged). Zero edits is normal - log and continue. Progress line: ` → applying shrink edits ({N} applied, {M} skipped)`
248
250
  4. **Re-run build + tests** (same build-queue lock + evidence-gate rule as Step 4). Any breakage -> revert shrink edits wholesale and proceed pre-shrink; the simplifier must never cost a green state.
249
251
  5. **Record tokens in the cost ledger** so Phase 7's Cost Breakdown captures the pass:
@@ -258,6 +260,19 @@ Scope guard: a single pass, never looped. Runs in a Short run too; component tas
258
260
 
259
261
  ---
260
262
 
263
+ #### Step 3.7 - Scope self-check (required handoff artifact)
264
+
265
+ Phase 4 cannot reconstruct why each file was touched or what was left out on purpose, so Dev states both before the handoff: write `$WORKTREE/.pipeline/scope-check.json` (`schemas/scope-check.schema.json`: one `{path, reason}` per file in the diff, `notDone[]` with `why` in `out-of-scope | follow-up | rejected-abstraction`, and the Step 3.6 `simplifier` rationales), then run the gate. **Record rules, consumers and the full contract: `$HOME/.claude/multi-agent-refs/features/scope-check.md`.**
266
+
267
+ ```bash
268
+ git -C "$WORKTREE" diff --name-only "origin/$BASE_BRANCH"...HEAD \
269
+ | node $HOME/.claude/scripts/scope-check-gate.mjs --check "$WORKTREE/.pipeline/scope-check.json" --diff-files -
270
+ ```
271
+
272
+ Exit 1 lists `unjustified[]`: complete the record once and re-run. A second exit 1 does not block; log `dev.scope_check=incomplete unjustified=<n>` and hand off, and Phase 4 shows the gate output to reviewers inside `<scope-self-check>`. Progress line: ` → checking scope self-check ({n} files, {m} not done)`.
273
+
274
+ ---
275
+
261
276
  #### Short pipeline (`state.onlyDevelop === true`)
262
277
 
263
278
  Set by the Phase 0 Step 7.5 depth picker, or by autopilot never (autopilot always runs Full). When it is true, Phase 3 runs self-contained with **Opus** (not Sonnet). No Phase 2 plan exists - the agent creates its own scope.