@mmerterden/multi-agent-pipeline 17.6.0 → 19.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (272) hide show
  1. package/CHANGELOG.md +310 -0
  2. package/README.md +76 -18
  3. package/README.tr.md +55 -16
  4. package/docs/adr/0002-instruction-driven-flag.md +1 -0
  5. package/docs/adr/0005-lazy-phase-docs.md +11 -1
  6. package/docs/adr/0008-installer-modularization-and-secret-leak-defense.md +1 -0
  7. package/docs/adr/0010-own-code-graph.md +1 -0
  8. package/docs/adr/0011-dormant-ci.md +25 -1
  9. package/docs/adr/0014-six-phase-consolidation.md +134 -0
  10. package/docs/adr/README.md +2 -1
  11. package/docs/architecture.md +37 -38
  12. package/docs/best-practices.md +1 -1
  13. package/docs/ecosystem.md +37 -26
  14. package/docs/engineering.md +1 -1
  15. package/docs/facts.json +45 -0
  16. package/docs/features.md +54 -53
  17. package/docs/performance.md +5 -5
  18. package/docs/recovery-guide.md +9 -9
  19. package/docs/server-readiness.md +188 -0
  20. package/docs/token-budget-history.md +3 -1
  21. package/index.js +18 -3
  22. package/install/_codex-agents.mjs +1 -1
  23. package/install/_common.mjs +42 -17
  24. package/install/_dev-only-files.mjs +8 -0
  25. package/install/_unattended-profile.mjs +113 -0
  26. package/install/index.mjs +48 -0
  27. package/install/templates/claude-hooks.json +1 -1
  28. package/install/templates/codex-instructions.md +1 -1
  29. package/install/templates/copilot-instructions.md +28 -28
  30. package/manifest.json +1065 -0
  31. package/package.json +6 -3
  32. package/pipeline/agents/dev-critic.md +3 -3
  33. package/pipeline/commands/figma-to-swiftui.md +1 -1
  34. package/pipeline/commands/multi-agent/SKILL.md +8 -8
  35. package/pipeline/commands/multi-agent/analysis/SKILL.md +9 -9
  36. package/pipeline/commands/multi-agent/autopilot/SKILL.md +7 -7
  37. package/pipeline/commands/multi-agent/channels/SKILL.md +15 -15
  38. package/pipeline/commands/multi-agent/diff-explain/SKILL.md +6 -6
  39. package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +1 -1
  40. package/pipeline/commands/multi-agent/graph/SKILL.md +1 -1
  41. package/pipeline/commands/multi-agent/help/SKILL.md +62 -62
  42. package/pipeline/commands/multi-agent/language/SKILL.md +2 -2
  43. package/pipeline/commands/multi-agent/local/SKILL.md +11 -11
  44. package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +13 -13
  45. package/pipeline/commands/multi-agent/log/SKILL.md +2 -2
  46. package/pipeline/commands/multi-agent/manual-test/SKILL.md +9 -9
  47. package/pipeline/commands/multi-agent/model/SKILL.md +69 -0
  48. package/pipeline/commands/multi-agent/refactor/SKILL.md +3 -3
  49. package/pipeline/commands/multi-agent/resume/SKILL.md +4 -4
  50. package/pipeline/commands/multi-agent/resume-local/SKILL.md +19 -17
  51. package/pipeline/commands/multi-agent/review/SKILL.md +1 -1
  52. package/pipeline/commands/multi-agent/route-off/SKILL.md +36 -0
  53. package/pipeline/commands/multi-agent/route-on/SKILL.md +74 -0
  54. package/pipeline/commands/multi-agent/route-status/SKILL.md +56 -0
  55. package/pipeline/commands/multi-agent/setup/SKILL.md +2 -2
  56. package/pipeline/commands/multi-agent/status/SKILL.md +54 -23
  57. package/pipeline/commands/multi-agent/steer/SKILL.md +2 -2
  58. package/pipeline/commands/multi-agent/sync/SKILL.md +12 -13
  59. package/pipeline/commands/multi-agent/test/SKILL.md +1 -1
  60. package/pipeline/lib/_jira-auth.sh +8 -0
  61. package/pipeline/lib/analysis-jira-write.sh +32 -0
  62. package/pipeline/lib/ask-choice.sh +13 -2
  63. package/pipeline/lib/autopilot-state.sh +8 -0
  64. package/pipeline/lib/credential-inventory.sh +1 -1
  65. package/pipeline/lib/fatal.mjs +129 -0
  66. package/pipeline/lib/fetch-fortify.sh +1 -1
  67. package/pipeline/lib/figma-mcp-refresh.sh +18 -0
  68. package/pipeline/lib/figma-screenshot.sh +18 -0
  69. package/pipeline/lib/invoked-directly.mjs +43 -0
  70. package/pipeline/lib/jira-publish.sh +42 -0
  71. package/pipeline/lib/md2confluence-v3.py +47 -0
  72. package/pipeline/lib/model-rung.sh +142 -0
  73. package/pipeline/lib/outbound-gate.mjs +175 -0
  74. package/pipeline/lib/phase-schema.mjs +88 -0
  75. package/pipeline/lib/plan-todos.sh +32 -11
  76. package/pipeline/lib/post-pr-review.sh +77 -8
  77. package/pipeline/lib/repo-hygiene.sh +8 -3
  78. package/pipeline/lib/require-jq.sh +40 -0
  79. package/pipeline/lib/route-state.sh +161 -0
  80. package/pipeline/lib/run-paths.sh +335 -0
  81. package/pipeline/multi-agent-refs/_account-picker.md +1 -1
  82. package/pipeline/multi-agent-refs/_dev-context.md +1 -1
  83. package/pipeline/multi-agent-refs/_input-parser.md +1 -1
  84. package/pipeline/multi-agent-refs/analysis/evidence.md +0 -9
  85. package/pipeline/multi-agent-refs/analysis/intake.md +1 -1
  86. package/pipeline/multi-agent-refs/analysis/locked.md +21 -22
  87. package/pipeline/multi-agent-refs/analysis/render.md +1 -1
  88. package/pipeline/multi-agent-refs/analysis/synthesis.md +12 -6
  89. package/pipeline/multi-agent-refs/android-guide.md +1 -1
  90. package/pipeline/multi-agent-refs/audit-guide.md +13 -13
  91. package/pipeline/multi-agent-refs/channels/issue-comment.md +2 -2
  92. package/pipeline/multi-agent-refs/channels/jira.md +3 -3
  93. package/pipeline/multi-agent-refs/channels/pr.md +4 -4
  94. package/pipeline/multi-agent-refs/channels/wiki.md +1 -1
  95. package/pipeline/multi-agent-refs/component-dispatch.md +3 -3
  96. package/pipeline/multi-agent-refs/cross-cli-contract.md +31 -6
  97. package/pipeline/multi-agent-refs/features/autopilot-circuit-breaker.md +74 -4
  98. package/pipeline/multi-agent-refs/features/code-graph.md +5 -5
  99. package/pipeline/multi-agent-refs/features/cost-analysis.md +93 -0
  100. package/pipeline/multi-agent-refs/features/design-conformance.md +1 -1
  101. package/pipeline/multi-agent-refs/features/dev-critic.md +3 -3
  102. package/pipeline/multi-agent-refs/features/doctor.md +47 -2
  103. package/pipeline/multi-agent-refs/features/external-context-injection.md +3 -3
  104. package/pipeline/multi-agent-refs/features/maturity-followup.md +3 -3
  105. package/pipeline/multi-agent-refs/features/model-fallback.md +5 -5
  106. package/pipeline/multi-agent-refs/features/plan-todos.md +1 -1
  107. package/pipeline/multi-agent-refs/features/repo-map.md +1 -1
  108. package/pipeline/multi-agent-refs/features/review-delta.md +3 -3
  109. package/pipeline/multi-agent-refs/features/review-multi-repo.md +1 -1
  110. package/pipeline/multi-agent-refs/features/scope-check.md +4 -4
  111. package/pipeline/multi-agent-refs/features/skill-conformance.md +2 -2
  112. package/pipeline/multi-agent-refs/features/stack-skill-routing.md +1 -1
  113. package/pipeline/multi-agent-refs/features/verify-by-test.md +4 -4
  114. package/pipeline/multi-agent-refs/features/verify.md +83 -0
  115. package/pipeline/multi-agent-refs/features/visual-evidence.md +19 -19
  116. package/pipeline/multi-agent-refs/features/worktree-finalize.md +6 -6
  117. package/pipeline/multi-agent-refs/issue-jira-triad.md +10 -10
  118. package/pipeline/multi-agent-refs/knowledge.md +11 -11
  119. package/pipeline/multi-agent-refs/multi-repo-integration-build.md +13 -13
  120. package/pipeline/multi-agent-refs/payload-contracts.md +8 -8
  121. package/pipeline/multi-agent-refs/phases/log-format.md +10 -10
  122. package/pipeline/multi-agent-refs/phases/modes.md +30 -30
  123. package/pipeline/multi-agent-refs/phases/operations.md +21 -10
  124. package/pipeline/multi-agent-refs/phases/phase-0-init.md +25 -25
  125. package/pipeline/multi-agent-refs/phases/phase-1-plan.md +599 -0
  126. package/pipeline/multi-agent-refs/phases/{phase-3-dev.md → phase-2-dev.md} +129 -49
  127. package/pipeline/multi-agent-refs/phases/{phase-4-review.md → phase-3-review.md} +225 -107
  128. package/pipeline/multi-agent-refs/phases/{phase-6-commit.md → phase-4-commit.md} +23 -23
  129. package/pipeline/multi-agent-refs/phases/{phase-7-report.md → phase-5-report.md} +29 -29
  130. package/pipeline/multi-agent-refs/phases.md +44 -48
  131. package/pipeline/multi-agent-refs/picker-contract.md +1 -1
  132. package/pipeline/multi-agent-refs/progress-contract.md +6 -6
  133. package/pipeline/multi-agent-refs/readiness-review.md +1 -1
  134. package/pipeline/multi-agent-refs/rules.md +7 -7
  135. package/pipeline/multi-agent-refs/swiftui-guide.md +2 -2
  136. package/pipeline/multi-agent-refs/tracker-contract.md +31 -32
  137. package/pipeline/multi-agent-refs/unattended-contract.md +129 -0
  138. package/pipeline/multi-agent-refs/wiki-capture.md +14 -14
  139. package/pipeline/preferences-template.json +9 -1
  140. package/pipeline/rules/outside-the-pipeline.md +1 -1
  141. package/pipeline/schemas/agent-state.schema.json +50 -50
  142. package/pipeline/schemas/analysis-output.schema.json +2 -2
  143. package/pipeline/schemas/autopilot-config.schema.json +1 -1
  144. package/pipeline/schemas/code-graph.schema.json +1 -1
  145. package/pipeline/schemas/criteria-manifest.schema.json +1 -1
  146. package/pipeline/schemas/dev-critic-output.schema.json +1 -1
  147. package/pipeline/schemas/diff-risk.schema.json +1 -1
  148. package/pipeline/schemas/migrations/prefs-2.4.0-to-2.5.0.mjs +2 -2
  149. package/pipeline/schemas/migrations/prefs-2.6.0-to-2.7.0.mjs +31 -0
  150. package/pipeline/schemas/migrations/state-2.1.0-to-2.2.0.mjs +129 -0
  151. package/pipeline/schemas/phases.json +105 -0
  152. package/pipeline/schemas/plan-todos.schema.json +5 -5
  153. package/pipeline/schemas/planning-output.schema.json +1 -1
  154. package/pipeline/schemas/prefs.schema.json +100 -56
  155. package/pipeline/schemas/reviewer-output.schema.json +3 -3
  156. package/pipeline/schemas/route-config.schema.json +74 -0
  157. package/pipeline/schemas/scope-check.schema.json +1 -1
  158. package/pipeline/schemas/test-gap.schema.json +1 -1
  159. package/pipeline/schemas/token-budget.json +12 -18
  160. package/pipeline/schemas/triage-output.schema.json +6 -6
  161. package/pipeline/scripts/README.md +3 -3
  162. package/pipeline/scripts/_code-graph.mjs +2 -2
  163. package/pipeline/scripts/_run-paths.mjs +372 -0
  164. package/pipeline/scripts/_smoke-root.sh +1 -1
  165. package/pipeline/scripts/aggregate-metrics.mjs +65 -65
  166. package/pipeline/scripts/autopilot-arming.mjs +2 -1
  167. package/pipeline/scripts/autopilot-intake.mjs +2 -1
  168. package/pipeline/scripts/autopilot-runner.mjs +206 -2
  169. package/pipeline/scripts/build-references.mjs +2 -1
  170. package/pipeline/scripts/build-stack-plugins.mjs +10 -2
  171. package/pipeline/scripts/capture-evidence.sh +7 -2
  172. package/pipeline/scripts/capture-flush.sh +8 -8
  173. package/pipeline/scripts/capture-resume.sh +3 -3
  174. package/pipeline/scripts/classify-plan-safety.mjs +3 -2
  175. package/pipeline/scripts/cost-analyze.mjs +600 -0
  176. package/pipeline/scripts/cost-budget-check.mjs +4 -12
  177. package/pipeline/scripts/council-view.mjs +2 -1
  178. package/pipeline/scripts/crush-json.mjs +2 -1
  179. package/pipeline/scripts/diff-explain.mjs +7 -10
  180. package/pipeline/scripts/diff-risk-score.mjs +2 -1
  181. package/pipeline/scripts/doctor.mjs +140 -6
  182. package/pipeline/scripts/evidence-gate.mjs +9 -3
  183. package/pipeline/scripts/feedback-send.mjs +12 -2
  184. package/pipeline/scripts/gc-abandoned.sh +32 -16
  185. package/pipeline/scripts/gc-tmp.sh +1 -1
  186. package/pipeline/scripts/gc-worktrees.sh +12 -5
  187. package/pipeline/scripts/gen-facts.mjs +175 -0
  188. package/pipeline/scripts/gen-mode-dispatch.mjs +32 -37
  189. package/pipeline/scripts/gen-ref-toc.mjs +1 -1
  190. package/pipeline/scripts/github-ssh-setup.sh +64 -7
  191. package/pipeline/scripts/graph-mermaid.mjs +4 -2
  192. package/pipeline/scripts/graph-report.mjs +1 -1
  193. package/pipeline/scripts/jira-attach.sh +1 -1
  194. package/pipeline/scripts/keychain-save.sh +101 -30
  195. package/pipeline/scripts/learn-from-transcripts.mjs +3 -2
  196. package/pipeline/scripts/learning-curve.mjs +36 -31
  197. package/pipeline/scripts/log-metric.sh +17 -4
  198. package/pipeline/scripts/make-manifest.mjs +199 -0
  199. package/pipeline/scripts/memory-save.sh +1 -1
  200. package/pipeline/scripts/migrate-prefs.mjs +24 -6
  201. package/pipeline/scripts/migrate-state.mjs +94 -4
  202. package/pipeline/scripts/phase-banner.sh +26 -22
  203. package/pipeline/scripts/phase-tracker.sh +48 -10
  204. package/pipeline/scripts/plan-coverage-gate.mjs +8 -4
  205. package/pipeline/scripts/pre-commit-check.sh +7 -0
  206. package/pipeline/scripts/pre-push-check.sh +7 -0
  207. package/pipeline/scripts/purge.sh +23 -6
  208. package/pipeline/scripts/render-agent-log-cost.sh +10 -3
  209. package/pipeline/scripts/render-cost-summary.sh +9 -2
  210. package/pipeline/scripts/render-work-summary.sh +14 -7
  211. package/pipeline/scripts/review-file-filter.mjs +5 -3
  212. package/pipeline/scripts/review-scope.mjs +2 -1
  213. package/pipeline/scripts/routine-registry.mjs +2 -1
  214. package/pipeline/scripts/run-aggregator.mjs +26 -20
  215. package/pipeline/scripts/run-metrics.mjs +4 -2
  216. package/pipeline/scripts/runs-index.mjs +353 -0
  217. package/pipeline/scripts/scorecard-snapshot.mjs +178 -0
  218. package/pipeline/scripts/search-logs.sh +18 -0
  219. package/pipeline/scripts/smoke-cross-cli-behavior.sh +6 -6
  220. package/pipeline/scripts/smoke-schema-validation.sh +26 -7
  221. package/pipeline/scripts/test-gap-scan.mjs +2 -1
  222. package/pipeline/scripts/test-integrity-gate.mjs +2 -1
  223. package/pipeline/scripts/token-budget-report.mjs +13 -2
  224. package/pipeline/scripts/triage-memory.mjs +2 -2
  225. package/pipeline/scripts/update-issue-progress.sh +56 -7
  226. package/pipeline/scripts/usage-report.mjs +12 -1
  227. package/pipeline/scripts/validate-analysis-doc.mjs +75 -18
  228. package/pipeline/scripts/validate-code-graph.mjs +6 -3
  229. package/pipeline/scripts/validate-complaint-doc.mjs +2 -1
  230. package/pipeline/scripts/validate-diff-risk.mjs +6 -3
  231. package/pipeline/scripts/validate-planning.mjs +1 -1
  232. package/pipeline/scripts/validate-reviewer.mjs +1 -1
  233. package/pipeline/scripts/validate-state.mjs +45 -5
  234. package/pipeline/scripts/validate-test-gap.mjs +6 -3
  235. package/pipeline/scripts/validate-triage.mjs +6 -4
  236. package/pipeline/scripts/verify-citations.mjs +4 -2
  237. package/pipeline/scripts/verify.mjs +327 -0
  238. package/pipeline/scripts/worktree-finalize.sh +18 -9
  239. package/pipeline/scripts/write-state.mjs +154 -15
  240. package/pipeline/skills/.skill-manifest.json +37 -21
  241. package/pipeline/skills/.skills-index.json +104 -5
  242. package/pipeline/skills/shared/README.md +15 -6
  243. package/pipeline/skills/shared/core/apple-archive-compliance/SKILL.md +2 -2
  244. package/pipeline/skills/shared/core/google-play-compliance/SKILL.md +2 -2
  245. package/pipeline/skills/shared/core/multi-agent/SKILL.md +69 -71
  246. package/pipeline/skills/shared/core/multi-agent-autopilot/SKILL.md +3 -3
  247. package/pipeline/skills/shared/core/multi-agent-channels/SKILL.md +14 -14
  248. package/pipeline/skills/shared/core/multi-agent-diff-explain/SKILL.md +5 -5
  249. package/pipeline/skills/shared/core/multi-agent-graph/SKILL.md +1 -1
  250. package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +25 -23
  251. package/pipeline/skills/shared/core/multi-agent-language/SKILL.md +2 -2
  252. package/pipeline/skills/shared/core/multi-agent-local/SKILL.md +2 -2
  253. package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +8 -8
  254. package/pipeline/skills/shared/core/multi-agent-manual-test/SKILL.md +6 -6
  255. package/pipeline/skills/shared/core/multi-agent-model/SKILL.md +71 -0
  256. package/pipeline/skills/shared/core/multi-agent-refactor/SKILL.md +3 -3
  257. package/pipeline/skills/shared/core/multi-agent-resume/SKILL.md +1 -1
  258. package/pipeline/skills/shared/core/multi-agent-resume-local/SKILL.md +7 -7
  259. package/pipeline/skills/shared/core/multi-agent-route-off/SKILL.md +39 -0
  260. package/pipeline/skills/shared/core/multi-agent-route-on/SKILL.md +76 -0
  261. package/pipeline/skills/shared/core/multi-agent-route-status/SKILL.md +59 -0
  262. package/pipeline/skills/shared/core/multi-agent-setup/SKILL.md +1 -1
  263. package/pipeline/skills/shared/core/multi-agent-status/SKILL.md +35 -11
  264. package/pipeline/skills/shared/core/multi-agent-steer/SKILL.md +2 -2
  265. package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +6 -5
  266. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/package_app.sh +4 -1
  267. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/setup_dev_signing.sh +4 -1
  268. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/sign-and-notarize.sh +2 -1
  269. package/pipeline/skills/skills-index.md +13 -4
  270. package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +0 -263
  271. package/pipeline/multi-agent-refs/phases/phase-2-planning.md +0 -344
  272. package/pipeline/multi-agent-refs/phases/phase-5-test.md +0 -182
@@ -54,7 +54,7 @@ ticket that referenced nothing. The run then produced a plan from a partial pict
54
54
  reported success, and the user found out by reading the output.
55
55
 
56
56
  So on exit `3`, classify the stderr and surface an `AskUserQuestion` before continuing.
57
- The classification is the same shape the Phase 6 remote gate uses, because the failures
57
+ The classification is the same shape the Phase 4 remote gate uses, because the failures
58
58
  are the same failures:
59
59
 
60
60
  | stderr signal | Diagnosis | Question offers |
@@ -78,14 +78,14 @@ Rules that make this actionable rather than decorative:
78
78
  "planned without Crashlytics, user's choice at <ts>" in the analysis doc's limitations
79
79
  section instead of leaving a silent hole.
80
80
  5. **Autopilot still reports.** Autopilot takes "continue without this source" without
81
- asking, but the skipped source is logged and carried into the Phase 7 report. An
81
+ asking, but the skipped source is logged and carried into the Phase 5 report. An
82
82
  autopilot run that quietly planned from a partial picture is the same defect with a
83
83
  flag on it.
84
84
 
85
85
  Failures remain non-fatal at Phase 1 by default - the agent still runs analysis. What
86
86
  changed is that the user learns about it while the choice is still theirs to make.
87
87
 
88
- **Graylog specifics.** `fetch-graylog.sh` degrades to empty on any network/VPN failure: it emits a normalized empty object and exits `0`, so an unreachable host reads as `{status:"skipped", reason:"vpn-unreachable"}` and never blocks. **Exiting `0` is not permission to stay quiet.** The payload carries `degraded: true` and a `degradeReason`; when it does, log one line naming the host and the reason, and carry it into the Phase 7 report. Logs are advisory, so this never asks a question and never blocks - but "I searched the logs and found nothing" and "I never reached the log server" are different statements, and only one of them is true. Only a genuine auth rejection on a reachable host exits `3` (marked `failed`, still non-fatal here). `state.graylogContext` is prepended to the analysis prompt inside the **Referenced External Sources** section as diagnostic context that is **advisory only** - the agent may use it to orient on a reported error, but code remains ground truth, and there is no Phase 4 review gate for logs.
88
+ **Graylog specifics.** `fetch-graylog.sh` degrades to empty on any network/VPN failure: it emits a normalized empty object and exits `0`, so an unreachable host reads as `{status:"skipped", reason:"vpn-unreachable"}` and never blocks. **Exiting `0` is not permission to stay quiet.** The payload carries `degraded: true` and a `degradeReason`; when it does, log one line naming the host and the reason, and carry it into the Phase 5 report. Logs are advisory, so this never asks a question and never blocks - but "I searched the logs and found nothing" and "I never reached the log server" are different statements, and only one of them is true. Only a genuine auth rejection on a reachable host exits `3` (marked `failed`, still non-fatal here). `state.graylogContext` is prepended to the analysis prompt inside the **Referenced External Sources** section as diagnostic context that is **advisory only** - the agent may use it to orient on a reported error, but code remains ground truth, and there is no Phase 3 review gate for logs.
89
89
 
90
90
  ## Prompt injection shape
91
91
 
@@ -128,15 +128,15 @@ maturity step has `currentPhase: 0`, so resuming would start at Phase 1 and skip
128
128
  the check - the halt would be permanent in the one direction that matters.
129
129
 
130
130
  So resume reads `state.waitingFor` first: when it names a step, the run re-enters
131
- THAT step rather than the next phase. `waitingFor` already existed and Phase 7's
131
+ THAT step rather than the next phase. `waitingFor` already existed and Phase 5's
132
132
  channels pause already documented itself as resumable through it
133
- (`phases/phase-7-report.md`), while `resume/SKILL.md` never mentioned the field -
133
+ (`phases/phase-5-report.md`), while `resume/SKILL.md` never mentioned the field -
134
134
  so that pause had the same gap and this fixes both.
135
135
 
136
136
  | `waitingFor` | Re-entry |
137
137
  |---|---|
138
138
  | `maturity` | Phase 0, the maturity step, with the item re-fetched |
139
- | `user-channels-choice` | Phase 7, the channels menu |
139
+ | `user-channels-choice` | Phase 5, the channels menu |
140
140
  | absent | `currentPhase + 1`, as before |
141
141
 
142
142
  `waitingFor` is cleared by the write that records the answer. A field that
@@ -42,7 +42,7 @@ architect/reviewer roles with no file edits.
42
42
  Personas and dispatch read the **rung** (`fable`, `opus`, `sonnet`, `haiku`). The wire
43
43
  model ID each rung resolves to lives in exactly two places: `scripts/cost-table.json`
44
44
  (`prices.<rung>.modelId`) and the per-host reviewer table in
45
- `phases/phase-4-review.md`. A generation move edits those; it never renames a rung and
45
+ `phases/phase-3-review.md`. A generation move edits those; it never renames a rung and
46
46
  never touches a persona file.
47
47
 
48
48
  Current resolution: `fable` -> `claude-fable-5`, `opus` -> `claude-opus-5`,
@@ -61,7 +61,7 @@ Three shifts on the current opus rung are worth knowing when reading a run that
61
61
  different rather than broken, and none of them are bugs in this contract:
62
62
 
63
63
  - **Longer user-facing output.** Effort is not the lever; prompt-level conciseness is.
64
- Phase 7 report length and reviewer prose are where this shows.
64
+ Phase 5 report length and reviewer prose are where this shows.
65
65
  - **Self-verification without being asked.** Explicit "double-check your work"
66
66
  scaffolding now causes over-verification rather than preventing under-verification.
67
67
  Phase 4's deterministic gates already carry that load.
@@ -162,7 +162,7 @@ prints it once instead:
162
162
  with `PHASE_MODEL_OVERRIDE=<floorModel>`. A failure at the floor (or when no
163
163
  floor is configured) falls through to the normal phase-error path (pause ->
164
164
  resume). Never silent-skip the persona. Each downgrade emits its own
165
- `model_fallback` metric line so a two-step degrade is visible in Phase 7.
165
+ `model_fallback` metric line so a two-step degrade is visible in Phase 5.
166
166
  3. **Cost budget ceiling (existing gate).** When `cost-budget-check.mjs` exits 11
167
167
  (exceeded) mid-run, the run already pauses per the cost-budget contract; on
168
168
  user-approved continue, `preferredModel` personas downgrade to `fallbackModel` for the
@@ -172,7 +172,7 @@ prints it once instead:
172
172
 
173
173
  Every fallback emits one `log-metric.sh` line (`metric: model_fallback`,
174
174
  attributes: `persona`, `from`, `to`, `trigger: date-gate|dispatch-error|budget`)
175
- and one agent-log line so Phase 7 reports show which phases ran degraded.
175
+ and one agent-log line so Phase 5 reports show which phases ran degraded.
176
176
  The cost ledger prices the dispatch at the model actually used (the tracker's
177
177
  per-phase `model` field already carries the override).
178
178
 
@@ -213,5 +213,5 @@ parent model and reasoning effort and silently discards the override, so a
213
213
  fallback that omits it appears to apply while changing nothing.
214
214
 
215
215
  The three model ids above live in exactly three places - this table,
216
- `pipeline/scripts/cost-table.json`, and the Phase 4 reviewer matrix - so an OpenAI
216
+ `pipeline/scripts/cost-table.json`, and the Phase 3 reviewer matrix - so an OpenAI
217
217
  rename is a three-file change.
@@ -1,6 +1,6 @@
1
1
  # Feature: Plan Todos Iteration (Phase 3)
2
2
 
3
- **Gated by `prefs.global.planTodos.enabled`** (default: `false`). When enabled and Phase 2 Step 4.5 emitted a `plan.todos[]` (conforming to `$HOME/.claude/schemas/plan-todos.schema.json`), Phase 3 iterates with the helper instead of walking `tasks[]` directly:
3
+ **Gated by `prefs.global.planTodos.enabled`** (default: `false`). When enabled and Phase 1 Step 11 emitted a `plan.todos[]` (conforming to `$HOME/.claude/schemas/plan-todos.schema.json`), Phase 2 iterates with the helper instead of walking `tasks[]` directly:
4
4
 
5
5
  ```bash
6
6
  while next=$(bash "$HOME/.claude/lib/plan-todos.sh" next "$TASK_ID"); [ -n "$next" ]; do
@@ -23,7 +23,7 @@ fi
23
23
  - **Deterministic** - same worktree state → same map. No embeddings, no network. Sub-second on monorepos.
24
24
  - **Advisory** - Explore agents may ignore the map if they have stronger signals (Phase 0 knowledge cache, explicit user context). Never gates the pipeline.
25
25
  - **Truncation-safe** - output is capped at `tokenBudget`; the script falls back to "no files extracted within budget" rather than over-spending.
26
- - **Cost ledger** - emit `phase-1.repo_map_emitted bytes=$(wc -c <<<"$REPO_MAP") budget=$BUDGET` so Phase 7 cost summary can show the size injected.
26
+ - **Cost ledger** - emit `phase-1.repo_map_emitted bytes=$(wc -c <<<"$REPO_MAP") budget=$BUDGET` so Phase 5 cost summary can show the size injected.
27
27
 
28
28
  ## When to enable
29
29
 
@@ -6,7 +6,7 @@ Gated by `prefs.global.autopilotCircuitBreaker` (`enabled` default true, `identi
6
6
 
7
7
  ## Files per round
8
8
 
9
- Phase 4 Step 3.2.1 writes `$WORKTREE/.pipeline/triage-round-<N>.json` (N = `state.reviewIterations | length`), validates it, annotates it and copies it to `$WORKTREE/triage-output.json`, the name Phase 7, `worktree-finalize.sh`, `render-work-summary.sh` and `diff-explain.mjs` read. A copy rather than a symlink because the salvage is `cp -R`, and `.pipeline/` is on the salvage list, so every round survives into `artifactsPath`. Step 3.7 rewrites the round file; the copy is repeated after it.
9
+ Phase 3 Step 3.2.1 writes `$WORKTREE/.pipeline/triage-round-<N>.json` (N = `state.reviewIterations | length`), validates it, annotates it and copies it to `$WORKTREE/triage-output.json`, the name Phase 5, `worktree-finalize.sh`, `render-work-summary.sh` and `diff-explain.mjs` read. A copy rather than a symlink because the salvage is `cp -R`, and `.pipeline/` is on the salvage list, so every round survives into `artifactsPath`. Step 3.7 rewrites the round file; the copy is repeated after it.
10
10
 
11
11
  ## Step 2.1 block: previous-round findings (iteration >= 2)
12
12
 
@@ -34,7 +34,7 @@ The cap of 40 entries drops `important` before `blocking` so the prefix stays bo
34
34
 
35
35
  ## Step 2.2 block: scope self-check (every iteration)
36
36
 
37
- Phase 3 Step 3.7 wrote `$WORKTREE/.pipeline/scope-check.json` (contract: `features/scope-check.md`). Render it so reviewers judge the diff against the dev's stated scope and do not re-propose what was rejected:
37
+ Phase 2 Step 3.7 wrote `$WORKTREE/.pipeline/scope-check.json` (contract: `features/scope-check.md`). Render it so reviewers judge the diff against the dev's stated scope and do not re-propose what was rejected:
38
38
 
39
39
  ```bash
40
40
  SCOPE_JSON="$WORKTREE/.pipeline/scope-check.json"
@@ -82,7 +82,7 @@ Same argument `shortId()` in `_retrieval.mjs` makes for corpus rows: derived fro
82
82
 
83
83
  ## Telemetry
84
84
 
85
- `review.delta` once per iteration >= 2: `iteration`, `new`, `still_present`, `resolved`, `plateau`, `tripped`. `run-metrics.mjs` reports `reviewDelta.stillPresentFinal`, `resolvedTotal` and `tripped`; Phase 7 renders the three counts per round.
85
+ `review.delta` once per iteration >= 2: `iteration`, `new`, `still_present`, `resolved`, `plateau`, `tripped`. `run-metrics.mjs` reports `reviewDelta.stillPresentFinal`, `resolvedTotal` and `tripped`; Phase 5 renders the three counts per round.
86
86
 
87
87
  ## Reference
88
88
 
@@ -1,6 +1,6 @@
1
1
  # Phase 4 Review - Multi-Repo Mode
2
2
 
3
- > Loaded on demand from `phases/phase-4-review.md`. Conditional on
3
+ > Loaded on demand from `phases/phase-3-review.md`. Conditional on
4
4
  > `state.projects[].length > 1`; a single-repo run (the common case) never needs
5
5
  > it and used to pay for it on every review.
6
6
 
@@ -1,6 +1,6 @@
1
- # Feature: Scope self-check (Phase 3 Step 3.7)
1
+ # Feature: Scope self-check (Phase 2 Step 3.7)
2
2
 
3
- **Pattern**: Phase 4 reconstructs everything from the diff. The one thing it cannot reconstruct is why each file was touched and what was left out on purpose, so Dev states both before the handoff, and a deterministic gate checks the statement against the real diff. The record also carries the code-simplifier rationales (Step 3.6) that used to be discarded, and it feeds two later consumers: the `<scope-self-check>` block in the Phase 4 reviewer prefix and the PR body in Phase 6.
3
+ **Pattern**: Phase 3 reconstructs everything from the diff. The one thing it cannot reconstruct is why each file was touched and what was left out on purpose, so Dev states both before the handoff, and a deterministic gate checks the statement against the real diff. The record also carries the code-simplifier rationales (Step 3.6) that used to be discarded, and it feeds two later consumers: the `<scope-self-check>` block in the Phase 3 reviewer prefix and the PR body in Phase 4.
4
4
 
5
5
  ## The record
6
6
 
@@ -33,8 +33,8 @@ Output `{ok, justified, unjustified[], unlisted[], notDone}`. Exit 1 lists `unju
33
33
 
34
34
  ## Consumers
35
35
 
36
- - Phase 4 Step 2.2 renders file reasons, unjustified files and `notDone[]` into the shared reviewer prefix (`features/review-delta.md`).
37
- - Phase 6 Step 3 builds the PR `## Changes` bullets from `files[].reason` and lists `notDone[]` under `## Related` as "Follow-ups not done in this PR", merged with the final triage `deferred[]` (`channels/pr.md`).
36
+ - Phase 3 Step 2.2 renders file reasons, unjustified files and `notDone[]` into the shared reviewer prefix (`features/review-delta.md`).
37
+ - Phase 4 Step 3 builds the PR `## Changes` bullets from `files[].reason` and lists `notDone[]` under `## Related` as "Follow-ups not done in this PR", merged with the final triage `deferred[]` (`channels/pr.md`).
38
38
 
39
39
  ## Reference
40
40
 
@@ -13,7 +13,7 @@
13
13
  - [Preference](#preference)
14
14
  <!-- /toc -->
15
15
 
16
- > **TLDR** - Phase 4 Step 1.78 resolves, deterministically and before any reviewer runs, WHAT the changed code was supposed to honour: declared rule registries scoped to the diff's languages, in-repo module guides, and the toolchains those registries delegate to. The result is `criteria-manifest.json`: a bounded set of rule IDs that becomes the denominator for "was this applied completely". Reviewers answer per rule ID. An ID that is neither checked nor explicitly waived fails the stage.
16
+ > **TLDR** - Phase 3 Step 1.78 resolves, deterministically and before any reviewer runs, WHAT the changed code was supposed to honour: declared rule registries scoped to the diff's languages, in-repo module guides, and the toolchains those registries delegate to. The result is `criteria-manifest.json`: a bounded set of rule IDs that becomes the denominator for "was this applied completely". Reviewers answer per rule ID. An ID that is neither checked nor explicitly waived fails the stage.
17
17
 
18
18
  ## Why this exists
19
19
 
@@ -88,7 +88,7 @@ Cheap, language-agnostic, and it answers "completely" head on - an exception i
88
88
  |---|---|
89
89
  | `detectedStack` (Phase 1) | language census of the diff by file extension. Describes what the diff CONTAINS, not what the repo is nominally built in, so one Objective-C bridging file in a Swift repo is classified correctly |
90
90
  | Phase 1 analysis summary (triage scope) | the task description plus the inline task list Phase 3 generated for itself |
91
- | Phase 2 plan | same inline task list |
91
+ | Phase 1 plan | same inline task list |
92
92
  | `state.evidence.figma[]` (Step 1.8) | recorded as `not-applicable (no Phase 1 evidence)` |
93
93
  | Step 2.8 visual conformance | recorded as `not-applicable (no Phase 1 evidence)` |
94
94
 
@@ -67,7 +67,7 @@ Append one `state.telemetry.skillCalls[]` entry per skill actually loaded:
67
67
 
68
68
  `routedBy` names the index and version that chose it. That is the difference between "the model happened to read a skill" and "the toolkit said this skill governs this task".
69
69
 
70
- What Phase 4 actually does with it, precisely: Step 1.78 lists these entries in the manifest under `ledger.routedByToolkit`, so a reviewer and the Phase 7 report can see which skills the project's own toolkit selected. It does **not** give them extra weight in the coverage maths. The deterministic resolver stays primary because an unrecorded load and no load are indistinguishable in state, and no `routedBy` tag changes that - the tag says who chose the skill, not that the code honoured it.
70
+ What Phase 4 actually does with it, precisely: Step 1.78 lists these entries in the manifest under `ledger.routedByToolkit`, so a reviewer and the Phase 5 report can see which skills the project's own toolkit selected. It does **not** give them extra weight in the coverage maths. The deterministic resolver stays primary because an unrecorded load and no load are indistinguishable in state, and no `routedBy` tag changes that - the tag says who chose the skill, not that the code honoured it.
71
71
 
72
72
  ## Failure modes, and why none of them halt
73
73
 
@@ -1,4 +1,4 @@
1
- # Feature: Verify-by-Test Triage (Phase 4 Step 3.7)
1
+ # Feature: Verify-by-Test Triage (Phase 3 Step 3.7)
2
2
 
3
3
  **Pattern**: reviewer findings are hypotheses; a failing repro test is proof. Adversarial-review research (Refute-or-Promote, 2026) found that plausible-but-wrong findings survive debate rounds but die on a single empirical test. This step converts the highest-stakes verdicts (accepted blocking) from judgment into evidence before the Phase 3 rework loop fires.
4
4
 
@@ -19,7 +19,7 @@
19
19
  | Passes on some runs, fails on others | `inconclusive` | Flake, not proof either way. Finding stays accepted blocking, `verification.note = "flaky: passed k/N"`, repro test file deleted, telemetry line gains `flaky=<count>`. |
20
20
  | Compile error / timeout / not unit-testable | `inconclusive` | Finding stays accepted blocking (judgment stands). Partial test deleted, cause in `verification.note`. |
21
21
 
22
- Downgrades go to `deferred`, never `rejected`: triage judged the issue real, and deferred items surface in the Phase 7 report for a human eye.
22
+ Downgrades go to `deferred`, never `rejected`: triage judged the issue real, and deferred items surface in the Phase 5 report for a human eye.
23
23
 
24
24
  ## Red-test handoff to Phase 3
25
25
 
@@ -27,7 +27,7 @@ Downgrades go to `deferred`, never `rejected`: triage judged the issue real, and
27
27
 
28
28
  ## Cleanup invariant
29
29
 
30
- After Step 3.7, the only uncommitted verifier artifacts are the confirmed repro tests listed in `redTests[]` (committed later with the fix) and logs under `$WORKTREE/.pipeline/` (outside Phase 6 commit scope). `not-reproduced` and `inconclusive` test files are always deleted.
30
+ After Step 3.7, the only uncommitted verifier artifacts are the confirmed repro tests listed in `redTests[]` (committed later with the fix) and logs under `$WORKTREE/.pipeline/` (outside Phase 4 commit scope). `not-reproduced` and `inconclusive` test files are always deleted.
31
31
 
32
32
  ## Telemetry
33
33
 
@@ -39,4 +39,4 @@ Adds one Sonnet call plus up to `maxFindings` single-test runs (and build-lock c
39
39
 
40
40
  ## Reference
41
41
 
42
- Wiring: `$HOME/.claude/multi-agent-refs/phases/phase-4-review.md` Step 3.7. Schema: `$HOME/.claude/schemas/triage-output.schema.json` v3.2.0 (`$defs.verification`). Evidence gate: `$HOME/.claude/scripts/evidence-gate.mjs`. Prefs: `prefs.global.verifyByTest` (including `repeatCount`) in `$HOME/.claude/schemas/prefs.schema.json`.
42
+ Wiring: `$HOME/.claude/multi-agent-refs/phases/phase-3-review.md` Step 3.7. Schema: `$HOME/.claude/schemas/triage-output.schema.json` v3.2.0 (`$defs.verification`). Evidence gate: `$HOME/.claude/scripts/evidence-gate.mjs`. Prefs: `prefs.global.verifyByTest` (including `repeatCount`) in `$HOME/.claude/schemas/prefs.schema.json`.
@@ -0,0 +1,83 @@
1
+ # verify - is this install the thing that was published
2
+
3
+ The install is a COPY. `install.js` writes the pipeline tree into `~/.claude`,
4
+ `~/.copilot` and `~/.codex`, and from that moment the two halves drift
5
+ independently. Both directions produce bugs that are hard to name:
6
+
7
+ - an edit made in the installed copy is a behaviour with no source, and the next
8
+ update silently reverts it;
9
+ - a file the installer failed to write is a script the docs describe and nobody
10
+ has, which reads as a documentation error.
11
+
12
+ `multi-agent-pipeline verify` answers both mechanically.
13
+
14
+ ```bash
15
+ npx @mmerterden/multi-agent-pipeline verify # package + install
16
+ npx @mmerterden/multi-agent-pipeline verify --package # package integrity only
17
+ npx @mmerterden/multi-agent-pipeline verify --install # install drift only
18
+ npx @mmerterden/multi-agent-pipeline verify --json
19
+ ```
20
+
21
+ | Code | Meaning |
22
+ |---|---|
23
+ | 0 | everything matches |
24
+ | 1 | a difference was found, named file by file |
25
+ | 2 | nothing to verify - a source checkout, or a version published before manifests existed |
26
+
27
+ Exit 2 is not a pass and not a failure. A dev checkout has no manifest by
28
+ design, and reporting that as either would be a lie in one direction or the
29
+ other.
30
+
31
+ ## The manifest
32
+
33
+ `manifest.json` is written at pack time by `prepack`, never committed. A
34
+ manifest in git is stale one commit after it is written, and a stale manifest
35
+ reports honest edits as tampering - which is worse than having none, because
36
+ people learn to ignore it.
37
+
38
+ The file list is not guessed. It comes from `npm pack --dry-run --json`, so by
39
+ construction it is the same set npm publishes, `files` globs and all. The gate
40
+ asserts the two counts agree, which is what catches a `files` entry and a
41
+ manifest that have stopped describing the same package.
42
+
43
+ Two things it cannot cover, said here rather than discovered later: it cannot
44
+ hash itself, and a signature over it does not authenticate the tarball.
45
+
46
+ ## What a green result proves, and what it does not
47
+
48
+ It proves the bytes match what the publisher recorded. It is not proof of WHO
49
+ published them. The manifest, the signature and the verifier all travel inside
50
+ the same tarball, so anyone able to rewrite one can rewrite the others.
51
+ Provenance belongs to npm's own integrity field.
52
+
53
+ What this does catch is the set of failures that actually happen: a damaged or
54
+ partial install, a file edited after install, and an update that did not land.
55
+
56
+ Signing is optional. `make-manifest.mjs --sign` reads an ed25519 private key
57
+ from the credential store (or `MULTI_AGENT_SIGNING_KEY` on a build host with no
58
+ store) and writes `manifest.sig`; `verify` checks it against
59
+ `MULTI_AGENT_SIGNING_PUBKEY` when one is pinned. Without a key it says "signed,
60
+ no public key to check it against" rather than claiming valid - a signature
61
+ nobody can check is not a signature that passed.
62
+
63
+ ## How each tree is compared
64
+
65
+ | Tree | Mode | Why |
66
+ |---|---|---|
67
+ | `scripts` | bytes | verbatim copy, minus the dev-only set |
68
+ | `lib` | bytes | verbatim copy |
69
+ | `multi-agent-refs` | bytes | verbatim copy |
70
+ | `agents` | bytes | verbatim copy |
71
+ | `commands/multi-agent` | presence | `install.js` rewrites each SKILL.md `description` into the user's `outputLanguage` |
72
+
73
+ Byte-comparing `commands/` reports every command as drift on a perfectly
74
+ healthy machine. Measured here: all 57 command files differ, and 56 of them
75
+ differ by nothing except the translated description. A report that is wrong by
76
+ default is a report nobody reads.
77
+
78
+ The dev-only filter matters just as much: smokes, linters and fixtures ship in
79
+ the package and are deliberately NOT installed. Without excluding them, `verify`
80
+ would report 252 files as "the installer skipped this".
81
+
82
+ `~/.copilot` and `~/.codex` are reported as present, not compared: the installer
83
+ rewrites paths for both on purpose, so a byte difference there is the design.
@@ -3,7 +3,7 @@
3
3
  <!-- toc -->
4
4
  - [1. When it is required](#1-when-it-is-required)
5
5
  - [2. Before - the reporter's screenshot, or nothing](#2-before---the-reporters-screenshot-or-nothing)
6
- - [3. After - Phase 3, not Phase 5](#3-after---phase-3-not-phase-5)
6
+ - [3. After - Phase 2, not the Phase 3 user test](#3-after---phase-2-not-the-phase-3-user-test)
7
7
  - [4. Video - the recording rides on a test run](#4-video---the-recording-rides-on-a-test-run)
8
8
  - [5. Size, and what happens when it does not fit](#5-size-and-what-happens-when-it-does-not-fit)
9
9
  - [5b. Where the artefacts live - the host](#5b-where-the-artefacts-live---the-host)
@@ -18,10 +18,10 @@ This contract makes the pipeline carry the picture: the state the reporter saw,
18
18
  the state the fix produces, and where possible a recording of the flow running.
19
19
 
20
20
  The artefacts live as **Jira attachments** and are referenced from two places -
21
- the Phase 7 Jira comment, where the picture belongs next to the work summary and
21
+ the Phase 5 Jira comment, where the picture belongs next to the work summary and
22
22
  the test scenarios, and the PR body, which tells the reviewer they exist.
23
23
 
24
- Consumers: Phase 3 (capture), Phase 5 (video, preferred host), Phase 6 (blocker),
24
+ Consumers: Phase 2 (capture), Phase 3 (video, user test), Phase 4 (blocker),
25
25
  `channels/jira.md` and `channels/pr.md` (render). Gate: `smoke-visual-evidence.sh`.
26
26
 
27
27
  ## 1. When it is required
@@ -49,16 +49,16 @@ artefact is absent, the absence is written down with its reason (section 5).
49
49
 
50
50
  **Phase 0 Step 7.7 writes `required`, `requiredBy` and `platform`**, before it
51
51
  probes. Naming the writer is not a formality: five phases read these three fields
52
- and nothing set them, so `required` was never true, Step 7.7 never ran, Phase 3
53
- never captured, and Phase 6's blocker never blocked. A contract with readers and
52
+ and nothing set them, so `required` was never true, Step 7.7 never ran, Phase 2
53
+ never captured, and Phase 4's blocker never blocked. A contract with readers and
54
54
  no writer reads exactly like a contract that is satisfied.
55
55
 
56
56
  Phase 0 has no diff, so it decides on what it actually knows - `taskType` and
57
57
  whether the intake carried a Figma reference - and that is enough for the whole
58
58
  design-work row. The `bugfix` row needs a changed-file list, so Phase 0 records
59
- it as provisional (`requiredBy: "bugfix (pending diff)"`) and **Phase 3 re-decides
59
+ it as provisional (`requiredBy: "bugfix (pending diff)"`) and **Phase 2 re-decides
60
60
  from the real diff** before capturing. A provisional `false` never suppresses the
61
- Phase 3 decision.
61
+ Phase 2 decision.
62
62
 
63
63
  ### 1b. The platform, and when there is not one
64
64
 
@@ -120,15 +120,15 @@ No image on the ticket means no "before". Record the gap
120
120
  (`before_missing: "ticket carries no image attachment"`) and continue. The
121
121
  section still renders, saying so.
122
122
 
123
- ## 3. After - Phase 3, not Phase 5
123
+ ## 3. After - Phase 2, not the Phase 3 user test
124
124
 
125
- Phase 5 is the natural home: the simulator is already up. It is also **dropped by
126
- every `autopilot` and `--local` entry** (`phase-5-test.md` TLDR), so a capture
125
+ The user test is the natural home: the simulator is already up. It is also **dropped by
126
+ every `autopilot` and `--local` entry** (`phase-3-review.md` TLDR), so a capture
127
127
  that lives only there produces nothing for unattended runs - which are exactly
128
128
  the runs where nobody watched the screen.
129
129
 
130
- The capture therefore happens in **Phase 3, after the build+test gate passes**,
131
- which is in every mode's phase set. Phase 5 may add richer evidence on top; it is
130
+ The capture therefore happens in **Phase 2, after the build+test gate passes**,
131
+ which is in every mode's phase set. The Phase 3 user test may add richer evidence on top; it is
132
132
  never the only source.
133
133
 
134
134
  ```bash
@@ -192,8 +192,8 @@ The recorder itself is `capture-evidence.sh video start|stop`, which writes
192
192
  `<task-id>[-<label>]-flow.mp4` into the evidence directory. The label is optional
193
193
  because one recording per task is the common case, and both renderers cite the
194
194
  unlabelled form. It is shell rather than MCP for three reasons: the sibling still capture already shells out,
195
- a host with no toolkit MCP registered still produces evidence, and Phase 3 and
196
- Phase 5 are exactly where an MCP call is contested.
195
+ a host with no toolkit MCP registered still produces evidence, and Phase 2 and
196
+ Phase 3 are exactly where an MCP call is contested.
197
197
 
198
198
  **iOS UI test targets are not named, they are detected.** A path containing
199
199
  `UITests` is not the signal: in the reference app 477 files sit under such a path
@@ -210,8 +210,8 @@ wearing a measurement's clothes.
210
210
 
211
211
  ### 4.4 Re-check before recording
212
212
 
213
- The probe runs at intake and the recording happens in Phase 3. A simulator booted
214
- then can be gone by the time the build goes green, so Phase 3 re-measures the
213
+ The probe runs at intake and the recording happens in Phase 2. A simulator booted
214
+ then can be gone by the time the build goes green, so Phase 2 re-measures the
215
215
  device row alone and downgrades the tier if it has to, recording the transition
216
216
  (`tier 1 -> 3: simulator no longer booted`). A tier taken from a stale
217
217
  measurement is a promise the run cannot keep.
@@ -232,7 +232,7 @@ never asserted against wall clock anywhere, because doing so fails a correct
232
232
  capture of a static screen.
233
233
 
234
234
  `visualEvidence.enabled` turns the whole feature off - capture, upload, both
235
- render sections and the Phase 6 blocker with it.
235
+ render sections and the Phase 4 blocker with it.
236
236
 
237
237
  ### 4.6 Web - the recorder IS the runner
238
238
 
@@ -283,7 +283,7 @@ Never silently attach nothing.
283
283
 
284
284
  ## 5b. Where the artefacts live - the host
285
285
 
286
- Resolved in Phase 6 Step 2.9, recorded as `state.visualEvidence.host`:
286
+ Resolved in Phase 4 Step 2.9, recorded as `state.visualEvidence.host`:
287
287
 
288
288
  | Order | Host | Condition | Stills | Video |
289
289
  |---|---|---|---|---|
@@ -342,7 +342,7 @@ where it lives - the images themselves stay on the ticket:
342
342
  ## 7. Blocker
343
343
 
344
344
  When section 1 says required and `state.visualEvidence` carries neither an
345
- artefact nor a recorded reason for its absence, **Phase 6 Step 3 blocks** - the
345
+ artefact nor a recorded reason for its absence, **Phase 4 Step 3 blocks** - the
346
346
  same shape as the `risk` section blocker. A recorded reason is enough to pass:
347
347
  the gate is against silence, not against an honest "no image on the ticket".
348
348
 
@@ -1,6 +1,6 @@
1
1
  # Worktree finalize - removing a task's worktree once its PR is open
2
2
 
3
- > **TLDR** - Phase 6 step 9. Once the PR exists the worktree is dead weight, so it is removed: artefacts are salvaged into the log dir first, the branch is kept and deliberately NOT checked out, and every destructive path is gated. Gated by `prefs.global.settings.worktreeAutoRemoveOnPr` (default **true**). Script: `worktree-finalize.sh`.
3
+ > **TLDR** - Phase 4 step 9. Once the PR exists the worktree is dead weight, so it is removed: artefacts are salvaged into the log dir first, the branch is kept and deliberately NOT checked out, and every destructive path is gated. Gated by `prefs.global.settings.worktreeAutoRemoveOnPr` (default **true**). Script: `worktree-finalize.sh`.
4
4
 
5
5
  ## Why it exists
6
6
 
@@ -21,12 +21,12 @@ Removing it at PR-open is only safe because of the salvage, so the two are one s
21
21
  | Condition | Why it blocks |
22
22
  |---|---|
23
23
  | `worktreePath == projectRoot` (`--local` mode) | there is no worktree; removing it would delete the user's checkout |
24
- | cwd is inside the worktree | a shell left on a deleted inode is worse than a leftover directory, and Phase 6 legitimately `cd`s into the worktree earlier |
24
+ | cwd is inside the worktree | a shell left on a deleted inode is worse than a leftover directory, and Phase 4 legitimately `cd`s into the worktree earlier |
25
25
  | not a registered worktree of the project root | a mistyped path must not delete an unrelated directory |
26
26
  | real uncommitted changes | never discarded; see the artefact carve-out below |
27
27
  | HEAD not on the remote | removing a worktree whose commits exist nowhere else is data loss, not cleanup |
28
28
 
29
- Exit codes: `0` removed, `3` skipped with a reason (report and continue to Phase 7), `1` usage error.
29
+ Exit codes: `0` removed, `3` skipped with a reason (report and continue to Phase 5), `1` usage error.
30
30
 
31
31
  ### The artefact carve-out, and why `--untracked-files=no` is wrong
32
32
 
@@ -49,10 +49,10 @@ This is why the removal is safe:
49
49
 
50
50
  | Consumer | Reads | Without salvage |
51
51
  |---|---|---|
52
- | Phase 7 triage-memory ingest | `triage-output.json` | `[ -f ]`-guarded, so it degrades **silently**: the triage corpus and learnings ledger stop being fed and no error appears. Since v16.20.0 Phase 4 writes `triage-output.json` itself (Step 3.2.1, the latest copy of `.pipeline/triage-round-<N>.json`), so the salvage is a second copy, not the only bridge |
53
- | Phase 7 learnings-ledger distill | same file | same silent degradation |
52
+ | Phase 5 triage-memory ingest | `triage-output.json` | `[ -f ]`-guarded, so it degrades **silently**: the triage corpus and learnings ledger stop being fed and no error appears. Since v16.20.0 Phase 4 writes `triage-output.json` itself (Step 3.2.1, the latest copy of `.pipeline/triage-round-<N>.json`), so the salvage is a second copy, not the only bridge |
53
+ | Phase 5 learnings-ledger distill | same file | same silent degradation |
54
54
  | `render-work-summary.sh` | `agent-state.json`, `phase-tracker.json` (falls back to `logs/multi-agent/<task>/tracker-state.json`, which survives removal) | loses the salvaged copies but keeps the tracker via the logs fallback |
55
- | `:resume` | `agent-state.json` | cannot continue a Phase 7 pause |
55
+ | `:resume` | `agent-state.json` | cannot continue a Phase 5 pause |
56
56
  | `:status`, `:log` | `agent-state.json` | the task becomes invisible |
57
57
 
58
58
  `state.worktreeRemovedAt` and `state.artifactsPath` record the outcome. The timestamp is what tells a reader that a worktree-less task was finished-and-tidied rather than killed - without it, a missing worktree is indistinguishable from a broken run.
@@ -3,15 +3,15 @@
3
3
  <!-- toc -->
4
4
  - [The triad at a glance](#the-triad-at-a-glance)
5
5
  - [Phase 0 auto-create policy](#phase-0-auto-create-policy)
6
- - [Phase 7 wiki → Jira comment](#phase-7-wiki-jira-comment)
6
+ - [Phase 5 wiki → Jira comment](#phase-5-wiki-jira-comment)
7
7
  - [Autopilot behaviour](#autopilot-behaviour)
8
8
  - [Preferences involved](#preferences-involved)
9
9
  - [Cross-CLI parity](#cross-cli-parity)
10
10
  <!-- /toc -->
11
11
 
12
- > **TLDR** - When a GitHub issue triggers the pipeline and has no Jira ID, the `autoJiraFromGithubIssue` policy decides whether to auto-create a Jira task (and patch the GitHub issue body with the new Jira link). Phase 7 then posts a humanizer'd wiki-content summary back as a Jira comment, closing the loop. Autopilot treats `ask` as `always`.
12
+ > **TLDR** - When a GitHub issue triggers the pipeline and has no Jira ID, the `autoJiraFromGithubIssue` policy decides whether to auto-create a Jira task (and patch the GitHub issue body with the new Jira link). Phase 5 then posts a humanizer'd wiki-content summary back as a Jira comment, closing the loop. Autopilot treats `ask` as `always`.
13
13
 
14
- This doc is referenced from `$HOME/.claude/multi-agent-refs/phases/phase-0-init.md` Step 1 (GitHub issue input) and `$HOME/.claude/multi-agent-refs/phases/phase-7-report.md` Step 2 (component wiki). Keeps the phase docs tight and gives the triad contract a stable home.
14
+ This doc is referenced from `$HOME/.claude/multi-agent-refs/phases/phase-0-init.md` Step 1 (GitHub issue input) and `$HOME/.claude/multi-agent-refs/phases/phase-5-report.md` Step 2 (component wiki). Keeps the phase docs tight and gives the triad contract a stable home.
15
15
 
16
16
  ## The triad at a glance
17
17
 
@@ -24,7 +24,7 @@ Jira issue (new, linked both ways)
24
24
  ▼ Phase 3 - component implementation
25
25
  wiki markdown in pipeline/skills/...
26
26
  │
27
- ▼ Phase 7 Report Step 2 - wiki capture + Jira comment post
27
+ ▼ Phase 5 Report Step 2 - wiki capture + Jira comment post
28
28
  Jira comment (humanizer'd component summary)
29
29
  ```
30
30
 
@@ -59,13 +59,13 @@ Decision table:
59
59
  5. On success: store new Jira key in `state.jiraId`. Emit `→ created Jira {newKey}`.
60
60
  6. Patch the GitHub issue body: append a line `Jira: [{newKey}]({jira.baseUrl}/browse/{newKey})` so the bidirectional link is visible on GitHub without refetching the pipeline state.
61
61
 
62
- **Skip path** (`never` or user declined): Branch falls back to `feature/GH{issueNo}-{kebab}`; Phase 7 later emits `→ no Jira linked (policy: {preference})`.
62
+ **Skip path** (`never` or user declined): Branch falls back to `feature/GH{issueNo}-{kebab}`; Phase 5 later emits `→ no Jira linked (policy: {preference})`.
63
63
 
64
64
  **Failure isolation:** Jira API 5xx or network error → log the error, warn the user, and **continue with no Jira link** (do not halt Phase 0). Rationale: a transient Jira outage should not block code work; user can link manually post-hoc.
65
65
 
66
- ## Phase 7 wiki → Jira comment
66
+ ## Phase 5 wiki → Jira comment
67
67
 
68
- Runs after Phase 7 Report Step 2 (wiki capture) has produced markdown and only when:
68
+ Runs after Phase 5 Report Step 2 (wiki capture) has produced markdown and only when:
69
69
 
70
70
  - `state.jiraId` is non-null (from Phase 0 scan or auto-create).
71
71
  - `prefs.global.wikiToJiraComment !== false` (default true).
@@ -78,9 +78,9 @@ Flow:
78
78
  3. Render as Jira comment: title line (`h3. Component docs - {componentName}`) + an overview paragraph + a link to the full page (wiki URL, if the adapter returned `pushedRemote`). The overview lines come from a Markdown file - convert them via the `channels/jira.md` table before POST, exactly like the title line already is, then run the whole body through `node "$HOME/.claude/scripts/jira-wiki-escape.mjs"` per that file's *Emoticon escaping* section.
79
79
  4. **Run through the humanizer skill** - same policy as Step 4 Confluence: user-facing content must read naturally.
80
80
  5. `POST {jira.baseUrl}/rest/api/2/issue/{jiraId}/comment`.
81
- 6. Log: `Phase 7: wiki summary posted to Jira {jiraId}`.
81
+ 6. Log: `Phase 5: wiki summary posted to Jira {jiraId}`.
82
82
 
83
- On failure: log + continue. Wiki is already written; the Jira comment is an augmentation. Don't regress Phase 7 over a Jira 5xx.
83
+ On failure: log + continue. Wiki is already written; the Jira comment is an augmentation. Don't regress Phase 5 over a Jira 5xx.
84
84
 
85
85
  ## Autopilot behaviour
86
86
 
@@ -96,7 +96,7 @@ Rationale: autopilot is explicit consent for side-effect creation; muting Jira a
96
96
  | key | type | default | controls |
97
97
  |---|---|---|---|
98
98
  | `prefs.global.autoJiraFromGithubIssue` | enum ask/always/never | `ask` | Phase 0 auto-create decision |
99
- | `prefs.global.wikiToJiraComment` | bool | `true` | Phase 7 wiki → Jira comment post |
99
+ | `prefs.global.wikiToJiraComment` | bool | `true` | Phase 5 wiki → Jira comment post |
100
100
  | `figmaConfig.jira.projectKey` | string | - | Target project for new issues (falls back to `prefs.global.defaultJiraKey`) |
101
101
  | `prefs.global.keychainMapping.jira` | string | - | Jira token keychain lookup |
102
102
 
@@ -30,11 +30,11 @@ $HOME/.claude/knowledge/
30
30
  ### How It Works
31
31
 
32
32
  ```
33
- Task 1 (first time): Phase 1 -> full Explore -> Phase 7 -> create knowledge
34
- Task 2: Phase 1 -> read knowledge -> targeted Explore -> Phase 7 -> update knowledge
35
- Task 3: Phase 1 -> read knowledge -> minimal Explore -> Phase 7 -> update knowledge
33
+ Task 1 (first time): Phase 1 -> full Explore -> Phase 5 -> create knowledge
34
+ Task 2: Phase 1 -> read knowledge -> targeted Explore -> Phase 5 -> update knowledge
35
+ Task 3: Phase 1 -> read knowledge -> minimal Explore -> Phase 5 -> update knowledge
36
36
  ...
37
- Task N: Phase 1 -> knowledge sufficient -> SKIP Explore -> Phase 7 -> update knowledge
37
+ Task N: Phase 1 -> knowledge sufficient -> SKIP Explore -> Phase 5 -> update knowledge
38
38
  ```
39
39
 
40
40
  Each task teaches the project a bit more. Token cost decreases over time.
@@ -45,7 +45,7 @@ Each task teaches the project a bit more. Token cost decreases over time.
45
45
  | ------------------------------------------- | --------------------- | ------------------------------------------ | ------------------------ |
46
46
  | `$PROJECT_ROOT/CLAUDE.md` | User | Project rules, build commands | Persistent (in git) |
47
47
  | Auto Memory (`~/.claude/projects/`) | Claude (automatic) | User preferences, feedback | Persistent (global) |
48
- | **Knowledge Base** (`~/.claude/knowledge/`) | Multi-agent (Phase 7) | Architecture, patterns, gotchas, decisions | Persistent (per-project) |
48
+ | **Knowledge Base** (`~/.claude/knowledge/`) | Multi-agent (Phase 5) | Architecture, patterns, gotchas, decisions | Persistent (per-project) |
49
49
 
50
50
  They all complement each other - no conflicts.
51
51
 
@@ -54,7 +54,7 @@ They all complement each other - no conflicts.
54
54
  A Short run skips Phase 1, but knowledge **is still read**:
55
55
 
56
56
  - Knowledge files are added to the prompt before the Opus agent starts in Phase 3
57
- - Knowledge capture is still performed in Phase 7 (simplified report)
57
+ - Knowledge capture is still performed in Phase 5 (simplified report)
58
58
 
59
59
  ### Knowledge Maintenance
60
60
 
@@ -107,18 +107,18 @@ When calling sub-agents, inject task-specific context from previous phases:
107
107
  ### Phase -> Agent -> Context Flow
108
108
 
109
109
  ```
110
- Phase 1: Analysis
110
+ Phase 1: Plan (analysis)
111
111
  -- Explore agents (parallel) -> codebase scan
112
112
 
113
- Phase 2: Planning
113
+ Phase 1: Plan
114
114
  -- Architecture review agent
115
115
  Context: task description + Phase 1 findings + affected modules
116
116
 
117
- Phase 3: Dev
117
+ Phase 2: Dev
118
118
  -- Sonnet agent (TDD)
119
- Context: Phase 2 plan + task dependencies
119
+ Context: Phase 1 plan + task dependencies
120
120
 
121
- Phase 4: Review (CLI-aware parallel + triage)
121
+ Phase 3: Review (CLI-aware parallel + triage)
122
122
  Claude Code (3 parallel):
123
123
  |-- Reviewer (opus) -> security + architecture
124
124
  +-- Reviewer (sonnet) -> quality + correctness + edge cases