@mmerterden/multi-agent-pipeline 17.6.0 → 19.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (272) hide show
  1. package/CHANGELOG.md +310 -0
  2. package/README.md +76 -18
  3. package/README.tr.md +55 -16
  4. package/docs/adr/0002-instruction-driven-flag.md +1 -0
  5. package/docs/adr/0005-lazy-phase-docs.md +11 -1
  6. package/docs/adr/0008-installer-modularization-and-secret-leak-defense.md +1 -0
  7. package/docs/adr/0010-own-code-graph.md +1 -0
  8. package/docs/adr/0011-dormant-ci.md +25 -1
  9. package/docs/adr/0014-six-phase-consolidation.md +134 -0
  10. package/docs/adr/README.md +2 -1
  11. package/docs/architecture.md +37 -38
  12. package/docs/best-practices.md +1 -1
  13. package/docs/ecosystem.md +37 -26
  14. package/docs/engineering.md +1 -1
  15. package/docs/facts.json +45 -0
  16. package/docs/features.md +54 -53
  17. package/docs/performance.md +5 -5
  18. package/docs/recovery-guide.md +9 -9
  19. package/docs/server-readiness.md +188 -0
  20. package/docs/token-budget-history.md +3 -1
  21. package/index.js +18 -3
  22. package/install/_codex-agents.mjs +1 -1
  23. package/install/_common.mjs +42 -17
  24. package/install/_dev-only-files.mjs +8 -0
  25. package/install/_unattended-profile.mjs +113 -0
  26. package/install/index.mjs +48 -0
  27. package/install/templates/claude-hooks.json +1 -1
  28. package/install/templates/codex-instructions.md +1 -1
  29. package/install/templates/copilot-instructions.md +28 -28
  30. package/manifest.json +1065 -0
  31. package/package.json +6 -3
  32. package/pipeline/agents/dev-critic.md +3 -3
  33. package/pipeline/commands/figma-to-swiftui.md +1 -1
  34. package/pipeline/commands/multi-agent/SKILL.md +8 -8
  35. package/pipeline/commands/multi-agent/analysis/SKILL.md +9 -9
  36. package/pipeline/commands/multi-agent/autopilot/SKILL.md +7 -7
  37. package/pipeline/commands/multi-agent/channels/SKILL.md +15 -15
  38. package/pipeline/commands/multi-agent/diff-explain/SKILL.md +6 -6
  39. package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +1 -1
  40. package/pipeline/commands/multi-agent/graph/SKILL.md +1 -1
  41. package/pipeline/commands/multi-agent/help/SKILL.md +62 -62
  42. package/pipeline/commands/multi-agent/language/SKILL.md +2 -2
  43. package/pipeline/commands/multi-agent/local/SKILL.md +11 -11
  44. package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +13 -13
  45. package/pipeline/commands/multi-agent/log/SKILL.md +2 -2
  46. package/pipeline/commands/multi-agent/manual-test/SKILL.md +9 -9
  47. package/pipeline/commands/multi-agent/model/SKILL.md +69 -0
  48. package/pipeline/commands/multi-agent/refactor/SKILL.md +3 -3
  49. package/pipeline/commands/multi-agent/resume/SKILL.md +4 -4
  50. package/pipeline/commands/multi-agent/resume-local/SKILL.md +19 -17
  51. package/pipeline/commands/multi-agent/review/SKILL.md +1 -1
  52. package/pipeline/commands/multi-agent/route-off/SKILL.md +36 -0
  53. package/pipeline/commands/multi-agent/route-on/SKILL.md +74 -0
  54. package/pipeline/commands/multi-agent/route-status/SKILL.md +56 -0
  55. package/pipeline/commands/multi-agent/setup/SKILL.md +2 -2
  56. package/pipeline/commands/multi-agent/status/SKILL.md +54 -23
  57. package/pipeline/commands/multi-agent/steer/SKILL.md +2 -2
  58. package/pipeline/commands/multi-agent/sync/SKILL.md +12 -13
  59. package/pipeline/commands/multi-agent/test/SKILL.md +1 -1
  60. package/pipeline/lib/_jira-auth.sh +8 -0
  61. package/pipeline/lib/analysis-jira-write.sh +32 -0
  62. package/pipeline/lib/ask-choice.sh +13 -2
  63. package/pipeline/lib/autopilot-state.sh +8 -0
  64. package/pipeline/lib/credential-inventory.sh +1 -1
  65. package/pipeline/lib/fatal.mjs +129 -0
  66. package/pipeline/lib/fetch-fortify.sh +1 -1
  67. package/pipeline/lib/figma-mcp-refresh.sh +18 -0
  68. package/pipeline/lib/figma-screenshot.sh +18 -0
  69. package/pipeline/lib/invoked-directly.mjs +43 -0
  70. package/pipeline/lib/jira-publish.sh +42 -0
  71. package/pipeline/lib/md2confluence-v3.py +47 -0
  72. package/pipeline/lib/model-rung.sh +142 -0
  73. package/pipeline/lib/outbound-gate.mjs +175 -0
  74. package/pipeline/lib/phase-schema.mjs +88 -0
  75. package/pipeline/lib/plan-todos.sh +32 -11
  76. package/pipeline/lib/post-pr-review.sh +77 -8
  77. package/pipeline/lib/repo-hygiene.sh +8 -3
  78. package/pipeline/lib/require-jq.sh +40 -0
  79. package/pipeline/lib/route-state.sh +161 -0
  80. package/pipeline/lib/run-paths.sh +335 -0
  81. package/pipeline/multi-agent-refs/_account-picker.md +1 -1
  82. package/pipeline/multi-agent-refs/_dev-context.md +1 -1
  83. package/pipeline/multi-agent-refs/_input-parser.md +1 -1
  84. package/pipeline/multi-agent-refs/analysis/evidence.md +0 -9
  85. package/pipeline/multi-agent-refs/analysis/intake.md +1 -1
  86. package/pipeline/multi-agent-refs/analysis/locked.md +21 -22
  87. package/pipeline/multi-agent-refs/analysis/render.md +1 -1
  88. package/pipeline/multi-agent-refs/analysis/synthesis.md +12 -6
  89. package/pipeline/multi-agent-refs/android-guide.md +1 -1
  90. package/pipeline/multi-agent-refs/audit-guide.md +13 -13
  91. package/pipeline/multi-agent-refs/channels/issue-comment.md +2 -2
  92. package/pipeline/multi-agent-refs/channels/jira.md +3 -3
  93. package/pipeline/multi-agent-refs/channels/pr.md +4 -4
  94. package/pipeline/multi-agent-refs/channels/wiki.md +1 -1
  95. package/pipeline/multi-agent-refs/component-dispatch.md +3 -3
  96. package/pipeline/multi-agent-refs/cross-cli-contract.md +31 -6
  97. package/pipeline/multi-agent-refs/features/autopilot-circuit-breaker.md +74 -4
  98. package/pipeline/multi-agent-refs/features/code-graph.md +5 -5
  99. package/pipeline/multi-agent-refs/features/cost-analysis.md +93 -0
  100. package/pipeline/multi-agent-refs/features/design-conformance.md +1 -1
  101. package/pipeline/multi-agent-refs/features/dev-critic.md +3 -3
  102. package/pipeline/multi-agent-refs/features/doctor.md +47 -2
  103. package/pipeline/multi-agent-refs/features/external-context-injection.md +3 -3
  104. package/pipeline/multi-agent-refs/features/maturity-followup.md +3 -3
  105. package/pipeline/multi-agent-refs/features/model-fallback.md +5 -5
  106. package/pipeline/multi-agent-refs/features/plan-todos.md +1 -1
  107. package/pipeline/multi-agent-refs/features/repo-map.md +1 -1
  108. package/pipeline/multi-agent-refs/features/review-delta.md +3 -3
  109. package/pipeline/multi-agent-refs/features/review-multi-repo.md +1 -1
  110. package/pipeline/multi-agent-refs/features/scope-check.md +4 -4
  111. package/pipeline/multi-agent-refs/features/skill-conformance.md +2 -2
  112. package/pipeline/multi-agent-refs/features/stack-skill-routing.md +1 -1
  113. package/pipeline/multi-agent-refs/features/verify-by-test.md +4 -4
  114. package/pipeline/multi-agent-refs/features/verify.md +83 -0
  115. package/pipeline/multi-agent-refs/features/visual-evidence.md +19 -19
  116. package/pipeline/multi-agent-refs/features/worktree-finalize.md +6 -6
  117. package/pipeline/multi-agent-refs/issue-jira-triad.md +10 -10
  118. package/pipeline/multi-agent-refs/knowledge.md +11 -11
  119. package/pipeline/multi-agent-refs/multi-repo-integration-build.md +13 -13
  120. package/pipeline/multi-agent-refs/payload-contracts.md +8 -8
  121. package/pipeline/multi-agent-refs/phases/log-format.md +10 -10
  122. package/pipeline/multi-agent-refs/phases/modes.md +30 -30
  123. package/pipeline/multi-agent-refs/phases/operations.md +21 -10
  124. package/pipeline/multi-agent-refs/phases/phase-0-init.md +25 -25
  125. package/pipeline/multi-agent-refs/phases/phase-1-plan.md +599 -0
  126. package/pipeline/multi-agent-refs/phases/{phase-3-dev.md → phase-2-dev.md} +129 -49
  127. package/pipeline/multi-agent-refs/phases/{phase-4-review.md → phase-3-review.md} +225 -107
  128. package/pipeline/multi-agent-refs/phases/{phase-6-commit.md → phase-4-commit.md} +23 -23
  129. package/pipeline/multi-agent-refs/phases/{phase-7-report.md → phase-5-report.md} +29 -29
  130. package/pipeline/multi-agent-refs/phases.md +44 -48
  131. package/pipeline/multi-agent-refs/picker-contract.md +1 -1
  132. package/pipeline/multi-agent-refs/progress-contract.md +6 -6
  133. package/pipeline/multi-agent-refs/readiness-review.md +1 -1
  134. package/pipeline/multi-agent-refs/rules.md +7 -7
  135. package/pipeline/multi-agent-refs/swiftui-guide.md +2 -2
  136. package/pipeline/multi-agent-refs/tracker-contract.md +31 -32
  137. package/pipeline/multi-agent-refs/unattended-contract.md +129 -0
  138. package/pipeline/multi-agent-refs/wiki-capture.md +14 -14
  139. package/pipeline/preferences-template.json +9 -1
  140. package/pipeline/rules/outside-the-pipeline.md +1 -1
  141. package/pipeline/schemas/agent-state.schema.json +50 -50
  142. package/pipeline/schemas/analysis-output.schema.json +2 -2
  143. package/pipeline/schemas/autopilot-config.schema.json +1 -1
  144. package/pipeline/schemas/code-graph.schema.json +1 -1
  145. package/pipeline/schemas/criteria-manifest.schema.json +1 -1
  146. package/pipeline/schemas/dev-critic-output.schema.json +1 -1
  147. package/pipeline/schemas/diff-risk.schema.json +1 -1
  148. package/pipeline/schemas/migrations/prefs-2.4.0-to-2.5.0.mjs +2 -2
  149. package/pipeline/schemas/migrations/prefs-2.6.0-to-2.7.0.mjs +31 -0
  150. package/pipeline/schemas/migrations/state-2.1.0-to-2.2.0.mjs +129 -0
  151. package/pipeline/schemas/phases.json +105 -0
  152. package/pipeline/schemas/plan-todos.schema.json +5 -5
  153. package/pipeline/schemas/planning-output.schema.json +1 -1
  154. package/pipeline/schemas/prefs.schema.json +100 -56
  155. package/pipeline/schemas/reviewer-output.schema.json +3 -3
  156. package/pipeline/schemas/route-config.schema.json +74 -0
  157. package/pipeline/schemas/scope-check.schema.json +1 -1
  158. package/pipeline/schemas/test-gap.schema.json +1 -1
  159. package/pipeline/schemas/token-budget.json +12 -18
  160. package/pipeline/schemas/triage-output.schema.json +6 -6
  161. package/pipeline/scripts/README.md +3 -3
  162. package/pipeline/scripts/_code-graph.mjs +2 -2
  163. package/pipeline/scripts/_run-paths.mjs +372 -0
  164. package/pipeline/scripts/_smoke-root.sh +1 -1
  165. package/pipeline/scripts/aggregate-metrics.mjs +65 -65
  166. package/pipeline/scripts/autopilot-arming.mjs +2 -1
  167. package/pipeline/scripts/autopilot-intake.mjs +2 -1
  168. package/pipeline/scripts/autopilot-runner.mjs +206 -2
  169. package/pipeline/scripts/build-references.mjs +2 -1
  170. package/pipeline/scripts/build-stack-plugins.mjs +10 -2
  171. package/pipeline/scripts/capture-evidence.sh +7 -2
  172. package/pipeline/scripts/capture-flush.sh +8 -8
  173. package/pipeline/scripts/capture-resume.sh +3 -3
  174. package/pipeline/scripts/classify-plan-safety.mjs +3 -2
  175. package/pipeline/scripts/cost-analyze.mjs +600 -0
  176. package/pipeline/scripts/cost-budget-check.mjs +4 -12
  177. package/pipeline/scripts/council-view.mjs +2 -1
  178. package/pipeline/scripts/crush-json.mjs +2 -1
  179. package/pipeline/scripts/diff-explain.mjs +7 -10
  180. package/pipeline/scripts/diff-risk-score.mjs +2 -1
  181. package/pipeline/scripts/doctor.mjs +140 -6
  182. package/pipeline/scripts/evidence-gate.mjs +9 -3
  183. package/pipeline/scripts/feedback-send.mjs +12 -2
  184. package/pipeline/scripts/gc-abandoned.sh +32 -16
  185. package/pipeline/scripts/gc-tmp.sh +1 -1
  186. package/pipeline/scripts/gc-worktrees.sh +12 -5
  187. package/pipeline/scripts/gen-facts.mjs +175 -0
  188. package/pipeline/scripts/gen-mode-dispatch.mjs +32 -37
  189. package/pipeline/scripts/gen-ref-toc.mjs +1 -1
  190. package/pipeline/scripts/github-ssh-setup.sh +64 -7
  191. package/pipeline/scripts/graph-mermaid.mjs +4 -2
  192. package/pipeline/scripts/graph-report.mjs +1 -1
  193. package/pipeline/scripts/jira-attach.sh +1 -1
  194. package/pipeline/scripts/keychain-save.sh +101 -30
  195. package/pipeline/scripts/learn-from-transcripts.mjs +3 -2
  196. package/pipeline/scripts/learning-curve.mjs +36 -31
  197. package/pipeline/scripts/log-metric.sh +17 -4
  198. package/pipeline/scripts/make-manifest.mjs +199 -0
  199. package/pipeline/scripts/memory-save.sh +1 -1
  200. package/pipeline/scripts/migrate-prefs.mjs +24 -6
  201. package/pipeline/scripts/migrate-state.mjs +94 -4
  202. package/pipeline/scripts/phase-banner.sh +26 -22
  203. package/pipeline/scripts/phase-tracker.sh +48 -10
  204. package/pipeline/scripts/plan-coverage-gate.mjs +8 -4
  205. package/pipeline/scripts/pre-commit-check.sh +7 -0
  206. package/pipeline/scripts/pre-push-check.sh +7 -0
  207. package/pipeline/scripts/purge.sh +23 -6
  208. package/pipeline/scripts/render-agent-log-cost.sh +10 -3
  209. package/pipeline/scripts/render-cost-summary.sh +9 -2
  210. package/pipeline/scripts/render-work-summary.sh +14 -7
  211. package/pipeline/scripts/review-file-filter.mjs +5 -3
  212. package/pipeline/scripts/review-scope.mjs +2 -1
  213. package/pipeline/scripts/routine-registry.mjs +2 -1
  214. package/pipeline/scripts/run-aggregator.mjs +26 -20
  215. package/pipeline/scripts/run-metrics.mjs +4 -2
  216. package/pipeline/scripts/runs-index.mjs +353 -0
  217. package/pipeline/scripts/scorecard-snapshot.mjs +178 -0
  218. package/pipeline/scripts/search-logs.sh +18 -0
  219. package/pipeline/scripts/smoke-cross-cli-behavior.sh +6 -6
  220. package/pipeline/scripts/smoke-schema-validation.sh +26 -7
  221. package/pipeline/scripts/test-gap-scan.mjs +2 -1
  222. package/pipeline/scripts/test-integrity-gate.mjs +2 -1
  223. package/pipeline/scripts/token-budget-report.mjs +13 -2
  224. package/pipeline/scripts/triage-memory.mjs +2 -2
  225. package/pipeline/scripts/update-issue-progress.sh +56 -7
  226. package/pipeline/scripts/usage-report.mjs +12 -1
  227. package/pipeline/scripts/validate-analysis-doc.mjs +75 -18
  228. package/pipeline/scripts/validate-code-graph.mjs +6 -3
  229. package/pipeline/scripts/validate-complaint-doc.mjs +2 -1
  230. package/pipeline/scripts/validate-diff-risk.mjs +6 -3
  231. package/pipeline/scripts/validate-planning.mjs +1 -1
  232. package/pipeline/scripts/validate-reviewer.mjs +1 -1
  233. package/pipeline/scripts/validate-state.mjs +45 -5
  234. package/pipeline/scripts/validate-test-gap.mjs +6 -3
  235. package/pipeline/scripts/validate-triage.mjs +6 -4
  236. package/pipeline/scripts/verify-citations.mjs +4 -2
  237. package/pipeline/scripts/verify.mjs +327 -0
  238. package/pipeline/scripts/worktree-finalize.sh +18 -9
  239. package/pipeline/scripts/write-state.mjs +154 -15
  240. package/pipeline/skills/.skill-manifest.json +37 -21
  241. package/pipeline/skills/.skills-index.json +104 -5
  242. package/pipeline/skills/shared/README.md +15 -6
  243. package/pipeline/skills/shared/core/apple-archive-compliance/SKILL.md +2 -2
  244. package/pipeline/skills/shared/core/google-play-compliance/SKILL.md +2 -2
  245. package/pipeline/skills/shared/core/multi-agent/SKILL.md +69 -71
  246. package/pipeline/skills/shared/core/multi-agent-autopilot/SKILL.md +3 -3
  247. package/pipeline/skills/shared/core/multi-agent-channels/SKILL.md +14 -14
  248. package/pipeline/skills/shared/core/multi-agent-diff-explain/SKILL.md +5 -5
  249. package/pipeline/skills/shared/core/multi-agent-graph/SKILL.md +1 -1
  250. package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +25 -23
  251. package/pipeline/skills/shared/core/multi-agent-language/SKILL.md +2 -2
  252. package/pipeline/skills/shared/core/multi-agent-local/SKILL.md +2 -2
  253. package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +8 -8
  254. package/pipeline/skills/shared/core/multi-agent-manual-test/SKILL.md +6 -6
  255. package/pipeline/skills/shared/core/multi-agent-model/SKILL.md +71 -0
  256. package/pipeline/skills/shared/core/multi-agent-refactor/SKILL.md +3 -3
  257. package/pipeline/skills/shared/core/multi-agent-resume/SKILL.md +1 -1
  258. package/pipeline/skills/shared/core/multi-agent-resume-local/SKILL.md +7 -7
  259. package/pipeline/skills/shared/core/multi-agent-route-off/SKILL.md +39 -0
  260. package/pipeline/skills/shared/core/multi-agent-route-on/SKILL.md +76 -0
  261. package/pipeline/skills/shared/core/multi-agent-route-status/SKILL.md +59 -0
  262. package/pipeline/skills/shared/core/multi-agent-setup/SKILL.md +1 -1
  263. package/pipeline/skills/shared/core/multi-agent-status/SKILL.md +35 -11
  264. package/pipeline/skills/shared/core/multi-agent-steer/SKILL.md +2 -2
  265. package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +6 -5
  266. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/package_app.sh +4 -1
  267. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/setup_dev_signing.sh +4 -1
  268. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/sign-and-notarize.sh +2 -1
  269. package/pipeline/skills/skills-index.md +13 -4
  270. package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +0 -263
  271. package/pipeline/multi-agent-refs/phases/phase-2-planning.md +0 -344
  272. package/pipeline/multi-agent-refs/phases/phase-5-test.md +0 -182
package/docs/features.md CHANGED
@@ -4,18 +4,19 @@ Comprehensive list of every feature the pipeline ships. The top-level `README.md
4
4
 
5
5
  ## Core Pipeline
6
6
 
7
- ### 8-Phase Orchestration (0-7)
7
+ ### 6-Phase Orchestration (0-5)
8
8
 
9
9
  ```
10
10
  Phase 0: Init Project selection, branch setup, identity, worktree
11
- Phase 1: Analysis Stack detection, codebase exploration (parallel Explore agents)
12
- Phase 2: Planning Task decomposition, architecture review, user approval
13
- Phase 3: Dev TDD cycle: test → code → build (Sonnet)
14
- Phase 4: Review Deterministic gates + parallel AI review + Fable triage
11
+ Phase 1: Plan Stack detection, codebase exploration (parallel Explore agents),
12
+ task decomposition, architecture review, user approval
13
+ Phase 2: Dev TDD cycle: test → code → build (Sonnet), then the Verify exit
14
+ gate: build · lint · tests · secrets, run once
15
+ Phase 3: Review Parallel AI review + Fable triage against Dev's logs, then the
16
+ optional user test + on-demand device audits
15
17
  (Claude Code: Fable + Opus + Sonnet · Copilot CLI: GPT-5.4 + Opus + Sonnet)
16
- Phase 5: Test Optional manual testing + on-demand device audits
17
- Phase 6: Commit Git commit, push, PR with default reviewers + draft/ready prompt
18
- Phase 7: Report External: Jira comment · Wiki + Figma screenshots · Confluence
18
+ Phase 4: Commit Git commit, push, PR with default reviewers + draft/ready prompt
19
+ Phase 5: Report External: Jira comment · Wiki + Figma screenshots · Confluence
19
20
  Internal: agent-log.md + Quality & Metrics + knowledge + memory
20
21
  ```
21
22
 
@@ -28,7 +29,7 @@ Each phase reads its own spec file under `pipeline/multi-agent-refs/phases/phase
28
29
  | `autopilot` | Skip all confirmation prompts; still fails safe on review blockers + build retries. |
29
30
  | `--local` | Answers the workspace question up front: no worktree, work directly in `$PROJECT_ROOT` on a local branch. |
30
31
 
31
- Neither the workspace nor the depth is a flag. **Where the branch lives** is asked at Phase 0 Step 5b - worktree (`.worktrees/{id}/`, your checkout untouched) or local (the project root, which drops Phase 5 because the user-test gate checks the change out of a worktree and there is none). `:local` and `--local` state it in advance; every autopilot entry resolves it to a worktree and never asks, because an unattended run commits and pushes from wherever it stands and doing that in the user's own checkout is what worktrees exist to prevent. `state.workspaceSource` records who decided - `localMode: false` alone is both "the user chose a worktree" and "nothing asked".
32
+ Neither the workspace nor the depth is a flag. **Where the branch lives** is asked at Phase 0 Step 5b - worktree (`.worktrees/{id}/`, your checkout untouched) or local (the project root, which drops the user test because that gate checks the change out of a worktree and there is none). `:local` and `--local` state it in advance; every autopilot entry resolves it to a worktree and never asks, because an unattended run commits and pushes from wherever it stands and doing that in the user's own checkout is what worktrees exist to prevent. `state.workspaceSource` records who decided - `localMode: false` alone is both "the user chose a worktree" and "nothing asked".
32
33
 
33
34
  Depth is not a flag either. `/multi-agent` and `/multi-agent:local` ask Full or Short at Phase 0 Step 7.5, recommending from the detected `taskType`; Short strips to Init → Dev(Opus self-contained) → Review → Test → Commit → Report. Autopilot never asks and always runs Full - "fast plus unattended" was removed in v16.0.0, because something has to choose when nobody is asked and unattended is the worst place to drop analysis and planning.
34
35
 
@@ -46,7 +47,7 @@ Uninstall preserves the whole layer - tokens, the reader that opens them, the ma
46
47
 
47
48
  A deterministic, LLM-free map of what a repo declares and what refers to what, extracted by regex over comment-stripped source into `~/.claude/knowledge/<project>/code-graph.json`. Four stacks build today (Swift, Kotlin/Java, TypeScript/JavaScript, Python); each is one rules file, and the engine is the same for all of them. Zero runtime dependencies, zero API cost, read-only on the repo.
48
49
 
49
- Phase 1 queries it to hand Explore a ranked starting file set instead of a full scan, and Phase 7 rebuilds it after the branch changed code - a rebuild is seconds, so staleness is a `baseCommit` comparison rather than a date heuristic. Off by default behind `prefs.global.codeGraph.enabled`; with it off the pipeline behaves exactly as before.
50
+ Phase 1 queries it to hand Explore a ranked starting file set instead of a full scan, and Phase 5 rebuilds it after the branch changed code - a rebuild is seconds, so staleness is a `baseCommit` comparison rather than a date heuristic. Off by default behind `prefs.global.codeGraph.enabled`; with it off the pipeline behaves exactly as before.
50
51
 
51
52
  `GRAPH_REPORT.md` also ends with **Symbols nothing else references**: symbols no other file in the repo names, split from the ones referenced only by their own tests. Candidates, never verdicts - the extractor is regex, not a parser, so the four classes that could not carry a reference edge either way (a kind outside the stack's `referenceKinds`, a name declared twice, a nested declaration, a test file) are counted and excluded rather than listed, and nothing gates on the result.
52
53
 
@@ -78,9 +79,9 @@ Stack skill sets ship as versioned plugins in the `multi-agent-plugins` marketpl
78
79
  /multi-agent:stack all # every stack plugin
79
80
  ```
80
81
 
81
- ### Package Manager Resolution (Phase 3, node-shaped stacks)
82
+ ### Package Manager Resolution (Phase 2, node-shaped stacks)
82
83
 
83
- Phase 3's web test arm and its build step used to type `npm`. A repo on pnpm, yarn or bun then failed in Phase 3 - with a worktree and a branch already created - or, worse, npm resolved against a lock file it does not own and the run continued on a tree the repo's own tooling would never have produced.
84
+ Phase 2's web test arm and its build step used to type `npm`. A repo on pnpm, yarn or bun then failed in Phase 2 - with a worktree and a branch already created - or, worse, npm resolved against a lock file it does not own and the run continued on a tree the repo's own tooling would never have produced.
84
85
 
85
86
  `scripts/package-manager.mjs` resolves it from the repo instead: `$MA_PACKAGE_MANAGER`, then `package.json#packageManager`, then a lock file, then npm - reported AS a default, never as evidence, because "npm because nothing said otherwise" and "npm because the repo committed a package-lock" are different answers. Node core only (ADR-0004): a resolver that shelled out would need a working install of the tool it is identifying. The walk goes up to the directory holding `.git` and stops there, so a monorepo's root lock file is found and a stray one in a home directory is not. Two lock files means a migration left one behind: the newest wins and both are named.
86
87
 
@@ -121,15 +122,15 @@ Result persisted to `agent-state.taskType`:
121
122
 
122
123
  | Type | Downstream effects |
123
124
  | ----------- | ----------------------------------------------------------------------------- |
124
- | `component` | Phase 3 dispatches to the marketplace component plugin (create-component) with SubPhase reporting |
125
- | `bugfix` | Phase 4 emphasizes test coverage + regression; Phase 6 uses `fix(...)` prefix |
126
- | `feature` | Standard TDD flow; Phase 6 uses `feat(...)` prefix |
127
- | `refactor` | Phase 4 emphasizes behavior preservation; Phase 6 uses `refactor(...)` prefix |
128
- | `chore` | Lightweight flow; Phase 6 uses `chore(...)` prefix |
125
+ | `component` | Phase 2 dispatches to the marketplace component plugin (create-component) with SubPhase reporting |
126
+ | `bugfix` | Phase 3 emphasizes test coverage + regression; Phase 4 uses `fix(...)` prefix |
127
+ | `feature` | Standard TDD flow; Phase 4 uses `feat(...)` prefix |
128
+ | `refactor` | Phase 3 emphasizes behavior preservation; Phase 4 uses `refactor(...)` prefix |
129
+ | `chore` | Lightweight flow; Phase 4 uses `chore(...)` prefix |
129
130
 
130
131
  ### SubPhase Convention
131
132
 
132
- When a specialized skill takes over a main pipeline phase, progress is reported as SubPhases (e.g. `SubPhase 3.0: Init`, `SubPhase 3.1: Gather`). The top-level pipeline stays fixed at 8 phases (0-7) - specialized work slots into its parent phase without inflating the count.
133
+ When a specialized skill takes over a main pipeline phase, progress is reported as SubPhases (e.g. `SubPhase 3.0: Init`, `SubPhase 3.1: Gather`). The top-level pipeline stays fixed at 6 phases (0-5) - specialized work slots into its parent phase without inflating the count.
133
134
 
134
135
  ## PR & Review Flow
135
136
 
@@ -141,14 +142,14 @@ When a specialized skill takes over a main pipeline phase, progress is reported
141
142
 
142
143
  ### Draft vs Ready Prompt
143
144
 
144
- Phase 6 asks `DRAFT or READY?` before creating the PR and persists the choice in `prefs.projects[].defaultPrMode`.
145
+ Phase 4 asks `DRAFT or READY?` before creating the PR and persists the choice in `prefs.projects[].defaultPrMode`.
145
146
 
146
147
  - Bitbucket: `draft: true` flag (DC 8.x+) with `[DRAFT]` title fallback for older servers.
147
148
  - GitHub: `gh pr create --draft` + `gh pr ready` for promotion.
148
149
 
149
150
  ### `channels` Command
150
151
 
151
- Multi-channel reporter - Phase 7 delegates to it, and it's also invocable post-hoc for fixes closed outside the pipeline:
152
+ Multi-channel reporter - Phase 5 delegates to it, and it's also invocable post-hoc for fixes closed outside the pipeline:
152
153
 
153
154
  ```bash
154
155
  /multi-agent:channels # current branch, current PR
@@ -170,7 +171,7 @@ Never auto-closes issues - uses `Ref: #N` / `Related: #N` / `See: PROJ-12345`, n
170
171
 
171
172
  ## Review Quality
172
173
 
173
- ### Deterministic Gates (Phase 4 Step 1)
174
+ ### Deterministic Gates (Phase 3 Step 1)
174
175
 
175
176
  Cheap, objective checks run BEFORE any AI token is spent:
176
177
 
@@ -181,23 +182,23 @@ Cheap, objective checks run BEFORE any AI token is spent:
181
182
 
182
183
  If any gate fails, fix first. Don't waste AI tokens reviewing broken code.
183
184
 
184
- ### Analysis Document Review (Phase 3.2 + 3.3)
185
+ ### Analysis Document Review (Phase 2.2 + 2.3)
185
186
 
186
187
  `/multi-agent:analysis` published behind a structural validator alone until v16.12.0: nothing read the
187
- document before it reached Confluence. Phase 3.2 now runs the same reviewer set and triage a code diff
188
+ document before it reached Confluence. Phase 2.2 now runs the same reviewer set and triage a code diff
188
189
  gets, on the draft, before the destination is even chosen. Its first question is what the run skipped -
189
190
  an input declared missing that nothing searched for, an open question about evidence nobody read, a gap
190
191
  with no owner, a scope call made without asking. A blocking finding returns to synthesis with dispatch
191
192
  closed; it never becomes an open question, because "the document is wrong" is not something to ask the
192
193
  reader.
193
194
 
194
- Phase 3.3 then sorts what is left: reachable evidence is searched (never asked about), decisions the
195
+ Phase 2.3 then sorts what is left: reachable evidence is searched (never asked about), decisions the
195
196
  user owns are asked with `AskUserQuestion`, and only genuinely external gaps enter the document as
196
197
  `AS-NN` rows with an owner. A gap carrying neither a `searched, not found` nor an `asked, external`
197
198
  stamp fails the dispatch gate. Autopilot runs both phases; only the asking degrades, into rows stamped
198
199
  `autopilot: could not ask`.
199
200
 
200
- ### CLI-Aware Parallel Review + Fable Triage (Phase 4 Steps 2-3)
201
+ ### CLI-Aware Parallel Review + Fable Triage (Phase 3 Steps 2-3)
201
202
 
202
203
  | Reviewer | Model | Focus | Where it runs |
203
204
  | ---------- | ------------------- | --------------------------------- | -------------------- |
@@ -207,7 +208,7 @@ stamp fails the dispatch gate. Autopilot runs both phases; only the asking degra
207
208
 
208
209
  The reviewer set is **CLI-aware**: Claude Code dispatches 3 reviewers in parallel (Fable + Opus + Sonnet - Opus fills the slot GPT-5.4 takes elsewhere); Copilot CLI dispatches all 3. Each returns structured JSON for deterministic aggregation. Cross-model diversity catches blind spots that any single model family would miss.
209
210
 
210
- **Fable Triage** (Phase 4 Step 3, Opus on Copilot CLI): Evaluates merged raw findings against task scope. Classifies each as `accepted` (fix now), `deferred` (out of scope, log for later), or `rejected` (false positive / noise). Only triage-accepted blocking items loop back to Phase 3.
211
+ **Fable Triage** (Phase 3 Step 3, Opus on Copilot CLI): Evaluates merged raw findings against task scope. Classifies each as `accepted` (fix now), `deferred` (out of scope, log for later), or `rejected` (false positive / noise). Only triage-accepted blocking items loop back to Phase 3.
211
212
 
212
213
  ### Runtime Triage Validator
213
214
 
@@ -224,13 +225,13 @@ After triage returns, output is validated by `validate-triage.mjs`:
224
225
 
225
226
  If triage returns `approved: false` but has no blocking items, the validator forces `approved: true`. Conversely, if `approved: true` but blocking items exist, it forces `approved: false`. Hardened with an `if`/`then` constraint in the schema itself.
226
227
 
227
- ### Verify-by-Test Triage (Phase 4 Step 3.7, opt-in)
228
+ ### Verify-by-Test Triage (Phase 3 Step 3.7, opt-in)
228
229
 
229
- A triage verdict is a judgment call; a failing repro test is proof. When `prefs.global.verifyByTest.enabled` is on, one verifier agent (default Sonnet) writes a minimal repro test per accepted blocking finding (cap: `maxFindings`=3) and runs only that test. Fails as predicted -> finding confirmed, the repro test becomes the Phase 3 rework RED test. Passes under `evidence-gate.mjs` -> finding downgraded to `deferred`. Compile error / timeout -> `inconclusive`, judgment stands. Timeout-bounded, never blocks. Full spec: `refs/features/verify-by-test.md`.
230
+ A triage verdict is a judgment call; a failing repro test is proof. When `prefs.global.verifyByTest.enabled` is on, one verifier agent (default Sonnet) writes a minimal repro test per accepted blocking finding (cap: `maxFindings`=3) and runs only that test. Fails as predicted -> finding confirmed, the repro test becomes the Phase 2 rework RED test. Passes under `evidence-gate.mjs` -> finding downgraded to `deferred`. Compile error / timeout -> `inconclusive`, judgment stands. Timeout-bounded, never blocks. Full spec: `refs/features/verify-by-test.md`.
230
231
 
231
232
  ### Immutable-Test Rule + `test_lines_removed` Signal
232
233
 
233
- Existing tests are immutable during a task: deleting, renaming, or weakening an assertion to reach green is a violation (`refs/rules.md`, Phase 3 GREEN step). A test changes only when the task changes the spec it encodes, named in the commit body. Deterministic backstop: `diff-risk-score.mjs` emits `test_lines_removed` (w=3.0) for any test-classified file whose diff removes more lines than it adds.
234
+ Existing tests are immutable during a task: deleting, renaming, or weakening an assertion to reach green is a violation (`refs/rules.md`, Phase 2 GREEN step). A test changes only when the task changes the spec it encodes, named in the commit body. Deterministic backstop: `diff-risk-score.mjs` emits `test_lines_removed` (w=3.0) for any test-classified file whose diff removes more lines than it adds.
234
235
 
235
236
  ### Update Check at Run Start
236
237
 
@@ -244,7 +245,7 @@ Phase 0 Step 0.6. Once per `ttlHours` window (cached, 3s-bounded curl to the npm
244
245
 
245
246
  Every phase transition appends a `## Handoff` block (Done / Remaining / Decisions / Open findings / Next) to `agent-log.md` - orchestrator-written from existing state, no LLM call. `/multi-agent:resume` and post-`/compact` re-grounding read the latest handoff first, so long runs re-enter from durable artifacts instead of conversation memory (fresh-context discipline from Anthropic's long-running-agent harness guidance).
246
247
 
247
- ### Accessibility Code Review (Phase 4 Step 1.5)
248
+ ### Accessibility Code Review (Phase 3 Step 1.5)
248
249
 
249
250
  If changes include UI files, reviewers check for:
250
251
 
@@ -256,16 +257,16 @@ Pure code analysis - no simulator needed. Device-level audits run in Phase 5 whe
256
257
 
257
258
  ### Status Enforcement
258
259
 
259
- Phase 3 treats the issue-tracker status update as a required step with a post-mutation verify step that re-reads the field and retries once on silent `VALIDATION` failures (e.g. stale Projects V2 option IDs after a board rebuild).
260
+ Phase 2 treats the issue-tracker status update as a required step with a post-mutation verify step that re-reads the field and retries once on silent `VALIDATION` failures (e.g. stale Projects V2 option IDs after a board rebuild).
260
261
 
261
262
  ## Safety & Hygiene
262
263
 
263
264
  - **Pre-Commit Secret Detection** (12 patterns): `PreToolUse` hook scans staged files for API keys/tokens, AWS access keys, private keys, `.env` files, service account JSON. Commit **blocked** if found.
264
265
  - **Read-Size Gate** (opt-in, `prefs.global.bulkRead.mode`): a `PreToolUse` hook inspects `Read` and the shell commands that read a file whole. In `observe` it only logs what it would have caught - the baseline you measure before routing anything. In `enforce` a file over `minLines` (default 350) is blocked and delegated to a haiku-rung worker (`bulk-read.sh`), which returns a line-numbered summary so the follow-up is a bounded `Read(offset:limit:)` instead of the whole file; the full text is parked under `.multi-agent/refs/`. The development phase and any file the run has already touched are exempt, because Claude Code's `Edit` requires its own `Read` first.
265
- - **Capture Hooks** (`SessionEnd`, `PreCompact`, `SessionStart`): every durable write used to live in Phase 7, the phase a run is least likely to reach. `SessionEnd` flushes a run that never got there; `PreCompact` flushes before an auto-compaction summarizes a long phase mid-flight, which is the same loss one level down; `SessionStart` prints at most two lines about an unfinished run. None calls a model, none reads a payload, and all exit 0 on every path - a hook that fails a session over bookkeeping is worse than the bookkeeping.
266
+ - **Capture Hooks** (`SessionEnd`, `PreCompact`, `SessionStart`): every durable write used to live in Phase 5, the phase a run is least likely to reach. `SessionEnd` flushes a run that never got there; `PreCompact` flushes before an auto-compaction summarizes a long phase mid-flight, which is the same loss one level down; `SessionStart` prints at most two lines about an unfinished run. None calls a model, none reads a payload, and all exit 0 on every path - a hook that fails a session over bookkeeping is worse than the bookkeeping.
266
267
  - **Operational Reporting** (`prefs.global.usageLog`): coarse run metadata - task id, phase, status, durations, token counts - and never prompts, code, diffs or absolute paths. The per-machine token is REQUESTED from the endpoint by `usage-register.mjs` (setup, update, and the Phase 0 exit gate as a backstop), is write-only, and lives in the OS credential store; prefs hold only the entry name and the switch. `usageLog.optOut: true` blocks registration permanently and is checked before the network call. An unreachable endpoint leaves reporting off with one line and exit 0 - a run is never failed over bookkeeping.
267
268
  - **Build Queue**: All `xcodebuild` calls acquire a lock. Each worktree uses own `-derivedDataPath`. Stale locks auto-clean after 15 min. Non-Xcode builds don't need the lock.
268
- - **Context Management**: `CLAUDE_AUTOCOMPACT_PCT_OVERRIDE=65` - compaction at 65% usage (prevents degradation in 8-phase sessions).
269
+ - **Context Management**: `CLAUDE_AUTOCOMPACT_PCT_OVERRIDE=65` - compaction at 65% usage (prevents degradation in 6-phase sessions).
269
270
  - **3-Iteration Hard Kill**: Any retry loop stops after 3 attempts, then pauses for user. No infinite loops.
270
271
 
271
272
  ## Testing & Quality
@@ -300,51 +301,51 @@ Per-phase token budgets prevent runaway sessions. If a phase exceeds its budget,
300
301
  - **Cost Telemetry**: Per-phase token cost tracking (`tokens_in`, `tokens_out`, `model`, `duration_ms`). Omitted fields handled gracefully.
301
302
  - **Phase Tracker**: Cross-CLI visual progress (current phase, elapsed time, iteration count).
302
303
  - **Phase Banner**: Terminal UI for phase transitions with Unicode box-drawing characters.
303
- - **Per-task Cost Breakdown in agent-log.md**: Phase 7 appends a 4-column block (Phase · Model · Tokens in/out · Est. USD) to every run's `agent-log.md`. Sourced from `phase-tracker.sh tokens` accumulators × `cost-table.json` prices. Independent of the channels-side `reportContent.costSummary` toggle. The `LOG_METRIC_FORWARD_TO_TRACKER=1` env flag mirrors `tokens_in`/`tokens_out`/`model` from `log-metric.sh` into the tracker so JSONL metrics and the cost block stay in sync from one call site.
304
+ - **Per-task Cost Breakdown in agent-log.md**: Phase 5 appends a 4-column block (Phase · Model · Tokens in/out · Est. USD) to every run's `agent-log.md`. Sourced from `phase-tracker.sh tokens` accumulators × `cost-table.json` prices. Independent of the channels-side `reportContent.costSummary` toggle. The `LOG_METRIC_FORWARD_TO_TRACKER=1` env flag mirrors `tokens_in`/`tokens_out`/`model` from `log-metric.sh` into the tracker so JSONL metrics and the cost block stay in sync from one call site.
304
305
 
305
306
  ### Diff Risk Scoring
306
307
 
307
- `pipeline/scripts/diff-risk-score.mjs` runs at Phase 4 Step 1.75 - before reviewer dispatch. Heuristic, deterministic, sub-second, no LLM. Top-N risk-ranked files inject into each reviewer's prompt as a `${PRIORITY_FILES}` block; reviewers read those files first but still review the entire diff.
308
+ `pipeline/scripts/diff-risk-score.mjs` runs at Phase 3 Step 1.75 - before reviewer dispatch. Heuristic, deterministic, sub-second, no LLM. Top-N risk-ranked files inject into each reviewer's prompt as a `${PRIORITY_FILES}` block; reviewers read those files first but still review the entire diff.
308
309
 
309
310
  Signals + weights: `security_path` ×3, `migration` ×4, `public_api` ×2, `no_test_change` ×2.5, `test_lines_removed` ×3 (test file shrinks - immutable-test backstop), `complexity_delta` ×1.5, `ui_critical` ×1.5, `loc_changed` ×1. Toggle via `prefs.global.diffRiskAdvisory` (default ON).
310
311
 
311
312
  ### Test Gap Detection
312
313
 
313
- `pipeline/scripts/test-gap-scan.mjs` runs at Phase 5 Step 0. Walks the diff for newly added public symbols and reports those with no paired test. Stack-specific rules ship for iOS, Android, Python, Node.js. iOS Views and Android `@Composable` symbols default to `important`; other public API additions to `suggestion`. Optional gating via `prefs.testGap.blockingThreshold` - when set, the report becomes a Phase 4 rework finding once `important + blocking` count exceeds the threshold.
314
+ `pipeline/scripts/test-gap-scan.mjs` runs at Phase 3 Step 0. Walks the diff for newly added public symbols and reports those with no paired test. Stack-specific rules ship for iOS, Android, Python, Node.js. iOS Views and Android `@Composable` symbols default to `important`; other public API additions to `suggestion`. Optional gating via `prefs.testGap.blockingThreshold` - when set, the report becomes a Phase 4 rework finding once `important + blocking` count exceeds the threshold.
314
315
 
315
316
  ### Visual Evidence (UI changes)
316
317
 
317
318
  A UI change carries its own picture. `state.visualEvidence.required` is decided mechanically from `taskType` plus the changed-file list, never from a reading of the task.
318
319
 
319
- **Stills.** The "before" is the reporter's own ticket attachment, harvested in Phase 0; the pipeline never rebuilds the old state to photograph it. The "after" is captured in Phase 3 right after the build goes green, not Phase 5, which autopilot and both local modes drop. `capture-evidence.sh` cleans the status bar and downscales to 1242px so two captures of one screen differ by the change and not by the clock.
320
+ **Stills.** The "before" is the reporter's own ticket attachment, harvested in Phase 0; the pipeline never rebuilds the old state to photograph it. The "after" is captured in Phase 2 right after the build goes green, not in the user test, which autopilot and both local modes drop. `capture-evidence.sh` cleans the status bar and downscales to 1242px so two captures of one screen differ by the change and not by the clock.
320
321
 
321
- **The flow video rides on a test run.** `probe-evidence-capability.sh` measures the UI test target, the tests matching this change, the device, the recorder and the MCP registration; Phase 0 Step 7.7 then asks the depth with the options built from that measurement, and a closed option keeps its row and states why. Tier 1 runs the repo's own UI test and records around it, tier 2 drives the flow through `agent_run_steps`, tier 3 records nothing and says so. The tier is re-checked before the recording starts, because a simulator booted at intake can be gone by Phase 3.
322
+ **The flow video rides on a test run.** `probe-evidence-capability.sh` measures the UI test target, the tests matching this change, the device, the recorder and the MCP registration; Phase 0 Step 7.7 then asks the depth with the options built from that measurement, and a closed option keeps its row and states why. Tier 1 runs the repo's own UI test and records around it, tier 2 drives the flow through `agent_run_steps`, tier 3 records nothing and says so. The tier is re-checked before the recording starts, because a simulator booted at intake can be gone by Phase 2.
322
323
 
323
324
  UI test detection keys on `XCUIApplication` rather than on a folder named `*UITests`: in a real app the overwhelming majority of files under such a path are snapshot tests, which never launch the app and would produce a still frame filed as a flow.
324
325
 
325
- **Where it lands.** Jira takes both stills and video as attachments. With no Jira the stills go to an orphan `evidence/<task-id>` branch and the PR body embeds them, or links them with a blob permalink when the repo is private (GitHub's image proxy has no credentials for a private repo, and a broken image reads as missing evidence). Phase 6 blocks when a required artefact is neither published nor explained; the gate is against silence, not against an honest "the ticket carries no image".
326
+ **Where it lands.** Jira takes both stills and video as attachments. With no Jira the stills go to an orphan `evidence/<task-id>` branch and the PR body embeds them, or links them with a blob permalink when the repo is private (GitHub's image proxy has no credentials for a private repo, and a broken image reads as missing evidence). Phase 4 blocks when a required artefact is neither published nor explained; the gate is against silence, not against an honest "the ticket carries no image".
326
327
 
327
328
  Toggle via `prefs.global.visualEvidence.enabled` (default ON), `visualEvidence.githubHost`, `visualEvidence.maxAttachmentMb`, `visualEvidence.maxVideoSeconds`, `prefs.global.testDepth.default`.
328
329
 
329
330
  ### Triage Memory
330
331
 
331
- Per-repo append-only JSONL corpus at `~/.claude/memory/multi-agent/<repo-slug>/triage-corpus.jsonl`. Phase 7 ingests every triage output (idempotent), Phase 1 enriches the analysis with similar past tasks, Phase 4 triage attaches prior-art hits to each raw finding with an explicit bias hedge. Token-overlap recall, zero deps, Node-18-compatible. `/multi-agent:search "<text>" --semantic` routes the query to the corpus instead of agent-log grep. Toggle via `prefs.global.priorArtEnrichment.enabled` (default ON).
332
+ Per-repo append-only JSONL corpus at `~/.claude/memory/multi-agent/<repo-slug>/triage-corpus.jsonl`. Phase 5 ingests every triage output (idempotent), Phase 1 enriches the analysis with similar past tasks, Phase 3 triage attaches prior-art hits to each raw finding with an explicit bias hedge. Token-overlap recall, zero deps, Node-18-compatible. `/multi-agent:search "<text>" --semantic` routes the query to the corpus instead of agent-log grep. Toggle via `prefs.global.priorArtEnrichment.enabled` (default ON).
332
333
 
333
334
  ## Learning
334
335
 
335
336
  ### Knowledge Base (per project)
336
337
 
337
- Incremental learning. Phase 7 captures architecture, patterns, gotchas, and decisions into `$HOME/.claude/knowledge/{project}/`. Phase 1 reads it on the next run. Token cost decreases over time as the base grows.
338
+ Incremental learning. Phase 5 captures architecture, patterns, gotchas, and decisions into `$HOME/.claude/knowledge/{project}/`. Phase 1 reads it on the next run. Token cost decreases over time as the base grows.
338
339
 
339
340
  ### Memory Capture (cross-session)
340
341
 
341
- Pipeline learns behavioral signals (feedback corrections, project constraints, external references). Phase 7 saves, Phase 1 injects. Max 3 new memories per run. Merge-over-duplicate. Stale memories verified before use.
342
+ Pipeline learns behavioral signals (feedback corrections, project constraints, external references). Phase 5 saves, Phase 1 injects. Max 3 new memories per run. Merge-over-duplicate. Stale memories verified before use.
342
343
 
343
344
  **What does NOT go in memory**: architecture, code patterns, build gotchas, design decisions - those belong in the knowledge base.
344
345
 
345
346
  ### Lesson Diagnosis (Reflexion)
346
347
 
347
- Phase 4's lesson-memory loop records the causal root cause of each fix (`--diagnosis`), not just the outcome: the verbal "why" that prevents recurrence (Reflexion). `learnings-ledger.mjs brief` renders it as `(why: ...)` back into Phase 1 + triage on the next run, so the reason re-enters the loop, not only the symptom.
348
+ Phase 3's lesson-memory loop records the causal root cause of each fix (`--diagnosis`), not just the outcome: the verbal "why" that prevents recurrence (Reflexion). `learnings-ledger.mjs brief` renders it as `(why: ...)` back into Phase 1 + triage on the next run, so the reason re-enters the loop, not only the symptom.
348
349
 
349
350
  ### Corpus Freshness Gate
350
351
 
@@ -366,7 +367,7 @@ Turn a recurring, project-specific job into a first-class `/multi-agent:<name>`
366
367
 
367
368
  ### Figma / Component Generation (dispatched to marketplace plugins)
368
369
 
369
- Component + Figma-to-code work is no longer bundled in this repo. When Phase 0 classifies a task as `component`, Phase 3 dispatches it to the per-stack marketplace plugins (`ai-ios-toolkit` / `ai-android-toolkit` in the `multi-agent-plugins` marketplace) via the Skill tool. The plugin's component skill generates `{Name}Configuration.swift`, `{Name}View.swift`, `{Name}+Modifiers.swift`, `{Name}.figma.swift`, and `FIGMA.md` with a variant matrix, then runs a 14-item pre-commit checklist covering design tokens, accessibility, tests, and Code Connect.
370
+ Component + Figma-to-code work is no longer bundled in this repo. When Phase 0 classifies a task as `component`, Phase 2 dispatches it to the per-stack marketplace plugins (`ai-ios-toolkit` / `ai-android-toolkit` in the `multi-agent-plugins` marketplace) via the Skill tool. The plugin's component skill generates `{Name}Configuration.swift`, `{Name}View.swift`, `{Name}+Modifiers.swift`, `{Name}.figma.swift`, and `FIGMA.md` with a variant matrix, then runs a 14-item pre-commit checklist covering design tokens, accessibility, tests, and Code Connect.
370
371
 
371
372
  The plugin's cross-cutting integration skills feed component detection + implementation when the design triggers them (content: form / price / ui-patterns; interaction: navigation / overlays / bottom-sheets). Each is native-SwiftUI-first and reads project specifics (token namespaces, component paths, UI systems) from `figma-config`, including the optional `ui.navigationSystem` / `ui.overlaySystem` / `ui.sheetSystem` hooks (absent -> stock SwiftUI), so the same capabilities work on any SwiftUI codebase. The plugin's evolve-component skill reconciles an existing component against current Figma (drift-heal) and additively extends it, behind a human gate.
372
373
 
@@ -376,20 +377,20 @@ Automated visual testing and compliance audits via direct Bash (no MCP server de
376
377
 
377
378
  | Audit | When | Command |
378
379
  | --------------------- | ------------------- | ---------------------------- |
379
- | iOS Accessibility | Phase 5, on request | `swift ui-tree-dumper.swift` |
380
- | Android Accessibility | Phase 5, on request | `adb shell uiautomator dump` |
381
- | iOS Biometric | Phase 5, auth flow | `xcrun simctl keychain` |
382
- | Android Launch Time | Phase 5, perf | `adb shell am start -W` |
383
- | iOS Archive | Phase 6, release | `codesign`, `plutil`, `nm` |
384
- | Android APK | Phase 6, release | `aapt2`, `apksigner` |
380
+ | iOS Accessibility | Phase 3, on request | `swift ui-tree-dumper.swift` |
381
+ | Android Accessibility | Phase 3, on request | `adb shell uiautomator dump` |
382
+ | iOS Biometric | Phase 3, auth flow | `xcrun simctl keychain` |
383
+ | Android Launch Time | Phase 3, perf | `adb shell am start -W` |
384
+ | iOS Archive | Phase 4, release | `codesign`, `plutil`, `nm` |
385
+ | Android APK | Phase 4, release | `aapt2`, `apksigner` |
385
386
 
386
387
  Audits are **on-demand** - triggered by user, never automatic.
387
388
 
388
389
  ### Jira + Confluence
389
390
 
390
- - Phase 3: transition issue to `In Progress` (verified post-mutation).
391
- - Phase 7: post analysis + test scenarios as Jira comment (Turkish by default, configurable).
392
- - Phase 7 (optional): create Confluence page under chosen parent, cached per project.
391
+ - Phase 2: transition issue to `In Progress` (verified post-mutation).
392
+ - Phase 5: post analysis + test scenarios as Jira comment (Turkish by default, configurable).
393
+ - Phase 5 (optional): create Confluence page under chosen parent, cached per project.
393
394
 
394
395
  ### Keychain
395
396
 
@@ -13,7 +13,7 @@ node pipeline/scripts/aggregate-metrics.mjs
13
13
  # Markdown table (for PR descriptions, wikis, dashboards)
14
14
  node pipeline/scripts/aggregate-metrics.mjs --markdown
15
15
 
16
- # JSON (for machine consumption - Phase 7 report uses this)
16
+ # JSON (for machine consumption - Phase 5 report uses this)
17
17
  node pipeline/scripts/aggregate-metrics.mjs --json
18
18
 
19
19
  # Filtered - only recent runs
@@ -77,7 +77,7 @@ _Source: ~/.claude/logs/multi-agent/metrics.jsonl · Events: 421 (0 parse errors
77
77
  - **`cycles per task avg` > 2.0** - triage is rejecting too many real findings
78
78
  or Phase 3 isn't converging. Inspect the edge-cases table.
79
79
  - **Most-common edge case = `over_rejection_guard_tripped`** - the triage
80
- prompt lost scope context. Look at `phase-4-review.md:57-91`.
80
+ prompt lost scope context. Look at `phase-3-review.md:57-91`.
81
81
  - **`p95` much higher than `avg`** - a few tasks are looping 3+ times. Usually
82
82
  means one of: bad acceptance criteria in Phase 2, tests that flake, or
83
83
  environment-dependent build failures.
@@ -103,11 +103,11 @@ Current totals (v3.5.0):
103
103
  Lazy loading keeps these off the model's context until each phase actually
104
104
  runs - the full 14 k total is never loaded at once.
105
105
 
106
- ## Embedding Metrics in Phase 7 Reports
106
+ ## Embedding Metrics in Phase 5 Reports
107
107
 
108
- Phase 7 automatically calls `aggregate-metrics.mjs --json` with the current
108
+ Phase 5 automatically calls `aggregate-metrics.mjs --json` with the current
109
109
  task id and embeds the last 30 days' summary into the report body. See the
110
- template in `phase-7-report.md`.
110
+ template in `phase-5-report.md`.
111
111
 
112
112
  ## Disabling Metrics
113
113
 
@@ -10,8 +10,8 @@ the answer across `modes.md`, `operations.md`, and the individual phase specs.
10
10
  | ------------------------------------------------------- | ----------------------------------------------- |
11
11
  | Pipeline paused/halted mid-run | [Resume a paused task](#resume-a-paused-task) |
12
12
  | Phase 3 build failed > 3 times | [Build retry exhausted](#build-retry-exhausted) |
13
- | Phase 4 triage returned exit 1 (invalid JSON) twice | [Triage fallback](#triage-fallback) |
14
- | Phase 4 triage returned exit 2 (over-rejection) | [Over-rejection](#over-rejection-guard) |
13
+ | Phase 3 triage returned exit 1 (invalid JSON) twice | [Triage fallback](#triage-fallback) |
14
+ | Phase 3 triage returned exit 2 (over-rejection) | [Over-rejection](#over-rejection-guard) |
15
15
  | Worktree already exists / dirty | [Worktree collisions](#worktree-collisions) |
16
16
  | `agent-state.json` corrupt or unreadable | [State corruption](#state-corruption) |
17
17
  | Wrong git identity committed | [Identity rewind](#identity-rewind) |
@@ -58,7 +58,7 @@ Path forward:
58
58
  environment and `resume` - the 3-retry counter resets.
59
59
  3. If the failure is logic (compile error in the generated code), edit the
60
60
  offending file in the worktree, then `resume` - Phase 3 re-runs build.
61
- 4. If the task itself is wrong-shaped (Phase 2 plan is infeasible), `kill` and
61
+ 4. If the task itself is wrong-shaped (Phase 1 plan is infeasible), `kill` and
62
62
  restart with a better-scoped prompt.
63
63
 
64
64
  ---
@@ -79,7 +79,7 @@ If exit 1 fires twice (fallback path):
79
79
  1. The pipeline logs `Phase 4: triage failed twice - fallback, all findings accepted as blocking`.
80
80
  2. All raw findings loop back into Phase 3 as if they were all real blockers.
81
81
  3. This is intentionally conservative - we'd rather over-fix than skip
82
- something real. You can manually mark noise in the Phase 6 PR description.
82
+ something real. You can manually mark noise in the Phase 4 PR description.
83
83
 
84
84
  ## Over-Rejection Guard
85
85
 
@@ -216,7 +216,7 @@ cat .worktrees/{id}/agent-state.json | grep status
216
216
  ## Instruction Fallback
217
217
 
218
218
  `state.instructionDriven=true` but the file at `state.instructionFiles.commit`
219
- is missing on disk. Phase 6 logs an error, sets
219
+ is missing on disk. Phase 4 logs an error, sets
220
220
  `state.instructionDrivenFallback=true`, and uses the standard commit path.
221
221
 
222
222
  This is rarely fatal - usually the instruction file was removed between
@@ -320,7 +320,7 @@ cp "$TRACKER_STATE" "$TRACKER_STATE.bak.$(date +%s)"
320
320
  # 3. Either restore from a recent valid snapshot in the same dir, or
321
321
  # reinitialize from the agent-log timeline:
322
322
  bash $HOME/.claude/scripts/phase-tracker.sh init "<task-id>"
323
- for p in 0:Init 1:Analysis 2:Planning 3:Dev 4:Review 5:Test 6:Commit 7:Report; do
323
+ for p in 0:Init 1:Plan 2:Dev 3:Review 4:Commit 5:Report; do
324
324
  bash $HOME/.claude/scripts/phase-tracker.sh add "${p%%:*}" "${p#*:}"
325
325
  done
326
326
  # Replay completed phases from agent-log.md headers
@@ -338,7 +338,7 @@ reinit. A footnote in `agent-log.md` documenting the rebuild is good practice.
338
338
  ## Token Rotation Mid-Pipeline (v8.0.0)
339
339
 
340
340
  A token (Jira / Bitbucket / GitHub / Vercel) was rotated while a long-running
341
- pipeline was active. Phase 6 / Phase 7 then fail with `401 Unauthorized` even
341
+ pipeline was active. Phase 4 / Phase 5 then fail with `401 Unauthorized` even
342
342
  though earlier phases worked.
343
343
 
344
344
  Fix:
@@ -362,7 +362,7 @@ echo "$NEW_TOKEN" | secret-tool store --label="$LABEL" account "$ACCOUNT" servic
362
362
  # 2. Sanity-check the doctor reports the new token's prefix
363
363
  bash pipeline/lib/credential-store.sh doctor
364
364
 
365
- # 3. Resume the pipeline - Phase 6/7 re-resolve the token on each invocation,
365
+ # 3. Resume the pipeline - Phase 4/5 re-resolve the token on each invocation,
366
366
  # so no state edit is required.
367
367
  /multi-agent:resume <task-id>
368
368
  ```
@@ -446,7 +446,7 @@ versions).
446
446
  ## Failed Push + Force-Push Decision (v8.0.0)
447
447
 
448
448
  `git push` rejects the branch (non-fast-forward, hook rejected, branch
449
- protection mismatch, etc.). The pipeline is paused at Phase 6 Step 3.
449
+ protection mismatch, etc.). The pipeline is paused at Phase 4 Step 3.
450
450
 
451
451
  Decision tree:
452
452
 
@@ -0,0 +1,188 @@
1
+ # Running multi-agent on a machine nobody is sitting at
2
+
3
+ This describes what a Mac has to have before an unattended run works, and what
4
+ breaks if it does not. **Nothing here is set up by reading it** - `install`
5
+ still writes nothing about permissions and no scheduler is loaded unless you
6
+ ask for one. That is deliberate: every item below changes what runs on a
7
+ machine without a human, and none of it should happen because someone ran an
8
+ installer.
9
+
10
+ Scope is macOS. Linux and Windows were removed on purpose (ADR-0012) and
11
+ nothing here reintroduces them.
12
+
13
+ The mechanical half of this page is `doctor --profile=server`, which checks
14
+ four of the conditions below and says which are missing. Read that first; this
15
+ page explains why each one matters and what to do about it.
16
+
17
+ ## The failure this page exists for
18
+
19
+ An unattended run does not fail loudly. It **stops without saying so**, and
20
+ every way it does that looks identical from outside: a process that is running,
21
+ printing nothing, finishing never.
22
+
23
+ Four causes, in the order they bite.
24
+
25
+ ## 1. The permission posture
26
+
27
+ autopilot spawns its child with `--permission-prompts none`. That stops Claude
28
+ Code from ASKING, and it does not grant anything - the tools the child then
29
+ calls still have to be allowed. On the maintainer's machine they are, because
30
+ `Bash`, `Edit`, `Write` and `Agent` sit in the global allow list from years of
31
+ interactive use. On a fresh machine they do not, and the first tool call ends
32
+ the run.
33
+
34
+ There is no DEFAULT that fixes this and should not be: writing a permission
35
+ posture into someone's settings during an install is the one action an
36
+ installer must never take silently. There is now a flag that asks for it:
37
+
38
+ ```bash
39
+ npx @mmerterden/multi-agent-pipeline install --unattended --dry-run # see it first
40
+ npx @mmerterden/multi-agent-pipeline install --unattended # write it
41
+ ```
42
+
43
+ It prints the whole profile with a reason per line before writing anything, and
44
+ it is additive: an entry you added by hand survives, a narrower `Bash(...)` rule
45
+ is kept alongside, unrelated settings are untouched, and a second run changes
46
+ nothing. A `settings.json` that does not parse is refused rather than
47
+ overwritten.
48
+
49
+ By hand, the same thing is:
50
+
51
+ ```jsonc
52
+ // ~/.claude/settings.json
53
+ {
54
+ "permissions": {
55
+ "allow": ["Bash", "Edit", "Write", "Agent"]
56
+ }
57
+ }
58
+ ```
59
+
60
+ Grant the narrowest set the work needs. A server that only reviews does not
61
+ need `Write` - though note that the shipped profile is deliberately broad,
62
+ because a run builds and tests whatever the target repo uses and a narrow list
63
+ is the one the first unfamiliar repo stops at.
64
+
65
+ `doctor --profile=server` reports this as `unattended-permissions`, against the
66
+ same rule the writer applies.
67
+
68
+ ## 2. The scheduler, and why it is a LaunchAgent
69
+
70
+ Something has to start the work. The template ships as a **LaunchAgent**
71
+ (`install/templates/multi-agent-autopilot.plist.template`), loaded in the
72
+ user's session at login - not a LaunchDaemon at boot.
73
+
74
+ That is a keychain decision, not a preference. Every credential the pipeline
75
+ reads lives in the login keychain, and the login keychain is **locked until the
76
+ user logs in**. A daemon started at boot gets a locked one: each credential
77
+ read fails, and those failures surface much later as 401s that blame the token.
78
+ The lock is the cause and nothing in the error says so.
79
+
80
+ So the machine needs to reach a logged-in session by itself:
81
+
82
+ - **System Settings → Users & Groups → Automatic login**, for the account that
83
+ owns the agent.
84
+ - A **keychain with no lock timeout** for that account, or the agent survives
85
+ a login and dies at the first idle period:
86
+ ```
87
+ security set-keychain-settings -l ~/Library/Keychains/login.keychain-db
88
+ ```
89
+ (no `-t`, so there is no timeout; `-l` still locks on sleep.)
90
+
91
+ Automatic login means physical access to the machine is access to the
92
+ credentials. On a server in a locked room that is the trade being made; on a
93
+ laptop it is not.
94
+
95
+ `doctor --profile=server` reports these as `scheduler` and `keychain-unlock`.
96
+
97
+ ## 3. The unattended contract
98
+
99
+ `MULTI_AGENT_UNATTENDED=1` is how a run states that nobody is watching.
100
+ `multi-agent-refs/unattended-contract.md` lists exactly which entry points read
101
+ it and what each resolves to - four of them, named, because a claim of coverage
102
+ that is not true is worse than a short list.
103
+
104
+ The variable does **not** grant permissions and does not suppress errors. It
105
+ only stops a process from waiting for an answer that is not coming.
106
+
107
+ `doctor --profile=server` reports this as `unattended-contract`.
108
+
109
+ ## 4. Which credential each phase needs
110
+
111
+ A server fails differently from a laptop here: an interactive run asks for a
112
+ missing token, an unattended one does not.
113
+
114
+ | Phase | Needs | Absent |
115
+ |---|---|---|
116
+ | 0 Init | `github` (issue intake), `jira` (ticket intake) | the run cannot resolve its input and stops at the picker |
117
+ | 1-2 Analysis, Plan | `figma` or `figma_mcp` when the task names a design | halts by contract rather than guessing at layout |
118
+ | 3 Dev | none beyond git access | - |
119
+ | 4 Review | `github` for PR comments | findings are computed and never posted |
120
+ | 5 Test | none | - |
121
+ | 6 Commit | git push credential | the branch stays local, which reads as "no PR yet" |
122
+ | 7 Report | `jira`, `confluence` as configured | the report is composed and dropped |
123
+
124
+ `credential-store.sh doctor` lists what is mapped. `doctor --probe` additionally
125
+ checks each one is alive, which costs network calls and is why it is opt-in.
126
+
127
+ ## 4b. What the runner does when nobody is looking
128
+
129
+ Three behaviours that are invisible on a laptop and decide whether a server
130
+ survives a month.
131
+
132
+ **The log.** launchd appends the runner's stdout to `runner.log` forever. Each
133
+ tick truncates it past `MA_AP_LOG_MAX_BYTES` (5MB by default), keeping the last
134
+ 256KB in `runner.log.1`. It truncates the SAME file rather than renaming it,
135
+ because launchd holds it open with `O_APPEND` - rename it and that descriptor
136
+ keeps writing into the renamed file while the new one stays empty, which is the
137
+ rotation that looks correct and silently stops logging.
138
+
139
+ **The breaker.** An expired token, a `claude` that no longer launches, a full
140
+ disk: every queued item fails the same way, minutes apart, until the queue is
141
+ empty and the ledger is a list of identical failures with no indication which
142
+ came first. After `MA_AP_BREAKER_LIMIT` consecutive attempts that produced
143
+ nothing (3 by default, 0 disables), the runner stops taking new work, writes the
144
+ reason into `queue.json` so `autopilot-status` shows it, and says what to check.
145
+ One successful attempt clears it. A run waiting for a person (`needs-input`) and
146
+ an item arming refused (`blocked-*`) are deliberately not failures.
147
+
148
+ **Telemetry.** `ticks.jsonl` gets one JSON line per tick. The human log answers
149
+ "what happened just now"; a server is only ever asked "how has this been
150
+ behaving for a week".
151
+
152
+ ## 5. Self-hosted CI runner
153
+
154
+ ADR-0011 rejected a self-hosted runner because "a self-hosted runner on a
155
+ public repository executes code from any fork's pull request". Half of that has
156
+ changed and half has not, and the difference matters before wiring a runner to
157
+ a machine that holds credentials.
158
+
159
+ Measured, not assumed: the repo is **private**, and it has **one fork**, owned
160
+ by another account. "Zero forks" is a claim worth not making - GitHub requires
161
+ approval before a fork's pull request runs a workflow, but that protection is a
162
+ SETTING, and a self-hosted runner turns "someone approved a workflow run by
163
+ reflex" into code on this machine with this keychain.
164
+
165
+ So: a runner here is defensible on a private single-owner repo, and it is
166
+ defensible only while fork-PR approval stays on and every approval is treated
167
+ as a decision rather than a formality.
168
+
169
+ The same machine can host the runner:
170
+
171
+ ```
172
+ mkdir ~/actions-runner && cd ~/actions-runner
173
+ # download and configure per GitHub's instructions for the repo
174
+ ./svc.sh install && ./svc.sh start
175
+ ```
176
+
177
+ Two things to keep in mind. The runner service runs in the same session as the
178
+ autopilot agent, so a long CI job and a pipeline run compete for the same CPU
179
+ and the same `~/.claude` state root. And the runner inherits that session's
180
+ keychain, which means a workflow can read the credentials - acceptable on a
181
+ single-owner machine, not on a shared one.
182
+
183
+ ## What this page does NOT cover
184
+
185
+ - Turning any of it on. Every step here is the operator's to take.
186
+ - Linux, Windows, systemd, Docker. Out of scope and staying there.
187
+ - Multi-user servers. Everything above assumes one account owns the install,
188
+ the keychain and the queue.