@mmerterden/multi-agent-pipeline 18.0.0 → 19.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (234) hide show
  1. package/CHANGELOG.md +287 -0
  2. package/README.md +36 -20
  3. package/README.tr.md +14 -16
  4. package/docs/adr/0002-instruction-driven-flag.md +1 -0
  5. package/docs/adr/0005-lazy-phase-docs.md +11 -1
  6. package/docs/adr/0008-installer-modularization-and-secret-leak-defense.md +1 -0
  7. package/docs/adr/0010-own-code-graph.md +1 -0
  8. package/docs/adr/0014-six-phase-consolidation.md +134 -0
  9. package/docs/adr/README.md +2 -1
  10. package/docs/architecture.md +37 -38
  11. package/docs/best-practices.md +1 -1
  12. package/docs/ecosystem.md +46 -27
  13. package/docs/engineering.md +1 -1
  14. package/docs/facts.json +61 -0
  15. package/docs/features.md +55 -54
  16. package/docs/performance.md +5 -5
  17. package/docs/recovery-guide.md +17 -17
  18. package/docs/token-budget-history.md +3 -1
  19. package/index.js +2 -2
  20. package/install/_codex-agents.mjs +1 -1
  21. package/install/templates/claude-hooks.json +1 -1
  22. package/install/templates/codex-instructions.md +1 -1
  23. package/install/templates/copilot-instructions.md +28 -28
  24. package/manifest.json +234 -216
  25. package/package.json +2 -2
  26. package/pipeline/agents/dev-critic.md +7 -7
  27. package/pipeline/commands/figma-to-swiftui.md +1 -1
  28. package/pipeline/commands/multi-agent/SKILL.md +9 -9
  29. package/pipeline/commands/multi-agent/analysis/SKILL.md +15 -15
  30. package/pipeline/commands/multi-agent/analysis-resolve/SKILL.md +2 -2
  31. package/pipeline/commands/multi-agent/autopilot/SKILL.md +7 -7
  32. package/pipeline/commands/multi-agent/channels/SKILL.md +15 -15
  33. package/pipeline/commands/multi-agent/diff-explain/SKILL.md +6 -6
  34. package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +1 -1
  35. package/pipeline/commands/multi-agent/graph/SKILL.md +1 -1
  36. package/pipeline/commands/multi-agent/help/SKILL.md +62 -62
  37. package/pipeline/commands/multi-agent/language/SKILL.md +2 -2
  38. package/pipeline/commands/multi-agent/local/SKILL.md +11 -11
  39. package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +13 -13
  40. package/pipeline/commands/multi-agent/log/SKILL.md +2 -2
  41. package/pipeline/commands/multi-agent/manual-test/SKILL.md +9 -9
  42. package/pipeline/commands/multi-agent/model/SKILL.md +69 -0
  43. package/pipeline/commands/multi-agent/refactor/SKILL.md +3 -3
  44. package/pipeline/commands/multi-agent/resume/SKILL.md +4 -4
  45. package/pipeline/commands/multi-agent/resume-local/SKILL.md +19 -17
  46. package/pipeline/commands/multi-agent/review/SKILL.md +2 -2
  47. package/pipeline/commands/multi-agent/review-analysis/SKILL.md +1 -1
  48. package/pipeline/commands/multi-agent/route-off/SKILL.md +36 -0
  49. package/pipeline/commands/multi-agent/route-on/SKILL.md +74 -0
  50. package/pipeline/commands/multi-agent/route-status/SKILL.md +56 -0
  51. package/pipeline/commands/multi-agent/setup/SKILL.md +2 -2
  52. package/pipeline/commands/multi-agent/status/SKILL.md +5 -5
  53. package/pipeline/commands/multi-agent/steer/SKILL.md +2 -2
  54. package/pipeline/commands/multi-agent/sync/SKILL.md +12 -13
  55. package/pipeline/commands/multi-agent/test/SKILL.md +1 -1
  56. package/pipeline/lib/credential-inventory.sh +1 -1
  57. package/pipeline/lib/fetch-fortify.sh +1 -1
  58. package/pipeline/lib/model-dispatch.sh +140 -0
  59. package/pipeline/lib/model-rung.sh +142 -0
  60. package/pipeline/lib/outbound-gate.mjs +14 -0
  61. package/pipeline/lib/phase-schema.mjs +88 -0
  62. package/pipeline/lib/plan-todos.sh +5 -5
  63. package/pipeline/lib/route-state.sh +161 -0
  64. package/pipeline/lib/run-paths.sh +2 -2
  65. package/pipeline/multi-agent-refs/_account-picker.md +1 -1
  66. package/pipeline/multi-agent-refs/_dev-context.md +6 -6
  67. package/pipeline/multi-agent-refs/_input-parser.md +1 -1
  68. package/pipeline/multi-agent-refs/analysis/evidence.md +2 -11
  69. package/pipeline/multi-agent-refs/analysis/intake.md +7 -7
  70. package/pipeline/multi-agent-refs/analysis/locked.md +48 -22
  71. package/pipeline/multi-agent-refs/analysis/redesign.md +1 -1
  72. package/pipeline/multi-agent-refs/analysis/render.md +10 -10
  73. package/pipeline/multi-agent-refs/analysis/resolve.md +1 -1
  74. package/pipeline/multi-agent-refs/analysis/review.md +2 -2
  75. package/pipeline/multi-agent-refs/analysis/synthesis.md +13 -7
  76. package/pipeline/multi-agent-refs/analysis-template-corporate.md +9 -9
  77. package/pipeline/multi-agent-refs/analysis-template.md +19 -19
  78. package/pipeline/multi-agent-refs/android-guide.md +1 -1
  79. package/pipeline/multi-agent-refs/audit-guide.md +13 -13
  80. package/pipeline/multi-agent-refs/channels/issue-comment.md +2 -2
  81. package/pipeline/multi-agent-refs/channels/jira.md +3 -3
  82. package/pipeline/multi-agent-refs/channels/pr.md +4 -4
  83. package/pipeline/multi-agent-refs/channels/wiki.md +1 -1
  84. package/pipeline/multi-agent-refs/component-dispatch.md +8 -8
  85. package/pipeline/multi-agent-refs/conventions-defaults.md +2 -2
  86. package/pipeline/multi-agent-refs/cross-cli-contract.md +31 -6
  87. package/pipeline/multi-agent-refs/features/analysis-jira.md +1 -1
  88. package/pipeline/multi-agent-refs/features/autopilot-circuit-breaker.md +4 -4
  89. package/pipeline/multi-agent-refs/features/code-graph.md +5 -5
  90. package/pipeline/multi-agent-refs/features/design-conformance.md +1 -1
  91. package/pipeline/multi-agent-refs/features/dev-critic.md +3 -3
  92. package/pipeline/multi-agent-refs/features/doctor.md +3 -3
  93. package/pipeline/multi-agent-refs/features/external-context-injection.md +3 -3
  94. package/pipeline/multi-agent-refs/features/maturity-followup.md +3 -3
  95. package/pipeline/multi-agent-refs/features/model-fallback.md +41 -5
  96. package/pipeline/multi-agent-refs/features/plan-todos.md +1 -1
  97. package/pipeline/multi-agent-refs/features/repo-map.md +1 -1
  98. package/pipeline/multi-agent-refs/features/review-delta.md +3 -3
  99. package/pipeline/multi-agent-refs/features/review-multi-repo.md +2 -2
  100. package/pipeline/multi-agent-refs/features/scope-check.md +4 -4
  101. package/pipeline/multi-agent-refs/features/skill-conformance.md +2 -2
  102. package/pipeline/multi-agent-refs/features/stack-skill-routing.md +1 -1
  103. package/pipeline/multi-agent-refs/features/url-enrichment.md +1 -1
  104. package/pipeline/multi-agent-refs/features/verify-by-test.md +4 -4
  105. package/pipeline/multi-agent-refs/features/visual-evidence.md +19 -19
  106. package/pipeline/multi-agent-refs/features/worktree-finalize.md +6 -6
  107. package/pipeline/multi-agent-refs/issue-jira-triad.md +10 -10
  108. package/pipeline/multi-agent-refs/knowledge.md +11 -11
  109. package/pipeline/multi-agent-refs/multi-repo-integration-build.md +13 -13
  110. package/pipeline/multi-agent-refs/payload-contracts.md +8 -8
  111. package/pipeline/multi-agent-refs/phases/log-format.md +10 -10
  112. package/pipeline/multi-agent-refs/phases/modes.md +30 -30
  113. package/pipeline/multi-agent-refs/phases/operations.md +8 -8
  114. package/pipeline/multi-agent-refs/phases/phase-0-init.md +24 -24
  115. package/pipeline/multi-agent-refs/phases/phase-1-plan.md +599 -0
  116. package/pipeline/multi-agent-refs/phases/{phase-3-dev.md → phase-2-dev.md} +129 -49
  117. package/pipeline/multi-agent-refs/phases/{phase-4-review.md → phase-3-review.md} +225 -107
  118. package/pipeline/multi-agent-refs/phases/{phase-6-commit.md → phase-4-commit.md} +23 -23
  119. package/pipeline/multi-agent-refs/phases/{phase-7-report.md → phase-5-report.md} +29 -29
  120. package/pipeline/multi-agent-refs/phases.md +44 -48
  121. package/pipeline/multi-agent-refs/picker-contract.md +1 -1
  122. package/pipeline/multi-agent-refs/progress-contract.md +6 -6
  123. package/pipeline/multi-agent-refs/readiness-review.md +1 -1
  124. package/pipeline/multi-agent-refs/rules.md +7 -7
  125. package/pipeline/multi-agent-refs/swiftui-guide.md +2 -2
  126. package/pipeline/multi-agent-refs/tracker-contract.md +31 -32
  127. package/pipeline/multi-agent-refs/wiki-capture.md +14 -14
  128. package/pipeline/preferences-template.json +9 -1
  129. package/pipeline/rules/figma-pipeline.md +8 -8
  130. package/pipeline/rules/outside-the-pipeline.md +1 -1
  131. package/pipeline/schemas/agent-state.schema.json +50 -50
  132. package/pipeline/schemas/analysis-output.schema.json +3 -3
  133. package/pipeline/schemas/analysis-spec.schema.json +2 -2
  134. package/pipeline/schemas/autopilot-config.schema.json +1 -1
  135. package/pipeline/schemas/code-graph.schema.json +1 -1
  136. package/pipeline/schemas/criteria-manifest.schema.json +1 -1
  137. package/pipeline/schemas/dev-critic-output.schema.json +1 -1
  138. package/pipeline/schemas/diff-risk.schema.json +1 -1
  139. package/pipeline/schemas/figma-project-config.schema.json +1 -1
  140. package/pipeline/schemas/migrations/prefs-2.4.0-to-2.5.0.mjs +2 -2
  141. package/pipeline/schemas/migrations/prefs-2.6.0-to-2.7.0.mjs +31 -0
  142. package/pipeline/schemas/migrations/state-2.1.0-to-2.2.0.mjs +129 -0
  143. package/pipeline/schemas/phases.json +105 -0
  144. package/pipeline/schemas/plan-todos.schema.json +5 -5
  145. package/pipeline/schemas/planning-output.schema.json +1 -1
  146. package/pipeline/schemas/prefs.schema.json +102 -58
  147. package/pipeline/schemas/reviewer-output.schema.json +3 -3
  148. package/pipeline/schemas/route-config.schema.json +74 -0
  149. package/pipeline/schemas/scope-check.schema.json +1 -1
  150. package/pipeline/schemas/secret-patterns.json +124 -0
  151. package/pipeline/schemas/test-gap.schema.json +1 -1
  152. package/pipeline/schemas/token-budget.json +12 -18
  153. package/pipeline/schemas/triage-output.schema.json +6 -6
  154. package/pipeline/scripts/README.md +3 -3
  155. package/pipeline/scripts/_code-graph.mjs +2 -2
  156. package/pipeline/scripts/_run-paths.mjs +2 -2
  157. package/pipeline/scripts/_smoke-root.sh +1 -1
  158. package/pipeline/scripts/aggregate-metrics.mjs +1 -1
  159. package/pipeline/scripts/build-references.mjs +2 -2
  160. package/pipeline/scripts/bulk-read.sh +10 -1
  161. package/pipeline/scripts/capture-flush.sh +8 -8
  162. package/pipeline/scripts/capture-resume.sh +3 -3
  163. package/pipeline/scripts/classify-plan-safety.mjs +1 -1
  164. package/pipeline/scripts/cost-table.json +8 -1
  165. package/pipeline/scripts/diff-explain.mjs +1 -1
  166. package/pipeline/scripts/doctor.mjs +3 -3
  167. package/pipeline/scripts/gc-abandoned.sh +3 -3
  168. package/pipeline/scripts/gc-tmp.sh +1 -1
  169. package/pipeline/scripts/gc-worktrees.sh +1 -1
  170. package/pipeline/scripts/gen-facts.mjs +280 -0
  171. package/pipeline/scripts/gen-mode-dispatch.mjs +32 -37
  172. package/pipeline/scripts/gen-ref-toc.mjs +1 -1
  173. package/pipeline/scripts/graph-report.mjs +1 -1
  174. package/pipeline/scripts/jira-attach.sh +1 -1
  175. package/pipeline/scripts/learn-from-transcripts.mjs +1 -1
  176. package/pipeline/scripts/learning-curve.mjs +2 -2
  177. package/pipeline/scripts/log-metric.sh +17 -4
  178. package/pipeline/scripts/memory-save.sh +1 -1
  179. package/pipeline/scripts/migrate-prefs.mjs +22 -5
  180. package/pipeline/scripts/phase-banner.sh +20 -20
  181. package/pipeline/scripts/phase-tracker.sh +12 -12
  182. package/pipeline/scripts/plan-coverage-gate.mjs +2 -2
  183. package/pipeline/scripts/pre-commit-check.sh +30 -1
  184. package/pipeline/scripts/render-agent-log-cost.sh +1 -1
  185. package/pipeline/scripts/render-work-summary.sh +3 -3
  186. package/pipeline/scripts/review-file-filter.mjs +1 -1
  187. package/pipeline/scripts/run-aggregator.mjs +13 -6
  188. package/pipeline/scripts/run-metrics.mjs +1 -1
  189. package/pipeline/scripts/runs-index.mjs +11 -1
  190. package/pipeline/scripts/scan-skills.sh +26 -0
  191. package/pipeline/scripts/smoke-cross-cli-behavior.sh +6 -6
  192. package/pipeline/scripts/smoke-schema-validation.sh +26 -7
  193. package/pipeline/scripts/token-budget-report.mjs +13 -2
  194. package/pipeline/scripts/triage-memory.mjs +2 -2
  195. package/pipeline/scripts/validate-analysis-doc.mjs +274 -43
  196. package/pipeline/scripts/validate-planning.mjs +1 -1
  197. package/pipeline/scripts/validate-reviewer.mjs +1 -1
  198. package/pipeline/scripts/validate-state.mjs +45 -5
  199. package/pipeline/scripts/validate-triage.mjs +3 -3
  200. package/pipeline/scripts/verify-citations.mjs +1 -1
  201. package/pipeline/scripts/worktree-finalize.sh +5 -5
  202. package/pipeline/scripts/write-state.mjs +32 -0
  203. package/pipeline/skills/.skill-manifest.json +38 -22
  204. package/pipeline/skills/.skills-index.json +49 -5
  205. package/pipeline/skills/shared/README.md +10 -6
  206. package/pipeline/skills/shared/core/apple-archive-compliance/SKILL.md +8 -8
  207. package/pipeline/skills/shared/core/google-play-compliance/SKILL.md +8 -8
  208. package/pipeline/skills/shared/core/multi-agent/SKILL.md +81 -82
  209. package/pipeline/skills/shared/core/multi-agent-autopilot/SKILL.md +3 -3
  210. package/pipeline/skills/shared/core/multi-agent-channels/SKILL.md +14 -14
  211. package/pipeline/skills/shared/core/multi-agent-diff-explain/SKILL.md +5 -5
  212. package/pipeline/skills/shared/core/multi-agent-graph/SKILL.md +1 -1
  213. package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +25 -23
  214. package/pipeline/skills/shared/core/multi-agent-language/SKILL.md +2 -2
  215. package/pipeline/skills/shared/core/multi-agent-local/SKILL.md +2 -2
  216. package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +8 -8
  217. package/pipeline/skills/shared/core/multi-agent-manual-test/SKILL.md +6 -6
  218. package/pipeline/skills/shared/core/multi-agent-model/SKILL.md +71 -0
  219. package/pipeline/skills/shared/core/multi-agent-refactor/SKILL.md +3 -3
  220. package/pipeline/skills/shared/core/multi-agent-resume/SKILL.md +1 -1
  221. package/pipeline/skills/shared/core/multi-agent-resume-local/SKILL.md +7 -7
  222. package/pipeline/skills/shared/core/multi-agent-route-off/SKILL.md +39 -0
  223. package/pipeline/skills/shared/core/multi-agent-route-on/SKILL.md +76 -0
  224. package/pipeline/skills/shared/core/multi-agent-route-status/SKILL.md +59 -0
  225. package/pipeline/skills/shared/core/multi-agent-setup/SKILL.md +1 -1
  226. package/pipeline/skills/shared/core/multi-agent-status/SKILL.md +5 -5
  227. package/pipeline/skills/shared/core/multi-agent-steer/SKILL.md +2 -2
  228. package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +6 -5
  229. package/pipeline/skills/shared/external/NOTICE-swift-ios-skills.md +1 -1
  230. package/pipeline/skills/shared/external/signal-community/SKILL.md +8 -1
  231. package/pipeline/skills/skills-index.md +8 -4
  232. package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +0 -263
  233. package/pipeline/multi-agent-refs/phases/phase-2-planning.md +0 -344
  234. package/pipeline/multi-agent-refs/phases/phase-5-test.md +0 -182
@@ -53,27 +53,26 @@ How It Works (Phase 0 - Interactive Flow):
53
53
  Pipeline (after Phase 0) - shown as visual cards in terminal:
54
54
 
55
55
  Phase 0: Init -> The 8 steps above
56
- Phase 1: Analysis -> Stack detection + codebase scan (Sonnet)
57
- Phase 2: Planning -> Task breakdown + architecture review + Plan Approval Gate
58
- (clarification max 2 rounds + approval loop - Full + interactive
59
- only; a Short run has no plan, autopilot may not ask)
60
- Phase 3: Dev -> TDD: test -> code -> build (Sonnet) + build queue
61
- Phase 4: Review -> Deterministic gates + parallel AI review + Fable triage
62
- (Claude Code: Fable + Opus + Sonnet · Copilot CLI: GPT-5.4 + Opus + Sonnet)
63
- Phase 5: Test -> Optional: switch to branch, test in Xcode
64
- (only /multi-agent has it; every autopilot and local entry drops it)
65
- Phase 6: Commit -> Commit -> push -> PR + issue body update (never auto-closes)
66
- Phase 7: Report -> Channels dispatcher (PR · Jira · Confluence · Wiki, multi-select)
56
+ Phase 1: Plan -> Stack detection + codebase scan, task breakdown, architecture
57
+ review, Plan Approval Gate (Opus; max 2 clarification rounds,
58
+ Full + interactive only - a Short run has no plan)
59
+ Phase 2: Dev -> TDD: test -> code -> build (Sonnet) + build queue, then the
60
+ Verify exit gate: build, lint, tests, secrets. The build runs
61
+ ONCE per run and its log carries into Review
62
+ Phase 3: Review -> Parallel AI review + Fable triage, then the optional user test
63
+ (Claude Code: Fable + Opus + Sonnet · Copilot: GPT-5.4 + Opus + Sonnet;
64
+ the user test is the interactive half: branch, test in Xcode)
65
+ Phase 4: Commit -> Commit -> push -> PR + issue body update (never auto-closes)
66
+ Phase 5: Report -> Channels dispatcher (PR · Jira · Confluence · Wiki, multi-select)
67
67
  + internal capture (agent-log · knowledge · memory)
68
68
 
69
- Autopilot always pauses at the Phase 7 channels menu (30-min timeout → session ends cleanly).
69
+ Autopilot always pauses at the Phase 5 channels menu (30-min timeout → session ends cleanly).
70
70
 
71
71
  Pipeline depth is asked at Phase 0 Step 7.5, not passed as a flag:
72
- Full Analysis -> Plan -> Dev(Sonnet) -> Review -> Test -> Commit -> Report
73
- Short Dev(Opus, self-contained) -> Review -> Test -> Commit -> Report
74
- Short skips Phases 1-2 only. Review is NEVER skipped.
75
- Recommended from taskType: bugfix/chore -> Short, feature/refactor/component -> Full.
76
- Autopilot never asks and always runs Full.
72
+ Full Plan -> Dev(Sonnet) -> Review -> Commit -> Report
73
+ Short Dev(Opus, self-contained) -> Review -> Commit -> Report
74
+ Short skips Phase 1 only; Review is NEVER skipped. Recommended from taskType:
75
+ bugfix/chore -> Short, feature/refactor/component -> Full. Autopilot always runs Full.
77
76
 
78
77
  Every step is logged. Error in any phase -> pause -> resume to continue.
79
78
 
@@ -90,10 +89,10 @@ Four pipeline entries:
90
89
 
91
90
  --local on the base command is the same as the :local entry.
92
91
  autopilot skips every confirmation INCLUDING the plan gate and the depth
93
- question, and auto commit/PR - EXCEPT the Phase 7 channels menu, which
92
+ question, and auto commit/PR - EXCEPT the Phase 5 channels menu, which
94
93
  always pauses.
95
94
 
96
- /multi-agent:resume-local [jira-id] [autopilot] Continue already-done LOCAL work: Review → Build+Test → PR → Jira analysis + test scenarios (no dev)
95
+ /multi-agent:resume-local [jira-id] [autopilot] Continue already-done LOCAL work: Review (build gate) → Commit/PR → Report (no dev)
97
96
 
98
97
  ------------------------------
99
98
 
@@ -126,7 +125,7 @@ Post-Hoc & Side-Channel:
126
125
  /multi-agent:complaint-analysis ["run-name"] Customer-complaint triage: Graylog evidence per trx/conv id + read-only repo correlation → client/bff root cause + fix plan + dev prompt, or core routing recommendation
127
126
  /multi-agent:build-optimize iOS-only Xcode build perf wrapper → benchmark + analyze + recommend-first .build-benchmark/optimization-plan.md
128
127
  /multi-agent:create-jira ["desc"] [figma-url] [swagger-url] Create a Jira Task/Bug/Story matching team conventions (asks type + mining + active sprint + auto-sizing sections + preview & approval)
129
- /multi-agent:diff-explain Map a Phase 4 triage finding back to specific diff lines
128
+ /multi-agent:diff-explain Map a Phase 3 triage finding back to specific diff lines
130
129
  /multi-agent:graph Build and query the repo code graph (symbols, imports, references), LLM-free
131
130
  /multi-agent:search Cross-task log search with smart ranking; --semantic queries triage corpus
132
131
  /multi-agent:scan Skill security scan against tiered pattern catalog
@@ -142,6 +141,10 @@ Setup & Maintenance:
142
141
 
143
142
  /multi-agent:setup First-run wizard - Keychain tokens + Git identity + language
144
143
  /multi-agent:language [en|tr] Show or set outputLanguage (promptLanguage stays English)
144
+ /multi-agent:model [on|off] Turn the top model rung (fable) on or off; keeps costBudget.pricingModel in step
145
+ /multi-agent:route-on Arm policy-driven model routing: strategy, scope, rules. Ships off
146
+ /multi-agent:route-off Disarm routing and KEEP the rules, so route-on does not re-ask
147
+ /multi-agent:route-status What is armed, which rule applies where, and the limit: subagents stay on Anthropic
145
148
  /multi-agent:stack [ids...] Enable stack plugin(s); multi-select: ids together (ios backend) or no arg -> native picker (common always on)
146
149
  /multi-agent:sync Sync ecosystem (Claude Code + Copilot CLI + pipeline + website + toolkit MCP)
147
150
  /multi-agent:update Pull latest pipeline + reinstall + run migrations
@@ -208,7 +211,7 @@ UI Testing (standalone - not part of pipeline phases):
208
211
  Uses xcrun simctl / adb (native, no external app needed).
209
212
  Booted simulator/emulator required. Auto-detects bundle ID from project.
210
213
 
211
- Manual Test (Phase 5 standalone - Xcode hint flow):
214
+ Manual Test (Phase 3 user-test step, standalone - Xcode hint flow):
212
215
 
213
216
  /multi-agent:manual-test Checkout task branch, print Xcode/SourceTree hints,
214
217
  /multi-agent:manual-test #N wait for your "ok" / "fix: ..." verdict.
@@ -250,21 +253,18 @@ Key Features:
250
253
  Multi-Repo Per-repo worktrees + identity, integration build before commit
251
254
  Identity Routing Git identity picked from the repo origin URL (corporate vs personal)
252
255
  Issue Safety Never auto-closes issues (4 approvals, GitHub + Jira)
253
- Store Compliance /multi-agent:test "store-ready" - 18-rule iOS audit (ITMS, Privacy Manifest,
254
- signing, debug leaks, IPv6, SDK list) + 21-rule Android audit
256
+ Store Compliance /multi-agent:test "store-ready" - 18-rule iOS audit (ITMS, Privacy Manifest, signing, debug leaks, IPv6, SDK) + 21-rule Android audit
255
257
  Bilingual EN + TR - outputLanguage toggles assistant explanations; promptLanguage is locked en
256
258
 
257
259
  Quality & Telemetry (advisory, on by default - flip prefs.global.* to disable):
258
260
 
259
- Diff Risk Score Phase 4 Step 1.75 ranks files before reviewer dispatch (security paths,
260
- migrations, no-test-change, complexity) - heuristic, sub-second
261
- Test Gap Report Phase 5 Step 0 surfaces public symbols added in this branch with no paired test
261
+ Diff Risk Score Phase 3 Step 1.75 ranks files before reviewer dispatch - heuristic, sub-second: security paths, migrations, no-test-change, complexity
262
+ Test Gap Report Phase 3 surfaces public symbols added in this branch with no paired test
262
263
  Visual Evidence before/after stills + flow video on UI changes; Step 7.7 asks the depth
263
- Cost Breakdown Phase 7 appends per-phase tokens + estimated USD to agent-log.md
264
- Triage Memory Phase 7 ingests accepted/deferred/rejected findings into a repo corpus
265
- Prior-Art Lookup Phase 1 + Phase 4 query the corpus for similar findings, inject as context
266
- Per-Persona Dispatch reads `preferredModel` from the persona file; override per call via
267
- PHASE_MODEL_OVERRIDE; ladder fable -> opus -> sonnet -> haiku
264
+ Cost Breakdown Phase 5 appends per-phase tokens + estimated USD to agent-log.md
265
+ Triage Memory Phase 5 ingests accepted/deferred/rejected findings into a repo corpus
266
+ Prior-Art Lookup Phase 1 + Phase 3 query the corpus for similar findings, inject as context
267
+ Per-Persona Dispatch reads `preferredModel` from the persona file; PHASE_MODEL_OVERRIDE overrides per call; ladder fable -> opus -> sonnet -> haiku
268
268
 
269
269
  ------------------------------
270
270
 
@@ -333,27 +333,26 @@ Nasıl Çalışır (Phase 0 - İnteraktif Akış):
333
333
  Pipeline (Phase 0'dan sonra) - terminalde görsel kart olarak görünür:
334
334
 
335
335
  Phase 0: Init -> Yukarıdaki 8 adım
336
- Phase 1: Analysis -> Stack tespiti + codebase taraması (Sonnet)
337
- Phase 2: Planning -> Task kırılımı + mimari inceleme + Plan Onay Kapısı
338
- (clarification max 2 tur + onay döngüsü - sadece Tam +
339
- etkileşimli; Kısa'da plan yok, autopilot soru soramaz)
340
- Phase 3: Dev -> TDD: test -> kod -> build (Sonnet) + build queue
341
- Phase 4: Review -> Deterministik kapılar + paralel AI review + Fable triage
342
- (Claude Code: Fable + Opus + Sonnet · Copilot CLI: GPT-5.4 + Opus + Sonnet)
343
- Phase 5: Test -> Opsiyonel: branch'e geç, Xcode'da test
344
- (yalnız /multi-agent'ta var; her autopilot ve local girişi düşürür)
345
- Phase 6: Commit -> Commit -> push -> PR + issue body güncelleme (hiç auto-close yok)
346
- Phase 7: Report -> Channels dispatcher (PR · Jira · Confluence · Wiki, multi-select)
336
+ Phase 1: Plan -> Stack tespiti + codebase taraması, task kırılımı, mimari
337
+ inceleme, Plan Onay Kapısı (Opus; max 2 clarification turu,
338
+ sadece Tam + etkileşimli - Kısa'da plan yok)
339
+ Phase 2: Dev -> TDD: test -> kod -> build (Sonnet) + build queue, sonra Verify
340
+ çıkış kapısı: build, lint, test, sır taraması. Build koşu başına
341
+ BİR kez çalışır, log'u Review'a devreder
342
+ Phase 3: Review -> Paralel AI review + Fable triage, sonra opsiyonel kullanıcı testi
343
+ (Claude Code: Fable + Opus + Sonnet · Copilot: GPT-5.4 + Opus + Sonnet;
344
+ kullanıcı testi etkileşimli yarısı: branch'e geç, Xcode'da test)
345
+ Phase 4: Commit -> Commit -> push -> PR + issue body güncelleme (hiç auto-close yok)
346
+ Phase 5: Report -> Channels dispatcher (PR · Jira · Confluence · Wiki, multi-select)
347
347
  + internal capture (agent-log · knowledge · memory)
348
348
 
349
- Autopilot Phase 7'deki channels menüsünde HER ZAMAN durur (30 dk timeout → session temiz biter).
349
+ Autopilot Phase 5'teki channels menüsünde HER ZAMAN durur (30 dk timeout → session temiz biter).
350
350
 
351
351
  Pipeline derinliği Faz 0 Adım 7.5'te sorulur, bayrakla geçilmez:
352
- Tam Analiz -> Plan -> Dev(Sonnet) -> Review -> Test -> Commit -> Report
353
- Kısa Dev(Opus, kendi kendine yeten) -> Review -> Test -> Commit -> Report
354
- Kısa yalnız Faz 1-2'yi atlar. Review ASLA atlanmaz.
355
- taskType'a göre önerilir: bugfix/chore -> Kısa, feature/refactor/component -> Tam.
356
- Autopilot hiç sormaz, her zaman Tam koşar.
352
+ Tam Plan -> Dev(Sonnet) -> Review -> Commit -> Report
353
+ Kısa Dev(Opus, kendi kendine yeten) -> Review -> Commit -> Report
354
+ Kısa yalnız Faz 1'i atlar; Review ASLA atlanmaz. taskType'a göre önerilir:
355
+ bugfix/chore -> Kısa, feature/refactor/component -> Tam. Autopilot her zaman Tam koşar.
357
356
 
358
357
  Her adım loglanır. Herhangi bir fazdaki hata -> pause -> resume ile devam et.
359
358
 
@@ -370,9 +369,9 @@ Dört pipeline girişi:
370
369
 
371
370
  Ana komuta --local eklemek :local girişiyle aynıdır.
372
371
  autopilot plan kapısı ve derinlik sorusu dahil her onayı atlar, otomatik
373
- commit/PR açar - İSTİSNA: Faz 7 channels menüsü, o her zaman durur.
372
+ commit/PR açar - İSTİSNA: Faz 5 channels menüsü, o her zaman durur.
374
373
 
375
- /multi-agent:resume-local [jira-id] [autopilot] Lokalde biten işi sürdür: Review → Build+Test → PR → Jira teknik analiz + test senaryoları (dev yok)
374
+ /multi-agent:resume-local [jira-id] [autopilot] Lokalde biten işi sürdür: Review (build kapısı) → Commit/PR → Report (dev yok)
376
375
 
377
376
  ------------------------------
378
377
 
@@ -404,7 +403,7 @@ Post-Hoc & Side-Channel:
404
403
  /multi-agent:complaint-analysis ["run-adı"] Müşteri şikayeti triyajı: trx/conv id ile Graylog kanıtı + salt-okunur repo eşleştirme → client/bff kök neden + fix planı + dev prompt'u, veya core'a yönlendirme önerisi
405
404
  /multi-agent:build-optimize iOS-only Xcode build performance wrapper → benchmark + analiz + recommend-first .build-benchmark/optimization-plan.md
406
405
  /multi-agent:create-jira ["açıklama"] [figma-url] [swagger-url] Takım standartlarına uygun Jira Task/Bug/Story oluştur (tip sorar + convention mining + aktif sprint + auto-sizing bölümler + önizleme & onay)
407
- /multi-agent:diff-explain Phase 4 triage bulgusunu diff satırlarına eşle
406
+ /multi-agent:diff-explain Phase 3 triage bulgusunu diff satırlarına eşle
408
407
  /multi-agent:graph Repo kod grafiğini kur ve sorgula (semboller, import'lar, referanslar), LLM'siz
409
408
  /multi-agent:search Task log'larında akıllı arama; --semantic triage corpus'unu sorgular
410
409
  /multi-agent:scan Skill güvenlik taraması (tiered pattern catalog)
@@ -419,6 +418,10 @@ Setup & Maintenance:
419
418
 
420
419
  /multi-agent:setup İlk kurulum sihirbazı - Keychain token + Git kimliği + dil
421
420
  /multi-agent:language [en|tr] outputLanguage'ı göster veya ayarla (promptLanguage İngilizce kalır)
421
+ /multi-agent:model [on|off] Üst model basamağını (fable) aç/kapat; costBudget.pricingModel'i birlikte taşır
422
+ /multi-agent:route-on Politika güdümlü model yönlendirmesini açar: strateji, kapsam, kurallar. Varsayılan kapalı
423
+ /multi-agent:route-off Yönlendirmeyi kapatır, kuralları SAKLAR; route-on aynı şeyi tekrar sormaz
424
+ /multi-agent:route-status Ne açık, hangi kural nerede geçerli, ve sınır: subagent'lar Anthropic'te kalır
422
425
  /multi-agent:stack [ids...] Stack plugin'lerini etkinleştir; çoklu seçim: id'ler yan yana (ios backend) ya da argümansız -> native picker (common hep açık)
423
426
  /multi-agent:sync Ekosistemi senkronize et (Claude Code + Copilot CLI + pipeline + website + multi-agent-toolkit MCP)
424
427
  /multi-agent:update En son pipeline'ı çek + reinstall + migration çalıştır
@@ -485,7 +488,7 @@ UI Testing (standalone - pipeline fazlarından bağımsız):
485
488
  xcrun simctl / adb kullanır (harici app gerekmez).
486
489
  Booted simulator/emulator şart. Bundle ID proje'den otomatik algılanır.
487
490
 
488
- Manuel Test (Phase 5 standalone - Xcode hint akışı):
491
+ Manuel Test (Phase 3 kullanıcı-testi adımı, standalone - Xcode hint akışı):
489
492
 
490
493
  /multi-agent:manual-test Task branch'ine checkout, Xcode/SourceTree hint basar,
491
494
  /multi-agent:manual-test #N "ok" / "fix: ..." yanıtını bekler.
@@ -527,20 +530,17 @@ Temel Özellikler:
527
530
  Multi-Repo Repo başına worktree, repo başına identity, commit öncesi entegrasyon build'i
528
531
  Identity Routing Repo origin URL'sinden git kimliği seçimi (kurumsal vs kişisel)
529
532
  Issue Safety Issue'lar asla auto-close edilmez (GitHub + Jira için 4 onay gerekir)
530
- Store Compliance /multi-agent:test "store-ready" - iOS için 18 kurallık audit (ITMS / Privacy Manifest /
531
- code signing / debug-tool leak / IPv6 / SDK list / vb.) + Android için 21 kurallık audit
533
+ Store Compliance /multi-agent:test "store-ready" - iOS 18 kurallık audit (ITMS / Privacy Manifest / signing / debug leak / IPv6 / SDK) + Android 21 kurallık audit
532
534
  Bilingual EN + TR - outputLanguage assistant açıklamasını değiştirir; promptLanguage en kilitli
533
535
 
534
536
  Quality & Telemetry (advisory, default açık - prefs.global.* ile kapatılabilir):
535
537
 
536
- Diff Risk Score Phase 4 Step 1.75 reviewer dispatch'tan önce dosyaları sıralar (güvenlik path'leri,
537
- schema migration, test eklenmemiş değişiklik, complexity delta) - heuristik, saniyenin altı
538
- Test Gap Report Phase 5 Step 0 - bu branch'te eklenmiş public sembol + eşleşen test yoksa raporlar
539
- Cost Breakdown Phase 7 agent-log.md'ye faz başına token (in/out) + tahmini USD ekler
540
- Triage Memory Phase 7 accepted/deferred/rejected bulguları repo başına corpus'a yazar
541
- Prior-Art Lookup Phase 1 + Phase 4 corpus'tan benzer geçmiş bulgu sorgular, ek context olarak enjekte
542
- Per-Persona Reviewer/agent dispatch persona dosyasından `preferredModel` okur;
543
- per-call override PHASE_MODEL_OVERRIDE ile; merdiven fable -> opus -> sonnet -> haiku
538
+ Diff Risk Score Phase 3 Step 1.75 reviewer dispatch'tan önce dosyaları sıralar - heuristik: güvenlik path'i, schema migration, testsiz değişiklik, complexity delta
539
+ Test Gap Report Phase 3 - bu branch'te eklenmiş public sembol + eşleşen test yoksa raporlar
540
+ Cost Breakdown Phase 5 agent-log.md'ye faz başına token (in/out) + tahmini USD ekler
541
+ Triage Memory Phase 5 accepted/deferred/rejected bulguları repo başına corpus'a yazar
542
+ Prior-Art Lookup Phase 1 + Phase 3 corpus'tan benzer geçmiş bulgu sorgular, ek context olarak enjekte
543
+ Per-Persona Dispatch persona dosyasından `preferredModel` okur; per-call override PHASE_MODEL_OVERRIDE; merdiven fable -> opus -> sonnet -> haiku
544
544
 
545
545
  ------------------------------
546
546
 
@@ -12,7 +12,7 @@ Two-axis language preference for the pipeline:
12
12
 
13
13
  | Field | Controls | Stored at | Mutability |
14
14
  |---|---|---|---|
15
- | `promptLanguage` | Interactive prompts during a pipeline run (account picker, project picker, dev-context picker, base-branch picker, branch-name picker, maturity ack, channels picker, Phase 5 test prompt, Phase 6 local-checkout prompt) | `prefs.global.promptLanguage` | **Fixed to `en`** - never toggled by this skill |
15
+ | `promptLanguage` | Interactive prompts during a pipeline run (account picker, project picker, dev-context picker, base-branch picker, branch-name picker, maturity ack, channels picker, Phase 3 test prompt, Phase 4 local-checkout prompt) | `prefs.global.promptLanguage` | **Fixed to `en`** - never toggled by this skill |
16
16
  | `outputLanguage` | Assistant's explanations, status updates, error messages, and pipeline-generated reports rendered to the user (NOT external payloads) | `prefs.global.outputLanguage` | Toggled by this skill |
17
17
 
18
18
  **Why promptLanguage is fixed:** it governs the picker's structural chrome only - the `AskUserQuestion` `header` chip, host error UI, and internal contract identifiers. The chip stays English because it is capped at 12 characters and most Turkish equivalents overflow it. Everything a user actually reads follows `outputLanguage`: the picker `question`, each option's `label`, and each option's `description`, per the canonical per-field matrix in `multi-agent-refs/rules.md`. A picker whose question or buttons are English on a Turkish run is a bug, not the contract. The host's own **Other** row is injected by the CLI and stays English on every run; nothing in the pipeline can localize it.
@@ -109,5 +109,5 @@ Render in the **new** outputLanguage:
109
109
 
110
110
  - `setup.md` Step 0 - first-run language picker (asks only outputLanguage; promptLanguage seeded as `en`)
111
111
  - `prefs.schema.json` - `global.promptLanguage` (fixed `"en"`) and `global.outputLanguage` definitions
112
- - `phase-5-test.md`, `phase-6-commit.md` - interactive prompts follow the per-field matrix: `question` + `label` + `description` in `outputLanguage`, `header` English
112
+ - `phase-3-review.md`, `phase-4-commit.md` - interactive prompts follow the per-field matrix: `question` + `label` + `description` in `outputLanguage`, `header` English
113
113
  - `help.md` - reads outputLanguage to render its own help text
@@ -28,12 +28,12 @@ Same phase count and order as the normal pipeline:
28
28
 
29
29
  ```
30
30
  Phase 0: Init → project detection, branch check, state (NO worktree)
31
- Phase 1: Analysis → codebase scan (parallel explore agents, Opus)
32
- Phase 2: Planning → task breakdown, Plan Approval Gate (approval loop)
33
- Phase 3: Dev → TDD (Sonnet), build queue
34
- Phase 4: Review → deterministic gates + parallel review + Fable triage
35
- Phase 6: Commit → pre-commit checkout prompt, commit + push + PR
36
- Phase 7: Report → Jira / Wiki / Confluence + log + knowledge/memory
31
+ Phase 1: Plan (analysis) → codebase scan (parallel explore agents, Opus)
32
+ Phase 1: Plan → task breakdown, Plan Approval Gate (approval loop)
33
+ Phase 2: Dev → TDD (Sonnet), build queue
34
+ Phase 3: Review → deterministic gates + parallel review + Fable triage
35
+ Phase 4: Commit → pre-commit checkout prompt, commit + push + PR
36
+ Phase 5: Report → Jira / Wiki / Confluence + log + knowledge/memory
37
37
  ```
38
38
 
39
39
  ## Delegation
@@ -73,11 +73,11 @@ bash $HOME/.claude/scripts/phase-tracker.sh update 0 in_progress
73
73
 
74
74
  # Phase 0 Step 7.5, immediately after the depth answer - the first moment this
75
75
  # mode knows its phase set. Full:
76
- for p in "1:Analysis" "2:Planning" "3:Dev" "4:Review" "6:Commit" "7:Report"; do
76
+ for p in "1:Plan" "2:Dev" "3:Review" "4:Commit" "5:Report"; do
77
77
  bash $HOME/.claude/scripts/phase-tracker.sh add "${p%%:*}" "${p#*:}"
78
78
  done
79
79
  # Short (Analysis and Planning are not run, so they get no tile at all):
80
- for p in "3:Dev" "4:Review" "6:Commit" "7:Report"; do
80
+ for p in "2:Dev" "3:Review" "4:Commit" "5:Report"; do
81
81
  bash $HOME/.claude/scripts/phase-tracker.sh add "${p%%:*}" "${p#*:}"
82
82
  done
83
83
  # Then the widget, narrowed to the phases that do not have a tile yet:
@@ -98,7 +98,7 @@ In Claude Code the agent MUST also drive the native TaskList widget so the user
98
98
 
99
99
  ```text
100
100
  # Register one tile per phase, capture the taskId, persist it:
101
- for each phase in 0:Init at Step -1, then 1:Analysis, 2:Planning, 3:Dev, 4:Review, 6:Commit, 7:Report (Full) or 3:Dev, 4:Review, 6:Commit, 7:Report (Short) at Step 7.5:
101
+ for each phase in 0:Init at Step -1, then 1:Plan, 2:Dev, 3:Review, 4:Commit, 5:Report (Full) or 2:Dev, 3:Review, 4:Commit, 5:Report (Short) at Step 7.5:
102
102
  TaskCreate({ subject: "Phase <N>: <Name>", activeForm: "<doing-form>" })
103
103
  -> returns taskId
104
104
  bash $HOME/.claude/scripts/phase-tracker.sh meta <N> tasklist_id "<taskId>"
@@ -115,11 +115,11 @@ TaskUpdate({ taskId: <saved>, status: "completed" })
115
115
  bash $HOME/.claude/scripts/phase-tracker.sh update <N> completed
116
116
  ```
117
117
 
118
- `--local` mode does NOT TaskCreate phases 5 - those are not part of the `--local` phase set (`0:Init 1:Analysis 2:Planning 3:Dev 4:Review 6:Commit 7:Report`). Only register tiles for the active set.
118
+ `--local` mode TaskCreates all 6 phases (no phase is skipped).
119
119
 
120
120
  #### TaskCreate ordering (strict)
121
121
 
122
- **All TaskCreate calls in a batch fire in strict phase-number order BEFORE any TaskUpdate is applied.** For `--local` that means: Phase 0 at Step -1, then the rest in ascending order at Step 7.5 (Phase 0 → Phase 1 → Phase 2 → Phase 3 → Phase 4 → Phase 6 → Phase 7 minus whatever the depth answer drops). The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks. Full ordering contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
122
+ **All TaskCreate calls in a batch fire in strict phase-number order BEFORE any TaskUpdate is applied.** For `--local` that means: Phase 0 at Step -1, then the rest in ascending order at Step 7.5 (Phase 0 → Phase 1 → Phase 2 → Phase 3 → Phase 4 → Phase 5 minus whatever the depth answer drops). The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks. Full ordering contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
123
123
 
124
124
  ### Visual channel - Copilot CLI / plain shell
125
125
 
@@ -1,6 +1,6 @@
1
1
  ---
2
- description: "Full pipeline + local + autopilot - no worktree, no confirmations; runs 7 phases (User Test skipped) end-to-end on the current branch. Use when the full pipeline should run on the current branch with no worktree and no prompts."
3
- description-tr: "Tam pipeline + lokal + autopilot - worktree yok, onay yok; 7 fazı (Kullanıcı Testi atlanır) mevcut branch üzerinde uçtan uca koşar."
2
+ description: "Full pipeline + local + autopilot - no worktree, no confirmations; runs all 6 phases end-to-end on the current branch. Use when the full pipeline should run on the current branch with no worktree and no prompts."
3
+ description-tr: "Tam pipeline + lokal + autopilot - worktree yok, onay yok; 6 fazı mevcut branch üzerinde uçtan uca koşar (kullanıcı testi Review'ın içinde, autopilot'ta atlanır)."
4
4
  argument-hint: '"task" - issue URL, Jira ID, free-text, or #id (for resume)'
5
5
  allowed-tools: Agent, Bash, Read, Write, Edit, Glob, Grep, TaskCreate, TaskUpdate, TaskList, TaskGet, AskUserQuestion, WebFetch, WebSearch, Skill
6
6
  ---
@@ -11,16 +11,16 @@ allowed-tools: Agent, Bash, Read, Write, Edit, Glob, Grep, TaskCreate, TaskUpdat
11
11
 
12
12
  > **Language (read FIRST)**: Before any status output, read `prefs.global.outputLanguage` and render every conversational line in it. `AskUserQuestion` renders its `question`, option `label`s and option `description`s in `outputLanguage`; only `header` stays English (<=12-char chip); external payload bodies follow `outputLanguage` too (identifiers, commit messages, branch names stay English). Full contract: `$HOME/.claude/multi-agent-refs/rules.md` "Language Application".
13
13
 
14
- Run the full pipeline **without a worktree** and **with every confirmation skipped**. Phase 5 (User Test) does not run (no worktree to check out from, and no interaction), so 7 phases execute. The `autopilot + local` combination - vs `/multi-agent:autopilot` the only difference is no worktree (work stays on the current branch); vs `/multi-agent:local` the only difference is zero interaction.
14
+ Run the full pipeline **without a worktree** and **with every confirmation skipped**. All six phases execute. The user test inside Phase 3 Review is skipped - there is no worktree to check out from and no interaction to have - but it stopped being a phase of its own in v19.0.0, so nothing is dropped from the set. The `autopilot + local` combination - vs `/multi-agent:autopilot` the only difference is no worktree (work stays on the current branch); vs `/multi-agent:local` the only difference is zero interaction.
15
15
 
16
16
  ## Matrix (which one should I use?)
17
17
 
18
18
  | Command | Pipeline | Worktree | Confirmations |
19
19
  |---|---|---|---|
20
- | `/multi-agent "task"` | Full 8 phases | ✅ | ✅ (interactive) |
21
- | `/multi-agent:autopilot "task"` | 7 phases (no User Test) | ✅ | ❌ (autopilot) |
22
- | `/multi-agent:local "task"` | 7 phases (no User Test) | ❌ | ✅ (interactive) |
23
- | **`/multi-agent:local-autopilot "task"`** | **7 phases (no User Test)** | **❌** | **❌ (autopilot)** |
20
+ | `/multi-agent "task"` | 6 phases, user test offered | ✅ | ✅ (interactive) |
21
+ | `/multi-agent:autopilot "task"` | 6 phases, user test skipped | ✅ | ❌ (autopilot) |
22
+ | `/multi-agent:local "task"` | 6 phases, user test offered | ❌ | ✅ (interactive) |
23
+ | **`/multi-agent:local-autopilot "task"`** | **6 phases, user test skipped** | **❌** | **❌ (autopilot)** |
24
24
 
25
25
  Depth is a separate axis, asked at Phase 0 Step 7.5 rather than encoded in the command name: Full runs every phase above, Short runs Dev → Review → Test → Commit → Report. The two autopilot rows never ask and always run Full.
26
26
 
@@ -28,7 +28,7 @@ Depth is a separate axis, asked at Phase 0 Step 7.5 rather than encoded in the c
28
28
 
29
29
  **vs `/multi-agent:autopilot`**: no worktree is created; work stays on the current branch. The task runs in the same editor / IDE session - no second folder.
30
30
 
31
- **vs `/multi-agent:local`**: Phase 2 Plan Approval Gate is skipped (the safety classifier still runs) and Phase 6 commit / PR runs without confirmation. Both modes already skip Phase 5 (User Test) because neither has a worktree, so the same seven phases execute - review / triage / build discipline are preserved.
31
+ **vs `/multi-agent:local`**: Phase 2 Plan Approval Gate is skipped (the safety classifier still runs) and Phase 4 commit / PR runs without confirmation. Both modes already skip Phase 5 (User Test) because neither has a worktree, so the same seven phases execute - review / triage / build discipline are preserved.
32
32
 
33
33
  ## When to use it
34
34
 
@@ -67,7 +67,7 @@ Depth is a separate axis, asked at Phase 0 Step 7.5 rather than encoded in the c
67
67
 
68
68
  ## Delegation
69
69
 
70
- Orchestrator routing: the routing table in `$HOME/.claude/commands/multi-agent/SKILL.md` resolves `local-autopilot` as the union of the `local` + `autopilot` mode mixins. Contract details: `$HOME/.claude/multi-agent-refs/phases/phase-0-init.md` Step 6 (local branch) + `$HOME/.claude/multi-agent-refs/phases/phase-2-planning.md` Step 5 (autopilot gate skip + safety classifier).
70
+ Orchestrator routing: the routing table in `$HOME/.claude/commands/multi-agent/SKILL.md` resolves `local-autopilot` as the union of the `local` + `autopilot` mode mixins. Contract details: `$HOME/.claude/multi-agent-refs/phases/phase-0-init.md` Step 6 (local branch) + `$HOME/.claude/multi-agent-refs/phases/phase-1-plan.md` Step 5 (autopilot gate skip + safety classifier).
71
71
  ## Required: outward-facing payload contracts
72
72
 
73
73
  Before writing anything outward-facing - PR body, Jira comment, Confluence page, closing report - load `$HOME/.claude/multi-agent-refs/payload-contracts.md`. It names the canonical section set for each payload, the markup dialect per surface (PR body is Markdown, Jira is wiki markup - mixing them is a defect), and the token/duration numbers the closing report must carry. Improvising a payload shape from memory is the most common failure of the short modes.
@@ -88,7 +88,7 @@ Two channels run in parallel at every phase boundary:
88
88
  ```bash
89
89
  # Phase 0, very first shell call (every CLI):
90
90
  bash $HOME/.claude/scripts/phase-tracker.sh init "$TASK_ID"
91
- for p in "0:Init" "1:Analysis" "2:Planning" "3:Dev" "4:Review" "6:Commit" "7:Report"; do
91
+ for p in "0:Init" "1:Plan" "2:Dev" "3:Review" "4:Commit" "5:Report"; do
92
92
  bash $HOME/.claude/scripts/phase-tracker.sh add "${p%%:*}" "${p#*:}"
93
93
  done
94
94
  bash $HOME/.claude/scripts/phase-tracker.sh update 0 in_progress
@@ -108,7 +108,7 @@ In Claude Code the agent MUST also drive the native TaskList widget so the user
108
108
 
109
109
  ```text
110
110
  # Register one tile per phase, capture the taskId, persist it:
111
- for each phase in 0:Init, 1:Analysis, 2:Planning, 3:Dev, 4:Review, 6:Commit, 7:Report:
111
+ for each phase in 0:Init, 1:Plan, 2:Dev, 3:Review, 4:Commit, 5:Report:
112
112
  TaskCreate({ subject: "Phase <N>: <Name>", activeForm: "<doing-form>" })
113
113
  -> returns taskId
114
114
  bash $HOME/.claude/scripts/phase-tracker.sh meta <N> tasklist_id "<taskId>"
@@ -125,11 +125,11 @@ TaskUpdate({ taskId: <saved>, status: "completed" })
125
125
  bash $HOME/.claude/scripts/phase-tracker.sh update <N> completed
126
126
  ```
127
127
 
128
- `--local autopilot` mode does NOT TaskCreate phases 5 - those are not part of the `--local autopilot` phase set (`0:Init 1:Analysis 2:Planning 3:Dev 4:Review 6:Commit 7:Report`). Only register tiles for the active set.
128
+ `--local autopilot` mode TaskCreates all 6 phases (no phase is skipped).
129
129
 
130
130
  #### TaskCreate ordering (strict)
131
131
 
132
- **All TaskCreate calls in a batch fire in strict phase-number order BEFORE any TaskUpdate is applied.** For `--local autopilot` that means: Phase 0 → Phase 1 → Phase 2 → Phase 3 → Phase 4 → Phase 6 → Phase 7. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks. Full ordering contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
132
+ **All TaskCreate calls in a batch fire in strict phase-number order BEFORE any TaskUpdate is applied.** For `--local autopilot` that means: Phase 0 → Phase 1 → Phase 2 → Phase 3 → Phase 4 → Phase 5. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks. Full ordering contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
133
133
 
134
134
  ### Visual channel - Copilot CLI / plain shell
135
135
 
@@ -15,7 +15,7 @@ Show the task's detailed agent-log.md report.
15
15
  1. **Parse task ID** - extract the `#N` form from the argument.
16
16
  - No argument → find the most recent (highest-ID) worktree.
17
17
 
18
- 2. **Find the worktree** - search the known repos. A task whose worktree was removed after its PR (Phase 6 finalize) is found under `$HOME/.claude/logs/multi-agent/<project>/<task-id>/` instead; the `agent-log.md` was always there, and `artifacts/agent-state.json` holds the state. Look there before reporting "task not found":
18
+ 2. **Find the worktree** - search the known repos. A task whose worktree was removed after its PR (Phase 4 finalize) is found under `$HOME/.claude/logs/multi-agent/<project>/<task-id>/` instead; the `agent-log.md` was always there, and `artifacts/agent-state.json` holds the state. Look there before reporting "task not found":
19
19
  ```bash
20
20
  find ~/my-ios-app/.worktrees/ ~/my-figma-app/.worktrees/ ~/my-ui-components/.worktrees/ -name "agent-state.json" -maxdepth 2 2>/dev/null
21
21
  ```
@@ -23,7 +23,7 @@ Show the task's detailed agent-log.md report.
23
23
 
24
24
  3. **Read agent-log.md** - display the file.
25
25
 
26
- 4. **Cost + durations come from the tracker, not only the log.** The tracker contract promises token/duration data "surfaced in `:log` reports", and agent-log.md only carries it when Phase 7 appended the Cost Breakdown. When that section is absent, render it live:
26
+ 4. **Cost + durations come from the tracker, not only the log.** The tracker contract promises token/duration data "surfaced in `:log` reports", and agent-log.md only carries it when Phase 5 appended the Cost Breakdown. When that section is absent, render it live:
27
27
  ```bash
28
28
  bash $HOME/.claude/scripts/render-agent-log-cost.sh "<task-id>" 2>/dev/null
29
29
  ```
@@ -1,9 +1,9 @@
1
1
  ---
2
- description: "Switch to the active task's branch and prepare it for manual testing in Xcode. Phase 5 standalone (the UI Bug Hunter lives at /multi-agent:test). Use when a finished change needs trying by hand on a device or simulator."
3
- description-tr: "Aktif görevin branch'ine geçer ve Xcode'da manuel test için hazırlar. Faz 5'in bağımsız hali (UI Bug Hunter /multi-agent:test'te)."
2
+ description: "Switch to the active task's branch and prepare it for manual testing in Xcode. Phase 3 user-test step, standalone (the UI Bug Hunter lives at /multi-agent:test). Use when a finished change needs trying by hand on a device or simulator."
3
+ description-tr: "Aktif görevin branch'ine geçer ve Xcode'da manuel test için hazırlar. Faz 3 kullanıcı-test adımının bağımsız hali (UI Bug Hunter /multi-agent:test'te)."
4
4
  ---
5
5
 
6
- # multi-agent manual-test - Phase 5 Manual Test Mode
6
+ # multi-agent manual-test - Phase 3 Manual Test Mode
7
7
 
8
8
  **Input**: $ARGUMENTS
9
9
 
@@ -13,7 +13,7 @@ Lets you switch to the task branch for manual testing in Xcode before the PR is
13
13
 
14
14
  1. **Find the task** - parse `#N` from the argument or pick the most recent task
15
15
 
16
- 2. **Check state** - the task must have completed Phase 3 or later
16
+ 2. **Check state** - the task must have completed Phase 2 (Dev) or later
17
17
  - Phase < 3 → "No code yet, run `resume #N` first"
18
18
 
19
19
  3. **Remove the worktree and switch to the branch**:
@@ -25,8 +25,8 @@ Lets you switch to the task branch for manual testing in Xcode before the PR is
25
25
 
26
26
  Then mark the tracker as waiting (this CONTINUES the task's existing `tracker-state.json`; never `init` here - see `$HOME/.claude/multi-agent-refs/tracker-contract.md` "Continuation runs"). On Claude Code in a fresh session, first rebuild the TaskList from state per the same contract:
27
27
  ```bash
28
- bash $HOME/.claude/scripts/phase-tracker.sh update 5 in_progress
29
- bash $HOME/.claude/scripts/phase-tracker.sh now 5 "awaiting local test (user)"
28
+ bash $HOME/.claude/scripts/phase-tracker.sh update 3 in_progress
29
+ bash $HOME/.claude/scripts/phase-tracker.sh now 3 "awaiting local test (user)"
30
30
  bash $HOME/.claude/scripts/phase-tracker.sh render
31
31
  ```
32
32
 
@@ -40,7 +40,7 @@ Lets you switch to the task branch for manual testing in Xcode before the PR is
40
40
  • Manual test on the simulator
41
41
 
42
42
  Reply:
43
- ✅ "ok" → continue to Phase 6 (commit)
43
+ ✅ "ok" → continue to Phase 4 (commit)
44
44
  ❌ "fix: ..." → recreate the worktree and apply the fix
45
45
  ```
46
46
 
@@ -51,5 +51,5 @@ Lets you switch to the task branch for manual testing in Xcode before the PR is
51
51
  ```json
52
52
  {"criteria":[{"spec":"<quote>","source":"analysis 15.2 | plan task 3 | user","observed":"<what was seen>","verdict":"pass|fail|not-tested","reason":"<required when not-tested>","screenshot":"<path or null>"}],"verdict":"passed|failed"}
53
53
  ```
54
- then run `node $HOME/.claude/scripts/evidence-gate.mjs --claim manual --status passed --evidence "$WORKTREE/.pipeline/manual-test.json"`, adding `--require-screenshot` when `state.visualEvidence.required` is true (a passing criterion then has to name a screenshot that is actually on disk). Exit 1 means the "ok" is not accepted: name the criterion that is missing evidence and wait for the next reply. Exit 0 → `phase-tracker.sh update 5 completed` + `phase-tracker.sh meta 5 Result "local test passed (user)"`, recreate the worktree, continue to Phase 6. Full contract: `$HOME/.claude/multi-agent-refs/phases/phase-5-test.md` step 5, whose "UI flow video, when Phase 3 produced none" block applies here too: invoked standalone, this command is the only phase that ran, so if `visualEvidence.required` is set and `visualEvidence.video.file` is empty, the recording has to happen here or nowhere.
55
- - **Fix needed** → `phase-tracker.sh now 5 "applying fix: <summary>"`, recreate the worktree, apply the fix
54
+ then run `node $HOME/.claude/scripts/evidence-gate.mjs --claim manual --status passed --evidence "$WORKTREE/.pipeline/manual-test.json"`, adding `--require-screenshot` when `state.visualEvidence.required` is true (a passing criterion then has to name a screenshot that is actually on disk). Exit 1 means the "ok" is not accepted: name the criterion that is missing evidence and wait for the next reply. Exit 0 → `phase-tracker.sh update 3 completed` + `phase-tracker.sh meta 3 Result "local test passed (user)"`, recreate the worktree, continue to Phase 4. Full contract: `$HOME/.claude/multi-agent-refs/phases/phase-3-review.md` step 5, whose "UI flow video, when Phase 2 produced none" block applies here too: invoked standalone, this command is the only phase that ran, so if `visualEvidence.required` is set and `visualEvidence.video.file` is empty, the recording has to happen here or nowhere.
55
+ - **Fix needed** → `phase-tracker.sh now 3 "applying fix: <summary>"`, recreate the worktree, apply the fix
@@ -0,0 +1,69 @@
1
+ ---
2
+ description: "Turn the top model rung on or off and keep the cost ledger's pricing in step with it. Use when asked to enable or disable Fable, or which model the pipeline dispatches."
3
+ description-tr: "Üst model basamağını açıp kapatır ve maliyet defterinin fiyatlamasını onunla birlikte taşır. Fable'ı aç/kapat sorulduğunda ya da pipeline hangi modeli çağırıyor sorulduğunda kullanılır."
4
+ argument-hint: "[on | off] - no argument reports the current state"
5
+ ---
6
+
7
+ # multi-agent model - which rung the pipeline dispatches on
8
+
9
+ The ladder is `fable -> opus -> sonnet -> haiku` and it does not change here.
10
+ What changes is whether the top rung is in play at all, which is
11
+ `prefs.global.modelFallback.fableEnabled` and ships `false`.
12
+
13
+ ```bash
14
+ bash "$HOME/.claude/lib/model-rung.sh" ${ARGUMENTS}
15
+ ```
16
+
17
+ ## Why this command exists
18
+
19
+ The switch existed before the command did. It could only be flipped by editing
20
+ preferences by hand, and `model-fallback.md` said so in a line most people never
21
+ reached. That is a knob with no handle.
22
+
23
+ It also never travelled alone. `prefs.global.costBudget.pricingModel` defaults to
24
+ `fable` so the estimate stays an upper bound; with the rung off, that default
25
+ prices every call above what it can cost, trips the budget ceiling early, and
26
+ triggers a downgrade nobody needed. The documentation asked the user to set both.
27
+ This command sets both, together:
28
+
29
+ | `fableEnabled` | `costBudget.pricingModel` |
30
+ |---|---|
31
+ | `true` | `fable` |
32
+ | `false` | `opus` |
33
+
34
+ Flipping one without the other is the defect this closes, so the pair moves as
35
+ one write or not at all.
36
+
37
+ ## What it reports with no argument
38
+
39
+ The current rung, the pricing model, and **what the switch means on this host** -
40
+ because it does not mean the same thing on all three:
41
+
42
+ | Host | What `off` does | Why |
43
+ |---|---|---|
44
+ | Claude Code | every `preferredModel: fable` persona (architects, reviewer 1, triage) starts on `opus` | the only host where the rung is Fable 5 |
45
+ | Copilot CLI | nothing | Fable 5 is not offered there; its personas never sat on this rung |
46
+ | Codex CLI | nothing, deliberately | the `fable` rung there means `gpt-5.6 @ xhigh`, a different model on a different account - switching it from a knob named after an Anthropic model would surprise a Codex user |
47
+
48
+ So on two of the three hosts this command is a **status report**, not a switch,
49
+ and it says which one it is rather than claiming a change it did not make.
50
+
51
+ ## The consequence it prints when turning the rung off
52
+
53
+ Phase 3's reviewer panel collapses from three models to two. Reviewer 1 lands on
54
+ `opus`, which Reviewer 2 already holds, and dispatching one model twice is not
55
+ cross-model review. `consensus.reviewerCount` records `2`, and the `unverified`
56
+ verdict rule matters more rather than less - two Anthropic models agreeing on a
57
+ judgment call was already weak evidence and there is now one fewer of them.
58
+ Triage also runs on `opus`, making it the same model as Reviewer 1; the Phase 3
59
+ Step 3 anonymisation requirement covers that case and is not optional here.
60
+
61
+ This is printed at the moment of the change, not left in a doc.
62
+
63
+ ## Related
64
+
65
+ - `/multi-agent:route-on` - policy-driven rung selection per persona or phase, a
66
+ separate feature that also ships off. This command decides whether a rung
67
+ EXISTS; that one decides which rung a given call picks.
68
+ - Full fallback contract, including the three failure triggers this switch is
69
+ deliberately not one of: `$HOME/.claude/multi-agent-refs/features/model-fallback.md`
@@ -110,7 +110,7 @@ Procedure:
110
110
 
111
111
  ## Step 0c: DEV-TOOLKIT - current MCP practice for the companion toolkit
112
112
 
113
- The pipeline's hands on devices and browsers are MCP tools served by a companion repo (`multi-agent-toolkit-mcp`): Phase 5 test, `manual-test`, `design-check` and `apple-archive-compliance` all call them, and several pipeline skills declare a minimum toolkit version (see `cross-cli-contract.md`). That repo therefore has to track the MCP field, not just its own README. This step researches what current practice is and audits the toolkit against it.
113
+ The pipeline's hands on devices and browsers are MCP tools served by a companion repo (`multi-agent-toolkit-mcp`): Phase 3 test, `manual-test`, `design-check` and `apple-archive-compliance` all call them, and several pipeline skills declare a minimum toolkit version (see `cross-cli-contract.md`). That repo therefore has to track the MCP field, not just its own README. This step researches what current practice is and audits the toolkit against it.
114
114
 
115
115
  Full procedure - resolution (configuration first, never a hardcoded path; skip when nothing resolves or `enabled` is false), the 5 research axes, the audit command block, and the band-E output table + rules - lives in `$HOME/.claude/multi-agent-refs/refactor/toolkit-research.md`. Read it before running this step.
116
116
 
@@ -145,8 +145,8 @@ Output (plan band F):
145
145
  ```
146
146
  | # | Error tag | Occurrences | Users | Usual phase | Versions | Root cause (file) | Fix | In plan? |
147
147
  |---|-----------|-------------|-------|-------------|----------|-------------------|-----|----------|
148
- | 1 | 4:reviewer-json-invalid | 12 | 3 | 4 | 14.x-15.x | reviewer prompt lets prose leak | tighten schema instruction in phase-4-review.md | Yes (P0) |
149
- | 2 | phase-3-failed | 5 | 2 | 3 | 15.0.x | build step misses a stack toolchain | add preflight in phase-3-dev.md | Yes (P1) |
148
+ | 1 | 4:reviewer-json-invalid | 12 | 3 | 4 | 14.x-15.x | reviewer prompt lets prose leak | tighten schema instruction in phase-3-review.md | Yes (P0) |
149
+ | 2 | phase-3-failed | 5 | 2 | 3 | 15.0.x | build step misses a stack toolchain | add preflight in phase-2-dev.md | Yes (P1) |
150
150
  ```
151
151
 
152
152
  Rules for this band:
@@ -19,7 +19,7 @@ Resume a paused or failed task from the last successful phase.
19
19
  2. **Read + validate state** - parse `agent-state.json`:
20
20
  - Validate first: `node $HOME/.claude/scripts/validate-state.mjs <state-file>` (resume-safety check, tolerant of legacy shapes). On non-zero exit, do NOT guess a phase - surface the errors and stop with `ERR: agent-state.json is unsafe to resume; inspect it or 'kill #N' and restart.`
21
21
  - Confirm the worktree (`worktreePath` / `projects[].worktreePath`) exists and is usable; if missing or locked, run the Phase 0 "Worktree stale-lock heal" before continuing.
22
- - **Unless `state.worktreeRemovedAt` is set.** Then the worktree was removed on purpose by Phase 6 once the PR opened, the branch is still local, and the artefacts live under `state.artifactsPath`. Do NOT heal or recreate it: read state from `artifactsPath`, and if the remaining work needs a checkout (a Phase 7 pause needs none), ask before moving the user's HEAD - they may be mid-work on another branch, which is exactly why the removal did not check the branch out.
22
+ - **Unless `state.worktreeRemovedAt` is set.** Then the worktree was removed on purpose by Phase 4 once the PR opened, the branch is still local, and the artefacts live under `state.artifactsPath`. Do NOT heal or recreate it: read state from `artifactsPath`, and if the remaining work needs a checkout (a Phase 5 pause needs none), ask before moving the user's HEAD - they may be mid-work on another branch, which is exactly why the removal did not check the branch out.
23
23
  - `currentPhase` - last completed phase
24
24
  - `status` - `paused` | `failed` | `in_progress`
25
25
  - `haltReason` - if set, show it so the user knows why the run stopped; clear it on successful re-entry
@@ -30,7 +30,7 @@ Resume a paused or failed task from the last successful phase.
30
30
  - **Handoff first (v10.8.0)**: read the LATEST `## Handoff` block in `agent-log.md` - it carries done/remaining/decisions/open-findings and the exact re-entry point (phase + subStep). When present, it is the primary context source; cross-check its `Next:` line against `state.currentPhase` and trust state on mismatch (state is the machine truth, handoff is the narrative).
31
31
  - Fall back to per-phase findings for logs written before v10.8 (no handoff blocks):
32
32
  - Phase 1 analysis → use it from Phase 2+
33
- - Phase 2 plan → use it from Phase 3+
33
+ - Phase 1 plan → use it from Phase 2+
34
34
  - Phase 3 code → already in the worktree
35
35
  - Recent `git log --oneline -10` in the worktree grounds what was actually committed vs. claimed.
36
36
 
@@ -42,13 +42,13 @@ Resume a paused or failed task from the last successful phase.
42
42
  run re-enters THAT step rather than the next phase. `currentPhase + 1` is the fallback,
43
43
  not the rule - a run that stopped mid-phase to ask a human has `currentPhase` pointing at
44
44
  the phase it is still inside, so resuming past it skips the question permanently. That
45
- was already true of Phase 7's channels pause, which documented itself as resumable
45
+ was already true of Phase 5's channels pause, which documented itself as resumable
46
46
  through this field while this file never mentioned it.
47
47
 
48
48
  | `waitingFor` | Re-entry |
49
49
  |---|---|
50
50
  | `maturity` | Phase 0, the maturity step, with the item **re-fetched** and re-scored - an edit is a reason to look again, never proof the gap closed (`$HOME/.claude/multi-agent-refs/features/maturity-followup.md`) |
51
- | `user-channels-choice` | Phase 7, the channels multi-select, with the stored `channelsInput` |
51
+ | `user-channels-choice` | Phase 5, the channels multi-select, with the stored `channelsInput` |
52
52
  | absent | `currentPhase + 1`, as before (same pipeline as the main multi-agent command) |
53
53
 
54
54
  Clear `waitingFor` in the same write that records the answer, the moment the step is