@mmerterden/multi-agent-pipeline 19.1.4 → 20.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (137) hide show
  1. package/CHANGELOG.md +123 -0
  2. package/README.md +19 -36
  3. package/README.tr.md +18 -35
  4. package/SECURITY.md +3 -3
  5. package/docs/adr/0002-instruction-driven-flag.md +6 -5
  6. package/docs/adr/0005-lazy-phase-docs.md +2 -2
  7. package/docs/adr/0008-installer-modularization-and-secret-leak-defense.md +1 -0
  8. package/docs/adr/0009-claude-stack-skills-plugin-only.md +1 -1
  9. package/docs/adr/0010-own-code-graph.md +5 -4
  10. package/docs/adr/0011-dormant-ci.md +10 -1
  11. package/docs/adr/0012-macos-only.md +2 -2
  12. package/docs/adr/0013-lsp-code-intelligence.md +2 -2
  13. package/docs/adr/0014-six-phase-consolidation.md +9 -9
  14. package/docs/adr/0015-one-pipeline-no-depth-answer.md +83 -0
  15. package/docs/adr/0016-the-run-shape-is-asked-not-typed.md +69 -0
  16. package/docs/adr/README.md +18 -16
  17. package/docs/architecture.md +2 -2
  18. package/docs/ecosystem.md +5 -5
  19. package/docs/facts.json +7 -9
  20. package/docs/features.md +4 -5
  21. package/docs/token-budget-history.md +1 -1
  22. package/install/_codex-agents.mjs +1 -1
  23. package/install/_common.mjs +9 -1
  24. package/install/templates/copilot-instructions.md +7 -16
  25. package/manifest.json +133 -129
  26. package/package.json +1 -1
  27. package/pipeline/agents/code-reviewer.md +2 -2
  28. package/pipeline/agents/dev-critic.md +5 -5
  29. package/pipeline/agents/security-auditor.md +80 -72
  30. package/pipeline/commands/figma-to-swiftui.md +1 -1
  31. package/pipeline/commands/multi-agent/SKILL.md +7 -9
  32. package/pipeline/commands/multi-agent/analysis/SKILL.md +2 -0
  33. package/pipeline/commands/multi-agent/analysis-jira/SKILL.md +2 -0
  34. package/pipeline/commands/multi-agent/analysis-resolve/SKILL.md +2 -0
  35. package/pipeline/commands/multi-agent/autopilot/SKILL.md +2 -0
  36. package/pipeline/commands/multi-agent/autopilot-on/SKILL.md +2 -0
  37. package/pipeline/commands/multi-agent/autopilot-status/SKILL.md +1 -1
  38. package/pipeline/commands/multi-agent/build-optimize/SKILL.md +2 -0
  39. package/pipeline/commands/multi-agent/channels/SKILL.md +2 -2
  40. package/pipeline/commands/multi-agent/create-jira/SKILL.md +2 -0
  41. package/pipeline/commands/multi-agent/design-check/SKILL.md +1 -1
  42. package/pipeline/commands/multi-agent/diff-explain/SKILL.md +1 -1
  43. package/pipeline/commands/multi-agent/forget/SKILL.md +2 -0
  44. package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +4 -2
  45. package/pipeline/commands/multi-agent/help/SKILL.md +23 -27
  46. package/pipeline/commands/multi-agent/ios-coding-standard/SKILL.md +5 -4
  47. package/pipeline/commands/multi-agent/issue/SKILL.md +2 -0
  48. package/pipeline/commands/multi-agent/jira/SKILL.md +2 -0
  49. package/pipeline/commands/multi-agent/language/SKILL.md +2 -0
  50. package/pipeline/commands/multi-agent/prune-logs/SKILL.md +2 -0
  51. package/pipeline/commands/multi-agent/purge/SKILL.md +2 -0
  52. package/pipeline/commands/multi-agent/resume/SKILL.md +177 -48
  53. package/pipeline/commands/multi-agent/save/SKILL.md +2 -0
  54. package/pipeline/commands/multi-agent/scan/SKILL.md +2 -2
  55. package/pipeline/commands/multi-agent/security-review/SKILL.md +52 -0
  56. package/pipeline/commands/multi-agent/stack/SKILL.md +2 -0
  57. package/pipeline/commands/multi-agent/sync/SKILL.md +7 -8
  58. package/pipeline/commands/multi-agent/test-screenshots/SKILL.md +2 -0
  59. package/pipeline/commands/multi-agent/uninstall/SKILL.md +2 -0
  60. package/pipeline/lib/repo-hygiene.sh +1 -1
  61. package/pipeline/multi-agent-refs/analysis/render.md +1 -1
  62. package/pipeline/multi-agent-refs/analysis/resolve.md +1 -1
  63. package/pipeline/multi-agent-refs/analysis/synthesis.md +1 -1
  64. package/pipeline/multi-agent-refs/analysis-template.md +1 -1
  65. package/pipeline/multi-agent-refs/component-dispatch.md +5 -13
  66. package/pipeline/multi-agent-refs/cross-cli-contract.md +14 -15
  67. package/pipeline/multi-agent-refs/features/external-context-injection.md +2 -0
  68. package/pipeline/multi-agent-refs/features/review-delta.md +1 -1
  69. package/pipeline/multi-agent-refs/features/review-multi-repo.md +3 -3
  70. package/pipeline/multi-agent-refs/features/security-audit.md +55 -0
  71. package/pipeline/multi-agent-refs/features/skill-conformance.md +1 -1
  72. package/pipeline/multi-agent-refs/features/visual-evidence.md +2 -1
  73. package/pipeline/multi-agent-refs/features/worktree-finalize.md +1 -1
  74. package/pipeline/multi-agent-refs/generate-issue.md +2 -0
  75. package/pipeline/multi-agent-refs/issue-jira-triad.md +2 -0
  76. package/pipeline/multi-agent-refs/keychain.md +2 -0
  77. package/pipeline/multi-agent-refs/knowledge.md +0 -7
  78. package/pipeline/multi-agent-refs/outside-the-pipeline.md +1 -1
  79. package/pipeline/multi-agent-refs/payload-contracts.md +1 -1
  80. package/pipeline/multi-agent-refs/phases/modes.md +33 -109
  81. package/pipeline/multi-agent-refs/phases/operations.md +2 -0
  82. package/pipeline/multi-agent-refs/phases/phase-0-init.md +23 -42
  83. package/pipeline/multi-agent-refs/phases/phase-1-plan.md +7 -18
  84. package/pipeline/multi-agent-refs/phases/phase-2-dev.md +13 -44
  85. package/pipeline/multi-agent-refs/phases/phase-3-review.md +28 -37
  86. package/pipeline/multi-agent-refs/phases/phase-4-commit.md +6 -6
  87. package/pipeline/multi-agent-refs/phases/phase-5-report.md +3 -3
  88. package/pipeline/multi-agent-refs/phases.md +9 -11
  89. package/pipeline/multi-agent-refs/progress-contract.md +1 -1
  90. package/pipeline/multi-agent-refs/readiness-review.md +2 -0
  91. package/pipeline/multi-agent-refs/rules.md +1 -1
  92. package/pipeline/multi-agent-refs/threat-model.md +39 -0
  93. package/pipeline/multi-agent-refs/tracker-contract.md +9 -40
  94. package/pipeline/multi-agent-refs/wiki-capture.md +3 -2
  95. package/pipeline/preferences-template.json +2 -2
  96. package/pipeline/rules/figma-pipeline.md +1 -1
  97. package/pipeline/schemas/agent-state.schema.json +28 -10
  98. package/pipeline/schemas/migrations/prefs-2.7.0-to-2.8.0.mjs +33 -0
  99. package/pipeline/schemas/phases.json +4 -26
  100. package/pipeline/schemas/prefs.schema.json +5 -9
  101. package/pipeline/schemas/reviewer-output.schema.json +99 -2
  102. package/pipeline/schemas/security-finding.schema.json +144 -0
  103. package/pipeline/scripts/_stack-routing.mjs +1 -0
  104. package/pipeline/scripts/cost-table.json +1 -1
  105. package/pipeline/scripts/gc-abandoned.sh +16 -9
  106. package/pipeline/scripts/gc-refs.sh +1 -1
  107. package/pipeline/scripts/gen-mode-dispatch.mjs +11 -41
  108. package/pipeline/scripts/migrate-prefs.mjs +18 -17
  109. package/pipeline/scripts/phase-tracker.sh +2 -2
  110. package/pipeline/scripts/phase0-exit-gate.mjs +1 -1
  111. package/pipeline/scripts/plan-coverage-gate.mjs +3 -3
  112. package/pipeline/scripts/render-work-summary.sh +7 -4
  113. package/pipeline/scripts/run-aggregator.mjs +1 -1
  114. package/pipeline/scripts/usage-report.mjs +0 -2
  115. package/pipeline/scripts/worktree-finalize.sh +2 -2
  116. package/pipeline/skills/.skill-manifest.json +17 -21
  117. package/pipeline/skills/.skills-index.json +6 -39
  118. package/pipeline/skills/shared/README.md +5 -8
  119. package/pipeline/skills/shared/core/multi-agent/SKILL.md +11 -15
  120. package/pipeline/skills/shared/core/multi-agent-autopilot-status/SKILL.md +1 -1
  121. package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +13 -16
  122. package/pipeline/skills/shared/core/multi-agent-ios-coding-standard/SKILL.md +2 -3
  123. package/pipeline/skills/shared/core/multi-agent-resume/SKILL.md +51 -15
  124. package/pipeline/skills/shared/core/multi-agent-scan/SKILL.md +2 -2
  125. package/pipeline/skills/shared/core/multi-agent-security-review/SKILL.md +29 -0
  126. package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +7 -7
  127. package/pipeline/skills/shared/external/security-review/SKILL.md +64 -0
  128. package/pipeline/skills/shared/external/security-review/references/owasp-mobile-top10-2024.md +53 -0
  129. package/pipeline/skills/shared/external/security-review/references/owasp-web-api-top10-2021.md +56 -0
  130. package/pipeline/skills/skills-index.md +3 -6
  131. package/pipeline/commands/multi-agent/local/SKILL.md +0 -132
  132. package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +0 -142
  133. package/pipeline/commands/multi-agent/resume-local/SKILL.md +0 -114
  134. package/pipeline/commands/security-review.md +0 -6
  135. package/pipeline/skills/shared/core/multi-agent-local/SKILL.md +0 -41
  136. package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +0 -55
  137. package/pipeline/skills/shared/core/multi-agent-resume-local/SKILL.md +0 -51
@@ -1,18 +1,16 @@
1
- > **TLDR** - four entry commands, two axes, and depth is no longer one of them:
1
+ > **TLDR** - four entry commands, one pipeline, two axes:
2
2
  >
3
- > - **Depth** is a picker, not a flag. Phase 0 Step 7.5 asks Full or Short and sets `state.onlyDevelop`. Short strips to Init → Dev(Opus self-contained) → Review → Test → Commit → Report. Review is never stripped.
4
- > - **`autopilot`** skips confirmations (Plan, Test, Commit, PR prompts) and therefore skips the depth question too, always running Full. Still fails safe on review blockers and build retries.
5
- > - **`--local`** means no worktree: work happens in `$PROJECT_ROOT` on a local branch.
3
+ > - **`autopilot`** skips confirmations (Plan, Test, Commit, PR prompts). Still fails safe on review blockers and build retries.
4
+ > - **Workspace** is the one question the run asks about its own shape: a worktree, or the project root on a local branch. Phase 0 Step 5b asks it; autopilot resolves it to a worktree and never asks.
6
5
  >
7
- > Fastest path is `/multi-agent:local` then Short. There is no fast-and-unattended combination any more, and that is deliberate - see "Pipeline depth".
6
+ > There is one pipeline and every mode runs its whole phase set. What a run costs follows the evidence the task carries, not an answer taken before the evidence exists.
8
7
 
9
8
  ## Autopilot Mode
10
9
 
11
10
  <!-- toc -->
12
11
  - [Autopilot Mode](#autopilot-mode)
13
- - [Pipeline depth (Full / Short)](#pipeline-depth-full-short)
14
12
  - [Analysis Mode (`/multi-agent:analysis`)](#analysis-mode-multi-agentanalysis)
15
- - [Local Mode (`--local`)](#local-mode---local)
13
+ - [Local Mode](#local-mode)
16
14
  <!-- /toc -->
17
15
 
18
16
  Autopilot mode skips interactive confirmations and runs the pipeline end-to-end autonomously.
@@ -27,15 +25,19 @@ Autopilot mode skips interactive confirmations and runs the pipeline end-to-end
27
25
  /multi-agent "LoginView dark mode fix" autopilot
28
26
  ```
29
27
 
30
- **What changes in autopilot (Tablo 1 - Full pipeline modes):**
28
+ **What changes in autopilot (Tablo 1)**
31
29
 
32
- | Phase | Normal (full) | autopilot (full) | local (full) | local autopilot (full) |
33
- | ------------------- | ------------------------------------------------------- | ------------------------------------------- | ------------------------------------------- | ------------------------------------------- |
34
- | Phase 1 (Plan Approval Gate) | Clarification (max 2 rounds) + approval loop - user: approve/abort/free-text edit | **Gate skip** - log the plan, proceed directly to Phase 2 (autopilot contract: zero interaction) | Same as Normal | Same as autopilot |
35
- | Phase 3 (User Test) | Interactive prompt ("Want to test?" -> wait) | Skip (autopilot suppresses interactive prompts) -> proceed directly to Phase 4 | **Not in the phase set** (the gate checks the change out of a worktree; local has none) | Not in the phase set |
36
- | Phase 4 (Commit) | "Want to commit?" -> wait | Auto commit + push | "Want to commit?" -> wait | Auto commit + push |
37
- | Phase 4 (PR) | "Want to open a PR?" -> wait | Auto create PR | "Want to open a PR?" -> wait | Auto create PR |
38
- | **Phase 5 (Channels)** | Multi-select channel + content menu | **STILL PAUSES** - see "Phase 5 autopilot exception" below | Multi-select (same as Normal) | **STILL PAUSES** (same as autopilot) |
30
+ Two axes, and only the first is a command: whether anything is confirmed
31
+ (`autopilot`), and where the work happens (the Step 5b answer). An autopilot run
32
+ always gets a worktree, so there is no unattended-local column.
33
+
34
+ | Phase | Interactive, worktree | Interactive, local | autopilot (always worktree) |
35
+ | --- | --- | --- | --- |
36
+ | Phase 1 (Plan Approval Gate) | Clarification (max 2 rounds) + approval loop - user: approve/abort/free-text edit | Same as worktree | **Gate skip** - log the plan, proceed directly to Phase 2 (autopilot contract: zero interaction) |
37
+ | Phase 3 (User Test step) | Interactive prompt ("Want to test?" -> wait) | **Not offered** - the step checks the change out of a worktree and there is none; the phase itself still runs | Skip (autopilot suppresses interactive prompts) -> proceed directly to Phase 4 |
38
+ | Phase 4 (Commit) | "Want to commit?" -> wait | "Want to commit?" -> wait | Auto commit + push |
39
+ | Phase 4 (PR) | "Want to open a PR?" -> wait | "Want to open a PR?" -> wait | Auto create PR |
40
+ | **Phase 5 (Channels)** | Multi-select channel + content menu | Multi-select (same) | **STILL PAUSES** - see "Phase 5 autopilot exception" |
39
41
 
40
42
  **What NEVER skips (even in autopilot):**
41
43
 
@@ -51,90 +53,20 @@ Autopilot mode skips interactive confirmations and runs the pipeline end-to-end
51
53
 
52
54
  The generic "zero-interaction" contract covers Phases 0-4 only. Phase 5 channels dispatch is the **single exception**:
53
55
 
54
- - ALL modes (`autopilot`, `--local autopilot`, and a Short run reaching Phase 5) pause at the channels multi-select menu.
56
+ - ALL modes, `autopilot` included, pause at the channels multi-select menu.
55
57
  - Menu pre-ticks from `prefs.global.reportChannels` + `prefs.global.reportContent` - user can accept with one keypress if prefs are stable.
56
58
  - **30-minute timeout** - if user does not respond, session ends cleanly:
57
59
  - External delivery aborted (no silent apply - prevents accidental Jira comments / Confluence pages).
58
60
  - Internal capture (`agent-log.md`, telemetry, knowledge base) STILL runs.
59
- - State persisted as `{phase: 7, waitingFor: "user-channels-choice", channelsTimeout: true}`.
61
+ - State persisted as `{phase: 5, waitingFor: "user-channels-choice", channelsTimeout: true}`.
60
62
  - Resume: `/multi-agent:resume <task-id>` re-opens menu with same inputs.
61
63
  - Post-hoc `/multi-agent:channels <task>` never times out - user invoked it explicitly.
62
64
 
63
65
  Full contract: `$HOME/.claude/multi-agent-refs/phases/phase-5-report.md` (Autopilot pause contract) + `commands/multi-agent/channels/SKILL.md`.
64
66
 
65
- ---
66
-
67
- ## Pipeline depth (Full / Short)
68
-
69
- Short strips the pipeline to what an already-scoped task needs: no deep analysis, no planning phase. The **Opus** dev agent reads the task scope and implements, and Phase 3 then reviews what it produced. Review is deliberately NOT part of the strip: analysis and planning shape work that has not happened yet, so a task the user has already scoped can skip them, while review judges work that now exists and has no substitute.
70
-
71
- **How it is chosen.** Phase 0 Step 7.5, after `taskType` is known:
72
-
73
- ```
74
- Bu is icin hangi pipeline? / Which pipeline for this task?
75
- 1. Tam / Full Analysis -> Plan -> Dev -> Review -> Test -> Commit -> Report
76
- 2. Kisa / Short Dev (self-contained, Opus) -> Review -> Test -> Commit -> Report
77
- ```
78
-
79
- `bugfix` and `chore` recommend Short; `feature`, `refactor` and `component` recommend Full. The recommendation is presented first and passed as `ASK_CHOICE_DEFAULT` on hosts without a native picker, because `ask-choice.sh` picks the first option on a non-TTY and option order is not a contract. Pass it as the **1-based index** (Full = 1, Short = 2), not as the label: labels render in `outputLanguage`, so a label-valued default matches nothing on a `tr` run.
80
-
81
- **Who is asked.** `/multi-agent` and `/multi-agent:local`. Both autopilot entries always run Full without asking.
82
-
83
- **Why there is no fast-and-unattended combination.** It existed until v16.0.0, and removing it was a real behaviour change, not a rename. Autopilot may not ask, so something has to choose, and unattended is the worst place to drop analysis and planning: nobody is watching to notice what the shortcut lost. A cron job or script that wants both now has to pick - stay unattended and pay for the full pipeline, or stay fast and have a person present.
84
-
85
- **Pipeline in a Short run:**
86
-
87
- ```
88
- Phase 0: Init -> Phase 2: Dev (self-contained) -> Phase 3: Review -> Phase 3: Review (user test) -> Phase 4: Commit -> Phase 5: Report
89
- ```
90
-
91
- **What changes (Tablo 2 - Short runs):**
92
-
93
- | Phase | Full | Short | Short + `--local` |
94
- | ------------------- | ----------------------------------------------- | ---------------------------------------------------------------------------------- | --------------------------------------------- |
95
- | Phase 0 (Init) | Full setup | Same - worktree, branch, state, and the depth question itself | Same - no worktree, branch on `$PROJECT_ROOT` |
96
- | Phase 1 (Analysis) | Parallel Explore agents + analysis document | **SKIP** - no tile is ever drawn for it (registration is deferred to Step 7.5) | **SKIP** |
97
- | Phase 1 (Planning) | TaskCreate + architecture review + **Plan Approval Gate** | **SKIP** (no plan means no plan gate) | **SKIP** |
98
- | Phase 2 (Dev) | Follows the Phase 1 plan, TDD cycle (Sonnet) | **Self-contained** (Opus): agent scans relevant files, implements with TDD, builds | Same, on the local branch |
99
- | Phase 3 (Review) | Parallel review + Fable triage (3 reviewers on every host: Claude Code Fable + Opus + Sonnet, Copilot GPT-5.4 + Opus + Sonnet) | **Same** - gates, parallel review, triage; blocking findings return to Phase 2 (cap 3) | **Same**, on the local branch diff |
100
- | Phase 3 (User Test) | Interactive prompt | **Interactive prompt** | Not in the set - no worktree to check out |
101
- | Phase 4 (Commit) | Commit + PR | Same - still asks | Same |
102
- | Phase 5 (Report) | Full report + channels multi-select | Simplified - no analysis section, review section IS present, channels menu still pauses | Same |
103
-
104
- **Phase 2 in a Short run (self-contained):**
105
-
106
- The **Opus** agent receives the task description (from Jira, GitHub issue, or free-text) and:
107
-
108
- 1. Quickly scans relevant files in the codebase (lightweight, not full Explore)
109
- 2. Implements the change following TDD cycle (test -> code -> build)
110
- 3. Runs build verification (max 3 retries on failure)
111
-
112
- No separate task breakdown - the agent handles scope autonomously.
113
-
114
- **State tracking**: `agent-state.json` gets `"onlyDevelop": true`. The tracker boots at Step -1 with Phase 0 alone and the rest of the tiles are registered at Step 7.5, once this answer says which phases the run actually has - so a Short run never draws an Analysis tile it will not use. Full contract: `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "Deferred registration".
115
-
116
- ### Intake warnings for a Short run
117
-
118
- **An analysis document was supplied.** Because Phase 1 and Phase 1 are skipped, no phase turns that document into a task breakdown. The doc becomes raw context for one Dev pass, and work comes out ordered by whatever the model read first: the bottom of the dependency chain lands, the screen wiring does not.
119
-
120
- The depth picker is where this is caught. When the intake carried an analysis document or a Figma reference, say so in the question itself rather than after the choice:
121
-
122
- ```
123
- This task carries an analysis document. Short skips Analysis and Planning, so
124
- the document will not be turned into a task breakdown.
125
- 1. Tam / Full (recommended here)
126
- 2. Kisa / Short (doc as context only)
127
- ```
128
-
129
- Autopilot never sees this: it runs Full.
130
-
131
- **The branch already carries the work.** When the change was developed outside the pipeline, or by hand, Short is still the wrong entry point: it will try to develop again. `/multi-agent:resume-local` puts the existing diff through the same review, adds a build+test success gate, opens the PR, and posts the Jira technical-analysis + test-scenario comment - without re-developing. Offer it when the working tree or branch is already ahead of the base with the task's changes.
132
-
133
- ---
134
-
135
67
  ## Analysis Mode (`/multi-agent:analysis`)
136
68
 
137
- Not a depth permutation - a different **output kind**. It produces the analysis document and stops; no code, no branch, no PR. The 6-phase contract holds, with four phases reinterpreted the same way a Short run reinterprets Phase 2.
69
+ A different **output kind**, not a shorter pipeline. It produces the analysis document and stops; no code, no branch, no PR. The 6-phase contract holds, with four phases reinterpreted the same way a Short run reinterprets Phase 2.
138
70
 
139
71
  ```
140
72
  Phase 0: Init -> Phase 1: Plan (analysis) -> Phase 1: Plan -> Phase 3: Review -> Phase 4: Publish -> Phase 5: Report
@@ -151,18 +83,19 @@ Phase 0: Init -> Phase 1: Plan (analysis) -> Phase 1: Plan -> Phase 3: Review ->
151
83
  | 6 Commit | **Publish** instead: Local file / Confluence / Jira. Locked 6 still forbids `git add` and `git commit` |
152
84
  | 7 Report | Channels, same as every mode |
153
85
 
154
- **No `local` or `autopilot` variant.** Worktree isolation buys nothing when no code is written, and the intake, the Pass B convention preview and the open-question resolution are interactive by nature; a zero-interaction analysis would be a document nobody agreed to.
86
+ **No `autopilot` variant.** Worktree isolation buys nothing when no code is written, and the intake, the Pass B convention preview and the open-question resolution are interactive by nature; a zero-interaction analysis would be a document nobody agreed to.
155
87
 
156
88
  **Where the document lands.** Drafts go to `/tmp/analysis-<slug>-<ts>/` first, then the Phase 4 picker decides: Local writes `<repo>/analysis/<feature>-<platform>.md` uncommitted, Confluence posts the page, Jira updates the description. In the full pipeline the same engine writes into the worktree instead and `prefs.global.analysisPhase.commitDoc` decides whether it rides along with the commit.
157
89
 
158
90
  ---
159
91
 
160
- ## Local Mode (`--local`)
92
+ ## Local Mode
161
93
 
162
94
  Local mode skips worktree creation - works directly on a local branch in the project root. Useful for single-task workflows or when worktrees cause issues.
163
95
 
164
- **Activation**: answer the Phase 0 Step 5b workspace question, or state it up
165
- front so the question resolves without being asked.
96
+ **Activation**: answer the Phase 0 Step 5b workspace question. There is no flag
97
+ and no command name for it - the answer is the only way in, which is why the
98
+ question carries what the choice costs.
166
99
 
167
100
  ```
168
101
  Bu is nerede kossun? / Where should this task run?
@@ -171,25 +104,16 @@ Bu is nerede kossun? / Where should this task run?
171
104
  ```
172
105
 
173
106
  Two genuine options, so it meets the two-option floor in `picker-contract.md` on
174
- its own. **Say what local costs inside the question**: Phase 3 is not in a local
175
- run's set (the user-test gate checks the change out of a worktree, and there is
176
- none), and uncommitted work in the project root is in the way of the checkout. A
177
- user choosing local should learn both before choosing, not after.
178
-
179
- The ways to state it up front:
180
-
181
- ```
182
- /multi-agent:local "PROJ-12345"
183
- /multi-agent:local "#316"
184
- /multi-agent:local "LoginView dark mode fix"
185
- /multi-agent:local-autopilot "PROJ-12345"
186
- /multi-agent "PROJ-12345" --local
187
- ```
107
+ its own. **Say what local costs inside the question**: the user-test STEP inside
108
+ Phase 3 checks the change out of a worktree, so a local run does not get that
109
+ offer (the phase itself still runs), and uncommitted work in the project root is
110
+ in the way of the checkout. A user choosing local should learn both before
111
+ choosing, not after.
188
112
 
189
113
  Every autopilot entry resolves this to **worktree** and never asks. That is not a
190
114
  skipped question: an unattended run commits and pushes from wherever it stands,
191
- and doing that in the user's own checkout is what worktrees exist to prevent.
192
- `:local-autopilot` is the explicit opt-out.
115
+ and doing that in the user's own checkout is what worktrees exist to prevent. An
116
+ unattended run therefore always gets its own checkout.
193
117
 
194
118
  **What changes in local mode:**
195
119
 
@@ -221,4 +145,4 @@ git -C $PROJECT_ROOT config user.email "{identity.email}"
221
145
  - Uncommitted changes in PROJECT_ROOT may conflict - pipeline warns if dirty
222
146
  - Build queue lock still applies for xcodebuild
223
147
 
224
- **Combinable**: `/multi-agent:local-autopilot` is the no-worktree unattended path; `/multi-agent:local` then Short is the no-worktree fast path.
148
+ Fast path.
@@ -17,6 +17,8 @@
17
17
  - [Phase Pipeline](#phase-pipeline)
18
18
  <!-- /toc -->
19
19
 
20
+ > **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
21
+
20
22
  Every task gets an auto-incremented short ID. Counter stored at `$HOME/.claude/logs/multi-agent/{project}/.counter` (persists across sessions).
21
23
 
22
24
  ```
@@ -14,23 +14,21 @@ $HOME/.claude/scripts/phase-tracker.sh tiles
14
14
  $HOME/.claude/scripts/phase-tracker.sh update 0 in_progress
15
15
  ```
16
16
 
17
- **Phase 0 only.** `/multi-agent` and `:local` do not know their phase set yet -
18
- depth decides it, and depth is Step 7.5 - so they register the rest there rather
19
- than drawing eight tiles the user has not chosen. Every other mode registers its
20
- whole set here, in phase-number order. Contract: `tracker-contract.md`,
21
- "Deferred registration".
17
+ **Every mode registers its whole set here**, in phase-number order. The phase
18
+ set is known at Phase 0 because there is one pipeline: no answer later in the
19
+ run can add or remove a phase. Contract: `tracker-contract.md`.
22
20
 
23
21
  `tiles` prints this host's widget-registration calls: **make them before continuing.** The card alone lands in collapsed tool output, so a run that skips them runs in silence. Contract: `tracker-contract.md`, "The card is not the widget".
24
22
 
25
23
  If `INPUT_TASK_ID` isn't known yet (free-text, project not selected), use a placeholder; rename later via `mv` once parsed in Step 1.
26
24
 
27
- Every subsequent phase (1-7) MUST call `phase-tracker.sh update <N> in_progress` on entry and `phase-tracker.sh update <N> completed|failed|skipped` on exit. Each `update` prints a `-- NEXT (required) --` block: act on it. Sub-phase milestones use `phase-tracker.sh sub <N> <subN> "<name>" <status>`. See `$HOME/.claude/multi-agent-refs/phases.md` "Visual Phase Tracker" for the full contract.
25
+ Every subsequent phase (1-5) MUST call `phase-tracker.sh update <N> in_progress` on entry and `phase-tracker.sh update <N> completed|failed|skipped` on exit. Each `update` prints a `-- NEXT (required) --` block: act on it. Sub-phase milestones use `phase-tracker.sh sub <N> <subN> "<name>" <status>`. See `$HOME/.claude/multi-agent-refs/phases.md` "Visual Phase Tracker" for the full contract.
28
26
 
29
- `update <N> completed` **exits 3** for phases 1-4 with no recorded spend: record `model` + `tokens`, or pass `--no-llm`, then re-run it. Contract: `tracker-contract.md`, "Accounting is a gate".
27
+ `update <N> completed` **exits 3** for phases 1-5 with no recorded spend: record `model` + `tokens`, or pass `--no-llm`, then re-run it. Contract: `tracker-contract.md`, "Accounting is a gate".
30
28
 
31
29
  ##### TaskCreate ordering on Claude Code (strict)
32
30
 
33
- On Claude Code, fire every `TaskCreate` in a registration batch in strict phase-number order BEFORE any `TaskUpdate` in that batch - which is the order `tiles` prints them in - and never register a phase whose number is below one already registered. A deferred batch (Step 7.5) therefore appends, it does not interleave. Full contract: `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
31
+ On Claude Code, fire every `TaskCreate` in a registration batch in strict phase-number order BEFORE any `TaskUpdate` in that batch - which is the order `tiles` prints them in - and never register a phase whose number is below one already registered. Full contract: `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
34
32
 
35
33
  ---
36
34
 
@@ -46,6 +44,13 @@ OUTPUT_LANG=$(jq -r '.global.outputLanguage // "en"' "$PREFS_FILE" 2>/dev/null |
46
44
 
47
45
  From this point on, everything the user reads renders in `$OUTPUT_LANG`: conversational lines, `AskUserQuestion` `question`/`label`/`description`, and external payload bodies (PR/Jira/Confluence). English stays only on `header`, commit messages, branch names, PR title prefixes, identifiers. Full matrix: `rules.md` "Language Application".
48
46
 
47
+ **Every picker in this phase prints its breadcrumb.** Phase 0 is one chain -
48
+ account, repos, base branch, maturity, dev-context, workspace - and each step
49
+ emits the narrator line `<localized: "Step i/n: what this step decides">` above
50
+ its question, per `picker-contract.md` "Step narration". A step that resolves
51
+ without asking (one account, no siblings) still prints its line with the
52
+ resolution noted, so the numbering reads continuously instead of jumping.
53
+
49
54
  **Model tier resolution** (same step, once per run): read `prefs.global.modelFallback`. If `fableEnabled` is `false`, every `preferredModel: fable` persona resolves to `opus` for this run and the Phase 3 Claude Code panel is 2 reviewers, not 3; print the one-line INFO. Then, if `premiumTierUntil` is set and in the past, apply the date-gate trigger - `preferredModel` personas dispatch on `fallbackModel`, with the one-line WARN. Both lines and the exact ordering: `$HOME/.claude/multi-agent-refs/features/model-fallback.md`. Dispatch-error and budget triggers apply per-dispatch later; nothing else to do here.
50
55
 
51
56
  **First-run guard**: After loading prefs, check if `keychainMapping` has at least one non-null value. If ALL values are null (template defaults - setup never ran), show:
@@ -323,8 +328,7 @@ above, recorded as `baseBranchSource: "input"`. Everything else asks, and
323
328
  One row is a normal outcome of the rule-5 filter and is **still asked**, with a
324
329
  real second option - `picker-contract.md`, "Two options or it is not a question".
325
330
 
326
- This holds in every mode. A Short run skips the *LLM* phases (Analysis, Planning); it
327
- does not skip Phase 0's pickers. Autopilot resolves them without prompting, which still
331
+ This holds in every mode. Autopilot resolves them without prompting, which still
328
332
  writes the fields; its base-branch resolution order is in `features/base-branch-evidence.md`.
329
333
 
330
334
  **TTL filter for recent branches**:
@@ -412,11 +416,11 @@ Branch name is deterministic - no user confirmation needed.
412
416
  #### Step 5b - Workspace (worktree or local)
413
417
 
414
418
  Ask where the branch lives - the wording, the two options and what local costs
415
- are in `modes.md`, "Local Mode". Here because Step 4 named the branch, Step 6b
416
- acts on the answer, and Step 7.5 needs the phase set it implies.
419
+ are in `modes.md`, "Local Mode". Here because Step 4 named the branch and Step
420
+ 6b acts on the answer.
417
421
 
418
- **Who is asked.** `/multi-agent` only (`workspaceSource: "asked"`). `:local` and
419
- `--local` state it up front (`command`).
422
+ **Who is asked.** Every interactive entry (`workspaceSource: "asked"`). There is
423
+ no flag and no command name that pre-answers it.
420
424
  Every autopilot entry resolves it to a worktree and never asks (`autopilot`).
421
425
 
422
426
  **Persist** `state.localMode` (semantics unchanged) and `state.workspaceSource`,
@@ -448,7 +452,7 @@ Log: `Identity: {identity.name} <{identity.email}>`
448
452
 
449
453
  1. `git -C $PROJECT_ROOT fetch origin`
450
454
 
451
- **If local** (Step 5b answered local, by question, by `:local` / `--local`):
455
+ **If local** (Step 5b answered local):
452
456
 
453
457
  ```bash
454
458
  if [ -n "$(git -C $PROJECT_ROOT status --porcelain)" ]; then
@@ -464,7 +468,7 @@ git -C $PROJECT_ROOT config user.email "{identity.email}"
464
468
 
465
469
  **If worktree** (Step 5b answered worktree, or autopilot resolved it): 2. Worktree path: Jira → `.worktrees/{jiraId}/`, GitHub → `.worktrees/GH{issueNo}/`, free-text → `.worktrees/task-{shortId}/` 3. **Heal stale admin state first** (see "Worktree stale-lock heal" below) and **apply the residue guard** (see "Worktree residue guard" below), then `git -C $PROJECT_ROOT worktree add {path} -b {branch} origin/{baseBranch}` (if exists: enter, pull) 4. Set identity: `git -C {worktree-path} config user.name/email` 5. Create log dir + `agent-log.md` + `agent-state.json` at `$HOME/.claude/logs/multi-agent/{project}/{task-id}/`, never inside the worktree:
466
470
 
467
- **Worktree location convention (cited by every other command):** always `{projectRoot}/.worktrees/{taskId}`, inside the repo, never under `$HOME`. `{taskId}` is the directory name from the rule above (`DC-<shortId>` for `/multi-agent:design-check`). `.worktrees` is fixed, not a preference: no `worktreeBasePath` key exists, and `gc-worktrees.sh`, `purge.sh`, the cost renderers and `usage-report.mjs` resolve `<repo>/.worktrees/` by name. Multi-repo tasks get one worktree per repo (the loop below); `--local` creates none and `worktreePath` is `$PROJECT_ROOT`.
471
+ **Worktree location convention (cited by every other command):** always `{projectRoot}/.worktrees/{taskId}`, inside the repo, never under `$HOME`. `{taskId}` is the directory name from the rule above (`DC-<shortId>` for `/multi-agent:design-check`). `.worktrees` is fixed, not a preference: no `worktreeBasePath` key exists, and `gc-worktrees.sh`, `purge.sh`, the cost renderers and `usage-report.mjs` resolve `<repo>/.worktrees/` by name. Multi-repo tasks get one worktree per repo (the loop below); a local answer creates none and `worktreePath` is `$PROJECT_ROOT`.
468
472
 
469
473
  **Worktree stale-lock heal (required before every `worktree add`):** a run killed mid-`worktree add` (OOM, SIGTERM, disk full) leaves a locked or broken admin entry under `.git/worktrees/{id}/`, so the retry fails with `fatal: '<path>' already exists`. Always run the heal first - it is a no-op on a clean repo:
470
474
 
@@ -584,30 +588,6 @@ Persist: `"taskType": "component" | "bugfix" | "feature" | "refactor" | "chore"`
584
588
 
585
589
  Log: `Phase 0 Step 7: taskType = {component|bugfix|feature|refactor|chore}`
586
590
 
587
- #### Step 7.5 - Pipeline depth (Full / Short)
588
-
589
- Ask the depth question from `$HOME/.claude/multi-agent-refs/phases/modes.md` "Pipeline depth" - it carries the wording, the per-`taskType` recommendation and the mode tables. Here because the recommendation needs `taskType` (Step 7), which needs the fetched issue (Step 1) and the branch (Step 3).
590
-
591
- **Who is asked.** `/multi-agent` and `/multi-agent:local` only. Both autopilot entries and analysis mode skip it; autopilot always runs Full.
592
-
593
- When the intake carried an analysis document or a Figma reference, say so **inside** the question: Short skips the only two phases that would turn that document into a task breakdown, and the user should learn that before choosing, not after.
594
-
595
- **Pass the default as a 1-based index, never a label.** `ask-choice.sh` takes the first option on a non-TTY, and a label-valued default matches nothing once the options render in `outputLanguage`. Reasoning: `modes.md`, "Pipeline depth".
596
-
597
- ```bash
598
- # Full is option 1, Short is option 2 (modes.md "Pipeline depth")
599
- DEPTH_DEFAULT_INDEX=1; [ "$DEPTH_RECOMMENDATION" = "short" ] && DEPTH_DEFAULT_INDEX=2
600
- ASK_CHOICE_DEFAULT="$DEPTH_DEFAULT_INDEX" \
601
- $HOME/.claude/lib/ask-choice.sh "<localized: 'Which pipeline for this task?'>" \
602
- "<localized: 'Full'>" "<localized: 'Short'>"
603
- ```
604
-
605
- **Persist.** Short sets `state.onlyDevelop = true`; Full leaves it `false`. The key is unchanged - only who sets it changed - so every downstream reader keeps working.
606
-
607
- **Now register the rest of the widget.** This answer is the first moment the phase set is known, so the remaining tiles are created here and not before - Full `1 2 3 4 5 6 7`, Short `3 4 5 6 7`, `:local` dropping 5 from either. `add` each, then `phase-tracker.sh tiles --new`, which emits TaskCreate only for phases that carry no tile yet. Contract: `tracker-contract.md`, "Deferred registration".
608
-
609
- Log: `Phase 0 Step 7.5: depth = {full|short} (recommended {full|short}, source {user|autopilot|default})`
610
-
611
591
  #### Step 7.6 - Test baseline (opt-in, `prefs.global.testBaseline.enabled`, default `false`)
612
592
 
613
593
  Phase 3 Gate 3 cannot tell an inherited red suite from one this run broke, so it blocks on someone else's bug or the agent "fixes" tests it never touched. Runs after Step 6, only when the stack has a test command; skipped in analysis mode.
@@ -644,7 +624,8 @@ Persist that file as `state.evidenceCapability`, then build the menu from it, ne
644
624
 
645
625
  A closed option keeps its row and prints the probe's reason verbatim. No `uiTestTargets` closes 2; `mcp` false closes 3; a missing `device` or `recorder` closes both. Targets present with no `matchingTests` leaves 2 open, warning that the whole UI suite will run. **When only option 1 is open, do not ask**: `testDepth = unit`, `testDepthSource = forced`, and the tier 3 gap is written.
646
626
 
647
- Default as a 1-based index, never a label - same rule and same reason as Step 7.5:
627
+ Default as a 1-based index, never a label: `ask-choice.sh` takes the first
628
+ option on a non-TTY, so a label-valued default silently becomes option 1.
648
629
 
649
630
  ```bash
650
631
  DEPTH_DEFAULT_INDEX=1
@@ -653,7 +634,7 @@ DEPTH_DEFAULT_INDEX=1
653
634
  ASK_CHOICE_DEFAULT="$DEPTH_DEFAULT_INDEX" $HOME/.claude/lib/ask-choice.sh ...
654
635
  ```
655
636
 
656
- Asked by `/multi-agent` and `:local`; autopilot reads `prefs.global.testDepth.default` and degrades to the best open option. Asked here, not Phase 3, which four of the eight modes drop.
637
+ Asked by every interactive entry; autopilot reads `prefs.global.testDepth.default` and degrades to the best open option. Asked here, not Phase 3, which four of the eight modes drop.
657
638
 
658
639
  Log: `Phase 0 Step 7.7: testDepth = {unit|unit+ui|unit+mcp} (source {user|autopilot|default|forced}), tier1/tier2 = {open|closed}`
659
640
 
@@ -1,6 +1,8 @@
1
1
  ### Phase 1: Plan (Opus)
2
2
 
3
- > **TLDR** - One phase, two halves. First the codebase is explored and the analysis document written (the design contract the rest of the run reads); then that document is decomposed into concrete tasks with file-level targets, risk grading and architecture review. The phase ends at the **Plan Approval Gate**: in normal mode the orchestrator asks structured clarification questions when scope is ambiguous (max 2 rounds), renders the plan, and loops on free-text edits until the user approves or aborts. The gate is **skipped entirely** for a Short run and for `autopilot`. Analysis and planning were Phases 1 and 2 until v19.0.0; the depth picker already skipped them as one unit (`state.onlyDevelop`), which is why they are one phase now.
3
+ > **TLDR** - One phase, two halves. First the codebase is explored and the analysis document written (the design contract the rest of the run reads); then that document is decomposed into concrete tasks with file-level targets, risk grading and architecture review. The phase ends at the **Plan Approval Gate**: in normal mode the orchestrator asks structured clarification questions when scope is ambiguous (max 2 rounds), renders the plan, and loops on free-text edits until the user approves or aborts. The gate is **skipped entirely** for `autopilot`, which has no one to ask.
4
+
5
+ > **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
4
6
 
5
7
  <!-- progress-contract: applied -->
6
8
  Progress emission per `$HOME/.claude/multi-agent-refs/progress-contract.md` - lines for each Explore dispatch, each finish, analyst synthesis start, `analysis.json` write.
@@ -441,10 +443,6 @@ Log: "Phase 1: Consistency - requirements:{N/N mapped} anchors:{ok|M unanchore
441
443
  - `state.autopilot === true` (autopilot contract: zero interaction)
442
444
  - Autopilot safety classifier returns `recommendPause: false` (see Step 5c below)
443
445
 
444
- OR:
445
-
446
- - `state.onlyDevelop === true` (Short pipeline: direct to Phase 2, no plan)
447
-
448
446
  In the skipped case, log `🧠 Phase 1: Plan - gate skipped ({mode}), proceeding to Phase 2` and go to Phase 2.
449
447
 
450
448
  ##### 5c - Autopilot safety classifier (runs before 5a/5b skip decision)
@@ -571,9 +569,8 @@ The pipeline shapes interact with the gate as follows. This table is the source
571
569
 
572
570
  | Mode | Clarification | Approval Loop | Safety Classifier | Notes |
573
571
  |---|---|---|---|---|
574
- | Full, interactive (`/multi-agent`, `:local`) | ✅ (max 2 rounds) | ✅ | - (redundant when a human approves) | Full gate |
575
- | Short (depth picker answered Short) | ❌ | ❌ | - | Phase 1-2 skipped at Step 7.5; Phase 2 starts immediately |
576
- | `autopilot`, `:local-autopilot` | ❌ | ❌ conditional | ✅ (if `autopilotSafetyGate !== false`) | Always Full, so a plan exists. Safe plans proceed silently; high-risk plans trigger a one-time manual approval. Log records the score. |
572
+ | Interactive (`/multi-agent`) | ✅ (max 2 rounds) | ✅ | - (redundant when a human approves) | Full gate |
573
+ | `/multi-agent:autopilot` | ❌ | ❌ conditional | ✅ (if `autopilotSafetyGate !== false`) | A plan always exists. Safe plans proceed silently; high-risk plans trigger a one-time manual approval. Log records the score. |
577
574
 
578
575
  **Why autopilot now has an escape hatch:** the old "zero interaction - fully trust the scope" contract held well for tightly-scoped batch workflows (figma component iteration over known-safe components) but broke in edge cases - schema migrations auto-merging, security-path drift going silent, delete-without-test sprawls. The safety classifier (Step 5c) is opt-out so the default protects against the edge-case cost; users running known-safe workflows can flip `prefs.global.autopilotSafetyGate = false` to restore pre-v7.0 behavior.
579
576
 
@@ -582,18 +579,10 @@ The pipeline shapes interact with the gate as follows. This table is the source
582
579
  After plan generation (and after each edit-loop iteration), forward the planning model's call totals so Phase 5's Cost Breakdown captures Phase 1 (`model=` names the rung that actually ran: `fable`, or `opus` after a fallback step):
583
580
 
584
581
  ```bash
585
- LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 2 plan.generated \
582
+ LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 1 plan.generated \
586
583
  model=fable tokens_in=$IN tokens_out=$OUT duration_ms=$DUR iteration=$N
587
584
  ```
588
585
 
589
586
  Best-effort. See `$HOME/.claude/multi-agent-refs/progress-contract.md#token-telemetry-forwarding`.
590
587
 
591
- ---
592
-
593
- ## Token telemetry - invoke after every LLM call
594
-
595
- ```bash
596
- bash $HOME/.claude/scripts/phase-tracker.sh tokens 2 <input_count> <output_count>
597
- ```
598
-
599
- Contract and rationale: `progress-contract.md` -> Token telemetry forwarding.
588
+ Both halves of this phase report into the same counter: see "Token telemetry" above.
@@ -1,6 +1,6 @@
1
1
  ### Phase 2: Dev (Sonnet)
2
2
 
3
- > **TLDR** - Sonnet executes the plan task-by-task with TDD (red→green→refactor). Required: issue-tracker status moved to "In Progress" before any code, with a post-mutation verify step (re-reads the field, retries once on silent VALIDATION failures). Build verification after each task (up to 3 retries). Build-queue lock serializes concurrent xcodebuild. A Short run (`state.onlyDevelop`) uses Opus self-contained, with no Phase 1 plan. `taskType === component` **short-circuits the TDD path** and delegates the whole phase to the enabled `ai-<platform>-toolkit` marketplace plugin's component skill (`create-component`, fallback `create-ui-component`) - see next subsection.
3
+ > **TLDR** - Sonnet executes the plan task-by-task with TDD (red→green→refactor). Required: issue-tracker status moved to "In Progress" before any code, with a post-mutation verify step (re-reads the field, retries once on silent VALIDATION failures). Build verification after each task (up to 3 retries). Build-queue lock serializes concurrent xcodebuild. `taskType === component` **short-circuits the TDD path** and delegates the whole phase to the enabled `ai-<platform>-toolkit` marketplace plugin's component skill (`create-component`, fallback `create-ui-component`) - see next subsection.
4
4
 
5
5
  ## Phase 2 Pre-flight (BLOCKING, v9.0.0)
6
6
 
@@ -8,11 +8,11 @@ Per Locked decision 30, Phase 2 Dev consumes the analysis document as the sole d
8
8
 
9
9
  Pre-flight steps (run in order, abort on failure).
10
10
 
11
- **Steps 1, 2, 3, 5 and 6 apply only when Phase 1 ran.** In a Short run (`state.onlyDevelop === true`) there is no analysis doc by design, so they are recorded `not-applicable (no Phase 1 in this mode)` and skipped - an unconditional abort there would make every fast mode impossible. Steps 4, 7, 8 and 9 apply in every mode.
11
+ **Steps 1, 2, 3, 5 and 6 read the analysis document.** Phase 1 always runs, so the document always exists; what varies is how much evidence it carries. A section the evidence did not support is absent by the Locked 2 omission rule, and a step whose section is absent is recorded `not-applicable (no <section> in this document)` rather than aborting. Steps 4, 7, 8 and 9 read nothing from it and apply always.
12
12
 
13
13
  1. **Analysis document presence** (Phase 1 modes only): read `state.analysis.docStatus` and `state.analysis.docPath[]`, both set by Phase 1 Step 4.
14
14
  - `produced` | `reused` -> read the active platform's file; multi-repo runs need one per selected repo.
15
- - `not-applicable` -> no document by design; record it for steps 1, 2, 3, 5, 6 and skip them, as a Short run does.
15
+ - `not-applicable` -> the analysis mode exported a document this run does not consume; record it for steps 1, 2, 3, 5, 6 and skip them.
16
16
  - **Abort**: `produced` but unreadable -> `ERR: analysis doc at <path> unreadable. Resume with /multi-agent:resume #N.` Producing it is Phase 1's job.
17
17
 
18
18
  2. **Parse YAML front-matter** into `state.analysis.frontMatter`: `feature`, `platform`, `language`, `mode`, `evidence_digest`, `template_version`. **Abort** when `template_version` < `v3`; there is no degraded mode.
@@ -49,7 +49,7 @@ The analysis document is the SOLE design source in Phase 2. Variant choices, pad
49
49
 
50
50
  #### Input contract
51
51
 
52
- Phase 2 consumes the Phase 1 output object conforming to `$HOME/.claude/schemas/planning-output.schema.json` - the task graph (`tasks[]` with `id`, `title`, `type`, `files`, and optional `dependsOn` / `acceptanceCriteria`) plus the architecture review notes. Tasks execute in dependency order; the schema's `dependsOn` field drives the ready-task picker. In a Short run (no Phase 1), Opus generates the equivalent task list inline before entering the loop below.
52
+ Phase 2 consumes the Phase 1 output object conforming to `$HOME/.claude/schemas/planning-output.schema.json` - the task graph (`tasks[]` with `id`, `title`, `type`, `files`, and optional `dependsOn` / `acceptanceCriteria`) plus the architecture review notes. Tasks execute in dependency order; the schema's `dependsOn` field drives the ready-task picker.
53
53
 
54
54
  **Plan Todo iteration (opt-in)**: gated by `prefs.global.planTodos.enabled` (default: `false`). When enabled and Phase 1 Step 11 emitted a `plan.todos[]`, Phase 2 iterates via `$HOME/.claude/lib/plan-todos.sh next/start/complete/fail` instead of walking `tasks[]` directly. When disabled, the loop walks `tasks[]` from `planning-output` - TDD contract is unchanged. Full helper loop + state semantics: `$HOME/.claude/multi-agent-refs/features/plan-todos.md`. A todo with `sourceTag: Reuse` binds the file analysis already found; `Modify` edits in place. Writing a new file over a `Reuse` step is a Locked 11 violation and Phase 3 flags it.
55
55
 
@@ -57,7 +57,7 @@ Phase 2 consumes the Phase 1 output object conforming to `$HOME/.claude/schemas/
57
57
 
58
58
  #### Component tasks - delegated dispatch (taskType === "component")
59
59
 
60
- When Phase 0 Step 7 classified the task as `component`, Phase 2 delegates the entire phase to the enabled `ai-<platform>-toolkit` marketplace plugin's component skill (`create-component`, fallback `create-ui-component`) via the Skill tool and does NOT run the TDD loop below. The dispatch layer passes the plugin skill the analysis Section 6 (Bileşen Envanteri) entry + Section 13.1 conventions for the named component as context. Because plugin skills do not write pipeline state, the **dispatch layer** (not the skill) owns `state.phases["2"].subphases[]`, recording a coarse component-build row - multi-agent's `phase-tracker` reads that array with no special case. Plugin resolution (dual-name), dispatch call, failure/resume, multi-repo, Short-run elisions, and the intentional cross-CLI divergence live in `$HOME/.claude/multi-agent-refs/component-dispatch.md` - read it before editing component-task behaviour here. Phase 2 still owns: progress line `-> dispatching create-component <name>`, `retryCount` cap at 3, and fallthrough to the TDD path when dispatch prerequisites are missing (`taskType` absent OR the plugin is not enabled in this repo -> log anomaly, halt or run TDD per component-dispatch.md).
60
+ When Phase 0 Step 7 classified the task as `component`, Phase 2 delegates the entire phase to the enabled `ai-<platform>-toolkit` marketplace plugin's component skill (`create-component`, fallback `create-ui-component`) via the Skill tool and does NOT run the TDD loop below. The dispatch layer passes the plugin skill the analysis Section 6 (Bileşen Envanteri) entry + Section 13.1 conventions for the named component as context. Because plugin skills do not write pipeline state, the **dispatch layer** (not the skill) owns `state.phases["2"].subphases[]`, recording a coarse component-build row - multi-agent's `phase-tracker` reads that array with no special case. Plugin resolution (dual-name), dispatch call, failure/resume, multi-repo, and the intentional cross-CLI divergence live in `$HOME/.claude/multi-agent-refs/component-dispatch.md` - read it before editing component-task behaviour here. Phase 2 still owns: progress line `-> dispatching create-component <name>`, `retryCount` cap at 3, and fallthrough to the TDD path when dispatch prerequisites are missing (`taskType` absent OR the plugin is not enabled in this repo -> log anomaly, halt or run TDD per component-dispatch.md).
61
61
 
62
62
  For non-component taskTypes (`bugfix`, `feature`, `refactor`, `chore`), continue with the standard TDD section below.
63
63
 
@@ -92,7 +92,7 @@ If the latest iteration has `triage.approved === true` AND `accepted === []`, Ph
92
92
  **Telemetry**: at the start of every re-entry, emit:
93
93
 
94
94
  ```bash
95
- $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 rework.started \
95
+ $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 2 rework.started \
96
96
  iteration=$ITERATION accepted_blocking=$BLOCKING accepted_important=$IMPORTANT
97
97
  ```
98
98
 
@@ -275,12 +275,12 @@ After the build/test green step and BEFORE Phase 3 handoff, run one diff-shrink
275
275
  5. **Record tokens in the cost ledger** so Phase 5's Cost Breakdown captures the pass:
276
276
 
277
277
  ```bash
278
- LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 dev.simplifier_pass \
278
+ LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 2 dev.simplifier_pass \
279
279
  model=sonnet tokens_in=$IN tokens_out=$OUT duration_ms=$DUR \
280
280
  edits_returned=$RET edits_applied=$APPLIED edits_skipped=$SKIPPED
281
281
  ```
282
282
 
283
- Scope guard: a single pass, never looped. Runs in a Short run too; component tasks (`taskType === "component"`) skip it (the figma skill owns its own checklist).
283
+ Scope guard: a single pass, never looped. Component tasks (`taskType === "component"`) skip it (the figma skill owns its own checklist).
284
284
 
285
285
  ---
286
286
 
@@ -297,37 +297,6 @@ Exit 1 lists `unjustified[]`: complete the record once and re-run. A second exit
297
297
 
298
298
  ---
299
299
 
300
- #### Short pipeline (`state.onlyDevelop === true`)
301
-
302
- Set by the Phase 0 Step 7.5 depth picker, or by autopilot never (autopilot always runs Full). When it is true, Phase 2 runs self-contained with **Opus** (not Sonnet). No Phase 1 plan exists - the agent creates its own scope.
303
-
304
- **Flow:**
305
- 1. Read task description (from Jira, GitHub issue, or free-text)
306
- 2. Lightweight file scan - grep/glob for relevant code (not full Explore agents)
307
- 3. Determine scope autonomously - no task breakdown, no user confirmation
308
- 4. Implement with TDD cycle (same RED→GREEN→REFACTOR as normal mode)
309
- 5. Build verification (same lock, same retry logic)
310
- 6. Intermediate commits (same WIP pattern)
311
-
312
- **Key differences from normal mode:**
313
-
314
- | Aspect | Full | Short |
315
- |--------|--------|----------|
316
- | Model | Sonnet | **Opus** |
317
- | Plan source | Phase 1 task list | Self-determined |
318
- | Task granularity | Per-plan-item | Agent decides |
319
- | Status updates | Per task item | Single in_progress → completed |
320
- | Scope confirmation | Phase 1 user approval | None (agent autonomous) |
321
- | Review of the result | Phase 3 | Phase 3 (same) |
322
-
323
- Because the agent determines its own scope here, Phase 3 is the only place that checks the result against anything external. Record every skill, plugin skill and guide consulted during this phase into `state.telemetry.skillCalls[]` with the files it was applied to - Phase 3 resolves the criteria set independently, and this record is what lets it tell "applied and honoured" from "never opened".
324
-
325
- **Never combined with autopilot.** Autopilot skips the depth question and runs Full, so `onlyDevelop` is false in every unattended run. "Fast plus unattended" was removed in v16.0.0 and no longer exists: something has to choose when nobody is asked, and unattended is the worst place to drop analysis and planning.
326
-
327
- **Tracker visibility during Opus dispatch**: on Claude Code the model switch to Opus happens via subagent dispatch, and the parent widget cannot move while an Agent call is in flight. Dispatch per task from the self-generated task list (never one monolithic call for the whole phase), set the pre-dispatch `activeForm` marker, and record tokens between chunks - full rules in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "Delegated phases".
328
-
329
- ---
330
-
331
300
  #### v2.1.0+ Multi-Repo Mode
332
301
 
333
302
  Active when `state.projects[].length > 1` (set by Phase 0 multi-select). Single-repo flow above is preserved verbatim - this section adds the deltas.
@@ -372,26 +341,26 @@ This closes the gap where an agent records "built" without ever producing build
372
341
 
373
342
  **Telemetry**: Per-repo metrics in addition to per-task metrics:
374
343
  ```bash
375
- $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 build.completed repo=common duration_ms=$D status=ok
376
- $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 build.completed repo=uicomponents duration_ms=$D status=ok
344
+ $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 2 build.completed repo=common duration_ms=$D status=ok
345
+ $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 2 build.completed repo=uicomponents duration_ms=$D status=ok
377
346
  ```
378
347
 
379
348
  **Token forwarding:** every TDD round (red, green, refactor) that hits the dev model MUST forward token totals into the tracker so Phase 5's Cost Breakdown captures Phase 2:
380
349
 
381
350
  ```bash
382
- LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 dev.tdd_round \
351
+ LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 2 dev.tdd_round \
383
352
  model=<sonnet|opus> step=<red|green|refactor> \
384
353
  tokens_in=$IN tokens_out=$OUT duration_ms=$DUR
385
354
  ```
386
355
 
387
- Model resolves from the active mode: a Full run uses Sonnet, a Short run uses Opus. Best-effort. See `$HOME/.claude/multi-agent-refs/progress-contract.md#token-telemetry-forwarding`.
356
+ Model is Sonnet unless routing names another rung (`multi-agent-refs/features/model-fallback.md`). Best-effort. See `$HOME/.claude/multi-agent-refs/progress-contract.md#token-telemetry-forwarding`.
388
357
 
389
358
  ---
390
359
 
391
360
  ## Token telemetry - invoke after every LLM call
392
361
 
393
362
  ```bash
394
- bash $HOME/.claude/scripts/phase-tracker.sh tokens 3 <input_count> <output_count>
363
+ bash $HOME/.claude/scripts/phase-tracker.sh tokens 2 <input_count> <output_count>
395
364
  ```
396
365
 
397
366
  Contract and rationale: `progress-contract.md` -> Token telemetry forwarding.