@mmerterden/multi-agent-pipeline 18.0.0 β†’ 19.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (234) hide show
  1. package/CHANGELOG.md +287 -0
  2. package/README.md +36 -20
  3. package/README.tr.md +14 -16
  4. package/docs/adr/0002-instruction-driven-flag.md +1 -0
  5. package/docs/adr/0005-lazy-phase-docs.md +11 -1
  6. package/docs/adr/0008-installer-modularization-and-secret-leak-defense.md +1 -0
  7. package/docs/adr/0010-own-code-graph.md +1 -0
  8. package/docs/adr/0014-six-phase-consolidation.md +134 -0
  9. package/docs/adr/README.md +2 -1
  10. package/docs/architecture.md +37 -38
  11. package/docs/best-practices.md +1 -1
  12. package/docs/ecosystem.md +46 -27
  13. package/docs/engineering.md +1 -1
  14. package/docs/facts.json +61 -0
  15. package/docs/features.md +55 -54
  16. package/docs/performance.md +5 -5
  17. package/docs/recovery-guide.md +17 -17
  18. package/docs/token-budget-history.md +3 -1
  19. package/index.js +2 -2
  20. package/install/_codex-agents.mjs +1 -1
  21. package/install/templates/claude-hooks.json +1 -1
  22. package/install/templates/codex-instructions.md +1 -1
  23. package/install/templates/copilot-instructions.md +28 -28
  24. package/manifest.json +234 -216
  25. package/package.json +2 -2
  26. package/pipeline/agents/dev-critic.md +7 -7
  27. package/pipeline/commands/figma-to-swiftui.md +1 -1
  28. package/pipeline/commands/multi-agent/SKILL.md +9 -9
  29. package/pipeline/commands/multi-agent/analysis/SKILL.md +15 -15
  30. package/pipeline/commands/multi-agent/analysis-resolve/SKILL.md +2 -2
  31. package/pipeline/commands/multi-agent/autopilot/SKILL.md +7 -7
  32. package/pipeline/commands/multi-agent/channels/SKILL.md +15 -15
  33. package/pipeline/commands/multi-agent/diff-explain/SKILL.md +6 -6
  34. package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +1 -1
  35. package/pipeline/commands/multi-agent/graph/SKILL.md +1 -1
  36. package/pipeline/commands/multi-agent/help/SKILL.md +62 -62
  37. package/pipeline/commands/multi-agent/language/SKILL.md +2 -2
  38. package/pipeline/commands/multi-agent/local/SKILL.md +11 -11
  39. package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +13 -13
  40. package/pipeline/commands/multi-agent/log/SKILL.md +2 -2
  41. package/pipeline/commands/multi-agent/manual-test/SKILL.md +9 -9
  42. package/pipeline/commands/multi-agent/model/SKILL.md +69 -0
  43. package/pipeline/commands/multi-agent/refactor/SKILL.md +3 -3
  44. package/pipeline/commands/multi-agent/resume/SKILL.md +4 -4
  45. package/pipeline/commands/multi-agent/resume-local/SKILL.md +19 -17
  46. package/pipeline/commands/multi-agent/review/SKILL.md +2 -2
  47. package/pipeline/commands/multi-agent/review-analysis/SKILL.md +1 -1
  48. package/pipeline/commands/multi-agent/route-off/SKILL.md +36 -0
  49. package/pipeline/commands/multi-agent/route-on/SKILL.md +74 -0
  50. package/pipeline/commands/multi-agent/route-status/SKILL.md +56 -0
  51. package/pipeline/commands/multi-agent/setup/SKILL.md +2 -2
  52. package/pipeline/commands/multi-agent/status/SKILL.md +5 -5
  53. package/pipeline/commands/multi-agent/steer/SKILL.md +2 -2
  54. package/pipeline/commands/multi-agent/sync/SKILL.md +12 -13
  55. package/pipeline/commands/multi-agent/test/SKILL.md +1 -1
  56. package/pipeline/lib/credential-inventory.sh +1 -1
  57. package/pipeline/lib/fetch-fortify.sh +1 -1
  58. package/pipeline/lib/model-dispatch.sh +140 -0
  59. package/pipeline/lib/model-rung.sh +142 -0
  60. package/pipeline/lib/outbound-gate.mjs +14 -0
  61. package/pipeline/lib/phase-schema.mjs +88 -0
  62. package/pipeline/lib/plan-todos.sh +5 -5
  63. package/pipeline/lib/route-state.sh +161 -0
  64. package/pipeline/lib/run-paths.sh +2 -2
  65. package/pipeline/multi-agent-refs/_account-picker.md +1 -1
  66. package/pipeline/multi-agent-refs/_dev-context.md +6 -6
  67. package/pipeline/multi-agent-refs/_input-parser.md +1 -1
  68. package/pipeline/multi-agent-refs/analysis/evidence.md +2 -11
  69. package/pipeline/multi-agent-refs/analysis/intake.md +7 -7
  70. package/pipeline/multi-agent-refs/analysis/locked.md +48 -22
  71. package/pipeline/multi-agent-refs/analysis/redesign.md +1 -1
  72. package/pipeline/multi-agent-refs/analysis/render.md +10 -10
  73. package/pipeline/multi-agent-refs/analysis/resolve.md +1 -1
  74. package/pipeline/multi-agent-refs/analysis/review.md +2 -2
  75. package/pipeline/multi-agent-refs/analysis/synthesis.md +13 -7
  76. package/pipeline/multi-agent-refs/analysis-template-corporate.md +9 -9
  77. package/pipeline/multi-agent-refs/analysis-template.md +19 -19
  78. package/pipeline/multi-agent-refs/android-guide.md +1 -1
  79. package/pipeline/multi-agent-refs/audit-guide.md +13 -13
  80. package/pipeline/multi-agent-refs/channels/issue-comment.md +2 -2
  81. package/pipeline/multi-agent-refs/channels/jira.md +3 -3
  82. package/pipeline/multi-agent-refs/channels/pr.md +4 -4
  83. package/pipeline/multi-agent-refs/channels/wiki.md +1 -1
  84. package/pipeline/multi-agent-refs/component-dispatch.md +8 -8
  85. package/pipeline/multi-agent-refs/conventions-defaults.md +2 -2
  86. package/pipeline/multi-agent-refs/cross-cli-contract.md +31 -6
  87. package/pipeline/multi-agent-refs/features/analysis-jira.md +1 -1
  88. package/pipeline/multi-agent-refs/features/autopilot-circuit-breaker.md +4 -4
  89. package/pipeline/multi-agent-refs/features/code-graph.md +5 -5
  90. package/pipeline/multi-agent-refs/features/design-conformance.md +1 -1
  91. package/pipeline/multi-agent-refs/features/dev-critic.md +3 -3
  92. package/pipeline/multi-agent-refs/features/doctor.md +3 -3
  93. package/pipeline/multi-agent-refs/features/external-context-injection.md +3 -3
  94. package/pipeline/multi-agent-refs/features/maturity-followup.md +3 -3
  95. package/pipeline/multi-agent-refs/features/model-fallback.md +41 -5
  96. package/pipeline/multi-agent-refs/features/plan-todos.md +1 -1
  97. package/pipeline/multi-agent-refs/features/repo-map.md +1 -1
  98. package/pipeline/multi-agent-refs/features/review-delta.md +3 -3
  99. package/pipeline/multi-agent-refs/features/review-multi-repo.md +2 -2
  100. package/pipeline/multi-agent-refs/features/scope-check.md +4 -4
  101. package/pipeline/multi-agent-refs/features/skill-conformance.md +2 -2
  102. package/pipeline/multi-agent-refs/features/stack-skill-routing.md +1 -1
  103. package/pipeline/multi-agent-refs/features/url-enrichment.md +1 -1
  104. package/pipeline/multi-agent-refs/features/verify-by-test.md +4 -4
  105. package/pipeline/multi-agent-refs/features/visual-evidence.md +19 -19
  106. package/pipeline/multi-agent-refs/features/worktree-finalize.md +6 -6
  107. package/pipeline/multi-agent-refs/issue-jira-triad.md +10 -10
  108. package/pipeline/multi-agent-refs/knowledge.md +11 -11
  109. package/pipeline/multi-agent-refs/multi-repo-integration-build.md +13 -13
  110. package/pipeline/multi-agent-refs/payload-contracts.md +8 -8
  111. package/pipeline/multi-agent-refs/phases/log-format.md +10 -10
  112. package/pipeline/multi-agent-refs/phases/modes.md +30 -30
  113. package/pipeline/multi-agent-refs/phases/operations.md +8 -8
  114. package/pipeline/multi-agent-refs/phases/phase-0-init.md +24 -24
  115. package/pipeline/multi-agent-refs/phases/phase-1-plan.md +599 -0
  116. package/pipeline/multi-agent-refs/phases/{phase-3-dev.md β†’ phase-2-dev.md} +129 -49
  117. package/pipeline/multi-agent-refs/phases/{phase-4-review.md β†’ phase-3-review.md} +225 -107
  118. package/pipeline/multi-agent-refs/phases/{phase-6-commit.md β†’ phase-4-commit.md} +23 -23
  119. package/pipeline/multi-agent-refs/phases/{phase-7-report.md β†’ phase-5-report.md} +29 -29
  120. package/pipeline/multi-agent-refs/phases.md +44 -48
  121. package/pipeline/multi-agent-refs/picker-contract.md +1 -1
  122. package/pipeline/multi-agent-refs/progress-contract.md +6 -6
  123. package/pipeline/multi-agent-refs/readiness-review.md +1 -1
  124. package/pipeline/multi-agent-refs/rules.md +7 -7
  125. package/pipeline/multi-agent-refs/swiftui-guide.md +2 -2
  126. package/pipeline/multi-agent-refs/tracker-contract.md +31 -32
  127. package/pipeline/multi-agent-refs/wiki-capture.md +14 -14
  128. package/pipeline/preferences-template.json +9 -1
  129. package/pipeline/rules/figma-pipeline.md +8 -8
  130. package/pipeline/rules/outside-the-pipeline.md +1 -1
  131. package/pipeline/schemas/agent-state.schema.json +50 -50
  132. package/pipeline/schemas/analysis-output.schema.json +3 -3
  133. package/pipeline/schemas/analysis-spec.schema.json +2 -2
  134. package/pipeline/schemas/autopilot-config.schema.json +1 -1
  135. package/pipeline/schemas/code-graph.schema.json +1 -1
  136. package/pipeline/schemas/criteria-manifest.schema.json +1 -1
  137. package/pipeline/schemas/dev-critic-output.schema.json +1 -1
  138. package/pipeline/schemas/diff-risk.schema.json +1 -1
  139. package/pipeline/schemas/figma-project-config.schema.json +1 -1
  140. package/pipeline/schemas/migrations/prefs-2.4.0-to-2.5.0.mjs +2 -2
  141. package/pipeline/schemas/migrations/prefs-2.6.0-to-2.7.0.mjs +31 -0
  142. package/pipeline/schemas/migrations/state-2.1.0-to-2.2.0.mjs +129 -0
  143. package/pipeline/schemas/phases.json +105 -0
  144. package/pipeline/schemas/plan-todos.schema.json +5 -5
  145. package/pipeline/schemas/planning-output.schema.json +1 -1
  146. package/pipeline/schemas/prefs.schema.json +102 -58
  147. package/pipeline/schemas/reviewer-output.schema.json +3 -3
  148. package/pipeline/schemas/route-config.schema.json +74 -0
  149. package/pipeline/schemas/scope-check.schema.json +1 -1
  150. package/pipeline/schemas/secret-patterns.json +124 -0
  151. package/pipeline/schemas/test-gap.schema.json +1 -1
  152. package/pipeline/schemas/token-budget.json +12 -18
  153. package/pipeline/schemas/triage-output.schema.json +6 -6
  154. package/pipeline/scripts/README.md +3 -3
  155. package/pipeline/scripts/_code-graph.mjs +2 -2
  156. package/pipeline/scripts/_run-paths.mjs +2 -2
  157. package/pipeline/scripts/_smoke-root.sh +1 -1
  158. package/pipeline/scripts/aggregate-metrics.mjs +1 -1
  159. package/pipeline/scripts/build-references.mjs +2 -2
  160. package/pipeline/scripts/bulk-read.sh +10 -1
  161. package/pipeline/scripts/capture-flush.sh +8 -8
  162. package/pipeline/scripts/capture-resume.sh +3 -3
  163. package/pipeline/scripts/classify-plan-safety.mjs +1 -1
  164. package/pipeline/scripts/cost-table.json +8 -1
  165. package/pipeline/scripts/diff-explain.mjs +1 -1
  166. package/pipeline/scripts/doctor.mjs +3 -3
  167. package/pipeline/scripts/gc-abandoned.sh +3 -3
  168. package/pipeline/scripts/gc-tmp.sh +1 -1
  169. package/pipeline/scripts/gc-worktrees.sh +1 -1
  170. package/pipeline/scripts/gen-facts.mjs +280 -0
  171. package/pipeline/scripts/gen-mode-dispatch.mjs +32 -37
  172. package/pipeline/scripts/gen-ref-toc.mjs +1 -1
  173. package/pipeline/scripts/graph-report.mjs +1 -1
  174. package/pipeline/scripts/jira-attach.sh +1 -1
  175. package/pipeline/scripts/learn-from-transcripts.mjs +1 -1
  176. package/pipeline/scripts/learning-curve.mjs +2 -2
  177. package/pipeline/scripts/log-metric.sh +17 -4
  178. package/pipeline/scripts/memory-save.sh +1 -1
  179. package/pipeline/scripts/migrate-prefs.mjs +22 -5
  180. package/pipeline/scripts/phase-banner.sh +20 -20
  181. package/pipeline/scripts/phase-tracker.sh +12 -12
  182. package/pipeline/scripts/plan-coverage-gate.mjs +2 -2
  183. package/pipeline/scripts/pre-commit-check.sh +30 -1
  184. package/pipeline/scripts/render-agent-log-cost.sh +1 -1
  185. package/pipeline/scripts/render-work-summary.sh +3 -3
  186. package/pipeline/scripts/review-file-filter.mjs +1 -1
  187. package/pipeline/scripts/run-aggregator.mjs +13 -6
  188. package/pipeline/scripts/run-metrics.mjs +1 -1
  189. package/pipeline/scripts/runs-index.mjs +11 -1
  190. package/pipeline/scripts/scan-skills.sh +26 -0
  191. package/pipeline/scripts/smoke-cross-cli-behavior.sh +6 -6
  192. package/pipeline/scripts/smoke-schema-validation.sh +26 -7
  193. package/pipeline/scripts/token-budget-report.mjs +13 -2
  194. package/pipeline/scripts/triage-memory.mjs +2 -2
  195. package/pipeline/scripts/validate-analysis-doc.mjs +274 -43
  196. package/pipeline/scripts/validate-planning.mjs +1 -1
  197. package/pipeline/scripts/validate-reviewer.mjs +1 -1
  198. package/pipeline/scripts/validate-state.mjs +45 -5
  199. package/pipeline/scripts/validate-triage.mjs +3 -3
  200. package/pipeline/scripts/verify-citations.mjs +1 -1
  201. package/pipeline/scripts/worktree-finalize.sh +5 -5
  202. package/pipeline/scripts/write-state.mjs +32 -0
  203. package/pipeline/skills/.skill-manifest.json +38 -22
  204. package/pipeline/skills/.skills-index.json +49 -5
  205. package/pipeline/skills/shared/README.md +10 -6
  206. package/pipeline/skills/shared/core/apple-archive-compliance/SKILL.md +8 -8
  207. package/pipeline/skills/shared/core/google-play-compliance/SKILL.md +8 -8
  208. package/pipeline/skills/shared/core/multi-agent/SKILL.md +81 -82
  209. package/pipeline/skills/shared/core/multi-agent-autopilot/SKILL.md +3 -3
  210. package/pipeline/skills/shared/core/multi-agent-channels/SKILL.md +14 -14
  211. package/pipeline/skills/shared/core/multi-agent-diff-explain/SKILL.md +5 -5
  212. package/pipeline/skills/shared/core/multi-agent-graph/SKILL.md +1 -1
  213. package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +25 -23
  214. package/pipeline/skills/shared/core/multi-agent-language/SKILL.md +2 -2
  215. package/pipeline/skills/shared/core/multi-agent-local/SKILL.md +2 -2
  216. package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +8 -8
  217. package/pipeline/skills/shared/core/multi-agent-manual-test/SKILL.md +6 -6
  218. package/pipeline/skills/shared/core/multi-agent-model/SKILL.md +71 -0
  219. package/pipeline/skills/shared/core/multi-agent-refactor/SKILL.md +3 -3
  220. package/pipeline/skills/shared/core/multi-agent-resume/SKILL.md +1 -1
  221. package/pipeline/skills/shared/core/multi-agent-resume-local/SKILL.md +7 -7
  222. package/pipeline/skills/shared/core/multi-agent-route-off/SKILL.md +39 -0
  223. package/pipeline/skills/shared/core/multi-agent-route-on/SKILL.md +76 -0
  224. package/pipeline/skills/shared/core/multi-agent-route-status/SKILL.md +59 -0
  225. package/pipeline/skills/shared/core/multi-agent-setup/SKILL.md +1 -1
  226. package/pipeline/skills/shared/core/multi-agent-status/SKILL.md +5 -5
  227. package/pipeline/skills/shared/core/multi-agent-steer/SKILL.md +2 -2
  228. package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +6 -5
  229. package/pipeline/skills/shared/external/NOTICE-swift-ios-skills.md +1 -1
  230. package/pipeline/skills/shared/external/signal-community/SKILL.md +8 -1
  231. package/pipeline/skills/skills-index.md +8 -4
  232. package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +0 -263
  233. package/pipeline/multi-agent-refs/phases/phase-2-planning.md +0 -344
  234. package/pipeline/multi-agent-refs/phases/phase-5-test.md +0 -182
@@ -1,263 +0,0 @@
1
- ### Phase 1: Analysis (Sonnet)
2
-
3
- > **TLDR** - Sonnet-driven codebase exploration: the `explorer` persona declares `preferredModel: sonnet` (haiku when the fallback ladder engages). Detects if the issue is already fixed (git blame, closed PRs), then launches parallel Explore sub-agents to map the affected code paths. Outputs: impact analysis, stack detection (auto-selects platform guide), relevant files, risk areas. Feeds Phase 2 planning.
4
-
5
- <!-- progress-contract: applied -->
6
- Progress emission per `$HOME/.claude/multi-agent-refs/progress-contract.md` - lines for each Explore dispatch, each finish, analyst synthesis start, `analysis.json` write.
7
-
8
- #### Step 0 - Prior Fix Detection
9
-
10
- Before any analysis, check if this issue was already fixed by someone else. Three signals (git commit grep on issue ID over 8 weeks, recent file-path history over 4 weeks, Jira/issue comments). On hit β†’ prompt: Verify (cherry-pick path) / Stop (cleanup + `stopped_already_fixed` state) / Continue. Miss β†’ Step 1. Full check commands, prompt block, and cleanup commands: `$HOME/.claude/multi-agent-refs/features/prior-fix-detection.md`.
11
-
12
- #### Step 1 - Knowledge Injection (cached context)
13
-
14
- Before launching Explore agents, check if project knowledge exists:
15
-
16
- ```
17
- $HOME/.claude/knowledge/{project-name}/
18
- architecture.md - file structure, module map, dependency graph
19
- patterns.md - patterns in use, conventions, idioms
20
- gotchas.md - encountered issues, edge cases and solutions
21
- decisions.md - architectural decisions and rationale (ADR-lite)
22
- ```
23
-
24
- Also read project-level CLAUDE.md if exists:
25
-
26
- - `$PROJECT_ROOT/CLAUDE.md`
27
- - `$PROJECT_ROOT/.claude/CLAUDE.md`
28
-
29
- **Per-repo memory injection (opt-in via `prefs.global.perRepoMemory`):**
30
-
31
- ```bash
32
- bash $HOME/.claude/scripts/memory-load.sh "$PROJECT_ROOT" "$TASK_TITLE $TASK_DESCRIPTION"
33
- ```
34
-
35
- Exit 0 with empty output = pref off or no memory on disk - skip. Otherwise the script emits a `<repo-memory path="...">...</repo-memory>` block of MEMORY.md pointers suitable for direct injection into the analysis prompt. Passing the task text ranks the pointers against it instead of printing the first thirty; individual memory files are read on-demand when a pointer looks relevant.
36
-
37
- **Durable knowledge (on by default via `prefs.global.learningsLedger.enabled`):** two blocks, and where each goes is part of the contract - see `$HOME/.claude/multi-agent-refs/prompt-assembly.md`.
38
-
39
- ```bash
40
- # HEAD of the prompt: task-independent, byte-stable, so it caches.
41
- node $HOME/.claude/scripts/learnings-ledger.mjs profile 2>/dev/null
42
- # END of the prompt, after the task text: ranked against this task.
43
- node $HOME/.claude/scripts/learnings-ledger.mjs brief --max "${prefs_learningsLedger_maxBriefEntries:-20}" \
44
- --task "$TASK_TITLE $TASK_DESCRIPTION" 2>/dev/null
45
- ```
46
-
47
- Exit 2 (empty ledger) = skip silently. Lines end with an `L:<id>` pointer; `show --id L:<id>` returns the full entry. Skip both when `injectIntoAnalysis = false`. Context, not commands - current scope decides. Then log the injection so recall quality stays measurable:
48
-
49
- ```bash
50
- bash $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 1 memory.injected \
51
- kind=<profile|task-relevant> rows=$N chars=$C
52
- ```
53
-
54
- **If knowledge files exist and are fresh** (modified within last 90 days - see knowledge.md staleness rules):
55
-
56
- 1. Read relevant knowledge files based on task description
57
- 2. Use knowledge to **narrow Explore scope** - instead of "very thorough" full scan, do targeted exploration of only unknown/changed areas
58
- 3. Log: "Phase 1: Knowledge injected ({N} files), targeted explore"
59
-
60
- **If no knowledge exists** (first time for this project):
61
-
62
- 1. Launch full Explore agents as before
63
- 2. Log: "Phase 1: Full explore (no cached knowledge)"
64
-
65
- #### Step 1.4 - Figma evidence capture (when task carries a Figma reference)
66
-
67
- When `state.contextLinks[]` or the task description contains a Figma reference, Phase 1 MUST collect the canonical evidence record. **Phase 0 Step 0.5 already resolved the tier** and the credential - read `state.figmaAccess.tier` and fetch with that tier's tool set; do not re-probe. Chain, tiers and halt conditions: `$HOME/.claude/multi-agent-refs/rules.md` "Figma Access Tier".
68
-
69
- What Phase 1 owns is the record. Every frame gets one entry in `state.evidence.figma[]`:
70
-
71
- | Field | Tier 1 (MCP) | Tier 2 (REST) | Tier 3 (screenshot) |
72
- |---|---|---|---|
73
- | `nodeId`, `screenshotUrl`, `tokens[]`, `textLayers[]` | required | required | required |
74
- | `codeConnectSnippets[]` | from `CodeConnectSnippet` blocks | from repo `*.figma.swift` / `*.figma.kt` keyed on `fileKey`+`nodeId`; empty -> Open Question | always `[]` -> forced Open Question |
75
- | `tier` | `1` | `2` | `3` |
76
-
77
- Halt if all three tiers fail; never substitute primitives or invent layout from prose.
78
-
79
- **Spacing goes in by token NAME, per atom - never a pixel number.** `tokens[]` must
80
- carry each frame's spacing/padding as Figma names them (`Spacing/12`, edge `4`), keyed
81
- to the atom. Phase 3 cannot call Figma, so what is missed here is gone: one run guessed
82
- `16` where the frame said `Spacing/12` and the sheet was rebuilt. A pixel number also
83
- cannot map back to a token. No spacing entries on a UI frame is a **capture failure**,
84
- not an empty frame - Open Question and halt. Canonical chain reference: `$HOME/.claude/rules/figma-pipeline.md` "MUST: Figma access - 3-tier fallback chain".
85
-
86
- **Telemetry (required for the no-MCP gate):** Tier 1 uses `mcp__claude_ai_Figma__*` tools. Every such MCP invocation MUST append an entry to `state.telemetry.mcpCalls[]` as `{ "tool": "<full mcp tool name>", "phase": 1, "timestamp": "<ISO-8601>" }`. This is the only phase permitted to record `phase: 1` (or `0`) entries; `smoke-no-mcp-in-dev-phases.sh` fails the run if any entry carries `phase >= 2`. Recording is what makes that BLOCKING contract enforceable - an MCP call left unrecorded defeats the gate, so record every one.
87
-
88
- Progress lines:
89
-
90
- ```
91
- β†’ figma evidence: <N> frames captured (tier=<n>, code-connect=<M>, open-questions=<K>)
92
- ```
93
-
94
- #### Step 1.45 - Reuse discovery (BLOCKING for new services, entities, mappers)
95
-
96
- Before proposing any new service call, entity or mapper, search for what already
97
- covers it. Record hits under `state.reuse[]` and cite them in the doc; proposing new
98
- code over a hit needs a one-line reason.
99
-
100
- Search for: a **wrapper** over the same endpoint (especially one supplying parameters
101
- the generated call leaves optional); an **entity** for the same concept (module's
102
- shared entities first, then siblings); a **mapper** over the same response; a **screen**
103
- doing the same interaction.
104
-
105
- Why blocking: one run proposed a new repository over an endpoint a sibling already
106
- wrapped **with its country parameter**, called the generated method without it, and
107
- re-invented an entity the module had. Half that branch's commits went to converging
108
- back. "Copy X and rename it" is the reuse answer, not a hint - name X's files.
109
-
110
- #### Step 1.5 - External Context Injection (`state.contextLinks[]`)
111
-
112
- Phase 0 Step 1b catalogued every typed external link from the task description into `state.contextLinks[]`. Phase 1 dispatches each entry to its matching fetcher (crashlytics, fortify, graylog, swagger, confluence, figma, generic-doc) and prepends results under a **Referenced External Sources** section in the analysis prompt - so the agent doesn't re-discover what the ticket already pointed at. `state.graylogContext` (advisory) and `state.relatedIssues[]` (sibling issues) are injected there too. Failures never fatal (a non-zero fetcher exit is marked skipped and the analysis still runs, exactly as for crashlytics); pending refs are advisories. Full dispatch table, exit-code handling, prompt injection shape, log line shape: `$HOME/.claude/multi-agent-refs/features/external-context-injection.md`.
113
-
114
- **Log line shape** (progress contract):
115
-
116
- ```
117
- β†’ context injection: total=<N>, fetched=<n>, pending=<n>, by-type={swagger:F/P, confluence:F/P, ...}
118
- ```
119
-
120
- #### Step 2 - Stack Detection
121
-
122
- Two questions shared one answer here, and the shared answer was wrong on both.
123
-
124
- **What the repo is BUILT WITH** - one owner, no marker list in this file:
125
-
126
- ```bash
127
- eval "$(bash $HOME/.claude/lib/stack-detect.sh "$PROJECT_ROOT")"
128
- ```
129
-
130
- Persist `MA_STACKS` as `state.stacks[]` (a subset of `ios android web backend`,
131
- always in that order) and `MA_STACK_WHY` as `state.stackWhy`. Empty is a real
132
- answer meaning no marker matched; `stackWhy` separates that from a directory that
133
- could not be read, and a caller that cannot tell those apart treats an unreadable
134
- repo as a language-free one. `state.stacks[]` routes toolkits, through
135
- `pluginsForStacks()` in `$HOME/.claude/scripts/_stack-routing.mjs`. Do not derive
136
- plugin names here.
137
-
138
- The table that used to live here scanned at depth 2, and `AndroidManifest.xml`
139
- sits at `<module>/src/main/` - depth 4 in every multi-module app - so **Android
140
- was never detected**. The owner goes to depth 5 for that one file, prunes
141
- submodules (a vendored checkout ships its own `Package.swift`, which reported a
142
- Compose app as iOS) and checks the root first so traversal order cannot decide.
143
-
144
- **What the code is WRITTEN IN** - a different axis with its own field.
145
- `state.detectedStack[]` stays the language answer (`ios`, `python`, `node`, `go`,
146
- `docker`, `monorepo`): `graph-build.mjs --stack` takes `ios|android|node|python|go`
147
- and has no notion of `web` or `backend`. Folding the two together is what made a
148
- Gradle-built JVM service read as an Android app.
149
-
150
- This informs:
151
-
152
- - Phase 3: which build/test commands to use
153
- - Phase 4: which deterministic gates and reviewer skills to load
154
- - Phase 6: which PR template fits best
155
-
156
- #### Step 2.5 - Repo Map Injection (advisory, opt-in)
157
-
158
- Gated by `prefs.global.repoMap.enabled` (default: `false`). When enabled, runs `$HOME/.claude/scripts/repo-map.mjs` and injects the budgeted result into each Explore prompt as `${REPO_MAP}`. Aider-style: deterministic, no embeddings, sub-second, advisory only. Full wiring (helper invocation, properties, when-to-enable): `$HOME/.claude/multi-agent-refs/features/repo-map.md`.
159
-
160
- #### Step 2.6 - Code Graph Injection (advisory, opt-in)
161
-
162
- Gated by `prefs.global.codeGraph.enabled` (default: `false`). With a rule file for `detectedStack`, Phase 1 refreshes the graph and queries it; `graph-affected` feeds `analysis.touchedAreas[]`. Zero API cost, read-only. Commands and measurements: `$HOME/.claude/multi-agent-refs/features/code-graph.md`.
163
-
164
- **A valid result REPLACES the opening sweep** rather than sitting beside it: Explore starts from those files and walks outward, with no broad `Glob`/`Grep` pass first - running both pays twice, and the second re-derives what the graph said. No graph (missing rule, stale, or disabled) leaves the previous behaviour untouched; a task naming a symbol or path is still grepped directly.
165
-
166
- #### Step 3 - Codebase Exploration
167
-
168
- Launch parallel Explore agents to scan codebase:
169
-
170
- - Related files to the task
171
- - Existing patterns and conventions
172
- - Potential impact areas
173
-
174
- Use `subagent_type: "Explore"` with thoroughness scaled to task size AND knowledge availability (first match wins):
175
-
176
- - A fresh code graph answered the task's query with a ranked file set (Step 2.6) β†’ "light" (the starting set is already narrowed; explore outward from it, do not re-scan)
177
- - `taskType` is `bugfix`/`chore` AND scope is small (single named file, or a referenced crash/stack frame that pinpoints the site) β†’ "light" (cheapest - scan only the named area + its direct callers)
178
- - Knowledge exists β†’ "medium" (targeted, cheaper)
179
- - No knowledge, or `taskType` is `feature`/`refactor`/`component` β†’ "very thorough" (full scan, first-time investment)
180
-
181
- The light tier keeps a one-line bug fix from triggering a full-repo scan; pairing it with the deterministic `taskType` (Phase 0 Step 7) prevents the cheap path from firing on feature work.
182
-
183
- **Dispatch resilience (required).** Explore agents run in parallel and the analyst synthesis waits on them, so a single stalled agent hangs the phase. Bound each Explore dispatch by a wall-clock budget (`EXPLORE_TIMEOUT_SECONDS`, default 180). If an agent has not returned by the budget: log `explore.timeout agent=<id>`, drop that agent's slice, and synthesize from the agents that did return. Proceed as long as at least one Explore agent returned; if zero returned, retry the cheapest single Explore once, then HALT with `ERR: no Explore agent returned within ${EXPLORE_TIMEOUT_SECONDS}s; resume with /multi-agent:resume #N.`. Never block indefinitely on a slow or dead dispatch.
184
-
185
- #### Step 4 - Analysis document (the design contract Phase 2 and Phase 3 demand)
186
-
187
- Phase 2 and Phase 3 pre-flights BLOCK on `analysis/<feature-slug>-<platform>.md`. **This step produces it.**
188
-
189
- **When it runs.** From signals Phase 0 already computed:
190
-
191
- | `taskType` | Figma reference in `state.contextLinks[]` | Document |
192
- |---|---|---|
193
- | `feature` Β· `refactor` Β· `component` | any | **produced** |
194
- | `bugfix` Β· `chore` | present | **produced** |
195
- | `bugfix` Β· `chore` | absent | **skipped** - record `state.analysis.docStatus = "not-applicable"` |
196
-
197
- `analysisPhase.forceFull` (default `false`) overrides the skip row, for a small change that must still leave a spec behind. `analysisPhase.mode` sets depth: `auto` is Lite on the skip-row shape and Full otherwise; `full` / `lite` pin it.
198
-
199
- A fresh document is also skipped when one already exists for this feature and platform AND its front-matter `evidence_digest` still matches (Locked 27 cache); record `docStatus = "reused"`.
200
-
201
- **How it runs.** Load `$HOME/.claude/multi-agent-refs/analysis/` on demand, in order: `locked.md` (the 31 binding decisions) then `evidence.md`, `synthesis.md`, `render.md`. Intake is NOT re-asked; platform, repos and account come from Phase 0 state. Autopilot auto-approves the Phase 2a convention preview and writes the local file, because it may not ask.
202
-
203
- **Where it lands.** `<worktree>/analysis/<feature-slug>-<platform>.md`, one file per platform in `state.analysisSpec.platforms[]`. Persist the paths to `state.analysis.docPath[]` and set `docStatus` to `produced` | `reused` | `not-applicable`. Also persist `state.run.lastAnalysisDigest` (this document's `evidence_digest`) and `state.run.analysisBaseCommit` (`git rev-parse HEAD`). Phase 3's freshness check reads both; unwritten, it has nothing to compare and passes silently. Whether the file is committed with the work is `prefs.global.analysisPhase.commitDoc` (default `true`), read at Phase 6.
204
-
205
- **The gate is not optional.** `node $HOME/.claude/scripts/validate-analysis-doc.mjs <file>` must exit 0 for every produced file; non-zero fails CLOSED like the JSON validator below (rework once, then halt).
206
-
207
- Progress line: `β†’ analysis doc: <produced|reused|not-applicable> (<N> platform, validator <pass|pass-after-rework>)`
208
-
209
- #### Output contract
210
-
211
- Two artefacts, both read downstream: `state.analysis` (the object below, for Phase 2 decomposition) and the Step 4 document (Phase 2 pre-flight + Phase 3's sole design source, Locked 30).
212
-
213
- Phase 1 produces an object conforming to `$HOME/.claude/schemas/analysis-output.schema.json` and persists it to `state.analysis`. Required fields (exact names per the schema): `stack` (detected stack identifier + primary language), `touchedAreas[]` (path + why), `risks[]` (existing-code hazards/observations the planner must respect - each `{risk, severity, mitigation}`; use an empty array when none), `summary` (one-paragraph human-readable). Phase 2 reads this object as its sole input - see `phase-2-planning.md`'s Input contract.
214
-
215
- **Required: validator gate (deterministic) - run on the persisted file immediately after the analysis object is produced; the validator's exit code decides, not the LLM turn:**
216
-
217
- ```bash
218
- ANALYSIS_FILE="$WORKTREE/.pipeline/analysis.json"
219
- mkdir -p "$(dirname "$ANALYSIS_FILE")"
220
- printf '%s' "$ANALYSIS_JSON" > "$ANALYSIS_FILE"
221
- node $HOME/.claude/scripts/validate-analysis.mjs "$ANALYSIS_FILE"
222
- ```
223
-
224
- Progress line: ` β†’ checking validator validate-analysis`
225
-
226
- Non-zero exit fails CLOSED: emit the validator stderr + `errors[]` verbatim, attempt ONE self-correction rework (re-invoke the explorer with the errors quoted, overwrite `$ANALYSIS_FILE`), re-run the validator. If it fails again -> HALT the phase with recovery hint: `ERR: analysis output failed validate-analysis.mjs twice. Inspect $ANALYSIS_FILE against $HOME/.claude/schemas/analysis-output.schema.json, then resume with /multi-agent:resume #N.` Record `agent-state.phases["1"].validator` (`pass` | `pass-after-rework` | `halted`).
227
-
228
- Log: "Phase 1: Analysis - stack:{stacks} lang:{detectedStack} | {N} files identified, {summary}"
229
-
230
- #### Telemetry - token forwarding
231
-
232
- Forward the explorer call's token totals into the tracker so Phase 7's Cost Breakdown captures Phase 1:
233
-
234
- ```bash
235
- LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 1 analysis.completed \
236
- model=sonnet tokens_in=$IN tokens_out=$OUT duration_ms=$DUR
237
- ```
238
-
239
- Best-effort. See `$HOME/.claude/multi-agent-refs/progress-contract.md#token-telemetry-forwarding` for the canonical contract.
240
-
241
- #### Prior-Art Enrichment (advisory)
242
-
243
- After the explorer returns its summary, consult the per-repo triage corpus for similar past tasks. Inject up to 3 matches into the analysis output as `priorArt[]` so Phase 2 planning can read them. Disabled when `prefs.global.priorArtEnrichment.enabled = false`.
244
-
245
- ```bash
246
- PRIOR=$(node $HOME/.claude/scripts/triage-memory.mjs query \
247
- --issue "$TASK_TITLE $TASK_DESCRIPTION" --top 3 2>/dev/null | jq -c '.hits // []')
248
- ```
249
-
250
- Hits are relevance-ranked, and a query matching nothing returns nothing. Each hit carries an `id`: `triage-memory.mjs show --id <id>` returns the full row.
251
-
252
- Treat hits as **context only** - they are past Phase 4 verdicts, not prescriptions. Useful when the new task touches the same files or symbols as a previous run.
253
-
254
- ---
255
-
256
- ## Token telemetry - invoke after every LLM call
257
-
258
- ```bash
259
- bash $HOME/.claude/scripts/phase-tracker.sh tokens 1 <input_count> <output_count>
260
- ```
261
-
262
- Contract and rationale: `progress-contract.md` -> Token telemetry forwarding.
263
-
@@ -1,344 +0,0 @@
1
- ### Phase 2: Planning (Fable)
2
-
3
- > **TLDR** - Fable decomposes the analysis (Opus when the fallback ladder engages) into concrete tasks with file-level targets, risk grading, and architecture review. Before Phase 3 a **Plan Approval Gate** runs in normal mode: if the Jira/issue description is ambiguous the orchestrator asks the user structured clarification questions (max 2 rounds) - once scope is clear it renders the plan and loops on free-text edit requests until the user approves or aborts. The gate is **skipped entirely** for a Short run and for `autopilot` (a Short run has no plan to approve; autopilot may not ask).
4
-
5
- <!-- progress-contract: applied -->
6
- Progress emission per `$HOME/.claude/multi-agent-refs/progress-contract.md` - lines for plan-draft start, clarification-ask, clarification-answer, plan render, plan-edit-request, plan-approved, plan-aborted.
7
-
8
- ## Phase 2 Pre-flight (BLOCKING, v9.0.0)
9
-
10
- Phase 2 Planning consumes the analysis document. **Figma** MCP / REST forbidden; the toolkit MCP is not.
11
-
12
- 1. **Analysis document presence**: read `state.analysis.docStatus`, which Phase 1 Step 4 set.
13
- - `produced` | `reused` -> the file is at `state.analysis.docPath[]`; continue with steps 2-4.
14
- - `not-applicable` -> no document by design (bugfix/chore, no Figma). Record it, skip steps 2-4, plan from `state.analysis` alone.
15
- - Unreadable or key missing -> `ERR: Phase 1 reported <status> but no analysis doc is readable. Resume with /multi-agent:resume #N.` Producing it is Phase 1's job; never send the user to another command.
16
-
17
- 2. **Parse YAML front-matter** into `state.analysis.frontMatter` and abort below `template_version: v3` - same contract as `phase-3-dev.md` step 2.
18
-
19
- 3. **Section coverage check**: verify these sections are non-empty (template v3 required sections per Locked 2):
20
- - Section 1 Summary, Section 2 Goals + Non-Goals, Section 4 User Stories, Section 9 API Contracts, Section 13 Architecture Plan, Section 14 Files to Add, Section 20 Risks, Section 21 References
21
- - Missing -> WARN, plan is allowed to proceed but Phase 4 reviewer flags it.
22
-
23
- 4. **Convert analysis tasks to plan**: Section 14 Files-to-Add becomes the seed task list. Carry each row's tag onto its todo as `sourceTag` (`Reuse` | `Add new` | `Modify`); it is an instruction Phase 3 follows and Phase 4 checks, not a label.
24
-
25
- 5. **MCP forbidden**: same rule as Phase 3.
26
-
27
- #### Input contract
28
-
29
- Phase 2 consumes the Phase 1 output object conforming to `$HOME/.claude/schemas/analysis-output.schema.json`. Read `state.analysis` (the explorer's return value) and treat its `touchedAreas` and `risks` arrays (plus `stack` and `summary`) as authoritative input - do not re-explore the codebase here. (Field names are exactly those in the analysis schema; the planner's own `targetFiles` belongs to `planning-output.schema.json`, not the analysis input.)
30
-
31
- #### Step 0.9 - Post-analysis confirmation (derive first, ask only what cannot be derived)
32
-
33
- Runs when `state.analysis.docStatus` is `produced` or `reused`, before any planning.
34
-
35
- **Derived is shown, not asked**: platform set, the seven convention groups, existing components (Code Connect + `uiComponents`), localization keys, analytics events, DI registration, test-method naming. **Only Section 20 rows are asked**, through `$HOME/.claude/multi-agent-refs/analysis/resolve.md` - one row, at most three source-labeled candidates, plus Defer. The engine never invents.
36
-
37
- ```
38
- TΓΌretildi (onay iΓ§in): platform=ios,android Β· 7 konvansiyon grubu (5 high, 2 medium)
39
- · 12 mevcut bileşen (9 Code Connect bağlı) · 34 lokalizasyon anahtarı
40
- Sorulacak: 4 aΓ§Δ±k soru (BΓΆlΓΌm 20)
41
- ```
42
-
43
- A corrected value rewrites its Pass B footnote as `^[user-override: resolved <date>]` (Locked 24). Deferred rows stay in `state.analysis.openQuestions[]` and Phase 4 flags them `review_blocking`. Here and not Phase 4 because Phase 4 runs after development - answered there is answered too late. Autopilot skips the asking and defers every row.
44
-
45
- #### Step 1 - Task Decomposition
46
-
47
- Break the work from Phase 1 analysis into discrete, implementable tasks:
48
-
49
- ```
50
- For each identified change area:
51
- 1. Define a task with clear scope (one file group or one logical change)
52
- 2. Estimate complexity: trivial (1 file) / moderate (2-5 files) / complex (6+ files)
53
- 3. Identify dependencies between tasks (which must complete before others)
54
- ```
55
-
56
- Create tasks using TaskCreate with imperative subject, description (what/which files/expected behavior), and `addBlockedBy` for dependencies.
57
-
58
- ##### Analysis citation requirement (every UI task)
59
-
60
- Every UI-touching task in the plan MUST cite:
61
-
62
- - The analysis Section 6 (Bileşen Envanteri) row it implements (one task per row when a single component covers several variants), AND
63
- - The canonical component name, sourced from the analysis doc:
64
- - **Code Connect mapping present**: take the component name verbatim from the matching repo `*.figma.swift` / `*.figma.kt` row referenced by analysis Section 6.
65
- - **Mapping absent**: cite "best-fit pending design review" and add a Risk row to the plan. The task description MUST also flag that Phase 4 will gate it as `review_blocking`.
66
-
67
- Tasks without an analysis Section 6 citation when the doc lists UI components are rejected at the plan-approval gate; the user is asked to either re-run `/multi-agent:analysis` to extend Section 6 or rescope the task to a non-UI change. Direct Figma fetches in Phase 2 are forbidden (Locked decision 30).
68
-
69
- Example task graph:
70
-
71
- ```
72
- Task 1: Add new token to common module (no deps)
73
- Task 2: Create ButtonConfiguration.swift (blocked by 1)
74
- Task 3: Create ButtonView.swift (blocked by 2)
75
- Task 4: Add ButtonView+Modifiers.swift (blocked by 3)
76
- Task 5: Write ViewInspector tests (blocked by 3)
77
- Task 6: Write snapshot tests (blocked by 3)
78
- ```
79
-
80
- #### Step 2 - Architecture Review (conditional)
81
-
82
- Trigger architecture review if ANY of these are true:
83
-
84
- - New module or package being created
85
- - Cross-module dependency being added
86
- - Public API surface changing
87
- - Data model / schema change
88
- - Navigation flow change
89
-
90
- If triggered:
91
-
92
- 1. Launch Agent with `subagent_type: "ios-architect"` (or `architecture` skill for non-iOS)
93
- 2. Provide: task list, affected files, proposed approach
94
- 3. Agent returns: recommendation, risks, alternative approaches
95
- 4. Incorporate recommendations into task descriptions
96
-
97
- If NOT triggered: skip, log "Phase 2: Architecture review - not needed (scope contained)"
98
-
99
- #### Step 3 - Development Approach per Task
100
-
101
- For each task, determine:
102
- | Approach | When | Example |
103
- |----------|------|---------|
104
- | **New file** | Feature doesn't exist | Create ButtonConfiguration.swift |
105
- | **Modify** | Extending existing code | Add property to existing Configuration |
106
- | **Refactor** | Restructuring without behavior change | Extract protocol from class |
107
- | **Fix** | Bug correction | Fix nil crash in edge case |
108
-
109
- Store approach in task metadata for Phase 3 agent.
110
-
111
- #### Step 4 - Skill Selection per Task
112
-
113
- Based on Phase 1 `detectedStack`, assign relevant skills:
114
-
115
- - iOS tasks -> `ai-ios-toolkit:*` SwiftUI skills, iOS patterns
116
- - Python tasks -> `ai-backend-toolkit:fastapi-pro`, `ai-backend-toolkit:api-patterns`
117
- - Node tasks -> `ai-backend-toolkit:nodejs-backend-patterns`
118
- - Security-sensitive -> `ai-backend-toolkit:api-security-best-practices`
119
- - Multi-submodule -> `ai-backend-toolkit:monorepo-architect`
120
-
121
- #### Output contract
122
-
123
- Phase 2 produces an object conforming to `$HOME/.claude/schemas/planning-output.schema.json` - `tasks[]` with `id`, `title`, `type`, `files`, plus optional `dependsOn` and `acceptanceCriteria`. Phase 3 reads `tasks[]` in dependency order; the schema's `dependsOn` field drives the ready-task picker.
124
-
125
- **Required: validator gate (deterministic) - run on the persisted file before the approval gate renders the plan; the validator's exit code decides, not the LLM turn:**
126
-
127
- ```bash
128
- PLAN_FILE="$WORKTREE/.pipeline/plan.json"
129
- mkdir -p "$(dirname "$PLAN_FILE")"
130
- printf '%s' "$PLAN_JSON" > "$PLAN_FILE"
131
- node $HOME/.claude/scripts/validate-planning.mjs "$PLAN_FILE"
132
- ```
133
-
134
- Progress line: ` β†’ checking validator validate-planning`
135
-
136
- Non-zero exit fails CLOSED: emit the validator stderr + `errors[]` verbatim, attempt ONE self-correction rework (re-invoke the planner with the errors quoted, overwrite `$PLAN_FILE`), re-run the validator. If it fails again -> HALT the phase (never enter Phase 3 with an invalid plan) with recovery hint: `ERR: plan output failed validate-planning.mjs twice. Inspect $PLAN_FILE against $HOME/.claude/schemas/planning-output.schema.json, then resume with /multi-agent:resume #N.` Record `agent-state.phases["2"].validator` (`pass` | `pass-after-rework` | `halted`).
137
-
138
- Log: "Phase 2: Plan - {N} tasks created, {M} with architecture review, validator:pass"
139
-
140
- #### Step 4.45 - Put the plan on the widget (required)
141
-
142
- ```bash
143
- printf '%s' "$PLAN_JSON" | bash "$HOME/.claude/scripts/phase-tracker.sh" plan 3
144
- ```
145
-
146
- Then do what its output asks - it prints a tile rebuild, because the widget
147
- orders by creation. `dependsOn` becomes `addBlockedBy`, so the widget answers
148
- "why has t3 not started". Required, unlike Step 4.5: the plan was always computed
149
- and stored, only invisible.
150
-
151
- #### Step 4.5 - Emit Plan Todo List (opt-in)
152
-
153
- **Gated by `prefs.global.planTodos.enabled`** (default: `false`). When enabled, after the planning-output JSON validates and BEFORE the approval gate, transform `tasks[]` into a structured Todo list conforming to `$HOME/.claude/schemas/plan-todos.schema.json` and persist into `agent-state.plan`. The plan is rendered as a live, always-visible Todo list.
154
-
155
- ```bash
156
- printf '%s' "$PLAN_JSON" | bash "$HOME/.claude/lib/plan-todos.sh" set "$TASK_ID" -
157
- ```
158
-
159
- `set` accepts a planning-output document and converts it itself. The conversion
160
- used to be written out here, which made the `tasks[]`-to-`todos[]` mapping a
161
- thing two files defined, and this was the copy nothing tested.
162
-
163
- Phase 3 (Dev) then iterates with `plan-todos.sh next "$TASK_ID"` until empty, calling `start` before each step and `complete` (with notes) or `fail`/`skip` after. Phase 4 (Review) reads the Todo list to verify all `completed` items map to diff hunks. Phase 7 (Report) renders `list` into the agent-log + PR body.
164
-
165
- **Why opt-in:** the existing `planning-output.schema.json` already drives Phase 3 dependency order - `plan.todos[]` is a richer surface (notes, durations, status transitions) but adds state writes per step. Off by default to keep the bare-bones flow unchanged; flip on for visibility into long features.
166
-
167
- #### Step 4.8 - Cross-artifact consistency check (required, before presenting the plan)
168
-
169
- Before the plan is rendered for approval (Step 5b) or silently accepted (the autopilot skip path), verify it against `state.analysis` (drifted plans are the root cause of "PR does not match the ticket"):
170
-
171
- 1. **Requirement coverage** - every analysis requirement (`touchedAreas[]` entry, Section 14 row, acceptance criterion) maps to at least one plan task.
172
- 2. **Anchor integrity** - No plan task without an analysis anchor (each task cites the `touchedAreas[].path`, Section 6 row, or `risks[]` mitigation it implements).
173
- 3. **Open-question carry-over** - every analysis open question lands in the plan (clarification item, risk row, or explicit descope note); none silently dropped.
174
-
175
- Progress line: ` β†’ checking plan-vs-analysis consistency (3 checks)`
176
-
177
- **On mismatch:** revise the plan ONCE (re-run Steps 1-4 with the gap list quoted, re-run the validator gate and this checklist). If gaps remain, do NOT loop: surface them in the Step 5b render under a `⚠️ Consistency gaps` banner (autopilot: log `plan.consistency_gaps={list}`). Persist `state.phases["2"].consistencyCheck = { "unmappedRequirements": [], "unanchoredTasks": [], "droppedOpenQuestions": [], "revised": true|false }`.
178
-
179
- Log: "Phase 2: Consistency - requirements:{N/N mapped} anchors:{ok|M unanchored} open-questions:{carried|K dropped}"
180
-
181
- #### Step 5 - Plan Approval Gate (normal mode + autopilot safety)
182
-
183
- **Scope guard - skip this step entirely when BOTH of these hold:**
184
-
185
- - `state.autopilot === true` (autopilot contract: zero interaction)
186
- - Autopilot safety classifier returns `recommendPause: false` (see Step 5c below)
187
-
188
- OR:
189
-
190
- - `state.onlyDevelop === true` (Short pipeline: direct to Phase 3, no plan)
191
-
192
- In the skipped case, log `🧠 Phase 2: Plan - gate skipped ({mode}), proceeding to Phase 3` and go to Phase 3.
193
-
194
- ##### 5c - Autopilot safety classifier (runs before 5a/5b skip decision)
195
-
196
- **Only relevant when `state.autopilot === true` and `prefs.global.autopilotSafetyGate !== false`.** The classifier protects autopilot's zero-interaction contract from edge cases where silent execution is genuinely dangerous (security-path touch, schema migration, many-file sprawl, delete-without-paired-test).
197
-
198
- ```bash
199
- verdict=$(node $HOME/.claude/scripts/classify-plan-safety.mjs <(echo "$PLAN_JSON"))
200
- pause=$(jq -r '.recommendPause' <<< "$verdict")
201
- score=$(jq -r '.score' <<< "$verdict")
202
- reasons=$(jq -r '.reasons[] | "β€’ \(.rule) (+\(.weight)): \(.detail)"' <<< "$verdict")
203
- ```
204
-
205
- | `recommendPause` | Action |
206
- |---|---|
207
- | `false` | Skip 5a/5b as before - autopilot proceeds to Phase 3. Log `🧠 Phase 2: Safety classifier - score {N}, autopilot proceeds`. |
208
- | `true` | Inject a one-time manual approval prompt even though we are in autopilot. Render the plan (5b shape) with the `reasons[]` list prepended. User sees `⚠️ Autopilot safety gate tripped (score {N})` banner and must explicitly choose Approve / Cancel (or edit via Other) in the 5b `AskUserQuestion` picker. Log `🧠 Phase 2: Safety classifier - score {N}, autopilot paused, reasons={rules}`. |
209
-
210
- **Why opt-out instead of opt-in:** the asymmetry favors pausing. A pause on a high-blast-radius plan costs seconds; a silent auto-merge of a bad one costs hours of rollback or a revert PR. Users running tightly-scoped batch workflows (e.g. figma component iteration over known-safe components) can set `prefs.global.autopilotSafetyGate = false` if they've validated the task class is safe.
211
-
212
- **Rules + weights** - see `classify-plan-safety.mjs` header comment for the canonical list. Summary: `file-count-high` (30) / `destructive-verb` (25) / `security-path` (35) / `delete-without-test` (30) / `schema-migration` (25) / `infrastructure` (20). Threshold: score β‰₯ 50 flips `recommendPause` true. Tuned so any single heavy signal or any two medium signals trigger the pause.
213
-
214
- **Telemetry:** emit `phase.plan.safety` OTel span (when `MULTI_AGENT_OTEL_SPANS=1`) with `{score, recommendPause, rules}` so post-hoc analysis can tune weights.
215
-
216
- Otherwise (normal mode), run the gate. The gate has **two modes** that chain: Clarification β†’ Approval. Each round emits a progress line and persists to `agent-state.json.phases["2"]` for audit.
217
-
218
- ##### 5a - Clarification Mode (conditional, max 2 rounds)
219
-
220
- Trigger if the plan Fable produced in Step 1-4 carries ANY ambiguity signal from Phase 1 analysis:
221
-
222
- | Signal | Check |
223
- |---|---|
224
- | Vague acceptance | Jira/issue description < 200 chars AND no `## Acceptance Criteria` section |
225
- | UI work, no design | Task touches `*View.swift` / `*Screen.kt` but Phase 1 captured no Figma URL |
226
- | API work, no contract | Task touches network/repository layer but no endpoint/OpenAPI reference in Phase 1 |
227
- | Ambiguous language | Phase 1 analysis flagged `ambiguityScore >= 2` (e.g. "improve", "fix", "update" with no object) |
228
- | Parent-story scope drift | `state.relatedIssues[]` names a sibling overlapping this scope |
229
-
230
- If any signal trips, DO NOT render the plan yet. Render structured questions:
231
-
232
- ```
233
- πŸ“‹ Development Plan - {taskId}
234
- ────────────────────────────────────
235
- ⚠️ There are points that need clarifying before drafting a plan for this item:
236
-
237
- 1. {question 1}
238
- 2. {question 2}
239
- 3. {question 3}
240
-
241
- Write your answers, then I will draft the plan.
242
- (or tell me you want to pause - the task can be resumed later)
243
- ```
244
-
245
- Wait for user response. On reply:
246
-
247
- - **User asks to pause / cancel** (intent, in any language) β†’ `state.status = "paused"`, log `🧠 Phase 2: Gate aborted (clarification)`, stop.
248
- - **Free-text answer** β†’ append to `state.phases["2"].clarificationAnswers`, bump `clarificationRounds`, re-run Steps 1-4 with answers as context, re-evaluate signals.
249
-
250
- **Cap: 2 rounds.** If `clarificationRounds === 2` and signals still trip, do NOT ask again - render the plan anyway with a `⚠️ best-effort (item still unclear)` banner in Mode 5b, and let the user decide via approval/edit. This prevents infinite grooming loops on chronically under-specified tickets.
251
-
252
- Persist to `state.phases["2"]`:
253
-
254
- ```json
255
- {
256
- "clarificationRounds": 1,
257
- "clarificationQuestions": ["Pasaport scan sonrasΔ±...", "Backend endpoint...", "Android kapsamda mΔ±?"],
258
- "clarificationAnswers": ["Silent reject. Toast yok.", "Backend hazΔ±r, v3.4 release'de.", "Sadece iOS bu task'ta."]
259
- }
260
- ```
261
-
262
- ##### 5b - Approval Loop (always runs after clarification resolves or is skipped)
263
-
264
- Render the plan:
265
-
266
- ```
267
- πŸ“‹ Development Plan - {taskId} {best-effort banner if 5a capped}
268
- ────────────────────────────────────
269
- Summary: {1-line summary}
270
- Approach: {1-paragraph approach}
271
- Risk: {low|medium|high} - {reason}
272
- Scope: {S|M|L} ({N} files, {K} new)
273
- Files to touch:
274
- - {path1}
275
- - {path2}
276
-
277
- Todos ({N} total):
278
- 1. {subject} - {approach}
279
- 2. {subject} - {approach} (blocked by #1)
280
- ...
281
-
282
- ```
283
-
284
- Then ask for the decision with a **native `AskUserQuestion` picker** (never a typed-keyword prompt):
285
-
286
- - `question`: "Do you approve this plan?" (rendered in `outputLanguage`)
287
- - `header`: "Plan" (English, <=12 chars)
288
- - `options` (label + description in `outputLanguage`; the names below are the option
289
- SEMANTICS, not strings to print):
290
- - option 1 - "Approve": proceed to development (Phase 3)
291
- - option 2 - "Cancel": pause the task; resume later with `:resume`
292
-
293
- The picker's built-in **Other** field is the free-text edit channel: the user types an edit request there (e.g. "also look at the auth service but keep LoginView out of scope") instead of typing a keyword. Per the `rules.md` matrix `question`, `label` and `description` all follow `outputLanguage`; only `header` stays English. Branch on which option was picked, never on its rendered text.
294
-
295
- Handle the selection:
296
-
297
- - **Option 1 (Approve)**:
298
- - Set `state.phases["2"].planApprovedAt = now()`, bump `planIterations` if not yet set (default 1)
299
- - Log `🧠 Phase 2: Plan approved (iterations={N}, clarificationRounds={M})`
300
- - Proceed to Phase 3
301
- - **Cancel**:
302
- - Set `state.status = "paused"`, log `🧠 Phase 2: Plan aborted by user`, stop
303
- - **Other (free-text edit request)**:
304
- - Treat the typed text as an edit request. Append to `state.phases["2"].planEditRequests`, bump `planIterations`
305
- - Pass the edit request + current plan to the planning model (Fable; Opus when the fallback ladder engages); it revises and returns a new plan (same schema, same validator)
306
- - Re-render the plan (5b), loop
307
-
308
- No hard cap on edit iterations - the user controls exit via the Approve / Cancel options. Between iterations, keep only the **latest plan** as canonical; previous renders are in the log for audit but do not re-enter the validator.
309
-
310
- **Validator**: every revised plan goes through `node $HOME/.claude/scripts/validate-planning.mjs -` before re-render. If validation fails after an edit, log `⚠️ Phase 2: Plan validator failed after edit request #N - retrying the planning model once` and retry once; on second failure surface the validator error to the user and go back to approval prompt with the pre-edit plan.
311
-
312
- #### Step 6 - Mode-specific short-circuit (reference)
313
-
314
- The pipeline shapes interact with the gate as follows. This table is the source of truth for the gate's mode-awareness - if behavior diverges in code, fix the code, not the table.
315
-
316
- | Mode | Clarification | Approval Loop | Safety Classifier | Notes |
317
- |---|---|---|---|---|
318
- | Full, interactive (`/multi-agent`, `:local`) | βœ… (max 2 rounds) | βœ… | - (redundant when a human approves) | Full gate |
319
- | Short (depth picker answered Short) | ❌ | ❌ | - | Phase 1-2 skipped at Step 7.5; Phase 3 starts immediately |
320
- | `autopilot`, `:local-autopilot` | ❌ | ❌ conditional | βœ… (if `autopilotSafetyGate !== false`) | Always Full, so a plan exists. Safe plans proceed silently; high-risk plans trigger a one-time manual approval. Log records the score. |
321
-
322
- **Why autopilot now has an escape hatch:** the old "zero interaction - fully trust the scope" contract held well for tightly-scoped batch workflows (figma component iteration over known-safe components) but broke in edge cases - schema migrations auto-merging, security-path drift going silent, delete-without-test sprawls. The safety classifier (Step 5c) is opt-out so the default protects against the edge-case cost; users running known-safe workflows can flip `prefs.global.autopilotSafetyGate = false` to restore pre-v7.0 behavior.
323
-
324
- #### Telemetry - token forwarding
325
-
326
- After plan generation (and after each edit-loop iteration), forward the planning model's call totals so Phase 7's Cost Breakdown captures Phase 2 (`model=` names the rung that actually ran: `fable`, or `opus` after a fallback step):
327
-
328
- ```bash
329
- LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 2 plan.generated \
330
- model=fable tokens_in=$IN tokens_out=$OUT duration_ms=$DUR iteration=$N
331
- ```
332
-
333
- Best-effort. See `$HOME/.claude/multi-agent-refs/progress-contract.md#token-telemetry-forwarding`.
334
-
335
- ---
336
-
337
- ## Token telemetry - invoke after every LLM call
338
-
339
- ```bash
340
- bash $HOME/.claude/scripts/phase-tracker.sh tokens 2 <input_count> <output_count>
341
- ```
342
-
343
- Contract and rationale: `progress-contract.md` -> Token telemetry forwarding.
344
-