@mmerterden/multi-agent-pipeline 17.6.0 → 19.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (272) hide show
  1. package/CHANGELOG.md +310 -0
  2. package/README.md +76 -18
  3. package/README.tr.md +55 -16
  4. package/docs/adr/0002-instruction-driven-flag.md +1 -0
  5. package/docs/adr/0005-lazy-phase-docs.md +11 -1
  6. package/docs/adr/0008-installer-modularization-and-secret-leak-defense.md +1 -0
  7. package/docs/adr/0010-own-code-graph.md +1 -0
  8. package/docs/adr/0011-dormant-ci.md +25 -1
  9. package/docs/adr/0014-six-phase-consolidation.md +134 -0
  10. package/docs/adr/README.md +2 -1
  11. package/docs/architecture.md +37 -38
  12. package/docs/best-practices.md +1 -1
  13. package/docs/ecosystem.md +37 -26
  14. package/docs/engineering.md +1 -1
  15. package/docs/facts.json +45 -0
  16. package/docs/features.md +54 -53
  17. package/docs/performance.md +5 -5
  18. package/docs/recovery-guide.md +9 -9
  19. package/docs/server-readiness.md +188 -0
  20. package/docs/token-budget-history.md +3 -1
  21. package/index.js +18 -3
  22. package/install/_codex-agents.mjs +1 -1
  23. package/install/_common.mjs +42 -17
  24. package/install/_dev-only-files.mjs +8 -0
  25. package/install/_unattended-profile.mjs +113 -0
  26. package/install/index.mjs +48 -0
  27. package/install/templates/claude-hooks.json +1 -1
  28. package/install/templates/codex-instructions.md +1 -1
  29. package/install/templates/copilot-instructions.md +28 -28
  30. package/manifest.json +1065 -0
  31. package/package.json +6 -3
  32. package/pipeline/agents/dev-critic.md +3 -3
  33. package/pipeline/commands/figma-to-swiftui.md +1 -1
  34. package/pipeline/commands/multi-agent/SKILL.md +8 -8
  35. package/pipeline/commands/multi-agent/analysis/SKILL.md +9 -9
  36. package/pipeline/commands/multi-agent/autopilot/SKILL.md +7 -7
  37. package/pipeline/commands/multi-agent/channels/SKILL.md +15 -15
  38. package/pipeline/commands/multi-agent/diff-explain/SKILL.md +6 -6
  39. package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +1 -1
  40. package/pipeline/commands/multi-agent/graph/SKILL.md +1 -1
  41. package/pipeline/commands/multi-agent/help/SKILL.md +62 -62
  42. package/pipeline/commands/multi-agent/language/SKILL.md +2 -2
  43. package/pipeline/commands/multi-agent/local/SKILL.md +11 -11
  44. package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +13 -13
  45. package/pipeline/commands/multi-agent/log/SKILL.md +2 -2
  46. package/pipeline/commands/multi-agent/manual-test/SKILL.md +9 -9
  47. package/pipeline/commands/multi-agent/model/SKILL.md +69 -0
  48. package/pipeline/commands/multi-agent/refactor/SKILL.md +3 -3
  49. package/pipeline/commands/multi-agent/resume/SKILL.md +4 -4
  50. package/pipeline/commands/multi-agent/resume-local/SKILL.md +19 -17
  51. package/pipeline/commands/multi-agent/review/SKILL.md +1 -1
  52. package/pipeline/commands/multi-agent/route-off/SKILL.md +36 -0
  53. package/pipeline/commands/multi-agent/route-on/SKILL.md +74 -0
  54. package/pipeline/commands/multi-agent/route-status/SKILL.md +56 -0
  55. package/pipeline/commands/multi-agent/setup/SKILL.md +2 -2
  56. package/pipeline/commands/multi-agent/status/SKILL.md +54 -23
  57. package/pipeline/commands/multi-agent/steer/SKILL.md +2 -2
  58. package/pipeline/commands/multi-agent/sync/SKILL.md +12 -13
  59. package/pipeline/commands/multi-agent/test/SKILL.md +1 -1
  60. package/pipeline/lib/_jira-auth.sh +8 -0
  61. package/pipeline/lib/analysis-jira-write.sh +32 -0
  62. package/pipeline/lib/ask-choice.sh +13 -2
  63. package/pipeline/lib/autopilot-state.sh +8 -0
  64. package/pipeline/lib/credential-inventory.sh +1 -1
  65. package/pipeline/lib/fatal.mjs +129 -0
  66. package/pipeline/lib/fetch-fortify.sh +1 -1
  67. package/pipeline/lib/figma-mcp-refresh.sh +18 -0
  68. package/pipeline/lib/figma-screenshot.sh +18 -0
  69. package/pipeline/lib/invoked-directly.mjs +43 -0
  70. package/pipeline/lib/jira-publish.sh +42 -0
  71. package/pipeline/lib/md2confluence-v3.py +47 -0
  72. package/pipeline/lib/model-rung.sh +142 -0
  73. package/pipeline/lib/outbound-gate.mjs +175 -0
  74. package/pipeline/lib/phase-schema.mjs +88 -0
  75. package/pipeline/lib/plan-todos.sh +32 -11
  76. package/pipeline/lib/post-pr-review.sh +77 -8
  77. package/pipeline/lib/repo-hygiene.sh +8 -3
  78. package/pipeline/lib/require-jq.sh +40 -0
  79. package/pipeline/lib/route-state.sh +161 -0
  80. package/pipeline/lib/run-paths.sh +335 -0
  81. package/pipeline/multi-agent-refs/_account-picker.md +1 -1
  82. package/pipeline/multi-agent-refs/_dev-context.md +1 -1
  83. package/pipeline/multi-agent-refs/_input-parser.md +1 -1
  84. package/pipeline/multi-agent-refs/analysis/evidence.md +0 -9
  85. package/pipeline/multi-agent-refs/analysis/intake.md +1 -1
  86. package/pipeline/multi-agent-refs/analysis/locked.md +21 -22
  87. package/pipeline/multi-agent-refs/analysis/render.md +1 -1
  88. package/pipeline/multi-agent-refs/analysis/synthesis.md +12 -6
  89. package/pipeline/multi-agent-refs/android-guide.md +1 -1
  90. package/pipeline/multi-agent-refs/audit-guide.md +13 -13
  91. package/pipeline/multi-agent-refs/channels/issue-comment.md +2 -2
  92. package/pipeline/multi-agent-refs/channels/jira.md +3 -3
  93. package/pipeline/multi-agent-refs/channels/pr.md +4 -4
  94. package/pipeline/multi-agent-refs/channels/wiki.md +1 -1
  95. package/pipeline/multi-agent-refs/component-dispatch.md +3 -3
  96. package/pipeline/multi-agent-refs/cross-cli-contract.md +31 -6
  97. package/pipeline/multi-agent-refs/features/autopilot-circuit-breaker.md +74 -4
  98. package/pipeline/multi-agent-refs/features/code-graph.md +5 -5
  99. package/pipeline/multi-agent-refs/features/cost-analysis.md +93 -0
  100. package/pipeline/multi-agent-refs/features/design-conformance.md +1 -1
  101. package/pipeline/multi-agent-refs/features/dev-critic.md +3 -3
  102. package/pipeline/multi-agent-refs/features/doctor.md +47 -2
  103. package/pipeline/multi-agent-refs/features/external-context-injection.md +3 -3
  104. package/pipeline/multi-agent-refs/features/maturity-followup.md +3 -3
  105. package/pipeline/multi-agent-refs/features/model-fallback.md +5 -5
  106. package/pipeline/multi-agent-refs/features/plan-todos.md +1 -1
  107. package/pipeline/multi-agent-refs/features/repo-map.md +1 -1
  108. package/pipeline/multi-agent-refs/features/review-delta.md +3 -3
  109. package/pipeline/multi-agent-refs/features/review-multi-repo.md +1 -1
  110. package/pipeline/multi-agent-refs/features/scope-check.md +4 -4
  111. package/pipeline/multi-agent-refs/features/skill-conformance.md +2 -2
  112. package/pipeline/multi-agent-refs/features/stack-skill-routing.md +1 -1
  113. package/pipeline/multi-agent-refs/features/verify-by-test.md +4 -4
  114. package/pipeline/multi-agent-refs/features/verify.md +83 -0
  115. package/pipeline/multi-agent-refs/features/visual-evidence.md +19 -19
  116. package/pipeline/multi-agent-refs/features/worktree-finalize.md +6 -6
  117. package/pipeline/multi-agent-refs/issue-jira-triad.md +10 -10
  118. package/pipeline/multi-agent-refs/knowledge.md +11 -11
  119. package/pipeline/multi-agent-refs/multi-repo-integration-build.md +13 -13
  120. package/pipeline/multi-agent-refs/payload-contracts.md +8 -8
  121. package/pipeline/multi-agent-refs/phases/log-format.md +10 -10
  122. package/pipeline/multi-agent-refs/phases/modes.md +30 -30
  123. package/pipeline/multi-agent-refs/phases/operations.md +21 -10
  124. package/pipeline/multi-agent-refs/phases/phase-0-init.md +25 -25
  125. package/pipeline/multi-agent-refs/phases/phase-1-plan.md +599 -0
  126. package/pipeline/multi-agent-refs/phases/{phase-3-dev.md → phase-2-dev.md} +129 -49
  127. package/pipeline/multi-agent-refs/phases/{phase-4-review.md → phase-3-review.md} +225 -107
  128. package/pipeline/multi-agent-refs/phases/{phase-6-commit.md → phase-4-commit.md} +23 -23
  129. package/pipeline/multi-agent-refs/phases/{phase-7-report.md → phase-5-report.md} +29 -29
  130. package/pipeline/multi-agent-refs/phases.md +44 -48
  131. package/pipeline/multi-agent-refs/picker-contract.md +1 -1
  132. package/pipeline/multi-agent-refs/progress-contract.md +6 -6
  133. package/pipeline/multi-agent-refs/readiness-review.md +1 -1
  134. package/pipeline/multi-agent-refs/rules.md +7 -7
  135. package/pipeline/multi-agent-refs/swiftui-guide.md +2 -2
  136. package/pipeline/multi-agent-refs/tracker-contract.md +31 -32
  137. package/pipeline/multi-agent-refs/unattended-contract.md +129 -0
  138. package/pipeline/multi-agent-refs/wiki-capture.md +14 -14
  139. package/pipeline/preferences-template.json +9 -1
  140. package/pipeline/rules/outside-the-pipeline.md +1 -1
  141. package/pipeline/schemas/agent-state.schema.json +50 -50
  142. package/pipeline/schemas/analysis-output.schema.json +2 -2
  143. package/pipeline/schemas/autopilot-config.schema.json +1 -1
  144. package/pipeline/schemas/code-graph.schema.json +1 -1
  145. package/pipeline/schemas/criteria-manifest.schema.json +1 -1
  146. package/pipeline/schemas/dev-critic-output.schema.json +1 -1
  147. package/pipeline/schemas/diff-risk.schema.json +1 -1
  148. package/pipeline/schemas/migrations/prefs-2.4.0-to-2.5.0.mjs +2 -2
  149. package/pipeline/schemas/migrations/prefs-2.6.0-to-2.7.0.mjs +31 -0
  150. package/pipeline/schemas/migrations/state-2.1.0-to-2.2.0.mjs +129 -0
  151. package/pipeline/schemas/phases.json +105 -0
  152. package/pipeline/schemas/plan-todos.schema.json +5 -5
  153. package/pipeline/schemas/planning-output.schema.json +1 -1
  154. package/pipeline/schemas/prefs.schema.json +100 -56
  155. package/pipeline/schemas/reviewer-output.schema.json +3 -3
  156. package/pipeline/schemas/route-config.schema.json +74 -0
  157. package/pipeline/schemas/scope-check.schema.json +1 -1
  158. package/pipeline/schemas/test-gap.schema.json +1 -1
  159. package/pipeline/schemas/token-budget.json +12 -18
  160. package/pipeline/schemas/triage-output.schema.json +6 -6
  161. package/pipeline/scripts/README.md +3 -3
  162. package/pipeline/scripts/_code-graph.mjs +2 -2
  163. package/pipeline/scripts/_run-paths.mjs +372 -0
  164. package/pipeline/scripts/_smoke-root.sh +1 -1
  165. package/pipeline/scripts/aggregate-metrics.mjs +65 -65
  166. package/pipeline/scripts/autopilot-arming.mjs +2 -1
  167. package/pipeline/scripts/autopilot-intake.mjs +2 -1
  168. package/pipeline/scripts/autopilot-runner.mjs +206 -2
  169. package/pipeline/scripts/build-references.mjs +2 -1
  170. package/pipeline/scripts/build-stack-plugins.mjs +10 -2
  171. package/pipeline/scripts/capture-evidence.sh +7 -2
  172. package/pipeline/scripts/capture-flush.sh +8 -8
  173. package/pipeline/scripts/capture-resume.sh +3 -3
  174. package/pipeline/scripts/classify-plan-safety.mjs +3 -2
  175. package/pipeline/scripts/cost-analyze.mjs +600 -0
  176. package/pipeline/scripts/cost-budget-check.mjs +4 -12
  177. package/pipeline/scripts/council-view.mjs +2 -1
  178. package/pipeline/scripts/crush-json.mjs +2 -1
  179. package/pipeline/scripts/diff-explain.mjs +7 -10
  180. package/pipeline/scripts/diff-risk-score.mjs +2 -1
  181. package/pipeline/scripts/doctor.mjs +140 -6
  182. package/pipeline/scripts/evidence-gate.mjs +9 -3
  183. package/pipeline/scripts/feedback-send.mjs +12 -2
  184. package/pipeline/scripts/gc-abandoned.sh +32 -16
  185. package/pipeline/scripts/gc-tmp.sh +1 -1
  186. package/pipeline/scripts/gc-worktrees.sh +12 -5
  187. package/pipeline/scripts/gen-facts.mjs +175 -0
  188. package/pipeline/scripts/gen-mode-dispatch.mjs +32 -37
  189. package/pipeline/scripts/gen-ref-toc.mjs +1 -1
  190. package/pipeline/scripts/github-ssh-setup.sh +64 -7
  191. package/pipeline/scripts/graph-mermaid.mjs +4 -2
  192. package/pipeline/scripts/graph-report.mjs +1 -1
  193. package/pipeline/scripts/jira-attach.sh +1 -1
  194. package/pipeline/scripts/keychain-save.sh +101 -30
  195. package/pipeline/scripts/learn-from-transcripts.mjs +3 -2
  196. package/pipeline/scripts/learning-curve.mjs +36 -31
  197. package/pipeline/scripts/log-metric.sh +17 -4
  198. package/pipeline/scripts/make-manifest.mjs +199 -0
  199. package/pipeline/scripts/memory-save.sh +1 -1
  200. package/pipeline/scripts/migrate-prefs.mjs +24 -6
  201. package/pipeline/scripts/migrate-state.mjs +94 -4
  202. package/pipeline/scripts/phase-banner.sh +26 -22
  203. package/pipeline/scripts/phase-tracker.sh +48 -10
  204. package/pipeline/scripts/plan-coverage-gate.mjs +8 -4
  205. package/pipeline/scripts/pre-commit-check.sh +7 -0
  206. package/pipeline/scripts/pre-push-check.sh +7 -0
  207. package/pipeline/scripts/purge.sh +23 -6
  208. package/pipeline/scripts/render-agent-log-cost.sh +10 -3
  209. package/pipeline/scripts/render-cost-summary.sh +9 -2
  210. package/pipeline/scripts/render-work-summary.sh +14 -7
  211. package/pipeline/scripts/review-file-filter.mjs +5 -3
  212. package/pipeline/scripts/review-scope.mjs +2 -1
  213. package/pipeline/scripts/routine-registry.mjs +2 -1
  214. package/pipeline/scripts/run-aggregator.mjs +26 -20
  215. package/pipeline/scripts/run-metrics.mjs +4 -2
  216. package/pipeline/scripts/runs-index.mjs +353 -0
  217. package/pipeline/scripts/scorecard-snapshot.mjs +178 -0
  218. package/pipeline/scripts/search-logs.sh +18 -0
  219. package/pipeline/scripts/smoke-cross-cli-behavior.sh +6 -6
  220. package/pipeline/scripts/smoke-schema-validation.sh +26 -7
  221. package/pipeline/scripts/test-gap-scan.mjs +2 -1
  222. package/pipeline/scripts/test-integrity-gate.mjs +2 -1
  223. package/pipeline/scripts/token-budget-report.mjs +13 -2
  224. package/pipeline/scripts/triage-memory.mjs +2 -2
  225. package/pipeline/scripts/update-issue-progress.sh +56 -7
  226. package/pipeline/scripts/usage-report.mjs +12 -1
  227. package/pipeline/scripts/validate-analysis-doc.mjs +75 -18
  228. package/pipeline/scripts/validate-code-graph.mjs +6 -3
  229. package/pipeline/scripts/validate-complaint-doc.mjs +2 -1
  230. package/pipeline/scripts/validate-diff-risk.mjs +6 -3
  231. package/pipeline/scripts/validate-planning.mjs +1 -1
  232. package/pipeline/scripts/validate-reviewer.mjs +1 -1
  233. package/pipeline/scripts/validate-state.mjs +45 -5
  234. package/pipeline/scripts/validate-test-gap.mjs +6 -3
  235. package/pipeline/scripts/validate-triage.mjs +6 -4
  236. package/pipeline/scripts/verify-citations.mjs +4 -2
  237. package/pipeline/scripts/verify.mjs +327 -0
  238. package/pipeline/scripts/worktree-finalize.sh +18 -9
  239. package/pipeline/scripts/write-state.mjs +154 -15
  240. package/pipeline/skills/.skill-manifest.json +37 -21
  241. package/pipeline/skills/.skills-index.json +104 -5
  242. package/pipeline/skills/shared/README.md +15 -6
  243. package/pipeline/skills/shared/core/apple-archive-compliance/SKILL.md +2 -2
  244. package/pipeline/skills/shared/core/google-play-compliance/SKILL.md +2 -2
  245. package/pipeline/skills/shared/core/multi-agent/SKILL.md +69 -71
  246. package/pipeline/skills/shared/core/multi-agent-autopilot/SKILL.md +3 -3
  247. package/pipeline/skills/shared/core/multi-agent-channels/SKILL.md +14 -14
  248. package/pipeline/skills/shared/core/multi-agent-diff-explain/SKILL.md +5 -5
  249. package/pipeline/skills/shared/core/multi-agent-graph/SKILL.md +1 -1
  250. package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +25 -23
  251. package/pipeline/skills/shared/core/multi-agent-language/SKILL.md +2 -2
  252. package/pipeline/skills/shared/core/multi-agent-local/SKILL.md +2 -2
  253. package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +8 -8
  254. package/pipeline/skills/shared/core/multi-agent-manual-test/SKILL.md +6 -6
  255. package/pipeline/skills/shared/core/multi-agent-model/SKILL.md +71 -0
  256. package/pipeline/skills/shared/core/multi-agent-refactor/SKILL.md +3 -3
  257. package/pipeline/skills/shared/core/multi-agent-resume/SKILL.md +1 -1
  258. package/pipeline/skills/shared/core/multi-agent-resume-local/SKILL.md +7 -7
  259. package/pipeline/skills/shared/core/multi-agent-route-off/SKILL.md +39 -0
  260. package/pipeline/skills/shared/core/multi-agent-route-on/SKILL.md +76 -0
  261. package/pipeline/skills/shared/core/multi-agent-route-status/SKILL.md +59 -0
  262. package/pipeline/skills/shared/core/multi-agent-setup/SKILL.md +1 -1
  263. package/pipeline/skills/shared/core/multi-agent-status/SKILL.md +35 -11
  264. package/pipeline/skills/shared/core/multi-agent-steer/SKILL.md +2 -2
  265. package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +6 -5
  266. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/package_app.sh +4 -1
  267. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/setup_dev_signing.sh +4 -1
  268. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/sign-and-notarize.sh +2 -1
  269. package/pipeline/skills/skills-index.md +13 -4
  270. package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +0 -263
  271. package/pipeline/multi-agent-refs/phases/phase-2-planning.md +0 -344
  272. package/pipeline/multi-agent-refs/phases/phase-5-test.md +0 -182
@@ -1,344 +0,0 @@
1
- ### Phase 2: Planning (Fable)
2
-
3
- > **TLDR** - Fable decomposes the analysis (Opus when the fallback ladder engages) into concrete tasks with file-level targets, risk grading, and architecture review. Before Phase 3 a **Plan Approval Gate** runs in normal mode: if the Jira/issue description is ambiguous the orchestrator asks the user structured clarification questions (max 2 rounds) - once scope is clear it renders the plan and loops on free-text edit requests until the user approves or aborts. The gate is **skipped entirely** for a Short run and for `autopilot` (a Short run has no plan to approve; autopilot may not ask).
4
-
5
- <!-- progress-contract: applied -->
6
- Progress emission per `$HOME/.claude/multi-agent-refs/progress-contract.md` - lines for plan-draft start, clarification-ask, clarification-answer, plan render, plan-edit-request, plan-approved, plan-aborted.
7
-
8
- ## Phase 2 Pre-flight (BLOCKING, v9.0.0)
9
-
10
- Phase 2 Planning consumes the analysis document. **Figma** MCP / REST forbidden; the toolkit MCP is not.
11
-
12
- 1. **Analysis document presence**: read `state.analysis.docStatus`, which Phase 1 Step 4 set.
13
- - `produced` | `reused` -> the file is at `state.analysis.docPath[]`; continue with steps 2-4.
14
- - `not-applicable` -> no document by design (bugfix/chore, no Figma). Record it, skip steps 2-4, plan from `state.analysis` alone.
15
- - Unreadable or key missing -> `ERR: Phase 1 reported <status> but no analysis doc is readable. Resume with /multi-agent:resume #N.` Producing it is Phase 1's job; never send the user to another command.
16
-
17
- 2. **Parse YAML front-matter** into `state.analysis.frontMatter` and abort below `template_version: v3` - same contract as `phase-3-dev.md` step 2.
18
-
19
- 3. **Section coverage check**: verify these sections are non-empty (template v3 required sections per Locked 2):
20
- - Section 1 Summary, Section 2 Goals + Non-Goals, Section 4 User Stories, Section 9 API Contracts, Section 13 Architecture Plan, Section 14 Files to Add, Section 20 Risks, Section 21 References
21
- - Missing -> WARN, plan is allowed to proceed but Phase 4 reviewer flags it.
22
-
23
- 4. **Convert analysis tasks to plan**: Section 14 Files-to-Add becomes the seed task list. Carry each row's tag onto its todo as `sourceTag` (`Reuse` | `Add new` | `Modify`); it is an instruction Phase 3 follows and Phase 4 checks, not a label.
24
-
25
- 5. **MCP forbidden**: same rule as Phase 3.
26
-
27
- #### Input contract
28
-
29
- Phase 2 consumes the Phase 1 output object conforming to `$HOME/.claude/schemas/analysis-output.schema.json`. Read `state.analysis` (the explorer's return value) and treat its `touchedAreas` and `risks` arrays (plus `stack` and `summary`) as authoritative input - do not re-explore the codebase here. (Field names are exactly those in the analysis schema; the planner's own `targetFiles` belongs to `planning-output.schema.json`, not the analysis input.)
30
-
31
- #### Step 0.9 - Post-analysis confirmation (derive first, ask only what cannot be derived)
32
-
33
- Runs when `state.analysis.docStatus` is `produced` or `reused`, before any planning.
34
-
35
- **Derived is shown, not asked**: platform set, the seven convention groups, existing components (Code Connect + `uiComponents`), localization keys, analytics events, DI registration, test-method naming. **Only Section 20 rows are asked**, through `$HOME/.claude/multi-agent-refs/analysis/resolve.md` - one row, at most three source-labeled candidates, plus Defer. The engine never invents.
36
-
37
- ```
38
- Türetildi (onay için): platform=ios,android · 7 konvansiyon grubu (5 high, 2 medium)
39
- · 12 mevcut bileşen (9 Code Connect bağlı) · 34 lokalizasyon anahtarı
40
- Sorulacak: 4 açık soru (Bölüm 20)
41
- ```
42
-
43
- A corrected value rewrites its Pass B footnote as `^[user-override: resolved <date>]` (Locked 24). Deferred rows stay in `state.analysis.openQuestions[]` and Phase 4 flags them `review_blocking`. Here and not Phase 4 because Phase 4 runs after development - answered there is answered too late. Autopilot skips the asking and defers every row.
44
-
45
- #### Step 1 - Task Decomposition
46
-
47
- Break the work from Phase 1 analysis into discrete, implementable tasks:
48
-
49
- ```
50
- For each identified change area:
51
- 1. Define a task with clear scope (one file group or one logical change)
52
- 2. Estimate complexity: trivial (1 file) / moderate (2-5 files) / complex (6+ files)
53
- 3. Identify dependencies between tasks (which must complete before others)
54
- ```
55
-
56
- Create tasks using TaskCreate with imperative subject, description (what/which files/expected behavior), and `addBlockedBy` for dependencies.
57
-
58
- ##### Analysis citation requirement (every UI task)
59
-
60
- Every UI-touching task in the plan MUST cite:
61
-
62
- - The analysis Section 6 (Bileşen Envanteri) row it implements (one task per row when a single component covers several variants), AND
63
- - The canonical component name, sourced from the analysis doc:
64
- - **Code Connect mapping present**: take the component name verbatim from the matching repo `*.figma.swift` / `*.figma.kt` row referenced by analysis Section 6.
65
- - **Mapping absent**: cite "best-fit pending design review" and add a Risk row to the plan. The task description MUST also flag that Phase 4 will gate it as `review_blocking`.
66
-
67
- Tasks without an analysis Section 6 citation when the doc lists UI components are rejected at the plan-approval gate; the user is asked to either re-run `/multi-agent:analysis` to extend Section 6 or rescope the task to a non-UI change. Direct Figma fetches in Phase 2 are forbidden (Locked decision 30).
68
-
69
- Example task graph:
70
-
71
- ```
72
- Task 1: Add new token to common module (no deps)
73
- Task 2: Create ButtonConfiguration.swift (blocked by 1)
74
- Task 3: Create ButtonView.swift (blocked by 2)
75
- Task 4: Add ButtonView+Modifiers.swift (blocked by 3)
76
- Task 5: Write ViewInspector tests (blocked by 3)
77
- Task 6: Write snapshot tests (blocked by 3)
78
- ```
79
-
80
- #### Step 2 - Architecture Review (conditional)
81
-
82
- Trigger architecture review if ANY of these are true:
83
-
84
- - New module or package being created
85
- - Cross-module dependency being added
86
- - Public API surface changing
87
- - Data model / schema change
88
- - Navigation flow change
89
-
90
- If triggered:
91
-
92
- 1. Launch Agent with `subagent_type: "ios-architect"` (or `architecture` skill for non-iOS)
93
- 2. Provide: task list, affected files, proposed approach
94
- 3. Agent returns: recommendation, risks, alternative approaches
95
- 4. Incorporate recommendations into task descriptions
96
-
97
- If NOT triggered: skip, log "Phase 2: Architecture review - not needed (scope contained)"
98
-
99
- #### Step 3 - Development Approach per Task
100
-
101
- For each task, determine:
102
- | Approach | When | Example |
103
- |----------|------|---------|
104
- | **New file** | Feature doesn't exist | Create ButtonConfiguration.swift |
105
- | **Modify** | Extending existing code | Add property to existing Configuration |
106
- | **Refactor** | Restructuring without behavior change | Extract protocol from class |
107
- | **Fix** | Bug correction | Fix nil crash in edge case |
108
-
109
- Store approach in task metadata for Phase 3 agent.
110
-
111
- #### Step 4 - Skill Selection per Task
112
-
113
- Based on Phase 1 `detectedStack`, assign relevant skills:
114
-
115
- - iOS tasks -> `ai-ios-toolkit:*` SwiftUI skills, iOS patterns
116
- - Python tasks -> `ai-backend-toolkit:fastapi-pro`, `ai-backend-toolkit:api-patterns`
117
- - Node tasks -> `ai-backend-toolkit:nodejs-backend-patterns`
118
- - Security-sensitive -> `ai-backend-toolkit:api-security-best-practices`
119
- - Multi-submodule -> `ai-backend-toolkit:monorepo-architect`
120
-
121
- #### Output contract
122
-
123
- Phase 2 produces an object conforming to `$HOME/.claude/schemas/planning-output.schema.json` - `tasks[]` with `id`, `title`, `type`, `files`, plus optional `dependsOn` and `acceptanceCriteria`. Phase 3 reads `tasks[]` in dependency order; the schema's `dependsOn` field drives the ready-task picker.
124
-
125
- **Required: validator gate (deterministic) - run on the persisted file before the approval gate renders the plan; the validator's exit code decides, not the LLM turn:**
126
-
127
- ```bash
128
- PLAN_FILE="$WORKTREE/.pipeline/plan.json"
129
- mkdir -p "$(dirname "$PLAN_FILE")"
130
- printf '%s' "$PLAN_JSON" > "$PLAN_FILE"
131
- node $HOME/.claude/scripts/validate-planning.mjs "$PLAN_FILE"
132
- ```
133
-
134
- Progress line: ` → checking validator validate-planning`
135
-
136
- Non-zero exit fails CLOSED: emit the validator stderr + `errors[]` verbatim, attempt ONE self-correction rework (re-invoke the planner with the errors quoted, overwrite `$PLAN_FILE`), re-run the validator. If it fails again -> HALT the phase (never enter Phase 3 with an invalid plan) with recovery hint: `ERR: plan output failed validate-planning.mjs twice. Inspect $PLAN_FILE against $HOME/.claude/schemas/planning-output.schema.json, then resume with /multi-agent:resume #N.` Record `agent-state.phases["2"].validator` (`pass` | `pass-after-rework` | `halted`).
137
-
138
- Log: "Phase 2: Plan - {N} tasks created, {M} with architecture review, validator:pass"
139
-
140
- #### Step 4.45 - Put the plan on the widget (required)
141
-
142
- ```bash
143
- printf '%s' "$PLAN_JSON" | bash "$HOME/.claude/scripts/phase-tracker.sh" plan 3
144
- ```
145
-
146
- Then do what its output asks - it prints a tile rebuild, because the widget
147
- orders by creation. `dependsOn` becomes `addBlockedBy`, so the widget answers
148
- "why has t3 not started". Required, unlike Step 4.5: the plan was always computed
149
- and stored, only invisible.
150
-
151
- #### Step 4.5 - Emit Plan Todo List (opt-in)
152
-
153
- **Gated by `prefs.global.planTodos.enabled`** (default: `false`). When enabled, after the planning-output JSON validates and BEFORE the approval gate, transform `tasks[]` into a structured Todo list conforming to `$HOME/.claude/schemas/plan-todos.schema.json` and persist into `agent-state.plan`. The plan is rendered as a live, always-visible Todo list.
154
-
155
- ```bash
156
- printf '%s' "$PLAN_JSON" | bash "$HOME/.claude/lib/plan-todos.sh" set "$TASK_ID" -
157
- ```
158
-
159
- `set` accepts a planning-output document and converts it itself. The conversion
160
- used to be written out here, which made the `tasks[]`-to-`todos[]` mapping a
161
- thing two files defined, and this was the copy nothing tested.
162
-
163
- Phase 3 (Dev) then iterates with `plan-todos.sh next "$TASK_ID"` until empty, calling `start` before each step and `complete` (with notes) or `fail`/`skip` after. Phase 4 (Review) reads the Todo list to verify all `completed` items map to diff hunks. Phase 7 (Report) renders `list` into the agent-log + PR body.
164
-
165
- **Why opt-in:** the existing `planning-output.schema.json` already drives Phase 3 dependency order - `plan.todos[]` is a richer surface (notes, durations, status transitions) but adds state writes per step. Off by default to keep the bare-bones flow unchanged; flip on for visibility into long features.
166
-
167
- #### Step 4.8 - Cross-artifact consistency check (required, before presenting the plan)
168
-
169
- Before the plan is rendered for approval (Step 5b) or silently accepted (the autopilot skip path), verify it against `state.analysis` (drifted plans are the root cause of "PR does not match the ticket"):
170
-
171
- 1. **Requirement coverage** - every analysis requirement (`touchedAreas[]` entry, Section 14 row, acceptance criterion) maps to at least one plan task.
172
- 2. **Anchor integrity** - No plan task without an analysis anchor (each task cites the `touchedAreas[].path`, Section 6 row, or `risks[]` mitigation it implements).
173
- 3. **Open-question carry-over** - every analysis open question lands in the plan (clarification item, risk row, or explicit descope note); none silently dropped.
174
-
175
- Progress line: ` → checking plan-vs-analysis consistency (3 checks)`
176
-
177
- **On mismatch:** revise the plan ONCE (re-run Steps 1-4 with the gap list quoted, re-run the validator gate and this checklist). If gaps remain, do NOT loop: surface them in the Step 5b render under a `⚠️ Consistency gaps` banner (autopilot: log `plan.consistency_gaps={list}`). Persist `state.phases["2"].consistencyCheck = { "unmappedRequirements": [], "unanchoredTasks": [], "droppedOpenQuestions": [], "revised": true|false }`.
178
-
179
- Log: "Phase 2: Consistency - requirements:{N/N mapped} anchors:{ok|M unanchored} open-questions:{carried|K dropped}"
180
-
181
- #### Step 5 - Plan Approval Gate (normal mode + autopilot safety)
182
-
183
- **Scope guard - skip this step entirely when BOTH of these hold:**
184
-
185
- - `state.autopilot === true` (autopilot contract: zero interaction)
186
- - Autopilot safety classifier returns `recommendPause: false` (see Step 5c below)
187
-
188
- OR:
189
-
190
- - `state.onlyDevelop === true` (Short pipeline: direct to Phase 3, no plan)
191
-
192
- In the skipped case, log `🧠 Phase 2: Plan - gate skipped ({mode}), proceeding to Phase 3` and go to Phase 3.
193
-
194
- ##### 5c - Autopilot safety classifier (runs before 5a/5b skip decision)
195
-
196
- **Only relevant when `state.autopilot === true` and `prefs.global.autopilotSafetyGate !== false`.** The classifier protects autopilot's zero-interaction contract from edge cases where silent execution is genuinely dangerous (security-path touch, schema migration, many-file sprawl, delete-without-paired-test).
197
-
198
- ```bash
199
- verdict=$(node $HOME/.claude/scripts/classify-plan-safety.mjs <(echo "$PLAN_JSON"))
200
- pause=$(jq -r '.recommendPause' <<< "$verdict")
201
- score=$(jq -r '.score' <<< "$verdict")
202
- reasons=$(jq -r '.reasons[] | "• \(.rule) (+\(.weight)): \(.detail)"' <<< "$verdict")
203
- ```
204
-
205
- | `recommendPause` | Action |
206
- |---|---|
207
- | `false` | Skip 5a/5b as before - autopilot proceeds to Phase 3. Log `🧠 Phase 2: Safety classifier - score {N}, autopilot proceeds`. |
208
- | `true` | Inject a one-time manual approval prompt even though we are in autopilot. Render the plan (5b shape) with the `reasons[]` list prepended. User sees `⚠️ Autopilot safety gate tripped (score {N})` banner and must explicitly choose Approve / Cancel (or edit via Other) in the 5b `AskUserQuestion` picker. Log `🧠 Phase 2: Safety classifier - score {N}, autopilot paused, reasons={rules}`. |
209
-
210
- **Why opt-out instead of opt-in:** the asymmetry favors pausing. A pause on a high-blast-radius plan costs seconds; a silent auto-merge of a bad one costs hours of rollback or a revert PR. Users running tightly-scoped batch workflows (e.g. figma component iteration over known-safe components) can set `prefs.global.autopilotSafetyGate = false` if they've validated the task class is safe.
211
-
212
- **Rules + weights** - see `classify-plan-safety.mjs` header comment for the canonical list. Summary: `file-count-high` (30) / `destructive-verb` (25) / `security-path` (35) / `delete-without-test` (30) / `schema-migration` (25) / `infrastructure` (20). Threshold: score ≥ 50 flips `recommendPause` true. Tuned so any single heavy signal or any two medium signals trigger the pause.
213
-
214
- **Telemetry:** emit `phase.plan.safety` OTel span (when `MULTI_AGENT_OTEL_SPANS=1`) with `{score, recommendPause, rules}` so post-hoc analysis can tune weights.
215
-
216
- Otherwise (normal mode), run the gate. The gate has **two modes** that chain: Clarification → Approval. Each round emits a progress line and persists to `agent-state.json.phases["2"]` for audit.
217
-
218
- ##### 5a - Clarification Mode (conditional, max 2 rounds)
219
-
220
- Trigger if the plan Fable produced in Step 1-4 carries ANY ambiguity signal from Phase 1 analysis:
221
-
222
- | Signal | Check |
223
- |---|---|
224
- | Vague acceptance | Jira/issue description < 200 chars AND no `## Acceptance Criteria` section |
225
- | UI work, no design | Task touches `*View.swift` / `*Screen.kt` but Phase 1 captured no Figma URL |
226
- | API work, no contract | Task touches network/repository layer but no endpoint/OpenAPI reference in Phase 1 |
227
- | Ambiguous language | Phase 1 analysis flagged `ambiguityScore >= 2` (e.g. "improve", "fix", "update" with no object) |
228
- | Parent-story scope drift | `state.relatedIssues[]` names a sibling overlapping this scope |
229
-
230
- If any signal trips, DO NOT render the plan yet. Render structured questions:
231
-
232
- ```
233
- 📋 Development Plan - {taskId}
234
- ────────────────────────────────────
235
- ⚠️ There are points that need clarifying before drafting a plan for this item:
236
-
237
- 1. {question 1}
238
- 2. {question 2}
239
- 3. {question 3}
240
-
241
- Write your answers, then I will draft the plan.
242
- (or tell me you want to pause - the task can be resumed later)
243
- ```
244
-
245
- Wait for user response. On reply:
246
-
247
- - **User asks to pause / cancel** (intent, in any language) → `state.status = "paused"`, log `🧠 Phase 2: Gate aborted (clarification)`, stop.
248
- - **Free-text answer** → append to `state.phases["2"].clarificationAnswers`, bump `clarificationRounds`, re-run Steps 1-4 with answers as context, re-evaluate signals.
249
-
250
- **Cap: 2 rounds.** If `clarificationRounds === 2` and signals still trip, do NOT ask again - render the plan anyway with a `⚠️ best-effort (item still unclear)` banner in Mode 5b, and let the user decide via approval/edit. This prevents infinite grooming loops on chronically under-specified tickets.
251
-
252
- Persist to `state.phases["2"]`:
253
-
254
- ```json
255
- {
256
- "clarificationRounds": 1,
257
- "clarificationQuestions": ["Pasaport scan sonrası...", "Backend endpoint...", "Android kapsamda mı?"],
258
- "clarificationAnswers": ["Silent reject. Toast yok.", "Backend hazır, v3.4 release'de.", "Sadece iOS bu task'ta."]
259
- }
260
- ```
261
-
262
- ##### 5b - Approval Loop (always runs after clarification resolves or is skipped)
263
-
264
- Render the plan:
265
-
266
- ```
267
- 📋 Development Plan - {taskId} {best-effort banner if 5a capped}
268
- ────────────────────────────────────
269
- Summary: {1-line summary}
270
- Approach: {1-paragraph approach}
271
- Risk: {low|medium|high} - {reason}
272
- Scope: {S|M|L} ({N} files, {K} new)
273
- Files to touch:
274
- - {path1}
275
- - {path2}
276
-
277
- Todos ({N} total):
278
- 1. {subject} - {approach}
279
- 2. {subject} - {approach} (blocked by #1)
280
- ...
281
-
282
- ```
283
-
284
- Then ask for the decision with a **native `AskUserQuestion` picker** (never a typed-keyword prompt):
285
-
286
- - `question`: "Do you approve this plan?" (rendered in `outputLanguage`)
287
- - `header`: "Plan" (English, <=12 chars)
288
- - `options` (label + description in `outputLanguage`; the names below are the option
289
- SEMANTICS, not strings to print):
290
- - option 1 - "Approve": proceed to development (Phase 3)
291
- - option 2 - "Cancel": pause the task; resume later with `:resume`
292
-
293
- The picker's built-in **Other** field is the free-text edit channel: the user types an edit request there (e.g. "also look at the auth service but keep LoginView out of scope") instead of typing a keyword. Per the `rules.md` matrix `question`, `label` and `description` all follow `outputLanguage`; only `header` stays English. Branch on which option was picked, never on its rendered text.
294
-
295
- Handle the selection:
296
-
297
- - **Option 1 (Approve)**:
298
- - Set `state.phases["2"].planApprovedAt = now()`, bump `planIterations` if not yet set (default 1)
299
- - Log `🧠 Phase 2: Plan approved (iterations={N}, clarificationRounds={M})`
300
- - Proceed to Phase 3
301
- - **Cancel**:
302
- - Set `state.status = "paused"`, log `🧠 Phase 2: Plan aborted by user`, stop
303
- - **Other (free-text edit request)**:
304
- - Treat the typed text as an edit request. Append to `state.phases["2"].planEditRequests`, bump `planIterations`
305
- - Pass the edit request + current plan to the planning model (Fable; Opus when the fallback ladder engages); it revises and returns a new plan (same schema, same validator)
306
- - Re-render the plan (5b), loop
307
-
308
- No hard cap on edit iterations - the user controls exit via the Approve / Cancel options. Between iterations, keep only the **latest plan** as canonical; previous renders are in the log for audit but do not re-enter the validator.
309
-
310
- **Validator**: every revised plan goes through `node $HOME/.claude/scripts/validate-planning.mjs -` before re-render. If validation fails after an edit, log `⚠️ Phase 2: Plan validator failed after edit request #N - retrying the planning model once` and retry once; on second failure surface the validator error to the user and go back to approval prompt with the pre-edit plan.
311
-
312
- #### Step 6 - Mode-specific short-circuit (reference)
313
-
314
- The pipeline shapes interact with the gate as follows. This table is the source of truth for the gate's mode-awareness - if behavior diverges in code, fix the code, not the table.
315
-
316
- | Mode | Clarification | Approval Loop | Safety Classifier | Notes |
317
- |---|---|---|---|---|
318
- | Full, interactive (`/multi-agent`, `:local`) | ✅ (max 2 rounds) | ✅ | - (redundant when a human approves) | Full gate |
319
- | Short (depth picker answered Short) | ❌ | ❌ | - | Phase 1-2 skipped at Step 7.5; Phase 3 starts immediately |
320
- | `autopilot`, `:local-autopilot` | ❌ | ❌ conditional | ✅ (if `autopilotSafetyGate !== false`) | Always Full, so a plan exists. Safe plans proceed silently; high-risk plans trigger a one-time manual approval. Log records the score. |
321
-
322
- **Why autopilot now has an escape hatch:** the old "zero interaction - fully trust the scope" contract held well for tightly-scoped batch workflows (figma component iteration over known-safe components) but broke in edge cases - schema migrations auto-merging, security-path drift going silent, delete-without-test sprawls. The safety classifier (Step 5c) is opt-out so the default protects against the edge-case cost; users running known-safe workflows can flip `prefs.global.autopilotSafetyGate = false` to restore pre-v7.0 behavior.
323
-
324
- #### Telemetry - token forwarding
325
-
326
- After plan generation (and after each edit-loop iteration), forward the planning model's call totals so Phase 7's Cost Breakdown captures Phase 2 (`model=` names the rung that actually ran: `fable`, or `opus` after a fallback step):
327
-
328
- ```bash
329
- LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 2 plan.generated \
330
- model=fable tokens_in=$IN tokens_out=$OUT duration_ms=$DUR iteration=$N
331
- ```
332
-
333
- Best-effort. See `$HOME/.claude/multi-agent-refs/progress-contract.md#token-telemetry-forwarding`.
334
-
335
- ---
336
-
337
- ## Token telemetry - invoke after every LLM call
338
-
339
- ```bash
340
- bash $HOME/.claude/scripts/phase-tracker.sh tokens 2 <input_count> <output_count>
341
- ```
342
-
343
- Contract and rationale: `progress-contract.md` -> Token telemetry forwarding.
344
-
@@ -1,182 +0,0 @@
1
- ### Phase 5: Test
2
-
3
- > **TLDR** - Optional test gate. Offers to boot the simulator/emulator (UI Bug Hunter) or hand off to the user for manual QA. Needs an interactive prompt AND a worktree checkout, so it is in the phase set of `/multi-agent` alone and dropped by every `autopilot` or `--local` entry. Depth does not affect it: a Short run still reaches Phase 5. If issues found, loops back to Phase 3.
4
-
5
- <!-- progress-contract: applied -->
6
- Progress emission per `$HOME/.claude/multi-agent-refs/progress-contract.md` - lines for local-test prompt render, user-answer capture, repo checkout (if selected).
7
-
8
- #### Step 0 - Test Gap Report (advisory)
9
-
10
- `state.testPolicy: none` → skip the gap scan (the gap IS the recorded policy) and run only pre-existing test targets; none → recorded no-op. Otherwise, before the local-checkout prompt, run the static test-gap detector. Heuristic, deterministic, no LLM, sub-second. The report ends up in `agent-log.md` under "Test Scenarios" and surfaces public symbols added in this branch that have no paired test.
11
-
12
- ```bash
13
- STACK=$(jq -r '.analysis.stack.primary // "unknown"' "$STATE_FILE")
14
- case "$STACK" in
15
- ios|swift) SCAN_STACK=ios ;;
16
- android|kotlin) SCAN_STACK=android ;;
17
- python) SCAN_STACK=python ;;
18
- node|typescript|js) SCAN_STACK=node ;;
19
- *) SCAN_STACK="" ;;
20
- esac
21
- if [ -n "$SCAN_STACK" ] && [ "${prefs_testGap_enabled:-true}" = "true" ]; then
22
- GAP_FLAGS=""
23
- [ "${prefs_testGap_scanTree:-false}" = "true" ] && GAP_FLAGS="$GAP_FLAGS --scan-tree"
24
- [ "${prefs_testGap_promoteSeverity:-false}" = "true" ] && GAP_FLAGS="$GAP_FLAGS --severity-promote"
25
- GAP_JSON=$(node $HOME/.claude/scripts/test-gap-scan.mjs \
26
- --base "$BASE_BRANCH" --head HEAD --stack "$SCAN_STACK" $GAP_FLAGS 2>/dev/null)
27
- echo "$GAP_JSON" | node $HOME/.claude/scripts/validate-test-gap.mjs - >/dev/null 2>&1 || GAP_JSON=""
28
- fi
29
- ```
30
-
31
- **What the report contains** (per `$HOME/.claude/schemas/test-gap.schema.json`):
32
-
33
- | Field | Meaning |
34
- |---|---|
35
- | `gaps[].sourcePath` | source file with the unprotected symbol |
36
- | `gaps[].symbol` | symbol name |
37
- | `gaps[].kind` | rule id (e.g. `public_func`, `composable_fun`, `named_export`) |
38
- | `gaps[].severity` | `blocking` / `important` / `suggestion` (see severity table below) |
39
- | `gaps[].expectedTestPaths` | likely test paths the user should land at, in priority order |
40
- | `gaps[].hint` | stack-specific testing reminder (e.g. swiftui-qa.md 3-layer) |
41
-
42
- **Severity defaults**:
43
-
44
- | Symbol kind | Severity |
45
- |---|---|
46
- | `composable_fun`, `view_struct`, `config_struct`, `interface`, `objc_export`, `public_proto` | important |
47
- | Other public API additions (`public_func`, `open_fun`, `named_export`, `default_export`, ...) | suggestion |
48
-
49
- **Gating** (opt-in): if `prefs.testGap.blockingThreshold` is set and `gapBySeverity.important + gapBySeverity.blocking` exceeds it, Phase 5 surfaces the report as a Phase 4 rework finding and loops back. **Default off** - gaps render as advisory under "Test Gap Report" only.
50
-
51
- **Telemetry**:
52
-
53
- ```bash
54
- LOG_METRIC_FORWARD_TO_TRACKER=0 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 5 test_gap.scanned \
55
- stack=$SCAN_STACK \
56
- sources=$(jq '.totals.sourcesScanned' <<< "$GAP_JSON") \
57
- gaps=$(jq '.totals.gapCount' <<< "$GAP_JSON")
58
- ```
59
-
60
- (No tracker forwarding - the scanner has no token cost.)
61
-
62
- **Figma reference panel (when `state.evidence.figma[]` is non-empty).** Before the local-checkout prompt, print a single block listing each captured frame so the user has a side-by-side reference during manual test:
63
-
64
- ```
65
- Figma evidence (tier=<n>):
66
- <fileKey>:<nodeId> <canonicalComponentName>
67
- screenshot: <screenshotUrl or local path>
68
- <fileKey>:<nodeId> <canonicalComponentName>
69
- screenshot: <screenshotUrl or local path>
70
- ```
71
-
72
- Tier 1 / Tier 2 records print `screenshotUrl` from the captured evidence (Tier 2 URLs expire after 30 days, re-fetch on the spot if needed). Tier 3 records print the local path to the user-attached screenshot. The block is informational; it never blocks the prompt.
73
-
74
- 1. Ask with a native `AskUserQuestion` picker (never a typed y/N prompt). The options MUST make the local-checkout side effect explicit - testing removes the worktree and checks the branch out into the main repo:
75
- - `question`: "Check out locally to test now?" (rendered in `outputLanguage`)
76
- - `header`: "Test" (English, <=12 chars)
77
- - `options`:
78
- - `{ label: "Test now", description: "Removes the worktree and checks the branch out into the main repo for Xcode / manual test" }`
79
- - `{ label: "Skip", description: "Stay in the worktree and go to Phase 6" }`
80
- - **Skip** → set `state.phases["5"].status = "skipped"` (so Phase 6 can offer the local-checkout prompt) → Phase 6
81
- - **Test now** → set `state.phases["5"].status = "in_progress"` → continue:
82
- 2. **Commit changes in worktree BEFORE removing** (WIP commit to preserve work):
83
- ```
84
- git -C {worktree-path} add -A
85
- git -C {worktree-path} commit -m "WIP: {jiraId} - changes for user test"
86
- ```
87
- 3. Remove worktree, checkout branch in main repo:
88
- ```
89
- git worktree remove .worktrees/{jiraId} --force
90
- git checkout {branch-name}
91
- ```
92
- Branch now has the WIP commit - all changes are preserved.
93
- 4. Show test instructions:
94
- ```
95
- Switched to branch: {branch-name}
96
- To test: Xcode -> build -> run -> manual test
97
- "ok" -> proceeds to Phase 6 (WIP commit will be replaced via git reset HEAD~1 + proper commit)
98
- "fix: ..." -> worktree is recreated, returns to Phase 3
99
- ```
100
- 5. Mark the tracker as waiting, then wait for the user response:
101
- ```bash
102
- bash $HOME/.claude/scripts/phase-tracker.sh now 5 "awaiting local test (user)"
103
- bash $HOME/.claude/scripts/phase-tracker.sh render
104
- ```
105
- The waiting state persists in `tracker-state.json` across the handoff; `/multi-agent:resume-local` and `/multi-agent:manual-test` CONTINUE this state file and never re-init it (`$HOME/.claude/multi-agent-refs/tracker-contract.md` "Continuation runs").
106
-
107
- **"ok" is a structured result, not a word.** Before "ok" is accepted, the run writes `$WORKTREE/.pipeline/manual-test.json`: one entry per acceptance criterion, the criteria taken from the analysis doc test plan (Section 15 / 20), the plan tasks, and the user's own words in the reply. Every criterion records what was seen; a criterion that was not tried says so with a reason.
108
- ```json
109
- {"criteria":[{"spec":"<quote>","source":"analysis 15.2 | plan task 3 | user","observed":"<what was seen>","verdict":"pass|fail|not-tested","reason":"<required when not-tested>","screenshot":"<path or null>"}],"verdict":"passed|failed"}
110
- ```
111
- Then gate it:
112
- ```bash
113
- node $HOME/.claude/scripts/evidence-gate.mjs --claim manual --status passed --evidence "$WORKTREE/.pipeline/manual-test.json"
114
- ```
115
- Add `--require-screenshot` when `state.visualEvidence.required` is true: a passing criterion then has to name a file that exists, because a `screenshot` key pointing nowhere is not evidence.
116
-
117
- Exit 1 means the "ok" is not accepted: tell the user which criterion is missing evidence (a `fail` verdict, or `not-tested` without a reason) and wait for the next reply. Exit 0 marks Phase 5 completed with `Result "local test passed (user)"`. The "fix: ..." path below is unchanged.
118
- 6. If fix needed:
119
- - Branch already has WIP commit (from step 2) - changes are safe
120
- - **Heal stale admin state first** (same contract as Phase 0 - step 3's
121
- `worktree remove` or an interrupted run can leave a stale entry, so a bare
122
- re-add fails with `already exists`/`already registered`):
123
- ```bash
124
- git -C "$PROJECT_ROOT" worktree prune 2>/dev/null || true
125
- if git -C "$PROJECT_ROOT" worktree list --porcelain | grep -qF "{worktree-path}"; then
126
- git -C "$PROJECT_ROOT" worktree unlock "{worktree-path}" 2>/dev/null || true
127
- fi
128
- ```
129
- Phase 0's repo residue guard is already in `.git/info/exclude` - no re-add.
130
- - Recreate worktree from branch: `git -C $PROJECT_ROOT worktree add {worktree-path} {branch}`
131
- - Re-set git identity: `git -C {worktree-path} config user.name/email` (from state)
132
- - Go back to Phase 3
133
- 7. Log: "Phase 5: Test - {result}"
134
-
135
- **CRITICAL**: Never remove a worktree with uncommitted changes. Always WIP commit first.
136
-
137
- #### Automated Device Checks (on-demand)
138
-
139
- Before or during user testing, run device-level audits via Bash if user requests. See `audit-guide.md` for commands.
140
-
141
- | Check | When | Command |
142
- | ------------------- | ----------------- | ----------------------------------------- |
143
- | UI flow video | `state.visualEvidence.required` AND Phase 3 recorded none | `capture-evidence.sh video start` -> drive the flow -> `video stop` -> `fit`. See below |
144
- | Accessibility audit | UI changes | `mcp__multi-agent-toolkit__{ios,android}_accessibility_audit` |
145
- | Biometric test | Auth flow changes | ios: `mcp__multi-agent-toolkit__ios_biometric` (android: manual) |
146
- | Launch time | Perf-sensitive changes | ios: app-launch instrument · android: `mcp__multi-agent-toolkit__android_launch_time` |
147
- | Visual test | Any UI changes | `/multi-agent test` (sim-test, both platforms) |
148
- | Snapshot regression | Component / pixel-stable UI changes | ios: `mcp__multi-agent-toolkit__ios_visual_diff` · android: `mcp__multi-agent-toolkit__android_screenshot` + compare |
149
- | Store screenshots | `taskType === screenshot` | ios: `ios_status_bar({preset: "clean"})` · android: `android_screenshot` |
150
-
151
- ##### UI flow video, when Phase 3 produced none
152
-
153
- Phase 3 Step 3.55 is the primary host and runs in every mode. Phase 5 is the richer one where it exists: the device is up and a person is watching, so the flow is one somebody confirmed. It adds, never replaces.
154
-
155
- Run only when `visualEvidence.required` and `visualEvidence.video.file` is empty: re-check the device with `$HOME/.claude/scripts/probe-evidence-capability.sh --only device`, then `$HOME/.claude/scripts/capture-evidence.sh video start` -> drive the flow (`run-ui-tests.sh run`, or the Phase 5 scenarios by hand) -> `video stop` -> `fit`. The cap is `visualEvidence.maxVideoSeconds`, read through `capture-evidence.sh limits` so one value serves both. Update `videoTier` and `videoTierReason` with what ran.
156
-
157
- When the intake answer was `unit` there is no recording here either: overriding it in a phase the user may not be watching makes the question decorative.
158
-
159
- Results included in Phase 7 report. MCP tools preferred when available - concise structured output, lower token cost.
160
-
161
- **Snapshot regression flow (optional):** when the task changes a stable component, capture a screenshot before the change (baseline) and after (current), then call `ios_visual_diff({baseline, current, max_diff_pct: 1.0})`. Threshold can be relaxed for animated / non-deterministic regions - keep `max_diff_pct ≤ 1.0` for static layouts.
162
-
163
- #### Security Audit (store-readiness)
164
-
165
- When the task touches authentication, keychain, network, or is scheduled for an imminent release, launch the `security-auditor` subagent to run an OWASP Mobile Top 10 pass plus App Store / Play Store compliance checks:
166
-
167
- ```
168
- Agent(subagent_type: "security-auditor", prompt: "<diff + context>")
169
- ```
170
-
171
- Returns severity-tagged findings (Critical / High / Medium). Critical items block Phase 6 just like Phase 4 blockers; High items are logged and surfaced in Phase 7 report. Skipped by default - opt-in for release branches or on explicit `/multi-agent "<task>" --audit` flag.
172
-
173
- #### Telemetry - token forwarding
174
-
175
- When the security-auditor or any other Phase 5 sub-agent runs, forward its token totals so Phase 7's Cost Breakdown captures Phase 5:
176
-
177
- ```bash
178
- LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 5 audit.completed \
179
- model=opus tokens_in=$IN tokens_out=$OUT duration_ms=$DUR
180
- ```
181
-
182
- If Phase 5 is purely user-driven (no sub-agent ran), no token forwarding is required and the cost block stays empty for this phase. Best-effort. See `$HOME/.claude/multi-agent-refs/progress-contract.md#token-telemetry-forwarding`.