@mmerterden/multi-agent-pipeline 18.0.0 → 19.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (234) hide show
  1. package/CHANGELOG.md +287 -0
  2. package/README.md +36 -20
  3. package/README.tr.md +14 -16
  4. package/docs/adr/0002-instruction-driven-flag.md +1 -0
  5. package/docs/adr/0005-lazy-phase-docs.md +11 -1
  6. package/docs/adr/0008-installer-modularization-and-secret-leak-defense.md +1 -0
  7. package/docs/adr/0010-own-code-graph.md +1 -0
  8. package/docs/adr/0014-six-phase-consolidation.md +134 -0
  9. package/docs/adr/README.md +2 -1
  10. package/docs/architecture.md +37 -38
  11. package/docs/best-practices.md +1 -1
  12. package/docs/ecosystem.md +46 -27
  13. package/docs/engineering.md +1 -1
  14. package/docs/facts.json +61 -0
  15. package/docs/features.md +55 -54
  16. package/docs/performance.md +5 -5
  17. package/docs/recovery-guide.md +17 -17
  18. package/docs/token-budget-history.md +3 -1
  19. package/index.js +2 -2
  20. package/install/_codex-agents.mjs +1 -1
  21. package/install/templates/claude-hooks.json +1 -1
  22. package/install/templates/codex-instructions.md +1 -1
  23. package/install/templates/copilot-instructions.md +28 -28
  24. package/manifest.json +234 -216
  25. package/package.json +2 -2
  26. package/pipeline/agents/dev-critic.md +7 -7
  27. package/pipeline/commands/figma-to-swiftui.md +1 -1
  28. package/pipeline/commands/multi-agent/SKILL.md +9 -9
  29. package/pipeline/commands/multi-agent/analysis/SKILL.md +15 -15
  30. package/pipeline/commands/multi-agent/analysis-resolve/SKILL.md +2 -2
  31. package/pipeline/commands/multi-agent/autopilot/SKILL.md +7 -7
  32. package/pipeline/commands/multi-agent/channels/SKILL.md +15 -15
  33. package/pipeline/commands/multi-agent/diff-explain/SKILL.md +6 -6
  34. package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +1 -1
  35. package/pipeline/commands/multi-agent/graph/SKILL.md +1 -1
  36. package/pipeline/commands/multi-agent/help/SKILL.md +62 -62
  37. package/pipeline/commands/multi-agent/language/SKILL.md +2 -2
  38. package/pipeline/commands/multi-agent/local/SKILL.md +11 -11
  39. package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +13 -13
  40. package/pipeline/commands/multi-agent/log/SKILL.md +2 -2
  41. package/pipeline/commands/multi-agent/manual-test/SKILL.md +9 -9
  42. package/pipeline/commands/multi-agent/model/SKILL.md +69 -0
  43. package/pipeline/commands/multi-agent/refactor/SKILL.md +3 -3
  44. package/pipeline/commands/multi-agent/resume/SKILL.md +4 -4
  45. package/pipeline/commands/multi-agent/resume-local/SKILL.md +19 -17
  46. package/pipeline/commands/multi-agent/review/SKILL.md +2 -2
  47. package/pipeline/commands/multi-agent/review-analysis/SKILL.md +1 -1
  48. package/pipeline/commands/multi-agent/route-off/SKILL.md +36 -0
  49. package/pipeline/commands/multi-agent/route-on/SKILL.md +74 -0
  50. package/pipeline/commands/multi-agent/route-status/SKILL.md +56 -0
  51. package/pipeline/commands/multi-agent/setup/SKILL.md +2 -2
  52. package/pipeline/commands/multi-agent/status/SKILL.md +5 -5
  53. package/pipeline/commands/multi-agent/steer/SKILL.md +2 -2
  54. package/pipeline/commands/multi-agent/sync/SKILL.md +12 -13
  55. package/pipeline/commands/multi-agent/test/SKILL.md +1 -1
  56. package/pipeline/lib/credential-inventory.sh +1 -1
  57. package/pipeline/lib/fetch-fortify.sh +1 -1
  58. package/pipeline/lib/model-dispatch.sh +140 -0
  59. package/pipeline/lib/model-rung.sh +142 -0
  60. package/pipeline/lib/outbound-gate.mjs +14 -0
  61. package/pipeline/lib/phase-schema.mjs +88 -0
  62. package/pipeline/lib/plan-todos.sh +5 -5
  63. package/pipeline/lib/route-state.sh +161 -0
  64. package/pipeline/lib/run-paths.sh +2 -2
  65. package/pipeline/multi-agent-refs/_account-picker.md +1 -1
  66. package/pipeline/multi-agent-refs/_dev-context.md +6 -6
  67. package/pipeline/multi-agent-refs/_input-parser.md +1 -1
  68. package/pipeline/multi-agent-refs/analysis/evidence.md +2 -11
  69. package/pipeline/multi-agent-refs/analysis/intake.md +7 -7
  70. package/pipeline/multi-agent-refs/analysis/locked.md +48 -22
  71. package/pipeline/multi-agent-refs/analysis/redesign.md +1 -1
  72. package/pipeline/multi-agent-refs/analysis/render.md +10 -10
  73. package/pipeline/multi-agent-refs/analysis/resolve.md +1 -1
  74. package/pipeline/multi-agent-refs/analysis/review.md +2 -2
  75. package/pipeline/multi-agent-refs/analysis/synthesis.md +13 -7
  76. package/pipeline/multi-agent-refs/analysis-template-corporate.md +9 -9
  77. package/pipeline/multi-agent-refs/analysis-template.md +19 -19
  78. package/pipeline/multi-agent-refs/android-guide.md +1 -1
  79. package/pipeline/multi-agent-refs/audit-guide.md +13 -13
  80. package/pipeline/multi-agent-refs/channels/issue-comment.md +2 -2
  81. package/pipeline/multi-agent-refs/channels/jira.md +3 -3
  82. package/pipeline/multi-agent-refs/channels/pr.md +4 -4
  83. package/pipeline/multi-agent-refs/channels/wiki.md +1 -1
  84. package/pipeline/multi-agent-refs/component-dispatch.md +8 -8
  85. package/pipeline/multi-agent-refs/conventions-defaults.md +2 -2
  86. package/pipeline/multi-agent-refs/cross-cli-contract.md +31 -6
  87. package/pipeline/multi-agent-refs/features/analysis-jira.md +1 -1
  88. package/pipeline/multi-agent-refs/features/autopilot-circuit-breaker.md +4 -4
  89. package/pipeline/multi-agent-refs/features/code-graph.md +5 -5
  90. package/pipeline/multi-agent-refs/features/design-conformance.md +1 -1
  91. package/pipeline/multi-agent-refs/features/dev-critic.md +3 -3
  92. package/pipeline/multi-agent-refs/features/doctor.md +3 -3
  93. package/pipeline/multi-agent-refs/features/external-context-injection.md +3 -3
  94. package/pipeline/multi-agent-refs/features/maturity-followup.md +3 -3
  95. package/pipeline/multi-agent-refs/features/model-fallback.md +41 -5
  96. package/pipeline/multi-agent-refs/features/plan-todos.md +1 -1
  97. package/pipeline/multi-agent-refs/features/repo-map.md +1 -1
  98. package/pipeline/multi-agent-refs/features/review-delta.md +3 -3
  99. package/pipeline/multi-agent-refs/features/review-multi-repo.md +2 -2
  100. package/pipeline/multi-agent-refs/features/scope-check.md +4 -4
  101. package/pipeline/multi-agent-refs/features/skill-conformance.md +2 -2
  102. package/pipeline/multi-agent-refs/features/stack-skill-routing.md +1 -1
  103. package/pipeline/multi-agent-refs/features/url-enrichment.md +1 -1
  104. package/pipeline/multi-agent-refs/features/verify-by-test.md +4 -4
  105. package/pipeline/multi-agent-refs/features/visual-evidence.md +19 -19
  106. package/pipeline/multi-agent-refs/features/worktree-finalize.md +6 -6
  107. package/pipeline/multi-agent-refs/issue-jira-triad.md +10 -10
  108. package/pipeline/multi-agent-refs/knowledge.md +11 -11
  109. package/pipeline/multi-agent-refs/multi-repo-integration-build.md +13 -13
  110. package/pipeline/multi-agent-refs/payload-contracts.md +8 -8
  111. package/pipeline/multi-agent-refs/phases/log-format.md +10 -10
  112. package/pipeline/multi-agent-refs/phases/modes.md +30 -30
  113. package/pipeline/multi-agent-refs/phases/operations.md +8 -8
  114. package/pipeline/multi-agent-refs/phases/phase-0-init.md +24 -24
  115. package/pipeline/multi-agent-refs/phases/phase-1-plan.md +599 -0
  116. package/pipeline/multi-agent-refs/phases/{phase-3-dev.md → phase-2-dev.md} +129 -49
  117. package/pipeline/multi-agent-refs/phases/{phase-4-review.md → phase-3-review.md} +225 -107
  118. package/pipeline/multi-agent-refs/phases/{phase-6-commit.md → phase-4-commit.md} +23 -23
  119. package/pipeline/multi-agent-refs/phases/{phase-7-report.md → phase-5-report.md} +29 -29
  120. package/pipeline/multi-agent-refs/phases.md +44 -48
  121. package/pipeline/multi-agent-refs/picker-contract.md +1 -1
  122. package/pipeline/multi-agent-refs/progress-contract.md +6 -6
  123. package/pipeline/multi-agent-refs/readiness-review.md +1 -1
  124. package/pipeline/multi-agent-refs/rules.md +7 -7
  125. package/pipeline/multi-agent-refs/swiftui-guide.md +2 -2
  126. package/pipeline/multi-agent-refs/tracker-contract.md +31 -32
  127. package/pipeline/multi-agent-refs/wiki-capture.md +14 -14
  128. package/pipeline/preferences-template.json +9 -1
  129. package/pipeline/rules/figma-pipeline.md +8 -8
  130. package/pipeline/rules/outside-the-pipeline.md +1 -1
  131. package/pipeline/schemas/agent-state.schema.json +50 -50
  132. package/pipeline/schemas/analysis-output.schema.json +3 -3
  133. package/pipeline/schemas/analysis-spec.schema.json +2 -2
  134. package/pipeline/schemas/autopilot-config.schema.json +1 -1
  135. package/pipeline/schemas/code-graph.schema.json +1 -1
  136. package/pipeline/schemas/criteria-manifest.schema.json +1 -1
  137. package/pipeline/schemas/dev-critic-output.schema.json +1 -1
  138. package/pipeline/schemas/diff-risk.schema.json +1 -1
  139. package/pipeline/schemas/figma-project-config.schema.json +1 -1
  140. package/pipeline/schemas/migrations/prefs-2.4.0-to-2.5.0.mjs +2 -2
  141. package/pipeline/schemas/migrations/prefs-2.6.0-to-2.7.0.mjs +31 -0
  142. package/pipeline/schemas/migrations/state-2.1.0-to-2.2.0.mjs +129 -0
  143. package/pipeline/schemas/phases.json +105 -0
  144. package/pipeline/schemas/plan-todos.schema.json +5 -5
  145. package/pipeline/schemas/planning-output.schema.json +1 -1
  146. package/pipeline/schemas/prefs.schema.json +102 -58
  147. package/pipeline/schemas/reviewer-output.schema.json +3 -3
  148. package/pipeline/schemas/route-config.schema.json +74 -0
  149. package/pipeline/schemas/scope-check.schema.json +1 -1
  150. package/pipeline/schemas/secret-patterns.json +124 -0
  151. package/pipeline/schemas/test-gap.schema.json +1 -1
  152. package/pipeline/schemas/token-budget.json +12 -18
  153. package/pipeline/schemas/triage-output.schema.json +6 -6
  154. package/pipeline/scripts/README.md +3 -3
  155. package/pipeline/scripts/_code-graph.mjs +2 -2
  156. package/pipeline/scripts/_run-paths.mjs +2 -2
  157. package/pipeline/scripts/_smoke-root.sh +1 -1
  158. package/pipeline/scripts/aggregate-metrics.mjs +1 -1
  159. package/pipeline/scripts/build-references.mjs +2 -2
  160. package/pipeline/scripts/bulk-read.sh +10 -1
  161. package/pipeline/scripts/capture-flush.sh +8 -8
  162. package/pipeline/scripts/capture-resume.sh +3 -3
  163. package/pipeline/scripts/classify-plan-safety.mjs +1 -1
  164. package/pipeline/scripts/cost-table.json +8 -1
  165. package/pipeline/scripts/diff-explain.mjs +1 -1
  166. package/pipeline/scripts/doctor.mjs +3 -3
  167. package/pipeline/scripts/gc-abandoned.sh +3 -3
  168. package/pipeline/scripts/gc-tmp.sh +1 -1
  169. package/pipeline/scripts/gc-worktrees.sh +1 -1
  170. package/pipeline/scripts/gen-facts.mjs +280 -0
  171. package/pipeline/scripts/gen-mode-dispatch.mjs +32 -37
  172. package/pipeline/scripts/gen-ref-toc.mjs +1 -1
  173. package/pipeline/scripts/graph-report.mjs +1 -1
  174. package/pipeline/scripts/jira-attach.sh +1 -1
  175. package/pipeline/scripts/learn-from-transcripts.mjs +1 -1
  176. package/pipeline/scripts/learning-curve.mjs +2 -2
  177. package/pipeline/scripts/log-metric.sh +17 -4
  178. package/pipeline/scripts/memory-save.sh +1 -1
  179. package/pipeline/scripts/migrate-prefs.mjs +22 -5
  180. package/pipeline/scripts/phase-banner.sh +20 -20
  181. package/pipeline/scripts/phase-tracker.sh +12 -12
  182. package/pipeline/scripts/plan-coverage-gate.mjs +2 -2
  183. package/pipeline/scripts/pre-commit-check.sh +30 -1
  184. package/pipeline/scripts/render-agent-log-cost.sh +1 -1
  185. package/pipeline/scripts/render-work-summary.sh +3 -3
  186. package/pipeline/scripts/review-file-filter.mjs +1 -1
  187. package/pipeline/scripts/run-aggregator.mjs +13 -6
  188. package/pipeline/scripts/run-metrics.mjs +1 -1
  189. package/pipeline/scripts/runs-index.mjs +11 -1
  190. package/pipeline/scripts/scan-skills.sh +26 -0
  191. package/pipeline/scripts/smoke-cross-cli-behavior.sh +6 -6
  192. package/pipeline/scripts/smoke-schema-validation.sh +26 -7
  193. package/pipeline/scripts/token-budget-report.mjs +13 -2
  194. package/pipeline/scripts/triage-memory.mjs +2 -2
  195. package/pipeline/scripts/validate-analysis-doc.mjs +274 -43
  196. package/pipeline/scripts/validate-planning.mjs +1 -1
  197. package/pipeline/scripts/validate-reviewer.mjs +1 -1
  198. package/pipeline/scripts/validate-state.mjs +45 -5
  199. package/pipeline/scripts/validate-triage.mjs +3 -3
  200. package/pipeline/scripts/verify-citations.mjs +1 -1
  201. package/pipeline/scripts/worktree-finalize.sh +5 -5
  202. package/pipeline/scripts/write-state.mjs +32 -0
  203. package/pipeline/skills/.skill-manifest.json +38 -22
  204. package/pipeline/skills/.skills-index.json +49 -5
  205. package/pipeline/skills/shared/README.md +10 -6
  206. package/pipeline/skills/shared/core/apple-archive-compliance/SKILL.md +8 -8
  207. package/pipeline/skills/shared/core/google-play-compliance/SKILL.md +8 -8
  208. package/pipeline/skills/shared/core/multi-agent/SKILL.md +81 -82
  209. package/pipeline/skills/shared/core/multi-agent-autopilot/SKILL.md +3 -3
  210. package/pipeline/skills/shared/core/multi-agent-channels/SKILL.md +14 -14
  211. package/pipeline/skills/shared/core/multi-agent-diff-explain/SKILL.md +5 -5
  212. package/pipeline/skills/shared/core/multi-agent-graph/SKILL.md +1 -1
  213. package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +25 -23
  214. package/pipeline/skills/shared/core/multi-agent-language/SKILL.md +2 -2
  215. package/pipeline/skills/shared/core/multi-agent-local/SKILL.md +2 -2
  216. package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +8 -8
  217. package/pipeline/skills/shared/core/multi-agent-manual-test/SKILL.md +6 -6
  218. package/pipeline/skills/shared/core/multi-agent-model/SKILL.md +71 -0
  219. package/pipeline/skills/shared/core/multi-agent-refactor/SKILL.md +3 -3
  220. package/pipeline/skills/shared/core/multi-agent-resume/SKILL.md +1 -1
  221. package/pipeline/skills/shared/core/multi-agent-resume-local/SKILL.md +7 -7
  222. package/pipeline/skills/shared/core/multi-agent-route-off/SKILL.md +39 -0
  223. package/pipeline/skills/shared/core/multi-agent-route-on/SKILL.md +76 -0
  224. package/pipeline/skills/shared/core/multi-agent-route-status/SKILL.md +59 -0
  225. package/pipeline/skills/shared/core/multi-agent-setup/SKILL.md +1 -1
  226. package/pipeline/skills/shared/core/multi-agent-status/SKILL.md +5 -5
  227. package/pipeline/skills/shared/core/multi-agent-steer/SKILL.md +2 -2
  228. package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +6 -5
  229. package/pipeline/skills/shared/external/NOTICE-swift-ios-skills.md +1 -1
  230. package/pipeline/skills/shared/external/signal-community/SKILL.md +8 -1
  231. package/pipeline/skills/skills-index.md +8 -4
  232. package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +0 -263
  233. package/pipeline/multi-agent-refs/phases/phase-2-planning.md +0 -344
  234. package/pipeline/multi-agent-refs/phases/phase-5-test.md +0 -182
@@ -1,182 +0,0 @@
1
- ### Phase 5: Test
2
-
3
- > **TLDR** - Optional test gate. Offers to boot the simulator/emulator (UI Bug Hunter) or hand off to the user for manual QA. Needs an interactive prompt AND a worktree checkout, so it is in the phase set of `/multi-agent` alone and dropped by every `autopilot` or `--local` entry. Depth does not affect it: a Short run still reaches Phase 5. If issues found, loops back to Phase 3.
4
-
5
- <!-- progress-contract: applied -->
6
- Progress emission per `$HOME/.claude/multi-agent-refs/progress-contract.md` - lines for local-test prompt render, user-answer capture, repo checkout (if selected).
7
-
8
- #### Step 0 - Test Gap Report (advisory)
9
-
10
- `state.testPolicy: none` → skip the gap scan (the gap IS the recorded policy) and run only pre-existing test targets; none → recorded no-op. Otherwise, before the local-checkout prompt, run the static test-gap detector. Heuristic, deterministic, no LLM, sub-second. The report ends up in `agent-log.md` under "Test Scenarios" and surfaces public symbols added in this branch that have no paired test.
11
-
12
- ```bash
13
- STACK=$(jq -r '.analysis.stack.primary // "unknown"' "$STATE_FILE")
14
- case "$STACK" in
15
- ios|swift) SCAN_STACK=ios ;;
16
- android|kotlin) SCAN_STACK=android ;;
17
- python) SCAN_STACK=python ;;
18
- node|typescript|js) SCAN_STACK=node ;;
19
- *) SCAN_STACK="" ;;
20
- esac
21
- if [ -n "$SCAN_STACK" ] && [ "${prefs_testGap_enabled:-true}" = "true" ]; then
22
- GAP_FLAGS=""
23
- [ "${prefs_testGap_scanTree:-false}" = "true" ] && GAP_FLAGS="$GAP_FLAGS --scan-tree"
24
- [ "${prefs_testGap_promoteSeverity:-false}" = "true" ] && GAP_FLAGS="$GAP_FLAGS --severity-promote"
25
- GAP_JSON=$(node $HOME/.claude/scripts/test-gap-scan.mjs \
26
- --base "$BASE_BRANCH" --head HEAD --stack "$SCAN_STACK" $GAP_FLAGS 2>/dev/null)
27
- echo "$GAP_JSON" | node $HOME/.claude/scripts/validate-test-gap.mjs - >/dev/null 2>&1 || GAP_JSON=""
28
- fi
29
- ```
30
-
31
- **What the report contains** (per `$HOME/.claude/schemas/test-gap.schema.json`):
32
-
33
- | Field | Meaning |
34
- |---|---|
35
- | `gaps[].sourcePath` | source file with the unprotected symbol |
36
- | `gaps[].symbol` | symbol name |
37
- | `gaps[].kind` | rule id (e.g. `public_func`, `composable_fun`, `named_export`) |
38
- | `gaps[].severity` | `blocking` / `important` / `suggestion` (see severity table below) |
39
- | `gaps[].expectedTestPaths` | likely test paths the user should land at, in priority order |
40
- | `gaps[].hint` | stack-specific testing reminder (e.g. swiftui-qa.md 3-layer) |
41
-
42
- **Severity defaults**:
43
-
44
- | Symbol kind | Severity |
45
- |---|---|
46
- | `composable_fun`, `view_struct`, `config_struct`, `interface`, `objc_export`, `public_proto` | important |
47
- | Other public API additions (`public_func`, `open_fun`, `named_export`, `default_export`, ...) | suggestion |
48
-
49
- **Gating** (opt-in): if `prefs.testGap.blockingThreshold` is set and `gapBySeverity.important + gapBySeverity.blocking` exceeds it, Phase 5 surfaces the report as a Phase 4 rework finding and loops back. **Default off** - gaps render as advisory under "Test Gap Report" only.
50
-
51
- **Telemetry**:
52
-
53
- ```bash
54
- LOG_METRIC_FORWARD_TO_TRACKER=0 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 5 test_gap.scanned \
55
- stack=$SCAN_STACK \
56
- sources=$(jq '.totals.sourcesScanned' <<< "$GAP_JSON") \
57
- gaps=$(jq '.totals.gapCount' <<< "$GAP_JSON")
58
- ```
59
-
60
- (No tracker forwarding - the scanner has no token cost.)
61
-
62
- **Figma reference panel (when `state.evidence.figma[]` is non-empty).** Before the local-checkout prompt, print a single block listing each captured frame so the user has a side-by-side reference during manual test:
63
-
64
- ```
65
- Figma evidence (tier=<n>):
66
- <fileKey>:<nodeId> <canonicalComponentName>
67
- screenshot: <screenshotUrl or local path>
68
- <fileKey>:<nodeId> <canonicalComponentName>
69
- screenshot: <screenshotUrl or local path>
70
- ```
71
-
72
- Tier 1 / Tier 2 records print `screenshotUrl` from the captured evidence (Tier 2 URLs expire after 30 days, re-fetch on the spot if needed). Tier 3 records print the local path to the user-attached screenshot. The block is informational; it never blocks the prompt.
73
-
74
- 1. Ask with a native `AskUserQuestion` picker (never a typed y/N prompt). The options MUST make the local-checkout side effect explicit - testing removes the worktree and checks the branch out into the main repo:
75
- - `question`: "Check out locally to test now?" (rendered in `outputLanguage`)
76
- - `header`: "Test" (English, <=12 chars)
77
- - `options`:
78
- - `{ label: "Test now", description: "Removes the worktree and checks the branch out into the main repo for Xcode / manual test" }`
79
- - `{ label: "Skip", description: "Stay in the worktree and go to Phase 6" }`
80
- - **Skip** → set `state.phases["5"].status = "skipped"` (so Phase 6 can offer the local-checkout prompt) → Phase 6
81
- - **Test now** → set `state.phases["5"].status = "in_progress"` → continue:
82
- 2. **Commit changes in worktree BEFORE removing** (WIP commit to preserve work):
83
- ```
84
- git -C {worktree-path} add -A
85
- git -C {worktree-path} commit -m "WIP: {jiraId} - changes for user test"
86
- ```
87
- 3. Remove worktree, checkout branch in main repo:
88
- ```
89
- git worktree remove .worktrees/{jiraId} --force
90
- git checkout {branch-name}
91
- ```
92
- Branch now has the WIP commit - all changes are preserved.
93
- 4. Show test instructions:
94
- ```
95
- Switched to branch: {branch-name}
96
- To test: Xcode -> build -> run -> manual test
97
- "ok" -> proceeds to Phase 6 (WIP commit will be replaced via git reset HEAD~1 + proper commit)
98
- "fix: ..." -> worktree is recreated, returns to Phase 3
99
- ```
100
- 5. Mark the tracker as waiting, then wait for the user response:
101
- ```bash
102
- bash $HOME/.claude/scripts/phase-tracker.sh now 5 "awaiting local test (user)"
103
- bash $HOME/.claude/scripts/phase-tracker.sh render
104
- ```
105
- The waiting state persists in `tracker-state.json` across the handoff; `/multi-agent:resume-local` and `/multi-agent:manual-test` CONTINUE this state file and never re-init it (`$HOME/.claude/multi-agent-refs/tracker-contract.md` "Continuation runs").
106
-
107
- **"ok" is a structured result, not a word.** Before "ok" is accepted, the run writes `$WORKTREE/.pipeline/manual-test.json`: one entry per acceptance criterion, the criteria taken from the analysis doc test plan (Section 15 / 20), the plan tasks, and the user's own words in the reply. Every criterion records what was seen; a criterion that was not tried says so with a reason.
108
- ```json
109
- {"criteria":[{"spec":"<quote>","source":"analysis 15.2 | plan task 3 | user","observed":"<what was seen>","verdict":"pass|fail|not-tested","reason":"<required when not-tested>","screenshot":"<path or null>"}],"verdict":"passed|failed"}
110
- ```
111
- Then gate it:
112
- ```bash
113
- node $HOME/.claude/scripts/evidence-gate.mjs --claim manual --status passed --evidence "$WORKTREE/.pipeline/manual-test.json"
114
- ```
115
- Add `--require-screenshot` when `state.visualEvidence.required` is true: a passing criterion then has to name a file that exists, because a `screenshot` key pointing nowhere is not evidence.
116
-
117
- Exit 1 means the "ok" is not accepted: tell the user which criterion is missing evidence (a `fail` verdict, or `not-tested` without a reason) and wait for the next reply. Exit 0 marks Phase 5 completed with `Result "local test passed (user)"`. The "fix: ..." path below is unchanged.
118
- 6. If fix needed:
119
- - Branch already has WIP commit (from step 2) - changes are safe
120
- - **Heal stale admin state first** (same contract as Phase 0 - step 3's
121
- `worktree remove` or an interrupted run can leave a stale entry, so a bare
122
- re-add fails with `already exists`/`already registered`):
123
- ```bash
124
- git -C "$PROJECT_ROOT" worktree prune 2>/dev/null || true
125
- if git -C "$PROJECT_ROOT" worktree list --porcelain | grep -qF "{worktree-path}"; then
126
- git -C "$PROJECT_ROOT" worktree unlock "{worktree-path}" 2>/dev/null || true
127
- fi
128
- ```
129
- Phase 0's repo residue guard is already in `.git/info/exclude` - no re-add.
130
- - Recreate worktree from branch: `git -C $PROJECT_ROOT worktree add {worktree-path} {branch}`
131
- - Re-set git identity: `git -C {worktree-path} config user.name/email` (from state)
132
- - Go back to Phase 3
133
- 7. Log: "Phase 5: Test - {result}"
134
-
135
- **CRITICAL**: Never remove a worktree with uncommitted changes. Always WIP commit first.
136
-
137
- #### Automated Device Checks (on-demand)
138
-
139
- Before or during user testing, run device-level audits via Bash if user requests. See `audit-guide.md` for commands.
140
-
141
- | Check | When | Command |
142
- | ------------------- | ----------------- | ----------------------------------------- |
143
- | UI flow video | `state.visualEvidence.required` AND Phase 3 recorded none | `capture-evidence.sh video start` -> drive the flow -> `video stop` -> `fit`. See below |
144
- | Accessibility audit | UI changes | `mcp__multi-agent-toolkit__{ios,android}_accessibility_audit` |
145
- | Biometric test | Auth flow changes | ios: `mcp__multi-agent-toolkit__ios_biometric` (android: manual) |
146
- | Launch time | Perf-sensitive changes | ios: app-launch instrument · android: `mcp__multi-agent-toolkit__android_launch_time` |
147
- | Visual test | Any UI changes | `/multi-agent test` (sim-test, both platforms) |
148
- | Snapshot regression | Component / pixel-stable UI changes | ios: `mcp__multi-agent-toolkit__ios_visual_diff` · android: `mcp__multi-agent-toolkit__android_screenshot` + compare |
149
- | Store screenshots | `taskType === screenshot` | ios: `ios_status_bar({preset: "clean"})` · android: `android_screenshot` |
150
-
151
- ##### UI flow video, when Phase 3 produced none
152
-
153
- Phase 3 Step 3.55 is the primary host and runs in every mode. Phase 5 is the richer one where it exists: the device is up and a person is watching, so the flow is one somebody confirmed. It adds, never replaces.
154
-
155
- Run only when `visualEvidence.required` and `visualEvidence.video.file` is empty: re-check the device with `$HOME/.claude/scripts/probe-evidence-capability.sh --only device`, then `$HOME/.claude/scripts/capture-evidence.sh video start` -> drive the flow (`run-ui-tests.sh run`, or the Phase 5 scenarios by hand) -> `video stop` -> `fit`. The cap is `visualEvidence.maxVideoSeconds`, read through `capture-evidence.sh limits` so one value serves both. Update `videoTier` and `videoTierReason` with what ran.
156
-
157
- When the intake answer was `unit` there is no recording here either: overriding it in a phase the user may not be watching makes the question decorative.
158
-
159
- Results included in Phase 7 report. MCP tools preferred when available - concise structured output, lower token cost.
160
-
161
- **Snapshot regression flow (optional):** when the task changes a stable component, capture a screenshot before the change (baseline) and after (current), then call `ios_visual_diff({baseline, current, max_diff_pct: 1.0})`. Threshold can be relaxed for animated / non-deterministic regions - keep `max_diff_pct ≤ 1.0` for static layouts.
162
-
163
- #### Security Audit (store-readiness)
164
-
165
- When the task touches authentication, keychain, network, or is scheduled for an imminent release, launch the `security-auditor` subagent to run an OWASP Mobile Top 10 pass plus App Store / Play Store compliance checks:
166
-
167
- ```
168
- Agent(subagent_type: "security-auditor", prompt: "<diff + context>")
169
- ```
170
-
171
- Returns severity-tagged findings (Critical / High / Medium). Critical items block Phase 6 just like Phase 4 blockers; High items are logged and surfaced in Phase 7 report. Skipped by default - opt-in for release branches or on explicit `/multi-agent "<task>" --audit` flag.
172
-
173
- #### Telemetry - token forwarding
174
-
175
- When the security-auditor or any other Phase 5 sub-agent runs, forward its token totals so Phase 7's Cost Breakdown captures Phase 5:
176
-
177
- ```bash
178
- LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 5 audit.completed \
179
- model=opus tokens_in=$IN tokens_out=$OUT duration_ms=$DUR
180
- ```
181
-
182
- If Phase 5 is purely user-driven (no sub-agent ran), no token forwarding is required and the cost block stays empty for this phase. Best-effort. See `$HOME/.claude/multi-agent-refs/progress-contract.md#token-telemetry-forwarding`.