@mmerterden/multi-agent-pipeline 17.6.0 → 19.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (272) hide show
  1. package/CHANGELOG.md +310 -0
  2. package/README.md +76 -18
  3. package/README.tr.md +55 -16
  4. package/docs/adr/0002-instruction-driven-flag.md +1 -0
  5. package/docs/adr/0005-lazy-phase-docs.md +11 -1
  6. package/docs/adr/0008-installer-modularization-and-secret-leak-defense.md +1 -0
  7. package/docs/adr/0010-own-code-graph.md +1 -0
  8. package/docs/adr/0011-dormant-ci.md +25 -1
  9. package/docs/adr/0014-six-phase-consolidation.md +134 -0
  10. package/docs/adr/README.md +2 -1
  11. package/docs/architecture.md +37 -38
  12. package/docs/best-practices.md +1 -1
  13. package/docs/ecosystem.md +37 -26
  14. package/docs/engineering.md +1 -1
  15. package/docs/facts.json +45 -0
  16. package/docs/features.md +54 -53
  17. package/docs/performance.md +5 -5
  18. package/docs/recovery-guide.md +9 -9
  19. package/docs/server-readiness.md +188 -0
  20. package/docs/token-budget-history.md +3 -1
  21. package/index.js +18 -3
  22. package/install/_codex-agents.mjs +1 -1
  23. package/install/_common.mjs +42 -17
  24. package/install/_dev-only-files.mjs +8 -0
  25. package/install/_unattended-profile.mjs +113 -0
  26. package/install/index.mjs +48 -0
  27. package/install/templates/claude-hooks.json +1 -1
  28. package/install/templates/codex-instructions.md +1 -1
  29. package/install/templates/copilot-instructions.md +28 -28
  30. package/manifest.json +1065 -0
  31. package/package.json +6 -3
  32. package/pipeline/agents/dev-critic.md +3 -3
  33. package/pipeline/commands/figma-to-swiftui.md +1 -1
  34. package/pipeline/commands/multi-agent/SKILL.md +8 -8
  35. package/pipeline/commands/multi-agent/analysis/SKILL.md +9 -9
  36. package/pipeline/commands/multi-agent/autopilot/SKILL.md +7 -7
  37. package/pipeline/commands/multi-agent/channels/SKILL.md +15 -15
  38. package/pipeline/commands/multi-agent/diff-explain/SKILL.md +6 -6
  39. package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +1 -1
  40. package/pipeline/commands/multi-agent/graph/SKILL.md +1 -1
  41. package/pipeline/commands/multi-agent/help/SKILL.md +62 -62
  42. package/pipeline/commands/multi-agent/language/SKILL.md +2 -2
  43. package/pipeline/commands/multi-agent/local/SKILL.md +11 -11
  44. package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +13 -13
  45. package/pipeline/commands/multi-agent/log/SKILL.md +2 -2
  46. package/pipeline/commands/multi-agent/manual-test/SKILL.md +9 -9
  47. package/pipeline/commands/multi-agent/model/SKILL.md +69 -0
  48. package/pipeline/commands/multi-agent/refactor/SKILL.md +3 -3
  49. package/pipeline/commands/multi-agent/resume/SKILL.md +4 -4
  50. package/pipeline/commands/multi-agent/resume-local/SKILL.md +19 -17
  51. package/pipeline/commands/multi-agent/review/SKILL.md +1 -1
  52. package/pipeline/commands/multi-agent/route-off/SKILL.md +36 -0
  53. package/pipeline/commands/multi-agent/route-on/SKILL.md +74 -0
  54. package/pipeline/commands/multi-agent/route-status/SKILL.md +56 -0
  55. package/pipeline/commands/multi-agent/setup/SKILL.md +2 -2
  56. package/pipeline/commands/multi-agent/status/SKILL.md +54 -23
  57. package/pipeline/commands/multi-agent/steer/SKILL.md +2 -2
  58. package/pipeline/commands/multi-agent/sync/SKILL.md +12 -13
  59. package/pipeline/commands/multi-agent/test/SKILL.md +1 -1
  60. package/pipeline/lib/_jira-auth.sh +8 -0
  61. package/pipeline/lib/analysis-jira-write.sh +32 -0
  62. package/pipeline/lib/ask-choice.sh +13 -2
  63. package/pipeline/lib/autopilot-state.sh +8 -0
  64. package/pipeline/lib/credential-inventory.sh +1 -1
  65. package/pipeline/lib/fatal.mjs +129 -0
  66. package/pipeline/lib/fetch-fortify.sh +1 -1
  67. package/pipeline/lib/figma-mcp-refresh.sh +18 -0
  68. package/pipeline/lib/figma-screenshot.sh +18 -0
  69. package/pipeline/lib/invoked-directly.mjs +43 -0
  70. package/pipeline/lib/jira-publish.sh +42 -0
  71. package/pipeline/lib/md2confluence-v3.py +47 -0
  72. package/pipeline/lib/model-rung.sh +142 -0
  73. package/pipeline/lib/outbound-gate.mjs +175 -0
  74. package/pipeline/lib/phase-schema.mjs +88 -0
  75. package/pipeline/lib/plan-todos.sh +32 -11
  76. package/pipeline/lib/post-pr-review.sh +77 -8
  77. package/pipeline/lib/repo-hygiene.sh +8 -3
  78. package/pipeline/lib/require-jq.sh +40 -0
  79. package/pipeline/lib/route-state.sh +161 -0
  80. package/pipeline/lib/run-paths.sh +335 -0
  81. package/pipeline/multi-agent-refs/_account-picker.md +1 -1
  82. package/pipeline/multi-agent-refs/_dev-context.md +1 -1
  83. package/pipeline/multi-agent-refs/_input-parser.md +1 -1
  84. package/pipeline/multi-agent-refs/analysis/evidence.md +0 -9
  85. package/pipeline/multi-agent-refs/analysis/intake.md +1 -1
  86. package/pipeline/multi-agent-refs/analysis/locked.md +21 -22
  87. package/pipeline/multi-agent-refs/analysis/render.md +1 -1
  88. package/pipeline/multi-agent-refs/analysis/synthesis.md +12 -6
  89. package/pipeline/multi-agent-refs/android-guide.md +1 -1
  90. package/pipeline/multi-agent-refs/audit-guide.md +13 -13
  91. package/pipeline/multi-agent-refs/channels/issue-comment.md +2 -2
  92. package/pipeline/multi-agent-refs/channels/jira.md +3 -3
  93. package/pipeline/multi-agent-refs/channels/pr.md +4 -4
  94. package/pipeline/multi-agent-refs/channels/wiki.md +1 -1
  95. package/pipeline/multi-agent-refs/component-dispatch.md +3 -3
  96. package/pipeline/multi-agent-refs/cross-cli-contract.md +31 -6
  97. package/pipeline/multi-agent-refs/features/autopilot-circuit-breaker.md +74 -4
  98. package/pipeline/multi-agent-refs/features/code-graph.md +5 -5
  99. package/pipeline/multi-agent-refs/features/cost-analysis.md +93 -0
  100. package/pipeline/multi-agent-refs/features/design-conformance.md +1 -1
  101. package/pipeline/multi-agent-refs/features/dev-critic.md +3 -3
  102. package/pipeline/multi-agent-refs/features/doctor.md +47 -2
  103. package/pipeline/multi-agent-refs/features/external-context-injection.md +3 -3
  104. package/pipeline/multi-agent-refs/features/maturity-followup.md +3 -3
  105. package/pipeline/multi-agent-refs/features/model-fallback.md +5 -5
  106. package/pipeline/multi-agent-refs/features/plan-todos.md +1 -1
  107. package/pipeline/multi-agent-refs/features/repo-map.md +1 -1
  108. package/pipeline/multi-agent-refs/features/review-delta.md +3 -3
  109. package/pipeline/multi-agent-refs/features/review-multi-repo.md +1 -1
  110. package/pipeline/multi-agent-refs/features/scope-check.md +4 -4
  111. package/pipeline/multi-agent-refs/features/skill-conformance.md +2 -2
  112. package/pipeline/multi-agent-refs/features/stack-skill-routing.md +1 -1
  113. package/pipeline/multi-agent-refs/features/verify-by-test.md +4 -4
  114. package/pipeline/multi-agent-refs/features/verify.md +83 -0
  115. package/pipeline/multi-agent-refs/features/visual-evidence.md +19 -19
  116. package/pipeline/multi-agent-refs/features/worktree-finalize.md +6 -6
  117. package/pipeline/multi-agent-refs/issue-jira-triad.md +10 -10
  118. package/pipeline/multi-agent-refs/knowledge.md +11 -11
  119. package/pipeline/multi-agent-refs/multi-repo-integration-build.md +13 -13
  120. package/pipeline/multi-agent-refs/payload-contracts.md +8 -8
  121. package/pipeline/multi-agent-refs/phases/log-format.md +10 -10
  122. package/pipeline/multi-agent-refs/phases/modes.md +30 -30
  123. package/pipeline/multi-agent-refs/phases/operations.md +21 -10
  124. package/pipeline/multi-agent-refs/phases/phase-0-init.md +25 -25
  125. package/pipeline/multi-agent-refs/phases/phase-1-plan.md +599 -0
  126. package/pipeline/multi-agent-refs/phases/{phase-3-dev.md → phase-2-dev.md} +129 -49
  127. package/pipeline/multi-agent-refs/phases/{phase-4-review.md → phase-3-review.md} +225 -107
  128. package/pipeline/multi-agent-refs/phases/{phase-6-commit.md → phase-4-commit.md} +23 -23
  129. package/pipeline/multi-agent-refs/phases/{phase-7-report.md → phase-5-report.md} +29 -29
  130. package/pipeline/multi-agent-refs/phases.md +44 -48
  131. package/pipeline/multi-agent-refs/picker-contract.md +1 -1
  132. package/pipeline/multi-agent-refs/progress-contract.md +6 -6
  133. package/pipeline/multi-agent-refs/readiness-review.md +1 -1
  134. package/pipeline/multi-agent-refs/rules.md +7 -7
  135. package/pipeline/multi-agent-refs/swiftui-guide.md +2 -2
  136. package/pipeline/multi-agent-refs/tracker-contract.md +31 -32
  137. package/pipeline/multi-agent-refs/unattended-contract.md +129 -0
  138. package/pipeline/multi-agent-refs/wiki-capture.md +14 -14
  139. package/pipeline/preferences-template.json +9 -1
  140. package/pipeline/rules/outside-the-pipeline.md +1 -1
  141. package/pipeline/schemas/agent-state.schema.json +50 -50
  142. package/pipeline/schemas/analysis-output.schema.json +2 -2
  143. package/pipeline/schemas/autopilot-config.schema.json +1 -1
  144. package/pipeline/schemas/code-graph.schema.json +1 -1
  145. package/pipeline/schemas/criteria-manifest.schema.json +1 -1
  146. package/pipeline/schemas/dev-critic-output.schema.json +1 -1
  147. package/pipeline/schemas/diff-risk.schema.json +1 -1
  148. package/pipeline/schemas/migrations/prefs-2.4.0-to-2.5.0.mjs +2 -2
  149. package/pipeline/schemas/migrations/prefs-2.6.0-to-2.7.0.mjs +31 -0
  150. package/pipeline/schemas/migrations/state-2.1.0-to-2.2.0.mjs +129 -0
  151. package/pipeline/schemas/phases.json +105 -0
  152. package/pipeline/schemas/plan-todos.schema.json +5 -5
  153. package/pipeline/schemas/planning-output.schema.json +1 -1
  154. package/pipeline/schemas/prefs.schema.json +100 -56
  155. package/pipeline/schemas/reviewer-output.schema.json +3 -3
  156. package/pipeline/schemas/route-config.schema.json +74 -0
  157. package/pipeline/schemas/scope-check.schema.json +1 -1
  158. package/pipeline/schemas/test-gap.schema.json +1 -1
  159. package/pipeline/schemas/token-budget.json +12 -18
  160. package/pipeline/schemas/triage-output.schema.json +6 -6
  161. package/pipeline/scripts/README.md +3 -3
  162. package/pipeline/scripts/_code-graph.mjs +2 -2
  163. package/pipeline/scripts/_run-paths.mjs +372 -0
  164. package/pipeline/scripts/_smoke-root.sh +1 -1
  165. package/pipeline/scripts/aggregate-metrics.mjs +65 -65
  166. package/pipeline/scripts/autopilot-arming.mjs +2 -1
  167. package/pipeline/scripts/autopilot-intake.mjs +2 -1
  168. package/pipeline/scripts/autopilot-runner.mjs +206 -2
  169. package/pipeline/scripts/build-references.mjs +2 -1
  170. package/pipeline/scripts/build-stack-plugins.mjs +10 -2
  171. package/pipeline/scripts/capture-evidence.sh +7 -2
  172. package/pipeline/scripts/capture-flush.sh +8 -8
  173. package/pipeline/scripts/capture-resume.sh +3 -3
  174. package/pipeline/scripts/classify-plan-safety.mjs +3 -2
  175. package/pipeline/scripts/cost-analyze.mjs +600 -0
  176. package/pipeline/scripts/cost-budget-check.mjs +4 -12
  177. package/pipeline/scripts/council-view.mjs +2 -1
  178. package/pipeline/scripts/crush-json.mjs +2 -1
  179. package/pipeline/scripts/diff-explain.mjs +7 -10
  180. package/pipeline/scripts/diff-risk-score.mjs +2 -1
  181. package/pipeline/scripts/doctor.mjs +140 -6
  182. package/pipeline/scripts/evidence-gate.mjs +9 -3
  183. package/pipeline/scripts/feedback-send.mjs +12 -2
  184. package/pipeline/scripts/gc-abandoned.sh +32 -16
  185. package/pipeline/scripts/gc-tmp.sh +1 -1
  186. package/pipeline/scripts/gc-worktrees.sh +12 -5
  187. package/pipeline/scripts/gen-facts.mjs +175 -0
  188. package/pipeline/scripts/gen-mode-dispatch.mjs +32 -37
  189. package/pipeline/scripts/gen-ref-toc.mjs +1 -1
  190. package/pipeline/scripts/github-ssh-setup.sh +64 -7
  191. package/pipeline/scripts/graph-mermaid.mjs +4 -2
  192. package/pipeline/scripts/graph-report.mjs +1 -1
  193. package/pipeline/scripts/jira-attach.sh +1 -1
  194. package/pipeline/scripts/keychain-save.sh +101 -30
  195. package/pipeline/scripts/learn-from-transcripts.mjs +3 -2
  196. package/pipeline/scripts/learning-curve.mjs +36 -31
  197. package/pipeline/scripts/log-metric.sh +17 -4
  198. package/pipeline/scripts/make-manifest.mjs +199 -0
  199. package/pipeline/scripts/memory-save.sh +1 -1
  200. package/pipeline/scripts/migrate-prefs.mjs +24 -6
  201. package/pipeline/scripts/migrate-state.mjs +94 -4
  202. package/pipeline/scripts/phase-banner.sh +26 -22
  203. package/pipeline/scripts/phase-tracker.sh +48 -10
  204. package/pipeline/scripts/plan-coverage-gate.mjs +8 -4
  205. package/pipeline/scripts/pre-commit-check.sh +7 -0
  206. package/pipeline/scripts/pre-push-check.sh +7 -0
  207. package/pipeline/scripts/purge.sh +23 -6
  208. package/pipeline/scripts/render-agent-log-cost.sh +10 -3
  209. package/pipeline/scripts/render-cost-summary.sh +9 -2
  210. package/pipeline/scripts/render-work-summary.sh +14 -7
  211. package/pipeline/scripts/review-file-filter.mjs +5 -3
  212. package/pipeline/scripts/review-scope.mjs +2 -1
  213. package/pipeline/scripts/routine-registry.mjs +2 -1
  214. package/pipeline/scripts/run-aggregator.mjs +26 -20
  215. package/pipeline/scripts/run-metrics.mjs +4 -2
  216. package/pipeline/scripts/runs-index.mjs +353 -0
  217. package/pipeline/scripts/scorecard-snapshot.mjs +178 -0
  218. package/pipeline/scripts/search-logs.sh +18 -0
  219. package/pipeline/scripts/smoke-cross-cli-behavior.sh +6 -6
  220. package/pipeline/scripts/smoke-schema-validation.sh +26 -7
  221. package/pipeline/scripts/test-gap-scan.mjs +2 -1
  222. package/pipeline/scripts/test-integrity-gate.mjs +2 -1
  223. package/pipeline/scripts/token-budget-report.mjs +13 -2
  224. package/pipeline/scripts/triage-memory.mjs +2 -2
  225. package/pipeline/scripts/update-issue-progress.sh +56 -7
  226. package/pipeline/scripts/usage-report.mjs +12 -1
  227. package/pipeline/scripts/validate-analysis-doc.mjs +75 -18
  228. package/pipeline/scripts/validate-code-graph.mjs +6 -3
  229. package/pipeline/scripts/validate-complaint-doc.mjs +2 -1
  230. package/pipeline/scripts/validate-diff-risk.mjs +6 -3
  231. package/pipeline/scripts/validate-planning.mjs +1 -1
  232. package/pipeline/scripts/validate-reviewer.mjs +1 -1
  233. package/pipeline/scripts/validate-state.mjs +45 -5
  234. package/pipeline/scripts/validate-test-gap.mjs +6 -3
  235. package/pipeline/scripts/validate-triage.mjs +6 -4
  236. package/pipeline/scripts/verify-citations.mjs +4 -2
  237. package/pipeline/scripts/verify.mjs +327 -0
  238. package/pipeline/scripts/worktree-finalize.sh +18 -9
  239. package/pipeline/scripts/write-state.mjs +154 -15
  240. package/pipeline/skills/.skill-manifest.json +37 -21
  241. package/pipeline/skills/.skills-index.json +104 -5
  242. package/pipeline/skills/shared/README.md +15 -6
  243. package/pipeline/skills/shared/core/apple-archive-compliance/SKILL.md +2 -2
  244. package/pipeline/skills/shared/core/google-play-compliance/SKILL.md +2 -2
  245. package/pipeline/skills/shared/core/multi-agent/SKILL.md +69 -71
  246. package/pipeline/skills/shared/core/multi-agent-autopilot/SKILL.md +3 -3
  247. package/pipeline/skills/shared/core/multi-agent-channels/SKILL.md +14 -14
  248. package/pipeline/skills/shared/core/multi-agent-diff-explain/SKILL.md +5 -5
  249. package/pipeline/skills/shared/core/multi-agent-graph/SKILL.md +1 -1
  250. package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +25 -23
  251. package/pipeline/skills/shared/core/multi-agent-language/SKILL.md +2 -2
  252. package/pipeline/skills/shared/core/multi-agent-local/SKILL.md +2 -2
  253. package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +8 -8
  254. package/pipeline/skills/shared/core/multi-agent-manual-test/SKILL.md +6 -6
  255. package/pipeline/skills/shared/core/multi-agent-model/SKILL.md +71 -0
  256. package/pipeline/skills/shared/core/multi-agent-refactor/SKILL.md +3 -3
  257. package/pipeline/skills/shared/core/multi-agent-resume/SKILL.md +1 -1
  258. package/pipeline/skills/shared/core/multi-agent-resume-local/SKILL.md +7 -7
  259. package/pipeline/skills/shared/core/multi-agent-route-off/SKILL.md +39 -0
  260. package/pipeline/skills/shared/core/multi-agent-route-on/SKILL.md +76 -0
  261. package/pipeline/skills/shared/core/multi-agent-route-status/SKILL.md +59 -0
  262. package/pipeline/skills/shared/core/multi-agent-setup/SKILL.md +1 -1
  263. package/pipeline/skills/shared/core/multi-agent-status/SKILL.md +35 -11
  264. package/pipeline/skills/shared/core/multi-agent-steer/SKILL.md +2 -2
  265. package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +6 -5
  266. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/package_app.sh +4 -1
  267. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/setup_dev_signing.sh +4 -1
  268. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/sign-and-notarize.sh +2 -1
  269. package/pipeline/skills/skills-index.md +13 -4
  270. package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +0 -263
  271. package/pipeline/multi-agent-refs/phases/phase-2-planning.md +0 -344
  272. package/pipeline/multi-agent-refs/phases/phase-5-test.md +0 -182
@@ -14,7 +14,7 @@
14
14
  - [Compliance Rules (maps to multi-agent-toolkit MCP audit tools)](#compliance-rules-maps-to-multi-agent-toolkit-mcp-audit-tools)
15
15
  <!-- /toc -->
16
16
 
17
- > **MUST: Figma MCP-first (BLOCKING).** If the task references any Figma frame (URL, node ID, or "from the design"), the Dev phase MUST call `mcp__claude_ai_Figma__get_design_context` for every frame BEFORE writing a single Composable line. Use the `CodeConnectSnippet` component name verbatim - no sound-alike substitutions. Authentication failure is not a skip path. Full rule, trigger conditions, and gate failure modes: `$HOME/.claude/rules/figma-pipeline.md` "MUST: Figma MCP-first (BLOCKING)". Phase wiring: `$HOME/.claude/multi-agent-refs/phases/phase-3-dev.md` "MUST: Figma MCP-first (BLOCKING pre-step)".
17
+ > **MUST: Figma MCP-first (BLOCKING).** If the task references any Figma frame (URL, node ID, or "from the design"), the Dev phase MUST call `mcp__claude_ai_Figma__get_design_context` for every frame BEFORE writing a single Composable line. Use the `CodeConnectSnippet` component name verbatim - no sound-alike substitutions. Authentication failure is not a skip path. Full rule, trigger conditions, and gate failure modes: `$HOME/.claude/rules/figma-pipeline.md` "MUST: Figma MCP-first (BLOCKING)". Phase wiring: `$HOME/.claude/multi-agent-refs/phases/phase-2-dev.md` "MUST: Figma MCP-first (BLOCKING pre-step)".
18
18
 
19
19
  When the task involves creating an Android UI component (Jetpack Compose), follow this architecture.
20
20
 
@@ -19,16 +19,16 @@ Standalone audit commands - runs directly via Bash, **no MCP server dependency
19
19
  **These audits are on-demand.** They run when:
20
20
 
21
21
  1. User explicitly requests: `/multi-agent test "accessibility"`, `/multi-agent test "store-ready"`
22
- 2. User asks during Phase 5: "run accessibility audit", "check store compliance"
22
+ 2. User asks during the Phase 3 user test: "run accessibility audit", "check store compliance"
23
23
  3. Pipeline suggests and user confirms: "UI changes detected - want to run accessibility audit?"
24
24
 
25
- **Pipeline never runs audits without user intent.** Phase 4 does code-level review (free, automatic). Phase 5 does device-level audit only when requested.
25
+ **Pipeline never runs audits without user intent.** Phase 3 does code-level review (free, automatic) and the device-level audit only when requested - both live in Review now, at different steps, which is why "Review ran" does not mean a device was touched.
26
26
 
27
27
  ---
28
28
 
29
29
  ### iOS Accessibility Audit
30
30
 
31
- **When**: Phase 5 - user requests, app is running on simulator
31
+ **When**: Phase 3 user test - user requests, app is running on simulator
32
32
  **What it checks**: Missing labels, small tap targets (<44pt), missing identifiers
33
33
 
34
34
  **How to run** (requires ui-tree-dumper.swift in project or ~/.claude/scripts/):
@@ -73,7 +73,7 @@ Accessibility Audit:
73
73
 
74
74
  ### Android Accessibility Audit
75
75
 
76
- **When**: Phase 5 - user requests, app is running on emulator
76
+ **When**: Phase 3 user test - user requests, app is running on emulator
77
77
  **What it checks**: Missing contentDescription, small touch targets (<48dp), missing resource-id
78
78
 
79
79
  ```bash
@@ -98,7 +98,7 @@ cat /tmp/_audit_ui.xml
98
98
 
99
99
  ### iOS Biometric Test
100
100
 
101
- **When**: Phase 5 - auth flow testing
101
+ **When**: Phase 3 user test - auth flow testing
102
102
 
103
103
  ```bash
104
104
  DEVICE_ID=$(xcrun simctl list devices booted -j | python3 -c "import sys,json; devs=json.load(sys.stdin)['devices']; print(next(d['udid'] for ds in devs.values() for d in ds if d['state']=='Booted'))")
@@ -119,7 +119,7 @@ xcrun simctl keychain $DEVICE_ID biometric-match --face --no-match
119
119
 
120
120
  ### Android Launch Time
121
121
 
122
- **When**: Phase 5 - performance baseline
122
+ **When**: Phase 3 user test - performance baseline
123
123
 
124
124
  ```bash
125
125
  # Force stop first (cold start)
@@ -143,7 +143,7 @@ adb shell am start -W -n {package_name}/.MainActivity 2>&1
143
143
 
144
144
  ### iOS Archive Audit (App Store Compliance)
145
145
 
146
- **When**: Phase 6 - release branches only, user confirms
146
+ **When**: Phase 4 - release branches only, user confirms
147
147
  **Input**: Path to .xcarchive
148
148
 
149
149
  ```bash
@@ -198,7 +198,7 @@ ls "$APP_DIR/Frameworks/" 2>/dev/null
198
198
 
199
199
  ### Android APK Audit (Play Store Compliance)
200
200
 
201
- **When**: Phase 6 - release branches only, user confirms
201
+ **When**: Phase 4 - release branches only, user confirms
202
202
  **Input**: Path to .apk
203
203
 
204
204
  ```bash
@@ -238,11 +238,11 @@ unzip -l "$APK" 2>/dev/null | grep "classes.*\.dex" | wc -l
238
238
 
239
239
  | Phase | What Happens | Method |
240
240
  | ---------------- | --------------------------------------------------- | ------------------------------- |
241
- | Phase 4 (Review) | Accessibility check - **code-level only** | AI reads source code, no device |
242
- | Phase 5 (Test) | Accessibility audit - **device-level, on-demand** | Bash commands above |
243
- | Phase 5 (Test) | Biometric test - **on-demand** | `xcrun simctl keychain` |
244
- | Phase 5 (Test) | Launch time - **on-demand** | `adb shell am start -W` |
245
- | Phase 6 (Commit) | Archive/APK audit - **release branches, on-demand** | Bash commands above |
241
+ | Phase 3 (Review) | Accessibility check - **code-level only** | AI reads source code, no device |
242
+ | Phase 3 (user test) | Accessibility audit - **device-level, on-demand** | Bash commands above |
243
+ | Phase 3 (user test) | Biometric test - **on-demand** | `xcrun simctl keychain` |
244
+ | Phase 3 (user test) | Launch time - **on-demand** | `adb shell am start -W` |
245
+ | Phase 4 (Commit) | Archive/APK audit - **release branches, on-demand** | Bash commands above |
246
246
 
247
247
  ### Graceful Degradation
248
248
 
@@ -124,10 +124,10 @@ Token resolution: `gh auth status` for the active GitHub account selected in Pha
124
124
 
125
125
  ## Pairing with the Progress flag updater
126
126
 
127
- Every issue comment post is paired with `$HOME/.claude/scripts/update-issue-progress.sh "$TASK_ID"` in the same Phase 7 step. Order: comment FIRST (so the timestamp marks the run), flags SECOND (so the body diff is one logical change).
127
+ Every issue comment post is paired with `$HOME/.claude/scripts/update-issue-progress.sh "$TASK_ID"` in the same Phase 5 step. Order: comment FIRST (so the timestamp marks the run), flags SECOND (so the body diff is one logical change).
128
128
 
129
129
  ```bash
130
- # Phase 7 Step 5 - issue channel
130
+ # Phase 5 Step 5 - issue channel
131
131
  gh issue comment "$ISSUE_NUMBER" --repo "$ORG/$REPO" --body-file /tmp/channels-$TASK_ID-issue.md
132
132
  bash $HOME/.claude/scripts/update-issue-progress.sh "$TASK_ID"
133
133
  ```
@@ -42,7 +42,7 @@ it is the PR body (`channels/pr.md`).
42
42
 
43
43
  **`summary`** - 2-5 sentences in `outputLanguage`. What changed, why, and the user-visible impact. No "we", no marketing tone. Past tense (the work is done at the time the comment goes up).
44
44
 
45
- **Visual evidence inside these sections.** When `state.visualEvidence` carries artefacts, they render INSIDE `summary` and `test_scenarios` - never as a section of their own, which the fixed section order forbids. Phase 6 Step 2.9 has already uploaded them and written the returned name to `visualEvidence.*[].jiraFilename`: reference that name, and call `jira-attach.sh <issue> <file>...` only for an artefact whose `jiraFilename` is absent. Uploading unconditionally here attaches every file a second time whenever the comment is re-rendered - which is exactly what a post-hoc `/multi-agent:channels` run does.
45
+ **Visual evidence inside these sections.** When `state.visualEvidence` carries artefacts, they render INSIDE `summary` and `test_scenarios` - never as a section of their own, which the fixed section order forbids. Phase 4 Step 2.9 has already uploaded them and written the returned name to `visualEvidence.*[].jiraFilename`: reference that name, and call `jira-attach.sh <issue> <file>...` only for an artefact whose `jiraFilename` is absent. Uploading unconditionally here attaches every file a second time whenever the comment is re-rendered - which is exactly what a post-hoc `/multi-agent:channels` run does.
46
46
 
47
47
  - `summary`, after its sentences: one line naming the pair in `outputLanguage` (`Düzeltme öncesi / Düzeltme sonrası`), then the thumbnails on the next line - `!<file>-before.png|thumbnail! !<file>-after.png|thumbnail!`.
48
48
  - `test_scenarios`, under the scenario the recording demonstrates: `!<file>-flow.mp4!` plus one line stating the tier used.
@@ -191,7 +191,7 @@ The adapter prepends the linked PR URL on the **first line** so reviewers can ju
191
191
  ```
192
192
  PR: https://github.com/<org>/<repo>/pull/4321
193
193
 
194
- Phase 4 review accepted 2 findings...
194
+ Phase 3 review accepted 2 findings...
195
195
  ```
196
196
 
197
197
  In multi-repo mode, `channels-multi-repo.sh render-jira <state> <body>` prepends a bulleted `* PR:` list - primary first, extras after. Single-repo tasks fall through to the unchanged body.
@@ -255,6 +255,6 @@ When the **Wiki** adapter writes pages on the same run AND `prefs.global.wikiToJ
255
255
  - UTF-8 in, UTF-8 out - the body file is UTF-8 and `--data-binary` ships its bytes verbatim. Never round-trip the body through `unicode_escape`, `latin-1`, or any re-encode step, and never hand-roll a Python/curl helper that re-decodes it: that mangles Turkish chars (ç ş ı ö ü ğ) into mojibake (`Çözüm` → `Ãözüm`). Use the `jq --rawfile` + `--data-binary @file` path above as-is. Same rule for the PR / Confluence / Wiki adapters.
256
256
  - Section order is fixed: `summary` → `test_scenarios` → `context_refs`. Never insert sections between them; never reorder.
257
257
  - Humanizer pass runs **after** body assembly and **before** wiki-markup conversion. Tone target: informal but technical. No marketing voice, no "we are excited", no "I have...".
258
- - **No decorative glyphs anywhere in the comment body** - neither emotive ( ) nor status glyphs ([done] [pending] failed skipped active ). The comment is plain technical prose. Status is written in words: `[done]` / `[pending]`, and `done · active · failed · skipped · pending` in the phase strip. This applies to every channel, not only Jira - see `channels/README` note in `phase-7-report.md`.
258
+ - **No decorative glyphs anywhere in the comment body** - neither emotive ( ) nor status glyphs ([done] [pending] failed skipped active ). The comment is plain technical prose. Status is written in words: `[done]` / `[pending]`, and `done · active · failed · skipped · pending` in the phase strip. This applies to every channel, not only Jira - see `channels/README` note in `phase-5-report.md`.
259
259
  - Body content language follows `prefs.global.outputLanguage`. Code identifiers, file paths, type names, branch names, and the wiki-markup syntax stay verbatim. The `promptLanguage="en"` lock means any LLM prompt that produces the body is in English; the body itself is then rendered in the user's language.
260
260
  - Adapter failures are non-blocking - a missing token or 401 returns `skipped`/`failed`, the loop keeps going for PR / Confluence / Wiki.
@@ -45,7 +45,7 @@ all - it is the section whose absence over there is the point.
45
45
 
46
46
  **`technical`** - the account for someone holding the diff: a short paragraph on the mechanism (what the code was actually doing wrong, or what the new code does), then the per-file bullets, then one line of diff stat (`2 files, 8 deletions, 0 insertions`). Where a change is safe for a reason that is not obvious from the diff - an equality relation preserved, an invariant kept, a call site left alone deliberately - that reason belongs here in a sentence, because it is the question the reviewer would otherwise ask in a comment. Tables are welcome when several symbols share a property worth listing side by side.
47
47
 
48
- The bullet list is one item per logically distinct change. Each bullet starts with the touched component and ends with a one-line "what". The source is `$WORKTREE/.pipeline/scope-check.json` `files[].reason` (Phase 3 Step 3.7): a file the dev could not justify there is a file this list cannot describe either, so the bullet quotes the gate output instead of inventing a reason. Use the stack's native file extensions / module paths - the example below shows the **shape**, not a stack lock-in:
48
+ The bullet list is one item per logically distinct change. Each bullet starts with the touched component and ends with a one-line "what". The source is `$WORKTREE/.pipeline/scope-check.json` `files[].reason` (Phase 2 Step 3.7): a file the dev could not justify there is a file this list cannot describe either, so the bullet quotes the gate output instead of inventing a reason. Use the stack's native file extensions / module paths - the example below shows the **shape**, not a stack lock-in:
49
49
 
50
50
  ```markdown
51
51
  ## Changes
@@ -135,7 +135,7 @@ The base sha matters because "it builds" is a claim about a merge base, and the
135
135
 
136
136
  Pick commands for the project's stack - the pipeline supports iOS (Swift/Xcode), Android (Gradle), web (npm/pnpm/yarn) and backend (pytest/jest/go test/etc). Multi-repo PRs (one PR per repo) emit the commands for that repo's stack only - never mix iOS + Android commands into a single PR body.
137
137
 
138
- **`risk`** - only when `state.diffRisk.signals` (Phase 4 Step 1.75) contains a high-stakes signal. Four fixed lines, each answered, never left as a placeholder; the source is Phase 1 `touchedAreas` plus the signals themselves, and when a signal is present the absence of this section is a Phase 6 Step 3 blocker:
138
+ **`risk`** - only when `state.diffRisk.signals` (Phase 3 Step 1.75) contains a high-stakes signal. Four fixed lines, each answered, never left as a placeholder; the source is Phase 1 `touchedAreas` plus the signals themselves, and when a signal is present the absence of this section is a Phase 4 Step 3 blocker:
139
139
 
140
140
  ```markdown
141
141
  ## Risk and Security
@@ -146,7 +146,7 @@ Pick commands for the project's stack - the pipeline supports iOS (Swift/Xcode
146
146
  - Rollback: feature flag <name> | git revert <sha> | none, and why
147
147
  ```
148
148
 
149
- **`visuals`** - only when `state.visualEvidence.required`. What this section can show depends on where the artefacts are hosted, which Phase 6 resolves into `state.visualEvidence.host`. Render the form for that host and no other.
149
+ **`visuals`** - only when `state.visualEvidence.required`. What this section can show depends on where the artefacts are hosted, which Phase 4 resolves into `state.visualEvidence.host`. Render the form for that host and no other.
150
150
 
151
151
  **`host: jira`.** Filenames, never URLs. A Jira attachment URL is auth-gated and renders as a broken image for anyone reading the PR outside a Jira session, and a broken image is worse than a filename because it looks like the evidence is missing.
152
152
 
@@ -192,7 +192,7 @@ Pick commands for the project's stack - the pipeline supports iOS (Swift/Xcode
192
192
 
193
193
  **Video is Jira-only.** On a GitHub-hosted run no recording is made and none is published: an mp4 behind a blob link is a download, not something a reviewer opens mid-review, and paying for a recording nobody watches is worse than saying plainly that there is none. The gap line carries that reason.
194
194
 
195
- Every `state.visualEvidence.gaps[]` entry becomes its own line with the reason instead of a filename (`- Before: none - the ticket carries no image attachment`). Phase 6 Step 3 blocks on a required artefact that is neither listed nor explained. Contract: `$HOME/.claude/multi-agent-refs/features/visual-evidence.md`.
195
+ Every `state.visualEvidence.gaps[]` entry becomes its own line with the reason instead of a filename (`- Before: none - the ticket carries no image attachment`). Phase 4 Step 3 blocks on a required artefact that is neither listed nor explained. Contract: `$HOME/.claude/multi-agent-refs/features/visual-evidence.md`.
196
196
 
197
197
  **`dependencies`** - only when `Package.swift` / `Podfile` / `build.gradle` / `package.json` changed. Each entry: `package@old → new - reason`.
198
198
 
@@ -64,4 +64,4 @@ Adapter implementation chooses the right git push target + commit message format
64
64
  - Screenshots cover light/dark + LTR/RTL - wiki adapter refuses to commit if any quadrant is missing.
65
65
  - Body content language follows `prefs.global.outputLanguage`. Code identifiers, file paths, type names, design token names, and Markdown formatting stay verbatim across languages (wiki pages are `.md` files - "wiki markup" in this doc set means Jira's dialect, which never appears here). The template (this doc) is English because `promptLanguage="en"` is locked.
66
66
  - No decorative/emotive emoji or smileys ( ) in the page prose. Only functional/structural marks a fixed template defines are allowed.
67
- - Autopilot always pauses at the channels menu (per `phase-7-report.md` autopilot contract) - even in autopilot mode the user gets to confirm wiki scope.
67
+ - Autopilot always pauses at the channels menu (per `phase-5-report.md` autopilot contract) - even in autopilot mode the user gets to confirm wiki scope.
@@ -13,7 +13,7 @@
13
13
 
14
14
  > **TLDR** - When `taskType === "component"` (Figma URL in task description or instruction-driven figma workflow), multi-agent Phase 3 **does not run the TDD loop**. It delegates the entire phase to the enabled `ai-<platform>-toolkit` **marketplace plugin's** component skill (`create-component`, falling back to `create-ui-component`) via the Skill tool. Implementation lives in the plugin; multi-agent's job is classification, dispatch, and state report. The pipeline no longer bundles its own `figma-to-component` orchestrator - component skills live in one place, the plugin marketplace.
15
15
 
16
- This doc is referenced from `$HOME/.claude/multi-agent-refs/phases/phase-3-dev.md`. Keeping it separate lets `phase-3-dev.md` remain tight (it's already the largest phase doc) and gives the orchestrator-report contract a stable URL for both Claude-side and Copilot-side implementations.
16
+ This doc is referenced from `$HOME/.claude/multi-agent-refs/phases/phase-2-dev.md`. Keeping it separate lets `phase-2-dev.md` remain tight (it's already the largest phase doc) and gives the orchestrator-report contract a stable URL for both Claude-side and Copilot-side implementations.
17
17
 
18
18
  ## Entry conditions
19
19
 
@@ -122,9 +122,9 @@ On failure (the plugin skill returns an unrecoverable build/test error, or the d
122
122
 
123
123
  ## Short-run behaviour
124
124
 
125
- When the Phase 0 Step 7.5 depth picker answered Short (`state.onlyDevelop === true`), the dispatch layer passes `mode: "dev"` so the plugin skill can elide unit tests and wiki (structural + snapshot still required; wiki deferred to Phase 7). If the plugin does not honor a `mode` hint, dispatch simply skips the post-build wiki step itself.
125
+ When the Phase 0 Step 7.5 depth picker answered Short (`state.onlyDevelop === true`), the dispatch layer passes `mode: "dev"` so the plugin skill can elide unit tests and wiki (structural + snapshot still required; wiki deferred to Phase 5). If the plugin does not honor a `mode` hint, dispatch simply skips the post-build wiki step itself.
126
126
 
127
- Phase 4 runs in a Short run as it does in a Full one, and its reviewer count is **not** Phase 3's concern - the Step 1.77 scope gate decides that from diff risk, independently of `mode`. What the dispatch layer owes Phase 4 is the record of which plugin skill it delegated to, appended to `state.telemetry.skillCalls[]`, so the review can check the delivered component against the criteria that skill imposes.
127
+ Phase 3 runs in a Short run as it does in a Full one, and its reviewer count is **not** Phase 3's concern - the Step 1.77 scope gate decides that from diff risk, independently of `mode`. What the dispatch layer owes Phase 4 is the record of which plugin skill it delegated to, appended to `state.telemetry.skillCalls[]`, so the review can check the delivered component against the criteria that skill imposes.
128
128
 
129
129
  ## Cross-CLI behaviour (intentional divergence)
130
130
 
@@ -1,9 +1,10 @@
1
1
  # Cross-CLI Contract (Claude Code · Copilot CLI · Codex CLI)
2
2
 
3
3
  <!-- toc -->
4
- - [1. Command Inventory (56 commands)](#1-command-inventory-56-commands)
4
+ - [1. Command Inventory (60 commands)](#1-command-inventory-60-commands)
5
5
  - [2. Canonical Placeholder Vocabulary](#2-canonical-placeholder-vocabulary)
6
6
  - [2.6 Intentional structural divergence - thin dispatcher vs inlined orchestrator](#26-intentional-structural-divergence---thin-dispatcher-vs-inlined-orchestrator)
7
+ - [2.7 One command, three different meanings: `/multi-agent:model`](#27-one-command-three-different-meanings-multi-agentmodel)
7
8
  - [3. Frontmatter Transform Rules (Claude ↔ Copilot)](#3-frontmatter-transform-rules-claude-copilot)
8
9
  - [4. Progress Signalling Parity](#4-progress-signalling-parity)
9
10
  - [5. Argument Parsing Invariants](#5-argument-parsing-invariants)
@@ -19,16 +20,17 @@
19
20
 
20
21
  ---
21
22
 
22
- ## 1. Command Inventory (56 commands)
23
+ ## 1. Command Inventory (60 commands)
23
24
 
24
25
  ```
25
26
  analysis, analysis-jira, analysis-resolve, autopilot, autopilot-off,
26
27
  autopilot-on, autopilot-status, build-optimize, channels, complaint-analysis,
27
28
  create-jira, design-check, diff-explain, doctor, feedback, forget,
28
29
  garbage-collect, graph, help, ios-coding-standard, issue, jira, kill,
29
- language, local, local-autopilot, log, manual-test, prune-logs,
30
+ language, local, local-autopilot, log, manual-test, model, prune-logs,
30
31
  prune-prompts, purge, refactor, resume, resume-local, review,
31
- review-analysis, review-issue, review-jira, routines, save, scan, search,
32
+ review-analysis, review-issue, review-jira, route-off, route-on,
33
+ route-status, routines, save, scan, search,
32
34
  setup, stack, status, steer, store-ready, sync, test, test-accessibility,
33
35
  test-dark-mode, test-dynamic-type, test-screenshots, testflight-validation,
34
36
  uninstall, update
@@ -199,7 +201,7 @@ on skill directories would demand exactly the layout that breaks it.
199
201
 
200
202
  **What must stay identical** (byte-level) across the two files:
201
203
 
202
- - Phase canonical labels (Phase 0-7)
204
+ - Phase canonical labels (Phase 0-5)
203
205
  - Input parsing table (Section 5 of this doc)
204
206
  - Routing decisions for each input type
205
207
  - Placeholder vocabulary (Section 2 of this doc)
@@ -216,7 +218,7 @@ Future changes that break an item in the "stay identical" list must update **bot
216
218
 
217
219
  ### Panel diversity per host
218
220
 
219
- Phase 4 runs three reviewers everywhere, but the diversity those three buy is not the
221
+ Phase 3 runs three reviewers everywhere, but the diversity those three buy is not the
220
222
  same on every host. Copilot CLI gets cross-VENDOR disagreement for free: GPT-5.4 sits
221
223
  beside two Claude models. Claude Code and Codex each run a one-vendor panel - three
222
224
  Anthropic models on one, three OpenAI models on the other - so the same three-way
@@ -236,6 +238,29 @@ members available there are closer to each other than Fable and Sonnet are. That
236
238
  weaker axis, not an equivalent one, and treating it as equivalent is the error this
237
239
  section exists to prevent.
238
240
 
241
+ ## 2.7 One command, three different meanings: `/multi-agent:model`
242
+
243
+ Most commands do the same thing on all three hosts. This one does not, and the
244
+ divergence is stated here rather than discovered at runtime, because a command
245
+ that quietly no-ops is worse than one that says it cannot act.
246
+
247
+ | Host | What `model on|off` does | Why |
248
+ |---|---|---|
249
+ | Claude Code | live switch: flips `modelFallback.fableEnabled` and realigns `costBudget.pricingModel` | the only host where the `fable` rung is Fable 5 |
250
+ | Copilot CLI | writes the preference, changes no dispatch | Fable 5 is not offered there; its personas never sat on this rung |
251
+ | Codex CLI | writes the preference, changes no dispatch, **deliberately** | the `fable` rung on Codex means `gpt-5.6 @ xhigh` - a different model on a different account. A knob named after an Anthropic model must not silently retune a Codex run |
252
+
253
+ `model-rung.sh` prints which of the three it is on every invocation. The rule it
254
+ follows: report the host's actual behaviour, never claim a change that did not
255
+ happen.
256
+
257
+ The routing trio (`route-on`, `route-off`, `route-status`) has no such split -
258
+ it writes `prefs.global.modelRouting` identically everywhere. Its honest limit is
259
+ different and applies on every host: a subagent cannot be dispatched to a
260
+ non-Anthropic model, because subagent dispatch belongs to the host. External
261
+ providers are reachable only where the pipeline makes the HTTP call itself
262
+ (`bulk-read.sh`, `research_ask`). `route-status` prints that on every run.
263
+
239
264
  ## 3. Frontmatter Transform Rules (Claude ↔ Copilot)
240
265
 
241
266
  Each file has a different frontmatter schema. The sync flow transforms between them:
@@ -1,6 +1,15 @@
1
1
  # Feature: Autopilot Circuit-Breaker
2
2
 
3
- **Pattern**: autopilot runs with zero interaction, which is exactly when a silent failure loop is most expensive - an agent can burn a budget re-attempting the same broken fix, or thrash between two phases, with nobody watching. A circuit-breaker converts "keep going no matter what" into "keep going until a defined unsafe condition, then halt and hand back to the user." This is the sanctioned autopilot pause (same class as the Phase 7 channels pause): the run stops, records why, and waits for an explicit `resume`.
3
+ <!-- toc -->
4
+ - [Trip conditions](#trip-conditions)
5
+ - [Wiring status](#wiring-status)
6
+ - [Action on trip](#action-on-trip)
7
+ - [Why this is the right autopilot exception](#why-this-is-the-right-autopilot-exception)
8
+ - [The runner-level breaker (a different failure, a different layer)](#the-runner-level-breaker-a-different-failure-a-different-layer)
9
+ - [Two more things a long-running runner needs](#two-more-things-a-long-running-runner-needs)
10
+ <!-- /toc -->
11
+
12
+ **Pattern**: autopilot runs with zero interaction, which is exactly when a silent failure loop is most expensive - an agent can burn a budget re-attempting the same broken fix, or thrash between two phases, with nobody watching. A circuit-breaker converts "keep going no matter what" into "keep going until a defined unsafe condition, then halt and hand back to the user." This is the sanctioned autopilot pause (same class as the Phase 5 channels pause): the run stops, records why, and waits for an explicit `resume`.
4
13
 
5
14
  **Gated by `prefs.global.autopilotCircuitBreaker`** (`enabled` default true, `identicalFindingCycles` default 2, `maxReworkCycles` default 3; `schemas/prefs.schema.json`). Halting is always safe, so the breaker itself defaults on. Disable per-run only with an explicit override. Complements, does not replace, the existing autopilot safety rules (build-fail max 3 retries, Phase 4 blocking-finding rework, destructive-op confirmations).
6
15
 
@@ -14,7 +23,7 @@ Any one trips the breaker. All are evaluated from `agent-state.json` + telemetry
14
23
  | 2 | **Identical repeated failure** | same normalized build-error signature, or the same Phase 4 finding fingerprint, recurs across consecutive Phase 3 rework cycles | 2 cycles |
15
24
  | 3 | **Rework storm** | Phase 4 -> Phase 3 rework cycles exceed the cap (distinct from the build-retry cap) | `maxReworkCycles` = 3 |
16
25
  | 4 | **Cost drift** | cumulative spend crosses the `costBudget` ceiling, or the projected next-phase spend would exceed it, after model-fallback has already downgraded | `costBudget` ceiling |
17
- | 5 | **Merge/rebase conflict** | Phase 6 push blocked by a conflict that requires history reconciliation | any conflict |
26
+ | 5 | **Merge/rebase conflict** | Phase 4 push blocked by a conflict that requires history reconciliation | any conflict |
18
27
 
19
28
  Trigger 2 is the key addition over the plain build-retry cap: a build can "fail differently" three times (legitimate iteration) or "fail identically" twice (stuck). Only the identical-failure case is a stall; the retry cap catches the rest.
20
29
 
@@ -22,12 +31,12 @@ Trigger 2 is the key addition over the plain build-retry cap: a build can "fail
22
31
 
23
32
  | Trigger | Evaluated by | Status |
24
33
  |---|---|---|
25
- | 2, finding half | `review-delta.mjs` exit 3 at Phase 4 Step 3.8: a blocking/important finding whose `fingerprint` (finding-fingerprint.mjs) stays in the accepted set for `identicalFindingCycles` consecutive rounds | **code** (v16.20.0) |
34
+ | 2, finding half | `review-delta.mjs` exit 3 at Phase 3 Step 3.8: a blocking/important finding whose `fingerprint` (finding-fingerprint.mjs) stays in the accepted set for `identicalFindingCycles` consecutive rounds | **code** (v16.20.0) |
26
35
  | 3 | Phase 3 re-entry item 6: the `retryCount === 3` hard-kill records the trip | **code** (v16.20.0) |
27
36
  | 2, build-error half | needs a build-log signature normaliser | documented behaviour, no script yet |
28
37
  | 1 | needs checkpoint-to-checkpoint artifact diffing | documented behaviour, no script yet |
29
38
  | 4 | belongs to `cost-budget-check.mjs` | documented behaviour, no script yet |
30
- | 5 | Phase 6 push | documented behaviour, no script yet |
39
+ | 5 | Phase 4 push | documented behaviour, no script yet |
31
40
 
32
41
  State shape: `state.circuitBreaker = {tripped, trigger, detail, checkpoint: {phase, step, iteration}, trippedAt, counters: {identicalFindingCycles, reworkCycles}}` (`schemas/agent-state.schema.json`). The per-round classification the finding half reads lives in `state.reviewIterations[i].delta` (`new`, `stillPresent`, `resolved`, `downgraded`, `recurrence`, `plateau`). `smoke-autopilot-circuit-breaker.sh` asserts the schema fields, the scripts and the phase wiring, not only this prose.
33
42
 
@@ -43,3 +52,64 @@ State shape: `state.circuitBreaker = {tripped, trigger, detail, checkpoint: {pha
43
52
  Autopilot's contract is "no interaction on the happy path." The circuit-breaker fires only off the happy path, where continuing unattended is the *less* safe choice: repeating a proven-broken action, or spending past a ceiling the user set, is not autonomy, it is a runaway. Halting with a precise reason is cheaper than the tokens (and trust) a silent loop burns.
44
53
 
45
54
  Inspired by autonomous-loop operators that gate on explicit stop conditions (no-progress, identical-stack-trace repetition, cost drift, conflict blocking), adapted to this pipeline's phase/rework/state model.
55
+
56
+ ## The runner-level breaker (a different failure, a different layer)
57
+
58
+ Everything above is the IN-RUN breaker: one run, one state file, triggers that
59
+ mean "this task is not converging". It cannot see the failure a server actually
60
+ hits, which is not about any one task.
61
+
62
+ An expired token, a `claude` binary that no longer launches, a full disk: every
63
+ item fails identically, a few minutes apart, until the queue is empty. Nothing
64
+ above notices, because each individual run failed for its own apparently local
65
+ reason. What is left afterwards is a ledger of failures with no indication of
66
+ which came first and no queue to resume.
67
+
68
+ `autopilot-runner.mjs` therefore refuses to take NEW work after
69
+ `MA_AP_BREAKER_LIMIT` consecutive attempts produced nothing (default 3, zero
70
+ disables). The count is CONSECUTIVE and read from `attempted.jsonl` rather than
71
+ kept as a counter, because a counter is a second piece of state that can
72
+ disagree with the ledger - and the ledger is what a person reads when they ask
73
+ why the runner stopped.
74
+
75
+ Two outcomes are deliberately not failures:
76
+
77
+ - `needs-input` is a run that did its work and is waiting for a person. That is
78
+ the system behaving correctly, and counting it would stop a machine whose only
79
+ problem is that somebody has not answered yet.
80
+ - `blocked-*` items never ran at all - arming refused them - so they say nothing
81
+ about whether a run would have worked.
82
+
83
+ When it opens, the reason goes into `queue.json.blockedReason` (so
84
+ `autopilot-status` shows it rather than reporting an idle queue), a
85
+ `breaker-open` line goes into `ticks.jsonl`, and the log names the outcome that
86
+ opened it plus the override.
87
+
88
+ It is HALF-OPEN rather than latched, and that distinction is load-bearing. The
89
+ breaker returns before the only writer of `attempted.jsonl`, so a version that
90
+ simply stopped could never record another attempt, the consecutive count could
91
+ never fall, and the runner would be off for good - while its own log said
92
+ "until an attempt succeeds". After `MA_AP_BREAKER_COOLDOWN_SEC` (30 min by
93
+ default) one probe is let through: a recovered machine succeeds and the count
94
+ resets on its own, and a still-broken one fails, becomes the newest attempt, and
95
+ closes the breaker for another cooldown. One wasted run per cooldown is the
96
+ price of not needing a person to notice.
97
+
98
+ ## Two more things a long-running runner needs
99
+
100
+ **The log.** launchd appends the runner's stdout to `runner.log` forever. On a
101
+ laptop that file is never read; on a machine ticking every few minutes for
102
+ months it is what fills the disk, and a full disk stops the runs the log existed
103
+ to record. Each tick truncates it past `MA_AP_LOG_MAX_BYTES` (default 5MB),
104
+ keeping the last 256KB in `runner.log.1`.
105
+
106
+ Truncating the same inode is the mechanism, not renaming: launchd holds the file
107
+ open with `O_APPEND`, so a rename leaves that descriptor writing into the renamed
108
+ file while the new one stays empty forever. That is the rotation that looks
109
+ correct and silently stops logging.
110
+
111
+ **Telemetry.** `ticks.jsonl` gets one JSON line per tick - what the tick did,
112
+ what it took, what came out, how many consecutive failures preceded it. The
113
+ human log answers "what happened just now"; this answers "how has this been
114
+ behaving for a week", which is the question a server raises and prose cannot be
115
+ asked.
@@ -1,8 +1,8 @@
1
- ## Code Graph (Phase 1 Step 2.6 + Phase 7 Step 3)
1
+ ## Code Graph (Phase 1 Step 2.6 + Phase 5 Step 3)
2
2
 
3
3
  <!-- toc -->
4
4
  - [Phase 1 Step 2.6 - query before dispatching Explore](#phase-1-step-26---query-before-dispatching-explore)
5
- - [Phase 7 Step 3 - refresh after the branch changed code](#phase-7-step-3---refresh-after-the-branch-changed-code)
5
+ - [Phase 5 Step 3 - refresh after the branch changed code](#phase-5-step-3---refresh-after-the-branch-changed-code)
6
6
  - [What the report answers that a file-level view cannot](#what-the-report-answers-that-a-file-level-view-cannot)
7
7
  - [The graph is drawable, and one place already asks for it](#the-graph-is-drawable-and-one-place-already-asks-for-it)
8
8
  <!-- /toc -->
@@ -10,7 +10,7 @@
10
10
  A deterministic, LLM-free map of what a repo declares and what refers to what,
11
11
  written to `~/.claude/knowledge/<project>/code-graph.json`. Gated by
12
12
  `prefs.global.codeGraph.enabled` (default `false`); with it off, Phase 1 and
13
- Phase 7 behave exactly as they did before. Design, trade and measurements:
13
+ Phase 5 behave exactly as they did before. Design, trade and measurements:
14
14
  `docs/adr/0010-own-code-graph.md`.
15
15
 
16
16
  ### Phase 1 Step 2.6 - query before dispatching Explore
@@ -38,7 +38,7 @@ at a graph that was never built. Pass it.
38
38
 
39
39
  Build when there is no graph at all, and when `baseCommit` no longer matches
40
40
  HEAD. The missing case is the one that matters in practice: `architecture.md`
41
- and its siblings are written in Phase 7, which is the phase a run is least
41
+ and its siblings are written in Phase 5, which is the phase a run is least
42
42
  likely to reach, so a repo can have a long history of tasks and an empty
43
43
  knowledge directory. The graph must not inherit that. `--status` exits 1 on a
44
44
  missing file, and that exit means build, not skip.
@@ -56,7 +56,7 @@ already spells out. Measured on a 4,300-file Swift app at a fixed 30k retrieval
56
56
  budget, it roughly doubled coverage at under half the cost on domain-word
57
57
  questions and lost narrowly to `grep -lw` on exact type names.
58
58
 
59
- ### Phase 7 Step 3 - refresh after the branch changed code
59
+ ### Phase 5 Step 3 - refresh after the branch changed code
60
60
 
61
61
  Runs when `prefs.global.codeGraph.enabled` and `prefs.global.codeGraph.autoRefresh`
62
62
  are both true.
@@ -0,0 +1,93 @@
1
+ # cost analysis - projection, anomaly, burn, diff
2
+
3
+ `cost-budget-check.mjs` answers one question, live, for one run: is THIS task
4
+ about to cross its ceiling. Two ways a budget actually empties are invisible to
5
+ it. A slow drift never trips any single run's ceiling. And one pathological
6
+ session can burn a week's worth in an hour while every individual run stays
7
+ comfortably under its cap.
8
+
9
+ `cost-analyze.mjs` answers the other four questions off one series.
10
+
11
+ | Sub-command | Question |
12
+ |---|---|
13
+ | `projection` | at this rate, what does a week / a month / a quarter cost, and when does a stated budget run out |
14
+ | `anomaly` | which day or session is out of family |
15
+ | `burn` | is the last day accelerating against the preceding week |
16
+ | `diff` | two saved snapshots, side by side |
17
+
18
+ ```bash
19
+ node "$HOME/.claude/scripts/cost-analyze.mjs" projection --days 14 [--monthly-usd 200]
20
+ node "$HOME/.claude/scripts/cost-analyze.mjs" anomaly --days 30 [--by day|session] [--threshold 3.5]
21
+ node "$HOME/.claude/scripts/cost-analyze.mjs" burn [--factor 3]
22
+ node "$HOME/.claude/scripts/cost-analyze.mjs" --save [--days 30]
23
+ node "$HOME/.claude/scripts/cost-analyze.mjs" diff [a.json b.json]
24
+ ```
25
+
26
+ Exit 0 means nothing to report, 10 means a finding, 2 is a usage error. Every
27
+ sub-command takes `--json`, and the human output and the JSON are rendered from
28
+ the same numbers.
29
+
30
+ ## Where the numbers come from, and what that makes them worth
31
+
32
+ The per-run token accumulators in `tracker-state.json` are the pipeline's own
33
+ ledger, and they are the obvious source. They are also empty: the schema has
34
+ `tokens_in` / `tokens_out` / `tokens_cached` per phase, they are written only
35
+ when a phase reports them, and on a real machine no tracker carries them at all.
36
+
37
+ The series that does exist is the host's own transcripts, one usage record per
38
+ assistant turn, under `~/.claude/projects/<slug>/<session>.jsonl`. So that is
39
+ what this reads, and three consequences follow that are stated here rather than
40
+ discovered later:
41
+
42
+ - It covers everything Claude Code did on this machine, not only pipeline runs.
43
+ - The figures are ESTIMATES priced from `cost-table.json`, at LIST price. On a
44
+ subscription they are the right number for comparing one day to another and
45
+ the wrong number to call a bill. The human output says so on every run.
46
+ - A host with no such transcripts (Copilot, Codex) reports `UNMEASURED`, never
47
+ zero. Zero would read as "you spent nothing" when it means "I could not see".
48
+
49
+ Records whose model is not in `cost-table.json` are COUNTED as unpriced and left
50
+ out of the total. Folding them in at zero would make an unknown model look free,
51
+ which is the direction that hurts.
52
+
53
+ Cache writes are priced at 1.25x input, derived rather than read from the table,
54
+ because `cost-table.json` carries a cache READ rate only. Dropping the component
55
+ would make every long-context session look cheap.
56
+
57
+ ## Why the median and the MAD, not the mean
58
+
59
+ The single expensive session `anomaly` exists to find is also the observation
60
+ that inflates a mean and a standard deviation - it hides inside the statistic
61
+ measured against it. With six ordinary days and one at fifty times the rest, a
62
+ mean-based z-score puts the outlier at about 2.3 sigma, under every usual
63
+ threshold. The median does not move for one observation, and neither does the
64
+ median absolute deviation.
65
+
66
+ The score is the modified z-score of Iglewicz and Hoaglin, `0.6745 * (x - median)
67
+ / MAD`, flagged at 3.5.
68
+
69
+ When more than half the values are identical the MAD is zero. That is not an
70
+ error and not a licence to divide by zero: the published fallback is the mean
71
+ absolute deviation scaled by 1.253314. When THAT is zero too there is no
72
+ dispersion at all, and the honest answer is that no anomaly can be called -
73
+ which is what it says.
74
+
75
+ Fewer than five points is refused outright, with the number it needs.
76
+
77
+ ## Why projection divides by the calendar window
78
+
79
+ Total over the window, divided by the WINDOW, not by the days that happen to
80
+ have data. Dividing by active days answers a different question - what a working
81
+ day costs - and then projecting a month as thirty working days overstates it by
82
+ roughly a third.
83
+
84
+ `--monthly-usd` is optional and there is no default. Without it the projection
85
+ is reported and the exhaustion date is not, with the reason given. Inventing a
86
+ budget to compare against would produce a number that looks measured.
87
+
88
+ ## Snapshots
89
+
90
+ `--save` writes `~/.claude/state/cost/<iso>.json`: the window, the total, the
91
+ per-day rate and the per-day breakdown. `diff` compares two, by path or the last
92
+ two by default, and says when the windows differ - the per-day rate is
93
+ comparable across different windows, the total is not.
@@ -21,7 +21,7 @@ Derived generically from an upstream corporate screen-QA tool (recorded in
21
21
  catalogs and product examples are not, and were not taken - section 7 says what
22
22
  replaces them.
23
23
 
24
- Consumers: `/multi-agent:design-check` (the runner), Phase 4 review when a UI
24
+ Consumers: `/multi-agent:design-check` (the runner), Phase 3 review when a UI
25
25
  diff is under review, and `features/visual-evidence.md` when a capture has to
26
26
  prove a fix. Gate: `smoke-design-conformance.sh`.
27
27
 
@@ -1,6 +1,6 @@
1
- # Feature: Dev Critic (Phase 3 Step 3.5 Evaluator-Optimizer)
1
+ # Feature: Dev Critic (Phase 2 Step 3.5 Evaluator-Optimizer)
2
2
 
3
- **Pattern**: Anthropic's evaluator-optimizer from "Building Effective Agents" (Dec 2024). Generator (Phase 3 Dev) and critic (`agents/dev-critic.md`) loop on deterministic criteria already written in `rules/*.md` before any Phase 4 reviewer call.
3
+ **Pattern**: Anthropic's evaluator-optimizer from "Building Effective Agents" (Dec 2024). Generator (Phase 2 Dev) and critic (`agents/dev-critic.md`) loop on deterministic criteria already written in `rules/*.md` before any Phase 3 reviewer call.
4
4
 
5
5
  **Gated by `prefs.global.devCritic.enabled`** (default: `false`). When enabled, after the generator's last edit and BEFORE Phase 4:
6
6
 
@@ -25,7 +25,7 @@
25
25
 
26
26
  ## Telemetry
27
27
 
28
- Each critic call emits `dev_critic.call` with `iteration`, `pass`, `gates_failed`, `blocking`, `important`, `duration_ms`, `tokens_in/out`. Phase 7 cost rollup lists these as `phase 3.5` line items so the net saving (Phase 4 reviewer/triage calls avoided) is measurable.
28
+ Each critic call emits `dev_critic.call` with `iteration`, `pass`, `gates_failed`, `blocking`, `important`, `duration_ms`, `tokens_in/out`. Phase 5 cost rollup lists these as `phase 3.5` line items so the net saving (Phase 3 reviewer/triage calls avoided) is measurable.
29
29
 
30
30
  ## Off by default reason
31
31
 
@@ -8,6 +8,7 @@
8
8
  - [The line shape](#the-line-shape)
9
9
  - [It recommends, it never fixes](#it-recommends-it-never-fixes)
10
10
  - [Checks](#checks)
11
+ - [The server profile (`--profile=server`)](#the-server-profile---profileserver)
11
12
  <!-- /toc -->
12
13
 
13
14
  Every check `/multi-agent:doctor` can report has a `### <id>` heading here, and
@@ -152,7 +153,7 @@ after the user has already answered the pickers.
152
153
  ### identity
153
154
 
154
155
  A git identity resolves for the account the run would commit as. WARN: the run
155
- reaches Phase 6 and stops there.
156
+ reaches Phase 4 and stops there.
156
157
 
157
158
  ### hook-coverage
158
159
 
@@ -245,10 +246,54 @@ file it was writing at the time.
245
246
 
246
247
  Worktrees left under `<repo>/.worktrees/` in the repository the caller is
247
248
  standing in. WARN at five or more, or at 2 GB. A finished task removes its own
248
- worktree at PR time, but a run that stops before Phase 6 never reaches that
249
+ worktree at PR time, but a run that stops before Phase 4 never reaches that
249
250
  step and nothing else collects it: the finalizer only runs on success, and
250
251
  `gc-worktrees` only sweeps entries git has already forgotten. Each survivor is
251
252
  a full second checkout, so the total is measured in gigabytes rather than
252
253
  megabytes. SKIP outside a git repository - a project name in prefs is a name,
253
254
  not a path, and guessing checkout locations to produce a number is how a
254
255
  diagnostic starts lying.
256
+
257
+ ## The server profile (`--profile=server`)
258
+
259
+ Four checks that only run under `--profile=server`. They are ADDITIVE: a
260
+ default run is byte-for-byte what it was, and `--list-checks` reports whichever
261
+ set the invocation would actually perform, because a list that disagrees with
262
+ the run is worse than no list.
263
+
264
+ None of them may BLOCK. The blocking set is closed at six and these are
265
+ readiness, not correctness: a machine that fails all four still runs the
266
+ pipeline perfectly well with someone watching it. What they catch is the
267
+ failure mode of an unwatched run - stopping without saying so.
268
+
269
+ ### unattended-contract
270
+
271
+ Is `multi-agent-refs/unattended-contract.md` installed. It is the only place
272
+ that says which entry points honour `MULTI_AGENT_UNATTENDED=1` and what each
273
+ one resolves to. Without it an operator setting up a server has to read the
274
+ scripts to find out, which is how the wrong assumption gets made.
275
+
276
+ ### unattended-permissions
277
+
278
+ Does `settings.json` carry `permissions.allow` entries covering Bash, Edit and
279
+ Write. autopilot spawns its child with `--permission-prompts none`, but the
280
+ tools that child then calls still have to be allowed, and on a fresh machine
281
+ they are not - the run stops at the first prompt with nobody there to answer
282
+ it, printing nothing. This is the single most likely reason a server sits idle
283
+ and looks healthy.
284
+
285
+ ### scheduler
286
+
287
+ Is a multi-agent launchd agent loaded. Something has to start the work; a
288
+ server with no agent and an empty queue looks exactly like a server that has
289
+ finished everything.
290
+
291
+ ### keychain-unlock
292
+
293
+ Is a default keychain reachable from this session. Every credential the
294
+ pipeline reads lives there, and a LaunchDaemon started before login sees a
295
+ LOCKED keychain: each fetch fails, and the failures surface much later as 401s
296
+ that blame the token rather than the lock. The check is a proxy - reading a
297
+ real credential would be a network call and a side effect - so it asks whether
298
+ a login keychain is reachable at all, which is the condition a boot-time
299
+ daemon fails.