@mmerterden/multi-agent-pipeline 17.6.0 → 19.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (272) hide show
  1. package/CHANGELOG.md +310 -0
  2. package/README.md +76 -18
  3. package/README.tr.md +55 -16
  4. package/docs/adr/0002-instruction-driven-flag.md +1 -0
  5. package/docs/adr/0005-lazy-phase-docs.md +11 -1
  6. package/docs/adr/0008-installer-modularization-and-secret-leak-defense.md +1 -0
  7. package/docs/adr/0010-own-code-graph.md +1 -0
  8. package/docs/adr/0011-dormant-ci.md +25 -1
  9. package/docs/adr/0014-six-phase-consolidation.md +134 -0
  10. package/docs/adr/README.md +2 -1
  11. package/docs/architecture.md +37 -38
  12. package/docs/best-practices.md +1 -1
  13. package/docs/ecosystem.md +37 -26
  14. package/docs/engineering.md +1 -1
  15. package/docs/facts.json +45 -0
  16. package/docs/features.md +54 -53
  17. package/docs/performance.md +5 -5
  18. package/docs/recovery-guide.md +9 -9
  19. package/docs/server-readiness.md +188 -0
  20. package/docs/token-budget-history.md +3 -1
  21. package/index.js +18 -3
  22. package/install/_codex-agents.mjs +1 -1
  23. package/install/_common.mjs +42 -17
  24. package/install/_dev-only-files.mjs +8 -0
  25. package/install/_unattended-profile.mjs +113 -0
  26. package/install/index.mjs +48 -0
  27. package/install/templates/claude-hooks.json +1 -1
  28. package/install/templates/codex-instructions.md +1 -1
  29. package/install/templates/copilot-instructions.md +28 -28
  30. package/manifest.json +1065 -0
  31. package/package.json +6 -3
  32. package/pipeline/agents/dev-critic.md +3 -3
  33. package/pipeline/commands/figma-to-swiftui.md +1 -1
  34. package/pipeline/commands/multi-agent/SKILL.md +8 -8
  35. package/pipeline/commands/multi-agent/analysis/SKILL.md +9 -9
  36. package/pipeline/commands/multi-agent/autopilot/SKILL.md +7 -7
  37. package/pipeline/commands/multi-agent/channels/SKILL.md +15 -15
  38. package/pipeline/commands/multi-agent/diff-explain/SKILL.md +6 -6
  39. package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +1 -1
  40. package/pipeline/commands/multi-agent/graph/SKILL.md +1 -1
  41. package/pipeline/commands/multi-agent/help/SKILL.md +62 -62
  42. package/pipeline/commands/multi-agent/language/SKILL.md +2 -2
  43. package/pipeline/commands/multi-agent/local/SKILL.md +11 -11
  44. package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +13 -13
  45. package/pipeline/commands/multi-agent/log/SKILL.md +2 -2
  46. package/pipeline/commands/multi-agent/manual-test/SKILL.md +9 -9
  47. package/pipeline/commands/multi-agent/model/SKILL.md +69 -0
  48. package/pipeline/commands/multi-agent/refactor/SKILL.md +3 -3
  49. package/pipeline/commands/multi-agent/resume/SKILL.md +4 -4
  50. package/pipeline/commands/multi-agent/resume-local/SKILL.md +19 -17
  51. package/pipeline/commands/multi-agent/review/SKILL.md +1 -1
  52. package/pipeline/commands/multi-agent/route-off/SKILL.md +36 -0
  53. package/pipeline/commands/multi-agent/route-on/SKILL.md +74 -0
  54. package/pipeline/commands/multi-agent/route-status/SKILL.md +56 -0
  55. package/pipeline/commands/multi-agent/setup/SKILL.md +2 -2
  56. package/pipeline/commands/multi-agent/status/SKILL.md +54 -23
  57. package/pipeline/commands/multi-agent/steer/SKILL.md +2 -2
  58. package/pipeline/commands/multi-agent/sync/SKILL.md +12 -13
  59. package/pipeline/commands/multi-agent/test/SKILL.md +1 -1
  60. package/pipeline/lib/_jira-auth.sh +8 -0
  61. package/pipeline/lib/analysis-jira-write.sh +32 -0
  62. package/pipeline/lib/ask-choice.sh +13 -2
  63. package/pipeline/lib/autopilot-state.sh +8 -0
  64. package/pipeline/lib/credential-inventory.sh +1 -1
  65. package/pipeline/lib/fatal.mjs +129 -0
  66. package/pipeline/lib/fetch-fortify.sh +1 -1
  67. package/pipeline/lib/figma-mcp-refresh.sh +18 -0
  68. package/pipeline/lib/figma-screenshot.sh +18 -0
  69. package/pipeline/lib/invoked-directly.mjs +43 -0
  70. package/pipeline/lib/jira-publish.sh +42 -0
  71. package/pipeline/lib/md2confluence-v3.py +47 -0
  72. package/pipeline/lib/model-rung.sh +142 -0
  73. package/pipeline/lib/outbound-gate.mjs +175 -0
  74. package/pipeline/lib/phase-schema.mjs +88 -0
  75. package/pipeline/lib/plan-todos.sh +32 -11
  76. package/pipeline/lib/post-pr-review.sh +77 -8
  77. package/pipeline/lib/repo-hygiene.sh +8 -3
  78. package/pipeline/lib/require-jq.sh +40 -0
  79. package/pipeline/lib/route-state.sh +161 -0
  80. package/pipeline/lib/run-paths.sh +335 -0
  81. package/pipeline/multi-agent-refs/_account-picker.md +1 -1
  82. package/pipeline/multi-agent-refs/_dev-context.md +1 -1
  83. package/pipeline/multi-agent-refs/_input-parser.md +1 -1
  84. package/pipeline/multi-agent-refs/analysis/evidence.md +0 -9
  85. package/pipeline/multi-agent-refs/analysis/intake.md +1 -1
  86. package/pipeline/multi-agent-refs/analysis/locked.md +21 -22
  87. package/pipeline/multi-agent-refs/analysis/render.md +1 -1
  88. package/pipeline/multi-agent-refs/analysis/synthesis.md +12 -6
  89. package/pipeline/multi-agent-refs/android-guide.md +1 -1
  90. package/pipeline/multi-agent-refs/audit-guide.md +13 -13
  91. package/pipeline/multi-agent-refs/channels/issue-comment.md +2 -2
  92. package/pipeline/multi-agent-refs/channels/jira.md +3 -3
  93. package/pipeline/multi-agent-refs/channels/pr.md +4 -4
  94. package/pipeline/multi-agent-refs/channels/wiki.md +1 -1
  95. package/pipeline/multi-agent-refs/component-dispatch.md +3 -3
  96. package/pipeline/multi-agent-refs/cross-cli-contract.md +31 -6
  97. package/pipeline/multi-agent-refs/features/autopilot-circuit-breaker.md +74 -4
  98. package/pipeline/multi-agent-refs/features/code-graph.md +5 -5
  99. package/pipeline/multi-agent-refs/features/cost-analysis.md +93 -0
  100. package/pipeline/multi-agent-refs/features/design-conformance.md +1 -1
  101. package/pipeline/multi-agent-refs/features/dev-critic.md +3 -3
  102. package/pipeline/multi-agent-refs/features/doctor.md +47 -2
  103. package/pipeline/multi-agent-refs/features/external-context-injection.md +3 -3
  104. package/pipeline/multi-agent-refs/features/maturity-followup.md +3 -3
  105. package/pipeline/multi-agent-refs/features/model-fallback.md +5 -5
  106. package/pipeline/multi-agent-refs/features/plan-todos.md +1 -1
  107. package/pipeline/multi-agent-refs/features/repo-map.md +1 -1
  108. package/pipeline/multi-agent-refs/features/review-delta.md +3 -3
  109. package/pipeline/multi-agent-refs/features/review-multi-repo.md +1 -1
  110. package/pipeline/multi-agent-refs/features/scope-check.md +4 -4
  111. package/pipeline/multi-agent-refs/features/skill-conformance.md +2 -2
  112. package/pipeline/multi-agent-refs/features/stack-skill-routing.md +1 -1
  113. package/pipeline/multi-agent-refs/features/verify-by-test.md +4 -4
  114. package/pipeline/multi-agent-refs/features/verify.md +83 -0
  115. package/pipeline/multi-agent-refs/features/visual-evidence.md +19 -19
  116. package/pipeline/multi-agent-refs/features/worktree-finalize.md +6 -6
  117. package/pipeline/multi-agent-refs/issue-jira-triad.md +10 -10
  118. package/pipeline/multi-agent-refs/knowledge.md +11 -11
  119. package/pipeline/multi-agent-refs/multi-repo-integration-build.md +13 -13
  120. package/pipeline/multi-agent-refs/payload-contracts.md +8 -8
  121. package/pipeline/multi-agent-refs/phases/log-format.md +10 -10
  122. package/pipeline/multi-agent-refs/phases/modes.md +30 -30
  123. package/pipeline/multi-agent-refs/phases/operations.md +21 -10
  124. package/pipeline/multi-agent-refs/phases/phase-0-init.md +25 -25
  125. package/pipeline/multi-agent-refs/phases/phase-1-plan.md +599 -0
  126. package/pipeline/multi-agent-refs/phases/{phase-3-dev.md → phase-2-dev.md} +129 -49
  127. package/pipeline/multi-agent-refs/phases/{phase-4-review.md → phase-3-review.md} +225 -107
  128. package/pipeline/multi-agent-refs/phases/{phase-6-commit.md → phase-4-commit.md} +23 -23
  129. package/pipeline/multi-agent-refs/phases/{phase-7-report.md → phase-5-report.md} +29 -29
  130. package/pipeline/multi-agent-refs/phases.md +44 -48
  131. package/pipeline/multi-agent-refs/picker-contract.md +1 -1
  132. package/pipeline/multi-agent-refs/progress-contract.md +6 -6
  133. package/pipeline/multi-agent-refs/readiness-review.md +1 -1
  134. package/pipeline/multi-agent-refs/rules.md +7 -7
  135. package/pipeline/multi-agent-refs/swiftui-guide.md +2 -2
  136. package/pipeline/multi-agent-refs/tracker-contract.md +31 -32
  137. package/pipeline/multi-agent-refs/unattended-contract.md +129 -0
  138. package/pipeline/multi-agent-refs/wiki-capture.md +14 -14
  139. package/pipeline/preferences-template.json +9 -1
  140. package/pipeline/rules/outside-the-pipeline.md +1 -1
  141. package/pipeline/schemas/agent-state.schema.json +50 -50
  142. package/pipeline/schemas/analysis-output.schema.json +2 -2
  143. package/pipeline/schemas/autopilot-config.schema.json +1 -1
  144. package/pipeline/schemas/code-graph.schema.json +1 -1
  145. package/pipeline/schemas/criteria-manifest.schema.json +1 -1
  146. package/pipeline/schemas/dev-critic-output.schema.json +1 -1
  147. package/pipeline/schemas/diff-risk.schema.json +1 -1
  148. package/pipeline/schemas/migrations/prefs-2.4.0-to-2.5.0.mjs +2 -2
  149. package/pipeline/schemas/migrations/prefs-2.6.0-to-2.7.0.mjs +31 -0
  150. package/pipeline/schemas/migrations/state-2.1.0-to-2.2.0.mjs +129 -0
  151. package/pipeline/schemas/phases.json +105 -0
  152. package/pipeline/schemas/plan-todos.schema.json +5 -5
  153. package/pipeline/schemas/planning-output.schema.json +1 -1
  154. package/pipeline/schemas/prefs.schema.json +100 -56
  155. package/pipeline/schemas/reviewer-output.schema.json +3 -3
  156. package/pipeline/schemas/route-config.schema.json +74 -0
  157. package/pipeline/schemas/scope-check.schema.json +1 -1
  158. package/pipeline/schemas/test-gap.schema.json +1 -1
  159. package/pipeline/schemas/token-budget.json +12 -18
  160. package/pipeline/schemas/triage-output.schema.json +6 -6
  161. package/pipeline/scripts/README.md +3 -3
  162. package/pipeline/scripts/_code-graph.mjs +2 -2
  163. package/pipeline/scripts/_run-paths.mjs +372 -0
  164. package/pipeline/scripts/_smoke-root.sh +1 -1
  165. package/pipeline/scripts/aggregate-metrics.mjs +65 -65
  166. package/pipeline/scripts/autopilot-arming.mjs +2 -1
  167. package/pipeline/scripts/autopilot-intake.mjs +2 -1
  168. package/pipeline/scripts/autopilot-runner.mjs +206 -2
  169. package/pipeline/scripts/build-references.mjs +2 -1
  170. package/pipeline/scripts/build-stack-plugins.mjs +10 -2
  171. package/pipeline/scripts/capture-evidence.sh +7 -2
  172. package/pipeline/scripts/capture-flush.sh +8 -8
  173. package/pipeline/scripts/capture-resume.sh +3 -3
  174. package/pipeline/scripts/classify-plan-safety.mjs +3 -2
  175. package/pipeline/scripts/cost-analyze.mjs +600 -0
  176. package/pipeline/scripts/cost-budget-check.mjs +4 -12
  177. package/pipeline/scripts/council-view.mjs +2 -1
  178. package/pipeline/scripts/crush-json.mjs +2 -1
  179. package/pipeline/scripts/diff-explain.mjs +7 -10
  180. package/pipeline/scripts/diff-risk-score.mjs +2 -1
  181. package/pipeline/scripts/doctor.mjs +140 -6
  182. package/pipeline/scripts/evidence-gate.mjs +9 -3
  183. package/pipeline/scripts/feedback-send.mjs +12 -2
  184. package/pipeline/scripts/gc-abandoned.sh +32 -16
  185. package/pipeline/scripts/gc-tmp.sh +1 -1
  186. package/pipeline/scripts/gc-worktrees.sh +12 -5
  187. package/pipeline/scripts/gen-facts.mjs +175 -0
  188. package/pipeline/scripts/gen-mode-dispatch.mjs +32 -37
  189. package/pipeline/scripts/gen-ref-toc.mjs +1 -1
  190. package/pipeline/scripts/github-ssh-setup.sh +64 -7
  191. package/pipeline/scripts/graph-mermaid.mjs +4 -2
  192. package/pipeline/scripts/graph-report.mjs +1 -1
  193. package/pipeline/scripts/jira-attach.sh +1 -1
  194. package/pipeline/scripts/keychain-save.sh +101 -30
  195. package/pipeline/scripts/learn-from-transcripts.mjs +3 -2
  196. package/pipeline/scripts/learning-curve.mjs +36 -31
  197. package/pipeline/scripts/log-metric.sh +17 -4
  198. package/pipeline/scripts/make-manifest.mjs +199 -0
  199. package/pipeline/scripts/memory-save.sh +1 -1
  200. package/pipeline/scripts/migrate-prefs.mjs +24 -6
  201. package/pipeline/scripts/migrate-state.mjs +94 -4
  202. package/pipeline/scripts/phase-banner.sh +26 -22
  203. package/pipeline/scripts/phase-tracker.sh +48 -10
  204. package/pipeline/scripts/plan-coverage-gate.mjs +8 -4
  205. package/pipeline/scripts/pre-commit-check.sh +7 -0
  206. package/pipeline/scripts/pre-push-check.sh +7 -0
  207. package/pipeline/scripts/purge.sh +23 -6
  208. package/pipeline/scripts/render-agent-log-cost.sh +10 -3
  209. package/pipeline/scripts/render-cost-summary.sh +9 -2
  210. package/pipeline/scripts/render-work-summary.sh +14 -7
  211. package/pipeline/scripts/review-file-filter.mjs +5 -3
  212. package/pipeline/scripts/review-scope.mjs +2 -1
  213. package/pipeline/scripts/routine-registry.mjs +2 -1
  214. package/pipeline/scripts/run-aggregator.mjs +26 -20
  215. package/pipeline/scripts/run-metrics.mjs +4 -2
  216. package/pipeline/scripts/runs-index.mjs +353 -0
  217. package/pipeline/scripts/scorecard-snapshot.mjs +178 -0
  218. package/pipeline/scripts/search-logs.sh +18 -0
  219. package/pipeline/scripts/smoke-cross-cli-behavior.sh +6 -6
  220. package/pipeline/scripts/smoke-schema-validation.sh +26 -7
  221. package/pipeline/scripts/test-gap-scan.mjs +2 -1
  222. package/pipeline/scripts/test-integrity-gate.mjs +2 -1
  223. package/pipeline/scripts/token-budget-report.mjs +13 -2
  224. package/pipeline/scripts/triage-memory.mjs +2 -2
  225. package/pipeline/scripts/update-issue-progress.sh +56 -7
  226. package/pipeline/scripts/usage-report.mjs +12 -1
  227. package/pipeline/scripts/validate-analysis-doc.mjs +75 -18
  228. package/pipeline/scripts/validate-code-graph.mjs +6 -3
  229. package/pipeline/scripts/validate-complaint-doc.mjs +2 -1
  230. package/pipeline/scripts/validate-diff-risk.mjs +6 -3
  231. package/pipeline/scripts/validate-planning.mjs +1 -1
  232. package/pipeline/scripts/validate-reviewer.mjs +1 -1
  233. package/pipeline/scripts/validate-state.mjs +45 -5
  234. package/pipeline/scripts/validate-test-gap.mjs +6 -3
  235. package/pipeline/scripts/validate-triage.mjs +6 -4
  236. package/pipeline/scripts/verify-citations.mjs +4 -2
  237. package/pipeline/scripts/verify.mjs +327 -0
  238. package/pipeline/scripts/worktree-finalize.sh +18 -9
  239. package/pipeline/scripts/write-state.mjs +154 -15
  240. package/pipeline/skills/.skill-manifest.json +37 -21
  241. package/pipeline/skills/.skills-index.json +104 -5
  242. package/pipeline/skills/shared/README.md +15 -6
  243. package/pipeline/skills/shared/core/apple-archive-compliance/SKILL.md +2 -2
  244. package/pipeline/skills/shared/core/google-play-compliance/SKILL.md +2 -2
  245. package/pipeline/skills/shared/core/multi-agent/SKILL.md +69 -71
  246. package/pipeline/skills/shared/core/multi-agent-autopilot/SKILL.md +3 -3
  247. package/pipeline/skills/shared/core/multi-agent-channels/SKILL.md +14 -14
  248. package/pipeline/skills/shared/core/multi-agent-diff-explain/SKILL.md +5 -5
  249. package/pipeline/skills/shared/core/multi-agent-graph/SKILL.md +1 -1
  250. package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +25 -23
  251. package/pipeline/skills/shared/core/multi-agent-language/SKILL.md +2 -2
  252. package/pipeline/skills/shared/core/multi-agent-local/SKILL.md +2 -2
  253. package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +8 -8
  254. package/pipeline/skills/shared/core/multi-agent-manual-test/SKILL.md +6 -6
  255. package/pipeline/skills/shared/core/multi-agent-model/SKILL.md +71 -0
  256. package/pipeline/skills/shared/core/multi-agent-refactor/SKILL.md +3 -3
  257. package/pipeline/skills/shared/core/multi-agent-resume/SKILL.md +1 -1
  258. package/pipeline/skills/shared/core/multi-agent-resume-local/SKILL.md +7 -7
  259. package/pipeline/skills/shared/core/multi-agent-route-off/SKILL.md +39 -0
  260. package/pipeline/skills/shared/core/multi-agent-route-on/SKILL.md +76 -0
  261. package/pipeline/skills/shared/core/multi-agent-route-status/SKILL.md +59 -0
  262. package/pipeline/skills/shared/core/multi-agent-setup/SKILL.md +1 -1
  263. package/pipeline/skills/shared/core/multi-agent-status/SKILL.md +35 -11
  264. package/pipeline/skills/shared/core/multi-agent-steer/SKILL.md +2 -2
  265. package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +6 -5
  266. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/package_app.sh +4 -1
  267. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/setup_dev_signing.sh +4 -1
  268. package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/sign-and-notarize.sh +2 -1
  269. package/pipeline/skills/skills-index.md +13 -4
  270. package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +0 -263
  271. package/pipeline/multi-agent-refs/phases/phase-2-planning.md +0 -344
  272. package/pipeline/multi-agent-refs/phases/phase-5-test.md +0 -182
package/CHANGELOG.md CHANGED
@@ -14,6 +14,316 @@ Internal file-layout changes that don't affect the slash-command surface are sti
14
14
 
15
15
  ---
16
16
 
17
+ ## [19.0.0] - 2026-09-18
18
+
19
+ Major, because phase numbers are the contract and they moved. Eight phases
20
+ became six: two of the eight were doing the same work twice, and the count
21
+ itself was guarded by nothing.
22
+
23
+ Full reasoning, mapping and rejected alternatives:
24
+ [ADR-0014](./docs/adr/0014-six-phase-consolidation.md).
25
+
26
+ ### Changed
27
+
28
+ - **Six phases.** `0 Init`, `1 Plan`, `2 Dev`, `3 Review`, `4 Commit`,
29
+ `5 Report`. Analysis and Planning were already one decision - the depth
30
+ picker skipped them together and `onlyDevelop` described them as one unit -
31
+ and they are one phase now. Review Stage 1 became Dev's exit gate, which is
32
+ what removes the second build: Dev built and tee'd a log, then Review built
33
+ again, and nothing consumed the difference. The user test moved inside
34
+ Review, keeping its waiting state.
35
+ - **`pipeline/schemas/phases.json` is the phase contract.** The list existed as
36
+ eight independent copies, none derived from another. The generator, the run
37
+ index, the metrics logger and the token budget read it now, and comparison
38
+ thresholds that used to be literals (`phase >= 6` for "waiting on you", the
39
+ Short-run boundary) are named fields in it.
40
+ - **`smoke-phase-contract.sh`** is the gate that never existed. The phase count
41
+ appeared in 91 places across 40 files with nothing holding any of them; the
42
+ command count, the jq count and the persona count all had gates. It derives
43
+ the count from the contract and checks the generator's output, the token
44
+ budget, the state-schema bounds, every shipped surface that states a count,
45
+ the progress fractions in sample output and the named thresholds. Verified
46
+ both ways: a deliberately wrong contract fails it.
47
+ - **`smoke-no-mcp-in-dev-phases.sh` keeps `phase >= 2`, and the gate now
48
+ asserts that it is unchanged.** Figma MCP was reachable only in Analysis (1);
49
+ Analysis is inside Plan (1). The permitted set `{0, 1}` is identical either
50
+ way, so the threshold surviving a renumbering is a result of the mapping
51
+ rather than an oversight - and an edit that "corrects" it would widen access.
52
+ - **Analysis has one pipeline.** Lite mode is removed. It chose sections from a
53
+ fixed list scored on three signals while Locked 2 chooses them from evidence,
54
+ and the two disagreed in both directions: a small feature with rich business
55
+ rules lost Section 15 because it was not on the list, and a feature with no
56
+ API contract kept Section 9 because it was. 37 Locked decisions became 36.
57
+ - **Locked 2's numbering clause was wrong and is corrected.** It said numbering
58
+ re-flows `1..N`; Locked 30 threads ids across the document BY NUMBER
59
+ (`Section 15.1`, `Section 4.4`), so a re-flowed document sends every one of
60
+ those references to the wrong section. No emitted document ever re-flowed.
61
+ A rendered section keeps its canonical number, gaps included, and the
62
+ validator now checks what actually holds: numbers inside the template range,
63
+ ascending, no repeats.
64
+
65
+ ### Added
66
+
67
+ - **`/multi-agent:model`** turns the top rung on or off AND realigns
68
+ `costBudget.pricingModel` in the same write. The switch existed; the command
69
+ did not, and the pricing field it must move with was left to the user to
70
+ remember. It reports what the switch means on the host it runs on: a live
71
+ switch on Claude Code, a status report on Copilot CLI and Codex CLI.
72
+ - **`/multi-agent:route-on` · `:route-off` · `:route-status`** - policy-driven
73
+ model routing, shipping disabled. `scope` has no `host-session` member and
74
+ the schema enforces that: rewriting the host's base URL would route the
75
+ user's whole session, including work unrelated to this pipeline. `route-off`
76
+ keeps the rules, so `route-on` does not re-ask. `route-status` prints the
77
+ honest limit every time - a subagent cannot be sent to a non-Anthropic model,
78
+ because subagent dispatch belongs to the host.
79
+ - **`docs/facts.json`**, generated by `pipeline/scripts/gen-facts.mjs`: the
80
+ phase, command, skill and tool counts, derived rather than written. The
81
+ website read its own copies and said "8 faz + 51 komut" while the repo had
82
+ six phases and sixty commands. The tool count is asked of the toolkit's own
83
+ `tools/list` rather than counted out of its source, because three tool
84
+ families live in modules the main file only spreads in - a regex over
85
+ `index.js` returns 2 when the answer is 99.
86
+ - **`prefs.schema.json` moves to 2.7.0, and the shipped template moves with
87
+ it.** The template had been left at 2.6.0, which
88
+ `smoke-schema-validation.sh` catches by design: a fresh install that starts
89
+ behind the migration target makes an old entry in the migrator's accepted set
90
+ load-bearing purely to rescue the template. 2.7.0 removes the Lite value from
91
+ `analysisPhase.mode` and declares `global.modelRouting`, which the template
92
+ now ships explicitly disabled rather than leaving absent - a default that is
93
+ written down is one a reader can find.
94
+ - **The description-surface ceiling moves 86,500 -> 88,000**, and this is where
95
+ that has to be said. Four commands with a shared/core twin each is eight
96
+ descriptions; the average held at 316 against its own 420 ceiling, which is
97
+ the condition the gate's convention names for a raise rather than a trim -
98
+ the surface grew because there are more commands, not wordier ones. The
99
+ fixed per-run load was a different answer: it went 122 bytes over its 60,000
100
+ ceiling and the bytes were reclaimed from prose rather than the ceiling
101
+ raised, which is what that gate's message asks for in as many words.
102
+ - **`smoke-six-phase-run.sh`** drives phase-tracker.sh through a synthetic run
103
+ and asserts what a live run would show: six tiles named from the contract, no
104
+ tile above 5, the `Phase 2 Dev` line shape, a sub-step that registers under
105
+ its parent phase rather than as a seventh tile, a token count and a start
106
+ timestamp for the cost and elapsed suffixes, and a state file that validates
107
+ at 0..5 while `currentPhase: 6` is rejected. It says plainly what it does not
108
+ cover: that Review does not build a second time is an assertion about a model
109
+ following a document, and only a live run's `.build.log` mtime can show it.
110
+ Verified both ways - a seventh phase in the contract fails it.
111
+ - **A facts gate on the website** (`tests/facts-consistency.test.ts`). It does
112
+ not check that the copied `facts.json` is fresh - CI has no pipeline
113
+ checkout, that is `sync-facts.mjs --check` on a machine that does. It checks
114
+ what actually failed: that no component states a phase count or a phase
115
+ number that disagrees with the contract. Copy may say "6 phases"; it may not
116
+ say a different number. Verified both ways.
117
+ - **`metrics.jsonl` carries `phaseSchema`.** The file is append-only across a
118
+ renumbering, so `phase: 3` means Dev in a pre-v19 row and Review in a post-v19
119
+ one. Lines without the field are generation 1. `pipeline/lib/phase-schema.mjs`
120
+ resolves both, and the two aggregators that compared phase numbers to literals
121
+ go through it.
122
+
123
+ ### Migration
124
+
125
+ - **`state-2.1.0-to-2.2.0.mjs`** maps `0→0, 1→1, 2→1, 3→2, 4→3, 5→3, 6→4, 7→5`.
126
+ Two sources can collide onto one `phases{}` key: furthest-along status wins,
127
+ `retryCount` takes the MAX (the schema caps it at 3, so a sum would emit an
128
+ invalid state), `files[]` union, earliest start, latest finish. It also
129
+ repairs four defects the live corpus already carried - `completed` →
130
+ `complete`, `awaiting-user-test-main-checkout` → `awaiting_input`, and
131
+ explicit defaults for a missing `currentPhase` or `status`. Measured on the
132
+ 62 real state files: **52 valid before, 62 after.**
133
+ - **`prefs-2.6.0-to-2.7.0.mjs`** rewrites `analysisPhase.mode` from `auto` or
134
+ `lite` to `full`. The key is kept rather than deleted, so a file that set it
135
+ stays valid.
136
+
137
+ ### Fixed
138
+
139
+ - `migrate-prefs.mjs` read its target version from a literal that had drifted
140
+ behind the schema. It reads the schema now, as do the two gates that were
141
+ checking against their own copies of it.
142
+ - `phases.md` carried a second token-budget table whose total said 17,000 while
143
+ the enforced file said 63,150 - wrong by a factor of four, for most of the
144
+ project's life, guarded by nothing. The numbers are gone; the enforced source
145
+ is named instead.
146
+ - `validate-analysis-doc.mjs` gated one half of Locked 2 and not the other. The
147
+ one emitted document available rendered `1..10, 12, 16, 20, 21` and nothing
148
+ looked.
149
+ - Eight pieces of dead code on the website, found by the linter rather than by
150
+ grep - which had already been wrong three times about this repo.
151
+ - **The phase bound is generation-aware, and it had to be.** Tightening
152
+ `currentPhase` to 0..5 marked every pre-v19 run log invalid - a run that
153
+ finished at phase 7 in October was correct when it was written, and a
154
+ validator that calls correct history invalid is one people learn to ignore.
155
+ A file stamped 2.2.0 or later is bounded 0..5; anything older, including the
156
+ files that predate stamping entirely, is bounded 0..7 and says so in the
157
+ error text. The same decision `metrics.jsonl` got: label the generation, do
158
+ not rewrite history. The tightening still bites where it matters - phase 7 on
159
+ a file claiming 2.2.0 is exactly what a skipped migration produces, and that
160
+ is rejected. Measured on the live corpus: 2 valid of 16 before, 11 of 16
161
+ after, and the 5 that remain were already invalid for reasons the plan had
162
+ recorded (no `currentPhase` at all, non-object phase values).
163
+ - `validate-state.mjs` did not check `retryCount`. The schema caps it at 3 and
164
+ four documents call 3 a hard kill, so `retryCount: 4` was a state that every
165
+ document forbade and every validator accepted. The bound is checked now. The
166
+ limit is named rather than overstated: this closes the validation boundary,
167
+ it does not stop the loop - that stays prose.
168
+ - **`phase-banner.sh` still had eight labels**, and it is the banner every
169
+ phase prints. Its table is bare words - `en:4) echo "Review"` - so all three
170
+ sweeps walked past it: they looked for `Phase 4 Review`, `4:Review` and
171
+ `Phase 4: Review`, and none of those spellings appear in it. It is now six
172
+ labels in both languages, and `smoke-phase-contract.sh` check 14 compares
173
+ every one against the contract and rejects a label above the last id, so the
174
+ one copy of the phase list that nothing derives is at least checked.
175
+ - A third spelling of a phase reference, `Phase N: Name`, which the first sweep
176
+ could not see: its rules matched `Phase 3 Dev` and the tracker tuple `3:Dev`,
177
+ and the colon form sits between them. It had left the canonical label table in
178
+ `skills/shared/core/multi-agent/SKILL.md` reading eight rows with six-phase
179
+ labels, and `/multi-agent:local` listing both a Phase 4 Review and a Phase 4
180
+ Commit. Fifteen files, corrected by name match rather than by number.
181
+ - The golden-task fixtures are named for the phases that produce them, so they
182
+ moved too: `phase-2-plan.json` -> `phase-1-plan.json`, `phase-4-review.json`
183
+ -> `phase-3-review.json`, `phase-4-triage.json` -> `phase-3-triage.json`.
184
+ - A doc sweep of 166 files in the repo and 68 in the source tree, none of which
185
+ the eight-phase plan had listed. Release history is deliberately excluded:
186
+ `CHANGELOG`, the `ROADMAP` "Previous Release" sections and
187
+ `docs/token-budget-history.md` keep the numbers their versions shipped with,
188
+ and four ADRs carry a pointer to ADR-0014 instead of being rewritten, because
189
+ an ADR records what was decided rather than what is true today.
190
+
191
+ ### Removed
192
+
193
+ - **The engagement page** (`src/app/_nisan`, its API routes and its admin
194
+ panel) on the website. The 16 RSVP rows were exported before anything was
195
+ deleted and **the `rsvp_entries` table is kept** - removing code does not
196
+ remove data, and dropping the table is a separate decision.
197
+
198
+ ---
199
+
200
+ ## [18.0.0] - 2026-09-17
201
+
202
+ Major, for two behaviour changes rather than a renamed command: a run's state now
203
+ has ONE canonical directory, and four publish paths refuse to send when the leak
204
+ gate cannot be loaded.
205
+
206
+ ### Added
207
+
208
+ - **`multi-agent-pipeline verify`** answers "is this install the thing that was
209
+ published". The install is a copy, and from the moment `install.js` writes it
210
+ the two halves drift independently: an edit in the installed tree is behaviour
211
+ with no source, and a file the installer skipped is a script the docs describe
212
+ and nobody has. `manifest.json` (SHA-256 per shipped file, version, source
213
+ commit) is built at pack time by `prepack` and never committed - a manifest in
214
+ git is stale one commit after it is written, and a stale manifest reports
215
+ honest edits as tampering.
216
+
217
+ `commands/` is compared by presence, not bytes, because `install.js` rewrites
218
+ each description into the user's `outputLanguage`: measured here, all 57
219
+ command files differ and 56 of them differ by nothing else. Dev-only files are
220
+ excluded, or 252 smokes and linters read as "the installer skipped this".
221
+
222
+ What a green result proves is stated in the output's own reference: the bytes
223
+ match what the publisher recorded. Not who published them - the manifest, the
224
+ signature and the verifier travel in the same tarball, so provenance belongs
225
+ to npm's integrity field. Signing is optional, and an unverifiable signature
226
+ says so rather than claiming valid.
227
+
228
+ - **`cost-analyze.mjs`** - projection, anomaly, burn and diff. `cost-budget-check`
229
+ watches one run against one ceiling, which is blind to both ways a budget
230
+ actually empties: a drift no single run trips, and one session that burns a
231
+ week in an hour while every run stays under its cap.
232
+
233
+ The series is not where the schema says it is. `tracker-state.json` carries
234
+ per-phase token fields that are written only when a phase reports them; of 98
235
+ trackers on a working machine, zero carry any. The dense series is the host's
236
+ own transcripts, so that is what this reads - and the consequences are printed
237
+ rather than buried: it covers everything Claude Code did on the machine, the
238
+ figures are LIST-price estimates rather than a bill, and a host with no
239
+ transcripts reports UNMEASURED instead of zero.
240
+
241
+ Anomalies use the median and the MAD, because the expensive session the check
242
+ exists to find is the observation that inflates a mean and a standard
243
+ deviation - it hides inside the statistic measured against it.
244
+
245
+ - **`install --unattended`** writes a documented permission profile, and prints
246
+ it with a reason per line before writing. autopilot passes
247
+ `--permission-prompts none`, which stops Claude Code asking and grants
248
+ nothing; on a fresh machine the run stops at the first tool call with no
249
+ prompt for anyone to answer, which is indistinguishable from an empty queue. A
250
+ default install still writes no permissions at all.
251
+
252
+ - **`doctor --profile=server`** adds four checks that only matter when nobody is
253
+ at the keyboard: the unattended contract, the permission posture, the
254
+ scheduler, and the keychain. The default run is unchanged - same 17 checks,
255
+ same verdict - because a laptop told it fails a server check learns to ignore
256
+ doctor.
257
+
258
+ - **`docs/server-readiness.md`** sets up nothing and says so first. It covers the
259
+ permission posture, why the scheduler is a LaunchAgent rather than a
260
+ LaunchDaemon (the keychain is locked until login, and that failure arrives
261
+ much later wearing a 401), and which credential each phase needs.
262
+
263
+ - **`scorecard-snapshot.mjs --diff`** keeps what the scorecard said and reports
264
+ what moved. Deliberately no 0-100 score: the scorecard reports twelve measured
265
+ metrics AND four it refuses to measure, and one figure would hide both halves.
266
+
267
+ ### Changed
268
+
269
+ - **A run's state has one canonical directory.** `{project}/{taskId}/` is
270
+ canonical and the flat `{taskId}/` is read for compatibility; every reader
271
+ resolves through `lib/run-paths.sh` / `scripts/_run-paths.mjs`, so a run that
272
+ exists in both layouts is counted once.
273
+
274
+ - **Five publish paths refuse rather than send when the leak gate is missing.**
275
+ Jira comments and descriptions, PR review bodies and inline comments, the
276
+ GitHub issue progress comment, the issue-CREATION path, and Confluence page
277
+ create/update all run their outbound text through `lib/outbound-gate.mjs`
278
+ first. A missing gate file refuses; opening the gate because the gate is not
279
+ there would be the one failure mode that matters. The count is part of the
280
+ gate: a sixth publisher cannot ship without the smoke's list naming it.
281
+
282
+ - **`set -e` on the four destructive scripts, and deliberately not on the two
283
+ collectors.** `gc-abandoned`, `gc-worktrees`, `purge` and `worktree-finalize`
284
+ delete things, so the command after an unnoticed failure is the dangerous one.
285
+ `pre-push-check` runs the gates and counts failures, and `pre-commit-check` is
286
+ a hook built out of greps that are supposed to find nothing - under `-e` the
287
+ first clean detector would end the scan and report "no secrets" for a file it
288
+ never finished reading.
289
+
290
+ ### Fixed
291
+
292
+ - **A script that dies now prints one line instead of a stack dump, and releases
293
+ what it held.** `lib/fatal.mjs` catches the sync throw, the rejection nobody
294
+ awaited and the throw from inside a callback; the last two are invisible to a
295
+ try/catch around `main()`. `write-state` releases its advisory lock on the way
296
+ out, so the next writer never has to judge a lock on age alone. EPIPE is
297
+ deliberately not fatal: `runs-index.mjs --json | head` is the ordinary way to
298
+ read a large output.
299
+
300
+ - **Build junk no longer reaches an install.** A CI runner shipped 147 files
301
+ where every developer tree shipped 146, for four rounds, and the extra was
302
+ `__pycache__/*.pyc` - untracked, so no diff of the source could show it, and
303
+ copied verbatim into all three install trees. `copyDir` now filters
304
+ `__pycache__`, `*.pyc` and `.DS_Store` on every path.
305
+
306
+ - **`jq` is required rather than optional on the nine paths that publish or
307
+ decide.** A missing `jq` renders as empty DATA and the work carries on with
308
+ it; those nine now exit 3.
309
+
310
+ - **The autopilot runner survives a month unwatched.** `runner.log` is truncated
311
+ in place past 5MB (renaming it leaves launchd's `O_APPEND` descriptor writing
312
+ into the renamed file while the new one stays empty - the rotation that looks
313
+ right and silently stops logging), three consecutive empty attempts stop new
314
+ work being taken, and one JSON line per tick goes to `ticks.jsonl`.
315
+
316
+ - **`github-ssh-setup.sh` no longer waits forever on a headless machine**, and
317
+ `write-state.mjs` gained `--if-rev=<n>` compare-and-swap so a second writer
318
+ cannot silently overwrite the first.
319
+
320
+ - **A direct-run guard that compared `import.meta.url` to `file://${argv[1]}`**
321
+ is false whenever argv[1] is not already resolved - a `/var` path that
322
+ resolves to `/private/var`, or the symlink `install --link` writes. The script
323
+ then does nothing at all, silently.
324
+
325
+ ---
326
+
17
327
  ## [17.6.0] - 2026-09-15
18
328
 
19
329
  ### Added
package/README.md CHANGED
@@ -8,16 +8,16 @@
8
8
 
9
9
  🇹🇷 Türkçe: [README.tr.md](./README.tr.md)
10
10
 
11
- An 8-phase AI development pipeline for **Claude Code**, **Copilot CLI** and **Codex CLI**. Drives a Jira issue or GitHub URL to a merged PR in one command - analysis → plan → TDD → review → test → commit → PR - with multi-repo orchestration, a plan-approval gate, CLI-aware parallel review, and store-compliance checks. Component and Figma-to-code work is dispatched to the per-stack marketplace plugins (iOS/SwiftUI, Android/Compose) rather than bundled, so component skills live in one place.
11
+ A 6-phase AI development pipeline for **Claude Code**, **Copilot CLI** and **Codex CLI**. Drives a Jira issue or GitHub URL to a merged PR in one command - analysis → plan → TDD → review → test → commit → PR - with multi-repo orchestration, a plan-approval gate, CLI-aware parallel review, and store-compliance checks. Component and Figma-to-code work is dispatched to the per-stack marketplace plugins (iOS/SwiftUI, Android/Compose) rather than bundled, so component skills live in one place.
12
12
 
13
13
  Runs natively on Claude Code, Copilot CLI and Codex CLI. macOS only. Zero runtime dependencies.
14
14
 
15
- 📐 **[Architecture diagrams](./docs/architecture.md)** - the 8-phase flow, operating modes, review/triage, Figma subphases, component layout. **[Ecosystem diagram](./docs/ecosystem.md)** - how this repo, the `multi-agent-plugins` marketplace and `multi-agent-toolkit-mcp` compose.
15
+ 📐 **[Architecture diagrams](./docs/architecture.md)** - the 6-phase flow, operating modes, review/triage, Figma subphases, component layout. **[Ecosystem diagram](./docs/ecosystem.md)** - how this repo, the `multi-agent-plugins` marketplace and `multi-agent-toolkit-mcp` compose.
16
16
 
17
17
  ### Prerequisites
18
18
 
19
19
  - **Node.js >= 20.11** - required; the pipeline's own tooling runs on it.
20
- - **`jq`** - optional but recommended. Thirteen shell helpers reach for it and skip in silence without it: cost summaries, the phase tracker's JSON reads, convention extraction, skill signing. The install prints a note when it is missing.
20
+ - **`jq`** - required for nine paths, optional for the rest. 82 shell files call it. The nine that publish or decide - the autopilot queue, Jira comments, PR reviews, issue updates, the plan file, both Figma fetchers, log search and Jira auth - now refuse with exit 3 rather than run, because a missing `jq` renders as empty DATA and the work carries on with it. Everywhere else it still degrades. The install prints a note when it is missing.
21
21
  - **`gh`** - for GitHub issue and PR work. Its built-in `--jq` is independent of the `jq` binary.
22
22
 
23
23
  ## Quick Start
@@ -33,6 +33,41 @@ npx @mmerterden/multi-agent-pipeline install --all # all three
33
33
  /multi-agent:setup # keychain token scan + git identity + default stack
34
34
  ```
35
35
 
36
+ ### No `npx` on that machine?
37
+
38
+ `npx` ships with npm, and npm ships with Node - so "npx: command not found" almost
39
+ always means Node is missing from that shell, not that anything is wrong with the
40
+ package. Check first, then pick the row that matches:
41
+
42
+ ```bash
43
+ node -v; npm -v; command -v node npm npx
44
+ ```
45
+
46
+ | What you see | What to do |
47
+ |---|---|
48
+ | nothing at all | Install Node >= 20.11: `brew install node`, or the LTS installer from nodejs.org |
49
+ | `node` works, `npx` does not | `npm i -g @mmerterden/multi-agent-pipeline` then `multi-agent-pipeline install --claude` |
50
+ | nvm is installed but the shell does not see it | `source ~/.nvm/nvm.sh && nvm use --lts`, or just open a new terminal |
51
+ | npm is ancient (< 5.2, which predates npx) | `npm i -g npm@latest`, or use the global-install row above |
52
+
53
+ And the path that needs neither `npx` nor a global install - clone and run the
54
+ installer directly:
55
+
56
+ ```bash
57
+ git clone https://github.com/mmerterden/multi-agent-pipeline.git
58
+ cd multi-agent-pipeline
59
+ node index.js install --claude # add --dry-run first to see what it would write
60
+ ```
61
+
62
+ `npm exec @mmerterden/multi-agent-pipeline install --claude` also works on any npm
63
+ 7+ without `npx` on `PATH`.
64
+
65
+ **One error that is not an npx problem.** This package declares `os: ["darwin"]`,
66
+ so npm refuses to install it anywhere else and says `npm ERR! notsup Unsupported
67
+ platform`. That is deliberate ([ADR-0012](./docs/adr/0012-macos-only.md)), not a
68
+ missing tool: every credential read shells `security`, every iOS build
69
+ `xcodebuild`, every piece of visual evidence `simctl`.
70
+
36
71
  Tool flags combine (`--claude --codex`). With no tool flag at all, the installer targets Claude Code only. Other flags: `--dry-run` (show what would be written, write nothing), `--platform=ios|android|all` (skip the stack skills you do not need), `--link` (symlink instead of copy, for local development).
37
72
 
38
73
  Run a task - the input type is auto-detected:
@@ -56,22 +91,45 @@ Update later with `/multi-agent:update`. Uninstall (tokens preserved) with `npx
56
91
 
57
92
  ## How it works
58
93
 
59
- One command runs up to 8 phases, with a gate between the risky ones. Phase 0
94
+ One command runs up to 6 phases, with a gate between the risky ones. Phase 0
60
95
  asks two questions that decide the shape of the rest - how deep the run goes
61
96
  (Full or Short) and where the branch lives (a worktree or your current
62
97
  checkout):
63
98
 
64
99
  - **0 · Init** - parse the input (Jira id / GitHub URL / free text), pick account + repo(s), fetch the issue, run a maturity check.
65
- - **1 · Analysis** - detect the stack, scan the codebase, map impact (Sonnet).
66
- - **2 · Plan** - write a task breakdown and **stop for your approval** before touching code.
67
- - **3 · Dev** - TDD: failing test → code → green, following the repo's style + the active stack skills.
68
- - **4 · Review** - deterministic gates (build / lint / test / secret-scan) must pass first, then a **CLI-aware parallel review** - Claude Code runs 3 models (Fable + Opus + Sonnet), Copilot CLI runs 3 (GPT-5.4 + Opus + Sonnet) - and a **Fable triage** keeps only actionable findings; blockers loop back to Phase 3.
69
- - **5 · Test** - build + run the suite; success is required (no faked passes).
70
- - **6 · Commit/PR** - conventional commit, push (must succeed), open a PR (`Ref: #N`, never auto-close).
71
- - **7 · Report** - technical summary + a Jira comment with test scenarios, posted through the channels layer.
100
+ - **1 · Plan** - detect the stack, scan the codebase and write the analysis document, then break it into tasks with file-level targets and **stop for your approval** before touching code. Analysis and planning were two phases until 19.0.0; the depth picker always skipped them together, because they are one decision. Codebase scanning runs on the explorer persona (Sonnet).
101
+ - **2 · Dev** - TDD: failing test → code → green, following the repo's style + the active stack skills. The phase ends at its own gate: build, lint, tests and a secret scan, run **once**. Review used to build again, and nothing consumed the difference.
102
+ - **3 · Review** - a **CLI-aware parallel review** against the logs Dev produced - Claude Code runs 3 models (Fable + Opus + Sonnet), Copilot CLI runs 3 (GPT-5.4 + Opus + Sonnet) - then a **Fable triage** keeps only actionable findings; blockers loop back to Phase 2. The optional user test lives here, keeping its waiting state.
103
+ - **4 · Commit/PR** - conventional commit, push (must succeed), open a PR (`Ref: #N`, never auto-close).
104
+ - **5 · Report** - technical summary + a Jira comment with test scenarios, posted through the channels layer. This is the one step autopilot still pauses at, in every mode.
72
105
 
73
106
  `/multi-agent:analysis` runs its own shorter chain and, since v16.12.0, reviews what it wrote before publishing it: the draft goes through the same three-reviewer set and triage as a code diff, a blocking finding returns it to synthesis with dispatch closed, and the gaps that survive are either searched, asked about, or recorded with an owner. It used to publish behind a structural validator alone.
74
107
 
108
+ ### 19.0.0: six phases, and a gate for the number
109
+
110
+ Two of the six phases were doing the same work twice. Dev built the project
111
+ and tee'd a log; Review opened by building it again. Analysis and Planning were
112
+ already one decision - the depth picker skipped them together and the state
113
+ schema described them as one unit. Six phases now, one build per run.
114
+
115
+ The other half of the change is that the count is finally guarded.
116
+ `smoke-phase-contract.sh` derives it from `pipeline/schemas/phases.json` and
117
+ holds every other copy to it: the generator's output, the token budget, the
118
+ state-schema bounds, the progress fractions in sample output, and the named
119
+ thresholds that used to be literals scattered across scripts. The phase count
120
+ appeared in 91 places across 40 files with nothing checking any of them, while
121
+ the command count, the jq count and the persona count all had gates.
122
+
123
+ Reasoning, mapping and rejected alternatives:
124
+ [ADR-0014](./docs/adr/0014-six-phase-consolidation.md).
125
+
126
+ ### 18.0.0: one state directory, and a way to ask whether your install is real
127
+
128
+ - **`multi-agent-pipeline verify`.** The install is a copy, and from the moment it is written the two halves drift independently: an edit in the installed tree is behaviour with no source, and a file the installer skipped is a script the docs describe and nobody has. `verify` compares both against a manifest built at pack time. What a green result proves is stated plainly - the bytes match what the publisher recorded, not who published them.
129
+ - **Cost, past the single run.** The per-task ceiling cannot see the two ways a budget actually empties: a drift that trips nothing, and one session that burns a week in an hour while every run stays under its cap. `cost-analyze` projects, finds days out of family by median absolute deviation, and reports acceleration - as LIST-price estimates, which it says on every run rather than in a footnote.
130
+ - **A run's state lives in one place.** `{project}/{taskId}/` is canonical, the flat layout is still read, and a run that exists in both is counted once.
131
+ - **Server readiness, entirely opt-in.** `doctor --profile=server` adds four checks that only matter when nobody is at the keyboard, `install --unattended` writes a permission profile after printing it, and the autopilot runner now survives a month unwatched. The default install writes no permissions and the default doctor run is unchanged, because a laptop told it fails a server check learns to ignore doctor.
132
+
75
133
  ### Your package manager, your hooks, your MCP surface
76
134
 
77
135
  Three smaller things in 17.6.0, each closing a gap where the pipeline assumed instead of looking:
@@ -108,8 +166,8 @@ The discipline behind all of this - bounded loops, evidence gates, token-budgete
108
166
 
109
167
  | Mode | Command | Flow |
110
168
  | --------- | ------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- |
111
- | Full | `/multi-agent "task"` | All 8 phases, interactive |
112
- | Autopilot | `/multi-agent:autopilot "task"` | 7 phases (interactive Test gate dropped), no confirmations |
169
+ | Full | `/multi-agent "task"` | All 6 phases, interactive |
170
+ | Autopilot | `/multi-agent:autopilot "task"` | 6 phases (interactive Test gate dropped), no confirmations |
113
171
  | Local | `/multi-agent:local "task"` | Full pipeline minus the interactive Test gate, current branch (no worktree) |
114
172
  | Depth | asked at Phase 0 Step 7.5 | Full (all phases) or Short (Dev → Review → Test → Commit → Report). Not a command name - `/multi-agent` and `:local` ask, both autopilot entries always run Full |
115
173
  | Ship | `/multi-agent:resume-local` | Run the review→test→commit→report tail over local work |
@@ -165,7 +223,7 @@ The widget follows the answer rather than predicting it: Phase 0 is the only til
165
223
  | `/multi-agent:review-jira` | Grade a Jira issue's readiness for development, comment the gaps |
166
224
  | `/multi-agent:review-issue` | Same grading for a GitHub issue |
167
225
  | `/multi-agent:review-analysis` | Review a written analysis document; findings cite the Locked rule they break |
168
- | `/multi-agent:diff-explain` | Map a Phase 4 triage finding back to the diff lines that caused it |
226
+ | `/multi-agent:diff-explain` | Map a Phase 3 triage finding back to the diff lines that caused it |
169
227
  | `/multi-agent:refactor` | Best-practice extraction + bug hunt + derived-skill drift + toolkit MCP research → one plan |
170
228
  | `/multi-agent:scan` | Skill security scan of local skill directories against a tiered pattern catalog |
171
229
  | `/multi-agent:prune-prompts` | Zero-base review of the always-on instruction footprint; keep / trial / delete per rule |
@@ -301,17 +359,17 @@ This enables the matching plugin (+ the shared `ai-common` plugin) in the repo's
301
359
 
302
360
  ## Tool support
303
361
 
304
- The pipeline runs natively on **Claude Code**, **Copilot CLI** and **Codex CLI** - all three install from the same `pipeline/` source and get the same 56 commands.
362
+ The pipeline runs natively on **Claude Code**, **Copilot CLI** and **Codex CLI** - all three install from the same `pipeline/` source and get the same 60 commands.
305
363
 
306
364
  | Tool | Flag | What it installs |
307
365
  | ----------- | -------------------- | ------------------------------------------------------------------------------------------------------ |
308
366
  | Claude Code | `--claude` (default) | slash commands + skills + agents + three `PreToolUse` hooks (secret scan, agent-guard, read-size gate) |
309
- | Copilot CLI | `--copilot` | instructions + 56 sub-command skills + scripts |
310
- | Codex CLI | `--codex` | one router skill + 56 specs as refs + 9 agent TOML + `AGENTS.md` block + `codex mcp add` |
367
+ | Copilot CLI | `--copilot` | instructions + 60 sub-command skills + scripts |
368
+ | Codex CLI | `--codex` | one router skill + 60 specs as refs + 9 agent TOML + `AGENTS.md` block + `codex mcp add` |
311
369
 
312
370
  Filter skills by stack with `--platform=ios\|android\|all`.
313
371
 
314
- **Why Codex gets one skill and not 56.** Codex assembles every discovered skill's name
372
+ **Why Codex gets one skill and not 60.** Codex assembles every discovered skill's name
315
373
  and description into a single prompt block and drops entries when it overflows, with no
316
374
  error. Measured on 0.145: installing one plugin that declares 142 skills surfaced only
317
375
  75 of them and evicted an unrelated user skill. So on Codex the pipeline ships a single
package/README.tr.md CHANGED
@@ -8,11 +8,11 @@
8
8
 
9
9
  🇬🇧 English: [README.md](./README.md)
10
10
 
11
- **Claude Code**, **Copilot CLI** ve **Codex CLI** için 8 fazlı bir AI geliştirme pipeline'ı. Bir Jira issue'sunu veya GitHub URL'sini tek komutla merge edilmiş bir PR'a dönüştürür - analiz → plan → TDD → review → test → commit → PR - çoklu-repo orkestrasyonu, bir plan-onay kapısı, CLI-farkında paralel review ve store-uyumluluk kontrolleriyle birlikte. Component ve Figma-to-code işleri paket içine gömülmek yerine stack başına marketplace plugin'lerine (iOS/SwiftUI, Android/Compose) devredilir, böylece component skill'leri tek bir yerde yaşar.
11
+ **Claude Code**, **Copilot CLI** ve **Codex CLI** için 6 fazlı bir AI geliştirme pipeline'ı. Bir Jira issue'sunu veya GitHub URL'sini tek komutla merge edilmiş bir PR'a dönüştürür - analiz → plan → TDD → review → test → commit → PR - çoklu-repo orkestrasyonu, bir plan-onay kapısı, CLI-farkında paralel review ve store-uyumluluk kontrolleriyle birlikte. Component ve Figma-to-code işleri paket içine gömülmek yerine stack başına marketplace plugin'lerine (iOS/SwiftUI, Android/Compose) devredilir, böylece component skill'leri tek bir yerde yaşar.
12
12
 
13
13
  Claude Code, Copilot CLI ve Codex CLI üzerinde native çalışır. Yalnızca macOS. Sıfır runtime dependency.
14
14
 
15
- 📐 **[Mimari diyagramları](./docs/architecture.md)** - 8 faz akışı, çalışma modları, review/triage, Figma subphase'leri, component yapısı. **[Ekosistem diyagramı](./docs/ecosystem.md)** - bu repo, `multi-agent-plugins` marketplace'i ve `multi-agent-toolkit-mcp`'nin nasıl bir araya geldiği.
15
+ 📐 **[Mimari diyagramları](./docs/architecture.md)** - 6 faz akışı, çalışma modları, review/triage, Figma subphase'leri, component yapısı. **[Ekosistem diyagramı](./docs/ecosystem.md)** - bu repo, `multi-agent-plugins` marketplace'i ve `multi-agent-toolkit-mcp`'nin nasıl bir araya geldiği.
16
16
 
17
17
  ### Önkoşullar
18
18
 
@@ -33,6 +33,40 @@ npx @mmerterden/multi-agent-pipeline install --all # üçü birden
33
33
  /multi-agent:setup # keychain token taraması + git kimliği + varsayılan stack
34
34
  ```
35
35
 
36
+ ### O makinede `npx` yoksa
37
+
38
+ `npx` npm ile, npm de Node ile gelir - yani "npx: command not found" neredeyse her
39
+ zaman o kabukta Node olmadığı anlamına gelir, pakette bir sorun olduğu değil. Önce
40
+ bak, sonra sana uyan satırı uygula:
41
+
42
+ ```bash
43
+ node -v; npm -v; command -v node npm npx
44
+ ```
45
+
46
+ | Gördüğün | Yapılacak |
47
+ |---|---|
48
+ | hiçbiri yok | Node >= 20.11 kur: `brew install node` ya da nodejs.org'dan LTS installer |
49
+ | `node` çalışıyor, `npx` çalışmıyor | `npm i -g @mmerterden/multi-agent-pipeline` sonra `multi-agent-pipeline install --claude` |
50
+ | nvm kurulu ama kabuk görmüyor | `source ~/.nvm/nvm.sh && nvm use --lts`, ya da yeni bir terminal aç |
51
+ | npm çok eski (< 5.2, npx'ten önceki sürümler) | `npm i -g npm@latest`, ya da üstteki global kurulum satırı |
52
+
53
+ Ne `npx` ne de global kurulum isteyen yol - klonla ve installer'ı doğrudan çalıştır:
54
+
55
+ ```bash
56
+ git clone https://github.com/mmerterden/multi-agent-pipeline.git
57
+ cd multi-agent-pipeline
58
+ node index.js install --claude # önce --dry-run ile ne yazacağını görebilirsin
59
+ ```
60
+
61
+ `npm exec @mmerterden/multi-agent-pipeline install --claude` de, `PATH`'te `npx`
62
+ olmayan her npm 7+ üzerinde çalışır.
63
+
64
+ **npx sorunu olmayan bir hata.** Bu paket `os: ["darwin"]` beyan eder; npm başka
65
+ hiçbir yerde kurmaz ve `npm ERR! notsup Unsupported platform` der. Bu kasıtlıdır
66
+ ([ADR-0012](./docs/adr/0012-macos-only.md)), eksik bir araç değil: her credential
67
+ okuması `security`, her iOS build'i `xcodebuild`, her görsel kanıt `simctl`
68
+ çağırıyor.
69
+
36
70
  Tool flag'leri birleştirilebilir (`--claude --codex`). Hiç tool flag'i verilmezse installer sadece Claude Code'u hedefler. Diğer flag'ler: `--dry-run` (ne yazılacağını gösterir, hiçbir şey yazmaz), `--platform=ios|android|all` (ihtiyacın olmayan stack skill'lerini atlar), `--link` (kopyalamak yerine symlink, lokal geliştirme için).
37
71
 
38
72
  Bir görev çalıştır - girdi tipi otomatik algılanır:
@@ -56,22 +90,27 @@ Sonra `/multi-agent:update` ile güncelle. Kaldırmak için (tokenlar korunur) `
56
90
 
57
91
  ## Nasıl çalışır
58
92
 
59
- Tek komut en fazla 8 fazı çalıştırır, riskli olanlar arasında bir kapı ile. Faz
93
+ Tek komut en fazla 6 fazı çalıştırır, riskli olanlar arasında bir kapı ile. Faz
60
94
  0 geri kalanın şeklini belirleyen iki soru sorar: koşu ne kadar derin olacak
61
95
  (Tam mı Kısa mı) ve branch nerede yaşayacak (worktree mi, mevcut checkout'un
62
96
  mu):
63
97
 
64
98
  - **0 · Init** - girdiyi ayrıştır (Jira id / GitHub URL / serbest metin), hesap + repo(lar) seç, issue'yu çek, maturity kontrolü yap.
65
- - **1 · Analysis** - stack'i tespit et, codebase'i tara, etkiyi haritala (Sonnet).
66
- - **2 · Plan** - bir görev kırılımı yaz ve koda dokunmadan önce **onayın için dur**.
67
- - **3 · Dev** - TDD: başarısız test → kod → yeşil, repo'nun stiline + aktif stack skill'lerine uyarak.
68
- - **4 · Review** - önce deterministik kapılar (build / lint / test / secret-scan) geçmeli, sonra bir **CLI-farkında paralel review** - Claude Code 3 model çalıştırır (Fable + Opus + Sonnet), Copilot CLI 3 (GPT-5.4 + Opus + Sonnet) - ve bir **Fable triage** sadece aksiyon alınabilir bulguları tutar; blocker'lar Phase 3'e geri döner.
69
- - **5 · Test** - build + suite'i çalıştır; başarı zorunlu (sahte pass yok).
70
- - **6 · Commit/PR** - conventional commit, push (başarılı olmalı), bir PR aç (`Ref: #N`, asla otomatik kapatma).
71
- - **7 · Report** - teknik özet + test senaryolarıyla bir Jira yorumu, channels katmanından gönderilir.
99
+ - **1 · Plan** - stack'i tespit et, codebase'i tara ve analiz dokümanını yaz; sonra onu dosya seviyesinde hedefleri olan görevlere böl ve koda dokunmadan önce **onayın için dur**. Analiz ve planlama 19.0.0'a kadar iki ayrı fazdı; derinlik seçici ikisini hep birlikte atlıyordu, çünkü tek bir karar. Codebase taraması explorer persona'sı üzerinde koşar (Sonnet).
100
+ - **2 · Dev** - TDD: başarısız test → kod → yeşil, repo'nun stiline + aktif stack skill'lerine uyarak. Faz kendi kapısında biter: build, lint, test ve sır taraması, **bir kez** koşar. Review eskiden ikinci kez build ediyordu ve aradaki farkı kimse okumuyordu.
101
+ - **3 · Review** - Dev'in ürettiği log'lara karşı **CLI-farkında paralel review** - Claude Code 3 model çalıştırır (Fable + Opus + Sonnet), Copilot CLI 3 (GPT-5.4 + Opus + Sonnet) - ve bir **Fable triage** sadece aksiyon alınabilir bulguları tutar; blocker'lar Phase 2'ye geri döner. Opsiyonel kullanıcı testi burada, bekleme durumunu koruyarak.
102
+ - **4 · Commit/PR** - conventional commit, push (başarılı olmalı), bir PR aç (`Ref: #N`, asla otomatik kapatma).
103
+ - **5 · Report** - teknik özet + test senaryolarıyla bir Jira yorumu, channels katmanından gönderilir. Autopilot'un her modda hâlâ durduğu tek adım bu.
72
104
 
73
105
  `/multi-agent:analysis` kendi kısa zincirini koşar ve v16.12.0'dan beri yazdığını yayınlamadan önce review ediyor: taslak, bir kod diff'iyle aynı üç-reviewer setinden ve triyajdan geçiyor, bloklayıcı bulgu dokümanı sentez fazına geri gönderip dispatch'i kapatıyor, hayatta kalan boşluklar ya aranıyor ya sana soruluyor ya da sahibiyle birlikte kayda giriyor. Önceden yalnızca yapısal bir validator'ın arkasından yayınlıyordu.
74
106
 
107
+ ### 18.0.0: tek bir durum dizini, ve kurulumun gerçekten o kurulum olup olmadığını sorma yolu
108
+
109
+ - **`multi-agent-pipeline verify`.** Kurulum bir kopyadır ve yazıldığı andan itibaren iki taraf birbirinden bağımsız kayar: kurulu ağaçtaki bir düzenleme kaynağı olmayan bir davranıştır, kurulumun atladığı dosya ise dokümanın anlattığı ama kimsede olmayan bir script. `verify` ikisini de paketleme anında üretilen bir manifest'e karşı karşılaştırır. Yeşil sonucun ne kanıtladığı açıkça yazılıdır: baytlar yayıncının kaydettiğiyle aynıdır, kimin yayınladığı değil.
110
+ - **Maliyet, tek koşunun ötesinde.** Görev başına tavan, bütçeyi gerçekten bitiren iki şeyi göremez: hiçbir koşuyu aşmayan kayma, ve her koşu tavanın altında kalırken bir saatte bir haftayı yakan tek oturum. `cost-analyze` projeksiyon yapar, medyan mutlak sapmayla aileden ayrılan günü bulur ve ivmeyi raporlar - liste fiyatından tahmin olarak, ve bunu dipnotta değil her koşuda söyler.
111
+ - **Bir koşunun durumu tek yerde.** `{project}/{taskId}/` kanonik, düz yerleşim hâlâ okunuyor, ve iki yerde birden duran koşu bir kez sayılıyor.
112
+ - **Sunucu hazırlığı, tamamen opt-in.** `doctor --profile=server` yalnızca klavyede kimse yokken anlamı olan dört kontrol ekler, `install --unattended` izin profilini yazmadan önce gösterir, autopilot runner artık bir ay gözetimsiz ayakta kalır. Varsayılan kurulum hiçbir izin yazmaz ve varsayılan doctor koşusu değişmez: sunucu kontrolünden kaldığı söylenen bir dizüstü, doctor'ı görmezden gelmeyi öğrenir.
113
+
75
114
  ### Paket yöneticisi, hook'lar ve MCP yüzeyi
76
115
 
77
116
  17.6.0'da üç küçük iş; üçü de pipeline'ın bakmak yerine varsaydığı bir yeri kapatıyor:
@@ -108,8 +147,8 @@ Bunun arkasındaki disiplin - sınırlı loop'lar, kanıt kapıları, token-büt
108
147
 
109
148
  | Mod | Komut | Akış |
110
149
  | --------- | ------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
111
- | Full | `/multi-agent "task"` | Tüm 8 faz, interaktif |
112
- | Autopilot | `/multi-agent:autopilot "task"` | 7 faz (interaktif Test kapısı atlanır), onaysız |
150
+ | Full | `/multi-agent "task"` | Tüm 6 faz, interaktif |
151
+ | Autopilot | `/multi-agent:autopilot "task"` | 6 faz (interaktif Test kapısı atlanır), onaysız |
113
152
  | Local | `/multi-agent:local "task"` | İnteraktif Test kapısı hariç tam pipeline, mevcut branch (worktree yok) |
114
153
  | Derinlik | Faz 0 Adım 7.5'te sorulur | Full (tüm fazlar) veya Short (Dev → Review → Test → Commit → Report). Komut adı değil - `/multi-agent` ve `:local` sorar, iki autopilot girişi de her zaman Full koşar |
115
154
  | Ship | `/multi-agent:resume-local` | Lokal iş üzerinde review→test→commit→report kuyruğunu çalıştır |
@@ -302,17 +341,17 @@ Bu, ilgili plugin'i (+ ortak `ai-common` plugin'ini) repo'nun `.claude/settings.
302
341
 
303
342
  ## Araç desteği
304
343
 
305
- Pipeline **Claude Code**, **Copilot CLI** ve **Codex CLI** üzerinde native çalışır - üçü de aynı `pipeline/` kaynağından kurulur ve aynı 56 komutu alır.
344
+ Pipeline **Claude Code**, **Copilot CLI** ve **Codex CLI** üzerinde native çalışır - üçü de aynı `pipeline/` kaynağından kurulur ve aynı 60 komutu alır.
306
345
 
307
346
  | Araç | Bayrak | Ne kurar |
308
347
  | ----------- | ----------------------- | ---------------------------------------------------------------------------------------------------------------- |
309
348
  | Claude Code | `--claude` (varsayılan) | slash komutları + skill'ler + agent'lar + üç `PreToolUse` hook'u (secret scan, agent-guard, okuma-boyutu geçidi) |
310
- | Copilot CLI | `--copilot` | talimatlar + 56 alt-komut skill'i + script'ler |
311
- | Codex CLI | `--codex` | bir router skill + ref olarak 56 spec + 9 agent TOML + `AGENTS.md` bloğu + `codex mcp add` |
349
+ | Copilot CLI | `--copilot` | talimatlar + 60 alt-komut skill'i + script'ler |
350
+ | Codex CLI | `--codex` | bir router skill + ref olarak 60 spec + 9 agent TOML + `AGENTS.md` bloğu + `codex mcp add` |
312
351
 
313
352
  Skill'leri stack'e göre filtrele: `--platform=ios\|android\|all`.
314
353
 
315
- **Codex neden 56 değil de tek bir skill alıyor.** Codex, keşfettiği her skill'in adını
354
+ **Codex neden 60 değil de tek bir skill alıyor.** Codex, keşfettiği her skill'in adını
316
355
  ve açıklamasını tek bir prompt bloğuna toplar ve blok taştığında girdileri hatasızca
317
356
  düşürür. 0.145 üzerinde ölçüldü: 142 skill deklare eden bir plugin kurulduğunda sadece
318
357
  75'i yüzeye çıktı ve alakasız bir kullanıcı skill'i tahliye edildi. Bu yüzden Codex'te
@@ -1,6 +1,7 @@
1
1
  # 2. `instructionDriven` flag as explicit pipeline fork
2
2
 
3
3
  **Status:** Accepted · 2025
4
+ > **Phase numbers below are the eight-phase ones.** [ADR-0014](./0014-six-phase-consolidation.md) renumbered the contract in v19.0.0 (Phase 6 Commit is now Phase 4, Phase 7 Report is now Phase 5). The decision this ADR records is unchanged; only the labels moved, and they are left as written because an ADR records what was decided.
4
5
 
5
6
  ## Context
6
7