@mmerterden/multi-agent-pipeline 19.1.3 → 20.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (138) hide show
  1. package/CHANGELOG.md +150 -7
  2. package/README.md +60 -50
  3. package/README.tr.md +55 -46
  4. package/docs/adr/0002-instruction-driven-flag.md +6 -5
  5. package/docs/adr/0005-lazy-phase-docs.md +2 -2
  6. package/docs/adr/0008-installer-modularization-and-secret-leak-defense.md +1 -0
  7. package/docs/adr/0009-claude-stack-skills-plugin-only.md +1 -1
  8. package/docs/adr/0010-own-code-graph.md +5 -4
  9. package/docs/adr/0012-macos-only.md +2 -2
  10. package/docs/adr/0013-lsp-code-intelligence.md +2 -2
  11. package/docs/adr/0014-six-phase-consolidation.md +9 -9
  12. package/docs/adr/0015-one-pipeline-no-depth-answer.md +83 -0
  13. package/docs/adr/0016-the-run-shape-is-asked-not-typed.md +69 -0
  14. package/docs/adr/README.md +18 -16
  15. package/docs/architecture.md +2 -2
  16. package/docs/ecosystem.md +8 -9
  17. package/docs/facts.json +5 -8
  18. package/docs/features.md +4 -5
  19. package/docs/token-budget-history.md +1 -1
  20. package/install/_common.mjs +14 -6
  21. package/install/_mcp-register.mjs +1 -1
  22. package/install/_plugin-skills.mjs +3 -4
  23. package/install/copilot.mjs +5 -5
  24. package/install/templates/copilot-instructions.md +7 -16
  25. package/manifest.json +135 -138
  26. package/package.json +1 -1
  27. package/pipeline/commands/multi-agent/SKILL.md +6 -8
  28. package/pipeline/commands/multi-agent/analysis/SKILL.md +2 -0
  29. package/pipeline/commands/multi-agent/analysis-jira/SKILL.md +2 -0
  30. package/pipeline/commands/multi-agent/analysis-resolve/SKILL.md +2 -0
  31. package/pipeline/commands/multi-agent/autopilot/SKILL.md +2 -0
  32. package/pipeline/commands/multi-agent/autopilot-on/SKILL.md +2 -0
  33. package/pipeline/commands/multi-agent/autopilot-status/SKILL.md +1 -1
  34. package/pipeline/commands/multi-agent/build-optimize/SKILL.md +2 -0
  35. package/pipeline/commands/multi-agent/channels/SKILL.md +1 -1
  36. package/pipeline/commands/multi-agent/create-jira/SKILL.md +2 -0
  37. package/pipeline/commands/multi-agent/design-check/SKILL.md +1 -1
  38. package/pipeline/commands/multi-agent/forget/SKILL.md +2 -0
  39. package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +4 -2
  40. package/pipeline/commands/multi-agent/help/SKILL.md +21 -27
  41. package/pipeline/commands/multi-agent/ios-coding-standard/SKILL.md +5 -4
  42. package/pipeline/commands/multi-agent/issue/SKILL.md +2 -0
  43. package/pipeline/commands/multi-agent/jira/SKILL.md +2 -0
  44. package/pipeline/commands/multi-agent/language/SKILL.md +2 -0
  45. package/pipeline/commands/multi-agent/prune-logs/SKILL.md +2 -0
  46. package/pipeline/commands/multi-agent/purge/SKILL.md +2 -0
  47. package/pipeline/commands/multi-agent/resume/SKILL.md +177 -48
  48. package/pipeline/commands/multi-agent/save/SKILL.md +2 -0
  49. package/pipeline/commands/multi-agent/setup/SKILL.md +4 -4
  50. package/pipeline/commands/multi-agent/stack/SKILL.md +2 -0
  51. package/pipeline/commands/multi-agent/sync/SKILL.md +6 -7
  52. package/pipeline/commands/multi-agent/test-screenshots/SKILL.md +2 -0
  53. package/pipeline/commands/multi-agent/uninstall/SKILL.md +2 -0
  54. package/pipeline/commands/sim-test.md +4 -4
  55. package/pipeline/lib/repo-hygiene.sh +1 -1
  56. package/pipeline/multi-agent-refs/analysis/locked.md +2 -2
  57. package/pipeline/multi-agent-refs/analysis/render.md +1 -1
  58. package/pipeline/multi-agent-refs/analysis/resolve.md +1 -1
  59. package/pipeline/multi-agent-refs/analysis/synthesis.md +1 -1
  60. package/pipeline/multi-agent-refs/analysis-template.md +1 -1
  61. package/pipeline/multi-agent-refs/channels/jira.md +8 -8
  62. package/pipeline/multi-agent-refs/component-dispatch.md +0 -8
  63. package/pipeline/multi-agent-refs/cross-cli-contract.md +10 -11
  64. package/pipeline/multi-agent-refs/features/base-branch-evidence.md +2 -2
  65. package/pipeline/multi-agent-refs/features/external-context-injection.md +2 -0
  66. package/pipeline/multi-agent-refs/features/review-delta.md +1 -1
  67. package/pipeline/multi-agent-refs/features/review-multi-repo.md +3 -3
  68. package/pipeline/multi-agent-refs/features/scope-check.md +1 -1
  69. package/pipeline/multi-agent-refs/features/skill-conformance.md +1 -1
  70. package/pipeline/multi-agent-refs/features/stack-skill-routing.md +1 -1
  71. package/pipeline/multi-agent-refs/features/visual-evidence.md +2 -1
  72. package/pipeline/multi-agent-refs/features/worktree-finalize.md +1 -1
  73. package/pipeline/multi-agent-refs/generate-issue.md +2 -0
  74. package/pipeline/multi-agent-refs/issue-jira-triad.md +2 -0
  75. package/pipeline/multi-agent-refs/keychain.md +2 -0
  76. package/pipeline/multi-agent-refs/knowledge.md +0 -7
  77. package/pipeline/multi-agent-refs/outside-the-pipeline.md +1 -1
  78. package/pipeline/multi-agent-refs/payload-contracts.md +1 -1
  79. package/pipeline/multi-agent-refs/phases/modes.md +32 -108
  80. package/pipeline/multi-agent-refs/phases/operations.md +3 -1
  81. package/pipeline/multi-agent-refs/phases/phase-0-init.md +23 -42
  82. package/pipeline/multi-agent-refs/phases/phase-1-plan.md +10 -21
  83. package/pipeline/multi-agent-refs/phases/phase-2-dev.md +13 -44
  84. package/pipeline/multi-agent-refs/phases/phase-3-review.md +19 -22
  85. package/pipeline/multi-agent-refs/phases/phase-4-commit.md +6 -6
  86. package/pipeline/multi-agent-refs/phases/phase-5-report.md +2 -2
  87. package/pipeline/multi-agent-refs/phases.md +9 -11
  88. package/pipeline/multi-agent-refs/progress-contract.md +1 -1
  89. package/pipeline/multi-agent-refs/readiness-review.md +2 -0
  90. package/pipeline/multi-agent-refs/rules.md +2 -2
  91. package/pipeline/multi-agent-refs/tracker-contract.md +9 -40
  92. package/pipeline/multi-agent-refs/wiki-capture.md +3 -2
  93. package/pipeline/preferences-template.json +2 -2
  94. package/pipeline/rules/figma-pipeline.md +1 -1
  95. package/pipeline/schemas/agent-state.schema.json +5 -10
  96. package/pipeline/schemas/migrations/prefs-2.7.0-to-2.8.0.mjs +33 -0
  97. package/pipeline/schemas/phases.json +3 -24
  98. package/pipeline/schemas/prefs.schema.json +5 -5
  99. package/pipeline/scripts/autopilot-runner.mjs +6 -7
  100. package/pipeline/scripts/build-references.mjs +3 -3
  101. package/pipeline/scripts/build-stack-plugins.mjs +1 -1
  102. package/pipeline/scripts/bulk-read.sh +6 -4
  103. package/pipeline/scripts/cost-table.json +1 -1
  104. package/pipeline/scripts/doctor.mjs +4 -4
  105. package/pipeline/scripts/gc-refs.sh +1 -1
  106. package/pipeline/scripts/gen-mode-dispatch.mjs +11 -41
  107. package/pipeline/scripts/learnings-ledger.mjs +1 -1
  108. package/pipeline/scripts/match-skills.mjs +4 -4
  109. package/pipeline/scripts/memory-load.sh +3 -3
  110. package/pipeline/scripts/migrate-prefs.mjs +18 -17
  111. package/pipeline/scripts/phase-tracker.sh +2 -2
  112. package/pipeline/scripts/phase0-exit-gate.mjs +1 -1
  113. package/pipeline/scripts/plan-coverage-gate.mjs +6 -6
  114. package/pipeline/scripts/run-aggregator.mjs +3 -3
  115. package/pipeline/scripts/runs-index.mjs +7 -7
  116. package/pipeline/scripts/scope-check-gate.mjs +1 -1
  117. package/pipeline/scripts/smoke-schema-validation.sh +9 -12
  118. package/pipeline/scripts/usage-report.mjs +5 -7
  119. package/pipeline/scripts/validate-analysis-doc.mjs +3 -3
  120. package/pipeline/scripts/worktree-finalize.sh +2 -2
  121. package/pipeline/scripts/write-state.mjs +22 -11
  122. package/pipeline/skills/.skill-manifest.json +9 -21
  123. package/pipeline/skills/.skills-index.json +6 -39
  124. package/pipeline/skills/shared/README.md +5 -8
  125. package/pipeline/skills/shared/core/multi-agent/SKILL.md +10 -13
  126. package/pipeline/skills/shared/core/multi-agent-autopilot-status/SKILL.md +1 -1
  127. package/pipeline/skills/shared/core/multi-agent-channels/SKILL.md +4 -5
  128. package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +13 -16
  129. package/pipeline/skills/shared/core/multi-agent-ios-coding-standard/SKILL.md +2 -3
  130. package/pipeline/skills/shared/core/multi-agent-resume/SKILL.md +51 -15
  131. package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +6 -6
  132. package/pipeline/skills/skills-index.md +3 -6
  133. package/pipeline/commands/multi-agent/local/SKILL.md +0 -132
  134. package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +0 -142
  135. package/pipeline/commands/multi-agent/resume-local/SKILL.md +0 -114
  136. package/pipeline/skills/shared/core/multi-agent-local/SKILL.md +0 -41
  137. package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +0 -55
  138. package/pipeline/skills/shared/core/multi-agent-resume-local/SKILL.md +0 -51
package/CHANGELOG.md CHANGED
@@ -14,6 +14,149 @@ Internal file-layout changes that don't affect the slash-command surface are sti
14
14
 
15
15
  ---
16
16
 
17
+ ## [20.0.0] - 2026-09-21
18
+
19
+ ### Removed
20
+
21
+ - **The Phase 0 Step 7.5 depth question, and `state.onlyDevelop` with it.** A
22
+ picker asked Full or Short before the run had read anything, and Short skipped
23
+ the plan phase outright - a guess about work not yet examined, recommended
24
+ from `taskType` alone. It was the same defect v19.0.0 removed from the
25
+ analysis phase when Lite went: a fixed selector overriding a rule that reads
26
+ evidence. There is one pipeline; every mode runs its whole phase set.
27
+ - **`/multi-agent:local`, `/multi-agent:local-autopilot` and the `--local`
28
+ flag.** Where the branch lives is the Phase 0 Step 5b question and nothing
29
+ else answers it. `autopilot` resolves it to a worktree without asking,
30
+ because an unattended run commits and pushes from wherever it stands.
31
+ - **`/multi-agent:resume-local`**, folded into `/multi-agent:resume`. One
32
+ command now lists both kinds of unfinished work - runs that stopped
33
+ mid-phase, and the current branch when it carries work no run produced - and
34
+ asks which to pick up. The tail it runs (Review → Commit → Report, no Plan
35
+ and no Dev) is unchanged; only its entry point moved.
36
+
37
+ Command inventory: 60 → 57.
38
+
39
+ ### Fixed
40
+
41
+ - **Every phase reported its telemetry under the number it had before the
42
+ six-phase merge.** 29 `log-metric.sh` and `phase-tracker.sh tokens` calls
43
+ across the phase docs carried literal ids from the eight-phase contract, so
44
+ Dev's spend landed on Review's tile, Review's on Commit's, and Commit's on a
45
+ number no phase has. The accounting gate then refused to close those phases,
46
+ because the phase whose completion it guarded had recorded nothing. Each doc
47
+ now emits under its own id, and `smoke-phase-telemetry-ids.sh` derives that
48
+ id from the file name rather than a list.
49
+ - **The accounting gate covered phases 1-4 by literal**, which under the new
50
+ numbering left Report ungated and, before the renumbering above, gated
51
+ Commit against a doc that recorded nowhere. `TRACKER_LLM_PHASES` now defaults
52
+ to every phase that dispatches a model, and the gate checks each covered
53
+ phase has a recording path.
54
+ - **`/multi-agent:doctor` was missing from the Turkish help catalog and
55
+ `:analysis-jira` from the Turkish one too.** Half the audience could not see
56
+ two commands that ship. `smoke-help-catalog.sh` now compares both language
57
+ blocks against the command tree, and fails a row pointing at a command that
58
+ does not ship.
59
+ - **`smoke-description-tr.sh` sampled a hardcoded command** that no longer
60
+ exists, so its round-trip check passed on a missing file. The sample is
61
+ derived from the tree.
62
+
63
+ ### Added
64
+
65
+ - **`smoke-picker-callers.sh`** - `AskUserQuestion` refuses a call declaring
66
+ fewer than two options and discards every question batched with it, so the
67
+ rules that prevent it have to reach whoever writes a picker. Every spec that
68
+ renders one now cites `picker-contract.md`, the gate holds that at 47
69
+ callers, and a third check fails any spec that narrates the one-candidate
70
+ skip the contract forbids. A fourth holds the Phase 0 chain to its step
71
+ breadcrumb, so a six-question intake keeps saying which step it is on.
72
+ - **`smoke-phase-telemetry-ids.sh`** and **`smoke-help-catalog.sh`**, described
73
+ under Fixed above.
74
+
75
+ ### Changed
76
+
77
+ - **Phase 1 always runs, and what it produces scales with the evidence.** The
78
+ Locked 2 omission rule already drops a section with nothing behind it, so a
79
+ one-line chore yields a short document rather than a skipped phase. Phase 2
80
+ steps 1, 2, 3, 5 and 6 read that document; a step whose section is absent
81
+ records `not-applicable (no <section> in this document)` instead of aborting.
82
+ - **The Plan Approval Gate has one skip left: `autopilot`**, and it is a skip
83
+ because there is nobody to ask, not because the run is a lighter kind of run.
84
+ - **Tracker registration is a single batch at Step -1.** Nothing is deferred:
85
+ no answer later in the run can add or remove a phase, so `gen-mode-dispatch.mjs`
86
+ loses its deferred branch and every generated tracker section shrinks to one
87
+ registration loop.
88
+ - **`global.resumeLocal` → `global.resume`**, preferences schema 2.8.0. The
89
+ `autoFix` value carries across; the key follows the command the tail now
90
+ enters through.
91
+ - **`phases.json` modes drop their `local` flag.** No mode set it once the two
92
+ local commands went, and a generator branch no caller can reach reads as a
93
+ capability. The generated dispatch blocks are byte-identical, so the drift
94
+ gate holds.
95
+ - Two gates now assert the ABSENCE: `smoke-pipeline-surface.sh` fails if Phase 0
96
+ grows a depth step or writes a depth key, or if any of the four downstream
97
+ readers reads one; `smoke-plan-approval-gate.sh` fails if the gate's scope
98
+ clause widens past autopilot.
99
+
100
+ ### Breaking
101
+
102
+ `state.onlyDevelop` is gone from a schema with `additionalProperties: false`, so
103
+ a state file written before v20 does not validate. The floor moves with
104
+ `dist-tags.required` rather than a migration: a run below it halts rather than
105
+ reading a record written under a contract that no longer exists.
106
+
107
+ Decision, alternatives and consequences: [ADR-0015](docs/adr/0015-one-pipeline-no-depth-answer.md) and [ADR-0016](docs/adr/0016-the-run-shape-is-asked-not-typed.md).
108
+
109
+ ---
110
+
111
+ ## [19.1.4] - 2026-09-21
112
+
113
+ ### Added
114
+
115
+ - **`smoke-gui-json-contract.sh` - the JSON a non-model consumer decodes stays
116
+ put.** `runs-index.mjs` and `update-check.sh` already had their output pinned;
117
+ `doctor.mjs --json` and `autopilot-status.sh --json` did not, and a renamed key
118
+ there breaks every consumer at once while the pipeline's own callers, which
119
+ read the human output, notice nothing. The gate drives both scripts and reads
120
+ their real stdout: `checks[]` with `{id, severity, step, problem, detail}` and
121
+ unique ids, `exit` an integer, severities from the documented set, the queue
122
+ triple as lists, the counters as numbers and `on` as a boolean. Adding a key
123
+ is deliberately not a failure - a decoder ignores what it does not know.
124
+ - **`smoke-no-confessional-prose.sh`** - shipped comments and docs describe the
125
+ code, not the project's own history. A reason a rule exists still ships; a
126
+ past-defect narration belongs in this file. Four allow-list entries are
127
+ anchored to the lines that carry them, so a stale exemption fails the gate.
128
+ - **A deterministic test for write-state's mid-write refusal.** The refusal
129
+ needs another writer to take the lock between acquiring it and renaming, which
130
+ on an idle machine is too narrow to hit and under load is a coin flip.
131
+ `WRITE_STATE_TEST_DELAY_MS` opens that window on purpose; the gate then asserts
132
+ both halves - exit 4, and nothing of that writer in the file.
133
+
134
+ ### Fixed
135
+
136
+ - **The concurrent-writer gate counted an honest exit 4 as an unexpected code.**
137
+ A writer whose lock is taken mid-write refuses and writes nothing, which is the
138
+ same class of outcome as an exit-2 lock timeout. It is now counted with its
139
+ invariant attached - a refusal that wrote nothing must leave nothing behind -
140
+ and both refusal counts are reported so a rising one stays visible.
141
+ - **`bulk-read.sh` claimed routing could send it outside the Anthropic ladder.**
142
+ It delegates to the `claude` CLI, which speaks that ladder and nothing else, so
143
+ the dispatcher's refusal of an external rung is the correct behaviour and the
144
+ comment was the wrong half. Both READMEs now state the limit: no call site
145
+ sends a model request itself, and a delegated read must not put a file's full
146
+ text outside the account that owns it.
147
+
148
+ ### Changed
149
+
150
+ - **Both READMEs gained a "which model answers" section**, listing
151
+ `/multi-agent:{route-on,route-off,route-status,model}` with the two lines that
152
+ matter: `scope` has no `host-session` value and that is a schema gate, and an
153
+ external rung is refused at dispatch.
154
+ - Confessional history removed from 57 shipped comments and doc paragraphs
155
+ across 40 files; each reason the rule exists was kept and rewritten with the
156
+ code as its subject.
157
+
158
+ ---
159
+
17
160
  ## [19.1.3] - 2026-09-21
18
161
 
19
162
  ### Fixed
@@ -101,6 +244,7 @@ them.
101
244
  is refused for a subagent, because subagent dispatch belongs to the host - the
102
245
  script says so on stderr instead of substituting an Anthropic rung and leaving
103
246
  the user believing a rule worked that never could.
247
+
104
248
  - `cost-table.json` rungs declare a `provider`. Without it every rung looks
105
249
  alike and the subagent limit above cannot be checked at all.
106
250
  - `smoke-model-dispatch.sh` (18 assertions). Half of them drive the router; the
@@ -495,7 +639,7 @@ gate cannot be loaded.
495
639
  and producing no effect.
496
640
 
497
641
  **Interactive runs ask at the step instead of ending at it**: open the item and
498
- fix it, continue without it (recording in `state.maturity.accepted[]` *which*
642
+ fix it, continue without it (recording in `state.maturity.accepted[]` _which_
499
643
  gap was waved through, which is what separates an informed continue from a
500
644
  skipped check), or abort. `askInteractively: false` restores the old halt.
501
645
 
@@ -519,7 +663,7 @@ gate cannot be loaded.
519
663
 
520
664
  **The rule that shapes the rest: an edit is a reason to look again, never proof
521
665
  the gap closed.** A reply reading "will do later" moves the timestamp and fixes
522
- nothing, so a changed item is re-fetched and re-scored and the *check* decides.
666
+ nothing, so a changed item is re-fetched and re-scored and the _check_ decides.
523
667
  Only a genuinely DIFFERENT gap set earns a second comment - otherwise an
524
668
  unattended queue turns an item into a wall of identical bot text. "Cannot tell
525
669
  whether it moved" resolves to re-check, never to wait, because folding unknown
@@ -886,6 +1030,7 @@ were the same shape: a chain fixed on one side and left broken on the other.
886
1030
  than attaching nothing because a present artefact does not get re-checked.
887
1031
  `fit` gained `webm`, which had been falling through to the pass-through arm and
888
1032
  reaching the uploader at full size.
1033
+
889
1034
  - **Continuous mode's node half still named one host.** 17.2.0 fixed the shell
890
1035
  side and closed the class. `autopilot-runner.mjs` was still resolving its three
891
1036
  siblings from `~/.claude/scripts`, and `autopilot-intake.mjs` looked for
@@ -1233,7 +1378,6 @@ The flow video existed as a contract with no recorder, no UI test ever ran, and
1233
1378
 
1234
1379
  - **`install/templates/` now reaches the user tree.** It shipped in the tarball and was never copied, so `setup` told people to merge `install/templates/claude-hooks.json` - a path that exists only in a checkout. From an install the instruction named a file the reader did not have, and nothing said so. The doctor's `hook-coverage` check would have been a permanent `SKIP` for the same reason.
1235
1380
 
1236
-
1237
1381
  ### Fixed
1238
1382
 
1239
1383
  - **The widget came back empty of everything except the phase name.** The native tile renders exactly one string, its subject, so whatever a reader wants from the card has to travel in that subject. It carried `Phase 1 Analysis` and nothing else, while the tracker already held the model, the elapsed time, the tokens and the cost for that phase and the fallback card printed all four. `phase-tracker.sh subjects [id]` now renders the same values in the one shape the host accepts, `tiles` builds its `TaskCreate` calls from it, and every boundary hint re-reads it so the numbers advance instead of freezing at creation time. A phase with nothing to report is still just its name - no empty separators.
@@ -1286,7 +1430,6 @@ The flow video existed as a contract with no recorder, no UI test ever ran, and
1286
1430
 
1287
1431
  - **`council-view.mjs`** renders what each reviewer found and what triage did with it. Everything it needs was already on the state and nothing displayed it; `run-metrics.mjs` reduced it to a ratio and the rows behind the ratio were invisible. De-anonymization is safe here because it runs after triage has ruled, and the gate pins that the map reaches no prompt. `/multi-agent:log` renders it; exit 2 means the run never reached Phase 4.
1288
1432
 
1289
-
1290
1433
  - **A validator summary that says which checks ran.** `validate-analysis-doc.mjs --report` prints every check by name with a verdict each: `ok`, a finding count, or `skipped: <reason>`. The gap it closes is specific - the traceability matrix only runs in the corporate profile, so on a global document it never executed and the output was byte-identical to "ran, found nothing". Attribution is positional and reconciles by construction, so a check added without its mark misnames a finding but can never lose one; the gate asserts the printed set equals the set the file defines. Default output is unchanged for gate callers.
1291
1434
 
1292
1435
  - **Provenance in the reports a human reads.** `evidence_digest` and `base_commit` were already in the document front-matter, on the side only a machine reads. The analysis Phase 5 report, the `review-analysis` verdict and `--report` now carry them too: a timestamp cannot separate two reports made the same day, and the first question anyone asks of an older verdict is which version of the document it judged.
@@ -1312,6 +1455,7 @@ The flow video existed as a contract with no recorder, no UI test ever ran, and
1312
1455
  The contract, its reasoning and the check table live in `analysis/redesign.md`, which `analysis/SKILL.md` does not name and which therefore loads only on a redesign run.
1313
1456
 
1314
1457
  - `ANALYSIS_CEILING` 155000 -> 158500, the largest single raise it has taken. What could not be moved out is the argument: the section skeletons belong in `analysis-template.md`, where every other conditional section (15.6, 16.2) is already defined; Locked 37 belongs in `analysis/locked.md`, because a decision recorded only where it is implemented is not locked; and the intake question belongs in `analysis/intake.md`, because the option is chosen before anything knows the run is a redesign. 5.9 kB stayed outside the count. Everything was compressed twice before the number was picked.
1458
+
1315
1459
  ### Fixed
1316
1460
 
1317
1461
  - `learn-from-transcripts.mjs` called `process.exit(0)` on the line after writing its `--json` result, so a large mining result was cut at the 64 KB pipe buffer - the exact defect the gate added in 16.27.0 exists to prevent, in the file that motivated it. The gate caught it; the two sites it named are now a `main()` with a return.
@@ -1368,12 +1512,11 @@ The flow video existed as a contract with no recorder, no UI test ever ran, and
1368
1512
 
1369
1513
  - **Setup merged one hook event and dropped the rest.** The instruction said to deep-merge `hooks.PreToolUse`, so the capture hooks would never have installed - the fix would have shipped inert.
1370
1514
 
1371
-
1372
1515
  ## [16.26.0] - 2026-09-09
1373
1516
 
1374
1517
  ### Fixed
1375
1518
 
1376
- - **The credential store mangled every multi-line secret, and the failure surfaced three layers away.** `security` picks its output format from the *content* of the value: anything it deems non-printable - a newline included - comes back as bare hex with no marker, and the reader took that at face value. A Firebase service-account JSON went in and `7b0a2020...` came out, which is neither JSON nor base64, so the Crashlytics fetcher reported a token-exchange error for a credential that was stored perfectly. The write side had the mirror defect: `security -i` is line-oriented, so a newline inside `-w` terminated the command mid-value and the item was never written at all.
1519
+ - **The credential store mangled every multi-line secret, and the failure surfaced three layers away.** `security` picks its output format from the _content_ of the value: anything it deems non-printable - a newline included - comes back as bare hex with no marker, and the reader took that at face value. A Firebase service-account JSON went in and `7b0a2020...` came out, which is neither JSON nor base64, so the Crashlytics fetcher reported a token-exchange error for a credential that was stored perfectly. The write side had the mirror defect: `security -i` is line-oriented, so a newline inside `-w` terminated the command mid-value and the item was never written at all.
1377
1520
 
1378
1521
  `keychain.py` now reads with `-g` and decodes on the explicit `0x` marker rather than guessing from shape, and writes with `-X <hex>` so no value can be cut at a newline. Lookup also walks all four attribute conventions (`-a`+`-l`, `-a`+`-s`, `-l`, `-s`) instead of one.
1379
1522
 
@@ -1397,7 +1540,7 @@ The flow video existed as a contract with no recorder, no UI test ever ran, and
1397
1540
 
1398
1541
  - **`firebase-app-discovery.sh` - the repo already knew the appIds.** Every Crashlytics call needs the opaque `1:<n>:ios:<hex>`, and a console URL carries only the bundle id, so the fetcher spent a Management API round trip resolving it on every run. `GoogleService-Info*.plist` and `google-services.json` name both ids; the script reads every match (a repo with several targets has several plists, and they do not all point at one project), skips build outputs, and emits entries shaped for `prefs.global.firebase.accounts[].apps[]`. The fetcher prefers that when present and falls through to the API when absent, so an install that never ran discovery behaves exactly as before.
1399
1542
 
1400
- - **An encoding-regression gate.** `smoke-keychain.sh` stored a short ASCII word, which is exactly the value class that never reproduced the bug. It now round-trips a multi-line service-account JSON (and asserts it still *parses*, not merely compares equal), a hex-looking string, embedded quotes and backslashes, non-ASCII, and leading/trailing spaces.
1543
+ - **An encoding-regression gate.** `smoke-keychain.sh` stored a short ASCII word, which is exactly the value class that never reproduced the bug. It now round-trips a multi-line service-account JSON (and asserts it still _parses_, not merely compares equal), a hex-looking string, embedded quotes and backslashes, non-ASCII, and leading/trailing spaces.
1401
1544
 
1402
1545
  - **A gate for the shipped-but-dead script.** `firebase-app-discovery.sh` was written, installed onto every machine by all three hosts, and invoked by nothing: its only mention sat inside a comment in another script. That is the declared-but-inert defect in its purest form - the docs describe a capability, the install carries the code, and nothing connects them. `smoke-consumer-smoke-surface.sh` now checks every shipped script for an actual invoker, and is deliberately picky about what counts: comment lines, JSON schema descriptions, CHANGELOG and ROADMAP entries all name a script without ever running it, and each of those, left in the corpus, made the check pass on the broken state while it was being built.
1403
1546
 
package/README.md CHANGED
@@ -43,12 +43,12 @@ package. Check first, then pick the row that matches:
43
43
  node -v; npm -v; command -v node npm npx
44
44
  ```
45
45
 
46
- | What you see | What to do |
47
- |---|---|
48
- | nothing at all | Install Node >= 20.11: `brew install node`, or the LTS installer from nodejs.org |
49
- | `node` works, `npx` does not | `npm i -g @mmerterden/multi-agent-pipeline` then `multi-agent-pipeline install --claude` |
50
- | nvm is installed but the shell does not see it | `source ~/.nvm/nvm.sh && nvm use --lts`, or just open a new terminal |
51
- | npm is ancient (< 5.2, which predates npx) | `npm i -g npm@latest`, or use the global-install row above |
46
+ | What you see | What to do |
47
+ | ---------------------------------------------- | ---------------------------------------------------------------------------------------- |
48
+ | nothing at all | Install Node >= 20.11: `brew install node`, or the LTS installer from nodejs.org |
49
+ | `node` works, `npx` does not | `npm i -g @mmerterden/multi-agent-pipeline` then `multi-agent-pipeline install --claude` |
50
+ | nvm is installed but the shell does not see it | `source ~/.nvm/nvm.sh && nvm use --lts`, or just open a new terminal |
51
+ | npm is ancient (< 5.2, which predates npx) | `npm i -g npm@latest`, or use the global-install row above |
52
52
 
53
53
  And the path that needs neither `npx` nor a global install - clone and run the
54
54
  installer directly:
@@ -83,7 +83,7 @@ Run a task - the input type is auto-detected:
83
83
 
84
84
  Every input runs the same short intake - **account → (repo) → maturity check → dev-context** - then enters Phase 0. A Jira id or GitHub URL is fetched and maturity-checked _before_ any code is written; free-text skips the fetch and goes straight to planning. Multi-repo tasks add extra repos at the dev-context step.
85
85
 
86
- Add `autopilot` to skip confirmations (e.g. `/multi-agent:autopilot "PROJ-1234"`). Neither the workspace nor the depth is a flag any more - they are two questions the run asks at Phase 0: where to run (worktree or your current checkout), then how deep (Full or Short). `--local` and `:local` answer the first up front.
86
+ Add `autopilot` to skip confirmations (e.g. `/multi-agent:autopilot "PROJ-1234"`). The workspace is not a flag either - it is the one question the run asks about its own shape at Phase 0: where to run, a worktree or your current checkout.
87
87
 
88
88
  Update later with `/multi-agent:update`. Uninstall (tokens preserved) with `npx @mmerterden/multi-agent-pipeline uninstall`.
89
89
 
@@ -91,13 +91,12 @@ Update later with `/multi-agent:update`. Uninstall (tokens preserved) with `npx
91
91
 
92
92
  ## How it works
93
93
 
94
- One command runs up to 6 phases, with a gate between the risky ones. Phase 0
95
- asks two questions that decide the shape of the rest - how deep the run goes
96
- (Full or Short) and where the branch lives (a worktree or your current
97
- checkout):
94
+ One command runs 6 phases, with a gate between the risky ones. Every run
95
+ carries the whole set; the one question Phase 0 asks about shape is where the
96
+ branch lives, a worktree or your current checkout:
98
97
 
99
98
  - **0 · Init** - parse the input (Jira id / GitHub URL / free text), pick account + repo(s), fetch the issue, run a maturity check.
100
- - **1 · Plan** - detect the stack, scan the codebase and write the analysis document, then break it into tasks with file-level targets and **stop for your approval** before touching code. Analysis and planning were two phases until 19.0.0; the depth picker always skipped them together, because they are one decision. Codebase scanning runs on the explorer persona (Sonnet).
99
+ - **1 · Plan** - detect the stack, scan the codebase and write the analysis document, then break it into tasks with file-level targets and **stop for your approval** before touching code. Analysis and planning were two phases until 19.0.0; they are one decision, so they are one phase. Codebase scanning runs on the explorer persona (Sonnet).
101
100
  - **2 · Dev** - TDD: failing test → code → green, following the repo's style + the active stack skills. The phase ends at its own gate: build, lint, tests and a secret scan, run **once**. Review used to build again, and nothing consumed the difference.
102
101
  - **3 · Review** - a **CLI-aware parallel review** against the logs Dev produced - Claude Code runs 3 models (Fable + Opus + Sonnet), Copilot CLI runs 3 (GPT-5.4 + Opus + Sonnet) - then a **Fable triage** keeps only actionable findings; blockers loop back to Phase 2. The optional user test lives here, keeping its waiting state.
103
102
  - **4 · Commit/PR** - conventional commit, push (must succeed), open a PR (`Ref: #N`, never auto-close).
@@ -109,16 +108,17 @@ checkout):
109
108
 
110
109
  Two of the six phases were doing the same work twice. Dev built the project
111
110
  and tee'd a log; Review opened by building it again. Analysis and Planning were
112
- already one decision - the depth picker skipped them together and the state
113
- schema described them as one unit. Six phases now, one build per run.
111
+ already one decision, and the state schema described them as one unit. Six
112
+ phases now, one build per run.
114
113
 
115
114
  The other half of the change is that the count is finally guarded.
116
115
  `smoke-phase-contract.sh` derives it from `pipeline/schemas/phases.json` and
117
116
  holds every other copy to it: the generator's output, the token budget, the
118
117
  state-schema bounds, the progress fractions in sample output, and the named
119
- thresholds that used to be literals scattered across scripts. The phase count
120
- appeared in 91 places across 40 files with nothing checking any of them, while
121
- the command count, the jq count and the persona count all had gates.
118
+ thresholds, none of which is a literal in a script any more. A phase count is
119
+ otherwise the kind of number that spreads across dozens of files with nothing
120
+ checking any of them, the way the command count, the jq count and the persona
121
+ count are each held by a gate.
122
122
 
123
123
  Reasoning, mapping and rejected alternatives:
124
124
  [ADR-0014](./docs/adr/0014-six-phase-consolidation.md).
@@ -156,7 +156,7 @@ A failed `git fetch` degrades loudly rather than silently: the list falls back t
156
156
 
157
157
  ### Workspace: worktree or local
158
158
 
159
- Phase 0 Step 5b asks where the branch lives. **Worktree** (`.worktrees/{id}/`) leaves your current checkout untouched; **Local** works in the project root on a new branch, which drops Phase 5 - the user-test gate checks the change out of a worktree and there is none - and needs the project root clean. `/multi-agent:local` and `--local` answer it up front. Every autopilot entry resolves it to a worktree without asking: an unattended run commits and pushes from wherever it stands, and doing that in your own checkout is what worktrees exist to prevent. `:local-autopilot` is the explicit opt-out.
159
+ Phase 0 Step 5b asks where the branch lives. **Worktree** (`.worktrees/{id}/`) leaves your current checkout untouched; **Local** works in the project root on a new branch, and needs the project root clean. It is a question, not a flag: there is no `:local` command and no `--local` switch, so the answer is always visible in the run rather than buried in how the run was typed. `autopilot` resolves it to a worktree without asking, because an unattended run commits and pushes from wherever it stands and doing that in your own checkout is what worktrees exist to prevent. The manual-test offer at the end of Review follows the same answer: a worktree run is asked whether to check the branch out and test it, a local run already is that checkout, so there is nothing to offer.
160
160
 
161
161
  Under the hood: each task runs in its own **git worktree** (or the current branch when you choose local), commits use the **git identity routed from the repo's origin URL**, and **multi-repo** tasks get per-repo worktrees plus an integration build. Tokens stay in the OS keychain; nothing is committed or logged. `/multi-agent:review` can also review an existing GitHub/Bitbucket PR - per-finding inline comments anchored to `file:line` + an explicit Approve / Needs-Work state.
162
162
 
@@ -167,39 +167,23 @@ The discipline behind all of this - bounded loops, evidence gates, token-budgete
167
167
  | Mode | Command | Flow |
168
168
  | --------- | ------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- |
169
169
  | Full | `/multi-agent "task"` | All 6 phases, interactive |
170
- | Autopilot | `/multi-agent:autopilot "task"` | 6 phases (interactive Test gate dropped), no confirmations |
171
- | Local | `/multi-agent:local "task"` | Full pipeline minus the interactive Test gate, current branch (no worktree) |
172
- | Depth | asked at Phase 0 Step 7.5 | Full (all phases) or Short (Dev → Review → Test → Commit → Report). Not a command name - `/multi-agent` and `:local` ask, both autopilot entries always run Full |
173
- | Ship | `/multi-agent:resume-local` | Run the review→test→commit→report tail over local work |
170
+ | Autopilot | `/multi-agent:autopilot "task"` | The same 6 phases, no confirmations; the workspace resolves to a worktree and the user-test gate inside Review is skipped |
174
171
  | Audit | `/multi-agent:design-check` | Mock-mode vs Figma conformance, local-only |
175
172
  | Audit | `/multi-agent:testflight-validation` | Pre-submission gates for a TestFlight build: static archive audit → Apple's `altool --validate-app` → Review-Guidelines check. Validates only, never uploads |
176
173
 
177
- Depth, autopilot and `--local` are the only knobs on the run itself; everything else is its own command. The full catalog is below.
178
- ### Pipeline depth: Full or Short
179
-
180
- Depth is the one question the run asks about its own shape. `/multi-agent` and `/multi-agent:local` ask it at Phase 0 Step 7.5 - after the issue is fetched and the task type is known, because that is what the recommendation is drawn from.
181
-
182
- - **Full** runs everything: Analysis reads the codebase and maps impact, Planning writes a task breakdown and stops for your approval, and Dev works from that plan.
183
- - **Short** starts at Dev: Init → Dev → Review → Test → Commit → Report. Analysis and Planning do not run, so there is no plan gate and no analysis document; Dev derives its own task list from the issue, and runs on **Opus** rather than Sonnet because it has no plan to follow. Review, the deterministic gates and the test suite are untouched - Short skips the thinking, never the proof.
184
-
185
- Short is right when you already know the fix and the file: a one-line guard, a copy change, a rename, a revert. It is wrong when the cause is still a hypothesis, when the task carries a Figma reference or an analysis document (Planning is what turns those into a breakdown), or when the change spans repos.
186
-
187
- The widget follows the answer rather than predicting it: Phase 0 is the only tile drawn before you choose, and a Short run never draws an Analysis tile at all. Both autopilot entries skip the question and always run Full. `agent-state.json` records which one ran as `onlyDevelop`.
188
-
174
+ `autopilot` is the only knob on the run itself; everything else is its own command. The full catalog is below.
189
175
 
190
176
  ## Commands
191
177
 
192
- `/multi-agent` plus 56 sub-commands. `/multi-agent:help` renders the same catalog in your terminal, in your `outputLanguage`.
178
+ `/multi-agent` plus 57 sub-commands. `/multi-agent:help` renders the same catalog in your terminal, in your `outputLanguage`.
193
179
 
194
180
  ### Pipeline entries
195
181
 
196
182
  | Command | What it does |
197
183
  | ------------------------------------- | ---------------------------------------------------------------------------------------------------- |
198
- | `/multi-agent "task"` | Full pipeline in a worktree. Asks Full or Short depth at Phase 0 |
199
- | `/multi-agent:local "task"` | Same pipeline on the current branch, no worktree |
200
- | `/multi-agent:autopilot "task"` | Worktree, no confirmations, always Full |
201
- | `/multi-agent:local-autopilot "task"` | Current branch, no confirmations, always Full |
202
- | `/multi-agent:resume-local` | Pipeline tail over work already done locally: Review → Build+Test → Commit/PR → Report. No dev phase |
184
+ | `/multi-agent "task"` | The pipeline; Phase 0 asks where the branch lives |
185
+ | `/multi-agent:autopilot "task"` | The same pipeline, unattended: worktree resolved, no confirmations |
186
+ | `/multi-agent:resume` | Unfinished work, either source: a stopped run picks up where it left off, a branch with no run behind it gets the tail (Review → Build+Test → Commit/PR → Report) |
203
187
 
204
188
  ### Task control
205
189
 
@@ -207,7 +191,7 @@ The widget follows the answer rather than predicting it: Phase 0 is the only til
207
191
  | --------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- |
208
192
  | `/multi-agent:status` | Every task's ID, phase, branch and state |
209
193
  | `/multi-agent:log [#N]` | Show a task's `agent-log.md` (most recent by default) |
210
- | `/multi-agent:resume [#N]` | Carry a stopped or failed task on from its last phase |
194
+ | `/multi-agent:resume [#N]` | Carry a stopped or failed task on from its last phase, or run the tail over a branch with no run behind it |
211
195
  | `/multi-agent:kill [#N]` | Stop a task, remove its worktree and branch |
212
196
  | `/multi-agent:steer #N "<instruction>"` | Correct a running task without stopping it; applied at the next phase boundary |
213
197
  | `/multi-agent:search` | Ranked search across every task log; `--semantic` queries the triage corpus |
@@ -300,11 +284,37 @@ Two compliance skills install on every host and back the store gates: `apple-arc
300
284
  Everything above starts when you start it. Continuous mode is the same pipeline
301
285
  picking work up on its own, on ONE machine you choose, from repos you choose.
302
286
 
303
- | Command | What it does |
304
- | --- | --- |
305
- | `/multi-agent:autopilot-on` | Pick the repos this machine watches. Labelled GitHub issues and assigned + labelled Jira items then run in a worktree and stop at an open PR. Re-run to change the list |
306
- | `/multi-agent:autopilot-status` | What is running and at which phase, what is queued, what is waiting for an answer, the PRs of the last day, and the rolling spend |
307
- | `/multi-agent:autopilot-off` | Remove the schedule. Work already running finishes; the repo selection is kept |
287
+ | Command | What it does |
288
+ | ------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
289
+ | `/multi-agent:autopilot-on` | Pick the repos this machine watches. Labelled GitHub issues and assigned + labelled Jira items then run in a worktree and stop at an open PR. Re-run to change the list |
290
+ | `/multi-agent:autopilot-status` | What is running and at which phase, what is queued, what is waiting for an answer, the PRs of the last day, and the rolling spend |
291
+ | `/multi-agent:autopilot-off` | Remove the schedule. Work already running finishes; the repo selection is kept |
292
+
293
+ ## Which model answers
294
+
295
+ The ladder is `fable -> opus -> sonnet -> haiku`, and a run walks it on failure.
296
+ Routing lets a policy pick the rung instead, per call site, and is off until you
297
+ turn it on.
298
+
299
+ | Command | What it does |
300
+ | --------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------- |
301
+ | `/multi-agent:route-on` | Pick a strategy, a scope and the rules that say which rung a call lands on. Validated against `route-config.schema.json` before anything is written |
302
+ | `/multi-agent:route-off` | Disarm it. The rules are KEPT, so turning it back on does not re-ask for the same configuration |
303
+ | `/multi-agent:route-status` | Whether it is armed, which rule applies where, which rung the last dispatches took, and what this run has cost |
304
+ | `/multi-agent:model` | Turn the top rung on or off, and move the cost ledger's pricing with it in the same step |
305
+
306
+ **The scope cannot be the whole session.** `scope` accepts `subagent`,
307
+ `bulk-read` and `research`; there is no `host-session` value and that is a schema
308
+ gate, not a convention. Owning the host's base URL would send every call you make
309
+ through a third layer, including work that has nothing to do with this pipeline.
310
+
311
+ **The honest limit, which `route-status` prints rather than hides.** A rung on a
312
+ non-Anthropic provider is refused at dispatch. No call site sends a model
313
+ request itself: the one that routes, `bulk-read.sh`, delegates to the `claude`
314
+ CLI, which speaks that ladder and nothing else. Sending a read to a third-party
315
+ provider would also put the file's full text outside the account that owns it,
316
+ so it is a decision a user makes rather than a default. Routing chooses inside
317
+ the ladder.
308
318
 
309
319
  **Nothing is on by default and nothing is added implicitly.** Installing the
310
320
  package writes no state and schedules nothing; `smoke-autopilot-default-off.sh`
@@ -359,17 +369,17 @@ This enables the matching plugin (+ the shared `ai-common` plugin) in the repo's
359
369
 
360
370
  ## Tool support
361
371
 
362
- The pipeline runs natively on **Claude Code**, **Copilot CLI** and **Codex CLI** - all three install from the same `pipeline/` source and get the same 60 commands.
372
+ The pipeline runs natively on **Claude Code**, **Copilot CLI** and **Codex CLI** - all three install from the same `pipeline/` source and get the same 57 commands.
363
373
 
364
374
  | Tool | Flag | What it installs |
365
375
  | ----------- | -------------------- | ------------------------------------------------------------------------------------------------------ |
366
376
  | Claude Code | `--claude` (default) | slash commands + skills + agents + three `PreToolUse` hooks (secret scan, agent-guard, read-size gate) |
367
- | Copilot CLI | `--copilot` | instructions + 60 sub-command skills + scripts |
368
- | Codex CLI | `--codex` | one router skill + 60 specs as refs + 9 agent TOML + `AGENTS.md` block + `codex mcp add` |
377
+ | Copilot CLI | `--copilot` | instructions + 57 sub-command skills + scripts |
378
+ | Codex CLI | `--codex` | one router skill + 57 specs as refs + 9 agent TOML + `AGENTS.md` block + `codex mcp add` |
369
379
 
370
380
  Filter skills by stack with `--platform=ios\|android\|all`.
371
381
 
372
- **Why Codex gets one skill and not 60.** Codex assembles every discovered skill's name
382
+ **Why Codex gets one skill and not 57.** Codex assembles every discovered skill's name
373
383
  and description into a single prompt block and drops entries when it overflows, with no
374
384
  error. Measured on 0.145: installing one plugin that declares 142 skills surfaced only
375
385
  75 of them and evicted an unrelated user skill. So on Codex the pipeline ships a single