@mmerterden/multi-agent-pipeline 17.5.0 → 17.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (41) hide show
  1. package/CHANGELOG.md +177 -0
  2. package/README.md +24 -0
  3. package/README.tr.md +24 -0
  4. package/docs/features.md +44 -3
  5. package/docs/token-budget-history.md +1 -1
  6. package/install/templates/claude-hooks.json +13 -1
  7. package/package.json +1 -1
  8. package/pipeline/commands/multi-agent/SKILL.md +1 -1
  9. package/pipeline/commands/multi-agent/feedback/SKILL.md +7 -1
  10. package/pipeline/commands/multi-agent/graph/SKILL.md +1 -1
  11. package/pipeline/commands/multi-agent/issue/SKILL.md +13 -1
  12. package/pipeline/commands/multi-agent/jira/SKILL.md +13 -1
  13. package/pipeline/commands/multi-agent/resume/SKILL.md +16 -1
  14. package/pipeline/commands/multi-agent/setup/SKILL.md +14 -16
  15. package/pipeline/commands/multi-agent/update/SKILL.md +13 -56
  16. package/pipeline/multi-agent-refs/features/code-graph.md +20 -0
  17. package/pipeline/multi-agent-refs/features/doctor.md +23 -0
  18. package/pipeline/multi-agent-refs/features/maturity-followup.md +166 -0
  19. package/pipeline/multi-agent-refs/features/package-manager.md +80 -0
  20. package/pipeline/multi-agent-refs/features/usage-reporting.md +79 -0
  21. package/pipeline/multi-agent-refs/features/verify-by-test.md +1 -1
  22. package/pipeline/multi-agent-refs/phases/phase-0-init.md +5 -2
  23. package/pipeline/multi-agent-refs/phases/phase-3-dev.md +8 -2
  24. package/pipeline/multi-agent-refs/phases/phase-4-review.md +1 -1
  25. package/pipeline/multi-agent-refs/picker-contract.md +1 -1
  26. package/pipeline/preferences-template.json +1 -1
  27. package/pipeline/schemas/agent-state.schema.json +122 -11
  28. package/pipeline/schemas/prefs.schema.json +35 -0
  29. package/pipeline/schemas/token-budget.json +2 -2
  30. package/pipeline/scripts/doctor.mjs +65 -0
  31. package/pipeline/scripts/feedback-send.mjs +1 -1
  32. package/pipeline/scripts/graph-report.mjs +155 -1
  33. package/pipeline/scripts/maturity-followup.mjs +294 -0
  34. package/pipeline/scripts/package-manager.mjs +310 -0
  35. package/pipeline/scripts/usage-register.mjs +271 -0
  36. package/pipeline/scripts/usage-report.mjs +2 -2
  37. package/pipeline/skills/.skill-manifest.json +5 -5
  38. package/pipeline/skills/shared/core/multi-agent-issue/SKILL.md +14 -0
  39. package/pipeline/skills/shared/core/multi-agent-jira/SKILL.md +14 -0
  40. package/pipeline/skills/shared/core/multi-agent-setup/SKILL.md +13 -0
  41. package/pipeline/skills/shared/core/multi-agent-update/SKILL.md +6 -0
package/CHANGELOG.md CHANGED
@@ -14,6 +14,183 @@ Internal file-layout changes that don't affect the slash-command surface are sti
14
14
 
15
15
  ---
16
16
 
17
+ ## [17.6.0] - 2026-09-15
18
+
19
+ ### Added
20
+
21
+ - **The maturity check now has an effect (`prefs.global.maturityFollowup`).** It
22
+ has always produced a machine-readable gap list - stable codes in `blockers[]`
23
+ and `warnings[]` - and then thrown most of it away. A blocker halted the run,
24
+ an autopilot queue moved to the next item, and the item stayed exactly as
25
+ immature as it was found. Nobody was told, so nothing changed, so the next
26
+ scan halted on the same item for the same reason. The check was doing its job
27
+ and producing no effect.
28
+
29
+ **Interactive runs ask at the step instead of ending at it**: open the item and
30
+ fix it, continue without it (recording in `state.maturity.accepted[]` *which*
31
+ gap was waved through, which is what separates an informed continue from a
32
+ skipped check), or abort. `askInteractively: false` restores the old halt.
33
+
34
+ **Autopilot can ask on the item itself**, behind
35
+ `autopilotCommentsOnIssue` (**off by default**, because it is an outward-facing
36
+ write). One comment naming what is missing, then a halt on the circuit breaker
37
+ with `state.waitingFor = "maturity"`. A question, never a state change: no
38
+ transition, no resolution, no assignee, no label, no close. `Ref:` never
39
+ `Closes:`. Copy in `outputLanguage`, and the gap wording is the fetcher's own
40
+ `maturity.summary` verbatim rather than a second copy of that table.
41
+
42
+ **"Have we already asked" is read off the ITEM, not off our state file.** An
43
+ autopilot scan is a new run with a fresh `agent-state.json`, so a state-only
44
+ record would make every scan a first ask - the wall of identical bot comments
45
+ this feature exists to prevent. The comment therefore carries its own gap set
46
+ on a last line (`multi-agent gaps: code,code`) and the next pass takes the
47
+ newest comment of ours. Neither the marker nor that line uses square brackets:
48
+ `[text]` is a link in Jira wiki markup, and Jira is where this comment is most
49
+ likely to land. The contract is carried on all three hosts - the Claude command
50
+ tree, and the `shared/core` skills that Copilot dispatches and Codex reads.
51
+
52
+ **The rule that shapes the rest: an edit is a reason to look again, never proof
53
+ the gap closed.** A reply reading "will do later" moves the timestamp and fixes
54
+ nothing, so a changed item is re-fetched and re-scored and the *check* decides.
55
+ Only a genuinely DIFFERENT gap set earns a second comment - otherwise an
56
+ unattended queue turns an item into a wall of identical bot text. "Cannot tell
57
+ whether it moved" resolves to re-check, never to wait, because folding unknown
58
+ into "nothing changed" parks a run forever on a tracker that omits the field.
59
+
60
+ Warnings still auto-continue under autopilot. Converting them to halts in a
61
+ release would stall queues overnight on items that ran fine yesterday;
62
+ `commentOnWarnings` raises them opt-in, and the gaps are recorded either way.
63
+
64
+ > Four of the additions below started as ideas in a public Claude Code
65
+ > configuration repo (`worldflowai/everything-claude-code`). No code was taken:
66
+ > that copy ships no LICENSE file and the GitHub API reports none, and the same
67
+ > call was already made twice here (ADR-0010 on `graphify`, and swiftlens). Each
68
+ > idea was re-derived against this pipeline's own constraints - zero runtime
69
+ > dependencies, macOS only, stdout is the MCP server's JSON-RPC channel - and
70
+ > most of what that repo does was already covered here, in a gated form.
71
+
72
+ - **The node stacks no longer type `npm` (`pipeline/scripts/package-manager.mjs`).**
73
+ Phase 3's web test arm and its build step were hardcoded, so a repo on pnpm,
74
+ yarn or bun failed in Phase 3 - with a worktree and a branch already created -
75
+ or, worse, npm resolved against a lock file it does not own and the run
76
+ continued on a tree the repo's own tooling would never have produced. The
77
+ manager is now resolved from the repo: `$MA_PACKAGE_MANAGER`, then
78
+ `package.json#packageManager`, then a lock file, then npm reported AS a
79
+ default rather than as evidence. Node core only, per ADR-0004. Two lock files
80
+ means a migration left one behind: the newest wins and both are named. Exit 3
81
+ means the repo declares no such script, which is the `--if-present` case
82
+ answered by an exit code instead of a flag whose support differs per manager.
83
+ The resolved name goes through the phase's `eval`, so it is held to the shape
84
+ a binary actually has and the gate proves it by eval'ing the produced line
85
+ with every manager stubbed out. iOS and Android are untouched.
86
+ `refs/features/package-manager.md`.
87
+
88
+ - **A `PreCompact` hook, so a compaction does not eat what a phase learned.**
89
+ `SessionEnd` capture exists because every durable write used to live in Phase
90
+ 7, the phase a run is least likely to reach. An auto-compaction is that same
91
+ failure one level down: it summarizes the conversation while a long Phase 3 or
92
+ Phase 4 is still running, and anything not yet flushed is gone before
93
+ `SessionEnd` ever fires. The template now wires `capture-flush.sh` there,
94
+ deliberately WITHOUT `--if-stale` - that test exists so a session exit does not
95
+ re-flush a finished run, and at a compaction what matters is whether anything
96
+ is unflushed, not whether the run finished. Both store writes are idempotent.
97
+ `doctor`'s `hook-coverage` check reports it as missing until the block is
98
+ merged, at no extra cost, because that check diffs the template generically.
99
+
100
+ - **`doctor` now counts the MCP servers this host has registered (`mcp-surface`).**
101
+ `mcp-registration` answers "is ours registered". Nobody was answering the
102
+ question the user never gets asked: every registered server sends its tool
103
+ list on every turn, they are added one at a time, and our own toolkit is 99
104
+ tools by itself. The check only ever REPORTS - INFO above
105
+ `prefs.global.mcpSurface.infoAbove` (default 8), `--explain` lists the names,
106
+ and it never blocks, never warns and never disables anything, because how many
107
+ servers are worth their context is the user's call and not a health failure.
108
+ The threshold is judgement, which is why it is a pref instead of a constant
109
+ nobody can see.
110
+
111
+ - **`GRAPH_REPORT.md` gained "Symbols nothing else references".** The report
112
+ found unconnected FILES; a file imported for one symbol while three of its
113
+ other exports were dead has edges, so those exports were invisible. The join
114
+ is a query over data the graph already held. Candidates, never verdicts, and
115
+ it gates nothing: ADR-0010 records that the extractor is regex over
116
+ comment-stripped source, not a parser, so dynamic dispatch, reflection,
117
+ string-keyed lookup and a public API consumed outside the repo all look like
118
+ dead code from here. Four classes are excluded and COUNTED rather than listed,
119
+ because they could not carry a reference edge however heavily used they are: a
120
+ kind outside the stack's `referenceKinds`, a name declared in more than one
121
+ place, a nested declaration, and anything in a test file. Symbols referenced
122
+ only from tests are listed separately - not dead code, but code whose only
123
+ consumer is its own test.
124
+
125
+ ### Fixed
126
+
127
+ - **`/multi-agent:resume` started from `currentPhase + 1` unconditionally**, so a
128
+ run that stopped mid-phase to ask a human resumed past the question. It reads
129
+ `state.waitingFor` first now. This was already broken for Phase 7's channels
130
+ pause, which documented itself as resumable through that field while
131
+ `resume/SKILL.md` never mentioned it - one fix, two callers.
132
+
133
+ - **Only one machine in the world was reporting usage, and nothing was red.**
134
+ Operational reporting needs `usageLog.enabled` AND a token that resolves, and
135
+ both were arranged in exactly one place: `/multi-agent:update`, as forty lines
136
+ of shell embedded in the skill. A user who installed the package, ran
137
+ `/multi-agent:setup` (which said registration happens in update) and worked for
138
+ weeks never registered, never reported, and an empty panel reads exactly like
139
+ nobody using the pipeline. Copilot and Codex were worse off still: their own
140
+ `setup` and `update` skills never mentioned registration at all.
141
+
142
+ It is now one script (`pipeline/scripts/usage-register.mjs`) called from the
143
+ five surfaces where a machine can first become real - setup and update on
144
+ Claude Code, the same two on the cross-CLI surface, and the Phase 0 exit gate
145
+ as the backstop for a machine that reached neither. Same fence as before, now
146
+ enforced in one place: the token is REQUESTED and write-only, it lands in the
147
+ credential store and never in a file, `usageLog.optOut: true` blocks everything
148
+ permanently and is checked before the network call, and an unreachable endpoint
149
+ leaves reporting off with one line and exit 0 - a caller is never failed over
150
+ bookkeeping. `--json` distinguishes "you opted out" from "we could not reach
151
+ it", because those are different facts about the same empty panel. A machine
152
+ that has a token but `enabled: false` (a run interrupted between the two
153
+ writes) is repaired rather than left silent. `refs/features/usage-reporting.md`,
154
+ `smoke-usage-register.sh`.
155
+
156
+ - **`state.maturity` was written by every issue-shaped run and forbidden by the
157
+ schema.** `agent-state.schema.json` has `additionalProperties: false`, and
158
+ `maturity` was not declared - it survived only on the grandfather list of
159
+ `smoke-state-keys-declared.sh`. Declared now, along with `maturityFollowup`
160
+ and `waitingFor`, whose two values (`maturity`, `user-channels-choice`) are the
161
+ two places a run pauses INSIDE a phase rather than between two. The grandfather
162
+ list is two entries shorter, and it only ever shrinks.
163
+
164
+ ---
165
+
166
+ ## [17.5.1] - 2026-09-15
167
+
168
+ ### Documentation
169
+
170
+ 17.5.0 shipped base-branch evidence and documented it in the feature ref and the
171
+ CHANGELOG only. `README.md`, `README.tr.md` and `docs/features.md` all ship in
172
+ the npm tarball, and `docs/features.md` calls itself the complete catalog - so
173
+ the published package described less than it contained, and the npm page showed
174
+ a release missing one of its two features.
175
+
176
+ - Both READMEs gain a "Base branch: evidence, then a question" section: what is
177
+ collected, that no field id or branch prefix is hardcoded, that an uncut
178
+ release is a note rather than an option, that a failed fetch degrades loudly,
179
+ and the pref that turns it off.
180
+ - `docs/features.md` gains the full section, and its modifier-flag table stops
181
+ presenting `--local` as the only route to local mode - the workspace is a
182
+ question at Step 5b now, which 17.5.0 changed without updating that table.
183
+ - `docs/features.md` said the toolkit has "80+ tools". It has 99.
184
+
185
+ ### Website
186
+
187
+ - The base-branch card, and three counts corrected against the filesystem:
188
+ `ai-ios-toolkit` 147 -> 134 skills (its own breakdown summed to 251 against a
189
+ stated total of 238), and the Remote Control entry's 35 -> 56 commands and
190
+ 263 -> 212 skills.
191
+
192
+ ---
193
+
17
194
  ## [17.5.0] - 2026-09-15
18
195
 
19
196
  Four defects from one live run of `/multi-agent`, all in Phase 0, all visible in
package/README.md CHANGED
@@ -72,6 +72,30 @@ checkout):
72
72
 
73
73
  `/multi-agent:analysis` runs its own shorter chain and, since v16.12.0, reviews what it wrote before publishing it: the draft goes through the same three-reviewer set and triage as a code diff, a blocking finding returns it to synthesis with dispatch closed, and the gaps that survive are either searched, asked about, or recorded with an owner. It used to publish behind a structural validator alone.
74
74
 
75
+ ### Your package manager, your hooks, your MCP surface
76
+
77
+ Three smaller things in 17.6.0, each closing a gap where the pipeline assumed instead of looking:
78
+
79
+ - **Phase 3 stopped typing `npm`.** A repo on pnpm, yarn or bun used to fail in the development phase, with a worktree and a branch already created. The manager is resolved from the repo now - an env override, then `package.json#packageManager`, then the lock file, then npm reported as a default rather than as evidence. iOS and Android are untouched.
80
+ - **A compaction no longer eats what a phase learned.** The capture hook ran at session end; an auto-compaction summarizes a long review or development phase while it is still running, and anything not yet written was gone before session end ever fired. The hooks template now flushes at `PreCompact` too.
81
+ - **`doctor` counts your MCP servers.** Every registered server sends its tool list on every turn and they are added one at a time, so nobody ever sees the total. It reports the count and nothing else: no warning, no blocking, no disabling.
82
+
83
+ ### Maturity: the check now has an effect
84
+
85
+ Phase 0 scores how ready the item is before anything is built, and until 17.6.0 it scored and stopped: a blocker halted the run, an autopilot queue moved on, and the item stayed exactly as immature as it was found. Nobody was told, so nothing changed, so the next scan halted on the same item for the same reason.
86
+
87
+ An interactive run now **asks at that step** - open the item and fix it, continue without it, or abort - and records which gap you waved through rather than just that you continued. An autopilot run can be told to **ask on the item itself** (`prefs.global.maturityFollowup.autopilotCommentsOnIssue`, off by default): one comment naming what is missing, then it stops and waits. Never a status change, never an assignee, never a close, `Ref:` and never `Closes:`.
88
+
89
+ Edit the item and the next scan re-checks it; if the gap closed, development starts on its own. An edit is only a reason to look again - a reply reading "will do later" moves the timestamp and fixes nothing, so the check re-runs against the new content and decides. The same question is never asked twice.
90
+
91
+ ### Base branch: evidence, then a question
92
+
93
+ Phase 0 Step 3 collects the candidates **with the evidence behind each one** before it asks. An issue carrying a target version, or linking a separate issue that represents the release, already names the branch on a repo whose release branches encode the version - so that becomes a ranked row whose description says why, next to rows that say "the repository's default branch" or "you used this last time". You still choose; the evidence only reorders.
94
+
95
+ Nothing is hardcoded: any Jira field whose schema resolves to `version` is read whatever the board calls it, and the release-branch naming convention is inferred from the refs that exist rather than read off a table - one repo yields `<prefix>/develop_<version>`, another `release-<version>`, out of the same code. A version whose branch has not been cut yet is reported as a note, never offered as an option you cannot check out.
96
+
97
+ A failed `git fetch` degrades loudly rather than silently: the list falls back to local refs, the question says so, it gains a retry row, and the exit gate refuses to close Phase 0 if a degraded run recorded its list as the remote's answer. Turn the whole thing off with `prefs.global.baseBranchEvidence.enabled: false` and Step 3 asks the way it always did.
98
+
75
99
  ### Workspace: worktree or local
76
100
 
77
101
  Phase 0 Step 5b asks where the branch lives. **Worktree** (`.worktrees/{id}/`) leaves your current checkout untouched; **Local** works in the project root on a new branch, which drops Phase 5 - the user-test gate checks the change out of a worktree and there is none - and needs the project root clean. `/multi-agent:local` and `--local` answer it up front. Every autopilot entry resolves it to a worktree without asking: an unattended run commits and pushes from wherever it stands, and doing that in your own checkout is what worktrees exist to prevent. `:local-autopilot` is the explicit opt-out.
package/README.tr.md CHANGED
@@ -72,6 +72,30 @@ mu):
72
72
 
73
73
  `/multi-agent:analysis` kendi kısa zincirini koşar ve v16.12.0'dan beri yazdığını yayınlamadan önce review ediyor: taslak, bir kod diff'iyle aynı üç-reviewer setinden ve triyajdan geçiyor, bloklayıcı bulgu dokümanı sentez fazına geri gönderip dispatch'i kapatıyor, hayatta kalan boşluklar ya aranıyor ya sana soruluyor ya da sahibiyle birlikte kayda giriyor. Önceden yalnızca yapısal bir validator'ın arkasından yayınlıyordu.
74
74
 
75
+ ### Paket yöneticisi, hook'lar ve MCP yüzeyi
76
+
77
+ 17.6.0'da üç küçük iş; üçü de pipeline'ın bakmak yerine varsaydığı bir yeri kapatıyor:
78
+
79
+ - **Faz 3 artık `npm` yazmıyor.** pnpm, yarn ya da bun kullanan bir repo geliştirme fazında, worktree ve branch zaten açılmışken patlıyordu. Yönetici artık reponun kendisinden çözülüyor: önce ortam değişkeni, sonra `package.json#packageManager`, sonra lock dosyası, sonra kanıt değil varsayılan olarak bildirilen npm. iOS ve Android'e dokunulmadı.
80
+ - **Compaction bir fazın öğrendiğini artık yemiyor.** Yakalama hook'u oturum sonunda koşuyordu; otomatik compaction ise uzun bir review ya da geliştirme fazını tam koşarken özetliyor ve henüz yazılmamış olan her şey oturum sonu hiç gelmeden kayboluyordu. Hook şablonu artık `PreCompact`'te de yazıyor.
81
+ - **`doctor` MCP sunucularını sayıyor.** Kayıtlı her sunucu her turda araç listesini gönderiyor ve sunucular teker teker ekleniyor, yani toplamı kimse görmüyor. Sadece sayıyı bildiriyor: uyarı yok, engelleme yok, kapatma yok.
82
+
83
+ ### Olgunluk: kontrolün artık bir sonucu var
84
+
85
+ Faz 0 hiçbir şey inşa edilmeden önce maddenin ne kadar hazır olduğunu puanlıyor, ve 17.6.0'a kadar puanlayıp duruyordu: bir blocker koşuyu durduruyor, autopilot sırası bir sonraki maddeye geçiyor, madde bulunduğu kadar olgunlaşmamış kalıyordu. Kimseye söylenmediği için hiçbir şey değişmiyor, bir sonraki tarama aynı maddede aynı sebeple duruyordu.
86
+
87
+ İnteraktif koşu artık **o adımda soruyor** - maddeyi açıp düzelt, onsuz devam et, ya da iptal - ve yalnızca devam ettiğini değil, hangi eksiği görmezden geldiğini kaydediyor. Autopilot koşusuna ise **maddenin kendisine sorması** söylenebiliyor (`prefs.global.maturityFollowup.autopilotCommentsOnIssue`, varsayılan kapalı): eksiği adıyla söyleyen tek bir yorum, sonra durup bekliyor. Durum değişmiyor, atanan kişi değişmiyor, kapatma yok; `Ref:` var, `Closes:` yok.
88
+
89
+ Maddeyi düzenlediğinde bir sonraki tarama yeniden kontrol ediyor; eksik kapandıysa geliştirme kendiliğinden başlıyor. Düzenleme yalnızca yeniden bakma sebebi - "sonra yaparım" diyen bir yorum da damgayı ilerletir ve hiçbir şeyi düzeltmez, o yüzden kontrol yeni içerik üzerinde yeniden koşup karar veriyor. Aynı soru iki kez sorulmuyor.
90
+
91
+ ### Temel branch: önce kanıt, sonra soru
92
+
93
+ Faz 0 Adım 3 sormadan önce adayları **her birinin arkasındaki kanıtla** topluyor. Hedef sürüm alanı taşıyan ya da release'i temsil eden ayrı bir maddeye bağlı bir issue, release dallarını sürümle adlandıran bir repoda branch'i zaten söylüyor - bu, nedenini açıklayan sıralı bir satır oluyor; yanında "reponun varsayılan dalı" ya da "geçen sefer bunu kullandın" diyen satırlar duruyor. Seçim yine senin; kanıt yalnızca sıralamayı değiştiriyor.
94
+
95
+ Hiçbir şey gömülü değil: şeması `version`'a çözülen her Jira alanı, board ona ne ad verirse versin okunuyor; release dalı konvansiyonu da tablodan değil, var olan ref'lerden çıkarılıyor - aynı koddan bir repoda `<prefix>/develop_<version>`, başkasında `release-<version>` çıkıyor. Dalı henüz açılmamış bir sürüm not olarak bildiriliyor, checkout edemeyeceğin bir seçenek olarak sunulmuyor.
96
+
97
+ Başarısız bir `git fetch` sessizce değil, yüksek sesle bozuluyor: liste lokal ref'lere düşüyor, soru bunu söylüyor, bir yeniden dene satırı kazanıyor, ve bozuk bir koşu listesini uzaktakinin cevabı diye kaydederse exit gate Faz 0'ı kapatmıyor. Tamamını kapatmak için `prefs.global.baseBranchEvidence.enabled: false` - Adım 3 eskisi gibi soruyor.
98
+
75
99
  ### Çalışma alanı: worktree mi lokal mi
76
100
 
77
101
  Faz 0 Adım 5b branch'in nerede yaşayacağını sorar. **Worktree** (`.worktrees/{id}/`) mevcut checkout'una dokunmaz; **Lokal** proje kökünde yeni bir branch'te çalışır, bu da Faz 5'i düşürür - kullanıcı-test kapısı değişikliği bir worktree'den checkout eder, ortada worktree yoktur - ve proje kökünün temiz olmasını ister. `/multi-agent:local` ve `--local` bu soruyu baştan cevaplar. Her autopilot girişi sormadan worktree'ye karar verir: gözetimsiz bir koşu nerede duruyorsa oradan commit'leyip push eder, ve bunu senin kendi checkout'unda yapmak tam olarak worktree'nin engellemek için var olduğu şeydir. `:local-autopilot` bunun açık opt-out'u.
package/docs/features.md CHANGED
@@ -26,9 +26,11 @@ Each phase reads its own spec file under `pipeline/multi-agent-refs/phases/phase
26
26
  | Flag | Effect |
27
27
  | ----------- | ----------------------------------------------------------------------------------- |
28
28
  | `autopilot` | Skip all confirmation prompts; still fails safe on review blockers + build retries. |
29
- | `--local` | No worktree - works directly in `$PROJECT_ROOT` on a local branch. |
29
+ | `--local` | Answers the workspace question up front: no worktree, work directly in `$PROJECT_ROOT` on a local branch. |
30
30
 
31
- Depth is not a flag. `/multi-agent` and `/multi-agent:local` ask Full or Short at Phase 0 Step 7.5, recommending from the detected `taskType`; Short strips to Init → Dev(Opus self-contained) → Review → Test → Commit → Report. Autopilot never asks and always runs Full - "fast plus unattended" was removed in v16.0.0, because something has to choose when nobody is asked and unattended is the worst place to drop analysis and planning.
31
+ Neither the workspace nor the depth is a flag. **Where the branch lives** is asked at Phase 0 Step 5b - worktree (`.worktrees/{id}/`, your checkout untouched) or local (the project root, which drops Phase 5 because the user-test gate checks the change out of a worktree and there is none). `:local` and `--local` state it in advance; every autopilot entry resolves it to a worktree and never asks, because an unattended run commits and pushes from wherever it stands and doing that in the user's own checkout is what worktrees exist to prevent. `state.workspaceSource` records who decided - `localMode: false` alone is both "the user chose a worktree" and "nothing asked".
32
+
33
+ Depth is not a flag either. `/multi-agent` and `/multi-agent:local` ask Full or Short at Phase 0 Step 7.5, recommending from the detected `taskType`; Short strips to Init → Dev(Opus self-contained) → Review → Test → Commit → Report. Autopilot never asks and always runs Full - "fast plus unattended" was removed in v16.0.0, because something has to choose when nobody is asked and unattended is the worst place to drop analysis and planning.
32
34
 
33
35
  ### Outside a Pipeline Run
34
36
 
@@ -36,7 +38,7 @@ The install is not only useful while `/multi-agent` is running. `rules/outside-t
36
38
 
37
39
  - **Onboarded service credentials.** Resolve the logical name through `credential-store.sh` and read the issue, page or log. Writes route through the pipeline commands, which carry the rules that make them safe - issues are never auto-closed, PR bodies use `Ref:`, outward prose goes through the humanizer.
38
40
  - **The stack skills enabled for this repo.** Each toolkit's own `index` skill routes. The pipeline reads the effective `enabledPlugins` rather than keeping a stack table, so a seventh toolkit needs no code change.
39
- - **The `multi-agent-toolkit` MCP.** 80+ tools for a running app.
41
+ - **The `multi-agent-toolkit` MCP.** 99 tools for a running app.
40
42
 
41
43
  Uninstall preserves the whole layer - tokens, the reader that opens them, the mapping that names them, the MCP registration. It is 1.5 kB of always-loaded text; the detail lives in a ref that loads on demand, and a gate keeps both under a ceiling because every byte there is paid by every session.
42
44
 
@@ -46,6 +48,8 @@ A deterministic, LLM-free map of what a repo declares and what refers to what, e
46
48
 
47
49
  Phase 1 queries it to hand Explore a ranked starting file set instead of a full scan, and Phase 7 rebuilds it after the branch changed code - a rebuild is seconds, so staleness is a `baseCommit` comparison rather than a date heuristic. Off by default behind `prefs.global.codeGraph.enabled`; with it off the pipeline behaves exactly as before.
48
50
 
51
+ `GRAPH_REPORT.md` also ends with **Symbols nothing else references**: symbols no other file in the repo names, split from the ones referenced only by their own tests. Candidates, never verdicts - the extractor is regex, not a parser, so the four classes that could not carry a reference edge either way (a kind outside the stack's `referenceKinds`, a name declared twice, a nested declaration, a test file) are counted and excluded rather than listed, and nothing gates on the result.
52
+
49
53
  Measured on a 4,300-file Swift app against a grep-and-read baseline at the same 30,000-token retrieval budget: 80.4% key-fact coverage at 18,465 tokens against 66.0% at 24,555. The gain is entirely in searches phrased in domain words (63.3% against 32.0%, at under half the cost). When the task already names an exact type, `grep -lw` is still slightly better and slightly cheaper, and the command says so rather than overselling. Reasoning, trade and limits: `docs/adr/0010-own-code-graph.md`.
50
54
 
51
55
  ### Stack Auto-Detection
@@ -74,6 +78,41 @@ Stack skill sets ship as versioned plugins in the `multi-agent-plugins` marketpl
74
78
  /multi-agent:stack all # every stack plugin
75
79
  ```
76
80
 
81
+ ### Package Manager Resolution (Phase 3, node-shaped stacks)
82
+
83
+ Phase 3's web test arm and its build step used to type `npm`. A repo on pnpm, yarn or bun then failed in Phase 3 - with a worktree and a branch already created - or, worse, npm resolved against a lock file it does not own and the run continued on a tree the repo's own tooling would never have produced.
84
+
85
+ `scripts/package-manager.mjs` resolves it from the repo instead: `$MA_PACKAGE_MANAGER`, then `package.json#packageManager`, then a lock file, then npm - reported AS a default, never as evidence, because "npm because nothing said otherwise" and "npm because the repo committed a package-lock" are different answers. Node core only (ADR-0004): a resolver that shelled out would need a working install of the tool it is identifying. The walk goes up to the directory holding `.git` and stops there, so a monorepo's root lock file is found and a stray one in a home directory is not. Two lock files means a migration left one behind: the newest wins and both are named.
86
+
87
+ Every manager gets the explicit `run` form (a script named `test` or `add` would otherwise lose to the builtin), only npm gets the `--` separator, and `bun run test` never `bun test`. Exit 3 means the repo declares no such script - the `--if-present` case, answered by an exit code rather than a flag whose support differs per manager. The resolved name is pasted into the phase's `eval`, so it is held to the shape a binary actually has, and the gate proves that by eval'ing the produced line with every manager stubbed out. iOS and Android are untouched. `refs/features/package-manager.md`.
88
+
89
+ ### Maturity Follow-Up (`prefs.global.maturityFollowup`)
90
+
91
+ The maturity check has always produced a machine-readable gap list - stable codes in `blockers[]` and `warnings[]` - and then thrown most of it away. A blocker halted the run, an autopilot queue moved to the next item, and the item stayed exactly as immature as it was found. Nobody was told, so nothing changed, so the next scan halted on the same item for the same reason.
92
+
93
+ **Interactive runs ask at the step rather than ending at it**: open the item and fix it, continue without it, or abort. Continuing records which gap was waved through in `state.maturity.accepted[]` - that is what separates an informed continue from a skipped check. An answer typed into a picker improves this run and leaves the item as immature for the next person, so the step offers to write the supplied content back, as a separately approved write.
94
+
95
+ **Autopilot can ask on the item itself**, behind `autopilotCommentsOnIssue` (off by default, because it is an outward-facing write). One comment naming what is missing, then a halt on the circuit breaker with `state.waitingFor = "maturity"` so `resume` re-enters that step. A question, never a state change: no transition, no resolution, no assignee, no label, no close; `Ref:` never `Closes:`; copy in `outputLanguage`, with the gap wording taken verbatim from the fetcher's own summary rather than re-derived.
96
+
97
+ **An edit is a reason to look again, never proof the gap closed.** A reply reading "will do later" moves the timestamp and fixes nothing, so a changed item is re-fetched and re-scored and the check decides. Only a different gap set earns a second comment; "cannot tell whether it moved" re-checks rather than waiting, because folding unknown into "nothing changed" parks a run forever on a tracker that omits the field.
98
+
99
+ Warnings still auto-continue under autopilot - converting them to halts would stall queues on items that ran fine yesterday. `commentOnWarnings` raises them opt-in.
100
+
101
+ ### Base-Branch Evidence (Phase 0 Step 3, `prefs.global.baseBranchEvidence.enabled`)
102
+
103
+ Step 3 used to ask one question with a list it could not vouch for. `git fetch origin` ran, its exit code was discarded, and `git branch -r` printed the remote-tracking cache either way - so on a restricted network a weeks-old local list was presented as the remote's answer, with nothing saying so. And the answer was usually derivable: an issue carrying a target version, or linking a separate issue that represents the release, already names the branch on a repo whose release branches encode the version.
104
+
105
+ `base-branch-candidates.mjs` collects candidates **with the evidence behind each one**, ranks them, and the picker row's description IS the evidence - "matches version 1.51.0 from the field Target Version" and "the repository's default branch" are different answers to the same question. A human still chooses; the derivation only reorders the rows.
106
+
107
+ Nothing is tabled, and that is the design rather than a detail:
108
+
109
+ - **No Jira field id is hardcoded.** A board's "target version" is a custom field whose id differs per instance. What is stable is the schema: any field resolving to type `version` is read, whatever it is called. A linked issue or parent whose own fix-version or summary names a version is the second source, and `baseBranchEvidence.preferLinkedRelease` raises it above the version field on boards where selecting the release issue is what opens the branch.
110
+ - **No branch prefix is tabled.** The release-branch template is inferred from the refs that exist, so one repo yields `<prefix>/develop_<version>` and another `release-<version>` out of the same code. The rule-5 filter carries a version alternative for the same reason: a word list that ranks is fine, one that discards is the prefix table this replaces.
111
+ - **A predicted branch is a note, never an option.** "That version has no branch on the remote yet" is a real answer; an option the user picks has to be checkoutable.
112
+ - **A failed fetch degrades loudly.** `refProvenance` rides on every candidate, the picker says the refs may be stale and gains a retry row, and `phase0-exit-gate.mjs` refuses to close Phase 0 if a degraded fetch recorded its list as `remote`.
113
+
114
+ Autopilot resolves `remembered` → `derived` → `default` and records which fired; `derived` requires issue evidence for the branch it chose. With `baseBranchEvidence.autopilotAsksOnIssue` (**off by default**) an ambiguous derivation posts one comment on the Jira or GitHub issue asking which branch, then halts on circuit-breaker trigger 6 and waits for `resume`. A question, never a state change: no transition, no close, `Ref:` never `Closes:`, copy in `outputLanguage`.
115
+
77
116
  ### Task Type Detection
78
117
 
79
118
  Phase 0 Step 9 classifies every task before Phase 1 starts. Deterministic priority order: Figma URL → instruction file path → git diff heuristic → Jira issue type → branch name → description keywords → user prompt (autopilot defaults to `feature`).
@@ -223,6 +262,8 @@ Phase 3 treats the issue-tracker status update as a required step with a post-mu
223
262
 
224
263
  - **Pre-Commit Secret Detection** (12 patterns): `PreToolUse` hook scans staged files for API keys/tokens, AWS access keys, private keys, `.env` files, service account JSON. Commit **blocked** if found.
225
264
  - **Read-Size Gate** (opt-in, `prefs.global.bulkRead.mode`): a `PreToolUse` hook inspects `Read` and the shell commands that read a file whole. In `observe` it only logs what it would have caught - the baseline you measure before routing anything. In `enforce` a file over `minLines` (default 350) is blocked and delegated to a haiku-rung worker (`bulk-read.sh`), which returns a line-numbered summary so the follow-up is a bounded `Read(offset:limit:)` instead of the whole file; the full text is parked under `.multi-agent/refs/`. The development phase and any file the run has already touched are exempt, because Claude Code's `Edit` requires its own `Read` first.
265
+ - **Capture Hooks** (`SessionEnd`, `PreCompact`, `SessionStart`): every durable write used to live in Phase 7, the phase a run is least likely to reach. `SessionEnd` flushes a run that never got there; `PreCompact` flushes before an auto-compaction summarizes a long phase mid-flight, which is the same loss one level down; `SessionStart` prints at most two lines about an unfinished run. None calls a model, none reads a payload, and all exit 0 on every path - a hook that fails a session over bookkeeping is worse than the bookkeeping.
266
+ - **Operational Reporting** (`prefs.global.usageLog`): coarse run metadata - task id, phase, status, durations, token counts - and never prompts, code, diffs or absolute paths. The per-machine token is REQUESTED from the endpoint by `usage-register.mjs` (setup, update, and the Phase 0 exit gate as a backstop), is write-only, and lives in the OS credential store; prefs hold only the entry name and the switch. `usageLog.optOut: true` blocks registration permanently and is checked before the network call. An unreachable endpoint leaves reporting off with one line and exit 0 - a run is never failed over bookkeeping.
226
267
  - **Build Queue**: All `xcodebuild` calls acquire a lock. Each worktree uses own `-derivedDataPath`. Stale locks auto-clean after 15 min. Non-Xcode builds don't need the lock.
227
268
  - **Context Management**: `CLAUDE_AUTOCOMPACT_PCT_OVERRIDE=65` - compaction at 65% usage (prevents degradation in 8-phase sessions).
228
269
  - **3-Iteration Hard Kill**: Any retry loop stops after 3 attempts, then pauses for user. No infinite loops.
@@ -19,4 +19,4 @@ back into. The ceiling can still be raised; it can no longer be raised for free.
19
19
 
20
20
  ---
21
21
 
22
- Token estimate = ceil(chars / 4). Per-phase budget rule: warn = current+10% (rounded to nearest 50), max = current+25%. Gives ~6 edit cycles of headroom before warn trips - intentionally quiet under normal maintenance, loud when a phase grows unusually. Only the active phase is loaded (lazy). Recalibrated at v10.0.0 after the validator/consistency/simplifier/lesson gate contracts landed in phases 1-4. Recalibrated again at v10.9.0 after the verify-by-test (Phase 4 Step 3.7), update-check (Phase 0 Step 0.6), immutable-test (Phase 3 GREEN) and redTests re-entry contracts landed - Step 3.7 prose was compressed to a pointer into refs/features/verify-by-test.md before the recalibration. Total bumped 50000 -> 51000 at v12.5.0 after the worktree residue/traversal-prune contract (Phase 0 + Phase 5 heal) and the Reflexion causal-diagnosis contract (Phase 4 lesson memory) landed; the prose was compressed first (161 tokens reclaimed) and every per-phase max still passes - only the aggregate needed room. Recalibrated again at v13.6.0 after the install-relative path correction: an instruction that names `pipeline/scripts/x` resolves only from a repo checkout, and a run happens in the user's worktree, so 157 references across these docs moved to `$HOME/.claude/...` at +5 bytes each - 196 tokens of pure correctness cost. Same discipline as before: prose was compressed FIRST (149 tokens reclaimed, by pointing Phase 1's Figma tier table at the Phase 0 probe that already resolved it and Phase 4's Codex constraints at the always-loaded AGENTS.md block), and only then were the budgets moved. Five warn lines had been permanently amber, which makes the amber tier useless as a signal, so every warn was reset to the documented current+10% and the four maxes that the new warn would have collided with were reset to current+25%. Aggregate 51000 -> 51500. Total bumped 51500 -> 52200 at v14.0.0 after Phase 4 Review entered the four --dev mode phase sets and the criteria-resolution contract (Step 1.78) landed. Same discipline as every prior bump: prose was compressed FIRST, 820 tokens reclaimed, before the number moved. Two of those compressions are structural rather than cosmetic - the hardcoded SwiftUI interaction list in Step 1.5 and the SwiftUI convention paragraph in Step 2.8 were transcriptions of rules that now live in a scoped registry, so keeping them here would have re-created the drift this release exists to remove, and the third moved the Step 1.78 full contract into refs/features/skill-conformance.md leaving a pointer. What remains is contract text that cannot be inferred: the manifest's four consumer-visible parts, the conformance checklist the reviewers must return, and the fail-closed semantics. Every per-phase max still passes (phase-4 12405/14750); only the aggregate needed room. Total bumped 52200 -> 52700 at v14.1.0 after two more contracts landed: stack skill routing (Phase 3 pre-flight step 9) and worktree finalize (Phase 6 step 9). Compression came first, as always, and twice: 224 tokens out of Phase 3 by pointing its criteria-ledger and routing steps at their feature files instead of restating them, and 190 out of Phase 6 by moving the finalize contract into refs/features/worktree-finalize.md and leaving the invocation plus the exit-3 semantics. Both new contracts follow the pattern the earlier ones set: the phase doc carries the call and the decision, the feature file carries the reasoning, and the feature files are outside this budget because it loops only the eight phase-N-* keys. Every per-phase max still passes (phase-3 7677/8950, phase-6 5223/6150 and both under warn); only the aggregate needed room. Total bumped 52700 -> 52750 for the Phase 0 Step 3 branch-persistence correction: the step wrote the legacy `projects[].branches` while the TTL filter two sections below read `global.recentBranches`, and both spots named a `{name, lastUsed}` shape the schema rejects (`branch` required, `additionalProperties: false`), so the recent-branch picker option could never populate and a literal implementation would have failed prefs validation. Naming the right target, the right key and the legacy field to avoid costs 41 tokens over the one line it replaces. Compression came first and was applied three times to the replacement text itself, from 120 tokens down to 66, by moving the rationale out of the phase doc entirely: the reasoning now lives where it is enforced, in the migrate-prefs carry-forward comment and the smoke-pref-migration f7 block, leaving the phase doc with only the instruction. 50 was the smallest step that clears it; phase-0-init sits at 10893/12400, far under its own max, so this is purely an aggregate ceiling. v15.0.0: total 52750 -> 53100, the stack-skill tables in phase-1/2/4 now carry plugin-namespaced names (ai-<stack>-toolkit:<skill>) - functional prefixes, ~170 tokens. v15.10.0: total 53350 -> 53950 for the memory-recall + context-offload contracts (Phase 1 two-block durable-knowledge injection and its telemetry, Phase 3 build-log offload pipe, Phase 4 ranked prior art, offload pipe and recall telemetry). Compression came first and twice, taking the new prose from 1168 tokens to 580: the reasoning behind the two blocks lives in multi-agent-refs/prompt-assembly.md and the reasoning behind the offload filter lives in the offload-ref.sh header, both outside this budget, so the phase docs carry only the call, the pref that gates it and the one fact an agent cannot infer - that the evidence gate still reads the whole build log, so offloading changes what is read, never what counts as a verified pass. Every per-phase max still passes (phase-3 7985/8950, phase-4 12997/14750); phase-3 and phase-4 crossed their warn lines and are left amber on purpose, because that is the signal that those two docs are the next ones needing structural compression rather than another bump. v15.13.0: total 53950 -> 54050 for the prefs-to-flag bridges. Five settings had shipped declared-but-inert: contextOffload.minLines and .tailLines (fixed in 15.11.0), learningsLedger.maxBriefEntries, and testGap.scanTree and .promoteSeverity - the last two declared in the schema AND implemented as flags in the scanner, with nothing in between reading the pref and passing the flag. Wiring three of them costs the phase docs 94 tokens, which is the wiring itself and not prose: two `--max` substitutions and a three-line GAP_FLAGS block. Compression came first and twice, as always: the rationale that would have sat in phase-5 now lives in the header of smoke-prefs-consumed.sh, the gate that makes this class fail a build instead of shipping, and a `--severity-promote` table row was dropped because the invocation above it now shows the flag and names the pref that triggers it, which the row did not. 100 was the smallest step that clears it. Every per-phase max still passes; phase-3 and phase-4 remain amber on purpose. v15.14.0: total 54050 -> 54400 for the supported-version gate. Phase 0 Step 0.6 stopped being purely advisory: a release can now publish an npm dist-tag `required` that names the oldest runnable version, and below it the run halts instead of nagging. What the phase doc has to carry is the part an agent cannot infer - the third stdout field, that the halt is identical in autopilot, and that the run must NOT continue on the freshly updated install because its docs were already loaded from the old version. Compression came first, as always, and took the new prose from 469 tokens to 337: the rationale for the floor, the exemption list, the fail-open rules and the `npm dist-tag add` recipe all moved to multi-agent-refs/rules.md "Supported Version Gate" (loaded by 25 commands, outside this budget) and to the header of require-supported-version.sh, leaving the phase doc with the call, the decision table and the halt. 350 was the smallest step that clears it. Every per-phase max still passes (phase-0-init 11230/12400); phase-3 and phase-4 remain amber on purpose. v15.17.0: total 54400 -> 54900 for the Phase 1 analysis-document step. Phase 2 and Phase 3 pre-flights had BLOCKED on `analysis/<feature>-<platform>.md` since v9.0.0 while nothing produced it, so a full run either aborted at Phase 2 or the model ignored its own BLOCKING contract; Step 4 is the producer. What the phase doc carries is only what cannot be inferred: the when-table (taskType x Figma reference), the four refs in load order, the two artefacts, and that the doc validator fails closed. Compression came first and took the step from 745 tokens to 497: the history of why the gap existed moved to the CHANGELOG, the per-ref one-line descriptions moved into the refs' own headers, and the autopilot carve-out collapsed to one clause. The 17.4k-token analysis engine itself is NOT in this budget - it moved out of commands/ into multi-agent-refs/analysis/{locked,evidence,synthesis,render}.md, loaded on demand, which also took analysis/SKILL.md from 18081 to 5974 tokens and retired its lint grace entry. 500 was the smallest step that clears it; phase-1-analysis sits at 4338/4600 and is amber on purpose, like phase-3 and phase-4. v15.18.0: total 54900 -> 55250 for analysis mode. Three phase docs gained a mode branch that cannot be inferred: Phase 4 reviews a document instead of a diff (validator, the one question reviewers answer, the open-question walk), and Phase 6 publishes instead of committing. Compression came first and was applied twice to the new prose and once to old: the Phase 4 branch went from 320 tokens to 180 and the Phase 6 branch from 190 to 120 by pointing at multi-agent-refs/analysis/{resolve,render}.md, which now hold the walks themselves, and the front-matter parse contract stopped being spelled out in both pre-flights. The analysis engine keeps leaving this budget rather than entering it: intake joined locked/evidence/synthesis/render/resolve in multi-agent-refs/analysis/, which is what let analysis/SKILL.md drop under the 6000 hard cap after its grace entry was retired. 350 was the smallest step that clears it; phase-4 and phase-6 are amber on purpose, as phase-1 and phase-3 already were. v15.20.0: total 55250 -> 55500 for the TDD bridge. Phase 3 pre-flight read the analysis doc's concept table and even said test method names come from it, while nothing read Section 15 - so the RED step invented tests and the analysis test matrix never reached development. Phase 3 step 5b now loads it into state.dev.testPlan[] and Phase 4 step 1.45 cross-checks every planned row against a real test, which is what turns "analysis quality is output quality" from a slogan into a finding. Compression came first on both blocks, 300 tokens down to 175, by dropping the enumerated failure modes to one line each and the rationale to one clause; the reasoning lives in the CHANGELOG. 250 was the smallest step that clears it. v15.21.0: total 55500 -> 55800 for the post-analysis confirmation. Phase 2 gained Step 0.9, the last human checkpoint before Phase 3: derived values are shown for confirmation and only Section 20 rows are asked, through the resolve engine that already exists in refs. It belongs here rather than Phase 4 because Phase 4 runs after development, where an answer arrives too late to change anything. Compression came first and twice, 430 tokens down to 250, by collapsing the derived-vs-asked explanation to one sentence each and moving the walk itself to multi-agent-refs/analysis/resolve.md, which Phase 4 and analysis-resolve already mount. 300 was the smallest step that clears it. v15.22.0: total 55800 -> 55900 for the analyst-toolkit hooks. Phase 1 Step 4 now names the two prefs that decide whether a document is produced at all and how deep it goes (forceFull, mode) - the first of those had shipped declared-but-inert and smoke-prefs-consumed caught it - and Phase 4 triage gained one clause: a finding that blames a third-party library asks evidence-github whether it is already open upstream, which turns it into a deferred item with a citation instead of Phase 3 rework on code that is not ours. Compression came first and three times, taking the new prose from 220 tokens to 110, and the Phase 1d evidence contract itself never entered this budget - it lives in multi-agent-refs/analysis/evidence.md beside the phases it belongs to. 100 was the smallest step that clears it, leaving 34 tokens of headroom. phase-4 stays amber and the debt named at v15.10.0 stands: it is the doc that needs structural compression rather than another bump, and the two candidates are the inline triage JSON shape and the 3.4 telemetry block, both of which restate something already authoritative elsewhere. v16.0.0: total 55900 -> 56350 for the depth picker. `--dev` and the four dev-* commands are gone; depth is Phase 0 Step 7.5, which costs phase-0-init a step it did not have. Compression came first and three times, taking the step from 530 tokens to 300: the question wording, the per-taskType recommendation and the mode tables all live in phases/modes.md (outside this budget), so the phase doc carries only what an agent cannot infer - that the step runs after Step 7 and why, who is exempt, that ASK_CHOICE_DEFAULT must be passed explicitly because ask-choice.sh takes the FIRST option on a non-TTY, and that Short flips the Phase 1/2 tiles late rather than pre-marking them. The phase-4 telemetry block named as compression debt at v15.22.0 was collapsed to an emit() helper (-27) and the four dev-* mode files left the tree entirely, but neither offsets a genuinely new phase step. 450 was the smallest step that clears it, leaving 119 tokens of headroom. phase-4 remains amber and its other named candidate, the inline triage JSON shape, was left alone on purpose: it is the prompt the triage agent is handed, not a restatement for readers. v16.2.0: total 56350 -> 56600 for the spec-freshness and reuse-tag contracts. Phase 3 step 3 had compared `state.run.lastAnalysisDigest` since it was written, against a key nothing ever set and that the state schema did not declare, so the staleness branch was unreachable and every run reported fresh by default. Phase 1 now persists the digest and a `base_commit` anchor, and step 3 gained the repo-drift half the digest cannot see: a reused document keeps a matching digest precisely because its evidence inputs did not change, while the code underneath it moved. The second contract is the Section 14 tag reaching development: Phase 2 carries it onto the todo as `sourceTag` and Phase 3 treats it as an instruction, which is what stops a Reuse row from being re-implemented. Compression came first and took the four additions from 380 tokens to 214, by moving every rationale clause out of the phase docs: why the commit anchor exists rather than a digest recomputation lives in this note and the CHANGELOG, and the schema descriptions carry the field semantics. The baseline had 9 tokens of headroom, so no addition of any size could have fit without a bump. 250 was the smallest step that clears it, leaving 45 tokens. phase-3 and phase-4 remain amber. v16.13.0: total 57600 -> 57700 for the code-graph injection and the fable-rung switch. Phase 1 gained Step 2.6 (query the graph, hand Explore a ranked starting set), Phase 7 gained the post-branch graph refresh, and Phase 0 Step 0 gained one line: a prefs switch that resolves every preferredModel: fable persona to opus for the run, which also collapses the Phase 4 Claude Code panel from three reviewers to two. Compression came first and mostly structurally: of roughly 1,630 tokens of new contract text, 1,310 never entered this budget at all - the whole code-graph contract lives in multi-agent-refs/features/code-graph.md (604) and the fable switch's scope table, per-host effects and cost-accounting consequence live in features/model-fallback.md (+707), leaving the phase docs with the call, the pref that gates it and the one fact an agent cannot infer. Phase 4 was compressed on top of that: its TLDR restated the reviewer matrix 270 lines below it, so 36 tokens came back and the doc nets +6 despite carrying two new clauses. One of those clauses is a correction rather than a feature - the consensus rule still said reviewerCount is 2 on Claude Code, which stopped being true when the third reviewer landed in 16.12.0, and the cross-CLI smoke never caught it because it reads the matrix line instead. 100 was the smallest step that clears it, leaving 54 tokens. phase-3 and phase-4 remain amber. v16.17.0: total 57700 -> 57850 for the platform-parity cross-check. Phase 4 gained Step 1.8: when dev-context carries a counterpart app repo, the review compares the change against the other platform on four axes. Compression came first and structurally, as always - of roughly 1,610 tokens of new contract text, 1,490 never entered this budget at all, because the four axes, the file cap, the graph-query recipe, the read-only prohibitions and the rule that an extractor miss may not be reported as an absence all live in multi-agent-refs/platform-parity.md. The step itself was then cut from ~200 tokens to 120 by deleting everything the ref already owns, leaving the trigger, the pointer and the two facts an agent must not infer: the counterpart repo is read-only, and parity findings are never blocking. The baseline had 13 tokens of headroom, so no addition of any size could have fit without a bump. 150 was the smallest step that clears it, leaving 35 tokens. phase-3 and phase-4 remain amber, and phase-4's structural-compression debt still stands. v16.20.0: phase-4-review max 14750 -> 15150 and total 58250 -> 60250 for the cross-round review delta, the scope self-check handoff and the circuit-breaker wiring. Compression came first and structurally: of roughly 3,900 tokens of new contract text, 2,700 never entered this budget at all - the previous-round-findings block, the scope-self-check block, the Step 3.8 state merge, telemetry and picker wording live in multi-agent-refs/features/review-delta.md, and the scope-check record rules and consumers in features/scope-check.md - so the phase docs carry the call, the pref that gates it and the exit table. The Phase 3 stability rule and the trigger-3 write were cut twice more before the bump; phase-3 stays under its max (8692/8950). Phase 4 is the first per-phase max raised since v10.9.0: the doc gained three steps that cannot be inferred (a per-round triage file, a prefix block that changes what reviewers report, and a halt condition), and its structural-compression debt (the inline triage JSON shape, named at v15.10.0) still stands and is the next candidate. 15150 and 60250 were the smallest steps that clear it, leaving 25 and 45 tokens. v16.23.0: phase-0-init max 12400 -> 12500 and total 60250 -> 60500 for the widget-registration call and the accounting gate. Phase 0 gained the `tiles` call and the exit-3 rule, Phase 7 gained the run report; together they are contract an agent cannot infer - which call registers this host's widget, and that a completion is refused without recorded spend. Compression came first and twice, taking the new prose from 472 tokens to 255: the per-host call list moved into tracker-contract.md "The card is not the widget" and the record-then-rerun recovery into "Accounting is a gate", both outside this budget, leaving the phase docs with the call and the one fact that cannot be looked up. 100 and 250 were the smallest steps that clear it, leaving 74 and 40 tokens. phase-3 and phase-4 remain amber. v16.24.0: total 60500 -> 60750 for visual evidence. Four phase docs gained one instruction each that cannot be inferred: Phase 0 keeps the issue's own images as the pre-fix evidence, Phase 3 captures the fixed state (there and not Phase 5, because every autopilot and --local entry drops Phase 5), Phase 5 hosts the flow recording when it runs, and Phase 6 blocks on a required artefact that is neither attached nor explained. Compression came first and twice, 42 tokens back, and the contract itself never entered this budget: the trigger matrix, the three video tiers, the size-degradation ladder and both render shapes live in multi-agent-refs/features/visual-evidence.md. phase-0-init cleared its own max without a bump. 250 was the smallest step that clears the aggregate. phase-3 and phase-4 remain amber. v17.0.0: phase-0-init max 13000 -> 13100 and total 62400 -> 62500 for the evidence-verdict writer. Phase 0 Step 7.7 probed with `--platform "$PLATFORM"`, a variable no phase document ever assigned, and it was gated on `visualEvidence.required`, which no phase document ever wrote - five readers, zero writers - so the step, Phase 3's capture and Phase 6's blocker were all unreachable and the pipeline reported nothing wrong. The step now writes the verdict and derives the platform from the stack, skipping the probe with a recorded reason when there is no device platform rather than passing the empty string the probe refuses with exit 2. Compression came first and three times, 30 tokens back from the Step 7.7 index rule that restated Step 7.5 verbatim and 55 from the new block itself; the reasoning never entered this budget, because who writes the verdict and how the platform is derived live in features/visual-evidence.md sections 1a and 1b. phase-3-dev max 9250 -> 9300 in the same change: it reads the platform back from state and re-decides the provisional verdict before capturing, which is the half of the fix that makes Phase 3 honest rather than merely reachable. Compressed three times first, 29 tokens back, by pointing its Phase-5 rationale and its tier mapping at visual-evidence.md sections 3 and 4.3 where both already live. 100, 50 and 150 were the smallest steps that clear it, leaving 24, 15 and 26 tokens. phase-3 and phase-4 remain amber. v17.1.0: total 62550 -> 62600 for the Figma/toolkit MCP distinction and web as an evidence platform. Three phase docs said "MCP forbidden" without qualifying it, while the gate that enforces it (smoke-no-mcp-in-dev-phases.sh) has always matched `figma` and nothing else - so the prose banned the screenshot, xcodebuild and UI-test tools that Phase 3 Steps 3.4 and 3.55 actually call, which is one way a run reaches Phase 7 with no evidence. Phase 0 gained one `web` arm in the platform derivation, now that run-ui-tests.sh has a web arm to derive it for. Compression came first and three times, taking the new prose from 175 tokens to 30: the reasoning moved to rules.md, whose own seven-row Figma phase matrix was deleted in the same pass because it duplicated rules/figma-pipeline.md "Phase access matrix" two paragraphs below this file's own instruction not to duplicate that rule file - 74 tokens back there, which is why phase-3-dev cleared its max without a bump (9298/9300). 50 was the smallest step that clears the aggregate, leaving 37 tokens. phase-3 and phase-4 remain amber. v17.1.0 (2): total 62600 -> 62700 for the plan reaching the task widget. Phase 2 computed tasks[], their order and their dependsOn[] edges, stored them, and used them to drive Phase 3's ready-task picker - and none of it was visible on the surface the user actually watches; the card had drawn sub-phases for releases, the widget never had. Phase 2 gains one call. Compression came first and twice: the tasks[]-to-todos[] jq blob left the phase doc for plan-todos.sh `set`, which now accepts a planning-output document directly (-28, and it removes a mapping two files defined, of which this was the untested copy), and the new step's own prose was cut from 116 tokens to 61. The parsing of the plan itself never entered this budget - it lives in phase-tracker.sh `plan`, which owns the sub-phase structure it writes. 100 was the smallest step that clears it, leaving 75 tokens. phase-3 and phase-4 remain amber. v17.5.0: phase-0-init max 13100 -> 13150 for the deferred widget registration, the dev-context step and two more exit-gate assertions. Compression came first and seven times, 492 tokens reclaimed before the number moved: the TLDR restated the step headings under it, the exit-gate list restated the script's own header comment, and the clarifier cost note, the branch-collision probe rationale, the fetch-fail host explanation, the multi-repo write-state race and the baseline unknown-vs-green warning all restate something already authoritative in agents/task-clarifier.md, phases/operations.md or phase0-exit-gate.mjs. The three additions are contract an agent cannot infer: that only Phase 0 is registered at Step -1 and the rest at 7.5 with `tiles --new`, that the dev-context picker runs between project and branch and writes siblings[] even when empty, and that a one-option AskUserQuestion is refused by the host along with every question batched with it. The reasoning for the last one never entered this budget - it lives in picker-contract.md "Two options or it is not a question". 50 was the smallest step that clears it, leaving 28 tokens; the aggregate needed no bump and sits at 62698/62700. phase-3 and phase-4 remain amber. v17.5.0 (2): phase-0-init max 13150 -> 13300 and total 62700 -> 62900 for the workspace question. Phase 0 gained Step 5b: where the branch lives was decided by a flag and by a command name, never by a question, so a user who saw a run improvise "Worktree / Lokal" once took it for a shipped picker and concluded /multi-agent:local was redundant - while it was the only way to reach local mode at all. Compression came first and three times, 156 tokens reclaimed before the number moved and the step itself cut from 426 tokens to 181: the ask-choice index rule in Step 7.5 restated modes.md "Pipeline depth" in full, and the question wording, the two options and what local costs now live in modes.md "Local Mode", outside this budget. What the phase doc carries is what an agent cannot infer - that it runs after Step 4 and before Step 6b and 7.5, who is exempt, and that autopilot resolves it to a worktree rather than skipping it, because an unattended commit in the user's own checkout is what worktrees exist to prevent. The state field is the other half: localMode alone cannot say whether anyone decided, since false is both a chosen worktree and one nothing asked about, and phase0-exit-gate.mjs now refuses to close Phase 0 without workspaceSource. 150 and 200 were the smallest steps that clear it, leaving 32 and 56 tokens. phase-3 and phase-4 remain amber. v17.5.0 (3): phase-0-init max 13300 -> 13400 and total 62900 -> 63000 for base-branch evidence. Step 3 asked one question with a list it could not vouch for: `git fetch origin` ran, its exit code was discarded, and `git branch -r` printed the remote-tracking cache either way, so a restricted network produced a weeks-old local list presented as the remote's answer. And the answer was usually derivable - an issue carrying a target version, or linking a separate issue that represents the release, already names the branch on a repo whose release branches encode the version - and nothing derived it. Rule 5 now captures the exit code and rule 5b hands the list plus the issue's version fields and links to base-branch-candidates.mjs. Compression came first and six times, 1,118 characters reclaimed across this step before the number moved: the one-row picker rule and the not-skippable rule both restated picker-contract.md, the autopilot resolution order and the whole rationale for admitting version-carrying branches moved into features/base-branch-evidence.md, and the fetch-fail persistence block was four sentences for one instruction. The feature's own reasoning never entered this budget at all - the evidence sources, the learned convention, the picker contract, the autopilot fencing and the state shape are ~1,900 tokens living in that ref. What the phase doc carries is the call and the two facts an agent cannot infer: that the fetch exit code decides whether the list may be called remote, and that the branch filter needs a version alternative because a word list that discards is the prefix table this replaces. 100 was the smallest step that clears both, leaving 51 and 61 tokens. phase-3 and phase-4 remain amber.
22
+ Token estimate = ceil(chars / 4). Per-phase budget rule: warn = current+10% (rounded to nearest 50), max = current+25%. Gives ~6 edit cycles of headroom before warn trips - intentionally quiet under normal maintenance, loud when a phase grows unusually. Only the active phase is loaded (lazy). Recalibrated at v10.0.0 after the validator/consistency/simplifier/lesson gate contracts landed in phases 1-4. Recalibrated again at v10.9.0 after the verify-by-test (Phase 4 Step 3.7), update-check (Phase 0 Step 0.6), immutable-test (Phase 3 GREEN) and redTests re-entry contracts landed - Step 3.7 prose was compressed to a pointer into refs/features/verify-by-test.md before the recalibration. Total bumped 50000 -> 51000 at v12.5.0 after the worktree residue/traversal-prune contract (Phase 0 + Phase 5 heal) and the Reflexion causal-diagnosis contract (Phase 4 lesson memory) landed; the prose was compressed first (161 tokens reclaimed) and every per-phase max still passes - only the aggregate needed room. Recalibrated again at v13.6.0 after the install-relative path correction: an instruction that names `pipeline/scripts/x` resolves only from a repo checkout, and a run happens in the user's worktree, so 157 references across these docs moved to `$HOME/.claude/...` at +5 bytes each - 196 tokens of pure correctness cost. Same discipline as before: prose was compressed FIRST (149 tokens reclaimed, by pointing Phase 1's Figma tier table at the Phase 0 probe that already resolved it and Phase 4's Codex constraints at the always-loaded AGENTS.md block), and only then were the budgets moved. Five warn lines had been permanently amber, which makes the amber tier useless as a signal, so every warn was reset to the documented current+10% and the four maxes that the new warn would have collided with were reset to current+25%. Aggregate 51000 -> 51500. Total bumped 51500 -> 52200 at v14.0.0 after Phase 4 Review entered the four --dev mode phase sets and the criteria-resolution contract (Step 1.78) landed. Same discipline as every prior bump: prose was compressed FIRST, 820 tokens reclaimed, before the number moved. Two of those compressions are structural rather than cosmetic - the hardcoded SwiftUI interaction list in Step 1.5 and the SwiftUI convention paragraph in Step 2.8 were transcriptions of rules that now live in a scoped registry, so keeping them here would have re-created the drift this release exists to remove, and the third moved the Step 1.78 full contract into refs/features/skill-conformance.md leaving a pointer. What remains is contract text that cannot be inferred: the manifest's four consumer-visible parts, the conformance checklist the reviewers must return, and the fail-closed semantics. Every per-phase max still passes (phase-4 12405/14750); only the aggregate needed room. Total bumped 52200 -> 52700 at v14.1.0 after two more contracts landed: stack skill routing (Phase 3 pre-flight step 9) and worktree finalize (Phase 6 step 9). Compression came first, as always, and twice: 224 tokens out of Phase 3 by pointing its criteria-ledger and routing steps at their feature files instead of restating them, and 190 out of Phase 6 by moving the finalize contract into refs/features/worktree-finalize.md and leaving the invocation plus the exit-3 semantics. Both new contracts follow the pattern the earlier ones set: the phase doc carries the call and the decision, the feature file carries the reasoning, and the feature files are outside this budget because it loops only the eight phase-N-* keys. Every per-phase max still passes (phase-3 7677/8950, phase-6 5223/6150 and both under warn); only the aggregate needed room. Total bumped 52700 -> 52750 for the Phase 0 Step 3 branch-persistence correction: the step wrote the legacy `projects[].branches` while the TTL filter two sections below read `global.recentBranches`, and both spots named a `{name, lastUsed}` shape the schema rejects (`branch` required, `additionalProperties: false`), so the recent-branch picker option could never populate and a literal implementation would have failed prefs validation. Naming the right target, the right key and the legacy field to avoid costs 41 tokens over the one line it replaces. Compression came first and was applied three times to the replacement text itself, from 120 tokens down to 66, by moving the rationale out of the phase doc entirely: the reasoning now lives where it is enforced, in the migrate-prefs carry-forward comment and the smoke-pref-migration f7 block, leaving the phase doc with only the instruction. 50 was the smallest step that clears it; phase-0-init sits at 10893/12400, far under its own max, so this is purely an aggregate ceiling. v15.0.0: total 52750 -> 53100, the stack-skill tables in phase-1/2/4 now carry plugin-namespaced names (ai-<stack>-toolkit:<skill>) - functional prefixes, ~170 tokens. v15.10.0: total 53350 -> 53950 for the memory-recall + context-offload contracts (Phase 1 two-block durable-knowledge injection and its telemetry, Phase 3 build-log offload pipe, Phase 4 ranked prior art, offload pipe and recall telemetry). Compression came first and twice, taking the new prose from 1168 tokens to 580: the reasoning behind the two blocks lives in multi-agent-refs/prompt-assembly.md and the reasoning behind the offload filter lives in the offload-ref.sh header, both outside this budget, so the phase docs carry only the call, the pref that gates it and the one fact an agent cannot infer - that the evidence gate still reads the whole build log, so offloading changes what is read, never what counts as a verified pass. Every per-phase max still passes (phase-3 7985/8950, phase-4 12997/14750); phase-3 and phase-4 crossed their warn lines and are left amber on purpose, because that is the signal that those two docs are the next ones needing structural compression rather than another bump. v15.13.0: total 53950 -> 54050 for the prefs-to-flag bridges. Five settings had shipped declared-but-inert: contextOffload.minLines and .tailLines (fixed in 15.11.0), learningsLedger.maxBriefEntries, and testGap.scanTree and .promoteSeverity - the last two declared in the schema AND implemented as flags in the scanner, with nothing in between reading the pref and passing the flag. Wiring three of them costs the phase docs 94 tokens, which is the wiring itself and not prose: two `--max` substitutions and a three-line GAP_FLAGS block. Compression came first and twice, as always: the rationale that would have sat in phase-5 now lives in the header of smoke-prefs-consumed.sh, the gate that makes this class fail a build instead of shipping, and a `--severity-promote` table row was dropped because the invocation above it now shows the flag and names the pref that triggers it, which the row did not. 100 was the smallest step that clears it. Every per-phase max still passes; phase-3 and phase-4 remain amber on purpose. v15.14.0: total 54050 -> 54400 for the supported-version gate. Phase 0 Step 0.6 stopped being purely advisory: a release can now publish an npm dist-tag `required` that names the oldest runnable version, and below it the run halts instead of nagging. What the phase doc has to carry is the part an agent cannot infer - the third stdout field, that the halt is identical in autopilot, and that the run must NOT continue on the freshly updated install because its docs were already loaded from the old version. Compression came first, as always, and took the new prose from 469 tokens to 337: the rationale for the floor, the exemption list, the fail-open rules and the `npm dist-tag add` recipe all moved to multi-agent-refs/rules.md "Supported Version Gate" (loaded by 25 commands, outside this budget) and to the header of require-supported-version.sh, leaving the phase doc with the call, the decision table and the halt. 350 was the smallest step that clears it. Every per-phase max still passes (phase-0-init 11230/12400); phase-3 and phase-4 remain amber on purpose. v15.17.0: total 54400 -> 54900 for the Phase 1 analysis-document step. Phase 2 and Phase 3 pre-flights had BLOCKED on `analysis/<feature>-<platform>.md` since v9.0.0 while nothing produced it, so a full run either aborted at Phase 2 or the model ignored its own BLOCKING contract; Step 4 is the producer. What the phase doc carries is only what cannot be inferred: the when-table (taskType x Figma reference), the four refs in load order, the two artefacts, and that the doc validator fails closed. Compression came first and took the step from 745 tokens to 497: the history of why the gap existed moved to the CHANGELOG, the per-ref one-line descriptions moved into the refs' own headers, and the autopilot carve-out collapsed to one clause. The 17.4k-token analysis engine itself is NOT in this budget - it moved out of commands/ into multi-agent-refs/analysis/{locked,evidence,synthesis,render}.md, loaded on demand, which also took analysis/SKILL.md from 18081 to 5974 tokens and retired its lint grace entry. 500 was the smallest step that clears it; phase-1-analysis sits at 4338/4600 and is amber on purpose, like phase-3 and phase-4. v15.18.0: total 54900 -> 55250 for analysis mode. Three phase docs gained a mode branch that cannot be inferred: Phase 4 reviews a document instead of a diff (validator, the one question reviewers answer, the open-question walk), and Phase 6 publishes instead of committing. Compression came first and was applied twice to the new prose and once to old: the Phase 4 branch went from 320 tokens to 180 and the Phase 6 branch from 190 to 120 by pointing at multi-agent-refs/analysis/{resolve,render}.md, which now hold the walks themselves, and the front-matter parse contract stopped being spelled out in both pre-flights. The analysis engine keeps leaving this budget rather than entering it: intake joined locked/evidence/synthesis/render/resolve in multi-agent-refs/analysis/, which is what let analysis/SKILL.md drop under the 6000 hard cap after its grace entry was retired. 350 was the smallest step that clears it; phase-4 and phase-6 are amber on purpose, as phase-1 and phase-3 already were. v15.20.0: total 55250 -> 55500 for the TDD bridge. Phase 3 pre-flight read the analysis doc's concept table and even said test method names come from it, while nothing read Section 15 - so the RED step invented tests and the analysis test matrix never reached development. Phase 3 step 5b now loads it into state.dev.testPlan[] and Phase 4 step 1.45 cross-checks every planned row against a real test, which is what turns "analysis quality is output quality" from a slogan into a finding. Compression came first on both blocks, 300 tokens down to 175, by dropping the enumerated failure modes to one line each and the rationale to one clause; the reasoning lives in the CHANGELOG. 250 was the smallest step that clears it. v15.21.0: total 55500 -> 55800 for the post-analysis confirmation. Phase 2 gained Step 0.9, the last human checkpoint before Phase 3: derived values are shown for confirmation and only Section 20 rows are asked, through the resolve engine that already exists in refs. It belongs here rather than Phase 4 because Phase 4 runs after development, where an answer arrives too late to change anything. Compression came first and twice, 430 tokens down to 250, by collapsing the derived-vs-asked explanation to one sentence each and moving the walk itself to multi-agent-refs/analysis/resolve.md, which Phase 4 and analysis-resolve already mount. 300 was the smallest step that clears it. v15.22.0: total 55800 -> 55900 for the analyst-toolkit hooks. Phase 1 Step 4 now names the two prefs that decide whether a document is produced at all and how deep it goes (forceFull, mode) - the first of those had shipped declared-but-inert and smoke-prefs-consumed caught it - and Phase 4 triage gained one clause: a finding that blames a third-party library asks evidence-github whether it is already open upstream, which turns it into a deferred item with a citation instead of Phase 3 rework on code that is not ours. Compression came first and three times, taking the new prose from 220 tokens to 110, and the Phase 1d evidence contract itself never entered this budget - it lives in multi-agent-refs/analysis/evidence.md beside the phases it belongs to. 100 was the smallest step that clears it, leaving 34 tokens of headroom. phase-4 stays amber and the debt named at v15.10.0 stands: it is the doc that needs structural compression rather than another bump, and the two candidates are the inline triage JSON shape and the 3.4 telemetry block, both of which restate something already authoritative elsewhere. v16.0.0: total 55900 -> 56350 for the depth picker. `--dev` and the four dev-* commands are gone; depth is Phase 0 Step 7.5, which costs phase-0-init a step it did not have. Compression came first and three times, taking the step from 530 tokens to 300: the question wording, the per-taskType recommendation and the mode tables all live in phases/modes.md (outside this budget), so the phase doc carries only what an agent cannot infer - that the step runs after Step 7 and why, who is exempt, that ASK_CHOICE_DEFAULT must be passed explicitly because ask-choice.sh takes the FIRST option on a non-TTY, and that Short flips the Phase 1/2 tiles late rather than pre-marking them. The phase-4 telemetry block named as compression debt at v15.22.0 was collapsed to an emit() helper (-27) and the four dev-* mode files left the tree entirely, but neither offsets a genuinely new phase step. 450 was the smallest step that clears it, leaving 119 tokens of headroom. phase-4 remains amber and its other named candidate, the inline triage JSON shape, was left alone on purpose: it is the prompt the triage agent is handed, not a restatement for readers. v16.2.0: total 56350 -> 56600 for the spec-freshness and reuse-tag contracts. Phase 3 step 3 had compared `state.run.lastAnalysisDigest` since it was written, against a key nothing ever set and that the state schema did not declare, so the staleness branch was unreachable and every run reported fresh by default. Phase 1 now persists the digest and a `base_commit` anchor, and step 3 gained the repo-drift half the digest cannot see: a reused document keeps a matching digest precisely because its evidence inputs did not change, while the code underneath it moved. The second contract is the Section 14 tag reaching development: Phase 2 carries it onto the todo as `sourceTag` and Phase 3 treats it as an instruction, which is what stops a Reuse row from being re-implemented. Compression came first and took the four additions from 380 tokens to 214, by moving every rationale clause out of the phase docs: why the commit anchor exists rather than a digest recomputation lives in this note and the CHANGELOG, and the schema descriptions carry the field semantics. The baseline had 9 tokens of headroom, so no addition of any size could have fit without a bump. 250 was the smallest step that clears it, leaving 45 tokens. phase-3 and phase-4 remain amber. v16.13.0: total 57600 -> 57700 for the code-graph injection and the fable-rung switch. Phase 1 gained Step 2.6 (query the graph, hand Explore a ranked starting set), Phase 7 gained the post-branch graph refresh, and Phase 0 Step 0 gained one line: a prefs switch that resolves every preferredModel: fable persona to opus for the run, which also collapses the Phase 4 Claude Code panel from three reviewers to two. Compression came first and mostly structurally: of roughly 1,630 tokens of new contract text, 1,310 never entered this budget at all - the whole code-graph contract lives in multi-agent-refs/features/code-graph.md (604) and the fable switch's scope table, per-host effects and cost-accounting consequence live in features/model-fallback.md (+707), leaving the phase docs with the call, the pref that gates it and the one fact an agent cannot infer. Phase 4 was compressed on top of that: its TLDR restated the reviewer matrix 270 lines below it, so 36 tokens came back and the doc nets +6 despite carrying two new clauses. One of those clauses is a correction rather than a feature - the consensus rule still said reviewerCount is 2 on Claude Code, which stopped being true when the third reviewer landed in 16.12.0, and the cross-CLI smoke never caught it because it reads the matrix line instead. 100 was the smallest step that clears it, leaving 54 tokens. phase-3 and phase-4 remain amber. v16.17.0: total 57700 -> 57850 for the platform-parity cross-check. Phase 4 gained Step 1.8: when dev-context carries a counterpart app repo, the review compares the change against the other platform on four axes. Compression came first and structurally, as always - of roughly 1,610 tokens of new contract text, 1,490 never entered this budget at all, because the four axes, the file cap, the graph-query recipe, the read-only prohibitions and the rule that an extractor miss may not be reported as an absence all live in multi-agent-refs/platform-parity.md. The step itself was then cut from ~200 tokens to 120 by deleting everything the ref already owns, leaving the trigger, the pointer and the two facts an agent must not infer: the counterpart repo is read-only, and parity findings are never blocking. The baseline had 13 tokens of headroom, so no addition of any size could have fit without a bump. 150 was the smallest step that clears it, leaving 35 tokens. phase-3 and phase-4 remain amber, and phase-4's structural-compression debt still stands. v16.20.0: phase-4-review max 14750 -> 15150 and total 58250 -> 60250 for the cross-round review delta, the scope self-check handoff and the circuit-breaker wiring. Compression came first and structurally: of roughly 3,900 tokens of new contract text, 2,700 never entered this budget at all - the previous-round-findings block, the scope-self-check block, the Step 3.8 state merge, telemetry and picker wording live in multi-agent-refs/features/review-delta.md, and the scope-check record rules and consumers in features/scope-check.md - so the phase docs carry the call, the pref that gates it and the exit table. The Phase 3 stability rule and the trigger-3 write were cut twice more before the bump; phase-3 stays under its max (8692/8950). Phase 4 is the first per-phase max raised since v10.9.0: the doc gained three steps that cannot be inferred (a per-round triage file, a prefix block that changes what reviewers report, and a halt condition), and its structural-compression debt (the inline triage JSON shape, named at v15.10.0) still stands and is the next candidate. 15150 and 60250 were the smallest steps that clear it, leaving 25 and 45 tokens. v16.23.0: phase-0-init max 12400 -> 12500 and total 60250 -> 60500 for the widget-registration call and the accounting gate. Phase 0 gained the `tiles` call and the exit-3 rule, Phase 7 gained the run report; together they are contract an agent cannot infer - which call registers this host's widget, and that a completion is refused without recorded spend. Compression came first and twice, taking the new prose from 472 tokens to 255: the per-host call list moved into tracker-contract.md "The card is not the widget" and the record-then-rerun recovery into "Accounting is a gate", both outside this budget, leaving the phase docs with the call and the one fact that cannot be looked up. 100 and 250 were the smallest steps that clear it, leaving 74 and 40 tokens. phase-3 and phase-4 remain amber. v16.24.0: total 60500 -> 60750 for visual evidence. Four phase docs gained one instruction each that cannot be inferred: Phase 0 keeps the issue's own images as the pre-fix evidence, Phase 3 captures the fixed state (there and not Phase 5, because every autopilot and --local entry drops Phase 5), Phase 5 hosts the flow recording when it runs, and Phase 6 blocks on a required artefact that is neither attached nor explained. Compression came first and twice, 42 tokens back, and the contract itself never entered this budget: the trigger matrix, the three video tiers, the size-degradation ladder and both render shapes live in multi-agent-refs/features/visual-evidence.md. phase-0-init cleared its own max without a bump. 250 was the smallest step that clears the aggregate. phase-3 and phase-4 remain amber. v17.0.0: phase-0-init max 13000 -> 13100 and total 62400 -> 62500 for the evidence-verdict writer. Phase 0 Step 7.7 probed with `--platform "$PLATFORM"`, a variable no phase document ever assigned, and it was gated on `visualEvidence.required`, which no phase document ever wrote - five readers, zero writers - so the step, Phase 3's capture and Phase 6's blocker were all unreachable and the pipeline reported nothing wrong. The step now writes the verdict and derives the platform from the stack, skipping the probe with a recorded reason when there is no device platform rather than passing the empty string the probe refuses with exit 2. Compression came first and three times, 30 tokens back from the Step 7.7 index rule that restated Step 7.5 verbatim and 55 from the new block itself; the reasoning never entered this budget, because who writes the verdict and how the platform is derived live in features/visual-evidence.md sections 1a and 1b. phase-3-dev max 9250 -> 9300 in the same change: it reads the platform back from state and re-decides the provisional verdict before capturing, which is the half of the fix that makes Phase 3 honest rather than merely reachable. Compressed three times first, 29 tokens back, by pointing its Phase-5 rationale and its tier mapping at visual-evidence.md sections 3 and 4.3 where both already live. 100, 50 and 150 were the smallest steps that clear it, leaving 24, 15 and 26 tokens. phase-3 and phase-4 remain amber. v17.1.0: total 62550 -> 62600 for the Figma/toolkit MCP distinction and web as an evidence platform. Three phase docs said "MCP forbidden" without qualifying it, while the gate that enforces it (smoke-no-mcp-in-dev-phases.sh) has always matched `figma` and nothing else - so the prose banned the screenshot, xcodebuild and UI-test tools that Phase 3 Steps 3.4 and 3.55 actually call, which is one way a run reaches Phase 7 with no evidence. Phase 0 gained one `web` arm in the platform derivation, now that run-ui-tests.sh has a web arm to derive it for. Compression came first and three times, taking the new prose from 175 tokens to 30: the reasoning moved to rules.md, whose own seven-row Figma phase matrix was deleted in the same pass because it duplicated rules/figma-pipeline.md "Phase access matrix" two paragraphs below this file's own instruction not to duplicate that rule file - 74 tokens back there, which is why phase-3-dev cleared its max without a bump (9298/9300). 50 was the smallest step that clears the aggregate, leaving 37 tokens. phase-3 and phase-4 remain amber. v17.1.0 (2): total 62600 -> 62700 for the plan reaching the task widget. Phase 2 computed tasks[], their order and their dependsOn[] edges, stored them, and used them to drive Phase 3's ready-task picker - and none of it was visible on the surface the user actually watches; the card had drawn sub-phases for releases, the widget never had. Phase 2 gains one call. Compression came first and twice: the tasks[]-to-todos[] jq blob left the phase doc for plan-todos.sh `set`, which now accepts a planning-output document directly (-28, and it removes a mapping two files defined, of which this was the untested copy), and the new step's own prose was cut from 116 tokens to 61. The parsing of the plan itself never entered this budget - it lives in phase-tracker.sh `plan`, which owns the sub-phase structure it writes. 100 was the smallest step that clears it, leaving 75 tokens. phase-3 and phase-4 remain amber. v17.5.0: phase-0-init max 13100 -> 13150 for the deferred widget registration, the dev-context step and two more exit-gate assertions. Compression came first and seven times, 492 tokens reclaimed before the number moved: the TLDR restated the step headings under it, the exit-gate list restated the script's own header comment, and the clarifier cost note, the branch-collision probe rationale, the fetch-fail host explanation, the multi-repo write-state race and the baseline unknown-vs-green warning all restate something already authoritative in agents/task-clarifier.md, phases/operations.md or phase0-exit-gate.mjs. The three additions are contract an agent cannot infer: that only Phase 0 is registered at Step -1 and the rest at 7.5 with `tiles --new`, that the dev-context picker runs between project and branch and writes siblings[] even when empty, and that a one-option AskUserQuestion is refused by the host along with every question batched with it. The reasoning for the last one never entered this budget - it lives in picker-contract.md "Two options or it is not a question". 50 was the smallest step that clears it, leaving 28 tokens; the aggregate needed no bump and sits at 62698/62700. phase-3 and phase-4 remain amber. v17.5.0 (2): phase-0-init max 13150 -> 13300 and total 62700 -> 62900 for the workspace question. Phase 0 gained Step 5b: where the branch lives was decided by a flag and by a command name, never by a question, so a user who saw a run improvise "Worktree / Lokal" once took it for a shipped picker and concluded /multi-agent:local was redundant - while it was the only way to reach local mode at all. Compression came first and three times, 156 tokens reclaimed before the number moved and the step itself cut from 426 tokens to 181: the ask-choice index rule in Step 7.5 restated modes.md "Pipeline depth" in full, and the question wording, the two options and what local costs now live in modes.md "Local Mode", outside this budget. What the phase doc carries is what an agent cannot infer - that it runs after Step 4 and before Step 6b and 7.5, who is exempt, and that autopilot resolves it to a worktree rather than skipping it, because an unattended commit in the user's own checkout is what worktrees exist to prevent. The state field is the other half: localMode alone cannot say whether anyone decided, since false is both a chosen worktree and one nothing asked about, and phase0-exit-gate.mjs now refuses to close Phase 0 without workspaceSource. 150 and 200 were the smallest steps that clear it, leaving 32 and 56 tokens. phase-3 and phase-4 remain amber. v17.5.0 (3): phase-0-init max 13300 -> 13400 and total 62900 -> 63000 for base-branch evidence. Step 3 asked one question with a list it could not vouch for: `git fetch origin` ran, its exit code was discarded, and `git branch -r` printed the remote-tracking cache either way, so a restricted network produced a weeks-old local list presented as the remote's answer. And the answer was usually derivable - an issue carrying a target version, or linking a separate issue that represents the release, already names the branch on a repo whose release branches encode the version - and nothing derived it. Rule 5 now captures the exit code and rule 5b hands the list plus the issue's version fields and links to base-branch-candidates.mjs. Compression came first and six times, 1,118 characters reclaimed across this step before the number moved: the one-row picker rule and the not-skippable rule both restated picker-contract.md, the autopilot resolution order and the whole rationale for admitting version-carrying branches moved into features/base-branch-evidence.md, and the fetch-fail persistence block was four sentences for one instruction. The feature's own reasoning never entered this budget at all - the evidence sources, the learned convention, the picker contract, the autopilot fencing and the state shape are ~1,900 tokens living in that ref. What the phase doc carries is the call and the two facts an agent cannot infer: that the fetch exit code decides whether the list may be called remote, and that the branch filter needs a version alternative because a word list that discards is the prefix table this replaces. 100 was the smallest step that clears both, leaving 51 and 61 tokens. phase-3 and phase-4 remain amber. v17.6.0: phase-3-dev max 9300 -> 9450 and total 63000 -> 63150 for package-manager resolution. The node-shaped arms typed `npm` into two command lines, so a repo on pnpm, yarn or bun failed in Phase 3 with a worktree and a branch already created - or, worse, npm resolved against a lock file it does not own and the run continued on a tree the repo's own tooling would never have produced. Both arms now call package-manager.mjs, which prints the line to run. A resolver invocation is inherently longer than the four characters it replaces, and phase-3-dev had 2 tokens of headroom (9298/9300), so no addition of any size could have fit without a bump. Compression came first and twice, taking the new prose from 430 characters to 160 and trimming the build arm's tail by 140: the resolution order, the two-lock-file rule, the per-manager command shapes, the exit-3 contract and the eval-safety argument all live in features/package-manager.md, outside this budget. What the phase doc carries is the call and the one fact an agent cannot infer - that exit 3 means the repo declares no such script, so the step is skipped and said, never substituted. phase-4-review was compressed rather than raised in the same pass: its Gate 3 comment named npm as though it were the only node test command, and the corrected line is 64 characters SHORTER than the one it replaces, so the doc nets -64 while becoming true (15150/15150). 150 was the smallest step that clears phase-3-dev, leaving 51 tokens; 150 clears the aggregate, leaving 133. phase-3 and phase-4 remain amber.
@@ -1,5 +1,5 @@
1
1
  {
2
- "_readme": "Recommended Claude Code hooks for multi-agent-pipeline. Merge the `hooks` object into your ~/.claude/settings.json to make these deterministic, OS-enforced PreToolUse gates real (exit 2 blocks the tool call) rather than prompt-level hopes. Three PreToolUse gates ship here: (1) a staged-diff secret scan on git commit (pre-commit-check.sh); (2) an agent-guard on git commit + git push (agent-guard.sh) that blocks AI/assistant attribution in commit messages and force-push to a protected branch (main/master/develop); (3) a read-size gate on Read and Bash (check-read-size.sh), which inspects Read plus the shell commands that read a file whole (cat/head/tail/sed) and returns immediately for everything else, which routes an oversized read to a cheap worker instead of the caller's own rung. The first two inspect what a run WRITES; the third inspects what it pays to READ, and it is inert until `prefs.global.bulkRead.mode` is set to observe or enforce - so merging this block changes nothing until you opt in. All three are self-contained, fail-open on internal error, never execute the inspected command, and need no run-specific arguments, which is why they are naturally PreToolUse hooks. The other deterministic gates (evidence, consensus, intent, learnings) take run-specific arguments and are phase-enforced by the pipeline instead. Two capture hooks ship alongside them, and they are the reason a killed run no longer loses what it learned: (4) SessionEnd runs capture-flush.sh --if-stale, which writes a run's triage findings and durable learnings into the per-repo stores when the run never reached Phase 7 - previously every persistent write lived in Phase 7, the phase a run is LEAST likely to reach; the same hook runs note-session.sh, which records the mechanical shape of a NON-pipeline session (tools used, commands that failed, calls the user refused) so work done outside a run stops vanishing. (5) SessionStart runs capture-resume.sh, which prints at most two lines: an unfinished run and how to resume it, and a stale pipeline-observation queue. Neither capture hook calls a model, neither reads a payload - note-session.sh keeps a command's first word and an exit code, never an argument or any output - and both exit 0 on every path, because a hook that fails a session over bookkeeping is worse than the bookkeeping it protects. multi-agent:setup offers to merge this block.",
2
+ "_readme": "Recommended Claude Code hooks for multi-agent-pipeline. Merge the `hooks` object into your ~/.claude/settings.json to make these deterministic, OS-enforced PreToolUse gates real (exit 2 blocks the tool call) rather than prompt-level hopes. Three PreToolUse gates ship here: (1) a staged-diff secret scan on git commit (pre-commit-check.sh); (2) an agent-guard on git commit + git push (agent-guard.sh) that blocks AI/assistant attribution in commit messages and force-push to a protected branch (main/master/develop); (3) a read-size gate on Read and Bash (check-read-size.sh), which inspects Read plus the shell commands that read a file whole (cat/head/tail/sed) and returns immediately for everything else, which routes an oversized read to a cheap worker instead of the caller's own rung. The first two inspect what a run WRITES; the third inspects what it pays to READ, and it is inert until `prefs.global.bulkRead.mode` is set to observe or enforce - so merging this block changes nothing until you opt in. All three are self-contained, fail-open on internal error, never execute the inspected command, and need no run-specific arguments, which is why they are naturally PreToolUse hooks. The other deterministic gates (evidence, consensus, intent, learnings) take run-specific arguments and are phase-enforced by the pipeline instead. Three capture hooks ship alongside them, and they are the reason a killed run no longer loses what it learned: (4) SessionEnd runs capture-flush.sh --if-stale, which writes a run's triage findings and durable learnings into the per-repo stores when the run never reached Phase 7 - previously every persistent write lived in Phase 7, the phase a run is LEAST likely to reach; the same hook runs note-session.sh, which records the mechanical shape of a NON-pipeline session (tools used, commands that failed, calls the user refused) so work done outside a run stops vanishing. (5) PreCompact runs capture-flush.sh - WITHOUT --if-stale, because the staleness test exists so a session exit does not re-flush a run that already finished, while a compaction is the moment un-flushed findings are actually at risk and both writes are idempotent, so the finished case costs two no-op writes. A compaction summarizes the conversation mid-phase: a long Phase 3 or Phase 4 can lose what it established before SessionEnd ever fires, which is the same failure that put the SessionEnd hook here one level down. (6) SessionStart runs capture-resume.sh, which prints at most two lines: an unfinished run and how to resume it, and a stale pipeline-observation queue. No capture hook calls a model, none reads a payload - note-session.sh keeps a command's first word and an exit code, never an argument or any output - and all exit 0 on every path, because a hook that fails a session over bookkeeping is worse than the bookkeeping it protects. multi-agent:setup offers to merge this block.",
3
3
  "hooks": {
4
4
  "PreToolUse": [
5
5
  {
@@ -42,6 +42,18 @@
42
42
  ]
43
43
  }
44
44
  ],
45
+ "PreCompact": [
46
+ {
47
+ "hooks": [
48
+ {
49
+ "type": "command",
50
+ "command": "bash $HOME/.claude/scripts/capture-flush.sh --quiet",
51
+ "timeout": 20,
52
+ "statusMessage": "Persisting what this run learned before compaction..."
53
+ }
54
+ ]
55
+ }
56
+ ],
45
57
  "SessionEnd": [
46
58
  {
47
59
  "hooks": [
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@mmerterden/multi-agent-pipeline",
3
- "version": "17.5.0",
3
+ "version": "17.6.0",
4
4
  "description": "8-phase AI development pipeline with full orchestration on Claude Code, Copilot CLI and Codex CLI. Analysis, planning, TDD, CLI-aware parallel review with consensus surfacing + Fable triage, default-FAIL evidence gates, secret + intent guards, per-phase cost ledger, persistent learnings memory, wiki generation, commit automation. Token-preserving uninstall.",
5
5
  "type": "module",
6
6
  "main": "index.js",
@@ -48,7 +48,7 @@ Classification schema lives in `$HOME/.claude/multi-agent-refs/_input-parser.md`
48
48
  | 7 | `issue` | full picker | account → repo (multi) → issue → maturity → dev-context |
49
49
  | 8 | Free-text | `freetext` | account → repo (single) → dev-context (maturity skip) |
50
50
 
51
- **Rule**: Whatever the type, **account is always asked** (autopilot picks a default). After issue fetch, **maturity check is mandatory** - blockers halt the pipeline. Picker `header` renders in English (the 12-char chip); the `question`, each option's `label` and each option's `description` render in `outputLanguage`, per the canonical matrix in `multi-agent-refs/rules.md`.
51
+ **Rule**: Whatever the type, **account is always asked** (autopilot picks a default). After issue fetch, **maturity check is mandatory** - a blocker asks before it halts, and an autopilot run can be told to ask on the item itself (`$HOME/.claude/multi-agent-refs/features/maturity-followup.md`). Picker `header` renders in English (the 12-char chip); the `question`, each option's `label` and each option's `description` render in `outputLanguage`, per the canonical matrix in `multi-agent-refs/rules.md`.
52
52
 
53
53
  Lib scripts (`~/.claude/lib/`):
54
54
  - `account-resolver.sh` - keychain account inventory
@@ -44,7 +44,13 @@ One message, sent to the maintainer's admin panel. This exists because the alter
44
44
 
45
45
  ## Auth and reachability
46
46
 
47
- Reuses the usage ingest token (`prefs.global.keychainMapping.usage_ingest`), so nothing extra has to be onboarded - `/multi-agent:update` registers one on first run. With no token, the command says so and names the command that fixes it rather than failing quietly.
47
+ Reuses the usage ingest token (`prefs.global.keychainMapping.usage_ingest`), so nothing extra has to be onboarded. When none resolves, register one first and then send:
48
+
49
+ ```bash
50
+ node "$HOME/.claude/scripts/usage-register.mjs" --feedback --quiet
51
+ ```
52
+
53
+ `--feedback` is what makes this work for a machine that opted out of telemetry: it mints the token and leaves `usageLog.enabled` untouched. If registration is unavailable (offline, endpoint down), the command says so and names the command that fixes it rather than failing quietly.
48
54
 
49
55
  `usageLog.optOut` does **not** silence this. Telemetry is passive collection and opting out of it is a real choice; feedback is a deliberate act by the person typing the command, and dropping a message somebody chose to send would be worse than not offering the command at all.
50
56
 
@@ -25,7 +25,7 @@ No worktree, no branch, no commit, no pipeline chaining.
25
25
  | `refresh` | same as `build` | A full rebuild takes seconds, so there is no separate incremental path |
26
26
  | `ask "<question>"` | `graph-query.mjs "<question>" --budget N` | Token-budgeted traversal; default budget 2000 |
27
27
  | `affected "<symbol>"` | `graph-affected.mjs "<symbol>" --depth N` | Reverse traversal: the blast radius of a change |
28
- | `report` | `graph-report.mjs` | Writes `GRAPH_REPORT.md` beside the graph |
28
+ | `report` | `graph-report.mjs` | Writes `GRAPH_REPORT.md` beside the graph: hubs, modules, external dependencies, unconnected files, and symbols no other file references (candidates only - the extractor is regex, not a parser, so nothing gates on that list) |
29
29
  | `status` | `graph-report.mjs --status` | One line: stack, scale, build time and whether `baseCommit` still matches HEAD. Never read the graph file yourself - it is 22MB on a large repo |
30
30
 
31
31
  With no argument, run `status`, then offer `build` when no graph exists and
@@ -56,10 +56,22 @@ After selection, inspect `maturity` from `~/.claude/lib/issue-fetcher.sh` (same
56
56
 
57
57
  | Outcome | Behavior |
58
58
  |---|---|
59
- | `blockers` non-empty (e.g. `description_empty`, `status_closed`) | **Halt** (autopilot too): show summary, advise "fix the issue first" |
59
+ | `blockers` non-empty (e.g. `description_empty`, `status_closed`) | **Ask, then halt** - see "Blockers" below. Default is still the halt |
60
60
  | only `warnings` | AskUserQuestion: show summary + "Continue?". Autopilot auto-continues; warnings logged to `agent-log.md` |
61
61
  | both empty | Continue silently |
62
62
 
63
+ **Blockers: ask, do not just stop.** `$HOME/.claude/multi-agent-refs/features/maturity-followup.md`
64
+ owns this. In short: an interactive run asks at this step (open the item and fix it /
65
+ continue without it, recording which gap was waved through / abort) instead of ending
66
+ with a summary; an autopilot run with `prefs.global.maturityFollowup.autopilotCommentsOnIssue`
67
+ on posts ONE comment on the item asking for what is missing, then halts on the circuit
68
+ breaker with `state.waitingFor = "maturity"` so `resume` re-enters THIS step with the item
69
+ re-fetched. Both defaults keep today's behaviour: the comment is off, the halt is the halt.
70
+ An edit is a reason to re-check, never proof the gap closed - the check re-runs on the new
71
+ content, and only a DIFFERENT gap set ever earns a second comment. Read the item's existing
72
+ comments FIRST and derive the prior ask from the newest one of ours (`priorFromComments`):
73
+ a scan is a new run with a fresh state file, so state alone would make every scan a first ask.
74
+
63
75
  `PROMPT_LANG=en` is passed to `issue-fetcher.sh` (`promptLanguage` is locked to `"en"`).
64
76
 
65
77
  ### [4/4] Dev context (`_dev-context`)