@mmerterden/multi-agent-pipeline 17.4.0 → 17.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +136 -0
- package/README.md +23 -5
- package/README.tr.md +23 -5
- package/docs/adr/0013-lsp-code-intelligence.md +102 -0
- package/docs/adr/README.md +1 -0
- package/docs/token-budget-history.md +1 -1
- package/install/templates/copilot-instructions.md +9 -3
- package/package.json +1 -1
- package/pipeline/commands/multi-agent/analysis/SKILL.md +3 -3
- package/pipeline/commands/multi-agent/autopilot/SKILL.md +3 -3
- package/pipeline/commands/multi-agent/autopilot-off/SKILL.md +5 -3
- package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/local/SKILL.md +17 -6
- package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +3 -3
- package/pipeline/lib/multi-repo-pipeline.sh +26 -0
- package/pipeline/multi-agent-refs/features/base-branch-evidence.md +222 -0
- package/pipeline/multi-agent-refs/features/code-intelligence.md +80 -0
- package/pipeline/multi-agent-refs/phases/modes.md +23 -3
- package/pipeline/multi-agent-refs/phases/phase-0-init.md +96 -71
- package/pipeline/multi-agent-refs/phases/phase-7-report.md +1 -1
- package/pipeline/multi-agent-refs/phases.md +7 -2
- package/pipeline/multi-agent-refs/picker-contract.md +37 -5
- package/pipeline/multi-agent-refs/tracker-contract.md +25 -14
- package/pipeline/schemas/agent-state.schema.json +88 -4
- package/pipeline/schemas/prefs.schema.json +22 -0
- package/pipeline/schemas/token-budget.json +2 -2
- package/pipeline/scripts/autopilot-runner.mjs +292 -45
- package/pipeline/scripts/base-branch-candidates.mjs +599 -0
- package/pipeline/scripts/gc-abandoned.sh +5 -3
- package/pipeline/scripts/gen-mode-dispatch.mjs +39 -16
- package/pipeline/scripts/phase-tracker.sh +39 -2
- package/pipeline/scripts/phase0-exit-gate.mjs +128 -0
- package/pipeline/scripts/verify-citations.mjs +84 -2
- package/pipeline/skills/.skill-manifest.json +2 -2
- package/pipeline/skills/shared/core/multi-agent/SKILL.md +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -14,6 +14,142 @@ Internal file-layout changes that don't affect the slash-command surface are sti
|
|
|
14
14
|
|
|
15
15
|
---
|
|
16
16
|
|
|
17
|
+
## [17.5.0] - 2026-09-15
|
|
18
|
+
|
|
19
|
+
Four defects from one live run of `/multi-agent`, all in Phase 0, all visible in
|
|
20
|
+
the same screenshot. The transcript was read before anything was changed, so each
|
|
21
|
+
of these names the exact call that produced it.
|
|
22
|
+
|
|
23
|
+
### Fixed
|
|
24
|
+
|
|
25
|
+
- **A one-option `AskUserQuestion` is refused by the host, and it takes every
|
|
26
|
+
question batched with it down unasked.** The run batched four questions -
|
|
27
|
+
maturity, base branch, depth, workspace. The remote had one PR-targetable
|
|
28
|
+
branch, so the base-branch question carried one option, and Claude Code's
|
|
29
|
+
schema requires two: `InputValidationError ... questions[1].options too_small`,
|
|
30
|
+
rendered in the terminal as `Invalid tool parameters`. All four were discarded.
|
|
31
|
+
The run then did what the host's own error text says to do - stated the single
|
|
32
|
+
path and continued - and printed "base branch: only candidate, continuing with
|
|
33
|
+
it". That line is why the base branch looked like it was never asked about: it
|
|
34
|
+
was asked, and the schema refused it.
|
|
35
|
+
`picker-contract.md` gains "Two options or it is not a question", which is the
|
|
36
|
+
missing half of "A single candidate is still a question" it has carried since
|
|
37
|
+
the beginning: a one-row filter is asked WITH a real escape as its second
|
|
38
|
+
option (unfiltered branch list, drop the filter, abort), never as a one-row
|
|
39
|
+
picker. A filler option is called out as worse than not asking.
|
|
40
|
+
- **The phase widget committed to eight phases before the user picked a depth.**
|
|
41
|
+
The tracker booted at Step -1 and registered all eight, so "8 tasks (0 done, 1
|
|
42
|
+
in progress, 7 open)" with Analysis and Planning tiles sat on screen beside the
|
|
43
|
+
question asking whether to run Analysis and Planning. Registration now splits:
|
|
44
|
+
Phase 0 at Step -1, so the run is never silent while Phase 0 works, and the
|
|
45
|
+
rest at Step 7.5 the moment depth names them. A Short run draws no Analysis
|
|
46
|
+
tile at all, rather than drawing one and flipping it to `[SKIPPED]`.
|
|
47
|
+
`phase-tracker.sh tiles --new` narrows the second batch to phases that carry no
|
|
48
|
+
`tasklist_id`, so the Phase 0 tile is never created twice. `tracker-contract.md`
|
|
49
|
+
"Late skip" is replaced by "Deferred registration".
|
|
50
|
+
- **The dev-context picker was unreachable from Phase 0.** `commands/multi-agent/SKILL.md`
|
|
51
|
+
has routed every input type through `_dev-context.md` since it existed, and
|
|
52
|
+
`phase-0-init.md` - the document a run actually follows step by step - never
|
|
53
|
+
mentioned it. The reported run left no `siblings` in `agent-state.json`, so
|
|
54
|
+
Phase 4's platform-parity cross-check had nothing to read and could not tell an
|
|
55
|
+
empty answer from no answer. Phase 0 gains **Step 2b**, between project
|
|
56
|
+
selection and the base branch, which is the order `picker-contract.md` already
|
|
57
|
+
required (a base branch is a property of a repo set).
|
|
58
|
+
|
|
59
|
+
### Added
|
|
60
|
+
|
|
61
|
+
- **Where the branch lives is a question now, not a flag (Phase 0 Step 5b).** It
|
|
62
|
+
was decided by `--local` and by a command name, never by asking. A run
|
|
63
|
+
improvised "Worktree / Lokal" once, a user took it for a shipped picker, and
|
|
64
|
+
concluded `/multi-agent:local` was redundant - while it was the only way to
|
|
65
|
+
reach local mode at all. No spec contained that question. It runs after Step 4
|
|
66
|
+
names the branch and before Step 6b creates the workspace and Step 7.5
|
|
67
|
+
registers the tiles for the resulting set. Autopilot **resolves** it to a
|
|
68
|
+
worktree rather than skipping it: an unattended run commits and pushes from
|
|
69
|
+
wherever it stands, and doing that in the user's own checkout is what
|
|
70
|
+
worktrees exist to prevent. `:local-autopilot` is the explicit opt-out.
|
|
71
|
+
`workspaceSource` is the record - `localMode: false` alone is both "the user
|
|
72
|
+
chose a worktree" and "nothing asked".
|
|
73
|
+
|
|
74
|
+
- **Base-branch evidence (Phase 0 Step 3), behind
|
|
75
|
+
`prefs.global.baseBranchEvidence.enabled`.** Step 3 asked one question with a
|
|
76
|
+
list it could not vouch for: `git fetch origin` ran, its exit code was
|
|
77
|
+
discarded, and `git branch -r` printed the remote-tracking cache either way -
|
|
78
|
+
so on a restricted network a weeks-old local list was presented as the
|
|
79
|
+
remote's answer with nothing saying so. And the answer was usually derivable:
|
|
80
|
+
an issue carrying a target version, or linking a separate issue that
|
|
81
|
+
represents the release, already names the branch on a repo whose release
|
|
82
|
+
branches encode the version.
|
|
83
|
+
|
|
84
|
+
`base-branch-candidates.mjs` collects candidates **with the evidence behind
|
|
85
|
+
each one** and ranks them; a human still chooses, and the evidence is the
|
|
86
|
+
picker row's description. No field id is hardcoded - any Jira field whose
|
|
87
|
+
schema resolves to `version` is read, whatever the board calls it - and no
|
|
88
|
+
branch prefix is tabled: the release-branch template is inferred from the refs
|
|
89
|
+
that exist, so one repo yields `<prefix>/develop_<version>` and another
|
|
90
|
+
`release-<version>` out of the same code. A predicted branch is a **note**,
|
|
91
|
+
never an option: "that version has no branch on the remote yet" is a real
|
|
92
|
+
answer, and an option the user picks has to be checkoutable.
|
|
93
|
+
|
|
94
|
+
The rule-5 branch filter gained a version alternative in the same change. It
|
|
95
|
+
had narrowed the list with `grep -E '(develop|release|main|master)'`, which is
|
|
96
|
+
the prefix table this feature exists not to have - and worse, because it
|
|
97
|
+
discards: a repo whose release branches read `stabilise-2.7` matched none of
|
|
98
|
+
the four words, so none of its branches reached the picker and the convention
|
|
99
|
+
was unlearnable. A word list that ranks is fine; one that filters is not.
|
|
100
|
+
|
|
101
|
+
Autopilot resolves `remembered` → `derived` → `default` and records which
|
|
102
|
+
fired. With `autopilotAsksOnIssue` (**off by default**) it may post **one**
|
|
103
|
+
comment on the Jira or GitHub issue asking which branch to develop from, then
|
|
104
|
+
halt on circuit-breaker trigger 6 and wait for `resume`. A question, never a
|
|
105
|
+
state change: no transition, no resolution, no assignee, no close, `Ref:` and
|
|
106
|
+
never `Closes:`, copy in `outputLanguage`. An autopilot that posts a question
|
|
107
|
+
and then answers it itself has not asked anything.
|
|
108
|
+
|
|
109
|
+
- `phase0-exit-gate.mjs` asserts two more fields, each the durable evidence a
|
|
110
|
+
step actually ran:
|
|
111
|
+
- `baseBranchSource` must be recorded, and an interactive run must record
|
|
112
|
+
`asked` or `input`. `remembered` and `default` are autopilot resolutions; an
|
|
113
|
+
interactive run that reaches for them skipped its picker. `baseBranch` alone
|
|
114
|
+
cannot make this assertion - a branch announced in prose and a branch the
|
|
115
|
+
user chose leave the same value behind, which is exactly what happened.
|
|
116
|
+
- `siblings` must be an array, `[]` included. The empty array is Step 2b's
|
|
117
|
+
record that it ran; absent means the picker never ran and a multi-repo task
|
|
118
|
+
silently became a single-repo one.
|
|
119
|
+
- `phase-tracker.sh tiles --new`: emit `TaskCreate` only for registered phases
|
|
120
|
+
that have no tile yet. Covered by `smoke-phase-tracker.sh` section 25,
|
|
121
|
+
including the negative control that the unnarrowed call still lists everything.
|
|
122
|
+
- Both READMEs document what Short actually is next to the modes table: which
|
|
123
|
+
phases it drops, what it does not drop (review, gates, tests), and when it is
|
|
124
|
+
the wrong answer. The table row naming it has existed since v16.0.0; nothing
|
|
125
|
+
explained it.
|
|
126
|
+
|
|
127
|
+
### Changed
|
|
128
|
+
|
|
129
|
+
- `gen-mode-dispatch.mjs` learns the deferred shape for the two modes that ask
|
|
130
|
+
the depth question (`full`, `local`, `full-local`); the other modes still
|
|
131
|
+
register their whole set at Step -1. `local/SKILL.md` regenerated.
|
|
132
|
+
- `smoke-pipeline-surface.sh` replaces the "Late skip" check with three: the
|
|
133
|
+
Step -1 block registers no phase set, it does register Phase 0, and Step 7.5
|
|
134
|
+
calls `tiles --new`.
|
|
135
|
+
- `smoke-mode-dispatch-drift.sh` accepts both phase-declaration spellings
|
|
136
|
+
(`"<id>:<name>"` in a loop, or `add <id> "<name>"` on its own line).
|
|
137
|
+
- Token budget: `phase-0-init` max 13100 -> 13150 after 492 tokens of compression
|
|
138
|
+
first, per `docs/token-budget-history.md`. The aggregate needed no bump.
|
|
139
|
+
|
|
140
|
+
### Investigated, not changed
|
|
141
|
+
|
|
142
|
+
- **`/multi-agent:local` is not redundant.** The picker does not ask
|
|
143
|
+
local-vs-worktree: the "Workspace: worktree or --local" question in the
|
|
144
|
+
reported run was improvised by the model, and no spec, ref or command file
|
|
145
|
+
contains it. `:local` is a real mode with its own phase set (no Phase 5, which
|
|
146
|
+
needs a worktree to check the change out from) and its own tracker
|
|
147
|
+
registration, and `:resume-local` depends on the same `localMode` state. Making
|
|
148
|
+
the improvised question official is a reasonable future change; removing the
|
|
149
|
+
command on the strength of a question nothing specifies is not.
|
|
150
|
+
|
|
151
|
+
---
|
|
152
|
+
|
|
17
153
|
## [17.4.0] - 2026-09-15
|
|
18
154
|
|
|
19
155
|
Phase 4 now answers two questions it could not answer before: does the evidence a
|
package/README.md
CHANGED
|
@@ -48,7 +48,7 @@ Run a task - the input type is auto-detected:
|
|
|
48
48
|
|
|
49
49
|
Every input runs the same short intake - **account → (repo) → maturity check → dev-context** - then enters Phase 0. A Jira id or GitHub URL is fetched and maturity-checked _before_ any code is written; free-text skips the fetch and goes straight to planning. Multi-repo tasks add extra repos at the dev-context step.
|
|
50
50
|
|
|
51
|
-
Add `autopilot` to skip confirmations
|
|
51
|
+
Add `autopilot` to skip confirmations (e.g. `/multi-agent:autopilot "PROJ-1234"`). Neither the workspace nor the depth is a flag any more - they are two questions the run asks at Phase 0: where to run (worktree or your current checkout), then how deep (Full or Short). `--local` and `:local` answer the first up front.
|
|
52
52
|
|
|
53
53
|
Update later with `/multi-agent:update`. Uninstall (tokens preserved) with `npx @mmerterden/multi-agent-pipeline uninstall`.
|
|
54
54
|
|
|
@@ -56,7 +56,10 @@ Update later with `/multi-agent:update`. Uninstall (tokens preserved) with `npx
|
|
|
56
56
|
|
|
57
57
|
## How it works
|
|
58
58
|
|
|
59
|
-
One command runs 8 phases, with a gate between the risky ones
|
|
59
|
+
One command runs up to 8 phases, with a gate between the risky ones. Phase 0
|
|
60
|
+
asks two questions that decide the shape of the rest - how deep the run goes
|
|
61
|
+
(Full or Short) and where the branch lives (a worktree or your current
|
|
62
|
+
checkout):
|
|
60
63
|
|
|
61
64
|
- **0 · Init** - parse the input (Jira id / GitHub URL / free text), pick account + repo(s), fetch the issue, run a maturity check.
|
|
62
65
|
- **1 · Analysis** - detect the stack, scan the codebase, map impact (Sonnet).
|
|
@@ -69,7 +72,11 @@ One command runs 8 phases, with a gate between the risky ones:
|
|
|
69
72
|
|
|
70
73
|
`/multi-agent:analysis` runs its own shorter chain and, since v16.12.0, reviews what it wrote before publishing it: the draft goes through the same three-reviewer set and triage as a code diff, a blocking finding returns it to synthesis with dispatch closed, and the gaps that survive are either searched, asked about, or recorded with an owner. It used to publish behind a structural validator alone.
|
|
71
74
|
|
|
72
|
-
|
|
75
|
+
### Workspace: worktree or local
|
|
76
|
+
|
|
77
|
+
Phase 0 Step 5b asks where the branch lives. **Worktree** (`.worktrees/{id}/`) leaves your current checkout untouched; **Local** works in the project root on a new branch, which drops Phase 5 - the user-test gate checks the change out of a worktree and there is none - and needs the project root clean. `/multi-agent:local` and `--local` answer it up front. Every autopilot entry resolves it to a worktree without asking: an unattended run commits and pushes from wherever it stands, and doing that in your own checkout is what worktrees exist to prevent. `:local-autopilot` is the explicit opt-out.
|
|
78
|
+
|
|
79
|
+
Under the hood: each task runs in its own **git worktree** (or the current branch when you choose local), commits use the **git identity routed from the repo's origin URL**, and **multi-repo** tasks get per-repo worktrees plus an integration build. Tokens stay in the OS keychain; nothing is committed or logged. `/multi-agent:review` can also review an existing GitHub/Bitbucket PR - per-finding inline comments anchored to `file:line` + an explicit Approve / Needs-Work state.
|
|
73
80
|
|
|
74
81
|
The discipline behind all of this - bounded loops, evidence gates, token-budgeted phase docs, immutable tests, fresh-context handoffs - is catalogued in [docs/engineering.md](./docs/engineering.md). The full feature list lives in [docs/features.md](./docs/features.md). How this repo, the `multi-agent-plugins` marketplace, and the `multi-agent-toolkit-mcp` server compose at install time and at run time is diagrammed in [docs/ecosystem.md](./docs/ecosystem.md).
|
|
75
82
|
|
|
@@ -86,6 +93,17 @@ The discipline behind all of this - bounded loops, evidence gates, token-budgete
|
|
|
86
93
|
| Audit | `/multi-agent:testflight-validation` | Pre-submission gates for a TestFlight build: static archive audit → Apple's `altool --validate-app` → Review-Guidelines check. Validates only, never uploads |
|
|
87
94
|
|
|
88
95
|
Depth, autopilot and `--local` are the only knobs on the run itself; everything else is its own command. The full catalog is below.
|
|
96
|
+
### Pipeline depth: Full or Short
|
|
97
|
+
|
|
98
|
+
Depth is the one question the run asks about its own shape. `/multi-agent` and `/multi-agent:local` ask it at Phase 0 Step 7.5 - after the issue is fetched and the task type is known, because that is what the recommendation is drawn from.
|
|
99
|
+
|
|
100
|
+
- **Full** runs everything: Analysis reads the codebase and maps impact, Planning writes a task breakdown and stops for your approval, and Dev works from that plan.
|
|
101
|
+
- **Short** starts at Dev: Init → Dev → Review → Test → Commit → Report. Analysis and Planning do not run, so there is no plan gate and no analysis document; Dev derives its own task list from the issue, and runs on **Opus** rather than Sonnet because it has no plan to follow. Review, the deterministic gates and the test suite are untouched - Short skips the thinking, never the proof.
|
|
102
|
+
|
|
103
|
+
Short is right when you already know the fix and the file: a one-line guard, a copy change, a rename, a revert. It is wrong when the cause is still a hypothesis, when the task carries a Figma reference or an analysis document (Planning is what turns those into a breakdown), or when the change spans repos.
|
|
104
|
+
|
|
105
|
+
The widget follows the answer rather than predicting it: Phase 0 is the only tile drawn before you choose, and a Short run never draws an Analysis tile at all. Both autopilot entries skip the question and always run Full. `agent-state.json` records which one ran as `onlyDevelop`.
|
|
106
|
+
|
|
89
107
|
|
|
90
108
|
## Commands
|
|
91
109
|
|
|
@@ -265,11 +283,11 @@ The pipeline runs natively on **Claude Code**, **Copilot CLI** and **Codex CLI**
|
|
|
265
283
|
| ----------- | -------------------- | ------------------------------------------------------------------------------------------------------ |
|
|
266
284
|
| Claude Code | `--claude` (default) | slash commands + skills + agents + three `PreToolUse` hooks (secret scan, agent-guard, read-size gate) |
|
|
267
285
|
| Copilot CLI | `--copilot` | instructions + 56 sub-command skills + scripts |
|
|
268
|
-
| Codex CLI | `--codex` | one router skill + 56 specs as refs +
|
|
286
|
+
| Codex CLI | `--codex` | one router skill + 56 specs as refs + 9 agent TOML + `AGENTS.md` block + `codex mcp add` |
|
|
269
287
|
|
|
270
288
|
Filter skills by stack with `--platform=ios\|android\|all`.
|
|
271
289
|
|
|
272
|
-
**Why Codex gets one skill and not
|
|
290
|
+
**Why Codex gets one skill and not 56.** Codex assembles every discovered skill's name
|
|
273
291
|
and description into a single prompt block and drops entries when it overflows, with no
|
|
274
292
|
error. Measured on 0.145: installing one plugin that declares 142 skills surfaced only
|
|
275
293
|
75 of them and evicted an unrelated user skill. So on Codex the pipeline ships a single
|
package/README.tr.md
CHANGED
|
@@ -48,7 +48,7 @@ Bir görev çalıştır - girdi tipi otomatik algılanır:
|
|
|
48
48
|
|
|
49
49
|
Her girdi aynı kısa intake'ten geçer - **hesap → (repo) → maturity kontrolü → dev-context** - sonra Phase 0'a girer. Bir Jira id'si veya GitHub URL'i hiçbir kod yazılmadan _önce_ çekilir ve maturity-kontrol edilir; serbest-metin bu çekimi atlayıp doğrudan planlamaya geçer. Çoklu-repo görevleri dev-context adımında ekstra repo ekler.
|
|
50
50
|
|
|
51
|
-
Onayları atlamak için `autopilot
|
|
51
|
+
Onayları atlamak için `autopilot` ekle (örn. `/multi-agent:autopilot "PROJ-1234"`). Derinlik de çalışma alanı da artık bayrak değil, koşunun sorduğu iki soru: `/multi-agent` Faz 0'da önce nerede koşacağını (worktree/lokal), sonra ne kadar derin koşacağını (Tam/Kısa) sorar. `--local` ve `:local` ilkini baştan cevaplar.
|
|
52
52
|
|
|
53
53
|
Sonra `/multi-agent:update` ile güncelle. Kaldırmak için (tokenlar korunur) `npx @mmerterden/multi-agent-pipeline uninstall`.
|
|
54
54
|
|
|
@@ -56,7 +56,10 @@ Sonra `/multi-agent:update` ile güncelle. Kaldırmak için (tokenlar korunur) `
|
|
|
56
56
|
|
|
57
57
|
## Nasıl çalışır
|
|
58
58
|
|
|
59
|
-
Tek komut 8 fazı çalıştırır, riskli olanlar arasında bir kapı ile
|
|
59
|
+
Tek komut en fazla 8 fazı çalıştırır, riskli olanlar arasında bir kapı ile. Faz
|
|
60
|
+
0 geri kalanın şeklini belirleyen iki soru sorar: koşu ne kadar derin olacak
|
|
61
|
+
(Tam mı Kısa mı) ve branch nerede yaşayacak (worktree mi, mevcut checkout'un
|
|
62
|
+
mu):
|
|
60
63
|
|
|
61
64
|
- **0 · Init** - girdiyi ayrıştır (Jira id / GitHub URL / serbest metin), hesap + repo(lar) seç, issue'yu çek, maturity kontrolü yap.
|
|
62
65
|
- **1 · Analysis** - stack'i tespit et, codebase'i tara, etkiyi haritala (Sonnet).
|
|
@@ -69,7 +72,11 @@ Tek komut 8 fazı çalıştırır, riskli olanlar arasında bir kapı ile:
|
|
|
69
72
|
|
|
70
73
|
`/multi-agent:analysis` kendi kısa zincirini koşar ve v16.12.0'dan beri yazdığını yayınlamadan önce review ediyor: taslak, bir kod diff'iyle aynı üç-reviewer setinden ve triyajdan geçiyor, bloklayıcı bulgu dokümanı sentez fazına geri gönderip dispatch'i kapatıyor, hayatta kalan boşluklar ya aranıyor ya sana soruluyor ya da sahibiyle birlikte kayda giriyor. Önceden yalnızca yapısal bir validator'ın arkasından yayınlıyordu.
|
|
71
74
|
|
|
72
|
-
|
|
75
|
+
### Çalışma alanı: worktree mi lokal mi
|
|
76
|
+
|
|
77
|
+
Faz 0 Adım 5b branch'in nerede yaşayacağını sorar. **Worktree** (`.worktrees/{id}/`) mevcut checkout'una dokunmaz; **Lokal** proje kökünde yeni bir branch'te çalışır, bu da Faz 5'i düşürür - kullanıcı-test kapısı değişikliği bir worktree'den checkout eder, ortada worktree yoktur - ve proje kökünün temiz olmasını ister. `/multi-agent:local` ve `--local` bu soruyu baştan cevaplar. Her autopilot girişi sormadan worktree'ye karar verir: gözetimsiz bir koşu nerede duruyorsa oradan commit'leyip push eder, ve bunu senin kendi checkout'unda yapmak tam olarak worktree'nin engellemek için var olduğu şeydir. `:local-autopilot` bunun açık opt-out'u.
|
|
78
|
+
|
|
79
|
+
Perde arkasında: her görev kendi **git worktree**'sinde çalışır (ya da lokal seçtiğinde mevcut branch'te), commit'ler **repo'nun origin URL'inden yönlendirilen git kimliğini** kullanır, ve **çoklu-repo** görevleri repo başına worktree artı bir integration build alır. Tokenlar OS keychain'de kalır; hiçbir şey commit edilmez ya da loglanmaz. `/multi-agent:review` mevcut bir GitHub/Bitbucket PR'ını da review edebilir - `file:line`'a bağlı bulgu-başına inline yorumlar + açık bir Approve / Needs-Work durumu.
|
|
73
80
|
|
|
74
81
|
Bunun arkasındaki disiplin - sınırlı loop'lar, kanıt kapıları, token-bütçeli faz dokümanları, değişmez testler, taze-context handoff'lar - [docs/engineering.md](./docs/engineering.md)'de kataloglanmıştır. Tam özellik listesi [docs/features.md](./docs/features.md)'te. Bu repo, `multi-agent-plugins` marketplace'i ve `multi-agent-toolkit-mcp` sunucusunun install zamanında ve run zamanında nasıl bir araya geldiği [docs/ecosystem.md](./docs/ecosystem.md)'de diyagramlanmıştır.
|
|
75
82
|
|
|
@@ -86,6 +93,17 @@ Bunun arkasındaki disiplin - sınırlı loop'lar, kanıt kapıları, token-büt
|
|
|
86
93
|
| Audit | `/multi-agent:testflight-validation` | TestFlight build için pre-submission kapıları: statik archive denetimi → Apple'ın `altool --validate-app`'i → Review-Guidelines kontrolü. Yalnızca doğrular, asla yüklemez |
|
|
87
94
|
|
|
88
95
|
Koşunun kendisinde ayarlanabilen tek şey derinlik, autopilot ve `--local`; geri kalan her şey kendi komutu. Tam katalog aşağıda.
|
|
96
|
+
### Pipeline derinliği: Tam mı Kısa mı
|
|
97
|
+
|
|
98
|
+
Derinlik, koşunun kendi şekli hakkında sorduğu tek soru. `/multi-agent` ve `/multi-agent:local` bunu Faz 0 Adım 7.5'te sorar - issue çekildikten ve görev tipi belirlendikten sonra, çünkü öneri onlardan çıkıyor.
|
|
99
|
+
|
|
100
|
+
- **Tam** hepsini koşar: Analiz kod tabanını okuyup etki alanını çıkarır, Planlama görev kırılımını yazıp onayını bekler, Dev de o plandan çalışır.
|
|
101
|
+
- **Kısa** doğrudan Dev'den başlar: Init → Dev → Review → Test → Commit → Report. Analiz ve Planlama koşmaz; plan kapısı ve analiz dokümanı yoktur, görev listesini Dev issue'dan kendi çıkarır ve takip edecek bir planı olmadığı için Sonnet yerine **Opus** üzerinde koşar. Review, deterministik kapılar ve test paketi aynen kalır - Kısa düşünmeyi atlar, kanıtı değil.
|
|
102
|
+
|
|
103
|
+
Kısa'yı düzeltmeyi ve dosyayı zaten biliyorsan seç: tek satırlık bir guard, bir metin değişikliği, bir yeniden adlandırma, bir geri alma. Neden hâlâ bir hipotezse, görevde Figma referansı ya da analiz dokümanı varsa (bunları kırılıma çeviren şey Planlama'dır) ya da değişiklik birden fazla repoya yayılıyorsa yanlış seçim olur.
|
|
104
|
+
|
|
105
|
+
Widget cevabı tahmin etmek yerine takip eder: sen seçmeden önce yalnızca Faz 0 karosu çizilir, Kısa koşu Analiz karosunu hiç çizmez. İki autopilot girişi de bu soruyu sormaz, her zaman Tam koşar. Hangisinin koştuğunu `agent-state.json` `onlyDevelop` alanında tutar.
|
|
106
|
+
|
|
89
107
|
|
|
90
108
|
## Komutlar
|
|
91
109
|
|
|
@@ -266,11 +284,11 @@ Pipeline **Claude Code**, **Copilot CLI** ve **Codex CLI** üzerinde native çal
|
|
|
266
284
|
| ----------- | ----------------------- | ---------------------------------------------------------------------------------------------------------------- |
|
|
267
285
|
| Claude Code | `--claude` (varsayılan) | slash komutları + skill'ler + agent'lar + üç `PreToolUse` hook'u (secret scan, agent-guard, okuma-boyutu geçidi) |
|
|
268
286
|
| Copilot CLI | `--copilot` | talimatlar + 56 alt-komut skill'i + script'ler |
|
|
269
|
-
| Codex CLI | `--codex` | bir router skill + ref olarak 56 spec +
|
|
287
|
+
| Codex CLI | `--codex` | bir router skill + ref olarak 56 spec + 9 agent TOML + `AGENTS.md` bloğu + `codex mcp add` |
|
|
270
288
|
|
|
271
289
|
Skill'leri stack'e göre filtrele: `--platform=ios\|android\|all`.
|
|
272
290
|
|
|
273
|
-
**Codex neden
|
|
291
|
+
**Codex neden 56 değil de tek bir skill alıyor.** Codex, keşfettiği her skill'in adını
|
|
274
292
|
ve açıklamasını tek bir prompt bloğuna toplar ve blok taştığında girdileri hatasızca
|
|
275
293
|
düşürür. 0.145 üzerinde ölçüldü: 142 skill deklare eden bir plugin kurulduğunda sadece
|
|
276
294
|
75'i yüzeye çıktı ve alakasız bir kullanıcı skill'i tahliye edildi. Bu yüzden Codex'te
|
|
@@ -0,0 +1,102 @@
|
|
|
1
|
+
# 13. A language server for the answers regex cannot give
|
|
2
|
+
|
|
3
|
+
**Status:** Accepted · 2026-09-15
|
|
4
|
+
|
|
5
|
+
## Context
|
|
6
|
+
|
|
7
|
+
ADR-0010 built the code graph and wrote down what it gave up:
|
|
8
|
+
|
|
9
|
+
> Regex over comment-stripped source is not a parser. Definitions and imports
|
|
10
|
+
> survive that trade; call graphs and type resolution do not.
|
|
11
|
+
|
|
12
|
+
It named two consequences precisely. A reference resolves only when a name maps
|
|
13
|
+
to exactly one declaration, so `affected` under-reports on duplicated names. And
|
|
14
|
+
only type-like symbols are reference targets, because a bare lowercase name
|
|
15
|
+
matched across files is almost never a call to that exact declaration - an early
|
|
16
|
+
build that allowed function targets filled the hub list with `with`, `localized`,
|
|
17
|
+
`size` and `name`.
|
|
18
|
+
|
|
19
|
+
Those are not defects in the graph. They are the price of a five-second pass over
|
|
20
|
+
a 4,300-file tree with no toolchain. But they mean the graph cannot check a
|
|
21
|
+
specific claim, and a specific claim is what a review makes.
|
|
22
|
+
|
|
23
|
+
Meanwhile `sourcekit-lsp` sits in every Xcode installation, already on `PATH`,
|
|
24
|
+
and answers exactly those questions.
|
|
25
|
+
|
|
26
|
+
`swiftlens/swiftlens` demonstrated the idea as an MCP server. It could not be
|
|
27
|
+
used: its licence is the "SwiftLens Non-Commercial Use License 1.0", which
|
|
28
|
+
prohibits *"Use by a company or organization for internal development or
|
|
29
|
+
production"*, and the repository has been archived since 2025-07-17. The
|
|
30
|
+
underlying technology is unencumbered - SourceKit-LSP is Apple's, Apache-2.0.
|
|
31
|
+
|
|
32
|
+
## Decision
|
|
33
|
+
|
|
34
|
+
Build the capability in the companion MCP server as a `code_*` family, taking
|
|
35
|
+
swiftlens as a design reference and never as a code source. Same posture
|
|
36
|
+
ADR-0010 took toward graphify.
|
|
37
|
+
|
|
38
|
+
The toolkit rather than the pipeline: ADR-0010 already said "an MCP server (we
|
|
39
|
+
ship our own toolkit)", the pipeline is held to zero runtime dependencies by
|
|
40
|
+
ADR-0004, and shelling out to platform tooling is the toolkit's daily work.
|
|
41
|
+
|
|
42
|
+
Eight tools, one family for both languages, the language inferred from the file
|
|
43
|
+
extension. Seven read-only. Nothing writes - an agent already has an editing
|
|
44
|
+
tool, and giving a read-only family a write verb is how "read-only" stops
|
|
45
|
+
meaning anything.
|
|
46
|
+
|
|
47
|
+
## What measurement decided, not judgement
|
|
48
|
+
|
|
49
|
+
Three things were measured before any of this was designed, and each changed it.
|
|
50
|
+
|
|
51
|
+
**An empty answer is the default, not the exception.** `initialize` advertises
|
|
52
|
+
`referencesProvider: true` whether or not references will ever be non-empty.
|
|
53
|
+
Against a two-file package the same query answered 0 at 0.7s and 3 at 5.8s, with
|
|
54
|
+
nothing changed but the background index finishing. Against a 658-file package
|
|
55
|
+
with dependencies it had not finished at 70s. So the tools track the index -
|
|
56
|
+
the server opens an `indexing.<uuid>` work-done token and reports begin, report
|
|
57
|
+
and end through `$/progress` - wait for it, and return `indexReady`. Reporting
|
|
58
|
+
an unindexed empty list as "no references" would read as "unused", and a review
|
|
59
|
+
would act on it.
|
|
60
|
+
|
|
61
|
+
**The Xcode umbrella project cannot be served without a build server, and it
|
|
62
|
+
barely matters.** `sourcekit-lsp` takes compiler arguments from a build system.
|
|
63
|
+
An `.xcodeproj` root with no `buildServer.json` gets fallback arguments:
|
|
64
|
+
cross-file answers are impossible and diagnostics are wrong rather than absent.
|
|
65
|
+
Both are marked `semantic: false` with the remedy. The measured app has 41
|
|
66
|
+
`Package.swift` roots and one `.swift` file under the umbrella project, so the
|
|
67
|
+
limitation is real and narrow. Nothing writes a build config into a user's
|
|
68
|
+
repository.
|
|
69
|
+
|
|
70
|
+
**`xcode-build-server` turned out not to be required.** The first design
|
|
71
|
+
assumed a pre-built index store had to be found on disk, most likely under
|
|
72
|
+
DerivedData. It does not: sourcekit-lsp performs its own background index build
|
|
73
|
+
into `.build/index-build`, with no prior `swift build` and no build server. That
|
|
74
|
+
removed an install step and a whole class of staleness handling.
|
|
75
|
+
|
|
76
|
+
## What we deliberately gave up
|
|
77
|
+
|
|
78
|
+
- **Kotlin is written but unverified here.** JetBrains `kotlin-lsp` is Alpha,
|
|
79
|
+
partially closed source, needs a separate install and a JVM, and its Android
|
|
80
|
+
Gradle Plugin support is experimental. The probe and the absence path are
|
|
81
|
+
tested; the server path is exercised by no gate. Answers carry
|
|
82
|
+
`confidence: "alpha"` and both READMEs say so, rather than letting a green
|
|
83
|
+
suite imply coverage.
|
|
84
|
+
- **No editing.** `swift_replace_symbol_body` and its relatives are not ported.
|
|
85
|
+
- **Two servers at a time, ten-minute idle eviction.** A language server with
|
|
86
|
+
background indexing is a large resident process. Holding four quietly is a bug
|
|
87
|
+
report.
|
|
88
|
+
- **The code graph is not replaced.** It stays the wide cheap pass; this is what
|
|
89
|
+
a specific claim is checked against. `features/code-intelligence.md` has the
|
|
90
|
+
table.
|
|
91
|
+
|
|
92
|
+
## Consequences
|
|
93
|
+
|
|
94
|
+
- 87 tools become 95; the toolkit goes to 3.10.0 (a new tool is a minor bump).
|
|
95
|
+
- Two new ship gates, because the usual ones cannot see this code: 14c requires
|
|
96
|
+
that `tools/code-intel` import only `spawn` and `execFileSync` and name no
|
|
97
|
+
shell - gate 14b protects shell templates and there are none here - and 14d
|
|
98
|
+
forbids touching `process.stdout` or inheriting child stdio, since a child
|
|
99
|
+
language server speaks JSON-RPC too and corrupting fd 1 takes all 95 tools
|
|
100
|
+
down rather than one. Both were negative-tested by breaking what they guard.
|
|
101
|
+
- Pipeline consumption is opt-in and lands behind `prefs`, the way the code
|
|
102
|
+
graph does.
|
package/docs/adr/README.md
CHANGED
|
@@ -22,6 +22,7 @@ Format: lightly adapted from [Michael Nygard's ADR template](https://cognitect.c
|
|
|
22
22
|
| [0010](./0010-own-code-graph.md) | Own code graph, referenced from graphify, not forked | Accepted |
|
|
23
23
|
| [0011](./0011-dormant-ci.md) | CI dormant in-repo; pre-push gate is primary | Accepted |
|
|
24
24
|
| [0012](./0012-macos-only.md) | macOS only; Linux and Windows support removed | Accepted |
|
|
25
|
+
| [0013](./0013-lsp-code-intelligence.md) | A language server for the answers regex cannot give | Accepted |
|
|
25
26
|
|
|
26
27
|
## Writing a New ADR
|
|
27
28
|
|
|
@@ -19,4 +19,4 @@ back into. The ceiling can still be raised; it can no longer be raised for free.
|
|
|
19
19
|
|
|
20
20
|
---
|
|
21
21
|
|
|
22
|
-
Token estimate = ceil(chars / 4). Per-phase budget rule: warn = current+10% (rounded to nearest 50), max = current+25%. Gives ~6 edit cycles of headroom before warn trips - intentionally quiet under normal maintenance, loud when a phase grows unusually. Only the active phase is loaded (lazy). Recalibrated at v10.0.0 after the validator/consistency/simplifier/lesson gate contracts landed in phases 1-4. Recalibrated again at v10.9.0 after the verify-by-test (Phase 4 Step 3.7), update-check (Phase 0 Step 0.6), immutable-test (Phase 3 GREEN) and redTests re-entry contracts landed - Step 3.7 prose was compressed to a pointer into refs/features/verify-by-test.md before the recalibration. Total bumped 50000 -> 51000 at v12.5.0 after the worktree residue/traversal-prune contract (Phase 0 + Phase 5 heal) and the Reflexion causal-diagnosis contract (Phase 4 lesson memory) landed; the prose was compressed first (161 tokens reclaimed) and every per-phase max still passes - only the aggregate needed room. Recalibrated again at v13.6.0 after the install-relative path correction: an instruction that names `pipeline/scripts/x` resolves only from a repo checkout, and a run happens in the user's worktree, so 157 references across these docs moved to `$HOME/.claude/...` at +5 bytes each - 196 tokens of pure correctness cost. Same discipline as before: prose was compressed FIRST (149 tokens reclaimed, by pointing Phase 1's Figma tier table at the Phase 0 probe that already resolved it and Phase 4's Codex constraints at the always-loaded AGENTS.md block), and only then were the budgets moved. Five warn lines had been permanently amber, which makes the amber tier useless as a signal, so every warn was reset to the documented current+10% and the four maxes that the new warn would have collided with were reset to current+25%. Aggregate 51000 -> 51500. Total bumped 51500 -> 52200 at v14.0.0 after Phase 4 Review entered the four --dev mode phase sets and the criteria-resolution contract (Step 1.78) landed. Same discipline as every prior bump: prose was compressed FIRST, 820 tokens reclaimed, before the number moved. Two of those compressions are structural rather than cosmetic - the hardcoded SwiftUI interaction list in Step 1.5 and the SwiftUI convention paragraph in Step 2.8 were transcriptions of rules that now live in a scoped registry, so keeping them here would have re-created the drift this release exists to remove, and the third moved the Step 1.78 full contract into refs/features/skill-conformance.md leaving a pointer. What remains is contract text that cannot be inferred: the manifest's four consumer-visible parts, the conformance checklist the reviewers must return, and the fail-closed semantics. Every per-phase max still passes (phase-4 12405/14750); only the aggregate needed room. Total bumped 52200 -> 52700 at v14.1.0 after two more contracts landed: stack skill routing (Phase 3 pre-flight step 9) and worktree finalize (Phase 6 step 9). Compression came first, as always, and twice: 224 tokens out of Phase 3 by pointing its criteria-ledger and routing steps at their feature files instead of restating them, and 190 out of Phase 6 by moving the finalize contract into refs/features/worktree-finalize.md and leaving the invocation plus the exit-3 semantics. Both new contracts follow the pattern the earlier ones set: the phase doc carries the call and the decision, the feature file carries the reasoning, and the feature files are outside this budget because it loops only the eight phase-N-* keys. Every per-phase max still passes (phase-3 7677/8950, phase-6 5223/6150 and both under warn); only the aggregate needed room. Total bumped 52700 -> 52750 for the Phase 0 Step 3 branch-persistence correction: the step wrote the legacy `projects[].branches` while the TTL filter two sections below read `global.recentBranches`, and both spots named a `{name, lastUsed}` shape the schema rejects (`branch` required, `additionalProperties: false`), so the recent-branch picker option could never populate and a literal implementation would have failed prefs validation. Naming the right target, the right key and the legacy field to avoid costs 41 tokens over the one line it replaces. Compression came first and was applied three times to the replacement text itself, from 120 tokens down to 66, by moving the rationale out of the phase doc entirely: the reasoning now lives where it is enforced, in the migrate-prefs carry-forward comment and the smoke-pref-migration f7 block, leaving the phase doc with only the instruction. 50 was the smallest step that clears it; phase-0-init sits at 10893/12400, far under its own max, so this is purely an aggregate ceiling. v15.0.0: total 52750 -> 53100, the stack-skill tables in phase-1/2/4 now carry plugin-namespaced names (ai-<stack>-toolkit:<skill>) - functional prefixes, ~170 tokens. v15.10.0: total 53350 -> 53950 for the memory-recall + context-offload contracts (Phase 1 two-block durable-knowledge injection and its telemetry, Phase 3 build-log offload pipe, Phase 4 ranked prior art, offload pipe and recall telemetry). Compression came first and twice, taking the new prose from 1168 tokens to 580: the reasoning behind the two blocks lives in multi-agent-refs/prompt-assembly.md and the reasoning behind the offload filter lives in the offload-ref.sh header, both outside this budget, so the phase docs carry only the call, the pref that gates it and the one fact an agent cannot infer - that the evidence gate still reads the whole build log, so offloading changes what is read, never what counts as a verified pass. Every per-phase max still passes (phase-3 7985/8950, phase-4 12997/14750); phase-3 and phase-4 crossed their warn lines and are left amber on purpose, because that is the signal that those two docs are the next ones needing structural compression rather than another bump. v15.13.0: total 53950 -> 54050 for the prefs-to-flag bridges. Five settings had shipped declared-but-inert: contextOffload.minLines and .tailLines (fixed in 15.11.0), learningsLedger.maxBriefEntries, and testGap.scanTree and .promoteSeverity - the last two declared in the schema AND implemented as flags in the scanner, with nothing in between reading the pref and passing the flag. Wiring three of them costs the phase docs 94 tokens, which is the wiring itself and not prose: two `--max` substitutions and a three-line GAP_FLAGS block. Compression came first and twice, as always: the rationale that would have sat in phase-5 now lives in the header of smoke-prefs-consumed.sh, the gate that makes this class fail a build instead of shipping, and a `--severity-promote` table row was dropped because the invocation above it now shows the flag and names the pref that triggers it, which the row did not. 100 was the smallest step that clears it. Every per-phase max still passes; phase-3 and phase-4 remain amber on purpose. v15.14.0: total 54050 -> 54400 for the supported-version gate. Phase 0 Step 0.6 stopped being purely advisory: a release can now publish an npm dist-tag `required` that names the oldest runnable version, and below it the run halts instead of nagging. What the phase doc has to carry is the part an agent cannot infer - the third stdout field, that the halt is identical in autopilot, and that the run must NOT continue on the freshly updated install because its docs were already loaded from the old version. Compression came first, as always, and took the new prose from 469 tokens to 337: the rationale for the floor, the exemption list, the fail-open rules and the `npm dist-tag add` recipe all moved to multi-agent-refs/rules.md "Supported Version Gate" (loaded by 25 commands, outside this budget) and to the header of require-supported-version.sh, leaving the phase doc with the call, the decision table and the halt. 350 was the smallest step that clears it. Every per-phase max still passes (phase-0-init 11230/12400); phase-3 and phase-4 remain amber on purpose. v15.17.0: total 54400 -> 54900 for the Phase 1 analysis-document step. Phase 2 and Phase 3 pre-flights had BLOCKED on `analysis/<feature>-<platform>.md` since v9.0.0 while nothing produced it, so a full run either aborted at Phase 2 or the model ignored its own BLOCKING contract; Step 4 is the producer. What the phase doc carries is only what cannot be inferred: the when-table (taskType x Figma reference), the four refs in load order, the two artefacts, and that the doc validator fails closed. Compression came first and took the step from 745 tokens to 497: the history of why the gap existed moved to the CHANGELOG, the per-ref one-line descriptions moved into the refs' own headers, and the autopilot carve-out collapsed to one clause. The 17.4k-token analysis engine itself is NOT in this budget - it moved out of commands/ into multi-agent-refs/analysis/{locked,evidence,synthesis,render}.md, loaded on demand, which also took analysis/SKILL.md from 18081 to 5974 tokens and retired its lint grace entry. 500 was the smallest step that clears it; phase-1-analysis sits at 4338/4600 and is amber on purpose, like phase-3 and phase-4. v15.18.0: total 54900 -> 55250 for analysis mode. Three phase docs gained a mode branch that cannot be inferred: Phase 4 reviews a document instead of a diff (validator, the one question reviewers answer, the open-question walk), and Phase 6 publishes instead of committing. Compression came first and was applied twice to the new prose and once to old: the Phase 4 branch went from 320 tokens to 180 and the Phase 6 branch from 190 to 120 by pointing at multi-agent-refs/analysis/{resolve,render}.md, which now hold the walks themselves, and the front-matter parse contract stopped being spelled out in both pre-flights. The analysis engine keeps leaving this budget rather than entering it: intake joined locked/evidence/synthesis/render/resolve in multi-agent-refs/analysis/, which is what let analysis/SKILL.md drop under the 6000 hard cap after its grace entry was retired. 350 was the smallest step that clears it; phase-4 and phase-6 are amber on purpose, as phase-1 and phase-3 already were. v15.20.0: total 55250 -> 55500 for the TDD bridge. Phase 3 pre-flight read the analysis doc's concept table and even said test method names come from it, while nothing read Section 15 - so the RED step invented tests and the analysis test matrix never reached development. Phase 3 step 5b now loads it into state.dev.testPlan[] and Phase 4 step 1.45 cross-checks every planned row against a real test, which is what turns "analysis quality is output quality" from a slogan into a finding. Compression came first on both blocks, 300 tokens down to 175, by dropping the enumerated failure modes to one line each and the rationale to one clause; the reasoning lives in the CHANGELOG. 250 was the smallest step that clears it. v15.21.0: total 55500 -> 55800 for the post-analysis confirmation. Phase 2 gained Step 0.9, the last human checkpoint before Phase 3: derived values are shown for confirmation and only Section 20 rows are asked, through the resolve engine that already exists in refs. It belongs here rather than Phase 4 because Phase 4 runs after development, where an answer arrives too late to change anything. Compression came first and twice, 430 tokens down to 250, by collapsing the derived-vs-asked explanation to one sentence each and moving the walk itself to multi-agent-refs/analysis/resolve.md, which Phase 4 and analysis-resolve already mount. 300 was the smallest step that clears it. v15.22.0: total 55800 -> 55900 for the analyst-toolkit hooks. Phase 1 Step 4 now names the two prefs that decide whether a document is produced at all and how deep it goes (forceFull, mode) - the first of those had shipped declared-but-inert and smoke-prefs-consumed caught it - and Phase 4 triage gained one clause: a finding that blames a third-party library asks evidence-github whether it is already open upstream, which turns it into a deferred item with a citation instead of Phase 3 rework on code that is not ours. Compression came first and three times, taking the new prose from 220 tokens to 110, and the Phase 1d evidence contract itself never entered this budget - it lives in multi-agent-refs/analysis/evidence.md beside the phases it belongs to. 100 was the smallest step that clears it, leaving 34 tokens of headroom. phase-4 stays amber and the debt named at v15.10.0 stands: it is the doc that needs structural compression rather than another bump, and the two candidates are the inline triage JSON shape and the 3.4 telemetry block, both of which restate something already authoritative elsewhere. v16.0.0: total 55900 -> 56350 for the depth picker. `--dev` and the four dev-* commands are gone; depth is Phase 0 Step 7.5, which costs phase-0-init a step it did not have. Compression came first and three times, taking the step from 530 tokens to 300: the question wording, the per-taskType recommendation and the mode tables all live in phases/modes.md (outside this budget), so the phase doc carries only what an agent cannot infer - that the step runs after Step 7 and why, who is exempt, that ASK_CHOICE_DEFAULT must be passed explicitly because ask-choice.sh takes the FIRST option on a non-TTY, and that Short flips the Phase 1/2 tiles late rather than pre-marking them. The phase-4 telemetry block named as compression debt at v15.22.0 was collapsed to an emit() helper (-27) and the four dev-* mode files left the tree entirely, but neither offsets a genuinely new phase step. 450 was the smallest step that clears it, leaving 119 tokens of headroom. phase-4 remains amber and its other named candidate, the inline triage JSON shape, was left alone on purpose: it is the prompt the triage agent is handed, not a restatement for readers. v16.2.0: total 56350 -> 56600 for the spec-freshness and reuse-tag contracts. Phase 3 step 3 had compared `state.run.lastAnalysisDigest` since it was written, against a key nothing ever set and that the state schema did not declare, so the staleness branch was unreachable and every run reported fresh by default. Phase 1 now persists the digest and a `base_commit` anchor, and step 3 gained the repo-drift half the digest cannot see: a reused document keeps a matching digest precisely because its evidence inputs did not change, while the code underneath it moved. The second contract is the Section 14 tag reaching development: Phase 2 carries it onto the todo as `sourceTag` and Phase 3 treats it as an instruction, which is what stops a Reuse row from being re-implemented. Compression came first and took the four additions from 380 tokens to 214, by moving every rationale clause out of the phase docs: why the commit anchor exists rather than a digest recomputation lives in this note and the CHANGELOG, and the schema descriptions carry the field semantics. The baseline had 9 tokens of headroom, so no addition of any size could have fit without a bump. 250 was the smallest step that clears it, leaving 45 tokens. phase-3 and phase-4 remain amber. v16.13.0: total 57600 -> 57700 for the code-graph injection and the fable-rung switch. Phase 1 gained Step 2.6 (query the graph, hand Explore a ranked starting set), Phase 7 gained the post-branch graph refresh, and Phase 0 Step 0 gained one line: a prefs switch that resolves every preferredModel: fable persona to opus for the run, which also collapses the Phase 4 Claude Code panel from three reviewers to two. Compression came first and mostly structurally: of roughly 1,630 tokens of new contract text, 1,310 never entered this budget at all - the whole code-graph contract lives in multi-agent-refs/features/code-graph.md (604) and the fable switch's scope table, per-host effects and cost-accounting consequence live in features/model-fallback.md (+707), leaving the phase docs with the call, the pref that gates it and the one fact an agent cannot infer. Phase 4 was compressed on top of that: its TLDR restated the reviewer matrix 270 lines below it, so 36 tokens came back and the doc nets +6 despite carrying two new clauses. One of those clauses is a correction rather than a feature - the consensus rule still said reviewerCount is 2 on Claude Code, which stopped being true when the third reviewer landed in 16.12.0, and the cross-CLI smoke never caught it because it reads the matrix line instead. 100 was the smallest step that clears it, leaving 54 tokens. phase-3 and phase-4 remain amber. v16.17.0: total 57700 -> 57850 for the platform-parity cross-check. Phase 4 gained Step 1.8: when dev-context carries a counterpart app repo, the review compares the change against the other platform on four axes. Compression came first and structurally, as always - of roughly 1,610 tokens of new contract text, 1,490 never entered this budget at all, because the four axes, the file cap, the graph-query recipe, the read-only prohibitions and the rule that an extractor miss may not be reported as an absence all live in multi-agent-refs/platform-parity.md. The step itself was then cut from ~200 tokens to 120 by deleting everything the ref already owns, leaving the trigger, the pointer and the two facts an agent must not infer: the counterpart repo is read-only, and parity findings are never blocking. The baseline had 13 tokens of headroom, so no addition of any size could have fit without a bump. 150 was the smallest step that clears it, leaving 35 tokens. phase-3 and phase-4 remain amber, and phase-4's structural-compression debt still stands. v16.20.0: phase-4-review max 14750 -> 15150 and total 58250 -> 60250 for the cross-round review delta, the scope self-check handoff and the circuit-breaker wiring. Compression came first and structurally: of roughly 3,900 tokens of new contract text, 2,700 never entered this budget at all - the previous-round-findings block, the scope-self-check block, the Step 3.8 state merge, telemetry and picker wording live in multi-agent-refs/features/review-delta.md, and the scope-check record rules and consumers in features/scope-check.md - so the phase docs carry the call, the pref that gates it and the exit table. The Phase 3 stability rule and the trigger-3 write were cut twice more before the bump; phase-3 stays under its max (8692/8950). Phase 4 is the first per-phase max raised since v10.9.0: the doc gained three steps that cannot be inferred (a per-round triage file, a prefix block that changes what reviewers report, and a halt condition), and its structural-compression debt (the inline triage JSON shape, named at v15.10.0) still stands and is the next candidate. 15150 and 60250 were the smallest steps that clear it, leaving 25 and 45 tokens. v16.23.0: phase-0-init max 12400 -> 12500 and total 60250 -> 60500 for the widget-registration call and the accounting gate. Phase 0 gained the `tiles` call and the exit-3 rule, Phase 7 gained the run report; together they are contract an agent cannot infer - which call registers this host's widget, and that a completion is refused without recorded spend. Compression came first and twice, taking the new prose from 472 tokens to 255: the per-host call list moved into tracker-contract.md "The card is not the widget" and the record-then-rerun recovery into "Accounting is a gate", both outside this budget, leaving the phase docs with the call and the one fact that cannot be looked up. 100 and 250 were the smallest steps that clear it, leaving 74 and 40 tokens. phase-3 and phase-4 remain amber. v16.24.0: total 60500 -> 60750 for visual evidence. Four phase docs gained one instruction each that cannot be inferred: Phase 0 keeps the issue's own images as the pre-fix evidence, Phase 3 captures the fixed state (there and not Phase 5, because every autopilot and --local entry drops Phase 5), Phase 5 hosts the flow recording when it runs, and Phase 6 blocks on a required artefact that is neither attached nor explained. Compression came first and twice, 42 tokens back, and the contract itself never entered this budget: the trigger matrix, the three video tiers, the size-degradation ladder and both render shapes live in multi-agent-refs/features/visual-evidence.md. phase-0-init cleared its own max without a bump. 250 was the smallest step that clears the aggregate. phase-3 and phase-4 remain amber. v17.0.0: phase-0-init max 13000 -> 13100 and total 62400 -> 62500 for the evidence-verdict writer. Phase 0 Step 7.7 probed with `--platform "$PLATFORM"`, a variable no phase document ever assigned, and it was gated on `visualEvidence.required`, which no phase document ever wrote - five readers, zero writers - so the step, Phase 3's capture and Phase 6's blocker were all unreachable and the pipeline reported nothing wrong. The step now writes the verdict and derives the platform from the stack, skipping the probe with a recorded reason when there is no device platform rather than passing the empty string the probe refuses with exit 2. Compression came first and three times, 30 tokens back from the Step 7.7 index rule that restated Step 7.5 verbatim and 55 from the new block itself; the reasoning never entered this budget, because who writes the verdict and how the platform is derived live in features/visual-evidence.md sections 1a and 1b. phase-3-dev max 9250 -> 9300 in the same change: it reads the platform back from state and re-decides the provisional verdict before capturing, which is the half of the fix that makes Phase 3 honest rather than merely reachable. Compressed three times first, 29 tokens back, by pointing its Phase-5 rationale and its tier mapping at visual-evidence.md sections 3 and 4.3 where both already live. 100, 50 and 150 were the smallest steps that clear it, leaving 24, 15 and 26 tokens. phase-3 and phase-4 remain amber. v17.1.0: total 62550 -> 62600 for the Figma/toolkit MCP distinction and web as an evidence platform. Three phase docs said "MCP forbidden" without qualifying it, while the gate that enforces it (smoke-no-mcp-in-dev-phases.sh) has always matched `figma` and nothing else - so the prose banned the screenshot, xcodebuild and UI-test tools that Phase 3 Steps 3.4 and 3.55 actually call, which is one way a run reaches Phase 7 with no evidence. Phase 0 gained one `web` arm in the platform derivation, now that run-ui-tests.sh has a web arm to derive it for. Compression came first and three times, taking the new prose from 175 tokens to 30: the reasoning moved to rules.md, whose own seven-row Figma phase matrix was deleted in the same pass because it duplicated rules/figma-pipeline.md "Phase access matrix" two paragraphs below this file's own instruction not to duplicate that rule file - 74 tokens back there, which is why phase-3-dev cleared its max without a bump (9298/9300). 50 was the smallest step that clears the aggregate, leaving 37 tokens. phase-3 and phase-4 remain amber. v17.1.0 (2): total 62600 -> 62700 for the plan reaching the task widget. Phase 2 computed tasks[], their order and their dependsOn[] edges, stored them, and used them to drive Phase 3's ready-task picker - and none of it was visible on the surface the user actually watches; the card had drawn sub-phases for releases, the widget never had. Phase 2 gains one call. Compression came first and twice: the tasks[]-to-todos[] jq blob left the phase doc for plan-todos.sh `set`, which now accepts a planning-output document directly (-28, and it removes a mapping two files defined, of which this was the untested copy), and the new step's own prose was cut from 116 tokens to 61. The parsing of the plan itself never entered this budget - it lives in phase-tracker.sh `plan`, which owns the sub-phase structure it writes. 100 was the smallest step that clears it, leaving 75 tokens. phase-3 and phase-4 remain amber.
|
|
22
|
+
Token estimate = ceil(chars / 4). Per-phase budget rule: warn = current+10% (rounded to nearest 50), max = current+25%. Gives ~6 edit cycles of headroom before warn trips - intentionally quiet under normal maintenance, loud when a phase grows unusually. Only the active phase is loaded (lazy). Recalibrated at v10.0.0 after the validator/consistency/simplifier/lesson gate contracts landed in phases 1-4. Recalibrated again at v10.9.0 after the verify-by-test (Phase 4 Step 3.7), update-check (Phase 0 Step 0.6), immutable-test (Phase 3 GREEN) and redTests re-entry contracts landed - Step 3.7 prose was compressed to a pointer into refs/features/verify-by-test.md before the recalibration. Total bumped 50000 -> 51000 at v12.5.0 after the worktree residue/traversal-prune contract (Phase 0 + Phase 5 heal) and the Reflexion causal-diagnosis contract (Phase 4 lesson memory) landed; the prose was compressed first (161 tokens reclaimed) and every per-phase max still passes - only the aggregate needed room. Recalibrated again at v13.6.0 after the install-relative path correction: an instruction that names `pipeline/scripts/x` resolves only from a repo checkout, and a run happens in the user's worktree, so 157 references across these docs moved to `$HOME/.claude/...` at +5 bytes each - 196 tokens of pure correctness cost. Same discipline as before: prose was compressed FIRST (149 tokens reclaimed, by pointing Phase 1's Figma tier table at the Phase 0 probe that already resolved it and Phase 4's Codex constraints at the always-loaded AGENTS.md block), and only then were the budgets moved. Five warn lines had been permanently amber, which makes the amber tier useless as a signal, so every warn was reset to the documented current+10% and the four maxes that the new warn would have collided with were reset to current+25%. Aggregate 51000 -> 51500. Total bumped 51500 -> 52200 at v14.0.0 after Phase 4 Review entered the four --dev mode phase sets and the criteria-resolution contract (Step 1.78) landed. Same discipline as every prior bump: prose was compressed FIRST, 820 tokens reclaimed, before the number moved. Two of those compressions are structural rather than cosmetic - the hardcoded SwiftUI interaction list in Step 1.5 and the SwiftUI convention paragraph in Step 2.8 were transcriptions of rules that now live in a scoped registry, so keeping them here would have re-created the drift this release exists to remove, and the third moved the Step 1.78 full contract into refs/features/skill-conformance.md leaving a pointer. What remains is contract text that cannot be inferred: the manifest's four consumer-visible parts, the conformance checklist the reviewers must return, and the fail-closed semantics. Every per-phase max still passes (phase-4 12405/14750); only the aggregate needed room. Total bumped 52200 -> 52700 at v14.1.0 after two more contracts landed: stack skill routing (Phase 3 pre-flight step 9) and worktree finalize (Phase 6 step 9). Compression came first, as always, and twice: 224 tokens out of Phase 3 by pointing its criteria-ledger and routing steps at their feature files instead of restating them, and 190 out of Phase 6 by moving the finalize contract into refs/features/worktree-finalize.md and leaving the invocation plus the exit-3 semantics. Both new contracts follow the pattern the earlier ones set: the phase doc carries the call and the decision, the feature file carries the reasoning, and the feature files are outside this budget because it loops only the eight phase-N-* keys. Every per-phase max still passes (phase-3 7677/8950, phase-6 5223/6150 and both under warn); only the aggregate needed room. Total bumped 52700 -> 52750 for the Phase 0 Step 3 branch-persistence correction: the step wrote the legacy `projects[].branches` while the TTL filter two sections below read `global.recentBranches`, and both spots named a `{name, lastUsed}` shape the schema rejects (`branch` required, `additionalProperties: false`), so the recent-branch picker option could never populate and a literal implementation would have failed prefs validation. Naming the right target, the right key and the legacy field to avoid costs 41 tokens over the one line it replaces. Compression came first and was applied three times to the replacement text itself, from 120 tokens down to 66, by moving the rationale out of the phase doc entirely: the reasoning now lives where it is enforced, in the migrate-prefs carry-forward comment and the smoke-pref-migration f7 block, leaving the phase doc with only the instruction. 50 was the smallest step that clears it; phase-0-init sits at 10893/12400, far under its own max, so this is purely an aggregate ceiling. v15.0.0: total 52750 -> 53100, the stack-skill tables in phase-1/2/4 now carry plugin-namespaced names (ai-<stack>-toolkit:<skill>) - functional prefixes, ~170 tokens. v15.10.0: total 53350 -> 53950 for the memory-recall + context-offload contracts (Phase 1 two-block durable-knowledge injection and its telemetry, Phase 3 build-log offload pipe, Phase 4 ranked prior art, offload pipe and recall telemetry). Compression came first and twice, taking the new prose from 1168 tokens to 580: the reasoning behind the two blocks lives in multi-agent-refs/prompt-assembly.md and the reasoning behind the offload filter lives in the offload-ref.sh header, both outside this budget, so the phase docs carry only the call, the pref that gates it and the one fact an agent cannot infer - that the evidence gate still reads the whole build log, so offloading changes what is read, never what counts as a verified pass. Every per-phase max still passes (phase-3 7985/8950, phase-4 12997/14750); phase-3 and phase-4 crossed their warn lines and are left amber on purpose, because that is the signal that those two docs are the next ones needing structural compression rather than another bump. v15.13.0: total 53950 -> 54050 for the prefs-to-flag bridges. Five settings had shipped declared-but-inert: contextOffload.minLines and .tailLines (fixed in 15.11.0), learningsLedger.maxBriefEntries, and testGap.scanTree and .promoteSeverity - the last two declared in the schema AND implemented as flags in the scanner, with nothing in between reading the pref and passing the flag. Wiring three of them costs the phase docs 94 tokens, which is the wiring itself and not prose: two `--max` substitutions and a three-line GAP_FLAGS block. Compression came first and twice, as always: the rationale that would have sat in phase-5 now lives in the header of smoke-prefs-consumed.sh, the gate that makes this class fail a build instead of shipping, and a `--severity-promote` table row was dropped because the invocation above it now shows the flag and names the pref that triggers it, which the row did not. 100 was the smallest step that clears it. Every per-phase max still passes; phase-3 and phase-4 remain amber on purpose. v15.14.0: total 54050 -> 54400 for the supported-version gate. Phase 0 Step 0.6 stopped being purely advisory: a release can now publish an npm dist-tag `required` that names the oldest runnable version, and below it the run halts instead of nagging. What the phase doc has to carry is the part an agent cannot infer - the third stdout field, that the halt is identical in autopilot, and that the run must NOT continue on the freshly updated install because its docs were already loaded from the old version. Compression came first, as always, and took the new prose from 469 tokens to 337: the rationale for the floor, the exemption list, the fail-open rules and the `npm dist-tag add` recipe all moved to multi-agent-refs/rules.md "Supported Version Gate" (loaded by 25 commands, outside this budget) and to the header of require-supported-version.sh, leaving the phase doc with the call, the decision table and the halt. 350 was the smallest step that clears it. Every per-phase max still passes (phase-0-init 11230/12400); phase-3 and phase-4 remain amber on purpose. v15.17.0: total 54400 -> 54900 for the Phase 1 analysis-document step. Phase 2 and Phase 3 pre-flights had BLOCKED on `analysis/<feature>-<platform>.md` since v9.0.0 while nothing produced it, so a full run either aborted at Phase 2 or the model ignored its own BLOCKING contract; Step 4 is the producer. What the phase doc carries is only what cannot be inferred: the when-table (taskType x Figma reference), the four refs in load order, the two artefacts, and that the doc validator fails closed. Compression came first and took the step from 745 tokens to 497: the history of why the gap existed moved to the CHANGELOG, the per-ref one-line descriptions moved into the refs' own headers, and the autopilot carve-out collapsed to one clause. The 17.4k-token analysis engine itself is NOT in this budget - it moved out of commands/ into multi-agent-refs/analysis/{locked,evidence,synthesis,render}.md, loaded on demand, which also took analysis/SKILL.md from 18081 to 5974 tokens and retired its lint grace entry. 500 was the smallest step that clears it; phase-1-analysis sits at 4338/4600 and is amber on purpose, like phase-3 and phase-4. v15.18.0: total 54900 -> 55250 for analysis mode. Three phase docs gained a mode branch that cannot be inferred: Phase 4 reviews a document instead of a diff (validator, the one question reviewers answer, the open-question walk), and Phase 6 publishes instead of committing. Compression came first and was applied twice to the new prose and once to old: the Phase 4 branch went from 320 tokens to 180 and the Phase 6 branch from 190 to 120 by pointing at multi-agent-refs/analysis/{resolve,render}.md, which now hold the walks themselves, and the front-matter parse contract stopped being spelled out in both pre-flights. The analysis engine keeps leaving this budget rather than entering it: intake joined locked/evidence/synthesis/render/resolve in multi-agent-refs/analysis/, which is what let analysis/SKILL.md drop under the 6000 hard cap after its grace entry was retired. 350 was the smallest step that clears it; phase-4 and phase-6 are amber on purpose, as phase-1 and phase-3 already were. v15.20.0: total 55250 -> 55500 for the TDD bridge. Phase 3 pre-flight read the analysis doc's concept table and even said test method names come from it, while nothing read Section 15 - so the RED step invented tests and the analysis test matrix never reached development. Phase 3 step 5b now loads it into state.dev.testPlan[] and Phase 4 step 1.45 cross-checks every planned row against a real test, which is what turns "analysis quality is output quality" from a slogan into a finding. Compression came first on both blocks, 300 tokens down to 175, by dropping the enumerated failure modes to one line each and the rationale to one clause; the reasoning lives in the CHANGELOG. 250 was the smallest step that clears it. v15.21.0: total 55500 -> 55800 for the post-analysis confirmation. Phase 2 gained Step 0.9, the last human checkpoint before Phase 3: derived values are shown for confirmation and only Section 20 rows are asked, through the resolve engine that already exists in refs. It belongs here rather than Phase 4 because Phase 4 runs after development, where an answer arrives too late to change anything. Compression came first and twice, 430 tokens down to 250, by collapsing the derived-vs-asked explanation to one sentence each and moving the walk itself to multi-agent-refs/analysis/resolve.md, which Phase 4 and analysis-resolve already mount. 300 was the smallest step that clears it. v15.22.0: total 55800 -> 55900 for the analyst-toolkit hooks. Phase 1 Step 4 now names the two prefs that decide whether a document is produced at all and how deep it goes (forceFull, mode) - the first of those had shipped declared-but-inert and smoke-prefs-consumed caught it - and Phase 4 triage gained one clause: a finding that blames a third-party library asks evidence-github whether it is already open upstream, which turns it into a deferred item with a citation instead of Phase 3 rework on code that is not ours. Compression came first and three times, taking the new prose from 220 tokens to 110, and the Phase 1d evidence contract itself never entered this budget - it lives in multi-agent-refs/analysis/evidence.md beside the phases it belongs to. 100 was the smallest step that clears it, leaving 34 tokens of headroom. phase-4 stays amber and the debt named at v15.10.0 stands: it is the doc that needs structural compression rather than another bump, and the two candidates are the inline triage JSON shape and the 3.4 telemetry block, both of which restate something already authoritative elsewhere. v16.0.0: total 55900 -> 56350 for the depth picker. `--dev` and the four dev-* commands are gone; depth is Phase 0 Step 7.5, which costs phase-0-init a step it did not have. Compression came first and three times, taking the step from 530 tokens to 300: the question wording, the per-taskType recommendation and the mode tables all live in phases/modes.md (outside this budget), so the phase doc carries only what an agent cannot infer - that the step runs after Step 7 and why, who is exempt, that ASK_CHOICE_DEFAULT must be passed explicitly because ask-choice.sh takes the FIRST option on a non-TTY, and that Short flips the Phase 1/2 tiles late rather than pre-marking them. The phase-4 telemetry block named as compression debt at v15.22.0 was collapsed to an emit() helper (-27) and the four dev-* mode files left the tree entirely, but neither offsets a genuinely new phase step. 450 was the smallest step that clears it, leaving 119 tokens of headroom. phase-4 remains amber and its other named candidate, the inline triage JSON shape, was left alone on purpose: it is the prompt the triage agent is handed, not a restatement for readers. v16.2.0: total 56350 -> 56600 for the spec-freshness and reuse-tag contracts. Phase 3 step 3 had compared `state.run.lastAnalysisDigest` since it was written, against a key nothing ever set and that the state schema did not declare, so the staleness branch was unreachable and every run reported fresh by default. Phase 1 now persists the digest and a `base_commit` anchor, and step 3 gained the repo-drift half the digest cannot see: a reused document keeps a matching digest precisely because its evidence inputs did not change, while the code underneath it moved. The second contract is the Section 14 tag reaching development: Phase 2 carries it onto the todo as `sourceTag` and Phase 3 treats it as an instruction, which is what stops a Reuse row from being re-implemented. Compression came first and took the four additions from 380 tokens to 214, by moving every rationale clause out of the phase docs: why the commit anchor exists rather than a digest recomputation lives in this note and the CHANGELOG, and the schema descriptions carry the field semantics. The baseline had 9 tokens of headroom, so no addition of any size could have fit without a bump. 250 was the smallest step that clears it, leaving 45 tokens. phase-3 and phase-4 remain amber. v16.13.0: total 57600 -> 57700 for the code-graph injection and the fable-rung switch. Phase 1 gained Step 2.6 (query the graph, hand Explore a ranked starting set), Phase 7 gained the post-branch graph refresh, and Phase 0 Step 0 gained one line: a prefs switch that resolves every preferredModel: fable persona to opus for the run, which also collapses the Phase 4 Claude Code panel from three reviewers to two. Compression came first and mostly structurally: of roughly 1,630 tokens of new contract text, 1,310 never entered this budget at all - the whole code-graph contract lives in multi-agent-refs/features/code-graph.md (604) and the fable switch's scope table, per-host effects and cost-accounting consequence live in features/model-fallback.md (+707), leaving the phase docs with the call, the pref that gates it and the one fact an agent cannot infer. Phase 4 was compressed on top of that: its TLDR restated the reviewer matrix 270 lines below it, so 36 tokens came back and the doc nets +6 despite carrying two new clauses. One of those clauses is a correction rather than a feature - the consensus rule still said reviewerCount is 2 on Claude Code, which stopped being true when the third reviewer landed in 16.12.0, and the cross-CLI smoke never caught it because it reads the matrix line instead. 100 was the smallest step that clears it, leaving 54 tokens. phase-3 and phase-4 remain amber. v16.17.0: total 57700 -> 57850 for the platform-parity cross-check. Phase 4 gained Step 1.8: when dev-context carries a counterpart app repo, the review compares the change against the other platform on four axes. Compression came first and structurally, as always - of roughly 1,610 tokens of new contract text, 1,490 never entered this budget at all, because the four axes, the file cap, the graph-query recipe, the read-only prohibitions and the rule that an extractor miss may not be reported as an absence all live in multi-agent-refs/platform-parity.md. The step itself was then cut from ~200 tokens to 120 by deleting everything the ref already owns, leaving the trigger, the pointer and the two facts an agent must not infer: the counterpart repo is read-only, and parity findings are never blocking. The baseline had 13 tokens of headroom, so no addition of any size could have fit without a bump. 150 was the smallest step that clears it, leaving 35 tokens. phase-3 and phase-4 remain amber, and phase-4's structural-compression debt still stands. v16.20.0: phase-4-review max 14750 -> 15150 and total 58250 -> 60250 for the cross-round review delta, the scope self-check handoff and the circuit-breaker wiring. Compression came first and structurally: of roughly 3,900 tokens of new contract text, 2,700 never entered this budget at all - the previous-round-findings block, the scope-self-check block, the Step 3.8 state merge, telemetry and picker wording live in multi-agent-refs/features/review-delta.md, and the scope-check record rules and consumers in features/scope-check.md - so the phase docs carry the call, the pref that gates it and the exit table. The Phase 3 stability rule and the trigger-3 write were cut twice more before the bump; phase-3 stays under its max (8692/8950). Phase 4 is the first per-phase max raised since v10.9.0: the doc gained three steps that cannot be inferred (a per-round triage file, a prefix block that changes what reviewers report, and a halt condition), and its structural-compression debt (the inline triage JSON shape, named at v15.10.0) still stands and is the next candidate. 15150 and 60250 were the smallest steps that clear it, leaving 25 and 45 tokens. v16.23.0: phase-0-init max 12400 -> 12500 and total 60250 -> 60500 for the widget-registration call and the accounting gate. Phase 0 gained the `tiles` call and the exit-3 rule, Phase 7 gained the run report; together they are contract an agent cannot infer - which call registers this host's widget, and that a completion is refused without recorded spend. Compression came first and twice, taking the new prose from 472 tokens to 255: the per-host call list moved into tracker-contract.md "The card is not the widget" and the record-then-rerun recovery into "Accounting is a gate", both outside this budget, leaving the phase docs with the call and the one fact that cannot be looked up. 100 and 250 were the smallest steps that clear it, leaving 74 and 40 tokens. phase-3 and phase-4 remain amber. v16.24.0: total 60500 -> 60750 for visual evidence. Four phase docs gained one instruction each that cannot be inferred: Phase 0 keeps the issue's own images as the pre-fix evidence, Phase 3 captures the fixed state (there and not Phase 5, because every autopilot and --local entry drops Phase 5), Phase 5 hosts the flow recording when it runs, and Phase 6 blocks on a required artefact that is neither attached nor explained. Compression came first and twice, 42 tokens back, and the contract itself never entered this budget: the trigger matrix, the three video tiers, the size-degradation ladder and both render shapes live in multi-agent-refs/features/visual-evidence.md. phase-0-init cleared its own max without a bump. 250 was the smallest step that clears the aggregate. phase-3 and phase-4 remain amber. v17.0.0: phase-0-init max 13000 -> 13100 and total 62400 -> 62500 for the evidence-verdict writer. Phase 0 Step 7.7 probed with `--platform "$PLATFORM"`, a variable no phase document ever assigned, and it was gated on `visualEvidence.required`, which no phase document ever wrote - five readers, zero writers - so the step, Phase 3's capture and Phase 6's blocker were all unreachable and the pipeline reported nothing wrong. The step now writes the verdict and derives the platform from the stack, skipping the probe with a recorded reason when there is no device platform rather than passing the empty string the probe refuses with exit 2. Compression came first and three times, 30 tokens back from the Step 7.7 index rule that restated Step 7.5 verbatim and 55 from the new block itself; the reasoning never entered this budget, because who writes the verdict and how the platform is derived live in features/visual-evidence.md sections 1a and 1b. phase-3-dev max 9250 -> 9300 in the same change: it reads the platform back from state and re-decides the provisional verdict before capturing, which is the half of the fix that makes Phase 3 honest rather than merely reachable. Compressed three times first, 29 tokens back, by pointing its Phase-5 rationale and its tier mapping at visual-evidence.md sections 3 and 4.3 where both already live. 100, 50 and 150 were the smallest steps that clear it, leaving 24, 15 and 26 tokens. phase-3 and phase-4 remain amber. v17.1.0: total 62550 -> 62600 for the Figma/toolkit MCP distinction and web as an evidence platform. Three phase docs said "MCP forbidden" without qualifying it, while the gate that enforces it (smoke-no-mcp-in-dev-phases.sh) has always matched `figma` and nothing else - so the prose banned the screenshot, xcodebuild and UI-test tools that Phase 3 Steps 3.4 and 3.55 actually call, which is one way a run reaches Phase 7 with no evidence. Phase 0 gained one `web` arm in the platform derivation, now that run-ui-tests.sh has a web arm to derive it for. Compression came first and three times, taking the new prose from 175 tokens to 30: the reasoning moved to rules.md, whose own seven-row Figma phase matrix was deleted in the same pass because it duplicated rules/figma-pipeline.md "Phase access matrix" two paragraphs below this file's own instruction not to duplicate that rule file - 74 tokens back there, which is why phase-3-dev cleared its max without a bump (9298/9300). 50 was the smallest step that clears the aggregate, leaving 37 tokens. phase-3 and phase-4 remain amber. v17.1.0 (2): total 62600 -> 62700 for the plan reaching the task widget. Phase 2 computed tasks[], their order and their dependsOn[] edges, stored them, and used them to drive Phase 3's ready-task picker - and none of it was visible on the surface the user actually watches; the card had drawn sub-phases for releases, the widget never had. Phase 2 gains one call. Compression came first and twice: the tasks[]-to-todos[] jq blob left the phase doc for plan-todos.sh `set`, which now accepts a planning-output document directly (-28, and it removes a mapping two files defined, of which this was the untested copy), and the new step's own prose was cut from 116 tokens to 61. The parsing of the plan itself never entered this budget - it lives in phase-tracker.sh `plan`, which owns the sub-phase structure it writes. 100 was the smallest step that clears it, leaving 75 tokens. phase-3 and phase-4 remain amber. v17.5.0: phase-0-init max 13100 -> 13150 for the deferred widget registration, the dev-context step and two more exit-gate assertions. Compression came first and seven times, 492 tokens reclaimed before the number moved: the TLDR restated the step headings under it, the exit-gate list restated the script's own header comment, and the clarifier cost note, the branch-collision probe rationale, the fetch-fail host explanation, the multi-repo write-state race and the baseline unknown-vs-green warning all restate something already authoritative in agents/task-clarifier.md, phases/operations.md or phase0-exit-gate.mjs. The three additions are contract an agent cannot infer: that only Phase 0 is registered at Step -1 and the rest at 7.5 with `tiles --new`, that the dev-context picker runs between project and branch and writes siblings[] even when empty, and that a one-option AskUserQuestion is refused by the host along with every question batched with it. The reasoning for the last one never entered this budget - it lives in picker-contract.md "Two options or it is not a question". 50 was the smallest step that clears it, leaving 28 tokens; the aggregate needed no bump and sits at 62698/62700. phase-3 and phase-4 remain amber. v17.5.0 (2): phase-0-init max 13150 -> 13300 and total 62700 -> 62900 for the workspace question. Phase 0 gained Step 5b: where the branch lives was decided by a flag and by a command name, never by a question, so a user who saw a run improvise "Worktree / Lokal" once took it for a shipped picker and concluded /multi-agent:local was redundant - while it was the only way to reach local mode at all. Compression came first and three times, 156 tokens reclaimed before the number moved and the step itself cut from 426 tokens to 181: the ask-choice index rule in Step 7.5 restated modes.md "Pipeline depth" in full, and the question wording, the two options and what local costs now live in modes.md "Local Mode", outside this budget. What the phase doc carries is what an agent cannot infer - that it runs after Step 4 and before Step 6b and 7.5, who is exempt, and that autopilot resolves it to a worktree rather than skipping it, because an unattended commit in the user's own checkout is what worktrees exist to prevent. The state field is the other half: localMode alone cannot say whether anyone decided, since false is both a chosen worktree and one nothing asked about, and phase0-exit-gate.mjs now refuses to close Phase 0 without workspaceSource. 150 and 200 were the smallest steps that clear it, leaving 32 and 56 tokens. phase-3 and phase-4 remain amber. v17.5.0 (3): phase-0-init max 13300 -> 13400 and total 62900 -> 63000 for base-branch evidence. Step 3 asked one question with a list it could not vouch for: `git fetch origin` ran, its exit code was discarded, and `git branch -r` printed the remote-tracking cache either way, so a restricted network produced a weeks-old local list presented as the remote's answer. And the answer was usually derivable - an issue carrying a target version, or linking a separate issue that represents the release, already names the branch on a repo whose release branches encode the version - and nothing derived it. Rule 5 now captures the exit code and rule 5b hands the list plus the issue's version fields and links to base-branch-candidates.mjs. Compression came first and six times, 1,118 characters reclaimed across this step before the number moved: the one-row picker rule and the not-skippable rule both restated picker-contract.md, the autopilot resolution order and the whole rationale for admitting version-carrying branches moved into features/base-branch-evidence.md, and the fetch-fail persistence block was four sentences for one instruction. The feature's own reasoning never entered this budget at all - the evidence sources, the learned convention, the picker contract, the autopilot fencing and the state shape are ~1,900 tokens living in that ref. What the phase doc carries is the call and the two facts an agent cannot infer: that the fetch exit code decides whether the list may be called remote, and that the branch filter needs a version alternative because a word list that discards is the prefix table this replaces. 100 was the smallest step that clears both, leaving 51 and 61 tokens. phase-3 and phase-4 remain amber.
|
|
@@ -118,12 +118,18 @@ collapses tool output, so a card left in stdout never reaches the user. **Reprin
|
|
|
118
118
|
inside your reply at every phase boundary.** `phase-tracker.sh tiles` says so and prints it.
|
|
119
119
|
|
|
120
120
|
```bash
|
|
121
|
-
# Bootstrap
|
|
121
|
+
# Bootstrap at Phase 0 start. The full pipeline and --local register Phase 0 alone
|
|
122
|
+
# here and the rest at Step 7.5, once the depth answer names them; every other mode
|
|
123
|
+
# registers its whole set now. See tracker-contract.md "Deferred registration".
|
|
122
124
|
bash ~/.copilot/scripts/phase-tracker.sh init "$TASK_ID"
|
|
123
|
-
|
|
125
|
+
bash ~/.copilot/scripts/phase-tracker.sh add 0 Init
|
|
126
|
+
bash ~/.copilot/scripts/phase-tracker.sh tiles
|
|
127
|
+
|
|
128
|
+
# Step 7.5, after the depth answer - Full below, Short drops 1:Analysis and 2:Planning:
|
|
129
|
+
for p in 1:Analysis 2:Planning 3:Dev 4:Review 5:Test 6:Commit 7:Report; do
|
|
124
130
|
bash ~/.copilot/scripts/phase-tracker.sh add "${p%%:*}" "${p#*:}"
|
|
125
131
|
done
|
|
126
|
-
bash ~/.copilot/scripts/phase-tracker.sh tiles
|
|
132
|
+
bash ~/.copilot/scripts/phase-tracker.sh tiles --new
|
|
127
133
|
|
|
128
134
|
# At each phase boundary - update status. Tracker stamps started_at on first
|
|
129
135
|
# transition to in_progress, completed_at on terminal status (completed/failed/skipped).
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@mmerterden/multi-agent-pipeline",
|
|
3
|
-
"version": "17.
|
|
3
|
+
"version": "17.5.0",
|
|
4
4
|
"description": "8-phase AI development pipeline with full orchestration on Claude Code, Copilot CLI and Codex CLI. Analysis, planning, TDD, CLI-aware parallel review with consensus surfacing + Fable triage, default-FAIL evidence gates, secret + intent guards, per-phase cost ledger, persistent learnings memory, wiki generation, commit automation. Token-preserving uninstall.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"main": "index.js",
|
|
@@ -169,10 +169,10 @@ bash $HOME/.claude/scripts/phase-tracker.sh tokens <N> <in> <out> [cached]
|
|
|
169
169
|
|
|
170
170
|
In Claude Code the agent MUST also drive the native TaskList widget so the user sees a sticky phase tile stack - this is the only progress signal Claude Code surfaces. Skipping these calls is the #1 source of "I don't see any phases" complaints.
|
|
171
171
|
|
|
172
|
-
**TaskCreate ordering (strict)**: All TaskCreate calls fire in strict phase-number order BEFORE any TaskUpdate
|
|
172
|
+
**TaskCreate ordering (strict)**: All TaskCreate calls in a registration batch fire in strict phase-number order BEFORE any TaskUpdate in that batch, and a later batch only ever appends phases numbered above everything already registered. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks (e.g. `1 ✓ · 2 ✓ · 4 ✓ · 0 ▶ · 3 ☐`) even when the underlying state is correct. Pre-marking phases as completed/skipped before Phase 0 starts is FORBIDDEN - register the tile in order, then flip status via TaskUpdate when the phase actually short-circuits. Full contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
|
|
173
173
|
|
|
174
174
|
```text
|
|
175
|
-
#
|
|
175
|
+
# Register one tile per phase, capture the taskId, persist it:
|
|
176
176
|
for each phase in 0:Init, 1:Analysis, 2:Planning, 4:Review, 6:Commit, 7:Report:
|
|
177
177
|
TaskCreate({ subject: "Phase <N>: <Name>", activeForm: "<doing-form>" })
|
|
178
178
|
-> returns taskId
|
|
@@ -194,7 +194,7 @@ bash $HOME/.claude/scripts/phase-tracker.sh update <N> completed
|
|
|
194
194
|
|
|
195
195
|
#### TaskCreate ordering (strict)
|
|
196
196
|
|
|
197
|
-
**All TaskCreate calls fire in strict phase-number order BEFORE any TaskUpdate is applied.** For `analysis` that means: Phase 0 → Phase 1 → Phase 2 → Phase 4 → Phase 6 → Phase 7. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks. Full ordering contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
|
|
197
|
+
**All TaskCreate calls in a batch fire in strict phase-number order BEFORE any TaskUpdate is applied.** For `analysis` that means: Phase 0 → Phase 1 → Phase 2 → Phase 4 → Phase 6 → Phase 7. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks. Full ordering contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
|
|
198
198
|
|
|
199
199
|
### Visual channel - Copilot CLI / plain shell
|
|
200
200
|
|
|
@@ -83,10 +83,10 @@ bash $HOME/.claude/scripts/phase-tracker.sh tokens <N> <in> <out> [cached]
|
|
|
83
83
|
|
|
84
84
|
In Claude Code the agent MUST also drive the native TaskList widget so the user sees a sticky phase tile stack - this is the only progress signal Claude Code surfaces. Skipping these calls is the #1 source of "I don't see any phases" complaints.
|
|
85
85
|
|
|
86
|
-
**TaskCreate ordering (strict)**: All TaskCreate calls fire in strict phase-number order BEFORE any TaskUpdate
|
|
86
|
+
**TaskCreate ordering (strict)**: All TaskCreate calls in a registration batch fire in strict phase-number order BEFORE any TaskUpdate in that batch, and a later batch only ever appends phases numbered above everything already registered. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks (e.g. `1 ✓ · 2 ✓ · 4 ✓ · 0 ▶ · 3 ☐`) even when the underlying state is correct. Pre-marking phases as completed/skipped before Phase 0 starts is FORBIDDEN - register the tile in order, then flip status via TaskUpdate when the phase actually short-circuits. Full contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
|
|
87
87
|
|
|
88
88
|
```text
|
|
89
|
-
#
|
|
89
|
+
# Register one tile per phase, capture the taskId, persist it:
|
|
90
90
|
for each phase in 0:Init, 1:Analysis, 2:Planning, 3:Dev, 4:Review, 6:Commit, 7:Report:
|
|
91
91
|
TaskCreate({ subject: "Phase <N>: <Name>", activeForm: "<doing-form>" })
|
|
92
92
|
-> returns taskId
|
|
@@ -108,7 +108,7 @@ bash $HOME/.claude/scripts/phase-tracker.sh update <N> completed
|
|
|
108
108
|
|
|
109
109
|
#### TaskCreate ordering (strict)
|
|
110
110
|
|
|
111
|
-
**All TaskCreate calls fire in strict phase-number order BEFORE any TaskUpdate is applied.** For `autopilot` that means: Phase 0 → Phase 1 → Phase 2 → Phase 3 → Phase 4 → Phase 6 → Phase 7. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks. Full ordering contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
|
|
111
|
+
**All TaskCreate calls in a batch fire in strict phase-number order BEFORE any TaskUpdate is applied.** For `autopilot` that means: Phase 0 → Phase 1 → Phase 2 → Phase 3 → Phase 4 → Phase 6 → Phase 7. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks. Full ordering contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
|
|
112
112
|
|
|
113
113
|
### Visual channel - Copilot CLI / plain shell
|
|
114
114
|
|