@mmerterden/multi-agent-pipeline 16.20.0 → 16.23.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (51) hide show
  1. package/CHANGELOG.md +231 -102
  2. package/README.md +6 -8
  3. package/README.tr.md +6 -8
  4. package/docs/architecture.md +3 -3
  5. package/docs/ecosystem.md +5 -5
  6. package/docs/features.md +1 -0
  7. package/install/templates/claude-hooks.json +12 -1
  8. package/install/templates/copilot-instructions.md +17 -2
  9. package/package.json +1 -1
  10. package/pipeline/agents/bulk-reader.md +57 -0
  11. package/pipeline/commands/multi-agent/SKILL.md +0 -5
  12. package/pipeline/commands/multi-agent/help/SKILL.md +0 -10
  13. package/pipeline/commands/multi-agent/ios-coding-standard/SKILL.md +2 -2
  14. package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +1 -1
  15. package/pipeline/commands/multi-agent/resume-local/SKILL.md +2 -2
  16. package/pipeline/commands/multi-agent/setup/SKILL.md +7 -5
  17. package/pipeline/commands/multi-agent/sync/SKILL.md +13 -13
  18. package/pipeline/multi-agent-refs/cross-cli-contract.md +10 -12
  19. package/pipeline/multi-agent-refs/phases/modes.md +1 -1
  20. package/pipeline/multi-agent-refs/phases/phase-0-init.md +7 -2
  21. package/pipeline/multi-agent-refs/phases/phase-3-dev.md +1 -1
  22. package/pipeline/multi-agent-refs/phases/phase-7-report.md +8 -1
  23. package/pipeline/multi-agent-refs/picker-contract.md +1 -1
  24. package/pipeline/multi-agent-refs/tracker-contract.md +46 -0
  25. package/pipeline/schemas/agent-state.schema.json +1 -1
  26. package/pipeline/schemas/bulk-read-output.schema.json +52 -0
  27. package/pipeline/schemas/prefs.schema.json +74 -19
  28. package/pipeline/schemas/token-budget.json +3 -3
  29. package/pipeline/scripts/bulk-read.sh +277 -0
  30. package/pipeline/scripts/check-read-size.py +335 -0
  31. package/pipeline/scripts/check-read-size.sh +86 -0
  32. package/pipeline/scripts/phase-tracker.sh +245 -3
  33. package/pipeline/scripts/pre-commit-check.sh +1 -0
  34. package/pipeline/scripts/uninstall.mjs +1 -0
  35. package/pipeline/skills/.skill-manifest.json +8 -24
  36. package/pipeline/skills/.skills-index.json +2 -46
  37. package/pipeline/skills/shared/README.md +4 -8
  38. package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +0 -8
  39. package/pipeline/skills/shared/core/multi-agent-ios-coding-standard/SKILL.md +1 -1
  40. package/pipeline/skills/shared/core/multi-agent-resume-local/SKILL.md +1 -1
  41. package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +12 -12
  42. package/pipeline/skills/shared/external/backlog/SKILL.md +10 -6
  43. package/pipeline/skills/skills-index.md +2 -6
  44. package/pipeline/commands/multi-agent/dev/SKILL.md +0 -17
  45. package/pipeline/commands/multi-agent/dev-autopilot/SKILL.md +0 -23
  46. package/pipeline/commands/multi-agent/dev-local/SKILL.md +0 -17
  47. package/pipeline/commands/multi-agent/dev-local-autopilot/SKILL.md +0 -21
  48. package/pipeline/skills/shared/core/multi-agent-dev/SKILL.md +0 -19
  49. package/pipeline/skills/shared/core/multi-agent-dev-autopilot/SKILL.md +0 -25
  50. package/pipeline/skills/shared/core/multi-agent-dev-local/SKILL.md +0 -19
  51. package/pipeline/skills/shared/core/multi-agent-dev-local-autopilot/SKILL.md +0 -23
@@ -103,6 +103,52 @@ For each phase in the current mode, IN PHASE-NUMBER ORDER (0 → 1 → 2 → ...
103
103
 
104
104
  The `tasklist_id` meta field is persisted so that `:resume` can rebuild the TaskList from the state file in a fresh session.
105
105
 
106
+ #### The card is not the widget
107
+
108
+ `phase-tracker.sh render` prints a bordered ANSI card to stdout, and every
109
+ supported host collapses tool output into a one-line summary (`Ran 8 shell
110
+ commands`). A run whose only progress signal is that card is a run the user
111
+ watches in silence - which is exactly what happened: the card was correct on
112
+ disk and invisible on screen.
113
+
114
+ So each host has a widget, and the tracker names the calls rather than assuming
115
+ the agent remembers them:
116
+
117
+ | Host | Widget | How it is registered |
118
+ |---|---|---|
119
+ | Claude Code | native TaskList tiles | `TaskCreate` per phase, then `TaskUpdate` per boundary |
120
+ | Codex CLI | native plan steps | one `update_plan` carrying the FULL step list, resent per boundary |
121
+ | Copilot CLI | none exists | the card itself, reprinted inside the agent's reply |
122
+
123
+ `phase-tracker.sh tiles` prints the registration calls for the host it runs on,
124
+ read from the phase set already registered by `add`. Every `update` then prints
125
+ a `-- NEXT (required) --` block with that host's mirror call and, on completion,
126
+ the narration line built from what was actually recorded. Both are printed to be
127
+ acted on, not read.
128
+
129
+ #### Accounting is a gate, not a request
130
+
131
+ `update <N> completed` exits **3** for phases 1-4 when nothing was recorded via
132
+ `tokens`. The alternative is what shipped for months: the docs asked for the
133
+ call, `smoke-tracker-tokens-invocation.sh` linted that the docs *mention* the
134
+ call, and no run made it - every phase priced at `-`, every completion line
135
+ printed without numbers, and every gate green. A lint over prose cannot observe
136
+ a runtime omission; the tool that would have received the call can.
137
+
138
+ The refusal is recoverable and loses no work:
139
+
140
+ ```bash
141
+ phase-tracker.sh model <N> <model_name>
142
+ phase-tracker.sh tokens <N> <in> <out> [cached]
143
+ phase-tracker.sh update <N> completed
144
+ ```
145
+
146
+ `model` resolves a full model id to its cost-table family (`claude-opus-5` →
147
+ `opus`) and says so out loud when it cannot, because an unpriced name is the
148
+ other way a phase silently costs nothing. A phase that truly ran no LLM call
149
+ completes with `--no-llm`, which is recorded and keeps it out of the run
150
+ report's cost-unavailable list.
151
+
106
152
  #### TaskCreate ordering (strict)
107
153
 
108
154
  **The native TaskList widget renders tiles in TaskCreate creation order, NOT by phase-number metadata.** Therefore:
@@ -244,7 +244,7 @@
244
244
  "onlyDevelop": {
245
245
  "type": "boolean",
246
246
  "default": false,
247
- "description": "Short pipeline - phases 1 and 2 are skipped. Set by the Phase 0 Step 7.5 depth picker as of v16.0.0 (it was the --dev flag before); the key and every reader of it are unchanged. Phase 4 Review still runs (v14.0.0+). Always false in an autopilot run, which never asks the depth question."
247
+ "description": "Short pipeline - phases 1 and 2 are skipped. Set by the Phase 0 Step 7.5 depth picker as of v16.0.0; the key and every reader of it are unchanged. Phase 4 Review still runs (v14.0.0+). Always false in an autopilot run, which never asks the depth question."
248
248
  },
249
249
  "localMode": {
250
250
  "type": "boolean",
@@ -0,0 +1,52 @@
1
+ {
2
+ "$schema": "http://json-schema.org/draft-07/schema#",
3
+ "$id": "https://example.com/pipeline/bulk-read-output.schema.json",
4
+ "title": "Bulk Read Output",
5
+ "description": "Schema for the bulk-reader worker's answer. Source: pipeline/agents/bulk-reader.md, invoked by pipeline/scripts/bulk-read.sh when check-read-size.sh blocks a whole-file read. Every positional field is a LINE NUMBER in the file as given to the worker, because a summary the caller cannot open is a summary the caller will pay to replace.",
6
+ "type": "object",
7
+ "additionalProperties": false,
8
+ "required": ["summary"],
9
+ "properties": {
10
+ "answer": {
11
+ "type": "string",
12
+ "description": "Direct answer to the question the caller asked, or a plain statement that this file cannot answer it. Never a guess - an unanswerable question is reported, not filled in."
13
+ },
14
+ "summary": {
15
+ "type": "string",
16
+ "description": "What the file is and does, 3-6 sentences. Descriptive only: no review, no judgement, no proposed change."
17
+ },
18
+ "symbols": {
19
+ "type": "array",
20
+ "description": "Top-level declarations the file contains, in file order.",
21
+ "items": {
22
+ "type": "object",
23
+ "additionalProperties": false,
24
+ "required": ["name", "line"],
25
+ "properties": {
26
+ "name": { "type": "string" },
27
+ "kind": { "type": "string", "enum": ["type", "func", "var", "extension", "other"] },
28
+ "line": { "type": "integer", "minimum": 1 }
29
+ }
30
+ }
31
+ },
32
+ "regions": {
33
+ "type": "array",
34
+ "maxItems": 8,
35
+ "description": "Spans worth opening next, most important first. The intended use is a bounded Read(offset:limit:) over one of these, which is both cheaper than the whole file and exact - and which the read gate allows through.",
36
+ "items": {
37
+ "type": "object",
38
+ "additionalProperties": false,
39
+ "required": ["start", "end"],
40
+ "properties": {
41
+ "why": { "type": "string", "description": "What a reader finds in this span." },
42
+ "start": { "type": "integer", "minimum": 1 },
43
+ "end": { "type": "integer", "minimum": 1 }
44
+ }
45
+ }
46
+ },
47
+ "truncated": {
48
+ "type": "boolean",
49
+ "description": "True when the worker did not see the whole file. A partial read summarized as a whole one is the failure this flag exists to make visible."
50
+ }
51
+ }
52
+ }
@@ -365,7 +365,7 @@
365
365
  },
366
366
  "multiRepoIntegrationHosts": {
367
367
  "type": "array",
368
- "description": "v5.6.0+. Learn-once registry of host projects that build together with a multi-repo combo (codegen producer → consumer + host integration project). Pipeline checks this registry in Phase 6 before commit; on match, auto-runs the host build. On miss for a ≥2-repo task, prompts the user once and persists the answer. See refs/multi-repo-integration-build.md for the full contract.",
368
+ "description": "v5.6.0+. Learn-once registry of host projects that build together with a multi-repo combo (codegen producer \u2192 consumer + host integration project). Pipeline checks this registry in Phase 6 before commit; on match, auto-runs the host build. On miss for a \u22652-repo task, prompts the user once and persists the answer. See refs/multi-repo-integration-build.md for the full contract.",
369
369
  "items": {
370
370
  "type": "object",
371
371
  "additionalProperties": false,
@@ -653,12 +653,12 @@
653
653
  "reportChannels": {
654
654
  "type": "object",
655
655
  "additionalProperties": false,
656
- "description": "v5.7+ - Phase 7 / /multi-agent:channels kanal seçimi default'ları. Multi-select menüde tick'li gelecek kanallar. Her kanal bağımsız boolean. Autopilot Phase 7'de ALWAYS pauses (30-min timeout) - bu değerler sadece menünün önceden seçili halini belirler.",
656
+ "description": "v5.7+ - Phase 7 / /multi-agent:channels kanal se\u00e7imi default'lar\u0131. Multi-select men\u00fcde tick'li gelecek kanallar. Her kanal ba\u011f\u0131ms\u0131z boolean. Autopilot Phase 7'de ALWAYS pauses (30-min timeout) - bu de\u011ferler sadece men\u00fcn\u00fcn \u00f6nceden se\u00e7ili halini belirler.",
657
657
  "properties": {
658
658
  "pr": {
659
659
  "type": "boolean",
660
660
  "default": true,
661
- "description": "PR description update (replace/append). Default ON - en yaygın kanal."
661
+ "description": "PR description update (replace/append). Default ON - en yayg\u0131n kanal."
662
662
  },
663
663
  "jira": {
664
664
  "type": "boolean",
@@ -668,12 +668,12 @@
668
668
  "confluence": {
669
669
  "type": "boolean",
670
670
  "default": false,
671
- "description": "Confluence page creation. Default OFF - bir kez parent page seçince LRU'dan öner."
671
+ "description": "Confluence page creation. Default OFF - bir kez parent page se\u00e7ince LRU'dan \u00f6ner."
672
672
  },
673
673
  "wiki": {
674
674
  "type": "boolean",
675
675
  "default": false,
676
- "description": "Component wiki pages (Case A scope multi-select). Default OFF - taskType=component + figmaConfig.wiki.enabled gerekli, yoksa menüde greyed out."
676
+ "description": "Component wiki pages (Case A scope multi-select). Default OFF - taskType=component + figmaConfig.wiki.enabled gerekli, yoksa men\u00fcde greyed out."
677
677
  }
678
678
  },
679
679
  "default": {
@@ -686,32 +686,32 @@
686
686
  "reportContent": {
687
687
  "type": "object",
688
688
  "additionalProperties": false,
689
- "description": "v5.7+ - Phase 7 / /multi-agent:channels içerik seçimi default'ları. Multi-select menüde tick'li gelecek content source'ları.",
689
+ "description": "v5.7+ - Phase 7 / /multi-agent:channels i\u00e7erik se\u00e7imi default'lar\u0131. Multi-select men\u00fcde tick'li gelecek content source'lar\u0131.",
690
690
  "properties": {
691
691
  "normalAnalysis": {
692
692
  "type": "boolean",
693
693
  "default": true,
694
- "description": "Phase 1+2+4 pipeline log'undan impact summary + risks + architectural decisions (yüksek seviye). Greyed out post-hoc çağrıda pipeline log yoksa."
694
+ "description": "Phase 1+2+4 pipeline log'undan impact summary + risks + architectural decisions (y\u00fcksek seviye). Greyed out post-hoc \u00e7a\u011fr\u0131da pipeline log yoksa."
695
695
  },
696
696
  "technicalAnalysis": {
697
697
  "type": "boolean",
698
698
  "default": false,
699
- "description": "Changes (değişen dosyalar gruplanıp ne/neden), Architecture (structural decisions), Dependencies (yeni import/framework/paket). PR body'deki 'Technical Details' bölümünün özeti; user'ın kanal seçimi PR içermediği durumlarda (ör. sadece Jira/Confluence) teknik içerik aktarmak istiyorsa devreye girer. Source: Phase 2 planning + Phase 3 dev log + PR diff stat."
699
+ "description": "Changes (de\u011fi\u015fen dosyalar gruplan\u0131p ne/neden), Architecture (structural decisions), Dependencies (yeni import/framework/paket). PR body'deki 'Technical Details' b\u00f6l\u00fcm\u00fcn\u00fcn \u00f6zeti; user'\u0131n kanal se\u00e7imi PR i\u00e7ermedi\u011fi durumlarda (\u00f6r. sadece Jira/Confluence) teknik i\u00e7erik aktarmak istiyorsa devreye girer. Source: Phase 2 planning + Phase 3 dev log + PR diff stat."
700
700
  },
701
701
  "testScenarios": {
702
702
  "type": "boolean",
703
703
  "default": true,
704
- "description": "Precondition / steps / expected tablosu (4-8 satır, user perspective). Pipeline log source."
704
+ "description": "Precondition / steps / expected tablosu (4-8 sat\u0131r, user perspective). Pipeline log source."
705
705
  },
706
706
  "autoDiff": {
707
707
  "type": "boolean",
708
708
  "default": false,
709
- "description": "PR diff'ten auto-generate özet (eski enrich behavior - root cause / solution / changed files / test scenarios). PR linked değilse greyed out."
709
+ "description": "PR diff'ten auto-generate \u00f6zet (eski enrich behavior - root cause / solution / changed files / test scenarios). PR linked de\u011filse greyed out."
710
710
  },
711
711
  "manualNote": {
712
712
  "type": "boolean",
713
713
  "default": false,
714
- "description": "Serbest metin paragraf (--message / --message-file). Her durumda seçilebilir."
714
+ "description": "Serbest metin paragraf (--message / --message-file). Her durumda se\u00e7ilebilir."
715
715
  },
716
716
  "costSummary": {
717
717
  "type": "boolean",
@@ -721,7 +721,7 @@
721
721
  "workSummary": {
722
722
  "type": "boolean",
723
723
  "default": false,
724
- "description": "v7.1.0+ - Executive 'Work Done' summary block. Distills the whole pipeline run into a single-screen section: task + branch + base + PR number, scope delivered (✅/⏳ per Phase 2 task), changed files with +/- counts (capped at 20 rows), review outcome (accepted/deferred/rejected counts + approved flag), and a one-line phase tick strip (0 Init ✅ · 1 Analysis ✅ · ...). Source: `agent-state.json` + `phase-tracker.json` + `git diff --numstat` between `baseBranch`...HEAD. Consumed by `render-work-summary.sh`. Greyed out if no state file exists for the task. Opt-in - off by default so baseline PR body stays unchanged."
724
+ "description": "v7.1.0+ - Executive 'Work Done' summary block. Distills the whole pipeline run into a single-screen section: task + branch + base + PR number, scope delivered (\u2705/\u23f3 per Phase 2 task), changed files with +/- counts (capped at 20 rows), review outcome (accepted/deferred/rejected counts + approved flag), and a one-line phase tick strip (0 Init \u2705 \u00b7 1 Analysis \u2705 \u00b7 ...). Source: `agent-state.json` + `phase-tracker.json` + `git diff --numstat` between `baseBranch`...HEAD. Consumed by `render-work-summary.sh`. Greyed out if no state file exists for the task. Opt-in - off by default so baseline PR body stays unchanged."
725
725
  }
726
726
  },
727
727
  "default": {
@@ -739,7 +739,7 @@
739
739
  "default": 1800,
740
740
  "minimum": 60,
741
741
  "maximum": 7200,
742
- "description": "v5.7+ - Phase 7'de autopilot always-pause menüsünde kullanıcı cevap vermezse session'ı sonlandırma süresi (saniye). Default 1800 (30 dk). Timeout'ta external delivery aborted, internal capture (agent-log, telemetry, knowledge) yine çalışır, session /multi-agent:resume ile devam ettirilebilir."
742
+ "description": "v5.7+ - Phase 7'de autopilot always-pause men\u00fcs\u00fcnde kullan\u0131c\u0131 cevap vermezse session'\u0131 sonland\u0131rma s\u00fcresi (saniye). Default 1800 (30 dk). Timeout'ta external delivery aborted, internal capture (agent-log, telemetry, knowledge) yine \u00e7al\u0131\u015f\u0131r, session /multi-agent:resume ile devam ettirilebilir."
743
743
  },
744
744
  "wikiScope": {
745
745
  "type": "array",
@@ -748,7 +748,7 @@
748
748
  "enum": ["main", "ios", "screenshots", "index"]
749
749
  },
750
750
  "default": ["main", "ios", "screenshots", "index"],
751
- "description": "v5.7+ - Wiki Case A scope multi-select default'u. Component wiki dispatch'inde hangi artifact'lar yazılacak: main (ana component sayfası), ios (iOS sub-page), screenshots (assets/ klasörü), index (_Sidebar.md + ComponentImplementationStatus.md). Legacy wikiDefault=true → [main,ios,screenshots,index] migration; wikiDefault=false → [] (empty array = Wiki adapter Case B menüsüne düşer)."
751
+ "description": "v5.7+ - Wiki Case A scope multi-select default'u. Component wiki dispatch'inde hangi artifact'lar yaz\u0131lacak: main (ana component sayfas\u0131), ios (iOS sub-page), screenshots (assets/ klas\u00f6r\u00fc), index (_Sidebar.md + ComponentImplementationStatus.md). Legacy wikiDefault=true \u2192 [main,ios,screenshots,index] migration; wikiDefault=false \u2192 [] (empty array = Wiki adapter Case B men\u00fcs\u00fcne d\u00fc\u015fer)."
752
752
  },
753
753
  "autoJiraFromGithubIssue": {
754
754
  "type": "string",
@@ -831,7 +831,7 @@
831
831
  "reviewDisagreementRound": {
832
832
  "type": "boolean",
833
833
  "default": false,
834
- "description": "v6.1.0+ - Phase 4 Step 2.5 rebuttal round. When reviewers disagree (mixed blocker/approved verdict), each reviewer is re-prompted with the others' opposing arguments for one additional round before triage. Lifts signal quality on ambiguous findings at ~1× Step 2 token cost. Off by default - flip for security-critical or release-branch reviews."
834
+ "description": "v6.1.0+ - Phase 4 Step 2.5 rebuttal round. When reviewers disagree (mixed blocker/approved verdict), each reviewer is re-prompted with the others' opposing arguments for one additional round before triage. Lifts signal quality on ambiguous findings at ~1\u00d7 Step 2 token cost. Off by default - flip for security-critical or release-branch reviews."
835
835
  },
836
836
  "analysisProfiles": {
837
837
  "type": "array",
@@ -1047,7 +1047,7 @@
1047
1047
  "minimum": 0,
1048
1048
  "maximum": 10,
1049
1049
  "default": 6,
1050
- "description": "Clarity threshold. Score ≥ threshold → proceed silently. Below → questions fire. 6 is the borderline 'the what is clear but the how is fuzzy' line."
1050
+ "description": "Clarity threshold. Score \u2265 threshold \u2192 proceed silently. Below \u2192 questions fire. 6 is the borderline 'the what is clear but the how is fuzzy' line."
1051
1051
  },
1052
1052
  "maxQuestions": {
1053
1053
  "type": "integer",
@@ -1215,7 +1215,7 @@
1215
1215
  "minimum": 30,
1216
1216
  "maximum": 86400,
1217
1217
  "default": 300,
1218
- "description": "Polling interval for --watch loop. Clamped to ≥30s to stay polite with GitHub rate limits."
1218
+ "description": "Polling interval for --watch loop. Clamped to \u226530s to stay polite with GitHub rate limits."
1219
1219
  },
1220
1220
  "labelFilter": {
1221
1221
  "type": "string",
@@ -1291,7 +1291,7 @@
1291
1291
  "devCritic": {
1292
1292
  "type": "object",
1293
1293
  "additionalProperties": false,
1294
- "description": "v8.6+ - Phase 3.5 evaluator-optimizer. After the Dev generator's last edit and BEFORE Phase 4 reviewers, dispatch agents/dev-critic.md (Sonnet by default) to run deterministic gates (build/lint/test/secrets) + the platform checklist (rules/*.md). Max 2 critic iterations, then escalate. Catches gate failures and checklist violations that would otherwise burn 2-3 Phase 4 reviewer calls + Opus triage. Off by default - introduces 1× Sonnet call per Dev iteration; flip on for feature work, security-touching paths, or multi-file refactors. Source: Anthropic 'Building Effective Agents' (Dec 2024) evaluator-optimizer pattern.",
1294
+ "description": "v8.6+ - Phase 3.5 evaluator-optimizer. After the Dev generator's last edit and BEFORE Phase 4 reviewers, dispatch agents/dev-critic.md (Sonnet by default) to run deterministic gates (build/lint/test/secrets) + the platform checklist (rules/*.md). Max 2 critic iterations, then escalate. Catches gate failures and checklist violations that would otherwise burn 2-3 Phase 4 reviewer calls + Opus triage. Off by default - introduces 1\u00d7 Sonnet call per Dev iteration; flip on for feature work, security-touching paths, or multi-file refactors. Source: Anthropic 'Building Effective Agents' (Dec 2024) evaluator-optimizer pattern.",
1295
1295
  "properties": {
1296
1296
  "enabled": {
1297
1297
  "type": "boolean",
@@ -1464,6 +1464,61 @@
1464
1464
  "tailLines": 20
1465
1465
  }
1466
1466
  },
1467
+ "bulkRead": {
1468
+ "type": "object",
1469
+ "additionalProperties": false,
1470
+ "description": "v16.21+ - Route a whole-file read that is too big to be worth the caller's rung to a cheap worker instead, and put its line-numbered summary in context rather than the file. The gate is pipeline/scripts/check-read-size.sh (PreToolUse, installed from install/templates/claude-hooks.json); the worker is pipeline/scripts/bulk-read.sh over the bulk-reader persona. Companion to contextOffload, which does the same for tool OUTPUT rather than for reads.",
1471
+ "properties": {
1472
+ "mode": {
1473
+ "type": "string",
1474
+ "enum": ["off", "observe", "enforce"],
1475
+ "default": "off",
1476
+ "description": "off = the hook is a no-op. observe = decide and log every inspected read, block nothing; this is the BASELINE MEASUREMENT, and running it before enforce is what makes a later saving claim checkable. enforce = block a whole-file read over minLines and tell the model to delegate it."
1477
+ },
1478
+ "minLines": {
1479
+ "type": "integer",
1480
+ "minimum": 1,
1481
+ "maximum": 100000,
1482
+ "default": 350,
1483
+ "description": "Files shorter than this are never inspected. A delegated read costs a round trip measured in tens of seconds, so a small file is cheaper read directly - the threshold is where that trade turns over."
1484
+ },
1485
+ "exemptPhases": {
1486
+ "type": "array",
1487
+ "items": {
1488
+ "type": "string"
1489
+ },
1490
+ "default": ["3"],
1491
+ "description": "Phases the gate never blocks in. Phase 3 is the shipped default and removing it breaks development: Claude Code's Edit requires the same file to have been Read first, so a gate that blocks reads while code is being changed blocks the change. The gate is for phases that read to UNDERSTAND."
1492
+ },
1493
+ "model": {
1494
+ "type": "string",
1495
+ "default": "haiku",
1496
+ "description": "Rung the delegated read runs on. The saving IS the gap between this rung and the caller's, so a worker raised to the caller's own rung saves nothing."
1497
+ },
1498
+ "timeoutSeconds": {
1499
+ "type": "integer",
1500
+ "minimum": 5,
1501
+ "maximum": 600,
1502
+ "default": 60,
1503
+ "description": "Wall-clock budget for one delegated read. On expiry the worker degrades and the caller is told to do a bounded read instead - never left waiting, and never handed an invented summary."
1504
+ },
1505
+ "maxBytes": {
1506
+ "type": "integer",
1507
+ "minimum": 1024,
1508
+ "maximum": 20971520,
1509
+ "default": 1048576,
1510
+ "description": "Ceiling past which a delegated read degrades instead of running. Delegation is not free: at some size the worker's own input bill approaches the read it replaced, and a file large enough to strain its window comes back truncated - a partial summary presented as a whole one is the failure mode this feature must never produce. Over the ceiling the caller is told to narrow first (grep, then a bounded read) rather than handed an expensive round trip to a worse answer."
1511
+ }
1512
+ },
1513
+ "default": {
1514
+ "mode": "off",
1515
+ "minLines": 350,
1516
+ "exemptPhases": ["3"],
1517
+ "model": "haiku",
1518
+ "timeoutSeconds": 60,
1519
+ "maxBytes": 1048576
1520
+ }
1521
+ },
1467
1522
  "testGap": {
1468
1523
  "type": "object",
1469
1524
  "additionalProperties": false,
@@ -1506,7 +1561,7 @@
1506
1561
  "autopilotSafetyGate": {
1507
1562
  "type": "boolean",
1508
1563
  "default": true,
1509
- "description": "v7.0.0+ - Phase 2 autopilot safety classifier. Before autopilot mode consumes the user's approval skip, run `classify-plan-safety.mjs` over the approved plan. If the heuristic score ≥ 50 (e.g. >15 files touched, or security-path touch, or delete-without-test, or schema migration) inject a ONE-TIME pause asking for explicit manual approval - even in autopilot. Default ON because the risk of skipping this gate is asymmetric: a pause on a high-blast-radius plan costs seconds; a silent auto-merge of a bad one costs hours of rollback. Flip to `false` only for tightly-scoped autopilot workflows (e.g. batch figma component iteration) where the task class is known-safe."
1564
+ "description": "v7.0.0+ - Phase 2 autopilot safety classifier. Before autopilot mode consumes the user's approval skip, run `classify-plan-safety.mjs` over the approved plan. If the heuristic score \u2265 50 (e.g. >15 files touched, or security-path touch, or delete-without-test, or schema migration) inject a ONE-TIME pause asking for explicit manual approval - even in autopilot. Default ON because the risk of skipping this gate is asymmetric: a pause on a high-blast-radius plan costs seconds; a silent auto-merge of a bad one costs hours of rollback. Flip to `false` only for tightly-scoped autopilot workflows (e.g. batch figma component iteration) where the task class is known-safe."
1510
1565
  },
1511
1566
  "dynamicSkillLoading": {
1512
1567
  "type": "boolean",
@@ -4,7 +4,7 @@
4
4
  "description": "Per-phase token budget for lazy-loaded pipeline docs. Enforced by smoke-token-budget.sh.",
5
5
  "phases": {
6
6
  "phase-0-init": {
7
- "max_tokens": 12400,
7
+ "max_tokens": 12500,
8
8
  "warn_tokens": 11900
9
9
  },
10
10
  "phase-1-analysis": {
@@ -36,6 +36,6 @@
36
36
  "warn_tokens": 5600
37
37
  }
38
38
  },
39
- "total_max_tokens": 60250,
40
- "note": "Token estimate = ceil(chars / 4). Per-phase budget rule: warn = current+10% (rounded to nearest 50), max = current+25%. Gives ~6 edit cycles of headroom before warn trips - intentionally quiet under normal maintenance, loud when a phase grows unusually. Only the active phase is loaded (lazy). Recalibrated at v10.0.0 after the validator/consistency/simplifier/lesson gate contracts landed in phases 1-4. Recalibrated again at v10.9.0 after the verify-by-test (Phase 4 Step 3.7), update-check (Phase 0 Step 0.6), immutable-test (Phase 3 GREEN) and redTests re-entry contracts landed - Step 3.7 prose was compressed to a pointer into refs/features/verify-by-test.md before the recalibration. Total bumped 50000 -> 51000 at v12.5.0 after the worktree residue/traversal-prune contract (Phase 0 + Phase 5 heal) and the Reflexion causal-diagnosis contract (Phase 4 lesson memory) landed; the prose was compressed first (161 tokens reclaimed) and every per-phase max still passes - only the aggregate needed room. Recalibrated again at v13.6.0 after the install-relative path correction: an instruction that names `pipeline/scripts/x` resolves only from a repo checkout, and a run happens in the user's worktree, so 157 references across these docs moved to `$HOME/.claude/...` at +5 bytes each - 196 tokens of pure correctness cost. Same discipline as before: prose was compressed FIRST (149 tokens reclaimed, by pointing Phase 1's Figma tier table at the Phase 0 probe that already resolved it and Phase 4's Codex constraints at the always-loaded AGENTS.md block), and only then were the budgets moved. Five warn lines had been permanently amber, which makes the amber tier useless as a signal, so every warn was reset to the documented current+10% and the four maxes that the new warn would have collided with were reset to current+25%. Aggregate 51000 -> 51500. Total bumped 51500 -> 52200 at v14.0.0 after Phase 4 Review entered the four --dev mode phase sets and the criteria-resolution contract (Step 1.78) landed. Same discipline as every prior bump: prose was compressed FIRST, 820 tokens reclaimed, before the number moved. Two of those compressions are structural rather than cosmetic - the hardcoded SwiftUI interaction list in Step 1.5 and the SwiftUI convention paragraph in Step 2.8 were transcriptions of rules that now live in a scoped registry, so keeping them here would have re-created the drift this release exists to remove, and the third moved the Step 1.78 full contract into refs/features/skill-conformance.md leaving a pointer. What remains is contract text that cannot be inferred: the manifest's four consumer-visible parts, the conformance checklist the reviewers must return, and the fail-closed semantics. Every per-phase max still passes (phase-4 12405/14750); only the aggregate needed room. Total bumped 52200 -> 52700 at v14.1.0 after two more contracts landed: stack skill routing (Phase 3 pre-flight step 9) and worktree finalize (Phase 6 step 9). Compression came first, as always, and twice: 224 tokens out of Phase 3 by pointing its criteria-ledger and routing steps at their feature files instead of restating them, and 190 out of Phase 6 by moving the finalize contract into refs/features/worktree-finalize.md and leaving the invocation plus the exit-3 semantics. Both new contracts follow the pattern the earlier ones set: the phase doc carries the call and the decision, the feature file carries the reasoning, and the feature files are outside this budget because it loops only the eight phase-N-* keys. Every per-phase max still passes (phase-3 7677/8950, phase-6 5223/6150 and both under warn); only the aggregate needed room. Total bumped 52700 -> 52750 for the Phase 0 Step 3 branch-persistence correction: the step wrote the legacy `projects[].branches` while the TTL filter two sections below read `global.recentBranches`, and both spots named a `{name, lastUsed}` shape the schema rejects (`branch` required, `additionalProperties: false`), so the recent-branch picker option could never populate and a literal implementation would have failed prefs validation. Naming the right target, the right key and the legacy field to avoid costs 41 tokens over the one line it replaces. Compression came first and was applied three times to the replacement text itself, from 120 tokens down to 66, by moving the rationale out of the phase doc entirely: the reasoning now lives where it is enforced, in the migrate-prefs carry-forward comment and the smoke-pref-migration f7 block, leaving the phase doc with only the instruction. 50 was the smallest step that clears it; phase-0-init sits at 10893/12400, far under its own max, so this is purely an aggregate ceiling. v15.0.0: total 52750 -> 53100, the stack-skill tables in phase-1/2/4 now carry plugin-namespaced names (ai-<stack>-toolkit:<skill>) - functional prefixes, ~170 tokens. v15.10.0: total 53350 -> 53950 for the memory-recall + context-offload contracts (Phase 1 two-block durable-knowledge injection and its telemetry, Phase 3 build-log offload pipe, Phase 4 ranked prior art, offload pipe and recall telemetry). Compression came first and twice, taking the new prose from 1168 tokens to 580: the reasoning behind the two blocks lives in multi-agent-refs/prompt-assembly.md and the reasoning behind the offload filter lives in the offload-ref.sh header, both outside this budget, so the phase docs carry only the call, the pref that gates it and the one fact an agent cannot infer - that the evidence gate still reads the whole build log, so offloading changes what is read, never what counts as a verified pass. Every per-phase max still passes (phase-3 7985/8950, phase-4 12997/14750); phase-3 and phase-4 crossed their warn lines and are left amber on purpose, because that is the signal that those two docs are the next ones needing structural compression rather than another bump. v15.13.0: total 53950 -> 54050 for the prefs-to-flag bridges. Five settings had shipped declared-but-inert: contextOffload.minLines and .tailLines (fixed in 15.11.0), learningsLedger.maxBriefEntries, and testGap.scanTree and .promoteSeverity - the last two declared in the schema AND implemented as flags in the scanner, with nothing in between reading the pref and passing the flag. Wiring three of them costs the phase docs 94 tokens, which is the wiring itself and not prose: two `--max` substitutions and a three-line GAP_FLAGS block. Compression came first and twice, as always: the rationale that would have sat in phase-5 now lives in the header of smoke-prefs-consumed.sh, the gate that makes this class fail a build instead of shipping, and a `--severity-promote` table row was dropped because the invocation above it now shows the flag and names the pref that triggers it, which the row did not. 100 was the smallest step that clears it. Every per-phase max still passes; phase-3 and phase-4 remain amber on purpose. v15.14.0: total 54050 -> 54400 for the supported-version gate. Phase 0 Step 0.6 stopped being purely advisory: a release can now publish an npm dist-tag `required` that names the oldest runnable version, and below it the run halts instead of nagging. What the phase doc has to carry is the part an agent cannot infer - the third stdout field, that the halt is identical in autopilot, and that the run must NOT continue on the freshly updated install because its docs were already loaded from the old version. Compression came first, as always, and took the new prose from 469 tokens to 337: the rationale for the floor, the exemption list, the fail-open rules and the `npm dist-tag add` recipe all moved to multi-agent-refs/rules.md \"Supported Version Gate\" (loaded by 25 commands, outside this budget) and to the header of require-supported-version.sh, leaving the phase doc with the call, the decision table and the halt. 350 was the smallest step that clears it. Every per-phase max still passes (phase-0-init 11230/12400); phase-3 and phase-4 remain amber on purpose. v15.17.0: total 54400 -> 54900 for the Phase 1 analysis-document step. Phase 2 and Phase 3 pre-flights had BLOCKED on `analysis/<feature>-<platform>.md` since v9.0.0 while nothing produced it, so a full run either aborted at Phase 2 or the model ignored its own BLOCKING contract; Step 4 is the producer. What the phase doc carries is only what cannot be inferred: the when-table (taskType x Figma reference), the four refs in load order, the two artefacts, and that the doc validator fails closed. Compression came first and took the step from 745 tokens to 497: the history of why the gap existed moved to the CHANGELOG, the per-ref one-line descriptions moved into the refs' own headers, and the autopilot carve-out collapsed to one clause. The 17.4k-token analysis engine itself is NOT in this budget - it moved out of commands/ into multi-agent-refs/analysis/{locked,evidence,synthesis,render}.md, loaded on demand, which also took analysis/SKILL.md from 18081 to 5974 tokens and retired its lint grace entry. 500 was the smallest step that clears it; phase-1-analysis sits at 4338/4600 and is amber on purpose, like phase-3 and phase-4. v15.18.0: total 54900 -> 55250 for analysis mode. Three phase docs gained a mode branch that cannot be inferred: Phase 4 reviews a document instead of a diff (validator, the one question reviewers answer, the open-question walk), and Phase 6 publishes instead of committing. Compression came first and was applied twice to the new prose and once to old: the Phase 4 branch went from 320 tokens to 180 and the Phase 6 branch from 190 to 120 by pointing at multi-agent-refs/analysis/{resolve,render}.md, which now hold the walks themselves, and the front-matter parse contract stopped being spelled out in both pre-flights. The analysis engine keeps leaving this budget rather than entering it: intake joined locked/evidence/synthesis/render/resolve in multi-agent-refs/analysis/, which is what let analysis/SKILL.md drop under the 6000 hard cap after its grace entry was retired. 350 was the smallest step that clears it; phase-4 and phase-6 are amber on purpose, as phase-1 and phase-3 already were. v15.20.0: total 55250 -> 55500 for the TDD bridge. Phase 3 pre-flight read the analysis doc's concept table and even said test method names come from it, while nothing read Section 15 - so the RED step invented tests and the analysis test matrix never reached development. Phase 3 step 5b now loads it into state.dev.testPlan[] and Phase 4 step 1.45 cross-checks every planned row against a real test, which is what turns \"analysis quality is output quality\" from a slogan into a finding. Compression came first on both blocks, 300 tokens down to 175, by dropping the enumerated failure modes to one line each and the rationale to one clause; the reasoning lives in the CHANGELOG. 250 was the smallest step that clears it. v15.21.0: total 55500 -> 55800 for the post-analysis confirmation. Phase 2 gained Step 0.9, the last human checkpoint before Phase 3: derived values are shown for confirmation and only Section 20 rows are asked, through the resolve engine that already exists in refs. It belongs here rather than Phase 4 because Phase 4 runs after development, where an answer arrives too late to change anything. Compression came first and twice, 430 tokens down to 250, by collapsing the derived-vs-asked explanation to one sentence each and moving the walk itself to multi-agent-refs/analysis/resolve.md, which Phase 4 and analysis-resolve already mount. 300 was the smallest step that clears it. v15.22.0: total 55800 -> 55900 for the analyst-toolkit hooks. Phase 1 Step 4 now names the two prefs that decide whether a document is produced at all and how deep it goes (forceFull, mode) - the first of those had shipped declared-but-inert and smoke-prefs-consumed caught it - and Phase 4 triage gained one clause: a finding that blames a third-party library asks evidence-github whether it is already open upstream, which turns it into a deferred item with a citation instead of Phase 3 rework on code that is not ours. Compression came first and three times, taking the new prose from 220 tokens to 110, and the Phase 1d evidence contract itself never entered this budget - it lives in multi-agent-refs/analysis/evidence.md beside the phases it belongs to. 100 was the smallest step that clears it, leaving 34 tokens of headroom. phase-4 stays amber and the debt named at v15.10.0 stands: it is the doc that needs structural compression rather than another bump, and the two candidates are the inline triage JSON shape and the 3.4 telemetry block, both of which restate something already authoritative elsewhere. v16.0.0: total 55900 -> 56350 for the depth picker. `--dev` and the four dev-* commands are gone; depth is Phase 0 Step 7.5, which costs phase-0-init a step it did not have. Compression came first and three times, taking the step from 530 tokens to 300: the question wording, the per-taskType recommendation and the mode tables all live in phases/modes.md (outside this budget), so the phase doc carries only what an agent cannot infer - that the step runs after Step 7 and why, who is exempt, that ASK_CHOICE_DEFAULT must be passed explicitly because ask-choice.sh takes the FIRST option on a non-TTY, and that Short flips the Phase 1/2 tiles late rather than pre-marking them. The phase-4 telemetry block named as compression debt at v15.22.0 was collapsed to an emit() helper (-27) and the four dev-* mode files left the tree entirely, but neither offsets a genuinely new phase step. 450 was the smallest step that clears it, leaving 119 tokens of headroom. phase-4 remains amber and its other named candidate, the inline triage JSON shape, was left alone on purpose: it is the prompt the triage agent is handed, not a restatement for readers. v16.2.0: total 56350 -> 56600 for the spec-freshness and reuse-tag contracts. Phase 3 step 3 had compared `state.run.lastAnalysisDigest` since it was written, against a key nothing ever set and that the state schema did not declare, so the staleness branch was unreachable and every run reported fresh by default. Phase 1 now persists the digest and a `base_commit` anchor, and step 3 gained the repo-drift half the digest cannot see: a reused document keeps a matching digest precisely because its evidence inputs did not change, while the code underneath it moved. The second contract is the Section 14 tag reaching development: Phase 2 carries it onto the todo as `sourceTag` and Phase 3 treats it as an instruction, which is what stops a Reuse row from being re-implemented. Compression came first and took the four additions from 380 tokens to 214, by moving every rationale clause out of the phase docs: why the commit anchor exists rather than a digest recomputation lives in this note and the CHANGELOG, and the schema descriptions carry the field semantics. The baseline had 9 tokens of headroom, so no addition of any size could have fit without a bump. 250 was the smallest step that clears it, leaving 45 tokens. phase-3 and phase-4 remain amber. v16.13.0: total 57600 -> 57700 for the code-graph injection and the fable-rung switch. Phase 1 gained Step 2.6 (query the graph, hand Explore a ranked starting set), Phase 7 gained the post-branch graph refresh, and Phase 0 Step 0 gained one line: a prefs switch that resolves every preferredModel: fable persona to opus for the run, which also collapses the Phase 4 Claude Code panel from three reviewers to two. Compression came first and mostly structurally: of roughly 1,630 tokens of new contract text, 1,310 never entered this budget at all - the whole code-graph contract lives in multi-agent-refs/features/code-graph.md (604) and the fable switch's scope table, per-host effects and cost-accounting consequence live in features/model-fallback.md (+707), leaving the phase docs with the call, the pref that gates it and the one fact an agent cannot infer. Phase 4 was compressed on top of that: its TLDR restated the reviewer matrix 270 lines below it, so 36 tokens came back and the doc nets +6 despite carrying two new clauses. One of those clauses is a correction rather than a feature - the consensus rule still said reviewerCount is 2 on Claude Code, which stopped being true when the third reviewer landed in 16.12.0, and the cross-CLI smoke never caught it because it reads the matrix line instead. 100 was the smallest step that clears it, leaving 54 tokens. phase-3 and phase-4 remain amber. v16.17.0: total 57700 -> 57850 for the platform-parity cross-check. Phase 4 gained Step 1.8: when dev-context carries a counterpart app repo, the review compares the change against the other platform on four axes. Compression came first and structurally, as always - of roughly 1,610 tokens of new contract text, 1,490 never entered this budget at all, because the four axes, the file cap, the graph-query recipe, the read-only prohibitions and the rule that an extractor miss may not be reported as an absence all live in multi-agent-refs/platform-parity.md. The step itself was then cut from ~200 tokens to 120 by deleting everything the ref already owns, leaving the trigger, the pointer and the two facts an agent must not infer: the counterpart repo is read-only, and parity findings are never blocking. The baseline had 13 tokens of headroom, so no addition of any size could have fit without a bump. 150 was the smallest step that clears it, leaving 35 tokens. phase-3 and phase-4 remain amber, and phase-4's structural-compression debt still stands. v16.20.0: phase-4-review max 14750 -> 15150 and total 58250 -> 60250 for the cross-round review delta, the scope self-check handoff and the circuit-breaker wiring. Compression came first and structurally: of roughly 3,900 tokens of new contract text, 2,700 never entered this budget at all - the previous-round-findings block, the scope-self-check block, the Step 3.8 state merge, telemetry and picker wording live in multi-agent-refs/features/review-delta.md, and the scope-check record rules and consumers in features/scope-check.md - so the phase docs carry the call, the pref that gates it and the exit table. The Phase 3 stability rule and the trigger-3 write were cut twice more before the bump; phase-3 stays under its max (8692/8950). Phase 4 is the first per-phase max raised since v10.9.0: the doc gained three steps that cannot be inferred (a per-round triage file, a prefix block that changes what reviewers report, and a halt condition), and its structural-compression debt (the inline triage JSON shape, named at v15.10.0) still stands and is the next candidate. 15150 and 60250 were the smallest steps that clear it, leaving 25 and 45 tokens."
39
+ "total_max_tokens": 60500,
40
+ "note": "Token estimate = ceil(chars / 4). Per-phase budget rule: warn = current+10% (rounded to nearest 50), max = current+25%. Gives ~6 edit cycles of headroom before warn trips - intentionally quiet under normal maintenance, loud when a phase grows unusually. Only the active phase is loaded (lazy). Recalibrated at v10.0.0 after the validator/consistency/simplifier/lesson gate contracts landed in phases 1-4. Recalibrated again at v10.9.0 after the verify-by-test (Phase 4 Step 3.7), update-check (Phase 0 Step 0.6), immutable-test (Phase 3 GREEN) and redTests re-entry contracts landed - Step 3.7 prose was compressed to a pointer into refs/features/verify-by-test.md before the recalibration. Total bumped 50000 -> 51000 at v12.5.0 after the worktree residue/traversal-prune contract (Phase 0 + Phase 5 heal) and the Reflexion causal-diagnosis contract (Phase 4 lesson memory) landed; the prose was compressed first (161 tokens reclaimed) and every per-phase max still passes - only the aggregate needed room. Recalibrated again at v13.6.0 after the install-relative path correction: an instruction that names `pipeline/scripts/x` resolves only from a repo checkout, and a run happens in the user's worktree, so 157 references across these docs moved to `$HOME/.claude/...` at +5 bytes each - 196 tokens of pure correctness cost. Same discipline as before: prose was compressed FIRST (149 tokens reclaimed, by pointing Phase 1's Figma tier table at the Phase 0 probe that already resolved it and Phase 4's Codex constraints at the always-loaded AGENTS.md block), and only then were the budgets moved. Five warn lines had been permanently amber, which makes the amber tier useless as a signal, so every warn was reset to the documented current+10% and the four maxes that the new warn would have collided with were reset to current+25%. Aggregate 51000 -> 51500. Total bumped 51500 -> 52200 at v14.0.0 after Phase 4 Review entered the four --dev mode phase sets and the criteria-resolution contract (Step 1.78) landed. Same discipline as every prior bump: prose was compressed FIRST, 820 tokens reclaimed, before the number moved. Two of those compressions are structural rather than cosmetic - the hardcoded SwiftUI interaction list in Step 1.5 and the SwiftUI convention paragraph in Step 2.8 were transcriptions of rules that now live in a scoped registry, so keeping them here would have re-created the drift this release exists to remove, and the third moved the Step 1.78 full contract into refs/features/skill-conformance.md leaving a pointer. What remains is contract text that cannot be inferred: the manifest's four consumer-visible parts, the conformance checklist the reviewers must return, and the fail-closed semantics. Every per-phase max still passes (phase-4 12405/14750); only the aggregate needed room. Total bumped 52200 -> 52700 at v14.1.0 after two more contracts landed: stack skill routing (Phase 3 pre-flight step 9) and worktree finalize (Phase 6 step 9). Compression came first, as always, and twice: 224 tokens out of Phase 3 by pointing its criteria-ledger and routing steps at their feature files instead of restating them, and 190 out of Phase 6 by moving the finalize contract into refs/features/worktree-finalize.md and leaving the invocation plus the exit-3 semantics. Both new contracts follow the pattern the earlier ones set: the phase doc carries the call and the decision, the feature file carries the reasoning, and the feature files are outside this budget because it loops only the eight phase-N-* keys. Every per-phase max still passes (phase-3 7677/8950, phase-6 5223/6150 and both under warn); only the aggregate needed room. Total bumped 52700 -> 52750 for the Phase 0 Step 3 branch-persistence correction: the step wrote the legacy `projects[].branches` while the TTL filter two sections below read `global.recentBranches`, and both spots named a `{name, lastUsed}` shape the schema rejects (`branch` required, `additionalProperties: false`), so the recent-branch picker option could never populate and a literal implementation would have failed prefs validation. Naming the right target, the right key and the legacy field to avoid costs 41 tokens over the one line it replaces. Compression came first and was applied three times to the replacement text itself, from 120 tokens down to 66, by moving the rationale out of the phase doc entirely: the reasoning now lives where it is enforced, in the migrate-prefs carry-forward comment and the smoke-pref-migration f7 block, leaving the phase doc with only the instruction. 50 was the smallest step that clears it; phase-0-init sits at 10893/12400, far under its own max, so this is purely an aggregate ceiling. v15.0.0: total 52750 -> 53100, the stack-skill tables in phase-1/2/4 now carry plugin-namespaced names (ai-<stack>-toolkit:<skill>) - functional prefixes, ~170 tokens. v15.10.0: total 53350 -> 53950 for the memory-recall + context-offload contracts (Phase 1 two-block durable-knowledge injection and its telemetry, Phase 3 build-log offload pipe, Phase 4 ranked prior art, offload pipe and recall telemetry). Compression came first and twice, taking the new prose from 1168 tokens to 580: the reasoning behind the two blocks lives in multi-agent-refs/prompt-assembly.md and the reasoning behind the offload filter lives in the offload-ref.sh header, both outside this budget, so the phase docs carry only the call, the pref that gates it and the one fact an agent cannot infer - that the evidence gate still reads the whole build log, so offloading changes what is read, never what counts as a verified pass. Every per-phase max still passes (phase-3 7985/8950, phase-4 12997/14750); phase-3 and phase-4 crossed their warn lines and are left amber on purpose, because that is the signal that those two docs are the next ones needing structural compression rather than another bump. v15.13.0: total 53950 -> 54050 for the prefs-to-flag bridges. Five settings had shipped declared-but-inert: contextOffload.minLines and .tailLines (fixed in 15.11.0), learningsLedger.maxBriefEntries, and testGap.scanTree and .promoteSeverity - the last two declared in the schema AND implemented as flags in the scanner, with nothing in between reading the pref and passing the flag. Wiring three of them costs the phase docs 94 tokens, which is the wiring itself and not prose: two `--max` substitutions and a three-line GAP_FLAGS block. Compression came first and twice, as always: the rationale that would have sat in phase-5 now lives in the header of smoke-prefs-consumed.sh, the gate that makes this class fail a build instead of shipping, and a `--severity-promote` table row was dropped because the invocation above it now shows the flag and names the pref that triggers it, which the row did not. 100 was the smallest step that clears it. Every per-phase max still passes; phase-3 and phase-4 remain amber on purpose. v15.14.0: total 54050 -> 54400 for the supported-version gate. Phase 0 Step 0.6 stopped being purely advisory: a release can now publish an npm dist-tag `required` that names the oldest runnable version, and below it the run halts instead of nagging. What the phase doc has to carry is the part an agent cannot infer - the third stdout field, that the halt is identical in autopilot, and that the run must NOT continue on the freshly updated install because its docs were already loaded from the old version. Compression came first, as always, and took the new prose from 469 tokens to 337: the rationale for the floor, the exemption list, the fail-open rules and the `npm dist-tag add` recipe all moved to multi-agent-refs/rules.md \"Supported Version Gate\" (loaded by 25 commands, outside this budget) and to the header of require-supported-version.sh, leaving the phase doc with the call, the decision table and the halt. 350 was the smallest step that clears it. Every per-phase max still passes (phase-0-init 11230/12400); phase-3 and phase-4 remain amber on purpose. v15.17.0: total 54400 -> 54900 for the Phase 1 analysis-document step. Phase 2 and Phase 3 pre-flights had BLOCKED on `analysis/<feature>-<platform>.md` since v9.0.0 while nothing produced it, so a full run either aborted at Phase 2 or the model ignored its own BLOCKING contract; Step 4 is the producer. What the phase doc carries is only what cannot be inferred: the when-table (taskType x Figma reference), the four refs in load order, the two artefacts, and that the doc validator fails closed. Compression came first and took the step from 745 tokens to 497: the history of why the gap existed moved to the CHANGELOG, the per-ref one-line descriptions moved into the refs' own headers, and the autopilot carve-out collapsed to one clause. The 17.4k-token analysis engine itself is NOT in this budget - it moved out of commands/ into multi-agent-refs/analysis/{locked,evidence,synthesis,render}.md, loaded on demand, which also took analysis/SKILL.md from 18081 to 5974 tokens and retired its lint grace entry. 500 was the smallest step that clears it; phase-1-analysis sits at 4338/4600 and is amber on purpose, like phase-3 and phase-4. v15.18.0: total 54900 -> 55250 for analysis mode. Three phase docs gained a mode branch that cannot be inferred: Phase 4 reviews a document instead of a diff (validator, the one question reviewers answer, the open-question walk), and Phase 6 publishes instead of committing. Compression came first and was applied twice to the new prose and once to old: the Phase 4 branch went from 320 tokens to 180 and the Phase 6 branch from 190 to 120 by pointing at multi-agent-refs/analysis/{resolve,render}.md, which now hold the walks themselves, and the front-matter parse contract stopped being spelled out in both pre-flights. The analysis engine keeps leaving this budget rather than entering it: intake joined locked/evidence/synthesis/render/resolve in multi-agent-refs/analysis/, which is what let analysis/SKILL.md drop under the 6000 hard cap after its grace entry was retired. 350 was the smallest step that clears it; phase-4 and phase-6 are amber on purpose, as phase-1 and phase-3 already were. v15.20.0: total 55250 -> 55500 for the TDD bridge. Phase 3 pre-flight read the analysis doc's concept table and even said test method names come from it, while nothing read Section 15 - so the RED step invented tests and the analysis test matrix never reached development. Phase 3 step 5b now loads it into state.dev.testPlan[] and Phase 4 step 1.45 cross-checks every planned row against a real test, which is what turns \"analysis quality is output quality\" from a slogan into a finding. Compression came first on both blocks, 300 tokens down to 175, by dropping the enumerated failure modes to one line each and the rationale to one clause; the reasoning lives in the CHANGELOG. 250 was the smallest step that clears it. v15.21.0: total 55500 -> 55800 for the post-analysis confirmation. Phase 2 gained Step 0.9, the last human checkpoint before Phase 3: derived values are shown for confirmation and only Section 20 rows are asked, through the resolve engine that already exists in refs. It belongs here rather than Phase 4 because Phase 4 runs after development, where an answer arrives too late to change anything. Compression came first and twice, 430 tokens down to 250, by collapsing the derived-vs-asked explanation to one sentence each and moving the walk itself to multi-agent-refs/analysis/resolve.md, which Phase 4 and analysis-resolve already mount. 300 was the smallest step that clears it. v15.22.0: total 55800 -> 55900 for the analyst-toolkit hooks. Phase 1 Step 4 now names the two prefs that decide whether a document is produced at all and how deep it goes (forceFull, mode) - the first of those had shipped declared-but-inert and smoke-prefs-consumed caught it - and Phase 4 triage gained one clause: a finding that blames a third-party library asks evidence-github whether it is already open upstream, which turns it into a deferred item with a citation instead of Phase 3 rework on code that is not ours. Compression came first and three times, taking the new prose from 220 tokens to 110, and the Phase 1d evidence contract itself never entered this budget - it lives in multi-agent-refs/analysis/evidence.md beside the phases it belongs to. 100 was the smallest step that clears it, leaving 34 tokens of headroom. phase-4 stays amber and the debt named at v15.10.0 stands: it is the doc that needs structural compression rather than another bump, and the two candidates are the inline triage JSON shape and the 3.4 telemetry block, both of which restate something already authoritative elsewhere. v16.0.0: total 55900 -> 56350 for the depth picker. `--dev` and the four dev-* commands are gone; depth is Phase 0 Step 7.5, which costs phase-0-init a step it did not have. Compression came first and three times, taking the step from 530 tokens to 300: the question wording, the per-taskType recommendation and the mode tables all live in phases/modes.md (outside this budget), so the phase doc carries only what an agent cannot infer - that the step runs after Step 7 and why, who is exempt, that ASK_CHOICE_DEFAULT must be passed explicitly because ask-choice.sh takes the FIRST option on a non-TTY, and that Short flips the Phase 1/2 tiles late rather than pre-marking them. The phase-4 telemetry block named as compression debt at v15.22.0 was collapsed to an emit() helper (-27) and the four dev-* mode files left the tree entirely, but neither offsets a genuinely new phase step. 450 was the smallest step that clears it, leaving 119 tokens of headroom. phase-4 remains amber and its other named candidate, the inline triage JSON shape, was left alone on purpose: it is the prompt the triage agent is handed, not a restatement for readers. v16.2.0: total 56350 -> 56600 for the spec-freshness and reuse-tag contracts. Phase 3 step 3 had compared `state.run.lastAnalysisDigest` since it was written, against a key nothing ever set and that the state schema did not declare, so the staleness branch was unreachable and every run reported fresh by default. Phase 1 now persists the digest and a `base_commit` anchor, and step 3 gained the repo-drift half the digest cannot see: a reused document keeps a matching digest precisely because its evidence inputs did not change, while the code underneath it moved. The second contract is the Section 14 tag reaching development: Phase 2 carries it onto the todo as `sourceTag` and Phase 3 treats it as an instruction, which is what stops a Reuse row from being re-implemented. Compression came first and took the four additions from 380 tokens to 214, by moving every rationale clause out of the phase docs: why the commit anchor exists rather than a digest recomputation lives in this note and the CHANGELOG, and the schema descriptions carry the field semantics. The baseline had 9 tokens of headroom, so no addition of any size could have fit without a bump. 250 was the smallest step that clears it, leaving 45 tokens. phase-3 and phase-4 remain amber. v16.13.0: total 57600 -> 57700 for the code-graph injection and the fable-rung switch. Phase 1 gained Step 2.6 (query the graph, hand Explore a ranked starting set), Phase 7 gained the post-branch graph refresh, and Phase 0 Step 0 gained one line: a prefs switch that resolves every preferredModel: fable persona to opus for the run, which also collapses the Phase 4 Claude Code panel from three reviewers to two. Compression came first and mostly structurally: of roughly 1,630 tokens of new contract text, 1,310 never entered this budget at all - the whole code-graph contract lives in multi-agent-refs/features/code-graph.md (604) and the fable switch's scope table, per-host effects and cost-accounting consequence live in features/model-fallback.md (+707), leaving the phase docs with the call, the pref that gates it and the one fact an agent cannot infer. Phase 4 was compressed on top of that: its TLDR restated the reviewer matrix 270 lines below it, so 36 tokens came back and the doc nets +6 despite carrying two new clauses. One of those clauses is a correction rather than a feature - the consensus rule still said reviewerCount is 2 on Claude Code, which stopped being true when the third reviewer landed in 16.12.0, and the cross-CLI smoke never caught it because it reads the matrix line instead. 100 was the smallest step that clears it, leaving 54 tokens. phase-3 and phase-4 remain amber. v16.17.0: total 57700 -> 57850 for the platform-parity cross-check. Phase 4 gained Step 1.8: when dev-context carries a counterpart app repo, the review compares the change against the other platform on four axes. Compression came first and structurally, as always - of roughly 1,610 tokens of new contract text, 1,490 never entered this budget at all, because the four axes, the file cap, the graph-query recipe, the read-only prohibitions and the rule that an extractor miss may not be reported as an absence all live in multi-agent-refs/platform-parity.md. The step itself was then cut from ~200 tokens to 120 by deleting everything the ref already owns, leaving the trigger, the pointer and the two facts an agent must not infer: the counterpart repo is read-only, and parity findings are never blocking. The baseline had 13 tokens of headroom, so no addition of any size could have fit without a bump. 150 was the smallest step that clears it, leaving 35 tokens. phase-3 and phase-4 remain amber, and phase-4's structural-compression debt still stands. v16.20.0: phase-4-review max 14750 -> 15150 and total 58250 -> 60250 for the cross-round review delta, the scope self-check handoff and the circuit-breaker wiring. Compression came first and structurally: of roughly 3,900 tokens of new contract text, 2,700 never entered this budget at all - the previous-round-findings block, the scope-self-check block, the Step 3.8 state merge, telemetry and picker wording live in multi-agent-refs/features/review-delta.md, and the scope-check record rules and consumers in features/scope-check.md - so the phase docs carry the call, the pref that gates it and the exit table. The Phase 3 stability rule and the trigger-3 write were cut twice more before the bump; phase-3 stays under its max (8692/8950). Phase 4 is the first per-phase max raised since v10.9.0: the doc gained three steps that cannot be inferred (a per-round triage file, a prefix block that changes what reviewers report, and a halt condition), and its structural-compression debt (the inline triage JSON shape, named at v15.10.0) still stands and is the next candidate. 15150 and 60250 were the smallest steps that clear it, leaving 25 and 45 tokens. v16.23.0: phase-0-init max 12400 -> 12500 and total 60250 -> 60500 for the widget-registration call and the accounting gate. Phase 0 gained the `tiles` call and the exit-3 rule, Phase 7 gained the run report; together they are contract an agent cannot infer - which call registers this host's widget, and that a completion is refused without recorded spend. Compression came first and twice, taking the new prose from 472 tokens to 255: the per-host call list moved into tracker-contract.md \"The card is not the widget\" and the record-then-rerun recovery into \"Accounting is a gate\", both outside this budget, leaving the phase docs with the call and the one fact that cannot be looked up. 100 and 250 were the smallest steps that clear it, leaving 74 and 40 tokens. phase-3 and phase-4 remain amber."
41
41
  }