@mmerterden/multi-agent-pipeline 16.12.0 → 16.13.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +20 -0
- package/README.md +4 -4
- package/README.tr.md +4 -4
- package/docs/adr/0010-own-code-graph.md +129 -0
- package/docs/adr/README.md +1 -0
- package/docs/architecture.md +2 -2
- package/docs/ecosystem.md +5 -5
- package/docs/features.md +8 -0
- package/package.json +1 -1
- package/pipeline/commands/multi-agent/graph/SKILL.md +105 -0
- package/pipeline/commands/multi-agent/help/SKILL.md +8 -8
- package/pipeline/commands/multi-agent/sync/SKILL.md +12 -12
- package/pipeline/commands/multi-agent/uninstall/SKILL.md +9 -7
- package/pipeline/multi-agent-refs/cross-cli-contract.md +10 -10
- package/pipeline/multi-agent-refs/features/code-graph.md +62 -0
- package/pipeline/multi-agent-refs/features/model-fallback.md +44 -2
- package/pipeline/multi-agent-refs/knowledge.md +6 -0
- package/pipeline/multi-agent-refs/phases/phase-0-init.md +1 -1
- package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +5 -0
- package/pipeline/multi-agent-refs/phases/phase-4-review.md +3 -3
- package/pipeline/multi-agent-refs/phases/phase-7-report.md +2 -0
- package/pipeline/preferences-template.json +2 -0
- package/pipeline/schemas/code-graph.schema.json +91 -0
- package/pipeline/schemas/prefs.schema.json +45 -0
- package/pipeline/schemas/token-budget.json +2 -2
- package/pipeline/scripts/_code-graph.mjs +518 -0
- package/pipeline/scripts/_path-match.mjs +87 -0
- package/pipeline/scripts/code-graph-rules/android.json +130 -0
- package/pipeline/scripts/code-graph-rules/ios.json +95 -0
- package/pipeline/scripts/code-graph-rules/node.json +151 -0
- package/pipeline/scripts/code-graph-rules/python.json +91 -0
- package/pipeline/scripts/graph-affected.mjs +161 -0
- package/pipeline/scripts/graph-build.mjs +157 -0
- package/pipeline/scripts/graph-query.mjs +191 -0
- package/pipeline/scripts/graph-report.mjs +237 -0
- package/pipeline/scripts/smoke-cross-cli-behavior.sh +7 -4
- package/pipeline/scripts/test-gap-rules/ios.json +38 -10
- package/pipeline/scripts/test-gap-scan.mjs +2 -21
- package/pipeline/scripts/uninstall.mjs +11 -2
- package/pipeline/scripts/validate-code-graph.mjs +174 -0
- package/pipeline/skills/.skills-index.json +14 -3
- package/pipeline/skills/shared/README.md +6 -5
- package/pipeline/skills/shared/core/multi-agent-graph/SKILL.md +106 -0
- package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +11 -11
- package/pipeline/skills/shared/core/multi-agent-uninstall/SKILL.md +4 -4
- package/pipeline/skills/skills-index.md +4 -3
|
@@ -6,18 +6,18 @@
|
|
|
6
6
|
|
|
7
7
|
---
|
|
8
8
|
|
|
9
|
-
## 1. Command Inventory (
|
|
9
|
+
## 1. Command Inventory (54 files, 50 live commands)
|
|
10
10
|
|
|
11
11
|
```
|
|
12
12
|
analysis, analysis-resolve, autopilot, build-optimize, channels,
|
|
13
13
|
complaint-analysis, create-jira, design-check, dev, dev-autopilot, dev-local,
|
|
14
|
-
dev-local-autopilot, diff-explain, feedback, forget, garbage-collect,
|
|
15
|
-
ios-coding-standard, issue, jira, kill, language, local,
|
|
16
|
-
log, manual-test, prune-logs, prune-prompts, purge,
|
|
17
|
-
resume-local, review, review-analysis, review-issue,
|
|
18
|
-
save, scan, search, setup, stack, status,
|
|
19
|
-
|
|
20
|
-
testflight-validation, uninstall, update
|
|
14
|
+
dev-local-autopilot, diff-explain, feedback, forget, garbage-collect,
|
|
15
|
+
graph, help, ios-coding-standard, issue, jira, kill, language, local,
|
|
16
|
+
local-autopilot, log, manual-test, prune-logs, prune-prompts, purge,
|
|
17
|
+
refactor, resume, resume-local, review, review-analysis, review-issue,
|
|
18
|
+
review-jira, routines, save, scan, search, setup, stack, status,
|
|
19
|
+
store-ready, sync, test, test-accessibility, test-dark-mode,
|
|
20
|
+
test-dynamic-type, test-screenshots, testflight-validation, uninstall, update
|
|
21
21
|
```
|
|
22
22
|
|
|
23
23
|
Categories:
|
|
@@ -27,12 +27,12 @@ Categories:
|
|
|
27
27
|
- **Pipeline entries**: `autopilot`, `local`, `local-autopilot` (plus the bare `/multi-agent` in the dispatcher). Depth is not a command: `/multi-agent` and `local` ask Full or Short at Phase 0 Step 7.5; the two autopilot entries never ask and always run Full.
|
|
28
28
|
- **Retired stubs** (v16.0.0, deleted next minor - they print a redirect and run no phase): `dev`, `dev-local` redirect to the picker entries with Short; `dev-autopilot`, `dev-local-autopilot` have no equivalent, because fast-plus-unattended no longer exists
|
|
29
29
|
- **Tail modes** (run the pipeline tail over already-done local work): `resume-local`
|
|
30
|
-
- **Ops commands** (one-shot, no worktree): `status`, `log`, `kill`, `purge`, `uninstall`, `resume`, `review`, `review-jira`, `review-issue`, `analysis`, `analysis-resolve`, `complaint-analysis`, `build-optimize`, `channels`, `scan`, `search`, `diff-explain`, `garbage-collect`, `prune-logs`, `prune-prompts`
|
|
30
|
+
- **Ops commands** (one-shot, no worktree): `status`, `log`, `kill`, `purge`, `uninstall`, `resume`, `review`, `review-jira`, `review-issue`, `analysis`, `analysis-resolve`, `complaint-analysis`, `build-optimize`, `channels`, `scan`, `search`, `diff-explain`, `garbage-collect`, `graph`, `prune-logs`, `prune-prompts`
|
|
31
31
|
- **Local audits** (worktree only to build; no commit, push, PR or channels): `design-check`, `testflight-validation`, `ios-coding-standard`. `testflight-validation` additionally never invokes `altool --upload-app` - a validation run must not be able to ship a build by accident.
|
|
32
32
|
- **Meta-ops**: `setup`, `sync`, `update`, `help`, `refactor`, `test`, `stack`, `manual-test`, `language`
|
|
33
33
|
- **Routines** (user-defined routine registry; the routines they create are local-only and never synced): `save`, `routines`, `forget`
|
|
34
34
|
|
|
35
|
-
The count is
|
|
35
|
+
The count is 54 files and 50 live commands until the four stubs are deleted, at which point both numbers become 50. A stub is still installed and still invocable, so counting it as absent would be wrong; counting it as a command would be worse.
|
|
36
36
|
|
|
37
37
|
> **Inventory drift is a contract violation.** Adding a slash command under `pipeline/commands/multi-agent/` without updating this list + its counterpart Copilot dir (`pipeline/skills/shared/core/multi-agent-<cmd>/`) is a merge blocker. `smoke-commands-skills-parity.sh` enforces command ↔ skill directory parity; `smoke-cross-cli-behavior.sh` enforces behavior parity. This doc is the authoritative command list - bump the count + table together.
|
|
38
38
|
|
|
@@ -0,0 +1,62 @@
|
|
|
1
|
+
## Code Graph (Phase 1 Step 2.6 + Phase 7 Step 3)
|
|
2
|
+
|
|
3
|
+
A deterministic, LLM-free map of what a repo declares and what refers to what,
|
|
4
|
+
written to `~/.claude/knowledge/<project>/code-graph.json`. Gated by
|
|
5
|
+
`prefs.global.codeGraph.enabled` (default `false`); with it off, Phase 1 and
|
|
6
|
+
Phase 7 behave exactly as they did before. Design, trade and measurements:
|
|
7
|
+
`docs/adr/0010-own-code-graph.md`.
|
|
8
|
+
|
|
9
|
+
### Phase 1 Step 2.6 - query before dispatching Explore
|
|
10
|
+
|
|
11
|
+
Runs only when a rule file exists for `detectedStack`. A stack with no rule file
|
|
12
|
+
is reported as unsupported, never guessed at.
|
|
13
|
+
|
|
14
|
+
```bash
|
|
15
|
+
GRAPH_PATH="$HOME/.claude/knowledge/$(basename "$PROJECT_ROOT")/code-graph.json"
|
|
16
|
+
|
|
17
|
+
node $HOME/.claude/scripts/graph-report.mjs --graph "$GRAPH_PATH" --status
|
|
18
|
+
node $HOME/.claude/scripts/graph-build.mjs --root "$PROJECT_ROOT" --stack "$STACK" \
|
|
19
|
+
--max-nodes "${prefs_codeGraph_maxNodes:-200000}" --json
|
|
20
|
+
node $HOME/.claude/scripts/validate-code-graph.mjs "$GRAPH_PATH"
|
|
21
|
+
node $HOME/.claude/scripts/graph-query.mjs "$TASK_TITLE $TASK_DESCRIPTION" \
|
|
22
|
+
--graph "$GRAPH_PATH" --budget "${prefs_codeGraph_queryBudget:-2000}" --json
|
|
23
|
+
node $HOME/.claude/scripts/graph-affected.mjs "<symbol the task names>" \
|
|
24
|
+
--graph "$GRAPH_PATH" --depth 2 --json
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
`GRAPH_PATH` is derived once and passed to every call. The query, affected and
|
|
28
|
+
report scripts default it from the CWD's basename, and a run happens in a
|
|
29
|
+
worktree whose basename need not equal the repo's, so a defaulted path can point
|
|
30
|
+
at a graph that was never built. Pass it.
|
|
31
|
+
|
|
32
|
+
Build only when `baseCommit` no longer matches HEAD. A build is not usable until
|
|
33
|
+
`validate-code-graph.mjs` exits 0: a graph whose edges point at missing nodes
|
|
34
|
+
truncates traversals silently, so a non-zero exit means skip the injection and
|
|
35
|
+
run Explore as if no graph existed.
|
|
36
|
+
|
|
37
|
+
`graph-query` output becomes the Explore agents' starting file set;
|
|
38
|
+
`graph-affected` output feeds `analysis.touchedAreas[]`.
|
|
39
|
+
|
|
40
|
+
Use it to narrow an open-ended search, not to replace a grep for a name the task
|
|
41
|
+
already spells out. Measured on a 4,300-file Swift app at a fixed 30k retrieval
|
|
42
|
+
budget, it roughly doubled coverage at under half the cost on domain-word
|
|
43
|
+
questions and lost narrowly to `grep -lw` on exact type names.
|
|
44
|
+
|
|
45
|
+
### Phase 7 Step 3 - refresh after the branch changed code
|
|
46
|
+
|
|
47
|
+
Runs when `prefs.global.codeGraph.enabled` and `prefs.global.codeGraph.autoRefresh`
|
|
48
|
+
are both true.
|
|
49
|
+
|
|
50
|
+
```bash
|
|
51
|
+
GRAPH_PATH="$HOME/.claude/knowledge/$(basename "$PROJECT_ROOT")/code-graph.json"
|
|
52
|
+
|
|
53
|
+
node $HOME/.claude/scripts/graph-build.mjs --root "$PROJECT_ROOT" --stack "$STACK" \
|
|
54
|
+
--max-nodes "${prefs_codeGraph_maxNodes:-200000}" --json
|
|
55
|
+
node $HOME/.claude/scripts/validate-code-graph.mjs "$GRAPH_PATH"
|
|
56
|
+
node $HOME/.claude/scripts/graph-report.mjs --graph "$GRAPH_PATH"
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
The rebuild costs no API tokens, so it runs every task rather than on a staleness
|
|
60
|
+
heuristic. A non-zero validator exit keeps the previous graph and logs
|
|
61
|
+
`knowledge.graph_invalid`; it never fails the run - a stale graph is a degraded
|
|
62
|
+
Phase 1, not a broken deliverable.
|
|
@@ -74,7 +74,7 @@ the override is per-dispatch and leaves files untouched.
|
|
|
74
74
|
## Prefs knob
|
|
75
75
|
|
|
76
76
|
`prefs.global.modelFallback` (template default below; absent knob = `enabled: true`
|
|
77
|
-
with no date gate; an absent `floorModel` defaults to `haiku`, so installs that
|
|
77
|
+
with no date gate; an absent `modelFallback.floorModel` defaults to `haiku`, so installs that
|
|
78
78
|
predate this field still get the second downgrade step):
|
|
79
79
|
|
|
80
80
|
```json
|
|
@@ -83,6 +83,7 @@ predate this field still get the second downgrade step):
|
|
|
83
83
|
"premiumTierUntil": null,
|
|
84
84
|
"fallbackModel": "sonnet",
|
|
85
85
|
"floorModel": "haiku",
|
|
86
|
+
"fableEnabled": true,
|
|
86
87
|
"onDispatchError": true
|
|
87
88
|
}
|
|
88
89
|
```
|
|
@@ -90,13 +91,54 @@ predate this field still get the second downgrade step):
|
|
|
90
91
|
| Field | Meaning |
|
|
91
92
|
|---|---|
|
|
92
93
|
| `enabled` | Master switch. `false` = always dispatch the persona's `preferredModel`, fail loudly on error. |
|
|
94
|
+
| `modelFallback.fableEnabled` | Whether the `fable` rung exists at all on Claude Code. `false` starts every `preferredModel: fable` persona on `opus`. See "Turning the fable rung off" below. |
|
|
93
95
|
| `premiumTierUntil` | ISO date (`YYYY-MM-DD`) or `null`. When set and today is **after** this date, every `preferredModel` dispatch is downgraded to `fallbackModel` unless the user re-confirms (see Date gate). Use when the top tier is included in a plan only until a known date. |
|
|
94
96
|
| `fallbackModel` | Target of the first downgrade step. Default `sonnet` (the next tier below opus). |
|
|
95
|
-
| `floorModel` | Last-resort tier when `fallbackModel` also fails to dispatch. Default `haiku`. Set to the same value as `fallbackModel` (or `null`) to disable the second step and halt after one downgrade. |
|
|
97
|
+
| `modelFallback.floorModel` | Last-resort tier when `fallbackModel` also fails to dispatch. Default `haiku`. Set to the same value as `fallbackModel` (or `null`) to disable the second step and halt after one downgrade. |
|
|
96
98
|
| `onDispatchError` | When `true`, a failed top-tier dispatch (model unavailable / quota / 4xx on model id) retries once on `fallbackModel`, and a failed `fallbackModel` dispatch retries once on `floorModel`, instead of aborting the phase. |
|
|
97
99
|
|
|
100
|
+
## Turning the fable rung off
|
|
101
|
+
|
|
102
|
+
`prefs.global.modelFallback.fableEnabled: false` is a **cost control, not a fallback**. The other three
|
|
103
|
+
triggers below react to something going wrong; this one asserts up front that a
|
|
104
|
+
rung is not in play, so there is no dispatch attempt, no error, and nothing to
|
|
105
|
+
recover from. That distinction is why it is checked before them.
|
|
106
|
+
|
|
107
|
+
Scope is Claude Code only, and the reason differs per host:
|
|
108
|
+
|
|
109
|
+
| Host | Effect of `fableEnabled: false` | Why |
|
|
110
|
+
|---|---|---|
|
|
111
|
+
| Claude Code | Every `preferredModel: fable` persona (`ios/android/backend-architect`, `code-reviewer`, triage) starts on `opus` | This is the only host where the rung means Fable 5 |
|
|
112
|
+
| Copilot CLI | none | Fable 5 is not offered there; its personas never sat on this rung |
|
|
113
|
+
| Codex CLI | none, deliberately | The `fable` rung there means `gpt-5.6 @ xhigh`, a different model on a different account. Switching it off from a knob named after an Anthropic model would surprise a Codex user, so it does not |
|
|
114
|
+
|
|
115
|
+
**Phase 4 loses a reviewer, and says so.** Reviewer 1 lands on `opus`, which
|
|
116
|
+
Reviewer 2 already holds. Dispatching one model twice is not cross-model review,
|
|
117
|
+
so the Claude Code panel collapses to two reviewers (`opus` + `sonnet`),
|
|
118
|
+
`consensus.reviewerCount` records `2`, and the `unverified` verdict rule matters
|
|
119
|
+
more, not less: two Anthropic models agreeing on a judgment call was already weak
|
|
120
|
+
evidence, and there is now one fewer of them. Triage also runs on `opus`, which
|
|
121
|
+
makes it the same model as Reviewer 1; the Step 3 anonymisation requirement
|
|
122
|
+
already covers that case and is not optional here.
|
|
123
|
+
|
|
124
|
+
**Cost accounting.** `prefs.global.costBudget.priceAt` defaults to `fable` to
|
|
125
|
+
keep the estimate an upper bound. With the rung off, that default prices every
|
|
126
|
+
call above what it can cost and trips the budget ceiling early, which then
|
|
127
|
+
triggers a downgrade nobody needed. Set `priceAt` to `opus` alongside
|
|
128
|
+
`fableEnabled: false`.
|
|
129
|
+
|
|
130
|
+
No log line is emitted per dispatch for this: it is a configured absence, not a
|
|
131
|
+
degrade event, and one line per persona per run would be noise. Phase 0 Step 0
|
|
132
|
+
prints it once instead:
|
|
133
|
+
`INFO: fable rung disabled by prefs; architect/reviewer/triage personas dispatch on opus (Phase 4 panel = 2 reviewers).`
|
|
134
|
+
|
|
98
135
|
## Triggers (checked in this order)
|
|
99
136
|
|
|
137
|
+
0. **Fable rung disabled (Phase 0 Step 0, once per run).** If
|
|
138
|
+
`fableEnabled: false`, resolve every `preferredModel: fable` persona to `opus`
|
|
139
|
+
for the whole run before any other trigger is considered. A later dispatch
|
|
140
|
+
error then walks `opus -> sonnet -> haiku` from there, exactly as it does for
|
|
141
|
+
a persona that declared `opus` in the first place.
|
|
100
142
|
1. **Date gate (Phase 0 Step 0, once per run).** If `premiumTierUntil` is set and
|
|
101
143
|
in the past, print one line:
|
|
102
144
|
`WARN: premium tier plan window ended <date>; preferredModel personas will dispatch on <fallbackModel>. Set prefs.global.modelFallback.premiumTierUntil to null to keep the top tier on usage credits.`
|
|
@@ -11,6 +11,8 @@ $HOME/.claude/knowledge/
|
|
|
11
11
|
patterns.md - naming, conventions, recurring structures
|
|
12
12
|
gotchas.md - build errors, edge cases, workarounds
|
|
13
13
|
decisions.md - architectural decisions and rationale (ADR-lite)
|
|
14
|
+
code-graph.json - symbols, imports and references, extracted without an LLM
|
|
15
|
+
GRAPH_REPORT.md - hubs, modules, external deps, unconnected files
|
|
14
16
|
another-project/
|
|
15
17
|
architecture.md
|
|
16
18
|
patterns.md
|
|
@@ -55,6 +57,10 @@ Knowledge files grow over time. Maintenance rules:
|
|
|
55
57
|
- Every entry has a `<!-- captured: {date}, task: {id} -->` tag
|
|
56
58
|
- 90-day-old entries are considered "stale" - not used without verification
|
|
57
59
|
- Stale check: orchestrator checks file mtime during Phase 1 knowledge injection
|
|
60
|
+
- The two graph files age differently, so they use a different signal: `code-graph.json`
|
|
61
|
+
is stale when its `baseCommit` no longer matches HEAD, not when it is 90 days old.
|
|
62
|
+
A repo nobody touched for a year has a year-old graph that is still exactly correct,
|
|
63
|
+
and a rebuild is seconds, so age is the wrong question to ask of it.
|
|
58
64
|
- Stale entries are added to the prompt with a "STALE - verify before relying" tag
|
|
59
65
|
- `prune-logs` does not touch knowledge - only deletes logs and state
|
|
60
66
|
- `purge` does not touch knowledge either - separate command: `/multi-agent clear-knowledge {project}`
|
|
@@ -37,7 +37,7 @@ OUTPUT_LANG=$(jq -r '.global.outputLanguage // "en"' "$PREFS_FILE" 2>/dev/null |
|
|
|
37
37
|
|
|
38
38
|
From this point on, everything the user reads renders in `$OUTPUT_LANG`: conversational lines, `AskUserQuestion` `question`/`label`/`description`, and external payload bodies (PR/Jira/Confluence). English stays only on `header`, commit messages, branch names, PR title prefixes, identifiers. Full matrix: `rules.md` "Language Application".
|
|
39
39
|
|
|
40
|
-
**Model
|
|
40
|
+
**Model tier resolution** (same step, once per run): read `prefs.global.modelFallback`. If `fableEnabled` is `false`, every `preferredModel: fable` persona resolves to `opus` for this run and the Phase 4 Claude Code panel is 2 reviewers, not 3; print the one-line INFO. Then, if `premiumTierUntil` is set and in the past, apply the date-gate trigger - `preferredModel` personas dispatch on `fallbackModel`, with the one-line WARN. Both lines and the exact ordering: `$HOME/.claude/multi-agent-refs/features/model-fallback.md`. Dispatch-error and budget triggers apply per-dispatch later; nothing else to do here.
|
|
41
41
|
|
|
42
42
|
**First-run guard**: After loading prefs, check if `keychainMapping` has at least one non-null value. If ALL values are null (template defaults - setup never ran), show:
|
|
43
43
|
```
|
|
@@ -143,6 +143,10 @@ This informs:
|
|
|
143
143
|
|
|
144
144
|
Gated by `prefs.global.repoMap.enabled` (default: `false`). When enabled, runs `$HOME/.claude/scripts/repo-map.mjs` and injects the budgeted result into each Explore prompt as `${REPO_MAP}`. Aider-style: deterministic, no embeddings, sub-second, advisory only. Full wiring (helper invocation, properties, when-to-enable): `$HOME/.claude/multi-agent-refs/features/repo-map.md`.
|
|
145
145
|
|
|
146
|
+
#### Step 2.6 - Code Graph Injection (advisory, opt-in)
|
|
147
|
+
|
|
148
|
+
Gated by `prefs.global.codeGraph.enabled` (default: `false`). With a rule file for `detectedStack`, Phase 1 refreshes the code graph, queries it, and hands Explore a ranked starting set instead of a full scan; `graph-affected` feeds `analysis.touchedAreas[]`. Zero API cost, read-only. Narrows an open-ended search; it does not replace a grep for a name the task already spells out. Commands, validator contract, measurements: `$HOME/.claude/multi-agent-refs/features/code-graph.md`.
|
|
149
|
+
|
|
146
150
|
#### Step 3 - Codebase Exploration
|
|
147
151
|
|
|
148
152
|
Launch parallel Explore agents to scan codebase:
|
|
@@ -153,6 +157,7 @@ Launch parallel Explore agents to scan codebase:
|
|
|
153
157
|
|
|
154
158
|
Use `subagent_type: "Explore"` with thoroughness scaled to task size AND knowledge availability (first match wins):
|
|
155
159
|
|
|
160
|
+
- A fresh code graph answered the task's query with a ranked file set (Step 2.6) → "light" (the starting set is already narrowed; explore outward from it, do not re-scan)
|
|
156
161
|
- `taskType` is `bugfix`/`chore` AND scope is small (single named file, or a referenced crash/stack frame that pinpoints the site) → "light" (cheapest - scan only the named area + its direct callers)
|
|
157
162
|
- Knowledge exists → "medium" (targeted, cheaper)
|
|
158
163
|
- No knowledge, or `taskType` is `feature`/`refactor`/`component` → "very thorough" (full scan, first-time investment)
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
### Phase 4: Review (deterministic gates + parallel + triage)
|
|
2
2
|
|
|
3
|
-
> **TLDR** - Three-stage review. Stage 1: deterministic gates (build + lint + test + secret scan) that MUST pass. Stage 2:
|
|
3
|
+
> **TLDR** - Three-stage review. Stage 1: deterministic gates (build + lint + test + secret scan) that MUST pass. Stage 2: 3 reviewers in parallel per host, second slot **CLI-aware** (Claude Code Fable + Opus + Sonnet; Copilot CLI GPT-5.4 + Opus + Sonnet). Stage 3: Fable triage (Opus on Copilot CLI) filters raw findings for false positives and out-of-scope items. Only triage-accepted blocking items loop back to Phase 3.
|
|
4
4
|
|
|
5
5
|
<!-- progress-contract: applied -->
|
|
6
6
|
Progress emission per `$HOME/.claude/multi-agent-refs/progress-contract.md` - lines for each gate, each reviewer dispatch + finish, triage start, triage verdict, fix dispatch.
|
|
@@ -275,7 +275,7 @@ Launch Agent instances **in parallel** using the shared `code-reviewer` subagent
|
|
|
275
275
|
| Reviewer 3 | `code-reviewer` | `claude-sonnet-5` | `claude-sonnet-5` | `gpt-5.6` @ `medium` | Quality + correctness + naming | `ai-backend-toolkit:clean-code`, stack-specific skill |
|
|
276
276
|
| Triage | triage persona | `claude-fable-5` | `claude-opus-5` | `gpt-5.6` @ `max` | Filter false positives + out-of-scope | - |
|
|
277
277
|
|
|
278
|
-
Reviewer count per host: **Claude Code 3, Copilot CLI 3, Codex CLI 3
|
|
278
|
+
Reviewer count per host: **Claude Code 3, Copilot CLI 3, Codex CLI 3** - **2** on Claude Code with the fable rung off, per `$HOME/.claude/multi-agent-refs/features/model-fallback.md`.
|
|
279
279
|
|
|
280
280
|
#### Codex CLI - two constraints that fail silently
|
|
281
281
|
|
|
@@ -577,7 +577,7 @@ Opt-in via `prefs.global.triageCrossCheck.enabled` (default `false`). Sampled ru
|
|
|
577
577
|
|
|
578
578
|
After the triage verdict is computed, populate `triage.consensus`:
|
|
579
579
|
|
|
580
|
-
1. `reviewerCount` =
|
|
580
|
+
1. `reviewerCount` = reviewers that actually dispatched, not the configured max: `3` normally, `2` with the fable rung off, lower if one timed out or was skipped.
|
|
581
581
|
2. Classify the iteration `verdict`:
|
|
582
582
|
- `unanimous-block` -> all reviewers returned at least one overlapping `blocking` finding.
|
|
583
583
|
- `split` -> reviewers disagreed on existence or severity of one or more findings (the Step 2.5 disagreement definition). List each split in `disagreements[]` with a `note` naming who held which position (e.g. "Fable blocking, Sonnet approved").
|
|
@@ -286,6 +286,8 @@ Extract reusable knowledge from this task and **append** to `$HOME/.claude/knowl
|
|
|
286
286
|
|
|
287
287
|
**First task** (no knowledge dir): `mkdir -p`, create initial architecture.md + patterns.md from Phase 1. Log: "Knowledge base initialized for {project}"
|
|
288
288
|
|
|
289
|
+
**Code graph refresh:** when `prefs.global.codeGraph.enabled` and `prefs.global.codeGraph.autoRefresh` are both true, rebuild the graph and re-render `GRAPH_REPORT.md` - the branch just changed the code the Phase 1 graph described. Costs no API tokens, never fails the run. Commands: `$HOME/.claude/multi-agent-refs/features/code-graph.md`.
|
|
290
|
+
|
|
289
291
|
**Save preferences**: Update prefs with Phase 7 selections (via channels return: `reportChannels`, `reportContent`, `confluenceUrls` per-project LRU).
|
|
290
292
|
|
|
291
293
|
**Per-repo memory synthesis (opt-in via `prefs.global.perRepoMemory`):**
|
|
@@ -0,0 +1,91 @@
|
|
|
1
|
+
{
|
|
2
|
+
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
|
3
|
+
"$id": "https://github.com/mmerterden/multi-agent-pipeline/schemas/code-graph.schema.json",
|
|
4
|
+
"title": "Code graph",
|
|
5
|
+
"description": "Deterministic, LLM-free code graph produced by graph-build.mjs and consumed by Phase 1 (Explore scope narrowing) and Phase 7 (knowledge capture). One file per project under ~/.claude/knowledge/<project>/code-graph.json.",
|
|
6
|
+
"type": "object",
|
|
7
|
+
"required": ["schemaVersion", "stack", "root", "generatedAt", "stats", "nodes", "edges"],
|
|
8
|
+
"additionalProperties": false,
|
|
9
|
+
"properties": {
|
|
10
|
+
"schemaVersion": {
|
|
11
|
+
"const": "1.0.0"
|
|
12
|
+
},
|
|
13
|
+
"stack": {
|
|
14
|
+
"type": "string",
|
|
15
|
+
"enum": ["ios", "android", "node", "python"]
|
|
16
|
+
},
|
|
17
|
+
"root": {
|
|
18
|
+
"type": "string",
|
|
19
|
+
"description": "Absolute repo root the graph was built from."
|
|
20
|
+
},
|
|
21
|
+
"baseCommit": {
|
|
22
|
+
"type": ["string", "null"],
|
|
23
|
+
"description": "git HEAD at build time. Phase 1 compares it against the current HEAD to decide staleness; null when the root is not a work tree."
|
|
24
|
+
},
|
|
25
|
+
"generatedAt": {
|
|
26
|
+
"type": "string",
|
|
27
|
+
"description": "ISO 8601 build timestamp."
|
|
28
|
+
},
|
|
29
|
+
"stats": {
|
|
30
|
+
"type": "object",
|
|
31
|
+
"required": ["files", "nodes", "edges"],
|
|
32
|
+
"additionalProperties": false,
|
|
33
|
+
"properties": {
|
|
34
|
+
"files": { "type": "integer", "minimum": 0 },
|
|
35
|
+
"nodes": { "type": "integer", "minimum": 0 },
|
|
36
|
+
"edges": { "type": "integer", "minimum": 0 }
|
|
37
|
+
}
|
|
38
|
+
},
|
|
39
|
+
"nodes": {
|
|
40
|
+
"type": "array",
|
|
41
|
+
"items": {
|
|
42
|
+
"type": "object",
|
|
43
|
+
"required": ["id", "kind", "name", "degree"],
|
|
44
|
+
"additionalProperties": false,
|
|
45
|
+
"properties": {
|
|
46
|
+
"id": {
|
|
47
|
+
"type": "string",
|
|
48
|
+
"description": "file:<relpath> | sym:<relpath>#<name> | module:<name>"
|
|
49
|
+
},
|
|
50
|
+
"kind": { "type": "string", "enum": ["file", "symbol", "module"] },
|
|
51
|
+
"name": { "type": "string" },
|
|
52
|
+
"path": { "type": ["string", "null"] },
|
|
53
|
+
"line": { "type": "integer", "minimum": 1 },
|
|
54
|
+
"symbolKind": { "type": "string" },
|
|
55
|
+
"nested": {
|
|
56
|
+
"type": "boolean",
|
|
57
|
+
"description": "Symbol nodes only. true when the declaration is indented, which in every stack these rules cover means it is declared inside another one. A nested symbol keeps its node and its defines edge, so it stays findable by name, but it is never a reference target: sealed cases and inner classes are named after the concept they model (Icon, Color, Success) and each is declared exactly once, so the ambiguity rule does not catch them."
|
|
58
|
+
},
|
|
59
|
+
"isTest": { "type": "boolean" },
|
|
60
|
+
"degree": { "type": "integer", "minimum": 0 }
|
|
61
|
+
}
|
|
62
|
+
}
|
|
63
|
+
},
|
|
64
|
+
"edges": {
|
|
65
|
+
"type": "array",
|
|
66
|
+
"items": {
|
|
67
|
+
"type": "object",
|
|
68
|
+
"required": ["from", "to", "kind"],
|
|
69
|
+
"additionalProperties": false,
|
|
70
|
+
"properties": {
|
|
71
|
+
"from": { "type": "string" },
|
|
72
|
+
"to": { "type": "string" },
|
|
73
|
+
"kind": { "type": "string", "enum": ["defines", "imports", "references"] }
|
|
74
|
+
}
|
|
75
|
+
}
|
|
76
|
+
},
|
|
77
|
+
"manifest": {
|
|
78
|
+
"type": "object",
|
|
79
|
+
"description": "relpath -> {size, mtimeMs}. Staleness signal only: a full rebuild takes seconds at corporate-repo scale, so the pipeline never rebuilds a subset.",
|
|
80
|
+
"additionalProperties": {
|
|
81
|
+
"type": "object",
|
|
82
|
+
"required": ["size", "mtimeMs"],
|
|
83
|
+
"additionalProperties": false,
|
|
84
|
+
"properties": {
|
|
85
|
+
"size": { "type": "integer", "minimum": 0 },
|
|
86
|
+
"mtimeMs": { "type": "integer", "minimum": 0 }
|
|
87
|
+
}
|
|
88
|
+
}
|
|
89
|
+
}
|
|
90
|
+
}
|
|
91
|
+
}
|
|
@@ -548,6 +548,16 @@
|
|
|
548
548
|
"type": "string",
|
|
549
549
|
"description": "Model id used when the premium tier is unavailable."
|
|
550
550
|
},
|
|
551
|
+
"floorModel": {
|
|
552
|
+
"type": ["string", "null"],
|
|
553
|
+
"default": "haiku",
|
|
554
|
+
"description": "Last-resort tier when fallbackModel also fails to dispatch. Set it equal to fallbackModel (or null) to stop after one downgrade step. Documented in multi-agent-refs/features/model-fallback.md since the two-step ladder landed; it was missing here, so a prefs file that followed that doc failed validation on an object that forbids extra keys."
|
|
555
|
+
},
|
|
556
|
+
"fableEnabled": {
|
|
557
|
+
"type": "boolean",
|
|
558
|
+
"default": true,
|
|
559
|
+
"description": "Whether the fable rung is available at all on Claude Code. false makes every persona that declares preferredModel: fable dispatch on opus from the first call, with no dispatch attempt on fable and no error to recover from - a cost control, not a fallback. Claude Code only: Copilot CLI does not offer Fable 5, and on Codex CLI the fable rung means gpt-5.6 @ xhigh, which this switch deliberately does not touch. Turning it off also collapses the Phase 4 Claude Code reviewer panel from three to two (Reviewer 1 lands on opus, which Reviewer 2 already holds), and consensus.reviewerCount records 2."
|
|
560
|
+
},
|
|
551
561
|
"onDispatchError": {
|
|
552
562
|
"type": "boolean",
|
|
553
563
|
"default": true,
|
|
@@ -1192,6 +1202,37 @@
|
|
|
1192
1202
|
}
|
|
1193
1203
|
}
|
|
1194
1204
|
},
|
|
1205
|
+
"codeGraph": {
|
|
1206
|
+
"type": "object",
|
|
1207
|
+
"additionalProperties": false,
|
|
1208
|
+
"description": "v16.1+ - deterministic, LLM-free code graph (symbols, imports, references) built by pipeline/scripts/graph-build.mjs into ~/.claude/knowledge/<project>/code-graph.json. Phase 1 Step 2.6 queries it to narrow the Explore starting set; Phase 7 Step 3 refreshes it after the branch changed code. Zero API cost, zero runtime dependencies (ADR-0004), read-only on the repo. Off by default: with it off the pipeline behaves exactly as it did before. Measured on a 4,300-file Swift app at a fixed 30k retrieval budget, it roughly doubled coverage at under half the token cost on domain-word searches, and was marginally worse than plain grep when the task already names an exact symbol.",
|
|
1209
|
+
"properties": {
|
|
1210
|
+
"enabled": {
|
|
1211
|
+
"type": "boolean",
|
|
1212
|
+
"default": false,
|
|
1213
|
+
"description": "Master switch. When false, Phase 1 Step 2.6 and the Phase 7 refresh are skipped entirely and no graph file is written."
|
|
1214
|
+
},
|
|
1215
|
+
"autoRefresh": {
|
|
1216
|
+
"type": "boolean",
|
|
1217
|
+
"default": true,
|
|
1218
|
+
"description": "Rebuild the graph in Phase 7 after the branch changed code. When false, the graph only refreshes when someone asks for it via /multi-agent:graph."
|
|
1219
|
+
},
|
|
1220
|
+
"maxNodes": {
|
|
1221
|
+
"type": "integer",
|
|
1222
|
+
"minimum": 1000,
|
|
1223
|
+
"maximum": 500000,
|
|
1224
|
+
"default": 200000,
|
|
1225
|
+
"description": "Upper bound on graph size. A build past this is reported and the graph is not written, so an unexpectedly huge tree (vendored sources, generated code) fails loudly instead of writing a multi-hundred-MB file."
|
|
1226
|
+
},
|
|
1227
|
+
"queryBudget": {
|
|
1228
|
+
"type": "integer",
|
|
1229
|
+
"minimum": 200,
|
|
1230
|
+
"maximum": 20000,
|
|
1231
|
+
"default": 2000,
|
|
1232
|
+
"description": "Token cap on a single graph-query traversal. The walk stops on budget, not on depth alone - an unbounded neighbourhood dump is the context-stuffing pattern this replaces."
|
|
1233
|
+
}
|
|
1234
|
+
}
|
|
1235
|
+
},
|
|
1195
1236
|
"devCritic": {
|
|
1196
1237
|
"type": "object",
|
|
1197
1238
|
"additionalProperties": false,
|
|
@@ -1433,6 +1474,10 @@
|
|
|
1433
1474
|
"type": "string",
|
|
1434
1475
|
"description": "Repo-relative directory holding our derived copy."
|
|
1435
1476
|
},
|
|
1477
|
+
"localRepo": {
|
|
1478
|
+
"type": "string",
|
|
1479
|
+
"description": "Repo that localPath is relative to, e.g. $HOME/multi-agent-plugins. Without it localPath reads as relative to the current repo, which is how both entries came to point at the UPSTREAM tree and stayed wrong through a rename: nothing in the config said which checkout held the derived copies."
|
|
1480
|
+
},
|
|
1436
1481
|
"upstreamMarketplace": {
|
|
1437
1482
|
"type": "string",
|
|
1438
1483
|
"description": "Installed marketplace name, resolved under ~/.claude/plugins/cache/<marketplace>/."
|
|
@@ -36,6 +36,6 @@
|
|
|
36
36
|
"warn_tokens": 5600
|
|
37
37
|
}
|
|
38
38
|
},
|
|
39
|
-
"total_max_tokens":
|
|
40
|
-
"note": "Token estimate = ceil(chars / 4). Per-phase budget rule: warn = current+10% (rounded to nearest 50), max = current+25%. Gives ~6 edit cycles of headroom before warn trips - intentionally quiet under normal maintenance, loud when a phase grows unusually. Only the active phase is loaded (lazy). Recalibrated at v10.0.0 after the validator/consistency/simplifier/lesson gate contracts landed in phases 1-4. Recalibrated again at v10.9.0 after the verify-by-test (Phase 4 Step 3.7), update-check (Phase 0 Step 0.6), immutable-test (Phase 3 GREEN) and redTests re-entry contracts landed - Step 3.7 prose was compressed to a pointer into refs/features/verify-by-test.md before the recalibration. Total bumped 50000 -> 51000 at v12.5.0 after the worktree residue/traversal-prune contract (Phase 0 + Phase 5 heal) and the Reflexion causal-diagnosis contract (Phase 4 lesson memory) landed; the prose was compressed first (161 tokens reclaimed) and every per-phase max still passes - only the aggregate needed room. Recalibrated again at v13.6.0 after the install-relative path correction: an instruction that names `pipeline/scripts/x` resolves only from a repo checkout, and a run happens in the user's worktree, so 157 references across these docs moved to `$HOME/.claude/...` at +5 bytes each - 196 tokens of pure correctness cost. Same discipline as before: prose was compressed FIRST (149 tokens reclaimed, by pointing Phase 1's Figma tier table at the Phase 0 probe that already resolved it and Phase 4's Codex constraints at the always-loaded AGENTS.md block), and only then were the budgets moved. Five warn lines had been permanently amber, which makes the amber tier useless as a signal, so every warn was reset to the documented current+10% and the four maxes that the new warn would have collided with were reset to current+25%. Aggregate 51000 -> 51500. Total bumped 51500 -> 52200 at v14.0.0 after Phase 4 Review entered the four --dev mode phase sets and the criteria-resolution contract (Step 1.78) landed. Same discipline as every prior bump: prose was compressed FIRST, 820 tokens reclaimed, before the number moved. Two of those compressions are structural rather than cosmetic - the hardcoded SwiftUI interaction list in Step 1.5 and the SwiftUI convention paragraph in Step 2.8 were transcriptions of rules that now live in a scoped registry, so keeping them here would have re-created the drift this release exists to remove, and the third moved the Step 1.78 full contract into refs/features/skill-conformance.md leaving a pointer. What remains is contract text that cannot be inferred: the manifest's four consumer-visible parts, the conformance checklist the reviewers must return, and the fail-closed semantics. Every per-phase max still passes (phase-4 12405/14750); only the aggregate needed room. Total bumped 52200 -> 52700 at v14.1.0 after two more contracts landed: stack skill routing (Phase 3 pre-flight step 9) and worktree finalize (Phase 6 step 9). Compression came first, as always, and twice: 224 tokens out of Phase 3 by pointing its criteria-ledger and routing steps at their feature files instead of restating them, and 190 out of Phase 6 by moving the finalize contract into refs/features/worktree-finalize.md and leaving the invocation plus the exit-3 semantics. Both new contracts follow the pattern the earlier ones set: the phase doc carries the call and the decision, the feature file carries the reasoning, and the feature files are outside this budget because it loops only the eight phase-N-* keys. Every per-phase max still passes (phase-3 7677/8950, phase-6 5223/6150 and both under warn); only the aggregate needed room. Total bumped 52700 -> 52750 for the Phase 0 Step 3 branch-persistence correction: the step wrote the legacy `projects[].branches` while the TTL filter two sections below read `global.recentBranches`, and both spots named a `{name, lastUsed}` shape the schema rejects (`branch` required, `additionalProperties: false`), so the recent-branch picker option could never populate and a literal implementation would have failed prefs validation. Naming the right target, the right key and the legacy field to avoid costs 41 tokens over the one line it replaces. Compression came first and was applied three times to the replacement text itself, from 120 tokens down to 66, by moving the rationale out of the phase doc entirely: the reasoning now lives where it is enforced, in the migrate-prefs carry-forward comment and the smoke-pref-migration f7 block, leaving the phase doc with only the instruction. 50 was the smallest step that clears it; phase-0-init sits at 10893/12400, far under its own max, so this is purely an aggregate ceiling. v15.0.0: total 52750 -> 53100, the stack-skill tables in phase-1/2/4 now carry plugin-namespaced names (ai-<stack>-toolkit:<skill>) - functional prefixes, ~170 tokens. v15.10.0: total 53350 -> 53950 for the memory-recall + context-offload contracts (Phase 1 two-block durable-knowledge injection and its telemetry, Phase 3 build-log offload pipe, Phase 4 ranked prior art, offload pipe and recall telemetry). Compression came first and twice, taking the new prose from 1168 tokens to 580: the reasoning behind the two blocks lives in multi-agent-refs/prompt-assembly.md and the reasoning behind the offload filter lives in the offload-ref.sh header, both outside this budget, so the phase docs carry only the call, the pref that gates it and the one fact an agent cannot infer - that the evidence gate still reads the whole build log, so offloading changes what is read, never what counts as a verified pass. Every per-phase max still passes (phase-3 7985/8950, phase-4 12997/14750); phase-3 and phase-4 crossed their warn lines and are left amber on purpose, because that is the signal that those two docs are the next ones needing structural compression rather than another bump. v15.13.0: total 53950 -> 54050 for the prefs-to-flag bridges. Five settings had shipped declared-but-inert: contextOffload.minLines and .tailLines (fixed in 15.11.0), learningsLedger.maxBriefEntries, and testGap.scanTree and .promoteSeverity - the last two declared in the schema AND implemented as flags in the scanner, with nothing in between reading the pref and passing the flag. Wiring three of them costs the phase docs 94 tokens, which is the wiring itself and not prose: two `--max` substitutions and a three-line GAP_FLAGS block. Compression came first and twice, as always: the rationale that would have sat in phase-5 now lives in the header of smoke-prefs-consumed.sh, the gate that makes this class fail a build instead of shipping, and a `--severity-promote` table row was dropped because the invocation above it now shows the flag and names the pref that triggers it, which the row did not. 100 was the smallest step that clears it. Every per-phase max still passes; phase-3 and phase-4 remain amber on purpose. v15.14.0: total 54050 -> 54400 for the supported-version gate. Phase 0 Step 0.6 stopped being purely advisory: a release can now publish an npm dist-tag `required` that names the oldest runnable version, and below it the run halts instead of nagging. What the phase doc has to carry is the part an agent cannot infer - the third stdout field, that the halt is identical in autopilot, and that the run must NOT continue on the freshly updated install because its docs were already loaded from the old version. Compression came first, as always, and took the new prose from 469 tokens to 337: the rationale for the floor, the exemption list, the fail-open rules and the `npm dist-tag add` recipe all moved to multi-agent-refs/rules.md \"Supported Version Gate\" (loaded by 25 commands, outside this budget) and to the header of require-supported-version.sh, leaving the phase doc with the call, the decision table and the halt. 350 was the smallest step that clears it. Every per-phase max still passes (phase-0-init 11230/12400); phase-3 and phase-4 remain amber on purpose. v15.17.0: total 54400 -> 54900 for the Phase 1 analysis-document step. Phase 2 and Phase 3 pre-flights had BLOCKED on `analysis/<feature>-<platform>.md` since v9.0.0 while nothing produced it, so a full run either aborted at Phase 2 or the model ignored its own BLOCKING contract; Step 4 is the producer. What the phase doc carries is only what cannot be inferred: the when-table (taskType x Figma reference), the four refs in load order, the two artefacts, and that the doc validator fails closed. Compression came first and took the step from 745 tokens to 497: the history of why the gap existed moved to the CHANGELOG, the per-ref one-line descriptions moved into the refs' own headers, and the autopilot carve-out collapsed to one clause. The 17.4k-token analysis engine itself is NOT in this budget - it moved out of commands/ into multi-agent-refs/analysis/{locked,evidence,synthesis,render}.md, loaded on demand, which also took analysis/SKILL.md from 18081 to 5974 tokens and retired its lint grace entry. 500 was the smallest step that clears it; phase-1-analysis sits at 4338/4600 and is amber on purpose, like phase-3 and phase-4. v15.18.0: total 54900 -> 55250 for analysis mode. Three phase docs gained a mode branch that cannot be inferred: Phase 4 reviews a document instead of a diff (validator, the one question reviewers answer, the open-question walk), and Phase 6 publishes instead of committing. Compression came first and was applied twice to the new prose and once to old: the Phase 4 branch went from 320 tokens to 180 and the Phase 6 branch from 190 to 120 by pointing at multi-agent-refs/analysis/{resolve,render}.md, which now hold the walks themselves, and the front-matter parse contract stopped being spelled out in both pre-flights. The analysis engine keeps leaving this budget rather than entering it: intake joined locked/evidence/synthesis/render/resolve in multi-agent-refs/analysis/, which is what let analysis/SKILL.md drop under the 6000 hard cap after its grace entry was retired. 350 was the smallest step that clears it; phase-4 and phase-6 are amber on purpose, as phase-1 and phase-3 already were. v15.20.0: total 55250 -> 55500 for the TDD bridge. Phase 3 pre-flight read the analysis doc's concept table and even said test method names come from it, while nothing read Section 15 - so the RED step invented tests and the analysis test matrix never reached development. Phase 3 step 5b now loads it into state.dev.testPlan[] and Phase 4 step 1.45 cross-checks every planned row against a real test, which is what turns \"analysis quality is output quality\" from a slogan into a finding. Compression came first on both blocks, 300 tokens down to 175, by dropping the enumerated failure modes to one line each and the rationale to one clause; the reasoning lives in the CHANGELOG. 250 was the smallest step that clears it. v15.21.0: total 55500 -> 55800 for the post-analysis confirmation. Phase 2 gained Step 0.9, the last human checkpoint before Phase 3: derived values are shown for confirmation and only Section 20 rows are asked, through the resolve engine that already exists in refs. It belongs here rather than Phase 4 because Phase 4 runs after development, where an answer arrives too late to change anything. Compression came first and twice, 430 tokens down to 250, by collapsing the derived-vs-asked explanation to one sentence each and moving the walk itself to multi-agent-refs/analysis/resolve.md, which Phase 4 and analysis-resolve already mount. 300 was the smallest step that clears it. v15.22.0: total 55800 -> 55900 for the analyst-toolkit hooks. Phase 1 Step 4 now names the two prefs that decide whether a document is produced at all and how deep it goes (forceFull, mode) - the first of those had shipped declared-but-inert and smoke-prefs-consumed caught it - and Phase 4 triage gained one clause: a finding that blames a third-party library asks evidence-github whether it is already open upstream, which turns it into a deferred item with a citation instead of Phase 3 rework on code that is not ours. Compression came first and three times, taking the new prose from 220 tokens to 110, and the Phase 1d evidence contract itself never entered this budget - it lives in multi-agent-refs/analysis/evidence.md beside the phases it belongs to. 100 was the smallest step that clears it, leaving 34 tokens of headroom. phase-4 stays amber and the debt named at v15.10.0 stands: it is the doc that needs structural compression rather than another bump, and the two candidates are the inline triage JSON shape and the 3.4 telemetry block, both of which restate something already authoritative elsewhere. v16.0.0: total 55900 -> 56350 for the depth picker. `--dev` and the four dev-* commands are gone; depth is Phase 0 Step 7.5, which costs phase-0-init a step it did not have. Compression came first and three times, taking the step from 530 tokens to 300: the question wording, the per-taskType recommendation and the mode tables all live in phases/modes.md (outside this budget), so the phase doc carries only what an agent cannot infer - that the step runs after Step 7 and why, who is exempt, that ASK_CHOICE_DEFAULT must be passed explicitly because ask-choice.sh takes the FIRST option on a non-TTY, and that Short flips the Phase 1/2 tiles late rather than pre-marking them. The phase-4 telemetry block named as compression debt at v15.22.0 was collapsed to an emit() helper (-27) and the four dev-* mode files left the tree entirely, but neither offsets a genuinely new phase step. 450 was the smallest step that clears it, leaving 119 tokens of headroom. phase-4 remains amber and its other named candidate, the inline triage JSON shape, was left alone on purpose: it is the prompt the triage agent is handed, not a restatement for readers. v16.2.0: total 56350 -> 56600 for the spec-freshness and reuse-tag contracts. Phase 3 step 3 had compared `state.run.lastAnalysisDigest` since it was written, against a key nothing ever set and that the state schema did not declare, so the staleness branch was unreachable and every run reported fresh by default. Phase 1 now persists the digest and a `base_commit` anchor, and step 3 gained the repo-drift half the digest cannot see: a reused document keeps a matching digest precisely because its evidence inputs did not change, while the code underneath it moved. The second contract is the Section 14 tag reaching development: Phase 2 carries it onto the todo as `sourceTag` and Phase 3 treats it as an instruction, which is what stops a Reuse row from being re-implemented. Compression came first and took the four additions from 380 tokens to 214, by moving every rationale clause out of the phase docs: why the commit anchor exists rather than a digest recomputation lives in this note and the CHANGELOG, and the schema descriptions carry the field semantics. The baseline had 9 tokens of headroom, so no addition of any size could have fit without a bump. 250 was the smallest step that clears it, leaving 45 tokens. phase-3 and phase-4 remain amber."
|
|
39
|
+
"total_max_tokens": 57700,
|
|
40
|
+
"note": "Token estimate = ceil(chars / 4). Per-phase budget rule: warn = current+10% (rounded to nearest 50), max = current+25%. Gives ~6 edit cycles of headroom before warn trips - intentionally quiet under normal maintenance, loud when a phase grows unusually. Only the active phase is loaded (lazy). Recalibrated at v10.0.0 after the validator/consistency/simplifier/lesson gate contracts landed in phases 1-4. Recalibrated again at v10.9.0 after the verify-by-test (Phase 4 Step 3.7), update-check (Phase 0 Step 0.6), immutable-test (Phase 3 GREEN) and redTests re-entry contracts landed - Step 3.7 prose was compressed to a pointer into refs/features/verify-by-test.md before the recalibration. Total bumped 50000 -> 51000 at v12.5.0 after the worktree residue/traversal-prune contract (Phase 0 + Phase 5 heal) and the Reflexion causal-diagnosis contract (Phase 4 lesson memory) landed; the prose was compressed first (161 tokens reclaimed) and every per-phase max still passes - only the aggregate needed room. Recalibrated again at v13.6.0 after the install-relative path correction: an instruction that names `pipeline/scripts/x` resolves only from a repo checkout, and a run happens in the user's worktree, so 157 references across these docs moved to `$HOME/.claude/...` at +5 bytes each - 196 tokens of pure correctness cost. Same discipline as before: prose was compressed FIRST (149 tokens reclaimed, by pointing Phase 1's Figma tier table at the Phase 0 probe that already resolved it and Phase 4's Codex constraints at the always-loaded AGENTS.md block), and only then were the budgets moved. Five warn lines had been permanently amber, which makes the amber tier useless as a signal, so every warn was reset to the documented current+10% and the four maxes that the new warn would have collided with were reset to current+25%. Aggregate 51000 -> 51500. Total bumped 51500 -> 52200 at v14.0.0 after Phase 4 Review entered the four --dev mode phase sets and the criteria-resolution contract (Step 1.78) landed. Same discipline as every prior bump: prose was compressed FIRST, 820 tokens reclaimed, before the number moved. Two of those compressions are structural rather than cosmetic - the hardcoded SwiftUI interaction list in Step 1.5 and the SwiftUI convention paragraph in Step 2.8 were transcriptions of rules that now live in a scoped registry, so keeping them here would have re-created the drift this release exists to remove, and the third moved the Step 1.78 full contract into refs/features/skill-conformance.md leaving a pointer. What remains is contract text that cannot be inferred: the manifest's four consumer-visible parts, the conformance checklist the reviewers must return, and the fail-closed semantics. Every per-phase max still passes (phase-4 12405/14750); only the aggregate needed room. Total bumped 52200 -> 52700 at v14.1.0 after two more contracts landed: stack skill routing (Phase 3 pre-flight step 9) and worktree finalize (Phase 6 step 9). Compression came first, as always, and twice: 224 tokens out of Phase 3 by pointing its criteria-ledger and routing steps at their feature files instead of restating them, and 190 out of Phase 6 by moving the finalize contract into refs/features/worktree-finalize.md and leaving the invocation plus the exit-3 semantics. Both new contracts follow the pattern the earlier ones set: the phase doc carries the call and the decision, the feature file carries the reasoning, and the feature files are outside this budget because it loops only the eight phase-N-* keys. Every per-phase max still passes (phase-3 7677/8950, phase-6 5223/6150 and both under warn); only the aggregate needed room. Total bumped 52700 -> 52750 for the Phase 0 Step 3 branch-persistence correction: the step wrote the legacy `projects[].branches` while the TTL filter two sections below read `global.recentBranches`, and both spots named a `{name, lastUsed}` shape the schema rejects (`branch` required, `additionalProperties: false`), so the recent-branch picker option could never populate and a literal implementation would have failed prefs validation. Naming the right target, the right key and the legacy field to avoid costs 41 tokens over the one line it replaces. Compression came first and was applied three times to the replacement text itself, from 120 tokens down to 66, by moving the rationale out of the phase doc entirely: the reasoning now lives where it is enforced, in the migrate-prefs carry-forward comment and the smoke-pref-migration f7 block, leaving the phase doc with only the instruction. 50 was the smallest step that clears it; phase-0-init sits at 10893/12400, far under its own max, so this is purely an aggregate ceiling. v15.0.0: total 52750 -> 53100, the stack-skill tables in phase-1/2/4 now carry plugin-namespaced names (ai-<stack>-toolkit:<skill>) - functional prefixes, ~170 tokens. v15.10.0: total 53350 -> 53950 for the memory-recall + context-offload contracts (Phase 1 two-block durable-knowledge injection and its telemetry, Phase 3 build-log offload pipe, Phase 4 ranked prior art, offload pipe and recall telemetry). Compression came first and twice, taking the new prose from 1168 tokens to 580: the reasoning behind the two blocks lives in multi-agent-refs/prompt-assembly.md and the reasoning behind the offload filter lives in the offload-ref.sh header, both outside this budget, so the phase docs carry only the call, the pref that gates it and the one fact an agent cannot infer - that the evidence gate still reads the whole build log, so offloading changes what is read, never what counts as a verified pass. Every per-phase max still passes (phase-3 7985/8950, phase-4 12997/14750); phase-3 and phase-4 crossed their warn lines and are left amber on purpose, because that is the signal that those two docs are the next ones needing structural compression rather than another bump. v15.13.0: total 53950 -> 54050 for the prefs-to-flag bridges. Five settings had shipped declared-but-inert: contextOffload.minLines and .tailLines (fixed in 15.11.0), learningsLedger.maxBriefEntries, and testGap.scanTree and .promoteSeverity - the last two declared in the schema AND implemented as flags in the scanner, with nothing in between reading the pref and passing the flag. Wiring three of them costs the phase docs 94 tokens, which is the wiring itself and not prose: two `--max` substitutions and a three-line GAP_FLAGS block. Compression came first and twice, as always: the rationale that would have sat in phase-5 now lives in the header of smoke-prefs-consumed.sh, the gate that makes this class fail a build instead of shipping, and a `--severity-promote` table row was dropped because the invocation above it now shows the flag and names the pref that triggers it, which the row did not. 100 was the smallest step that clears it. Every per-phase max still passes; phase-3 and phase-4 remain amber on purpose. v15.14.0: total 54050 -> 54400 for the supported-version gate. Phase 0 Step 0.6 stopped being purely advisory: a release can now publish an npm dist-tag `required` that names the oldest runnable version, and below it the run halts instead of nagging. What the phase doc has to carry is the part an agent cannot infer - the third stdout field, that the halt is identical in autopilot, and that the run must NOT continue on the freshly updated install because its docs were already loaded from the old version. Compression came first, as always, and took the new prose from 469 tokens to 337: the rationale for the floor, the exemption list, the fail-open rules and the `npm dist-tag add` recipe all moved to multi-agent-refs/rules.md \"Supported Version Gate\" (loaded by 25 commands, outside this budget) and to the header of require-supported-version.sh, leaving the phase doc with the call, the decision table and the halt. 350 was the smallest step that clears it. Every per-phase max still passes (phase-0-init 11230/12400); phase-3 and phase-4 remain amber on purpose. v15.17.0: total 54400 -> 54900 for the Phase 1 analysis-document step. Phase 2 and Phase 3 pre-flights had BLOCKED on `analysis/<feature>-<platform>.md` since v9.0.0 while nothing produced it, so a full run either aborted at Phase 2 or the model ignored its own BLOCKING contract; Step 4 is the producer. What the phase doc carries is only what cannot be inferred: the when-table (taskType x Figma reference), the four refs in load order, the two artefacts, and that the doc validator fails closed. Compression came first and took the step from 745 tokens to 497: the history of why the gap existed moved to the CHANGELOG, the per-ref one-line descriptions moved into the refs' own headers, and the autopilot carve-out collapsed to one clause. The 17.4k-token analysis engine itself is NOT in this budget - it moved out of commands/ into multi-agent-refs/analysis/{locked,evidence,synthesis,render}.md, loaded on demand, which also took analysis/SKILL.md from 18081 to 5974 tokens and retired its lint grace entry. 500 was the smallest step that clears it; phase-1-analysis sits at 4338/4600 and is amber on purpose, like phase-3 and phase-4. v15.18.0: total 54900 -> 55250 for analysis mode. Three phase docs gained a mode branch that cannot be inferred: Phase 4 reviews a document instead of a diff (validator, the one question reviewers answer, the open-question walk), and Phase 6 publishes instead of committing. Compression came first and was applied twice to the new prose and once to old: the Phase 4 branch went from 320 tokens to 180 and the Phase 6 branch from 190 to 120 by pointing at multi-agent-refs/analysis/{resolve,render}.md, which now hold the walks themselves, and the front-matter parse contract stopped being spelled out in both pre-flights. The analysis engine keeps leaving this budget rather than entering it: intake joined locked/evidence/synthesis/render/resolve in multi-agent-refs/analysis/, which is what let analysis/SKILL.md drop under the 6000 hard cap after its grace entry was retired. 350 was the smallest step that clears it; phase-4 and phase-6 are amber on purpose, as phase-1 and phase-3 already were. v15.20.0: total 55250 -> 55500 for the TDD bridge. Phase 3 pre-flight read the analysis doc's concept table and even said test method names come from it, while nothing read Section 15 - so the RED step invented tests and the analysis test matrix never reached development. Phase 3 step 5b now loads it into state.dev.testPlan[] and Phase 4 step 1.45 cross-checks every planned row against a real test, which is what turns \"analysis quality is output quality\" from a slogan into a finding. Compression came first on both blocks, 300 tokens down to 175, by dropping the enumerated failure modes to one line each and the rationale to one clause; the reasoning lives in the CHANGELOG. 250 was the smallest step that clears it. v15.21.0: total 55500 -> 55800 for the post-analysis confirmation. Phase 2 gained Step 0.9, the last human checkpoint before Phase 3: derived values are shown for confirmation and only Section 20 rows are asked, through the resolve engine that already exists in refs. It belongs here rather than Phase 4 because Phase 4 runs after development, where an answer arrives too late to change anything. Compression came first and twice, 430 tokens down to 250, by collapsing the derived-vs-asked explanation to one sentence each and moving the walk itself to multi-agent-refs/analysis/resolve.md, which Phase 4 and analysis-resolve already mount. 300 was the smallest step that clears it. v15.22.0: total 55800 -> 55900 for the analyst-toolkit hooks. Phase 1 Step 4 now names the two prefs that decide whether a document is produced at all and how deep it goes (forceFull, mode) - the first of those had shipped declared-but-inert and smoke-prefs-consumed caught it - and Phase 4 triage gained one clause: a finding that blames a third-party library asks evidence-github whether it is already open upstream, which turns it into a deferred item with a citation instead of Phase 3 rework on code that is not ours. Compression came first and three times, taking the new prose from 220 tokens to 110, and the Phase 1d evidence contract itself never entered this budget - it lives in multi-agent-refs/analysis/evidence.md beside the phases it belongs to. 100 was the smallest step that clears it, leaving 34 tokens of headroom. phase-4 stays amber and the debt named at v15.10.0 stands: it is the doc that needs structural compression rather than another bump, and the two candidates are the inline triage JSON shape and the 3.4 telemetry block, both of which restate something already authoritative elsewhere. v16.0.0: total 55900 -> 56350 for the depth picker. `--dev` and the four dev-* commands are gone; depth is Phase 0 Step 7.5, which costs phase-0-init a step it did not have. Compression came first and three times, taking the step from 530 tokens to 300: the question wording, the per-taskType recommendation and the mode tables all live in phases/modes.md (outside this budget), so the phase doc carries only what an agent cannot infer - that the step runs after Step 7 and why, who is exempt, that ASK_CHOICE_DEFAULT must be passed explicitly because ask-choice.sh takes the FIRST option on a non-TTY, and that Short flips the Phase 1/2 tiles late rather than pre-marking them. The phase-4 telemetry block named as compression debt at v15.22.0 was collapsed to an emit() helper (-27) and the four dev-* mode files left the tree entirely, but neither offsets a genuinely new phase step. 450 was the smallest step that clears it, leaving 119 tokens of headroom. phase-4 remains amber and its other named candidate, the inline triage JSON shape, was left alone on purpose: it is the prompt the triage agent is handed, not a restatement for readers. v16.2.0: total 56350 -> 56600 for the spec-freshness and reuse-tag contracts. Phase 3 step 3 had compared `state.run.lastAnalysisDigest` since it was written, against a key nothing ever set and that the state schema did not declare, so the staleness branch was unreachable and every run reported fresh by default. Phase 1 now persists the digest and a `base_commit` anchor, and step 3 gained the repo-drift half the digest cannot see: a reused document keeps a matching digest precisely because its evidence inputs did not change, while the code underneath it moved. The second contract is the Section 14 tag reaching development: Phase 2 carries it onto the todo as `sourceTag` and Phase 3 treats it as an instruction, which is what stops a Reuse row from being re-implemented. Compression came first and took the four additions from 380 tokens to 214, by moving every rationale clause out of the phase docs: why the commit anchor exists rather than a digest recomputation lives in this note and the CHANGELOG, and the schema descriptions carry the field semantics. The baseline had 9 tokens of headroom, so no addition of any size could have fit without a bump. 250 was the smallest step that clears it, leaving 45 tokens. phase-3 and phase-4 remain amber. v16.13.0: total 57600 -> 57700 for the code-graph injection and the fable-rung switch. Phase 1 gained Step 2.6 (query the graph, hand Explore a ranked starting set), Phase 7 gained the post-branch graph refresh, and Phase 0 Step 0 gained one line: a prefs switch that resolves every preferredModel: fable persona to opus for the run, which also collapses the Phase 4 Claude Code panel from three reviewers to two. Compression came first and mostly structurally: of roughly 1,630 tokens of new contract text, 1,310 never entered this budget at all - the whole code-graph contract lives in multi-agent-refs/features/code-graph.md (604) and the fable switch's scope table, per-host effects and cost-accounting consequence live in features/model-fallback.md (+707), leaving the phase docs with the call, the pref that gates it and the one fact an agent cannot infer. Phase 4 was compressed on top of that: its TLDR restated the reviewer matrix 270 lines below it, so 36 tokens came back and the doc nets +6 despite carrying two new clauses. One of those clauses is a correction rather than a feature - the consensus rule still said reviewerCount is 2 on Claude Code, which stopped being true when the third reviewer landed in 16.12.0, and the cross-CLI smoke never caught it because it reads the matrix line instead. 100 was the smallest step that clears it, leaving 54 tokens. phase-3 and phase-4 remain amber."
|
|
41
41
|
}
|