@mmerterden/multi-agent-pipeline 17.0.0 → 17.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +159 -0
- package/README.md +56 -4
- package/README.tr.md +57 -4
- package/docs/architecture.md +3 -3
- package/docs/ecosystem.md +5 -5
- package/docs/token-budget-history.md +22 -0
- package/install/_dev-only-files.mjs +1 -0
- package/install/codex.mjs +18 -1
- package/install/copilot.mjs +17 -1
- package/install/templates/multi-agent-autopilot.plist.template +79 -0
- package/package.json +1 -1
- package/pipeline/commands/multi-agent/autopilot-off/SKILL.md +64 -0
- package/pipeline/commands/multi-agent/autopilot-on/SKILL.md +181 -0
- package/pipeline/commands/multi-agent/autopilot-status/SKILL.md +74 -0
- package/pipeline/commands/multi-agent/channels/SKILL.md +41 -12
- package/pipeline/commands/multi-agent/help/SKILL.md +41 -35
- package/pipeline/commands/multi-agent/manual-test/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/sync/SKILL.md +10 -9
- package/pipeline/commands/multi-agent/update/SKILL.md +1 -1
- package/pipeline/lib/autopilot-activation.sh +117 -0
- package/pipeline/lib/autopilot-state.sh +184 -0
- package/pipeline/lib/issue-fetcher.sh +18 -1
- package/pipeline/lib/plan-todos.sh +18 -0
- package/pipeline/multi-agent-refs/_dev-context.md +10 -0
- package/pipeline/multi-agent-refs/analysis/redesign.md +8 -0
- package/pipeline/multi-agent-refs/analysis/review.md +9 -0
- package/pipeline/multi-agent-refs/android-guide.md +14 -0
- package/pipeline/multi-agent-refs/audit-guide.md +12 -0
- package/pipeline/multi-agent-refs/backend-guide.md +10 -0
- package/pipeline/multi-agent-refs/channels/confluence.md +11 -0
- package/pipeline/multi-agent-refs/channels/issue-comment.md +12 -0
- package/pipeline/multi-agent-refs/channels/jira.md +90 -20
- package/pipeline/multi-agent-refs/channels/pr-review-actions.md +13 -0
- package/pipeline/multi-agent-refs/channels/pr.md +76 -19
- package/pipeline/multi-agent-refs/component-dispatch.md +11 -0
- package/pipeline/multi-agent-refs/component-generation.md +11 -0
- package/pipeline/multi-agent-refs/conventions-defaults.md +15 -0
- package/pipeline/multi-agent-refs/cross-cli-contract.md +49 -5
- package/pipeline/multi-agent-refs/features/analysis-jira.md +11 -0
- package/pipeline/multi-agent-refs/features/design-conformance.md +10 -0
- package/pipeline/multi-agent-refs/features/doctor.md +10 -0
- package/pipeline/multi-agent-refs/features/external-context-injection.md +7 -0
- package/pipeline/multi-agent-refs/features/jira-context.md +9 -0
- package/pipeline/multi-agent-refs/features/model-fallback.md +10 -0
- package/pipeline/multi-agent-refs/features/skill-conformance.md +13 -0
- package/pipeline/multi-agent-refs/features/url-enrichment.md +9 -0
- package/pipeline/multi-agent-refs/features/visual-evidence.md +61 -7
- package/pipeline/multi-agent-refs/generate-issue.md +7 -0
- package/pipeline/multi-agent-refs/issue-jira-triad.md +9 -0
- package/pipeline/multi-agent-refs/knowledge.md +6 -0
- package/pipeline/multi-agent-refs/multi-repo-integration-build.md +13 -0
- package/pipeline/multi-agent-refs/phases/modes.md +7 -0
- package/pipeline/multi-agent-refs/phases/operations.md +9 -0
- package/pipeline/multi-agent-refs/phases/phase-0-init.md +2 -2
- package/pipeline/multi-agent-refs/phases/phase-2-planning.md +17 -15
- package/pipeline/multi-agent-refs/phases/phase-3-dev.md +1 -1
- package/pipeline/multi-agent-refs/phases/phase-6-commit.md +1 -1
- package/pipeline/multi-agent-refs/phases.md +11 -0
- package/pipeline/multi-agent-refs/picker-contract.md +12 -0
- package/pipeline/multi-agent-refs/platform-parity.md +10 -0
- package/pipeline/multi-agent-refs/progress-contract.md +10 -0
- package/pipeline/multi-agent-refs/readiness-review.md +7 -1
- package/pipeline/multi-agent-refs/rules.md +3 -11
- package/pipeline/multi-agent-refs/setup/firebase.md +9 -0
- package/pipeline/multi-agent-refs/swiftui-guide.md +17 -0
- package/pipeline/multi-agent-refs/tracker-contract.md +44 -0
- package/pipeline/multi-agent-refs/web-guide.md +10 -0
- package/pipeline/multi-agent-refs/wiki-capture.md +11 -0
- package/pipeline/schemas/autopilot-config.schema.json +149 -0
- package/pipeline/schemas/prefs.schema.json +4 -0
- package/pipeline/schemas/token-budget.json +10 -19
- package/pipeline/scripts/autopilot-arming.mjs +147 -0
- package/pipeline/scripts/autopilot-intake.mjs +387 -0
- package/pipeline/scripts/autopilot-menubar.swift +361 -0
- package/pipeline/scripts/autopilot-runner.mjs +354 -0
- package/pipeline/scripts/autopilot-status.sh +213 -0
- package/pipeline/scripts/capture-evidence.sh +79 -11
- package/pipeline/scripts/gen-ref-toc.mjs +279 -0
- package/pipeline/scripts/jira-search.sh +70 -0
- package/pipeline/scripts/phase-tracker.sh +134 -12
- package/pipeline/scripts/probe-evidence-capability.sh +27 -3
- package/pipeline/scripts/run-ui-tests.sh +113 -4
- package/pipeline/skills/.skill-manifest.json +16 -4
- package/pipeline/skills/shared/core/multi-agent-autopilot-off/SKILL.md +67 -0
- package/pipeline/skills/shared/core/multi-agent-autopilot-on/SKILL.md +146 -0
- package/pipeline/skills/shared/core/multi-agent-autopilot-status/SKILL.md +64 -0
- package/pipeline/skills/shared/core/multi-agent-channels/SKILL.md +62 -11
- package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +9 -8
|
@@ -7,7 +7,7 @@ Progress emission per `$HOME/.claude/multi-agent-refs/progress-contract.md` -
|
|
|
7
7
|
|
|
8
8
|
## Phase 2 Pre-flight (BLOCKING, v9.0.0)
|
|
9
9
|
|
|
10
|
-
Phase 2 Planning consumes the analysis document. MCP forbidden.
|
|
10
|
+
Phase 2 Planning consumes the analysis document. **Figma** MCP / REST forbidden; the toolkit MCP is not.
|
|
11
11
|
|
|
12
12
|
1. **Analysis document presence**: read `state.analysis.docStatus`, which Phase 1 Step 4 set.
|
|
13
13
|
- `produced` | `reused` -> the file is at `state.analysis.docPath[]`; continue with steps 2-4.
|
|
@@ -137,27 +137,29 @@ Non-zero exit fails CLOSED: emit the validator stderr + `errors[]` verbatim, att
|
|
|
137
137
|
|
|
138
138
|
Log: "Phase 2: Plan - {N} tasks created, {M} with architecture review, validator:pass"
|
|
139
139
|
|
|
140
|
+
#### Step 4.45 - Put the plan on the widget (required)
|
|
141
|
+
|
|
142
|
+
```bash
|
|
143
|
+
printf '%s' "$PLAN_JSON" | bash "$HOME/.claude/scripts/phase-tracker.sh" plan 3
|
|
144
|
+
```
|
|
145
|
+
|
|
146
|
+
Then do what its output asks - it prints a tile rebuild, because the widget
|
|
147
|
+
orders by creation. `dependsOn` becomes `addBlockedBy`, so the widget answers
|
|
148
|
+
"why has t3 not started". Required, unlike Step 4.5: the plan was always computed
|
|
149
|
+
and stored, only invisible.
|
|
150
|
+
|
|
140
151
|
#### Step 4.5 - Emit Plan Todo List (opt-in)
|
|
141
152
|
|
|
142
153
|
**Gated by `prefs.global.planTodos.enabled`** (default: `false`). When enabled, after the planning-output JSON validates and BEFORE the approval gate, transform `tasks[]` into a structured Todo list conforming to `$HOME/.claude/schemas/plan-todos.schema.json` and persist into `agent-state.plan`. The plan is rendered as a live, always-visible Todo list.
|
|
143
154
|
|
|
144
155
|
```bash
|
|
145
|
-
|
|
146
|
-
{
|
|
147
|
-
title: .summary,
|
|
148
|
-
todos: [ .tasks[] | {
|
|
149
|
-
id: .id,
|
|
150
|
-
task: .subject,
|
|
151
|
-
status: "pending",
|
|
152
|
-
deps: (.dependsOn // .blockedBy // []),
|
|
153
|
-
estimatedMinutes: (.estimatedMinutes // null)
|
|
154
|
-
} | with_entries(select(.value != null)) ]
|
|
155
|
-
}
|
|
156
|
-
' <<<"$PLAN_JSON")
|
|
157
|
-
|
|
158
|
-
bash "$HOME/.claude/lib/plan-todos.sh" set "$TASK_ID" "$TODO_BLOB"
|
|
156
|
+
printf '%s' "$PLAN_JSON" | bash "$HOME/.claude/lib/plan-todos.sh" set "$TASK_ID" -
|
|
159
157
|
```
|
|
160
158
|
|
|
159
|
+
`set` accepts a planning-output document and converts it itself. The conversion
|
|
160
|
+
used to be written out here, which made the `tasks[]`-to-`todos[]` mapping a
|
|
161
|
+
thing two files defined, and this was the copy nothing tested.
|
|
162
|
+
|
|
161
163
|
Phase 3 (Dev) then iterates with `plan-todos.sh next "$TASK_ID"` until empty, calling `start` before each step and `complete` (with notes) or `fail`/`skip` after. Phase 4 (Review) reads the Todo list to verify all `completed` items map to diff hunks. Phase 7 (Report) renders `list` into the agent-log + PR body.
|
|
162
164
|
|
|
163
165
|
**Why opt-in:** the existing `planning-output.schema.json` already drives Phase 3 dependency order - `plan.todos[]` is a richer surface (notes, durations, status transitions) but adds state writes per step. Off by default to keep the bare-bones flow unchanged; flip on for visibility into long features.
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
## Phase 3 Pre-flight (BLOCKING, v9.0.0)
|
|
6
6
|
|
|
7
|
-
Per Locked decision 30, Phase 3 Dev consumes the analysis document as the sole design source. MCP /
|
|
7
|
+
Per Locked decision 30, Phase 3 Dev consumes the analysis document as the sole design source. **Figma** MCP / REST / URL fetches are forbidden; the toolkit MCP is not, and Steps 3.4 and 3.55 call it.
|
|
8
8
|
|
|
9
9
|
Pre-flight steps (run in order, abort on failure).
|
|
10
10
|
|
|
@@ -159,7 +159,7 @@ Log: `Phase 6 Step 2.9: evidence host = {jira|github-public|github-private|none}
|
|
|
159
159
|
|
|
160
160
|
Generate a structured PR description based on task type. The PR body targets **code reviewers** - it should be technical: what changed, why, architecture decisions, how to verify.
|
|
161
161
|
|
|
162
|
-
Two inputs are read from state and the worktree before writing, not recalled from the conversation: `$WORKTREE/.pipeline/scope-check.json` (Phase 3 Step 3.7) supplies the `##
|
|
162
|
+
Two inputs are read from state and the worktree before writing, not recalled from the conversation: `$WORKTREE/.pipeline/scope-check.json` (Phase 3 Step 3.7) supplies the `## Technical Explanation` bullets from `files[].reason` and the "Follow-ups not done in this PR" list under `## Related` from `notDone[]`; `state.diffRisk.signals` (Phase 4 Step 1.75) decides whether the conditional `## Risk and Security` section is required. When a high-stakes signal is present and the section is missing, this step blocks until it is written; a placeholder answer ("TBD") counts as missing. `state.visualEvidence.required` blocks the same way: every required artefact is either attached or carries a recorded reason in `gaps[]` - the gate is against silence, not against an honest "no image on the ticket".
|
|
163
163
|
|
|
164
164
|
**required**: Run all generated text (PR body, commit message) through the `humanizer` skill before posting. This removes AI-generated patterns (inflated language, filler phrases, repetitive structure) and makes the output sound like a developer wrote it.
|
|
165
165
|
|
|
@@ -1,5 +1,16 @@
|
|
|
1
1
|
# Multi-Agent Pipeline - Phase Reference
|
|
2
2
|
|
|
3
|
+
<!-- toc -->
|
|
4
|
+
- [Phase Files](#phase-files)
|
|
5
|
+
- [Pipeline Flow](#pipeline-flow)
|
|
6
|
+
- [Phase entry - pending steer (every phase, every mode)](#phase-entry---pending-steer-every-phase-every-mode)
|
|
7
|
+
- [Visual Phase Tracker](#visual-phase-tracker)
|
|
8
|
+
- [Preferences File](#preferences-file)
|
|
9
|
+
- [Host Configuration](#host-configuration)
|
|
10
|
+
- [Token Budget](#token-budget)
|
|
11
|
+
- [SubPhase Convention](#subphase-convention)
|
|
12
|
+
<!-- /toc -->
|
|
13
|
+
|
|
3
14
|
## Phase Files
|
|
4
15
|
|
|
5
16
|
| Phase | File |
|
|
@@ -1,5 +1,17 @@
|
|
|
1
1
|
# Picker Contract (cross-platform single-choice abstraction)
|
|
2
2
|
|
|
3
|
+
<!-- toc -->
|
|
4
|
+
- [The abstract primitive](#the-abstract-primitive)
|
|
5
|
+
- [Step narration (breadcrumb)](#step-narration-breadcrumb)
|
|
6
|
+
- [Per-platform rendering (degradation ladder)](#per-platform-rendering-degradation-ladder)
|
|
7
|
+
- [Universal fallback: `pipeline/lib/ask-choice.sh`](#universal-fallback-pipelinelibask-choicesh)
|
|
8
|
+
- [Localized labels: what the caller owns](#localized-labels-what-the-caller-owns)
|
|
9
|
+
- [Order: project, then repo, then branch](#order-project-then-repo-then-branch)
|
|
10
|
+
- [A single candidate is still a question](#a-single-candidate-is-still-a-question)
|
|
11
|
+
- [Autopilot / non-interactive contract](#autopilot-non-interactive-contract)
|
|
12
|
+
- [Deterministic gates note](#deterministic-gates-note)
|
|
13
|
+
<!-- /toc -->
|
|
14
|
+
|
|
3
15
|
> Native `AskUserQuestion` is a Claude-Code primitive; Copilot CLI has no
|
|
4
16
|
> agent-invokable choice picker. This contract defines ONE abstract "ask the
|
|
5
17
|
> user to choose" primitive that each supported CLI renders to the best
|
|
@@ -4,6 +4,16 @@ description: "Internal - cross-platform parity cross-check for Phase 4 and /mu
|
|
|
4
4
|
|
|
5
5
|
# Platform parity - compare the change against the other platform's repo
|
|
6
6
|
|
|
7
|
+
<!-- toc -->
|
|
8
|
+
- [When it runs](#when-it-runs)
|
|
9
|
+
- [Read-only, without exception](#read-only-without-exception)
|
|
10
|
+
- [Locating the counterpart, deterministically](#locating-the-counterpart-deterministically)
|
|
11
|
+
- [What is compared](#what-is-compared)
|
|
12
|
+
- [What a finding may not claim](#what-a-finding-may-not-claim)
|
|
13
|
+
- [Output](#output)
|
|
14
|
+
- [Severity](#severity)
|
|
15
|
+
<!-- /toc -->
|
|
16
|
+
|
|
7
17
|
A feature that exists on iOS and Android is written twice, and the two copies
|
|
8
18
|
drift. The drift is invisible from inside one repo: the iOS diff is
|
|
9
19
|
self-consistent, the tests pass, and nobody notices that the Android screen
|
|
@@ -1,5 +1,15 @@
|
|
|
1
1
|
# Progress Line Contract
|
|
2
2
|
|
|
3
|
+
<!-- toc -->
|
|
4
|
+
- [Why](#why)
|
|
5
|
+
- [Line shape](#line-shape)
|
|
6
|
+
- [When to emit](#when-to-emit)
|
|
7
|
+
- [Verbosity](#verbosity)
|
|
8
|
+
- [Telemetry](#telemetry)
|
|
9
|
+
- [Phase adoption marker](#phase-adoption-marker)
|
|
10
|
+
- [Cross-CLI parity](#cross-cli-parity)
|
|
11
|
+
<!-- /toc -->
|
|
12
|
+
|
|
3
13
|
> **TLDR** - Every non-trivial pipeline action emits an inline "I'm doing X right now" line so the user always knows what's happening. Immediate flush, no batching. One line per action, standard shape. Autopilot prefers verbose. Low overhead. Mirrored in telemetry as `progress.step` events.
|
|
4
14
|
|
|
5
15
|
This contract is consumed by every phase (0-7) and every dispatched sub-skill (including the marketplace component toolkits when enabled). It is enforced by `smoke-progress-contract.sh` - any phase doc that drops the contract marker or diverges from the line shape fails CI.
|
|
@@ -58,7 +58,13 @@ default: option 1
|
|
|
58
58
|
Comment body: a short "Multi-agent readiness review" heading, the score + verdict, then the gap list grouped Blockers / Warnings / Gaps, each with a one-line fix suggestion. Tone rules (same as `channels/issue-comment.md`): no AI/Claude/Copilot attribution, no "generated by", no em-dash/section-sign, plain ASCII; if referencing another item use `Ref: #N` never `Closes/Fixes`. READY items get a short "ready to pick up" confirmation instead of a gap list.
|
|
59
59
|
|
|
60
60
|
Provider dispatch:
|
|
61
|
-
- **Jira** (`review-jira`):
|
|
61
|
+
- **Jira** (`review-jira`): convert markdown to wiki markup, then post it with
|
|
62
|
+
`bash "$HOME/.claude/lib/jira-publish.sh" --issue "$KEY" --body-file "$F" --target comment`,
|
|
63
|
+
per `channels/jira.md` - that file owns the Jira posting contract and this is
|
|
64
|
+
the same write path a Phase 7 report takes.
|
|
65
|
+
That script escapes the body and resolves the token itself; do not build the
|
|
66
|
+
`curl`. A readiness comment quotes the ticket's own text back at it, so it is
|
|
67
|
+
the body most likely to contain a `:)` sequence Jira would render as a smiley.
|
|
62
68
|
- **GitHub** (`review-issue`): `gh issue comment "$N" --repo "$org/$repo" --body-file <file>` per `channels/issue-comment.md` (auth via the Phase 0 `gh` account).
|
|
63
69
|
|
|
64
70
|
On post failure, surface the failing endpoint on stderr; the chat verdict remains the source of truth.
|
|
@@ -213,17 +213,9 @@ Full chain definition, REST endpoints, URL parsing, Code Connect snippet rules,
|
|
|
213
213
|
|
|
214
214
|
Per Locked decision 30 of `/multi-agent:analysis` and the parallel rule in `$HOME/.claude/rules/figma-pipeline.md`, Figma MCP / REST is allowed only in the analysis phase. Phase 2 through Phase 7 in every orchestrator mode (Full or Short, `--local`, autopilot) consume the analysis document + repo Code Connect mappings.
|
|
215
215
|
|
|
216
|
-
|
|
217
|
-
|
|
218
|
-
|
|
219
|
-
| Phase 2 Planning | forbidden | forbidden | analysis/<feature>-<platform>.md Section 6, 14 |
|
|
220
|
-
| Phase 3 Dev | forbidden | forbidden | analysis doc Section 5, 6, 7, 13 + Code Connect *.figma.swift |
|
|
221
|
-
| Phase 4 Review | forbidden | forbidden | analysis doc Section 21 References citations |
|
|
222
|
-
| Phase 5 Test | forbidden | forbidden | analysis doc Section 13.6 + 15.2 variant subset |
|
|
223
|
-
| Phase 6 Commit | forbidden | forbidden | analysis doc URL in PR body |
|
|
224
|
-
| Phase 7 Report | forbidden | forbidden | analysis doc embedded in channel artefacts |
|
|
225
|
-
|
|
226
|
-
Violation: smoke gate `smoke-no-mcp-in-dev-phases.sh` fails the run.
|
|
216
|
+
The ban is **Figma**-shaped, not MCP-shaped: the multi-agent-toolkit MCP (simulator, screenshot, xcodebuild, accessibility) is unaffected in every phase, and `smoke-no-mcp-in-dev-phases.sh` agrees - it matches `figma` in the tool name and nothing else. A bare "MCP forbidden" has been read as banning the screenshot and UI-test tools, which is how a run reaches Phase 7 with no evidence.
|
|
217
|
+
|
|
218
|
+
The per-phase matrix (which phase may fetch Figma, and what its sole design source is instead) lives in `rules/figma-pipeline.md` "Phase access matrix". It was copied here as a seven-row table for several releases, two paragraphs below this file's own instruction not to duplicate that rule file. Violation of either copy: `smoke-no-mcp-in-dev-phases.sh` fails the run.
|
|
227
219
|
|
|
228
220
|
Memory: [[mcp-only-in-analysis]]
|
|
229
221
|
|
|
@@ -1,5 +1,14 @@
|
|
|
1
1
|
# Firebase / Crashlytics Onboarding (setup Step 3c)
|
|
2
2
|
|
|
3
|
+
<!-- toc -->
|
|
4
|
+
- [1. Why there are three ways in](#1-why-there-are-three-ways-in)
|
|
5
|
+
- [2. Tier 1 - the service account](#2-tier-1---the-service-account)
|
|
6
|
+
- [3. Tier 2 - interactive session plus MCP](#3-tier-2---interactive-session-plus-mcp)
|
|
7
|
+
- [4. appId discovery](#4-appid-discovery)
|
|
8
|
+
- [5. v1alpha, and what to do when it breaks](#5-v1alpha-and-what-to-do-when-it-breaks)
|
|
9
|
+
- [6. Skipping](#6-skipping)
|
|
10
|
+
<!-- /toc -->
|
|
11
|
+
|
|
3
12
|
Loaded on demand by `/multi-agent:setup` Step 3c (optional, any platform). The SKILL.md carries the step intro; this file is the full flow.
|
|
4
13
|
|
|
5
14
|
Runs inside Step 3 alongside the other missing credentials. A user who already
|
|
@@ -1,5 +1,22 @@
|
|
|
1
1
|
## SwiftUI Component Generation Guide (Generic)
|
|
2
2
|
|
|
3
|
+
<!-- toc -->
|
|
4
|
+
- [Component Architecture: Configuration / View / Modifiers](#component-architecture-configuration-view-modifiers)
|
|
5
|
+
- [Simple vs Complex Decision](#simple-vs-complex-decision)
|
|
6
|
+
- [Configuration Purity Rules](#configuration-purity-rules)
|
|
7
|
+
- [Fluent Modifier Pattern](#fluent-modifier-pattern)
|
|
8
|
+
- [Token Discipline](#token-discipline)
|
|
9
|
+
- [Variant-Driven Implementation](#variant-driven-implementation)
|
|
10
|
+
- [Nested Component Handling](#nested-component-handling)
|
|
11
|
+
- [Accessibility Requirements](#accessibility-requirements)
|
|
12
|
+
- [Preview Best Practices](#preview-best-practices)
|
|
13
|
+
- [3-Layer Test Strategy](#3-layer-test-strategy)
|
|
14
|
+
- [Build Verification](#build-verification)
|
|
15
|
+
- [Component Quality Checklist](#component-quality-checklist)
|
|
16
|
+
- [Compliance Rules (maps to multi-agent-toolkit MCP audit tools)](#compliance-rules-maps-to-multi-agent-toolkit-mcp-audit-tools)
|
|
17
|
+
- [Figma URL Given](#figma-url-given)
|
|
18
|
+
<!-- /toc -->
|
|
19
|
+
|
|
3
20
|
> **MUST: Figma MCP-first (BLOCKING).** If the task references any Figma frame (URL, node ID, or "from the design"), the Dev phase MUST call `mcp__claude_ai_Figma__get_design_context` for every frame BEFORE writing a single UI line. Use the `CodeConnectSnippet` component name verbatim - no sound-alike substitutions. Authentication failure is not a skip path. Full rule, trigger conditions, and gate failure modes: `$HOME/.claude/rules/figma-pipeline.md` "MUST: Figma MCP-first (BLOCKING)". Phase wiring: `$HOME/.claude/multi-agent-refs/phases/phase-3-dev.md` "MUST: Figma MCP-first (BLOCKING pre-step)".
|
|
4
21
|
|
|
5
22
|
When the task involves creating a SwiftUI component (any project), follow this architecture.
|
|
@@ -4,6 +4,18 @@ description: "Phase tracker mandatory contract - every mode depends on this. N
|
|
|
4
4
|
|
|
5
5
|
# Phase Tracker - Mandatory Contract
|
|
6
6
|
|
|
7
|
+
<!-- toc -->
|
|
8
|
+
- [Why two channels](#why-two-channels)
|
|
9
|
+
- [Visual channel - chosen by the agent based on the host CLI](#visual-channel---chosen-by-the-agent-based-on-the-host-cli)
|
|
10
|
+
- [Call pattern](#call-pattern)
|
|
11
|
+
- [The plan is part of the list](#the-plan-is-part-of-the-list)
|
|
12
|
+
- [Resume behaviour](#resume-behaviour)
|
|
13
|
+
- [Continuation runs (finish / manual-test)](#continuation-runs-finish-manual-test)
|
|
14
|
+
- [Anti-patterns](#anti-patterns)
|
|
15
|
+
- [Verification](#verification)
|
|
16
|
+
- [Cross-reference](#cross-reference)
|
|
17
|
+
<!-- /toc -->
|
|
18
|
+
|
|
7
19
|
> **TLDR** - At every phase boundary the agent does two things: (1) writes to the state file via `phase-tracker.sh` (identical on every CLI), (2) drives the visual channel for the CLI it's running in - `TaskCreate`/`TaskUpdate` in Claude Code, `phase-tracker.sh render` in every other CLI. The state file alone is not enough; the user must see the phases progress.
|
|
8
20
|
|
|
9
21
|
## Why two channels
|
|
@@ -368,6 +380,38 @@ bash $HOME/.claude/scripts/phase-tracker.sh sub <N> <sub_id> "<sub name>" in_pro
|
|
|
368
380
|
TaskUpdate({taskId: <phase_task>, activeForm: "Running explorer: repo-map"})
|
|
369
381
|
```
|
|
370
382
|
|
|
383
|
+
## The plan is part of the list
|
|
384
|
+
|
|
385
|
+
Phase 2 computes `tasks[]`, their order and their `dependsOn[]` edges, stores
|
|
386
|
+
them, and uses them to drive Phase 3's ready-task picker. For releases the card
|
|
387
|
+
drew them as sub-phases and the widget did not, which meant the one surface the
|
|
388
|
+
user actually watches was the one place the plan did not exist.
|
|
389
|
+
|
|
390
|
+
Phase 2 Step 4.45 calls `phase-tracker.sh plan 3` with the planning-output
|
|
391
|
+
document on stdin. Each task becomes a sub-phase of the phase that will execute
|
|
392
|
+
it, `pending`, carrying its `dependsOn[]` as `deps`. Phase 3 moves them with
|
|
393
|
+
`sub` as it works.
|
|
394
|
+
|
|
395
|
+
**A rebuild, not an append.** Creation order is the only ordering the native
|
|
396
|
+
widget has - there is no parent field and no insert - so steps arriving at
|
|
397
|
+
Phase 2 cannot be appended without landing after Phase 7. `tiles` detects that
|
|
398
|
+
sub-steps exist and asks for the list to be deleted and recreated. That is one
|
|
399
|
+
rebuild at one boundary, and it is the same thing Resume already does below for
|
|
400
|
+
a different reason.
|
|
401
|
+
|
|
402
|
+
**Per host, with what each one actually has:**
|
|
403
|
+
|
|
404
|
+
| Host | How the plan arrives | Dependencies |
|
|
405
|
+
|---|---|---|
|
|
406
|
+
| Claude Code | `TaskCreate` per row, the steps indented inside the subject string | `TaskUpdate(..., addBlockedBy: [...])`, a second pass because a tile cannot be blocked by one that does not exist yet |
|
|
407
|
+
| Codex | `update_plan` re-read from `subjects`, which now includes the steps | No dependency concept. The `(bekliyor: t1)` suffix inside the step text is where that survives |
|
|
408
|
+
| Copilot CLI | The card, which has drawn sub-phases with tree glyphs since they existed | Same suffix, on the card row |
|
|
409
|
+
|
|
410
|
+
The dependency lines name tiles by their subject, never by the plan's own `t1` /
|
|
411
|
+
`t2` ids: `addBlockedBy` takes the ids `TaskCreate` returned, which no shell
|
|
412
|
+
script can know. Printing `addBlockedBy: ["t1"]` would be a line that looks
|
|
413
|
+
executable and is not, which is worse than one that admits what it needs.
|
|
414
|
+
|
|
371
415
|
## Resume behaviour
|
|
372
416
|
|
|
373
417
|
When `/multi-agent:resume <task_id>` is called:
|
|
@@ -1,5 +1,15 @@
|
|
|
1
1
|
## Web Development Guide
|
|
2
2
|
|
|
3
|
+
<!-- toc -->
|
|
4
|
+
- [Component Architecture](#component-architecture)
|
|
5
|
+
- [React Component Pattern](#react-component-pattern)
|
|
6
|
+
- [State Management](#state-management)
|
|
7
|
+
- [Token Discipline](#token-discipline)
|
|
8
|
+
- [Accessibility](#accessibility)
|
|
9
|
+
- [Testing](#testing)
|
|
10
|
+
- [Quality Checklist](#quality-checklist)
|
|
11
|
+
<!-- /toc -->
|
|
12
|
+
|
|
3
13
|
When the task involves web development (React, Next.js, Vue), follow these patterns.
|
|
4
14
|
|
|
5
15
|
### Component Architecture
|
|
@@ -1,5 +1,16 @@
|
|
|
1
1
|
# Component Wiki Capture (channels.md Wiki adapter)
|
|
2
2
|
|
|
3
|
+
<!-- toc -->
|
|
4
|
+
- [Applicability](#applicability)
|
|
5
|
+
- [Case A - preconditions met (scope multi-select)](#case-a---preconditions-met-scope-multi-select)
|
|
6
|
+
- [Case B - preconditions missing (actionable menu)](#case-b---preconditions-missing-actionable-menu)
|
|
7
|
+
- [Legacy prompt + preference flow (pre-v5.7, still supported for backward compat)](#legacy-prompt-preference-flow-pre-v57-still-supported-for-backward-compat)
|
|
8
|
+
- [Dispatch](#dispatch)
|
|
9
|
+
- [Skip conditions (explicit log lines)](#skip-conditions-explicit-log-lines)
|
|
10
|
+
- [Success log](#success-log)
|
|
11
|
+
- [Cross-CLI parity](#cross-cli-parity)
|
|
12
|
+
<!-- /toc -->
|
|
13
|
+
|
|
3
14
|
> **TLDR** - Component tasks can auto-generate wiki docs + Figma screenshots. The Wiki adapter is invoked from `/multi-agent:channels` (Phase 7 delegates, or user invokes post-hoc). Four layouts supported (`submodule`, `in-repo`, `github-wiki`, `separate-repo`) - adapter picked from `figmaConfig.wiki.mode`. Non-blocking: failures log a warning and channels continues to other adapters. The Wiki adapter supports scope multi-select (Case A) and a precondition-failure menu (Case B) - see below.
|
|
4
15
|
|
|
5
16
|
This doc is referenced from `commands/multi-agent/channels/SKILL.md` (Wiki adapter) and indirectly from `$HOME/.claude/multi-agent-refs/phases/phase-7-report.md` (which delegates all external delivery to channels). Keeping it separate keeps both files under their token budgets and gives the contract a stable location for Claude-side + Copilot-side implementations.
|
|
@@ -0,0 +1,149 @@
|
|
|
1
|
+
{
|
|
2
|
+
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
|
3
|
+
"$id": "https://github.com/mmerterden/multi-agent-pipeline/pipeline/schemas/autopilot-config.schema.json",
|
|
4
|
+
"title": "Autopilot configuration",
|
|
5
|
+
"description": "Continuous mode on ONE machine. Lives at ~/.claude/autopilot/config.json and deliberately NOT in multi-agent-preferences.json: that file has additionalProperties false and a migration chain, so a block there would push a key into every local user's preferences forever, including everyone who never turns this on. This file is the repo SELECTION, not the on/off state: autopilot-off keeps it so turning the mode back on does not re-ask which repos. On/off is whether launchd holds the job, because that is the thing that actually makes work happen. A clean install writes neither - there is no `enabled: false` default, because a file written at install time is a file that can be wrong.",
|
|
6
|
+
"type": "object",
|
|
7
|
+
"additionalProperties": false,
|
|
8
|
+
"required": ["version", "repos"],
|
|
9
|
+
"properties": {
|
|
10
|
+
"version": {
|
|
11
|
+
"type": "integer",
|
|
12
|
+
"const": 1,
|
|
13
|
+
"description": "Config shape version. Bumped only when a field changes meaning, never for additions."
|
|
14
|
+
},
|
|
15
|
+
"createdAt": {
|
|
16
|
+
"type": "string",
|
|
17
|
+
"format": "date-time"
|
|
18
|
+
},
|
|
19
|
+
"updatedAt": {
|
|
20
|
+
"type": "string",
|
|
21
|
+
"format": "date-time"
|
|
22
|
+
},
|
|
23
|
+
"repos": {
|
|
24
|
+
"type": "array",
|
|
25
|
+
"description": "Repos this machine picks work up from. Empty is valid and means the mode is configured but idle - the picker was opened and everything was unchecked. No repo is ever added implicitly.",
|
|
26
|
+
"items": {
|
|
27
|
+
"type": "object",
|
|
28
|
+
"additionalProperties": false,
|
|
29
|
+
"required": ["nameWithOwner", "localPath"],
|
|
30
|
+
"properties": {
|
|
31
|
+
"nameWithOwner": {
|
|
32
|
+
"type": "string",
|
|
33
|
+
"description": "owner/repo as GitHub reports it."
|
|
34
|
+
},
|
|
35
|
+
"localPath": {
|
|
36
|
+
"type": "string",
|
|
37
|
+
"description": "The checkout on this machine. A repo with write permission but no local checkout is listed in the picker as unavailable WITH that reason, never silently dropped - the measured set is 65 repos with push rights and far fewer checkouts, so silent dropping would look like a permission problem."
|
|
38
|
+
},
|
|
39
|
+
"sources": {
|
|
40
|
+
"type": "array",
|
|
41
|
+
"description": "Where items come from for this repo. Both may be on.",
|
|
42
|
+
"items": {
|
|
43
|
+
"enum": ["github", "jira"]
|
|
44
|
+
},
|
|
45
|
+
"uniqueItems": true,
|
|
46
|
+
"default": ["github"]
|
|
47
|
+
},
|
|
48
|
+
"githubLabel": {
|
|
49
|
+
"type": "string",
|
|
50
|
+
"default": "agent-queue",
|
|
51
|
+
"description": "Only issues carrying this label are picked up. `gh issue list --label` does the filtering server-side."
|
|
52
|
+
},
|
|
53
|
+
"jiraJql": {
|
|
54
|
+
"type": "string",
|
|
55
|
+
"description": "Full JQL, so an instance that cannot use labels is a config change rather than a redesign. The label form is sugar for: assignee = currentUser() AND labels = \"<label>\" AND resolution = EMPTY AND status not in (Done, Closed, Cancelled) ORDER BY priority DESC, created ASC. Jira labels are a single global namespace anyone can write to, so on a shared instance prefer a qualified name, and note that the Labels field must be on the edit screen for the issue types in play."
|
|
56
|
+
},
|
|
57
|
+
"jiraLabel": {
|
|
58
|
+
"type": "string",
|
|
59
|
+
"default": "agent-queue",
|
|
60
|
+
"description": "Sugar for jiraJql. Ignored when jiraJql is set."
|
|
61
|
+
},
|
|
62
|
+
"rank": {
|
|
63
|
+
"type": "integer",
|
|
64
|
+
"minimum": 1,
|
|
65
|
+
"maximum": 9,
|
|
66
|
+
"description": "Optional explicit repo precedence. Absent means ordered by the normal rules."
|
|
67
|
+
}
|
|
68
|
+
}
|
|
69
|
+
}
|
|
70
|
+
},
|
|
71
|
+
"slots": {
|
|
72
|
+
"type": "integer",
|
|
73
|
+
"minimum": 1,
|
|
74
|
+
"maximum": 8,
|
|
75
|
+
"default": 1,
|
|
76
|
+
"description": "How many items run at once. Per-repo concurrency is always 1 whatever this says, because two runs in one checkout contend on .git/index.lock - so slots > 1 means that many DIFFERENT repos. Starts at 1 on purpose; autopilot-status prints the ceiling this machine's RAM, cores and free disk would allow so raising it is a measured decision rather than a guess."
|
|
77
|
+
},
|
|
78
|
+
"scanIntervalSeconds": {
|
|
79
|
+
"type": "integer",
|
|
80
|
+
"minimum": 60,
|
|
81
|
+
"maximum": 3600,
|
|
82
|
+
"default": 120,
|
|
83
|
+
"description": "How often a tick looks for new work. 120s means an issue is picked up within two minutes of being labelled. The scan itself is cheap - one `gh issue list` per repo plus one Jira search - so at five repos this is about 150 GitHub calls an hour against a 5000/hour limit. launchd does NOT wake a sleeping Mac, it only fires on the next wake, which is why the mode holds a sleep assertion while on AC and releases it on battery."
|
|
84
|
+
},
|
|
85
|
+
"maxAttempts": {
|
|
86
|
+
"type": "integer",
|
|
87
|
+
"minimum": 1,
|
|
88
|
+
"default": 2,
|
|
89
|
+
"description": "How many times a died or build-failed item is re-queued before it is left alone. Retry is part of intake rather than a separate mechanism; without it one transient failure drops an item silently."
|
|
90
|
+
},
|
|
91
|
+
"maxConsecutivePerRepo": {
|
|
92
|
+
"type": "integer",
|
|
93
|
+
"minimum": 1,
|
|
94
|
+
"default": 5,
|
|
95
|
+
"description": "How many items from one repo may run back to back before another repo gets a turn. Grouping by repo is what lets items 2..N reuse the code graph, repo map and knowledge index that are rebuilt per run today, so the default is deliberately generous - but without a ceiling a repo with forty labelled issues would hold the queue until it was empty, and every other repo would look broken."
|
|
96
|
+
},
|
|
97
|
+
"askOnLowMaturity": {
|
|
98
|
+
"type": "boolean",
|
|
99
|
+
"default": true,
|
|
100
|
+
"description": "An item that fails the readiness review gets ONE comment listing the open questions, and waits. This is an outward-facing write to a real ticket, so it is a flag rather than an assumption. When false the item is skipped with the reason recorded and nothing is posted."
|
|
101
|
+
},
|
|
102
|
+
"maxAskRounds": {
|
|
103
|
+
"type": "integer",
|
|
104
|
+
"minimum": 1,
|
|
105
|
+
"default": 2,
|
|
106
|
+
"description": "Ceiling on how many times one item can be asked about, so a permanently vague ticket cannot be commented on forever. The gap fingerprint prevents repeating the SAME question; this prevents an endless series of different ones."
|
|
107
|
+
},
|
|
108
|
+
"costCeilingUsd": {
|
|
109
|
+
"type": "number",
|
|
110
|
+
"minimum": 0,
|
|
111
|
+
"default": 25,
|
|
112
|
+
"description": "Rolling 24-hour spend ceiling. Dispatch refuses past it and says so; running work is never killed mid-item. There is deliberately no cap on the number of PRs - the queue is bounded by this and by machine capacity, not by a daily item count."
|
|
113
|
+
},
|
|
114
|
+
"channels": {
|
|
115
|
+
"type": "object",
|
|
116
|
+
"additionalProperties": false,
|
|
117
|
+
"description": "Phase 7 pauses for channel selection in every attended mode, by design. Unattended there is nobody to ask, so the answer comes from here instead and Phase 7 completes without prompting. Without this the queue would generate its own backlog: every finished item parked waiting for a menu.",
|
|
118
|
+
"properties": {
|
|
119
|
+
"mode": {
|
|
120
|
+
"enum": ["none", "pr-body", "configured"],
|
|
121
|
+
"default": "pr-body",
|
|
122
|
+
"description": "none: no report anywhere. pr-body: the report goes in the PR body only, no external writes. configured: use the list below."
|
|
123
|
+
},
|
|
124
|
+
"targets": {
|
|
125
|
+
"type": "array",
|
|
126
|
+
"items": {
|
|
127
|
+
"enum": ["jira", "confluence", "wiki", "pr"]
|
|
128
|
+
},
|
|
129
|
+
"uniqueItems": true
|
|
130
|
+
}
|
|
131
|
+
}
|
|
132
|
+
},
|
|
133
|
+
"reportStatus": {
|
|
134
|
+
"type": "boolean",
|
|
135
|
+
"default": true,
|
|
136
|
+
"description": "Send a small heartbeat to the admin panel: counts and states only. Never item keys, repo names, branch names or PR titles - the existing telemetry hashes even the repo name, and carrying corporate ticket keys into that database would break the protection it was built for. A failed heartbeat is silent and never blocks a tick."
|
|
137
|
+
},
|
|
138
|
+
"sendItemKeys": {
|
|
139
|
+
"type": "boolean",
|
|
140
|
+
"default": false,
|
|
141
|
+
"description": "Opt-in to also send item identifiers to the panel. Off by default and shown once with an example of what would be sent, because the same pipeline carries corporate identifiers."
|
|
142
|
+
},
|
|
143
|
+
"depthRouter": {
|
|
144
|
+
"enum": ["off", "auto"],
|
|
145
|
+
"default": "off",
|
|
146
|
+
"description": "off: every item runs Full. auto: a deterministic router may pick Short for a bugfix or chore under the blast-radius threshold, and writes its reasoning into the PR body so the shortcut is visible to the reviewer. Off until a dozen items have been watched on Full."
|
|
147
|
+
}
|
|
148
|
+
}
|
|
149
|
+
}
|
|
@@ -1871,6 +1871,10 @@
|
|
|
1871
1871
|
"default": 60,
|
|
1872
1872
|
"description": "Recording cap. A flow needing longer is a debugging session, not a review artefact. Android's `screenrecord` has its own ceiling of 180s that no setting can lift, so capture-evidence.sh clamps to it and says when it did: one preference honoured on one platform and silently halved on the other is worse than a stated limit."
|
|
1873
1873
|
},
|
|
1874
|
+
"webBaseUrl": {
|
|
1875
|
+
"type": "string",
|
|
1876
|
+
"description": "The address a web capture points the browser at. Web is the one evidence platform that needs one: a simulator is already showing something, a dev server has to be pointed at. Absent, `capture-evidence.sh after --platform web` reports a gap with that reason instead of guessing a port. Per-run override: `--url`."
|
|
1877
|
+
},
|
|
1874
1878
|
"githubHost": {
|
|
1875
1879
|
"type": "string",
|
|
1876
1880
|
"enum": ["branch", "off"],
|
|
@@ -1,41 +1,32 @@
|
|
|
1
1
|
{
|
|
2
2
|
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
|
3
3
|
"$id": "https://github.com/mmerterden/multi-agent-pipeline/pipeline/schemas/token-budget.json",
|
|
4
|
-
"description": "Per-phase token
|
|
4
|
+
"description": "Per-phase token ceilings for the lazy-loaded pipeline docs, enforced by smoke-token-budget.sh. Only the ACTIVE phase is loaded at run time and nothing truncates a phase document, so these numbers govern what may be written, not what a run receives. Two rules, and the split is the point. `max_tokens` and `total_max_tokens` are committed constants a human owns: computing them would end the gate, because a ceiling that is always current x k can never fail. The warn tier is NOT stored - the gate derives it per phase from that phase's own git history (median + 3*MAD of its historical per-commit deltas, which is robust to the one 980-token release that makes sigma meaningless on phase-4), because a soft line maintained by hand rots, and this one rotted twice: it was reset at v13.6.0 when five lines had gone permanently amber, and six of eight were amber again by v17.1.0. A ceiling more than 25% above the measurement is stale and fails, which is the 'only ratchets down' rule the SKILL.md grace list always had and this budget never did. Change history: docs/token-budget-history.md.",
|
|
5
5
|
"phases": {
|
|
6
6
|
"phase-0-init": {
|
|
7
|
-
"max_tokens": 13100
|
|
8
|
-
"warn_tokens": 12400
|
|
7
|
+
"max_tokens": 13100
|
|
9
8
|
},
|
|
10
9
|
"phase-1-analysis": {
|
|
11
|
-
"max_tokens": 4600
|
|
12
|
-
"warn_tokens": 4050
|
|
10
|
+
"max_tokens": 4600
|
|
13
11
|
},
|
|
14
12
|
"phase-2-planning": {
|
|
15
|
-
"max_tokens": 6500
|
|
16
|
-
"warn_tokens": 5650
|
|
13
|
+
"max_tokens": 6500
|
|
17
14
|
},
|
|
18
15
|
"phase-3-dev": {
|
|
19
|
-
"max_tokens": 9300
|
|
20
|
-
"warn_tokens": 8600
|
|
16
|
+
"max_tokens": 9300
|
|
21
17
|
},
|
|
22
18
|
"phase-4-review": {
|
|
23
|
-
"max_tokens": 15150
|
|
24
|
-
"warn_tokens": 12950
|
|
19
|
+
"max_tokens": 15150
|
|
25
20
|
},
|
|
26
21
|
"phase-5-test": {
|
|
27
|
-
"max_tokens": 3050
|
|
28
|
-
"warn_tokens": 2850
|
|
22
|
+
"max_tokens": 3050
|
|
29
23
|
},
|
|
30
24
|
"phase-6-commit": {
|
|
31
|
-
"max_tokens": 6550
|
|
32
|
-
"warn_tokens": 6100
|
|
25
|
+
"max_tokens": 6550
|
|
33
26
|
},
|
|
34
27
|
"phase-7-report": {
|
|
35
|
-
"max_tokens": 6350
|
|
36
|
-
"warn_tokens": 5600
|
|
28
|
+
"max_tokens": 6350
|
|
37
29
|
}
|
|
38
30
|
},
|
|
39
|
-
"total_max_tokens":
|
|
40
|
-
"note": "Token estimate = ceil(chars / 4). Per-phase budget rule: warn = current+10% (rounded to nearest 50), max = current+25%. Gives ~6 edit cycles of headroom before warn trips - intentionally quiet under normal maintenance, loud when a phase grows unusually. Only the active phase is loaded (lazy). Recalibrated at v10.0.0 after the validator/consistency/simplifier/lesson gate contracts landed in phases 1-4. Recalibrated again at v10.9.0 after the verify-by-test (Phase 4 Step 3.7), update-check (Phase 0 Step 0.6), immutable-test (Phase 3 GREEN) and redTests re-entry contracts landed - Step 3.7 prose was compressed to a pointer into refs/features/verify-by-test.md before the recalibration. Total bumped 50000 -> 51000 at v12.5.0 after the worktree residue/traversal-prune contract (Phase 0 + Phase 5 heal) and the Reflexion causal-diagnosis contract (Phase 4 lesson memory) landed; the prose was compressed first (161 tokens reclaimed) and every per-phase max still passes - only the aggregate needed room. Recalibrated again at v13.6.0 after the install-relative path correction: an instruction that names `pipeline/scripts/x` resolves only from a repo checkout, and a run happens in the user's worktree, so 157 references across these docs moved to `$HOME/.claude/...` at +5 bytes each - 196 tokens of pure correctness cost. Same discipline as before: prose was compressed FIRST (149 tokens reclaimed, by pointing Phase 1's Figma tier table at the Phase 0 probe that already resolved it and Phase 4's Codex constraints at the always-loaded AGENTS.md block), and only then were the budgets moved. Five warn lines had been permanently amber, which makes the amber tier useless as a signal, so every warn was reset to the documented current+10% and the four maxes that the new warn would have collided with were reset to current+25%. Aggregate 51000 -> 51500. Total bumped 51500 -> 52200 at v14.0.0 after Phase 4 Review entered the four --dev mode phase sets and the criteria-resolution contract (Step 1.78) landed. Same discipline as every prior bump: prose was compressed FIRST, 820 tokens reclaimed, before the number moved. Two of those compressions are structural rather than cosmetic - the hardcoded SwiftUI interaction list in Step 1.5 and the SwiftUI convention paragraph in Step 2.8 were transcriptions of rules that now live in a scoped registry, so keeping them here would have re-created the drift this release exists to remove, and the third moved the Step 1.78 full contract into refs/features/skill-conformance.md leaving a pointer. What remains is contract text that cannot be inferred: the manifest's four consumer-visible parts, the conformance checklist the reviewers must return, and the fail-closed semantics. Every per-phase max still passes (phase-4 12405/14750); only the aggregate needed room. Total bumped 52200 -> 52700 at v14.1.0 after two more contracts landed: stack skill routing (Phase 3 pre-flight step 9) and worktree finalize (Phase 6 step 9). Compression came first, as always, and twice: 224 tokens out of Phase 3 by pointing its criteria-ledger and routing steps at their feature files instead of restating them, and 190 out of Phase 6 by moving the finalize contract into refs/features/worktree-finalize.md and leaving the invocation plus the exit-3 semantics. Both new contracts follow the pattern the earlier ones set: the phase doc carries the call and the decision, the feature file carries the reasoning, and the feature files are outside this budget because it loops only the eight phase-N-* keys. Every per-phase max still passes (phase-3 7677/8950, phase-6 5223/6150 and both under warn); only the aggregate needed room. Total bumped 52700 -> 52750 for the Phase 0 Step 3 branch-persistence correction: the step wrote the legacy `projects[].branches` while the TTL filter two sections below read `global.recentBranches`, and both spots named a `{name, lastUsed}` shape the schema rejects (`branch` required, `additionalProperties: false`), so the recent-branch picker option could never populate and a literal implementation would have failed prefs validation. Naming the right target, the right key and the legacy field to avoid costs 41 tokens over the one line it replaces. Compression came first and was applied three times to the replacement text itself, from 120 tokens down to 66, by moving the rationale out of the phase doc entirely: the reasoning now lives where it is enforced, in the migrate-prefs carry-forward comment and the smoke-pref-migration f7 block, leaving the phase doc with only the instruction. 50 was the smallest step that clears it; phase-0-init sits at 10893/12400, far under its own max, so this is purely an aggregate ceiling. v15.0.0: total 52750 -> 53100, the stack-skill tables in phase-1/2/4 now carry plugin-namespaced names (ai-<stack>-toolkit:<skill>) - functional prefixes, ~170 tokens. v15.10.0: total 53350 -> 53950 for the memory-recall + context-offload contracts (Phase 1 two-block durable-knowledge injection and its telemetry, Phase 3 build-log offload pipe, Phase 4 ranked prior art, offload pipe and recall telemetry). Compression came first and twice, taking the new prose from 1168 tokens to 580: the reasoning behind the two blocks lives in multi-agent-refs/prompt-assembly.md and the reasoning behind the offload filter lives in the offload-ref.sh header, both outside this budget, so the phase docs carry only the call, the pref that gates it and the one fact an agent cannot infer - that the evidence gate still reads the whole build log, so offloading changes what is read, never what counts as a verified pass. Every per-phase max still passes (phase-3 7985/8950, phase-4 12997/14750); phase-3 and phase-4 crossed their warn lines and are left amber on purpose, because that is the signal that those two docs are the next ones needing structural compression rather than another bump. v15.13.0: total 53950 -> 54050 for the prefs-to-flag bridges. Five settings had shipped declared-but-inert: contextOffload.minLines and .tailLines (fixed in 15.11.0), learningsLedger.maxBriefEntries, and testGap.scanTree and .promoteSeverity - the last two declared in the schema AND implemented as flags in the scanner, with nothing in between reading the pref and passing the flag. Wiring three of them costs the phase docs 94 tokens, which is the wiring itself and not prose: two `--max` substitutions and a three-line GAP_FLAGS block. Compression came first and twice, as always: the rationale that would have sat in phase-5 now lives in the header of smoke-prefs-consumed.sh, the gate that makes this class fail a build instead of shipping, and a `--severity-promote` table row was dropped because the invocation above it now shows the flag and names the pref that triggers it, which the row did not. 100 was the smallest step that clears it. Every per-phase max still passes; phase-3 and phase-4 remain amber on purpose. v15.14.0: total 54050 -> 54400 for the supported-version gate. Phase 0 Step 0.6 stopped being purely advisory: a release can now publish an npm dist-tag `required` that names the oldest runnable version, and below it the run halts instead of nagging. What the phase doc has to carry is the part an agent cannot infer - the third stdout field, that the halt is identical in autopilot, and that the run must NOT continue on the freshly updated install because its docs were already loaded from the old version. Compression came first, as always, and took the new prose from 469 tokens to 337: the rationale for the floor, the exemption list, the fail-open rules and the `npm dist-tag add` recipe all moved to multi-agent-refs/rules.md \"Supported Version Gate\" (loaded by 25 commands, outside this budget) and to the header of require-supported-version.sh, leaving the phase doc with the call, the decision table and the halt. 350 was the smallest step that clears it. Every per-phase max still passes (phase-0-init 11230/12400); phase-3 and phase-4 remain amber on purpose. v15.17.0: total 54400 -> 54900 for the Phase 1 analysis-document step. Phase 2 and Phase 3 pre-flights had BLOCKED on `analysis/<feature>-<platform>.md` since v9.0.0 while nothing produced it, so a full run either aborted at Phase 2 or the model ignored its own BLOCKING contract; Step 4 is the producer. What the phase doc carries is only what cannot be inferred: the when-table (taskType x Figma reference), the four refs in load order, the two artefacts, and that the doc validator fails closed. Compression came first and took the step from 745 tokens to 497: the history of why the gap existed moved to the CHANGELOG, the per-ref one-line descriptions moved into the refs' own headers, and the autopilot carve-out collapsed to one clause. The 17.4k-token analysis engine itself is NOT in this budget - it moved out of commands/ into multi-agent-refs/analysis/{locked,evidence,synthesis,render}.md, loaded on demand, which also took analysis/SKILL.md from 18081 to 5974 tokens and retired its lint grace entry. 500 was the smallest step that clears it; phase-1-analysis sits at 4338/4600 and is amber on purpose, like phase-3 and phase-4. v15.18.0: total 54900 -> 55250 for analysis mode. Three phase docs gained a mode branch that cannot be inferred: Phase 4 reviews a document instead of a diff (validator, the one question reviewers answer, the open-question walk), and Phase 6 publishes instead of committing. Compression came first and was applied twice to the new prose and once to old: the Phase 4 branch went from 320 tokens to 180 and the Phase 6 branch from 190 to 120 by pointing at multi-agent-refs/analysis/{resolve,render}.md, which now hold the walks themselves, and the front-matter parse contract stopped being spelled out in both pre-flights. The analysis engine keeps leaving this budget rather than entering it: intake joined locked/evidence/synthesis/render/resolve in multi-agent-refs/analysis/, which is what let analysis/SKILL.md drop under the 6000 hard cap after its grace entry was retired. 350 was the smallest step that clears it; phase-4 and phase-6 are amber on purpose, as phase-1 and phase-3 already were. v15.20.0: total 55250 -> 55500 for the TDD bridge. Phase 3 pre-flight read the analysis doc's concept table and even said test method names come from it, while nothing read Section 15 - so the RED step invented tests and the analysis test matrix never reached development. Phase 3 step 5b now loads it into state.dev.testPlan[] and Phase 4 step 1.45 cross-checks every planned row against a real test, which is what turns \"analysis quality is output quality\" from a slogan into a finding. Compression came first on both blocks, 300 tokens down to 175, by dropping the enumerated failure modes to one line each and the rationale to one clause; the reasoning lives in the CHANGELOG. 250 was the smallest step that clears it. v15.21.0: total 55500 -> 55800 for the post-analysis confirmation. Phase 2 gained Step 0.9, the last human checkpoint before Phase 3: derived values are shown for confirmation and only Section 20 rows are asked, through the resolve engine that already exists in refs. It belongs here rather than Phase 4 because Phase 4 runs after development, where an answer arrives too late to change anything. Compression came first and twice, 430 tokens down to 250, by collapsing the derived-vs-asked explanation to one sentence each and moving the walk itself to multi-agent-refs/analysis/resolve.md, which Phase 4 and analysis-resolve already mount. 300 was the smallest step that clears it. v15.22.0: total 55800 -> 55900 for the analyst-toolkit hooks. Phase 1 Step 4 now names the two prefs that decide whether a document is produced at all and how deep it goes (forceFull, mode) - the first of those had shipped declared-but-inert and smoke-prefs-consumed caught it - and Phase 4 triage gained one clause: a finding that blames a third-party library asks evidence-github whether it is already open upstream, which turns it into a deferred item with a citation instead of Phase 3 rework on code that is not ours. Compression came first and three times, taking the new prose from 220 tokens to 110, and the Phase 1d evidence contract itself never entered this budget - it lives in multi-agent-refs/analysis/evidence.md beside the phases it belongs to. 100 was the smallest step that clears it, leaving 34 tokens of headroom. phase-4 stays amber and the debt named at v15.10.0 stands: it is the doc that needs structural compression rather than another bump, and the two candidates are the inline triage JSON shape and the 3.4 telemetry block, both of which restate something already authoritative elsewhere. v16.0.0: total 55900 -> 56350 for the depth picker. `--dev` and the four dev-* commands are gone; depth is Phase 0 Step 7.5, which costs phase-0-init a step it did not have. Compression came first and three times, taking the step from 530 tokens to 300: the question wording, the per-taskType recommendation and the mode tables all live in phases/modes.md (outside this budget), so the phase doc carries only what an agent cannot infer - that the step runs after Step 7 and why, who is exempt, that ASK_CHOICE_DEFAULT must be passed explicitly because ask-choice.sh takes the FIRST option on a non-TTY, and that Short flips the Phase 1/2 tiles late rather than pre-marking them. The phase-4 telemetry block named as compression debt at v15.22.0 was collapsed to an emit() helper (-27) and the four dev-* mode files left the tree entirely, but neither offsets a genuinely new phase step. 450 was the smallest step that clears it, leaving 119 tokens of headroom. phase-4 remains amber and its other named candidate, the inline triage JSON shape, was left alone on purpose: it is the prompt the triage agent is handed, not a restatement for readers. v16.2.0: total 56350 -> 56600 for the spec-freshness and reuse-tag contracts. Phase 3 step 3 had compared `state.run.lastAnalysisDigest` since it was written, against a key nothing ever set and that the state schema did not declare, so the staleness branch was unreachable and every run reported fresh by default. Phase 1 now persists the digest and a `base_commit` anchor, and step 3 gained the repo-drift half the digest cannot see: a reused document keeps a matching digest precisely because its evidence inputs did not change, while the code underneath it moved. The second contract is the Section 14 tag reaching development: Phase 2 carries it onto the todo as `sourceTag` and Phase 3 treats it as an instruction, which is what stops a Reuse row from being re-implemented. Compression came first and took the four additions from 380 tokens to 214, by moving every rationale clause out of the phase docs: why the commit anchor exists rather than a digest recomputation lives in this note and the CHANGELOG, and the schema descriptions carry the field semantics. The baseline had 9 tokens of headroom, so no addition of any size could have fit without a bump. 250 was the smallest step that clears it, leaving 45 tokens. phase-3 and phase-4 remain amber. v16.13.0: total 57600 -> 57700 for the code-graph injection and the fable-rung switch. Phase 1 gained Step 2.6 (query the graph, hand Explore a ranked starting set), Phase 7 gained the post-branch graph refresh, and Phase 0 Step 0 gained one line: a prefs switch that resolves every preferredModel: fable persona to opus for the run, which also collapses the Phase 4 Claude Code panel from three reviewers to two. Compression came first and mostly structurally: of roughly 1,630 tokens of new contract text, 1,310 never entered this budget at all - the whole code-graph contract lives in multi-agent-refs/features/code-graph.md (604) and the fable switch's scope table, per-host effects and cost-accounting consequence live in features/model-fallback.md (+707), leaving the phase docs with the call, the pref that gates it and the one fact an agent cannot infer. Phase 4 was compressed on top of that: its TLDR restated the reviewer matrix 270 lines below it, so 36 tokens came back and the doc nets +6 despite carrying two new clauses. One of those clauses is a correction rather than a feature - the consensus rule still said reviewerCount is 2 on Claude Code, which stopped being true when the third reviewer landed in 16.12.0, and the cross-CLI smoke never caught it because it reads the matrix line instead. 100 was the smallest step that clears it, leaving 54 tokens. phase-3 and phase-4 remain amber. v16.17.0: total 57700 -> 57850 for the platform-parity cross-check. Phase 4 gained Step 1.8: when dev-context carries a counterpart app repo, the review compares the change against the other platform on four axes. Compression came first and structurally, as always - of roughly 1,610 tokens of new contract text, 1,490 never entered this budget at all, because the four axes, the file cap, the graph-query recipe, the read-only prohibitions and the rule that an extractor miss may not be reported as an absence all live in multi-agent-refs/platform-parity.md. The step itself was then cut from ~200 tokens to 120 by deleting everything the ref already owns, leaving the trigger, the pointer and the two facts an agent must not infer: the counterpart repo is read-only, and parity findings are never blocking. The baseline had 13 tokens of headroom, so no addition of any size could have fit without a bump. 150 was the smallest step that clears it, leaving 35 tokens. phase-3 and phase-4 remain amber, and phase-4's structural-compression debt still stands. v16.20.0: phase-4-review max 14750 -> 15150 and total 58250 -> 60250 for the cross-round review delta, the scope self-check handoff and the circuit-breaker wiring. Compression came first and structurally: of roughly 3,900 tokens of new contract text, 2,700 never entered this budget at all - the previous-round-findings block, the scope-self-check block, the Step 3.8 state merge, telemetry and picker wording live in multi-agent-refs/features/review-delta.md, and the scope-check record rules and consumers in features/scope-check.md - so the phase docs carry the call, the pref that gates it and the exit table. The Phase 3 stability rule and the trigger-3 write were cut twice more before the bump; phase-3 stays under its max (8692/8950). Phase 4 is the first per-phase max raised since v10.9.0: the doc gained three steps that cannot be inferred (a per-round triage file, a prefix block that changes what reviewers report, and a halt condition), and its structural-compression debt (the inline triage JSON shape, named at v15.10.0) still stands and is the next candidate. 15150 and 60250 were the smallest steps that clear it, leaving 25 and 45 tokens. v16.23.0: phase-0-init max 12400 -> 12500 and total 60250 -> 60500 for the widget-registration call and the accounting gate. Phase 0 gained the `tiles` call and the exit-3 rule, Phase 7 gained the run report; together they are contract an agent cannot infer - which call registers this host's widget, and that a completion is refused without recorded spend. Compression came first and twice, taking the new prose from 472 tokens to 255: the per-host call list moved into tracker-contract.md \"The card is not the widget\" and the record-then-rerun recovery into \"Accounting is a gate\", both outside this budget, leaving the phase docs with the call and the one fact that cannot be looked up. 100 and 250 were the smallest steps that clear it, leaving 74 and 40 tokens. phase-3 and phase-4 remain amber. v16.24.0: total 60500 -> 60750 for visual evidence. Four phase docs gained one instruction each that cannot be inferred: Phase 0 keeps the issue's own images as the pre-fix evidence, Phase 3 captures the fixed state (there and not Phase 5, because every autopilot and --local entry drops Phase 5), Phase 5 hosts the flow recording when it runs, and Phase 6 blocks on a required artefact that is neither attached nor explained. Compression came first and twice, 42 tokens back, and the contract itself never entered this budget: the trigger matrix, the three video tiers, the size-degradation ladder and both render shapes live in multi-agent-refs/features/visual-evidence.md. phase-0-init cleared its own max without a bump. 250 was the smallest step that clears the aggregate. phase-3 and phase-4 remain amber. v17.0.0: phase-0-init max 13000 -> 13100 and total 62400 -> 62500 for the evidence-verdict writer. Phase 0 Step 7.7 probed with `--platform \"$PLATFORM\"`, a variable no phase document ever assigned, and it was gated on `visualEvidence.required`, which no phase document ever wrote - five readers, zero writers - so the step, Phase 3's capture and Phase 6's blocker were all unreachable and the pipeline reported nothing wrong. The step now writes the verdict and derives the platform from the stack, skipping the probe with a recorded reason when there is no device platform rather than passing the empty string the probe refuses with exit 2. Compression came first and three times, 30 tokens back from the Step 7.7 index rule that restated Step 7.5 verbatim and 55 from the new block itself; the reasoning never entered this budget, because who writes the verdict and how the platform is derived live in features/visual-evidence.md sections 1a and 1b. phase-3-dev max 9250 -> 9300 in the same change: it reads the platform back from state and re-decides the provisional verdict before capturing, which is the half of the fix that makes Phase 3 honest rather than merely reachable. Compressed three times first, 29 tokens back, by pointing its Phase-5 rationale and its tier mapping at visual-evidence.md sections 3 and 4.3 where both already live. 100, 50 and 150 were the smallest steps that clear it, leaving 24, 15 and 26 tokens. phase-3 and phase-4 remain amber."
|
|
31
|
+
"total_max_tokens": 62700
|
|
41
32
|
}
|