@mmerterden/multi-agent-pipeline 19.1.4 → 20.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +123 -0
- package/README.md +19 -36
- package/README.tr.md +18 -35
- package/SECURITY.md +3 -3
- package/docs/adr/0002-instruction-driven-flag.md +6 -5
- package/docs/adr/0005-lazy-phase-docs.md +2 -2
- package/docs/adr/0008-installer-modularization-and-secret-leak-defense.md +1 -0
- package/docs/adr/0009-claude-stack-skills-plugin-only.md +1 -1
- package/docs/adr/0010-own-code-graph.md +5 -4
- package/docs/adr/0011-dormant-ci.md +10 -1
- package/docs/adr/0012-macos-only.md +2 -2
- package/docs/adr/0013-lsp-code-intelligence.md +2 -2
- package/docs/adr/0014-six-phase-consolidation.md +9 -9
- package/docs/adr/0015-one-pipeline-no-depth-answer.md +83 -0
- package/docs/adr/0016-the-run-shape-is-asked-not-typed.md +69 -0
- package/docs/adr/README.md +18 -16
- package/docs/architecture.md +2 -2
- package/docs/ecosystem.md +5 -5
- package/docs/facts.json +7 -9
- package/docs/features.md +4 -5
- package/docs/token-budget-history.md +1 -1
- package/install/_codex-agents.mjs +1 -1
- package/install/_common.mjs +9 -1
- package/install/templates/copilot-instructions.md +7 -16
- package/manifest.json +133 -129
- package/package.json +1 -1
- package/pipeline/agents/code-reviewer.md +2 -2
- package/pipeline/agents/dev-critic.md +5 -5
- package/pipeline/agents/security-auditor.md +80 -72
- package/pipeline/commands/figma-to-swiftui.md +1 -1
- package/pipeline/commands/multi-agent/SKILL.md +7 -9
- package/pipeline/commands/multi-agent/analysis/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/analysis-jira/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/analysis-resolve/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/autopilot/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/autopilot-on/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/autopilot-status/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/build-optimize/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/channels/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/create-jira/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/design-check/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/diff-explain/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/forget/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +4 -2
- package/pipeline/commands/multi-agent/help/SKILL.md +23 -27
- package/pipeline/commands/multi-agent/ios-coding-standard/SKILL.md +5 -4
- package/pipeline/commands/multi-agent/issue/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/jira/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/language/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/prune-logs/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/purge/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/resume/SKILL.md +177 -48
- package/pipeline/commands/multi-agent/save/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/scan/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/security-review/SKILL.md +52 -0
- package/pipeline/commands/multi-agent/stack/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/sync/SKILL.md +7 -8
- package/pipeline/commands/multi-agent/test-screenshots/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/uninstall/SKILL.md +2 -0
- package/pipeline/lib/repo-hygiene.sh +1 -1
- package/pipeline/multi-agent-refs/analysis/render.md +1 -1
- package/pipeline/multi-agent-refs/analysis/resolve.md +1 -1
- package/pipeline/multi-agent-refs/analysis/synthesis.md +1 -1
- package/pipeline/multi-agent-refs/analysis-template.md +1 -1
- package/pipeline/multi-agent-refs/component-dispatch.md +5 -13
- package/pipeline/multi-agent-refs/cross-cli-contract.md +14 -15
- package/pipeline/multi-agent-refs/features/external-context-injection.md +2 -0
- package/pipeline/multi-agent-refs/features/review-delta.md +1 -1
- package/pipeline/multi-agent-refs/features/review-multi-repo.md +3 -3
- package/pipeline/multi-agent-refs/features/security-audit.md +55 -0
- package/pipeline/multi-agent-refs/features/skill-conformance.md +1 -1
- package/pipeline/multi-agent-refs/features/visual-evidence.md +2 -1
- package/pipeline/multi-agent-refs/features/worktree-finalize.md +1 -1
- package/pipeline/multi-agent-refs/generate-issue.md +2 -0
- package/pipeline/multi-agent-refs/issue-jira-triad.md +2 -0
- package/pipeline/multi-agent-refs/keychain.md +2 -0
- package/pipeline/multi-agent-refs/knowledge.md +0 -7
- package/pipeline/multi-agent-refs/outside-the-pipeline.md +1 -1
- package/pipeline/multi-agent-refs/payload-contracts.md +1 -1
- package/pipeline/multi-agent-refs/phases/modes.md +33 -109
- package/pipeline/multi-agent-refs/phases/operations.md +2 -0
- package/pipeline/multi-agent-refs/phases/phase-0-init.md +23 -42
- package/pipeline/multi-agent-refs/phases/phase-1-plan.md +7 -18
- package/pipeline/multi-agent-refs/phases/phase-2-dev.md +13 -44
- package/pipeline/multi-agent-refs/phases/phase-3-review.md +28 -37
- package/pipeline/multi-agent-refs/phases/phase-4-commit.md +6 -6
- package/pipeline/multi-agent-refs/phases/phase-5-report.md +3 -3
- package/pipeline/multi-agent-refs/phases.md +9 -11
- package/pipeline/multi-agent-refs/progress-contract.md +1 -1
- package/pipeline/multi-agent-refs/readiness-review.md +2 -0
- package/pipeline/multi-agent-refs/rules.md +1 -1
- package/pipeline/multi-agent-refs/threat-model.md +39 -0
- package/pipeline/multi-agent-refs/tracker-contract.md +9 -40
- package/pipeline/multi-agent-refs/wiki-capture.md +3 -2
- package/pipeline/preferences-template.json +2 -2
- package/pipeline/rules/figma-pipeline.md +1 -1
- package/pipeline/schemas/agent-state.schema.json +28 -10
- package/pipeline/schemas/migrations/prefs-2.7.0-to-2.8.0.mjs +33 -0
- package/pipeline/schemas/phases.json +4 -26
- package/pipeline/schemas/prefs.schema.json +5 -9
- package/pipeline/schemas/reviewer-output.schema.json +99 -2
- package/pipeline/schemas/security-finding.schema.json +144 -0
- package/pipeline/scripts/_stack-routing.mjs +1 -0
- package/pipeline/scripts/cost-table.json +1 -1
- package/pipeline/scripts/gc-abandoned.sh +16 -9
- package/pipeline/scripts/gc-refs.sh +1 -1
- package/pipeline/scripts/gen-mode-dispatch.mjs +11 -41
- package/pipeline/scripts/migrate-prefs.mjs +18 -17
- package/pipeline/scripts/phase-tracker.sh +2 -2
- package/pipeline/scripts/phase0-exit-gate.mjs +1 -1
- package/pipeline/scripts/plan-coverage-gate.mjs +3 -3
- package/pipeline/scripts/render-work-summary.sh +7 -4
- package/pipeline/scripts/run-aggregator.mjs +1 -1
- package/pipeline/scripts/usage-report.mjs +0 -2
- package/pipeline/scripts/worktree-finalize.sh +2 -2
- package/pipeline/skills/.skill-manifest.json +17 -21
- package/pipeline/skills/.skills-index.json +6 -39
- package/pipeline/skills/shared/README.md +5 -8
- package/pipeline/skills/shared/core/multi-agent/SKILL.md +11 -15
- package/pipeline/skills/shared/core/multi-agent-autopilot-status/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +13 -16
- package/pipeline/skills/shared/core/multi-agent-ios-coding-standard/SKILL.md +2 -3
- package/pipeline/skills/shared/core/multi-agent-resume/SKILL.md +51 -15
- package/pipeline/skills/shared/core/multi-agent-scan/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-security-review/SKILL.md +29 -0
- package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +7 -7
- package/pipeline/skills/shared/external/security-review/SKILL.md +64 -0
- package/pipeline/skills/shared/external/security-review/references/owasp-mobile-top10-2024.md +53 -0
- package/pipeline/skills/shared/external/security-review/references/owasp-web-api-top10-2021.md +56 -0
- package/pipeline/skills/skills-index.md +3 -6
- package/pipeline/commands/multi-agent/local/SKILL.md +0 -132
- package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +0 -142
- package/pipeline/commands/multi-agent/resume-local/SKILL.md +0 -114
- package/pipeline/commands/security-review.md +0 -6
- package/pipeline/skills/shared/core/multi-agent-local/SKILL.md +0 -41
- package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +0 -55
- package/pipeline/skills/shared/core/multi-agent-resume-local/SKILL.md +0 -51
|
@@ -1,18 +1,16 @@
|
|
|
1
|
-
> **TLDR** - four entry commands,
|
|
1
|
+
> **TLDR** - four entry commands, one pipeline, two axes:
|
|
2
2
|
>
|
|
3
|
-
> -
|
|
4
|
-
> -
|
|
5
|
-
> - **`--local`** means no worktree: work happens in `$PROJECT_ROOT` on a local branch.
|
|
3
|
+
> - **`autopilot`** skips confirmations (Plan, Test, Commit, PR prompts). Still fails safe on review blockers and build retries.
|
|
4
|
+
> - **Workspace** is the one question the run asks about its own shape: a worktree, or the project root on a local branch. Phase 0 Step 5b asks it; autopilot resolves it to a worktree and never asks.
|
|
6
5
|
>
|
|
7
|
-
>
|
|
6
|
+
> There is one pipeline and every mode runs its whole phase set. What a run costs follows the evidence the task carries, not an answer taken before the evidence exists.
|
|
8
7
|
|
|
9
8
|
## Autopilot Mode
|
|
10
9
|
|
|
11
10
|
<!-- toc -->
|
|
12
11
|
- [Autopilot Mode](#autopilot-mode)
|
|
13
|
-
- [Pipeline depth (Full / Short)](#pipeline-depth-full-short)
|
|
14
12
|
- [Analysis Mode (`/multi-agent:analysis`)](#analysis-mode-multi-agentanalysis)
|
|
15
|
-
- [Local Mode
|
|
13
|
+
- [Local Mode](#local-mode)
|
|
16
14
|
<!-- /toc -->
|
|
17
15
|
|
|
18
16
|
Autopilot mode skips interactive confirmations and runs the pipeline end-to-end autonomously.
|
|
@@ -27,15 +25,19 @@ Autopilot mode skips interactive confirmations and runs the pipeline end-to-end
|
|
|
27
25
|
/multi-agent "LoginView dark mode fix" autopilot
|
|
28
26
|
```
|
|
29
27
|
|
|
30
|
-
**What changes in autopilot (Tablo 1
|
|
28
|
+
**What changes in autopilot (Tablo 1)**
|
|
31
29
|
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
| Phase
|
|
37
|
-
|
|
|
38
|
-
|
|
|
30
|
+
Two axes, and only the first is a command: whether anything is confirmed
|
|
31
|
+
(`autopilot`), and where the work happens (the Step 5b answer). An autopilot run
|
|
32
|
+
always gets a worktree, so there is no unattended-local column.
|
|
33
|
+
|
|
34
|
+
| Phase | Interactive, worktree | Interactive, local | autopilot (always worktree) |
|
|
35
|
+
| --- | --- | --- | --- |
|
|
36
|
+
| Phase 1 (Plan Approval Gate) | Clarification (max 2 rounds) + approval loop - user: approve/abort/free-text edit | Same as worktree | **Gate skip** - log the plan, proceed directly to Phase 2 (autopilot contract: zero interaction) |
|
|
37
|
+
| Phase 3 (User Test step) | Interactive prompt ("Want to test?" -> wait) | **Not offered** - the step checks the change out of a worktree and there is none; the phase itself still runs | Skip (autopilot suppresses interactive prompts) -> proceed directly to Phase 4 |
|
|
38
|
+
| Phase 4 (Commit) | "Want to commit?" -> wait | "Want to commit?" -> wait | Auto commit + push |
|
|
39
|
+
| Phase 4 (PR) | "Want to open a PR?" -> wait | "Want to open a PR?" -> wait | Auto create PR |
|
|
40
|
+
| **Phase 5 (Channels)** | Multi-select channel + content menu | Multi-select (same) | **STILL PAUSES** - see "Phase 5 autopilot exception" |
|
|
39
41
|
|
|
40
42
|
**What NEVER skips (even in autopilot):**
|
|
41
43
|
|
|
@@ -51,90 +53,20 @@ Autopilot mode skips interactive confirmations and runs the pipeline end-to-end
|
|
|
51
53
|
|
|
52
54
|
The generic "zero-interaction" contract covers Phases 0-4 only. Phase 5 channels dispatch is the **single exception**:
|
|
53
55
|
|
|
54
|
-
- ALL modes
|
|
56
|
+
- ALL modes, `autopilot` included, pause at the channels multi-select menu.
|
|
55
57
|
- Menu pre-ticks from `prefs.global.reportChannels` + `prefs.global.reportContent` - user can accept with one keypress if prefs are stable.
|
|
56
58
|
- **30-minute timeout** - if user does not respond, session ends cleanly:
|
|
57
59
|
- External delivery aborted (no silent apply - prevents accidental Jira comments / Confluence pages).
|
|
58
60
|
- Internal capture (`agent-log.md`, telemetry, knowledge base) STILL runs.
|
|
59
|
-
- State persisted as `{phase:
|
|
61
|
+
- State persisted as `{phase: 5, waitingFor: "user-channels-choice", channelsTimeout: true}`.
|
|
60
62
|
- Resume: `/multi-agent:resume <task-id>` re-opens menu with same inputs.
|
|
61
63
|
- Post-hoc `/multi-agent:channels <task>` never times out - user invoked it explicitly.
|
|
62
64
|
|
|
63
65
|
Full contract: `$HOME/.claude/multi-agent-refs/phases/phase-5-report.md` (Autopilot pause contract) + `commands/multi-agent/channels/SKILL.md`.
|
|
64
66
|
|
|
65
|
-
---
|
|
66
|
-
|
|
67
|
-
## Pipeline depth (Full / Short)
|
|
68
|
-
|
|
69
|
-
Short strips the pipeline to what an already-scoped task needs: no deep analysis, no planning phase. The **Opus** dev agent reads the task scope and implements, and Phase 3 then reviews what it produced. Review is deliberately NOT part of the strip: analysis and planning shape work that has not happened yet, so a task the user has already scoped can skip them, while review judges work that now exists and has no substitute.
|
|
70
|
-
|
|
71
|
-
**How it is chosen.** Phase 0 Step 7.5, after `taskType` is known:
|
|
72
|
-
|
|
73
|
-
```
|
|
74
|
-
Bu is icin hangi pipeline? / Which pipeline for this task?
|
|
75
|
-
1. Tam / Full Analysis -> Plan -> Dev -> Review -> Test -> Commit -> Report
|
|
76
|
-
2. Kisa / Short Dev (self-contained, Opus) -> Review -> Test -> Commit -> Report
|
|
77
|
-
```
|
|
78
|
-
|
|
79
|
-
`bugfix` and `chore` recommend Short; `feature`, `refactor` and `component` recommend Full. The recommendation is presented first and passed as `ASK_CHOICE_DEFAULT` on hosts without a native picker, because `ask-choice.sh` picks the first option on a non-TTY and option order is not a contract. Pass it as the **1-based index** (Full = 1, Short = 2), not as the label: labels render in `outputLanguage`, so a label-valued default matches nothing on a `tr` run.
|
|
80
|
-
|
|
81
|
-
**Who is asked.** `/multi-agent` and `/multi-agent:local`. Both autopilot entries always run Full without asking.
|
|
82
|
-
|
|
83
|
-
**Why there is no fast-and-unattended combination.** It existed until v16.0.0, and removing it was a real behaviour change, not a rename. Autopilot may not ask, so something has to choose, and unattended is the worst place to drop analysis and planning: nobody is watching to notice what the shortcut lost. A cron job or script that wants both now has to pick - stay unattended and pay for the full pipeline, or stay fast and have a person present.
|
|
84
|
-
|
|
85
|
-
**Pipeline in a Short run:**
|
|
86
|
-
|
|
87
|
-
```
|
|
88
|
-
Phase 0: Init -> Phase 2: Dev (self-contained) -> Phase 3: Review -> Phase 3: Review (user test) -> Phase 4: Commit -> Phase 5: Report
|
|
89
|
-
```
|
|
90
|
-
|
|
91
|
-
**What changes (Tablo 2 - Short runs):**
|
|
92
|
-
|
|
93
|
-
| Phase | Full | Short | Short + `--local` |
|
|
94
|
-
| ------------------- | ----------------------------------------------- | ---------------------------------------------------------------------------------- | --------------------------------------------- |
|
|
95
|
-
| Phase 0 (Init) | Full setup | Same - worktree, branch, state, and the depth question itself | Same - no worktree, branch on `$PROJECT_ROOT` |
|
|
96
|
-
| Phase 1 (Analysis) | Parallel Explore agents + analysis document | **SKIP** - no tile is ever drawn for it (registration is deferred to Step 7.5) | **SKIP** |
|
|
97
|
-
| Phase 1 (Planning) | TaskCreate + architecture review + **Plan Approval Gate** | **SKIP** (no plan means no plan gate) | **SKIP** |
|
|
98
|
-
| Phase 2 (Dev) | Follows the Phase 1 plan, TDD cycle (Sonnet) | **Self-contained** (Opus): agent scans relevant files, implements with TDD, builds | Same, on the local branch |
|
|
99
|
-
| Phase 3 (Review) | Parallel review + Fable triage (3 reviewers on every host: Claude Code Fable + Opus + Sonnet, Copilot GPT-5.4 + Opus + Sonnet) | **Same** - gates, parallel review, triage; blocking findings return to Phase 2 (cap 3) | **Same**, on the local branch diff |
|
|
100
|
-
| Phase 3 (User Test) | Interactive prompt | **Interactive prompt** | Not in the set - no worktree to check out |
|
|
101
|
-
| Phase 4 (Commit) | Commit + PR | Same - still asks | Same |
|
|
102
|
-
| Phase 5 (Report) | Full report + channels multi-select | Simplified - no analysis section, review section IS present, channels menu still pauses | Same |
|
|
103
|
-
|
|
104
|
-
**Phase 2 in a Short run (self-contained):**
|
|
105
|
-
|
|
106
|
-
The **Opus** agent receives the task description (from Jira, GitHub issue, or free-text) and:
|
|
107
|
-
|
|
108
|
-
1. Quickly scans relevant files in the codebase (lightweight, not full Explore)
|
|
109
|
-
2. Implements the change following TDD cycle (test -> code -> build)
|
|
110
|
-
3. Runs build verification (max 3 retries on failure)
|
|
111
|
-
|
|
112
|
-
No separate task breakdown - the agent handles scope autonomously.
|
|
113
|
-
|
|
114
|
-
**State tracking**: `agent-state.json` gets `"onlyDevelop": true`. The tracker boots at Step -1 with Phase 0 alone and the rest of the tiles are registered at Step 7.5, once this answer says which phases the run actually has - so a Short run never draws an Analysis tile it will not use. Full contract: `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "Deferred registration".
|
|
115
|
-
|
|
116
|
-
### Intake warnings for a Short run
|
|
117
|
-
|
|
118
|
-
**An analysis document was supplied.** Because Phase 1 and Phase 1 are skipped, no phase turns that document into a task breakdown. The doc becomes raw context for one Dev pass, and work comes out ordered by whatever the model read first: the bottom of the dependency chain lands, the screen wiring does not.
|
|
119
|
-
|
|
120
|
-
The depth picker is where this is caught. When the intake carried an analysis document or a Figma reference, say so in the question itself rather than after the choice:
|
|
121
|
-
|
|
122
|
-
```
|
|
123
|
-
This task carries an analysis document. Short skips Analysis and Planning, so
|
|
124
|
-
the document will not be turned into a task breakdown.
|
|
125
|
-
1. Tam / Full (recommended here)
|
|
126
|
-
2. Kisa / Short (doc as context only)
|
|
127
|
-
```
|
|
128
|
-
|
|
129
|
-
Autopilot never sees this: it runs Full.
|
|
130
|
-
|
|
131
|
-
**The branch already carries the work.** When the change was developed outside the pipeline, or by hand, Short is still the wrong entry point: it will try to develop again. `/multi-agent:resume-local` puts the existing diff through the same review, adds a build+test success gate, opens the PR, and posts the Jira technical-analysis + test-scenario comment - without re-developing. Offer it when the working tree or branch is already ahead of the base with the task's changes.
|
|
132
|
-
|
|
133
|
-
---
|
|
134
|
-
|
|
135
67
|
## Analysis Mode (`/multi-agent:analysis`)
|
|
136
68
|
|
|
137
|
-
|
|
69
|
+
A different **output kind**, not a shorter pipeline. It produces the analysis document and stops; no code, no branch, no PR. The 6-phase contract holds, with four phases reinterpreted the same way a Short run reinterprets Phase 2.
|
|
138
70
|
|
|
139
71
|
```
|
|
140
72
|
Phase 0: Init -> Phase 1: Plan (analysis) -> Phase 1: Plan -> Phase 3: Review -> Phase 4: Publish -> Phase 5: Report
|
|
@@ -151,18 +83,19 @@ Phase 0: Init -> Phase 1: Plan (analysis) -> Phase 1: Plan -> Phase 3: Review ->
|
|
|
151
83
|
| 6 Commit | **Publish** instead: Local file / Confluence / Jira. Locked 6 still forbids `git add` and `git commit` |
|
|
152
84
|
| 7 Report | Channels, same as every mode |
|
|
153
85
|
|
|
154
|
-
**No `
|
|
86
|
+
**No `autopilot` variant.** Worktree isolation buys nothing when no code is written, and the intake, the Pass B convention preview and the open-question resolution are interactive by nature; a zero-interaction analysis would be a document nobody agreed to.
|
|
155
87
|
|
|
156
88
|
**Where the document lands.** Drafts go to `/tmp/analysis-<slug>-<ts>/` first, then the Phase 4 picker decides: Local writes `<repo>/analysis/<feature>-<platform>.md` uncommitted, Confluence posts the page, Jira updates the description. In the full pipeline the same engine writes into the worktree instead and `prefs.global.analysisPhase.commitDoc` decides whether it rides along with the commit.
|
|
157
89
|
|
|
158
90
|
---
|
|
159
91
|
|
|
160
|
-
## Local Mode
|
|
92
|
+
## Local Mode
|
|
161
93
|
|
|
162
94
|
Local mode skips worktree creation - works directly on a local branch in the project root. Useful for single-task workflows or when worktrees cause issues.
|
|
163
95
|
|
|
164
|
-
**Activation**: answer the Phase 0 Step 5b workspace question
|
|
165
|
-
|
|
96
|
+
**Activation**: answer the Phase 0 Step 5b workspace question. There is no flag
|
|
97
|
+
and no command name for it - the answer is the only way in, which is why the
|
|
98
|
+
question carries what the choice costs.
|
|
166
99
|
|
|
167
100
|
```
|
|
168
101
|
Bu is nerede kossun? / Where should this task run?
|
|
@@ -171,25 +104,16 @@ Bu is nerede kossun? / Where should this task run?
|
|
|
171
104
|
```
|
|
172
105
|
|
|
173
106
|
Two genuine options, so it meets the two-option floor in `picker-contract.md` on
|
|
174
|
-
its own. **Say what local costs inside the question**:
|
|
175
|
-
|
|
176
|
-
|
|
177
|
-
user choosing local should learn both before
|
|
178
|
-
|
|
179
|
-
The ways to state it up front:
|
|
180
|
-
|
|
181
|
-
```
|
|
182
|
-
/multi-agent:local "PROJ-12345"
|
|
183
|
-
/multi-agent:local "#316"
|
|
184
|
-
/multi-agent:local "LoginView dark mode fix"
|
|
185
|
-
/multi-agent:local-autopilot "PROJ-12345"
|
|
186
|
-
/multi-agent "PROJ-12345" --local
|
|
187
|
-
```
|
|
107
|
+
its own. **Say what local costs inside the question**: the user-test STEP inside
|
|
108
|
+
Phase 3 checks the change out of a worktree, so a local run does not get that
|
|
109
|
+
offer (the phase itself still runs), and uncommitted work in the project root is
|
|
110
|
+
in the way of the checkout. A user choosing local should learn both before
|
|
111
|
+
choosing, not after.
|
|
188
112
|
|
|
189
113
|
Every autopilot entry resolves this to **worktree** and never asks. That is not a
|
|
190
114
|
skipped question: an unattended run commits and pushes from wherever it stands,
|
|
191
|
-
and doing that in the user's own checkout is what worktrees exist to prevent.
|
|
192
|
-
|
|
115
|
+
and doing that in the user's own checkout is what worktrees exist to prevent. An
|
|
116
|
+
unattended run therefore always gets its own checkout.
|
|
193
117
|
|
|
194
118
|
**What changes in local mode:**
|
|
195
119
|
|
|
@@ -221,4 +145,4 @@ git -C $PROJECT_ROOT config user.email "{identity.email}"
|
|
|
221
145
|
- Uncommitted changes in PROJECT_ROOT may conflict - pipeline warns if dirty
|
|
222
146
|
- Build queue lock still applies for xcodebuild
|
|
223
147
|
|
|
224
|
-
|
|
148
|
+
Fast path.
|
|
@@ -17,6 +17,8 @@
|
|
|
17
17
|
- [Phase Pipeline](#phase-pipeline)
|
|
18
18
|
<!-- /toc -->
|
|
19
19
|
|
|
20
|
+
> **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
|
|
21
|
+
|
|
20
22
|
Every task gets an auto-incremented short ID. Counter stored at `$HOME/.claude/logs/multi-agent/{project}/.counter` (persists across sessions).
|
|
21
23
|
|
|
22
24
|
```
|
|
@@ -14,23 +14,21 @@ $HOME/.claude/scripts/phase-tracker.sh tiles
|
|
|
14
14
|
$HOME/.claude/scripts/phase-tracker.sh update 0 in_progress
|
|
15
15
|
```
|
|
16
16
|
|
|
17
|
-
**
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
whole set here, in phase-number order. Contract: `tracker-contract.md`,
|
|
21
|
-
"Deferred registration".
|
|
17
|
+
**Every mode registers its whole set here**, in phase-number order. The phase
|
|
18
|
+
set is known at Phase 0 because there is one pipeline: no answer later in the
|
|
19
|
+
run can add or remove a phase. Contract: `tracker-contract.md`.
|
|
22
20
|
|
|
23
21
|
`tiles` prints this host's widget-registration calls: **make them before continuing.** The card alone lands in collapsed tool output, so a run that skips them runs in silence. Contract: `tracker-contract.md`, "The card is not the widget".
|
|
24
22
|
|
|
25
23
|
If `INPUT_TASK_ID` isn't known yet (free-text, project not selected), use a placeholder; rename later via `mv` once parsed in Step 1.
|
|
26
24
|
|
|
27
|
-
Every subsequent phase (1-
|
|
25
|
+
Every subsequent phase (1-5) MUST call `phase-tracker.sh update <N> in_progress` on entry and `phase-tracker.sh update <N> completed|failed|skipped` on exit. Each `update` prints a `-- NEXT (required) --` block: act on it. Sub-phase milestones use `phase-tracker.sh sub <N> <subN> "<name>" <status>`. See `$HOME/.claude/multi-agent-refs/phases.md` "Visual Phase Tracker" for the full contract.
|
|
28
26
|
|
|
29
|
-
`update <N> completed` **exits 3** for phases 1-
|
|
27
|
+
`update <N> completed` **exits 3** for phases 1-5 with no recorded spend: record `model` + `tokens`, or pass `--no-llm`, then re-run it. Contract: `tracker-contract.md`, "Accounting is a gate".
|
|
30
28
|
|
|
31
29
|
##### TaskCreate ordering on Claude Code (strict)
|
|
32
30
|
|
|
33
|
-
On Claude Code, fire every `TaskCreate` in a registration batch in strict phase-number order BEFORE any `TaskUpdate` in that batch - which is the order `tiles` prints them in - and never register a phase whose number is below one already registered.
|
|
31
|
+
On Claude Code, fire every `TaskCreate` in a registration batch in strict phase-number order BEFORE any `TaskUpdate` in that batch - which is the order `tiles` prints them in - and never register a phase whose number is below one already registered. Full contract: `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
|
|
34
32
|
|
|
35
33
|
---
|
|
36
34
|
|
|
@@ -46,6 +44,13 @@ OUTPUT_LANG=$(jq -r '.global.outputLanguage // "en"' "$PREFS_FILE" 2>/dev/null |
|
|
|
46
44
|
|
|
47
45
|
From this point on, everything the user reads renders in `$OUTPUT_LANG`: conversational lines, `AskUserQuestion` `question`/`label`/`description`, and external payload bodies (PR/Jira/Confluence). English stays only on `header`, commit messages, branch names, PR title prefixes, identifiers. Full matrix: `rules.md` "Language Application".
|
|
48
46
|
|
|
47
|
+
**Every picker in this phase prints its breadcrumb.** Phase 0 is one chain -
|
|
48
|
+
account, repos, base branch, maturity, dev-context, workspace - and each step
|
|
49
|
+
emits the narrator line `<localized: "Step i/n: what this step decides">` above
|
|
50
|
+
its question, per `picker-contract.md` "Step narration". A step that resolves
|
|
51
|
+
without asking (one account, no siblings) still prints its line with the
|
|
52
|
+
resolution noted, so the numbering reads continuously instead of jumping.
|
|
53
|
+
|
|
49
54
|
**Model tier resolution** (same step, once per run): read `prefs.global.modelFallback`. If `fableEnabled` is `false`, every `preferredModel: fable` persona resolves to `opus` for this run and the Phase 3 Claude Code panel is 2 reviewers, not 3; print the one-line INFO. Then, if `premiumTierUntil` is set and in the past, apply the date-gate trigger - `preferredModel` personas dispatch on `fallbackModel`, with the one-line WARN. Both lines and the exact ordering: `$HOME/.claude/multi-agent-refs/features/model-fallback.md`. Dispatch-error and budget triggers apply per-dispatch later; nothing else to do here.
|
|
50
55
|
|
|
51
56
|
**First-run guard**: After loading prefs, check if `keychainMapping` has at least one non-null value. If ALL values are null (template defaults - setup never ran), show:
|
|
@@ -323,8 +328,7 @@ above, recorded as `baseBranchSource: "input"`. Everything else asks, and
|
|
|
323
328
|
One row is a normal outcome of the rule-5 filter and is **still asked**, with a
|
|
324
329
|
real second option - `picker-contract.md`, "Two options or it is not a question".
|
|
325
330
|
|
|
326
|
-
This holds in every mode.
|
|
327
|
-
does not skip Phase 0's pickers. Autopilot resolves them without prompting, which still
|
|
331
|
+
This holds in every mode. Autopilot resolves them without prompting, which still
|
|
328
332
|
writes the fields; its base-branch resolution order is in `features/base-branch-evidence.md`.
|
|
329
333
|
|
|
330
334
|
**TTL filter for recent branches**:
|
|
@@ -412,11 +416,11 @@ Branch name is deterministic - no user confirmation needed.
|
|
|
412
416
|
#### Step 5b - Workspace (worktree or local)
|
|
413
417
|
|
|
414
418
|
Ask where the branch lives - the wording, the two options and what local costs
|
|
415
|
-
are in `modes.md`, "Local Mode". Here because Step 4 named the branch
|
|
416
|
-
acts on the answer
|
|
419
|
+
are in `modes.md`, "Local Mode". Here because Step 4 named the branch and Step
|
|
420
|
+
6b acts on the answer.
|
|
417
421
|
|
|
418
|
-
**Who is asked.**
|
|
419
|
-
|
|
422
|
+
**Who is asked.** Every interactive entry (`workspaceSource: "asked"`). There is
|
|
423
|
+
no flag and no command name that pre-answers it.
|
|
420
424
|
Every autopilot entry resolves it to a worktree and never asks (`autopilot`).
|
|
421
425
|
|
|
422
426
|
**Persist** `state.localMode` (semantics unchanged) and `state.workspaceSource`,
|
|
@@ -448,7 +452,7 @@ Log: `Identity: {identity.name} <{identity.email}>`
|
|
|
448
452
|
|
|
449
453
|
1. `git -C $PROJECT_ROOT fetch origin`
|
|
450
454
|
|
|
451
|
-
**If local** (Step 5b answered local
|
|
455
|
+
**If local** (Step 5b answered local):
|
|
452
456
|
|
|
453
457
|
```bash
|
|
454
458
|
if [ -n "$(git -C $PROJECT_ROOT status --porcelain)" ]; then
|
|
@@ -464,7 +468,7 @@ git -C $PROJECT_ROOT config user.email "{identity.email}"
|
|
|
464
468
|
|
|
465
469
|
**If worktree** (Step 5b answered worktree, or autopilot resolved it): 2. Worktree path: Jira → `.worktrees/{jiraId}/`, GitHub → `.worktrees/GH{issueNo}/`, free-text → `.worktrees/task-{shortId}/` 3. **Heal stale admin state first** (see "Worktree stale-lock heal" below) and **apply the residue guard** (see "Worktree residue guard" below), then `git -C $PROJECT_ROOT worktree add {path} -b {branch} origin/{baseBranch}` (if exists: enter, pull) 4. Set identity: `git -C {worktree-path} config user.name/email` 5. Create log dir + `agent-log.md` + `agent-state.json` at `$HOME/.claude/logs/multi-agent/{project}/{task-id}/`, never inside the worktree:
|
|
466
470
|
|
|
467
|
-
**Worktree location convention (cited by every other command):** always `{projectRoot}/.worktrees/{taskId}`, inside the repo, never under `$HOME`. `{taskId}` is the directory name from the rule above (`DC-<shortId>` for `/multi-agent:design-check`). `.worktrees` is fixed, not a preference: no `worktreeBasePath` key exists, and `gc-worktrees.sh`, `purge.sh`, the cost renderers and `usage-report.mjs` resolve `<repo>/.worktrees/` by name. Multi-repo tasks get one worktree per repo (the loop below);
|
|
471
|
+
**Worktree location convention (cited by every other command):** always `{projectRoot}/.worktrees/{taskId}`, inside the repo, never under `$HOME`. `{taskId}` is the directory name from the rule above (`DC-<shortId>` for `/multi-agent:design-check`). `.worktrees` is fixed, not a preference: no `worktreeBasePath` key exists, and `gc-worktrees.sh`, `purge.sh`, the cost renderers and `usage-report.mjs` resolve `<repo>/.worktrees/` by name. Multi-repo tasks get one worktree per repo (the loop below); a local answer creates none and `worktreePath` is `$PROJECT_ROOT`.
|
|
468
472
|
|
|
469
473
|
**Worktree stale-lock heal (required before every `worktree add`):** a run killed mid-`worktree add` (OOM, SIGTERM, disk full) leaves a locked or broken admin entry under `.git/worktrees/{id}/`, so the retry fails with `fatal: '<path>' already exists`. Always run the heal first - it is a no-op on a clean repo:
|
|
470
474
|
|
|
@@ -584,30 +588,6 @@ Persist: `"taskType": "component" | "bugfix" | "feature" | "refactor" | "chore"`
|
|
|
584
588
|
|
|
585
589
|
Log: `Phase 0 Step 7: taskType = {component|bugfix|feature|refactor|chore}`
|
|
586
590
|
|
|
587
|
-
#### Step 7.5 - Pipeline depth (Full / Short)
|
|
588
|
-
|
|
589
|
-
Ask the depth question from `$HOME/.claude/multi-agent-refs/phases/modes.md` "Pipeline depth" - it carries the wording, the per-`taskType` recommendation and the mode tables. Here because the recommendation needs `taskType` (Step 7), which needs the fetched issue (Step 1) and the branch (Step 3).
|
|
590
|
-
|
|
591
|
-
**Who is asked.** `/multi-agent` and `/multi-agent:local` only. Both autopilot entries and analysis mode skip it; autopilot always runs Full.
|
|
592
|
-
|
|
593
|
-
When the intake carried an analysis document or a Figma reference, say so **inside** the question: Short skips the only two phases that would turn that document into a task breakdown, and the user should learn that before choosing, not after.
|
|
594
|
-
|
|
595
|
-
**Pass the default as a 1-based index, never a label.** `ask-choice.sh` takes the first option on a non-TTY, and a label-valued default matches nothing once the options render in `outputLanguage`. Reasoning: `modes.md`, "Pipeline depth".
|
|
596
|
-
|
|
597
|
-
```bash
|
|
598
|
-
# Full is option 1, Short is option 2 (modes.md "Pipeline depth")
|
|
599
|
-
DEPTH_DEFAULT_INDEX=1; [ "$DEPTH_RECOMMENDATION" = "short" ] && DEPTH_DEFAULT_INDEX=2
|
|
600
|
-
ASK_CHOICE_DEFAULT="$DEPTH_DEFAULT_INDEX" \
|
|
601
|
-
$HOME/.claude/lib/ask-choice.sh "<localized: 'Which pipeline for this task?'>" \
|
|
602
|
-
"<localized: 'Full'>" "<localized: 'Short'>"
|
|
603
|
-
```
|
|
604
|
-
|
|
605
|
-
**Persist.** Short sets `state.onlyDevelop = true`; Full leaves it `false`. The key is unchanged - only who sets it changed - so every downstream reader keeps working.
|
|
606
|
-
|
|
607
|
-
**Now register the rest of the widget.** This answer is the first moment the phase set is known, so the remaining tiles are created here and not before - Full `1 2 3 4 5 6 7`, Short `3 4 5 6 7`, `:local` dropping 5 from either. `add` each, then `phase-tracker.sh tiles --new`, which emits TaskCreate only for phases that carry no tile yet. Contract: `tracker-contract.md`, "Deferred registration".
|
|
608
|
-
|
|
609
|
-
Log: `Phase 0 Step 7.5: depth = {full|short} (recommended {full|short}, source {user|autopilot|default})`
|
|
610
|
-
|
|
611
591
|
#### Step 7.6 - Test baseline (opt-in, `prefs.global.testBaseline.enabled`, default `false`)
|
|
612
592
|
|
|
613
593
|
Phase 3 Gate 3 cannot tell an inherited red suite from one this run broke, so it blocks on someone else's bug or the agent "fixes" tests it never touched. Runs after Step 6, only when the stack has a test command; skipped in analysis mode.
|
|
@@ -644,7 +624,8 @@ Persist that file as `state.evidenceCapability`, then build the menu from it, ne
|
|
|
644
624
|
|
|
645
625
|
A closed option keeps its row and prints the probe's reason verbatim. No `uiTestTargets` closes 2; `mcp` false closes 3; a missing `device` or `recorder` closes both. Targets present with no `matchingTests` leaves 2 open, warning that the whole UI suite will run. **When only option 1 is open, do not ask**: `testDepth = unit`, `testDepthSource = forced`, and the tier 3 gap is written.
|
|
646
626
|
|
|
647
|
-
Default as a 1-based index, never a label -
|
|
627
|
+
Default as a 1-based index, never a label: `ask-choice.sh` takes the first
|
|
628
|
+
option on a non-TTY, so a label-valued default silently becomes option 1.
|
|
648
629
|
|
|
649
630
|
```bash
|
|
650
631
|
DEPTH_DEFAULT_INDEX=1
|
|
@@ -653,7 +634,7 @@ DEPTH_DEFAULT_INDEX=1
|
|
|
653
634
|
ASK_CHOICE_DEFAULT="$DEPTH_DEFAULT_INDEX" $HOME/.claude/lib/ask-choice.sh ...
|
|
654
635
|
```
|
|
655
636
|
|
|
656
|
-
Asked by
|
|
637
|
+
Asked by every interactive entry; autopilot reads `prefs.global.testDepth.default` and degrades to the best open option. Asked here, not Phase 3, which four of the eight modes drop.
|
|
657
638
|
|
|
658
639
|
Log: `Phase 0 Step 7.7: testDepth = {unit|unit+ui|unit+mcp} (source {user|autopilot|default|forced}), tier1/tier2 = {open|closed}`
|
|
659
640
|
|
|
@@ -1,6 +1,8 @@
|
|
|
1
1
|
### Phase 1: Plan (Opus)
|
|
2
2
|
|
|
3
|
-
> **TLDR** - One phase, two halves. First the codebase is explored and the analysis document written (the design contract the rest of the run reads); then that document is decomposed into concrete tasks with file-level targets, risk grading and architecture review. The phase ends at the **Plan Approval Gate**: in normal mode the orchestrator asks structured clarification questions when scope is ambiguous (max 2 rounds), renders the plan, and loops on free-text edits until the user approves or aborts. The gate is **skipped entirely** for
|
|
3
|
+
> **TLDR** - One phase, two halves. First the codebase is explored and the analysis document written (the design contract the rest of the run reads); then that document is decomposed into concrete tasks with file-level targets, risk grading and architecture review. The phase ends at the **Plan Approval Gate**: in normal mode the orchestrator asks structured clarification questions when scope is ambiguous (max 2 rounds), renders the plan, and loops on free-text edits until the user approves or aborts. The gate is **skipped entirely** for `autopilot`, which has no one to ask.
|
|
4
|
+
|
|
5
|
+
> **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
|
|
4
6
|
|
|
5
7
|
<!-- progress-contract: applied -->
|
|
6
8
|
Progress emission per `$HOME/.claude/multi-agent-refs/progress-contract.md` - lines for each Explore dispatch, each finish, analyst synthesis start, `analysis.json` write.
|
|
@@ -441,10 +443,6 @@ Log: "Phase 1: Consistency - requirements:{N/N mapped} anchors:{ok|M unanchore
|
|
|
441
443
|
- `state.autopilot === true` (autopilot contract: zero interaction)
|
|
442
444
|
- Autopilot safety classifier returns `recommendPause: false` (see Step 5c below)
|
|
443
445
|
|
|
444
|
-
OR:
|
|
445
|
-
|
|
446
|
-
- `state.onlyDevelop === true` (Short pipeline: direct to Phase 2, no plan)
|
|
447
|
-
|
|
448
446
|
In the skipped case, log `🧠 Phase 1: Plan - gate skipped ({mode}), proceeding to Phase 2` and go to Phase 2.
|
|
449
447
|
|
|
450
448
|
##### 5c - Autopilot safety classifier (runs before 5a/5b skip decision)
|
|
@@ -571,9 +569,8 @@ The pipeline shapes interact with the gate as follows. This table is the source
|
|
|
571
569
|
|
|
572
570
|
| Mode | Clarification | Approval Loop | Safety Classifier | Notes |
|
|
573
571
|
|---|---|---|---|---|
|
|
574
|
-
|
|
|
575
|
-
|
|
|
576
|
-
| `autopilot`, `:local-autopilot` | ❌ | ❌ conditional | ✅ (if `autopilotSafetyGate !== false`) | Always Full, so a plan exists. Safe plans proceed silently; high-risk plans trigger a one-time manual approval. Log records the score. |
|
|
572
|
+
| Interactive (`/multi-agent`) | ✅ (max 2 rounds) | ✅ | - (redundant when a human approves) | Full gate |
|
|
573
|
+
| `/multi-agent:autopilot` | ❌ | ❌ conditional | ✅ (if `autopilotSafetyGate !== false`) | A plan always exists. Safe plans proceed silently; high-risk plans trigger a one-time manual approval. Log records the score. |
|
|
577
574
|
|
|
578
575
|
**Why autopilot now has an escape hatch:** the old "zero interaction - fully trust the scope" contract held well for tightly-scoped batch workflows (figma component iteration over known-safe components) but broke in edge cases - schema migrations auto-merging, security-path drift going silent, delete-without-test sprawls. The safety classifier (Step 5c) is opt-out so the default protects against the edge-case cost; users running known-safe workflows can flip `prefs.global.autopilotSafetyGate = false` to restore pre-v7.0 behavior.
|
|
579
576
|
|
|
@@ -582,18 +579,10 @@ The pipeline shapes interact with the gate as follows. This table is the source
|
|
|
582
579
|
After plan generation (and after each edit-loop iteration), forward the planning model's call totals so Phase 5's Cost Breakdown captures Phase 1 (`model=` names the rung that actually ran: `fable`, or `opus` after a fallback step):
|
|
583
580
|
|
|
584
581
|
```bash
|
|
585
|
-
LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID"
|
|
582
|
+
LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 1 plan.generated \
|
|
586
583
|
model=fable tokens_in=$IN tokens_out=$OUT duration_ms=$DUR iteration=$N
|
|
587
584
|
```
|
|
588
585
|
|
|
589
586
|
Best-effort. See `$HOME/.claude/multi-agent-refs/progress-contract.md#token-telemetry-forwarding`.
|
|
590
587
|
|
|
591
|
-
|
|
592
|
-
|
|
593
|
-
## Token telemetry - invoke after every LLM call
|
|
594
|
-
|
|
595
|
-
```bash
|
|
596
|
-
bash $HOME/.claude/scripts/phase-tracker.sh tokens 2 <input_count> <output_count>
|
|
597
|
-
```
|
|
598
|
-
|
|
599
|
-
Contract and rationale: `progress-contract.md` -> Token telemetry forwarding.
|
|
588
|
+
Both halves of this phase report into the same counter: see "Token telemetry" above.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
### Phase 2: Dev (Sonnet)
|
|
2
2
|
|
|
3
|
-
> **TLDR** - Sonnet executes the plan task-by-task with TDD (red→green→refactor). Required: issue-tracker status moved to "In Progress" before any code, with a post-mutation verify step (re-reads the field, retries once on silent VALIDATION failures). Build verification after each task (up to 3 retries). Build-queue lock serializes concurrent xcodebuild.
|
|
3
|
+
> **TLDR** - Sonnet executes the plan task-by-task with TDD (red→green→refactor). Required: issue-tracker status moved to "In Progress" before any code, with a post-mutation verify step (re-reads the field, retries once on silent VALIDATION failures). Build verification after each task (up to 3 retries). Build-queue lock serializes concurrent xcodebuild. `taskType === component` **short-circuits the TDD path** and delegates the whole phase to the enabled `ai-<platform>-toolkit` marketplace plugin's component skill (`create-component`, fallback `create-ui-component`) - see next subsection.
|
|
4
4
|
|
|
5
5
|
## Phase 2 Pre-flight (BLOCKING, v9.0.0)
|
|
6
6
|
|
|
@@ -8,11 +8,11 @@ Per Locked decision 30, Phase 2 Dev consumes the analysis document as the sole d
|
|
|
8
8
|
|
|
9
9
|
Pre-flight steps (run in order, abort on failure).
|
|
10
10
|
|
|
11
|
-
**Steps 1, 2, 3, 5 and 6
|
|
11
|
+
**Steps 1, 2, 3, 5 and 6 read the analysis document.** Phase 1 always runs, so the document always exists; what varies is how much evidence it carries. A section the evidence did not support is absent by the Locked 2 omission rule, and a step whose section is absent is recorded `not-applicable (no <section> in this document)` rather than aborting. Steps 4, 7, 8 and 9 read nothing from it and apply always.
|
|
12
12
|
|
|
13
13
|
1. **Analysis document presence** (Phase 1 modes only): read `state.analysis.docStatus` and `state.analysis.docPath[]`, both set by Phase 1 Step 4.
|
|
14
14
|
- `produced` | `reused` -> read the active platform's file; multi-repo runs need one per selected repo.
|
|
15
|
-
- `not-applicable` ->
|
|
15
|
+
- `not-applicable` -> the analysis mode exported a document this run does not consume; record it for steps 1, 2, 3, 5, 6 and skip them.
|
|
16
16
|
- **Abort**: `produced` but unreadable -> `ERR: analysis doc at <path> unreadable. Resume with /multi-agent:resume #N.` Producing it is Phase 1's job.
|
|
17
17
|
|
|
18
18
|
2. **Parse YAML front-matter** into `state.analysis.frontMatter`: `feature`, `platform`, `language`, `mode`, `evidence_digest`, `template_version`. **Abort** when `template_version` < `v3`; there is no degraded mode.
|
|
@@ -49,7 +49,7 @@ The analysis document is the SOLE design source in Phase 2. Variant choices, pad
|
|
|
49
49
|
|
|
50
50
|
#### Input contract
|
|
51
51
|
|
|
52
|
-
Phase 2 consumes the Phase 1 output object conforming to `$HOME/.claude/schemas/planning-output.schema.json` - the task graph (`tasks[]` with `id`, `title`, `type`, `files`, and optional `dependsOn` / `acceptanceCriteria`) plus the architecture review notes. Tasks execute in dependency order; the schema's `dependsOn` field drives the ready-task picker.
|
|
52
|
+
Phase 2 consumes the Phase 1 output object conforming to `$HOME/.claude/schemas/planning-output.schema.json` - the task graph (`tasks[]` with `id`, `title`, `type`, `files`, and optional `dependsOn` / `acceptanceCriteria`) plus the architecture review notes. Tasks execute in dependency order; the schema's `dependsOn` field drives the ready-task picker.
|
|
53
53
|
|
|
54
54
|
**Plan Todo iteration (opt-in)**: gated by `prefs.global.planTodos.enabled` (default: `false`). When enabled and Phase 1 Step 11 emitted a `plan.todos[]`, Phase 2 iterates via `$HOME/.claude/lib/plan-todos.sh next/start/complete/fail` instead of walking `tasks[]` directly. When disabled, the loop walks `tasks[]` from `planning-output` - TDD contract is unchanged. Full helper loop + state semantics: `$HOME/.claude/multi-agent-refs/features/plan-todos.md`. A todo with `sourceTag: Reuse` binds the file analysis already found; `Modify` edits in place. Writing a new file over a `Reuse` step is a Locked 11 violation and Phase 3 flags it.
|
|
55
55
|
|
|
@@ -57,7 +57,7 @@ Phase 2 consumes the Phase 1 output object conforming to `$HOME/.claude/schemas/
|
|
|
57
57
|
|
|
58
58
|
#### Component tasks - delegated dispatch (taskType === "component")
|
|
59
59
|
|
|
60
|
-
When Phase 0 Step 7 classified the task as `component`, Phase 2 delegates the entire phase to the enabled `ai-<platform>-toolkit` marketplace plugin's component skill (`create-component`, fallback `create-ui-component`) via the Skill tool and does NOT run the TDD loop below. The dispatch layer passes the plugin skill the analysis Section 6 (Bileşen Envanteri) entry + Section 13.1 conventions for the named component as context. Because plugin skills do not write pipeline state, the **dispatch layer** (not the skill) owns `state.phases["2"].subphases[]`, recording a coarse component-build row - multi-agent's `phase-tracker` reads that array with no special case. Plugin resolution (dual-name), dispatch call, failure/resume, multi-repo,
|
|
60
|
+
When Phase 0 Step 7 classified the task as `component`, Phase 2 delegates the entire phase to the enabled `ai-<platform>-toolkit` marketplace plugin's component skill (`create-component`, fallback `create-ui-component`) via the Skill tool and does NOT run the TDD loop below. The dispatch layer passes the plugin skill the analysis Section 6 (Bileşen Envanteri) entry + Section 13.1 conventions for the named component as context. Because plugin skills do not write pipeline state, the **dispatch layer** (not the skill) owns `state.phases["2"].subphases[]`, recording a coarse component-build row - multi-agent's `phase-tracker` reads that array with no special case. Plugin resolution (dual-name), dispatch call, failure/resume, multi-repo, and the intentional cross-CLI divergence live in `$HOME/.claude/multi-agent-refs/component-dispatch.md` - read it before editing component-task behaviour here. Phase 2 still owns: progress line `-> dispatching create-component <name>`, `retryCount` cap at 3, and fallthrough to the TDD path when dispatch prerequisites are missing (`taskType` absent OR the plugin is not enabled in this repo -> log anomaly, halt or run TDD per component-dispatch.md).
|
|
61
61
|
|
|
62
62
|
For non-component taskTypes (`bugfix`, `feature`, `refactor`, `chore`), continue with the standard TDD section below.
|
|
63
63
|
|
|
@@ -92,7 +92,7 @@ If the latest iteration has `triage.approved === true` AND `accepted === []`, Ph
|
|
|
92
92
|
**Telemetry**: at the start of every re-entry, emit:
|
|
93
93
|
|
|
94
94
|
```bash
|
|
95
|
-
$HOME/.claude/scripts/log-metric.sh "$TASK_ID"
|
|
95
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 2 rework.started \
|
|
96
96
|
iteration=$ITERATION accepted_blocking=$BLOCKING accepted_important=$IMPORTANT
|
|
97
97
|
```
|
|
98
98
|
|
|
@@ -275,12 +275,12 @@ After the build/test green step and BEFORE Phase 3 handoff, run one diff-shrink
|
|
|
275
275
|
5. **Record tokens in the cost ledger** so Phase 5's Cost Breakdown captures the pass:
|
|
276
276
|
|
|
277
277
|
```bash
|
|
278
|
-
LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID"
|
|
278
|
+
LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 2 dev.simplifier_pass \
|
|
279
279
|
model=sonnet tokens_in=$IN tokens_out=$OUT duration_ms=$DUR \
|
|
280
280
|
edits_returned=$RET edits_applied=$APPLIED edits_skipped=$SKIPPED
|
|
281
281
|
```
|
|
282
282
|
|
|
283
|
-
Scope guard: a single pass, never looped.
|
|
283
|
+
Scope guard: a single pass, never looped. Component tasks (`taskType === "component"`) skip it (the figma skill owns its own checklist).
|
|
284
284
|
|
|
285
285
|
---
|
|
286
286
|
|
|
@@ -297,37 +297,6 @@ Exit 1 lists `unjustified[]`: complete the record once and re-run. A second exit
|
|
|
297
297
|
|
|
298
298
|
---
|
|
299
299
|
|
|
300
|
-
#### Short pipeline (`state.onlyDevelop === true`)
|
|
301
|
-
|
|
302
|
-
Set by the Phase 0 Step 7.5 depth picker, or by autopilot never (autopilot always runs Full). When it is true, Phase 2 runs self-contained with **Opus** (not Sonnet). No Phase 1 plan exists - the agent creates its own scope.
|
|
303
|
-
|
|
304
|
-
**Flow:**
|
|
305
|
-
1. Read task description (from Jira, GitHub issue, or free-text)
|
|
306
|
-
2. Lightweight file scan - grep/glob for relevant code (not full Explore agents)
|
|
307
|
-
3. Determine scope autonomously - no task breakdown, no user confirmation
|
|
308
|
-
4. Implement with TDD cycle (same RED→GREEN→REFACTOR as normal mode)
|
|
309
|
-
5. Build verification (same lock, same retry logic)
|
|
310
|
-
6. Intermediate commits (same WIP pattern)
|
|
311
|
-
|
|
312
|
-
**Key differences from normal mode:**
|
|
313
|
-
|
|
314
|
-
| Aspect | Full | Short |
|
|
315
|
-
|--------|--------|----------|
|
|
316
|
-
| Model | Sonnet | **Opus** |
|
|
317
|
-
| Plan source | Phase 1 task list | Self-determined |
|
|
318
|
-
| Task granularity | Per-plan-item | Agent decides |
|
|
319
|
-
| Status updates | Per task item | Single in_progress → completed |
|
|
320
|
-
| Scope confirmation | Phase 1 user approval | None (agent autonomous) |
|
|
321
|
-
| Review of the result | Phase 3 | Phase 3 (same) |
|
|
322
|
-
|
|
323
|
-
Because the agent determines its own scope here, Phase 3 is the only place that checks the result against anything external. Record every skill, plugin skill and guide consulted during this phase into `state.telemetry.skillCalls[]` with the files it was applied to - Phase 3 resolves the criteria set independently, and this record is what lets it tell "applied and honoured" from "never opened".
|
|
324
|
-
|
|
325
|
-
**Never combined with autopilot.** Autopilot skips the depth question and runs Full, so `onlyDevelop` is false in every unattended run. "Fast plus unattended" was removed in v16.0.0 and no longer exists: something has to choose when nobody is asked, and unattended is the worst place to drop analysis and planning.
|
|
326
|
-
|
|
327
|
-
**Tracker visibility during Opus dispatch**: on Claude Code the model switch to Opus happens via subagent dispatch, and the parent widget cannot move while an Agent call is in flight. Dispatch per task from the self-generated task list (never one monolithic call for the whole phase), set the pre-dispatch `activeForm` marker, and record tokens between chunks - full rules in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "Delegated phases".
|
|
328
|
-
|
|
329
|
-
---
|
|
330
|
-
|
|
331
300
|
#### v2.1.0+ Multi-Repo Mode
|
|
332
301
|
|
|
333
302
|
Active when `state.projects[].length > 1` (set by Phase 0 multi-select). Single-repo flow above is preserved verbatim - this section adds the deltas.
|
|
@@ -372,26 +341,26 @@ This closes the gap where an agent records "built" without ever producing build
|
|
|
372
341
|
|
|
373
342
|
**Telemetry**: Per-repo metrics in addition to per-task metrics:
|
|
374
343
|
```bash
|
|
375
|
-
$HOME/.claude/scripts/log-metric.sh "$TASK_ID"
|
|
376
|
-
$HOME/.claude/scripts/log-metric.sh "$TASK_ID"
|
|
344
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 2 build.completed repo=common duration_ms=$D status=ok
|
|
345
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 2 build.completed repo=uicomponents duration_ms=$D status=ok
|
|
377
346
|
```
|
|
378
347
|
|
|
379
348
|
**Token forwarding:** every TDD round (red, green, refactor) that hits the dev model MUST forward token totals into the tracker so Phase 5's Cost Breakdown captures Phase 2:
|
|
380
349
|
|
|
381
350
|
```bash
|
|
382
|
-
LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID"
|
|
351
|
+
LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 2 dev.tdd_round \
|
|
383
352
|
model=<sonnet|opus> step=<red|green|refactor> \
|
|
384
353
|
tokens_in=$IN tokens_out=$OUT duration_ms=$DUR
|
|
385
354
|
```
|
|
386
355
|
|
|
387
|
-
Model
|
|
356
|
+
Model is Sonnet unless routing names another rung (`multi-agent-refs/features/model-fallback.md`). Best-effort. See `$HOME/.claude/multi-agent-refs/progress-contract.md#token-telemetry-forwarding`.
|
|
388
357
|
|
|
389
358
|
---
|
|
390
359
|
|
|
391
360
|
## Token telemetry - invoke after every LLM call
|
|
392
361
|
|
|
393
362
|
```bash
|
|
394
|
-
bash $HOME/.claude/scripts/phase-tracker.sh tokens
|
|
363
|
+
bash $HOME/.claude/scripts/phase-tracker.sh tokens 2 <input_count> <output_count>
|
|
395
364
|
```
|
|
396
365
|
|
|
397
366
|
Contract and rationale: `progress-contract.md` -> Token telemetry forwarding.
|