@mmerterden/multi-agent-pipeline 19.1.3 → 20.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +150 -7
- package/README.md +60 -50
- package/README.tr.md +55 -46
- package/docs/adr/0002-instruction-driven-flag.md +6 -5
- package/docs/adr/0005-lazy-phase-docs.md +2 -2
- package/docs/adr/0008-installer-modularization-and-secret-leak-defense.md +1 -0
- package/docs/adr/0009-claude-stack-skills-plugin-only.md +1 -1
- package/docs/adr/0010-own-code-graph.md +5 -4
- package/docs/adr/0012-macos-only.md +2 -2
- package/docs/adr/0013-lsp-code-intelligence.md +2 -2
- package/docs/adr/0014-six-phase-consolidation.md +9 -9
- package/docs/adr/0015-one-pipeline-no-depth-answer.md +83 -0
- package/docs/adr/0016-the-run-shape-is-asked-not-typed.md +69 -0
- package/docs/adr/README.md +18 -16
- package/docs/architecture.md +2 -2
- package/docs/ecosystem.md +8 -9
- package/docs/facts.json +5 -8
- package/docs/features.md +4 -5
- package/docs/token-budget-history.md +1 -1
- package/install/_common.mjs +14 -6
- package/install/_mcp-register.mjs +1 -1
- package/install/_plugin-skills.mjs +3 -4
- package/install/copilot.mjs +5 -5
- package/install/templates/copilot-instructions.md +7 -16
- package/manifest.json +135 -138
- package/package.json +1 -1
- package/pipeline/commands/multi-agent/SKILL.md +6 -8
- package/pipeline/commands/multi-agent/analysis/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/analysis-jira/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/analysis-resolve/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/autopilot/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/autopilot-on/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/autopilot-status/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/build-optimize/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/channels/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/create-jira/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/design-check/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/forget/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +4 -2
- package/pipeline/commands/multi-agent/help/SKILL.md +21 -27
- package/pipeline/commands/multi-agent/ios-coding-standard/SKILL.md +5 -4
- package/pipeline/commands/multi-agent/issue/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/jira/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/language/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/prune-logs/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/purge/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/resume/SKILL.md +177 -48
- package/pipeline/commands/multi-agent/save/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/setup/SKILL.md +4 -4
- package/pipeline/commands/multi-agent/stack/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/sync/SKILL.md +6 -7
- package/pipeline/commands/multi-agent/test-screenshots/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/uninstall/SKILL.md +2 -0
- package/pipeline/commands/sim-test.md +4 -4
- package/pipeline/lib/repo-hygiene.sh +1 -1
- package/pipeline/multi-agent-refs/analysis/locked.md +2 -2
- package/pipeline/multi-agent-refs/analysis/render.md +1 -1
- package/pipeline/multi-agent-refs/analysis/resolve.md +1 -1
- package/pipeline/multi-agent-refs/analysis/synthesis.md +1 -1
- package/pipeline/multi-agent-refs/analysis-template.md +1 -1
- package/pipeline/multi-agent-refs/channels/jira.md +8 -8
- package/pipeline/multi-agent-refs/component-dispatch.md +0 -8
- package/pipeline/multi-agent-refs/cross-cli-contract.md +10 -11
- package/pipeline/multi-agent-refs/features/base-branch-evidence.md +2 -2
- package/pipeline/multi-agent-refs/features/external-context-injection.md +2 -0
- package/pipeline/multi-agent-refs/features/review-delta.md +1 -1
- package/pipeline/multi-agent-refs/features/review-multi-repo.md +3 -3
- package/pipeline/multi-agent-refs/features/scope-check.md +1 -1
- package/pipeline/multi-agent-refs/features/skill-conformance.md +1 -1
- package/pipeline/multi-agent-refs/features/stack-skill-routing.md +1 -1
- package/pipeline/multi-agent-refs/features/visual-evidence.md +2 -1
- package/pipeline/multi-agent-refs/features/worktree-finalize.md +1 -1
- package/pipeline/multi-agent-refs/generate-issue.md +2 -0
- package/pipeline/multi-agent-refs/issue-jira-triad.md +2 -0
- package/pipeline/multi-agent-refs/keychain.md +2 -0
- package/pipeline/multi-agent-refs/knowledge.md +0 -7
- package/pipeline/multi-agent-refs/outside-the-pipeline.md +1 -1
- package/pipeline/multi-agent-refs/payload-contracts.md +1 -1
- package/pipeline/multi-agent-refs/phases/modes.md +32 -108
- package/pipeline/multi-agent-refs/phases/operations.md +3 -1
- package/pipeline/multi-agent-refs/phases/phase-0-init.md +23 -42
- package/pipeline/multi-agent-refs/phases/phase-1-plan.md +10 -21
- package/pipeline/multi-agent-refs/phases/phase-2-dev.md +13 -44
- package/pipeline/multi-agent-refs/phases/phase-3-review.md +19 -22
- package/pipeline/multi-agent-refs/phases/phase-4-commit.md +6 -6
- package/pipeline/multi-agent-refs/phases/phase-5-report.md +2 -2
- package/pipeline/multi-agent-refs/phases.md +9 -11
- package/pipeline/multi-agent-refs/progress-contract.md +1 -1
- package/pipeline/multi-agent-refs/readiness-review.md +2 -0
- package/pipeline/multi-agent-refs/rules.md +2 -2
- package/pipeline/multi-agent-refs/tracker-contract.md +9 -40
- package/pipeline/multi-agent-refs/wiki-capture.md +3 -2
- package/pipeline/preferences-template.json +2 -2
- package/pipeline/rules/figma-pipeline.md +1 -1
- package/pipeline/schemas/agent-state.schema.json +5 -10
- package/pipeline/schemas/migrations/prefs-2.7.0-to-2.8.0.mjs +33 -0
- package/pipeline/schemas/phases.json +3 -24
- package/pipeline/schemas/prefs.schema.json +5 -5
- package/pipeline/scripts/autopilot-runner.mjs +6 -7
- package/pipeline/scripts/build-references.mjs +3 -3
- package/pipeline/scripts/build-stack-plugins.mjs +1 -1
- package/pipeline/scripts/bulk-read.sh +6 -4
- package/pipeline/scripts/cost-table.json +1 -1
- package/pipeline/scripts/doctor.mjs +4 -4
- package/pipeline/scripts/gc-refs.sh +1 -1
- package/pipeline/scripts/gen-mode-dispatch.mjs +11 -41
- package/pipeline/scripts/learnings-ledger.mjs +1 -1
- package/pipeline/scripts/match-skills.mjs +4 -4
- package/pipeline/scripts/memory-load.sh +3 -3
- package/pipeline/scripts/migrate-prefs.mjs +18 -17
- package/pipeline/scripts/phase-tracker.sh +2 -2
- package/pipeline/scripts/phase0-exit-gate.mjs +1 -1
- package/pipeline/scripts/plan-coverage-gate.mjs +6 -6
- package/pipeline/scripts/run-aggregator.mjs +3 -3
- package/pipeline/scripts/runs-index.mjs +7 -7
- package/pipeline/scripts/scope-check-gate.mjs +1 -1
- package/pipeline/scripts/smoke-schema-validation.sh +9 -12
- package/pipeline/scripts/usage-report.mjs +5 -7
- package/pipeline/scripts/validate-analysis-doc.mjs +3 -3
- package/pipeline/scripts/worktree-finalize.sh +2 -2
- package/pipeline/scripts/write-state.mjs +22 -11
- package/pipeline/skills/.skill-manifest.json +9 -21
- package/pipeline/skills/.skills-index.json +6 -39
- package/pipeline/skills/shared/README.md +5 -8
- package/pipeline/skills/shared/core/multi-agent/SKILL.md +10 -13
- package/pipeline/skills/shared/core/multi-agent-autopilot-status/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-channels/SKILL.md +4 -5
- package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +13 -16
- package/pipeline/skills/shared/core/multi-agent-ios-coding-standard/SKILL.md +2 -3
- package/pipeline/skills/shared/core/multi-agent-resume/SKILL.md +51 -15
- package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +6 -6
- package/pipeline/skills/skills-index.md +3 -6
- package/pipeline/commands/multi-agent/local/SKILL.md +0 -132
- package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +0 -142
- package/pipeline/commands/multi-agent/resume-local/SKILL.md +0 -114
- package/pipeline/skills/shared/core/multi-agent-local/SKILL.md +0 -41
- package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +0 -55
- package/pipeline/skills/shared/core/multi-agent-resume-local/SKILL.md +0 -51
|
@@ -7,6 +7,8 @@
|
|
|
7
7
|
- [Log line shape](#log-line-shape)
|
|
8
8
|
<!-- /toc -->
|
|
9
9
|
|
|
10
|
+
> **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
|
|
11
|
+
|
|
10
12
|
Phase 0 Step 1b catalogued every typed external link from the task description into `state.contextLinks[]`. Phase 1 dispatches each entry to its matching fetcher and prepends the result to the analysis prompt under a **Referenced External Sources** section, so the agent doesn't re-discover what the ticket already pointed at.
|
|
11
13
|
|
|
12
14
|
```bash
|
|
@@ -63,7 +63,7 @@ DELTA_JSON=$(node $HOME/.claude/scripts/review-delta.mjs --rounds-dir "$WORKTREE
|
|
|
63
63
|
jq -c --argjson d "$DELTA_JSON" --argjson i "$((ITERATION-1))" \
|
|
64
64
|
'{reviewIterations: (.reviewIterations | .[$i] += {delta: ($d + {computedAt: (now | todate)})})}' "$STATE_FILE" \
|
|
65
65
|
| node $HOME/.claude/scripts/write-state.mjs "$STATE_FILE"
|
|
66
|
-
$HOME/.claude/scripts/log-metric.sh "$TASK_ID"
|
|
66
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 review.delta iteration=$ITERATION \
|
|
67
67
|
new=$(jq '.counts.new // 0' <<< "$DELTA_JSON") still_present=$(jq '.counts.stillPresent // 0' <<< "$DELTA_JSON") \
|
|
68
68
|
resolved=$(jq '.counts.resolved // 0' <<< "$DELTA_JSON") plateau=$(jq '.plateau // false' <<< "$DELTA_JSON") tripped=$([ "$DELTA_RC" -eq 3 ] && echo true || echo false)
|
|
69
69
|
```
|
|
@@ -57,9 +57,9 @@ And log `review.diff_truncated repo=<name> bytes_dropped=<N>`. Triage receives t
|
|
|
57
57
|
|
|
58
58
|
**Telemetry**: Per-repo build/test gate timings + a single combined review/triage call set:
|
|
59
59
|
```bash
|
|
60
|
-
$HOME/.claude/scripts/log-metric.sh "$TASK_ID"
|
|
61
|
-
$HOME/.claude/scripts/log-metric.sh "$TASK_ID"
|
|
62
|
-
$HOME/.claude/scripts/log-metric.sh "$TASK_ID"
|
|
60
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 gate.build repo=common status=pass duration_ms=$D
|
|
61
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 gate.build repo=uicomponents status=pass duration_ms=$D
|
|
62
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 review.combined_diff repos=2 bytes=$BYTES truncated=false
|
|
63
63
|
```
|
|
64
64
|
|
|
65
65
|
---
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Feature: Scope self-check (Phase 2 Step 3.7)
|
|
2
2
|
|
|
3
|
-
**Pattern**: Phase 3 reconstructs everything from the diff. The one thing it cannot reconstruct is why each file was touched and what was left out on purpose, so Dev states both before the handoff, and a deterministic gate checks the statement against the real diff. The record also carries the code-simplifier rationales (Step 3.6) that
|
|
3
|
+
**Pattern**: Phase 3 reconstructs everything from the diff. The one thing it cannot reconstruct is why each file was touched and what was left out on purpose, so Dev states both before the handoff, and a deterministic gate checks the statement against the real diff. The record also carries the code-simplifier rationales (Step 3.6) that would otherwise be discarded, and it feeds two later consumers: the `<scope-self-check>` block in the Phase 3 reviewer prefix and the PR body in Phase 4.
|
|
4
4
|
|
|
5
5
|
## The record
|
|
6
6
|
|
|
@@ -22,7 +22,7 @@ Phase 4 used to select its review criteria from `detectedStack`, a string Phase
|
|
|
22
22
|
- The dev side and the review side could disagree about the standard without either noticing. A component built by a plugin skill was reviewed against `clean-code`.
|
|
23
23
|
- "Did it do this correctly?" had no fixed denominator, so the only available answer was "it looks fine", and a reviewer that opened nothing produced the same output as a reviewer that checked everything.
|
|
24
24
|
|
|
25
|
-
|
|
25
|
+
The sharper case is a document whose evidence supported few sections: the sole input to the old selection logic was then nearly empty.
|
|
26
26
|
|
|
27
27
|
## The four rules that make it work
|
|
28
28
|
|
|
@@ -41,7 +41,7 @@ Nothing enabled is not an error: a repo whose stack was never selected legitimat
|
|
|
41
41
|
no toolkit. Record the no-op with the enabled set that was read, so "none applied" is
|
|
42
42
|
distinguishable from "never looked".
|
|
43
43
|
|
|
44
|
-
**Not enabled is not an error here**, unlike component dispatch: a repo whose stack was never selected legitimately has no toolkit, and halting would make the pipeline unusable there. Record the no-op and continue.
|
|
44
|
+
**Not enabled is not an error here**, unlike component dispatch: a repo whose stack was never selected legitimately has no toolkit, and halting would make the pipeline unusable there. Record the no-op and continue. Every stack now has a toolkit, so "no toolkit" means the stack was never selected, not that none exists for it.
|
|
45
45
|
|
|
46
46
|
Two marketplaces may ship the same toolkit name (a public one and a corporate one). Resolve whichever is enabled and record its **name and version** in the ledger entry, because the routing table and the skill set differ between versions - a finding that cites a skill has to be traceable to the version that defined it.
|
|
47
47
|
|
|
@@ -123,7 +123,8 @@ section still renders, saying so.
|
|
|
123
123
|
## 3. After - Phase 2, not the Phase 3 user test
|
|
124
124
|
|
|
125
125
|
The user test is the natural home: the simulator is already up. It is also **dropped by
|
|
126
|
-
every `autopilot`
|
|
126
|
+
every `autopilot` entry and by any run whose
|
|
127
|
+
workspace is local** (`phase-3-review.md` TLDR), so a capture
|
|
127
128
|
that lives only there produces nothing for unattended runs - which are exactly
|
|
128
129
|
the runs where nobody watched the screen.
|
|
129
130
|
|
|
@@ -20,7 +20,7 @@ Removing it at PR-open is only safe because of the salvage, so the two are one s
|
|
|
20
20
|
|
|
21
21
|
| Condition | Why it blocks |
|
|
22
22
|
|---|---|
|
|
23
|
-
| `worktreePath == projectRoot` (
|
|
23
|
+
| `worktreePath == projectRoot` (local workspace) | there is no worktree; removing it would delete the user's checkout |
|
|
24
24
|
| cwd is inside the worktree | a shell left on a deleted inode is worse than a leftover directory, and Phase 4 legitimately `cd`s into the worktree earlier |
|
|
25
25
|
| not a registered worktree of the project root | a mistyped path must not delete an unrelated directory |
|
|
26
26
|
| real uncommitted changes | never discarded; see the artefact carve-out below |
|
|
@@ -9,6 +9,8 @@
|
|
|
9
9
|
|
|
10
10
|
> **TLDR** - Shared 12-step flow for `/multi-agent:create-jira`. Asks the issue type (Task / Bug / Story), mines the target project's existing same-type issues to learn team conventions, detects the active sprint, drafts a standards-compliant issue from a fixed standard template with auto-sizing sections, asks the user about every genuinely unknown field, renders a full preview, and creates the Jira issue only after explicit approval. Creates exactly one Jira issue per run - no branches, no commits, no worktrees.
|
|
11
11
|
|
|
12
|
+
> **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
|
|
13
|
+
|
|
12
14
|
Consumed by `create-jira/SKILL.md`. This ref is never invoked directly.
|
|
13
15
|
|
|
14
16
|
## Hard rules (must not regress)
|
|
@@ -9,6 +9,8 @@
|
|
|
9
9
|
- [Cross-CLI parity](#cross-cli-parity)
|
|
10
10
|
<!-- /toc -->
|
|
11
11
|
|
|
12
|
+
> **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
|
|
13
|
+
|
|
12
14
|
> **TLDR** - When a GitHub issue triggers the pipeline and has no Jira ID, the `autoJiraFromGithubIssue` policy decides whether to auto-create a Jira task (and patch the GitHub issue body with the new Jira link). Phase 5 then posts a humanizer'd wiki-content summary back as a Jira comment, closing the loop. Autopilot treats `ask` as `always`.
|
|
13
15
|
|
|
14
16
|
This doc is referenced from `$HOME/.claude/multi-agent-refs/phases/phase-0-init.md` Step 1 (GitHub issue input) and `$HOME/.claude/multi-agent-refs/phases/phase-5-report.md` Step 2 (component wiki). Keeps the phase docs tight and gives the triad contract a stable home.
|
|
@@ -1,5 +1,7 @@
|
|
|
1
1
|
## Keychain Token Registry
|
|
2
2
|
|
|
3
|
+
> **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
|
|
4
|
+
|
|
3
5
|
Tokens live in the platform-native credential store (macOS Keychain / Linux libsecret / Windows Credential Manager). Key names are resolved via **preferences mapping** - never hardcoded.
|
|
4
6
|
|
|
5
7
|
### Rule 1 - Never prompt for a token *value* mid-run
|
|
@@ -49,13 +49,6 @@ Each task teaches the project a bit more. Token cost decreases over time.
|
|
|
49
49
|
|
|
50
50
|
They all complement each other - no conflicts.
|
|
51
51
|
|
|
52
|
-
### Knowledge in a Short run
|
|
53
|
-
|
|
54
|
-
A Short run skips Phase 1, but knowledge **is still read**:
|
|
55
|
-
|
|
56
|
-
- Knowledge files are added to the prompt before the Opus agent starts in Phase 3
|
|
57
|
-
- Knowledge capture is still performed in Phase 5 (simplified report)
|
|
58
|
-
|
|
59
52
|
### Knowledge Maintenance
|
|
60
53
|
|
|
61
54
|
Knowledge files grow over time. Maintenance rules:
|
|
@@ -47,7 +47,7 @@ and a review between fetch and action; an ordinary session has neither.
|
|
|
47
47
|
| Read an issue, page, log, crash, scan result | Resolve the credential and fetch |
|
|
48
48
|
| Comment on Jira, edit an issue, move a board column | `/multi-agent:channels` |
|
|
49
49
|
| Create an issue | `/multi-agent:create-jira` |
|
|
50
|
-
| Open or update a PR | a pipeline run, or `/multi-agent:resume
|
|
50
|
+
| Open or update a PR | a pipeline run, or `/multi-agent:resume` |
|
|
51
51
|
|
|
52
52
|
The split is not bureaucracy. Outward writes carry rules that live in those commands:
|
|
53
53
|
issues are never auto-closed (four approvals), PR bodies use `Ref:` and never
|
|
@@ -64,4 +64,4 @@ Non-empty `UNTRACKED` → both the agent-log report and the closing chat summary
|
|
|
64
64
|
|
|
65
65
|
## Fast modes are not exempt
|
|
66
66
|
|
|
67
|
-
|
|
67
|
+
The autopilot and local entries skip the interactive test gate. They run Phase 4 and Phase 5 **unchanged**. A short pipeline is not a licence for an improvised payload shape, a missing Test Scenarios section, or a report without numbers.
|
|
@@ -1,18 +1,16 @@
|
|
|
1
|
-
> **TLDR** - four entry commands,
|
|
1
|
+
> **TLDR** - four entry commands, one pipeline, two axes:
|
|
2
2
|
>
|
|
3
|
-
> -
|
|
4
|
-
> -
|
|
5
|
-
> - **`--local`** means no worktree: work happens in `$PROJECT_ROOT` on a local branch.
|
|
3
|
+
> - **`autopilot`** skips confirmations (Plan, Test, Commit, PR prompts). Still fails safe on review blockers and build retries.
|
|
4
|
+
> - **Workspace** is the one question the run asks about its own shape: a worktree, or the project root on a local branch. Phase 0 Step 5b asks it; autopilot resolves it to a worktree and never asks.
|
|
6
5
|
>
|
|
7
|
-
>
|
|
6
|
+
> There is one pipeline and every mode runs its whole phase set. What a run costs follows the evidence the task carries, not an answer taken before the evidence exists.
|
|
8
7
|
|
|
9
8
|
## Autopilot Mode
|
|
10
9
|
|
|
11
10
|
<!-- toc -->
|
|
12
11
|
- [Autopilot Mode](#autopilot-mode)
|
|
13
|
-
- [Pipeline depth (Full / Short)](#pipeline-depth-full-short)
|
|
14
12
|
- [Analysis Mode (`/multi-agent:analysis`)](#analysis-mode-multi-agentanalysis)
|
|
15
|
-
- [Local Mode
|
|
13
|
+
- [Local Mode](#local-mode)
|
|
16
14
|
<!-- /toc -->
|
|
17
15
|
|
|
18
16
|
Autopilot mode skips interactive confirmations and runs the pipeline end-to-end autonomously.
|
|
@@ -27,15 +25,19 @@ Autopilot mode skips interactive confirmations and runs the pipeline end-to-end
|
|
|
27
25
|
/multi-agent "LoginView dark mode fix" autopilot
|
|
28
26
|
```
|
|
29
27
|
|
|
30
|
-
**What changes in autopilot (Tablo 1
|
|
28
|
+
**What changes in autopilot (Tablo 1)**
|
|
31
29
|
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
| Phase
|
|
37
|
-
|
|
|
38
|
-
|
|
|
30
|
+
Two axes, and only the first is a command: whether anything is confirmed
|
|
31
|
+
(`autopilot`), and where the work happens (the Step 5b answer). An autopilot run
|
|
32
|
+
always gets a worktree, so there is no unattended-local column.
|
|
33
|
+
|
|
34
|
+
| Phase | Interactive, worktree | Interactive, local | autopilot (always worktree) |
|
|
35
|
+
| --- | --- | --- | --- |
|
|
36
|
+
| Phase 1 (Plan Approval Gate) | Clarification (max 2 rounds) + approval loop - user: approve/abort/free-text edit | Same as worktree | **Gate skip** - log the plan, proceed directly to Phase 2 (autopilot contract: zero interaction) |
|
|
37
|
+
| Phase 3 (User Test step) | Interactive prompt ("Want to test?" -> wait) | **Not offered** - the step checks the change out of a worktree and there is none; the phase itself still runs | Skip (autopilot suppresses interactive prompts) -> proceed directly to Phase 4 |
|
|
38
|
+
| Phase 4 (Commit) | "Want to commit?" -> wait | "Want to commit?" -> wait | Auto commit + push |
|
|
39
|
+
| Phase 4 (PR) | "Want to open a PR?" -> wait | "Want to open a PR?" -> wait | Auto create PR |
|
|
40
|
+
| **Phase 5 (Channels)** | Multi-select channel + content menu | Multi-select (same) | **STILL PAUSES** - see "Phase 5 autopilot exception" |
|
|
39
41
|
|
|
40
42
|
**What NEVER skips (even in autopilot):**
|
|
41
43
|
|
|
@@ -51,7 +53,7 @@ Autopilot mode skips interactive confirmations and runs the pipeline end-to-end
|
|
|
51
53
|
|
|
52
54
|
The generic "zero-interaction" contract covers Phases 0-4 only. Phase 5 channels dispatch is the **single exception**:
|
|
53
55
|
|
|
54
|
-
- ALL modes
|
|
56
|
+
- ALL modes, `autopilot` included, pause at the channels multi-select menu.
|
|
55
57
|
- Menu pre-ticks from `prefs.global.reportChannels` + `prefs.global.reportContent` - user can accept with one keypress if prefs are stable.
|
|
56
58
|
- **30-minute timeout** - if user does not respond, session ends cleanly:
|
|
57
59
|
- External delivery aborted (no silent apply - prevents accidental Jira comments / Confluence pages).
|
|
@@ -62,79 +64,9 @@ The generic "zero-interaction" contract covers Phases 0-4 only. Phase 5 channels
|
|
|
62
64
|
|
|
63
65
|
Full contract: `$HOME/.claude/multi-agent-refs/phases/phase-5-report.md` (Autopilot pause contract) + `commands/multi-agent/channels/SKILL.md`.
|
|
64
66
|
|
|
65
|
-
---
|
|
66
|
-
|
|
67
|
-
## Pipeline depth (Full / Short)
|
|
68
|
-
|
|
69
|
-
Short strips the pipeline to what an already-scoped task needs: no deep analysis, no planning phase. The **Opus** dev agent reads the task scope and implements, and Phase 3 then reviews what it produced. Review is deliberately NOT part of the strip: analysis and planning shape work that has not happened yet, so a task the user has already scoped can skip them, while review judges work that now exists and has no substitute.
|
|
70
|
-
|
|
71
|
-
**How it is chosen.** Phase 0 Step 7.5, after `taskType` is known:
|
|
72
|
-
|
|
73
|
-
```
|
|
74
|
-
Bu is icin hangi pipeline? / Which pipeline for this task?
|
|
75
|
-
1. Tam / Full Analysis -> Plan -> Dev -> Review -> Test -> Commit -> Report
|
|
76
|
-
2. Kisa / Short Dev (self-contained, Opus) -> Review -> Test -> Commit -> Report
|
|
77
|
-
```
|
|
78
|
-
|
|
79
|
-
`bugfix` and `chore` recommend Short; `feature`, `refactor` and `component` recommend Full. The recommendation is presented first and passed as `ASK_CHOICE_DEFAULT` on hosts without a native picker, because `ask-choice.sh` picks the first option on a non-TTY and option order is not a contract. Pass it as the **1-based index** (Full = 1, Short = 2), not as the label: labels render in `outputLanguage`, so a label-valued default matches nothing on a `tr` run.
|
|
80
|
-
|
|
81
|
-
**Who is asked.** `/multi-agent` and `/multi-agent:local`. Both autopilot entries always run Full without asking.
|
|
82
|
-
|
|
83
|
-
**Why there is no fast-and-unattended combination.** It existed until v16.0.0, and removing it was a real behaviour change, not a rename. Autopilot may not ask, so something has to choose, and unattended is the worst place to drop analysis and planning: nobody is watching to notice what the shortcut lost. A cron job or script that wants both now has to pick - stay unattended and pay for the full pipeline, or stay fast and have a person present.
|
|
84
|
-
|
|
85
|
-
**Pipeline in a Short run:**
|
|
86
|
-
|
|
87
|
-
```
|
|
88
|
-
Phase 0: Init -> Phase 2: Dev (self-contained) -> Phase 3: Review -> Phase 3: Review (user test) -> Phase 4: Commit -> Phase 5: Report
|
|
89
|
-
```
|
|
90
|
-
|
|
91
|
-
**What changes (Tablo 2 - Short runs):**
|
|
92
|
-
|
|
93
|
-
| Phase | Full | Short | Short + `--local` |
|
|
94
|
-
| ------------------- | ----------------------------------------------- | ---------------------------------------------------------------------------------- | --------------------------------------------- |
|
|
95
|
-
| Phase 0 (Init) | Full setup | Same - worktree, branch, state, and the depth question itself | Same - no worktree, branch on `$PROJECT_ROOT` |
|
|
96
|
-
| Phase 1 (Analysis) | Parallel Explore agents + analysis document | **SKIP** - no tile is ever drawn for it (registration is deferred to Step 7.5) | **SKIP** |
|
|
97
|
-
| Phase 1 (Planning) | TaskCreate + architecture review + **Plan Approval Gate** | **SKIP** (no plan means no plan gate) | **SKIP** |
|
|
98
|
-
| Phase 2 (Dev) | Follows the Phase 1 plan, TDD cycle (Sonnet) | **Self-contained** (Opus): agent scans relevant files, implements with TDD, builds | Same, on the local branch |
|
|
99
|
-
| Phase 3 (Review) | Parallel review + Fable triage (3 reviewers on every host: Claude Code Fable + Opus + Sonnet, Copilot GPT-5.4 + Opus + Sonnet) | **Same** - gates, parallel review, triage; blocking findings return to Phase 2 (cap 3) | **Same**, on the local branch diff |
|
|
100
|
-
| Phase 3 (User Test) | Interactive prompt | **Interactive prompt** | Not in the set - no worktree to check out |
|
|
101
|
-
| Phase 4 (Commit) | Commit + PR | Same - still asks | Same |
|
|
102
|
-
| Phase 5 (Report) | Full report + channels multi-select | Simplified - no analysis section, review section IS present, channels menu still pauses | Same |
|
|
103
|
-
|
|
104
|
-
**Phase 2 in a Short run (self-contained):**
|
|
105
|
-
|
|
106
|
-
The **Opus** agent receives the task description (from Jira, GitHub issue, or free-text) and:
|
|
107
|
-
|
|
108
|
-
1. Quickly scans relevant files in the codebase (lightweight, not full Explore)
|
|
109
|
-
2. Implements the change following TDD cycle (test -> code -> build)
|
|
110
|
-
3. Runs build verification (max 3 retries on failure)
|
|
111
|
-
|
|
112
|
-
No separate task breakdown - the agent handles scope autonomously.
|
|
113
|
-
|
|
114
|
-
**State tracking**: `agent-state.json` gets `"onlyDevelop": true`. The tracker boots at Step -1 with Phase 0 alone and the rest of the tiles are registered at Step 7.5, once this answer says which phases the run actually has - so a Short run never draws an Analysis tile it will not use. Full contract: `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "Deferred registration".
|
|
115
|
-
|
|
116
|
-
### Intake warnings for a Short run
|
|
117
|
-
|
|
118
|
-
**An analysis document was supplied.** Because Phase 1 and Phase 1 are skipped, no phase turns that document into a task breakdown. The doc becomes raw context for one Dev pass, and work comes out ordered by whatever the model read first: the bottom of the dependency chain lands, the screen wiring does not.
|
|
119
|
-
|
|
120
|
-
The depth picker is where this is caught. When the intake carried an analysis document or a Figma reference, say so in the question itself rather than after the choice:
|
|
121
|
-
|
|
122
|
-
```
|
|
123
|
-
This task carries an analysis document. Short skips Analysis and Planning, so
|
|
124
|
-
the document will not be turned into a task breakdown.
|
|
125
|
-
1. Tam / Full (recommended here)
|
|
126
|
-
2. Kisa / Short (doc as context only)
|
|
127
|
-
```
|
|
128
|
-
|
|
129
|
-
Autopilot never sees this: it runs Full.
|
|
130
|
-
|
|
131
|
-
**The branch already carries the work.** When the change was developed outside the pipeline, or by hand, Short is still the wrong entry point: it will try to develop again. `/multi-agent:resume-local` puts the existing diff through the same review, adds a build+test success gate, opens the PR, and posts the Jira technical-analysis + test-scenario comment - without re-developing. Offer it when the working tree or branch is already ahead of the base with the task's changes.
|
|
132
|
-
|
|
133
|
-
---
|
|
134
|
-
|
|
135
67
|
## Analysis Mode (`/multi-agent:analysis`)
|
|
136
68
|
|
|
137
|
-
|
|
69
|
+
A different **output kind**, not a shorter pipeline. It produces the analysis document and stops; no code, no branch, no PR. The 6-phase contract holds, with four phases reinterpreted the same way a Short run reinterprets Phase 2.
|
|
138
70
|
|
|
139
71
|
```
|
|
140
72
|
Phase 0: Init -> Phase 1: Plan (analysis) -> Phase 1: Plan -> Phase 3: Review -> Phase 4: Publish -> Phase 5: Report
|
|
@@ -151,18 +83,19 @@ Phase 0: Init -> Phase 1: Plan (analysis) -> Phase 1: Plan -> Phase 3: Review ->
|
|
|
151
83
|
| 6 Commit | **Publish** instead: Local file / Confluence / Jira. Locked 6 still forbids `git add` and `git commit` |
|
|
152
84
|
| 7 Report | Channels, same as every mode |
|
|
153
85
|
|
|
154
|
-
**No `
|
|
86
|
+
**No `autopilot` variant.** Worktree isolation buys nothing when no code is written, and the intake, the Pass B convention preview and the open-question resolution are interactive by nature; a zero-interaction analysis would be a document nobody agreed to.
|
|
155
87
|
|
|
156
88
|
**Where the document lands.** Drafts go to `/tmp/analysis-<slug>-<ts>/` first, then the Phase 4 picker decides: Local writes `<repo>/analysis/<feature>-<platform>.md` uncommitted, Confluence posts the page, Jira updates the description. In the full pipeline the same engine writes into the worktree instead and `prefs.global.analysisPhase.commitDoc` decides whether it rides along with the commit.
|
|
157
89
|
|
|
158
90
|
---
|
|
159
91
|
|
|
160
|
-
## Local Mode
|
|
92
|
+
## Local Mode
|
|
161
93
|
|
|
162
94
|
Local mode skips worktree creation - works directly on a local branch in the project root. Useful for single-task workflows or when worktrees cause issues.
|
|
163
95
|
|
|
164
|
-
**Activation**: answer the Phase 0 Step 5b workspace question
|
|
165
|
-
|
|
96
|
+
**Activation**: answer the Phase 0 Step 5b workspace question. There is no flag
|
|
97
|
+
and no command name for it - the answer is the only way in, which is why the
|
|
98
|
+
question carries what the choice costs.
|
|
166
99
|
|
|
167
100
|
```
|
|
168
101
|
Bu is nerede kossun? / Where should this task run?
|
|
@@ -171,25 +104,16 @@ Bu is nerede kossun? / Where should this task run?
|
|
|
171
104
|
```
|
|
172
105
|
|
|
173
106
|
Two genuine options, so it meets the two-option floor in `picker-contract.md` on
|
|
174
|
-
its own. **Say what local costs inside the question**:
|
|
175
|
-
|
|
176
|
-
|
|
177
|
-
user choosing local should learn both before
|
|
178
|
-
|
|
179
|
-
The ways to state it up front:
|
|
180
|
-
|
|
181
|
-
```
|
|
182
|
-
/multi-agent:local "PROJ-12345"
|
|
183
|
-
/multi-agent:local "#316"
|
|
184
|
-
/multi-agent:local "LoginView dark mode fix"
|
|
185
|
-
/multi-agent:local-autopilot "PROJ-12345"
|
|
186
|
-
/multi-agent "PROJ-12345" --local
|
|
187
|
-
```
|
|
107
|
+
its own. **Say what local costs inside the question**: the user-test STEP inside
|
|
108
|
+
Phase 3 checks the change out of a worktree, so a local run does not get that
|
|
109
|
+
offer (the phase itself still runs), and uncommitted work in the project root is
|
|
110
|
+
in the way of the checkout. A user choosing local should learn both before
|
|
111
|
+
choosing, not after.
|
|
188
112
|
|
|
189
113
|
Every autopilot entry resolves this to **worktree** and never asks. That is not a
|
|
190
114
|
skipped question: an unattended run commits and pushes from wherever it stands,
|
|
191
|
-
and doing that in the user's own checkout is what worktrees exist to prevent.
|
|
192
|
-
|
|
115
|
+
and doing that in the user's own checkout is what worktrees exist to prevent. An
|
|
116
|
+
unattended run therefore always gets its own checkout.
|
|
193
117
|
|
|
194
118
|
**What changes in local mode:**
|
|
195
119
|
|
|
@@ -221,4 +145,4 @@ git -C $PROJECT_ROOT config user.email "{identity.email}"
|
|
|
221
145
|
- Uncommitted changes in PROJECT_ROOT may conflict - pipeline warns if dirty
|
|
222
146
|
- Build queue lock still applies for xcodebuild
|
|
223
147
|
|
|
224
|
-
|
|
148
|
+
Fast path.
|
|
@@ -17,6 +17,8 @@
|
|
|
17
17
|
- [Phase Pipeline](#phase-pipeline)
|
|
18
18
|
<!-- /toc -->
|
|
19
19
|
|
|
20
|
+
> **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
|
|
21
|
+
|
|
20
22
|
Every task gets an auto-incremented short ID. Counter stored at `$HOME/.claude/logs/multi-agent/{project}/.counter` (persists across sessions).
|
|
21
23
|
|
|
22
24
|
```
|
|
@@ -176,7 +178,7 @@ This keeps orchestrator context lean and enables programmatic routing.
|
|
|
176
178
|
bash $HOME/.claude/scripts/capture-flush.sh --state "$STATE_FILE" --quiet
|
|
177
179
|
```
|
|
178
180
|
|
|
179
|
-
The durable stores
|
|
181
|
+
The durable stores are NOT written only in Phase 5, which is the phase a run is LEAST likely to reach: a run killed in Phase 2 threw away every finding it had established, and the next run on the same repo paid to rediscover it. The flush is idempotent (the second call through writes 0 rows), costs no API tokens, calls no model, and never fails the transition. Phase 5 is now the LAST flush rather than the only one. The trigger is deliberately mechanical - a phase transition, not a model noticing that a moment qualifies.
|
|
180
182
|
- *Compaction trigger.* If conversation context exceeds ~50%, run `/compact` preserving "modified files, plan, open review findings, current phase + sub-step" before continuing. Don't wait for auto-compaction near the limit - it triggers exactly when context is worst and is lossy. After compaction, re-read `agent-state.json` AND the latest `## Handoff` block in `agent-log.md` to re-ground.
|
|
181
183
|
|
|
182
184
|
**Handoff block (v10.8.0)**: the structured artifact the phase-boundary checkpoint appends to `agent-log.md`. Written by the orchestrator from state it already holds - no agent dispatch, no extra LLM call. Cap at ~15 lines; the latest block is authoritative (earlier ones are history). This is the fresh-context re-entry contract: a resume or post-compaction session rebuilds working context from the latest handoff + `agent-state.json` + git log, never from conversation memory.
|
|
@@ -14,23 +14,21 @@ $HOME/.claude/scripts/phase-tracker.sh tiles
|
|
|
14
14
|
$HOME/.claude/scripts/phase-tracker.sh update 0 in_progress
|
|
15
15
|
```
|
|
16
16
|
|
|
17
|
-
**
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
whole set here, in phase-number order. Contract: `tracker-contract.md`,
|
|
21
|
-
"Deferred registration".
|
|
17
|
+
**Every mode registers its whole set here**, in phase-number order. The phase
|
|
18
|
+
set is known at Phase 0 because there is one pipeline: no answer later in the
|
|
19
|
+
run can add or remove a phase. Contract: `tracker-contract.md`.
|
|
22
20
|
|
|
23
21
|
`tiles` prints this host's widget-registration calls: **make them before continuing.** The card alone lands in collapsed tool output, so a run that skips them runs in silence. Contract: `tracker-contract.md`, "The card is not the widget".
|
|
24
22
|
|
|
25
23
|
If `INPUT_TASK_ID` isn't known yet (free-text, project not selected), use a placeholder; rename later via `mv` once parsed in Step 1.
|
|
26
24
|
|
|
27
|
-
Every subsequent phase (1-
|
|
25
|
+
Every subsequent phase (1-5) MUST call `phase-tracker.sh update <N> in_progress` on entry and `phase-tracker.sh update <N> completed|failed|skipped` on exit. Each `update` prints a `-- NEXT (required) --` block: act on it. Sub-phase milestones use `phase-tracker.sh sub <N> <subN> "<name>" <status>`. See `$HOME/.claude/multi-agent-refs/phases.md` "Visual Phase Tracker" for the full contract.
|
|
28
26
|
|
|
29
|
-
`update <N> completed` **exits 3** for phases 1-
|
|
27
|
+
`update <N> completed` **exits 3** for phases 1-5 with no recorded spend: record `model` + `tokens`, or pass `--no-llm`, then re-run it. Contract: `tracker-contract.md`, "Accounting is a gate".
|
|
30
28
|
|
|
31
29
|
##### TaskCreate ordering on Claude Code (strict)
|
|
32
30
|
|
|
33
|
-
On Claude Code, fire every `TaskCreate` in a registration batch in strict phase-number order BEFORE any `TaskUpdate` in that batch - which is the order `tiles` prints them in - and never register a phase whose number is below one already registered.
|
|
31
|
+
On Claude Code, fire every `TaskCreate` in a registration batch in strict phase-number order BEFORE any `TaskUpdate` in that batch - which is the order `tiles` prints them in - and never register a phase whose number is below one already registered. Full contract: `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
|
|
34
32
|
|
|
35
33
|
---
|
|
36
34
|
|
|
@@ -46,6 +44,13 @@ OUTPUT_LANG=$(jq -r '.global.outputLanguage // "en"' "$PREFS_FILE" 2>/dev/null |
|
|
|
46
44
|
|
|
47
45
|
From this point on, everything the user reads renders in `$OUTPUT_LANG`: conversational lines, `AskUserQuestion` `question`/`label`/`description`, and external payload bodies (PR/Jira/Confluence). English stays only on `header`, commit messages, branch names, PR title prefixes, identifiers. Full matrix: `rules.md` "Language Application".
|
|
48
46
|
|
|
47
|
+
**Every picker in this phase prints its breadcrumb.** Phase 0 is one chain -
|
|
48
|
+
account, repos, base branch, maturity, dev-context, workspace - and each step
|
|
49
|
+
emits the narrator line `<localized: "Step i/n: what this step decides">` above
|
|
50
|
+
its question, per `picker-contract.md` "Step narration". A step that resolves
|
|
51
|
+
without asking (one account, no siblings) still prints its line with the
|
|
52
|
+
resolution noted, so the numbering reads continuously instead of jumping.
|
|
53
|
+
|
|
49
54
|
**Model tier resolution** (same step, once per run): read `prefs.global.modelFallback`. If `fableEnabled` is `false`, every `preferredModel: fable` persona resolves to `opus` for this run and the Phase 3 Claude Code panel is 2 reviewers, not 3; print the one-line INFO. Then, if `premiumTierUntil` is set and in the past, apply the date-gate trigger - `preferredModel` personas dispatch on `fallbackModel`, with the one-line WARN. Both lines and the exact ordering: `$HOME/.claude/multi-agent-refs/features/model-fallback.md`. Dispatch-error and budget triggers apply per-dispatch later; nothing else to do here.
|
|
50
55
|
|
|
51
56
|
**First-run guard**: After loading prefs, check if `keychainMapping` has at least one non-null value. If ALL values are null (template defaults - setup never ran), show:
|
|
@@ -323,8 +328,7 @@ above, recorded as `baseBranchSource: "input"`. Everything else asks, and
|
|
|
323
328
|
One row is a normal outcome of the rule-5 filter and is **still asked**, with a
|
|
324
329
|
real second option - `picker-contract.md`, "Two options or it is not a question".
|
|
325
330
|
|
|
326
|
-
This holds in every mode.
|
|
327
|
-
does not skip Phase 0's pickers. Autopilot resolves them without prompting, which still
|
|
331
|
+
This holds in every mode. Autopilot resolves them without prompting, which still
|
|
328
332
|
writes the fields; its base-branch resolution order is in `features/base-branch-evidence.md`.
|
|
329
333
|
|
|
330
334
|
**TTL filter for recent branches**:
|
|
@@ -412,11 +416,11 @@ Branch name is deterministic - no user confirmation needed.
|
|
|
412
416
|
#### Step 5b - Workspace (worktree or local)
|
|
413
417
|
|
|
414
418
|
Ask where the branch lives - the wording, the two options and what local costs
|
|
415
|
-
are in `modes.md`, "Local Mode". Here because Step 4 named the branch
|
|
416
|
-
acts on the answer
|
|
419
|
+
are in `modes.md`, "Local Mode". Here because Step 4 named the branch and Step
|
|
420
|
+
6b acts on the answer.
|
|
417
421
|
|
|
418
|
-
**Who is asked.**
|
|
419
|
-
|
|
422
|
+
**Who is asked.** Every interactive entry (`workspaceSource: "asked"`). There is
|
|
423
|
+
no flag and no command name that pre-answers it.
|
|
420
424
|
Every autopilot entry resolves it to a worktree and never asks (`autopilot`).
|
|
421
425
|
|
|
422
426
|
**Persist** `state.localMode` (semantics unchanged) and `state.workspaceSource`,
|
|
@@ -448,7 +452,7 @@ Log: `Identity: {identity.name} <{identity.email}>`
|
|
|
448
452
|
|
|
449
453
|
1. `git -C $PROJECT_ROOT fetch origin`
|
|
450
454
|
|
|
451
|
-
**If local** (Step 5b answered local
|
|
455
|
+
**If local** (Step 5b answered local):
|
|
452
456
|
|
|
453
457
|
```bash
|
|
454
458
|
if [ -n "$(git -C $PROJECT_ROOT status --porcelain)" ]; then
|
|
@@ -464,7 +468,7 @@ git -C $PROJECT_ROOT config user.email "{identity.email}"
|
|
|
464
468
|
|
|
465
469
|
**If worktree** (Step 5b answered worktree, or autopilot resolved it): 2. Worktree path: Jira → `.worktrees/{jiraId}/`, GitHub → `.worktrees/GH{issueNo}/`, free-text → `.worktrees/task-{shortId}/` 3. **Heal stale admin state first** (see "Worktree stale-lock heal" below) and **apply the residue guard** (see "Worktree residue guard" below), then `git -C $PROJECT_ROOT worktree add {path} -b {branch} origin/{baseBranch}` (if exists: enter, pull) 4. Set identity: `git -C {worktree-path} config user.name/email` 5. Create log dir + `agent-log.md` + `agent-state.json` at `$HOME/.claude/logs/multi-agent/{project}/{task-id}/`, never inside the worktree:
|
|
466
470
|
|
|
467
|
-
**Worktree location convention (cited by every other command):** always `{projectRoot}/.worktrees/{taskId}`, inside the repo, never under `$HOME`. `{taskId}` is the directory name from the rule above (`DC-<shortId>` for `/multi-agent:design-check`). `.worktrees` is fixed, not a preference: no `worktreeBasePath` key exists, and `gc-worktrees.sh`, `purge.sh`, the cost renderers and `usage-report.mjs` resolve `<repo>/.worktrees/` by name. Multi-repo tasks get one worktree per repo (the loop below);
|
|
471
|
+
**Worktree location convention (cited by every other command):** always `{projectRoot}/.worktrees/{taskId}`, inside the repo, never under `$HOME`. `{taskId}` is the directory name from the rule above (`DC-<shortId>` for `/multi-agent:design-check`). `.worktrees` is fixed, not a preference: no `worktreeBasePath` key exists, and `gc-worktrees.sh`, `purge.sh`, the cost renderers and `usage-report.mjs` resolve `<repo>/.worktrees/` by name. Multi-repo tasks get one worktree per repo (the loop below); a local answer creates none and `worktreePath` is `$PROJECT_ROOT`.
|
|
468
472
|
|
|
469
473
|
**Worktree stale-lock heal (required before every `worktree add`):** a run killed mid-`worktree add` (OOM, SIGTERM, disk full) leaves a locked or broken admin entry under `.git/worktrees/{id}/`, so the retry fails with `fatal: '<path>' already exists`. Always run the heal first - it is a no-op on a clean repo:
|
|
470
474
|
|
|
@@ -584,30 +588,6 @@ Persist: `"taskType": "component" | "bugfix" | "feature" | "refactor" | "chore"`
|
|
|
584
588
|
|
|
585
589
|
Log: `Phase 0 Step 7: taskType = {component|bugfix|feature|refactor|chore}`
|
|
586
590
|
|
|
587
|
-
#### Step 7.5 - Pipeline depth (Full / Short)
|
|
588
|
-
|
|
589
|
-
Ask the depth question from `$HOME/.claude/multi-agent-refs/phases/modes.md` "Pipeline depth" - it carries the wording, the per-`taskType` recommendation and the mode tables. Here because the recommendation needs `taskType` (Step 7), which needs the fetched issue (Step 1) and the branch (Step 3).
|
|
590
|
-
|
|
591
|
-
**Who is asked.** `/multi-agent` and `/multi-agent:local` only. Both autopilot entries and analysis mode skip it; autopilot always runs Full.
|
|
592
|
-
|
|
593
|
-
When the intake carried an analysis document or a Figma reference, say so **inside** the question: Short skips the only two phases that would turn that document into a task breakdown, and the user should learn that before choosing, not after.
|
|
594
|
-
|
|
595
|
-
**Pass the default as a 1-based index, never a label.** `ask-choice.sh` takes the first option on a non-TTY, and a label-valued default matches nothing once the options render in `outputLanguage`. Reasoning: `modes.md`, "Pipeline depth".
|
|
596
|
-
|
|
597
|
-
```bash
|
|
598
|
-
# Full is option 1, Short is option 2 (modes.md "Pipeline depth")
|
|
599
|
-
DEPTH_DEFAULT_INDEX=1; [ "$DEPTH_RECOMMENDATION" = "short" ] && DEPTH_DEFAULT_INDEX=2
|
|
600
|
-
ASK_CHOICE_DEFAULT="$DEPTH_DEFAULT_INDEX" \
|
|
601
|
-
$HOME/.claude/lib/ask-choice.sh "<localized: 'Which pipeline for this task?'>" \
|
|
602
|
-
"<localized: 'Full'>" "<localized: 'Short'>"
|
|
603
|
-
```
|
|
604
|
-
|
|
605
|
-
**Persist.** Short sets `state.onlyDevelop = true`; Full leaves it `false`. The key is unchanged - only who sets it changed - so every downstream reader keeps working.
|
|
606
|
-
|
|
607
|
-
**Now register the rest of the widget.** This answer is the first moment the phase set is known, so the remaining tiles are created here and not before - Full `1 2 3 4 5 6 7`, Short `3 4 5 6 7`, `:local` dropping 5 from either. `add` each, then `phase-tracker.sh tiles --new`, which emits TaskCreate only for phases that carry no tile yet. Contract: `tracker-contract.md`, "Deferred registration".
|
|
608
|
-
|
|
609
|
-
Log: `Phase 0 Step 7.5: depth = {full|short} (recommended {full|short}, source {user|autopilot|default})`
|
|
610
|
-
|
|
611
591
|
#### Step 7.6 - Test baseline (opt-in, `prefs.global.testBaseline.enabled`, default `false`)
|
|
612
592
|
|
|
613
593
|
Phase 3 Gate 3 cannot tell an inherited red suite from one this run broke, so it blocks on someone else's bug or the agent "fixes" tests it never touched. Runs after Step 6, only when the stack has a test command; skipped in analysis mode.
|
|
@@ -644,7 +624,8 @@ Persist that file as `state.evidenceCapability`, then build the menu from it, ne
|
|
|
644
624
|
|
|
645
625
|
A closed option keeps its row and prints the probe's reason verbatim. No `uiTestTargets` closes 2; `mcp` false closes 3; a missing `device` or `recorder` closes both. Targets present with no `matchingTests` leaves 2 open, warning that the whole UI suite will run. **When only option 1 is open, do not ask**: `testDepth = unit`, `testDepthSource = forced`, and the tier 3 gap is written.
|
|
646
626
|
|
|
647
|
-
Default as a 1-based index, never a label -
|
|
627
|
+
Default as a 1-based index, never a label: `ask-choice.sh` takes the first
|
|
628
|
+
option on a non-TTY, so a label-valued default silently becomes option 1.
|
|
648
629
|
|
|
649
630
|
```bash
|
|
650
631
|
DEPTH_DEFAULT_INDEX=1
|
|
@@ -653,7 +634,7 @@ DEPTH_DEFAULT_INDEX=1
|
|
|
653
634
|
ASK_CHOICE_DEFAULT="$DEPTH_DEFAULT_INDEX" $HOME/.claude/lib/ask-choice.sh ...
|
|
654
635
|
```
|
|
655
636
|
|
|
656
|
-
Asked by
|
|
637
|
+
Asked by every interactive entry; autopilot reads `prefs.global.testDepth.default` and degrades to the best open option. Asked here, not Phase 3, which four of the eight modes drop.
|
|
657
638
|
|
|
658
639
|
Log: `Phase 0 Step 7.7: testDepth = {unit|unit+ui|unit+mcp} (source {user|autopilot|default|forced}), tier1/tier2 = {open|closed}`
|
|
659
640
|
|
|
@@ -1,6 +1,8 @@
|
|
|
1
1
|
### Phase 1: Plan (Opus)
|
|
2
2
|
|
|
3
|
-
> **TLDR** - One phase, two halves. First the codebase is explored and the analysis document written (the design contract the rest of the run reads); then that document is decomposed into concrete tasks with file-level targets, risk grading and architecture review. The phase ends at the **Plan Approval Gate**: in normal mode the orchestrator asks structured clarification questions when scope is ambiguous (max 2 rounds), renders the plan, and loops on free-text edits until the user approves or aborts. The gate is **skipped entirely** for
|
|
3
|
+
> **TLDR** - One phase, two halves. First the codebase is explored and the analysis document written (the design contract the rest of the run reads); then that document is decomposed into concrete tasks with file-level targets, risk grading and architecture review. The phase ends at the **Plan Approval Gate**: in normal mode the orchestrator asks structured clarification questions when scope is ambiguous (max 2 rounds), renders the plan, and loops on free-text edits until the user approves or aborts. The gate is **skipped entirely** for `autopilot`, which has no one to ask.
|
|
4
|
+
|
|
5
|
+
> **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
|
|
4
6
|
|
|
5
7
|
<!-- progress-contract: applied -->
|
|
6
8
|
Progress emission per `$HOME/.claude/multi-agent-refs/progress-contract.md` - lines for each Explore dispatch, each finish, analyst synthesis start, `analysis.json` write.
|
|
@@ -412,9 +414,9 @@ and stored, only invisible.
|
|
|
412
414
|
printf '%s' "$PLAN_JSON" | bash "$HOME/.claude/lib/plan-todos.sh" set "$TASK_ID" -
|
|
413
415
|
```
|
|
414
416
|
|
|
415
|
-
`set` accepts a planning-output document and converts it itself.
|
|
416
|
-
|
|
417
|
-
|
|
417
|
+
`set` accepts a planning-output document and converts it itself. Writing the
|
|
418
|
+
conversion out here too would make the `tasks[]`-to-`todos[]` mapping a thing two
|
|
419
|
+
files define, and the copy in prose is the one nothing tests.
|
|
418
420
|
|
|
419
421
|
Phase 2 (Dev) then iterates with `plan-todos.sh next "$TASK_ID"` until empty, calling `start` before each step and `complete` (with notes) or `fail`/`skip` after. Phase 3 (Review) reads the Todo list to verify all `completed` items map to diff hunks. Phase 5 (Report) renders `list` into the agent-log + PR body.
|
|
420
422
|
|
|
@@ -441,10 +443,6 @@ Log: "Phase 1: Consistency - requirements:{N/N mapped} anchors:{ok|M unanchore
|
|
|
441
443
|
- `state.autopilot === true` (autopilot contract: zero interaction)
|
|
442
444
|
- Autopilot safety classifier returns `recommendPause: false` (see Step 5c below)
|
|
443
445
|
|
|
444
|
-
OR:
|
|
445
|
-
|
|
446
|
-
- `state.onlyDevelop === true` (Short pipeline: direct to Phase 2, no plan)
|
|
447
|
-
|
|
448
446
|
In the skipped case, log `🧠 Phase 1: Plan - gate skipped ({mode}), proceeding to Phase 2` and go to Phase 2.
|
|
449
447
|
|
|
450
448
|
##### 5c - Autopilot safety classifier (runs before 5a/5b skip decision)
|
|
@@ -571,9 +569,8 @@ The pipeline shapes interact with the gate as follows. This table is the source
|
|
|
571
569
|
|
|
572
570
|
| Mode | Clarification | Approval Loop | Safety Classifier | Notes |
|
|
573
571
|
|---|---|---|---|---|
|
|
574
|
-
|
|
|
575
|
-
|
|
|
576
|
-
| `autopilot`, `:local-autopilot` | ❌ | ❌ conditional | ✅ (if `autopilotSafetyGate !== false`) | Always Full, so a plan exists. Safe plans proceed silently; high-risk plans trigger a one-time manual approval. Log records the score. |
|
|
572
|
+
| Interactive (`/multi-agent`) | ✅ (max 2 rounds) | ✅ | - (redundant when a human approves) | Full gate |
|
|
573
|
+
| `/multi-agent:autopilot` | ❌ | ❌ conditional | ✅ (if `autopilotSafetyGate !== false`) | A plan always exists. Safe plans proceed silently; high-risk plans trigger a one-time manual approval. Log records the score. |
|
|
577
574
|
|
|
578
575
|
**Why autopilot now has an escape hatch:** the old "zero interaction - fully trust the scope" contract held well for tightly-scoped batch workflows (figma component iteration over known-safe components) but broke in edge cases - schema migrations auto-merging, security-path drift going silent, delete-without-test sprawls. The safety classifier (Step 5c) is opt-out so the default protects against the edge-case cost; users running known-safe workflows can flip `prefs.global.autopilotSafetyGate = false` to restore pre-v7.0 behavior.
|
|
579
576
|
|
|
@@ -582,18 +579,10 @@ The pipeline shapes interact with the gate as follows. This table is the source
|
|
|
582
579
|
After plan generation (and after each edit-loop iteration), forward the planning model's call totals so Phase 5's Cost Breakdown captures Phase 1 (`model=` names the rung that actually ran: `fable`, or `opus` after a fallback step):
|
|
583
580
|
|
|
584
581
|
```bash
|
|
585
|
-
LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID"
|
|
582
|
+
LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 1 plan.generated \
|
|
586
583
|
model=fable tokens_in=$IN tokens_out=$OUT duration_ms=$DUR iteration=$N
|
|
587
584
|
```
|
|
588
585
|
|
|
589
586
|
Best-effort. See `$HOME/.claude/multi-agent-refs/progress-contract.md#token-telemetry-forwarding`.
|
|
590
587
|
|
|
591
|
-
|
|
592
|
-
|
|
593
|
-
## Token telemetry - invoke after every LLM call
|
|
594
|
-
|
|
595
|
-
```bash
|
|
596
|
-
bash $HOME/.claude/scripts/phase-tracker.sh tokens 2 <input_count> <output_count>
|
|
597
|
-
```
|
|
598
|
-
|
|
599
|
-
Contract and rationale: `progress-contract.md` -> Token telemetry forwarding.
|
|
588
|
+
Both halves of this phase report into the same counter: see "Token telemetry" above.
|