@mmerterden/multi-agent-pipeline 15.16.0 → 16.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +124 -0
- package/README.md +2 -2
- package/README.tr.md +1 -2
- package/docs/architecture.md +2 -2
- package/docs/ecosystem.md +9 -6
- package/docs/features.md +2 -3
- package/install/_plugin-skills.mjs +28 -2
- package/install/templates/copilot-instructions.md +8 -7
- package/package.json +1 -1
- package/pipeline/commands/multi-agent/SKILL.md +5 -6
- package/pipeline/commands/multi-agent/analysis/SKILL.md +77 -588
- package/pipeline/commands/multi-agent/analysis-resolve/SKILL.md +5 -69
- package/pipeline/commands/multi-agent/channels/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/complaint-analysis/SKILL.md +5 -4
- package/pipeline/commands/multi-agent/dev/SKILL.md +8 -280
- package/pipeline/commands/multi-agent/dev-autopilot/SKILL.md +12 -124
- package/pipeline/commands/multi-agent/dev-local/SKILL.md +8 -111
- package/pipeline/commands/multi-agent/dev-local-autopilot/SKILL.md +11 -113
- package/pipeline/commands/multi-agent/help/SKILL.md +61 -56
- package/pipeline/commands/multi-agent/ios-coding-standard/SKILL.md +4 -2
- package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +2 -4
- package/pipeline/commands/multi-agent/refactor/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/resume-local/SKILL.md +3 -3
- package/pipeline/commands/multi-agent/setup/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/stack/SKILL.md +10 -9
- package/pipeline/commands/multi-agent/store-ready/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/sync/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/update/SKILL.md +1 -1
- package/pipeline/lib/context-link-extractor.sh +38 -0
- package/pipeline/lib/fetch-document.sh +190 -0
- package/pipeline/multi-agent-refs/_dev-context.md +4 -0
- package/pipeline/multi-agent-refs/analysis/evidence.md +213 -0
- package/pipeline/multi-agent-refs/analysis/intake.md +167 -0
- package/pipeline/multi-agent-refs/analysis/locked.md +53 -0
- package/pipeline/multi-agent-refs/analysis/render.md +133 -0
- package/pipeline/multi-agent-refs/analysis/resolve.md +76 -0
- package/pipeline/multi-agent-refs/analysis/synthesis.md +98 -0
- package/pipeline/multi-agent-refs/analysis-template.md +58 -11
- package/pipeline/multi-agent-refs/complaint-analysis-template.md +1 -1
- package/pipeline/multi-agent-refs/component-dispatch.md +5 -5
- package/pipeline/multi-agent-refs/cross-cli-contract.md +9 -7
- package/pipeline/multi-agent-refs/features/skill-conformance.md +1 -1
- package/pipeline/multi-agent-refs/features/url-enrichment.md +13 -3
- package/pipeline/multi-agent-refs/knowledge.md +2 -2
- package/pipeline/multi-agent-refs/payload-contracts.md +1 -1
- package/pipeline/multi-agent-refs/phases/modes.md +73 -53
- package/pipeline/multi-agent-refs/phases/phase-0-init.md +21 -2
- package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +26 -0
- package/pipeline/multi-agent-refs/phases/phase-2-planning.md +26 -12
- package/pipeline/multi-agent-refs/phases/phase-3-dev.md +17 -18
- package/pipeline/multi-agent-refs/phases/phase-4-review.md +29 -7
- package/pipeline/multi-agent-refs/phases/phase-5-test.md +1 -1
- package/pipeline/multi-agent-refs/phases/phase-6-commit.md +8 -0
- package/pipeline/multi-agent-refs/phases/phase-7-report.md +1 -1
- package/pipeline/multi-agent-refs/phases.md +9 -7
- package/pipeline/multi-agent-refs/progress-contract.md +1 -1
- package/pipeline/multi-agent-refs/readiness-review.md +2 -0
- package/pipeline/multi-agent-refs/rules.md +2 -2
- package/pipeline/multi-agent-refs/tracker-contract.md +32 -13
- package/pipeline/multi-agent-refs/wiki-capture.md +2 -2
- package/pipeline/schemas/agent-state.schema.json +3 -3
- package/pipeline/schemas/analysis-output.schema.json +17 -0
- package/pipeline/schemas/analysis-spec.schema.json +69 -1
- package/pipeline/schemas/complaint-analysis-spec.schema.json +1 -1
- package/pipeline/schemas/prefs.schema.json +47 -0
- package/pipeline/schemas/token-budget.json +2 -2
- package/pipeline/scripts/_stack-routing.mjs +17 -12
- package/pipeline/scripts/build-stack-plugins.mjs +16 -6
- package/pipeline/scripts/cost-table.json +1 -1
- package/pipeline/scripts/gen-mode-dispatch.mjs +20 -30
- package/pipeline/scripts/run-aggregator.mjs +1 -1
- package/pipeline/scripts/validate-analysis-doc.mjs +36 -3
- package/pipeline/skills/.skills-index.json +40 -7
- package/pipeline/skills/shared/README.md +12 -9
- package/pipeline/skills/shared/core/multi-agent/SKILL.md +9 -5
- package/pipeline/skills/shared/core/multi-agent-dev/SKILL.md +7 -61
- package/pipeline/skills/shared/core/multi-agent-dev-autopilot/SKILL.md +11 -51
- package/pipeline/skills/shared/core/multi-agent-dev-local/SKILL.md +7 -33
- package/pipeline/skills/shared/core/multi-agent-dev-local-autopilot/SKILL.md +10 -38
- package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +28 -18
- package/pipeline/skills/shared/core/multi-agent-ios-coding-standard/SKILL.md +4 -2
- package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +2 -4
- package/pipeline/skills/shared/core/multi-agent-refactor/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-resume-local/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-stack/SKILL.md +10 -9
- package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +2 -2
- package/pipeline/skills/shared/external/evidence-github/SKILL.md +45 -0
- package/pipeline/skills/shared/external/evidence-registry/SKILL.md +33 -0
- package/pipeline/skills/shared/external/signal-community/SKILL.md +44 -0
- package/pipeline/skills/skills-index.md +10 -7
|
@@ -21,24 +21,13 @@ Companion command to `/multi-agent:analysis`. Takes an `analysis/<feature>-<plat
|
|
|
21
21
|
- **Locked 24 (Pass B footnote).** When a resolution overrides a convention cell in Section 13 (or any per-platform projection), the cell's footnote is rewritten to `^[user-override: resolved via analysis-resolve <date>]`.
|
|
22
22
|
- **Locked 30 (no MCP outside analysis phase).** This command makes NO Figma MCP calls, no `api.figma.com` requests, no `figma.com/design/...` fetches. A row that genuinely needs new design information gets only `Defer` plus a printed recommendation to re-run `/multi-agent:analysis` for that input. Repo reads and `~/.claude/lib/extract-conventions.sh` re-runs are allowed (they are local evidence, not design fetches).
|
|
23
23
|
|
|
24
|
-
##
|
|
25
|
-
|
|
26
|
-
1. **One question per `AskUserQuestion` call.** Never batch Section 20 rows. Sequential resolution keeps each decision explicit and traceable.
|
|
27
|
-
2. **Up to 3 source-labeled candidates plus Defer = 4 options max.** Each candidate's `description` carries its source label: `From evidence - <doc section / citation>`, `From repo - <file:line>`, `AI reasoned - <one-line rationale>`. The auto-provided Other accepts free text and stop tokens.
|
|
28
|
-
3. **Three sources, never blended.**
|
|
29
|
-
- **From evidence**: the answer is already derivable from the doc's own sections and their citations (Sections 5, 6, 9, 13, 21) or from cached `state.analysisSpec.evidence.*` when the session still holds it. Cite the section or evidence bucket.
|
|
30
|
-
- **From repo**: the batched Phase 1 repo lookup (one Explore subagent) returned an answer with a `file:line` citation.
|
|
31
|
-
- **AI reasoned**: contextual reasoning grounded in the loaded doc (feature scope, platform, surrounding sections). Requires a one-line rationale; skip if the basis is no stronger than guessing.
|
|
32
|
-
4. **Never invent.** If no source produces a credible candidate, offer only Defer and Other.
|
|
33
|
-
5. **Stop tokens** (case-insensitive, typed into Other): `stop`, `pause`, `dur`, `kes`. Halt immediately, save current state, jump to the Finalize phase.
|
|
34
|
-
6. **Follow-up questions surface immediately.** If an answer creates a new question ("use a new endpoint" leads to "endpoint name?"), insert a new `Acik / Open` row into Section 20 directly after the current row and process it next, not at queue end.
|
|
35
|
-
7. **Save after every answer.** Each resolution writes the full doc to disk atomically before the next question. Interruption loses nothing.
|
|
36
|
-
8. **Status enum is fixed:** `Acik / Open`, `Girdi bekleniyor / Pending input`, `Karar verildi / Decided` (render the side matching the doc's `language`). Only the Status cell and a trailing `[Cozum: <value> - <source>]` / `[Resolved: <value> - <source>]` fragment on the question cell are edited in Section 20; Owner is never synthesized or changed.
|
|
37
|
-
9. **Sibling propagation is explicit, never silent.** Per-platform sibling files (front-matter `siblings:`) are touched only in the Finalize phase, only for rows whose question text matches verbatim, and only after the user approves one summary question. Platform-specific rows (convention fallbacks, per-platform reuse rows) never propagate.
|
|
38
|
-
10. **`evidence_digest` stays untouched.** Resolving questions does not change the evidence inputs; the front-matter digest and `generated` timestamp are preserved. Only the Section 23 changelog records the revision.
|
|
24
|
+
## Engine
|
|
39
25
|
|
|
40
|
-
|
|
26
|
+
The resolution engine - this command's own Locked decisions, the batched repo lookup, the sequential resolution loop and the finalize pass - lives in `$HOME/.claude/multi-agent-refs/analysis/resolve.md`. Pipeline Phase 4 (analysis mode) and Phase 2 (full pipeline) mount the same file, so there is one walk with three entry points.
|
|
27
|
+
|
|
28
|
+
Inherited decisions: `$HOME/.claude/multi-agent-refs/analysis/locked.md`.
|
|
41
29
|
|
|
30
|
+
## Input
|
|
42
31
|
- `$ARGUMENTS` - optional doc path. If empty, auto-glob at Phase 0.
|
|
43
32
|
- `--autonomous` - auto-pick the strongest candidate where exactly one source is credible; Defer otherwise. No per-row prompts; final report still printed.
|
|
44
33
|
|
|
@@ -60,59 +49,6 @@ Companion command to `/multi-agent:analysis`. Takes an `analysis/<feature>-<plat
|
|
|
60
49
|
|
|
61
50
|
**Echo** one line: `Resolve: <doc> | repo: <path> | mode: <interactive|autonomous>. Parsing Section 20.`
|
|
62
51
|
|
|
63
|
-
### Phase 1 - Parse and batched repo lookup
|
|
64
|
-
|
|
65
|
-
1. Read the doc end-to-end. Capture:
|
|
66
|
-
- Front-matter: `feature`, `platform`, `language`, `mode`, `siblings[]`.
|
|
67
|
-
- The Risks and Open Questions table: ordered rows `{index, question, owner, status}`.
|
|
68
|
-
- Context for reasoning: Summary, Goals, API Contracts, Architecture Plan (concept table + footnotes), Files to Add, References.
|
|
69
|
-
2. Filter to rows whose status is `Acik / Open` or `Girdi bekleniyor / Pending input`. If zero remain, report `Section 20 has no open rows.` and stop.
|
|
70
|
-
3. **Classify each row** by its auto-population fingerprint:
|
|
71
|
-
|
|
72
|
-
| Class | Fingerprint (from analysis.md auto-population rules) | Extra candidate source |
|
|
73
|
-
|---|---|---|
|
|
74
|
-
| `convention-fallback` | row text contains `convention fallback applied` or cites `conventions-defaults.md:C<n>-` (Locked 23) | re-run `~/.claude/lib/extract-conventions.sh <repo> <platform>` for that one field; offer the default vs the fresh extraction |
|
|
75
|
-
| `reuse-vs-new` | row text matches `existing <X> candidate found; reuse or document why a new one is needed` (Locked 11) | repo lookup reads the candidate `file:line` and summarizes its fit |
|
|
76
|
-
| `citation-tbd` | row originated from a `[label TBD - see Open Questions]` downgrade (Locked 3) | repo localization keys / Figma evidence already cited elsewhere in the doc |
|
|
77
|
-
| `legacy-decision` | row phrased `current code does X; should the new feature keep, change, or drop this?` (Locked 4) | repo lookup confirms current behavior with `file:line` |
|
|
78
|
-
| `standards-conflict` | row cites a binding standards source (Locked 8) | quote the binding section verbatim from References |
|
|
79
|
-
| `design-gap` | row needs new Figma/design information (tier-3 forced questions, missing variants per Locked 12/19) | NONE - Locked 30; offer only Defer + the re-run recommendation |
|
|
80
|
-
| `generic` | anything else | all three standard sources |
|
|
81
|
-
|
|
82
|
-
4. **Batched repo lookup** (single Explore subagent, one call regardless of row count; skip if every row classified `design-gap`):
|
|
83
|
-
> "For the feature `<feature>` on platform `<platform>`, answer the following questions from the repo at `<repo-root>`. Return one entry per question ID with either an answer plus `file:line` citation OR `not-found`. Scope reading to paths the questions name plus their direct dependencies; do not roam, do not propose changes.
|
|
84
|
-
> Q1: <row 1 text> ..."
|
|
85
|
-
5. Cache per-question answers keyed by row index. These feed the `From repo` candidate in Phase 2.
|
|
86
|
-
|
|
87
|
-
### Phase 2 - Sequential resolution loop
|
|
88
|
-
|
|
89
|
-
For each open row in source order:
|
|
90
|
-
|
|
91
|
-
1. **Generate candidates** per the row's class (max 3). Skip any source with no credible output. For `design-gap` rows, skip candidate generation entirely.
|
|
92
|
-
2. **Ask** (interactive mode) - single AskUserQuestion:
|
|
93
|
-
- `header`: the target section anchor, max 12 chars (e.g. `S13 Arch`, `S9 API`, `S10 L10n`, `Conventions`)
|
|
94
|
-
- `question`: the row text verbatim, prefixed `[R-<index>]` (follows the doc, so it may be Turkish; that is correct here - the row IS the doc content)
|
|
95
|
-
- `options`: up to 3 candidates + `Defer - keep open`. Candidate descriptions carry the source label per New Locked 2.
|
|
96
|
-
- Autonomous mode: pick the single credible candidate if exactly one source produced one; otherwise Defer. Log each auto-decision as one line.
|
|
97
|
-
3. **Apply**:
|
|
98
|
-
- **Concrete answer**: locate the target body section from the row's anchor or content (convention rows target the Architecture Plan concept table; reuse rows target Files to Add + Architecture Plan; API rows target API Contracts; etc.).
|
|
99
|
-
- `From evidence` / `AI reasoned` / user Other: write as primary content in the doc's `language`, ASCII punctuation, citations preserved.
|
|
100
|
-
- `From repo` describing current behavior: write as `> Legacy reference: <text> (file:line)` under the target entry (Locked 4). A repo answer that the user adopts as the forward decision is primary content WITH its `file:line` citation.
|
|
101
|
-
- `convention-fallback`: update the convention cell, rewrite its footnote to `^[user-override: resolved via analysis-resolve <YYYY-MM-DD>]` (Locked 24).
|
|
102
|
-
- Update the Section 20 row: Status to `Karar verildi / Decided`, append ` [Cozum: <one-line> - <source>]` (tr) / ` [Resolved: <one-line> - <source>]` (en) to the question cell. Do not delete the row; the table stays a decision log.
|
|
103
|
-
- **Defer**: Status to `Girdi bekleniyor / Pending input`. Row otherwise untouched.
|
|
104
|
-
- **Stop token**: save, jump to Phase 3.
|
|
105
|
-
- **Follow-up**: insert the new `Acik / Open` row right after the current one (New Locked 6) and process it next.
|
|
106
|
-
4. **Save the full doc to disk** before the next iteration (New Locked 7).
|
|
107
|
-
|
|
108
|
-
### Phase 3 - Finalize
|
|
109
|
-
|
|
110
|
-
1. **Sibling propagation** (only if front-matter `siblings[]` is non-empty AND at least one resolved row's question text appears verbatim with an open status in a sibling file on disk): one AskUserQuestion, `header: "Siblings"`, question `<localized: "N resolved rows also appear open in sibling file(s) <list>. Apply the same resolutions there?">`, options `Apply to all listed` / `Skip siblings`. On apply: repeat the Phase 2 apply step per matching row per sibling, then give each touched sibling its own changelog row. Platform-specific classes (`convention-fallback`, `reuse-vs-new`) are excluded from matching (New Locked 9).
|
|
111
|
-
2. **Changelog**: append one row to the Changelog section of every touched file: next version (integer scheme `v1 -> v2`; dotted scheme bumps the minor), today's date, author `analysis-resolve`, change `Resolved <N> of <M> Section 20 rows; <K> deferred`.
|
|
112
|
-
3. **Punctuation gate**: run the Locked 7 verification grep over every touched file; fix any hit before reporting.
|
|
113
|
-
4. **Report** (in `outputLanguage`, max 12 lines): doc path(s) + new version, counts (resolved / deferred / follow-ups created / still open), `design-gap` rows that need a `/multi-agent:analysis` re-run (list inputs to re-supply), and - if open rows remain - a reminder that re-running this command resumes where it left off (state is the doc itself; no separate state file).
|
|
114
|
-
|
|
115
|
-
**Stop. No commit, no branch, no dispatch.** The user reviews the diff and commits manually (Locked 6).
|
|
116
52
|
|
|
117
53
|
## Reusable refs
|
|
118
54
|
|
|
@@ -245,7 +245,7 @@ Emitted only when `reportContent.workSummary === true` and at least one of (`age
|
|
|
245
245
|
1. **Task header** - `taskId`, `branch`, `baseBranch`, `prNumber` from `agent-state.json` (or explicit flags for post-hoc invocation).
|
|
246
246
|
2. **Scope delivered** - Phase 2 `planTodos[]` / `tasks[]` rendered as `done` (status=done) or `pending` (anything else) rows. Task id + title shown; `(deferred - rationale)` appended if the task's `status` is `"deferred"`.
|
|
247
247
|
3. **Changed files** - `git -C $WORKTREE diff --numstat $baseBranch...HEAD`. Shows `` `path` (+add / -del) `` per row, capped at 20 with a `_... +N more files not shown_` footer when exceeded. Total adds/dels + file count in section header.
|
|
248
|
-
4. **Review outcome** - from `reviewConsensus` (pre-v6.1) or `phases["4"].triage`: `{accepted} accepted · {deferred} deferred · {rejected} rejected · approved={bool}`. Hidden entirely when all three buckets are empty (normal for
|
|
248
|
+
4. **Review outcome** - from `reviewConsensus` (pre-v6.1) or `phases["4"].triage`: `{accepted} accepted · {deferred} deferred · {rejected} rejected · approved={bool}`. Hidden entirely when all three buckets are empty (normal for a run that never reached Phase 4).
|
|
249
249
|
5. **Phase tick strip** - single line from the tracker state (`render-work-summary.sh` resolves worktree/artifacts copies, then `$HOME/.claude/logs/multi-agent/{taskId}/tracker-state.json`): `0 Init [done] · 1 Analysis [done] · 2 Planning [done] · 3 Dev [done] · 4 Review [done] · 5 Test skipped · 6 Commit [done] · 7 Report active`. Marks: `done` completed · `active` in_progress · `failed` failed · `skipped` skipped · `·` pending.
|
|
250
250
|
|
|
251
251
|
**Output template:**
|
|
@@ -28,7 +28,7 @@ Triage of customer-reported errors across the layers the team owns (client apps
|
|
|
28
28
|
10. **Verdict citation discipline.** Every `client`/`bff` verdict cites at least one Graylog message (timestamp + source) AND one repo evidence row (`file:line`). Anything less is `insufficient-evidence`. Every `core` verdict cites the Graylog message that names the upstream service.
|
|
29
29
|
11. **One complaint batch per run.** Mixed batches spanning unrelated products are user error: surface it and ask to split.
|
|
30
30
|
12. **No auto-commit.** The local report is written to the working tree; the user commits it themselves if they want.
|
|
31
|
-
13. **Client/bff verdicts carry a development handoff.** Every `client`/`bff` verdict renders a fix plan grounded in the EXISTING architecture (the `repoEvidence[]` files are the reference: name the concrete files/components to touch, reuse-first, no invented structures) plus a ready-to-run dev prompt (English, one fenced block, `/multi-agent
|
|
31
|
+
13. **Client/bff verdicts carry a development handoff.** Every `client`/`bff` verdict renders a fix plan grounded in the EXISTING architecture (the `repoEvidence[]` files are the reference: name the concrete files/components to touch, reuse-first, no invented structures) plus a ready-to-run dev prompt (English, one fenced block, `/multi-agent`-compatible). The handoff is part of the report - this command still never runs dev itself (Locked 3). Core and insufficient-evidence verdicts never get a fix plan (Locked 5).
|
|
32
32
|
|
|
33
33
|
## Steps
|
|
34
34
|
|
|
@@ -120,7 +120,7 @@ Verdict shape: `verdict: {category, layer, confidence: high|medium|low, rational
|
|
|
120
120
|
**Development handoff (client/bff only, Locked 13)**: for each `client`/`bff` verdict, derive from the `repoEvidence[]` rows:
|
|
121
121
|
|
|
122
122
|
- **Fix plan**: 2-5 numbered steps referencing the existing architecture by `file:line` - which service/view/handler to change, what to reuse (reuse-first: prefer extending the cited components over adding new ones), which tests to add. No speculative rewrites.
|
|
123
|
-
- **Dev prompt**: one fenced English block the user can paste into `/multi-agent
|
|
123
|
+
- **Dev prompt**: one fenced English block the user can paste into `/multi-agent` (or a Jira description): complaint summary, root cause, the cited files, the fix plan steps, and the acceptance check. Include the complaint id (`[C-NN]`) for traceability.
|
|
124
124
|
|
|
125
125
|
Set `phase: "drafting"`.
|
|
126
126
|
|
|
@@ -128,7 +128,7 @@ Set `phase: "drafting"`.
|
|
|
128
128
|
|
|
129
129
|
1. Render the report per `$HOME/.claude/multi-agent-refs/complaint-analysis-template.md` (8 fixed sections; single-language body in `outputLanguage`; verdict tokens English per Locked 9) to `/tmp/complaint-analysis-<run-slug>-<UTC-iso8601>/report.md`. Store `outputs.draftDir`.
|
|
130
130
|
2. **Humanizer pass (MANDATORY: actually invoke the `ai-common-toolkit:humanizer` skill; the punctuation grep alone does NOT satisfy this)** with `language: <tr|en>`, `tone: technical-explanatory`, `stripFancyPunctuation: true`. Diacritics preserved (Locked 8).
|
|
131
|
-
3. Punctuation gate: `grep -P
|
|
131
|
+
3. Punctuation gate: `node $HOME/.claude/scripts/validate-complaint-doc.mjs <draft>` reports no banned-punctuation error. It checks the policy in Node, so the same result holds on macOS, Linux and Windows; `grep -P` is absent from BSD grep and would never run there.
|
|
132
132
|
4. Show the draft path + size to the user. Set `phase: "awaiting_output_decision"`.
|
|
133
133
|
|
|
134
134
|
### Phase 4.5 - Output destination picker
|
|
@@ -158,7 +158,7 @@ Follow-ups: Confluence -> ask parent page (Other, LRU recents from `prefs.projec
|
|
|
158
158
|
|
|
159
159
|
Then print the summary in `outputLanguage`: complaint count, verdict counts (`X client / Y bff / Z core / W insufficient-evidence`), degraded services, output paths/URLs. When any `core` verdict exists, add: <localized: "N complaint(s) route to the core team - see the routing section before forwarding.">
|
|
160
160
|
|
|
161
|
-
**Stop. Do not chain into
|
|
161
|
+
**Stop. Do not chain into a dev run. Do not open a worktree. Do not create a branch.** Set `phase: "done"`.
|
|
162
162
|
|
|
163
163
|
### Resume contract
|
|
164
164
|
|
|
@@ -178,6 +178,7 @@ Then print the summary in `outputLanguage`: complaint count, verdict counts (`X
|
|
|
178
178
|
| `$HOME/.claude/multi-agent-refs/keychain.md` | Phase 1 non-critical credential handling |
|
|
179
179
|
| `$HOME/.claude/multi-agent-refs/complaint-analysis-template.md` | Phase 4 report template (8 sections) |
|
|
180
180
|
| `ai-common-toolkit:humanizer` | Phase 4 tone pass |
|
|
181
|
+
| `ai-analyst-toolkit:signal-community` | Phase 1 corroboration, optional. Advisory only: a matching report outside the company shows the complaint is not one user's device, and lands in the risks section with its link, never as a root cause. |
|
|
181
182
|
| `$HOME/.claude/scripts/validate-complaint-doc.mjs` | Phase 5 pre-dispatch gate |
|
|
182
183
|
| `$HOME/.claude/multi-agent-refs/channels/confluence.md`, `channels/jira.md`, `~/.claude/lib/md2confluence-v3.py` | Phase 5 dispatch |
|
|
183
184
|
| `$HOME/.claude/schemas/complaint-analysis-spec.schema.json` | State contract |
|
|
@@ -1,289 +1,17 @@
|
|
|
1
1
|
---
|
|
2
|
-
description: "
|
|
3
|
-
description-tr: "
|
|
2
|
+
description: "Removed in v16.0.0. Depth is a question the run asks, not a command name: start /multi-agent and answer Short. Invoke only to see that redirect."
|
|
3
|
+
description-tr: "v16.0.0'da kaldırıldı. Pipeline derinliği artık komut adı değil: /multi-agent çalıştırıp derinlik sorusunda Kısa seçin."
|
|
4
4
|
---
|
|
5
5
|
|
|
6
|
-
# multi-agent dev -
|
|
6
|
+
# multi-agent dev - Removed in v16.0.0
|
|
7
7
|
|
|
8
|
-
|
|
8
|
+
Until v16.0.0, `--dev` and the four `dev-*` commands existed only to express one boolean: skip Analysis and Planning. That boolean is now a question, asked once at Phase 0 Step 7.5, where the recommendation can come from the detected `taskType`.
|
|
9
9
|
|
|
10
|
-
> **Language (read FIRST)**: Two language axes, independent:
|
|
11
|
-
>
|
|
12
|
-
> 1. **`prefs.global.promptLanguage` is always `en`** - instructions to the LLM (this file, every skill spec, every slash-command body, agent system prompts) stay in English. Do not translate spec docs.
|
|
13
|
-
> 2. **`prefs.global.outputLanguage` is user-selectable** - applies to (a) every conversational line the agent writes back, and (b) every user-facing external payload: PR description body, Jira comment, Confluence body, Wiki body.
|
|
14
|
-
>
|
|
15
|
-
> Regardless of `outputLanguage`, these stay English (interop / convention):
|
|
16
|
-
> - `AskUserQuestion` `label` + `header` only (UI contract) - the `question` and each option's `description` follow `outputLanguage`; a picker whose question is English on a Turkish run is a bug, not the contract
|
|
17
|
-
> - Commit message subject/body (git convention)
|
|
18
|
-
> - Branch names (`feature/`, `bugfix/`, ...)
|
|
19
|
-
> - Code identifiers, file paths, log lines
|
|
20
|
-
> - PR title prefix (`feat:`, `fix:`, ...) - only the body honors `outputLanguage`
|
|
21
|
-
>
|
|
22
|
-
> Full contract: `$HOME/.claude/multi-agent-refs/rules.md` "Language Application".
|
|
23
|
-
|
|
24
|
-
A 6-phase fast pipeline: Init → Dev → Review → Test → Commit → Report. Development runs on the Opus model (top intelligence tier).
|
|
25
|
-
|
|
26
|
-
## When to use it
|
|
27
|
-
- Small changes, bug fixes, quick features
|
|
28
|
-
- Simple tasks that don't need the full 8-phase pipeline
|
|
29
|
-
- Work that is already scoped, so analysis and planning buy nothing - the result is still reviewed
|
|
30
|
-
|
|
31
|
-
## Pipeline
|
|
32
|
-
|
|
33
|
-
```
|
|
34
|
-
Phase 0: Init → project detection, worktree, branch, identity (full picker)
|
|
35
|
-
Phase 3: Dev → direct development on Opus (TDD optional)
|
|
36
|
-
Phase 4: Review → deterministic gates + parallel review + triage
|
|
37
|
-
Phase 5: Test → User Test (interactive simulator)
|
|
38
|
-
Phase 6: Commit → commit + push + PR
|
|
39
|
-
Phase 7: Report → channels (Jira / Confluence / PR / Wiki)
|
|
40
|
-
```
|
|
41
|
-
|
|
42
|
-
`--dev` skips Phase 1 (Analysis) and Phase 2 (Planning + Approval Gate). Phase 3 Dev runs on **Opus**. Phase 0 picker, Phase 4 review, Phase 5 test, Phase 6 commit, and Phase 7 channels are identical to the full pipeline.
|
|
43
|
-
|
|
44
|
-
## What `--dev` does NOT skip (required)
|
|
45
|
-
|
|
46
|
-
`--dev` skips the two phases that front-load a task (Phase 1 Analysis, Phase 2 Planning + Approval) and runs Phase 3 Dev on Opus. Phase 0 (Init), Phase 4 (Review), Phase 5 (Test), Phase 6 (Commit), and Phase 7 (Report + channels) are byte-for-byte the same as the full pipeline. User-facing prompts **always run** - they exist for audit, not speed.
|
|
47
|
-
|
|
48
|
-
**Review is not a fast-mode casualty.** Skipping analysis means the run has less context, not that the output deserves less scrutiny: Phase 4 runs its deterministic gates, dispatches the parallel reviewers and triages their findings exactly as the full pipeline does. Accepted blocking findings loop back to Phase 3, capped at 3 iterations. Because Phase 1 never ran, Phase 4 substitutes a deterministic stack classifier for `detectedStack` and records every Phase-1-dependent step as `not-applicable`, so a step that could not run is distinguishable from a step that passed.
|
|
49
|
-
|
|
50
|
-
### Required user prompts (Phase 0)
|
|
51
|
-
|
|
52
|
-
The agent CANNOT make these Phase 0 decisions automatically; it suggests, then waits for confirmation:
|
|
53
|
-
|
|
54
|
-
1. Issue / Jira ID confirm
|
|
55
|
-
2. **Account picker - provider-aware**: First detect the git remote provider via `git remote get-url origin`, then list provider-appropriate accounts. Do not default to `gh auth` when the remote is non-GitHub.
|
|
56
|
-
- `bitbucket.*` → Bitbucket account (parse user from remote URL: `https://USER:TOKEN@host/...`; if multiple stored tokens in OS keychain, pick from those)
|
|
57
|
-
- `github.com` → `gh auth status` accounts
|
|
58
|
-
- `gitlab.*` → GitLab account (similar URL parse / keychain)
|
|
59
|
-
- Unknown host → list all stored credentials and ask
|
|
60
|
-
- Show the chosen provider in the picker label (e.g. "Bitbucket account?" not "GitHub account?")
|
|
61
|
-
3. Project root (state inheritance is FORBIDDEN - picker runs every run)
|
|
62
|
-
4. dev-context picker (extra repos / submodules - read `.gitmodules` and suggest)
|
|
63
|
-
5. **Remote reachability gate (run before branch picker)**: Test remote with
|
|
64
|
-
`git ls-remote --heads origin <baseBranch>` (5s timeout), capturing **stderr**.
|
|
65
|
-
|
|
66
|
-
**Classify the failure before naming a cause.** A non-zero exit is not evidence
|
|
67
|
-
of a network problem, and telling the user to enable a VPN for an auth failure
|
|
68
|
-
sends them to fix something that was never broken. Match stderr, in this order:
|
|
69
|
-
|
|
70
|
-
| stderr contains | Cause | What to offer |
|
|
71
|
-
|---|---|---|
|
|
72
|
-
| `could not read Password`, `Authentication failed`, `terminal prompts disabled`, `Invalid username or password`, `403` | **credential**, not network | the remote URL usually embeds a username with no credential-helper entry. Offer: store the PAT in the credential helper (`git credential approve`, or re-run `/multi-agent:setup`), or switch the remote to SSH. A VPN cannot fix this. |
|
|
73
|
-
| `Could not resolve host`, `Operation timed out`, `Connection refused`, `Network is unreachable`, or the 5s timeout fired with no output | **network** | the VPN/DNS prompt below |
|
|
74
|
-
| `Repository not found`, `does not appear to be a git repository`, `404` | **wrong remote** | show `git remote -v` and ask which remote is correct |
|
|
75
|
-
| anything else | **unknown** | print the stderr line verbatim and ask; never assert a cause you did not observe |
|
|
76
|
-
|
|
77
|
-
Outcomes:
|
|
78
|
-
- **Reachable** → proceed to base branch picker, fetch latest, create worktree
|
|
79
|
-
from `origin/<baseBranch>`
|
|
80
|
-
- **Credential / wrong remote** → STOP with the matching remedy above. Do **not**
|
|
81
|
-
offer "continue from the local ref": the base being stale is unrelated to the
|
|
82
|
-
actual failure, so accepting staleness here trades correctness for nothing.
|
|
83
|
-
- **Network** → STOP and ask:
|
|
84
|
-
```
|
|
85
|
-
[N/Total] Remote gate
|
|
86
|
-
Observed: <the stderr line, verbatim>
|
|
87
|
-
Classified: <host> unreachable (network).
|
|
88
|
-
Options:
|
|
89
|
-
1. Enable VPN and retry (recommended - branch will be fresh)
|
|
90
|
-
2. Continue from local origin/<baseBranch> ref (may be stale)
|
|
91
|
-
3. Cancel
|
|
92
|
-
Confirm? [1/2/3]
|
|
93
|
-
```
|
|
94
|
-
- User picks `2` → log warning + record `"baseFetchStatus": "cached-stale"` in
|
|
95
|
-
`agent-state.json`, proceed from local ref. Phase 6 push needs network anyway,
|
|
96
|
-
so re-prompt there if still unreachable.
|
|
97
|
-
- Option `1` (fetch succeeded) records `"fresh"`; a local-branch base records
|
|
98
|
-
`"local-branch"`; option `3` records `"aborted"`. The field name and this
|
|
99
|
-
four-value vocabulary are what `phase0-exit-gate.mjs` reads, so a run that writes
|
|
100
|
-
anything else cannot close Phase 0.
|
|
101
|
-
|
|
102
|
-
Always show what was **observed** next to what was **classified**. The previous
|
|
103
|
-
wording asserted `Detected: <host> unreachable (VPN/DNS)` for every failure mode,
|
|
104
|
-
including a plain missing-credential error that returns in under a second.
|
|
105
|
-
6. Base branch picker (show recents, take an explicit pick)
|
|
106
|
-
7. Branch name (suggest, allow edit)
|
|
107
|
-
8. Maturity flags acknowledge (read from the issue body's Progress table)
|
|
108
|
-
|
|
109
|
-
### STOP-AND-CONFIRM format
|
|
110
|
-
|
|
111
|
-
For each decision:
|
|
112
|
-
|
|
113
|
-
```
|
|
114
|
-
[N/Total] <Decision name>
|
|
115
|
-
Detected: <suggestion> [+ reason]
|
|
116
|
-
Recent: <last N values used, if any>
|
|
117
|
-
Alternatives: <other options, if any>
|
|
118
|
-
|
|
119
|
-
Confirm? [y/<number>/other/cancel]
|
|
120
|
-
→ <waiting for user input - agent stays SILENT>
|
|
121
|
-
```
|
|
122
|
-
|
|
123
|
-
**Rules:**
|
|
124
|
-
- Confirmation is required even when there is only one option
|
|
125
|
-
- Empty Enter ≠ confirmation; require an explicit answer (`y`, `yes`, `1`, ...)
|
|
126
|
-
- State inheritance is FORBIDDEN - the previous `agent-state.json` of the same task ID is not a template
|
|
127
|
-
- Cancel always halts
|
|
128
|
-
- One prompt per decision - bundling is FORBIDDEN
|
|
129
|
-
|
|
130
|
-
### Required commit / report prompts
|
|
131
|
-
|
|
132
|
-
- Phase 6 → "Check out locally to test? 1 = Yes (WIP) / 2 = No (continue)" (default 2)
|
|
133
|
-
- Phase 6 → When pushing directly to a shared branch (`iteration/develop`, `develop`, `main`), a final confirm via a native `AskUserQuestion` picker (never a typed y/n): `question` "This pushes to iteration/develop with no CI gate. Push anyway?" (rendered in `outputLanguage`); `header` "Push" (English, <=12 chars); `options`: `{ label: "Push", description: "Push directly to the shared branch with no CI gate" }`, `{ label: "Cancel", description: "Do not push; stop here" }`. Anything other than **Push** cancels.
|
|
134
|
-
- Phase 7 → Channels prompt: Jira comment / Confluence / PR description / (optional) Wiki - multi-select
|
|
135
|
-
|
|
136
|
-
### Phase 6: PR creation contract
|
|
137
|
-
|
|
138
|
-
When opening a PR, the agent must:
|
|
139
|
-
|
|
140
|
-
1. **Always add default reviewers**: GET the host's default-reviewers endpoint (Bitbucket Server: `/rest/default-reviewers/1.0/projects/{projectKey}/repos/{repoSlug}/reviewers?sourceRepoId=...&targetRepoId=...&sourceRefId=...&targetRefId=...`; GitHub: `CODEOWNERS` / branch protection required reviewers). Exclude the author from the list. Attach via the create-PR payload's `reviewers` array (Bitbucket Server) or via the `review_requests` API (GitHub) at creation time.
|
|
141
|
-
2. **Never wipe reviewers on update**: Bitbucket Server `PUT /pull-requests/{id}` REPLACES the entire resource. If the payload omits `reviewers`, the field is cleared. Two safe patterns:
|
|
142
|
-
- **Preferred**: do not PUT-update; pre-render the final description so the create-PR call is one-shot.
|
|
143
|
-
- **If you must PUT**: re-fetch the current PR, copy the existing `reviewers` array into the PUT payload along with `version` + the field you're changing.
|
|
144
|
-
- **Or use sub-resource endpoints**: `POST /pull-requests/{id}/participants` (Bitbucket) adds individual reviewers without touching the rest.
|
|
145
|
-
3. **Verify after every PR write**: re-fetch the PR and assert `len(reviewers) >= len(defaultReviewers) - 1` (minus author). If it dropped, re-add via the participants endpoint immediately.
|
|
146
|
-
4. **PR description language**: render in `prefs.global.outputLanguage`. Required sections - Summary, Changed Files, Behavior / Flow, Test Plan, Build Verification (or their localized equivalents). Match the Jira comment structure for cross-channel consistency.
|
|
147
|
-
|
|
148
|
-
### Build "pre-existing" claim rule
|
|
149
|
-
|
|
150
|
-
Before reporting a build failure as "pre-existing, not my problem", reproduce it on the baseline:
|
|
151
|
-
|
|
152
|
-
1. `git stash push -u -m "wip-baseline-check"` (stash local changes)
|
|
153
|
-
2. `git checkout <baseline-sha>` (one commit before any of yours)
|
|
154
|
-
3. Run the same build command - does the failure repeat?
|
|
155
|
-
4. If it repeats: pre-existing is confirmed, report it as such
|
|
156
|
-
5. If it does not repeat: the bug is YOURS - return and fix it
|
|
157
|
-
6. `git checkout -` and `git stash pop`
|
|
158
|
-
|
|
159
|
-
No "pre-existing" claim is allowed without a baseline reproduce.
|
|
160
|
-
|
|
161
|
-
## Skipped phases (definitive list)
|
|
162
|
-
|
|
163
|
-
- Phase 1 (Analysis → Opus deep-think + explore agents)
|
|
164
|
-
- Phase 2 → Plan Approval Gate including clarification + approval
|
|
165
|
-
|
|
166
|
-
That is the whole list. Phase 0 (Init full picker), Phase 4 (Review), Phase 5 (User Test), Phase 6 (Commit), and Phase 7 (Report + channels) run exactly as in the full pipeline. Phase 3 Dev model is **Opus**.
|
|
167
|
-
|
|
168
|
-
## Steps
|
|
169
|
-
|
|
170
|
-
1. **Parse input** - same multi-agent input formats (GitHub Issue URL, Jira ID, free-text)
|
|
171
|
-
|
|
172
|
-
2. **Phase 0: Init** - same as the full pipeline (full picker: account, project, dev-context, base branch, branch name, maturity check, identity, worktree, state)
|
|
173
|
-
- Add `"mode": "dev"` to `agent-state.json`
|
|
174
|
-
|
|
175
|
-
3. **Phase 3: Dev** - on the **Opus model**, with `taskType`-based dispatch:
|
|
176
|
-
|
|
177
|
-
**3a. If `state.taskType === "component"` (figma URL detected at Phase 0 Step 7):**
|
|
178
|
-
- Run the **full figma 17-substep orchestrator** - the same one the full pipeline uses. `--dev` does NOT shrink this. Sub-phases:
|
|
179
|
-
```
|
|
180
|
-
3.0-Init → 3.1-Gather → 3.2A-TestingIDs → 3.2B-Localization → 3.2C-Accessibility → 3.2D-Analytics
|
|
181
|
-
→ 3.3-TokenMapping → 3.4A-Configuration → 3.4B-View → 3.4C-Docs → 3.4D-Preview → 3.4E-Modifiers
|
|
182
|
-
→ 3.5A-Structural (ViewInspector) → 3.5B-Snapshot → 3.5C-Unit
|
|
183
|
-
→ 3.6-CodeConnect → 3.7-Wiki → 3.8-Cleanup
|
|
184
|
-
```
|
|
185
|
-
- **Tests are mandatory** (3.5A + 3.5B + 3.5C). The "tests optional in dev mode" exception below applies to non-component tasks only.
|
|
186
|
-
- Code Connect publish (3.6) is mandatory if `figmaConfig.codeConnect.enabled` (default true).
|
|
187
|
-
- Wiki page generation (3.7) is mandatory if `figmaConfig.wiki.enabled`.
|
|
188
|
-
|
|
189
|
-
**3b. Else (free-text task, bug fix, refactor, generic feature):**
|
|
190
|
-
- Brief analysis (no explore agent - read files directly)
|
|
191
|
-
- Write code + verify build
|
|
192
|
-
- Write tests if needed (not mandatory in dev mode for non-component work)
|
|
193
|
-
|
|
194
|
-
The dispatch is keyed on `agent-state.taskType`, written by Phase 0 Step 7 (`feedback_phase0_tasktype.md`). If the issue body contains a Figma URL, the type is `"component"` and 3a applies - no exception, even in `--dev`.
|
|
195
|
-
|
|
196
|
-
4. **Phase 4: Review** - same as the full pipeline, per `$HOME/.claude/multi-agent-refs/phases/phase-4-review.md`: deterministic gates (build / lint / test / secrets / evidence), the parallel reviewer set (CLI-aware), and triage. Accepted blocking findings return to step 3 for rework, capped at 3 iterations per `operations.md`. Phase-1-dependent steps (1.8 Figma evidence, 2.8 visual conformance) are recorded as `not-applicable (no Phase 1 evidence)` rather than skipped silently.
|
|
197
|
-
|
|
198
|
-
5. **Phase 5: User Test** - same as the full pipeline (interactive simulator / local-test prompt)
|
|
199
|
-
|
|
200
|
-
6. **Phase 6: Commit** - same as the full pipeline (commit + push + PR + shared-branch final confirm). Phase 6 refuses to commit while an accepted blocking finding is unresolved.
|
|
201
|
-
|
|
202
|
-
7. **Phase 7: Report** - same as the full pipeline (channels prompt: Jira / Confluence / PR description / Wiki)
|
|
203
|
-
|
|
204
|
-
## Differences vs full pipeline
|
|
205
|
-
|
|
206
|
-
| Aspect | Full (`/multi-agent`) | Dev (`/multi-agent:dev`) |
|
|
207
|
-
|---------|---------------------|------------------------|
|
|
208
|
-
| Phases | 8 (0-7) | 6 (0, 3, 4, 5, 6, 7) - 1, 2 skipped |
|
|
209
|
-
| Phase 3 Dev model | Sonnet | **Opus** |
|
|
210
|
-
| Phase 1 Analysis (explore agents) | ✅ | ❌ |
|
|
211
|
-
| Phase 2 Planning + Approval Gate | ✅ | ❌ |
|
|
212
|
-
| Phase 4 Review (parallel + triage) | ✅ | ✅ (same) |
|
|
213
|
-
| Phase 4 criteria source | Phase 1 `detectedStack` | deterministic stack classifier |
|
|
214
|
-
| Phase 0 picker (account, repos, branch, ...) | ✅ | ✅ (same) |
|
|
215
|
-
| Phase 5 User Test | ✅ | ✅ (same) |
|
|
216
|
-
| Phase 7 channels (Jira / Confluence / PR / Wiki) | ✅ | ✅ (same) |
|
|
217
|
-
| Duration | ~10-15 min | ~7-10 min |
|
|
218
|
-
## Intake warnings (`--dev` family)
|
|
219
|
-
|
|
220
|
-
Two checks belong at the top of every `--dev` run and are specified once in `$HOME/.claude/multi-agent-refs/phases/modes.md` "Intake warnings shared by the whole `--dev` family": an analysis document supplied to a mode that skips Analysis and Planning, and a branch that already carries the work (which wants `/multi-agent:resume-local`, not a second Dev pass). Read that section rather than reasoning about it from scratch.
|
|
221
|
-
|
|
222
|
-
## Required: outward-facing payload contracts
|
|
223
|
-
|
|
224
|
-
Before writing anything outward-facing - PR body, Jira comment, Confluence page, closing report - load `$HOME/.claude/multi-agent-refs/payload-contracts.md`. It names the canonical section set for each payload, the markup dialect per surface (PR body is Markdown, Jira is wiki markup - mixing them is a defect), and the token/duration numbers the closing report must carry. Improvising a payload shape from memory is the most common failure of the short modes.
|
|
225
|
-
|
|
226
|
-
## Required: Phase Tracker Contract
|
|
227
|
-
|
|
228
|
-
**The phase tracker is mandatory** - the agent cannot skip it. Full spec: [`$HOME/.claude/multi-agent-refs/tracker-contract.md`]($HOME/.claude/multi-agent-refs/tracker-contract.md).
|
|
229
|
-
|
|
230
|
-
Two channels run in parallel at every phase boundary:
|
|
231
|
-
|
|
232
|
-
1. **State channel** (every CLI, identical): `phase-tracker.sh` writes to `tracker-state.json`. Drives `:resume`, `:log`, `:status`.
|
|
233
|
-
2. **Visual channel** (CLI-specific): native widget on Claude Code, ANSI render on every other CLI. Without it the user sees no phase progress.
|
|
234
|
-
|
|
235
|
-
```bash
|
|
236
|
-
# Phase 0, very first shell call (every CLI):
|
|
237
|
-
bash $HOME/.claude/scripts/phase-tracker.sh init "$TASK_ID"
|
|
238
|
-
for p in "0:Init" "3:Dev" "4:Review" "5:Test" "6:Commit" "7:Report"; do
|
|
239
|
-
bash $HOME/.claude/scripts/phase-tracker.sh add "${p%%:*}" "${p#*:}"
|
|
240
|
-
done
|
|
241
|
-
bash $HOME/.claude/scripts/phase-tracker.sh update 0 in_progress
|
|
242
|
-
|
|
243
|
-
# Every phase boundary (every CLI):
|
|
244
|
-
bash $HOME/.claude/scripts/phase-tracker.sh update <N> in_progress|completed|failed|skipped
|
|
245
|
-
|
|
246
|
-
# After every LLM call (every CLI):
|
|
247
|
-
bash $HOME/.claude/scripts/phase-tracker.sh tokens <N> <in> <out> [cached]
|
|
248
10
|
```
|
|
249
|
-
|
|
250
|
-
|
|
251
|
-
|
|
252
|
-
In Claude Code the agent MUST also drive the native TaskList widget so the user sees a sticky phase tile stack - this is the only progress signal Claude Code surfaces. Skipping these calls is the #1 source of "I don't see any phases" complaints.
|
|
253
|
-
|
|
254
|
-
**TaskCreate ordering (strict)**: All TaskCreate calls fire in strict phase-number order BEFORE any TaskUpdate is applied. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks (e.g. `1 ✓ · 2 ✓ · 4 ✓ · 0 ▶ · 3 ☐`) even when the underlying state is correct. Pre-marking phases as completed/skipped before Phase 0 starts is FORBIDDEN - register the tile in order, then flip status via TaskUpdate when the phase actually short-circuits. Full contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
|
|
255
|
-
|
|
256
|
-
```text
|
|
257
|
-
# Phase 0 startup - register one tile per phase (0..N), capture the taskId, persist it:
|
|
258
|
-
for each phase in 0:Init, 3:Dev, 4:Review, 5:Test, 6:Commit, 7:Report:
|
|
259
|
-
TaskCreate({ subject: "Phase <N>: <Name>", activeForm: "<doing-form>" })
|
|
260
|
-
-> returns taskId
|
|
261
|
-
bash $HOME/.claude/scripts/phase-tracker.sh meta <N> tasklist_id "<taskId>"
|
|
262
|
-
|
|
263
|
-
# Phase entry - flip the tile to in_progress alongside the state update:
|
|
264
|
-
TaskUpdate({ taskId: <saved>, status: "in_progress" })
|
|
265
|
-
bash $HOME/.claude/scripts/phase-tracker.sh update <N> in_progress
|
|
266
|
-
|
|
267
|
-
# Active sub-step inside a phase - update activeForm so the spinner header reflects what's happening now:
|
|
268
|
-
TaskUpdate({ taskId: <saved>, activeForm: "Editing TopBarView.swift" })
|
|
269
|
-
|
|
270
|
-
# Phase exit - flip to completed/failed/skipped on both channels:
|
|
271
|
-
TaskUpdate({ taskId: <saved>, status: "completed" })
|
|
272
|
-
bash $HOME/.claude/scripts/phase-tracker.sh update <N> completed
|
|
11
|
+
/multi-agent:dev has been removed. Run /multi-agent and choose Short at the
|
|
12
|
+
depth question - same pipeline, one picker step earlier.
|
|
273
13
|
```
|
|
274
14
|
|
|
275
|
-
|
|
276
|
-
|
|
277
|
-
#### TaskCreate ordering (strict)
|
|
278
|
-
|
|
279
|
-
**All TaskCreate calls fire in strict phase-number order BEFORE any TaskUpdate is applied.** For `--dev` that means: Phase 0 → Phase 3 → Phase 4 → Phase 5 → Phase 6 → Phase 7. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks. Full ordering contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
|
|
280
|
-
|
|
281
|
-
### Visual channel - Copilot CLI / plain shell
|
|
282
|
-
|
|
283
|
-
These CLIs have no TaskList widget. After every state change the agent calls render, which prints a bordered ANSI card as the last tool result so the user sees an updated phase table:
|
|
284
|
-
|
|
285
|
-
```bash
|
|
286
|
-
bash $HOME/.claude/scripts/phase-tracker.sh render
|
|
287
|
-
```
|
|
15
|
+
Print the line above, then continue at the named entry with the same `$ARGUMENTS`. Do not run any phase from this file: it carries no tracker contract and no phase set.
|
|
288
16
|
|
|
289
|
-
|
|
17
|
+
This stub exists so the old name fails loudly and usefully for one minor release instead of silently doing nothing. It is deleted in the next one.
|
|
@@ -1,135 +1,23 @@
|
|
|
1
1
|
---
|
|
2
|
-
description: "
|
|
3
|
-
description-tr: "
|
|
2
|
+
description: "Removed in v16.0.0 and not replaced: an unsupervised run is always the full pipeline now. Invoke only to be pointed at /multi-agent:autopilot or the interactive picker."
|
|
3
|
+
description-tr: "v16.0.0'da kaldırıldı, birebir karşılığı yok: gözetimsiz koşular artık her zaman tam pipeline. /multi-agent:autopilot veya /multi-agent'a yönlendirir."
|
|
4
4
|
---
|
|
5
5
|
|
|
6
|
-
# multi-agent dev
|
|
6
|
+
# multi-agent dev-autopilot - Removed in v16.0.0
|
|
7
7
|
|
|
8
|
-
|
|
8
|
+
This one has no equivalent, and the removal is deliberate rather than a rename.
|
|
9
9
|
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
Dev mode + Autopilot combined: a 4-phase pipeline with no confirmations, end-to-end autonomous.
|
|
13
|
-
|
|
14
|
-
## When to use it
|
|
15
|
-
- Small, trusted, urgent changes
|
|
16
|
-
- Bug fixes, copy updates, config tweaks
|
|
17
|
-
- "Just do it, commit, open the PR" scenarios
|
|
18
|
-
|
|
19
|
-
## Pipeline
|
|
10
|
+
Depth is a question now, and autopilot may not ask questions. That leaves one choice to make on the user's behalf, and Full is the honest default: skipping analysis and planning in a run nobody is watching is where a wrong shortcut costs the most, because there is no one there to notice what it lost.
|
|
20
11
|
|
|
21
12
|
```
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
Phase 7: Report → short terminal summary
|
|
13
|
+
/multi-agent:dev-autopilot has been removed. "Fast plus unattended" no longer
|
|
14
|
+
exists as a combination.
|
|
15
|
+
1. /multi-agent:autopilot - unattended, full pipeline (slower, more tokens)
|
|
16
|
+
2. /multi-agent - short pipeline, but asks the depth question
|
|
27
17
|
```
|
|
28
18
|
|
|
29
|
-
|
|
30
|
-
**Skipped confirmations**: Plan Approval Gate (clarification + approval loop), test prompt, commit prompt, PR prompt.
|
|
31
|
-
|
|
32
|
-
> If you want to talk through the plan, do not use this mode - pick `/multi-agent "task"` (normal) or any mode without `autopilot`. The Dev+Autopilot contract is: "ask nothing."
|
|
33
|
-
|
|
34
|
-
**Review runs here too, and asking nothing does not mean accepting anything.** Autopilot consumes the *confirmation* prompts, not the *quality* gates: Phase 4 reviews the diff and auto-fixes accepted blocking findings without asking. The contract holds because no question is put to the user - the run only halts when it cannot converge, which is the rework-storm trigger in `$HOME/.claude/multi-agent-refs/features/autopilot-circuit-breaker.md` (`maxReworkCycles = 3`).
|
|
35
|
-
|
|
36
|
-
## Steps
|
|
37
|
-
|
|
38
|
-
1. **Parse input** - standard multi-agent input formats (Issue URL, Jira ID, free text)
|
|
39
|
-
2. **Phase 0: Init** - set `"mode": "dev", "autopilot": true` in `agent-state.json`
|
|
40
|
-
3. **Phase 3: Dev** - write code directly on `claude-opus-5` and verify the build
|
|
41
|
-
4. **Phase 4: Review** - gates, parallel reviewers, triage. Accepted blocking findings are fixed automatically (back to step 3), no prompt
|
|
42
|
-
5. **Phase 6: Commit** - auto commit + push + PR
|
|
43
|
-
6. **Phase 7: Report** - terminal summary
|
|
44
|
-
|
|
45
|
-
## Safety
|
|
46
|
-
|
|
47
|
-
- Build failure → 3 retries; if it still fails → **pause** (ask the user)
|
|
48
|
-
- Review rework → 3 cycles; if blocking findings survive → **pause** (circuit breaker, no auto-commit)
|
|
49
|
-
- Kill / Purge → always asks for confirmation (destructive)
|
|
50
|
-
|
|
51
|
-
## Differences
|
|
52
|
-
|
|
53
|
-
| | Full | Dev | Autopilot | **Dev+Autopilot** |
|
|
54
|
-
|--|------|-----|-----------|-------------------|
|
|
55
|
-
| Phases | 8 | 6 | 8 | **5** |
|
|
56
|
-
| Model | Sonnet | Opus | Sonnet | **Opus** |
|
|
57
|
-
| Plan Approval Gate | ✅ (clarification + approval) | ❌ | ❌ | **❌** |
|
|
58
|
-
| Confirmations (test / commit / PR) | Yes | Yes | No | **No** |
|
|
59
|
-
| Review | Parallel + triage (CLI-aware) | Parallel + triage (CLI-aware) | Parallel + triage (CLI-aware) | **Parallel + triage, auto-fix** |
|
|
60
|
-
| Blocking finding | Fix loop, then ask | Fix loop, then ask | Auto-fix, breaker at 3 | **Auto-fix, breaker at 3** |
|
|
61
|
-
| Estimated duration | ~12 min | ~7 min | ~10 min | **~5 min** |
|
|
62
|
-
## Intake warnings (`--dev` family)
|
|
63
|
-
|
|
64
|
-
Two checks belong at the top of every `--dev` run and are specified once in `$HOME/.claude/multi-agent-refs/phases/modes.md` "Intake warnings shared by the whole `--dev` family": an analysis document supplied to a mode that skips Analysis and Planning, and a branch that already carries the work (which wants `/multi-agent:resume-local`, not a second Dev pass). Read that section rather than reasoning about it from scratch.
|
|
65
|
-
|
|
66
|
-
## Required: outward-facing payload contracts
|
|
67
|
-
|
|
68
|
-
Before writing anything outward-facing - PR body, Jira comment, Confluence page, closing report - load `$HOME/.claude/multi-agent-refs/payload-contracts.md`. It names the canonical section set for each payload, the markup dialect per surface (PR body is Markdown, Jira is wiki markup - mixing them is a defect), and the token/duration numbers the closing report must carry. Improvising a payload shape from memory is the most common failure of the short modes.
|
|
69
|
-
|
|
70
|
-
## Required: Phase Tracker Contract
|
|
71
|
-
|
|
72
|
-
**The phase tracker is mandatory** - the agent cannot skip it. Full spec: [`$HOME/.claude/multi-agent-refs/tracker-contract.md`]($HOME/.claude/multi-agent-refs/tracker-contract.md).
|
|
73
|
-
|
|
74
|
-
> **Autopilot mode:** user confirmations are skipped. The tracker is still mandatory - autopilot agent calls cannot skip it; skipping breaks `smoke-tracker-contract.sh`.
|
|
75
|
-
|
|
76
|
-
Two channels run in parallel at every phase boundary:
|
|
77
|
-
|
|
78
|
-
1. **State channel** (every CLI, identical): `phase-tracker.sh` writes to `tracker-state.json`. Drives `:resume`, `:log`, `:status`.
|
|
79
|
-
2. **Visual channel** (CLI-specific): native widget on Claude Code, ANSI render on every other CLI. Without it the user sees no phase progress.
|
|
19
|
+
If a cron job or script calls this name it needs editing, and the two options trade differently: option 1 stays unattended and will take longer and cost more per run, option 2 stays fast and needs a person.
|
|
80
20
|
|
|
81
|
-
|
|
82
|
-
# Phase 0, very first shell call (every CLI):
|
|
83
|
-
bash $HOME/.claude/scripts/phase-tracker.sh init "$TASK_ID"
|
|
84
|
-
for p in "0:Init" "3:Dev" "4:Review" "6:Commit" "7:Report"; do
|
|
85
|
-
bash $HOME/.claude/scripts/phase-tracker.sh add "${p%%:*}" "${p#*:}"
|
|
86
|
-
done
|
|
87
|
-
bash $HOME/.claude/scripts/phase-tracker.sh update 0 in_progress
|
|
88
|
-
|
|
89
|
-
# Every phase boundary (every CLI):
|
|
90
|
-
bash $HOME/.claude/scripts/phase-tracker.sh update <N> in_progress|completed|failed|skipped
|
|
91
|
-
|
|
92
|
-
# After every LLM call (every CLI):
|
|
93
|
-
bash $HOME/.claude/scripts/phase-tracker.sh tokens <N> <in> <out> [cached]
|
|
94
|
-
```
|
|
95
|
-
|
|
96
|
-
### Visual channel - Claude Code (native TaskList widget, required)
|
|
97
|
-
|
|
98
|
-
In Claude Code the agent MUST also drive the native TaskList widget so the user sees a sticky phase tile stack - this is the only progress signal Claude Code surfaces. Skipping these calls is the #1 source of "I don't see any phases" complaints.
|
|
99
|
-
|
|
100
|
-
**TaskCreate ordering (strict)**: All TaskCreate calls fire in strict phase-number order BEFORE any TaskUpdate is applied. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks (e.g. `1 ✓ · 2 ✓ · 4 ✓ · 0 ▶ · 3 ☐`) even when the underlying state is correct. Pre-marking phases as completed/skipped before Phase 0 starts is FORBIDDEN - register the tile in order, then flip status via TaskUpdate when the phase actually short-circuits. Full contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
|
|
101
|
-
|
|
102
|
-
```text
|
|
103
|
-
# Phase 0 startup - register one tile per phase (0..N), capture the taskId, persist it:
|
|
104
|
-
for each phase in 0:Init, 3:Dev, 4:Review, 6:Commit, 7:Report:
|
|
105
|
-
TaskCreate({ subject: "Phase <N>: <Name>", activeForm: "<doing-form>" })
|
|
106
|
-
-> returns taskId
|
|
107
|
-
bash $HOME/.claude/scripts/phase-tracker.sh meta <N> tasklist_id "<taskId>"
|
|
108
|
-
|
|
109
|
-
# Phase entry - flip the tile to in_progress alongside the state update:
|
|
110
|
-
TaskUpdate({ taskId: <saved>, status: "in_progress" })
|
|
111
|
-
bash $HOME/.claude/scripts/phase-tracker.sh update <N> in_progress
|
|
112
|
-
|
|
113
|
-
# Active sub-step inside a phase - update activeForm so the spinner header reflects what's happening now:
|
|
114
|
-
TaskUpdate({ taskId: <saved>, activeForm: "Editing TopBarView.swift" })
|
|
115
|
-
|
|
116
|
-
# Phase exit - flip to completed/failed/skipped on both channels:
|
|
117
|
-
TaskUpdate({ taskId: <saved>, status: "completed" })
|
|
118
|
-
bash $HOME/.claude/scripts/phase-tracker.sh update <N> completed
|
|
119
|
-
```
|
|
120
|
-
|
|
121
|
-
`--dev autopilot` mode does NOT TaskCreate phases 1/2/5 - those are not part of the `--dev autopilot` phase set (`0:Init 3:Dev 4:Review 6:Commit 7:Report`). Only register tiles for the active set.
|
|
122
|
-
|
|
123
|
-
#### TaskCreate ordering (strict)
|
|
124
|
-
|
|
125
|
-
**All TaskCreate calls fire in strict phase-number order BEFORE any TaskUpdate is applied.** For `--dev autopilot` that means: Phase 0 → Phase 3 → Phase 4 → Phase 6 → Phase 7. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks. Full ordering contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
|
|
126
|
-
|
|
127
|
-
### Visual channel - Copilot CLI / plain shell
|
|
128
|
-
|
|
129
|
-
These CLIs have no TaskList widget. After every state change the agent calls render, which prints a bordered ANSI card as the last tool result so the user sees an updated phase table:
|
|
130
|
-
|
|
131
|
-
```bash
|
|
132
|
-
bash $HOME/.claude/scripts/phase-tracker.sh render
|
|
133
|
-
```
|
|
21
|
+
Print the block above and stop. Do not pick an entry on the user's behalf: the two options differ in what gets skipped and who is watching, and that is their call. Autopilot passed to this stub still stops here - a zero-interaction contract does not authorise choosing a different pipeline than the one that was asked for.
|
|
134
22
|
|
|
135
|
-
|
|
23
|
+
This stub exists so the old name fails loudly and usefully for one minor release instead of silently doing nothing. It is deleted in the next one.
|