@mmerterden/multi-agent-pipeline 17.3.0 → 17.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (54) hide show
  1. package/CHANGELOG.md +203 -0
  2. package/README.md +23 -5
  3. package/README.tr.md +23 -5
  4. package/docs/adr/0013-lsp-code-intelligence.md +102 -0
  5. package/docs/adr/README.md +1 -0
  6. package/docs/token-budget-history.md +1 -1
  7. package/install/templates/copilot-instructions.md +9 -3
  8. package/package.json +1 -1
  9. package/pipeline/agents/code-reviewer.md +35 -1
  10. package/pipeline/commands/multi-agent/analysis/SKILL.md +3 -3
  11. package/pipeline/commands/multi-agent/autopilot/SKILL.md +3 -3
  12. package/pipeline/commands/multi-agent/autopilot-off/SKILL.md +5 -3
  13. package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +1 -1
  14. package/pipeline/commands/multi-agent/local/SKILL.md +17 -6
  15. package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +3 -3
  16. package/pipeline/lib/multi-repo-pipeline.sh +26 -0
  17. package/pipeline/multi-agent-refs/analysis/locked.md +4 -4
  18. package/pipeline/multi-agent-refs/analysis/render.md +2 -1
  19. package/pipeline/multi-agent-refs/channels/pr.md +26 -0
  20. package/pipeline/multi-agent-refs/cross-cli-contract.md +22 -0
  21. package/pipeline/multi-agent-refs/features/base-branch-evidence.md +222 -0
  22. package/pipeline/multi-agent-refs/features/code-graph.md +40 -0
  23. package/pipeline/multi-agent-refs/features/code-intelligence.md +80 -0
  24. package/pipeline/multi-agent-refs/features/design-conformance.md +14 -0
  25. package/pipeline/multi-agent-refs/features/review-file-set.md +132 -0
  26. package/pipeline/multi-agent-refs/phases/modes.md +23 -3
  27. package/pipeline/multi-agent-refs/phases/phase-0-init.md +96 -71
  28. package/pipeline/multi-agent-refs/phases/phase-4-review.md +31 -23
  29. package/pipeline/multi-agent-refs/phases/phase-7-report.md +1 -1
  30. package/pipeline/multi-agent-refs/phases.md +7 -2
  31. package/pipeline/multi-agent-refs/picker-contract.md +37 -5
  32. package/pipeline/multi-agent-refs/tracker-contract.md +25 -14
  33. package/pipeline/schemas/agent-state.schema.json +88 -4
  34. package/pipeline/schemas/prefs.schema.json +22 -0
  35. package/pipeline/schemas/review-file-exclusions.json +137 -0
  36. package/pipeline/schemas/reviewer-output.schema.json +27 -1
  37. package/pipeline/schemas/token-budget.json +2 -2
  38. package/pipeline/scripts/autopilot-runner.mjs +292 -45
  39. package/pipeline/scripts/base-branch-candidates.mjs +599 -0
  40. package/pipeline/scripts/diff-risk-score.mjs +1 -36
  41. package/pipeline/scripts/gc-abandoned.sh +5 -3
  42. package/pipeline/scripts/gen-mode-dispatch.mjs +39 -16
  43. package/pipeline/scripts/git-path.mjs +63 -0
  44. package/pipeline/scripts/glob-match.mjs +62 -0
  45. package/pipeline/scripts/graph-mermaid.mjs +251 -0
  46. package/pipeline/scripts/phase-tracker.sh +39 -2
  47. package/pipeline/scripts/phase0-exit-gate.mjs +128 -0
  48. package/pipeline/scripts/review-file-filter.mjs +180 -0
  49. package/pipeline/scripts/skill-conformance.mjs +1 -31
  50. package/pipeline/scripts/validate-analysis-doc.mjs +53 -0
  51. package/pipeline/scripts/validate-reviewer.mjs +90 -1
  52. package/pipeline/scripts/verify-citations.mjs +428 -0
  53. package/pipeline/skills/.skill-manifest.json +2 -2
  54. package/pipeline/skills/shared/core/multi-agent/SKILL.md +1 -1
@@ -0,0 +1,80 @@
1
+ ## Code Intelligence (multi-agent-toolkit `code_*`)
2
+
3
+ Compiler-grade answers about Swift and Kotlin source, from the language server
4
+ each platform already ships. Eight tools in the companion MCP server
5
+ (`@mmerterden/multi-agent-toolkit-mcp` 3.10.0 and later), all read-only except
6
+ `code_server_reset`.
7
+
8
+ ### What it is for, and what `code-graph.md` already covers
9
+
10
+ These two answer different questions and neither replaces the other.
11
+
12
+ | Question | Use |
13
+ |---|---|
14
+ | Where does this concept live in a repo I do not know | `code-graph.md` - one cheap pass over the whole tree, token-budgeted |
15
+ | What depends on this area, roughly | `graph-affected` |
16
+ | Is THIS exact symbol referenced, and where | `mcp__multi-agent-toolkit__code_references` |
17
+ | What type is this, really | `mcp__multi-agent-toolkit__code_hover` |
18
+ | Does this file compile, without a full build | `mcp__multi-agent-toolkit__code_diagnostics` |
19
+
20
+ `docs/adr/0010-own-code-graph.md` states the trade the graph makes in as many
21
+ words: *"Regex over comment-stripped source is not a parser. Definitions and
22
+ imports survive that trade; call graphs and type resolution do not."* It names
23
+ two consequences - a name declared in two files is dropped rather than fanned
24
+ out, so `affected` under-reports on duplicated names, and only type-like symbols
25
+ are reference targets. Those are exactly the two the language server answers
26
+ exactly. So the graph stays the wide, cheap first pass, and `code_*` is what a
27
+ specific claim is checked against.
28
+
29
+ ### The one thing to know before trusting an answer
30
+
31
+ An empty result is not the same as a negative result, and on a cold repository
32
+ it is the likelier of the two. Measured on Swift 6.3.1 against a two-file
33
+ package: the same `references` query answered 0 at 0.7s and 3 at 5.8s, with
34
+ nothing changed except that the background index had finished. On a 658-file
35
+ package with dependencies the first index had not finished at 70s.
36
+
37
+ Every index-backed result therefore carries `indexReady`, and the tools wait for
38
+ the index rather than answering early. When `indexReady` is false, an empty list
39
+ means "not indexed yet". Never report it as "unused".
40
+
41
+ `code_index_status` answers whether this machine can answer at all, and it
42
+ succeeds even when nothing is installed - it is the tool to run first on an
43
+ unfamiliar machine.
44
+
45
+ ### Root decides what an answer is worth
46
+
47
+ `semantic` and `buildSettingsSource` come back on every result.
48
+
49
+ | Root | definition | references | diagnostics |
50
+ |---|---|---|---|
51
+ | `Package.swift` | yes | yes, after the index settles | yes |
52
+ | `buildServer.json` | yes | yes | yes |
53
+ | `.xcodeproj`, no build server | same-file only | no, `semantic: false` | misleading, marked non-semantic |
54
+
55
+ A modular app is mostly packages, so this matters less than it reads: the iOS
56
+ app measured here has 41 `Package.swift` roots and a single `.swift` file under
57
+ the umbrella project.
58
+
59
+ ### Where a run uses them
60
+
61
+ All three are opt-in and none is on by default.
62
+
63
+ 1. **Phase 4, checking a citation's claim.** `verify-citations.mjs` already
64
+ proves a cited `file:line` exists. `code_hover` at that position proves the
65
+ line holds what the finding says it holds, and `code_references` falsifies a
66
+ "nothing handles this" claim.
67
+ 2. **Phase 1, sharpening an impact estimate.** When `graph-affected` reports on
68
+ a name that the graph itself flagged as ambiguous, `code_references` on the
69
+ one symbol the task names resolves it exactly.
70
+ 3. **Phase 3, fast feedback.** `code_diagnostics` on the file just edited,
71
+ without waiting for `xcodebuild`.
72
+
73
+ ### Kotlin
74
+
75
+ Runs on JetBrains `kotlin-lsp`: Alpha, partially closed source, Android Gradle
76
+ Plugin support experimental, and installed separately
77
+ (`brew tap JetBrains/utils && brew install kotlin-lsp`, JVM 17+). Answers carry
78
+ `confidence: "alpha"`, and the toolkit's own README states that no gate in that
79
+ repository exercises the Kotlin server. Treat Kotlin answers as a lead, not as
80
+ evidence, until that changes.
@@ -1,6 +1,7 @@
1
1
  # Design conformance - the component walk, and why a glance is not a pass
2
2
 
3
3
  <!-- toc -->
4
+ - [0. Why this runs as a gate](#0-why-this-runs-as-a-gate)
4
5
  - [1. Enumerate first, then fill every cell](#1-enumerate-first-then-fill-every-cell)
5
6
  - [2. Measure, never read the token](#2-measure-never-read-the-token)
6
7
  - [3. Adaptive per-component convergence](#3-adaptive-per-component-convergence)
@@ -24,6 +25,19 @@ Consumers: `/multi-agent:design-check` (the runner), Phase 4 review when a UI
24
25
  diff is under review, and `features/visual-evidence.md` when a capture has to
25
26
  prove a fix. Gate: `smoke-design-conformance.sh`.
26
27
 
28
+ ## 0. Why this runs as a gate
29
+
30
+ `design-check` existed as a command for a while with no phase invoking it, so the only
31
+ thing standing between a build and visual drift was the user opening the app and
32
+ looking. On one run that produced 16pt padding where the frame said `Spacing/12`, and a
33
+ full sheet rebuild afterwards.
34
+
35
+ The reason it cannot be advice is structural, not historical: a reviewer reading a diff
36
+ cannot see spacing. Every other Phase 4 check reads text and reasons about text; this
37
+ one is the only thing in the pipeline that compares a rendered result against the
38
+ design it was drawn from. Left optional, it is the check that gets skipped on exactly
39
+ the runs that are in a hurry, which are the runs that produce drift.
40
+
27
41
  ## 1. Enumerate first, then fill every cell
28
42
 
29
43
  Do **not** walk by finding, and do not walk by headline. Build the inventory
@@ -0,0 +1,132 @@
1
+ # The review file set - what was read, and what was not, on the record
2
+
3
+ <!-- toc -->
4
+ - [Why this exists](#why-this-exists)
5
+ - [The set is fixed before the reviewer sees the diff](#the-set-is-fixed-before-the-reviewer-sees-the-diff)
6
+ - [Every exclusion names the pattern that produced it](#every-exclusion-names-the-pattern-that-produced-it)
7
+ - [The reviewer answers for each file](#the-reviewer-answers-for-each-file)
8
+ - [Invocation](#invocation)
9
+ - [What this is not](#what-this-is-not)
10
+ <!-- /toc -->
11
+
12
+ > The denominator for Phase 4's other axis. Loaded on demand by
13
+ > `/multi-agent:review` and by pipeline Phase 4.
14
+
15
+ ## Why this exists
16
+
17
+ Phase 4 had a size cap and no exclusion list. When the diff exceeds the phase
18
+ token allowance the cap truncates the LARGEST files first, so a regenerated
19
+ lockfile or a snapshot dump is not merely wasted budget: it is the thing that
20
+ survives while real code is cut.
21
+
22
+ The second half is worse and quieter. A reviewer that opened one file of ten and
23
+ a reviewer that read all ten and found nothing return the identical
24
+ `{"findings": [], "approved": true}`. Nothing in the pipeline could tell them
25
+ apart, so "no findings" has been carrying two meanings at once.
26
+
27
+ Both halves are one fix: decide what is worth reading BEFORE the cap decides
28
+ what fits, and make the reviewer answer for each file it was given.
29
+
30
+ ## The set is fixed before the reviewer sees the diff
31
+
32
+ The same rule `selectedRules[]` follows, for the same reason. After a model has
33
+ seen the diff, "I did not open that one" and "there was nothing there" become
34
+ the same sentence, and whichever one is cheaper to say is the one that gets
35
+ said. So the set is computed from `git diff --name-only`, written to
36
+ `.pipeline/review-files.json`, and never recomputed inside the round.
37
+
38
+ ```bash
39
+ git -C "$WORKTREE" diff --name-only "$BASE_BRANCH"...HEAD \
40
+ | node $HOME/.claude/scripts/review-file-filter.mjs \
41
+ > "$WORKTREE/.pipeline/review-files.json"
42
+ ```
43
+
44
+ Report shape:
45
+
46
+ ```json
47
+ {
48
+ "reviewed": ["src/App.swift"],
49
+ "excluded": [
50
+ {
51
+ "path": "package-lock.json",
52
+ "reason": "lockfile - resolved by the package manager, not written by hand",
53
+ "pattern": "**/package-lock.json"
54
+ }
55
+ ],
56
+ "total": 2,
57
+ "patternsSource": ".../schemas/review-file-exclusions.json",
58
+ "patternCount": 45
59
+ }
60
+ ```
61
+
62
+ `reviewed` is what goes into the diff cap and into the reviewer prompt, as a
63
+ `${REVIEW_FILES}` block beside `${CRITERIA}` in the shared cache prefix. It has to
64
+ be in the prompt: a reviewer asked to account for a set it was never shown can
65
+ only guess, and `fileCoverage` would then fail on every single dispatch - a gate
66
+ that always fires is a gate that gets switched off. `excluded` goes into the run
67
+ report, never into silence.
68
+
69
+ ## Every exclusion names the pattern that produced it
70
+
71
+ A file that disappears between the diff and the review is indistinguishable from
72
+ a file nobody found anything in - which is the exact confusion this whole
73
+ feature exists to remove, so reintroducing it in the filter would be
74
+ self-defeating. Each excluded row carries both the human reason and the glob
75
+ that matched, so a reviewer, a PR reader or a future maintainer can dispute the
76
+ call rather than discover it.
77
+
78
+ The pattern list is data, in `schemas/review-file-exclusions.json`, and it is
79
+ generic: generated trees, lockfiles, recorded snapshots, vendored source, build
80
+ output, binary assets. No stack, project or company name appears in it. A
81
+ pattern with no reason invalidates the whole list rather than being defaulted -
82
+ the default would be exactly the sentence the caller is supposed to print.
83
+
84
+ **It fails open, on purpose.** An unreadable or malformed pattern file yields
85
+ every file reviewed, the reason on stderr, and exit 2. Failing closed would
86
+ review nothing and report a clean run.
87
+
88
+ ## The reviewer answers for each file
89
+
90
+ `reviewer-output.schema.json` (v1.3.0) carries `fileCoverage[]`: one row per
91
+ path in `reviewed`, `{path, verdict: reviewed|skipped, reason}`. There is
92
+ deliberately no `partial` - a file read in part is read, and what was not
93
+ understood belongs in a finding.
94
+
95
+ `validate-reviewer.mjs --coverage <report>` enforces it, catching the same three
96
+ failures the conformance checklist catches on the rule axis:
97
+
98
+ | Failure | Why it matters |
99
+ |---|---|
100
+ | a file in the set with no row | silently unread, and the empty `findings[]` reads as clean |
101
+ | a row for a path outside the set | an answer about something the reviewer was not given, the same shape as a hallucinated rule ID |
102
+ | `skipped` with no reason | a drop with no cause is indistinguishable from a read |
103
+
104
+ `skipped` is legitimate and expected: a file past the diff cap, a file whose
105
+ content the host truncated. What it may not be is unexplained. "Not relevant" is
106
+ a review decision and belongs in a verdict of `reviewed`, not a skip.
107
+
108
+ An empty `reviewed` set (a diff that is entirely lockfiles) demands no checklist
109
+ at all. Requiring an empty array there would fail honest output, and the
110
+ filter's `excluded[]` is what carries that information onward.
111
+
112
+ ## Invocation
113
+
114
+ ```bash
115
+ node $HOME/.claude/scripts/validate-reviewer.mjs "$REVIEWER_FILE" \
116
+ --criteria "$WORKTREE/.pipeline/criteria-manifest.json" \
117
+ --coverage "$WORKTREE/.pipeline/review-files.json"
118
+ ```
119
+
120
+ Without `--coverage` the field stays optional, so every existing caller keeps
121
+ working unchanged. With it, the checklist is enforced and exit 1 takes the same
122
+ single self-correction rework the rest of the validator gate takes.
123
+
124
+ ## What this is not
125
+
126
+ It is not a relevance filter. Nothing here decides that a file is uninteresting;
127
+ it decides that a file is not human-authored source, which is a mechanical
128
+ question with a mechanical answer. The moment a pattern starts encoding "we
129
+ probably do not care about this directory", the list has become a way to hide
130
+ work, and the near-miss assertions in `smoke-review-file-filter.sh`
131
+ (`CodeGenerator.swift`, `generated-report.md`, `distribution/`, `buildSrc/`) are
132
+ what fail when it does.
@@ -93,7 +93,7 @@ Phase 0: Init -> Phase 3: Dev (self-contained) -> Phase 4: Review -> Phase 5: Te
93
93
  | Phase | Full | Short | Short + `--local` |
94
94
  | ------------------- | ----------------------------------------------- | ---------------------------------------------------------------------------------- | --------------------------------------------- |
95
95
  | Phase 0 (Init) | Full setup | Same - worktree, branch, state, and the depth question itself | Same - no worktree, branch on `$PROJECT_ROOT` |
96
- | Phase 1 (Analysis) | Parallel Explore agents + analysis document | **SKIP** - tile flips to `skipped` at Step 7.5 | **SKIP** |
96
+ | Phase 1 (Analysis) | Parallel Explore agents + analysis document | **SKIP** - no tile is ever drawn for it (registration is deferred to Step 7.5) | **SKIP** |
97
97
  | Phase 2 (Planning) | TaskCreate + architecture review + **Plan Approval Gate** | **SKIP** (no plan means no plan gate) | **SKIP** |
98
98
  | Phase 3 (Dev) | Follows the Phase 2 plan, TDD cycle (Sonnet) | **Self-contained** (Opus): agent scans relevant files, implements with TDD, builds | Same, on the local branch |
99
99
  | Phase 4 (Review) | Parallel review + Fable triage (3 reviewers on every host: Claude Code Fable + Opus + Sonnet, Copilot GPT-5.4 + Opus + Sonnet) | **Same** - gates, parallel review, triage; blocking findings return to Phase 3 (cap 3) | **Same**, on the local branch diff |
@@ -111,7 +111,7 @@ The **Opus** agent receives the task description (from Jira, GitHub issue, or fr
111
111
 
112
112
  No separate task breakdown - the agent handles scope autonomously.
113
113
 
114
- **State tracking**: `agent-state.json` gets `"onlyDevelop": true`. Because the tracker boots at Step -1 and the answer arrives at Step 7.5, Phases 1 and 2 are registered `pending` with everything else and flipped to `skipped` when the answer lands; pre-marking is forbidden and would scramble the tile order. Full contract: `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "Late skip".
114
+ **State tracking**: `agent-state.json` gets `"onlyDevelop": true`. The tracker boots at Step -1 with Phase 0 alone and the rest of the tiles are registered at Step 7.5, once this answer says which phases the run actually has - so a Short run never draws an Analysis tile it will not use. Full contract: `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "Deferred registration".
115
115
 
116
116
  ### Intake warnings for a Short run
117
117
 
@@ -161,7 +161,22 @@ Phase 0: Init -> Phase 1: Analysis -> Phase 2: Planning -> Phase 4: Review -> Ph
161
161
 
162
162
  Local mode skips worktree creation - works directly on a local branch in the project root. Useful for single-task workflows or when worktrees cause issues.
163
163
 
164
- **Activation**: the `:local` entries, or `--local` on the base command:
164
+ **Activation**: answer the Phase 0 Step 5b workspace question, or state it up
165
+ front so the question resolves without being asked.
166
+
167
+ ```
168
+ Bu is nerede kossun? / Where should this task run?
169
+ 1. Worktree .worktrees/{id}/ - your current checkout stays untouched
170
+ 2. Lokal / Local the project root, on a new branch - no Phase 5
171
+ ```
172
+
173
+ Two genuine options, so it meets the two-option floor in `picker-contract.md` on
174
+ its own. **Say what local costs inside the question**: Phase 5 is not in a local
175
+ run's set (the user-test gate checks the change out of a worktree, and there is
176
+ none), and uncommitted work in the project root is in the way of the checkout. A
177
+ user choosing local should learn both before choosing, not after.
178
+
179
+ The ways to state it up front:
165
180
 
166
181
  ```
167
182
  /multi-agent:local "PROJ-12345"
@@ -171,6 +186,11 @@ Local mode skips worktree creation - works directly on a local branch in the p
171
186
  /multi-agent "PROJ-12345" --local
172
187
  ```
173
188
 
189
+ Every autopilot entry resolves this to **worktree** and never asks. That is not a
190
+ skipped question: an unattended run commits and pushes from wherever it stands,
191
+ and doing that in the user's own checkout is what worktrees exist to prevent.
192
+ `:local-autopilot` is the explicit opt-out.
193
+
174
194
  **What changes in local mode:**
175
195
 
176
196
  | Phase | Normal (worktree) | Local |
@@ -1,6 +1,6 @@
1
1
  ### Phase 0: Init
2
2
 
3
- > **TLDR** - 8 sequential steps: load prefs → parse input (Jira/GitHub/free-text) → select project → detect remote + pick base branch → optional design input → branch name → git identity → instruction files → create worktree/local branch + agent-state.json. Every input type (Jira URL, GitHub issue, free-text) flows through the same 8 steps.
3
+ > **TLDR** - every input type flows through the same numbered steps below, in order, and ends at the exit gate.
4
4
 
5
5
  #### Step −1 - Bootstrap the cross-CLI tracker (FIRST thing in every run)
6
6
 
@@ -9,13 +9,17 @@ Before anything else - initialize the visual tracker so the user sees the pipe
9
9
  ```bash
10
10
  TASK_ID="${INPUT_TASK_ID:-pipeline-$(date +%Y%m%d-%H%M%S)}"
11
11
  $HOME/.claude/scripts/phase-tracker.sh init "$TASK_ID"
12
- for p in 0:Init 1:Analysis 2:Planning 3:Dev 4:Review 5:Test 6:Commit 7:Report; do
13
- $HOME/.claude/scripts/phase-tracker.sh add "${p%%:*}" "${p#*:}"
14
- done
12
+ $HOME/.claude/scripts/phase-tracker.sh add 0 Init
15
13
  $HOME/.claude/scripts/phase-tracker.sh tiles
16
14
  $HOME/.claude/scripts/phase-tracker.sh update 0 in_progress
17
15
  ```
18
16
 
17
+ **Phase 0 only.** `/multi-agent` and `:local` do not know their phase set yet -
18
+ depth decides it, and depth is Step 7.5 - so they register the rest there rather
19
+ than drawing eight tiles the user has not chosen. Every other mode registers its
20
+ whole set here, in phase-number order. Contract: `tracker-contract.md`,
21
+ "Deferred registration".
22
+
19
23
  `tiles` prints this host's widget-registration calls: **make them before continuing.** The card alone lands in collapsed tool output, so a run that skips them runs in silence. Contract: `tracker-contract.md`, "The card is not the widget".
20
24
 
21
25
  If `INPUT_TASK_ID` isn't known yet (free-text, project not selected), use a placeholder; rename later via `mv` once parsed in Step 1.
@@ -26,7 +30,7 @@ Every subsequent phase (1-7) MUST call `phase-tracker.sh update <N> in_progress`
26
30
 
27
31
  ##### TaskCreate ordering on Claude Code (strict)
28
32
 
29
- On Claude Code, fire all `TaskCreate` calls in strict phase-number order (0 → 7) BEFORE any `TaskUpdate` - which is the order `tiles` prints them in. Full contract: `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
33
+ On Claude Code, fire every `TaskCreate` in a registration batch in strict phase-number order BEFORE any `TaskUpdate` in that batch - which is the order `tiles` prints them in - and never register a phase whose number is below one already registered. A deferred batch (Step 7.5) therefore appends, it does not interleave. Full contract: `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
30
34
 
31
35
  ---
32
36
 
@@ -261,58 +265,76 @@ Scan `$HOME` (maxdepth 2) for project markers (`.xcodeproj`, `Package.swift`, `b
261
265
 
262
266
  **Test policy (once per project).** Resolve `prefs.projects[{slug}].testPolicy` → `state.testPolicy`; missing → native picker "Should development write tests here?" (`header`: "Tests"): `tdd` (recommended) / `tests-after` / `none`; persist. Autopilot without a record: `tdd`, noted.
263
267
 
268
+ #### Step 2b - Dev context (extra repos)
269
+
270
+ Run `$HOME/.claude/multi-agent-refs/_dev-context.md`: `.gitmodules` submodules
271
+ plus `prefs.projects[{project}].webRepos[]`, editable ones pre-selected,
272
+ read-only siblings listed. It runs on every input type, direct-ID included, and
273
+ an empty submit is a valid answer meaning primary repo only. Here rather than
274
+ later because a base branch is a property of a repo set (`picker-contract.md`,
275
+ "Order: project, then repo, then branch").
276
+
277
+ Persist `state.siblings[]` even when empty: the empty array is the record that
278
+ the step ran. Absent, Phase 4's parity cross-check cannot tell "no siblings"
279
+ from "never asked", and the exit gate below fails.
280
+
264
281
  #### Step 3 - Remote Detection + Branch Selection
265
282
 
266
283
  1. **Check preferences first**: If `prefs.projects[{project}].remoteType` exists, use cached value.
267
284
  2. Read remote: `git -C $PROJECT_ROOT remote get-url origin`
268
285
  3. Detect type: `github.com` → github (`gh` CLI), `{BITBUCKET_HOST}` → bitbucket (API + `keychainMapping.bitbucket_token`), other → generic-git. Save to `prefs.projects[{project}].remoteType`.
269
286
  4. **Skip if baseBranch already set** (from Step 1). Otherwise:
270
- 5. Fetch + list PR-targetable branches:
287
+ 5. Fetch + list PR-targetable branches. **Capture the exit code** - it decides
288
+ whether the list may be called the remote's answer:
271
289
  ```bash
272
- git -C $PROJECT_ROOT fetch origin
290
+ git -C $PROJECT_ROOT fetch origin; FETCH_RC=$?
273
291
  git -C $PROJECT_ROOT branch -r --sort=-committerdate \
274
- | grep -E '(develop|release|main|master)' \
275
- | grep -v -E '(feature/|bugfix/|fix/|hotfix/|chore/)'
292
+ | grep -v -E '(feature/|bugfix/|fix/|hotfix/|chore/)' \
293
+ | grep -E '([0-9]+[._][0-9]+|develop|release|main|master)'
276
294
  ```
277
- 6. Sort: `develop*` first, then `release/*`, then `main`/`master`. Surface through the
278
- **native picker** per `picker-contract.md`, recent branch first and marked
279
- `(Recommended)`:
295
+ **Keep the version alternative.** A branch carrying a version token is a
296
+ candidate whatever it is called; the family words only RANK. Why that
297
+ distinction matters: `features/base-branch-evidence.md`, "A filter is not a
298
+ ranking".
299
+ 5b. With `prefs.global.baseBranchEvidence` on, pipe that list plus the issue's
300
+ version fields and links through `$HOME/.claude/scripts/base-branch-candidates.mjs`,
301
+ which ranks the candidates and records the evidence behind each. It is pure, so
302
+ Step 3 owns every git and network call. Contract - autopilot resolution order and
303
+ the fenced ask-on-the-issue path included:
304
+ `$HOME/.claude/multi-agent-refs/features/base-branch-evidence.md`. Off: rule 6 alone.
305
+ 6. Sort: `develop*` first, then `release/*`, then `main`/`master` (the collector's
306
+ order when it ran). Surface through the **native picker** per `picker-contract.md`,
307
+ recent branch first and marked `(Recommended)`, each row carrying its evidence:
280
308
 
281
309
  ```
282
310
  header: "Base branch"
283
311
  options: origin/develop (Recommended, reused from last run) | origin/main | release/8.4.0 | Other
284
312
  ```
285
- 7. User picks → store as `baseBranch`, and append `{branch, lastUsed, count?}` to
313
+ 7. User picks → store as `baseBranch` with `baseBranchSource: "asked"`, and append `{branch, lastUsed, count?}` to
286
314
  `prefs.global.recentBranches[{projectKey}]` (dedup by `branch`, cap 10) - what the TTL
287
315
  filter below reads. Key is `branch`, not `name`. Never the legacy
288
316
  `projects[{project}].branches`.
289
317
 
290
318
  **MUST: this step is not skippable (BLOCKING).** The only legitimate skip is rule 4
291
- above - `baseBranch` already supplied in the input. Everything else asks. A run once
292
- took a Jira ID and implemented straight onto whatever the local checkout was pointing
293
- at, with neither the project nor the branch picker ever shown; nothing failed, so
294
- nothing surfaced it. `phase0-exit-gate.mjs` now refuses to close Phase 0 unless
295
- `agent-state.json` carries `baseBranch` and `baseFetchStatus`, so a skipped picker is a
296
- gate failure rather than a silent default.
319
+ above, recorded as `baseBranchSource: "input"`. Everything else asks, and
320
+ `phase0-exit-gate.mjs` refuses to close Phase 0 without `baseBranch`,
321
+ `baseFetchStatus` and `baseBranchSource` - a skipped picker fails the gate.
297
322
 
298
- One row is a normal outcome of the rule-5 filter and is **still asked** - see
299
- `picker-contract.md` "A single candidate is still a question".
323
+ One row is a normal outcome of the rule-5 filter and is **still asked**, with a
324
+ real second option - `picker-contract.md`, "Two options or it is not a question".
300
325
 
301
326
  This holds in every mode. A Short run skips the *LLM* phases (Analysis, Planning); it
302
327
  does not skip Phase 0's pickers. Autopilot resolves them without prompting, which still
303
- writes the fields. For the base branch it reads `recentBranches[{projectKey}]` (most
304
- recent inside the TTL, still on the remote) before the rule-6 order, recording which
305
- fired in `state.baseBranchSource`.
328
+ writes the fields; its base-branch resolution order is in `features/base-branch-evidence.md`.
306
329
 
307
330
  **TTL filter for recent branches**:
308
331
 
309
- - `prefs.global.recentBranches[{projectKey}][]` carries `{branch, lastUsed, count?}`. Filter to those whose `lastUsed` is within `settings.branchTtlDays` (default 15).
310
- - Stale entries (>TTL) are pruned in-place during the read - keeps the picker uncluttered without a separate cleanup pass.
332
+ - `prefs.global.recentBranches[{projectKey}][]` carries `{branch, lastUsed, count?}`. Keep those whose `lastUsed` is within `settings.branchTtlDays` (default 15); prune the rest in place during the read.
311
333
  - The filtered "Recent" list precedes the fresh `git branch -r` list; cap at 5 visible recent entries.
312
334
 
313
335
  **Fetch-fail handling** (replaces silent `git fetch origin` failure):
314
336
 
315
- The legacy `git -C $PROJECT_ROOT fetch origin` step (line 110 above) MUST not silently fall back to a stale cached ref. On non-zero exit, surface the **native picker** (per `picker-contract.md`) with 4 options:
337
+ `FETCH_RC` non-zero MUST not silently fall back to a stale cached ref. Surface the **native picker** (per `picker-contract.md`) with 4 options:
316
338
 
317
339
  ```
318
340
  question: "git fetch failed for {project} - the base ref may be stale. How should I proceed?"
@@ -325,19 +347,13 @@ options:
325
347
  Abort no worktree, no branch, no state file
326
348
  ```
327
349
 
328
- **Say which is which.** The question must name the corporate host when the remote points
329
- at one - a `{BITBUCKET_HOST}` remote failing to resolve is almost always the VPN, and
330
- telling the user that is the difference between a five-second fix and a run built on a
331
- month-old ref. "Connect VPN and retry" re-runs the fetch and re-enters this picker if it
332
- fails again; it is a real retry, not a label.
350
+ **Say which is which.** Name the host when the remote points at one - an unresolvable
351
+ `{BITBUCKET_HOST}` is almost always the VPN - and make the retry real: it re-runs the
352
+ fetch and re-enters this picker on a second failure.
333
353
 
334
- Persist user choice in `agent-state.json.baseFetchStatus` ∈ `"fresh" | "cached-stale" | "local-branch" | "aborted"`. On any non-fresh choice, log:
335
- ```
336
- ⚠️ Base ref stale (fetch fail @ {ts}, choice: {cached-stale|local-branch})
337
- ```
338
- Phase 6 (commit/push) MUST re-attempt `git fetch origin` before push; if successful, prompt to rebase before pushing.
354
+ Persist the choice in `agent-state.json.baseFetchStatus` ∈ `"fresh" | "cached-stale" | "local-branch" | "aborted"`, and on any non-fresh one log `⚠️ Base ref stale (fetch fail @ {ts}, choice: {status})`. Phase 6 re-attempts `git fetch origin` before push and prompts to rebase if it succeeds.
339
355
 
340
- In multi-repo mode, the prompt fires per-repo. Choosing `[4] Abort` for any single repo aborts the entire task (atomic - no partial worktrees).
356
+ In multi-repo mode the prompt fires per repo, and Abort on any one aborts the whole task (atomic - no partial worktrees).
341
357
 
342
358
  #### Step 4 - Branch Naming (automatic)
343
359
 
@@ -365,10 +381,8 @@ Branch name is deterministic - no user confirmation needed.
365
381
 
366
382
  **Collision handling** (automatic - no prompt):
367
383
  - Probe local + remote for existing branch. **Distinguish "no such ref" from "the
368
- probe failed"**: with `2>/dev/null` and an empty-output test they look identical,
369
- so an auth or network failure reads as "no collision" and the run creates a
370
- branch that already exists on the remote - surfacing as a rejected push at
371
- Phase 6, far from its cause.
384
+ probe failed"**: with `2>/dev/null` and an empty-output test a failed probe reads
385
+ as "no collision", and the duplicate branch surfaces as a rejected push at Phase 6.
372
386
  ```bash
373
387
  LOCAL_HIT=$(git -C "$root" rev-parse --verify --quiet "refs/heads/$branch")
374
388
  REMOTE_ERR=$(git -C "$root" ls-remote --exit-code --heads origin "$branch" 2>&1 >/dev/null)
@@ -395,6 +409,21 @@ Branch name is deterministic - no user confirmation needed.
395
409
  3. Map instruction files to `"instructionFiles": { "start": "...", "validate": "...", "dev": "...", "commit": "..." }`
396
410
  4. Instruction-driven → later phases read SKILL.md; no instructions → standard phases
397
411
 
412
+ #### Step 5b - Workspace (worktree or local)
413
+
414
+ Ask where the branch lives - the wording, the two options and what local costs
415
+ are in `modes.md`, "Local Mode". Here because Step 4 named the branch, Step 6b
416
+ acts on the answer, and Step 7.5 needs the phase set it implies.
417
+
418
+ **Who is asked.** `/multi-agent` only (`workspaceSource: "asked"`). `:local` and
419
+ `--local` state it up front (`command`).
420
+ Every autopilot entry resolves it to a worktree and never asks (`autopilot`).
421
+
422
+ **Persist** `state.localMode` (semantics unchanged) and `state.workspaceSource`,
423
+ which is what separates a worktree the user chose from one nothing asked about.
424
+
425
+ Log: `Phase 0 Step 5b: workspace = {worktree|local} (source {asked|command|autopilot})`
426
+
398
427
  #### Step 6 - Branch + Workspace Setup
399
428
 
400
429
  **6a. Resolve git identity** (automatic, no prompt):
@@ -419,7 +448,7 @@ Log: `Identity: {identity.name} <{identity.email}>`
419
448
 
420
449
  1. `git -C $PROJECT_ROOT fetch origin`
421
450
 
422
- **If `--local` mode** (no worktree):
451
+ **If local** (Step 5b answered local, by question, by `:local` / `--local`):
423
452
 
424
453
  ```bash
425
454
  if [ -n "$(git -C $PROJECT_ROOT status --porcelain)" ]; then
@@ -430,9 +459,10 @@ git -C $PROJECT_ROOT config user.name "{identity.name}"
430
459
  git -C $PROJECT_ROOT config user.email "{identity.email}"
431
460
  ```
432
461
 
433
- `worktreePath` = `$PROJECT_ROOT`, `localMode` = `true`.
462
+ `worktreePath` = `$PROJECT_ROOT`, `localMode` = `true`. Step 5b already recorded
463
+ `workspaceSource`; this step does not decide, it applies.
434
464
 
435
- **If normal mode** (worktree - default): 2. Worktree path: Jira → `.worktrees/{jiraId}/`, GitHub → `.worktrees/GH{issueNo}/`, free-text → `.worktrees/task-{shortId}/` 3. **Heal stale admin state first** (see "Worktree stale-lock heal" below) and **apply the residue guard** (see "Worktree residue guard" below), then `git -C $PROJECT_ROOT worktree add {path} -b {branch} origin/{baseBranch}` (if exists: enter, pull) 4. Set identity: `git -C {worktree-path} config user.name/email` 5. Create log dir + `agent-log.md` + `agent-state.json` at `$HOME/.claude/logs/multi-agent/{project}/{task-id}/`, never inside the worktree:
465
+ **If worktree** (Step 5b answered worktree, or autopilot resolved it): 2. Worktree path: Jira → `.worktrees/{jiraId}/`, GitHub → `.worktrees/GH{issueNo}/`, free-text → `.worktrees/task-{shortId}/` 3. **Heal stale admin state first** (see "Worktree stale-lock heal" below) and **apply the residue guard** (see "Worktree residue guard" below), then `git -C $PROJECT_ROOT worktree add {path} -b {branch} origin/{baseBranch}` (if exists: enter, pull) 4. Set identity: `git -C {worktree-path} config user.name/email` 5. Create log dir + `agent-log.md` + `agent-state.json` at `$HOME/.claude/logs/multi-agent/{project}/{task-id}/`, never inside the worktree:
436
466
 
437
467
  **Worktree location convention (cited by every other command):** always `{projectRoot}/.worktrees/{taskId}`, inside the repo, never under `$HOME`. `{taskId}` is the directory name from the rule above (`DC-<shortId>` for `/multi-agent:design-check`). `.worktrees` is fixed, not a preference: no `worktreeBasePath` key exists, and `gc-worktrees.sh`, `purge.sh`, the cost renderers and `usage-report.mjs` resolve `<repo>/.worktrees/` by name. Multi-repo tasks get one worktree per repo (the loop below); `--local` creates none and `worktreePath` is `$PROJECT_ROOT`.
438
468
 
@@ -484,7 +514,7 @@ done
484
514
 
485
515
  State file in multi-repo mode:
486
516
  - Single shared `agent-state.json` lives at `$HOME/.claude/logs/multi-agent/{first-project}/{task-id}/agent-state.json` (anchored on the first repo for back-compat with `multi-agent log`/`status` commands)
487
- - **One shared file + per-repo writers is exactly the race `write-state.mjs` exists for.** Every update to it (here and in every later phase) goes through `node $HOME/.claude/scripts/write-state.mjs` per the required mechanism in `operations.md` "Writing `agent-state.json`". A read-modify-write from two repos in the same loop drops one repo's `projects[]` entry.
517
+ - Every update to it, here and in every later phase, goes through `node $HOME/.claude/scripts/write-state.mjs` - the required mechanism in `operations.md` "Writing `agent-state.json`", and the race a per-repo read-modify-write loses `projects[]` entries to.
488
518
  - `state.projects[]` holds per-repo `{name, root, worktreePath, branch, baseBranch, identity, platform, baseFetchStatus, commit, pr, pushAttempts, buildStatus}` - see `agent-state.schema.json`
489
519
  - Scalar fields (`project`, `projectRoot`, `worktreePath`, `branch`, `baseBranch`, `identity`) mirror `projects[0]` so legacy phases that read scalars keep working
490
520
  - Atomicity: if any repo's worktree creation fails (collision aborted, fetch aborted, disk full), roll back already-created worktrees: `git -C $proj worktree remove --force $WT_PATH; git -C $proj branch -D $BRANCH`. Never leave a partial multi-repo state.
@@ -562,7 +592,7 @@ Ask the depth question from `$HOME/.claude/multi-agent-refs/phases/modes.md` "Pi
562
592
 
563
593
  When the intake carried an analysis document or a Figma reference, say so **inside** the question: Short skips the only two phases that would turn that document into a task breakdown, and the user should learn that before choosing, not after.
564
594
 
565
- **Pass the default explicitly, as an index.** `ask-choice.sh` picks the FIRST option on a non-TTY, so relying on option order breaks the first time someone reorders them for readability, on the host where nobody is watching. The default must be the 1-based index, never the label: labels follow `outputLanguage` (`rules.md` matrix), so `ASK_CHOICE_DEFAULT="Full"` matches nothing once the options render as `Tam` / `Kisa`, falls through, and silently takes option 1 on a non-TTY - the exact failure this step exists to prevent.
595
+ **Pass the default as a 1-based index, never a label.** `ask-choice.sh` takes the first option on a non-TTY, and a label-valued default matches nothing once the options render in `outputLanguage`. Reasoning: `modes.md`, "Pipeline depth".
566
596
 
567
597
  ```bash
568
598
  # Full is option 1, Short is option 2 (modes.md "Pipeline depth")
@@ -572,7 +602,9 @@ ASK_CHOICE_DEFAULT="$DEPTH_DEFAULT_INDEX" \
572
602
  "<localized: 'Full'>" "<localized: 'Short'>"
573
603
  ```
574
604
 
575
- **Persist.** Short sets `state.onlyDevelop = true`; Full leaves it `false`. The key is unchanged - only who sets it changed - so every downstream reader keeps working. Short also flips the Phase 1 and Phase 2 tiles to `skipped` (tracker-contract.md, "Late skip"); pre-marking is forbidden.
605
+ **Persist.** Short sets `state.onlyDevelop = true`; Full leaves it `false`. The key is unchanged - only who sets it changed - so every downstream reader keeps working.
606
+
607
+ **Now register the rest of the widget.** This answer is the first moment the phase set is known, so the remaining tiles are created here and not before - Full `1 2 3 4 5 6 7`, Short `3 4 5 6 7`, `:local` dropping 5 from either. `add` each, then `phase-tracker.sh tiles --new`, which emits TaskCreate only for phases that carry no tile yet. Contract: `tracker-contract.md`, "Deferred registration".
576
608
 
577
609
  Log: `Phase 0 Step 7.5: depth = {full|short} (recommended {full|short}, source {user|autopilot|default})`
578
610
 
@@ -585,7 +617,7 @@ BASELINE_LOG="$WORKTREE/.baseline-test.log"
585
617
  timeout "${prefs_testBaseline_timeoutSeconds:-600}" <same-test-command-as-Phase-4-Gate-3> 2>&1 | tee "$BASELINE_LOG"
586
618
  ```
587
619
 
588
- Persist `state.baseline.tests` with `command`, `capturedAt`, `logPath` and exactly one status: `green` (passed), `red` + `failing[]` (failed, names parsed), `red` + empty `failing[]` (failed, names unparseable), `unknown` (no test command, `timeout` fired, or flag off). Folding `unknown` into `green` would let a skipped baseline read as a clean tree, which is the failure this record exists to prevent.
620
+ Persist `state.baseline.tests` with `command`, `capturedAt`, `logPath` and exactly one status: `green` (passed), `red` + `failing[]` (failed, names parsed), `red` + empty `failing[]` (failed, names unparseable), `unknown` (no test command, `timeout` fired, or flag off). Never fold `unknown` into `green`: a skipped baseline would read as a clean tree.
589
621
 
590
622
  Log: `Phase 0 Step 7.6: test baseline = {green|red|unknown} ({N} pre-existing failures)`
591
623
 
@@ -606,9 +638,7 @@ if [ -n "$EVIDENCE_PLATFORM" ]; then
606
638
  fi
607
639
  ```
608
640
 
609
- Empty is an outcome, not a failure: backend has no device, so the probe does not run and `evidenceCapability.skippedReason` says so. Web does: the browser the runner drives is the device (4.6). No `--changed` yet, by the same token.
610
-
611
- One run, both forms: stdout is `EVIDENCE_*` (shell-quoted, so the eval is safe), and the same measurement lands as JSON.
641
+ Empty is an outcome, not a failure: backend has no device, so the probe does not run and `evidenceCapability.skippedReason` says so. Web does - the browser the runner drives is the device (4.6). No `--changed` yet. One run, both forms: stdout is `EVIDENCE_*` (shell-quoted, so the eval is safe), and the same measurement lands as JSON.
612
642
 
613
643
  Persist that file as `state.evidenceCapability`, then build the menu from it, never from a reading of the repo: `1. Sadece unit test` / `2. Unit + UI test, ekran kaydiyla` (tier 1) / `3. Unit + MCP ile akis kaydi` (tier 2).
614
644
 
@@ -647,11 +677,7 @@ Log: `Phase 0 Step 7.7: testDepth = {unit|unit+ui|unit+mcp} (source {user|autopi
647
677
 
648
678
  5. Phase 1 Analysis reads `state.clarification.userAnswers` (when present) as additional context - fold answers into the Explore prompt so downstream phases inherit the resolution.
649
679
 
650
- **Cost:** ~$0.0025 per Haiku call. The pipeline's other expensive phases (Phase 4 reviewers, Phase 3 Sonnet codegen) far outweigh this - the value is avoiding the ~30 min wasted when Phase 3 builds the wrong thing because Phase 0 didn't ask.
651
-
652
- **Reference:** see `$HOME/.claude/agents/task-clarifier.md` for the full scoring rubric and question-quality rules.
653
-
654
- **Why this fits Phase 0 (not a new phase):** clarification doesn't change what code gets written - it changes what gets understood before code is written. Phase 0 already collects identity / project / branch / maturity; ambiguity scoring fits naturally as the last contextual gate.
680
+ **Reference:** `$HOME/.claude/agents/task-clarifier.md` - scoring rubric, question-quality rules and the ~$0.0025-per-call cost note.
655
681
 
656
682
  #### Telemetry
657
683
 
@@ -686,20 +712,19 @@ node "$HOME/.claude/scripts/usage-report.mjs" --task-id "$TASK_ID" >/dev/null 2>
686
712
  The second line reports the run as started: reporting only from Phase 7 reported
687
713
  only runs that finish, and few do. Phase 7 upserts the same key over it.
688
714
 
689
- It asserts three things, each of which has failed silently in a real run:
715
+ It asserts five things, each of which has failed silently in a real run:
690
716
 
691
- 1. **`agent-state.json` exists.** A run once reported Phase 0 `completed` with only
692
- `tracker-state.json` on disk. Every later phase then reasons from fields that are
693
- not there.
694
- 2. **`taskType` is set.** Phase 3 branches on it (Step 7). Absent, a Figma-driven
695
- screen is dispatched as generic development, skipping the stack plugin's
696
- token-compliance check, Code Connect publish and component review. That run
697
- guessed `16` where the frame said `Spacing/12`, and half its commits were rework.
717
+ 1. **`agent-state.json` exists.** Every later phase reasons from it.
718
+ 2. **`taskType` is set.** Phase 3 branches on it (Step 7).
698
719
  3. **A Figma reference forces `taskType: "component"`, and `figmaAccess.tier` is
699
- recorded.** Without the tier, a later phase cannot tell "the design was confirmed"
700
- from "the design was never fetched" - which is exactly when spacing gets guessed.
720
+ recorded**, so a later phase can tell a confirmed design from an unfetched one.
721
+ 4. **`baseBranchSource` is recorded**, and an interactive run recorded `asked` or
722
+ `input` - the only field that separates a branch that was chosen from one that
723
+ was announced.
724
+ 5. **`siblings` is an array**, empty included: Step 2b's record that it ran.
725
+
726
+ Each failed silently in a real run; the script header names which.
701
727
 
702
- A failure is a halt, not a warning. Fix the state and re-run the gate; the phase
703
- stays `in_progress` until it passes. **Never** mark Phase 0 completed on the grounds
704
- that its steps ran - the gate checks the output, and the output is what Phase 3
705
- consumes.
728
+ A failure is a halt, not a warning: fix the state, re-run the gate, and leave the phase
729
+ `in_progress` until it passes. Never close Phase 0 on the grounds that its steps ran -
730
+ the gate checks the output, which is what Phase 3 consumes.