@mmerterden/multi-agent-pipeline 17.4.0 → 17.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (35) hide show
  1. package/CHANGELOG.md +136 -0
  2. package/README.md +23 -5
  3. package/README.tr.md +23 -5
  4. package/docs/adr/0013-lsp-code-intelligence.md +102 -0
  5. package/docs/adr/README.md +1 -0
  6. package/docs/token-budget-history.md +1 -1
  7. package/install/templates/copilot-instructions.md +9 -3
  8. package/package.json +1 -1
  9. package/pipeline/commands/multi-agent/analysis/SKILL.md +3 -3
  10. package/pipeline/commands/multi-agent/autopilot/SKILL.md +3 -3
  11. package/pipeline/commands/multi-agent/autopilot-off/SKILL.md +5 -3
  12. package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +1 -1
  13. package/pipeline/commands/multi-agent/local/SKILL.md +17 -6
  14. package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +3 -3
  15. package/pipeline/lib/multi-repo-pipeline.sh +26 -0
  16. package/pipeline/multi-agent-refs/features/base-branch-evidence.md +222 -0
  17. package/pipeline/multi-agent-refs/features/code-intelligence.md +80 -0
  18. package/pipeline/multi-agent-refs/phases/modes.md +23 -3
  19. package/pipeline/multi-agent-refs/phases/phase-0-init.md +96 -71
  20. package/pipeline/multi-agent-refs/phases/phase-7-report.md +1 -1
  21. package/pipeline/multi-agent-refs/phases.md +7 -2
  22. package/pipeline/multi-agent-refs/picker-contract.md +37 -5
  23. package/pipeline/multi-agent-refs/tracker-contract.md +25 -14
  24. package/pipeline/schemas/agent-state.schema.json +88 -4
  25. package/pipeline/schemas/prefs.schema.json +22 -0
  26. package/pipeline/schemas/token-budget.json +2 -2
  27. package/pipeline/scripts/autopilot-runner.mjs +292 -45
  28. package/pipeline/scripts/base-branch-candidates.mjs +599 -0
  29. package/pipeline/scripts/gc-abandoned.sh +5 -3
  30. package/pipeline/scripts/gen-mode-dispatch.mjs +39 -16
  31. package/pipeline/scripts/phase-tracker.sh +39 -2
  32. package/pipeline/scripts/phase0-exit-gate.mjs +128 -0
  33. package/pipeline/scripts/verify-citations.mjs +84 -2
  34. package/pipeline/skills/.skill-manifest.json +2 -2
  35. package/pipeline/skills/shared/core/multi-agent/SKILL.md +1 -1
@@ -45,9 +45,11 @@ pkill -f "$HOME/.claude/autopilot/bin/menubar" 2>/dev/null || true
45
45
 
46
46
  Without `--now` nothing here runs. With it, the child session is stopped and the
47
47
  run is marked `abandoned`, its worktree is removed **unless it holds uncommitted
48
- work**, in which case the work is stashed to `autopilot/abandoned/<task-id>` and
49
- the worktree is kept. Same contract as `gc-abandoned.sh`; losing a day of edits
50
- is worse than 750 MB.
48
+ work**, in which case the work goes into a stash entry labelled
49
+ `autopilot/abandoned/<task-id>` (find it with `git stash list`) and the worktree
50
+ is kept. Same contract as `gc-abandoned.sh` and `autopilot-runner.mjs`; losing a
51
+ day of edits is worse than 750 MB. No branch is created - one made after
52
+ `stash push` would point at HEAD and contain none of the work.
51
53
 
52
54
  ### 5. Keep the selection
53
55
 
@@ -115,7 +115,7 @@ deletes nothing until you confirm.
115
115
  | Line | Meaning |
116
116
  |---|---|
117
117
  | `would remove` / `would remove worktree` | reapable: stopped past its age limit, or finished with its worktree left behind |
118
- | `would stash` | uncommitted work found - it is stashed to `autopilot/abandoned/<task-id>` and the worktree is **kept** |
118
+ | `would stash` | uncommitted work found - it goes into a stash entry labelled `autopilot/abandoned/<task-id>` (`git stash list`) and the worktree is **kept** |
119
119
  | `would mark abandoned` | the state is stale but there is no worktree left to remove |
120
120
  | `unattributed` | a worktree with no run state. **Never removed, whatever the flags.** `.worktrees/` is not exclusively ours, and nothing distinguishes a hand-made worktree from a pipeline one whose log was pruned. Report the list and its size; the user decides |
121
121
 
@@ -67,10 +67,21 @@ Two channels run in parallel at every phase boundary:
67
67
  ```bash
68
68
  # Phase 0, very first shell call (every CLI):
69
69
  bash $HOME/.claude/scripts/phase-tracker.sh init "$TASK_ID"
70
- for p in "0:Init" "1:Analysis" "2:Planning" "3:Dev" "4:Review" "6:Commit" "7:Report"; do
70
+ bash $HOME/.claude/scripts/phase-tracker.sh add 0 "Init"
71
+ bash $HOME/.claude/scripts/phase-tracker.sh tiles
72
+ bash $HOME/.claude/scripts/phase-tracker.sh update 0 in_progress
73
+
74
+ # Phase 0 Step 7.5, immediately after the depth answer - the first moment this
75
+ # mode knows its phase set. Full:
76
+ for p in "1:Analysis" "2:Planning" "3:Dev" "4:Review" "6:Commit" "7:Report"; do
71
77
  bash $HOME/.claude/scripts/phase-tracker.sh add "${p%%:*}" "${p#*:}"
72
78
  done
73
- bash $HOME/.claude/scripts/phase-tracker.sh update 0 in_progress
79
+ # Short (Analysis and Planning are not run, so they get no tile at all):
80
+ for p in "3:Dev" "4:Review" "6:Commit" "7:Report"; do
81
+ bash $HOME/.claude/scripts/phase-tracker.sh add "${p%%:*}" "${p#*:}"
82
+ done
83
+ # Then the widget, narrowed to the phases that do not have a tile yet:
84
+ bash $HOME/.claude/scripts/phase-tracker.sh tiles --new
74
85
 
75
86
  # Every phase boundary (every CLI):
76
87
  bash $HOME/.claude/scripts/phase-tracker.sh update <N> in_progress|completed|failed|skipped
@@ -83,11 +94,11 @@ bash $HOME/.claude/scripts/phase-tracker.sh tokens <N> <in> <out> [cached]
83
94
 
84
95
  In Claude Code the agent MUST also drive the native TaskList widget so the user sees a sticky phase tile stack - this is the only progress signal Claude Code surfaces. Skipping these calls is the #1 source of "I don't see any phases" complaints.
85
96
 
86
- **TaskCreate ordering (strict)**: All TaskCreate calls fire in strict phase-number order BEFORE any TaskUpdate is applied. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks (e.g. `1 ✓ · 2 ✓ · 4 ✓ · 0 ▶ · 3 ☐`) even when the underlying state is correct. Pre-marking phases as completed/skipped before Phase 0 starts is FORBIDDEN - register the tile in order, then flip status via TaskUpdate when the phase actually short-circuits. Full contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
97
+ **TaskCreate ordering (strict)**: All TaskCreate calls in a registration batch fire in strict phase-number order BEFORE any TaskUpdate in that batch, and a later batch only ever appends phases numbered above everything already registered. This mode registers in two batches (Step -1, then Step 7.5), so `tiles --new` narrows the second one and the Phase 0 tile is never created twice. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks (e.g. `1 ✓ · 2 ✓ · 4 ✓ · 0 ▶ · 3 ☐`) even when the underlying state is correct. Pre-marking phases as completed/skipped before Phase 0 starts is FORBIDDEN - register the tile in order, then flip status via TaskUpdate when the phase actually short-circuits. Full contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
87
98
 
88
99
  ```text
89
- # Phase 0 startup - register one tile per phase (0..N), capture the taskId, persist it:
90
- for each phase in 0:Init, 1:Analysis, 2:Planning, 3:Dev, 4:Review, 6:Commit, 7:Report:
100
+ # Register one tile per phase, capture the taskId, persist it:
101
+ for each phase in 0:Init at Step -1, then 1:Analysis, 2:Planning, 3:Dev, 4:Review, 6:Commit, 7:Report (Full) or 3:Dev, 4:Review, 6:Commit, 7:Report (Short) at Step 7.5:
91
102
  TaskCreate({ subject: "Phase <N>: <Name>", activeForm: "<doing-form>" })
92
103
  -> returns taskId
93
104
  bash $HOME/.claude/scripts/phase-tracker.sh meta <N> tasklist_id "<taskId>"
@@ -108,7 +119,7 @@ bash $HOME/.claude/scripts/phase-tracker.sh update <N> completed
108
119
 
109
120
  #### TaskCreate ordering (strict)
110
121
 
111
- **All TaskCreate calls fire in strict phase-number order BEFORE any TaskUpdate is applied.** For `--local` that means: Phase 0 → Phase 1 → Phase 2 → Phase 3 → Phase 4 → Phase 6 → Phase 7. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks. Full ordering contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
122
+ **All TaskCreate calls in a batch fire in strict phase-number order BEFORE any TaskUpdate is applied.** For `--local` that means: Phase 0 at Step -1, then the rest in ascending order at Step 7.5 (Phase 0 → Phase 1 → Phase 2 → Phase 3 → Phase 4 → Phase 6 → Phase 7 minus whatever the depth answer drops). The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks. Full ordering contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
112
123
 
113
124
  ### Visual channel - Copilot CLI / plain shell
114
125
 
@@ -104,10 +104,10 @@ bash $HOME/.claude/scripts/phase-tracker.sh tokens <N> <in> <out> [cached]
104
104
 
105
105
  In Claude Code the agent MUST also drive the native TaskList widget so the user sees a sticky phase tile stack - this is the only progress signal Claude Code surfaces. Skipping these calls is the #1 source of "I don't see any phases" complaints.
106
106
 
107
- **TaskCreate ordering (strict)**: All TaskCreate calls fire in strict phase-number order BEFORE any TaskUpdate is applied. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks (e.g. `1 ✓ · 2 ✓ · 4 ✓ · 0 ▶ · 3 ☐`) even when the underlying state is correct. Pre-marking phases as completed/skipped before Phase 0 starts is FORBIDDEN - register the tile in order, then flip status via TaskUpdate when the phase actually short-circuits. Full contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
107
+ **TaskCreate ordering (strict)**: All TaskCreate calls in a registration batch fire in strict phase-number order BEFORE any TaskUpdate in that batch, and a later batch only ever appends phases numbered above everything already registered. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks (e.g. `1 ✓ · 2 ✓ · 4 ✓ · 0 ▶ · 3 ☐`) even when the underlying state is correct. Pre-marking phases as completed/skipped before Phase 0 starts is FORBIDDEN - register the tile in order, then flip status via TaskUpdate when the phase actually short-circuits. Full contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
108
108
 
109
109
  ```text
110
- # Phase 0 startup - register one tile per phase (0..N), capture the taskId, persist it:
110
+ # Register one tile per phase, capture the taskId, persist it:
111
111
  for each phase in 0:Init, 1:Analysis, 2:Planning, 3:Dev, 4:Review, 6:Commit, 7:Report:
112
112
  TaskCreate({ subject: "Phase <N>: <Name>", activeForm: "<doing-form>" })
113
113
  -> returns taskId
@@ -129,7 +129,7 @@ bash $HOME/.claude/scripts/phase-tracker.sh update <N> completed
129
129
 
130
130
  #### TaskCreate ordering (strict)
131
131
 
132
- **All TaskCreate calls fire in strict phase-number order BEFORE any TaskUpdate is applied.** For `--local autopilot` that means: Phase 0 → Phase 1 → Phase 2 → Phase 3 → Phase 4 → Phase 6 → Phase 7. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks. Full ordering contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
132
+ **All TaskCreate calls in a batch fire in strict phase-number order BEFORE any TaskUpdate is applied.** For `--local autopilot` that means: Phase 0 → Phase 1 → Phase 2 → Phase 3 → Phase 4 → Phase 6 → Phase 7. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks. Full ordering contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
133
133
 
134
134
  ### Visual channel - Copilot CLI / plain shell
135
135
 
@@ -292,6 +292,12 @@ agent_state = {
292
292
  "worktreePath": os.path.join(worktree_root, primary["name"]),
293
293
  "branch": state["branch"],
294
294
  "baseBranch": state.get("baseBranch", "develop"),
295
+ "baseBranchSource": state.get("baseBranchSource", "asked"),
296
+ # The multi-repo bridge always builds worktrees - that is what it is for -
297
+ # so the workspace was never in question here. It is recorded anyway,
298
+ # because phase0-exit-gate.mjs refuses to close Phase 0 on a state file
299
+ # that cannot say who decided.
300
+ "workspaceSource": state.get("workspaceSource", "command"),
295
301
  "remoteType": primary.get("provider", "github"),
296
302
  "currentPhase": 0,
297
303
  "status": "in_progress",
@@ -308,6 +314,26 @@ related = issue.get("relatedIssues") or []
308
314
  if isinstance(related, list) and related:
309
315
  agent_state["relatedIssues"] = related
310
316
 
317
+ # Always written, `[]` included: the empty array is the record that the
318
+ # dev-context step ran, and phase0-exit-gate.mjs refuses to close Phase 0
319
+ # without it. Only repos this run will NOT modify belong here - the extras that
320
+ # get a worktree are already in projects[].
321
+ STACKS = {"ios", "android", "node", "python", "go"}
322
+ worktreed = {r.get("name") for r in all_repos}
323
+ siblings = []
324
+ for r in (state.get("readonlySiblings") or []) + extras:
325
+ name = r.get("name")
326
+ if not name or name in worktreed:
327
+ continue
328
+ stack = r.get("stack", "unknown")
329
+ siblings.append({
330
+ "name": name,
331
+ "root": r.get("localPath") or None,
332
+ "stack": stack if stack in STACKS else "unknown",
333
+ "canPush": bool(r.get("canPush", False)),
334
+ })
335
+ agent_state["siblings"] = siblings
336
+
311
337
  logs_dir = os.path.expanduser(os.path.join("~/.claude/logs/multi-agent", primary["name"], state["taskId"]))
312
338
  os.makedirs(logs_dir, exist_ok=True)
313
339
  out_path = os.path.join(logs_dir, "agent-state.json")
@@ -0,0 +1,222 @@
1
+ # Feature: Base-Branch Evidence
2
+
3
+ <!-- toc -->
4
+ - [1. The fetch is a fact, not a formality](#1-the-fetch-is-a-fact-not-a-formality)
5
+ - [2. Evidence sources](#2-evidence-sources)
6
+ - [2b. A filter is not a ranking](#2b-a-filter-is-not-a-ranking)
7
+ - [3. The convention is learned from the remote](#3-the-convention-is-learned-from-the-remote)
8
+ - [4. The picker](#4-the-picker)
9
+ - [5. Autopilot](#5-autopilot)
10
+ - [6. State](#6-state)
11
+ <!-- /toc -->
12
+
13
+ **Pattern**: Phase 0 Step 3 used to ask one question with a list it could not
14
+ vouch for. `git fetch origin` ran, its exit code was ignored, and `git branch -r`
15
+ printed the remote-tracking cache either way - so on a restricted network a
16
+ weeks-old local list was presented as the remote's answer, with nothing in the
17
+ output saying so. And the answer was usually derivable: an issue that carries a
18
+ target version, or that links a separate issue representing the release, already
19
+ names the branch on a repo whose release branches encode the version. Nothing
20
+ derived it.
21
+
22
+ The shape here is **evidence collection, then a picker** - deliberately not a
23
+ rule engine. Every candidate carries why it is a candidate; the ranking orders
24
+ them; a human chooses. A rule engine would have to be right, and the inputs
25
+ (field names, branch spellings, board conventions) differ per repo and change
26
+ under it. An evidence list only has to be honest.
27
+
28
+ Implemented by `$HOME/.claude/scripts/base-branch-candidates.mjs` (pure, no git,
29
+ no network - it is handed ref lists and issue evidence). Gated by `prefs.global.baseBranchEvidence.enabled`
30
+ (default `true`; off falls back to the rule-6 sort order, still through the picker). Asserted by `smoke-base-branch-evidence.sh`
31
+ and `test/base-branch-candidates.test.mjs`.
32
+
33
+ ## 1. The fetch is a fact, not a formality
34
+
35
+ ```bash
36
+ git -C "$PROJECT_ROOT" fetch origin; FETCH_RC=$?
37
+ ```
38
+
39
+ `FETCH_RC` decides `refProvenance`, which is a required input to the collector
40
+ and rides on every candidate as its own evidence row:
41
+
42
+ | `FETCH_RC` | `refProvenance` | Ref list | What the user is told |
43
+ |---|---|---|---|
44
+ | 0 | `remote` | `git branch -r` | nothing extra; the list is current |
45
+ | non-zero | `local` | `git branch -r` (cache) **and** `git branch` | the picker question itself says the fetch failed and these refs may be stale |
46
+
47
+ A degraded list is still a list - falling back is correct, hiding it is not. The
48
+ degraded picker gains a **Retry the fetch** row, and the fetch-fail picker
49
+ already in Step 3 (Connect VPN / cached / local branch / abort) still owns the
50
+ `baseFetchStatus` value. The two fit together: that picker records *what the run
51
+ is working from*, this one records *what the candidate list is worth*.
52
+
53
+ Local-only refs never silently become remote ones. `state.baseBranchEvidence.refProvenance`
54
+ must be `local` whenever `baseFetchStatus` is `cached-stale` or `local-branch`,
55
+ and `phase0-exit-gate.mjs` refuses to close Phase 0 otherwise. That assertion is
56
+ the whole of problem 1: the run may degrade, it may not misreport.
57
+
58
+ ## 2. Evidence sources
59
+
60
+ | Kind | Where it comes from | Weight |
61
+ |---|---|---|
62
+ | `issue-version` | a version-typed field on the issue whose value matches a branch on the ref list | 100 exact, 60 same major.minor, x0.6 for an affects-version field |
63
+ | `linked-release` | a linked issue or parent whose own fix-version or summary names a version that matches a branch | 90 exact (110 when `baseBranchEvidence.preferLinkedRelease`) |
64
+ | `version-convention` | the release-branch template inferred from the ref list | annotation only, no points |
65
+ | `recent` | `prefs.global.recentBranches[{projectKey}]`, inside `settings.branchTtlDays` | 40 + up to 10 for recency |
66
+ | `repo-default` | `origin/HEAD` | 20 |
67
+ | `sort-order` | the develop / release / main families Step 3 always sorted by | 15 / 10 / 5 |
68
+ | `ref-provenance` | the fetch outcome above | annotation only, on every candidate |
69
+
70
+ ### Field discovery, not a field table
71
+
72
+ No field id is hardcoded, and none can be: a board's "target version" is a
73
+ custom field whose id differs per Jira instance. What is stable is the *schema*.
74
+ `GET /rest/api/2/field` returns `[{id, name, custom, schema: {type, items, custom}}]`,
75
+ and any field whose schema resolves to `version` - directly, or as `array` with
76
+ `items: "version"` - is read, whatever it is called. The two system fields
77
+ (`fixVersions`, `versions`) are read unconditionally; an affects-version field
78
+ is weighted lower than a fix-version one because it describes where the bug was
79
+ seen, not where the fix lands.
80
+
81
+ `GET /rest/api/2/issue/<key>?expand=names` returns the same display names inline
82
+ and saves the second call when only labels are needed.
83
+
84
+ ### The linked release issue
85
+
86
+ `fields.issuelinks[]` carries `{type: {inward, outward}, inwardIssue, outwardIssue}`
87
+ and `fields.parent` carries the parent of a sub-task. On a board where opening a
88
+ development sub-task requires selecting the related release issue, the link is
89
+ the authoritative answer and the version field is the corroboration - so
90
+ `prefs.global.baseBranchEvidence.preferLinkedRelease` (default `false`) raises the
91
+ linked-release weight above the version-field one rather than adding a second rule. A linked issue counts as
92
+ a release issue when a version can be read out of its own fix-version field or
93
+ its summary; the issue *type name* is not consulted, because type names are
94
+ per-board copy.
95
+
96
+ GitHub's analogue is the milestone title, read the same way.
97
+
98
+ ## 2b. A filter is not a ranking
99
+
100
+ Step 3's rule-5 list used to be narrowed with `grep -E '(develop|release|main|master)'`,
101
+ which is the prefix table this whole feature exists not to have - and worse than a
102
+ table, because it DISCARDS. A repo whose release branches read `stabilise-2.7`
103
+ matched none of the four words, so none of its branches reached the picker, the
104
+ convention-learner had nothing to learn from, and the version on the issue could
105
+ never match anything. The feature would have been inert on exactly the repos it
106
+ was written for.
107
+
108
+ The rule is the distinction: **a word list that ranks is fine, a word list that
109
+ filters is not.** Ranking only reorders rows the user can still see past; filtering
110
+ removes answers with nothing saying so. So the four family words stay where they
111
+ belong - rule 6's sort and this collector's 15/10/5 `sort-order` weights - and the
112
+ filter gained a version alternative that admits any branch carrying a version
113
+ token, whatever it is called.
114
+
115
+ The excluded prefixes (`feature/`, `bugfix/`, `fix/`, `hotfix/`, `chore/`) are a
116
+ different thing again: those are the task branches **this pipeline creates itself**,
117
+ in Step 4. Excluding your own output is not a convention assumption.
118
+
119
+ ## 3. The convention is learned from the remote
120
+
121
+ A table of branch prefixes would be wrong for most repos the day it was written.
122
+ So the template is inferred from the refs that exist: every branch carrying a
123
+ version token is reduced to its shape by replacing the token with `<version>`,
124
+ and the most common shape is this repo's convention. A repo whose release
125
+ branches read `<prefix>/develop_<version>` produces that template; a repo that
126
+ spells them `release-<version>` produces that one; a repo with no versioned
127
+ branches produces `null` and the whole source goes quiet.
128
+
129
+ The inference is reported with its member count (`learned from 4 branch(es)`),
130
+ because an inference from one branch and an inference from a dozen are different
131
+ claims and the picker row should not flatten them.
132
+
133
+ **A predicted branch is never offered.** When the learned template predicts a
134
+ name that is not on the ref list, that is a note, not an option:
135
+
136
+ ```
137
+ version 1.51.0 (field "Target Version") has no branch on the remote;
138
+ the convention <prefix>/develop_<version> learned from 4 branch(es) would spell it
139
+ <prefix>/develop_1.51.0
140
+ ```
141
+
142
+ That sentence is more useful than a candidate would be - it tells the user the
143
+ release branch has not been cut yet, which is a real answer to "which base?" -
144
+ and an option the user picks has to be checkoutable. Offering a name that does
145
+ not exist just moves the failure to Step 8, where it reads as a git error.
146
+
147
+ ## 4. The picker
148
+
149
+ `toPickerOptions()` builds the rows, and the **evidence is the description**. A
150
+ row reading `matches version 1.51.0 from the issue field "Target Version"` and a
151
+ row reading `the repository's default branch` are different answers to the same
152
+ question; a picker that hides which one it is cannot be chosen on its merits.
153
+
154
+ The two-option floor is enforced inside that function rather than left to the
155
+ caller: it never returns fewer than two rows, because `AskUserQuestion` refuses a
156
+ question with fewer than two declared options *and discards every question
157
+ batched with it*, and the host's injected Other row does not count toward the
158
+ schema minimum. The escape row (`Show all branches`) is a genuine second choice,
159
+ not an `OK` button - see picker-contract.md, "Two options or it is not a
160
+ question".
161
+
162
+ **Interactive runs always ask.** A derived candidate is a better-ordered list,
163
+ never a skipped question. `baseBranchSource` stays `asked`, because a human
164
+ answered; the derivation is recorded separately in `state.baseBranchEvidence` so
165
+ it stays readable afterwards.
166
+
167
+ ## 5. Autopilot
168
+
169
+ Autopilot cannot be asked anything, so it resolves in this order and records
170
+ which rule fired in `baseBranchSource`:
171
+
172
+ 1. `remembered` - `recentBranches` inside the TTL and still on the ref list (memory outranks the default, per the picker contract).
173
+ 2. `derived` - the top candidate carries `issue-version` or `linked-release` evidence **and** no other candidate ties its score. Ambiguity is not resolved by coin-flip.
174
+ 3. `default` - the sort order.
175
+
176
+ `derived` is an autopilot-only value, exactly like `remembered` and `default`;
177
+ an interactive run recording it has skipped its picker, and the exit gate fails
178
+ it. The gate additionally refuses `derived` unless `state.baseBranchEvidence`
179
+ actually contains issue-derived evidence for the chosen branch - a `derived` that
180
+ derived from nothing is `default` wearing a hat.
181
+
182
+ ### Asking on the issue (`prefs.global.baseBranchEvidence.autopilotAsksOnIssue`, default OFF)
183
+
184
+ When the derivation is **ambiguous** and nothing is remembered, autopilot can
185
+ post one comment on the Jira issue or GitHub issue asking which branch to
186
+ develop from, then halt.
187
+
188
+ This is an outward-facing write, so it is fenced:
189
+
190
+ - **Off by default.** Nothing posts unless the user turned it on for this project.
191
+ - **A question, never a state change.** No transition, no resolution, no assignee change, no label change, no close - ever. The standing rule that the pipeline never auto-closes an issue is not relaxed by this feature, and a comment is the only write it is allowed to make.
192
+ - **Human-facing copy follows `outputLanguage`**, like every other comment the pipeline writes.
193
+ - **`Ref:`, never `Closes:`/`Fixes:`/`Resolves:`** in the body, so no platform-side automation reads it as an instruction.
194
+ - **It renders the same evidence the picker would have shown**, candidate by candidate, so the person answering sees what the run saw.
195
+ - **Then the run halts.** Posting a question and continuing on a guess is worse than not asking: the guess lands in a branch while the question sits unanswered. So the comment trips the circuit breaker (trigger 6), which is the sanctioned autopilot pause - state recorded, one actionable line printed, waiting for an explicit `resume`. Phase 0 does not close and no worktree is created.
196
+
197
+ The nearest-safe reading of "post a comment asking which branch" is therefore:
198
+ ask once, change nothing, stop. An autopilot that posts a question and then
199
+ answers it itself has not asked anything.
200
+
201
+ ## 6. State
202
+
203
+ ```jsonc
204
+ "baseBranchSource": "asked" | "input" | "remembered" | "default" | "derived",
205
+ "baseBranchEvidence": {
206
+ "refProvenance": "remote" | "local", // required whenever this object exists
207
+ "chosen": "<branch>",
208
+ "ambiguous": false,
209
+ "convention": { "template": "<prefix>/develop_<version>", "members": 4 },
210
+ "candidates": [
211
+ { "branch": "<branch>", "score": 115,
212
+ "evidence": [{ "kind": "issue-version", "detail": "..." }] }
213
+ ],
214
+ "notes": ["..."],
215
+ "askedOnIssue": { "target": "PROJ-1234", "url": "...", "at": "..." }
216
+ }
217
+ ```
218
+
219
+ `baseBranchEvidence` is required when `baseBranchSource` is `derived`, and when
220
+ `baseFetchStatus` is `cached-stale` or `local-branch`. Everywhere else it is
221
+ optional: a run that took the base from the task reference has no evidence to
222
+ record and should not be made to invent some.
@@ -0,0 +1,80 @@
1
+ ## Code Intelligence (multi-agent-toolkit `code_*`)
2
+
3
+ Compiler-grade answers about Swift and Kotlin source, from the language server
4
+ each platform already ships. Eight tools in the companion MCP server
5
+ (`@mmerterden/multi-agent-toolkit-mcp` 3.10.0 and later), all read-only except
6
+ `code_server_reset`.
7
+
8
+ ### What it is for, and what `code-graph.md` already covers
9
+
10
+ These two answer different questions and neither replaces the other.
11
+
12
+ | Question | Use |
13
+ |---|---|
14
+ | Where does this concept live in a repo I do not know | `code-graph.md` - one cheap pass over the whole tree, token-budgeted |
15
+ | What depends on this area, roughly | `graph-affected` |
16
+ | Is THIS exact symbol referenced, and where | `mcp__multi-agent-toolkit__code_references` |
17
+ | What type is this, really | `mcp__multi-agent-toolkit__code_hover` |
18
+ | Does this file compile, without a full build | `mcp__multi-agent-toolkit__code_diagnostics` |
19
+
20
+ `docs/adr/0010-own-code-graph.md` states the trade the graph makes in as many
21
+ words: *"Regex over comment-stripped source is not a parser. Definitions and
22
+ imports survive that trade; call graphs and type resolution do not."* It names
23
+ two consequences - a name declared in two files is dropped rather than fanned
24
+ out, so `affected` under-reports on duplicated names, and only type-like symbols
25
+ are reference targets. Those are exactly the two the language server answers
26
+ exactly. So the graph stays the wide, cheap first pass, and `code_*` is what a
27
+ specific claim is checked against.
28
+
29
+ ### The one thing to know before trusting an answer
30
+
31
+ An empty result is not the same as a negative result, and on a cold repository
32
+ it is the likelier of the two. Measured on Swift 6.3.1 against a two-file
33
+ package: the same `references` query answered 0 at 0.7s and 3 at 5.8s, with
34
+ nothing changed except that the background index had finished. On a 658-file
35
+ package with dependencies the first index had not finished at 70s.
36
+
37
+ Every index-backed result therefore carries `indexReady`, and the tools wait for
38
+ the index rather than answering early. When `indexReady` is false, an empty list
39
+ means "not indexed yet". Never report it as "unused".
40
+
41
+ `code_index_status` answers whether this machine can answer at all, and it
42
+ succeeds even when nothing is installed - it is the tool to run first on an
43
+ unfamiliar machine.
44
+
45
+ ### Root decides what an answer is worth
46
+
47
+ `semantic` and `buildSettingsSource` come back on every result.
48
+
49
+ | Root | definition | references | diagnostics |
50
+ |---|---|---|---|
51
+ | `Package.swift` | yes | yes, after the index settles | yes |
52
+ | `buildServer.json` | yes | yes | yes |
53
+ | `.xcodeproj`, no build server | same-file only | no, `semantic: false` | misleading, marked non-semantic |
54
+
55
+ A modular app is mostly packages, so this matters less than it reads: the iOS
56
+ app measured here has 41 `Package.swift` roots and a single `.swift` file under
57
+ the umbrella project.
58
+
59
+ ### Where a run uses them
60
+
61
+ All three are opt-in and none is on by default.
62
+
63
+ 1. **Phase 4, checking a citation's claim.** `verify-citations.mjs` already
64
+ proves a cited `file:line` exists. `code_hover` at that position proves the
65
+ line holds what the finding says it holds, and `code_references` falsifies a
66
+ "nothing handles this" claim.
67
+ 2. **Phase 1, sharpening an impact estimate.** When `graph-affected` reports on
68
+ a name that the graph itself flagged as ambiguous, `code_references` on the
69
+ one symbol the task names resolves it exactly.
70
+ 3. **Phase 3, fast feedback.** `code_diagnostics` on the file just edited,
71
+ without waiting for `xcodebuild`.
72
+
73
+ ### Kotlin
74
+
75
+ Runs on JetBrains `kotlin-lsp`: Alpha, partially closed source, Android Gradle
76
+ Plugin support experimental, and installed separately
77
+ (`brew tap JetBrains/utils && brew install kotlin-lsp`, JVM 17+). Answers carry
78
+ `confidence: "alpha"`, and the toolkit's own README states that no gate in that
79
+ repository exercises the Kotlin server. Treat Kotlin answers as a lead, not as
80
+ evidence, until that changes.
@@ -93,7 +93,7 @@ Phase 0: Init -> Phase 3: Dev (self-contained) -> Phase 4: Review -> Phase 5: Te
93
93
  | Phase | Full | Short | Short + `--local` |
94
94
  | ------------------- | ----------------------------------------------- | ---------------------------------------------------------------------------------- | --------------------------------------------- |
95
95
  | Phase 0 (Init) | Full setup | Same - worktree, branch, state, and the depth question itself | Same - no worktree, branch on `$PROJECT_ROOT` |
96
- | Phase 1 (Analysis) | Parallel Explore agents + analysis document | **SKIP** - tile flips to `skipped` at Step 7.5 | **SKIP** |
96
+ | Phase 1 (Analysis) | Parallel Explore agents + analysis document | **SKIP** - no tile is ever drawn for it (registration is deferred to Step 7.5) | **SKIP** |
97
97
  | Phase 2 (Planning) | TaskCreate + architecture review + **Plan Approval Gate** | **SKIP** (no plan means no plan gate) | **SKIP** |
98
98
  | Phase 3 (Dev) | Follows the Phase 2 plan, TDD cycle (Sonnet) | **Self-contained** (Opus): agent scans relevant files, implements with TDD, builds | Same, on the local branch |
99
99
  | Phase 4 (Review) | Parallel review + Fable triage (3 reviewers on every host: Claude Code Fable + Opus + Sonnet, Copilot GPT-5.4 + Opus + Sonnet) | **Same** - gates, parallel review, triage; blocking findings return to Phase 3 (cap 3) | **Same**, on the local branch diff |
@@ -111,7 +111,7 @@ The **Opus** agent receives the task description (from Jira, GitHub issue, or fr
111
111
 
112
112
  No separate task breakdown - the agent handles scope autonomously.
113
113
 
114
- **State tracking**: `agent-state.json` gets `"onlyDevelop": true`. Because the tracker boots at Step -1 and the answer arrives at Step 7.5, Phases 1 and 2 are registered `pending` with everything else and flipped to `skipped` when the answer lands; pre-marking is forbidden and would scramble the tile order. Full contract: `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "Late skip".
114
+ **State tracking**: `agent-state.json` gets `"onlyDevelop": true`. The tracker boots at Step -1 with Phase 0 alone and the rest of the tiles are registered at Step 7.5, once this answer says which phases the run actually has - so a Short run never draws an Analysis tile it will not use. Full contract: `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "Deferred registration".
115
115
 
116
116
  ### Intake warnings for a Short run
117
117
 
@@ -161,7 +161,22 @@ Phase 0: Init -> Phase 1: Analysis -> Phase 2: Planning -> Phase 4: Review -> Ph
161
161
 
162
162
  Local mode skips worktree creation - works directly on a local branch in the project root. Useful for single-task workflows or when worktrees cause issues.
163
163
 
164
- **Activation**: the `:local` entries, or `--local` on the base command:
164
+ **Activation**: answer the Phase 0 Step 5b workspace question, or state it up
165
+ front so the question resolves without being asked.
166
+
167
+ ```
168
+ Bu is nerede kossun? / Where should this task run?
169
+ 1. Worktree .worktrees/{id}/ - your current checkout stays untouched
170
+ 2. Lokal / Local the project root, on a new branch - no Phase 5
171
+ ```
172
+
173
+ Two genuine options, so it meets the two-option floor in `picker-contract.md` on
174
+ its own. **Say what local costs inside the question**: Phase 5 is not in a local
175
+ run's set (the user-test gate checks the change out of a worktree, and there is
176
+ none), and uncommitted work in the project root is in the way of the checkout. A
177
+ user choosing local should learn both before choosing, not after.
178
+
179
+ The ways to state it up front:
165
180
 
166
181
  ```
167
182
  /multi-agent:local "PROJ-12345"
@@ -171,6 +186,11 @@ Local mode skips worktree creation - works directly on a local branch in the p
171
186
  /multi-agent "PROJ-12345" --local
172
187
  ```
173
188
 
189
+ Every autopilot entry resolves this to **worktree** and never asks. That is not a
190
+ skipped question: an unattended run commits and pushes from wherever it stands,
191
+ and doing that in the user's own checkout is what worktrees exist to prevent.
192
+ `:local-autopilot` is the explicit opt-out.
193
+
174
194
  **What changes in local mode:**
175
195
 
176
196
  | Phase | Normal (worktree) | Local |