@mmerterden/multi-agent-pipeline 17.3.0 → 17.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (54) hide show
  1. package/CHANGELOG.md +203 -0
  2. package/README.md +23 -5
  3. package/README.tr.md +23 -5
  4. package/docs/adr/0013-lsp-code-intelligence.md +102 -0
  5. package/docs/adr/README.md +1 -0
  6. package/docs/token-budget-history.md +1 -1
  7. package/install/templates/copilot-instructions.md +9 -3
  8. package/package.json +1 -1
  9. package/pipeline/agents/code-reviewer.md +35 -1
  10. package/pipeline/commands/multi-agent/analysis/SKILL.md +3 -3
  11. package/pipeline/commands/multi-agent/autopilot/SKILL.md +3 -3
  12. package/pipeline/commands/multi-agent/autopilot-off/SKILL.md +5 -3
  13. package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +1 -1
  14. package/pipeline/commands/multi-agent/local/SKILL.md +17 -6
  15. package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +3 -3
  16. package/pipeline/lib/multi-repo-pipeline.sh +26 -0
  17. package/pipeline/multi-agent-refs/analysis/locked.md +4 -4
  18. package/pipeline/multi-agent-refs/analysis/render.md +2 -1
  19. package/pipeline/multi-agent-refs/channels/pr.md +26 -0
  20. package/pipeline/multi-agent-refs/cross-cli-contract.md +22 -0
  21. package/pipeline/multi-agent-refs/features/base-branch-evidence.md +222 -0
  22. package/pipeline/multi-agent-refs/features/code-graph.md +40 -0
  23. package/pipeline/multi-agent-refs/features/code-intelligence.md +80 -0
  24. package/pipeline/multi-agent-refs/features/design-conformance.md +14 -0
  25. package/pipeline/multi-agent-refs/features/review-file-set.md +132 -0
  26. package/pipeline/multi-agent-refs/phases/modes.md +23 -3
  27. package/pipeline/multi-agent-refs/phases/phase-0-init.md +96 -71
  28. package/pipeline/multi-agent-refs/phases/phase-4-review.md +31 -23
  29. package/pipeline/multi-agent-refs/phases/phase-7-report.md +1 -1
  30. package/pipeline/multi-agent-refs/phases.md +7 -2
  31. package/pipeline/multi-agent-refs/picker-contract.md +37 -5
  32. package/pipeline/multi-agent-refs/tracker-contract.md +25 -14
  33. package/pipeline/schemas/agent-state.schema.json +88 -4
  34. package/pipeline/schemas/prefs.schema.json +22 -0
  35. package/pipeline/schemas/review-file-exclusions.json +137 -0
  36. package/pipeline/schemas/reviewer-output.schema.json +27 -1
  37. package/pipeline/schemas/token-budget.json +2 -2
  38. package/pipeline/scripts/autopilot-runner.mjs +292 -45
  39. package/pipeline/scripts/base-branch-candidates.mjs +599 -0
  40. package/pipeline/scripts/diff-risk-score.mjs +1 -36
  41. package/pipeline/scripts/gc-abandoned.sh +5 -3
  42. package/pipeline/scripts/gen-mode-dispatch.mjs +39 -16
  43. package/pipeline/scripts/git-path.mjs +63 -0
  44. package/pipeline/scripts/glob-match.mjs +62 -0
  45. package/pipeline/scripts/graph-mermaid.mjs +251 -0
  46. package/pipeline/scripts/phase-tracker.sh +39 -2
  47. package/pipeline/scripts/phase0-exit-gate.mjs +128 -0
  48. package/pipeline/scripts/review-file-filter.mjs +180 -0
  49. package/pipeline/scripts/skill-conformance.mjs +1 -31
  50. package/pipeline/scripts/validate-analysis-doc.mjs +53 -0
  51. package/pipeline/scripts/validate-reviewer.mjs +90 -1
  52. package/pipeline/scripts/verify-citations.mjs +428 -0
  53. package/pipeline/skills/.skill-manifest.json +2 -2
  54. package/pipeline/skills/shared/core/multi-agent/SKILL.md +1 -1
@@ -240,8 +240,7 @@ $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.criteria \
240
240
  Output conforms to `$HOME/.claude/schemas/criteria-manifest.schema.json`. The four parts that matter downstream:
241
241
 
242
242
  - **`selectedRules[]` is the denominator.** Rule IDs from every registry whose declared `scope` matches the diff, persisted BEFORE the reviewers run so the set cannot be renegotiated after one has seen the diff. That is what makes "applied completely" answerable rather than "looks fine": a reviewer finding nothing must still return a verdict per ID. Discovery is declared via `standards-registry:` frontmatter, so no stack-specific skill is named here.
243
- - **`coverage.declaredGaps[]` + `droppedReasons`.** An uncovered language, and every out-of-scope rule, is reported with a reason. "No rule applied", "every rule passed" and "the rule set was narrowed" must never render the same.
244
- - **`ledger`.** `state.telemetry.skillCalls[]` corroborates only; `ledger.source` defaults to `derived` and coverage is never computed from self-report. A declared skill the resolver cannot bind is flagged.
243
+ - **`coverage.declaredGaps[]`, `droppedReasons`, `ledger`.** Every gap, dropped rule and unbindable declared skill is reported with a reason; coverage is never computed from self-report. Rationale: the contract above.
245
244
  - **`findings[]`.** Reviewer-shaped, merging at Step 3.0 alongside test-integrity: expired / unexplained / unknown-ID exception markers, plus any registry whose delegated linter is not wired here (those rules are unverified, so reporting no violations reports that nothing was measured).
246
245
 
247
246
  **Any non-zero exit halts** (1 = setup error incl. a bad `--skills-root`; 2 = no root resolved, unparseable registry, or a declared path escaping its skill dir): continuing would drop a rule set from the denominator, and since the reviewer validator skips the checklist when zero rules were selected, the run would report clean over criteria never loaded. A coverage gap halts only under `prefs.global.skillConformance.blockOnCoverageGap`. No opt-out for the stage or the exception-expiry check, on the same grounds as Step 1.76.
@@ -266,13 +265,25 @@ Visual-fidelity mismatches against the captured screenshot are BLOCKING findings
266
265
 
267
266
  When `state.figmaAccess.tier === 3` (user-attached screenshot, no Code Connect snippet), the reviewer additionally sets `findings[i].severity = "blocking"` and `findings[i].tag = "review_blocking_tier3"` on every UI atom that lacks a confirmed canonical-component mapping. The triage step preserves these findings unless the user has explicitly cleared the open question.
268
267
 
268
+ #### Step 1.85 - Review file set (the denominator, required)
269
+
270
+ Decides what the reviewers read before the cap decides what fits: the cap truncates the largest files first, so a lockfile survives while real code is cut. Zero LLM. Full contract: `$HOME/.claude/multi-agent-refs/features/review-file-set.md`.
271
+
272
+ ```bash
273
+ git -C "$WORKTREE" diff --name-only "$BASE_BRANCH"...HEAD \
274
+ | node $HOME/.claude/scripts/review-file-filter.mjs \
275
+ > "$WORKTREE/.pipeline/review-files.json"
276
+ ```
277
+
278
+ `reviewed[]` is the denominator, fixed here for the same reason `selectedRules[]` is. `excluded[]` reaches the run report with its reason and the glob that matched. Exit 2 means the pattern list is unreadable and everything is reviewed: continue, logging `review.file_filter_failed`.
279
+
269
280
  #### Step 1.9 - Context economy (cache prefix + diff cap)
270
281
 
271
282
  Phase 4 sends the same diff to every reviewer and then to triage, so the diff is the dominant token cost. Two measures keep it bounded:
272
283
 
273
- **Shared cache prefix.** Build the reviewer and triage prompts so the large invariant context - the full diff, the `${CRITERIA}` block from Step 1.78, the Phase 1 analysis summary, the Phase 2 plan - is a byte-identical leading block across all dispatches in this iteration. Only the per-reviewer focus + skill line varies, and it goes AFTER the shared block. `${CRITERIA}` goes in the prefix, identical for every reviewer: subsetting it per reviewer would invalidate the prefix for the whole panel and re-bill the largest block in the phase. Per-reviewer emphasis stays a one-line pointer in the suffix. When the host supports prompt caching, the 2nd/3rd reviewer and the triage call then read that prefix at the discounted cache-read rate instead of re-billing it as fresh input. Forward the host-reported cache-read count as `tokens_cached` per the Token telemetry contract so the saving lands in the cost ledger. The `<scope-self-check>` block and, from iteration 2, the `<previous-round-findings>` block (Step 2.1) close the shared block, after the plan and before the per-reviewer suffix.
284
+ **Shared cache prefix.** Build the reviewer and triage prompts so the large invariant context - the full diff, the `${CRITERIA}` block from Step 1.78, the `${REVIEW_FILES}` list from Step 1.85, the Phase 1 analysis summary, the Phase 2 plan - is a byte-identical leading block across all dispatches in this iteration. Only the per-reviewer focus + skill line varies, and it goes AFTER the shared block. `${CRITERIA}` goes in the prefix, identical for every reviewer: subsetting it per reviewer would invalidate the prefix for the whole panel and re-bill the largest block in the phase. Per-reviewer emphasis stays a one-line pointer in the suffix. When the host supports prompt caching, the 2nd/3rd reviewer and the triage call then read that prefix at the discounted cache-read rate instead of re-billing it as fresh input. Forward the host-reported cache-read count as `tokens_cached` per the Token telemetry contract so the saving lands in the cost ledger. The `<scope-self-check>` block and, from iteration 2, the `<previous-round-findings>` block (Step 2.1) close the shared block, after the plan and before the per-reviewer suffix.
274
285
 
275
- **Single-repo diff cap.** If the diff exceeds the Phase 4 token allowance (`token-budget.json`), truncate the largest files and append a footer `[truncated - full diff in file://$WORKTREE/.review-diff.txt]`, writing the full diff to that path. Reviewers and triage receive the same capped view + the marker so they can flag "review the full diff manually." Log `review.diff_truncated bytes_dropped=<N>`. (Multi-repo already caps the combined diff at 80% of budget; this is the single-repo equivalent.)
286
+ **Single-repo diff cap.** Applied to the Step 1.85 `reviewed[]` set only. If it exceeds the Phase 4 token allowance (`token-budget.json`), truncate the largest files and append a footer `[truncated - full diff in file://$WORKTREE/.review-diff.txt]`, writing the full diff to that path. Reviewers and triage receive the same capped view + the marker so they can flag "review the full diff manually." Log `review.diff_truncated bytes_dropped=<N>`. (Multi-repo already caps the combined diff at 80% of budget; this is the single-repo equivalent.)
276
287
 
277
288
  #### Step 2 - Parallel AI Review (CLI-aware reviewer set)
278
289
 
@@ -291,9 +302,8 @@ Reviewer count per host: **Claude Code 3, Copilot CLI 3, Codex CLI 3** - **2**
291
302
 
292
303
  #### Codex CLI - two constraints that fail silently
293
304
 
294
- Both were measured against Codex 0.145, not inferred, and both produce a review that
295
- looks like it ran. The full statement lives in the managed block at `~/.codex/AGENTS.md`
296
- (always loaded on that host, so it is not restated here):
305
+ Both measured against Codex 0.145, both produce a review that looks like it ran. Full
306
+ statement: the managed block in `~/.codex/AGENTS.md`, always loaded on that host.
297
307
 
298
308
  1. **`fork_turns: "none"` on every `spawn_agent` that sets `model` or
299
309
  `reasoning_effort`** - a full-history fork discards the override and collapses the
@@ -304,15 +314,10 @@ looks like it ran. The full statement lives in the managed block at `~/.codex/AG
304
314
  Sub-agent delegation itself is authorized by that same managed block; without it Phase 4
305
315
  degrades to a single in-thread review.
306
316
 
307
- **Single-vendor caveat.** Every Codex reviewer is an OpenAI model, and every Claude
308
- Code reviewer is an Anthropic model, so the cross-vendor disagreement that Copilot CLI
309
- gets for free (GPT-5.4 beside two Claude models) is absent on both. The diversity budget
310
- shifts to model generation, reasoning effort and persona focus: on Codex Reviewer 1 runs
311
- `xhigh` on security and architecture, Reviewer 2 runs a different model family member
312
- on edge cases, Reviewer 3 runs `medium` on quality; on Claude Code the three slots are
313
- three different Claude tiers. Treat consensus among a single-vendor panel as weaker
314
- evidence than the same consensus on Copilot CLI, and say so in the triage note when all
315
- three agree on a borderline finding.
317
+ **Single-vendor caveat.** Claude Code and Codex both run a one-vendor panel, so their
318
+ consensus is weaker evidence than Copilot CLI's; say so in the triage note on a
319
+ borderline finding. Where the diversity budget goes instead:
320
+ `cross-cli-contract.md`, "Panel diversity per host".
316
321
 
317
322
  Each reviewer inherits the `code-reviewer` agent's focus areas (Security, Architecture, Quality, Performance) and output contract. The orchestrator overrides only the model and the stack-specific skill per-reviewer - no prompt duplication.
318
323
 
@@ -331,7 +336,7 @@ Each reviewer inherits the `code-reviewer` agent's focus areas (Security, Archit
331
336
 
332
337
  ##### 2.1 Previous-round findings (iteration >= 2) and 2.2 scope self-check (every iteration)
333
338
 
334
- A reviewer has no memory of the round before, so it rediscovers last round's findings in new words. From iteration 2, render the previous round's accepted blocking/important findings (`.pipeline/triage-round-$((ITERATION-1)).json`, max 40) into a `<previous-round-findings>` block at the end of the shared prefix: a still-present issue is reported with the SAME fingerprint and the current line, a fixed one is omitted, anything new leaves `fingerprint` unset. Every iteration also renders `.pipeline/scope-check.json` (Phase 3 Step 3.7) plus `scope-check-gate.mjs --advisory` output as `<scope-self-check>`: file reasons, unjustified files, and `notDone[]` (never re-raised as findings); a missing record logs `review.scope_check=missing`. Block text and recipes: `$HOME/.claude/multi-agent-refs/features/review-delta.md`.
339
+ From iteration 2, render the previous round's accepted blocking/important findings (`.pipeline/triage-round-$((ITERATION-1)).json`, max 40) into a `<previous-round-findings>` block at the end of the shared prefix. Every iteration also renders `.pipeline/scope-check.json` (Phase 3 Step 3.7) plus `scope-check-gate.mjs --advisory` output as `<scope-self-check>`: file reasons, unjustified files, and `notDone[]` (never re-raised as findings); a missing record logs `review.scope_check=missing`. Block text and recipes: `$HOME/.claude/multi-agent-refs/features/review-delta.md`.
335
340
 
336
341
  #### Step 2.8 - Visual conformance gate (component / screen work only)
337
342
 
@@ -350,11 +355,7 @@ disk with `Code Connect: Not published` in Figma means the binding does not exis
350
355
  anyone but the author. Assert the publish step ran; an unpublished binding is a
351
356
  blocking finding.
352
357
 
353
- Why this is a gate and not advice: `design-check` existed as a command for a while
354
- with **no phase invoking it**, so the only thing standing between a build and visual
355
- drift was the user opening the app and looking. On one run that produced 16pt padding
356
- where the frame said `Spacing/12`, and a full sheet rebuild afterwards. A reviewer
357
- reading a diff cannot see spacing; something has to compare against the design.
358
+ Why this is a gate and not advice: `design-conformance.md`, "Why this runs as a gate".
358
359
 
359
360
  Skip only when the diff has no UI change. Record the outcome in
360
361
  `consensus.visualConformance` so Phase 7 reports whether it ran.
@@ -374,10 +375,11 @@ Step 2 produces N reviewer-output objects (one per dispatched reviewer), each co
374
375
  ```json
375
376
  {"findings":[{"severity":"blocking|important|suggestion","file":"...","line":N,"issue":"...","fix":"...","ruleId":"SEC-01","criteriaSource":"ios-coding-standard"}],
376
377
  "conformance":[{"ruleId":"SEC-01","verdict":"conformant|violated|not-applicable","file":"...","line":N,"reason":"..."}],
378
+ "fileCoverage":[{"path":"src/App.swift","verdict":"reviewed|skipped","reason":"..."}],
377
379
  "approved":true|false}
378
380
  ```
379
381
 
380
- `ruleId` + `criteriaSource` appear on a finding that cites a rule from `${CRITERIA}`. `conformance` is required whenever Step 1.78 selected at least one rule, with exactly one row per selected ID and none outside the set.
382
+ `ruleId` + `criteriaSource` appear on a finding that cites a rule from `${CRITERIA}`. `conformance` is required whenever Step 1.78 selected at least one rule, and `fileCoverage` whenever Step 1.85 left at least one file in `reviewed[]`: one row per ID, one row per path, none outside either set.
381
383
 
382
384
  **Required: validator gate (deterministic) - run immediately after each reviewer returns, before merging findings.** Persist each reviewer's output and validate the file - the validator's exit code decides, not the LLM turn:
383
385
 
@@ -386,9 +388,15 @@ REVIEWER_FILE="$WORKTREE/.pipeline/reviewer-$N.json"
386
388
  printf '%s' "$REVIEWER_JSON" > "$REVIEWER_FILE"
387
389
  node $HOME/.claude/scripts/validate-reviewer.mjs "$REVIEWER_FILE" \
388
390
  --criteria "$WORKTREE/.pipeline/criteria-manifest.json" \
391
+ --coverage "$WORKTREE/.pipeline/review-files.json" \
392
+ && node $HOME/.claude/scripts/verify-citations.mjs "$REVIEWER_FILE" --repo "$WORKTREE" --worktree \
389
393
  && node $HOME/.claude/scripts/finding-fingerprint.mjs annotate --in-place "$REVIEWER_FILE"
390
394
  ```
391
395
 
396
+ `verify-citations.mjs` resolves each finding's `file:line` against the checkout,
397
+ not a commit: a round's fix is uncommitted, and HEAD would call it invented.
398
+ Exit 1 takes the single rework below; exit 2 means not a repository.
399
+
392
400
  Progress line: ` → checking validator validate-reviewer ({reviewer})`
393
401
 
394
402
  `finding-fingerprint.mjs` stamps each finding with its cross-round id once the validator passes; an echoed one is kept, and anonymization leaves it intact.
@@ -323,7 +323,7 @@ If user does not respond at the channels multi-select menu within 30 minutes (wa
323
323
  1. Channels command aborts its interactive prompt, returns `{status: "timeout"}`.
324
324
  2. Phase 7 records `phase-tracker.sh sub 7 1 "Channels dispatch" timeout`.
325
325
  3. **Internal capture still runs** - Steps 2 + 3 write `agent-log.md` (with `channels: timeout` in summary), emit telemetry, update knowledge base.
326
- 4. Session exits cleanly. State persisted: `state.phase=7, state.waitingFor="user-channels-choice", state.channelsTimeout=true`.
326
+ 4. Session exits cleanly. State persisted: `state.currentPhase=7, state.waitingFor="user-channels-choice", state.channelsTimeout=true`.
327
327
  5. Resume contract: `/multi-agent:resume <task-id>` re-opens the channels menu with the original state bundle (pipeline log, PR metadata, prefs pre-ticks).
328
328
 
329
329
  Rationale: silently posting defaults to Jira / Confluence after a timeout would leak wrong-tone content to external systems. Hard-stop + resume is the safer policy.
@@ -90,7 +90,7 @@ Two channels run in parallel at every phase boundary. Both are required in their
90
90
 
91
91
  ### Tracker bootstrap (Phase 0, mandatory)
92
92
 
93
- Phase 0 MUST initialize the tracker and register all 8 phases:
93
+ Phase 0 MUST initialize the tracker and register the active mode's phase set:
94
94
 
95
95
  ```bash
96
96
  $HOME/.claude/scripts/phase-tracker.sh init "$TASK_ID"
@@ -101,6 +101,11 @@ done
101
101
 
102
102
  This produces an initial card stack printed by both CLIs.
103
103
 
104
+ `/multi-agent` and `/multi-agent:local` are the exception: they do not know their
105
+ set here, because depth decides it at Step 7.5. They register `0:Init` alone, then
106
+ the rest once the answer lands, and call `tiles --new` for the second batch. Full
107
+ contract: `tracker-contract.md`, "Deferred registration".
108
+
104
109
  ### Tracker updates (every phase boundary)
105
110
 
106
111
  As each phase enters/exits:
@@ -162,7 +167,7 @@ TaskUpdate({ taskId: <saved>, status: "completed" })
162
167
  bash phase-tracker.sh update <N> completed
163
168
  ```
164
169
 
165
- A phase outside the command's set gets no TaskCreate at all. Depth is different: it is not known at registration time, because the tracker boots at Step -1 and the depth question runs at Step 7.5, so a Short run registers Phases 1 and 2 like any other and flips them to `skipped` when the answer lands. Pre-marking them before Phase 0 produces visually scrambled tile stacks - see the ordering rule below.
170
+ A phase outside the command's set gets no TaskCreate at all. Depth is different: it is not known at registration time, because the tracker boots at Step -1 and the depth question runs at Step 7.5. So registration splits - Phase 0 alone at Step -1, the rest at Step 7.5 once the answer says which phases the run has. A Short run never draws an Analysis tile it will not use. Registering all eight and flipping 1 and 2 to `skipped` is what this replaced, in v17.5.0: it put an eight-tile widget on screen beside the question asking whether to run two of them - see the ordering rule below.
166
171
 
167
172
  **(strict) TaskCreate ordering**: All TaskCreate calls MUST fire in strict phase-number order BEFORE any TaskUpdate is applied. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks (e.g. `1 ✓ · 2 ✓ · 4 ✓ · 0 ▶ · 3 ☐`) even when the underlying state is correct. Pre-marking phases as completed/skipped before Phase 0 starts is FORBIDDEN - register the tile in order with default `pending` status, then flip status via TaskUpdate when the phase actually short-circuits. Full contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
168
173
 
@@ -8,6 +8,7 @@
8
8
  - [Localized labels: what the caller owns](#localized-labels-what-the-caller-owns)
9
9
  - [Order: project, then repo, then branch](#order-project-then-repo-then-branch)
10
10
  - [A single candidate is still a question](#a-single-candidate-is-still-a-question)
11
+ - [Two options or it is not a question](#two-options-or-it-is-not-a-question)
11
12
  - [Autopilot / non-interactive contract](#autopilot-non-interactive-contract)
12
13
  - [Deterministic gates note](#deterministic-gates-note)
13
14
  <!-- /toc -->
@@ -102,15 +103,46 @@ input with a second question.
102
103
  ## A single candidate is still a question
103
104
 
104
105
  The number of options never authorises a skip. A filter that leaves one row has
105
- narrowed the world; it has not decided anything, and the host's **Other** row is a real
106
- choice on every picker - a branch the filter excluded, an account the probe missed, a
107
- repo git does not know about. "There was only one option, so I picked it" is a skipped
108
- picker, and announcing the pick in prose first is the same skip with a sentence in front
109
- of it.
106
+ narrowed the world; it has not decided anything. "There was only one option, so I
107
+ picked it" is a skipped picker, and announcing the pick in prose first is the same
108
+ skip with a sentence in front of it.
110
109
 
111
110
  This is the failure that is hardest to see afterwards, because the transcript reads like
112
111
  a decision was made. Only the picker's absence records that the user was never asked.
113
112
 
113
+ The one row is asked by giving the question a real second option, not by sending a
114
+ one-row picker - see the next section for why that is not the same thing.
115
+
116
+ ## Two options or it is not a question
117
+
118
+ Claude Code's `AskUserQuestion` refuses a question with fewer than two declared
119
+ options, and it refuses the **whole call**: every other question batched with it is
120
+ discarded unasked, and the host's reply says not to retry and not to invent a filler
121
+ option. The host's **Other** row does not rescue it - the schema counts declared
122
+ options, and Other is injected afterwards.
123
+
124
+ This has already cost a run. A base-branch question with one remote candidate was
125
+ batched with the maturity, depth and workspace questions; the call was rejected, the
126
+ branch was announced in prose instead ("only candidate, continuing with it"), and the
127
+ three surviving questions had to be re-asked. The section above was followed to the
128
+ letter and produced the skip it exists to prevent.
129
+
130
+ So a one-candidate picker is asked with a genuine escape as its second option:
131
+
132
+ | One candidate | Second option that makes it a question |
133
+ |---|---|
134
+ | base branch | "Pick another branch" - re-opens with the unfiltered `git branch -r` list |
135
+ | account, repo, module | "Show all" - re-opens with the filter dropped |
136
+ | a destructive step | "Abort" |
137
+
138
+ A second option that is not a choice ("OK", "Continue") is worse than not asking: it
139
+ manufactures consent. When nothing genuine can be offered, do not ask - say which
140
+ single path is being taken and continue.
141
+
142
+ `ask-choice.sh` accepts one option and always has, so this floor is a Claude Code
143
+ fact the shell path does not share. Write the picker to the floor anyway: one spec
144
+ text drives both hosts.
145
+
114
146
  ## Autopilot / non-interactive contract
115
147
 
116
148
  In autopilot, `ask_choice` resolves to `default` (or the safe first option) without prompting - identical to how the native gates auto-proceed today. A picker is only surfaced for genuinely ambiguous or destructive decisions, matching the maturity-check model.
@@ -94,7 +94,7 @@ Phases by mode:
94
94
  | `/multi-agent:autopilot`, `/multi-agent:local-autopilot` | 0,1,2,3,4,6,7 (always Full; autopilot drops the interactive Phase 5 gate) |
95
95
  | `/multi-agent:analysis` | 0,1,2,4,6,7 (no code is written, so no Dev and no Test) |
96
96
 
97
- What changed in v16.0.0: the two picker entries (`/multi-agent`, `:local`) register their FULL set even when the run turns out to be Short, and there is a timing reason. The tracker boots at Step -1, the first thing in every run, while the depth picker cannot run before Step 7.5 - its recommendation needs `taskType`, which needs the fetched issue and the branch. So Phases 1 and 2 are registered `pending` like any other and flipped to `skipped` at 7.5 if Short is chosen. See "Late skip" below.
97
+ The two picker entries (`/multi-agent`, `:local`) register in two batches, because at Step -1 they do not yet know which set is theirs: see "Deferred registration" below. Every other mode registers its whole set at Step -1.
98
98
 
99
99
  Register each phase:
100
100
 
@@ -199,29 +199,40 @@ Mode-specific phase sets:
199
199
 
200
200
  | Mode | TaskCreate set (in order) |
201
201
  |---|---|
202
- | `/multi-agent` | 0 → 1 → 2 → 3 → 4 → 5 → 6 → 7 (all 8; 1 and 2 flip to skipped at Step 7.5 if the user picks Short) |
203
- | `:local` | 0 → 1 → 2 → 3 → 4 → 6 → 7 (same late skip; Phase 5 is not in the set at all) |
202
+ | `/multi-agent` | 0 at Step -1; then at Step 7.5 either 1 → 2 → 3 → 4 → 5 → 6 → 7 (Full) or 3 → 4 → 5 → 6 → 7 (Short) |
203
+ | `:local` | same two batches, with Phase 5 in neither |
204
204
  | `:autopilot`, `:local-autopilot` | 0 → 1 → 2 → 3 → 4 → 6 → 7 (7 phases - always Full, and the interactive Phase 5 gate is dropped) |
205
205
  | `:analysis` | 0 → 1 → 2 → 4 → 6 → 7 (6 phases - no code is written, so 3 and 5 are not in the set) |
206
206
 
207
207
  A phase outside the mode's set gets no TaskCreate at all; the `[SKIPPED]` pattern applies only to a phase that IS in the set and short-circuits at runtime. Phase 4 is in every mode's set as of v14.0.0. The authoritative per-mode set is the `for p in ...` init block in each mode's own entry doc, generated by `gen-mode-dispatch.mjs`; this table mirrors those blocks.
208
208
 
209
- #### Late skip - the depth picker
209
+ #### Deferred registration - the depth picker
210
210
 
211
- The two picker entries cannot know their phase set at Step -1, and the ordering rule above forbids pre-marking. The contract already has the answer, and it is the only permitted one: register the tile in order with the default `pending` status, then flip it when the phase actually short-circuits.
211
+ `/multi-agent` and `:local` cannot know their phase set at Step -1. Depth decides it, and the depth picker cannot run before Step 7.5: its recommendation needs `taskType`, which needs the fetched issue and the branch.
212
+
213
+ Until v17.5.0 they registered all eight anyway and flipped 1 and 2 to `skipped` at 7.5. That put a widget reading "8 tasks, 7 open - Phase 1 Analysis, Phase 2 Planning, ..." on screen *beside* the question asking whether to run Analysis and Planning at all, and a Short answer then contradicted a list the user had just been shown. The widget was asserting a shape the run had not chosen.
214
+
215
+ So registration splits at the moment the shape is known:
212
216
 
213
217
  ```text
214
- # Step -1, before anything else: all eight, in order, all pending
215
- TaskCreate(Phase 0) ... TaskCreate(Phase 7)
216
-
217
- # Step 7.5, after the depth answer. Short only:
218
- TaskUpdate(taskId₁, status="completed", activeForm="[SKIPPED]")
219
- TaskUpdate(taskId₂, status="completed", activeForm="[SKIPPED]")
220
- bash $HOME/.claude/scripts/phase-tracker.sh update 1 skipped
221
- bash $HOME/.claude/scripts/phase-tracker.sh update 2 skipped
218
+ # Step -1, first thing in the run: Phase 0 only. It is the one phase that is
219
+ # certain, and the run is never silent while Phase 0 does its work.
220
+ bash $HOME/.claude/scripts/phase-tracker.sh add 0 Init
221
+ bash $HOME/.claude/scripts/phase-tracker.sh tiles # -> TaskCreate(Phase 0)
222
+ bash $HOME/.claude/scripts/phase-tracker.sh update 0 in_progress
223
+
224
+ # Step 7.5, immediately after the depth answer:
225
+ # Full -> 1 2 3 4 5 6 7 Short -> 3 4 5 6 7
226
+ # :local drops 5 from either (no worktree to check out from)
227
+ for p in "3:Dev" "4:Review" "5:Test" "6:Commit" "7:Report"; do
228
+ bash $HOME/.claude/scripts/phase-tracker.sh add "${p%%:*}" "${p#*:}"
229
+ done
230
+ bash $HOME/.claude/scripts/phase-tracker.sh tiles --new # -> TaskCreate for the new tiles only
222
231
  ```
223
232
 
224
- Order holds because every tile was created before any update. Nothing is pre-marked: at creation time the run genuinely does not know, and the flip happens at the moment it learns.
233
+ `tiles --new` emits `TaskCreate` only for phases that carry no `tasklist_id` yet, so the Phase 0 tile is not created twice. It is the same ordering rule, applied per batch: every tile in a batch is created in ascending phase order, and a deferred batch only ever appends phases numbered above everything already registered. Nothing is pre-marked, and a phase the run will not execute never gets a tile at all.
234
+
235
+ A phase that IS registered and short-circuits later still flips with `[SKIPPED]` - autopilot suppressing Phase 5, for instance. That is a runtime outcome, not an unknown set.
225
236
 
226
237
  **Enforcement**: `smoke-tasklist-ordering.sh` scans the dispatcher (`commands/multi-agent/SKILL.md`) and every mode entry point doc (`commands/multi-agent/{autopilot,local,local-autopilot,analysis,resume-local}/SKILL.md` + the Copilot full-inline orchestrator mirror) for the explicit "in phase-number order" rule. Inventory drift fails the smoke.
227
238
 
@@ -76,10 +76,94 @@
76
76
  "type": "string",
77
77
  "description": "PR target branch (e.g. develop, main)."
78
78
  },
79
+ "baseFetchStatus": {
80
+ "type": "string",
81
+ "enum": ["fresh", "cached-stale", "local-branch", "aborted"],
82
+ "description": "What the base ref is worth. fresh = git fetch origin succeeded; cached-stale = the fetch failed and the user chose the remote-tracking cache; local-branch = the fetch failed and the user chose the local branch; aborted = the user stopped the run at the fetch-fail picker. phase0-exit-gate.mjs has required this field since v17.0, while this schema forbade it under additionalProperties: false - so a state that satisfied the gate failed validation and vice versa. Declared here as of v17.5.0."
83
+ },
79
84
  "baseBranchSource": {
80
85
  "type": "string",
81
- "enum": ["asked", "input", "remembered", "default"],
82
- "description": "How baseBranch was decided. asked = the user answered the Step 3 picker; input = it arrived with the task reference; remembered = autopilot took the most recent entry in prefs.global.recentBranches still inside the TTL and still on the remote; default = autopilot fell back to the develop/release/main sort order. An autopilot run cannot be asked anything, so recording which rule fired is what keeps it readable afterwards."
86
+ "enum": ["asked", "input", "remembered", "default", "derived"],
87
+ "description": "How baseBranch was decided. asked = the user answered the Step 3 picker; input = it arrived with the task reference; remembered = autopilot took the most recent entry in prefs.global.recentBranches still inside the TTL and still on the remote; default = autopilot fell back to the develop/release/main sort order; derived = autopilot took the top base-branch-candidates.mjs candidate, which carried issue-version or linked-release evidence and tied with nothing. The last three are autopilot resolutions: an autopilot run cannot be asked anything, so recording which rule fired is what keeps it readable afterwards, and an interactive run that records one has skipped its picker. A derived branch that a human then confirmed is still asked - the derivation is recorded in baseBranchEvidence, not in this field."
88
+ },
89
+ "baseBranchEvidence": {
90
+ "type": "object",
91
+ "additionalProperties": false,
92
+ "description": "v17.5.0+ - what Step 3 knew when it chose the base branch (refs/features/base-branch-evidence.md). Required when baseBranchSource is derived, and whenever baseFetchStatus is cached-stale or local-branch: a run may degrade to local refs, it may not report a local-only list as the remote's answer.",
93
+ "required": ["refProvenance"],
94
+ "properties": {
95
+ "refProvenance": {
96
+ "type": "string",
97
+ "enum": ["remote", "local"],
98
+ "description": "Where the candidate ref list came from. local means the fetch failed and the list is the local cache plus local heads - possibly stale, possibly incomplete."
99
+ },
100
+ "chosen": { "type": "string" },
101
+ "ambiguous": {
102
+ "type": "boolean",
103
+ "description": "Two or more candidates tied at the top score. Autopilot may not record derived when this is true."
104
+ },
105
+ "convention": {
106
+ "type": ["object", "null"],
107
+ "additionalProperties": false,
108
+ "description": "The release-branch template inferred from the refs that exist, never from a built-in table.",
109
+ "properties": {
110
+ "template": { "type": "string" },
111
+ "members": { "type": "integer", "minimum": 0 }
112
+ }
113
+ },
114
+ "candidates": {
115
+ "type": "array",
116
+ "items": {
117
+ "type": "object",
118
+ "additionalProperties": false,
119
+ "required": ["branch"],
120
+ "properties": {
121
+ "branch": { "type": "string" },
122
+ "score": { "type": "number" },
123
+ "refs": { "type": "array", "items": { "type": "string" } },
124
+ "evidence": {
125
+ "type": "array",
126
+ "items": {
127
+ "type": "object",
128
+ "additionalProperties": false,
129
+ "required": ["kind"],
130
+ "properties": {
131
+ "kind": {
132
+ "type": "string",
133
+ "enum": [
134
+ "issue-version",
135
+ "linked-release",
136
+ "version-convention",
137
+ "recent",
138
+ "repo-default",
139
+ "sort-order",
140
+ "ref-provenance"
141
+ ]
142
+ },
143
+ "detail": { "type": "string" }
144
+ }
145
+ }
146
+ }
147
+ }
148
+ }
149
+ },
150
+ "notes": { "type": "array", "items": { "type": "string" } },
151
+ "askedOnIssue": {
152
+ "type": ["object", "null"],
153
+ "additionalProperties": false,
154
+ "description": "The one comment autopilot is allowed to post when the derivation is ambiguous, gated by prefs.global.baseBranchEvidence.autopilotAsksOnIssue (default false). A question, never a state change: no transition, no close, no assignee. Posting it trips circuit-breaker trigger 6 and the run waits for resume.",
155
+ "properties": {
156
+ "target": { "type": "string" },
157
+ "url": { "type": "string" },
158
+ "at": { "type": "string", "format": "date-time" }
159
+ }
160
+ }
161
+ }
162
+ },
163
+ "workspaceSource": {
164
+ "type": "string",
165
+ "enum": ["asked", "command", "autopilot"],
166
+ "description": "Who decided where the branch lives. asked = the user answered the Step 5b workspace picker; command = :local / --local / :local-autopilot stated it up front, or a flow that only ever builds worktrees; autopilot = resolved to a worktree without asking, because an unattended commit in the user's own checkout is what worktrees prevent. localMode alone cannot say: false is both a chosen worktree and one nothing asked about."
83
167
  },
84
168
  "remoteType": {
85
169
  "type": "string",
@@ -883,7 +967,7 @@
883
967
  "circuitBreaker": {
884
968
  "type": "object",
885
969
  "additionalProperties": false,
886
- "description": "Autopilot circuit-breaker record (refs/features/autopilot-circuit-breaker.md). Written only when a trigger trips: trigger 2 by Phase 4 Step 3.8 (a mandate finding survived identicalFindingCycles rework cycles), trigger 3 by the Phase 3 re-entry hard-kill. /multi-agent:resume clears tripped and keeps counters.",
970
+ "description": "Autopilot circuit-breaker record (refs/features/autopilot-circuit-breaker.md). Written only when a trigger trips: trigger 2 by Phase 4 Step 3.8 (a mandate finding survived identicalFindingCycles rework cycles), trigger 3 by the Phase 3 re-entry hard-kill, trigger 6 by Phase 0 Step 3 when autopilot posted the base-branch question on the issue and must not answer it itself. /multi-agent:resume clears tripped and keeps counters.",
887
971
  "required": ["tripped"],
888
972
  "properties": {
889
973
  "tripped": {
@@ -892,7 +976,7 @@
892
976
  "trigger": {
893
977
  "type": ["integer", "null"],
894
978
  "minimum": 1,
895
- "maximum": 5
979
+ "maximum": 6
896
980
  },
897
981
  "detail": {
898
982
  "type": "string"
@@ -969,6 +969,28 @@
969
969
  }
970
970
  }
971
971
  },
972
+ "baseBranchEvidence": {
973
+ "type": "object",
974
+ "additionalProperties": false,
975
+ "description": "v17.5.0+ - Phase 0 Step 3 base-branch evidence collection (refs/features/base-branch-evidence.md). Candidates are collected with the evidence behind them (issue version field, linked release issue, the branch convention learned from the refs that exist, recent branches, repo default), ranked, and shown. Interactive runs always ask; the evidence only reorders the rows.",
976
+ "properties": {
977
+ "enabled": {
978
+ "type": "boolean",
979
+ "default": true,
980
+ "description": "Collect issue-derived candidates at all. Off falls back to the develop/release/main sort order, which is still surfaced through the picker."
981
+ },
982
+ "preferLinkedRelease": {
983
+ "type": "boolean",
984
+ "default": false,
985
+ "description": "On boards where opening a development sub-task requires selecting the related release issue, that link is the authoritative base-branch answer and the version field is corroboration. Raises the linked-release weight above the version-field one rather than adding a second rule."
986
+ },
987
+ "autopilotAsksOnIssue": {
988
+ "type": "boolean",
989
+ "default": false,
990
+ "description": "OFF by default, and it is an outward-facing write. When on, an autopilot run whose base-branch derivation is ambiguous posts ONE comment on the Jira or GitHub issue asking which branch to develop from, then halts on circuit-breaker trigger 6 and waits for resume. A question, never a state change: no transition, no resolution, no assignee, no close. The body uses Ref:, never Closes:/Fixes:/Resolves:, and its human-facing copy follows outputLanguage. Autopilot never posts the question and then answers it itself."
991
+ }
992
+ }
993
+ },
972
994
  "autopilotCircuitBreaker": {
973
995
  "type": "object",
974
996
  "additionalProperties": false,
@@ -0,0 +1,137 @@
1
+ {
2
+ "$comment": "Not a JSON Schema: the data the review file filter reads. Lives beside token-budget.json for the same reason - it is a budget-shaped decision the pipeline owns, versioned with the code that consumes it.",
3
+ "version": 1,
4
+ "patterns": [
5
+ {
6
+ "glob": "**/*.generated.*",
7
+ "reason": "generated file - review the generator, not its output"
8
+ },
9
+ {
10
+ "glob": "**/Generated/**",
11
+ "reason": "generated tree - review the generator, not its output"
12
+ },
13
+ {
14
+ "glob": "**/generated/**",
15
+ "reason": "generated tree - review the generator, not its output"
16
+ },
17
+ { "glob": "**/*.pb.go", "reason": "generated file - review the generator, not its output" },
18
+ { "glob": "**/*_pb2.py", "reason": "generated file - review the generator, not its output" },
19
+
20
+ {
21
+ "glob": "**/package-lock.json",
22
+ "reason": "lockfile - resolved by the package manager, not written by hand"
23
+ },
24
+ {
25
+ "glob": "**/yarn.lock",
26
+ "reason": "lockfile - resolved by the package manager, not written by hand"
27
+ },
28
+ {
29
+ "glob": "**/pnpm-lock.yaml",
30
+ "reason": "lockfile - resolved by the package manager, not written by hand"
31
+ },
32
+ {
33
+ "glob": "**/Package.resolved",
34
+ "reason": "lockfile - resolved by the package manager, not written by hand"
35
+ },
36
+ {
37
+ "glob": "**/Podfile.lock",
38
+ "reason": "lockfile - resolved by the package manager, not written by hand"
39
+ },
40
+ {
41
+ "glob": "**/Gemfile.lock",
42
+ "reason": "lockfile - resolved by the package manager, not written by hand"
43
+ },
44
+ {
45
+ "glob": "**/poetry.lock",
46
+ "reason": "lockfile - resolved by the package manager, not written by hand"
47
+ },
48
+ {
49
+ "glob": "**/Cargo.lock",
50
+ "reason": "lockfile - resolved by the package manager, not written by hand"
51
+ },
52
+ {
53
+ "glob": "**/gradle.lockfile",
54
+ "reason": "lockfile - resolved by the package manager, not written by hand"
55
+ },
56
+
57
+ {
58
+ "glob": "**/__snapshots__/**",
59
+ "reason": "recorded snapshot - the assertion is the test that records it"
60
+ },
61
+ {
62
+ "glob": "**/*.snap",
63
+ "reason": "recorded snapshot - the assertion is the test that records it"
64
+ },
65
+ {
66
+ "glob": "**/__Snapshots__/**",
67
+ "reason": "recorded snapshot - the assertion is the test that records it"
68
+ },
69
+ {
70
+ "glob": "**/ReferenceImages/**",
71
+ "reason": "recorded snapshot - the assertion is the test that records it"
72
+ },
73
+
74
+ {
75
+ "glob": "vendor/**",
76
+ "reason": "vendored third-party source - not ours to change in this diff"
77
+ },
78
+ {
79
+ "glob": "**/node_modules/**",
80
+ "reason": "vendored third-party source - not ours to change in this diff"
81
+ },
82
+ {
83
+ "glob": "Pods/**",
84
+ "reason": "vendored third-party source - not ours to change in this diff"
85
+ },
86
+ {
87
+ "glob": "**/third_party/**",
88
+ "reason": "vendored third-party source - not ours to change in this diff"
89
+ },
90
+
91
+ {
92
+ "glob": "**/*.min.js",
93
+ "reason": "build output - unreadable by design, and the source is elsewhere in the diff"
94
+ },
95
+ {
96
+ "glob": "**/*.min.css",
97
+ "reason": "build output - unreadable by design, and the source is elsewhere in the diff"
98
+ },
99
+ {
100
+ "glob": "**/*.map",
101
+ "reason": "build output - unreadable by design, and the source is elsewhere in the diff"
102
+ },
103
+ {
104
+ "glob": "dist/**",
105
+ "reason": "build output - unreadable by design, and the source is elsewhere in the diff"
106
+ },
107
+ {
108
+ "glob": "build/**",
109
+ "reason": "build output - unreadable by design, and the source is elsewhere in the diff"
110
+ },
111
+ {
112
+ "glob": ".build/**",
113
+ "reason": "build output - unreadable by design, and the source is elsewhere in the diff"
114
+ },
115
+ {
116
+ "glob": "**/DerivedData/**",
117
+ "reason": "build output - unreadable by design, and the source is elsewhere in the diff"
118
+ },
119
+
120
+ { "glob": "**/*.png", "reason": "binary asset - a text reviewer cannot read it" },
121
+ { "glob": "**/*.jpg", "reason": "binary asset - a text reviewer cannot read it" },
122
+ { "glob": "**/*.jpeg", "reason": "binary asset - a text reviewer cannot read it" },
123
+ { "glob": "**/*.gif", "reason": "binary asset - a text reviewer cannot read it" },
124
+ { "glob": "**/*.webp", "reason": "binary asset - a text reviewer cannot read it" },
125
+ { "glob": "**/*.pdf", "reason": "binary asset - a text reviewer cannot read it" },
126
+ { "glob": "**/*.zip", "reason": "binary asset - a text reviewer cannot read it" },
127
+ { "glob": "**/*.ipa", "reason": "binary asset - a text reviewer cannot read it" },
128
+ { "glob": "**/*.apk", "reason": "binary asset - a text reviewer cannot read it" },
129
+ { "glob": "**/*.aab", "reason": "binary asset - a text reviewer cannot read it" },
130
+ { "glob": "**/*.mp4", "reason": "binary asset - a text reviewer cannot read it" },
131
+ { "glob": "**/*.mov", "reason": "binary asset - a text reviewer cannot read it" },
132
+ { "glob": "**/*.ttf", "reason": "binary asset - a text reviewer cannot read it" },
133
+ { "glob": "**/*.otf", "reason": "binary asset - a text reviewer cannot read it" },
134
+ { "glob": "**/*.woff", "reason": "binary asset - a text reviewer cannot read it" },
135
+ { "glob": "**/*.woff2", "reason": "binary asset - a text reviewer cannot read it" }
136
+ ]
137
+ }