@mmerterden/multi-agent-pipeline 16.28.0 → 16.30.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (52) hide show
  1. package/CHANGELOG.md +119 -2
  2. package/README.md +4 -4
  3. package/README.tr.md +3 -3
  4. package/docs/architecture.md +3 -3
  5. package/docs/ecosystem.md +5 -5
  6. package/docs/features.md +14 -0
  7. package/install/claude.mjs +17 -0
  8. package/package.json +1 -1
  9. package/pipeline/commands/multi-agent/analysis-jira/SKILL.md +93 -0
  10. package/pipeline/commands/multi-agent/design-check/SKILL.md +6 -5
  11. package/pipeline/commands/multi-agent/doctor/SKILL.md +78 -0
  12. package/pipeline/commands/multi-agent/help/SKILL.md +15 -12
  13. package/pipeline/commands/multi-agent/manual-test/SKILL.md +1 -1
  14. package/pipeline/commands/multi-agent/setup/SKILL.md +14 -1
  15. package/pipeline/commands/multi-agent/sync/SKILL.md +12 -9
  16. package/pipeline/commands/multi-agent/update/SKILL.md +12 -0
  17. package/pipeline/lib/_jira-auth.sh +99 -0
  18. package/pipeline/lib/analysis-jira-write.sh +203 -0
  19. package/pipeline/lib/issue-fetcher.sh +4 -4
  20. package/pipeline/multi-agent-refs/analysis/render.md +1 -1
  21. package/pipeline/multi-agent-refs/channels/pr.md +37 -1
  22. package/pipeline/multi-agent-refs/cross-cli-contract.md +3 -3
  23. package/pipeline/multi-agent-refs/features/analysis-jira.md +128 -0
  24. package/pipeline/multi-agent-refs/features/doctor.md +197 -0
  25. package/pipeline/multi-agent-refs/features/model-fallback.md +2 -2
  26. package/pipeline/multi-agent-refs/features/visual-evidence.md +103 -20
  27. package/pipeline/multi-agent-refs/phases/phase-0-init.md +38 -7
  28. package/pipeline/multi-agent-refs/phases/phase-3-dev.md +13 -1
  29. package/pipeline/multi-agent-refs/phases/phase-5-test.md +11 -1
  30. package/pipeline/multi-agent-refs/phases/phase-6-commit.md +23 -0
  31. package/pipeline/multi-agent-refs/picker-contract.md +35 -0
  32. package/pipeline/multi-agent-refs/tracker-contract.md +5 -1
  33. package/pipeline/preferences-template.json +1 -1
  34. package/pipeline/schemas/agent-state.schema.json +84 -1
  35. package/pipeline/schemas/analysis-spec.schema.json +336 -95
  36. package/pipeline/schemas/prefs.schema.json +80 -3
  37. package/pipeline/schemas/token-budget.json +10 -10
  38. package/pipeline/scripts/analysis-story-tree.mjs +441 -0
  39. package/pipeline/scripts/capture-evidence.sh +170 -5
  40. package/pipeline/scripts/doctor.mjs +758 -0
  41. package/pipeline/scripts/evidence-gate.mjs +31 -2
  42. package/pipeline/scripts/phase-tracker.sh +97 -17
  43. package/pipeline/scripts/probe-evidence-capability.sh +250 -0
  44. package/pipeline/scripts/run-ui-tests.sh +380 -0
  45. package/pipeline/scripts/scan-agent-config.sh +48 -10
  46. package/pipeline/scripts/skill-siblings.mjs +1 -1
  47. package/pipeline/skills/shared/core/multi-agent-analysis-jira/SKILL.md +94 -0
  48. package/pipeline/skills/shared/core/multi-agent-doctor/SKILL.md +79 -0
  49. package/pipeline/skills/shared/core/multi-agent-manual-test/SKILL.md +10 -1
  50. package/pipeline/skills/shared/core/multi-agent-setup/SKILL.md +13 -0
  51. package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +9 -6
  52. package/pipeline/skills/shared/core/multi-agent-update/SKILL.md +18 -0
@@ -0,0 +1,197 @@
1
+ # doctor - the check registry
2
+
3
+ Every check `/multi-agent:doctor` can report has a `### <id>` heading here, and
4
+ `doctor.mjs --list-checks` prints exactly the same set. The equality is checked
5
+ in both directions by `smoke-doctor.sh`: a check that ships without an entry
6
+ cannot be released, and an entry whose check was deleted cannot linger as orphan
7
+ prose. That is the mechanism that keeps this file from rotting into a list of
8
+ things the tool used to do.
9
+
10
+ ## What the exit code means
11
+
12
+ | Code | Meaning |
13
+ |---|---|
14
+ | 0 | healthy - nothing above INFO |
15
+ | 1 | degraded - at least one WARN, no BLOCK |
16
+ | 2 | blocked - at least one BLOCK |
17
+ | 3 | usage error |
18
+ | 4 | indeterminate - the installed layout could not be resolved, so nothing was checked |
19
+
20
+ 4 exists so that "I could not look" never borrows the exit code of "I looked and
21
+ it is fine". A consumer that treats 0 as healthy would otherwise read a broken
22
+ resolver as a clean bill.
23
+
24
+ ## What BLOCK means, exactly
25
+
26
+ **A state in which a run will fail or leak. Not a state in which it will merely
27
+ be worse.** Without a closed definition "blocked" grows until nobody respects
28
+ exit 2, so only five checks can produce it: `install-present`, `script-surface`,
29
+ `state-writable`, `prefs-valid` and `embedded-credentials`. Every other check
30
+ tops out at WARN however bad it looks.
31
+
32
+ ## Who calls it, and what they do with the code
33
+
34
+ | Caller | When | On exit 2 |
35
+ |---|---|---|
36
+ | `/multi-agent:setup` | before the first question, and again at the end | the closing run is the evidence for or against "setup complete" |
37
+ | `/multi-agent:update` | after the install | report that the update left a state a run will fail from, rather than "updated" |
38
+ | `/multi-agent:sync` | step 0 | **stop.** Syncing a blocked install copies one fault onto five surfaces, and the copies are what people then debug |
39
+
40
+ `/multi-agent:sync` stops on exit 4 as well: the layout did not resolve, so
41
+ nothing was checked, and syncing from an unknown state is worse than not syncing
42
+ at all. Exit 1 never stops anything - a degraded install still syncs correctly,
43
+ and a gate that blocked on every warning would be routed around within a week.
44
+
45
+ ## The four severities
46
+
47
+ | Severity | Meaning |
48
+ |---|---|
49
+ | BLOCK | a run will fail or leak |
50
+ | WARN | a run will work and be worse: a degraded capability, a stale copy, a rejected credential |
51
+ | INFO | a capability the user chose not to enable; nothing to fix |
52
+ | SKIP | not checked, and why |
53
+
54
+ `SKIP` is printed in the same list, in the same position, and the summary always
55
+ carries its count (`... 2 not checked, 15 checks`). Absence is a finding: a run
56
+ without `--probe` exits 0 **and** prints `SKIP credential-liveness`, so "healthy"
57
+ and "not looked at" are never the same output.
58
+
59
+ The INFO / WARN boundary is already decided in `lib/credential-inventory.sh` and
60
+ this tool only renders it: an unmapped key is INFO because it is a capability the
61
+ user chose not to enable, while `auth-rejected`, `malformed` and
62
+ `tier-1-no-grant` are WARN because the configuration made a claim the service
63
+ refused. **A missing optional capability is information; a mapped credential the
64
+ service rejected is a warning. The difference is whether the configuration
65
+ asserted anything.**
66
+
67
+ ## The line shape
68
+
69
+ ```
70
+ <SEV> <id> - <what is wrong> - <single imperative step>
71
+ ```
72
+
73
+ The third field is one step, and its first word comes from a closed list:
74
+ `run`, `set`, `map`, `revoke`, `install`, `remove`, `free`, `export`, `merge`,
75
+ `record`, `re-run`. `consider`, `may want` and `should probably` are refused by
76
+ the gate. One problem, one step: a reader with five suggestions does nothing.
77
+
78
+ ## It recommends, it never fixes
79
+
80
+ Nothing here rewrites a remote URL, edits `settings.json` or touches a token.
81
+ A remote whose URL carries a token may be the only credential that repo has, the
82
+ remote may be a mirror a script depends on verbatim, and doctor can run inside a
83
+ checkout the user does not own. Silent repair breaks all three.
84
+
85
+ For an embedded credential the honest single step is **not** "hide it". A token
86
+ that reached `.git/config` is already burned: it is in the shell history and
87
+ readable by anything that can read the working tree. The step is to revoke it at
88
+ its host; everything after that is behind `--explain`.
89
+
90
+ ---
91
+
92
+ ## Checks
93
+
94
+ ### install-present
95
+
96
+ The installed tree exists and carries the subtrees a run reads: `commands/`,
97
+ `multi-agent-refs/`, `scripts/`, `lib/`, `schemas/`. BLOCK when any is missing -
98
+ a run cannot start without them.
99
+
100
+ ### install-version
101
+
102
+ The installed `.pipeline-version` matches the package version resolved from the
103
+ checkout or the published package. WARN on a mismatch: the run works, it just
104
+ is not the version the user thinks they have.
105
+
106
+ ### script-surface
107
+
108
+ Every script a shipped command names by path exists in the installed tree. BLOCK:
109
+ the failure lands mid-run, at the call, with the phase already half done.
110
+
111
+ ### skill-siblings
112
+
113
+ Every command has its copies in the host trees this machine installs. WARN, never
114
+ BLOCK: an unsynced host is a fact about the machine, and `skill-siblings.mjs`
115
+ itself reports the installed layout as "authored side not checked" rather than
116
+ claiming a repo verdict it cannot reach.
117
+
118
+ ### state-writable
119
+
120
+ `$HOME/.claude/logs/multi-agent` exists or can be created, and a file can be
121
+ written there. BLOCK: a run that cannot record its state cannot be resumed,
122
+ reported or priced, and it discovers this at Phase 0 after the pickers.
123
+
124
+ ### prefs-valid
125
+
126
+ `multi-agent-preferences.json` parses and satisfies `schemas/prefs.schema.json`.
127
+ BLOCK: Phase 0 reads it before anything else, and a malformed file fails the run
128
+ after the user has already answered the pickers.
129
+
130
+ ### identity
131
+
132
+ A git identity resolves for the account the run would commit as. WARN: the run
133
+ reaches Phase 6 and stops there.
134
+
135
+ ### hook-coverage
136
+
137
+ The blocking `PreToolUse` gates from `templates/claude-hooks.json` are present in
138
+ `settings.json`. WARN. SKIP when the template is not installed - which was the
139
+ permanent state until the installer began copying `templates/`, and is exactly
140
+ the case this severity exists to make visible.
141
+
142
+ ### credential-mapping
143
+
144
+ Which logical keys are mapped, and whether each mapped value is well formed.
145
+ Unmapped is INFO. Malformed is WARN.
146
+
147
+ ### credential-liveness
148
+
149
+ One cheap authenticated request per mapped credential. WARN on `auth-rejected`,
150
+ `tier-1-no-grant` or `unreachable` - but with **different steps**, because a
151
+ refused credential and a host that never answered need different actions.
152
+ `credential-inventory.sh` already draws that line ("unreachable - no response at
153
+ all - on a corporate host, almost always the VPN"), and telling someone to
154
+ re-onboard a token that was never the problem is the exact failure the keychain
155
+ rules warn about one layer up. **SKIP unless `--probe` is passed**, and the
156
+ skip is printed, because a network check is the user's decision to spend.
157
+
158
+ ### embedded-credentials
159
+
160
+ A token in a git remote URL or a tracked config file. BLOCK: this one leaks
161
+ rather than fails.
162
+
163
+ The report carries the repo path, the config key, the **host**, a shape label and
164
+ a length bucket. Never the value, and not the prefix either: `ghp_` plus a length
165
+ is already a fingerprint, and this line is printed to a terminal that is often
166
+ shared, which is the whole reason the check exists. The URL is parsed inside a
167
+ `python3` heredoc reading stdin, so no shell variable ever holds it.
168
+
169
+ ### task-tools
170
+
171
+ Whether this session's model carries `TaskCreate` / `TaskUpdate`. Claude Code
172
+ provides them by default only on Claude 3.x, Opus 4 through 4.7, Sonnet 4 through
173
+ 4.6 and Haiku 4.5; on any newer model they are absent unless the user opts in.
174
+ INFO, with the opt-in as the step.
175
+
176
+ A script cannot answer this - only the agent knows its own tool list - so the
177
+ caller passes `--task-tools=yes|no` and the default is SKIP. Reporting "absent"
178
+ from a script that never looked would be the same defect this check is about.
179
+
180
+ ### mcp-registration
181
+
182
+ Whether `multi-agent-toolkit` is registered as an MCP server. INFO when it is
183
+ not: the tools it provides are optional and a run without them is smaller, not
184
+ wrong.
185
+
186
+ **Registered is not the same as working.** With `--probe` the server is started
187
+ and its tools counted; serving none is WARN. Without `--probe` this reports what
188
+ is configured, and says so. The distinction is not theoretical: a half-extracted
189
+ package in the npx cache left the server dying on `Cannot find module` at
190
+ startup, which the client surfaces only as `CONNECTION_CLOSED`, and a check that
191
+ read the registration and stopped would have called that healthy.
192
+
193
+ ### disk-space
194
+
195
+ Free space on the volume holding `$HOME`. WARN under 2 GB: a worktree plus a
196
+ build is the largest thing a run writes, and ENOSPC mid-run corrupts the state
197
+ file it was writing at the time.
@@ -83,7 +83,7 @@ predate this field still get the second downgrade step):
83
83
  "premiumTierUntil": null,
84
84
  "fallbackModel": "sonnet",
85
85
  "floorModel": "haiku",
86
- "fableEnabled": true,
86
+ "fableEnabled": false,
87
87
  "onDispatchError": true
88
88
  }
89
89
  ```
@@ -91,7 +91,7 @@ predate this field still get the second downgrade step):
91
91
  | Field | Meaning |
92
92
  |---|---|
93
93
  | `enabled` | Master switch. `false` = always dispatch the persona's `preferredModel`, fail loudly on error. |
94
- | `modelFallback.fableEnabled` | Whether the `fable` rung exists at all on Claude Code. `false` starts every `preferredModel: fable` persona on `opus`. See "Turning the fable rung off" below. |
94
+ | `modelFallback.fableEnabled` | Whether the `fable` rung exists at all on Claude Code. **Ships `false`.** `false` starts every `preferredModel: fable` persona on `opus`; set it to `true` to opt the top rung back in. A cost control that defaults to on is not a control. See "Turning the fable rung off" below. |
95
95
  | `premiumTierUntil` | ISO date (`YYYY-MM-DD`) or `null`. When set and today is **after** this date, every `preferredModel` dispatch is downgraded to `fallbackModel` unless the user re-confirms (see Date gate). Use when the top tier is included in a plan only until a known date. |
96
96
  | `fallbackModel` | Target of the first downgrade step. Default `sonnet` (the next tier below opus). |
97
97
  | `modelFallback.floorModel` | Last-resort tier when `fallbackModel` also fails to dispatch. Default `haiku`. Set to the same value as `fallbackModel` (or `null`) to disable the second step and halt after one downgrade. |
@@ -78,32 +78,86 @@ at most 1242px wide, and writes
78
78
  A clean status bar is not cosmetic: without it two captures of the same screen
79
79
  differ by the clock, which makes every "after" look like a change.
80
80
 
81
- ## 4. Video - three tiers, resolved once per repo
81
+ ## 4. Video - the recording rides on a test run
82
82
 
83
- The flow source, in order: a UI test flow written on the ticket wins; otherwise
84
- the Phase 5 test scenarios are the flow. Human-written beats generated, the same
85
- rule the "before" follows.
83
+ The flow video is not recorded on its own. It wraps something that drives the
84
+ screen, and what drives it is chosen by the user at intake, because running a UI
85
+ suite costs minutes and that is their time to spend.
86
86
 
87
- The tier is probed once and stored in `state.visualEvidence.videoTier`:
87
+ ### 4.1 Probe first
88
88
 
89
- | Tier | Condition | Mechanism |
90
- |---|---|---|
91
- | 1 | The repo has a UI test target (XCUITest / Espresso) **and** it can be seeded with test data | Run that test, record the screen for its duration. Most faithful, and it is a test that already runs in CI |
92
- | 2 | A runnable build exists | Drive the flow with `mcp__multi-agent-toolkit__agent_run_steps` and record with `ios_record_video` / `android_record_screen`. **No test target needed** |
93
- | 3 | No runnable build (library / SPM package / backend) | Skip, with the reason recorded |
89
+ `probe-evidence-capability.sh` measures, before the question is asked, what this
90
+ machine and this repo can actually do: the UI test target, the tests matching
91
+ this change, the device, the recorder CLI, and whether the toolkit MCP is
92
+ registered. Detection is delegated to `run-ui-tests.sh detect`, which is also
93
+ what runs the tests, so there is one implementation of the answer.
94
+
95
+ Two rules the probe keeps, and the reason for each:
96
+
97
+ - **Absence carries a reason.** `no booted simulator, but one is available to
98
+ boot` and `no iOS simulator available on this machine` close the same menu row
99
+ and ask the user for completely different things.
100
+ - **Unmeasurable is `null`, never `false`.** With no `adb` on the PATH the probe
101
+ cannot see whether a device is attached; reporting that as "no device" is a
102
+ negative nobody looked for, and it reads exactly like one somebody checked.
94
103
 
95
- Tier 2 is why "can every repo do this" answers yes in practice: any app that can
96
- be launched can be driven. A repo without a UI test target loses fidelity, not
97
- the recording.
104
+ ### 4.2 Then the question
98
105
 
99
- Preferred host is **Phase 5** - the device is already up and a human is present
100
- to confirm the flow is the right one. In a mode without Phase 5, tier 2 runs at
101
- the end of Phase 3 instead. Tier 1 always runs where the test runs.
106
+ Phase 0 Step 7.7 asks the depth, with the options built from the probe. A closed
107
+ option stays on the list carrying its reason; when every option but the first is
108
+ closed, nothing is asked and `testDepthSource` records `forced`. The full menu
109
+ rules are in `phases/phase-0-init.md`.
102
110
 
103
- Duration is capped by `visualEvidence.maxVideoSeconds` (default 60), read via
104
- `capture-evidence.sh limits` so the number lives in one place rather than in two
105
- documents. A flow that needs longer is not a review artefact, it is a debugging
106
- session.
111
+ ### 4.3 Then the recording
112
+
113
+ | `testDepth` | Tier | What drives the screen |
114
+ |---|---|---|
115
+ | `unit+ui` | 1 | The repo's own UI test, selected by the changed files. Most faithful, and it is a test that already runs in CI |
116
+ | `unit+mcp` | 2 | `mcp__multi-agent-toolkit__agent_run_steps` drives the flow. No test target needed |
117
+ | `unit` | 3 | Nothing. No recording, and the reason is the user's own answer |
118
+
119
+ The recorder itself is `capture-evidence.sh video start|stop`, which writes
120
+ `<task-id>[-<label>]-flow.mp4` into the evidence directory. The label is optional
121
+ because one recording per task is the common case, and both renderers cite the
122
+ unlabelled form. It is shell rather than MCP for three reasons: the sibling still capture already shells out,
123
+ a host with no toolkit MCP registered still produces evidence, and Phase 3 and
124
+ Phase 5 are exactly where an MCP call is contested.
125
+
126
+ **iOS UI test targets are not named, they are detected.** A path containing
127
+ `UITests` is not the signal: in the reference app 477 files sit under such a path
128
+ and exactly 2 drive the UI, the rest being snapshot tests that render a view and
129
+ compare pixels without ever launching the app. The signal is `XCUIApplication`,
130
+ the only API that drives another process's interface, which is precisely the
131
+ precondition a flow recording has. On Android the signal is the `androidTest`
132
+ source set, which is instrumentation by definition.
133
+
134
+ **A real app has many candidates** - 17 UI test directories on the reference iOS
135
+ app, 8 instrumentation modules on the Android one. Detection reports the whole
136
+ set and lets the changed files choose; taking the first off a `find` is a guess
137
+ wearing a measurement's clothes.
138
+
139
+ ### 4.4 Re-check before recording
140
+
141
+ The probe runs at intake and the recording happens in Phase 3. A simulator booted
142
+ then can be gone by the time the build goes green, so Phase 3 re-measures the
143
+ device row alone and downgrades the tier if it has to, recording the transition
144
+ (`tier 1 -> 3: simulator no longer booted`). A tier taken from a stale
145
+ measurement is a promise the run cannot keep.
146
+
147
+ ### 4.5 Duration, and the recording that shows nothing
148
+
149
+ The cap is `visualEvidence.maxVideoSeconds` (default 60), read through
150
+ `capture-evidence.sh limits` so one value serves the phase doc and the recorder.
151
+ Android's `screenrecord` has its own ceiling of 180 seconds that no setting can
152
+ lift, so the script clamps to it and says when it did: one preference honoured on
153
+ one platform and silently halved on the other is worse than a stated limit.
154
+
155
+ Both recorders encode on change. A flow over a screen that never moved therefore
156
+ produces a valid two-frame mp4 a fraction of a second long - not a broken file,
157
+ but not evidence of a flow either. `video stop` says so on stderr when the result
158
+ is under a second, and the caller records that as a gap reason. The duration is
159
+ never asserted against wall clock anywhere, because doing so fails a correct
160
+ capture of a static screen.
107
161
 
108
162
  `visualEvidence.enabled` turns the whole feature off - capture, upload, both
109
163
  render sections and the Phase 6 blocker with it.
@@ -125,6 +179,35 @@ Order of attempts, each one recorded:
125
179
 
126
180
  Never silently attach nothing.
127
181
 
182
+ ## 5b. Where the artefacts live - the host
183
+
184
+ Resolved in Phase 6 Step 2.9, recorded as `state.visualEvidence.host`:
185
+
186
+ | Order | Host | Condition | Stills | Video |
187
+ |---|---|---|---|---|
188
+ | 1 | `jira` | `jiraId` present | Jira attachments | Jira attachment |
189
+ | 2 | `github-public` | no Jira, GitHub remote, public repo | `evidence/<task-id>` orphan branch, embedded in the PR body | none |
190
+ | 3 | `github-private` | no Jira, GitHub remote, private repo | same branch, blob permalink in the PR body | none |
191
+ | 4 | `none` | anything else, or `githubHost: off` | not published, gap recorded | none |
192
+
193
+ **GitHub cannot be given a file.** There is no API that attaches an image to an
194
+ issue or a pull request; the web uploader posts to an endpoint that needs a
195
+ browser session, so no token can drive it. The PR body can only point at
196
+ something already hosted, and an orphan branch is the one mechanism a script has
197
+ that neither touches the PR diff nor publishes a release.
198
+
199
+ **A private repo cannot show an inline image.** GitHub renders markdown images
200
+ through its own proxy, which carries no credentials for a private repo, so an
201
+ embedded `raw.githubusercontent.com` URL renders broken for every reader
202
+ including the author. That is worse than a link, because a broken image looks
203
+ like missing evidence. The private variant therefore links rather than embeds.
204
+
205
+ **Video is Jira-only, by decision.** On a GitHub-hosted run none is recorded at
206
+ all. An mp4 behind a blob link is a download rather than something a reviewer
207
+ opens mid-review, and paying minutes of UI-test time for a recording nobody
208
+ watches is worse than saying plainly that there is none. The gap line carries
209
+ that reason.
210
+
128
211
  ## 6. Rendering
129
212
 
130
213
  ### Jira comment (`channels/jira.md`)
@@ -275,11 +275,8 @@ Scan `$HOME` (maxdepth 2) for project markers (`.xcodeproj`, `Package.swift`, `b
275
275
  | grep -v -E '(feature/|bugfix/|fix/|hotfix/|chore/)'
276
276
  ```
277
277
  6. Sort: `develop*` first, then `release/*`, then `main`/`master`. Surface through the
278
- **native picker** per `picker-contract.md` (`AskUserQuestion` on Claude Code,
279
- `ask-choice.sh` on Copilot CLI) - `question` + `description` in `outputLanguage`,
280
- `label` = the branch name verbatim (a proper noun, never translated) with the recent
281
- branch first and marked `(Recommended)`. The ASCII sketch
282
- below is what the options carry, not a menu to print:
278
+ **native picker** per `picker-contract.md`, recent branch first and marked
279
+ `(Recommended)`:
283
280
 
284
281
  ```
285
282
  header: "Base branch"
@@ -298,9 +295,14 @@ nothing surfaced it. `phase0-exit-gate.mjs` now refuses to close Phase 0 unless
298
295
  `agent-state.json` carries `baseBranch` and `baseFetchStatus`, so a skipped picker is a
299
296
  gate failure rather than a silent default.
300
297
 
298
+ One row is a normal outcome of the rule-5 filter and is **still asked** - see
299
+ `picker-contract.md` "A single candidate is still a question".
300
+
301
301
  This holds in every mode. A Short run skips the *LLM* phases (Analysis, Planning); it
302
- does not skip Phase 0's pickers. Autopilot resolves them to their defaults without prompting,
303
- which still writes the fields - it does not leave them unset.
302
+ does not skip Phase 0's pickers. Autopilot resolves them without prompting, which still
303
+ writes the fields. For the base branch it reads `recentBranches[{projectKey}]` (most
304
+ recent inside the TTL, still on the remote) before the rule-6 order, recording which
305
+ fired in `state.baseBranchSource`.
304
306
 
305
307
  **TTL filter for recent branches**:
306
308
 
@@ -587,6 +589,35 @@ Persist `state.baseline.tests` with `command`, `capturedAt`, `logPath` and exact
587
589
 
588
590
  Log: `Phase 0 Step 7.6: test baseline = {green|red|unknown} ({N} pre-existing failures)`
589
591
 
592
+ #### Step 7.7 - Evidence capability, then test depth
593
+
594
+ Probe, then ask, then run. Skipped unless `state.visualEvidence.required` and `visualEvidence.enabled` is not `false`. Reasoning: `features/visual-evidence.md` section 4.
595
+
596
+ ```bash
597
+ eval "$(bash $HOME/.claude/scripts/probe-evidence-capability.sh \
598
+ --platform "$PLATFORM" --repo "$WORKTREE" --changed "$CHANGED_CSV" \
599
+ --json-out "$WORKTREE/.pipeline/evidence-capability.json")"
600
+ ```
601
+
602
+ One run, both forms: stdout is `EVIDENCE_*` (shell-quoted, so the eval is safe), and the same measurement lands as JSON.
603
+
604
+ Persist that file as `state.evidenceCapability`, then build the menu from it, never from a reading of the repo: `1. Sadece unit test` / `2. Unit + UI test, ekran kaydiyla` (tier 1) / `3. Unit + MCP ile akis kaydi` (tier 2).
605
+
606
+ A closed option keeps its row and prints the probe's reason verbatim. No `uiTestTargets` closes 2; `mcp` false closes 3; a missing `device` or `recorder` closes both. Targets present with no `matchingTests` leaves 2 open, warning that the whole UI suite will run. **When only option 1 is open, do not ask**: `testDepth = unit`, `testDepthSource = forced`, and the tier 3 gap is written.
607
+
608
+ Pass the default as the **1-based index**, never the label (`rules.md`: labels render in `outputLanguage`, so a label default matches nothing on a `tr` run and `ask-choice.sh` takes option 1 on a non-TTY):
609
+
610
+ ```bash
611
+ DEPTH_DEFAULT_INDEX=1
612
+ [ "$EVIDENCE_TIER1" = "open" ] && [ -n "$EVIDENCE_MATCHING_TESTS" ] && DEPTH_DEFAULT_INDEX=2
613
+ [ "$DEPTH_DEFAULT_INDEX" = "1" ] && [ "$EVIDENCE_TIER2" = "open" ] && DEPTH_DEFAULT_INDEX=3
614
+ ASK_CHOICE_DEFAULT="$DEPTH_DEFAULT_INDEX" $HOME/.claude/lib/ask-choice.sh ...
615
+ ```
616
+
617
+ Asked by `/multi-agent` and `:local`; autopilot reads `prefs.global.testDepth.default` and degrades to the best open option. Asked here, not Phase 5, which four of the eight modes drop.
618
+
619
+ Log: `Phase 0 Step 7.7: testDepth = {unit|unit+ui|unit+mcp} (source {user|autopilot|default|forced}), tier1/tier2 = {open|closed}`
620
+
590
621
  #### Step 8 - Clarification (opt-in, runs AFTER maturity, BEFORE Phase 1)
591
622
 
592
623
  **Gated by `prefs.global.clarifyAmbiguous.enabled`** (default: `false`). When enabled and `state.maturity.status != "blocker"`:
@@ -236,7 +236,19 @@ Gated by `prefs.global.devCritic.enabled` (default: `false`). When enabled, afte
236
236
 
237
237
  #### Step 3.55 - Visual evidence capture (UI changes only)
238
238
 
239
- When `state.visualEvidence.required`, capture the fixed state with the build that just went green: `capture-evidence.sh after --task "$TASK_ID" --platform "$PLATFORM" --label <slug>`. Here, not Phase 5, which every autopilot and `--local` entry drops. Exit 4 is a gap, not a failure. See `$HOME/.claude/multi-agent-refs/features/visual-evidence.md`.
239
+ When `state.visualEvidence.required`, capture with the build that just went green. Here, not Phase 5, which every autopilot and `--local` entry drops. Exit 4 anywhere below is a gap, not a failure. Contract: `features/visual-evidence.md`.
240
+
241
+ 1. **Still**: `$HOME/.claude/scripts/capture-evidence.sh after --task "$TASK_ID" --platform "$PLATFORM" --label <slug>`.
242
+ 2. **Re-check the device.** `evidenceCapability` was measured at intake; a simulator booted then can be gone now. Re-measure the one volatile row with `$HOME/.claude/scripts/probe-evidence-capability.sh --platform "$PLATFORM" --repo "$WORKTREE" --only device`, which skips the repo scan the full probe does; gone means fall to the next open tier and write `visualEvidence.videoTierReason` as the transition (`tier 1 -> 3: simulator no longer booted`).
243
+ 3. **Recording**, by `state.testDepth`:
244
+ - `unit+ui`: `video start` -> `$HOME/.claude/scripts/run-ui-tests.sh run --platform "$PLATFORM" --repo "$WORKTREE" --changed "$CHANGED_CSV"` -> `video stop`. Runner exit 3 (no matching test) or 4 (no target) -> stop, discard, fall to tier 2. Exit 1 is a red UI test: keep the recording, it shows the failure.
245
+ - `unit+mcp`: `video start` -> drive with `mcp__multi-agent-toolkit__agent_run_steps` -> `video stop`. The only MCP-dependent path; MCP absent -> tier 3 gap.
246
+ - `unit`: no recording; the gap reason is the user's own answer.
247
+ 4. **Fit**: `$HOME/.claude/scripts/capture-evidence.sh fit --file <path>` per artefact. Quality degrades, the artefact is never dropped.
248
+
249
+ `video stop` warns when the recording is under a second: both recorders encode on change, so a screen that never moved yields a valid two-frame file. Keep it, record the warning as a gap reason, never present a still as a flow.
250
+
251
+ Persist `state.uiTest` and the artefact entries. A UI test that ran is subject to the default-FAIL rule like the build: run `evidence-gate.mjs --claim test --status passed --evidence "$WORKTREE/.pipeline/ui-test.log"` first, because a runner that died before reaching the tests also exits non-zero.
240
252
 
241
253
  #### Step 3.6 - Code-simplifier pass (required diff shrink, before Phase 4 handoff)
242
254
 
@@ -112,6 +112,8 @@ Tier 1 / Tier 2 records print `screenshotUrl` from the captured evidence (Tier 2
112
112
  ```bash
113
113
  node $HOME/.claude/scripts/evidence-gate.mjs --claim manual --status passed --evidence "$WORKTREE/.pipeline/manual-test.json"
114
114
  ```
115
+ Add `--require-screenshot` when `state.visualEvidence.required` is true: a passing criterion then has to name a file that exists, because a `screenshot` key pointing nowhere is not evidence.
116
+
115
117
  Exit 1 means the "ok" is not accepted: tell the user which criterion is missing evidence (a `fail` verdict, or `not-tested` without a reason) and wait for the next reply. Exit 0 marks Phase 5 completed with `Result "local test passed (user)"`. The "fix: ..." path below is unchanged.
116
118
  6. If fix needed:
117
119
  - Branch already has WIP commit (from step 2) - changes are safe
@@ -138,7 +140,7 @@ Before or during user testing, run device-level audits via Bash if user requests
138
140
 
139
141
  | Check | When | Command |
140
142
  | ------------------- | ----------------- | ----------------------------------------- |
141
- | UI flow video | `state.visualEvidence.required` | 3 tiers, cap from `capture-evidence.sh limits` (`visualEvidence.maxVideoSeconds`) - `$HOME/.claude/multi-agent-refs/features/visual-evidence.md` |
143
+ | UI flow video | `state.visualEvidence.required` AND Phase 3 recorded none | `capture-evidence.sh video start` -> drive the flow -> `video stop` -> `fit`. See below |
142
144
  | Accessibility audit | UI changes | `mcp__multi-agent-toolkit__{ios,android}_accessibility_audit` |
143
145
  | Biometric test | Auth flow changes | ios: `mcp__multi-agent-toolkit__ios_biometric` (android: manual) |
144
146
  | Launch time | Perf-sensitive changes | ios: app-launch instrument · android: `mcp__multi-agent-toolkit__android_launch_time` |
@@ -146,6 +148,14 @@ Before or during user testing, run device-level audits via Bash if user requests
146
148
  | Snapshot regression | Component / pixel-stable UI changes | ios: `mcp__multi-agent-toolkit__ios_visual_diff` · android: `mcp__multi-agent-toolkit__android_screenshot` + compare |
147
149
  | Store screenshots | `taskType === screenshot` | ios: `ios_status_bar({preset: "clean"})` · android: `android_screenshot` |
148
150
 
151
+ ##### UI flow video, when Phase 3 produced none
152
+
153
+ Phase 3 Step 3.55 is the primary host and runs in every mode. Phase 5 is the richer one where it exists: the device is up and a person is watching, so the flow is one somebody confirmed. It adds, never replaces.
154
+
155
+ Run only when `visualEvidence.required` and `visualEvidence.video.file` is empty: re-check the device with `$HOME/.claude/scripts/probe-evidence-capability.sh --only device`, then `$HOME/.claude/scripts/capture-evidence.sh video start` -> drive the flow (`run-ui-tests.sh run`, or the Phase 5 scenarios by hand) -> `video stop` -> `fit`. The cap is `visualEvidence.maxVideoSeconds`, read through `capture-evidence.sh limits` so one value serves both. Update `videoTier` and `videoTierReason` with what ran.
156
+
157
+ When the intake answer was `unit` there is no recording here either: overriding it in a phase the user may not be watching makes the question decorative.
158
+
149
159
  Results included in Phase 7 report. MCP tools preferred when available - concise structured output, lower token cost.
150
160
 
151
161
  **Snapshot regression flow (optional):** when the task changes a stable component, capture a screenshot before the change (baseline) and after (current), then call `ios_visual_diff({baseline, current, max_diff_pct: 1.0})`. Threshold can be relaxed for animated / non-deterministic regions - keep `max_diff_pct ≤ 1.0` for static layouts.
@@ -130,6 +130,29 @@ Branch **deterministically**, no implicit fallback. Read `agent-state.json` and
130
130
  11. **NEVER close or resolve the issue** - neither GitHub Issue nor Jira. Issues require team review (4 approvals) before closing. Only post a comment with commit/PR URLs.
131
131
  12. Log: "Phase 6: Commit {sha} - PR #{number}, worktree {removed|kept: <reason>}"
132
132
 
133
+ #### Step 2.9 - Resolve the evidence host (UI changes only)
134
+
135
+ Runs when `state.visualEvidence.required`, before the PR body is written, because the body renders a different shape per host. Record `state.visualEvidence.host` and `hostReason`:
136
+
137
+ 1. **`jira`** - `state.jiraId` present. Upload stills and video with `jira-attach.sh "$JIRA_ID" <files>`; it prints `<filename>\t<url>` per file into `before[]` / `after[]` / `video`.
138
+ 2. **`github-public` / `github-private`** - no Jira, `visualEvidence.githubHost` is `branch`, remote is GitHub. Push the stills to the orphan branch below, then pick the variant from `gh repo view --json isPrivate`.
139
+ 3. **`none`** - anything else, including `githubHost: off`. Publish nothing, record the reason, let the `gaps[]` rule carry it.
140
+
141
+ No API attaches a file to a GitHub issue or PR (the web uploader needs a browser session), so the PR body can only point at something already hosted. An orphan branch is the one mechanism a script has that neither touches the PR diff nor publishes a release.
142
+
143
+ ```bash
144
+ EVB="evidence/$TASK_ID"
145
+ git -C "$WORKTREE" worktree add --detach "$TMP_EV" 2>/dev/null
146
+ git -C "$TMP_EV" checkout --orphan "$EVB" && git -C "$TMP_EV" rm -rf . >/dev/null 2>&1 || true
147
+ cp "$WORKTREE"/.pipeline/evidence/*.png "$TMP_EV"/ 2>/dev/null || true
148
+ git -C "$TMP_EV" add -A && git -C "$TMP_EV" commit -m "evidence: $TASK_ID"
149
+ git -C "$TMP_EV" push -u origin "$EVB"
150
+ ```
151
+
152
+ Stills only. On a GitHub-hosted run no video is recorded (Step 3.55), and an mp4 behind a blob link is a download rather than something a reviewer opens. A failed push is not a phase failure: drop to `host: none` with the error as `hostReason`.
153
+
154
+ Log: `Phase 6 Step 2.9: evidence host = {jira|github-public|github-private|none} ({reason})`
155
+
133
156
  #### Step 3 - PR Description (technical detail for reviewers)
134
157
 
135
158
  **Section set + markup dialect: `channels/pr.md` - read it first.** Phase 7 channels replaces this body with that section set, so build to it. Required reading: `payload-contracts.md`.
@@ -75,10 +75,45 @@ identifies which option, not what the button said.
75
75
  The host injects its own **Other** free-text row in English on every run. Nothing in
76
76
  the pipeline can localize it, and no option may depend on its wording.
77
77
 
78
+ ## Order: project, then repo, then branch
79
+
80
+ A picker may only be asked once everything it depends on is settled, and the
81
+ dependency runs one way: a base branch is a property of a repo, and a repo is a
82
+ property of a project. Asking for a branch before the dev-context repo set is
83
+ known means the answer was given about a repo the run had not chosen yet, and
84
+ nothing downstream can tell that apart from a correct answer.
85
+
86
+ So: project selection, then `_dev-context.md`, then the base branch. A step whose
87
+ input is not yet resolved waits; it does not guess and it does not resolve its own
88
+ input with a second question.
89
+
90
+ ## A single candidate is still a question
91
+
92
+ The number of options never authorises a skip. A filter that leaves one row has
93
+ narrowed the world; it has not decided anything, and the host's **Other** row is a real
94
+ choice on every picker - a branch the filter excluded, an account the probe missed, a
95
+ repo git does not know about. "There was only one option, so I picked it" is a skipped
96
+ picker, and announcing the pick in prose first is the same skip with a sentence in front
97
+ of it.
98
+
99
+ This is the failure that is hardest to see afterwards, because the transcript reads like
100
+ a decision was made. Only the picker's absence records that the user was never asked.
101
+
78
102
  ## Autopilot / non-interactive contract
79
103
 
80
104
  In autopilot, `ask_choice` resolves to `default` (or the safe first option) without prompting - identical to how the native gates auto-proceed today. A picker is only surfaced for genuinely ambiguous or destructive decisions, matching the maturity-check model.
81
105
 
106
+ **Memory outranks the default.** Where the pipeline has recorded what this user chose
107
+ last time for this project - `prefs.global.recentBranches[{projectKey}]` is the one that
108
+ exists today - autopilot resolves from that record first, and only falls back to the
109
+ option order when the record is empty, stale past its TTL, or names something that no
110
+ longer exists. A remembered choice is evidence about this user and this repo; an option
111
+ order is a guess that happens to be sorted. Where the two agree nothing changes, and
112
+ where they disagree the remembered one is the answer with a reason behind it.
113
+
114
+ The run records which rule fired (`remembered` or `default`). An autopilot run cannot be
115
+ asked anything, so the only thing that keeps it accountable is being readable afterwards.
116
+
82
117
  ## Deterministic gates note
83
118
 
84
119
  Claude Code's `PreToolUse` exit-2 hooks are the HARD blocking gates. Three ship, none needing run-specific arguments so they are naturally hookable: (1) `pre-commit-check.sh` scans the staged diff on every `git commit` and blocks on a detected secret; (2) `agent-guard.sh` runs on `git commit` + `git push` and blocks AI/assistant attribution in a commit message and force-push to a protected branch (main/master/develop); (3) `check-read-size.sh` runs on `Read` and on the shell commands that read a file whole, and routes an oversized read to a cheap worker (`bulk-read.sh`) instead of the caller's own rung. The first two inspect what a run WRITES; the third inspects what it pays to READ, and it is inert until `bulkRead.mode` is set to `observe` or `enforce`, so merging the block changes nothing until the user opts in. Its `observe` mode blocks nothing and only logs, which is how the baseline is measured before anything is routed. All three are self-contained, fail-open on internal error, and never execute the inspected command. Two capture hooks ship in the same block and block nothing: `SessionEnd` runs `capture-flush.sh --if-stale` (writing a killed run's findings into the per-repo stores, since every durable write used to live in Phase 7 - the phase a run is least likely to reach) plus `note-session.sh` (the mechanical shape of a non-pipeline session: tools used, commands that failed, calls the user refused - never an argument, never any output), and `SessionStart` runs `capture-resume.sh`, at most two lines about an unfinished run and a stale observation queue. Neither calls a model; both exit 0 on every path. The recommended hook block ships at `install/templates/claude-hooks.json`; `multi-agent:setup` offers to merge it into `~/.claude/settings.json`. The other deterministic gates (evidence, consensus, intent, learnings) are invoked by the pipeline phases with per-run arguments (a build-log path, the triage JSON, the free-text input), so they are phase-enforced by contract, not OS-hookable.
@@ -32,7 +32,7 @@ Visual mechanism per CLI:
32
32
 
33
33
  | CLI | What the agent calls at every phase boundary |
34
34
  |---|---|
35
- | **claude-code** | `TaskCreate({subject, activeForm})` then `TaskUpdate({status, activeForm})`. Native sticky widget; ⏺ tiles, spinner header. |
35
+ | **claude-code** | `TaskCreate({subject, activeForm})` then `TaskUpdate({status, activeForm})`. Native sticky widget; ⏺ tiles, spinner header. **Not in every session** - see below. |
36
36
  | **copilot** | Inline call: `bash phase-tracker.sh render`. The bordered ANSI card lands as the last tool result in the chat. |
37
37
  | **codex** | The native `update_plan` tool: one plan step per phase, `status: pending \| in_progress \| completed`. Rewrite the whole step list on each boundary - the tool takes the full plan, not a delta. |
38
38
  | **generic** (plain shell, Git Bash, WSL, tmux) | Same bash render - the bordered ANSI card prints to terminal stdout in place. |
@@ -41,6 +41,10 @@ Visual mechanism per CLI:
41
41
 
42
42
  **Claude Code only**: in addition, `TaskCreate` / `TaskUpdate` native tool calls → the sticky widget pins the phase stack in the user's view.
43
43
 
44
+ **The task tools depend on the MODEL, not the CLI version.** Claude Code provides `TaskCreate`, `TaskUpdate`, `TaskGet`, `TaskList` and `TodoWrite` by default only on Claude 3.x, Opus 4 through 4.7, Sonnet 4 through 4.6 and Haiku 4.5; on any newer model it leaves them out unless the user opts in. That default arrived in Claude Code v2.1.268, and this contract was written when the tools were universal, so nothing noticed the change: a run wrote its tracker state correctly, advanced every phase, and drew nothing for the whole run.
45
+
46
+ So the branch is taken by the agent, which knows its own tool list, and never by the shell, which cannot see it. Without the tools the bordered card IS the widget and has to be reprinted inside the reply at every boundary, exactly as on Copilot CLI. Say once, in `outputLanguage`, that the native widget returns with `CLAUDE_CODE_ENABLE_TODO_TOOLS=1 claude` (or `claude --allowedTools TaskCreate`); a user staring at a missing widget needs the command, not the diagnosis. A subagent inherits the session's tool set even when it runs a different model, so this is a session-level fact and not a per-phase one.
47
+
44
48
  ### Codex specifics
45
49
 
46
50
  Two constraints on `update_plan`, both of which fail quietly if ignored:
@@ -63,7 +63,7 @@
63
63
  "premiumTierUntil": null,
64
64
  "fallbackModel": "sonnet",
65
65
  "floorModel": "haiku",
66
- "fableEnabled": true,
66
+ "fableEnabled": false,
67
67
  "onDispatchError": true
68
68
  },
69
69
  "costBudget": {