@mmerterden/multi-agent-pipeline 16.29.0 → 16.31.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (40) hide show
  1. package/CHANGELOG.md +82 -0
  2. package/docs/features.md +14 -0
  3. package/install/_common.mjs +45 -4
  4. package/package.json +1 -1
  5. package/pipeline/commands/multi-agent/design-check/SKILL.md +6 -5
  6. package/pipeline/commands/multi-agent/help/SKILL.md +13 -12
  7. package/pipeline/commands/multi-agent/manual-test/SKILL.md +1 -1
  8. package/pipeline/commands/multi-agent/sync/SKILL.md +3 -4
  9. package/pipeline/lib/credential-inventory.sh +1 -0
  10. package/pipeline/lib/repo-hygiene.sh +164 -0
  11. package/pipeline/lib/vercel-deploy.sh +41 -22
  12. package/pipeline/multi-agent-refs/channels/pr.md +37 -1
  13. package/pipeline/multi-agent-refs/features/doctor.md +12 -0
  14. package/pipeline/multi-agent-refs/features/model-fallback.md +2 -2
  15. package/pipeline/multi-agent-refs/features/visual-evidence.md +103 -20
  16. package/pipeline/multi-agent-refs/keychain.md +1 -0
  17. package/pipeline/multi-agent-refs/knowledge.md +1 -1
  18. package/pipeline/multi-agent-refs/phases/phase-0-init.md +36 -7
  19. package/pipeline/multi-agent-refs/phases/phase-2-planning.md +1 -1
  20. package/pipeline/multi-agent-refs/phases/phase-3-dev.md +14 -2
  21. package/pipeline/multi-agent-refs/phases/phase-5-test.md +12 -2
  22. package/pipeline/multi-agent-refs/phases/phase-6-commit.md +23 -0
  23. package/pipeline/preferences-template.json +2 -1
  24. package/pipeline/schemas/agent-state.schema.json +79 -1
  25. package/pipeline/schemas/prefs.schema.json +24 -1
  26. package/pipeline/schemas/token-budget.json +10 -10
  27. package/pipeline/scripts/bulk-read.sh +13 -5
  28. package/pipeline/scripts/capture-evidence.sh +170 -5
  29. package/pipeline/scripts/doctor.mjs +56 -1
  30. package/pipeline/scripts/evidence-gate.mjs +31 -2
  31. package/pipeline/scripts/gc-tmp.sh +30 -0
  32. package/pipeline/scripts/gc-worktrees.sh +32 -10
  33. package/pipeline/scripts/offload-ref.sh +13 -5
  34. package/pipeline/scripts/probe-evidence-capability.sh +250 -0
  35. package/pipeline/scripts/purge.sh +11 -0
  36. package/pipeline/scripts/run-ui-tests.sh +380 -0
  37. package/pipeline/scripts/worktree-finalize.sh +9 -0
  38. package/pipeline/skills/.skill-manifest.json +19 -11
  39. package/pipeline/skills/shared/core/multi-agent-manual-test/SKILL.md +10 -1
  40. package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +2 -1
@@ -63,7 +63,9 @@ Multi-repo PRs (one PR per repo) emit verification commands for that repo's stac
63
63
  - Rollback: feature flag <name> | git revert <sha> | none, and why
64
64
  ```
65
65
 
66
- **`visuals`** - only when `state.visualEvidence.required`. The images live on the Jira issue as attachments; this section exists so a reviewer opening the PR knows they are there and what each one shows. Filenames, not raw URLs - a Jira attachment URL is auth-gated and renders as a broken image for anyone reading the PR outside a Jira session.
66
+ **`visuals`** - only when `state.visualEvidence.required`. What this section can show depends on where the artefacts are hosted, which Phase 6 resolves into `state.visualEvidence.host`. Render the form for that host and no other.
67
+
68
+ **`host: jira`.** Filenames, never URLs. A Jira attachment URL is auth-gated and renders as a broken image for anyone reading the PR outside a Jira session, and a broken image is worse than a filename because it looks like the evidence is missing.
67
69
 
68
70
  ```markdown
69
71
  ## Visual Evidence
@@ -73,6 +75,40 @@ Multi-repo PRs (one PR per repo) emit verification commands for that repo's stac
73
75
  - Flow video: `<flow-filename>`, tier <N> (attached to PROJ-XXXXX)
74
76
  ```
75
77
 
78
+ **`host: github-public`.** The stills are on the `evidence/<task-id>` branch, so they embed and the reviewer sees them without leaving the PR:
79
+
80
+ ```markdown
81
+ ## Visual Evidence
82
+
83
+ **Before**
84
+
85
+ ![before](https://raw.githubusercontent.com/<owner>/<repo>/evidence/<task-id>/<before-filename>)
86
+
87
+ **After**
88
+
89
+ ![after](https://raw.githubusercontent.com/<owner>/<repo>/evidence/<task-id>/<after-filename>)
90
+ ```
91
+
92
+ **`host: github-private`.** Same branch, but a link rather than an embed. GitHub renders markdown images through its own proxy, which has no credentials for a private repo, so an embedded raw URL renders broken for every reader including the author. A blob link opens the image for anyone who can already see the repo:
93
+
94
+ ```markdown
95
+ ## Visual Evidence
96
+
97
+ - Before: [<before-filename>](https://github.com/<owner>/<repo>/blob/evidence/<task-id>/<before-filename>)
98
+ - After: [<after-filename>](https://github.com/<owner>/<repo>/blob/evidence/<task-id>/<after-filename>)
99
+ ```
100
+
101
+ **`host: none`.** Filenames plus the artefact directory, and the reason there is no host:
102
+
103
+ ```markdown
104
+ ## Visual Evidence
105
+
106
+ - After: `<after-filename>` (run artefacts: `<artifactsPath>`)
107
+ - Not published: <hostReason>
108
+ ```
109
+
110
+ **Video is Jira-only.** On a GitHub-hosted run no recording is made and none is published: an mp4 behind a blob link is a download, not something a reviewer opens mid-review, and paying for a recording nobody watches is worse than saying plainly that there is none. The gap line carries that reason.
111
+
76
112
  Every `state.visualEvidence.gaps[]` entry becomes its own line with the reason instead of a filename (`- Before: none - the ticket carries no image attachment`). Phase 6 Step 3 blocks on a required artefact that is neither listed nor explained. Contract: `$HOME/.claude/multi-agent-refs/features/visual-evidence.md`.
77
113
 
78
114
  **`dependencies`** - only when `Package.swift` / `Podfile` / `build.gradle` / `package.json` changed. Each entry: `package@old → new - reason`.
@@ -195,3 +195,15 @@ read the registration and stopped would have called that healthy.
195
195
  Free space on the volume holding `$HOME`. WARN under 2 GB: a worktree plus a
196
196
  build is the largest thing a run writes, and ENOSPC mid-run corrupts the state
197
197
  file it was writing at the time.
198
+
199
+ ### worktree-residue
200
+
201
+ Worktrees left under `<repo>/.worktrees/` in the repository the caller is
202
+ standing in. WARN at five or more, or at 2 GB. A finished task removes its own
203
+ worktree at PR time, but a run that stops before Phase 6 never reaches that
204
+ step and nothing else collects it: the finalizer only runs on success, and
205
+ `gc-worktrees` only sweeps entries git has already forgotten. Each survivor is
206
+ a full second checkout, so the total is measured in gigabytes rather than
207
+ megabytes. SKIP outside a git repository - a project name in prefs is a name,
208
+ not a path, and guessing checkout locations to produce a number is how a
209
+ diagnostic starts lying.
@@ -121,10 +121,10 @@ evidence, and there is now one fewer of them. Triage also runs on `opus`, which
121
121
  makes it the same model as Reviewer 1; the Step 3 anonymisation requirement
122
122
  already covers that case and is not optional here.
123
123
 
124
- **Cost accounting.** `prefs.global.costBudget.priceAt` defaults to `fable` to
124
+ **Cost accounting.** `prefs.global.costBudget.pricingModel` defaults to `fable` to
125
125
  keep the estimate an upper bound. With the rung off, that default prices every
126
126
  call above what it can cost and trips the budget ceiling early, which then
127
- triggers a downgrade nobody needed. Set `priceAt` to `opus` alongside
127
+ triggers a downgrade nobody needed. Set `pricingModel` to `opus` alongside
128
128
  `fableEnabled: false`.
129
129
 
130
130
  No log line is emitted per dispatch for this: it is a configured absence, not a
@@ -78,32 +78,86 @@ at most 1242px wide, and writes
78
78
  A clean status bar is not cosmetic: without it two captures of the same screen
79
79
  differ by the clock, which makes every "after" look like a change.
80
80
 
81
- ## 4. Video - three tiers, resolved once per repo
81
+ ## 4. Video - the recording rides on a test run
82
82
 
83
- The flow source, in order: a UI test flow written on the ticket wins; otherwise
84
- the Phase 5 test scenarios are the flow. Human-written beats generated, the same
85
- rule the "before" follows.
83
+ The flow video is not recorded on its own. It wraps something that drives the
84
+ screen, and what drives it is chosen by the user at intake, because running a UI
85
+ suite costs minutes and that is their time to spend.
86
86
 
87
- The tier is probed once and stored in `state.visualEvidence.videoTier`:
87
+ ### 4.1 Probe first
88
88
 
89
- | Tier | Condition | Mechanism |
90
- |---|---|---|
91
- | 1 | The repo has a UI test target (XCUITest / Espresso) **and** it can be seeded with test data | Run that test, record the screen for its duration. Most faithful, and it is a test that already runs in CI |
92
- | 2 | A runnable build exists | Drive the flow with `mcp__multi-agent-toolkit__agent_run_steps` and record with `ios_record_video` / `android_record_screen`. **No test target needed** |
93
- | 3 | No runnable build (library / SPM package / backend) | Skip, with the reason recorded |
89
+ `probe-evidence-capability.sh` measures, before the question is asked, what this
90
+ machine and this repo can actually do: the UI test target, the tests matching
91
+ this change, the device, the recorder CLI, and whether the toolkit MCP is
92
+ registered. Detection is delegated to `run-ui-tests.sh detect`, which is also
93
+ what runs the tests, so there is one implementation of the answer.
94
+
95
+ Two rules the probe keeps, and the reason for each:
96
+
97
+ - **Absence carries a reason.** `no booted simulator, but one is available to
98
+ boot` and `no iOS simulator available on this machine` close the same menu row
99
+ and ask the user for completely different things.
100
+ - **Unmeasurable is `null`, never `false`.** With no `adb` on the PATH the probe
101
+ cannot see whether a device is attached; reporting that as "no device" is a
102
+ negative nobody looked for, and it reads exactly like one somebody checked.
94
103
 
95
- Tier 2 is why "can every repo do this" answers yes in practice: any app that can
96
- be launched can be driven. A repo without a UI test target loses fidelity, not
97
- the recording.
104
+ ### 4.2 Then the question
98
105
 
99
- Preferred host is **Phase 5** - the device is already up and a human is present
100
- to confirm the flow is the right one. In a mode without Phase 5, tier 2 runs at
101
- the end of Phase 3 instead. Tier 1 always runs where the test runs.
106
+ Phase 0 Step 7.7 asks the depth, with the options built from the probe. A closed
107
+ option stays on the list carrying its reason; when every option but the first is
108
+ closed, nothing is asked and `testDepthSource` records `forced`. The full menu
109
+ rules are in `phases/phase-0-init.md`.
102
110
 
103
- Duration is capped by `visualEvidence.maxVideoSeconds` (default 60), read via
104
- `capture-evidence.sh limits` so the number lives in one place rather than in two
105
- documents. A flow that needs longer is not a review artefact, it is a debugging
106
- session.
111
+ ### 4.3 Then the recording
112
+
113
+ | `testDepth` | Tier | What drives the screen |
114
+ |---|---|---|
115
+ | `unit+ui` | 1 | The repo's own UI test, selected by the changed files. Most faithful, and it is a test that already runs in CI |
116
+ | `unit+mcp` | 2 | `mcp__multi-agent-toolkit__agent_run_steps` drives the flow. No test target needed |
117
+ | `unit` | 3 | Nothing. No recording, and the reason is the user's own answer |
118
+
119
+ The recorder itself is `capture-evidence.sh video start|stop`, which writes
120
+ `<task-id>[-<label>]-flow.mp4` into the evidence directory. The label is optional
121
+ because one recording per task is the common case, and both renderers cite the
122
+ unlabelled form. It is shell rather than MCP for three reasons: the sibling still capture already shells out,
123
+ a host with no toolkit MCP registered still produces evidence, and Phase 3 and
124
+ Phase 5 are exactly where an MCP call is contested.
125
+
126
+ **iOS UI test targets are not named, they are detected.** A path containing
127
+ `UITests` is not the signal: in the reference app 477 files sit under such a path
128
+ and exactly 2 drive the UI, the rest being snapshot tests that render a view and
129
+ compare pixels without ever launching the app. The signal is `XCUIApplication`,
130
+ the only API that drives another process's interface, which is precisely the
131
+ precondition a flow recording has. On Android the signal is the `androidTest`
132
+ source set, which is instrumentation by definition.
133
+
134
+ **A real app has many candidates** - 17 UI test directories on the reference iOS
135
+ app, 8 instrumentation modules on the Android one. Detection reports the whole
136
+ set and lets the changed files choose; taking the first off a `find` is a guess
137
+ wearing a measurement's clothes.
138
+
139
+ ### 4.4 Re-check before recording
140
+
141
+ The probe runs at intake and the recording happens in Phase 3. A simulator booted
142
+ then can be gone by the time the build goes green, so Phase 3 re-measures the
143
+ device row alone and downgrades the tier if it has to, recording the transition
144
+ (`tier 1 -> 3: simulator no longer booted`). A tier taken from a stale
145
+ measurement is a promise the run cannot keep.
146
+
147
+ ### 4.5 Duration, and the recording that shows nothing
148
+
149
+ The cap is `visualEvidence.maxVideoSeconds` (default 60), read through
150
+ `capture-evidence.sh limits` so one value serves the phase doc and the recorder.
151
+ Android's `screenrecord` has its own ceiling of 180 seconds that no setting can
152
+ lift, so the script clamps to it and says when it did: one preference honoured on
153
+ one platform and silently halved on the other is worse than a stated limit.
154
+
155
+ Both recorders encode on change. A flow over a screen that never moved therefore
156
+ produces a valid two-frame mp4 a fraction of a second long - not a broken file,
157
+ but not evidence of a flow either. `video stop` says so on stderr when the result
158
+ is under a second, and the caller records that as a gap reason. The duration is
159
+ never asserted against wall clock anywhere, because doing so fails a correct
160
+ capture of a static screen.
107
161
 
108
162
  `visualEvidence.enabled` turns the whole feature off - capture, upload, both
109
163
  render sections and the Phase 6 blocker with it.
@@ -125,6 +179,35 @@ Order of attempts, each one recorded:
125
179
 
126
180
  Never silently attach nothing.
127
181
 
182
+ ## 5b. Where the artefacts live - the host
183
+
184
+ Resolved in Phase 6 Step 2.9, recorded as `state.visualEvidence.host`:
185
+
186
+ | Order | Host | Condition | Stills | Video |
187
+ |---|---|---|---|---|
188
+ | 1 | `jira` | `jiraId` present | Jira attachments | Jira attachment |
189
+ | 2 | `github-public` | no Jira, GitHub remote, public repo | `evidence/<task-id>` orphan branch, embedded in the PR body | none |
190
+ | 3 | `github-private` | no Jira, GitHub remote, private repo | same branch, blob permalink in the PR body | none |
191
+ | 4 | `none` | anything else, or `githubHost: off` | not published, gap recorded | none |
192
+
193
+ **GitHub cannot be given a file.** There is no API that attaches an image to an
194
+ issue or a pull request; the web uploader posts to an endpoint that needs a
195
+ browser session, so no token can drive it. The PR body can only point at
196
+ something already hosted, and an orphan branch is the one mechanism a script has
197
+ that neither touches the PR diff nor publishes a release.
198
+
199
+ **A private repo cannot show an inline image.** GitHub renders markdown images
200
+ through its own proxy, which carries no credentials for a private repo, so an
201
+ embedded `raw.githubusercontent.com` URL renders broken for every reader
202
+ including the author. That is worse than a link, because a broken image looks
203
+ like missing evidence. The private variant therefore links rather than embeds.
204
+
205
+ **Video is Jira-only, by decision.** On a GitHub-hosted run none is recorded at
206
+ all. An mp4 behind a blob link is a download rather than something a reviewer
207
+ opens mid-review, and paying minutes of UI-test time for a recording nobody
208
+ watches is worse than saying plainly that there is none. The gap line carries
209
+ that reason.
210
+
128
211
  ## 6. Rendering
129
212
 
130
213
  ### Jira comment (`channels/jira.md`)
@@ -158,6 +158,7 @@ The shell driver auto-delegates to `~/.claude/scripts/keychain.py` on macOS / Li
158
158
  | `graylog_test` | Graylog (test) | `${USER}_Graylog_Test_Access_Token` | API Token, optional - unset falls back to `graylog` | Same page on the TEST instance; optional |
159
159
  | `firebase` | Firebase | `${USER}_Firebase_Access_Json` | Service-account JSON, as issued. One key per Firebase project; extras are named `..._Json_<projectId>` and listed in `global.firebase.accounts[]` | Firebase Console -> Project settings -> Service accounts -> Generate new private key |
160
160
  | `jenkins` | Jenkins CI | `${USER}_Jenkins_Access_Token` | API Token | Jenkins -> User -> Configure -> API Token |
161
+ | `vercel` | Vercel deploys | `${USER}_Vercel_Access_Token` | Access Token | Vercel -> Account Settings -> Tokens. Read by `vercel-deploy.sh`; never passed on argv |
161
162
  | `appstore_connect_key_id` | App Store Connect | `${USER}_AppStoreConnect_Key_Id` | Identifier, not a secret | App Store Connect -> Users and Access -> Integrations -> App Store Connect API |
162
163
  | `appstore_connect_issuer_id` | App Store Connect | `${USER}_AppStoreConnect_Issuer_Id` | Identifier, not a secret | Same page as the key id; one issuer id per team |
163
164
  | `appstore_connect_apple_id` | App Store Connect | `${USER}_AppStoreConnect_Apple_Id` | Email address | The Apple ID itself; no generation step |
@@ -63,7 +63,7 @@ Knowledge files grow over time. Maintenance rules:
63
63
  and a rebuild is seconds, so age is the wrong question to ask of it.
64
64
  - Stale entries are added to the prompt with a "STALE - verify before relying" tag
65
65
  - `prune-logs` does not touch knowledge - only deletes logs and state
66
- - `purge` does not touch knowledge either - separate command: `/multi-agent clear-knowledge {project}`
66
+ - `purge` does not touch knowledge either. Today the only removal path is `/multi-agent:uninstall --all-data`, which drops every project's knowledge at once; there is no per-project clear yet
67
67
 
68
68
  ---
69
69
 
@@ -447,14 +447,15 @@ fi
447
447
 
448
448
  If the path is still registered and healthy after the heal, enter + pull instead of re-adding (existing behavior). Only when `worktree add` still fails after the heal do the rollback / collision flow in the multi-repo block below.
449
449
 
450
- **Worktree residue guard (required once per repo, idempotent):** every worktree dir holds a `.git` file, so a blanket `git add -A` in the parent tree records `.worktrees/{id}` as a gitlink that pollutes the branch. Keep `.worktrees/` out of the index via the clone-local exclude file (never committed, so no PR noise):
450
+ **Repo residue guard (required once per repo, idempotent):** keep every path the pipeline writes - `.worktrees/`, `.pipeline/`, `.multi-agent/` and the loose logs - out of the index via the clone-local exclude file (never committed, so no PR noise). One call; the rationale is in the lib header:
451
451
 
452
452
  ```bash
453
- ex="$(git -C "$PROJECT_ROOT" rev-parse --path-format=absolute --git-common-dir)/info/exclude"
454
- mkdir -p "$(dirname "$ex")"
455
- grep -qxF '.worktrees/' "$ex" 2>/dev/null || printf '.worktrees/\n' >> "$ex"
453
+ . "$HOME/.claude/lib/repo-hygiene.sh"
454
+ ma_hygiene_ensure_exclusions "$PROJECT_ROOT"
456
455
  ```
457
456
 
457
+ Rewritten on each call, so an older checkout picks up new entries; lines a human added are never touched.
458
+
458
459
  **Traversal-prune contract:** the exclude guard covers only the git INDEX, not
459
460
  filesystem scans. Since each worktree is a full checkout, an unpruned tree walk
460
461
  double-processes every file and can re-stage gitlinks. So `.worktrees` joins the
@@ -468,14 +469,13 @@ scan, `shadow-git.sh` excludes, and the shared walkers (`extract-conventions.sh`
468
469
  When `state.projects[].length > 1`, repeat steps 2-4 **serially per repo** (worktrees are cheap; serial keeps git index sane and surfaces collisions one at a time):
469
470
 
470
471
  ```bash
472
+ . "$HOME/.claude/lib/repo-hygiene.sh"
471
473
  for proj in "${PROJECTS[@]}"; do
472
474
  WT_PATH="$proj/.worktrees/$BRANCH_DIR/"
473
475
  git -C "$proj" worktree prune 2>/dev/null || true # heal stale admin state
474
476
  git -C "$proj" worktree list --porcelain | grep -qF "$WT_PATH" \
475
477
  && git -C "$proj" worktree unlock "$WT_PATH" 2>/dev/null || true
476
- ex="$(git -C "$proj" rev-parse --path-format=absolute --git-common-dir)/info/exclude"
477
- mkdir -p "$(dirname "$ex")"
478
- grep -qxF '.worktrees/' "$ex" 2>/dev/null || printf '.worktrees/\n' >> "$ex" # residue guard
478
+ ma_hygiene_ensure_exclusions "$proj" # residue guard
479
479
  git -C "$proj" worktree add "$WT_PATH" -b "$BRANCH" "origin/$BASE_BRANCH"
480
480
  git -C "$WT_PATH" config user.name "${proj_identity_name}"
481
481
  git -C "$WT_PATH" config user.email "${proj_identity_email}"
@@ -589,6 +589,35 @@ Persist `state.baseline.tests` with `command`, `capturedAt`, `logPath` and exact
589
589
 
590
590
  Log: `Phase 0 Step 7.6: test baseline = {green|red|unknown} ({N} pre-existing failures)`
591
591
 
592
+ #### Step 7.7 - Evidence capability, then test depth
593
+
594
+ Probe, then ask, then run. Skipped unless `state.visualEvidence.required` and `visualEvidence.enabled` is not `false`. Reasoning: `features/visual-evidence.md` section 4.
595
+
596
+ ```bash
597
+ eval "$(bash $HOME/.claude/scripts/probe-evidence-capability.sh \
598
+ --platform "$PLATFORM" --repo "$WORKTREE" --changed "$CHANGED_CSV" \
599
+ --json-out "$WORKTREE/.pipeline/evidence-capability.json")"
600
+ ```
601
+
602
+ One run, both forms: stdout is `EVIDENCE_*` (shell-quoted, so the eval is safe), and the same measurement lands as JSON.
603
+
604
+ Persist that file as `state.evidenceCapability`, then build the menu from it, never from a reading of the repo: `1. Sadece unit test` / `2. Unit + UI test, ekran kaydiyla` (tier 1) / `3. Unit + MCP ile akis kaydi` (tier 2).
605
+
606
+ A closed option keeps its row and prints the probe's reason verbatim. No `uiTestTargets` closes 2; `mcp` false closes 3; a missing `device` or `recorder` closes both. Targets present with no `matchingTests` leaves 2 open, warning that the whole UI suite will run. **When only option 1 is open, do not ask**: `testDepth = unit`, `testDepthSource = forced`, and the tier 3 gap is written.
607
+
608
+ Pass the default as the **1-based index**, never the label (`rules.md`: labels render in `outputLanguage`, so a label default matches nothing on a `tr` run and `ask-choice.sh` takes option 1 on a non-TTY):
609
+
610
+ ```bash
611
+ DEPTH_DEFAULT_INDEX=1
612
+ [ "$EVIDENCE_TIER1" = "open" ] && [ -n "$EVIDENCE_MATCHING_TESTS" ] && DEPTH_DEFAULT_INDEX=2
613
+ [ "$DEPTH_DEFAULT_INDEX" = "1" ] && [ "$EVIDENCE_TIER2" = "open" ] && DEPTH_DEFAULT_INDEX=3
614
+ ASK_CHOICE_DEFAULT="$DEPTH_DEFAULT_INDEX" $HOME/.claude/lib/ask-choice.sh ...
615
+ ```
616
+
617
+ Asked by `/multi-agent` and `:local`; autopilot reads `prefs.global.testDepth.default` and degrades to the best open option. Asked here, not Phase 5, which four of the eight modes drop.
618
+
619
+ Log: `Phase 0 Step 7.7: testDepth = {unit|unit+ui|unit+mcp} (source {user|autopilot|default|forced}), tier1/tier2 = {open|closed}`
620
+
592
621
  #### Step 8 - Clarification (opt-in, runs AFTER maturity, BEFORE Phase 1)
593
622
 
594
623
  **Gated by `prefs.global.clarifyAmbiguous.enabled`** (default: `false`). When enabled and `state.maturity.status != "blocker"`:
@@ -120,7 +120,7 @@ Based on Phase 1 `detectedStack`, assign relevant skills:
120
120
 
121
121
  #### Output contract
122
122
 
123
- Phase 2 produces an object conforming to `$HOME/.claude/schemas/planning-output.schema.json` - `tasks[]` with `id`, `subject`, `targetFiles`, `complexity`, `blockedBy`, plus optional `architectureNotes` and `mode`. Phase 3 reads `tasks[]` in dependency order; the schema's `blockedBy` field drives the ready-task picker.
123
+ Phase 2 produces an object conforming to `$HOME/.claude/schemas/planning-output.schema.json` - `tasks[]` with `id`, `title`, `type`, `files`, plus optional `dependsOn` and `acceptanceCriteria`. Phase 3 reads `tasks[]` in dependency order; the schema's `dependsOn` field drives the ready-task picker.
124
124
 
125
125
  **Required: validator gate (deterministic) - run on the persisted file before the approval gate renders the plan; the validator's exit code decides, not the LLM turn:**
126
126
 
@@ -49,7 +49,7 @@ The analysis document is the SOLE design source in Phase 3. Variant choices, pad
49
49
 
50
50
  #### Input contract
51
51
 
52
- Phase 3 consumes the Phase 2 output object conforming to `$HOME/.claude/schemas/planning-output.schema.json` - the task graph (`tasks[]` with `id`, `subject`, `targetFiles`, `complexity`, `blockedBy`) plus the architecture review notes. Tasks execute in dependency order; the schema's `blockedBy` field drives the ready-task picker. In a Short run (no Phase 2), Opus generates the equivalent task list inline before entering the loop below.
52
+ Phase 3 consumes the Phase 2 output object conforming to `$HOME/.claude/schemas/planning-output.schema.json` - the task graph (`tasks[]` with `id`, `title`, `type`, `files`, and optional `dependsOn` / `acceptanceCriteria`) plus the architecture review notes. Tasks execute in dependency order; the schema's `dependsOn` field drives the ready-task picker. In a Short run (no Phase 2), Opus generates the equivalent task list inline before entering the loop below.
53
53
 
54
54
  **Plan Todo iteration (opt-in)**: gated by `prefs.global.planTodos.enabled` (default: `false`). When enabled and Phase 2 Step 4.5 emitted a `plan.todos[]`, Phase 3 iterates via `$HOME/.claude/lib/plan-todos.sh next/start/complete/fail` instead of walking `tasks[]` directly. When disabled, the loop walks `tasks[]` from `planning-output` - TDD contract is unchanged. Full helper loop + state semantics: `$HOME/.claude/multi-agent-refs/features/plan-todos.md`. A todo with `sourceTag: Reuse` binds the file analysis already found; `Modify` edits in place. Writing a new file over a `Reuse` step is a Locked 11 violation and Phase 4 flags it.
55
55
 
@@ -236,7 +236,19 @@ Gated by `prefs.global.devCritic.enabled` (default: `false`). When enabled, afte
236
236
 
237
237
  #### Step 3.55 - Visual evidence capture (UI changes only)
238
238
 
239
- When `state.visualEvidence.required`, capture the fixed state with the build that just went green: `capture-evidence.sh after --task "$TASK_ID" --platform "$PLATFORM" --label <slug>`. Here, not Phase 5, which every autopilot and `--local` entry drops. Exit 4 is a gap, not a failure. See `$HOME/.claude/multi-agent-refs/features/visual-evidence.md`.
239
+ When `state.visualEvidence.required`, capture with the build that just went green. Here, not Phase 5, which every autopilot and `--local` entry drops. Exit 4 anywhere below is a gap, not a failure. Contract: `features/visual-evidence.md`.
240
+
241
+ 1. **Still**: `$HOME/.claude/scripts/capture-evidence.sh after --task "$TASK_ID" --platform "$PLATFORM" --label <slug>`.
242
+ 2. **Re-check the device.** `evidenceCapability` was measured at intake; a simulator booted then can be gone now. Re-measure the one volatile row with `$HOME/.claude/scripts/probe-evidence-capability.sh --platform "$PLATFORM" --repo "$WORKTREE" --only device`, which skips the repo scan the full probe does; gone means fall to the next open tier and write `visualEvidence.videoTierReason` as the transition (`tier 1 -> 3: simulator no longer booted`).
243
+ 3. **Recording**, by `state.testDepth`:
244
+ - `unit+ui`: `video start` -> `$HOME/.claude/scripts/run-ui-tests.sh run --platform "$PLATFORM" --repo "$WORKTREE" --changed "$CHANGED_CSV"` -> `video stop`. Runner exit 3 (no matching test) or 4 (no target) -> stop, discard, fall to tier 2. Exit 1 is a red UI test: keep the recording, it shows the failure.
245
+ - `unit+mcp`: `video start` -> drive with `mcp__multi-agent-toolkit__agent_run_steps` -> `video stop`. The only MCP-dependent path; MCP absent -> tier 3 gap.
246
+ - `unit`: no recording; the gap reason is the user's own answer.
247
+ 4. **Fit**: `$HOME/.claude/scripts/capture-evidence.sh fit --file <path>` per artefact. Quality degrades, the artefact is never dropped.
248
+
249
+ `video stop` warns when the recording is under a second: both recorders encode on change, so a screen that never moved yields a valid two-frame file. Keep it, record the warning as a gap reason, never present a still as a flow.
250
+
251
+ Persist `state.uiTest` and the artefact entries. A UI test that ran is subject to the default-FAIL rule like the build: run `evidence-gate.mjs --claim test --status passed --evidence "$WORKTREE/.pipeline/ui-test.log"` first, because a runner that died before reaching the tests also exits non-zero.
240
252
 
241
253
  #### Step 3.6 - Code-simplifier pass (required diff shrink, before Phase 4 handoff)
242
254
 
@@ -112,6 +112,8 @@ Tier 1 / Tier 2 records print `screenshotUrl` from the captured evidence (Tier 2
112
112
  ```bash
113
113
  node $HOME/.claude/scripts/evidence-gate.mjs --claim manual --status passed --evidence "$WORKTREE/.pipeline/manual-test.json"
114
114
  ```
115
+ Add `--require-screenshot` when `state.visualEvidence.required` is true: a passing criterion then has to name a file that exists, because a `screenshot` key pointing nowhere is not evidence.
116
+
115
117
  Exit 1 means the "ok" is not accepted: tell the user which criterion is missing evidence (a `fail` verdict, or `not-tested` without a reason) and wait for the next reply. Exit 0 marks Phase 5 completed with `Result "local test passed (user)"`. The "fix: ..." path below is unchanged.
116
118
  6. If fix needed:
117
119
  - Branch already has WIP commit (from step 2) - changes are safe
@@ -124,7 +126,7 @@ Tier 1 / Tier 2 records print `screenshotUrl` from the captured evidence (Tier 2
124
126
  git -C "$PROJECT_ROOT" worktree unlock "{worktree-path}" 2>/dev/null || true
125
127
  fi
126
128
  ```
127
- Phase 0's `.worktrees/` residue guard is already in `.git/info/exclude` - no re-add.
129
+ Phase 0's repo residue guard is already in `.git/info/exclude` - no re-add.
128
130
  - Recreate worktree from branch: `git -C $PROJECT_ROOT worktree add {worktree-path} {branch}`
129
131
  - Re-set git identity: `git -C {worktree-path} config user.name/email` (from state)
130
132
  - Go back to Phase 3
@@ -138,7 +140,7 @@ Before or during user testing, run device-level audits via Bash if user requests
138
140
 
139
141
  | Check | When | Command |
140
142
  | ------------------- | ----------------- | ----------------------------------------- |
141
- | UI flow video | `state.visualEvidence.required` | 3 tiers, cap from `capture-evidence.sh limits` (`visualEvidence.maxVideoSeconds`) - `$HOME/.claude/multi-agent-refs/features/visual-evidence.md` |
143
+ | UI flow video | `state.visualEvidence.required` AND Phase 3 recorded none | `capture-evidence.sh video start` -> drive the flow -> `video stop` -> `fit`. See below |
142
144
  | Accessibility audit | UI changes | `mcp__multi-agent-toolkit__{ios,android}_accessibility_audit` |
143
145
  | Biometric test | Auth flow changes | ios: `mcp__multi-agent-toolkit__ios_biometric` (android: manual) |
144
146
  | Launch time | Perf-sensitive changes | ios: app-launch instrument · android: `mcp__multi-agent-toolkit__android_launch_time` |
@@ -146,6 +148,14 @@ Before or during user testing, run device-level audits via Bash if user requests
146
148
  | Snapshot regression | Component / pixel-stable UI changes | ios: `mcp__multi-agent-toolkit__ios_visual_diff` · android: `mcp__multi-agent-toolkit__android_screenshot` + compare |
147
149
  | Store screenshots | `taskType === screenshot` | ios: `ios_status_bar({preset: "clean"})` · android: `android_screenshot` |
148
150
 
151
+ ##### UI flow video, when Phase 3 produced none
152
+
153
+ Phase 3 Step 3.55 is the primary host and runs in every mode. Phase 5 is the richer one where it exists: the device is up and a person is watching, so the flow is one somebody confirmed. It adds, never replaces.
154
+
155
+ Run only when `visualEvidence.required` and `visualEvidence.video.file` is empty: re-check the device with `$HOME/.claude/scripts/probe-evidence-capability.sh --only device`, then `$HOME/.claude/scripts/capture-evidence.sh video start` -> drive the flow (`run-ui-tests.sh run`, or the Phase 5 scenarios by hand) -> `video stop` -> `fit`. The cap is `visualEvidence.maxVideoSeconds`, read through `capture-evidence.sh limits` so one value serves both. Update `videoTier` and `videoTierReason` with what ran.
156
+
157
+ When the intake answer was `unit` there is no recording here either: overriding it in a phase the user may not be watching makes the question decorative.
158
+
149
159
  Results included in Phase 7 report. MCP tools preferred when available - concise structured output, lower token cost.
150
160
 
151
161
  **Snapshot regression flow (optional):** when the task changes a stable component, capture a screenshot before the change (baseline) and after (current), then call `ios_visual_diff({baseline, current, max_diff_pct: 1.0})`. Threshold can be relaxed for animated / non-deterministic regions - keep `max_diff_pct ≤ 1.0` for static layouts.
@@ -130,6 +130,29 @@ Branch **deterministically**, no implicit fallback. Read `agent-state.json` and
130
130
  11. **NEVER close or resolve the issue** - neither GitHub Issue nor Jira. Issues require team review (4 approvals) before closing. Only post a comment with commit/PR URLs.
131
131
  12. Log: "Phase 6: Commit {sha} - PR #{number}, worktree {removed|kept: <reason>}"
132
132
 
133
+ #### Step 2.9 - Resolve the evidence host (UI changes only)
134
+
135
+ Runs when `state.visualEvidence.required`, before the PR body is written, because the body renders a different shape per host. Record `state.visualEvidence.host` and `hostReason`:
136
+
137
+ 1. **`jira`** - `state.jiraId` present. Upload stills and video with `jira-attach.sh "$JIRA_ID" <files>`; it prints `<filename>\t<url>` per file into `before[]` / `after[]` / `video`.
138
+ 2. **`github-public` / `github-private`** - no Jira, `visualEvidence.githubHost` is `branch`, remote is GitHub. Push the stills to the orphan branch below, then pick the variant from `gh repo view --json isPrivate`.
139
+ 3. **`none`** - anything else, including `githubHost: off`. Publish nothing, record the reason, let the `gaps[]` rule carry it.
140
+
141
+ No API attaches a file to a GitHub issue or PR (the web uploader needs a browser session), so the PR body can only point at something already hosted. An orphan branch is the one mechanism a script has that neither touches the PR diff nor publishes a release.
142
+
143
+ ```bash
144
+ EVB="evidence/$TASK_ID"
145
+ git -C "$WORKTREE" worktree add --detach "$TMP_EV" 2>/dev/null
146
+ git -C "$TMP_EV" checkout --orphan "$EVB" && git -C "$TMP_EV" rm -rf . >/dev/null 2>&1 || true
147
+ cp "$WORKTREE"/.pipeline/evidence/*.png "$TMP_EV"/ 2>/dev/null || true
148
+ git -C "$TMP_EV" add -A && git -C "$TMP_EV" commit -m "evidence: $TASK_ID"
149
+ git -C "$TMP_EV" push -u origin "$EVB"
150
+ ```
151
+
152
+ Stills only. On a GitHub-hosted run no video is recorded (Step 3.55), and an mp4 behind a blob link is a download rather than something a reviewer opens. A failed push is not a phase failure: drop to `host: none` with the error as `hostReason`.
153
+
154
+ Log: `Phase 6 Step 2.9: evidence host = {jira|github-public|github-private|none} ({reason})`
155
+
133
156
  #### Step 3 - PR Description (technical detail for reviewers)
134
157
 
135
158
  **Section set + markup dialect: `channels/pr.md` - read it first.** Phase 7 channels replaces this body with that section set, so build to it. Required reading: `payload-contracts.md`.
@@ -15,7 +15,8 @@
15
15
  "graylog_test": null,
16
16
  "usage_ingest": null,
17
17
  "firebase": null,
18
- "jenkins": null
18
+ "jenkins": null,
19
+ "vercel": null
19
20
  },
20
21
  "tokenScripts": {},
21
22
  "platformIdentityRouting": {},
@@ -1122,6 +1122,71 @@
1122
1122
  "type": ["string", "null"],
1123
1123
  "description": "Directory the worktree's artefacts were salvaged into before removal (agent-state, phase-tracker, triage-output, .pipeline/, build+test logs, review diff). Phase 7 and :resume read from here when worktreePath is gone."
1124
1124
  },
1125
+ "testDepth": {
1126
+ "type": ["string", "null"],
1127
+ "enum": ["unit", "unit+ui", "unit+mcp", null],
1128
+ "description": "How far the run tests, answered at intake because Phase 5 is absent from four of the eight modes and a question asked where it cannot be reached is a question nobody answers. `unit+ui` runs the repo's own UI test and records the screen around it; `unit+mcp` drives the flow through the toolkit MCP instead. The options offered are built from evidenceCapability, never from the model's reading of the repo."
1129
+ },
1130
+ "testDepthSource": {
1131
+ "type": ["string", "null"],
1132
+ "enum": ["user", "autopilot", "default", "forced", null],
1133
+ "description": "Who chose. `forced` means only one option was open, so nothing was asked - recorded rather than passed off as the user's answer."
1134
+ },
1135
+ "evidenceCapability": {
1136
+ "type": ["object", "null"],
1137
+ "additionalProperties": true,
1138
+ "description": "What this machine and this repo can actually produce, measured by probe-evidence-capability.sh BEFORE the test-depth question. Every absent value carries its reason, so a closed option can say why instead of vanishing from the menu; a value that could not be measured is null with a reason, never false, because a probe that did not look and a probe that found nothing are different facts. Contract: multi-agent-refs/features/visual-evidence.md.",
1139
+ "properties": {
1140
+ "platform": { "type": "string", "enum": ["ios", "android", "web", "other"] },
1141
+ "uiTestTarget": {
1142
+ "type": ["string", "null"],
1143
+ "description": "The single chosen target, empty while several candidates exist and no match picks one."
1144
+ },
1145
+ "uiTestTargets": {
1146
+ "type": "array",
1147
+ "items": { "type": "string" },
1148
+ "description": "Every candidate. A real app has many: the reference iOS app has one XCUITest bundle among 477 files that merely sit under a *UITests path, and the reference Android app has eight instrumentation source sets."
1149
+ },
1150
+ "uiTestTargetReason": { "type": ["string", "null"] },
1151
+ "matchingTests": {
1152
+ "type": "array",
1153
+ "items": { "type": "string" },
1154
+ "description": "Tests that mention a changed file's name. A heuristic, and treated as one: an empty set falls to the next tier rather than concluding the screen is untested."
1155
+ },
1156
+ "matchingTestsReason": { "type": ["string", "null"] },
1157
+ "device": { "type": ["string", "null"] },
1158
+ "deviceReason": {
1159
+ "type": ["string", "null"],
1160
+ "description": "'no booted simulator, but one is available to boot' and 'no iOS simulator available on this machine' are different problems with different fixes, and the user can act on only one of them."
1161
+ },
1162
+ "recorder": { "type": ["boolean", "null"] },
1163
+ "recorderReason": { "type": ["string", "null"] },
1164
+ "mcp": { "type": ["boolean", "null"] },
1165
+ "mcpReason": { "type": ["string", "null"] },
1166
+ "tier1": {
1167
+ "type": "string",
1168
+ "enum": ["open", "closed", "unknown"],
1169
+ "description": "Whether the depth menu may offer tier 1. `unknown` means the target was not probed (a --only device re-check), which is not the same as closed and must not be rendered as one."
1170
+ },
1171
+ "tier2": { "type": "string", "enum": ["open", "closed", "unknown"] }
1172
+ }
1173
+ },
1174
+ "uiTest": {
1175
+ "type": ["object", "null"],
1176
+ "additionalProperties": true,
1177
+ "description": "The UI test run that produced the tier 1 recording. Subject to the same default-FAIL rule as the build: a zero exit code alone is not a pass, the log is the evidence, and evidence-gate.mjs reads it.",
1178
+ "properties": {
1179
+ "ran": { "type": "boolean" },
1180
+ "target": { "type": ["string", "null"] },
1181
+ "selected": { "type": "array", "items": { "type": "string" } },
1182
+ "status": { "type": ["string", "null"], "enum": ["passed", "failed", "not-run", null] },
1183
+ "notRunReason": {
1184
+ "type": ["string", "null"],
1185
+ "description": "Why it did not run: no target, no matching test, no device. Each is a reason to fall to the next video tier, never a phase failure."
1186
+ },
1187
+ "logPath": { "type": ["string", "null"] }
1188
+ }
1189
+ },
1125
1190
  "visualEvidence": {
1126
1191
  "type": ["object", "null"],
1127
1192
  "additionalProperties": false,
@@ -1198,10 +1263,23 @@
1198
1263
  },
1199
1264
  "description": "Captured in Phase 3 after the build+test gate, because Phase 5 is dropped by every autopilot and --local entry."
1200
1265
  },
1266
+ "host": {
1267
+ "type": ["string", "null"],
1268
+ "enum": ["jira", "github-public", "github-private", "none", null],
1269
+ "description": "Where the artefacts are published, resolved in Phase 6. `jira` attaches both stills and video. `github-public` pushes the stills to the evidence branch and embeds them in the PR body. `github-private` pushes the same stills but the PR carries a blob permalink instead of an inline image, because GitHub's image proxy cannot fetch a private repo's raw URL and an embedded one renders broken for every reader. `none` publishes nothing and records the gap. Video is Jira-only by decision: without an attachment host there is nothing a recording can be attached to."
1270
+ },
1271
+ "hostReason": {
1272
+ "type": ["string", "null"],
1273
+ "description": "Why this host and not the one above it in the order. A host of `none` with no reason is the silence the Phase 6 blocker exists to catch."
1274
+ },
1201
1275
  "videoTier": {
1202
1276
  "type": ["integer", "null"],
1203
1277
  "enum": [1, 2, 3, null],
1204
- "description": "1 = the repo's own UI test target drove the flow, 2 = MCP-driven flow, 3 = no runnable build."
1278
+ "description": "1 = the repo's own UI test target drove the flow, 2 = MCP-driven flow, 3 = no recording. Resolved from the capability probe, then RE-CHECKED at capture time: a device booted at intake can be gone by Phase 3, and a tier recorded from a stale measurement is a promise the run cannot keep."
1279
+ },
1280
+ "videoTierReason": {
1281
+ "type": ["string", "null"],
1282
+ "description": "Which rule produced the tier, and the tier it came down from when it was downgraded at capture time, e.g. 'tier 1 -> 2: simulator no longer booted'."
1205
1283
  },
1206
1284
  "video": {
1207
1285
  "type": "object",
@@ -198,6 +198,10 @@
198
198
  "type": ["string", "null"],
199
199
  "description": "NPM registry token (npm.pkg.github.com or npmjs.com). Used by package publish flow."
200
200
  },
201
+ "vercel": {
202
+ "type": ["string", "null"],
203
+ "description": "Keychain item holding the Vercel access token. vercel-deploy.sh resolves it through credential-store.sh so the value never reaches argv, where the CLI would leak it on a retry. Absent = deploys need VERCEL_TOKEN in the environment."
204
+ },
201
205
  "appstore_connect_key_id": {
202
206
  "type": ["string", "null"],
203
207
  "description": "App Store Connect API key ID. Tier 1 of the App Store Connect access chain, used by /multi-agent:store-ready Gate 2 (and its iOS alias /multi-agent:testflight-validation). An identifier rather than a secret; mapped anyway so every credential is read through the same layer. Creating an API key needs an Admin or App Manager role, which is why Tier 2 exists."
@@ -1865,7 +1869,26 @@
1865
1869
  "minimum": 5,
1866
1870
  "maximum": 300,
1867
1871
  "default": 60,
1868
- "description": "Recording cap. A flow needing longer is a debugging session, not a review artefact."
1872
+ "description": "Recording cap. A flow needing longer is a debugging session, not a review artefact. Android's `screenrecord` has its own ceiling of 180s that no setting can lift, so capture-evidence.sh clamps to it and says when it did: one preference honoured on one platform and silently halved on the other is worse than a stated limit."
1873
+ },
1874
+ "githubHost": {
1875
+ "type": "string",
1876
+ "enum": ["branch", "off"],
1877
+ "default": "branch",
1878
+ "description": "Where stills go on a run with no Jira. `branch` pushes them to an orphan `evidence/<task-id>` branch so the PR body can show them; GitHub has no API that attaches a file to an issue or a PR, and the web uploader needs a browser session, so a branch is the only mechanism a script has. `off` means no upload, the PR names the artefact directory, and the absence is recorded as a gap. Video is never pushed here: an mp4 behind a blob link is a download, not evidence a reviewer will open."
1879
+ }
1880
+ }
1881
+ },
1882
+ "testDepth": {
1883
+ "type": "object",
1884
+ "additionalProperties": false,
1885
+ "description": "How far a run tests by default. The question is asked at intake (Phase 0), not in Phase 5, because Phase 5 is not in the phase set for autopilot or either local mode. Autopilot never asks and reads `default` instead; the run records which rule fired.",
1886
+ "properties": {
1887
+ "default": {
1888
+ "type": "string",
1889
+ "enum": ["unit", "unit+ui", "unit+mcp"],
1890
+ "default": "unit+ui",
1891
+ "description": "Used by both autopilot entries, and as the preselected option for everyone else. An option the capability probe found closed is never taken from here: a default that cannot run is not a default."
1869
1892
  }
1870
1893
  }
1871
1894
  },