@mmerterden/multi-agent-pipeline 12.10.0 → 13.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (42) hide show
  1. package/CHANGELOG.md +179 -0
  2. package/README.md +24 -7
  3. package/index.js +5 -2
  4. package/install/_codex-agents.mjs +211 -0
  5. package/install/_codex-instructions.mjs +33 -0
  6. package/install/_managed-block.mjs +99 -0
  7. package/install/codex.mjs +478 -0
  8. package/install/copilot.mjs +34 -80
  9. package/install/index.mjs +25 -9
  10. package/install/templates/codex-instructions.md +45 -0
  11. package/package.json +5 -3
  12. package/pipeline/claude-md-template.md +1 -0
  13. package/pipeline/commands/multi-agent/SKILL.md +1 -1
  14. package/pipeline/commands/multi-agent/dev/SKILL.md +31 -6
  15. package/pipeline/commands/multi-agent/finish/SKILL.md +1 -1
  16. package/pipeline/commands/multi-agent/language/SKILL.md +2 -2
  17. package/pipeline/commands/multi-agent/setup/SKILL.md +69 -2
  18. package/pipeline/commands/multi-agent/sync/SKILL.md +128 -5
  19. package/pipeline/commands/multi-agent/testflight-validation/SKILL.md +219 -0
  20. package/pipeline/commands/multi-agent/update/SKILL.md +7 -4
  21. package/pipeline/multi-agent-refs/_input-parser.md +1 -1
  22. package/pipeline/multi-agent-refs/cross-cli-contract.md +51 -17
  23. package/pipeline/multi-agent-refs/features/model-fallback.md +29 -0
  24. package/pipeline/multi-agent-refs/features/review-multi-repo.md +1 -1
  25. package/pipeline/multi-agent-refs/phases/log-format.md +1 -1
  26. package/pipeline/multi-agent-refs/phases/phase-0-init.md +14 -2
  27. package/pipeline/multi-agent-refs/phases/phase-4-review.md +36 -5
  28. package/pipeline/multi-agent-refs/progress-contract.md +1 -1
  29. package/pipeline/multi-agent-refs/tracker-contract.md +17 -1
  30. package/pipeline/schemas/prefs.schema.json +296 -62
  31. package/pipeline/schemas/reviewer-output.schema.json +1 -1
  32. package/pipeline/schemas/triage-output.schema.json +1 -1
  33. package/pipeline/scripts/cost-table.json +15 -1
  34. package/pipeline/scripts/output-quality-check.sh +34 -0
  35. package/pipeline/scripts/phase-tracker.sh +79 -1
  36. package/pipeline/scripts/smoke-cross-cli-behavior.sh +25 -10
  37. package/pipeline/scripts/uninstall.mjs +105 -9
  38. package/pipeline/scripts/update-check.sh +2 -1
  39. package/pipeline/skills/shared/core/multi-agent-language/SKILL.md +2 -2
  40. package/pipeline/skills/shared/core/multi-agent-setup/SKILL.md +48 -1
  41. package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +87 -6
  42. package/pipeline/skills/shared/core/multi-agent-testflight-validation/SKILL.md +120 -0
@@ -0,0 +1,219 @@
1
+ ---
2
+ description: "Pre-submission validation for a TestFlight / App Store build (iOS, local-only). Three gates: static archive audit, Apple's own `altool --validate-app`, and a Review-Guidelines check. ITMS codes are mapped to the rule each implies. Validates only, never uploads. Use when a build is about to go to TestFlight, or a submission was rejected and you need why."
3
+ description-tr: "TestFlight / App Store yüklemesi öncesi doğrulama (iOS, yalnızca lokal). Repo + branch seç, sonra ya build'i sen ver ya da koşu archive alsın; üç kapıyı geç: statik 18-kurallı archive denetimi, Apple'ın kendi `altool --validate-app`'i, ve App Store Review Guidelines'a karşı guideline incelemesi. Her kapı için ayrı verdict, ITMS kodları ilgili kurala eşlenmiş. Sadece doğrular - asla yüklemez."
4
+ argument-hint: "[repo] - empty = pick from prefs; repo name or path; --ipa=<path>; --archive=<path>; --resume"
5
+ allowed-tools: Agent, Bash, Read, Write, Edit, Glob, Grep, TaskCreate, TaskUpdate, AskUserQuestion, Skill, mcp__dev-toolkit__ios_app_store_audit, mcp__dev-toolkit__ios_export_ipa, mcp__dev-toolkit__ios_testflight_validate, mcp__dev-toolkit__ios_xcodebuild, mcp__dev-toolkit__ios_xcresult
6
+ ---
7
+
8
+ # /multi-agent:testflight-validation - pre-submission validation
9
+
10
+ Catch, before you upload, what App Store Connect would send back after you do.
11
+
12
+ **Local-only.** No commits, no push, no PR, no channels. The worktree exists only
13
+ to archive without touching your working tree.
14
+
15
+ **It never uploads.** Only `--validate-app` is ever invoked, never `--upload-app`.
16
+ A validation run must not be able to ship a build by accident.
17
+
18
+ ## Why three gates and not one
19
+
20
+ Each gate sees something the others structurally cannot. Reporting one of them as
21
+ "the check" is how a build passes locally and gets rejected anyway.
22
+
23
+ | Gate | What runs | Needs | Sees | Blind to |
24
+ |---|---|---|---|---|
25
+ | **1. Static** | `ios_app_store_audit` (18 rules, real ITMS codes) | an `.xcarchive` | privacy manifest, required-reason API, Info.plist, code signing, entitlements, embedded SDK, IPv6, debug-tool leak, binary size | anything that depends on the App Store Connect account |
26
+ | **2. Authoritative** | `ios_testflight_validate` → `altool --validate-app` | an `.ipa` + credentials | unregistered bundle ID, profile that does not match the app record, **a version+build pair already used**, entitlements not provisioned for the App ID | the Review Guidelines - Apple's validator does not read them |
27
+ | **3. Guideline** | `app-store-review` skill + repo evidence | repo checkout | ATT flow, privacy policy, account deletion, IAP rules, purpose-string wording, permission justification | anything not visible in source |
28
+
29
+ Gate 2 is the only one that asks Apple, and Gate 3 is the only one that covers the
30
+ rejections a human reviewer writes. Most "we passed validation and still got
31
+ rejected" cases are Gate 3 findings.
32
+
33
+ ## Step 0 - parse input
34
+
35
+ | Input | Meaning |
36
+ |---|---|
37
+ | (empty) | ask which repo (Step 1) |
38
+ | `my-ios-app` or a path | that repo |
39
+ | `--ipa=<path>` | Mode B with an `.ipa`; Gate 1 cannot run (see Step 3) |
40
+ | `--archive=<path>` | Mode B with an `.xcarchive`; all three gates run |
41
+ | `--resume` | continue the last run from its state file |
42
+
43
+ State lives at `$HOME/.claude/logs/multi-agent/<task_id>/agent-state.json` with
44
+ `taskId = TFV-<repo>-<yyyymmddHHMM>`. Register phases with the tracker
45
+ (`$HOME/.claude/multi-agent-refs/tracker-contract.md`) so `:resume` and `:status`
46
+ work like any other run.
47
+
48
+ ## Step 1 - pickers (native, always)
49
+
50
+ Use `AskUserQuestion` for every step - never a numbered text menu. Questions and
51
+ descriptions render in `prefs.global.outputLanguage`; `label` and `header` stay
52
+ English, per `$HOME/.claude/multi-agent-refs/picker-contract.md`. Print the
53
+ `Step <i>/<n>: <what this decides>` breadcrumb for each.
54
+
55
+ 1. **Repo** - from `prefs.projects` where the stack is iOS. A single match
56
+ auto-resolves (say so in the breadcrumb, do not silently skip the step).
57
+ 2. **Branch** - the branch to validate. Resolution order:
58
+ - `git fetch --prune` first, capturing **stderr**. **If the fetch fails, do not
59
+ silently fall back to a cached ref**, and **classify before naming a cause** -
60
+ the same rule as the `/multi-agent:dev` remote gate:
61
+
62
+ | stderr contains | Cause | Remedy |
63
+ |---|---|---|
64
+ | `could not read Password`, `Authentication failed`, `403` | credential | store the PAT in the credential helper or switch the remote to SSH. **A VPN cannot fix this**, and the base ref being stale is unrelated to what broke - do not offer the cached-ref fallback. |
65
+ | `Could not resolve host`, `Operation timed out`, `Connection refused` | network | retry / continue on the cached ref with an explicit warning / switch remote / abort, per `$HOME/.claude/multi-agent-refs/rules.md` |
66
+ | `Repository not found`, `404` | wrong remote | show `git remote -v` and ask |
67
+
68
+ Always print the observed stderr line next to the classification. Asserting
69
+ `unreachable (VPN/DNS)` for a missing-credential error that returns in under a
70
+ second sends the user to fix something that was never broken.
71
+ - Offer the current branch, the default branch, and any `release/*` /
72
+ `tkdevelop/*` heads.
73
+ 3. **Mode** - how the build is obtained:
74
+ - `Supply a build` (Mode B, default) - fastest, no signing needed in-run.
75
+ - `Archive from this branch` (Mode A) - needs a distribution certificate and
76
+ profile in the keychain, and takes as long as a release archive.
77
+
78
+ ## Step 2 - pre-flight, before anything expensive
79
+
80
+ Report every line; a missing prerequisite is a halt, not a warning.
81
+
82
+ ```bash
83
+ xcrun --find altool >/dev/null 2>&1 || echo "MISSING: altool (install Xcode)"
84
+ xcodebuild -version | head -1
85
+ ```
86
+
87
+ Then resolve credentials, and **state which tier is active in the report**:
88
+
89
+ | Tier | Source | Effect |
90
+ |---|---|---|
91
+ | 1 | ASC API key - key id + issuer id from the keychain via `prefs.global.keychainMapping`, `.p8` at `~/.appstoreconnect/private_keys/AuthKey_<keyId>.p8` | Gate 2 runs |
92
+ | 2 | Apple ID + app-specific password, referenced as a keychain item | Gate 2 runs |
93
+ | 3 | neither | **Gate 2 reports `SKIPPED`, and the run says so in the verdict line** |
94
+
95
+ Credentials come from `/multi-agent:setup`; never prompt for a secret value in
96
+ chat. If nothing is configured, tell the user which of the two tiers they can set
97
+ up and that tier 2 needs no elevated App Store Connect role.
98
+
99
+ Multi-provider accounts need `--provider-public-id`. When it is not in prefs, run
100
+ `ios_testflight_validate({list_providers: true})` once and ask which provider.
101
+
102
+ ## Step 3 - obtain the build
103
+
104
+ ### Mode B - a build you supply
105
+
106
+ - `.xcarchive` → Gate 1 runs on it. To reach Gate 2 the archive must be exported,
107
+ so run `ios_export_ipa` (see Mode A step 3 for the signing inputs).
108
+ - `.ipa` only → **Gate 1 is reported `SKIPPED (needs .xcarchive)`.** The static
109
+ audit reads archive structure that an `.ipa` does not carry. Do not present a
110
+ two-gate run as a full pass; say which gate did not run and why, and offer to
111
+ re-run with the archive.
112
+
113
+ ### Mode A - archive from the branch
114
+
115
+ 1. Worktree at `{projectRoot}/{worktreeBasePath}/{taskId}` on the chosen branch.
116
+ **Never under `$HOME`**, never a direct checkout of the main working tree.
117
+ 2. Resolve the scheme and workspace/project from prefs; ask if ambiguous.
118
+ 3. Archive:
119
+ `ios_xcodebuild({workspace|project, scheme, action: "archive", configuration: "Release", destination: "generic/platform=iOS"})`
120
+ Note the destination: the simulator default would produce an archive that
121
+ cannot be exported for distribution.
122
+ 4. Export:
123
+ `ios_export_ipa({archive_path, output_dir, method: "app-store-connect", team_id, provisioning_profiles?, signing_style?})`
124
+ Leave `allow_provisioning_updates` off unless the user asks: it lets xcodebuild
125
+ create or modify profiles in the developer account, which a validation run has
126
+ no business doing.
127
+ 5. A failed export halts with the parsed errors. The usual causes are a missing
128
+ distribution certificate, a profile that does not match the bundle ID, or
129
+ `signing_style: "manual"` with no `provisioning_profiles` map.
130
+
131
+ ## Step 4 - Gate 1, static audit
132
+
133
+ `ios_app_store_audit({archive_path, rules: "all"})`.
134
+
135
+ `error` findings are blocking; `warning` is advisory. Group the output by severity
136
+ and keep each finding's ITMS code - Gate 2 may return the same code, and seeing
137
+ it in both places tells the user it is real rather than a heuristic.
138
+
139
+ ## Step 5 - Gate 2, Apple's own validation
140
+
141
+ `ios_testflight_validate({ipa_path, platform: "ios", <credential args>})`.
142
+
143
+ Render the verdict exactly as returned:
144
+
145
+ - `PASS` - Apple accepted the binary for delivery.
146
+ - `FAIL` - list each issue with its ITMS code, the mapped guideline, and the hint.
147
+ - `SKIPPED` - print the reason. **Never render this as a pass.** The verdict line
148
+ for the whole run must read `2 of 3 gates cleared, 1 skipped`, not `passed`.
149
+
150
+ Gate 2 is the only gate that catches a build number already used - the most
151
+ common wasted upload - so when it fails on that, say so plainly and name the next
152
+ free build number.
153
+
154
+ ## Step 6 - Gate 3, guideline review
155
+
156
+ Load the `app-store-review` skill and review the repo against it. This is the gate
157
+ that catches what a human reviewer rejects, so it reads source, not the binary:
158
+
159
+ | Area | Evidence to gather |
160
+ |---|---|
161
+ | Purpose strings | every `NS*UsageDescription` in Info.plist - present, specific, user-facing, and matching what the code actually does with the data |
162
+ | Privacy manifest | `PrivacyInfo.xcprivacy` exists, declares required-reason APIs, and matches the SDKs actually linked |
163
+ | Tracking | if any tracking API or SDK is present, an ATT prompt exists and runs before collection |
164
+ | Account deletion | if the app creates accounts, an in-app deletion path exists (guideline 5.1.1(v)) |
165
+ | Privacy policy | reachable in-app and in the metadata |
166
+ | IAP | anything unlocking features goes through StoreKit, with no external purchase path |
167
+ | Sign in with Apple | present when a third-party social login is offered |
168
+
169
+ For each: `pass` / `fail` / `not-applicable` with the evidence path that justifies
170
+ it. `not-applicable` needs a reason - an unexamined area is not a pass.
171
+
172
+ ## Step 7 - report
173
+
174
+ Write to `~/TestFlightChecks/<repo>-<branch>-<timestamp>/report.md` and print a
175
+ summary. Structure:
176
+
177
+ ```
178
+ Verdict: <N> of 3 gates cleared[, <M> skipped]
179
+ Build: <ipa or archive path> · <bundle id> <version> (<build>)
180
+ Auth: tier <1|2|none> · <method>
181
+
182
+ Gate 1 static audit PASS | FAIL (<n> blocking, <n> advisory) | SKIPPED (<reason>)
183
+ Gate 2 Apple validation PASS | FAIL (<n> issues) | SKIPPED (<reason>)
184
+ Gate 3 guideline review PASS | FAIL (<n> findings) | <n> not-applicable
185
+
186
+ Blocking - fix before uploading
187
+ [ITMS-90683] Info.plist: NSCameraUsageDescription missing
188
+ guideline 5.1.1 Data Collection and Storage
189
+ <hint>
190
+ <file:line>
191
+
192
+ Advisory
193
+ ...
194
+
195
+ Not run
196
+ Gate 1: needs an .xcarchive; only an .ipa was supplied
197
+ ```
198
+
199
+ Rules for the report:
200
+
201
+ - A skipped gate is never folded into the pass count. The verdict line states the
202
+ skip.
203
+ - Every blocking finding carries a file path or an ITMS code. A finding the user
204
+ cannot act on is noise.
205
+ - No AI or assistant attribution anywhere, per
206
+ `$HOME/.claude/rules/git-conventions.md`.
207
+ - Real newlines, no HTML entities, per
208
+ `$HOME/.claude/rules/pipeline-output-formatting.md`.
209
+
210
+ ## Step 8 - offer the next action, do not take it
211
+
212
+ Print, and stop:
213
+
214
+ - the exact `xcrun altool --upload-app` command for when the gates are clear, so
215
+ uploading stays an explicit human act
216
+ - `/multi-agent:testflight-validation --resume` to re-run after fixes
217
+ - `/multi-agent:fix-bug` when Gate 3 produced code-level findings
218
+
219
+ Never upload, never bump the build number, never commit.
@@ -52,15 +52,18 @@ Update the pipeline in one command. Existing preferences are preserved; only ski
52
52
  fi
53
53
  ```
54
54
 
55
- 4c. **Prune retired adapter files** (Cursor / Antigravity / Codex / Copilot Chat were removed in v10.7.0 - the pipeline now targets Claude Code + Copilot CLI only):
55
+ 4c. **Prune retired adapter files** (Cursor / Antigravity / Copilot Chat were removed in v10.7.0 - the pipeline targets Claude Code, Copilot CLI and Codex CLI):
56
56
  ```bash
57
- # Global Codex adapter prompt is safe to remove (no longer produced).
58
- rm -f "$HOME/.codex/prompts/multi-agent.md"
59
- echo " -> pruned retired global Codex adapter prompt (if present)"
60
57
  # Per-project adapter files (.cursor/, .agent/, .github/copilot-instructions.md)
61
58
  # live in your repos and are left untouched - remove them manually if you like.
59
+ echo " -> no global adapter files to prune"
62
60
  ```
63
61
 
62
+ > **Do NOT delete `$HOME/.codex/prompts/multi-agent.md`.** Releases up to
63
+ > v12.11.0 removed it here as a retired v9.7.0 adapter leftover. Codex CLI is a
64
+ > supported target again as of v13.0.0 and the installer writes that file, so
65
+ > deleting it silently breaks the `/multi-agent` slash command on Codex.
66
+
64
67
  5. **Migrate preferences** (if there is an old schema):
65
68
  ```bash
66
69
  if [ -f "$HOME/.claude/scripts/migrate-prefs.mjs" ]; then
@@ -6,7 +6,7 @@ description: "Internal - input type detection for multi-agent dispatcher."
6
6
 
7
7
  The top-level `multi-agent` command classifies user arguments per the rules below.
8
8
 
9
- > **Language**: Schema reference only - no user-facing text. Picker UI strings always English (`promptLanguage` is locked to `"en"`). Commit/PR/Jira payloads stay English.
9
+ > **Language**: Schema reference only - no user-facing text. Picker `label` + `header` always English (`promptLanguage` is locked to `"en"`); `question` + `description` follow `outputLanguage` per the `rules.md` matrix. Commit/PR/Jira payloads stay English.
10
10
 
11
11
  ## Type Table
12
12
 
@@ -1,19 +1,19 @@
1
- # Cross-CLI Contract (Claude Code ↔ Copilot CLI)
1
+ # Cross-CLI Contract (Claude Code · Copilot CLI · Codex CLI)
2
2
 
3
3
  > **Non-negotiable**. Any change that breaks this contract blocks merge. Validated by `smoke-cross-cli-behavior.sh`.
4
4
 
5
- **Purpose**: every pipeline command must produce identical artifacts (state, logs, outputs) and respect identical placeholder vocabulary regardless of which CLI invokes it. This file is the source of truth for "what must stay the same."
5
+ **Purpose**: every pipeline command must produce identical artifacts (state, logs, outputs) and respect identical placeholder vocabulary regardless of which of the three host CLIs invokes it. This file is the source of truth for "what must stay the same."
6
6
 
7
7
  ---
8
8
 
9
- ## 1. Command Inventory (42 commands)
9
+ ## 1. Command Inventory (43 commands)
10
10
 
11
11
  ```
12
12
  analysis, analysis-resolve, autopilot, build-optimize, channels, create-jira, design-check, dev,
13
13
  dev-autopilot, dev-local, dev-local-autopilot, diff-explain, finish, forget, garbage-collect,
14
14
  help, issue, jira, kill, language, local,
15
15
  local-autopilot, log, manual-test, prune-logs, purge, refactor, resume, review, review-issue, review-jira,
16
- routines, save, scan, search, setup, stack, status, sync, test, uninstall, update
16
+ routines, save, scan, search, setup, stack, status, sync, test, testflight-validation, uninstall, update
17
17
  ```
18
18
 
19
19
  Categories:
@@ -24,6 +24,7 @@ Categories:
24
24
  - **Fast modes** (Init -> Dev(Opus) -> Commit -> Report): `dev`, `dev-autopilot`, `dev-local`, `dev-local-autopilot`
25
25
  - **Tail modes** (run the pipeline tail over already-done local work): `finish`
26
26
  - **Ops commands** (one-shot, no worktree): `status`, `log`, `kill`, `purge`, `uninstall`, `resume`, `review`, `review-jira`, `review-issue`, `analysis`, `analysis-resolve`, `build-optimize`, `channels`, `scan`, `search`, `diff-explain`, `garbage-collect`, `prune-logs`
27
+ - **Local audits** (worktree only to build; no commit, push, PR or channels): `design-check`, `testflight-validation`. `testflight-validation` additionally never invokes `altool --upload-app` - a validation run must not be able to ship a build by accident.
27
28
  - **Meta-ops**: `setup`, `sync`, `update`, `help`, `refactor`, `test`, `stack`, `manual-test`, `language`
28
29
  - **Routines** (user-defined routine registry; the routines they create are local-only and never synced): `save`, `routines`, `forget`
29
30
 
@@ -110,14 +111,43 @@ Legacy names found during the v3.7 audit - replaced per this contract.
110
111
 
111
112
  ---
112
113
 
113
- ## 2.6 Intentional structural divergence - `multi-agent.md` vs `multi-agent/SKILL.md`
114
+ ## 2.6 Intentional structural divergence - thin dispatcher vs inlined orchestrator
114
115
 
115
- The top-level orchestrator has two deliberately different shapes on the two CLIs. This is **not drift** - it reflects a real capability gap between the host CLIs, and auditors must not flag it as a parity violation.
116
+ The top-level orchestrator has two deliberately different shapes. This is **not
117
+ drift** - it reflects a real capability gap between the host CLIs, and auditors
118
+ must not flag it as a parity violation.
116
119
 
117
120
  | Surface | File | Shape | Why |
118
121
  |---|---|---|---|
119
- | Claude Code (colon-form) | `pipeline/commands/multi-agent.md` | ~300 lines - thin dispatcher that routes to `$HOME/.claude/multi-agent-refs/phases/phase-N-*.md` on demand | Claude Code supports lazy-loaded reference files (`@refs/phases/...`), so only the active phase's docs enter the context window. Saves tokens. |
120
- | Copilot CLI (dash-form) | `pipeline/skills/shared/core/multi-agent/SKILL.md` | ~830 lines - full inline orchestrator with all 8 phase specs embedded | Copilot CLI loads the whole SKILL.md into context once the skill is dispatched; it doesn't have an equivalent of Claude's ref-file loading pattern. Embedding keeps behavior identical without relying on a feature Copilot lacks. |
122
+ | Claude Code (colon-form) | `pipeline/commands/multi-agent/SKILL.md` | ~300 lines - thin dispatcher that routes to `$HOME/.claude/multi-agent-refs/phases/phase-N-*.md` on demand | Claude Code lazy-loads reference files, so only the active phase's docs enter the context window. Saves tokens. |
123
+ | Copilot CLI (dash-form) | `pipeline/skills/shared/core/multi-agent/SKILL.md` | ~830 lines - full inline orchestrator with all 8 phase specs embedded | Copilot CLI loads the whole SKILL.md once the skill is dispatched; it has no equivalent of Claude's ref-file loading. Embedding keeps behavior identical without relying on a feature Copilot lacks. |
124
+ | Codex CLI | installed as `~/.codex/skills/multi-agent/SKILL.md`, generated from the Claude dispatcher | thin dispatcher, refs under `~/.codex/multi-agent-refs/` | Codex reads files on demand, so it takes the Claude shape - **and it has to.** See below. |
125
+
126
+ ### Why Codex takes the thin-dispatcher shape, and must keep it
127
+
128
+ Codex assembles every discovered skill's name + description into a single prompt
129
+ block and **silently drops entries when that block overflows**. Measured against
130
+ Codex 0.145 during the install design: installing one plugin that declares 142
131
+ skills took the block from 11 skills / 4,710 bytes to 83 skills / 22,111 bytes -
132
+ only **75 of the 142** surfaced, **and an unrelated user-scope skill was evicted**.
133
+ Removing the plugin brought it back.
134
+
135
+ So shipping the 43 sub-commands as peer skills on Codex would silently lose
136
+ pipeline commands next to any stack toolkit, with no error anywhere. The pipeline
137
+ therefore contributes **exactly one** skill on Codex (`multi-agent`) and keeps the
138
+ 43 sub-command specs as reference files that cost nothing until read.
139
+
140
+ **Do not "fix" this by adding per-command skills on Codex.** The layout is
141
+ capability-derived, and `smoke-install-layout.sh` fails if the Codex skills tree
142
+ gains a second pipeline entry.
143
+
144
+ ### Parity axis differs per host
145
+
146
+ Claude Code and Copilot CLI are compared on their **skill directory sets**. Codex is
147
+ compared on its **ref set**: the 43 command specs must all exist under
148
+ `~/.codex/multi-agent-refs/commands/<cmd>/SKILL.md`, and
149
+ `smoke-codex-install.sh` asserts the count against the source tree. Comparing Codex
150
+ on skill directories would demand exactly the layout that breaks it.
121
151
 
122
152
  **What must stay identical** (byte-level) across the two files:
123
153
 
@@ -175,13 +205,17 @@ argument-hint: "<input hint>"
175
205
 
176
206
  ### 4.1 TaskCreate ↔ phase-tracker.sh
177
207
 
178
- | Concept | Claude Code | Copilot CLI |
179
- |---|---|---|
180
- | Register a phase | `TaskCreate` tool call with subject/description | `phase-tracker.sh add <N> <name>` |
181
- | Mark a phase in-progress | `TaskUpdate` → `in_progress` | `phase-tracker.sh update <N> in_progress` |
182
- | Mark a phase complete | `TaskUpdate` → `completed` | `phase-tracker.sh update <N> completed` |
183
- | Mark a phase failed | `TaskUpdate` → `completed` + log failure in agent-log | `phase-tracker.sh update <N> failed` |
184
- | Sub-phase progress | TaskCreate with `addBlockedBy` | `phase-tracker.sh sub <N> <sub> <name> <status>` |
208
+ | Concept | Claude Code | Copilot CLI | Codex CLI |
209
+ |---|---|---|---|
210
+ | Register a phase | `TaskCreate` tool call with subject/description | `phase-tracker.sh add <N> <name>` | `update_plan` step, `status: pending` |
211
+ | Mark a phase in-progress | `TaskUpdate` → `in_progress` | `phase-tracker.sh update <N> in_progress` | `update_plan` step → `in_progress` |
212
+ | Mark a phase complete | `TaskUpdate` → `completed` | `phase-tracker.sh update <N> completed` | `update_plan` step → `completed` |
213
+ | Mark a phase failed | `TaskUpdate` → `completed` + log failure in agent-log | `phase-tracker.sh update <N> failed` | `update_plan` step → `completed` + failure in agent-log |
214
+ | Sub-phase progress | TaskCreate with `addBlockedBy` | `phase-tracker.sh sub <N> <sub> <name> <status>` | `phase-tracker.sh sub ...` (the plan tool has no nesting) |
215
+
216
+ `update_plan` takes the FULL step list, not a delta, so each boundary rewrites the
217
+ whole plan. It must never be called in parallel with another tool, and it is
218
+ unavailable in Codex plan mode - fall back to `phase-tracker.sh render` there.
185
219
 
186
220
  ### 4.2 Contract
187
221
 
@@ -227,9 +261,9 @@ Modifier flags are orthogonal and compose:
227
261
 
228
262
  ## 7. Platform Guards (macOS / Linux / WSL)
229
263
 
230
- Commands invoked via Copilot CLI may run on Linux or WSL. Claude Code currently macOS-only. Shell code in `pipeline/skills/shared/core/` command files and in `pipeline/lib` / `pipeline/scripts` must be portable.
264
+ Commands invoked via Copilot CLI or Codex CLI may run on Linux or WSL. Claude Code currently macOS-only. Shell code in `pipeline/skills/shared/core/` command files and in `pipeline/lib` / `pipeline/scripts` must be portable.
231
265
 
232
- For Keychain I/O the canonical path is **`~/.claude/lib/credential-store.sh`** (or `~/.copilot/lib/credential-store.sh` on Copilot installs). The shell driver detects platform internally and auto-delegates to `keychain.py` (Python helper, macOS / Linux) or PowerShell `CredentialManager` (Windows). Call sites stay platform-agnostic - no per-OS branching needed.
266
+ For Keychain I/O the canonical path is **`~/.claude/lib/credential-store.sh`** (or `~/.copilot/lib/credential-store.sh` / `~/.codex/lib/credential-store.sh` on those installs). The shell driver detects platform internally and auto-delegates to `keychain.py` (Python helper, macOS / Linux) or PowerShell `CredentialManager` (Windows). Call sites stay platform-agnostic - no per-OS branching needed.
233
267
 
234
268
  | Purpose | Canonical (cross-platform) | Underlying backend (for reference / debugging) |
235
269
  |---|---|---|
@@ -101,3 +101,32 @@ per-phase `model` field already carries the override).
101
101
  reviewer models (GPT-5.4 + Opus + Sonnet - Fable 5 is not offered there) and
102
102
  does not use this persona ladder. Only Claude Code dispatches Reviewer-1 on
103
103
  Fable.
104
+
105
+ ## Codex CLI
106
+
107
+ Codex offers no Anthropic models, so the tier names map onto OpenAI models plus a
108
+ reasoning effort - effort carries the depth distinction that the model id carries
109
+ on Claude Code. The map is applied at install time by
110
+ `install/_codex-agents.mjs`, which writes `model` + `model_reasoning_effort` into
111
+ each `~/.codex/agents/<persona>.toml`:
112
+
113
+ | Tier | Codex model | Effort |
114
+ |---|---|---|
115
+ | `fable` | `gpt-5.6` | `xhigh` |
116
+ | `opus` | `gpt-5.6` | `high` |
117
+ | `sonnet` | `gpt-5.4` | `medium` |
118
+ | `haiku` | `gpt-5.6-terra` | `low` |
119
+
120
+ Ladder on Codex: `gpt-5.6 @ xhigh -> gpt-5.6 @ high -> gpt-5.4 -> gpt-5.6-terra`.
121
+ The first step down lowers effort rather than switching model, which is the
122
+ cheapest useful degradation when the top tier is rate-limited rather than
123
+ unavailable.
124
+
125
+ **Per-dispatch override on Codex** goes through `spawn_agent`, and it MUST pass
126
+ `fork_turns: "none"` (or a positive integer). A full-history fork inherits the
127
+ parent model and reasoning effort and silently discards the override, so a
128
+ fallback that omits it appears to apply while changing nothing.
129
+
130
+ The three model ids above live in exactly three places - this table,
131
+ `pipeline/scripts/cost-table.json`, and the Phase 4 reviewer matrix - so an OpenAI
132
+ rename is a three-file change.
@@ -20,7 +20,7 @@ for proj in $(jq -r '.projects[] | "\(.name)\t\(.worktreePath)\t\(.baseBranch)"'
20
20
  done
21
21
  ```
22
22
 
23
- Same reviewer set (Fable-or-Opus / GPT-5.4 / Sonnet) receive `COMBINED_DIFF` with a multi-repo prefix in the system prompt:
23
+ Same reviewer set (per host: Fable+Sonnet on Claude Code, Opus/GPT-5.4/Sonnet on Copilot CLI, gpt-5.6/gpt-5.4/gpt-5.6 on Codex CLI) receive `COMBINED_DIFF` with a multi-repo prefix in the system prompt:
24
24
 
25
25
  ```
26
26
  This is a multi-repo task spanning {N} repos: {repo names}.
@@ -101,7 +101,7 @@ Every phase that dispatches a billable LLM agent MUST forward its token totals t
101
101
 
102
102
  ```bash
103
103
  pipeline/scripts/log-metric.sh "$TASK_ID" <phase-id> <event> \
104
- model=<fable|opus|sonnet|haiku|gpt-5.4> tokens_in=$IN tokens_out=$OUT tokens_cached=$CACHED duration_ms=$DUR
104
+ model=<fable|opus|sonnet|haiku|gpt-5.4|gpt-5.6|gpt-5.6-terra> tokens_in=$IN tokens_out=$OUT tokens_cached=$CACHED duration_ms=$DUR
105
105
  LOG_METRIC_FORWARD_TO_TRACKER=1 pipeline/scripts/log-metric.sh "$TASK_ID" <phase-id> tokens \
106
106
  model=<...> tokens_in=$IN tokens_out=$OUT tokens_cached=$CACHED
107
107
  ```
@@ -300,11 +300,23 @@ Branch name is deterministic - no user confirmation needed.
300
300
  6. If empty after kebab (e.g. all-emoji title) → fall back to `task-{shortId}`
301
301
 
302
302
  **Collision handling** (automatic - no prompt):
303
- - Probe local + remote for existing branch:
303
+ - Probe local + remote for existing branch. **Distinguish "no such ref" from "the
304
+ probe failed"**: with `2>/dev/null` and an empty-output test they look identical,
305
+ so an auth or network failure reads as "no collision" and the run creates a
306
+ branch that already exists on the remote - surfacing as a rejected push at
307
+ Phase 6, far from its cause.
304
308
  ```bash
305
309
  LOCAL_HIT=$(git -C "$root" rev-parse --verify --quiet "refs/heads/$branch")
306
- REMOTE_HIT=$(git -C "$root" ls-remote --exit-code --heads origin "$branch" 2>/dev/null)
310
+ REMOTE_ERR=$(git -C "$root" ls-remote --exit-code --heads origin "$branch" 2>&1 >/dev/null)
311
+ REMOTE_RC=$?
312
+ # 0 = ref exists (collision) · 2 = no matching ref (authoritative "free")
313
+ # anything else = the probe itself failed; REMOTE_ERR holds why
307
314
  ```
315
+ - `REMOTE_RC` is 0 or 2 → treat as authoritative
316
+ - `REMOTE_RC` is anything else → the remote answer is **unknown**, not "free". Log
317
+ `Remote collision probe failed: <REMOTE_ERR>`, fall back to the local check only,
318
+ and record `"remoteCollisionProbe": "failed"` in `agent-state.json` so Phase 6
319
+ expects a possible non-fast-forward and re-checks before pushing.
308
320
  - No collision → use as-is
309
321
  - Collision found → append `-v2`, `-v3`, etc. until unique:
310
322
  `bugfix/ABC-12345` exists → `bugfix/ABC-12345-v2`
@@ -218,11 +218,42 @@ Launch Agent instances **in parallel** using the shared `code-reviewer` subagent
218
218
 
219
219
  **Scope from Step 1.77.** `$REVIEW_SCOPE == "single"` → dispatch **Reviewer 1 only**, and skip Step 2.5 + 3.6 (both no-ops with one reviewer). `"full"` (default + fail-safe) → the whole set below. Either way record the count in `consensus.reviewerCount`.
220
220
 
221
- | Reviewer | subagent_type | Model | Focus | Skills Referenced | Where it runs |
222
- | ---------- | ----------------- | ------------------- | --------------------------------- | --------------------------------------------- | -------------------- |
223
- | Reviewer 1 | `code-reviewer` | `claude-fable-5` (Claude Code) / `claude-opus-4-8` (Copilot CLI) | Deep security + architecture | `api-security-best-practices`, `architecture` | Both CLIs |
224
- | Reviewer 2 | `code-reviewer` | `gpt-5.4` | Edge cases, different perspective | cross-model diversity | **Copilot CLI only** |
225
- | Reviewer 3 | `code-reviewer` | `claude-sonnet-4-6` | Quality + correctness + naming | `clean-code`, stack-specific skill | Both CLIs |
221
+ | Reviewer | subagent_type | Claude Code | Copilot CLI | Codex CLI | Focus | Skills Referenced |
222
+ | ---------- | --------------- | --- | --- | --- | --- | --- |
223
+ | Reviewer 1 | `code-reviewer` | `claude-fable-5` | `claude-opus-4-8` | `gpt-5.6` @ `xhigh` | Deep security + architecture | `api-security-best-practices`, `architecture` |
224
+ | Reviewer 2 | `code-reviewer` | (not dispatched) | `gpt-5.4` | `gpt-5.4` @ `high` | Edge cases, different perspective | cross-model diversity |
225
+ | Reviewer 3 | `code-reviewer` | `claude-sonnet-4-6` | `claude-sonnet-4-6` | `gpt-5.6` @ `medium` | Quality + correctness + naming | `clean-code`, stack-specific skill |
226
+ | Triage | triage persona | `claude-fable-5` | `claude-opus-4-8` | `gpt-5.6` @ `max` | Filter false positives + out-of-scope | - |
227
+
228
+ Reviewer count per host: **Claude Code 2, Copilot CLI 3, Codex CLI 3**.
229
+
230
+ #### Codex CLI - two constraints that fail silently
231
+
232
+ Both were measured against Codex 0.145, not inferred. Getting either wrong produces
233
+ a review that looks like it ran.
234
+
235
+ 1. **Every `spawn_agent` that sets `model` or `reasoning_effort` MUST pass
236
+ `fork_turns: "none"`** (or a positive integer). Codex states that a full-history
237
+ fork "inherits the parent model and reasoning effort and does not accept
238
+ overrides" - so omitting `fork_turns` silently collapses the whole panel onto
239
+ the orchestrator's model, with three reviewers reporting from one perspective
240
+ and no error anywhere.
241
+ 2. **Three concurrent children is the ceiling.** Codex advertises 4 concurrency
242
+ slots *including the orchestrator*, so a 3-reviewer panel saturates it and a 4th
243
+ child queues instead of running in parallel. The reviewer count on Codex is
244
+ capability-derived, not a preference.
245
+
246
+ Codex also refuses to spawn sub-agents at all unless instructions ask for it
247
+ explicitly; the pipeline's `~/.codex/AGENTS.md` managed block carries that
248
+ authorization. Without it Phase 4 degrades to a single in-thread review.
249
+
250
+ **Single-vendor caveat.** Every Codex reviewer is an OpenAI model, so the
251
+ cross-vendor disagreement that Claude Code and Copilot CLI get for free is absent.
252
+ The diversity budget shifts to reasoning effort and persona focus: Reviewer 1 runs
253
+ `xhigh` on security and architecture, Reviewer 2 runs a different model family
254
+ member on edge cases, Reviewer 3 runs `medium` on quality. Treat consensus among
255
+ them as weaker evidence than the same consensus on a two-vendor host, and say so
256
+ in the triage note when all three agree on a borderline finding.
226
257
 
227
258
  Each reviewer inherits the `code-reviewer` agent's focus areas (Security, Architecture, Quality, Performance) and output contract. The orchestrator overrides only the model and the stack-specific skill per-reviewer - no prompt duplication.
228
259
 
@@ -129,7 +129,7 @@ Every phase that dispatches a billable LLM agent MUST forward the call's token t
129
129
 
130
130
  ```bash
131
131
  LOG_METRIC_FORWARD_TO_TRACKER=1 pipeline/scripts/log-metric.sh "$TASK_ID" <phase-id> <event> \
132
- model=<fable|opus|sonnet|haiku|gpt-5.4> \
132
+ model=<fable|opus|sonnet|haiku|gpt-5.4|gpt-5.6|gpt-5.6-terra> \
133
133
  tokens_in=$IN tokens_out=$OUT duration_ms=$DUR
134
134
  ```
135
135
 
@@ -24,7 +24,8 @@ The agent detects which CLI it's running in and uses the appropriate visual mech
24
24
  ```
25
25
  1. system prompt mentions "Claude Code" → claude-code
26
26
  2. system prompt mentions "Copilot" / "GitHub Copilot" → copilot
27
- 3. None of the above → generic (bash stdout)
27
+ 3. system prompt mentions "Codex" → codex
28
+ 4. None of the above → generic (bash stdout)
28
29
  ```
29
30
 
30
31
  Visual mechanism per CLI:
@@ -33,12 +34,27 @@ Visual mechanism per CLI:
33
34
  |---|---|
34
35
  | **claude-code** | `TaskCreate({subject, activeForm})` then `TaskUpdate({status, activeForm})`. Native sticky widget; ⏺ tiles, spinner header. |
35
36
  | **copilot** | Inline call: `bash phase-tracker.sh render`. The bordered ANSI card lands as the last tool result in the chat. |
37
+ | **codex** | The native `update_plan` tool: one plan step per phase, `status: pending \| in_progress \| completed`. Rewrite the whole step list on each boundary - the tool takes the full plan, not a delta. |
36
38
  | **generic** (plain shell, Git Bash, WSL, tmux) | Same bash render - the bordered ANSI card prints to terminal stdout in place. |
37
39
 
38
40
  **Common to every CLI**: `phase-tracker.sh add/update/meta/tokens` calls run identically → the state file is always correct, and `:resume` / `:log` / `:status` work on every CLI.
39
41
 
40
42
  **Claude Code only**: in addition, `TaskCreate` / `TaskUpdate` native tool calls → the sticky widget pins the phase stack in the user's view.
41
43
 
44
+ ### Codex specifics
45
+
46
+ Two constraints on `update_plan`, both of which fail quietly if ignored:
47
+
48
+ - **Never call it in parallel with another tool.** Codex states this explicitly. A
49
+ phase boundary that batches `update_plan` alongside other calls loses the update.
50
+ - **It is unavailable in Codex plan mode.** Fall back to
51
+ `bash phase-tracker.sh render` there rather than skipping the visual channel.
52
+
53
+ The plan step is the phase label, not a restatement of the work: `Phase 4 - Review`
54
+ in the step list, with detail going to the agent log. A plan that mirrors the phase
55
+ list is legible; one that mirrors the task list duplicates what the log already
56
+ holds.
57
+
42
58
  ## Call pattern
43
59
 
44
60
  At every phase boundary the agent takes **two steps**: