@mmerterden/multi-agent-pipeline 12.11.0 → 13.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +179 -0
- package/README.md +24 -7
- package/index.js +5 -2
- package/install/_codex-agents.mjs +211 -0
- package/install/_codex-instructions.mjs +33 -0
- package/install/_managed-block.mjs +99 -0
- package/install/codex.mjs +478 -0
- package/install/copilot.mjs +34 -80
- package/install/index.mjs +25 -9
- package/install/templates/codex-instructions.md +45 -0
- package/package.json +5 -3
- package/pipeline/claude-md-template.md +1 -0
- package/pipeline/commands/multi-agent/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/dev/SKILL.md +52 -6
- package/pipeline/commands/multi-agent/dev-local/SKILL.md +21 -0
- package/pipeline/commands/multi-agent/finish/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/language/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/setup/SKILL.md +69 -2
- package/pipeline/commands/multi-agent/sync/SKILL.md +128 -5
- package/pipeline/commands/multi-agent/testflight-validation/SKILL.md +219 -0
- package/pipeline/commands/multi-agent/update/SKILL.md +7 -4
- package/pipeline/multi-agent-refs/_input-parser.md +1 -1
- package/pipeline/multi-agent-refs/component-dispatch.md +40 -7
- package/pipeline/multi-agent-refs/cross-cli-contract.md +51 -17
- package/pipeline/multi-agent-refs/features/model-fallback.md +29 -0
- package/pipeline/multi-agent-refs/features/review-multi-repo.md +1 -1
- package/pipeline/multi-agent-refs/phases/log-format.md +1 -1
- package/pipeline/multi-agent-refs/phases/phase-0-init.md +43 -2
- package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +24 -1
- package/pipeline/multi-agent-refs/phases/phase-3-dev.md +32 -0
- package/pipeline/multi-agent-refs/phases/phase-4-review.md +62 -5
- package/pipeline/multi-agent-refs/progress-contract.md +1 -1
- package/pipeline/multi-agent-refs/tracker-contract.md +17 -1
- package/pipeline/schemas/prefs.schema.json +296 -62
- package/pipeline/schemas/reviewer-output.schema.json +1 -1
- package/pipeline/schemas/triage-output.schema.json +1 -1
- package/pipeline/scripts/cost-table.json +15 -1
- package/pipeline/scripts/phase0-exit-gate.mjs +185 -0
- package/pipeline/scripts/smoke-cross-cli-behavior.sh +25 -10
- package/pipeline/scripts/uninstall.mjs +105 -9
- package/pipeline/scripts/update-check.sh +2 -1
- package/pipeline/skills/shared/core/multi-agent-dev/SKILL.md +21 -0
- package/pipeline/skills/shared/core/multi-agent-dev-local/SKILL.md +21 -0
- package/pipeline/skills/shared/core/multi-agent-language/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-setup/SKILL.md +48 -1
- package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +87 -6
- package/pipeline/skills/shared/core/multi-agent-testflight-validation/SKILL.md +120 -0
|
@@ -0,0 +1,219 @@
|
|
|
1
|
+
---
|
|
2
|
+
description: "Pre-submission validation for a TestFlight / App Store build (iOS, local-only). Three gates: static archive audit, Apple's own `altool --validate-app`, and a Review-Guidelines check. ITMS codes are mapped to the rule each implies. Validates only, never uploads. Use when a build is about to go to TestFlight, or a submission was rejected and you need why."
|
|
3
|
+
description-tr: "TestFlight / App Store yüklemesi öncesi doğrulama (iOS, yalnızca lokal). Repo + branch seç, sonra ya build'i sen ver ya da koşu archive alsın; üç kapıyı geç: statik 18-kurallı archive denetimi, Apple'ın kendi `altool --validate-app`'i, ve App Store Review Guidelines'a karşı guideline incelemesi. Her kapı için ayrı verdict, ITMS kodları ilgili kurala eşlenmiş. Sadece doğrular - asla yüklemez."
|
|
4
|
+
argument-hint: "[repo] - empty = pick from prefs; repo name or path; --ipa=<path>; --archive=<path>; --resume"
|
|
5
|
+
allowed-tools: Agent, Bash, Read, Write, Edit, Glob, Grep, TaskCreate, TaskUpdate, AskUserQuestion, Skill, mcp__dev-toolkit__ios_app_store_audit, mcp__dev-toolkit__ios_export_ipa, mcp__dev-toolkit__ios_testflight_validate, mcp__dev-toolkit__ios_xcodebuild, mcp__dev-toolkit__ios_xcresult
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
# /multi-agent:testflight-validation - pre-submission validation
|
|
9
|
+
|
|
10
|
+
Catch, before you upload, what App Store Connect would send back after you do.
|
|
11
|
+
|
|
12
|
+
**Local-only.** No commits, no push, no PR, no channels. The worktree exists only
|
|
13
|
+
to archive without touching your working tree.
|
|
14
|
+
|
|
15
|
+
**It never uploads.** Only `--validate-app` is ever invoked, never `--upload-app`.
|
|
16
|
+
A validation run must not be able to ship a build by accident.
|
|
17
|
+
|
|
18
|
+
## Why three gates and not one
|
|
19
|
+
|
|
20
|
+
Each gate sees something the others structurally cannot. Reporting one of them as
|
|
21
|
+
"the check" is how a build passes locally and gets rejected anyway.
|
|
22
|
+
|
|
23
|
+
| Gate | What runs | Needs | Sees | Blind to |
|
|
24
|
+
|---|---|---|---|---|
|
|
25
|
+
| **1. Static** | `ios_app_store_audit` (18 rules, real ITMS codes) | an `.xcarchive` | privacy manifest, required-reason API, Info.plist, code signing, entitlements, embedded SDK, IPv6, debug-tool leak, binary size | anything that depends on the App Store Connect account |
|
|
26
|
+
| **2. Authoritative** | `ios_testflight_validate` → `altool --validate-app` | an `.ipa` + credentials | unregistered bundle ID, profile that does not match the app record, **a version+build pair already used**, entitlements not provisioned for the App ID | the Review Guidelines - Apple's validator does not read them |
|
|
27
|
+
| **3. Guideline** | `app-store-review` skill + repo evidence | repo checkout | ATT flow, privacy policy, account deletion, IAP rules, purpose-string wording, permission justification | anything not visible in source |
|
|
28
|
+
|
|
29
|
+
Gate 2 is the only one that asks Apple, and Gate 3 is the only one that covers the
|
|
30
|
+
rejections a human reviewer writes. Most "we passed validation and still got
|
|
31
|
+
rejected" cases are Gate 3 findings.
|
|
32
|
+
|
|
33
|
+
## Step 0 - parse input
|
|
34
|
+
|
|
35
|
+
| Input | Meaning |
|
|
36
|
+
|---|---|
|
|
37
|
+
| (empty) | ask which repo (Step 1) |
|
|
38
|
+
| `my-ios-app` or a path | that repo |
|
|
39
|
+
| `--ipa=<path>` | Mode B with an `.ipa`; Gate 1 cannot run (see Step 3) |
|
|
40
|
+
| `--archive=<path>` | Mode B with an `.xcarchive`; all three gates run |
|
|
41
|
+
| `--resume` | continue the last run from its state file |
|
|
42
|
+
|
|
43
|
+
State lives at `$HOME/.claude/logs/multi-agent/<task_id>/agent-state.json` with
|
|
44
|
+
`taskId = TFV-<repo>-<yyyymmddHHMM>`. Register phases with the tracker
|
|
45
|
+
(`$HOME/.claude/multi-agent-refs/tracker-contract.md`) so `:resume` and `:status`
|
|
46
|
+
work like any other run.
|
|
47
|
+
|
|
48
|
+
## Step 1 - pickers (native, always)
|
|
49
|
+
|
|
50
|
+
Use `AskUserQuestion` for every step - never a numbered text menu. Questions and
|
|
51
|
+
descriptions render in `prefs.global.outputLanguage`; `label` and `header` stay
|
|
52
|
+
English, per `$HOME/.claude/multi-agent-refs/picker-contract.md`. Print the
|
|
53
|
+
`Step <i>/<n>: <what this decides>` breadcrumb for each.
|
|
54
|
+
|
|
55
|
+
1. **Repo** - from `prefs.projects` where the stack is iOS. A single match
|
|
56
|
+
auto-resolves (say so in the breadcrumb, do not silently skip the step).
|
|
57
|
+
2. **Branch** - the branch to validate. Resolution order:
|
|
58
|
+
- `git fetch --prune` first, capturing **stderr**. **If the fetch fails, do not
|
|
59
|
+
silently fall back to a cached ref**, and **classify before naming a cause** -
|
|
60
|
+
the same rule as the `/multi-agent:dev` remote gate:
|
|
61
|
+
|
|
62
|
+
| stderr contains | Cause | Remedy |
|
|
63
|
+
|---|---|---|
|
|
64
|
+
| `could not read Password`, `Authentication failed`, `403` | credential | store the PAT in the credential helper or switch the remote to SSH. **A VPN cannot fix this**, and the base ref being stale is unrelated to what broke - do not offer the cached-ref fallback. |
|
|
65
|
+
| `Could not resolve host`, `Operation timed out`, `Connection refused` | network | retry / continue on the cached ref with an explicit warning / switch remote / abort, per `$HOME/.claude/multi-agent-refs/rules.md` |
|
|
66
|
+
| `Repository not found`, `404` | wrong remote | show `git remote -v` and ask |
|
|
67
|
+
|
|
68
|
+
Always print the observed stderr line next to the classification. Asserting
|
|
69
|
+
`unreachable (VPN/DNS)` for a missing-credential error that returns in under a
|
|
70
|
+
second sends the user to fix something that was never broken.
|
|
71
|
+
- Offer the current branch, the default branch, and any `release/*` /
|
|
72
|
+
`tkdevelop/*` heads.
|
|
73
|
+
3. **Mode** - how the build is obtained:
|
|
74
|
+
- `Supply a build` (Mode B, default) - fastest, no signing needed in-run.
|
|
75
|
+
- `Archive from this branch` (Mode A) - needs a distribution certificate and
|
|
76
|
+
profile in the keychain, and takes as long as a release archive.
|
|
77
|
+
|
|
78
|
+
## Step 2 - pre-flight, before anything expensive
|
|
79
|
+
|
|
80
|
+
Report every line; a missing prerequisite is a halt, not a warning.
|
|
81
|
+
|
|
82
|
+
```bash
|
|
83
|
+
xcrun --find altool >/dev/null 2>&1 || echo "MISSING: altool (install Xcode)"
|
|
84
|
+
xcodebuild -version | head -1
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
Then resolve credentials, and **state which tier is active in the report**:
|
|
88
|
+
|
|
89
|
+
| Tier | Source | Effect |
|
|
90
|
+
|---|---|---|
|
|
91
|
+
| 1 | ASC API key - key id + issuer id from the keychain via `prefs.global.keychainMapping`, `.p8` at `~/.appstoreconnect/private_keys/AuthKey_<keyId>.p8` | Gate 2 runs |
|
|
92
|
+
| 2 | Apple ID + app-specific password, referenced as a keychain item | Gate 2 runs |
|
|
93
|
+
| 3 | neither | **Gate 2 reports `SKIPPED`, and the run says so in the verdict line** |
|
|
94
|
+
|
|
95
|
+
Credentials come from `/multi-agent:setup`; never prompt for a secret value in
|
|
96
|
+
chat. If nothing is configured, tell the user which of the two tiers they can set
|
|
97
|
+
up and that tier 2 needs no elevated App Store Connect role.
|
|
98
|
+
|
|
99
|
+
Multi-provider accounts need `--provider-public-id`. When it is not in prefs, run
|
|
100
|
+
`ios_testflight_validate({list_providers: true})` once and ask which provider.
|
|
101
|
+
|
|
102
|
+
## Step 3 - obtain the build
|
|
103
|
+
|
|
104
|
+
### Mode B - a build you supply
|
|
105
|
+
|
|
106
|
+
- `.xcarchive` → Gate 1 runs on it. To reach Gate 2 the archive must be exported,
|
|
107
|
+
so run `ios_export_ipa` (see Mode A step 3 for the signing inputs).
|
|
108
|
+
- `.ipa` only → **Gate 1 is reported `SKIPPED (needs .xcarchive)`.** The static
|
|
109
|
+
audit reads archive structure that an `.ipa` does not carry. Do not present a
|
|
110
|
+
two-gate run as a full pass; say which gate did not run and why, and offer to
|
|
111
|
+
re-run with the archive.
|
|
112
|
+
|
|
113
|
+
### Mode A - archive from the branch
|
|
114
|
+
|
|
115
|
+
1. Worktree at `{projectRoot}/{worktreeBasePath}/{taskId}` on the chosen branch.
|
|
116
|
+
**Never under `$HOME`**, never a direct checkout of the main working tree.
|
|
117
|
+
2. Resolve the scheme and workspace/project from prefs; ask if ambiguous.
|
|
118
|
+
3. Archive:
|
|
119
|
+
`ios_xcodebuild({workspace|project, scheme, action: "archive", configuration: "Release", destination: "generic/platform=iOS"})`
|
|
120
|
+
Note the destination: the simulator default would produce an archive that
|
|
121
|
+
cannot be exported for distribution.
|
|
122
|
+
4. Export:
|
|
123
|
+
`ios_export_ipa({archive_path, output_dir, method: "app-store-connect", team_id, provisioning_profiles?, signing_style?})`
|
|
124
|
+
Leave `allow_provisioning_updates` off unless the user asks: it lets xcodebuild
|
|
125
|
+
create or modify profiles in the developer account, which a validation run has
|
|
126
|
+
no business doing.
|
|
127
|
+
5. A failed export halts with the parsed errors. The usual causes are a missing
|
|
128
|
+
distribution certificate, a profile that does not match the bundle ID, or
|
|
129
|
+
`signing_style: "manual"` with no `provisioning_profiles` map.
|
|
130
|
+
|
|
131
|
+
## Step 4 - Gate 1, static audit
|
|
132
|
+
|
|
133
|
+
`ios_app_store_audit({archive_path, rules: "all"})`.
|
|
134
|
+
|
|
135
|
+
`error` findings are blocking; `warning` is advisory. Group the output by severity
|
|
136
|
+
and keep each finding's ITMS code - Gate 2 may return the same code, and seeing
|
|
137
|
+
it in both places tells the user it is real rather than a heuristic.
|
|
138
|
+
|
|
139
|
+
## Step 5 - Gate 2, Apple's own validation
|
|
140
|
+
|
|
141
|
+
`ios_testflight_validate({ipa_path, platform: "ios", <credential args>})`.
|
|
142
|
+
|
|
143
|
+
Render the verdict exactly as returned:
|
|
144
|
+
|
|
145
|
+
- `PASS` - Apple accepted the binary for delivery.
|
|
146
|
+
- `FAIL` - list each issue with its ITMS code, the mapped guideline, and the hint.
|
|
147
|
+
- `SKIPPED` - print the reason. **Never render this as a pass.** The verdict line
|
|
148
|
+
for the whole run must read `2 of 3 gates cleared, 1 skipped`, not `passed`.
|
|
149
|
+
|
|
150
|
+
Gate 2 is the only gate that catches a build number already used - the most
|
|
151
|
+
common wasted upload - so when it fails on that, say so plainly and name the next
|
|
152
|
+
free build number.
|
|
153
|
+
|
|
154
|
+
## Step 6 - Gate 3, guideline review
|
|
155
|
+
|
|
156
|
+
Load the `app-store-review` skill and review the repo against it. This is the gate
|
|
157
|
+
that catches what a human reviewer rejects, so it reads source, not the binary:
|
|
158
|
+
|
|
159
|
+
| Area | Evidence to gather |
|
|
160
|
+
|---|---|
|
|
161
|
+
| Purpose strings | every `NS*UsageDescription` in Info.plist - present, specific, user-facing, and matching what the code actually does with the data |
|
|
162
|
+
| Privacy manifest | `PrivacyInfo.xcprivacy` exists, declares required-reason APIs, and matches the SDKs actually linked |
|
|
163
|
+
| Tracking | if any tracking API or SDK is present, an ATT prompt exists and runs before collection |
|
|
164
|
+
| Account deletion | if the app creates accounts, an in-app deletion path exists (guideline 5.1.1(v)) |
|
|
165
|
+
| Privacy policy | reachable in-app and in the metadata |
|
|
166
|
+
| IAP | anything unlocking features goes through StoreKit, with no external purchase path |
|
|
167
|
+
| Sign in with Apple | present when a third-party social login is offered |
|
|
168
|
+
|
|
169
|
+
For each: `pass` / `fail` / `not-applicable` with the evidence path that justifies
|
|
170
|
+
it. `not-applicable` needs a reason - an unexamined area is not a pass.
|
|
171
|
+
|
|
172
|
+
## Step 7 - report
|
|
173
|
+
|
|
174
|
+
Write to `~/TestFlightChecks/<repo>-<branch>-<timestamp>/report.md` and print a
|
|
175
|
+
summary. Structure:
|
|
176
|
+
|
|
177
|
+
```
|
|
178
|
+
Verdict: <N> of 3 gates cleared[, <M> skipped]
|
|
179
|
+
Build: <ipa or archive path> · <bundle id> <version> (<build>)
|
|
180
|
+
Auth: tier <1|2|none> · <method>
|
|
181
|
+
|
|
182
|
+
Gate 1 static audit PASS | FAIL (<n> blocking, <n> advisory) | SKIPPED (<reason>)
|
|
183
|
+
Gate 2 Apple validation PASS | FAIL (<n> issues) | SKIPPED (<reason>)
|
|
184
|
+
Gate 3 guideline review PASS | FAIL (<n> findings) | <n> not-applicable
|
|
185
|
+
|
|
186
|
+
Blocking - fix before uploading
|
|
187
|
+
[ITMS-90683] Info.plist: NSCameraUsageDescription missing
|
|
188
|
+
guideline 5.1.1 Data Collection and Storage
|
|
189
|
+
<hint>
|
|
190
|
+
<file:line>
|
|
191
|
+
|
|
192
|
+
Advisory
|
|
193
|
+
...
|
|
194
|
+
|
|
195
|
+
Not run
|
|
196
|
+
Gate 1: needs an .xcarchive; only an .ipa was supplied
|
|
197
|
+
```
|
|
198
|
+
|
|
199
|
+
Rules for the report:
|
|
200
|
+
|
|
201
|
+
- A skipped gate is never folded into the pass count. The verdict line states the
|
|
202
|
+
skip.
|
|
203
|
+
- Every blocking finding carries a file path or an ITMS code. A finding the user
|
|
204
|
+
cannot act on is noise.
|
|
205
|
+
- No AI or assistant attribution anywhere, per
|
|
206
|
+
`$HOME/.claude/rules/git-conventions.md`.
|
|
207
|
+
- Real newlines, no HTML entities, per
|
|
208
|
+
`$HOME/.claude/rules/pipeline-output-formatting.md`.
|
|
209
|
+
|
|
210
|
+
## Step 8 - offer the next action, do not take it
|
|
211
|
+
|
|
212
|
+
Print, and stop:
|
|
213
|
+
|
|
214
|
+
- the exact `xcrun altool --upload-app` command for when the gates are clear, so
|
|
215
|
+
uploading stays an explicit human act
|
|
216
|
+
- `/multi-agent:testflight-validation --resume` to re-run after fixes
|
|
217
|
+
- `/multi-agent:fix-bug` when Gate 3 produced code-level findings
|
|
218
|
+
|
|
219
|
+
Never upload, never bump the build number, never commit.
|
|
@@ -52,15 +52,18 @@ Update the pipeline in one command. Existing preferences are preserved; only ski
|
|
|
52
52
|
fi
|
|
53
53
|
```
|
|
54
54
|
|
|
55
|
-
4c. **Prune retired adapter files** (Cursor / Antigravity /
|
|
55
|
+
4c. **Prune retired adapter files** (Cursor / Antigravity / Copilot Chat were removed in v10.7.0 - the pipeline targets Claude Code, Copilot CLI and Codex CLI):
|
|
56
56
|
```bash
|
|
57
|
-
# Global Codex adapter prompt is safe to remove (no longer produced).
|
|
58
|
-
rm -f "$HOME/.codex/prompts/multi-agent.md"
|
|
59
|
-
echo " -> pruned retired global Codex adapter prompt (if present)"
|
|
60
57
|
# Per-project adapter files (.cursor/, .agent/, .github/copilot-instructions.md)
|
|
61
58
|
# live in your repos and are left untouched - remove them manually if you like.
|
|
59
|
+
echo " -> no global adapter files to prune"
|
|
62
60
|
```
|
|
63
61
|
|
|
62
|
+
> **Do NOT delete `$HOME/.codex/prompts/multi-agent.md`.** Releases up to
|
|
63
|
+
> v12.11.0 removed it here as a retired v9.7.0 adapter leftover. Codex CLI is a
|
|
64
|
+
> supported target again as of v13.0.0 and the installer writes that file, so
|
|
65
|
+
> deleting it silently breaks the `/multi-agent` slash command on Codex.
|
|
66
|
+
|
|
64
67
|
5. **Migrate preferences** (if there is an old schema):
|
|
65
68
|
```bash
|
|
66
69
|
if [ -f "$HOME/.claude/scripts/migrate-prefs.mjs" ]; then
|
|
@@ -6,7 +6,7 @@ description: "Internal - input type detection for multi-agent dispatcher."
|
|
|
6
6
|
|
|
7
7
|
The top-level `multi-agent` command classifies user arguments per the rules below.
|
|
8
8
|
|
|
9
|
-
> **Language**: Schema reference only - no user-facing text. Picker
|
|
9
|
+
> **Language**: Schema reference only - no user-facing text. Picker `label` + `header` always English (`promptLanguage` is locked to `"en"`); `question` + `description` follow `outputLanguage` per the `rules.md` matrix. Commit/PR/Jira payloads stay English.
|
|
10
10
|
|
|
11
11
|
## Type Table
|
|
12
12
|
|
|
@@ -11,20 +11,53 @@ Phase 3 checks, in order:
|
|
|
11
11
|
1. `agent-state.json` has `taskType: "component"` (set by Phase 0 Step 7).
|
|
12
12
|
2. `agent-state.json` has a non-null `figmaUrl`.
|
|
13
13
|
|
|
14
|
-
Either missing →
|
|
14
|
+
Either missing → **HALT.** Log the anomaly to `agent-log.md` ("component dispatch
|
|
15
|
+
expected but state incomplete: <which field>") and stop with a user-visible error
|
|
16
|
+
naming the missing field and pointing at `phase0-exit-gate.mjs`.
|
|
17
|
+
|
|
18
|
+
> Earlier wording sent an incomplete-state component task down the generic TDD path,
|
|
19
|
+
> which contradicted the sentence that followed it: taking the generic path **is**
|
|
20
|
+
> skipping the Figma work.
|
|
21
|
+
> It also authorised the exact degradation that broke a real run - Phase 0 never
|
|
22
|
+
> wrote `agent-state.json`, so `taskType` was absent, so a Figma-driven screen was
|
|
23
|
+
> built through the generic path with no token-compliance check, no Code Connect
|
|
24
|
+
> publish and no component review. Spacing came out `16` where the frame said
|
|
25
|
+
> `Spacing/12`, and half the branch's commits were rework.
|
|
26
|
+
>
|
|
27
|
+
> The Phase 0 exit gate now prevents reaching Phase 3 in that state at all; this
|
|
28
|
+
> halt is the second line of defence. A component task that cannot be dispatched as
|
|
29
|
+
> one must fail loudly, because the generic path produces artefacts that look
|
|
30
|
+
> finished and are not.
|
|
15
31
|
|
|
16
32
|
## Plugin skill resolution
|
|
17
33
|
|
|
34
|
+
Scope first, then platform. A screen and a single component are different jobs and
|
|
35
|
+
the plugin ships a skill for each; routing a screen to the component skill is why one
|
|
36
|
+
run produced entities and a mapper but left the screen half-wired.
|
|
37
|
+
|
|
38
|
+
| `state.componentScope` | Meaning | iOS skill | Android skill |
|
|
39
|
+
|---|---|---|---|
|
|
40
|
+
| `screen` (default when the frame is a full screen, or the task names a screen) | Full clean-architecture vertical: Entity → Repository → Mapper → UseCase → LocalizedText → AnalyticsTracking → CoordinatorEvent → ViewModel → Scene → Preview, then verify | `ai-ios-engineering-toolkit:create-screen` | `ai-android-engineering-toolkit:create-screen` |
|
|
41
|
+
| `component` | One reusable UI component (Configuration / View / +Modifiers / Code Connect) | `ai-ios-engineering-toolkit:create-component` (fallback `create-ui-component`) | `ai-android-engineering-toolkit:create-component` (fallback `create-ui-component`) |
|
|
42
|
+
| `evolve` | Change an existing component | `evolve-component` (fallback `evolve-ui-component`) | same |
|
|
43
|
+
|
|
18
44
|
```
|
|
19
|
-
project.platform → component skill (enabled marketplace plugin, Claude Code)
|
|
20
|
-
──────────────────────────────────────────────────────────────────────────────
|
|
21
|
-
ios → Skill: ai-ios-engineering-toolkit:create-component
|
|
22
|
-
(fallback: ai-ios-engineering-toolkit:create-ui-component)
|
|
23
|
-
android → Skill: ai-android-engineering-toolkit:create-component
|
|
24
|
-
(fallback: ai-android-engineering-toolkit:create-ui-component)
|
|
25
45
|
web, multi-* → HALT with clear error (no web target)
|
|
26
46
|
```
|
|
27
47
|
|
|
48
|
+
Phase 0 Step 7 sets `state.componentScope` alongside `taskType`: a Figma frame that
|
|
49
|
+
is a full screen, or a task whose title names a screen, is `screen`; a frame that is
|
|
50
|
+
a single atom is `component`. When it cannot be decided, ask - do not default to
|
|
51
|
+
`component`, because the screen path is a superset and the component path silently
|
|
52
|
+
omits the wiring.
|
|
53
|
+
|
|
54
|
+
**Pre-implementation validation is not optional on iOS.** Before the create skill
|
|
55
|
+
runs, dispatch `ai-ios-engineering-toolkit:figma-validate` for the frame. It checks
|
|
56
|
+
registry presence, Code Connect strategy, **design token compliance**, dependency
|
|
57
|
+
readiness, atomic scope and already-implemented status in about ten seconds. Those
|
|
58
|
+
are precisely the checks whose absence produced guessed spacing and an unpublished
|
|
59
|
+
Code Connect binding. A `figma-validate` failure halts the dispatch.
|
|
60
|
+
|
|
28
61
|
**Dual-name resolution.** The public (`multi-agent-plugins`) and a corporate/private marketplace named the same skill differently - `create-component` vs `create-ui-component`. Dispatch tries `create-component` first; if it is not available in the current repo, tries `create-ui-component`. (Same dual-name rule applies when `taskType` maps to evolve → `evolve-component`/`evolve-ui-component`, or fix → `fix-bug`.)
|
|
29
62
|
|
|
30
63
|
If **neither** resolves, the platform's `ai-<platform>-engineering-toolkit` plugin is not enabled in this repo. **Halt with a user-visible error**: "component task requires the ai-<platform>-engineering-toolkit plugin enabled in this repo (`.claude/settings.local.json`)." Do not silently fall back to TDD - a component task ran through the bugfix path would produce wrong artefacts.
|
|
@@ -1,19 +1,19 @@
|
|
|
1
|
-
# Cross-CLI Contract (Claude Code
|
|
1
|
+
# Cross-CLI Contract (Claude Code · Copilot CLI · Codex CLI)
|
|
2
2
|
|
|
3
3
|
> **Non-negotiable**. Any change that breaks this contract blocks merge. Validated by `smoke-cross-cli-behavior.sh`.
|
|
4
4
|
|
|
5
|
-
**Purpose**: every pipeline command must produce identical artifacts (state, logs, outputs) and respect identical placeholder vocabulary regardless of which
|
|
5
|
+
**Purpose**: every pipeline command must produce identical artifacts (state, logs, outputs) and respect identical placeholder vocabulary regardless of which of the three host CLIs invokes it. This file is the source of truth for "what must stay the same."
|
|
6
6
|
|
|
7
7
|
---
|
|
8
8
|
|
|
9
|
-
## 1. Command Inventory (
|
|
9
|
+
## 1. Command Inventory (43 commands)
|
|
10
10
|
|
|
11
11
|
```
|
|
12
12
|
analysis, analysis-resolve, autopilot, build-optimize, channels, create-jira, design-check, dev,
|
|
13
13
|
dev-autopilot, dev-local, dev-local-autopilot, diff-explain, finish, forget, garbage-collect,
|
|
14
14
|
help, issue, jira, kill, language, local,
|
|
15
15
|
local-autopilot, log, manual-test, prune-logs, purge, refactor, resume, review, review-issue, review-jira,
|
|
16
|
-
routines, save, scan, search, setup, stack, status, sync, test, uninstall, update
|
|
16
|
+
routines, save, scan, search, setup, stack, status, sync, test, testflight-validation, uninstall, update
|
|
17
17
|
```
|
|
18
18
|
|
|
19
19
|
Categories:
|
|
@@ -24,6 +24,7 @@ Categories:
|
|
|
24
24
|
- **Fast modes** (Init -> Dev(Opus) -> Commit -> Report): `dev`, `dev-autopilot`, `dev-local`, `dev-local-autopilot`
|
|
25
25
|
- **Tail modes** (run the pipeline tail over already-done local work): `finish`
|
|
26
26
|
- **Ops commands** (one-shot, no worktree): `status`, `log`, `kill`, `purge`, `uninstall`, `resume`, `review`, `review-jira`, `review-issue`, `analysis`, `analysis-resolve`, `build-optimize`, `channels`, `scan`, `search`, `diff-explain`, `garbage-collect`, `prune-logs`
|
|
27
|
+
- **Local audits** (worktree only to build; no commit, push, PR or channels): `design-check`, `testflight-validation`. `testflight-validation` additionally never invokes `altool --upload-app` - a validation run must not be able to ship a build by accident.
|
|
27
28
|
- **Meta-ops**: `setup`, `sync`, `update`, `help`, `refactor`, `test`, `stack`, `manual-test`, `language`
|
|
28
29
|
- **Routines** (user-defined routine registry; the routines they create are local-only and never synced): `save`, `routines`, `forget`
|
|
29
30
|
|
|
@@ -110,14 +111,43 @@ Legacy names found during the v3.7 audit - replaced per this contract.
|
|
|
110
111
|
|
|
111
112
|
---
|
|
112
113
|
|
|
113
|
-
## 2.6 Intentional structural divergence -
|
|
114
|
+
## 2.6 Intentional structural divergence - thin dispatcher vs inlined orchestrator
|
|
114
115
|
|
|
115
|
-
The top-level orchestrator has two deliberately different shapes
|
|
116
|
+
The top-level orchestrator has two deliberately different shapes. This is **not
|
|
117
|
+
drift** - it reflects a real capability gap between the host CLIs, and auditors
|
|
118
|
+
must not flag it as a parity violation.
|
|
116
119
|
|
|
117
120
|
| Surface | File | Shape | Why |
|
|
118
121
|
|---|---|---|---|
|
|
119
|
-
| Claude Code (colon-form) | `pipeline/commands/multi-agent.md` | ~300 lines - thin dispatcher that routes to `$HOME/.claude/multi-agent-refs/phases/phase-N-*.md` on demand | Claude Code
|
|
120
|
-
| Copilot CLI (dash-form) | `pipeline/skills/shared/core/multi-agent/SKILL.md` | ~830 lines - full inline orchestrator with all 8 phase specs embedded | Copilot CLI loads the whole SKILL.md
|
|
122
|
+
| Claude Code (colon-form) | `pipeline/commands/multi-agent/SKILL.md` | ~300 lines - thin dispatcher that routes to `$HOME/.claude/multi-agent-refs/phases/phase-N-*.md` on demand | Claude Code lazy-loads reference files, so only the active phase's docs enter the context window. Saves tokens. |
|
|
123
|
+
| Copilot CLI (dash-form) | `pipeline/skills/shared/core/multi-agent/SKILL.md` | ~830 lines - full inline orchestrator with all 8 phase specs embedded | Copilot CLI loads the whole SKILL.md once the skill is dispatched; it has no equivalent of Claude's ref-file loading. Embedding keeps behavior identical without relying on a feature Copilot lacks. |
|
|
124
|
+
| Codex CLI | installed as `~/.codex/skills/multi-agent/SKILL.md`, generated from the Claude dispatcher | thin dispatcher, refs under `~/.codex/multi-agent-refs/` | Codex reads files on demand, so it takes the Claude shape - **and it has to.** See below. |
|
|
125
|
+
|
|
126
|
+
### Why Codex takes the thin-dispatcher shape, and must keep it
|
|
127
|
+
|
|
128
|
+
Codex assembles every discovered skill's name + description into a single prompt
|
|
129
|
+
block and **silently drops entries when that block overflows**. Measured against
|
|
130
|
+
Codex 0.145 during the install design: installing one plugin that declares 142
|
|
131
|
+
skills took the block from 11 skills / 4,710 bytes to 83 skills / 22,111 bytes -
|
|
132
|
+
only **75 of the 142** surfaced, **and an unrelated user-scope skill was evicted**.
|
|
133
|
+
Removing the plugin brought it back.
|
|
134
|
+
|
|
135
|
+
So shipping the 43 sub-commands as peer skills on Codex would silently lose
|
|
136
|
+
pipeline commands next to any stack toolkit, with no error anywhere. The pipeline
|
|
137
|
+
therefore contributes **exactly one** skill on Codex (`multi-agent`) and keeps the
|
|
138
|
+
43 sub-command specs as reference files that cost nothing until read.
|
|
139
|
+
|
|
140
|
+
**Do not "fix" this by adding per-command skills on Codex.** The layout is
|
|
141
|
+
capability-derived, and `smoke-install-layout.sh` fails if the Codex skills tree
|
|
142
|
+
gains a second pipeline entry.
|
|
143
|
+
|
|
144
|
+
### Parity axis differs per host
|
|
145
|
+
|
|
146
|
+
Claude Code and Copilot CLI are compared on their **skill directory sets**. Codex is
|
|
147
|
+
compared on its **ref set**: the 43 command specs must all exist under
|
|
148
|
+
`~/.codex/multi-agent-refs/commands/<cmd>/SKILL.md`, and
|
|
149
|
+
`smoke-codex-install.sh` asserts the count against the source tree. Comparing Codex
|
|
150
|
+
on skill directories would demand exactly the layout that breaks it.
|
|
121
151
|
|
|
122
152
|
**What must stay identical** (byte-level) across the two files:
|
|
123
153
|
|
|
@@ -175,13 +205,17 @@ argument-hint: "<input hint>"
|
|
|
175
205
|
|
|
176
206
|
### 4.1 TaskCreate ↔ phase-tracker.sh
|
|
177
207
|
|
|
178
|
-
| Concept | Claude Code | Copilot CLI |
|
|
179
|
-
|
|
180
|
-
| Register a phase | `TaskCreate` tool call with subject/description | `phase-tracker.sh add <N> <name>` |
|
|
181
|
-
| Mark a phase in-progress | `TaskUpdate` → `in_progress` | `phase-tracker.sh update <N> in_progress` |
|
|
182
|
-
| Mark a phase complete | `TaskUpdate` → `completed` | `phase-tracker.sh update <N> completed` |
|
|
183
|
-
| Mark a phase failed | `TaskUpdate` → `completed` + log failure in agent-log | `phase-tracker.sh update <N> failed` |
|
|
184
|
-
| Sub-phase progress | TaskCreate with `addBlockedBy` | `phase-tracker.sh sub <N> <sub> <name> <status>` |
|
|
208
|
+
| Concept | Claude Code | Copilot CLI | Codex CLI |
|
|
209
|
+
|---|---|---|---|
|
|
210
|
+
| Register a phase | `TaskCreate` tool call with subject/description | `phase-tracker.sh add <N> <name>` | `update_plan` step, `status: pending` |
|
|
211
|
+
| Mark a phase in-progress | `TaskUpdate` → `in_progress` | `phase-tracker.sh update <N> in_progress` | `update_plan` step → `in_progress` |
|
|
212
|
+
| Mark a phase complete | `TaskUpdate` → `completed` | `phase-tracker.sh update <N> completed` | `update_plan` step → `completed` |
|
|
213
|
+
| Mark a phase failed | `TaskUpdate` → `completed` + log failure in agent-log | `phase-tracker.sh update <N> failed` | `update_plan` step → `completed` + failure in agent-log |
|
|
214
|
+
| Sub-phase progress | TaskCreate with `addBlockedBy` | `phase-tracker.sh sub <N> <sub> <name> <status>` | `phase-tracker.sh sub ...` (the plan tool has no nesting) |
|
|
215
|
+
|
|
216
|
+
`update_plan` takes the FULL step list, not a delta, so each boundary rewrites the
|
|
217
|
+
whole plan. It must never be called in parallel with another tool, and it is
|
|
218
|
+
unavailable in Codex plan mode - fall back to `phase-tracker.sh render` there.
|
|
185
219
|
|
|
186
220
|
### 4.2 Contract
|
|
187
221
|
|
|
@@ -227,9 +261,9 @@ Modifier flags are orthogonal and compose:
|
|
|
227
261
|
|
|
228
262
|
## 7. Platform Guards (macOS / Linux / WSL)
|
|
229
263
|
|
|
230
|
-
Commands invoked via Copilot CLI may run on Linux or WSL. Claude Code currently macOS-only. Shell code in `pipeline/skills/shared/core/` command files and in `pipeline/lib` / `pipeline/scripts` must be portable.
|
|
264
|
+
Commands invoked via Copilot CLI or Codex CLI may run on Linux or WSL. Claude Code currently macOS-only. Shell code in `pipeline/skills/shared/core/` command files and in `pipeline/lib` / `pipeline/scripts` must be portable.
|
|
231
265
|
|
|
232
|
-
For Keychain I/O the canonical path is **`~/.claude/lib/credential-store.sh`** (or `~/.copilot/lib/credential-store.sh` on
|
|
266
|
+
For Keychain I/O the canonical path is **`~/.claude/lib/credential-store.sh`** (or `~/.copilot/lib/credential-store.sh` / `~/.codex/lib/credential-store.sh` on those installs). The shell driver detects platform internally and auto-delegates to `keychain.py` (Python helper, macOS / Linux) or PowerShell `CredentialManager` (Windows). Call sites stay platform-agnostic - no per-OS branching needed.
|
|
233
267
|
|
|
234
268
|
| Purpose | Canonical (cross-platform) | Underlying backend (for reference / debugging) |
|
|
235
269
|
|---|---|---|
|
|
@@ -101,3 +101,32 @@ per-phase `model` field already carries the override).
|
|
|
101
101
|
reviewer models (GPT-5.4 + Opus + Sonnet - Fable 5 is not offered there) and
|
|
102
102
|
does not use this persona ladder. Only Claude Code dispatches Reviewer-1 on
|
|
103
103
|
Fable.
|
|
104
|
+
|
|
105
|
+
## Codex CLI
|
|
106
|
+
|
|
107
|
+
Codex offers no Anthropic models, so the tier names map onto OpenAI models plus a
|
|
108
|
+
reasoning effort - effort carries the depth distinction that the model id carries
|
|
109
|
+
on Claude Code. The map is applied at install time by
|
|
110
|
+
`install/_codex-agents.mjs`, which writes `model` + `model_reasoning_effort` into
|
|
111
|
+
each `~/.codex/agents/<persona>.toml`:
|
|
112
|
+
|
|
113
|
+
| Tier | Codex model | Effort |
|
|
114
|
+
|---|---|---|
|
|
115
|
+
| `fable` | `gpt-5.6` | `xhigh` |
|
|
116
|
+
| `opus` | `gpt-5.6` | `high` |
|
|
117
|
+
| `sonnet` | `gpt-5.4` | `medium` |
|
|
118
|
+
| `haiku` | `gpt-5.6-terra` | `low` |
|
|
119
|
+
|
|
120
|
+
Ladder on Codex: `gpt-5.6 @ xhigh -> gpt-5.6 @ high -> gpt-5.4 -> gpt-5.6-terra`.
|
|
121
|
+
The first step down lowers effort rather than switching model, which is the
|
|
122
|
+
cheapest useful degradation when the top tier is rate-limited rather than
|
|
123
|
+
unavailable.
|
|
124
|
+
|
|
125
|
+
**Per-dispatch override on Codex** goes through `spawn_agent`, and it MUST pass
|
|
126
|
+
`fork_turns: "none"` (or a positive integer). A full-history fork inherits the
|
|
127
|
+
parent model and reasoning effort and silently discards the override, so a
|
|
128
|
+
fallback that omits it appears to apply while changing nothing.
|
|
129
|
+
|
|
130
|
+
The three model ids above live in exactly three places - this table,
|
|
131
|
+
`pipeline/scripts/cost-table.json`, and the Phase 4 reviewer matrix - so an OpenAI
|
|
132
|
+
rename is a three-file change.
|
|
@@ -20,7 +20,7 @@ for proj in $(jq -r '.projects[] | "\(.name)\t\(.worktreePath)\t\(.baseBranch)"'
|
|
|
20
20
|
done
|
|
21
21
|
```
|
|
22
22
|
|
|
23
|
-
Same reviewer set (Fable
|
|
23
|
+
Same reviewer set (per host: Fable+Sonnet on Claude Code, Opus/GPT-5.4/Sonnet on Copilot CLI, gpt-5.6/gpt-5.4/gpt-5.6 on Codex CLI) receive `COMBINED_DIFF` with a multi-repo prefix in the system prompt:
|
|
24
24
|
|
|
25
25
|
```
|
|
26
26
|
This is a multi-repo task spanning {N} repos: {repo names}.
|
|
@@ -101,7 +101,7 @@ Every phase that dispatches a billable LLM agent MUST forward its token totals t
|
|
|
101
101
|
|
|
102
102
|
```bash
|
|
103
103
|
pipeline/scripts/log-metric.sh "$TASK_ID" <phase-id> <event> \
|
|
104
|
-
model=<fable|opus|sonnet|haiku|gpt-5.4> tokens_in=$IN tokens_out=$OUT tokens_cached=$CACHED duration_ms=$DUR
|
|
104
|
+
model=<fable|opus|sonnet|haiku|gpt-5.4|gpt-5.6|gpt-5.6-terra> tokens_in=$IN tokens_out=$OUT tokens_cached=$CACHED duration_ms=$DUR
|
|
105
105
|
LOG_METRIC_FORWARD_TO_TRACKER=1 pipeline/scripts/log-metric.sh "$TASK_ID" <phase-id> tokens \
|
|
106
106
|
model=<...> tokens_in=$IN tokens_out=$OUT tokens_cached=$CACHED
|
|
107
107
|
```
|
|
@@ -300,11 +300,23 @@ Branch name is deterministic - no user confirmation needed.
|
|
|
300
300
|
6. If empty after kebab (e.g. all-emoji title) → fall back to `task-{shortId}`
|
|
301
301
|
|
|
302
302
|
**Collision handling** (automatic - no prompt):
|
|
303
|
-
- Probe local + remote for existing branch
|
|
303
|
+
- Probe local + remote for existing branch. **Distinguish "no such ref" from "the
|
|
304
|
+
probe failed"**: with `2>/dev/null` and an empty-output test they look identical,
|
|
305
|
+
so an auth or network failure reads as "no collision" and the run creates a
|
|
306
|
+
branch that already exists on the remote - surfacing as a rejected push at
|
|
307
|
+
Phase 6, far from its cause.
|
|
304
308
|
```bash
|
|
305
309
|
LOCAL_HIT=$(git -C "$root" rev-parse --verify --quiet "refs/heads/$branch")
|
|
306
|
-
|
|
310
|
+
REMOTE_ERR=$(git -C "$root" ls-remote --exit-code --heads origin "$branch" 2>&1 >/dev/null)
|
|
311
|
+
REMOTE_RC=$?
|
|
312
|
+
# 0 = ref exists (collision) · 2 = no matching ref (authoritative "free")
|
|
313
|
+
# anything else = the probe itself failed; REMOTE_ERR holds why
|
|
307
314
|
```
|
|
315
|
+
- `REMOTE_RC` is 0 or 2 → treat as authoritative
|
|
316
|
+
- `REMOTE_RC` is anything else → the remote answer is **unknown**, not "free". Log
|
|
317
|
+
`Remote collision probe failed: <REMOTE_ERR>`, fall back to the local check only,
|
|
318
|
+
and record `"remoteCollisionProbe": "failed"` in `agent-state.json` so Phase 6
|
|
319
|
+
expects a possible non-fast-forward and re-checks before pushing.
|
|
308
320
|
- No collision → use as-is
|
|
309
321
|
- Collision found → append `-v2`, `-v3`, etc. until unique:
|
|
310
322
|
`bugfix/ABC-12345` exists → `bugfix/ABC-12345-v2`
|
|
@@ -519,3 +531,32 @@ Phase 7 cost rollup carries this as a `phase 0` line item so the user sees ambig
|
|
|
519
531
|
**Progress (per `$HOME/.claude/multi-agent-refs/progress-contract.md`):** emit one `→ <verb> <object>` line for each of: `→ parsing input`, `→ checking token <service>`, `→ scanning project <root>`, `→ creating worktree <repo>`, `→ binding identity <name>`, `→ writing state`. When `clarifyAmbiguous.enabled`, also emit `→ scoring task ambiguity` before Step 8 and `→ asking clarifying questions <N>` when `stopAndAsk` fires.
|
|
520
532
|
|
|
521
533
|
**Save preferences**: Write updated prefs to `$HOME/.claude/multi-agent-preferences.json` with all Phase 0 selections.
|
|
534
|
+
|
|
535
|
+
---
|
|
536
|
+
|
|
537
|
+
#### Phase 0 exit gate (BLOCKING - run before marking the phase completed)
|
|
538
|
+
|
|
539
|
+
Phase 0 owns `agent-state.json`. Do not call
|
|
540
|
+
`phase-tracker.sh update 0 completed` until this gate passes:
|
|
541
|
+
|
|
542
|
+
```bash
|
|
543
|
+
node "$HOME/.claude/scripts/phase0-exit-gate.mjs" "$TASK_ID" --input "$ORIGINAL_INPUT"
|
|
544
|
+
```
|
|
545
|
+
|
|
546
|
+
It asserts three things, each of which has failed silently in a real run:
|
|
547
|
+
|
|
548
|
+
1. **`agent-state.json` exists.** A run once reported Phase 0 `completed` with only
|
|
549
|
+
`tracker-state.json` on disk. Every later phase then reasons from fields that are
|
|
550
|
+
not there.
|
|
551
|
+
2. **`taskType` is set.** Phase 3 branches on it (Step 7). Absent, a Figma-driven
|
|
552
|
+
screen is dispatched as generic development, skipping the stack plugin's
|
|
553
|
+
token-compliance check, Code Connect publish and component review. That run
|
|
554
|
+
guessed `16` where the frame said `Spacing/12`, and half its commits were rework.
|
|
555
|
+
3. **A Figma reference forces `taskType: "component"`, and `figmaAccess.tier` is
|
|
556
|
+
recorded.** Without the tier, a later phase cannot tell "the design was confirmed"
|
|
557
|
+
from "the design was never fetched" - which is exactly when spacing gets guessed.
|
|
558
|
+
|
|
559
|
+
A failure is a halt, not a warning. Fix the state and re-run the gate; the phase
|
|
560
|
+
stays `in_progress` until it passes. **Never** mark Phase 0 completed on the grounds
|
|
561
|
+
that its steps ran - the gate checks the output, and the output is what Phase 3
|
|
562
|
+
consumes.
|
|
@@ -63,7 +63,14 @@ When `state.contextLinks[]` or the task description contains a Figma reference,
|
|
|
63
63
|
| 2 (REST) | `GET /v1/files/{fileKey}/nodes?ids={nodeId}` + `GET /v1/images/{fileKey}?ids={nodeId}&format=png&scale=2`, PAT via `~/.claude/lib/credential-store.sh get <logical-key>` (logical key = `prefs.global.keychainMapping.figma_pat`); canonical component resolved from repo `*.figma.swift` / `*.figma.kt` mapping keyed by `fileKey` + `nodeId` | same shape, but `codeConnectSnippets[]` is empty when repo mapping is absent (record an Open Question), `tier: 2` |
|
|
64
64
|
| 3 (screenshot) | User-attached screenshot stored alongside task evidence | degraded record: `codeConnectSnippets: []`, forced Open Question, `tier: 3` |
|
|
65
65
|
|
|
66
|
-
Persist results under `state.evidence.figma[]`. Halt the run if all three tiers fail; never substitute primitives or invent layout from prose.
|
|
66
|
+
Persist results under `state.evidence.figma[]`. Halt the run if all three tiers fail; never substitute primitives or invent layout from prose.
|
|
67
|
+
|
|
68
|
+
**Spacing goes in by token NAME, per atom - never a pixel number.** `tokens[]` must
|
|
69
|
+
carry each frame's spacing/padding as Figma names them (`Spacing/12`, edge `4`), keyed
|
|
70
|
+
to the atom. Phase 3 cannot call Figma, so what is missed here is gone: one run guessed
|
|
71
|
+
`16` where the frame said `Spacing/12` and the sheet was rebuilt. A pixel number also
|
|
72
|
+
cannot map back to a token. No spacing entries on a UI frame is a **capture failure**,
|
|
73
|
+
not an empty frame - Open Question and halt. Canonical chain reference: `pipeline/rules/figma-pipeline.md` "MUST: Figma access - 3-tier fallback chain".
|
|
67
74
|
|
|
68
75
|
**Telemetry (required for the no-MCP gate):** Tier 1 uses `mcp__claude_ai_Figma__*` tools. Every such MCP invocation MUST append an entry to `state.telemetry.mcpCalls[]` as `{ "tool": "<full mcp tool name>", "phase": 1, "timestamp": "<ISO-8601>" }`. This is the only phase permitted to record `phase: 1` (or `0`) entries; `smoke-no-mcp-in-dev-phases.sh` fails the run if any entry carries `phase >= 2`. Recording is what makes that BLOCKING contract enforceable - an MCP call left unrecorded defeats the gate, so record every one.
|
|
69
76
|
|
|
@@ -73,6 +80,22 @@ Progress lines:
|
|
|
73
80
|
→ figma evidence: <N> frames captured (tier=<n>, code-connect=<M>, open-questions=<K>)
|
|
74
81
|
```
|
|
75
82
|
|
|
83
|
+
#### Step 1.45 - Reuse discovery (BLOCKING for new services, entities, mappers)
|
|
84
|
+
|
|
85
|
+
Before proposing any new service call, entity or mapper, search for what already
|
|
86
|
+
covers it. Record hits under `state.reuse[]` and cite them in the doc; proposing new
|
|
87
|
+
code over a hit needs a one-line reason.
|
|
88
|
+
|
|
89
|
+
Search for: a **wrapper** over the same endpoint (especially one supplying parameters
|
|
90
|
+
the generated call leaves optional); an **entity** for the same concept (module's
|
|
91
|
+
shared entities first, then siblings); a **mapper** over the same response; a **screen**
|
|
92
|
+
doing the same interaction.
|
|
93
|
+
|
|
94
|
+
Why blocking: one run proposed a new repository over an endpoint a sibling already
|
|
95
|
+
wrapped **with its country parameter**, called the generated method without it, and
|
|
96
|
+
re-invented an entity the module had. Half that branch's commits went to converging
|
|
97
|
+
back. "Copy X and rename it" is the reuse answer, not a hint - name X's files.
|
|
98
|
+
|
|
76
99
|
#### Step 1.5 - External Context Injection (`state.contextLinks[]`)
|
|
77
100
|
|
|
78
101
|
Phase 0 Step 1b catalogued every typed external link from the task description into `state.contextLinks[]`. Phase 1 dispatches each entry to its matching fetcher (crashlytics, fortify, graylog, swagger, confluence, figma, generic-doc) and prepends results under a **Referenced External Sources** section in the analysis prompt - so the agent doesn't re-discover what the ticket already pointed at. `state.graylogContext` is injected there too, as diagnostic context (advisory only). Failures never fatal (a non-zero fetcher exit is marked skipped and the analysis still runs, exactly as for crashlytics); pending refs are advisories. Full dispatch table, exit-code handling, prompt injection shape, log line shape: `$HOME/.claude/multi-agent-refs/features/external-context-injection.md`.
|
|
@@ -367,3 +367,35 @@ bash $HOME/.claude/scripts/phase-tracker.sh tokens 3 <input_count> <output_count
|
|
|
367
367
|
The tracker accumulates the totals additively, so multiple calls in the same phase compound. The render output then shows live cost on the active phase tile (e.g. `Phase 3 Dev 2m 14s · 12.4k tok`). This satisfies the contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` and the `smoke-tracker-tokens-invocation.sh` enforcement gate. Skipping this call is the #1 cause of "I can't see how much it cost" complaints.
|
|
368
368
|
|
|
369
369
|
If you do not have access to the model's reported token counts, pass best-effort estimates derived from input length / output length - partial cost data is better than none.
|
|
370
|
+
|
|
371
|
+
|
|
372
|
+
#### Generated trees are not yours to edit
|
|
373
|
+
|
|
374
|
+
Many repos generate part of their source: a service client from an OpenAPI spec, mock
|
|
375
|
+
scenario indexes, localization keys, testing identifiers, design tokens. A generated
|
|
376
|
+
file is regenerated on the next build, so an edit there is lost silently, and the
|
|
377
|
+
matching hand-authored tree is the one that takes the change.
|
|
378
|
+
|
|
379
|
+
Before writing into any path, check whether it is generated:
|
|
380
|
+
|
|
381
|
+
```bash
|
|
382
|
+
# a Generated/ segment, or a header saying so, is the signal
|
|
383
|
+
find . -type d -name Generated -not -path './.*' | head
|
|
384
|
+
grep -rl "DO NOT EDIT\|auto-generated\|Generated by" --include="*.swift" --include="*.kt" . | head
|
|
385
|
+
```
|
|
386
|
+
|
|
387
|
+
The pairing is usually `Generated/<x>` for output and `Custom<X>/` or
|
|
388
|
+
`CustomSources/` for input. Two concrete shapes seen in the wild:
|
|
389
|
+
|
|
390
|
+
| Want to | Wrong place | Right place |
|
|
391
|
+
|---|---|---|
|
|
392
|
+
| add a mock fixture / named scenario | a `Fixtures/` file under a generated tree | the repo's custom fixture tree, plus registering the scenario in the generated index the build reads |
|
|
393
|
+
| add or change a service endpoint | the generated client method | the OpenAPI source the generator consumes, then regenerate |
|
|
394
|
+
|
|
395
|
+
One run wrote a mock fixture into the generated fixtures tree; the fix commit moved it
|
|
396
|
+
to the custom tree and registered the scenario in the generated index. Same content,
|
|
397
|
+
wrong side of the generator, and the Debug menu never showed it.
|
|
398
|
+
|
|
399
|
+
When the analysis doc has not recorded which trees are generated, that is a Phase 1
|
|
400
|
+
gap - say so rather than guessing, since guessing wrong is invisible until the next
|
|
401
|
+
regeneration.
|