@mmerterden/multi-agent-pipeline 13.5.0 → 14.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (119) hide show
  1. package/CHANGELOG.md +243 -0
  2. package/README.md +3 -3
  3. package/docs/features.md +1 -1
  4. package/install/_common.mjs +73 -0
  5. package/install/_mcp-register.mjs +70 -31
  6. package/install/_plugin-skills.mjs +73 -14
  7. package/install/claude.mjs +28 -4
  8. package/install/codex.mjs +33 -2
  9. package/install/copilot.mjs +145 -9
  10. package/install/index.mjs +10 -6
  11. package/install/templates/copilot-instructions.md +1 -1
  12. package/package.json +1 -1
  13. package/pipeline/agents/code-reviewer.md +58 -1
  14. package/pipeline/commands/multi-agent/SKILL.md +7 -5
  15. package/pipeline/commands/multi-agent/analysis/SKILL.md +7 -7
  16. package/pipeline/commands/multi-agent/analysis-resolve/SKILL.md +1 -1
  17. package/pipeline/commands/multi-agent/build-optimize/SKILL.md +7 -7
  18. package/pipeline/commands/multi-agent/channels/SKILL.md +5 -5
  19. package/pipeline/commands/multi-agent/dev/SKILL.md +23 -18
  20. package/pipeline/commands/multi-agent/dev-autopilot/SKILL.md +19 -13
  21. package/pipeline/commands/multi-agent/dev-local/SKILL.md +14 -12
  22. package/pipeline/commands/multi-agent/dev-local-autopilot/SKILL.md +17 -12
  23. package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +1 -1
  24. package/pipeline/commands/multi-agent/help/SKILL.md +4 -4
  25. package/pipeline/commands/multi-agent/ios-coding-standard/SKILL.md +2 -2
  26. package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +4 -4
  27. package/pipeline/commands/multi-agent/resume/SKILL.md +1 -1
  28. package/pipeline/commands/multi-agent/review/SKILL.md +5 -5
  29. package/pipeline/commands/multi-agent/scan/SKILL.md +1 -1
  30. package/pipeline/commands/multi-agent/search/SKILL.md +1 -1
  31. package/pipeline/commands/multi-agent/setup/SKILL.md +6 -6
  32. package/pipeline/commands/multi-agent/{finish → ship}/SKILL.md +12 -12
  33. package/pipeline/commands/multi-agent/testflight-validation/SKILL.md +1 -1
  34. package/pipeline/commands/multi-agent/update/SKILL.md +5 -2
  35. package/pipeline/commands/sim-test.md +2 -2
  36. package/pipeline/lib/credential-store-resolver.sh +16 -0
  37. package/pipeline/lib/credential-store.sh +47 -4
  38. package/pipeline/lib/fetch-figma-annotations.sh +26 -28
  39. package/pipeline/lib/figma-screenshot.sh +28 -39
  40. package/pipeline/lib/figma-token.sh +63 -0
  41. package/pipeline/multi-agent-refs/analysis-template.md +1 -1
  42. package/pipeline/multi-agent-refs/android-guide.md +1 -1
  43. package/pipeline/multi-agent-refs/channels/issue-comment.md +1 -1
  44. package/pipeline/multi-agent-refs/component-dispatch.md +2 -2
  45. package/pipeline/multi-agent-refs/cross-cli-contract.md +4 -4
  46. package/pipeline/multi-agent-refs/features/dev-critic.md +2 -2
  47. package/pipeline/multi-agent-refs/features/model-fallback.md +35 -2
  48. package/pipeline/multi-agent-refs/features/plan-todos.md +1 -1
  49. package/pipeline/multi-agent-refs/features/repo-map.md +1 -1
  50. package/pipeline/multi-agent-refs/features/review-multi-repo.md +3 -3
  51. package/pipeline/multi-agent-refs/features/shadow-git.md +1 -1
  52. package/pipeline/multi-agent-refs/features/skill-conformance.md +116 -0
  53. package/pipeline/multi-agent-refs/features/verify-by-test.md +1 -1
  54. package/pipeline/multi-agent-refs/generate-issue.md +1 -1
  55. package/pipeline/multi-agent-refs/multi-repo-integration-build.md +1 -1
  56. package/pipeline/multi-agent-refs/phases/log-format.md +4 -4
  57. package/pipeline/multi-agent-refs/phases/modes.md +7 -7
  58. package/pipeline/multi-agent-refs/phases/phase-0-init.md +13 -11
  59. package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +17 -15
  60. package/pipeline/multi-agent-refs/phases/phase-2-planning.md +7 -7
  61. package/pipeline/multi-agent-refs/phases/phase-3-dev.md +28 -13
  62. package/pipeline/multi-agent-refs/phases/phase-4-review.md +90 -58
  63. package/pipeline/multi-agent-refs/phases/phase-5-test.md +7 -7
  64. package/pipeline/multi-agent-refs/phases/phase-6-commit.md +8 -8
  65. package/pipeline/multi-agent-refs/phases/phase-7-report.md +8 -8
  66. package/pipeline/multi-agent-refs/phases.md +13 -13
  67. package/pipeline/multi-agent-refs/progress-contract.md +2 -2
  68. package/pipeline/multi-agent-refs/rules.md +7 -5
  69. package/pipeline/multi-agent-refs/swiftui-guide.md +1 -1
  70. package/pipeline/multi-agent-refs/tracker-contract.md +16 -15
  71. package/pipeline/preferences-template.json +7 -1
  72. package/pipeline/rules/figma-pipeline.md +2 -2
  73. package/pipeline/schemas/agent-state.schema.json +333 -79
  74. package/pipeline/schemas/criteria-manifest.schema.json +228 -0
  75. package/pipeline/schemas/migrations/prefs-2.4.0-to-2.5.0.mjs +64 -0
  76. package/pipeline/schemas/prefs.schema.json +118 -262
  77. package/pipeline/schemas/reviewer-output.schema.json +48 -3
  78. package/pipeline/schemas/token-budget.json +34 -10
  79. package/pipeline/schemas/triage-output.schema.json +112 -27
  80. package/pipeline/scripts/cost-table.json +7 -4
  81. package/pipeline/scripts/gc-worktrees.sh +1 -1
  82. package/pipeline/scripts/gen-mode-dispatch.mjs +6 -6
  83. package/pipeline/scripts/match-skills.mjs +37 -4
  84. package/pipeline/scripts/migrate-prefs.mjs +88 -17
  85. package/pipeline/scripts/phase-tracker.sh +14 -3
  86. package/pipeline/scripts/pre-commit-check.sh +49 -2
  87. package/pipeline/scripts/skill-conformance.mjs +960 -0
  88. package/pipeline/scripts/smoke-schema-validation.sh +17 -4
  89. package/pipeline/scripts/uninstall.mjs +35 -9
  90. package/pipeline/scripts/validate-reviewer.mjs +108 -1
  91. package/pipeline/skills/.skill-manifest.json +1 -1
  92. package/pipeline/skills/.skills-index.json +36 -9
  93. package/pipeline/skills/shared/README.md +15 -12
  94. package/pipeline/skills/shared/core/apple-archive-compliance/SKILL.md +1 -0
  95. package/pipeline/skills/shared/core/apple-archive-compliance/references/rules.yml +167 -0
  96. package/pipeline/skills/shared/core/google-play-compliance/SKILL.md +1 -0
  97. package/pipeline/skills/shared/core/google-play-compliance/references/rules.yml +184 -0
  98. package/pipeline/skills/shared/core/multi-agent/SKILL.md +10 -10
  99. package/pipeline/skills/shared/core/multi-agent-analysis/SKILL.md +4 -4
  100. package/pipeline/skills/shared/core/multi-agent-analysis-resolve/SKILL.md +3 -3
  101. package/pipeline/skills/shared/core/multi-agent-build-optimize/SKILL.md +2 -2
  102. package/pipeline/skills/shared/core/multi-agent-create-jira/SKILL.md +1 -1
  103. package/pipeline/skills/shared/core/multi-agent-dev/SKILL.md +6 -5
  104. package/pipeline/skills/shared/core/multi-agent-dev-autopilot/SKILL.md +7 -6
  105. package/pipeline/skills/shared/core/multi-agent-dev-local/SKILL.md +4 -3
  106. package/pipeline/skills/shared/core/multi-agent-dev-local-autopilot/SKILL.md +2 -1
  107. package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +2 -2
  108. package/pipeline/skills/shared/core/multi-agent-ios-coding-standard/SKILL.md +1 -1
  109. package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +4 -4
  110. package/pipeline/skills/shared/core/multi-agent-review/SKILL.md +5 -5
  111. package/pipeline/skills/shared/core/multi-agent-scan/SKILL.md +1 -1
  112. package/pipeline/skills/shared/core/multi-agent-search/SKILL.md +1 -1
  113. package/pipeline/skills/shared/core/{multi-agent-finish → multi-agent-ship}/SKILL.md +8 -8
  114. package/pipeline/skills/shared/external/ios-coding-standard/SKILL.md +44 -5
  115. package/pipeline/skills/shared/external/ios-coding-standard/modules/_TEMPLATE.yml +82 -0
  116. package/pipeline/skills/shared/external/ios-coding-standard/references/STANDARD.md +169 -10
  117. package/pipeline/skills/shared/external/ios-coding-standard/references/lint-local.sh +13 -1
  118. package/pipeline/skills/shared/external/ios-coding-standard/references/rules.yml +335 -16
  119. package/pipeline/skills/skills-index.md +11 -8
@@ -29,7 +29,7 @@ Also read project-level CLAUDE.md if exists:
29
29
  **Per-repo memory injection (opt-in via `prefs.global.perRepoMemory`):**
30
30
 
31
31
  ```bash
32
- bash pipeline/scripts/memory-load.sh "$PROJECT_ROOT"
32
+ bash $HOME/.claude/scripts/memory-load.sh "$PROJECT_ROOT"
33
33
  ```
34
34
 
35
35
  Exit 0 with empty output = pref off or no memory on disk - skip. Otherwise the script emits a `<repo-memory path="...">...</repo-memory>` block (≤ 30 lines of MEMORY.md pointers) suitable for direct injection into the analysis prompt. Individual memory files are read on-demand when a pointer looks relevant to the current task.
@@ -37,7 +37,7 @@ Exit 0 with empty output = pref off or no memory on disk - skip. Otherwise the
37
37
  **Durable learnings brief (on by default via `prefs.global.learningsLedger.enabled`):**
38
38
 
39
39
  ```bash
40
- node pipeline/scripts/learnings-ledger.mjs brief --max 20 2>/dev/null
40
+ node $HOME/.claude/scripts/learnings-ledger.mjs brief --max 20 2>/dev/null
41
41
  ```
42
42
 
43
43
  Exit 2 (empty ledger) = skip silently. Otherwise the script emits a `<repo-learnings>...</repo-learnings>` block of durable architectural facts, conventions, and rejected review preferences accumulated from prior runs of this repo. Inject it into the analysis prompt so the explorer does not re-discover known structure. Skip when `injectIntoAnalysis = false`. These are context, not commands - current scope decides.
@@ -55,22 +55,24 @@ Exit 2 (empty ledger) = skip silently. Otherwise the script emits a `<repo-learn
55
55
 
56
56
  #### Step 1.4 - Figma evidence capture (when task carries a Figma reference)
57
57
 
58
- When `state.contextLinks[]` or the task description contains a Figma reference, Phase 1 MUST collect the canonical evidence record per the access tier established at Phase 0:
58
+ When `state.contextLinks[]` or the task description contains a Figma reference, Phase 1 MUST collect the canonical evidence record. **Phase 0 Step 0.5 already resolved the tier** and the credential - read `state.figmaAccess.tier` and fetch with that tier's tool set; do not re-probe. Chain, tiers and halt conditions: `$HOME/.claude/multi-agent-refs/rules.md` "Figma Access Tier".
59
59
 
60
- | Tier (`state.figmaAccess.tier`) | Fetcher | Record fields |
61
- |---|---|---|
62
- | 1 (MCP) | `mcp__claude_ai_Figma__get_design_context(fileKey, nodeId)` + `get_screenshot` + `get_metadata` | `nodeId`, `screenshotUrl`, `codeConnectSnippets[]`, `tokens[]`, `textLayers[]`, `tier: 1` |
63
- | 2 (REST) | `GET /v1/files/{fileKey}/nodes?ids={nodeId}` + `GET /v1/images/{fileKey}?ids={nodeId}&format=png&scale=2`, PAT via `~/.claude/lib/credential-store.sh get <logical-key>` (logical key = `prefs.global.keychainMapping.figma_pat`); canonical component resolved from repo `*.figma.swift` / `*.figma.kt` mapping keyed by `fileKey` + `nodeId` | same shape, but `codeConnectSnippets[]` is empty when repo mapping is absent (record an Open Question), `tier: 2` |
64
- | 3 (screenshot) | User-attached screenshot stored alongside task evidence | degraded record: `codeConnectSnippets: []`, forced Open Question, `tier: 3` |
60
+ What Phase 1 owns is the record. Every frame gets one entry in `state.evidence.figma[]`:
65
61
 
66
- Persist results under `state.evidence.figma[]`. Halt the run if all three tiers fail; never substitute primitives or invent layout from prose.
62
+ | Field | Tier 1 (MCP) | Tier 2 (REST) | Tier 3 (screenshot) |
63
+ |---|---|---|---|
64
+ | `nodeId`, `screenshotUrl`, `tokens[]`, `textLayers[]` | required | required | required |
65
+ | `codeConnectSnippets[]` | from `CodeConnectSnippet` blocks | from repo `*.figma.swift` / `*.figma.kt` keyed on `fileKey`+`nodeId`; empty -> Open Question | always `[]` -> forced Open Question |
66
+ | `tier` | `1` | `2` | `3` |
67
+
68
+ Halt if all three tiers fail; never substitute primitives or invent layout from prose.
67
69
 
68
70
  **Spacing goes in by token NAME, per atom - never a pixel number.** `tokens[]` must
69
71
  carry each frame's spacing/padding as Figma names them (`Spacing/12`, edge `4`), keyed
70
72
  to the atom. Phase 3 cannot call Figma, so what is missed here is gone: one run guessed
71
73
  `16` where the frame said `Spacing/12` and the sheet was rebuilt. A pixel number also
72
74
  cannot map back to a token. No spacing entries on a UI frame is a **capture failure**,
73
- not an empty frame - Open Question and halt. Canonical chain reference: `pipeline/rules/figma-pipeline.md` "MUST: Figma access - 3-tier fallback chain".
75
+ not an empty frame - Open Question and halt. Canonical chain reference: `$HOME/.claude/rules/figma-pipeline.md` "MUST: Figma access - 3-tier fallback chain".
74
76
 
75
77
  **Telemetry (required for the no-MCP gate):** Tier 1 uses `mcp__claude_ai_Figma__*` tools. Every such MCP invocation MUST append an entry to `state.telemetry.mcpCalls[]` as `{ "tool": "<full mcp tool name>", "phase": 1, "timestamp": "<ISO-8601>" }`. This is the only phase permitted to record `phase: 1` (or `0`) entries; `smoke-no-mcp-in-dev-phases.sh` fails the run if any entry carries `phase >= 2`. Recording is what makes that BLOCKING contract enforceable - an MCP call left unrecorded defeats the gate, so record every one.
76
78
 
@@ -130,7 +132,7 @@ This informs:
130
132
 
131
133
  #### Step 2.5 - Repo Map Injection (advisory, opt-in)
132
134
 
133
- Gated by `prefs.global.repoMap.enabled` (default: `false`). When enabled, runs `pipeline/scripts/repo-map.mjs` and injects the budgeted result into each Explore prompt as `${REPO_MAP}`. Aider-style: deterministic, no embeddings, sub-second, advisory only. Full wiring (helper invocation, properties, when-to-enable): `$HOME/.claude/multi-agent-refs/features/repo-map.md`.
135
+ Gated by `prefs.global.repoMap.enabled` (default: `false`). When enabled, runs `$HOME/.claude/scripts/repo-map.mjs` and injects the budgeted result into each Explore prompt as `${REPO_MAP}`. Aider-style: deterministic, no embeddings, sub-second, advisory only. Full wiring (helper invocation, properties, when-to-enable): `$HOME/.claude/multi-agent-refs/features/repo-map.md`.
134
136
 
135
137
  #### Step 3 - Codebase Exploration
136
138
 
@@ -152,7 +154,7 @@ The light tier keeps a one-line bug fix from triggering a full-repo scan; pairin
152
154
 
153
155
  #### Output contract
154
156
 
155
- Phase 1 produces an object conforming to `pipeline/schemas/analysis-output.schema.json` and persists it to `state.analysis`. Required fields (exact names per the schema): `stack` (detected stack identifier + primary language), `touchedAreas[]` (path + why), `risks[]` (existing-code hazards/observations the planner must respect - each `{risk, severity, mitigation}`; use an empty array when none), `summary` (one-paragraph human-readable). Phase 2 reads this object as its sole input - see `phase-2-planning.md`'s Input contract.
157
+ Phase 1 produces an object conforming to `$HOME/.claude/schemas/analysis-output.schema.json` and persists it to `state.analysis`. Required fields (exact names per the schema): `stack` (detected stack identifier + primary language), `touchedAreas[]` (path + why), `risks[]` (existing-code hazards/observations the planner must respect - each `{risk, severity, mitigation}`; use an empty array when none), `summary` (one-paragraph human-readable). Phase 2 reads this object as its sole input - see `phase-2-planning.md`'s Input contract.
156
158
 
157
159
  **Required: validator gate (deterministic) - run on the persisted file immediately after the analysis object is produced; the validator's exit code decides, not the LLM turn:**
158
160
 
@@ -165,7 +167,7 @@ node $HOME/.claude/scripts/validate-analysis.mjs "$ANALYSIS_FILE"
165
167
 
166
168
  Progress line: ` → checking validator validate-analysis`
167
169
 
168
- Non-zero exit fails CLOSED: emit the validator stderr + `errors[]` verbatim, attempt ONE self-correction rework (re-invoke the explorer with the errors quoted, overwrite `$ANALYSIS_FILE`), re-run the validator. If it fails again -> HALT the phase with recovery hint: `ERR: analysis output failed validate-analysis.mjs twice. Inspect $ANALYSIS_FILE against pipeline/schemas/analysis-output.schema.json, then resume with /multi-agent:resume #N.` Record `agent-state.phases["1"].validator` (`pass` | `pass-after-rework` | `halted`).
170
+ Non-zero exit fails CLOSED: emit the validator stderr + `errors[]` verbatim, attempt ONE self-correction rework (re-invoke the explorer with the errors quoted, overwrite `$ANALYSIS_FILE`), re-run the validator. If it fails again -> HALT the phase with recovery hint: `ERR: analysis output failed validate-analysis.mjs twice. Inspect $ANALYSIS_FILE against $HOME/.claude/schemas/analysis-output.schema.json, then resume with /multi-agent:resume #N.` Record `agent-state.phases["1"].validator` (`pass` | `pass-after-rework` | `halted`).
169
171
 
170
172
  Log: "Phase 1: Analysis - stack:{detectedStack} | {N} files identified, {summary}"
171
173
 
@@ -174,7 +176,7 @@ Log: "Phase 1: Analysis - stack:{detectedStack} | {N} files identified, {summa
174
176
  Forward the explorer/Opus call's token totals into the tracker so Phase 7's Cost Breakdown captures Phase 1:
175
177
 
176
178
  ```bash
177
- LOG_METRIC_FORWARD_TO_TRACKER=1 pipeline/scripts/log-metric.sh "$TASK_ID" 1 analysis.completed \
179
+ LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 1 analysis.completed \
178
180
  model=opus tokens_in=$IN tokens_out=$OUT duration_ms=$DUR
179
181
  ```
180
182
 
@@ -185,7 +187,7 @@ Best-effort. See `$HOME/.claude/multi-agent-refs/progress-contract.md#token-tele
185
187
  After the explorer returns its summary, consult the per-repo triage corpus for similar past tasks. Inject up to 3 matches into the analysis output as `priorArt[]` so Phase 2 planning can read them. Disabled when `prefs.global.priorArtEnrichment.enabled = false`.
186
188
 
187
189
  ```bash
188
- PRIOR=$(node pipeline/scripts/triage-memory.mjs query \
190
+ PRIOR=$(node $HOME/.claude/scripts/triage-memory.mjs query \
189
191
  --issue "$TASK_TITLE $TASK_DESCRIPTION" --top 3 2>/dev/null | jq -c '.hits // []')
190
192
  ```
191
193
 
@@ -25,7 +25,7 @@ Phase 2 Planning consumes the analysis document. MCP forbidden.
25
25
 
26
26
  #### Input contract
27
27
 
28
- Phase 2 consumes the Phase 1 output object conforming to `pipeline/schemas/analysis-output.schema.json`. Read `state.analysis` (the explorer's return value) and treat its `touchedAreas` and `risks` arrays (plus `stack` and `summary`) as authoritative input - do not re-explore the codebase here. (Field names are exactly those in the analysis schema; the planner's own `targetFiles` belongs to `planning-output.schema.json`, not the analysis input.)
28
+ Phase 2 consumes the Phase 1 output object conforming to `$HOME/.claude/schemas/analysis-output.schema.json`. Read `state.analysis` (the explorer's return value) and treat its `touchedAreas` and `risks` arrays (plus `stack` and `summary`) as authoritative input - do not re-explore the codebase here. (Field names are exactly those in the analysis schema; the planner's own `targetFiles` belongs to `planning-output.schema.json`, not the analysis input.)
29
29
 
30
30
  #### Step 1 - Task Decomposition
31
31
 
@@ -105,7 +105,7 @@ Based on Phase 1 `detectedStack`, assign relevant skills:
105
105
 
106
106
  #### Output contract
107
107
 
108
- Phase 2 produces an object conforming to `pipeline/schemas/planning-output.schema.json` - `tasks[]` with `id`, `subject`, `targetFiles`, `complexity`, `blockedBy`, plus optional `architectureNotes` and `mode`. Phase 3 reads `tasks[]` in dependency order; the schema's `blockedBy` field drives the ready-task picker.
108
+ Phase 2 produces an object conforming to `$HOME/.claude/schemas/planning-output.schema.json` - `tasks[]` with `id`, `subject`, `targetFiles`, `complexity`, `blockedBy`, plus optional `architectureNotes` and `mode`. Phase 3 reads `tasks[]` in dependency order; the schema's `blockedBy` field drives the ready-task picker.
109
109
 
110
110
  **Required: validator gate (deterministic) - run on the persisted file before the approval gate renders the plan; the validator's exit code decides, not the LLM turn:**
111
111
 
@@ -118,13 +118,13 @@ node $HOME/.claude/scripts/validate-planning.mjs "$PLAN_FILE"
118
118
 
119
119
  Progress line: ` → checking validator validate-planning`
120
120
 
121
- Non-zero exit fails CLOSED: emit the validator stderr + `errors[]` verbatim, attempt ONE self-correction rework (re-invoke the planner with the errors quoted, overwrite `$PLAN_FILE`), re-run the validator. If it fails again -> HALT the phase (never enter Phase 3 with an invalid plan) with recovery hint: `ERR: plan output failed validate-planning.mjs twice. Inspect $PLAN_FILE against pipeline/schemas/planning-output.schema.json, then resume with /multi-agent:resume #N.` Record `agent-state.phases["2"].validator` (`pass` | `pass-after-rework` | `halted`).
121
+ Non-zero exit fails CLOSED: emit the validator stderr + `errors[]` verbatim, attempt ONE self-correction rework (re-invoke the planner with the errors quoted, overwrite `$PLAN_FILE`), re-run the validator. If it fails again -> HALT the phase (never enter Phase 3 with an invalid plan) with recovery hint: `ERR: plan output failed validate-planning.mjs twice. Inspect $PLAN_FILE against $HOME/.claude/schemas/planning-output.schema.json, then resume with /multi-agent:resume #N.` Record `agent-state.phases["2"].validator` (`pass` | `pass-after-rework` | `halted`).
122
122
 
123
123
  Log: "Phase 2: Plan - {N} tasks created, {M} with architecture review, validator:pass"
124
124
 
125
125
  #### Step 4.5 - Emit Plan Todo List (opt-in)
126
126
 
127
- **Gated by `prefs.global.planTodos.enabled`** (default: `false`). When enabled, after the planning-output JSON validates and BEFORE the approval gate, transform `tasks[]` into a structured Todo list conforming to `pipeline/schemas/plan-todos.schema.json` and persist into `agent-state.plan`. The plan is rendered as a live, always-visible Todo list.
127
+ **Gated by `prefs.global.planTodos.enabled`** (default: `false`). When enabled, after the planning-output JSON validates and BEFORE the approval gate, transform `tasks[]` into a structured Todo list conforming to `$HOME/.claude/schemas/plan-todos.schema.json` and persist into `agent-state.plan`. The plan is rendered as a live, always-visible Todo list.
128
128
 
129
129
  ```bash
130
130
  TODO_BLOB=$(jq '
@@ -179,7 +179,7 @@ In the skipped case, log `🧠 Phase 2: Plan - gate skipped ({mode}), proceedi
179
179
  **Only relevant when `state.autopilot === true` and `prefs.global.autopilotSafetyGate !== false`.** The classifier protects autopilot's zero-interaction contract from edge cases where silent execution is genuinely dangerous (security-path touch, schema migration, many-file sprawl, delete-without-paired-test).
180
180
 
181
181
  ```bash
182
- verdict=$(node pipeline/scripts/classify-plan-safety.mjs <(echo "$PLAN_JSON"))
182
+ verdict=$(node $HOME/.claude/scripts/classify-plan-safety.mjs <(echo "$PLAN_JSON"))
183
183
  pause=$(jq -r '.recommendPause' <<< "$verdict")
184
184
  score=$(jq -r '.score' <<< "$verdict")
185
185
  reasons=$(jq -r '.reasons[] | "• \(.rule) (+\(.weight)): \(.detail)"' <<< "$verdict")
@@ -289,7 +289,7 @@ Handle the selection:
289
289
 
290
290
  No hard cap on edit iterations - the user controls exit via the Approve / Cancel options. Between iterations, keep only the **latest plan** as canonical; previous renders are in the log for audit but do not re-enter the validator.
291
291
 
292
- **Validator**: every revised plan goes through `node pipeline/scripts/validate-planning.mjs -` before re-render. If validation fails after an edit, log `⚠️ Phase 2: Plan validator failed after edit request #N - retrying Opus once` and retry once; on second failure surface the validator error to the user and go back to approval prompt with the pre-edit plan.
292
+ **Validator**: every revised plan goes through `node $HOME/.claude/scripts/validate-planning.mjs -` before re-render. If validation fails after an edit, log `⚠️ Phase 2: Plan validator failed after edit request #N - retrying Opus once` and retry once; on second failure surface the validator error to the user and go back to approval prompt with the pre-edit plan.
293
293
 
294
294
  #### Step 6 - Mode-specific short-circuit (reference)
295
295
 
@@ -309,7 +309,7 @@ The 4 pipeline modes interact with the gate as follows. This table is the source
309
309
  After plan generation (and after each edit-loop iteration), forward Opus call totals so Phase 7's Cost Breakdown captures Phase 2:
310
310
 
311
311
  ```bash
312
- LOG_METRIC_FORWARD_TO_TRACKER=1 pipeline/scripts/log-metric.sh "$TASK_ID" 2 plan.generated \
312
+ LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 2 plan.generated \
313
313
  model=opus tokens_in=$IN tokens_out=$OUT duration_ms=$DUR iteration=$N
314
314
  ```
315
315
 
@@ -6,9 +6,11 @@
6
6
 
7
7
  Per Locked decision 30, Phase 3 Dev consumes the analysis document as the sole design source. MCP / Figma REST / Figma URL fetches are forbidden.
8
8
 
9
- Pre-flight steps (run in order, abort on failure):
9
+ Pre-flight steps (run in order, abort on failure).
10
10
 
11
- 1. **Analysis document presence**: locate `analysis/<feature-slug>-<platform>.md` for the active platform.
11
+ **Steps 1, 2, 3, 5 and 6 apply only when Phase 1 ran.** In the `--dev` family (`state.onlyDevelop === true`) there is no analysis document by design, so those steps are recorded as `not-applicable (no Phase 1 in this mode)` and skipped - an unconditional abort here would make every fast mode impossible, which is the contradiction the modes have always carried in practice. Steps 4, 7 and 8 apply in every mode.
12
+
13
+ 1. **Analysis document presence** (Phase 1 modes only): locate `analysis/<feature-slug>-<platform>.md` for the active platform.
12
14
  - Path resolution: `state.run.repoPath` + `/analysis/<feature>-<platform>.md`
13
15
  - Multi-repo runs: each selected repo MUST have the file under its working tree
14
16
  - **Abort condition**: file missing -> exit with `ERR: analysis doc missing for <platform>. Run /multi-agent:analysis "<feature>" first to generate it.`
@@ -28,6 +30,16 @@ Pre-flight steps (run in order, abort on failure):
28
30
 
29
31
  7. **MCP forbidden**: any attempt to call `mcp__claude_ai_Figma__*` in Phase 3 is a violation. The smoke gate `smoke-no-mcp-in-dev-phases.sh` (see CHANGELOG [9.0.0]) checks `state.telemetry.mcpCalls[]` after run and fails if Phase 3 contributed any entry.
30
32
 
33
+ 8. **Criteria ledger (required, every mode)**: the moment this phase consults a skill, a marketplace plugin skill, a stack guide or a module `CLAUDE.md` in order to write code, append an entry to `state.telemetry.skillCalls[]`:
34
+
35
+ ```json
36
+ {"skill": "ios-coding-standard", "phase": 3, "targetFiles": ["Sources/Login/LoginViewModel.swift"], "timestamp": "<ISO-8601>"}
37
+ ```
38
+
39
+ `targetFiles` is required: a skill applied to the wrong files still reads as "applied" without it, so coverage could not be attributed. Append at the moment of consultation, not retrospectively at the end of the phase.
40
+
41
+ This is a **self-report**, and Phase 4 treats it as such. Step 1.78 resolves the criteria independently and defaults `ledger.source` to `derived`; a skill named here that the resolver cannot bind to a changed file is flagged rather than believed. The record still earns its keep, because it is the only signal that separates "this skill was applied to the wrong files" from "this skill was never opened" - and because a phase that has to name what it followed tends to follow something.
42
+
31
43
  The analysis document is the SOLE design source in Phase 3. Variant choices, padding values, color tokens, copy strings, accessibility identifiers, and test method names all come from the rendered Pass B cells. If something is missing in the analysis doc, the fix is to re-run `/multi-agent:analysis`, not to fetch from Figma.
32
44
 
33
45
  <!-- progress-contract: applied -->
@@ -36,11 +48,11 @@ The analysis document is the SOLE design source in Phase 3. Variant choices, pad
36
48
 
37
49
  #### Input contract
38
50
 
39
- Phase 3 consumes the Phase 2 output object conforming to `pipeline/schemas/planning-output.schema.json` - the task graph (`tasks[]` with `id`, `subject`, `targetFiles`, `complexity`, `blockedBy`) plus the architecture review notes. Tasks execute in dependency order; the schema's `blockedBy` field drives the ready-task picker. In `--dev` mode (no Phase 2), Sonnet/Opus generates the equivalent task list inline before entering the loop below.
51
+ Phase 3 consumes the Phase 2 output object conforming to `$HOME/.claude/schemas/planning-output.schema.json` - the task graph (`tasks[]` with `id`, `subject`, `targetFiles`, `complexity`, `blockedBy`) plus the architecture review notes. Tasks execute in dependency order; the schema's `blockedBy` field drives the ready-task picker. In `--dev` mode (no Phase 2), Sonnet/Opus generates the equivalent task list inline before entering the loop below.
40
52
 
41
- **Plan Todo iteration (opt-in)**: gated by `prefs.global.planTodos.enabled` (default: `false`). When enabled and Phase 2 Step 4.5 emitted a `plan.todos[]`, Phase 3 iterates via `pipeline/lib/plan-todos.sh next/start/complete/fail` instead of walking `tasks[]` directly. When disabled, the loop walks `tasks[]` from `planning-output` - TDD contract is unchanged. Full helper loop + state semantics: `$HOME/.claude/multi-agent-refs/features/plan-todos.md`.
53
+ **Plan Todo iteration (opt-in)**: gated by `prefs.global.planTodos.enabled` (default: `false`). When enabled and Phase 2 Step 4.5 emitted a `plan.todos[]`, Phase 3 iterates via `$HOME/.claude/lib/plan-todos.sh next/start/complete/fail` instead of walking `tasks[]` directly. When disabled, the loop walks `tasks[]` from `planning-output` - TDD contract is unchanged. Full helper loop + state semantics: `$HOME/.claude/multi-agent-refs/features/plan-todos.md`.
42
54
 
43
- **Shadow-Git checkpoints (opt-in)**: gated by `prefs.global.shadowGit.enabled` (default: `false`). When enabled, the orchestrator snapshots the worktree via `pipeline/lib/shadow-git.sh` so sub-phase rollback is possible without touching the project's real `.git` history. Lifecycle: `shadow-git.sh init` (Phase 0 baseline), `shadow-git.sh snapshot` (per step after `plan-todos complete`), `shadow-git.sh restore <sha> --files` (rollback). Modes: `per-todo-step` (default) or `per-tool-call`. Full wiring + storage cap: `$HOME/.claude/multi-agent-refs/features/shadow-git.md`.
55
+ **Shadow-Git checkpoints (opt-in)**: gated by `prefs.global.shadowGit.enabled` (default: `false`). When enabled, the orchestrator snapshots the worktree via `$HOME/.claude/lib/shadow-git.sh` so sub-phase rollback is possible without touching the project's real `.git` history. Lifecycle: `shadow-git.sh init` (Phase 0 baseline), `shadow-git.sh snapshot` (per step after `plan-todos complete`), `shadow-git.sh restore <sha> --files` (rollback). Modes: `per-todo-step` (default) or `per-tool-call`. Full wiring + storage cap: `$HOME/.claude/multi-agent-refs/features/shadow-git.md`.
44
56
 
45
57
  #### Component tasks - delegated dispatch (taskType === "component")
46
58
 
@@ -78,7 +90,7 @@ If the latest iteration has `triage.approved === true` AND `accepted === []`, Ph
78
90
  **Telemetry**: at the start of every re-entry, emit:
79
91
 
80
92
  ```bash
81
- pipeline/scripts/log-metric.sh "$TASK_ID" 3 rework.started \
93
+ $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 rework.started \
82
94
  iteration=$ITERATION accepted_blocking=$BLOCKING accepted_important=$IMPORTANT
83
95
  ```
84
96
 
@@ -241,7 +253,7 @@ release_build_lock() { rm -rf "$BUILD_LOCK"; }
241
253
 
242
254
  #### Step 3.5 - Dev Critic (Evaluator-Optimizer, opt-in)
243
255
 
244
- Gated by `prefs.global.devCritic.enabled` (default: `false`). When enabled, after the generator's last edit and BEFORE Phase 4: dispatch `dev-critic` sub-agent (Sonnet), run 4 deterministic gates (build / lint / test / secrets), then platform checklist. STRICT loop cap - **max 2 iterations** (round 2 re-checks round-1 failures only); round 3+ returns `escalate: true`. Severity routing: `blocking` → generator must fix; `important` → SHOULD fix or pass through to Phase 4; `suggestion` → generator's judgement. Full agent contract, schema, telemetry, when-to-enable: `$HOME/.claude/multi-agent-refs/features/dev-critic.md` + `pipeline/agents/dev-critic.md`.
256
+ Gated by `prefs.global.devCritic.enabled` (default: `false`). When enabled, after the generator's last edit and BEFORE Phase 4: dispatch `dev-critic` sub-agent (Sonnet), run 4 deterministic gates (build / lint / test / secrets), then platform checklist. STRICT loop cap - **max 2 iterations** (round 2 re-checks round-1 failures only); round 3+ returns `escalate: true`. Severity routing: `blocking` → generator must fix; `important` → SHOULD fix or pass through to Phase 4; `suggestion` → generator's judgement. Full agent contract, schema, telemetry, when-to-enable: `$HOME/.claude/multi-agent-refs/features/dev-critic.md` + `$HOME/.claude/agents/dev-critic.md`.
245
257
 
246
258
  ---
247
259
 
@@ -262,7 +274,7 @@ After the build/test green step and BEFORE Phase 4 handoff, run one diff-shrink
262
274
  5. **Record tokens in the cost ledger** so Phase 7's Cost Breakdown captures the pass:
263
275
 
264
276
  ```bash
265
- LOG_METRIC_FORWARD_TO_TRACKER=1 pipeline/scripts/log-metric.sh "$TASK_ID" 3 dev.simplifier_pass \
277
+ LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 dev.simplifier_pass \
266
278
  model=sonnet tokens_in=$IN tokens_out=$OUT duration_ms=$DUR \
267
279
  edits_returned=$RET edits_applied=$APPLIED edits_skipped=$SKIPPED
268
280
  ```
@@ -292,8 +304,11 @@ When `agent-state.json.onlyDevelop === true`, Phase 3 runs self-contained with *
292
304
  | Task granularity | Per-plan-item | Agent decides |
293
305
  | Status updates | Per task item | Single in_progress → completed |
294
306
  | Scope confirmation | Phase 2 user approval | None (agent autonomous) |
307
+ | Review of the result | Phase 4 | Phase 4 (same) |
308
+
309
+ Because the agent determines its own scope here, Phase 4 is the only place that checks the result against anything external. Record every skill, plugin skill and guide consulted during this phase into `state.telemetry.skillCalls[]` with the files it was applied to - Phase 4 resolves the criteria set independently, and this record is what lets it tell "applied and honoured" from "never opened".
295
310
 
296
- **Combinable with autopilot**: `--dev autopilot` = zero interaction. Init → Dev → auto-commit → auto-PR → Report.
311
+ **Combinable with autopilot**: `--dev autopilot` = zero interaction. Init → Dev → Review (auto-fix) → auto-commit → auto-PR → Report.
297
312
 
298
313
  **Tracker visibility during Opus dispatch**: on Claude Code the model switch to Opus happens via subagent dispatch, and the parent widget cannot move while an Agent call is in flight. Dispatch per task from the self-generated task list (never one monolithic call for the whole phase), set the pre-dispatch `activeForm` marker, and record tokens between chunks - full rules in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "Delegated phases".
299
314
 
@@ -333,21 +348,21 @@ If a todo has no `repo` tag in multi-repo mode → log warning + ask user, do no
333
348
  **Recording a pass (default-FAIL evidence gate):** before setting `buildStatus.ok = true`, the build output must be tee'd to a log and that log must substantiate the success - a zero exit code alone is not trusted. Run the evidence gate; on exit 1, do NOT record a pass:
334
349
  ```bash
335
350
  <build-command> 2>&1 | tee "$WORKTREE/.build.log"
336
- node pipeline/scripts/evidence-gate.mjs --claim build --status passed --evidence "$WORKTREE/.build.log" \
351
+ node $HOME/.claude/scripts/evidence-gate.mjs --claim build --status passed --evidence "$WORKTREE/.build.log" \
337
352
  || { echo "build pass unverified - treat as failure"; /* keep buildStatus.ok=false */ }
338
353
  ```
339
354
  This closes the gap where an agent records "built" without ever producing build output.
340
355
 
341
356
  **Telemetry**: Per-repo metrics in addition to per-task metrics:
342
357
  ```bash
343
- pipeline/scripts/log-metric.sh "$TASK_ID" 3 build.completed repo=common duration_ms=$D status=ok
344
- pipeline/scripts/log-metric.sh "$TASK_ID" 3 build.completed repo=uicomponents duration_ms=$D status=ok
358
+ $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 build.completed repo=common duration_ms=$D status=ok
359
+ $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 build.completed repo=uicomponents duration_ms=$D status=ok
345
360
  ```
346
361
 
347
362
  **Token forwarding:** every TDD round (red, green, refactor) that hits the dev model MUST forward token totals into the tracker so Phase 7's Cost Breakdown captures Phase 3:
348
363
 
349
364
  ```bash
350
- LOG_METRIC_FORWARD_TO_TRACKER=1 pipeline/scripts/log-metric.sh "$TASK_ID" 3 dev.tdd_round \
365
+ LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 dev.tdd_round \
351
366
  model=<sonnet|opus> step=<red|green|refactor> \
352
367
  tokens_in=$IN tokens_out=$OUT duration_ms=$DUR
353
368
  ```