@mmerterden/multi-agent-pipeline 13.5.1 → 14.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +232 -0
- package/README.md +3 -3
- package/docs/features.md +1 -1
- package/install/_common.mjs +73 -0
- package/install/_mcp-register.mjs +70 -31
- package/install/_plugin-skills.mjs +73 -14
- package/install/claude.mjs +28 -4
- package/install/codex.mjs +33 -2
- package/install/copilot.mjs +145 -9
- package/install/index.mjs +10 -6
- package/install/templates/copilot-instructions.md +1 -1
- package/package.json +1 -1
- package/pipeline/agents/code-reviewer.md +58 -1
- package/pipeline/commands/multi-agent/SKILL.md +7 -5
- package/pipeline/commands/multi-agent/analysis/SKILL.md +7 -7
- package/pipeline/commands/multi-agent/analysis-resolve/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/build-optimize/SKILL.md +7 -7
- package/pipeline/commands/multi-agent/channels/SKILL.md +5 -5
- package/pipeline/commands/multi-agent/dev/SKILL.md +23 -18
- package/pipeline/commands/multi-agent/dev-autopilot/SKILL.md +19 -13
- package/pipeline/commands/multi-agent/dev-local/SKILL.md +14 -12
- package/pipeline/commands/multi-agent/dev-local-autopilot/SKILL.md +17 -12
- package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/help/SKILL.md +4 -4
- package/pipeline/commands/multi-agent/ios-coding-standard/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +4 -4
- package/pipeline/commands/multi-agent/resume/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/review/SKILL.md +5 -5
- package/pipeline/commands/multi-agent/scan/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/search/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/setup/SKILL.md +6 -6
- package/pipeline/commands/multi-agent/{finish → ship}/SKILL.md +12 -12
- package/pipeline/commands/multi-agent/testflight-validation/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/update/SKILL.md +5 -2
- package/pipeline/commands/sim-test.md +2 -2
- package/pipeline/lib/credential-store-resolver.sh +16 -0
- package/pipeline/lib/credential-store.sh +47 -4
- package/pipeline/lib/fetch-figma-annotations.sh +26 -28
- package/pipeline/lib/figma-screenshot.sh +28 -39
- package/pipeline/lib/figma-token.sh +63 -0
- package/pipeline/multi-agent-refs/analysis-template.md +1 -1
- package/pipeline/multi-agent-refs/android-guide.md +1 -1
- package/pipeline/multi-agent-refs/channels/issue-comment.md +1 -1
- package/pipeline/multi-agent-refs/component-dispatch.md +2 -2
- package/pipeline/multi-agent-refs/cross-cli-contract.md +4 -4
- package/pipeline/multi-agent-refs/features/dev-critic.md +2 -2
- package/pipeline/multi-agent-refs/features/model-fallback.md +35 -2
- package/pipeline/multi-agent-refs/features/plan-todos.md +1 -1
- package/pipeline/multi-agent-refs/features/repo-map.md +1 -1
- package/pipeline/multi-agent-refs/features/review-multi-repo.md +3 -3
- package/pipeline/multi-agent-refs/features/shadow-git.md +1 -1
- package/pipeline/multi-agent-refs/features/skill-conformance.md +116 -0
- package/pipeline/multi-agent-refs/features/verify-by-test.md +1 -1
- package/pipeline/multi-agent-refs/generate-issue.md +1 -1
- package/pipeline/multi-agent-refs/multi-repo-integration-build.md +1 -1
- package/pipeline/multi-agent-refs/phases/log-format.md +4 -4
- package/pipeline/multi-agent-refs/phases/modes.md +7 -7
- package/pipeline/multi-agent-refs/phases/phase-0-init.md +13 -11
- package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +17 -15
- package/pipeline/multi-agent-refs/phases/phase-2-planning.md +7 -7
- package/pipeline/multi-agent-refs/phases/phase-3-dev.md +28 -13
- package/pipeline/multi-agent-refs/phases/phase-4-review.md +90 -58
- package/pipeline/multi-agent-refs/phases/phase-5-test.md +7 -7
- package/pipeline/multi-agent-refs/phases/phase-6-commit.md +8 -8
- package/pipeline/multi-agent-refs/phases/phase-7-report.md +8 -8
- package/pipeline/multi-agent-refs/phases.md +13 -13
- package/pipeline/multi-agent-refs/progress-contract.md +2 -2
- package/pipeline/multi-agent-refs/rules.md +7 -5
- package/pipeline/multi-agent-refs/swiftui-guide.md +1 -1
- package/pipeline/multi-agent-refs/tracker-contract.md +16 -15
- package/pipeline/preferences-template.json +7 -1
- package/pipeline/rules/figma-pipeline.md +2 -2
- package/pipeline/schemas/agent-state.schema.json +333 -79
- package/pipeline/schemas/criteria-manifest.schema.json +228 -0
- package/pipeline/schemas/migrations/prefs-2.4.0-to-2.5.0.mjs +64 -0
- package/pipeline/schemas/prefs.schema.json +118 -262
- package/pipeline/schemas/reviewer-output.schema.json +48 -3
- package/pipeline/schemas/token-budget.json +34 -10
- package/pipeline/schemas/triage-output.schema.json +112 -27
- package/pipeline/scripts/cost-table.json +7 -4
- package/pipeline/scripts/gc-worktrees.sh +1 -1
- package/pipeline/scripts/gen-mode-dispatch.mjs +6 -6
- package/pipeline/scripts/match-skills.mjs +37 -4
- package/pipeline/scripts/migrate-prefs.mjs +88 -17
- package/pipeline/scripts/phase-tracker.sh +14 -3
- package/pipeline/scripts/pre-commit-check.sh +49 -2
- package/pipeline/scripts/skill-conformance.mjs +960 -0
- package/pipeline/scripts/smoke-schema-validation.sh +17 -4
- package/pipeline/scripts/uninstall.mjs +35 -9
- package/pipeline/scripts/validate-reviewer.mjs +108 -1
- package/pipeline/skills/.skill-manifest.json +1 -1
- package/pipeline/skills/.skills-index.json +36 -9
- package/pipeline/skills/shared/README.md +15 -12
- package/pipeline/skills/shared/core/apple-archive-compliance/SKILL.md +1 -0
- package/pipeline/skills/shared/core/apple-archive-compliance/references/rules.yml +167 -0
- package/pipeline/skills/shared/core/google-play-compliance/SKILL.md +1 -0
- package/pipeline/skills/shared/core/google-play-compliance/references/rules.yml +184 -0
- package/pipeline/skills/shared/core/multi-agent/SKILL.md +10 -10
- package/pipeline/skills/shared/core/multi-agent-analysis/SKILL.md +4 -4
- package/pipeline/skills/shared/core/multi-agent-analysis-resolve/SKILL.md +3 -3
- package/pipeline/skills/shared/core/multi-agent-build-optimize/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-create-jira/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-dev/SKILL.md +6 -5
- package/pipeline/skills/shared/core/multi-agent-dev-autopilot/SKILL.md +7 -6
- package/pipeline/skills/shared/core/multi-agent-dev-local/SKILL.md +4 -3
- package/pipeline/skills/shared/core/multi-agent-dev-local-autopilot/SKILL.md +2 -1
- package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-ios-coding-standard/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +4 -4
- package/pipeline/skills/shared/core/multi-agent-review/SKILL.md +5 -5
- package/pipeline/skills/shared/core/multi-agent-scan/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-search/SKILL.md +1 -1
- package/pipeline/skills/shared/core/{multi-agent-finish → multi-agent-ship}/SKILL.md +8 -8
- package/pipeline/skills/shared/external/ios-coding-standard/SKILL.md +44 -5
- package/pipeline/skills/shared/external/ios-coding-standard/modules/_TEMPLATE.yml +82 -0
- package/pipeline/skills/shared/external/ios-coding-standard/references/STANDARD.md +169 -10
- package/pipeline/skills/shared/external/ios-coding-standard/references/rules.yml +335 -16
- package/pipeline/skills/skills-index.md +11 -8
|
@@ -16,14 +16,14 @@ xcodebuild ... 2>&1 | tee "$WORKTREE/.build.log"
|
|
|
16
16
|
# Gate 3: Tests pass (xcodebuild test/gradle test/pytest/npm test) - tee output to a log
|
|
17
17
|
<test-command> 2>&1 | tee "$WORKTREE/.test.log"
|
|
18
18
|
# Gate 4: Secrets - run the scanner against the staged diff
|
|
19
|
-
bash
|
|
19
|
+
bash $HOME/.claude/scripts/pre-commit-check.sh
|
|
20
20
|
```
|
|
21
21
|
|
|
22
22
|
**Default-FAIL evidence gate (required before recording any pass):** a green exit code is not enough - the captured log must actually show success. Before marking build/test passed, run the evidence gate against the tee'd log; it fails CLOSED when the log is missing, empty, or shows failure markers:
|
|
23
23
|
|
|
24
24
|
```bash
|
|
25
|
-
node
|
|
26
|
-
node
|
|
25
|
+
node $HOME/.claude/scripts/evidence-gate.mjs --claim build --status passed --evidence "$WORKTREE/.build.log" || BUILD_PASS=false
|
|
26
|
+
node $HOME/.claude/scripts/evidence-gate.mjs --claim test --status passed --evidence "$WORKTREE/.test.log" || TEST_PASS=false
|
|
27
27
|
```
|
|
28
28
|
|
|
29
29
|
This prevents a false "it built" claim with no log behind it. On exit 1, treat the gate as failed (do NOT proceed to AI review) and surface the gate's `reason`.
|
|
@@ -81,13 +81,7 @@ If changes include UI files (iOS: `*View.swift`, `*Screen.swift`, `*Cell.swift`;
|
|
|
81
81
|
- Missing safe area / keyboard avoidance → **important**
|
|
82
82
|
- Hardcoded colors instead of system/semantic colors → **suggestion**
|
|
83
83
|
|
|
84
|
-
**iOS - SwiftUI interaction & accessibility conventions
|
|
85
|
-
- A reusable component that **routes itself** (hardcoded `NavigationLink(destination: ConcreteScreen())`, calls a router) instead of emitting a typed `Output` → **important** (breaks reuse).
|
|
86
|
-
- A reusable component that **presents its own app-level overlay** (toast/alert/modal) instead of emitting an intent for the caller → **important** (double-fires when composed).
|
|
87
|
-
- `AnyView` routed through an overlay center / data-modal path (rich content should be a local `.sheet/.modal { }`) → **important**.
|
|
88
|
-
- Loading HUD not **ref-counted** / missing `end` on an error path / no min-show (flash) → **important**.
|
|
89
|
-
- A pinned bottom CTA built as a presented sheet instead of `.safeAreaInset(edge:.bottom)`, or detents that aren't ascending / use magic-number heights → **suggestion**.
|
|
90
|
-
- **Accessibility over-annotation** (redundant `.isButton` on `Button` / `.isStaticText` on `Text`, `.accessibilityLabel` duplicating visible text, a per-enum-variant a11y key for non-interactive visual state, hints on self-explanatory actions) → **suggestion**; **missing** decorative-image hiding or interactive-state `.accessibilityValue` → **important**.
|
|
84
|
+
**iOS - SwiftUI interaction & accessibility conventions.** Gated to changed SwiftUI files. These are rule-registry territory as of v14.0.0, not a list transcribed here: Step 1.78 resolves them from whichever registry declares SwiftUI scope, so the criteria and their severities live in one place instead of drifting between this doc and the skill. Reviewers receive the resolved rule IDs. Native-SwiftUI-first unless the project's `figma-config` `ui.*` declares a custom system, in which case check against that system. Reference skills, when no registry covers the change: `figma-navigation`, `figma-overlays`, `figma-bottom-sheets`, `figma-to-swiftui`.
|
|
91
85
|
|
|
92
86
|
**Android - Material Design compliance** (skills: `compose-components`, `android-architecture`):
|
|
93
87
|
- Non-Material3 component when M3 equivalent exists → **suggestion**
|
|
@@ -102,7 +96,7 @@ Device-level audit runs in Phase 5 (`audit-guide.md`).
|
|
|
102
96
|
|
|
103
97
|
#### Step 1.6 - Repo Map Injection (advisory, opt-in)
|
|
104
98
|
|
|
105
|
-
**Gated by `prefs.global.repoMap.enabled`** (default: `false`). Same pattern as Phase 1 Step 2.5 - runs
|
|
99
|
+
**Gated by `prefs.global.repoMap.enabled`** (default: `false`). Same pattern as Phase 1 Step 2.5 - runs `$HOME/.claude/scripts/repo-map.mjs` and injects the result as `${REPO_MAP}` into each reviewer's prompt context. Reviewers treat it the same way the priority files block is treated: advisory hint, not gospel.
|
|
106
100
|
|
|
107
101
|
Phase 4 reuses the cached map from Phase 1 when both phases run in the same task (avoid recomputing the same scan twice). The orchestrator caches under `state.repoMap.<sha>` keyed by `git rev-parse HEAD`; cache miss → regenerate. When `enabled=false`, both phases skip the script entirely and no `${REPO_MAP}` placeholder appears in prompts (no empty-string artefact).
|
|
108
102
|
|
|
@@ -115,16 +109,16 @@ Before dispatching reviewers, run the deterministic diff risk scorer and inject
|
|
|
115
109
|
Score the diff **once, in full** (no `--top`): Steps 1.76/1.77 need every scored file (a shrinking test file ranked 20th; whether ANY file is high-stakes). The top-N hint is derived from that report, not a second git walk.
|
|
116
110
|
|
|
117
111
|
```bash
|
|
118
|
-
RISK_FULL=$(node
|
|
112
|
+
RISK_FULL=$(node $HOME/.claude/scripts/diff-risk-score.mjs \
|
|
119
113
|
--base "$BASE_BRANCH" --head HEAD \
|
|
120
114
|
--task-id "$TASK_ID" 2>/dev/null)
|
|
121
|
-
echo "$RISK_FULL" | node
|
|
115
|
+
echo "$RISK_FULL" | node $HOME/.claude/scripts/validate-diff-risk.mjs - >/dev/null 2>&1 || RISK_FULL=""
|
|
122
116
|
|
|
123
117
|
# Priority hint for the reviewer prompts: top 5 of the full report.
|
|
124
118
|
RISK_JSON=$([ -n "$RISK_FULL" ] && jq -c '.files |= (sort_by(-.score) | .[:5])' <<< "$RISK_FULL" || echo "")
|
|
125
119
|
```
|
|
126
120
|
|
|
127
|
-
**Signals & weights** (see
|
|
121
|
+
**Signals & weights** (see `$HOME/.claude/schemas/diff-risk.schema.json`):
|
|
128
122
|
|
|
129
123
|
| Signal | Weight | Triggers when |
|
|
130
124
|
|---|---|---|
|
|
@@ -142,13 +136,13 @@ RISK_JSON=$([ -n "$RISK_FULL" ] && jq -c '.files |= (sort_by(-.score) | .[:5])'
|
|
|
142
136
|
**Gate behavior**: this step is **never blocking**. If risk scoring fails (git error, parse error, validator rejection), continue with no priority hint - reviewers receive the full diff in their default order. Failures are logged via metrics:
|
|
143
137
|
|
|
144
138
|
```bash
|
|
145
|
-
[ -z "$RISK_JSON" ] &&
|
|
139
|
+
[ -z "$RISK_JSON" ] && $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.diff_risk_skipped reason=$REASON
|
|
146
140
|
```
|
|
147
141
|
|
|
148
142
|
On success, emit a single summary metric:
|
|
149
143
|
|
|
150
144
|
```bash
|
|
151
|
-
|
|
145
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.diff_risk \
|
|
152
146
|
top_files=$(jq '.files | length' <<< "$RISK_JSON") \
|
|
153
147
|
max_score=$(jq '.totals.max_score' <<< "$RISK_JSON") \
|
|
154
148
|
loc_added=$(jq '.totals.loc_added' <<< "$RISK_JSON")
|
|
@@ -161,9 +155,9 @@ pipeline/scripts/log-metric.sh "$TASK_ID" 4 review.diff_risk \
|
|
|
161
155
|
Step 1.75 uses `test_lines_removed` as an advisory hint only - too weak for what it detects: a suite made green by deleting tests instead of fixing code. This turns the signal into blocking findings triage must adjudicate. Pure function of the full report, no git, no LLM.
|
|
162
156
|
|
|
163
157
|
```bash
|
|
164
|
-
TEST_INTEGRITY_JSON=$(printf '%s' "$RISK_FULL" | node
|
|
158
|
+
TEST_INTEGRITY_JSON=$(printf '%s' "$RISK_FULL" | node $HOME/.claude/scripts/test-integrity-gate.mjs 2>/dev/null || echo "")
|
|
165
159
|
TI_COUNT=$(jq -r '.count // 0' <<< "${TEST_INTEGRITY_JSON:-{\}}" 2>/dev/null || echo 0)
|
|
166
|
-
[ "$TI_COUNT" -gt 0 ] &&
|
|
160
|
+
[ "$TI_COUNT" -gt 0 ] && $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.test_integrity findings="$TI_COUNT"
|
|
167
161
|
```
|
|
168
162
|
|
|
169
163
|
`findings[]` are reviewer-shaped (`test_integrity`, `blocking`), so they merge into the reviewer findings at Step 3.0 and need no triage-prompt or `validate-triage.mjs` change. Triage keeps each blocking unless the removal is justified per the immutable-test rule (spec changed AND commit body names the test) → `deferred[]`.
|
|
@@ -175,18 +169,51 @@ The gate never blocks the phase; it *emits* blocking findings. Empty or unreadab
|
|
|
175
169
|
On a trivial diff every reviewer agrees and the extra models plus triage are paid for nothing. Reviewer count comes from the same report, no LLM.
|
|
176
170
|
|
|
177
171
|
```bash
|
|
178
|
-
SCOPE_JSON=$(printf '%s' "$RISK_FULL" | node
|
|
172
|
+
SCOPE_JSON=$(printf '%s' "$RISK_FULL" | node $HOME/.claude/scripts/review-scope.mjs 2>/dev/null \
|
|
179
173
|
|| echo '{"scope":"full","reason":"no risk report - failing safe"}')
|
|
180
174
|
REVIEW_SCOPE=$(jq -r '.scope // "full"' <<< "$SCOPE_JSON")
|
|
181
|
-
|
|
175
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.scope scope="$REVIEW_SCOPE"
|
|
182
176
|
```
|
|
183
177
|
|
|
184
178
|
`single` (Reviewer 1 only) requires **all** of: churn <= 20 lines, `totals.max_score` < 3.0, and no `security_path` / `migration` / `public_api` / `no_test_change` / `test_lines_removed` on any file. Anything else → `full`.
|
|
185
179
|
|
|
186
180
|
Fails safe in one direction only: empty report, parse failure or validator rejection all resolve to `full`, because skipping a reviewer trades coverage for cost. `consensus.reviewerCount` records what actually ran, so a single-reviewer run never reads as cross-model agreement. Opt-out: `prefs.global.reviewScopeGate = false` forces `full`.
|
|
187
181
|
|
|
182
|
+
#### Step 1.78 - Criteria resolution (skill conformance, required)
|
|
183
|
+
|
|
184
|
+
Resolves WHAT the changed code was supposed to honour, before any reviewer sees it. Zero LLM. Full contract: [`$HOME/.claude/multi-agent-refs/features/skill-conformance.md`]($HOME/.claude/multi-agent-refs/features/skill-conformance.md).
|
|
185
|
+
|
|
186
|
+
```bash
|
|
187
|
+
node $HOME/.claude/scripts/skill-conformance.mjs \
|
|
188
|
+
--diff "$WORKTREE/.review-diff.txt" --state "$WORKTREE/agent-state.json" \
|
|
189
|
+
--repo "$WORKTREE" --out "$WORKTREE/.pipeline/criteria-manifest.json"
|
|
190
|
+
CRIT_RC=$?
|
|
191
|
+
CRITERIA=$(cat "$WORKTREE/.pipeline/criteria-manifest.json" 2>/dev/null || echo '{}')
|
|
192
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.criteria \
|
|
193
|
+
rules="$(jq -r '.selectedRuleCount // 0' <<< "$CRITERIA")" \
|
|
194
|
+
ledger="$(jq -r '.ledger.source // "derived"' <<< "$CRITERIA")"
|
|
195
|
+
[ "$CRIT_RC" != "0" ] && HALT "criteria could not be resolved (rc=$CRIT_RC) - see resolutionFailure / unparseableRegistries"
|
|
196
|
+
```
|
|
197
|
+
|
|
198
|
+
Output conforms to `$HOME/.claude/schemas/criteria-manifest.schema.json`. The four parts that matter downstream:
|
|
199
|
+
|
|
200
|
+
- **`selectedRules[]` is the denominator.** Rule IDs from every registry whose declared `scope` matches the diff, persisted BEFORE the reviewers run so the set cannot be renegotiated after one has seen the diff. That is what makes "applied completely" answerable rather than "looks fine": a reviewer finding nothing must still return a verdict per ID. Discovery is declared via `standards-registry:` frontmatter, so no stack-specific skill is named here.
|
|
201
|
+
- **`coverage.declaredGaps[]` + `droppedReasons`.** An uncovered language, and every out-of-scope rule, is reported with a reason. "No rule applied", "every rule passed" and "the rule set was narrowed" must never render the same.
|
|
202
|
+
- **`ledger`.** `state.telemetry.skillCalls[]` corroborates only; `ledger.source` defaults to `derived` and coverage is never computed from self-report. A declared skill the resolver cannot bind is flagged.
|
|
203
|
+
- **`findings[]`.** Reviewer-shaped, merging at Step 3.0 alongside test-integrity: expired / unexplained / unknown-ID exception markers, plus any registry whose delegated linter is not wired here (those rules are unverified, so reporting no violations reports that nothing was measured).
|
|
204
|
+
|
|
205
|
+
**Any non-zero exit halts** (1 = setup error incl. a bad `--skills-root`; 2 = no root resolved, unparseable registry, or a declared path escaping its skill dir): continuing would drop a rule set from the denominator, and since the reviewer validator skips the checklist when zero rules were selected, the run would report clean over criteria never loaded. A coverage gap halts only under `prefs.global.skillConformance.blockOnCoverageGap`. No opt-out for the stage or the exception-expiry check, on the same grounds as Step 1.76.
|
|
206
|
+
|
|
188
207
|
#### Step 1.8 - Figma visual-fidelity context (when task carries a Figma reference)
|
|
189
208
|
|
|
209
|
+
**Dev-mode inputs.** Phases 1 and 2 never run in the `--dev` family, so three inputs this phase was written around are absent. Substitute them and RECORD the substitution - a step that could not run and a step that passed must not read the same, or the completeness claim cannot be checked:
|
|
210
|
+
|
|
211
|
+
| Absent input | Substitute |
|
|
212
|
+
|---|---|
|
|
213
|
+
| `detectedStack` (Phase 1 Step 2) | the language census in `criteria-manifest.json` → `languages`, from the diff's file extensions |
|
|
214
|
+
| Phase 1 analysis summary + Phase 2 plan (triage scope, cache prefix) | the task description plus the inline task list Phase 3 generated for itself |
|
|
215
|
+
| `state.evidence.figma[]` (this step), Step 2.8 visual conformance | nothing. Record each as `not-applicable (no Phase 1 evidence)` in `consensus.visualConformance`; never silently omit |
|
|
216
|
+
|
|
190
217
|
When `state.evidence.figma[]` is non-empty, the reviewer subagents MUST receive the captured screenshot URLs / paths and the canonical-component name (from `state.evidence.figma[i].screenshotUrl` and `state.evidence.figma[i].codeConnectSnippets[0].componentName` when present) so they can compare visual fidelity. Pass them inline in each reviewer prompt under a `## Figma evidence` block, one row per frame.
|
|
191
218
|
|
|
192
219
|
Visual-fidelity mismatches against the captured screenshot are BLOCKING findings, not nits:
|
|
@@ -208,7 +235,7 @@ When `state.figmaAccess.tier === 3` (user-attached screenshot, no Code Connect s
|
|
|
208
235
|
|
|
209
236
|
Phase 4 sends the same diff to every reviewer and then to triage, so the diff is the dominant token cost. Two measures keep it bounded:
|
|
210
237
|
|
|
211
|
-
**Shared cache prefix.** Build the reviewer and triage prompts so the large invariant context - the full diff, the Phase 1 analysis summary, the Phase 2 plan - is a byte-identical leading block across all dispatches in this iteration. Only the per-reviewer focus + skill line varies, and it goes AFTER the shared block. When the host supports prompt caching, the 2nd/3rd reviewer and the triage call then read that prefix at the discounted cache-read rate instead of re-billing it as fresh input. Forward the host-reported cache-read count as `tokens_cached` per the Token telemetry contract so the saving lands in the cost ledger.
|
|
238
|
+
**Shared cache prefix.** Build the reviewer and triage prompts so the large invariant context - the full diff, the `${CRITERIA}` block from Step 1.78, the Phase 1 analysis summary, the Phase 2 plan - is a byte-identical leading block across all dispatches in this iteration. Only the per-reviewer focus + skill line varies, and it goes AFTER the shared block. `${CRITERIA}` goes in the prefix, identical for every reviewer: subsetting it per reviewer would invalidate the prefix for the whole panel and re-bill the largest block in the phase. Per-reviewer emphasis stays a one-line pointer in the suffix. When the host supports prompt caching, the 2nd/3rd reviewer and the triage call then read that prefix at the discounted cache-read rate instead of re-billing it as fresh input. Forward the host-reported cache-read count as `tokens_cached` per the Token telemetry contract so the saving lands in the cost ledger.
|
|
212
239
|
|
|
213
240
|
**Single-repo diff cap.** If the diff exceeds the Phase 4 token allowance (`token-budget.json`), truncate the largest files and append a footer `[truncated - full diff in file://$WORKTREE/.review-diff.txt]`, writing the full diff to that path. Reviewers and triage receive the same capped view + the marker so they can flag "review the full diff manually." Log `review.diff_truncated bytes_dropped=<N>`. (Multi-repo already caps the combined diff at 80% of budget; this is the single-repo equivalent.)
|
|
214
241
|
|
|
@@ -220,32 +247,27 @@ Launch Agent instances **in parallel** using the shared `code-reviewer` subagent
|
|
|
220
247
|
|
|
221
248
|
| Reviewer | subagent_type | Claude Code | Copilot CLI | Codex CLI | Focus | Skills Referenced |
|
|
222
249
|
| ---------- | --------------- | --- | --- | --- | --- | --- |
|
|
223
|
-
| Reviewer 1 | `code-reviewer` | `claude-fable-5` | `claude-opus-
|
|
250
|
+
| Reviewer 1 | `code-reviewer` | `claude-fable-5` | `claude-opus-5` | `gpt-5.6` @ `xhigh` | Deep security + architecture | `api-security-best-practices`, `architecture` |
|
|
224
251
|
| Reviewer 2 | `code-reviewer` | (not dispatched) | `gpt-5.4` | `gpt-5.4` @ `high` | Edge cases, different perspective | cross-model diversity |
|
|
225
|
-
| Reviewer 3 | `code-reviewer` | `claude-sonnet-
|
|
226
|
-
| Triage | triage persona | `claude-fable-5` | `claude-opus-
|
|
252
|
+
| Reviewer 3 | `code-reviewer` | `claude-sonnet-5` | `claude-sonnet-5` | `gpt-5.6` @ `medium` | Quality + correctness + naming | `clean-code`, stack-specific skill |
|
|
253
|
+
| Triage | triage persona | `claude-fable-5` | `claude-opus-5` | `gpt-5.6` @ `max` | Filter false positives + out-of-scope | - |
|
|
227
254
|
|
|
228
255
|
Reviewer count per host: **Claude Code 2, Copilot CLI 3, Codex CLI 3**.
|
|
229
256
|
|
|
230
257
|
#### Codex CLI - two constraints that fail silently
|
|
231
258
|
|
|
232
|
-
Both were measured against Codex 0.145, not inferred
|
|
233
|
-
|
|
259
|
+
Both were measured against Codex 0.145, not inferred, and both produce a review that
|
|
260
|
+
looks like it ran. The full statement lives in the managed block at `~/.codex/AGENTS.md`
|
|
261
|
+
(always loaded on that host, so it is not restated here):
|
|
234
262
|
|
|
235
|
-
1.
|
|
236
|
-
`
|
|
237
|
-
|
|
238
|
-
|
|
239
|
-
the
|
|
240
|
-
and no error anywhere.
|
|
241
|
-
2. **Three concurrent children is the ceiling.** Codex advertises 4 concurrency
|
|
242
|
-
slots *including the orchestrator*, so a 3-reviewer panel saturates it and a 4th
|
|
243
|
-
child queues instead of running in parallel. The reviewer count on Codex is
|
|
244
|
-
capability-derived, not a preference.
|
|
263
|
+
1. **`fork_turns: "none"` on every `spawn_agent` that sets `model` or
|
|
264
|
+
`reasoning_effort`** - a full-history fork discards the override and collapses the
|
|
265
|
+
panel onto one model, silently.
|
|
266
|
+
2. **Three concurrent children is the ceiling** - 4 slots including the orchestrator,
|
|
267
|
+
so the reviewer count on Codex is capability-derived, not a preference.
|
|
245
268
|
|
|
246
|
-
|
|
247
|
-
|
|
248
|
-
authorization. Without it Phase 4 degrades to a single in-thread review.
|
|
269
|
+
Sub-agent delegation itself is authorized by that same managed block; without it Phase 4
|
|
270
|
+
degrades to a single in-thread review.
|
|
249
271
|
|
|
250
272
|
**Single-vendor caveat.** Every Codex reviewer is an OpenAI model, so the
|
|
251
273
|
cross-vendor disagreement that Claude Code and Copilot CLI get for free is absent.
|
|
@@ -257,7 +279,7 @@ in the triage note when all three agree on a borderline finding.
|
|
|
257
279
|
|
|
258
280
|
Each reviewer inherits the `code-reviewer` agent's focus areas (Security, Architecture, Quality, Performance) and output contract. The orchestrator overrides only the model and the stack-specific skill per-reviewer - no prompt duplication.
|
|
259
281
|
|
|
260
|
-
**Model override wiring:** `code-reviewer.md` declares `preferredModel: fable`, so Reviewer 1 uses the persona default (Fable 5). Reviewer 2 (Copilot-only, `gpt-5.4`) and Reviewer 3 (`claude-sonnet-
|
|
282
|
+
**Model override wiring:** `code-reviewer.md` declares `preferredModel: fable`, so Reviewer 1 uses the persona default (Fable 5). Reviewer 2 (Copilot-only, `gpt-5.4`) and Reviewer 3 (`claude-sonnet-5`) set `PHASE_MODEL_OVERRIDE=<model>` before dispatch - the orchestrator exports `CLAUDE_CODE_SUBAGENT_MODEL` on Claude Code, or passes `--model` on Copilot CLI. Full precedence rule: `skills/shared/core/multi-agent/SKILL.md#agent-dispatch--per-persona-model-routing-v610`. Fable dispatches are subject to the fallback contract (`$HOME/.claude/multi-agent-refs/features/model-fallback.md`): dispatch-error retry walks `fable -> opus -> sonnet` and budget-ceiling downgrade.
|
|
261
283
|
|
|
262
284
|
**Stack-specific skills loaded per reviewer** (from Phase 1 `detectedStack`). On Claude Code, Reviewer 2 (GPT-5.4) is not dispatched - its skill column is ignored. On Copilot CLI all three columns are used.
|
|
263
285
|
|
|
@@ -298,33 +320,38 @@ reading a diff cannot see spacing; something has to compare against the design.
|
|
|
298
320
|
Skip only when the diff has no UI change. Record the outcome in
|
|
299
321
|
`consensus.visualConformance` so Phase 7 reports whether it ran.
|
|
300
322
|
|
|
301
|
-
**iOS/Swift - interaction & convention checks (conditional).**
|
|
323
|
+
**iOS/Swift - interaction & convention checks (conditional).** Step 1.78 resolves these. Where no registry covers the change, reviewers fall back to the analysis doc (Section 14 Code Connect mapping) and, when `ai-ios-engineering-toolkit` is enabled, that plugin's navigation / overlay / bottom-sheet + accessibility conventions.
|
|
302
324
|
|
|
303
|
-
**Module review guides (conditional, all stacks).**
|
|
325
|
+
**Module review guides (conditional, all stacks).** Step 1.78 resolves them into `criteria-manifest.json` → `moduleGuides`. Inject with the directive: read each guide, apply its rules to the changed files under its directory - a guide governs only its own subtree, and its violations are findings triaged like any other. Same contract as `/multi-agent:review` Step 2b.
|
|
304
326
|
|
|
305
327
|
**Dispatch timeout (required, mirrors triage 3.3).** Reviewers run in parallel and triage waits on all of them, so one stalled reviewer hangs the phase. Bound each reviewer dispatch by `REVIEWER_TIMEOUT_SECONDS` (default 180). If a reviewer has not returned by the budget: log `review.reviewer_timeout reviewer=<name>`, treat that reviewer as absent, and proceed to triage with the reviewers that did return. The merged-findings count and `consensus.reviewerCount` reflect only the reviewers that returned. If **zero** reviewers return, retry Reviewer 1 once; on a second total failure HALT with `ERR: no reviewer returned within ${REVIEWER_TIMEOUT_SECONDS}s; resume with /multi-agent:resume #N.`. The Step 2.5 rebuttal round uses the same per-dispatch timeout. Never block indefinitely on a slow or dead reviewer dispatch.
|
|
306
328
|
|
|
307
329
|
#### Output contract - reviewer step
|
|
308
330
|
|
|
309
|
-
Step 2 produces N reviewer-output objects (one per dispatched reviewer), each conforming to
|
|
331
|
+
Step 2 produces N reviewer-output objects (one per dispatched reviewer), each conforming to `$HOME/.claude/schemas/reviewer-output.schema.json`. They are persisted to `state.reviewIterations[<iteration>].reviewers[]` and consumed by Step 3 (Fable triage) - never by Phase 6 directly. The triage step (below) is the producer of the only review artifact Phase 6 reads, conforming to `$HOME/.claude/schemas/triage-output.schema.json`.
|
|
310
332
|
|
|
311
|
-
**Subagent return format** - each reviewer returns JSON conforming to
|
|
333
|
+
**Subagent return format** - each reviewer returns JSON conforming to `$HOME/.claude/schemas/reviewer-output.schema.json`:
|
|
312
334
|
|
|
313
335
|
```json
|
|
314
|
-
{"findings":[{"severity":"blocking|important|suggestion","file":"...","line":N,"issue":"...","fix":"..."
|
|
336
|
+
{"findings":[{"severity":"blocking|important|suggestion","file":"...","line":N,"issue":"...","fix":"...","ruleId":"SEC-01","criteriaSource":"ios-coding-standard"}],
|
|
337
|
+
"conformance":[{"ruleId":"SEC-01","verdict":"conformant|violated|not-applicable","file":"...","line":N,"reason":"..."}],
|
|
338
|
+
"approved":true|false}
|
|
315
339
|
```
|
|
316
340
|
|
|
341
|
+
`ruleId` + `criteriaSource` appear on a finding that cites a rule from `${CRITERIA}`. `conformance` is required whenever Step 1.78 selected at least one rule, with exactly one row per selected ID and none outside the set.
|
|
342
|
+
|
|
317
343
|
**Required: validator gate (deterministic) - run immediately after each reviewer returns, before merging findings.** Persist each reviewer's output and validate the file - the validator's exit code decides, not the LLM turn:
|
|
318
344
|
|
|
319
345
|
```bash
|
|
320
346
|
REVIEWER_FILE="$WORKTREE/.pipeline/reviewer-$N.json"
|
|
321
347
|
printf '%s' "$REVIEWER_JSON" > "$REVIEWER_FILE"
|
|
322
|
-
node $HOME/.claude/scripts/validate-reviewer.mjs "$REVIEWER_FILE"
|
|
348
|
+
node $HOME/.claude/scripts/validate-reviewer.mjs "$REVIEWER_FILE" \
|
|
349
|
+
--criteria "$WORKTREE/.pipeline/criteria-manifest.json"
|
|
323
350
|
```
|
|
324
351
|
|
|
325
352
|
Progress line: ` → checking validator validate-reviewer ({reviewer})`
|
|
326
353
|
|
|
327
|
-
Exit 0 = valid. Exit 2 = contradiction (approved=true with blocking findings) - flip `approved` to `false`, continue. Exit 1 = malformed; gate protocol (fails CLOSED, same handling as the evidence gate): emit the validator stderr + `errors[]` verbatim, attempt ONE self-correction rework (re-invoke that reviewer with the errors quoted, overwrite the file), re-run the validator. If it fails again -> HALT the phase (no merge, no triage). Recovery hint: `ERR: reviewer output failed validate-reviewer.mjs twice. Inspect $REVIEWER_FILE against
|
|
354
|
+
Exit 0 = valid. Exit 2 = contradiction (approved=true with blocking findings) - flip `approved` to `false`, continue. With `--criteria`, exit 1 also covers the conformance checklist: a selected rule ID with no verdict, a verdict for an ID that was never selected, a `conformant` row with no file evidence, or a `violated` row with no matching finding. Those are the four ways a review can look complete without being complete, and the validator is what makes the checklist more than decoration - it is hand-written and does not apply `additionalProperties`, so an unchecked array would otherwise pass. Exit 1 = malformed; gate protocol (fails CLOSED, same handling as the evidence gate): emit the validator stderr + `errors[]` verbatim, attempt ONE self-correction rework (re-invoke that reviewer with the errors quoted, overwrite the file), re-run the validator. If it fails again -> HALT the phase (no merge, no triage). Recovery hint: `ERR: reviewer output failed validate-reviewer.mjs twice. Inspect $REVIEWER_FILE against $HOME/.claude/schemas/reviewer-output.schema.json, then resume with /multi-agent:resume #N.`
|
|
328
355
|
|
|
329
356
|
#### Step 2.5 - Disagreement-round loop (opt-in)
|
|
330
357
|
|
|
@@ -379,7 +406,7 @@ PRIOR_ART="["
|
|
|
379
406
|
for finding in $(jq -c '.findings[]' <<< "$MERGED_FINDINGS"); do
|
|
380
407
|
issue=$(jq -r '.issue' <<< "$finding")
|
|
381
408
|
file=$(jq -r '.file' <<< "$finding")
|
|
382
|
-
hits=$(node
|
|
409
|
+
hits=$(node $HOME/.claude/scripts/triage-memory.mjs query \
|
|
383
410
|
--issue "$issue" --file-glob "$(dirname "$file")/*" --top 3 2>/dev/null \
|
|
384
411
|
| jq -c '.hits // []')
|
|
385
412
|
PRIOR_ART="$PRIOR_ART$hits,"
|
|
@@ -394,7 +421,7 @@ The triage prompt MUST include a hedge: *"prior-art entries are context, not com
|
|
|
394
421
|
**Rejected-preference brief (on by default via `prefs.global.learningsLedger.injectIntoTriage`).** Inject the durable rejected-preference list so triage does not re-accept a suggestion the team already rejected on this repo:
|
|
395
422
|
|
|
396
423
|
```bash
|
|
397
|
-
node
|
|
424
|
+
node $HOME/.claude/scripts/learnings-ledger.mjs brief --max 20 2>/dev/null
|
|
398
425
|
```
|
|
399
426
|
|
|
400
427
|
Exit 2 (empty ledger) skips silently. The `## Rejected review preferences` section names patterns the team chose not to act on; triage should lean toward `rejected` for a finding that restates one - but the same hedge applies (context, not command; a genuinely new instance can still be accepted).
|
|
@@ -410,11 +437,16 @@ For each finding, decide:
|
|
|
410
437
|
- DEFERRED: real issue but out of current task scope → log for later, do not block
|
|
411
438
|
- REJECTED: false positive, duplicate, style-only nitpick, or already correct
|
|
412
439
|
|
|
440
|
+
A finding carrying a `ruleId` cites a written rule from a registry the project
|
|
441
|
+
adopted, so it is not a matter of taste: it can be DEFERRED or REJECTED as out of
|
|
442
|
+
scope for THIS task, but "I would have written it differently" is not available as
|
|
443
|
+
a reason. Preserve `ruleId` and `criteriaSource` on every finding you keep.
|
|
444
|
+
|
|
413
445
|
#### Output contract - triage step
|
|
414
446
|
|
|
415
|
-
Step 3 produces a single triage-output object conforming to
|
|
447
|
+
Step 3 produces a single triage-output object conforming to `$HOME/.claude/schemas/triage-output.schema.json` and persists it to `state.reviewIterations[<iteration>].triage`. This is the **only** Phase 4 artifact Phase 6 reads. Phase 6 commits MUST cite only `accepted` findings that were resolved; `deferred` items get linked in the PR description as follow-up work; `rejected` items never appear in any user-facing output.
|
|
416
448
|
|
|
417
|
-
Return ONLY valid JSON conforming to
|
|
449
|
+
Return ONLY valid JSON conforming to $HOME/.claude/schemas/triage-output.schema.json:
|
|
418
450
|
{
|
|
419
451
|
"accepted": [{ "severity": "blocking|important|suggestion", "file": "...", "line": N, "issue": "...", "fix": "...", "reviewer": "fable|opus|sonnet|gpt" }],
|
|
420
452
|
"deferred": [{ "finding": {...}, "reason": "..." }],
|
|
@@ -440,7 +472,7 @@ Progress line: ` → checking validator validate-triage`
|
|
|
440
472
|
| Exit | Meaning | Action |
|
|
441
473
|
| ----- | ---------------------------- | ------------------------------------------------- |
|
|
442
474
|
| **0** | Valid and clean | Act on triage output as-is |
|
|
443
|
-
| **1** | Invalid structure | Gate protocol (strict, fails CLOSED): (a) emit the validator's stderr and `errors[]` JSON into the log verbatim; (b) attempt ONE self-correction rework - re-prompt triage with the errors quoted, overwrite `$TRIAGE_FILE`; (c) re-run the validator. Second failure -> HALT the phase with the recovery hint: `ERR: triage output failed validate-triage.mjs twice. Inspect $TRIAGE_FILE against
|
|
475
|
+
| **1** | Invalid structure | Gate protocol (strict, fails CLOSED): (a) emit the validator's stderr and `errors[]` JSON into the log verbatim; (b) attempt ONE self-correction rework - re-prompt triage with the errors quoted, overwrite `$TRIAGE_FILE`; (c) re-run the validator. Second failure -> HALT the phase with the recovery hint: `ERR: triage output failed validate-triage.mjs twice. Inspect $TRIAGE_FILE against $HOME/.claude/schemas/triage-output.schema.json, fix or regenerate it, then resume with /multi-agent:resume #N.` |
|
|
444
476
|
| **2** | Over-rejection guard tripped | Pause for human confirm (autopilot: log + accept) |
|
|
445
477
|
| **3** | Contradiction auto-corrected | Proceed with `result.corrected` |
|
|
446
478
|
|
|
@@ -461,13 +493,13 @@ Failure fallback (timeout >120s, or agent crash before any JSON is produced): re
|
|
|
461
493
|
Emit metrics per review pass for Phase 7 cost rollup:
|
|
462
494
|
|
|
463
495
|
```bash
|
|
464
|
-
LOG_METRIC_FORWARD_TO_TRACKER=1
|
|
496
|
+
LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.reviewer_call model=fable duration_ms=$R1_DURATION tokens_in=$R1_IN tokens_out=$R1_OUT # model=opus on Copilot CLI
|
|
465
497
|
# GPT-5.4 metric emitted only on Copilot CLI (skip on Claude Code):
|
|
466
498
|
[ "${CLI_HOST:-claude}" = "copilot" ] && \
|
|
467
|
-
LOG_METRIC_FORWARD_TO_TRACKER=1
|
|
468
|
-
LOG_METRIC_FORWARD_TO_TRACKER=1
|
|
469
|
-
LOG_METRIC_FORWARD_TO_TRACKER=1
|
|
470
|
-
|
|
499
|
+
LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.reviewer_call model=gpt-5.4 duration_ms=$GPT_DURATION tokens_in=$GPT_IN tokens_out=$GPT_OUT
|
|
500
|
+
LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.reviewer_call model=sonnet duration_ms=$SONNET_DURATION tokens_in=$SONNET_IN tokens_out=$SONNET_OUT
|
|
501
|
+
LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.triage_call model=fable duration_ms=$TRIAGE_DURATION tokens_in=$TRIAGE_IN tokens_out=$TRIAGE_OUT
|
|
502
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.completed raw_count=$RAW accepted=$ACC deferred=$DEF rejected=$REJ approved=$APPROVED duration_ms=$DURATION
|
|
471
503
|
```
|
|
472
504
|
|
|
473
505
|
`LOG_METRIC_FORWARD_TO_TRACKER=1` mirrors `tokens_in`/`tokens_out`/`model` into `phase-tracker.sh` (see `$HOME/.claude/multi-agent-refs/progress-contract.md#token-telemetry-forwarding-v83`). On non-zero validator exit (1/2/3), also emit `triage.edge_case` with cause. Omit `tokens_in`/`tokens_out` if unavailable. Best-effort - never fails the pipeline.
|
|
@@ -515,7 +547,7 @@ Act **only on triage.accepted**:
|
|
|
515
547
|
At the end of every fix/rework round (each Phase 3 re-entry that resolved accepted findings, including the final one), append ONE one-line root-cause lesson per resolved blocking/important finding to the existing learnings ledger (`learnings-ledger.mjs` - the store Phase 1 and triage already replay; never invent a parallel store):
|
|
516
548
|
|
|
517
549
|
```bash
|
|
518
|
-
node
|
|
550
|
+
node $HOME/.claude/scripts/learnings-ledger.mjs add --kind fact \
|
|
519
551
|
--statement "<one-line rule/what-to-do so this class of finding does not recur (<=140 chars)>" \
|
|
520
552
|
--diagnosis "<the causal WHY the failure happened, one line>" \
|
|
521
553
|
--scope "<file-or-area glob from the finding>" --task "$TASK_ID"
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
### Phase 5: Test
|
|
2
2
|
|
|
3
|
-
> **TLDR** - Optional test gate. Offers to boot the simulator/emulator (UI Bug Hunter) or hand off to the user for manual QA.
|
|
3
|
+
> **TLDR** - Optional test gate. Offers to boot the simulator/emulator (UI Bug Hunter) or hand off to the user for manual QA. Needs an interactive prompt AND a worktree checkout, so it runs in `--dev` and the full pipeline but is dropped by every `autopilot` or `--local` variant. If issues found, loops back to Phase 3.
|
|
4
4
|
|
|
5
5
|
<!-- progress-contract: applied -->
|
|
6
6
|
Progress emission per `$HOME/.claude/multi-agent-refs/progress-contract.md` - lines for local-test prompt render, user-answer capture, repo checkout (if selected).
|
|
@@ -19,13 +19,13 @@ case "$STACK" in
|
|
|
19
19
|
*) SCAN_STACK="" ;;
|
|
20
20
|
esac
|
|
21
21
|
if [ -n "$SCAN_STACK" ] && [ "${prefs_testGap_enabled:-true}" = "true" ]; then
|
|
22
|
-
GAP_JSON=$(node
|
|
22
|
+
GAP_JSON=$(node $HOME/.claude/scripts/test-gap-scan.mjs \
|
|
23
23
|
--base "$BASE_BRANCH" --head HEAD --stack "$SCAN_STACK" 2>/dev/null)
|
|
24
|
-
echo "$GAP_JSON" | node
|
|
24
|
+
echo "$GAP_JSON" | node $HOME/.claude/scripts/validate-test-gap.mjs - >/dev/null 2>&1 || GAP_JSON=""
|
|
25
25
|
fi
|
|
26
26
|
```
|
|
27
27
|
|
|
28
|
-
**What the report contains** (per
|
|
28
|
+
**What the report contains** (per `$HOME/.claude/schemas/test-gap.schema.json`):
|
|
29
29
|
|
|
30
30
|
| Field | Meaning |
|
|
31
31
|
|---|---|
|
|
@@ -49,7 +49,7 @@ fi
|
|
|
49
49
|
**Telemetry**:
|
|
50
50
|
|
|
51
51
|
```bash
|
|
52
|
-
LOG_METRIC_FORWARD_TO_TRACKER=0
|
|
52
|
+
LOG_METRIC_FORWARD_TO_TRACKER=0 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 5 test_gap.scanned \
|
|
53
53
|
stack=$SCAN_STACK \
|
|
54
54
|
sources=$(jq '.totals.sourcesScanned' <<< "$GAP_JSON") \
|
|
55
55
|
gaps=$(jq '.totals.gapCount' <<< "$GAP_JSON")
|
|
@@ -100,7 +100,7 @@ Tier 1 / Tier 2 records print `screenshotUrl` from the captured evidence (Tier 2
|
|
|
100
100
|
bash $HOME/.claude/scripts/phase-tracker.sh now 5 "awaiting local test (user)"
|
|
101
101
|
bash $HOME/.claude/scripts/phase-tracker.sh render
|
|
102
102
|
```
|
|
103
|
-
The waiting state persists in `tracker-state.json` across the handoff; `/multi-agent:
|
|
103
|
+
The waiting state persists in `tracker-state.json` across the handoff; `/multi-agent:ship` and `/multi-agent:manual-test` CONTINUE this state file and never re-init it (`$HOME/.claude/multi-agent-refs/tracker-contract.md` "Continuation runs").
|
|
104
104
|
6. If fix needed:
|
|
105
105
|
- Branch already has WIP commit (from step 2) - changes are safe
|
|
106
106
|
- **Heal stale admin state first** (same contract as Phase 0 - step 3's
|
|
@@ -152,7 +152,7 @@ Returns severity-tagged findings (Critical / High / Medium). Critical items bloc
|
|
|
152
152
|
When the security-auditor or any other Phase 5 sub-agent runs, forward its token totals so Phase 7's Cost Breakdown captures Phase 5:
|
|
153
153
|
|
|
154
154
|
```bash
|
|
155
|
-
LOG_METRIC_FORWARD_TO_TRACKER=1
|
|
155
|
+
LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 5 audit.completed \
|
|
156
156
|
model=opus tokens_in=$IN tokens_out=$OUT duration_ms=$DUR
|
|
157
157
|
```
|
|
158
158
|
|
|
@@ -15,9 +15,9 @@ Progress emission per `$HOME/.claude/multi-agent-refs/progress-contract.md` -
|
|
|
15
15
|
|
|
16
16
|
#### Input contract
|
|
17
17
|
|
|
18
|
-
Phase 6 consumes the latest Phase 4 triage output object conforming to
|
|
18
|
+
Phase 6 consumes the latest Phase 4 triage output object conforming to `$HOME/.claude/schemas/triage-output.schema.json` (`accepted`, `deferred`, `rejected` buckets) plus the working-tree mutations from Phase 3. The commit body cites only `accepted` findings that were resolved; `deferred` items must be linked in the PR description as follow-up work, never silently dropped. If the triage output reports any unresolved blocking accepted finding, Phase 6 refuses to commit - the pipeline returns to Phase 3 for rework.
|
|
19
19
|
|
|
20
|
-
The "Build passes" checklist item below is evidence-gated, not self-asserted: Phase 6 trusts the `buildStatus.ok` that Phase 3 / Phase 4 recorded through
|
|
20
|
+
The "Build passes" checklist item below is evidence-gated, not self-asserted: Phase 6 trusts the `buildStatus.ok` that Phase 3 / Phase 4 recorded through `$HOME/.claude/scripts/evidence-gate.mjs` (a pass without a substantiating build/test log is treated as unverified and blocks the commit). The secret scanner (`pre-commit-check.sh`) runs on the staged diff as the final pre-commit gate.
|
|
21
21
|
|
|
22
22
|
#### Step 0 - Multi-Repo Integration Build
|
|
23
23
|
|
|
@@ -322,19 +322,19 @@ A task ending with one or more `pushStatus === "skipped"` repos does NOT bump `r
|
|
|
322
322
|
|
|
323
323
|
```bash
|
|
324
324
|
for proj in $(jq -r '.projects[].name' "$STATE_FILE"); do
|
|
325
|
-
|
|
326
|
-
|
|
327
|
-
|
|
325
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 6 commit.created repo=$proj sha=$SHA
|
|
326
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 6 push.attempted repo=$proj attempts=$N status=$STATUS
|
|
327
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 6 pr.opened repo=$proj url=$URL number=$N
|
|
328
328
|
done
|
|
329
|
-
|
|
329
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 6 multi_repo.completed repos=$REPOS skipped=$SKIPPED
|
|
330
330
|
```
|
|
331
331
|
|
|
332
332
|
**Token forwarding:** the commit-message and PR-body generators run on a model. Forward those calls into the tracker so Phase 7's Cost Breakdown captures Phase 6:
|
|
333
333
|
|
|
334
334
|
```bash
|
|
335
|
-
LOG_METRIC_FORWARD_TO_TRACKER=1
|
|
335
|
+
LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 6 commit.message_generated \
|
|
336
336
|
model=<sonnet|opus> tokens_in=$IN tokens_out=$OUT duration_ms=$DUR
|
|
337
|
-
LOG_METRIC_FORWARD_TO_TRACKER=1
|
|
337
|
+
LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 6 pr.body_generated \
|
|
338
338
|
model=<sonnet|opus> tokens_in=$IN tokens_out=$OUT duration_ms=$DUR
|
|
339
339
|
```
|
|
340
340
|
|
|
@@ -147,7 +147,7 @@ Write `agent-log.md` with ALL sections:
|
|
|
147
147
|
|
|
148
148
|
## Plan Todo Rollup (when `prefs.global.planTodos.enabled`)
|
|
149
149
|
|
|
150
|
-
Rendered via
|
|
150
|
+
Rendered via `$HOME/.claude/lib/plan-todos.sh list "$TASK_ID"`. Inline the resulting markdown checklist verbatim - one line per todo with status marker ([x] / [~] / [/] / [!]), ID, task text, and notes when present. Status summary one-liner ("4/6 done, 1 skipped, 1 failed") goes at the top from `plan-todos.sh status`.
|
|
151
151
|
|
|
152
152
|
Skipped sections: when `planTodos.enabled` is false or no `plan.todos[]` was emitted, the Plan Todo block is omitted (no empty header). Phase 3 task-by-task progress still appears in the Timeline section either way.
|
|
153
153
|
|
|
@@ -187,9 +187,9 @@ Skipped sections: when `planTodos.enabled` is false or no `plan.todos[]` was emi
|
|
|
187
187
|
**Telemetry emission** (mandatory): forward the phase's own LLM spend (humanizer + report compose calls) to the tracker, then emit the final event:
|
|
188
188
|
|
|
189
189
|
```bash
|
|
190
|
-
LOG_METRIC_FORWARD_TO_TRACKER=1
|
|
190
|
+
LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 7 report.compose \
|
|
191
191
|
model=$REPORT_MODEL tokens_in=$R_IN tokens_out=$R_OUT duration_ms=$R_DUR
|
|
192
|
-
|
|
192
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 7 task.completed \
|
|
193
193
|
phases=$PHASE_COUNT review_cycles=$CYCLES lang=$PROMPT_LANG \
|
|
194
194
|
channels_pr=$PR_STATUS channels_jira=$JIRA_STATUS \
|
|
195
195
|
channels_confluence=$CONF_STATUS channels_wiki=$WIKI_STATUS \
|
|
@@ -203,7 +203,7 @@ pipeline/scripts/log-metric.sh "$TASK_ID" 7 task.completed \
|
|
|
203
203
|
**Cost Breakdown emission (mandatory):** as part of agent-log compose, append the per-task cost block produced by `render-agent-log-cost.sh`. Best-effort - exit 2 (no tracker data) silently skipped:
|
|
204
204
|
|
|
205
205
|
```bash
|
|
206
|
-
COST_BLOCK=$(bash
|
|
206
|
+
COST_BLOCK=$(bash $HOME/.claude/scripts/render-agent-log-cost.sh "$TASK_ID" 2>/dev/null) && \
|
|
207
207
|
printf '\n%s\n' "$COST_BLOCK" >> "$AGENT_LOG"
|
|
208
208
|
```
|
|
209
209
|
|
|
@@ -214,7 +214,7 @@ This is independent of the channels-side `reportContent.costSummary` (which gate
|
|
|
214
214
|
```bash
|
|
215
215
|
TRIAGE_PATH="$WORKTREE/triage-output.json"
|
|
216
216
|
if [ -f "$TRIAGE_PATH" ]; then
|
|
217
|
-
node
|
|
217
|
+
node $HOME/.claude/scripts/triage-memory.mjs ingest \
|
|
218
218
|
--triage "$TRIAGE_PATH" \
|
|
219
219
|
--task-id "$TASK_ID" \
|
|
220
220
|
--task-title "$TASK_TITLE" \
|
|
@@ -228,11 +228,11 @@ Best-effort. The corpus is JSONL at `~/.claude/memory/multi-agent/<repo-slug>/tr
|
|
|
228
228
|
|
|
229
229
|
```bash
|
|
230
230
|
if [ -f "$TRIAGE_PATH" ]; then
|
|
231
|
-
node
|
|
231
|
+
node $HOME/.claude/scripts/learnings-ledger.mjs from-triage \
|
|
232
232
|
--triage "$TRIAGE_PATH" --task "$TASK_ID" >/dev/null 2>&1 || true
|
|
233
233
|
fi
|
|
234
234
|
# Optionally capture a durable architectural fact the analysis established:
|
|
235
|
-
# node
|
|
235
|
+
# node $HOME/.claude/scripts/learnings-ledger.mjs add --kind fact \
|
|
236
236
|
# --statement "<one-line fact>" --scope "<path-glob>" --task "$TASK_ID" >/dev/null 2>&1 || true
|
|
237
237
|
```
|
|
238
238
|
|
|
@@ -286,7 +286,7 @@ for entry in $(jq -c '.[]' memory-synthesis.json); do
|
|
|
286
286
|
body=$(jq -r '.body' <<< "$entry")
|
|
287
287
|
slug=$(jq -r '.slug' <<< "$entry")
|
|
288
288
|
type=$(jq -r '.type' <<< "$entry")
|
|
289
|
-
printf '%s\n' "$body" | bash
|
|
289
|
+
printf '%s\n' "$body" | bash $HOME/.claude/scripts/memory-save.sh "$PROJECT_ROOT" "$type" "$slug" -
|
|
290
290
|
done
|
|
291
291
|
```
|
|
292
292
|
|
|
@@ -20,7 +20,7 @@
|
|
|
20
20
|
|
|
21
21
|
```
|
|
22
22
|
Normal: 0-Init -> 1-Analysis -> 2-Planning -> 3-Dev -> 4-Review -> 5-Test -> 6-Commit -> 7-Report
|
|
23
|
-
--dev: 0-Init ------------------------------------------> 3-Dev
|
|
23
|
+
--dev: 0-Init ------------------------------------------> 3-Dev -> 4-Review -> 5-Test -> 6-Commit -> 7-Report
|
|
24
24
|
--local: Same as above but no worktree - works directly on local branch
|
|
25
25
|
```
|
|
26
26
|
|
|
@@ -42,9 +42,9 @@ Two channels run in parallel at every phase boundary. Both are required in their
|
|
|
42
42
|
Phase 0 MUST initialize the tracker and register all 8 phases:
|
|
43
43
|
|
|
44
44
|
```bash
|
|
45
|
-
|
|
45
|
+
$HOME/.claude/scripts/phase-tracker.sh init "$TASK_ID"
|
|
46
46
|
for p in 0:Init 1:Analysis 2:Planning 3:Dev 4:Review 5:Test 6:Commit 7:Report; do
|
|
47
|
-
|
|
47
|
+
$HOME/.claude/scripts/phase-tracker.sh add "${p%%:*}" "${p#*:}"
|
|
48
48
|
done
|
|
49
49
|
```
|
|
50
50
|
|
|
@@ -55,20 +55,20 @@ This produces an initial card stack printed by both CLIs.
|
|
|
55
55
|
As each phase enters/exits:
|
|
56
56
|
|
|
57
57
|
```bash
|
|
58
|
-
|
|
58
|
+
$HOME/.claude/scripts/phase-tracker.sh update <N> in_progress # phase starts
|
|
59
59
|
# ...do the work...
|
|
60
|
-
|
|
60
|
+
$HOME/.claude/scripts/phase-tracker.sh update <N> completed # phase ends OK
|
|
61
61
|
# or:
|
|
62
|
-
|
|
63
|
-
|
|
62
|
+
$HOME/.claude/scripts/phase-tracker.sh update <N> failed # phase failed
|
|
63
|
+
$HOME/.claude/scripts/phase-tracker.sh update <N> skipped # --dev mode etc.
|
|
64
64
|
```
|
|
65
65
|
|
|
66
66
|
For sub-phase progress (e.g. Phase 1's parallel Explore agents, Phase 4's reviewer dispatch + triage + validator gate, Phase 6's commit + push + PR sub-steps):
|
|
67
67
|
|
|
68
68
|
```bash
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
69
|
+
$HOME/.claude/scripts/phase-tracker.sh sub <N> 1 "<sub name>" pending # register
|
|
70
|
+
$HOME/.claude/scripts/phase-tracker.sh sub <N> 1 "<sub name>" in_progress # advance
|
|
71
|
+
$HOME/.claude/scripts/phase-tracker.sh sub <N> 1 "<sub name>" completed # done
|
|
72
72
|
```
|
|
73
73
|
|
|
74
74
|
State persists at `$HOME/.claude/logs/multi-agent/<task_id>/tracker-state.json` - atomic writes mean concurrent worktrees don't corrupt each other.
|
|
@@ -78,9 +78,9 @@ State persists at `$HOME/.claude/logs/multi-agent/<task_id>/tracker-state.json`
|
|
|
78
78
|
Call alongside tracker update for extra emphasis:
|
|
79
79
|
|
|
80
80
|
```bash
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
|
|
81
|
+
$HOME/.claude/scripts/phase-banner.sh start <N> "<name>" "<one-line detail>"
|
|
82
|
+
$HOME/.claude/scripts/phase-banner.sh end <N> done "<name>" "<short result>"
|
|
83
|
+
$HOME/.claude/scripts/phase-banner.sh sub <N> 2 "<sub>" "<detail>"
|
|
84
84
|
```
|
|
85
85
|
|
|
86
86
|
Banner status enum on `end`: `done` | `failed` | `skipped`. Anything else exits 64.
|
|
@@ -128,7 +128,7 @@ Overhead: ~40 bytes / line + one millisecond clock read. Negligible vs. any of t
|
|
|
128
128
|
Every phase that dispatches a billable LLM agent MUST forward the call's token totals to the tracker so the agent-log Cost Breakdown stays complete. The single canonical call shape is:
|
|
129
129
|
|
|
130
130
|
```bash
|
|
131
|
-
LOG_METRIC_FORWARD_TO_TRACKER=1
|
|
131
|
+
LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" <phase-id> <event> \
|
|
132
132
|
model=<fable|opus|sonnet|haiku|gpt-5.4|gpt-5.6|gpt-5.6-terra> \
|
|
133
133
|
tokens_in=$IN tokens_out=$OUT duration_ms=$DUR
|
|
134
134
|
```
|
|
@@ -142,7 +142,7 @@ This contract is enforced by `smoke-agent-log-cost.sh` (forwarder unit test) and
|
|
|
142
142
|
When `prefs.global.costBudget.enabled` is true, run the budget check once at the END of every phase, right after that phase's token totals have been forwarded to the tracker:
|
|
143
143
|
|
|
144
144
|
```bash
|
|
145
|
-
|
|
145
|
+
$HOME/.claude/scripts/cost-budget-check.mjs --task-id "$TASK_ID" \
|
|
146
146
|
--prefs "$HOME/.claude/multi-agent-preferences.json"
|
|
147
147
|
```
|
|
148
148
|
|