@mmerterden/multi-agent-pipeline 16.19.0 → 16.20.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (47) hide show
  1. package/CHANGELOG.md +23 -0
  2. package/install/_codex-agents.mjs +2 -2
  3. package/install/_dev-only-files.mjs +1 -0
  4. package/package.json +4 -3
  5. package/pipeline/agents/android-architect.md +1 -0
  6. package/pipeline/agents/backend-architect.md +1 -0
  7. package/pipeline/agents/code-reviewer.md +1 -0
  8. package/pipeline/agents/dev-critic.md +2 -1
  9. package/pipeline/agents/explorer.md +1 -0
  10. package/pipeline/agents/ios-architect.md +1 -0
  11. package/pipeline/agents/security-auditor.md +1 -0
  12. package/pipeline/agents/task-clarifier.md +1 -0
  13. package/pipeline/commands/multi-agent/manual-test/SKILL.md +5 -1
  14. package/pipeline/commands/multi-agent/resume/SKILL.md +1 -0
  15. package/pipeline/commands/multi-agent/store-ready/SKILL.md +22 -0
  16. package/pipeline/commands/sim-test.md +10 -5
  17. package/pipeline/multi-agent-refs/channels/pr.md +22 -4
  18. package/pipeline/multi-agent-refs/features/autopilot-circuit-breaker.md +15 -2
  19. package/pipeline/multi-agent-refs/features/review-delta.md +89 -0
  20. package/pipeline/multi-agent-refs/features/scope-check.md +41 -0
  21. package/pipeline/multi-agent-refs/features/verify-by-test.md +6 -5
  22. package/pipeline/multi-agent-refs/features/worktree-finalize.md +1 -1
  23. package/pipeline/multi-agent-refs/payload-contracts.md +1 -1
  24. package/pipeline/multi-agent-refs/phases/phase-3-dev.md +17 -2
  25. package/pipeline/multi-agent-refs/phases/phase-4-review.md +46 -8
  26. package/pipeline/multi-agent-refs/phases/phase-5-test.md +10 -0
  27. package/pipeline/multi-agent-refs/phases/phase-6-commit.md +2 -0
  28. package/pipeline/multi-agent-refs/phases/phase-7-report.md +4 -2
  29. package/pipeline/multi-agent-refs/rules.md +2 -2
  30. package/pipeline/schemas/agent-state.schema.json +129 -0
  31. package/pipeline/schemas/dev-critic-output.schema.json +5 -0
  32. package/pipeline/schemas/prefs.schema.json +47 -0
  33. package/pipeline/schemas/reviewer-output.schema.json +7 -2
  34. package/pipeline/schemas/scope-check.schema.json +55 -0
  35. package/pipeline/schemas/token-budget.json +3 -3
  36. package/pipeline/schemas/triage-output.schema.json +12 -2
  37. package/pipeline/scripts/README.md +3 -2
  38. package/pipeline/scripts/_fingerprint.mjs +173 -0
  39. package/pipeline/scripts/evidence-gate.mjs +73 -5
  40. package/pipeline/scripts/finding-fingerprint.mjs +101 -0
  41. package/pipeline/scripts/review-delta.mjs +217 -0
  42. package/pipeline/scripts/run-metrics.mjs +20 -0
  43. package/pipeline/scripts/scope-check-gate.mjs +90 -0
  44. package/pipeline/scripts/smoke-cross-cli-behavior.sh +11 -4
  45. package/pipeline/scripts/validate-reviewer.mjs +6 -0
  46. package/pipeline/scripts/validate-triage.mjs +20 -0
  47. package/pipeline/skills/shared/core/multi-agent-store-ready/SKILL.md +5 -0
@@ -84,7 +84,8 @@ When re-entering from Phase 4:
84
84
  3. `accepted.suggestion` items are applied opportunistically (no TDD loop required) unless the user asked for suggestions to be treated strictly.
85
85
  4. `deferred` items are NOT actioned in this re-entry - they surface in Phase 7's "Follow-up items" section.
86
86
  5. `rejected` items are never touched. Log their IDs + triage reasons for audit only.
87
- 6. After rework, increment `state.phases["3"].retryCount`; hard-kill at `retryCount === 3` and escalate to the user. Do not loop indefinitely.
87
+ 6. After rework, increment `state.phases["3"].retryCount`; hard-kill at `retryCount === 3` and escalate to the user. Do not loop indefinitely. When `prefs.global.autopilotCircuitBreaker.enabled` is true (default), record the hard-kill before escalating as `state.circuitBreaker = {tripped: true, trigger: 3, detail: "rework cycles exhausted", checkpoint: {phase: 3, step: "re-entry", iteration: <n>}, trippedAt, counters: {reworkCycles: <n>}}`: the rework-storm trigger; `autopilotCircuitBreaker.maxReworkCycles` (default 3) is bounded above by this cap.
88
+ 7. When `state.reviewIterations[-1].delta` exists, `stillPresent[]` findings come first in the task list, quoted `STILL PRESENT after round N-1's fix`; repeating the previous attempt is the loop the circuit-breaker stops.
88
89
 
89
90
  If the latest iteration has `triage.approved === true` AND `accepted === []`, Phase 3 was entered by mistake - log the anomaly and return to Phase 5.
90
91
 
@@ -155,6 +156,7 @@ For each task (respecting dependency order):
155
156
  **REFACTOR (if needed):**
156
157
  - Only if duplication or naming is poor
157
158
  - Re-run tests after refactor → still GREEN
159
+ - **Stability (required):** run every test added or changed in this diff `prefs.global.testStability.repeatCount` times (default 3; 1 disables) with the single-test invocation above. Disagreeing outcomes are not GREEN: log `test.flake_signal file=<f> passed=<k> of=<N>` and fix the test or the code first. A pass only on retry is a flake signal, not a pass. Same rule on the Phase 4 rework re-entry.
158
160
 
159
161
  **Target resolution** (auto-detect once per project, cache in `agent-state.json`; ios resolves scheme + simulator, android resolves module + variant, backend/web need none):
160
162
  ```bash
@@ -243,7 +245,7 @@ After the build/test green step and BEFORE Phase 4 handoff, run one diff-shrink
243
245
  - **Over-abstraction** - single-call-site protocols/wrappers/helpers from this diff
244
246
 
245
247
  Progress line: ` → dispatching code-simplifier diff-shrink`
246
- 2. **Return contract** - a list of shrink edits `{ "file", "lines", "edit", "rationale", "risk": "safe|unsafe" }`; the subagent never writes files.
248
+ 2. **Return contract** - a list of shrink edits `{ "file", "lines", "edit", "rationale", "risk": "safe|unsafe" }`; the subagent never writes files. Keep the `rationale` strings: Step 3.7 writes them into `scope-check.json`.
247
249
  3. **Apply safe edits only** - skip `unsafe` (logged). Zero edits is normal - log and continue. Progress line: ` → applying shrink edits ({N} applied, {M} skipped)`
248
250
  4. **Re-run build + tests** (same build-queue lock + evidence-gate rule as Step 4). Any breakage -> revert shrink edits wholesale and proceed pre-shrink; the simplifier must never cost a green state.
249
251
  5. **Record tokens in the cost ledger** so Phase 7's Cost Breakdown captures the pass:
@@ -258,6 +260,19 @@ Scope guard: a single pass, never looped. Runs in a Short run too; component tas
258
260
 
259
261
  ---
260
262
 
263
+ #### Step 3.7 - Scope self-check (required handoff artifact)
264
+
265
+ Phase 4 cannot reconstruct why each file was touched or what was left out on purpose, so Dev states both before the handoff: write `$WORKTREE/.pipeline/scope-check.json` (`schemas/scope-check.schema.json`: one `{path, reason}` per file in the diff, `notDone[]` with `why` in `out-of-scope | follow-up | rejected-abstraction`, and the Step 3.6 `simplifier` rationales), then run the gate. **Record rules, consumers and the full contract: `$HOME/.claude/multi-agent-refs/features/scope-check.md`.**
266
+
267
+ ```bash
268
+ git -C "$WORKTREE" diff --name-only "origin/$BASE_BRANCH"...HEAD \
269
+ | node $HOME/.claude/scripts/scope-check-gate.mjs --check "$WORKTREE/.pipeline/scope-check.json" --diff-files -
270
+ ```
271
+
272
+ Exit 1 lists `unjustified[]`: complete the record once and re-run. A second exit 1 does not block; log `dev.scope_check=incomplete unjustified=<n>` and hand off, and Phase 4 shows the gate output to reviewers inside `<scope-self-check>`. Progress line: ` → checking scope self-check ({n} files, {m} not done)`.
273
+
274
+ ---
275
+
261
276
  #### Short pipeline (`state.onlyDevelop === true`)
262
277
 
263
278
  Set by the Phase 0 Step 7.5 depth picker, or by autopilot never (autopilot always runs Full). When it is true, Phase 3 runs self-contained with **Opus** (not Sonnet). No Phase 2 plan exists - the agent creates its own scope.
@@ -152,6 +152,12 @@ echo "$RISK_FULL" | node $HOME/.claude/scripts/validate-diff-risk.mjs - >/dev/nu
152
152
  RISK_JSON=$([ -n "$RISK_FULL" ] && jq -c '.files |= (sort_by(-.score) | .[:5])' <<< "$RISK_FULL" || echo "")
153
153
  ```
154
154
 
155
+ Persist the totals as `state.diffRisk` (Phase 6 `risk` section, Phase 7, `run-metrics.mjs` read them):
156
+
157
+ ```bash
158
+ [ -n "$RISK_FULL" ] && jq -c '{diffRisk: (.totals + {signals: ([.files[].signals[]?.name] | unique)})}' <<< "$RISK_FULL" | node $HOME/.claude/scripts/write-state.mjs "$STATE_FILE"
159
+ ```
160
+
155
161
  **Signals & weights** (see `$HOME/.claude/schemas/diff-risk.schema.json`):
156
162
 
157
163
  | Signal | Weight | Triggers when |
@@ -179,7 +185,9 @@ On success, emit a single summary metric:
179
185
  $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.diff_risk \
180
186
  top_files=$(jq '.files | length' <<< "$RISK_JSON") \
181
187
  max_score=$(jq '.totals.max_score' <<< "$RISK_JSON") \
182
- loc_added=$(jq '.totals.loc_added' <<< "$RISK_JSON")
188
+ loc_added=$(jq '.totals.loc_added' <<< "$RISK_JSON") \
189
+ loc_removed=$(jq '.totals.loc_removed' <<< "$RISK_JSON") \
190
+ files=$(jq '.totals.files' <<< "$RISK_JSON")
183
191
  ```
184
192
 
185
193
  **Opt-out**: `prefs.global.diffRiskAdvisory = false` skips this step entirely (no script invocation, no priority block injection). Default `true` because the cost is bounded and the signal-to-noise has been measured against the golden-task fixture set.
@@ -269,7 +277,7 @@ When `state.figmaAccess.tier === 3` (user-attached screenshot, no Code Connect s
269
277
 
270
278
  Phase 4 sends the same diff to every reviewer and then to triage, so the diff is the dominant token cost. Two measures keep it bounded:
271
279
 
272
- **Shared cache prefix.** Build the reviewer and triage prompts so the large invariant context - the full diff, the `${CRITERIA}` block from Step 1.78, the Phase 1 analysis summary, the Phase 2 plan - is a byte-identical leading block across all dispatches in this iteration. Only the per-reviewer focus + skill line varies, and it goes AFTER the shared block. `${CRITERIA}` goes in the prefix, identical for every reviewer: subsetting it per reviewer would invalidate the prefix for the whole panel and re-bill the largest block in the phase. Per-reviewer emphasis stays a one-line pointer in the suffix. When the host supports prompt caching, the 2nd/3rd reviewer and the triage call then read that prefix at the discounted cache-read rate instead of re-billing it as fresh input. Forward the host-reported cache-read count as `tokens_cached` per the Token telemetry contract so the saving lands in the cost ledger.
280
+ **Shared cache prefix.** Build the reviewer and triage prompts so the large invariant context - the full diff, the `${CRITERIA}` block from Step 1.78, the Phase 1 analysis summary, the Phase 2 plan - is a byte-identical leading block across all dispatches in this iteration. Only the per-reviewer focus + skill line varies, and it goes AFTER the shared block. `${CRITERIA}` goes in the prefix, identical for every reviewer: subsetting it per reviewer would invalidate the prefix for the whole panel and re-bill the largest block in the phase. Per-reviewer emphasis stays a one-line pointer in the suffix. When the host supports prompt caching, the 2nd/3rd reviewer and the triage call then read that prefix at the discounted cache-read rate instead of re-billing it as fresh input. Forward the host-reported cache-read count as `tokens_cached` per the Token telemetry contract so the saving lands in the cost ledger. The `<scope-self-check>` block and, from iteration 2, the `<previous-round-findings>` block (Step 2.1) close the shared block, after the plan and before the per-reviewer suffix.
273
281
 
274
282
  **Single-repo diff cap.** If the diff exceeds the Phase 4 token allowance (`token-budget.json`), truncate the largest files and append a footer `[truncated - full diff in file://$WORKTREE/.review-diff.txt]`, writing the full diff to that path. Reviewers and triage receive the same capped view + the marker so they can flag "review the full diff manually." Log `review.diff_truncated bytes_dropped=<N>`. (Multi-repo already caps the combined diff at 80% of budget; this is the single-repo equivalent.)
275
283
 
@@ -328,7 +336,9 @@ Each reviewer inherits the `code-reviewer` agent's focus areas (Security, Archit
328
336
  | Docker | `ai-backend-toolkit:docker-expert` | `ai-backend-toolkit:docker-expert` | `ai-backend-toolkit:ci-cd-pipelines` |
329
337
  | Generic | `security-review` | `ai-backend-toolkit:clean-code` | `ai-backend-toolkit:clean-code` |
330
338
 
331
- Skills are injected into reviewer prompt context - the reviewer uses them as reference, not as commands.
339
+ ##### 2.1 Previous-round findings (iteration >= 2) and 2.2 scope self-check (every iteration)
340
+
341
+ A reviewer has no memory of the round before, so it rediscovers last round's findings in new words. From iteration 2, render the previous round's accepted blocking/important findings (`.pipeline/triage-round-$((ITERATION-1)).json`, max 40) into a `<previous-round-findings>` block at the end of the shared prefix: a still-present issue is reported with the SAME fingerprint and the current line, a fixed one is omitted, anything new leaves `fingerprint` unset. Every iteration also renders `.pipeline/scope-check.json` (Phase 3 Step 3.7) plus `scope-check-gate.mjs --advisory` output as `<scope-self-check>`: file reasons, unjustified files, and `notDone[]` (never re-raised as findings); a missing record logs `review.scope_check=missing`. Block text and recipes: `$HOME/.claude/multi-agent-refs/features/review-delta.md`.
332
342
 
333
343
  #### Step 2.8 - Visual conformance gate (component / screen work only)
334
344
 
@@ -382,11 +392,14 @@ Step 2 produces N reviewer-output objects (one per dispatched reviewer), each co
382
392
  REVIEWER_FILE="$WORKTREE/.pipeline/reviewer-$N.json"
383
393
  printf '%s' "$REVIEWER_JSON" > "$REVIEWER_FILE"
384
394
  node $HOME/.claude/scripts/validate-reviewer.mjs "$REVIEWER_FILE" \
385
- --criteria "$WORKTREE/.pipeline/criteria-manifest.json"
395
+ --criteria "$WORKTREE/.pipeline/criteria-manifest.json" \
396
+ && node $HOME/.claude/scripts/finding-fingerprint.mjs annotate --in-place "$REVIEWER_FILE"
386
397
  ```
387
398
 
388
399
  Progress line: ` → checking validator validate-reviewer ({reviewer})`
389
400
 
401
+ `finding-fingerprint.mjs` stamps each finding with its cross-round id once the validator passes; an echoed one is kept, and anonymization leaves it intact.
402
+
390
403
  Exit 0 = valid. Exit 2 = contradiction (approved=true with blocking findings) - flip `approved` to `false`, continue. With `--criteria`, exit 1 also covers the conformance checklist: a selected rule ID with no verdict, a verdict for an ID that was never selected, a `conformant` row with no file evidence, or a `violated` row with no matching finding. Those are the four ways a review can look complete without being complete, and the validator is what makes the checklist more than decoration - it is hand-written and does not apply `additionalProperties`, so an unchecked array would otherwise pass. Exit 1 = malformed; gate protocol (fails CLOSED, same handling as the evidence gate): emit the validator stderr + `errors[]` verbatim, attempt ONE self-correction rework (re-invoke that reviewer with the errors quoted, overwrite the file), re-run the validator. If it fails again -> HALT the phase (no merge, no triage). Recovery hint: `ERR: reviewer output failed validate-reviewer.mjs twice. Inspect $REVIEWER_FILE against $HOME/.claude/schemas/reviewer-output.schema.json, then resume with /multi-agent:resume #N.`
391
404
 
392
405
  #### Step 2.5 - Disagreement-round loop (opt-in)
@@ -499,6 +512,7 @@ bash $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 memory.hit rows=$CITED_COU
499
512
  ```
500
513
  You are the Review Triage agent. Three reviewers (fewer on a single-scope run) returned findings on this diff.
501
514
  Your job: separate signal from noise. Do NOT add new findings. Do NOT re-review code.
515
+ Preserve `fingerprint` verbatim on every finding you carry through; it is the finding's identity across rounds.
502
516
 
503
517
  For each finding, decide:
504
518
  - ACCEPTED: real issue, in scope, must be fixed now
@@ -529,14 +543,19 @@ Return ONLY valid JSON conforming to $HOME/.claude/schemas/triage-output.schema.
529
543
  Run on the persisted file immediately after the triage agent returns, before acting on the verdict; the validator's exit code decides, not the LLM turn:
530
544
 
531
545
  ```bash
532
- TRIAGE_FILE="$WORKTREE/.pipeline/triage.json"
546
+ ITERATION=$(jq '.reviewIterations | length' "$STATE_FILE")
547
+ TRIAGE_FILE="$WORKTREE/.pipeline/triage-round-$ITERATION.json"
533
548
  mkdir -p "$(dirname "$TRIAGE_FILE")"
534
549
  printf '%s' "$TRIAGE_JSON" > "$TRIAGE_FILE"
535
- node $HOME/.claude/scripts/validate-triage.mjs "$TRIAGE_FILE"
550
+ node $HOME/.claude/scripts/validate-triage.mjs "$TRIAGE_FILE" \
551
+ && node $HOME/.claude/scripts/finding-fingerprint.mjs annotate --in-place "$TRIAGE_FILE" \
552
+ && cp "$TRIAGE_FILE" "$WORKTREE/triage-output.json"
536
553
  ```
537
554
 
538
555
  Progress line: ` → checking validator validate-triage`
539
556
 
557
+ One file per round; `triage-output.json` is the latest copy every downstream reader expects (Phase 7, finalize, work-summary, diff-explain). Step 3.7 rewrites `$TRIAGE_FILE`; repeat the `cp` after it.
558
+
540
559
  | Exit | Meaning | Action |
541
560
  | ----- | ---------------------------- | ------------------------------------------------- |
542
561
  | **0** | Valid and clean | Act on triage output as-is |
@@ -607,7 +626,26 @@ After the triage verdict is computed, populate `triage.consensus`:
607
626
 
608
627
  A triage verdict is judgment; a failing repro test is proof. Runs only when `prefs.global.verifyByTest.enabled` is `true` AND `accepted` contains a `blocking` finding; otherwise skip silently. **Full contract (verdict table, cleanup invariant, prompts): `$HOME/.claude/multi-agent-refs/features/verify-by-test.md` - read it before executing this step.**
609
628
 
610
- Compressed flow: dispatch ONE verifier agent (model `verifyByTest.model`, default `sonnet`) for up to `maxFindings` (default 3) accepted blocking findings. Per finding it writes ONE minimal repro test and runs ONLY that test (Phase 3 single-test invocation, build lock, log tee'd to `$WORKTREE/.pipeline/verify-<i>.test.log`). Outcomes: test FAILS as predicted -> `confirmed`, finding stays blocking and the test is KEPT in `redTests[]` as the Phase 3 rework RED test; test PASSES -> `not-reproduced` ONLY if `evidence-gate.mjs --claim test --status passed` exits 0 on the log, finding moves to `deferred[]`, test deleted; compile error / timeout / not unit-testable -> `inconclusive`, judgment verdict stands. Stamp findings with `verification` (schema v3.2.0), persist `state.reviewIterations[-1].verifyByTest = {attempted, confirmed, downgraded, inconclusive, redTests[]}`, recompute `approved`, re-run `validate-triage.mjs` under the 3.2.1 gate. Whole step bounded by `stepTimeoutSec` (default 600); on breach or crash remaining findings keep judgment verdicts - never blocks. Telemetry per 3.4: `review.verify_by_test attempted= confirmed= downgraded= inconclusive= duration_ms=`.
629
+ Compressed flow: dispatch ONE verifier agent (model `verifyByTest.model`, default `sonnet`) for up to `maxFindings` (default 3) accepted blocking findings. Per finding it writes ONE minimal repro test and runs ONLY that test (Phase 3 single-test invocation, build lock, log tee'd to `$WORKTREE/.pipeline/verify-<i>.test.log`). Outcomes: test FAILS as predicted -> `confirmed`, finding stays blocking and the test is KEPT in `redTests[]` as the Phase 3 rework RED test; test PASSES on every one of `verifyByTest.repeatCount` runs (default 3) -> `not-reproduced` ONLY if `evidence-gate.mjs --claim test --status passed` exits 0 on each log; a run that disagrees with the others -> `inconclusive` with `flaky: passed k/N`, finding moves to `deferred[]`, test deleted; compile error / timeout / not unit-testable -> `inconclusive`, judgment verdict stands. Stamp findings with `verification` (schema v3.2.0), persist `state.reviewIterations[-1].verifyByTest = {attempted, confirmed, downgraded, inconclusive, redTests[]}`, recompute `approved`, re-run `validate-triage.mjs` under the 3.2.1 gate. Whole step bounded by `stepTimeoutSec` (default 600); on breach or crash remaining findings keep judgment verdicts - never blocks. Telemetry per 3.4: `review.verify_by_test attempted= confirmed= downgraded= inconclusive= duration_ms=`.
630
+
631
+ ##### 3.8 Cross-round delta + circuit-breaker trigger 2 (iteration >= 2)
632
+
633
+ **Full contract (state merge, telemetry line, picker wording): `$HOME/.claude/multi-agent-refs/features/review-delta.md`.**
634
+
635
+ ```bash
636
+ TRIP=$(jq -r '.global.autopilotCircuitBreaker.identicalFindingCycles // 2' "$PREFS_FILE")
637
+ DELTA_JSON=$(node $HOME/.claude/scripts/review-delta.mjs --rounds-dir "$WORKTREE/.pipeline" --iteration "$ITERATION" --trip-cycles "$TRIP"); DELTA_RC=$?
638
+ ```
639
+
640
+ Merge the JSON into `state.reviewIterations[-1].delta` via `write-state.mjs`; emit `review.delta iteration= new= still_present= resolved= plateau= tripped=`. Progress line: ` → comparing round {N} vs {N-1} (new={n} still={n} resolved={n})`.
641
+
642
+ | Exit | Action |
643
+ |---|---|
644
+ | **0** | Continue to Step 4 (also iteration 1 and a missing previous round). |
645
+ | **3** | Autopilot with `prefs.global.autopilotCircuitBreaker.enabled` (default true): write `state.circuitBreaker = {tripped: true, trigger: 2, detail, checkpoint: {phase: 4, step: "3.8", iteration: N}, trippedAt, counters}`, then the `operations.md` halt protocol with `haltReason="4:circuit-breaker:identical-finding"`. Interactive: show `delta.stillPresent`, ask `Continue rework` / `Escalate to me` / `Accept as deferred`. |
646
+ | **1** | Log `review.delta_skipped reason=invalid`, continue; the delta never blocks on its own failure. |
647
+
648
+ `plateau` is logged, not acted on; trigger 3 (the rework cap) is recorded by the Phase 3 re-entry.
611
649
 
612
650
  #### Step 4 - Consensus + Action (triage-driven)
613
651
 
@@ -615,7 +653,7 @@ If `triage.consensus.verdict` is `split` or `unverified`, surface `consensus.dis
615
653
 
616
654
  Act **only on triage.accepted**:
617
655
 
618
- - **accepted.blocking** → back to Phase 3 (max 3 iterations, with reflection prompt citing only accepted items). When Step 3.7 ran and `state.reviewIterations[-1].verifyByTest.redTests[]` is non-empty, the reflection prompt cites each red test: "a failing repro test already exists at <testRef>; make it green; do not delete or weaken it."
656
+ - **accepted.blocking** → back to Phase 3 (max 3 iterations, with reflection prompt citing only accepted items). The reflection prompt names findings by `fingerprint` and quotes `delta.stillPresent` first, marked `STILL PRESENT after round N-1's fix`. When Step 3.7 ran and `state.reviewIterations[-1].verifyByTest.redTests[]` is non-empty, the reflection prompt cites each red test: "a failing repro test already exists at <testRef>; make it green; do not delete or weaken it."
619
657
  - **accepted.important** → fix and re-review
620
658
  - **accepted.suggestion** → apply if reasonable
621
659
  - **deferred** → append to Phase 7 report as "follow-up items" (do not block)
@@ -103,6 +103,16 @@ Tier 1 / Tier 2 records print `screenshotUrl` from the captured evidence (Tier 2
103
103
  bash $HOME/.claude/scripts/phase-tracker.sh render
104
104
  ```
105
105
  The waiting state persists in `tracker-state.json` across the handoff; `/multi-agent:resume-local` and `/multi-agent:manual-test` CONTINUE this state file and never re-init it (`$HOME/.claude/multi-agent-refs/tracker-contract.md` "Continuation runs").
106
+
107
+ **"ok" is a structured result, not a word.** Before "ok" is accepted, the run writes `$WORKTREE/.pipeline/manual-test.json`: one entry per acceptance criterion, the criteria taken from the analysis doc test plan (Section 15 / 20), the plan tasks, and the user's own words in the reply. Every criterion records what was seen; a criterion that was not tried says so with a reason.
108
+ ```json
109
+ {"criteria":[{"spec":"<quote>","source":"analysis 15.2 | plan task 3 | user","observed":"<what was seen>","verdict":"pass|fail|not-tested","reason":"<required when not-tested>","screenshot":"<path or null>"}],"verdict":"passed|failed"}
110
+ ```
111
+ Then gate it:
112
+ ```bash
113
+ node $HOME/.claude/scripts/evidence-gate.mjs --claim manual --status passed --evidence "$WORKTREE/.pipeline/manual-test.json"
114
+ ```
115
+ Exit 1 means the "ok" is not accepted: tell the user which criterion is missing evidence (a `fail` verdict, or `not-tested` without a reason) and wait for the next reply. Exit 0 marks Phase 5 completed with `Result "local test passed (user)"`. The "fix: ..." path below is unchanged.
106
116
  6. If fix needed:
107
117
  - Branch already has WIP commit (from step 2) - changes are safe
108
118
  - **Heal stale admin state first** (same contract as Phase 0 - step 3's
@@ -136,6 +136,8 @@ Branch **deterministically**, no implicit fallback. Read `agent-state.json` and
136
136
 
137
137
  Generate a structured PR description based on task type. The PR body targets **code reviewers** - it should be technical: what changed, why, architecture decisions, how to verify.
138
138
 
139
+ Two inputs are read from state and the worktree before writing, not recalled from the conversation: `$WORKTREE/.pipeline/scope-check.json` (Phase 3 Step 3.7) supplies the `## Changes` bullets from `files[].reason` and the "Follow-ups not done in this PR" list under `## Related` from `notDone[]`; `state.diffRisk.signals` (Phase 4 Step 1.75) decides whether the conditional `## Risk and Security` section is required. When a high-stakes signal is present and the section is missing, this step blocks until it is written; a placeholder answer ("TBD") counts as missing.
140
+
139
141
  **required**: Run all generated text (PR body, commit message) through the `humanizer` skill before posting. This removes AI-generated patterns (inflated language, filler phrases, repetitive structure) and makes the output sound like a developer wrote it.
140
142
 
141
143
  **IMPORTANT - Issue auto-close prevention**: Never use these keywords in PR title, body, or commit messages - they auto-close the issue on merge:
@@ -163,8 +163,10 @@ Skipped sections: when `planTodos.enabled` is false or no `plan.todos[]` was emi
163
163
 
164
164
  ## Review Iterations
165
165
 
166
- | Iteration | Blocking | Important | Suggestion | Decision |
167
- | --------- | -------- | --------- | ---------- | -------- |
166
+ | Iteration | Blocking | Important | Suggestion | Decision | Still present | Resolved | New |
167
+ | --------- | -------- | --------- | ---------- | -------- | ------------- | -------- | --- |
168
+
169
+ (the last three columns come from `reviewIterations[i].delta.counts` (Phase 4 Step 3.8) and read `-` on iteration 1 and on runs that predate the delta; a tripped circuit-breaker is one extra line under the table: `Circuit-breaker tripped: trigger {n}, {detail}`)
168
170
 
169
171
  ## Files Changed
170
172
 
@@ -52,10 +52,10 @@ This is the single source of truth. When a contributor or model is unsure where
52
52
  - **NEVER** commit without passing build (all gates in Phase 4 Step 1 must be green).
53
53
  - **NEVER** commit without passing review (at least one AI reviewer must return `approved: true` with no blocking findings).
54
54
  - **NEVER** skip tests. Every public method, every error path, every edge case.
55
- - **NEVER** delete, rename, or weaken an existing test to get a green run. Existing tests are immutable during a task. A test may change only when the task itself changes the spec that test encodes, and the commit body must name the changed test and the spec change. Deterministic backstop: the `test_lines_removed` diff-risk signal (Phase 4 Step 1.75) flags test files that shrink.
55
+ - **NEVER** delete, rename, or weaken an existing test to get a green run. Existing tests are immutable during a task; one may change only when the task changes the spec it encodes, and the commit body names both. Deterministic backstop: the `test_lines_removed` diff-risk signal (Phase 4 Step 1.75) flags test files that shrink. A pass only on retry is a flake signal, not a pass: Phase 3 repeats new and changed tests (`testStability.repeatCount`, default 3) and logs `test.flake_signal` when runs disagree.
56
56
  - **Follow existing code style and conventions.** Read neighbor files before writing new ones - match naming, structure, import order.
57
57
  - **Use design tokens, no magic numbers.** `16` → `.Spacing.spacing16`. `#E31837` → `Color.Primary.primary`. `.font(.system(size: 14))` → `.typographyStyle(.body1)`.
58
- - **Design system primitives before custom views.** Before writing a new SwiftUI / Compose / React / View / Configuration triplet inside a domain or feature module, grep the shared component library (project-specific path, e.g. `Common/UIComponents/`, `core-ui/`, `packages/ui/`) for an existing primitive that solves the same problem. New domain-level wrappers, custom modals, custom buttons, or hand-rolled toasts are forbidden when the design system already has an equivalent. If the primitive exists but lacks a modifier (placeholder, size, error binding), **add the modifier to the primitive** in its `+Modifiers` extension - do not fork the primitive into the consumer domain. The Figma `CodeConnectSnippet` is the authoritative pointer to which primitive to use.
58
+ - **Design system primitives before custom views.** Before writing a new View / Configuration triplet in a feature module, grep the shared component library (e.g. `Common/UIComponents/`, `core-ui/`, `packages/ui/`) for a primitive that already solves it. Domain-level wrappers, custom modals, buttons or toasts are forbidden when the design system has an equivalent. If the primitive exists but lacks a modifier (placeholder, size, error binding), **add the modifier to the primitive** in its `+Modifiers` extension; never fork it into the consumer domain. The Figma `CodeConnectSnippet` is the authoritative pointer to which primitive to use.
59
59
 
60
60
  ## Swift-Specific Rules
61
61
 
@@ -650,6 +650,90 @@
650
650
  "type": "string",
651
651
  "enum": ["fix", "accept", "escalate"]
652
652
  },
653
+ "delta": {
654
+ "type": "object",
655
+ "additionalProperties": true,
656
+ "description": "Written by Phase 4 Step 3.8 from review-delta.mjs: how this round's accepted findings relate to the previous round's rework mandate. Absent on iteration 1 and on runs from before the delta existed.",
657
+ "properties": {
658
+ "previousIteration": { "type": "integer", "minimum": 1 },
659
+ "new": {
660
+ "type": "array",
661
+ "items": {
662
+ "type": "object",
663
+ "additionalProperties": true,
664
+ "properties": {
665
+ "fingerprint": { "type": "string" },
666
+ "severity": { "type": ["string", "null"] },
667
+ "file": { "type": ["string", "null"] },
668
+ "issue": { "type": ["string", "null"] }
669
+ }
670
+ }
671
+ },
672
+ "stillPresent": {
673
+ "type": "array",
674
+ "items": {
675
+ "type": "object",
676
+ "additionalProperties": true,
677
+ "properties": {
678
+ "fingerprint": { "type": "string" },
679
+ "severity": { "type": ["string", "null"] },
680
+ "file": { "type": ["string", "null"] },
681
+ "issue": { "type": ["string", "null"] }
682
+ }
683
+ }
684
+ },
685
+ "resolved": {
686
+ "type": "array",
687
+ "items": {
688
+ "type": "object",
689
+ "additionalProperties": true,
690
+ "properties": {
691
+ "fingerprint": { "type": "string" },
692
+ "severity": { "type": ["string", "null"] },
693
+ "file": { "type": ["string", "null"] },
694
+ "issue": { "type": ["string", "null"] }
695
+ }
696
+ }
697
+ },
698
+ "downgraded": {
699
+ "type": "array",
700
+ "items": {
701
+ "type": "object",
702
+ "additionalProperties": true,
703
+ "properties": {
704
+ "fingerprint": { "type": "string" },
705
+ "severity": { "type": ["string", "null"] },
706
+ "file": { "type": ["string", "null"] },
707
+ "issue": { "type": ["string", "null"] }
708
+ }
709
+ }
710
+ },
711
+ "stillPresentBlocking": {
712
+ "type": "array",
713
+ "items": {
714
+ "type": "object",
715
+ "additionalProperties": true,
716
+ "properties": {
717
+ "fingerprint": { "type": "string" },
718
+ "severity": { "type": ["string", "null"] },
719
+ "file": { "type": ["string", "null"] },
720
+ "issue": { "type": ["string", "null"] }
721
+ }
722
+ }
723
+ },
724
+ "recurrence": {
725
+ "type": "object",
726
+ "additionalProperties": { "type": "integer", "minimum": 1 },
727
+ "description": "fingerprint -> consecutive rework cycles the finding has survived."
728
+ },
729
+ "plateau": {
730
+ "type": "boolean",
731
+ "description": "The still-present set is unchanged from the previous delta and non-empty."
732
+ },
733
+ "tripped": { "type": "boolean" },
734
+ "computedAt": { "type": "string", "format": "date-time" }
735
+ }
736
+ },
653
737
  "reviewers": {
654
738
  "type": "array",
655
739
  "description": "One entry per reviewer dispatch that RETURNED. Typed because two consumers depend on the shape: anonymize-findings.mjs needs model+findings to build the label map, and run-metrics.mjs reports acceptedRatio per reviewer. Extra keys are allowed; nothing is required, so a run written before this shape existed still validates and surfaces as model \"unknown\" rather than failing.",
@@ -699,6 +783,51 @@
699
783
  }
700
784
  }
701
785
  },
786
+ "circuitBreaker": {
787
+ "type": "object",
788
+ "additionalProperties": false,
789
+ "description": "Autopilot circuit-breaker record (refs/features/autopilot-circuit-breaker.md). Written only when a trigger trips: trigger 2 by Phase 4 Step 3.8 (a mandate finding survived identicalFindingCycles rework cycles), trigger 3 by the Phase 3 re-entry hard-kill. /multi-agent:resume clears tripped and keeps counters.",
790
+ "required": ["tripped"],
791
+ "properties": {
792
+ "tripped": { "type": "boolean" },
793
+ "trigger": { "type": ["integer", "null"], "minimum": 1, "maximum": 5 },
794
+ "detail": { "type": "string" },
795
+ "checkpoint": {
796
+ "type": "object",
797
+ "additionalProperties": false,
798
+ "properties": {
799
+ "phase": { "type": "integer", "minimum": 0, "maximum": 7 },
800
+ "step": { "type": "string" },
801
+ "iteration": { "type": "integer", "minimum": 1 }
802
+ }
803
+ },
804
+ "trippedAt": { "type": "string", "format": "date-time" },
805
+ "counters": {
806
+ "type": "object",
807
+ "additionalProperties": false,
808
+ "properties": {
809
+ "identicalFindingCycles": { "type": "integer", "minimum": 0 },
810
+ "reworkCycles": { "type": "integer", "minimum": 0 }
811
+ }
812
+ }
813
+ }
814
+ },
815
+ "diffRisk": {
816
+ "type": "object",
817
+ "additionalProperties": true,
818
+ "description": "Totals from diff-risk-score.mjs, persisted by Phase 4 Step 1.75 so Phase 6 (PR risk section), Phase 7 and run-metrics.mjs read the same numbers the review scope decision used.",
819
+ "properties": {
820
+ "files": { "type": "integer", "minimum": 0 },
821
+ "loc_added": { "type": "integer", "minimum": 0 },
822
+ "loc_removed": { "type": "integer", "minimum": 0 },
823
+ "max_score": { "type": "number" },
824
+ "signals": {
825
+ "type": "array",
826
+ "items": { "type": "string" },
827
+ "description": "Distinct signal names seen across files (security_path, migration, public_api, no_test_change, test_lines_removed, ...)."
828
+ }
829
+ }
830
+ },
702
831
  "confluenceSpace": {
703
832
  "type": ["string", "null"],
704
833
  "description": "Cached Confluence space key for the project - avoids re-asking on every run."
@@ -100,6 +100,11 @@
100
100
  "description": "Citation. Format: 'rules/<file>.md#<anchor>' or 'gate/<name>'."
101
101
  },
102
102
  "issue": { "type": "string", "description": "Short, one-sentence description." },
103
+ "fingerprint": {
104
+ "type": "string",
105
+ "pattern": "^F:[0-9a-f]{8}$",
106
+ "description": "Stable id computed by finding-fingerprint.mjs --kind dev-critic. Round-2 findings carry the round-1 fingerprint; a round-2 finding with none is the scope creep the loop contract forbids."
107
+ },
103
108
  "fix": {
104
109
  "type": "string",
105
110
  "description": "Concrete suggestion the generator can act on without re-reading the rule."
@@ -927,6 +927,53 @@
927
927
  "maximum": 1800,
928
928
  "default": 600,
929
929
  "description": "Wall-clock budget for the whole Step 3.7 pass. On breach, remaining findings keep judgment-only verdicts and the pipeline proceeds (never blocks)."
930
+ },
931
+ "repeatCount": {
932
+ "type": "integer",
933
+ "minimum": 1,
934
+ "maximum": 10,
935
+ "default": 3,
936
+ "description": "v16.20+ - how many times a repro test must pass before a finding is downgraded to not-reproduced. Any red run makes the verdict inconclusive with a flaky note. 1 disables the repeat."
937
+ }
938
+ }
939
+ },
940
+ "autopilotCircuitBreaker": {
941
+ "type": "object",
942
+ "additionalProperties": false,
943
+ "description": "v16.20+ - autopilot circuit-breaker (refs/features/autopilot-circuit-breaker.md). Trigger 2 (a finding that survives consecutive rework cycles) and trigger 3 (the rework-cycle cap) are evaluated in code; the other triggers stay documented behaviour. When a trigger trips in autopilot the run halts visibly with state.circuitBreaker set; interactive modes show the still-present list and ask instead.",
944
+ "properties": {
945
+ "enabled": {
946
+ "type": "boolean",
947
+ "default": true,
948
+ "description": "Master switch for the code-evaluated triggers."
949
+ },
950
+ "identicalFindingCycles": {
951
+ "type": "integer",
952
+ "minimum": 1,
953
+ "maximum": 5,
954
+ "default": 2,
955
+ "description": "Consecutive rework cycles a blocking/important finding may survive before trigger 2 trips (review-delta.mjs --trip-cycles)."
956
+ },
957
+ "maxReworkCycles": {
958
+ "type": "integer",
959
+ "minimum": 1,
960
+ "maximum": 3,
961
+ "default": 3,
962
+ "description": "Phase 4 -> Phase 3 rework cycles before trigger 3 trips. Bounded above by the Phase 3 retryCount hard-kill."
963
+ }
964
+ }
965
+ },
966
+ "testStability": {
967
+ "type": "object",
968
+ "additionalProperties": false,
969
+ "description": "v16.20+ - a test that passes only on retry is a flake signal, not a pass. Phase 3 runs new and changed tests repeatCount times before GREEN; disagreeing outcomes block and are recorded as test.flake_signal.",
970
+ "properties": {
971
+ "repeatCount": {
972
+ "type": "integer",
973
+ "minimum": 1,
974
+ "maximum": 10,
975
+ "default": 3,
976
+ "description": "Runs per new or changed test. 1 disables the repeat."
930
977
  }
931
978
  }
932
979
  },
@@ -1,9 +1,9 @@
1
1
  {
2
2
  "$schema": "https://json-schema.org/draft/2020-12/schema",
3
3
  "$id": "https://github.com/mmerterden/multi-agent-pipeline/pipeline/schemas/reviewer-output.schema.json",
4
- "version": "1.1.0",
4
+ "version": "1.2.0",
5
5
  "title": "Multi-Agent Pipeline - Phase 4 reviewer output",
6
- "description": "Contract for a single code-reviewer subagent's JSON output in Phase 4 Step 2. Every host dispatches 3 parallel reviewers; the middle slot is CLI-aware: Claude Code (Fable, Opus, Sonnet); Copilot CLI (Opus, GPT-5.4, Sonnet); Codex CLI dispatches 3 (gpt-5.6 at xhigh, gpt-5.4, gpt-5.6 at medium). Every reviewer must return an object matching this shape before Opus triage merges them. v1.1.0 adds the rule-ID conformance checklist: when the orchestrator supplies a ${CRITERIA} block (Phase 4 Step 1.78), the reviewer must return one conformance row per selected rule ID. Findings alone cannot answer 'was this applied completely' - a reviewer that opened nothing returns the same empty findings array as one that checked everything.",
6
+ "description": "Contract for a single code-reviewer subagent's JSON output in Phase 4 Step 2. Every host dispatches 3 parallel reviewers; the middle slot is CLI-aware: Claude Code (Fable, Opus, Sonnet); Copilot CLI (Opus, GPT-5.4, Sonnet); Codex CLI dispatches 3 (gpt-5.6 at xhigh, gpt-5.4, gpt-5.6 at medium). Every reviewer must return an object matching this shape before Opus triage merges them. v1.1.0 adds the rule-ID conformance checklist: when the orchestrator supplies a ${CRITERIA} block (Phase 4 Step 1.78), the reviewer must return one conformance row per selected rule ID. Findings alone cannot answer 'was this applied completely' - a reviewer that opened nothing returns the same empty findings array as one that checked everything. v1.2.0 adds the optional per-finding fingerprint (Phase 4 Step 2.1): the stable id a finding keeps across review rounds.",
7
7
  "type": "object",
8
8
  "additionalProperties": false,
9
9
  "required": ["findings", "approved"],
@@ -97,6 +97,11 @@
97
97
  "type": "string",
98
98
  "minLength": 1,
99
99
  "description": "Which criteria source the rule came from: a registry name, a module-guide path, or 'exception-marker-audit'. Lets Phase 7 attribute findings to the standard that produced them."
100
+ },
101
+ "fingerprint": {
102
+ "type": "string",
103
+ "pattern": "^F:[0-9a-f]{8}$",
104
+ "description": "Stable cross-round identity of the finding, computed by finding-fingerprint.mjs from (file, ruleId or normalized issue text). Never includes the line. On iteration >= 2 a reviewer that recognises an entry from <previous-round-findings> echoes its fingerprint; otherwise leave it unset and the script fills it in."
100
105
  }
101
106
  }
102
107
  }
@@ -0,0 +1,55 @@
1
+ {
2
+ "$schema": "https://json-schema.org/draft/2020-12/schema",
3
+ "$id": "https://github.com/mmerterden/multi-agent-pipeline/pipeline/schemas/scope-check.schema.json",
4
+ "version": "1.0.0",
5
+ "title": "Multi-Agent Pipeline - Phase 3 scope self-check",
6
+ "description": "Written by Phase 3 Step 3.7 to $WORKTREE/.pipeline/scope-check.json before the Phase 4 handoff: one stated reason per file in the diff, the changes deliberately not made, and the code-simplifier rationales. scope-check-gate.mjs compares files[] with the real diff; Phase 4 injects the record as <scope-self-check>; Phase 6 builds the PR Changes bullets and the follow-up list from it.",
7
+ "type": "object",
8
+ "additionalProperties": false,
9
+ "required": ["version", "taskId", "files"],
10
+ "properties": {
11
+ "version": { "type": "string", "const": "1.0.0" },
12
+ "taskId": { "type": "string", "minLength": 1 },
13
+ "files": {
14
+ "type": "array",
15
+ "items": {
16
+ "type": "object",
17
+ "additionalProperties": false,
18
+ "required": ["path", "reason"],
19
+ "properties": {
20
+ "path": {
21
+ "type": "string",
22
+ "minLength": 1,
23
+ "description": "Repo-relative path as it appears in git diff --name-only."
24
+ },
25
+ "reason": {
26
+ "type": "string",
27
+ "minLength": 4,
28
+ "description": "Why this task needs this exact file. Names the requirement, never 'while I was here'."
29
+ }
30
+ }
31
+ }
32
+ },
33
+ "notDone": {
34
+ "type": "array",
35
+ "items": {
36
+ "type": "object",
37
+ "additionalProperties": false,
38
+ "required": ["what", "why"],
39
+ "properties": {
40
+ "what": { "type": "string", "minLength": 4 },
41
+ "why": { "type": "string", "enum": ["out-of-scope", "follow-up", "rejected-abstraction"] }
42
+ }
43
+ }
44
+ },
45
+ "simplifier": {
46
+ "type": "object",
47
+ "additionalProperties": false,
48
+ "properties": {
49
+ "applied": { "type": "integer", "minimum": 0 },
50
+ "skipped": { "type": "integer", "minimum": 0 },
51
+ "rationales": { "type": "array", "items": { "type": "string" } }
52
+ }
53
+ }
54
+ }
55
+ }