@mmerterden/multi-agent-pipeline 19.0.0 → 19.1.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (61) hide show
  1. package/CHANGELOG.md +121 -0
  2. package/README.md +3 -3
  3. package/docs/ecosystem.md +12 -4
  4. package/docs/facts.json +20 -4
  5. package/docs/features.md +1 -1
  6. package/docs/recovery-guide.md +8 -8
  7. package/manifest.json +64 -62
  8. package/package.json +1 -1
  9. package/pipeline/agents/dev-critic.md +4 -4
  10. package/pipeline/commands/multi-agent/SKILL.md +1 -1
  11. package/pipeline/commands/multi-agent/analysis/SKILL.md +6 -6
  12. package/pipeline/commands/multi-agent/analysis-resolve/SKILL.md +2 -2
  13. package/pipeline/commands/multi-agent/resume-local/SKILL.md +1 -1
  14. package/pipeline/commands/multi-agent/review/SKILL.md +1 -1
  15. package/pipeline/commands/multi-agent/review-analysis/SKILL.md +1 -1
  16. package/pipeline/lib/model-dispatch.sh +140 -0
  17. package/pipeline/lib/outbound-gate.mjs +14 -0
  18. package/pipeline/multi-agent-refs/_dev-context.md +5 -5
  19. package/pipeline/multi-agent-refs/analysis/evidence.md +2 -2
  20. package/pipeline/multi-agent-refs/analysis/intake.md +6 -6
  21. package/pipeline/multi-agent-refs/analysis/locked.md +27 -0
  22. package/pipeline/multi-agent-refs/analysis/redesign.md +1 -1
  23. package/pipeline/multi-agent-refs/analysis/render.md +9 -9
  24. package/pipeline/multi-agent-refs/analysis/resolve.md +1 -1
  25. package/pipeline/multi-agent-refs/analysis/review.md +2 -2
  26. package/pipeline/multi-agent-refs/analysis/synthesis.md +2 -2
  27. package/pipeline/multi-agent-refs/analysis-template-corporate.md +9 -9
  28. package/pipeline/multi-agent-refs/analysis-template.md +19 -19
  29. package/pipeline/multi-agent-refs/component-dispatch.md +5 -5
  30. package/pipeline/multi-agent-refs/conventions-defaults.md +2 -2
  31. package/pipeline/multi-agent-refs/features/analysis-jira.md +1 -1
  32. package/pipeline/multi-agent-refs/features/doctor.md +1 -1
  33. package/pipeline/multi-agent-refs/features/model-fallback.md +36 -0
  34. package/pipeline/multi-agent-refs/features/review-multi-repo.md +1 -1
  35. package/pipeline/multi-agent-refs/features/url-enrichment.md +1 -1
  36. package/pipeline/multi-agent-refs/phases/phase-1-plan.md +2 -2
  37. package/pipeline/multi-agent-refs/phases/phase-3-review.md +2 -2
  38. package/pipeline/preferences-template.json +5 -0
  39. package/pipeline/rules/figma-pipeline.md +8 -8
  40. package/pipeline/schemas/analysis-output.schema.json +1 -1
  41. package/pipeline/schemas/analysis-spec.schema.json +2 -2
  42. package/pipeline/schemas/figma-project-config.schema.json +1 -1
  43. package/pipeline/schemas/prefs.schema.json +2 -2
  44. package/pipeline/schemas/secret-patterns.json +124 -0
  45. package/pipeline/scripts/build-references.mjs +2 -2
  46. package/pipeline/scripts/bulk-read.sh +10 -1
  47. package/pipeline/scripts/cost-table.json +8 -1
  48. package/pipeline/scripts/doctor.mjs +1 -1
  49. package/pipeline/scripts/gen-facts.mjs +112 -7
  50. package/pipeline/scripts/phase-tracker.sh +5 -5
  51. package/pipeline/scripts/pre-commit-check.sh +30 -1
  52. package/pipeline/scripts/scan-skills.sh +26 -0
  53. package/pipeline/scripts/validate-analysis-doc.mjs +201 -26
  54. package/pipeline/scripts/verify-citations.mjs +1 -1
  55. package/pipeline/scripts/write-state.mjs +32 -0
  56. package/pipeline/skills/.skill-manifest.json +5 -5
  57. package/pipeline/skills/shared/core/apple-archive-compliance/SKILL.md +6 -6
  58. package/pipeline/skills/shared/core/google-play-compliance/SKILL.md +6 -6
  59. package/pipeline/skills/shared/core/multi-agent/SKILL.md +14 -13
  60. package/pipeline/skills/shared/external/NOTICE-swift-ios-skills.md +1 -1
  61. package/pipeline/skills/shared/external/signal-community/SKILL.md +8 -1
package/CHANGELOG.md CHANGED
@@ -14,6 +14,127 @@ Internal file-layout changes that don't affect the slash-command surface are sti
14
14
 
15
15
  ---
16
16
 
17
+ ## [19.1.1] - 2026-09-21
18
+
19
+ ### Changed
20
+
21
+ - **`required` floor raised to 19.1.0.** An install below it halts at Phase 0
22
+ Step 0.6, runs the update flow and asks for the command to be re-issued. The
23
+ reason is the write-state window closed in 19.1.0: a writer could report
24
+ success while its update was overwritten, and a floor exists for exactly the
25
+ case where an older install is not behind but wrong.
26
+ - `preferences-template.json` carries the `updateCheck` block explicitly
27
+ (`enabled`, `ttlHours`, `autoUpdate`). The schema defaults were already what
28
+ the phase doc describes and `smoke-update-check.sh` already pins them, but a
29
+ fresh install wrote no block at all, so the knob that decides whether a run
30
+ updates itself was not visible in the file a user opens to change it.
31
+
32
+ ---
33
+
34
+ ## [19.1.0] - 2026-09-20
35
+
36
+ Minor: the routing preference 19.0.0 shipped now reaches dispatch, and the
37
+ toolkit's three new tool families are reflected everywhere the pipeline counts
38
+ them.
39
+
40
+ ### Added
41
+
42
+ - **`pipeline/lib/model-dispatch.sh` - routing that changes the answer.**
43
+ 19.0.0 shipped `prefs.global.modelRouting` with a schema, four commands and a
44
+ status report, and nothing consulted it at dispatch time. A preference that is
45
+ written, validated and displayed but never read is worse than a missing one:
46
+ `route-status` said routing was on, the rules looked applied, and every call
47
+ went where it always went. Two call sites now ask: the subagent dispatch
48
+ contract in `skills/shared/core/multi-agent/SKILL.md`, and `bulk-read.sh`.
49
+
50
+ Precedence is `PHASE_MODEL_OVERRIDE` > a matching rule > persona
51
+ `preferredModel` > the global default. Routing sits below the per-dispatch
52
+ override on purpose: Phase 3 makes Reviewer 3 sonnet so the three reviewers
53
+ disagree, and a policy that could overrule that would turn a deliberate choice
54
+ into a suggestion.
55
+
56
+ The script never fails and never returns an empty rung - a router that can die
57
+ turns every call site into a place the run can die, for a feature that ships
58
+ disabled. Missing prefs, missing jq, unparseable JSON, an out-of-scope call
59
+ site and an unknown rung all return the caller's default, exit 0.
60
+
61
+ Two limits are enforced rather than documented. A rule preferring `fable`
62
+ falls past it while `modelFallback.fableEnabled` is false, so
63
+ `/multi-agent:model off` keeps meaning what it says. And a non-Anthropic rung
64
+ is refused for a subagent, because subagent dispatch belongs to the host - the
65
+ script says so on stderr instead of substituting an Anthropic rung and leaving
66
+ the user believing a rule worked that never could.
67
+ - `cost-table.json` rungs declare a `provider`. Without it every rung looks
68
+ alike and the subagent limit above cannot be checked at all.
69
+ - `smoke-model-dispatch.sh` (18 assertions). Half of them drive the router; the
70
+ other half assert the call sites invoke it, because a correct router nothing
71
+ calls is the same outage with better internals - which is exactly what 19.0.0
72
+ shipped.
73
+ - **Six analysis Locked decisions gained a gate.** 13 (Section 4 scenarios are
74
+ Gherkin), 14 (a goal owes a paired non-goal), 15 (a new non-SVG asset owes a
75
+ rationale), 17 (Section 9 may not write "other errors" in place of a status
76
+ code), 18 (a screenshot is embedded, not linked to a host that outlives
77
+ nothing), and 21 (References is the last numbered section). Each checks a
78
+ shape the template prescribes and keys off a structure only a real document
79
+ carries, so a minimal fixture is skipped rather than failed. The Gate status
80
+ count moves 11 -> 18, and the 18 that stay prose now say WHY in three groups:
81
+ run behaviour no document records, evidence gathering that happened before
82
+ rendering, and judgement about meaning.
83
+ - Nine anchors below Locked 25 in `smoke-locked-citations.sh`. The existing
84
+ anchors all sat above 25, which is where the v19 removal re-flowed the
85
+ numbering - and that is precisely why the older drift survived.
86
+
87
+ ### Fixed
88
+
89
+ - **Nine drifted Locked citations in `analysis-template.md`.** It cited 13 for
90
+ paired goals (14), 12 for Gherkin (13), 14 for the SVG default (15), 16 for
91
+ exhaustive response variants (17), 21 for the concept layer (22), 28 for the
92
+ SwiftUI preview (27), 18 for variant drilling (19), and "17 + 18" for the
93
+ design reference (18 + 19). A tenth cited Locked 17 for analytics PII, which
94
+ no decision covers at all - the rule stays, the number goes, because a number
95
+ a reader cannot look up is worse than none. The new anchors hold all of them.
96
+ - **Locked 33 was enforced all along and documented as prose.**
97
+ `build-references.mjs --check` runs its coverage gate and
98
+ `smoke-build-references.sh` tests it, but neither named the decision in a
99
+ failure message and the attribution check reads messages. Both name it now.
100
+ The same trap caught the new Locked 21 check on its first run: the attribution
101
+ scan reads an error message up to the first semicolon, and the message had one
102
+ in the middle.
103
+ - **write-state verifies the write after the rename, not only before it.** No
104
+ POSIX call renames a file conditionally on still holding a lock, so the
105
+ `stillOurs` check is a time-of-check and the rename is the time-of-use. The
106
+ writer now reads the file back and compares the `rev` on disk against the one
107
+ it just wrote; a different rev means another writer's rename landed on top,
108
+ and the writer exits 4 instead of 0. Reporting success while losing an update
109
+ is the one outcome that script exists to prevent, and a window it could not
110
+ see was the one place that could still happen. The clobber branch has no test:
111
+ staging it from a shell needs a hook inside the writer, and a test-only hook
112
+ in the file that guards state is the worse trade.
113
+
114
+ - **Stale phase numbers in eleven files.** `/multi-agent:review` recorded its
115
+ standalone runs under phase id 4; a tracker example in `phase-3-review.md`
116
+ drew `Phase 3 Dev`; `rules/figma-pipeline.md` carried the whole eight-phase
117
+ access matrix, including two rows for phases that no longer exist. Historical
118
+ files (CHANGELOG, ROADMAP entries, ADRs, migration headers) were left alone -
119
+ they narrate what was true then - and so were the separate phase namespaces
120
+ that `analysis/SKILL.md` and the Figma component flow use.
121
+
122
+ ### Changed
123
+
124
+ - Toolkit counts follow `multi-agent-toolkit-mcp` 3.13.0: **115 tools in 13
125
+ categories**, up from 99 in 10. The new families are a full-text context index
126
+ (FTS5 + bm25 over offloaded payloads), provider-backed research, and video key
127
+ frames. `docs/ecosystem.md` gained their rows, `docs/facts.json` regenerated,
128
+ and the MCP-server context-cost note in `doctor` and the prefs schema now
129
+ names the real number.
130
+ - `signal-community` records the second search path. The community-signal skill
131
+ carried a parity exemption because web search is not guaranteed on every host;
132
+ `research_search` runs over the MCP channel every host already speaks, so the
133
+ exemption now names the host's own search specifically and points at the
134
+ toolkit path as the preferred one.
135
+
136
+ ---
137
+
17
138
  ## [19.0.0] - 2026-09-18
18
139
 
19
140
  Major, because phase numbers are the contract and they moved. Eight phases
package/README.md CHANGED
@@ -17,7 +17,7 @@ Runs natively on Claude Code, Copilot CLI and Codex CLI. macOS only. Zero runtim
17
17
  ### Prerequisites
18
18
 
19
19
  - **Node.js >= 20.11** - required; the pipeline's own tooling runs on it.
20
- - **`jq`** - required for nine paths, optional for the rest. 82 shell files call it. The nine that publish or decide - the autopilot queue, Jira comments, PR reviews, issue updates, the plan file, both Figma fetchers, log search and Jira auth - now refuse with exit 3 rather than run, because a missing `jq` renders as empty DATA and the work carries on with it. Everywhere else it still degrades. The install prints a note when it is missing.
20
+ - **`jq`** - required for nine paths, optional for the rest. 84 shell files call it. The nine that publish or decide - the autopilot queue, Jira comments, PR reviews, issue updates, the plan file, both Figma fetchers, log search and Jira auth - now refuse with exit 3 rather than run, because a missing `jq` renders as empty DATA and the work carries on with it. Everywhere else it still degrades. The install prints a note when it is missing.
21
21
  - **`gh`** - for GitHub issue and PR work. Its built-in `--jq` is independent of the `jq` binary.
22
22
 
23
23
  ## Quick Start
@@ -134,7 +134,7 @@ Reasoning, mapping and rejected alternatives:
134
134
 
135
135
  Three smaller things in 17.6.0, each closing a gap where the pipeline assumed instead of looking:
136
136
 
137
- - **Phase 3 stopped typing `npm`.** A repo on pnpm, yarn or bun used to fail in the development phase, with a worktree and a branch already created. The manager is resolved from the repo now - an env override, then `package.json#packageManager`, then the lock file, then npm reported as a default rather than as evidence. iOS and Android are untouched.
137
+ - **Phase 2 stopped typing `npm`.** A repo on pnpm, yarn or bun used to fail in the development phase, with a worktree and a branch already created. The manager is resolved from the repo now - an env override, then `package.json#packageManager`, then the lock file, then npm reported as a default rather than as evidence. iOS and Android are untouched.
138
138
  - **A compaction no longer eats what a phase learned.** The capture hook ran at session end; an auto-compaction summarizes a long review or development phase while it is still running, and anything not yet written was gone before session end ever fired. The hooks template now flushes at `PreCompact` too.
139
139
  - **`doctor` counts your MCP servers.** Every registered server sends its tool list on every turn and they are added one at a time, so nobody ever sees the total. It reports the count and nothing else: no warning, no blocking, no disabling.
140
140
 
@@ -246,7 +246,7 @@ The widget follows the answer rather than predicting it: Phase 0 is the only til
246
246
  | `/multi-agent:test-accessibility` | VoiceOver labels, sub-44pt tap targets, contrast, traits |
247
247
  | `/multi-agent:test-dynamic-type` | Re-walk every screen at XL through accessibility-XL, report truncation |
248
248
  | `/multi-agent:test-screenshots [locale]` | App Store screenshot set in a locale (defaults to `tr`) |
249
- | `/multi-agent:manual-test` | Phase 5 standalone: check out the task branch and prepare it for Xcode |
249
+ | `/multi-agent:manual-test` | Phase 3 standalone: check out the task branch and prepare it for Xcode |
250
250
 
251
251
  ### Design, build and store
252
252
 
package/docs/ecosystem.md CHANGED
@@ -7,7 +7,7 @@ separately, wired together at install time and at run time:
7
7
  |---|---|---|
8
8
  | **`multi-agent-pipeline`** (this repo) | Orchestration: the 6-phase flow, the 60 slash commands, quality gates, review/triage, cross-CLI parity | npm package (`@mmerterden/multi-agent-pipeline`), installs itself onto Claude Code / Copilot CLI / Codex CLI |
9
9
  | **`multi-agent-plugins`** | Stack knowledge: per-platform component/lifecycle skills (iOS, Android, Frontend, Backend) + shared knowledge | Claude Code marketplace, 5 independently-versioned plugins |
10
- | **`multi-agent-toolkit-mcp`** | The pipeline's hands on devices and browsers: 99 MCP tools across 10 categories (simulator/emulator control, memory, crash diagnostics, accessibility audit, store compliance, web automation, Figma-vs-mock design audit, code intelligence, wallet passes, an agent-DSL batch runner) | npm package, registered as a standard stdio MCP server on every host |
10
+ | **`multi-agent-toolkit-mcp`** | The pipeline's hands on devices and browsers: 115 MCP tools across 13 categories (simulator/emulator control, memory, crash diagnostics, accessibility audit, store compliance, web automation, Figma-vs-mock design audit, code intelligence, wallet passes, an agent-DSL batch runner, a full-text context index, provider-backed research, video key frames) | npm package, registered as a standard stdio MCP server on every host |
11
11
 
12
12
  None of the three depends on the others at the code level. They compose through two
13
13
  narrow contracts: the **Skill tool** (pipeline → plugin, at Phase 2) and the **MCP
@@ -211,12 +211,17 @@ that step's own contract).
211
211
 
212
212
  ### multi-agent-toolkit-mcp's tools, by category
213
213
 
214
- 99 tools in 10 categories, counted from the server's own `tools/list` response at
215
- toolkit 3.12.0 rather than from a README. The table stood at 87 across 8
214
+ 115 tools in 13 categories, counted from the server's own `tools/list` response
215
+ at toolkit 3.13.0 rather than from a README. The table stood at 87 across 8
216
216
  categories for several releases because Code Intelligence and Wallet Passes
217
217
  shipped without anyone adding their rows, and nothing here was checked against
218
218
  the server - which is why the count now names its source.
219
219
 
220
+ The families can be narrowed per host: `MCP_TOOLKIT_CAPS` serves only the named
221
+ ones, and unset serves all 115. A frontend repo that sets `web,code,context`
222
+ sees 29 tools instead of 115, which matters because every registered server's
223
+ tool list is charged against the context window on every turn.
224
+
220
225
  | Category | Tools | Primary pipeline consumers |
221
226
  |---|---|---|
222
227
  | Device Control | 59 | `/multi-agent:test`, `test-dark-mode`, `test-accessibility`, `test-dynamic-type`, `test-screenshots`, `manual-test`, `design-check` |
@@ -224,11 +229,14 @@ the server - which is why the count now names its source.
224
229
  | Crash Diagnostics | 2 (`ios_list_crashes`, `android_list_crashes`) | `/multi-agent:test` full scenario (end-of-run crash sweep), outside-the-pipeline sessions |
225
230
  | Accessibility Audit | 3 (`ios_accessibility_audit`, `android_accessibility_audit`, `ios_accessibility_audit_deep`) | `/multi-agent:test` accessibility scenario, `test-accessibility` |
226
231
  | Store Compliance | 5 | `store-ready`, `testflight-validation`, `apple-archive-compliance` skill, Phase 3 Security Auditor |
227
- | Web Automation | 8 | frontend-stack UI testing (via `test`) |
232
+ | Web Automation | 18 | frontend-stack UI testing (via `test`); snapshot `ref` targeting, console/network capture, readable extraction and a bounded same-origin crawl |
228
233
  | Design Audit | 6 | `design-check` (mock-mode vs Figma conformance) |
229
234
  | Autonomous Agent DSL | 2 (`agent_run_steps`, `agent_query_output`) | any skill that needs a scripted multi-step device flow in one round trip |
230
235
  | Code Intelligence | 8 (`code_definition`, `code_references`, `code_hover`, `code_diagnostics`, `code_document_symbols`, `code_workspace_symbols`, `code_index_status`, `code_server_reset`) | `/multi-agent:refactor`, Phase 3 reviewers needing a real symbol graph rather than grep |
231
236
  | Wallet Passes | 4 (`pass_build`, `pass_validate`, `pass_inspect`, `pass_certificates`) | pass-kit work outside the pipeline; no pipeline phase consumes them |
237
+ | Context Index | 3 (`context_index`, `context_search`, `context_get`) | any skill holding an offloaded payload too large to read; ranks passages instead of grepping lines |
238
+ | Research | 2 (`research_search`, `research_ask`) | closes the parity exemption that had web search outside the MCP channel |
239
+ | Media | 1 (`media_frames`) | Phase 2 visual evidence: turns a screen recording into frames a reviewer and a model can both look at |
232
240
 
233
241
  ---
234
242
 
package/docs/facts.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "$comment": "Generated by pipeline/scripts/gen-facts.mjs. Do not hand-edit: the site reads this, and a number edited here instead of at its source is the drift this file removes.",
3
- "generatedAt": "2026-09-17",
4
- "version": "19.0.0",
3
+ "generatedAt": "2026-09-20",
4
+ "version": "19.1.1",
5
5
  "phaseSchema": 2,
6
6
  "phases": [
7
7
  {
@@ -40,6 +40,22 @@
40
40
  },
41
41
  "commandCount": 60,
42
42
  "skillCount": 153,
43
- "toolCount": 99,
44
- "toolkitVersion": "3.12.0"
43
+ "skillCountAll": 216,
44
+ "agentCount": 9,
45
+ "hosts": ["Claude Code", "Codex CLI", "Copilot CLI"],
46
+ "toolCount": 115,
47
+ "toolCategories": {
48
+ "ios": 40,
49
+ "android": 31,
50
+ "web": 18,
51
+ "code": 8,
52
+ "design": 6,
53
+ "pass": 4,
54
+ "context": 3,
55
+ "agent": 2,
56
+ "research": 2,
57
+ "media": 1
58
+ },
59
+ "pluginSkillCount": 238,
60
+ "toolkitVersion": "3.13.0"
45
61
  }
package/docs/features.md CHANGED
@@ -39,7 +39,7 @@ The install is not only useful while `/multi-agent` is running. `rules/outside-t
39
39
 
40
40
  - **Onboarded service credentials.** Resolve the logical name through `credential-store.sh` and read the issue, page or log. Writes route through the pipeline commands, which carry the rules that make them safe - issues are never auto-closed, PR bodies use `Ref:`, outward prose goes through the humanizer.
41
41
  - **The stack skills enabled for this repo.** Each toolkit's own `index` skill routes. The pipeline reads the effective `enabledPlugins` rather than keeping a stack table, so a seventh toolkit needs no code change.
42
- - **The `multi-agent-toolkit` MCP.** 99 tools for a running app.
42
+ - **The `multi-agent-toolkit` MCP.** 115 tools for a running app.
43
43
 
44
44
  Uninstall preserves the whole layer - tokens, the reader that opens them, the mapping that names them, the MCP registration. It is 1.5 kB of always-loaded text; the detail lives in a ref that loads on demand, and a gate keeps both under a ceiling because every byte there is paid by every session.
45
45
 
@@ -9,7 +9,7 @@ the answer across `modes.md`, `operations.md`, and the individual phase specs.
9
9
  | Symptom | Jump to |
10
10
  | ------------------------------------------------------- | ----------------------------------------------- |
11
11
  | Pipeline paused/halted mid-run | [Resume a paused task](#resume-a-paused-task) |
12
- | Phase 3 build failed > 3 times | [Build retry exhausted](#build-retry-exhausted) |
12
+ | Phase 2 build failed > 3 times | [Build retry exhausted](#build-retry-exhausted) |
13
13
  | Phase 3 triage returned exit 1 (invalid JSON) twice | [Triage fallback](#triage-fallback) |
14
14
  | Phase 3 triage returned exit 2 (over-rejection) | [Over-rejection](#over-rejection-guard) |
15
15
  | Worktree already exists / dirty | [Worktree collisions](#worktree-collisions) |
@@ -47,7 +47,7 @@ so `resume` continues in autopilot.
47
47
 
48
48
  ## Build Retry Exhausted
49
49
 
50
- Phase 3 retries the build up to **3 times** on each TDD cycle's green step. On
50
+ Phase 2 retries the build up to **3 times** on each TDD cycle's green step. On
51
51
  the 4th consecutive failure the pipeline **pauses** (even in autopilot) and
52
52
  surfaces the error. Autopilot does not loop infinitely - this is intentional.
53
53
 
@@ -57,7 +57,7 @@ Path forward:
57
57
  2. If the failure is environmental (missing Xcode sim, wrong JDK), fix the
58
58
  environment and `resume` - the 3-retry counter resets.
59
59
  3. If the failure is logic (compile error in the generated code), edit the
60
- offending file in the worktree, then `resume` - Phase 3 re-runs build.
60
+ offending file in the worktree, then `resume` - Phase 2 re-runs build.
61
61
  4. If the task itself is wrong-shaped (Phase 1 plan is infeasible), `kill` and
62
62
  restart with a better-scoped prompt.
63
63
 
@@ -69,15 +69,15 @@ Path forward:
69
69
 
70
70
  | Exit | Meaning | Pipeline action |
71
71
  | ---- | ------------------------ | ---------------------------------------------- |
72
- | 0 | Valid & clean | Proceed to Phase 5 |
72
+ | 0 | Valid & clean | Proceed to Phase 4 |
73
73
  | 1 | Invalid JSON / schema | Retry triage ONCE. On second exit-1, fallback: treat ALL raw findings as accepted. |
74
74
  | 2 | Over-rejection (>80%) | Pause for human confirmation |
75
75
  | 3 | Contradiction corrected | Use `result.corrected`, continue |
76
76
 
77
77
  If exit 1 fires twice (fallback path):
78
78
 
79
- 1. The pipeline logs `Phase 4: triage failed twice - fallback, all findings accepted as blocking`.
80
- 2. All raw findings loop back into Phase 3 as if they were all real blockers.
79
+ 1. The pipeline logs `Phase 3: triage failed twice - fallback, all findings accepted as blocking`.
80
+ 2. All raw findings loop back into Phase 2 as if they were all real blockers.
81
81
  3. This is intentionally conservative - we'd rather over-fix than skip
82
82
  something real. You can manually mark noise in the Phase 4 PR description.
83
83
 
@@ -91,8 +91,8 @@ asks a human to look at the raw findings vs triage output.
91
91
 
92
92
  - Open `agent-log.md`, find the triage invocation.
93
93
  - Compare raw findings (logged) with rejected items.
94
- - If the rejections are correct, run `resume` with `--accept-overrejection` (Phase 4 re-enters validation with the guard soft-failed).
95
- - If not, run `resume` - Phase 4 re-runs triage with the guard still active.
94
+ - If the rejections are correct, run `resume` with `--accept-overrejection` (Phase 3 re-enters validation with the guard soft-failed).
95
+ - If not, run `resume` - Phase 3 re-runs triage with the guard still active.
96
96
 
97
97
  Autopilot behavior: logs the warning, accepts the triage output, continues -
98
98
  the autopilot contract values throughput over human-in-loop.