@mmerterden/multi-agent-pipeline 19.0.0 → 19.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +104 -0
- package/README.md +3 -3
- package/docs/ecosystem.md +12 -4
- package/docs/facts.json +20 -4
- package/docs/features.md +1 -1
- package/docs/recovery-guide.md +8 -8
- package/manifest.json +63 -61
- package/package.json +1 -1
- package/pipeline/agents/dev-critic.md +4 -4
- package/pipeline/commands/multi-agent/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/analysis/SKILL.md +6 -6
- package/pipeline/commands/multi-agent/analysis-resolve/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/resume-local/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/review/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/review-analysis/SKILL.md +1 -1
- package/pipeline/lib/model-dispatch.sh +140 -0
- package/pipeline/lib/outbound-gate.mjs +14 -0
- package/pipeline/multi-agent-refs/_dev-context.md +5 -5
- package/pipeline/multi-agent-refs/analysis/evidence.md +2 -2
- package/pipeline/multi-agent-refs/analysis/intake.md +6 -6
- package/pipeline/multi-agent-refs/analysis/locked.md +27 -0
- package/pipeline/multi-agent-refs/analysis/redesign.md +1 -1
- package/pipeline/multi-agent-refs/analysis/render.md +9 -9
- package/pipeline/multi-agent-refs/analysis/resolve.md +1 -1
- package/pipeline/multi-agent-refs/analysis/review.md +2 -2
- package/pipeline/multi-agent-refs/analysis/synthesis.md +2 -2
- package/pipeline/multi-agent-refs/analysis-template-corporate.md +9 -9
- package/pipeline/multi-agent-refs/analysis-template.md +19 -19
- package/pipeline/multi-agent-refs/component-dispatch.md +5 -5
- package/pipeline/multi-agent-refs/conventions-defaults.md +2 -2
- package/pipeline/multi-agent-refs/features/analysis-jira.md +1 -1
- package/pipeline/multi-agent-refs/features/doctor.md +1 -1
- package/pipeline/multi-agent-refs/features/model-fallback.md +36 -0
- package/pipeline/multi-agent-refs/features/review-multi-repo.md +1 -1
- package/pipeline/multi-agent-refs/features/url-enrichment.md +1 -1
- package/pipeline/multi-agent-refs/phases/phase-1-plan.md +2 -2
- package/pipeline/multi-agent-refs/phases/phase-3-review.md +2 -2
- package/pipeline/rules/figma-pipeline.md +8 -8
- package/pipeline/schemas/analysis-output.schema.json +1 -1
- package/pipeline/schemas/analysis-spec.schema.json +2 -2
- package/pipeline/schemas/figma-project-config.schema.json +1 -1
- package/pipeline/schemas/prefs.schema.json +2 -2
- package/pipeline/schemas/secret-patterns.json +124 -0
- package/pipeline/scripts/build-references.mjs +2 -2
- package/pipeline/scripts/bulk-read.sh +10 -1
- package/pipeline/scripts/cost-table.json +8 -1
- package/pipeline/scripts/doctor.mjs +1 -1
- package/pipeline/scripts/gen-facts.mjs +112 -7
- package/pipeline/scripts/phase-tracker.sh +5 -5
- package/pipeline/scripts/pre-commit-check.sh +30 -1
- package/pipeline/scripts/scan-skills.sh +26 -0
- package/pipeline/scripts/validate-analysis-doc.mjs +201 -26
- package/pipeline/scripts/verify-citations.mjs +1 -1
- package/pipeline/scripts/write-state.mjs +32 -0
- package/pipeline/skills/.skill-manifest.json +5 -5
- package/pipeline/skills/shared/core/apple-archive-compliance/SKILL.md +6 -6
- package/pipeline/skills/shared/core/google-play-compliance/SKILL.md +6 -6
- package/pipeline/skills/shared/core/multi-agent/SKILL.md +14 -13
- package/pipeline/skills/shared/external/NOTICE-swift-ios-skills.md +1 -1
- package/pipeline/skills/shared/external/signal-community/SKILL.md +8 -1
package/CHANGELOG.md
CHANGED
|
@@ -14,6 +14,110 @@ Internal file-layout changes that don't affect the slash-command surface are sti
|
|
|
14
14
|
|
|
15
15
|
---
|
|
16
16
|
|
|
17
|
+
## [19.1.0] - 2026-09-20
|
|
18
|
+
|
|
19
|
+
Minor: the routing preference 19.0.0 shipped now reaches dispatch, and the
|
|
20
|
+
toolkit's three new tool families are reflected everywhere the pipeline counts
|
|
21
|
+
them.
|
|
22
|
+
|
|
23
|
+
### Added
|
|
24
|
+
|
|
25
|
+
- **`pipeline/lib/model-dispatch.sh` - routing that changes the answer.**
|
|
26
|
+
19.0.0 shipped `prefs.global.modelRouting` with a schema, four commands and a
|
|
27
|
+
status report, and nothing consulted it at dispatch time. A preference that is
|
|
28
|
+
written, validated and displayed but never read is worse than a missing one:
|
|
29
|
+
`route-status` said routing was on, the rules looked applied, and every call
|
|
30
|
+
went where it always went. Two call sites now ask: the subagent dispatch
|
|
31
|
+
contract in `skills/shared/core/multi-agent/SKILL.md`, and `bulk-read.sh`.
|
|
32
|
+
|
|
33
|
+
Precedence is `PHASE_MODEL_OVERRIDE` > a matching rule > persona
|
|
34
|
+
`preferredModel` > the global default. Routing sits below the per-dispatch
|
|
35
|
+
override on purpose: Phase 3 makes Reviewer 3 sonnet so the three reviewers
|
|
36
|
+
disagree, and a policy that could overrule that would turn a deliberate choice
|
|
37
|
+
into a suggestion.
|
|
38
|
+
|
|
39
|
+
The script never fails and never returns an empty rung - a router that can die
|
|
40
|
+
turns every call site into a place the run can die, for a feature that ships
|
|
41
|
+
disabled. Missing prefs, missing jq, unparseable JSON, an out-of-scope call
|
|
42
|
+
site and an unknown rung all return the caller's default, exit 0.
|
|
43
|
+
|
|
44
|
+
Two limits are enforced rather than documented. A rule preferring `fable`
|
|
45
|
+
falls past it while `modelFallback.fableEnabled` is false, so
|
|
46
|
+
`/multi-agent:model off` keeps meaning what it says. And a non-Anthropic rung
|
|
47
|
+
is refused for a subagent, because subagent dispatch belongs to the host - the
|
|
48
|
+
script says so on stderr instead of substituting an Anthropic rung and leaving
|
|
49
|
+
the user believing a rule worked that never could.
|
|
50
|
+
- `cost-table.json` rungs declare a `provider`. Without it every rung looks
|
|
51
|
+
alike and the subagent limit above cannot be checked at all.
|
|
52
|
+
- `smoke-model-dispatch.sh` (18 assertions). Half of them drive the router; the
|
|
53
|
+
other half assert the call sites invoke it, because a correct router nothing
|
|
54
|
+
calls is the same outage with better internals - which is exactly what 19.0.0
|
|
55
|
+
shipped.
|
|
56
|
+
- **Six analysis Locked decisions gained a gate.** 13 (Section 4 scenarios are
|
|
57
|
+
Gherkin), 14 (a goal owes a paired non-goal), 15 (a new non-SVG asset owes a
|
|
58
|
+
rationale), 17 (Section 9 may not write "other errors" in place of a status
|
|
59
|
+
code), 18 (a screenshot is embedded, not linked to a host that outlives
|
|
60
|
+
nothing), and 21 (References is the last numbered section). Each checks a
|
|
61
|
+
shape the template prescribes and keys off a structure only a real document
|
|
62
|
+
carries, so a minimal fixture is skipped rather than failed. The Gate status
|
|
63
|
+
count moves 11 -> 18, and the 18 that stay prose now say WHY in three groups:
|
|
64
|
+
run behaviour no document records, evidence gathering that happened before
|
|
65
|
+
rendering, and judgement about meaning.
|
|
66
|
+
- Nine anchors below Locked 25 in `smoke-locked-citations.sh`. The existing
|
|
67
|
+
anchors all sat above 25, which is where the v19 removal re-flowed the
|
|
68
|
+
numbering - and that is precisely why the older drift survived.
|
|
69
|
+
|
|
70
|
+
### Fixed
|
|
71
|
+
|
|
72
|
+
- **Nine drifted Locked citations in `analysis-template.md`.** It cited 13 for
|
|
73
|
+
paired goals (14), 12 for Gherkin (13), 14 for the SVG default (15), 16 for
|
|
74
|
+
exhaustive response variants (17), 21 for the concept layer (22), 28 for the
|
|
75
|
+
SwiftUI preview (27), 18 for variant drilling (19), and "17 + 18" for the
|
|
76
|
+
design reference (18 + 19). A tenth cited Locked 17 for analytics PII, which
|
|
77
|
+
no decision covers at all - the rule stays, the number goes, because a number
|
|
78
|
+
a reader cannot look up is worse than none. The new anchors hold all of them.
|
|
79
|
+
- **Locked 33 was enforced all along and documented as prose.**
|
|
80
|
+
`build-references.mjs --check` runs its coverage gate and
|
|
81
|
+
`smoke-build-references.sh` tests it, but neither named the decision in a
|
|
82
|
+
failure message and the attribution check reads messages. Both name it now.
|
|
83
|
+
The same trap caught the new Locked 21 check on its first run: the attribution
|
|
84
|
+
scan reads an error message up to the first semicolon, and the message had one
|
|
85
|
+
in the middle.
|
|
86
|
+
- **write-state verifies the write after the rename, not only before it.** No
|
|
87
|
+
POSIX call renames a file conditionally on still holding a lock, so the
|
|
88
|
+
`stillOurs` check is a time-of-check and the rename is the time-of-use. The
|
|
89
|
+
writer now reads the file back and compares the `rev` on disk against the one
|
|
90
|
+
it just wrote; a different rev means another writer's rename landed on top,
|
|
91
|
+
and the writer exits 4 instead of 0. Reporting success while losing an update
|
|
92
|
+
is the one outcome that script exists to prevent, and a window it could not
|
|
93
|
+
see was the one place that could still happen. The clobber branch has no test:
|
|
94
|
+
staging it from a shell needs a hook inside the writer, and a test-only hook
|
|
95
|
+
in the file that guards state is the worse trade.
|
|
96
|
+
|
|
97
|
+
- **Stale phase numbers in eleven files.** `/multi-agent:review` recorded its
|
|
98
|
+
standalone runs under phase id 4; a tracker example in `phase-3-review.md`
|
|
99
|
+
drew `Phase 3 Dev`; `rules/figma-pipeline.md` carried the whole eight-phase
|
|
100
|
+
access matrix, including two rows for phases that no longer exist. Historical
|
|
101
|
+
files (CHANGELOG, ROADMAP entries, ADRs, migration headers) were left alone -
|
|
102
|
+
they narrate what was true then - and so were the separate phase namespaces
|
|
103
|
+
that `analysis/SKILL.md` and the Figma component flow use.
|
|
104
|
+
|
|
105
|
+
### Changed
|
|
106
|
+
|
|
107
|
+
- Toolkit counts follow `multi-agent-toolkit-mcp` 3.13.0: **115 tools in 13
|
|
108
|
+
categories**, up from 99 in 10. The new families are a full-text context index
|
|
109
|
+
(FTS5 + bm25 over offloaded payloads), provider-backed research, and video key
|
|
110
|
+
frames. `docs/ecosystem.md` gained their rows, `docs/facts.json` regenerated,
|
|
111
|
+
and the MCP-server context-cost note in `doctor` and the prefs schema now
|
|
112
|
+
names the real number.
|
|
113
|
+
- `signal-community` records the second search path. The community-signal skill
|
|
114
|
+
carried a parity exemption because web search is not guaranteed on every host;
|
|
115
|
+
`research_search` runs over the MCP channel every host already speaks, so the
|
|
116
|
+
exemption now names the host's own search specifically and points at the
|
|
117
|
+
toolkit path as the preferred one.
|
|
118
|
+
|
|
119
|
+
---
|
|
120
|
+
|
|
17
121
|
## [19.0.0] - 2026-09-18
|
|
18
122
|
|
|
19
123
|
Major, because phase numbers are the contract and they moved. Eight phases
|
package/README.md
CHANGED
|
@@ -17,7 +17,7 @@ Runs natively on Claude Code, Copilot CLI and Codex CLI. macOS only. Zero runtim
|
|
|
17
17
|
### Prerequisites
|
|
18
18
|
|
|
19
19
|
- **Node.js >= 20.11** - required; the pipeline's own tooling runs on it.
|
|
20
|
-
- **`jq`** - required for nine paths, optional for the rest.
|
|
20
|
+
- **`jq`** - required for nine paths, optional for the rest. 84 shell files call it. The nine that publish or decide - the autopilot queue, Jira comments, PR reviews, issue updates, the plan file, both Figma fetchers, log search and Jira auth - now refuse with exit 3 rather than run, because a missing `jq` renders as empty DATA and the work carries on with it. Everywhere else it still degrades. The install prints a note when it is missing.
|
|
21
21
|
- **`gh`** - for GitHub issue and PR work. Its built-in `--jq` is independent of the `jq` binary.
|
|
22
22
|
|
|
23
23
|
## Quick Start
|
|
@@ -134,7 +134,7 @@ Reasoning, mapping and rejected alternatives:
|
|
|
134
134
|
|
|
135
135
|
Three smaller things in 17.6.0, each closing a gap where the pipeline assumed instead of looking:
|
|
136
136
|
|
|
137
|
-
- **Phase
|
|
137
|
+
- **Phase 2 stopped typing `npm`.** A repo on pnpm, yarn or bun used to fail in the development phase, with a worktree and a branch already created. The manager is resolved from the repo now - an env override, then `package.json#packageManager`, then the lock file, then npm reported as a default rather than as evidence. iOS and Android are untouched.
|
|
138
138
|
- **A compaction no longer eats what a phase learned.** The capture hook ran at session end; an auto-compaction summarizes a long review or development phase while it is still running, and anything not yet written was gone before session end ever fired. The hooks template now flushes at `PreCompact` too.
|
|
139
139
|
- **`doctor` counts your MCP servers.** Every registered server sends its tool list on every turn and they are added one at a time, so nobody ever sees the total. It reports the count and nothing else: no warning, no blocking, no disabling.
|
|
140
140
|
|
|
@@ -246,7 +246,7 @@ The widget follows the answer rather than predicting it: Phase 0 is the only til
|
|
|
246
246
|
| `/multi-agent:test-accessibility` | VoiceOver labels, sub-44pt tap targets, contrast, traits |
|
|
247
247
|
| `/multi-agent:test-dynamic-type` | Re-walk every screen at XL through accessibility-XL, report truncation |
|
|
248
248
|
| `/multi-agent:test-screenshots [locale]` | App Store screenshot set in a locale (defaults to `tr`) |
|
|
249
|
-
| `/multi-agent:manual-test` | Phase
|
|
249
|
+
| `/multi-agent:manual-test` | Phase 3 standalone: check out the task branch and prepare it for Xcode |
|
|
250
250
|
|
|
251
251
|
### Design, build and store
|
|
252
252
|
|
package/docs/ecosystem.md
CHANGED
|
@@ -7,7 +7,7 @@ separately, wired together at install time and at run time:
|
|
|
7
7
|
|---|---|---|
|
|
8
8
|
| **`multi-agent-pipeline`** (this repo) | Orchestration: the 6-phase flow, the 60 slash commands, quality gates, review/triage, cross-CLI parity | npm package (`@mmerterden/multi-agent-pipeline`), installs itself onto Claude Code / Copilot CLI / Codex CLI |
|
|
9
9
|
| **`multi-agent-plugins`** | Stack knowledge: per-platform component/lifecycle skills (iOS, Android, Frontend, Backend) + shared knowledge | Claude Code marketplace, 5 independently-versioned plugins |
|
|
10
|
-
| **`multi-agent-toolkit-mcp`** | The pipeline's hands on devices and browsers:
|
|
10
|
+
| **`multi-agent-toolkit-mcp`** | The pipeline's hands on devices and browsers: 115 MCP tools across 13 categories (simulator/emulator control, memory, crash diagnostics, accessibility audit, store compliance, web automation, Figma-vs-mock design audit, code intelligence, wallet passes, an agent-DSL batch runner, a full-text context index, provider-backed research, video key frames) | npm package, registered as a standard stdio MCP server on every host |
|
|
11
11
|
|
|
12
12
|
None of the three depends on the others at the code level. They compose through two
|
|
13
13
|
narrow contracts: the **Skill tool** (pipeline → plugin, at Phase 2) and the **MCP
|
|
@@ -211,12 +211,17 @@ that step's own contract).
|
|
|
211
211
|
|
|
212
212
|
### multi-agent-toolkit-mcp's tools, by category
|
|
213
213
|
|
|
214
|
-
|
|
215
|
-
toolkit 3.
|
|
214
|
+
115 tools in 13 categories, counted from the server's own `tools/list` response
|
|
215
|
+
at toolkit 3.13.0 rather than from a README. The table stood at 87 across 8
|
|
216
216
|
categories for several releases because Code Intelligence and Wallet Passes
|
|
217
217
|
shipped without anyone adding their rows, and nothing here was checked against
|
|
218
218
|
the server - which is why the count now names its source.
|
|
219
219
|
|
|
220
|
+
The families can be narrowed per host: `MCP_TOOLKIT_CAPS` serves only the named
|
|
221
|
+
ones, and unset serves all 115. A frontend repo that sets `web,code,context`
|
|
222
|
+
sees 29 tools instead of 115, which matters because every registered server's
|
|
223
|
+
tool list is charged against the context window on every turn.
|
|
224
|
+
|
|
220
225
|
| Category | Tools | Primary pipeline consumers |
|
|
221
226
|
|---|---|---|
|
|
222
227
|
| Device Control | 59 | `/multi-agent:test`, `test-dark-mode`, `test-accessibility`, `test-dynamic-type`, `test-screenshots`, `manual-test`, `design-check` |
|
|
@@ -224,11 +229,14 @@ the server - which is why the count now names its source.
|
|
|
224
229
|
| Crash Diagnostics | 2 (`ios_list_crashes`, `android_list_crashes`) | `/multi-agent:test` full scenario (end-of-run crash sweep), outside-the-pipeline sessions |
|
|
225
230
|
| Accessibility Audit | 3 (`ios_accessibility_audit`, `android_accessibility_audit`, `ios_accessibility_audit_deep`) | `/multi-agent:test` accessibility scenario, `test-accessibility` |
|
|
226
231
|
| Store Compliance | 5 | `store-ready`, `testflight-validation`, `apple-archive-compliance` skill, Phase 3 Security Auditor |
|
|
227
|
-
| Web Automation |
|
|
232
|
+
| Web Automation | 18 | frontend-stack UI testing (via `test`); snapshot `ref` targeting, console/network capture, readable extraction and a bounded same-origin crawl |
|
|
228
233
|
| Design Audit | 6 | `design-check` (mock-mode vs Figma conformance) |
|
|
229
234
|
| Autonomous Agent DSL | 2 (`agent_run_steps`, `agent_query_output`) | any skill that needs a scripted multi-step device flow in one round trip |
|
|
230
235
|
| Code Intelligence | 8 (`code_definition`, `code_references`, `code_hover`, `code_diagnostics`, `code_document_symbols`, `code_workspace_symbols`, `code_index_status`, `code_server_reset`) | `/multi-agent:refactor`, Phase 3 reviewers needing a real symbol graph rather than grep |
|
|
231
236
|
| Wallet Passes | 4 (`pass_build`, `pass_validate`, `pass_inspect`, `pass_certificates`) | pass-kit work outside the pipeline; no pipeline phase consumes them |
|
|
237
|
+
| Context Index | 3 (`context_index`, `context_search`, `context_get`) | any skill holding an offloaded payload too large to read; ranks passages instead of grepping lines |
|
|
238
|
+
| Research | 2 (`research_search`, `research_ask`) | closes the parity exemption that had web search outside the MCP channel |
|
|
239
|
+
| Media | 1 (`media_frames`) | Phase 2 visual evidence: turns a screen recording into frames a reviewer and a model can both look at |
|
|
232
240
|
|
|
233
241
|
---
|
|
234
242
|
|
package/docs/facts.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"$comment": "Generated by pipeline/scripts/gen-facts.mjs. Do not hand-edit: the site reads this, and a number edited here instead of at its source is the drift this file removes.",
|
|
3
|
-
"generatedAt": "2026-09-
|
|
4
|
-
"version": "19.
|
|
3
|
+
"generatedAt": "2026-09-20",
|
|
4
|
+
"version": "19.1.0",
|
|
5
5
|
"phaseSchema": 2,
|
|
6
6
|
"phases": [
|
|
7
7
|
{
|
|
@@ -40,6 +40,22 @@
|
|
|
40
40
|
},
|
|
41
41
|
"commandCount": 60,
|
|
42
42
|
"skillCount": 153,
|
|
43
|
-
"
|
|
44
|
-
"
|
|
43
|
+
"skillCountAll": 216,
|
|
44
|
+
"agentCount": 9,
|
|
45
|
+
"hosts": ["Claude Code", "Codex CLI", "Copilot CLI"],
|
|
46
|
+
"toolCount": 115,
|
|
47
|
+
"toolCategories": {
|
|
48
|
+
"ios": 40,
|
|
49
|
+
"android": 31,
|
|
50
|
+
"web": 18,
|
|
51
|
+
"code": 8,
|
|
52
|
+
"design": 6,
|
|
53
|
+
"pass": 4,
|
|
54
|
+
"context": 3,
|
|
55
|
+
"agent": 2,
|
|
56
|
+
"research": 2,
|
|
57
|
+
"media": 1
|
|
58
|
+
},
|
|
59
|
+
"pluginSkillCount": 238,
|
|
60
|
+
"toolkitVersion": "3.13.0"
|
|
45
61
|
}
|
package/docs/features.md
CHANGED
|
@@ -39,7 +39,7 @@ The install is not only useful while `/multi-agent` is running. `rules/outside-t
|
|
|
39
39
|
|
|
40
40
|
- **Onboarded service credentials.** Resolve the logical name through `credential-store.sh` and read the issue, page or log. Writes route through the pipeline commands, which carry the rules that make them safe - issues are never auto-closed, PR bodies use `Ref:`, outward prose goes through the humanizer.
|
|
41
41
|
- **The stack skills enabled for this repo.** Each toolkit's own `index` skill routes. The pipeline reads the effective `enabledPlugins` rather than keeping a stack table, so a seventh toolkit needs no code change.
|
|
42
|
-
- **The `multi-agent-toolkit` MCP.**
|
|
42
|
+
- **The `multi-agent-toolkit` MCP.** 115 tools for a running app.
|
|
43
43
|
|
|
44
44
|
Uninstall preserves the whole layer - tokens, the reader that opens them, the mapping that names them, the MCP registration. It is 1.5 kB of always-loaded text; the detail lives in a ref that loads on demand, and a gate keeps both under a ceiling because every byte there is paid by every session.
|
|
45
45
|
|
package/docs/recovery-guide.md
CHANGED
|
@@ -9,7 +9,7 @@ the answer across `modes.md`, `operations.md`, and the individual phase specs.
|
|
|
9
9
|
| Symptom | Jump to |
|
|
10
10
|
| ------------------------------------------------------- | ----------------------------------------------- |
|
|
11
11
|
| Pipeline paused/halted mid-run | [Resume a paused task](#resume-a-paused-task) |
|
|
12
|
-
| Phase
|
|
12
|
+
| Phase 2 build failed > 3 times | [Build retry exhausted](#build-retry-exhausted) |
|
|
13
13
|
| Phase 3 triage returned exit 1 (invalid JSON) twice | [Triage fallback](#triage-fallback) |
|
|
14
14
|
| Phase 3 triage returned exit 2 (over-rejection) | [Over-rejection](#over-rejection-guard) |
|
|
15
15
|
| Worktree already exists / dirty | [Worktree collisions](#worktree-collisions) |
|
|
@@ -47,7 +47,7 @@ so `resume` continues in autopilot.
|
|
|
47
47
|
|
|
48
48
|
## Build Retry Exhausted
|
|
49
49
|
|
|
50
|
-
Phase
|
|
50
|
+
Phase 2 retries the build up to **3 times** on each TDD cycle's green step. On
|
|
51
51
|
the 4th consecutive failure the pipeline **pauses** (even in autopilot) and
|
|
52
52
|
surfaces the error. Autopilot does not loop infinitely - this is intentional.
|
|
53
53
|
|
|
@@ -57,7 +57,7 @@ Path forward:
|
|
|
57
57
|
2. If the failure is environmental (missing Xcode sim, wrong JDK), fix the
|
|
58
58
|
environment and `resume` - the 3-retry counter resets.
|
|
59
59
|
3. If the failure is logic (compile error in the generated code), edit the
|
|
60
|
-
offending file in the worktree, then `resume` - Phase
|
|
60
|
+
offending file in the worktree, then `resume` - Phase 2 re-runs build.
|
|
61
61
|
4. If the task itself is wrong-shaped (Phase 1 plan is infeasible), `kill` and
|
|
62
62
|
restart with a better-scoped prompt.
|
|
63
63
|
|
|
@@ -69,15 +69,15 @@ Path forward:
|
|
|
69
69
|
|
|
70
70
|
| Exit | Meaning | Pipeline action |
|
|
71
71
|
| ---- | ------------------------ | ---------------------------------------------- |
|
|
72
|
-
| 0 | Valid & clean | Proceed to Phase
|
|
72
|
+
| 0 | Valid & clean | Proceed to Phase 4 |
|
|
73
73
|
| 1 | Invalid JSON / schema | Retry triage ONCE. On second exit-1, fallback: treat ALL raw findings as accepted. |
|
|
74
74
|
| 2 | Over-rejection (>80%) | Pause for human confirmation |
|
|
75
75
|
| 3 | Contradiction corrected | Use `result.corrected`, continue |
|
|
76
76
|
|
|
77
77
|
If exit 1 fires twice (fallback path):
|
|
78
78
|
|
|
79
|
-
1. The pipeline logs `Phase
|
|
80
|
-
2. All raw findings loop back into Phase
|
|
79
|
+
1. The pipeline logs `Phase 3: triage failed twice - fallback, all findings accepted as blocking`.
|
|
80
|
+
2. All raw findings loop back into Phase 2 as if they were all real blockers.
|
|
81
81
|
3. This is intentionally conservative - we'd rather over-fix than skip
|
|
82
82
|
something real. You can manually mark noise in the Phase 4 PR description.
|
|
83
83
|
|
|
@@ -91,8 +91,8 @@ asks a human to look at the raw findings vs triage output.
|
|
|
91
91
|
|
|
92
92
|
- Open `agent-log.md`, find the triage invocation.
|
|
93
93
|
- Compare raw findings (logged) with rejected items.
|
|
94
|
-
- If the rejections are correct, run `resume` with `--accept-overrejection` (Phase
|
|
95
|
-
- If not, run `resume` - Phase
|
|
94
|
+
- If the rejections are correct, run `resume` with `--accept-overrejection` (Phase 3 re-enters validation with the guard soft-failed).
|
|
95
|
+
- If not, run `resume` - Phase 3 re-runs triage with the guard still active.
|
|
96
96
|
|
|
97
97
|
Autopilot behavior: logs the warning, accepts the triage output, continues -
|
|
98
98
|
the autopilot contract values throughput over human-in-loop.
|