@mmerterden/multi-agent-pipeline 18.0.0 → 19.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +287 -0
- package/README.md +36 -20
- package/README.tr.md +14 -16
- package/docs/adr/0002-instruction-driven-flag.md +1 -0
- package/docs/adr/0005-lazy-phase-docs.md +11 -1
- package/docs/adr/0008-installer-modularization-and-secret-leak-defense.md +1 -0
- package/docs/adr/0010-own-code-graph.md +1 -0
- package/docs/adr/0014-six-phase-consolidation.md +134 -0
- package/docs/adr/README.md +2 -1
- package/docs/architecture.md +37 -38
- package/docs/best-practices.md +1 -1
- package/docs/ecosystem.md +46 -27
- package/docs/engineering.md +1 -1
- package/docs/facts.json +61 -0
- package/docs/features.md +55 -54
- package/docs/performance.md +5 -5
- package/docs/recovery-guide.md +17 -17
- package/docs/token-budget-history.md +3 -1
- package/index.js +2 -2
- package/install/_codex-agents.mjs +1 -1
- package/install/templates/claude-hooks.json +1 -1
- package/install/templates/codex-instructions.md +1 -1
- package/install/templates/copilot-instructions.md +28 -28
- package/manifest.json +234 -216
- package/package.json +2 -2
- package/pipeline/agents/dev-critic.md +7 -7
- package/pipeline/commands/figma-to-swiftui.md +1 -1
- package/pipeline/commands/multi-agent/SKILL.md +9 -9
- package/pipeline/commands/multi-agent/analysis/SKILL.md +15 -15
- package/pipeline/commands/multi-agent/analysis-resolve/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/autopilot/SKILL.md +7 -7
- package/pipeline/commands/multi-agent/channels/SKILL.md +15 -15
- package/pipeline/commands/multi-agent/diff-explain/SKILL.md +6 -6
- package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/graph/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/help/SKILL.md +62 -62
- package/pipeline/commands/multi-agent/language/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/local/SKILL.md +11 -11
- package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +13 -13
- package/pipeline/commands/multi-agent/log/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/manual-test/SKILL.md +9 -9
- package/pipeline/commands/multi-agent/model/SKILL.md +69 -0
- package/pipeline/commands/multi-agent/refactor/SKILL.md +3 -3
- package/pipeline/commands/multi-agent/resume/SKILL.md +4 -4
- package/pipeline/commands/multi-agent/resume-local/SKILL.md +19 -17
- package/pipeline/commands/multi-agent/review/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/review-analysis/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/route-off/SKILL.md +36 -0
- package/pipeline/commands/multi-agent/route-on/SKILL.md +74 -0
- package/pipeline/commands/multi-agent/route-status/SKILL.md +56 -0
- package/pipeline/commands/multi-agent/setup/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/status/SKILL.md +5 -5
- package/pipeline/commands/multi-agent/steer/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/sync/SKILL.md +12 -13
- package/pipeline/commands/multi-agent/test/SKILL.md +1 -1
- package/pipeline/lib/credential-inventory.sh +1 -1
- package/pipeline/lib/fetch-fortify.sh +1 -1
- package/pipeline/lib/model-dispatch.sh +140 -0
- package/pipeline/lib/model-rung.sh +142 -0
- package/pipeline/lib/outbound-gate.mjs +14 -0
- package/pipeline/lib/phase-schema.mjs +88 -0
- package/pipeline/lib/plan-todos.sh +5 -5
- package/pipeline/lib/route-state.sh +161 -0
- package/pipeline/lib/run-paths.sh +2 -2
- package/pipeline/multi-agent-refs/_account-picker.md +1 -1
- package/pipeline/multi-agent-refs/_dev-context.md +6 -6
- package/pipeline/multi-agent-refs/_input-parser.md +1 -1
- package/pipeline/multi-agent-refs/analysis/evidence.md +2 -11
- package/pipeline/multi-agent-refs/analysis/intake.md +7 -7
- package/pipeline/multi-agent-refs/analysis/locked.md +48 -22
- package/pipeline/multi-agent-refs/analysis/redesign.md +1 -1
- package/pipeline/multi-agent-refs/analysis/render.md +10 -10
- package/pipeline/multi-agent-refs/analysis/resolve.md +1 -1
- package/pipeline/multi-agent-refs/analysis/review.md +2 -2
- package/pipeline/multi-agent-refs/analysis/synthesis.md +13 -7
- package/pipeline/multi-agent-refs/analysis-template-corporate.md +9 -9
- package/pipeline/multi-agent-refs/analysis-template.md +19 -19
- package/pipeline/multi-agent-refs/android-guide.md +1 -1
- package/pipeline/multi-agent-refs/audit-guide.md +13 -13
- package/pipeline/multi-agent-refs/channels/issue-comment.md +2 -2
- package/pipeline/multi-agent-refs/channels/jira.md +3 -3
- package/pipeline/multi-agent-refs/channels/pr.md +4 -4
- package/pipeline/multi-agent-refs/channels/wiki.md +1 -1
- package/pipeline/multi-agent-refs/component-dispatch.md +8 -8
- package/pipeline/multi-agent-refs/conventions-defaults.md +2 -2
- package/pipeline/multi-agent-refs/cross-cli-contract.md +31 -6
- package/pipeline/multi-agent-refs/features/analysis-jira.md +1 -1
- package/pipeline/multi-agent-refs/features/autopilot-circuit-breaker.md +4 -4
- package/pipeline/multi-agent-refs/features/code-graph.md +5 -5
- package/pipeline/multi-agent-refs/features/design-conformance.md +1 -1
- package/pipeline/multi-agent-refs/features/dev-critic.md +3 -3
- package/pipeline/multi-agent-refs/features/doctor.md +3 -3
- package/pipeline/multi-agent-refs/features/external-context-injection.md +3 -3
- package/pipeline/multi-agent-refs/features/maturity-followup.md +3 -3
- package/pipeline/multi-agent-refs/features/model-fallback.md +41 -5
- package/pipeline/multi-agent-refs/features/plan-todos.md +1 -1
- package/pipeline/multi-agent-refs/features/repo-map.md +1 -1
- package/pipeline/multi-agent-refs/features/review-delta.md +3 -3
- package/pipeline/multi-agent-refs/features/review-multi-repo.md +2 -2
- package/pipeline/multi-agent-refs/features/scope-check.md +4 -4
- package/pipeline/multi-agent-refs/features/skill-conformance.md +2 -2
- package/pipeline/multi-agent-refs/features/stack-skill-routing.md +1 -1
- package/pipeline/multi-agent-refs/features/url-enrichment.md +1 -1
- package/pipeline/multi-agent-refs/features/verify-by-test.md +4 -4
- package/pipeline/multi-agent-refs/features/visual-evidence.md +19 -19
- package/pipeline/multi-agent-refs/features/worktree-finalize.md +6 -6
- package/pipeline/multi-agent-refs/issue-jira-triad.md +10 -10
- package/pipeline/multi-agent-refs/knowledge.md +11 -11
- package/pipeline/multi-agent-refs/multi-repo-integration-build.md +13 -13
- package/pipeline/multi-agent-refs/payload-contracts.md +8 -8
- package/pipeline/multi-agent-refs/phases/log-format.md +10 -10
- package/pipeline/multi-agent-refs/phases/modes.md +30 -30
- package/pipeline/multi-agent-refs/phases/operations.md +8 -8
- package/pipeline/multi-agent-refs/phases/phase-0-init.md +24 -24
- package/pipeline/multi-agent-refs/phases/phase-1-plan.md +599 -0
- package/pipeline/multi-agent-refs/phases/{phase-3-dev.md → phase-2-dev.md} +129 -49
- package/pipeline/multi-agent-refs/phases/{phase-4-review.md → phase-3-review.md} +225 -107
- package/pipeline/multi-agent-refs/phases/{phase-6-commit.md → phase-4-commit.md} +23 -23
- package/pipeline/multi-agent-refs/phases/{phase-7-report.md → phase-5-report.md} +29 -29
- package/pipeline/multi-agent-refs/phases.md +44 -48
- package/pipeline/multi-agent-refs/picker-contract.md +1 -1
- package/pipeline/multi-agent-refs/progress-contract.md +6 -6
- package/pipeline/multi-agent-refs/readiness-review.md +1 -1
- package/pipeline/multi-agent-refs/rules.md +7 -7
- package/pipeline/multi-agent-refs/swiftui-guide.md +2 -2
- package/pipeline/multi-agent-refs/tracker-contract.md +31 -32
- package/pipeline/multi-agent-refs/wiki-capture.md +14 -14
- package/pipeline/preferences-template.json +9 -1
- package/pipeline/rules/figma-pipeline.md +8 -8
- package/pipeline/rules/outside-the-pipeline.md +1 -1
- package/pipeline/schemas/agent-state.schema.json +50 -50
- package/pipeline/schemas/analysis-output.schema.json +3 -3
- package/pipeline/schemas/analysis-spec.schema.json +2 -2
- package/pipeline/schemas/autopilot-config.schema.json +1 -1
- package/pipeline/schemas/code-graph.schema.json +1 -1
- package/pipeline/schemas/criteria-manifest.schema.json +1 -1
- package/pipeline/schemas/dev-critic-output.schema.json +1 -1
- package/pipeline/schemas/diff-risk.schema.json +1 -1
- package/pipeline/schemas/figma-project-config.schema.json +1 -1
- package/pipeline/schemas/migrations/prefs-2.4.0-to-2.5.0.mjs +2 -2
- package/pipeline/schemas/migrations/prefs-2.6.0-to-2.7.0.mjs +31 -0
- package/pipeline/schemas/migrations/state-2.1.0-to-2.2.0.mjs +129 -0
- package/pipeline/schemas/phases.json +105 -0
- package/pipeline/schemas/plan-todos.schema.json +5 -5
- package/pipeline/schemas/planning-output.schema.json +1 -1
- package/pipeline/schemas/prefs.schema.json +102 -58
- package/pipeline/schemas/reviewer-output.schema.json +3 -3
- package/pipeline/schemas/route-config.schema.json +74 -0
- package/pipeline/schemas/scope-check.schema.json +1 -1
- package/pipeline/schemas/secret-patterns.json +124 -0
- package/pipeline/schemas/test-gap.schema.json +1 -1
- package/pipeline/schemas/token-budget.json +12 -18
- package/pipeline/schemas/triage-output.schema.json +6 -6
- package/pipeline/scripts/README.md +3 -3
- package/pipeline/scripts/_code-graph.mjs +2 -2
- package/pipeline/scripts/_run-paths.mjs +2 -2
- package/pipeline/scripts/_smoke-root.sh +1 -1
- package/pipeline/scripts/aggregate-metrics.mjs +1 -1
- package/pipeline/scripts/build-references.mjs +2 -2
- package/pipeline/scripts/bulk-read.sh +10 -1
- package/pipeline/scripts/capture-flush.sh +8 -8
- package/pipeline/scripts/capture-resume.sh +3 -3
- package/pipeline/scripts/classify-plan-safety.mjs +1 -1
- package/pipeline/scripts/cost-table.json +8 -1
- package/pipeline/scripts/diff-explain.mjs +1 -1
- package/pipeline/scripts/doctor.mjs +3 -3
- package/pipeline/scripts/gc-abandoned.sh +3 -3
- package/pipeline/scripts/gc-tmp.sh +1 -1
- package/pipeline/scripts/gc-worktrees.sh +1 -1
- package/pipeline/scripts/gen-facts.mjs +280 -0
- package/pipeline/scripts/gen-mode-dispatch.mjs +32 -37
- package/pipeline/scripts/gen-ref-toc.mjs +1 -1
- package/pipeline/scripts/graph-report.mjs +1 -1
- package/pipeline/scripts/jira-attach.sh +1 -1
- package/pipeline/scripts/learn-from-transcripts.mjs +1 -1
- package/pipeline/scripts/learning-curve.mjs +2 -2
- package/pipeline/scripts/log-metric.sh +17 -4
- package/pipeline/scripts/memory-save.sh +1 -1
- package/pipeline/scripts/migrate-prefs.mjs +22 -5
- package/pipeline/scripts/phase-banner.sh +20 -20
- package/pipeline/scripts/phase-tracker.sh +12 -12
- package/pipeline/scripts/plan-coverage-gate.mjs +2 -2
- package/pipeline/scripts/pre-commit-check.sh +30 -1
- package/pipeline/scripts/render-agent-log-cost.sh +1 -1
- package/pipeline/scripts/render-work-summary.sh +3 -3
- package/pipeline/scripts/review-file-filter.mjs +1 -1
- package/pipeline/scripts/run-aggregator.mjs +13 -6
- package/pipeline/scripts/run-metrics.mjs +1 -1
- package/pipeline/scripts/runs-index.mjs +11 -1
- package/pipeline/scripts/scan-skills.sh +26 -0
- package/pipeline/scripts/smoke-cross-cli-behavior.sh +6 -6
- package/pipeline/scripts/smoke-schema-validation.sh +26 -7
- package/pipeline/scripts/token-budget-report.mjs +13 -2
- package/pipeline/scripts/triage-memory.mjs +2 -2
- package/pipeline/scripts/validate-analysis-doc.mjs +274 -43
- package/pipeline/scripts/validate-planning.mjs +1 -1
- package/pipeline/scripts/validate-reviewer.mjs +1 -1
- package/pipeline/scripts/validate-state.mjs +45 -5
- package/pipeline/scripts/validate-triage.mjs +3 -3
- package/pipeline/scripts/verify-citations.mjs +1 -1
- package/pipeline/scripts/worktree-finalize.sh +5 -5
- package/pipeline/scripts/write-state.mjs +32 -0
- package/pipeline/skills/.skill-manifest.json +38 -22
- package/pipeline/skills/.skills-index.json +49 -5
- package/pipeline/skills/shared/README.md +10 -6
- package/pipeline/skills/shared/core/apple-archive-compliance/SKILL.md +8 -8
- package/pipeline/skills/shared/core/google-play-compliance/SKILL.md +8 -8
- package/pipeline/skills/shared/core/multi-agent/SKILL.md +81 -82
- package/pipeline/skills/shared/core/multi-agent-autopilot/SKILL.md +3 -3
- package/pipeline/skills/shared/core/multi-agent-channels/SKILL.md +14 -14
- package/pipeline/skills/shared/core/multi-agent-diff-explain/SKILL.md +5 -5
- package/pipeline/skills/shared/core/multi-agent-graph/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +25 -23
- package/pipeline/skills/shared/core/multi-agent-language/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-local/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +8 -8
- package/pipeline/skills/shared/core/multi-agent-manual-test/SKILL.md +6 -6
- package/pipeline/skills/shared/core/multi-agent-model/SKILL.md +71 -0
- package/pipeline/skills/shared/core/multi-agent-refactor/SKILL.md +3 -3
- package/pipeline/skills/shared/core/multi-agent-resume/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-resume-local/SKILL.md +7 -7
- package/pipeline/skills/shared/core/multi-agent-route-off/SKILL.md +39 -0
- package/pipeline/skills/shared/core/multi-agent-route-on/SKILL.md +76 -0
- package/pipeline/skills/shared/core/multi-agent-route-status/SKILL.md +59 -0
- package/pipeline/skills/shared/core/multi-agent-setup/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-status/SKILL.md +5 -5
- package/pipeline/skills/shared/core/multi-agent-steer/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +6 -5
- package/pipeline/skills/shared/external/NOTICE-swift-ios-skills.md +1 -1
- package/pipeline/skills/shared/external/signal-community/SKILL.md +8 -1
- package/pipeline/skills/skills-index.md +8 -4
- package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +0 -263
- package/pipeline/multi-agent-refs/phases/phase-2-planning.md +0 -344
- package/pipeline/multi-agent-refs/phases/phase-5-test.md +0 -182
|
@@ -153,7 +153,7 @@ after the user has already answered the pickers.
|
|
|
153
153
|
### identity
|
|
154
154
|
|
|
155
155
|
A git identity resolves for the account the run would commit as. WARN: the run
|
|
156
|
-
reaches Phase
|
|
156
|
+
reaches Phase 4 and stops there.
|
|
157
157
|
|
|
158
158
|
### hook-coverage
|
|
159
159
|
|
|
@@ -228,7 +228,7 @@ is countable.
|
|
|
228
228
|
|
|
229
229
|
The reason it exists: every registered server sends its tool list on every turn,
|
|
230
230
|
the user adds them one at a time, and nobody ever sees the running total - our
|
|
231
|
-
own toolkit contributes
|
|
231
|
+
own toolkit contributes 115 tools by itself. This is the same argument that made
|
|
232
232
|
the pre-run context budget a measured number rather than an intention.
|
|
233
233
|
|
|
234
234
|
It only reports. It never disables a server, never blocks and never warns: how
|
|
@@ -246,7 +246,7 @@ file it was writing at the time.
|
|
|
246
246
|
|
|
247
247
|
Worktrees left under `<repo>/.worktrees/` in the repository the caller is
|
|
248
248
|
standing in. WARN at five or more, or at 2 GB. A finished task removes its own
|
|
249
|
-
worktree at PR time, but a run that stops before Phase
|
|
249
|
+
worktree at PR time, but a run that stops before Phase 4 never reaches that
|
|
250
250
|
step and nothing else collects it: the finalizer only runs on success, and
|
|
251
251
|
`gc-worktrees` only sweeps entries git has already forgotten. Each survivor is
|
|
252
252
|
a full second checkout, so the total is measured in gigabytes rather than
|
|
@@ -54,7 +54,7 @@ ticket that referenced nothing. The run then produced a plan from a partial pict
|
|
|
54
54
|
reported success, and the user found out by reading the output.
|
|
55
55
|
|
|
56
56
|
So on exit `3`, classify the stderr and surface an `AskUserQuestion` before continuing.
|
|
57
|
-
The classification is the same shape the Phase
|
|
57
|
+
The classification is the same shape the Phase 4 remote gate uses, because the failures
|
|
58
58
|
are the same failures:
|
|
59
59
|
|
|
60
60
|
| stderr signal | Diagnosis | Question offers |
|
|
@@ -78,14 +78,14 @@ Rules that make this actionable rather than decorative:
|
|
|
78
78
|
"planned without Crashlytics, user's choice at <ts>" in the analysis doc's limitations
|
|
79
79
|
section instead of leaving a silent hole.
|
|
80
80
|
5. **Autopilot still reports.** Autopilot takes "continue without this source" without
|
|
81
|
-
asking, but the skipped source is logged and carried into the Phase
|
|
81
|
+
asking, but the skipped source is logged and carried into the Phase 5 report. An
|
|
82
82
|
autopilot run that quietly planned from a partial picture is the same defect with a
|
|
83
83
|
flag on it.
|
|
84
84
|
|
|
85
85
|
Failures remain non-fatal at Phase 1 by default - the agent still runs analysis. What
|
|
86
86
|
changed is that the user learns about it while the choice is still theirs to make.
|
|
87
87
|
|
|
88
|
-
**Graylog specifics.** `fetch-graylog.sh` degrades to empty on any network/VPN failure: it emits a normalized empty object and exits `0`, so an unreachable host reads as `{status:"skipped", reason:"vpn-unreachable"}` and never blocks. **Exiting `0` is not permission to stay quiet.** The payload carries `degraded: true` and a `degradeReason`; when it does, log one line naming the host and the reason, and carry it into the Phase
|
|
88
|
+
**Graylog specifics.** `fetch-graylog.sh` degrades to empty on any network/VPN failure: it emits a normalized empty object and exits `0`, so an unreachable host reads as `{status:"skipped", reason:"vpn-unreachable"}` and never blocks. **Exiting `0` is not permission to stay quiet.** The payload carries `degraded: true` and a `degradeReason`; when it does, log one line naming the host and the reason, and carry it into the Phase 5 report. Logs are advisory, so this never asks a question and never blocks - but "I searched the logs and found nothing" and "I never reached the log server" are different statements, and only one of them is true. Only a genuine auth rejection on a reachable host exits `3` (marked `failed`, still non-fatal here). `state.graylogContext` is prepended to the analysis prompt inside the **Referenced External Sources** section as diagnostic context that is **advisory only** - the agent may use it to orient on a reported error, but code remains ground truth, and there is no Phase 3 review gate for logs.
|
|
89
89
|
|
|
90
90
|
## Prompt injection shape
|
|
91
91
|
|
|
@@ -128,15 +128,15 @@ maturity step has `currentPhase: 0`, so resuming would start at Phase 1 and skip
|
|
|
128
128
|
the check - the halt would be permanent in the one direction that matters.
|
|
129
129
|
|
|
130
130
|
So resume reads `state.waitingFor` first: when it names a step, the run re-enters
|
|
131
|
-
THAT step rather than the next phase. `waitingFor` already existed and Phase
|
|
131
|
+
THAT step rather than the next phase. `waitingFor` already existed and Phase 5's
|
|
132
132
|
channels pause already documented itself as resumable through it
|
|
133
|
-
(`phases/phase-
|
|
133
|
+
(`phases/phase-5-report.md`), while `resume/SKILL.md` never mentioned the field -
|
|
134
134
|
so that pause had the same gap and this fixes both.
|
|
135
135
|
|
|
136
136
|
| `waitingFor` | Re-entry |
|
|
137
137
|
|---|---|
|
|
138
138
|
| `maturity` | Phase 0, the maturity step, with the item re-fetched |
|
|
139
|
-
| `user-channels-choice` | Phase
|
|
139
|
+
| `user-channels-choice` | Phase 5, the channels menu |
|
|
140
140
|
| absent | `currentPhase + 1`, as before |
|
|
141
141
|
|
|
142
142
|
`waitingFor` is cleared by the write that records the answer. A field that
|
|
@@ -6,6 +6,7 @@
|
|
|
6
6
|
- [Turning the fable rung off](#turning-the-fable-rung-off)
|
|
7
7
|
- [Triggers (checked in this order)](#triggers-checked-in-this-order)
|
|
8
8
|
- [Logging](#logging)
|
|
9
|
+
- [Routing at dispatch](#routing-at-dispatch)
|
|
9
10
|
- [Non-goals](#non-goals)
|
|
10
11
|
- [Codex CLI](#codex-cli)
|
|
11
12
|
<!-- /toc -->
|
|
@@ -42,7 +43,7 @@ architect/reviewer roles with no file edits.
|
|
|
42
43
|
Personas and dispatch read the **rung** (`fable`, `opus`, `sonnet`, `haiku`). The wire
|
|
43
44
|
model ID each rung resolves to lives in exactly two places: `scripts/cost-table.json`
|
|
44
45
|
(`prices.<rung>.modelId`) and the per-host reviewer table in
|
|
45
|
-
`phases/phase-
|
|
46
|
+
`phases/phase-3-review.md`. A generation move edits those; it never renames a rung and
|
|
46
47
|
never touches a persona file.
|
|
47
48
|
|
|
48
49
|
Current resolution: `fable` -> `claude-fable-5`, `opus` -> `claude-opus-5`,
|
|
@@ -61,7 +62,7 @@ Three shifts on the current opus rung are worth knowing when reading a run that
|
|
|
61
62
|
different rather than broken, and none of them are bugs in this contract:
|
|
62
63
|
|
|
63
64
|
- **Longer user-facing output.** Effort is not the lever; prompt-level conciseness is.
|
|
64
|
-
Phase
|
|
65
|
+
Phase 5 report length and reviewer prose are where this shows.
|
|
65
66
|
- **Self-verification without being asked.** Explicit "double-check your work"
|
|
66
67
|
scaffolding now causes over-verification rather than preventing under-verification.
|
|
67
68
|
Phase 4's deterministic gates already carry that load.
|
|
@@ -162,7 +163,7 @@ prints it once instead:
|
|
|
162
163
|
with `PHASE_MODEL_OVERRIDE=<floorModel>`. A failure at the floor (or when no
|
|
163
164
|
floor is configured) falls through to the normal phase-error path (pause ->
|
|
164
165
|
resume). Never silent-skip the persona. Each downgrade emits its own
|
|
165
|
-
`model_fallback` metric line so a two-step degrade is visible in Phase
|
|
166
|
+
`model_fallback` metric line so a two-step degrade is visible in Phase 5.
|
|
166
167
|
3. **Cost budget ceiling (existing gate).** When `cost-budget-check.mjs` exits 11
|
|
167
168
|
(exceeded) mid-run, the run already pauses per the cost-budget contract; on
|
|
168
169
|
user-approved continue, `preferredModel` personas downgrade to `fallbackModel` for the
|
|
@@ -172,10 +173,45 @@ prints it once instead:
|
|
|
172
173
|
|
|
173
174
|
Every fallback emits one `log-metric.sh` line (`metric: model_fallback`,
|
|
174
175
|
attributes: `persona`, `from`, `to`, `trigger: date-gate|dispatch-error|budget`)
|
|
175
|
-
and one agent-log line so Phase
|
|
176
|
+
and one agent-log line so Phase 5 reports show which phases ran degraded.
|
|
176
177
|
The cost ledger prices the dispatch at the model actually used (the tracker's
|
|
177
178
|
per-phase `model` field already carries the override).
|
|
178
179
|
|
|
180
|
+
## Routing at dispatch
|
|
181
|
+
|
|
182
|
+
`$HOME/.claude/lib/model-dispatch.sh` turns `prefs.global.modelRouting` into a
|
|
183
|
+
rung. Two call sites ask it: the subagent dispatch contract in
|
|
184
|
+
`skills/shared/core/multi-agent/SKILL.md`, and `bulk-read.sh`.
|
|
185
|
+
|
|
186
|
+
Routing is consulted only when the phase named no model. A per-dispatch
|
|
187
|
+
`PHASE_MODEL_OVERRIDE` is a decision the phase made for a reason - Phase 3 sets
|
|
188
|
+
Reviewer 3 to sonnet so the three reviewers disagree - and a policy able to
|
|
189
|
+
overrule it would turn a deliberate choice into a suggestion.
|
|
190
|
+
|
|
191
|
+
Two limits are enforced by the script rather than written here and hoped for:
|
|
192
|
+
|
|
193
|
+
- **A rule may not turn the fable rung on.** With `modelFallback.fableEnabled`
|
|
194
|
+
false, a rule preferring `fable` falls past it to the next rung. Otherwise
|
|
195
|
+
`/multi-agent:model off` would have a second door.
|
|
196
|
+
- **A subagent stays inside the Anthropic ladder.** Subagent dispatch belongs to
|
|
197
|
+
the host, so a rule preferring a non-Anthropic rung cannot be honoured there;
|
|
198
|
+
the script says so on stderr and moves on rather than substituting an
|
|
199
|
+
Anthropic rung and leaving the user believing a rule worked that never could.
|
|
200
|
+
`bulk-read.sh` is a call site the pipeline makes itself, so the same rung IS
|
|
201
|
+
allowed there. That asymmetry is why `/multi-agent:route-status` reports scope
|
|
202
|
+
per call site instead of one on/off.
|
|
203
|
+
|
|
204
|
+
The script never fails and never returns an empty rung: a missing preferences
|
|
205
|
+
file, a missing `jq`, unparseable JSON, an out-of-scope call site and a rung
|
|
206
|
+
absent from `cost-table.json` all return the caller's default and exit 0. A
|
|
207
|
+
router that can die would turn every call site into a place the run can die, for
|
|
208
|
+
a feature that ships disabled.
|
|
209
|
+
|
|
210
|
+
`cost-table.json` rows carry a `provider`, which is what tells the two layers
|
|
211
|
+
apart. `smoke-model-dispatch.sh` drives the script and also asserts that the two
|
|
212
|
+
call sites invoke it: a correct router nothing calls is the same outage with
|
|
213
|
+
better internals.
|
|
214
|
+
|
|
179
215
|
## Non-goals
|
|
180
216
|
|
|
181
217
|
- No automatic re-upgrade mid-run (a run that fell back stays fallen back -
|
|
@@ -213,5 +249,5 @@ parent model and reasoning effort and silently discards the override, so a
|
|
|
213
249
|
fallback that omits it appears to apply while changing nothing.
|
|
214
250
|
|
|
215
251
|
The three model ids above live in exactly three places - this table,
|
|
216
|
-
`pipeline/scripts/cost-table.json`, and the Phase
|
|
252
|
+
`pipeline/scripts/cost-table.json`, and the Phase 3 reviewer matrix - so an OpenAI
|
|
217
253
|
rename is a three-file change.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Feature: Plan Todos Iteration (Phase 3)
|
|
2
2
|
|
|
3
|
-
**Gated by `prefs.global.planTodos.enabled`** (default: `false`). When enabled and Phase
|
|
3
|
+
**Gated by `prefs.global.planTodos.enabled`** (default: `false`). When enabled and Phase 1 Step 11 emitted a `plan.todos[]` (conforming to `$HOME/.claude/schemas/plan-todos.schema.json`), Phase 2 iterates with the helper instead of walking `tasks[]` directly:
|
|
4
4
|
|
|
5
5
|
```bash
|
|
6
6
|
while next=$(bash "$HOME/.claude/lib/plan-todos.sh" next "$TASK_ID"); [ -n "$next" ]; do
|
|
@@ -23,7 +23,7 @@ fi
|
|
|
23
23
|
- **Deterministic** - same worktree state → same map. No embeddings, no network. Sub-second on monorepos.
|
|
24
24
|
- **Advisory** - Explore agents may ignore the map if they have stronger signals (Phase 0 knowledge cache, explicit user context). Never gates the pipeline.
|
|
25
25
|
- **Truncation-safe** - output is capped at `tokenBudget`; the script falls back to "no files extracted within budget" rather than over-spending.
|
|
26
|
-
- **Cost ledger** - emit `phase-1.repo_map_emitted bytes=$(wc -c <<<"$REPO_MAP") budget=$BUDGET` so Phase
|
|
26
|
+
- **Cost ledger** - emit `phase-1.repo_map_emitted bytes=$(wc -c <<<"$REPO_MAP") budget=$BUDGET` so Phase 5 cost summary can show the size injected.
|
|
27
27
|
|
|
28
28
|
## When to enable
|
|
29
29
|
|
|
@@ -6,7 +6,7 @@ Gated by `prefs.global.autopilotCircuitBreaker` (`enabled` default true, `identi
|
|
|
6
6
|
|
|
7
7
|
## Files per round
|
|
8
8
|
|
|
9
|
-
Phase
|
|
9
|
+
Phase 3 Step 3.2.1 writes `$WORKTREE/.pipeline/triage-round-<N>.json` (N = `state.reviewIterations | length`), validates it, annotates it and copies it to `$WORKTREE/triage-output.json`, the name Phase 5, `worktree-finalize.sh`, `render-work-summary.sh` and `diff-explain.mjs` read. A copy rather than a symlink because the salvage is `cp -R`, and `.pipeline/` is on the salvage list, so every round survives into `artifactsPath`. Step 3.7 rewrites the round file; the copy is repeated after it.
|
|
10
10
|
|
|
11
11
|
## Step 2.1 block: previous-round findings (iteration >= 2)
|
|
12
12
|
|
|
@@ -34,7 +34,7 @@ The cap of 40 entries drops `important` before `blocking` so the prefix stays bo
|
|
|
34
34
|
|
|
35
35
|
## Step 2.2 block: scope self-check (every iteration)
|
|
36
36
|
|
|
37
|
-
Phase
|
|
37
|
+
Phase 2 Step 3.7 wrote `$WORKTREE/.pipeline/scope-check.json` (contract: `features/scope-check.md`). Render it so reviewers judge the diff against the dev's stated scope and do not re-propose what was rejected:
|
|
38
38
|
|
|
39
39
|
```bash
|
|
40
40
|
SCOPE_JSON="$WORKTREE/.pipeline/scope-check.json"
|
|
@@ -82,7 +82,7 @@ Same argument `shortId()` in `_retrieval.mjs` makes for corpus rows: derived fro
|
|
|
82
82
|
|
|
83
83
|
## Telemetry
|
|
84
84
|
|
|
85
|
-
`review.delta` once per iteration >= 2: `iteration`, `new`, `still_present`, `resolved`, `plateau`, `tripped`. `run-metrics.mjs` reports `reviewDelta.stillPresentFinal`, `resolvedTotal` and `tripped`; Phase
|
|
85
|
+
`review.delta` once per iteration >= 2: `iteration`, `new`, `still_present`, `resolved`, `plateau`, `tripped`. `run-metrics.mjs` reports `reviewDelta.stillPresentFinal`, `resolvedTotal` and `tripped`; Phase 5 renders the three counts per round.
|
|
86
86
|
|
|
87
87
|
## Reference
|
|
88
88
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
|
-
# Phase
|
|
1
|
+
# Phase 3 Review - Multi-Repo Mode
|
|
2
2
|
|
|
3
|
-
> Loaded on demand from `phases/phase-
|
|
3
|
+
> Loaded on demand from `phases/phase-3-review.md`. Conditional on
|
|
4
4
|
> `state.projects[].length > 1`; a single-repo run (the common case) never needs
|
|
5
5
|
> it and used to pay for it on every review.
|
|
6
6
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
|
-
# Feature: Scope self-check (Phase
|
|
1
|
+
# Feature: Scope self-check (Phase 2 Step 3.7)
|
|
2
2
|
|
|
3
|
-
**Pattern**: Phase
|
|
3
|
+
**Pattern**: Phase 3 reconstructs everything from the diff. The one thing it cannot reconstruct is why each file was touched and what was left out on purpose, so Dev states both before the handoff, and a deterministic gate checks the statement against the real diff. The record also carries the code-simplifier rationales (Step 3.6) that used to be discarded, and it feeds two later consumers: the `<scope-self-check>` block in the Phase 3 reviewer prefix and the PR body in Phase 4.
|
|
4
4
|
|
|
5
5
|
## The record
|
|
6
6
|
|
|
@@ -33,8 +33,8 @@ Output `{ok, justified, unjustified[], unlisted[], notDone}`. Exit 1 lists `unju
|
|
|
33
33
|
|
|
34
34
|
## Consumers
|
|
35
35
|
|
|
36
|
-
- Phase
|
|
37
|
-
- Phase
|
|
36
|
+
- Phase 3 Step 2.2 renders file reasons, unjustified files and `notDone[]` into the shared reviewer prefix (`features/review-delta.md`).
|
|
37
|
+
- Phase 4 Step 3 builds the PR `## Changes` bullets from `files[].reason` and lists `notDone[]` under `## Related` as "Follow-ups not done in this PR", merged with the final triage `deferred[]` (`channels/pr.md`).
|
|
38
38
|
|
|
39
39
|
## Reference
|
|
40
40
|
|
|
@@ -13,7 +13,7 @@
|
|
|
13
13
|
- [Preference](#preference)
|
|
14
14
|
<!-- /toc -->
|
|
15
15
|
|
|
16
|
-
> **TLDR** - Phase
|
|
16
|
+
> **TLDR** - Phase 3 Step 1.78 resolves, deterministically and before any reviewer runs, WHAT the changed code was supposed to honour: declared rule registries scoped to the diff's languages, in-repo module guides, and the toolchains those registries delegate to. The result is `criteria-manifest.json`: a bounded set of rule IDs that becomes the denominator for "was this applied completely". Reviewers answer per rule ID. An ID that is neither checked nor explicitly waived fails the stage.
|
|
17
17
|
|
|
18
18
|
## Why this exists
|
|
19
19
|
|
|
@@ -88,7 +88,7 @@ Cheap, language-agnostic, and it answers "completely" head on - an exception i
|
|
|
88
88
|
|---|---|
|
|
89
89
|
| `detectedStack` (Phase 1) | language census of the diff by file extension. Describes what the diff CONTAINS, not what the repo is nominally built in, so one Objective-C bridging file in a Swift repo is classified correctly |
|
|
90
90
|
| Phase 1 analysis summary (triage scope) | the task description plus the inline task list Phase 3 generated for itself |
|
|
91
|
-
| Phase
|
|
91
|
+
| Phase 1 plan | same inline task list |
|
|
92
92
|
| `state.evidence.figma[]` (Step 1.8) | recorded as `not-applicable (no Phase 1 evidence)` |
|
|
93
93
|
| Step 2.8 visual conformance | recorded as `not-applicable (no Phase 1 evidence)` |
|
|
94
94
|
|
|
@@ -67,7 +67,7 @@ Append one `state.telemetry.skillCalls[]` entry per skill actually loaded:
|
|
|
67
67
|
|
|
68
68
|
`routedBy` names the index and version that chose it. That is the difference between "the model happened to read a skill" and "the toolkit said this skill governs this task".
|
|
69
69
|
|
|
70
|
-
What Phase 4 actually does with it, precisely: Step 1.78 lists these entries in the manifest under `ledger.routedByToolkit`, so a reviewer and the Phase
|
|
70
|
+
What Phase 4 actually does with it, precisely: Step 1.78 lists these entries in the manifest under `ledger.routedByToolkit`, so a reviewer and the Phase 5 report can see which skills the project's own toolkit selected. It does **not** give them extra weight in the coverage maths. The deterministic resolver stays primary because an unrecorded load and no load are indistinguishable in state, and no `routedBy` tag changes that - the tag says who chose the skill, not that the code honoured it.
|
|
71
71
|
|
|
72
72
|
## Failure modes, and why none of them halt
|
|
73
73
|
|
|
@@ -77,7 +77,7 @@ Two entry shapes reach this step, and `metadata.source` says which:
|
|
|
77
77
|
- Issue list / single issue: if `issueInstanceId` in URL → `GET /ssc/api/v1/projectVersions/<versionId>/issues?q=issueInstanceId:"<id>"&fields=id,issueName,friority,fullFileName,lineNumber,analyzer,kingdom`. Otherwise fetch top open issue of the version.
|
|
78
78
|
- Issue details + recommendation: `GET /ssc/api/v1/issueDetails/<issueId>` → description, recommendation, code snippet (if the artifact is retained).
|
|
79
79
|
- Store as `state.fortifyFinding = { host, versionId, versionName, projectName, issueId?, issueName?, severity, friority, kingdom, analyzer, file, line, description, recommendation, codeSnippet?, fetchedAt }`.
|
|
80
|
-
- Phase
|
|
80
|
+
- Phase 1 Plan injects `state.fortifyFinding` into the architectural-review prompt under a **Known Security Finding** section, and the generated task breakdown includes a task addressing the finding (naming convention: `sec(fortify-<issueId>): <issueName> at <file>:<line>`).
|
|
81
81
|
- Soft-fail: 401 → re-run Token Save Flow for `fortify`; 403 → warn and skip; network timeout → respect VPN fallback (skip enrichment, mark `state.fortifyFinding = { skipped: "vpn-unreachable" }`).
|
|
82
82
|
|
|
83
83
|
##### Step 1b.3 - Graylog deep fetch (runs when `type == "graylog"` present in `state.contextLinks[]`)
|
|
@@ -1,4 +1,4 @@
|
|
|
1
|
-
# Feature: Verify-by-Test Triage (Phase
|
|
1
|
+
# Feature: Verify-by-Test Triage (Phase 3 Step 3.7)
|
|
2
2
|
|
|
3
3
|
**Pattern**: reviewer findings are hypotheses; a failing repro test is proof. Adversarial-review research (Refute-or-Promote, 2026) found that plausible-but-wrong findings survive debate rounds but die on a single empirical test. This step converts the highest-stakes verdicts (accepted blocking) from judgment into evidence before the Phase 3 rework loop fires.
|
|
4
4
|
|
|
@@ -19,7 +19,7 @@
|
|
|
19
19
|
| Passes on some runs, fails on others | `inconclusive` | Flake, not proof either way. Finding stays accepted blocking, `verification.note = "flaky: passed k/N"`, repro test file deleted, telemetry line gains `flaky=<count>`. |
|
|
20
20
|
| Compile error / timeout / not unit-testable | `inconclusive` | Finding stays accepted blocking (judgment stands). Partial test deleted, cause in `verification.note`. |
|
|
21
21
|
|
|
22
|
-
Downgrades go to `deferred`, never `rejected`: triage judged the issue real, and deferred items surface in the Phase
|
|
22
|
+
Downgrades go to `deferred`, never `rejected`: triage judged the issue real, and deferred items surface in the Phase 5 report for a human eye.
|
|
23
23
|
|
|
24
24
|
## Red-test handoff to Phase 3
|
|
25
25
|
|
|
@@ -27,7 +27,7 @@ Downgrades go to `deferred`, never `rejected`: triage judged the issue real, and
|
|
|
27
27
|
|
|
28
28
|
## Cleanup invariant
|
|
29
29
|
|
|
30
|
-
After Step 3.7, the only uncommitted verifier artifacts are the confirmed repro tests listed in `redTests[]` (committed later with the fix) and logs under `$WORKTREE/.pipeline/` (outside Phase
|
|
30
|
+
After Step 3.7, the only uncommitted verifier artifacts are the confirmed repro tests listed in `redTests[]` (committed later with the fix) and logs under `$WORKTREE/.pipeline/` (outside Phase 4 commit scope). `not-reproduced` and `inconclusive` test files are always deleted.
|
|
31
31
|
|
|
32
32
|
## Telemetry
|
|
33
33
|
|
|
@@ -39,4 +39,4 @@ Adds one Sonnet call plus up to `maxFindings` single-test runs (and build-lock c
|
|
|
39
39
|
|
|
40
40
|
## Reference
|
|
41
41
|
|
|
42
|
-
Wiring: `$HOME/.claude/multi-agent-refs/phases/phase-
|
|
42
|
+
Wiring: `$HOME/.claude/multi-agent-refs/phases/phase-3-review.md` Step 3.7. Schema: `$HOME/.claude/schemas/triage-output.schema.json` v3.2.0 (`$defs.verification`). Evidence gate: `$HOME/.claude/scripts/evidence-gate.mjs`. Prefs: `prefs.global.verifyByTest` (including `repeatCount`) in `$HOME/.claude/schemas/prefs.schema.json`.
|
|
@@ -3,7 +3,7 @@
|
|
|
3
3
|
<!-- toc -->
|
|
4
4
|
- [1. When it is required](#1-when-it-is-required)
|
|
5
5
|
- [2. Before - the reporter's screenshot, or nothing](#2-before---the-reporters-screenshot-or-nothing)
|
|
6
|
-
- [3. After - Phase
|
|
6
|
+
- [3. After - Phase 2, not the Phase 3 user test](#3-after---phase-2-not-the-phase-3-user-test)
|
|
7
7
|
- [4. Video - the recording rides on a test run](#4-video---the-recording-rides-on-a-test-run)
|
|
8
8
|
- [5. Size, and what happens when it does not fit](#5-size-and-what-happens-when-it-does-not-fit)
|
|
9
9
|
- [5b. Where the artefacts live - the host](#5b-where-the-artefacts-live---the-host)
|
|
@@ -18,10 +18,10 @@ This contract makes the pipeline carry the picture: the state the reporter saw,
|
|
|
18
18
|
the state the fix produces, and where possible a recording of the flow running.
|
|
19
19
|
|
|
20
20
|
The artefacts live as **Jira attachments** and are referenced from two places -
|
|
21
|
-
the Phase
|
|
21
|
+
the Phase 5 Jira comment, where the picture belongs next to the work summary and
|
|
22
22
|
the test scenarios, and the PR body, which tells the reviewer they exist.
|
|
23
23
|
|
|
24
|
-
Consumers: Phase
|
|
24
|
+
Consumers: Phase 2 (capture), Phase 3 (video, user test), Phase 4 (blocker),
|
|
25
25
|
`channels/jira.md` and `channels/pr.md` (render). Gate: `smoke-visual-evidence.sh`.
|
|
26
26
|
|
|
27
27
|
## 1. When it is required
|
|
@@ -49,16 +49,16 @@ artefact is absent, the absence is written down with its reason (section 5).
|
|
|
49
49
|
|
|
50
50
|
**Phase 0 Step 7.7 writes `required`, `requiredBy` and `platform`**, before it
|
|
51
51
|
probes. Naming the writer is not a formality: five phases read these three fields
|
|
52
|
-
and nothing set them, so `required` was never true, Step 7.7 never ran, Phase
|
|
53
|
-
never captured, and Phase
|
|
52
|
+
and nothing set them, so `required` was never true, Step 7.7 never ran, Phase 2
|
|
53
|
+
never captured, and Phase 4's blocker never blocked. A contract with readers and
|
|
54
54
|
no writer reads exactly like a contract that is satisfied.
|
|
55
55
|
|
|
56
56
|
Phase 0 has no diff, so it decides on what it actually knows - `taskType` and
|
|
57
57
|
whether the intake carried a Figma reference - and that is enough for the whole
|
|
58
58
|
design-work row. The `bugfix` row needs a changed-file list, so Phase 0 records
|
|
59
|
-
it as provisional (`requiredBy: "bugfix (pending diff)"`) and **Phase
|
|
59
|
+
it as provisional (`requiredBy: "bugfix (pending diff)"`) and **Phase 2 re-decides
|
|
60
60
|
from the real diff** before capturing. A provisional `false` never suppresses the
|
|
61
|
-
Phase
|
|
61
|
+
Phase 2 decision.
|
|
62
62
|
|
|
63
63
|
### 1b. The platform, and when there is not one
|
|
64
64
|
|
|
@@ -120,15 +120,15 @@ No image on the ticket means no "before". Record the gap
|
|
|
120
120
|
(`before_missing: "ticket carries no image attachment"`) and continue. The
|
|
121
121
|
section still renders, saying so.
|
|
122
122
|
|
|
123
|
-
## 3. After - Phase
|
|
123
|
+
## 3. After - Phase 2, not the Phase 3 user test
|
|
124
124
|
|
|
125
|
-
|
|
126
|
-
every `autopilot` and `--local` entry** (`phase-
|
|
125
|
+
The user test is the natural home: the simulator is already up. It is also **dropped by
|
|
126
|
+
every `autopilot` and `--local` entry** (`phase-3-review.md` TLDR), so a capture
|
|
127
127
|
that lives only there produces nothing for unattended runs - which are exactly
|
|
128
128
|
the runs where nobody watched the screen.
|
|
129
129
|
|
|
130
|
-
The capture therefore happens in **Phase
|
|
131
|
-
which is in every mode's phase set. Phase
|
|
130
|
+
The capture therefore happens in **Phase 2, after the build+test gate passes**,
|
|
131
|
+
which is in every mode's phase set. The Phase 3 user test may add richer evidence on top; it is
|
|
132
132
|
never the only source.
|
|
133
133
|
|
|
134
134
|
```bash
|
|
@@ -192,8 +192,8 @@ The recorder itself is `capture-evidence.sh video start|stop`, which writes
|
|
|
192
192
|
`<task-id>[-<label>]-flow.mp4` into the evidence directory. The label is optional
|
|
193
193
|
because one recording per task is the common case, and both renderers cite the
|
|
194
194
|
unlabelled form. It is shell rather than MCP for three reasons: the sibling still capture already shells out,
|
|
195
|
-
a host with no toolkit MCP registered still produces evidence, and Phase
|
|
196
|
-
Phase
|
|
195
|
+
a host with no toolkit MCP registered still produces evidence, and Phase 2 and
|
|
196
|
+
Phase 3 are exactly where an MCP call is contested.
|
|
197
197
|
|
|
198
198
|
**iOS UI test targets are not named, they are detected.** A path containing
|
|
199
199
|
`UITests` is not the signal: in the reference app 477 files sit under such a path
|
|
@@ -210,8 +210,8 @@ wearing a measurement's clothes.
|
|
|
210
210
|
|
|
211
211
|
### 4.4 Re-check before recording
|
|
212
212
|
|
|
213
|
-
The probe runs at intake and the recording happens in Phase
|
|
214
|
-
then can be gone by the time the build goes green, so Phase
|
|
213
|
+
The probe runs at intake and the recording happens in Phase 2. A simulator booted
|
|
214
|
+
then can be gone by the time the build goes green, so Phase 2 re-measures the
|
|
215
215
|
device row alone and downgrades the tier if it has to, recording the transition
|
|
216
216
|
(`tier 1 -> 3: simulator no longer booted`). A tier taken from a stale
|
|
217
217
|
measurement is a promise the run cannot keep.
|
|
@@ -232,7 +232,7 @@ never asserted against wall clock anywhere, because doing so fails a correct
|
|
|
232
232
|
capture of a static screen.
|
|
233
233
|
|
|
234
234
|
`visualEvidence.enabled` turns the whole feature off - capture, upload, both
|
|
235
|
-
render sections and the Phase
|
|
235
|
+
render sections and the Phase 4 blocker with it.
|
|
236
236
|
|
|
237
237
|
### 4.6 Web - the recorder IS the runner
|
|
238
238
|
|
|
@@ -283,7 +283,7 @@ Never silently attach nothing.
|
|
|
283
283
|
|
|
284
284
|
## 5b. Where the artefacts live - the host
|
|
285
285
|
|
|
286
|
-
Resolved in Phase
|
|
286
|
+
Resolved in Phase 4 Step 2.9, recorded as `state.visualEvidence.host`:
|
|
287
287
|
|
|
288
288
|
| Order | Host | Condition | Stills | Video |
|
|
289
289
|
|---|---|---|---|---|
|
|
@@ -342,7 +342,7 @@ where it lives - the images themselves stay on the ticket:
|
|
|
342
342
|
## 7. Blocker
|
|
343
343
|
|
|
344
344
|
When section 1 says required and `state.visualEvidence` carries neither an
|
|
345
|
-
artefact nor a recorded reason for its absence, **Phase
|
|
345
|
+
artefact nor a recorded reason for its absence, **Phase 4 Step 3 blocks** - the
|
|
346
346
|
same shape as the `risk` section blocker. A recorded reason is enough to pass:
|
|
347
347
|
the gate is against silence, not against an honest "no image on the ticket".
|
|
348
348
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Worktree finalize - removing a task's worktree once its PR is open
|
|
2
2
|
|
|
3
|
-
> **TLDR** - Phase
|
|
3
|
+
> **TLDR** - Phase 4 step 9. Once the PR exists the worktree is dead weight, so it is removed: artefacts are salvaged into the log dir first, the branch is kept and deliberately NOT checked out, and every destructive path is gated. Gated by `prefs.global.settings.worktreeAutoRemoveOnPr` (default **true**). Script: `worktree-finalize.sh`.
|
|
4
4
|
|
|
5
5
|
## Why it exists
|
|
6
6
|
|
|
@@ -21,12 +21,12 @@ Removing it at PR-open is only safe because of the salvage, so the two are one s
|
|
|
21
21
|
| Condition | Why it blocks |
|
|
22
22
|
|---|---|
|
|
23
23
|
| `worktreePath == projectRoot` (`--local` mode) | there is no worktree; removing it would delete the user's checkout |
|
|
24
|
-
| cwd is inside the worktree | a shell left on a deleted inode is worse than a leftover directory, and Phase
|
|
24
|
+
| cwd is inside the worktree | a shell left on a deleted inode is worse than a leftover directory, and Phase 4 legitimately `cd`s into the worktree earlier |
|
|
25
25
|
| not a registered worktree of the project root | a mistyped path must not delete an unrelated directory |
|
|
26
26
|
| real uncommitted changes | never discarded; see the artefact carve-out below |
|
|
27
27
|
| HEAD not on the remote | removing a worktree whose commits exist nowhere else is data loss, not cleanup |
|
|
28
28
|
|
|
29
|
-
Exit codes: `0` removed, `3` skipped with a reason (report and continue to Phase
|
|
29
|
+
Exit codes: `0` removed, `3` skipped with a reason (report and continue to Phase 5), `1` usage error.
|
|
30
30
|
|
|
31
31
|
### The artefact carve-out, and why `--untracked-files=no` is wrong
|
|
32
32
|
|
|
@@ -49,10 +49,10 @@ This is why the removal is safe:
|
|
|
49
49
|
|
|
50
50
|
| Consumer | Reads | Without salvage |
|
|
51
51
|
|---|---|---|
|
|
52
|
-
| Phase
|
|
53
|
-
| Phase
|
|
52
|
+
| Phase 5 triage-memory ingest | `triage-output.json` | `[ -f ]`-guarded, so it degrades **silently**: the triage corpus and learnings ledger stop being fed and no error appears. Since v16.20.0 Phase 4 writes `triage-output.json` itself (Step 3.2.1, the latest copy of `.pipeline/triage-round-<N>.json`), so the salvage is a second copy, not the only bridge |
|
|
53
|
+
| Phase 5 learnings-ledger distill | same file | same silent degradation |
|
|
54
54
|
| `render-work-summary.sh` | `agent-state.json`, `phase-tracker.json` (falls back to `logs/multi-agent/<task>/tracker-state.json`, which survives removal) | loses the salvaged copies but keeps the tracker via the logs fallback |
|
|
55
|
-
| `:resume` | `agent-state.json` | cannot continue a Phase
|
|
55
|
+
| `:resume` | `agent-state.json` | cannot continue a Phase 5 pause |
|
|
56
56
|
| `:status`, `:log` | `agent-state.json` | the task becomes invisible |
|
|
57
57
|
|
|
58
58
|
`state.worktreeRemovedAt` and `state.artifactsPath` record the outcome. The timestamp is what tells a reader that a worktree-less task was finished-and-tidied rather than killed - without it, a missing worktree is indistinguishable from a broken run.
|
|
@@ -3,15 +3,15 @@
|
|
|
3
3
|
<!-- toc -->
|
|
4
4
|
- [The triad at a glance](#the-triad-at-a-glance)
|
|
5
5
|
- [Phase 0 auto-create policy](#phase-0-auto-create-policy)
|
|
6
|
-
- [Phase
|
|
6
|
+
- [Phase 5 wiki → Jira comment](#phase-5-wiki-jira-comment)
|
|
7
7
|
- [Autopilot behaviour](#autopilot-behaviour)
|
|
8
8
|
- [Preferences involved](#preferences-involved)
|
|
9
9
|
- [Cross-CLI parity](#cross-cli-parity)
|
|
10
10
|
<!-- /toc -->
|
|
11
11
|
|
|
12
|
-
> **TLDR** - When a GitHub issue triggers the pipeline and has no Jira ID, the `autoJiraFromGithubIssue` policy decides whether to auto-create a Jira task (and patch the GitHub issue body with the new Jira link). Phase
|
|
12
|
+
> **TLDR** - When a GitHub issue triggers the pipeline and has no Jira ID, the `autoJiraFromGithubIssue` policy decides whether to auto-create a Jira task (and patch the GitHub issue body with the new Jira link). Phase 5 then posts a humanizer'd wiki-content summary back as a Jira comment, closing the loop. Autopilot treats `ask` as `always`.
|
|
13
13
|
|
|
14
|
-
This doc is referenced from `$HOME/.claude/multi-agent-refs/phases/phase-0-init.md` Step 1 (GitHub issue input) and `$HOME/.claude/multi-agent-refs/phases/phase-
|
|
14
|
+
This doc is referenced from `$HOME/.claude/multi-agent-refs/phases/phase-0-init.md` Step 1 (GitHub issue input) and `$HOME/.claude/multi-agent-refs/phases/phase-5-report.md` Step 2 (component wiki). Keeps the phase docs tight and gives the triad contract a stable home.
|
|
15
15
|
|
|
16
16
|
## The triad at a glance
|
|
17
17
|
|
|
@@ -24,7 +24,7 @@ Jira issue (new, linked both ways)
|
|
|
24
24
|
▼ Phase 3 - component implementation
|
|
25
25
|
wiki markdown in pipeline/skills/...
|
|
26
26
|
│
|
|
27
|
-
▼ Phase
|
|
27
|
+
▼ Phase 5 Report Step 2 - wiki capture + Jira comment post
|
|
28
28
|
Jira comment (humanizer'd component summary)
|
|
29
29
|
```
|
|
30
30
|
|
|
@@ -59,13 +59,13 @@ Decision table:
|
|
|
59
59
|
5. On success: store new Jira key in `state.jiraId`. Emit `→ created Jira {newKey}`.
|
|
60
60
|
6. Patch the GitHub issue body: append a line `Jira: [{newKey}]({jira.baseUrl}/browse/{newKey})` so the bidirectional link is visible on GitHub without refetching the pipeline state.
|
|
61
61
|
|
|
62
|
-
**Skip path** (`never` or user declined): Branch falls back to `feature/GH{issueNo}-{kebab}`; Phase
|
|
62
|
+
**Skip path** (`never` or user declined): Branch falls back to `feature/GH{issueNo}-{kebab}`; Phase 5 later emits `→ no Jira linked (policy: {preference})`.
|
|
63
63
|
|
|
64
64
|
**Failure isolation:** Jira API 5xx or network error → log the error, warn the user, and **continue with no Jira link** (do not halt Phase 0). Rationale: a transient Jira outage should not block code work; user can link manually post-hoc.
|
|
65
65
|
|
|
66
|
-
## Phase
|
|
66
|
+
## Phase 5 wiki → Jira comment
|
|
67
67
|
|
|
68
|
-
Runs after Phase
|
|
68
|
+
Runs after Phase 5 Report Step 2 (wiki capture) has produced markdown and only when:
|
|
69
69
|
|
|
70
70
|
- `state.jiraId` is non-null (from Phase 0 scan or auto-create).
|
|
71
71
|
- `prefs.global.wikiToJiraComment !== false` (default true).
|
|
@@ -78,9 +78,9 @@ Flow:
|
|
|
78
78
|
3. Render as Jira comment: title line (`h3. Component docs - {componentName}`) + an overview paragraph + a link to the full page (wiki URL, if the adapter returned `pushedRemote`). The overview lines come from a Markdown file - convert them via the `channels/jira.md` table before POST, exactly like the title line already is, then run the whole body through `node "$HOME/.claude/scripts/jira-wiki-escape.mjs"` per that file's *Emoticon escaping* section.
|
|
79
79
|
4. **Run through the humanizer skill** - same policy as Step 4 Confluence: user-facing content must read naturally.
|
|
80
80
|
5. `POST {jira.baseUrl}/rest/api/2/issue/{jiraId}/comment`.
|
|
81
|
-
6. Log: `Phase
|
|
81
|
+
6. Log: `Phase 5: wiki summary posted to Jira {jiraId}`.
|
|
82
82
|
|
|
83
|
-
On failure: log + continue. Wiki is already written; the Jira comment is an augmentation. Don't regress Phase
|
|
83
|
+
On failure: log + continue. Wiki is already written; the Jira comment is an augmentation. Don't regress Phase 5 over a Jira 5xx.
|
|
84
84
|
|
|
85
85
|
## Autopilot behaviour
|
|
86
86
|
|
|
@@ -96,7 +96,7 @@ Rationale: autopilot is explicit consent for side-effect creation; muting Jira a
|
|
|
96
96
|
| key | type | default | controls |
|
|
97
97
|
|---|---|---|---|
|
|
98
98
|
| `prefs.global.autoJiraFromGithubIssue` | enum ask/always/never | `ask` | Phase 0 auto-create decision |
|
|
99
|
-
| `prefs.global.wikiToJiraComment` | bool | `true` | Phase
|
|
99
|
+
| `prefs.global.wikiToJiraComment` | bool | `true` | Phase 5 wiki → Jira comment post |
|
|
100
100
|
| `figmaConfig.jira.projectKey` | string | - | Target project for new issues (falls back to `prefs.global.defaultJiraKey`) |
|
|
101
101
|
| `prefs.global.keychainMapping.jira` | string | - | Jira token keychain lookup |
|
|
102
102
|
|
|
@@ -30,11 +30,11 @@ $HOME/.claude/knowledge/
|
|
|
30
30
|
### How It Works
|
|
31
31
|
|
|
32
32
|
```
|
|
33
|
-
Task 1 (first time): Phase 1 -> full Explore -> Phase
|
|
34
|
-
Task 2: Phase 1 -> read knowledge -> targeted Explore -> Phase
|
|
35
|
-
Task 3: Phase 1 -> read knowledge -> minimal Explore -> Phase
|
|
33
|
+
Task 1 (first time): Phase 1 -> full Explore -> Phase 5 -> create knowledge
|
|
34
|
+
Task 2: Phase 1 -> read knowledge -> targeted Explore -> Phase 5 -> update knowledge
|
|
35
|
+
Task 3: Phase 1 -> read knowledge -> minimal Explore -> Phase 5 -> update knowledge
|
|
36
36
|
...
|
|
37
|
-
Task N: Phase 1 -> knowledge sufficient -> SKIP Explore -> Phase
|
|
37
|
+
Task N: Phase 1 -> knowledge sufficient -> SKIP Explore -> Phase 5 -> update knowledge
|
|
38
38
|
```
|
|
39
39
|
|
|
40
40
|
Each task teaches the project a bit more. Token cost decreases over time.
|
|
@@ -45,7 +45,7 @@ Each task teaches the project a bit more. Token cost decreases over time.
|
|
|
45
45
|
| ------------------------------------------- | --------------------- | ------------------------------------------ | ------------------------ |
|
|
46
46
|
| `$PROJECT_ROOT/CLAUDE.md` | User | Project rules, build commands | Persistent (in git) |
|
|
47
47
|
| Auto Memory (`~/.claude/projects/`) | Claude (automatic) | User preferences, feedback | Persistent (global) |
|
|
48
|
-
| **Knowledge Base** (`~/.claude/knowledge/`) | Multi-agent (Phase
|
|
48
|
+
| **Knowledge Base** (`~/.claude/knowledge/`) | Multi-agent (Phase 5) | Architecture, patterns, gotchas, decisions | Persistent (per-project) |
|
|
49
49
|
|
|
50
50
|
They all complement each other - no conflicts.
|
|
51
51
|
|
|
@@ -54,7 +54,7 @@ They all complement each other - no conflicts.
|
|
|
54
54
|
A Short run skips Phase 1, but knowledge **is still read**:
|
|
55
55
|
|
|
56
56
|
- Knowledge files are added to the prompt before the Opus agent starts in Phase 3
|
|
57
|
-
- Knowledge capture is still performed in Phase
|
|
57
|
+
- Knowledge capture is still performed in Phase 5 (simplified report)
|
|
58
58
|
|
|
59
59
|
### Knowledge Maintenance
|
|
60
60
|
|
|
@@ -107,18 +107,18 @@ When calling sub-agents, inject task-specific context from previous phases:
|
|
|
107
107
|
### Phase -> Agent -> Context Flow
|
|
108
108
|
|
|
109
109
|
```
|
|
110
|
-
Phase 1:
|
|
110
|
+
Phase 1: Plan (analysis)
|
|
111
111
|
-- Explore agents (parallel) -> codebase scan
|
|
112
112
|
|
|
113
|
-
Phase
|
|
113
|
+
Phase 1: Plan
|
|
114
114
|
-- Architecture review agent
|
|
115
115
|
Context: task description + Phase 1 findings + affected modules
|
|
116
116
|
|
|
117
|
-
Phase
|
|
117
|
+
Phase 2: Dev
|
|
118
118
|
-- Sonnet agent (TDD)
|
|
119
|
-
Context: Phase
|
|
119
|
+
Context: Phase 1 plan + task dependencies
|
|
120
120
|
|
|
121
|
-
Phase
|
|
121
|
+
Phase 3: Review (CLI-aware parallel + triage)
|
|
122
122
|
Claude Code (3 parallel):
|
|
123
123
|
|-- Reviewer (opus) -> security + architecture
|
|
124
124
|
+-- Reviewer (sonnet) -> quality + correctness + edge cases
|