@mmerterden/multi-agent-pipeline 18.0.0 → 19.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +183 -0
- package/README.md +34 -18
- package/README.tr.md +14 -16
- package/docs/adr/0002-instruction-driven-flag.md +1 -0
- package/docs/adr/0005-lazy-phase-docs.md +11 -1
- package/docs/adr/0008-installer-modularization-and-secret-leak-defense.md +1 -0
- package/docs/adr/0010-own-code-graph.md +1 -0
- package/docs/adr/0014-six-phase-consolidation.md +134 -0
- package/docs/adr/README.md +2 -1
- package/docs/architecture.md +37 -38
- package/docs/best-practices.md +1 -1
- package/docs/ecosystem.md +37 -26
- package/docs/engineering.md +1 -1
- package/docs/facts.json +45 -0
- package/docs/features.md +54 -53
- package/docs/performance.md +5 -5
- package/docs/recovery-guide.md +9 -9
- package/docs/token-budget-history.md +3 -1
- package/index.js +2 -2
- package/install/_codex-agents.mjs +1 -1
- package/install/templates/claude-hooks.json +1 -1
- package/install/templates/codex-instructions.md +1 -1
- package/install/templates/copilot-instructions.md +28 -28
- package/manifest.json +209 -193
- package/package.json +2 -2
- package/pipeline/agents/dev-critic.md +3 -3
- package/pipeline/commands/figma-to-swiftui.md +1 -1
- package/pipeline/commands/multi-agent/SKILL.md +8 -8
- package/pipeline/commands/multi-agent/analysis/SKILL.md +9 -9
- package/pipeline/commands/multi-agent/autopilot/SKILL.md +7 -7
- package/pipeline/commands/multi-agent/channels/SKILL.md +15 -15
- package/pipeline/commands/multi-agent/diff-explain/SKILL.md +6 -6
- package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/graph/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/help/SKILL.md +62 -62
- package/pipeline/commands/multi-agent/language/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/local/SKILL.md +11 -11
- package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +13 -13
- package/pipeline/commands/multi-agent/log/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/manual-test/SKILL.md +9 -9
- package/pipeline/commands/multi-agent/model/SKILL.md +69 -0
- package/pipeline/commands/multi-agent/refactor/SKILL.md +3 -3
- package/pipeline/commands/multi-agent/resume/SKILL.md +4 -4
- package/pipeline/commands/multi-agent/resume-local/SKILL.md +19 -17
- package/pipeline/commands/multi-agent/review/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/route-off/SKILL.md +36 -0
- package/pipeline/commands/multi-agent/route-on/SKILL.md +74 -0
- package/pipeline/commands/multi-agent/route-status/SKILL.md +56 -0
- package/pipeline/commands/multi-agent/setup/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/status/SKILL.md +5 -5
- package/pipeline/commands/multi-agent/steer/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/sync/SKILL.md +12 -13
- package/pipeline/commands/multi-agent/test/SKILL.md +1 -1
- package/pipeline/lib/credential-inventory.sh +1 -1
- package/pipeline/lib/fetch-fortify.sh +1 -1
- package/pipeline/lib/model-rung.sh +142 -0
- package/pipeline/lib/phase-schema.mjs +88 -0
- package/pipeline/lib/plan-todos.sh +5 -5
- package/pipeline/lib/route-state.sh +161 -0
- package/pipeline/lib/run-paths.sh +2 -2
- package/pipeline/multi-agent-refs/_account-picker.md +1 -1
- package/pipeline/multi-agent-refs/_dev-context.md +1 -1
- package/pipeline/multi-agent-refs/_input-parser.md +1 -1
- package/pipeline/multi-agent-refs/analysis/evidence.md +0 -9
- package/pipeline/multi-agent-refs/analysis/intake.md +1 -1
- package/pipeline/multi-agent-refs/analysis/locked.md +21 -22
- package/pipeline/multi-agent-refs/analysis/render.md +1 -1
- package/pipeline/multi-agent-refs/analysis/synthesis.md +12 -6
- package/pipeline/multi-agent-refs/android-guide.md +1 -1
- package/pipeline/multi-agent-refs/audit-guide.md +13 -13
- package/pipeline/multi-agent-refs/channels/issue-comment.md +2 -2
- package/pipeline/multi-agent-refs/channels/jira.md +3 -3
- package/pipeline/multi-agent-refs/channels/pr.md +4 -4
- package/pipeline/multi-agent-refs/channels/wiki.md +1 -1
- package/pipeline/multi-agent-refs/component-dispatch.md +3 -3
- package/pipeline/multi-agent-refs/cross-cli-contract.md +31 -6
- package/pipeline/multi-agent-refs/features/autopilot-circuit-breaker.md +4 -4
- package/pipeline/multi-agent-refs/features/code-graph.md +5 -5
- package/pipeline/multi-agent-refs/features/design-conformance.md +1 -1
- package/pipeline/multi-agent-refs/features/dev-critic.md +3 -3
- package/pipeline/multi-agent-refs/features/doctor.md +2 -2
- package/pipeline/multi-agent-refs/features/external-context-injection.md +3 -3
- package/pipeline/multi-agent-refs/features/maturity-followup.md +3 -3
- package/pipeline/multi-agent-refs/features/model-fallback.md +5 -5
- package/pipeline/multi-agent-refs/features/plan-todos.md +1 -1
- package/pipeline/multi-agent-refs/features/repo-map.md +1 -1
- package/pipeline/multi-agent-refs/features/review-delta.md +3 -3
- package/pipeline/multi-agent-refs/features/review-multi-repo.md +1 -1
- package/pipeline/multi-agent-refs/features/scope-check.md +4 -4
- package/pipeline/multi-agent-refs/features/skill-conformance.md +2 -2
- package/pipeline/multi-agent-refs/features/stack-skill-routing.md +1 -1
- package/pipeline/multi-agent-refs/features/verify-by-test.md +4 -4
- package/pipeline/multi-agent-refs/features/visual-evidence.md +19 -19
- package/pipeline/multi-agent-refs/features/worktree-finalize.md +6 -6
- package/pipeline/multi-agent-refs/issue-jira-triad.md +10 -10
- package/pipeline/multi-agent-refs/knowledge.md +11 -11
- package/pipeline/multi-agent-refs/multi-repo-integration-build.md +13 -13
- package/pipeline/multi-agent-refs/payload-contracts.md +8 -8
- package/pipeline/multi-agent-refs/phases/log-format.md +10 -10
- package/pipeline/multi-agent-refs/phases/modes.md +30 -30
- package/pipeline/multi-agent-refs/phases/operations.md +8 -8
- package/pipeline/multi-agent-refs/phases/phase-0-init.md +24 -24
- package/pipeline/multi-agent-refs/phases/phase-1-plan.md +599 -0
- package/pipeline/multi-agent-refs/phases/{phase-3-dev.md → phase-2-dev.md} +129 -49
- package/pipeline/multi-agent-refs/phases/{phase-4-review.md → phase-3-review.md} +225 -107
- package/pipeline/multi-agent-refs/phases/{phase-6-commit.md → phase-4-commit.md} +23 -23
- package/pipeline/multi-agent-refs/phases/{phase-7-report.md → phase-5-report.md} +29 -29
- package/pipeline/multi-agent-refs/phases.md +44 -48
- package/pipeline/multi-agent-refs/picker-contract.md +1 -1
- package/pipeline/multi-agent-refs/progress-contract.md +6 -6
- package/pipeline/multi-agent-refs/readiness-review.md +1 -1
- package/pipeline/multi-agent-refs/rules.md +7 -7
- package/pipeline/multi-agent-refs/swiftui-guide.md +2 -2
- package/pipeline/multi-agent-refs/tracker-contract.md +31 -32
- package/pipeline/multi-agent-refs/wiki-capture.md +14 -14
- package/pipeline/preferences-template.json +9 -1
- package/pipeline/rules/outside-the-pipeline.md +1 -1
- package/pipeline/schemas/agent-state.schema.json +50 -50
- package/pipeline/schemas/analysis-output.schema.json +2 -2
- package/pipeline/schemas/autopilot-config.schema.json +1 -1
- package/pipeline/schemas/code-graph.schema.json +1 -1
- package/pipeline/schemas/criteria-manifest.schema.json +1 -1
- package/pipeline/schemas/dev-critic-output.schema.json +1 -1
- package/pipeline/schemas/diff-risk.schema.json +1 -1
- package/pipeline/schemas/migrations/prefs-2.4.0-to-2.5.0.mjs +2 -2
- package/pipeline/schemas/migrations/prefs-2.6.0-to-2.7.0.mjs +31 -0
- package/pipeline/schemas/migrations/state-2.1.0-to-2.2.0.mjs +129 -0
- package/pipeline/schemas/phases.json +105 -0
- package/pipeline/schemas/plan-todos.schema.json +5 -5
- package/pipeline/schemas/planning-output.schema.json +1 -1
- package/pipeline/schemas/prefs.schema.json +100 -56
- package/pipeline/schemas/reviewer-output.schema.json +3 -3
- package/pipeline/schemas/route-config.schema.json +74 -0
- package/pipeline/schemas/scope-check.schema.json +1 -1
- package/pipeline/schemas/test-gap.schema.json +1 -1
- package/pipeline/schemas/token-budget.json +12 -18
- package/pipeline/schemas/triage-output.schema.json +6 -6
- package/pipeline/scripts/README.md +3 -3
- package/pipeline/scripts/_code-graph.mjs +2 -2
- package/pipeline/scripts/_run-paths.mjs +2 -2
- package/pipeline/scripts/_smoke-root.sh +1 -1
- package/pipeline/scripts/aggregate-metrics.mjs +1 -1
- package/pipeline/scripts/capture-flush.sh +8 -8
- package/pipeline/scripts/capture-resume.sh +3 -3
- package/pipeline/scripts/classify-plan-safety.mjs +1 -1
- package/pipeline/scripts/diff-explain.mjs +1 -1
- package/pipeline/scripts/doctor.mjs +2 -2
- package/pipeline/scripts/gc-abandoned.sh +3 -3
- package/pipeline/scripts/gc-tmp.sh +1 -1
- package/pipeline/scripts/gc-worktrees.sh +1 -1
- package/pipeline/scripts/gen-facts.mjs +175 -0
- package/pipeline/scripts/gen-mode-dispatch.mjs +32 -37
- package/pipeline/scripts/gen-ref-toc.mjs +1 -1
- package/pipeline/scripts/graph-report.mjs +1 -1
- package/pipeline/scripts/jira-attach.sh +1 -1
- package/pipeline/scripts/learn-from-transcripts.mjs +1 -1
- package/pipeline/scripts/learning-curve.mjs +2 -2
- package/pipeline/scripts/log-metric.sh +17 -4
- package/pipeline/scripts/memory-save.sh +1 -1
- package/pipeline/scripts/migrate-prefs.mjs +22 -5
- package/pipeline/scripts/phase-banner.sh +20 -20
- package/pipeline/scripts/phase-tracker.sh +7 -7
- package/pipeline/scripts/plan-coverage-gate.mjs +2 -2
- package/pipeline/scripts/render-agent-log-cost.sh +1 -1
- package/pipeline/scripts/render-work-summary.sh +3 -3
- package/pipeline/scripts/review-file-filter.mjs +1 -1
- package/pipeline/scripts/run-aggregator.mjs +13 -6
- package/pipeline/scripts/run-metrics.mjs +1 -1
- package/pipeline/scripts/runs-index.mjs +11 -1
- package/pipeline/scripts/smoke-cross-cli-behavior.sh +6 -6
- package/pipeline/scripts/smoke-schema-validation.sh +26 -7
- package/pipeline/scripts/token-budget-report.mjs +13 -2
- package/pipeline/scripts/triage-memory.mjs +2 -2
- package/pipeline/scripts/validate-analysis-doc.mjs +73 -17
- package/pipeline/scripts/validate-planning.mjs +1 -1
- package/pipeline/scripts/validate-reviewer.mjs +1 -1
- package/pipeline/scripts/validate-state.mjs +45 -5
- package/pipeline/scripts/validate-triage.mjs +3 -3
- package/pipeline/scripts/worktree-finalize.sh +5 -5
- package/pipeline/skills/.skill-manifest.json +37 -21
- package/pipeline/skills/.skills-index.json +49 -5
- package/pipeline/skills/shared/README.md +10 -6
- package/pipeline/skills/shared/core/apple-archive-compliance/SKILL.md +2 -2
- package/pipeline/skills/shared/core/google-play-compliance/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent/SKILL.md +69 -71
- package/pipeline/skills/shared/core/multi-agent-autopilot/SKILL.md +3 -3
- package/pipeline/skills/shared/core/multi-agent-channels/SKILL.md +14 -14
- package/pipeline/skills/shared/core/multi-agent-diff-explain/SKILL.md +5 -5
- package/pipeline/skills/shared/core/multi-agent-graph/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +25 -23
- package/pipeline/skills/shared/core/multi-agent-language/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-local/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +8 -8
- package/pipeline/skills/shared/core/multi-agent-manual-test/SKILL.md +6 -6
- package/pipeline/skills/shared/core/multi-agent-model/SKILL.md +71 -0
- package/pipeline/skills/shared/core/multi-agent-refactor/SKILL.md +3 -3
- package/pipeline/skills/shared/core/multi-agent-resume/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-resume-local/SKILL.md +7 -7
- package/pipeline/skills/shared/core/multi-agent-route-off/SKILL.md +39 -0
- package/pipeline/skills/shared/core/multi-agent-route-on/SKILL.md +76 -0
- package/pipeline/skills/shared/core/multi-agent-route-status/SKILL.md +59 -0
- package/pipeline/skills/shared/core/multi-agent-setup/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-status/SKILL.md +5 -5
- package/pipeline/skills/shared/core/multi-agent-steer/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +6 -5
- package/pipeline/skills/skills-index.md +8 -4
- package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +0 -263
- package/pipeline/multi-agent-refs/phases/phase-2-planning.md +0 -344
- package/pipeline/multi-agent-refs/phases/phase-5-test.md +0 -182
|
@@ -1,4 +1,4 @@
|
|
|
1
|
-
# Feature: Verify-by-Test Triage (Phase
|
|
1
|
+
# Feature: Verify-by-Test Triage (Phase 3 Step 3.7)
|
|
2
2
|
|
|
3
3
|
**Pattern**: reviewer findings are hypotheses; a failing repro test is proof. Adversarial-review research (Refute-or-Promote, 2026) found that plausible-but-wrong findings survive debate rounds but die on a single empirical test. This step converts the highest-stakes verdicts (accepted blocking) from judgment into evidence before the Phase 3 rework loop fires.
|
|
4
4
|
|
|
@@ -19,7 +19,7 @@
|
|
|
19
19
|
| Passes on some runs, fails on others | `inconclusive` | Flake, not proof either way. Finding stays accepted blocking, `verification.note = "flaky: passed k/N"`, repro test file deleted, telemetry line gains `flaky=<count>`. |
|
|
20
20
|
| Compile error / timeout / not unit-testable | `inconclusive` | Finding stays accepted blocking (judgment stands). Partial test deleted, cause in `verification.note`. |
|
|
21
21
|
|
|
22
|
-
Downgrades go to `deferred`, never `rejected`: triage judged the issue real, and deferred items surface in the Phase
|
|
22
|
+
Downgrades go to `deferred`, never `rejected`: triage judged the issue real, and deferred items surface in the Phase 5 report for a human eye.
|
|
23
23
|
|
|
24
24
|
## Red-test handoff to Phase 3
|
|
25
25
|
|
|
@@ -27,7 +27,7 @@ Downgrades go to `deferred`, never `rejected`: triage judged the issue real, and
|
|
|
27
27
|
|
|
28
28
|
## Cleanup invariant
|
|
29
29
|
|
|
30
|
-
After Step 3.7, the only uncommitted verifier artifacts are the confirmed repro tests listed in `redTests[]` (committed later with the fix) and logs under `$WORKTREE/.pipeline/` (outside Phase
|
|
30
|
+
After Step 3.7, the only uncommitted verifier artifacts are the confirmed repro tests listed in `redTests[]` (committed later with the fix) and logs under `$WORKTREE/.pipeline/` (outside Phase 4 commit scope). `not-reproduced` and `inconclusive` test files are always deleted.
|
|
31
31
|
|
|
32
32
|
## Telemetry
|
|
33
33
|
|
|
@@ -39,4 +39,4 @@ Adds one Sonnet call plus up to `maxFindings` single-test runs (and build-lock c
|
|
|
39
39
|
|
|
40
40
|
## Reference
|
|
41
41
|
|
|
42
|
-
Wiring: `$HOME/.claude/multi-agent-refs/phases/phase-
|
|
42
|
+
Wiring: `$HOME/.claude/multi-agent-refs/phases/phase-3-review.md` Step 3.7. Schema: `$HOME/.claude/schemas/triage-output.schema.json` v3.2.0 (`$defs.verification`). Evidence gate: `$HOME/.claude/scripts/evidence-gate.mjs`. Prefs: `prefs.global.verifyByTest` (including `repeatCount`) in `$HOME/.claude/schemas/prefs.schema.json`.
|
|
@@ -3,7 +3,7 @@
|
|
|
3
3
|
<!-- toc -->
|
|
4
4
|
- [1. When it is required](#1-when-it-is-required)
|
|
5
5
|
- [2. Before - the reporter's screenshot, or nothing](#2-before---the-reporters-screenshot-or-nothing)
|
|
6
|
-
- [3. After - Phase
|
|
6
|
+
- [3. After - Phase 2, not the Phase 3 user test](#3-after---phase-2-not-the-phase-3-user-test)
|
|
7
7
|
- [4. Video - the recording rides on a test run](#4-video---the-recording-rides-on-a-test-run)
|
|
8
8
|
- [5. Size, and what happens when it does not fit](#5-size-and-what-happens-when-it-does-not-fit)
|
|
9
9
|
- [5b. Where the artefacts live - the host](#5b-where-the-artefacts-live---the-host)
|
|
@@ -18,10 +18,10 @@ This contract makes the pipeline carry the picture: the state the reporter saw,
|
|
|
18
18
|
the state the fix produces, and where possible a recording of the flow running.
|
|
19
19
|
|
|
20
20
|
The artefacts live as **Jira attachments** and are referenced from two places -
|
|
21
|
-
the Phase
|
|
21
|
+
the Phase 5 Jira comment, where the picture belongs next to the work summary and
|
|
22
22
|
the test scenarios, and the PR body, which tells the reviewer they exist.
|
|
23
23
|
|
|
24
|
-
Consumers: Phase
|
|
24
|
+
Consumers: Phase 2 (capture), Phase 3 (video, user test), Phase 4 (blocker),
|
|
25
25
|
`channels/jira.md` and `channels/pr.md` (render). Gate: `smoke-visual-evidence.sh`.
|
|
26
26
|
|
|
27
27
|
## 1. When it is required
|
|
@@ -49,16 +49,16 @@ artefact is absent, the absence is written down with its reason (section 5).
|
|
|
49
49
|
|
|
50
50
|
**Phase 0 Step 7.7 writes `required`, `requiredBy` and `platform`**, before it
|
|
51
51
|
probes. Naming the writer is not a formality: five phases read these three fields
|
|
52
|
-
and nothing set them, so `required` was never true, Step 7.7 never ran, Phase
|
|
53
|
-
never captured, and Phase
|
|
52
|
+
and nothing set them, so `required` was never true, Step 7.7 never ran, Phase 2
|
|
53
|
+
never captured, and Phase 4's blocker never blocked. A contract with readers and
|
|
54
54
|
no writer reads exactly like a contract that is satisfied.
|
|
55
55
|
|
|
56
56
|
Phase 0 has no diff, so it decides on what it actually knows - `taskType` and
|
|
57
57
|
whether the intake carried a Figma reference - and that is enough for the whole
|
|
58
58
|
design-work row. The `bugfix` row needs a changed-file list, so Phase 0 records
|
|
59
|
-
it as provisional (`requiredBy: "bugfix (pending diff)"`) and **Phase
|
|
59
|
+
it as provisional (`requiredBy: "bugfix (pending diff)"`) and **Phase 2 re-decides
|
|
60
60
|
from the real diff** before capturing. A provisional `false` never suppresses the
|
|
61
|
-
Phase
|
|
61
|
+
Phase 2 decision.
|
|
62
62
|
|
|
63
63
|
### 1b. The platform, and when there is not one
|
|
64
64
|
|
|
@@ -120,15 +120,15 @@ No image on the ticket means no "before". Record the gap
|
|
|
120
120
|
(`before_missing: "ticket carries no image attachment"`) and continue. The
|
|
121
121
|
section still renders, saying so.
|
|
122
122
|
|
|
123
|
-
## 3. After - Phase
|
|
123
|
+
## 3. After - Phase 2, not the Phase 3 user test
|
|
124
124
|
|
|
125
|
-
|
|
126
|
-
every `autopilot` and `--local` entry** (`phase-
|
|
125
|
+
The user test is the natural home: the simulator is already up. It is also **dropped by
|
|
126
|
+
every `autopilot` and `--local` entry** (`phase-3-review.md` TLDR), so a capture
|
|
127
127
|
that lives only there produces nothing for unattended runs - which are exactly
|
|
128
128
|
the runs where nobody watched the screen.
|
|
129
129
|
|
|
130
|
-
The capture therefore happens in **Phase
|
|
131
|
-
which is in every mode's phase set. Phase
|
|
130
|
+
The capture therefore happens in **Phase 2, after the build+test gate passes**,
|
|
131
|
+
which is in every mode's phase set. The Phase 3 user test may add richer evidence on top; it is
|
|
132
132
|
never the only source.
|
|
133
133
|
|
|
134
134
|
```bash
|
|
@@ -192,8 +192,8 @@ The recorder itself is `capture-evidence.sh video start|stop`, which writes
|
|
|
192
192
|
`<task-id>[-<label>]-flow.mp4` into the evidence directory. The label is optional
|
|
193
193
|
because one recording per task is the common case, and both renderers cite the
|
|
194
194
|
unlabelled form. It is shell rather than MCP for three reasons: the sibling still capture already shells out,
|
|
195
|
-
a host with no toolkit MCP registered still produces evidence, and Phase
|
|
196
|
-
Phase
|
|
195
|
+
a host with no toolkit MCP registered still produces evidence, and Phase 2 and
|
|
196
|
+
Phase 3 are exactly where an MCP call is contested.
|
|
197
197
|
|
|
198
198
|
**iOS UI test targets are not named, they are detected.** A path containing
|
|
199
199
|
`UITests` is not the signal: in the reference app 477 files sit under such a path
|
|
@@ -210,8 +210,8 @@ wearing a measurement's clothes.
|
|
|
210
210
|
|
|
211
211
|
### 4.4 Re-check before recording
|
|
212
212
|
|
|
213
|
-
The probe runs at intake and the recording happens in Phase
|
|
214
|
-
then can be gone by the time the build goes green, so Phase
|
|
213
|
+
The probe runs at intake and the recording happens in Phase 2. A simulator booted
|
|
214
|
+
then can be gone by the time the build goes green, so Phase 2 re-measures the
|
|
215
215
|
device row alone and downgrades the tier if it has to, recording the transition
|
|
216
216
|
(`tier 1 -> 3: simulator no longer booted`). A tier taken from a stale
|
|
217
217
|
measurement is a promise the run cannot keep.
|
|
@@ -232,7 +232,7 @@ never asserted against wall clock anywhere, because doing so fails a correct
|
|
|
232
232
|
capture of a static screen.
|
|
233
233
|
|
|
234
234
|
`visualEvidence.enabled` turns the whole feature off - capture, upload, both
|
|
235
|
-
render sections and the Phase
|
|
235
|
+
render sections and the Phase 4 blocker with it.
|
|
236
236
|
|
|
237
237
|
### 4.6 Web - the recorder IS the runner
|
|
238
238
|
|
|
@@ -283,7 +283,7 @@ Never silently attach nothing.
|
|
|
283
283
|
|
|
284
284
|
## 5b. Where the artefacts live - the host
|
|
285
285
|
|
|
286
|
-
Resolved in Phase
|
|
286
|
+
Resolved in Phase 4 Step 2.9, recorded as `state.visualEvidence.host`:
|
|
287
287
|
|
|
288
288
|
| Order | Host | Condition | Stills | Video |
|
|
289
289
|
|---|---|---|---|---|
|
|
@@ -342,7 +342,7 @@ where it lives - the images themselves stay on the ticket:
|
|
|
342
342
|
## 7. Blocker
|
|
343
343
|
|
|
344
344
|
When section 1 says required and `state.visualEvidence` carries neither an
|
|
345
|
-
artefact nor a recorded reason for its absence, **Phase
|
|
345
|
+
artefact nor a recorded reason for its absence, **Phase 4 Step 3 blocks** - the
|
|
346
346
|
same shape as the `risk` section blocker. A recorded reason is enough to pass:
|
|
347
347
|
the gate is against silence, not against an honest "no image on the ticket".
|
|
348
348
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Worktree finalize - removing a task's worktree once its PR is open
|
|
2
2
|
|
|
3
|
-
> **TLDR** - Phase
|
|
3
|
+
> **TLDR** - Phase 4 step 9. Once the PR exists the worktree is dead weight, so it is removed: artefacts are salvaged into the log dir first, the branch is kept and deliberately NOT checked out, and every destructive path is gated. Gated by `prefs.global.settings.worktreeAutoRemoveOnPr` (default **true**). Script: `worktree-finalize.sh`.
|
|
4
4
|
|
|
5
5
|
## Why it exists
|
|
6
6
|
|
|
@@ -21,12 +21,12 @@ Removing it at PR-open is only safe because of the salvage, so the two are one s
|
|
|
21
21
|
| Condition | Why it blocks |
|
|
22
22
|
|---|---|
|
|
23
23
|
| `worktreePath == projectRoot` (`--local` mode) | there is no worktree; removing it would delete the user's checkout |
|
|
24
|
-
| cwd is inside the worktree | a shell left on a deleted inode is worse than a leftover directory, and Phase
|
|
24
|
+
| cwd is inside the worktree | a shell left on a deleted inode is worse than a leftover directory, and Phase 4 legitimately `cd`s into the worktree earlier |
|
|
25
25
|
| not a registered worktree of the project root | a mistyped path must not delete an unrelated directory |
|
|
26
26
|
| real uncommitted changes | never discarded; see the artefact carve-out below |
|
|
27
27
|
| HEAD not on the remote | removing a worktree whose commits exist nowhere else is data loss, not cleanup |
|
|
28
28
|
|
|
29
|
-
Exit codes: `0` removed, `3` skipped with a reason (report and continue to Phase
|
|
29
|
+
Exit codes: `0` removed, `3` skipped with a reason (report and continue to Phase 5), `1` usage error.
|
|
30
30
|
|
|
31
31
|
### The artefact carve-out, and why `--untracked-files=no` is wrong
|
|
32
32
|
|
|
@@ -49,10 +49,10 @@ This is why the removal is safe:
|
|
|
49
49
|
|
|
50
50
|
| Consumer | Reads | Without salvage |
|
|
51
51
|
|---|---|---|
|
|
52
|
-
| Phase
|
|
53
|
-
| Phase
|
|
52
|
+
| Phase 5 triage-memory ingest | `triage-output.json` | `[ -f ]`-guarded, so it degrades **silently**: the triage corpus and learnings ledger stop being fed and no error appears. Since v16.20.0 Phase 4 writes `triage-output.json` itself (Step 3.2.1, the latest copy of `.pipeline/triage-round-<N>.json`), so the salvage is a second copy, not the only bridge |
|
|
53
|
+
| Phase 5 learnings-ledger distill | same file | same silent degradation |
|
|
54
54
|
| `render-work-summary.sh` | `agent-state.json`, `phase-tracker.json` (falls back to `logs/multi-agent/<task>/tracker-state.json`, which survives removal) | loses the salvaged copies but keeps the tracker via the logs fallback |
|
|
55
|
-
| `:resume` | `agent-state.json` | cannot continue a Phase
|
|
55
|
+
| `:resume` | `agent-state.json` | cannot continue a Phase 5 pause |
|
|
56
56
|
| `:status`, `:log` | `agent-state.json` | the task becomes invisible |
|
|
57
57
|
|
|
58
58
|
`state.worktreeRemovedAt` and `state.artifactsPath` record the outcome. The timestamp is what tells a reader that a worktree-less task was finished-and-tidied rather than killed - without it, a missing worktree is indistinguishable from a broken run.
|
|
@@ -3,15 +3,15 @@
|
|
|
3
3
|
<!-- toc -->
|
|
4
4
|
- [The triad at a glance](#the-triad-at-a-glance)
|
|
5
5
|
- [Phase 0 auto-create policy](#phase-0-auto-create-policy)
|
|
6
|
-
- [Phase
|
|
6
|
+
- [Phase 5 wiki → Jira comment](#phase-5-wiki-jira-comment)
|
|
7
7
|
- [Autopilot behaviour](#autopilot-behaviour)
|
|
8
8
|
- [Preferences involved](#preferences-involved)
|
|
9
9
|
- [Cross-CLI parity](#cross-cli-parity)
|
|
10
10
|
<!-- /toc -->
|
|
11
11
|
|
|
12
|
-
> **TLDR** - When a GitHub issue triggers the pipeline and has no Jira ID, the `autoJiraFromGithubIssue` policy decides whether to auto-create a Jira task (and patch the GitHub issue body with the new Jira link). Phase
|
|
12
|
+
> **TLDR** - When a GitHub issue triggers the pipeline and has no Jira ID, the `autoJiraFromGithubIssue` policy decides whether to auto-create a Jira task (and patch the GitHub issue body with the new Jira link). Phase 5 then posts a humanizer'd wiki-content summary back as a Jira comment, closing the loop. Autopilot treats `ask` as `always`.
|
|
13
13
|
|
|
14
|
-
This doc is referenced from `$HOME/.claude/multi-agent-refs/phases/phase-0-init.md` Step 1 (GitHub issue input) and `$HOME/.claude/multi-agent-refs/phases/phase-
|
|
14
|
+
This doc is referenced from `$HOME/.claude/multi-agent-refs/phases/phase-0-init.md` Step 1 (GitHub issue input) and `$HOME/.claude/multi-agent-refs/phases/phase-5-report.md` Step 2 (component wiki). Keeps the phase docs tight and gives the triad contract a stable home.
|
|
15
15
|
|
|
16
16
|
## The triad at a glance
|
|
17
17
|
|
|
@@ -24,7 +24,7 @@ Jira issue (new, linked both ways)
|
|
|
24
24
|
▼ Phase 3 - component implementation
|
|
25
25
|
wiki markdown in pipeline/skills/...
|
|
26
26
|
│
|
|
27
|
-
▼ Phase
|
|
27
|
+
▼ Phase 5 Report Step 2 - wiki capture + Jira comment post
|
|
28
28
|
Jira comment (humanizer'd component summary)
|
|
29
29
|
```
|
|
30
30
|
|
|
@@ -59,13 +59,13 @@ Decision table:
|
|
|
59
59
|
5. On success: store new Jira key in `state.jiraId`. Emit `→ created Jira {newKey}`.
|
|
60
60
|
6. Patch the GitHub issue body: append a line `Jira: [{newKey}]({jira.baseUrl}/browse/{newKey})` so the bidirectional link is visible on GitHub without refetching the pipeline state.
|
|
61
61
|
|
|
62
|
-
**Skip path** (`never` or user declined): Branch falls back to `feature/GH{issueNo}-{kebab}`; Phase
|
|
62
|
+
**Skip path** (`never` or user declined): Branch falls back to `feature/GH{issueNo}-{kebab}`; Phase 5 later emits `→ no Jira linked (policy: {preference})`.
|
|
63
63
|
|
|
64
64
|
**Failure isolation:** Jira API 5xx or network error → log the error, warn the user, and **continue with no Jira link** (do not halt Phase 0). Rationale: a transient Jira outage should not block code work; user can link manually post-hoc.
|
|
65
65
|
|
|
66
|
-
## Phase
|
|
66
|
+
## Phase 5 wiki → Jira comment
|
|
67
67
|
|
|
68
|
-
Runs after Phase
|
|
68
|
+
Runs after Phase 5 Report Step 2 (wiki capture) has produced markdown and only when:
|
|
69
69
|
|
|
70
70
|
- `state.jiraId` is non-null (from Phase 0 scan or auto-create).
|
|
71
71
|
- `prefs.global.wikiToJiraComment !== false` (default true).
|
|
@@ -78,9 +78,9 @@ Flow:
|
|
|
78
78
|
3. Render as Jira comment: title line (`h3. Component docs - {componentName}`) + an overview paragraph + a link to the full page (wiki URL, if the adapter returned `pushedRemote`). The overview lines come from a Markdown file - convert them via the `channels/jira.md` table before POST, exactly like the title line already is, then run the whole body through `node "$HOME/.claude/scripts/jira-wiki-escape.mjs"` per that file's *Emoticon escaping* section.
|
|
79
79
|
4. **Run through the humanizer skill** - same policy as Step 4 Confluence: user-facing content must read naturally.
|
|
80
80
|
5. `POST {jira.baseUrl}/rest/api/2/issue/{jiraId}/comment`.
|
|
81
|
-
6. Log: `Phase
|
|
81
|
+
6. Log: `Phase 5: wiki summary posted to Jira {jiraId}`.
|
|
82
82
|
|
|
83
|
-
On failure: log + continue. Wiki is already written; the Jira comment is an augmentation. Don't regress Phase
|
|
83
|
+
On failure: log + continue. Wiki is already written; the Jira comment is an augmentation. Don't regress Phase 5 over a Jira 5xx.
|
|
84
84
|
|
|
85
85
|
## Autopilot behaviour
|
|
86
86
|
|
|
@@ -96,7 +96,7 @@ Rationale: autopilot is explicit consent for side-effect creation; muting Jira a
|
|
|
96
96
|
| key | type | default | controls |
|
|
97
97
|
|---|---|---|---|
|
|
98
98
|
| `prefs.global.autoJiraFromGithubIssue` | enum ask/always/never | `ask` | Phase 0 auto-create decision |
|
|
99
|
-
| `prefs.global.wikiToJiraComment` | bool | `true` | Phase
|
|
99
|
+
| `prefs.global.wikiToJiraComment` | bool | `true` | Phase 5 wiki → Jira comment post |
|
|
100
100
|
| `figmaConfig.jira.projectKey` | string | - | Target project for new issues (falls back to `prefs.global.defaultJiraKey`) |
|
|
101
101
|
| `prefs.global.keychainMapping.jira` | string | - | Jira token keychain lookup |
|
|
102
102
|
|
|
@@ -30,11 +30,11 @@ $HOME/.claude/knowledge/
|
|
|
30
30
|
### How It Works
|
|
31
31
|
|
|
32
32
|
```
|
|
33
|
-
Task 1 (first time): Phase 1 -> full Explore -> Phase
|
|
34
|
-
Task 2: Phase 1 -> read knowledge -> targeted Explore -> Phase
|
|
35
|
-
Task 3: Phase 1 -> read knowledge -> minimal Explore -> Phase
|
|
33
|
+
Task 1 (first time): Phase 1 -> full Explore -> Phase 5 -> create knowledge
|
|
34
|
+
Task 2: Phase 1 -> read knowledge -> targeted Explore -> Phase 5 -> update knowledge
|
|
35
|
+
Task 3: Phase 1 -> read knowledge -> minimal Explore -> Phase 5 -> update knowledge
|
|
36
36
|
...
|
|
37
|
-
Task N: Phase 1 -> knowledge sufficient -> SKIP Explore -> Phase
|
|
37
|
+
Task N: Phase 1 -> knowledge sufficient -> SKIP Explore -> Phase 5 -> update knowledge
|
|
38
38
|
```
|
|
39
39
|
|
|
40
40
|
Each task teaches the project a bit more. Token cost decreases over time.
|
|
@@ -45,7 +45,7 @@ Each task teaches the project a bit more. Token cost decreases over time.
|
|
|
45
45
|
| ------------------------------------------- | --------------------- | ------------------------------------------ | ------------------------ |
|
|
46
46
|
| `$PROJECT_ROOT/CLAUDE.md` | User | Project rules, build commands | Persistent (in git) |
|
|
47
47
|
| Auto Memory (`~/.claude/projects/`) | Claude (automatic) | User preferences, feedback | Persistent (global) |
|
|
48
|
-
| **Knowledge Base** (`~/.claude/knowledge/`) | Multi-agent (Phase
|
|
48
|
+
| **Knowledge Base** (`~/.claude/knowledge/`) | Multi-agent (Phase 5) | Architecture, patterns, gotchas, decisions | Persistent (per-project) |
|
|
49
49
|
|
|
50
50
|
They all complement each other - no conflicts.
|
|
51
51
|
|
|
@@ -54,7 +54,7 @@ They all complement each other - no conflicts.
|
|
|
54
54
|
A Short run skips Phase 1, but knowledge **is still read**:
|
|
55
55
|
|
|
56
56
|
- Knowledge files are added to the prompt before the Opus agent starts in Phase 3
|
|
57
|
-
- Knowledge capture is still performed in Phase
|
|
57
|
+
- Knowledge capture is still performed in Phase 5 (simplified report)
|
|
58
58
|
|
|
59
59
|
### Knowledge Maintenance
|
|
60
60
|
|
|
@@ -107,18 +107,18 @@ When calling sub-agents, inject task-specific context from previous phases:
|
|
|
107
107
|
### Phase -> Agent -> Context Flow
|
|
108
108
|
|
|
109
109
|
```
|
|
110
|
-
Phase 1:
|
|
110
|
+
Phase 1: Plan (analysis)
|
|
111
111
|
-- Explore agents (parallel) -> codebase scan
|
|
112
112
|
|
|
113
|
-
Phase
|
|
113
|
+
Phase 1: Plan
|
|
114
114
|
-- Architecture review agent
|
|
115
115
|
Context: task description + Phase 1 findings + affected modules
|
|
116
116
|
|
|
117
|
-
Phase
|
|
117
|
+
Phase 2: Dev
|
|
118
118
|
-- Sonnet agent (TDD)
|
|
119
|
-
Context: Phase
|
|
119
|
+
Context: Phase 1 plan + task dependencies
|
|
120
120
|
|
|
121
|
-
Phase
|
|
121
|
+
Phase 3: Review (CLI-aware parallel + triage)
|
|
122
122
|
Claude Code (3 parallel):
|
|
123
123
|
|-- Reviewer (opus) -> security + architecture
|
|
124
124
|
+-- Reviewer (sonnet) -> quality + correctness + edge cases
|
|
@@ -7,7 +7,7 @@
|
|
|
7
7
|
- [Auto-run steps (when host is learned)](#auto-run-steps-when-host-is-learned)
|
|
8
8
|
- [Error evaluation contract](#error-evaluation-contract)
|
|
9
9
|
- [Autopilot behavior](#autopilot-behavior)
|
|
10
|
-
- [Phase
|
|
10
|
+
- [Phase 5 knowledge capture](#phase-5-knowledge-capture)
|
|
11
11
|
- [Generic rule (for copilot-instructions.md)](#generic-rule-for-copilot-instructionsmd)
|
|
12
12
|
- [Schema](#schema)
|
|
13
13
|
- [Smoke coverage](#smoke-coverage)
|
|
@@ -27,7 +27,7 @@ Codegen outputs (identifiers, localization keys, tokens) in Repo A get reference
|
|
|
27
27
|
|
|
28
28
|
## When this rule fires
|
|
29
29
|
|
|
30
|
-
In Phase
|
|
30
|
+
In Phase 4 (Commit & PR), **before** the pre-commit local checkout prompt, if `state.projects.length >= 2`:
|
|
31
31
|
|
|
32
32
|
1. Compute `repoSet` = sorted array of touched repo names.
|
|
33
33
|
2. Look up `prefs.global.multiRepoIntegrationHosts` for an entry whose `repoSet` (sorted) equals this combo.
|
|
@@ -100,7 +100,7 @@ Persist to `prefs.global.multiRepoIntegrationHosts[]`:
|
|
|
100
100
|
|
|
101
101
|
## Auto-run steps (when host is learned)
|
|
102
102
|
|
|
103
|
-
Progress-contract line `→ integrating <host-scheme> with <N> submodules`. Tracker sub-step on Phase
|
|
103
|
+
Progress-contract line `→ integrating <host-scheme> with <N> submodules`. Tracker sub-step on Phase 4 to show live status: `phase-tracker.sh sub 4 0 "Integration build" in_progress`.
|
|
104
104
|
|
|
105
105
|
```bash
|
|
106
106
|
HOST="$(jq -r --arg key "$repoSet_sorted_joined" \
|
|
@@ -126,12 +126,12 @@ BUILD_ERRORS=$(eval "$BUILD_CMD" 2>&1 | grep -E "error:|FAILURE:|error FS" || tr
|
|
|
126
126
|
|
|
127
127
|
# 4. Evaluate + decide
|
|
128
128
|
if [ -z "$BUILD_ERRORS" ]; then
|
|
129
|
-
phase-tracker.sh sub
|
|
130
|
-
log "Phase
|
|
129
|
+
phase-tracker.sh sub 4 0 "Integration build" completed
|
|
130
|
+
log "Phase 4.0: Integration build - clean"
|
|
131
131
|
# proceed to pre-commit checkout prompt + commit
|
|
132
132
|
else
|
|
133
|
-
phase-tracker.sh sub
|
|
134
|
-
log "Phase
|
|
133
|
+
phase-tracker.sh sub 4 0 "Integration build" failed
|
|
134
|
+
log "Phase 4.0: Integration build - NEW errors detected"
|
|
135
135
|
# show errors, ask user
|
|
136
136
|
fi
|
|
137
137
|
```
|
|
@@ -150,13 +150,13 @@ Three outcomes after the build:
|
|
|
150
150
|
Integration build failed with N new errors. Pipeline can go back to
|
|
151
151
|
Phase 3 to fix, or you can fix manually and tell the pipeline to retry.
|
|
152
152
|
|
|
153
|
-
[1] Return to Phase
|
|
153
|
+
[1] Return to Phase 2: Dev with these errors as input (auto-fix attempt)
|
|
154
154
|
[2] Pause - I'll fix manually, then /multi-agent resume
|
|
155
|
-
[3] Override - proceed to commit anyway (logs a warning in Phase
|
|
155
|
+
[3] Override - proceed to commit anyway (logs a warning in Phase 5 report)
|
|
156
156
|
|
|
157
157
|
Select [1-3]:
|
|
158
158
|
```
|
|
159
|
-
3. **Pre-existing errors (unrelated to this task)** → if the pipeline has a baseline error count (from a clean pre-change build), compare deltas. When `current_errors <= baseline`, treat as **no new errors** and proceed; log the pre-existing count in Phase
|
|
159
|
+
3. **Pre-existing errors (unrelated to this task)** → if the pipeline has a baseline error count (from a clean pre-change build), compare deltas. When `current_errors <= baseline`, treat as **no new errors** and proceed; log the pre-existing count in Phase 5 Report.
|
|
160
160
|
|
|
161
161
|
---
|
|
162
162
|
|
|
@@ -164,13 +164,13 @@ Three outcomes after the build:
|
|
|
164
164
|
|
|
165
165
|
Autopilot skips the "Override / pause" prompt. If `count` < 3 on this combo (new / low-confidence), autopilot treats a build failure as BLOCKING and returns to Phase 3 automatically (option 1). If `count >= 3` and `lastResult == success`, a new failure is treated as blocking the same way - but with lower surprise since we've seen this combo succeed before.
|
|
166
166
|
|
|
167
|
-
The learn-once prompt itself also skips under autopilot: instead, autopilot logs `"Phase
|
|
167
|
+
The learn-once prompt itself also skips under autopilot: instead, autopilot logs `"Phase 4.0: SKIPPED - no learned host for this combo, autopilot refuses to prompt. Run in normal mode once to teach."` and proceeds. This is explicit and recoverable.
|
|
168
168
|
|
|
169
169
|
---
|
|
170
170
|
|
|
171
|
-
## Phase
|
|
171
|
+
## Phase 5 knowledge capture
|
|
172
172
|
|
|
173
|
-
When a learn-once prompt writes a new entry, Phase
|
|
173
|
+
When a learn-once prompt writes a new entry, Phase 5 Step 5 (knowledge + memory) ALSO writes a project-scoped memory:
|
|
174
174
|
|
|
175
175
|
```
|
|
176
176
|
Type: reference
|
|
@@ -1,26 +1,26 @@
|
|
|
1
1
|
---
|
|
2
|
-
description: "Canonical required-reading list for outward-facing payloads (PR body, Jira comment, closing report) plus the markup dialect per surface. Loaded by every mode that runs Phase
|
|
2
|
+
description: "Canonical required-reading list for outward-facing payloads (PR body, Jira comment, closing report) plus the markup dialect per surface. Loaded by every mode that runs Phase 4 or Phase 5."
|
|
3
3
|
---
|
|
4
4
|
|
|
5
5
|
# Outward-facing payload contracts
|
|
6
6
|
|
|
7
7
|
> Every mode that opens a PR, comments on a tracker, or closes out a run reads this file first. It does not restate the contracts - it names them, so no mode has to carry its own copy and drift from the others.
|
|
8
8
|
|
|
9
|
-
## Read before Phase
|
|
9
|
+
## Read before Phase 4
|
|
10
10
|
|
|
11
11
|
| Read | Before | Governs |
|
|
12
12
|
|---|---|---|
|
|
13
13
|
| [`channels/pr.md`]($HOME/.claude/multi-agent-refs/channels/pr.md) | assembling the PR body | fixed section set (`summary` → `changes` → `architecture` cond. → `verification` → `risk` cond. → `dependencies` cond. → `related`), Markdown-only rule, reviewer-preserving Bitbucket PUT payload |
|
|
14
|
-
| [`phases/phase-
|
|
14
|
+
| [`phases/phase-4-commit.md`]($HOME/.claude/multi-agent-refs/phases/phase-4-commit.md) | committing | commit convention, default-reviewer fetch, draft/ready prompt, push-must-succeed loop |
|
|
15
15
|
| [`rules.md`]($HOME/.claude/multi-agent-refs/rules.md) "External System Outputs" | any REST payload | real newlines, no HTML entities, no hand-rolled JSON, markup dialect per surface |
|
|
16
16
|
|
|
17
|
-
## Read before Phase
|
|
17
|
+
## Read before Phase 5
|
|
18
18
|
|
|
19
19
|
| Read | Before | Governs |
|
|
20
20
|
|---|---|---|
|
|
21
21
|
| [`channels/jira.md`]($HOME/.claude/multi-agent-refs/channels/jira.md) | posting the Jira comment | fixed section set incl. **Test Scenarios** (Given/When/Then, always present), markdown→wiki conversion table |
|
|
22
22
|
| [`channels/confluence.md`]($HOME/.claude/multi-agent-refs/channels/confluence.md) | writing a Confluence page | storage-format conversion, endpoint flavor |
|
|
23
|
-
| [`phases/phase-
|
|
23
|
+
| [`phases/phase-5-report.md`]($HOME/.claude/multi-agent-refs/phases/phase-5-report.md) | closing out | Timeline + Agent Activity + Cost Breakdown tables, missing-telemetry disclosure |
|
|
24
24
|
| [`tracker-contract.md`]($HOME/.claude/multi-agent-refs/tracker-contract.md) | every phase boundary | per-phase token narration, completion tile suffix |
|
|
25
25
|
|
|
26
26
|
## Markup dialect per surface
|
|
@@ -44,12 +44,12 @@ A payload that is in the right language but the wrong dialect is a defect of the
|
|
|
44
44
|
|
|
45
45
|
The run ends with the phase tracker glyph block **and** the numbers behind it - per-phase duration and token spend, plus totals.
|
|
46
46
|
|
|
47
|
-
- Record spend as you go: `phase-tracker.sh tokens <N> <in> <out> [cached]` after **every** LLM call, including each Phase
|
|
47
|
+
- Record spend as you go: `phase-tracker.sh tokens <N> <in> <out> [cached]` after **every** LLM call, including each Phase 3 reviewer subagent and each Phase 3 chunk. Counts are additive and nothing reconstructs them after the fact.
|
|
48
48
|
- Tag the model once per phase (`phase-tracker.sh model <N> <name>`) or the cost helper cannot price it and prints `-`.
|
|
49
49
|
- Durations come from the phase timestamps and survive a missed `tokens` call; token spend does not.
|
|
50
50
|
- If any phase has no token data, name those phases and say their cost is unavailable. Never print a report whose cost section is simply absent - the reader cannot tell "cheap run" from "nobody recorded it".
|
|
51
51
|
|
|
52
|
-
### Missing-telemetry disclosure (Phase
|
|
52
|
+
### Missing-telemetry disclosure (Phase 5, required)
|
|
53
53
|
|
|
54
54
|
Before composing the report, compute which phases carry no token data:
|
|
55
55
|
|
|
@@ -64,4 +64,4 @@ Non-empty `UNTRACKED` → both the agent-log report and the closing chat summary
|
|
|
64
64
|
|
|
65
65
|
## Fast modes are not exempt
|
|
66
66
|
|
|
67
|
-
A Short run skips Analysis and Planning, and the autopilot and local entries skip the interactive test gate. They run Phase
|
|
67
|
+
A Short run skips Analysis and Planning, and the autopilot and local entries skip the interactive test gate. They run Phase 4 and Phase 5 **unchanged**. A short pipeline is not a licence for an improvised payload shape, a missing Test Scenarios section, or a report without numbers.
|
|
@@ -33,13 +33,13 @@ Create at `$HOME/.claude/logs/multi-agent/{project}/{task-id}/agent-log.md`:
|
|
|
33
33
|
## Phase Duration Distribution
|
|
34
34
|
|
|
35
35
|
Phase 0: Init ████░░░░░░░░░░░░ 5s
|
|
36
|
-
Phase 1:
|
|
37
|
-
Phase
|
|
38
|
-
Phase
|
|
39
|
-
Phase
|
|
40
|
-
Phase
|
|
41
|
-
Phase
|
|
42
|
-
Phase
|
|
36
|
+
Phase 1: Plan (analysis) ██████░░░░░░░░░░ 12s
|
|
37
|
+
Phase 1: Plan █████░░░░░░░░░░░ 8s
|
|
38
|
+
Phase 2: Dev ████████████████ 2m 35s
|
|
39
|
+
Phase 3: Review ████████░░░░░░░░ 37s
|
|
40
|
+
Phase 3: Review (user test) ██░░░░░░░░░░░░░░ (user wait)
|
|
41
|
+
Phase 4: Commit ██░░░░░░░░░░░░░░ 3s
|
|
42
|
+
Phase 5: Report █░░░░░░░░░░░░░░░ 2s
|
|
43
43
|
|
|
44
44
|
## Review Iterations
|
|
45
45
|
|
|
@@ -62,7 +62,7 @@ Verdict: unverified (3 reviewers)
|
|
|
62
62
|
|
|
63
63
|
## Cost Breakdown
|
|
64
64
|
|
|
65
|
-
(emit by Phase
|
|
65
|
+
(emit by Phase 5 via `$HOME/.claude/scripts/render-agent-log-cost.sh <task-id>`. Renders unconditionally on every run. If the renderer exits 2 (no tracker data + no OTel spans), Phase 5 omits this section without failing the run.)
|
|
66
66
|
|
|
67
67
|
| Phase | Model | Tokens in | Tokens out | Est. USD |
|
|
68
68
|
| ----- | ----- | --------- | ---------- | -------- |
|
|
@@ -86,7 +86,7 @@ Verdict: unverified (3 reviewers)
|
|
|
86
86
|
|
|
87
87
|
### Cost Breakdown - emission contract
|
|
88
88
|
|
|
89
|
-
Phase
|
|
89
|
+
Phase 5 MUST attempt to render the Cost Breakdown section as part of the agent-log compose step:
|
|
90
90
|
|
|
91
91
|
```bash
|
|
92
92
|
COST_BLOCK=$(bash $HOME/.claude/scripts/render-agent-log-cost.sh "$TASK_ID" 2>/dev/null) && \
|
|
@@ -97,7 +97,7 @@ Emission is best-effort - exit 2 (no data) is silently skipped. Never fail the
|
|
|
97
97
|
|
|
98
98
|
### Tokens telemetry - phase responsibility
|
|
99
99
|
|
|
100
|
-
Every phase that dispatches a billable LLM agent MUST forward its token totals to the tracker. The minimal contract (already enforced via `smoke-tracker-contract.sh` for Phase
|
|
100
|
+
Every phase that dispatches a billable LLM agent MUST forward its token totals to the tracker. The minimal contract (already enforced via `smoke-tracker-contract.sh` for Phase 3):
|
|
101
101
|
|
|
102
102
|
```bash
|
|
103
103
|
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" <phase-id> <event> \
|