@mmerterden/multi-agent-pipeline 18.0.0 → 19.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +183 -0
- package/README.md +34 -18
- package/README.tr.md +14 -16
- package/docs/adr/0002-instruction-driven-flag.md +1 -0
- package/docs/adr/0005-lazy-phase-docs.md +11 -1
- package/docs/adr/0008-installer-modularization-and-secret-leak-defense.md +1 -0
- package/docs/adr/0010-own-code-graph.md +1 -0
- package/docs/adr/0014-six-phase-consolidation.md +134 -0
- package/docs/adr/README.md +2 -1
- package/docs/architecture.md +37 -38
- package/docs/best-practices.md +1 -1
- package/docs/ecosystem.md +37 -26
- package/docs/engineering.md +1 -1
- package/docs/facts.json +45 -0
- package/docs/features.md +54 -53
- package/docs/performance.md +5 -5
- package/docs/recovery-guide.md +9 -9
- package/docs/token-budget-history.md +3 -1
- package/index.js +2 -2
- package/install/_codex-agents.mjs +1 -1
- package/install/templates/claude-hooks.json +1 -1
- package/install/templates/codex-instructions.md +1 -1
- package/install/templates/copilot-instructions.md +28 -28
- package/manifest.json +209 -193
- package/package.json +2 -2
- package/pipeline/agents/dev-critic.md +3 -3
- package/pipeline/commands/figma-to-swiftui.md +1 -1
- package/pipeline/commands/multi-agent/SKILL.md +8 -8
- package/pipeline/commands/multi-agent/analysis/SKILL.md +9 -9
- package/pipeline/commands/multi-agent/autopilot/SKILL.md +7 -7
- package/pipeline/commands/multi-agent/channels/SKILL.md +15 -15
- package/pipeline/commands/multi-agent/diff-explain/SKILL.md +6 -6
- package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/graph/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/help/SKILL.md +62 -62
- package/pipeline/commands/multi-agent/language/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/local/SKILL.md +11 -11
- package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +13 -13
- package/pipeline/commands/multi-agent/log/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/manual-test/SKILL.md +9 -9
- package/pipeline/commands/multi-agent/model/SKILL.md +69 -0
- package/pipeline/commands/multi-agent/refactor/SKILL.md +3 -3
- package/pipeline/commands/multi-agent/resume/SKILL.md +4 -4
- package/pipeline/commands/multi-agent/resume-local/SKILL.md +19 -17
- package/pipeline/commands/multi-agent/review/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/route-off/SKILL.md +36 -0
- package/pipeline/commands/multi-agent/route-on/SKILL.md +74 -0
- package/pipeline/commands/multi-agent/route-status/SKILL.md +56 -0
- package/pipeline/commands/multi-agent/setup/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/status/SKILL.md +5 -5
- package/pipeline/commands/multi-agent/steer/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/sync/SKILL.md +12 -13
- package/pipeline/commands/multi-agent/test/SKILL.md +1 -1
- package/pipeline/lib/credential-inventory.sh +1 -1
- package/pipeline/lib/fetch-fortify.sh +1 -1
- package/pipeline/lib/model-rung.sh +142 -0
- package/pipeline/lib/phase-schema.mjs +88 -0
- package/pipeline/lib/plan-todos.sh +5 -5
- package/pipeline/lib/route-state.sh +161 -0
- package/pipeline/lib/run-paths.sh +2 -2
- package/pipeline/multi-agent-refs/_account-picker.md +1 -1
- package/pipeline/multi-agent-refs/_dev-context.md +1 -1
- package/pipeline/multi-agent-refs/_input-parser.md +1 -1
- package/pipeline/multi-agent-refs/analysis/evidence.md +0 -9
- package/pipeline/multi-agent-refs/analysis/intake.md +1 -1
- package/pipeline/multi-agent-refs/analysis/locked.md +21 -22
- package/pipeline/multi-agent-refs/analysis/render.md +1 -1
- package/pipeline/multi-agent-refs/analysis/synthesis.md +12 -6
- package/pipeline/multi-agent-refs/android-guide.md +1 -1
- package/pipeline/multi-agent-refs/audit-guide.md +13 -13
- package/pipeline/multi-agent-refs/channels/issue-comment.md +2 -2
- package/pipeline/multi-agent-refs/channels/jira.md +3 -3
- package/pipeline/multi-agent-refs/channels/pr.md +4 -4
- package/pipeline/multi-agent-refs/channels/wiki.md +1 -1
- package/pipeline/multi-agent-refs/component-dispatch.md +3 -3
- package/pipeline/multi-agent-refs/cross-cli-contract.md +31 -6
- package/pipeline/multi-agent-refs/features/autopilot-circuit-breaker.md +4 -4
- package/pipeline/multi-agent-refs/features/code-graph.md +5 -5
- package/pipeline/multi-agent-refs/features/design-conformance.md +1 -1
- package/pipeline/multi-agent-refs/features/dev-critic.md +3 -3
- package/pipeline/multi-agent-refs/features/doctor.md +2 -2
- package/pipeline/multi-agent-refs/features/external-context-injection.md +3 -3
- package/pipeline/multi-agent-refs/features/maturity-followup.md +3 -3
- package/pipeline/multi-agent-refs/features/model-fallback.md +5 -5
- package/pipeline/multi-agent-refs/features/plan-todos.md +1 -1
- package/pipeline/multi-agent-refs/features/repo-map.md +1 -1
- package/pipeline/multi-agent-refs/features/review-delta.md +3 -3
- package/pipeline/multi-agent-refs/features/review-multi-repo.md +1 -1
- package/pipeline/multi-agent-refs/features/scope-check.md +4 -4
- package/pipeline/multi-agent-refs/features/skill-conformance.md +2 -2
- package/pipeline/multi-agent-refs/features/stack-skill-routing.md +1 -1
- package/pipeline/multi-agent-refs/features/verify-by-test.md +4 -4
- package/pipeline/multi-agent-refs/features/visual-evidence.md +19 -19
- package/pipeline/multi-agent-refs/features/worktree-finalize.md +6 -6
- package/pipeline/multi-agent-refs/issue-jira-triad.md +10 -10
- package/pipeline/multi-agent-refs/knowledge.md +11 -11
- package/pipeline/multi-agent-refs/multi-repo-integration-build.md +13 -13
- package/pipeline/multi-agent-refs/payload-contracts.md +8 -8
- package/pipeline/multi-agent-refs/phases/log-format.md +10 -10
- package/pipeline/multi-agent-refs/phases/modes.md +30 -30
- package/pipeline/multi-agent-refs/phases/operations.md +8 -8
- package/pipeline/multi-agent-refs/phases/phase-0-init.md +24 -24
- package/pipeline/multi-agent-refs/phases/phase-1-plan.md +599 -0
- package/pipeline/multi-agent-refs/phases/{phase-3-dev.md → phase-2-dev.md} +129 -49
- package/pipeline/multi-agent-refs/phases/{phase-4-review.md → phase-3-review.md} +225 -107
- package/pipeline/multi-agent-refs/phases/{phase-6-commit.md → phase-4-commit.md} +23 -23
- package/pipeline/multi-agent-refs/phases/{phase-7-report.md → phase-5-report.md} +29 -29
- package/pipeline/multi-agent-refs/phases.md +44 -48
- package/pipeline/multi-agent-refs/picker-contract.md +1 -1
- package/pipeline/multi-agent-refs/progress-contract.md +6 -6
- package/pipeline/multi-agent-refs/readiness-review.md +1 -1
- package/pipeline/multi-agent-refs/rules.md +7 -7
- package/pipeline/multi-agent-refs/swiftui-guide.md +2 -2
- package/pipeline/multi-agent-refs/tracker-contract.md +31 -32
- package/pipeline/multi-agent-refs/wiki-capture.md +14 -14
- package/pipeline/preferences-template.json +9 -1
- package/pipeline/rules/outside-the-pipeline.md +1 -1
- package/pipeline/schemas/agent-state.schema.json +50 -50
- package/pipeline/schemas/analysis-output.schema.json +2 -2
- package/pipeline/schemas/autopilot-config.schema.json +1 -1
- package/pipeline/schemas/code-graph.schema.json +1 -1
- package/pipeline/schemas/criteria-manifest.schema.json +1 -1
- package/pipeline/schemas/dev-critic-output.schema.json +1 -1
- package/pipeline/schemas/diff-risk.schema.json +1 -1
- package/pipeline/schemas/migrations/prefs-2.4.0-to-2.5.0.mjs +2 -2
- package/pipeline/schemas/migrations/prefs-2.6.0-to-2.7.0.mjs +31 -0
- package/pipeline/schemas/migrations/state-2.1.0-to-2.2.0.mjs +129 -0
- package/pipeline/schemas/phases.json +105 -0
- package/pipeline/schemas/plan-todos.schema.json +5 -5
- package/pipeline/schemas/planning-output.schema.json +1 -1
- package/pipeline/schemas/prefs.schema.json +100 -56
- package/pipeline/schemas/reviewer-output.schema.json +3 -3
- package/pipeline/schemas/route-config.schema.json +74 -0
- package/pipeline/schemas/scope-check.schema.json +1 -1
- package/pipeline/schemas/test-gap.schema.json +1 -1
- package/pipeline/schemas/token-budget.json +12 -18
- package/pipeline/schemas/triage-output.schema.json +6 -6
- package/pipeline/scripts/README.md +3 -3
- package/pipeline/scripts/_code-graph.mjs +2 -2
- package/pipeline/scripts/_run-paths.mjs +2 -2
- package/pipeline/scripts/_smoke-root.sh +1 -1
- package/pipeline/scripts/aggregate-metrics.mjs +1 -1
- package/pipeline/scripts/capture-flush.sh +8 -8
- package/pipeline/scripts/capture-resume.sh +3 -3
- package/pipeline/scripts/classify-plan-safety.mjs +1 -1
- package/pipeline/scripts/diff-explain.mjs +1 -1
- package/pipeline/scripts/doctor.mjs +2 -2
- package/pipeline/scripts/gc-abandoned.sh +3 -3
- package/pipeline/scripts/gc-tmp.sh +1 -1
- package/pipeline/scripts/gc-worktrees.sh +1 -1
- package/pipeline/scripts/gen-facts.mjs +175 -0
- package/pipeline/scripts/gen-mode-dispatch.mjs +32 -37
- package/pipeline/scripts/gen-ref-toc.mjs +1 -1
- package/pipeline/scripts/graph-report.mjs +1 -1
- package/pipeline/scripts/jira-attach.sh +1 -1
- package/pipeline/scripts/learn-from-transcripts.mjs +1 -1
- package/pipeline/scripts/learning-curve.mjs +2 -2
- package/pipeline/scripts/log-metric.sh +17 -4
- package/pipeline/scripts/memory-save.sh +1 -1
- package/pipeline/scripts/migrate-prefs.mjs +22 -5
- package/pipeline/scripts/phase-banner.sh +20 -20
- package/pipeline/scripts/phase-tracker.sh +7 -7
- package/pipeline/scripts/plan-coverage-gate.mjs +2 -2
- package/pipeline/scripts/render-agent-log-cost.sh +1 -1
- package/pipeline/scripts/render-work-summary.sh +3 -3
- package/pipeline/scripts/review-file-filter.mjs +1 -1
- package/pipeline/scripts/run-aggregator.mjs +13 -6
- package/pipeline/scripts/run-metrics.mjs +1 -1
- package/pipeline/scripts/runs-index.mjs +11 -1
- package/pipeline/scripts/smoke-cross-cli-behavior.sh +6 -6
- package/pipeline/scripts/smoke-schema-validation.sh +26 -7
- package/pipeline/scripts/token-budget-report.mjs +13 -2
- package/pipeline/scripts/triage-memory.mjs +2 -2
- package/pipeline/scripts/validate-analysis-doc.mjs +73 -17
- package/pipeline/scripts/validate-planning.mjs +1 -1
- package/pipeline/scripts/validate-reviewer.mjs +1 -1
- package/pipeline/scripts/validate-state.mjs +45 -5
- package/pipeline/scripts/validate-triage.mjs +3 -3
- package/pipeline/scripts/worktree-finalize.sh +5 -5
- package/pipeline/skills/.skill-manifest.json +37 -21
- package/pipeline/skills/.skills-index.json +49 -5
- package/pipeline/skills/shared/README.md +10 -6
- package/pipeline/skills/shared/core/apple-archive-compliance/SKILL.md +2 -2
- package/pipeline/skills/shared/core/google-play-compliance/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent/SKILL.md +69 -71
- package/pipeline/skills/shared/core/multi-agent-autopilot/SKILL.md +3 -3
- package/pipeline/skills/shared/core/multi-agent-channels/SKILL.md +14 -14
- package/pipeline/skills/shared/core/multi-agent-diff-explain/SKILL.md +5 -5
- package/pipeline/skills/shared/core/multi-agent-graph/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +25 -23
- package/pipeline/skills/shared/core/multi-agent-language/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-local/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +8 -8
- package/pipeline/skills/shared/core/multi-agent-manual-test/SKILL.md +6 -6
- package/pipeline/skills/shared/core/multi-agent-model/SKILL.md +71 -0
- package/pipeline/skills/shared/core/multi-agent-refactor/SKILL.md +3 -3
- package/pipeline/skills/shared/core/multi-agent-resume/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-resume-local/SKILL.md +7 -7
- package/pipeline/skills/shared/core/multi-agent-route-off/SKILL.md +39 -0
- package/pipeline/skills/shared/core/multi-agent-route-on/SKILL.md +76 -0
- package/pipeline/skills/shared/core/multi-agent-route-status/SKILL.md +59 -0
- package/pipeline/skills/shared/core/multi-agent-setup/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-status/SKILL.md +5 -5
- package/pipeline/skills/shared/core/multi-agent-steer/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +6 -5
- package/pipeline/skills/skills-index.md +8 -4
- package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +0 -263
- package/pipeline/multi-agent-refs/phases/phase-2-planning.md +0 -344
- package/pipeline/multi-agent-refs/phases/phase-5-test.md +0 -182
|
@@ -1,10 +1,10 @@
|
|
|
1
|
-
### Phase
|
|
1
|
+
### Phase 2: Dev (Sonnet)
|
|
2
2
|
|
|
3
|
-
> **TLDR** - Sonnet executes the plan task-by-task with TDD (red→green→refactor). Required: issue-tracker status moved to "In Progress" before any code, with a post-mutation verify step (re-reads the field, retries once on silent VALIDATION failures). Build verification after each task (up to 3 retries). Build-queue lock serializes concurrent xcodebuild. A Short run (`state.onlyDevelop`) uses Opus self-contained, with no Phase
|
|
3
|
+
> **TLDR** - Sonnet executes the plan task-by-task with TDD (red→green→refactor). Required: issue-tracker status moved to "In Progress" before any code, with a post-mutation verify step (re-reads the field, retries once on silent VALIDATION failures). Build verification after each task (up to 3 retries). Build-queue lock serializes concurrent xcodebuild. A Short run (`state.onlyDevelop`) uses Opus self-contained, with no Phase 1 plan. `taskType === component` **short-circuits the TDD path** and delegates the whole phase to the enabled `ai-<platform>-toolkit` marketplace plugin's component skill (`create-component`, fallback `create-ui-component`) - see next subsection.
|
|
4
4
|
|
|
5
|
-
## Phase
|
|
5
|
+
## Phase 2 Pre-flight (BLOCKING, v9.0.0)
|
|
6
6
|
|
|
7
|
-
Per Locked decision 30, Phase
|
|
7
|
+
Per Locked decision 30, Phase 2 Dev consumes the analysis document as the sole design source. **Figma** MCP / REST / URL fetches are forbidden; the toolkit MCP is not, and Steps 3.4 and 3.55 call it.
|
|
8
8
|
|
|
9
9
|
Pre-flight steps (run in order, abort on failure).
|
|
10
10
|
|
|
@@ -17,19 +17,19 @@ Pre-flight steps (run in order, abort on failure).
|
|
|
17
17
|
|
|
18
18
|
2. **Parse YAML front-matter** into `state.analysis.frontMatter`: `feature`, `platform`, `language`, `mode`, `evidence_digest`, `template_version`. **Abort** when `template_version` < `v3`; there is no degraded mode.
|
|
19
19
|
|
|
20
|
-
3. **Spec freshness**: two checks, both non-blocking, both flagged for Phase
|
|
20
|
+
3. **Spec freshness**: two checks, both non-blocking, both flagged for Phase 3.
|
|
21
21
|
- **Digest**: `state.run.lastAnalysisDigest` against `frontMatter.evidence_digest`. Mismatch = written from different evidence. Key absent -> `not-verifiable`, never `fresh`.
|
|
22
22
|
- **Repo drift**: `git diff --name-only <state.run.analysisBaseCommit>..HEAD` intersected with Section 14 paths. Non-empty = the repo moved under the spec. This is the one that fires on a `reused` doc, where the digest still matches because the evidence inputs did not change.
|
|
23
23
|
|
|
24
|
-
4. **Code Connect mapping lookup**: for each component named in analysis Section 6 (Bileşen Envanteri), search the repo for matching `*.figma.swift` / `*.figma.kt` files. Persist hits to `state.dev.codeConnect[<componentName>]`. Missing mappings are not blockers but the Phase
|
|
24
|
+
4. **Code Connect mapping lookup**: for each component named in analysis Section 6 (Bileşen Envanteri), search the repo for matching `*.figma.swift` / `*.figma.kt` files. Persist hits to `state.dev.codeConnect[<componentName>]`. Missing mappings are not blockers but the Phase 3 reviewer will flag them.
|
|
25
25
|
|
|
26
|
-
5. **Standards binding citations**: read `analysis Section 21 References` table. For each row with `Rol: bağlayıcı / binding`, persist the source path into `state.dev.standardsBindings[]`. Phase
|
|
26
|
+
5. **Standards binding citations**: read `analysis Section 21 References` table. For each row with `Rol: bağlayıcı / binding`, persist the source path into `state.dev.standardsBindings[]`. Phase 2 Dev tasks MUST cite at least one binding source per architectural decision (pre-existing Locked decision 8).
|
|
27
27
|
|
|
28
28
|
5b. **Test plan handoff (the RED input)**: read `analysis Section 15` into `state.dev.testPlan[]` (15.1 unit rows with name/arrange/act/expected/`BR-` id, 15.2 snapshot variants, 15.6 UI flows, 15.7 manual scenarios). **RED writes these tests, not invented ones** - development is TDD, so the analysis matrix is literally the first thing written. A row too vague to write from is an analysis defect: open a Section 20 question, do not improvise.
|
|
29
29
|
|
|
30
|
-
6. **Conventions handoff**: read `analysis Section 13.1 Concept Table` (Pass B output with footnotes). Persist concept-to-realization mapping into `state.dev.conventions[<concept>]`. Phase
|
|
30
|
+
6. **Conventions handoff**: read `analysis Section 13.1 Concept Table` (Pass B output with footnotes). Persist concept-to-realization mapping into `state.dev.conventions[<concept>]`. Phase 2 implementation uses these names verbatim, and a non-empty `sharedUtilities` bucket is binding: bind those formatters/rule-facades/tokens, never hand-roll a duplicate (e.g., if Section 13.1 says "State holder: PassengerFlightViewModel", Phase 2 names the class exactly `PassengerFlightViewModel`).
|
|
31
31
|
|
|
32
|
-
7. **MCP forbidden**: calling `mcp__claude_ai_Figma__*` in Phase
|
|
32
|
+
7. **MCP forbidden**: calling `mcp__claude_ai_Figma__*` in Phase 2 is a violation. `smoke-no-mcp-in-dev-phases.sh` reads `state.telemetry.mcpCalls[]` after the run and fails if Phase 2 contributed an entry.
|
|
33
33
|
|
|
34
34
|
8. **Criteria ledger (required, every mode)**: the moment this phase consults a skill, a marketplace plugin skill, a stack guide or a module `CLAUDE.md` in order to write code, append an entry to `state.telemetry.skillCalls[]`:
|
|
35
35
|
|
|
@@ -37,11 +37,11 @@ Pre-flight steps (run in order, abort on failure).
|
|
|
37
37
|
{"skill": "ios-coding-standard", "phase": 3, "targetFiles": ["Sources/Login/LoginViewModel.swift"], "timestamp": "<ISO-8601>"}
|
|
38
38
|
```
|
|
39
39
|
|
|
40
|
-
`targetFiles` is required - without it a skill applied to the wrong files still reads as "applied". Append at the moment of consultation, not at the end of the phase. Phase
|
|
40
|
+
`targetFiles` is required - without it a skill applied to the wrong files still reads as "applied". Append at the moment of consultation, not at the end of the phase. Phase 3 Step 1.78 treats this as self-report only and resolves criteria independently; it is the one signal separating "applied to the wrong files" from "never opened".
|
|
41
41
|
|
|
42
42
|
9. **Stack skill routing (every `taskType`, when a stack toolkit plugin is enabled)**: ask each enabled toolkit's own `index` skill which skills govern this task, load them BEFORE writing code, and record each into `state.telemetry.skillCalls[]` with `routedBy: "<toolkit>:index@<version>"`. Candidates are the effective `enabledPlugins`, not a stack table. The routing table stays in the plugin - a copy here would be the stale one, and `rules/outside-the-pipeline.md` runs the same routing outside a run. A screen-creation task loads the routed toolkit's `workflow/create-screen` when one exists. No toolkit, or none enabled, is a recorded no-op, not a halt. Contract: [`features/stack-skill-routing.md`]($HOME/.claude/multi-agent-refs/features/stack-skill-routing.md).
|
|
43
43
|
|
|
44
|
-
The analysis document is the SOLE design source in Phase
|
|
44
|
+
The analysis document is the SOLE design source in Phase 2. Variant choices, padding values, color tokens, copy strings, accessibility identifiers, and test method names all come from the rendered Pass B cells. If something is missing in the analysis doc, the fix is to re-run `/multi-agent:analysis`, not to fetch from Figma.
|
|
45
45
|
|
|
46
46
|
<!-- progress-contract: applied -->
|
|
47
47
|
|
|
@@ -49,15 +49,15 @@ The analysis document is the SOLE design source in Phase 3. Variant choices, pad
|
|
|
49
49
|
|
|
50
50
|
#### Input contract
|
|
51
51
|
|
|
52
|
-
Phase
|
|
52
|
+
Phase 2 consumes the Phase 1 output object conforming to `$HOME/.claude/schemas/planning-output.schema.json` - the task graph (`tasks[]` with `id`, `title`, `type`, `files`, and optional `dependsOn` / `acceptanceCriteria`) plus the architecture review notes. Tasks execute in dependency order; the schema's `dependsOn` field drives the ready-task picker. In a Short run (no Phase 1), Opus generates the equivalent task list inline before entering the loop below.
|
|
53
53
|
|
|
54
|
-
**Plan Todo iteration (opt-in)**: gated by `prefs.global.planTodos.enabled` (default: `false`). When enabled and Phase
|
|
54
|
+
**Plan Todo iteration (opt-in)**: gated by `prefs.global.planTodos.enabled` (default: `false`). When enabled and Phase 1 Step 11 emitted a `plan.todos[]`, Phase 2 iterates via `$HOME/.claude/lib/plan-todos.sh next/start/complete/fail` instead of walking `tasks[]` directly. When disabled, the loop walks `tasks[]` from `planning-output` - TDD contract is unchanged. Full helper loop + state semantics: `$HOME/.claude/multi-agent-refs/features/plan-todos.md`. A todo with `sourceTag: Reuse` binds the file analysis already found; `Modify` edits in place. Writing a new file over a `Reuse` step is a Locked 11 violation and Phase 3 flags it.
|
|
55
55
|
|
|
56
56
|
**Shadow-Git checkpoints (opt-in)**: gated by `prefs.global.shadowGit.enabled` (default: `false`). When enabled, the orchestrator snapshots the worktree via `$HOME/.claude/lib/shadow-git.sh` so sub-phase rollback is possible without touching the project's real `.git` history. Lifecycle: `shadow-git.sh init` (Phase 0 baseline), `shadow-git.sh snapshot` (per step after `plan-todos complete`), `shadow-git.sh restore <sha> --files` (rollback). Modes: `per-todo-step` (default) or `per-tool-call`. Full wiring + storage cap: `$HOME/.claude/multi-agent-refs/features/shadow-git.md`.
|
|
57
57
|
|
|
58
58
|
#### Component tasks - delegated dispatch (taskType === "component")
|
|
59
59
|
|
|
60
|
-
When Phase 0 Step 7 classified the task as `component`, Phase
|
|
60
|
+
When Phase 0 Step 7 classified the task as `component`, Phase 2 delegates the entire phase to the enabled `ai-<platform>-toolkit` marketplace plugin's component skill (`create-component`, fallback `create-ui-component`) via the Skill tool and does NOT run the TDD loop below. The dispatch layer passes the plugin skill the analysis Section 6 (Bileşen Envanteri) entry + Section 13.1 conventions for the named component as context. Because plugin skills do not write pipeline state, the **dispatch layer** (not the skill) owns `state.phases["2"].subphases[]`, recording a coarse component-build row - multi-agent's `phase-tracker` reads that array with no special case. Plugin resolution (dual-name), dispatch call, failure/resume, multi-repo, Short-run elisions, and the intentional cross-CLI divergence live in `$HOME/.claude/multi-agent-refs/component-dispatch.md` - read it before editing component-task behaviour here. Phase 2 still owns: progress line `-> dispatching create-component <name>`, `retryCount` cap at 3, and fallthrough to the TDD path when dispatch prerequisites are missing (`taskType` absent OR the plugin is not enabled in this repo -> log anomaly, halt or run TDD per component-dispatch.md).
|
|
61
61
|
|
|
62
62
|
For non-component taskTypes (`bugfix`, `feature`, `refactor`, `chore`), continue with the standard TDD section below.
|
|
63
63
|
|
|
@@ -67,27 +67,27 @@ Applies to EVERY task whose analysis doc carries design ground truth (Section 6
|
|
|
67
67
|
|
|
68
68
|
1. **1:1, not interpretation.** The captured frames are the reference for every UI line. Layout, variant states, copy, and colours come from the analysis doc's measured values - never eyeballed, never "close enough".
|
|
69
69
|
2. **Code Connect decides the component.** A mapped component (`CodeConnectSnippet` / repo `*.figma.swift` / `*.figma.kt` row) MUST be used verbatim - sound-alike substitutes are forbidden; a missing modifier is added to the component's `+Modifiers` extension, never forked. No mapping exists -> write a NEW component per the Configuration/View/Modifiers architecture; never inline ad-hoc UI into the consumer screen.
|
|
70
|
-
3. **Inter-component spacing is part of the design.** Gaps, paddings, and alignment BETWEEN components must match the frame's measured values, mapped to spacing tokens (`Spacing*`) - never invented numbers. A spacing/layout deviation is a `review_blocking` finding in Phase
|
|
70
|
+
3. **Inter-component spacing is part of the design.** Gaps, paddings, and alignment BETWEEN components must match the frame's measured values, mapped to spacing tokens (`Spacing*`) - never invented numbers. A spacing/layout deviation is a `review_blocking` finding in Phase 3 Step 1.8, not a nitpick.
|
|
71
71
|
|
|
72
72
|
Missing design data (variant, padding, copy) → HALT and instruct the user to re-run `/multi-agent:analysis`; never guess and never call Figma MCP from this phase.
|
|
73
73
|
|
|
74
|
-
#### Re-entry from Phase
|
|
74
|
+
#### Re-entry from Phase 3 triage
|
|
75
75
|
|
|
76
|
-
Phase
|
|
76
|
+
Phase 2 runs twice in the pipeline lifetime: first for initial development, then optionally for rework after Phase 3 review. **Phase 2 never acts on raw reviewer output.** It only consumes `triage.accepted` findings - Fable triage in Phase 3 already filtered false-positives, deferred out-of-scope items, and rejected noise.
|
|
77
77
|
|
|
78
|
-
When re-entering from Phase
|
|
78
|
+
When re-entering from Phase 3:
|
|
79
79
|
|
|
80
80
|
1. Read `state.reviewIterations[-1]` - the latest iteration's **accepted** list.
|
|
81
81
|
2. For each `accepted` finding where `severity === "blocking"` or `severity === "important"`:
|
|
82
|
-
- Treat it exactly like a Phase
|
|
82
|
+
- Treat it exactly like a Phase 1 plan task (TDD cycle, build verify).
|
|
83
83
|
- Quote the original `issue` + `fix` text in the task description so the dev agent has full context.
|
|
84
84
|
3. `accepted.suggestion` items are applied opportunistically (no TDD loop required) unless the user asked for suggestions to be treated strictly.
|
|
85
|
-
4. `deferred` items are NOT actioned in this re-entry - they surface in Phase
|
|
85
|
+
4. `deferred` items are NOT actioned in this re-entry - they surface in Phase 5's "Follow-up items" section.
|
|
86
86
|
5. `rejected` items are never touched. Log their IDs + triage reasons for audit only.
|
|
87
|
-
6. After rework, increment `state.phases["
|
|
87
|
+
6. After rework, increment `state.phases["2"].retryCount`; hard-kill at `retryCount === 3` and escalate to the user. Do not loop indefinitely. When `prefs.global.autopilotCircuitBreaker.enabled` is true (default), record the hard-kill before escalating as `state.circuitBreaker = {tripped: true, trigger: 3, detail: "rework cycles exhausted", checkpoint: {phase: 3, step: "re-entry", iteration: <n>}, trippedAt, counters: {reworkCycles: <n>}}`: the rework-storm trigger; `autopilotCircuitBreaker.maxReworkCycles` (default 3) is bounded above by this cap.
|
|
88
88
|
7. When `state.reviewIterations[-1].delta` exists, `stillPresent[]` findings come first in the task list, quoted `STILL PRESENT after round N-1's fix`; repeating the previous attempt is the loop the circuit-breaker stops.
|
|
89
89
|
|
|
90
|
-
If the latest iteration has `triage.approved === true` AND `accepted === []`, Phase
|
|
90
|
+
If the latest iteration has `triage.approved === true` AND `accepted === []`, Phase 2 was entered by mistake - log the anomaly and return to Phase 3.
|
|
91
91
|
|
|
92
92
|
**Telemetry**: at the start of every re-entry, emit:
|
|
93
93
|
|
|
@@ -105,7 +105,7 @@ For each task (respecting dependency order):
|
|
|
105
105
|
- **Where it applies**:
|
|
106
106
|
- Jira: transition issue to "In Progress" via REST API (`/rest/api/2/issue/{id}/transitions`)
|
|
107
107
|
- GitHub Projects V2: `updateProjectV2ItemFieldValue` on the Status field (singleSelectOptionId)
|
|
108
|
-
- Internal state: `agent-state.json` → `phases["
|
|
108
|
+
- Internal state: `agent-state.json` → `phases["2"].status = "in_progress"`
|
|
109
109
|
- **Post-update verify** (catches silent VALIDATION errors from stale option IDs after board rebuilds):
|
|
110
110
|
- Re-read the Status field after mutation.
|
|
111
111
|
- If the value is NOT `In Progress` (Jira) / `In Develop` (Projects V2), re-fetch current option IDs and retry ONCE.
|
|
@@ -117,7 +117,7 @@ For each task (respecting dependency order):
|
|
|
117
117
|
3. **TDD cycle** (Launch Agent with `model: "sonnet"`) - gated by `state.testPolicy`: `tdd` = the loop below; `tests-after` = skip RED, author the same tests once green; `none` = author no tests (existing ones never weakened), report "tests not written - project policy" rather than a gap.
|
|
118
118
|
|
|
119
119
|
**RED - Write ONE failing test first:**
|
|
120
|
-
- **Rework re-entry**: if `state.reviewIterations[-1].verifyByTest.redTests[]` exists (Phase
|
|
120
|
+
- **Rework re-entry**: if `state.reviewIterations[-1].verifyByTest.redTests[]` exists (Phase 3 Step 3.7 ran), those failing repro tests ARE the RED step for their findings - make them green, do not write a duplicate failing test and do not delete or weaken them. See `$HOME/.claude/multi-agent-refs/features/verify-by-test.md`.
|
|
121
121
|
- Test framework: use whatever the project already uses. Detect from existing test files:
|
|
122
122
|
- `@Test` / `#expect` → Swift Testing
|
|
123
123
|
- `XCTestCase` / `XCTAssert` → XCTest
|
|
@@ -146,7 +146,7 @@ For each task (respecting dependency order):
|
|
|
146
146
|
|
|
147
147
|
**GREEN - Minimal code to pass:**
|
|
148
148
|
- Smallest change that makes the test green - no extras
|
|
149
|
-
- **Existing tests are immutable**: deleting, renaming, skipping, or weakening an existing assertion to reach green is a violation. A test may change only when the task itself changes the spec that test encodes, and the commit body must name the changed test and the spec change. (Deterministic backstop: the `test_lines_removed` signal in Phase
|
|
149
|
+
- **Existing tests are immutable**: deleting, renaming, skipping, or weakening an existing assertion to reach green is a violation. A test may change only when the task itself changes the spec that test encodes, and the commit body must name the changed test and the spec change. (Deterministic backstop: the `test_lines_removed` signal in Phase 3 Step 1.75 flags test files that shrink.)
|
|
150
150
|
- Run the same test again → must PASS
|
|
151
151
|
- Run full test suite → no regressions:
|
|
152
152
|
```bash
|
|
@@ -162,7 +162,7 @@ For each task (respecting dependency order):
|
|
|
162
162
|
**REFACTOR (if needed):**
|
|
163
163
|
- Only if duplication or naming is poor
|
|
164
164
|
- Re-run tests after refactor → still GREEN
|
|
165
|
-
- **Stability (required):** run every test added or changed in this diff `prefs.global.testStability.repeatCount` times (default 3; 1 disables) with the single-test invocation above. Disagreeing outcomes are not GREEN: log `test.flake_signal file=<f> passed=<k> of=<N>` and fix the test or the code first. A pass only on retry is a flake signal, not a pass. Same rule on the Phase
|
|
165
|
+
- **Stability (required):** run every test added or changed in this diff `prefs.global.testStability.repeatCount` times (default 3; 1 disables) with the single-test invocation above. Disagreeing outcomes are not GREEN: log `test.flake_signal file=<f> passed=<k> of=<N>` and fix the test or the code first. A pass only on retry is a flake signal, not a pass. Same rule on the Phase 3 rework re-entry.
|
|
166
166
|
|
|
167
167
|
**Target resolution** (auto-detect once per project, cache in `agent-state.json`; ios resolves scheme + simulator, android resolves module + variant, backend/web need none):
|
|
168
168
|
```bash
|
|
@@ -185,9 +185,9 @@ For each task (respecting dependency order):
|
|
|
185
185
|
git -C "{worktreePath}" add -A
|
|
186
186
|
git -C "{worktreePath}" commit -m "wip({scope}): {task summary} [{jiraId}]"
|
|
187
187
|
```
|
|
188
|
-
WIP commits are squashed in Phase
|
|
188
|
+
WIP commits are squashed in Phase 4 before PR. This prevents work loss if a later task fails or the session crashes.
|
|
189
189
|
7. Update task status: `completed` (apply the same verify rule - re-read the Status field and retry once on mismatch)
|
|
190
|
-
8. Log: "Phase
|
|
190
|
+
8. Log: "Phase 2: {task} - Test + Build passed"
|
|
191
191
|
|
|
192
192
|
---
|
|
193
193
|
|
|
@@ -226,9 +226,9 @@ release_build_lock() { rm -rf "$BUILD_LOCK"; }
|
|
|
226
226
|
|
|
227
227
|
**This applies to ALL xcodebuild calls in the pipeline:**
|
|
228
228
|
|
|
229
|
-
- Phase
|
|
230
|
-
- Phase
|
|
231
|
-
- Phase
|
|
229
|
+
- Phase 2 Step 4 (build after development)
|
|
230
|
+
- Phase 3 Step 1 Gate 1 (build gate before review)
|
|
231
|
+
- Phase 3 Step 1 Gate 3 (test gate before review)
|
|
232
232
|
|
|
233
233
|
**Android**: same lock discipline - the Gradle daemon and `build/` outputs contend across parallel worktrees. **Backend/web** (Python, Node.js): no lock needed - these build/test in parallel without conflicts.
|
|
234
234
|
|
|
@@ -236,13 +236,13 @@ release_build_lock() { rm -rf "$BUILD_LOCK"; }
|
|
|
236
236
|
|
|
237
237
|
#### Step 3.5 - Dev Critic (Evaluator-Optimizer, opt-in)
|
|
238
238
|
|
|
239
|
-
Gated by `prefs.global.devCritic.enabled` (default: `false`). When enabled, after the generator's last edit and BEFORE Phase
|
|
239
|
+
Gated by `prefs.global.devCritic.enabled` (default: `false`). When enabled, after the generator's last edit and BEFORE Phase 3: dispatch `dev-critic` sub-agent (Sonnet), run 4 deterministic gates (build / lint / test / secrets), then platform checklist. STRICT loop cap - **max 2 iterations** (round 2 re-checks round-1 failures only); round 3+ returns `escalate: true`. Severity routing: `blocking` → generator must fix; `important` → SHOULD fix or pass through to Phase 3; `suggestion` → generator's judgement. Full agent contract, schema, telemetry, when-to-enable: `$HOME/.claude/multi-agent-refs/features/dev-critic.md` + `$HOME/.claude/agents/dev-critic.md`.
|
|
240
240
|
|
|
241
241
|
---
|
|
242
242
|
|
|
243
243
|
#### Step 3.55 - Visual evidence capture (UI changes only)
|
|
244
244
|
|
|
245
|
-
Re-decide `visualEvidence.required` from the real diff first - Phase 0's `bugfix` verdict is provisional (section 1a). When required, capture with the build that just went green - here, not Phase
|
|
245
|
+
Re-decide `visualEvidence.required` from the real diff first - Phase 0's `bugfix` verdict is provisional (section 1a). When required, capture with the build that just went green - here, not Phase 3 (section 3). Exit 4 anywhere below is a gap, not a failure. Contract: `features/visual-evidence.md`.
|
|
246
246
|
|
|
247
247
|
`EVIDENCE_PLATFORM=$(jq -r '.visualEvidence.platform // ""' "$STATE_FILE")`. Empty closes every tier below with that reason; passing it on is exit 2.
|
|
248
248
|
|
|
@@ -258,11 +258,11 @@ Re-decide `visualEvidence.required` from the real diff first - Phase 0's `bugfix
|
|
|
258
258
|
|
|
259
259
|
Persist `state.uiTest` and the artefact entries. A UI test that ran is subject to the default-FAIL rule like the build: run `evidence-gate.mjs --claim test --status passed --evidence "$WORKTREE/.pipeline/ui-test.log"` first, because a runner that died before reaching the tests also exits non-zero.
|
|
260
260
|
|
|
261
|
-
#### Step 3.6 - Code-simplifier pass (required diff shrink, before Phase
|
|
261
|
+
#### Step 3.6 - Code-simplifier pass (required diff shrink, before Phase 3 handoff)
|
|
262
262
|
|
|
263
|
-
After the build/test green step and BEFORE Phase
|
|
263
|
+
After the build/test green step and BEFORE Phase 3 handoff, run one diff-shrink round (bloated diffs waste Phase 3 reviewer tokens and unfocus the PR).
|
|
264
264
|
|
|
265
|
-
1. **Dispatch ONE subagent** (`subagent_type: "general-purpose"`, model: `sonnet`); input is ONLY the working diff (`git -C "$WORKTREE" diff "origin/$BASE_BRANCH"...HEAD`) + task summary. No correctness re-review (Phase
|
|
265
|
+
1. **Dispatch ONE subagent** (`subagent_type: "general-purpose"`, model: `sonnet`); input is ONLY the working diff (`git -C "$WORKTREE" diff "origin/$BASE_BRANCH"...HEAD`) + task summary. No correctness re-review (Phase 3's job), no new features. Exactly four smells:
|
|
266
266
|
- **Comment bloat** - comments restating code, scaffolding notes, TODO chatter
|
|
267
267
|
- **Unrelated rewrites** - hunks outside task scope (re-formatting, unrequested rename sweeps)
|
|
268
268
|
- **Dead code** - unused symbols, unreachable branches, commented-out blocks from this diff
|
|
@@ -272,7 +272,7 @@ After the build/test green step and BEFORE Phase 4 handoff, run one diff-shrink
|
|
|
272
272
|
2. **Return contract** - a list of shrink edits `{ "file", "lines", "edit", "rationale", "risk": "safe|unsafe" }`; the subagent never writes files. Keep the `rationale` strings: Step 3.7 writes them into `scope-check.json`.
|
|
273
273
|
3. **Apply safe edits only** - skip `unsafe` (logged). Zero edits is normal - log and continue. Progress line: ` → applying shrink edits ({N} applied, {M} skipped)`
|
|
274
274
|
4. **Re-run build + tests** (same build-queue lock + evidence-gate rule as Step 4). Any breakage -> revert shrink edits wholesale and proceed pre-shrink; the simplifier must never cost a green state.
|
|
275
|
-
5. **Record tokens in the cost ledger** so Phase
|
|
275
|
+
5. **Record tokens in the cost ledger** so Phase 5's Cost Breakdown captures the pass:
|
|
276
276
|
|
|
277
277
|
```bash
|
|
278
278
|
LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 dev.simplifier_pass \
|
|
@@ -286,20 +286,20 @@ Scope guard: a single pass, never looped. Runs in a Short run too; component tas
|
|
|
286
286
|
|
|
287
287
|
#### Step 3.7 - Scope self-check (required handoff artifact)
|
|
288
288
|
|
|
289
|
-
Phase
|
|
289
|
+
Phase 3 cannot reconstruct why each file was touched or what was left out on purpose, so Dev states both before the handoff: write `$WORKTREE/.pipeline/scope-check.json` (`schemas/scope-check.schema.json`: one `{path, reason}` per file in the diff, `notDone[]` with `why` in `out-of-scope | follow-up | rejected-abstraction`, and the Step 3.6 `simplifier` rationales), then run the gate. **Record rules, consumers and the full contract: `$HOME/.claude/multi-agent-refs/features/scope-check.md`.**
|
|
290
290
|
|
|
291
291
|
```bash
|
|
292
292
|
git -C "$WORKTREE" diff --name-only "origin/$BASE_BRANCH"...HEAD \
|
|
293
293
|
| node $HOME/.claude/scripts/scope-check-gate.mjs --check "$WORKTREE/.pipeline/scope-check.json" --diff-files -
|
|
294
294
|
```
|
|
295
295
|
|
|
296
|
-
Exit 1 lists `unjustified[]`: complete the record once and re-run. A second exit 1 does not block; log `dev.scope_check=incomplete unjustified=<n>` and hand off, and Phase
|
|
296
|
+
Exit 1 lists `unjustified[]`: complete the record once and re-run. A second exit 1 does not block; log `dev.scope_check=incomplete unjustified=<n>` and hand off, and Phase 3 shows the gate output to reviewers inside `<scope-self-check>`. Progress line: ` → checking scope self-check ({n} files, {m} not done)`.
|
|
297
297
|
|
|
298
298
|
---
|
|
299
299
|
|
|
300
300
|
#### Short pipeline (`state.onlyDevelop === true`)
|
|
301
301
|
|
|
302
|
-
Set by the Phase 0 Step 7.5 depth picker, or by autopilot never (autopilot always runs Full). When it is true, Phase
|
|
302
|
+
Set by the Phase 0 Step 7.5 depth picker, or by autopilot never (autopilot always runs Full). When it is true, Phase 2 runs self-contained with **Opus** (not Sonnet). No Phase 1 plan exists - the agent creates its own scope.
|
|
303
303
|
|
|
304
304
|
**Flow:**
|
|
305
305
|
1. Read task description (from Jira, GitHub issue, or free-text)
|
|
@@ -314,13 +314,13 @@ Set by the Phase 0 Step 7.5 depth picker, or by autopilot never (autopilot alway
|
|
|
314
314
|
| Aspect | Full | Short |
|
|
315
315
|
|--------|--------|----------|
|
|
316
316
|
| Model | Sonnet | **Opus** |
|
|
317
|
-
| Plan source | Phase
|
|
317
|
+
| Plan source | Phase 1 task list | Self-determined |
|
|
318
318
|
| Task granularity | Per-plan-item | Agent decides |
|
|
319
319
|
| Status updates | Per task item | Single in_progress → completed |
|
|
320
|
-
| Scope confirmation | Phase
|
|
321
|
-
| Review of the result | Phase
|
|
320
|
+
| Scope confirmation | Phase 1 user approval | None (agent autonomous) |
|
|
321
|
+
| Review of the result | Phase 3 | Phase 3 (same) |
|
|
322
322
|
|
|
323
|
-
Because the agent determines its own scope here, Phase
|
|
323
|
+
Because the agent determines its own scope here, Phase 3 is the only place that checks the result against anything external. Record every skill, plugin skill and guide consulted during this phase into `state.telemetry.skillCalls[]` with the files it was applied to - Phase 3 resolves the criteria set independently, and this record is what lets it tell "applied and honoured" from "never opened".
|
|
324
324
|
|
|
325
325
|
**Never combined with autopilot.** Autopilot skips the depth question and runs Full, so `onlyDevelop` is false in every unattended run. "Fast plus unattended" was removed in v16.0.0 and no longer exists: something has to choose when nobody is asked, and unattended is the worst place to drop analysis and planning.
|
|
326
326
|
|
|
@@ -332,7 +332,7 @@ Because the agent determines its own scope here, Phase 4 is the only place that
|
|
|
332
332
|
|
|
333
333
|
Active when `state.projects[].length > 1` (set by Phase 0 multi-select). Single-repo flow above is preserved verbatim - this section adds the deltas.
|
|
334
334
|
|
|
335
|
-
**Todo tagging**: Phase
|
|
335
|
+
**Todo tagging**: Phase 1 plan items in multi-repo mode carry a `repo` field naming the target project root. The dev agent uses this tag to dispatch each todo:
|
|
336
336
|
|
|
337
337
|
```json
|
|
338
338
|
{
|
|
@@ -352,12 +352,12 @@ If a todo has no `repo` tag in multi-repo mode → log warning + ask user, do no
|
|
|
352
352
|
|
|
353
353
|
**Build serialization**:
|
|
354
354
|
- Xcode worktrees still share `BUILD_LOCK = /tmp/claude-xcodebuild.lock` - multi-repo just means more callers contending for the same lock; the existing `mkdir`-based queue handles it correctly.
|
|
355
|
-
- Non-Xcode repos may build in parallel across worktrees - but ONLY when builds are independent. If repo B's build depends on repo A's freshly-published artifact (e.g. SPM local path dependency), serialize them in dependency order. Plan items declare dependencies via Phase
|
|
355
|
+
- Non-Xcode repos may build in parallel across worktrees - but ONLY when builds are independent. If repo B's build depends on repo A's freshly-published artifact (e.g. SPM local path dependency), serialize them in dependency order. Plan items declare dependencies via Phase 1's `dependsOn` field; multi-repo dispatch respects that order.
|
|
356
356
|
|
|
357
357
|
**Failure isolation**: A build/test failure in one repo's worktree must NOT silently leave the other repos in a half-built state. On failure:
|
|
358
358
|
1. Capture the failure into `state.projects[i].buildStatus = { ok: false, attempts: N, lastError: "..." }`
|
|
359
|
-
2. The ENTIRE Phase
|
|
360
|
-
3. After 3 attempts, escalate; do NOT proceed to Phase
|
|
359
|
+
2. The ENTIRE Phase 2 retries within that repo's worktree (max 3, same as single-repo)
|
|
360
|
+
3. After 3 attempts, escalate; do NOT proceed to Phase 3 with any unbuilt repo
|
|
361
361
|
|
|
362
362
|
**Recording a pass (default-FAIL evidence gate):** before setting `buildStatus.ok = true`, the build output must be tee'd to a log and that log must substantiate the success - a zero exit code alone is not trusted. Run the evidence gate; on exit 1, do NOT record a pass:
|
|
363
363
|
```bash
|
|
@@ -376,7 +376,7 @@ $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 build.completed repo=common dur
|
|
|
376
376
|
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 build.completed repo=uicomponents duration_ms=$D status=ok
|
|
377
377
|
```
|
|
378
378
|
|
|
379
|
-
**Token forwarding:** every TDD round (red, green, refactor) that hits the dev model MUST forward token totals into the tracker so Phase
|
|
379
|
+
**Token forwarding:** every TDD round (red, green, refactor) that hits the dev model MUST forward token totals into the tracker so Phase 5's Cost Breakdown captures Phase 2:
|
|
380
380
|
|
|
381
381
|
```bash
|
|
382
382
|
LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 dev.tdd_round \
|
|
@@ -427,3 +427,83 @@ wrong side of the generator, and the Debug menu never showed it.
|
|
|
427
427
|
When the analysis doc has not recorded which trees are generated, that is a Phase 1
|
|
428
428
|
gap - say so rather than guessing, since guessing wrong is invisible until the next
|
|
429
429
|
regeneration.
|
|
430
|
+
|
|
431
|
+
---
|
|
432
|
+
|
|
433
|
+
## Exit gate: Verify (BLOCKING)
|
|
434
|
+
|
|
435
|
+
These four deterministic gates ran at the top of Review until v19.0.0, which meant
|
|
436
|
+
the build ran twice: once here at the end of Dev, once again as Review's Stage 1.
|
|
437
|
+
They now run once, at this phase's exit, and Review inherits the logs instead of
|
|
438
|
+
regenerating them. A failure here does not reach Review at all.
|
|
439
|
+
|
|
440
|
+
#### Step 1 - Deterministic Gates (run BEFORE AI review)
|
|
441
|
+
|
|
442
|
+
If any gate fails → fix first, don't waste AI tokens reviewing broken code.
|
|
443
|
+
|
|
444
|
+
```bash
|
|
445
|
+
# Gate 1: Build (xcodebuild/gradle assemble/tsc/py compile - stack-dependent; Xcode uses the build queue lock, see Phase 2) - tee output to a log
|
|
446
|
+
<build-command> 2>&1 | tee "$WORKTREE/.build.log"
|
|
447
|
+
# Gate 2: Lint (swiftlint/ktlint/ruff/eslint - stack-dependent)
|
|
448
|
+
# Gate 3: Tests pass (xcodebuild/gradle/pytest/the resolved node command) - tee output to a log
|
|
449
|
+
<test-command> 2>&1 | tee "$WORKTREE/.test.log"
|
|
450
|
+
# Gate 4: Secrets - run the scanner against the staged diff
|
|
451
|
+
bash $HOME/.claude/scripts/pre-commit-check.sh
|
|
452
|
+
```
|
|
453
|
+
|
|
454
|
+
**Default-FAIL evidence gate (required before recording any pass):** a green exit code is not enough - the captured log must actually show success. Before marking build/test passed, run the evidence gate against the tee'd log; it fails CLOSED when the log is missing, empty, or shows failure markers:
|
|
455
|
+
|
|
456
|
+
```bash
|
|
457
|
+
node $HOME/.claude/scripts/evidence-gate.mjs --claim build --status passed --evidence "$WORKTREE/.build.log" || BUILD_PASS=false
|
|
458
|
+
node $HOME/.claude/scripts/evidence-gate.mjs --claim test --status passed --evidence "$WORKTREE/.test.log" || TEST_PASS=false
|
|
459
|
+
```
|
|
460
|
+
|
|
461
|
+
This prevents a false "it built" claim with no log behind it. On exit 1, treat the gate as failed (do NOT proceed to AI review) and surface the gate's `reason`.
|
|
462
|
+
|
|
463
|
+
**Inherited failures (when `state.baseline.tests` exists).** Phase 0 Step 7.6 recorded whether the suite was already red, so Gate 3 blocks on what this work broke, not what it walked into:
|
|
464
|
+
|
|
465
|
+
| baseline status | Gate 3 |
|
|
466
|
+
|---|---|
|
|
467
|
+
| `green` | unchanged; every failure is this run's |
|
|
468
|
+
| `red` + `failing[]` | subtract those ids. Nothing left -> pass, logged `test:pass (inherited {N})`. A NEW failure still blocks. |
|
|
469
|
+
| `red`, empty `failing[]` | do NOT pass and do NOT silently block: report `test:inherited-red (not attributable, <logPath>)` and ask. Inventing a set here masks regressions. |
|
|
470
|
+
| `unknown` / absent | unchanged from today |
|
|
471
|
+
|
|
472
|
+
The subtraction never widens: match on identifier only, and when identifiers cannot be compared fall to the `not attributable` row.
|
|
473
|
+
|
|
474
|
+
**Gate results:**
|
|
475
|
+
|
|
476
|
+
- All pass (including the evidence gate) -> proceed to AI review
|
|
477
|
+
- Any fail -> fix immediately, re-run gates (no AI review until clean)
|
|
478
|
+
Log: "Phase 3: Gates - build:{pass/fail} lint:{pass/fail} test:{pass/fail/inherited-red} secrets:{clean/found} evidence:{ok/unverified}"
|
|
479
|
+
|
|
480
|
+
##### Gate 5 - Fortify SSC findings (runs when `state.contextLinks[]` contains a `fortify` entry, or when `prefs.global.fortify.alwaysCheck === true`)
|
|
481
|
+
|
|
482
|
+
If the task description referenced a Fortify version, or named a bare issue instance id, `~/.claude/lib/fetch-fortify.sh` already populated `state.fortifyFinding` in Phase 0 (`alwaysCheck` needs `prefs.global.fortify.versionIds` to know what to scan). Phase 3 reuses that payload and applies the deterministic gate:
|
|
483
|
+
|
|
484
|
+
```bash
|
|
485
|
+
gate=$(jq -r '.fortifyFinding.gateOutcome // empty' "$STATE_FILE")
|
|
486
|
+
if [ -z "$gate" ]; then
|
|
487
|
+
# No fortify entry in contextLinks; gate is N/A.
|
|
488
|
+
echo "→ fortify gate: n/a (no Fortify URL referenced)"
|
|
489
|
+
elif [ "$(jq -r '.fortifyFinding.gateOutcome.blocking' "$STATE_FILE")" = "true" ]; then
|
|
490
|
+
reason=$(jq -r '.fortifyFinding.gateOutcome.reason' "$STATE_FILE")
|
|
491
|
+
critical=$(jq -r '.fortifyFinding.severityCounts.Critical' "$STATE_FILE")
|
|
492
|
+
echo "→ fortify gate: BLOCKED ($reason, critical=$critical)"
|
|
493
|
+
# Treat as a deterministic gate failure - fix the critical findings before AI review.
|
|
494
|
+
exit 1
|
|
495
|
+
else
|
|
496
|
+
high=$(jq -r '.fortifyFinding.severityCounts.High // 0' "$STATE_FILE")
|
|
497
|
+
echo "→ fortify gate: pass (high=$high warnings carry into the channel summary)"
|
|
498
|
+
fi
|
|
499
|
+
```
|
|
500
|
+
|
|
501
|
+
Gate semantics:
|
|
502
|
+
|
|
503
|
+
| `gateOutcome.reason` | Phase 3 action |
|
|
504
|
+
|---|---|
|
|
505
|
+
| `critical-findings` | **Blocking** - Phase 2 re-dispatches with the finding list; pipeline does not advance until Critical count is 0. |
|
|
506
|
+
| `high-findings-warning` | Pass-with-warning - High findings appear in the channels summary (`## Security Scan (Fortify)` section in PR + Jira); they don't block the merge. |
|
|
507
|
+
| `clean` | Pass silently - section omitted from the channels summary. |
|
|
508
|
+
|
|
509
|
+
When `state.fortifyFinding.status === "skipped"` (token missing, VPN unreachable, or host mismatch) the gate logs `→ fortify gate: skipped (<reason>)` and never blocks - the user already saw the structured Save Flow signal at Phase 0 and chose to proceed.
|