@mmerterden/multi-agent-pipeline 13.5.1 → 14.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +232 -0
- package/README.md +3 -3
- package/docs/features.md +1 -1
- package/install/_common.mjs +73 -0
- package/install/_mcp-register.mjs +70 -31
- package/install/_plugin-skills.mjs +73 -14
- package/install/claude.mjs +28 -4
- package/install/codex.mjs +33 -2
- package/install/copilot.mjs +145 -9
- package/install/index.mjs +10 -6
- package/install/templates/copilot-instructions.md +1 -1
- package/package.json +1 -1
- package/pipeline/agents/code-reviewer.md +58 -1
- package/pipeline/commands/multi-agent/SKILL.md +7 -5
- package/pipeline/commands/multi-agent/analysis/SKILL.md +7 -7
- package/pipeline/commands/multi-agent/analysis-resolve/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/build-optimize/SKILL.md +7 -7
- package/pipeline/commands/multi-agent/channels/SKILL.md +5 -5
- package/pipeline/commands/multi-agent/dev/SKILL.md +23 -18
- package/pipeline/commands/multi-agent/dev-autopilot/SKILL.md +19 -13
- package/pipeline/commands/multi-agent/dev-local/SKILL.md +14 -12
- package/pipeline/commands/multi-agent/dev-local-autopilot/SKILL.md +17 -12
- package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/help/SKILL.md +4 -4
- package/pipeline/commands/multi-agent/ios-coding-standard/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +4 -4
- package/pipeline/commands/multi-agent/resume/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/review/SKILL.md +5 -5
- package/pipeline/commands/multi-agent/scan/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/search/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/setup/SKILL.md +6 -6
- package/pipeline/commands/multi-agent/{finish → ship}/SKILL.md +12 -12
- package/pipeline/commands/multi-agent/testflight-validation/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/update/SKILL.md +5 -2
- package/pipeline/commands/sim-test.md +2 -2
- package/pipeline/lib/credential-store-resolver.sh +16 -0
- package/pipeline/lib/credential-store.sh +47 -4
- package/pipeline/lib/fetch-figma-annotations.sh +26 -28
- package/pipeline/lib/figma-screenshot.sh +28 -39
- package/pipeline/lib/figma-token.sh +63 -0
- package/pipeline/multi-agent-refs/analysis-template.md +1 -1
- package/pipeline/multi-agent-refs/android-guide.md +1 -1
- package/pipeline/multi-agent-refs/channels/issue-comment.md +1 -1
- package/pipeline/multi-agent-refs/component-dispatch.md +2 -2
- package/pipeline/multi-agent-refs/cross-cli-contract.md +4 -4
- package/pipeline/multi-agent-refs/features/dev-critic.md +2 -2
- package/pipeline/multi-agent-refs/features/model-fallback.md +35 -2
- package/pipeline/multi-agent-refs/features/plan-todos.md +1 -1
- package/pipeline/multi-agent-refs/features/repo-map.md +1 -1
- package/pipeline/multi-agent-refs/features/review-multi-repo.md +3 -3
- package/pipeline/multi-agent-refs/features/shadow-git.md +1 -1
- package/pipeline/multi-agent-refs/features/skill-conformance.md +116 -0
- package/pipeline/multi-agent-refs/features/verify-by-test.md +1 -1
- package/pipeline/multi-agent-refs/generate-issue.md +1 -1
- package/pipeline/multi-agent-refs/multi-repo-integration-build.md +1 -1
- package/pipeline/multi-agent-refs/phases/log-format.md +4 -4
- package/pipeline/multi-agent-refs/phases/modes.md +7 -7
- package/pipeline/multi-agent-refs/phases/phase-0-init.md +13 -11
- package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +17 -15
- package/pipeline/multi-agent-refs/phases/phase-2-planning.md +7 -7
- package/pipeline/multi-agent-refs/phases/phase-3-dev.md +28 -13
- package/pipeline/multi-agent-refs/phases/phase-4-review.md +90 -58
- package/pipeline/multi-agent-refs/phases/phase-5-test.md +7 -7
- package/pipeline/multi-agent-refs/phases/phase-6-commit.md +8 -8
- package/pipeline/multi-agent-refs/phases/phase-7-report.md +8 -8
- package/pipeline/multi-agent-refs/phases.md +13 -13
- package/pipeline/multi-agent-refs/progress-contract.md +2 -2
- package/pipeline/multi-agent-refs/rules.md +7 -5
- package/pipeline/multi-agent-refs/swiftui-guide.md +1 -1
- package/pipeline/multi-agent-refs/tracker-contract.md +16 -15
- package/pipeline/preferences-template.json +7 -1
- package/pipeline/rules/figma-pipeline.md +2 -2
- package/pipeline/schemas/agent-state.schema.json +333 -79
- package/pipeline/schemas/criteria-manifest.schema.json +228 -0
- package/pipeline/schemas/migrations/prefs-2.4.0-to-2.5.0.mjs +64 -0
- package/pipeline/schemas/prefs.schema.json +118 -262
- package/pipeline/schemas/reviewer-output.schema.json +48 -3
- package/pipeline/schemas/token-budget.json +34 -10
- package/pipeline/schemas/triage-output.schema.json +112 -27
- package/pipeline/scripts/cost-table.json +7 -4
- package/pipeline/scripts/gc-worktrees.sh +1 -1
- package/pipeline/scripts/gen-mode-dispatch.mjs +6 -6
- package/pipeline/scripts/match-skills.mjs +37 -4
- package/pipeline/scripts/migrate-prefs.mjs +88 -17
- package/pipeline/scripts/phase-tracker.sh +14 -3
- package/pipeline/scripts/pre-commit-check.sh +49 -2
- package/pipeline/scripts/skill-conformance.mjs +960 -0
- package/pipeline/scripts/smoke-schema-validation.sh +17 -4
- package/pipeline/scripts/uninstall.mjs +35 -9
- package/pipeline/scripts/validate-reviewer.mjs +108 -1
- package/pipeline/skills/.skill-manifest.json +1 -1
- package/pipeline/skills/.skills-index.json +36 -9
- package/pipeline/skills/shared/README.md +15 -12
- package/pipeline/skills/shared/core/apple-archive-compliance/SKILL.md +1 -0
- package/pipeline/skills/shared/core/apple-archive-compliance/references/rules.yml +167 -0
- package/pipeline/skills/shared/core/google-play-compliance/SKILL.md +1 -0
- package/pipeline/skills/shared/core/google-play-compliance/references/rules.yml +184 -0
- package/pipeline/skills/shared/core/multi-agent/SKILL.md +10 -10
- package/pipeline/skills/shared/core/multi-agent-analysis/SKILL.md +4 -4
- package/pipeline/skills/shared/core/multi-agent-analysis-resolve/SKILL.md +3 -3
- package/pipeline/skills/shared/core/multi-agent-build-optimize/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-create-jira/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-dev/SKILL.md +6 -5
- package/pipeline/skills/shared/core/multi-agent-dev-autopilot/SKILL.md +7 -6
- package/pipeline/skills/shared/core/multi-agent-dev-local/SKILL.md +4 -3
- package/pipeline/skills/shared/core/multi-agent-dev-local-autopilot/SKILL.md +2 -1
- package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-ios-coding-standard/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +4 -4
- package/pipeline/skills/shared/core/multi-agent-review/SKILL.md +5 -5
- package/pipeline/skills/shared/core/multi-agent-scan/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-search/SKILL.md +1 -1
- package/pipeline/skills/shared/core/{multi-agent-finish → multi-agent-ship}/SKILL.md +8 -8
- package/pipeline/skills/shared/external/ios-coding-standard/SKILL.md +44 -5
- package/pipeline/skills/shared/external/ios-coding-standard/modules/_TEMPLATE.yml +82 -0
- package/pipeline/skills/shared/external/ios-coding-standard/references/STANDARD.md +169 -10
- package/pipeline/skills/shared/external/ios-coding-standard/references/rules.yml +335 -16
- package/pipeline/skills/skills-index.md +11 -8
|
@@ -0,0 +1,63 @@
|
|
|
1
|
+
#!/bin/bash
|
|
2
|
+
#
|
|
3
|
+
# figma-token.sh
|
|
4
|
+
# Single resolution point for the Tier 2 Figma Personal Access Token.
|
|
5
|
+
#
|
|
6
|
+
# Sourced by every Tier 2 fetcher (fetch-figma-annotations.sh,
|
|
7
|
+
# figma-screenshot.sh). It exists because those two each carried their own copy
|
|
8
|
+
# of this lookup: when `migrate-prefs.mjs` consolidated
|
|
9
|
+
# `keychainMapping.figma_pat` into `keychainMapping.figma` and deleted the old
|
|
10
|
+
# key, only the migration was updated. Both fetchers kept reading the deleted
|
|
11
|
+
# key, so every migrated install reported `missing-token` while a valid PAT sat
|
|
12
|
+
# under the new name - and the error text told the user to map the key the
|
|
13
|
+
# migration had just removed. One copy of the lookup cannot drift from itself.
|
|
14
|
+
#
|
|
15
|
+
# Tier 2 is only reached when Tier 1 (Figma MCP) is unreachable; see the Figma
|
|
16
|
+
# Access Tier rule in multi-agent-refs/rules.md for the chain and the
|
|
17
|
+
# expired-token decision that gates the move to Tier 3.
|
|
18
|
+
#
|
|
19
|
+
# Resolution order, first non-empty wins:
|
|
20
|
+
# 1. credential-store.sh get figma - the canonical logical key
|
|
21
|
+
# 2. credential-store.sh get figma_pat - pre-v13.6 installs that never migrated
|
|
22
|
+
# 3. $FIGMA_PAT - env fallback for CI and one-shot runs
|
|
23
|
+
#
|
|
24
|
+
# Logical keys are passed through verbatim. credential-store.sh owns the
|
|
25
|
+
# `prefs.global.keychainMapping` indirection, so resolving the mapping here as
|
|
26
|
+
# well would reintroduce the duplication this file exists to remove. No literal
|
|
27
|
+
# Keychain service name appears here, which keeps the Synced Command Hygiene
|
|
28
|
+
# rule satisfied.
|
|
29
|
+
#
|
|
30
|
+
# Requires `$CRED_STORE` to already be resolved (via credential-store-resolver.sh).
|
|
31
|
+
# An unset or non-executable CRED_STORE is tolerated: resolution falls through to
|
|
32
|
+
# the env fallback so a keychain-less CI box still works.
|
|
33
|
+
|
|
34
|
+
# The canonical key first, the legacy key second. Space-separated so the loop
|
|
35
|
+
# below stays POSIX-ish and works under bash 3.2 (stock macOS).
|
|
36
|
+
FIGMA_TOKEN_LOGICAL_KEYS="${FIGMA_TOKEN_LOGICAL_KEYS:-figma figma_pat}"
|
|
37
|
+
|
|
38
|
+
# resolve_figma_token
|
|
39
|
+
# Prints the token on stdout and returns 0, or prints nothing and returns 1.
|
|
40
|
+
resolve_figma_token() {
|
|
41
|
+
local key tok
|
|
42
|
+
if [ -n "${CRED_STORE:-}" ] && [ -x "${CRED_STORE:-}" ]; then
|
|
43
|
+
for key in $FIGMA_TOKEN_LOGICAL_KEYS; do
|
|
44
|
+
tok=$("$CRED_STORE" get "$key" 2>/dev/null || true)
|
|
45
|
+
if [ -n "$tok" ]; then
|
|
46
|
+
printf '%s' "$tok"
|
|
47
|
+
return 0
|
|
48
|
+
fi
|
|
49
|
+
done
|
|
50
|
+
fi
|
|
51
|
+
if [ -n "${FIGMA_PAT:-}" ]; then
|
|
52
|
+
printf '%s' "$FIGMA_PAT"
|
|
53
|
+
return 0
|
|
54
|
+
fi
|
|
55
|
+
return 1
|
|
56
|
+
}
|
|
57
|
+
|
|
58
|
+
# figma_token_remediation
|
|
59
|
+
# The one place the "how do I fix this" sentence is written, so a fetcher can
|
|
60
|
+
# never name a key the migration deleted.
|
|
61
|
+
figma_token_remediation() {
|
|
62
|
+
printf 'no Figma PAT for Tier 2. Map prefs.global.keychainMapping.figma (run /multi-agent:setup), or export FIGMA_PAT.'
|
|
63
|
+
}
|
|
@@ -10,7 +10,7 @@ description: "v3 canonical template for /multi-agent:analysis. 23 main sections
|
|
|
10
10
|
|
|
11
11
|
> **Language**: This file is read as a system prompt. Prose stays English. Example tables and headings carry bilingual TR / EN scaffolds; the runtime renderer picks the right column per `outputLanguage`.
|
|
12
12
|
|
|
13
|
-
> **Locked decisions that govern this template**: the canonical numbered list lives in
|
|
13
|
+
> **Locked decisions that govern this template**: the canonical numbered list lives in `$HOME/.claude/commands/multi-agent/analysis/SKILL.md` (cite decisions by label, not by a number duplicated here, to avoid drift). The template's structure is shaped chiefly by: one-feature-per-run, section omission rule, citation discipline (incl. annotation-as-copy), forward-looking spec, humanizer punctuation policy, standards binding, per-platform output split, repo-evidence reuse-first, Figma 3-tier access, Gherkin user stories, Goals + Non-Goals paired, SVG default, Files-to-Add tag, API response variants exhaustive, screenshots embedded, all Figma variants drilled, localization mode (ownership-aware), References at bottom, platform-agnostic + Pass B render, convention extraction, Pass B footnote mandatory, Lite mode, SwiftUI Preview block (iOS), variant usage explicit, analysis self-contained (no MCP downstream), and business-rule to acceptance-criterion to test traceability.
|
|
14
14
|
|
|
15
15
|
## Mode selection - Full vs Lite
|
|
16
16
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
## Android/Kotlin Component Generation Guide
|
|
2
2
|
|
|
3
|
-
> **MUST: Figma MCP-first (BLOCKING).** If the task references any Figma frame (URL, node ID, or "from the design"), the Dev phase MUST call `mcp__claude_ai_Figma__get_design_context` for every frame BEFORE writing a single Composable line. Use the `CodeConnectSnippet` component name verbatim - no sound-alike substitutions. Authentication failure is not a skip path. Full rule, trigger conditions, and gate failure modes:
|
|
3
|
+
> **MUST: Figma MCP-first (BLOCKING).** If the task references any Figma frame (URL, node ID, or "from the design"), the Dev phase MUST call `mcp__claude_ai_Figma__get_design_context` for every frame BEFORE writing a single Composable line. Use the `CodeConnectSnippet` component name verbatim - no sound-alike substitutions. Authentication failure is not a skip path. Full rule, trigger conditions, and gate failure modes: `$HOME/.claude/rules/figma-pipeline.md` "MUST: Figma MCP-first (BLOCKING)". Phase wiring: `$HOME/.claude/multi-agent-refs/phases/phase-3-dev.md` "MUST: Figma MCP-first (BLOCKING pre-step)".
|
|
4
4
|
|
|
5
5
|
When the task involves creating an Android UI component (Jetpack Compose), follow this architecture.
|
|
6
6
|
|
|
@@ -112,7 +112,7 @@ Token resolution: `gh auth status` for the active GitHub account selected in Pha
|
|
|
112
112
|
|
|
113
113
|
## Pairing with the Progress flag updater
|
|
114
114
|
|
|
115
|
-
Every issue comment post is paired with
|
|
115
|
+
Every issue comment post is paired with `$HOME/.claude/scripts/update-issue-progress.sh "$TASK_ID"` in the same Phase 7 step. Order: comment FIRST (so the timestamp marks the run), flags SECOND (so the body diff is one logical change).
|
|
116
116
|
|
|
117
117
|
```bash
|
|
118
118
|
# Phase 7 Step 5 - issue channel
|
|
@@ -111,9 +111,9 @@ On failure (the plugin skill returns an unrecoverable build/test error, or the d
|
|
|
111
111
|
|
|
112
112
|
## `--dev` mode behaviour
|
|
113
113
|
|
|
114
|
-
When multi-agent is invoked with `--dev` (fast path: Init → Dev → Commit → Report), the dispatch layer passes `mode: "dev"` so the plugin skill can elide unit tests and wiki (structural + snapshot still required; wiki deferred to Phase 7). If the plugin does not honor a `mode` hint, dispatch simply skips the post-build wiki step itself.
|
|
114
|
+
When multi-agent is invoked with `--dev` (fast path: Init → Dev → Review → Commit → Report), the dispatch layer passes `mode: "dev"` so the plugin skill can elide unit tests and wiki (structural + snapshot still required; wiki deferred to Phase 7). If the plugin does not honor a `mode` hint, dispatch simply skips the post-build wiki step itself.
|
|
115
115
|
|
|
116
|
-
Phase 4 reviewer
|
|
116
|
+
Phase 4 runs in `--dev` as it does in the full pipeline, and its reviewer count is **not** Phase 3's concern - the Step 1.77 scope gate decides that from diff risk, independently of `mode`. What the dispatch layer owes Phase 4 is the record of which plugin skill it delegated to, appended to `state.telemetry.skillCalls[]`, so the review can check the delivered component against the criteria that skill imposes.
|
|
117
117
|
|
|
118
118
|
## Cross-CLI behaviour (intentional divergence)
|
|
119
119
|
|
|
@@ -10,10 +10,10 @@
|
|
|
10
10
|
|
|
11
11
|
```
|
|
12
12
|
analysis, analysis-resolve, autopilot, build-optimize, channels, create-jira, design-check, dev,
|
|
13
|
-
dev-autopilot, dev-local, dev-local-autopilot, diff-explain,
|
|
13
|
+
dev-autopilot, dev-local, dev-local-autopilot, diff-explain, forget, garbage-collect,
|
|
14
14
|
help, ios-coding-standard, issue, jira, kill, language, local,
|
|
15
15
|
local-autopilot, log, manual-test, prune-logs, purge, refactor, resume, review, review-issue, review-jira,
|
|
16
|
-
routines, save, scan, search, setup, stack, status, sync, test, testflight-validation, uninstall, update
|
|
16
|
+
routines, save, scan, search, setup, ship, stack, status, sync, test, testflight-validation, uninstall, update
|
|
17
17
|
```
|
|
18
18
|
|
|
19
19
|
Categories:
|
|
@@ -21,8 +21,8 @@ Categories:
|
|
|
21
21
|
- **Interactive pickers** (single-purpose, not modes): `jira`, `issue`
|
|
22
22
|
- **Issue generator** (one-shot, no worktree, asks type Task/Bug/Story, hard approval gate before create): `create-jira`
|
|
23
23
|
- **Full 8-phase modes**: `autopilot`, `local`, `local-autopilot`
|
|
24
|
-
- **Fast modes** (Init -> Dev(Opus) -> Commit -> Report): `dev`, `dev-autopilot`, `dev-local`, `dev-local-autopilot`
|
|
25
|
-
- **Tail modes** (run the pipeline tail over already-done local work): `
|
|
24
|
+
- **Fast modes** (Init -> Dev(Opus) -> Review -> Commit -> Report): `dev`, `dev-autopilot`, `dev-local`, `dev-local-autopilot`
|
|
25
|
+
- **Tail modes** (run the pipeline tail over already-done local work): `ship`
|
|
26
26
|
- **Ops commands** (one-shot, no worktree): `status`, `log`, `kill`, `purge`, `uninstall`, `resume`, `review`, `review-jira`, `review-issue`, `analysis`, `analysis-resolve`, `build-optimize`, `channels`, `scan`, `search`, `diff-explain`, `garbage-collect`, `prune-logs`
|
|
27
27
|
- **Local audits** (worktree only to build; no commit, push, PR or channels): `design-check`, `testflight-validation`, `ios-coding-standard`. `testflight-validation` additionally never invokes `altool --upload-app` - a validation run must not be able to ship a build by accident.
|
|
28
28
|
- **Meta-ops**: `setup`, `sync`, `update`, `help`, `refactor`, `test`, `stack`, `manual-test`, `language`
|
|
@@ -7,7 +7,7 @@
|
|
|
7
7
|
1. Dispatch `dev-critic` sub-agent (Sonnet by default, tools: `Read, Grep, Glob, Bash`).
|
|
8
8
|
2. Critic runs 4 deterministic gates (build / lint / test / secrets) - failure of any → `pass: false` with `blocking` finding.
|
|
9
9
|
3. If gates green, critic walks the platform checklist (iOS 13-item / Android Kotlin / Backend generic - selected by Phase 1 `detectedStack`).
|
|
10
|
-
4. Returns schema-validated JSON (
|
|
10
|
+
4. Returns schema-validated JSON (`$HOME/.claude/schemas/dev-critic-output.schema.json`): `{pass, iteration, gates, findings[], escalate}`.
|
|
11
11
|
|
|
12
12
|
## Loop cap (STRICT) - max 2 iterations
|
|
13
13
|
|
|
@@ -41,4 +41,4 @@ Phase 4 is parallelization-with-voting - good for *adversarial* perspectives (
|
|
|
41
41
|
|
|
42
42
|
## Reference
|
|
43
43
|
|
|
44
|
-
See
|
|
44
|
+
See `$HOME/.claude/agents/dev-critic.md` for the full agent specification (gates, checklist enumeration, output schema, severity semantics).
|
|
@@ -10,7 +10,7 @@ without ever editing persona files at runtime.
|
|
|
10
10
|
> **Fable 5 available again (2026-07).** `claude-fable-5` is the top tier once
|
|
11
11
|
> more. Architect (`ios/android/backend-architect`), Reviewer-1 (`code-reviewer`),
|
|
12
12
|
> and triage personas declare `preferredModel: fable` - the deepest-reasoning
|
|
13
|
-
> roles. **If Fable 5 is unavailable, the first fallback is opus (`claude-opus-
|
|
13
|
+
> roles. **If Fable 5 is unavailable, the first fallback is opus (`claude-opus-5`).**
|
|
14
14
|
> Security and other `preferredModel: opus` personas keep opus as their top tier.
|
|
15
15
|
> The `premiumTierUntil` date gate stays in the contract as a generic mechanism
|
|
16
16
|
> for any plan-window-limited premium tier.
|
|
@@ -24,9 +24,42 @@ fable -> opus -> sonnet -> haiku
|
|
|
24
24
|
The orchestrator walks **one step down per trigger** from whichever tier a
|
|
25
25
|
persona prefers: a `fable` persona degrades `fable -> opus` first, then
|
|
26
26
|
`opus -> sonnet`, then `sonnet -> haiku`; an `opus` persona starts at
|
|
27
|
-
`opus -> sonnet`. So a Fable 5 outage transparently promotes Opus
|
|
27
|
+
`opus -> sonnet`. So a Fable 5 outage transparently promotes Opus 5 into the
|
|
28
28
|
architect/reviewer roles with no file edits.
|
|
29
29
|
|
|
30
|
+
### Rung names are the contract; model IDs are not
|
|
31
|
+
|
|
32
|
+
Personas and dispatch read the **rung** (`fable`, `opus`, `sonnet`, `haiku`). The wire
|
|
33
|
+
model ID each rung resolves to lives in exactly two places: `scripts/cost-table.json`
|
|
34
|
+
(`prices.<rung>.modelId`) and the per-host reviewer table in
|
|
35
|
+
`phases/phase-4-review.md`. A generation move edits those; it never renames a rung and
|
|
36
|
+
never touches a persona file.
|
|
37
|
+
|
|
38
|
+
Current resolution: `fable` -> `claude-fable-5`, `opus` -> `claude-opus-5`,
|
|
39
|
+
`sonnet` -> `claude-sonnet-5`, `haiku` -> `claude-haiku-4-5`.
|
|
40
|
+
|
|
41
|
+
**Rate limits do not follow the rung.** Claude Opus 5 draws on a pool separate from
|
|
42
|
+
the combined Opus 4.x pool, so a generation move does not inherit the previous
|
|
43
|
+
generation's headroom. Budget-ceiling downgrades below are unaffected - they read the
|
|
44
|
+
cost ledger, not the provider's quota - but a 429 on the opus rung after a generation
|
|
45
|
+
move is a limits question, not a fallback-contract question.
|
|
46
|
+
|
|
47
|
+
### Behavioural re-tuning the generation move implies
|
|
48
|
+
|
|
49
|
+
The pipeline's reviewer and dev prompts were tuned against the previous generation.
|
|
50
|
+
Three shifts on the current opus rung are worth knowing when reading a run that looks
|
|
51
|
+
different rather than broken, and none of them are bugs in this contract:
|
|
52
|
+
|
|
53
|
+
- **Longer user-facing output.** Effort is not the lever; prompt-level conciseness is.
|
|
54
|
+
Phase 7 report length and reviewer prose are where this shows.
|
|
55
|
+
- **Self-verification without being asked.** Explicit "double-check your work"
|
|
56
|
+
scaffolding now causes over-verification rather than preventing under-verification.
|
|
57
|
+
Phase 4's deterministic gates already carry that load.
|
|
58
|
+
- **Readier subagent delegation.** The previous generation under-reached and needed
|
|
59
|
+
encouragement; this one does not. Phase 1's parallel scan and Phase 4's reviewer
|
|
60
|
+
panel are already bounded by explicit counts, which is the right shape - keep them
|
|
61
|
+
bounded rather than adding "delegate more" guidance.
|
|
62
|
+
|
|
30
63
|
One step down per trigger, walking the ladder until a tier dispatches or the
|
|
31
64
|
floor (`haiku`) is reached. `haiku` is the last-resort floor: a run that reaches
|
|
32
65
|
it is heavily degraded but still makes progress instead of hard-halting because
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Feature: Plan Todos Iteration (Phase 3)
|
|
2
2
|
|
|
3
|
-
**Gated by `prefs.global.planTodos.enabled`** (default: `false`). When enabled and Phase 2 Step 4.5 emitted a `plan.todos[]` (conforming to
|
|
3
|
+
**Gated by `prefs.global.planTodos.enabled`** (default: `false`). When enabled and Phase 2 Step 4.5 emitted a `plan.todos[]` (conforming to `$HOME/.claude/schemas/plan-todos.schema.json`), Phase 3 iterates with the helper instead of walking `tasks[]` directly:
|
|
4
4
|
|
|
5
5
|
```bash
|
|
6
6
|
while next=$(bash "$HOME/.claude/lib/plan-todos.sh" next "$TASK_ID"); [ -n "$next" ]; do
|
|
@@ -14,7 +14,7 @@ if [ "$(jq -r '.global.repoMap.enabled // false' "$PREFS")" = "true" ]; then
|
|
|
14
14
|
args=(--root "$WORKTREE" --budget "$BUDGET" --top "$TOP" --format md)
|
|
15
15
|
[ -n "$INCLUDE" ] && args+=(--include "$INCLUDE")
|
|
16
16
|
[ -n "$EXCLUDE" ] && args+=(--exclude "$EXCLUDE")
|
|
17
|
-
REPO_MAP=$(node
|
|
17
|
+
REPO_MAP=$(node $HOME/.claude/scripts/repo-map.mjs "${args[@]}" 2>/dev/null || echo "")
|
|
18
18
|
fi
|
|
19
19
|
```
|
|
20
20
|
|
|
@@ -57,9 +57,9 @@ And log `review.diff_truncated repo=<name> bytes_dropped=<N>`. Triage receives t
|
|
|
57
57
|
|
|
58
58
|
**Telemetry**: Per-repo build/test gate timings + a single combined review/triage call set:
|
|
59
59
|
```bash
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
60
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 gate.build repo=common status=pass duration_ms=$D
|
|
61
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 gate.build repo=uicomponents status=pass duration_ms=$D
|
|
62
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.combined_diff repos=2 bytes=$BYTES truncated=false
|
|
63
63
|
```
|
|
64
64
|
|
|
65
65
|
---
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Feature: Shadow-Git Checkpoints (Phase 3)
|
|
2
2
|
|
|
3
|
-
**Gated by `prefs.global.shadowGit.enabled`** (default: `false`). The orchestrator snapshots the worktree via
|
|
3
|
+
**Gated by `prefs.global.shadowGit.enabled`** (default: `false`). The orchestrator snapshots the worktree via `$HOME/.claude/lib/shadow-git.sh` so sub-phase rollback is possible without polluting the project's real `.git` history.
|
|
4
4
|
|
|
5
5
|
```bash
|
|
6
6
|
# Phase 0 (one-time per task): initialize shadow repo + baseline snapshot.
|
|
@@ -0,0 +1,116 @@
|
|
|
1
|
+
# Skill conformance - reviewing against the criteria the work was built to
|
|
2
|
+
|
|
3
|
+
> **TLDR** - Phase 4 Step 1.78 resolves, deterministically and before any reviewer runs, WHAT the changed code was supposed to honour: declared rule registries scoped to the diff's languages, in-repo module guides, and the toolchains those registries delegate to. The result is `criteria-manifest.json`: a bounded set of rule IDs that becomes the denominator for "was this applied completely". Reviewers answer per rule ID. An ID that is neither checked nor explicitly waived fails the stage.
|
|
4
|
+
|
|
5
|
+
## Why this exists
|
|
6
|
+
|
|
7
|
+
Phase 4 used to select its review criteria from `detectedStack`, a string Phase 1 computed before a line of code existed, and Phase 3 never recorded which skills it actually applied. Two consequences:
|
|
8
|
+
|
|
9
|
+
- The dev side and the review side could disagree about the standard without either noticing. A component built by a plugin skill was reviewed against `clean-code`.
|
|
10
|
+
- "Did it do this correctly?" had no fixed denominator, so the only available answer was "it looks fine", and a reviewer that opened nothing produced the same output as a reviewer that checked everything.
|
|
11
|
+
|
|
12
|
+
The `--dev` family made this sharper: those modes have no Phase 1 at all, so the sole input to the old selection logic was absent.
|
|
13
|
+
|
|
14
|
+
## The four rules that make it work
|
|
15
|
+
|
|
16
|
+
1. **The denominator is written before the reviewers run.** `criteria-manifest.json` lands on disk in Step 1.78. Reviewers receive the selected rule IDs and must return a verdict per ID. A reviewer cannot narrow the set after seeing the diff, and a no-findings result still has to say what it checked.
|
|
17
|
+
2. **The resolver is primary; self-report only corroborates.** Phase 3 appends to `state.telemetry.skillCalls[]`, but coverage is never computed from it. An unrecorded consultation and no consultation are byte-identical in state, so a percentage derived from self-report reads green over an empty set. `ledger.source` therefore defaults to `derived`, and a declared entry the resolver could not bind to a changed file is FLAGGED rather than believed.
|
|
18
|
+
3. **Absent, empty and zero are three states.** `selfReport: absent` (no key), `empty` (key is `[]`), `declared` (has entries). Collapsing them makes the completeness claim unfalsifiable.
|
|
19
|
+
4. **The pipeline names no stack-specific skill.** Registries opt in; see below.
|
|
20
|
+
|
|
21
|
+
## Registry discovery is declared, never sniffed
|
|
22
|
+
|
|
23
|
+
A skill becomes a standards registry by saying so in its own frontmatter:
|
|
24
|
+
|
|
25
|
+
```yaml
|
|
26
|
+
standards-registry: references/rules.yml
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
`skill-conformance.mjs` reads frontmatter across the installed skills trees and loads what declared itself. Nothing in the pipeline names `ios-coding-standard`, `apple-archive-compliance` or any other skill, so a future UIKit, Objective-C, Kotlin or backend registry drops in with zero pipeline change - and a registry that is absent produces a declared coverage gap rather than a silent pass.
|
|
30
|
+
|
|
31
|
+
**The skills root differs per host**, so discovery is a bounded walk over candidate roots (`<install>/skills`, `<install>/multi-agent-refs/skills` for Codex, `<repo>/pipeline/skills`), installed layouts first. `install/copilot.mjs` copies `scripts/` byte-for-byte with no path rewrite, so a hardcoded `~/.claude/skills` is inert on two of the three hosts. That bug has already shipped here once: dynamic skill loading exited 1 on every real install while passing a smoke that ran from the repo. `skillsRootsSearched` is recorded in the manifest so an empty result is attributable to a root rather than to an absence of registries.
|
|
32
|
+
|
|
33
|
+
## Scope is required, and it is what makes this stack-generic
|
|
34
|
+
|
|
35
|
+
Every registry declares the languages and paths its rules may be applied to:
|
|
36
|
+
|
|
37
|
+
```yaml
|
|
38
|
+
scope:
|
|
39
|
+
languages: [swift]
|
|
40
|
+
paths: ["**/*.swift"]
|
|
41
|
+
excludePaths: ["**/Generated/**"]
|
|
42
|
+
notCovered:
|
|
43
|
+
objective-c: "no ObjC rules exist here; report as a coverage gap"
|
|
44
|
+
```
|
|
45
|
+
|
|
46
|
+
Per-rule `scope:` narrows this further and never widens it.
|
|
47
|
+
|
|
48
|
+
Without this the 99 Swift rules of the iOS registry would be applied to an Objective-C or UIKit file, which manufactures findings, buries the real ones, and teaches the reader to distrust the run - strictly worse than declaring no coverage. Every dropped rule is counted with its reason in `droppedReasons`, because silently narrowing the rule set is indistinguishable from the code passing it.
|
|
49
|
+
|
|
50
|
+
A registry that declares no scope loads as `reference-only`: readable context for the reviewers, contributing no rule IDs to the denominator.
|
|
51
|
+
|
|
52
|
+
## What is deterministic here, and what is deliberately not
|
|
53
|
+
|
|
54
|
+
**Deterministic (this gate):** which registries apply, which rule IDs are in scope, which changed files no criteria cover, whether the delegated toolchain is wired, and the exception-marker audit.
|
|
55
|
+
|
|
56
|
+
**Not deterministic, on purpose:** the rules themselves. This gate does NOT execute registry `mechanism` patterns. Across the 99-rule iOS registry the distribution is 47 `judgement` / 41 `lint` / 10 `scan` / 1 `format`, only 43 rules carry `mechanism` at all, and the field is prose with an embedded pattern (`swiftlint file_length, function_body_length`, `custom regex per module`) whose scope is written in English. Perhaps 12-15 have an extractable pattern. Running them unscoped floods a review with doc-comment matches - which the registry's own guidance says, requiring each hit be opened and confirmed. The properly engineered version already exists as `references/swiftlint.draft.yml`, 36 regex rules with path scoping.
|
|
57
|
+
|
|
58
|
+
So instead: **if a registry names a toolchain, check whether the repo wired it, run it when it is there, and report its absence as a finding.** That is stack-generic for free, because swiftlint, ktlint, detekt, eslint and ruff all speak rule IDs. An unwired toolchain ranks above most individual violations it would have caught: those rules are unverified, and a review reporting no violations for them is reporting that nothing was measured.
|
|
59
|
+
|
|
60
|
+
## The one bespoke scan: exception markers
|
|
61
|
+
|
|
62
|
+
Cheap, language-agnostic, and it answers "completely" head on - an exception is the author's own claim that a rule does not apply here, so an expired or unexplained one is something the reviewers would otherwise take on trust. The marker template is read from the registry's `exception_marker`, never hardcoded, so a registry using a different comment syntax still works.
|
|
63
|
+
|
|
64
|
+
| Condition | Severity |
|
|
65
|
+
|---|---|
|
|
66
|
+
| Exception expired (`expiry < today`) | blocking |
|
|
67
|
+
| Exception names an ID absent from every registry in scope | blocking |
|
|
68
|
+
| Exception with no rule ID | blocking |
|
|
69
|
+
| Exception with no expiry date | important |
|
|
70
|
+
| Exception with no reason | important |
|
|
71
|
+
|
|
72
|
+
## Dev-mode substitutes (Phases 1 and 2 never ran)
|
|
73
|
+
|
|
74
|
+
| Full-pipeline input | Dev-mode substitute |
|
|
75
|
+
|---|---|
|
|
76
|
+
| `detectedStack` (Phase 1) | language census of the diff by file extension. Describes what the diff CONTAINS, not what the repo is nominally built in, so one Objective-C bridging file in a Swift repo is classified correctly |
|
|
77
|
+
| Phase 1 analysis summary (triage scope) | the task description plus the inline task list Phase 3 generated for itself |
|
|
78
|
+
| Phase 2 plan | same inline task list |
|
|
79
|
+
| `state.evidence.figma[]` (Step 1.8) | recorded as `not-applicable (no Phase 1 evidence)` |
|
|
80
|
+
| Step 2.8 visual conformance | recorded as `not-applicable (no Phase 1 evidence)` |
|
|
81
|
+
|
|
82
|
+
Recording `not-applicable` rather than skipping is the point: "silently skipped" and "out of scope with a reason" must be distinguishable, or the completeness claim cannot be checked.
|
|
83
|
+
|
|
84
|
+
## Invocation
|
|
85
|
+
|
|
86
|
+
```bash
|
|
87
|
+
node $HOME/.claude/scripts/skill-conformance.mjs \
|
|
88
|
+
--diff "$WORKTREE/.review-diff.txt" \
|
|
89
|
+
--state "$WORKTREE/agent-state.json" \
|
|
90
|
+
--repo "$WORKTREE" \
|
|
91
|
+
--out "$WORKTREE/.pipeline/criteria-manifest.json"
|
|
92
|
+
```
|
|
93
|
+
|
|
94
|
+
Exit codes:
|
|
95
|
+
|
|
96
|
+
| Code | Meaning |
|
|
97
|
+
|---|---|
|
|
98
|
+
| `0` | manifest written; may still carry findings or declared coverage gaps |
|
|
99
|
+
| `1` | setup error: no diff, unreadable state, or an explicit `--skills-root` that does not exist |
|
|
100
|
+
| `2` | **fail closed**: no skills root resolved at all, a declared registry could not be parsed, or a declared registry path escaped its skill directory |
|
|
101
|
+
|
|
102
|
+
Any non-zero code halts before the reviewers. Continuing on `2` would drop a whole rule set from the denominator; and because `validate-reviewer.mjs` skips the checklist when zero rules were selected, an empty resolution would otherwise render as a clean review over criteria that were never loaded. An empty registry set inside a root that DOES exist is the different, legitimate case: declared coverage gap, exit 0.
|
|
103
|
+
|
|
104
|
+
`skillsRootsSearched` lists every candidate root that was probed with an `existed` flag, not only the ones found - otherwise the field is empty in exactly the case it exists to explain.
|
|
105
|
+
|
|
106
|
+
Output validates against `$HOME/.claude/schemas/criteria-manifest.schema.json`.
|
|
107
|
+
|
|
108
|
+
## Handoff to the reviewers
|
|
109
|
+
|
|
110
|
+
The manifest's selected rule IDs and registry file paths go into the **shared, cacheable prompt prefix** as one `${CRITERIA}` block, byte-identical for every reviewer (Step 1.9 requires an identical leading block; per-reviewer criteria subsets would invalidate the prefix for the whole panel and pay full input rate on the largest block in the phase). Reviewers read the registry YAML natively - the pipeline hands over a path, not a parse.
|
|
111
|
+
|
|
112
|
+
Deterministic findings merge at Step 3.0 alongside the test-integrity set, so triage adjudicates them rather than never seeing them. A finding citing a registry rule ID carries evidence a reviewer opinion does not, and the triage prompt says so: an ID-citing finding is not dismissible as a matter of taste, though it can still be out of scope for this task.
|
|
113
|
+
|
|
114
|
+
## Preference
|
|
115
|
+
|
|
116
|
+
`prefs.global.skillConformance.blockOnCoverageGap` (default **false**). A coverage gap on a stack with no registry would otherwise block every non-iOS run from day one. There is deliberately no opt-out for the stage itself or for the exception-expiry check, on the same grounds as Step 1.76: a run that can switch off its own anti-reward-hacking control cannot be trusted to report a pass.
|
|
@@ -38,4 +38,4 @@ Adds one Sonnet call plus up to `maxFindings` single-test runs (and build-lock c
|
|
|
38
38
|
|
|
39
39
|
## Reference
|
|
40
40
|
|
|
41
|
-
Wiring: `$HOME/.claude/multi-agent-refs/phases/phase-4-review.md` Step 3.7. Schema:
|
|
41
|
+
Wiring: `$HOME/.claude/multi-agent-refs/phases/phase-4-review.md` Step 3.7. Schema: `$HOME/.claude/schemas/triage-output.schema.json` v3.2.0 (`$defs.verification`). Evidence gate: `$HOME/.claude/scripts/evidence-gate.mjs`. Prefs: `prefs.global.verifyByTest` in `$HOME/.claude/schemas/prefs.schema.json`.
|
|
@@ -140,7 +140,7 @@ curl -s -H "Authorization: Bearer $TOKEN" \
|
|
|
140
140
|
Each source is failure-isolated and never blocks issue creation.
|
|
141
141
|
|
|
142
142
|
**6a. Figma** (only when `FIGMA_URL` is set):
|
|
143
|
-
Follow the 3-tier chain from `rules/figma-pipeline.md`: Tier 1 MCP (`get_design_context` / `get_screenshot`, one re-auth retry) → Tier 2 REST (`keychainMapping.
|
|
143
|
+
Follow the 3-tier chain from `rules/figma-pipeline.md`: Tier 1 MCP (`get_design_context` / `get_screenshot`, one re-auth retry) → Tier 2 REST (`keychainMapping.figma`) → Tier 3 ask the user for a screenshot or proceed link-only.
|
|
144
144
|
|
|
145
145
|
> This command is standalone (same class as `/multi-agent:analysis`), so Figma MCP use here is allowed - it is not a dev-phase violation under the "No MCP outside analysis phase" rule.
|
|
146
146
|
|
|
@@ -195,7 +195,7 @@ Memories that reference integration hosts NEVER get auto-pruned - even if a co
|
|
|
195
195
|
|
|
196
196
|
## Schema
|
|
197
197
|
|
|
198
|
-
`prefs.schema.json` adds `global.multiRepoIntegrationHosts` - see
|
|
198
|
+
`prefs.schema.json` adds `global.multiRepoIntegrationHosts` - see `$HOME/.claude/schemas/prefs.schema.json` for the authoritative definition. Required fields: `repoSet` (array ≥2), one of `(hostPath + platform)` OR `(noHost: true)`. Optional: `hostScheme`, `submodulePaths`, `resolveCommand`, `buildCommand`, `lastUsed`, `count`, `lastResult`.
|
|
199
199
|
|
|
200
200
|
## Smoke coverage
|
|
201
201
|
|
|
@@ -62,7 +62,7 @@ Verdict: unverified (2 reviewers)
|
|
|
62
62
|
|
|
63
63
|
## Cost Breakdown
|
|
64
64
|
|
|
65
|
-
(emit by Phase 7 via
|
|
65
|
+
(emit by Phase 7 via `$HOME/.claude/scripts/render-agent-log-cost.sh <task-id>`. Renders unconditionally on every run. If the renderer exits 2 (no tracker data + no OTel spans), Phase 7 omits this section without failing the run.)
|
|
66
66
|
|
|
67
67
|
| Phase | Model | Tokens in | Tokens out | Est. USD |
|
|
68
68
|
| ----- | ----- | --------- | ---------- | -------- |
|
|
@@ -89,7 +89,7 @@ Verdict: unverified (2 reviewers)
|
|
|
89
89
|
Phase 7 MUST attempt to render the Cost Breakdown section as part of the agent-log compose step:
|
|
90
90
|
|
|
91
91
|
```bash
|
|
92
|
-
COST_BLOCK=$(bash
|
|
92
|
+
COST_BLOCK=$(bash $HOME/.claude/scripts/render-agent-log-cost.sh "$TASK_ID" 2>/dev/null) && \
|
|
93
93
|
printf '%s\n' "$COST_BLOCK" >> "$AGENT_LOG"
|
|
94
94
|
```
|
|
95
95
|
|
|
@@ -100,9 +100,9 @@ Emission is best-effort - exit 2 (no data) is silently skipped. Never fail the
|
|
|
100
100
|
Every phase that dispatches a billable LLM agent MUST forward its token totals to the tracker. The minimal contract (already enforced via `smoke-tracker-contract.sh` for Phase 4):
|
|
101
101
|
|
|
102
102
|
```bash
|
|
103
|
-
|
|
103
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" <phase-id> <event> \
|
|
104
104
|
model=<fable|opus|sonnet|haiku|gpt-5.4|gpt-5.6|gpt-5.6-terra> tokens_in=$IN tokens_out=$OUT tokens_cached=$CACHED duration_ms=$DUR
|
|
105
|
-
LOG_METRIC_FORWARD_TO_TRACKER=1
|
|
105
|
+
LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" <phase-id> tokens \
|
|
106
106
|
model=<...> tokens_in=$IN tokens_out=$OUT tokens_cached=$CACHED
|
|
107
107
|
```
|
|
108
108
|
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
> **TLDR** - Three orthogonal modifier flags that can combine freely:
|
|
2
2
|
>
|
|
3
3
|
> - `autopilot` → skips confirmations (Plan, Test, Commit, PR prompts), still fails safe on review blockers and build retries.
|
|
4
|
-
> - `--dev` → strips to Init → Dev(Opus self-contained) → Test → Commit → Report. No Phase 1/2
|
|
4
|
+
> - `--dev` → strips to Init → Dev(Opus self-contained) → Review → Test → Commit → Report. No Phase 1/2. Phase 4 Review runs in every `--dev` combination; Phase 5 Test is an interactive prompt (suppressed only by `autopilot`).
|
|
5
5
|
> - `--local` → no worktree, works directly in `$PROJECT_ROOT` on a local branch.
|
|
6
6
|
>
|
|
7
7
|
> Combined: `--dev --local autopilot` = fastest, least-friction path.
|
|
@@ -59,7 +59,7 @@ Full contract: `$HOME/.claude/multi-agent-refs/phases/phase-7-report.md` (Autopi
|
|
|
59
59
|
|
|
60
60
|
## Dev-Only Mode (`--dev`)
|
|
61
61
|
|
|
62
|
-
Dev-only mode strips the pipeline to
|
|
62
|
+
Dev-only mode strips the pipeline to what a scoped task needs: no deep analysis, no planning phase. The **Opus** dev agent directly analyzes the task scope and implements, and Phase 4 then reviews what it produced. Review is deliberately NOT part of the strip: analysis and planning shape work that has not happened yet, so a task the user has already scoped can skip them, while review judges work that now exists and has no substitute.
|
|
63
63
|
|
|
64
64
|
**Activation**: Append `--dev` to any pipeline command:
|
|
65
65
|
|
|
@@ -75,7 +75,7 @@ Dev-only mode strips the pipeline to minimum: no deep analysis, no planning phas
|
|
|
75
75
|
**Pipeline in dev-only mode:**
|
|
76
76
|
|
|
77
77
|
```
|
|
78
|
-
Phase 0: Init -> Phase 3: Dev (self-contained) -> Phase 5: Test -> Phase 6: Commit -> Phase 7: Report
|
|
78
|
+
Phase 0: Init -> Phase 3: Dev (self-contained) -> Phase 4: Review -> Phase 5: Test -> Phase 6: Commit -> Phase 7: Report
|
|
79
79
|
```
|
|
80
80
|
|
|
81
81
|
**What changes in dev-only (Tablo 2 - Dev-Only modes):**
|
|
@@ -86,10 +86,10 @@ Phase 0: Init -> Phase 3: Dev (self-contained) -> Phase 5: Test -> Phase 6: Comm
|
|
|
86
86
|
| Phase 1 (Analysis) | Parallel Explore agents | **SKIP** | **SKIP** | **SKIP** | **SKIP** |
|
|
87
87
|
| Phase 2 (Planning) | TaskCreate + architecture review + **Plan Approval Gate** (clarification max 2 rounds + approval loop) | **SKIP** (plan gate not applicable - fast path) | **SKIP** | **SKIP** | **SKIP** |
|
|
88
88
|
| Phase 3 (Dev) | Follows Phase 2 plan, TDD cycle (Sonnet) | **Self-contained** (Opus): agent scans relevant files, implements with TDD, builds | Same as `--dev` | Same as `--dev` | Same as `--dev` |
|
|
89
|
-
| Phase 4 (Review) | Parallel review + Fable triage (Claude: 2-model / Copilot: 3-model) | **
|
|
89
|
+
| Phase 4 (Review) | Parallel review + Fable triage (Claude: 2-model / Copilot: 3-model) | **Same** - gates, parallel review, triage; blocking findings return to Phase 3 (cap 3) | **Same**, blocking findings auto-fixed; circuit breaker halts at 3 cycles | **Same as `--dev`**, on the local branch diff | **Same as `--dev autopilot`** |
|
|
90
90
|
| Phase 5 (User Test) | Interactive prompt (ask user to test) | **Interactive prompt** (ask user to test) | **Skip** (autopilot suppresses interactive prompts) | **Interactive prompt** (same as `--dev`) | **Skip** (same as `--dev autopilot`) |
|
|
91
91
|
| Phase 6 (Commit) | Commit + PR | Same - still asks (unless autopilot) | Auto commit + push + PR | Same as `--dev` | Auto commit + push + PR |
|
|
92
|
-
| Phase 7 (Report) | Full report + channels multi-select | Simplified - no
|
|
92
|
+
| Phase 7 (Report) | Full report + channels multi-select | Simplified - no analysis section, review section IS present, BUT **channels menu still pauses** | Channels menu **STILL PAUSES** | Same as `--dev` | Channels menu **STILL PAUSES** |
|
|
93
93
|
|
|
94
94
|
**Phase 3 in dev-only mode (self-contained):**
|
|
95
95
|
|
|
@@ -101,9 +101,9 @@ The **Opus** agent receives the task description (from Jira, GitHub issue, or fr
|
|
|
101
101
|
|
|
102
102
|
No separate task breakdown - the agent handles scope autonomously.
|
|
103
103
|
|
|
104
|
-
**State tracking**: `agent-state.json` gets `"onlyDevelop": true`. Phases 1
|
|
104
|
+
**State tracking**: `agent-state.json` gets `"onlyDevelop": true`. Phases 1 and 2 are marked `"skipped"` instead of `"pending"`. Phase 4 stays `"pending"` and runs. Phase 5 stays `"pending"` and runs interactively unless `autopilot` is also set, in which case it is marked `"skipped"`.
|
|
105
105
|
|
|
106
|
-
**Combinable with autopilot**: `--dev autopilot` = fastest path. Init -> Dev -> auto-commit -> auto-PR -> Report (Phase 5 skipped). Zero user interaction (except build failures after 3 retries).
|
|
106
|
+
**Combinable with autopilot**: `--dev autopilot` = fastest path. Init -> Dev -> Review (auto-fix) -> auto-commit -> auto-PR -> Report (Phase 5 skipped). Zero user interaction (except build failures after 3 retries, and review findings that survive 3 rework cycles).
|
|
107
107
|
|
|
108
108
|
---
|
|
109
109
|
|
|
@@ -8,11 +8,11 @@ Before anything else - initialize the visual tracker so the user sees the pipe
|
|
|
8
8
|
|
|
9
9
|
```bash
|
|
10
10
|
TASK_ID="${INPUT_TASK_ID:-pipeline-$(date +%Y%m%d-%H%M%S)}"
|
|
11
|
-
|
|
11
|
+
$HOME/.claude/scripts/phase-tracker.sh init "$TASK_ID"
|
|
12
12
|
for p in 0:Init 1:Analysis 2:Planning 3:Dev 4:Review 5:Test 6:Commit 7:Report; do
|
|
13
|
-
|
|
13
|
+
$HOME/.claude/scripts/phase-tracker.sh add "${p%%:*}" "${p#*:}"
|
|
14
14
|
done
|
|
15
|
-
|
|
15
|
+
$HOME/.claude/scripts/phase-tracker.sh update 0 in_progress
|
|
16
16
|
```
|
|
17
17
|
|
|
18
18
|
If `INPUT_TASK_ID` isn't known yet (free-text, project not selected), use a placeholder; rename later via `mv` once parsed in Step 1.
|
|
@@ -85,22 +85,24 @@ When Step 1 input parsing surfaces a Figma reference (URL, node ID, or "from the
|
|
|
85
85
|
|
|
86
86
|
Probe order:
|
|
87
87
|
|
|
88
|
-
1. **Tier 1 (Figma MCP)**:
|
|
89
|
-
2. **Tier 2 (Figma REST)**: when Tier 1 fails, resolve the PAT via `~/.claude/lib/credential-store.sh get <logical-key>` where `<logical-key>` = `prefs.global.keychainMapping.
|
|
88
|
+
1. **Tier 1 (Figma MCP)**: check the host serves `mcp__claude_ai_Figma__*` before probing. Absent → set `state.figmaAccess.tier1Unavailable = "host"` and fall through to Tier 2 with no probe, no re-auth retry, no MCP-token question. Present → probe `get_metadata(fileKey, nodeId)` on the first frame; on auth failure run `authenticate` + `complete_authentication` and retry once, and only a *second* failure raises the recreate-or-continue question. Success → `state.figmaAccess.tier = 1`.
|
|
89
|
+
2. **Tier 2 (Figma REST)**: when Tier 1 fails, resolve the PAT via `~/.claude/lib/credential-store.sh get <logical-key>` where `<logical-key>` = `prefs.global.keychainMapping.figma`. Probe `GET https://api.figma.com/v1/files/{fileKey}/nodes?ids={nodeId}` with header `X-Figma-Token: $TOKEN`. HTTP 200 → `state.figmaAccess.tier = 2`. Token missing / 401 / 403 → fall through.
|
|
90
90
|
3. **Tier 3 (User-attached screenshot)**: when Tiers 1 + 2 both fail, scan the task payload for inline screenshots or attachments. Present → `state.figmaAccess.tier = 3` and `state.figmaAccess.reviewBlocking = true` (Phase 4 enforces this).
|
|
91
91
|
4. **Halt**: all three tiers fail → emit a single AskUserQuestion asking the user how to proceed (provide PAT, paste a screenshot, abort). Never proceed with text-derived guesses.
|
|
92
92
|
|
|
93
|
+
Record the cause, not just the downshift. `tier1Unavailable = "host"` means the tier never existed here - routine on Copilot and Codex, where the installer registers only `dev-toolkit`. `"auth"` means it existed and the credential failed, which on Claude Code points at a dead `figma_mcp` token worth surfacing in Phase 7. Conflating them costs two wasted MCP round trips and a question the user cannot act on. On those two hosts a mapped `figma` PAT is the primary path, not a fallback.
|
|
94
|
+
|
|
93
95
|
Log the resolved tier in the agent log:
|
|
94
96
|
|
|
95
97
|
```
|
|
96
98
|
→ figma access tier: <1|2|3>
|
|
97
99
|
```
|
|
98
100
|
|
|
99
|
-
Full chain definition, REST endpoints, URL parsing, canonical-component contract:
|
|
101
|
+
Full chain definition, REST endpoints, URL parsing, canonical-component contract: `$HOME/.claude/rules/figma-pipeline.md` "MUST: Figma access - 3-tier fallback chain (BLOCKING, pipeline-wide)". Do not duplicate it here.
|
|
100
102
|
|
|
101
103
|
#### Step 0.6 - Update check (advisory, opt-out via `prefs.global.updateCheck.enabled`)
|
|
102
104
|
|
|
103
|
-
Run `bash
|
|
105
|
+
Run `bash $HOME/.claude/scripts/update-check.sh` (cached per `updateCheck.ttlHours`, default 24h; 3s-bounded; every failure path silent). Empty output → continue. Output `<local>|<latest>` means a newer version exists:
|
|
104
106
|
|
|
105
107
|
- **Interactive**: log `→ update available: v<local> -> v<latest>`, ask ONE AskUserQuestion - **Update now** (recommended) / **Continue without updating**. On *Update now* (or `updateCheck.autoUpdate: true`, which skips the question): run the `/multi-agent:update` flow, log `→ updated to v<latest>`, continue the run (note in the log: already-loaded phase docs finish this run on the old version; full effect next session). On *Continue*: no re-ask until the TTL expires.
|
|
106
108
|
- **Autopilot**: never ask (zero-interaction contract). Log `→ update available: ... (log-only; run /multi-agent:update)` and continue - unless `autoUpdate: true`, then update silently first.
|
|
@@ -189,7 +191,7 @@ Classify and fetch external data:
|
|
|
189
191
|
|
|
190
192
|
0. **Intent guard (conceptual-vs-edit)** - gated by `prefs.global.intentGuard.enabled` (default `true`). Before any project selection, worktree, or Jira prompt, classify the input:
|
|
191
193
|
```bash
|
|
192
|
-
INTENT=$(bash
|
|
194
|
+
INTENT=$(bash $HOME/.claude/lib/classify-intent.sh "$DESCRIPTION")
|
|
193
195
|
```
|
|
194
196
|
- `question` -> the user asked something conceptual, not a task to implement. Do NOT create a branch/worktree/Jira. Surface a picker (picker-contract): **Answer here** (default) / **Treat as a task**. On "Answer here" (autopilot default for `question`), answer the question directly in chat and end the run cleanly - no dev chain, no commits. On "Treat as a task", fall through to step 1 below.
|
|
195
197
|
- `ambiguous` or `task` -> proceed to step 1 (normal task flow). Ambiguous input is treated as a task; the guard never blocks an actionable request.
|
|
@@ -526,7 +528,7 @@ Log: `Phase 0 Step 7: taskType = {component|bugfix|feature|refactor|chore}`
|
|
|
526
528
|
**Gated by `prefs.global.clarifyAmbiguous.enabled`** (default: `false`). When enabled and `state.maturity.status != "blocker"`:
|
|
527
529
|
|
|
528
530
|
1. Dispatch `agents/task-clarifier.md` (Haiku by default) with the task title + body + acceptance + maturity warnings already on `agent-state`.
|
|
529
|
-
2. The agent returns JSON conforming to
|
|
531
|
+
2. The agent returns JSON conforming to `$HOME/.claude/schemas/clarify-output.schema.json` - `clarityScore` (0-10), `questions[]`, `stopAndAsk`.
|
|
530
532
|
3. If `clarityScore >= prefs.clarifyAmbiguous.minScoreToProceed` (default 6) or `stopAndAsk == false` → write `state.clarification` (score + rationale, no questions), proceed to Phase 1 silently.
|
|
531
533
|
4. If `stopAndAsk == true`:
|
|
532
534
|
|
|
@@ -543,7 +545,7 @@ Log: `Phase 0 Step 7: taskType = {component|bugfix|feature|refactor|chore}`
|
|
|
543
545
|
|
|
544
546
|
**Cost:** ~$0.0025 per Haiku call. The pipeline's other expensive phases (Phase 4 reviewers, Phase 3 Sonnet codegen) far outweigh this - the value is avoiding the ~30 min wasted when Phase 3 builds the wrong thing because Phase 0 didn't ask.
|
|
545
547
|
|
|
546
|
-
**Reference:** see
|
|
548
|
+
**Reference:** see `$HOME/.claude/agents/task-clarifier.md` for the full scoring rubric and question-quality rules.
|
|
547
549
|
|
|
548
550
|
**Why this fits Phase 0 (not a new phase):** clarification doesn't change what code gets written - it changes what gets understood before code is written. Phase 0 already collects identity / project / branch / maturity; ambiguity scoring fits naturally as the last contextual gate.
|
|
549
551
|
|
|
@@ -552,7 +554,7 @@ Log: `Phase 0 Step 7: taskType = {component|bugfix|feature|refactor|chore}`
|
|
|
552
554
|
After each clarifier call:
|
|
553
555
|
|
|
554
556
|
```bash
|
|
555
|
-
LOG_METRIC_FORWARD_TO_TRACKER=1
|
|
557
|
+
LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 0 clarify.call \
|
|
556
558
|
model=haiku score=$SCORE questions=$Q stop_and_ask=$STOP autopilot_mode=$AP \
|
|
557
559
|
duration_ms=$D tokens_in=$TI tokens_out=$TO
|
|
558
560
|
```
|