@mmerterden/multi-agent-pipeline 13.5.1 → 14.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +317 -0
- package/README.md +3 -3
- package/docs/features.md +1 -1
- package/install/_common.mjs +73 -0
- package/install/_mcp-register.mjs +70 -31
- package/install/_plugin-skills.mjs +73 -14
- package/install/claude.mjs +28 -4
- package/install/codex.mjs +33 -2
- package/install/copilot.mjs +145 -9
- package/install/index.mjs +10 -6
- package/install/templates/copilot-instructions.md +1 -1
- package/package.json +1 -1
- package/pipeline/agents/code-reviewer.md +58 -1
- package/pipeline/commands/multi-agent/SKILL.md +7 -5
- package/pipeline/commands/multi-agent/analysis/SKILL.md +7 -7
- package/pipeline/commands/multi-agent/analysis-resolve/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/build-optimize/SKILL.md +7 -7
- package/pipeline/commands/multi-agent/channels/SKILL.md +5 -5
- package/pipeline/commands/multi-agent/dev/SKILL.md +23 -18
- package/pipeline/commands/multi-agent/dev-autopilot/SKILL.md +19 -13
- package/pipeline/commands/multi-agent/dev-local/SKILL.md +14 -12
- package/pipeline/commands/multi-agent/dev-local-autopilot/SKILL.md +17 -12
- package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/help/SKILL.md +4 -4
- package/pipeline/commands/multi-agent/ios-coding-standard/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +4 -4
- package/pipeline/commands/multi-agent/log/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/resume/SKILL.md +2 -1
- package/pipeline/commands/multi-agent/review/SKILL.md +5 -5
- package/pipeline/commands/multi-agent/scan/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/search/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/setup/SKILL.md +6 -6
- package/pipeline/commands/multi-agent/{finish → ship}/SKILL.md +12 -12
- package/pipeline/commands/multi-agent/status/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/sync/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/testflight-validation/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/update/SKILL.md +5 -2
- package/pipeline/commands/sim-test.md +2 -2
- package/pipeline/lib/credential-store-resolver.sh +16 -0
- package/pipeline/lib/credential-store.sh +47 -4
- package/pipeline/lib/fetch-figma-annotations.sh +26 -28
- package/pipeline/lib/figma-screenshot.sh +28 -39
- package/pipeline/lib/figma-token.sh +63 -0
- package/pipeline/multi-agent-refs/analysis-template.md +1 -1
- package/pipeline/multi-agent-refs/android-guide.md +1 -1
- package/pipeline/multi-agent-refs/channels/issue-comment.md +1 -1
- package/pipeline/multi-agent-refs/component-dispatch.md +2 -2
- package/pipeline/multi-agent-refs/cross-cli-contract.md +4 -4
- package/pipeline/multi-agent-refs/features/dev-critic.md +2 -2
- package/pipeline/multi-agent-refs/features/model-fallback.md +35 -2
- package/pipeline/multi-agent-refs/features/plan-todos.md +1 -1
- package/pipeline/multi-agent-refs/features/repo-map.md +1 -1
- package/pipeline/multi-agent-refs/features/review-multi-repo.md +3 -3
- package/pipeline/multi-agent-refs/features/shadow-git.md +1 -1
- package/pipeline/multi-agent-refs/features/skill-conformance.md +116 -0
- package/pipeline/multi-agent-refs/features/stack-skill-routing.md +72 -0
- package/pipeline/multi-agent-refs/features/verify-by-test.md +1 -1
- package/pipeline/multi-agent-refs/features/worktree-finalize.md +66 -0
- package/pipeline/multi-agent-refs/generate-issue.md +1 -1
- package/pipeline/multi-agent-refs/multi-repo-integration-build.md +1 -1
- package/pipeline/multi-agent-refs/phases/log-format.md +4 -4
- package/pipeline/multi-agent-refs/phases/modes.md +7 -7
- package/pipeline/multi-agent-refs/phases/phase-0-init.md +13 -11
- package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +17 -15
- package/pipeline/multi-agent-refs/phases/phase-2-planning.md +7 -7
- package/pipeline/multi-agent-refs/phases/phase-3-dev.md +29 -14
- package/pipeline/multi-agent-refs/phases/phase-4-review.md +90 -58
- package/pipeline/multi-agent-refs/phases/phase-5-test.md +7 -7
- package/pipeline/multi-agent-refs/phases/phase-6-commit.md +25 -12
- package/pipeline/multi-agent-refs/phases/phase-7-report.md +12 -9
- package/pipeline/multi-agent-refs/phases.md +13 -13
- package/pipeline/multi-agent-refs/progress-contract.md +2 -2
- package/pipeline/multi-agent-refs/rules.md +7 -5
- package/pipeline/multi-agent-refs/swiftui-guide.md +1 -1
- package/pipeline/multi-agent-refs/tracker-contract.md +16 -15
- package/pipeline/preferences-template.json +9 -2
- package/pipeline/rules/figma-pipeline.md +2 -2
- package/pipeline/schemas/agent-state.schema.json +346 -79
- package/pipeline/schemas/criteria-manifest.schema.json +398 -0
- package/pipeline/schemas/migrations/prefs-2.4.0-to-2.5.0.mjs +64 -0
- package/pipeline/schemas/prefs.schema.json +123 -262
- package/pipeline/schemas/reviewer-output.schema.json +48 -3
- package/pipeline/schemas/token-budget.json +34 -10
- package/pipeline/schemas/triage-output.schema.json +112 -27
- package/pipeline/scripts/cost-table.json +7 -4
- package/pipeline/scripts/gc-worktrees.sh +3 -2
- package/pipeline/scripts/gen-mode-dispatch.mjs +6 -6
- package/pipeline/scripts/match-skills.mjs +37 -4
- package/pipeline/scripts/migrate-prefs.mjs +89 -17
- package/pipeline/scripts/phase-tracker.sh +14 -3
- package/pipeline/scripts/pre-commit-check.sh +49 -2
- package/pipeline/scripts/render-work-summary.sh +51 -3
- package/pipeline/scripts/skill-conformance.mjs +970 -0
- package/pipeline/scripts/smoke-schema-validation.sh +17 -4
- package/pipeline/scripts/test-integrity-gate.mjs +10 -2
- package/pipeline/scripts/uninstall.mjs +35 -9
- package/pipeline/scripts/validate-reviewer.mjs +108 -1
- package/pipeline/scripts/worktree-finalize.sh +299 -0
- package/pipeline/skills/.skill-manifest.json +1 -1
- package/pipeline/skills/.skills-index.json +36 -9
- package/pipeline/skills/shared/README.md +15 -12
- package/pipeline/skills/shared/core/apple-archive-compliance/SKILL.md +1 -0
- package/pipeline/skills/shared/core/apple-archive-compliance/references/rules.yml +167 -0
- package/pipeline/skills/shared/core/google-play-compliance/SKILL.md +1 -0
- package/pipeline/skills/shared/core/google-play-compliance/references/rules.yml +184 -0
- package/pipeline/skills/shared/core/multi-agent/SKILL.md +10 -10
- package/pipeline/skills/shared/core/multi-agent-analysis/SKILL.md +4 -4
- package/pipeline/skills/shared/core/multi-agent-analysis-resolve/SKILL.md +3 -3
- package/pipeline/skills/shared/core/multi-agent-build-optimize/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-create-jira/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-dev/SKILL.md +6 -5
- package/pipeline/skills/shared/core/multi-agent-dev-autopilot/SKILL.md +7 -6
- package/pipeline/skills/shared/core/multi-agent-dev-local/SKILL.md +4 -3
- package/pipeline/skills/shared/core/multi-agent-dev-local-autopilot/SKILL.md +2 -1
- package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-ios-coding-standard/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +4 -4
- package/pipeline/skills/shared/core/multi-agent-review/SKILL.md +5 -5
- package/pipeline/skills/shared/core/multi-agent-scan/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-search/SKILL.md +1 -1
- package/pipeline/skills/shared/core/{multi-agent-finish → multi-agent-ship}/SKILL.md +8 -8
- package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +2 -2
- package/pipeline/skills/shared/external/ios-coding-standard/SKILL.md +44 -5
- package/pipeline/skills/shared/external/ios-coding-standard/modules/_TEMPLATE.yml +82 -0
- package/pipeline/skills/shared/external/ios-coding-standard/references/STANDARD.md +169 -10
- package/pipeline/skills/shared/external/ios-coding-standard/references/rules.yml +335 -16
- package/pipeline/skills/skills-index.md +11 -8
|
@@ -23,10 +23,10 @@
|
|
|
23
23
|
# Tier 2 (Figma REST + PAT) -> primary path here
|
|
24
24
|
# Tier 3 (user screenshot) -> not handled; exit 3 with guidance
|
|
25
25
|
#
|
|
26
|
-
# Token source:
|
|
27
|
-
#
|
|
28
|
-
# FIGMA_PAT environment variable. No literal keychain service name
|
|
29
|
-
# embedded in this script.
|
|
26
|
+
# Token source: figma-token.sh, which resolves the canonical `figma` logical
|
|
27
|
+
# key through credential-store.sh (falling back to the legacy `figma_pat` key,
|
|
28
|
+
# then to the FIGMA_PAT environment variable). No literal keychain service name
|
|
29
|
+
# is embedded in this script.
|
|
30
30
|
#
|
|
31
31
|
# Exit codes:
|
|
32
32
|
# 0 success
|
|
@@ -41,7 +41,6 @@ set -euo pipefail
|
|
|
41
41
|
|
|
42
42
|
# --- Globals ----------------------------------------------------------------
|
|
43
43
|
|
|
44
|
-
PREFS="$HOME/.claude/multi-agent-preferences.json"
|
|
45
44
|
# Locate the resolver with an existence check, not a `.`-chain: sourcing a missing file
|
|
46
45
|
# aborts the shell under `set -e`, `||` included, so a chain skips both its later
|
|
47
46
|
# candidates and its trailing `|| true`. A missing store is tolerated here - CRED_STORE
|
|
@@ -61,6 +60,26 @@ for _cred_resolver in \
|
|
|
61
60
|
done
|
|
62
61
|
unset _cred_resolver
|
|
63
62
|
CRED_STORE="${CRED_STORE:-}"
|
|
63
|
+
|
|
64
|
+
# Tier 2 token lookup lives in one file so it cannot drift per fetcher. Same
|
|
65
|
+
# existence-check discipline as the resolver loop above, and the same tolerance:
|
|
66
|
+
# without it, resolution still works through the FIGMA_PAT env fallback.
|
|
67
|
+
for _figma_token_lib in \
|
|
68
|
+
"$(cd "$(dirname "${BASH_SOURCE[0]:-$0}")" && pwd)/figma-token.sh" \
|
|
69
|
+
"$HOME/.claude/lib/figma-token.sh" \
|
|
70
|
+
"$HOME/.copilot/lib/figma-token.sh" \
|
|
71
|
+
"$HOME/.codex/lib/figma-token.sh"; do
|
|
72
|
+
[ -f "$_figma_token_lib" ] || continue
|
|
73
|
+
# shellcheck source=/dev/null
|
|
74
|
+
. "$_figma_token_lib" 2>/dev/null || true
|
|
75
|
+
if command -v resolve_figma_token >/dev/null 2>&1; then break; fi
|
|
76
|
+
done
|
|
77
|
+
unset _figma_token_lib
|
|
78
|
+
if ! command -v resolve_figma_token >/dev/null 2>&1; then
|
|
79
|
+
resolve_figma_token() { [ -n "${FIGMA_PAT:-}" ] && printf '%s' "$FIGMA_PAT"; }
|
|
80
|
+
figma_token_remediation() { printf 'figma-token.sh not found - reinstall the pipeline, or export FIGMA_PAT.'; }
|
|
81
|
+
fi
|
|
82
|
+
|
|
64
83
|
FIGMA_API="https://api.figma.com/v1"
|
|
65
84
|
PARALLEL_DL=4
|
|
66
85
|
HTTP_TIMEOUT=60
|
|
@@ -105,8 +124,8 @@ Options:
|
|
|
105
124
|
--help Show this help text and exit.
|
|
106
125
|
|
|
107
126
|
Token resolution:
|
|
108
|
-
Read via credential-store.sh using
|
|
109
|
-
|
|
127
|
+
Read via credential-store.sh using prefs.global.keychainMapping.figma
|
|
128
|
+
(legacy figma_pat still honoured). Falls back to FIGMA_PAT env.
|
|
110
129
|
|
|
111
130
|
Output:
|
|
112
131
|
PNG files plus manifest.json in the output directory.
|
|
@@ -151,35 +170,6 @@ print(f"{file_key}\t{node}")
|
|
|
151
170
|
PY
|
|
152
171
|
}
|
|
153
172
|
|
|
154
|
-
resolve_token() {
|
|
155
|
-
# 1) Try the keychain via credential-store.sh, keyed by prefs mapping.
|
|
156
|
-
local token_key=""
|
|
157
|
-
if [ -f "$PREFS" ] && command -v python3 >/dev/null 2>&1; then
|
|
158
|
-
token_key=$(python3 -c "
|
|
159
|
-
import json, sys
|
|
160
|
-
try:
|
|
161
|
-
p = json.load(open('$PREFS'))
|
|
162
|
-
print(p.get('global', {}).get('keychainMapping', {}).get('figma_pat') or '')
|
|
163
|
-
except Exception:
|
|
164
|
-
print('')
|
|
165
|
-
" 2>/dev/null || true)
|
|
166
|
-
fi
|
|
167
|
-
if [ -n "$token_key" ] && [ -x "$CRED_STORE" ]; then
|
|
168
|
-
local tok
|
|
169
|
-
tok=$("$CRED_STORE" get "$token_key" 2>/dev/null || true)
|
|
170
|
-
if [ -n "$tok" ]; then
|
|
171
|
-
printf '%s' "$tok"
|
|
172
|
-
return 0
|
|
173
|
-
fi
|
|
174
|
-
fi
|
|
175
|
-
# 2) Env fallback so CI / one-shot use still works.
|
|
176
|
-
if [ -n "${FIGMA_PAT:-}" ]; then
|
|
177
|
-
printf '%s' "$FIGMA_PAT"
|
|
178
|
-
return 0
|
|
179
|
-
fi
|
|
180
|
-
return 1
|
|
181
|
-
}
|
|
182
|
-
|
|
183
173
|
# curl_figma <output> <url>
|
|
184
174
|
# Honours retries: 429 retries once after a 5s pause; 5xx uses exponential
|
|
185
175
|
# backoff (1s, 2s, 4s). Writes the body to <output>, returns the HTTP code on
|
|
@@ -409,11 +399,10 @@ require_cmd python3
|
|
|
409
399
|
# Tier 1 (Figma MCP) is not reachable from a bash script. Stay on Tier 2 (REST).
|
|
410
400
|
log "INFO: Tier 1 (Figma MCP) not available from shell; using Tier 2 (REST)."
|
|
411
401
|
|
|
412
|
-
FIGMA_TOKEN=$(
|
|
402
|
+
FIGMA_TOKEN=$(resolve_figma_token || true)
|
|
413
403
|
if [ -z "${FIGMA_TOKEN:-}" ]; then
|
|
414
404
|
log "Tier 3 (user-provided screenshot required): no PAT resolved."
|
|
415
|
-
log "Hint:
|
|
416
|
-
log " prefs.global.keychainMapping.figma_pat, or export FIGMA_PAT."
|
|
405
|
+
log "Hint: $(figma_token_remediation)"
|
|
417
406
|
exit 3
|
|
418
407
|
fi
|
|
419
408
|
|
|
@@ -0,0 +1,63 @@
|
|
|
1
|
+
#!/bin/bash
|
|
2
|
+
#
|
|
3
|
+
# figma-token.sh
|
|
4
|
+
# Single resolution point for the Tier 2 Figma Personal Access Token.
|
|
5
|
+
#
|
|
6
|
+
# Sourced by every Tier 2 fetcher (fetch-figma-annotations.sh,
|
|
7
|
+
# figma-screenshot.sh). It exists because those two each carried their own copy
|
|
8
|
+
# of this lookup: when `migrate-prefs.mjs` consolidated
|
|
9
|
+
# `keychainMapping.figma_pat` into `keychainMapping.figma` and deleted the old
|
|
10
|
+
# key, only the migration was updated. Both fetchers kept reading the deleted
|
|
11
|
+
# key, so every migrated install reported `missing-token` while a valid PAT sat
|
|
12
|
+
# under the new name - and the error text told the user to map the key the
|
|
13
|
+
# migration had just removed. One copy of the lookup cannot drift from itself.
|
|
14
|
+
#
|
|
15
|
+
# Tier 2 is only reached when Tier 1 (Figma MCP) is unreachable; see the Figma
|
|
16
|
+
# Access Tier rule in multi-agent-refs/rules.md for the chain and the
|
|
17
|
+
# expired-token decision that gates the move to Tier 3.
|
|
18
|
+
#
|
|
19
|
+
# Resolution order, first non-empty wins:
|
|
20
|
+
# 1. credential-store.sh get figma - the canonical logical key
|
|
21
|
+
# 2. credential-store.sh get figma_pat - pre-v13.6 installs that never migrated
|
|
22
|
+
# 3. $FIGMA_PAT - env fallback for CI and one-shot runs
|
|
23
|
+
#
|
|
24
|
+
# Logical keys are passed through verbatim. credential-store.sh owns the
|
|
25
|
+
# `prefs.global.keychainMapping` indirection, so resolving the mapping here as
|
|
26
|
+
# well would reintroduce the duplication this file exists to remove. No literal
|
|
27
|
+
# Keychain service name appears here, which keeps the Synced Command Hygiene
|
|
28
|
+
# rule satisfied.
|
|
29
|
+
#
|
|
30
|
+
# Requires `$CRED_STORE` to already be resolved (via credential-store-resolver.sh).
|
|
31
|
+
# An unset or non-executable CRED_STORE is tolerated: resolution falls through to
|
|
32
|
+
# the env fallback so a keychain-less CI box still works.
|
|
33
|
+
|
|
34
|
+
# The canonical key first, the legacy key second. Space-separated so the loop
|
|
35
|
+
# below stays POSIX-ish and works under bash 3.2 (stock macOS).
|
|
36
|
+
FIGMA_TOKEN_LOGICAL_KEYS="${FIGMA_TOKEN_LOGICAL_KEYS:-figma figma_pat}"
|
|
37
|
+
|
|
38
|
+
# resolve_figma_token
|
|
39
|
+
# Prints the token on stdout and returns 0, or prints nothing and returns 1.
|
|
40
|
+
resolve_figma_token() {
|
|
41
|
+
local key tok
|
|
42
|
+
if [ -n "${CRED_STORE:-}" ] && [ -x "${CRED_STORE:-}" ]; then
|
|
43
|
+
for key in $FIGMA_TOKEN_LOGICAL_KEYS; do
|
|
44
|
+
tok=$("$CRED_STORE" get "$key" 2>/dev/null || true)
|
|
45
|
+
if [ -n "$tok" ]; then
|
|
46
|
+
printf '%s' "$tok"
|
|
47
|
+
return 0
|
|
48
|
+
fi
|
|
49
|
+
done
|
|
50
|
+
fi
|
|
51
|
+
if [ -n "${FIGMA_PAT:-}" ]; then
|
|
52
|
+
printf '%s' "$FIGMA_PAT"
|
|
53
|
+
return 0
|
|
54
|
+
fi
|
|
55
|
+
return 1
|
|
56
|
+
}
|
|
57
|
+
|
|
58
|
+
# figma_token_remediation
|
|
59
|
+
# The one place the "how do I fix this" sentence is written, so a fetcher can
|
|
60
|
+
# never name a key the migration deleted.
|
|
61
|
+
figma_token_remediation() {
|
|
62
|
+
printf 'no Figma PAT for Tier 2. Map prefs.global.keychainMapping.figma (run /multi-agent:setup), or export FIGMA_PAT.'
|
|
63
|
+
}
|
|
@@ -10,7 +10,7 @@ description: "v3 canonical template for /multi-agent:analysis. 23 main sections
|
|
|
10
10
|
|
|
11
11
|
> **Language**: This file is read as a system prompt. Prose stays English. Example tables and headings carry bilingual TR / EN scaffolds; the runtime renderer picks the right column per `outputLanguage`.
|
|
12
12
|
|
|
13
|
-
> **Locked decisions that govern this template**: the canonical numbered list lives in
|
|
13
|
+
> **Locked decisions that govern this template**: the canonical numbered list lives in `$HOME/.claude/commands/multi-agent/analysis/SKILL.md` (cite decisions by label, not by a number duplicated here, to avoid drift). The template's structure is shaped chiefly by: one-feature-per-run, section omission rule, citation discipline (incl. annotation-as-copy), forward-looking spec, humanizer punctuation policy, standards binding, per-platform output split, repo-evidence reuse-first, Figma 3-tier access, Gherkin user stories, Goals + Non-Goals paired, SVG default, Files-to-Add tag, API response variants exhaustive, screenshots embedded, all Figma variants drilled, localization mode (ownership-aware), References at bottom, platform-agnostic + Pass B render, convention extraction, Pass B footnote mandatory, Lite mode, SwiftUI Preview block (iOS), variant usage explicit, analysis self-contained (no MCP downstream), and business-rule to acceptance-criterion to test traceability.
|
|
14
14
|
|
|
15
15
|
## Mode selection - Full vs Lite
|
|
16
16
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
## Android/Kotlin Component Generation Guide
|
|
2
2
|
|
|
3
|
-
> **MUST: Figma MCP-first (BLOCKING).** If the task references any Figma frame (URL, node ID, or "from the design"), the Dev phase MUST call `mcp__claude_ai_Figma__get_design_context` for every frame BEFORE writing a single Composable line. Use the `CodeConnectSnippet` component name verbatim - no sound-alike substitutions. Authentication failure is not a skip path. Full rule, trigger conditions, and gate failure modes:
|
|
3
|
+
> **MUST: Figma MCP-first (BLOCKING).** If the task references any Figma frame (URL, node ID, or "from the design"), the Dev phase MUST call `mcp__claude_ai_Figma__get_design_context` for every frame BEFORE writing a single Composable line. Use the `CodeConnectSnippet` component name verbatim - no sound-alike substitutions. Authentication failure is not a skip path. Full rule, trigger conditions, and gate failure modes: `$HOME/.claude/rules/figma-pipeline.md` "MUST: Figma MCP-first (BLOCKING)". Phase wiring: `$HOME/.claude/multi-agent-refs/phases/phase-3-dev.md` "MUST: Figma MCP-first (BLOCKING pre-step)".
|
|
4
4
|
|
|
5
5
|
When the task involves creating an Android UI component (Jetpack Compose), follow this architecture.
|
|
6
6
|
|
|
@@ -112,7 +112,7 @@ Token resolution: `gh auth status` for the active GitHub account selected in Pha
|
|
|
112
112
|
|
|
113
113
|
## Pairing with the Progress flag updater
|
|
114
114
|
|
|
115
|
-
Every issue comment post is paired with
|
|
115
|
+
Every issue comment post is paired with `$HOME/.claude/scripts/update-issue-progress.sh "$TASK_ID"` in the same Phase 7 step. Order: comment FIRST (so the timestamp marks the run), flags SECOND (so the body diff is one logical change).
|
|
116
116
|
|
|
117
117
|
```bash
|
|
118
118
|
# Phase 7 Step 5 - issue channel
|
|
@@ -111,9 +111,9 @@ On failure (the plugin skill returns an unrecoverable build/test error, or the d
|
|
|
111
111
|
|
|
112
112
|
## `--dev` mode behaviour
|
|
113
113
|
|
|
114
|
-
When multi-agent is invoked with `--dev` (fast path: Init → Dev → Commit → Report), the dispatch layer passes `mode: "dev"` so the plugin skill can elide unit tests and wiki (structural + snapshot still required; wiki deferred to Phase 7). If the plugin does not honor a `mode` hint, dispatch simply skips the post-build wiki step itself.
|
|
114
|
+
When multi-agent is invoked with `--dev` (fast path: Init → Dev → Review → Commit → Report), the dispatch layer passes `mode: "dev"` so the plugin skill can elide unit tests and wiki (structural + snapshot still required; wiki deferred to Phase 7). If the plugin does not honor a `mode` hint, dispatch simply skips the post-build wiki step itself.
|
|
115
115
|
|
|
116
|
-
Phase 4 reviewer
|
|
116
|
+
Phase 4 runs in `--dev` as it does in the full pipeline, and its reviewer count is **not** Phase 3's concern - the Step 1.77 scope gate decides that from diff risk, independently of `mode`. What the dispatch layer owes Phase 4 is the record of which plugin skill it delegated to, appended to `state.telemetry.skillCalls[]`, so the review can check the delivered component against the criteria that skill imposes.
|
|
117
117
|
|
|
118
118
|
## Cross-CLI behaviour (intentional divergence)
|
|
119
119
|
|
|
@@ -10,10 +10,10 @@
|
|
|
10
10
|
|
|
11
11
|
```
|
|
12
12
|
analysis, analysis-resolve, autopilot, build-optimize, channels, create-jira, design-check, dev,
|
|
13
|
-
dev-autopilot, dev-local, dev-local-autopilot, diff-explain,
|
|
13
|
+
dev-autopilot, dev-local, dev-local-autopilot, diff-explain, forget, garbage-collect,
|
|
14
14
|
help, ios-coding-standard, issue, jira, kill, language, local,
|
|
15
15
|
local-autopilot, log, manual-test, prune-logs, purge, refactor, resume, review, review-issue, review-jira,
|
|
16
|
-
routines, save, scan, search, setup, stack, status, sync, test, testflight-validation, uninstall, update
|
|
16
|
+
routines, save, scan, search, setup, ship, stack, status, sync, test, testflight-validation, uninstall, update
|
|
17
17
|
```
|
|
18
18
|
|
|
19
19
|
Categories:
|
|
@@ -21,8 +21,8 @@ Categories:
|
|
|
21
21
|
- **Interactive pickers** (single-purpose, not modes): `jira`, `issue`
|
|
22
22
|
- **Issue generator** (one-shot, no worktree, asks type Task/Bug/Story, hard approval gate before create): `create-jira`
|
|
23
23
|
- **Full 8-phase modes**: `autopilot`, `local`, `local-autopilot`
|
|
24
|
-
- **Fast modes** (Init -> Dev(Opus) -> Commit -> Report): `dev`, `dev-autopilot`, `dev-local`, `dev-local-autopilot`
|
|
25
|
-
- **Tail modes** (run the pipeline tail over already-done local work): `
|
|
24
|
+
- **Fast modes** (Init -> Dev(Opus) -> Review -> Commit -> Report): `dev`, `dev-autopilot`, `dev-local`, `dev-local-autopilot`
|
|
25
|
+
- **Tail modes** (run the pipeline tail over already-done local work): `ship`
|
|
26
26
|
- **Ops commands** (one-shot, no worktree): `status`, `log`, `kill`, `purge`, `uninstall`, `resume`, `review`, `review-jira`, `review-issue`, `analysis`, `analysis-resolve`, `build-optimize`, `channels`, `scan`, `search`, `diff-explain`, `garbage-collect`, `prune-logs`
|
|
27
27
|
- **Local audits** (worktree only to build; no commit, push, PR or channels): `design-check`, `testflight-validation`, `ios-coding-standard`. `testflight-validation` additionally never invokes `altool --upload-app` - a validation run must not be able to ship a build by accident.
|
|
28
28
|
- **Meta-ops**: `setup`, `sync`, `update`, `help`, `refactor`, `test`, `stack`, `manual-test`, `language`
|
|
@@ -7,7 +7,7 @@
|
|
|
7
7
|
1. Dispatch `dev-critic` sub-agent (Sonnet by default, tools: `Read, Grep, Glob, Bash`).
|
|
8
8
|
2. Critic runs 4 deterministic gates (build / lint / test / secrets) - failure of any → `pass: false` with `blocking` finding.
|
|
9
9
|
3. If gates green, critic walks the platform checklist (iOS 13-item / Android Kotlin / Backend generic - selected by Phase 1 `detectedStack`).
|
|
10
|
-
4. Returns schema-validated JSON (
|
|
10
|
+
4. Returns schema-validated JSON (`$HOME/.claude/schemas/dev-critic-output.schema.json`): `{pass, iteration, gates, findings[], escalate}`.
|
|
11
11
|
|
|
12
12
|
## Loop cap (STRICT) - max 2 iterations
|
|
13
13
|
|
|
@@ -41,4 +41,4 @@ Phase 4 is parallelization-with-voting - good for *adversarial* perspectives (
|
|
|
41
41
|
|
|
42
42
|
## Reference
|
|
43
43
|
|
|
44
|
-
See
|
|
44
|
+
See `$HOME/.claude/agents/dev-critic.md` for the full agent specification (gates, checklist enumeration, output schema, severity semantics).
|
|
@@ -10,7 +10,7 @@ without ever editing persona files at runtime.
|
|
|
10
10
|
> **Fable 5 available again (2026-07).** `claude-fable-5` is the top tier once
|
|
11
11
|
> more. Architect (`ios/android/backend-architect`), Reviewer-1 (`code-reviewer`),
|
|
12
12
|
> and triage personas declare `preferredModel: fable` - the deepest-reasoning
|
|
13
|
-
> roles. **If Fable 5 is unavailable, the first fallback is opus (`claude-opus-
|
|
13
|
+
> roles. **If Fable 5 is unavailable, the first fallback is opus (`claude-opus-5`).**
|
|
14
14
|
> Security and other `preferredModel: opus` personas keep opus as their top tier.
|
|
15
15
|
> The `premiumTierUntil` date gate stays in the contract as a generic mechanism
|
|
16
16
|
> for any plan-window-limited premium tier.
|
|
@@ -24,9 +24,42 @@ fable -> opus -> sonnet -> haiku
|
|
|
24
24
|
The orchestrator walks **one step down per trigger** from whichever tier a
|
|
25
25
|
persona prefers: a `fable` persona degrades `fable -> opus` first, then
|
|
26
26
|
`opus -> sonnet`, then `sonnet -> haiku`; an `opus` persona starts at
|
|
27
|
-
`opus -> sonnet`. So a Fable 5 outage transparently promotes Opus
|
|
27
|
+
`opus -> sonnet`. So a Fable 5 outage transparently promotes Opus 5 into the
|
|
28
28
|
architect/reviewer roles with no file edits.
|
|
29
29
|
|
|
30
|
+
### Rung names are the contract; model IDs are not
|
|
31
|
+
|
|
32
|
+
Personas and dispatch read the **rung** (`fable`, `opus`, `sonnet`, `haiku`). The wire
|
|
33
|
+
model ID each rung resolves to lives in exactly two places: `scripts/cost-table.json`
|
|
34
|
+
(`prices.<rung>.modelId`) and the per-host reviewer table in
|
|
35
|
+
`phases/phase-4-review.md`. A generation move edits those; it never renames a rung and
|
|
36
|
+
never touches a persona file.
|
|
37
|
+
|
|
38
|
+
Current resolution: `fable` -> `claude-fable-5`, `opus` -> `claude-opus-5`,
|
|
39
|
+
`sonnet` -> `claude-sonnet-5`, `haiku` -> `claude-haiku-4-5`.
|
|
40
|
+
|
|
41
|
+
**Rate limits do not follow the rung.** Claude Opus 5 draws on a pool separate from
|
|
42
|
+
the combined Opus 4.x pool, so a generation move does not inherit the previous
|
|
43
|
+
generation's headroom. Budget-ceiling downgrades below are unaffected - they read the
|
|
44
|
+
cost ledger, not the provider's quota - but a 429 on the opus rung after a generation
|
|
45
|
+
move is a limits question, not a fallback-contract question.
|
|
46
|
+
|
|
47
|
+
### Behavioural re-tuning the generation move implies
|
|
48
|
+
|
|
49
|
+
The pipeline's reviewer and dev prompts were tuned against the previous generation.
|
|
50
|
+
Three shifts on the current opus rung are worth knowing when reading a run that looks
|
|
51
|
+
different rather than broken, and none of them are bugs in this contract:
|
|
52
|
+
|
|
53
|
+
- **Longer user-facing output.** Effort is not the lever; prompt-level conciseness is.
|
|
54
|
+
Phase 7 report length and reviewer prose are where this shows.
|
|
55
|
+
- **Self-verification without being asked.** Explicit "double-check your work"
|
|
56
|
+
scaffolding now causes over-verification rather than preventing under-verification.
|
|
57
|
+
Phase 4's deterministic gates already carry that load.
|
|
58
|
+
- **Readier subagent delegation.** The previous generation under-reached and needed
|
|
59
|
+
encouragement; this one does not. Phase 1's parallel scan and Phase 4's reviewer
|
|
60
|
+
panel are already bounded by explicit counts, which is the right shape - keep them
|
|
61
|
+
bounded rather than adding "delegate more" guidance.
|
|
62
|
+
|
|
30
63
|
One step down per trigger, walking the ladder until a tier dispatches or the
|
|
31
64
|
floor (`haiku`) is reached. `haiku` is the last-resort floor: a run that reaches
|
|
32
65
|
it is heavily degraded but still makes progress instead of hard-halting because
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Feature: Plan Todos Iteration (Phase 3)
|
|
2
2
|
|
|
3
|
-
**Gated by `prefs.global.planTodos.enabled`** (default: `false`). When enabled and Phase 2 Step 4.5 emitted a `plan.todos[]` (conforming to
|
|
3
|
+
**Gated by `prefs.global.planTodos.enabled`** (default: `false`). When enabled and Phase 2 Step 4.5 emitted a `plan.todos[]` (conforming to `$HOME/.claude/schemas/plan-todos.schema.json`), Phase 3 iterates with the helper instead of walking `tasks[]` directly:
|
|
4
4
|
|
|
5
5
|
```bash
|
|
6
6
|
while next=$(bash "$HOME/.claude/lib/plan-todos.sh" next "$TASK_ID"); [ -n "$next" ]; do
|
|
@@ -14,7 +14,7 @@ if [ "$(jq -r '.global.repoMap.enabled // false' "$PREFS")" = "true" ]; then
|
|
|
14
14
|
args=(--root "$WORKTREE" --budget "$BUDGET" --top "$TOP" --format md)
|
|
15
15
|
[ -n "$INCLUDE" ] && args+=(--include "$INCLUDE")
|
|
16
16
|
[ -n "$EXCLUDE" ] && args+=(--exclude "$EXCLUDE")
|
|
17
|
-
REPO_MAP=$(node
|
|
17
|
+
REPO_MAP=$(node $HOME/.claude/scripts/repo-map.mjs "${args[@]}" 2>/dev/null || echo "")
|
|
18
18
|
fi
|
|
19
19
|
```
|
|
20
20
|
|
|
@@ -57,9 +57,9 @@ And log `review.diff_truncated repo=<name> bytes_dropped=<N>`. Triage receives t
|
|
|
57
57
|
|
|
58
58
|
**Telemetry**: Per-repo build/test gate timings + a single combined review/triage call set:
|
|
59
59
|
```bash
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
60
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 gate.build repo=common status=pass duration_ms=$D
|
|
61
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 gate.build repo=uicomponents status=pass duration_ms=$D
|
|
62
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 4 review.combined_diff repos=2 bytes=$BYTES truncated=false
|
|
63
63
|
```
|
|
64
64
|
|
|
65
65
|
---
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Feature: Shadow-Git Checkpoints (Phase 3)
|
|
2
2
|
|
|
3
|
-
**Gated by `prefs.global.shadowGit.enabled`** (default: `false`). The orchestrator snapshots the worktree via
|
|
3
|
+
**Gated by `prefs.global.shadowGit.enabled`** (default: `false`). The orchestrator snapshots the worktree via `$HOME/.claude/lib/shadow-git.sh` so sub-phase rollback is possible without polluting the project's real `.git` history.
|
|
4
4
|
|
|
5
5
|
```bash
|
|
6
6
|
# Phase 0 (one-time per task): initialize shadow repo + baseline snapshot.
|
|
@@ -0,0 +1,116 @@
|
|
|
1
|
+
# Skill conformance - reviewing against the criteria the work was built to
|
|
2
|
+
|
|
3
|
+
> **TLDR** - Phase 4 Step 1.78 resolves, deterministically and before any reviewer runs, WHAT the changed code was supposed to honour: declared rule registries scoped to the diff's languages, in-repo module guides, and the toolchains those registries delegate to. The result is `criteria-manifest.json`: a bounded set of rule IDs that becomes the denominator for "was this applied completely". Reviewers answer per rule ID. An ID that is neither checked nor explicitly waived fails the stage.
|
|
4
|
+
|
|
5
|
+
## Why this exists
|
|
6
|
+
|
|
7
|
+
Phase 4 used to select its review criteria from `detectedStack`, a string Phase 1 computed before a line of code existed, and Phase 3 never recorded which skills it actually applied. Two consequences:
|
|
8
|
+
|
|
9
|
+
- The dev side and the review side could disagree about the standard without either noticing. A component built by a plugin skill was reviewed against `clean-code`.
|
|
10
|
+
- "Did it do this correctly?" had no fixed denominator, so the only available answer was "it looks fine", and a reviewer that opened nothing produced the same output as a reviewer that checked everything.
|
|
11
|
+
|
|
12
|
+
The `--dev` family made this sharper: those modes have no Phase 1 at all, so the sole input to the old selection logic was absent.
|
|
13
|
+
|
|
14
|
+
## The four rules that make it work
|
|
15
|
+
|
|
16
|
+
1. **The denominator is written before the reviewers run.** `criteria-manifest.json` lands on disk in Step 1.78. Reviewers receive the selected rule IDs and must return a verdict per ID. A reviewer cannot narrow the set after seeing the diff, and a no-findings result still has to say what it checked.
|
|
17
|
+
2. **The resolver is primary; self-report only corroborates.** Phase 3 appends to `state.telemetry.skillCalls[]`, but coverage is never computed from it. An unrecorded consultation and no consultation are byte-identical in state, so a percentage derived from self-report reads green over an empty set. `ledger.source` therefore defaults to `derived`, and a declared entry the resolver could not bind to a changed file is FLAGGED rather than believed.
|
|
18
|
+
3. **Absent, empty and zero are three states.** `selfReport: absent` (no key), `empty` (key is `[]`), `declared` (has entries). Collapsing them makes the completeness claim unfalsifiable.
|
|
19
|
+
4. **The pipeline names no stack-specific skill.** Registries opt in; see below.
|
|
20
|
+
|
|
21
|
+
## Registry discovery is declared, never sniffed
|
|
22
|
+
|
|
23
|
+
A skill becomes a standards registry by saying so in its own frontmatter:
|
|
24
|
+
|
|
25
|
+
```yaml
|
|
26
|
+
standards-registry: references/rules.yml
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
`skill-conformance.mjs` reads frontmatter across the installed skills trees and loads what declared itself. Nothing in the pipeline names `ios-coding-standard`, `apple-archive-compliance` or any other skill, so a future UIKit, Objective-C, Kotlin or backend registry drops in with zero pipeline change - and a registry that is absent produces a declared coverage gap rather than a silent pass.
|
|
30
|
+
|
|
31
|
+
**The skills root differs per host**, so discovery is a bounded walk over candidate roots (`<install>/skills`, `<install>/multi-agent-refs/skills` for Codex, `<repo>/pipeline/skills`), installed layouts first. `install/copilot.mjs` copies `scripts/` byte-for-byte with no path rewrite, so a hardcoded `~/.claude/skills` is inert on two of the three hosts. That bug has already shipped here once: dynamic skill loading exited 1 on every real install while passing a smoke that ran from the repo. `skillsRootsSearched` is recorded in the manifest so an empty result is attributable to a root rather than to an absence of registries.
|
|
32
|
+
|
|
33
|
+
## Scope is required, and it is what makes this stack-generic
|
|
34
|
+
|
|
35
|
+
Every registry declares the languages and paths its rules may be applied to:
|
|
36
|
+
|
|
37
|
+
```yaml
|
|
38
|
+
scope:
|
|
39
|
+
languages: [swift]
|
|
40
|
+
paths: ["**/*.swift"]
|
|
41
|
+
excludePaths: ["**/Generated/**"]
|
|
42
|
+
notCovered:
|
|
43
|
+
objective-c: "no ObjC rules exist here; report as a coverage gap"
|
|
44
|
+
```
|
|
45
|
+
|
|
46
|
+
Per-rule `scope:` narrows this further and never widens it.
|
|
47
|
+
|
|
48
|
+
Without this the 99 Swift rules of the iOS registry would be applied to an Objective-C or UIKit file, which manufactures findings, buries the real ones, and teaches the reader to distrust the run - strictly worse than declaring no coverage. Every dropped rule is counted with its reason in `droppedReasons`, because silently narrowing the rule set is indistinguishable from the code passing it.
|
|
49
|
+
|
|
50
|
+
A registry that declares no scope loads as `reference-only`: readable context for the reviewers, contributing no rule IDs to the denominator.
|
|
51
|
+
|
|
52
|
+
## What is deterministic here, and what is deliberately not
|
|
53
|
+
|
|
54
|
+
**Deterministic (this gate):** which registries apply, which rule IDs are in scope, which changed files no criteria cover, whether the delegated toolchain is wired, and the exception-marker audit.
|
|
55
|
+
|
|
56
|
+
**Not deterministic, on purpose:** the rules themselves. This gate does NOT execute registry `mechanism` patterns. Across the 99-rule iOS registry the distribution is 47 `judgement` / 41 `lint` / 10 `scan` / 1 `format`, only 43 rules carry `mechanism` at all, and the field is prose with an embedded pattern (`swiftlint file_length, function_body_length`, `custom regex per module`) whose scope is written in English. Perhaps 12-15 have an extractable pattern. Running them unscoped floods a review with doc-comment matches - which the registry's own guidance says, requiring each hit be opened and confirmed. The properly engineered version already exists as `references/swiftlint.draft.yml`, 36 regex rules with path scoping.
|
|
57
|
+
|
|
58
|
+
So instead: **if a registry names a toolchain, check whether the repo wired it, run it when it is there, and report its absence as a finding.** That is stack-generic for free, because swiftlint, ktlint, detekt, eslint and ruff all speak rule IDs. An unwired toolchain ranks above most individual violations it would have caught: those rules are unverified, and a review reporting no violations for them is reporting that nothing was measured.
|
|
59
|
+
|
|
60
|
+
## The one bespoke scan: exception markers
|
|
61
|
+
|
|
62
|
+
Cheap, language-agnostic, and it answers "completely" head on - an exception is the author's own claim that a rule does not apply here, so an expired or unexplained one is something the reviewers would otherwise take on trust. The marker template is read from the registry's `exception_marker`, never hardcoded, so a registry using a different comment syntax still works.
|
|
63
|
+
|
|
64
|
+
| Condition | Severity |
|
|
65
|
+
|---|---|
|
|
66
|
+
| Exception expired (`expiry < today`) | blocking |
|
|
67
|
+
| Exception names an ID absent from every registry in scope | blocking |
|
|
68
|
+
| Exception with no rule ID | blocking |
|
|
69
|
+
| Exception with no expiry date | important |
|
|
70
|
+
| Exception with no reason | important |
|
|
71
|
+
|
|
72
|
+
## Dev-mode substitutes (Phases 1 and 2 never ran)
|
|
73
|
+
|
|
74
|
+
| Full-pipeline input | Dev-mode substitute |
|
|
75
|
+
|---|---|
|
|
76
|
+
| `detectedStack` (Phase 1) | language census of the diff by file extension. Describes what the diff CONTAINS, not what the repo is nominally built in, so one Objective-C bridging file in a Swift repo is classified correctly |
|
|
77
|
+
| Phase 1 analysis summary (triage scope) | the task description plus the inline task list Phase 3 generated for itself |
|
|
78
|
+
| Phase 2 plan | same inline task list |
|
|
79
|
+
| `state.evidence.figma[]` (Step 1.8) | recorded as `not-applicable (no Phase 1 evidence)` |
|
|
80
|
+
| Step 2.8 visual conformance | recorded as `not-applicable (no Phase 1 evidence)` |
|
|
81
|
+
|
|
82
|
+
Recording `not-applicable` rather than skipping is the point: "silently skipped" and "out of scope with a reason" must be distinguishable, or the completeness claim cannot be checked.
|
|
83
|
+
|
|
84
|
+
## Invocation
|
|
85
|
+
|
|
86
|
+
```bash
|
|
87
|
+
node $HOME/.claude/scripts/skill-conformance.mjs \
|
|
88
|
+
--diff "$WORKTREE/.review-diff.txt" \
|
|
89
|
+
--state "$WORKTREE/agent-state.json" \
|
|
90
|
+
--repo "$WORKTREE" \
|
|
91
|
+
--out "$WORKTREE/.pipeline/criteria-manifest.json"
|
|
92
|
+
```
|
|
93
|
+
|
|
94
|
+
Exit codes:
|
|
95
|
+
|
|
96
|
+
| Code | Meaning |
|
|
97
|
+
|---|---|
|
|
98
|
+
| `0` | manifest written; may still carry findings or declared coverage gaps |
|
|
99
|
+
| `1` | setup error: no diff, unreadable state, or an explicit `--skills-root` that does not exist |
|
|
100
|
+
| `2` | **fail closed**: no skills root resolved at all, a declared registry could not be parsed, or a declared registry path escaped its skill directory |
|
|
101
|
+
|
|
102
|
+
Any non-zero code halts before the reviewers. Continuing on `2` would drop a whole rule set from the denominator; and because `validate-reviewer.mjs` skips the checklist when zero rules were selected, an empty resolution would otherwise render as a clean review over criteria that were never loaded. An empty registry set inside a root that DOES exist is the different, legitimate case: declared coverage gap, exit 0.
|
|
103
|
+
|
|
104
|
+
`skillsRootsSearched` lists every candidate root that was probed with an `existed` flag, not only the ones found - otherwise the field is empty in exactly the case it exists to explain.
|
|
105
|
+
|
|
106
|
+
Output validates against `$HOME/.claude/schemas/criteria-manifest.schema.json`.
|
|
107
|
+
|
|
108
|
+
## Handoff to the reviewers
|
|
109
|
+
|
|
110
|
+
The manifest's selected rule IDs and registry file paths go into the **shared, cacheable prompt prefix** as one `${CRITERIA}` block, byte-identical for every reviewer (Step 1.9 requires an identical leading block; per-reviewer criteria subsets would invalidate the prefix for the whole panel and pay full input rate on the largest block in the phase). Reviewers read the registry YAML natively - the pipeline hands over a path, not a parse.
|
|
111
|
+
|
|
112
|
+
Deterministic findings merge at Step 3.0 alongside the test-integrity set, so triage adjudicates them rather than never seeing them. A finding citing a registry rule ID carries evidence a reviewer opinion does not, and the triage prompt says so: an ID-citing finding is not dismissible as a matter of taste, though it can still be out of scope for this task.
|
|
113
|
+
|
|
114
|
+
## Preference
|
|
115
|
+
|
|
116
|
+
`prefs.global.skillConformance.blockOnCoverageGap` (default **false**). A coverage gap on a stack with no registry would otherwise block every non-iOS run from day one. There is deliberately no opt-out for the stage itself or for the exception-expiry check, on the same grounds as Step 1.76: a run that can switch off its own anti-reward-hacking control cannot be trusted to report a pass.
|
|
@@ -0,0 +1,72 @@
|
|
|
1
|
+
# Stack skill routing - letting the toolkit plugin choose its own skills
|
|
2
|
+
|
|
3
|
+
> **TLDR** - When a stack toolkit plugin is enabled, Phase 3 asks that plugin's own `index` skill which of its skills apply to this task, loads them before writing code, and records each into `state.telemetry.skillCalls[]`. The routing table lives in the plugin; the pipeline copies none of it.
|
|
4
|
+
|
|
5
|
+
## Why this exists
|
|
6
|
+
|
|
7
|
+
Phase 3 dispatched to the toolkit plugin for exactly one case, `taskType === "component"` (see `component-dispatch.md`). Every other task - `bugfix`, `feature`, `refactor`, `chore` - had no skill dispatch at all: whichever skills the host happened to surface by description match were the ones that got used, and nothing recorded or required any of them.
|
|
8
|
+
|
|
9
|
+
That is the dev-side half of the gap `features/skill-conformance.md` closes on the review side. Review now asks "was this built to the rules it was supposed to follow"; without this step, the answer for a non-component task was "there were no declared rules, because nobody chose any".
|
|
10
|
+
|
|
11
|
+
The fix is not a routing table in the pipeline. Each `ai-<platform>-engineering-toolkit` already ships one: an `index` skill whose description says *"Load this first when unsure which skill applies"*, holding a 30-plus row intent-to-skill map maintained alongside the skills it points at. A second copy in this repo would drift the moment the plugin shipped a new skill, and the pipeline's copy would be the stale one.
|
|
12
|
+
|
|
13
|
+
So the pipeline's job is to **ask**, not to know.
|
|
14
|
+
|
|
15
|
+
## When it runs
|
|
16
|
+
|
|
17
|
+
Phase 3 pre-flight, before any code is written, for **every** `taskType`. Component tasks keep their dedicated dispatch in `component-dispatch.md`; this step runs in addition, because the reference skills (architecture, naming, file placement, tokens) apply to a component build too.
|
|
18
|
+
|
|
19
|
+
## Resolution
|
|
20
|
+
|
|
21
|
+
Platform comes from the same mapping component dispatch uses, so the two cannot disagree:
|
|
22
|
+
|
|
23
|
+
| `state.platform` / detected stack | Toolkit |
|
|
24
|
+
|---|---|
|
|
25
|
+
| ios, swift | `ai-ios-engineering-toolkit` |
|
|
26
|
+
| android, kotlin | `ai-android-engineering-toolkit` |
|
|
27
|
+
| anything else | no toolkit - step is a recorded no-op |
|
|
28
|
+
|
|
29
|
+
The toolkit is enabled per repo (`.claude/settings.local.json` / `~/.claude/settings.json` `enabledPlugins`). **Not enabled is not an error here**, unlike component dispatch: a backend or web repo legitimately has no toolkit, and halting would make the pipeline unusable outside mobile. Record the no-op and continue.
|
|
30
|
+
|
|
31
|
+
Two marketplaces may ship the same toolkit name (a public one and a corporate one). Resolve whichever is enabled and record its **name and version** in the ledger entry, because the routing table and the skill set differ between versions - a finding that cites a skill has to be traceable to the version that defined it.
|
|
32
|
+
|
|
33
|
+
## The call
|
|
34
|
+
|
|
35
|
+
```text
|
|
36
|
+
Skill(<toolkit>:index, args: "<task title + one-line intent>")
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
The index returns which `reference/` and `workflow/` skills apply. Load each via the Skill tool before writing code. A typical task pulls one workflow skill plus one or more reference skills.
|
|
40
|
+
|
|
41
|
+
Emit one progress line per loaded skill per `progress-contract.md`, so the user can see which standards the run bound itself to rather than inferring it afterwards.
|
|
42
|
+
|
|
43
|
+
## Recording - what makes this checkable
|
|
44
|
+
|
|
45
|
+
Append one `state.telemetry.skillCalls[]` entry per skill actually loaded:
|
|
46
|
+
|
|
47
|
+
```json
|
|
48
|
+
{"skill": "ai-ios-engineering-toolkit:reference/architecture", "phase": 3,
|
|
49
|
+
"targetFiles": ["Domains/Checkin/Sources/CheckinScene.swift"],
|
|
50
|
+
"routedBy": "ai-ios-engineering-toolkit:index@0.13.0", "timestamp": "<ISO-8601>"}
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
`routedBy` names the index and version that chose it. That is the difference between "the model happened to read a skill" and "the toolkit said this skill governs this task".
|
|
54
|
+
|
|
55
|
+
What Phase 4 actually does with it, precisely: Step 1.78 lists these entries in the manifest under `ledger.routedByToolkit`, so a reviewer and the Phase 7 report can see which skills the project's own toolkit selected. It does **not** give them extra weight in the coverage maths. The deterministic resolver stays primary because an unrecorded load and no load are indistinguishable in state, and no `routedBy` tag changes that - the tag says who chose the skill, not that the code honoured it.
|
|
56
|
+
|
|
57
|
+
## Failure modes, and why none of them halt
|
|
58
|
+
|
|
59
|
+
| Situation | Behaviour |
|
|
60
|
+
|---|---|
|
|
61
|
+
| No toolkit for this stack | recorded no-op, continue |
|
|
62
|
+
| Toolkit not enabled in this repo | recorded no-op, continue (component dispatch still halts for its own case) |
|
|
63
|
+
| `index` resolves but routes to a skill that does not exist in this version | record the miss with the version, load the rest, continue. A stale row in a plugin's table must not stop a run |
|
|
64
|
+
| `index` itself does not resolve | record and fall back to the host's own description matching, which is the pre-v14.1.0 behaviour - no worse than before |
|
|
65
|
+
|
|
66
|
+
Nothing here blocks Phase 3. What is downstream is visibility, not enforcement: routed skills appear in the manifest's `ledger.routedByToolkit`, and a task that recorded nothing shows up as `ledgerSource: derived` with its coverage gap stated. Enforcement over rule IDs is the registry's job (`features/skill-conformance.md`), not this step's.
|
|
67
|
+
|
|
68
|
+
## What this deliberately does NOT do
|
|
69
|
+
|
|
70
|
+
- It does not decide which skills apply. Copying the plugin's routing into this repo would put the authoritative table in the wrong place and guarantee drift.
|
|
71
|
+
- It does not fail a run for a missing skill. The pipeline's contract is to ask and record, not to require that a third-party plugin be complete.
|
|
72
|
+
- It does not replace `component-dispatch.md`. That path owns the component build itself, including the `figma-validate` pre-check and the halt-on-incomplete-state rule.
|
|
@@ -38,4 +38,4 @@ Adds one Sonnet call plus up to `maxFindings` single-test runs (and build-lock c
|
|
|
38
38
|
|
|
39
39
|
## Reference
|
|
40
40
|
|
|
41
|
-
Wiring: `$HOME/.claude/multi-agent-refs/phases/phase-4-review.md` Step 3.7. Schema:
|
|
41
|
+
Wiring: `$HOME/.claude/multi-agent-refs/phases/phase-4-review.md` Step 3.7. Schema: `$HOME/.claude/schemas/triage-output.schema.json` v3.2.0 (`$defs.verification`). Evidence gate: `$HOME/.claude/scripts/evidence-gate.mjs`. Prefs: `prefs.global.verifyByTest` in `$HOME/.claude/schemas/prefs.schema.json`.
|
|
@@ -0,0 +1,66 @@
|
|
|
1
|
+
# Worktree finalize - removing a task's worktree once its PR is open
|
|
2
|
+
|
|
3
|
+
> **TLDR** - Phase 6 step 9. Once the PR exists the worktree is dead weight, so it is removed: artefacts are salvaged into the log dir first, the branch is kept and deliberately NOT checked out, and every destructive path is gated. Gated by `prefs.global.settings.worktreeAutoRemoveOnPr` (default **true**). Script: `worktree-finalize.sh`.
|
|
4
|
+
|
|
5
|
+
## Why it exists
|
|
6
|
+
|
|
7
|
+
A finished task's worktree is a full second checkout that nobody needs after the PR is open, and removing it is the step people forget. `.worktrees/` then accumulates copies of the repo until `/multi-agent:kill` or `:garbage-collect` is run by hand.
|
|
8
|
+
|
|
9
|
+
Removing it at PR-open is only safe because of the salvage, so the two are one step and not two.
|
|
10
|
+
|
|
11
|
+
## What it will not do
|
|
12
|
+
|
|
13
|
+
**No `git checkout` of the task branch.** `git worktree remove` leaves the branch as an ordinary local branch: the ref, its commits, and the ability to `git checkout <branch>` later all survive untouched. Checking it out here would move the user's HEAD out from under them and can collide with their own uncommitted work on another branch. Phase 5 removes-then-checks-out on purpose, because it is handing the branch over for manual testing; this step is not.
|
|
14
|
+
|
|
15
|
+
**No `git branch -D`.** The branch is the deliverable.
|
|
16
|
+
|
|
17
|
+
**No `--force`, ever.** `git worktree remove` refusing is a safety feature. The clean-tree check runs before it, so a refusal at that point means something unexpected (a lock, a submodule, permissions) and forcing past unexpected dirt is how work gets lost.
|
|
18
|
+
|
|
19
|
+
## Preconditions - each one skips with a reason, none is an error
|
|
20
|
+
|
|
21
|
+
| Condition | Why it blocks |
|
|
22
|
+
|---|---|
|
|
23
|
+
| `worktreePath == projectRoot` (`--local` mode) | there is no worktree; removing it would delete the user's checkout |
|
|
24
|
+
| cwd is inside the worktree | a shell left on a deleted inode is worse than a leftover directory, and Phase 6 legitimately `cd`s into the worktree earlier |
|
|
25
|
+
| not a registered worktree of the project root | a mistyped path must not delete an unrelated directory |
|
|
26
|
+
| real uncommitted changes | never discarded; see the artefact carve-out below |
|
|
27
|
+
| HEAD not on the remote | removing a worktree whose commits exist nowhere else is data loss, not cleanup |
|
|
28
|
+
|
|
29
|
+
Exit codes: `0` removed, `3` skipped with a reason (report and continue to Phase 7), `1` usage error.
|
|
30
|
+
|
|
31
|
+
### The artefact carve-out, and why `--untracked-files=no` is wrong
|
|
32
|
+
|
|
33
|
+
The pipeline's own artefacts live inside the worktree and are untracked, so a raw `git status --porcelain` is never empty at PR-open. Left unhandled the removal would never fire and the feature would look implemented while doing nothing.
|
|
34
|
+
|
|
35
|
+
So exactly these paths are forgiven, and nothing else:
|
|
36
|
+
|
|
37
|
+
```
|
|
38
|
+
agent-state.json phase-tracker.json triage-output.json
|
|
39
|
+
.review-diff.txt .build.log .test.log .pipeline/
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
Suppressing all untracked files instead (`--untracked-files=no`) would have been shorter and wrong: a source file the developer created but never `git add`ed is invisible to it, and that file would be destroyed silently.
|
|
43
|
+
|
|
44
|
+
## Salvage
|
|
45
|
+
|
|
46
|
+
Copied into `$HOME/.claude/logs/multi-agent/<project>/<task-id>/artifacts/` before removal, each only if present - a task that never reached Phase 4 has no triage output and that is not an error.
|
|
47
|
+
|
|
48
|
+
This is why the removal is safe:
|
|
49
|
+
|
|
50
|
+
| Consumer | Reads | Without salvage |
|
|
51
|
+
|---|---|---|
|
|
52
|
+
| Phase 7 triage-memory ingest | `triage-output.json` | `[ -f ]`-guarded, so it degrades **silently**: the triage corpus and learnings ledger stop being fed and no error appears |
|
|
53
|
+
| Phase 7 learnings-ledger distill | same file | same silent degradation |
|
|
54
|
+
| `render-work-summary.sh` | `agent-state.json`, `phase-tracker.json` | exits 2, so the Work Summary vanishes from the PR body and the Jira comment |
|
|
55
|
+
| `:resume` | `agent-state.json` | cannot continue a Phase 7 pause |
|
|
56
|
+
| `:status`, `:log` | `agent-state.json` | the task becomes invisible |
|
|
57
|
+
|
|
58
|
+
`state.worktreeRemovedAt` and `state.artifactsPath` record the outcome. The timestamp is what tells a reader that a worktree-less task was finished-and-tidied rather than killed - without it, a missing worktree is indistinguishable from a broken run.
|
|
59
|
+
|
|
60
|
+
## Multi-repo
|
|
61
|
+
|
|
62
|
+
Run serially per repo, and only **after** `update_sibling_links`: that function issues an update per PR and the loop `cd`s per repo, so removing repo 1's worktree mid-loop breaks repos 2..N.
|
|
63
|
+
|
|
64
|
+
## Interaction with the existing removal sites
|
|
65
|
+
|
|
66
|
+
`gc-worktrees.sh` and `/multi-agent:garbage-collect` never touch a registered, healthy worktree - they sweep orphans. Both name this step as the owner of finishing-a-task removal, which is now true rather than aspirational.
|