@mmerterden/multi-agent-pipeline 19.1.4 → 20.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (137) hide show
  1. package/CHANGELOG.md +123 -0
  2. package/README.md +19 -36
  3. package/README.tr.md +18 -35
  4. package/SECURITY.md +3 -3
  5. package/docs/adr/0002-instruction-driven-flag.md +6 -5
  6. package/docs/adr/0005-lazy-phase-docs.md +2 -2
  7. package/docs/adr/0008-installer-modularization-and-secret-leak-defense.md +1 -0
  8. package/docs/adr/0009-claude-stack-skills-plugin-only.md +1 -1
  9. package/docs/adr/0010-own-code-graph.md +5 -4
  10. package/docs/adr/0011-dormant-ci.md +10 -1
  11. package/docs/adr/0012-macos-only.md +2 -2
  12. package/docs/adr/0013-lsp-code-intelligence.md +2 -2
  13. package/docs/adr/0014-six-phase-consolidation.md +9 -9
  14. package/docs/adr/0015-one-pipeline-no-depth-answer.md +83 -0
  15. package/docs/adr/0016-the-run-shape-is-asked-not-typed.md +69 -0
  16. package/docs/adr/README.md +18 -16
  17. package/docs/architecture.md +2 -2
  18. package/docs/ecosystem.md +5 -5
  19. package/docs/facts.json +7 -9
  20. package/docs/features.md +4 -5
  21. package/docs/token-budget-history.md +1 -1
  22. package/install/_codex-agents.mjs +1 -1
  23. package/install/_common.mjs +9 -1
  24. package/install/templates/copilot-instructions.md +7 -16
  25. package/manifest.json +133 -129
  26. package/package.json +1 -1
  27. package/pipeline/agents/code-reviewer.md +2 -2
  28. package/pipeline/agents/dev-critic.md +5 -5
  29. package/pipeline/agents/security-auditor.md +80 -72
  30. package/pipeline/commands/figma-to-swiftui.md +1 -1
  31. package/pipeline/commands/multi-agent/SKILL.md +7 -9
  32. package/pipeline/commands/multi-agent/analysis/SKILL.md +2 -0
  33. package/pipeline/commands/multi-agent/analysis-jira/SKILL.md +2 -0
  34. package/pipeline/commands/multi-agent/analysis-resolve/SKILL.md +2 -0
  35. package/pipeline/commands/multi-agent/autopilot/SKILL.md +2 -0
  36. package/pipeline/commands/multi-agent/autopilot-on/SKILL.md +2 -0
  37. package/pipeline/commands/multi-agent/autopilot-status/SKILL.md +1 -1
  38. package/pipeline/commands/multi-agent/build-optimize/SKILL.md +2 -0
  39. package/pipeline/commands/multi-agent/channels/SKILL.md +2 -2
  40. package/pipeline/commands/multi-agent/create-jira/SKILL.md +2 -0
  41. package/pipeline/commands/multi-agent/design-check/SKILL.md +1 -1
  42. package/pipeline/commands/multi-agent/diff-explain/SKILL.md +1 -1
  43. package/pipeline/commands/multi-agent/forget/SKILL.md +2 -0
  44. package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +4 -2
  45. package/pipeline/commands/multi-agent/help/SKILL.md +23 -27
  46. package/pipeline/commands/multi-agent/ios-coding-standard/SKILL.md +5 -4
  47. package/pipeline/commands/multi-agent/issue/SKILL.md +2 -0
  48. package/pipeline/commands/multi-agent/jira/SKILL.md +2 -0
  49. package/pipeline/commands/multi-agent/language/SKILL.md +2 -0
  50. package/pipeline/commands/multi-agent/prune-logs/SKILL.md +2 -0
  51. package/pipeline/commands/multi-agent/purge/SKILL.md +2 -0
  52. package/pipeline/commands/multi-agent/resume/SKILL.md +177 -48
  53. package/pipeline/commands/multi-agent/save/SKILL.md +2 -0
  54. package/pipeline/commands/multi-agent/scan/SKILL.md +2 -2
  55. package/pipeline/commands/multi-agent/security-review/SKILL.md +52 -0
  56. package/pipeline/commands/multi-agent/stack/SKILL.md +2 -0
  57. package/pipeline/commands/multi-agent/sync/SKILL.md +7 -8
  58. package/pipeline/commands/multi-agent/test-screenshots/SKILL.md +2 -0
  59. package/pipeline/commands/multi-agent/uninstall/SKILL.md +2 -0
  60. package/pipeline/lib/repo-hygiene.sh +1 -1
  61. package/pipeline/multi-agent-refs/analysis/render.md +1 -1
  62. package/pipeline/multi-agent-refs/analysis/resolve.md +1 -1
  63. package/pipeline/multi-agent-refs/analysis/synthesis.md +1 -1
  64. package/pipeline/multi-agent-refs/analysis-template.md +1 -1
  65. package/pipeline/multi-agent-refs/component-dispatch.md +5 -13
  66. package/pipeline/multi-agent-refs/cross-cli-contract.md +14 -15
  67. package/pipeline/multi-agent-refs/features/external-context-injection.md +2 -0
  68. package/pipeline/multi-agent-refs/features/review-delta.md +1 -1
  69. package/pipeline/multi-agent-refs/features/review-multi-repo.md +3 -3
  70. package/pipeline/multi-agent-refs/features/security-audit.md +55 -0
  71. package/pipeline/multi-agent-refs/features/skill-conformance.md +1 -1
  72. package/pipeline/multi-agent-refs/features/visual-evidence.md +2 -1
  73. package/pipeline/multi-agent-refs/features/worktree-finalize.md +1 -1
  74. package/pipeline/multi-agent-refs/generate-issue.md +2 -0
  75. package/pipeline/multi-agent-refs/issue-jira-triad.md +2 -0
  76. package/pipeline/multi-agent-refs/keychain.md +2 -0
  77. package/pipeline/multi-agent-refs/knowledge.md +0 -7
  78. package/pipeline/multi-agent-refs/outside-the-pipeline.md +1 -1
  79. package/pipeline/multi-agent-refs/payload-contracts.md +1 -1
  80. package/pipeline/multi-agent-refs/phases/modes.md +33 -109
  81. package/pipeline/multi-agent-refs/phases/operations.md +2 -0
  82. package/pipeline/multi-agent-refs/phases/phase-0-init.md +23 -42
  83. package/pipeline/multi-agent-refs/phases/phase-1-plan.md +7 -18
  84. package/pipeline/multi-agent-refs/phases/phase-2-dev.md +13 -44
  85. package/pipeline/multi-agent-refs/phases/phase-3-review.md +28 -37
  86. package/pipeline/multi-agent-refs/phases/phase-4-commit.md +6 -6
  87. package/pipeline/multi-agent-refs/phases/phase-5-report.md +3 -3
  88. package/pipeline/multi-agent-refs/phases.md +9 -11
  89. package/pipeline/multi-agent-refs/progress-contract.md +1 -1
  90. package/pipeline/multi-agent-refs/readiness-review.md +2 -0
  91. package/pipeline/multi-agent-refs/rules.md +1 -1
  92. package/pipeline/multi-agent-refs/threat-model.md +39 -0
  93. package/pipeline/multi-agent-refs/tracker-contract.md +9 -40
  94. package/pipeline/multi-agent-refs/wiki-capture.md +3 -2
  95. package/pipeline/preferences-template.json +2 -2
  96. package/pipeline/rules/figma-pipeline.md +1 -1
  97. package/pipeline/schemas/agent-state.schema.json +28 -10
  98. package/pipeline/schemas/migrations/prefs-2.7.0-to-2.8.0.mjs +33 -0
  99. package/pipeline/schemas/phases.json +4 -26
  100. package/pipeline/schemas/prefs.schema.json +5 -9
  101. package/pipeline/schemas/reviewer-output.schema.json +99 -2
  102. package/pipeline/schemas/security-finding.schema.json +144 -0
  103. package/pipeline/scripts/_stack-routing.mjs +1 -0
  104. package/pipeline/scripts/cost-table.json +1 -1
  105. package/pipeline/scripts/gc-abandoned.sh +16 -9
  106. package/pipeline/scripts/gc-refs.sh +1 -1
  107. package/pipeline/scripts/gen-mode-dispatch.mjs +11 -41
  108. package/pipeline/scripts/migrate-prefs.mjs +18 -17
  109. package/pipeline/scripts/phase-tracker.sh +2 -2
  110. package/pipeline/scripts/phase0-exit-gate.mjs +1 -1
  111. package/pipeline/scripts/plan-coverage-gate.mjs +3 -3
  112. package/pipeline/scripts/render-work-summary.sh +7 -4
  113. package/pipeline/scripts/run-aggregator.mjs +1 -1
  114. package/pipeline/scripts/usage-report.mjs +0 -2
  115. package/pipeline/scripts/worktree-finalize.sh +2 -2
  116. package/pipeline/skills/.skill-manifest.json +17 -21
  117. package/pipeline/skills/.skills-index.json +6 -39
  118. package/pipeline/skills/shared/README.md +5 -8
  119. package/pipeline/skills/shared/core/multi-agent/SKILL.md +11 -15
  120. package/pipeline/skills/shared/core/multi-agent-autopilot-status/SKILL.md +1 -1
  121. package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +13 -16
  122. package/pipeline/skills/shared/core/multi-agent-ios-coding-standard/SKILL.md +2 -3
  123. package/pipeline/skills/shared/core/multi-agent-resume/SKILL.md +51 -15
  124. package/pipeline/skills/shared/core/multi-agent-scan/SKILL.md +2 -2
  125. package/pipeline/skills/shared/core/multi-agent-security-review/SKILL.md +29 -0
  126. package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +7 -7
  127. package/pipeline/skills/shared/external/security-review/SKILL.md +64 -0
  128. package/pipeline/skills/shared/external/security-review/references/owasp-mobile-top10-2024.md +53 -0
  129. package/pipeline/skills/shared/external/security-review/references/owasp-web-api-top10-2021.md +56 -0
  130. package/pipeline/skills/skills-index.md +3 -6
  131. package/pipeline/commands/multi-agent/local/SKILL.md +0 -132
  132. package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +0 -142
  133. package/pipeline/commands/multi-agent/resume-local/SKILL.md +0 -114
  134. package/pipeline/commands/security-review.md +0 -6
  135. package/pipeline/skills/shared/core/multi-agent-local/SKILL.md +0 -41
  136. package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +0 -55
  137. package/pipeline/skills/shared/core/multi-agent-resume-local/SKILL.md +0 -51
@@ -1,31 +1,67 @@
1
1
  ---
2
2
  name: multi-agent-resume
3
3
  language: en
4
- description: "Resume a stopped or failed task from the phase where it left off. Use when a task stopped or failed and should carry on from where it left off."
4
+ description: "Pick up unfinished work: a pipeline run that stopped mid-phase, or work already written on the current branch that never went through the pipeline. Use when a task stopped or failed, or when hand-written local work needs review, build, PR and reporting."
5
5
  user-invocable: true
6
- argument-hint: "[#id] - optional: task ID (e.g. #2). If omitted, the most recent paused task is used."
6
+ argument-hint: "[#id | PROJ-12345] [--base <branch>] [autopilot] - no argument: pick from the list"
7
7
  ---
8
8
 
9
- # multi-agent resume - Resume Paused Task
9
+ # multi-agent resume - Continue Unfinished Work
10
10
 
11
11
  **Input**: $ARGUMENTS
12
12
 
13
- Resume a paused or failed task from the last successful phase.
13
+ Unfinished work arrives from two directions, and which one applies is a property of the repository, not of how the command was typed:
14
+
15
+ - **A tracked run stopped.** It has an `agent-state.json`, a phase it was inside, and possibly a worktree. It resumes from where it stopped.
16
+ - **Work exists on the current branch with no run behind it.** Nothing to resume, so the **pipeline tail** runs over the diff. Plan and Dev never run; the diff is the Dev output.
17
+
18
+ Step 1 finds both and asks, so the user does not have to know which case they are in before asking the question.
19
+
20
+ ## Input
21
+
22
+ ```bash
23
+ multi-agent resume # list everything resumable, pick one
24
+ multi-agent resume #2 # a tracked run by task number
25
+ multi-agent resume PROJ-12345 # a tracked run by Jira id, or bind the tail to that id
26
+ multi-agent resume --base develop # tail path: override the base branch for the diff
27
+ multi-agent resume autopilot # no gate prompts: auto-fix, auto-PR, auto-comment
28
+ ```
14
29
 
15
30
  ## Steps
16
31
 
17
- 1. **Find task** - Parse `#N` from the argument, or find the most recent worktree with `status != "done"`
32
+ 1. **Build the resumable list, then ask.** Source A: every `agent-state.json` under `$HOME/.claude/logs/multi-agent/{project}/` with `status != "done"`, each row showing task id, the phase it stopped inside, its workspace (worktree path, or `local`), age, and `haltReason` when set. Source B: the current branch, when `git diff <base>...HEAD` plus working-tree changes is non-empty **and** no Source A run owns that branch. An argument naming a run selects it directly; an argument naming nothing resumable is an error, never a silent fall-through to the tail. Nothing in either source: `ERR: nothing to resume` - do not ask.
33
+
34
+ Ask per `$HOME/.claude/multi-agent-refs/picker-contract.md` (Copilot CLI: `$HOME/.claude/lib/ask-choice.sh`). Breadcrumb `Step 1/1: which unfinished work to pick up` above the picker, in `outputLanguage`; labels carry the task id and branch name verbatim, and the run branches on which row was selected rather than on the label text. A single row is still a question and needs a genuine second option: **Show every run** (drops the `status != "done"` filter) when the only row is a stopped run, **Pick a stopped run instead** when the only row is the branch diff.
35
+
36
+ 2. **Tracked run - read and validate state.** `validate-state.mjs` first; on non-zero exit do NOT guess a phase. Read `currentPhase`, `status`, `haltReason`, `circuitBreaker`, `autopilot`. Heal a stale worktree lock before continuing - unless `worktreeRemovedAt` is set, in which case the artefacts live under `artifactsPath` and the worktree must not be recreated.
37
+
38
+ 3. **Tracked run - load context** from durable artifacts, never from conversation memory: the latest `## Handoff` block in `agent-log.md` first, per-phase findings as the fallback, `git log --oneline -10` to ground what was actually committed.
39
+
40
+ 4. **Tracked run - rebuild the phase tiles** from `tracker-state.json`, never re-initialized. Copilot CLI: a single `phase-tracker.sh render`.
41
+
42
+ 5. **Tracked run - continue.** Read `state.waitingFor` FIRST: when it names a step (`maturity`, `user-channels-choice`, `local-test`), the run re-enters THAT step. `currentPhase + 1` is the fallback, not the rule. Clear `waitingFor` in the same write that records the answer.
43
+
44
+ 6. **Untracked branch work - resolve context.** Base branch: `--base` → `figma-config.project.baseBranch` → `develop` → upstream/merge-base. Jira id from the argument or the branch name. State under `$HOME/.claude/logs/multi-agent/{project}/{taskId}/`. No worktree is created.
45
+
46
+ 7. **Untracked branch work - run the tail.**
47
+
48
+ ```
49
+ Phase 0: Init → project/branch detect, base + diff, Jira id, state (NO worktree)
50
+ Phase 3: Review → the Verify gate first (stack-aware build + existing tests; SUCCESS required), then
51
+ parallel review (GPT + Opus + Sonnet) + triage
52
+ Phase 4: Commit → commit remaining changes + push + open PR if none exists
53
+ Phase 5: Report → technical analysis + Jira comment with test scenarios (channels: Jira / PR / Confluence / Wiki)
54
+ ```
55
+
56
+ The build that Dev's exit gate would have run happens inside Review instead, because there is no Dev run to inherit a log from.
18
57
 
19
- 2. **Read state** - Parse `agent-state.json`:
20
- - `currentPhase` - last completed phase
21
- - `status` - `paused` | `failed` | `in_progress`
22
- - `autopilot` - preserve the mode
58
+ ## Notes
23
59
 
24
- 3. **Load context** - Read the findings of previous phases from `agent-log.md`:
25
- - Phase 1 analysis → use in Phase 2+
26
- - Phase 1 plan → use in Phase 2+
27
- - Phase 3 code → already present in the worktree
60
+ - Build+Test is the automated success gate (the interactive device user-test is `multi-agent:manual-test`). If the repo has no tests, it reports "no tests present" - never fabricates results.
61
+ - `autopilot` (or `prefs.global.resume.autoFix == true`) auto-fixes triage-accepted blocking/important findings and re-reviews the fix before advancing.
62
+ - Commit/PR follows house rules: conventional message, `Ref: #N` (never Closes/Fixes), NO AI/bot attribution. PR opened only if one does not already exist.
63
+ - Full phase contract lives in the Claude Code command `commands/multi-agent/resume/SKILL.md`; this skill is the Copilot-CLI counterpart.
28
64
 
29
- 4. **Continue the pipeline** - Start from where it left off (same pipeline as the main multi-agent command)
65
+ ## Required: outward-facing payload contracts
30
66
 
31
- 5. **Log**: `🔄 Resumed PROJ-{id} from Phase {N}`
67
+ Before writing anything outward-facing - PR body, Jira comment, Confluence page, closing report - load `$HOME/.claude/multi-agent-refs/payload-contracts.md`. It names the canonical section set for each payload, the markup dialect per surface (PR body is Markdown, Jira is wiki markup - mixing them is a defect), and the token/duration numbers the closing report must carry. Improvising a payload shape from memory is the most common failure of a run that starts in the middle.
@@ -3,7 +3,7 @@ name: multi-agent-scan
3
3
  language: en
4
4
  description: "Skill security scan: walks local skill directories against a tiered pattern catalog. Use when local skill directories need checking for unsafe or unexpected content."
5
5
  user-invocable: true
6
- argument-hint: "[--strict] [--target <path>] - optional: --strict enables CI failure mode, --target picks a custom directory"
6
+ argument-hint: "[--strict] [--root PATH] - optional: --strict enables strict exit codes, --root picks a custom directory"
7
7
  ---
8
8
 
9
9
  # multi-agent-scan
@@ -55,7 +55,7 @@ multi-agent-scan --root ~/.copilot/skills
55
55
  ## Integration
56
56
 
57
57
  - **install.js** pre-deploy hook (automatic, warn-only high-threshold)
58
- - **CI** `.github/workflows/smoke.yml` step (strict)
58
+ - **Smoke suite** `smoke-skill-scan.sh` runs it in strict mode under `npm test`
59
59
  - **Standalone** this skill
60
60
 
61
61
  Cross-CLI parity: `/multi-agent:scan` on Claude Code, `multi-agent-scan` on Copilot CLI. Same source script, same behavior.
@@ -0,0 +1,29 @@
1
+ ---
2
+ name: multi-agent-security-review
3
+ language: en
4
+ description: "Run a standalone defensive, static security review of a diff, branch or repo: run-scoped threat model, reviewer-shaped findings joined to OWASP + CWE with CVSS scoring, evidence and before/after fixes, plus an offline dependency inventory. No live target, no payloads. Use when reviewing code for security outside a full pipeline run, auditing dependencies, or preparing a branch for a security sign-off."
5
+ user-invocable: true
6
+ argument-hint: "[#N | repo#N | PR-URL | branch | path] - a PR, a local branch, or a path to scope the review. If omitted: the current branch diff against its base."
7
+ ---
8
+
9
+ # multi-agent security-review - standalone defensive static review
10
+
11
+ **Input**: `$ARGUMENTS`
12
+
13
+ The same security audit Phase 3 runs at Step 2.7, invoked on its own. Defensive and static: reads code, config and dependency manifests; never runs the target, fires a payload, or reaches a live host.
14
+
15
+ ## Scope
16
+
17
+ Resolve from `$ARGUMENTS`, same shapes as `/multi-agent:review`: `#N` / `repo#N` / PR URL → that PR's diff; a branch → its diff against base; a path → files under it; omitted → the current branch diff. Cap the diff as Phase 3 Step 1.9 does.
18
+
19
+ ## Steps
20
+
21
+ 1. **Threat model** - produce `.pipeline/threat-model.md` (four sections) if absent, else read it; mirror to `state.threatModel`. Contract: `threat-model.md`.
22
+ 2. **Method** - load `ai-common-toolkit:security-review` for the OWASP walk (Web/API Top 10 2021, Mobile Top 10 2024) + CWE join. This command orchestrates, it does not re-derive the method.
23
+ 3. **Findings** - dispatch the `security-auditor`; it returns a `reviewer-output.schema.json` object whose `findings[]` carry the `security` envelope (`security-finding.schema.json`). Score every vector with the toolkit `security_cvss_score`; validate with `validate-reviewer.mjs` (one rework then halt).
24
+ 4. **Dependencies** - run `security_dep_inventory` on the lockfiles; hand the inventory to `ai-analyst-toolkit:evidence-registry` for known CVEs when it is registered. A vulnerable dependency is an `A06:2021` finding with the advisory CWE + CVE. This command contacts nothing itself.
25
+ 5. **Report** - write `.pipeline/security-findings.json` and a human summary (count by severity; each blocking/important finding with OWASP id, CWE, CVSS band, evidence, before/after fix). An empty list with `approved: true` is a clean review.
26
+
27
+ ## Boundaries
28
+
29
+ Inside a run the same audit is Phase 3 Step 2.7 (`features/security-audit.md`), triggered by the `security_path` signal; this is the standalone entry point. Not the `store-ready` device pass, not a secret scanner (`pre-commit-check.sh` covers secrets). No writes to Jira / GitHub / Confluence. Autopilot reviews the current branch, writes the report, opens no PR.
@@ -32,8 +32,8 @@ Run all steps automatically:
32
32
  ```
33
33
  Step 0: DOCTOR node $HOME/.claude/scripts/doctor.mjs - exit 2 or 4 STOPS the sync
34
34
  Step 1: DETECT Compare timestamps, find stale targets
35
- Step 2: COPILOT Claude Code -> Copilot CLI (instructions + 60 sub-command skills)
36
- Step 2b: CODEX Claude Code -> Codex CLI (1 router skill + 60 specs as refs + 8 agent TOML)
35
+ Step 2: COPILOT Claude Code -> Copilot CLI (instructions + 58 sub-command skills)
36
+ Step 2b: CODEX Claude Code -> Codex CLI (1 router skill + 57 specs as refs + 8 agent TOML)
37
37
  Step 3: REPO Claude Code -> pipeline repo (genericized, personal data scrub)
38
38
  Step 3d: DEV-TOOLKIT Companion MCP server -> detect movement, ship gates, commit + publish
39
39
  Step 4: WEBSITE Version + phase/model counts -> {website-host} (i18n + projects.ts)
@@ -63,7 +63,7 @@ If nothing is stale -> report "All targets up to date" and stop.
63
63
  - `copilot-instructions.md`: general development instructions + pipeline summary section
64
64
  - `multi-agent-pipeline/pipeline/`: generic open-source version (NO personal data)
65
65
  3. **Sync shared sections** (Claude <-> Copilot):
66
- - Pipeline entries table (base / :local / :autopilot / :local-autopilot) + the Full-or-Short depth question
66
+ - Pipeline entries table (base / :autopilot)
67
67
  - Project detection (URL-based + cwd-based)
68
68
  - Figma pipeline flow
69
69
  - Git conventions (author, branch, commit format)
@@ -228,18 +228,18 @@ When invoked with the `release` argument:
228
228
  |-------------|-------------|
229
229
  | `~/.claude/commands/multi-agent/{cmd}/SKILL.md` | `~/.copilot/skills/multi-agent-{cmd}/SKILL.md` |
230
230
 
231
- **60 commands are synced** (canonical inventory - must match `cross-cli-contract.md` section 1; drift = contract violation):
231
+ **58 commands are synced** (canonical inventory - must match `cross-cli-contract.md` section 1; drift = contract violation):
232
232
 
233
233
  ```
234
234
  analysis, analysis-jira, analysis-resolve, autopilot, autopilot-off,
235
235
  autopilot-on, autopilot-status, build-optimize, channels, complaint-analysis,
236
236
  create-jira, design-check, diff-explain, doctor, feedback, forget,
237
237
  garbage-collect, graph, help, ios-coding-standard, issue, jira, kill,
238
- language, local, local-autopilot, log, manual-test, model, prune-logs,
239
- prune-prompts, purge, refactor, resume, resume-local, review,
238
+ language, log, manual-test, model, prune-logs,
239
+ prune-prompts, purge, refactor, resume, review,
240
240
  review-analysis, review-issue, review-jira, route-off, route-on,
241
241
  route-status, routines, save, scan, search,
242
- setup, stack, status, steer, store-ready, sync, test, test-accessibility,
242
+ security-review, setup, stack, status, steer, store-ready, sync, test, test-accessibility,
243
243
  test-dark-mode, test-dynamic-type, test-screenshots, testflight-validation,
244
244
  uninstall, update
245
245
  ```
@@ -0,0 +1,64 @@
1
+ ---
2
+ name: security-review
3
+ description: "Run a defensive, static security review of a diff or a repo: build a run-scoped threat model, find vulnerabilities joined to OWASP + CWE with CVSS scoring and evidence, and emit reviewer-shaped findings with before/after fixes. Use when reviewing code for security, auditing a dependency set, preparing a release for a security sign-off, or when a diff touches auth, crypto, input handling, networking or secrets."
4
+ risk: low
5
+ source: multi-agent-pipeline
6
+ date_added: "2026-09-21"
7
+ ---
8
+
9
+ # Security Review (defensive, static)
10
+
11
+ ## Overview
12
+
13
+ A static, read-only security review. It reads code, configuration and dependency manifests and reports vulnerabilities the way the pipeline can act on them: each finding joined to an OWASP category and a CWE, scored with CVSS 3.1, backed by evidence and its counterevidence, and paired with a before/after fix. It never runs the target, fires a payload, or reaches a live host - the posture is defensive and offline. The output is reviewer-shaped, so a blocking finding merges into triage and blocks the commit like any other blocker.
14
+
15
+ This is stack-neutral. The same method covers a Swift or Kotlin mobile diff, a TypeScript / Python / Go / Java backend, and a web front-end. What changes per stack is the OWASP catalog (Mobile Top 10 for apps, Web/API Top 10 for services and sites) and the sink vocabulary, not the method.
16
+
17
+ ## When to Use This Skill
18
+
19
+ - Reviewing a diff whose paths touch authentication, authorization, crypto, input handling, networking, deserialization, file access or secrets.
20
+ - Auditing a dependency set for known-vulnerable packages before a release.
21
+ - Preparing a branch for a security sign-off, or answering "is this change safe to ship".
22
+ - Any run where `/multi-agent:security-review` is invoked, or Phase 3 fires the `security_path` signal.
23
+
24
+ ## How It Works
25
+
26
+ ### Step 1: Build the run-scoped threat model
27
+
28
+ Before any finding, establish the four-section threat model for THIS change and write it to `.pipeline/threat-model.md` (contract: `threat-model.md` in the pipeline refs). Attacker, trust boundaries, attack surface, severity calibration. Every finding's severity is calibrated against it, and a `blocking` finding must trace to a named assumption in it. A diff with no plausible attacker gets a short model and findings that cap at `suggestion`.
29
+
30
+ ### Step 2: Map the surface to a standard
31
+
32
+ Join every candidate finding to a checkable standard, never a bare opinion:
33
+
34
+ - **Web / API** - OWASP Top 10 2021 (`references/owasp-web-api-top10-2021.md`), with the CWE most associated with each category.
35
+ - **Mobile** - OWASP Mobile Top 10 2024 (`references/owasp-mobile-top10-2024.md`).
36
+
37
+ Walk the checklist against the changed surface only. A finding outside the touched surface is out of scope unless the diff made it reachable, and then you say how.
38
+
39
+ ### Step 3: Score with CVSS, and keep the number honest
40
+
41
+ For each finding, write a CVSS 3.1 base vector and compute the score with the `security_cvss_score` toolkit tool - never by hand, so the score cannot drift from the vector. The band sets the reviewer severity: critical or high -> `blocking`, medium -> `important`, low or none -> `suggestion`.
42
+
43
+ Every finding carries:
44
+
45
+ - **evidence** - what in the code proves it, cited by `file:line`. No evidence, no finding.
46
+ - **counterevidence** - what would disprove it, or the condition that makes it a false positive. Blank asserts there is none.
47
+ - **confidence** - `high|medium|low`. Low does not mean silent; it means report with the counterevidence.
48
+ - **severityChangeConditions** - the assumption the score rests on, so a reader disputes the assumption, not the number.
49
+
50
+ ### Step 4: Inventory dependencies (offline)
51
+
52
+ Run `security_dep_inventory` on the lockfiles in scope to get a normalized `{ecosystem, name, version}` list. Hand that list to a known-CVE lookup (the analyst toolkit's evidence-registry) as a separate step - this skill contacts nothing itself. A vulnerable dependency is `A06:2021` with the advisory's CWE and CVE, on the manifest file at line 0.
53
+
54
+ ### Step 5: Emit reviewer-shaped findings
55
+
56
+ Emit one object conforming to `reviewer-output.schema.json`, each finding also conforming to `security-finding.schema.json` (the `security` block). Include the before/after fix as `security.remediationDiff`, and `security.fixVerification` naming the empirical check a human or a later dynamic pass would run - static review cannot fire the exploit itself. `approved` is `false` if any finding is `blocking`.
57
+
58
+ An empty findings list with `approved: true` is the right answer for a clean diff. Do not invent findings to look thorough, and do not raise a theoretical risk with no path from the threat model's attacker above `suggestion`.
59
+
60
+ ## What This Skill Does Not Do
61
+
62
+ - No live target, no payloads, no exploitation, no proxy, no browser-driven attack. Defensive and static only.
63
+ - No secret scanning of its own: the pipeline already has that surface (`pre-commit-check.sh` / the egress gate, held in parity against `secret-patterns.json`). Call it; do not hand-roll a fourth list.
64
+ - No store-policy audit: that is the separate `store-ready` device pass under `/multi-agent:test`.
@@ -0,0 +1,53 @@
1
+ # OWASP Mobile Top 10 2024 - review checklist with CWE mapping
2
+
3
+ The category id (`M1:2024` .. `M10:2024`) goes in a finding's `security.owaspCategory`; the CWE that names the actual weakness goes in `security.cwe`. For the fuller mobile catalog that ties these to MASVS / MASTG and the store-compliance rules, the swift-security compliance mapping is the deeper reference; this is the review checklist.
4
+
5
+ ## M1:2024 - Improper Credential Usage
6
+
7
+ Check: hardcoded API keys / passwords / tokens, credentials in source or resource files, credentials in logs, a secret shipped in the binary.
8
+ CWEs: CWE-798 (hardcoded credentials), CWE-259, CWE-522.
9
+
10
+ ## M2:2024 - Inadequate Supply Chain Security
11
+
12
+ Check: a dependency at a version with a known advisory, an unpinned SDK, a build step pulling an unverified artifact. This is the `security_dep_inventory` -> CVE path (`A06:2021`'s mobile analog).
13
+ CWEs: CWE-1104, plus the advisory's CWE.
14
+
15
+ ## M3:2024 - Insecure Authentication/Authorization
16
+
17
+ Check: auth decided client-side, a token accepted without verification, missing authorization on a sensitive action, biometric gate that only hides UI without protecting data.
18
+ CWEs: CWE-287, CWE-306 (missing authentication), CWE-862, CWE-863.
19
+
20
+ ## M4:2024 - Insufficient Input/Output Validation
21
+
22
+ Check: untrusted input into a SQL/content-provider query, a WebView `evaluateJavascript` / `postMessage` handler trusting page content, deep-link parameters used without validation, format-string or path built from input.
23
+ CWEs: CWE-20, CWE-79 (WebView XSS), CWE-89, CWE-22.
24
+
25
+ ## M5:2024 - Insecure Communication
26
+
27
+ Check: HTTP instead of HTTPS, disabled ATS / cleartext-traffic permitted, `TrustManager` that accepts all certs, missing certificate pinning on a sensitive endpoint, ignored TLS errors.
28
+ CWEs: CWE-319 (cleartext), CWE-295 (improper certificate validation).
29
+
30
+ ## M6:2024 - Inadequate Privacy Controls
31
+
32
+ Check: PII collected without need, location/contacts/identifiers sent off-device without disclosure, tracking before consent, PII in logs or analytics events.
33
+ CWEs: CWE-359, CWE-200, CWE-532.
34
+
35
+ ## M7:2024 - Insufficient Binary Protections
36
+
37
+ Check: no tamper/integrity check where the threat model needs one, debug symbols or verbose logging left in a release build, an easily-reversible secret embedded in the binary. Judge against the threat model - most apps do not need anti-reversing, and a `blocking` here needs a named attacker.
38
+ CWEs: CWE-656, CWE-489 (debug code left in).
39
+
40
+ ## M8:2024 - Security Misconfiguration
41
+
42
+ Check: an exported Android component with no permission, `android:debuggable=true` or `allowBackup=true` on sensitive data, an overly-broad entitlement, a permissive `network_security_config`, default or weak settings.
43
+ CWEs: CWE-16, CWE-276 (incorrect default permissions), CWE-926 (improper export).
44
+
45
+ ## M9:2024 - Insecure Data Storage
46
+
47
+ Check: sensitive data in `UserDefaults` / `SharedPreferences` / plist / plain files instead of the Keychain / Keystore, a database without encryption, a cache holding secrets, pasteboard leakage.
48
+ CWEs: CWE-312 (cleartext storage), CWE-922 (insecure storage of sensitive info).
49
+
50
+ ## M10:2024 - Insufficient Cryptography
51
+
52
+ Check: weak or deprecated algorithm (MD5/SHA1 for integrity, DES/ECB), a hardcoded key/IV, a home-rolled cipher, a predictable random source for a security purpose.
53
+ CWEs: CWE-327 (broken/risky algorithm), CWE-326, CWE-330 (insufficient randomness), CWE-338.
@@ -0,0 +1,56 @@
1
+ # OWASP Top 10 2021 (Web / API) - review checklist with CWE mapping
2
+
3
+ The category id (`A01:2021` .. `A10:2021`) goes in a finding's `security.owaspCategory`; the CWE most associated with the specific weakness goes in `security.cwe`. One weakness per finding. The CWEs listed per category are the common ones, not the whole set - pick the one that names the actual defect.
4
+
5
+ ## A01:2021 - Broken Access Control
6
+
7
+ Check: missing authorization on a route or action, IDOR (object id from the request trusted without an ownership check), path traversal, forced browsing, CORS misconfiguration allowing credentialed cross-origin reads, privilege escalation through a mass-assignable field.
8
+ CWEs: CWE-284, CWE-285, CWE-639 (IDOR), CWE-862 (missing authorization), CWE-863 (incorrect authorization), CWE-22 (path traversal), CWE-352 (CSRF).
9
+ Evidence to cite: the handler that reads an id from input and queries without a `where owner = current_user` predicate; a route with no auth middleware.
10
+
11
+ ## A02:2021 - Cryptographic Failures
12
+
13
+ Check: secrets or PII sent or stored in clear, weak or deprecated algorithms (MD5, SHA1 for passwords, DES, ECB), hardcoded keys, missing TLS, disabled certificate validation, predictable IVs/nonces, passwords hashed without a slow KDF (bcrypt/scrypt/argon2).
14
+ CWEs: CWE-311 (missing encryption), CWE-319 (cleartext transmission), CWE-327 (broken/risky algorithm), CWE-326 (inadequate strength), CWE-798 (hardcoded credentials), CWE-916 (weak password hash).
15
+
16
+ ## A03:2021 - Injection
17
+
18
+ Check: SQL/NoSQL/ORM query built by string concatenation of untrusted input, OS command built from input, LDAP/XPath injection, unsanitized input reflected into HTML (XSS), template injection, header injection.
19
+ CWEs: CWE-89 (SQL), CWE-78 (OS command), CWE-79 (XSS), CWE-90 (LDAP), CWE-94 (code injection), CWE-943 (NoSQL/query).
20
+ Evidence: the untrusted source and the sink on the same path, with no parameterization or encoding between.
21
+
22
+ ## A04:2021 - Insecure Design
23
+
24
+ Check: a missing control the design needed - no rate limit on a credential endpoint, no anti-automation on a costly action, trust placed in a client-supplied value that decides server behaviour, a workflow that can be replayed.
25
+ CWEs: CWE-73, CWE-183, CWE-209 (info leak by design), CWE-256, CWE-501 (trust boundary violation), CWE-522.
26
+
27
+ ## A05:2021 - Security Misconfiguration
28
+
29
+ Check: debug or verbose errors in production, default credentials, an unnecessary feature or port enabled, permissive CORS, missing security headers (CSP, HSTS, X-Content-Type-Options), directory listing, an overly-permissive cloud bucket or IAM policy in config.
30
+ CWEs: CWE-16, CWE-611 (XXE), CWE-732 (incorrect permissions), CWE-1032, CWE-756.
31
+
32
+ ## A06:2021 - Vulnerable and Outdated Components
33
+
34
+ Check: a dependency at a version with a known advisory. This is the `security_dep_inventory` -> CVE-lookup path. Carry the `cve` and the advisory's CWE; put the finding on the manifest at line 0.
35
+ CWEs: CWE-1104, plus the advisory's own CWE.
36
+
37
+ ## A07:2021 - Identification and Authentication Failures
38
+
39
+ Check: credential stuffing possible (no throttle/lockout), weak password policy, session id in the URL, session not rotated on login, missing or weak MFA, JWT with `alg:none` accepted or signature not verified, long-lived non-revocable tokens.
40
+ CWEs: CWE-287 (improper auth), CWE-297, CWE-384 (session fixation), CWE-521 (weak password), CWE-613 (insufficient expiration), CWE-347 (improper signature verification).
41
+
42
+ ## A08:2021 - Software and Data Integrity Failures
43
+
44
+ Check: insecure deserialization of untrusted data, an update or plugin loaded without signature verification, a CI/CD step pulling an unpinned or unverified artifact, client-side data trusted without integrity check.
45
+ CWEs: CWE-502 (deserialization), CWE-345, CWE-494 (download without integrity check), CWE-829.
46
+
47
+ ## A09:2021 - Security Logging and Monitoring Failures
48
+
49
+ Check: security-relevant events not logged (auth failures, access-control denials, high-value actions), logs containing secrets or PII, no alerting path. Note it as a `suggestion`/`important` gap, rarely `blocking` on its own.
50
+ CWEs: CWE-778 (insufficient logging), CWE-532 (secrets in logs), CWE-223.
51
+
52
+ ## A10:2021 - Server-Side Request Forgery (SSRF)
53
+
54
+ Check: the server fetches a URL built from user input without an allowlist, letting a caller reach internal services, cloud metadata endpoints, or the loopback interface.
55
+ CWEs: CWE-918.
56
+ Evidence: the input-derived URL reaching an HTTP client with no host allowlist or scheme restriction.
@@ -3,7 +3,7 @@
3
3
  > Auto-generated by `pipeline/scripts/build-skills-index.mjs` - do not hand-edit.
4
4
  > Regenerate with `node pipeline/scripts/build-skills-index.mjs`.
5
5
 
6
- **216 skills** across 2 groups.
6
+ **213 skills** across 2 groups.
7
7
 
8
8
  | Group | Name | Platform | Description |
9
9
  |-------|------|----------|-------------|
@@ -117,17 +117,14 @@
117
117
  | core | `multi-agent-jira` | - | List open Jira issues, pick one, and launch the multi-agent pipeline. Use when a Jira issue should be picked up and started without knowing |
118
118
  | core | `multi-agent-kill` | - | Stop the given task, then remove its worktree and branch. Asks for confirmation. Use when a running or stuck task should be stopped and its |
119
119
  | core | `multi-agent-language` | - | Toggle outputLanguage (assistant explanations, picker questions, PR/Jira/Confluence bodies). promptLanguage is fixed to English; commit mess |
120
- | core | `multi-agent-local` | - | Full pipeline in local mode - no worktree, runs directly on the current branch. Use when the full pipeline should run on the current branc |
121
- | core | `multi-agent-local-autopilot` | - | Full pipeline + local + autopilot - no worktree, no confirmations, all 6 phases run end-to-end on the current branch. Use when the full pi |
122
120
  | core | `multi-agent-log` | - | Show the agent-log.md for the given task. With no ID, shows the most recent task. Use when asked what a task did, or to read its log. |
123
- | core | `multi-agent-manual-test` | - | Switch to the active task's branch and prepare it for manual testing in Xcode. Phase 5 standalone (the UI Bug Hunter lives at multi-agent-te |
121
+ | core | `multi-agent-manual-test` | - | Switch to the active task's branch and prepare it for manual testing in Xcode. Phase 3 user-test step, standalone (the UI Bug Hunter lives a |
124
122
  | core | `multi-agent-model` | - | Turn the top model rung on or off and keep the cost ledger's pricing in step with it. Use when asked to enable or disable Fable, or which mo |
125
123
  | core | `multi-agent-prune-logs` | - | Delete per-task project logs under ~/.claude/logs/multi-agent (filter by age/project/task). Audit trail + metrics are preserved. Dry-run fir |
126
124
  | core | `multi-agent-prune-prompts` | - | Zero-base prompt review: measure the always-on instruction footprint, classify every rule block, propose keep/trial-removal/delete; applies |
127
125
  | core | `multi-agent-purge` | - | ⚠️ Wipes every worktree, branch, log, and state file. Irreversible; asks for double confirmation. Use when every worktree, branch, log and s |
128
126
  | core | `multi-agent-refactor` | - | Analyse the project: extract adapted best-practices, hunt real bugs + improvement areas, check upstream drift of derived skills, research th |
129
- | core | `multi-agent-resume` | - | Resume a stopped or failed task from the phase where it left off. Use when a task stopped or failed and should carry on from where it left o |
130
- | core | `multi-agent-resume-local` | - | Continue already-done LOCAL work through the pipeline tail: Review (with its build gate) → Commit/PR → Report (technical analysis + Jira tes |
127
+ | core | `multi-agent-resume` | - | Pick up unfinished work: a pipeline run that stopped mid-phase, or work already written on the current branch that never went through the pi |
131
128
  | core | `multi-agent-review` | - | Run parallel review on a branch diff or a Pull Request: 3 models on Claude Code (Fable + Opus + Sonnet), 3 models on Copilot CLI (GPT + Opus |
132
129
  | core | `multi-agent-review-analysis` | - | Review a written analysis document instead of a diff: resolve it from a path, a Confluence page or a Jira issue, run the deterministic gates |
133
130
  | core | `multi-agent-review-issue` | - | Assess whether a GitHub issue is ready for multi-agent development: fetch it, grade scope / acceptance criteria / repro / design / API / sta |
@@ -1,132 +0,0 @@
1
- ---
2
- description: "Full pipeline in local mode - no worktree, runs directly on the current branch. Use when the full pipeline should run on the current branch without creating a worktree."
3
- description-tr: "Tam pipeline lokal modda - worktree yok, doğrudan mevcut branch üzerinde çalışır."
4
- allowed-tools: Agent, Bash, Read, Write, Edit, Glob, Grep, TaskCreate, TaskUpdate, TaskList, TaskGet, AskUserQuestion, WebFetch, WebSearch, Skill
5
- ---
6
-
7
- # multi-agent local - Full Pipeline, Local Branch
8
-
9
- > **Language (read FIRST)**: Before any status output, read `prefs.global.outputLanguage` and render every conversational line in it. `AskUserQuestion` renders its `question`, option `label`s and option `description`s in `outputLanguage`; only `header` stays English (<=12-char chip); external payload bodies follow `outputLanguage` too (identifiers, commit messages, branch names stay English). Full contract: `$HOME/.claude/multi-agent-refs/rules.md` "Language Application".
10
-
11
- The full pipeline in normal mode (Plan Approval Gate + parallel review + triage included), running on the **current branch with no worktree**. Dedicated alias for the `multi-agent "task" --local` flag form. Phase 5 (User Test) is skipped because there is no worktree to check the change out from; the other seven phases all run.
12
-
13
- ## When to use it
14
-
15
- - You don't want a separate worktree open - keeps the editor / IDE in one folder
16
- - You're already on the right branch and just want pipeline discipline
17
- - Small project or prototype - the worktree overhead is unnecessary
18
-
19
- ## When NOT to use it
20
-
21
- - Multiple parallel tasks at the same time - you lose worktree isolation
22
- - Long-iteration changes on a production repo - branch-switch friction shows up
23
- - Multi-repo tasks - `--local` is locked to a single repo
24
-
25
- ## Pipeline
26
-
27
- Same phase count and order as the normal pipeline:
28
-
29
- ```
30
- Phase 0: Init → project detection, branch check, state (NO worktree)
31
- Phase 1: Plan (analysis) → codebase scan (parallel explore agents, Opus)
32
- Phase 1: Plan → task breakdown, Plan Approval Gate (approval loop)
33
- Phase 2: Dev → TDD (Sonnet), build queue
34
- Phase 3: Review → deterministic gates + parallel review + Fable triage
35
- Phase 4: Commit → pre-commit checkout prompt, commit + push + PR
36
- Phase 5: Report → Jira / Wiki / Confluence + log + knowledge/memory
37
- ```
38
-
39
- ## Delegation
40
-
41
- This command routes to the orchestrator with the `--local` flag set. The Phase 0-7 contract from `$HOME/.claude/multi-agent-refs/phases/phase-0-init.md` and the later phase docs applies as-is - only the worktree step is skipped, and `state.projects[*].worktreePath` stays `null`.
42
-
43
- Read the routing table in `$HOME/.claude/commands/multi-agent/SKILL.md` and apply Phase 0 Step 8 in local mode (no worktree: continue on the current branch, state file under `.claude/logs/multi-agent/{project}/{taskId}/`).
44
-
45
- ## Examples
46
-
47
- ```bash
48
- /multi-agent:local "PROJ-12345" # Jira
49
- /multi-agent:local "#42" # GitHub issue
50
- /multi-agent:local "LoginView dark mode fix" # Free-text
51
- ```
52
- ## Required: outward-facing payload contracts
53
-
54
- Before writing anything outward-facing - PR body, Jira comment, Confluence page, closing report - load `$HOME/.claude/multi-agent-refs/payload-contracts.md`. It names the canonical section set for each payload, the markup dialect per surface (PR body is Markdown, Jira is wiki markup - mixing them is a defect), and the token/duration numbers the closing report must carry. Improvising a payload shape from memory is the most common failure of the short modes.
55
-
56
- ## Required: Phase Tracker Contract
57
-
58
- **The phase tracker is mandatory** - the agent cannot skip it. Full spec: [`$HOME/.claude/multi-agent-refs/tracker-contract.md`]($HOME/.claude/multi-agent-refs/tracker-contract.md).
59
-
60
- > **Local mode:** no worktree is created, work happens on the current branch. Phase 0 Init still calls `init` - the `--local` flag is stored in tracker-state.json, and `:resume` restores the correct CWD.
61
-
62
- Two channels run in parallel at every phase boundary:
63
-
64
- 1. **State channel** (every CLI, identical): `phase-tracker.sh` writes to `tracker-state.json`. Drives `:resume`, `:log`, `:status`.
65
- 2. **Visual channel** (CLI-specific): native widget on Claude Code, ANSI render on every other CLI. Without it the user sees no phase progress.
66
-
67
- ```bash
68
- # Phase 0, very first shell call (every CLI):
69
- bash $HOME/.claude/scripts/phase-tracker.sh init "$TASK_ID"
70
- bash $HOME/.claude/scripts/phase-tracker.sh add 0 "Init"
71
- bash $HOME/.claude/scripts/phase-tracker.sh tiles
72
- bash $HOME/.claude/scripts/phase-tracker.sh update 0 in_progress
73
-
74
- # Phase 0 Step 7.5, immediately after the depth answer - the first moment this
75
- # mode knows its phase set. Full:
76
- for p in "1:Plan" "2:Dev" "3:Review" "4:Commit" "5:Report"; do
77
- bash $HOME/.claude/scripts/phase-tracker.sh add "${p%%:*}" "${p#*:}"
78
- done
79
- # Short (Analysis and Planning are not run, so they get no tile at all):
80
- for p in "2:Dev" "3:Review" "4:Commit" "5:Report"; do
81
- bash $HOME/.claude/scripts/phase-tracker.sh add "${p%%:*}" "${p#*:}"
82
- done
83
- # Then the widget, narrowed to the phases that do not have a tile yet:
84
- bash $HOME/.claude/scripts/phase-tracker.sh tiles --new
85
-
86
- # Every phase boundary (every CLI):
87
- bash $HOME/.claude/scripts/phase-tracker.sh update <N> in_progress|completed|failed|skipped
88
-
89
- # After every LLM call (every CLI):
90
- bash $HOME/.claude/scripts/phase-tracker.sh tokens <N> <in> <out> [cached]
91
- ```
92
-
93
- ### Visual channel - Claude Code (native TaskList widget, required)
94
-
95
- In Claude Code the agent MUST also drive the native TaskList widget so the user sees a sticky phase tile stack - this is the only progress signal Claude Code surfaces. Skipping these calls is the #1 source of "I don't see any phases" complaints.
96
-
97
- **TaskCreate ordering (strict)**: All TaskCreate calls in a registration batch fire in strict phase-number order BEFORE any TaskUpdate in that batch, and a later batch only ever appends phases numbered above everything already registered. This mode registers in two batches (Step -1, then Step 7.5), so `tiles --new` narrows the second one and the Phase 0 tile is never created twice. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks (e.g. `1 ✓ · 2 ✓ · 4 ✓ · 0 ▶ · 3 ☐`) even when the underlying state is correct. Pre-marking phases as completed/skipped before Phase 0 starts is FORBIDDEN - register the tile in order, then flip status via TaskUpdate when the phase actually short-circuits. Full contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
98
-
99
- ```text
100
- # Register one tile per phase, capture the taskId, persist it:
101
- for each phase in 0:Init at Step -1, then 1:Plan, 2:Dev, 3:Review, 4:Commit, 5:Report (Full) or 2:Dev, 3:Review, 4:Commit, 5:Report (Short) at Step 7.5:
102
- TaskCreate({ subject: "Phase <N>: <Name>", activeForm: "<doing-form>" })
103
- -> returns taskId
104
- bash $HOME/.claude/scripts/phase-tracker.sh meta <N> tasklist_id "<taskId>"
105
-
106
- # Phase entry - flip the tile to in_progress alongside the state update:
107
- TaskUpdate({ taskId: <saved>, status: "in_progress" })
108
- bash $HOME/.claude/scripts/phase-tracker.sh update <N> in_progress
109
-
110
- # Active sub-step inside a phase - update activeForm so the spinner header reflects what's happening now:
111
- TaskUpdate({ taskId: <saved>, activeForm: "Editing TopBarView.swift" })
112
-
113
- # Phase exit - flip to completed/failed/skipped on both channels:
114
- TaskUpdate({ taskId: <saved>, status: "completed" })
115
- bash $HOME/.claude/scripts/phase-tracker.sh update <N> completed
116
- ```
117
-
118
- `--local` mode TaskCreates all 6 phases (no phase is skipped).
119
-
120
- #### TaskCreate ordering (strict)
121
-
122
- **All TaskCreate calls in a batch fire in strict phase-number order BEFORE any TaskUpdate is applied.** For `--local` that means: Phase 0 at Step -1, then the rest in ascending order at Step 7.5 (Phase 0 → Phase 1 → Phase 2 → Phase 3 → Phase 4 → Phase 5 minus whatever the depth answer drops). The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks. Full ordering contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".
123
-
124
- ### Visual channel - Copilot CLI / plain shell
125
-
126
- These CLIs have no TaskList widget. After every state change the agent calls render, which prints a bordered ANSI card as the last tool result so the user sees an updated phase table:
127
-
128
- ```bash
129
- bash $HOME/.claude/scripts/phase-tracker.sh render
130
- ```
131
-
132
- Do NOT call TaskCreate on these CLIs - the tool does not exist and the call fails.