@mmerterden/multi-agent-pipeline 19.1.4 → 20.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (137) hide show
  1. package/CHANGELOG.md +123 -0
  2. package/README.md +19 -36
  3. package/README.tr.md +18 -35
  4. package/SECURITY.md +3 -3
  5. package/docs/adr/0002-instruction-driven-flag.md +6 -5
  6. package/docs/adr/0005-lazy-phase-docs.md +2 -2
  7. package/docs/adr/0008-installer-modularization-and-secret-leak-defense.md +1 -0
  8. package/docs/adr/0009-claude-stack-skills-plugin-only.md +1 -1
  9. package/docs/adr/0010-own-code-graph.md +5 -4
  10. package/docs/adr/0011-dormant-ci.md +10 -1
  11. package/docs/adr/0012-macos-only.md +2 -2
  12. package/docs/adr/0013-lsp-code-intelligence.md +2 -2
  13. package/docs/adr/0014-six-phase-consolidation.md +9 -9
  14. package/docs/adr/0015-one-pipeline-no-depth-answer.md +83 -0
  15. package/docs/adr/0016-the-run-shape-is-asked-not-typed.md +69 -0
  16. package/docs/adr/README.md +18 -16
  17. package/docs/architecture.md +2 -2
  18. package/docs/ecosystem.md +5 -5
  19. package/docs/facts.json +7 -9
  20. package/docs/features.md +4 -5
  21. package/docs/token-budget-history.md +1 -1
  22. package/install/_codex-agents.mjs +1 -1
  23. package/install/_common.mjs +9 -1
  24. package/install/templates/copilot-instructions.md +7 -16
  25. package/manifest.json +133 -129
  26. package/package.json +1 -1
  27. package/pipeline/agents/code-reviewer.md +2 -2
  28. package/pipeline/agents/dev-critic.md +5 -5
  29. package/pipeline/agents/security-auditor.md +80 -72
  30. package/pipeline/commands/figma-to-swiftui.md +1 -1
  31. package/pipeline/commands/multi-agent/SKILL.md +7 -9
  32. package/pipeline/commands/multi-agent/analysis/SKILL.md +2 -0
  33. package/pipeline/commands/multi-agent/analysis-jira/SKILL.md +2 -0
  34. package/pipeline/commands/multi-agent/analysis-resolve/SKILL.md +2 -0
  35. package/pipeline/commands/multi-agent/autopilot/SKILL.md +2 -0
  36. package/pipeline/commands/multi-agent/autopilot-on/SKILL.md +2 -0
  37. package/pipeline/commands/multi-agent/autopilot-status/SKILL.md +1 -1
  38. package/pipeline/commands/multi-agent/build-optimize/SKILL.md +2 -0
  39. package/pipeline/commands/multi-agent/channels/SKILL.md +2 -2
  40. package/pipeline/commands/multi-agent/create-jira/SKILL.md +2 -0
  41. package/pipeline/commands/multi-agent/design-check/SKILL.md +1 -1
  42. package/pipeline/commands/multi-agent/diff-explain/SKILL.md +1 -1
  43. package/pipeline/commands/multi-agent/forget/SKILL.md +2 -0
  44. package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +4 -2
  45. package/pipeline/commands/multi-agent/help/SKILL.md +23 -27
  46. package/pipeline/commands/multi-agent/ios-coding-standard/SKILL.md +5 -4
  47. package/pipeline/commands/multi-agent/issue/SKILL.md +2 -0
  48. package/pipeline/commands/multi-agent/jira/SKILL.md +2 -0
  49. package/pipeline/commands/multi-agent/language/SKILL.md +2 -0
  50. package/pipeline/commands/multi-agent/prune-logs/SKILL.md +2 -0
  51. package/pipeline/commands/multi-agent/purge/SKILL.md +2 -0
  52. package/pipeline/commands/multi-agent/resume/SKILL.md +177 -48
  53. package/pipeline/commands/multi-agent/save/SKILL.md +2 -0
  54. package/pipeline/commands/multi-agent/scan/SKILL.md +2 -2
  55. package/pipeline/commands/multi-agent/security-review/SKILL.md +52 -0
  56. package/pipeline/commands/multi-agent/stack/SKILL.md +2 -0
  57. package/pipeline/commands/multi-agent/sync/SKILL.md +7 -8
  58. package/pipeline/commands/multi-agent/test-screenshots/SKILL.md +2 -0
  59. package/pipeline/commands/multi-agent/uninstall/SKILL.md +2 -0
  60. package/pipeline/lib/repo-hygiene.sh +1 -1
  61. package/pipeline/multi-agent-refs/analysis/render.md +1 -1
  62. package/pipeline/multi-agent-refs/analysis/resolve.md +1 -1
  63. package/pipeline/multi-agent-refs/analysis/synthesis.md +1 -1
  64. package/pipeline/multi-agent-refs/analysis-template.md +1 -1
  65. package/pipeline/multi-agent-refs/component-dispatch.md +5 -13
  66. package/pipeline/multi-agent-refs/cross-cli-contract.md +14 -15
  67. package/pipeline/multi-agent-refs/features/external-context-injection.md +2 -0
  68. package/pipeline/multi-agent-refs/features/review-delta.md +1 -1
  69. package/pipeline/multi-agent-refs/features/review-multi-repo.md +3 -3
  70. package/pipeline/multi-agent-refs/features/security-audit.md +55 -0
  71. package/pipeline/multi-agent-refs/features/skill-conformance.md +1 -1
  72. package/pipeline/multi-agent-refs/features/visual-evidence.md +2 -1
  73. package/pipeline/multi-agent-refs/features/worktree-finalize.md +1 -1
  74. package/pipeline/multi-agent-refs/generate-issue.md +2 -0
  75. package/pipeline/multi-agent-refs/issue-jira-triad.md +2 -0
  76. package/pipeline/multi-agent-refs/keychain.md +2 -0
  77. package/pipeline/multi-agent-refs/knowledge.md +0 -7
  78. package/pipeline/multi-agent-refs/outside-the-pipeline.md +1 -1
  79. package/pipeline/multi-agent-refs/payload-contracts.md +1 -1
  80. package/pipeline/multi-agent-refs/phases/modes.md +33 -109
  81. package/pipeline/multi-agent-refs/phases/operations.md +2 -0
  82. package/pipeline/multi-agent-refs/phases/phase-0-init.md +23 -42
  83. package/pipeline/multi-agent-refs/phases/phase-1-plan.md +7 -18
  84. package/pipeline/multi-agent-refs/phases/phase-2-dev.md +13 -44
  85. package/pipeline/multi-agent-refs/phases/phase-3-review.md +28 -37
  86. package/pipeline/multi-agent-refs/phases/phase-4-commit.md +6 -6
  87. package/pipeline/multi-agent-refs/phases/phase-5-report.md +3 -3
  88. package/pipeline/multi-agent-refs/phases.md +9 -11
  89. package/pipeline/multi-agent-refs/progress-contract.md +1 -1
  90. package/pipeline/multi-agent-refs/readiness-review.md +2 -0
  91. package/pipeline/multi-agent-refs/rules.md +1 -1
  92. package/pipeline/multi-agent-refs/threat-model.md +39 -0
  93. package/pipeline/multi-agent-refs/tracker-contract.md +9 -40
  94. package/pipeline/multi-agent-refs/wiki-capture.md +3 -2
  95. package/pipeline/preferences-template.json +2 -2
  96. package/pipeline/rules/figma-pipeline.md +1 -1
  97. package/pipeline/schemas/agent-state.schema.json +28 -10
  98. package/pipeline/schemas/migrations/prefs-2.7.0-to-2.8.0.mjs +33 -0
  99. package/pipeline/schemas/phases.json +4 -26
  100. package/pipeline/schemas/prefs.schema.json +5 -9
  101. package/pipeline/schemas/reviewer-output.schema.json +99 -2
  102. package/pipeline/schemas/security-finding.schema.json +144 -0
  103. package/pipeline/scripts/_stack-routing.mjs +1 -0
  104. package/pipeline/scripts/cost-table.json +1 -1
  105. package/pipeline/scripts/gc-abandoned.sh +16 -9
  106. package/pipeline/scripts/gc-refs.sh +1 -1
  107. package/pipeline/scripts/gen-mode-dispatch.mjs +11 -41
  108. package/pipeline/scripts/migrate-prefs.mjs +18 -17
  109. package/pipeline/scripts/phase-tracker.sh +2 -2
  110. package/pipeline/scripts/phase0-exit-gate.mjs +1 -1
  111. package/pipeline/scripts/plan-coverage-gate.mjs +3 -3
  112. package/pipeline/scripts/render-work-summary.sh +7 -4
  113. package/pipeline/scripts/run-aggregator.mjs +1 -1
  114. package/pipeline/scripts/usage-report.mjs +0 -2
  115. package/pipeline/scripts/worktree-finalize.sh +2 -2
  116. package/pipeline/skills/.skill-manifest.json +17 -21
  117. package/pipeline/skills/.skills-index.json +6 -39
  118. package/pipeline/skills/shared/README.md +5 -8
  119. package/pipeline/skills/shared/core/multi-agent/SKILL.md +11 -15
  120. package/pipeline/skills/shared/core/multi-agent-autopilot-status/SKILL.md +1 -1
  121. package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +13 -16
  122. package/pipeline/skills/shared/core/multi-agent-ios-coding-standard/SKILL.md +2 -3
  123. package/pipeline/skills/shared/core/multi-agent-resume/SKILL.md +51 -15
  124. package/pipeline/skills/shared/core/multi-agent-scan/SKILL.md +2 -2
  125. package/pipeline/skills/shared/core/multi-agent-security-review/SKILL.md +29 -0
  126. package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +7 -7
  127. package/pipeline/skills/shared/external/security-review/SKILL.md +64 -0
  128. package/pipeline/skills/shared/external/security-review/references/owasp-mobile-top10-2024.md +53 -0
  129. package/pipeline/skills/shared/external/security-review/references/owasp-web-api-top10-2021.md +56 -0
  130. package/pipeline/skills/skills-index.md +3 -6
  131. package/pipeline/commands/multi-agent/local/SKILL.md +0 -132
  132. package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +0 -142
  133. package/pipeline/commands/multi-agent/resume-local/SKILL.md +0 -114
  134. package/pipeline/commands/security-review.md +0 -6
  135. package/pipeline/skills/shared/core/multi-agent-local/SKILL.md +0 -41
  136. package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +0 -55
  137. package/pipeline/skills/shared/core/multi-agent-resume-local/SKILL.md +0 -51
@@ -0,0 +1,33 @@
1
+ /**
2
+ * prefs-2.7.0-to-2.8.0.mjs - preferences migration
3
+ *
4
+ * v20.0.0 folds `/multi-agent:resume-local` into `/multi-agent:resume`. One
5
+ * command now asks which unfinished work to pick up: a run that stopped
6
+ * mid-phase, or a branch carrying work no run ever produced. The preference
7
+ * that decides whether triage-accepted findings are auto-fixed applies to the
8
+ * second path, so it moves with it: `global.resumeLocal` -> `global.resume`.
9
+ *
10
+ * The value is carried, never reset. Someone who turned auto-fix on for the
11
+ * tail meant it for exactly the work this path still handles.
12
+ *
13
+ * @param {object} data - parsed multi-agent-preferences.json
14
+ * @returns {object} - migrated data with schemaVersion "2.8.0"
15
+ */
16
+ export default function migrate(data) {
17
+ const out = JSON.parse(JSON.stringify(data));
18
+
19
+ out.schemaVersion = "2.8.0";
20
+
21
+ if (!out.global || typeof out.global !== "object") return out;
22
+
23
+ const legacy = out.global.resumeLocal;
24
+ if (!out.global.resume || typeof out.global.resume !== "object") {
25
+ out.global.resume = {};
26
+ }
27
+ if (typeof out.global.resume.autoFix !== "boolean") {
28
+ out.global.resume.autoFix = typeof legacy?.autoFix === "boolean" ? legacy.autoFix : false;
29
+ }
30
+ delete out.global.resumeLocal;
31
+
32
+ return out;
33
+ }
@@ -6,7 +6,6 @@
6
6
  "phaseSchema": 2,
7
7
  "phaseSchemaNote": "The phase vocabulary generation. metrics.jsonl is append-only and v19.0.0 renumbered the phases, so every line written from v19.0.0 on carries this number and a line without the field is generation 1. Aggregators pick their name table from it; without it phase 3 means Dev in old rows and Review in new ones and no reader can tell them apart.",
8
8
  "legacyPhaseNames": {
9
- "note": "Generation 1 phase names, kept so an aggregator can label a historical metrics.jsonl row correctly instead of printing the current name for a number that meant something else. Read-only history: nothing emits these any more. The generation 1 -> 2 number map is not repeated here, it is the `was` array on each phase above.",
10
9
  "0": "Init",
11
10
  "1": "Analysis",
12
11
  "2": "Planning",
@@ -14,7 +13,8 @@
14
13
  "4": "Review",
15
14
  "5": "Test",
16
15
  "6": "Commit",
17
- "7": "Report"
16
+ "7": "Report",
17
+ "note": "Generation 1 phase names, kept so an aggregator can label a historical metrics.jsonl row correctly instead of printing the current name for a number that meant something else. Read-only history: nothing emits these any more. The generation 1 -> 2 number map is not repeated here, it is the `was` array on each phase above."
18
18
  },
19
19
  "phases": [
20
20
  {
@@ -64,42 +64,20 @@
64
64
  "modes": {
65
65
  "full": {
66
66
  "phases": [0, 1, 2, 3, 4, 5],
67
- "local": false,
68
- "autopilot": false,
69
- "depth": true
67
+ "autopilot": false
70
68
  },
71
69
  "autopilot": {
72
70
  "phases": [0, 1, 2, 3, 4, 5],
73
- "local": false,
74
- "autopilot": true
75
- },
76
- "local": {
77
- "phases": [0, 1, 2, 3, 4, 5],
78
- "local": true,
79
- "autopilot": false,
80
- "depth": true
81
- },
82
- "local-autopilot": {
83
- "phases": [0, 1, 2, 3, 4, 5],
84
- "local": true,
85
71
  "autopilot": true
86
72
  },
87
- "full-local": {
88
- "phases": [0, 1, 2, 3, 4, 5],
89
- "local": true,
90
- "autopilot": false,
91
- "depth": true
92
- },
93
73
  "analysis": {
94
74
  "phases": [0, 1, 3, 4, 5],
95
- "local": false,
96
75
  "autopilot": false
97
76
  }
98
77
  },
99
78
  "thresholds": {
100
79
  "waitingFromPhase": 4,
101
80
  "mcpAllowedThroughPhase": 1,
102
- "shortRunFromPhase": 2,
103
- "note": "Phase numbers other code compares against, named here so a renumbering moves them with the contract. waitingFromPhase: runs-index.mjs groups a run as 'waiting on you' from Commit on. mcpAllowedThroughPhase: Figma MCP is reachable only through Plan; smoke-no-mcp-in-dev-phases.sh fails any recorded call at a higher phase. Its literal threshold is unchanged from the eight-phase contract on purpose - Analysis was 1 and is now inside Plan, also 1, so the permitted set {0,1} is identical. shortRunFromPhase: the depth picker's Short set starts here."
81
+ "note": "Phase numbers other code compares against, named here so a renumbering moves them with the contract. waitingFromPhase: runs-index.mjs groups a run as 'waiting on you' from Commit on. mcpAllowedThroughPhase: Figma MCP is reachable only through Plan; smoke-no-mcp-in-dev-phases.sh fails any recorded call at a higher phase. Its literal threshold is unchanged from the eight-phase contract on purpose - Analysis was 1 and is now inside Plan, also 1, so the permitted set {0,1} is identical."
104
82
  }
105
83
  }
@@ -14,8 +14,8 @@
14
14
  "properties": {
15
15
  "schemaVersion": {
16
16
  "type": "string",
17
- "enum": ["2.0.0", "2.1.0", "2.2.0", "2.3.0", "2.4.0", "2.5.0", "2.6.0", "2.7.0"],
18
- "description": "v2.0.0: pre-v3.7. v2.1.0: v3.7+ adds identities[].servicePatMap, platformIdentityRouting, recentGroups, recentBranches, serviceStatus, settings, expanded keychainMapping. v2.2.0: v6.0.0 formalizes v5.7 / v5.8 additions (reportChannels, reportContent with technicalAnalysis, wikiScope, autopilotReportTimeoutSeconds) that had been running as 2.1.0 sub-migrations without a proper version bump. v2.5.0: v14.0.0 adds global.skillConformance (Phase 4 criteria resolution) and declares global.ship.autoFix, which the tail command's spec had referenced as global.finish.autoFix without ever declaring it. v2.6.0: v15.0.0 renames global.ship to global.resumeLocal (/multi-agent:ship -> :resume-local). v2.7.0: v19.0.0 drops the analysis Lite mode (global.analysisPhase.mode keeps a single value 'full') and adds global.modelRouting, which ships disabled."
17
+ "enum": ["2.0.0", "2.1.0", "2.2.0", "2.3.0", "2.4.0", "2.5.0", "2.6.0", "2.7.0", "2.8.0"],
18
+ "description": "v2.0.0: pre-v3.7. v2.1.0: v3.7+ adds identities[].servicePatMap, platformIdentityRouting, recentGroups, recentBranches, serviceStatus, settings, expanded keychainMapping. v2.2.0: v6.0.0 formalizes v5.7 / v5.8 additions (reportChannels, reportContent with technicalAnalysis, wikiScope, autopilotReportTimeoutSeconds) that had been running as 2.1.0 sub-migrations without a proper version bump. v2.5.0: v14.0.0 adds global.skillConformance (Phase 4 criteria resolution) and declares global.ship.autoFix, which the tail command's spec had referenced as global.finish.autoFix without ever declaring it. v2.6.0: v15.0.0 renames global.ship to global.resumeLocal (/multi-agent:ship -> :resume-local). v2.7.0: v19.0.0 drops the analysis Lite mode (global.analysisPhase.mode keeps a single value 'full') and adds global.modelRouting, which ships disabled. v2.8.0: v20.0.0 folds :resume-local into :resume and renames global.resumeLocal to global.resume."
19
19
  },
20
20
  "global": {
21
21
  "type": "object",
@@ -1923,15 +1923,15 @@
1923
1923
  }
1924
1924
  }
1925
1925
  },
1926
- "resumeLocal": {
1926
+ "resume": {
1927
1927
  "type": "object",
1928
1928
  "additionalProperties": false,
1929
- "description": "/multi-agent:resume-local (formerly :ship, :finish) - the tail that runs review + build/test + PR + report over work already on the branch.",
1929
+ "description": "/multi-agent:resume - picking up unfinished work, either a run that stopped mid-phase or a branch carrying work no run produced. These keys apply to the second path, which runs review + build/test + PR + report over the branch diff.",
1930
1930
  "properties": {
1931
1931
  "autoFix": {
1932
1932
  "type": "boolean",
1933
1933
  "default": false,
1934
- "description": "When true, resume-local auto-fixes triage-accepted blocking/important findings and re-reviews instead of asking. Equivalent to passing `autopilot` on every run. Default false: resume-local operates on work the user wrote by hand, so silently rewriting it is the surprising option."
1934
+ "description": "When true, the tail path auto-fixes triage-accepted blocking/important findings and re-reviews instead of asking. Equivalent to passing `autopilot` on every run. Default false: this path operates on work the user wrote by hand, so silently rewriting it is the surprising option."
1935
1935
  }
1936
1936
  }
1937
1937
  },
@@ -2140,10 +2140,6 @@
2140
2140
  "format": "date-time",
2141
2141
  "description": "Timestamp of last task start against this project."
2142
2142
  },
2143
- "componentDevWorkflow": {
2144
- "type": "boolean",
2145
- "description": "Project follows the component (Configuration / View / Modifiers) development workflow, so Phase 2 dispatches the component skills instead of the standard TDD flow."
2146
- },
2147
2143
  "figmaConfigPath": {
2148
2144
  "type": "string",
2149
2145
  "description": "Path to per-project figma-config.json (for Figma pipeline projects)."
@@ -1,9 +1,9 @@
1
1
  {
2
2
  "$schema": "https://json-schema.org/draft/2020-12/schema",
3
3
  "$id": "https://github.com/mmerterden/multi-agent-pipeline/pipeline/schemas/reviewer-output.schema.json",
4
- "version": "1.3.0",
4
+ "version": "1.4.0",
5
5
  "title": "Multi-Agent Pipeline - Phase 3 reviewer output",
6
- "description": "Contract for a single code-reviewer subagent's JSON output in Phase 4 Step 2. Every host dispatches 3 parallel reviewers; the middle slot is CLI-aware: Claude Code (Fable, Opus, Sonnet); Copilot CLI (Opus, GPT-5.4, Sonnet); Codex CLI dispatches 3 (gpt-5.6 at xhigh, gpt-5.4, gpt-5.6 at medium). Every reviewer must return an object matching this shape before Opus triage merges them. v1.1.0 adds the rule-ID conformance checklist: when the orchestrator supplies a ${CRITERIA} block (Phase 3 Step 1.78), the reviewer must return one conformance row per selected rule ID. Findings alone cannot answer 'was this applied completely' - a reviewer that opened nothing returns the same empty findings array as one that checked everything. v1.2.0 adds the optional per-finding fingerprint (Phase 3 Step 2.1): the stable id a finding keeps across review rounds.",
6
+ "description": "Contract for a single code-reviewer subagent's JSON output in Phase 4 Step 2. Every host dispatches 3 parallel reviewers; the middle slot is CLI-aware: Claude Code (Fable, Opus, Sonnet); Copilot CLI (Opus, GPT-5.4, Sonnet); Codex CLI dispatches 3 (gpt-5.6 at xhigh, gpt-5.4, gpt-5.6 at medium). Every reviewer must return an object matching this shape before Opus triage merges them. v1.1.0 adds the rule-ID conformance checklist: when the orchestrator supplies a ${CRITERIA} block (Phase 3 Step 1.78), the reviewer must return one conformance row per selected rule ID. Findings alone cannot answer 'was this applied completely' - a reviewer that opened nothing returns the same empty findings array as one that checked everything. v1.2.0 adds the optional per-finding fingerprint (Phase 3 Step 2.1): the stable id a finding keeps across review rounds. v1.4.0 adds the optional per-finding `security` envelope (OWASP + CWE + CVSS + evidence + remediation): the security-auditor emits reviewer-output objects whose findings carry it, so a blocking security finding merges at Step 3.0 and blocks Phase 4 like any reviewer blocker. General reviewers omit it. Its shape mirrors security-finding.schema.json, kept in step by smoke-security-schema-parity.sh.",
7
7
  "type": "object",
8
8
  "additionalProperties": false,
9
9
  "required": ["findings", "approved"],
@@ -128,6 +128,103 @@
128
128
  "type": "string",
129
129
  "pattern": "^F:[0-9a-f]{8}$",
130
130
  "description": "Stable cross-round identity of the finding, computed by finding-fingerprint.mjs from (file, ruleId or normalized issue text). Never includes the line. On iteration >= 2 a reviewer that recognises an entry from <previous-round-findings> echoes its fingerprint; otherwise leave it unset and the script fills it in."
131
+ },
132
+ "security": {
133
+ "type": "object",
134
+ "additionalProperties": false,
135
+ "required": ["owaspCategory", "cwe", "cvss", "evidence", "confidence", "remediation"],
136
+ "description": "OPTIONAL security envelope. Present when the finding comes from the security-auditor (Phase 3 Step 3.0 merge) or /multi-agent:security-review; absent on a general code-reviewer finding. Its shape is the security block of security-finding.schema.json, kept in step by smoke-security-schema-parity.sh (self-contained, no cross-file $ref). validate-reviewer.mjs ignores it; triage carries it through unchanged so Phase 5 can report CVSS + CWE + remediation.",
137
+ "properties": {
138
+ "owaspCategory": {
139
+ "type": "string",
140
+ "pattern": "^(A(0[1-9]|10):2021|M[1-9]:2024|M10:2024)( .+)?$",
141
+ "description": "OWASP Top 10 2021 id (web/API) or OWASP Mobile Top 10 2024 id, optionally followed by a human title."
142
+ },
143
+ "cwe": {
144
+ "type": "string",
145
+ "pattern": "^CWE-[0-9]{1,5}$",
146
+ "description": "The specific CWE weakness id. One per finding."
147
+ },
148
+ "cve": {
149
+ "type": "string",
150
+ "pattern": "^CVE-[0-9]{4}-[0-9]{4,}$",
151
+ "description": "A published CVE, for a known-vulnerable dependency finding."
152
+ },
153
+ "cvss": {
154
+ "type": "object",
155
+ "additionalProperties": false,
156
+ "required": ["vector", "baseScore", "band"],
157
+ "description": "CVSS 3.1 base metrics; baseScore and band are computed from the vector by security_cvss_score.",
158
+ "properties": {
159
+ "vector": {
160
+ "type": "string",
161
+ "pattern": "^CVSS:3[.]1/AV:[NALP]/AC:[LH]/PR:[NLH]/UI:[NR]/S:[UC]/C:[NLH]/I:[NLH]/A:[NLH]$",
162
+ "description": "Full CVSS 3.1 base vector."
163
+ },
164
+ "baseScore": {
165
+ "type": "number",
166
+ "minimum": 0,
167
+ "maximum": 10,
168
+ "description": "0.0 .. 10.0 base score."
169
+ },
170
+ "band": {
171
+ "type": "string",
172
+ "enum": ["none", "low", "medium", "high", "critical"],
173
+ "description": "Qualitative band; maps to the finding severity (critical/high -> blocking, medium -> important, low/none -> suggestion)."
174
+ }
175
+ }
176
+ },
177
+ "evidence": {
178
+ "type": "string",
179
+ "minLength": 8,
180
+ "description": "What in the code proves the finding, cited by file:line."
181
+ },
182
+ "counterevidence": {
183
+ "type": "string",
184
+ "description": "What would disprove it, or the condition under which it is a false positive."
185
+ },
186
+ "confidence": {
187
+ "type": "string",
188
+ "enum": ["high", "medium", "low"],
189
+ "description": "How sure the auditor is the finding is real."
190
+ },
191
+ "confidenceRationale": {
192
+ "type": "string",
193
+ "description": "Why the confidence is what it is."
194
+ },
195
+ "severityChangeConditions": {
196
+ "type": "string",
197
+ "description": "The fact that, if learned, would move the severity."
198
+ },
199
+ "remediation": {
200
+ "type": "string",
201
+ "minLength": 8,
202
+ "description": "Remediation steps in prose."
203
+ },
204
+ "remediationDiff": {
205
+ "type": "object",
206
+ "additionalProperties": false,
207
+ "required": ["before", "after"],
208
+ "description": "The fix as a before/after pair.",
209
+ "properties": {
210
+ "before": { "type": "string" },
211
+ "after": { "type": "string" }
212
+ }
213
+ },
214
+ "endpoint": {
215
+ "type": "string",
216
+ "description": "The HTTP route or RPC method for a service-layer finding."
217
+ },
218
+ "method": {
219
+ "type": "string",
220
+ "enum": ["GET", "POST", "PUT", "PATCH", "DELETE", "HEAD", "OPTIONS"],
221
+ "description": "HTTP method for an endpoint finding."
222
+ },
223
+ "fixVerification": {
224
+ "type": "string",
225
+ "description": "How to confirm the fix worked - the test to add or check to run."
226
+ }
227
+ }
131
228
  }
132
229
  }
133
230
  }
@@ -0,0 +1,144 @@
1
+ {
2
+ "$schema": "https://json-schema.org/draft/2020-12/schema",
3
+ "$id": "https://github.com/mmerterden/multi-agent-pipeline/pipeline/schemas/security-finding.schema.json",
4
+ "version": "1.0.0",
5
+ "title": "Multi-Agent Pipeline - security finding",
6
+ "description": "Contract for one finding produced by the security-auditor subagent (Phase 3) and the /multi-agent:security-review command. A security finding IS a reviewer-output.schema.json finding - same core fields (severity, file, line, issue, fix) - carrying one extra `security` block. That is exactly why it works: the auditor emits reviewer-output objects whose findings have the block, they merge into the Phase 3 reviewer set at Step 3.0, and a blocking one blocks Phase 4 like any reviewer blocker. reviewer-output.schema.json declares the same `security` property as OPTIONAL on its finding; here it is REQUIRED, which is the whole difference between a general finding and a security finding. The severity enum is the reviewer's (blocking|important|suggestion), never a private Critical/High/Medium scale: severity is derived from security.cvss.band by the rule on the `severity` property, and a severity that contradicts the band is the one inconsistency this contract forbids. The two schemas are kept in step by smoke-security-schema-parity.sh, not by a cross-file $ref, matching the self-contained style of the other pipeline schemas.",
7
+ "type": "object",
8
+ "additionalProperties": false,
9
+ "required": ["severity", "file", "line", "issue", "fix", "security"],
10
+ "properties": {
11
+ "severity": {
12
+ "type": "string",
13
+ "enum": ["blocking", "important", "suggestion"],
14
+ "description": "The reviewer severity, derived from security.cvss.band so this finding filters through triage like any other: critical or high band -> blocking; medium -> important; low or none -> suggestion. It is not a free choice - a severity that disagrees with the band is the one inconsistency this schema exists to forbid."
15
+ },
16
+ "file": {
17
+ "type": "string",
18
+ "minLength": 1,
19
+ "description": "Path relative to repo root. The auditor must not invent a file that is not in the tree."
20
+ },
21
+ "line": {
22
+ "type": "integer",
23
+ "minimum": 0,
24
+ "description": "Line number. 0 = whole-file, configuration-level, or dependency-manifest finding."
25
+ },
26
+ "issue": {
27
+ "type": "string",
28
+ "minLength": 4,
29
+ "description": "What is wrong, in one sentence. The vulnerability, not the fix."
30
+ },
31
+ "fix": {
32
+ "type": "string",
33
+ "minLength": 4,
34
+ "description": "Concrete remediation in one line, for the reviewer-shaped view. The full before/after diff lives in security.remediationDiff."
35
+ },
36
+ "ruleId": {
37
+ "type": "string",
38
+ "minLength": 1,
39
+ "description": "Stable id of a cited rule (e.g. SEC-03) when the finding comes from a standards registry supplied to the auditor. The author can look the rule up and argue with it rather than with the auditor."
40
+ },
41
+ "criteriaSource": {
42
+ "type": "string",
43
+ "minLength": 1,
44
+ "description": "Which source the rule came from: a registry name, a compliance catalog (apple-archive-compliance / google-play-compliance), a reference path, or 'threat-model' when the finding is derived from the run-scoped threat model rather than a fixed rule."
45
+ },
46
+ "security": {
47
+ "type": "object",
48
+ "additionalProperties": false,
49
+ "required": ["owaspCategory", "cwe", "cvss", "evidence", "confidence", "remediation"],
50
+ "description": "The security envelope. Present on every security finding; optional on a general reviewer finding.",
51
+ "properties": {
52
+ "owaspCategory": {
53
+ "type": "string",
54
+ "pattern": "^(A(0[1-9]|10):2021|M[1-9]:2024|M10:2024)( .+)?$",
55
+ "description": "The OWASP category: web/API uses OWASP Top 10 2021 ids (A01:2021 .. A10:2021), mobile uses OWASP Mobile Top 10 2024 ids (M1:2024 .. M10:2024). An optional human title may follow the id, e.g. 'A01:2021 Broken Access Control'."
56
+ },
57
+ "cwe": {
58
+ "type": "string",
59
+ "pattern": "^CWE-[0-9]{1,5}$",
60
+ "description": "The specific weakness, as a validated CWE id (CWE-89, CWE-798, ...). One weakness per finding; if a line has two, it is two findings."
61
+ },
62
+ "cve": {
63
+ "type": "string",
64
+ "pattern": "^CVE-[0-9]{4}-[0-9]{4,}$",
65
+ "description": "A published CVE, when the finding is a known-vulnerable dependency rather than first-party code. Set by the dependency-audit path, not by static code review."
66
+ },
67
+ "cvss": {
68
+ "type": "object",
69
+ "additionalProperties": false,
70
+ "required": ["vector", "baseScore", "band"],
71
+ "description": "CVSS 3.1 base metrics. The auditor writes the vector; security_cvss_score (toolkit) computes baseScore and band from it, so the two cannot drift from the vector by hand.",
72
+ "properties": {
73
+ "vector": {
74
+ "type": "string",
75
+ "pattern": "^CVSS:3[.]1/AV:[NALP]/AC:[LH]/PR:[NLH]/UI:[NR]/S:[UC]/C:[NLH]/I:[NLH]/A:[NLH]$",
76
+ "description": "A full CVSS 3.1 base vector string. All eight base metrics are required; temporal and environmental metrics are out of scope for a static review."
77
+ },
78
+ "baseScore": {
79
+ "type": "number",
80
+ "minimum": 0,
81
+ "maximum": 10,
82
+ "description": "The CVSS 3.1 base score computed from the vector by security_cvss_score. 0.0 .. 10.0, one decimal."
83
+ },
84
+ "band": {
85
+ "type": "string",
86
+ "enum": ["none", "low", "medium", "high", "critical"],
87
+ "description": "The qualitative band of baseScore: none 0.0, low 0.1-3.9, medium 4.0-6.9, high 7.0-8.9, critical 9.0-10.0. This is what maps to `severity`."
88
+ }
89
+ }
90
+ },
91
+ "evidence": {
92
+ "type": "string",
93
+ "minLength": 8,
94
+ "description": "What in the code proves the finding, quoted or cited by file:line. A static finding with no evidence is a guess; this is what a reader checks before agreeing."
95
+ },
96
+ "counterevidence": {
97
+ "type": "string",
98
+ "description": "What in the code would DISPROVE the finding, or the condition under which it is a false positive - a guard elsewhere, a framework default, an unreachable path. Stating it keeps confidence honest; leaving it blank asserts there is none."
99
+ },
100
+ "confidence": {
101
+ "type": "string",
102
+ "enum": ["high", "medium", "low"],
103
+ "description": "How sure the auditor is that this is real. Low confidence does not mean do-not-report; it means report with the counterevidence and let triage weigh it."
104
+ },
105
+ "confidenceRationale": {
106
+ "type": "string",
107
+ "description": "One line on why the confidence is what it is - what was and was not verifiable from the code alone."
108
+ },
109
+ "severityChangeConditions": {
110
+ "type": "string",
111
+ "description": "The fact that, if learned, would move the severity: 'critical if this endpoint is unauthenticated in production; medium if it is admin-only'. Names the assumption the score rests on."
112
+ },
113
+ "remediation": {
114
+ "type": "string",
115
+ "minLength": 8,
116
+ "description": "The remediation steps in prose - what to change and why it closes the weakness. Actionable, not 'sanitize input'."
117
+ },
118
+ "remediationDiff": {
119
+ "type": "object",
120
+ "additionalProperties": false,
121
+ "required": ["before", "after"],
122
+ "description": "The fix as a before/after pair the author can read as a diff. Optional - some findings are configuration or process changes with no single code hunk - but preferred for first-party code.",
123
+ "properties": {
124
+ "before": { "type": "string", "description": "The vulnerable code as it stands." },
125
+ "after": { "type": "string", "description": "The same region after the fix." }
126
+ }
127
+ },
128
+ "endpoint": {
129
+ "type": "string",
130
+ "description": "The HTTP route or RPC method the finding sits on, when it is a service-layer issue. Empty for non-network findings."
131
+ },
132
+ "method": {
133
+ "type": "string",
134
+ "enum": ["GET", "POST", "PUT", "PATCH", "DELETE", "HEAD", "OPTIONS"],
135
+ "description": "The HTTP method for an endpoint finding."
136
+ },
137
+ "fixVerification": {
138
+ "type": "string",
139
+ "description": "How to confirm the fix worked - the test to add or the check to run. Static review cannot fire an exploit, so this is the empirical step a human or a later dynamic pass takes."
140
+ }
141
+ }
142
+ }
143
+ }
144
+ }
@@ -28,6 +28,7 @@ export const COMMON_SKILLS = [
28
28
  "skill-creator",
29
29
  "backlog",
30
30
  "localization-reuse-map",
31
+ "security-review",
31
32
  ];
32
33
 
33
34
  // The analyst plugin is the second always-enabled toolkit. It carries outside
@@ -16,7 +16,7 @@
16
16
  "cacheReadPerMtok": 0.5,
17
17
  "modelId": "claude-opus-5",
18
18
  "provider": "anthropic",
19
- "note": "Second tier - dev phase on a Short run, Reviewer 1 and triage on Copilot CLI, and the opus rung of the fable -> opus -> sonnet fallback ladder. Same rate as the Opus 4.8 it replaces, so the ledger needed no reprice on the generation move. Claude Opus 5 draws on a rate-limit pool SEPARATE from the combined Opus 4.x pool - moving traffic here neither frees headroom on the old bucket nor inherits it."
19
+ "note": "Second tier - Reviewer 1 and triage on Copilot CLI, and the opus rung of the fable -> opus -> sonnet fallback ladder. Same rate as the Opus 4.8 it replaces, so the ledger needed no reprice on the generation move. Claude Opus 5 draws on a rate-limit pool SEPARATE from the combined Opus 4.x pool - moving traffic here neither frees headroom on the old bucket nor inherits it."
20
20
  },
21
21
  "sonnet": {
22
22
  "inPerMtok": 3.0,
@@ -11,9 +11,9 @@
11
11
  # Three things keep this from being a foot-gun, and each is a rule rather than a
12
12
  # heuristic:
13
13
  #
14
- # 1. A run WAITING FOR YOU is never reaped. An open PR, Phase 4 or 7, or
15
- # `status: awaiting_input` means the work landed and the pipeline is
16
- # holding for an answer by design. That is finished work, not residue.
14
+ # 1. A run WAITING FOR YOU is never reaped. An open PR, the Commit or Report
15
+ # phase, or `status: awaiting_input` means the work landed and the pipeline
16
+ # is holding for an answer by design. That is finished work, not residue.
17
17
  # 2. A path that is not strictly inside `<repo>/.worktrees/` is never removed.
18
18
  # This is not theoretical: four state files on that machine record
19
19
  # `worktreePath` as the REPO ROOT, so a sweep that trusted the field would
@@ -96,6 +96,13 @@ case "$DAYS$PHASE0_DAYS" in *[!0-9]*) echo "gc-abandoned: --days and --phase0-da
96
96
  command -v jq >/dev/null 2>&1 || { echo "gc-abandoned: jq is required"; exit 0; }
97
97
  [ -d "$LOGS" ] || { echo "gc-abandoned: nothing to do (no state at $LOGS)"; exit 0; }
98
98
 
99
+ # The phase from which a stopped run is "waiting on you", not residue, is the
100
+ # contract's threshold - Commit onward - not a literal pinned to one numbering.
101
+ # runs-index.mjs groups the same way from the same field.
102
+ SELF_DIR=$(cd "$(dirname "$0")" && pwd)
103
+ WAITING_FROM=$(jq -r '.thresholds.waitingFromPhase' "$SELF_DIR/../schemas/phases.json" 2>/dev/null || echo 4)
104
+ case "$WAITING_FROM" in *[!0-9]*|"") WAITING_FROM=4 ;; esac
105
+
99
106
  # GNU form FIRST. `stat -f` is a valid GNU flag (--file-system) that succeeds and
100
107
  # prints a mount point, so BSD-first silently returns a number that is not a
101
108
  # timestamp. ADR-0012 lists this among the constructs that look cross-platform
@@ -142,7 +149,7 @@ while IFS= read -r state; do
142
149
  wt=$(jq -r '.worktreePath // ""' "$state" 2>/dev/null || true)
143
150
 
144
151
  # Rule 1: waiting for you is not residue.
145
- if [ "$status" = "awaiting_input" ] || [ -n "$pr" ] || [ "$phase" = "6" ] || [ "$phase" = "7" ]; then
152
+ if [ "$status" = "awaiting_input" ] || [ -n "$pr" ] || { [ "$phase" != "?" ] && [ "$phase" -ge "$WAITING_FROM" ] 2>/dev/null; }; then
146
153
  skipped_waiting=$((skipped_waiting + 1))
147
154
  continue
148
155
  fi
@@ -271,10 +278,10 @@ if [ "$STATE_ONLY" -eq 0 ] && [ -d "$REPOS" ]; then
271
278
  ph=$(printf '%s' "$row" | cut -f5)
272
279
 
273
280
  # Rule 1 again: landed work is not residue, whichever pass finds it. Order
274
- # matters here - the phase-4/5 test is a proxy for "still waiting", and a
275
- # run whose status already says it FINISHED is not waiting for anyone. A
276
- # complete run sitting at phase 7 was read as waiting and never reaped,
277
- # which is how three finished worktrees held 5.1 GB indefinitely.
281
+ # matters here - the waiting-phase test is a proxy for "still waiting", and
282
+ # a run whose status already says it FINISHED is not waiting for anyone. A
283
+ # complete run sitting at the report phase was read as waiting and never
284
+ # reaped, which is how three finished worktrees held 5.1 GB indefinitely.
278
285
  case "$st" in
279
286
  complete | completed | failed) ;;
280
287
  awaiting_input)
@@ -282,7 +289,7 @@ if [ "$STATE_ONLY" -eq 0 ] && [ -d "$REPOS" ]; then
282
289
  in_progress | paused)
283
290
  continue ;; # pass 1 owns these; do not report them twice
284
291
  *)
285
- if [ -n "$pr" ] || [ "$ph" = "6" ] || [ "$ph" = "7" ]; then
292
+ if [ -n "$pr" ] || { [ "$ph" != "?" ] && [ "$ph" -ge "$WAITING_FROM" ] 2>/dev/null; }; then
286
293
  skipped_waiting=$((skipped_waiting + 1)); continue
287
294
  fi ;;
288
295
  esac
@@ -4,7 +4,7 @@
4
4
  # `offload-ref.sh` parks build logs, diffs and test output under
5
5
  # <root>/.multi-agent/refs/<node_id>.md so a phase prompt can carry a pointer
6
6
  # instead of the whole log. In worktree modes that directory dies with the
7
- # worktree. In the --local modes there is no worktree: the refs land in the
7
+ # worktree. With a local workspace there is no worktree: the refs land in the
8
8
  # real checkout and nothing ever removes them. They are gitignored, so they are
9
9
  # invisible to `git status` and grow without bound.
10
10
  #
@@ -16,12 +16,9 @@ import { readFileSync } from "node:fs";
16
16
  * node pipeline/scripts/gen-mode-dispatch.mjs --mode=autopilot # 0..5
17
17
  * node pipeline/scripts/gen-mode-dispatch.mjs --mode=analysis # 0/1/3/4/5 (no Dev)
18
18
  *
19
- * v16.0.0 removed the four dev-* modes. Depth is no longer a command name: the
20
- * Phase 0 Step 7.5 picker asks Full or Short and sets `state.onlyDevelop`. That
21
- * answer arrives long after the tracker boots at Step -1, so v17.5.0 splits
22
- * registration for the two modes that ask it (`full`, `local`): Phase 0 at Step -1,
23
- * the rest once depth has named them. Everything else still registers its whole
24
- * set up front, and a generated phase set is per-COMMAND, never per-depth.
19
+ * There is one pipeline. A mode's phase set is a property of the COMMAND and is
20
+ * known before the tracker boots, so every mode registers its whole set at
21
+ * Step -1 and no registration is deferred.
25
22
  *
26
23
  * Companion smoke `smoke-mode-dispatch-drift.sh` regenerates the section for
27
24
  * each mode file, diffs against the on-disk content, and fails on drift.
@@ -63,9 +60,7 @@ const MODES = Object.fromEntries(
63
60
  name,
64
61
  {
65
62
  phases: modePhases(name),
66
- local: m.local,
67
63
  autopilot: m.autopilot,
68
- ...(m.depth ? { depth: true } : {}),
69
64
  },
70
65
  ]),
71
66
  );
@@ -86,9 +81,6 @@ const phaseSequence = activeIds.map((id) => `Phase ${id}`).join(" → ");
86
81
  const MODE_LABELS = {
87
82
  autopilot: "`autopilot`",
88
83
  full: "full-pipeline",
89
- local: "`--local`",
90
- "local-autopilot": "`--local autopilot`",
91
- "full-local": "full-pipeline + `--local`",
92
84
  };
93
85
  const modeLabel = MODE_LABELS[MODE] ?? `\`${MODE}\``;
94
86
 
@@ -98,46 +90,24 @@ const skipNote =
98
90
  ? `${modeLabel} mode does NOT TaskCreate phases ${skippedIds.join("/")} - those are not part of the ${modeLabel} phase set (\`${spec.phases.join(" ")}\`). Only register tiles for the active set.`
99
91
  : `${modeLabel} mode TaskCreates all ${PHASES.length} phases (no phase is skipped).`;
100
92
 
101
- const orderingNote = `**All TaskCreate calls in a batch fire in strict phase-number order BEFORE any TaskUpdate is applied.** For ${modeLabel} that means: ${spec.depth ? `Phase 0 at Step -1, then the rest in ascending order at Step 7.5 (${phaseSequence} minus whatever the depth answer drops)` : phaseSequence}. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks. Full ordering contract in \`$HOME/.claude/multi-agent-refs/tracker-contract.md\` section "TaskCreate ordering (strict)".`;
102
-
103
- const localCaveat = spec.local
104
- ? `\n> **Local mode:** no worktree is created, work happens on the current branch. Phase 0 Init still calls \`init\` - the \`--local\` flag is stored in tracker-state.json, and \`:resume\` restores the correct CWD.\n`
105
- : "";
93
+ const orderingNote = `**All TaskCreate calls in a batch fire in strict phase-number order BEFORE any TaskUpdate is applied.** For ${modeLabel} that means: ${phaseSequence}. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks. Full ordering contract in \`$HOME/.claude/multi-agent-refs/tracker-contract.md\` section "TaskCreate ordering (strict)".`;
106
94
 
107
95
  const banner = spec.autopilot
108
96
  ? `\n> **Autopilot mode:** user confirmations are skipped. The tracker is still mandatory - autopilot agent calls cannot skip it; skipping breaks \`smoke-tracker-contract.sh\`.\n`
109
97
  : "";
110
98
 
111
- // Registration shape. A mode that asks the depth question does not know its phase
112
- // set at Step -1 (the answer needs taskType, which needs the fetched issue and the
113
- // branch), so it registers Phase 0 there and the rest at Step 7.5. Drawing eight
114
- // tiles beside the question that decides whether two of them run is the failure this
115
- // split exists to remove.
116
- const fullSet = spec.phases.filter((p) => p !== "0:Init");
117
- const shortSet = fullSet.filter((p) => Number(p.split(":")[0]) >= 2);
99
+ // Registration shape. Every mode knows its whole phase set at Step -1, so every
100
+ // tile is created there and nothing is appended later.
118
101
  const loopFor = (list) =>
119
102
  `for p in ${list.map((x) => `"${x}"`).join(" ")}; do\n bash $HOME/.claude/scripts/phase-tracker.sh add "\${p%%:*}" "\${p#*:}"\ndone`;
120
103
 
121
- const initLoop = spec.depth
122
- ? 'bash $HOME/.claude/scripts/phase-tracker.sh add 0 "Init"\nbash $HOME/.claude/scripts/phase-tracker.sh tiles'
123
- : `${loopFor(spec.phases)}`;
124
-
125
- const deferredBlock = spec.depth
126
- ? `
127
- # Phase 0 Step 7.5, immediately after the depth answer - the first moment this
128
- # mode knows its phase set. Full:
129
- ${loopFor(fullSet)}
130
- # Short (Analysis and Planning are not run, so they get no tile at all):
131
- ${loopFor(shortSet)}
132
- # Then the widget, narrowed to the phases that do not have a tile yet:
133
- bash $HOME/.claude/scripts/phase-tracker.sh tiles --new
134
- `
135
- : "";
104
+ const initLoop = loopFor(spec.phases);
105
+ const deferredBlock = "";
136
106
 
137
107
  const out = `## Required: Phase Tracker Contract
138
108
 
139
109
  **The phase tracker is mandatory** - the agent cannot skip it. Full spec: [\`$HOME/.claude/multi-agent-refs/tracker-contract.md\`]($HOME/.claude/multi-agent-refs/tracker-contract.md).
140
- ${banner}${localCaveat}
110
+ ${banner}
141
111
  Two channels run in parallel at every phase boundary:
142
112
 
143
113
  1. **State channel** (every CLI, identical): \`phase-tracker.sh\` writes to \`tracker-state.json\`. Drives \`:resume\`, \`:log\`, \`:status\`.
@@ -160,11 +130,11 @@ bash $HOME/.claude/scripts/phase-tracker.sh tokens <N> <in> <out> [cached]
160
130
 
161
131
  In Claude Code the agent MUST also drive the native TaskList widget so the user sees a sticky phase tile stack - this is the only progress signal Claude Code surfaces. Skipping these calls is the #1 source of "I don't see any phases" complaints.
162
132
 
163
- **TaskCreate ordering (strict)**: All TaskCreate calls in a registration batch fire in strict phase-number order BEFORE any TaskUpdate in that batch, and a later batch only ever appends phases numbered above everything already registered.${spec.depth ? " This mode registers in two batches (Step -1, then Step 7.5), so `tiles --new` narrows the second one and the Phase 0 tile is never created twice." : ""} The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks (e.g. \`1 ✓ · 2 ✓ · 4 ✓ · 0 ▶ · 3 ☐\`) even when the underlying state is correct. Pre-marking phases as completed/skipped before Phase 0 starts is FORBIDDEN - register the tile in order, then flip status via TaskUpdate when the phase actually short-circuits. Full contract in \`$HOME/.claude/multi-agent-refs/tracker-contract.md\` section "TaskCreate ordering (strict)".
133
+ **TaskCreate ordering (strict)**: All TaskCreate calls in a registration batch fire in strict phase-number order BEFORE any TaskUpdate in that batch, and a later batch only ever appends phases numbered above everything already registered. The native widget renders by creation order, not by phase number - out-of-order calls produce visually scrambled tile stacks (e.g. \`1 ✓ · 2 ✓ · 4 ✓ · 0 ▶ · 3 ☐\`) even when the underlying state is correct. Pre-marking phases as completed/skipped before Phase 0 starts is FORBIDDEN - register the tile in order, then flip status via TaskUpdate when the phase actually short-circuits. Full contract in \`$HOME/.claude/multi-agent-refs/tracker-contract.md\` section "TaskCreate ordering (strict)".
164
134
 
165
135
  \`\`\`text
166
136
  # Register one tile per phase, capture the taskId, persist it:
167
- for each phase in ${spec.depth ? `0:Init at Step -1, then ${fullSet.join(", ")} (Full) or ${shortSet.join(", ")} (Short) at Step 7.5` : spec.phases.join(", ")}:
137
+ for each phase in ${spec.phases.join(", ")}:
168
138
  TaskCreate({ subject: "Phase <N>: <Name>", activeForm: "<doing-form>" })
169
139
  -> returns taskId
170
140
  bash $HOME/.claude/scripts/phase-tracker.sh meta <N> tasklist_id "<taskId>"