@mmerterden/multi-agent-pipeline 17.1.0 → 17.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (74) hide show
  1. package/CHANGELOG.md +194 -0
  2. package/README.md +7 -0
  3. package/README.tr.md +7 -0
  4. package/docs/token-budget-history.md +22 -0
  5. package/install/_dev-only-files.mjs +1 -0
  6. package/install/codex.mjs +18 -1
  7. package/install/copilot.mjs +17 -1
  8. package/package.json +1 -1
  9. package/pipeline/agents/code-reviewer.md +35 -1
  10. package/pipeline/commands/multi-agent/autopilot-on/SKILL.md +9 -1
  11. package/pipeline/lib/autopilot-state.sh +34 -0
  12. package/pipeline/multi-agent-refs/_dev-context.md +10 -0
  13. package/pipeline/multi-agent-refs/analysis/locked.md +4 -4
  14. package/pipeline/multi-agent-refs/analysis/redesign.md +8 -0
  15. package/pipeline/multi-agent-refs/analysis/render.md +2 -1
  16. package/pipeline/multi-agent-refs/analysis/review.md +9 -0
  17. package/pipeline/multi-agent-refs/android-guide.md +14 -0
  18. package/pipeline/multi-agent-refs/audit-guide.md +12 -0
  19. package/pipeline/multi-agent-refs/backend-guide.md +10 -0
  20. package/pipeline/multi-agent-refs/channels/confluence.md +11 -0
  21. package/pipeline/multi-agent-refs/channels/issue-comment.md +12 -0
  22. package/pipeline/multi-agent-refs/channels/jira.md +10 -0
  23. package/pipeline/multi-agent-refs/channels/pr-review-actions.md +13 -0
  24. package/pipeline/multi-agent-refs/channels/pr.md +37 -0
  25. package/pipeline/multi-agent-refs/component-dispatch.md +11 -0
  26. package/pipeline/multi-agent-refs/component-generation.md +11 -0
  27. package/pipeline/multi-agent-refs/conventions-defaults.md +15 -0
  28. package/pipeline/multi-agent-refs/cross-cli-contract.md +65 -0
  29. package/pipeline/multi-agent-refs/features/analysis-jira.md +11 -0
  30. package/pipeline/multi-agent-refs/features/code-graph.md +40 -0
  31. package/pipeline/multi-agent-refs/features/design-conformance.md +24 -0
  32. package/pipeline/multi-agent-refs/features/doctor.md +10 -0
  33. package/pipeline/multi-agent-refs/features/external-context-injection.md +7 -0
  34. package/pipeline/multi-agent-refs/features/jira-context.md +9 -0
  35. package/pipeline/multi-agent-refs/features/model-fallback.md +10 -0
  36. package/pipeline/multi-agent-refs/features/review-file-set.md +132 -0
  37. package/pipeline/multi-agent-refs/features/skill-conformance.md +13 -0
  38. package/pipeline/multi-agent-refs/features/url-enrichment.md +9 -0
  39. package/pipeline/multi-agent-refs/features/visual-evidence.md +42 -0
  40. package/pipeline/multi-agent-refs/generate-issue.md +7 -0
  41. package/pipeline/multi-agent-refs/issue-jira-triad.md +9 -0
  42. package/pipeline/multi-agent-refs/knowledge.md +6 -0
  43. package/pipeline/multi-agent-refs/multi-repo-integration-build.md +13 -0
  44. package/pipeline/multi-agent-refs/phases/modes.md +7 -0
  45. package/pipeline/multi-agent-refs/phases/operations.md +9 -0
  46. package/pipeline/multi-agent-refs/phases/phase-0-init.md +1 -1
  47. package/pipeline/multi-agent-refs/phases/phase-4-review.md +31 -23
  48. package/pipeline/multi-agent-refs/phases.md +11 -0
  49. package/pipeline/multi-agent-refs/picker-contract.md +12 -0
  50. package/pipeline/multi-agent-refs/platform-parity.md +10 -0
  51. package/pipeline/multi-agent-refs/progress-contract.md +10 -0
  52. package/pipeline/multi-agent-refs/setup/firebase.md +9 -0
  53. package/pipeline/multi-agent-refs/swiftui-guide.md +17 -0
  54. package/pipeline/multi-agent-refs/tracker-contract.md +12 -0
  55. package/pipeline/multi-agent-refs/web-guide.md +10 -0
  56. package/pipeline/multi-agent-refs/wiki-capture.md +11 -0
  57. package/pipeline/schemas/prefs.schema.json +4 -0
  58. package/pipeline/schemas/review-file-exclusions.json +137 -0
  59. package/pipeline/schemas/reviewer-output.schema.json +27 -1
  60. package/pipeline/schemas/token-budget.json +10 -19
  61. package/pipeline/scripts/autopilot-intake.mjs +5 -1
  62. package/pipeline/scripts/autopilot-runner.mjs +6 -1
  63. package/pipeline/scripts/autopilot-status.sh +3 -2
  64. package/pipeline/scripts/capture-evidence.sh +79 -11
  65. package/pipeline/scripts/diff-risk-score.mjs +1 -36
  66. package/pipeline/scripts/gen-ref-toc.mjs +279 -0
  67. package/pipeline/scripts/git-path.mjs +63 -0
  68. package/pipeline/scripts/glob-match.mjs +62 -0
  69. package/pipeline/scripts/graph-mermaid.mjs +251 -0
  70. package/pipeline/scripts/review-file-filter.mjs +180 -0
  71. package/pipeline/scripts/skill-conformance.mjs +1 -31
  72. package/pipeline/scripts/validate-analysis-doc.mjs +53 -0
  73. package/pipeline/scripts/validate-reviewer.mjs +90 -1
  74. package/pipeline/scripts/verify-citations.mjs +346 -0
@@ -1,5 +1,19 @@
1
1
  ## Android/Kotlin Component Generation Guide
2
2
 
3
+ <!-- toc -->
4
+ - [Component Architecture: State / Screen / Content](#component-architecture-state-screen-content)
5
+ - [Simple vs Complex Decision](#simple-vs-complex-decision)
6
+ - [State Pattern](#state-pattern)
7
+ - [Token Discipline](#token-discipline)
8
+ - [Stability for Performance](#stability-for-performance)
9
+ - [Accessibility](#accessibility)
10
+ - [Preview Best Practices](#preview-best-practices)
11
+ - [Testing](#testing)
12
+ - [Build Verification](#build-verification)
13
+ - [Component Quality Checklist](#component-quality-checklist)
14
+ - [Compliance Rules (maps to multi-agent-toolkit MCP audit tools)](#compliance-rules-maps-to-multi-agent-toolkit-mcp-audit-tools)
15
+ <!-- /toc -->
16
+
3
17
  > **MUST: Figma MCP-first (BLOCKING).** If the task references any Figma frame (URL, node ID, or "from the design"), the Dev phase MUST call `mcp__claude_ai_Figma__get_design_context` for every frame BEFORE writing a single Composable line. Use the `CodeConnectSnippet` component name verbatim - no sound-alike substitutions. Authentication failure is not a skip path. Full rule, trigger conditions, and gate failure modes: `$HOME/.claude/rules/figma-pipeline.md` "MUST: Figma MCP-first (BLOCKING)". Phase wiring: `$HOME/.claude/multi-agent-refs/phases/phase-3-dev.md` "MUST: Figma MCP-first (BLOCKING pre-step)".
4
18
 
5
19
  When the task involves creating an Android UI component (Jetpack Compose), follow this architecture.
@@ -1,5 +1,17 @@
1
1
  ## Audit & Quality Tools Guide
2
2
 
3
+ <!-- toc -->
4
+ - [Trigger Model](#trigger-model)
5
+ - [iOS Accessibility Audit](#ios-accessibility-audit)
6
+ - [Android Accessibility Audit](#android-accessibility-audit)
7
+ - [iOS Biometric Test](#ios-biometric-test)
8
+ - [Android Launch Time](#android-launch-time)
9
+ - [iOS Archive Audit (App Store Compliance)](#ios-archive-audit-app-store-compliance)
10
+ - [Android APK Audit (Play Store Compliance)](#android-apk-audit-play-store-compliance)
11
+ - [Integration with Pipeline Phases](#integration-with-pipeline-phases)
12
+ - [Graceful Degradation](#graceful-degradation)
13
+ <!-- /toc -->
14
+
3
15
  Standalone audit commands - runs directly via Bash, **no MCP server dependency**. These are the same checks that multi-agent-toolkit-mcp provides as MCP tools, but embedded here as pipeline skills.
4
16
 
5
17
  ### Trigger Model
@@ -1,5 +1,15 @@
1
1
  ## Backend API Development Guide
2
2
 
3
+ <!-- toc -->
4
+ - [API Architecture](#api-architecture)
5
+ - [Python/FastAPI Pattern](#pythonfastapi-pattern)
6
+ - [Node.js/Express Pattern](#nodejsexpress-pattern)
7
+ - [Error Handling](#error-handling)
8
+ - [Security Checklist](#security-checklist)
9
+ - [Testing](#testing)
10
+ - [Quality Checklist](#quality-checklist)
11
+ <!-- /toc -->
12
+
3
13
  When the task involves backend development (Python/FastAPI, Node.js/Express, Go), follow these patterns.
4
14
 
5
15
  ### API Architecture
@@ -1,5 +1,16 @@
1
1
  # Channel adapter - Confluence page
2
2
 
3
+ <!-- toc -->
4
+ - [Required body structure](#required-body-structure)
5
+ - [Token check](#token-check)
6
+ - [Parent page resolution](#parent-page-resolution)
7
+ - [Page title](#page-title)
8
+ - [Body conversion (markdown → storage format)](#body-conversion-markdown-storage-format)
9
+ - [POST contract](#post-contract)
10
+ - [Recents persistence](#recents-persistence)
11
+ - [Hard rules (must not regress)](#hard-rules-must-not-regress)
12
+ <!-- /toc -->
13
+
3
14
  > Detailed contract for the `confluence` channel of `/multi-agent:channels`. Split out of `channels.md` in v8.0.0; the parent doc keeps a one-line summary and a link here.
4
15
 
5
16
  The Confluence adapter creates or updates a Confluence page under a chosen parent. Like the Jira adapter, it runs **once** per invocation - the page lives at the primary repo's component slug in multi-repo mode.
@@ -1,5 +1,17 @@
1
1
  # Channel adapter - GitHub Issue comment
2
2
 
3
+ <!-- toc -->
4
+ - [When this fires](#when-this-fires)
5
+ - [The hard rule](#the-hard-rule)
6
+ - [Required body structure](#required-body-structure)
7
+ - [Hard prohibitions](#hard-prohibitions)
8
+ - [Compaction policy](#compaction-policy)
9
+ - [API contract](#api-contract)
10
+ - [Pairing with the Progress flag updater](#pairing-with-the-progress-flag-updater)
11
+ - [Drift detection](#drift-detection)
12
+ - [Hard rules (must not regress)](#hard-rules-must-not-regress)
13
+ <!-- /toc -->
14
+
3
15
  > Canonical template for the `issue` channel of `/multi-agent:channels`. Every successful run that touched a tracked GitHub issue MUST post one comment using this template - no exceptions, no "state-only" shortcuts.
4
16
 
5
17
  ## When this fires
@@ -1,5 +1,15 @@
1
1
  # Channel adapter - Jira comment
2
2
 
3
+ <!-- toc -->
4
+ - [Required body structure](#required-body-structure)
5
+ - [Wiki markup conversion](#wiki-markup-conversion)
6
+ - [Cross-link injection](#cross-link-injection)
7
+ - [Token resolution](#token-resolution)
8
+ - [POST contract](#post-contract)
9
+ - [Wiki → Jira auto-link triad](#wiki-jira-auto-link-triad)
10
+ - [Hard rules (must not regress)](#hard-rules-must-not-regress)
11
+ <!-- /toc -->
12
+
3
13
  > Detailed contract for the `jira` channel of `/multi-agent:channels`. Split out of `channels.md` in v8.0.0; the parent doc keeps a one-line summary and a link here.
4
14
 
5
15
  The Jira adapter posts a comment on the linked issue. The Jira ticket is shared by all repos in a multi-repo task, so the adapter runs **once** per channels invocation regardless of how many PR targets are dispatched in parallel.
@@ -1,5 +1,18 @@
1
1
  # Channel adapter - Pull Request review actions
2
2
 
3
+ <!-- toc -->
4
+ - [When this fires](#when-this-fires)
5
+ - [The decision rule](#the-decision-rule)
6
+ - [Inline comment template (per finding)](#inline-comment-template-per-finding)
7
+ - [Decision endpoints](#decision-endpoints)
8
+ - [Order of operations](#order-of-operations)
9
+ - [Hard prohibitions](#hard-prohibitions)
10
+ - [Idempotency](#idempotency)
11
+ - [Provider auto-detection](#provider-auto-detection)
12
+ - [Pairing with the chat summary](#pairing-with-the-chat-summary)
13
+ - [Drift detection](#drift-detection)
14
+ <!-- /toc -->
15
+
3
16
  > Canonical contract for the PR-actions branch of `/multi-agent:review`. The PR-side artifacts are **per-finding inline comments + an explicit approve / needs-work decision** - never a single monolithic description comment.
4
17
 
5
18
  ## When this fires
@@ -1,5 +1,16 @@
1
1
  # Channel adapter - PR description
2
2
 
3
+ <!-- toc -->
4
+ - [Required body structure](#required-body-structure)
5
+ - [Markup dialect (per surface, not per pipeline)](#markup-dialect-per-surface-not-per-pipeline)
6
+ - [Behaviour by remote](#behaviour-by-remote)
7
+ - [Reviewer-preserving Bitbucket payload (required)](#reviewer-preserving-bitbucket-payload-required)
8
+ - [Multi-repo cross-links](#multi-repo-cross-links)
9
+ - [Version mismatch handling](#version-mismatch-handling)
10
+ - [Flags that affect this adapter](#flags-that-affect-this-adapter)
11
+ - [Hard rules (must not regress)](#hard-rules-must-not-regress)
12
+ <!-- /toc -->
13
+
3
14
  > Detailed contract for the `pr` channel of `/multi-agent:channels`. Split out of `channels.md` in v8.0.0; the parent doc keeps a one-line summary and a link here.
4
15
 
5
16
  The PR adapter rewrites the pull request description with the body assembled in Step 5 of `channels.md`. Default behaviour is **replace**; `--append` opt-in preserves existing content.
@@ -66,6 +77,32 @@ Across stacks the same shape produces, for example: `LoginView.swift - ...` (i
66
77
  <none, or which service, contract or channel>
67
78
  ```
68
79
 
80
+ Part 3 is the one part of this body that is MEASURED rather than recalled. The
81
+ code graph already answers it, so draw the answer instead of re-typing it:
82
+
83
+ ```bash
84
+ node "$HOME/.claude/scripts/graph-mermaid.mjs" "<changed symbol[,symbol]>"
85
+ ```
86
+
87
+ Append the fenced block it prints under part 3, above the prose. GitHub renders
88
+ mermaid natively in pull requests, so this costs no renderer and no plugin. The
89
+ prose stays: the diagram says which symbols the change reaches, the sentence says
90
+ which screens and flows a tester must open, and neither answers the other.
91
+
92
+ Exit 1 means the repo has no graph yet (`/multi-agent:graph` builds it) or the
93
+ symbol is not in it. That is a gap with a reason, not a failure: write the prose
94
+ alone and say the graph was unavailable. Never hand-draw the diagram - a drawn
95
+ blast radius nobody measured is worse than none, because a diagram is read as
96
+ fact.
97
+
98
+ The commit line the script prints stays with it. A graph built before the change
99
+ draws the radius of an older tree, and the reader has no other way to notice.
100
+
101
+ This is a GitHub-only section. `channels/jira.md` has no mermaid handling at all:
102
+ a fence there converts to a literal `{code:mermaid}` block, so the Jira impact
103
+ section keeps its prose. Confluence renders it through the `ac:name="mermaid"`
104
+ macro (`md2confluence-v3.py`) when the space has the plugin.
105
+
69
106
  When the change deliberately fixes part of a wider problem, a closing **Risk and remaining scope** paragraph names what is still open and why it was left - a reviewer who can see the rest of the pattern in the repo will ask otherwise, and the honest answer is cheaper written down than defended in a thread.
70
107
 
71
108
  **`test_scenarios`** - the same titled-scenario shape the Jira adapter uses, so the tester reads one list on both surfaces, with symbols allowed here:
@@ -1,5 +1,16 @@
1
1
  # Component Dispatch (Phase 3 short-circuit)
2
2
 
3
+ <!-- toc -->
4
+ - [Entry conditions](#entry-conditions)
5
+ - [Plugin skill resolution](#plugin-skill-resolution)
6
+ - [Dispatch call](#dispatch-call)
7
+ - [Subphase contract (dispatch-layer owned)](#subphase-contract-dispatch-layer-owned)
8
+ - [Multi-repo report](#multi-repo-report)
9
+ - [Failure + resume](#failure-resume)
10
+ - [Short-run behaviour](#short-run-behaviour)
11
+ - [Cross-CLI behaviour (intentional divergence)](#cross-cli-behaviour-intentional-divergence)
12
+ <!-- /toc -->
13
+
3
14
  > **TLDR** - When `taskType === "component"` (Figma URL in task description or instruction-driven figma workflow), multi-agent Phase 3 **does not run the TDD loop**. It delegates the entire phase to the enabled `ai-<platform>-toolkit` **marketplace plugin's** component skill (`create-component`, falling back to `create-ui-component`) via the Skill tool. Implementation lives in the plugin; multi-agent's job is classification, dispatch, and state report. The pipeline no longer bundles its own `figma-to-component` orchestrator - component skills live in one place, the plugin marketplace.
4
15
 
5
16
  This doc is referenced from `$HOME/.claude/multi-agent-refs/phases/phase-3-dev.md`. Keeping it separate lets `phase-3-dev.md` remain tight (it's already the largest phase doc) and gives the orchestrator-report contract a stable URL for both Claude-side and Copilot-side implementations.
@@ -1,5 +1,16 @@
1
1
  # Component Generation Guide (generic)
2
2
 
3
+ <!-- toc -->
4
+ - [Component Architecture: Configuration / View / Modifiers](#component-architecture-configuration-view-modifiers)
5
+ - [Configuration Purity Rule](#configuration-purity-rule)
6
+ - [View Implementation](#view-implementation)
7
+ - [Modifier Pattern](#modifier-pattern)
8
+ - [Simple vs Complex Decision](#simple-vs-complex-decision)
9
+ - [3-Layer Test Strategy](#3-layer-test-strategy)
10
+ - [Component Checklist (Before Commit)](#component-checklist-before-commit)
11
+ - [When a Figma URL Is Provided](#when-a-figma-url-is-provided)
12
+ <!-- /toc -->
13
+
3
14
  > Lifted out of `core/multi-agent/SKILL.md`, where it was loaded on every
4
15
  > run of every mode. It applies only to a task that generates a UI
5
16
  > component from a design, so it now loads when that path is taken.
@@ -4,6 +4,21 @@ description: "Convention fallback defaults for /multi-agent:analysis Phase 2b Pa
4
4
 
5
5
  # Convention Defaults - Pass B Fallback Reference
6
6
 
7
+ <!-- toc -->
8
+ - [How the fallback chain works](#how-the-fallback-chain-works)
9
+ - [C1 - Folder Structure](#c1---folder-structure)
10
+ - [C2 - Class Naming](#c2---class-naming)
11
+ - [C3 - UI State Model](#c3---ui-state-model)
12
+ - [C4 - Test Method Naming](#c4---test-method-naming)
13
+ - [C5 - Accessibility Identifier](#c5---accessibility-identifier)
14
+ - [C6 - Localization Key](#c6---localization-key)
15
+ - [C8 - SwiftUI Preview macro (iOS only)](#c8---swiftui-preview-macro-ios-only)
16
+ - [C7 - Dependency Injection](#c7---dependency-injection)
17
+ - [Risk row template (Section 20)](#risk-row-template-section-20)
18
+ - [Maintenance](#maintenance)
19
+ - [Locked decisions that govern this file](#locked-decisions-that-govern-this-file)
20
+ <!-- /toc -->
21
+
7
22
  `/multi-agent:analysis` Phase 1c extracts conventions from each selected repo (folder structure, class naming, state model, test naming, accessibility identifier, localization key, DI registration). When `confidence == "none"` AND the standards binding source (`evidence.standards[]`) does not provide an explicit rule, Pass B falls back to the platform defaults catalogued here. Every applied default emits a row in Section 20 Risks of the rendered document.
8
23
 
9
24
  > **Language**: This file is read as a system prompt. Prose stays English. Examples carry generic placeholder names (`Foo`, `Bar`); the runtime substitutes the actual feature slug.
@@ -1,5 +1,18 @@
1
1
  # Cross-CLI Contract (Claude Code · Copilot CLI · Codex CLI)
2
2
 
3
+ <!-- toc -->
4
+ - [1. Command Inventory (56 commands)](#1-command-inventory-56-commands)
5
+ - [2. Canonical Placeholder Vocabulary](#2-canonical-placeholder-vocabulary)
6
+ - [2.6 Intentional structural divergence - thin dispatcher vs inlined orchestrator](#26-intentional-structural-divergence---thin-dispatcher-vs-inlined-orchestrator)
7
+ - [3. Frontmatter Transform Rules (Claude ↔ Copilot)](#3-frontmatter-transform-rules-claude-copilot)
8
+ - [4. Progress Signalling Parity](#4-progress-signalling-parity)
9
+ - [5. Argument Parsing Invariants](#5-argument-parsing-invariants)
10
+ - [6. Output Format Expectations](#6-output-format-expectations)
11
+ - [7. Platform Guards (macOS)](#7-platform-guards-macos)
12
+ - [8. Enforcement](#8-enforcement)
13
+ - [9. Change Control](#9-change-control)
14
+ <!-- /toc -->
15
+
3
16
  > **Non-negotiable**. Any change that breaks this contract blocks merge. Validated by `smoke-cross-cli-behavior.sh`.
4
17
 
5
18
  **Purpose**: every pipeline command must produce identical artifacts (state, logs, outputs) and respect identical placeholder vocabulary regardless of which of the three host CLIs invokes it. This file is the source of truth for "what must stay the same."
@@ -201,6 +214,28 @@ on skill directories would demand exactly the layout that breaks it.
201
214
 
202
215
  Future changes that break an item in the "stay identical" list must update **both** files in the same commit. `smoke-cross-cli-behavior.sh` enforces the identity-preserving axis (input parsing, routing, output shape); structural differences are left to manual review because enforcing them would require forcing the files to the same shape, which we intentionally don't want.
203
216
 
217
+ ### Panel diversity per host
218
+
219
+ Phase 4 runs three reviewers everywhere, but the diversity those three buy is not the
220
+ same on every host. Copilot CLI gets cross-VENDOR disagreement for free: GPT-5.4 sits
221
+ beside two Claude models. Claude Code and Codex each run a one-vendor panel - three
222
+ Anthropic models on one, three OpenAI models on the other - so the same three-way
223
+ agreement is weaker evidence there, and Phase 4 says so in the triage note on a
224
+ borderline finding.
225
+
226
+ Where the budget goes instead, when vendor diversity is unavailable:
227
+
228
+ | Host | Reviewer 1 | Reviewer 2 | Reviewer 3 |
229
+ |---|---|---|---|
230
+ | Copilot CLI | Fable/Opus, security + architecture | GPT-5.4, edge cases (cross-vendor) | Sonnet, quality |
231
+ | Claude Code | Fable, security + architecture | Opus, edge cases | Sonnet, quality |
232
+ | Codex | `xhigh`, security + architecture | a different family member, edge cases | `medium`, quality |
233
+
234
+ On Codex the axis is reasoning effort as much as model identity, because the family
235
+ members available there are closer to each other than Fable and Sonnet are. That is a
236
+ weaker axis, not an equivalent one, and treating it as equivalent is the error this
237
+ section exists to prevent.
238
+
204
239
  ## 3. Frontmatter Transform Rules (Claude ↔ Copilot)
205
240
 
206
241
  Each file has a different frontmatter schema. The sync flow transforms between them:
@@ -264,6 +299,35 @@ Every phase boundary MUST call EITHER TaskCreate/Update (Claude) OR `phase-track
264
299
 
265
300
  `phase-banner.sh` runs on both CLIs. Same output format, no Claude-only or Copilot-only flair.
266
301
 
302
+ ### 4.4 Continuous mode
303
+
304
+ `autopilot-on`, `autopilot-off` and `autopilot-status` ship to all three hosts and
305
+ behave identically there, because the thing they control is not a CLI feature:
306
+ it is a launchd user agent whose tick spawns a `claude --bg` child regardless of
307
+ which CLI you typed the command in. Continuous mode therefore requires the
308
+ `claude` binary on `PATH` on every host, Copilot and Codex included, and there is
309
+ no Copilot or Codex equivalent of the child.
310
+
311
+ | Concept | Claude Code | Copilot CLI | Codex CLI |
312
+ |---|---|---|---|
313
+ | Turn the mode on | `/multi-agent:autopilot-on` | `/multi-agent-autopilot-on` | `multi-agent autopilot-on` via the router skill |
314
+ | Scripts, libs, templates | `~/.claude/{scripts,lib,templates}` | `~/.copilot/...` | `~/.codex/...` |
315
+ | Which one is read | `ma_ap_asset <rel>` resolves the caller's own tree first, then the other two | same | same |
316
+ | Queue and config state | `~/.claude/autopilot/` | `~/.claude/autopilot/` | `~/.claude/autopilot/` |
317
+ | In-flight rows on the widget | `autopilot-status --subjects` into `TaskUpdate` | into the reprinted card | into `update_plan` |
318
+
319
+ The state row is the one that looks wrong and is not. `~/.claude/autopilot/` is
320
+ shared across hosts on purpose, exactly like `logs/`, `knowledge/` and
321
+ `multi-agent-preferences.json` (2.6, "shared state is deliberately NOT
322
+ retargeted"): two CLIs on one machine must read ONE queue. A per-host state root
323
+ would give a Copilot session a second, invisible queue and the same ticket would
324
+ be taken twice.
325
+
326
+ Everything that is NOT state resolves per host, and that half had to be fixed:
327
+ `templates/` was laid down only by `install/claude.mjs`, so `autopilot-on` on a
328
+ Codex-only machine rendered a launchd job from a file the host did not have.
329
+ `smoke-autopilot-hosts.sh` holds the line.
330
+
267
331
  ---
268
332
 
269
333
  ## 5. Argument Parsing Invariants
@@ -338,6 +402,7 @@ installs on.
338
402
 
339
403
  This contract is validated by:
340
404
 
405
+ - `smoke-autopilot-hosts.sh` - asserts continuous mode resolves its scripts, libs and plist template into whichever host tree is installed, and that the state root is NOT retargeted (4.4)
341
406
  - `smoke-cross-cli-behavior.sh` - asserts every command behaves identically, pulls from Section 2 (placeholder vocab), Section 5 (argument parsing), Section 6 (output formats); also regression-locks the 8-persona agent deployment
342
407
  - `smoke-commands-skills-parity.sh` (two assertions per command) - enforces colon-form command ↔ dash-form skill directory parity
343
408
  - `smoke-compliance-skills.sh` - enforces store-compliance skill catalog + 4 consumer wiring
@@ -1,5 +1,16 @@
1
1
  # analysis-jira - an analysis document, read as work
2
2
 
3
+ <!-- toc -->
4
+ - [The marker gate runs first, before any network call](#the-marker-gate-runs-first-before-any-network-call)
5
+ - [Coverage is two-way, and the second direction is the useful one](#coverage-is-two-way-and-the-second-direction-is-the-useful-one)
6
+ - [An unverifiable run is allowed; looking verified is not](#an-unverifiable-run-is-allowed-looking-verified-is-not)
7
+ - [Identity is a label, not a title](#identity-is-a-label-not-a-title)
8
+ - [The write is ledgered](#the-write-is-ledgered)
9
+ - [An existing node is skipped, never updated](#an-existing-node-is-skipped-never-updated)
10
+ - [Every site-specific name is a VALUE, never a schema key](#every-site-specific-name-is-a-value-never-a-schema-key)
11
+ - [Auth](#auth)
12
+ <!-- /toc -->
13
+
3
14
  `/multi-agent:analysis-jira` turns a rendered analysis document into a Jira tree.
4
15
  Two files do it, and the split is the design:
5
16
 
@@ -1,5 +1,11 @@
1
1
  ## Code Graph (Phase 1 Step 2.6 + Phase 7 Step 3)
2
2
 
3
+ <!-- toc -->
4
+ - [Phase 1 Step 2.6 - query before dispatching Explore](#phase-1-step-26---query-before-dispatching-explore)
5
+ - [Phase 7 Step 3 - refresh after the branch changed code](#phase-7-step-3---refresh-after-the-branch-changed-code)
6
+ - [The graph is drawable, and one place already asks for it](#the-graph-is-drawable-and-one-place-already-asks-for-it)
7
+ <!-- /toc -->
8
+
3
9
  A deterministic, LLM-free map of what a repo declares and what refers to what,
4
10
  written to `~/.claude/knowledge/<project>/code-graph.json`. Gated by
5
11
  `prefs.global.codeGraph.enabled` (default `false`); with it off, Phase 1 and
@@ -67,3 +73,37 @@ The rebuild costs no API tokens, so it runs every task rather than on a stalenes
67
73
  heuristic. A non-zero validator exit keeps the previous graph and logs
68
74
  `knowledge.graph_invalid`; it never fails the run - a stale graph is a degraded
69
75
  Phase 1, not a broken deliverable.
76
+
77
+ ### The graph is drawable, and one place already asks for it
78
+
79
+ The PR body's Impact Analysis, part 3, asks which symbols and files a change
80
+ reaches. That is `graph-affected.mjs`'s question, and until now the answer was
81
+ re-typed as prose by a model while the measurement sat on disk unread.
82
+
83
+ ```bash
84
+ node $HOME/.claude/scripts/graph-mermaid.mjs "<symbol[,symbol]>" [--depth N] [--max-nodes N]
85
+ ```
86
+
87
+ It emits a fenced `flowchart` and nothing else - no renderer, no plugin, no
88
+ dependency, because mermaid is text and GitHub renders it natively in pull
89
+ requests, issues and markdown files. Traversal is not reimplemented: `findByName`
90
+ and `affected` are imported from `graph-affected.mjs`, so the diagram and the
91
+ text report cannot disagree about what is affected.
92
+
93
+ Three properties that are enforced rather than promised
94
+ (`smoke-graph-mermaid.sh`):
95
+
96
+ - Every drawn node and edge resolves back into `code-graph.json`, with the edge
97
+ kind it claims. A diagram is read as fact and checked less than prose, so an
98
+ invented edge is the expensive failure.
99
+ - Over `--max-nodes` the leftover count is printed inside the diagram, not
100
+ dropped. A small picture of a large blast radius reads as reassurance.
101
+ - The graph's `baseCommit` is printed beside it. A graph built before the change
102
+ draws an older tree, and nothing else in the PR would reveal that.
103
+
104
+ Exit 1 with a reason on stderr means no graph or no such symbol. The caller
105
+ records the gap and writes the prose alone; it never hand-draws a replacement.
106
+
107
+ Jira is not a target: its renderer turns the fence into a literal
108
+ `{code:mermaid}` block. Confluence renders it through the `ac:name="mermaid"`
109
+ macro when the space carries the plugin (`channels/confluence.md`).
@@ -1,5 +1,16 @@
1
1
  # Design conformance - the component walk, and why a glance is not a pass
2
2
 
3
+ <!-- toc -->
4
+ - [0. Why this runs as a gate](#0-why-this-runs-as-a-gate)
5
+ - [1. Enumerate first, then fill every cell](#1-enumerate-first-then-fill-every-cell)
6
+ - [2. Measure, never read the token](#2-measure-never-read-the-token)
7
+ - [3. Adaptive per-component convergence](#3-adaptive-per-component-convergence)
8
+ - [4. Scope tags - how an item is verified](#4-scope-tags---how-an-item-is-verified)
9
+ - [5. The catalog](#5-the-catalog)
10
+ - [6. Output](#6-output)
11
+ - [7. What a token catalog is, and why none ships here](#7-what-a-token-catalog-is-and-why-none-ships-here)
12
+ <!-- /toc -->
13
+
3
14
  A design audit that reads a screen top to bottom and reports what looks wrong
4
15
  finds about one defect per element: the wrong font on a label, and not that same
5
16
  label's wrong colour and wrong inset. This file is the catalog, and the three
@@ -14,6 +25,19 @@ Consumers: `/multi-agent:design-check` (the runner), Phase 4 review when a UI
14
25
  diff is under review, and `features/visual-evidence.md` when a capture has to
15
26
  prove a fix. Gate: `smoke-design-conformance.sh`.
16
27
 
28
+ ## 0. Why this runs as a gate
29
+
30
+ `design-check` existed as a command for a while with no phase invoking it, so the only
31
+ thing standing between a build and visual drift was the user opening the app and
32
+ looking. On one run that produced 16pt padding where the frame said `Spacing/12`, and a
33
+ full sheet rebuild afterwards.
34
+
35
+ The reason it cannot be advice is structural, not historical: a reviewer reading a diff
36
+ cannot see spacing. Every other Phase 4 check reads text and reasons about text; this
37
+ one is the only thing in the pipeline that compares a rendered result against the
38
+ design it was drawn from. Left optional, it is the check that gets skipped on exactly
39
+ the runs that are in a hurry, which are the runs that produce drift.
40
+
17
41
  ## 1. Enumerate first, then fill every cell
18
42
 
19
43
  Do **not** walk by finding, and do not walk by headline. Build the inventory
@@ -1,5 +1,15 @@
1
1
  # doctor - the check registry
2
2
 
3
+ <!-- toc -->
4
+ - [What the exit code means](#what-the-exit-code-means)
5
+ - [What BLOCK means, exactly](#what-block-means-exactly)
6
+ - [Who calls it, and what they do with the code](#who-calls-it-and-what-they-do-with-the-code)
7
+ - [The four severities](#the-four-severities)
8
+ - [The line shape](#the-line-shape)
9
+ - [It recommends, it never fixes](#it-recommends-it-never-fixes)
10
+ - [Checks](#checks)
11
+ <!-- /toc -->
12
+
3
13
  Every check `/multi-agent:doctor` can report has a `### <id>` heading here, and
4
14
  `doctor.mjs --list-checks` prints exactly the same set. The equality is checked
5
15
  in both directions by `smoke-doctor.sh`: a check that ships without an entry
@@ -1,5 +1,12 @@
1
1
  # Feature: External Context Injection (Phase 1 Step 1.5)
2
2
 
3
+ <!-- toc -->
4
+ - [Dispatch table](#dispatch-table)
5
+ - [Exit code handling](#exit-code-handling)
6
+ - [Prompt injection shape](#prompt-injection-shape)
7
+ - [Log line shape](#log-line-shape)
8
+ <!-- /toc -->
9
+
3
10
  Phase 0 Step 1b catalogued every typed external link from the task description into `state.contextLinks[]`. Phase 1 dispatches each entry to its matching fetcher and prepends the result to the analysis prompt under a **Referenced External Sources** section, so the agent doesn't re-discover what the ticket already pointed at.
4
11
 
5
12
  ```bash
@@ -1,5 +1,14 @@
1
1
  # Related-issue context at intake
2
2
 
3
+ <!-- toc -->
4
+ - [What it costs](#what-it-costs)
5
+ - [Shape](#shape)
6
+ - [Settings](#settings)
7
+ - [Maturity](#maturity)
8
+ - [Where it goes](#where-it-goes)
9
+ - [Not included](#not-included)
10
+ <!-- /toc -->
11
+
3
12
  A development sub-task is often filed with no description of its own. The
4
13
  requirement sits on the parent, and the rest of the picture - the analysis, the
5
14
  test scope - sits on the sibling sub-tasks beside it. The fetcher already read
@@ -1,5 +1,15 @@
1
1
  # Model Fallback Contract
2
2
 
3
+ <!-- toc -->
4
+ - [Tier ladder](#tier-ladder)
5
+ - [Prefs knob](#prefs-knob)
6
+ - [Turning the fable rung off](#turning-the-fable-rung-off)
7
+ - [Triggers (checked in this order)](#triggers-checked-in-this-order)
8
+ - [Logging](#logging)
9
+ - [Non-goals](#non-goals)
10
+ - [Codex CLI](#codex-cli)
11
+ <!-- /toc -->
12
+
3
13
  > Contract last revised in **v10.6.0** (Fable 5 restored as top tier). The version tag here tracks the last substantive change to this contract, not the pipeline release.
4
14
 
5
15
  Personas route to the top available intelligence tier they declare in
@@ -0,0 +1,132 @@
1
+ # The review file set - what was read, and what was not, on the record
2
+
3
+ <!-- toc -->
4
+ - [Why this exists](#why-this-exists)
5
+ - [The set is fixed before the reviewer sees the diff](#the-set-is-fixed-before-the-reviewer-sees-the-diff)
6
+ - [Every exclusion names the pattern that produced it](#every-exclusion-names-the-pattern-that-produced-it)
7
+ - [The reviewer answers for each file](#the-reviewer-answers-for-each-file)
8
+ - [Invocation](#invocation)
9
+ - [What this is not](#what-this-is-not)
10
+ <!-- /toc -->
11
+
12
+ > The denominator for Phase 4's other axis. Loaded on demand by
13
+ > `/multi-agent:review` and by pipeline Phase 4.
14
+
15
+ ## Why this exists
16
+
17
+ Phase 4 had a size cap and no exclusion list. When the diff exceeds the phase
18
+ token allowance the cap truncates the LARGEST files first, so a regenerated
19
+ lockfile or a snapshot dump is not merely wasted budget: it is the thing that
20
+ survives while real code is cut.
21
+
22
+ The second half is worse and quieter. A reviewer that opened one file of ten and
23
+ a reviewer that read all ten and found nothing return the identical
24
+ `{"findings": [], "approved": true}`. Nothing in the pipeline could tell them
25
+ apart, so "no findings" has been carrying two meanings at once.
26
+
27
+ Both halves are one fix: decide what is worth reading BEFORE the cap decides
28
+ what fits, and make the reviewer answer for each file it was given.
29
+
30
+ ## The set is fixed before the reviewer sees the diff
31
+
32
+ The same rule `selectedRules[]` follows, for the same reason. After a model has
33
+ seen the diff, "I did not open that one" and "there was nothing there" become
34
+ the same sentence, and whichever one is cheaper to say is the one that gets
35
+ said. So the set is computed from `git diff --name-only`, written to
36
+ `.pipeline/review-files.json`, and never recomputed inside the round.
37
+
38
+ ```bash
39
+ git -C "$WORKTREE" diff --name-only "$BASE_BRANCH"...HEAD \
40
+ | node $HOME/.claude/scripts/review-file-filter.mjs \
41
+ > "$WORKTREE/.pipeline/review-files.json"
42
+ ```
43
+
44
+ Report shape:
45
+
46
+ ```json
47
+ {
48
+ "reviewed": ["src/App.swift"],
49
+ "excluded": [
50
+ {
51
+ "path": "package-lock.json",
52
+ "reason": "lockfile - resolved by the package manager, not written by hand",
53
+ "pattern": "**/package-lock.json"
54
+ }
55
+ ],
56
+ "total": 2,
57
+ "patternsSource": ".../schemas/review-file-exclusions.json",
58
+ "patternCount": 45
59
+ }
60
+ ```
61
+
62
+ `reviewed` is what goes into the diff cap and into the reviewer prompt, as a
63
+ `${REVIEW_FILES}` block beside `${CRITERIA}` in the shared cache prefix. It has to
64
+ be in the prompt: a reviewer asked to account for a set it was never shown can
65
+ only guess, and `fileCoverage` would then fail on every single dispatch - a gate
66
+ that always fires is a gate that gets switched off. `excluded` goes into the run
67
+ report, never into silence.
68
+
69
+ ## Every exclusion names the pattern that produced it
70
+
71
+ A file that disappears between the diff and the review is indistinguishable from
72
+ a file nobody found anything in - which is the exact confusion this whole
73
+ feature exists to remove, so reintroducing it in the filter would be
74
+ self-defeating. Each excluded row carries both the human reason and the glob
75
+ that matched, so a reviewer, a PR reader or a future maintainer can dispute the
76
+ call rather than discover it.
77
+
78
+ The pattern list is data, in `schemas/review-file-exclusions.json`, and it is
79
+ generic: generated trees, lockfiles, recorded snapshots, vendored source, build
80
+ output, binary assets. No stack, project or company name appears in it. A
81
+ pattern with no reason invalidates the whole list rather than being defaulted -
82
+ the default would be exactly the sentence the caller is supposed to print.
83
+
84
+ **It fails open, on purpose.** An unreadable or malformed pattern file yields
85
+ every file reviewed, the reason on stderr, and exit 2. Failing closed would
86
+ review nothing and report a clean run.
87
+
88
+ ## The reviewer answers for each file
89
+
90
+ `reviewer-output.schema.json` (v1.3.0) carries `fileCoverage[]`: one row per
91
+ path in `reviewed`, `{path, verdict: reviewed|skipped, reason}`. There is
92
+ deliberately no `partial` - a file read in part is read, and what was not
93
+ understood belongs in a finding.
94
+
95
+ `validate-reviewer.mjs --coverage <report>` enforces it, catching the same three
96
+ failures the conformance checklist catches on the rule axis:
97
+
98
+ | Failure | Why it matters |
99
+ |---|---|
100
+ | a file in the set with no row | silently unread, and the empty `findings[]` reads as clean |
101
+ | a row for a path outside the set | an answer about something the reviewer was not given, the same shape as a hallucinated rule ID |
102
+ | `skipped` with no reason | a drop with no cause is indistinguishable from a read |
103
+
104
+ `skipped` is legitimate and expected: a file past the diff cap, a file whose
105
+ content the host truncated. What it may not be is unexplained. "Not relevant" is
106
+ a review decision and belongs in a verdict of `reviewed`, not a skip.
107
+
108
+ An empty `reviewed` set (a diff that is entirely lockfiles) demands no checklist
109
+ at all. Requiring an empty array there would fail honest output, and the
110
+ filter's `excluded[]` is what carries that information onward.
111
+
112
+ ## Invocation
113
+
114
+ ```bash
115
+ node $HOME/.claude/scripts/validate-reviewer.mjs "$REVIEWER_FILE" \
116
+ --criteria "$WORKTREE/.pipeline/criteria-manifest.json" \
117
+ --coverage "$WORKTREE/.pipeline/review-files.json"
118
+ ```
119
+
120
+ Without `--coverage` the field stays optional, so every existing caller keeps
121
+ working unchanged. With it, the checklist is enforced and exit 1 takes the same
122
+ single self-correction rework the rest of the validator gate takes.
123
+
124
+ ## What this is not
125
+
126
+ It is not a relevance filter. Nothing here decides that a file is uninteresting;
127
+ it decides that a file is not human-authored source, which is a mechanical
128
+ question with a mechanical answer. The moment a pattern starts encoding "we
129
+ probably do not care about this directory", the list has become a way to hide
130
+ work, and the near-miss assertions in `smoke-review-file-filter.sh`
131
+ (`CodeGenerator.swift`, `generated-report.md`, `distribution/`, `buildSrc/`) are
132
+ what fail when it does.
@@ -1,5 +1,18 @@
1
1
  # Skill conformance - reviewing against the criteria the work was built to
2
2
 
3
+ <!-- toc -->
4
+ - [Why this exists](#why-this-exists)
5
+ - [The four rules that make it work](#the-four-rules-that-make-it-work)
6
+ - [Registry discovery is declared, never sniffed](#registry-discovery-is-declared-never-sniffed)
7
+ - [Scope is required, and it is what makes this stack-generic](#scope-is-required-and-it-is-what-makes-this-stack-generic)
8
+ - [What is deterministic here, and what is deliberately not](#what-is-deterministic-here-and-what-is-deliberately-not)
9
+ - [The one bespoke scan: exception markers](#the-one-bespoke-scan-exception-markers)
10
+ - [Dev-mode substitutes (Phases 1 and 2 never ran)](#dev-mode-substitutes-phases-1-and-2-never-ran)
11
+ - [Invocation](#invocation)
12
+ - [Handoff to the reviewers](#handoff-to-the-reviewers)
13
+ - [Preference](#preference)
14
+ <!-- /toc -->
15
+
3
16
  > **TLDR** - Phase 4 Step 1.78 resolves, deterministically and before any reviewer runs, WHAT the changed code was supposed to honour: declared rule registries scoped to the diff's languages, in-repo module guides, and the toolchains those registries delegate to. The result is `criteria-manifest.json`: a bounded set of rule IDs that becomes the denominator for "was this applied completely". Reviewers answer per rule ID. An ID that is neither checked nor explicitly waived fails the stage.
4
17
 
5
18
  ## Why this exists