@mmerterden/multi-agent-pipeline 17.1.0 → 17.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (57) hide show
  1. package/CHANGELOG.md +127 -0
  2. package/README.md +7 -0
  3. package/README.tr.md +7 -0
  4. package/docs/token-budget-history.md +22 -0
  5. package/install/_dev-only-files.mjs +1 -0
  6. package/install/codex.mjs +18 -1
  7. package/install/copilot.mjs +17 -1
  8. package/package.json +1 -1
  9. package/pipeline/commands/multi-agent/autopilot-on/SKILL.md +9 -1
  10. package/pipeline/lib/autopilot-state.sh +34 -0
  11. package/pipeline/multi-agent-refs/_dev-context.md +10 -0
  12. package/pipeline/multi-agent-refs/analysis/redesign.md +8 -0
  13. package/pipeline/multi-agent-refs/analysis/review.md +9 -0
  14. package/pipeline/multi-agent-refs/android-guide.md +14 -0
  15. package/pipeline/multi-agent-refs/audit-guide.md +12 -0
  16. package/pipeline/multi-agent-refs/backend-guide.md +10 -0
  17. package/pipeline/multi-agent-refs/channels/confluence.md +11 -0
  18. package/pipeline/multi-agent-refs/channels/issue-comment.md +12 -0
  19. package/pipeline/multi-agent-refs/channels/jira.md +10 -0
  20. package/pipeline/multi-agent-refs/channels/pr-review-actions.md +13 -0
  21. package/pipeline/multi-agent-refs/channels/pr.md +11 -0
  22. package/pipeline/multi-agent-refs/component-dispatch.md +11 -0
  23. package/pipeline/multi-agent-refs/component-generation.md +11 -0
  24. package/pipeline/multi-agent-refs/conventions-defaults.md +15 -0
  25. package/pipeline/multi-agent-refs/cross-cli-contract.md +43 -0
  26. package/pipeline/multi-agent-refs/features/analysis-jira.md +11 -0
  27. package/pipeline/multi-agent-refs/features/design-conformance.md +10 -0
  28. package/pipeline/multi-agent-refs/features/doctor.md +10 -0
  29. package/pipeline/multi-agent-refs/features/external-context-injection.md +7 -0
  30. package/pipeline/multi-agent-refs/features/jira-context.md +9 -0
  31. package/pipeline/multi-agent-refs/features/model-fallback.md +10 -0
  32. package/pipeline/multi-agent-refs/features/skill-conformance.md +13 -0
  33. package/pipeline/multi-agent-refs/features/url-enrichment.md +9 -0
  34. package/pipeline/multi-agent-refs/features/visual-evidence.md +42 -0
  35. package/pipeline/multi-agent-refs/generate-issue.md +7 -0
  36. package/pipeline/multi-agent-refs/issue-jira-triad.md +9 -0
  37. package/pipeline/multi-agent-refs/knowledge.md +6 -0
  38. package/pipeline/multi-agent-refs/multi-repo-integration-build.md +13 -0
  39. package/pipeline/multi-agent-refs/phases/modes.md +7 -0
  40. package/pipeline/multi-agent-refs/phases/operations.md +9 -0
  41. package/pipeline/multi-agent-refs/phases/phase-0-init.md +1 -1
  42. package/pipeline/multi-agent-refs/phases.md +11 -0
  43. package/pipeline/multi-agent-refs/picker-contract.md +12 -0
  44. package/pipeline/multi-agent-refs/platform-parity.md +10 -0
  45. package/pipeline/multi-agent-refs/progress-contract.md +10 -0
  46. package/pipeline/multi-agent-refs/setup/firebase.md +9 -0
  47. package/pipeline/multi-agent-refs/swiftui-guide.md +17 -0
  48. package/pipeline/multi-agent-refs/tracker-contract.md +12 -0
  49. package/pipeline/multi-agent-refs/web-guide.md +10 -0
  50. package/pipeline/multi-agent-refs/wiki-capture.md +11 -0
  51. package/pipeline/schemas/prefs.schema.json +4 -0
  52. package/pipeline/schemas/token-budget.json +10 -19
  53. package/pipeline/scripts/autopilot-intake.mjs +5 -1
  54. package/pipeline/scripts/autopilot-runner.mjs +6 -1
  55. package/pipeline/scripts/autopilot-status.sh +3 -2
  56. package/pipeline/scripts/capture-evidence.sh +79 -11
  57. package/pipeline/scripts/gen-ref-toc.mjs +279 -0
@@ -1,5 +1,16 @@
1
1
  # Component Generation Guide (generic)
2
2
 
3
+ <!-- toc -->
4
+ - [Component Architecture: Configuration / View / Modifiers](#component-architecture-configuration-view-modifiers)
5
+ - [Configuration Purity Rule](#configuration-purity-rule)
6
+ - [View Implementation](#view-implementation)
7
+ - [Modifier Pattern](#modifier-pattern)
8
+ - [Simple vs Complex Decision](#simple-vs-complex-decision)
9
+ - [3-Layer Test Strategy](#3-layer-test-strategy)
10
+ - [Component Checklist (Before Commit)](#component-checklist-before-commit)
11
+ - [When a Figma URL Is Provided](#when-a-figma-url-is-provided)
12
+ <!-- /toc -->
13
+
3
14
  > Lifted out of `core/multi-agent/SKILL.md`, where it was loaded on every
4
15
  > run of every mode. It applies only to a task that generates a UI
5
16
  > component from a design, so it now loads when that path is taken.
@@ -4,6 +4,21 @@ description: "Convention fallback defaults for /multi-agent:analysis Phase 2b Pa
4
4
 
5
5
  # Convention Defaults - Pass B Fallback Reference
6
6
 
7
+ <!-- toc -->
8
+ - [How the fallback chain works](#how-the-fallback-chain-works)
9
+ - [C1 - Folder Structure](#c1---folder-structure)
10
+ - [C2 - Class Naming](#c2---class-naming)
11
+ - [C3 - UI State Model](#c3---ui-state-model)
12
+ - [C4 - Test Method Naming](#c4---test-method-naming)
13
+ - [C5 - Accessibility Identifier](#c5---accessibility-identifier)
14
+ - [C6 - Localization Key](#c6---localization-key)
15
+ - [C8 - SwiftUI Preview macro (iOS only)](#c8---swiftui-preview-macro-ios-only)
16
+ - [C7 - Dependency Injection](#c7---dependency-injection)
17
+ - [Risk row template (Section 20)](#risk-row-template-section-20)
18
+ - [Maintenance](#maintenance)
19
+ - [Locked decisions that govern this file](#locked-decisions-that-govern-this-file)
20
+ <!-- /toc -->
21
+
7
22
  `/multi-agent:analysis` Phase 1c extracts conventions from each selected repo (folder structure, class naming, state model, test naming, accessibility identifier, localization key, DI registration). When `confidence == "none"` AND the standards binding source (`evidence.standards[]`) does not provide an explicit rule, Pass B falls back to the platform defaults catalogued here. Every applied default emits a row in Section 20 Risks of the rendered document.
8
23
 
9
24
  > **Language**: This file is read as a system prompt. Prose stays English. Examples carry generic placeholder names (`Foo`, `Bar`); the runtime substitutes the actual feature slug.
@@ -1,5 +1,18 @@
1
1
  # Cross-CLI Contract (Claude Code · Copilot CLI · Codex CLI)
2
2
 
3
+ <!-- toc -->
4
+ - [1. Command Inventory (56 commands)](#1-command-inventory-56-commands)
5
+ - [2. Canonical Placeholder Vocabulary](#2-canonical-placeholder-vocabulary)
6
+ - [2.6 Intentional structural divergence - thin dispatcher vs inlined orchestrator](#26-intentional-structural-divergence---thin-dispatcher-vs-inlined-orchestrator)
7
+ - [3. Frontmatter Transform Rules (Claude ↔ Copilot)](#3-frontmatter-transform-rules-claude-copilot)
8
+ - [4. Progress Signalling Parity](#4-progress-signalling-parity)
9
+ - [5. Argument Parsing Invariants](#5-argument-parsing-invariants)
10
+ - [6. Output Format Expectations](#6-output-format-expectations)
11
+ - [7. Platform Guards (macOS)](#7-platform-guards-macos)
12
+ - [8. Enforcement](#8-enforcement)
13
+ - [9. Change Control](#9-change-control)
14
+ <!-- /toc -->
15
+
3
16
  > **Non-negotiable**. Any change that breaks this contract blocks merge. Validated by `smoke-cross-cli-behavior.sh`.
4
17
 
5
18
  **Purpose**: every pipeline command must produce identical artifacts (state, logs, outputs) and respect identical placeholder vocabulary regardless of which of the three host CLIs invokes it. This file is the source of truth for "what must stay the same."
@@ -264,6 +277,35 @@ Every phase boundary MUST call EITHER TaskCreate/Update (Claude) OR `phase-track
264
277
 
265
278
  `phase-banner.sh` runs on both CLIs. Same output format, no Claude-only or Copilot-only flair.
266
279
 
280
+ ### 4.4 Continuous mode
281
+
282
+ `autopilot-on`, `autopilot-off` and `autopilot-status` ship to all three hosts and
283
+ behave identically there, because the thing they control is not a CLI feature:
284
+ it is a launchd user agent whose tick spawns a `claude --bg` child regardless of
285
+ which CLI you typed the command in. Continuous mode therefore requires the
286
+ `claude` binary on `PATH` on every host, Copilot and Codex included, and there is
287
+ no Copilot or Codex equivalent of the child.
288
+
289
+ | Concept | Claude Code | Copilot CLI | Codex CLI |
290
+ |---|---|---|---|
291
+ | Turn the mode on | `/multi-agent:autopilot-on` | `/multi-agent-autopilot-on` | `multi-agent autopilot-on` via the router skill |
292
+ | Scripts, libs, templates | `~/.claude/{scripts,lib,templates}` | `~/.copilot/...` | `~/.codex/...` |
293
+ | Which one is read | `ma_ap_asset <rel>` resolves the caller's own tree first, then the other two | same | same |
294
+ | Queue and config state | `~/.claude/autopilot/` | `~/.claude/autopilot/` | `~/.claude/autopilot/` |
295
+ | In-flight rows on the widget | `autopilot-status --subjects` into `TaskUpdate` | into the reprinted card | into `update_plan` |
296
+
297
+ The state row is the one that looks wrong and is not. `~/.claude/autopilot/` is
298
+ shared across hosts on purpose, exactly like `logs/`, `knowledge/` and
299
+ `multi-agent-preferences.json` (2.6, "shared state is deliberately NOT
300
+ retargeted"): two CLIs on one machine must read ONE queue. A per-host state root
301
+ would give a Copilot session a second, invisible queue and the same ticket would
302
+ be taken twice.
303
+
304
+ Everything that is NOT state resolves per host, and that half had to be fixed:
305
+ `templates/` was laid down only by `install/claude.mjs`, so `autopilot-on` on a
306
+ Codex-only machine rendered a launchd job from a file the host did not have.
307
+ `smoke-autopilot-hosts.sh` holds the line.
308
+
267
309
  ---
268
310
 
269
311
  ## 5. Argument Parsing Invariants
@@ -338,6 +380,7 @@ installs on.
338
380
 
339
381
  This contract is validated by:
340
382
 
383
+ - `smoke-autopilot-hosts.sh` - asserts continuous mode resolves its scripts, libs and plist template into whichever host tree is installed, and that the state root is NOT retargeted (4.4)
341
384
  - `smoke-cross-cli-behavior.sh` - asserts every command behaves identically, pulls from Section 2 (placeholder vocab), Section 5 (argument parsing), Section 6 (output formats); also regression-locks the 8-persona agent deployment
342
385
  - `smoke-commands-skills-parity.sh` (two assertions per command) - enforces colon-form command ↔ dash-form skill directory parity
343
386
  - `smoke-compliance-skills.sh` - enforces store-compliance skill catalog + 4 consumer wiring
@@ -1,5 +1,16 @@
1
1
  # analysis-jira - an analysis document, read as work
2
2
 
3
+ <!-- toc -->
4
+ - [The marker gate runs first, before any network call](#the-marker-gate-runs-first-before-any-network-call)
5
+ - [Coverage is two-way, and the second direction is the useful one](#coverage-is-two-way-and-the-second-direction-is-the-useful-one)
6
+ - [An unverifiable run is allowed; looking verified is not](#an-unverifiable-run-is-allowed-looking-verified-is-not)
7
+ - [Identity is a label, not a title](#identity-is-a-label-not-a-title)
8
+ - [The write is ledgered](#the-write-is-ledgered)
9
+ - [An existing node is skipped, never updated](#an-existing-node-is-skipped-never-updated)
10
+ - [Every site-specific name is a VALUE, never a schema key](#every-site-specific-name-is-a-value-never-a-schema-key)
11
+ - [Auth](#auth)
12
+ <!-- /toc -->
13
+
3
14
  `/multi-agent:analysis-jira` turns a rendered analysis document into a Jira tree.
4
15
  Two files do it, and the split is the design:
5
16
 
@@ -1,5 +1,15 @@
1
1
  # Design conformance - the component walk, and why a glance is not a pass
2
2
 
3
+ <!-- toc -->
4
+ - [1. Enumerate first, then fill every cell](#1-enumerate-first-then-fill-every-cell)
5
+ - [2. Measure, never read the token](#2-measure-never-read-the-token)
6
+ - [3. Adaptive per-component convergence](#3-adaptive-per-component-convergence)
7
+ - [4. Scope tags - how an item is verified](#4-scope-tags---how-an-item-is-verified)
8
+ - [5. The catalog](#5-the-catalog)
9
+ - [6. Output](#6-output)
10
+ - [7. What a token catalog is, and why none ships here](#7-what-a-token-catalog-is-and-why-none-ships-here)
11
+ <!-- /toc -->
12
+
3
13
  A design audit that reads a screen top to bottom and reports what looks wrong
4
14
  finds about one defect per element: the wrong font on a label, and not that same
5
15
  label's wrong colour and wrong inset. This file is the catalog, and the three
@@ -1,5 +1,15 @@
1
1
  # doctor - the check registry
2
2
 
3
+ <!-- toc -->
4
+ - [What the exit code means](#what-the-exit-code-means)
5
+ - [What BLOCK means, exactly](#what-block-means-exactly)
6
+ - [Who calls it, and what they do with the code](#who-calls-it-and-what-they-do-with-the-code)
7
+ - [The four severities](#the-four-severities)
8
+ - [The line shape](#the-line-shape)
9
+ - [It recommends, it never fixes](#it-recommends-it-never-fixes)
10
+ - [Checks](#checks)
11
+ <!-- /toc -->
12
+
3
13
  Every check `/multi-agent:doctor` can report has a `### <id>` heading here, and
4
14
  `doctor.mjs --list-checks` prints exactly the same set. The equality is checked
5
15
  in both directions by `smoke-doctor.sh`: a check that ships without an entry
@@ -1,5 +1,12 @@
1
1
  # Feature: External Context Injection (Phase 1 Step 1.5)
2
2
 
3
+ <!-- toc -->
4
+ - [Dispatch table](#dispatch-table)
5
+ - [Exit code handling](#exit-code-handling)
6
+ - [Prompt injection shape](#prompt-injection-shape)
7
+ - [Log line shape](#log-line-shape)
8
+ <!-- /toc -->
9
+
3
10
  Phase 0 Step 1b catalogued every typed external link from the task description into `state.contextLinks[]`. Phase 1 dispatches each entry to its matching fetcher and prepends the result to the analysis prompt under a **Referenced External Sources** section, so the agent doesn't re-discover what the ticket already pointed at.
4
11
 
5
12
  ```bash
@@ -1,5 +1,14 @@
1
1
  # Related-issue context at intake
2
2
 
3
+ <!-- toc -->
4
+ - [What it costs](#what-it-costs)
5
+ - [Shape](#shape)
6
+ - [Settings](#settings)
7
+ - [Maturity](#maturity)
8
+ - [Where it goes](#where-it-goes)
9
+ - [Not included](#not-included)
10
+ <!-- /toc -->
11
+
3
12
  A development sub-task is often filed with no description of its own. The
4
13
  requirement sits on the parent, and the rest of the picture - the analysis, the
5
14
  test scope - sits on the sibling sub-tasks beside it. The fetcher already read
@@ -1,5 +1,15 @@
1
1
  # Model Fallback Contract
2
2
 
3
+ <!-- toc -->
4
+ - [Tier ladder](#tier-ladder)
5
+ - [Prefs knob](#prefs-knob)
6
+ - [Turning the fable rung off](#turning-the-fable-rung-off)
7
+ - [Triggers (checked in this order)](#triggers-checked-in-this-order)
8
+ - [Logging](#logging)
9
+ - [Non-goals](#non-goals)
10
+ - [Codex CLI](#codex-cli)
11
+ <!-- /toc -->
12
+
3
13
  > Contract last revised in **v10.6.0** (Fable 5 restored as top tier). The version tag here tracks the last substantive change to this contract, not the pipeline release.
4
14
 
5
15
  Personas route to the top available intelligence tier they declare in
@@ -1,5 +1,18 @@
1
1
  # Skill conformance - reviewing against the criteria the work was built to
2
2
 
3
+ <!-- toc -->
4
+ - [Why this exists](#why-this-exists)
5
+ - [The four rules that make it work](#the-four-rules-that-make-it-work)
6
+ - [Registry discovery is declared, never sniffed](#registry-discovery-is-declared-never-sniffed)
7
+ - [Scope is required, and it is what makes this stack-generic](#scope-is-required-and-it-is-what-makes-this-stack-generic)
8
+ - [What is deterministic here, and what is deliberately not](#what-is-deterministic-here-and-what-is-deliberately-not)
9
+ - [The one bespoke scan: exception markers](#the-one-bespoke-scan-exception-markers)
10
+ - [Dev-mode substitutes (Phases 1 and 2 never ran)](#dev-mode-substitutes-phases-1-and-2-never-ran)
11
+ - [Invocation](#invocation)
12
+ - [Handoff to the reviewers](#handoff-to-the-reviewers)
13
+ - [Preference](#preference)
14
+ <!-- /toc -->
15
+
3
16
  > **TLDR** - Phase 4 Step 1.78 resolves, deterministically and before any reviewer runs, WHAT the changed code was supposed to honour: declared rule registries scoped to the diff's languages, in-repo module guides, and the toolchains those registries delegate to. The result is `criteria-manifest.json`: a bounded set of rule IDs that becomes the denominator for "was this applied completely". Reviewers answer per rule ID. An ID that is neither checked nor explicitly waived fails the stage.
4
17
 
5
18
  ## Why this exists
@@ -1,5 +1,14 @@
1
1
  # Step 1b - URL Enrichment (Phase 0)
2
2
 
3
+ <!-- toc -->
4
+ - [Step 1b.0 - Extract context links (always runs)](#step-1b0---extract-context-links-always-runs)
5
+ - [Step 1b.1 - Firebase Crashlytics deep fetch (runs when `type == "crashlytics"` present in `state.contextLinks[]`)](#step-1b1---firebase-crashlytics-deep-fetch-runs-when-type-crashlytics-present-in-statecontextlinks)
6
+ - [Step 1b.2 - Fortify SSC deep fetch (runs when `type == "fortify"` present in `state.contextLinks[]`)](#step-1b2---fortify-ssc-deep-fetch-runs-when-type-fortify-present-in-statecontextlinks)
7
+ - [Step 1b.3 - Graylog deep fetch (runs when `type == "graylog"` present in `state.contextLinks[]`)](#step-1b3---graylog-deep-fetch-runs-when-type-graylog-present-in-statecontextlinks)
8
+ - [Step 1b.3b - Document fetch (runs when `type == "document"` present)](#step-1b3b---document-fetch-runs-when-type-document-present)
9
+ - [Step 1b.4 - Other link types (catalogued only at Phase 0; fetched at Phase 1)](#step-1b4---other-link-types-catalogued-only-at-phase-0-fetched-at-phase-1)
10
+ <!-- /toc -->
11
+
3
12
  > Loaded on demand from `phases/phase-0-init.md` Step 1b. It lives here rather
4
13
  > than inline because it applies only to a task that carries URLs: a bare Jira
5
14
  > ID, a GitHub issue number, or a free-text task never needs any of it, yet every
@@ -1,5 +1,17 @@
1
1
  # Visual evidence - before/after screenshots and the UI flow video
2
2
 
3
+ <!-- toc -->
4
+ - [1. When it is required](#1-when-it-is-required)
5
+ - [2. Before - the reporter's screenshot, or nothing](#2-before---the-reporters-screenshot-or-nothing)
6
+ - [3. After - Phase 3, not Phase 5](#3-after---phase-3-not-phase-5)
7
+ - [4. Video - the recording rides on a test run](#4-video---the-recording-rides-on-a-test-run)
8
+ - [5. Size, and what happens when it does not fit](#5-size-and-what-happens-when-it-does-not-fit)
9
+ - [5b. Where the artefacts live - the host](#5b-where-the-artefacts-live---the-host)
10
+ - [6. Rendering](#6-rendering)
11
+ - [7. Blocker](#7-blocker)
12
+ - [8. State](#8-state)
13
+ <!-- /toc -->
14
+
3
15
  A UI fix that reads as three changed files in a diff is not reviewable. The
4
16
  reviewer cannot see what was wrong, and the tester cannot see what to look for.
5
17
  This contract makes the pipeline carry the picture: the state the reporter saw,
@@ -222,6 +234,36 @@ capture of a static screen.
222
234
  `visualEvidence.enabled` turns the whole feature off - capture, upload, both
223
235
  render sections and the Phase 6 blocker with it.
224
236
 
237
+ ### 4.6 Web - the recorder IS the runner
238
+
239
+ Web is a real evidence platform: the schema has it in `visualEvidence.platform`,
240
+ the probe opens tier 1 and tier 2 for it, and `run-ui-tests.sh` has a web arm.
241
+ What it does NOT share with iOS and Android is the recording model, and forcing
242
+ it to would record the same run twice - Playwright and Cypress write the video
243
+ themselves.
244
+
245
+ | Step | iOS / Android | Web |
246
+ |---|---|---|
247
+ | `after` | screenshot the booted device | drive a browser at a URL |
248
+ | the URL | not needed, a simulator already shows something | `--url`, else `prefs.global.visualEvidence.webBaseUrl`, else a gap with that reason |
249
+ | `video start` | spawn a recorder, hold its pid | write a marker with the current time; spawn nothing |
250
+ | `video stop` | kill the recorder, pull the file | harvest the newest video under `test-results/` or `cypress/videos/` written AFTER the marker |
251
+ | container | mp4 already | webm re-encoded to h264, for the same reason the iOS arm re-encodes: that is what the Jira preview and the PR body play |
252
+
253
+ The marker is the part that matters and the part that is easy to leave out. Without
254
+ a timestamp to compare against, `stop` harvests whatever video is lying around -
255
+ and a video from yesterday's run attached as today's evidence is worse than no
256
+ video, because an artefact that is present does not get re-checked.
257
+
258
+ Two honest gaps, both exit 4 with a reason rather than a failure: no address to
259
+ point at, and a project whose runner config does not record video at all. Neither
260
+ is a phase failure; both are recorded and shown.
261
+
262
+ Every browser call goes through `npx --no-install`, never a bare `npx`. A bare one
263
+ DOWNLOADS the browser stack, which turns "this project has no browser tooling"
264
+ into a silent network fetch. `run-ui-tests.sh` carries the same rule for the same
265
+ reason, and it is the line that keeps the pipeline's own dependency list empty.
266
+
225
267
  ## 5. Size, and what happens when it does not fit
226
268
 
227
269
  Jira's attachment ceiling is an instance setting, so it is a preference:
@@ -1,5 +1,12 @@
1
1
  # Generate Issue - Shared Flow (generate)
2
2
 
3
+ <!-- toc -->
4
+ - [Hard rules (must not regress)](#hard-rules-must-not-regress)
5
+ - [Standard templates](#standard-templates)
6
+ - [Flow](#flow)
7
+ - [Error paths](#error-paths)
8
+ <!-- /toc -->
9
+
3
10
  > **TLDR** - Shared 12-step flow for `/multi-agent:create-jira`. Asks the issue type (Task / Bug / Story), mines the target project's existing same-type issues to learn team conventions, detects the active sprint, drafts a standards-compliant issue from a fixed standard template with auto-sizing sections, asks the user about every genuinely unknown field, renders a full preview, and creates the Jira issue only after explicit approval. Creates exactly one Jira issue per run - no branches, no commits, no worktrees.
4
11
 
5
12
  Consumed by `create-jira/SKILL.md`. This ref is never invoked directly.
@@ -1,5 +1,14 @@
1
1
  # Issue → Jira → Wiki Triad
2
2
 
3
+ <!-- toc -->
4
+ - [The triad at a glance](#the-triad-at-a-glance)
5
+ - [Phase 0 auto-create policy](#phase-0-auto-create-policy)
6
+ - [Phase 7 wiki → Jira comment](#phase-7-wiki-jira-comment)
7
+ - [Autopilot behaviour](#autopilot-behaviour)
8
+ - [Preferences involved](#preferences-involved)
9
+ - [Cross-CLI parity](#cross-cli-parity)
10
+ <!-- /toc -->
11
+
3
12
  > **TLDR** - When a GitHub issue triggers the pipeline and has no Jira ID, the `autoJiraFromGithubIssue` policy decides whether to auto-create a Jira task (and patch the GitHub issue body with the new Jira link). Phase 7 then posts a humanizer'd wiki-content summary back as a Jira comment, closing the loop. Autopilot treats `ask` as `always`.
4
13
 
5
14
  This doc is referenced from `$HOME/.claude/multi-agent-refs/phases/phase-0-init.md` Step 1 (GitHub issue input) and `$HOME/.claude/multi-agent-refs/phases/phase-7-report.md` Step 2 (component wiki). Keeps the phase docs tight and gives the triad contract a stable home.
@@ -1,5 +1,11 @@
1
1
  ## Project Knowledge Base
2
2
 
3
+ <!-- toc -->
4
+ - [Project Knowledge Base](#project-knowledge-base)
5
+ - [Retrieval order, and what each layer costs](#retrieval-order-and-what-each-layer-costs)
6
+ - [Skill Injection Strategy](#skill-injection-strategy)
7
+ <!-- /toc -->
8
+
3
9
  An incrementally growing knowledge system for each project, enabling knowledge transfer across sessions.
4
10
 
5
11
  ### Directory Structure
@@ -1,5 +1,18 @@
1
1
  # Multi-Repo Integration Build - Learn Once, Auto-Apply
2
2
 
3
+ <!-- toc -->
4
+ - [Why this exists](#why-this-exists)
5
+ - [When this rule fires](#when-this-rule-fires)
6
+ - [Learn-once prompt (first encounter of a combo)](#learn-once-prompt-first-encounter-of-a-combo)
7
+ - [Auto-run steps (when host is learned)](#auto-run-steps-when-host-is-learned)
8
+ - [Error evaluation contract](#error-evaluation-contract)
9
+ - [Autopilot behavior](#autopilot-behavior)
10
+ - [Phase 7 knowledge capture](#phase-7-knowledge-capture)
11
+ - [Generic rule (for copilot-instructions.md)](#generic-rule-for-copilot-instructionsmd)
12
+ - [Schema](#schema)
13
+ - [Smoke coverage](#smoke-coverage)
14
+ <!-- /toc -->
15
+
3
16
  > **TLDR** - When a task touches ≥2 repos that have a producer→consumer dependency (e.g. shared codegen library + consuming UI library), the pipeline MUST build the **host project** that integrates them before commit/PR. Codegen mismatches (nested vs flat keys, missing entries, overwritten files) only surface when the full dependency chain builds together. Building repos in isolation gives false confidence.
4
17
  >
5
18
  > The pipeline **learns** the host per repo-combo once, persists to `prefs.global.multiRepoIntegrationHosts`, and auto-applies on subsequent runs.
@@ -8,6 +8,13 @@
8
8
 
9
9
  ## Autopilot Mode
10
10
 
11
+ <!-- toc -->
12
+ - [Autopilot Mode](#autopilot-mode)
13
+ - [Pipeline depth (Full / Short)](#pipeline-depth-full-short)
14
+ - [Analysis Mode (`/multi-agent:analysis`)](#analysis-mode-multi-agentanalysis)
15
+ - [Local Mode (`--local`)](#local-mode---local)
16
+ <!-- /toc -->
17
+
11
18
  Autopilot mode skips interactive confirmations and runs the pipeline end-to-end autonomously.
12
19
 
13
20
  **Activation**: Add `autopilot` flag to any pipeline command:
@@ -8,6 +8,15 @@
8
8
 
9
9
  ## Task ID System
10
10
 
11
+ <!-- toc -->
12
+ - [Task ID System](#task-id-system)
13
+ - [Kill Logic](#kill-logic)
14
+ - [Clear Logs](#clear-logs)
15
+ - [Purge (Full Reset)](#purge-full-reset)
16
+ - [Resume Logic](#resume-logic)
17
+ - [Phase Pipeline](#phase-pipeline)
18
+ <!-- /toc -->
19
+
11
20
  Every task gets an auto-incremented short ID. Counter stored at `$HOME/.claude/logs/multi-agent/{project}/.counter` (persists across sessions).
12
21
 
13
22
  ```
@@ -606,7 +606,7 @@ if [ -n "$EVIDENCE_PLATFORM" ]; then
606
606
  fi
607
607
  ```
608
608
 
609
- Empty is an outcome, not a failure: no simulator on web or backend, so the probe does not run and `evidenceCapability.skippedReason` says so. No `--changed` yet, by the same token.
609
+ Empty is an outcome, not a failure: backend has no device, so the probe does not run and `evidenceCapability.skippedReason` says so. Web does: the browser the runner drives is the device (4.6). No `--changed` yet, by the same token.
610
610
 
611
611
  One run, both forms: stdout is `EVIDENCE_*` (shell-quoted, so the eval is safe), and the same measurement lands as JSON.
612
612
 
@@ -1,5 +1,16 @@
1
1
  # Multi-Agent Pipeline - Phase Reference
2
2
 
3
+ <!-- toc -->
4
+ - [Phase Files](#phase-files)
5
+ - [Pipeline Flow](#pipeline-flow)
6
+ - [Phase entry - pending steer (every phase, every mode)](#phase-entry---pending-steer-every-phase-every-mode)
7
+ - [Visual Phase Tracker](#visual-phase-tracker)
8
+ - [Preferences File](#preferences-file)
9
+ - [Host Configuration](#host-configuration)
10
+ - [Token Budget](#token-budget)
11
+ - [SubPhase Convention](#subphase-convention)
12
+ <!-- /toc -->
13
+
3
14
  ## Phase Files
4
15
 
5
16
  | Phase | File |
@@ -1,5 +1,17 @@
1
1
  # Picker Contract (cross-platform single-choice abstraction)
2
2
 
3
+ <!-- toc -->
4
+ - [The abstract primitive](#the-abstract-primitive)
5
+ - [Step narration (breadcrumb)](#step-narration-breadcrumb)
6
+ - [Per-platform rendering (degradation ladder)](#per-platform-rendering-degradation-ladder)
7
+ - [Universal fallback: `pipeline/lib/ask-choice.sh`](#universal-fallback-pipelinelibask-choicesh)
8
+ - [Localized labels: what the caller owns](#localized-labels-what-the-caller-owns)
9
+ - [Order: project, then repo, then branch](#order-project-then-repo-then-branch)
10
+ - [A single candidate is still a question](#a-single-candidate-is-still-a-question)
11
+ - [Autopilot / non-interactive contract](#autopilot-non-interactive-contract)
12
+ - [Deterministic gates note](#deterministic-gates-note)
13
+ <!-- /toc -->
14
+
3
15
  > Native `AskUserQuestion` is a Claude-Code primitive; Copilot CLI has no
4
16
  > agent-invokable choice picker. This contract defines ONE abstract "ask the
5
17
  > user to choose" primitive that each supported CLI renders to the best
@@ -4,6 +4,16 @@ description: "Internal - cross-platform parity cross-check for Phase 4 and /mu
4
4
 
5
5
  # Platform parity - compare the change against the other platform's repo
6
6
 
7
+ <!-- toc -->
8
+ - [When it runs](#when-it-runs)
9
+ - [Read-only, without exception](#read-only-without-exception)
10
+ - [Locating the counterpart, deterministically](#locating-the-counterpart-deterministically)
11
+ - [What is compared](#what-is-compared)
12
+ - [What a finding may not claim](#what-a-finding-may-not-claim)
13
+ - [Output](#output)
14
+ - [Severity](#severity)
15
+ <!-- /toc -->
16
+
7
17
  A feature that exists on iOS and Android is written twice, and the two copies
8
18
  drift. The drift is invisible from inside one repo: the iOS diff is
9
19
  self-consistent, the tests pass, and nobody notices that the Android screen
@@ -1,5 +1,15 @@
1
1
  # Progress Line Contract
2
2
 
3
+ <!-- toc -->
4
+ - [Why](#why)
5
+ - [Line shape](#line-shape)
6
+ - [When to emit](#when-to-emit)
7
+ - [Verbosity](#verbosity)
8
+ - [Telemetry](#telemetry)
9
+ - [Phase adoption marker](#phase-adoption-marker)
10
+ - [Cross-CLI parity](#cross-cli-parity)
11
+ <!-- /toc -->
12
+
3
13
  > **TLDR** - Every non-trivial pipeline action emits an inline "I'm doing X right now" line so the user always knows what's happening. Immediate flush, no batching. One line per action, standard shape. Autopilot prefers verbose. Low overhead. Mirrored in telemetry as `progress.step` events.
4
14
 
5
15
  This contract is consumed by every phase (0-7) and every dispatched sub-skill (including the marketplace component toolkits when enabled). It is enforced by `smoke-progress-contract.sh` - any phase doc that drops the contract marker or diverges from the line shape fails CI.
@@ -1,5 +1,14 @@
1
1
  # Firebase / Crashlytics Onboarding (setup Step 3c)
2
2
 
3
+ <!-- toc -->
4
+ - [1. Why there are three ways in](#1-why-there-are-three-ways-in)
5
+ - [2. Tier 1 - the service account](#2-tier-1---the-service-account)
6
+ - [3. Tier 2 - interactive session plus MCP](#3-tier-2---interactive-session-plus-mcp)
7
+ - [4. appId discovery](#4-appid-discovery)
8
+ - [5. v1alpha, and what to do when it breaks](#5-v1alpha-and-what-to-do-when-it-breaks)
9
+ - [6. Skipping](#6-skipping)
10
+ <!-- /toc -->
11
+
3
12
  Loaded on demand by `/multi-agent:setup` Step 3c (optional, any platform). The SKILL.md carries the step intro; this file is the full flow.
4
13
 
5
14
  Runs inside Step 3 alongside the other missing credentials. A user who already
@@ -1,5 +1,22 @@
1
1
  ## SwiftUI Component Generation Guide (Generic)
2
2
 
3
+ <!-- toc -->
4
+ - [Component Architecture: Configuration / View / Modifiers](#component-architecture-configuration-view-modifiers)
5
+ - [Simple vs Complex Decision](#simple-vs-complex-decision)
6
+ - [Configuration Purity Rules](#configuration-purity-rules)
7
+ - [Fluent Modifier Pattern](#fluent-modifier-pattern)
8
+ - [Token Discipline](#token-discipline)
9
+ - [Variant-Driven Implementation](#variant-driven-implementation)
10
+ - [Nested Component Handling](#nested-component-handling)
11
+ - [Accessibility Requirements](#accessibility-requirements)
12
+ - [Preview Best Practices](#preview-best-practices)
13
+ - [3-Layer Test Strategy](#3-layer-test-strategy)
14
+ - [Build Verification](#build-verification)
15
+ - [Component Quality Checklist](#component-quality-checklist)
16
+ - [Compliance Rules (maps to multi-agent-toolkit MCP audit tools)](#compliance-rules-maps-to-multi-agent-toolkit-mcp-audit-tools)
17
+ - [Figma URL Given](#figma-url-given)
18
+ <!-- /toc -->
19
+
3
20
  > **MUST: Figma MCP-first (BLOCKING).** If the task references any Figma frame (URL, node ID, or "from the design"), the Dev phase MUST call `mcp__claude_ai_Figma__get_design_context` for every frame BEFORE writing a single UI line. Use the `CodeConnectSnippet` component name verbatim - no sound-alike substitutions. Authentication failure is not a skip path. Full rule, trigger conditions, and gate failure modes: `$HOME/.claude/rules/figma-pipeline.md` "MUST: Figma MCP-first (BLOCKING)". Phase wiring: `$HOME/.claude/multi-agent-refs/phases/phase-3-dev.md` "MUST: Figma MCP-first (BLOCKING pre-step)".
4
21
 
5
22
  When the task involves creating a SwiftUI component (any project), follow this architecture.
@@ -4,6 +4,18 @@ description: "Phase tracker mandatory contract - every mode depends on this. N
4
4
 
5
5
  # Phase Tracker - Mandatory Contract
6
6
 
7
+ <!-- toc -->
8
+ - [Why two channels](#why-two-channels)
9
+ - [Visual channel - chosen by the agent based on the host CLI](#visual-channel---chosen-by-the-agent-based-on-the-host-cli)
10
+ - [Call pattern](#call-pattern)
11
+ - [The plan is part of the list](#the-plan-is-part-of-the-list)
12
+ - [Resume behaviour](#resume-behaviour)
13
+ - [Continuation runs (finish / manual-test)](#continuation-runs-finish-manual-test)
14
+ - [Anti-patterns](#anti-patterns)
15
+ - [Verification](#verification)
16
+ - [Cross-reference](#cross-reference)
17
+ <!-- /toc -->
18
+
7
19
  > **TLDR** - At every phase boundary the agent does two things: (1) writes to the state file via `phase-tracker.sh` (identical on every CLI), (2) drives the visual channel for the CLI it's running in - `TaskCreate`/`TaskUpdate` in Claude Code, `phase-tracker.sh render` in every other CLI. The state file alone is not enough; the user must see the phases progress.
8
20
 
9
21
  ## Why two channels
@@ -1,5 +1,15 @@
1
1
  ## Web Development Guide
2
2
 
3
+ <!-- toc -->
4
+ - [Component Architecture](#component-architecture)
5
+ - [React Component Pattern](#react-component-pattern)
6
+ - [State Management](#state-management)
7
+ - [Token Discipline](#token-discipline)
8
+ - [Accessibility](#accessibility)
9
+ - [Testing](#testing)
10
+ - [Quality Checklist](#quality-checklist)
11
+ <!-- /toc -->
12
+
3
13
  When the task involves web development (React, Next.js, Vue), follow these patterns.
4
14
 
5
15
  ### Component Architecture
@@ -1,5 +1,16 @@
1
1
  # Component Wiki Capture (channels.md Wiki adapter)
2
2
 
3
+ <!-- toc -->
4
+ - [Applicability](#applicability)
5
+ - [Case A - preconditions met (scope multi-select)](#case-a---preconditions-met-scope-multi-select)
6
+ - [Case B - preconditions missing (actionable menu)](#case-b---preconditions-missing-actionable-menu)
7
+ - [Legacy prompt + preference flow (pre-v5.7, still supported for backward compat)](#legacy-prompt-preference-flow-pre-v57-still-supported-for-backward-compat)
8
+ - [Dispatch](#dispatch)
9
+ - [Skip conditions (explicit log lines)](#skip-conditions-explicit-log-lines)
10
+ - [Success log](#success-log)
11
+ - [Cross-CLI parity](#cross-cli-parity)
12
+ <!-- /toc -->
13
+
3
14
  > **TLDR** - Component tasks can auto-generate wiki docs + Figma screenshots. The Wiki adapter is invoked from `/multi-agent:channels` (Phase 7 delegates, or user invokes post-hoc). Four layouts supported (`submodule`, `in-repo`, `github-wiki`, `separate-repo`) - adapter picked from `figmaConfig.wiki.mode`. Non-blocking: failures log a warning and channels continues to other adapters. The Wiki adapter supports scope multi-select (Case A) and a precondition-failure menu (Case B) - see below.
4
15
 
5
16
  This doc is referenced from `commands/multi-agent/channels/SKILL.md` (Wiki adapter) and indirectly from `$HOME/.claude/multi-agent-refs/phases/phase-7-report.md` (which delegates all external delivery to channels). Keeping it separate keeps both files under their token budgets and gives the contract a stable location for Claude-side + Copilot-side implementations.
@@ -1871,6 +1871,10 @@
1871
1871
  "default": 60,
1872
1872
  "description": "Recording cap. A flow needing longer is a debugging session, not a review artefact. Android's `screenrecord` has its own ceiling of 180s that no setting can lift, so capture-evidence.sh clamps to it and says when it did: one preference honoured on one platform and silently halved on the other is worse than a stated limit."
1873
1873
  },
1874
+ "webBaseUrl": {
1875
+ "type": "string",
1876
+ "description": "The address a web capture points the browser at. Web is the one evidence platform that needs one: a simulator is already showing something, a dev server has to be pointed at. Absent, `capture-evidence.sh after --platform web` reports a gap with that reason instead of guessing a port. Per-run override: `--url`."
1877
+ },
1874
1878
  "githubHost": {
1875
1879
  "type": "string",
1876
1880
  "enum": ["branch", "off"],