@mmerterden/multi-agent-pipeline 16.31.1 → 17.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (69) hide show
  1. package/CHANGELOG.md +110 -0
  2. package/README.md +148 -103
  3. package/README.tr.md +149 -103
  4. package/docs/adr/0011-dormant-ci.md +10 -1
  5. package/docs/adr/0012-macos-only.md +98 -0
  6. package/docs/adr/README.md +1 -0
  7. package/docs/architecture.md +3 -3
  8. package/docs/ecosystem.md +5 -5
  9. package/docs/engineering.md +1 -1
  10. package/index.js +26 -0
  11. package/install/_dev-only-files.mjs +0 -1
  12. package/install/index.mjs +10 -0
  13. package/install/templates/multi-agent-autopilot.plist.template +79 -0
  14. package/package.json +5 -3
  15. package/pipeline/commands/multi-agent/autopilot-off/SKILL.md +64 -0
  16. package/pipeline/commands/multi-agent/autopilot-on/SKILL.md +173 -0
  17. package/pipeline/commands/multi-agent/autopilot-status/SKILL.md +74 -0
  18. package/pipeline/commands/multi-agent/channels/SKILL.md +41 -12
  19. package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +40 -3
  20. package/pipeline/commands/multi-agent/help/SKILL.md +43 -37
  21. package/pipeline/commands/multi-agent/manual-test/SKILL.md +1 -1
  22. package/pipeline/commands/multi-agent/setup/SKILL.md +15 -7
  23. package/pipeline/commands/multi-agent/stack/SKILL.md +31 -32
  24. package/pipeline/commands/multi-agent/status/SKILL.md +17 -1
  25. package/pipeline/commands/multi-agent/sync/SKILL.md +34 -28
  26. package/pipeline/commands/multi-agent/update/SKILL.md +1 -1
  27. package/pipeline/lib/autopilot-activation.sh +117 -0
  28. package/pipeline/lib/autopilot-state.sh +150 -0
  29. package/pipeline/lib/issue-fetcher.sh +18 -1
  30. package/pipeline/lib/plan-todos.sh +18 -0
  31. package/pipeline/lib/stack-detect.sh +200 -0
  32. package/pipeline/multi-agent-refs/channels/jira.md +80 -20
  33. package/pipeline/multi-agent-refs/channels/pr.md +65 -19
  34. package/pipeline/multi-agent-refs/cross-cli-contract.md +35 -15
  35. package/pipeline/multi-agent-refs/features/doctor.md +15 -3
  36. package/pipeline/multi-agent-refs/features/visual-evidence.md +61 -1
  37. package/pipeline/multi-agent-refs/phases/phase-0-init.md +14 -5
  38. package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +26 -12
  39. package/pipeline/multi-agent-refs/phases/phase-2-planning.md +17 -15
  40. package/pipeline/multi-agent-refs/phases/phase-3-dev.md +7 -5
  41. package/pipeline/multi-agent-refs/phases/phase-6-commit.md +1 -1
  42. package/pipeline/multi-agent-refs/phases/phase-7-report.md +2 -3
  43. package/pipeline/multi-agent-refs/readiness-review.md +7 -1
  44. package/pipeline/multi-agent-refs/rules.md +3 -11
  45. package/pipeline/multi-agent-refs/tracker-contract.md +32 -0
  46. package/pipeline/schemas/agent-state.schema.json +99 -25
  47. package/pipeline/schemas/autopilot-config.schema.json +149 -0
  48. package/pipeline/schemas/token-budget.json +4 -4
  49. package/pipeline/scripts/_stack-routing.mjs +91 -0
  50. package/pipeline/scripts/autopilot-arming.mjs +147 -0
  51. package/pipeline/scripts/autopilot-intake.mjs +383 -0
  52. package/pipeline/scripts/autopilot-menubar.swift +361 -0
  53. package/pipeline/scripts/autopilot-runner.mjs +349 -0
  54. package/pipeline/scripts/autopilot-status.sh +212 -0
  55. package/pipeline/scripts/capture-resume.sh +76 -14
  56. package/pipeline/scripts/doctor.mjs +26 -3
  57. package/pipeline/scripts/gc-abandoned.sh +352 -0
  58. package/pipeline/scripts/jira-search.sh +70 -0
  59. package/pipeline/scripts/phase-tracker.sh +134 -12
  60. package/pipeline/scripts/probe-evidence-capability.sh +27 -3
  61. package/pipeline/scripts/run-ui-tests.sh +113 -4
  62. package/pipeline/scripts/usage-report.mjs +5 -5
  63. package/pipeline/skills/.skill-manifest.json +16 -4
  64. package/pipeline/skills/shared/core/multi-agent-autopilot-off/SKILL.md +67 -0
  65. package/pipeline/skills/shared/core/multi-agent-autopilot-on/SKILL.md +146 -0
  66. package/pipeline/skills/shared/core/multi-agent-autopilot-status/SKILL.md +64 -0
  67. package/pipeline/skills/shared/core/multi-agent-channels/SKILL.md +62 -11
  68. package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +9 -8
  69. package/pipeline/scripts/gate-linux.sh +0 -62
package/CHANGELOG.md CHANGED
@@ -14,6 +14,116 @@ Internal file-layout changes that don't affect the slash-command surface are sti
14
14
 
15
15
  ---
16
16
 
17
+ ## [17.1.0] - 2026-09-14
18
+
19
+ Continuous mode, and three things that were computed correctly and shown to
20
+ nobody: the plan's own steps, a web repo's UI tests, and the queue's status when
21
+ one field arrived as a number.
22
+
23
+ ### Added
24
+
25
+ - **Continuous mode.** `/multi-agent:autopilot-on` picks the repos ONE machine watches; labelled GitHub issues and assigned + labelled Jira items then run in a worktree and stop at an open PR. `:autopilot-status` is the single producer the terminal, the menu bar indicator and the session hook all render, `:autopilot-off` removes the schedule and keeps the selection. Nothing is on by default: installing writes no state and schedules nothing, and `smoke-autopilot-default-off.sh` fails the build if that changes.
26
+
27
+ The ordering is deterministic with no model call - explicit rank, repo grouping, priority, then OLDEST first, which is the opposite of the `jira` picker on purpose: a human wants what just landed, an unattended queue must not starve what has been waiting. Two preconditions are checked before any item is taken: a rolling 24-hour spend total, and a per-source credential gate, so a dead Jira token stops Jira items without stopping GitHub ones.
28
+
29
+ A menu bar indicator draws the same `status.json` in the top right when `swiftc` is present. ActivityKit is `@available(macOS, unavailable)` - there is no Live Activity on a Mac - so this is an `NSStatusItem`, built from source on demand rather than shipped as a binary that would need signing.
30
+
31
+ - **UI tests on web.** `run-ui-tests.sh` rejected every platform but `ios|android`, so a repo with a full Playwright suite reported "no UI test target" and the PR body said UI tests had not run - a statement true of the runner and false of the repo. Detection keys on the browser-driving import (`@playwright/test`, `cy.visit(`), not on a directory called `e2e`: the same distinction that keeps 475 iOS snapshot tests from being counted as UI tests.
32
+
33
+ ### Changed
34
+
35
+ - **The Jira comment and the PR body stopped being the same document.** Jira now carries Geliştirme Özeti, Test Senaryoları, Etki Analizi and Bağlantılar with the PR link on line 1 and no identifiers, file paths or diff hunks anywhere in it. The PR keeps all of that and gains Teknik Açıklama, Etki Analizi and Build - the build command, its result and the base sha it ran against. Given/When/Then is gone from the scenarios: it reads as translated English to the person running them, who wants a titled list they can follow with the app open.
36
+
37
+ ### Fixed
38
+
39
+ - **The menu bar indicator read "autopilot kapalı" with items in flight.** `phase` was declared `String?` while the producer passes the queue's own value through `jq` untouched, so a numeric phase made Swift's `Decodable` throw - and the throw did not lose one field, it failed the whole document. A silent blank is the worst failure a status indicator has, because it is indistinguishable from good news.
40
+
41
+ - **`ma_ap_boottime` returned the microseconds.** `.*sec = ` is greedy and walks past `sec` into `usec`, so the helper produced a six-digit number that looked plausible and never equalled the same fact read anywhere else. It surfaced as a live runner being declared dead.
42
+
43
+ - **Three phase documents said "MCP forbidden" without qualifying it**, while the gate that enforces it has always matched `figma` and nothing else. The prose therefore banned the screenshot, xcodebuild and UI-test tools that Phase 3 itself calls, which is one way a run reaches Phase 7 with no evidence.
44
+
45
+ - **`channels/jira.md` told Phase 7 to upload evidence Phase 6 had already uploaded**, so re-rendering a comment attached every file a second time.
46
+
47
+ ---
48
+
49
+ ## [17.0.0] - 2026-09-14
50
+
51
+ Five things in this release were not working, and four of them looked like they
52
+ were. A platform we advertised and never verified; a stack answer that sent
53
+ nearly every repo to the iOS toolkit; an evidence step gated on a field no phase
54
+ wrote, passing an argument no phase assigned; and one word, "in progress",
55
+ covering a finished job, a crashed one and a run that never started. The common
56
+ shape is a contract with readers and no writer, which reads exactly like a
57
+ contract that is satisfied.
58
+
59
+ ### Fixed
60
+
61
+ - **"In progress" covered three different situations and offered one answer.** Measured: 20 runs read `in_progress` and every one was over a day old, the oldest 157 days. They were not one failure. 3 had their PR already open and were holding at Phase 6/7, where the pipeline pauses BY DESIGN for channel selection - finished work. 11 were left at a Phase 0 question, which is almost entirely questions, so nothing was ever built. 6 stopped mid-development, the only group that is actually broken. The session hook reported the newest of the 20 as "stopped at Phase N" and said nothing about the other 19, in the same words whether the pipeline was waiting for the user or had crashed.
62
+
63
+ `agent-state` gains `awaiting_input`, distinct from both `in_progress` and `paused`: the run is not broken and not resumable by a retry, it is finished with what it can do alone. A state file in the wild had already invented `awaiting-user-test-main-checkout` to say this. Phase 7 writes it where it pauses, the session hook and `/multi-agent:status` group on it, and each group offers its own action - resuming a Phase 0 question rebuilds nothing, and garbage-collecting an open PR throws away landed work.
64
+
65
+ A run with no status is now left out of all three groups. Eight such files exist here, six with no phase either, and calling them dead is the same false claim as calling a waiting run dead.
66
+
67
+ - **`.pr` is an object in the schema and a bare URL string in three state files on disk.** `jq '.pr.url'` on a string errors, and with stderr suppressed that reads as "no PR" - so a run whose PR was already open could be classified as dead. Both readers now read the shape instead of assuming it.
68
+
69
+ ### Added
70
+
71
+ - **`/multi-agent:garbage-collect --abandoned`**, backed by `gc-abandoned.sh`. Nothing collected these: `gc-worktrees.sh` skips REGISTERED worktrees on purpose, because a registered worktree belongs to a live run - and a run that stopped never stops being live. 28 worktrees across 7 repos, 23 GB.
72
+
73
+ Two passes, because the state files are not where the disk is: only 5 of those 28 belong to a run that still reads `in_progress`, while 19 have no state file at all and 3 belong to runs that finished and outlived their worktree. A state-driven sweep alone reaches 4% of the problem.
74
+
75
+ Three rules, each written against a shape found on the machine rather than imagined. A run waiting on you is never reaped. A path outside `<repo>/.worktrees/` is never removed - four state files record the REPO ROOT as their `worktreePath`, so a sweep that trusted the field would have deleted a checkout. Uncommitted work is stashed to `autopilot/abandoned/<task-id>` and the worktree is kept, because losing a day of edits is worse than 750 MB.
76
+
77
+ And the rule that decides what the tool is for: **a worktree with no run state is reported, never removed.** `.worktrees/` is not exclusively ours - this machine holds `174`, `pr-4051` and `task-1` there, hand-made, one in active use - and nothing distinguishes those from a pipeline worktree whose log was pruned. They are the 11 GB, and naming them with their sizes is worth more than a rule that guesses.
78
+
79
+ Dry-run by default, `--yes` applies, state salvaged to `artifacts/` before anything goes.
80
+
81
+ ### Fixed
82
+
83
+ - **The visual-evidence chain had five readers and no writer, so none of it ever ran.** Phase 0 Step 7.7 is gated on `state.visualEvidence.required`; Phase 3 captures when it is true; Phase 5 records the flow video; Phase 6 blocks on a required artefact that is neither attached nor explained. No phase document ever wrote `required`, `requiredBy` or `platform`, so the verdict was never true, and every consumer below it read a contract that looked satisfied. Step 7.7 now writes the verdict and is named in `features/visual-evidence.md` section 1a as its only writer. Phase 0 has no diff, so the `bugfix` row is provisional and Phase 3 re-decides from the real one.
84
+
85
+ - **The probe was called with `--platform "$PLATFORM"`, a variable no phase document assigns.** The probe refuses an empty value with exit 2, so the call could not have worked on its best day - it went unnoticed only because the step above was unreachable. The platform now comes from the stack (section 1b), which is the right source: it selects which device tooling to probe, and that is a property of the repo rather than of the diff. iOS wins a tie, recorded in `stackWhy`. A repo with no device platform does not get an empty string passed through; the probe does not run and `evidenceCapability.skippedReason` says why.
86
+
87
+ - `smoke-evidence-probe.sh` now fails a `--platform` argument whose variable is not assigned in the same file. Deliberately narrow: more than a hundred variables appear in these documents without an assignment and most are fine, because an agent fills `$WORKTREE` from context. `--platform` is not inferable and its consumer enforces a closed set with a hard exit, which is what makes the rule about this flag rather than about shell hygiene.
88
+
89
+ ### Fixed
90
+
91
+ - **Nearly every checkout on the development machine was routed to the iOS toolkit** - 22 of 25, measured. Two causes, measured: `setup` wrote `ai-ios-toolkit` whenever its marker list matched nothing, and nothing else ever revisited the answer, so a Next.js site and several Node CLIs inherited it. The default is gone; an undetected repo now gets the two stack-independent toolkits and a recorded reason, which is an answer rather than a guess.
92
+
93
+ - **Android was never detected.** Phase 1's marker scan ran at `maxdepth 2`, and `AndroidManifest.xml` sits at `<module>/src/main/` - depth 4 in every multi-module app. The scan also read a generated `.next/package.json` as a repo's manifest, and reported a Compose app as iOS because the app vendors a submodule that ships a `Package.swift`.
94
+
95
+ ### Added
96
+
97
+ - **`pipeline/lib/stack-detect.sh`** is now the one owner of "what is this repo built with". Deterministic file markers, never a model call: root before anything deeper so `find` traversal order cannot decide the answer, submodule and build-output pruning, depth 5 for the Android manifest, and `MA_STACK_WHY` separating "no marker matched" from "could not be read" - a caller that cannot tell those apart treats an unreadable repo as a language-free one. Verified against every repo on the development machine and seven fixtures.
98
+
99
+ - **`pluginsForStacks()` in `scripts/_stack-routing.mjs`** is the one owner of "which plugins does that want". The mapping was written out by hand in four places, each a chance to get the single asymmetric name wrong - the detector says `web`, the marketplace ships `ai-frontend-toolkit`. Derived names are checked against the locally installed marketplace manifest, so a toolkit that was never shipped is caught at derivation instead of presenting as one silently missing from the session. An unreadable manifest routes by convention and reports `manifestVerified: false`; it never fails closed.
100
+
101
+ - `state.stacks[]` and `state.stackWhy` for the stack axis. `state.detectedStack[]` stays the LANGUAGE axis and is now declared in the schema, which had `additionalProperties: false` and no such property while the phase documents instructed the write. The two axes are deliberately not merged: `graph-build.mjs --stack` accepts `ios|android|node|python|go` and has no notion of `web` or `backend`, and conflating them is what let a Gradle-built JVM service read as an Android app.
102
+
103
+ - **`smoke-stack-detect.sh`**, whose golden-output assertion is what made the consolidation safe: every argument the four hand-written tables accepted must still resolve to exactly the plugins they listed, or the gate fails.
104
+
105
+ ### Removed
106
+
107
+ - **Linux and Windows support.** The package advertised three platforms and verified one. Every credential read shells `security`, every iOS build `xcodebuild`, every piece of visual evidence `simctl`, so a run elsewhere did not degrade into a smaller run - it failed partway through with a worktree and a branch already created. ADR-0011 already recorded that CI had not executed since 2026-07-02, and its own status block called the Linux path "smoke-skipped" and the Windows path "unverified".
108
+
109
+ The removal is a declaration, not a deletion. `package.json` declares `os: ["darwin"]`, which npm enforces on the root package: a non-macOS `npm ci` ends in `EBADPLATFORM` before a file is written. `index.js` and `install/index.mjs` refuse a non-darwin host with one line naming the requirement; `uninstall`, `help` and `--version` stay reachable, because somebody who installed before this gate must still be able to remove it. `MULTI_AGENT_ALLOW_NON_DARWIN=1` overrides and says on stderr that nothing on that path is tested.
110
+
111
+ The dead Linux and Windows branches are left in place deliberately. They cost nothing at runtime, and seven constructs in them look like cross-platform scaffolding while being load-bearing on macOS - the `grep -P` ban, the `sha256sum || shasum` ordering, the `stat -c || stat -f` chains among them. ADR-0012 enumerates all seven so the follow-up cleanup does not break the BSD path.
112
+
113
+ - `pipeline/scripts/gate-linux.sh` and the `npm run gate:linux` script. Every workflow now runs on a macOS runner; `test.yml` loses its Ubuntu leg and its `windows-test` job, since with `os: ["darwin"]` enforced those jobs could no longer `npm ci` this tree at all.
114
+
115
+ ### Added
116
+
117
+ - **`doctor` check `host-platform`**, which runs first and blocks. First, so an unsupported host reads one honest line instead of a cascade whose common cause is never stated. The blocking set grows from five to six and is documented in `features/doctor.md`.
118
+
119
+ - **`smoke-macos-only.sh`** asserts each declaration separately, because declarations rot silently: the `os` field, both entry gates, the uninstall exemption, the doctor registration and its first position, that no user-facing doc promises another platform, and that no workflow targets a runner that cannot install this package.
120
+
121
+ ### Fixed
122
+
123
+ - **`smoke-doctor.sh` computed the documented blocking set and never compared it to anything**, so the ref and the code were free to disagree - and they did, the moment `host-platform` was added. Worse, the computation used `sed -n '/a\|b/p'`, and BSD sed reads `\|` as a literal pipe, so the value had been empty on every macOS run since it was written. Both halves are fixed and the comparison is now an assertion.
124
+
125
+ - **The step-verb mutation test mutated the first matching literal in `doctor.mjs`'s source**, which stopped being a step the probe run reaches as soon as a check that passes on macOS was registered above `install-present`: the engine was never asked, and the gate reported it had gone soft. The step to mutate is now read back from the probe run's own output, so it stays unpinned to any particular check while being guaranteed to execute.
126
+
17
127
  ## [16.31.1] - 2026-09-13
18
128
 
19
129
  ### Fixed
package/README.md CHANGED
@@ -10,7 +10,7 @@
10
10
 
11
11
  An 8-phase AI development pipeline for **Claude Code**, **Copilot CLI** and **Codex CLI**. Drives a Jira issue or GitHub URL to a merged PR in one command - analysis → plan → TDD → review → test → commit → PR - with multi-repo orchestration, a plan-approval gate, CLI-aware parallel review, and store-compliance checks. Component and Figma-to-code work is dispatched to the per-stack marketplace plugins (iOS/SwiftUI, Android/Compose) rather than bundled, so component skills live in one place.
12
12
 
13
- Runs natively on Claude Code, Copilot CLI and Codex CLI. macOS / Linux / Windows. Zero runtime dependencies.
13
+ Runs natively on Claude Code, Copilot CLI and Codex CLI. macOS only. Zero runtime dependencies.
14
14
 
15
15
  📐 **[Architecture diagrams](./docs/architecture.md)** - the 8-phase flow, operating modes, review/triage, Figma subphases, component layout. **[Ecosystem diagram](./docs/ecosystem.md)** - how this repo, the `multi-agent-plugins` marketplace and `multi-agent-toolkit-mcp` compose.
16
16
 
@@ -46,7 +46,7 @@ Run a task - the input type is auto-detected:
46
46
  /multi-agent:issue # browse unassigned GitHub issues → pick
47
47
  ```
48
48
 
49
- Every input runs the same short intake - **account → (repo) → maturity check → dev-context** - then enters Phase 0. A Jira id or GitHub URL is fetched and maturity-checked *before* any code is written; free-text skips the fetch and goes straight to planning. Multi-repo tasks add extra repos at the dev-context step.
49
+ Every input runs the same short intake - **account → (repo) → maturity check → dev-context** - then enters Phase 0. A Jira id or GitHub URL is fetched and maturity-checked _before_ any code is written; free-text skips the fetch and goes straight to planning. Multi-repo tasks add extra repos at the dev-context step.
50
50
 
51
51
  Add `autopilot` to skip confirmations, or `--local` to work on the current branch without a worktree (e.g. `/multi-agent:autopilot "PROJ-1234"`). Pipeline depth is a question the run asks, not a flag: `/multi-agent` and `/multi-agent:local` offer Full or Short at Phase 0.
52
52
 
@@ -75,117 +75,117 @@ The discipline behind all of this - bounded loops, evidence gates, token-budgete
75
75
 
76
76
  ## Modes
77
77
 
78
- | Mode | Command | Flow |
79
- |---|---|---|
80
- | Full | `/multi-agent "task"` | All 8 phases, interactive |
81
- | Autopilot | `/multi-agent:autopilot "task"` | 7 phases (interactive Test gate dropped), no confirmations |
82
- | Local | `/multi-agent:local "task"` | Full pipeline minus the interactive Test gate, current branch (no worktree) |
83
- | Depth | asked at Phase 0 Step 7.5 | Full (all phases) or Short (Dev → Review → Test → Commit → Report). Not a command name - `/multi-agent` and `:local` ask, both autopilot entries always run Full |
84
- | Ship | `/multi-agent:resume-local` | Run the review→test→commit→report tail over local work |
85
- | Audit | `/multi-agent:design-check` | Mock-mode vs Figma conformance, local-only |
86
- | Audit | `/multi-agent:testflight-validation` | Pre-submission gates for a TestFlight build: static archive audit → Apple's `altool --validate-app` → Review-Guidelines check. Validates only, never uploads |
78
+ | Mode | Command | Flow |
79
+ | --------- | ------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- |
80
+ | Full | `/multi-agent "task"` | All 8 phases, interactive |
81
+ | Autopilot | `/multi-agent:autopilot "task"` | 7 phases (interactive Test gate dropped), no confirmations |
82
+ | Local | `/multi-agent:local "task"` | Full pipeline minus the interactive Test gate, current branch (no worktree) |
83
+ | Depth | asked at Phase 0 Step 7.5 | Full (all phases) or Short (Dev → Review → Test → Commit → Report). Not a command name - `/multi-agent` and `:local` ask, both autopilot entries always run Full |
84
+ | Ship | `/multi-agent:resume-local` | Run the review→test→commit→report tail over local work |
85
+ | Audit | `/multi-agent:design-check` | Mock-mode vs Figma conformance, local-only |
86
+ | Audit | `/multi-agent:testflight-validation` | Pre-submission gates for a TestFlight build: static archive audit → Apple's `altool --validate-app` → Review-Guidelines check. Validates only, never uploads |
87
87
 
88
88
  Depth, autopilot and `--local` are the only knobs on the run itself; everything else is its own command. The full catalog is below.
89
89
 
90
90
  ## Commands
91
91
 
92
- `/multi-agent` plus 53 sub-commands. `/multi-agent:help` renders the same catalog in your terminal, in your `outputLanguage`.
92
+ `/multi-agent` plus 56 sub-commands. `/multi-agent:help` renders the same catalog in your terminal, in your `outputLanguage`.
93
93
 
94
94
  ### Pipeline entries
95
95
 
96
- | Command | What it does |
97
- |---|---|
98
- | `/multi-agent "task"` | Full pipeline in a worktree. Asks Full or Short depth at Phase 0 |
99
- | `/multi-agent:local "task"` | Same pipeline on the current branch, no worktree |
100
- | `/multi-agent:autopilot "task"` | Worktree, no confirmations, always Full |
101
- | `/multi-agent:local-autopilot "task"` | Current branch, no confirmations, always Full |
102
- | `/multi-agent:resume-local` | Pipeline tail over work already done locally: Review → Build+Test → Commit/PR → Report. No dev phase |
96
+ | Command | What it does |
97
+ | ------------------------------------- | ---------------------------------------------------------------------------------------------------- |
98
+ | `/multi-agent "task"` | Full pipeline in a worktree. Asks Full or Short depth at Phase 0 |
99
+ | `/multi-agent:local "task"` | Same pipeline on the current branch, no worktree |
100
+ | `/multi-agent:autopilot "task"` | Worktree, no confirmations, always Full |
101
+ | `/multi-agent:local-autopilot "task"` | Current branch, no confirmations, always Full |
102
+ | `/multi-agent:resume-local` | Pipeline tail over work already done locally: Review → Build+Test → Commit/PR → Report. No dev phase |
103
103
 
104
104
  ### Task control
105
105
 
106
- | Command | What it does |
107
- |---|---|
108
- | `/multi-agent:status` | Every task's ID, phase, branch and state |
109
- | `/multi-agent:log [#N]` | Show a task's `agent-log.md` (most recent by default) |
110
- | `/multi-agent:resume [#N]` | Carry a stopped or failed task on from its last phase |
111
- | `/multi-agent:kill [#N]` | Stop a task, remove its worktree and branch |
112
- | `/multi-agent:steer #N "<instruction>"` | Correct a running task without stopping it; applied at the next phase boundary |
113
- | `/multi-agent:search` | Ranked search across every task log; `--semantic` queries the triage corpus |
114
- | `/multi-agent:garbage-collect` | Sweep leftover scratch, orphan worktrees and offloaded payloads. Dry-run first |
115
- | `/multi-agent:prune-logs` | Delete per-task logs by age / project / task. Audit trail and metrics kept |
116
- | `/multi-agent:purge` | Wipe every worktree, branch, log and state file. Double confirmation |
106
+ | Command | What it does |
107
+ | --------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- |
108
+ | `/multi-agent:status` | Every task's ID, phase, branch and state |
109
+ | `/multi-agent:log [#N]` | Show a task's `agent-log.md` (most recent by default) |
110
+ | `/multi-agent:resume [#N]` | Carry a stopped or failed task on from its last phase |
111
+ | `/multi-agent:kill [#N]` | Stop a task, remove its worktree and branch |
112
+ | `/multi-agent:steer #N "<instruction>"` | Correct a running task without stopping it; applied at the next phase boundary |
113
+ | `/multi-agent:search` | Ranked search across every task log; `--semantic` queries the triage corpus |
114
+ | `/multi-agent:garbage-collect` | Sweep leftover scratch, orphan worktrees and offloaded payloads. `--abandoned` also reaps runs that stopped and were never picked back up. Dry-run first |
115
+ | `/multi-agent:prune-logs` | Delete per-task logs by age / project / task. Audit trail and metrics kept |
116
+ | `/multi-agent:purge` | Wipe every worktree, branch, log and state file. Double confirmation |
117
117
 
118
118
  ### Review
119
119
 
120
- | Command | What it does |
121
- |---|---|
122
- | `/multi-agent:review` | Parallel review of a branch diff or a PR; inline comments + approve/needs-work on PR input |
123
- | `/multi-agent:review-jira` | Grade a Jira issue's readiness for development, comment the gaps |
124
- | `/multi-agent:review-issue` | Same grading for a GitHub issue |
125
- | `/multi-agent:review-analysis` | Review a written analysis document; findings cite the Locked rule they break |
126
- | `/multi-agent:diff-explain` | Map a Phase 4 triage finding back to the diff lines that caused it |
127
- | `/multi-agent:refactor` | Best-practice extraction + bug hunt + derived-skill drift + toolkit MCP research → one plan |
128
- | `/multi-agent:scan` | Skill security scan of local skill directories against a tiered pattern catalog |
129
- | `/multi-agent:prune-prompts` | Zero-base review of the always-on instruction footprint; keep / trial / delete per rule |
130
- | `/multi-agent:ios-coding-standard` | Audit an iOS module against the 99-rule registry, produce a remediation plan |
120
+ | Command | What it does |
121
+ | ---------------------------------- | ------------------------------------------------------------------------------------------- |
122
+ | `/multi-agent:review` | Parallel review of a branch diff or a PR; inline comments + approve/needs-work on PR input |
123
+ | `/multi-agent:review-jira` | Grade a Jira issue's readiness for development, comment the gaps |
124
+ | `/multi-agent:review-issue` | Same grading for a GitHub issue |
125
+ | `/multi-agent:review-analysis` | Review a written analysis document; findings cite the Locked rule they break |
126
+ | `/multi-agent:diff-explain` | Map a Phase 4 triage finding back to the diff lines that caused it |
127
+ | `/multi-agent:refactor` | Best-practice extraction + bug hunt + derived-skill drift + toolkit MCP research → one plan |
128
+ | `/multi-agent:scan` | Skill security scan of local skill directories against a tiered pattern catalog |
129
+ | `/multi-agent:prune-prompts` | Zero-base review of the always-on instruction footprint; keep / trial / delete per rule |
130
+ | `/multi-agent:ios-coding-standard` | Audit an iOS module against the 99-rule registry, produce a remediation plan |
131
131
 
132
132
  ### Analysis
133
133
 
134
- | Command | What it does |
135
- |---|---|
136
- | `/multi-agent:analysis` | Standalone feature spec: global (23-section handoff) or corporate (IG/UC/FG) profile |
137
- | `/multi-agent:analysis-resolve` | Answer an analysis doc's Section 20 open questions one row at a time |
134
+ | Command | What it does |
135
+ | --------------------------------- | ----------------------------------------------------------------------------------------- |
136
+ | `/multi-agent:analysis` | Standalone feature spec: global (23-section handoff) or corporate (IG/UC/FG) profile |
137
+ | `/multi-agent:analysis-resolve` | Answer an analysis doc's Section 20 open questions one row at a time |
138
138
  | `/multi-agent:complaint-analysis` | Customer-complaint triage with Graylog evidence: client / bff root cause, or core routing |
139
139
 
140
140
  ### Testing on a device
141
141
 
142
- | Command | What it does |
143
- |---|---|
144
- | `/multi-agent:test` | UI Bug Hunter on a booted simulator or emulator: screenshot, tap, analyze |
145
- | `/multi-agent:test-dark-mode` | Walk every screen light then dark, report contrast and colour bugs |
146
- | `/multi-agent:test-accessibility` | VoiceOver labels, sub-44pt tap targets, contrast, traits |
147
- | `/multi-agent:test-dynamic-type` | Re-walk every screen at XL through accessibility-XL, report truncation |
148
- | `/multi-agent:test-screenshots [locale]` | App Store screenshot set in a locale (defaults to `tr`) |
149
- | `/multi-agent:manual-test` | Phase 5 standalone: check out the task branch and prepare it for Xcode |
142
+ | Command | What it does |
143
+ | ---------------------------------------- | ------------------------------------------------------------------------- |
144
+ | `/multi-agent:test` | UI Bug Hunter on a booted simulator or emulator: screenshot, tap, analyze |
145
+ | `/multi-agent:test-dark-mode` | Walk every screen light then dark, report contrast and colour bugs |
146
+ | `/multi-agent:test-accessibility` | VoiceOver labels, sub-44pt tap targets, contrast, traits |
147
+ | `/multi-agent:test-dynamic-type` | Re-walk every screen at XL through accessibility-XL, report truncation |
148
+ | `/multi-agent:test-screenshots [locale]` | App Store screenshot set in a locale (defaults to `tr`) |
149
+ | `/multi-agent:manual-test` | Phase 5 standalone: check out the task branch and prepare it for Xcode |
150
150
 
151
151
  ### Design, build and store
152
152
 
153
- | Command | What it does |
154
- |---|---|
155
- | `/multi-agent:design-check` | Mock-mode vs Figma audit with a coverage gate; annotated HTML + PDF report |
156
- | `/multi-agent:store-ready` | Pre-submission gates for iOS and Android: package audit, store validation, policy review |
157
- | `/multi-agent:testflight-validation` | iOS-pinned alias of `store-ready`. Validates only, never uploads |
158
- | `/multi-agent:build-optimize` | Benchmark an Xcode build, run the analyzers, produce a recommend-first plan |
153
+ | Command | What it does |
154
+ | ------------------------------------ | ---------------------------------------------------------------------------------------- |
155
+ | `/multi-agent:design-check` | Mock-mode vs Figma audit with a coverage gate; annotated HTML + PDF report |
156
+ | `/multi-agent:store-ready` | Pre-submission gates for iOS and Android: package audit, store validation, policy review |
157
+ | `/multi-agent:testflight-validation` | iOS-pinned alias of `store-ready`. Validates only, never uploads |
158
+ | `/multi-agent:build-optimize` | Benchmark an Xcode build, run the analyzers, produce a recommend-first plan |
159
159
 
160
160
  ### Tickets and reporting
161
161
 
162
- | Command | What it does |
163
- |---|---|
164
- | `/multi-agent:jira` | Browse your open Jira issues → pick → branch → mode → launch |
165
- | `/multi-agent:issue` | Browse unassigned GitHub issues → pick → auto-assign → launch |
166
- | `/multi-agent:create-jira` | Draft a Task / Bug / Story to the project's own conventions, preview before create |
167
- | `/multi-agent:channels` | Post the multi-channel report: Jira, Confluence, Wiki, PR description, board status |
168
- | `/multi-agent:feedback` | Send one message to the maintainer. Only your text is sent, no logs or paths |
162
+ | Command | What it does |
163
+ | -------------------------- | ----------------------------------------------------------------------------------- |
164
+ | `/multi-agent:jira` | Browse your open Jira issues → pick → branch → mode → launch |
165
+ | `/multi-agent:issue` | Browse unassigned GitHub issues → pick → auto-assign → launch |
166
+ | `/multi-agent:create-jira` | Draft a Task / Bug / Story to the project's own conventions, preview before create |
167
+ | `/multi-agent:channels` | Post the multi-channel report: Jira, Confluence, Wiki, PR description, board status |
168
+ | `/multi-agent:feedback` | Send one message to the maintainer. Only your text is sent, no logs or paths |
169
169
 
170
170
  ### Your own routines
171
171
 
172
- | Command | What it does |
173
- |---|---|
174
- | `/multi-agent:save [name]` | Save a recurring job as a reusable `/multi-agent:<name>`. Local-only, never synced |
175
- | `/multi-agent:routines` | List your saved routines and what each does |
176
- | `/multi-agent:forget [name]` | Remove a saved routine and its registry entry |
172
+ | Command | What it does |
173
+ | ---------------------------- | ---------------------------------------------------------------------------------- |
174
+ | `/multi-agent:save [name]` | Save a recurring job as a reusable `/multi-agent:<name>`. Local-only, never synced |
175
+ | `/multi-agent:routines` | List your saved routines and what each does |
176
+ | `/multi-agent:forget [name]` | Remove a saved routine and its registry entry |
177
177
 
178
178
  ### Setup and maintenance
179
179
 
180
- | Command | What it does |
181
- |---|---|
182
- | `/multi-agent:setup` | First-run wizard: keychain token discovery, git identity, pipeline preparation |
183
- | `/multi-agent:stack [ids]` | Enable the marketplace plugin(s) for this repo. Multi-select |
184
- | `/multi-agent:language [en\|tr]` | Show or set `outputLanguage`; `promptLanguage` stays English |
185
- | `/multi-agent:sync` | One-shot sync: Claude Code, Copilot CLI, pipeline repo, website, toolkit MCP |
186
- | `/multi-agent:update` | Update to the latest published npm release and run migrations |
187
- | `/multi-agent:uninstall` | Remove the pipeline from every CLI. Keychain tokens always left intact |
188
- | `/multi-agent:help` | This catalog, in the terminal, in your `outputLanguage` |
180
+ | Command | What it does |
181
+ | -------------------------------- | ------------------------------------------------------------------------------ |
182
+ | `/multi-agent:setup` | First-run wizard: keychain token discovery, git identity, pipeline preparation |
183
+ | `/multi-agent:stack [ids]` | Enable the marketplace plugin(s) for this repo. Multi-select |
184
+ | `/multi-agent:language [en\|tr]` | Show or set `outputLanguage`; `promptLanguage` stays English |
185
+ | `/multi-agent:sync` | One-shot sync: Claude Code, Copilot CLI, pipeline repo, website, toolkit MCP |
186
+ | `/multi-agent:update` | Update to the latest published npm release and run migrations |
187
+ | `/multi-agent:uninstall` | Remove the pipeline from every CLI. Keychain tokens always left intact |
188
+ | `/multi-agent:help` | This catalog, in the terminal, in your `outputLanguage` |
189
189
 
190
190
  ### Subagents
191
191
 
@@ -195,6 +195,51 @@ Eight are installed alongside the commands and dispatched by the phases: `explor
195
195
 
196
196
  Two compliance skills install on every host and back the store gates: `apple-archive-compliance` (18-rule Apple review scan with ITMS code mapping) and `google-play-compliance` (21-rule Play policy catalog with Console error codes). Everything else stack-shaped - SwiftUI, Compose, backend, frontend - comes from the marketplace plugins described below.
197
197
 
198
+ ## Continuous mode
199
+
200
+ Everything above starts when you start it. Continuous mode is the same pipeline
201
+ picking work up on its own, on ONE machine you choose, from repos you choose.
202
+
203
+ | Command | What it does |
204
+ | --- | --- |
205
+ | `/multi-agent:autopilot-on` | Pick the repos this machine watches. Labelled GitHub issues and assigned + labelled Jira items then run in a worktree and stop at an open PR. Re-run to change the list |
206
+ | `/multi-agent:autopilot-status` | What is running and at which phase, what is queued, what is waiting for an answer, the PRs of the last day, and the rolling spend |
207
+ | `/multi-agent:autopilot-off` | Remove the schedule. Work already running finishes; the repo selection is kept |
208
+
209
+ **Nothing is on by default and nothing is added implicitly.** Installing the
210
+ package writes no state and schedules nothing; `smoke-autopilot-default-off.sh`
211
+ fails the build if that ever changes. A label is a filter, not a gate - anyone
212
+ who can open an issue in a repo you have push on could add one - so the gate is
213
+ the picker, and it is per machine.
214
+
215
+ The biggest win is not parallelism. Measured here, the median run is 44 minutes,
216
+ but an item that finishes at 14:00 waits until you sit down again: overnight that
217
+ is 16 hours against 40 minutes. Continuous mode removes the waiting, not the work.
218
+
219
+ **What it will not do.** It does not merge - the runner stops at an open PR and
220
+ the decision stays yours. It does not touch an attended run: per-repo concurrency
221
+ is always 1, so the queue steps around a repo you are working in rather than
222
+ competing for `.git/index.lock`. There is no cap on PRs; the bounds are
223
+ `costCeilingUsd` over a rolling 24 hours and what the machine can hold.
224
+
225
+ **A menu bar indicator**, when `swiftc` is present, draws the same `status.json`
226
+ in the top right and refreshes on its own: one row per item with its id, Full or
227
+ Short, the phase as a fraction, the elapsed time and the stack. A row disappears
228
+ the moment the item finishes and reappears under Reports with its PR. It only
229
+ draws - it cannot start, stop or change a run. ActivityKit is unavailable on
230
+ macOS, so this is an `NSStatusItem`, built from source on demand rather than
231
+ shipped as a binary that would need signing.
232
+
233
+ **It survives a restart with no command to run.** launchd loads the job at
234
+ **login**, not at boot, and that is correct rather than a limitation: the login
235
+ keychain is what unlocks the tokens, so a tick that fired before login could not
236
+ reach Jira or GitHub anyway. There is no `autopilot-resume` - a command you have
237
+ to remember is a queue that silently stops when you forget it. `doctor` reports
238
+ the real failure instead: configured, but launchd holds no job.
239
+
240
+ Sleep is held **only on AC**. A queue that flattens a laptop off the charger is a
241
+ bug; on battery the assertion is released and work resumes when you plug in.
242
+
198
243
  ## Stacks
199
244
 
200
245
  Stack skills ship as versioned plugins in the [`mmerterden/multi-agent-plugins`](https://github.com/mmerterden/multi-agent-plugins) marketplace. Select a stack per-repo:
@@ -207,13 +252,13 @@ This enables the matching plugin (+ the shared `ai-common` plugin) in the repo's
207
252
 
208
253
  ## Tool support
209
254
 
210
- The pipeline runs natively on **Claude Code**, **Copilot CLI** and **Codex CLI** - all three install from the same `pipeline/` source and get the same 53 commands.
255
+ The pipeline runs natively on **Claude Code**, **Copilot CLI** and **Codex CLI** - all three install from the same `pipeline/` source and get the same 56 commands.
211
256
 
212
- | Tool | Flag | What it installs |
213
- |---|---|---|
257
+ | Tool | Flag | What it installs |
258
+ | ----------- | -------------------- | ------------------------------------------------------------------------------------------------------ |
214
259
  | Claude Code | `--claude` (default) | slash commands + skills + agents + three `PreToolUse` hooks (secret scan, agent-guard, read-size gate) |
215
- | Copilot CLI | `--copilot` | instructions + 53 sub-command skills + scripts |
216
- | Codex CLI | `--codex` | one router skill + 53 specs as refs + 8 agent TOML + `AGENTS.md` block + `codex mcp add` |
260
+ | Copilot CLI | `--copilot` | instructions + 56 sub-command skills + scripts |
261
+ | Codex CLI | `--codex` | one router skill + 56 specs as refs + 8 agent TOML + `AGENTS.md` block + `codex mcp add` |
217
262
 
218
263
  Filter skills by stack with `--platform=ios\|android\|all`.
219
264
 
@@ -232,20 +277,20 @@ the one host whose panel spans two vendors, and the triage note says so.
232
277
 
233
278
  ## Tokens & integrations
234
279
 
235
- `setup` scans your OS keychain and maps each token by a **logical name** (e.g. `jira`) to its real keychain entry - the pipeline resolves tokens through that mapping (`credential-store.sh`), so literal keychain names never appear in synced files. Tokens stay in the keychain (macOS Keychain / Windows Credential Manager / Linux libsecret), are **never committed or logged**, and are all **optional** - the pipeline asks for any it needs at Phase 0.
236
-
237
- | Token | Used for | Phase |
238
- |---|---|---|
239
- | `jira` | fetch the issue · post the report comment | 0, 7 |
240
- | `github` | issues · PRs · `gh` auth | 0, 6 |
241
- | `bitbucket` | PR create/update (reviewer-preserving) · diff | 6 |
242
- | `confluence` | publish analysis / wiki pages | 7 |
243
- | `figma` + `figma_mcp` | fetch design context | analysis only |
244
- | `fortify` | security-scan findings gate | 4 |
245
- | `firebase` | Firebase service-account JSON for Firebase projects | as needed |
246
- | `jenkins` | CI trigger / status | build / deploy |
247
- | `npm` | package publish (mostly CI) | release |
248
- | `appstore_connect_*` | TestFlight / App Store pre-submission validation (optional, iOS) | `testflight-validation` |
280
+ `setup` scans your OS keychain and maps each token by a **logical name** (e.g. `jira`) to its real keychain entry - the pipeline resolves tokens through that mapping (`credential-store.sh`), so literal keychain names never appear in synced files. Tokens stay in the macOS Keychain, are **never committed or logged**, and are all **optional** - the pipeline asks for any it needs at Phase 0.
281
+
282
+ | Token | Used for | Phase |
283
+ | --------------------- | ---------------------------------------------------------------- | ----------------------- |
284
+ | `jira` | fetch the issue · post the report comment | 0, 7 |
285
+ | `github` | issues · PRs · `gh` auth | 0, 6 |
286
+ | `bitbucket` | PR create/update (reviewer-preserving) · diff | 6 |
287
+ | `confluence` | publish analysis / wiki pages | 7 |
288
+ | `figma` + `figma_mcp` | fetch design context | analysis only |
289
+ | `fortify` | security-scan findings gate | 4 |
290
+ | `firebase` | Firebase service-account JSON for Firebase projects | as needed |
291
+ | `jenkins` | CI trigger / status | build / deploy |
292
+ | `npm` | package publish (mostly CI) | release |
293
+ | `appstore_connect_*` | TestFlight / App Store pre-submission validation (optional, iOS) | `testflight-validation` |
249
294
 
250
295
  The **secret scan** runs as a `PreToolUse` hook on Claude Code (hard-blocks a commit on a hit) and as a pre-push check elsewhere.
251
296
 
@@ -261,13 +306,13 @@ Uninstall preserves this layer: the tokens, the reader that opens them, the mapp
261
306
 
262
307
  ## Platform support
263
308
 
264
- Runs on **macOS**, **Linux**, and **Windows** (Git Bash / WSL). Shell and credential access go through a platform-agnostic layer - the keychain resolves automatically to **macOS Keychain**, **Linux libsecret** (`secret-tool`), or **Windows Credential Manager**, and scripts fall back between BSD and GNU tool variants. Node.js 20.11+ (tested on 20 and 22).
309
+ Runs on **macOS** only. The package declares `os: ["darwin"]`, so `npm` refuses to install it elsewhere rather than letting a run fail halfway through: every credential read shells `security`, every iOS build `xcodebuild`, every piece of visual evidence `simctl`. Node.js 20.11+ (tested on 20 and 22). Reasoning: [ADR-0012](docs/adr/0012-macos-only.md).
265
310
 
266
311
  ## Companion repos
267
312
 
268
- | Repo | What it is |
269
- |---|---|
270
- | [`mmerterden/multi-agent-plugins`](https://github.com/mmerterden/multi-agent-plugins) | Marketplace of per-stack skill toolkits (iOS / Android / Frontend / Backend + common). `/multi-agent:stack` enables the matching plugin. |
313
+ | Repo | What it is |
314
+ | --------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
315
+ | [`mmerterden/multi-agent-plugins`](https://github.com/mmerterden/multi-agent-plugins) | Marketplace of per-stack skill toolkits (iOS / Android / Frontend / Backend + common). `/multi-agent:stack` enables the matching plugin. |
271
316
  | [`mmerterden/multi-agent-toolkit-mcp`](https://github.com/mmerterden/multi-agent-toolkit-mcp) | MCP server for UI testing / simulator capture / xcodebuild - powers the Phase 5 UI Bug Hunter. Published on the public npm registry as [`@mmerterden/multi-agent-toolkit-mcp`](https://www.npmjs.com/package/@mmerterden/multi-agent-toolkit-mcp); the installer registers it with each CLI for you, so `npx` resolves it with no extra configuration. |
272
317
 
273
318
  ## License