@mmerterden/multi-agent-pipeline 16.31.1 → 17.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +110 -0
- package/README.md +148 -103
- package/README.tr.md +149 -103
- package/docs/adr/0011-dormant-ci.md +10 -1
- package/docs/adr/0012-macos-only.md +98 -0
- package/docs/adr/README.md +1 -0
- package/docs/architecture.md +3 -3
- package/docs/ecosystem.md +5 -5
- package/docs/engineering.md +1 -1
- package/index.js +26 -0
- package/install/_dev-only-files.mjs +0 -1
- package/install/index.mjs +10 -0
- package/install/templates/multi-agent-autopilot.plist.template +79 -0
- package/package.json +5 -3
- package/pipeline/commands/multi-agent/autopilot-off/SKILL.md +64 -0
- package/pipeline/commands/multi-agent/autopilot-on/SKILL.md +173 -0
- package/pipeline/commands/multi-agent/autopilot-status/SKILL.md +74 -0
- package/pipeline/commands/multi-agent/channels/SKILL.md +41 -12
- package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +40 -3
- package/pipeline/commands/multi-agent/help/SKILL.md +43 -37
- package/pipeline/commands/multi-agent/manual-test/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/setup/SKILL.md +15 -7
- package/pipeline/commands/multi-agent/stack/SKILL.md +31 -32
- package/pipeline/commands/multi-agent/status/SKILL.md +17 -1
- package/pipeline/commands/multi-agent/sync/SKILL.md +34 -28
- package/pipeline/commands/multi-agent/update/SKILL.md +1 -1
- package/pipeline/lib/autopilot-activation.sh +117 -0
- package/pipeline/lib/autopilot-state.sh +150 -0
- package/pipeline/lib/issue-fetcher.sh +18 -1
- package/pipeline/lib/plan-todos.sh +18 -0
- package/pipeline/lib/stack-detect.sh +200 -0
- package/pipeline/multi-agent-refs/channels/jira.md +80 -20
- package/pipeline/multi-agent-refs/channels/pr.md +65 -19
- package/pipeline/multi-agent-refs/cross-cli-contract.md +35 -15
- package/pipeline/multi-agent-refs/features/doctor.md +15 -3
- package/pipeline/multi-agent-refs/features/visual-evidence.md +61 -1
- package/pipeline/multi-agent-refs/phases/phase-0-init.md +14 -5
- package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +26 -12
- package/pipeline/multi-agent-refs/phases/phase-2-planning.md +17 -15
- package/pipeline/multi-agent-refs/phases/phase-3-dev.md +7 -5
- package/pipeline/multi-agent-refs/phases/phase-6-commit.md +1 -1
- package/pipeline/multi-agent-refs/phases/phase-7-report.md +2 -3
- package/pipeline/multi-agent-refs/readiness-review.md +7 -1
- package/pipeline/multi-agent-refs/rules.md +3 -11
- package/pipeline/multi-agent-refs/tracker-contract.md +32 -0
- package/pipeline/schemas/agent-state.schema.json +99 -25
- package/pipeline/schemas/autopilot-config.schema.json +149 -0
- package/pipeline/schemas/token-budget.json +4 -4
- package/pipeline/scripts/_stack-routing.mjs +91 -0
- package/pipeline/scripts/autopilot-arming.mjs +147 -0
- package/pipeline/scripts/autopilot-intake.mjs +383 -0
- package/pipeline/scripts/autopilot-menubar.swift +361 -0
- package/pipeline/scripts/autopilot-runner.mjs +349 -0
- package/pipeline/scripts/autopilot-status.sh +212 -0
- package/pipeline/scripts/capture-resume.sh +76 -14
- package/pipeline/scripts/doctor.mjs +26 -3
- package/pipeline/scripts/gc-abandoned.sh +352 -0
- package/pipeline/scripts/jira-search.sh +70 -0
- package/pipeline/scripts/phase-tracker.sh +134 -12
- package/pipeline/scripts/probe-evidence-capability.sh +27 -3
- package/pipeline/scripts/run-ui-tests.sh +113 -4
- package/pipeline/scripts/usage-report.mjs +5 -5
- package/pipeline/skills/.skill-manifest.json +16 -4
- package/pipeline/skills/shared/core/multi-agent-autopilot-off/SKILL.md +67 -0
- package/pipeline/skills/shared/core/multi-agent-autopilot-on/SKILL.md +146 -0
- package/pipeline/skills/shared/core/multi-agent-autopilot-status/SKILL.md +64 -0
- package/pipeline/skills/shared/core/multi-agent-channels/SKILL.md +62 -11
- package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +9 -8
- package/pipeline/scripts/gate-linux.sh +0 -62
package/CHANGELOG.md
CHANGED
|
@@ -14,6 +14,116 @@ Internal file-layout changes that don't affect the slash-command surface are sti
|
|
|
14
14
|
|
|
15
15
|
---
|
|
16
16
|
|
|
17
|
+
## [17.1.0] - 2026-09-14
|
|
18
|
+
|
|
19
|
+
Continuous mode, and three things that were computed correctly and shown to
|
|
20
|
+
nobody: the plan's own steps, a web repo's UI tests, and the queue's status when
|
|
21
|
+
one field arrived as a number.
|
|
22
|
+
|
|
23
|
+
### Added
|
|
24
|
+
|
|
25
|
+
- **Continuous mode.** `/multi-agent:autopilot-on` picks the repos ONE machine watches; labelled GitHub issues and assigned + labelled Jira items then run in a worktree and stop at an open PR. `:autopilot-status` is the single producer the terminal, the menu bar indicator and the session hook all render, `:autopilot-off` removes the schedule and keeps the selection. Nothing is on by default: installing writes no state and schedules nothing, and `smoke-autopilot-default-off.sh` fails the build if that changes.
|
|
26
|
+
|
|
27
|
+
The ordering is deterministic with no model call - explicit rank, repo grouping, priority, then OLDEST first, which is the opposite of the `jira` picker on purpose: a human wants what just landed, an unattended queue must not starve what has been waiting. Two preconditions are checked before any item is taken: a rolling 24-hour spend total, and a per-source credential gate, so a dead Jira token stops Jira items without stopping GitHub ones.
|
|
28
|
+
|
|
29
|
+
A menu bar indicator draws the same `status.json` in the top right when `swiftc` is present. ActivityKit is `@available(macOS, unavailable)` - there is no Live Activity on a Mac - so this is an `NSStatusItem`, built from source on demand rather than shipped as a binary that would need signing.
|
|
30
|
+
|
|
31
|
+
- **UI tests on web.** `run-ui-tests.sh` rejected every platform but `ios|android`, so a repo with a full Playwright suite reported "no UI test target" and the PR body said UI tests had not run - a statement true of the runner and false of the repo. Detection keys on the browser-driving import (`@playwright/test`, `cy.visit(`), not on a directory called `e2e`: the same distinction that keeps 475 iOS snapshot tests from being counted as UI tests.
|
|
32
|
+
|
|
33
|
+
### Changed
|
|
34
|
+
|
|
35
|
+
- **The Jira comment and the PR body stopped being the same document.** Jira now carries Geliştirme Özeti, Test Senaryoları, Etki Analizi and Bağlantılar with the PR link on line 1 and no identifiers, file paths or diff hunks anywhere in it. The PR keeps all of that and gains Teknik Açıklama, Etki Analizi and Build - the build command, its result and the base sha it ran against. Given/When/Then is gone from the scenarios: it reads as translated English to the person running them, who wants a titled list they can follow with the app open.
|
|
36
|
+
|
|
37
|
+
### Fixed
|
|
38
|
+
|
|
39
|
+
- **The menu bar indicator read "autopilot kapalı" with items in flight.** `phase` was declared `String?` while the producer passes the queue's own value through `jq` untouched, so a numeric phase made Swift's `Decodable` throw - and the throw did not lose one field, it failed the whole document. A silent blank is the worst failure a status indicator has, because it is indistinguishable from good news.
|
|
40
|
+
|
|
41
|
+
- **`ma_ap_boottime` returned the microseconds.** `.*sec = ` is greedy and walks past `sec` into `usec`, so the helper produced a six-digit number that looked plausible and never equalled the same fact read anywhere else. It surfaced as a live runner being declared dead.
|
|
42
|
+
|
|
43
|
+
- **Three phase documents said "MCP forbidden" without qualifying it**, while the gate that enforces it has always matched `figma` and nothing else. The prose therefore banned the screenshot, xcodebuild and UI-test tools that Phase 3 itself calls, which is one way a run reaches Phase 7 with no evidence.
|
|
44
|
+
|
|
45
|
+
- **`channels/jira.md` told Phase 7 to upload evidence Phase 6 had already uploaded**, so re-rendering a comment attached every file a second time.
|
|
46
|
+
|
|
47
|
+
---
|
|
48
|
+
|
|
49
|
+
## [17.0.0] - 2026-09-14
|
|
50
|
+
|
|
51
|
+
Five things in this release were not working, and four of them looked like they
|
|
52
|
+
were. A platform we advertised and never verified; a stack answer that sent
|
|
53
|
+
nearly every repo to the iOS toolkit; an evidence step gated on a field no phase
|
|
54
|
+
wrote, passing an argument no phase assigned; and one word, "in progress",
|
|
55
|
+
covering a finished job, a crashed one and a run that never started. The common
|
|
56
|
+
shape is a contract with readers and no writer, which reads exactly like a
|
|
57
|
+
contract that is satisfied.
|
|
58
|
+
|
|
59
|
+
### Fixed
|
|
60
|
+
|
|
61
|
+
- **"In progress" covered three different situations and offered one answer.** Measured: 20 runs read `in_progress` and every one was over a day old, the oldest 157 days. They were not one failure. 3 had their PR already open and were holding at Phase 6/7, where the pipeline pauses BY DESIGN for channel selection - finished work. 11 were left at a Phase 0 question, which is almost entirely questions, so nothing was ever built. 6 stopped mid-development, the only group that is actually broken. The session hook reported the newest of the 20 as "stopped at Phase N" and said nothing about the other 19, in the same words whether the pipeline was waiting for the user or had crashed.
|
|
62
|
+
|
|
63
|
+
`agent-state` gains `awaiting_input`, distinct from both `in_progress` and `paused`: the run is not broken and not resumable by a retry, it is finished with what it can do alone. A state file in the wild had already invented `awaiting-user-test-main-checkout` to say this. Phase 7 writes it where it pauses, the session hook and `/multi-agent:status` group on it, and each group offers its own action - resuming a Phase 0 question rebuilds nothing, and garbage-collecting an open PR throws away landed work.
|
|
64
|
+
|
|
65
|
+
A run with no status is now left out of all three groups. Eight such files exist here, six with no phase either, and calling them dead is the same false claim as calling a waiting run dead.
|
|
66
|
+
|
|
67
|
+
- **`.pr` is an object in the schema and a bare URL string in three state files on disk.** `jq '.pr.url'` on a string errors, and with stderr suppressed that reads as "no PR" - so a run whose PR was already open could be classified as dead. Both readers now read the shape instead of assuming it.
|
|
68
|
+
|
|
69
|
+
### Added
|
|
70
|
+
|
|
71
|
+
- **`/multi-agent:garbage-collect --abandoned`**, backed by `gc-abandoned.sh`. Nothing collected these: `gc-worktrees.sh` skips REGISTERED worktrees on purpose, because a registered worktree belongs to a live run - and a run that stopped never stops being live. 28 worktrees across 7 repos, 23 GB.
|
|
72
|
+
|
|
73
|
+
Two passes, because the state files are not where the disk is: only 5 of those 28 belong to a run that still reads `in_progress`, while 19 have no state file at all and 3 belong to runs that finished and outlived their worktree. A state-driven sweep alone reaches 4% of the problem.
|
|
74
|
+
|
|
75
|
+
Three rules, each written against a shape found on the machine rather than imagined. A run waiting on you is never reaped. A path outside `<repo>/.worktrees/` is never removed - four state files record the REPO ROOT as their `worktreePath`, so a sweep that trusted the field would have deleted a checkout. Uncommitted work is stashed to `autopilot/abandoned/<task-id>` and the worktree is kept, because losing a day of edits is worse than 750 MB.
|
|
76
|
+
|
|
77
|
+
And the rule that decides what the tool is for: **a worktree with no run state is reported, never removed.** `.worktrees/` is not exclusively ours - this machine holds `174`, `pr-4051` and `task-1` there, hand-made, one in active use - and nothing distinguishes those from a pipeline worktree whose log was pruned. They are the 11 GB, and naming them with their sizes is worth more than a rule that guesses.
|
|
78
|
+
|
|
79
|
+
Dry-run by default, `--yes` applies, state salvaged to `artifacts/` before anything goes.
|
|
80
|
+
|
|
81
|
+
### Fixed
|
|
82
|
+
|
|
83
|
+
- **The visual-evidence chain had five readers and no writer, so none of it ever ran.** Phase 0 Step 7.7 is gated on `state.visualEvidence.required`; Phase 3 captures when it is true; Phase 5 records the flow video; Phase 6 blocks on a required artefact that is neither attached nor explained. No phase document ever wrote `required`, `requiredBy` or `platform`, so the verdict was never true, and every consumer below it read a contract that looked satisfied. Step 7.7 now writes the verdict and is named in `features/visual-evidence.md` section 1a as its only writer. Phase 0 has no diff, so the `bugfix` row is provisional and Phase 3 re-decides from the real one.
|
|
84
|
+
|
|
85
|
+
- **The probe was called with `--platform "$PLATFORM"`, a variable no phase document assigns.** The probe refuses an empty value with exit 2, so the call could not have worked on its best day - it went unnoticed only because the step above was unreachable. The platform now comes from the stack (section 1b), which is the right source: it selects which device tooling to probe, and that is a property of the repo rather than of the diff. iOS wins a tie, recorded in `stackWhy`. A repo with no device platform does not get an empty string passed through; the probe does not run and `evidenceCapability.skippedReason` says why.
|
|
86
|
+
|
|
87
|
+
- `smoke-evidence-probe.sh` now fails a `--platform` argument whose variable is not assigned in the same file. Deliberately narrow: more than a hundred variables appear in these documents without an assignment and most are fine, because an agent fills `$WORKTREE` from context. `--platform` is not inferable and its consumer enforces a closed set with a hard exit, which is what makes the rule about this flag rather than about shell hygiene.
|
|
88
|
+
|
|
89
|
+
### Fixed
|
|
90
|
+
|
|
91
|
+
- **Nearly every checkout on the development machine was routed to the iOS toolkit** - 22 of 25, measured. Two causes, measured: `setup` wrote `ai-ios-toolkit` whenever its marker list matched nothing, and nothing else ever revisited the answer, so a Next.js site and several Node CLIs inherited it. The default is gone; an undetected repo now gets the two stack-independent toolkits and a recorded reason, which is an answer rather than a guess.
|
|
92
|
+
|
|
93
|
+
- **Android was never detected.** Phase 1's marker scan ran at `maxdepth 2`, and `AndroidManifest.xml` sits at `<module>/src/main/` - depth 4 in every multi-module app. The scan also read a generated `.next/package.json` as a repo's manifest, and reported a Compose app as iOS because the app vendors a submodule that ships a `Package.swift`.
|
|
94
|
+
|
|
95
|
+
### Added
|
|
96
|
+
|
|
97
|
+
- **`pipeline/lib/stack-detect.sh`** is now the one owner of "what is this repo built with". Deterministic file markers, never a model call: root before anything deeper so `find` traversal order cannot decide the answer, submodule and build-output pruning, depth 5 for the Android manifest, and `MA_STACK_WHY` separating "no marker matched" from "could not be read" - a caller that cannot tell those apart treats an unreadable repo as a language-free one. Verified against every repo on the development machine and seven fixtures.
|
|
98
|
+
|
|
99
|
+
- **`pluginsForStacks()` in `scripts/_stack-routing.mjs`** is the one owner of "which plugins does that want". The mapping was written out by hand in four places, each a chance to get the single asymmetric name wrong - the detector says `web`, the marketplace ships `ai-frontend-toolkit`. Derived names are checked against the locally installed marketplace manifest, so a toolkit that was never shipped is caught at derivation instead of presenting as one silently missing from the session. An unreadable manifest routes by convention and reports `manifestVerified: false`; it never fails closed.
|
|
100
|
+
|
|
101
|
+
- `state.stacks[]` and `state.stackWhy` for the stack axis. `state.detectedStack[]` stays the LANGUAGE axis and is now declared in the schema, which had `additionalProperties: false` and no such property while the phase documents instructed the write. The two axes are deliberately not merged: `graph-build.mjs --stack` accepts `ios|android|node|python|go` and has no notion of `web` or `backend`, and conflating them is what let a Gradle-built JVM service read as an Android app.
|
|
102
|
+
|
|
103
|
+
- **`smoke-stack-detect.sh`**, whose golden-output assertion is what made the consolidation safe: every argument the four hand-written tables accepted must still resolve to exactly the plugins they listed, or the gate fails.
|
|
104
|
+
|
|
105
|
+
### Removed
|
|
106
|
+
|
|
107
|
+
- **Linux and Windows support.** The package advertised three platforms and verified one. Every credential read shells `security`, every iOS build `xcodebuild`, every piece of visual evidence `simctl`, so a run elsewhere did not degrade into a smaller run - it failed partway through with a worktree and a branch already created. ADR-0011 already recorded that CI had not executed since 2026-07-02, and its own status block called the Linux path "smoke-skipped" and the Windows path "unverified".
|
|
108
|
+
|
|
109
|
+
The removal is a declaration, not a deletion. `package.json` declares `os: ["darwin"]`, which npm enforces on the root package: a non-macOS `npm ci` ends in `EBADPLATFORM` before a file is written. `index.js` and `install/index.mjs` refuse a non-darwin host with one line naming the requirement; `uninstall`, `help` and `--version` stay reachable, because somebody who installed before this gate must still be able to remove it. `MULTI_AGENT_ALLOW_NON_DARWIN=1` overrides and says on stderr that nothing on that path is tested.
|
|
110
|
+
|
|
111
|
+
The dead Linux and Windows branches are left in place deliberately. They cost nothing at runtime, and seven constructs in them look like cross-platform scaffolding while being load-bearing on macOS - the `grep -P` ban, the `sha256sum || shasum` ordering, the `stat -c || stat -f` chains among them. ADR-0012 enumerates all seven so the follow-up cleanup does not break the BSD path.
|
|
112
|
+
|
|
113
|
+
- `pipeline/scripts/gate-linux.sh` and the `npm run gate:linux` script. Every workflow now runs on a macOS runner; `test.yml` loses its Ubuntu leg and its `windows-test` job, since with `os: ["darwin"]` enforced those jobs could no longer `npm ci` this tree at all.
|
|
114
|
+
|
|
115
|
+
### Added
|
|
116
|
+
|
|
117
|
+
- **`doctor` check `host-platform`**, which runs first and blocks. First, so an unsupported host reads one honest line instead of a cascade whose common cause is never stated. The blocking set grows from five to six and is documented in `features/doctor.md`.
|
|
118
|
+
|
|
119
|
+
- **`smoke-macos-only.sh`** asserts each declaration separately, because declarations rot silently: the `os` field, both entry gates, the uninstall exemption, the doctor registration and its first position, that no user-facing doc promises another platform, and that no workflow targets a runner that cannot install this package.
|
|
120
|
+
|
|
121
|
+
### Fixed
|
|
122
|
+
|
|
123
|
+
- **`smoke-doctor.sh` computed the documented blocking set and never compared it to anything**, so the ref and the code were free to disagree - and they did, the moment `host-platform` was added. Worse, the computation used `sed -n '/a\|b/p'`, and BSD sed reads `\|` as a literal pipe, so the value had been empty on every macOS run since it was written. Both halves are fixed and the comparison is now an assertion.
|
|
124
|
+
|
|
125
|
+
- **The step-verb mutation test mutated the first matching literal in `doctor.mjs`'s source**, which stopped being a step the probe run reaches as soon as a check that passes on macOS was registered above `install-present`: the engine was never asked, and the gate reported it had gone soft. The step to mutate is now read back from the probe run's own output, so it stays unpinned to any particular check while being guaranteed to execute.
|
|
126
|
+
|
|
17
127
|
## [16.31.1] - 2026-09-13
|
|
18
128
|
|
|
19
129
|
### Fixed
|
package/README.md
CHANGED
|
@@ -10,7 +10,7 @@
|
|
|
10
10
|
|
|
11
11
|
An 8-phase AI development pipeline for **Claude Code**, **Copilot CLI** and **Codex CLI**. Drives a Jira issue or GitHub URL to a merged PR in one command - analysis → plan → TDD → review → test → commit → PR - with multi-repo orchestration, a plan-approval gate, CLI-aware parallel review, and store-compliance checks. Component and Figma-to-code work is dispatched to the per-stack marketplace plugins (iOS/SwiftUI, Android/Compose) rather than bundled, so component skills live in one place.
|
|
12
12
|
|
|
13
|
-
Runs natively on Claude Code, Copilot CLI and Codex CLI. macOS
|
|
13
|
+
Runs natively on Claude Code, Copilot CLI and Codex CLI. macOS only. Zero runtime dependencies.
|
|
14
14
|
|
|
15
15
|
📐 **[Architecture diagrams](./docs/architecture.md)** - the 8-phase flow, operating modes, review/triage, Figma subphases, component layout. **[Ecosystem diagram](./docs/ecosystem.md)** - how this repo, the `multi-agent-plugins` marketplace and `multi-agent-toolkit-mcp` compose.
|
|
16
16
|
|
|
@@ -46,7 +46,7 @@ Run a task - the input type is auto-detected:
|
|
|
46
46
|
/multi-agent:issue # browse unassigned GitHub issues → pick
|
|
47
47
|
```
|
|
48
48
|
|
|
49
|
-
Every input runs the same short intake - **account → (repo) → maturity check → dev-context** - then enters Phase 0. A Jira id or GitHub URL is fetched and maturity-checked
|
|
49
|
+
Every input runs the same short intake - **account → (repo) → maturity check → dev-context** - then enters Phase 0. A Jira id or GitHub URL is fetched and maturity-checked _before_ any code is written; free-text skips the fetch and goes straight to planning. Multi-repo tasks add extra repos at the dev-context step.
|
|
50
50
|
|
|
51
51
|
Add `autopilot` to skip confirmations, or `--local` to work on the current branch without a worktree (e.g. `/multi-agent:autopilot "PROJ-1234"`). Pipeline depth is a question the run asks, not a flag: `/multi-agent` and `/multi-agent:local` offer Full or Short at Phase 0.
|
|
52
52
|
|
|
@@ -75,117 +75,117 @@ The discipline behind all of this - bounded loops, evidence gates, token-budgete
|
|
|
75
75
|
|
|
76
76
|
## Modes
|
|
77
77
|
|
|
78
|
-
| Mode
|
|
79
|
-
|
|
80
|
-
| Full
|
|
81
|
-
| Autopilot | `/multi-agent:autopilot "task"`
|
|
82
|
-
| Local
|
|
83
|
-
| Depth
|
|
84
|
-
| Ship
|
|
85
|
-
| Audit
|
|
86
|
-
| Audit
|
|
78
|
+
| Mode | Command | Flow |
|
|
79
|
+
| --------- | ------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
80
|
+
| Full | `/multi-agent "task"` | All 8 phases, interactive |
|
|
81
|
+
| Autopilot | `/multi-agent:autopilot "task"` | 7 phases (interactive Test gate dropped), no confirmations |
|
|
82
|
+
| Local | `/multi-agent:local "task"` | Full pipeline minus the interactive Test gate, current branch (no worktree) |
|
|
83
|
+
| Depth | asked at Phase 0 Step 7.5 | Full (all phases) or Short (Dev → Review → Test → Commit → Report). Not a command name - `/multi-agent` and `:local` ask, both autopilot entries always run Full |
|
|
84
|
+
| Ship | `/multi-agent:resume-local` | Run the review→test→commit→report tail over local work |
|
|
85
|
+
| Audit | `/multi-agent:design-check` | Mock-mode vs Figma conformance, local-only |
|
|
86
|
+
| Audit | `/multi-agent:testflight-validation` | Pre-submission gates for a TestFlight build: static archive audit → Apple's `altool --validate-app` → Review-Guidelines check. Validates only, never uploads |
|
|
87
87
|
|
|
88
88
|
Depth, autopilot and `--local` are the only knobs on the run itself; everything else is its own command. The full catalog is below.
|
|
89
89
|
|
|
90
90
|
## Commands
|
|
91
91
|
|
|
92
|
-
`/multi-agent` plus
|
|
92
|
+
`/multi-agent` plus 56 sub-commands. `/multi-agent:help` renders the same catalog in your terminal, in your `outputLanguage`.
|
|
93
93
|
|
|
94
94
|
### Pipeline entries
|
|
95
95
|
|
|
96
|
-
| Command
|
|
97
|
-
|
|
98
|
-
| `/multi-agent "task"`
|
|
99
|
-
| `/multi-agent:local "task"`
|
|
100
|
-
| `/multi-agent:autopilot "task"`
|
|
101
|
-
| `/multi-agent:local-autopilot "task"` | Current branch, no confirmations, always Full
|
|
102
|
-
| `/multi-agent:resume-local`
|
|
96
|
+
| Command | What it does |
|
|
97
|
+
| ------------------------------------- | ---------------------------------------------------------------------------------------------------- |
|
|
98
|
+
| `/multi-agent "task"` | Full pipeline in a worktree. Asks Full or Short depth at Phase 0 |
|
|
99
|
+
| `/multi-agent:local "task"` | Same pipeline on the current branch, no worktree |
|
|
100
|
+
| `/multi-agent:autopilot "task"` | Worktree, no confirmations, always Full |
|
|
101
|
+
| `/multi-agent:local-autopilot "task"` | Current branch, no confirmations, always Full |
|
|
102
|
+
| `/multi-agent:resume-local` | Pipeline tail over work already done locally: Review → Build+Test → Commit/PR → Report. No dev phase |
|
|
103
103
|
|
|
104
104
|
### Task control
|
|
105
105
|
|
|
106
|
-
| Command
|
|
107
|
-
|
|
108
|
-
| `/multi-agent:status`
|
|
109
|
-
| `/multi-agent:log [#N]`
|
|
110
|
-
| `/multi-agent:resume [#N]`
|
|
111
|
-
| `/multi-agent:kill [#N]`
|
|
112
|
-
| `/multi-agent:steer #N "<instruction>"` | Correct a running task without stopping it; applied at the next phase boundary
|
|
113
|
-
| `/multi-agent:search`
|
|
114
|
-
| `/multi-agent:garbage-collect`
|
|
115
|
-
| `/multi-agent:prune-logs`
|
|
116
|
-
| `/multi-agent:purge`
|
|
106
|
+
| Command | What it does |
|
|
107
|
+
| --------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
108
|
+
| `/multi-agent:status` | Every task's ID, phase, branch and state |
|
|
109
|
+
| `/multi-agent:log [#N]` | Show a task's `agent-log.md` (most recent by default) |
|
|
110
|
+
| `/multi-agent:resume [#N]` | Carry a stopped or failed task on from its last phase |
|
|
111
|
+
| `/multi-agent:kill [#N]` | Stop a task, remove its worktree and branch |
|
|
112
|
+
| `/multi-agent:steer #N "<instruction>"` | Correct a running task without stopping it; applied at the next phase boundary |
|
|
113
|
+
| `/multi-agent:search` | Ranked search across every task log; `--semantic` queries the triage corpus |
|
|
114
|
+
| `/multi-agent:garbage-collect` | Sweep leftover scratch, orphan worktrees and offloaded payloads. `--abandoned` also reaps runs that stopped and were never picked back up. Dry-run first |
|
|
115
|
+
| `/multi-agent:prune-logs` | Delete per-task logs by age / project / task. Audit trail and metrics kept |
|
|
116
|
+
| `/multi-agent:purge` | Wipe every worktree, branch, log and state file. Double confirmation |
|
|
117
117
|
|
|
118
118
|
### Review
|
|
119
119
|
|
|
120
|
-
| Command
|
|
121
|
-
|
|
122
|
-
| `/multi-agent:review`
|
|
123
|
-
| `/multi-agent:review-jira`
|
|
124
|
-
| `/multi-agent:review-issue`
|
|
125
|
-
| `/multi-agent:review-analysis`
|
|
126
|
-
| `/multi-agent:diff-explain`
|
|
127
|
-
| `/multi-agent:refactor`
|
|
128
|
-
| `/multi-agent:scan`
|
|
129
|
-
| `/multi-agent:prune-prompts`
|
|
130
|
-
| `/multi-agent:ios-coding-standard` | Audit an iOS module against the 99-rule registry, produce a remediation plan
|
|
120
|
+
| Command | What it does |
|
|
121
|
+
| ---------------------------------- | ------------------------------------------------------------------------------------------- |
|
|
122
|
+
| `/multi-agent:review` | Parallel review of a branch diff or a PR; inline comments + approve/needs-work on PR input |
|
|
123
|
+
| `/multi-agent:review-jira` | Grade a Jira issue's readiness for development, comment the gaps |
|
|
124
|
+
| `/multi-agent:review-issue` | Same grading for a GitHub issue |
|
|
125
|
+
| `/multi-agent:review-analysis` | Review a written analysis document; findings cite the Locked rule they break |
|
|
126
|
+
| `/multi-agent:diff-explain` | Map a Phase 4 triage finding back to the diff lines that caused it |
|
|
127
|
+
| `/multi-agent:refactor` | Best-practice extraction + bug hunt + derived-skill drift + toolkit MCP research → one plan |
|
|
128
|
+
| `/multi-agent:scan` | Skill security scan of local skill directories against a tiered pattern catalog |
|
|
129
|
+
| `/multi-agent:prune-prompts` | Zero-base review of the always-on instruction footprint; keep / trial / delete per rule |
|
|
130
|
+
| `/multi-agent:ios-coding-standard` | Audit an iOS module against the 99-rule registry, produce a remediation plan |
|
|
131
131
|
|
|
132
132
|
### Analysis
|
|
133
133
|
|
|
134
|
-
| Command
|
|
135
|
-
|
|
136
|
-
| `/multi-agent:analysis`
|
|
137
|
-
| `/multi-agent:analysis-resolve`
|
|
134
|
+
| Command | What it does |
|
|
135
|
+
| --------------------------------- | ----------------------------------------------------------------------------------------- |
|
|
136
|
+
| `/multi-agent:analysis` | Standalone feature spec: global (23-section handoff) or corporate (IG/UC/FG) profile |
|
|
137
|
+
| `/multi-agent:analysis-resolve` | Answer an analysis doc's Section 20 open questions one row at a time |
|
|
138
138
|
| `/multi-agent:complaint-analysis` | Customer-complaint triage with Graylog evidence: client / bff root cause, or core routing |
|
|
139
139
|
|
|
140
140
|
### Testing on a device
|
|
141
141
|
|
|
142
|
-
| Command
|
|
143
|
-
|
|
144
|
-
| `/multi-agent:test`
|
|
145
|
-
| `/multi-agent:test-dark-mode`
|
|
146
|
-
| `/multi-agent:test-accessibility`
|
|
147
|
-
| `/multi-agent:test-dynamic-type`
|
|
148
|
-
| `/multi-agent:test-screenshots [locale]` | App Store screenshot set in a locale (defaults to `tr`)
|
|
149
|
-
| `/multi-agent:manual-test`
|
|
142
|
+
| Command | What it does |
|
|
143
|
+
| ---------------------------------------- | ------------------------------------------------------------------------- |
|
|
144
|
+
| `/multi-agent:test` | UI Bug Hunter on a booted simulator or emulator: screenshot, tap, analyze |
|
|
145
|
+
| `/multi-agent:test-dark-mode` | Walk every screen light then dark, report contrast and colour bugs |
|
|
146
|
+
| `/multi-agent:test-accessibility` | VoiceOver labels, sub-44pt tap targets, contrast, traits |
|
|
147
|
+
| `/multi-agent:test-dynamic-type` | Re-walk every screen at XL through accessibility-XL, report truncation |
|
|
148
|
+
| `/multi-agent:test-screenshots [locale]` | App Store screenshot set in a locale (defaults to `tr`) |
|
|
149
|
+
| `/multi-agent:manual-test` | Phase 5 standalone: check out the task branch and prepare it for Xcode |
|
|
150
150
|
|
|
151
151
|
### Design, build and store
|
|
152
152
|
|
|
153
|
-
| Command
|
|
154
|
-
|
|
155
|
-
| `/multi-agent:design-check`
|
|
156
|
-
| `/multi-agent:store-ready`
|
|
157
|
-
| `/multi-agent:testflight-validation` | iOS-pinned alias of `store-ready`. Validates only, never uploads
|
|
158
|
-
| `/multi-agent:build-optimize`
|
|
153
|
+
| Command | What it does |
|
|
154
|
+
| ------------------------------------ | ---------------------------------------------------------------------------------------- |
|
|
155
|
+
| `/multi-agent:design-check` | Mock-mode vs Figma audit with a coverage gate; annotated HTML + PDF report |
|
|
156
|
+
| `/multi-agent:store-ready` | Pre-submission gates for iOS and Android: package audit, store validation, policy review |
|
|
157
|
+
| `/multi-agent:testflight-validation` | iOS-pinned alias of `store-ready`. Validates only, never uploads |
|
|
158
|
+
| `/multi-agent:build-optimize` | Benchmark an Xcode build, run the analyzers, produce a recommend-first plan |
|
|
159
159
|
|
|
160
160
|
### Tickets and reporting
|
|
161
161
|
|
|
162
|
-
| Command
|
|
163
|
-
|
|
164
|
-
| `/multi-agent:jira`
|
|
165
|
-
| `/multi-agent:issue`
|
|
166
|
-
| `/multi-agent:create-jira` | Draft a Task / Bug / Story to the project's own conventions, preview before create
|
|
167
|
-
| `/multi-agent:channels`
|
|
168
|
-
| `/multi-agent:feedback`
|
|
162
|
+
| Command | What it does |
|
|
163
|
+
| -------------------------- | ----------------------------------------------------------------------------------- |
|
|
164
|
+
| `/multi-agent:jira` | Browse your open Jira issues → pick → branch → mode → launch |
|
|
165
|
+
| `/multi-agent:issue` | Browse unassigned GitHub issues → pick → auto-assign → launch |
|
|
166
|
+
| `/multi-agent:create-jira` | Draft a Task / Bug / Story to the project's own conventions, preview before create |
|
|
167
|
+
| `/multi-agent:channels` | Post the multi-channel report: Jira, Confluence, Wiki, PR description, board status |
|
|
168
|
+
| `/multi-agent:feedback` | Send one message to the maintainer. Only your text is sent, no logs or paths |
|
|
169
169
|
|
|
170
170
|
### Your own routines
|
|
171
171
|
|
|
172
|
-
| Command
|
|
173
|
-
|
|
174
|
-
| `/multi-agent:save [name]`
|
|
175
|
-
| `/multi-agent:routines`
|
|
176
|
-
| `/multi-agent:forget [name]` | Remove a saved routine and its registry entry
|
|
172
|
+
| Command | What it does |
|
|
173
|
+
| ---------------------------- | ---------------------------------------------------------------------------------- |
|
|
174
|
+
| `/multi-agent:save [name]` | Save a recurring job as a reusable `/multi-agent:<name>`. Local-only, never synced |
|
|
175
|
+
| `/multi-agent:routines` | List your saved routines and what each does |
|
|
176
|
+
| `/multi-agent:forget [name]` | Remove a saved routine and its registry entry |
|
|
177
177
|
|
|
178
178
|
### Setup and maintenance
|
|
179
179
|
|
|
180
|
-
| Command
|
|
181
|
-
|
|
182
|
-
| `/multi-agent:setup`
|
|
183
|
-
| `/multi-agent:stack [ids]`
|
|
184
|
-
| `/multi-agent:language [en\|tr]` | Show or set `outputLanguage`; `promptLanguage` stays English
|
|
185
|
-
| `/multi-agent:sync`
|
|
186
|
-
| `/multi-agent:update`
|
|
187
|
-
| `/multi-agent:uninstall`
|
|
188
|
-
| `/multi-agent:help`
|
|
180
|
+
| Command | What it does |
|
|
181
|
+
| -------------------------------- | ------------------------------------------------------------------------------ |
|
|
182
|
+
| `/multi-agent:setup` | First-run wizard: keychain token discovery, git identity, pipeline preparation |
|
|
183
|
+
| `/multi-agent:stack [ids]` | Enable the marketplace plugin(s) for this repo. Multi-select |
|
|
184
|
+
| `/multi-agent:language [en\|tr]` | Show or set `outputLanguage`; `promptLanguage` stays English |
|
|
185
|
+
| `/multi-agent:sync` | One-shot sync: Claude Code, Copilot CLI, pipeline repo, website, toolkit MCP |
|
|
186
|
+
| `/multi-agent:update` | Update to the latest published npm release and run migrations |
|
|
187
|
+
| `/multi-agent:uninstall` | Remove the pipeline from every CLI. Keychain tokens always left intact |
|
|
188
|
+
| `/multi-agent:help` | This catalog, in the terminal, in your `outputLanguage` |
|
|
189
189
|
|
|
190
190
|
### Subagents
|
|
191
191
|
|
|
@@ -195,6 +195,51 @@ Eight are installed alongside the commands and dispatched by the phases: `explor
|
|
|
195
195
|
|
|
196
196
|
Two compliance skills install on every host and back the store gates: `apple-archive-compliance` (18-rule Apple review scan with ITMS code mapping) and `google-play-compliance` (21-rule Play policy catalog with Console error codes). Everything else stack-shaped - SwiftUI, Compose, backend, frontend - comes from the marketplace plugins described below.
|
|
197
197
|
|
|
198
|
+
## Continuous mode
|
|
199
|
+
|
|
200
|
+
Everything above starts when you start it. Continuous mode is the same pipeline
|
|
201
|
+
picking work up on its own, on ONE machine you choose, from repos you choose.
|
|
202
|
+
|
|
203
|
+
| Command | What it does |
|
|
204
|
+
| --- | --- |
|
|
205
|
+
| `/multi-agent:autopilot-on` | Pick the repos this machine watches. Labelled GitHub issues and assigned + labelled Jira items then run in a worktree and stop at an open PR. Re-run to change the list |
|
|
206
|
+
| `/multi-agent:autopilot-status` | What is running and at which phase, what is queued, what is waiting for an answer, the PRs of the last day, and the rolling spend |
|
|
207
|
+
| `/multi-agent:autopilot-off` | Remove the schedule. Work already running finishes; the repo selection is kept |
|
|
208
|
+
|
|
209
|
+
**Nothing is on by default and nothing is added implicitly.** Installing the
|
|
210
|
+
package writes no state and schedules nothing; `smoke-autopilot-default-off.sh`
|
|
211
|
+
fails the build if that ever changes. A label is a filter, not a gate - anyone
|
|
212
|
+
who can open an issue in a repo you have push on could add one - so the gate is
|
|
213
|
+
the picker, and it is per machine.
|
|
214
|
+
|
|
215
|
+
The biggest win is not parallelism. Measured here, the median run is 44 minutes,
|
|
216
|
+
but an item that finishes at 14:00 waits until you sit down again: overnight that
|
|
217
|
+
is 16 hours against 40 minutes. Continuous mode removes the waiting, not the work.
|
|
218
|
+
|
|
219
|
+
**What it will not do.** It does not merge - the runner stops at an open PR and
|
|
220
|
+
the decision stays yours. It does not touch an attended run: per-repo concurrency
|
|
221
|
+
is always 1, so the queue steps around a repo you are working in rather than
|
|
222
|
+
competing for `.git/index.lock`. There is no cap on PRs; the bounds are
|
|
223
|
+
`costCeilingUsd` over a rolling 24 hours and what the machine can hold.
|
|
224
|
+
|
|
225
|
+
**A menu bar indicator**, when `swiftc` is present, draws the same `status.json`
|
|
226
|
+
in the top right and refreshes on its own: one row per item with its id, Full or
|
|
227
|
+
Short, the phase as a fraction, the elapsed time and the stack. A row disappears
|
|
228
|
+
the moment the item finishes and reappears under Reports with its PR. It only
|
|
229
|
+
draws - it cannot start, stop or change a run. ActivityKit is unavailable on
|
|
230
|
+
macOS, so this is an `NSStatusItem`, built from source on demand rather than
|
|
231
|
+
shipped as a binary that would need signing.
|
|
232
|
+
|
|
233
|
+
**It survives a restart with no command to run.** launchd loads the job at
|
|
234
|
+
**login**, not at boot, and that is correct rather than a limitation: the login
|
|
235
|
+
keychain is what unlocks the tokens, so a tick that fired before login could not
|
|
236
|
+
reach Jira or GitHub anyway. There is no `autopilot-resume` - a command you have
|
|
237
|
+
to remember is a queue that silently stops when you forget it. `doctor` reports
|
|
238
|
+
the real failure instead: configured, but launchd holds no job.
|
|
239
|
+
|
|
240
|
+
Sleep is held **only on AC**. A queue that flattens a laptop off the charger is a
|
|
241
|
+
bug; on battery the assertion is released and work resumes when you plug in.
|
|
242
|
+
|
|
198
243
|
## Stacks
|
|
199
244
|
|
|
200
245
|
Stack skills ship as versioned plugins in the [`mmerterden/multi-agent-plugins`](https://github.com/mmerterden/multi-agent-plugins) marketplace. Select a stack per-repo:
|
|
@@ -207,13 +252,13 @@ This enables the matching plugin (+ the shared `ai-common` plugin) in the repo's
|
|
|
207
252
|
|
|
208
253
|
## Tool support
|
|
209
254
|
|
|
210
|
-
The pipeline runs natively on **Claude Code**, **Copilot CLI** and **Codex CLI** - all three install from the same `pipeline/` source and get the same
|
|
255
|
+
The pipeline runs natively on **Claude Code**, **Copilot CLI** and **Codex CLI** - all three install from the same `pipeline/` source and get the same 56 commands.
|
|
211
256
|
|
|
212
|
-
| Tool
|
|
213
|
-
|
|
257
|
+
| Tool | Flag | What it installs |
|
|
258
|
+
| ----------- | -------------------- | ------------------------------------------------------------------------------------------------------ |
|
|
214
259
|
| Claude Code | `--claude` (default) | slash commands + skills + agents + three `PreToolUse` hooks (secret scan, agent-guard, read-size gate) |
|
|
215
|
-
| Copilot CLI | `--copilot`
|
|
216
|
-
| Codex CLI
|
|
260
|
+
| Copilot CLI | `--copilot` | instructions + 56 sub-command skills + scripts |
|
|
261
|
+
| Codex CLI | `--codex` | one router skill + 56 specs as refs + 8 agent TOML + `AGENTS.md` block + `codex mcp add` |
|
|
217
262
|
|
|
218
263
|
Filter skills by stack with `--platform=ios\|android\|all`.
|
|
219
264
|
|
|
@@ -232,20 +277,20 @@ the one host whose panel spans two vendors, and the triage note says so.
|
|
|
232
277
|
|
|
233
278
|
## Tokens & integrations
|
|
234
279
|
|
|
235
|
-
`setup` scans your OS keychain and maps each token by a **logical name** (e.g. `jira`) to its real keychain entry - the pipeline resolves tokens through that mapping (`credential-store.sh`), so literal keychain names never appear in synced files. Tokens stay in the
|
|
236
|
-
|
|
237
|
-
| Token
|
|
238
|
-
|
|
239
|
-
| `jira`
|
|
240
|
-
| `github`
|
|
241
|
-
| `bitbucket`
|
|
242
|
-
| `confluence`
|
|
243
|
-
| `figma` + `figma_mcp` | fetch design context
|
|
244
|
-
| `fortify`
|
|
245
|
-
| `firebase`
|
|
246
|
-
| `jenkins`
|
|
247
|
-
| `npm`
|
|
248
|
-
| `appstore_connect_*`
|
|
280
|
+
`setup` scans your OS keychain and maps each token by a **logical name** (e.g. `jira`) to its real keychain entry - the pipeline resolves tokens through that mapping (`credential-store.sh`), so literal keychain names never appear in synced files. Tokens stay in the macOS Keychain, are **never committed or logged**, and are all **optional** - the pipeline asks for any it needs at Phase 0.
|
|
281
|
+
|
|
282
|
+
| Token | Used for | Phase |
|
|
283
|
+
| --------------------- | ---------------------------------------------------------------- | ----------------------- |
|
|
284
|
+
| `jira` | fetch the issue · post the report comment | 0, 7 |
|
|
285
|
+
| `github` | issues · PRs · `gh` auth | 0, 6 |
|
|
286
|
+
| `bitbucket` | PR create/update (reviewer-preserving) · diff | 6 |
|
|
287
|
+
| `confluence` | publish analysis / wiki pages | 7 |
|
|
288
|
+
| `figma` + `figma_mcp` | fetch design context | analysis only |
|
|
289
|
+
| `fortify` | security-scan findings gate | 4 |
|
|
290
|
+
| `firebase` | Firebase service-account JSON for Firebase projects | as needed |
|
|
291
|
+
| `jenkins` | CI trigger / status | build / deploy |
|
|
292
|
+
| `npm` | package publish (mostly CI) | release |
|
|
293
|
+
| `appstore_connect_*` | TestFlight / App Store pre-submission validation (optional, iOS) | `testflight-validation` |
|
|
249
294
|
|
|
250
295
|
The **secret scan** runs as a `PreToolUse` hook on Claude Code (hard-blocks a commit on a hit) and as a pre-push check elsewhere.
|
|
251
296
|
|
|
@@ -261,13 +306,13 @@ Uninstall preserves this layer: the tokens, the reader that opens them, the mapp
|
|
|
261
306
|
|
|
262
307
|
## Platform support
|
|
263
308
|
|
|
264
|
-
Runs on **macOS
|
|
309
|
+
Runs on **macOS** only. The package declares `os: ["darwin"]`, so `npm` refuses to install it elsewhere rather than letting a run fail halfway through: every credential read shells `security`, every iOS build `xcodebuild`, every piece of visual evidence `simctl`. Node.js 20.11+ (tested on 20 and 22). Reasoning: [ADR-0012](docs/adr/0012-macos-only.md).
|
|
265
310
|
|
|
266
311
|
## Companion repos
|
|
267
312
|
|
|
268
|
-
| Repo
|
|
269
|
-
|
|
270
|
-
| [`mmerterden/multi-agent-plugins`](https://github.com/mmerterden/multi-agent-plugins)
|
|
313
|
+
| Repo | What it is |
|
|
314
|
+
| --------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
|
315
|
+
| [`mmerterden/multi-agent-plugins`](https://github.com/mmerterden/multi-agent-plugins) | Marketplace of per-stack skill toolkits (iOS / Android / Frontend / Backend + common). `/multi-agent:stack` enables the matching plugin. |
|
|
271
316
|
| [`mmerterden/multi-agent-toolkit-mcp`](https://github.com/mmerterden/multi-agent-toolkit-mcp) | MCP server for UI testing / simulator capture / xcodebuild - powers the Phase 5 UI Bug Hunter. Published on the public npm registry as [`@mmerterden/multi-agent-toolkit-mcp`](https://www.npmjs.com/package/@mmerterden/multi-agent-toolkit-mcp); the installer registers it with each CLI for you, so `npx` resolves it with no extra configuration. |
|
|
272
317
|
|
|
273
318
|
## License
|