@mmerterden/multi-agent-pipeline 17.0.0 → 17.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +32 -0
- package/README.md +49 -4
- package/README.tr.md +50 -4
- package/docs/architecture.md +3 -3
- package/docs/ecosystem.md +5 -5
- package/install/templates/multi-agent-autopilot.plist.template +79 -0
- package/package.json +1 -1
- package/pipeline/commands/multi-agent/autopilot-off/SKILL.md +64 -0
- package/pipeline/commands/multi-agent/autopilot-on/SKILL.md +173 -0
- package/pipeline/commands/multi-agent/autopilot-status/SKILL.md +74 -0
- package/pipeline/commands/multi-agent/channels/SKILL.md +41 -12
- package/pipeline/commands/multi-agent/help/SKILL.md +41 -35
- package/pipeline/commands/multi-agent/manual-test/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/sync/SKILL.md +10 -9
- package/pipeline/commands/multi-agent/update/SKILL.md +1 -1
- package/pipeline/lib/autopilot-activation.sh +117 -0
- package/pipeline/lib/autopilot-state.sh +150 -0
- package/pipeline/lib/issue-fetcher.sh +18 -1
- package/pipeline/lib/plan-todos.sh +18 -0
- package/pipeline/multi-agent-refs/channels/jira.md +80 -20
- package/pipeline/multi-agent-refs/channels/pr.md +65 -19
- package/pipeline/multi-agent-refs/cross-cli-contract.md +6 -5
- package/pipeline/multi-agent-refs/features/visual-evidence.md +19 -7
- package/pipeline/multi-agent-refs/phases/phase-0-init.md +1 -1
- package/pipeline/multi-agent-refs/phases/phase-2-planning.md +17 -15
- package/pipeline/multi-agent-refs/phases/phase-3-dev.md +1 -1
- package/pipeline/multi-agent-refs/phases/phase-6-commit.md +1 -1
- package/pipeline/multi-agent-refs/readiness-review.md +7 -1
- package/pipeline/multi-agent-refs/rules.md +3 -11
- package/pipeline/multi-agent-refs/tracker-contract.md +32 -0
- package/pipeline/schemas/autopilot-config.schema.json +149 -0
- package/pipeline/schemas/token-budget.json +2 -2
- package/pipeline/scripts/autopilot-arming.mjs +147 -0
- package/pipeline/scripts/autopilot-intake.mjs +383 -0
- package/pipeline/scripts/autopilot-menubar.swift +361 -0
- package/pipeline/scripts/autopilot-runner.mjs +349 -0
- package/pipeline/scripts/autopilot-status.sh +212 -0
- package/pipeline/scripts/jira-search.sh +70 -0
- package/pipeline/scripts/phase-tracker.sh +134 -12
- package/pipeline/scripts/probe-evidence-capability.sh +27 -3
- package/pipeline/scripts/run-ui-tests.sh +113 -4
- package/pipeline/skills/.skill-manifest.json +16 -4
- package/pipeline/skills/shared/core/multi-agent-autopilot-off/SKILL.md +67 -0
- package/pipeline/skills/shared/core/multi-agent-autopilot-on/SKILL.md +146 -0
- package/pipeline/skills/shared/core/multi-agent-autopilot-status/SKILL.md +64 -0
- package/pipeline/skills/shared/core/multi-agent-channels/SKILL.md +62 -11
- package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +9 -8
package/CHANGELOG.md
CHANGED
|
@@ -14,6 +14,38 @@ Internal file-layout changes that don't affect the slash-command surface are sti
|
|
|
14
14
|
|
|
15
15
|
---
|
|
16
16
|
|
|
17
|
+
## [17.1.0] - 2026-09-14
|
|
18
|
+
|
|
19
|
+
Continuous mode, and three things that were computed correctly and shown to
|
|
20
|
+
nobody: the plan's own steps, a web repo's UI tests, and the queue's status when
|
|
21
|
+
one field arrived as a number.
|
|
22
|
+
|
|
23
|
+
### Added
|
|
24
|
+
|
|
25
|
+
- **Continuous mode.** `/multi-agent:autopilot-on` picks the repos ONE machine watches; labelled GitHub issues and assigned + labelled Jira items then run in a worktree and stop at an open PR. `:autopilot-status` is the single producer the terminal, the menu bar indicator and the session hook all render, `:autopilot-off` removes the schedule and keeps the selection. Nothing is on by default: installing writes no state and schedules nothing, and `smoke-autopilot-default-off.sh` fails the build if that changes.
|
|
26
|
+
|
|
27
|
+
The ordering is deterministic with no model call - explicit rank, repo grouping, priority, then OLDEST first, which is the opposite of the `jira` picker on purpose: a human wants what just landed, an unattended queue must not starve what has been waiting. Two preconditions are checked before any item is taken: a rolling 24-hour spend total, and a per-source credential gate, so a dead Jira token stops Jira items without stopping GitHub ones.
|
|
28
|
+
|
|
29
|
+
A menu bar indicator draws the same `status.json` in the top right when `swiftc` is present. ActivityKit is `@available(macOS, unavailable)` - there is no Live Activity on a Mac - so this is an `NSStatusItem`, built from source on demand rather than shipped as a binary that would need signing.
|
|
30
|
+
|
|
31
|
+
- **UI tests on web.** `run-ui-tests.sh` rejected every platform but `ios|android`, so a repo with a full Playwright suite reported "no UI test target" and the PR body said UI tests had not run - a statement true of the runner and false of the repo. Detection keys on the browser-driving import (`@playwright/test`, `cy.visit(`), not on a directory called `e2e`: the same distinction that keeps 475 iOS snapshot tests from being counted as UI tests.
|
|
32
|
+
|
|
33
|
+
### Changed
|
|
34
|
+
|
|
35
|
+
- **The Jira comment and the PR body stopped being the same document.** Jira now carries Geliştirme Özeti, Test Senaryoları, Etki Analizi and Bağlantılar with the PR link on line 1 and no identifiers, file paths or diff hunks anywhere in it. The PR keeps all of that and gains Teknik Açıklama, Etki Analizi and Build - the build command, its result and the base sha it ran against. Given/When/Then is gone from the scenarios: it reads as translated English to the person running them, who wants a titled list they can follow with the app open.
|
|
36
|
+
|
|
37
|
+
### Fixed
|
|
38
|
+
|
|
39
|
+
- **The menu bar indicator read "autopilot kapalı" with items in flight.** `phase` was declared `String?` while the producer passes the queue's own value through `jq` untouched, so a numeric phase made Swift's `Decodable` throw - and the throw did not lose one field, it failed the whole document. A silent blank is the worst failure a status indicator has, because it is indistinguishable from good news.
|
|
40
|
+
|
|
41
|
+
- **`ma_ap_boottime` returned the microseconds.** `.*sec = ` is greedy and walks past `sec` into `usec`, so the helper produced a six-digit number that looked plausible and never equalled the same fact read anywhere else. It surfaced as a live runner being declared dead.
|
|
42
|
+
|
|
43
|
+
- **Three phase documents said "MCP forbidden" without qualifying it**, while the gate that enforces it has always matched `figma` and nothing else. The prose therefore banned the screenshot, xcodebuild and UI-test tools that Phase 3 itself calls, which is one way a run reaches Phase 7 with no evidence.
|
|
44
|
+
|
|
45
|
+
- **`channels/jira.md` told Phase 7 to upload evidence Phase 6 had already uploaded**, so re-rendering a comment attached every file a second time.
|
|
46
|
+
|
|
47
|
+
---
|
|
48
|
+
|
|
17
49
|
## [17.0.0] - 2026-09-14
|
|
18
50
|
|
|
19
51
|
Five things in this release were not working, and four of them looked like they
|
package/README.md
CHANGED
|
@@ -89,7 +89,7 @@ Depth, autopilot and `--local` are the only knobs on the run itself; everything
|
|
|
89
89
|
|
|
90
90
|
## Commands
|
|
91
91
|
|
|
92
|
-
`/multi-agent` plus
|
|
92
|
+
`/multi-agent` plus 56 sub-commands. `/multi-agent:help` renders the same catalog in your terminal, in your `outputLanguage`.
|
|
93
93
|
|
|
94
94
|
### Pipeline entries
|
|
95
95
|
|
|
@@ -195,6 +195,51 @@ Eight are installed alongside the commands and dispatched by the phases: `explor
|
|
|
195
195
|
|
|
196
196
|
Two compliance skills install on every host and back the store gates: `apple-archive-compliance` (18-rule Apple review scan with ITMS code mapping) and `google-play-compliance` (21-rule Play policy catalog with Console error codes). Everything else stack-shaped - SwiftUI, Compose, backend, frontend - comes from the marketplace plugins described below.
|
|
197
197
|
|
|
198
|
+
## Continuous mode
|
|
199
|
+
|
|
200
|
+
Everything above starts when you start it. Continuous mode is the same pipeline
|
|
201
|
+
picking work up on its own, on ONE machine you choose, from repos you choose.
|
|
202
|
+
|
|
203
|
+
| Command | What it does |
|
|
204
|
+
| --- | --- |
|
|
205
|
+
| `/multi-agent:autopilot-on` | Pick the repos this machine watches. Labelled GitHub issues and assigned + labelled Jira items then run in a worktree and stop at an open PR. Re-run to change the list |
|
|
206
|
+
| `/multi-agent:autopilot-status` | What is running and at which phase, what is queued, what is waiting for an answer, the PRs of the last day, and the rolling spend |
|
|
207
|
+
| `/multi-agent:autopilot-off` | Remove the schedule. Work already running finishes; the repo selection is kept |
|
|
208
|
+
|
|
209
|
+
**Nothing is on by default and nothing is added implicitly.** Installing the
|
|
210
|
+
package writes no state and schedules nothing; `smoke-autopilot-default-off.sh`
|
|
211
|
+
fails the build if that ever changes. A label is a filter, not a gate - anyone
|
|
212
|
+
who can open an issue in a repo you have push on could add one - so the gate is
|
|
213
|
+
the picker, and it is per machine.
|
|
214
|
+
|
|
215
|
+
The biggest win is not parallelism. Measured here, the median run is 44 minutes,
|
|
216
|
+
but an item that finishes at 14:00 waits until you sit down again: overnight that
|
|
217
|
+
is 16 hours against 40 minutes. Continuous mode removes the waiting, not the work.
|
|
218
|
+
|
|
219
|
+
**What it will not do.** It does not merge - the runner stops at an open PR and
|
|
220
|
+
the decision stays yours. It does not touch an attended run: per-repo concurrency
|
|
221
|
+
is always 1, so the queue steps around a repo you are working in rather than
|
|
222
|
+
competing for `.git/index.lock`. There is no cap on PRs; the bounds are
|
|
223
|
+
`costCeilingUsd` over a rolling 24 hours and what the machine can hold.
|
|
224
|
+
|
|
225
|
+
**A menu bar indicator**, when `swiftc` is present, draws the same `status.json`
|
|
226
|
+
in the top right and refreshes on its own: one row per item with its id, Full or
|
|
227
|
+
Short, the phase as a fraction, the elapsed time and the stack. A row disappears
|
|
228
|
+
the moment the item finishes and reappears under Reports with its PR. It only
|
|
229
|
+
draws - it cannot start, stop or change a run. ActivityKit is unavailable on
|
|
230
|
+
macOS, so this is an `NSStatusItem`, built from source on demand rather than
|
|
231
|
+
shipped as a binary that would need signing.
|
|
232
|
+
|
|
233
|
+
**It survives a restart with no command to run.** launchd loads the job at
|
|
234
|
+
**login**, not at boot, and that is correct rather than a limitation: the login
|
|
235
|
+
keychain is what unlocks the tokens, so a tick that fired before login could not
|
|
236
|
+
reach Jira or GitHub anyway. There is no `autopilot-resume` - a command you have
|
|
237
|
+
to remember is a queue that silently stops when you forget it. `doctor` reports
|
|
238
|
+
the real failure instead: configured, but launchd holds no job.
|
|
239
|
+
|
|
240
|
+
Sleep is held **only on AC**. A queue that flattens a laptop off the charger is a
|
|
241
|
+
bug; on battery the assertion is released and work resumes when you plug in.
|
|
242
|
+
|
|
198
243
|
## Stacks
|
|
199
244
|
|
|
200
245
|
Stack skills ship as versioned plugins in the [`mmerterden/multi-agent-plugins`](https://github.com/mmerterden/multi-agent-plugins) marketplace. Select a stack per-repo:
|
|
@@ -207,13 +252,13 @@ This enables the matching plugin (+ the shared `ai-common` plugin) in the repo's
|
|
|
207
252
|
|
|
208
253
|
## Tool support
|
|
209
254
|
|
|
210
|
-
The pipeline runs natively on **Claude Code**, **Copilot CLI** and **Codex CLI** - all three install from the same `pipeline/` source and get the same
|
|
255
|
+
The pipeline runs natively on **Claude Code**, **Copilot CLI** and **Codex CLI** - all three install from the same `pipeline/` source and get the same 56 commands.
|
|
211
256
|
|
|
212
257
|
| Tool | Flag | What it installs |
|
|
213
258
|
| ----------- | -------------------- | ------------------------------------------------------------------------------------------------------ |
|
|
214
259
|
| Claude Code | `--claude` (default) | slash commands + skills + agents + three `PreToolUse` hooks (secret scan, agent-guard, read-size gate) |
|
|
215
|
-
| Copilot CLI | `--copilot` | instructions +
|
|
216
|
-
| Codex CLI | `--codex` | one router skill +
|
|
260
|
+
| Copilot CLI | `--copilot` | instructions + 56 sub-command skills + scripts |
|
|
261
|
+
| Codex CLI | `--codex` | one router skill + 56 specs as refs + 8 agent TOML + `AGENTS.md` block + `codex mcp add` |
|
|
217
262
|
|
|
218
263
|
Filter skills by stack with `--platform=ios\|android\|all`.
|
|
219
264
|
|
package/README.tr.md
CHANGED
|
@@ -89,7 +89,7 @@ Koşunun kendisinde ayarlanabilen tek şey derinlik, autopilot ve `--local`; ger
|
|
|
89
89
|
|
|
90
90
|
## Komutlar
|
|
91
91
|
|
|
92
|
-
`/multi-agent` ve
|
|
92
|
+
`/multi-agent` ve 56 alt komut. `/multi-agent:help` aynı katalogu terminalde, `outputLanguage` ayarına göre gösterir.
|
|
93
93
|
|
|
94
94
|
### Pipeline girişleri
|
|
95
95
|
|
|
@@ -195,6 +195,52 @@ Komutlarla birlikte sekiz agent kurulur ve fazlar bunları çağırır: `explore
|
|
|
195
195
|
|
|
196
196
|
İki uyumluluk skill'i her host'a kurulur ve store kapılarını besler: `apple-archive-compliance` (ITMS kod eşlemeli 18 kurallı Apple review taraması) ve `google-play-compliance` (Console hata kodlu 21 kurallı Play politika kataloğu). Stack'e bağlı geri kalan her şey - SwiftUI, Compose, backend, frontend - aşağıda anlatılan marketplace plugin'lerinden gelir.
|
|
197
197
|
|
|
198
|
+
## Sürekli mod
|
|
199
|
+
|
|
200
|
+
Yukarıdaki her şey sen başlattığında başlar. Sürekli mod, aynı pipeline'ın
|
|
201
|
+
**senin seçtiğin tek makinede**, **senin seçtiğin repolardan** işi kendi başına
|
|
202
|
+
alması.
|
|
203
|
+
|
|
204
|
+
| Komut | Ne yapar |
|
|
205
|
+
| --- | --- |
|
|
206
|
+
| `/multi-agent:autopilot-on` | Bu makinenin izleyeceği repoları seçersin. Etiketli GitHub issue'ları ve sana atanmış + etiketli Jira maddeleri worktree'de koşar, açık PR'da durur. Listeyi değiştirmek için tekrar çalıştır |
|
|
207
|
+
| `/multi-agent:autopilot-status` | Ne koşuyor hangi fazda, sırada ne var, ne cevap bekliyor, son bir günün PR'ları ve yuvarlanan harcama |
|
|
208
|
+
| `/multi-agent:autopilot-off` | Zamanlamayı kaldırır. Koşan iş biter; repo seçimi saklanır |
|
|
209
|
+
|
|
210
|
+
**Varsayılan olarak hiçbir şey açık değil ve hiçbir repo kendiliğinden eklenmez.**
|
|
211
|
+
Paketi kurmak ne durum yazar ne zamanlama kurar; bu bir gün değişirse
|
|
212
|
+
`smoke-autopilot-default-off.sh` derlemeyi düşürür. Etiket bir filtredir, kapı
|
|
213
|
+
değil - push yetkin olan bir repoda issue açabilen herkes etiket ekleyebilir -
|
|
214
|
+
o yüzden kapı picker'dır ve makine başınadır.
|
|
215
|
+
|
|
216
|
+
En büyük kazanç paralellik değil. Ölçüm: medyan koşu 44 dakika, ama 14:00'te
|
|
217
|
+
biten bir madde sen masaya oturana kadar bekler; gece boyunca bu 40 dakikaya
|
|
218
|
+
karşı **16 saat**. Sürekli mod işi değil, beklemeyi kaldırıyor.
|
|
219
|
+
|
|
220
|
+
**Yapmayacakları.** Merge etmez - koşucu açık PR'da durur, karar sende kalır.
|
|
221
|
+
Senin elle koşturduğun bir işe dokunmaz: repo başına eşzamanlılık her zaman 1,
|
|
222
|
+
yani kuyruk senin çalıştığın repoyu `.git/index.lock` için yarışmak yerine atlar.
|
|
223
|
+
PR sayısına tavan yok; sınırlar 24 saatlik yuvarlanan `costCeilingUsd` ve
|
|
224
|
+
makinenin kapasitesi.
|
|
225
|
+
|
|
226
|
+
**`swiftc` varsa menü çubuğu göstergesi** aynı `status.json`'ı sağ üstte çizer ve
|
|
227
|
+
kendi kendine tazelenir: madde başına tek satır - id, Full ya da Short, kesir
|
|
228
|
+
olarak faz, geçen süre, stack. Madde biter bitmez satır kaybolur ve PR'ıyla
|
|
229
|
+
Raporlar altında görünür. Yalnızca **çizer**; bir koşuyu başlatamaz, durduramaz,
|
|
230
|
+
değiştiremez. ActivityKit macOS'ta yok, o yüzden bu bir `NSStatusItem` ve
|
|
231
|
+
imzalanması gerekecek bir ikili olarak değil, yerinde kaynaktan derleniyor.
|
|
232
|
+
|
|
233
|
+
**Yeniden başlatmadan sonra çalıştırman gereken komut yok.** launchd işi
|
|
234
|
+
**girişte** yükler, boot'ta değil - ve bu bir eksiklik değil doğrusu: token'ları
|
|
235
|
+
açan şey login keychain'in açılması, giriş öncesi ateşlenen bir tick zaten ne
|
|
236
|
+
Jira'ya ne GitHub'a bağlanabilirdi. `autopilot-resume` diye bir komut yok -
|
|
237
|
+
hatırlanması gereken komut, unutulduğunda sessizce duran kuyruk demektir.
|
|
238
|
+
Bunun yerine `doctor` gerçek arızayı söyler: kurulu, ama launchd'de iş yok.
|
|
239
|
+
|
|
240
|
+
Uyku **yalnızca prizdeyken** engellenir. Prizde olmayan bir laptop'u boşaltan
|
|
241
|
+
kuyruk hatadır; bataryada assertion bırakılır ve prize takınca kaldığı yerden
|
|
242
|
+
devam eder.
|
|
243
|
+
|
|
198
244
|
## Stack'ler
|
|
199
245
|
|
|
200
246
|
Stack skill'leri [`mmerterden/multi-agent-plugins`](https://github.com/mmerterden/multi-agent-plugins) marketplace'inde versiyonlu plugin'ler olarak gönderilir. Repo başına bir stack seç:
|
|
@@ -207,13 +253,13 @@ Bu, ilgili plugin'i (+ ortak `ai-common` plugin'ini) repo'nun `.claude/settings.
|
|
|
207
253
|
|
|
208
254
|
## Araç desteği
|
|
209
255
|
|
|
210
|
-
Pipeline **Claude Code**, **Copilot CLI** ve **Codex CLI** üzerinde native çalışır - üçü de aynı `pipeline/` kaynağından kurulur ve aynı
|
|
256
|
+
Pipeline **Claude Code**, **Copilot CLI** ve **Codex CLI** üzerinde native çalışır - üçü de aynı `pipeline/` kaynağından kurulur ve aynı 56 komutu alır.
|
|
211
257
|
|
|
212
258
|
| Araç | Bayrak | Ne kurar |
|
|
213
259
|
| ----------- | ----------------------- | ---------------------------------------------------------------------------------------------------------------- |
|
|
214
260
|
| Claude Code | `--claude` (varsayılan) | slash komutları + skill'ler + agent'lar + üç `PreToolUse` hook'u (secret scan, agent-guard, okuma-boyutu geçidi) |
|
|
215
|
-
| Copilot CLI | `--copilot` | talimatlar +
|
|
216
|
-
| Codex CLI | `--codex` | bir router skill + ref olarak
|
|
261
|
+
| Copilot CLI | `--copilot` | talimatlar + 56 alt-komut skill'i + script'ler |
|
|
262
|
+
| Codex CLI | `--codex` | bir router skill + ref olarak 56 spec + 8 agent TOML + `AGENTS.md` bloğu + `codex mcp add` |
|
|
217
263
|
|
|
218
264
|
Skill'leri stack'e göre filtrele: `--platform=ios\|android\|all`.
|
|
219
265
|
|
package/docs/architecture.md
CHANGED
|
@@ -117,7 +117,7 @@ graph TB
|
|
|
117
117
|
end
|
|
118
118
|
|
|
119
119
|
subgraph "Pipeline Specs"
|
|
120
|
-
CMD[commands/<br/>
|
|
120
|
+
CMD[commands/<br/>56 command files]
|
|
121
121
|
AGT[agents/<br/>8 agent personas]
|
|
122
122
|
RUL[rules/<br/>12 domain rules]
|
|
123
123
|
PHS[multi-agent-refs/phases/<br/>phase specs + contracts]
|
|
@@ -169,8 +169,8 @@ revisions of this diagram - Codex CLI and the two independently-shipped repos
|
|
|
169
169
|
```mermaid
|
|
170
170
|
graph TD
|
|
171
171
|
CC["Claude Code<br/>(source of truth)"]
|
|
172
|
-
COP["Copilot CLI<br/>(instructions +
|
|
173
|
-
COD["Codex CLI<br/>(1 router skill +
|
|
172
|
+
COP["Copilot CLI<br/>(instructions + 56 skills)"]
|
|
173
|
+
COD["Codex CLI<br/>(1 router skill + 56 refs)"]
|
|
174
174
|
REPO["Pipeline Repo<br/>(npm package)"]
|
|
175
175
|
WEB["Website"]
|
|
176
176
|
PLUGREPO["multi-agent-plugins<br/>(5 stack plugins, own repo)"]
|
package/docs/ecosystem.md
CHANGED
|
@@ -5,7 +5,7 @@ separately, wired together at install time and at run time:
|
|
|
5
5
|
|
|
6
6
|
| Repo | What it owns | Ships as |
|
|
7
7
|
|---|---|---|
|
|
8
|
-
| **`multi-agent-pipeline`** (this repo) | Orchestration: the 8-phase flow, the
|
|
8
|
+
| **`multi-agent-pipeline`** (this repo) | Orchestration: the 8-phase flow, the 56 slash commands, quality gates, review/triage, cross-CLI parity | npm package (`@mmerterden/multi-agent-pipeline`), installs itself onto Claude Code / Copilot CLI / Codex CLI |
|
|
9
9
|
| **`multi-agent-plugins`** | Stack knowledge: per-platform component/lifecycle skills (iOS, Android, Frontend, Backend) + shared knowledge | Claude Code marketplace, 5 independently-versioned plugins |
|
|
10
10
|
| **`multi-agent-toolkit-mcp`** | The pipeline's hands on devices and browsers: 80 MCP tools across 6 categories (simulator/emulator control, accessibility audit, store compliance, web automation, Figma-vs-mock design audit, an agent-DSL batch runner) | npm package, registered as a standard stdio MCP server on every host |
|
|
11
11
|
|
|
@@ -18,7 +18,7 @@ Either can be swapped or removed without touching the other two's source.
|
|
|
18
18
|
graph LR
|
|
19
19
|
subgraph PIPE ["multi-agent-pipeline (orchestrator)"]
|
|
20
20
|
direction TB
|
|
21
|
-
PHASES["8 phases ·
|
|
21
|
+
PHASES["8 phases · 56 commands"]
|
|
22
22
|
GATES["deterministic gates + review triage"]
|
|
23
23
|
end
|
|
24
24
|
|
|
@@ -64,8 +64,8 @@ only those:
|
|
|
64
64
|
graph TD
|
|
65
65
|
CC["Claude Code<br/>~/.claude/commands/multi-agent/<br/>(source of truth)"]
|
|
66
66
|
|
|
67
|
-
CC -->|"Step 2: copy + reformat<br/>
|
|
68
|
-
CC -->|"Step 2b: transform<br/>(install.js --codex)"| COD["Codex CLI<br/>1 router skill +
|
|
67
|
+
CC -->|"Step 2: copy + reformat<br/>56 sub-command skills"| COP["Copilot CLI<br/>~/.copilot/skills/"]
|
|
68
|
+
CC -->|"Step 2b: transform<br/>(install.js --codex)"| COD["Codex CLI<br/>1 router skill + 56 refs<br/>+ 8 agent TOML"]
|
|
69
69
|
CC -->|"Step 3: genericize<br/>(strip personal data)"| REPO["multi-agent-pipeline repo<br/>pipeline/"]
|
|
70
70
|
CC -->|"Step 4: version + feature sync"| WEB["Website<br/>projects.ts / i18n.tsx"]
|
|
71
71
|
|
|
@@ -153,7 +153,7 @@ measurements behind this table):
|
|
|
153
153
|
|
|
154
154
|
| | Claude Code | Copilot CLI | Codex CLI |
|
|
155
155
|
|---|---|---|---|
|
|
156
|
-
| **Pipeline commands** |
|
|
156
|
+
| **Pipeline commands** | 56 slash-command skills, native | 56 skills, `multi-agent-{cmd}` naming, copied in | 1 router skill (`multi-agent`) + 56 command specs as reference files - Codex silently truncates its skills block past a few dozen entries, so sub-commands are not peer skills here |
|
|
157
157
|
| **Stack plugins** | Marketplace plugin, loaded natively, resolved by `.claude/settings.json` enabled-list | Enabled plugin's authored skills copied flat into `~/.copilot/skills/`; `knowledge/` **not** re-copied (already delivered via `shared/external`) | Copied as reference files under `~/.codex/multi-agent-refs/skills/`, plugin-prefixed on name clash (e.g. `architecture` → `ai-ios-toolkit-architecture`) |
|
|
158
158
|
| **Component dispatch (Phase 3)** | Marketplace plugin's `create-component`/`create-screen` skill via the Skill tool | No plugin loader - the enabled stack plugin's authored skills (incl. `create-component`) are copied flat into `~/.copilot/skills/` at install time (the old frozen `figma-*` copies are pruned, they were never a fallback) | Not part of the enforced parity axis; classification + state-shape must match, skill *inventory* does not |
|
|
159
159
|
| **multi-agent-toolkit-mcp** | `claude mcp add multi-agent-toolkit -- npx -y @mmerterden/multi-agent-toolkit-mcp` | `copilot mcp add multi-agent-toolkit -- npx -y @mmerterden/multi-agent-toolkit-mcp` | `codex mcp add multi-agent-toolkit -- npx -y @mmerterden/multi-agent-toolkit-mcp` (skipped with a warning if `codex` isn't on `PATH`) |
|
|
@@ -0,0 +1,79 @@
|
|
|
1
|
+
<?xml version="1.0" encoding="UTF-8"?>
|
|
2
|
+
<!--
|
|
3
|
+
multi-agent autopilot - the schedule for continuous mode.
|
|
4
|
+
|
|
5
|
+
Rendered and installed ONLY by /multi-agent:autopilot-on, after an explicit
|
|
6
|
+
confirmation that names this file's path. A clean install writes nothing here;
|
|
7
|
+
smoke-autopilot-default-off.sh fails if it ever does.
|
|
8
|
+
|
|
9
|
+
A user agent, not a daemon, and that is the correct choice rather than a
|
|
10
|
+
limitation: it loads at LOGIN. Before login the keychain is locked, so a tick
|
|
11
|
+
that fired at boot could not read a single token - it would take items it could
|
|
12
|
+
not work on and burn them. Automatic login means nothing to do after a restart.
|
|
13
|
+
|
|
14
|
+
StartInterval does NOT wake a sleeping Mac; it fires on the next wake. The mode
|
|
15
|
+
therefore holds its own sleep assertion while on AC and releases it on battery,
|
|
16
|
+
so an overnight queue keeps moving without flattening a laptop off the charger.
|
|
17
|
+
|
|
18
|
+
{{PLACEHOLDERS}} are substituted by autopilot-on: NODE, TICK, LABEL, ROOT,
|
|
19
|
+
INTERVAL.
|
|
20
|
+
-->
|
|
21
|
+
<plist version="1.0">
|
|
22
|
+
<dict>
|
|
23
|
+
<key>Label</key>
|
|
24
|
+
<string>{{LABEL}}</string>
|
|
25
|
+
|
|
26
|
+
<key>ProgramArguments</key>
|
|
27
|
+
<array>
|
|
28
|
+
<string>{{NODE}}</string>
|
|
29
|
+
<string>{{TICK}}</string>
|
|
30
|
+
</array>
|
|
31
|
+
|
|
32
|
+
<!-- Load at login and run one tick immediately, so turning the mode on does
|
|
33
|
+
something visible rather than waiting out the first interval. -->
|
|
34
|
+
<key>RunAtLoad</key>
|
|
35
|
+
<true/>
|
|
36
|
+
|
|
37
|
+
<key>StartInterval</key>
|
|
38
|
+
<integer>{{INTERVAL}}</integer>
|
|
39
|
+
|
|
40
|
+
<!-- One tick at a time. Overlapping ticks would each claim the queue head;
|
|
41
|
+
the runner's own pid+sessionId check is the real mutual exclusion, and
|
|
42
|
+
this keeps the common case from ever reaching it. -->
|
|
43
|
+
<key>AbandonProcessGroup</key>
|
|
44
|
+
<false/>
|
|
45
|
+
|
|
46
|
+
<key>ProcessType</key>
|
|
47
|
+
<string>Background</string>
|
|
48
|
+
|
|
49
|
+
<!-- Nice to the interactive session. The point of this mode is work that
|
|
50
|
+
happens while you do something else, so it must never be what makes the
|
|
51
|
+
machine feel slow. -->
|
|
52
|
+
<key>LowPriorityIO</key>
|
|
53
|
+
<true/>
|
|
54
|
+
<key>Nice</key>
|
|
55
|
+
<integer>5</integer>
|
|
56
|
+
|
|
57
|
+
<!-- Both streams go to the runner's own log, inside the 0700 state directory.
|
|
58
|
+
Not /tmp: a tick's output names tickets, repos and branches, and on a
|
|
59
|
+
shared machine that is somebody's roadmap. Every file the runner writes
|
|
60
|
+
is scanned for credentials before it is kept. -->
|
|
61
|
+
<key>StandardOutPath</key>
|
|
62
|
+
<string>{{ROOT}}/runner.log</string>
|
|
63
|
+
<key>StandardErrorPath</key>
|
|
64
|
+
<string>{{ROOT}}/runner.log</string>
|
|
65
|
+
|
|
66
|
+
<key>EnvironmentVariables</key>
|
|
67
|
+
<dict>
|
|
68
|
+
<key>PATH</key>
|
|
69
|
+
<string>/usr/local/bin:/opt/homebrew/bin:/usr/bin:/bin:/usr/sbin:/sbin</string>
|
|
70
|
+
<!-- No token is set here, deliberately. launchd plists are world-readable
|
|
71
|
+
and get backed up; credentials are read from the keychain at tick time
|
|
72
|
+
through credential-store.sh, which streams them over stdin so they never
|
|
73
|
+
reach argv either. -->
|
|
74
|
+
</dict>
|
|
75
|
+
|
|
76
|
+
<key>WorkingDirectory</key>
|
|
77
|
+
<string>{{ROOT}}</string>
|
|
78
|
+
</dict>
|
|
79
|
+
</plist>
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@mmerterden/multi-agent-pipeline",
|
|
3
|
-
"version": "17.
|
|
3
|
+
"version": "17.1.0",
|
|
4
4
|
"description": "8-phase AI development pipeline with full orchestration on Claude Code, Copilot CLI and Codex CLI. Analysis, planning, TDD, CLI-aware parallel review with consensus surfacing + Fable triage, default-FAIL evidence gates, secret + intent guards, per-phase cost ledger, persistent learnings memory, wiki generation, commit automation. Token-preserving uninstall.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"main": "index.js",
|
|
@@ -0,0 +1,64 @@
|
|
|
1
|
+
---
|
|
2
|
+
description: "Turn off continuous mode on this machine: the schedule is removed, work already running finishes, the repo selection is kept. Use when the machine should stop picking work up."
|
|
3
|
+
description-tr: "Bu makinede sürekli modu kapatır: zamanlama kaldırılır, koşan iş biter, repo seçimi saklanır."
|
|
4
|
+
argument-hint: "[--now] - --now also stops the item currently running"
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# multi-agent autopilot-off - stop picking work up
|
|
8
|
+
|
|
9
|
+
Removes the schedule. **Work already in flight finishes** unless you pass
|
|
10
|
+
`--now`: an item stopped mid-development leaves a worktree and a half-written
|
|
11
|
+
branch, which is the exact state this release spent its time cleaning up.
|
|
12
|
+
|
|
13
|
+
## Steps
|
|
14
|
+
|
|
15
|
+
### 1. Say what is in flight before stopping anything
|
|
16
|
+
|
|
17
|
+
```bash
|
|
18
|
+
bash "$HOME/.claude/scripts/autopilot-status.sh" --short
|
|
19
|
+
```
|
|
20
|
+
|
|
21
|
+
When something is running, name it and how long it has been going. "Stopped"
|
|
22
|
+
means something different when an item is 2 minutes from a PR than when the queue
|
|
23
|
+
is idle, and the user is the one who knows which.
|
|
24
|
+
|
|
25
|
+
### 2. Remove the schedule
|
|
26
|
+
|
|
27
|
+
```bash
|
|
28
|
+
L="com.multi-agent.autopilot"
|
|
29
|
+
launchctl bootout "gui/$(id -u)/$L" 2>/dev/null || launchctl unload "$HOME/Library/LaunchAgents/$L.plist" 2>/dev/null
|
|
30
|
+
rm -f "$HOME/Library/LaunchAgents/$L.plist"
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
The plist is removed rather than left disabled, because its presence is the
|
|
34
|
+
definition of on: a disabled-but-present job is a third state nobody can read
|
|
35
|
+
from the outside.
|
|
36
|
+
|
|
37
|
+
### 3. Release the sleep assertion and the indicator
|
|
38
|
+
|
|
39
|
+
```bash
|
|
40
|
+
pkill -f 'caffeinate -i -w' 2>/dev/null || true # only the one this mode holds
|
|
41
|
+
pkill -f "$HOME/.claude/autopilot/bin/menubar" 2>/dev/null || true
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
### 4. `--now` only: end the in-flight item deliberately
|
|
45
|
+
|
|
46
|
+
Without `--now` nothing here runs. With it, the child session is stopped and the
|
|
47
|
+
run is marked `abandoned`, its worktree is removed **unless it holds uncommitted
|
|
48
|
+
work**, in which case the work is stashed to `autopilot/abandoned/<task-id>` and
|
|
49
|
+
the worktree is kept. Same contract as `gc-abandoned.sh`; losing a day of edits
|
|
50
|
+
is worse than 750 MB.
|
|
51
|
+
|
|
52
|
+
### 5. Keep the selection
|
|
53
|
+
|
|
54
|
+
`~/.claude/autopilot/config.json`, `queue.json` and `attempted.jsonl` all stay.
|
|
55
|
+
Turning the mode back on must not re-ask which repos, and the attempt history is
|
|
56
|
+
what stops an item that already failed twice from being retried forever.
|
|
57
|
+
|
|
58
|
+
To forget the selection as well: `rm -rf ~/.claude/autopilot`. Say that rather
|
|
59
|
+
than doing it - "off" and "forget everything" are different requests.
|
|
60
|
+
|
|
61
|
+
### 6. Report
|
|
62
|
+
|
|
63
|
+
State that it is off, what was left running or finishing, and that the repo
|
|
64
|
+
selection is kept. If an item was stashed, name the branch.
|
|
@@ -0,0 +1,173 @@
|
|
|
1
|
+
---
|
|
2
|
+
description: "Turn on continuous mode on THIS machine: pick the repos it watches; labelled GitHub issues and assigned+labelled Jira items then run in a worktree and stop at an open PR. Use when the machine should pick work up unattended."
|
|
3
|
+
description-tr: "Bu makinede sürekli modu açar: izlenecek repoları seçersin, etiketli GitHub issue'ları ve sana atanmış + etiketli Jira maddeleri worktree'de geliştirilip PR'da durur. Repo listesini değiştirmek için aynı komutu tekrar çalıştır."
|
|
4
|
+
argument-hint: "(no arguments - opens the repo picker)"
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# multi-agent autopilot-on - continuous mode, on this machine
|
|
8
|
+
|
|
9
|
+
Turns on the mode that picks work up without being asked. **No repo is included
|
|
10
|
+
by default and none is ever added implicitly** - you choose, here, every time.
|
|
11
|
+
|
|
12
|
+
Why a picker rather than a label alone: measured on one machine, 66 repos grant
|
|
13
|
+
push and a good half belong to other people. A label is a filter; anyone who can
|
|
14
|
+
open an issue in a repo you happen to have push on could add one. The gate is
|
|
15
|
+
this selection, and it is per machine.
|
|
16
|
+
|
|
17
|
+
## Steps
|
|
18
|
+
|
|
19
|
+
### 1. Prerequisites, stated before anything is written
|
|
20
|
+
|
|
21
|
+
```bash
|
|
22
|
+
gh auth status >/dev/null 2>&1 || echo "BLOCKED: gh is not authenticated - run 'gh auth login'"
|
|
23
|
+
command -v jq >/dev/null 2>&1 || echo "BLOCKED: jq is required"
|
|
24
|
+
node "$HOME/.claude/scripts/doctor.mjs" >/dev/null 2>&1; [ "$?" -ge 2 ] && echo "BLOCKED: doctor reports a blocking problem - fix it first"
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
Any BLOCKED line stops here and is reported. Arming unattended work on an install
|
|
28
|
+
that `doctor` calls broken produces a queue of failures, one per item.
|
|
29
|
+
|
|
30
|
+
### 2. The repo list, both halves of it
|
|
31
|
+
|
|
32
|
+
```bash
|
|
33
|
+
bash "$HOME/.claude/lib/autopilot-activation.sh" list
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
Two states come back and **both are shown**:
|
|
37
|
+
|
|
38
|
+
- `eligible` - push rights and a checkout on this machine
|
|
39
|
+
- `unavailable` - push rights, no local checkout, listed WITH that reason
|
|
40
|
+
|
|
41
|
+
Never hide the second list. Dropping 51 of 66 rows silently looks exactly like a
|
|
42
|
+
permissions problem, and the user goes hunting for a token fault that does not
|
|
43
|
+
exist.
|
|
44
|
+
|
|
45
|
+
### 3. Pick, with the current selection already checked
|
|
46
|
+
|
|
47
|
+
Read `~/.claude/autopilot/config.json` if it exists, then one `AskUserQuestion`
|
|
48
|
+
(`multiSelect: true`) over the **eligible** rows:
|
|
49
|
+
|
|
50
|
+
- `question` (in `outputLanguage`): which repos this machine should pick work up from
|
|
51
|
+
- `header`: "Repos" (English, UI contract)
|
|
52
|
+
- `options`: one per eligible repo, currently-selected ones marked "(seçili)" / "(selected)"
|
|
53
|
+
|
|
54
|
+
**This is how repos are added and removed.** Checking a new one adds it;
|
|
55
|
+
unchecking an existing one removes it. There is no `add-repo` / `remove-repo`
|
|
56
|
+
command, because two entry points to one list is two lists.
|
|
57
|
+
|
|
58
|
+
Unavailable repos are printed above the question as a plain list with their
|
|
59
|
+
reason, not offered as options.
|
|
60
|
+
|
|
61
|
+
### 4. Sources and label, per selected repo
|
|
62
|
+
|
|
63
|
+
One `AskUserQuestion` per newly added repo (already-configured ones keep their
|
|
64
|
+
answers - re-running must not re-ask what was already settled):
|
|
65
|
+
|
|
66
|
+
- Sources: `GitHub issues` / `Jira` / both
|
|
67
|
+
- Label: default **`agent-queue`** on both sides, editable
|
|
68
|
+
|
|
69
|
+
The same name works on both sides and both mechanisms already exist:
|
|
70
|
+
|
|
71
|
+
| Source | Query | Extra condition |
|
|
72
|
+
|---|---|---|
|
|
73
|
+
| GitHub | `gh issue list --repo <r> --label <label> --state open` | none - the label is the whole gate |
|
|
74
|
+
| Jira | `assignee = currentUser() AND labels = "<label>" AND resolution = EMPTY AND status not in (Done, Closed, Cancelled)` | **assigned to you**, so a label somebody else adds is not enough |
|
|
75
|
+
|
|
76
|
+
Two Jira caveats worth saying out loud when Jira is selected: labels are one
|
|
77
|
+
global namespace across the whole instance and anyone can write them, so on a
|
|
78
|
+
shared instance prefer a qualified name (`agent-queue-<team>`); and the Labels
|
|
79
|
+
field has to be on the edit screen for the issue types in play. An instance where
|
|
80
|
+
neither holds is a `jiraJql` config change - a saved filter, a component - not a
|
|
81
|
+
redesign.
|
|
82
|
+
|
|
83
|
+
Run the query before saving it, and show the count:
|
|
84
|
+
|
|
85
|
+
```bash
|
|
86
|
+
bash "$HOME/.claude/scripts/jira-search.sh" --jql "$JQL" --max 5 | jq '.total, [.issues[].key]'
|
|
87
|
+
```
|
|
88
|
+
|
|
89
|
+
A JQL that returns nothing is not an error and is usually correct on day one
|
|
90
|
+
(nothing is labelled yet), but a JQL that errors is a config that will be silent
|
|
91
|
+
forever - the Labels field missing from an edit screen looks exactly like an
|
|
92
|
+
empty queue from the outside. Saying "0 items, query valid" and "query rejected"
|
|
93
|
+
differently here is the difference between a two-minute fix and a week of
|
|
94
|
+
wondering why nothing happens.
|
|
95
|
+
|
|
96
|
+
### 5. Write the config
|
|
97
|
+
|
|
98
|
+
Validate against `schemas/autopilot-config.schema.json`, then write
|
|
99
|
+
`~/.claude/autopilot/config.json` (0600, in a 0700 directory) via
|
|
100
|
+
`ma_ap_write` in `lib/autopilot-state.sh`. Defaults that are not asked:
|
|
101
|
+
`slots: 1`, `scanIntervalSeconds: 120`, `maxAttempts: 2`,
|
|
102
|
+
`askOnLowMaturity: true`, `costCeilingUsd: 25`, `depthRouter: "off"`.
|
|
103
|
+
|
|
104
|
+
### 6. Build the menu bar indicator (best effort)
|
|
105
|
+
|
|
106
|
+
```bash
|
|
107
|
+
SRC="$HOME/.claude/scripts/autopilot-menubar.swift"
|
|
108
|
+
OUT="$HOME/.claude/autopilot/bin/menubar"
|
|
109
|
+
if command -v swiftc >/dev/null 2>&1; then
|
|
110
|
+
swiftc -O "$SRC" -o "$OUT" 2>/dev/null && chmod 700 "$OUT"
|
|
111
|
+
fi
|
|
112
|
+
```
|
|
113
|
+
|
|
114
|
+
Built from source on demand, never shipped as a binary: the package contains only
|
|
115
|
+
text today, and a compiled artifact in it would need signing and notarization to
|
|
116
|
+
be distributable. No `swiftc` means no indicator and nothing else changes - say
|
|
117
|
+
so in one line rather than failing. `autopilot-status` works either way.
|
|
118
|
+
|
|
119
|
+
### 7. Confirm the schedule, naming the file
|
|
120
|
+
|
|
121
|
+
Render `templates/multi-agent-autopilot.plist.template` by substituting its five
|
|
122
|
+
placeholders, then show the resolved **path** and interval and `AskUserQuestion`:
|
|
123
|
+
|
|
124
|
+
```bash
|
|
125
|
+
TPL="$HOME/.claude/templates/multi-agent-autopilot.plist.template"
|
|
126
|
+
DEST="$HOME/Library/LaunchAgents/com.multi-agent.autopilot.plist"
|
|
127
|
+
INTERVAL=$(bash "$HOME/.claude/scripts/autopilot-status.sh" --json | jq -r '.scanIntervalSeconds // 120')
|
|
128
|
+
sed -e "s|{{NODE}}|$(command -v node)|g" \
|
|
129
|
+
-e "s|{{TICK}}|$HOME/.claude/scripts/autopilot-runner.mjs|g" \
|
|
130
|
+
-e "s|{{LABEL}}|com.multi-agent.autopilot|g" \
|
|
131
|
+
-e "s|{{ROOT}}|$HOME/.claude/autopilot|g" \
|
|
132
|
+
-e "s|{{INTERVAL}}|$INTERVAL|g" \
|
|
133
|
+
"$TPL" > "$DEST"
|
|
134
|
+
```
|
|
135
|
+
|
|
136
|
+
`command -v node` rather than the bare word: launchd does not read a login
|
|
137
|
+
shell, so its `PATH` is the short one in the plist's own `EnvironmentVariables`,
|
|
138
|
+
and a Homebrew or nvm node that is not on that path would make every tick fail
|
|
139
|
+
with a message nobody is watching for.
|
|
140
|
+
|
|
141
|
+
- `question`: install the schedule at `~/Library/LaunchAgents/com.multi-agent.autopilot.plist` and start picking work up?
|
|
142
|
+
- options: `{ label: "Start" }` / `{ label: "Configure only", description: "Save the repo selection, do not schedule anything yet" }`
|
|
143
|
+
|
|
144
|
+
"Configure only" is a real answer and the safe default to offer first: the queue
|
|
145
|
+
can be watched with `autopilot-status` for a while before anything runs.
|
|
146
|
+
|
|
147
|
+
On **Start**:
|
|
148
|
+
|
|
149
|
+
```bash
|
|
150
|
+
launchctl bootstrap "gui/$(id -u)" "$HOME/Library/LaunchAgents/com.multi-agent.autopilot.plist" 2>/dev/null \
|
|
151
|
+
|| launchctl load "$HOME/Library/LaunchAgents/com.multi-agent.autopilot.plist"
|
|
152
|
+
```
|
|
153
|
+
|
|
154
|
+
### 8. Report what is now true
|
|
155
|
+
|
|
156
|
+
Say all five, because each one is something a user has asked about after the
|
|
157
|
+
fact: which repos and which labels; that it survives a restart with no command to
|
|
158
|
+
run (launchd loads at **login** - not boot, and that is correct, because the login
|
|
159
|
+
keychain is what unlocks the tokens, so a run before login could not reach Jira or
|
|
160
|
+
GitHub anyway); that the machine will be kept awake **only on AC**; the slot count
|
|
161
|
+
and the ceiling this machine would allow (`ma_ap_slot_ceiling`); and that
|
|
162
|
+
`/multi-agent:autopilot-off` is the only way to stop it.
|
|
163
|
+
|
|
164
|
+
## What this does not do
|
|
165
|
+
|
|
166
|
+
- **No cap on PRs.** As many items as carry the label become as many PRs. The
|
|
167
|
+
bounds are `costCeilingUsd` over a rolling 24 hours and what the machine can
|
|
168
|
+
hold - never a daily count.
|
|
169
|
+
- **Does not merge.** The runner stops at an open PR; the merge decision stays
|
|
170
|
+
yours.
|
|
171
|
+
- **Does not touch an attended run.** Both are worktrees and per-repo concurrency
|
|
172
|
+
is 1, so the queue skips a repo you are working in rather than competing for
|
|
173
|
+
`.git/index.lock`.
|