@mmerterden/multi-agent-pipeline 16.31.1 → 17.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +78 -0
- package/README.md +101 -101
- package/README.tr.md +101 -101
- package/docs/adr/0011-dormant-ci.md +10 -1
- package/docs/adr/0012-macos-only.md +98 -0
- package/docs/adr/README.md +1 -0
- package/docs/engineering.md +1 -1
- package/index.js +26 -0
- package/install/_dev-only-files.mjs +0 -1
- package/install/index.mjs +10 -0
- package/package.json +5 -3
- package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +40 -3
- package/pipeline/commands/multi-agent/help/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/setup/SKILL.md +15 -7
- package/pipeline/commands/multi-agent/stack/SKILL.md +31 -32
- package/pipeline/commands/multi-agent/status/SKILL.md +17 -1
- package/pipeline/commands/multi-agent/sync/SKILL.md +24 -19
- package/pipeline/lib/stack-detect.sh +200 -0
- package/pipeline/multi-agent-refs/cross-cli-contract.md +29 -10
- package/pipeline/multi-agent-refs/features/doctor.md +15 -3
- package/pipeline/multi-agent-refs/features/visual-evidence.md +49 -1
- package/pipeline/multi-agent-refs/phases/phase-0-init.md +14 -5
- package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +26 -12
- package/pipeline/multi-agent-refs/phases/phase-3-dev.md +6 -4
- package/pipeline/multi-agent-refs/phases/phase-7-report.md +2 -3
- package/pipeline/schemas/agent-state.schema.json +99 -25
- package/pipeline/schemas/token-budget.json +4 -4
- package/pipeline/scripts/_stack-routing.mjs +91 -0
- package/pipeline/scripts/capture-resume.sh +76 -14
- package/pipeline/scripts/doctor.mjs +26 -3
- package/pipeline/scripts/gc-abandoned.sh +352 -0
- package/pipeline/scripts/usage-report.mjs +5 -5
- package/pipeline/scripts/gate-linux.sh +0 -62
package/CHANGELOG.md
CHANGED
|
@@ -14,6 +14,84 @@ Internal file-layout changes that don't affect the slash-command surface are sti
|
|
|
14
14
|
|
|
15
15
|
---
|
|
16
16
|
|
|
17
|
+
## [17.0.0] - 2026-09-14
|
|
18
|
+
|
|
19
|
+
Five things in this release were not working, and four of them looked like they
|
|
20
|
+
were. A platform we advertised and never verified; a stack answer that sent
|
|
21
|
+
nearly every repo to the iOS toolkit; an evidence step gated on a field no phase
|
|
22
|
+
wrote, passing an argument no phase assigned; and one word, "in progress",
|
|
23
|
+
covering a finished job, a crashed one and a run that never started. The common
|
|
24
|
+
shape is a contract with readers and no writer, which reads exactly like a
|
|
25
|
+
contract that is satisfied.
|
|
26
|
+
|
|
27
|
+
### Fixed
|
|
28
|
+
|
|
29
|
+
- **"In progress" covered three different situations and offered one answer.** Measured: 20 runs read `in_progress` and every one was over a day old, the oldest 157 days. They were not one failure. 3 had their PR already open and were holding at Phase 6/7, where the pipeline pauses BY DESIGN for channel selection - finished work. 11 were left at a Phase 0 question, which is almost entirely questions, so nothing was ever built. 6 stopped mid-development, the only group that is actually broken. The session hook reported the newest of the 20 as "stopped at Phase N" and said nothing about the other 19, in the same words whether the pipeline was waiting for the user or had crashed.
|
|
30
|
+
|
|
31
|
+
`agent-state` gains `awaiting_input`, distinct from both `in_progress` and `paused`: the run is not broken and not resumable by a retry, it is finished with what it can do alone. A state file in the wild had already invented `awaiting-user-test-main-checkout` to say this. Phase 7 writes it where it pauses, the session hook and `/multi-agent:status` group on it, and each group offers its own action - resuming a Phase 0 question rebuilds nothing, and garbage-collecting an open PR throws away landed work.
|
|
32
|
+
|
|
33
|
+
A run with no status is now left out of all three groups. Eight such files exist here, six with no phase either, and calling them dead is the same false claim as calling a waiting run dead.
|
|
34
|
+
|
|
35
|
+
- **`.pr` is an object in the schema and a bare URL string in three state files on disk.** `jq '.pr.url'` on a string errors, and with stderr suppressed that reads as "no PR" - so a run whose PR was already open could be classified as dead. Both readers now read the shape instead of assuming it.
|
|
36
|
+
|
|
37
|
+
### Added
|
|
38
|
+
|
|
39
|
+
- **`/multi-agent:garbage-collect --abandoned`**, backed by `gc-abandoned.sh`. Nothing collected these: `gc-worktrees.sh` skips REGISTERED worktrees on purpose, because a registered worktree belongs to a live run - and a run that stopped never stops being live. 28 worktrees across 7 repos, 23 GB.
|
|
40
|
+
|
|
41
|
+
Two passes, because the state files are not where the disk is: only 5 of those 28 belong to a run that still reads `in_progress`, while 19 have no state file at all and 3 belong to runs that finished and outlived their worktree. A state-driven sweep alone reaches 4% of the problem.
|
|
42
|
+
|
|
43
|
+
Three rules, each written against a shape found on the machine rather than imagined. A run waiting on you is never reaped. A path outside `<repo>/.worktrees/` is never removed - four state files record the REPO ROOT as their `worktreePath`, so a sweep that trusted the field would have deleted a checkout. Uncommitted work is stashed to `autopilot/abandoned/<task-id>` and the worktree is kept, because losing a day of edits is worse than 750 MB.
|
|
44
|
+
|
|
45
|
+
And the rule that decides what the tool is for: **a worktree with no run state is reported, never removed.** `.worktrees/` is not exclusively ours - this machine holds `174`, `pr-4051` and `task-1` there, hand-made, one in active use - and nothing distinguishes those from a pipeline worktree whose log was pruned. They are the 11 GB, and naming them with their sizes is worth more than a rule that guesses.
|
|
46
|
+
|
|
47
|
+
Dry-run by default, `--yes` applies, state salvaged to `artifacts/` before anything goes.
|
|
48
|
+
|
|
49
|
+
### Fixed
|
|
50
|
+
|
|
51
|
+
- **The visual-evidence chain had five readers and no writer, so none of it ever ran.** Phase 0 Step 7.7 is gated on `state.visualEvidence.required`; Phase 3 captures when it is true; Phase 5 records the flow video; Phase 6 blocks on a required artefact that is neither attached nor explained. No phase document ever wrote `required`, `requiredBy` or `platform`, so the verdict was never true, and every consumer below it read a contract that looked satisfied. Step 7.7 now writes the verdict and is named in `features/visual-evidence.md` section 1a as its only writer. Phase 0 has no diff, so the `bugfix` row is provisional and Phase 3 re-decides from the real one.
|
|
52
|
+
|
|
53
|
+
- **The probe was called with `--platform "$PLATFORM"`, a variable no phase document assigns.** The probe refuses an empty value with exit 2, so the call could not have worked on its best day - it went unnoticed only because the step above was unreachable. The platform now comes from the stack (section 1b), which is the right source: it selects which device tooling to probe, and that is a property of the repo rather than of the diff. iOS wins a tie, recorded in `stackWhy`. A repo with no device platform does not get an empty string passed through; the probe does not run and `evidenceCapability.skippedReason` says why.
|
|
54
|
+
|
|
55
|
+
- `smoke-evidence-probe.sh` now fails a `--platform` argument whose variable is not assigned in the same file. Deliberately narrow: more than a hundred variables appear in these documents without an assignment and most are fine, because an agent fills `$WORKTREE` from context. `--platform` is not inferable and its consumer enforces a closed set with a hard exit, which is what makes the rule about this flag rather than about shell hygiene.
|
|
56
|
+
|
|
57
|
+
### Fixed
|
|
58
|
+
|
|
59
|
+
- **Nearly every checkout on the development machine was routed to the iOS toolkit** - 22 of 25, measured. Two causes, measured: `setup` wrote `ai-ios-toolkit` whenever its marker list matched nothing, and nothing else ever revisited the answer, so a Next.js site and several Node CLIs inherited it. The default is gone; an undetected repo now gets the two stack-independent toolkits and a recorded reason, which is an answer rather than a guess.
|
|
60
|
+
|
|
61
|
+
- **Android was never detected.** Phase 1's marker scan ran at `maxdepth 2`, and `AndroidManifest.xml` sits at `<module>/src/main/` - depth 4 in every multi-module app. The scan also read a generated `.next/package.json` as a repo's manifest, and reported a Compose app as iOS because the app vendors a submodule that ships a `Package.swift`.
|
|
62
|
+
|
|
63
|
+
### Added
|
|
64
|
+
|
|
65
|
+
- **`pipeline/lib/stack-detect.sh`** is now the one owner of "what is this repo built with". Deterministic file markers, never a model call: root before anything deeper so `find` traversal order cannot decide the answer, submodule and build-output pruning, depth 5 for the Android manifest, and `MA_STACK_WHY` separating "no marker matched" from "could not be read" - a caller that cannot tell those apart treats an unreadable repo as a language-free one. Verified against every repo on the development machine and seven fixtures.
|
|
66
|
+
|
|
67
|
+
- **`pluginsForStacks()` in `scripts/_stack-routing.mjs`** is the one owner of "which plugins does that want". The mapping was written out by hand in four places, each a chance to get the single asymmetric name wrong - the detector says `web`, the marketplace ships `ai-frontend-toolkit`. Derived names are checked against the locally installed marketplace manifest, so a toolkit that was never shipped is caught at derivation instead of presenting as one silently missing from the session. An unreadable manifest routes by convention and reports `manifestVerified: false`; it never fails closed.
|
|
68
|
+
|
|
69
|
+
- `state.stacks[]` and `state.stackWhy` for the stack axis. `state.detectedStack[]` stays the LANGUAGE axis and is now declared in the schema, which had `additionalProperties: false` and no such property while the phase documents instructed the write. The two axes are deliberately not merged: `graph-build.mjs --stack` accepts `ios|android|node|python|go` and has no notion of `web` or `backend`, and conflating them is what let a Gradle-built JVM service read as an Android app.
|
|
70
|
+
|
|
71
|
+
- **`smoke-stack-detect.sh`**, whose golden-output assertion is what made the consolidation safe: every argument the four hand-written tables accepted must still resolve to exactly the plugins they listed, or the gate fails.
|
|
72
|
+
|
|
73
|
+
### Removed
|
|
74
|
+
|
|
75
|
+
- **Linux and Windows support.** The package advertised three platforms and verified one. Every credential read shells `security`, every iOS build `xcodebuild`, every piece of visual evidence `simctl`, so a run elsewhere did not degrade into a smaller run - it failed partway through with a worktree and a branch already created. ADR-0011 already recorded that CI had not executed since 2026-07-02, and its own status block called the Linux path "smoke-skipped" and the Windows path "unverified".
|
|
76
|
+
|
|
77
|
+
The removal is a declaration, not a deletion. `package.json` declares `os: ["darwin"]`, which npm enforces on the root package: a non-macOS `npm ci` ends in `EBADPLATFORM` before a file is written. `index.js` and `install/index.mjs` refuse a non-darwin host with one line naming the requirement; `uninstall`, `help` and `--version` stay reachable, because somebody who installed before this gate must still be able to remove it. `MULTI_AGENT_ALLOW_NON_DARWIN=1` overrides and says on stderr that nothing on that path is tested.
|
|
78
|
+
|
|
79
|
+
The dead Linux and Windows branches are left in place deliberately. They cost nothing at runtime, and seven constructs in them look like cross-platform scaffolding while being load-bearing on macOS - the `grep -P` ban, the `sha256sum || shasum` ordering, the `stat -c || stat -f` chains among them. ADR-0012 enumerates all seven so the follow-up cleanup does not break the BSD path.
|
|
80
|
+
|
|
81
|
+
- `pipeline/scripts/gate-linux.sh` and the `npm run gate:linux` script. Every workflow now runs on a macOS runner; `test.yml` loses its Ubuntu leg and its `windows-test` job, since with `os: ["darwin"]` enforced those jobs could no longer `npm ci` this tree at all.
|
|
82
|
+
|
|
83
|
+
### Added
|
|
84
|
+
|
|
85
|
+
- **`doctor` check `host-platform`**, which runs first and blocks. First, so an unsupported host reads one honest line instead of a cascade whose common cause is never stated. The blocking set grows from five to six and is documented in `features/doctor.md`.
|
|
86
|
+
|
|
87
|
+
- **`smoke-macos-only.sh`** asserts each declaration separately, because declarations rot silently: the `os` field, both entry gates, the uninstall exemption, the doctor registration and its first position, that no user-facing doc promises another platform, and that no workflow targets a runner that cannot install this package.
|
|
88
|
+
|
|
89
|
+
### Fixed
|
|
90
|
+
|
|
91
|
+
- **`smoke-doctor.sh` computed the documented blocking set and never compared it to anything**, so the ref and the code were free to disagree - and they did, the moment `host-platform` was added. Worse, the computation used `sed -n '/a\|b/p'`, and BSD sed reads `\|` as a literal pipe, so the value had been empty on every macOS run since it was written. Both halves are fixed and the comparison is now an assertion.
|
|
92
|
+
|
|
93
|
+
- **The step-verb mutation test mutated the first matching literal in `doctor.mjs`'s source**, which stopped being a step the probe run reaches as soon as a check that passes on macOS was registered above `install-present`: the engine was never asked, and the gate reported it had gone soft. The step to mutate is now read back from the probe run's own output, so it stays unpinned to any particular check while being guaranteed to execute.
|
|
94
|
+
|
|
17
95
|
## [16.31.1] - 2026-09-13
|
|
18
96
|
|
|
19
97
|
### Fixed
|
package/README.md
CHANGED
|
@@ -10,7 +10,7 @@
|
|
|
10
10
|
|
|
11
11
|
An 8-phase AI development pipeline for **Claude Code**, **Copilot CLI** and **Codex CLI**. Drives a Jira issue or GitHub URL to a merged PR in one command - analysis → plan → TDD → review → test → commit → PR - with multi-repo orchestration, a plan-approval gate, CLI-aware parallel review, and store-compliance checks. Component and Figma-to-code work is dispatched to the per-stack marketplace plugins (iOS/SwiftUI, Android/Compose) rather than bundled, so component skills live in one place.
|
|
12
12
|
|
|
13
|
-
Runs natively on Claude Code, Copilot CLI and Codex CLI. macOS
|
|
13
|
+
Runs natively on Claude Code, Copilot CLI and Codex CLI. macOS only. Zero runtime dependencies.
|
|
14
14
|
|
|
15
15
|
📐 **[Architecture diagrams](./docs/architecture.md)** - the 8-phase flow, operating modes, review/triage, Figma subphases, component layout. **[Ecosystem diagram](./docs/ecosystem.md)** - how this repo, the `multi-agent-plugins` marketplace and `multi-agent-toolkit-mcp` compose.
|
|
16
16
|
|
|
@@ -46,7 +46,7 @@ Run a task - the input type is auto-detected:
|
|
|
46
46
|
/multi-agent:issue # browse unassigned GitHub issues → pick
|
|
47
47
|
```
|
|
48
48
|
|
|
49
|
-
Every input runs the same short intake - **account → (repo) → maturity check → dev-context** - then enters Phase 0. A Jira id or GitHub URL is fetched and maturity-checked
|
|
49
|
+
Every input runs the same short intake - **account → (repo) → maturity check → dev-context** - then enters Phase 0. A Jira id or GitHub URL is fetched and maturity-checked _before_ any code is written; free-text skips the fetch and goes straight to planning. Multi-repo tasks add extra repos at the dev-context step.
|
|
50
50
|
|
|
51
51
|
Add `autopilot` to skip confirmations, or `--local` to work on the current branch without a worktree (e.g. `/multi-agent:autopilot "PROJ-1234"`). Pipeline depth is a question the run asks, not a flag: `/multi-agent` and `/multi-agent:local` offer Full or Short at Phase 0.
|
|
52
52
|
|
|
@@ -75,15 +75,15 @@ The discipline behind all of this - bounded loops, evidence gates, token-budgete
|
|
|
75
75
|
|
|
76
76
|
## Modes
|
|
77
77
|
|
|
78
|
-
| Mode
|
|
79
|
-
|
|
80
|
-
| Full
|
|
81
|
-
| Autopilot | `/multi-agent:autopilot "task"`
|
|
82
|
-
| Local
|
|
83
|
-
| Depth
|
|
84
|
-
| Ship
|
|
85
|
-
| Audit
|
|
86
|
-
| Audit
|
|
78
|
+
| Mode | Command | Flow |
|
|
79
|
+
| --------- | ------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
80
|
+
| Full | `/multi-agent "task"` | All 8 phases, interactive |
|
|
81
|
+
| Autopilot | `/multi-agent:autopilot "task"` | 7 phases (interactive Test gate dropped), no confirmations |
|
|
82
|
+
| Local | `/multi-agent:local "task"` | Full pipeline minus the interactive Test gate, current branch (no worktree) |
|
|
83
|
+
| Depth | asked at Phase 0 Step 7.5 | Full (all phases) or Short (Dev → Review → Test → Commit → Report). Not a command name - `/multi-agent` and `:local` ask, both autopilot entries always run Full |
|
|
84
|
+
| Ship | `/multi-agent:resume-local` | Run the review→test→commit→report tail over local work |
|
|
85
|
+
| Audit | `/multi-agent:design-check` | Mock-mode vs Figma conformance, local-only |
|
|
86
|
+
| Audit | `/multi-agent:testflight-validation` | Pre-submission gates for a TestFlight build: static archive audit → Apple's `altool --validate-app` → Review-Guidelines check. Validates only, never uploads |
|
|
87
87
|
|
|
88
88
|
Depth, autopilot and `--local` are the only knobs on the run itself; everything else is its own command. The full catalog is below.
|
|
89
89
|
|
|
@@ -93,99 +93,99 @@ Depth, autopilot and `--local` are the only knobs on the run itself; everything
|
|
|
93
93
|
|
|
94
94
|
### Pipeline entries
|
|
95
95
|
|
|
96
|
-
| Command
|
|
97
|
-
|
|
98
|
-
| `/multi-agent "task"`
|
|
99
|
-
| `/multi-agent:local "task"`
|
|
100
|
-
| `/multi-agent:autopilot "task"`
|
|
101
|
-
| `/multi-agent:local-autopilot "task"` | Current branch, no confirmations, always Full
|
|
102
|
-
| `/multi-agent:resume-local`
|
|
96
|
+
| Command | What it does |
|
|
97
|
+
| ------------------------------------- | ---------------------------------------------------------------------------------------------------- |
|
|
98
|
+
| `/multi-agent "task"` | Full pipeline in a worktree. Asks Full or Short depth at Phase 0 |
|
|
99
|
+
| `/multi-agent:local "task"` | Same pipeline on the current branch, no worktree |
|
|
100
|
+
| `/multi-agent:autopilot "task"` | Worktree, no confirmations, always Full |
|
|
101
|
+
| `/multi-agent:local-autopilot "task"` | Current branch, no confirmations, always Full |
|
|
102
|
+
| `/multi-agent:resume-local` | Pipeline tail over work already done locally: Review → Build+Test → Commit/PR → Report. No dev phase |
|
|
103
103
|
|
|
104
104
|
### Task control
|
|
105
105
|
|
|
106
|
-
| Command
|
|
107
|
-
|
|
108
|
-
| `/multi-agent:status`
|
|
109
|
-
| `/multi-agent:log [#N]`
|
|
110
|
-
| `/multi-agent:resume [#N]`
|
|
111
|
-
| `/multi-agent:kill [#N]`
|
|
112
|
-
| `/multi-agent:steer #N "<instruction>"` | Correct a running task without stopping it; applied at the next phase boundary
|
|
113
|
-
| `/multi-agent:search`
|
|
114
|
-
| `/multi-agent:garbage-collect`
|
|
115
|
-
| `/multi-agent:prune-logs`
|
|
116
|
-
| `/multi-agent:purge`
|
|
106
|
+
| Command | What it does |
|
|
107
|
+
| --------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
108
|
+
| `/multi-agent:status` | Every task's ID, phase, branch and state |
|
|
109
|
+
| `/multi-agent:log [#N]` | Show a task's `agent-log.md` (most recent by default) |
|
|
110
|
+
| `/multi-agent:resume [#N]` | Carry a stopped or failed task on from its last phase |
|
|
111
|
+
| `/multi-agent:kill [#N]` | Stop a task, remove its worktree and branch |
|
|
112
|
+
| `/multi-agent:steer #N "<instruction>"` | Correct a running task without stopping it; applied at the next phase boundary |
|
|
113
|
+
| `/multi-agent:search` | Ranked search across every task log; `--semantic` queries the triage corpus |
|
|
114
|
+
| `/multi-agent:garbage-collect` | Sweep leftover scratch, orphan worktrees and offloaded payloads. `--abandoned` also reaps runs that stopped and were never picked back up. Dry-run first |
|
|
115
|
+
| `/multi-agent:prune-logs` | Delete per-task logs by age / project / task. Audit trail and metrics kept |
|
|
116
|
+
| `/multi-agent:purge` | Wipe every worktree, branch, log and state file. Double confirmation |
|
|
117
117
|
|
|
118
118
|
### Review
|
|
119
119
|
|
|
120
|
-
| Command
|
|
121
|
-
|
|
122
|
-
| `/multi-agent:review`
|
|
123
|
-
| `/multi-agent:review-jira`
|
|
124
|
-
| `/multi-agent:review-issue`
|
|
125
|
-
| `/multi-agent:review-analysis`
|
|
126
|
-
| `/multi-agent:diff-explain`
|
|
127
|
-
| `/multi-agent:refactor`
|
|
128
|
-
| `/multi-agent:scan`
|
|
129
|
-
| `/multi-agent:prune-prompts`
|
|
130
|
-
| `/multi-agent:ios-coding-standard` | Audit an iOS module against the 99-rule registry, produce a remediation plan
|
|
120
|
+
| Command | What it does |
|
|
121
|
+
| ---------------------------------- | ------------------------------------------------------------------------------------------- |
|
|
122
|
+
| `/multi-agent:review` | Parallel review of a branch diff or a PR; inline comments + approve/needs-work on PR input |
|
|
123
|
+
| `/multi-agent:review-jira` | Grade a Jira issue's readiness for development, comment the gaps |
|
|
124
|
+
| `/multi-agent:review-issue` | Same grading for a GitHub issue |
|
|
125
|
+
| `/multi-agent:review-analysis` | Review a written analysis document; findings cite the Locked rule they break |
|
|
126
|
+
| `/multi-agent:diff-explain` | Map a Phase 4 triage finding back to the diff lines that caused it |
|
|
127
|
+
| `/multi-agent:refactor` | Best-practice extraction + bug hunt + derived-skill drift + toolkit MCP research → one plan |
|
|
128
|
+
| `/multi-agent:scan` | Skill security scan of local skill directories against a tiered pattern catalog |
|
|
129
|
+
| `/multi-agent:prune-prompts` | Zero-base review of the always-on instruction footprint; keep / trial / delete per rule |
|
|
130
|
+
| `/multi-agent:ios-coding-standard` | Audit an iOS module against the 99-rule registry, produce a remediation plan |
|
|
131
131
|
|
|
132
132
|
### Analysis
|
|
133
133
|
|
|
134
|
-
| Command
|
|
135
|
-
|
|
136
|
-
| `/multi-agent:analysis`
|
|
137
|
-
| `/multi-agent:analysis-resolve`
|
|
134
|
+
| Command | What it does |
|
|
135
|
+
| --------------------------------- | ----------------------------------------------------------------------------------------- |
|
|
136
|
+
| `/multi-agent:analysis` | Standalone feature spec: global (23-section handoff) or corporate (IG/UC/FG) profile |
|
|
137
|
+
| `/multi-agent:analysis-resolve` | Answer an analysis doc's Section 20 open questions one row at a time |
|
|
138
138
|
| `/multi-agent:complaint-analysis` | Customer-complaint triage with Graylog evidence: client / bff root cause, or core routing |
|
|
139
139
|
|
|
140
140
|
### Testing on a device
|
|
141
141
|
|
|
142
|
-
| Command
|
|
143
|
-
|
|
144
|
-
| `/multi-agent:test`
|
|
145
|
-
| `/multi-agent:test-dark-mode`
|
|
146
|
-
| `/multi-agent:test-accessibility`
|
|
147
|
-
| `/multi-agent:test-dynamic-type`
|
|
148
|
-
| `/multi-agent:test-screenshots [locale]` | App Store screenshot set in a locale (defaults to `tr`)
|
|
149
|
-
| `/multi-agent:manual-test`
|
|
142
|
+
| Command | What it does |
|
|
143
|
+
| ---------------------------------------- | ------------------------------------------------------------------------- |
|
|
144
|
+
| `/multi-agent:test` | UI Bug Hunter on a booted simulator or emulator: screenshot, tap, analyze |
|
|
145
|
+
| `/multi-agent:test-dark-mode` | Walk every screen light then dark, report contrast and colour bugs |
|
|
146
|
+
| `/multi-agent:test-accessibility` | VoiceOver labels, sub-44pt tap targets, contrast, traits |
|
|
147
|
+
| `/multi-agent:test-dynamic-type` | Re-walk every screen at XL through accessibility-XL, report truncation |
|
|
148
|
+
| `/multi-agent:test-screenshots [locale]` | App Store screenshot set in a locale (defaults to `tr`) |
|
|
149
|
+
| `/multi-agent:manual-test` | Phase 5 standalone: check out the task branch and prepare it for Xcode |
|
|
150
150
|
|
|
151
151
|
### Design, build and store
|
|
152
152
|
|
|
153
|
-
| Command
|
|
154
|
-
|
|
155
|
-
| `/multi-agent:design-check`
|
|
156
|
-
| `/multi-agent:store-ready`
|
|
157
|
-
| `/multi-agent:testflight-validation` | iOS-pinned alias of `store-ready`. Validates only, never uploads
|
|
158
|
-
| `/multi-agent:build-optimize`
|
|
153
|
+
| Command | What it does |
|
|
154
|
+
| ------------------------------------ | ---------------------------------------------------------------------------------------- |
|
|
155
|
+
| `/multi-agent:design-check` | Mock-mode vs Figma audit with a coverage gate; annotated HTML + PDF report |
|
|
156
|
+
| `/multi-agent:store-ready` | Pre-submission gates for iOS and Android: package audit, store validation, policy review |
|
|
157
|
+
| `/multi-agent:testflight-validation` | iOS-pinned alias of `store-ready`. Validates only, never uploads |
|
|
158
|
+
| `/multi-agent:build-optimize` | Benchmark an Xcode build, run the analyzers, produce a recommend-first plan |
|
|
159
159
|
|
|
160
160
|
### Tickets and reporting
|
|
161
161
|
|
|
162
|
-
| Command
|
|
163
|
-
|
|
164
|
-
| `/multi-agent:jira`
|
|
165
|
-
| `/multi-agent:issue`
|
|
166
|
-
| `/multi-agent:create-jira` | Draft a Task / Bug / Story to the project's own conventions, preview before create
|
|
167
|
-
| `/multi-agent:channels`
|
|
168
|
-
| `/multi-agent:feedback`
|
|
162
|
+
| Command | What it does |
|
|
163
|
+
| -------------------------- | ----------------------------------------------------------------------------------- |
|
|
164
|
+
| `/multi-agent:jira` | Browse your open Jira issues → pick → branch → mode → launch |
|
|
165
|
+
| `/multi-agent:issue` | Browse unassigned GitHub issues → pick → auto-assign → launch |
|
|
166
|
+
| `/multi-agent:create-jira` | Draft a Task / Bug / Story to the project's own conventions, preview before create |
|
|
167
|
+
| `/multi-agent:channels` | Post the multi-channel report: Jira, Confluence, Wiki, PR description, board status |
|
|
168
|
+
| `/multi-agent:feedback` | Send one message to the maintainer. Only your text is sent, no logs or paths |
|
|
169
169
|
|
|
170
170
|
### Your own routines
|
|
171
171
|
|
|
172
|
-
| Command
|
|
173
|
-
|
|
174
|
-
| `/multi-agent:save [name]`
|
|
175
|
-
| `/multi-agent:routines`
|
|
176
|
-
| `/multi-agent:forget [name]` | Remove a saved routine and its registry entry
|
|
172
|
+
| Command | What it does |
|
|
173
|
+
| ---------------------------- | ---------------------------------------------------------------------------------- |
|
|
174
|
+
| `/multi-agent:save [name]` | Save a recurring job as a reusable `/multi-agent:<name>`. Local-only, never synced |
|
|
175
|
+
| `/multi-agent:routines` | List your saved routines and what each does |
|
|
176
|
+
| `/multi-agent:forget [name]` | Remove a saved routine and its registry entry |
|
|
177
177
|
|
|
178
178
|
### Setup and maintenance
|
|
179
179
|
|
|
180
|
-
| Command
|
|
181
|
-
|
|
182
|
-
| `/multi-agent:setup`
|
|
183
|
-
| `/multi-agent:stack [ids]`
|
|
184
|
-
| `/multi-agent:language [en\|tr]` | Show or set `outputLanguage`; `promptLanguage` stays English
|
|
185
|
-
| `/multi-agent:sync`
|
|
186
|
-
| `/multi-agent:update`
|
|
187
|
-
| `/multi-agent:uninstall`
|
|
188
|
-
| `/multi-agent:help`
|
|
180
|
+
| Command | What it does |
|
|
181
|
+
| -------------------------------- | ------------------------------------------------------------------------------ |
|
|
182
|
+
| `/multi-agent:setup` | First-run wizard: keychain token discovery, git identity, pipeline preparation |
|
|
183
|
+
| `/multi-agent:stack [ids]` | Enable the marketplace plugin(s) for this repo. Multi-select |
|
|
184
|
+
| `/multi-agent:language [en\|tr]` | Show or set `outputLanguage`; `promptLanguage` stays English |
|
|
185
|
+
| `/multi-agent:sync` | One-shot sync: Claude Code, Copilot CLI, pipeline repo, website, toolkit MCP |
|
|
186
|
+
| `/multi-agent:update` | Update to the latest published npm release and run migrations |
|
|
187
|
+
| `/multi-agent:uninstall` | Remove the pipeline from every CLI. Keychain tokens always left intact |
|
|
188
|
+
| `/multi-agent:help` | This catalog, in the terminal, in your `outputLanguage` |
|
|
189
189
|
|
|
190
190
|
### Subagents
|
|
191
191
|
|
|
@@ -209,11 +209,11 @@ This enables the matching plugin (+ the shared `ai-common` plugin) in the repo's
|
|
|
209
209
|
|
|
210
210
|
The pipeline runs natively on **Claude Code**, **Copilot CLI** and **Codex CLI** - all three install from the same `pipeline/` source and get the same 53 commands.
|
|
211
211
|
|
|
212
|
-
| Tool
|
|
213
|
-
|
|
212
|
+
| Tool | Flag | What it installs |
|
|
213
|
+
| ----------- | -------------------- | ------------------------------------------------------------------------------------------------------ |
|
|
214
214
|
| Claude Code | `--claude` (default) | slash commands + skills + agents + three `PreToolUse` hooks (secret scan, agent-guard, read-size gate) |
|
|
215
|
-
| Copilot CLI | `--copilot`
|
|
216
|
-
| Codex CLI
|
|
215
|
+
| Copilot CLI | `--copilot` | instructions + 53 sub-command skills + scripts |
|
|
216
|
+
| Codex CLI | `--codex` | one router skill + 53 specs as refs + 8 agent TOML + `AGENTS.md` block + `codex mcp add` |
|
|
217
217
|
|
|
218
218
|
Filter skills by stack with `--platform=ios\|android\|all`.
|
|
219
219
|
|
|
@@ -232,20 +232,20 @@ the one host whose panel spans two vendors, and the triage note says so.
|
|
|
232
232
|
|
|
233
233
|
## Tokens & integrations
|
|
234
234
|
|
|
235
|
-
`setup` scans your OS keychain and maps each token by a **logical name** (e.g. `jira`) to its real keychain entry - the pipeline resolves tokens through that mapping (`credential-store.sh`), so literal keychain names never appear in synced files. Tokens stay in the
|
|
236
|
-
|
|
237
|
-
| Token
|
|
238
|
-
|
|
239
|
-
| `jira`
|
|
240
|
-
| `github`
|
|
241
|
-
| `bitbucket`
|
|
242
|
-
| `confluence`
|
|
243
|
-
| `figma` + `figma_mcp` | fetch design context
|
|
244
|
-
| `fortify`
|
|
245
|
-
| `firebase`
|
|
246
|
-
| `jenkins`
|
|
247
|
-
| `npm`
|
|
248
|
-
| `appstore_connect_*`
|
|
235
|
+
`setup` scans your OS keychain and maps each token by a **logical name** (e.g. `jira`) to its real keychain entry - the pipeline resolves tokens through that mapping (`credential-store.sh`), so literal keychain names never appear in synced files. Tokens stay in the macOS Keychain, are **never committed or logged**, and are all **optional** - the pipeline asks for any it needs at Phase 0.
|
|
236
|
+
|
|
237
|
+
| Token | Used for | Phase |
|
|
238
|
+
| --------------------- | ---------------------------------------------------------------- | ----------------------- |
|
|
239
|
+
| `jira` | fetch the issue · post the report comment | 0, 7 |
|
|
240
|
+
| `github` | issues · PRs · `gh` auth | 0, 6 |
|
|
241
|
+
| `bitbucket` | PR create/update (reviewer-preserving) · diff | 6 |
|
|
242
|
+
| `confluence` | publish analysis / wiki pages | 7 |
|
|
243
|
+
| `figma` + `figma_mcp` | fetch design context | analysis only |
|
|
244
|
+
| `fortify` | security-scan findings gate | 4 |
|
|
245
|
+
| `firebase` | Firebase service-account JSON for Firebase projects | as needed |
|
|
246
|
+
| `jenkins` | CI trigger / status | build / deploy |
|
|
247
|
+
| `npm` | package publish (mostly CI) | release |
|
|
248
|
+
| `appstore_connect_*` | TestFlight / App Store pre-submission validation (optional, iOS) | `testflight-validation` |
|
|
249
249
|
|
|
250
250
|
The **secret scan** runs as a `PreToolUse` hook on Claude Code (hard-blocks a commit on a hit) and as a pre-push check elsewhere.
|
|
251
251
|
|
|
@@ -261,13 +261,13 @@ Uninstall preserves this layer: the tokens, the reader that opens them, the mapp
|
|
|
261
261
|
|
|
262
262
|
## Platform support
|
|
263
263
|
|
|
264
|
-
Runs on **macOS
|
|
264
|
+
Runs on **macOS** only. The package declares `os: ["darwin"]`, so `npm` refuses to install it elsewhere rather than letting a run fail halfway through: every credential read shells `security`, every iOS build `xcodebuild`, every piece of visual evidence `simctl`. Node.js 20.11+ (tested on 20 and 22). Reasoning: [ADR-0012](docs/adr/0012-macos-only.md).
|
|
265
265
|
|
|
266
266
|
## Companion repos
|
|
267
267
|
|
|
268
|
-
| Repo
|
|
269
|
-
|
|
270
|
-
| [`mmerterden/multi-agent-plugins`](https://github.com/mmerterden/multi-agent-plugins)
|
|
268
|
+
| Repo | What it is |
|
|
269
|
+
| --------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
|
270
|
+
| [`mmerterden/multi-agent-plugins`](https://github.com/mmerterden/multi-agent-plugins) | Marketplace of per-stack skill toolkits (iOS / Android / Frontend / Backend + common). `/multi-agent:stack` enables the matching plugin. |
|
|
271
271
|
| [`mmerterden/multi-agent-toolkit-mcp`](https://github.com/mmerterden/multi-agent-toolkit-mcp) | MCP server for UI testing / simulator capture / xcodebuild - powers the Phase 5 UI Bug Hunter. Published on the public npm registry as [`@mmerterden/multi-agent-toolkit-mcp`](https://www.npmjs.com/package/@mmerterden/multi-agent-toolkit-mcp); the installer registers it with each CLI for you, so `npx` resolves it with no extra configuration. |
|
|
272
272
|
|
|
273
273
|
## License
|