@mmerterden/multi-agent-pipeline 14.2.2 → 15.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (132) hide show
  1. package/CHANGELOG.md +186 -6
  2. package/README.md +19 -12
  3. package/README.tr.md +19 -12
  4. package/SECURITY.md +43 -0
  5. package/docs/FIGMA_PIPELINE.md +3 -3
  6. package/docs/adr/0006-skills-core-external-split.md +1 -1
  7. package/docs/adr/0007-multi-tool-adapter-framework.md +1 -1
  8. package/docs/adr/0009-claude-stack-skills-plugin-only.md +31 -0
  9. package/docs/adr/README.md +1 -0
  10. package/docs/architecture.md +13 -13
  11. package/docs/ecosystem.md +31 -31
  12. package/docs/features.md +5 -5
  13. package/index.js +6 -1
  14. package/install/_codex-agents.mjs +11 -2
  15. package/install/_common.mjs +109 -3
  16. package/install/_dev-only-files.mjs +0 -1
  17. package/install/_platform-filter.mjs +54 -113
  18. package/install/_plugin-skills.mjs +36 -36
  19. package/install/claude.mjs +251 -61
  20. package/install/codex.mjs +28 -6
  21. package/install/copilot.mjs +69 -9
  22. package/install/index.mjs +9 -3
  23. package/install/templates/codex-instructions.md +1 -1
  24. package/install/templates/copilot-instructions.md +3 -3
  25. package/package.json +2 -3
  26. package/pipeline/commands/multi-agent/SKILL.md +2 -0
  27. package/pipeline/commands/multi-agent/analysis/SKILL.md +3 -3
  28. package/pipeline/commands/multi-agent/analysis-resolve/SKILL.md +2 -2
  29. package/pipeline/commands/multi-agent/build-optimize/SKILL.md +9 -9
  30. package/pipeline/commands/multi-agent/channels/SKILL.md +1 -1
  31. package/pipeline/commands/multi-agent/complaint-analysis/SKILL.md +186 -0
  32. package/pipeline/commands/multi-agent/dev/SKILL.md +1 -1
  33. package/pipeline/commands/multi-agent/dev-autopilot/SKILL.md +1 -1
  34. package/pipeline/commands/multi-agent/dev-local/SKILL.md +1 -1
  35. package/pipeline/commands/multi-agent/dev-local-autopilot/SKILL.md +1 -1
  36. package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +1 -1
  37. package/pipeline/commands/multi-agent/help/SKILL.md +19 -4
  38. package/pipeline/commands/multi-agent/ios-coding-standard/SKILL.md +2 -2
  39. package/pipeline/commands/multi-agent/jira/SKILL.md +1 -1
  40. package/pipeline/commands/multi-agent/prune-prompts/SKILL.md +81 -0
  41. package/pipeline/commands/multi-agent/refactor/SKILL.md +36 -1
  42. package/pipeline/commands/multi-agent/resume/SKILL.md +1 -1
  43. package/pipeline/commands/multi-agent/{ship → resume-local}/SKILL.md +8 -8
  44. package/pipeline/commands/multi-agent/scan/SKILL.md +1 -1
  45. package/pipeline/commands/multi-agent/setup/SKILL.md +5 -5
  46. package/pipeline/commands/multi-agent/stack/SKILL.md +62 -40
  47. package/pipeline/commands/multi-agent/store-ready/SKILL.md +3 -3
  48. package/pipeline/commands/multi-agent/sync/SKILL.md +18 -11
  49. package/pipeline/commands/multi-agent/testflight-validation/SKILL.md +1 -1
  50. package/pipeline/commands/multi-agent/uninstall/SKILL.md +2 -0
  51. package/pipeline/commands/multi-agent/update/SKILL.md +4 -4
  52. package/pipeline/lib/issue-fetcher.sh +1 -1
  53. package/pipeline/lib/parse-complaints.sh +316 -0
  54. package/pipeline/multi-agent-refs/channels/wiki.md +3 -3
  55. package/pipeline/multi-agent-refs/complaint-analysis-template.md +99 -0
  56. package/pipeline/multi-agent-refs/component-dispatch.md +6 -6
  57. package/pipeline/multi-agent-refs/cross-cli-contract.md +16 -16
  58. package/pipeline/multi-agent-refs/features/external-context-injection.md +1 -1
  59. package/pipeline/multi-agent-refs/features/stack-skill-routing.md +5 -5
  60. package/pipeline/multi-agent-refs/generate-issue.md +1 -1
  61. package/pipeline/multi-agent-refs/phases/modes.md +1 -1
  62. package/pipeline/multi-agent-refs/phases/operations.md +7 -1
  63. package/pipeline/multi-agent-refs/phases/phase-0-init.md +1 -1
  64. package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +7 -7
  65. package/pipeline/multi-agent-refs/phases/phase-2-planning.md +5 -5
  66. package/pipeline/multi-agent-refs/phases/phase-3-dev.md +3 -3
  67. package/pipeline/multi-agent-refs/phases/phase-4-review.md +12 -12
  68. package/pipeline/multi-agent-refs/phases/phase-5-test.md +1 -1
  69. package/pipeline/multi-agent-refs/phases/phase-7-report.md +6 -0
  70. package/pipeline/multi-agent-refs/tracker-contract.md +3 -2
  71. package/pipeline/multi-agent-refs/wiki-capture.md +2 -2
  72. package/pipeline/preferences-template.json +18 -5
  73. package/pipeline/rules/figma-pipeline.md +2 -2
  74. package/pipeline/schemas/agent-state.schema.json +1 -1
  75. package/pipeline/schemas/complaint-analysis-spec.schema.json +216 -0
  76. package/pipeline/schemas/migrations/prefs-2.5.0-to-2.6.0.mjs +46 -0
  77. package/pipeline/schemas/prefs.schema.json +296 -66
  78. package/pipeline/schemas/token-budget.json +2 -2
  79. package/pipeline/scripts/README.md +4 -3
  80. package/pipeline/scripts/_stack-routing.mjs +79 -0
  81. package/pipeline/scripts/audit-log-rotate.sh +4 -1
  82. package/pipeline/scripts/build-skills-index.mjs +11 -0
  83. package/pipeline/scripts/build-stack-plugins.mjs +28 -60
  84. package/pipeline/scripts/check-derived-drift.mjs +55 -28
  85. package/pipeline/scripts/gc-worktrees.sh +4 -1
  86. package/pipeline/scripts/gen-skills-index.mjs +1 -1
  87. package/pipeline/scripts/match-skills.mjs +12 -2
  88. package/pipeline/scripts/migrate-prefs.mjs +33 -21
  89. package/pipeline/scripts/phase-tracker.sh +32 -5
  90. package/pipeline/scripts/phase0-exit-gate.mjs +3 -2
  91. package/pipeline/scripts/run-aggregator.mjs +7 -2
  92. package/pipeline/scripts/scan-agent-config.sh +1 -1
  93. package/pipeline/scripts/skill-conformance.mjs +165 -30
  94. package/pipeline/scripts/smoke-cross-cli-behavior.sh +1 -1
  95. package/pipeline/scripts/test-gap-rules/android.json +25 -0
  96. package/pipeline/scripts/test-gap-rules/ios.json +34 -0
  97. package/pipeline/scripts/test-gap-rules/node.json +29 -0
  98. package/pipeline/scripts/test-gap-rules/python.json +25 -0
  99. package/pipeline/scripts/uninstall.mjs +160 -11
  100. package/pipeline/scripts/usage-report.mjs +426 -0
  101. package/pipeline/scripts/validate-complaint-doc.mjs +250 -0
  102. package/pipeline/scripts/validate-reviewer.mjs +9 -3
  103. package/pipeline/skills/.skill-manifest.json +156 -108
  104. package/pipeline/skills/.skills-index.json +449 -12
  105. package/pipeline/skills/shared/README.md +14 -10
  106. package/pipeline/skills/shared/core/multi-agent-analysis-resolve/SKILL.md +1 -1
  107. package/pipeline/skills/shared/core/multi-agent-build-optimize/SKILL.md +1 -1
  108. package/pipeline/skills/shared/core/multi-agent-complaint-analysis/SKILL.md +49 -0
  109. package/pipeline/skills/shared/core/multi-agent-dev/SKILL.md +1 -1
  110. package/pipeline/skills/shared/core/multi-agent-dev-autopilot/SKILL.md +1 -1
  111. package/pipeline/skills/shared/core/multi-agent-dev-local/SKILL.md +1 -1
  112. package/pipeline/skills/shared/core/multi-agent-dev-local-autopilot/SKILL.md +1 -1
  113. package/pipeline/skills/shared/core/multi-agent-ios-coding-standard/SKILL.md +2 -2
  114. package/pipeline/skills/shared/core/multi-agent-prune-prompts/SKILL.md +83 -0
  115. package/pipeline/skills/shared/core/multi-agent-refactor/SKILL.md +153 -90
  116. package/pipeline/skills/shared/core/{multi-agent-ship → multi-agent-resume-local}/SKILL.md +6 -6
  117. package/pipeline/skills/shared/core/multi-agent-stack/SKILL.md +89 -22
  118. package/pipeline/skills/shared/core/multi-agent-store-ready/SKILL.md +1 -1
  119. package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +8 -8
  120. package/pipeline/skills/shared/core/multi-agent-testflight-validation/SKILL.md +1 -1
  121. package/pipeline/skills/shared/core/multi-agent-update/SKILL.md +1 -1
  122. package/pipeline/skills/shared/external/ios-coding-standard/modules/_TEMPLATE.yml +2 -2
  123. package/pipeline/skills/shared/external/ios-coding-standard/references/rules.yml +368 -33
  124. package/pipeline/skills/shared/external/ios-coding-standard/references/swiftlint.draft.yml +1 -2
  125. package/pipeline/skills/shared/external/ios-coding-standard/scripts/check_structure.py +765 -0
  126. package/pipeline/skills/shared/external/ios-module-structure/SKILL.md +75 -0
  127. package/pipeline/skills/shared/external/ios-module-structure/modules/_TEMPLATE.yml +131 -0
  128. package/pipeline/skills/shared/external/ios-module-structure/references/rules.yml +559 -0
  129. package/pipeline/skills/shared/external/ios-module-structure/scripts/check_structure.py +765 -0
  130. package/pipeline/skills/shared/external/localization-reuse-map/example-mapping.json +53 -10
  131. package/pipeline/skills/shared/external/localization-reuse-map/reference/sources-and-recipes.md +4 -3
  132. package/pipeline/skills/skills-index.md +7 -4
@@ -3,14 +3,14 @@ name: multi-agent-refactor
3
3
  language: en
4
4
  description: "Analyse the project: extract adapted best-practices, hunt real bugs + improvement areas, check upstream drift of derived skills, research the companion dev-toolkit MCP server against current MCP practice, draft one plan, take approval, develop, then ask whether to sync. Use when asked to review a project for bugs, gaps or improvements, or to check whether derived skills have drifted from upstream."
5
5
  user-invocable: true
6
- argument-hint: 'bugs | best-practices | drift | dev-toolkit | security | tests | performance | docs | ci | deps'
6
+ argument-hint: 'bugs | best-practices | drift | dev-toolkit | run-errors | security | tests | performance | docs | ci | deps'
7
7
  ---
8
8
 
9
9
  # Multi-Agent Refactor
10
10
 
11
11
  **One command. Best-practices + Bug hunt + Upstream drift + Dev-toolkit -> Score -> Plan -> Approval -> Develop -> Sync.**
12
12
 
13
- Deep-analyses the current project, extracts the global best-practices worth adopting (adapted to our stack), hunts real bugs and improvement areas, checks whether any skills we derived from an upstream source have drifted, researches the companion dev-toolkit MCP server against current MCP practice, scores everything, drafts a single prioritized plan, takes approval, applies the approved items, and asks about sync when done.
13
+ Deep-analyses the current project, extracts the global best-practices worth adopting (adapted to our stack), hunts real bugs and improvement areas, checks whether any skills we derived from an upstream source have drifted, researches the companion dev-toolkit MCP server against current MCP practice, scores everything, drafts a single prioritized plan, asks the user for approval, applies the approved items, and asks about sync at the end.
14
14
 
15
15
  **Input**: $ARGUMENTS (optional - area to focus on: "security", "performance", "tests", "bugs", "best-practices", "drift", "dev-toolkit", etc.)
16
16
 
@@ -20,14 +20,15 @@ Deep-analyses the current project, extracts the global best-practices worth adop
20
20
  Step 0: BEST-PRACTICES Research the field, extract the best approaches, ADAPT them to our stack -> plan band A
21
21
  Step 0b: DRIFT Check upstream-derived skills for updates we have not pulled -> plan band D
22
22
  Step 0c: DEV-TOOLKIT Research current MCP practice + audit the companion dev-toolkit repo -> plan band E
23
- Step 1: SCAN Scan the project structure (files, LOC, dependencies, CI, tests)
23
+ Step 0d: RUN-ERRORS Read the local run-error ledger, rank recurring failures -> plan band F
24
+ Step 1: SCAN Walk the project structure (files, LOC, dependencies, CI, tests)
24
25
  Step 2: ANALYZE 10 categories + an explicit BUG HUNT (real defects, not just scores) -> plan bands B, C
25
- Step 3: SCORE Each category out of 10, total /100
26
- Step 4: PLAN Merge bands A (best-practice) + B (bugs) + C (improvements) + D (drift) + E (dev-toolkit)
27
- Step 5: ASK "Shall I start developing?" - take approval
28
- Step 6: IMPLEMENT Apply approved items in order (lint, test, commit)
29
- Step 7: VERIFY Confirm all tests + lint pass
30
- Step 8: ASK SYNC "Shall I run multi-agent-sync?" - take approval
26
+ Step 3: SCORE Each category /10, overall /100
27
+ Step 4: PLAN Merge bands A (best-practice) + B (bugs) + C (improvements) + D (drift) + E (dev-toolkit) + F (run-errors)
28
+ Step 5: ASK "Start the work?" - wait for user approval
29
+ Step 6: IMPLEMENT Apply approved items one by one (lint, test, commit)
30
+ Step 7: VERIFY Confirm every test + lint passes
31
+ Step 8: ASK SYNC "Run /multi-agent:sync?" - wait for user approval
31
32
  ```
32
33
 
33
34
  Plan bands (all five feed the single Step 4 table):
@@ -36,19 +37,24 @@ Plan bands (all five feed the single Step 4 table):
36
37
  - **C - Improvement**: quality/perf/DX gaps surfaced by the 10-category analysis.
37
38
  - **D - Drift**: upstream updates to skills we derived from an external source.
38
39
  - **E - Dev-toolkit**: current-practice gaps in the companion dev-toolkit MCP server (the pipeline's device and browser hands), applied in that repo.
40
+ - **F - Run-errors**: recurring real failures the pipeline actually hit across past runs, read from the local run-error ledger. These are lived evidence, not speculation - a failure that recurs across users or tasks is a prioritized improvement area.
39
41
 
40
42
  ## Step 0: BEST-PRACTICES - research the field, adapt to us
41
43
 
42
44
  Do not copy features blindly. The goal is the best *approaches*, reshaped so they are applicable to THIS project's stack, size, and constraints.
43
45
 
44
- 1. **Identify the domain** (read README + package/build manifests): kind of project, stack, users.
46
+ 1. **Identify the domain** of the current project (read README + package/build manifests): what kind of project is it (CLI, SDK, app, pipeline, library), what stack, who uses it.
45
47
  2. **Research the best-in-class** for that exact domain:
46
48
  - **GitHub**: search the domain terms + `stars:>50`; read the top repos' READMEs, architecture docs, CI configs, test layout.
47
- - **X / Twitter**: search the domain terms + "best practices" / "we switched to" / "lesson learned"; capture what practitioners recommend right now (WebFetch/WebSearch on `x.com` / `twitter.com` threads).
49
+ - **X / Twitter**: search the domain terms + "best practices" / "we switched to" / "lesson learned"; capture what practitioners and tool authors actually recommend right now (WebFetch/WebSearch on `x.com` / `twitter.com` threads).
48
50
  - **Reddit**: relevant subreddits for real-world pain points and adopted patterns.
49
- - **Web**: "<domain> best practices <current-year>", official style guides, platform guidance.
50
- 3. **For each candidate approach, record**: what it is + source, the concrete benefit, the **adaptation** (exactly how it looks in OUR repo: which file/dir/script and what changes), effort (L/M/H), impact (L/M/H/Critical).
51
- 4. **Discard** approaches that do not fit our stack/scale/constraints - and say WHY.
51
+ - **Web**: "<domain> best practices <current-year>", official style guides, and the platform's own guidance.
52
+ 3. **For each candidate approach, record**:
53
+ - what it is, who does it well (source link)
54
+ - the concrete benefit
55
+ - **adaptation**: exactly how it would look in OUR repo (which file/dir/script), and what has to change for it to fit - not a generic "add tests" but "add a flow-assertion helper in `test/helpers/`"
56
+ - effort (Low/Med/High) and impact (Low/Med/High/Critical)
57
+ 4. **Discard** approaches that do not fit our stack, our scale, or our constraints - and say WHY (so the user sees the filter working, not just the survivors).
52
58
 
53
59
  Output (plan band A):
54
60
  ```
@@ -79,21 +85,21 @@ The upstream mapping is **configuration, never hardcoded** (it can reference pri
79
85
  }
80
86
  ```
81
87
 
82
- **The plugin cache is a mirror, not the authority.** `~/.claude/plugins/cache/<marketplace>/<plugin>/<version>/` holds only what the last `claude marketplace update` fetched. Treating it as the current upstream version is how this step reported "up to date" while the derivation was four releases behind: the cache sat at 0.2.1 and upstream was at 0.4.1. Resolve in the order below, and never let the cache alone produce an "up to date" verdict.
88
+ **The plugin cache is a mirror, not the authority.** `~/.claude/plugins/cache/<marketplace>/<plugin>/<version>/` only holds whatever the last `claude marketplace update` fetched. Reading it as the current upstream version is how this step reported "up to date" while the derivation was four releases behind: the cache sat at 0.2.1 and upstream was at 0.4.1. Resolve in the order below and never let the cache alone produce an "up to date" verdict.
83
89
 
84
- **Which manifest carries the version** is `upstreamVersionSource`, default `marketplace.json` - that is what a marketplace consumer resolves. Some upstreams keep per-plugin `plugin.json` versions deliberately unused, so reading those records a version nobody ships. One entry here was recorded at 0.7.0 from `plugin.json` while the upstream marketplace said 0.6.0 and its CHANGELOG said plainly that `plugin.json` is not used.
90
+ **Which manifest carries the version** is `upstreamVersionSource`, default `marketplace.json`. That is what a marketplace consumer actually resolves, and some upstreams keep per-plugin `plugin.json` versions deliberately unused - reading those records a version nobody ships. An entry here was recorded at 0.7.0 from `plugin.json` while the upstream marketplace said 0.6.0 and its CHANGELOG stated in as many words that `plugin.json` is not used.
85
91
 
86
92
  Procedure:
87
- 1. If `global.derivedSkillSources` is missing or empty -> **skip** this step and report "no derived-skill sources configured". Never invent a source.
88
- 2. For each entry, resolve the current upstream version, stopping at the first source that answers:
89
- - **`upstreamLocalClone`** if set: read `<clone>/.claude-plugin/<upstreamVersionSource>`, and report when the clone is itself behind its remote so a stale working copy is not silently trusted.
90
- - **`upstreamRepoUrl`**: read the same manifest over the API (`gh api` / `WebFetch`). A private upstream can 404 for the active account even though it exists; that is unreachable, not "no drift".
91
- - **Plugin cache**: last resort. When it is the only source that answered, report **`unverified (cache only)`**, never "up to date", and add a plan item to configure `upstreamLocalClone`.
92
- - If nothing is reachable, record "upstream unreachable" and move on (do not fail the run).
93
+ 1. If `global.derivedSkillSources` is missing or empty -> **skip** this step and report "no derived-skill sources configured" (nothing to check). Never invent a source.
94
+ 2. For each entry, resolve the current upstream version, in this order, stopping at the first that answers:
95
+ - **`upstreamLocalClone`** if set: read `<clone>/.claude-plugin/<upstreamVersionSource>`. Also run `git -C <clone> fetch --dry-run` (or compare against `@{u}`) and say so when the clone is itself behind, so a stale working copy is not silently trusted either.
96
+ - **`upstreamRepoUrl`**: read the same manifest over the API (`gh api` / `WebFetch`). A private upstream can 404 for the currently active account even when the repo exists - that is an unreachable result, not a "no drift" result.
97
+ - **Plugin cache** under `~/.claude/plugins/cache/<upstreamMarketplace>/<upstreamPlugin>/*/`: last resort only. When the cache is the only source that answered, report the entry as **`unverified (cache only)`**, never as "up to date", and add a plan item to configure `upstreamLocalClone`.
98
+ - If nothing is reachable, record the entry as "upstream unreachable" and move on (do not fail the whole run).
93
99
  3. Compare the resolved upstream version to `derivedFromVersion`:
94
100
  - equal, from an authoritative source -> "up to date" (no drift)
95
101
  - equal, from the cache only -> "unverified (cache only)"
96
- - newer -> **drift**: read the CHANGELOG entries between the two versions, and diff each `upstreamSkills` SKILL.md (+ templates) against our `localPath` copy. Summarize what changed, ignoring entries that touch only skills outside `upstreamSkills`.
102
+ - newer -> **drift**: read the CHANGELOG entries between the two versions, and diff each `upstreamSkills` SKILL.md (+ any templates) against our `localPath` copy. Summarize what changed upstream (bug fixes, new sections, new templates, renamed inputs). Ignore changelog entries that only touch skills outside `upstreamSkills` - they are not ours to port.
97
103
  4. Emit the drift table (plan band D):
98
104
 
99
105
  ```
@@ -102,40 +108,63 @@ Procedure:
102
108
  | <label> | <path> | 0.2.1 | 0.3.0 | YES | <changelog + diff summary> |
103
109
  ```
104
110
 
105
- 5. For each drifted entry, add a band-D plan item: "port upstream <plugin> <version> changes into <localPath>". Do not auto-apply - it goes through Step 5 approval; after porting, bump the entry's `derivedFromVersion`.
111
+ 5. For each drifted entry, add a band-D plan item: "port upstream <plugin> <version> changes into <localPath>", with the specific skills to update. Do not auto-apply upstream changes - they go through Step 5 approval like everything else, and after porting, bump the entry's `derivedFromVersion`.
106
112
 
107
113
  ## Step 0c: DEV-TOOLKIT - current MCP practice for the companion toolkit
108
114
 
109
- The pipeline's device and browser hands are MCP tools served by a companion repo (`dev-toolkit-mcp`), and several pipeline skills declare a minimum toolkit version. This step researches current MCP practice and audits that repo against it.
115
+ The pipeline's hands on devices and browsers are MCP tools served by a companion repo (`dev-toolkit-mcp`): Phase 5 test, `manual-test`, `design-check` and `apple-archive-compliance` all call them, and several pipeline skills declare a minimum toolkit version (see `cross-cli-contract.md`). That repo therefore has to track the MCP field, not just its own README. This step researches what current practice is and audits the toolkit against it.
110
116
 
111
117
  **Resolution** - configuration first, never a hardcoded path:
112
118
 
113
- 1. `prefs.global.devToolkit`: `{ enabled, label, localPath, mcpServerName, packageName, registry, repoUrl }`.
114
- 2. If unset, auto-detect from the MCP registration: `mcpServers` in `~/.claude.json` (plus `projects[*].mcpServers`) and `~/.claude/settings.json`; for a stdio `node` entry take `dirname(args[0])`, and accept it only if that directory is a git repo whose `package.json` depends on `@modelcontextprotocol/sdk`.
115
- 3. Neither resolves (or `enabled: false`) -> skip and report "no dev-toolkit configured". Never guess a path, never clone.
119
+ 1. `prefs.global.devToolkit` in `~/.claude/multi-agent-preferences.json`:
120
+
121
+ ```jsonc
122
+ {
123
+ "enabled": true,
124
+ "label": "<human name>",
125
+ "localPath": "$HOME/<repo-dir>", // the companion repo working copy
126
+ "mcpServerName": "<registered MCP server name>",
127
+ "packageName": "@<scope>/<package>",
128
+ "registry": "github-packages", // github-packages | npmjs | none
129
+ "repoUrl": "https://github.com/<owner>/<repo>"
130
+ }
131
+ ```
132
+
133
+ 2. If unset, auto-detect from the MCP registration: read `mcpServers` in `~/.claude.json` (including each `projects[*].mcpServers`) and in `~/.claude/settings.json`; for a stdio entry whose command is `node`, take `dirname(args[0])`. Accept it only when that directory is a git repo whose `package.json` depends on `@modelcontextprotocol/sdk`.
134
+ 3. If neither resolves, skip this step and report "no dev-toolkit configured". Never guess a path, never clone.
135
+ 4. `enabled: false` skips the step.
116
136
 
117
137
  **Research axes** - a finding without a source link is not a finding:
118
138
 
119
139
  | # | Axis | Where to look | What to extract |
120
140
  |---|------|---------------|-----------------|
121
- | 1 | MCP protocol | spec revisions + `@modelcontextprotocol/sdk` releases | features released since the pinned SDK that the server does not use: tool annotations, `outputSchema` + structured content, resource links, progress + cancellation, `tools/list_changed`, pagination |
122
- | 2 | Host clients | Claude Code / Copilot CLI / Cursor / Antigravity docs | description budget, tool-count ceilings, naming, output size limits, permission ergonomics |
123
- | 3 | Peer servers | GitHub search on the same domain + `stars:>50` | surfaces we lack, conventions peers converged on, what to discard |
124
- | 4 | Wrapped tooling | `simctl`, `idb`, `adb`, `xcodebuild`, Playwright notes, Apple ITMS + review guidelines | deprecated flags in use, new capabilities worth a tool, audit rules that changed |
125
- | 5 | Field practice | X / Twitter, Reddit, MCP community | what server authors changed recently (transport, output-token diets, error shape) |
141
+ | 1 | MCP protocol | spec revisions + `@modelcontextprotocol/sdk` releases | protocol features released since the pinned SDK range that the server does not use yet: tool annotations (`readOnlyHint` / `destructiveHint` / `idempotentHint` / `openWorldHint`), `outputSchema` + structured content, resource links in results, progress + cancellation, `tools/list_changed`, pagination, elicitation |
142
+ | 2 | Host clients | Claude Code / Copilot CLI / Cursor / Antigravity docs + release notes | per-tool description budget, tool-count ceilings, naming conventions, image and output size limits, permission / allowlist ergonomics |
143
+ | 3 | Peer servers | GitHub search on the same domain terms + `stars:>50` | tool surfaces we lack, conventions peers converged on, and what to discard as out of scope |
144
+ | 4 | Wrapped tooling | `xcrun simctl help`, `idb`, `adb`, `xcodebuild`, Playwright release notes, Apple ITMS + App Store Review Guidelines | deprecated flags still in use, new capabilities worth a tool, audit rules that changed |
145
+ | 5 | Field practice | X / Twitter, Reddit, MCP community threads | what server authors actually changed recently (transport choice, output-token diets, sandboxing, error shape) |
126
146
 
127
- **Audit the toolkit** - run the checks, do not assume:
147
+ **Audit the toolkit against the findings** - run the checks, do not assume:
128
148
 
129
149
  ```bash
130
150
  DT="<resolved localPath>"
131
- node --check "$DT/index.js"; find "$DT/tools" -name "*.js" -exec node --check {} \;
132
- grep -rn "console\.log(" "$DT/index.js" "$DT/tools" || echo "stdout clean" # stdout = JSON-RPC channel
133
- grep -nE "[0-9]+ tools" "$DT/README.md" "$DT/package.json" # advertised count vs reality
134
- node -p "require('$DT/package.json').files.join('\n')"; ls -d "$DT"/tools/*/ # files[] covers runtime dirs
151
+ node --check "$DT/index.js"
152
+ find "$DT/tools" -name "*.js" -type f -exec node --check {} \;
153
+
154
+ # stdout carries the JSON-RPC frames: a stray stdout write corrupts the stream
155
+ grep -rn "console\.log(" "$DT/index.js" "$DT/tools" || echo "stdout clean"
156
+
157
+ # advertised tool counts vs reality (README header + package.json description)
158
+ grep -nE "[0-9]+ tools" "$DT/README.md" "$DT/package.json"
159
+
160
+ # packaging: every runtime directory must be inside files[]
161
+ node -p "require('$DT/package.json').files.join('\n')"
162
+ ls -d "$DT"/tools/*/
163
+
135
164
  cd "$DT" && npm outdated; npm audit --omit=dev 2>/dev/null | tail -20
136
165
  ```
137
166
 
138
- Also check: every tool has a description + `inputSchema`; token-heavy results (screenshots, UI trees) are truncated or file-backed; failures return an error result, not a throw; `CHANGELOG.md`, CI and tests exist.
167
+ Also check: every tool carries a description and an `inputSchema`; token-heavy results (screenshots, UI trees, logs) are truncated or written to a file path instead of inlined; failures return an error result with an actionable message instead of throwing; `engines.node` matches what the SDK needs; `CHANGELOG.md`, a CI workflow and a test harness exist.
139
168
 
140
169
  Output (plan band E):
141
170
 
@@ -143,40 +172,74 @@ Output (plan band E):
143
172
  | # | Axis | Finding | Source | Adaptation in the toolkit (file) | Effort | Impact | In plan? |
144
173
  |---|------|---------|--------|----------------------------------|--------|--------|----------|
145
174
  | 1 | Protocol | read-only tools carry no annotations | <spec link> | add `annotations` to the read-only tools in index.js | Low | Medium | Yes (P1) |
146
- | 2 | Peer servers | peer exposes <surface> | <repo link> | does not fit: outside the pipeline's phases | - | - | No |
175
+ | 2 | Wrapped tooling | uses a simctl flag removed in Xcode <v> | <release notes> | switch tools/<family>/<file>.js to <new flag> | Low | High | Yes (P0) |
176
+ | 3 | Peer servers | peer exposes <surface> | <repo link> | does not fit: outside the pipeline's phases | - | - | No |
177
+ ```
178
+
179
+ Rules for this band:
180
+
181
+ - Band-E work lands in the toolkit repo, never mirrored into this one. Shipping it is `/multi-agent:sync` Step 3d.
182
+ - A finding that changes the tool surface (new / renamed / removed tool) pairs with a pipeline-side item: bump the minimum toolkit version wherever a pipeline skill declares one.
183
+ - If the current working directory IS the toolkit repo, skip band E and let bands A/B/C cover it - never report the same finding twice.
184
+
185
+ ## Step 0d: RUN-ERRORS - what the pipeline actually failed on
186
+
187
+ Before speculating about improvements, read what real runs already failed on. When usage logging is enabled, every terminated run appends its errors to a local ledger:
188
+
189
+ ```bash
190
+ LEDGER="$HOME/.claude/logs/multi-agent/errors-ledger.jsonl"
191
+ [ -f "$LEDGER" ] || echo "no run-error ledger yet - skip band F"
192
+ ```
193
+
194
+ Each line is one run: `{ t, id, u, c, rp, ph, v, errs[] }`. The `errs[]` entries are cause tags (`<phase>:<cause>` halt reasons, `phase-<id>-failed`, `run-failed`).
195
+
196
+ 1. Read the ledger (best-effort; a missing or unreadable file means skip band F, not a failure).
197
+ 2. Group by error tag. For each tag compute: occurrences, distinct users affected, distinct repos, the phase it usually strikes, first + last seen, and which pipeline versions it spans (a tag that persists across versions is unfixed; one that stopped at a version is already resolved - do not re-raise it).
198
+ 3. Rank by `occurrences x users-affected`. A failure that recurs across users or tasks is lived evidence, not a hypothesis - it outranks a speculative improvement.
199
+ 4. For each surviving tag, trace it to the phase doc / script that emits that cause and propose the concrete fix.
200
+
201
+ This band mirrors the admin dashboard's "Gelişim alanları" panel, but reads the local ledger so it needs no auth and works offline. If `$ARGUMENTS` names a focus area, still read the ledger - a recurring run error in that area is the strongest possible signal.
202
+
203
+ Output (plan band F):
204
+
205
+ ```
206
+ | # | Error tag | Occurrences | Users | Usual phase | Versions | Root cause (file) | Fix | In plan? |
207
+ |---|-----------|-------------|-------|-------------|----------|-------------------|-----|----------|
208
+ | 1 | 4:reviewer-json-invalid | 12 | 3 | 4 | 14.x-15.x | reviewer prompt lets prose leak | tighten schema instruction in phase-4-review.md | Yes (P0) |
209
+ | 2 | phase-3-failed | 5 | 2 | 3 | 15.0.x | build step misses a stack toolchain | add preflight in phase-3-dev.md | Yes (P1) |
147
210
  ```
148
211
 
149
212
  Rules for this band:
150
213
 
151
- - Band-E work lands in the toolkit repo, never mirrored into this one. Shipping it is `multi-agent-sync` Step 3d.
152
- - A finding that changes the tool surface pairs with a pipeline-side item: bump the minimum toolkit version wherever a pipeline skill declares one.
153
- - If the current working directory IS the toolkit repo, skip band E and let bands A/B/C cover it.
214
+ - The ledger is evidence of the past, not a spec. A tag that stopped recurring after a version bump is resolved - report it as resolved, do not add a plan item.
215
+ - Never quote a user's identity as blame. The `u` field is for counting distinct affected users, not for naming anyone in the plan.
216
+ - If the ledger is empty or absent, skip band F silently - it is additive signal, never a gate.
154
217
 
155
218
  ## Step 1: SCAN
156
219
 
157
220
  ```
158
- - File count, LOC (cloc or wc -l)
221
+ - file count, LOC (cloc or wc -l)
159
222
  - package.json / Package.swift / build.gradle analysis
160
- - Dependency count (runtime vs dev)
223
+ - dependency count (runtime vs dev)
161
224
  - Is there a CI/CD pipeline? (.github/workflows/, fastlane/, Makefile)
162
- - Is there a test setup? How many tests? Coverage?
163
- - Is there a linter/formatter configuration?
164
- - README, LICENSE, SECURITY.md, CHANGELOG.md present?
165
- - Git status: branch count, last commit date, tags
225
+ - Is there test infrastructure? How many tests? Coverage?
226
+ - Is a linter/formatter config in place?
227
+ - Do README, LICENSE, SECURITY.md, CHANGELOG.md exist?
228
+ - Git state: branch count, last commit date, tags
166
229
  ```
167
230
 
168
231
  ## Step 2: ANALYZE - 10 categories + BUG HUNT
169
232
 
170
233
  ### 2a. 10 categories (quality lens)
171
234
 
172
- | # | Category | What is checked |
173
- |---|----------|-----------------|
235
+ | # | Category | What to check |
236
+ |---|----------|---------------|
174
237
  | 1 | **Architecture** | Layer separation, modularity, SOLID, dependency direction |
175
- | 2 | **Code Quality** | Naming, magic number, dead code, complexity, DRY |
176
- | 3 | **Security** | Hardcoded secret, input validation, Keychain, ATS, ATT |
177
- | 4 | **Testing** | Coverage, test pyramid, edge case, mock/stub quality |
238
+ | 2 | **Code Quality** | Naming, magic numbers, dead code, complexity, DRY |
239
+ | 3 | **Security** | Hardcoded secrets, input validation, Keychain, ATS, ATT |
240
+ | 4 | **Tests** | Coverage, test pyramid, edge cases, mock/stub quality |
178
241
  | 5 | **CI/CD** | Build matrix, lint job, release automation, artifact caching |
179
- | 6 | **Documentation** | README quality, JSDoc/Swift doc, API docs, CHANGELOG |
242
+ | 6 | **Docs** | README quality, JSDoc / Swift docs, API docs, CHANGELOG |
180
243
  | 7 | **Performance** | Bundle size, lazy loading, memory leaks, unnecessary re-renders |
181
244
  | 8 | **Accessibility** | a11y labels, tap targets, VoiceOver, Dynamic Type |
182
245
  | 9 | **Dependency Management** | Outdated deps, vulnerability scan, lockfile, version pinning |
@@ -186,9 +249,9 @@ Rules for this band:
186
249
 
187
250
  Category scores measure quality; they do not find the bug that ships. Run an explicit defect hunt IN ADDITION to the scoring:
188
251
 
189
- - Dispatch focused review agents over the highest-risk surfaces: recently changed files, error/edge-path handling, concurrency, external input parsing, resource cleanup, off-by-one boundaries.
252
+ - Dispatch focused review agents (Agent tool, `code-reviewer` persona where available) over the highest-risk surfaces: recently changed files, error/edge-path handling, concurrency, external input parsing, resource cleanup, off-by-one boundaries.
190
253
  - For each candidate defect, record: **file:line**, the concrete failure scenario (inputs/state -> wrong output/crash), and severity.
191
- - **Verify before reporting**: only list a bug you can trace to a real failure path; drop the plausible-but-unprovable ones.
254
+ - **Verify before reporting**: only list a bug you can trace to a real failure path; drop the plausible-but-unprovable ones. Match the bar you would apply to a third party's code.
192
255
 
193
256
  Output (plan band B) - every confirmed defect becomes a P0/P1 item:
194
257
  ```
@@ -201,7 +264,7 @@ Improvement areas surfaced by 2a that are not outright bugs become band-C items.
201
264
 
202
265
  ## Step 3: SCORE
203
266
 
204
- Each category is scored out of 10. Output format:
267
+ Each category is scored out of 10. Output:
205
268
 
206
269
  ```
207
270
  +---------------------+-------+
@@ -210,9 +273,9 @@ Each category is scored out of 10. Output format:
210
273
  | Architecture | 9/10 |
211
274
  | Code Quality | 8/10 |
212
275
  | Security | 9/10 |
213
- | Testing | 7/10 |
276
+ | Tests | 7/10 |
214
277
  | CI/CD | 8/10 |
215
- | Documentation | 7/10 |
278
+ | Docs | 7/10 |
216
279
  | Performance | 9/10 |
217
280
  | Accessibility | 6/10 |
218
281
  | Dependency Mgmt | 8/10 |
@@ -224,13 +287,13 @@ Each category is scored out of 10. Output format:
224
287
 
225
288
  ## Step 4: PLAN - one merged, prioritized table
226
289
 
227
- Merge all five bands into a single plan. Tag each row with its band (A best-practice / B bug / C improvement / D drift / E dev-toolkit).
290
+ Merge all five bands into a single plan. Tag each row with its band (A best-practice / B bug / C improvement / D drift / E dev-toolkit) so the source is visible.
228
291
 
229
292
  ```
230
293
  | # | Priority | Band | Category | Item | Impact |
231
294
  |---|----------|------|----------|------|--------|
232
295
  | 1 | P0 | B | Security | Remove hardcoded API key (src/x:12) | Critical |
233
- | 2 | P0 | B | Testing | Fix nil-deref on empty response (src/y:42) | High |
296
+ | 2 | P0 | B | Tests | Fix nil-deref on empty response (src/y:42) | High |
234
297
  | 3 | P1 | A | CI/CD | Adopt matrix build (adapt: .github/workflows/ci.yml) | Medium |
235
298
  | 4 | P1 | D | Skills | Port upstream <plugin> 0.3.0 fixes into <localPath> | Medium |
236
299
  | 5 | P1 | E | Toolkit | Add read-only annotations to the dev-toolkit device tools | Medium |
@@ -238,39 +301,39 @@ Merge all five bands into a single plan. Tag each row with its band (A best-prac
238
301
  ```
239
302
 
240
303
  Priority levels:
241
- - **P0**: Bug, security hole, broken functionality - fix without fail
242
- - **P1**: Clear quality improvement or high-value adopted best-practice / drift port - should be done
243
- - **P2**: Nice-to-have - if there is time
304
+ - **P0**: bugs, security holes, broken functionality - must fix
305
+ - **P1**: clear quality improvement or high-value adopted best-practice / drift port - should be done
306
+ - **P2**: nice-to-have - if time allows
244
307
 
245
- ## Step 5: ASK - approval from the user
308
+ ## Step 5: ASK - user approval
246
309
 
247
310
  After showing the plan, ask:
248
311
 
249
- > "I found X items (P0: N, P1: M, P2: K) across bugs, best-practices, improvements, upstream drift, and dev-toolkit practice. Shall I start developing?"
312
+ > "Found X items (P0: N, P1: M, P2: K) across bugs, best-practices, improvements, upstream drift, and dev-toolkit practice. Should I start the work?"
250
313
 
251
314
  Options:
252
- - "Yes, do all of them"
253
- - "Do only the P0 ones"
254
- - "Do only P0 + P1"
255
- - "Only the bugs (band B)"
315
+ - "Yes, do all"
316
+ - "Only P0"
317
+ - "P0 + P1 only"
318
+ - "Only bugs (band B)"
256
319
  - "Only the dev-toolkit (band E)"
257
- - (the user can make a specific selection)
320
+ - (the user may pick specific items)
258
321
 
259
- **NEVER start developing without approval.**
322
+ **Never start the work without approval.**
260
323
 
261
324
  ## Step 6: IMPLEMENT
262
325
 
263
- Apply approved items in order:
326
+ Apply the approved items one by one:
264
327
 
265
328
  1. For each item:
266
- - Make the change (for band D, port the upstream diff into `localPath`, then bump that entry's `derivedFromVersion`)
267
- - Band-E items are applied inside the toolkit repo, never mirrored here: edit there, re-run its gates (`node --check`, `tools/list` handshake, advertised tool count matching reality), commit there. Publishing is `multi-agent-sync` Step 3d.
268
- - Run the relevant tests
269
- - If successful, move to the next
270
- - If it fails, revert and notify the user
329
+ - apply the change (for band D, port the upstream diff into `localPath`, then bump that entry's `derivedFromVersion` in preferences)
330
+ - band-E items are applied inside the toolkit repo, never mirrored here: edit there, re-run its gates (`node --check`, the `tools/list` handshake, advertised tool count matching reality), then commit there with that repo's own convention. Publishing is `/multi-agent:sync` Step 3d - do not publish from this step.
331
+ - run the relevant tests
332
+ - on success, move to the next
333
+ - on failure, roll back and notify the user
271
334
  2. After all changes are done:
272
- - Run the full lint + test suite
273
- - If successful, commit
335
+ - run the full lint + test suite
336
+ - on success, commit
274
337
 
275
338
  Commit format: `refactor(scope): {short description}`
276
339
 
@@ -288,22 +351,22 @@ echo "Lint: PASS/FAIL"
288
351
  echo "Test: PASS/FAIL (X/Y passed)"
289
352
  ```
290
353
 
291
- If any band-E item was applied, verify the toolkit repo too: syntax-check every file it loads, handshake the server and confirm `tools/list` still answers, and confirm the advertised tool counts match the count the server reports.
354
+ If any band-E item was applied, verify the toolkit repo too - syntax check every file it loads, handshake the server and confirm `tools/list` still answers, and confirm the advertised tool counts (README header, `package.json` description) match the count the server reports.
292
355
 
293
356
  ## Step 8: ASK SYNC
294
357
 
295
- After all approved items are applied, ask:
358
+ After every approved item is applied, ask:
296
359
 
297
- > "Development complete. Shall I run multi-agent-sync?"
360
+ > "Work is complete. Should I run /multi-agent:sync?"
298
361
 
299
362
  Options:
300
- - "Yes" -> run the `/sync` command (full ecosystem sync - its Step 3d ships any band-E work in the toolkit repo)
301
- - "No" -> report and finish
302
- - "Only commit + push" -> push only the current repo without sync
363
+ - "Yes" -> run the `/multi-agent:sync` command (full ecosystem sync - its Step 3d ships any band-E work in the toolkit repo)
364
+ - "No" -> emit a report and stop
365
+ - "Commit + push only" -> push the current repo without sync
303
366
 
304
367
  ## Focus filter
305
368
 
306
- If $ARGUMENTS is specified, focus on that band/category only:
369
+ If $ARGUMENTS is set, focus on that band/category only:
307
370
 
308
371
  | Input | Focus |
309
372
  |-------|-------|
@@ -313,8 +376,8 @@ If $ARGUMENTS is specified, focus on that band/category only:
313
376
  | `dev-toolkit` | Companion MCP toolkit research + audit only (band E) |
314
377
  | `security` | Security analysis only |
315
378
  | `tests` | Test coverage and quality only |
316
- | `performance` | Performance optimization only |
317
- | `docs` | Documentation only |
379
+ | `performance` | Performance optimisation only |
380
+ | `docs` | Docs only |
318
381
  | `ci` | CI/CD pipeline only |
319
382
  | `deps` | Dependency updates only |
320
383
  | (empty) | Everything (default) |
@@ -1,11 +1,11 @@
1
1
  ---
2
- name: multi-agent-ship
2
+ name: multi-agent-resume-local
3
3
  language: en
4
4
  description: "Continue already-done LOCAL work through the pipeline tail: Review → Build+Test → Commit/PR → Report (technical analysis + Jira test-scenario comment). No dev phase. Use when local work is already done and only review, build, commit and reporting remain."
5
5
  user-invocable: true
6
6
  ---
7
7
 
8
- # multi-agent ship - Take Existing Branch Work Through the Pipeline Tail
8
+ # multi-agent resume-local - Take Existing Branch Work Through the Pipeline Tail
9
9
 
10
10
  You already wrote (and maybe hand-tested) the change on the current branch, or committed it outside the pipeline entirely. `ship` picks up from there and runs the **pipeline tail** over that existing work in one command, without re-developing. (As of v14.0.0 the `--dev` family reviews its own output, so this is for work with no pipeline run behind it.)
11
11
 
@@ -34,10 +34,10 @@ Phases 1-3 (Analysis / Planning / Dev) are skipped by design - the branch's lo
34
34
  ## Input
35
35
 
36
36
  ```bash
37
- multi-agent ship # current branch vs base; Jira id from branch name
38
- multi-agent ship PROJ-12345 # explicit Jira id for the Phase 7 comment
39
- multi-agent ship --base develop # override base branch for the diff
40
- multi-agent ship autopilot # no gate prompts: auto-fix, auto-PR, auto-comment
37
+ multi-agent resume-local # current branch vs base; Jira id from branch name
38
+ multi-agent resume-local PROJ-12345 # explicit Jira id for the Phase 7 comment
39
+ multi-agent resume-local --base develop # override base branch for the diff
40
+ multi-agent resume-local autopilot # no gate prompts: auto-fix, auto-PR, auto-comment
41
41
  ```
42
42
 
43
43
  ## Notes
@@ -1,50 +1,117 @@
1
1
  ---
2
2
  name: multi-agent-stack
3
3
  language: en
4
- description: "Select the active stack for this repo by enabling the matching marketplace plugin(s) in .claude/settings.json (ios/android/mobile/backend/frontend/fullstack/all). Use when a repo's stack changed or the wrong plugins are enabled for it."
4
+ description: "Select the active stack(s) for this repo by enabling the matching marketplace plugin(s) in .claude/settings.json. Multi-select: pass several stacks (ios backend) or pick them in the native picker. Use when a repo's stack changed or the wrong plugins are enabled for it."
5
5
  user-invocable: true
6
6
  ---
7
7
 
8
- # multi-agent stack - Select Stack via Plugin Enablement
8
+ # multi-agent stack - Select Stack(s) via Plugin Enablement
9
9
 
10
- Stack skills ship as plugins in the `{owner}/multi-agent-plugins` marketplace. Selecting a stack = **enabling the matching plugin(s)** in the current repo's `.claude/settings.json` `enabledPlugins`. The `ai-common-engineering-toolkit` (accessibility audit, humanizer, Firebase) is always enabled alongside the stack plugin.
10
+ Stack skills ship as plugins in the `{owner}/multi-agent-plugins` marketplace. Selecting a stack = **enabling the matching plugin(s)** in the current repo's `.claude/settings.json` `enabledPlugins`. The `ai-common-toolkit` (accessibility audit, humanizer, Firebase) is always enabled alongside the stack plugin(s).
11
11
 
12
- This replaces the old `stack-swap.sh` mechanic that physically moved skill directories. No SessionStart hook, no directory shuffling - enablement is declarative, per-repo, and versioned.
12
+ On Claude Code the marketplace plugins are the ONLY source of stack skills - nothing is copied into `~/.claude/skills` anymore. A stack that is not enabled here is simply absent from the session. Copilot CLI and Codex CLI have no plugin loader; they receive a local copy filtered to the enabled stacks at install time, which is why step 5 below offers to refresh those copies after a change.
13
+
14
+ This replaces the old `stack-swap.sh` mechanic that physically moved skill directories in `~/.claude/skills/`. There is no SessionStart hook and no directory shuffling - enablement is declarative, per-repo, and versioned.
13
15
 
14
16
  ## Usage
15
17
 
16
18
  ```bash
17
- multi-agent-stack # show which plugins are enabled here
18
- multi-agent-stack ios # SwiftUI toolkit + common
19
- multi-agent-stack android # Compose toolkit + common
20
- multi-agent-stack mobile # iOS + Android + common
21
- multi-agent-stack backend # Python / Node spec-driven toolkit + common
22
- multi-agent-stack frontend # React / TSX toolkit + common
23
- multi-agent-stack fullstack # frontend + backend + common
24
- multi-agent-stack all # all four stack toolkits + common
19
+ /multi-agent:stack # no arg → native multi-select picker (current state pre-noted)
20
+ /multi-agent:stack ios # SwiftUI toolkit + common
21
+ /multi-agent:stack ios backend # any combination, space-separated
22
+ /multi-agent:stack android # Compose toolkit + common
23
+ /multi-agent:stack mobile # alias: ios + android + common
24
+ /multi-agent:stack backend # Python / Node spec-driven toolkit + common
25
+ /multi-agent:stack frontend # React / TSX toolkit + common (alias: web)
26
+ /multi-agent:stack fullstack # alias: frontend + backend + common
27
+ /multi-agent:stack all # all four stack toolkits + common
25
28
  ```
26
29
 
27
30
  ## Stack → plugin map
28
31
 
29
32
  | Stack | Plugins enabled (all `@multi-agent-plugins`) |
30
33
  |---|---|
31
- | `ios` | `ai-common-engineering-toolkit`, `ai-ios-engineering-toolkit` |
32
- | `android` | `ai-common-engineering-toolkit`, `ai-android-engineering-toolkit` |
33
- | `frontend` | `ai-common-engineering-toolkit`, `ai-frontend-engineering-toolkit` |
34
- | `backend` | `ai-common-engineering-toolkit`, `ai-backend-toolkit` |
35
- | `mobile` | common + `ai-ios-engineering-toolkit` + `ai-android-engineering-toolkit` |
36
- | `fullstack` | common + `ai-frontend-engineering-toolkit` + `ai-backend-toolkit` |
34
+ | `ios` | `ai-common-toolkit`, `ai-ios-toolkit` |
35
+ | `android` | `ai-common-toolkit`, `ai-android-toolkit` |
36
+ | `frontend` / `web` | `ai-common-toolkit`, `ai-frontend-toolkit` |
37
+ | `backend` | `ai-common-toolkit`, `ai-backend-toolkit` |
38
+ | `mobile` | common + `ai-ios-toolkit` + `ai-android-toolkit` |
39
+ | `fullstack` | common + `ai-frontend-toolkit` + `ai-backend-toolkit` |
37
40
  | `all` | common + all four stack toolkits |
38
41
 
42
+ Multiple args union their plugin sets: `ios backend` → common + iOS + backend.
43
+
39
44
  ## Behaviour
40
45
 
41
- 1. **No arg → status mode.** Read `.claude/settings.json` (repo) + the global settings and print which `@multi-agent-plugins` plugins are enabled. Modify nothing.
42
- 2. **Arg present → enable mode.** Ensure the marketplace is known (`claude marketplace add {owner}/multi-agent-plugins`), then write the matching `enabledPlugins` entries into the current repo's `.claude/settings.json`, setting the stack toolkits that don't belong to `false`. `ai-common-engineering-toolkit` is always `true`.
43
- 3. **Unknown arg** → show the list above; do not guess.
46
+ 1. **No arg → native multi-select picker.** Read `.claude/settings.json` (repo) + `~/.claude/settings.json` (global) to learn the current state, then ask with `AskUserQuestion` (`multiSelect: true`) - NEVER a numbered text menu:
47
+ - `question` (in `outputLanguage`): which stacks should be active in this repo, noting the currently enabled ones
48
+ - `header`: "Stacks" (English, UI contract)
49
+ - `options` (4): `ios` / `android` / `frontend (web)` / `backend`, each `description` (in `outputLanguage`) naming the plugin it enables and marking the ones already enabled with "(currently on)" / "(şu an açık)"
50
+ - Map each selected label back to its canonical arg before step 2: the token before any parenthesis (`frontend (web)` → `frontend`).
51
+ - Empty selection or cancel → **status mode**: print the currently enabled `@multi-agent-plugins` plugins and exit without modifying anything. Status mode never runs the Implementation block below - that block is enable-mode only.
52
+ 2. **Arg(s) present → enable mode.** Accept multiple space-separated stacks. Resolve each through the alias table (`web`→`frontend`, `mobile`→`ios android`, `fullstack`→`frontend backend`, `all`→every stack), union the plugin sets, then write the **current repo's** `.claude/settings.json`.
53
+ 3. **Write rules:**
54
+ - every plugin in the union is set to `true`
55
+ - every stack toolkit **not** in the union is set to `false` (leave non-`@multi-agent-plugins` entries untouched)
56
+ - `ai-common-toolkit@multi-agent-plugins` is always `true`
57
+ - **legacy-key cleanup**: delete any `ai-*-engineering-toolkit@multi-agent-plugins` keys - those plugin names were retired by the `ai-<stack>-toolkit` rename and a stale `true` there enables a plugin that no longer exists in the marketplace
58
+ 4. **Unknown arg** → show the table above; do not guess.
59
+ 5. **Copilot/Codex refresh offer.** Their local skill copies are filtered to the enabled stacks at install time, so after a change ask (single `AskUserQuestion`, not silent): "Refresh Copilot/Codex local copies now?" - Yes runs `node <pipelineRepo>/install.js --copilot --codex`, No leaves them stale with a one-line warning naming the command to run later. Skip this question entirely when neither `~/.copilot` nor `~/.codex` exists.
60
+
61
+ ## Implementation
62
+
63
+ ```bash
64
+ REPO_SETTINGS=".claude/settings.json"
65
+ MP="multi-agent-plugins"
66
+
67
+ # --- ensure the marketplace is known (idempotent) -------------------------
68
+ if ! claude marketplace list 2>/dev/null | grep -q "$MP"; then
69
+ claude marketplace add {owner}/multi-agent-plugins 2>/dev/null \
70
+ || echo "note: add the marketplace once with: claude marketplace add {owner}/multi-agent-plugins"
71
+ fi
72
+
73
+ COMMON="ai-common-toolkit@${MP}"
74
+ IOS="ai-ios-toolkit@${MP}"
75
+ ANDROID="ai-android-toolkit@${MP}"
76
+ FRONTEND="ai-frontend-toolkit@${MP}"
77
+ BACKEND="ai-backend-toolkit@${MP}"
78
+
79
+ # Enable-mode only: with zero args the write rules below would set every stack
80
+ # toolkit to false (nothing but $COMMON is in $ON), silently wiping the repo's
81
+ # selection. No args belongs to the picker / status path (Behaviour step 1).
82
+ if [ "$#" -eq 0 ]; then
83
+ echo "No stack given - nothing changed. Pass stacks (ios backend ...) or use the picker."
84
+ exit 0
85
+ fi
86
+
87
+ # resolve every arg through the alias table, union the ON set
88
+ ON="$COMMON"
89
+ for ARG in "$@"; do
90
+ case "$ARG" in
91
+ ios) ON="$ON $IOS" ;;
92
+ android) ON="$ON $ANDROID" ;;
93
+ frontend|web) ON="$ON $FRONTEND" ;;
94
+ "frontend (web)") ON="$ON $FRONTEND" ;;
95
+ backend) ON="$ON $BACKEND" ;;
96
+ mobile) ON="$ON $IOS $ANDROID" ;;
97
+ fullstack) ON="$ON $FRONTEND $BACKEND" ;;
98
+ all) ON="$ON $IOS $ANDROID $FRONTEND $BACKEND" ;;
99
+ *) echo "Unknown stack '$ARG'. One of: ios android mobile backend frontend web fullstack all"; exit 1 ;;
100
+ esac
101
+ done
102
+ ```
103
+
104
+ After resolving `$ON`, edit `.claude/settings.json` (create `{ "enabledPlugins": {} }` if absent) so that:
105
+ - every plugin in `$ON` is set to `true`,
106
+ - every stack toolkit **not** in `$ON` is set to `false` (leave non-`@multi-agent-plugins` entries untouched),
107
+ - `ai-common-toolkit@multi-agent-plugins` is always `true`,
108
+ - every `ai-*-engineering-toolkit@multi-agent-plugins` key is **deleted** (retired names).
109
+
110
+ Use the Read + Edit/Write tools (JSON must stay valid). Then print the resulting `enabledPlugins` block and run step 5 (Copilot/Codex refresh offer).
44
111
 
45
112
  ## Notes
46
113
 
47
114
  - Enablement is per-repo and declarative - commit `.claude/settings.json` so teammates get the same stack.
48
- - Restart the conversation to pick up newly enabled plugins.
115
+ - Restart the conversation (or reload the window) for Claude Code to pick up newly enabled plugins.
49
116
  - Pipeline Phase 1 stack detection is independent (it reads project files); `stack` only sets which plugin skill set is active.
50
117
  - The old `stack-swap.sh` skill-dir swap has been removed; stack selection is entirely plugin enablement.
@@ -37,7 +37,7 @@ doc is the contract.
37
37
  |---|---|---|
38
38
  | 1 Static | `ios_app_store_audit` 18 rules, needs an `.xcarchive` | `android_apk_audit` + `google-play-compliance` 21 rules, needs an `.aab` |
39
39
  | 2 Authoritative | `ios_testflight_validate` → `altool --validate-app`, needs credentials | `SKIPPED` - Play's authoritative check is server-side only and no client ships here |
40
- | 3 Policy | `app-store-review` skill vs repo source | `play-store-review` skill vs repo source |
40
+ | 3 Policy | `ai-ios-toolkit:app-store-review` skill vs repo source | `ai-android-toolkit:play-store-review` skill vs repo source |
41
41
 
42
42
  Gate 2's asymmetry is reported as an asymmetry. An Android run clears at most 2 of 3
43
43
  gates and must never print `passed`. A skipped gate is never folded into the pass