@mmerterden/multi-agent-pipeline 19.1.2 → 19.1.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (43) hide show
  1. package/CHANGELOG.md +80 -7
  2. package/README.md +42 -15
  3. package/README.tr.md +38 -12
  4. package/docs/ecosystem.md +3 -4
  5. package/docs/facts.json +4 -4
  6. package/install/_common.mjs +5 -5
  7. package/install/_mcp-register.mjs +1 -1
  8. package/install/_plugin-skills.mjs +3 -4
  9. package/install/copilot.mjs +5 -5
  10. package/manifest.json +45 -45
  11. package/package.json +1 -1
  12. package/pipeline/commands/multi-agent/setup/SKILL.md +4 -4
  13. package/pipeline/commands/sim-test.md +4 -4
  14. package/pipeline/lib/model-dispatch.sh +16 -10
  15. package/pipeline/multi-agent-refs/analysis/locked.md +2 -2
  16. package/pipeline/multi-agent-refs/channels/jira.md +8 -8
  17. package/pipeline/multi-agent-refs/features/base-branch-evidence.md +2 -2
  18. package/pipeline/multi-agent-refs/features/model-fallback.md +11 -7
  19. package/pipeline/multi-agent-refs/features/scope-check.md +1 -1
  20. package/pipeline/multi-agent-refs/features/stack-skill-routing.md +1 -1
  21. package/pipeline/multi-agent-refs/phases/operations.md +1 -1
  22. package/pipeline/multi-agent-refs/phases/phase-1-plan.md +3 -3
  23. package/pipeline/multi-agent-refs/rules.md +1 -1
  24. package/pipeline/schemas/route-config.schema.json +1 -1
  25. package/pipeline/scripts/autopilot-runner.mjs +6 -7
  26. package/pipeline/scripts/build-references.mjs +3 -3
  27. package/pipeline/scripts/build-stack-plugins.mjs +1 -1
  28. package/pipeline/scripts/bulk-read.sh +6 -4
  29. package/pipeline/scripts/doctor.mjs +4 -4
  30. package/pipeline/scripts/learnings-ledger.mjs +1 -1
  31. package/pipeline/scripts/match-skills.mjs +4 -4
  32. package/pipeline/scripts/memory-load.sh +3 -3
  33. package/pipeline/scripts/plan-coverage-gate.mjs +3 -3
  34. package/pipeline/scripts/run-aggregator.mjs +2 -2
  35. package/pipeline/scripts/runs-index.mjs +7 -7
  36. package/pipeline/scripts/scope-check-gate.mjs +1 -1
  37. package/pipeline/scripts/smoke-schema-validation.sh +9 -12
  38. package/pipeline/scripts/usage-report.mjs +5 -5
  39. package/pipeline/scripts/validate-analysis-doc.mjs +3 -3
  40. package/pipeline/scripts/write-state.mjs +22 -11
  41. package/pipeline/skills/.skill-manifest.json +3 -3
  42. package/pipeline/skills/shared/core/multi-agent/SKILL.md +2 -2
  43. package/pipeline/skills/shared/core/multi-agent-channels/SKILL.md +4 -5
package/CHANGELOG.md CHANGED
@@ -14,6 +14,79 @@ Internal file-layout changes that don't affect the slash-command surface are sti
14
14
 
15
15
  ---
16
16
 
17
+ ## [19.1.4] - 2026-09-21
18
+
19
+ ### Added
20
+
21
+ - **`smoke-gui-json-contract.sh` - the JSON a non-model consumer decodes stays
22
+ put.** `runs-index.mjs` and `update-check.sh` already had their output pinned;
23
+ `doctor.mjs --json` and `autopilot-status.sh --json` did not, and a renamed key
24
+ there breaks every consumer at once while the pipeline's own callers, which
25
+ read the human output, notice nothing. The gate drives both scripts and reads
26
+ their real stdout: `checks[]` with `{id, severity, step, problem, detail}` and
27
+ unique ids, `exit` an integer, severities from the documented set, the queue
28
+ triple as lists, the counters as numbers and `on` as a boolean. Adding a key
29
+ is deliberately not a failure - a decoder ignores what it does not know.
30
+ - **`smoke-no-confessional-prose.sh`** - shipped comments and docs describe the
31
+ code, not the project's own history. A reason a rule exists still ships; a
32
+ past-defect narration belongs in this file. Four allow-list entries are
33
+ anchored to the lines that carry them, so a stale exemption fails the gate.
34
+ - **A deterministic test for write-state's mid-write refusal.** The refusal
35
+ needs another writer to take the lock between acquiring it and renaming, which
36
+ on an idle machine is too narrow to hit and under load is a coin flip.
37
+ `WRITE_STATE_TEST_DELAY_MS` opens that window on purpose; the gate then asserts
38
+ both halves - exit 4, and nothing of that writer in the file.
39
+
40
+ ### Fixed
41
+
42
+ - **The concurrent-writer gate counted an honest exit 4 as an unexpected code.**
43
+ A writer whose lock is taken mid-write refuses and writes nothing, which is the
44
+ same class of outcome as an exit-2 lock timeout. It is now counted with its
45
+ invariant attached - a refusal that wrote nothing must leave nothing behind -
46
+ and both refusal counts are reported so a rising one stays visible.
47
+ - **`bulk-read.sh` claimed routing could send it outside the Anthropic ladder.**
48
+ It delegates to the `claude` CLI, which speaks that ladder and nothing else, so
49
+ the dispatcher's refusal of an external rung is the correct behaviour and the
50
+ comment was the wrong half. Both READMEs now state the limit: no call site
51
+ sends a model request itself, and a delegated read must not put a file's full
52
+ text outside the account that owns it.
53
+
54
+ ### Changed
55
+
56
+ - **Both READMEs gained a "which model answers" section**, listing
57
+ `/multi-agent:{route-on,route-off,route-status,model}` with the two lines that
58
+ matter: `scope` has no `host-session` value and that is a schema gate, and an
59
+ external rung is refused at dispatch.
60
+ - Confessional history removed from 57 shipped comments and doc paragraphs
61
+ across 40 files; each reason the rule exists was kept and rewritten with the
62
+ code as its subject.
63
+
64
+ ---
65
+
66
+ ## [19.1.3] - 2026-09-21
67
+
68
+ ### Fixed
69
+
70
+ - **Routing no longer hands back a rung the caller cannot dispatch to.** The
71
+ router permitted a non-Anthropic rung at `bulk-read`, and `bulk-read.sh`
72
+ passes the rung straight to `claude -p --model <rung>` - a CLI that speaks the
73
+ Anthropic ladder and nothing else. `research_ask` never consulted the router
74
+ at all. So no call site could honour an external rung, and the permission was
75
+ empty: a rule naming one produced a failed call rather than a cheaper one.
76
+
77
+ `model-dispatch.sh` now refuses a non-Anthropic rung at every call site and
78
+ falls to the NEXT rung the rule named, not to the caller's default - the user
79
+ asked for a preference order, and only the part that cannot be honoured is
80
+ dropped. `provider` in `cost-table.json` stays, because it is what makes the
81
+ refusal checkable and is the seam a future call site would use.
82
+
83
+ `smoke-model-dispatch.sh` asserts the reason rather than the rule:
84
+ `bulk-read.sh` still dispatches through the Anthropic CLI. The day a call site
85
+ grows its own provider path, that assertion is what says the refusal may be
86
+ lifted.
87
+
88
+ ---
89
+
17
90
  ## [19.1.2] - 2026-09-21
18
91
 
19
92
  ### Added
@@ -77,6 +150,7 @@ them.
77
150
  is refused for a subagent, because subagent dispatch belongs to the host - the
78
151
  script says so on stderr instead of substituting an Anthropic rung and leaving
79
152
  the user believing a rule worked that never could.
153
+
80
154
  - `cost-table.json` rungs declare a `provider`. Without it every rung looks
81
155
  alike and the subagent limit above cannot be checked at all.
82
156
  - `smoke-model-dispatch.sh` (18 assertions). Half of them drive the router; the
@@ -471,7 +545,7 @@ gate cannot be loaded.
471
545
  and producing no effect.
472
546
 
473
547
  **Interactive runs ask at the step instead of ending at it**: open the item and
474
- fix it, continue without it (recording in `state.maturity.accepted[]` *which*
548
+ fix it, continue without it (recording in `state.maturity.accepted[]` _which_
475
549
  gap was waved through, which is what separates an informed continue from a
476
550
  skipped check), or abort. `askInteractively: false` restores the old halt.
477
551
 
@@ -495,7 +569,7 @@ gate cannot be loaded.
495
569
 
496
570
  **The rule that shapes the rest: an edit is a reason to look again, never proof
497
571
  the gap closed.** A reply reading "will do later" moves the timestamp and fixes
498
- nothing, so a changed item is re-fetched and re-scored and the *check* decides.
572
+ nothing, so a changed item is re-fetched and re-scored and the _check_ decides.
499
573
  Only a genuinely DIFFERENT gap set earns a second comment - otherwise an
500
574
  unattended queue turns an item into a wall of identical bot text. "Cannot tell
501
575
  whether it moved" resolves to re-check, never to wait, because folding unknown
@@ -862,6 +936,7 @@ were the same shape: a chain fixed on one side and left broken on the other.
862
936
  than attaching nothing because a present artefact does not get re-checked.
863
937
  `fit` gained `webm`, which had been falling through to the pass-through arm and
864
938
  reaching the uploader at full size.
939
+
865
940
  - **Continuous mode's node half still named one host.** 17.2.0 fixed the shell
866
941
  side and closed the class. `autopilot-runner.mjs` was still resolving its three
867
942
  siblings from `~/.claude/scripts`, and `autopilot-intake.mjs` looked for
@@ -1209,7 +1284,6 @@ The flow video existed as a contract with no recorder, no UI test ever ran, and
1209
1284
 
1210
1285
  - **`install/templates/` now reaches the user tree.** It shipped in the tarball and was never copied, so `setup` told people to merge `install/templates/claude-hooks.json` - a path that exists only in a checkout. From an install the instruction named a file the reader did not have, and nothing said so. The doctor's `hook-coverage` check would have been a permanent `SKIP` for the same reason.
1211
1286
 
1212
-
1213
1287
  ### Fixed
1214
1288
 
1215
1289
  - **The widget came back empty of everything except the phase name.** The native tile renders exactly one string, its subject, so whatever a reader wants from the card has to travel in that subject. It carried `Phase 1 Analysis` and nothing else, while the tracker already held the model, the elapsed time, the tokens and the cost for that phase and the fallback card printed all four. `phase-tracker.sh subjects [id]` now renders the same values in the one shape the host accepts, `tiles` builds its `TaskCreate` calls from it, and every boundary hint re-reads it so the numbers advance instead of freezing at creation time. A phase with nothing to report is still just its name - no empty separators.
@@ -1262,7 +1336,6 @@ The flow video existed as a contract with no recorder, no UI test ever ran, and
1262
1336
 
1263
1337
  - **`council-view.mjs`** renders what each reviewer found and what triage did with it. Everything it needs was already on the state and nothing displayed it; `run-metrics.mjs` reduced it to a ratio and the rows behind the ratio were invisible. De-anonymization is safe here because it runs after triage has ruled, and the gate pins that the map reaches no prompt. `/multi-agent:log` renders it; exit 2 means the run never reached Phase 4.
1264
1338
 
1265
-
1266
1339
  - **A validator summary that says which checks ran.** `validate-analysis-doc.mjs --report` prints every check by name with a verdict each: `ok`, a finding count, or `skipped: <reason>`. The gap it closes is specific - the traceability matrix only runs in the corporate profile, so on a global document it never executed and the output was byte-identical to "ran, found nothing". Attribution is positional and reconciles by construction, so a check added without its mark misnames a finding but can never lose one; the gate asserts the printed set equals the set the file defines. Default output is unchanged for gate callers.
1267
1340
 
1268
1341
  - **Provenance in the reports a human reads.** `evidence_digest` and `base_commit` were already in the document front-matter, on the side only a machine reads. The analysis Phase 5 report, the `review-analysis` verdict and `--report` now carry them too: a timestamp cannot separate two reports made the same day, and the first question anyone asks of an older verdict is which version of the document it judged.
@@ -1288,6 +1361,7 @@ The flow video existed as a contract with no recorder, no UI test ever ran, and
1288
1361
  The contract, its reasoning and the check table live in `analysis/redesign.md`, which `analysis/SKILL.md` does not name and which therefore loads only on a redesign run.
1289
1362
 
1290
1363
  - `ANALYSIS_CEILING` 155000 -> 158500, the largest single raise it has taken. What could not be moved out is the argument: the section skeletons belong in `analysis-template.md`, where every other conditional section (15.6, 16.2) is already defined; Locked 37 belongs in `analysis/locked.md`, because a decision recorded only where it is implemented is not locked; and the intake question belongs in `analysis/intake.md`, because the option is chosen before anything knows the run is a redesign. 5.9 kB stayed outside the count. Everything was compressed twice before the number was picked.
1364
+
1291
1365
  ### Fixed
1292
1366
 
1293
1367
  - `learn-from-transcripts.mjs` called `process.exit(0)` on the line after writing its `--json` result, so a large mining result was cut at the 64 KB pipe buffer - the exact defect the gate added in 16.27.0 exists to prevent, in the file that motivated it. The gate caught it; the two sites it named are now a `main()` with a return.
@@ -1344,12 +1418,11 @@ The flow video existed as a contract with no recorder, no UI test ever ran, and
1344
1418
 
1345
1419
  - **Setup merged one hook event and dropped the rest.** The instruction said to deep-merge `hooks.PreToolUse`, so the capture hooks would never have installed - the fix would have shipped inert.
1346
1420
 
1347
-
1348
1421
  ## [16.26.0] - 2026-09-09
1349
1422
 
1350
1423
  ### Fixed
1351
1424
 
1352
- - **The credential store mangled every multi-line secret, and the failure surfaced three layers away.** `security` picks its output format from the *content* of the value: anything it deems non-printable - a newline included - comes back as bare hex with no marker, and the reader took that at face value. A Firebase service-account JSON went in and `7b0a2020...` came out, which is neither JSON nor base64, so the Crashlytics fetcher reported a token-exchange error for a credential that was stored perfectly. The write side had the mirror defect: `security -i` is line-oriented, so a newline inside `-w` terminated the command mid-value and the item was never written at all.
1425
+ - **The credential store mangled every multi-line secret, and the failure surfaced three layers away.** `security` picks its output format from the _content_ of the value: anything it deems non-printable - a newline included - comes back as bare hex with no marker, and the reader took that at face value. A Firebase service-account JSON went in and `7b0a2020...` came out, which is neither JSON nor base64, so the Crashlytics fetcher reported a token-exchange error for a credential that was stored perfectly. The write side had the mirror defect: `security -i` is line-oriented, so a newline inside `-w` terminated the command mid-value and the item was never written at all.
1353
1426
 
1354
1427
  `keychain.py` now reads with `-g` and decodes on the explicit `0x` marker rather than guessing from shape, and writes with `-X <hex>` so no value can be cut at a newline. Lookup also walks all four attribute conventions (`-a`+`-l`, `-a`+`-s`, `-l`, `-s`) instead of one.
1355
1428
 
@@ -1373,7 +1446,7 @@ The flow video existed as a contract with no recorder, no UI test ever ran, and
1373
1446
 
1374
1447
  - **`firebase-app-discovery.sh` - the repo already knew the appIds.** Every Crashlytics call needs the opaque `1:<n>:ios:<hex>`, and a console URL carries only the bundle id, so the fetcher spent a Management API round trip resolving it on every run. `GoogleService-Info*.plist` and `google-services.json` name both ids; the script reads every match (a repo with several targets has several plists, and they do not all point at one project), skips build outputs, and emits entries shaped for `prefs.global.firebase.accounts[].apps[]`. The fetcher prefers that when present and falls through to the API when absent, so an install that never ran discovery behaves exactly as before.
1375
1448
 
1376
- - **An encoding-regression gate.** `smoke-keychain.sh` stored a short ASCII word, which is exactly the value class that never reproduced the bug. It now round-trips a multi-line service-account JSON (and asserts it still *parses*, not merely compares equal), a hex-looking string, embedded quotes and backslashes, non-ASCII, and leading/trailing spaces.
1449
+ - **An encoding-regression gate.** `smoke-keychain.sh` stored a short ASCII word, which is exactly the value class that never reproduced the bug. It now round-trips a multi-line service-account JSON (and asserts it still _parses_, not merely compares equal), a hex-looking string, embedded quotes and backslashes, non-ASCII, and leading/trailing spaces.
1377
1450
 
1378
1451
  - **A gate for the shipped-but-dead script.** `firebase-app-discovery.sh` was written, installed onto every machine by all three hosts, and invoked by nothing: its only mention sat inside a comment in another script. That is the declared-but-inert defect in its purest form - the docs describe a capability, the install carries the code, and nothing connects them. `smoke-consumer-smoke-surface.sh` now checks every shipped script for an actual invoker, and is deliberately picky about what counts: comment lines, JSON schema descriptions, CHANGELOG and ROADMAP entries all name a script without ever running it, and each of those, left in the corpus, made the check pass on the broken state while it was being built.
1379
1452
 
package/README.md CHANGED
@@ -43,12 +43,12 @@ package. Check first, then pick the row that matches:
43
43
  node -v; npm -v; command -v node npm npx
44
44
  ```
45
45
 
46
- | What you see | What to do |
47
- |---|---|
48
- | nothing at all | Install Node >= 20.11: `brew install node`, or the LTS installer from nodejs.org |
49
- | `node` works, `npx` does not | `npm i -g @mmerterden/multi-agent-pipeline` then `multi-agent-pipeline install --claude` |
50
- | nvm is installed but the shell does not see it | `source ~/.nvm/nvm.sh && nvm use --lts`, or just open a new terminal |
51
- | npm is ancient (< 5.2, which predates npx) | `npm i -g npm@latest`, or use the global-install row above |
46
+ | What you see | What to do |
47
+ | ---------------------------------------------- | ---------------------------------------------------------------------------------------- |
48
+ | nothing at all | Install Node >= 20.11: `brew install node`, or the LTS installer from nodejs.org |
49
+ | `node` works, `npx` does not | `npm i -g @mmerterden/multi-agent-pipeline` then `multi-agent-pipeline install --claude` |
50
+ | nvm is installed but the shell does not see it | `source ~/.nvm/nvm.sh && nvm use --lts`, or just open a new terminal |
51
+ | npm is ancient (< 5.2, which predates npx) | `npm i -g npm@latest`, or use the global-install row above |
52
52
 
53
53
  And the path that needs neither `npx` nor a global install - clone and run the
54
54
  installer directly:
@@ -116,9 +116,10 @@ The other half of the change is that the count is finally guarded.
116
116
  `smoke-phase-contract.sh` derives it from `pipeline/schemas/phases.json` and
117
117
  holds every other copy to it: the generator's output, the token budget, the
118
118
  state-schema bounds, the progress fractions in sample output, and the named
119
- thresholds that used to be literals scattered across scripts. The phase count
120
- appeared in 91 places across 40 files with nothing checking any of them, while
121
- the command count, the jq count and the persona count all had gates.
119
+ thresholds, none of which is a literal in a script any more. A phase count is
120
+ otherwise the kind of number that spreads across dozens of files with nothing
121
+ checking any of them, the way the command count, the jq count and the persona
122
+ count are each held by a gate.
122
123
 
123
124
  Reasoning, mapping and rejected alternatives:
124
125
  [ADR-0014](./docs/adr/0014-six-phase-consolidation.md).
@@ -175,6 +176,7 @@ The discipline behind all of this - bounded loops, evidence gates, token-budgete
175
176
  | Audit | `/multi-agent:testflight-validation` | Pre-submission gates for a TestFlight build: static archive audit → Apple's `altool --validate-app` → Review-Guidelines check. Validates only, never uploads |
176
177
 
177
178
  Depth, autopilot and `--local` are the only knobs on the run itself; everything else is its own command. The full catalog is below.
179
+
178
180
  ### Pipeline depth: Full or Short
179
181
 
180
182
  Depth is the one question the run asks about its own shape. `/multi-agent` and `/multi-agent:local` ask it at Phase 0 Step 7.5 - after the issue is fetched and the task type is known, because that is what the recommendation is drawn from.
@@ -186,7 +188,6 @@ Short is right when you already know the fix and the file: a one-line guard, a c
186
188
 
187
189
  The widget follows the answer rather than predicting it: Phase 0 is the only tile drawn before you choose, and a Short run never draws an Analysis tile at all. Both autopilot entries skip the question and always run Full. `agent-state.json` records which one ran as `onlyDevelop`.
188
190
 
189
-
190
191
  ## Commands
191
192
 
192
193
  `/multi-agent` plus 56 sub-commands. `/multi-agent:help` renders the same catalog in your terminal, in your `outputLanguage`.
@@ -300,11 +301,37 @@ Two compliance skills install on every host and back the store gates: `apple-arc
300
301
  Everything above starts when you start it. Continuous mode is the same pipeline
301
302
  picking work up on its own, on ONE machine you choose, from repos you choose.
302
303
 
303
- | Command | What it does |
304
- | --- | --- |
305
- | `/multi-agent:autopilot-on` | Pick the repos this machine watches. Labelled GitHub issues and assigned + labelled Jira items then run in a worktree and stop at an open PR. Re-run to change the list |
306
- | `/multi-agent:autopilot-status` | What is running and at which phase, what is queued, what is waiting for an answer, the PRs of the last day, and the rolling spend |
307
- | `/multi-agent:autopilot-off` | Remove the schedule. Work already running finishes; the repo selection is kept |
304
+ | Command | What it does |
305
+ | ------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
306
+ | `/multi-agent:autopilot-on` | Pick the repos this machine watches. Labelled GitHub issues and assigned + labelled Jira items then run in a worktree and stop at an open PR. Re-run to change the list |
307
+ | `/multi-agent:autopilot-status` | What is running and at which phase, what is queued, what is waiting for an answer, the PRs of the last day, and the rolling spend |
308
+ | `/multi-agent:autopilot-off` | Remove the schedule. Work already running finishes; the repo selection is kept |
309
+
310
+ ## Which model answers
311
+
312
+ The ladder is `fable -> opus -> sonnet -> haiku`, and a run walks it on failure.
313
+ Routing lets a policy pick the rung instead, per call site, and is off until you
314
+ turn it on.
315
+
316
+ | Command | What it does |
317
+ | --------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------- |
318
+ | `/multi-agent:route-on` | Pick a strategy, a scope and the rules that say which rung a call lands on. Validated against `route-config.schema.json` before anything is written |
319
+ | `/multi-agent:route-off` | Disarm it. The rules are KEPT, so turning it back on does not re-ask for the same configuration |
320
+ | `/multi-agent:route-status` | Whether it is armed, which rule applies where, which rung the last dispatches took, and what this run has cost |
321
+ | `/multi-agent:model` | Turn the top rung on or off, and move the cost ledger's pricing with it in the same step |
322
+
323
+ **The scope cannot be the whole session.** `scope` accepts `subagent`,
324
+ `bulk-read` and `research`; there is no `host-session` value and that is a schema
325
+ gate, not a convention. Owning the host's base URL would send every call you make
326
+ through a third layer, including work that has nothing to do with this pipeline.
327
+
328
+ **The honest limit, which `route-status` prints rather than hides.** A rung on a
329
+ non-Anthropic provider is refused at dispatch. No call site sends a model
330
+ request itself: the one that routes, `bulk-read.sh`, delegates to the `claude`
331
+ CLI, which speaks that ladder and nothing else. Sending a read to a third-party
332
+ provider would also put the file's full text outside the account that owns it,
333
+ so it is a decision a user makes rather than a default. Routing chooses inside
334
+ the ladder.
308
335
 
309
336
  **Nothing is on by default and nothing is added implicitly.** Installing the
310
337
  package writes no state and schedules nothing; `smoke-autopilot-default-off.sh`
package/README.tr.md CHANGED
@@ -43,12 +43,12 @@ bak, sonra sana uyan satırı uygula:
43
43
  node -v; npm -v; command -v node npm npx
44
44
  ```
45
45
 
46
- | Gördüğün | Yapılacak |
47
- |---|---|
48
- | hiçbiri yok | Node >= 20.11 kur: `brew install node` ya da nodejs.org'dan LTS installer |
49
- | `node` çalışıyor, `npx` çalışmıyor | `npm i -g @mmerterden/multi-agent-pipeline` sonra `multi-agent-pipeline install --claude` |
50
- | nvm kurulu ama kabuk görmüyor | `source ~/.nvm/nvm.sh && nvm use --lts`, ya da yeni bir terminal aç |
51
- | npm çok eski (< 5.2, npx'ten önceki sürümler) | `npm i -g npm@latest`, ya da üstteki global kurulum satırı |
46
+ | Gördüğün | Yapılacak |
47
+ | --------------------------------------------- | ----------------------------------------------------------------------------------------- |
48
+ | hiçbiri yok | Node >= 20.11 kur: `brew install node` ya da nodejs.org'dan LTS installer |
49
+ | `node` çalışıyor, `npx` çalışmıyor | `npm i -g @mmerterden/multi-agent-pipeline` sonra `multi-agent-pipeline install --claude` |
50
+ | nvm kurulu ama kabuk görmüyor | `source ~/.nvm/nvm.sh && nvm use --lts`, ya da yeni bir terminal aç |
51
+ | npm çok eski (< 5.2, npx'ten önceki sürümler) | `npm i -g npm@latest`, ya da üstteki global kurulum satırı |
52
52
 
53
53
  Ne `npx` ne de global kurulum isteyen yol - klonla ve installer'ı doğrudan çalıştır:
54
54
 
@@ -156,6 +156,7 @@ Bunun arkasındaki disiplin - sınırlı loop'lar, kanıt kapıları, token-büt
156
156
  | Audit | `/multi-agent:testflight-validation` | TestFlight build için pre-submission kapıları: statik archive denetimi → Apple'ın `altool --validate-app`'i → Review-Guidelines kontrolü. Yalnızca doğrular, asla yüklemez |
157
157
 
158
158
  Koşunun kendisinde ayarlanabilen tek şey derinlik, autopilot ve `--local`; geri kalan her şey kendi komutu. Tam katalog aşağıda.
159
+
159
160
  ### Pipeline derinliği: Tam mı Kısa mı
160
161
 
161
162
  Derinlik, koşunun kendi şekli hakkında sorduğu tek soru. `/multi-agent` ve `/multi-agent:local` bunu Faz 0 Adım 7.5'te sorar - issue çekildikten ve görev tipi belirlendikten sonra, çünkü öneri onlardan çıkıyor.
@@ -167,7 +168,6 @@ Kısa'yı düzeltmeyi ve dosyayı zaten biliyorsan seç: tek satırlık bir guar
167
168
 
168
169
  Widget cevabı tahmin etmek yerine takip eder: sen seçmeden önce yalnızca Faz 0 karosu çizilir, Kısa koşu Analiz karosunu hiç çizmez. İki autopilot girişi de bu soruyu sormaz, her zaman Tam koşar. Hangisinin koştuğunu `agent-state.json` `onlyDevelop` alanında tutar.
169
170
 
170
-
171
171
  ## Komutlar
172
172
 
173
173
  `/multi-agent` ve 56 alt komut. `/multi-agent:help` aynı katalogu terminalde, `outputLanguage` ayarına göre gösterir.
@@ -282,11 +282,37 @@ Yukarıdaki her şey sen başlattığında başlar. Sürekli mod, aynı pipeline
282
282
  **senin seçtiğin tek makinede**, **senin seçtiğin repolardan** işi kendi başına
283
283
  alması.
284
284
 
285
- | Komut | Ne yapar |
286
- | --- | --- |
287
- | `/multi-agent:autopilot-on` | Bu makinenin izleyeceği repoları seçersin. Etiketli GitHub issue'ları ve sana atanmış + etiketli Jira maddeleri worktree'de koşar, açık PR'da durur. Listeyi değiştirmek için tekrar çalıştır |
288
- | `/multi-agent:autopilot-status` | Ne koşuyor hangi fazda, sırada ne var, ne cevap bekliyor, son bir günün PR'ları ve yuvarlanan harcama |
289
- | `/multi-agent:autopilot-off` | Zamanlamayı kaldırır. Koşan iş biter; repo seçimi saklanır |
285
+ | Komut | Ne yapar |
286
+ | ------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
287
+ | `/multi-agent:autopilot-on` | Bu makinenin izleyeceği repoları seçersin. Etiketli GitHub issue'ları ve sana atanmış + etiketli Jira maddeleri worktree'de koşar, açık PR'da durur. Listeyi değiştirmek için tekrar çalıştır |
288
+ | `/multi-agent:autopilot-status` | Ne koşuyor hangi fazda, sırada ne var, ne cevap bekliyor, son bir günün PR'ları ve yuvarlanan harcama |
289
+ | `/multi-agent:autopilot-off` | Zamanlamayı kaldırır. Koşan iş biter; repo seçimi saklanır |
290
+
291
+ ## Hangi model cevaplıyor
292
+
293
+ Merdiven `fable -> opus -> sonnet -> haiku`; koşu başarısızlıkta merdiveni iner.
294
+ Yönlendirme basamağı çağrı yeri başına bir politikaya seçtirir ve sen açana
295
+ kadar kapalıdır.
296
+
297
+ | Komut | Ne yapar |
298
+ | --------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- |
299
+ | `/multi-agent:route-on` | Strateji, kapsam ve hangi çağrının hangi basamağa düşeceğini söyleyen kuralları sorar. Hiçbir şey yazılmadan önce `route-config.schema.json`'a karşı doğrulanır |
300
+ | `/multi-agent:route-off` | Kapatır. Kurallar SAKLANIR, tekrar açıldığında aynı yapılandırma yeniden sorulmaz |
301
+ | `/multi-agent:route-status` | Açık mı, hangi kural nerede geçerli, son çağrılar hangi basamağa gitti, bu koşu ne tuttu |
302
+ | `/multi-agent:model` | Üst basamağı açıp kapatır ve maliyet defterinin fiyatlamasını aynı adımda onunla taşır |
303
+
304
+ **Kapsam oturumun tamamı olamaz.** `scope` yalnız `subagent`, `bulk-read` ve
305
+ `research` kabul eder; `host-session` diye bir değer yoktur ve bu bir teamül
306
+ değil, şema kapısıdır. Host'un base URL'ini sahiplenmek, bu pipeline'la ilgisi
307
+ olmayan işler dahil yaptığın her çağrıyı üçüncü bir katmandan geçirirdi.
308
+
309
+ **Dürüst sınır, ve `route-status` bunu gizlemek yerine yazar.** Anthropic
310
+ dışı bir sağlayıcıdaki basamak dağıtımda reddedilir. Hiçbir çağrı yeri model
311
+ isteğini kendisi atmıyor: yönlendirilen tek yer olan `bulk-read.sh` işi `claude`
312
+ CLI'ına devrediyor, o da yalnız bu merdiveni konuşuyor. Bir okumayı üçüncü bir
313
+ sağlayıcıya göndermek ayrıca dosyanın tam metnini onu sahiplenen hesabın dışına
314
+ çıkarırdı; bu yüzden varsayılan değil, kullanıcının verdiği bir karardır.
315
+ Yönlendirme merdivenin içinde seçim yapar.
290
316
 
291
317
  **Varsayılan olarak hiçbir şey açık değil ve hiçbir repo kendiliğinden eklenmez.**
292
318
  Paketi kurmak ne durum yazar ne zamanlama kurar; bu bir gün değişirse
package/docs/ecosystem.md CHANGED
@@ -212,10 +212,9 @@ that step's own contract).
212
212
  ### multi-agent-toolkit-mcp's tools, by category
213
213
 
214
214
  115 tools in 13 categories, counted from the server's own `tools/list` response
215
- at toolkit 3.13.0 rather than from a README. The table stood at 87 across 8
216
- categories for several releases because Code Intelligence and Wallet Passes
217
- shipped without anyone adding their rows, and nothing here was checked against
218
- the server - which is why the count now names its source.
215
+ at toolkit 3.13.1 rather than from a README. Three families are defined in
216
+ modules the main file only spreads in, so a count taken from the source reads
217
+ low; the response is the only place the whole surface appears at once.
219
218
 
220
219
  The families can be narrowed per host: `MCP_TOOLKIT_CAPS` serves only the named
221
220
  ones, and unset serves all 115. A frontend repo that sets `web,code,context`
package/docs/facts.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "$comment": "Generated by pipeline/scripts/gen-facts.mjs. Do not hand-edit: the site reads this, and a number edited here instead of at its source is the drift this file removes.",
3
- "generatedAt": "2026-09-20",
4
- "version": "19.1.2",
3
+ "generatedAt": "2026-09-21",
4
+ "version": "19.1.4",
5
5
  "phaseSchema": 2,
6
6
  "phases": [
7
7
  {
@@ -56,6 +56,6 @@
56
56
  "research": 2,
57
57
  "media": 1
58
58
  },
59
- "pluginSkillCount": 238,
60
- "toolkitVersion": "3.13.0"
59
+ "pluginSkillCount": 239,
60
+ "toolkitVersion": "3.13.1"
61
61
  }
@@ -485,11 +485,11 @@ export function pruneOrphanSkillFiles(skillsDir) {
485
485
  *
486
486
  * `.skills-index.json` is what `match-skills.mjs` reads when
487
487
  * `prefs.global.dynamicSkillLoading` is on; `skills-index.md` is the human-readable
488
- * twin. Both used to be copied ONLY under `--index-only`, while the preferences
489
- * schema stated that a full install always ships the index. It did not, so turning
490
- * the pref on after a normal install produced `cannot read index` and a silent
491
- * fall back to eager loading. One helper, called from every target, so the two
492
- * install modes cannot disagree about it again.
488
+ * twin. Both ship on a full install, not only under `--index-only`: the
489
+ * preferences schema states that a full install always carries the index, and an
490
+ * install that skips it makes turning the pref on afterwards produce `cannot read
491
+ * index` and a silent fall back to eager loading. One helper, called from every
492
+ * target, so the two install modes cannot disagree about it.
493
493
  *
494
494
  * @param {string} pipelineSrc - the repo's `pipeline/` directory
495
495
  * @param {string} skillsDest - the host's skills directory
@@ -152,7 +152,7 @@ function saysAlreadyExists(text) {
152
152
  * automatically a failure, and a zero exit is not automatically a fresh write - which
153
153
  * is why the "already exists" test runs on both paths. Reporting an existing server as
154
154
  * "skipped MCP registration (Command failed)" would send the user to fix something that
155
- * is already correct, and that is exactly what the first version of this did.
155
+ * is already correct.
156
156
  *
157
157
  * An existing entry under the same name is removed and re-added so its spec follows
158
158
  * the installed version (the entry is installer-owned; the reason is at the remove
@@ -4,9 +4,8 @@
4
4
  * Claude Code loads `{owner}/multi-agent-plugins` natively, so it gets all of a
5
5
  * stack plugin's skills. The other two hosts do not:
6
6
  *
7
- * - **Copilot CLI** has no plugin loader at all. The contract used to say it kept
8
- * standalone `figma-*` copies as a "frozen fallback", but `install/copilot.mjs`
9
- * pruned exactly those directories, so in practice Copilot had **zero**
7
+ * - **Copilot CLI** has no plugin loader at all. `install/copilot.mjs` prunes the
8
+ * standalone `figma-*` directories, so Copilot carries **zero**
10
9
  * plugin-authored skills - `create-screen`, `figma-validate`, `figma-review`,
11
10
  * `component`, `state` and `navigation` were all absent. A component task on
12
11
  * Copilot therefore had nothing to dispatch to.
@@ -221,7 +220,7 @@ export function installAuthoredPluginSkills(opts) {
221
220
  .map((e) => ({ name: e.name, from: join(groupDir, e.name) }));
222
221
 
223
222
  for (const { name, from } of entries) {
224
- // A plugin skill whose name a PIPELINE skill already owns used to be dropped.
223
+ // A plugin skill whose name a PIPELINE skill already owns is KEPT.
225
224
  // That is not a duplicate: the pipeline's `architecture` is a generic ADR
226
225
  // framework while the iOS plugin's is that stack's structural rules, and the
227
226
  // same holds for `backlog`. Claude Code reaches both because its loader
@@ -86,11 +86,11 @@ export function installCopilot(ctx) {
86
86
  /**
87
87
  * `$HOME/.claude/...` rewrites applied to skill files copied into the Copilot tree.
88
88
  *
89
- * Copilot's skills used to be copied byte-for-byte, so a Copilot-only install
90
- * shipped 15 references to `~/.claude/scripts` and `~/.claude/lib` - trees that
91
- * install never creates. Every one of those was a dangling path: the skill told the
92
- * agent to run a script at a location that does not exist on that machine. Codex has
93
- * had this rewrite since its installer was written; Copilot simply never got it.
89
+ * Copied byte-for-byte, a Copilot-only install would carry every `~/.claude/scripts`
90
+ * and `~/.claude/lib` reference in the skill text - trees that install never creates
91
+ * on that host. Each one is a dangling path: the skill tells the agent to run a
92
+ * script at a location that does not exist on that machine. Codex's installer
93
+ * applies the same rewrite for the same reason.
94
94
  *
95
95
  * Only the trees Copilot actually OWNS are rewritten:
96
96
  *