@mmerterden/multi-agent-pipeline 19.1.2 → 19.1.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +80 -7
- package/README.md +42 -15
- package/README.tr.md +38 -12
- package/docs/ecosystem.md +3 -4
- package/docs/facts.json +4 -4
- package/install/_common.mjs +5 -5
- package/install/_mcp-register.mjs +1 -1
- package/install/_plugin-skills.mjs +3 -4
- package/install/copilot.mjs +5 -5
- package/manifest.json +45 -45
- package/package.json +1 -1
- package/pipeline/commands/multi-agent/setup/SKILL.md +4 -4
- package/pipeline/commands/sim-test.md +4 -4
- package/pipeline/lib/model-dispatch.sh +16 -10
- package/pipeline/multi-agent-refs/analysis/locked.md +2 -2
- package/pipeline/multi-agent-refs/channels/jira.md +8 -8
- package/pipeline/multi-agent-refs/features/base-branch-evidence.md +2 -2
- package/pipeline/multi-agent-refs/features/model-fallback.md +11 -7
- package/pipeline/multi-agent-refs/features/scope-check.md +1 -1
- package/pipeline/multi-agent-refs/features/stack-skill-routing.md +1 -1
- package/pipeline/multi-agent-refs/phases/operations.md +1 -1
- package/pipeline/multi-agent-refs/phases/phase-1-plan.md +3 -3
- package/pipeline/multi-agent-refs/rules.md +1 -1
- package/pipeline/schemas/route-config.schema.json +1 -1
- package/pipeline/scripts/autopilot-runner.mjs +6 -7
- package/pipeline/scripts/build-references.mjs +3 -3
- package/pipeline/scripts/build-stack-plugins.mjs +1 -1
- package/pipeline/scripts/bulk-read.sh +6 -4
- package/pipeline/scripts/doctor.mjs +4 -4
- package/pipeline/scripts/learnings-ledger.mjs +1 -1
- package/pipeline/scripts/match-skills.mjs +4 -4
- package/pipeline/scripts/memory-load.sh +3 -3
- package/pipeline/scripts/plan-coverage-gate.mjs +3 -3
- package/pipeline/scripts/run-aggregator.mjs +2 -2
- package/pipeline/scripts/runs-index.mjs +7 -7
- package/pipeline/scripts/scope-check-gate.mjs +1 -1
- package/pipeline/scripts/smoke-schema-validation.sh +9 -12
- package/pipeline/scripts/usage-report.mjs +5 -5
- package/pipeline/scripts/validate-analysis-doc.mjs +3 -3
- package/pipeline/scripts/write-state.mjs +22 -11
- package/pipeline/skills/.skill-manifest.json +3 -3
- package/pipeline/skills/shared/core/multi-agent/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-channels/SKILL.md +4 -5
package/CHANGELOG.md
CHANGED
|
@@ -14,6 +14,79 @@ Internal file-layout changes that don't affect the slash-command surface are sti
|
|
|
14
14
|
|
|
15
15
|
---
|
|
16
16
|
|
|
17
|
+
## [19.1.4] - 2026-09-21
|
|
18
|
+
|
|
19
|
+
### Added
|
|
20
|
+
|
|
21
|
+
- **`smoke-gui-json-contract.sh` - the JSON a non-model consumer decodes stays
|
|
22
|
+
put.** `runs-index.mjs` and `update-check.sh` already had their output pinned;
|
|
23
|
+
`doctor.mjs --json` and `autopilot-status.sh --json` did not, and a renamed key
|
|
24
|
+
there breaks every consumer at once while the pipeline's own callers, which
|
|
25
|
+
read the human output, notice nothing. The gate drives both scripts and reads
|
|
26
|
+
their real stdout: `checks[]` with `{id, severity, step, problem, detail}` and
|
|
27
|
+
unique ids, `exit` an integer, severities from the documented set, the queue
|
|
28
|
+
triple as lists, the counters as numbers and `on` as a boolean. Adding a key
|
|
29
|
+
is deliberately not a failure - a decoder ignores what it does not know.
|
|
30
|
+
- **`smoke-no-confessional-prose.sh`** - shipped comments and docs describe the
|
|
31
|
+
code, not the project's own history. A reason a rule exists still ships; a
|
|
32
|
+
past-defect narration belongs in this file. Four allow-list entries are
|
|
33
|
+
anchored to the lines that carry them, so a stale exemption fails the gate.
|
|
34
|
+
- **A deterministic test for write-state's mid-write refusal.** The refusal
|
|
35
|
+
needs another writer to take the lock between acquiring it and renaming, which
|
|
36
|
+
on an idle machine is too narrow to hit and under load is a coin flip.
|
|
37
|
+
`WRITE_STATE_TEST_DELAY_MS` opens that window on purpose; the gate then asserts
|
|
38
|
+
both halves - exit 4, and nothing of that writer in the file.
|
|
39
|
+
|
|
40
|
+
### Fixed
|
|
41
|
+
|
|
42
|
+
- **The concurrent-writer gate counted an honest exit 4 as an unexpected code.**
|
|
43
|
+
A writer whose lock is taken mid-write refuses and writes nothing, which is the
|
|
44
|
+
same class of outcome as an exit-2 lock timeout. It is now counted with its
|
|
45
|
+
invariant attached - a refusal that wrote nothing must leave nothing behind -
|
|
46
|
+
and both refusal counts are reported so a rising one stays visible.
|
|
47
|
+
- **`bulk-read.sh` claimed routing could send it outside the Anthropic ladder.**
|
|
48
|
+
It delegates to the `claude` CLI, which speaks that ladder and nothing else, so
|
|
49
|
+
the dispatcher's refusal of an external rung is the correct behaviour and the
|
|
50
|
+
comment was the wrong half. Both READMEs now state the limit: no call site
|
|
51
|
+
sends a model request itself, and a delegated read must not put a file's full
|
|
52
|
+
text outside the account that owns it.
|
|
53
|
+
|
|
54
|
+
### Changed
|
|
55
|
+
|
|
56
|
+
- **Both READMEs gained a "which model answers" section**, listing
|
|
57
|
+
`/multi-agent:{route-on,route-off,route-status,model}` with the two lines that
|
|
58
|
+
matter: `scope` has no `host-session` value and that is a schema gate, and an
|
|
59
|
+
external rung is refused at dispatch.
|
|
60
|
+
- Confessional history removed from 57 shipped comments and doc paragraphs
|
|
61
|
+
across 40 files; each reason the rule exists was kept and rewritten with the
|
|
62
|
+
code as its subject.
|
|
63
|
+
|
|
64
|
+
---
|
|
65
|
+
|
|
66
|
+
## [19.1.3] - 2026-09-21
|
|
67
|
+
|
|
68
|
+
### Fixed
|
|
69
|
+
|
|
70
|
+
- **Routing no longer hands back a rung the caller cannot dispatch to.** The
|
|
71
|
+
router permitted a non-Anthropic rung at `bulk-read`, and `bulk-read.sh`
|
|
72
|
+
passes the rung straight to `claude -p --model <rung>` - a CLI that speaks the
|
|
73
|
+
Anthropic ladder and nothing else. `research_ask` never consulted the router
|
|
74
|
+
at all. So no call site could honour an external rung, and the permission was
|
|
75
|
+
empty: a rule naming one produced a failed call rather than a cheaper one.
|
|
76
|
+
|
|
77
|
+
`model-dispatch.sh` now refuses a non-Anthropic rung at every call site and
|
|
78
|
+
falls to the NEXT rung the rule named, not to the caller's default - the user
|
|
79
|
+
asked for a preference order, and only the part that cannot be honoured is
|
|
80
|
+
dropped. `provider` in `cost-table.json` stays, because it is what makes the
|
|
81
|
+
refusal checkable and is the seam a future call site would use.
|
|
82
|
+
|
|
83
|
+
`smoke-model-dispatch.sh` asserts the reason rather than the rule:
|
|
84
|
+
`bulk-read.sh` still dispatches through the Anthropic CLI. The day a call site
|
|
85
|
+
grows its own provider path, that assertion is what says the refusal may be
|
|
86
|
+
lifted.
|
|
87
|
+
|
|
88
|
+
---
|
|
89
|
+
|
|
17
90
|
## [19.1.2] - 2026-09-21
|
|
18
91
|
|
|
19
92
|
### Added
|
|
@@ -77,6 +150,7 @@ them.
|
|
|
77
150
|
is refused for a subagent, because subagent dispatch belongs to the host - the
|
|
78
151
|
script says so on stderr instead of substituting an Anthropic rung and leaving
|
|
79
152
|
the user believing a rule worked that never could.
|
|
153
|
+
|
|
80
154
|
- `cost-table.json` rungs declare a `provider`. Without it every rung looks
|
|
81
155
|
alike and the subagent limit above cannot be checked at all.
|
|
82
156
|
- `smoke-model-dispatch.sh` (18 assertions). Half of them drive the router; the
|
|
@@ -471,7 +545,7 @@ gate cannot be loaded.
|
|
|
471
545
|
and producing no effect.
|
|
472
546
|
|
|
473
547
|
**Interactive runs ask at the step instead of ending at it**: open the item and
|
|
474
|
-
fix it, continue without it (recording in `state.maturity.accepted[]`
|
|
548
|
+
fix it, continue without it (recording in `state.maturity.accepted[]` _which_
|
|
475
549
|
gap was waved through, which is what separates an informed continue from a
|
|
476
550
|
skipped check), or abort. `askInteractively: false` restores the old halt.
|
|
477
551
|
|
|
@@ -495,7 +569,7 @@ gate cannot be loaded.
|
|
|
495
569
|
|
|
496
570
|
**The rule that shapes the rest: an edit is a reason to look again, never proof
|
|
497
571
|
the gap closed.** A reply reading "will do later" moves the timestamp and fixes
|
|
498
|
-
nothing, so a changed item is re-fetched and re-scored and the
|
|
572
|
+
nothing, so a changed item is re-fetched and re-scored and the _check_ decides.
|
|
499
573
|
Only a genuinely DIFFERENT gap set earns a second comment - otherwise an
|
|
500
574
|
unattended queue turns an item into a wall of identical bot text. "Cannot tell
|
|
501
575
|
whether it moved" resolves to re-check, never to wait, because folding unknown
|
|
@@ -862,6 +936,7 @@ were the same shape: a chain fixed on one side and left broken on the other.
|
|
|
862
936
|
than attaching nothing because a present artefact does not get re-checked.
|
|
863
937
|
`fit` gained `webm`, which had been falling through to the pass-through arm and
|
|
864
938
|
reaching the uploader at full size.
|
|
939
|
+
|
|
865
940
|
- **Continuous mode's node half still named one host.** 17.2.0 fixed the shell
|
|
866
941
|
side and closed the class. `autopilot-runner.mjs` was still resolving its three
|
|
867
942
|
siblings from `~/.claude/scripts`, and `autopilot-intake.mjs` looked for
|
|
@@ -1209,7 +1284,6 @@ The flow video existed as a contract with no recorder, no UI test ever ran, and
|
|
|
1209
1284
|
|
|
1210
1285
|
- **`install/templates/` now reaches the user tree.** It shipped in the tarball and was never copied, so `setup` told people to merge `install/templates/claude-hooks.json` - a path that exists only in a checkout. From an install the instruction named a file the reader did not have, and nothing said so. The doctor's `hook-coverage` check would have been a permanent `SKIP` for the same reason.
|
|
1211
1286
|
|
|
1212
|
-
|
|
1213
1287
|
### Fixed
|
|
1214
1288
|
|
|
1215
1289
|
- **The widget came back empty of everything except the phase name.** The native tile renders exactly one string, its subject, so whatever a reader wants from the card has to travel in that subject. It carried `Phase 1 Analysis` and nothing else, while the tracker already held the model, the elapsed time, the tokens and the cost for that phase and the fallback card printed all four. `phase-tracker.sh subjects [id]` now renders the same values in the one shape the host accepts, `tiles` builds its `TaskCreate` calls from it, and every boundary hint re-reads it so the numbers advance instead of freezing at creation time. A phase with nothing to report is still just its name - no empty separators.
|
|
@@ -1262,7 +1336,6 @@ The flow video existed as a contract with no recorder, no UI test ever ran, and
|
|
|
1262
1336
|
|
|
1263
1337
|
- **`council-view.mjs`** renders what each reviewer found and what triage did with it. Everything it needs was already on the state and nothing displayed it; `run-metrics.mjs` reduced it to a ratio and the rows behind the ratio were invisible. De-anonymization is safe here because it runs after triage has ruled, and the gate pins that the map reaches no prompt. `/multi-agent:log` renders it; exit 2 means the run never reached Phase 4.
|
|
1264
1338
|
|
|
1265
|
-
|
|
1266
1339
|
- **A validator summary that says which checks ran.** `validate-analysis-doc.mjs --report` prints every check by name with a verdict each: `ok`, a finding count, or `skipped: <reason>`. The gap it closes is specific - the traceability matrix only runs in the corporate profile, so on a global document it never executed and the output was byte-identical to "ran, found nothing". Attribution is positional and reconciles by construction, so a check added without its mark misnames a finding but can never lose one; the gate asserts the printed set equals the set the file defines. Default output is unchanged for gate callers.
|
|
1267
1340
|
|
|
1268
1341
|
- **Provenance in the reports a human reads.** `evidence_digest` and `base_commit` were already in the document front-matter, on the side only a machine reads. The analysis Phase 5 report, the `review-analysis` verdict and `--report` now carry them too: a timestamp cannot separate two reports made the same day, and the first question anyone asks of an older verdict is which version of the document it judged.
|
|
@@ -1288,6 +1361,7 @@ The flow video existed as a contract with no recorder, no UI test ever ran, and
|
|
|
1288
1361
|
The contract, its reasoning and the check table live in `analysis/redesign.md`, which `analysis/SKILL.md` does not name and which therefore loads only on a redesign run.
|
|
1289
1362
|
|
|
1290
1363
|
- `ANALYSIS_CEILING` 155000 -> 158500, the largest single raise it has taken. What could not be moved out is the argument: the section skeletons belong in `analysis-template.md`, where every other conditional section (15.6, 16.2) is already defined; Locked 37 belongs in `analysis/locked.md`, because a decision recorded only where it is implemented is not locked; and the intake question belongs in `analysis/intake.md`, because the option is chosen before anything knows the run is a redesign. 5.9 kB stayed outside the count. Everything was compressed twice before the number was picked.
|
|
1364
|
+
|
|
1291
1365
|
### Fixed
|
|
1292
1366
|
|
|
1293
1367
|
- `learn-from-transcripts.mjs` called `process.exit(0)` on the line after writing its `--json` result, so a large mining result was cut at the 64 KB pipe buffer - the exact defect the gate added in 16.27.0 exists to prevent, in the file that motivated it. The gate caught it; the two sites it named are now a `main()` with a return.
|
|
@@ -1344,12 +1418,11 @@ The flow video existed as a contract with no recorder, no UI test ever ran, and
|
|
|
1344
1418
|
|
|
1345
1419
|
- **Setup merged one hook event and dropped the rest.** The instruction said to deep-merge `hooks.PreToolUse`, so the capture hooks would never have installed - the fix would have shipped inert.
|
|
1346
1420
|
|
|
1347
|
-
|
|
1348
1421
|
## [16.26.0] - 2026-09-09
|
|
1349
1422
|
|
|
1350
1423
|
### Fixed
|
|
1351
1424
|
|
|
1352
|
-
- **The credential store mangled every multi-line secret, and the failure surfaced three layers away.** `security` picks its output format from the
|
|
1425
|
+
- **The credential store mangled every multi-line secret, and the failure surfaced three layers away.** `security` picks its output format from the _content_ of the value: anything it deems non-printable - a newline included - comes back as bare hex with no marker, and the reader took that at face value. A Firebase service-account JSON went in and `7b0a2020...` came out, which is neither JSON nor base64, so the Crashlytics fetcher reported a token-exchange error for a credential that was stored perfectly. The write side had the mirror defect: `security -i` is line-oriented, so a newline inside `-w` terminated the command mid-value and the item was never written at all.
|
|
1353
1426
|
|
|
1354
1427
|
`keychain.py` now reads with `-g` and decodes on the explicit `0x` marker rather than guessing from shape, and writes with `-X <hex>` so no value can be cut at a newline. Lookup also walks all four attribute conventions (`-a`+`-l`, `-a`+`-s`, `-l`, `-s`) instead of one.
|
|
1355
1428
|
|
|
@@ -1373,7 +1446,7 @@ The flow video existed as a contract with no recorder, no UI test ever ran, and
|
|
|
1373
1446
|
|
|
1374
1447
|
- **`firebase-app-discovery.sh` - the repo already knew the appIds.** Every Crashlytics call needs the opaque `1:<n>:ios:<hex>`, and a console URL carries only the bundle id, so the fetcher spent a Management API round trip resolving it on every run. `GoogleService-Info*.plist` and `google-services.json` name both ids; the script reads every match (a repo with several targets has several plists, and they do not all point at one project), skips build outputs, and emits entries shaped for `prefs.global.firebase.accounts[].apps[]`. The fetcher prefers that when present and falls through to the API when absent, so an install that never ran discovery behaves exactly as before.
|
|
1375
1448
|
|
|
1376
|
-
- **An encoding-regression gate.** `smoke-keychain.sh` stored a short ASCII word, which is exactly the value class that never reproduced the bug. It now round-trips a multi-line service-account JSON (and asserts it still
|
|
1449
|
+
- **An encoding-regression gate.** `smoke-keychain.sh` stored a short ASCII word, which is exactly the value class that never reproduced the bug. It now round-trips a multi-line service-account JSON (and asserts it still _parses_, not merely compares equal), a hex-looking string, embedded quotes and backslashes, non-ASCII, and leading/trailing spaces.
|
|
1377
1450
|
|
|
1378
1451
|
- **A gate for the shipped-but-dead script.** `firebase-app-discovery.sh` was written, installed onto every machine by all three hosts, and invoked by nothing: its only mention sat inside a comment in another script. That is the declared-but-inert defect in its purest form - the docs describe a capability, the install carries the code, and nothing connects them. `smoke-consumer-smoke-surface.sh` now checks every shipped script for an actual invoker, and is deliberately picky about what counts: comment lines, JSON schema descriptions, CHANGELOG and ROADMAP entries all name a script without ever running it, and each of those, left in the corpus, made the check pass on the broken state while it was being built.
|
|
1379
1452
|
|
package/README.md
CHANGED
|
@@ -43,12 +43,12 @@ package. Check first, then pick the row that matches:
|
|
|
43
43
|
node -v; npm -v; command -v node npm npx
|
|
44
44
|
```
|
|
45
45
|
|
|
46
|
-
| What you see
|
|
47
|
-
|
|
48
|
-
| nothing at all
|
|
49
|
-
| `node` works, `npx` does not
|
|
50
|
-
| nvm is installed but the shell does not see it | `source ~/.nvm/nvm.sh && nvm use --lts`, or just open a new terminal
|
|
51
|
-
| npm is ancient (< 5.2, which predates npx)
|
|
46
|
+
| What you see | What to do |
|
|
47
|
+
| ---------------------------------------------- | ---------------------------------------------------------------------------------------- |
|
|
48
|
+
| nothing at all | Install Node >= 20.11: `brew install node`, or the LTS installer from nodejs.org |
|
|
49
|
+
| `node` works, `npx` does not | `npm i -g @mmerterden/multi-agent-pipeline` then `multi-agent-pipeline install --claude` |
|
|
50
|
+
| nvm is installed but the shell does not see it | `source ~/.nvm/nvm.sh && nvm use --lts`, or just open a new terminal |
|
|
51
|
+
| npm is ancient (< 5.2, which predates npx) | `npm i -g npm@latest`, or use the global-install row above |
|
|
52
52
|
|
|
53
53
|
And the path that needs neither `npx` nor a global install - clone and run the
|
|
54
54
|
installer directly:
|
|
@@ -116,9 +116,10 @@ The other half of the change is that the count is finally guarded.
|
|
|
116
116
|
`smoke-phase-contract.sh` derives it from `pipeline/schemas/phases.json` and
|
|
117
117
|
holds every other copy to it: the generator's output, the token budget, the
|
|
118
118
|
state-schema bounds, the progress fractions in sample output, and the named
|
|
119
|
-
thresholds
|
|
120
|
-
|
|
121
|
-
the command count, the jq count and the persona
|
|
119
|
+
thresholds, none of which is a literal in a script any more. A phase count is
|
|
120
|
+
otherwise the kind of number that spreads across dozens of files with nothing
|
|
121
|
+
checking any of them, the way the command count, the jq count and the persona
|
|
122
|
+
count are each held by a gate.
|
|
122
123
|
|
|
123
124
|
Reasoning, mapping and rejected alternatives:
|
|
124
125
|
[ADR-0014](./docs/adr/0014-six-phase-consolidation.md).
|
|
@@ -175,6 +176,7 @@ The discipline behind all of this - bounded loops, evidence gates, token-budgete
|
|
|
175
176
|
| Audit | `/multi-agent:testflight-validation` | Pre-submission gates for a TestFlight build: static archive audit → Apple's `altool --validate-app` → Review-Guidelines check. Validates only, never uploads |
|
|
176
177
|
|
|
177
178
|
Depth, autopilot and `--local` are the only knobs on the run itself; everything else is its own command. The full catalog is below.
|
|
179
|
+
|
|
178
180
|
### Pipeline depth: Full or Short
|
|
179
181
|
|
|
180
182
|
Depth is the one question the run asks about its own shape. `/multi-agent` and `/multi-agent:local` ask it at Phase 0 Step 7.5 - after the issue is fetched and the task type is known, because that is what the recommendation is drawn from.
|
|
@@ -186,7 +188,6 @@ Short is right when you already know the fix and the file: a one-line guard, a c
|
|
|
186
188
|
|
|
187
189
|
The widget follows the answer rather than predicting it: Phase 0 is the only tile drawn before you choose, and a Short run never draws an Analysis tile at all. Both autopilot entries skip the question and always run Full. `agent-state.json` records which one ran as `onlyDevelop`.
|
|
188
190
|
|
|
189
|
-
|
|
190
191
|
## Commands
|
|
191
192
|
|
|
192
193
|
`/multi-agent` plus 56 sub-commands. `/multi-agent:help` renders the same catalog in your terminal, in your `outputLanguage`.
|
|
@@ -300,11 +301,37 @@ Two compliance skills install on every host and back the store gates: `apple-arc
|
|
|
300
301
|
Everything above starts when you start it. Continuous mode is the same pipeline
|
|
301
302
|
picking work up on its own, on ONE machine you choose, from repos you choose.
|
|
302
303
|
|
|
303
|
-
| Command
|
|
304
|
-
|
|
|
305
|
-
| `/multi-agent:autopilot-on`
|
|
306
|
-
| `/multi-agent:autopilot-status` | What is running and at which phase, what is queued, what is waiting for an answer, the PRs of the last day, and the rolling spend
|
|
307
|
-
| `/multi-agent:autopilot-off`
|
|
304
|
+
| Command | What it does |
|
|
305
|
+
| ------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
306
|
+
| `/multi-agent:autopilot-on` | Pick the repos this machine watches. Labelled GitHub issues and assigned + labelled Jira items then run in a worktree and stop at an open PR. Re-run to change the list |
|
|
307
|
+
| `/multi-agent:autopilot-status` | What is running and at which phase, what is queued, what is waiting for an answer, the PRs of the last day, and the rolling spend |
|
|
308
|
+
| `/multi-agent:autopilot-off` | Remove the schedule. Work already running finishes; the repo selection is kept |
|
|
309
|
+
|
|
310
|
+
## Which model answers
|
|
311
|
+
|
|
312
|
+
The ladder is `fable -> opus -> sonnet -> haiku`, and a run walks it on failure.
|
|
313
|
+
Routing lets a policy pick the rung instead, per call site, and is off until you
|
|
314
|
+
turn it on.
|
|
315
|
+
|
|
316
|
+
| Command | What it does |
|
|
317
|
+
| --------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
318
|
+
| `/multi-agent:route-on` | Pick a strategy, a scope and the rules that say which rung a call lands on. Validated against `route-config.schema.json` before anything is written |
|
|
319
|
+
| `/multi-agent:route-off` | Disarm it. The rules are KEPT, so turning it back on does not re-ask for the same configuration |
|
|
320
|
+
| `/multi-agent:route-status` | Whether it is armed, which rule applies where, which rung the last dispatches took, and what this run has cost |
|
|
321
|
+
| `/multi-agent:model` | Turn the top rung on or off, and move the cost ledger's pricing with it in the same step |
|
|
322
|
+
|
|
323
|
+
**The scope cannot be the whole session.** `scope` accepts `subagent`,
|
|
324
|
+
`bulk-read` and `research`; there is no `host-session` value and that is a schema
|
|
325
|
+
gate, not a convention. Owning the host's base URL would send every call you make
|
|
326
|
+
through a third layer, including work that has nothing to do with this pipeline.
|
|
327
|
+
|
|
328
|
+
**The honest limit, which `route-status` prints rather than hides.** A rung on a
|
|
329
|
+
non-Anthropic provider is refused at dispatch. No call site sends a model
|
|
330
|
+
request itself: the one that routes, `bulk-read.sh`, delegates to the `claude`
|
|
331
|
+
CLI, which speaks that ladder and nothing else. Sending a read to a third-party
|
|
332
|
+
provider would also put the file's full text outside the account that owns it,
|
|
333
|
+
so it is a decision a user makes rather than a default. Routing chooses inside
|
|
334
|
+
the ladder.
|
|
308
335
|
|
|
309
336
|
**Nothing is on by default and nothing is added implicitly.** Installing the
|
|
310
337
|
package writes no state and schedules nothing; `smoke-autopilot-default-off.sh`
|
package/README.tr.md
CHANGED
|
@@ -43,12 +43,12 @@ bak, sonra sana uyan satırı uygula:
|
|
|
43
43
|
node -v; npm -v; command -v node npm npx
|
|
44
44
|
```
|
|
45
45
|
|
|
46
|
-
| Gördüğün
|
|
47
|
-
|
|
48
|
-
| hiçbiri yok
|
|
49
|
-
| `node` çalışıyor, `npx` çalışmıyor
|
|
50
|
-
| nvm kurulu ama kabuk görmüyor
|
|
51
|
-
| npm çok eski (< 5.2, npx'ten önceki sürümler) | `npm i -g npm@latest`, ya da üstteki global kurulum satırı
|
|
46
|
+
| Gördüğün | Yapılacak |
|
|
47
|
+
| --------------------------------------------- | ----------------------------------------------------------------------------------------- |
|
|
48
|
+
| hiçbiri yok | Node >= 20.11 kur: `brew install node` ya da nodejs.org'dan LTS installer |
|
|
49
|
+
| `node` çalışıyor, `npx` çalışmıyor | `npm i -g @mmerterden/multi-agent-pipeline` sonra `multi-agent-pipeline install --claude` |
|
|
50
|
+
| nvm kurulu ama kabuk görmüyor | `source ~/.nvm/nvm.sh && nvm use --lts`, ya da yeni bir terminal aç |
|
|
51
|
+
| npm çok eski (< 5.2, npx'ten önceki sürümler) | `npm i -g npm@latest`, ya da üstteki global kurulum satırı |
|
|
52
52
|
|
|
53
53
|
Ne `npx` ne de global kurulum isteyen yol - klonla ve installer'ı doğrudan çalıştır:
|
|
54
54
|
|
|
@@ -156,6 +156,7 @@ Bunun arkasındaki disiplin - sınırlı loop'lar, kanıt kapıları, token-büt
|
|
|
156
156
|
| Audit | `/multi-agent:testflight-validation` | TestFlight build için pre-submission kapıları: statik archive denetimi → Apple'ın `altool --validate-app`'i → Review-Guidelines kontrolü. Yalnızca doğrular, asla yüklemez |
|
|
157
157
|
|
|
158
158
|
Koşunun kendisinde ayarlanabilen tek şey derinlik, autopilot ve `--local`; geri kalan her şey kendi komutu. Tam katalog aşağıda.
|
|
159
|
+
|
|
159
160
|
### Pipeline derinliği: Tam mı Kısa mı
|
|
160
161
|
|
|
161
162
|
Derinlik, koşunun kendi şekli hakkında sorduğu tek soru. `/multi-agent` ve `/multi-agent:local` bunu Faz 0 Adım 7.5'te sorar - issue çekildikten ve görev tipi belirlendikten sonra, çünkü öneri onlardan çıkıyor.
|
|
@@ -167,7 +168,6 @@ Kısa'yı düzeltmeyi ve dosyayı zaten biliyorsan seç: tek satırlık bir guar
|
|
|
167
168
|
|
|
168
169
|
Widget cevabı tahmin etmek yerine takip eder: sen seçmeden önce yalnızca Faz 0 karosu çizilir, Kısa koşu Analiz karosunu hiç çizmez. İki autopilot girişi de bu soruyu sormaz, her zaman Tam koşar. Hangisinin koştuğunu `agent-state.json` `onlyDevelop` alanında tutar.
|
|
169
170
|
|
|
170
|
-
|
|
171
171
|
## Komutlar
|
|
172
172
|
|
|
173
173
|
`/multi-agent` ve 56 alt komut. `/multi-agent:help` aynı katalogu terminalde, `outputLanguage` ayarına göre gösterir.
|
|
@@ -282,11 +282,37 @@ Yukarıdaki her şey sen başlattığında başlar. Sürekli mod, aynı pipeline
|
|
|
282
282
|
**senin seçtiğin tek makinede**, **senin seçtiğin repolardan** işi kendi başına
|
|
283
283
|
alması.
|
|
284
284
|
|
|
285
|
-
| Komut
|
|
286
|
-
|
|
|
287
|
-
| `/multi-agent:autopilot-on`
|
|
288
|
-
| `/multi-agent:autopilot-status` | Ne koşuyor hangi fazda, sırada ne var, ne cevap bekliyor, son bir günün PR'ları ve yuvarlanan harcama
|
|
289
|
-
| `/multi-agent:autopilot-off`
|
|
285
|
+
| Komut | Ne yapar |
|
|
286
|
+
| ------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
287
|
+
| `/multi-agent:autopilot-on` | Bu makinenin izleyeceği repoları seçersin. Etiketli GitHub issue'ları ve sana atanmış + etiketli Jira maddeleri worktree'de koşar, açık PR'da durur. Listeyi değiştirmek için tekrar çalıştır |
|
|
288
|
+
| `/multi-agent:autopilot-status` | Ne koşuyor hangi fazda, sırada ne var, ne cevap bekliyor, son bir günün PR'ları ve yuvarlanan harcama |
|
|
289
|
+
| `/multi-agent:autopilot-off` | Zamanlamayı kaldırır. Koşan iş biter; repo seçimi saklanır |
|
|
290
|
+
|
|
291
|
+
## Hangi model cevaplıyor
|
|
292
|
+
|
|
293
|
+
Merdiven `fable -> opus -> sonnet -> haiku`; koşu başarısızlıkta merdiveni iner.
|
|
294
|
+
Yönlendirme basamağı çağrı yeri başına bir politikaya seçtirir ve sen açana
|
|
295
|
+
kadar kapalıdır.
|
|
296
|
+
|
|
297
|
+
| Komut | Ne yapar |
|
|
298
|
+
| --------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
299
|
+
| `/multi-agent:route-on` | Strateji, kapsam ve hangi çağrının hangi basamağa düşeceğini söyleyen kuralları sorar. Hiçbir şey yazılmadan önce `route-config.schema.json`'a karşı doğrulanır |
|
|
300
|
+
| `/multi-agent:route-off` | Kapatır. Kurallar SAKLANIR, tekrar açıldığında aynı yapılandırma yeniden sorulmaz |
|
|
301
|
+
| `/multi-agent:route-status` | Açık mı, hangi kural nerede geçerli, son çağrılar hangi basamağa gitti, bu koşu ne tuttu |
|
|
302
|
+
| `/multi-agent:model` | Üst basamağı açıp kapatır ve maliyet defterinin fiyatlamasını aynı adımda onunla taşır |
|
|
303
|
+
|
|
304
|
+
**Kapsam oturumun tamamı olamaz.** `scope` yalnız `subagent`, `bulk-read` ve
|
|
305
|
+
`research` kabul eder; `host-session` diye bir değer yoktur ve bu bir teamül
|
|
306
|
+
değil, şema kapısıdır. Host'un base URL'ini sahiplenmek, bu pipeline'la ilgisi
|
|
307
|
+
olmayan işler dahil yaptığın her çağrıyı üçüncü bir katmandan geçirirdi.
|
|
308
|
+
|
|
309
|
+
**Dürüst sınır, ve `route-status` bunu gizlemek yerine yazar.** Anthropic
|
|
310
|
+
dışı bir sağlayıcıdaki basamak dağıtımda reddedilir. Hiçbir çağrı yeri model
|
|
311
|
+
isteğini kendisi atmıyor: yönlendirilen tek yer olan `bulk-read.sh` işi `claude`
|
|
312
|
+
CLI'ına devrediyor, o da yalnız bu merdiveni konuşuyor. Bir okumayı üçüncü bir
|
|
313
|
+
sağlayıcıya göndermek ayrıca dosyanın tam metnini onu sahiplenen hesabın dışına
|
|
314
|
+
çıkarırdı; bu yüzden varsayılan değil, kullanıcının verdiği bir karardır.
|
|
315
|
+
Yönlendirme merdivenin içinde seçim yapar.
|
|
290
316
|
|
|
291
317
|
**Varsayılan olarak hiçbir şey açık değil ve hiçbir repo kendiliğinden eklenmez.**
|
|
292
318
|
Paketi kurmak ne durum yazar ne zamanlama kurar; bu bir gün değişirse
|
package/docs/ecosystem.md
CHANGED
|
@@ -212,10 +212,9 @@ that step's own contract).
|
|
|
212
212
|
### multi-agent-toolkit-mcp's tools, by category
|
|
213
213
|
|
|
214
214
|
115 tools in 13 categories, counted from the server's own `tools/list` response
|
|
215
|
-
at toolkit 3.13.
|
|
216
|
-
|
|
217
|
-
|
|
218
|
-
the server - which is why the count now names its source.
|
|
215
|
+
at toolkit 3.13.1 rather than from a README. Three families are defined in
|
|
216
|
+
modules the main file only spreads in, so a count taken from the source reads
|
|
217
|
+
low; the response is the only place the whole surface appears at once.
|
|
219
218
|
|
|
220
219
|
The families can be narrowed per host: `MCP_TOOLKIT_CAPS` serves only the named
|
|
221
220
|
ones, and unset serves all 115. A frontend repo that sets `web,code,context`
|
package/docs/facts.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"$comment": "Generated by pipeline/scripts/gen-facts.mjs. Do not hand-edit: the site reads this, and a number edited here instead of at its source is the drift this file removes.",
|
|
3
|
-
"generatedAt": "2026-09-
|
|
4
|
-
"version": "19.1.
|
|
3
|
+
"generatedAt": "2026-09-21",
|
|
4
|
+
"version": "19.1.4",
|
|
5
5
|
"phaseSchema": 2,
|
|
6
6
|
"phases": [
|
|
7
7
|
{
|
|
@@ -56,6 +56,6 @@
|
|
|
56
56
|
"research": 2,
|
|
57
57
|
"media": 1
|
|
58
58
|
},
|
|
59
|
-
"pluginSkillCount":
|
|
60
|
-
"toolkitVersion": "3.13.
|
|
59
|
+
"pluginSkillCount": 239,
|
|
60
|
+
"toolkitVersion": "3.13.1"
|
|
61
61
|
}
|
package/install/_common.mjs
CHANGED
|
@@ -485,11 +485,11 @@ export function pruneOrphanSkillFiles(skillsDir) {
|
|
|
485
485
|
*
|
|
486
486
|
* `.skills-index.json` is what `match-skills.mjs` reads when
|
|
487
487
|
* `prefs.global.dynamicSkillLoading` is on; `skills-index.md` is the human-readable
|
|
488
|
-
* twin. Both
|
|
489
|
-
* schema
|
|
490
|
-
* the pref on
|
|
491
|
-
* fall back to eager loading. One helper, called from every
|
|
492
|
-
* install modes cannot disagree about it
|
|
488
|
+
* twin. Both ship on a full install, not only under `--index-only`: the
|
|
489
|
+
* preferences schema states that a full install always carries the index, and an
|
|
490
|
+
* install that skips it makes turning the pref on afterwards produce `cannot read
|
|
491
|
+
* index` and a silent fall back to eager loading. One helper, called from every
|
|
492
|
+
* target, so the two install modes cannot disagree about it.
|
|
493
493
|
*
|
|
494
494
|
* @param {string} pipelineSrc - the repo's `pipeline/` directory
|
|
495
495
|
* @param {string} skillsDest - the host's skills directory
|
|
@@ -152,7 +152,7 @@ function saysAlreadyExists(text) {
|
|
|
152
152
|
* automatically a failure, and a zero exit is not automatically a fresh write - which
|
|
153
153
|
* is why the "already exists" test runs on both paths. Reporting an existing server as
|
|
154
154
|
* "skipped MCP registration (Command failed)" would send the user to fix something that
|
|
155
|
-
* is already correct
|
|
155
|
+
* is already correct.
|
|
156
156
|
*
|
|
157
157
|
* An existing entry under the same name is removed and re-added so its spec follows
|
|
158
158
|
* the installed version (the entry is installer-owned; the reason is at the remove
|
|
@@ -4,9 +4,8 @@
|
|
|
4
4
|
* Claude Code loads `{owner}/multi-agent-plugins` natively, so it gets all of a
|
|
5
5
|
* stack plugin's skills. The other two hosts do not:
|
|
6
6
|
*
|
|
7
|
-
* - **Copilot CLI** has no plugin loader at all.
|
|
8
|
-
* standalone `figma-*`
|
|
9
|
-
* pruned exactly those directories, so in practice Copilot had **zero**
|
|
7
|
+
* - **Copilot CLI** has no plugin loader at all. `install/copilot.mjs` prunes the
|
|
8
|
+
* standalone `figma-*` directories, so Copilot carries **zero**
|
|
10
9
|
* plugin-authored skills - `create-screen`, `figma-validate`, `figma-review`,
|
|
11
10
|
* `component`, `state` and `navigation` were all absent. A component task on
|
|
12
11
|
* Copilot therefore had nothing to dispatch to.
|
|
@@ -221,7 +220,7 @@ export function installAuthoredPluginSkills(opts) {
|
|
|
221
220
|
.map((e) => ({ name: e.name, from: join(groupDir, e.name) }));
|
|
222
221
|
|
|
223
222
|
for (const { name, from } of entries) {
|
|
224
|
-
// A plugin skill whose name a PIPELINE skill already owns
|
|
223
|
+
// A plugin skill whose name a PIPELINE skill already owns is KEPT.
|
|
225
224
|
// That is not a duplicate: the pipeline's `architecture` is a generic ADR
|
|
226
225
|
// framework while the iOS plugin's is that stack's structural rules, and the
|
|
227
226
|
// same holds for `backlog`. Claude Code reaches both because its loader
|
package/install/copilot.mjs
CHANGED
|
@@ -86,11 +86,11 @@ export function installCopilot(ctx) {
|
|
|
86
86
|
/**
|
|
87
87
|
* `$HOME/.claude/...` rewrites applied to skill files copied into the Copilot tree.
|
|
88
88
|
*
|
|
89
|
-
*
|
|
90
|
-
*
|
|
91
|
-
*
|
|
92
|
-
*
|
|
93
|
-
*
|
|
89
|
+
* Copied byte-for-byte, a Copilot-only install would carry every `~/.claude/scripts`
|
|
90
|
+
* and `~/.claude/lib` reference in the skill text - trees that install never creates
|
|
91
|
+
* on that host. Each one is a dangling path: the skill tells the agent to run a
|
|
92
|
+
* script at a location that does not exist on that machine. Codex's installer
|
|
93
|
+
* applies the same rewrite for the same reason.
|
|
94
94
|
*
|
|
95
95
|
* Only the trees Copilot actually OWNS are rewritten:
|
|
96
96
|
*
|