@mmerterden/multi-agent-pipeline 17.6.0 → 19.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +310 -0
- package/README.md +76 -18
- package/README.tr.md +55 -16
- package/docs/adr/0002-instruction-driven-flag.md +1 -0
- package/docs/adr/0005-lazy-phase-docs.md +11 -1
- package/docs/adr/0008-installer-modularization-and-secret-leak-defense.md +1 -0
- package/docs/adr/0010-own-code-graph.md +1 -0
- package/docs/adr/0011-dormant-ci.md +25 -1
- package/docs/adr/0014-six-phase-consolidation.md +134 -0
- package/docs/adr/README.md +2 -1
- package/docs/architecture.md +37 -38
- package/docs/best-practices.md +1 -1
- package/docs/ecosystem.md +37 -26
- package/docs/engineering.md +1 -1
- package/docs/facts.json +45 -0
- package/docs/features.md +54 -53
- package/docs/performance.md +5 -5
- package/docs/recovery-guide.md +9 -9
- package/docs/server-readiness.md +188 -0
- package/docs/token-budget-history.md +3 -1
- package/index.js +18 -3
- package/install/_codex-agents.mjs +1 -1
- package/install/_common.mjs +42 -17
- package/install/_dev-only-files.mjs +8 -0
- package/install/_unattended-profile.mjs +113 -0
- package/install/index.mjs +48 -0
- package/install/templates/claude-hooks.json +1 -1
- package/install/templates/codex-instructions.md +1 -1
- package/install/templates/copilot-instructions.md +28 -28
- package/manifest.json +1065 -0
- package/package.json +6 -3
- package/pipeline/agents/dev-critic.md +3 -3
- package/pipeline/commands/figma-to-swiftui.md +1 -1
- package/pipeline/commands/multi-agent/SKILL.md +8 -8
- package/pipeline/commands/multi-agent/analysis/SKILL.md +9 -9
- package/pipeline/commands/multi-agent/autopilot/SKILL.md +7 -7
- package/pipeline/commands/multi-agent/channels/SKILL.md +15 -15
- package/pipeline/commands/multi-agent/diff-explain/SKILL.md +6 -6
- package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/graph/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/help/SKILL.md +62 -62
- package/pipeline/commands/multi-agent/language/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/local/SKILL.md +11 -11
- package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +13 -13
- package/pipeline/commands/multi-agent/log/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/manual-test/SKILL.md +9 -9
- package/pipeline/commands/multi-agent/model/SKILL.md +69 -0
- package/pipeline/commands/multi-agent/refactor/SKILL.md +3 -3
- package/pipeline/commands/multi-agent/resume/SKILL.md +4 -4
- package/pipeline/commands/multi-agent/resume-local/SKILL.md +19 -17
- package/pipeline/commands/multi-agent/review/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/route-off/SKILL.md +36 -0
- package/pipeline/commands/multi-agent/route-on/SKILL.md +74 -0
- package/pipeline/commands/multi-agent/route-status/SKILL.md +56 -0
- package/pipeline/commands/multi-agent/setup/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/status/SKILL.md +54 -23
- package/pipeline/commands/multi-agent/steer/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/sync/SKILL.md +12 -13
- package/pipeline/commands/multi-agent/test/SKILL.md +1 -1
- package/pipeline/lib/_jira-auth.sh +8 -0
- package/pipeline/lib/analysis-jira-write.sh +32 -0
- package/pipeline/lib/ask-choice.sh +13 -2
- package/pipeline/lib/autopilot-state.sh +8 -0
- package/pipeline/lib/credential-inventory.sh +1 -1
- package/pipeline/lib/fatal.mjs +129 -0
- package/pipeline/lib/fetch-fortify.sh +1 -1
- package/pipeline/lib/figma-mcp-refresh.sh +18 -0
- package/pipeline/lib/figma-screenshot.sh +18 -0
- package/pipeline/lib/invoked-directly.mjs +43 -0
- package/pipeline/lib/jira-publish.sh +42 -0
- package/pipeline/lib/md2confluence-v3.py +47 -0
- package/pipeline/lib/model-rung.sh +142 -0
- package/pipeline/lib/outbound-gate.mjs +175 -0
- package/pipeline/lib/phase-schema.mjs +88 -0
- package/pipeline/lib/plan-todos.sh +32 -11
- package/pipeline/lib/post-pr-review.sh +77 -8
- package/pipeline/lib/repo-hygiene.sh +8 -3
- package/pipeline/lib/require-jq.sh +40 -0
- package/pipeline/lib/route-state.sh +161 -0
- package/pipeline/lib/run-paths.sh +335 -0
- package/pipeline/multi-agent-refs/_account-picker.md +1 -1
- package/pipeline/multi-agent-refs/_dev-context.md +1 -1
- package/pipeline/multi-agent-refs/_input-parser.md +1 -1
- package/pipeline/multi-agent-refs/analysis/evidence.md +0 -9
- package/pipeline/multi-agent-refs/analysis/intake.md +1 -1
- package/pipeline/multi-agent-refs/analysis/locked.md +21 -22
- package/pipeline/multi-agent-refs/analysis/render.md +1 -1
- package/pipeline/multi-agent-refs/analysis/synthesis.md +12 -6
- package/pipeline/multi-agent-refs/android-guide.md +1 -1
- package/pipeline/multi-agent-refs/audit-guide.md +13 -13
- package/pipeline/multi-agent-refs/channels/issue-comment.md +2 -2
- package/pipeline/multi-agent-refs/channels/jira.md +3 -3
- package/pipeline/multi-agent-refs/channels/pr.md +4 -4
- package/pipeline/multi-agent-refs/channels/wiki.md +1 -1
- package/pipeline/multi-agent-refs/component-dispatch.md +3 -3
- package/pipeline/multi-agent-refs/cross-cli-contract.md +31 -6
- package/pipeline/multi-agent-refs/features/autopilot-circuit-breaker.md +74 -4
- package/pipeline/multi-agent-refs/features/code-graph.md +5 -5
- package/pipeline/multi-agent-refs/features/cost-analysis.md +93 -0
- package/pipeline/multi-agent-refs/features/design-conformance.md +1 -1
- package/pipeline/multi-agent-refs/features/dev-critic.md +3 -3
- package/pipeline/multi-agent-refs/features/doctor.md +47 -2
- package/pipeline/multi-agent-refs/features/external-context-injection.md +3 -3
- package/pipeline/multi-agent-refs/features/maturity-followup.md +3 -3
- package/pipeline/multi-agent-refs/features/model-fallback.md +5 -5
- package/pipeline/multi-agent-refs/features/plan-todos.md +1 -1
- package/pipeline/multi-agent-refs/features/repo-map.md +1 -1
- package/pipeline/multi-agent-refs/features/review-delta.md +3 -3
- package/pipeline/multi-agent-refs/features/review-multi-repo.md +1 -1
- package/pipeline/multi-agent-refs/features/scope-check.md +4 -4
- package/pipeline/multi-agent-refs/features/skill-conformance.md +2 -2
- package/pipeline/multi-agent-refs/features/stack-skill-routing.md +1 -1
- package/pipeline/multi-agent-refs/features/verify-by-test.md +4 -4
- package/pipeline/multi-agent-refs/features/verify.md +83 -0
- package/pipeline/multi-agent-refs/features/visual-evidence.md +19 -19
- package/pipeline/multi-agent-refs/features/worktree-finalize.md +6 -6
- package/pipeline/multi-agent-refs/issue-jira-triad.md +10 -10
- package/pipeline/multi-agent-refs/knowledge.md +11 -11
- package/pipeline/multi-agent-refs/multi-repo-integration-build.md +13 -13
- package/pipeline/multi-agent-refs/payload-contracts.md +8 -8
- package/pipeline/multi-agent-refs/phases/log-format.md +10 -10
- package/pipeline/multi-agent-refs/phases/modes.md +30 -30
- package/pipeline/multi-agent-refs/phases/operations.md +21 -10
- package/pipeline/multi-agent-refs/phases/phase-0-init.md +25 -25
- package/pipeline/multi-agent-refs/phases/phase-1-plan.md +599 -0
- package/pipeline/multi-agent-refs/phases/{phase-3-dev.md → phase-2-dev.md} +129 -49
- package/pipeline/multi-agent-refs/phases/{phase-4-review.md → phase-3-review.md} +225 -107
- package/pipeline/multi-agent-refs/phases/{phase-6-commit.md → phase-4-commit.md} +23 -23
- package/pipeline/multi-agent-refs/phases/{phase-7-report.md → phase-5-report.md} +29 -29
- package/pipeline/multi-agent-refs/phases.md +44 -48
- package/pipeline/multi-agent-refs/picker-contract.md +1 -1
- package/pipeline/multi-agent-refs/progress-contract.md +6 -6
- package/pipeline/multi-agent-refs/readiness-review.md +1 -1
- package/pipeline/multi-agent-refs/rules.md +7 -7
- package/pipeline/multi-agent-refs/swiftui-guide.md +2 -2
- package/pipeline/multi-agent-refs/tracker-contract.md +31 -32
- package/pipeline/multi-agent-refs/unattended-contract.md +129 -0
- package/pipeline/multi-agent-refs/wiki-capture.md +14 -14
- package/pipeline/preferences-template.json +9 -1
- package/pipeline/rules/outside-the-pipeline.md +1 -1
- package/pipeline/schemas/agent-state.schema.json +50 -50
- package/pipeline/schemas/analysis-output.schema.json +2 -2
- package/pipeline/schemas/autopilot-config.schema.json +1 -1
- package/pipeline/schemas/code-graph.schema.json +1 -1
- package/pipeline/schemas/criteria-manifest.schema.json +1 -1
- package/pipeline/schemas/dev-critic-output.schema.json +1 -1
- package/pipeline/schemas/diff-risk.schema.json +1 -1
- package/pipeline/schemas/migrations/prefs-2.4.0-to-2.5.0.mjs +2 -2
- package/pipeline/schemas/migrations/prefs-2.6.0-to-2.7.0.mjs +31 -0
- package/pipeline/schemas/migrations/state-2.1.0-to-2.2.0.mjs +129 -0
- package/pipeline/schemas/phases.json +105 -0
- package/pipeline/schemas/plan-todos.schema.json +5 -5
- package/pipeline/schemas/planning-output.schema.json +1 -1
- package/pipeline/schemas/prefs.schema.json +100 -56
- package/pipeline/schemas/reviewer-output.schema.json +3 -3
- package/pipeline/schemas/route-config.schema.json +74 -0
- package/pipeline/schemas/scope-check.schema.json +1 -1
- package/pipeline/schemas/test-gap.schema.json +1 -1
- package/pipeline/schemas/token-budget.json +12 -18
- package/pipeline/schemas/triage-output.schema.json +6 -6
- package/pipeline/scripts/README.md +3 -3
- package/pipeline/scripts/_code-graph.mjs +2 -2
- package/pipeline/scripts/_run-paths.mjs +372 -0
- package/pipeline/scripts/_smoke-root.sh +1 -1
- package/pipeline/scripts/aggregate-metrics.mjs +65 -65
- package/pipeline/scripts/autopilot-arming.mjs +2 -1
- package/pipeline/scripts/autopilot-intake.mjs +2 -1
- package/pipeline/scripts/autopilot-runner.mjs +206 -2
- package/pipeline/scripts/build-references.mjs +2 -1
- package/pipeline/scripts/build-stack-plugins.mjs +10 -2
- package/pipeline/scripts/capture-evidence.sh +7 -2
- package/pipeline/scripts/capture-flush.sh +8 -8
- package/pipeline/scripts/capture-resume.sh +3 -3
- package/pipeline/scripts/classify-plan-safety.mjs +3 -2
- package/pipeline/scripts/cost-analyze.mjs +600 -0
- package/pipeline/scripts/cost-budget-check.mjs +4 -12
- package/pipeline/scripts/council-view.mjs +2 -1
- package/pipeline/scripts/crush-json.mjs +2 -1
- package/pipeline/scripts/diff-explain.mjs +7 -10
- package/pipeline/scripts/diff-risk-score.mjs +2 -1
- package/pipeline/scripts/doctor.mjs +140 -6
- package/pipeline/scripts/evidence-gate.mjs +9 -3
- package/pipeline/scripts/feedback-send.mjs +12 -2
- package/pipeline/scripts/gc-abandoned.sh +32 -16
- package/pipeline/scripts/gc-tmp.sh +1 -1
- package/pipeline/scripts/gc-worktrees.sh +12 -5
- package/pipeline/scripts/gen-facts.mjs +175 -0
- package/pipeline/scripts/gen-mode-dispatch.mjs +32 -37
- package/pipeline/scripts/gen-ref-toc.mjs +1 -1
- package/pipeline/scripts/github-ssh-setup.sh +64 -7
- package/pipeline/scripts/graph-mermaid.mjs +4 -2
- package/pipeline/scripts/graph-report.mjs +1 -1
- package/pipeline/scripts/jira-attach.sh +1 -1
- package/pipeline/scripts/keychain-save.sh +101 -30
- package/pipeline/scripts/learn-from-transcripts.mjs +3 -2
- package/pipeline/scripts/learning-curve.mjs +36 -31
- package/pipeline/scripts/log-metric.sh +17 -4
- package/pipeline/scripts/make-manifest.mjs +199 -0
- package/pipeline/scripts/memory-save.sh +1 -1
- package/pipeline/scripts/migrate-prefs.mjs +24 -6
- package/pipeline/scripts/migrate-state.mjs +94 -4
- package/pipeline/scripts/phase-banner.sh +26 -22
- package/pipeline/scripts/phase-tracker.sh +48 -10
- package/pipeline/scripts/plan-coverage-gate.mjs +8 -4
- package/pipeline/scripts/pre-commit-check.sh +7 -0
- package/pipeline/scripts/pre-push-check.sh +7 -0
- package/pipeline/scripts/purge.sh +23 -6
- package/pipeline/scripts/render-agent-log-cost.sh +10 -3
- package/pipeline/scripts/render-cost-summary.sh +9 -2
- package/pipeline/scripts/render-work-summary.sh +14 -7
- package/pipeline/scripts/review-file-filter.mjs +5 -3
- package/pipeline/scripts/review-scope.mjs +2 -1
- package/pipeline/scripts/routine-registry.mjs +2 -1
- package/pipeline/scripts/run-aggregator.mjs +26 -20
- package/pipeline/scripts/run-metrics.mjs +4 -2
- package/pipeline/scripts/runs-index.mjs +353 -0
- package/pipeline/scripts/scorecard-snapshot.mjs +178 -0
- package/pipeline/scripts/search-logs.sh +18 -0
- package/pipeline/scripts/smoke-cross-cli-behavior.sh +6 -6
- package/pipeline/scripts/smoke-schema-validation.sh +26 -7
- package/pipeline/scripts/test-gap-scan.mjs +2 -1
- package/pipeline/scripts/test-integrity-gate.mjs +2 -1
- package/pipeline/scripts/token-budget-report.mjs +13 -2
- package/pipeline/scripts/triage-memory.mjs +2 -2
- package/pipeline/scripts/update-issue-progress.sh +56 -7
- package/pipeline/scripts/usage-report.mjs +12 -1
- package/pipeline/scripts/validate-analysis-doc.mjs +75 -18
- package/pipeline/scripts/validate-code-graph.mjs +6 -3
- package/pipeline/scripts/validate-complaint-doc.mjs +2 -1
- package/pipeline/scripts/validate-diff-risk.mjs +6 -3
- package/pipeline/scripts/validate-planning.mjs +1 -1
- package/pipeline/scripts/validate-reviewer.mjs +1 -1
- package/pipeline/scripts/validate-state.mjs +45 -5
- package/pipeline/scripts/validate-test-gap.mjs +6 -3
- package/pipeline/scripts/validate-triage.mjs +6 -4
- package/pipeline/scripts/verify-citations.mjs +4 -2
- package/pipeline/scripts/verify.mjs +327 -0
- package/pipeline/scripts/worktree-finalize.sh +18 -9
- package/pipeline/scripts/write-state.mjs +154 -15
- package/pipeline/skills/.skill-manifest.json +37 -21
- package/pipeline/skills/.skills-index.json +104 -5
- package/pipeline/skills/shared/README.md +15 -6
- package/pipeline/skills/shared/core/apple-archive-compliance/SKILL.md +2 -2
- package/pipeline/skills/shared/core/google-play-compliance/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent/SKILL.md +69 -71
- package/pipeline/skills/shared/core/multi-agent-autopilot/SKILL.md +3 -3
- package/pipeline/skills/shared/core/multi-agent-channels/SKILL.md +14 -14
- package/pipeline/skills/shared/core/multi-agent-diff-explain/SKILL.md +5 -5
- package/pipeline/skills/shared/core/multi-agent-graph/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +25 -23
- package/pipeline/skills/shared/core/multi-agent-language/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-local/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +8 -8
- package/pipeline/skills/shared/core/multi-agent-manual-test/SKILL.md +6 -6
- package/pipeline/skills/shared/core/multi-agent-model/SKILL.md +71 -0
- package/pipeline/skills/shared/core/multi-agent-refactor/SKILL.md +3 -3
- package/pipeline/skills/shared/core/multi-agent-resume/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-resume-local/SKILL.md +7 -7
- package/pipeline/skills/shared/core/multi-agent-route-off/SKILL.md +39 -0
- package/pipeline/skills/shared/core/multi-agent-route-on/SKILL.md +76 -0
- package/pipeline/skills/shared/core/multi-agent-route-status/SKILL.md +59 -0
- package/pipeline/skills/shared/core/multi-agent-setup/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-status/SKILL.md +35 -11
- package/pipeline/skills/shared/core/multi-agent-steer/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +6 -5
- package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/package_app.sh +4 -1
- package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/setup_dev_signing.sh +4 -1
- package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/sign-and-notarize.sh +2 -1
- package/pipeline/skills/skills-index.md +13 -4
- package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +0 -263
- package/pipeline/multi-agent-refs/phases/phase-2-planning.md +0 -344
- package/pipeline/multi-agent-refs/phases/phase-5-test.md +0 -182
package/CHANGELOG.md
CHANGED
|
@@ -14,6 +14,316 @@ Internal file-layout changes that don't affect the slash-command surface are sti
|
|
|
14
14
|
|
|
15
15
|
---
|
|
16
16
|
|
|
17
|
+
## [19.0.0] - 2026-09-18
|
|
18
|
+
|
|
19
|
+
Major, because phase numbers are the contract and they moved. Eight phases
|
|
20
|
+
became six: two of the eight were doing the same work twice, and the count
|
|
21
|
+
itself was guarded by nothing.
|
|
22
|
+
|
|
23
|
+
Full reasoning, mapping and rejected alternatives:
|
|
24
|
+
[ADR-0014](./docs/adr/0014-six-phase-consolidation.md).
|
|
25
|
+
|
|
26
|
+
### Changed
|
|
27
|
+
|
|
28
|
+
- **Six phases.** `0 Init`, `1 Plan`, `2 Dev`, `3 Review`, `4 Commit`,
|
|
29
|
+
`5 Report`. Analysis and Planning were already one decision - the depth
|
|
30
|
+
picker skipped them together and `onlyDevelop` described them as one unit -
|
|
31
|
+
and they are one phase now. Review Stage 1 became Dev's exit gate, which is
|
|
32
|
+
what removes the second build: Dev built and tee'd a log, then Review built
|
|
33
|
+
again, and nothing consumed the difference. The user test moved inside
|
|
34
|
+
Review, keeping its waiting state.
|
|
35
|
+
- **`pipeline/schemas/phases.json` is the phase contract.** The list existed as
|
|
36
|
+
eight independent copies, none derived from another. The generator, the run
|
|
37
|
+
index, the metrics logger and the token budget read it now, and comparison
|
|
38
|
+
thresholds that used to be literals (`phase >= 6` for "waiting on you", the
|
|
39
|
+
Short-run boundary) are named fields in it.
|
|
40
|
+
- **`smoke-phase-contract.sh`** is the gate that never existed. The phase count
|
|
41
|
+
appeared in 91 places across 40 files with nothing holding any of them; the
|
|
42
|
+
command count, the jq count and the persona count all had gates. It derives
|
|
43
|
+
the count from the contract and checks the generator's output, the token
|
|
44
|
+
budget, the state-schema bounds, every shipped surface that states a count,
|
|
45
|
+
the progress fractions in sample output and the named thresholds. Verified
|
|
46
|
+
both ways: a deliberately wrong contract fails it.
|
|
47
|
+
- **`smoke-no-mcp-in-dev-phases.sh` keeps `phase >= 2`, and the gate now
|
|
48
|
+
asserts that it is unchanged.** Figma MCP was reachable only in Analysis (1);
|
|
49
|
+
Analysis is inside Plan (1). The permitted set `{0, 1}` is identical either
|
|
50
|
+
way, so the threshold surviving a renumbering is a result of the mapping
|
|
51
|
+
rather than an oversight - and an edit that "corrects" it would widen access.
|
|
52
|
+
- **Analysis has one pipeline.** Lite mode is removed. It chose sections from a
|
|
53
|
+
fixed list scored on three signals while Locked 2 chooses them from evidence,
|
|
54
|
+
and the two disagreed in both directions: a small feature with rich business
|
|
55
|
+
rules lost Section 15 because it was not on the list, and a feature with no
|
|
56
|
+
API contract kept Section 9 because it was. 37 Locked decisions became 36.
|
|
57
|
+
- **Locked 2's numbering clause was wrong and is corrected.** It said numbering
|
|
58
|
+
re-flows `1..N`; Locked 30 threads ids across the document BY NUMBER
|
|
59
|
+
(`Section 15.1`, `Section 4.4`), so a re-flowed document sends every one of
|
|
60
|
+
those references to the wrong section. No emitted document ever re-flowed.
|
|
61
|
+
A rendered section keeps its canonical number, gaps included, and the
|
|
62
|
+
validator now checks what actually holds: numbers inside the template range,
|
|
63
|
+
ascending, no repeats.
|
|
64
|
+
|
|
65
|
+
### Added
|
|
66
|
+
|
|
67
|
+
- **`/multi-agent:model`** turns the top rung on or off AND realigns
|
|
68
|
+
`costBudget.pricingModel` in the same write. The switch existed; the command
|
|
69
|
+
did not, and the pricing field it must move with was left to the user to
|
|
70
|
+
remember. It reports what the switch means on the host it runs on: a live
|
|
71
|
+
switch on Claude Code, a status report on Copilot CLI and Codex CLI.
|
|
72
|
+
- **`/multi-agent:route-on` · `:route-off` · `:route-status`** - policy-driven
|
|
73
|
+
model routing, shipping disabled. `scope` has no `host-session` member and
|
|
74
|
+
the schema enforces that: rewriting the host's base URL would route the
|
|
75
|
+
user's whole session, including work unrelated to this pipeline. `route-off`
|
|
76
|
+
keeps the rules, so `route-on` does not re-ask. `route-status` prints the
|
|
77
|
+
honest limit every time - a subagent cannot be sent to a non-Anthropic model,
|
|
78
|
+
because subagent dispatch belongs to the host.
|
|
79
|
+
- **`docs/facts.json`**, generated by `pipeline/scripts/gen-facts.mjs`: the
|
|
80
|
+
phase, command, skill and tool counts, derived rather than written. The
|
|
81
|
+
website read its own copies and said "8 faz + 51 komut" while the repo had
|
|
82
|
+
six phases and sixty commands. The tool count is asked of the toolkit's own
|
|
83
|
+
`tools/list` rather than counted out of its source, because three tool
|
|
84
|
+
families live in modules the main file only spreads in - a regex over
|
|
85
|
+
`index.js` returns 2 when the answer is 99.
|
|
86
|
+
- **`prefs.schema.json` moves to 2.7.0, and the shipped template moves with
|
|
87
|
+
it.** The template had been left at 2.6.0, which
|
|
88
|
+
`smoke-schema-validation.sh` catches by design: a fresh install that starts
|
|
89
|
+
behind the migration target makes an old entry in the migrator's accepted set
|
|
90
|
+
load-bearing purely to rescue the template. 2.7.0 removes the Lite value from
|
|
91
|
+
`analysisPhase.mode` and declares `global.modelRouting`, which the template
|
|
92
|
+
now ships explicitly disabled rather than leaving absent - a default that is
|
|
93
|
+
written down is one a reader can find.
|
|
94
|
+
- **The description-surface ceiling moves 86,500 -> 88,000**, and this is where
|
|
95
|
+
that has to be said. Four commands with a shared/core twin each is eight
|
|
96
|
+
descriptions; the average held at 316 against its own 420 ceiling, which is
|
|
97
|
+
the condition the gate's convention names for a raise rather than a trim -
|
|
98
|
+
the surface grew because there are more commands, not wordier ones. The
|
|
99
|
+
fixed per-run load was a different answer: it went 122 bytes over its 60,000
|
|
100
|
+
ceiling and the bytes were reclaimed from prose rather than the ceiling
|
|
101
|
+
raised, which is what that gate's message asks for in as many words.
|
|
102
|
+
- **`smoke-six-phase-run.sh`** drives phase-tracker.sh through a synthetic run
|
|
103
|
+
and asserts what a live run would show: six tiles named from the contract, no
|
|
104
|
+
tile above 5, the `Phase 2 Dev` line shape, a sub-step that registers under
|
|
105
|
+
its parent phase rather than as a seventh tile, a token count and a start
|
|
106
|
+
timestamp for the cost and elapsed suffixes, and a state file that validates
|
|
107
|
+
at 0..5 while `currentPhase: 6` is rejected. It says plainly what it does not
|
|
108
|
+
cover: that Review does not build a second time is an assertion about a model
|
|
109
|
+
following a document, and only a live run's `.build.log` mtime can show it.
|
|
110
|
+
Verified both ways - a seventh phase in the contract fails it.
|
|
111
|
+
- **A facts gate on the website** (`tests/facts-consistency.test.ts`). It does
|
|
112
|
+
not check that the copied `facts.json` is fresh - CI has no pipeline
|
|
113
|
+
checkout, that is `sync-facts.mjs --check` on a machine that does. It checks
|
|
114
|
+
what actually failed: that no component states a phase count or a phase
|
|
115
|
+
number that disagrees with the contract. Copy may say "6 phases"; it may not
|
|
116
|
+
say a different number. Verified both ways.
|
|
117
|
+
- **`metrics.jsonl` carries `phaseSchema`.** The file is append-only across a
|
|
118
|
+
renumbering, so `phase: 3` means Dev in a pre-v19 row and Review in a post-v19
|
|
119
|
+
one. Lines without the field are generation 1. `pipeline/lib/phase-schema.mjs`
|
|
120
|
+
resolves both, and the two aggregators that compared phase numbers to literals
|
|
121
|
+
go through it.
|
|
122
|
+
|
|
123
|
+
### Migration
|
|
124
|
+
|
|
125
|
+
- **`state-2.1.0-to-2.2.0.mjs`** maps `0→0, 1→1, 2→1, 3→2, 4→3, 5→3, 6→4, 7→5`.
|
|
126
|
+
Two sources can collide onto one `phases{}` key: furthest-along status wins,
|
|
127
|
+
`retryCount` takes the MAX (the schema caps it at 3, so a sum would emit an
|
|
128
|
+
invalid state), `files[]` union, earliest start, latest finish. It also
|
|
129
|
+
repairs four defects the live corpus already carried - `completed` →
|
|
130
|
+
`complete`, `awaiting-user-test-main-checkout` → `awaiting_input`, and
|
|
131
|
+
explicit defaults for a missing `currentPhase` or `status`. Measured on the
|
|
132
|
+
62 real state files: **52 valid before, 62 after.**
|
|
133
|
+
- **`prefs-2.6.0-to-2.7.0.mjs`** rewrites `analysisPhase.mode` from `auto` or
|
|
134
|
+
`lite` to `full`. The key is kept rather than deleted, so a file that set it
|
|
135
|
+
stays valid.
|
|
136
|
+
|
|
137
|
+
### Fixed
|
|
138
|
+
|
|
139
|
+
- `migrate-prefs.mjs` read its target version from a literal that had drifted
|
|
140
|
+
behind the schema. It reads the schema now, as do the two gates that were
|
|
141
|
+
checking against their own copies of it.
|
|
142
|
+
- `phases.md` carried a second token-budget table whose total said 17,000 while
|
|
143
|
+
the enforced file said 63,150 - wrong by a factor of four, for most of the
|
|
144
|
+
project's life, guarded by nothing. The numbers are gone; the enforced source
|
|
145
|
+
is named instead.
|
|
146
|
+
- `validate-analysis-doc.mjs` gated one half of Locked 2 and not the other. The
|
|
147
|
+
one emitted document available rendered `1..10, 12, 16, 20, 21` and nothing
|
|
148
|
+
looked.
|
|
149
|
+
- Eight pieces of dead code on the website, found by the linter rather than by
|
|
150
|
+
grep - which had already been wrong three times about this repo.
|
|
151
|
+
- **The phase bound is generation-aware, and it had to be.** Tightening
|
|
152
|
+
`currentPhase` to 0..5 marked every pre-v19 run log invalid - a run that
|
|
153
|
+
finished at phase 7 in October was correct when it was written, and a
|
|
154
|
+
validator that calls correct history invalid is one people learn to ignore.
|
|
155
|
+
A file stamped 2.2.0 or later is bounded 0..5; anything older, including the
|
|
156
|
+
files that predate stamping entirely, is bounded 0..7 and says so in the
|
|
157
|
+
error text. The same decision `metrics.jsonl` got: label the generation, do
|
|
158
|
+
not rewrite history. The tightening still bites where it matters - phase 7 on
|
|
159
|
+
a file claiming 2.2.0 is exactly what a skipped migration produces, and that
|
|
160
|
+
is rejected. Measured on the live corpus: 2 valid of 16 before, 11 of 16
|
|
161
|
+
after, and the 5 that remain were already invalid for reasons the plan had
|
|
162
|
+
recorded (no `currentPhase` at all, non-object phase values).
|
|
163
|
+
- `validate-state.mjs` did not check `retryCount`. The schema caps it at 3 and
|
|
164
|
+
four documents call 3 a hard kill, so `retryCount: 4` was a state that every
|
|
165
|
+
document forbade and every validator accepted. The bound is checked now. The
|
|
166
|
+
limit is named rather than overstated: this closes the validation boundary,
|
|
167
|
+
it does not stop the loop - that stays prose.
|
|
168
|
+
- **`phase-banner.sh` still had eight labels**, and it is the banner every
|
|
169
|
+
phase prints. Its table is bare words - `en:4) echo "Review"` - so all three
|
|
170
|
+
sweeps walked past it: they looked for `Phase 4 Review`, `4:Review` and
|
|
171
|
+
`Phase 4: Review`, and none of those spellings appear in it. It is now six
|
|
172
|
+
labels in both languages, and `smoke-phase-contract.sh` check 14 compares
|
|
173
|
+
every one against the contract and rejects a label above the last id, so the
|
|
174
|
+
one copy of the phase list that nothing derives is at least checked.
|
|
175
|
+
- A third spelling of a phase reference, `Phase N: Name`, which the first sweep
|
|
176
|
+
could not see: its rules matched `Phase 3 Dev` and the tracker tuple `3:Dev`,
|
|
177
|
+
and the colon form sits between them. It had left the canonical label table in
|
|
178
|
+
`skills/shared/core/multi-agent/SKILL.md` reading eight rows with six-phase
|
|
179
|
+
labels, and `/multi-agent:local` listing both a Phase 4 Review and a Phase 4
|
|
180
|
+
Commit. Fifteen files, corrected by name match rather than by number.
|
|
181
|
+
- The golden-task fixtures are named for the phases that produce them, so they
|
|
182
|
+
moved too: `phase-2-plan.json` -> `phase-1-plan.json`, `phase-4-review.json`
|
|
183
|
+
-> `phase-3-review.json`, `phase-4-triage.json` -> `phase-3-triage.json`.
|
|
184
|
+
- A doc sweep of 166 files in the repo and 68 in the source tree, none of which
|
|
185
|
+
the eight-phase plan had listed. Release history is deliberately excluded:
|
|
186
|
+
`CHANGELOG`, the `ROADMAP` "Previous Release" sections and
|
|
187
|
+
`docs/token-budget-history.md` keep the numbers their versions shipped with,
|
|
188
|
+
and four ADRs carry a pointer to ADR-0014 instead of being rewritten, because
|
|
189
|
+
an ADR records what was decided rather than what is true today.
|
|
190
|
+
|
|
191
|
+
### Removed
|
|
192
|
+
|
|
193
|
+
- **The engagement page** (`src/app/_nisan`, its API routes and its admin
|
|
194
|
+
panel) on the website. The 16 RSVP rows were exported before anything was
|
|
195
|
+
deleted and **the `rsvp_entries` table is kept** - removing code does not
|
|
196
|
+
remove data, and dropping the table is a separate decision.
|
|
197
|
+
|
|
198
|
+
---
|
|
199
|
+
|
|
200
|
+
## [18.0.0] - 2026-09-17
|
|
201
|
+
|
|
202
|
+
Major, for two behaviour changes rather than a renamed command: a run's state now
|
|
203
|
+
has ONE canonical directory, and four publish paths refuse to send when the leak
|
|
204
|
+
gate cannot be loaded.
|
|
205
|
+
|
|
206
|
+
### Added
|
|
207
|
+
|
|
208
|
+
- **`multi-agent-pipeline verify`** answers "is this install the thing that was
|
|
209
|
+
published". The install is a copy, and from the moment `install.js` writes it
|
|
210
|
+
the two halves drift independently: an edit in the installed tree is behaviour
|
|
211
|
+
with no source, and a file the installer skipped is a script the docs describe
|
|
212
|
+
and nobody has. `manifest.json` (SHA-256 per shipped file, version, source
|
|
213
|
+
commit) is built at pack time by `prepack` and never committed - a manifest in
|
|
214
|
+
git is stale one commit after it is written, and a stale manifest reports
|
|
215
|
+
honest edits as tampering.
|
|
216
|
+
|
|
217
|
+
`commands/` is compared by presence, not bytes, because `install.js` rewrites
|
|
218
|
+
each description into the user's `outputLanguage`: measured here, all 57
|
|
219
|
+
command files differ and 56 of them differ by nothing else. Dev-only files are
|
|
220
|
+
excluded, or 252 smokes and linters read as "the installer skipped this".
|
|
221
|
+
|
|
222
|
+
What a green result proves is stated in the output's own reference: the bytes
|
|
223
|
+
match what the publisher recorded. Not who published them - the manifest, the
|
|
224
|
+
signature and the verifier travel in the same tarball, so provenance belongs
|
|
225
|
+
to npm's integrity field. Signing is optional, and an unverifiable signature
|
|
226
|
+
says so rather than claiming valid.
|
|
227
|
+
|
|
228
|
+
- **`cost-analyze.mjs`** - projection, anomaly, burn and diff. `cost-budget-check`
|
|
229
|
+
watches one run against one ceiling, which is blind to both ways a budget
|
|
230
|
+
actually empties: a drift no single run trips, and one session that burns a
|
|
231
|
+
week in an hour while every run stays under its cap.
|
|
232
|
+
|
|
233
|
+
The series is not where the schema says it is. `tracker-state.json` carries
|
|
234
|
+
per-phase token fields that are written only when a phase reports them; of 98
|
|
235
|
+
trackers on a working machine, zero carry any. The dense series is the host's
|
|
236
|
+
own transcripts, so that is what this reads - and the consequences are printed
|
|
237
|
+
rather than buried: it covers everything Claude Code did on the machine, the
|
|
238
|
+
figures are LIST-price estimates rather than a bill, and a host with no
|
|
239
|
+
transcripts reports UNMEASURED instead of zero.
|
|
240
|
+
|
|
241
|
+
Anomalies use the median and the MAD, because the expensive session the check
|
|
242
|
+
exists to find is the observation that inflates a mean and a standard
|
|
243
|
+
deviation - it hides inside the statistic measured against it.
|
|
244
|
+
|
|
245
|
+
- **`install --unattended`** writes a documented permission profile, and prints
|
|
246
|
+
it with a reason per line before writing. autopilot passes
|
|
247
|
+
`--permission-prompts none`, which stops Claude Code asking and grants
|
|
248
|
+
nothing; on a fresh machine the run stops at the first tool call with no
|
|
249
|
+
prompt for anyone to answer, which is indistinguishable from an empty queue. A
|
|
250
|
+
default install still writes no permissions at all.
|
|
251
|
+
|
|
252
|
+
- **`doctor --profile=server`** adds four checks that only matter when nobody is
|
|
253
|
+
at the keyboard: the unattended contract, the permission posture, the
|
|
254
|
+
scheduler, and the keychain. The default run is unchanged - same 17 checks,
|
|
255
|
+
same verdict - because a laptop told it fails a server check learns to ignore
|
|
256
|
+
doctor.
|
|
257
|
+
|
|
258
|
+
- **`docs/server-readiness.md`** sets up nothing and says so first. It covers the
|
|
259
|
+
permission posture, why the scheduler is a LaunchAgent rather than a
|
|
260
|
+
LaunchDaemon (the keychain is locked until login, and that failure arrives
|
|
261
|
+
much later wearing a 401), and which credential each phase needs.
|
|
262
|
+
|
|
263
|
+
- **`scorecard-snapshot.mjs --diff`** keeps what the scorecard said and reports
|
|
264
|
+
what moved. Deliberately no 0-100 score: the scorecard reports twelve measured
|
|
265
|
+
metrics AND four it refuses to measure, and one figure would hide both halves.
|
|
266
|
+
|
|
267
|
+
### Changed
|
|
268
|
+
|
|
269
|
+
- **A run's state has one canonical directory.** `{project}/{taskId}/` is
|
|
270
|
+
canonical and the flat `{taskId}/` is read for compatibility; every reader
|
|
271
|
+
resolves through `lib/run-paths.sh` / `scripts/_run-paths.mjs`, so a run that
|
|
272
|
+
exists in both layouts is counted once.
|
|
273
|
+
|
|
274
|
+
- **Five publish paths refuse rather than send when the leak gate is missing.**
|
|
275
|
+
Jira comments and descriptions, PR review bodies and inline comments, the
|
|
276
|
+
GitHub issue progress comment, the issue-CREATION path, and Confluence page
|
|
277
|
+
create/update all run their outbound text through `lib/outbound-gate.mjs`
|
|
278
|
+
first. A missing gate file refuses; opening the gate because the gate is not
|
|
279
|
+
there would be the one failure mode that matters. The count is part of the
|
|
280
|
+
gate: a sixth publisher cannot ship without the smoke's list naming it.
|
|
281
|
+
|
|
282
|
+
- **`set -e` on the four destructive scripts, and deliberately not on the two
|
|
283
|
+
collectors.** `gc-abandoned`, `gc-worktrees`, `purge` and `worktree-finalize`
|
|
284
|
+
delete things, so the command after an unnoticed failure is the dangerous one.
|
|
285
|
+
`pre-push-check` runs the gates and counts failures, and `pre-commit-check` is
|
|
286
|
+
a hook built out of greps that are supposed to find nothing - under `-e` the
|
|
287
|
+
first clean detector would end the scan and report "no secrets" for a file it
|
|
288
|
+
never finished reading.
|
|
289
|
+
|
|
290
|
+
### Fixed
|
|
291
|
+
|
|
292
|
+
- **A script that dies now prints one line instead of a stack dump, and releases
|
|
293
|
+
what it held.** `lib/fatal.mjs` catches the sync throw, the rejection nobody
|
|
294
|
+
awaited and the throw from inside a callback; the last two are invisible to a
|
|
295
|
+
try/catch around `main()`. `write-state` releases its advisory lock on the way
|
|
296
|
+
out, so the next writer never has to judge a lock on age alone. EPIPE is
|
|
297
|
+
deliberately not fatal: `runs-index.mjs --json | head` is the ordinary way to
|
|
298
|
+
read a large output.
|
|
299
|
+
|
|
300
|
+
- **Build junk no longer reaches an install.** A CI runner shipped 147 files
|
|
301
|
+
where every developer tree shipped 146, for four rounds, and the extra was
|
|
302
|
+
`__pycache__/*.pyc` - untracked, so no diff of the source could show it, and
|
|
303
|
+
copied verbatim into all three install trees. `copyDir` now filters
|
|
304
|
+
`__pycache__`, `*.pyc` and `.DS_Store` on every path.
|
|
305
|
+
|
|
306
|
+
- **`jq` is required rather than optional on the nine paths that publish or
|
|
307
|
+
decide.** A missing `jq` renders as empty DATA and the work carries on with
|
|
308
|
+
it; those nine now exit 3.
|
|
309
|
+
|
|
310
|
+
- **The autopilot runner survives a month unwatched.** `runner.log` is truncated
|
|
311
|
+
in place past 5MB (renaming it leaves launchd's `O_APPEND` descriptor writing
|
|
312
|
+
into the renamed file while the new one stays empty - the rotation that looks
|
|
313
|
+
right and silently stops logging), three consecutive empty attempts stop new
|
|
314
|
+
work being taken, and one JSON line per tick goes to `ticks.jsonl`.
|
|
315
|
+
|
|
316
|
+
- **`github-ssh-setup.sh` no longer waits forever on a headless machine**, and
|
|
317
|
+
`write-state.mjs` gained `--if-rev=<n>` compare-and-swap so a second writer
|
|
318
|
+
cannot silently overwrite the first.
|
|
319
|
+
|
|
320
|
+
- **A direct-run guard that compared `import.meta.url` to `file://${argv[1]}`**
|
|
321
|
+
is false whenever argv[1] is not already resolved - a `/var` path that
|
|
322
|
+
resolves to `/private/var`, or the symlink `install --link` writes. The script
|
|
323
|
+
then does nothing at all, silently.
|
|
324
|
+
|
|
325
|
+
---
|
|
326
|
+
|
|
17
327
|
## [17.6.0] - 2026-09-15
|
|
18
328
|
|
|
19
329
|
### Added
|
package/README.md
CHANGED
|
@@ -8,16 +8,16 @@
|
|
|
8
8
|
|
|
9
9
|
🇹🇷 Türkçe: [README.tr.md](./README.tr.md)
|
|
10
10
|
|
|
11
|
-
|
|
11
|
+
A 6-phase AI development pipeline for **Claude Code**, **Copilot CLI** and **Codex CLI**. Drives a Jira issue or GitHub URL to a merged PR in one command - analysis → plan → TDD → review → test → commit → PR - with multi-repo orchestration, a plan-approval gate, CLI-aware parallel review, and store-compliance checks. Component and Figma-to-code work is dispatched to the per-stack marketplace plugins (iOS/SwiftUI, Android/Compose) rather than bundled, so component skills live in one place.
|
|
12
12
|
|
|
13
13
|
Runs natively on Claude Code, Copilot CLI and Codex CLI. macOS only. Zero runtime dependencies.
|
|
14
14
|
|
|
15
|
-
📐 **[Architecture diagrams](./docs/architecture.md)** - the
|
|
15
|
+
📐 **[Architecture diagrams](./docs/architecture.md)** - the 6-phase flow, operating modes, review/triage, Figma subphases, component layout. **[Ecosystem diagram](./docs/ecosystem.md)** - how this repo, the `multi-agent-plugins` marketplace and `multi-agent-toolkit-mcp` compose.
|
|
16
16
|
|
|
17
17
|
### Prerequisites
|
|
18
18
|
|
|
19
19
|
- **Node.js >= 20.11** - required; the pipeline's own tooling runs on it.
|
|
20
|
-
- **`jq`** - optional
|
|
20
|
+
- **`jq`** - required for nine paths, optional for the rest. 82 shell files call it. The nine that publish or decide - the autopilot queue, Jira comments, PR reviews, issue updates, the plan file, both Figma fetchers, log search and Jira auth - now refuse with exit 3 rather than run, because a missing `jq` renders as empty DATA and the work carries on with it. Everywhere else it still degrades. The install prints a note when it is missing.
|
|
21
21
|
- **`gh`** - for GitHub issue and PR work. Its built-in `--jq` is independent of the `jq` binary.
|
|
22
22
|
|
|
23
23
|
## Quick Start
|
|
@@ -33,6 +33,41 @@ npx @mmerterden/multi-agent-pipeline install --all # all three
|
|
|
33
33
|
/multi-agent:setup # keychain token scan + git identity + default stack
|
|
34
34
|
```
|
|
35
35
|
|
|
36
|
+
### No `npx` on that machine?
|
|
37
|
+
|
|
38
|
+
`npx` ships with npm, and npm ships with Node - so "npx: command not found" almost
|
|
39
|
+
always means Node is missing from that shell, not that anything is wrong with the
|
|
40
|
+
package. Check first, then pick the row that matches:
|
|
41
|
+
|
|
42
|
+
```bash
|
|
43
|
+
node -v; npm -v; command -v node npm npx
|
|
44
|
+
```
|
|
45
|
+
|
|
46
|
+
| What you see | What to do |
|
|
47
|
+
|---|---|
|
|
48
|
+
| nothing at all | Install Node >= 20.11: `brew install node`, or the LTS installer from nodejs.org |
|
|
49
|
+
| `node` works, `npx` does not | `npm i -g @mmerterden/multi-agent-pipeline` then `multi-agent-pipeline install --claude` |
|
|
50
|
+
| nvm is installed but the shell does not see it | `source ~/.nvm/nvm.sh && nvm use --lts`, or just open a new terminal |
|
|
51
|
+
| npm is ancient (< 5.2, which predates npx) | `npm i -g npm@latest`, or use the global-install row above |
|
|
52
|
+
|
|
53
|
+
And the path that needs neither `npx` nor a global install - clone and run the
|
|
54
|
+
installer directly:
|
|
55
|
+
|
|
56
|
+
```bash
|
|
57
|
+
git clone https://github.com/mmerterden/multi-agent-pipeline.git
|
|
58
|
+
cd multi-agent-pipeline
|
|
59
|
+
node index.js install --claude # add --dry-run first to see what it would write
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
`npm exec @mmerterden/multi-agent-pipeline install --claude` also works on any npm
|
|
63
|
+
7+ without `npx` on `PATH`.
|
|
64
|
+
|
|
65
|
+
**One error that is not an npx problem.** This package declares `os: ["darwin"]`,
|
|
66
|
+
so npm refuses to install it anywhere else and says `npm ERR! notsup Unsupported
|
|
67
|
+
platform`. That is deliberate ([ADR-0012](./docs/adr/0012-macos-only.md)), not a
|
|
68
|
+
missing tool: every credential read shells `security`, every iOS build
|
|
69
|
+
`xcodebuild`, every piece of visual evidence `simctl`.
|
|
70
|
+
|
|
36
71
|
Tool flags combine (`--claude --codex`). With no tool flag at all, the installer targets Claude Code only. Other flags: `--dry-run` (show what would be written, write nothing), `--platform=ios|android|all` (skip the stack skills you do not need), `--link` (symlink instead of copy, for local development).
|
|
37
72
|
|
|
38
73
|
Run a task - the input type is auto-detected:
|
|
@@ -56,22 +91,45 @@ Update later with `/multi-agent:update`. Uninstall (tokens preserved) with `npx
|
|
|
56
91
|
|
|
57
92
|
## How it works
|
|
58
93
|
|
|
59
|
-
One command runs up to
|
|
94
|
+
One command runs up to 6 phases, with a gate between the risky ones. Phase 0
|
|
60
95
|
asks two questions that decide the shape of the rest - how deep the run goes
|
|
61
96
|
(Full or Short) and where the branch lives (a worktree or your current
|
|
62
97
|
checkout):
|
|
63
98
|
|
|
64
99
|
- **0 · Init** - parse the input (Jira id / GitHub URL / free text), pick account + repo(s), fetch the issue, run a maturity check.
|
|
65
|
-
- **1 ·
|
|
66
|
-
- **2 ·
|
|
67
|
-
- **3 ·
|
|
68
|
-
- **4 ·
|
|
69
|
-
- **5 ·
|
|
70
|
-
- **6 · Commit/PR** - conventional commit, push (must succeed), open a PR (`Ref: #N`, never auto-close).
|
|
71
|
-
- **7 · Report** - technical summary + a Jira comment with test scenarios, posted through the channels layer.
|
|
100
|
+
- **1 · Plan** - detect the stack, scan the codebase and write the analysis document, then break it into tasks with file-level targets and **stop for your approval** before touching code. Analysis and planning were two phases until 19.0.0; the depth picker always skipped them together, because they are one decision. Codebase scanning runs on the explorer persona (Sonnet).
|
|
101
|
+
- **2 · Dev** - TDD: failing test → code → green, following the repo's style + the active stack skills. The phase ends at its own gate: build, lint, tests and a secret scan, run **once**. Review used to build again, and nothing consumed the difference.
|
|
102
|
+
- **3 · Review** - a **CLI-aware parallel review** against the logs Dev produced - Claude Code runs 3 models (Fable + Opus + Sonnet), Copilot CLI runs 3 (GPT-5.4 + Opus + Sonnet) - then a **Fable triage** keeps only actionable findings; blockers loop back to Phase 2. The optional user test lives here, keeping its waiting state.
|
|
103
|
+
- **4 · Commit/PR** - conventional commit, push (must succeed), open a PR (`Ref: #N`, never auto-close).
|
|
104
|
+
- **5 · Report** - technical summary + a Jira comment with test scenarios, posted through the channels layer. This is the one step autopilot still pauses at, in every mode.
|
|
72
105
|
|
|
73
106
|
`/multi-agent:analysis` runs its own shorter chain and, since v16.12.0, reviews what it wrote before publishing it: the draft goes through the same three-reviewer set and triage as a code diff, a blocking finding returns it to synthesis with dispatch closed, and the gaps that survive are either searched, asked about, or recorded with an owner. It used to publish behind a structural validator alone.
|
|
74
107
|
|
|
108
|
+
### 19.0.0: six phases, and a gate for the number
|
|
109
|
+
|
|
110
|
+
Two of the six phases were doing the same work twice. Dev built the project
|
|
111
|
+
and tee'd a log; Review opened by building it again. Analysis and Planning were
|
|
112
|
+
already one decision - the depth picker skipped them together and the state
|
|
113
|
+
schema described them as one unit. Six phases now, one build per run.
|
|
114
|
+
|
|
115
|
+
The other half of the change is that the count is finally guarded.
|
|
116
|
+
`smoke-phase-contract.sh` derives it from `pipeline/schemas/phases.json` and
|
|
117
|
+
holds every other copy to it: the generator's output, the token budget, the
|
|
118
|
+
state-schema bounds, the progress fractions in sample output, and the named
|
|
119
|
+
thresholds that used to be literals scattered across scripts. The phase count
|
|
120
|
+
appeared in 91 places across 40 files with nothing checking any of them, while
|
|
121
|
+
the command count, the jq count and the persona count all had gates.
|
|
122
|
+
|
|
123
|
+
Reasoning, mapping and rejected alternatives:
|
|
124
|
+
[ADR-0014](./docs/adr/0014-six-phase-consolidation.md).
|
|
125
|
+
|
|
126
|
+
### 18.0.0: one state directory, and a way to ask whether your install is real
|
|
127
|
+
|
|
128
|
+
- **`multi-agent-pipeline verify`.** The install is a copy, and from the moment it is written the two halves drift independently: an edit in the installed tree is behaviour with no source, and a file the installer skipped is a script the docs describe and nobody has. `verify` compares both against a manifest built at pack time. What a green result proves is stated plainly - the bytes match what the publisher recorded, not who published them.
|
|
129
|
+
- **Cost, past the single run.** The per-task ceiling cannot see the two ways a budget actually empties: a drift that trips nothing, and one session that burns a week in an hour while every run stays under its cap. `cost-analyze` projects, finds days out of family by median absolute deviation, and reports acceleration - as LIST-price estimates, which it says on every run rather than in a footnote.
|
|
130
|
+
- **A run's state lives in one place.** `{project}/{taskId}/` is canonical, the flat layout is still read, and a run that exists in both is counted once.
|
|
131
|
+
- **Server readiness, entirely opt-in.** `doctor --profile=server` adds four checks that only matter when nobody is at the keyboard, `install --unattended` writes a permission profile after printing it, and the autopilot runner now survives a month unwatched. The default install writes no permissions and the default doctor run is unchanged, because a laptop told it fails a server check learns to ignore doctor.
|
|
132
|
+
|
|
75
133
|
### Your package manager, your hooks, your MCP surface
|
|
76
134
|
|
|
77
135
|
Three smaller things in 17.6.0, each closing a gap where the pipeline assumed instead of looking:
|
|
@@ -108,8 +166,8 @@ The discipline behind all of this - bounded loops, evidence gates, token-budgete
|
|
|
108
166
|
|
|
109
167
|
| Mode | Command | Flow |
|
|
110
168
|
| --------- | ------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
111
|
-
| Full | `/multi-agent "task"` | All
|
|
112
|
-
| Autopilot | `/multi-agent:autopilot "task"` |
|
|
169
|
+
| Full | `/multi-agent "task"` | All 6 phases, interactive |
|
|
170
|
+
| Autopilot | `/multi-agent:autopilot "task"` | 6 phases (interactive Test gate dropped), no confirmations |
|
|
113
171
|
| Local | `/multi-agent:local "task"` | Full pipeline minus the interactive Test gate, current branch (no worktree) |
|
|
114
172
|
| Depth | asked at Phase 0 Step 7.5 | Full (all phases) or Short (Dev → Review → Test → Commit → Report). Not a command name - `/multi-agent` and `:local` ask, both autopilot entries always run Full |
|
|
115
173
|
| Ship | `/multi-agent:resume-local` | Run the review→test→commit→report tail over local work |
|
|
@@ -165,7 +223,7 @@ The widget follows the answer rather than predicting it: Phase 0 is the only til
|
|
|
165
223
|
| `/multi-agent:review-jira` | Grade a Jira issue's readiness for development, comment the gaps |
|
|
166
224
|
| `/multi-agent:review-issue` | Same grading for a GitHub issue |
|
|
167
225
|
| `/multi-agent:review-analysis` | Review a written analysis document; findings cite the Locked rule they break |
|
|
168
|
-
| `/multi-agent:diff-explain` | Map a Phase
|
|
226
|
+
| `/multi-agent:diff-explain` | Map a Phase 3 triage finding back to the diff lines that caused it |
|
|
169
227
|
| `/multi-agent:refactor` | Best-practice extraction + bug hunt + derived-skill drift + toolkit MCP research → one plan |
|
|
170
228
|
| `/multi-agent:scan` | Skill security scan of local skill directories against a tiered pattern catalog |
|
|
171
229
|
| `/multi-agent:prune-prompts` | Zero-base review of the always-on instruction footprint; keep / trial / delete per rule |
|
|
@@ -301,17 +359,17 @@ This enables the matching plugin (+ the shared `ai-common` plugin) in the repo's
|
|
|
301
359
|
|
|
302
360
|
## Tool support
|
|
303
361
|
|
|
304
|
-
The pipeline runs natively on **Claude Code**, **Copilot CLI** and **Codex CLI** - all three install from the same `pipeline/` source and get the same
|
|
362
|
+
The pipeline runs natively on **Claude Code**, **Copilot CLI** and **Codex CLI** - all three install from the same `pipeline/` source and get the same 60 commands.
|
|
305
363
|
|
|
306
364
|
| Tool | Flag | What it installs |
|
|
307
365
|
| ----------- | -------------------- | ------------------------------------------------------------------------------------------------------ |
|
|
308
366
|
| Claude Code | `--claude` (default) | slash commands + skills + agents + three `PreToolUse` hooks (secret scan, agent-guard, read-size gate) |
|
|
309
|
-
| Copilot CLI | `--copilot` | instructions +
|
|
310
|
-
| Codex CLI | `--codex` | one router skill +
|
|
367
|
+
| Copilot CLI | `--copilot` | instructions + 60 sub-command skills + scripts |
|
|
368
|
+
| Codex CLI | `--codex` | one router skill + 60 specs as refs + 9 agent TOML + `AGENTS.md` block + `codex mcp add` |
|
|
311
369
|
|
|
312
370
|
Filter skills by stack with `--platform=ios\|android\|all`.
|
|
313
371
|
|
|
314
|
-
**Why Codex gets one skill and not
|
|
372
|
+
**Why Codex gets one skill and not 60.** Codex assembles every discovered skill's name
|
|
315
373
|
and description into a single prompt block and drops entries when it overflows, with no
|
|
316
374
|
error. Measured on 0.145: installing one plugin that declares 142 skills surfaced only
|
|
317
375
|
75 of them and evicted an unrelated user skill. So on Codex the pipeline ships a single
|
package/README.tr.md
CHANGED
|
@@ -8,11 +8,11 @@
|
|
|
8
8
|
|
|
9
9
|
🇬🇧 English: [README.md](./README.md)
|
|
10
10
|
|
|
11
|
-
**Claude Code**, **Copilot CLI** ve **Codex CLI** için
|
|
11
|
+
**Claude Code**, **Copilot CLI** ve **Codex CLI** için 6 fazlı bir AI geliştirme pipeline'ı. Bir Jira issue'sunu veya GitHub URL'sini tek komutla merge edilmiş bir PR'a dönüştürür - analiz → plan → TDD → review → test → commit → PR - çoklu-repo orkestrasyonu, bir plan-onay kapısı, CLI-farkında paralel review ve store-uyumluluk kontrolleriyle birlikte. Component ve Figma-to-code işleri paket içine gömülmek yerine stack başına marketplace plugin'lerine (iOS/SwiftUI, Android/Compose) devredilir, böylece component skill'leri tek bir yerde yaşar.
|
|
12
12
|
|
|
13
13
|
Claude Code, Copilot CLI ve Codex CLI üzerinde native çalışır. Yalnızca macOS. Sıfır runtime dependency.
|
|
14
14
|
|
|
15
|
-
📐 **[Mimari diyagramları](./docs/architecture.md)** -
|
|
15
|
+
📐 **[Mimari diyagramları](./docs/architecture.md)** - 6 faz akışı, çalışma modları, review/triage, Figma subphase'leri, component yapısı. **[Ekosistem diyagramı](./docs/ecosystem.md)** - bu repo, `multi-agent-plugins` marketplace'i ve `multi-agent-toolkit-mcp`'nin nasıl bir araya geldiği.
|
|
16
16
|
|
|
17
17
|
### Önkoşullar
|
|
18
18
|
|
|
@@ -33,6 +33,40 @@ npx @mmerterden/multi-agent-pipeline install --all # üçü birden
|
|
|
33
33
|
/multi-agent:setup # keychain token taraması + git kimliği + varsayılan stack
|
|
34
34
|
```
|
|
35
35
|
|
|
36
|
+
### O makinede `npx` yoksa
|
|
37
|
+
|
|
38
|
+
`npx` npm ile, npm de Node ile gelir - yani "npx: command not found" neredeyse her
|
|
39
|
+
zaman o kabukta Node olmadığı anlamına gelir, pakette bir sorun olduğu değil. Önce
|
|
40
|
+
bak, sonra sana uyan satırı uygula:
|
|
41
|
+
|
|
42
|
+
```bash
|
|
43
|
+
node -v; npm -v; command -v node npm npx
|
|
44
|
+
```
|
|
45
|
+
|
|
46
|
+
| Gördüğün | Yapılacak |
|
|
47
|
+
|---|---|
|
|
48
|
+
| hiçbiri yok | Node >= 20.11 kur: `brew install node` ya da nodejs.org'dan LTS installer |
|
|
49
|
+
| `node` çalışıyor, `npx` çalışmıyor | `npm i -g @mmerterden/multi-agent-pipeline` sonra `multi-agent-pipeline install --claude` |
|
|
50
|
+
| nvm kurulu ama kabuk görmüyor | `source ~/.nvm/nvm.sh && nvm use --lts`, ya da yeni bir terminal aç |
|
|
51
|
+
| npm çok eski (< 5.2, npx'ten önceki sürümler) | `npm i -g npm@latest`, ya da üstteki global kurulum satırı |
|
|
52
|
+
|
|
53
|
+
Ne `npx` ne de global kurulum isteyen yol - klonla ve installer'ı doğrudan çalıştır:
|
|
54
|
+
|
|
55
|
+
```bash
|
|
56
|
+
git clone https://github.com/mmerterden/multi-agent-pipeline.git
|
|
57
|
+
cd multi-agent-pipeline
|
|
58
|
+
node index.js install --claude # önce --dry-run ile ne yazacağını görebilirsin
|
|
59
|
+
```
|
|
60
|
+
|
|
61
|
+
`npm exec @mmerterden/multi-agent-pipeline install --claude` de, `PATH`'te `npx`
|
|
62
|
+
olmayan her npm 7+ üzerinde çalışır.
|
|
63
|
+
|
|
64
|
+
**npx sorunu olmayan bir hata.** Bu paket `os: ["darwin"]` beyan eder; npm başka
|
|
65
|
+
hiçbir yerde kurmaz ve `npm ERR! notsup Unsupported platform` der. Bu kasıtlıdır
|
|
66
|
+
([ADR-0012](./docs/adr/0012-macos-only.md)), eksik bir araç değil: her credential
|
|
67
|
+
okuması `security`, her iOS build'i `xcodebuild`, her görsel kanıt `simctl`
|
|
68
|
+
çağırıyor.
|
|
69
|
+
|
|
36
70
|
Tool flag'leri birleştirilebilir (`--claude --codex`). Hiç tool flag'i verilmezse installer sadece Claude Code'u hedefler. Diğer flag'ler: `--dry-run` (ne yazılacağını gösterir, hiçbir şey yazmaz), `--platform=ios|android|all` (ihtiyacın olmayan stack skill'lerini atlar), `--link` (kopyalamak yerine symlink, lokal geliştirme için).
|
|
37
71
|
|
|
38
72
|
Bir görev çalıştır - girdi tipi otomatik algılanır:
|
|
@@ -56,22 +90,27 @@ Sonra `/multi-agent:update` ile güncelle. Kaldırmak için (tokenlar korunur) `
|
|
|
56
90
|
|
|
57
91
|
## Nasıl çalışır
|
|
58
92
|
|
|
59
|
-
Tek komut en fazla
|
|
93
|
+
Tek komut en fazla 6 fazı çalıştırır, riskli olanlar arasında bir kapı ile. Faz
|
|
60
94
|
0 geri kalanın şeklini belirleyen iki soru sorar: koşu ne kadar derin olacak
|
|
61
95
|
(Tam mı Kısa mı) ve branch nerede yaşayacak (worktree mi, mevcut checkout'un
|
|
62
96
|
mu):
|
|
63
97
|
|
|
64
98
|
- **0 · Init** - girdiyi ayrıştır (Jira id / GitHub URL / serbest metin), hesap + repo(lar) seç, issue'yu çek, maturity kontrolü yap.
|
|
65
|
-
- **1 ·
|
|
66
|
-
- **2 ·
|
|
67
|
-
- **3 · Dev** -
|
|
68
|
-
- **4 ·
|
|
69
|
-
- **5 ·
|
|
70
|
-
- **6 · Commit/PR** - conventional commit, push (başarılı olmalı), bir PR aç (`Ref: #N`, asla otomatik kapatma).
|
|
71
|
-
- **7 · Report** - teknik özet + test senaryolarıyla bir Jira yorumu, channels katmanından gönderilir.
|
|
99
|
+
- **1 · Plan** - stack'i tespit et, codebase'i tara ve analiz dokümanını yaz; sonra onu dosya seviyesinde hedefleri olan görevlere böl ve koda dokunmadan önce **onayın için dur**. Analiz ve planlama 19.0.0'a kadar iki ayrı fazdı; derinlik seçici ikisini hep birlikte atlıyordu, çünkü tek bir karar. Codebase taraması explorer persona'sı üzerinde koşar (Sonnet).
|
|
100
|
+
- **2 · Dev** - TDD: başarısız test → kod → yeşil, repo'nun stiline + aktif stack skill'lerine uyarak. Faz kendi kapısında biter: build, lint, test ve sır taraması, **bir kez** koşar. Review eskiden ikinci kez build ediyordu ve aradaki farkı kimse okumuyordu.
|
|
101
|
+
- **3 · Review** - Dev'in ürettiği log'lara karşı **CLI-farkında paralel review** - Claude Code 3 model çalıştırır (Fable + Opus + Sonnet), Copilot CLI 3 (GPT-5.4 + Opus + Sonnet) - ve bir **Fable triage** sadece aksiyon alınabilir bulguları tutar; blocker'lar Phase 2'ye geri döner. Opsiyonel kullanıcı testi burada, bekleme durumunu koruyarak.
|
|
102
|
+
- **4 · Commit/PR** - conventional commit, push (başarılı olmalı), bir PR aç (`Ref: #N`, asla otomatik kapatma).
|
|
103
|
+
- **5 · Report** - teknik özet + test senaryolarıyla bir Jira yorumu, channels katmanından gönderilir. Autopilot'un her modda hâlâ durduğu tek adım bu.
|
|
72
104
|
|
|
73
105
|
`/multi-agent:analysis` kendi kısa zincirini koşar ve v16.12.0'dan beri yazdığını yayınlamadan önce review ediyor: taslak, bir kod diff'iyle aynı üç-reviewer setinden ve triyajdan geçiyor, bloklayıcı bulgu dokümanı sentez fazına geri gönderip dispatch'i kapatıyor, hayatta kalan boşluklar ya aranıyor ya sana soruluyor ya da sahibiyle birlikte kayda giriyor. Önceden yalnızca yapısal bir validator'ın arkasından yayınlıyordu.
|
|
74
106
|
|
|
107
|
+
### 18.0.0: tek bir durum dizini, ve kurulumun gerçekten o kurulum olup olmadığını sorma yolu
|
|
108
|
+
|
|
109
|
+
- **`multi-agent-pipeline verify`.** Kurulum bir kopyadır ve yazıldığı andan itibaren iki taraf birbirinden bağımsız kayar: kurulu ağaçtaki bir düzenleme kaynağı olmayan bir davranıştır, kurulumun atladığı dosya ise dokümanın anlattığı ama kimsede olmayan bir script. `verify` ikisini de paketleme anında üretilen bir manifest'e karşı karşılaştırır. Yeşil sonucun ne kanıtladığı açıkça yazılıdır: baytlar yayıncının kaydettiğiyle aynıdır, kimin yayınladığı değil.
|
|
110
|
+
- **Maliyet, tek koşunun ötesinde.** Görev başına tavan, bütçeyi gerçekten bitiren iki şeyi göremez: hiçbir koşuyu aşmayan kayma, ve her koşu tavanın altında kalırken bir saatte bir haftayı yakan tek oturum. `cost-analyze` projeksiyon yapar, medyan mutlak sapmayla aileden ayrılan günü bulur ve ivmeyi raporlar - liste fiyatından tahmin olarak, ve bunu dipnotta değil her koşuda söyler.
|
|
111
|
+
- **Bir koşunun durumu tek yerde.** `{project}/{taskId}/` kanonik, düz yerleşim hâlâ okunuyor, ve iki yerde birden duran koşu bir kez sayılıyor.
|
|
112
|
+
- **Sunucu hazırlığı, tamamen opt-in.** `doctor --profile=server` yalnızca klavyede kimse yokken anlamı olan dört kontrol ekler, `install --unattended` izin profilini yazmadan önce gösterir, autopilot runner artık bir ay gözetimsiz ayakta kalır. Varsayılan kurulum hiçbir izin yazmaz ve varsayılan doctor koşusu değişmez: sunucu kontrolünden kaldığı söylenen bir dizüstü, doctor'ı görmezden gelmeyi öğrenir.
|
|
113
|
+
|
|
75
114
|
### Paket yöneticisi, hook'lar ve MCP yüzeyi
|
|
76
115
|
|
|
77
116
|
17.6.0'da üç küçük iş; üçü de pipeline'ın bakmak yerine varsaydığı bir yeri kapatıyor:
|
|
@@ -108,8 +147,8 @@ Bunun arkasındaki disiplin - sınırlı loop'lar, kanıt kapıları, token-büt
|
|
|
108
147
|
|
|
109
148
|
| Mod | Komut | Akış |
|
|
110
149
|
| --------- | ------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
111
|
-
| Full | `/multi-agent "task"` | Tüm
|
|
112
|
-
| Autopilot | `/multi-agent:autopilot "task"` |
|
|
150
|
+
| Full | `/multi-agent "task"` | Tüm 6 faz, interaktif |
|
|
151
|
+
| Autopilot | `/multi-agent:autopilot "task"` | 6 faz (interaktif Test kapısı atlanır), onaysız |
|
|
113
152
|
| Local | `/multi-agent:local "task"` | İnteraktif Test kapısı hariç tam pipeline, mevcut branch (worktree yok) |
|
|
114
153
|
| Derinlik | Faz 0 Adım 7.5'te sorulur | Full (tüm fazlar) veya Short (Dev → Review → Test → Commit → Report). Komut adı değil - `/multi-agent` ve `:local` sorar, iki autopilot girişi de her zaman Full koşar |
|
|
115
154
|
| Ship | `/multi-agent:resume-local` | Lokal iş üzerinde review→test→commit→report kuyruğunu çalıştır |
|
|
@@ -302,17 +341,17 @@ Bu, ilgili plugin'i (+ ortak `ai-common` plugin'ini) repo'nun `.claude/settings.
|
|
|
302
341
|
|
|
303
342
|
## Araç desteği
|
|
304
343
|
|
|
305
|
-
Pipeline **Claude Code**, **Copilot CLI** ve **Codex CLI** üzerinde native çalışır - üçü de aynı `pipeline/` kaynağından kurulur ve aynı
|
|
344
|
+
Pipeline **Claude Code**, **Copilot CLI** ve **Codex CLI** üzerinde native çalışır - üçü de aynı `pipeline/` kaynağından kurulur ve aynı 60 komutu alır.
|
|
306
345
|
|
|
307
346
|
| Araç | Bayrak | Ne kurar |
|
|
308
347
|
| ----------- | ----------------------- | ---------------------------------------------------------------------------------------------------------------- |
|
|
309
348
|
| Claude Code | `--claude` (varsayılan) | slash komutları + skill'ler + agent'lar + üç `PreToolUse` hook'u (secret scan, agent-guard, okuma-boyutu geçidi) |
|
|
310
|
-
| Copilot CLI | `--copilot` | talimatlar +
|
|
311
|
-
| Codex CLI | `--codex` | bir router skill + ref olarak
|
|
349
|
+
| Copilot CLI | `--copilot` | talimatlar + 60 alt-komut skill'i + script'ler |
|
|
350
|
+
| Codex CLI | `--codex` | bir router skill + ref olarak 60 spec + 9 agent TOML + `AGENTS.md` bloğu + `codex mcp add` |
|
|
312
351
|
|
|
313
352
|
Skill'leri stack'e göre filtrele: `--platform=ios\|android\|all`.
|
|
314
353
|
|
|
315
|
-
**Codex neden
|
|
354
|
+
**Codex neden 60 değil de tek bir skill alıyor.** Codex, keşfettiği her skill'in adını
|
|
316
355
|
ve açıklamasını tek bir prompt bloğuna toplar ve blok taştığında girdileri hatasızca
|
|
317
356
|
düşürür. 0.145 üzerinde ölçüldü: 142 skill deklare eden bir plugin kurulduğunda sadece
|
|
318
357
|
75'i yüzeye çıktı ve alakasız bir kullanıcı skill'i tahliye edildi. Bu yüzden Codex'te
|
|
@@ -1,6 +1,7 @@
|
|
|
1
1
|
# 2. `instructionDriven` flag as explicit pipeline fork
|
|
2
2
|
|
|
3
3
|
**Status:** Accepted · 2025
|
|
4
|
+
> **Phase numbers below are the eight-phase ones.** [ADR-0014](./0014-six-phase-consolidation.md) renumbered the contract in v19.0.0 (Phase 6 Commit is now Phase 4, Phase 7 Report is now Phase 5). The decision this ADR records is unchanged; only the labels moved, and they are left as written because an ADR records what was decided.
|
|
4
5
|
|
|
5
6
|
## Context
|
|
6
7
|
|