@mmerterden/multi-agent-pipeline 19.1.4 → 20.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +123 -0
- package/README.md +19 -36
- package/README.tr.md +18 -35
- package/SECURITY.md +3 -3
- package/docs/adr/0002-instruction-driven-flag.md +6 -5
- package/docs/adr/0005-lazy-phase-docs.md +2 -2
- package/docs/adr/0008-installer-modularization-and-secret-leak-defense.md +1 -0
- package/docs/adr/0009-claude-stack-skills-plugin-only.md +1 -1
- package/docs/adr/0010-own-code-graph.md +5 -4
- package/docs/adr/0011-dormant-ci.md +10 -1
- package/docs/adr/0012-macos-only.md +2 -2
- package/docs/adr/0013-lsp-code-intelligence.md +2 -2
- package/docs/adr/0014-six-phase-consolidation.md +9 -9
- package/docs/adr/0015-one-pipeline-no-depth-answer.md +83 -0
- package/docs/adr/0016-the-run-shape-is-asked-not-typed.md +69 -0
- package/docs/adr/README.md +18 -16
- package/docs/architecture.md +2 -2
- package/docs/ecosystem.md +5 -5
- package/docs/facts.json +7 -9
- package/docs/features.md +4 -5
- package/docs/token-budget-history.md +1 -1
- package/install/_codex-agents.mjs +1 -1
- package/install/_common.mjs +9 -1
- package/install/templates/copilot-instructions.md +7 -16
- package/manifest.json +133 -129
- package/package.json +1 -1
- package/pipeline/agents/code-reviewer.md +2 -2
- package/pipeline/agents/dev-critic.md +5 -5
- package/pipeline/agents/security-auditor.md +80 -72
- package/pipeline/commands/figma-to-swiftui.md +1 -1
- package/pipeline/commands/multi-agent/SKILL.md +7 -9
- package/pipeline/commands/multi-agent/analysis/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/analysis-jira/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/analysis-resolve/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/autopilot/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/autopilot-on/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/autopilot-status/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/build-optimize/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/channels/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/create-jira/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/design-check/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/diff-explain/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/forget/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +4 -2
- package/pipeline/commands/multi-agent/help/SKILL.md +23 -27
- package/pipeline/commands/multi-agent/ios-coding-standard/SKILL.md +5 -4
- package/pipeline/commands/multi-agent/issue/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/jira/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/language/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/prune-logs/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/purge/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/resume/SKILL.md +177 -48
- package/pipeline/commands/multi-agent/save/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/scan/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/security-review/SKILL.md +52 -0
- package/pipeline/commands/multi-agent/stack/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/sync/SKILL.md +7 -8
- package/pipeline/commands/multi-agent/test-screenshots/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/uninstall/SKILL.md +2 -0
- package/pipeline/lib/repo-hygiene.sh +1 -1
- package/pipeline/multi-agent-refs/analysis/render.md +1 -1
- package/pipeline/multi-agent-refs/analysis/resolve.md +1 -1
- package/pipeline/multi-agent-refs/analysis/synthesis.md +1 -1
- package/pipeline/multi-agent-refs/analysis-template.md +1 -1
- package/pipeline/multi-agent-refs/component-dispatch.md +5 -13
- package/pipeline/multi-agent-refs/cross-cli-contract.md +14 -15
- package/pipeline/multi-agent-refs/features/external-context-injection.md +2 -0
- package/pipeline/multi-agent-refs/features/review-delta.md +1 -1
- package/pipeline/multi-agent-refs/features/review-multi-repo.md +3 -3
- package/pipeline/multi-agent-refs/features/security-audit.md +55 -0
- package/pipeline/multi-agent-refs/features/skill-conformance.md +1 -1
- package/pipeline/multi-agent-refs/features/visual-evidence.md +2 -1
- package/pipeline/multi-agent-refs/features/worktree-finalize.md +1 -1
- package/pipeline/multi-agent-refs/generate-issue.md +2 -0
- package/pipeline/multi-agent-refs/issue-jira-triad.md +2 -0
- package/pipeline/multi-agent-refs/keychain.md +2 -0
- package/pipeline/multi-agent-refs/knowledge.md +0 -7
- package/pipeline/multi-agent-refs/outside-the-pipeline.md +1 -1
- package/pipeline/multi-agent-refs/payload-contracts.md +1 -1
- package/pipeline/multi-agent-refs/phases/modes.md +33 -109
- package/pipeline/multi-agent-refs/phases/operations.md +2 -0
- package/pipeline/multi-agent-refs/phases/phase-0-init.md +23 -42
- package/pipeline/multi-agent-refs/phases/phase-1-plan.md +7 -18
- package/pipeline/multi-agent-refs/phases/phase-2-dev.md +13 -44
- package/pipeline/multi-agent-refs/phases/phase-3-review.md +28 -37
- package/pipeline/multi-agent-refs/phases/phase-4-commit.md +6 -6
- package/pipeline/multi-agent-refs/phases/phase-5-report.md +3 -3
- package/pipeline/multi-agent-refs/phases.md +9 -11
- package/pipeline/multi-agent-refs/progress-contract.md +1 -1
- package/pipeline/multi-agent-refs/readiness-review.md +2 -0
- package/pipeline/multi-agent-refs/rules.md +1 -1
- package/pipeline/multi-agent-refs/threat-model.md +39 -0
- package/pipeline/multi-agent-refs/tracker-contract.md +9 -40
- package/pipeline/multi-agent-refs/wiki-capture.md +3 -2
- package/pipeline/preferences-template.json +2 -2
- package/pipeline/rules/figma-pipeline.md +1 -1
- package/pipeline/schemas/agent-state.schema.json +28 -10
- package/pipeline/schemas/migrations/prefs-2.7.0-to-2.8.0.mjs +33 -0
- package/pipeline/schemas/phases.json +4 -26
- package/pipeline/schemas/prefs.schema.json +5 -9
- package/pipeline/schemas/reviewer-output.schema.json +99 -2
- package/pipeline/schemas/security-finding.schema.json +144 -0
- package/pipeline/scripts/_stack-routing.mjs +1 -0
- package/pipeline/scripts/cost-table.json +1 -1
- package/pipeline/scripts/gc-abandoned.sh +16 -9
- package/pipeline/scripts/gc-refs.sh +1 -1
- package/pipeline/scripts/gen-mode-dispatch.mjs +11 -41
- package/pipeline/scripts/migrate-prefs.mjs +18 -17
- package/pipeline/scripts/phase-tracker.sh +2 -2
- package/pipeline/scripts/phase0-exit-gate.mjs +1 -1
- package/pipeline/scripts/plan-coverage-gate.mjs +3 -3
- package/pipeline/scripts/render-work-summary.sh +7 -4
- package/pipeline/scripts/run-aggregator.mjs +1 -1
- package/pipeline/scripts/usage-report.mjs +0 -2
- package/pipeline/scripts/worktree-finalize.sh +2 -2
- package/pipeline/skills/.skill-manifest.json +17 -21
- package/pipeline/skills/.skills-index.json +6 -39
- package/pipeline/skills/shared/README.md +5 -8
- package/pipeline/skills/shared/core/multi-agent/SKILL.md +11 -15
- package/pipeline/skills/shared/core/multi-agent-autopilot-status/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +13 -16
- package/pipeline/skills/shared/core/multi-agent-ios-coding-standard/SKILL.md +2 -3
- package/pipeline/skills/shared/core/multi-agent-resume/SKILL.md +51 -15
- package/pipeline/skills/shared/core/multi-agent-scan/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-security-review/SKILL.md +29 -0
- package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +7 -7
- package/pipeline/skills/shared/external/security-review/SKILL.md +64 -0
- package/pipeline/skills/shared/external/security-review/references/owasp-mobile-top10-2024.md +53 -0
- package/pipeline/skills/shared/external/security-review/references/owasp-web-api-top10-2021.md +56 -0
- package/pipeline/skills/skills-index.md +3 -6
- package/pipeline/commands/multi-agent/local/SKILL.md +0 -132
- package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +0 -142
- package/pipeline/commands/multi-agent/resume-local/SKILL.md +0 -114
- package/pipeline/commands/security-review.md +0 -6
- package/pipeline/skills/shared/core/multi-agent-local/SKILL.md +0 -41
- package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +0 -55
- package/pipeline/skills/shared/core/multi-agent-resume-local/SKILL.md +0 -51
|
@@ -65,8 +65,8 @@ Run every step automatically:
|
|
|
65
65
|
```
|
|
66
66
|
Step 0: DOCTOR doctor.mjs - exit 2 or 4 stops the sync
|
|
67
67
|
Step 1.5: DETECT Compare timestamps, find stale targets
|
|
68
|
-
Step 2: COPILOT Claude Code -> Copilot CLI (instructions +
|
|
69
|
-
Step 2b: CODEX Claude Code -> Codex CLI (1 router skill +
|
|
68
|
+
Step 2: COPILOT Claude Code -> Copilot CLI (instructions + 58 sub-command skills)
|
|
69
|
+
Step 2b: CODEX Claude Code -> Codex CLI (1 router skill + 57 specs as refs + 8 agent TOML)
|
|
70
70
|
Step 3: REPO Claude Code -> pipeline repo (genericized, personal data scrub, bash -n on all sh)
|
|
71
71
|
Step 3c: PLUGINS pipeline shared/external -> multi-agent-plugins marketplace (rebuild knowledge/,
|
|
72
72
|
bump changed plugins' patch version, commit + push the plugins repo)
|
|
@@ -90,7 +90,7 @@ Step 0 gate rules and why: `features/doctor.md`.
|
|
|
90
90
|
- `copilot-instructions.md`: general development instructions + pipeline summary section
|
|
91
91
|
- `multi-agent-pipeline/pipeline/`: generic open-source version (NO personal data)
|
|
92
92
|
3. **Sync the shared sections** (Claude ↔ Copilot):
|
|
93
|
-
- Pipeline entries table (base / :
|
|
93
|
+
- Pipeline entries table (base / :autopilot)
|
|
94
94
|
- Project detection (URL-based + cwd-based)
|
|
95
95
|
- Figma pipeline flow
|
|
96
96
|
- Git conventions (author, branch, commit format)
|
|
@@ -494,16 +494,15 @@ same 56 specs as reference files rather than as peer skills, via Step 2b - see
|
|
|
494
494
|
|-------------|-------------|
|
|
495
495
|
| `~/.claude/commands/multi-agent/{cmd}/SKILL.md` | `~/.copilot/skills/multi-agent-{cmd}/SKILL.md` |
|
|
496
496
|
|
|
497
|
-
**
|
|
497
|
+
**58 commands are synced** (canonical inventory - must match `cross-cli-contract.md` section 1; drift = contract violation):
|
|
498
498
|
|
|
499
499
|
```
|
|
500
500
|
analysis, analysis-jira, analysis-resolve, autopilot, autopilot-off, autopilot-on,
|
|
501
501
|
autopilot-status, build-optimize, channels, complaint-analysis, create-jira,
|
|
502
502
|
design-check, diff-explain, doctor, feedback, forget, garbage-collect, graph, help,
|
|
503
|
-
ios-coding-standard, issue, jira, kill, language,
|
|
504
|
-
manual-test, model, prune-logs, prune-prompts, purge, refactor, resume,
|
|
505
|
-
|
|
506
|
-
routines, save, scan, search, setup, stack, status, steer, store-ready, sync, test,
|
|
503
|
+
ios-coding-standard, issue, jira, kill, language, log,
|
|
504
|
+
manual-test, model, prune-logs, prune-prompts, purge, refactor, resume, review, review-analysis, review-issue, review-jira, route-off, route-on, route-status,
|
|
505
|
+
routines, save, scan, search, security-review, setup, stack, status, steer, store-ready, sync, test,
|
|
507
506
|
test-accessibility, test-dark-mode, test-dynamic-type, test-screenshots,
|
|
508
507
|
testflight-validation, uninstall, update
|
|
509
508
|
```
|
|
@@ -6,6 +6,8 @@ argument-hint: "[locale] - boş = tr; örn. tr | en | de | ar"
|
|
|
6
6
|
|
|
7
7
|
# /multi-agent:test-screenshots - Locale screenshot set
|
|
8
8
|
|
|
9
|
+
> **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
|
|
10
|
+
|
|
9
11
|
Fixed-scenario alias, **parameterised by locale**. The scenario is pinned; the
|
|
10
12
|
language is not, because one command per language would put a parameter in a name.
|
|
11
13
|
|
|
@@ -7,6 +7,8 @@ argument-hint: "[--dry-run] [--all-data] [--claude] [--copilot] [--target=<path>
|
|
|
7
7
|
|
|
8
8
|
# multi-agent uninstall - Token-Preserving Uninstall
|
|
9
9
|
|
|
10
|
+
> **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
|
|
11
|
+
|
|
10
12
|
Uninstalls the pipeline itself from the system. **This is different from `:purge`:**
|
|
11
13
|
|
|
12
14
|
| Command | What it removes |
|
|
@@ -8,7 +8,7 @@
|
|
|
8
8
|
# (.multi-agent/). Only the first was ever excluded, and only by a one-line
|
|
9
9
|
# guard duplicated across two phase docs and gc-worktrees.sh. The other two
|
|
10
10
|
# were never excluded anywhere, which is invisible in worktree mode (the
|
|
11
|
-
# worktree is deleted) and permanent
|
|
11
|
+
# worktree is deleted) and permanent with a local workspace (nothing deletes it).
|
|
12
12
|
#
|
|
13
13
|
# Functions:
|
|
14
14
|
# ma_hygiene_ensure_exclusions <repo-root> Write the managed block into
|
|
@@ -64,7 +64,7 @@ Phase 3.2 produced the list. Sort every entry and act:
|
|
|
64
64
|
| Bucket | Meaning | Action |
|
|
65
65
|
|---|---|---|
|
|
66
66
|
| A. Not searched | The evidence is reachable and this run did not look | Run the search; it closes, and never becomes a question |
|
|
67
|
-
| B. The user knows | Channel scope, which service backs a screen, what is in scope | Ask, `AskUserQuestion`, at most 4 per call |
|
|
67
|
+
| B. The user knows | Channel scope, which service backs a screen, what is in scope | Ask, `AskUserQuestion` per `picker-contract.md`, at most 4 per call |
|
|
68
68
|
| C. External | Final copy, a legal basis, a contract nobody has written | Record as `AS-NN` in Section 20 |
|
|
69
69
|
|
|
70
70
|
**Bucket A must be empty before Phase 3.5.** A gap reaches C only with a stamp:
|
|
@@ -6,7 +6,7 @@
|
|
|
6
6
|
|
|
7
7
|
## New Locked decisions (this command)
|
|
8
8
|
|
|
9
|
-
1. **One question per `AskUserQuestion` call
|
|
9
|
+
1. **One question per `AskUserQuestion` call**, shaped by `picker-contract.md`. Never batch Section 20 rows. Sequential resolution keeps each decision explicit and traceable.
|
|
10
10
|
2. **Up to 3 source-labeled candidates plus Defer = 4 options max.** Each candidate's `description` carries its source label: `From evidence - <doc section / citation>`, `From repo - <file:line>`, `AI reasoned - <one-line rationale>`. The auto-provided Other accepts free text and stop tokens.
|
|
11
11
|
3. **Three sources, never blended.**
|
|
12
12
|
- **From evidence**: the answer is already derivable from the doc's own sections and their citations (Sections 5, 6, 9, 13, 21) or from cached `state.analysisSpec.evidence.*` when the session still holds it. Cite the section or evidence bucket.
|
|
@@ -34,7 +34,7 @@ synthesizedSections = {
|
|
|
34
34
|
|
|
35
35
|
#### Phase 2a - Pass B preview (Locked 25)
|
|
36
36
|
|
|
37
|
-
Before Pass B renders any file, present the resolved convention table to the user via `AskUserQuestion
|
|
37
|
+
Before Pass B renders any file, present the resolved convention table to the user via `AskUserQuestion` (`picker-contract.md`). The table is one row per concept, one column per selected platform.
|
|
38
38
|
|
|
39
39
|
Example (iOS + Android selected):
|
|
40
40
|
|
|
@@ -830,7 +830,7 @@ Framework: iOS XCUITest (`waitForExistence`, identifier-driven); Android Compose
|
|
|
830
830
|
|
|
831
831
|
### 15.7 Manuel test senaryoları / Manual test scenarios
|
|
832
832
|
|
|
833
|
-
The scenarios a person runs by hand. Same format the pipeline already posts as the Jira test-scenario comment (`/multi-agent:resume
|
|
833
|
+
The scenarios a person runs by hand. Same format the pipeline already posts as the Jira test-scenario comment (`/multi-agent:resume`), defined once and read by both. Omitted entirely when the feature has none (Locked 2); never rendered as an empty table.
|
|
834
834
|
|
|
835
835
|
Each row is executable by someone who did not write the feature: no "verify it works", no implied setup. Cover the happy path, at least one boundary, and every failure mode that reaches the user - the same four-way split Locked 30 requires of unit tests.
|
|
836
836
|
|
|
@@ -7,11 +7,10 @@
|
|
|
7
7
|
- [Subphase contract (dispatch-layer owned)](#subphase-contract-dispatch-layer-owned)
|
|
8
8
|
- [Multi-repo report](#multi-repo-report)
|
|
9
9
|
- [Failure + resume](#failure-resume)
|
|
10
|
-
- [Short-run behaviour](#short-run-behaviour)
|
|
11
10
|
- [Cross-CLI behaviour (intentional divergence)](#cross-cli-behaviour-intentional-divergence)
|
|
12
11
|
<!-- /toc -->
|
|
13
12
|
|
|
14
|
-
> **TLDR** - When `taskType === "component"` (Figma URL in task description or instruction-driven figma workflow), multi-agent Phase
|
|
13
|
+
> **TLDR** - When `taskType === "component"` (Figma URL in task description or instruction-driven figma workflow), multi-agent Phase 2 **does not run the TDD loop**. It delegates the entire phase to the enabled `ai-<platform>-toolkit` **marketplace plugin's** component skill (`create-component`, falling back to `create-ui-component`) via the Skill tool. Implementation lives in the plugin; multi-agent's job is classification, dispatch, and state report. The pipeline no longer bundles its own `figma-to-component` orchestrator - component skills live in one place, the plugin marketplace.
|
|
15
14
|
|
|
16
15
|
This doc is referenced from `$HOME/.claude/multi-agent-refs/phases/phase-2-dev.md`. Keeping it separate lets `phase-2-dev.md` remain tight (it's already the largest phase doc) and gives the orchestrator-report contract a stable URL for both Claude-side and Copilot-side implementations.
|
|
17
16
|
|
|
@@ -79,13 +78,12 @@ Emit `→ dispatching create-component <componentName>` via the progress contrac
|
|
|
79
78
|
|
|
80
79
|
- `figmaUrl` - primary argument (the plugin skill's `<figma-url>`)
|
|
81
80
|
- the component name + the analysis Section 6 (Bileşen Envanteri) row + Section 13.1 conventions as context (the plugin skill does not re-read the pipeline's analysis doc on its own - pass what it needs)
|
|
82
|
-
- `mode` - `"dev"` vs `"full"`, from `state.onlyDevelop`, so the plugin can elide tests/wiki on a Short run
|
|
83
81
|
|
|
84
82
|
Plugin skills are user-facing lifecycle skills; they do **not** accept an `agentState` path and do **not** write `agent-state.json`. State tracking therefore moves to the dispatch layer (next section).
|
|
85
83
|
|
|
86
84
|
## Subphase contract (dispatch-layer owned)
|
|
87
85
|
|
|
88
|
-
Because the plugin skill does not write pipeline state, **multi-agent's dispatch layer owns `state.phases["
|
|
86
|
+
Because the plugin skill does not write pipeline state, **multi-agent's dispatch layer owns `state.phases["2"].subphases[]`**, not the skill. On dispatch, seed one coarse component-build subphase; on return, finalize it:
|
|
89
87
|
|
|
90
88
|
```json
|
|
91
89
|
{
|
|
@@ -115,17 +113,11 @@ Multi-agent does **not** split the component task into per-repo Phase 2 runs -
|
|
|
115
113
|
|
|
116
114
|
On failure (the plugin skill returns an unrecoverable build/test error, or the dispatch hits retry cap):
|
|
117
115
|
|
|
118
|
-
1. The dispatch layer persists `{error, buildLog, testLog}` to `state.phases["
|
|
119
|
-
2. Multi-agent increments `state.phases["
|
|
116
|
+
1. The dispatch layer persists `{error, buildLog, testLog}` to `state.phases["2"].errors[]`.
|
|
117
|
+
2. Multi-agent increments `state.phases["2"].retryCount`.
|
|
120
118
|
3. Retry re-invokes the plugin skill (it is idempotent on an existing component; it reconciles rather than duplicating). There is no bundled `phase-<N>` resume anymore.
|
|
121
119
|
4. Hard kill at `retryCount === 3` → surface the errors to the user, halt Phase 2. Do not loop indefinitely.
|
|
122
120
|
|
|
123
|
-
## Short-run behaviour
|
|
124
|
-
|
|
125
|
-
When the Phase 0 Step 7.5 depth picker answered Short (`state.onlyDevelop === true`), the dispatch layer passes `mode: "dev"` so the plugin skill can elide unit tests and wiki (structural + snapshot still required; wiki deferred to Phase 5). If the plugin does not honor a `mode` hint, dispatch simply skips the post-build wiki step itself.
|
|
126
|
-
|
|
127
|
-
Phase 3 runs in a Short run as it does in a Full one, and its reviewer count is **not** Phase 3's concern - the Step 1.77 scope gate decides that from diff risk, independently of `mode`. What the dispatch layer owes Phase 4 is the record of which plugin skill it delegated to, appended to `state.telemetry.skillCalls[]`, so the review can check the delivered component against the criteria that skill imposes.
|
|
128
|
-
|
|
129
121
|
## Cross-CLI behaviour (intentional divergence)
|
|
130
122
|
|
|
131
123
|
Component dispatch is **no longer byte-identical across CLIs** and that is by design (see `cross-cli-contract.md` section 1.1):
|
|
@@ -133,4 +125,4 @@ Component dispatch is **no longer byte-identical across CLIs** and that is by de
|
|
|
133
125
|
- **Claude Code**: dispatches to the enabled `ai-<platform>-toolkit` marketplace plugin via the Skill tool (this doc).
|
|
134
126
|
- **Copilot CLI**: has no plugin loader; it continues to use its standalone `~/.copilot/skills/figma-*` skill copies (frozen fallback). Copilot's resolution + progress lines follow those local skills.
|
|
135
127
|
|
|
136
|
-
What still MUST match across CLIs: the `taskType === "component"` classification, the `state.phases["
|
|
128
|
+
What still MUST match across CLIs: the `taskType === "component"` classification, the `state.phases["2"].subphases[]` shape the dispatch layer writes, and the Short-run elision semantics. Figma-skill *inventory* parity is no longer enforced. `smoke-cross-cli-behavior.sh` asserts only the classification + state-shape axis for components.
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Cross-CLI Contract (Claude Code · Copilot CLI · Codex CLI)
|
|
2
2
|
|
|
3
3
|
<!-- toc -->
|
|
4
|
-
- [1. Command Inventory (
|
|
4
|
+
- [1. Command Inventory (58 commands)](#1-command-inventory-58-commands)
|
|
5
5
|
- [2. Canonical Placeholder Vocabulary](#2-canonical-placeholder-vocabulary)
|
|
6
6
|
- [2.6 Intentional structural divergence - thin dispatcher vs inlined orchestrator](#26-intentional-structural-divergence---thin-dispatcher-vs-inlined-orchestrator)
|
|
7
7
|
- [2.7 One command, three different meanings: `/multi-agent:model`](#27-one-command-three-different-meanings-multi-agentmodel)
|
|
@@ -20,18 +20,18 @@
|
|
|
20
20
|
|
|
21
21
|
---
|
|
22
22
|
|
|
23
|
-
## 1. Command Inventory (
|
|
23
|
+
## 1. Command Inventory (58 commands)
|
|
24
24
|
|
|
25
25
|
```
|
|
26
26
|
analysis, analysis-jira, analysis-resolve, autopilot, autopilot-off,
|
|
27
27
|
autopilot-on, autopilot-status, build-optimize, channels, complaint-analysis,
|
|
28
28
|
create-jira, design-check, diff-explain, doctor, feedback, forget,
|
|
29
29
|
garbage-collect, graph, help, ios-coding-standard, issue, jira, kill,
|
|
30
|
-
language,
|
|
31
|
-
prune-prompts, purge, refactor, resume,
|
|
30
|
+
language, log, manual-test, model, prune-logs,
|
|
31
|
+
prune-prompts, purge, refactor, resume, review,
|
|
32
32
|
review-analysis, review-issue, review-jira, route-off, route-on,
|
|
33
33
|
route-status, routines, save, scan, search,
|
|
34
|
-
setup, stack, status, steer, store-ready, sync, test, test-accessibility,
|
|
34
|
+
security-review, setup, stack, status, steer, store-ready, sync, test, test-accessibility,
|
|
35
35
|
test-dark-mode, test-dynamic-type, test-screenshots, testflight-validation,
|
|
36
36
|
uninstall, update
|
|
37
37
|
```
|
|
@@ -40,8 +40,8 @@ Categories:
|
|
|
40
40
|
|
|
41
41
|
- **Interactive pickers** (single-purpose, not modes): `jira`, `issue`
|
|
42
42
|
- **Issue generator** (one-shot, no worktree, asks type Task/Bug/Story, hard approval gate before create): `create-jira`
|
|
43
|
-
- **Pipeline entries**: `autopilot
|
|
44
|
-
- **Tail
|
|
43
|
+
- **Pipeline entries**: `autopilot` (plus the bare `/multi-agent` in the dispatcher). The phase set is a property of the command, and every mode runs its whole set.
|
|
44
|
+
- **Tail path** (the pipeline tail over work already on a branch, no run behind it): `resume`
|
|
45
45
|
- **Ops commands** (one-shot, no worktree): `status`, `log`, `kill`, `steer`, `purge`, `uninstall`, `resume`, `review`, `review-jira`, `review-issue`, `analysis`, `analysis-resolve`, `complaint-analysis`, `build-optimize`, `channels`, `scan`, `search`, `diff-explain`, `garbage-collect`, `graph`, `prune-logs`, `prune-prompts`
|
|
46
46
|
- **Local audits** (worktree only to build; no commit, push, PR or channels): `design-check`, `testflight-validation`, `ios-coding-standard`. `testflight-validation` additionally never invokes `altool --upload-app` - a validation run must not be able to ship a build by accident.
|
|
47
47
|
- **Meta-ops**: `setup`, `sync`, `update`, `help`, `refactor`, `test`, `stack`, `manual-test`, `language`
|
|
@@ -86,7 +86,7 @@ Three rules make that work:
|
|
|
86
86
|
> absent, and a component task on Copilot had nothing to dispatch to. The contract
|
|
87
87
|
> documented a safety net the installer deleted.
|
|
88
88
|
|
|
89
|
-
What still MUST match across CLIs for component tasks: the `taskType === "component"` classification, the `state.phases["
|
|
89
|
+
What still MUST match across CLIs for component tasks: the `taskType === "component"` classification, the `state.phases["2"].subphases[]` shape the dispatch layer writes, and the Short-run elision semantics. Resolution/routing is Claude-plugin vs Copilot-local-copy by design.
|
|
90
90
|
|
|
91
91
|
### 1.2 Store-compliance skills
|
|
92
92
|
|
|
@@ -97,7 +97,7 @@ Two parallel shared/core skills under `pipeline/skills/shared/core/` wrap extern
|
|
|
97
97
|
| `apple-archive-compliance` | `pipeline/skills/shared/core/apple-archive-compliance/` | `ios_app_store_audit` MCP tool (in `@mmerterden/multi-agent-toolkit-mcp` ≥ v3.0.0) | 18 (Apple ITMS + App Store Review Guidelines) |
|
|
98
98
|
| `google-play-compliance` | `pipeline/skills/shared/core/google-play-compliance/` | bundletool + aapt2 + apksigner | 21 (4 categories: Technical / Security / Privacy / Hygiene) |
|
|
99
99
|
|
|
100
|
-
Both skills are wired to 4 consumers: `/multi-agent:test "store-ready"` (primary), Phase
|
|
100
|
+
Both skills are wired to 4 consumers: `/multi-agent:test "store-ready"` (primary), Phase 3 Security Auditor (`pipeline/agents/security-auditor.md`), `/multi-agent:review` + SKILL.md counterpart, `/multi-agent:channels` PR-body auto-augmentation. Contract enforced by `smoke-compliance-skills.sh`.
|
|
101
101
|
|
|
102
102
|
### 1.3 Figma-skill routing from multi-agent (superseded)
|
|
103
103
|
|
|
@@ -222,7 +222,7 @@ Phase 3 runs three reviewers everywhere, but the diversity those three buy is no
|
|
|
222
222
|
same on every host. Copilot CLI gets cross-VENDOR disagreement for free: GPT-5.4 sits
|
|
223
223
|
beside two Claude models. Claude Code and Codex each run a one-vendor panel - three
|
|
224
224
|
Anthropic models on one, three OpenAI models on the other - so the same three-way
|
|
225
|
-
agreement is weaker evidence there, and Phase
|
|
225
|
+
agreement is weaker evidence there, and Phase 3 says so in the triage note on a
|
|
226
226
|
borderline finding.
|
|
227
227
|
|
|
228
228
|
Where the budget goes instead, when vendor diversity is unavailable:
|
|
@@ -367,11 +367,10 @@ Given the same user input, both CLIs must produce identical parsed state:
|
|
|
367
367
|
| `#42` or bare `42` | `input.type = "github-issue-number"`, `input.issueNo = 42` |
|
|
368
368
|
| Any other string | `input.type = "free-text"`, `input.summary = <string>` |
|
|
369
369
|
|
|
370
|
-
|
|
371
|
-
- `autopilot` - skip confirmations
|
|
372
|
-
|
|
373
|
-
|
|
374
|
-
- Combinations: `--local autopilot`. Short plus autopilot is NOT a combination: autopilot never asks the depth question and always runs Full.
|
|
370
|
+
One modifier exists:
|
|
371
|
+
- `autopilot` - skip confirmations, and resolve the workspace to a worktree without asking.
|
|
372
|
+
|
|
373
|
+
The workspace is not a flag. Where the branch lives is a Phase 0 Step 5b question on attended runs.
|
|
375
374
|
|
|
376
375
|
---
|
|
377
376
|
|
|
@@ -7,6 +7,8 @@
|
|
|
7
7
|
- [Log line shape](#log-line-shape)
|
|
8
8
|
<!-- /toc -->
|
|
9
9
|
|
|
10
|
+
> **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
|
|
11
|
+
|
|
10
12
|
Phase 0 Step 1b catalogued every typed external link from the task description into `state.contextLinks[]`. Phase 1 dispatches each entry to its matching fetcher and prepends the result to the analysis prompt under a **Referenced External Sources** section, so the agent doesn't re-discover what the ticket already pointed at.
|
|
11
13
|
|
|
12
14
|
```bash
|
|
@@ -63,7 +63,7 @@ DELTA_JSON=$(node $HOME/.claude/scripts/review-delta.mjs --rounds-dir "$WORKTREE
|
|
|
63
63
|
jq -c --argjson d "$DELTA_JSON" --argjson i "$((ITERATION-1))" \
|
|
64
64
|
'{reviewIterations: (.reviewIterations | .[$i] += {delta: ($d + {computedAt: (now | todate)})})}' "$STATE_FILE" \
|
|
65
65
|
| node $HOME/.claude/scripts/write-state.mjs "$STATE_FILE"
|
|
66
|
-
$HOME/.claude/scripts/log-metric.sh "$TASK_ID"
|
|
66
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 review.delta iteration=$ITERATION \
|
|
67
67
|
new=$(jq '.counts.new // 0' <<< "$DELTA_JSON") still_present=$(jq '.counts.stillPresent // 0' <<< "$DELTA_JSON") \
|
|
68
68
|
resolved=$(jq '.counts.resolved // 0' <<< "$DELTA_JSON") plateau=$(jq '.plateau // false' <<< "$DELTA_JSON") tripped=$([ "$DELTA_RC" -eq 3 ] && echo true || echo false)
|
|
69
69
|
```
|
|
@@ -57,9 +57,9 @@ And log `review.diff_truncated repo=<name> bytes_dropped=<N>`. Triage receives t
|
|
|
57
57
|
|
|
58
58
|
**Telemetry**: Per-repo build/test gate timings + a single combined review/triage call set:
|
|
59
59
|
```bash
|
|
60
|
-
$HOME/.claude/scripts/log-metric.sh "$TASK_ID"
|
|
61
|
-
$HOME/.claude/scripts/log-metric.sh "$TASK_ID"
|
|
62
|
-
$HOME/.claude/scripts/log-metric.sh "$TASK_ID"
|
|
60
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 gate.build repo=common status=pass duration_ms=$D
|
|
61
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 gate.build repo=uicomponents status=pass duration_ms=$D
|
|
62
|
+
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 review.combined_diff repos=2 bytes=$BYTES truncated=false
|
|
63
63
|
```
|
|
64
64
|
|
|
65
65
|
---
|
|
@@ -0,0 +1,55 @@
|
|
|
1
|
+
# Security audit (Phase 3 Step 2.7)
|
|
2
|
+
|
|
3
|
+
The mechanics of the conditional security audit that runs inside the Phase 3 review window and merges at Step 3.0. The phase doc carries the trigger and the merge; the rest is here.
|
|
4
|
+
|
|
5
|
+
## Trigger (deterministic, no new classifier)
|
|
6
|
+
|
|
7
|
+
Run the audit when any of these holds:
|
|
8
|
+
|
|
9
|
+
- Step 1.75 scored `security_path` on a file in the diff (the same signal that already forces `full` review at Step 1.77).
|
|
10
|
+
- The base branch is a release branch.
|
|
11
|
+
- The run is the standalone `/multi-agent:security-review` command, which dispatches straight to this step.
|
|
12
|
+
|
|
13
|
+
There is no `--audit` flag: the trigger is the diff, not a word the user has to remember. When none of these holds, the audit does not run, `$SECURITY_AUDIT_JSON` stays empty, and the Step 3.0 merge is a no-op.
|
|
14
|
+
|
|
15
|
+
## Threat model first
|
|
16
|
+
|
|
17
|
+
The auditor reads `.pipeline/threat-model.md` if Phase 1 or a prior step wrote it, and produces it (four sections) if absent. Contract: `$HOME/.claude/multi-agent-refs/threat-model.md`. It is keyed to repo+branch and mirrored to `state.threatModel` so a resume reuses it rather than re-deriving it.
|
|
18
|
+
|
|
19
|
+
## Dispatch
|
|
20
|
+
|
|
21
|
+
One `Agent(subagent_type: "security-auditor")` on the same capped diff the reviewers saw, plus the threat model and the resolved `${CRITERIA}` block:
|
|
22
|
+
|
|
23
|
+
```
|
|
24
|
+
Agent(subagent_type: "security-auditor", prompt: "<threat-model + capped diff + ${CRITERIA}>")
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
Persona contract: `~/.claude/agents/security-auditor.md`.
|
|
28
|
+
|
|
29
|
+
## Output is reviewer-shaped, and that is the point
|
|
30
|
+
|
|
31
|
+
The auditor returns one object conforming to `$HOME/.claude/schemas/reviewer-output.schema.json`, whose `findings[]` each carry the `security` envelope of `security-finding.schema.json`: OWASP category, CWE, CVSS vector + band, evidence, counterevidence, confidence, remediation, and the before/after fix. Severity is the reviewer enum, derived from the CVSS band:
|
|
32
|
+
|
|
33
|
+
| CVSS band | severity |
|
|
34
|
+
| --------- | -------- |
|
|
35
|
+
| critical / high | blocking |
|
|
36
|
+
| medium | important |
|
|
37
|
+
| low / none | suggestion |
|
|
38
|
+
|
|
39
|
+
Compute `security.cvss.baseScore` + `band` with the toolkit `security_cvss_score` tool, not by hand, so the score cannot drift from the vector.
|
|
40
|
+
|
|
41
|
+
## Validate, persist, merge
|
|
42
|
+
|
|
43
|
+
Validate with the same gate protocol as a reviewer - the exit code decides, not the LLM turn:
|
|
44
|
+
|
|
45
|
+
```bash
|
|
46
|
+
printf '%s' "$SECURITY_AUDIT_JSON" | node "$HOME/.claude/scripts/validate-reviewer.mjs" -
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
On validator failure: one self-correction rework, then HALT the phase (identical to the reviewer output-contract gate). Persist the object to `state.reviewIterations[<iteration>].securityAudit` and hold it in `$SECURITY_AUDIT_JSON`.
|
|
50
|
+
|
|
51
|
+
Because the output is reviewer-shaped, the Step 3.0 merge appends its `findings[]` alongside the test-integrity findings. A `blocking` security finding then reaches triage, and a triage-accepted blocker blocks Phase 4 - the "critical security items block the commit" contract is wired here, not merely stated.
|
|
52
|
+
|
|
53
|
+
## Store-compliance cross-reference
|
|
54
|
+
|
|
55
|
+
On store-relevant diffs the auditor also cites the Apple ITMS / Google Play catalog rule on the finding via `ruleId` + `criteriaSource` (`apple-archive-compliance` / `google-play-compliance`). This is annotation on the same finding, not a second pass; the device-level `store-ready` Gate under `/multi-agent:test` stays separate.
|
|
@@ -22,7 +22,7 @@ Phase 4 used to select its review criteria from `detectedStack`, a string Phase
|
|
|
22
22
|
- The dev side and the review side could disagree about the standard without either noticing. A component built by a plugin skill was reviewed against `clean-code`.
|
|
23
23
|
- "Did it do this correctly?" had no fixed denominator, so the only available answer was "it looks fine", and a reviewer that opened nothing produced the same output as a reviewer that checked everything.
|
|
24
24
|
|
|
25
|
-
|
|
25
|
+
The sharper case is a document whose evidence supported few sections: the sole input to the old selection logic was then nearly empty.
|
|
26
26
|
|
|
27
27
|
## The four rules that make it work
|
|
28
28
|
|
|
@@ -123,7 +123,8 @@ section still renders, saying so.
|
|
|
123
123
|
## 3. After - Phase 2, not the Phase 3 user test
|
|
124
124
|
|
|
125
125
|
The user test is the natural home: the simulator is already up. It is also **dropped by
|
|
126
|
-
every `autopilot`
|
|
126
|
+
every `autopilot` entry and by any run whose
|
|
127
|
+
workspace is local** (`phase-3-review.md` TLDR), so a capture
|
|
127
128
|
that lives only there produces nothing for unattended runs - which are exactly
|
|
128
129
|
the runs where nobody watched the screen.
|
|
129
130
|
|
|
@@ -20,7 +20,7 @@ Removing it at PR-open is only safe because of the salvage, so the two are one s
|
|
|
20
20
|
|
|
21
21
|
| Condition | Why it blocks |
|
|
22
22
|
|---|---|
|
|
23
|
-
| `worktreePath == projectRoot` (
|
|
23
|
+
| `worktreePath == projectRoot` (local workspace) | there is no worktree; removing it would delete the user's checkout |
|
|
24
24
|
| cwd is inside the worktree | a shell left on a deleted inode is worse than a leftover directory, and Phase 4 legitimately `cd`s into the worktree earlier |
|
|
25
25
|
| not a registered worktree of the project root | a mistyped path must not delete an unrelated directory |
|
|
26
26
|
| real uncommitted changes | never discarded; see the artefact carve-out below |
|
|
@@ -9,6 +9,8 @@
|
|
|
9
9
|
|
|
10
10
|
> **TLDR** - Shared 12-step flow for `/multi-agent:create-jira`. Asks the issue type (Task / Bug / Story), mines the target project's existing same-type issues to learn team conventions, detects the active sprint, drafts a standards-compliant issue from a fixed standard template with auto-sizing sections, asks the user about every genuinely unknown field, renders a full preview, and creates the Jira issue only after explicit approval. Creates exactly one Jira issue per run - no branches, no commits, no worktrees.
|
|
11
11
|
|
|
12
|
+
> **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
|
|
13
|
+
|
|
12
14
|
Consumed by `create-jira/SKILL.md`. This ref is never invoked directly.
|
|
13
15
|
|
|
14
16
|
## Hard rules (must not regress)
|
|
@@ -9,6 +9,8 @@
|
|
|
9
9
|
- [Cross-CLI parity](#cross-cli-parity)
|
|
10
10
|
<!-- /toc -->
|
|
11
11
|
|
|
12
|
+
> **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
|
|
13
|
+
|
|
12
14
|
> **TLDR** - When a GitHub issue triggers the pipeline and has no Jira ID, the `autoJiraFromGithubIssue` policy decides whether to auto-create a Jira task (and patch the GitHub issue body with the new Jira link). Phase 5 then posts a humanizer'd wiki-content summary back as a Jira comment, closing the loop. Autopilot treats `ask` as `always`.
|
|
13
15
|
|
|
14
16
|
This doc is referenced from `$HOME/.claude/multi-agent-refs/phases/phase-0-init.md` Step 1 (GitHub issue input) and `$HOME/.claude/multi-agent-refs/phases/phase-5-report.md` Step 2 (component wiki). Keeps the phase docs tight and gives the triad contract a stable home.
|
|
@@ -1,5 +1,7 @@
|
|
|
1
1
|
## Keychain Token Registry
|
|
2
2
|
|
|
3
|
+
> **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
|
|
4
|
+
|
|
3
5
|
Tokens live in the platform-native credential store (macOS Keychain / Linux libsecret / Windows Credential Manager). Key names are resolved via **preferences mapping** - never hardcoded.
|
|
4
6
|
|
|
5
7
|
### Rule 1 - Never prompt for a token *value* mid-run
|
|
@@ -49,13 +49,6 @@ Each task teaches the project a bit more. Token cost decreases over time.
|
|
|
49
49
|
|
|
50
50
|
They all complement each other - no conflicts.
|
|
51
51
|
|
|
52
|
-
### Knowledge in a Short run
|
|
53
|
-
|
|
54
|
-
A Short run skips Phase 1, but knowledge **is still read**:
|
|
55
|
-
|
|
56
|
-
- Knowledge files are added to the prompt before the Opus agent starts in Phase 3
|
|
57
|
-
- Knowledge capture is still performed in Phase 5 (simplified report)
|
|
58
|
-
|
|
59
52
|
### Knowledge Maintenance
|
|
60
53
|
|
|
61
54
|
Knowledge files grow over time. Maintenance rules:
|
|
@@ -47,7 +47,7 @@ and a review between fetch and action; an ordinary session has neither.
|
|
|
47
47
|
| Read an issue, page, log, crash, scan result | Resolve the credential and fetch |
|
|
48
48
|
| Comment on Jira, edit an issue, move a board column | `/multi-agent:channels` |
|
|
49
49
|
| Create an issue | `/multi-agent:create-jira` |
|
|
50
|
-
| Open or update a PR | a pipeline run, or `/multi-agent:resume
|
|
50
|
+
| Open or update a PR | a pipeline run, or `/multi-agent:resume` |
|
|
51
51
|
|
|
52
52
|
The split is not bureaucracy. Outward writes carry rules that live in those commands:
|
|
53
53
|
issues are never auto-closed (four approvals), PR bodies use `Ref:` and never
|
|
@@ -64,4 +64,4 @@ Non-empty `UNTRACKED` → both the agent-log report and the closing chat summary
|
|
|
64
64
|
|
|
65
65
|
## Fast modes are not exempt
|
|
66
66
|
|
|
67
|
-
|
|
67
|
+
The autopilot and local entries skip the interactive test gate. They run Phase 4 and Phase 5 **unchanged**. A short pipeline is not a licence for an improvised payload shape, a missing Test Scenarios section, or a report without numbers.
|