@mmerterden/multi-agent-pipeline 17.6.0 → 19.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +310 -0
- package/README.md +76 -18
- package/README.tr.md +55 -16
- package/docs/adr/0002-instruction-driven-flag.md +1 -0
- package/docs/adr/0005-lazy-phase-docs.md +11 -1
- package/docs/adr/0008-installer-modularization-and-secret-leak-defense.md +1 -0
- package/docs/adr/0010-own-code-graph.md +1 -0
- package/docs/adr/0011-dormant-ci.md +25 -1
- package/docs/adr/0014-six-phase-consolidation.md +134 -0
- package/docs/adr/README.md +2 -1
- package/docs/architecture.md +37 -38
- package/docs/best-practices.md +1 -1
- package/docs/ecosystem.md +37 -26
- package/docs/engineering.md +1 -1
- package/docs/facts.json +45 -0
- package/docs/features.md +54 -53
- package/docs/performance.md +5 -5
- package/docs/recovery-guide.md +9 -9
- package/docs/server-readiness.md +188 -0
- package/docs/token-budget-history.md +3 -1
- package/index.js +18 -3
- package/install/_codex-agents.mjs +1 -1
- package/install/_common.mjs +42 -17
- package/install/_dev-only-files.mjs +8 -0
- package/install/_unattended-profile.mjs +113 -0
- package/install/index.mjs +48 -0
- package/install/templates/claude-hooks.json +1 -1
- package/install/templates/codex-instructions.md +1 -1
- package/install/templates/copilot-instructions.md +28 -28
- package/manifest.json +1065 -0
- package/package.json +6 -3
- package/pipeline/agents/dev-critic.md +3 -3
- package/pipeline/commands/figma-to-swiftui.md +1 -1
- package/pipeline/commands/multi-agent/SKILL.md +8 -8
- package/pipeline/commands/multi-agent/analysis/SKILL.md +9 -9
- package/pipeline/commands/multi-agent/autopilot/SKILL.md +7 -7
- package/pipeline/commands/multi-agent/channels/SKILL.md +15 -15
- package/pipeline/commands/multi-agent/diff-explain/SKILL.md +6 -6
- package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/graph/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/help/SKILL.md +62 -62
- package/pipeline/commands/multi-agent/language/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/local/SKILL.md +11 -11
- package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +13 -13
- package/pipeline/commands/multi-agent/log/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/manual-test/SKILL.md +9 -9
- package/pipeline/commands/multi-agent/model/SKILL.md +69 -0
- package/pipeline/commands/multi-agent/refactor/SKILL.md +3 -3
- package/pipeline/commands/multi-agent/resume/SKILL.md +4 -4
- package/pipeline/commands/multi-agent/resume-local/SKILL.md +19 -17
- package/pipeline/commands/multi-agent/review/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/route-off/SKILL.md +36 -0
- package/pipeline/commands/multi-agent/route-on/SKILL.md +74 -0
- package/pipeline/commands/multi-agent/route-status/SKILL.md +56 -0
- package/pipeline/commands/multi-agent/setup/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/status/SKILL.md +54 -23
- package/pipeline/commands/multi-agent/steer/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/sync/SKILL.md +12 -13
- package/pipeline/commands/multi-agent/test/SKILL.md +1 -1
- package/pipeline/lib/_jira-auth.sh +8 -0
- package/pipeline/lib/analysis-jira-write.sh +32 -0
- package/pipeline/lib/ask-choice.sh +13 -2
- package/pipeline/lib/autopilot-state.sh +8 -0
- package/pipeline/lib/credential-inventory.sh +1 -1
- package/pipeline/lib/fatal.mjs +129 -0
- package/pipeline/lib/fetch-fortify.sh +1 -1
- package/pipeline/lib/figma-mcp-refresh.sh +18 -0
- package/pipeline/lib/figma-screenshot.sh +18 -0
- package/pipeline/lib/invoked-directly.mjs +43 -0
- package/pipeline/lib/jira-publish.sh +42 -0
- package/pipeline/lib/md2confluence-v3.py +47 -0
- package/pipeline/lib/model-rung.sh +142 -0
- package/pipeline/lib/outbound-gate.mjs +175 -0
- package/pipeline/lib/phase-schema.mjs +88 -0
- package/pipeline/lib/plan-todos.sh +32 -11
- package/pipeline/lib/post-pr-review.sh +77 -8
- package/pipeline/lib/repo-hygiene.sh +8 -3
- package/pipeline/lib/require-jq.sh +40 -0
- package/pipeline/lib/route-state.sh +161 -0
- package/pipeline/lib/run-paths.sh +335 -0
- package/pipeline/multi-agent-refs/_account-picker.md +1 -1
- package/pipeline/multi-agent-refs/_dev-context.md +1 -1
- package/pipeline/multi-agent-refs/_input-parser.md +1 -1
- package/pipeline/multi-agent-refs/analysis/evidence.md +0 -9
- package/pipeline/multi-agent-refs/analysis/intake.md +1 -1
- package/pipeline/multi-agent-refs/analysis/locked.md +21 -22
- package/pipeline/multi-agent-refs/analysis/render.md +1 -1
- package/pipeline/multi-agent-refs/analysis/synthesis.md +12 -6
- package/pipeline/multi-agent-refs/android-guide.md +1 -1
- package/pipeline/multi-agent-refs/audit-guide.md +13 -13
- package/pipeline/multi-agent-refs/channels/issue-comment.md +2 -2
- package/pipeline/multi-agent-refs/channels/jira.md +3 -3
- package/pipeline/multi-agent-refs/channels/pr.md +4 -4
- package/pipeline/multi-agent-refs/channels/wiki.md +1 -1
- package/pipeline/multi-agent-refs/component-dispatch.md +3 -3
- package/pipeline/multi-agent-refs/cross-cli-contract.md +31 -6
- package/pipeline/multi-agent-refs/features/autopilot-circuit-breaker.md +74 -4
- package/pipeline/multi-agent-refs/features/code-graph.md +5 -5
- package/pipeline/multi-agent-refs/features/cost-analysis.md +93 -0
- package/pipeline/multi-agent-refs/features/design-conformance.md +1 -1
- package/pipeline/multi-agent-refs/features/dev-critic.md +3 -3
- package/pipeline/multi-agent-refs/features/doctor.md +47 -2
- package/pipeline/multi-agent-refs/features/external-context-injection.md +3 -3
- package/pipeline/multi-agent-refs/features/maturity-followup.md +3 -3
- package/pipeline/multi-agent-refs/features/model-fallback.md +5 -5
- package/pipeline/multi-agent-refs/features/plan-todos.md +1 -1
- package/pipeline/multi-agent-refs/features/repo-map.md +1 -1
- package/pipeline/multi-agent-refs/features/review-delta.md +3 -3
- package/pipeline/multi-agent-refs/features/review-multi-repo.md +1 -1
- package/pipeline/multi-agent-refs/features/scope-check.md +4 -4
- package/pipeline/multi-agent-refs/features/skill-conformance.md +2 -2
- package/pipeline/multi-agent-refs/features/stack-skill-routing.md +1 -1
- package/pipeline/multi-agent-refs/features/verify-by-test.md +4 -4
- package/pipeline/multi-agent-refs/features/verify.md +83 -0
- package/pipeline/multi-agent-refs/features/visual-evidence.md +19 -19
- package/pipeline/multi-agent-refs/features/worktree-finalize.md +6 -6
- package/pipeline/multi-agent-refs/issue-jira-triad.md +10 -10
- package/pipeline/multi-agent-refs/knowledge.md +11 -11
- package/pipeline/multi-agent-refs/multi-repo-integration-build.md +13 -13
- package/pipeline/multi-agent-refs/payload-contracts.md +8 -8
- package/pipeline/multi-agent-refs/phases/log-format.md +10 -10
- package/pipeline/multi-agent-refs/phases/modes.md +30 -30
- package/pipeline/multi-agent-refs/phases/operations.md +21 -10
- package/pipeline/multi-agent-refs/phases/phase-0-init.md +25 -25
- package/pipeline/multi-agent-refs/phases/phase-1-plan.md +599 -0
- package/pipeline/multi-agent-refs/phases/{phase-3-dev.md → phase-2-dev.md} +129 -49
- package/pipeline/multi-agent-refs/phases/{phase-4-review.md → phase-3-review.md} +225 -107
- package/pipeline/multi-agent-refs/phases/{phase-6-commit.md → phase-4-commit.md} +23 -23
- package/pipeline/multi-agent-refs/phases/{phase-7-report.md → phase-5-report.md} +29 -29
- package/pipeline/multi-agent-refs/phases.md +44 -48
- package/pipeline/multi-agent-refs/picker-contract.md +1 -1
- package/pipeline/multi-agent-refs/progress-contract.md +6 -6
- package/pipeline/multi-agent-refs/readiness-review.md +1 -1
- package/pipeline/multi-agent-refs/rules.md +7 -7
- package/pipeline/multi-agent-refs/swiftui-guide.md +2 -2
- package/pipeline/multi-agent-refs/tracker-contract.md +31 -32
- package/pipeline/multi-agent-refs/unattended-contract.md +129 -0
- package/pipeline/multi-agent-refs/wiki-capture.md +14 -14
- package/pipeline/preferences-template.json +9 -1
- package/pipeline/rules/outside-the-pipeline.md +1 -1
- package/pipeline/schemas/agent-state.schema.json +50 -50
- package/pipeline/schemas/analysis-output.schema.json +2 -2
- package/pipeline/schemas/autopilot-config.schema.json +1 -1
- package/pipeline/schemas/code-graph.schema.json +1 -1
- package/pipeline/schemas/criteria-manifest.schema.json +1 -1
- package/pipeline/schemas/dev-critic-output.schema.json +1 -1
- package/pipeline/schemas/diff-risk.schema.json +1 -1
- package/pipeline/schemas/migrations/prefs-2.4.0-to-2.5.0.mjs +2 -2
- package/pipeline/schemas/migrations/prefs-2.6.0-to-2.7.0.mjs +31 -0
- package/pipeline/schemas/migrations/state-2.1.0-to-2.2.0.mjs +129 -0
- package/pipeline/schemas/phases.json +105 -0
- package/pipeline/schemas/plan-todos.schema.json +5 -5
- package/pipeline/schemas/planning-output.schema.json +1 -1
- package/pipeline/schemas/prefs.schema.json +100 -56
- package/pipeline/schemas/reviewer-output.schema.json +3 -3
- package/pipeline/schemas/route-config.schema.json +74 -0
- package/pipeline/schemas/scope-check.schema.json +1 -1
- package/pipeline/schemas/test-gap.schema.json +1 -1
- package/pipeline/schemas/token-budget.json +12 -18
- package/pipeline/schemas/triage-output.schema.json +6 -6
- package/pipeline/scripts/README.md +3 -3
- package/pipeline/scripts/_code-graph.mjs +2 -2
- package/pipeline/scripts/_run-paths.mjs +372 -0
- package/pipeline/scripts/_smoke-root.sh +1 -1
- package/pipeline/scripts/aggregate-metrics.mjs +65 -65
- package/pipeline/scripts/autopilot-arming.mjs +2 -1
- package/pipeline/scripts/autopilot-intake.mjs +2 -1
- package/pipeline/scripts/autopilot-runner.mjs +206 -2
- package/pipeline/scripts/build-references.mjs +2 -1
- package/pipeline/scripts/build-stack-plugins.mjs +10 -2
- package/pipeline/scripts/capture-evidence.sh +7 -2
- package/pipeline/scripts/capture-flush.sh +8 -8
- package/pipeline/scripts/capture-resume.sh +3 -3
- package/pipeline/scripts/classify-plan-safety.mjs +3 -2
- package/pipeline/scripts/cost-analyze.mjs +600 -0
- package/pipeline/scripts/cost-budget-check.mjs +4 -12
- package/pipeline/scripts/council-view.mjs +2 -1
- package/pipeline/scripts/crush-json.mjs +2 -1
- package/pipeline/scripts/diff-explain.mjs +7 -10
- package/pipeline/scripts/diff-risk-score.mjs +2 -1
- package/pipeline/scripts/doctor.mjs +140 -6
- package/pipeline/scripts/evidence-gate.mjs +9 -3
- package/pipeline/scripts/feedback-send.mjs +12 -2
- package/pipeline/scripts/gc-abandoned.sh +32 -16
- package/pipeline/scripts/gc-tmp.sh +1 -1
- package/pipeline/scripts/gc-worktrees.sh +12 -5
- package/pipeline/scripts/gen-facts.mjs +175 -0
- package/pipeline/scripts/gen-mode-dispatch.mjs +32 -37
- package/pipeline/scripts/gen-ref-toc.mjs +1 -1
- package/pipeline/scripts/github-ssh-setup.sh +64 -7
- package/pipeline/scripts/graph-mermaid.mjs +4 -2
- package/pipeline/scripts/graph-report.mjs +1 -1
- package/pipeline/scripts/jira-attach.sh +1 -1
- package/pipeline/scripts/keychain-save.sh +101 -30
- package/pipeline/scripts/learn-from-transcripts.mjs +3 -2
- package/pipeline/scripts/learning-curve.mjs +36 -31
- package/pipeline/scripts/log-metric.sh +17 -4
- package/pipeline/scripts/make-manifest.mjs +199 -0
- package/pipeline/scripts/memory-save.sh +1 -1
- package/pipeline/scripts/migrate-prefs.mjs +24 -6
- package/pipeline/scripts/migrate-state.mjs +94 -4
- package/pipeline/scripts/phase-banner.sh +26 -22
- package/pipeline/scripts/phase-tracker.sh +48 -10
- package/pipeline/scripts/plan-coverage-gate.mjs +8 -4
- package/pipeline/scripts/pre-commit-check.sh +7 -0
- package/pipeline/scripts/pre-push-check.sh +7 -0
- package/pipeline/scripts/purge.sh +23 -6
- package/pipeline/scripts/render-agent-log-cost.sh +10 -3
- package/pipeline/scripts/render-cost-summary.sh +9 -2
- package/pipeline/scripts/render-work-summary.sh +14 -7
- package/pipeline/scripts/review-file-filter.mjs +5 -3
- package/pipeline/scripts/review-scope.mjs +2 -1
- package/pipeline/scripts/routine-registry.mjs +2 -1
- package/pipeline/scripts/run-aggregator.mjs +26 -20
- package/pipeline/scripts/run-metrics.mjs +4 -2
- package/pipeline/scripts/runs-index.mjs +353 -0
- package/pipeline/scripts/scorecard-snapshot.mjs +178 -0
- package/pipeline/scripts/search-logs.sh +18 -0
- package/pipeline/scripts/smoke-cross-cli-behavior.sh +6 -6
- package/pipeline/scripts/smoke-schema-validation.sh +26 -7
- package/pipeline/scripts/test-gap-scan.mjs +2 -1
- package/pipeline/scripts/test-integrity-gate.mjs +2 -1
- package/pipeline/scripts/token-budget-report.mjs +13 -2
- package/pipeline/scripts/triage-memory.mjs +2 -2
- package/pipeline/scripts/update-issue-progress.sh +56 -7
- package/pipeline/scripts/usage-report.mjs +12 -1
- package/pipeline/scripts/validate-analysis-doc.mjs +75 -18
- package/pipeline/scripts/validate-code-graph.mjs +6 -3
- package/pipeline/scripts/validate-complaint-doc.mjs +2 -1
- package/pipeline/scripts/validate-diff-risk.mjs +6 -3
- package/pipeline/scripts/validate-planning.mjs +1 -1
- package/pipeline/scripts/validate-reviewer.mjs +1 -1
- package/pipeline/scripts/validate-state.mjs +45 -5
- package/pipeline/scripts/validate-test-gap.mjs +6 -3
- package/pipeline/scripts/validate-triage.mjs +6 -4
- package/pipeline/scripts/verify-citations.mjs +4 -2
- package/pipeline/scripts/verify.mjs +327 -0
- package/pipeline/scripts/worktree-finalize.sh +18 -9
- package/pipeline/scripts/write-state.mjs +154 -15
- package/pipeline/skills/.skill-manifest.json +37 -21
- package/pipeline/skills/.skills-index.json +104 -5
- package/pipeline/skills/shared/README.md +15 -6
- package/pipeline/skills/shared/core/apple-archive-compliance/SKILL.md +2 -2
- package/pipeline/skills/shared/core/google-play-compliance/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent/SKILL.md +69 -71
- package/pipeline/skills/shared/core/multi-agent-autopilot/SKILL.md +3 -3
- package/pipeline/skills/shared/core/multi-agent-channels/SKILL.md +14 -14
- package/pipeline/skills/shared/core/multi-agent-diff-explain/SKILL.md +5 -5
- package/pipeline/skills/shared/core/multi-agent-graph/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +25 -23
- package/pipeline/skills/shared/core/multi-agent-language/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-local/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +8 -8
- package/pipeline/skills/shared/core/multi-agent-manual-test/SKILL.md +6 -6
- package/pipeline/skills/shared/core/multi-agent-model/SKILL.md +71 -0
- package/pipeline/skills/shared/core/multi-agent-refactor/SKILL.md +3 -3
- package/pipeline/skills/shared/core/multi-agent-resume/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-resume-local/SKILL.md +7 -7
- package/pipeline/skills/shared/core/multi-agent-route-off/SKILL.md +39 -0
- package/pipeline/skills/shared/core/multi-agent-route-on/SKILL.md +76 -0
- package/pipeline/skills/shared/core/multi-agent-route-status/SKILL.md +59 -0
- package/pipeline/skills/shared/core/multi-agent-setup/SKILL.md +1 -1
- package/pipeline/skills/shared/core/multi-agent-status/SKILL.md +35 -11
- package/pipeline/skills/shared/core/multi-agent-steer/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +6 -5
- package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/package_app.sh +4 -1
- package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/setup_dev_signing.sh +4 -1
- package/pipeline/skills/shared/external/macos-spm-app-packaging/assets/templates/sign-and-notarize.sh +2 -1
- package/pipeline/skills/skills-index.md +13 -4
- package/pipeline/multi-agent-refs/phases/phase-1-analysis.md +0 -263
- package/pipeline/multi-agent-refs/phases/phase-2-planning.md +0 -344
- package/pipeline/multi-agent-refs/phases/phase-5-test.md +0 -182
package/docs/features.md
CHANGED
|
@@ -4,18 +4,19 @@ Comprehensive list of every feature the pipeline ships. The top-level `README.md
|
|
|
4
4
|
|
|
5
5
|
## Core Pipeline
|
|
6
6
|
|
|
7
|
-
###
|
|
7
|
+
### 6-Phase Orchestration (0-5)
|
|
8
8
|
|
|
9
9
|
```
|
|
10
10
|
Phase 0: Init Project selection, branch setup, identity, worktree
|
|
11
|
-
Phase 1:
|
|
12
|
-
|
|
13
|
-
Phase
|
|
14
|
-
|
|
11
|
+
Phase 1: Plan Stack detection, codebase exploration (parallel Explore agents),
|
|
12
|
+
task decomposition, architecture review, user approval
|
|
13
|
+
Phase 2: Dev TDD cycle: test → code → build (Sonnet), then the Verify exit
|
|
14
|
+
gate: build · lint · tests · secrets, run once
|
|
15
|
+
Phase 3: Review Parallel AI review + Fable triage against Dev's logs, then the
|
|
16
|
+
optional user test + on-demand device audits
|
|
15
17
|
(Claude Code: Fable + Opus + Sonnet · Copilot CLI: GPT-5.4 + Opus + Sonnet)
|
|
16
|
-
Phase
|
|
17
|
-
Phase
|
|
18
|
-
Phase 7: Report External: Jira comment · Wiki + Figma screenshots · Confluence
|
|
18
|
+
Phase 4: Commit Git commit, push, PR with default reviewers + draft/ready prompt
|
|
19
|
+
Phase 5: Report External: Jira comment · Wiki + Figma screenshots · Confluence
|
|
19
20
|
Internal: agent-log.md + Quality & Metrics + knowledge + memory
|
|
20
21
|
```
|
|
21
22
|
|
|
@@ -28,7 +29,7 @@ Each phase reads its own spec file under `pipeline/multi-agent-refs/phases/phase
|
|
|
28
29
|
| `autopilot` | Skip all confirmation prompts; still fails safe on review blockers + build retries. |
|
|
29
30
|
| `--local` | Answers the workspace question up front: no worktree, work directly in `$PROJECT_ROOT` on a local branch. |
|
|
30
31
|
|
|
31
|
-
Neither the workspace nor the depth is a flag. **Where the branch lives** is asked at Phase 0 Step 5b - worktree (`.worktrees/{id}/`, your checkout untouched) or local (the project root, which drops
|
|
32
|
+
Neither the workspace nor the depth is a flag. **Where the branch lives** is asked at Phase 0 Step 5b - worktree (`.worktrees/{id}/`, your checkout untouched) or local (the project root, which drops the user test because that gate checks the change out of a worktree and there is none). `:local` and `--local` state it in advance; every autopilot entry resolves it to a worktree and never asks, because an unattended run commits and pushes from wherever it stands and doing that in the user's own checkout is what worktrees exist to prevent. `state.workspaceSource` records who decided - `localMode: false` alone is both "the user chose a worktree" and "nothing asked".
|
|
32
33
|
|
|
33
34
|
Depth is not a flag either. `/multi-agent` and `/multi-agent:local` ask Full or Short at Phase 0 Step 7.5, recommending from the detected `taskType`; Short strips to Init → Dev(Opus self-contained) → Review → Test → Commit → Report. Autopilot never asks and always runs Full - "fast plus unattended" was removed in v16.0.0, because something has to choose when nobody is asked and unattended is the worst place to drop analysis and planning.
|
|
34
35
|
|
|
@@ -46,7 +47,7 @@ Uninstall preserves the whole layer - tokens, the reader that opens them, the ma
|
|
|
46
47
|
|
|
47
48
|
A deterministic, LLM-free map of what a repo declares and what refers to what, extracted by regex over comment-stripped source into `~/.claude/knowledge/<project>/code-graph.json`. Four stacks build today (Swift, Kotlin/Java, TypeScript/JavaScript, Python); each is one rules file, and the engine is the same for all of them. Zero runtime dependencies, zero API cost, read-only on the repo.
|
|
48
49
|
|
|
49
|
-
Phase 1 queries it to hand Explore a ranked starting file set instead of a full scan, and Phase
|
|
50
|
+
Phase 1 queries it to hand Explore a ranked starting file set instead of a full scan, and Phase 5 rebuilds it after the branch changed code - a rebuild is seconds, so staleness is a `baseCommit` comparison rather than a date heuristic. Off by default behind `prefs.global.codeGraph.enabled`; with it off the pipeline behaves exactly as before.
|
|
50
51
|
|
|
51
52
|
`GRAPH_REPORT.md` also ends with **Symbols nothing else references**: symbols no other file in the repo names, split from the ones referenced only by their own tests. Candidates, never verdicts - the extractor is regex, not a parser, so the four classes that could not carry a reference edge either way (a kind outside the stack's `referenceKinds`, a name declared twice, a nested declaration, a test file) are counted and excluded rather than listed, and nothing gates on the result.
|
|
52
53
|
|
|
@@ -78,9 +79,9 @@ Stack skill sets ship as versioned plugins in the `multi-agent-plugins` marketpl
|
|
|
78
79
|
/multi-agent:stack all # every stack plugin
|
|
79
80
|
```
|
|
80
81
|
|
|
81
|
-
### Package Manager Resolution (Phase
|
|
82
|
+
### Package Manager Resolution (Phase 2, node-shaped stacks)
|
|
82
83
|
|
|
83
|
-
Phase
|
|
84
|
+
Phase 2's web test arm and its build step used to type `npm`. A repo on pnpm, yarn or bun then failed in Phase 2 - with a worktree and a branch already created - or, worse, npm resolved against a lock file it does not own and the run continued on a tree the repo's own tooling would never have produced.
|
|
84
85
|
|
|
85
86
|
`scripts/package-manager.mjs` resolves it from the repo instead: `$MA_PACKAGE_MANAGER`, then `package.json#packageManager`, then a lock file, then npm - reported AS a default, never as evidence, because "npm because nothing said otherwise" and "npm because the repo committed a package-lock" are different answers. Node core only (ADR-0004): a resolver that shelled out would need a working install of the tool it is identifying. The walk goes up to the directory holding `.git` and stops there, so a monorepo's root lock file is found and a stray one in a home directory is not. Two lock files means a migration left one behind: the newest wins and both are named.
|
|
86
87
|
|
|
@@ -121,15 +122,15 @@ Result persisted to `agent-state.taskType`:
|
|
|
121
122
|
|
|
122
123
|
| Type | Downstream effects |
|
|
123
124
|
| ----------- | ----------------------------------------------------------------------------- |
|
|
124
|
-
| `component` | Phase
|
|
125
|
-
| `bugfix` | Phase
|
|
126
|
-
| `feature` | Standard TDD flow; Phase
|
|
127
|
-
| `refactor` | Phase
|
|
128
|
-
| `chore` | Lightweight flow; Phase
|
|
125
|
+
| `component` | Phase 2 dispatches to the marketplace component plugin (create-component) with SubPhase reporting |
|
|
126
|
+
| `bugfix` | Phase 3 emphasizes test coverage + regression; Phase 4 uses `fix(...)` prefix |
|
|
127
|
+
| `feature` | Standard TDD flow; Phase 4 uses `feat(...)` prefix |
|
|
128
|
+
| `refactor` | Phase 3 emphasizes behavior preservation; Phase 4 uses `refactor(...)` prefix |
|
|
129
|
+
| `chore` | Lightweight flow; Phase 4 uses `chore(...)` prefix |
|
|
129
130
|
|
|
130
131
|
### SubPhase Convention
|
|
131
132
|
|
|
132
|
-
When a specialized skill takes over a main pipeline phase, progress is reported as SubPhases (e.g. `SubPhase 3.0: Init`, `SubPhase 3.1: Gather`). The top-level pipeline stays fixed at
|
|
133
|
+
When a specialized skill takes over a main pipeline phase, progress is reported as SubPhases (e.g. `SubPhase 3.0: Init`, `SubPhase 3.1: Gather`). The top-level pipeline stays fixed at 6 phases (0-5) - specialized work slots into its parent phase without inflating the count.
|
|
133
134
|
|
|
134
135
|
## PR & Review Flow
|
|
135
136
|
|
|
@@ -141,14 +142,14 @@ When a specialized skill takes over a main pipeline phase, progress is reported
|
|
|
141
142
|
|
|
142
143
|
### Draft vs Ready Prompt
|
|
143
144
|
|
|
144
|
-
Phase
|
|
145
|
+
Phase 4 asks `DRAFT or READY?` before creating the PR and persists the choice in `prefs.projects[].defaultPrMode`.
|
|
145
146
|
|
|
146
147
|
- Bitbucket: `draft: true` flag (DC 8.x+) with `[DRAFT]` title fallback for older servers.
|
|
147
148
|
- GitHub: `gh pr create --draft` + `gh pr ready` for promotion.
|
|
148
149
|
|
|
149
150
|
### `channels` Command
|
|
150
151
|
|
|
151
|
-
Multi-channel reporter - Phase
|
|
152
|
+
Multi-channel reporter - Phase 5 delegates to it, and it's also invocable post-hoc for fixes closed outside the pipeline:
|
|
152
153
|
|
|
153
154
|
```bash
|
|
154
155
|
/multi-agent:channels # current branch, current PR
|
|
@@ -170,7 +171,7 @@ Never auto-closes issues - uses `Ref: #N` / `Related: #N` / `See: PROJ-12345`, n
|
|
|
170
171
|
|
|
171
172
|
## Review Quality
|
|
172
173
|
|
|
173
|
-
### Deterministic Gates (Phase
|
|
174
|
+
### Deterministic Gates (Phase 3 Step 1)
|
|
174
175
|
|
|
175
176
|
Cheap, objective checks run BEFORE any AI token is spent:
|
|
176
177
|
|
|
@@ -181,23 +182,23 @@ Cheap, objective checks run BEFORE any AI token is spent:
|
|
|
181
182
|
|
|
182
183
|
If any gate fails, fix first. Don't waste AI tokens reviewing broken code.
|
|
183
184
|
|
|
184
|
-
### Analysis Document Review (Phase
|
|
185
|
+
### Analysis Document Review (Phase 2.2 + 2.3)
|
|
185
186
|
|
|
186
187
|
`/multi-agent:analysis` published behind a structural validator alone until v16.12.0: nothing read the
|
|
187
|
-
document before it reached Confluence. Phase
|
|
188
|
+
document before it reached Confluence. Phase 2.2 now runs the same reviewer set and triage a code diff
|
|
188
189
|
gets, on the draft, before the destination is even chosen. Its first question is what the run skipped -
|
|
189
190
|
an input declared missing that nothing searched for, an open question about evidence nobody read, a gap
|
|
190
191
|
with no owner, a scope call made without asking. A blocking finding returns to synthesis with dispatch
|
|
191
192
|
closed; it never becomes an open question, because "the document is wrong" is not something to ask the
|
|
192
193
|
reader.
|
|
193
194
|
|
|
194
|
-
Phase
|
|
195
|
+
Phase 2.3 then sorts what is left: reachable evidence is searched (never asked about), decisions the
|
|
195
196
|
user owns are asked with `AskUserQuestion`, and only genuinely external gaps enter the document as
|
|
196
197
|
`AS-NN` rows with an owner. A gap carrying neither a `searched, not found` nor an `asked, external`
|
|
197
198
|
stamp fails the dispatch gate. Autopilot runs both phases; only the asking degrades, into rows stamped
|
|
198
199
|
`autopilot: could not ask`.
|
|
199
200
|
|
|
200
|
-
### CLI-Aware Parallel Review + Fable Triage (Phase
|
|
201
|
+
### CLI-Aware Parallel Review + Fable Triage (Phase 3 Steps 2-3)
|
|
201
202
|
|
|
202
203
|
| Reviewer | Model | Focus | Where it runs |
|
|
203
204
|
| ---------- | ------------------- | --------------------------------- | -------------------- |
|
|
@@ -207,7 +208,7 @@ stamp fails the dispatch gate. Autopilot runs both phases; only the asking degra
|
|
|
207
208
|
|
|
208
209
|
The reviewer set is **CLI-aware**: Claude Code dispatches 3 reviewers in parallel (Fable + Opus + Sonnet - Opus fills the slot GPT-5.4 takes elsewhere); Copilot CLI dispatches all 3. Each returns structured JSON for deterministic aggregation. Cross-model diversity catches blind spots that any single model family would miss.
|
|
209
210
|
|
|
210
|
-
**Fable Triage** (Phase
|
|
211
|
+
**Fable Triage** (Phase 3 Step 3, Opus on Copilot CLI): Evaluates merged raw findings against task scope. Classifies each as `accepted` (fix now), `deferred` (out of scope, log for later), or `rejected` (false positive / noise). Only triage-accepted blocking items loop back to Phase 3.
|
|
211
212
|
|
|
212
213
|
### Runtime Triage Validator
|
|
213
214
|
|
|
@@ -224,13 +225,13 @@ After triage returns, output is validated by `validate-triage.mjs`:
|
|
|
224
225
|
|
|
225
226
|
If triage returns `approved: false` but has no blocking items, the validator forces `approved: true`. Conversely, if `approved: true` but blocking items exist, it forces `approved: false`. Hardened with an `if`/`then` constraint in the schema itself.
|
|
226
227
|
|
|
227
|
-
### Verify-by-Test Triage (Phase
|
|
228
|
+
### Verify-by-Test Triage (Phase 3 Step 3.7, opt-in)
|
|
228
229
|
|
|
229
|
-
A triage verdict is a judgment call; a failing repro test is proof. When `prefs.global.verifyByTest.enabled` is on, one verifier agent (default Sonnet) writes a minimal repro test per accepted blocking finding (cap: `maxFindings`=3) and runs only that test. Fails as predicted -> finding confirmed, the repro test becomes the Phase
|
|
230
|
+
A triage verdict is a judgment call; a failing repro test is proof. When `prefs.global.verifyByTest.enabled` is on, one verifier agent (default Sonnet) writes a minimal repro test per accepted blocking finding (cap: `maxFindings`=3) and runs only that test. Fails as predicted -> finding confirmed, the repro test becomes the Phase 2 rework RED test. Passes under `evidence-gate.mjs` -> finding downgraded to `deferred`. Compile error / timeout -> `inconclusive`, judgment stands. Timeout-bounded, never blocks. Full spec: `refs/features/verify-by-test.md`.
|
|
230
231
|
|
|
231
232
|
### Immutable-Test Rule + `test_lines_removed` Signal
|
|
232
233
|
|
|
233
|
-
Existing tests are immutable during a task: deleting, renaming, or weakening an assertion to reach green is a violation (`refs/rules.md`, Phase
|
|
234
|
+
Existing tests are immutable during a task: deleting, renaming, or weakening an assertion to reach green is a violation (`refs/rules.md`, Phase 2 GREEN step). A test changes only when the task changes the spec it encodes, named in the commit body. Deterministic backstop: `diff-risk-score.mjs` emits `test_lines_removed` (w=3.0) for any test-classified file whose diff removes more lines than it adds.
|
|
234
235
|
|
|
235
236
|
### Update Check at Run Start
|
|
236
237
|
|
|
@@ -244,7 +245,7 @@ Phase 0 Step 0.6. Once per `ttlHours` window (cached, 3s-bounded curl to the npm
|
|
|
244
245
|
|
|
245
246
|
Every phase transition appends a `## Handoff` block (Done / Remaining / Decisions / Open findings / Next) to `agent-log.md` - orchestrator-written from existing state, no LLM call. `/multi-agent:resume` and post-`/compact` re-grounding read the latest handoff first, so long runs re-enter from durable artifacts instead of conversation memory (fresh-context discipline from Anthropic's long-running-agent harness guidance).
|
|
246
247
|
|
|
247
|
-
### Accessibility Code Review (Phase
|
|
248
|
+
### Accessibility Code Review (Phase 3 Step 1.5)
|
|
248
249
|
|
|
249
250
|
If changes include UI files, reviewers check for:
|
|
250
251
|
|
|
@@ -256,16 +257,16 @@ Pure code analysis - no simulator needed. Device-level audits run in Phase 5 whe
|
|
|
256
257
|
|
|
257
258
|
### Status Enforcement
|
|
258
259
|
|
|
259
|
-
Phase
|
|
260
|
+
Phase 2 treats the issue-tracker status update as a required step with a post-mutation verify step that re-reads the field and retries once on silent `VALIDATION` failures (e.g. stale Projects V2 option IDs after a board rebuild).
|
|
260
261
|
|
|
261
262
|
## Safety & Hygiene
|
|
262
263
|
|
|
263
264
|
- **Pre-Commit Secret Detection** (12 patterns): `PreToolUse` hook scans staged files for API keys/tokens, AWS access keys, private keys, `.env` files, service account JSON. Commit **blocked** if found.
|
|
264
265
|
- **Read-Size Gate** (opt-in, `prefs.global.bulkRead.mode`): a `PreToolUse` hook inspects `Read` and the shell commands that read a file whole. In `observe` it only logs what it would have caught - the baseline you measure before routing anything. In `enforce` a file over `minLines` (default 350) is blocked and delegated to a haiku-rung worker (`bulk-read.sh`), which returns a line-numbered summary so the follow-up is a bounded `Read(offset:limit:)` instead of the whole file; the full text is parked under `.multi-agent/refs/`. The development phase and any file the run has already touched are exempt, because Claude Code's `Edit` requires its own `Read` first.
|
|
265
|
-
- **Capture Hooks** (`SessionEnd`, `PreCompact`, `SessionStart`): every durable write used to live in Phase
|
|
266
|
+
- **Capture Hooks** (`SessionEnd`, `PreCompact`, `SessionStart`): every durable write used to live in Phase 5, the phase a run is least likely to reach. `SessionEnd` flushes a run that never got there; `PreCompact` flushes before an auto-compaction summarizes a long phase mid-flight, which is the same loss one level down; `SessionStart` prints at most two lines about an unfinished run. None calls a model, none reads a payload, and all exit 0 on every path - a hook that fails a session over bookkeeping is worse than the bookkeeping.
|
|
266
267
|
- **Operational Reporting** (`prefs.global.usageLog`): coarse run metadata - task id, phase, status, durations, token counts - and never prompts, code, diffs or absolute paths. The per-machine token is REQUESTED from the endpoint by `usage-register.mjs` (setup, update, and the Phase 0 exit gate as a backstop), is write-only, and lives in the OS credential store; prefs hold only the entry name and the switch. `usageLog.optOut: true` blocks registration permanently and is checked before the network call. An unreachable endpoint leaves reporting off with one line and exit 0 - a run is never failed over bookkeeping.
|
|
267
268
|
- **Build Queue**: All `xcodebuild` calls acquire a lock. Each worktree uses own `-derivedDataPath`. Stale locks auto-clean after 15 min. Non-Xcode builds don't need the lock.
|
|
268
|
-
- **Context Management**: `CLAUDE_AUTOCOMPACT_PCT_OVERRIDE=65` - compaction at 65% usage (prevents degradation in
|
|
269
|
+
- **Context Management**: `CLAUDE_AUTOCOMPACT_PCT_OVERRIDE=65` - compaction at 65% usage (prevents degradation in 6-phase sessions).
|
|
269
270
|
- **3-Iteration Hard Kill**: Any retry loop stops after 3 attempts, then pauses for user. No infinite loops.
|
|
270
271
|
|
|
271
272
|
## Testing & Quality
|
|
@@ -300,51 +301,51 @@ Per-phase token budgets prevent runaway sessions. If a phase exceeds its budget,
|
|
|
300
301
|
- **Cost Telemetry**: Per-phase token cost tracking (`tokens_in`, `tokens_out`, `model`, `duration_ms`). Omitted fields handled gracefully.
|
|
301
302
|
- **Phase Tracker**: Cross-CLI visual progress (current phase, elapsed time, iteration count).
|
|
302
303
|
- **Phase Banner**: Terminal UI for phase transitions with Unicode box-drawing characters.
|
|
303
|
-
- **Per-task Cost Breakdown in agent-log.md**: Phase
|
|
304
|
+
- **Per-task Cost Breakdown in agent-log.md**: Phase 5 appends a 4-column block (Phase · Model · Tokens in/out · Est. USD) to every run's `agent-log.md`. Sourced from `phase-tracker.sh tokens` accumulators × `cost-table.json` prices. Independent of the channels-side `reportContent.costSummary` toggle. The `LOG_METRIC_FORWARD_TO_TRACKER=1` env flag mirrors `tokens_in`/`tokens_out`/`model` from `log-metric.sh` into the tracker so JSONL metrics and the cost block stay in sync from one call site.
|
|
304
305
|
|
|
305
306
|
### Diff Risk Scoring
|
|
306
307
|
|
|
307
|
-
`pipeline/scripts/diff-risk-score.mjs` runs at Phase
|
|
308
|
+
`pipeline/scripts/diff-risk-score.mjs` runs at Phase 3 Step 1.75 - before reviewer dispatch. Heuristic, deterministic, sub-second, no LLM. Top-N risk-ranked files inject into each reviewer's prompt as a `${PRIORITY_FILES}` block; reviewers read those files first but still review the entire diff.
|
|
308
309
|
|
|
309
310
|
Signals + weights: `security_path` ×3, `migration` ×4, `public_api` ×2, `no_test_change` ×2.5, `test_lines_removed` ×3 (test file shrinks - immutable-test backstop), `complexity_delta` ×1.5, `ui_critical` ×1.5, `loc_changed` ×1. Toggle via `prefs.global.diffRiskAdvisory` (default ON).
|
|
310
311
|
|
|
311
312
|
### Test Gap Detection
|
|
312
313
|
|
|
313
|
-
`pipeline/scripts/test-gap-scan.mjs` runs at Phase
|
|
314
|
+
`pipeline/scripts/test-gap-scan.mjs` runs at Phase 3 Step 0. Walks the diff for newly added public symbols and reports those with no paired test. Stack-specific rules ship for iOS, Android, Python, Node.js. iOS Views and Android `@Composable` symbols default to `important`; other public API additions to `suggestion`. Optional gating via `prefs.testGap.blockingThreshold` - when set, the report becomes a Phase 4 rework finding once `important + blocking` count exceeds the threshold.
|
|
314
315
|
|
|
315
316
|
### Visual Evidence (UI changes)
|
|
316
317
|
|
|
317
318
|
A UI change carries its own picture. `state.visualEvidence.required` is decided mechanically from `taskType` plus the changed-file list, never from a reading of the task.
|
|
318
319
|
|
|
319
|
-
**Stills.** The "before" is the reporter's own ticket attachment, harvested in Phase 0; the pipeline never rebuilds the old state to photograph it. The "after" is captured in Phase
|
|
320
|
+
**Stills.** The "before" is the reporter's own ticket attachment, harvested in Phase 0; the pipeline never rebuilds the old state to photograph it. The "after" is captured in Phase 2 right after the build goes green, not in the user test, which autopilot and both local modes drop. `capture-evidence.sh` cleans the status bar and downscales to 1242px so two captures of one screen differ by the change and not by the clock.
|
|
320
321
|
|
|
321
|
-
**The flow video rides on a test run.** `probe-evidence-capability.sh` measures the UI test target, the tests matching this change, the device, the recorder and the MCP registration; Phase 0 Step 7.7 then asks the depth with the options built from that measurement, and a closed option keeps its row and states why. Tier 1 runs the repo's own UI test and records around it, tier 2 drives the flow through `agent_run_steps`, tier 3 records nothing and says so. The tier is re-checked before the recording starts, because a simulator booted at intake can be gone by Phase
|
|
322
|
+
**The flow video rides on a test run.** `probe-evidence-capability.sh` measures the UI test target, the tests matching this change, the device, the recorder and the MCP registration; Phase 0 Step 7.7 then asks the depth with the options built from that measurement, and a closed option keeps its row and states why. Tier 1 runs the repo's own UI test and records around it, tier 2 drives the flow through `agent_run_steps`, tier 3 records nothing and says so. The tier is re-checked before the recording starts, because a simulator booted at intake can be gone by Phase 2.
|
|
322
323
|
|
|
323
324
|
UI test detection keys on `XCUIApplication` rather than on a folder named `*UITests`: in a real app the overwhelming majority of files under such a path are snapshot tests, which never launch the app and would produce a still frame filed as a flow.
|
|
324
325
|
|
|
325
|
-
**Where it lands.** Jira takes both stills and video as attachments. With no Jira the stills go to an orphan `evidence/<task-id>` branch and the PR body embeds them, or links them with a blob permalink when the repo is private (GitHub's image proxy has no credentials for a private repo, and a broken image reads as missing evidence). Phase
|
|
326
|
+
**Where it lands.** Jira takes both stills and video as attachments. With no Jira the stills go to an orphan `evidence/<task-id>` branch and the PR body embeds them, or links them with a blob permalink when the repo is private (GitHub's image proxy has no credentials for a private repo, and a broken image reads as missing evidence). Phase 4 blocks when a required artefact is neither published nor explained; the gate is against silence, not against an honest "the ticket carries no image".
|
|
326
327
|
|
|
327
328
|
Toggle via `prefs.global.visualEvidence.enabled` (default ON), `visualEvidence.githubHost`, `visualEvidence.maxAttachmentMb`, `visualEvidence.maxVideoSeconds`, `prefs.global.testDepth.default`.
|
|
328
329
|
|
|
329
330
|
### Triage Memory
|
|
330
331
|
|
|
331
|
-
Per-repo append-only JSONL corpus at `~/.claude/memory/multi-agent/<repo-slug>/triage-corpus.jsonl`. Phase
|
|
332
|
+
Per-repo append-only JSONL corpus at `~/.claude/memory/multi-agent/<repo-slug>/triage-corpus.jsonl`. Phase 5 ingests every triage output (idempotent), Phase 1 enriches the analysis with similar past tasks, Phase 3 triage attaches prior-art hits to each raw finding with an explicit bias hedge. Token-overlap recall, zero deps, Node-18-compatible. `/multi-agent:search "<text>" --semantic` routes the query to the corpus instead of agent-log grep. Toggle via `prefs.global.priorArtEnrichment.enabled` (default ON).
|
|
332
333
|
|
|
333
334
|
## Learning
|
|
334
335
|
|
|
335
336
|
### Knowledge Base (per project)
|
|
336
337
|
|
|
337
|
-
Incremental learning. Phase
|
|
338
|
+
Incremental learning. Phase 5 captures architecture, patterns, gotchas, and decisions into `$HOME/.claude/knowledge/{project}/`. Phase 1 reads it on the next run. Token cost decreases over time as the base grows.
|
|
338
339
|
|
|
339
340
|
### Memory Capture (cross-session)
|
|
340
341
|
|
|
341
|
-
Pipeline learns behavioral signals (feedback corrections, project constraints, external references). Phase
|
|
342
|
+
Pipeline learns behavioral signals (feedback corrections, project constraints, external references). Phase 5 saves, Phase 1 injects. Max 3 new memories per run. Merge-over-duplicate. Stale memories verified before use.
|
|
342
343
|
|
|
343
344
|
**What does NOT go in memory**: architecture, code patterns, build gotchas, design decisions - those belong in the knowledge base.
|
|
344
345
|
|
|
345
346
|
### Lesson Diagnosis (Reflexion)
|
|
346
347
|
|
|
347
|
-
Phase
|
|
348
|
+
Phase 3's lesson-memory loop records the causal root cause of each fix (`--diagnosis`), not just the outcome: the verbal "why" that prevents recurrence (Reflexion). `learnings-ledger.mjs brief` renders it as `(why: ...)` back into Phase 1 + triage on the next run, so the reason re-enters the loop, not only the symptom.
|
|
348
349
|
|
|
349
350
|
### Corpus Freshness Gate
|
|
350
351
|
|
|
@@ -366,7 +367,7 @@ Turn a recurring, project-specific job into a first-class `/multi-agent:<name>`
|
|
|
366
367
|
|
|
367
368
|
### Figma / Component Generation (dispatched to marketplace plugins)
|
|
368
369
|
|
|
369
|
-
Component + Figma-to-code work is no longer bundled in this repo. When Phase 0 classifies a task as `component`, Phase
|
|
370
|
+
Component + Figma-to-code work is no longer bundled in this repo. When Phase 0 classifies a task as `component`, Phase 2 dispatches it to the per-stack marketplace plugins (`ai-ios-toolkit` / `ai-android-toolkit` in the `multi-agent-plugins` marketplace) via the Skill tool. The plugin's component skill generates `{Name}Configuration.swift`, `{Name}View.swift`, `{Name}+Modifiers.swift`, `{Name}.figma.swift`, and `FIGMA.md` with a variant matrix, then runs a 14-item pre-commit checklist covering design tokens, accessibility, tests, and Code Connect.
|
|
370
371
|
|
|
371
372
|
The plugin's cross-cutting integration skills feed component detection + implementation when the design triggers them (content: form / price / ui-patterns; interaction: navigation / overlays / bottom-sheets). Each is native-SwiftUI-first and reads project specifics (token namespaces, component paths, UI systems) from `figma-config`, including the optional `ui.navigationSystem` / `ui.overlaySystem` / `ui.sheetSystem` hooks (absent -> stock SwiftUI), so the same capabilities work on any SwiftUI codebase. The plugin's evolve-component skill reconciles an existing component against current Figma (drift-heal) and additively extends it, behind a human gate.
|
|
372
373
|
|
|
@@ -376,20 +377,20 @@ Automated visual testing and compliance audits via direct Bash (no MCP server de
|
|
|
376
377
|
|
|
377
378
|
| Audit | When | Command |
|
|
378
379
|
| --------------------- | ------------------- | ---------------------------- |
|
|
379
|
-
| iOS Accessibility | Phase
|
|
380
|
-
| Android Accessibility | Phase
|
|
381
|
-
| iOS Biometric | Phase
|
|
382
|
-
| Android Launch Time | Phase
|
|
383
|
-
| iOS Archive | Phase
|
|
384
|
-
| Android APK | Phase
|
|
380
|
+
| iOS Accessibility | Phase 3, on request | `swift ui-tree-dumper.swift` |
|
|
381
|
+
| Android Accessibility | Phase 3, on request | `adb shell uiautomator dump` |
|
|
382
|
+
| iOS Biometric | Phase 3, auth flow | `xcrun simctl keychain` |
|
|
383
|
+
| Android Launch Time | Phase 3, perf | `adb shell am start -W` |
|
|
384
|
+
| iOS Archive | Phase 4, release | `codesign`, `plutil`, `nm` |
|
|
385
|
+
| Android APK | Phase 4, release | `aapt2`, `apksigner` |
|
|
385
386
|
|
|
386
387
|
Audits are **on-demand** - triggered by user, never automatic.
|
|
387
388
|
|
|
388
389
|
### Jira + Confluence
|
|
389
390
|
|
|
390
|
-
- Phase
|
|
391
|
-
- Phase
|
|
392
|
-
- Phase
|
|
391
|
+
- Phase 2: transition issue to `In Progress` (verified post-mutation).
|
|
392
|
+
- Phase 5: post analysis + test scenarios as Jira comment (Turkish by default, configurable).
|
|
393
|
+
- Phase 5 (optional): create Confluence page under chosen parent, cached per project.
|
|
393
394
|
|
|
394
395
|
### Keychain
|
|
395
396
|
|
package/docs/performance.md
CHANGED
|
@@ -13,7 +13,7 @@ node pipeline/scripts/aggregate-metrics.mjs
|
|
|
13
13
|
# Markdown table (for PR descriptions, wikis, dashboards)
|
|
14
14
|
node pipeline/scripts/aggregate-metrics.mjs --markdown
|
|
15
15
|
|
|
16
|
-
# JSON (for machine consumption - Phase
|
|
16
|
+
# JSON (for machine consumption - Phase 5 report uses this)
|
|
17
17
|
node pipeline/scripts/aggregate-metrics.mjs --json
|
|
18
18
|
|
|
19
19
|
# Filtered - only recent runs
|
|
@@ -77,7 +77,7 @@ _Source: ~/.claude/logs/multi-agent/metrics.jsonl · Events: 421 (0 parse errors
|
|
|
77
77
|
- **`cycles per task avg` > 2.0** - triage is rejecting too many real findings
|
|
78
78
|
or Phase 3 isn't converging. Inspect the edge-cases table.
|
|
79
79
|
- **Most-common edge case = `over_rejection_guard_tripped`** - the triage
|
|
80
|
-
prompt lost scope context. Look at `phase-
|
|
80
|
+
prompt lost scope context. Look at `phase-3-review.md:57-91`.
|
|
81
81
|
- **`p95` much higher than `avg`** - a few tasks are looping 3+ times. Usually
|
|
82
82
|
means one of: bad acceptance criteria in Phase 2, tests that flake, or
|
|
83
83
|
environment-dependent build failures.
|
|
@@ -103,11 +103,11 @@ Current totals (v3.5.0):
|
|
|
103
103
|
Lazy loading keeps these off the model's context until each phase actually
|
|
104
104
|
runs - the full 14 k total is never loaded at once.
|
|
105
105
|
|
|
106
|
-
## Embedding Metrics in Phase
|
|
106
|
+
## Embedding Metrics in Phase 5 Reports
|
|
107
107
|
|
|
108
|
-
Phase
|
|
108
|
+
Phase 5 automatically calls `aggregate-metrics.mjs --json` with the current
|
|
109
109
|
task id and embeds the last 30 days' summary into the report body. See the
|
|
110
|
-
template in `phase-
|
|
110
|
+
template in `phase-5-report.md`.
|
|
111
111
|
|
|
112
112
|
## Disabling Metrics
|
|
113
113
|
|
package/docs/recovery-guide.md
CHANGED
|
@@ -10,8 +10,8 @@ the answer across `modes.md`, `operations.md`, and the individual phase specs.
|
|
|
10
10
|
| ------------------------------------------------------- | ----------------------------------------------- |
|
|
11
11
|
| Pipeline paused/halted mid-run | [Resume a paused task](#resume-a-paused-task) |
|
|
12
12
|
| Phase 3 build failed > 3 times | [Build retry exhausted](#build-retry-exhausted) |
|
|
13
|
-
| Phase
|
|
14
|
-
| Phase
|
|
13
|
+
| Phase 3 triage returned exit 1 (invalid JSON) twice | [Triage fallback](#triage-fallback) |
|
|
14
|
+
| Phase 3 triage returned exit 2 (over-rejection) | [Over-rejection](#over-rejection-guard) |
|
|
15
15
|
| Worktree already exists / dirty | [Worktree collisions](#worktree-collisions) |
|
|
16
16
|
| `agent-state.json` corrupt or unreadable | [State corruption](#state-corruption) |
|
|
17
17
|
| Wrong git identity committed | [Identity rewind](#identity-rewind) |
|
|
@@ -58,7 +58,7 @@ Path forward:
|
|
|
58
58
|
environment and `resume` - the 3-retry counter resets.
|
|
59
59
|
3. If the failure is logic (compile error in the generated code), edit the
|
|
60
60
|
offending file in the worktree, then `resume` - Phase 3 re-runs build.
|
|
61
|
-
4. If the task itself is wrong-shaped (Phase
|
|
61
|
+
4. If the task itself is wrong-shaped (Phase 1 plan is infeasible), `kill` and
|
|
62
62
|
restart with a better-scoped prompt.
|
|
63
63
|
|
|
64
64
|
---
|
|
@@ -79,7 +79,7 @@ If exit 1 fires twice (fallback path):
|
|
|
79
79
|
1. The pipeline logs `Phase 4: triage failed twice - fallback, all findings accepted as blocking`.
|
|
80
80
|
2. All raw findings loop back into Phase 3 as if they were all real blockers.
|
|
81
81
|
3. This is intentionally conservative - we'd rather over-fix than skip
|
|
82
|
-
something real. You can manually mark noise in the Phase
|
|
82
|
+
something real. You can manually mark noise in the Phase 4 PR description.
|
|
83
83
|
|
|
84
84
|
## Over-Rejection Guard
|
|
85
85
|
|
|
@@ -216,7 +216,7 @@ cat .worktrees/{id}/agent-state.json | grep status
|
|
|
216
216
|
## Instruction Fallback
|
|
217
217
|
|
|
218
218
|
`state.instructionDriven=true` but the file at `state.instructionFiles.commit`
|
|
219
|
-
is missing on disk. Phase
|
|
219
|
+
is missing on disk. Phase 4 logs an error, sets
|
|
220
220
|
`state.instructionDrivenFallback=true`, and uses the standard commit path.
|
|
221
221
|
|
|
222
222
|
This is rarely fatal - usually the instruction file was removed between
|
|
@@ -320,7 +320,7 @@ cp "$TRACKER_STATE" "$TRACKER_STATE.bak.$(date +%s)"
|
|
|
320
320
|
# 3. Either restore from a recent valid snapshot in the same dir, or
|
|
321
321
|
# reinitialize from the agent-log timeline:
|
|
322
322
|
bash $HOME/.claude/scripts/phase-tracker.sh init "<task-id>"
|
|
323
|
-
for p in 0:Init 1:
|
|
323
|
+
for p in 0:Init 1:Plan 2:Dev 3:Review 4:Commit 5:Report; do
|
|
324
324
|
bash $HOME/.claude/scripts/phase-tracker.sh add "${p%%:*}" "${p#*:}"
|
|
325
325
|
done
|
|
326
326
|
# Replay completed phases from agent-log.md headers
|
|
@@ -338,7 +338,7 @@ reinit. A footnote in `agent-log.md` documenting the rebuild is good practice.
|
|
|
338
338
|
## Token Rotation Mid-Pipeline (v8.0.0)
|
|
339
339
|
|
|
340
340
|
A token (Jira / Bitbucket / GitHub / Vercel) was rotated while a long-running
|
|
341
|
-
pipeline was active. Phase
|
|
341
|
+
pipeline was active. Phase 4 / Phase 5 then fail with `401 Unauthorized` even
|
|
342
342
|
though earlier phases worked.
|
|
343
343
|
|
|
344
344
|
Fix:
|
|
@@ -362,7 +362,7 @@ echo "$NEW_TOKEN" | secret-tool store --label="$LABEL" account "$ACCOUNT" servic
|
|
|
362
362
|
# 2. Sanity-check the doctor reports the new token's prefix
|
|
363
363
|
bash pipeline/lib/credential-store.sh doctor
|
|
364
364
|
|
|
365
|
-
# 3. Resume the pipeline - Phase
|
|
365
|
+
# 3. Resume the pipeline - Phase 4/5 re-resolve the token on each invocation,
|
|
366
366
|
# so no state edit is required.
|
|
367
367
|
/multi-agent:resume <task-id>
|
|
368
368
|
```
|
|
@@ -446,7 +446,7 @@ versions).
|
|
|
446
446
|
## Failed Push + Force-Push Decision (v8.0.0)
|
|
447
447
|
|
|
448
448
|
`git push` rejects the branch (non-fast-forward, hook rejected, branch
|
|
449
|
-
protection mismatch, etc.). The pipeline is paused at Phase
|
|
449
|
+
protection mismatch, etc.). The pipeline is paused at Phase 4 Step 3.
|
|
450
450
|
|
|
451
451
|
Decision tree:
|
|
452
452
|
|
|
@@ -0,0 +1,188 @@
|
|
|
1
|
+
# Running multi-agent on a machine nobody is sitting at
|
|
2
|
+
|
|
3
|
+
This describes what a Mac has to have before an unattended run works, and what
|
|
4
|
+
breaks if it does not. **Nothing here is set up by reading it** - `install`
|
|
5
|
+
still writes nothing about permissions and no scheduler is loaded unless you
|
|
6
|
+
ask for one. That is deliberate: every item below changes what runs on a
|
|
7
|
+
machine without a human, and none of it should happen because someone ran an
|
|
8
|
+
installer.
|
|
9
|
+
|
|
10
|
+
Scope is macOS. Linux and Windows were removed on purpose (ADR-0012) and
|
|
11
|
+
nothing here reintroduces them.
|
|
12
|
+
|
|
13
|
+
The mechanical half of this page is `doctor --profile=server`, which checks
|
|
14
|
+
four of the conditions below and says which are missing. Read that first; this
|
|
15
|
+
page explains why each one matters and what to do about it.
|
|
16
|
+
|
|
17
|
+
## The failure this page exists for
|
|
18
|
+
|
|
19
|
+
An unattended run does not fail loudly. It **stops without saying so**, and
|
|
20
|
+
every way it does that looks identical from outside: a process that is running,
|
|
21
|
+
printing nothing, finishing never.
|
|
22
|
+
|
|
23
|
+
Four causes, in the order they bite.
|
|
24
|
+
|
|
25
|
+
## 1. The permission posture
|
|
26
|
+
|
|
27
|
+
autopilot spawns its child with `--permission-prompts none`. That stops Claude
|
|
28
|
+
Code from ASKING, and it does not grant anything - the tools the child then
|
|
29
|
+
calls still have to be allowed. On the maintainer's machine they are, because
|
|
30
|
+
`Bash`, `Edit`, `Write` and `Agent` sit in the global allow list from years of
|
|
31
|
+
interactive use. On a fresh machine they do not, and the first tool call ends
|
|
32
|
+
the run.
|
|
33
|
+
|
|
34
|
+
There is no DEFAULT that fixes this and should not be: writing a permission
|
|
35
|
+
posture into someone's settings during an install is the one action an
|
|
36
|
+
installer must never take silently. There is now a flag that asks for it:
|
|
37
|
+
|
|
38
|
+
```bash
|
|
39
|
+
npx @mmerterden/multi-agent-pipeline install --unattended --dry-run # see it first
|
|
40
|
+
npx @mmerterden/multi-agent-pipeline install --unattended # write it
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
It prints the whole profile with a reason per line before writing anything, and
|
|
44
|
+
it is additive: an entry you added by hand survives, a narrower `Bash(...)` rule
|
|
45
|
+
is kept alongside, unrelated settings are untouched, and a second run changes
|
|
46
|
+
nothing. A `settings.json` that does not parse is refused rather than
|
|
47
|
+
overwritten.
|
|
48
|
+
|
|
49
|
+
By hand, the same thing is:
|
|
50
|
+
|
|
51
|
+
```jsonc
|
|
52
|
+
// ~/.claude/settings.json
|
|
53
|
+
{
|
|
54
|
+
"permissions": {
|
|
55
|
+
"allow": ["Bash", "Edit", "Write", "Agent"]
|
|
56
|
+
}
|
|
57
|
+
}
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
Grant the narrowest set the work needs. A server that only reviews does not
|
|
61
|
+
need `Write` - though note that the shipped profile is deliberately broad,
|
|
62
|
+
because a run builds and tests whatever the target repo uses and a narrow list
|
|
63
|
+
is the one the first unfamiliar repo stops at.
|
|
64
|
+
|
|
65
|
+
`doctor --profile=server` reports this as `unattended-permissions`, against the
|
|
66
|
+
same rule the writer applies.
|
|
67
|
+
|
|
68
|
+
## 2. The scheduler, and why it is a LaunchAgent
|
|
69
|
+
|
|
70
|
+
Something has to start the work. The template ships as a **LaunchAgent**
|
|
71
|
+
(`install/templates/multi-agent-autopilot.plist.template`), loaded in the
|
|
72
|
+
user's session at login - not a LaunchDaemon at boot.
|
|
73
|
+
|
|
74
|
+
That is a keychain decision, not a preference. Every credential the pipeline
|
|
75
|
+
reads lives in the login keychain, and the login keychain is **locked until the
|
|
76
|
+
user logs in**. A daemon started at boot gets a locked one: each credential
|
|
77
|
+
read fails, and those failures surface much later as 401s that blame the token.
|
|
78
|
+
The lock is the cause and nothing in the error says so.
|
|
79
|
+
|
|
80
|
+
So the machine needs to reach a logged-in session by itself:
|
|
81
|
+
|
|
82
|
+
- **System Settings → Users & Groups → Automatic login**, for the account that
|
|
83
|
+
owns the agent.
|
|
84
|
+
- A **keychain with no lock timeout** for that account, or the agent survives
|
|
85
|
+
a login and dies at the first idle period:
|
|
86
|
+
```
|
|
87
|
+
security set-keychain-settings -l ~/Library/Keychains/login.keychain-db
|
|
88
|
+
```
|
|
89
|
+
(no `-t`, so there is no timeout; `-l` still locks on sleep.)
|
|
90
|
+
|
|
91
|
+
Automatic login means physical access to the machine is access to the
|
|
92
|
+
credentials. On a server in a locked room that is the trade being made; on a
|
|
93
|
+
laptop it is not.
|
|
94
|
+
|
|
95
|
+
`doctor --profile=server` reports these as `scheduler` and `keychain-unlock`.
|
|
96
|
+
|
|
97
|
+
## 3. The unattended contract
|
|
98
|
+
|
|
99
|
+
`MULTI_AGENT_UNATTENDED=1` is how a run states that nobody is watching.
|
|
100
|
+
`multi-agent-refs/unattended-contract.md` lists exactly which entry points read
|
|
101
|
+
it and what each resolves to - four of them, named, because a claim of coverage
|
|
102
|
+
that is not true is worse than a short list.
|
|
103
|
+
|
|
104
|
+
The variable does **not** grant permissions and does not suppress errors. It
|
|
105
|
+
only stops a process from waiting for an answer that is not coming.
|
|
106
|
+
|
|
107
|
+
`doctor --profile=server` reports this as `unattended-contract`.
|
|
108
|
+
|
|
109
|
+
## 4. Which credential each phase needs
|
|
110
|
+
|
|
111
|
+
A server fails differently from a laptop here: an interactive run asks for a
|
|
112
|
+
missing token, an unattended one does not.
|
|
113
|
+
|
|
114
|
+
| Phase | Needs | Absent |
|
|
115
|
+
|---|---|---|
|
|
116
|
+
| 0 Init | `github` (issue intake), `jira` (ticket intake) | the run cannot resolve its input and stops at the picker |
|
|
117
|
+
| 1-2 Analysis, Plan | `figma` or `figma_mcp` when the task names a design | halts by contract rather than guessing at layout |
|
|
118
|
+
| 3 Dev | none beyond git access | - |
|
|
119
|
+
| 4 Review | `github` for PR comments | findings are computed and never posted |
|
|
120
|
+
| 5 Test | none | - |
|
|
121
|
+
| 6 Commit | git push credential | the branch stays local, which reads as "no PR yet" |
|
|
122
|
+
| 7 Report | `jira`, `confluence` as configured | the report is composed and dropped |
|
|
123
|
+
|
|
124
|
+
`credential-store.sh doctor` lists what is mapped. `doctor --probe` additionally
|
|
125
|
+
checks each one is alive, which costs network calls and is why it is opt-in.
|
|
126
|
+
|
|
127
|
+
## 4b. What the runner does when nobody is looking
|
|
128
|
+
|
|
129
|
+
Three behaviours that are invisible on a laptop and decide whether a server
|
|
130
|
+
survives a month.
|
|
131
|
+
|
|
132
|
+
**The log.** launchd appends the runner's stdout to `runner.log` forever. Each
|
|
133
|
+
tick truncates it past `MA_AP_LOG_MAX_BYTES` (5MB by default), keeping the last
|
|
134
|
+
256KB in `runner.log.1`. It truncates the SAME file rather than renaming it,
|
|
135
|
+
because launchd holds it open with `O_APPEND` - rename it and that descriptor
|
|
136
|
+
keeps writing into the renamed file while the new one stays empty, which is the
|
|
137
|
+
rotation that looks correct and silently stops logging.
|
|
138
|
+
|
|
139
|
+
**The breaker.** An expired token, a `claude` that no longer launches, a full
|
|
140
|
+
disk: every queued item fails the same way, minutes apart, until the queue is
|
|
141
|
+
empty and the ledger is a list of identical failures with no indication which
|
|
142
|
+
came first. After `MA_AP_BREAKER_LIMIT` consecutive attempts that produced
|
|
143
|
+
nothing (3 by default, 0 disables), the runner stops taking new work, writes the
|
|
144
|
+
reason into `queue.json` so `autopilot-status` shows it, and says what to check.
|
|
145
|
+
One successful attempt clears it. A run waiting for a person (`needs-input`) and
|
|
146
|
+
an item arming refused (`blocked-*`) are deliberately not failures.
|
|
147
|
+
|
|
148
|
+
**Telemetry.** `ticks.jsonl` gets one JSON line per tick. The human log answers
|
|
149
|
+
"what happened just now"; a server is only ever asked "how has this been
|
|
150
|
+
behaving for a week".
|
|
151
|
+
|
|
152
|
+
## 5. Self-hosted CI runner
|
|
153
|
+
|
|
154
|
+
ADR-0011 rejected a self-hosted runner because "a self-hosted runner on a
|
|
155
|
+
public repository executes code from any fork's pull request". Half of that has
|
|
156
|
+
changed and half has not, and the difference matters before wiring a runner to
|
|
157
|
+
a machine that holds credentials.
|
|
158
|
+
|
|
159
|
+
Measured, not assumed: the repo is **private**, and it has **one fork**, owned
|
|
160
|
+
by another account. "Zero forks" is a claim worth not making - GitHub requires
|
|
161
|
+
approval before a fork's pull request runs a workflow, but that protection is a
|
|
162
|
+
SETTING, and a self-hosted runner turns "someone approved a workflow run by
|
|
163
|
+
reflex" into code on this machine with this keychain.
|
|
164
|
+
|
|
165
|
+
So: a runner here is defensible on a private single-owner repo, and it is
|
|
166
|
+
defensible only while fork-PR approval stays on and every approval is treated
|
|
167
|
+
as a decision rather than a formality.
|
|
168
|
+
|
|
169
|
+
The same machine can host the runner:
|
|
170
|
+
|
|
171
|
+
```
|
|
172
|
+
mkdir ~/actions-runner && cd ~/actions-runner
|
|
173
|
+
# download and configure per GitHub's instructions for the repo
|
|
174
|
+
./svc.sh install && ./svc.sh start
|
|
175
|
+
```
|
|
176
|
+
|
|
177
|
+
Two things to keep in mind. The runner service runs in the same session as the
|
|
178
|
+
autopilot agent, so a long CI job and a pipeline run compete for the same CPU
|
|
179
|
+
and the same `~/.claude` state root. And the runner inherits that session's
|
|
180
|
+
keychain, which means a workflow can read the credentials - acceptable on a
|
|
181
|
+
single-owner machine, not on a shared one.
|
|
182
|
+
|
|
183
|
+
## What this page does NOT cover
|
|
184
|
+
|
|
185
|
+
- Turning any of it on. Every step here is the operator's to take.
|
|
186
|
+
- Linux, Windows, systemd, Docker. Out of scope and staying there.
|
|
187
|
+
- Multi-user servers. Everything above assumes one account owns the install,
|
|
188
|
+
the keychain and the queue.
|