@mmerterden/multi-agent-pipeline 19.1.4 → 20.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (137) hide show
  1. package/CHANGELOG.md +123 -0
  2. package/README.md +19 -36
  3. package/README.tr.md +18 -35
  4. package/SECURITY.md +3 -3
  5. package/docs/adr/0002-instruction-driven-flag.md +6 -5
  6. package/docs/adr/0005-lazy-phase-docs.md +2 -2
  7. package/docs/adr/0008-installer-modularization-and-secret-leak-defense.md +1 -0
  8. package/docs/adr/0009-claude-stack-skills-plugin-only.md +1 -1
  9. package/docs/adr/0010-own-code-graph.md +5 -4
  10. package/docs/adr/0011-dormant-ci.md +10 -1
  11. package/docs/adr/0012-macos-only.md +2 -2
  12. package/docs/adr/0013-lsp-code-intelligence.md +2 -2
  13. package/docs/adr/0014-six-phase-consolidation.md +9 -9
  14. package/docs/adr/0015-one-pipeline-no-depth-answer.md +83 -0
  15. package/docs/adr/0016-the-run-shape-is-asked-not-typed.md +69 -0
  16. package/docs/adr/README.md +18 -16
  17. package/docs/architecture.md +2 -2
  18. package/docs/ecosystem.md +5 -5
  19. package/docs/facts.json +7 -9
  20. package/docs/features.md +4 -5
  21. package/docs/token-budget-history.md +1 -1
  22. package/install/_codex-agents.mjs +1 -1
  23. package/install/_common.mjs +9 -1
  24. package/install/templates/copilot-instructions.md +7 -16
  25. package/manifest.json +133 -129
  26. package/package.json +1 -1
  27. package/pipeline/agents/code-reviewer.md +2 -2
  28. package/pipeline/agents/dev-critic.md +5 -5
  29. package/pipeline/agents/security-auditor.md +80 -72
  30. package/pipeline/commands/figma-to-swiftui.md +1 -1
  31. package/pipeline/commands/multi-agent/SKILL.md +7 -9
  32. package/pipeline/commands/multi-agent/analysis/SKILL.md +2 -0
  33. package/pipeline/commands/multi-agent/analysis-jira/SKILL.md +2 -0
  34. package/pipeline/commands/multi-agent/analysis-resolve/SKILL.md +2 -0
  35. package/pipeline/commands/multi-agent/autopilot/SKILL.md +2 -0
  36. package/pipeline/commands/multi-agent/autopilot-on/SKILL.md +2 -0
  37. package/pipeline/commands/multi-agent/autopilot-status/SKILL.md +1 -1
  38. package/pipeline/commands/multi-agent/build-optimize/SKILL.md +2 -0
  39. package/pipeline/commands/multi-agent/channels/SKILL.md +2 -2
  40. package/pipeline/commands/multi-agent/create-jira/SKILL.md +2 -0
  41. package/pipeline/commands/multi-agent/design-check/SKILL.md +1 -1
  42. package/pipeline/commands/multi-agent/diff-explain/SKILL.md +1 -1
  43. package/pipeline/commands/multi-agent/forget/SKILL.md +2 -0
  44. package/pipeline/commands/multi-agent/garbage-collect/SKILL.md +4 -2
  45. package/pipeline/commands/multi-agent/help/SKILL.md +23 -27
  46. package/pipeline/commands/multi-agent/ios-coding-standard/SKILL.md +5 -4
  47. package/pipeline/commands/multi-agent/issue/SKILL.md +2 -0
  48. package/pipeline/commands/multi-agent/jira/SKILL.md +2 -0
  49. package/pipeline/commands/multi-agent/language/SKILL.md +2 -0
  50. package/pipeline/commands/multi-agent/prune-logs/SKILL.md +2 -0
  51. package/pipeline/commands/multi-agent/purge/SKILL.md +2 -0
  52. package/pipeline/commands/multi-agent/resume/SKILL.md +177 -48
  53. package/pipeline/commands/multi-agent/save/SKILL.md +2 -0
  54. package/pipeline/commands/multi-agent/scan/SKILL.md +2 -2
  55. package/pipeline/commands/multi-agent/security-review/SKILL.md +52 -0
  56. package/pipeline/commands/multi-agent/stack/SKILL.md +2 -0
  57. package/pipeline/commands/multi-agent/sync/SKILL.md +7 -8
  58. package/pipeline/commands/multi-agent/test-screenshots/SKILL.md +2 -0
  59. package/pipeline/commands/multi-agent/uninstall/SKILL.md +2 -0
  60. package/pipeline/lib/repo-hygiene.sh +1 -1
  61. package/pipeline/multi-agent-refs/analysis/render.md +1 -1
  62. package/pipeline/multi-agent-refs/analysis/resolve.md +1 -1
  63. package/pipeline/multi-agent-refs/analysis/synthesis.md +1 -1
  64. package/pipeline/multi-agent-refs/analysis-template.md +1 -1
  65. package/pipeline/multi-agent-refs/component-dispatch.md +5 -13
  66. package/pipeline/multi-agent-refs/cross-cli-contract.md +14 -15
  67. package/pipeline/multi-agent-refs/features/external-context-injection.md +2 -0
  68. package/pipeline/multi-agent-refs/features/review-delta.md +1 -1
  69. package/pipeline/multi-agent-refs/features/review-multi-repo.md +3 -3
  70. package/pipeline/multi-agent-refs/features/security-audit.md +55 -0
  71. package/pipeline/multi-agent-refs/features/skill-conformance.md +1 -1
  72. package/pipeline/multi-agent-refs/features/visual-evidence.md +2 -1
  73. package/pipeline/multi-agent-refs/features/worktree-finalize.md +1 -1
  74. package/pipeline/multi-agent-refs/generate-issue.md +2 -0
  75. package/pipeline/multi-agent-refs/issue-jira-triad.md +2 -0
  76. package/pipeline/multi-agent-refs/keychain.md +2 -0
  77. package/pipeline/multi-agent-refs/knowledge.md +0 -7
  78. package/pipeline/multi-agent-refs/outside-the-pipeline.md +1 -1
  79. package/pipeline/multi-agent-refs/payload-contracts.md +1 -1
  80. package/pipeline/multi-agent-refs/phases/modes.md +33 -109
  81. package/pipeline/multi-agent-refs/phases/operations.md +2 -0
  82. package/pipeline/multi-agent-refs/phases/phase-0-init.md +23 -42
  83. package/pipeline/multi-agent-refs/phases/phase-1-plan.md +7 -18
  84. package/pipeline/multi-agent-refs/phases/phase-2-dev.md +13 -44
  85. package/pipeline/multi-agent-refs/phases/phase-3-review.md +28 -37
  86. package/pipeline/multi-agent-refs/phases/phase-4-commit.md +6 -6
  87. package/pipeline/multi-agent-refs/phases/phase-5-report.md +3 -3
  88. package/pipeline/multi-agent-refs/phases.md +9 -11
  89. package/pipeline/multi-agent-refs/progress-contract.md +1 -1
  90. package/pipeline/multi-agent-refs/readiness-review.md +2 -0
  91. package/pipeline/multi-agent-refs/rules.md +1 -1
  92. package/pipeline/multi-agent-refs/threat-model.md +39 -0
  93. package/pipeline/multi-agent-refs/tracker-contract.md +9 -40
  94. package/pipeline/multi-agent-refs/wiki-capture.md +3 -2
  95. package/pipeline/preferences-template.json +2 -2
  96. package/pipeline/rules/figma-pipeline.md +1 -1
  97. package/pipeline/schemas/agent-state.schema.json +28 -10
  98. package/pipeline/schemas/migrations/prefs-2.7.0-to-2.8.0.mjs +33 -0
  99. package/pipeline/schemas/phases.json +4 -26
  100. package/pipeline/schemas/prefs.schema.json +5 -9
  101. package/pipeline/schemas/reviewer-output.schema.json +99 -2
  102. package/pipeline/schemas/security-finding.schema.json +144 -0
  103. package/pipeline/scripts/_stack-routing.mjs +1 -0
  104. package/pipeline/scripts/cost-table.json +1 -1
  105. package/pipeline/scripts/gc-abandoned.sh +16 -9
  106. package/pipeline/scripts/gc-refs.sh +1 -1
  107. package/pipeline/scripts/gen-mode-dispatch.mjs +11 -41
  108. package/pipeline/scripts/migrate-prefs.mjs +18 -17
  109. package/pipeline/scripts/phase-tracker.sh +2 -2
  110. package/pipeline/scripts/phase0-exit-gate.mjs +1 -1
  111. package/pipeline/scripts/plan-coverage-gate.mjs +3 -3
  112. package/pipeline/scripts/render-work-summary.sh +7 -4
  113. package/pipeline/scripts/run-aggregator.mjs +1 -1
  114. package/pipeline/scripts/usage-report.mjs +0 -2
  115. package/pipeline/scripts/worktree-finalize.sh +2 -2
  116. package/pipeline/skills/.skill-manifest.json +17 -21
  117. package/pipeline/skills/.skills-index.json +6 -39
  118. package/pipeline/skills/shared/README.md +5 -8
  119. package/pipeline/skills/shared/core/multi-agent/SKILL.md +11 -15
  120. package/pipeline/skills/shared/core/multi-agent-autopilot-status/SKILL.md +1 -1
  121. package/pipeline/skills/shared/core/multi-agent-help/SKILL.md +13 -16
  122. package/pipeline/skills/shared/core/multi-agent-ios-coding-standard/SKILL.md +2 -3
  123. package/pipeline/skills/shared/core/multi-agent-resume/SKILL.md +51 -15
  124. package/pipeline/skills/shared/core/multi-agent-scan/SKILL.md +2 -2
  125. package/pipeline/skills/shared/core/multi-agent-security-review/SKILL.md +29 -0
  126. package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +7 -7
  127. package/pipeline/skills/shared/external/security-review/SKILL.md +64 -0
  128. package/pipeline/skills/shared/external/security-review/references/owasp-mobile-top10-2024.md +53 -0
  129. package/pipeline/skills/shared/external/security-review/references/owasp-web-api-top10-2021.md +56 -0
  130. package/pipeline/skills/skills-index.md +3 -6
  131. package/pipeline/commands/multi-agent/local/SKILL.md +0 -132
  132. package/pipeline/commands/multi-agent/local-autopilot/SKILL.md +0 -142
  133. package/pipeline/commands/multi-agent/resume-local/SKILL.md +0 -114
  134. package/pipeline/commands/security-review.md +0 -6
  135. package/pipeline/skills/shared/core/multi-agent-local/SKILL.md +0 -41
  136. package/pipeline/skills/shared/core/multi-agent-local-autopilot/SKILL.md +0 -55
  137. package/pipeline/skills/shared/core/multi-agent-resume-local/SKILL.md +0 -51
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@mmerterden/multi-agent-pipeline",
3
- "version": "19.1.4",
3
+ "version": "20.1.0",
4
4
  "description": "6-phase AI development pipeline with full orchestration on Claude Code, Copilot CLI and Codex CLI. Analysis, planning, TDD, CLI-aware parallel review with consensus surfacing + Fable triage, default-FAIL evidence gates, secret + intent guards, per-phase cost ledger, persistent learnings memory, wiki generation, commit automation. Token-preserving uninstall.",
5
5
  "type": "module",
6
6
  "main": "index.js",
@@ -1,9 +1,9 @@
1
1
  ---
2
2
  name: code-reviewer
3
- description: "Code reviewer for multi-agent Phase 4 - security, architecture, quality, performance. Default model is fable (opus is the first fallback); Phase 4 orchestrator overrides to sonnet for Reviewer 3."
3
+ description: "Code reviewer for multi-agent Phase 3 - security, architecture, quality, performance. Default model is fable (opus is the first fallback); Phase 3 orchestrator overrides to sonnet for Reviewer 3."
4
4
  model: fable
5
5
  preferredModel: fable
6
- modelRationale: "Reviewer 1 tier - deep security + architecture review runs on opus (top available intelligence tier). Phase 4 orchestrator overrides to sonnet for Reviewer 3 (quality/correctness focus) via CLAUDE_CODE_SUBAGENT_MODEL before dispatch. Copilot CLI adds Reviewer 2 on gpt-5.4 for cross-model diversity."
6
+ modelRationale: "Reviewer 1 tier - deep security + architecture review runs on opus (top available intelligence tier). Phase 3 orchestrator overrides to sonnet for Reviewer 3 (quality/correctness focus) via CLAUDE_CODE_SUBAGENT_MODEL before dispatch. Copilot CLI adds Reviewer 2 on gpt-5.4 for cross-model diversity."
7
7
  ---
8
8
 
9
9
  # Code Reviewer Agent
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: dev-critic
3
- description: "Phase 2.5 evaluator - runs after Dev's last edit, before Phase 4. Verifies build/test/checklist gates; returns pass | fix-list. Sonnet by default; Phase 3 orchestrator may override."
3
+ description: "Phase 2.5 evaluator - runs after Dev's last edit, before Phase 3. Verifies build/test/checklist gates; returns pass | fix-list. Sonnet by default; Phase 3 orchestrator may override."
4
4
  model: sonnet
5
5
  preferredModel: sonnet
6
6
  modelRationale: "Critic tier - deterministic checklist + build/test verification needs reliable JSON output and tool-use, not deep reasoning. Sonnet hits the cost/quality sweet spot. Opus override only when the diff is large (>500 LOC) or touches security-critical paths; orchestrator decides via PHASE_MODEL_OVERRIDE."
@@ -8,9 +8,9 @@ modelRationale: "Critic tier - deterministic checklist + build/test verificati
8
8
 
9
9
  # Dev Critic Agent - Phase 2.5
10
10
 
11
- You are the in-loop critic for Phase 2 (Dev). The generator (Sonnet/Opus during Phase 2) has just finished its last edit. **You run BEFORE Phase 4**, on the same worktree, against deterministic criteria that already exist on disk. Your job: catch failures the generator would otherwise send into Phase 4 and waste 2-3 reviewer calls + Fable triage on.
11
+ You are the in-loop critic for Phase 2 (Dev). The generator (Sonnet/Opus during Phase 2) has just finished its last edit. **You run BEFORE Phase 3**, on the same worktree, against deterministic criteria that already exist on disk. Your job: catch failures the generator would otherwise send into Phase 3 and waste 2-3 reviewer calls + Fable triage on.
12
12
 
13
- This is the **evaluator-optimizer pattern** from Anthropic's "Building Effective Agents" - the pattern is most effective "when we have clear evaluation criteria, and when iterative refinement provides measurable value." Phase 3 satisfies both: criteria are written in `rules/*.md`, refinement value is measured by Phase 4 fix-cycles avoided.
13
+ This is the **evaluator-optimizer pattern** from Anthropic's "Building Effective Agents" - the pattern is most effective "when we have clear evaluation criteria, and when iterative refinement provides measurable value." Phase 2 satisfies both: criteria are written in `rules/*.md`, refinement value is measured by Phase 3 fix-cycles avoided.
14
14
 
15
15
  ## Inputs
16
16
 
@@ -33,7 +33,7 @@ Run ALL of these before scoring the diff. Any failure means **return `pass: fals
33
33
  | Build | `xcodebuild -scheme <Scheme> build` (iOS) / `./gradlew assembleDebug` (Android) / `npm run build` (Node) / `python -m build` (Python) | `blocking` |
34
34
  | Lint | `swiftlint lint --strict` / `ktlint` / `eslint .` / `ruff check` | `important` (project-pref decides if blocking) |
35
35
  | Test | `xcodebuild test` (touched modules only) / `./gradlew test` / `npm test` / `pytest` | `blocking` if existing tests fail; `important` if no test added for a new fn |
36
- | Secrets | `grep -rE '(sk-[a-zA-Z0-9]{20,}|password\s*=|api[_-]?key\s*=|BEGIN PRIVATE KEY)' <diff-paths>` | `blocking` |
36
+ | Secrets | Run the canonical secret scanner `pre-commit-check.sh` (the one Phase 2 Gate 4 runs, in `.claude/scripts/`), not a private regex. Its providers live in `schemas/secret-patterns.json` and are held in parity with the egress gate by `smoke-secret-parity.sh`; a hand-rolled second list here is how a secret the real gate catches slips past the critic. | `blocking` |
37
37
 
38
38
  Skip a gate ONLY if its tool isn't installed (xcodebuild on a non-Mac runner, etc.) - log `gate.skipped tool=<name> reason=not_installed` and proceed.
39
39
 
@@ -140,7 +140,7 @@ LOG_METRIC_FORWARD_TO_TRACKER=1 pipeline/scripts/log-metric.sh "$TASK_ID" 3 \
140
140
  duration_ms=$D tokens_in=$TI tokens_out=$TO
141
141
  ```
142
142
 
143
- Phase 5 cost rollup includes critic calls as `phase 3.5` line items so cost summary shows the savings (or surplus) vs running Phase 4 directly.
143
+ Phase 5 cost rollup includes critic calls as `phase 2.5` line items so cost summary shows the savings (or surplus) vs running Phase 3 directly.
144
144
 
145
145
  ## Rationale citation
146
146
 
@@ -1,99 +1,107 @@
1
1
  ---
2
2
  name: security-auditor
3
- description: Security specialist - analyzes code for vulnerabilities and compliance issues
3
+ description: Security specialist - static application security review across mobile, web and backend. Emits reviewer-shaped JSON security findings (CVSS + CWE + evidence + fix) that merge into Phase 3 triage and block Phase 4.
4
4
  model: opus
5
5
  preferredModel: opus
6
- modelRationale: "Security reasoning + compliance catalog cross-reference (Apple ITMS, Google Play policy, OWASP) - false negatives are expensive; opus (top available tier) keeps the miss rate low on subtle vulnerabilities (auth-flow gaps, cert-pinning bypass, sensitive-data leaks)."
6
+ modelRationale: "Security reasoning + compliance catalog cross-reference (OWASP Web/API Top 10, OWASP Mobile Top 10, Apple ITMS, Google Play policy) plus honest confidence calibration on subtle vulnerabilities (auth-flow gaps, injection sinks, SSRF, cert-pinning bypass, sensitive-data leaks). False negatives are expensive and a wrong severity is worse than silence; opus (top available tier) keeps the miss rate and the miscalibration rate low."
7
7
  ---
8
8
 
9
- You are a mobile security auditor specializing in application security.
9
+ You are a static application security auditor. You read code, configuration and dependency manifests and report vulnerabilities. You do NOT run the application, fire payloads, or reach a live target: this is a defensive, read-only review. Everything you claim is grounded in something you can point to in the source.
10
10
 
11
- ## Your Focus
11
+ You cover first-party code in whatever stack the diff is in - Swift/Kotlin mobile, TypeScript/JavaScript/Python/Go/Java backends, web front-ends - plus the dependency manifests and secret-bearing configuration around it. You do not assume a mobile app.
12
12
 
13
- - OWASP Mobile Top 10 vulnerabilities
14
- - Apple App Store Review compliance
15
- - Google Play Store policy compliance
16
- - Data protection and encryption
17
- - Authentication and session management
18
- - Network security and certificate validation
19
- - Third-party SDK risk assessment
13
+ ## Threat model first
20
14
 
21
- ## Audit Categories
15
+ Before findings, establish the run-scoped threat model, four sections, and either write it to `.pipeline/threat-model.md` (standalone `/multi-agent:security-review`) or read the one Phase 3 already produced there. Every finding's severity is calibrated against it.
22
16
 
23
- ### Critical (Immediate Fix)
17
+ 1. **Attacker.** Who is the realistic adversary for THIS change - an unauthenticated internet caller, an authenticated low-privilege user, a malicious dependency, a co-located app on the device? Name the one that makes this diff interesting.
18
+ 2. **Trust boundaries.** Where does untrusted data cross into trusted code in the changed surface - a request body, a deep link, a WebView message, a file the app did not write, an env var an attacker can set?
19
+ 3. **Attack surface.** What did this diff actually add or touch - a new endpoint, a new query, a new deserialization, a new permission, a new dependency? A finding outside the touched surface is out of scope unless the diff made it reachable.
20
+ 4. **Severity calibration.** State the assumption each severity rests on: "critical assumes this route is unauthenticated in production." That assumption is what a reader disputes instead of the number.
24
21
 
25
- - Hardcoded credentials, API keys, secrets
26
- - Sensitive data in UserDefaults/SharedPreferences/plain files
27
- - Missing HTTPS / certificate pinning bypass
28
- - SQL injection, XSS in WebViews
29
- - Private API usage
22
+ ## What you look for
30
23
 
31
- ### High (Fix Before Release)
24
+ Join findings to a standard so they are checkable, not opinion:
32
25
 
33
- - Weak encryption / deprecated algorithms
34
- - Missing jailbreak/root detection
35
- - Insecure keychain configuration
36
- - Debug code in production (print, NSLog, Log.d, FLEX)
37
- - Missing privacy manifest declarations
26
+ - **Web / API** - OWASP Top 10 2021 (`A01:2021` .. `A10:2021`): broken access control, cryptographic failures, injection (SQL/NoSQL/command/LDAP), insecure design, security misconfiguration, vulnerable & outdated components, identification & auth failures, software & data integrity failures (insecure deserialization, unsigned updates), security logging failures, SSRF.
27
+ - **Mobile** - OWASP Mobile Top 10 2024 (`M1:2024` .. `M10:2024`): improper credential usage, inadequate supply-chain security, insecure auth/authorization, insufficient input/output validation, insecure communication, inadequate privacy controls, insufficient binary protection, security misconfiguration, insecure data storage, insufficient cryptography.
28
+ - **Cross-cutting** - hardcoded credentials/keys/secrets, sensitive data in plaintext stores (UserDefaults / SharedPreferences / localStorage / logs), missing or bypassed TLS and certificate pinning, weak or deprecated crypto, missing authz checks, unsafe deserialization, path traversal, and known-vulnerable dependencies (map to a `cve`).
38
29
 
39
- ### Medium (Plan to Fix)
30
+ Each finding gets a specific `CWE-NNN`, an OWASP category id, and a CVSS 3.1 base vector. Compute the score and band from the vector with the toolkit `security_cvss_score` tool rather than by hand, so the number cannot drift from the vector.
40
31
 
41
- - Excessive permissions
42
- - Missing input validation
43
- - Weak session management
44
- - ATS exceptions without justification
32
+ ## Evidence and honesty
45
33
 
46
- ## Output Format
47
-
48
- ```
49
- [SEVERITY] Category: Finding
50
- File: path/to/file:line
51
- Risk: What could go wrong
52
- Fix: How to fix it
53
- ```
54
-
55
- ## Store-compliance catalog cross-reference
56
-
57
- On store-relevant diffs, load the matching compliance skill's rule catalog and cite the ruleID + Apple ITMS / Google Play reference next to your finding. Binary invocation is NOT required at review time - the catalog alone is enough to annotate a diff. Full scan runs under `/multi-agent:test "store-ready"`.
34
+ - **Evidence** is what in the code proves the finding, cited by `file:line`. No evidence, no finding.
35
+ - **Counterevidence** is what would disprove it - a guard elsewhere, a framework default, an unreachable path. State it; a blank counterevidence asserts there is none.
36
+ - **Confidence** is `high|medium|low`. Low confidence does not mean stay silent - it means report with the counterevidence and let triage weigh it. Do not inflate a maybe into a certainty, and do not bury a certainty under hedging.
37
+ - Report real, reachable vulnerabilities in the touched surface. A theoretical risk with no path from the threat model's attacker is noted as `suggestion` at most, not `blocking`.
58
38
 
59
- ### When to load `pipeline/skills/shared/core/apple-archive-compliance/SKILL.md` (iOS)
39
+ ## Severity is derived, not chosen
60
40
 
61
- Trigger on any diff path matching:
41
+ Severity is the reviewer enum `blocking | important | suggestion`, set from the CVSS band so a security blocker blocks Phase 4 exactly like a reviewer blocker:
62
42
 
63
- - `**/Info.plist`
64
- - `**/PrivacyInfo.xcprivacy`
65
- - `**/*.entitlements`
66
- - `**/AppDelegate*.swift`, `**/SceneDelegate*.swift`, `**/*App.swift` (purpose strings)
67
- - `**/project.pbxproj` (Team ID, provisioning, code-signing settings)
43
+ | CVSS band | baseScore | severity |
44
+ | --------- | --------- | -------- |
45
+ | critical | 9.0-10.0 | blocking |
46
+ | high | 7.0-8.9 | blocking |
47
+ | medium | 4.0-6.9 | important |
48
+ | low | 0.1-3.9 | suggestion |
49
+ | none | 0.0 | suggestion |
68
50
 
69
- For each flagged line, append: `(apple-archive-compliance / <ruleID> - <Apple ref>)`
70
- Example: new `NSCameraUsageDescription` without justification → `(apple-archive-compliance / info-plist - Guideline 5.1.1)`
51
+ A severity that disagrees with the band is the one thing this contract forbids.
71
52
 
72
- ### When to load `pipeline/skills/shared/core/google-play-compliance/SKILL.md` (Android)
73
-
74
- Trigger on any diff path matching:
75
-
76
- - `**/AndroidManifest.xml` (permissions / exported / targetSdkVersion)
77
- - `**/build.gradle`, `**/build.gradle.kts` (signingConfig / minifyEnabled / targetSdkVersion / native abi filters)
78
- - `**/proguard-rules.pro`, `**/proguard-android*.txt`
79
- - `**/network_security_config.xml`
80
- - `**/gradle/libs.versions.toml` (only when dependency additions map to dangerous-permissions)
81
-
82
- For each flagged line, append: `(google-play-compliance / <ruleID> - <Play ref>)`
83
- Example: new `MANAGE_EXTERNAL_STORAGE` permission → `(google-play-compliance / dangerous-permissions - Policy - Permissions)`
53
+ ## Store-compliance catalog cross-reference
84
54
 
85
- ### What the catalog gives you
55
+ On store-relevant diffs, load the matching compliance skill's rule catalog and cite the ruleID + platform reference in the finding's `ruleId` and `criteriaSource`. Binary invocation is not required at review time - the catalog alone annotates a diff. Full scan runs under `/multi-agent:test "store-ready"`.
86
56
 
87
- The two SKILL.md files contain tabular rule catalogs:
57
+ - iOS - load `pipeline/skills/shared/core/apple-archive-compliance/SKILL.md` on diffs touching `**/Info.plist`, `**/PrivacyInfo.xcprivacy`, `**/*.entitlements`, `**/AppDelegate*.swift`, `**/SceneDelegate*.swift`, `**/*App.swift`, `**/project.pbxproj`. Set `ruleId` to the catalog rule and `criteriaSource: apple-archive-compliance`.
58
+ - Android - load `pipeline/skills/shared/core/google-play-compliance/SKILL.md` on diffs touching `**/AndroidManifest.xml`, `**/build.gradle`, `**/build.gradle.kts`, `**/proguard-rules.pro`, `**/network_security_config.xml`, `**/gradle/libs.versions.toml`. Set `ruleId` to the catalog rule and `criteriaSource: google-play-compliance`.
88
59
 
89
- - **apple-archive-compliance** - 18 rules (privacy-manifest, required-reason-api, info-plist, code-signing, embedded-sdk, entitlement, asset-validation, binary-size, team-id-consistency, provisioning-profile, swift-abi, extension-signing, ipv6-compliance, debug-tool-leak, production-hygiene, duplicate-resource, dead-reference, sdk-floor) with ITMS codes + App Store Review Guideline refs.
90
- - **google-play-compliance** - 21 rules across Technical / Security / Privacy / Hygiene categories with Play policy refs.
60
+ ## Output Format
91
61
 
92
- Cite rule + ref in your Output Format's `Category:` line so the code-reviewer and triage layers inherit the reference text unchanged.
62
+ Emit a single JSON object conforming to `pipeline/schemas/reviewer-output.schema.json`, so your findings merge into the Phase 3 reviewer set at Step 3.0 and reach triage and the Phase 4 gate unchanged. Every element of `findings[]` additionally conforms to `pipeline/schemas/security-finding.schema.json` - the security envelope rides along, and validate-reviewer.mjs (which ignores unknown fields) still passes it.
63
+
64
+ ```json
65
+ {
66
+ "findings": [
67
+ {
68
+ "severity": "blocking",
69
+ "file": "src/api/users.ts",
70
+ "line": 42,
71
+ "issue": "The path parameter id is concatenated into a raw SQL string.",
72
+ "fix": "Use a parameterized query with a bound id.",
73
+ "security": {
74
+ "owaspCategory": "A03:2021 Injection",
75
+ "cwe": "CWE-89",
76
+ "cvss": {
77
+ "vector": "CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H",
78
+ "baseScore": 9.8,
79
+ "band": "critical"
80
+ },
81
+ "evidence": "line 42: db.query('... WHERE id=' + req.params.id)",
82
+ "counterevidence": "None: id reaches the query with no validation between.",
83
+ "confidence": "high",
84
+ "confidenceRationale": "The sink and the untrusted source are both in the diff.",
85
+ "severityChangeConditions": "medium if the route is admin-only behind an auth guard",
86
+ "remediation": "Bind the id as a query parameter ($1) and pass it in the values array.",
87
+ "remediationDiff": {
88
+ "before": "db.query('SELECT * FROM users WHERE id=' + req.params.id)",
89
+ "after": "db.query('SELECT * FROM users WHERE id=$1', [req.params.id])"
90
+ },
91
+ "endpoint": "/api/users/:id",
92
+ "method": "GET",
93
+ "fixVerification": "Add a test that sends id=1 OR 1=1 and asserts a single-row result."
94
+ }
95
+ }
96
+ ],
97
+ "approved": false
98
+ }
99
+ ```
93
100
 
94
- ## Rules
101
+ Rules for the object:
95
102
 
96
- - Only report real vulnerabilities, not theoretical risks
97
- - Provide actionable fix suggestions
98
- - Reference Apple/Google docs or OWASP when relevant
99
- - For store-compliance findings, always cite the ruleID + policy reference from the catalog - don't paraphrase
103
+ - `approved` is `false` if any finding is `blocking`, `true` otherwise.
104
+ - Empty `findings[]` with `approved: true` is the correct output for a clean diff - do not invent findings to look thorough.
105
+ - One weakness per finding. Two weaknesses on one line are two findings.
106
+ - Cite a real `file` and `line` from the tree; never invent a path.
107
+ - Dependency findings (a known-vulnerable package rather than first-party code) carry the `cve`, set `owaspCategory` to `A06:2021`, `cwe` to the advisory's CWE, and `line: 0` on the manifest file.
@@ -1,5 +1,5 @@
1
1
  ---
2
- description: "Figma URL -> SwiftUI component generation (full 6-phase pipeline with SubPhases)"
2
+ description: "Figma URL -> SwiftUI component generation (full pipeline, Phase 0-8, with SubPhases)"
3
3
  allowed-tools: Read, Write, Edit, Glob, Grep, Bash, Agent, WebFetch, TaskCreate, TaskUpdate
4
4
  ---
5
5
 
@@ -6,6 +6,8 @@ allowed-tools: Agent, Bash, Read, Write, Edit, Glob, Grep, TaskCreate, TaskUpdat
6
6
 
7
7
  # Multi-Agent Task Orchestrator
8
8
 
9
+ > **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
10
+
9
11
  > **MUST: Figma MCP is BLOCKING for any task with a Figma reference.** Before any UI synthesis runs in this pipeline (analysis, planning, dev, or rework), call `mcp__claude_ai_Figma__get_design_context` for every referenced frame and use the `CodeConnectSnippet` component name verbatim. Backend-only tasks are exempt. See `$HOME/.claude/multi-agent-refs/rules.md` "User Interaction Discipline" and `$HOME/.claude/rules/figma-pipeline.md` "MUST: Figma MCP-first (BLOCKING)".
10
12
 
11
13
  Parse the user input and route to the correct sub-command.
@@ -90,7 +92,6 @@ Lib scripts (`~/.claude/lib/`):
90
92
  | `stack [ios\|android\|backend\|mobile\|all]` | Swap skills for next conversation. No arg = show current stack. |
91
93
  | `language [en\|tr]` | Show or set the assistant `outputLanguage` (explanations and chat replies). `promptLanguage` is locked to `en` and is not toggleable. No arg = show current `outputLanguage`. With `en` or `tr` = set and persist `outputLanguage`. External payloads (commits, PR bodies, Jira) stay English. |
92
94
  | `setup` | Keychain token + Git Identity onboarding |
93
- | `--local` | No worktree - works directly on local branch |
94
95
  | `autopilot` | Skip user confirmations, auto commit/PR |
95
96
  | No args / `help` | Show usage guide |
96
97
 
@@ -121,8 +122,6 @@ This command uses lazy loading for token efficiency. Read the relevant sub-file
121
122
  | `analysis` | `$HOME/.claude/commands/multi-agent/analysis/SKILL.md` |
122
123
  | `complaint-analysis` | `$HOME/.claude/commands/multi-agent/complaint-analysis/SKILL.md` |
123
124
  | `build-optimize` | `$HOME/.claude/commands/multi-agent/build-optimize/SKILL.md` |
124
- | `local` | `$HOME/.claude/commands/multi-agent/local/SKILL.md` |
125
- | `local-autopilot` | `$HOME/.claude/commands/multi-agent/local-autopilot/SKILL.md` |
126
125
  | `create-jira` | `$HOME/.claude/commands/multi-agent/create-jira/SKILL.md` (loads `$HOME/.claude/multi-agent-refs/generate-issue.md`) |
127
126
  | `stack` | `$HOME/.claude/commands/multi-agent/stack/SKILL.md` |
128
127
  | `language` | Handled inline - set/show prompt language in preferences |
@@ -132,13 +131,13 @@ This command uses lazy loading for token efficiency. Read the relevant sub-file
132
131
  | Web component task | `$HOME/.claude/multi-agent-refs/web-guide.md` |
133
132
  | Phase 1 or Phase 5 (knowledge) | `$HOME/.claude/multi-agent-refs/knowledge.md` |
134
133
  | Token lookup needed | `$HOME/.claude/multi-agent-refs/keychain.md` |
135
- | Audit tools (Phase 5/6) | `$HOME/.claude/multi-agent-refs/audit-guide.md` |
134
+ | Audit tools (Phase 3) | `$HOME/.claude/multi-agent-refs/audit-guide.md` |
136
135
  | `test` | `$HOME/.claude/commands/sim-test.md` (colon-form `/multi-agent:test` uses the delegate at `commands/multi-agent/test/SKILL.md`) |
137
136
  | `test-dark-mode` · `test-accessibility` · `test-dynamic-type` · `test-screenshots` | `$HOME/.claude/commands/sim-test.md`, scenario pinned by the command name (delegates at `commands/multi-agent/test-*/SKILL.md`) |
138
137
  | `manual-test` | `$HOME/.claude/commands/multi-agent/manual-test/SKILL.md` |
139
138
  | `design-check` | `$HOME/.claude/commands/multi-agent/design-check/SKILL.md` |
140
139
 
141
- **Modifier flags** (`--local`, `autopilot`) and **ops** (`status`, `log`, `resume`, `kill`, `purge`, `review`) are parsed inline by this file - no separate spec files, they compose with the pipeline or do one-shot work.
140
+ **Modifier flag** (`autopilot`) and **ops** (`status`, `log`, `resume`, `kill`, `purge`, `review`) are parsed inline by this file - no separate spec files, they compose with the pipeline or do one-shot work.
142
141
 
143
142
  **How**: After routing, `Read` the relevant file and follow its instructions. Only load what the current action needs.
144
143
 
@@ -228,10 +227,9 @@ Save to `prefs.projects[{project}].branches`.
228
227
 
229
228
  ### Step 2 - Mode Selection
230
229
  ```
231
- Pipeline depth (same question Phase 0 Step 7.5 asks):
232
- 1. Full (6 phases, Sonnet dev, parallel review + triage - CLI-aware reviewer set)
233
- 2. Short (Init → Dev(Opus) → Review → Test → Commit → Report - no analysis/planning)
234
- Select [1/2]:
230
+ One pipeline, 6 phases, Sonnet dev, parallel review + triage (CLI-aware
231
+ reviewer set). The questions Phase 0 asks are the workspace and autopilot,
232
+ never how much of the pipeline runs.
235
233
  ```
236
234
 
237
235
  ### Step 3 - Autopilot
@@ -6,6 +6,8 @@ argument-hint: "[\"<analysis-name>\"] [--no-cache] [--preview-conventions]"
6
6
 
7
7
  # multi-agent analysis - Feature Spec Analysis (v3)
8
8
 
9
+ > **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
10
+
9
11
  This command is **independent** from the orchestrator's Phase 1 analysis (which is a stack/findings detector inside `/multi-agent`). It produces a stakeholder-ready, platform-agnostic feature-spec document - up to 23 sections, trimmed to whatever the evidence supports - before any implementation starts. Each per-platform file is rendered by projecting concept-layer content onto repo-extracted conventions (Pass B).
10
12
 
11
13
  **Side-effect contract**: the command may write a local markdown file, post a Confluence page, or update a Jira description - but **never** creates branches, worktrees, commits, or PRs. It stops at the document.
@@ -8,6 +8,8 @@ not-for: create-jira, jira
8
8
 
9
9
  # multi-agent analysis-jira
10
10
 
11
+ > **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
12
+
11
13
  **Input**: $ARGUMENTS
12
14
 
13
15
  Reads an analysis document as a work breakdown and creates the tree. It never
@@ -6,6 +6,8 @@ argument-hint: "[path/to/analysis/<feature>-<platform>.md] [--autonomous]"
6
6
 
7
7
  # multi-agent analysis-resolve - Open Question Resolver for Analysis Docs
8
8
 
9
+ > **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
10
+
9
11
  Companion command to `/multi-agent:analysis`. Takes an `analysis/<feature>-<platform>.md` (template v3) and walks the user through the open rows of Section 20 (Riskler ve Acik Sorular / Risks and Open Questions) one row at a time, proposing source-labeled answer candidates and merging the chosen answer into the proper body section. The pattern follows a spec-resolver approach (walk the open rows, propose source-labeled answer candidates, merge the chosen answer) adapted to the analysis doc contract.
10
12
 
11
13
  **Core invariant - read this twice:** the analysis doc is authoritative and forward-looking. The resolver never invents an answer; if no source produces a credible candidate, the only options offered are Defer and Other. Each Section 20 row is its own decision; never blend candidates across rows.
@@ -6,6 +6,8 @@ argument-hint: '"task" - issue URL, Jira ID, free-text, or #id (for resume)'
6
6
 
7
7
  # multi-agent autopilot - Autonomous Pipeline
8
8
 
9
+ > **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
10
+
9
11
  **Input**: $ARGUMENTS
10
12
 
11
13
  > **Language (read FIRST)**: Before any status output, read `prefs.global.outputLanguage` and render every conversational line in it. `AskUserQuestion` renders its `question`, option `label`s and option `description`s in `outputLanguage`; only `header` stays English (<=12-char chip); external payload bodies follow `outputLanguage` too (identifiers, commit messages, branch names stay English). Full contract: `$HOME/.claude/multi-agent-refs/rules.md` "Language Application".
@@ -6,6 +6,8 @@ argument-hint: "(no arguments - opens the repo picker)"
6
6
 
7
7
  # multi-agent autopilot-on - continuous mode, on this machine
8
8
 
9
+ > **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
10
+
9
11
  Turns on the mode that picks work up without being asked. **No repo is included
10
12
  by default and none is ever added implicitly** - you choose, here, every time.
11
13
 
@@ -34,7 +34,7 @@ not a fault.
34
34
 
35
35
  When `swiftc` was available at `autopilot-on` time, `~/.claude/autopilot/bin/menubar`
36
36
  shows the same thing in the top right, refreshing on its own: one row per item
37
- with its id, whether it is running Full or Short, the phase as a fraction, the
37
+ with its id, the phase as a fraction, the
38
38
  elapsed time and the stack, then the queue, then what is waiting, then the PRs of
39
39
  the last day - each one clickable.
40
40
 
@@ -6,6 +6,8 @@ argument-hint: "(none - operates on current repo)"
6
6
 
7
7
  # multi-agent build-optimize - Xcode Build Performance Wrapper
8
8
 
9
+ > **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
10
+
9
11
  Wraps the vendored `xcode-build-orchestrator` skill (and its 5 specialist analyzers) so iOS users do not have to remember the orchestrator skill name. Recommend-first: no project files are modified without explicit developer approval. All output lands in `.build-benchmark/` in the current repo.
10
12
 
11
13
  > **Language**: Per `$HOME/.claude/multi-agent-refs/rules.md` Language Application matrix - instruction prose stays English (this file is read as a system prompt). `AskUserQuestion.question`, `.options[].label` and `.options[].description` follow `outputLanguage`; only `header` stays English (<=12-char chip). The emitted benchmark plan body (the upstream skill controls it) is English by upstream convention.
@@ -288,7 +288,7 @@ Emitted only when `reportContent.workSummary === true` and at least one of (`age
288
288
  - 1 accepted · 0 deferred · 0 rejected · approved=true
289
289
 
290
290
  #### Phases
291
- - 0 Init [done] · 1 Analysis [done] · 2 Planning [done] · 3 Dev [done] · 4 Review [done] · 5 Test skipped · 6 Commit [done] · 7 Report active
291
+ - 0 Init [done] · 1 Plan [done] · 2 Dev [done] · 3 Review [done] · 4 Commit [done] · 5 Report active
292
292
  ```
293
293
 
294
294
  **Post-hoc invocation** (task already finished, agent-state.json may be archived): pass `--branch` + `--base-branch` explicitly - the script still emits task header + changed-files + phases sections (scope + review outcome are state-dependent and will be absent if state is gone).
@@ -355,7 +355,7 @@ Applying the Jira table to the PR body is the recurring failure of this step: it
355
355
 
356
356
  Secondary artifacts ride along as one-line intents at the bottom of the preview: Wiki page list, wiki → Jira triad comment (when Wiki is selected and the triad is active per `$HOME/.claude/multi-agent-refs/issue-jira-triad.md`), Board status move (`{from} → {to}`), GitHub issue comment + Progress-flag sync (Phase 5 Step 1.5 - same gate, see `$HOME/.claude/multi-agent-refs/phases/phase-5-report.md`).
357
357
 
358
- **Then a single AskUserQuestion** (`question`/`label`/`description` in `outputLanguage`; `header` English):
358
+ **Then a single AskUserQuestion** (`picker-contract.md`; `question`/`label`/`description` in `outputLanguage`, `header` English):
359
359
 
360
360
  | Option | Behavior |
361
361
  |---|---|
@@ -6,6 +6,8 @@ argument-hint: "[\"<free-text description>\"] [figma-url] [swagger-url] - all
6
6
 
7
7
  # multi-agent create-jira - Standards-Compliant Jira Issue Creator
8
8
 
9
+ > **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
10
+
9
11
  **Input**: $ARGUMENTS
10
12
 
11
13
  Creates exactly one Jira issue, and only after explicit approval. The issue type (Task / Bug / Story) is asked at the start of every run. No branches, no commits, no worktrees, no pipeline chaining. The draft learns the target project's conventions from its recent same-type issues (summary format, labels, components, priority norms, test-scenario style) and offers active-sprint placement when a sprint is running.
@@ -67,7 +67,7 @@ In Claude Code also drive the native TaskList widget: fire all `TaskCreate` call
67
67
 
68
68
  ## STOP-AND-CONFIRM format
69
69
 
70
- Every Phase 0 / Phase 2 decision uses a native `AskUserQuestion` picker (numbered text menus are forbidden - `feedback_native_picker_ui`). Print a `Step <i>/<n>: <what it decides>` breadcrumb before each picker. Confirmation is required even with a single option; state inheritance from a previous run is FORBIDDEN.
70
+ Every Phase 0 / Phase 2 decision uses a native `AskUserQuestion` picker per `picker-contract.md` (numbered text menus are forbidden). Print a `Step <i>/<n>: <what it decides>` breadcrumb before each picker. Confirmation is required even with a single option; state inheritance from a previous run is FORBIDDEN.
71
71
 
72
72
  ---
73
73
 
@@ -12,7 +12,7 @@ Maps the Phase 3 triage output (`accepted` / `deferred` / `rejected`) onto the b
12
12
 
13
13
  ## When to use it
14
14
 
15
- - Between Phase 4 and Phase 5/6: answers "5 findings - which one belongs to which file?"
15
+ - Between Phase 4 and Phase 5: answers "5 findings - which one belongs to which file?"
16
16
  - Phase 5 reporting: a "review summary with code context" block to embed in the Wiki / Confluence output.
17
17
  - Post-hoc audit: for a closed PR, a "what did review flag, and how did the code change" diff.
18
18
 
@@ -7,6 +7,8 @@ allowed-tools: Bash, Read, AskUserQuestion
7
7
 
8
8
  # multi-agent forget - Remove a Routine
9
9
 
10
+ > **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
11
+
10
12
  Delete a routine the user saved via `/multi-agent:save`. Backed by `$HOME/.claude/scripts/routine-registry.mjs`.
11
13
 
12
14
  ## Steps
@@ -6,6 +6,8 @@ argument-hint: "[--older-than=<minutes>] [--abandoned [--days=N]] [--yes] - dr
6
6
 
7
7
  # multi-agent garbage-collect - Sweep /tmp Scratch
8
8
 
9
+ > **Pickers follow** `$HOME/.claude/multi-agent-refs/picker-contract.md`: never a one-option call, and branch on the option selected, not on its text.
10
+
9
11
  Each run writes ephemeral scratch under `/tmp`: picker state
10
12
  (`multi-agent-picker-*`), PR/review diffs (`multi-agent-review-*`), issue
11
13
  progress bodies (`issue-progress-*`), channel payloads (`channels-*`),
@@ -62,7 +64,7 @@ deletes nothing until you confirm.
62
64
  residue in the index (`.worktrees/{id}` recorded as a "Subproject commit"
63
65
  entry by a pre-guard `git add -A`), and a missing `.worktrees/` line in
64
66
  `.git/info/exclude`. Registered, healthy worktrees are NEVER touched -
65
- those belong to `/multi-agent:resume-local` / `/multi-agent:kill`.
67
+ those belong to `/multi-agent:resume` / `/multi-agent:kill`.
66
68
 
67
69
  - Output says `nothing to do` -> skip silently, no question.
68
70
  - Otherwise surface a second `AskUserQuestion` (in `outputLanguage`):
@@ -81,7 +83,7 @@ deletes nothing until you confirm.
81
83
 
82
84
  `offload-ref.sh` parks build logs, diffs and test output under
83
85
  `.multi-agent/refs/<node_id>.md`. In worktree modes that directory dies with
84
- the worktree; in the `--local` modes it lands in the real checkout, and
86
+ the worktree; when the run chose a local workspace it lands in the real checkout, and
85
87
  because it is gitignored it never shows up in `git status`. Only files
86
88
  matching the node-id shape are touched - anything else in that directory is
87
89
  left alone, and files inside the grace window are spared so a sweep cannot