@mmerterden/multi-agent-pipeline 20.0.0 → 20.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (90) hide show
  1. package/CHANGELOG.md +55 -0
  2. package/README.md +5 -5
  3. package/README.tr.md +5 -5
  4. package/SECURITY.md +3 -3
  5. package/docs/adr/0011-dormant-ci.md +10 -1
  6. package/docs/architecture.md +2 -2
  7. package/docs/ecosystem.md +5 -5
  8. package/docs/facts.json +8 -7
  9. package/install/_codex-agents.mjs +1 -1
  10. package/manifest.json +92 -64
  11. package/package.json +1 -1
  12. package/pipeline/agents/code-reviewer.md +2 -2
  13. package/pipeline/agents/dev-critic.md +5 -5
  14. package/pipeline/agents/security-auditor.md +80 -72
  15. package/pipeline/commands/figma-to-swiftui.md +1 -1
  16. package/pipeline/commands/multi-agent/SKILL.md +1 -1
  17. package/pipeline/commands/multi-agent/channels/SKILL.md +1 -1
  18. package/pipeline/commands/multi-agent/diff-explain/SKILL.md +1 -1
  19. package/pipeline/commands/multi-agent/help/SKILL.md +2 -0
  20. package/pipeline/commands/multi-agent/scan/SKILL.md +2 -2
  21. package/pipeline/commands/multi-agent/security-review/SKILL.md +52 -0
  22. package/pipeline/commands/multi-agent/sync/SKILL.md +3 -3
  23. package/pipeline/multi-agent-refs/component-dispatch.md +5 -5
  24. package/pipeline/multi-agent-refs/cross-cli-contract.md +6 -6
  25. package/pipeline/multi-agent-refs/features/security-audit.md +55 -0
  26. package/pipeline/multi-agent-refs/phases/modes.md +1 -1
  27. package/pipeline/multi-agent-refs/phases/phase-3-review.md +9 -15
  28. package/pipeline/multi-agent-refs/phases/phase-5-report.md +1 -1
  29. package/pipeline/multi-agent-refs/threat-model.md +39 -0
  30. package/pipeline/schemas/agent-state.schema.json +23 -0
  31. package/pipeline/schemas/phases.json +1 -2
  32. package/pipeline/schemas/prefs.schema.json +0 -4
  33. package/pipeline/schemas/reviewer-output.schema.json +99 -2
  34. package/pipeline/schemas/security-finding.schema.json +144 -0
  35. package/pipeline/scripts/_stack-routing.mjs +1 -0
  36. package/pipeline/scripts/gc-abandoned.sh +16 -9
  37. package/pipeline/scripts/render-work-summary.sh +7 -4
  38. package/pipeline/skills/.skill-manifest.json +47 -23
  39. package/pipeline/skills/.skills-index.json +75 -9
  40. package/pipeline/skills/shared/README.md +13 -7
  41. package/pipeline/skills/shared/core/multi-agent/SKILL.md +3 -4
  42. package/pipeline/skills/shared/core/multi-agent-scan/SKILL.md +2 -2
  43. package/pipeline/skills/shared/core/multi-agent-security-review/SKILL.md +29 -0
  44. package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +3 -3
  45. package/pipeline/skills/shared/external/android-architecture/SKILL.md +71 -0
  46. package/pipeline/skills/shared/external/android-architecture/references/patterns.md +142 -0
  47. package/pipeline/skills/shared/external/android-build-quality-gates/SKILL.md +314 -0
  48. package/pipeline/skills/shared/external/android-build-quality-gates/references/patterns.md +432 -0
  49. package/pipeline/skills/shared/external/android-datastore/SKILL.md +236 -0
  50. package/pipeline/skills/shared/external/android-datastore/references/patterns.md +297 -0
  51. package/pipeline/skills/shared/external/android-design-tokens-codegen/SKILL.md +249 -0
  52. package/pipeline/skills/shared/external/android-design-tokens-codegen/references/patterns.md +270 -0
  53. package/pipeline/skills/shared/external/android-jetpack-compose-expert/SKILL.md +62 -0
  54. package/pipeline/skills/shared/external/android-mvi-viewmodel/SKILL.md +255 -0
  55. package/pipeline/skills/shared/external/android-mvi-viewmodel/references/patterns.md +257 -0
  56. package/pipeline/skills/shared/external/android-performance/SKILL.md +86 -602
  57. package/pipeline/skills/shared/external/android-performance/references/patterns.md +659 -0
  58. package/pipeline/skills/shared/external/android-security/SKILL.md +117 -430
  59. package/pipeline/skills/shared/external/android-security/references/patterns.md +690 -0
  60. package/pipeline/skills/shared/external/{android_ui_verification → android-ui-verification}/SKILL.md +1 -1
  61. package/pipeline/skills/shared/external/api-security-best-practices/SKILL.md +35 -733
  62. package/pipeline/skills/shared/external/api-security-best-practices/references/auth.md +299 -0
  63. package/pipeline/skills/shared/external/api-security-best-practices/references/input-validation.md +255 -0
  64. package/pipeline/skills/shared/external/api-security-best-practices/references/rate-limiting.md +167 -0
  65. package/pipeline/skills/shared/external/app-intents/SKILL.md +39 -174
  66. package/pipeline/skills/shared/external/app-intents/references/appintents-advanced.md +178 -0
  67. package/pipeline/skills/shared/external/compose-components/SKILL.md +48 -0
  68. package/pipeline/skills/shared/external/compose-components/references/patterns.md +200 -0
  69. package/pipeline/skills/shared/external/compose-navigation/SKILL.md +66 -3
  70. package/pipeline/skills/shared/external/compose-navigation/references/patterns.md +191 -0
  71. package/pipeline/skills/shared/external/compose-testing/SKILL.md +107 -397
  72. package/pipeline/skills/shared/external/compose-testing/references/patterns.md +631 -0
  73. package/pipeline/skills/shared/external/gradle-kotlin-dsl/SKILL.md +121 -449
  74. package/pipeline/skills/shared/external/gradle-kotlin-dsl/references/patterns.md +715 -0
  75. package/pipeline/skills/shared/external/kotlin-coroutines-expert/SKILL.md +143 -0
  76. package/pipeline/skills/shared/external/mapkit-location/SKILL.md +27 -102
  77. package/pipeline/skills/shared/external/mapkit-location/references/mapkit-patterns.md +42 -0
  78. package/pipeline/skills/shared/external/retrofit-networking/SKILL.md +94 -383
  79. package/pipeline/skills/shared/external/retrofit-networking/references/patterns.md +640 -0
  80. package/pipeline/skills/shared/external/room-database/SKILL.md +101 -440
  81. package/pipeline/skills/shared/external/room-database/references/patterns.md +614 -0
  82. package/pipeline/skills/shared/external/security-review/SKILL.md +64 -0
  83. package/pipeline/skills/shared/external/security-review/references/owasp-mobile-top10-2024.md +53 -0
  84. package/pipeline/skills/shared/external/security-review/references/owasp-web-api-top10-2021.md +56 -0
  85. package/pipeline/skills/shared/external/storekit/SKILL.md +69 -343
  86. package/pipeline/skills/shared/external/storekit/references/core-patterns.md +371 -0
  87. package/pipeline/skills/shared/external/widgetkit/SKILL.md +25 -101
  88. package/pipeline/skills/shared/external/widgetkit/references/widgetkit-advanced.md +107 -0
  89. package/pipeline/skills/skills-index.md +8 -2
  90. package/pipeline/commands/security-review.md +0 -6
@@ -1,99 +1,107 @@
1
1
  ---
2
2
  name: security-auditor
3
- description: Security specialist - analyzes code for vulnerabilities and compliance issues
3
+ description: Security specialist - static application security review across mobile, web and backend. Emits reviewer-shaped JSON security findings (CVSS + CWE + evidence + fix) that merge into Phase 3 triage and block Phase 4.
4
4
  model: opus
5
5
  preferredModel: opus
6
- modelRationale: "Security reasoning + compliance catalog cross-reference (Apple ITMS, Google Play policy, OWASP) - false negatives are expensive; opus (top available tier) keeps the miss rate low on subtle vulnerabilities (auth-flow gaps, cert-pinning bypass, sensitive-data leaks)."
6
+ modelRationale: "Security reasoning + compliance catalog cross-reference (OWASP Web/API Top 10, OWASP Mobile Top 10, Apple ITMS, Google Play policy) plus honest confidence calibration on subtle vulnerabilities (auth-flow gaps, injection sinks, SSRF, cert-pinning bypass, sensitive-data leaks). False negatives are expensive and a wrong severity is worse than silence; opus (top available tier) keeps the miss rate and the miscalibration rate low."
7
7
  ---
8
8
 
9
- You are a mobile security auditor specializing in application security.
9
+ You are a static application security auditor. You read code, configuration and dependency manifests and report vulnerabilities. You do NOT run the application, fire payloads, or reach a live target: this is a defensive, read-only review. Everything you claim is grounded in something you can point to in the source.
10
10
 
11
- ## Your Focus
11
+ You cover first-party code in whatever stack the diff is in - Swift/Kotlin mobile, TypeScript/JavaScript/Python/Go/Java backends, web front-ends - plus the dependency manifests and secret-bearing configuration around it. You do not assume a mobile app.
12
12
 
13
- - OWASP Mobile Top 10 vulnerabilities
14
- - Apple App Store Review compliance
15
- - Google Play Store policy compliance
16
- - Data protection and encryption
17
- - Authentication and session management
18
- - Network security and certificate validation
19
- - Third-party SDK risk assessment
13
+ ## Threat model first
20
14
 
21
- ## Audit Categories
15
+ Before findings, establish the run-scoped threat model, four sections, and either write it to `.pipeline/threat-model.md` (standalone `/multi-agent:security-review`) or read the one Phase 3 already produced there. Every finding's severity is calibrated against it.
22
16
 
23
- ### Critical (Immediate Fix)
17
+ 1. **Attacker.** Who is the realistic adversary for THIS change - an unauthenticated internet caller, an authenticated low-privilege user, a malicious dependency, a co-located app on the device? Name the one that makes this diff interesting.
18
+ 2. **Trust boundaries.** Where does untrusted data cross into trusted code in the changed surface - a request body, a deep link, a WebView message, a file the app did not write, an env var an attacker can set?
19
+ 3. **Attack surface.** What did this diff actually add or touch - a new endpoint, a new query, a new deserialization, a new permission, a new dependency? A finding outside the touched surface is out of scope unless the diff made it reachable.
20
+ 4. **Severity calibration.** State the assumption each severity rests on: "critical assumes this route is unauthenticated in production." That assumption is what a reader disputes instead of the number.
24
21
 
25
- - Hardcoded credentials, API keys, secrets
26
- - Sensitive data in UserDefaults/SharedPreferences/plain files
27
- - Missing HTTPS / certificate pinning bypass
28
- - SQL injection, XSS in WebViews
29
- - Private API usage
22
+ ## What you look for
30
23
 
31
- ### High (Fix Before Release)
24
+ Join findings to a standard so they are checkable, not opinion:
32
25
 
33
- - Weak encryption / deprecated algorithms
34
- - Missing jailbreak/root detection
35
- - Insecure keychain configuration
36
- - Debug code in production (print, NSLog, Log.d, FLEX)
37
- - Missing privacy manifest declarations
26
+ - **Web / API** - OWASP Top 10 2021 (`A01:2021` .. `A10:2021`): broken access control, cryptographic failures, injection (SQL/NoSQL/command/LDAP), insecure design, security misconfiguration, vulnerable & outdated components, identification & auth failures, software & data integrity failures (insecure deserialization, unsigned updates), security logging failures, SSRF.
27
+ - **Mobile** - OWASP Mobile Top 10 2024 (`M1:2024` .. `M10:2024`): improper credential usage, inadequate supply-chain security, insecure auth/authorization, insufficient input/output validation, insecure communication, inadequate privacy controls, insufficient binary protection, security misconfiguration, insecure data storage, insufficient cryptography.
28
+ - **Cross-cutting** - hardcoded credentials/keys/secrets, sensitive data in plaintext stores (UserDefaults / SharedPreferences / localStorage / logs), missing or bypassed TLS and certificate pinning, weak or deprecated crypto, missing authz checks, unsafe deserialization, path traversal, and known-vulnerable dependencies (map to a `cve`).
38
29
 
39
- ### Medium (Plan to Fix)
30
+ Each finding gets a specific `CWE-NNN`, an OWASP category id, and a CVSS 3.1 base vector. Compute the score and band from the vector with the toolkit `security_cvss_score` tool rather than by hand, so the number cannot drift from the vector.
40
31
 
41
- - Excessive permissions
42
- - Missing input validation
43
- - Weak session management
44
- - ATS exceptions without justification
32
+ ## Evidence and honesty
45
33
 
46
- ## Output Format
47
-
48
- ```
49
- [SEVERITY] Category: Finding
50
- File: path/to/file:line
51
- Risk: What could go wrong
52
- Fix: How to fix it
53
- ```
54
-
55
- ## Store-compliance catalog cross-reference
56
-
57
- On store-relevant diffs, load the matching compliance skill's rule catalog and cite the ruleID + Apple ITMS / Google Play reference next to your finding. Binary invocation is NOT required at review time - the catalog alone is enough to annotate a diff. Full scan runs under `/multi-agent:test "store-ready"`.
34
+ - **Evidence** is what in the code proves the finding, cited by `file:line`. No evidence, no finding.
35
+ - **Counterevidence** is what would disprove it - a guard elsewhere, a framework default, an unreachable path. State it; a blank counterevidence asserts there is none.
36
+ - **Confidence** is `high|medium|low`. Low confidence does not mean stay silent - it means report with the counterevidence and let triage weigh it. Do not inflate a maybe into a certainty, and do not bury a certainty under hedging.
37
+ - Report real, reachable vulnerabilities in the touched surface. A theoretical risk with no path from the threat model's attacker is noted as `suggestion` at most, not `blocking`.
58
38
 
59
- ### When to load `pipeline/skills/shared/core/apple-archive-compliance/SKILL.md` (iOS)
39
+ ## Severity is derived, not chosen
60
40
 
61
- Trigger on any diff path matching:
41
+ Severity is the reviewer enum `blocking | important | suggestion`, set from the CVSS band so a security blocker blocks Phase 4 exactly like a reviewer blocker:
62
42
 
63
- - `**/Info.plist`
64
- - `**/PrivacyInfo.xcprivacy`
65
- - `**/*.entitlements`
66
- - `**/AppDelegate*.swift`, `**/SceneDelegate*.swift`, `**/*App.swift` (purpose strings)
67
- - `**/project.pbxproj` (Team ID, provisioning, code-signing settings)
43
+ | CVSS band | baseScore | severity |
44
+ | --------- | --------- | -------- |
45
+ | critical | 9.0-10.0 | blocking |
46
+ | high | 7.0-8.9 | blocking |
47
+ | medium | 4.0-6.9 | important |
48
+ | low | 0.1-3.9 | suggestion |
49
+ | none | 0.0 | suggestion |
68
50
 
69
- For each flagged line, append: `(apple-archive-compliance / <ruleID> - <Apple ref>)`
70
- Example: new `NSCameraUsageDescription` without justification → `(apple-archive-compliance / info-plist - Guideline 5.1.1)`
51
+ A severity that disagrees with the band is the one thing this contract forbids.
71
52
 
72
- ### When to load `pipeline/skills/shared/core/google-play-compliance/SKILL.md` (Android)
73
-
74
- Trigger on any diff path matching:
75
-
76
- - `**/AndroidManifest.xml` (permissions / exported / targetSdkVersion)
77
- - `**/build.gradle`, `**/build.gradle.kts` (signingConfig / minifyEnabled / targetSdkVersion / native abi filters)
78
- - `**/proguard-rules.pro`, `**/proguard-android*.txt`
79
- - `**/network_security_config.xml`
80
- - `**/gradle/libs.versions.toml` (only when dependency additions map to dangerous-permissions)
81
-
82
- For each flagged line, append: `(google-play-compliance / <ruleID> - <Play ref>)`
83
- Example: new `MANAGE_EXTERNAL_STORAGE` permission → `(google-play-compliance / dangerous-permissions - Policy - Permissions)`
53
+ ## Store-compliance catalog cross-reference
84
54
 
85
- ### What the catalog gives you
55
+ On store-relevant diffs, load the matching compliance skill's rule catalog and cite the ruleID + platform reference in the finding's `ruleId` and `criteriaSource`. Binary invocation is not required at review time - the catalog alone annotates a diff. Full scan runs under `/multi-agent:test "store-ready"`.
86
56
 
87
- The two SKILL.md files contain tabular rule catalogs:
57
+ - iOS - load `pipeline/skills/shared/core/apple-archive-compliance/SKILL.md` on diffs touching `**/Info.plist`, `**/PrivacyInfo.xcprivacy`, `**/*.entitlements`, `**/AppDelegate*.swift`, `**/SceneDelegate*.swift`, `**/*App.swift`, `**/project.pbxproj`. Set `ruleId` to the catalog rule and `criteriaSource: apple-archive-compliance`.
58
+ - Android - load `pipeline/skills/shared/core/google-play-compliance/SKILL.md` on diffs touching `**/AndroidManifest.xml`, `**/build.gradle`, `**/build.gradle.kts`, `**/proguard-rules.pro`, `**/network_security_config.xml`, `**/gradle/libs.versions.toml`. Set `ruleId` to the catalog rule and `criteriaSource: google-play-compliance`.
88
59
 
89
- - **apple-archive-compliance** - 18 rules (privacy-manifest, required-reason-api, info-plist, code-signing, embedded-sdk, entitlement, asset-validation, binary-size, team-id-consistency, provisioning-profile, swift-abi, extension-signing, ipv6-compliance, debug-tool-leak, production-hygiene, duplicate-resource, dead-reference, sdk-floor) with ITMS codes + App Store Review Guideline refs.
90
- - **google-play-compliance** - 21 rules across Technical / Security / Privacy / Hygiene categories with Play policy refs.
60
+ ## Output Format
91
61
 
92
- Cite rule + ref in your Output Format's `Category:` line so the code-reviewer and triage layers inherit the reference text unchanged.
62
+ Emit a single JSON object conforming to `pipeline/schemas/reviewer-output.schema.json`, so your findings merge into the Phase 3 reviewer set at Step 3.0 and reach triage and the Phase 4 gate unchanged. Every element of `findings[]` additionally conforms to `pipeline/schemas/security-finding.schema.json` - the security envelope rides along, and validate-reviewer.mjs (which ignores unknown fields) still passes it.
63
+
64
+ ```json
65
+ {
66
+ "findings": [
67
+ {
68
+ "severity": "blocking",
69
+ "file": "src/api/users.ts",
70
+ "line": 42,
71
+ "issue": "The path parameter id is concatenated into a raw SQL string.",
72
+ "fix": "Use a parameterized query with a bound id.",
73
+ "security": {
74
+ "owaspCategory": "A03:2021 Injection",
75
+ "cwe": "CWE-89",
76
+ "cvss": {
77
+ "vector": "CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H",
78
+ "baseScore": 9.8,
79
+ "band": "critical"
80
+ },
81
+ "evidence": "line 42: db.query('... WHERE id=' + req.params.id)",
82
+ "counterevidence": "None: id reaches the query with no validation between.",
83
+ "confidence": "high",
84
+ "confidenceRationale": "The sink and the untrusted source are both in the diff.",
85
+ "severityChangeConditions": "medium if the route is admin-only behind an auth guard",
86
+ "remediation": "Bind the id as a query parameter ($1) and pass it in the values array.",
87
+ "remediationDiff": {
88
+ "before": "db.query('SELECT * FROM users WHERE id=' + req.params.id)",
89
+ "after": "db.query('SELECT * FROM users WHERE id=$1', [req.params.id])"
90
+ },
91
+ "endpoint": "/api/users/:id",
92
+ "method": "GET",
93
+ "fixVerification": "Add a test that sends id=1 OR 1=1 and asserts a single-row result."
94
+ }
95
+ }
96
+ ],
97
+ "approved": false
98
+ }
99
+ ```
93
100
 
94
- ## Rules
101
+ Rules for the object:
95
102
 
96
- - Only report real vulnerabilities, not theoretical risks
97
- - Provide actionable fix suggestions
98
- - Reference Apple/Google docs or OWASP when relevant
99
- - For store-compliance findings, always cite the ruleID + policy reference from the catalog - don't paraphrase
103
+ - `approved` is `false` if any finding is `blocking`, `true` otherwise.
104
+ - Empty `findings[]` with `approved: true` is the correct output for a clean diff - do not invent findings to look thorough.
105
+ - One weakness per finding. Two weaknesses on one line are two findings.
106
+ - Cite a real `file` and `line` from the tree; never invent a path.
107
+ - Dependency findings (a known-vulnerable package rather than first-party code) carry the `cve`, set `owaspCategory` to `A06:2021`, `cwe` to the advisory's CWE, and `line: 0` on the manifest file.
@@ -1,5 +1,5 @@
1
1
  ---
2
- description: "Figma URL -> SwiftUI component generation (full 6-phase pipeline with SubPhases)"
2
+ description: "Figma URL -> SwiftUI component generation (full pipeline, Phase 0-8, with SubPhases)"
3
3
  allowed-tools: Read, Write, Edit, Glob, Grep, Bash, Agent, WebFetch, TaskCreate, TaskUpdate
4
4
  ---
5
5
 
@@ -131,7 +131,7 @@ This command uses lazy loading for token efficiency. Read the relevant sub-file
131
131
  | Web component task | `$HOME/.claude/multi-agent-refs/web-guide.md` |
132
132
  | Phase 1 or Phase 5 (knowledge) | `$HOME/.claude/multi-agent-refs/knowledge.md` |
133
133
  | Token lookup needed | `$HOME/.claude/multi-agent-refs/keychain.md` |
134
- | Audit tools (Phase 5/6) | `$HOME/.claude/multi-agent-refs/audit-guide.md` |
134
+ | Audit tools (Phase 3) | `$HOME/.claude/multi-agent-refs/audit-guide.md` |
135
135
  | `test` | `$HOME/.claude/commands/sim-test.md` (colon-form `/multi-agent:test` uses the delegate at `commands/multi-agent/test/SKILL.md`) |
136
136
  | `test-dark-mode` · `test-accessibility` · `test-dynamic-type` · `test-screenshots` | `$HOME/.claude/commands/sim-test.md`, scenario pinned by the command name (delegates at `commands/multi-agent/test-*/SKILL.md`) |
137
137
  | `manual-test` | `$HOME/.claude/commands/multi-agent/manual-test/SKILL.md` |
@@ -288,7 +288,7 @@ Emitted only when `reportContent.workSummary === true` and at least one of (`age
288
288
  - 1 accepted · 0 deferred · 0 rejected · approved=true
289
289
 
290
290
  #### Phases
291
- - 0 Init [done] · 1 Analysis [done] · 2 Planning [done] · 3 Dev [done] · 4 Review [done] · 5 Test skipped · 6 Commit [done] · 7 Report active
291
+ - 0 Init [done] · 1 Plan [done] · 2 Dev [done] · 3 Review [done] · 4 Commit [done] · 5 Report active
292
292
  ```
293
293
 
294
294
  **Post-hoc invocation** (task already finished, agent-state.json may be archived): pass `--branch` + `--base-branch` explicitly - the script still emits task header + changed-files + phases sections (scope + review outcome are state-dependent and will be absent if state is gone).
@@ -12,7 +12,7 @@ Maps the Phase 3 triage output (`accepted` / `deferred` / `rejected`) onto the b
12
12
 
13
13
  ## When to use it
14
14
 
15
- - Between Phase 4 and Phase 5/6: answers "5 findings - which one belongs to which file?"
15
+ - Between Phase 4 and Phase 5: answers "5 findings - which one belongs to which file?"
16
16
  - Phase 5 reporting: a "review summary with code context" block to embed in the Wiki / Confluence output.
17
17
  - Post-hoc audit: for a closed PR, a "what did review flag, and how did the code change" diff.
18
18
 
@@ -126,6 +126,7 @@ Post-Hoc & Side-Channel:
126
126
  /multi-agent:graph Build and query the repo code graph (symbols, imports, references), LLM-free
127
127
  /multi-agent:search Cross-task log search with smart ranking; --semantic queries triage corpus
128
128
  /multi-agent:scan Skill security scan against tiered pattern catalog
129
+ /multi-agent:security-review Standalone defensive static security review: threat model, OWASP + CWE + CVSS findings, dep inventory. No live target.
129
130
  /multi-agent:doctor Would a run work here? Layout, prefs, credentials, hooks; exit code is the verdict
130
131
  /multi-agent:refactor Best practices + bug hunt + upstream drift + toolkit MCP research -> one plan, approval, dev + sync
131
132
  /multi-agent:refactor backlog Decide the friction already recorded about the pipeline itself, nothing re-derived
@@ -401,6 +402,7 @@ Post-Hoc & Side-Channel:
401
402
  /multi-agent:graph Repo kod grafiğini kur ve sorgula (semboller, import'lar, referanslar), LLM'siz
402
403
  /multi-agent:search Task log'larında akıllı arama; --semantic triage corpus'unu sorgular
403
404
  /multi-agent:scan Skill güvenlik taraması (tiered pattern catalog)
405
+ /multi-agent:security-review Bağımsız savunma amaçlı statik güvenlik incelemesi: tehdit modeli, OWASP + CWE + CVSS bulguları, bağımlılık envanteri. Canlı hedef yok.
404
406
  /multi-agent:doctor Burada bir koşu çalışır mı? Yerleşim, prefs, kimlik bilgileri, hook'lar; çıkış kodu kararı verir
405
407
  /multi-agent:refactor Uyarlanmış best-practice + bug avı + upstream-drift + multi-agent-toolkit MCP araştırması -> tek plan, onay, dev + sync
406
408
  /multi-agent:refactor backlog Pipeline'ın kendisi hakkında kaydedilmiş sürtünmeyi karara bağlar, sıfırdan türetmez
@@ -66,10 +66,10 @@ Known legitimate domains are embedded in the scanner (github, anthropic, figma,
66
66
 
67
67
  ## Smoke - self-verification
68
68
 
69
- `pipeline/scripts/smoke-skill-scan.sh` confirms the scanner triggers on positive fixtures for every severity and produces no false positives on the real tree. This is a maintainer-repo gate (smoke scripts are excluded from the npm package): it runs from a pipeline checkout and in CI, where it must be 13/13 green before a team rollout.
69
+ `pipeline/scripts/smoke-skill-scan.sh` confirms the scanner triggers on positive fixtures for every severity and produces no false positives on the real tree. This is a maintainer-repo gate (smoke scripts are excluded from the npm package): it runs from a pipeline checkout and in CI, where it must be green before a team rollout.
70
70
 
71
71
  ## Integration points
72
72
 
73
73
  - **install.js pre-deploy hook** - every `install.js --all` runs an automatic high-threshold warn-only scan
74
- - **CI** - `.github/workflows/smoke.yml` step (strict mode, critical findings block the PR)
74
+ - **Smoke suite** - `smoke-skill-scan.sh` runs the scanner in strict mode under `npm test`, so a critical pattern fails the suite
75
75
  - **Standalone** - `/multi-agent:scan` (this command)
@@ -0,0 +1,52 @@
1
+ ---
2
+ description: "Run a standalone defensive, static security review of a diff, a branch or a repo: build a run-scoped threat model, emit reviewer-shaped findings joined to OWASP + CWE with CVSS scoring, evidence and before/after fixes, and inventory dependencies offline. No live target, no payloads. Use when reviewing code for security outside a full pipeline run, auditing a dependency set, or preparing a branch for a security sign-off."
3
+ description-tr: "Bir diff, branch veya repo üzerinde bağımsız, savunma amaçlı statik güvenlik incelemesi koşar: çalışmaya özel tehdit modeli kurar, OWASP + CWE'ye bağlı, CVSS puanlı, kanıtlı ve önce/sonra düzeltmeli reviewer biçiminde bulgular üretir, bağımlılıkları çevrimdışı envanterler. Canlı hedef yok, payload yok."
4
+ argument-hint: "[#N | repo#N | PR-URL | branch | path] - optional: a PR, a local branch, or a path to scope the review. If omitted (interactive), reviews the current branch diff against its base."
5
+ ---
6
+
7
+ # multi-agent security-review - standalone defensive static review
8
+
9
+ **Input**: $ARGUMENTS
10
+
11
+ The same security audit Phase 3 runs at Step 2.7, invoked on its own. Defensive and static: it reads code, configuration and dependency manifests and reports vulnerabilities. It never runs the target, fires a payload, or reaches a live host.
12
+
13
+ ## Scope resolution
14
+
15
+ Resolve what to review from `$ARGUMENTS`, same input shapes as `/multi-agent:review`:
16
+
17
+ - `#N` / `repo#N` / a PR URL (GitHub or Bitbucket Server) → that PR's diff.
18
+ - a local branch name → its diff against the base branch.
19
+ - a path → the files under it.
20
+ - omitted (interactive) → the current branch diff against its base; autopilot uses the current branch.
21
+
22
+ Compute the diff once and cap it the way Phase 3 Step 1.9 does; write the capped diff to `.pipeline/security-diff.txt` when it exceeds the budget.
23
+
24
+ ## Steps
25
+
26
+ ### 1. Threat model
27
+
28
+ Produce `.pipeline/threat-model.md` if absent (four sections: attacker, trust boundaries, attack surface, severity calibration), or read the existing one. Contract: `$HOME/.claude/multi-agent-refs/threat-model.md`. Mirror to `state.threatModel`.
29
+
30
+ ### 2. Method (the plugin skill)
31
+
32
+ Load the always-on `ai-common-toolkit:security-review` skill for the method: walk the OWASP checklist (Web/API Top 10 2021 for services and sites, Mobile Top 10 2024 for apps) against the changed surface, join each finding to a CWE, and score it. This command orchestrates; it does not re-derive the method.
33
+
34
+ ### 3. Findings (the security-auditor)
35
+
36
+ Dispatch `Agent(subagent_type: "security-auditor")` on the capped diff plus the threat model. It returns one `reviewer-output.schema.json` object whose `findings[]` each carry the `security` envelope of `security-finding.schema.json`. Compute every `security.cvss.baseScore` + `band` with the toolkit `security_cvss_score` tool, so a score cannot drift from its vector. Validate the object with `$HOME/.claude/scripts/validate-reviewer.mjs`; on failure, one self-correction rework then halt.
37
+
38
+ ### 4. Dependencies (offline)
39
+
40
+ Run the toolkit `security_dep_inventory` tool on the lockfiles in scope to get a normalized `{ecosystem, name, version}` inventory. When the analyst toolkit is registered, hand that inventory to `ai-analyst-toolkit:evidence-registry` for a known-CVE lookup; a vulnerable dependency becomes an `A06:2021` finding with the advisory's CWE and CVE. This command contacts nothing itself.
41
+
42
+ ### 5. Report
43
+
44
+ Write the findings to `.pipeline/security-findings.json` (the reviewer-output object) and render a human summary: count by severity, each blocking and important finding with its OWASP id, CWE, CVSS band, evidence, and before/after fix. An empty findings list with `approved: true` is the correct result for a clean review; do not invent findings.
45
+
46
+ ## Relationship to the pipeline
47
+
48
+ - Inside a run, the same audit is Phase 3 Step 2.7 (`$HOME/.claude/multi-agent-refs/features/security-audit.md`), triggered by the `security_path` diff-risk signal. This command is the standalone entry point; both share the persona, the schemas and the toolkit tools.
49
+ - This is not the `store-ready` device pass under `/multi-agent:test`, and it is not a secret scanner: the pipeline's `pre-commit-check.sh` already covers secrets.
50
+ - No writes to Jira, GitHub or Confluence: routing findings to a ticket is a separate, approval-gated step.
51
+
52
+ **Autopilot**: reviews the current branch, writes the report, and never opens a PR or comments.
@@ -65,7 +65,7 @@ Run every step automatically:
65
65
  ```
66
66
  Step 0: DOCTOR doctor.mjs - exit 2 or 4 stops the sync
67
67
  Step 1.5: DETECT Compare timestamps, find stale targets
68
- Step 2: COPILOT Claude Code -> Copilot CLI (instructions + 57 sub-command skills)
68
+ Step 2: COPILOT Claude Code -> Copilot CLI (instructions + 58 sub-command skills)
69
69
  Step 2b: CODEX Claude Code -> Codex CLI (1 router skill + 57 specs as refs + 8 agent TOML)
70
70
  Step 3: REPO Claude Code -> pipeline repo (genericized, personal data scrub, bash -n on all sh)
71
71
  Step 3c: PLUGINS pipeline shared/external -> multi-agent-plugins marketplace (rebuild knowledge/,
@@ -494,7 +494,7 @@ same 56 specs as reference files rather than as peer skills, via Step 2b - see
494
494
  |-------------|-------------|
495
495
  | `~/.claude/commands/multi-agent/{cmd}/SKILL.md` | `~/.copilot/skills/multi-agent-{cmd}/SKILL.md` |
496
496
 
497
- **57 commands are synced** (canonical inventory - must match `cross-cli-contract.md` section 1; drift = contract violation):
497
+ **58 commands are synced** (canonical inventory - must match `cross-cli-contract.md` section 1; drift = contract violation):
498
498
 
499
499
  ```
500
500
  analysis, analysis-jira, analysis-resolve, autopilot, autopilot-off, autopilot-on,
@@ -502,7 +502,7 @@ autopilot-status, build-optimize, channels, complaint-analysis, create-jira,
502
502
  design-check, diff-explain, doctor, feedback, forget, garbage-collect, graph, help,
503
503
  ios-coding-standard, issue, jira, kill, language, log,
504
504
  manual-test, model, prune-logs, prune-prompts, purge, refactor, resume, review, review-analysis, review-issue, review-jira, route-off, route-on, route-status,
505
- routines, save, scan, search, setup, stack, status, steer, store-ready, sync, test,
505
+ routines, save, scan, search, security-review, setup, stack, status, steer, store-ready, sync, test,
506
506
  test-accessibility, test-dark-mode, test-dynamic-type, test-screenshots,
507
507
  testflight-validation, uninstall, update
508
508
  ```
@@ -10,7 +10,7 @@
10
10
  - [Cross-CLI behaviour (intentional divergence)](#cross-cli-behaviour-intentional-divergence)
11
11
  <!-- /toc -->
12
12
 
13
- > **TLDR** - When `taskType === "component"` (Figma URL in task description or instruction-driven figma workflow), multi-agent Phase 3 **does not run the TDD loop**. It delegates the entire phase to the enabled `ai-<platform>-toolkit` **marketplace plugin's** component skill (`create-component`, falling back to `create-ui-component`) via the Skill tool. Implementation lives in the plugin; multi-agent's job is classification, dispatch, and state report. The pipeline no longer bundles its own `figma-to-component` orchestrator - component skills live in one place, the plugin marketplace.
13
+ > **TLDR** - When `taskType === "component"` (Figma URL in task description or instruction-driven figma workflow), multi-agent Phase 2 **does not run the TDD loop**. It delegates the entire phase to the enabled `ai-<platform>-toolkit` **marketplace plugin's** component skill (`create-component`, falling back to `create-ui-component`) via the Skill tool. Implementation lives in the plugin; multi-agent's job is classification, dispatch, and state report. The pipeline no longer bundles its own `figma-to-component` orchestrator - component skills live in one place, the plugin marketplace.
14
14
 
15
15
  This doc is referenced from `$HOME/.claude/multi-agent-refs/phases/phase-2-dev.md`. Keeping it separate lets `phase-2-dev.md` remain tight (it's already the largest phase doc) and gives the orchestrator-report contract a stable URL for both Claude-side and Copilot-side implementations.
16
16
 
@@ -83,7 +83,7 @@ Plugin skills are user-facing lifecycle skills; they do **not** accept an `agent
83
83
 
84
84
  ## Subphase contract (dispatch-layer owned)
85
85
 
86
- Because the plugin skill does not write pipeline state, **multi-agent's dispatch layer owns `state.phases["3"].subphases[]`**, not the skill. On dispatch, seed one coarse component-build subphase; on return, finalize it:
86
+ Because the plugin skill does not write pipeline state, **multi-agent's dispatch layer owns `state.phases["2"].subphases[]`**, not the skill. On dispatch, seed one coarse component-build subphase; on return, finalize it:
87
87
 
88
88
  ```json
89
89
  {
@@ -113,8 +113,8 @@ Multi-agent does **not** split the component task into per-repo Phase 2 runs -
113
113
 
114
114
  On failure (the plugin skill returns an unrecoverable build/test error, or the dispatch hits retry cap):
115
115
 
116
- 1. The dispatch layer persists `{error, buildLog, testLog}` to `state.phases["3"].errors[]`.
117
- 2. Multi-agent increments `state.phases["3"].retryCount`.
116
+ 1. The dispatch layer persists `{error, buildLog, testLog}` to `state.phases["2"].errors[]`.
117
+ 2. Multi-agent increments `state.phases["2"].retryCount`.
118
118
  3. Retry re-invokes the plugin skill (it is idempotent on an existing component; it reconciles rather than duplicating). There is no bundled `phase-<N>` resume anymore.
119
119
  4. Hard kill at `retryCount === 3` → surface the errors to the user, halt Phase 2. Do not loop indefinitely.
120
120
 
@@ -125,4 +125,4 @@ Component dispatch is **no longer byte-identical across CLIs** and that is by de
125
125
  - **Claude Code**: dispatches to the enabled `ai-<platform>-toolkit` marketplace plugin via the Skill tool (this doc).
126
126
  - **Copilot CLI**: has no plugin loader; it continues to use its standalone `~/.copilot/skills/figma-*` skill copies (frozen fallback). Copilot's resolution + progress lines follow those local skills.
127
127
 
128
- What still MUST match across CLIs: the `taskType === "component"` classification, the `state.phases["3"].subphases[]` shape the dispatch layer writes, and the Short-run elision semantics. Figma-skill *inventory* parity is no longer enforced. `smoke-cross-cli-behavior.sh` asserts only the classification + state-shape axis for components.
128
+ What still MUST match across CLIs: the `taskType === "component"` classification, the `state.phases["2"].subphases[]` shape the dispatch layer writes, and the Short-run elision semantics. Figma-skill *inventory* parity is no longer enforced. `smoke-cross-cli-behavior.sh` asserts only the classification + state-shape axis for components.
@@ -1,7 +1,7 @@
1
1
  # Cross-CLI Contract (Claude Code · Copilot CLI · Codex CLI)
2
2
 
3
3
  <!-- toc -->
4
- - [1. Command Inventory (57 commands)](#1-command-inventory-57-commands)
4
+ - [1. Command Inventory (58 commands)](#1-command-inventory-58-commands)
5
5
  - [2. Canonical Placeholder Vocabulary](#2-canonical-placeholder-vocabulary)
6
6
  - [2.6 Intentional structural divergence - thin dispatcher vs inlined orchestrator](#26-intentional-structural-divergence---thin-dispatcher-vs-inlined-orchestrator)
7
7
  - [2.7 One command, three different meanings: `/multi-agent:model`](#27-one-command-three-different-meanings-multi-agentmodel)
@@ -20,7 +20,7 @@
20
20
 
21
21
  ---
22
22
 
23
- ## 1. Command Inventory (57 commands)
23
+ ## 1. Command Inventory (58 commands)
24
24
 
25
25
  ```
26
26
  analysis, analysis-jira, analysis-resolve, autopilot, autopilot-off,
@@ -31,7 +31,7 @@ language, log, manual-test, model, prune-logs,
31
31
  prune-prompts, purge, refactor, resume, review,
32
32
  review-analysis, review-issue, review-jira, route-off, route-on,
33
33
  route-status, routines, save, scan, search,
34
- setup, stack, status, steer, store-ready, sync, test, test-accessibility,
34
+ security-review, setup, stack, status, steer, store-ready, sync, test, test-accessibility,
35
35
  test-dark-mode, test-dynamic-type, test-screenshots, testflight-validation,
36
36
  uninstall, update
37
37
  ```
@@ -86,7 +86,7 @@ Three rules make that work:
86
86
  > absent, and a component task on Copilot had nothing to dispatch to. The contract
87
87
  > documented a safety net the installer deleted.
88
88
 
89
- What still MUST match across CLIs for component tasks: the `taskType === "component"` classification, the `state.phases["3"].subphases[]` shape the dispatch layer writes, and the Short-run elision semantics. Resolution/routing is Claude-plugin vs Copilot-local-copy by design.
89
+ What still MUST match across CLIs for component tasks: the `taskType === "component"` classification, the `state.phases["2"].subphases[]` shape the dispatch layer writes, and the Short-run elision semantics. Resolution/routing is Claude-plugin vs Copilot-local-copy by design.
90
90
 
91
91
  ### 1.2 Store-compliance skills
92
92
 
@@ -97,7 +97,7 @@ Two parallel shared/core skills under `pipeline/skills/shared/core/` wrap extern
97
97
  | `apple-archive-compliance` | `pipeline/skills/shared/core/apple-archive-compliance/` | `ios_app_store_audit` MCP tool (in `@mmerterden/multi-agent-toolkit-mcp` ≥ v3.0.0) | 18 (Apple ITMS + App Store Review Guidelines) |
98
98
  | `google-play-compliance` | `pipeline/skills/shared/core/google-play-compliance/` | bundletool + aapt2 + apksigner | 21 (4 categories: Technical / Security / Privacy / Hygiene) |
99
99
 
100
- Both skills are wired to 4 consumers: `/multi-agent:test "store-ready"` (primary), Phase 4 Security Auditor (`pipeline/agents/security-auditor.md`), `/multi-agent:review` + SKILL.md counterpart, `/multi-agent:channels` PR-body auto-augmentation. Contract enforced by `smoke-compliance-skills.sh`.
100
+ Both skills are wired to 4 consumers: `/multi-agent:test "store-ready"` (primary), Phase 3 Security Auditor (`pipeline/agents/security-auditor.md`), `/multi-agent:review` + SKILL.md counterpart, `/multi-agent:channels` PR-body auto-augmentation. Contract enforced by `smoke-compliance-skills.sh`.
101
101
 
102
102
  ### 1.3 Figma-skill routing from multi-agent (superseded)
103
103
 
@@ -222,7 +222,7 @@ Phase 3 runs three reviewers everywhere, but the diversity those three buy is no
222
222
  same on every host. Copilot CLI gets cross-VENDOR disagreement for free: GPT-5.4 sits
223
223
  beside two Claude models. Claude Code and Codex each run a one-vendor panel - three
224
224
  Anthropic models on one, three OpenAI models on the other - so the same three-way
225
- agreement is weaker evidence there, and Phase 4 says so in the triage note on a
225
+ agreement is weaker evidence there, and Phase 3 says so in the triage note on a
226
226
  borderline finding.
227
227
 
228
228
  Where the budget goes instead, when vendor diversity is unavailable:
@@ -0,0 +1,55 @@
1
+ # Security audit (Phase 3 Step 2.7)
2
+
3
+ The mechanics of the conditional security audit that runs inside the Phase 3 review window and merges at Step 3.0. The phase doc carries the trigger and the merge; the rest is here.
4
+
5
+ ## Trigger (deterministic, no new classifier)
6
+
7
+ Run the audit when any of these holds:
8
+
9
+ - Step 1.75 scored `security_path` on a file in the diff (the same signal that already forces `full` review at Step 1.77).
10
+ - The base branch is a release branch.
11
+ - The run is the standalone `/multi-agent:security-review` command, which dispatches straight to this step.
12
+
13
+ There is no `--audit` flag: the trigger is the diff, not a word the user has to remember. When none of these holds, the audit does not run, `$SECURITY_AUDIT_JSON` stays empty, and the Step 3.0 merge is a no-op.
14
+
15
+ ## Threat model first
16
+
17
+ The auditor reads `.pipeline/threat-model.md` if Phase 1 or a prior step wrote it, and produces it (four sections) if absent. Contract: `$HOME/.claude/multi-agent-refs/threat-model.md`. It is keyed to repo+branch and mirrored to `state.threatModel` so a resume reuses it rather than re-deriving it.
18
+
19
+ ## Dispatch
20
+
21
+ One `Agent(subagent_type: "security-auditor")` on the same capped diff the reviewers saw, plus the threat model and the resolved `${CRITERIA}` block:
22
+
23
+ ```
24
+ Agent(subagent_type: "security-auditor", prompt: "<threat-model + capped diff + ${CRITERIA}>")
25
+ ```
26
+
27
+ Persona contract: `~/.claude/agents/security-auditor.md`.
28
+
29
+ ## Output is reviewer-shaped, and that is the point
30
+
31
+ The auditor returns one object conforming to `$HOME/.claude/schemas/reviewer-output.schema.json`, whose `findings[]` each carry the `security` envelope of `security-finding.schema.json`: OWASP category, CWE, CVSS vector + band, evidence, counterevidence, confidence, remediation, and the before/after fix. Severity is the reviewer enum, derived from the CVSS band:
32
+
33
+ | CVSS band | severity |
34
+ | --------- | -------- |
35
+ | critical / high | blocking |
36
+ | medium | important |
37
+ | low / none | suggestion |
38
+
39
+ Compute `security.cvss.baseScore` + `band` with the toolkit `security_cvss_score` tool, not by hand, so the score cannot drift from the vector.
40
+
41
+ ## Validate, persist, merge
42
+
43
+ Validate with the same gate protocol as a reviewer - the exit code decides, not the LLM turn:
44
+
45
+ ```bash
46
+ printf '%s' "$SECURITY_AUDIT_JSON" | node "$HOME/.claude/scripts/validate-reviewer.mjs" -
47
+ ```
48
+
49
+ On validator failure: one self-correction rework, then HALT the phase (identical to the reviewer output-contract gate). Persist the object to `state.reviewIterations[<iteration>].securityAudit` and hold it in `$SECURITY_AUDIT_JSON`.
50
+
51
+ Because the output is reviewer-shaped, the Step 3.0 merge appends its `findings[]` alongside the test-integrity findings. A `blocking` security finding then reaches triage, and a triage-accepted blocker blocks Phase 4 - the "critical security items block the commit" contract is wired here, not merely stated.
52
+
53
+ ## Store-compliance cross-reference
54
+
55
+ On store-relevant diffs the auditor also cites the Apple ITMS / Google Play catalog rule on the finding via `ruleId` + `criteriaSource` (`apple-archive-compliance` / `google-play-compliance`). This is annotation on the same finding, not a second pass; the device-level `store-ready` Gate under `/multi-agent:test` stays separate.
@@ -58,7 +58,7 @@ The generic "zero-interaction" contract covers Phases 0-4 only. Phase 5 channels
58
58
  - **30-minute timeout** - if user does not respond, session ends cleanly:
59
59
  - External delivery aborted (no silent apply - prevents accidental Jira comments / Confluence pages).
60
60
  - Internal capture (`agent-log.md`, telemetry, knowledge base) STILL runs.
61
- - State persisted as `{phase: 7, waitingFor: "user-channels-choice", channelsTimeout: true}`.
61
+ - State persisted as `{phase: 5, waitingFor: "user-channels-choice", channelsTimeout: true}`.
62
62
  - Resume: `/multi-agent:resume <task-id>` re-opens menu with same inputs.
63
63
  - Post-hoc `/multi-agent:channels <task>` never times out - user invoked it explicitly.
64
64
 
@@ -260,7 +260,7 @@ Each reviewer inherits the `code-reviewer` agent's focus areas (Security, Archit
260
260
  | Python | `ai-backend-toolkit:api-security-best-practices` | `ai-backend-toolkit:fastapi-pro` | `ai-backend-toolkit:python-patterns` |
261
261
  | Node.js | `ai-backend-toolkit:api-security-best-practices` | `ai-backend-toolkit:nodejs-backend-patterns` | `ai-frontend-toolkit:typescript-patterns` |
262
262
  | Docker | `ai-backend-toolkit:docker-expert` | `ai-backend-toolkit:docker-expert` | `ai-backend-toolkit:ci-cd-pipelines` |
263
- | Generic | `security-review` | `ai-backend-toolkit:clean-code` | `ai-backend-toolkit:clean-code` |
263
+ | Generic | `ai-common-toolkit:security-review` | `ai-backend-toolkit:clean-code` | `ai-backend-toolkit:clean-code` |
264
264
 
265
265
  ##### 2.1 Previous-round findings (iteration >= 2) and 2.2 scope self-check (every iteration)
266
266
 
@@ -298,6 +298,10 @@ Skip only when the diff has no UI change. Record the outcome in
298
298
 
299
299
  Step 2 produces N reviewer-output objects (one per dispatched reviewer), each conforming to `$HOME/.claude/schemas/reviewer-output.schema.json`. They are persisted to `state.reviewIterations[<iteration>].reviewers[]` and consumed by Step 3 (Fable triage) - never by Phase 4 directly. The triage step (below) is the producer of the only review artifact Phase 4 reads, conforming to `$HOME/.claude/schemas/triage-output.schema.json`.
300
300
 
301
+ #### Step 2.7 - Security audit (conditional, produces reviewer-shaped findings)
302
+
303
+ Runs when Step 1.75 scored `security_path`, on a release branch, or when `/multi-agent:security-review` dispatches here. The `security-auditor` returns reviewer `findings[]`, each with a `security` envelope (`security-finding.schema.json`) and severity from the CVSS band, held in `$SECURITY_AUDIT_JSON` for the merge so a `blocking` one reaches triage and blocks Phase 4. Mechanics: `~/.claude/multi-agent-refs/features/security-audit.md`.
304
+
301
305
  **Subagent return format** - each reviewer returns JSON conforming to `$HOME/.claude/schemas/reviewer-output.schema.json`:
302
306
 
303
307
  ```json
@@ -371,14 +375,14 @@ ANON=$(jq -n --argjson r "$REVIEWERS_JSON" --arg t "$TASK_ID" --argjson i "$ITER
371
375
 
372
376
  `$REVIEWERS_JSON` is `state.reviewIterations[i].reviewers`. Findings come back with `foundBy: "Source A|B|C"` and every identity key removed. Persist the map to `state.reviewIterations[i].anonymizationMap` for Phase 5 per-reviewer telemetry, and **never put the map in a prompt**.
373
377
 
374
- Then append the Step 1.76 test-integrity findings, so they are adjudicated rather than never seen:
378
+ Then append the Step 1.76 test-integrity and Step 2.7 security-audit findings, so they are adjudicated rather than never seen:
375
379
 
376
380
  ```bash
377
- MERGED=$(jq -s '.[0] + (.[1].findings // [])' \
378
- <(printf '%s' "$ANON") <(printf '%s' "${TEST_INTEGRITY_JSON:-{\}}"))
381
+ MERGED=$(jq -s '.[0] + (.[1].findings // []) + (.[2].findings // [])' \
382
+ <(printf '%s' "$ANON") <(printf '%s' "${TEST_INTEGRITY_JSON:-{\}}") <(printf '%s' "${SECURITY_AUDIT_JSON:-{\}}"))
379
383
  ```
380
384
 
381
- Deterministic findings keep `tag: test_integrity` and carry no `foundBy`: a reviewer finding may be a hallucination, a gate finding is a fact.
385
+ Deterministic findings keep `tag: test_integrity` and no `foundBy`: a reviewer finding may hallucinate, a gate finding is a fact. Security-audit findings carry their `security` envelope and `foundBy: "security-auditor"`; an empty `$SECURITY_AUDIT_JSON` contributes nothing.
382
386
 
383
387
  ##### 3.1 Short-circuit: no findings
384
388
 
@@ -788,16 +792,6 @@ Results included in Phase 5 report. MCP tools preferred when available - conci
788
792
 
789
793
  **Snapshot regression flow (optional):** when the task changes a stable component, capture a screenshot before the change (baseline) and after (current), then call `ios_visual_diff({baseline, current, max_diff_pct: 1.0})`. Threshold can be relaxed for animated / non-deterministic regions - keep `max_diff_pct ≤ 1.0` for static layouts.
790
794
 
791
- #### Security Audit (store-readiness)
792
-
793
- When the task touches authentication, keychain, network, or is scheduled for an imminent release, launch the `security-auditor` subagent to run an OWASP Mobile Top 10 pass plus App Store / Play Store compliance checks:
794
-
795
- ```
796
- Agent(subagent_type: "security-auditor", prompt: "<diff + context>")
797
- ```
798
-
799
- Returns severity-tagged findings (Critical / High / Medium). Critical items block Phase 4 just like Phase 3 blockers; High items are logged and surfaced in Phase 5 report. Skipped by default - opt-in for release branches or on explicit `/multi-agent "<task>" --audit` flag.
800
-
801
795
  #### Telemetry - token forwarding
802
796
 
803
797
  When the security-auditor or any other Phase 3 sub-agent runs, forward its token totals so Phase 5's Cost Breakdown captures Phase 3:
@@ -25,7 +25,7 @@ On entry: `phase-tracker.sh sub 5 <N> "<name>" in_progress`. On exit: `completed
25
25
  Phase 5 is the single exception to the autopilot zero-interaction rule: every mode, Full or Short, attended or autopilot, pauses at the channels multi-select menu. Full contract in `$HOME/.claude/multi-agent-refs/phases/modes.md`:
26
26
 
27
27
  - **30-min timeout** - if user does not respond, session ends cleanly. External delivery is aborted (no silent apply of defaults - prevents accidental Jira comments / Confluence pages). Internal capture (Steps 2 + 3 below) STILL runs so `agent-log.md` + knowledge base are persisted.
28
- - **Resumable** - state written as `{status: "awaiting_input", phase: 7, waitingFor: "user-channels-choice", channelsInput: <state-bundle>}`. User can `/multi-agent:resume <task-id>` any time later; channels menu re-opens with same inputs.
28
+ - **Resumable** - state written as `{status: "awaiting_input", phase: 5, waitingFor: "user-channels-choice", channelsInput: <state-bundle>}`. User can `/multi-agent:resume <task-id>` any time later; channels menu re-opens with same inputs.
29
29
  - **Timeout log line:** `Phase 5: channels menu timeout (30 min) - session ended, resume with /multi-agent:resume {taskId}`.
30
30
 
31
31
  ---
@@ -0,0 +1,39 @@
1
+ # Threat model contract
2
+
3
+ The run-scoped threat model is the frame every security finding is calibrated against. It is produced by the security-auditor at Phase 3 Step 2.7 (or by `/multi-agent:security-review` running standalone), read by any later step that revisits security, and mirrored to `state.threatModel.path` so a resume re-uses it instead of re-deriving it.
4
+
5
+ ## Where it lives
6
+
7
+ `.pipeline/threat-model.md` in the run's worktree, keyed to repo + branch. One per run. If a fresh one already exists for this repo+branch (a prior step or a Phase 1 pass wrote it), read it; do not overwrite. If it is absent, produce it before emitting any finding.
8
+
9
+ Mirror the resolved path and a content hash to `state.threatModel = { path, sha, producedAt, producedBy }` so resume and Phase 5 can find it without re-reading the tree.
10
+
11
+ ## The four sections (all required, in order)
12
+
13
+ A threat model with a missing section is not a threat model; the auditor treats a missing section as "produce it," not "skip it."
14
+
15
+ ### 1. Attacker
16
+
17
+ The one realistic adversary for THIS change. Name it concretely: an unauthenticated internet caller, an authenticated low-privilege user, a malicious or compromised dependency, a co-located app on the device, a user with physical access. Not "attackers" in the abstract - the specific actor whose capability makes this diff interesting. If the diff has no plausible attacker, say so; that is a valid, short threat model and most findings then cap at `suggestion`.
18
+
19
+ ### 2. Trust boundaries
20
+
21
+ Where untrusted data crosses into trusted code within the changed surface: a request body or query parameter, a deep link or universal link, a WebView `postMessage`, a file the app did not write, an environment variable an attacker can set, a response from a third-party service treated as safe. List the boundaries the diff touches, each with the `file:line` where the crossing happens.
22
+
23
+ ### 3. Attack surface
24
+
25
+ What this diff actually added or touched: a new endpoint, a new query or ORM call, a new deserialization, a new permission, a new dependency, a new crypto usage, a new storage write. A finding outside this surface is out of scope unless the diff made it reachable - and if it did, say how. This section is what keeps the audit anchored to the change instead of drifting into a whole-repo review.
26
+
27
+ ### 4. Severity calibration
28
+
29
+ The assumption each severity rests on, stated so a reader disputes the assumption rather than the number. "Critical assumes this route is unauthenticated in production; if it is admin-only the same finding is medium." "High assumes the secret reaches a log that ships off-device." Every `blocking` finding must trace to an assumption named here; a blocker resting on an unstated assumption is miscalibrated.
30
+
31
+ ## Shape
32
+
33
+ Plain Markdown, four `##` sections with those names, human-readable. It is evidence for a person and context for the auditor, not a machine artifact - no schema. Keep it short: a page, not a report. Cite `file:line` where a boundary or surface item has one.
34
+
35
+ ## What it is not
36
+
37
+ - Not a whole-repo model. It is scoped to the diff under review.
38
+ - Not a dynamic test plan. This is a static, read-only posture: no live target, no payloads, no exploitation. `fixVerification` on a finding names the empirical check a human or a later dynamic pass would run; the threat model does not run it.
39
+ - Not embedded in the analysis document. The security-review command must run standalone, so the model lives in its own file and is produced on demand.