@mmerterden/multi-agent-pipeline 20.0.0 → 20.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +55 -0
- package/README.md +5 -5
- package/README.tr.md +5 -5
- package/SECURITY.md +3 -3
- package/docs/adr/0011-dormant-ci.md +10 -1
- package/docs/architecture.md +2 -2
- package/docs/ecosystem.md +5 -5
- package/docs/facts.json +8 -7
- package/install/_codex-agents.mjs +1 -1
- package/manifest.json +92 -64
- package/package.json +1 -1
- package/pipeline/agents/code-reviewer.md +2 -2
- package/pipeline/agents/dev-critic.md +5 -5
- package/pipeline/agents/security-auditor.md +80 -72
- package/pipeline/commands/figma-to-swiftui.md +1 -1
- package/pipeline/commands/multi-agent/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/channels/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/diff-explain/SKILL.md +1 -1
- package/pipeline/commands/multi-agent/help/SKILL.md +2 -0
- package/pipeline/commands/multi-agent/scan/SKILL.md +2 -2
- package/pipeline/commands/multi-agent/security-review/SKILL.md +52 -0
- package/pipeline/commands/multi-agent/sync/SKILL.md +3 -3
- package/pipeline/multi-agent-refs/component-dispatch.md +5 -5
- package/pipeline/multi-agent-refs/cross-cli-contract.md +6 -6
- package/pipeline/multi-agent-refs/features/security-audit.md +55 -0
- package/pipeline/multi-agent-refs/phases/modes.md +1 -1
- package/pipeline/multi-agent-refs/phases/phase-3-review.md +9 -15
- package/pipeline/multi-agent-refs/phases/phase-5-report.md +1 -1
- package/pipeline/multi-agent-refs/threat-model.md +39 -0
- package/pipeline/schemas/agent-state.schema.json +23 -0
- package/pipeline/schemas/phases.json +1 -2
- package/pipeline/schemas/prefs.schema.json +0 -4
- package/pipeline/schemas/reviewer-output.schema.json +99 -2
- package/pipeline/schemas/security-finding.schema.json +144 -0
- package/pipeline/scripts/_stack-routing.mjs +1 -0
- package/pipeline/scripts/gc-abandoned.sh +16 -9
- package/pipeline/scripts/render-work-summary.sh +7 -4
- package/pipeline/skills/.skill-manifest.json +47 -23
- package/pipeline/skills/.skills-index.json +75 -9
- package/pipeline/skills/shared/README.md +13 -7
- package/pipeline/skills/shared/core/multi-agent/SKILL.md +3 -4
- package/pipeline/skills/shared/core/multi-agent-scan/SKILL.md +2 -2
- package/pipeline/skills/shared/core/multi-agent-security-review/SKILL.md +29 -0
- package/pipeline/skills/shared/core/multi-agent-sync/SKILL.md +3 -3
- package/pipeline/skills/shared/external/android-architecture/SKILL.md +71 -0
- package/pipeline/skills/shared/external/android-architecture/references/patterns.md +142 -0
- package/pipeline/skills/shared/external/android-build-quality-gates/SKILL.md +314 -0
- package/pipeline/skills/shared/external/android-build-quality-gates/references/patterns.md +432 -0
- package/pipeline/skills/shared/external/android-datastore/SKILL.md +236 -0
- package/pipeline/skills/shared/external/android-datastore/references/patterns.md +297 -0
- package/pipeline/skills/shared/external/android-design-tokens-codegen/SKILL.md +249 -0
- package/pipeline/skills/shared/external/android-design-tokens-codegen/references/patterns.md +270 -0
- package/pipeline/skills/shared/external/android-jetpack-compose-expert/SKILL.md +62 -0
- package/pipeline/skills/shared/external/android-mvi-viewmodel/SKILL.md +255 -0
- package/pipeline/skills/shared/external/android-mvi-viewmodel/references/patterns.md +257 -0
- package/pipeline/skills/shared/external/android-performance/SKILL.md +86 -602
- package/pipeline/skills/shared/external/android-performance/references/patterns.md +659 -0
- package/pipeline/skills/shared/external/android-security/SKILL.md +117 -430
- package/pipeline/skills/shared/external/android-security/references/patterns.md +690 -0
- package/pipeline/skills/shared/external/{android_ui_verification → android-ui-verification}/SKILL.md +1 -1
- package/pipeline/skills/shared/external/api-security-best-practices/SKILL.md +35 -733
- package/pipeline/skills/shared/external/api-security-best-practices/references/auth.md +299 -0
- package/pipeline/skills/shared/external/api-security-best-practices/references/input-validation.md +255 -0
- package/pipeline/skills/shared/external/api-security-best-practices/references/rate-limiting.md +167 -0
- package/pipeline/skills/shared/external/app-intents/SKILL.md +39 -174
- package/pipeline/skills/shared/external/app-intents/references/appintents-advanced.md +178 -0
- package/pipeline/skills/shared/external/compose-components/SKILL.md +48 -0
- package/pipeline/skills/shared/external/compose-components/references/patterns.md +200 -0
- package/pipeline/skills/shared/external/compose-navigation/SKILL.md +66 -3
- package/pipeline/skills/shared/external/compose-navigation/references/patterns.md +191 -0
- package/pipeline/skills/shared/external/compose-testing/SKILL.md +107 -397
- package/pipeline/skills/shared/external/compose-testing/references/patterns.md +631 -0
- package/pipeline/skills/shared/external/gradle-kotlin-dsl/SKILL.md +121 -449
- package/pipeline/skills/shared/external/gradle-kotlin-dsl/references/patterns.md +715 -0
- package/pipeline/skills/shared/external/kotlin-coroutines-expert/SKILL.md +143 -0
- package/pipeline/skills/shared/external/mapkit-location/SKILL.md +27 -102
- package/pipeline/skills/shared/external/mapkit-location/references/mapkit-patterns.md +42 -0
- package/pipeline/skills/shared/external/retrofit-networking/SKILL.md +94 -383
- package/pipeline/skills/shared/external/retrofit-networking/references/patterns.md +640 -0
- package/pipeline/skills/shared/external/room-database/SKILL.md +101 -440
- package/pipeline/skills/shared/external/room-database/references/patterns.md +614 -0
- package/pipeline/skills/shared/external/security-review/SKILL.md +64 -0
- package/pipeline/skills/shared/external/security-review/references/owasp-mobile-top10-2024.md +53 -0
- package/pipeline/skills/shared/external/security-review/references/owasp-web-api-top10-2021.md +56 -0
- package/pipeline/skills/shared/external/storekit/SKILL.md +69 -343
- package/pipeline/skills/shared/external/storekit/references/core-patterns.md +371 -0
- package/pipeline/skills/shared/external/widgetkit/SKILL.md +25 -101
- package/pipeline/skills/shared/external/widgetkit/references/widgetkit-advanced.md +107 -0
- package/pipeline/skills/skills-index.md +8 -2
- package/pipeline/commands/security-review.md +0 -6
|
@@ -1,99 +1,107 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: security-auditor
|
|
3
|
-
description: Security specialist -
|
|
3
|
+
description: Security specialist - static application security review across mobile, web and backend. Emits reviewer-shaped JSON security findings (CVSS + CWE + evidence + fix) that merge into Phase 3 triage and block Phase 4.
|
|
4
4
|
model: opus
|
|
5
5
|
preferredModel: opus
|
|
6
|
-
modelRationale: "Security reasoning + compliance catalog cross-reference (
|
|
6
|
+
modelRationale: "Security reasoning + compliance catalog cross-reference (OWASP Web/API Top 10, OWASP Mobile Top 10, Apple ITMS, Google Play policy) plus honest confidence calibration on subtle vulnerabilities (auth-flow gaps, injection sinks, SSRF, cert-pinning bypass, sensitive-data leaks). False negatives are expensive and a wrong severity is worse than silence; opus (top available tier) keeps the miss rate and the miscalibration rate low."
|
|
7
7
|
---
|
|
8
8
|
|
|
9
|
-
You are a
|
|
9
|
+
You are a static application security auditor. You read code, configuration and dependency manifests and report vulnerabilities. You do NOT run the application, fire payloads, or reach a live target: this is a defensive, read-only review. Everything you claim is grounded in something you can point to in the source.
|
|
10
10
|
|
|
11
|
-
|
|
11
|
+
You cover first-party code in whatever stack the diff is in - Swift/Kotlin mobile, TypeScript/JavaScript/Python/Go/Java backends, web front-ends - plus the dependency manifests and secret-bearing configuration around it. You do not assume a mobile app.
|
|
12
12
|
|
|
13
|
-
|
|
14
|
-
- Apple App Store Review compliance
|
|
15
|
-
- Google Play Store policy compliance
|
|
16
|
-
- Data protection and encryption
|
|
17
|
-
- Authentication and session management
|
|
18
|
-
- Network security and certificate validation
|
|
19
|
-
- Third-party SDK risk assessment
|
|
13
|
+
## Threat model first
|
|
20
14
|
|
|
21
|
-
|
|
15
|
+
Before findings, establish the run-scoped threat model, four sections, and either write it to `.pipeline/threat-model.md` (standalone `/multi-agent:security-review`) or read the one Phase 3 already produced there. Every finding's severity is calibrated against it.
|
|
22
16
|
|
|
23
|
-
|
|
17
|
+
1. **Attacker.** Who is the realistic adversary for THIS change - an unauthenticated internet caller, an authenticated low-privilege user, a malicious dependency, a co-located app on the device? Name the one that makes this diff interesting.
|
|
18
|
+
2. **Trust boundaries.** Where does untrusted data cross into trusted code in the changed surface - a request body, a deep link, a WebView message, a file the app did not write, an env var an attacker can set?
|
|
19
|
+
3. **Attack surface.** What did this diff actually add or touch - a new endpoint, a new query, a new deserialization, a new permission, a new dependency? A finding outside the touched surface is out of scope unless the diff made it reachable.
|
|
20
|
+
4. **Severity calibration.** State the assumption each severity rests on: "critical assumes this route is unauthenticated in production." That assumption is what a reader disputes instead of the number.
|
|
24
21
|
|
|
25
|
-
|
|
26
|
-
- Sensitive data in UserDefaults/SharedPreferences/plain files
|
|
27
|
-
- Missing HTTPS / certificate pinning bypass
|
|
28
|
-
- SQL injection, XSS in WebViews
|
|
29
|
-
- Private API usage
|
|
22
|
+
## What you look for
|
|
30
23
|
|
|
31
|
-
|
|
24
|
+
Join findings to a standard so they are checkable, not opinion:
|
|
32
25
|
|
|
33
|
-
-
|
|
34
|
-
-
|
|
35
|
-
-
|
|
36
|
-
- Debug code in production (print, NSLog, Log.d, FLEX)
|
|
37
|
-
- Missing privacy manifest declarations
|
|
26
|
+
- **Web / API** - OWASP Top 10 2021 (`A01:2021` .. `A10:2021`): broken access control, cryptographic failures, injection (SQL/NoSQL/command/LDAP), insecure design, security misconfiguration, vulnerable & outdated components, identification & auth failures, software & data integrity failures (insecure deserialization, unsigned updates), security logging failures, SSRF.
|
|
27
|
+
- **Mobile** - OWASP Mobile Top 10 2024 (`M1:2024` .. `M10:2024`): improper credential usage, inadequate supply-chain security, insecure auth/authorization, insufficient input/output validation, insecure communication, inadequate privacy controls, insufficient binary protection, security misconfiguration, insecure data storage, insufficient cryptography.
|
|
28
|
+
- **Cross-cutting** - hardcoded credentials/keys/secrets, sensitive data in plaintext stores (UserDefaults / SharedPreferences / localStorage / logs), missing or bypassed TLS and certificate pinning, weak or deprecated crypto, missing authz checks, unsafe deserialization, path traversal, and known-vulnerable dependencies (map to a `cve`).
|
|
38
29
|
|
|
39
|
-
|
|
30
|
+
Each finding gets a specific `CWE-NNN`, an OWASP category id, and a CVSS 3.1 base vector. Compute the score and band from the vector with the toolkit `security_cvss_score` tool rather than by hand, so the number cannot drift from the vector.
|
|
40
31
|
|
|
41
|
-
|
|
42
|
-
- Missing input validation
|
|
43
|
-
- Weak session management
|
|
44
|
-
- ATS exceptions without justification
|
|
32
|
+
## Evidence and honesty
|
|
45
33
|
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
File: path/to/file:line
|
|
51
|
-
Risk: What could go wrong
|
|
52
|
-
Fix: How to fix it
|
|
53
|
-
```
|
|
54
|
-
|
|
55
|
-
## Store-compliance catalog cross-reference
|
|
56
|
-
|
|
57
|
-
On store-relevant diffs, load the matching compliance skill's rule catalog and cite the ruleID + Apple ITMS / Google Play reference next to your finding. Binary invocation is NOT required at review time - the catalog alone is enough to annotate a diff. Full scan runs under `/multi-agent:test "store-ready"`.
|
|
34
|
+
- **Evidence** is what in the code proves the finding, cited by `file:line`. No evidence, no finding.
|
|
35
|
+
- **Counterevidence** is what would disprove it - a guard elsewhere, a framework default, an unreachable path. State it; a blank counterevidence asserts there is none.
|
|
36
|
+
- **Confidence** is `high|medium|low`. Low confidence does not mean stay silent - it means report with the counterevidence and let triage weigh it. Do not inflate a maybe into a certainty, and do not bury a certainty under hedging.
|
|
37
|
+
- Report real, reachable vulnerabilities in the touched surface. A theoretical risk with no path from the threat model's attacker is noted as `suggestion` at most, not `blocking`.
|
|
58
38
|
|
|
59
|
-
|
|
39
|
+
## Severity is derived, not chosen
|
|
60
40
|
|
|
61
|
-
|
|
41
|
+
Severity is the reviewer enum `blocking | important | suggestion`, set from the CVSS band so a security blocker blocks Phase 4 exactly like a reviewer blocker:
|
|
62
42
|
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
-
|
|
66
|
-
|
|
67
|
-
|
|
43
|
+
| CVSS band | baseScore | severity |
|
|
44
|
+
| --------- | --------- | -------- |
|
|
45
|
+
| critical | 9.0-10.0 | blocking |
|
|
46
|
+
| high | 7.0-8.9 | blocking |
|
|
47
|
+
| medium | 4.0-6.9 | important |
|
|
48
|
+
| low | 0.1-3.9 | suggestion |
|
|
49
|
+
| none | 0.0 | suggestion |
|
|
68
50
|
|
|
69
|
-
|
|
70
|
-
Example: new `NSCameraUsageDescription` without justification → `(apple-archive-compliance / info-plist - Guideline 5.1.1)`
|
|
51
|
+
A severity that disagrees with the band is the one thing this contract forbids.
|
|
71
52
|
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
Trigger on any diff path matching:
|
|
75
|
-
|
|
76
|
-
- `**/AndroidManifest.xml` (permissions / exported / targetSdkVersion)
|
|
77
|
-
- `**/build.gradle`, `**/build.gradle.kts` (signingConfig / minifyEnabled / targetSdkVersion / native abi filters)
|
|
78
|
-
- `**/proguard-rules.pro`, `**/proguard-android*.txt`
|
|
79
|
-
- `**/network_security_config.xml`
|
|
80
|
-
- `**/gradle/libs.versions.toml` (only when dependency additions map to dangerous-permissions)
|
|
81
|
-
|
|
82
|
-
For each flagged line, append: `(google-play-compliance / <ruleID> - <Play ref>)`
|
|
83
|
-
Example: new `MANAGE_EXTERNAL_STORAGE` permission → `(google-play-compliance / dangerous-permissions - Policy - Permissions)`
|
|
53
|
+
## Store-compliance catalog cross-reference
|
|
84
54
|
|
|
85
|
-
|
|
55
|
+
On store-relevant diffs, load the matching compliance skill's rule catalog and cite the ruleID + platform reference in the finding's `ruleId` and `criteriaSource`. Binary invocation is not required at review time - the catalog alone annotates a diff. Full scan runs under `/multi-agent:test "store-ready"`.
|
|
86
56
|
|
|
87
|
-
|
|
57
|
+
- iOS - load `pipeline/skills/shared/core/apple-archive-compliance/SKILL.md` on diffs touching `**/Info.plist`, `**/PrivacyInfo.xcprivacy`, `**/*.entitlements`, `**/AppDelegate*.swift`, `**/SceneDelegate*.swift`, `**/*App.swift`, `**/project.pbxproj`. Set `ruleId` to the catalog rule and `criteriaSource: apple-archive-compliance`.
|
|
58
|
+
- Android - load `pipeline/skills/shared/core/google-play-compliance/SKILL.md` on diffs touching `**/AndroidManifest.xml`, `**/build.gradle`, `**/build.gradle.kts`, `**/proguard-rules.pro`, `**/network_security_config.xml`, `**/gradle/libs.versions.toml`. Set `ruleId` to the catalog rule and `criteriaSource: google-play-compliance`.
|
|
88
59
|
|
|
89
|
-
|
|
90
|
-
- **google-play-compliance** - 21 rules across Technical / Security / Privacy / Hygiene categories with Play policy refs.
|
|
60
|
+
## Output Format
|
|
91
61
|
|
|
92
|
-
|
|
62
|
+
Emit a single JSON object conforming to `pipeline/schemas/reviewer-output.schema.json`, so your findings merge into the Phase 3 reviewer set at Step 3.0 and reach triage and the Phase 4 gate unchanged. Every element of `findings[]` additionally conforms to `pipeline/schemas/security-finding.schema.json` - the security envelope rides along, and validate-reviewer.mjs (which ignores unknown fields) still passes it.
|
|
63
|
+
|
|
64
|
+
```json
|
|
65
|
+
{
|
|
66
|
+
"findings": [
|
|
67
|
+
{
|
|
68
|
+
"severity": "blocking",
|
|
69
|
+
"file": "src/api/users.ts",
|
|
70
|
+
"line": 42,
|
|
71
|
+
"issue": "The path parameter id is concatenated into a raw SQL string.",
|
|
72
|
+
"fix": "Use a parameterized query with a bound id.",
|
|
73
|
+
"security": {
|
|
74
|
+
"owaspCategory": "A03:2021 Injection",
|
|
75
|
+
"cwe": "CWE-89",
|
|
76
|
+
"cvss": {
|
|
77
|
+
"vector": "CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H",
|
|
78
|
+
"baseScore": 9.8,
|
|
79
|
+
"band": "critical"
|
|
80
|
+
},
|
|
81
|
+
"evidence": "line 42: db.query('... WHERE id=' + req.params.id)",
|
|
82
|
+
"counterevidence": "None: id reaches the query with no validation between.",
|
|
83
|
+
"confidence": "high",
|
|
84
|
+
"confidenceRationale": "The sink and the untrusted source are both in the diff.",
|
|
85
|
+
"severityChangeConditions": "medium if the route is admin-only behind an auth guard",
|
|
86
|
+
"remediation": "Bind the id as a query parameter ($1) and pass it in the values array.",
|
|
87
|
+
"remediationDiff": {
|
|
88
|
+
"before": "db.query('SELECT * FROM users WHERE id=' + req.params.id)",
|
|
89
|
+
"after": "db.query('SELECT * FROM users WHERE id=$1', [req.params.id])"
|
|
90
|
+
},
|
|
91
|
+
"endpoint": "/api/users/:id",
|
|
92
|
+
"method": "GET",
|
|
93
|
+
"fixVerification": "Add a test that sends id=1 OR 1=1 and asserts a single-row result."
|
|
94
|
+
}
|
|
95
|
+
}
|
|
96
|
+
],
|
|
97
|
+
"approved": false
|
|
98
|
+
}
|
|
99
|
+
```
|
|
93
100
|
|
|
94
|
-
|
|
101
|
+
Rules for the object:
|
|
95
102
|
|
|
96
|
-
-
|
|
97
|
-
-
|
|
98
|
-
-
|
|
99
|
-
-
|
|
103
|
+
- `approved` is `false` if any finding is `blocking`, `true` otherwise.
|
|
104
|
+
- Empty `findings[]` with `approved: true` is the correct output for a clean diff - do not invent findings to look thorough.
|
|
105
|
+
- One weakness per finding. Two weaknesses on one line are two findings.
|
|
106
|
+
- Cite a real `file` and `line` from the tree; never invent a path.
|
|
107
|
+
- Dependency findings (a known-vulnerable package rather than first-party code) carry the `cve`, set `owaspCategory` to `A06:2021`, `cwe` to the advisory's CWE, and `line: 0` on the manifest file.
|
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
---
|
|
2
|
-
description: "Figma URL -> SwiftUI component generation (full
|
|
2
|
+
description: "Figma URL -> SwiftUI component generation (full pipeline, Phase 0-8, with SubPhases)"
|
|
3
3
|
allowed-tools: Read, Write, Edit, Glob, Grep, Bash, Agent, WebFetch, TaskCreate, TaskUpdate
|
|
4
4
|
---
|
|
5
5
|
|
|
@@ -131,7 +131,7 @@ This command uses lazy loading for token efficiency. Read the relevant sub-file
|
|
|
131
131
|
| Web component task | `$HOME/.claude/multi-agent-refs/web-guide.md` |
|
|
132
132
|
| Phase 1 or Phase 5 (knowledge) | `$HOME/.claude/multi-agent-refs/knowledge.md` |
|
|
133
133
|
| Token lookup needed | `$HOME/.claude/multi-agent-refs/keychain.md` |
|
|
134
|
-
| Audit tools (Phase
|
|
134
|
+
| Audit tools (Phase 3) | `$HOME/.claude/multi-agent-refs/audit-guide.md` |
|
|
135
135
|
| `test` | `$HOME/.claude/commands/sim-test.md` (colon-form `/multi-agent:test` uses the delegate at `commands/multi-agent/test/SKILL.md`) |
|
|
136
136
|
| `test-dark-mode` · `test-accessibility` · `test-dynamic-type` · `test-screenshots` | `$HOME/.claude/commands/sim-test.md`, scenario pinned by the command name (delegates at `commands/multi-agent/test-*/SKILL.md`) |
|
|
137
137
|
| `manual-test` | `$HOME/.claude/commands/multi-agent/manual-test/SKILL.md` |
|
|
@@ -288,7 +288,7 @@ Emitted only when `reportContent.workSummary === true` and at least one of (`age
|
|
|
288
288
|
- 1 accepted · 0 deferred · 0 rejected · approved=true
|
|
289
289
|
|
|
290
290
|
#### Phases
|
|
291
|
-
- 0 Init [done] · 1
|
|
291
|
+
- 0 Init [done] · 1 Plan [done] · 2 Dev [done] · 3 Review [done] · 4 Commit [done] · 5 Report active
|
|
292
292
|
```
|
|
293
293
|
|
|
294
294
|
**Post-hoc invocation** (task already finished, agent-state.json may be archived): pass `--branch` + `--base-branch` explicitly - the script still emits task header + changed-files + phases sections (scope + review outcome are state-dependent and will be absent if state is gone).
|
|
@@ -12,7 +12,7 @@ Maps the Phase 3 triage output (`accepted` / `deferred` / `rejected`) onto the b
|
|
|
12
12
|
|
|
13
13
|
## When to use it
|
|
14
14
|
|
|
15
|
-
- Between Phase 4 and Phase 5
|
|
15
|
+
- Between Phase 4 and Phase 5: answers "5 findings - which one belongs to which file?"
|
|
16
16
|
- Phase 5 reporting: a "review summary with code context" block to embed in the Wiki / Confluence output.
|
|
17
17
|
- Post-hoc audit: for a closed PR, a "what did review flag, and how did the code change" diff.
|
|
18
18
|
|
|
@@ -126,6 +126,7 @@ Post-Hoc & Side-Channel:
|
|
|
126
126
|
/multi-agent:graph Build and query the repo code graph (symbols, imports, references), LLM-free
|
|
127
127
|
/multi-agent:search Cross-task log search with smart ranking; --semantic queries triage corpus
|
|
128
128
|
/multi-agent:scan Skill security scan against tiered pattern catalog
|
|
129
|
+
/multi-agent:security-review Standalone defensive static security review: threat model, OWASP + CWE + CVSS findings, dep inventory. No live target.
|
|
129
130
|
/multi-agent:doctor Would a run work here? Layout, prefs, credentials, hooks; exit code is the verdict
|
|
130
131
|
/multi-agent:refactor Best practices + bug hunt + upstream drift + toolkit MCP research -> one plan, approval, dev + sync
|
|
131
132
|
/multi-agent:refactor backlog Decide the friction already recorded about the pipeline itself, nothing re-derived
|
|
@@ -401,6 +402,7 @@ Post-Hoc & Side-Channel:
|
|
|
401
402
|
/multi-agent:graph Repo kod grafiğini kur ve sorgula (semboller, import'lar, referanslar), LLM'siz
|
|
402
403
|
/multi-agent:search Task log'larında akıllı arama; --semantic triage corpus'unu sorgular
|
|
403
404
|
/multi-agent:scan Skill güvenlik taraması (tiered pattern catalog)
|
|
405
|
+
/multi-agent:security-review Bağımsız savunma amaçlı statik güvenlik incelemesi: tehdit modeli, OWASP + CWE + CVSS bulguları, bağımlılık envanteri. Canlı hedef yok.
|
|
404
406
|
/multi-agent:doctor Burada bir koşu çalışır mı? Yerleşim, prefs, kimlik bilgileri, hook'lar; çıkış kodu kararı verir
|
|
405
407
|
/multi-agent:refactor Uyarlanmış best-practice + bug avı + upstream-drift + multi-agent-toolkit MCP araştırması -> tek plan, onay, dev + sync
|
|
406
408
|
/multi-agent:refactor backlog Pipeline'ın kendisi hakkında kaydedilmiş sürtünmeyi karara bağlar, sıfırdan türetmez
|
|
@@ -66,10 +66,10 @@ Known legitimate domains are embedded in the scanner (github, anthropic, figma,
|
|
|
66
66
|
|
|
67
67
|
## Smoke - self-verification
|
|
68
68
|
|
|
69
|
-
`pipeline/scripts/smoke-skill-scan.sh` confirms the scanner triggers on positive fixtures for every severity and produces no false positives on the real tree. This is a maintainer-repo gate (smoke scripts are excluded from the npm package): it runs from a pipeline checkout and in CI, where it must be
|
|
69
|
+
`pipeline/scripts/smoke-skill-scan.sh` confirms the scanner triggers on positive fixtures for every severity and produces no false positives on the real tree. This is a maintainer-repo gate (smoke scripts are excluded from the npm package): it runs from a pipeline checkout and in CI, where it must be green before a team rollout.
|
|
70
70
|
|
|
71
71
|
## Integration points
|
|
72
72
|
|
|
73
73
|
- **install.js pre-deploy hook** - every `install.js --all` runs an automatic high-threshold warn-only scan
|
|
74
|
-
- **
|
|
74
|
+
- **Smoke suite** - `smoke-skill-scan.sh` runs the scanner in strict mode under `npm test`, so a critical pattern fails the suite
|
|
75
75
|
- **Standalone** - `/multi-agent:scan` (this command)
|
|
@@ -0,0 +1,52 @@
|
|
|
1
|
+
---
|
|
2
|
+
description: "Run a standalone defensive, static security review of a diff, a branch or a repo: build a run-scoped threat model, emit reviewer-shaped findings joined to OWASP + CWE with CVSS scoring, evidence and before/after fixes, and inventory dependencies offline. No live target, no payloads. Use when reviewing code for security outside a full pipeline run, auditing a dependency set, or preparing a branch for a security sign-off."
|
|
3
|
+
description-tr: "Bir diff, branch veya repo üzerinde bağımsız, savunma amaçlı statik güvenlik incelemesi koşar: çalışmaya özel tehdit modeli kurar, OWASP + CWE'ye bağlı, CVSS puanlı, kanıtlı ve önce/sonra düzeltmeli reviewer biçiminde bulgular üretir, bağımlılıkları çevrimdışı envanterler. Canlı hedef yok, payload yok."
|
|
4
|
+
argument-hint: "[#N | repo#N | PR-URL | branch | path] - optional: a PR, a local branch, or a path to scope the review. If omitted (interactive), reviews the current branch diff against its base."
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# multi-agent security-review - standalone defensive static review
|
|
8
|
+
|
|
9
|
+
**Input**: $ARGUMENTS
|
|
10
|
+
|
|
11
|
+
The same security audit Phase 3 runs at Step 2.7, invoked on its own. Defensive and static: it reads code, configuration and dependency manifests and reports vulnerabilities. It never runs the target, fires a payload, or reaches a live host.
|
|
12
|
+
|
|
13
|
+
## Scope resolution
|
|
14
|
+
|
|
15
|
+
Resolve what to review from `$ARGUMENTS`, same input shapes as `/multi-agent:review`:
|
|
16
|
+
|
|
17
|
+
- `#N` / `repo#N` / a PR URL (GitHub or Bitbucket Server) → that PR's diff.
|
|
18
|
+
- a local branch name → its diff against the base branch.
|
|
19
|
+
- a path → the files under it.
|
|
20
|
+
- omitted (interactive) → the current branch diff against its base; autopilot uses the current branch.
|
|
21
|
+
|
|
22
|
+
Compute the diff once and cap it the way Phase 3 Step 1.9 does; write the capped diff to `.pipeline/security-diff.txt` when it exceeds the budget.
|
|
23
|
+
|
|
24
|
+
## Steps
|
|
25
|
+
|
|
26
|
+
### 1. Threat model
|
|
27
|
+
|
|
28
|
+
Produce `.pipeline/threat-model.md` if absent (four sections: attacker, trust boundaries, attack surface, severity calibration), or read the existing one. Contract: `$HOME/.claude/multi-agent-refs/threat-model.md`. Mirror to `state.threatModel`.
|
|
29
|
+
|
|
30
|
+
### 2. Method (the plugin skill)
|
|
31
|
+
|
|
32
|
+
Load the always-on `ai-common-toolkit:security-review` skill for the method: walk the OWASP checklist (Web/API Top 10 2021 for services and sites, Mobile Top 10 2024 for apps) against the changed surface, join each finding to a CWE, and score it. This command orchestrates; it does not re-derive the method.
|
|
33
|
+
|
|
34
|
+
### 3. Findings (the security-auditor)
|
|
35
|
+
|
|
36
|
+
Dispatch `Agent(subagent_type: "security-auditor")` on the capped diff plus the threat model. It returns one `reviewer-output.schema.json` object whose `findings[]` each carry the `security` envelope of `security-finding.schema.json`. Compute every `security.cvss.baseScore` + `band` with the toolkit `security_cvss_score` tool, so a score cannot drift from its vector. Validate the object with `$HOME/.claude/scripts/validate-reviewer.mjs`; on failure, one self-correction rework then halt.
|
|
37
|
+
|
|
38
|
+
### 4. Dependencies (offline)
|
|
39
|
+
|
|
40
|
+
Run the toolkit `security_dep_inventory` tool on the lockfiles in scope to get a normalized `{ecosystem, name, version}` inventory. When the analyst toolkit is registered, hand that inventory to `ai-analyst-toolkit:evidence-registry` for a known-CVE lookup; a vulnerable dependency becomes an `A06:2021` finding with the advisory's CWE and CVE. This command contacts nothing itself.
|
|
41
|
+
|
|
42
|
+
### 5. Report
|
|
43
|
+
|
|
44
|
+
Write the findings to `.pipeline/security-findings.json` (the reviewer-output object) and render a human summary: count by severity, each blocking and important finding with its OWASP id, CWE, CVSS band, evidence, and before/after fix. An empty findings list with `approved: true` is the correct result for a clean review; do not invent findings.
|
|
45
|
+
|
|
46
|
+
## Relationship to the pipeline
|
|
47
|
+
|
|
48
|
+
- Inside a run, the same audit is Phase 3 Step 2.7 (`$HOME/.claude/multi-agent-refs/features/security-audit.md`), triggered by the `security_path` diff-risk signal. This command is the standalone entry point; both share the persona, the schemas and the toolkit tools.
|
|
49
|
+
- This is not the `store-ready` device pass under `/multi-agent:test`, and it is not a secret scanner: the pipeline's `pre-commit-check.sh` already covers secrets.
|
|
50
|
+
- No writes to Jira, GitHub or Confluence: routing findings to a ticket is a separate, approval-gated step.
|
|
51
|
+
|
|
52
|
+
**Autopilot**: reviews the current branch, writes the report, and never opens a PR or comments.
|
|
@@ -65,7 +65,7 @@ Run every step automatically:
|
|
|
65
65
|
```
|
|
66
66
|
Step 0: DOCTOR doctor.mjs - exit 2 or 4 stops the sync
|
|
67
67
|
Step 1.5: DETECT Compare timestamps, find stale targets
|
|
68
|
-
Step 2: COPILOT Claude Code -> Copilot CLI (instructions +
|
|
68
|
+
Step 2: COPILOT Claude Code -> Copilot CLI (instructions + 58 sub-command skills)
|
|
69
69
|
Step 2b: CODEX Claude Code -> Codex CLI (1 router skill + 57 specs as refs + 8 agent TOML)
|
|
70
70
|
Step 3: REPO Claude Code -> pipeline repo (genericized, personal data scrub, bash -n on all sh)
|
|
71
71
|
Step 3c: PLUGINS pipeline shared/external -> multi-agent-plugins marketplace (rebuild knowledge/,
|
|
@@ -494,7 +494,7 @@ same 56 specs as reference files rather than as peer skills, via Step 2b - see
|
|
|
494
494
|
|-------------|-------------|
|
|
495
495
|
| `~/.claude/commands/multi-agent/{cmd}/SKILL.md` | `~/.copilot/skills/multi-agent-{cmd}/SKILL.md` |
|
|
496
496
|
|
|
497
|
-
**
|
|
497
|
+
**58 commands are synced** (canonical inventory - must match `cross-cli-contract.md` section 1; drift = contract violation):
|
|
498
498
|
|
|
499
499
|
```
|
|
500
500
|
analysis, analysis-jira, analysis-resolve, autopilot, autopilot-off, autopilot-on,
|
|
@@ -502,7 +502,7 @@ autopilot-status, build-optimize, channels, complaint-analysis, create-jira,
|
|
|
502
502
|
design-check, diff-explain, doctor, feedback, forget, garbage-collect, graph, help,
|
|
503
503
|
ios-coding-standard, issue, jira, kill, language, log,
|
|
504
504
|
manual-test, model, prune-logs, prune-prompts, purge, refactor, resume, review, review-analysis, review-issue, review-jira, route-off, route-on, route-status,
|
|
505
|
-
routines, save, scan, search, setup, stack, status, steer, store-ready, sync, test,
|
|
505
|
+
routines, save, scan, search, security-review, setup, stack, status, steer, store-ready, sync, test,
|
|
506
506
|
test-accessibility, test-dark-mode, test-dynamic-type, test-screenshots,
|
|
507
507
|
testflight-validation, uninstall, update
|
|
508
508
|
```
|
|
@@ -10,7 +10,7 @@
|
|
|
10
10
|
- [Cross-CLI behaviour (intentional divergence)](#cross-cli-behaviour-intentional-divergence)
|
|
11
11
|
<!-- /toc -->
|
|
12
12
|
|
|
13
|
-
> **TLDR** - When `taskType === "component"` (Figma URL in task description or instruction-driven figma workflow), multi-agent Phase
|
|
13
|
+
> **TLDR** - When `taskType === "component"` (Figma URL in task description or instruction-driven figma workflow), multi-agent Phase 2 **does not run the TDD loop**. It delegates the entire phase to the enabled `ai-<platform>-toolkit` **marketplace plugin's** component skill (`create-component`, falling back to `create-ui-component`) via the Skill tool. Implementation lives in the plugin; multi-agent's job is classification, dispatch, and state report. The pipeline no longer bundles its own `figma-to-component` orchestrator - component skills live in one place, the plugin marketplace.
|
|
14
14
|
|
|
15
15
|
This doc is referenced from `$HOME/.claude/multi-agent-refs/phases/phase-2-dev.md`. Keeping it separate lets `phase-2-dev.md` remain tight (it's already the largest phase doc) and gives the orchestrator-report contract a stable URL for both Claude-side and Copilot-side implementations.
|
|
16
16
|
|
|
@@ -83,7 +83,7 @@ Plugin skills are user-facing lifecycle skills; they do **not** accept an `agent
|
|
|
83
83
|
|
|
84
84
|
## Subphase contract (dispatch-layer owned)
|
|
85
85
|
|
|
86
|
-
Because the plugin skill does not write pipeline state, **multi-agent's dispatch layer owns `state.phases["
|
|
86
|
+
Because the plugin skill does not write pipeline state, **multi-agent's dispatch layer owns `state.phases["2"].subphases[]`**, not the skill. On dispatch, seed one coarse component-build subphase; on return, finalize it:
|
|
87
87
|
|
|
88
88
|
```json
|
|
89
89
|
{
|
|
@@ -113,8 +113,8 @@ Multi-agent does **not** split the component task into per-repo Phase 2 runs -
|
|
|
113
113
|
|
|
114
114
|
On failure (the plugin skill returns an unrecoverable build/test error, or the dispatch hits retry cap):
|
|
115
115
|
|
|
116
|
-
1. The dispatch layer persists `{error, buildLog, testLog}` to `state.phases["
|
|
117
|
-
2. Multi-agent increments `state.phases["
|
|
116
|
+
1. The dispatch layer persists `{error, buildLog, testLog}` to `state.phases["2"].errors[]`.
|
|
117
|
+
2. Multi-agent increments `state.phases["2"].retryCount`.
|
|
118
118
|
3. Retry re-invokes the plugin skill (it is idempotent on an existing component; it reconciles rather than duplicating). There is no bundled `phase-<N>` resume anymore.
|
|
119
119
|
4. Hard kill at `retryCount === 3` → surface the errors to the user, halt Phase 2. Do not loop indefinitely.
|
|
120
120
|
|
|
@@ -125,4 +125,4 @@ Component dispatch is **no longer byte-identical across CLIs** and that is by de
|
|
|
125
125
|
- **Claude Code**: dispatches to the enabled `ai-<platform>-toolkit` marketplace plugin via the Skill tool (this doc).
|
|
126
126
|
- **Copilot CLI**: has no plugin loader; it continues to use its standalone `~/.copilot/skills/figma-*` skill copies (frozen fallback). Copilot's resolution + progress lines follow those local skills.
|
|
127
127
|
|
|
128
|
-
What still MUST match across CLIs: the `taskType === "component"` classification, the `state.phases["
|
|
128
|
+
What still MUST match across CLIs: the `taskType === "component"` classification, the `state.phases["2"].subphases[]` shape the dispatch layer writes, and the Short-run elision semantics. Figma-skill *inventory* parity is no longer enforced. `smoke-cross-cli-behavior.sh` asserts only the classification + state-shape axis for components.
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Cross-CLI Contract (Claude Code · Copilot CLI · Codex CLI)
|
|
2
2
|
|
|
3
3
|
<!-- toc -->
|
|
4
|
-
- [1. Command Inventory (
|
|
4
|
+
- [1. Command Inventory (58 commands)](#1-command-inventory-58-commands)
|
|
5
5
|
- [2. Canonical Placeholder Vocabulary](#2-canonical-placeholder-vocabulary)
|
|
6
6
|
- [2.6 Intentional structural divergence - thin dispatcher vs inlined orchestrator](#26-intentional-structural-divergence---thin-dispatcher-vs-inlined-orchestrator)
|
|
7
7
|
- [2.7 One command, three different meanings: `/multi-agent:model`](#27-one-command-three-different-meanings-multi-agentmodel)
|
|
@@ -20,7 +20,7 @@
|
|
|
20
20
|
|
|
21
21
|
---
|
|
22
22
|
|
|
23
|
-
## 1. Command Inventory (
|
|
23
|
+
## 1. Command Inventory (58 commands)
|
|
24
24
|
|
|
25
25
|
```
|
|
26
26
|
analysis, analysis-jira, analysis-resolve, autopilot, autopilot-off,
|
|
@@ -31,7 +31,7 @@ language, log, manual-test, model, prune-logs,
|
|
|
31
31
|
prune-prompts, purge, refactor, resume, review,
|
|
32
32
|
review-analysis, review-issue, review-jira, route-off, route-on,
|
|
33
33
|
route-status, routines, save, scan, search,
|
|
34
|
-
setup, stack, status, steer, store-ready, sync, test, test-accessibility,
|
|
34
|
+
security-review, setup, stack, status, steer, store-ready, sync, test, test-accessibility,
|
|
35
35
|
test-dark-mode, test-dynamic-type, test-screenshots, testflight-validation,
|
|
36
36
|
uninstall, update
|
|
37
37
|
```
|
|
@@ -86,7 +86,7 @@ Three rules make that work:
|
|
|
86
86
|
> absent, and a component task on Copilot had nothing to dispatch to. The contract
|
|
87
87
|
> documented a safety net the installer deleted.
|
|
88
88
|
|
|
89
|
-
What still MUST match across CLIs for component tasks: the `taskType === "component"` classification, the `state.phases["
|
|
89
|
+
What still MUST match across CLIs for component tasks: the `taskType === "component"` classification, the `state.phases["2"].subphases[]` shape the dispatch layer writes, and the Short-run elision semantics. Resolution/routing is Claude-plugin vs Copilot-local-copy by design.
|
|
90
90
|
|
|
91
91
|
### 1.2 Store-compliance skills
|
|
92
92
|
|
|
@@ -97,7 +97,7 @@ Two parallel shared/core skills under `pipeline/skills/shared/core/` wrap extern
|
|
|
97
97
|
| `apple-archive-compliance` | `pipeline/skills/shared/core/apple-archive-compliance/` | `ios_app_store_audit` MCP tool (in `@mmerterden/multi-agent-toolkit-mcp` ≥ v3.0.0) | 18 (Apple ITMS + App Store Review Guidelines) |
|
|
98
98
|
| `google-play-compliance` | `pipeline/skills/shared/core/google-play-compliance/` | bundletool + aapt2 + apksigner | 21 (4 categories: Technical / Security / Privacy / Hygiene) |
|
|
99
99
|
|
|
100
|
-
Both skills are wired to 4 consumers: `/multi-agent:test "store-ready"` (primary), Phase
|
|
100
|
+
Both skills are wired to 4 consumers: `/multi-agent:test "store-ready"` (primary), Phase 3 Security Auditor (`pipeline/agents/security-auditor.md`), `/multi-agent:review` + SKILL.md counterpart, `/multi-agent:channels` PR-body auto-augmentation. Contract enforced by `smoke-compliance-skills.sh`.
|
|
101
101
|
|
|
102
102
|
### 1.3 Figma-skill routing from multi-agent (superseded)
|
|
103
103
|
|
|
@@ -222,7 +222,7 @@ Phase 3 runs three reviewers everywhere, but the diversity those three buy is no
|
|
|
222
222
|
same on every host. Copilot CLI gets cross-VENDOR disagreement for free: GPT-5.4 sits
|
|
223
223
|
beside two Claude models. Claude Code and Codex each run a one-vendor panel - three
|
|
224
224
|
Anthropic models on one, three OpenAI models on the other - so the same three-way
|
|
225
|
-
agreement is weaker evidence there, and Phase
|
|
225
|
+
agreement is weaker evidence there, and Phase 3 says so in the triage note on a
|
|
226
226
|
borderline finding.
|
|
227
227
|
|
|
228
228
|
Where the budget goes instead, when vendor diversity is unavailable:
|
|
@@ -0,0 +1,55 @@
|
|
|
1
|
+
# Security audit (Phase 3 Step 2.7)
|
|
2
|
+
|
|
3
|
+
The mechanics of the conditional security audit that runs inside the Phase 3 review window and merges at Step 3.0. The phase doc carries the trigger and the merge; the rest is here.
|
|
4
|
+
|
|
5
|
+
## Trigger (deterministic, no new classifier)
|
|
6
|
+
|
|
7
|
+
Run the audit when any of these holds:
|
|
8
|
+
|
|
9
|
+
- Step 1.75 scored `security_path` on a file in the diff (the same signal that already forces `full` review at Step 1.77).
|
|
10
|
+
- The base branch is a release branch.
|
|
11
|
+
- The run is the standalone `/multi-agent:security-review` command, which dispatches straight to this step.
|
|
12
|
+
|
|
13
|
+
There is no `--audit` flag: the trigger is the diff, not a word the user has to remember. When none of these holds, the audit does not run, `$SECURITY_AUDIT_JSON` stays empty, and the Step 3.0 merge is a no-op.
|
|
14
|
+
|
|
15
|
+
## Threat model first
|
|
16
|
+
|
|
17
|
+
The auditor reads `.pipeline/threat-model.md` if Phase 1 or a prior step wrote it, and produces it (four sections) if absent. Contract: `$HOME/.claude/multi-agent-refs/threat-model.md`. It is keyed to repo+branch and mirrored to `state.threatModel` so a resume reuses it rather than re-deriving it.
|
|
18
|
+
|
|
19
|
+
## Dispatch
|
|
20
|
+
|
|
21
|
+
One `Agent(subagent_type: "security-auditor")` on the same capped diff the reviewers saw, plus the threat model and the resolved `${CRITERIA}` block:
|
|
22
|
+
|
|
23
|
+
```
|
|
24
|
+
Agent(subagent_type: "security-auditor", prompt: "<threat-model + capped diff + ${CRITERIA}>")
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
Persona contract: `~/.claude/agents/security-auditor.md`.
|
|
28
|
+
|
|
29
|
+
## Output is reviewer-shaped, and that is the point
|
|
30
|
+
|
|
31
|
+
The auditor returns one object conforming to `$HOME/.claude/schemas/reviewer-output.schema.json`, whose `findings[]` each carry the `security` envelope of `security-finding.schema.json`: OWASP category, CWE, CVSS vector + band, evidence, counterevidence, confidence, remediation, and the before/after fix. Severity is the reviewer enum, derived from the CVSS band:
|
|
32
|
+
|
|
33
|
+
| CVSS band | severity |
|
|
34
|
+
| --------- | -------- |
|
|
35
|
+
| critical / high | blocking |
|
|
36
|
+
| medium | important |
|
|
37
|
+
| low / none | suggestion |
|
|
38
|
+
|
|
39
|
+
Compute `security.cvss.baseScore` + `band` with the toolkit `security_cvss_score` tool, not by hand, so the score cannot drift from the vector.
|
|
40
|
+
|
|
41
|
+
## Validate, persist, merge
|
|
42
|
+
|
|
43
|
+
Validate with the same gate protocol as a reviewer - the exit code decides, not the LLM turn:
|
|
44
|
+
|
|
45
|
+
```bash
|
|
46
|
+
printf '%s' "$SECURITY_AUDIT_JSON" | node "$HOME/.claude/scripts/validate-reviewer.mjs" -
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
On validator failure: one self-correction rework, then HALT the phase (identical to the reviewer output-contract gate). Persist the object to `state.reviewIterations[<iteration>].securityAudit` and hold it in `$SECURITY_AUDIT_JSON`.
|
|
50
|
+
|
|
51
|
+
Because the output is reviewer-shaped, the Step 3.0 merge appends its `findings[]` alongside the test-integrity findings. A `blocking` security finding then reaches triage, and a triage-accepted blocker blocks Phase 4 - the "critical security items block the commit" contract is wired here, not merely stated.
|
|
52
|
+
|
|
53
|
+
## Store-compliance cross-reference
|
|
54
|
+
|
|
55
|
+
On store-relevant diffs the auditor also cites the Apple ITMS / Google Play catalog rule on the finding via `ruleId` + `criteriaSource` (`apple-archive-compliance` / `google-play-compliance`). This is annotation on the same finding, not a second pass; the device-level `store-ready` Gate under `/multi-agent:test` stays separate.
|
|
@@ -58,7 +58,7 @@ The generic "zero-interaction" contract covers Phases 0-4 only. Phase 5 channels
|
|
|
58
58
|
- **30-minute timeout** - if user does not respond, session ends cleanly:
|
|
59
59
|
- External delivery aborted (no silent apply - prevents accidental Jira comments / Confluence pages).
|
|
60
60
|
- Internal capture (`agent-log.md`, telemetry, knowledge base) STILL runs.
|
|
61
|
-
- State persisted as `{phase:
|
|
61
|
+
- State persisted as `{phase: 5, waitingFor: "user-channels-choice", channelsTimeout: true}`.
|
|
62
62
|
- Resume: `/multi-agent:resume <task-id>` re-opens menu with same inputs.
|
|
63
63
|
- Post-hoc `/multi-agent:channels <task>` never times out - user invoked it explicitly.
|
|
64
64
|
|
|
@@ -260,7 +260,7 @@ Each reviewer inherits the `code-reviewer` agent's focus areas (Security, Archit
|
|
|
260
260
|
| Python | `ai-backend-toolkit:api-security-best-practices` | `ai-backend-toolkit:fastapi-pro` | `ai-backend-toolkit:python-patterns` |
|
|
261
261
|
| Node.js | `ai-backend-toolkit:api-security-best-practices` | `ai-backend-toolkit:nodejs-backend-patterns` | `ai-frontend-toolkit:typescript-patterns` |
|
|
262
262
|
| Docker | `ai-backend-toolkit:docker-expert` | `ai-backend-toolkit:docker-expert` | `ai-backend-toolkit:ci-cd-pipelines` |
|
|
263
|
-
| Generic | `security-review` | `ai-backend-toolkit:clean-code` | `ai-backend-toolkit:clean-code` |
|
|
263
|
+
| Generic | `ai-common-toolkit:security-review` | `ai-backend-toolkit:clean-code` | `ai-backend-toolkit:clean-code` |
|
|
264
264
|
|
|
265
265
|
##### 2.1 Previous-round findings (iteration >= 2) and 2.2 scope self-check (every iteration)
|
|
266
266
|
|
|
@@ -298,6 +298,10 @@ Skip only when the diff has no UI change. Record the outcome in
|
|
|
298
298
|
|
|
299
299
|
Step 2 produces N reviewer-output objects (one per dispatched reviewer), each conforming to `$HOME/.claude/schemas/reviewer-output.schema.json`. They are persisted to `state.reviewIterations[<iteration>].reviewers[]` and consumed by Step 3 (Fable triage) - never by Phase 4 directly. The triage step (below) is the producer of the only review artifact Phase 4 reads, conforming to `$HOME/.claude/schemas/triage-output.schema.json`.
|
|
300
300
|
|
|
301
|
+
#### Step 2.7 - Security audit (conditional, produces reviewer-shaped findings)
|
|
302
|
+
|
|
303
|
+
Runs when Step 1.75 scored `security_path`, on a release branch, or when `/multi-agent:security-review` dispatches here. The `security-auditor` returns reviewer `findings[]`, each with a `security` envelope (`security-finding.schema.json`) and severity from the CVSS band, held in `$SECURITY_AUDIT_JSON` for the merge so a `blocking` one reaches triage and blocks Phase 4. Mechanics: `~/.claude/multi-agent-refs/features/security-audit.md`.
|
|
304
|
+
|
|
301
305
|
**Subagent return format** - each reviewer returns JSON conforming to `$HOME/.claude/schemas/reviewer-output.schema.json`:
|
|
302
306
|
|
|
303
307
|
```json
|
|
@@ -371,14 +375,14 @@ ANON=$(jq -n --argjson r "$REVIEWERS_JSON" --arg t "$TASK_ID" --argjson i "$ITER
|
|
|
371
375
|
|
|
372
376
|
`$REVIEWERS_JSON` is `state.reviewIterations[i].reviewers`. Findings come back with `foundBy: "Source A|B|C"` and every identity key removed. Persist the map to `state.reviewIterations[i].anonymizationMap` for Phase 5 per-reviewer telemetry, and **never put the map in a prompt**.
|
|
373
377
|
|
|
374
|
-
Then append the Step 1.76 test-integrity findings, so they are adjudicated rather than never seen:
|
|
378
|
+
Then append the Step 1.76 test-integrity and Step 2.7 security-audit findings, so they are adjudicated rather than never seen:
|
|
375
379
|
|
|
376
380
|
```bash
|
|
377
|
-
MERGED=$(jq -s '.[0] + (.[1].findings // [])' \
|
|
378
|
-
<(printf '%s' "$ANON") <(printf '%s' "${TEST_INTEGRITY_JSON:-{\}}"))
|
|
381
|
+
MERGED=$(jq -s '.[0] + (.[1].findings // []) + (.[2].findings // [])' \
|
|
382
|
+
<(printf '%s' "$ANON") <(printf '%s' "${TEST_INTEGRITY_JSON:-{\}}") <(printf '%s' "${SECURITY_AUDIT_JSON:-{\}}"))
|
|
379
383
|
```
|
|
380
384
|
|
|
381
|
-
Deterministic findings keep `tag: test_integrity` and
|
|
385
|
+
Deterministic findings keep `tag: test_integrity` and no `foundBy`: a reviewer finding may hallucinate, a gate finding is a fact. Security-audit findings carry their `security` envelope and `foundBy: "security-auditor"`; an empty `$SECURITY_AUDIT_JSON` contributes nothing.
|
|
382
386
|
|
|
383
387
|
##### 3.1 Short-circuit: no findings
|
|
384
388
|
|
|
@@ -788,16 +792,6 @@ Results included in Phase 5 report. MCP tools preferred when available - conci
|
|
|
788
792
|
|
|
789
793
|
**Snapshot regression flow (optional):** when the task changes a stable component, capture a screenshot before the change (baseline) and after (current), then call `ios_visual_diff({baseline, current, max_diff_pct: 1.0})`. Threshold can be relaxed for animated / non-deterministic regions - keep `max_diff_pct ≤ 1.0` for static layouts.
|
|
790
794
|
|
|
791
|
-
#### Security Audit (store-readiness)
|
|
792
|
-
|
|
793
|
-
When the task touches authentication, keychain, network, or is scheduled for an imminent release, launch the `security-auditor` subagent to run an OWASP Mobile Top 10 pass plus App Store / Play Store compliance checks:
|
|
794
|
-
|
|
795
|
-
```
|
|
796
|
-
Agent(subagent_type: "security-auditor", prompt: "<diff + context>")
|
|
797
|
-
```
|
|
798
|
-
|
|
799
|
-
Returns severity-tagged findings (Critical / High / Medium). Critical items block Phase 4 just like Phase 3 blockers; High items are logged and surfaced in Phase 5 report. Skipped by default - opt-in for release branches or on explicit `/multi-agent "<task>" --audit` flag.
|
|
800
|
-
|
|
801
795
|
#### Telemetry - token forwarding
|
|
802
796
|
|
|
803
797
|
When the security-auditor or any other Phase 3 sub-agent runs, forward its token totals so Phase 5's Cost Breakdown captures Phase 3:
|
|
@@ -25,7 +25,7 @@ On entry: `phase-tracker.sh sub 5 <N> "<name>" in_progress`. On exit: `completed
|
|
|
25
25
|
Phase 5 is the single exception to the autopilot zero-interaction rule: every mode, Full or Short, attended or autopilot, pauses at the channels multi-select menu. Full contract in `$HOME/.claude/multi-agent-refs/phases/modes.md`:
|
|
26
26
|
|
|
27
27
|
- **30-min timeout** - if user does not respond, session ends cleanly. External delivery is aborted (no silent apply of defaults - prevents accidental Jira comments / Confluence pages). Internal capture (Steps 2 + 3 below) STILL runs so `agent-log.md` + knowledge base are persisted.
|
|
28
|
-
- **Resumable** - state written as `{status: "awaiting_input", phase:
|
|
28
|
+
- **Resumable** - state written as `{status: "awaiting_input", phase: 5, waitingFor: "user-channels-choice", channelsInput: <state-bundle>}`. User can `/multi-agent:resume <task-id>` any time later; channels menu re-opens with same inputs.
|
|
29
29
|
- **Timeout log line:** `Phase 5: channels menu timeout (30 min) - session ended, resume with /multi-agent:resume {taskId}`.
|
|
30
30
|
|
|
31
31
|
---
|
|
@@ -0,0 +1,39 @@
|
|
|
1
|
+
# Threat model contract
|
|
2
|
+
|
|
3
|
+
The run-scoped threat model is the frame every security finding is calibrated against. It is produced by the security-auditor at Phase 3 Step 2.7 (or by `/multi-agent:security-review` running standalone), read by any later step that revisits security, and mirrored to `state.threatModel.path` so a resume re-uses it instead of re-deriving it.
|
|
4
|
+
|
|
5
|
+
## Where it lives
|
|
6
|
+
|
|
7
|
+
`.pipeline/threat-model.md` in the run's worktree, keyed to repo + branch. One per run. If a fresh one already exists for this repo+branch (a prior step or a Phase 1 pass wrote it), read it; do not overwrite. If it is absent, produce it before emitting any finding.
|
|
8
|
+
|
|
9
|
+
Mirror the resolved path and a content hash to `state.threatModel = { path, sha, producedAt, producedBy }` so resume and Phase 5 can find it without re-reading the tree.
|
|
10
|
+
|
|
11
|
+
## The four sections (all required, in order)
|
|
12
|
+
|
|
13
|
+
A threat model with a missing section is not a threat model; the auditor treats a missing section as "produce it," not "skip it."
|
|
14
|
+
|
|
15
|
+
### 1. Attacker
|
|
16
|
+
|
|
17
|
+
The one realistic adversary for THIS change. Name it concretely: an unauthenticated internet caller, an authenticated low-privilege user, a malicious or compromised dependency, a co-located app on the device, a user with physical access. Not "attackers" in the abstract - the specific actor whose capability makes this diff interesting. If the diff has no plausible attacker, say so; that is a valid, short threat model and most findings then cap at `suggestion`.
|
|
18
|
+
|
|
19
|
+
### 2. Trust boundaries
|
|
20
|
+
|
|
21
|
+
Where untrusted data crosses into trusted code within the changed surface: a request body or query parameter, a deep link or universal link, a WebView `postMessage`, a file the app did not write, an environment variable an attacker can set, a response from a third-party service treated as safe. List the boundaries the diff touches, each with the `file:line` where the crossing happens.
|
|
22
|
+
|
|
23
|
+
### 3. Attack surface
|
|
24
|
+
|
|
25
|
+
What this diff actually added or touched: a new endpoint, a new query or ORM call, a new deserialization, a new permission, a new dependency, a new crypto usage, a new storage write. A finding outside this surface is out of scope unless the diff made it reachable - and if it did, say how. This section is what keeps the audit anchored to the change instead of drifting into a whole-repo review.
|
|
26
|
+
|
|
27
|
+
### 4. Severity calibration
|
|
28
|
+
|
|
29
|
+
The assumption each severity rests on, stated so a reader disputes the assumption rather than the number. "Critical assumes this route is unauthenticated in production; if it is admin-only the same finding is medium." "High assumes the secret reaches a log that ships off-device." Every `blocking` finding must trace to an assumption named here; a blocker resting on an unstated assumption is miscalibrated.
|
|
30
|
+
|
|
31
|
+
## Shape
|
|
32
|
+
|
|
33
|
+
Plain Markdown, four `##` sections with those names, human-readable. It is evidence for a person and context for the auditor, not a machine artifact - no schema. Keep it short: a page, not a report. Cite `file:line` where a boundary or surface item has one.
|
|
34
|
+
|
|
35
|
+
## What it is not
|
|
36
|
+
|
|
37
|
+
- Not a whole-repo model. It is scoped to the diff under review.
|
|
38
|
+
- Not a dynamic test plan. This is a static, read-only posture: no live target, no payloads, no exploitation. `fixVerification` on a finding names the empirical check a human or a later dynamic pass would run; the threat model does not run it.
|
|
39
|
+
- Not embedded in the analysis document. The security-review command must run standalone, so the model lives in its own file and is produced on demand.
|