@rune-kit/rune 2.10.0 → 2.12.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -21
- package/README.md +65 -6
- package/commands/rune.md +168 -168
- package/compiler/__tests__/detect-invariants.test.js +136 -0
- package/compiler/__tests__/doctor-mesh.test.js +229 -0
- package/compiler/__tests__/hook-dispatch.test.js +91 -0
- package/compiler/__tests__/hooks-antigravity.test.js +118 -0
- package/compiler/__tests__/hooks-cursor.test.js +139 -0
- package/compiler/__tests__/hooks-install.test.js +305 -0
- package/compiler/__tests__/hooks-merge.test.js +204 -0
- package/compiler/__tests__/hooks-tiers.test.js +519 -0
- package/compiler/__tests__/hooks-windsurf.test.js +115 -0
- package/compiler/__tests__/inject-claude-md.test.js +152 -0
- package/compiler/__tests__/load-invariants.test.js +408 -0
- package/compiler/__tests__/onboard-invariants.test.js +240 -0
- package/compiler/adapters/hooks/antigravity.js +140 -0
- package/compiler/adapters/hooks/claude.js +166 -0
- package/compiler/adapters/hooks/cursor.js +191 -0
- package/compiler/adapters/hooks/index.js +82 -0
- package/compiler/adapters/hooks/tier-emitter.js +182 -0
- package/compiler/adapters/hooks/windsurf.js +202 -0
- package/compiler/bin/rune.js +196 -6
- package/compiler/commands/hook-dispatch.js +87 -0
- package/compiler/commands/hooks/install.js +120 -0
- package/compiler/commands/hooks/merge.js +211 -0
- package/compiler/commands/hooks/presets.js +116 -0
- package/compiler/commands/hooks/status.js +112 -0
- package/compiler/commands/hooks/tiers.js +221 -0
- package/compiler/commands/hooks/uninstall.js +94 -0
- package/compiler/doctor.js +236 -0
- package/contexts/dev.md +34 -34
- package/contexts/research.md +43 -43
- package/contexts/review.md +55 -55
- package/extensions/ai-ml/PACK.md +88 -88
- package/extensions/ai-ml/skills/ai-agents.md +172 -172
- package/extensions/ai-ml/skills/code-sandbox.md +187 -187
- package/extensions/ai-ml/skills/deep-research.md +146 -146
- package/extensions/ai-ml/skills/embedding-search.md +66 -66
- package/extensions/ai-ml/skills/fine-tuning-guide.md +74 -74
- package/extensions/ai-ml/skills/llm-architect.md +125 -125
- package/extensions/ai-ml/skills/llm-integration.md +64 -64
- package/extensions/ai-ml/skills/prompt-patterns.md +72 -72
- package/extensions/ai-ml/skills/rag-patterns.md +66 -66
- package/extensions/ai-ml/skills/web-extraction.md +114 -114
- package/extensions/analytics/PACK.md +92 -92
- package/extensions/analytics/skills/ab-testing.md +72 -72
- package/extensions/analytics/skills/dashboard-patterns.md +83 -83
- package/extensions/analytics/skills/data-validation.md +68 -68
- package/extensions/analytics/skills/funnel-analysis.md +81 -81
- package/extensions/analytics/skills/sql-patterns.md +57 -57
- package/extensions/analytics/skills/statistical-analysis.md +79 -79
- package/extensions/analytics/skills/tracking-setup.md +71 -71
- package/extensions/backend/PACK.md +104 -104
- package/extensions/backend/skills/api-patterns.md +84 -84
- package/extensions/backend/skills/async-pipeline.md +193 -193
- package/extensions/backend/skills/auth-patterns.md +97 -97
- package/extensions/backend/skills/background-jobs.md +133 -133
- package/extensions/backend/skills/caching-patterns.md +108 -108
- package/extensions/backend/skills/cli-generation.md +133 -133
- package/extensions/backend/skills/database-patterns.md +87 -87
- package/extensions/backend/skills/middleware-patterns.md +104 -104
- package/extensions/chrome-ext/PACK.md +93 -93
- package/extensions/chrome-ext/skills/cws-preflight.md +143 -143
- package/extensions/chrome-ext/skills/cws-publish.md +104 -104
- package/extensions/chrome-ext/skills/ext-ai-integration.md +251 -251
- package/extensions/chrome-ext/skills/ext-messaging.md +139 -139
- package/extensions/chrome-ext/skills/ext-storage.md +133 -133
- package/extensions/chrome-ext/skills/mv3-scaffold.md +164 -164
- package/extensions/content/PACK.md +96 -96
- package/extensions/content/skills/blog-patterns.md +88 -88
- package/extensions/content/skills/cms-integration.md +131 -131
- package/extensions/content/skills/content-scoring.md +107 -107
- package/extensions/content/skills/i18n.md +83 -83
- package/extensions/content/skills/mdx-authoring.md +137 -137
- package/extensions/content/skills/reference.md +1014 -1014
- package/extensions/content/skills/seo-patterns.md +67 -67
- package/extensions/content/skills/video-repurpose.md +153 -153
- package/extensions/devops/PACK.md +101 -101
- package/extensions/devops/skills/chaos-testing.md +67 -67
- package/extensions/devops/skills/ci-cd.md +75 -75
- package/extensions/devops/skills/docker.md +58 -58
- package/extensions/devops/skills/edge-serverless.md +163 -163
- package/extensions/devops/skills/infra-as-code.md +158 -158
- package/extensions/devops/skills/kubernetes.md +110 -110
- package/extensions/devops/skills/monitoring.md +57 -57
- package/extensions/devops/skills/server-setup.md +64 -64
- package/extensions/devops/skills/ssl-domain.md +42 -42
- package/extensions/ecommerce/PACK.md +116 -116
- package/extensions/ecommerce/skills/cart-system.md +79 -79
- package/extensions/ecommerce/skills/inventory-mgmt.md +102 -102
- package/extensions/ecommerce/skills/order-management.md +126 -126
- package/extensions/ecommerce/skills/payment-integration.md +472 -472
- package/extensions/ecommerce/skills/shopify-dev.md +69 -69
- package/extensions/ecommerce/skills/subscription-billing.md +93 -93
- package/extensions/ecommerce/skills/tax-compliance.md +117 -117
- package/extensions/gamedev/PACK.md +142 -142
- package/extensions/gamedev/skills/asset-pipeline.md +74 -74
- package/extensions/gamedev/skills/audio-system.md +129 -129
- package/extensions/gamedev/skills/camera-system.md +87 -87
- package/extensions/gamedev/skills/ecs.md +98 -98
- package/extensions/gamedev/skills/game-loops.md +72 -72
- package/extensions/gamedev/skills/input-system.md +199 -199
- package/extensions/gamedev/skills/multiplayer.md +180 -180
- package/extensions/gamedev/skills/particles.md +105 -105
- package/extensions/gamedev/skills/physics-engine.md +89 -89
- package/extensions/gamedev/skills/scene-management.md +146 -146
- package/extensions/gamedev/skills/threejs-patterns.md +90 -90
- package/extensions/gamedev/skills/webgl.md +71 -71
- package/extensions/mobile/PACK.md +106 -106
- package/extensions/mobile/skills/app-store-connect.md +152 -152
- package/extensions/mobile/skills/app-store-prep.md +66 -66
- package/extensions/mobile/skills/deep-linking.md +109 -109
- package/extensions/mobile/skills/flutter.md +60 -60
- package/extensions/mobile/skills/ios-build-pipeline.md +142 -142
- package/extensions/mobile/skills/native-bridge.md +66 -66
- package/extensions/mobile/skills/ota-updates.md +97 -97
- package/extensions/mobile/skills/push-notifications.md +111 -111
- package/extensions/mobile/skills/react-native.md +82 -82
- package/extensions/saas/PACK.md +116 -116
- package/extensions/saas/skills/billing-integration.md +200 -200
- package/extensions/saas/skills/feature-flags.md +130 -130
- package/extensions/saas/skills/multi-tenant.md +103 -103
- package/extensions/saas/skills/onboarding-flow.md +139 -139
- package/extensions/saas/skills/subscription-flow.md +95 -95
- package/extensions/saas/skills/team-management.md +144 -144
- package/extensions/security/PACK.md +99 -99
- package/extensions/security/skills/api-security.md +140 -140
- package/extensions/security/skills/compliance.md +68 -68
- package/extensions/security/skills/owasp-audit.md +64 -64
- package/extensions/security/skills/pentest-patterns.md +77 -77
- package/extensions/security/skills/secret-mgmt.md +65 -65
- package/extensions/security/skills/supply-chain.md +65 -65
- package/extensions/trading/PACK.md +80 -80
- package/extensions/trading/skills/chart-components.md +55 -55
- package/extensions/trading/skills/experiment-loop.md +125 -125
- package/extensions/trading/skills/fintech-patterns.md +47 -47
- package/extensions/trading/skills/indicator-library.md +58 -58
- package/extensions/trading/skills/quant-analysis.md +111 -111
- package/extensions/trading/skills/realtime-data.md +58 -58
- package/extensions/trading/skills/trade-logic.md +104 -104
- package/extensions/ui/PACK.md +130 -130
- package/extensions/ui/skills/a11y-audit.md +91 -91
- package/extensions/ui/skills/animation-patterns.md +127 -127
- package/extensions/ui/skills/component-patterns.md +100 -100
- package/extensions/ui/skills/design-decision.md +108 -108
- package/extensions/ui/skills/design-system.md +68 -68
- package/extensions/ui/skills/landing-patterns.md +155 -155
- package/extensions/ui/skills/palette-picker.md +173 -173
- package/extensions/ui/skills/react-health.md +90 -90
- package/extensions/ui/skills/type-system.md +125 -125
- package/extensions/ui/skills/web-vitals.md +153 -153
- package/extensions/zalo/PACK.md +145 -145
- package/extensions/zalo/skills/zalo-oa-mcp.md +317 -317
- package/extensions/zalo/skills/zalo-oa-messaging.md +429 -429
- package/extensions/zalo/skills/zalo-oa-setup.md +236 -236
- package/extensions/zalo/skills/zalo-oa-webhook.md +189 -189
- package/extensions/zalo/skills/zalo-personal-messaging.md +194 -194
- package/extensions/zalo/skills/zalo-personal-setup.md +153 -153
- package/extensions/zalo/skills/zalo-rate-guard.md +219 -219
- package/hooks/auto-format/index.cjs +48 -48
- package/hooks/hooks.json +111 -111
- package/hooks/post-session-reflect/index.cjs +189 -189
- package/hooks/pre-compact/index.cjs +95 -95
- package/hooks/run-hook.cmd +1 -1
- package/hooks/secrets-scan/index.cjs +100 -100
- package/hooks/session-start/index.cjs +71 -71
- package/hooks/typecheck/index.cjs +65 -65
- package/package.json +63 -63
- package/references/ui-pro-max-data/LICENSE-UI-PRO-MAX +21 -21
- package/references/ui-pro-max-data/charts.csv +26 -26
- package/references/ui-pro-max-data/colors.csv +161 -161
- package/references/ui-pro-max-data/styles.csv +68 -68
- package/references/ui-pro-max-data/typography.csv +74 -74
- package/references/ui-pro-max-data/ui-reasoning.csv +162 -162
- package/references/ui-pro-max-data/ux-guidelines.csv +99 -99
- package/skills/adversary/SKILL.md +283 -283
- package/skills/asset-creator/SKILL.md +157 -157
- package/skills/audit/SKILL.md +147 -2
- package/skills/autopsy/SKILL.md +335 -335
- package/skills/ba/SKILL.md +85 -1
- package/skills/brainstorm/SKILL.md +380 -342
- package/skills/browser-pilot/SKILL.md +169 -168
- package/skills/constraint-check/SKILL.md +165 -165
- package/skills/context-engine/SKILL.md +408 -404
- package/skills/cook/SKILL.md +917 -863
- package/skills/db/SKILL.md +273 -273
- package/skills/debug/SKILL.md +465 -465
- package/skills/dependency-doctor/SKILL.md +265 -235
- package/skills/deploy/SKILL.md +274 -231
- package/skills/design/DESIGN-REFERENCE.md +365 -365
- package/skills/design/SKILL.md +590 -589
- package/skills/doc-processor/SKILL.md +254 -254
- package/skills/docs/SKILL.md +374 -374
- package/skills/docs-seeker/SKILL.md +178 -177
- package/skills/fix/SKILL.md +332 -330
- package/skills/git/SKILL.md +339 -339
- package/skills/hallucination-guard/SKILL.md +220 -219
- package/skills/incident/SKILL.md +254 -253
- package/skills/integrity-check/SKILL.md +169 -169
- package/skills/journal/SKILL.md +241 -240
- package/skills/launch/SKILL.md +344 -344
- package/skills/logic-guardian/SKILL.md +269 -251
- package/skills/marketing/SKILL.md +351 -289
- package/skills/mcp-builder/SKILL.md +425 -425
- package/skills/neural-memory/SKILL.md +359 -362
- package/skills/onboard/SKILL.md +432 -403
- package/skills/onboard/references/invariants-template.md +76 -0
- package/skills/onboard/scripts/detect-invariants.js +439 -0
- package/skills/onboard/scripts/inject-claude-md.js +150 -0
- package/skills/onboard/scripts/onboard-invariants.js +194 -0
- package/skills/perf/SKILL.md +347 -346
- package/skills/plan/SKILL.md +435 -428
- package/skills/preflight/SKILL.md +415 -415
- package/skills/problem-solver/SKILL.md +380 -284
- package/skills/rescue/SKILL.md +474 -474
- package/skills/research/SKILL.md +4 -0
- package/skills/retro/SKILL.md +3 -1
- package/skills/review/SKILL.md +614 -588
- package/skills/review-intake/SKILL.md +249 -249
- package/skills/safeguard/SKILL.md +200 -200
- package/skills/sast/SKILL.md +190 -190
- package/skills/scaffold/SKILL.md +328 -287
- package/skills/scope-guard/SKILL.md +183 -180
- package/skills/scout/SKILL.md +269 -263
- package/skills/sentinel/SKILL.md +384 -381
- package/skills/sentinel-env/SKILL.md +254 -254
- package/skills/sequential-thinking/SKILL.md +234 -234
- package/skills/session-bridge/SKILL.md +595 -543
- package/skills/session-bridge/scripts/load-invariants.js +397 -0
- package/skills/skill-forge/SKILL.md +581 -581
- package/skills/skill-router/SKILL.md +3 -0
- package/skills/slides/SKILL.md +19 -0
- package/skills/surgeon/SKILL.md +215 -215
- package/skills/team/SKILL.md +557 -537
- package/skills/test/SKILL.md +620 -614
- package/skills/trend-scout/SKILL.md +145 -145
- package/skills/verification/SKILL.md +334 -326
- package/skills/video-creator/SKILL.md +201 -201
- package/skills/watchdog/SKILL.md +168 -168
- package/skills/worktree/SKILL.md +140 -140
|
@@ -1,283 +1,283 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: adversary
|
|
3
|
-
description: Pre-implementation red-team analysis. Challenges plans before code is written — finds edge cases, security holes, scalability bottlenecks, error propagation risks, and integration conflicts. Catches flaws at plan time (10x cheaper than post-implementation).
|
|
4
|
-
metadata:
|
|
5
|
-
author: runedev
|
|
6
|
-
version: "0.1.0"
|
|
7
|
-
layer: L2
|
|
8
|
-
model: opus
|
|
9
|
-
group: quality
|
|
10
|
-
tools: "Read, Glob, Grep"
|
|
11
|
-
---
|
|
12
|
-
|
|
13
|
-
# adversary
|
|
14
|
-
|
|
15
|
-
## Purpose
|
|
16
|
-
|
|
17
|
-
Pre-implementation adversarial analysis. After a plan is approved but BEFORE code is written, adversary stress-tests the plan across 5 dimensions: edge cases, security, scalability, error propagation, and integration risk. It does NOT fix or redesign — it reports weaknesses so the plan can be hardened before implementation begins.
|
|
18
|
-
|
|
19
|
-
This fills the only gap in the plan-to-ship pipeline: all other quality skills (review, preflight, sentinel) operate AFTER code exists. Catching a flaw in a plan costs minutes; catching it in implementation costs hours.
|
|
20
|
-
|
|
21
|
-
<HARD-GATE>
|
|
22
|
-
adversary MUST NOT approve a plan without at least one specific challenge per dimension analyzed.
|
|
23
|
-
A report that says "plan looks solid" without concrete attack vectors is NOT a red-team analysis.
|
|
24
|
-
Every finding MUST reference the specific plan section, file, or assumption it challenges.
|
|
25
|
-
</HARD-GATE>
|
|
26
|
-
|
|
27
|
-
## Triggers
|
|
28
|
-
|
|
29
|
-
- Called by `cook` Phase 2.5 — after plan approved, before Phase 3 (TEST)
|
|
30
|
-
- `/rune adversary` — manual red-team analysis of any plan or design document
|
|
31
|
-
- Auto-trigger: when plan files are created in `.rune/` or `docs/plans/`
|
|
32
|
-
|
|
33
|
-
## Calls (outbound)
|
|
34
|
-
|
|
35
|
-
- `sentinel` (L2): deep security scan when adversary identifies auth/crypto/payment attack vectors in the plan
|
|
36
|
-
- `perf` (L2): scalability analysis when adversary identifies potential bottleneck patterns
|
|
37
|
-
- `scout` (L2): find existing code that might conflict with planned changes
|
|
38
|
-
- `docs-seeker` (L3): verify framework/API assumptions in the plan are correct and current
|
|
39
|
-
- `hallucination-guard` (L3): verify that APIs, packages, or patterns referenced in the plan actually exist
|
|
40
|
-
|
|
41
|
-
## Called By (inbound)
|
|
42
|
-
|
|
43
|
-
- `cook` (L1): Phase 2.5 — after plan approval, before TDD
|
|
44
|
-
- `plan` (L2): optional post-step for critical features
|
|
45
|
-
- `team` (L1): when decomposing large tasks, adversary validates the decomposition
|
|
46
|
-
- User: `/rune adversary` direct invocation
|
|
47
|
-
|
|
48
|
-
## Cross-Hub Connections
|
|
49
|
-
|
|
50
|
-
- `adversary` ← `cook` — plan produced → adversary challenges it → hardened plan feeds Phase 3
|
|
51
|
-
- `adversary` → `sentinel` — security attack vector identified → sentinel validates depth
|
|
52
|
-
- `adversary` → `perf` — scalability concern raised → perf quantifies the bottleneck
|
|
53
|
-
- `adversary` → `scout` — integration risk flagged → scout finds affected code
|
|
54
|
-
- `adversary` → `plan` — CRITICAL findings → plan revises before implementation
|
|
55
|
-
|
|
56
|
-
## Execution
|
|
57
|
-
|
|
58
|
-
### Step 0: Load Context
|
|
59
|
-
|
|
60
|
-
1. Read the plan document (from `.rune/features/<name>/plan.md`, phase file, or user-specified path)
|
|
61
|
-
2. Read the requirements document if it exists (`.rune/features/<name>/requirements.md` from BA)
|
|
62
|
-
3. Use `scout` to identify existing code files that the plan will touch or depend on
|
|
63
|
-
4. Identify the plan's core assumptions — what MUST be true for this plan to work?
|
|
64
|
-
|
|
65
|
-
### Step 1: Edge Case Analysis
|
|
66
|
-
|
|
67
|
-
Challenge the plan's handling of boundary conditions.
|
|
68
|
-
|
|
69
|
-
For each input/output/state transition in the plan, ask:
|
|
70
|
-
- **Empty/zero**: What happens with no data, zero items, empty strings, null users?
|
|
71
|
-
- **Overflow**: What happens at MAX — 10K items, 1MB payload, 1000 concurrent users?
|
|
72
|
-
- **Race conditions**: What if two operations happen simultaneously? Can state become inconsistent?
|
|
73
|
-
- **Partial failure**: What if step 3 of 5 fails? Is there rollback? Or orphaned state?
|
|
74
|
-
- **Invalid combinations**: What input combinations are technically possible but semantically nonsensical?
|
|
75
|
-
|
|
76
|
-
```
|
|
77
|
-
EDGE_CASE_TEMPLATE:
|
|
78
|
-
- Scenario: [specific edge case]
|
|
79
|
-
- Plan assumption: [what the plan assumes]
|
|
80
|
-
- Attack: [how this breaks]
|
|
81
|
-
- Impact: [what fails — data loss, crash, wrong result, security breach]
|
|
82
|
-
- Remediation: [1-sentence fix suggestion]
|
|
83
|
-
```
|
|
84
|
-
|
|
85
|
-
### Step 2: Security Attack Vectors
|
|
86
|
-
|
|
87
|
-
Analyze the plan for security weaknesses BEFORE any code exists.
|
|
88
|
-
|
|
89
|
-
- **Input trust boundaries**: Where does the plan accept external input? Is validation specified?
|
|
90
|
-
- **Authentication gaps**: Does the plan assume auth exists? Are there unprotected routes or actions?
|
|
91
|
-
- **Data exposure**: Could the planned API responses leak sensitive fields? Are there over-fetching risks?
|
|
92
|
-
- **Privilege escalation**: Can a normal user reach admin functionality through the planned flow?
|
|
93
|
-
- **Injection surfaces**: Does the plan involve dynamic queries, template rendering, or shell commands?
|
|
94
|
-
- **Dependency risk**: Does the plan introduce new dependencies? Are they well-maintained and trusted?
|
|
95
|
-
|
|
96
|
-
If any auth, crypto, or payment logic is in the plan: MUST call `rune:sentinel` for deep analysis.
|
|
97
|
-
|
|
98
|
-
```
|
|
99
|
-
SECURITY_TEMPLATE:
|
|
100
|
-
- Vector: [attack type — OWASP category if applicable]
|
|
101
|
-
- Entry point: [which part of the plan is vulnerable]
|
|
102
|
-
- Exploit scenario: [how an attacker would use this]
|
|
103
|
-
- Severity: CRITICAL | HIGH | MEDIUM
|
|
104
|
-
- Remediation: [what the plan should specify to prevent this]
|
|
105
|
-
```
|
|
106
|
-
|
|
107
|
-
### Step 3: Scalability Stress Test
|
|
108
|
-
|
|
109
|
-
Project the plan forward — what happens at 10x and 100x scale?
|
|
110
|
-
|
|
111
|
-
- **N+1 queries**: Does the plan describe data fetching that will create N+1 database calls?
|
|
112
|
-
- **Missing pagination**: Does the plan handle lists without specifying limits?
|
|
113
|
-
- **Synchronous bottlenecks**: Are there blocking operations in the hot path?
|
|
114
|
-
- **Cache invalidation**: If caching is planned, what happens when data changes? Stale reads?
|
|
115
|
-
- **State growth**: Does the plan accumulate state (in-memory, database, file system) without cleanup?
|
|
116
|
-
- **External service limits**: Does the plan account for rate limits on third-party APIs?
|
|
117
|
-
|
|
118
|
-
If bottleneck patterns detected: call `rune:perf` for quantitative analysis.
|
|
119
|
-
|
|
120
|
-
```
|
|
121
|
-
SCALE_TEMPLATE:
|
|
122
|
-
- Bottleneck: [what breaks at scale]
|
|
123
|
-
- Current plan: [what the plan specifies]
|
|
124
|
-
- At 10x: [what happens]
|
|
125
|
-
- At 100x: [what happens]
|
|
126
|
-
- Remediation: [what to add to the plan]
|
|
127
|
-
```
|
|
128
|
-
|
|
129
|
-
### Step 4: Error Propagation Analysis
|
|
130
|
-
|
|
131
|
-
Trace failure paths through the planned system.
|
|
132
|
-
|
|
133
|
-
- **Cascade failures**: If Service A fails, does the plan specify what happens to B, C, D?
|
|
134
|
-
- **Retry storms**: Does the plan include retries? Could retries amplify the failure?
|
|
135
|
-
- **Silent failures**: Are there operations that could fail without anyone knowing?
|
|
136
|
-
- **Inconsistent state**: If a multi-step operation fails midway, is the data left in a valid state?
|
|
137
|
-
- **User experience**: When things fail, what does the user see? Is there a degraded mode?
|
|
138
|
-
- **Recovery path**: After failure + fix, can the system resume? Or does it require manual intervention?
|
|
139
|
-
|
|
140
|
-
```
|
|
141
|
-
ERROR_TEMPLATE:
|
|
142
|
-
- Failure point: [where in the plan]
|
|
143
|
-
- Propagation: [what else breaks]
|
|
144
|
-
- User impact: [what the user experiences]
|
|
145
|
-
- Recovery: [how to get back to good state]
|
|
146
|
-
- Missing in plan: [what the plan should specify]
|
|
147
|
-
```
|
|
148
|
-
|
|
149
|
-
### Step 5: Integration Risk Assessment
|
|
150
|
-
|
|
151
|
-
Check for conflicts with existing code and architecture.
|
|
152
|
-
|
|
153
|
-
- Use `rune:scout` to find all files the plan will modify or depend on
|
|
154
|
-
- **Breaking changes**: Does the plan modify shared interfaces, types, or APIs that other code depends on?
|
|
155
|
-
- **Migration gaps**: Does the plan require database migrations? Are they reversible?
|
|
156
|
-
- **Configuration drift**: Does the plan add new environment variables, feature flags, or config files?
|
|
157
|
-
- **Test invalidation**: Will existing tests break from the planned changes?
|
|
158
|
-
- **Deployment ordering**: Does the plan require specific deployment sequence? (DB first, then API, then frontend?)
|
|
159
|
-
|
|
160
|
-
```
|
|
161
|
-
INTEGRATION_TEMPLATE:
|
|
162
|
-
- Conflict: [what clashes]
|
|
163
|
-
- Existing code: [file:line that would be affected]
|
|
164
|
-
- Plan assumption: [what the plan assumes about existing code]
|
|
165
|
-
- Reality: [what the existing code actually does]
|
|
166
|
-
- Remediation: [how to resolve the conflict]
|
|
167
|
-
```
|
|
168
|
-
|
|
169
|
-
### Step 6: Verdict and Report
|
|
170
|
-
|
|
171
|
-
Synthesize all findings into an actionable report.
|
|
172
|
-
|
|
173
|
-
**Before reporting, apply rigor filter:**
|
|
174
|
-
- Only report findings you can justify with specific references to the plan or codebase
|
|
175
|
-
- Do NOT report theoretical concerns that require 3+ unlikely conditions to trigger
|
|
176
|
-
- Prioritize findings that would cause the MOST wasted implementation time if discovered later
|
|
177
|
-
- Consolidate related findings — "auth is underspecified" not 5 separate auth findings
|
|
178
|
-
|
|
179
|
-
**Verdict logic:**
|
|
180
|
-
- Any CRITICAL finding → **REVISE** (plan must be updated before Phase 3)
|
|
181
|
-
- 3+ HIGH findings → **REVISE**
|
|
182
|
-
- HIGH findings with clear remediations → **HARDEN** (add remediations to plan, then proceed)
|
|
183
|
-
- Only MEDIUM/LOW findings → **PROCEED** (note findings for implementation awareness)
|
|
184
|
-
|
|
185
|
-
After reporting:
|
|
186
|
-
- If verdict is REVISE: return to `plan` with findings attached as constraints
|
|
187
|
-
- If verdict is HARDEN: present remediations to user for plan update
|
|
188
|
-
- If verdict is PROCEED: pass findings to cook Phase 3 as implementation notes
|
|
189
|
-
|
|
190
|
-
## Output Format
|
|
191
|
-
|
|
192
|
-
```
|
|
193
|
-
## Adversary Report: [feature/plan name]
|
|
194
|
-
- **Plan analyzed**: [path to plan file]
|
|
195
|
-
- **Dimensions checked**: [which of the 5 were relevant]
|
|
196
|
-
- **Findings**: [count by severity]
|
|
197
|
-
- **Verdict**: REVISE | HARDEN | PROCEED
|
|
198
|
-
|
|
199
|
-
### CRITICAL
|
|
200
|
-
- [ADV-001] [dimension]: [description with plan reference]
|
|
201
|
-
- Attack: [how this breaks]
|
|
202
|
-
- Remediation: [specific fix]
|
|
203
|
-
|
|
204
|
-
### HIGH
|
|
205
|
-
- [ADV-002] [dimension]: [description with plan reference]
|
|
206
|
-
- Attack: [how this breaks]
|
|
207
|
-
- Remediation: [specific fix]
|
|
208
|
-
|
|
209
|
-
### MEDIUM
|
|
210
|
-
- [ADV-003] [dimension]: [description]
|
|
211
|
-
|
|
212
|
-
### Strength Notes
|
|
213
|
-
- [what the plan does well — adversary is harsh but fair]
|
|
214
|
-
|
|
215
|
-
### Verdict
|
|
216
|
-
[Summary: why REVISE/HARDEN/PROCEED, what to do next]
|
|
217
|
-
```
|
|
218
|
-
|
|
219
|
-
## Workflow Modes
|
|
220
|
-
|
|
221
|
-
### Full Red-Team (default)
|
|
222
|
-
All 5 dimensions analyzed. Used for new features, architectural changes, security-sensitive plans.
|
|
223
|
-
|
|
224
|
-
### Quick Challenge (for smaller plans)
|
|
225
|
-
Skip Steps 3-4 (scalability, error propagation). Focus on edge cases, security, and integration.
|
|
226
|
-
Trigger: plan modifies < 3 files AND no auth/payment/data logic.
|
|
227
|
-
|
|
228
|
-
### Security-Focused
|
|
229
|
-
Steps 2 and 5 only (security + integration). Used when `sentinel` requests adversarial pre-analysis.
|
|
230
|
-
Trigger: plan involves auth, crypto, payment, or user data handling.
|
|
231
|
-
|
|
232
|
-
## Constraints
|
|
233
|
-
|
|
234
|
-
1. MUST challenge every plan — no rubber-stamping. At minimum, one finding per analyzed dimension
|
|
235
|
-
2. MUST NOT modify the plan or write code — adversary is read-only analysis
|
|
236
|
-
3. MUST reference specific plan sections or existing code for every finding
|
|
237
|
-
4. MUST escalate to sentinel when auth/crypto/payment attack vectors are identified
|
|
238
|
-
5. MUST use concrete attack scenarios, not vague warnings ("could be a problem" is NOT a finding)
|
|
239
|
-
6. MUST NOT block on MEDIUM/LOW findings — only CRITICAL and HIGH trigger REVISE verdict
|
|
240
|
-
7. MUST include Strength Notes — adversary finds weaknesses AND acknowledges what's well-designed
|
|
241
|
-
|
|
242
|
-
## Mesh Gates
|
|
243
|
-
|
|
244
|
-
| Gate | Requires | If Missing |
|
|
245
|
-
|------|----------|------------|
|
|
246
|
-
| Plan Gate | A plan document exists (from plan skill or user-provided) | Cannot run — ask for plan first |
|
|
247
|
-
| Codebase Gate | Access to existing codebase (for integration checks) | Skip Step 5, note in report |
|
|
248
|
-
|
|
249
|
-
## Sharp Edges
|
|
250
|
-
|
|
251
|
-
| Failure Mode | Severity | Mitigation |
|
|
252
|
-
|---|---|---|
|
|
253
|
-
| Over-challenging — nitpicking every line of the plan | HIGH | Rigor filter: only findings you can justify with specific references. Skip theoretical 3+ condition chains |
|
|
254
|
-
| False security alarms — flagging secure patterns as vulnerable | HIGH | Call sentinel for validation before reporting security findings as CRITICAL |
|
|
255
|
-
| Analysis paralysis — too many findings block all progress | MEDIUM | Max 3 CRITICAL + 5 HIGH. If more found, consolidate or prioritize top impact |
|
|
256
|
-
| Missing context — challenging plan without understanding existing codebase | HIGH | Step 0 MUST load existing code context via scout before challenging |
|
|
257
|
-
| Scope creep — reviewing existing code quality instead of plan quality | MEDIUM | Adversary reviews THE PLAN, not the codebase. Existing code is context only |
|
|
258
|
-
| Redundancy with review/preflight — duplicating post-implementation checks | MEDIUM | Adversary operates PRE-implementation only. Never run adversary on existing code |
|
|
259
|
-
|
|
260
|
-
## Done When
|
|
261
|
-
|
|
262
|
-
- All relevant dimensions analyzed (minimum: edge cases + security + integration)
|
|
263
|
-
- Every finding references specific plan section or codebase file
|
|
264
|
-
- Security-sensitive plans escalated to sentinel (or confirmed not security-relevant)
|
|
265
|
-
- Verdict rendered: REVISE, HARDEN, or PROCEED
|
|
266
|
-
- Findings formatted for consumption by cook Phase 3 (if PROCEED) or plan (if REVISE)
|
|
267
|
-
- Strength Notes section acknowledges well-designed aspects of the plan
|
|
268
|
-
|
|
269
|
-
## Returns
|
|
270
|
-
|
|
271
|
-
| Artifact | Format | Location |
|
|
272
|
-
|----------|--------|----------|
|
|
273
|
-
| Adversary Report | Markdown | inline (stdout) |
|
|
274
|
-
| Threat findings | Structured list (CRITICAL/HIGH/MEDIUM) | inline |
|
|
275
|
-
| Risk matrix per dimension | Table | inline |
|
|
276
|
-
| Verdict + remediation list | Markdown | inline |
|
|
277
|
-
| Hardened plan notes (if PROCEED) | Text | passed to cook Phase 3 |
|
|
278
|
-
|
|
279
|
-
## Cost Profile
|
|
280
|
-
|
|
281
|
-
~4000-8000 tokens input (plan + codebase context), ~2000-3000 tokens output. Opus model for adversarial depth. Runs once per feature plan — high cost justified by preventing wasted implementation cycles.
|
|
282
|
-
|
|
283
|
-
**Scope guardrail:** adversary reviews THE PLAN only — never audits existing codebase quality or rewrites code.
|
|
1
|
+
---
|
|
2
|
+
name: adversary
|
|
3
|
+
description: Pre-implementation red-team analysis. Challenges plans before code is written — finds edge cases, security holes, scalability bottlenecks, error propagation risks, and integration conflicts. Catches flaws at plan time (10x cheaper than post-implementation).
|
|
4
|
+
metadata:
|
|
5
|
+
author: runedev
|
|
6
|
+
version: "0.1.0"
|
|
7
|
+
layer: L2
|
|
8
|
+
model: opus
|
|
9
|
+
group: quality
|
|
10
|
+
tools: "Read, Glob, Grep"
|
|
11
|
+
---
|
|
12
|
+
|
|
13
|
+
# adversary
|
|
14
|
+
|
|
15
|
+
## Purpose
|
|
16
|
+
|
|
17
|
+
Pre-implementation adversarial analysis. After a plan is approved but BEFORE code is written, adversary stress-tests the plan across 5 dimensions: edge cases, security, scalability, error propagation, and integration risk. It does NOT fix or redesign — it reports weaknesses so the plan can be hardened before implementation begins.
|
|
18
|
+
|
|
19
|
+
This fills the only gap in the plan-to-ship pipeline: all other quality skills (review, preflight, sentinel) operate AFTER code exists. Catching a flaw in a plan costs minutes; catching it in implementation costs hours.
|
|
20
|
+
|
|
21
|
+
<HARD-GATE>
|
|
22
|
+
adversary MUST NOT approve a plan without at least one specific challenge per dimension analyzed.
|
|
23
|
+
A report that says "plan looks solid" without concrete attack vectors is NOT a red-team analysis.
|
|
24
|
+
Every finding MUST reference the specific plan section, file, or assumption it challenges.
|
|
25
|
+
</HARD-GATE>
|
|
26
|
+
|
|
27
|
+
## Triggers
|
|
28
|
+
|
|
29
|
+
- Called by `cook` Phase 2.5 — after plan approved, before Phase 3 (TEST)
|
|
30
|
+
- `/rune adversary` — manual red-team analysis of any plan or design document
|
|
31
|
+
- Auto-trigger: when plan files are created in `.rune/` or `docs/plans/`
|
|
32
|
+
|
|
33
|
+
## Calls (outbound)
|
|
34
|
+
|
|
35
|
+
- `sentinel` (L2): deep security scan when adversary identifies auth/crypto/payment attack vectors in the plan
|
|
36
|
+
- `perf` (L2): scalability analysis when adversary identifies potential bottleneck patterns
|
|
37
|
+
- `scout` (L2): find existing code that might conflict with planned changes
|
|
38
|
+
- `docs-seeker` (L3): verify framework/API assumptions in the plan are correct and current
|
|
39
|
+
- `hallucination-guard` (L3): verify that APIs, packages, or patterns referenced in the plan actually exist
|
|
40
|
+
|
|
41
|
+
## Called By (inbound)
|
|
42
|
+
|
|
43
|
+
- `cook` (L1): Phase 2.5 — after plan approval, before TDD
|
|
44
|
+
- `plan` (L2): optional post-step for critical features
|
|
45
|
+
- `team` (L1): when decomposing large tasks, adversary validates the decomposition
|
|
46
|
+
- User: `/rune adversary` direct invocation
|
|
47
|
+
|
|
48
|
+
## Cross-Hub Connections
|
|
49
|
+
|
|
50
|
+
- `adversary` ← `cook` — plan produced → adversary challenges it → hardened plan feeds Phase 3
|
|
51
|
+
- `adversary` → `sentinel` — security attack vector identified → sentinel validates depth
|
|
52
|
+
- `adversary` → `perf` — scalability concern raised → perf quantifies the bottleneck
|
|
53
|
+
- `adversary` → `scout` — integration risk flagged → scout finds affected code
|
|
54
|
+
- `adversary` → `plan` — CRITICAL findings → plan revises before implementation
|
|
55
|
+
|
|
56
|
+
## Execution
|
|
57
|
+
|
|
58
|
+
### Step 0: Load Context
|
|
59
|
+
|
|
60
|
+
1. Read the plan document (from `.rune/features/<name>/plan.md`, phase file, or user-specified path)
|
|
61
|
+
2. Read the requirements document if it exists (`.rune/features/<name>/requirements.md` from BA)
|
|
62
|
+
3. Use `scout` to identify existing code files that the plan will touch or depend on
|
|
63
|
+
4. Identify the plan's core assumptions — what MUST be true for this plan to work?
|
|
64
|
+
|
|
65
|
+
### Step 1: Edge Case Analysis
|
|
66
|
+
|
|
67
|
+
Challenge the plan's handling of boundary conditions.
|
|
68
|
+
|
|
69
|
+
For each input/output/state transition in the plan, ask:
|
|
70
|
+
- **Empty/zero**: What happens with no data, zero items, empty strings, null users?
|
|
71
|
+
- **Overflow**: What happens at MAX — 10K items, 1MB payload, 1000 concurrent users?
|
|
72
|
+
- **Race conditions**: What if two operations happen simultaneously? Can state become inconsistent?
|
|
73
|
+
- **Partial failure**: What if step 3 of 5 fails? Is there rollback? Or orphaned state?
|
|
74
|
+
- **Invalid combinations**: What input combinations are technically possible but semantically nonsensical?
|
|
75
|
+
|
|
76
|
+
```
|
|
77
|
+
EDGE_CASE_TEMPLATE:
|
|
78
|
+
- Scenario: [specific edge case]
|
|
79
|
+
- Plan assumption: [what the plan assumes]
|
|
80
|
+
- Attack: [how this breaks]
|
|
81
|
+
- Impact: [what fails — data loss, crash, wrong result, security breach]
|
|
82
|
+
- Remediation: [1-sentence fix suggestion]
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
### Step 2: Security Attack Vectors
|
|
86
|
+
|
|
87
|
+
Analyze the plan for security weaknesses BEFORE any code exists.
|
|
88
|
+
|
|
89
|
+
- **Input trust boundaries**: Where does the plan accept external input? Is validation specified?
|
|
90
|
+
- **Authentication gaps**: Does the plan assume auth exists? Are there unprotected routes or actions?
|
|
91
|
+
- **Data exposure**: Could the planned API responses leak sensitive fields? Are there over-fetching risks?
|
|
92
|
+
- **Privilege escalation**: Can a normal user reach admin functionality through the planned flow?
|
|
93
|
+
- **Injection surfaces**: Does the plan involve dynamic queries, template rendering, or shell commands?
|
|
94
|
+
- **Dependency risk**: Does the plan introduce new dependencies? Are they well-maintained and trusted?
|
|
95
|
+
|
|
96
|
+
If any auth, crypto, or payment logic is in the plan: MUST call `rune:sentinel` for deep analysis.
|
|
97
|
+
|
|
98
|
+
```
|
|
99
|
+
SECURITY_TEMPLATE:
|
|
100
|
+
- Vector: [attack type — OWASP category if applicable]
|
|
101
|
+
- Entry point: [which part of the plan is vulnerable]
|
|
102
|
+
- Exploit scenario: [how an attacker would use this]
|
|
103
|
+
- Severity: CRITICAL | HIGH | MEDIUM
|
|
104
|
+
- Remediation: [what the plan should specify to prevent this]
|
|
105
|
+
```
|
|
106
|
+
|
|
107
|
+
### Step 3: Scalability Stress Test
|
|
108
|
+
|
|
109
|
+
Project the plan forward — what happens at 10x and 100x scale?
|
|
110
|
+
|
|
111
|
+
- **N+1 queries**: Does the plan describe data fetching that will create N+1 database calls?
|
|
112
|
+
- **Missing pagination**: Does the plan handle lists without specifying limits?
|
|
113
|
+
- **Synchronous bottlenecks**: Are there blocking operations in the hot path?
|
|
114
|
+
- **Cache invalidation**: If caching is planned, what happens when data changes? Stale reads?
|
|
115
|
+
- **State growth**: Does the plan accumulate state (in-memory, database, file system) without cleanup?
|
|
116
|
+
- **External service limits**: Does the plan account for rate limits on third-party APIs?
|
|
117
|
+
|
|
118
|
+
If bottleneck patterns detected: call `rune:perf` for quantitative analysis.
|
|
119
|
+
|
|
120
|
+
```
|
|
121
|
+
SCALE_TEMPLATE:
|
|
122
|
+
- Bottleneck: [what breaks at scale]
|
|
123
|
+
- Current plan: [what the plan specifies]
|
|
124
|
+
- At 10x: [what happens]
|
|
125
|
+
- At 100x: [what happens]
|
|
126
|
+
- Remediation: [what to add to the plan]
|
|
127
|
+
```
|
|
128
|
+
|
|
129
|
+
### Step 4: Error Propagation Analysis
|
|
130
|
+
|
|
131
|
+
Trace failure paths through the planned system.
|
|
132
|
+
|
|
133
|
+
- **Cascade failures**: If Service A fails, does the plan specify what happens to B, C, D?
|
|
134
|
+
- **Retry storms**: Does the plan include retries? Could retries amplify the failure?
|
|
135
|
+
- **Silent failures**: Are there operations that could fail without anyone knowing?
|
|
136
|
+
- **Inconsistent state**: If a multi-step operation fails midway, is the data left in a valid state?
|
|
137
|
+
- **User experience**: When things fail, what does the user see? Is there a degraded mode?
|
|
138
|
+
- **Recovery path**: After failure + fix, can the system resume? Or does it require manual intervention?
|
|
139
|
+
|
|
140
|
+
```
|
|
141
|
+
ERROR_TEMPLATE:
|
|
142
|
+
- Failure point: [where in the plan]
|
|
143
|
+
- Propagation: [what else breaks]
|
|
144
|
+
- User impact: [what the user experiences]
|
|
145
|
+
- Recovery: [how to get back to good state]
|
|
146
|
+
- Missing in plan: [what the plan should specify]
|
|
147
|
+
```
|
|
148
|
+
|
|
149
|
+
### Step 5: Integration Risk Assessment
|
|
150
|
+
|
|
151
|
+
Check for conflicts with existing code and architecture.
|
|
152
|
+
|
|
153
|
+
- Use `rune:scout` to find all files the plan will modify or depend on
|
|
154
|
+
- **Breaking changes**: Does the plan modify shared interfaces, types, or APIs that other code depends on?
|
|
155
|
+
- **Migration gaps**: Does the plan require database migrations? Are they reversible?
|
|
156
|
+
- **Configuration drift**: Does the plan add new environment variables, feature flags, or config files?
|
|
157
|
+
- **Test invalidation**: Will existing tests break from the planned changes?
|
|
158
|
+
- **Deployment ordering**: Does the plan require specific deployment sequence? (DB first, then API, then frontend?)
|
|
159
|
+
|
|
160
|
+
```
|
|
161
|
+
INTEGRATION_TEMPLATE:
|
|
162
|
+
- Conflict: [what clashes]
|
|
163
|
+
- Existing code: [file:line that would be affected]
|
|
164
|
+
- Plan assumption: [what the plan assumes about existing code]
|
|
165
|
+
- Reality: [what the existing code actually does]
|
|
166
|
+
- Remediation: [how to resolve the conflict]
|
|
167
|
+
```
|
|
168
|
+
|
|
169
|
+
### Step 6: Verdict and Report
|
|
170
|
+
|
|
171
|
+
Synthesize all findings into an actionable report.
|
|
172
|
+
|
|
173
|
+
**Before reporting, apply rigor filter:**
|
|
174
|
+
- Only report findings you can justify with specific references to the plan or codebase
|
|
175
|
+
- Do NOT report theoretical concerns that require 3+ unlikely conditions to trigger
|
|
176
|
+
- Prioritize findings that would cause the MOST wasted implementation time if discovered later
|
|
177
|
+
- Consolidate related findings — "auth is underspecified" not 5 separate auth findings
|
|
178
|
+
|
|
179
|
+
**Verdict logic:**
|
|
180
|
+
- Any CRITICAL finding → **REVISE** (plan must be updated before Phase 3)
|
|
181
|
+
- 3+ HIGH findings → **REVISE**
|
|
182
|
+
- HIGH findings with clear remediations → **HARDEN** (add remediations to plan, then proceed)
|
|
183
|
+
- Only MEDIUM/LOW findings → **PROCEED** (note findings for implementation awareness)
|
|
184
|
+
|
|
185
|
+
After reporting:
|
|
186
|
+
- If verdict is REVISE: return to `plan` with findings attached as constraints
|
|
187
|
+
- If verdict is HARDEN: present remediations to user for plan update
|
|
188
|
+
- If verdict is PROCEED: pass findings to cook Phase 3 as implementation notes
|
|
189
|
+
|
|
190
|
+
## Output Format
|
|
191
|
+
|
|
192
|
+
```
|
|
193
|
+
## Adversary Report: [feature/plan name]
|
|
194
|
+
- **Plan analyzed**: [path to plan file]
|
|
195
|
+
- **Dimensions checked**: [which of the 5 were relevant]
|
|
196
|
+
- **Findings**: [count by severity]
|
|
197
|
+
- **Verdict**: REVISE | HARDEN | PROCEED
|
|
198
|
+
|
|
199
|
+
### CRITICAL
|
|
200
|
+
- [ADV-001] [dimension]: [description with plan reference]
|
|
201
|
+
- Attack: [how this breaks]
|
|
202
|
+
- Remediation: [specific fix]
|
|
203
|
+
|
|
204
|
+
### HIGH
|
|
205
|
+
- [ADV-002] [dimension]: [description with plan reference]
|
|
206
|
+
- Attack: [how this breaks]
|
|
207
|
+
- Remediation: [specific fix]
|
|
208
|
+
|
|
209
|
+
### MEDIUM
|
|
210
|
+
- [ADV-003] [dimension]: [description]
|
|
211
|
+
|
|
212
|
+
### Strength Notes
|
|
213
|
+
- [what the plan does well — adversary is harsh but fair]
|
|
214
|
+
|
|
215
|
+
### Verdict
|
|
216
|
+
[Summary: why REVISE/HARDEN/PROCEED, what to do next]
|
|
217
|
+
```
|
|
218
|
+
|
|
219
|
+
## Workflow Modes
|
|
220
|
+
|
|
221
|
+
### Full Red-Team (default)
|
|
222
|
+
All 5 dimensions analyzed. Used for new features, architectural changes, security-sensitive plans.
|
|
223
|
+
|
|
224
|
+
### Quick Challenge (for smaller plans)
|
|
225
|
+
Skip Steps 3-4 (scalability, error propagation). Focus on edge cases, security, and integration.
|
|
226
|
+
Trigger: plan modifies < 3 files AND no auth/payment/data logic.
|
|
227
|
+
|
|
228
|
+
### Security-Focused
|
|
229
|
+
Steps 2 and 5 only (security + integration). Used when `sentinel` requests adversarial pre-analysis.
|
|
230
|
+
Trigger: plan involves auth, crypto, payment, or user data handling.
|
|
231
|
+
|
|
232
|
+
## Constraints
|
|
233
|
+
|
|
234
|
+
1. MUST challenge every plan — no rubber-stamping. At minimum, one finding per analyzed dimension
|
|
235
|
+
2. MUST NOT modify the plan or write code — adversary is read-only analysis
|
|
236
|
+
3. MUST reference specific plan sections or existing code for every finding
|
|
237
|
+
4. MUST escalate to sentinel when auth/crypto/payment attack vectors are identified
|
|
238
|
+
5. MUST use concrete attack scenarios, not vague warnings ("could be a problem" is NOT a finding)
|
|
239
|
+
6. MUST NOT block on MEDIUM/LOW findings — only CRITICAL and HIGH trigger REVISE verdict
|
|
240
|
+
7. MUST include Strength Notes — adversary finds weaknesses AND acknowledges what's well-designed
|
|
241
|
+
|
|
242
|
+
## Mesh Gates
|
|
243
|
+
|
|
244
|
+
| Gate | Requires | If Missing |
|
|
245
|
+
|------|----------|------------|
|
|
246
|
+
| Plan Gate | A plan document exists (from plan skill or user-provided) | Cannot run — ask for plan first |
|
|
247
|
+
| Codebase Gate | Access to existing codebase (for integration checks) | Skip Step 5, note in report |
|
|
248
|
+
|
|
249
|
+
## Sharp Edges
|
|
250
|
+
|
|
251
|
+
| Failure Mode | Severity | Mitigation |
|
|
252
|
+
|---|---|---|
|
|
253
|
+
| Over-challenging — nitpicking every line of the plan | HIGH | Rigor filter: only findings you can justify with specific references. Skip theoretical 3+ condition chains |
|
|
254
|
+
| False security alarms — flagging secure patterns as vulnerable | HIGH | Call sentinel for validation before reporting security findings as CRITICAL |
|
|
255
|
+
| Analysis paralysis — too many findings block all progress | MEDIUM | Max 3 CRITICAL + 5 HIGH. If more found, consolidate or prioritize top impact |
|
|
256
|
+
| Missing context — challenging plan without understanding existing codebase | HIGH | Step 0 MUST load existing code context via scout before challenging |
|
|
257
|
+
| Scope creep — reviewing existing code quality instead of plan quality | MEDIUM | Adversary reviews THE PLAN, not the codebase. Existing code is context only |
|
|
258
|
+
| Redundancy with review/preflight — duplicating post-implementation checks | MEDIUM | Adversary operates PRE-implementation only. Never run adversary on existing code |
|
|
259
|
+
|
|
260
|
+
## Done When
|
|
261
|
+
|
|
262
|
+
- All relevant dimensions analyzed (minimum: edge cases + security + integration)
|
|
263
|
+
- Every finding references specific plan section or codebase file
|
|
264
|
+
- Security-sensitive plans escalated to sentinel (or confirmed not security-relevant)
|
|
265
|
+
- Verdict rendered: REVISE, HARDEN, or PROCEED
|
|
266
|
+
- Findings formatted for consumption by cook Phase 3 (if PROCEED) or plan (if REVISE)
|
|
267
|
+
- Strength Notes section acknowledges well-designed aspects of the plan
|
|
268
|
+
|
|
269
|
+
## Returns
|
|
270
|
+
|
|
271
|
+
| Artifact | Format | Location |
|
|
272
|
+
|----------|--------|----------|
|
|
273
|
+
| Adversary Report | Markdown | inline (stdout) |
|
|
274
|
+
| Threat findings | Structured list (CRITICAL/HIGH/MEDIUM) | inline |
|
|
275
|
+
| Risk matrix per dimension | Table | inline |
|
|
276
|
+
| Verdict + remediation list | Markdown | inline |
|
|
277
|
+
| Hardened plan notes (if PROCEED) | Text | passed to cook Phase 3 |
|
|
278
|
+
|
|
279
|
+
## Cost Profile
|
|
280
|
+
|
|
281
|
+
~4000-8000 tokens input (plan + codebase context), ~2000-3000 tokens output. Opus model for adversarial depth. Runs once per feature plan — high cost justified by preventing wasted implementation cycles.
|
|
282
|
+
|
|
283
|
+
**Scope guardrail:** adversary reviews THE PLAN only — never audits existing codebase quality or rewrites code.
|