@rune-kit/rune 2.8.0 → 2.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (287) hide show
  1. package/LICENSE +21 -21
  2. package/README.md +68 -34
  3. package/agents/adversary.md +27 -0
  4. package/agents/architect.md +19 -29
  5. package/agents/asset-creator.md +18 -4
  6. package/agents/audit.md +25 -4
  7. package/agents/autopsy.md +19 -4
  8. package/agents/ba.md +35 -0
  9. package/agents/brainstorm.md +31 -4
  10. package/agents/browser-pilot.md +21 -4
  11. package/agents/coder.md +21 -29
  12. package/agents/completion-gate.md +20 -4
  13. package/agents/constraint-check.md +18 -4
  14. package/agents/context-engine.md +22 -4
  15. package/agents/context-pack.md +32 -0
  16. package/agents/cook.md +41 -4
  17. package/agents/db.md +19 -4
  18. package/agents/debug.md +33 -4
  19. package/agents/dependency-doctor.md +20 -4
  20. package/agents/deploy.md +27 -4
  21. package/agents/design.md +22 -4
  22. package/agents/doc-processor.md +27 -0
  23. package/agents/docs-seeker.md +19 -4
  24. package/agents/docs.md +31 -0
  25. package/agents/fix.md +37 -4
  26. package/agents/git.md +29 -0
  27. package/agents/hallucination-guard.md +20 -4
  28. package/agents/incident.md +21 -4
  29. package/agents/integrity-check.md +18 -4
  30. package/agents/journal.md +19 -4
  31. package/agents/launch.md +32 -4
  32. package/agents/logic-guardian.md +26 -11
  33. package/agents/marketing.md +23 -4
  34. package/agents/mcp-builder.md +26 -0
  35. package/agents/neural-memory.md +30 -0
  36. package/agents/onboard.md +22 -4
  37. package/agents/perf.md +21 -4
  38. package/agents/plan.md +29 -4
  39. package/agents/preflight.md +22 -4
  40. package/agents/problem-solver.md +20 -4
  41. package/agents/rescue.md +23 -4
  42. package/agents/research.md +19 -4
  43. package/agents/researcher.md +19 -29
  44. package/agents/retro.md +32 -0
  45. package/agents/review-intake.md +20 -4
  46. package/agents/review.md +32 -4
  47. package/agents/reviewer.md +20 -28
  48. package/agents/safeguard.md +19 -4
  49. package/agents/sast.md +18 -4
  50. package/agents/scaffold.md +41 -0
  51. package/agents/scanner.md +19 -28
  52. package/agents/scope-guard.md +18 -4
  53. package/agents/scout.md +23 -4
  54. package/agents/sentinel-env.md +26 -0
  55. package/agents/sentinel.md +33 -4
  56. package/agents/sequential-thinking.md +20 -4
  57. package/agents/session-bridge.md +24 -4
  58. package/agents/skill-forge.md +22 -4
  59. package/agents/skill-router.md +26 -4
  60. package/agents/slides.md +24 -0
  61. package/agents/surgeon.md +19 -4
  62. package/agents/team.md +30 -4
  63. package/agents/test.md +36 -4
  64. package/agents/trend-scout.md +17 -4
  65. package/agents/verification.md +20 -4
  66. package/agents/video-creator.md +20 -4
  67. package/agents/watchdog.md +19 -4
  68. package/agents/worktree.md +17 -4
  69. package/commands/rune.md +168 -168
  70. package/compiler/__tests__/analytics.test.js +370 -0
  71. package/compiler/adapters/openclaw.js +2 -2
  72. package/compiler/analytics.js +385 -0
  73. package/compiler/bin/rune.js +68 -2
  74. package/compiler/dashboard.js +883 -0
  75. package/compiler/transforms/branding.js +1 -1
  76. package/contexts/dev.md +34 -34
  77. package/contexts/research.md +43 -43
  78. package/contexts/review.md +55 -55
  79. package/extensions/ai-ml/PACK.md +88 -88
  80. package/extensions/ai-ml/skills/ai-agents.md +172 -172
  81. package/extensions/ai-ml/skills/code-sandbox.md +187 -187
  82. package/extensions/ai-ml/skills/deep-research.md +146 -146
  83. package/extensions/ai-ml/skills/embedding-search.md +66 -66
  84. package/extensions/ai-ml/skills/fine-tuning-guide.md +74 -74
  85. package/extensions/ai-ml/skills/llm-architect.md +125 -125
  86. package/extensions/ai-ml/skills/llm-integration.md +64 -64
  87. package/extensions/ai-ml/skills/prompt-patterns.md +72 -72
  88. package/extensions/ai-ml/skills/rag-patterns.md +66 -66
  89. package/extensions/ai-ml/skills/web-extraction.md +114 -114
  90. package/extensions/analytics/PACK.md +92 -92
  91. package/extensions/analytics/skills/ab-testing.md +72 -72
  92. package/extensions/analytics/skills/dashboard-patterns.md +83 -83
  93. package/extensions/analytics/skills/data-validation.md +68 -68
  94. package/extensions/analytics/skills/funnel-analysis.md +81 -81
  95. package/extensions/analytics/skills/sql-patterns.md +57 -57
  96. package/extensions/analytics/skills/statistical-analysis.md +79 -79
  97. package/extensions/analytics/skills/tracking-setup.md +71 -71
  98. package/extensions/backend/PACK.md +104 -104
  99. package/extensions/backend/skills/api-patterns.md +84 -84
  100. package/extensions/backend/skills/async-pipeline.md +193 -193
  101. package/extensions/backend/skills/auth-patterns.md +97 -97
  102. package/extensions/backend/skills/background-jobs.md +133 -133
  103. package/extensions/backend/skills/caching-patterns.md +108 -108
  104. package/extensions/backend/skills/cli-generation.md +133 -133
  105. package/extensions/backend/skills/database-patterns.md +87 -87
  106. package/extensions/backend/skills/middleware-patterns.md +104 -104
  107. package/extensions/chrome-ext/PACK.md +93 -93
  108. package/extensions/chrome-ext/skills/cws-preflight.md +143 -143
  109. package/extensions/chrome-ext/skills/cws-publish.md +104 -104
  110. package/extensions/chrome-ext/skills/ext-ai-integration.md +251 -251
  111. package/extensions/chrome-ext/skills/ext-messaging.md +139 -139
  112. package/extensions/chrome-ext/skills/ext-storage.md +133 -133
  113. package/extensions/chrome-ext/skills/mv3-scaffold.md +164 -164
  114. package/extensions/content/PACK.md +96 -96
  115. package/extensions/content/skills/blog-patterns.md +88 -88
  116. package/extensions/content/skills/cms-integration.md +131 -131
  117. package/extensions/content/skills/content-scoring.md +107 -107
  118. package/extensions/content/skills/i18n.md +83 -83
  119. package/extensions/content/skills/mdx-authoring.md +137 -137
  120. package/extensions/content/skills/reference.md +1014 -1014
  121. package/extensions/content/skills/seo-patterns.md +67 -67
  122. package/extensions/content/skills/video-repurpose.md +153 -153
  123. package/extensions/devops/PACK.md +101 -101
  124. package/extensions/devops/skills/chaos-testing.md +67 -67
  125. package/extensions/devops/skills/ci-cd.md +75 -75
  126. package/extensions/devops/skills/docker.md +58 -58
  127. package/extensions/devops/skills/edge-serverless.md +163 -163
  128. package/extensions/devops/skills/infra-as-code.md +158 -158
  129. package/extensions/devops/skills/kubernetes.md +110 -110
  130. package/extensions/devops/skills/monitoring.md +57 -57
  131. package/extensions/devops/skills/server-setup.md +64 -64
  132. package/extensions/devops/skills/ssl-domain.md +42 -42
  133. package/extensions/ecommerce/PACK.md +116 -116
  134. package/extensions/ecommerce/skills/cart-system.md +79 -79
  135. package/extensions/ecommerce/skills/inventory-mgmt.md +102 -102
  136. package/extensions/ecommerce/skills/order-management.md +126 -126
  137. package/extensions/ecommerce/skills/payment-integration.md +472 -472
  138. package/extensions/ecommerce/skills/shopify-dev.md +69 -69
  139. package/extensions/ecommerce/skills/subscription-billing.md +93 -93
  140. package/extensions/ecommerce/skills/tax-compliance.md +117 -117
  141. package/extensions/gamedev/PACK.md +142 -142
  142. package/extensions/gamedev/skills/asset-pipeline.md +74 -74
  143. package/extensions/gamedev/skills/audio-system.md +129 -129
  144. package/extensions/gamedev/skills/camera-system.md +87 -87
  145. package/extensions/gamedev/skills/ecs.md +98 -98
  146. package/extensions/gamedev/skills/game-loops.md +72 -72
  147. package/extensions/gamedev/skills/input-system.md +199 -199
  148. package/extensions/gamedev/skills/multiplayer.md +180 -180
  149. package/extensions/gamedev/skills/particles.md +105 -105
  150. package/extensions/gamedev/skills/physics-engine.md +89 -89
  151. package/extensions/gamedev/skills/scene-management.md +146 -146
  152. package/extensions/gamedev/skills/threejs-patterns.md +90 -90
  153. package/extensions/gamedev/skills/webgl.md +71 -71
  154. package/extensions/mobile/PACK.md +106 -106
  155. package/extensions/mobile/skills/app-store-connect.md +152 -152
  156. package/extensions/mobile/skills/app-store-prep.md +66 -66
  157. package/extensions/mobile/skills/deep-linking.md +109 -109
  158. package/extensions/mobile/skills/flutter.md +60 -60
  159. package/extensions/mobile/skills/ios-build-pipeline.md +142 -142
  160. package/extensions/mobile/skills/native-bridge.md +66 -66
  161. package/extensions/mobile/skills/ota-updates.md +97 -97
  162. package/extensions/mobile/skills/push-notifications.md +111 -111
  163. package/extensions/mobile/skills/react-native.md +82 -82
  164. package/extensions/saas/PACK.md +116 -116
  165. package/extensions/saas/skills/billing-integration.md +200 -200
  166. package/extensions/saas/skills/feature-flags.md +130 -130
  167. package/extensions/saas/skills/multi-tenant.md +103 -103
  168. package/extensions/saas/skills/onboarding-flow.md +139 -139
  169. package/extensions/saas/skills/subscription-flow.md +95 -95
  170. package/extensions/saas/skills/team-management.md +144 -144
  171. package/extensions/security/PACK.md +99 -99
  172. package/extensions/security/skills/api-security.md +140 -140
  173. package/extensions/security/skills/compliance.md +68 -68
  174. package/extensions/security/skills/owasp-audit.md +64 -64
  175. package/extensions/security/skills/pentest-patterns.md +77 -77
  176. package/extensions/security/skills/secret-mgmt.md +65 -65
  177. package/extensions/security/skills/supply-chain.md +65 -65
  178. package/extensions/trading/PACK.md +80 -80
  179. package/extensions/trading/skills/chart-components.md +55 -55
  180. package/extensions/trading/skills/experiment-loop.md +125 -125
  181. package/extensions/trading/skills/fintech-patterns.md +47 -47
  182. package/extensions/trading/skills/indicator-library.md +58 -58
  183. package/extensions/trading/skills/quant-analysis.md +111 -111
  184. package/extensions/trading/skills/realtime-data.md +58 -58
  185. package/extensions/trading/skills/trade-logic.md +104 -104
  186. package/extensions/ui/PACK.md +130 -130
  187. package/extensions/ui/skills/a11y-audit.md +91 -91
  188. package/extensions/ui/skills/animation-patterns.md +127 -106
  189. package/extensions/ui/skills/component-patterns.md +100 -75
  190. package/extensions/ui/skills/design-decision.md +108 -108
  191. package/extensions/ui/skills/design-system.md +68 -68
  192. package/extensions/ui/skills/landing-patterns.md +155 -155
  193. package/extensions/ui/skills/palette-picker.md +173 -173
  194. package/extensions/ui/skills/react-health.md +90 -90
  195. package/extensions/ui/skills/type-system.md +125 -125
  196. package/extensions/ui/skills/web-vitals.md +153 -153
  197. package/extensions/zalo/PACK.md +145 -145
  198. package/extensions/zalo/skills/zalo-oa-mcp.md +317 -317
  199. package/extensions/zalo/skills/zalo-oa-messaging.md +429 -429
  200. package/extensions/zalo/skills/zalo-oa-setup.md +236 -236
  201. package/extensions/zalo/skills/zalo-oa-webhook.md +189 -189
  202. package/extensions/zalo/skills/zalo-personal-messaging.md +194 -194
  203. package/extensions/zalo/skills/zalo-personal-setup.md +153 -153
  204. package/extensions/zalo/skills/zalo-rate-guard.md +219 -219
  205. package/hooks/auto-format/index.cjs +48 -48
  206. package/hooks/context-watch/index.cjs +95 -68
  207. package/hooks/hooks.json +111 -111
  208. package/hooks/metrics-collector/index.cjs +86 -42
  209. package/hooks/post-session-reflect/index.cjs +189 -153
  210. package/hooks/pre-compact/index.cjs +95 -95
  211. package/hooks/run-hook.cmd +1 -1
  212. package/hooks/secrets-scan/index.cjs +100 -100
  213. package/hooks/session-start/index.cjs +71 -65
  214. package/hooks/typecheck/index.cjs +65 -65
  215. package/package.json +63 -63
  216. package/references/ui-pro-max-data/LICENSE-UI-PRO-MAX +21 -21
  217. package/references/ui-pro-max-data/charts.csv +26 -26
  218. package/references/ui-pro-max-data/colors.csv +161 -161
  219. package/references/ui-pro-max-data/styles.csv +68 -68
  220. package/references/ui-pro-max-data/typography.csv +74 -74
  221. package/references/ui-pro-max-data/ui-reasoning.csv +162 -162
  222. package/references/ui-pro-max-data/ux-guidelines.csv +99 -99
  223. package/skills/adversary/SKILL.md +283 -283
  224. package/skills/asset-creator/SKILL.md +157 -157
  225. package/skills/audit/SKILL.md +148 -2
  226. package/skills/autopsy/SKILL.md +335 -259
  227. package/skills/autopsy/references/repo-analysis-patterns.md +113 -0
  228. package/skills/ba/SKILL.md +72 -2
  229. package/skills/brainstorm/SKILL.md +342 -341
  230. package/skills/browser-pilot/SKILL.md +168 -168
  231. package/skills/constraint-check/SKILL.md +165 -165
  232. package/skills/context-engine/SKILL.md +404 -404
  233. package/skills/cook/SKILL.md +917 -834
  234. package/skills/cook/references/output-format.md +33 -0
  235. package/skills/db/SKILL.md +273 -272
  236. package/skills/debug/SKILL.md +465 -443
  237. package/skills/dependency-doctor/SKILL.md +265 -235
  238. package/skills/deploy/SKILL.md +274 -231
  239. package/skills/design/DESIGN-REFERENCE.md +365 -365
  240. package/skills/design/SKILL.md +589 -482
  241. package/skills/doc-processor/SKILL.md +254 -254
  242. package/skills/docs/SKILL.md +374 -373
  243. package/skills/docs-seeker/SKILL.md +177 -177
  244. package/skills/fix/SKILL.md +330 -308
  245. package/skills/git/SKILL.md +339 -339
  246. package/skills/graft/SKILL.md +352 -0
  247. package/skills/graft/references/challenge-framework.md +98 -0
  248. package/skills/graft/references/mode-decision.md +44 -0
  249. package/skills/hallucination-guard/SKILL.md +219 -219
  250. package/skills/incident/SKILL.md +254 -251
  251. package/skills/integrity-check/SKILL.md +169 -169
  252. package/skills/journal/SKILL.md +240 -238
  253. package/skills/launch/SKILL.md +344 -342
  254. package/skills/logic-guardian/SKILL.md +251 -251
  255. package/skills/marketing/SKILL.md +290 -245
  256. package/skills/mcp-builder/SKILL.md +425 -423
  257. package/skills/mcp-builder/references/auto-discovery-pattern.md +169 -0
  258. package/skills/neural-memory/SKILL.md +362 -362
  259. package/skills/onboard/SKILL.md +404 -403
  260. package/skills/perf/SKILL.md +346 -346
  261. package/skills/plan/SKILL.md +433 -370
  262. package/skills/plan/references/feature-map.md +84 -0
  263. package/skills/preflight/SKILL.md +415 -396
  264. package/skills/problem-solver/SKILL.md +380 -284
  265. package/skills/rescue/SKILL.md +474 -450
  266. package/skills/retro/SKILL.md +5 -1
  267. package/skills/review/SKILL.md +612 -535
  268. package/skills/review-intake/SKILL.md +249 -249
  269. package/skills/safeguard/SKILL.md +200 -200
  270. package/skills/sast/SKILL.md +190 -190
  271. package/skills/scaffold/SKILL.md +328 -286
  272. package/skills/scope-guard/SKILL.md +180 -162
  273. package/skills/scout/SKILL.md +263 -263
  274. package/skills/sentinel/SKILL.md +382 -353
  275. package/skills/sentinel-env/SKILL.md +254 -254
  276. package/skills/sequential-thinking/SKILL.md +234 -234
  277. package/skills/session-bridge/SKILL.md +543 -397
  278. package/skills/skill-forge/SKILL.md +581 -539
  279. package/skills/skill-router/{skill.md → SKILL.md} +30 -2
  280. package/skills/surgeon/SKILL.md +215 -215
  281. package/skills/team/SKILL.md +556 -514
  282. package/skills/test/SKILL.md +614 -587
  283. package/skills/trend-scout/SKILL.md +145 -145
  284. package/skills/verification/SKILL.md +326 -325
  285. package/skills/video-creator/SKILL.md +201 -201
  286. package/skills/watchdog/SKILL.md +168 -168
  287. package/skills/worktree/SKILL.md +140 -140
@@ -1,187 +1,187 @@
1
- ---
2
- name: "code-sandbox"
3
- pack: "@rune/ai-ml"
4
- description: "Secure code execution for AI agents — sandboxed environments for running LLM-generated code safely with container isolation, resource limits, and timeout enforcement."
5
- model: sonnet
6
- tools: [Read, Edit, Write, Grep, Glob, Bash]
7
- ---
8
-
9
- # code-sandbox
10
-
11
- Secure code execution for AI agents — sandboxed environments for running LLM-generated code safely. Covers container isolation, resource limits, timeout enforcement, file system boundaries, and output capture for code interpreter, CI/CD, and interactive development use cases.
12
-
13
- #### Workflow
14
-
15
- **Step 1 — Assess execution requirements**
16
- Determine what kind of code the agent needs to run:
17
-
18
- | Use Case | Isolation Level | Runtime |
19
- |---|---|---|
20
- | Code interpreter (data analysis, math) | High — untrusted code | Python + pandas/numpy |
21
- | Build/test pipeline | Medium — project code | Node.js / Python with project deps |
22
- | Interactive preview (web app) | Medium — expose HTTP port | Node.js + browser preview |
23
- | Shell commands (file ops, git) | Low — trusted context | System shell with path restrictions |
24
-
25
- **Step 2 — Configure sandbox environment**
26
- Emit sandbox configuration based on use case:
27
-
28
- ```typescript
29
- // Sandbox factory — select isolation level by use case
30
- interface SandboxConfig {
31
- language: 'python' | 'javascript' | 'typescript';
32
- timeout: number; // max execution time in ms
33
- memoryLimit: number; // max memory in MB
34
- networkAccess: boolean;
35
- fileSystemRoot: string; // restricted working directory
36
- allowedModules: string[];
37
- }
38
-
39
- const SANDBOX_PRESETS: Record<string, SandboxConfig> = {
40
- 'code-interpreter': {
41
- language: 'python',
42
- timeout: 30_000,
43
- memoryLimit: 256,
44
- networkAccess: false,
45
- fileSystemRoot: '/workspace',
46
- allowedModules: ['pandas', 'numpy', 'matplotlib', 'scipy', 'json', 'csv', 'math'],
47
- },
48
- 'build-test': {
49
- language: 'typescript',
50
- timeout: 120_000,
51
- memoryLimit: 512,
52
- networkAccess: true, // needs npm registry
53
- fileSystemRoot: '/project',
54
- allowedModules: ['*'], // project dependencies
55
- },
56
- 'preview': {
57
- language: 'javascript',
58
- timeout: 300_000,
59
- memoryLimit: 256,
60
- networkAccess: true,
61
- fileSystemRoot: '/app',
62
- allowedModules: ['*'],
63
- },
64
- };
65
- ```
66
-
67
- **Step 3 — Implement execution with resource limits**
68
- Emit code execution wrapper with safety boundaries:
69
-
70
- ```typescript
71
- // Docker-based sandbox execution
72
- import { spawn } from 'child_process';
73
-
74
- interface ExecutionResult {
75
- stdout: string;
76
- stderr: string;
77
- exitCode: number;
78
- durationMs: number;
79
- timedOut: boolean;
80
- }
81
-
82
- async function executeInSandbox(
83
- code: string,
84
- config: SandboxConfig
85
- ): Promise<ExecutionResult> {
86
- const start = Date.now();
87
-
88
- // Write code to temp file in sandbox root
89
- const codePath = `${config.fileSystemRoot}/run.${config.language === 'python' ? 'py' : 'ts'}`;
90
- await writeFile(codePath, code);
91
-
92
- const proc = spawn('docker', [
93
- 'run', '--rm',
94
- '--memory', `${config.memoryLimit}m`,
95
- '--cpus', '1',
96
- '--network', config.networkAccess ? 'bridge' : 'none',
97
- '--read-only',
98
- '--tmpfs', '/tmp:size=64m',
99
- '-v', `${config.fileSystemRoot}:/workspace:ro`,
100
- '-w', '/workspace',
101
- `sandbox-${config.language}:latest`,
102
- config.language === 'python' ? 'python' : 'npx tsx',
103
- `/workspace/run.${config.language === 'python' ? 'py' : 'ts'}`,
104
- ]);
105
-
106
- let stdout = '';
107
- let stderr = '';
108
- let timedOut = false;
109
-
110
- proc.stdout.on('data', (d) => { stdout += d.toString(); });
111
- proc.stderr.on('data', (d) => { stderr += d.toString(); });
112
-
113
- const timeout = setTimeout(() => {
114
- timedOut = true;
115
- proc.kill('SIGKILL');
116
- }, config.timeout);
117
-
118
- const exitCode = await new Promise<number>((resolve) => {
119
- proc.on('close', (code) => {
120
- clearTimeout(timeout);
121
- resolve(code ?? 1);
122
- });
123
- });
124
-
125
- return { stdout, stderr, exitCode, durationMs: Date.now() - start, timedOut };
126
- }
127
- ```
128
-
129
- **Step 4 — Code interpreter mode (stateful sessions)**
130
- For multi-turn code execution where variables persist between runs:
131
-
132
- ```typescript
133
- // Stateful code interpreter — variables persist across executions
134
- interface CodeSession {
135
- id: string;
136
- language: 'python' | 'javascript';
137
- history: { code: string; result: ExecutionResult }[];
138
- }
139
-
140
- async function runInSession(
141
- session: CodeSession,
142
- code: string
143
- ): Promise<ExecutionResult> {
144
- // Python: use exec() with persistent globals dict
145
- // JavaScript: use Node.js vm module with persistent context
146
- const wrappedCode = session.language === 'python'
147
- ? `exec(${JSON.stringify(code)}, _globals)`
148
- : code;
149
-
150
- const result = await executeInSandbox(wrappedCode, SANDBOX_PRESETS['code-interpreter']);
151
-
152
- // Append to history (immutable update)
153
- session.history = [...session.history, { code, result }];
154
-
155
- return result;
156
- }
157
-
158
- // Rich output capture — not just stdout
159
- interface RichOutput {
160
- text?: string;
161
- images?: { data: string; mimeType: string }[]; // base64 encoded
162
- tables?: { headers: string[]; rows: string[][] }[];
163
- error?: string;
164
- }
165
- ```
166
-
167
- **Step 5 — Security boundaries**
168
- Enforce isolation guarantees:
169
-
170
- | Boundary | Enforcement |
171
- |---|---|
172
- | File system | Read-only mount + tmpfs for temp files. No access to host filesystem. |
173
- | Network | `--network none` for code interpreter. Whitelist for build/test. |
174
- | Memory | Docker `--memory` limit. OOM killed if exceeded. |
175
- | CPU | Docker `--cpus` limit. Prevents crypto mining / infinite loops. |
176
- | Time | Kill process after timeout. Return partial output. |
177
- | Secrets | Never mount env vars or secrets into sandbox container. |
178
- | Output size | Cap stdout/stderr at 1MB. Truncate with `[output truncated]` marker. |
179
-
180
- #### Sharp Edges
181
-
182
- | Failure Mode | Mitigation |
183
- |---|---|
184
- | Sandbox escape via Docker vulnerability | Pin Docker version; use rootless Docker; consider gVisor/Firecracker for high-security |
185
- | Code writes to /tmp exhausting disk | Use `--tmpfs` with size limit (64MB default) |
186
- | Infinite loop inside sandbox hangs API | Hard timeout with SIGKILL — never rely on SIGTERM alone |
187
- | Stateful session grows unbounded memory | Limit session history to last 50 executions; reset context on overflow |
1
+ ---
2
+ name: "code-sandbox"
3
+ pack: "@rune/ai-ml"
4
+ description: "Secure code execution for AI agents — sandboxed environments for running LLM-generated code safely with container isolation, resource limits, and timeout enforcement."
5
+ model: sonnet
6
+ tools: [Read, Edit, Write, Grep, Glob, Bash]
7
+ ---
8
+
9
+ # code-sandbox
10
+
11
+ Secure code execution for AI agents — sandboxed environments for running LLM-generated code safely. Covers container isolation, resource limits, timeout enforcement, file system boundaries, and output capture for code interpreter, CI/CD, and interactive development use cases.
12
+
13
+ #### Workflow
14
+
15
+ **Step 1 — Assess execution requirements**
16
+ Determine what kind of code the agent needs to run:
17
+
18
+ | Use Case | Isolation Level | Runtime |
19
+ |---|---|---|
20
+ | Code interpreter (data analysis, math) | High — untrusted code | Python + pandas/numpy |
21
+ | Build/test pipeline | Medium — project code | Node.js / Python with project deps |
22
+ | Interactive preview (web app) | Medium — expose HTTP port | Node.js + browser preview |
23
+ | Shell commands (file ops, git) | Low — trusted context | System shell with path restrictions |
24
+
25
+ **Step 2 — Configure sandbox environment**
26
+ Emit sandbox configuration based on use case:
27
+
28
+ ```typescript
29
+ // Sandbox factory — select isolation level by use case
30
+ interface SandboxConfig {
31
+ language: 'python' | 'javascript' | 'typescript';
32
+ timeout: number; // max execution time in ms
33
+ memoryLimit: number; // max memory in MB
34
+ networkAccess: boolean;
35
+ fileSystemRoot: string; // restricted working directory
36
+ allowedModules: string[];
37
+ }
38
+
39
+ const SANDBOX_PRESETS: Record<string, SandboxConfig> = {
40
+ 'code-interpreter': {
41
+ language: 'python',
42
+ timeout: 30_000,
43
+ memoryLimit: 256,
44
+ networkAccess: false,
45
+ fileSystemRoot: '/workspace',
46
+ allowedModules: ['pandas', 'numpy', 'matplotlib', 'scipy', 'json', 'csv', 'math'],
47
+ },
48
+ 'build-test': {
49
+ language: 'typescript',
50
+ timeout: 120_000,
51
+ memoryLimit: 512,
52
+ networkAccess: true, // needs npm registry
53
+ fileSystemRoot: '/project',
54
+ allowedModules: ['*'], // project dependencies
55
+ },
56
+ 'preview': {
57
+ language: 'javascript',
58
+ timeout: 300_000,
59
+ memoryLimit: 256,
60
+ networkAccess: true,
61
+ fileSystemRoot: '/app',
62
+ allowedModules: ['*'],
63
+ },
64
+ };
65
+ ```
66
+
67
+ **Step 3 — Implement execution with resource limits**
68
+ Emit code execution wrapper with safety boundaries:
69
+
70
+ ```typescript
71
+ // Docker-based sandbox execution
72
+ import { spawn } from 'child_process';
73
+
74
+ interface ExecutionResult {
75
+ stdout: string;
76
+ stderr: string;
77
+ exitCode: number;
78
+ durationMs: number;
79
+ timedOut: boolean;
80
+ }
81
+
82
+ async function executeInSandbox(
83
+ code: string,
84
+ config: SandboxConfig
85
+ ): Promise<ExecutionResult> {
86
+ const start = Date.now();
87
+
88
+ // Write code to temp file in sandbox root
89
+ const codePath = `${config.fileSystemRoot}/run.${config.language === 'python' ? 'py' : 'ts'}`;
90
+ await writeFile(codePath, code);
91
+
92
+ const proc = spawn('docker', [
93
+ 'run', '--rm',
94
+ '--memory', `${config.memoryLimit}m`,
95
+ '--cpus', '1',
96
+ '--network', config.networkAccess ? 'bridge' : 'none',
97
+ '--read-only',
98
+ '--tmpfs', '/tmp:size=64m',
99
+ '-v', `${config.fileSystemRoot}:/workspace:ro`,
100
+ '-w', '/workspace',
101
+ `sandbox-${config.language}:latest`,
102
+ config.language === 'python' ? 'python' : 'npx tsx',
103
+ `/workspace/run.${config.language === 'python' ? 'py' : 'ts'}`,
104
+ ]);
105
+
106
+ let stdout = '';
107
+ let stderr = '';
108
+ let timedOut = false;
109
+
110
+ proc.stdout.on('data', (d) => { stdout += d.toString(); });
111
+ proc.stderr.on('data', (d) => { stderr += d.toString(); });
112
+
113
+ const timeout = setTimeout(() => {
114
+ timedOut = true;
115
+ proc.kill('SIGKILL');
116
+ }, config.timeout);
117
+
118
+ const exitCode = await new Promise<number>((resolve) => {
119
+ proc.on('close', (code) => {
120
+ clearTimeout(timeout);
121
+ resolve(code ?? 1);
122
+ });
123
+ });
124
+
125
+ return { stdout, stderr, exitCode, durationMs: Date.now() - start, timedOut };
126
+ }
127
+ ```
128
+
129
+ **Step 4 — Code interpreter mode (stateful sessions)**
130
+ For multi-turn code execution where variables persist between runs:
131
+
132
+ ```typescript
133
+ // Stateful code interpreter — variables persist across executions
134
+ interface CodeSession {
135
+ id: string;
136
+ language: 'python' | 'javascript';
137
+ history: { code: string; result: ExecutionResult }[];
138
+ }
139
+
140
+ async function runInSession(
141
+ session: CodeSession,
142
+ code: string
143
+ ): Promise<ExecutionResult> {
144
+ // Python: use exec() with persistent globals dict
145
+ // JavaScript: use Node.js vm module with persistent context
146
+ const wrappedCode = session.language === 'python'
147
+ ? `exec(${JSON.stringify(code)}, _globals)`
148
+ : code;
149
+
150
+ const result = await executeInSandbox(wrappedCode, SANDBOX_PRESETS['code-interpreter']);
151
+
152
+ // Append to history (immutable update)
153
+ session.history = [...session.history, { code, result }];
154
+
155
+ return result;
156
+ }
157
+
158
+ // Rich output capture — not just stdout
159
+ interface RichOutput {
160
+ text?: string;
161
+ images?: { data: string; mimeType: string }[]; // base64 encoded
162
+ tables?: { headers: string[]; rows: string[][] }[];
163
+ error?: string;
164
+ }
165
+ ```
166
+
167
+ **Step 5 — Security boundaries**
168
+ Enforce isolation guarantees:
169
+
170
+ | Boundary | Enforcement |
171
+ |---|---|
172
+ | File system | Read-only mount + tmpfs for temp files. No access to host filesystem. |
173
+ | Network | `--network none` for code interpreter. Whitelist for build/test. |
174
+ | Memory | Docker `--memory` limit. OOM killed if exceeded. |
175
+ | CPU | Docker `--cpus` limit. Prevents crypto mining / infinite loops. |
176
+ | Time | Kill process after timeout. Return partial output. |
177
+ | Secrets | Never mount env vars or secrets into sandbox container. |
178
+ | Output size | Cap stdout/stderr at 1MB. Truncate with `[output truncated]` marker. |
179
+
180
+ #### Sharp Edges
181
+
182
+ | Failure Mode | Mitigation |
183
+ |---|---|
184
+ | Sandbox escape via Docker vulnerability | Pin Docker version; use rootless Docker; consider gVisor/Firecracker for high-security |
185
+ | Code writes to /tmp exhausting disk | Use `--tmpfs` with size limit (64MB default) |
186
+ | Infinite loop inside sandbox hangs API | Hard timeout with SIGKILL — never rely on SIGTERM alone |
187
+ | Stateful session grows unbounded memory | Limit session history to last 50 executions; reset context on overflow |
@@ -1,146 +1,146 @@
1
- ---
2
- name: "deep-research"
3
- pack: "@rune/ai-ml"
4
- description: "Iterative AI research loop that converges on comprehensive answers — search, analyze, identify gaps, repeat. Outputs synthesized report with source attribution."
5
- model: sonnet
6
- tools: [Read, Edit, Write, Grep, Glob, Bash]
7
- ---
8
-
9
- # deep-research
10
-
11
- Iterative AI research loop that converges on comprehensive answers. Search → analyze → identify gaps → search again. Bounded by depth, time, and URL limits. Outputs synthesized report with source attribution.
12
-
13
- #### Workflow
14
-
15
- **Step 1 — Initialize research state**
16
- ```typescript
17
- interface ResearchState {
18
- query: string;
19
- findings: Finding[]; // max 50 most recent (memory bound)
20
- gaps: string[]; // what we still don't know
21
- seenUrls: Set<string>; // dedup
22
- failedQueries: number; // convergence signal
23
- depth: number; // current iteration
24
- maxDepth: number; // hard limit (default: 10)
25
- maxUrls: number; // hard limit (default: 100)
26
- maxTimeMs: number; // hard limit (default: 300_000 = 5 min)
27
- startedAt: number;
28
- activityLog: ActivityEntry[]; // for progress streaming
29
- }
30
-
31
- interface Finding {
32
- content: string;
33
- sourceUrl: string;
34
- relevance: number; // 0-1
35
- extractedAt: number;
36
- }
37
- ```
38
-
39
- **Step 2 — Generate search queries from current state**
40
- Each iteration, LLM generates 3 search queries based on:
41
- - Original research question
42
- - Current findings (what we know)
43
- - Current gaps (what we don't know)
44
-
45
- ```typescript
46
- const queryPrompt = `Given the research question: "${state.query}"
47
- Current findings: ${summarizeFindings(state.findings)}
48
- Knowledge gaps: ${state.gaps.join(', ')}
49
-
50
- Generate 3 specific search queries that would fill the most important gaps.
51
- Avoid queries similar to: ${state.seenQueries.join(', ')}`;
52
- ```
53
-
54
- **Step 3 — Search and deduplicate**
55
- Execute queries in parallel → collect URLs → filter against `seenUrls` → scrape new URLs → extract relevant content.
56
-
57
- ```typescript
58
- async function searchAndExtract(queries: string[], state: ResearchState): Promise<Finding[]> {
59
- // Parallel search
60
- const allResults = await Promise.all(queries.map(q => webSearch(q, { limit: 10 })));
61
- const urls = deduplicateUrls(allResults.flat(), state.seenUrls);
62
-
63
- // Mark as seen immediately (even before scraping)
64
- for (const url of urls) state.seenUrls.add(url);
65
-
66
- // Scrape and extract in parallel (with concurrency limit)
67
- const findings = await pMap(urls, async (url) => {
68
- const content = await scrapeAndClean(url);
69
- const relevance = await scoreRelevance(content, state.query);
70
- return { content: summarize(content, 500), sourceUrl: url, relevance, extractedAt: Date.now() };
71
- }, { concurrency: 5 });
72
-
73
- return findings.filter(f => f.relevance > 0.3); // threshold
74
- }
75
- ```
76
-
77
- **Step 4 — Analyze findings and detect gaps**
78
- LLM analyzes new findings against existing knowledge:
79
- ```typescript
80
- interface AnalysisResult {
81
- newInsights: string[];
82
- updatedGaps: string[];
83
- shouldContinue: boolean;
84
- nextSearchTopic: string | null;
85
- confidence: number; // 0-1: how complete is our understanding?
86
- }
87
- ```
88
-
89
- **Step 5 — Check convergence criteria**
90
- Stop when ANY of:
91
- - `depth >= maxDepth`
92
- - `seenUrls.size >= maxUrls`
93
- - `Date.now() - startedAt >= maxTimeMs`
94
- - `gaps.length === 0` (all gaps filled)
95
- - `failedQueries >= 3` consecutive (no new information available)
96
- - `confidence >= 0.9` (LLM believes research is comprehensive)
97
-
98
- **Step 6 — Synthesize final report**
99
- ```typescript
100
- interface ResearchReport {
101
- question: string;
102
- answer: string; // comprehensive markdown synthesis
103
- confidence: number;
104
- sources: Array<{
105
- url: string;
106
- title: string;
107
- relevance: number;
108
- citedIn: string[]; // which sections cite this source
109
- }>;
110
- methodology: {
111
- totalIterations: number;
112
- urlsExamined: number;
113
- findingsCount: number;
114
- timeElapsed: number;
115
- remainingGaps: string[];
116
- };
117
- }
118
- ```
119
-
120
- Memory management: keep only 50 most recent findings to avoid context explosion. Summarize older findings into a "background knowledge" string before dropping them.
121
-
122
- #### Example
123
-
124
- ```typescript
125
- // Usage
126
- const report = await deepResearch({
127
- query: 'What are the best practices for implementing RAG in production in 2026?',
128
- maxDepth: 8,
129
- maxUrls: 50,
130
- maxTimeMs: 180_000, // 3 minutes
131
- onProgress: (entry) => console.log(`[${entry.depth}] ${entry.action}: ${entry.detail}`),
132
- });
133
-
134
- // Output: comprehensive report with 15-30 sources, gap analysis, confidence score
135
- ```
136
-
137
- #### Sharp Edges
138
-
139
- | Failure Mode | Mitigation |
140
- |---|---|
141
- | Research loop runs forever (no convergence) | Hard limits on depth, URLs, and time; monitor `failedQueries` counter |
142
- | LLM generates duplicate search queries | Track seen queries; include exclusion list in prompt |
143
- | Memory explosion from accumulating findings | Cap at 50 findings; summarize oldest into background knowledge string |
144
- | Low-quality sources pollute findings | Relevance threshold (0.3); domain blocklist for known low-quality sites |
145
- | Rate limiting on search API | Per-provider rate limiter; fallback to alternative search provider |
146
- | Circular research (keeps finding same information) | Track `confidence` — if stable for 3 iterations, force stop |
1
+ ---
2
+ name: "deep-research"
3
+ pack: "@rune/ai-ml"
4
+ description: "Iterative AI research loop that converges on comprehensive answers — search, analyze, identify gaps, repeat. Outputs synthesized report with source attribution."
5
+ model: sonnet
6
+ tools: [Read, Edit, Write, Grep, Glob, Bash]
7
+ ---
8
+
9
+ # deep-research
10
+
11
+ Iterative AI research loop that converges on comprehensive answers. Search → analyze → identify gaps → search again. Bounded by depth, time, and URL limits. Outputs synthesized report with source attribution.
12
+
13
+ #### Workflow
14
+
15
+ **Step 1 — Initialize research state**
16
+ ```typescript
17
+ interface ResearchState {
18
+ query: string;
19
+ findings: Finding[]; // max 50 most recent (memory bound)
20
+ gaps: string[]; // what we still don't know
21
+ seenUrls: Set<string>; // dedup
22
+ failedQueries: number; // convergence signal
23
+ depth: number; // current iteration
24
+ maxDepth: number; // hard limit (default: 10)
25
+ maxUrls: number; // hard limit (default: 100)
26
+ maxTimeMs: number; // hard limit (default: 300_000 = 5 min)
27
+ startedAt: number;
28
+ activityLog: ActivityEntry[]; // for progress streaming
29
+ }
30
+
31
+ interface Finding {
32
+ content: string;
33
+ sourceUrl: string;
34
+ relevance: number; // 0-1
35
+ extractedAt: number;
36
+ }
37
+ ```
38
+
39
+ **Step 2 — Generate search queries from current state**
40
+ Each iteration, LLM generates 3 search queries based on:
41
+ - Original research question
42
+ - Current findings (what we know)
43
+ - Current gaps (what we don't know)
44
+
45
+ ```typescript
46
+ const queryPrompt = `Given the research question: "${state.query}"
47
+ Current findings: ${summarizeFindings(state.findings)}
48
+ Knowledge gaps: ${state.gaps.join(', ')}
49
+
50
+ Generate 3 specific search queries that would fill the most important gaps.
51
+ Avoid queries similar to: ${state.seenQueries.join(', ')}`;
52
+ ```
53
+
54
+ **Step 3 — Search and deduplicate**
55
+ Execute queries in parallel → collect URLs → filter against `seenUrls` → scrape new URLs → extract relevant content.
56
+
57
+ ```typescript
58
+ async function searchAndExtract(queries: string[], state: ResearchState): Promise<Finding[]> {
59
+ // Parallel search
60
+ const allResults = await Promise.all(queries.map(q => webSearch(q, { limit: 10 })));
61
+ const urls = deduplicateUrls(allResults.flat(), state.seenUrls);
62
+
63
+ // Mark as seen immediately (even before scraping)
64
+ for (const url of urls) state.seenUrls.add(url);
65
+
66
+ // Scrape and extract in parallel (with concurrency limit)
67
+ const findings = await pMap(urls, async (url) => {
68
+ const content = await scrapeAndClean(url);
69
+ const relevance = await scoreRelevance(content, state.query);
70
+ return { content: summarize(content, 500), sourceUrl: url, relevance, extractedAt: Date.now() };
71
+ }, { concurrency: 5 });
72
+
73
+ return findings.filter(f => f.relevance > 0.3); // threshold
74
+ }
75
+ ```
76
+
77
+ **Step 4 — Analyze findings and detect gaps**
78
+ LLM analyzes new findings against existing knowledge:
79
+ ```typescript
80
+ interface AnalysisResult {
81
+ newInsights: string[];
82
+ updatedGaps: string[];
83
+ shouldContinue: boolean;
84
+ nextSearchTopic: string | null;
85
+ confidence: number; // 0-1: how complete is our understanding?
86
+ }
87
+ ```
88
+
89
+ **Step 5 — Check convergence criteria**
90
+ Stop when ANY of:
91
+ - `depth >= maxDepth`
92
+ - `seenUrls.size >= maxUrls`
93
+ - `Date.now() - startedAt >= maxTimeMs`
94
+ - `gaps.length === 0` (all gaps filled)
95
+ - `failedQueries >= 3` consecutive (no new information available)
96
+ - `confidence >= 0.9` (LLM believes research is comprehensive)
97
+
98
+ **Step 6 — Synthesize final report**
99
+ ```typescript
100
+ interface ResearchReport {
101
+ question: string;
102
+ answer: string; // comprehensive markdown synthesis
103
+ confidence: number;
104
+ sources: Array<{
105
+ url: string;
106
+ title: string;
107
+ relevance: number;
108
+ citedIn: string[]; // which sections cite this source
109
+ }>;
110
+ methodology: {
111
+ totalIterations: number;
112
+ urlsExamined: number;
113
+ findingsCount: number;
114
+ timeElapsed: number;
115
+ remainingGaps: string[];
116
+ };
117
+ }
118
+ ```
119
+
120
+ Memory management: keep only 50 most recent findings to avoid context explosion. Summarize older findings into a "background knowledge" string before dropping them.
121
+
122
+ #### Example
123
+
124
+ ```typescript
125
+ // Usage
126
+ const report = await deepResearch({
127
+ query: 'What are the best practices for implementing RAG in production in 2026?',
128
+ maxDepth: 8,
129
+ maxUrls: 50,
130
+ maxTimeMs: 180_000, // 3 minutes
131
+ onProgress: (entry) => console.log(`[${entry.depth}] ${entry.action}: ${entry.detail}`),
132
+ });
133
+
134
+ // Output: comprehensive report with 15-30 sources, gap analysis, confidence score
135
+ ```
136
+
137
+ #### Sharp Edges
138
+
139
+ | Failure Mode | Mitigation |
140
+ |---|---|
141
+ | Research loop runs forever (no convergence) | Hard limits on depth, URLs, and time; monitor `failedQueries` counter |
142
+ | LLM generates duplicate search queries | Track seen queries; include exclusion list in prompt |
143
+ | Memory explosion from accumulating findings | Cap at 50 findings; summarize oldest into background knowledge string |
144
+ | Low-quality sources pollute findings | Relevance threshold (0.3); domain blocklist for known low-quality sites |
145
+ | Rate limiting on search API | Per-provider rate limiter; fallback to alternative search provider |
146
+ | Circular research (keeps finding same information) | Track `confidence` — if stable for 3 iterations, force stop |