@wardby/cli 0.2.1 → 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (555) hide show
  1. package/.env.example +53 -4
  2. package/README.md +52 -13
  3. package/dist/claude-coding-worker/driver.d.ts +6 -1
  4. package/dist/claude-coding-worker/driver.js +27 -1
  5. package/dist/claude-coding-worker/main.js +2 -0
  6. package/dist/claude-coding-worker/tool-socket.d.ts +9 -0
  7. package/dist/claude-coding-worker/tool-socket.js +26 -0
  8. package/dist/cli-help.d.ts +1 -1
  9. package/dist/cli-help.js +14 -4
  10. package/dist/cli.d.ts +1 -1
  11. package/dist/cli.js +284 -36
  12. package/dist/coding/base-commit.d.ts +6 -0
  13. package/dist/coding/base-commit.js +12 -0
  14. package/dist/coding/collect-exclude.d.ts +25 -0
  15. package/dist/coding/collect-exclude.js +76 -0
  16. package/dist/coding/profile.d.ts +67 -22
  17. package/dist/coding/profile.js +58 -25
  18. package/dist/coding/protected-path-wording.d.ts +24 -0
  19. package/dist/coding/protected-path-wording.js +55 -0
  20. package/dist/coding/protected-paths.d.ts +31 -0
  21. package/dist/coding/protected-paths.js +48 -0
  22. package/dist/coding/protocol.d.ts +91 -4
  23. package/dist/coding/protocol.js +123 -12
  24. package/dist/coding/provider.d.ts +6 -1
  25. package/dist/coding/provider.js +16 -9
  26. package/dist/coding/registry/adapters.d.ts +2 -0
  27. package/dist/coding/registry/adapters.js +9 -0
  28. package/dist/coding/registry/allowlist.d.ts +9 -0
  29. package/dist/coding/registry/allowlist.js +28 -0
  30. package/dist/coding/registry/json-scan.d.ts +55 -0
  31. package/dist/coding/registry/json-scan.js +191 -0
  32. package/dist/coding/registry/lockfiles.d.ts +6 -0
  33. package/dist/coding/registry/lockfiles.js +89 -0
  34. package/dist/coding/registry/npm-lockfile.d.ts +9 -0
  35. package/dist/coding/registry/npm-lockfile.js +73 -0
  36. package/dist/coding/registry/npm-plan.d.ts +29 -0
  37. package/dist/coding/registry/npm-plan.js +289 -0
  38. package/dist/coding/registry/npm.d.ts +7 -0
  39. package/dist/coding/registry/npm.js +252 -0
  40. package/dist/coding/registry/pypi.d.ts +4 -0
  41. package/dist/coding/registry/pypi.js +210 -0
  42. package/dist/coding/registry/report.d.ts +37 -0
  43. package/dist/coding/registry/report.js +31 -0
  44. package/dist/coding/registry/token.d.ts +4 -0
  45. package/dist/coding/registry/token.js +7 -0
  46. package/dist/coding/registry/types.d.ts +286 -0
  47. package/dist/coding/registry/types.js +19 -0
  48. package/dist/coding/registry/worker-config.d.ts +14 -0
  49. package/dist/coding/registry/worker-config.js +31 -0
  50. package/dist/coding/services/builtins.d.ts +22 -0
  51. package/dist/coding/services/builtins.js +85 -0
  52. package/dist/coding/services/catalog.d.ts +443 -0
  53. package/dist/coding/services/catalog.js +159 -0
  54. package/dist/coding/services/declaration.d.ts +11 -0
  55. package/dist/coding/services/declaration.js +114 -0
  56. package/dist/coding/services/note.d.ts +8 -0
  57. package/dist/coding/services/note.js +15 -0
  58. package/dist/coding/services/resolve.d.ts +25 -0
  59. package/dist/coding/services/resolve.js +50 -0
  60. package/dist/coding/services/wording.d.ts +25 -0
  61. package/dist/coding/services/wording.js +70 -0
  62. package/dist/coding-proxy/main.js +24 -4
  63. package/dist/coding-worker/artifact.d.ts +6 -0
  64. package/dist/coding-worker/debug-trace.d.ts +37 -0
  65. package/dist/coding-worker/debug-trace.js +117 -0
  66. package/dist/coding-worker/driver.d.ts +13 -1
  67. package/dist/coding-worker/driver.js +102 -13
  68. package/dist/coding-worker/errors.js +8 -0
  69. package/dist/coding-worker/main.js +10 -2
  70. package/dist/coding-worker/sdk.d.ts +2 -2
  71. package/dist/coding-worker/sdk.js +7 -2
  72. package/dist/coding-worker/types.d.ts +3 -0
  73. package/dist/config/providers.d.ts +66 -0
  74. package/dist/config/providers.js +137 -0
  75. package/dist/core/attribution.d.ts +101 -0
  76. package/dist/core/attribution.js +208 -0
  77. package/dist/core/budget-groups.d.ts +120 -10
  78. package/dist/core/budget-groups.js +133 -24
  79. package/dist/core/budget-wording.d.ts +22 -0
  80. package/dist/core/budget-wording.js +55 -0
  81. package/dist/core/coding-queue.d.ts +3 -0
  82. package/dist/core/coding-queue.js +6 -2
  83. package/dist/core/coding-service-status.d.ts +11 -0
  84. package/dist/core/coding-service-status.js +17 -0
  85. package/dist/core/cost-report.d.ts +88 -0
  86. package/dist/core/cost-report.js +248 -0
  87. package/dist/core/datastores.js +18 -2
  88. package/dist/core/db.d.ts +5 -1
  89. package/dist/core/db.js +8 -2
  90. package/dist/core/dispatch.d.ts +74 -7
  91. package/dist/core/dispatch.js +382 -47
  92. package/dist/core/engine-native.js +44 -14
  93. package/dist/core/glob.d.ts +10 -0
  94. package/dist/core/glob.js +33 -0
  95. package/dist/core/grants.d.ts +86 -0
  96. package/dist/core/grants.js +126 -0
  97. package/dist/core/host-events.d.ts +72 -0
  98. package/dist/core/host-events.js +597 -0
  99. package/dist/core/host-identity-links.d.ts +54 -0
  100. package/dist/core/host-identity-links.js +189 -0
  101. package/dist/core/host-status.d.ts +78 -0
  102. package/dist/core/host-status.js +228 -0
  103. package/dist/core/in-flight-runs.d.ts +8 -0
  104. package/dist/core/in-flight-runs.js +56 -0
  105. package/dist/core/issue-bridge.d.ts +60 -0
  106. package/dist/core/issue-bridge.js +189 -0
  107. package/dist/core/issue-dedupe.d.ts +70 -0
  108. package/dist/core/issue-dedupe.js +255 -0
  109. package/dist/core/issue-events.d.ts +42 -0
  110. package/dist/core/issue-events.js +155 -0
  111. package/dist/core/issue-status.d.ts +29 -0
  112. package/dist/core/issue-status.js +241 -0
  113. package/dist/core/issue-tracker-tools.d.ts +64 -0
  114. package/dist/core/issue-tracker-tools.js +850 -0
  115. package/dist/core/model-usage.d.ts +10 -0
  116. package/dist/core/model-usage.js +24 -0
  117. package/dist/core/provider-wording.d.ts +12 -0
  118. package/dist/core/provider-wording.js +44 -0
  119. package/dist/core/reconciler.d.ts +42 -2
  120. package/dist/core/reconciler.js +97 -2
  121. package/dist/core/repo-access.d.ts +99 -0
  122. package/dist/core/repo-access.js +136 -0
  123. package/dist/core/review-host-checks.d.ts +11 -0
  124. package/dist/core/review-host-checks.js +41 -0
  125. package/dist/core/review-host-tools.d.ts +47 -0
  126. package/dist/core/review-host-tools.js +354 -0
  127. package/dist/core/run-heartbeat.d.ts +27 -0
  128. package/dist/core/run-heartbeat.js +54 -0
  129. package/dist/core/run-pricing.d.ts +61 -0
  130. package/dist/core/run-pricing.js +56 -0
  131. package/dist/core/runner.d.ts +33 -8
  132. package/dist/core/runner.js +410 -62
  133. package/dist/core/scheduler.d.ts +4 -1
  134. package/dist/core/scheduler.js +3 -2
  135. package/dist/core/secrets.d.ts +11 -2
  136. package/dist/core/secrets.js +25 -6
  137. package/dist/core/self-defects.d.ts +80 -0
  138. package/dist/core/self-defects.js +180 -0
  139. package/dist/core/subagent-memory-tools.d.ts +1 -1
  140. package/dist/core/subagent-memory-tools.js +16 -3
  141. package/dist/core/tool-admin.d.ts +81 -0
  142. package/dist/core/tool-admin.js +129 -0
  143. package/dist/core/tool-names.d.ts +42 -0
  144. package/dist/core/tool-names.js +67 -0
  145. package/dist/core/untrusted-content.d.ts +32 -0
  146. package/dist/core/untrusted-content.js +72 -0
  147. package/dist/core/webhooks.d.ts +17 -2
  148. package/dist/core/webhooks.js +42 -3
  149. package/dist/generated/prisma/browser.d.ts +166 -0
  150. package/dist/generated/prisma/client.d.ts +166 -0
  151. package/dist/generated/prisma/commonInputTypes.d.ts +152 -52
  152. package/dist/generated/prisma/enums.d.ts +13 -0
  153. package/dist/generated/prisma/enums.js +12 -1
  154. package/dist/generated/prisma/internal/class.d.ts +242 -0
  155. package/dist/generated/prisma/internal/class.js +4 -4
  156. package/dist/generated/prisma/internal/prismaNamespace.d.ts +2571 -577
  157. package/dist/generated/prisma/internal/prismaNamespace.js +312 -6
  158. package/dist/generated/prisma/internal/prismaNamespaceBrowser.d.ts +328 -0
  159. package/dist/generated/prisma/internal/prismaNamespaceBrowser.js +312 -6
  160. package/dist/generated/prisma/models/Agent.d.ts +682 -1
  161. package/dist/generated/prisma/models/AgentIssueProject.d.ts +1838 -0
  162. package/dist/generated/prisma/models/AgentIssueProject.js +1 -0
  163. package/dist/generated/prisma/models/AgentRepository.d.ts +1425 -0
  164. package/dist/generated/prisma/models/AgentRepository.js +1 -0
  165. package/dist/generated/prisma/models/AgentTool.d.ts +95 -1
  166. package/dist/generated/prisma/models/AuthUser.d.ts +56 -1
  167. package/dist/generated/prisma/models/CodingAgentProfile.d.ts +258 -7
  168. package/dist/generated/prisma/models/CodingProxySession.d.ts +195 -2
  169. package/dist/generated/prisma/models/CodingRun.d.ts +1405 -95
  170. package/dist/generated/prisma/models/CodingRunServiceStatus.d.ts +1404 -0
  171. package/dist/generated/prisma/models/CodingRunServiceStatus.js +1 -0
  172. package/dist/generated/prisma/models/CodingService.d.ts +1348 -0
  173. package/dist/generated/prisma/models/CodingService.js +1 -0
  174. package/dist/generated/prisma/models/HostEventDelivery.d.ts +946 -0
  175. package/dist/generated/prisma/models/HostEventDelivery.js +1 -0
  176. package/dist/generated/prisma/models/HostIdentity.d.ts +1232 -0
  177. package/dist/generated/prisma/models/HostIdentity.js +1 -0
  178. package/dist/generated/prisma/models/HostIdentityLinkRequest.d.ts +1473 -0
  179. package/dist/generated/prisma/models/HostIdentityLinkRequest.js +1 -0
  180. package/dist/generated/prisma/models/IssueFingerprint.d.ts +1183 -0
  181. package/dist/generated/prisma/models/IssueFingerprint.js +1 -0
  182. package/dist/generated/prisma/models/IssuePullRequest.d.ts +1255 -0
  183. package/dist/generated/prisma/models/IssuePullRequest.js +1 -0
  184. package/dist/generated/prisma/models/ModelCatalogEntry.d.ts +1322 -0
  185. package/dist/generated/prisma/models/ModelCatalogEntry.js +1 -0
  186. package/dist/generated/prisma/models/Principal.d.ts +455 -0
  187. package/dist/generated/prisma/models/RegistryAllowance.d.ts +1148 -0
  188. package/dist/generated/prisma/models/RegistryAllowance.js +1 -0
  189. package/dist/generated/prisma/models/RegistryApprovedVersion.d.ts +1219 -0
  190. package/dist/generated/prisma/models/RegistryApprovedVersion.js +1 -0
  191. package/dist/generated/prisma/models/RegistryFetch.d.ts +1428 -0
  192. package/dist/generated/prisma/models/RegistryFetch.js +1 -0
  193. package/dist/generated/prisma/models/RegistryPlanRefusal.d.ts +1294 -0
  194. package/dist/generated/prisma/models/RegistryPlanRefusal.js +1 -0
  195. package/dist/generated/prisma/models/RegistryVersionFact.d.ts +1085 -0
  196. package/dist/generated/prisma/models/RegistryVersionFact.js +1 -0
  197. package/dist/generated/prisma/models/ResourceGrant.d.ts +1437 -0
  198. package/dist/generated/prisma/models/ResourceGrant.js +1 -0
  199. package/dist/generated/prisma/models/Run.d.ts +1402 -82
  200. package/dist/generated/prisma/models/RunAttribution.d.ts +1259 -0
  201. package/dist/generated/prisma/models/RunAttribution.js +1 -0
  202. package/dist/generated/prisma/models/RunHostCheck.d.ts +1239 -0
  203. package/dist/generated/prisma/models/RunHostCheck.js +1 -0
  204. package/dist/generated/prisma/models/RunHostStatus.d.ts +1315 -0
  205. package/dist/generated/prisma/models/RunHostStatus.js +1 -0
  206. package/dist/generated/prisma/models/RunIssueStatus.d.ts +1199 -0
  207. package/dist/generated/prisma/models/RunIssueStatus.js +1 -0
  208. package/dist/generated/prisma/models/RunModelUsage.d.ts +1316 -0
  209. package/dist/generated/prisma/models/RunModelUsage.js +1 -0
  210. package/dist/generated/prisma/models/Tool.d.ts +15 -3
  211. package/dist/generated/prisma/models/WorkItem.d.ts +1408 -0
  212. package/dist/generated/prisma/models/WorkItem.js +1 -0
  213. package/dist/generated/prisma/models.d.ts +22 -0
  214. package/dist/help/build.d.ts +1 -0
  215. package/dist/help/build.js +9 -0
  216. package/dist/help/catalog.d.ts +24 -0
  217. package/dist/help/catalog.js +160 -0
  218. package/dist/help/cli.d.ts +2 -0
  219. package/dist/help/cli.js +65 -0
  220. package/dist/help/runtime.d.ts +3 -0
  221. package/dist/help/runtime.js +44 -0
  222. package/dist/help/search.d.ts +9 -0
  223. package/dist/help/search.js +104 -0
  224. package/dist/help-index.json +999 -0
  225. package/dist/import/cli-args.js +3 -2
  226. package/dist/import/create.d.ts +5 -0
  227. package/dist/import/create.js +67 -13
  228. package/dist/import/index.js +19 -9
  229. package/dist/import/neutral-schema.d.ts +18 -18
  230. package/dist/knowledge/check.d.ts +13 -0
  231. package/dist/knowledge/check.js +69 -0
  232. package/dist/knowledge/cli.d.ts +14 -0
  233. package/dist/knowledge/cli.js +67 -0
  234. package/dist/knowledge/concept.d.ts +54 -0
  235. package/dist/knowledge/concept.js +78 -0
  236. package/dist/knowledge/note.d.ts +11 -0
  237. package/dist/knowledge/note.js +39 -0
  238. package/dist/knowledge/relevance.d.ts +11 -0
  239. package/dist/knowledge/relevance.js +14 -0
  240. package/dist/knowledge/span-hash.d.ts +3 -0
  241. package/dist/knowledge/span-hash.js +16 -0
  242. package/dist/mcp/auth/access.d.ts +72 -0
  243. package/dist/mcp/auth/access.js +58 -0
  244. package/dist/mcp/auth/grants-cli.d.ts +149 -0
  245. package/dist/mcp/auth/grants-cli.js +518 -0
  246. package/dist/mcp/auth/host-account-cli.d.ts +2 -0
  247. package/dist/mcp/auth/host-account-cli.js +47 -0
  248. package/dist/mcp/auth/ownership.d.ts +25 -51
  249. package/dist/mcp/auth/ownership.js +19 -14
  250. package/dist/mcp/auth/repo-authorization.d.ts +22 -0
  251. package/dist/mcp/auth/repo-authorization.js +51 -0
  252. package/dist/mcp/auth/resource-server.d.ts +33 -2
  253. package/dist/mcp/auth/resource-server.js +99 -3
  254. package/dist/mcp/auth/self-hosted/browser.js +2 -2
  255. package/dist/mcp/auth/self-hosted/cli.js +47 -13
  256. package/dist/mcp/auth/self-hosted/credentials.d.ts +23 -5
  257. package/dist/mcp/auth/self-hosted/credentials.js +62 -4
  258. package/dist/mcp/auth/self-hosted/session.d.ts +7 -6
  259. package/dist/mcp/context.d.ts +30 -1
  260. package/dist/mcp/errors.d.ts +25 -6
  261. package/dist/mcp/errors.js +98 -0
  262. package/dist/mcp/host-events/deliveries.d.ts +9 -0
  263. package/dist/mcp/host-events/deliveries.js +17 -0
  264. package/dist/mcp/host-events/github-ingress.d.ts +38 -0
  265. package/dist/mcp/host-events/github-ingress.js +83 -0
  266. package/dist/mcp/host-events/github-user-callback.d.ts +18 -0
  267. package/dist/mcp/host-events/github-user-callback.js +50 -0
  268. package/dist/mcp/host-events/jira-ingress.d.ts +29 -0
  269. package/dist/mcp/host-events/jira-ingress.js +92 -0
  270. package/dist/mcp/index.d.ts +10 -1
  271. package/dist/mcp/index.js +187 -19
  272. package/dist/mcp/server.js +32 -11
  273. package/dist/mcp/tools/agents.js +519 -50
  274. package/dist/mcp/tools/budget-groups.js +3 -3
  275. package/dist/mcp/tools/cost-report.d.ts +8 -0
  276. package/dist/mcp/tools/cost-report.js +60 -0
  277. package/dist/mcp/tools/datastore.js +22 -15
  278. package/dist/mcp/tools/grants.d.ts +2 -0
  279. package/dist/mcp/tools/grants.js +239 -0
  280. package/dist/mcp/tools/help.d.ts +5 -0
  281. package/dist/mcp/tools/help.js +67 -0
  282. package/dist/mcp/tools/host-accounts.d.ts +2 -0
  283. package/dist/mcp/tools/host-accounts.js +114 -0
  284. package/dist/mcp/tools/issue-projects.d.ts +2 -0
  285. package/dist/mcp/tools/issue-projects.js +238 -0
  286. package/dist/mcp/tools/memory.d.ts +7 -1
  287. package/dist/mcp/tools/memory.js +5 -5
  288. package/dist/mcp/tools/model-catalog.d.ts +22 -0
  289. package/dist/mcp/tools/model-catalog.js +423 -0
  290. package/dist/mcp/tools/repositories.d.ts +2 -0
  291. package/dist/mcp/tools/repositories.js +182 -0
  292. package/dist/mcp/tools/runs.d.ts +6 -0
  293. package/dist/mcp/tools/runs.js +60 -6
  294. package/dist/mcp/tools/scheduling.js +3 -6
  295. package/dist/mcp/tools/secrets.js +18 -6
  296. package/dist/mcp/tools/services.d.ts +2 -0
  297. package/dist/mcp/tools/services.js +222 -0
  298. package/dist/mcp/tools/subagents.js +52 -16
  299. package/dist/mcp/tools/tools.d.ts +41 -0
  300. package/dist/mcp/tools/tools.js +201 -40
  301. package/dist/mcp/tools/trigger.js +65 -7
  302. package/dist/mcp/tools/webhooks.js +15 -4
  303. package/dist/mcp/transport/streamable-http.d.ts +15 -0
  304. package/dist/mcp/transport/streamable-http.js +55 -3
  305. package/dist/mcp/webhooks/ingress.d.ts +2 -1
  306. package/dist/mcp/webhooks/ingress.js +9 -2
  307. package/dist/providers/auth/delegating.d.ts +19 -0
  308. package/dist/providers/auth/delegating.js +71 -0
  309. package/dist/providers/auth/index.d.ts +2 -0
  310. package/dist/providers/auth/index.js +11 -0
  311. package/dist/providers/auth/self-hosted.d.ts +10 -2
  312. package/dist/providers/auth/self-hosted.js +55 -6
  313. package/dist/providers/auth/types.d.ts +6 -0
  314. package/dist/providers/coding-proxy/memory-ledger.d.ts +4 -1
  315. package/dist/providers/coding-proxy/memory-ledger.js +31 -3
  316. package/dist/providers/coding-proxy/metering.d.ts +6 -1
  317. package/dist/providers/coding-proxy/metering.js +41 -4
  318. package/dist/providers/coding-proxy/prisma-ledger.d.ts +17 -0
  319. package/dist/providers/coding-proxy/prisma-ledger.js +106 -8
  320. package/dist/providers/coding-proxy/proxy.d.ts +25 -2
  321. package/dist/providers/coding-proxy/proxy.js +561 -38
  322. package/dist/providers/coding-proxy/registry/audit.d.ts +131 -0
  323. package/dist/providers/coding-proxy/registry/audit.js +380 -0
  324. package/dist/providers/coding-proxy/registry/bounded-fetch.d.ts +16 -0
  325. package/dist/providers/coding-proxy/registry/bounded-fetch.js +77 -0
  326. package/dist/providers/coding-proxy/registry/plan.d.ts +58 -0
  327. package/dist/providers/coding-proxy/registry/plan.js +304 -0
  328. package/dist/providers/coding-proxy/registry/prisma-store.d.ts +34 -0
  329. package/dist/providers/coding-proxy/registry/prisma-store.js +151 -0
  330. package/dist/providers/coding-proxy/registry/service.d.ts +213 -0
  331. package/dist/providers/coding-proxy/registry/service.js +1137 -0
  332. package/dist/providers/coding-proxy/registry/store.d.ts +127 -0
  333. package/dist/providers/coding-proxy/registry/store.js +70 -0
  334. package/dist/providers/coding-proxy/runtime.d.ts +17 -0
  335. package/dist/providers/coding-proxy/runtime.js +63 -0
  336. package/dist/providers/coding-proxy/secure-fetch.js +0 -1
  337. package/dist/providers/coding-proxy/server.d.ts +11 -0
  338. package/dist/providers/coding-proxy/server.js +163 -0
  339. package/dist/providers/coding-proxy/types.d.ts +33 -1
  340. package/dist/providers/coding-proxy/types.js +12 -1
  341. package/dist/providers/engine/types.d.ts +29 -0
  342. package/dist/providers/executor/build.d.ts +2 -2
  343. package/dist/providers/executor/composition.d.ts +3 -0
  344. package/dist/providers/executor/composition.js +28 -1
  345. package/dist/providers/executor/container.d.ts +134 -4
  346. package/dist/providers/executor/container.js +394 -37
  347. package/dist/providers/executor/dbos.d.ts +4 -3
  348. package/dist/providers/executor/dbos.js +7 -5
  349. package/dist/providers/executor/in-process.d.ts +2 -3
  350. package/dist/providers/executor/routing.d.ts +13 -0
  351. package/dist/providers/executor/routing.js +18 -0
  352. package/dist/providers/executor/types.d.ts +34 -0
  353. package/dist/providers/issue-tracker/adf.d.ts +31 -0
  354. package/dist/providers/issue-tracker/adf.js +181 -0
  355. package/dist/providers/issue-tracker/index.d.ts +5 -0
  356. package/dist/providers/issue-tracker/index.js +12 -0
  357. package/dist/providers/issue-tracker/jira-client.d.ts +41 -0
  358. package/dist/providers/issue-tracker/jira-client.js +151 -0
  359. package/dist/providers/issue-tracker/jira-events.d.ts +3 -0
  360. package/dist/providers/issue-tracker/jira-events.js +98 -0
  361. package/dist/providers/issue-tracker/jira.d.ts +116 -0
  362. package/dist/providers/issue-tracker/jira.js +502 -0
  363. package/dist/providers/issue-tracker/types.d.ts +269 -0
  364. package/dist/providers/issue-tracker/types.js +16 -0
  365. package/dist/providers/jobs/claude-tool-setup.d.ts +23 -0
  366. package/dist/providers/jobs/claude-tool-setup.js +50 -0
  367. package/dist/providers/jobs/collect-prune.d.ts +7 -0
  368. package/dist/providers/jobs/collect-prune.js +27 -0
  369. package/dist/providers/jobs/docker-isolation.d.ts +48 -3
  370. package/dist/providers/jobs/docker-isolation.js +213 -22
  371. package/dist/providers/jobs/docker-services.d.ts +35 -0
  372. package/dist/providers/jobs/docker-services.js +191 -0
  373. package/dist/providers/jobs/docker.d.ts +53 -2
  374. package/dist/providers/jobs/docker.js +307 -32
  375. package/dist/providers/jobs/fake-kubernetes-api.d.ts +1 -0
  376. package/dist/providers/jobs/fake-kubernetes-api.js +12 -3
  377. package/dist/providers/jobs/kubernetes-isolation.d.ts +17 -2
  378. package/dist/providers/jobs/kubernetes-isolation.js +238 -57
  379. package/dist/providers/jobs/kubernetes-platform.d.ts +10 -4
  380. package/dist/providers/jobs/kubernetes-platform.js +11 -5
  381. package/dist/providers/jobs/kubernetes-preflight.js +3 -0
  382. package/dist/providers/jobs/kubernetes.d.ts +24 -2
  383. package/dist/providers/jobs/kubernetes.js +153 -22
  384. package/dist/providers/jobs/service-state.d.ts +22 -0
  385. package/dist/providers/jobs/service-state.js +17 -0
  386. package/dist/providers/jobs/types.d.ts +18 -0
  387. package/dist/providers/llm/anthropic.d.ts +3 -3
  388. package/dist/providers/llm/anthropic.js +3 -5
  389. package/dist/providers/llm/bedrock.d.ts +3 -3
  390. package/dist/providers/llm/bedrock.js +3 -8
  391. package/dist/providers/llm/catalog-lookup.d.ts +10 -0
  392. package/dist/providers/llm/catalog-lookup.js +15 -0
  393. package/dist/providers/llm/catalog-shipped.d.ts +18 -0
  394. package/dist/providers/llm/catalog-shipped.js +197 -0
  395. package/dist/providers/llm/catalog-store.d.ts +58 -0
  396. package/dist/providers/llm/catalog-store.js +138 -0
  397. package/dist/providers/llm/catalog-types.d.ts +66 -0
  398. package/dist/providers/llm/catalog-types.js +64 -0
  399. package/dist/providers/llm/catalog.d.ts +61 -0
  400. package/dist/providers/llm/catalog.js +147 -0
  401. package/dist/providers/llm/claude-messages.d.ts +5 -1
  402. package/dist/providers/llm/claude-messages.js +1 -0
  403. package/dist/providers/llm/claude-provider.d.ts +10 -12
  404. package/dist/providers/llm/claude-provider.js +15 -6
  405. package/dist/providers/llm/index.d.ts +9 -6
  406. package/dist/providers/llm/index.js +8 -5
  407. package/dist/providers/llm/openai.d.ts +14 -5
  408. package/dist/providers/llm/openai.js +24 -14
  409. package/dist/providers/llm/pricing-core.d.ts +5 -3
  410. package/dist/providers/llm/registration.js +8 -12
  411. package/dist/providers/llm/routing.d.ts +21 -11
  412. package/dist/providers/llm/routing.js +47 -13
  413. package/dist/providers/llm/types.d.ts +12 -0
  414. package/dist/providers/llm/types.js +8 -1
  415. package/dist/providers/review-host/diff-lines.d.ts +16 -0
  416. package/dist/providers/review-host/diff-lines.js +59 -0
  417. package/dist/providers/review-host/github-events.d.ts +6 -0
  418. package/dist/providers/review-host/github-events.js +226 -0
  419. package/dist/providers/review-host/github-user-auth.d.ts +38 -0
  420. package/dist/providers/review-host/github-user-auth.js +128 -0
  421. package/dist/providers/review-host/github.d.ts +57 -0
  422. package/dist/providers/review-host/github.js +568 -0
  423. package/dist/providers/review-host/index.d.ts +12 -0
  424. package/dist/providers/review-host/index.js +30 -0
  425. package/dist/providers/review-host/review-format.d.ts +19 -0
  426. package/dist/providers/review-host/review-format.js +59 -0
  427. package/dist/providers/review-host/types.d.ts +309 -0
  428. package/dist/providers/review-host/types.js +24 -0
  429. package/dist/providers/vcs/git.d.ts +16 -6
  430. package/dist/providers/vcs/git.js +78 -34
  431. package/dist/providers/vcs/github.d.ts +94 -3
  432. package/dist/providers/vcs/github.js +223 -17
  433. package/dist/providers/vcs/types.d.ts +48 -4
  434. package/dist/quickstart/index.d.ts +2 -0
  435. package/dist/quickstart/index.js +36 -8
  436. package/dist/sandbox/fetch-policy.d.ts +20 -2
  437. package/dist/sandbox/fetch-policy.js +64 -3
  438. package/dist/sandbox/host-functions.d.ts +9 -1
  439. package/dist/sandbox/host-functions.js +12 -5
  440. package/dist/serve.js +8 -2
  441. package/dist/viewer/api-schema.d.ts +2757 -0
  442. package/dist/viewer/api-schema.js +165 -0
  443. package/dist/viewer/build-schemas.d.ts +2 -0
  444. package/dist/viewer/build-schemas.js +18 -0
  445. package/dist/viewer/event-bus.d.ts +38 -0
  446. package/dist/viewer/event-bus.js +232 -0
  447. package/dist/viewer/graph.d.ts +40 -0
  448. package/dist/viewer/graph.js +243 -0
  449. package/dist/viewer/http.d.ts +30 -0
  450. package/dist/viewer/http.js +133 -0
  451. package/dist/viewer/run-detail.d.ts +4 -0
  452. package/dist/viewer/run-detail.js +61 -0
  453. package/dist/wardby-bin.js +11 -0
  454. package/docs/README.md +40 -0
  455. package/docs/agent-recipes.md +383 -0
  456. package/docs/architecture-runtime.md +90 -0
  457. package/docs/assets/brand/wardby-icon-512.png +0 -0
  458. package/docs/assets/brand/wardby-icon.svg +16 -0
  459. package/docs/assets/brand/wardby-mascot-profile-512.png +0 -0
  460. package/docs/assets/brand/wardby-mascot.png +0 -0
  461. package/docs/assets/brand/wardby-mascot.svg +5 -0
  462. package/docs/assets/wardby-workflow.svg +106 -0
  463. package/docs/code-review-agents.md +508 -0
  464. package/docs/coding-agent-setup.md +175 -0
  465. package/docs/coding-packages.md +455 -0
  466. package/docs/coding-services.md +300 -0
  467. package/docs/coding-worker-byo-images.md +98 -0
  468. package/docs/coding-worker-isolation.md +1068 -0
  469. package/docs/getting-started-gke.md +632 -0
  470. package/docs/getting-started-identity-provider.md +319 -0
  471. package/docs/getting-started.md +147 -0
  472. package/docs/jira-agents.md +649 -0
  473. package/docs/knowledge.md +387 -0
  474. package/docs/models.md +221 -0
  475. package/docs/observability.md +53 -0
  476. package/docs/release-verification.md +66 -0
  477. package/docs/security-deployment.md +667 -0
  478. package/docs/viewer-api.md +142 -0
  479. package/help/admin-viewer.md +39 -0
  480. package/help/agent-recipes.md +173 -0
  481. package/help/architecture-agent.md +189 -0
  482. package/help/builder-agent.md +80 -0
  483. package/help/code-review-agents.md +37 -0
  484. package/help/coding-packages.md +31 -0
  485. package/help/coding-services.md +71 -0
  486. package/help/cost-attribution.md +67 -0
  487. package/help/creating-agents.md +90 -0
  488. package/help/deploy-gke.md +45 -0
  489. package/help/deployment-targets.md +39 -0
  490. package/help/errors/budget-group-exhausted.md +25 -0
  491. package/help/errors/docker-isolation-unsupported.md +26 -0
  492. package/help/errors/model-unavailable.md +63 -0
  493. package/help/errors/protected-path.md +45 -0
  494. package/help/errors/repo-access.md +27 -0
  495. package/help/errors/service-declaration-invalid.md +27 -0
  496. package/help/errors/service-declaration-unavailable.md +25 -0
  497. package/help/errors/service-launcher-unsupported.md +30 -0
  498. package/help/errors/service-not-allowed.md +27 -0
  499. package/help/errors/service-unknown.md +23 -0
  500. package/help/errors/service-unready.md +49 -0
  501. package/help/getting-started.md +32 -0
  502. package/help/github.md +48 -0
  503. package/help/identity-and-access.md +44 -0
  504. package/help/jira.md +135 -0
  505. package/help/knowledge.md +47 -0
  506. package/help/mcp.md +30 -0
  507. package/help/models.md +90 -0
  508. package/help/native-capabilities.md +30 -0
  509. package/help/observability.md +41 -0
  510. package/help/operating-agents.md +36 -0
  511. package/help/security.md +27 -0
  512. package/help/troubleshooting/budgets.md +32 -0
  513. package/help/troubleshooting/coding-workers.md +47 -0
  514. package/help/troubleshooting/repository-access.md +27 -0
  515. package/package.json +14 -3
  516. package/prisma/migrations/20260925010000_coding_collect_exclude/migration.sql +8 -0
  517. package/prisma/migrations/20260925015000_allowed_egress_default/migration.sql +8 -0
  518. package/prisma/migrations/20260925020000_coding_package_registry/migration.sql +55 -0
  519. package/prisma/migrations/20260925030000_tool_name_per_owner/migration.sql +11 -0
  520. package/prisma/migrations/20260926010000_run_trigger_host_event/migration.sql +7 -0
  521. package/prisma/migrations/20260926020000_code_review_hosts/migration.sql +53 -0
  522. package/prisma/migrations/20260926030000_auth_user_roles/migration.sql +9 -0
  523. package/prisma/migrations/20260926040000_agent_effort/migration.sql +6 -0
  524. package/prisma/migrations/20260926050000_registry_lockfile_plan/migration.sql +32 -0
  525. package/prisma/migrations/20260926050000_repo_access_authorization/migration.sql +82 -0
  526. package/prisma/migrations/20260926060000_registry_plan_refusal/migration.sql +19 -0
  527. package/prisma/migrations/20260926100000_registry_plan_refusal_published_at/migration.sql +5 -0
  528. package/prisma/migrations/20260926190000_run_host_status/migration.sql +18 -0
  529. package/prisma/migrations/20260926210000_run_host_status_at_dispatch/migration.sql +5 -0
  530. package/prisma/migrations/20260927010000_resource_grants/migration.sql +72 -0
  531. package/prisma/migrations/20260927020000_proxy_session_budget_exhausted/migration.sql +3 -0
  532. package/prisma/migrations/20260927030000_coding_debug_trace/migration.sql +4 -0
  533. package/prisma/migrations/20260927040000_proxy_session_upstream_failure/migration.sql +3 -0
  534. package/prisma/migrations/20260927050000_coding_run_services/migration.sql +74 -0
  535. package/prisma/migrations/20260928000000_coding_run_tool_image/migration.sql +5 -0
  536. package/prisma/migrations/20260930000000_jira_issue_projects/migration.sql +34 -0
  537. package/prisma/migrations/20261001000000_jira_phase2_allowlists/migration.sql +3 -0
  538. package/prisma/migrations/20261001010000_jira_link_types_allowlist/migration.sql +2 -0
  539. package/prisma/migrations/20261002000000_jira_coding_bridge/migration.sql +28 -0
  540. package/prisma/migrations/20261002010000_jira_issue_creation/migration.sql +25 -0
  541. package/prisma/migrations/20261003000000_issue_cost_attribution/migration.sql +56 -0
  542. package/prisma/migrations/20261003010000_coding_run_service_status/migration.sql +23 -0
  543. package/prisma/migrations/20261003020000_viewer_notify/migration.sql +54 -0
  544. package/prisma/migrations/20261003030000_viewer_notify_fixes/migration.sql +47 -0
  545. package/prisma/migrations/20261003040000_viewer_indexes/migration.sql +12 -0
  546. package/prisma/migrations/20261004000000_model_catalog/migration.sql +26 -0
  547. package/prisma/schema.prisma +663 -21
  548. package/dist/mcp/tools/models.d.ts +0 -8
  549. package/dist/mcp/tools/models.js +0 -15
  550. package/dist/providers/llm/pricing-anthropic.d.ts +0 -4
  551. package/dist/providers/llm/pricing-anthropic.js +0 -30
  552. package/dist/providers/llm/pricing-bedrock-claude.d.ts +0 -12
  553. package/dist/providers/llm/pricing-bedrock-claude.js +0 -37
  554. package/dist/providers/llm/pricing.d.ts +0 -30
  555. package/dist/providers/llm/pricing.js +0 -74
@@ -0,0 +1,1068 @@
1
+ # Coding Worker Isolation
2
+
3
+ Wardby supports isolated Codex execution with Docker or Kubernetes and
4
+ credential-separated Claude Code execution with Docker. Both launchers fail
5
+ closed when the effective runtime does not match the reviewed policy.
6
+
7
+ ## Security Boundary
8
+
9
+ Coding-agent repositories and instructions are untrusted. The Docker daemon,
10
+ host kernel, Wardby control plane, immutable worker image, and dedicated coding
11
+ proxy are trusted. Containers are defense in depth rather than a VM boundary;
12
+ production should run the Docker host on a dedicated worker node or VM with no
13
+ production credentials beyond those required by the proxy.
14
+
15
+ The Codex worker has one network attachment: a unique per-run internal bridge using
16
+ Docker's isolated gateway mode. It has no default external route, published
17
+ port, host mapping, custom DNS server, or direct connection to the control
18
+ plane. A dedicated proxy container is attached to both that internal network
19
+ as `wardby-proxy` and an external network. No other service may join the run
20
+ network.
21
+
22
+ For a coding run with a package allowlist, the same proxy container also
23
+ serves the coding package registry (npm and PyPI) on the same
24
+ `wardby-proxy:8787` port, over a separate, registry-only token derived from
25
+ the run's capability; npm and pip never receive the run's model-API
26
+ capability. See [Installing packages in coding runs](coding-packages.md).
27
+ Registry mode serves both providers: Claude Code's tool runner gets that
28
+ registry-only token's settings from the trusted launcher, delivered as
29
+ `WARDBY_TOOL_SETUP` — never the run capability. Claude Code runs support the
30
+ `node` and `node-python` toolchains; the toolchain selects the tool-runner
31
+ image, which is resolved when the run is dispatched and kept for the whole run.
32
+
33
+ Claude Code uses a credential-separated composite job. Its agent container
34
+ holds the run capability and is attached only to the proxy network; it never
35
+ mounts the repository. Its tool runner container mounts the workspace and
36
+ shares that same run network (for a run with services, through the network
37
+ keeper's namespace; see below), so it can reach the coding proxy too, but it
38
+ holds no model capability or provider credential of its own — only the
39
+ registry-only settings and service variables in `WARDBY_TOOL_SETUP`. The two talk over a private
40
+ Unix socket (`/run/wardby/tool/runner.sock`). The trade is deliberate: the
41
+ tool runner can reach the proxy, but the proxy accepts nothing from it for a
42
+ model call. Both containers, the socket volume, keeper, network, and
43
+ artifacts are attested and cleaned as one persisted handle. On Docker, the tool
44
+ runner gets a fixed 0.25 CPU, a third of the run's memory (clamped between
45
+ 128 and 512 MiB), and 64 PIDs, and the agent gets the rest, so a Claude Code
46
+ run needs `cpus ≥ 0.35`, `memoryMb ≥ 256`, and `pids ≥ 96` (the default
47
+ `CODING_PIDS` is 128).
48
+
49
+ **Upgrading with Claude Code runs in flight (Docker launcher).** Let in-flight
50
+ Claude Code runs finish, or stop them, before upgrading the control plane. A
51
+ run launched by a different version fails the new version's container
52
+ attestation and is reported lost.
53
+
54
+ The proxy accepts a run-scoped capability, resolves only exact configured HTTPS
55
+ hostnames, rejects IP literals and every private, loopback, link-local,
56
+ documentation, transition, multicast, and metadata address, rejects mixed DNS
57
+ answers, and pins the vetted address into the socket lookup. Redirects are
58
+ denied. Injected fetch implementations are a test seam and must not be used in
59
+ production composition.
60
+
61
+ The Codex worker retries a failed model request, or a response stream that
62
+ drops, up to three times each before it fails the run. The Claude worker
63
+ likewise retries a failed model request up to three times. Every retry is a
64
+ new request through the proxy, so it is checked against the run's budget
65
+ again.
66
+ If the stream still fails, the run's diagnostic (in the control-plane log,
67
+ under the run's `coding_diag_…` id) names the cause as a fixed code, never the
68
+ error text: `job_coding_stream_proxy_denied` (401/403 from the proxy),
69
+ `job_coding_stream_rate_limited` (429), `job_coding_stream_upstream_error`
70
+ (5xx), `job_coding_stream_timeout`, `job_coding_stream_proxy_unreachable`
71
+ (the worker could not connect), `job_coding_stream_agent_exited` (the agent
72
+ process exited without an HTTP error), or `job_coding_stream_failed` when none
73
+ of these match. A refusal because the run's budget is used up ends the run as
74
+ out of budget instead.
75
+
76
+ The proxy side of the same request is in the coding proxy's log as
77
+ `audit.*` events (`audit.request.reserved`, `audit.response.completed`,
78
+ `audit.request.uncertain`, …), with ids, models, amounts and fixed reason
79
+ codes only. When the proxy has to end a stream early, `audit.request.uncertain`
80
+ says why: `upstream_failed:<code>` when the model API reported a failure (the
81
+ failure event is passed on to the worker unchanged), or a code such as
82
+ `terminal_usage_missing`, with the upstream response's content type and
83
+ encoding. To see the error text behind a code, turn on a
84
+ [debug trace](#debug-trace) for the agent.
85
+
86
+ **Model-provider failures.** The proxy also remembers the first failure the
87
+ model provider reported for a run: the error code of a failed stream, or of a
88
+ rejected request (its HTTP status as `http_<status>` when the response names no
89
+ code). When the run's worker job then fails, the run ends `failed` with the
90
+ error `coding_provider_<class>` and the failure category `provider_<class>`,
91
+ where the class is:
92
+
93
+ | Category | Provider codes | What the host is told |
94
+ | ----------------------- | --------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- |
95
+ | `provider_quota` | codes naming a quota, spend limit, billing, funds or credit (`insufficient_quota`, `project_spend_limit_exceeded`, …) | "The model provider refused the request: its account has reached a spending or quota limit. An operator needs to raise the limit with the provider, then retry." |
96
+ | `provider_rate_limited` | `rate_limit_exceeded` and other rate-limit codes | "The model provider is rate-limiting requests. Try again later." |
97
+ | `provider_unavailable` | `server_error`, `overloaded`, `overloaded_error`, `service_unavailable`, `api_error`, `http_5xx` | "The model provider reported an outage or overload. Try again later." |
98
+ | `provider_rejected` | anything else | "The model provider rejected the request." |
99
+
100
+ The host sees that sentence in the continuation status comment
101
+ (`❌ <agent> could not run: …`) and in an @-mention's status comment
102
+ (`❌ A sub-run could not reach the model: …`); it never sees the provider's
103
+ code or message. `provider_quota` means the **provider account** behind the
104
+ coding credential is out of quota or has hit its spend limit — raise the limit
105
+ with the provider. It is not wardby's budget: a run that uses up its own
106
+ wardby budget still ends `budget_exhausted`, and that wins when both happen.
107
+
108
+ For alerting, the control plane logs one `warn` line per such run with
109
+ `event: "coding.provider_failure"`, the `runId`, the `class`, and the raw
110
+ `upstreamCode`.
111
+
112
+ ### Model request allowlist
113
+
114
+ The network policy only helps if the one reachable destination, the proxy,
115
+ cannot be asked to reach somewhere else. Model APIs can do that on the
116
+ caller's behalf: an OpenAI hosted `mcp` tool or a remote `input_image` URL
117
+ makes OpenAI's servers contact any host, and hosted-tool fees and non-default
118
+ service tiers are billed outside the metered tokens. Code in a Codex worker can
119
+ read the run capability, so the proxy validates every request body against an
120
+ allowlist before it resolves a credential, and refuses anything else with a
121
+ `400` (`src/providers/coding-proxy/proxy.ts`). Each refusal is also written to
122
+ the proxy audit log as `request.rejected` with the run id and the refusal
123
+ code, so a smuggling attempt is attributable to its run. JSON nested too
124
+ deeply to process is refused with `request_nesting_too_deep`.
125
+
126
+ - **Anthropic Messages** (`parseAnthropicRequest`): a fixed set of keys,
127
+ text/`tool_use`/`tool_result`/thinking blocks, exactly the two Wardby tool
128
+ names, the reviewed beta values, the thinking mode the run's model catalog
129
+ entry names, and only effort levels that entry lists in `efforts` (top-level,
130
+ or on an effort-only system message for models that change effort per turn).
131
+ - **OpenAI Responses** (`parseOpenAiRequest`), built from what the pinned
132
+ Codex CLI actually sends (recorded in
133
+ `src/providers/coding-proxy/fixtures/codex-<version>-responses-requests.json`
134
+ and exercised end to end by `src/coding-worker/codex-compatibility.test.ts`):
135
+ - Top-level keys: `model`, `instructions`, `input`, `tools`, `tool_choice`,
136
+ `parallel_tool_calls`, `reasoning`, `store`, `stream`, `include`,
137
+ `prompt_cache_key`, `text`, `client_metadata`, `max_output_tokens`,
138
+ `background`, `service_tier`. Any other key, including `prompt`,
139
+ `previous_response_id`, `conversation` and `metadata`, is refused with
140
+ `openai_request_key_not_allowed:<key>`.
141
+ - Tools: `function` and `custom` definitions, and `namespace` groups of
142
+ them, either in `tools` or in Codex's `additional_tools` input item. Every
143
+ other tool type, including all hosted tools (`web_search`, `mcp`,
144
+ `code_interpreter`, `image_generation`, `file_search`, `local_shell`, and
145
+ so on), is refused with `openai_tool_not_allowed:<type>`. Tool names are
146
+ checked for shape (`[A-Za-z0-9_-]{1,128}`) but not against a fixed list:
147
+ they change with the model and the Codex version (for example `exec` in
148
+ code mode against `exec_command` otherwise), and a client-side tool runs
149
+ inside the worker, where the container boundary already governs it.
150
+ `tool_choice` may only be `auto`, `none` or `required`.
151
+ - Input items: `message` (`input_text`/`input_image` for user, developer
152
+ and system; `output_text` for assistant), `reasoning`, `function_call`,
153
+ `function_call_output`, `custom_tool_call`, `custom_tool_call_output`,
154
+ `agent_message` and `additional_tools`, each with only the keys Codex
155
+ sends. Anything else (`input_file`, `input_audio`, `item_reference`,
156
+ replayed hosted-tool calls, `compaction`, `local_shell_call`) is refused
157
+ with `openai_input_not_allowed:<type>`. An `input_image` must be an inline
158
+ `data:image/{png,jpeg,gif,webp};base64,` URL, as Codex's `view_image`
159
+ produces; a remote URL or a `file_id` is refused with
160
+ `openai_remote_input_not_allowed`.
161
+ - Pinned values: `include` only `reasoning.encrypted_content`; `reasoning`
162
+ only `effort`, `summary` and `context` with known values; `text` only
163
+ `verbosity` and a `text` or `json_schema` format; `service_tier` only
164
+ unset or `default` (`service_tier_not_allowed`; `auto` is refused because
165
+ it defers to the OpenAI project's own tier, which may be priority);
166
+ `client_metadata` only the string-valued keys Codex sends
167
+ (`openai_client_metadata_key_not_allowed:<key>`). The proxy still
168
+ forces `store: false` and `background: false` and adds a
169
+ `max_output_tokens` ceiling when Codex omits it.
170
+
171
+ Upgrading the pinned `@openai/codex-sdk` means re-recording that fixture
172
+ against a local fake upstream and rerunning the compatibility test. A Codex
173
+ release that sends a new key or item type fails closed at the proxy rather
174
+ than silently widening what reaches OpenAI.
175
+
176
+ The proxy also binds a second listener, the **deny port** (`8788`,
177
+ `CODING_PROXY_DENY_PORT`), which serves nothing: it accepts a connection,
178
+ sends no bytes and closes it immediately
179
+ (`src/providers/coding-proxy/deny-port.ts`). It exists so a run pod can prove
180
+ its own NetworkPolicy is enforced — see "The enforcement gate" below.
181
+
182
+ The deny port ships to **every** deployment, not just Kubernetes:
183
+ `startConfiguredCodingProxy` is the only proxy entry point, so a Docker
184
+ deployment's proxy also binds `0.0.0.0:8788` (and refuses to start if it
185
+ cannot). No host port is published for it, so there is no port conflict. In
186
+ Docker mode a run container can reach it, since a Docker network has no
187
+ port-level policy — that is harmless, because the listener accepts the
188
+ connection, sends nothing and closes it. It is not a leak; it is a fact about
189
+ reachability that the Kubernetes launcher turns into evidence.
190
+
191
+ ## Debug trace
192
+
193
+ When a fixed code is not enough to tell why a coding agent's runs fail, an
194
+ admin can turn on a **debug trace** for that agent for a limited time:
195
+
196
+ ```
197
+ update_agent { "id": "<agent id>", "codingProfile": { "debugTraceMinutes": 30 } }
198
+ ```
199
+
200
+ `debugTraceMinutes` is 1 to 1440 and needs the `agents:admin` scope (the admin
201
+ role), as `workerImageRef` does; `null` turns the trace off early. Every change
202
+ is written to the control-plane log as a `coding.debug_trace.set` audit line.
203
+ `get_agent` shows the expiry as `codingProfile.debugTraceUntil`. The trace
204
+ expires on its own: each coding run dispatched before the expiry is traced for
205
+ its whole life, and `get_run` shows `debugTrace: true` for it. Runs dispatched
206
+ afterwards are not.
207
+
208
+ A traced run's Codex worker writes every Codex stream event, and the full text
209
+ of a stream failure (including the agent process's own error output and cause
210
+ chain), to its **own log**, one JSON line each under a `debugTrace` key. On
211
+ Kubernetes that is the run pod's `worker` container log; on GKE, Cloud Logging
212
+ keeps it after the pod is deleted. With the Docker launcher it is the worker
213
+ container's log (`docker logs`), for as long as the container exists. Nothing from the trace goes to the database,
214
+ GitHub, the run's error, or the control-plane log. Token-shaped values (the run
215
+ capability, bearer tokens, API keys, GitHub tokens) are redacted, each line is
216
+ capped at 16 KiB and each run at 2 MiB.
217
+
218
+ The trace can contain prompts, model output and repository content. Turn it on
219
+ only while diagnosing a failure, for as short a time as you can, and treat the
220
+ pod logs of traced runs as sensitive. Tracing needs a worker driver image that
221
+ supports it; a worker that predates it rejects a traced run's input as
222
+ invalid.
223
+
224
+ ## Container Policy
225
+
226
+ `src/providers/jobs/docker-isolation.ts` is the canonical policy builder and
227
+ startup attestation layer. `DockerJobLauncher` executes its argument arrays
228
+ directly with `spawn`; it never invokes a shell.
229
+
230
+ The worker policy requires:
231
+
232
+ - An immutable `sha256:` image ID or repository digest with `--pull never`.
233
+ - UID/GID `10001:10001`, all capabilities dropped, no new privileges, Docker's
234
+ built-in seccomp profile, private cgroup and PID namespaces, and no host IPC.
235
+ - A read-only root filesystem with bounded `noexec,nosuid,nodev` tmpfs mounts
236
+ for `/tmp` and `/home/wardby`. The agent's own temporary files (`TMPDIR`)
237
+ are not among them: they go to the workspace's `.cache/tmp`, which is
238
+ disk-backed and counts toward the run's `workspaceDiskMb`, not this small
239
+ in-memory scratch — a real `pip install` or `npm install` unpacks and builds
240
+ in `TMPDIR`, which the tmpfs is usually too small to hold. `.cache` is
241
+ never part of the collected diff or pull request.
242
+ - Exact CPU, memory, equal memory+swap, PID, shared-memory, disk, and wall-clock
243
+ limits. Equal memory and memory+swap disables additional swap allowance.
244
+ - No devices, device requests, bind mounts, extra groups, custom DNS, extra
245
+ hosts, published ports, or restart policy.
246
+ - At most 2 MiB of local Docker logs and a cooperative SIGTERM grace period
247
+ before forced termination.
248
+
249
+ A run with services (see [coding-services.md](coding-services.md)) adds two
250
+ kinds of container, both labelled and attested like the others before they
251
+ start:
252
+
253
+ - A **network keeper** (`wardby-netns-<token>`): the worker image running an
254
+ idle `node` process as `10001:10001`, read-only, all capabilities dropped,
255
+ no new privileges, built-in seccomp, 64 MiB and 32 PIDs, with no mounts, on
256
+ the run's internal network. It owns the run's network namespace.
257
+ - One container per service (`wardby-svc-<token>-<name>`): the catalog image
258
+ by digest, as `10001:10001`, with a read-only root filesystem, all
259
+ capabilities dropped, no new privileges, built-in seccomp, private cgroup
260
+ namespace, private IPC with 64 MiB of shared memory, the catalog's CPU,
261
+ memory (plus its tmpfs disk and shared memory), 512 PIDs, a bounded tmpfs at
262
+ its data path and each writable path, its catalog `serviceEnv`, and no
263
+ published ports, mounts, devices or restart policy.
264
+
265
+ A Codex run's worker and every service use `--network container:<network
266
+ keeper>`, so a service answers the worker on `127.0.0.1` and the worker still
267
+ reaches the proxy by its alias on the internal network. Nothing else about the
268
+ worker changes.
269
+
270
+ For a Claude Code run it is the tool runner, which runs the agent's shell
271
+ commands, that joins the network keeper's namespace: it reaches every service
272
+ on `127.0.0.1` and the proxy by its alias, and it is created only after every
273
+ service is ready. The agent container stays on the run's internal network,
274
+ unchanged; it reaches the tool runner over the Unix socket in their shared
275
+ storage volume, which does not depend on either container's network. The
276
+ launcher attests the tool runner's network before it starts and on every
277
+ status check: the network keeper's namespace and no network of its own.
278
+
279
+ A service may listen on every interface of the network keeper's namespace, so
280
+ anything on the run's internal network (the proxy, and for a Claude Code run
281
+ the agent container) can address it there, just as every container in a
282
+ Kubernetes run's pod shares the services' namespace. Services are disposable
283
+ test fixtures with catalog credentials; do not treat them as a boundary.
284
+
285
+ The capability value is inherited from the trusted launcher's child-process
286
+ environment with `--env WARDBY_RUN_CAPABILITY`; it is never included in command
287
+ arguments. Docker administrators can still inspect container environment, so
288
+ daemon access remains privileged and must be tightly restricted.
289
+
290
+ ## Ephemeral Storage
291
+
292
+ Each run receives one quota-bounded local tmpfs volume. A hardened, no-network
293
+ keeper container holds the volume open from preparation through result
294
+ collection. It creates exactly four private subdirectories and emits
295
+ `wardby_storage_ready` before the launcher may seed them.
296
+
297
+ The worker sees only these volume subpaths:
298
+
299
+ - `/workspace`: read-write checkout files.
300
+ - Git metadata remains in the trusted keeper volume and is not mounted into the worker.
301
+ - `/run/wardby/input`: read-only, validated input artifact.
302
+ - `/run/wardby/output`: read-write result artifact.
303
+
304
+ There are no production host bind mounts. The Docker JobLauncher transfers data
305
+ through the keeper with Docker copy/archive APIs, validates it before launch and
306
+ after collection, and stops the keeper only after collection. Stopping the last
307
+ container that mounts this local tmpfs intentionally destroys the run data.
308
+ The quota is RAM-backed; operators must bound aggregate concurrent `diskMb`
309
+ allocations at the host scheduler as well as per run.
310
+
311
+ ## Required Lifecycle
312
+
313
+ 1. Validate `JobSpec`, image digest, proxy identity, and host support.
314
+ 2. Create and inspect the internal network and tmpfs volume.
315
+ 3. Create, inspect, and start the keeper; wait for its readiness marker.
316
+ 4. Seed the four fixed storage areas through the keeper, never a bind mount.
317
+ 5. Attach the dedicated proxy and attest that it is both internally and
318
+ externally connected.
319
+ 6. For a run with services only: create, inspect and start the network
320
+ keeper; then, for each service in turn, use the host's copy of its image or
321
+ pull it by digest, create and inspect its container in the keeper's
322
+ namespace, start it, and run its readiness command until it passes or the
323
+ run fails with `coding_service_unready:<name>`. All of this shares one
324
+ 120-second start-up limit (image pulls, each bounded to 5 minutes, are not
325
+ counted) and never runs past the run's deadline.
326
+ A Claude Code run's tool runner is created, inspected and started only
327
+ after this step, in the network keeper's namespace.
328
+ 7. Create the worker with the run capability supplied only in the child
329
+ environment; inspect every effective control before start.
330
+ 8. Start the worker and enforce `deadlineMs`. Send SIGTERM at expiry, then
331
+ SIGKILL after `stopGraceSeconds` if it remains alive.
332
+ 9. Cancel the proxy session, collect and validate bounded output, and re-read
333
+ authoritative usage before any repository publication.
334
+ 10. For `changes_ready` only, copy the worker workspace into a new host staging
335
+ directory, reject special files, nested `.git`, escaping symlinks, and size
336
+ or entry-limit violations, then atomically replace the trusted checkout.
337
+ 11. Revalidate protected paths, Git configuration, branch ancestry, remotes,
338
+ and budget; create one controlled commit, push one deterministic branch,
339
+ and create or find one draft pull request.
340
+ 12. Persist the typed coding result and terminal run status in one transaction,
341
+ then remove the worker, any service containers and network keeper, the
342
+ keeper, network, volume, input artifact, and VCS workspace. `no_changes`
343
+ and `budget_exhausted` never push.
344
+
345
+ `ContainerExecutor` treats a durable proxy session without a durable job handle
346
+ as ambiguous provisioning and never relaunches it. A persisted handle is the
347
+ only recovery path. Duplicate starts, terminal collection, Git finalization,
348
+ and cleanup converge on the same handle, branch, commit, pull request, usage,
349
+ and status.
350
+
351
+ Any missing host feature, unsupported network option, failed inspection,
352
+ unexpected mount/network/environment, or cleanup ambiguity is the fixed
353
+ `docker_isolation_unsupported` failure. Production must not fall back to a
354
+ weaker profile.
355
+
356
+ ## Control Plane Configuration
357
+
358
+ Set `JOB_LAUNCHER=docker`, `CODING_WORKER_IMAGE` to an immutable repository
359
+ digest or Docker local image ID, and `CODING_PROXY_CONTAINER` to the dedicated proxy container name.
360
+ For Claude Code, also set `CODING_CLAUDE_WORKER_IMAGE` and
361
+ `CODING_CLAUDE_TOOL_RUNNER_IMAGE` to their immutable IDs.
362
+ `VCS_WORK_ROOT`, `CODING_JOB_STATE_ROOT`, and `CODING_ARTIFACT_ROOT` must be
363
+ trusted host-only directories. Resource limits are controlled by
364
+ `CODING_CPUS`, `CODING_MEMORY_MB`, `CODING_PIDS`, and `CODING_DISK_MB`.
365
+ `CODING_MAX_DISK_MB` (must be an integer between 64 and 32768 and at least
366
+ the effective `CODING_DISK_MB`; **defaults to the effective `CODING_DISK_MB`
367
+ itself**, not a larger number) is the operator ceiling on the per-agent
368
+ `workspaceDiskMb` coding-profile field described below — without it, any
369
+ `agents:write` caller could size a run's workspace disk up to 32 GiB
370
+ (Docker: a RAM-backed tmpfs; Kubernetes: an `emptyDir`), times
371
+ `CODING_MAX_CONCURRENT`, on every run. Defaulting the ceiling to the
372
+ existing disk size means upgrading to this branch changes nothing for a
373
+ deployment that doesn't set `CODING_MAX_DISK_MB`: `workspaceDiskMb` stays
374
+ inert until an operator explicitly raises the ceiling above `CODING_DISK_MB`.
375
+
376
+ A coding agent's profile carries an optional `workspaceDiskMb` (MiB; `null`
377
+ means "use the deployment default `CODING_DISK_MB`"). It is snapshotted onto
378
+ the `CodingRun` at dispatch time, so a later profile edit never changes an
379
+ in-flight run's size, and it is capped by `CODING_MAX_DISK_MB`: a run whose
380
+ snapshotted `workspaceDiskMb` exceeds the ceiling fails with
381
+ `coding_workspace_disk_exceeds_limit`
382
+ (`src/providers/executor/container.ts`, the `jobSpec()` check). This check
383
+ runs after the workspace has already been cloned onto the control plane, so
384
+ an over-ceiling agent still pays for a clone before failing, and the failure
385
+ reaches the run record only as the sanitized `coding_failure_workspace:<id>`
386
+ (the same generic bucketing every workspace/git-related failure gets) — not
387
+ a distinctly labeled "refused" outcome, and not currently logged anywhere
388
+ more diagnosable on the control plane. This is a known rough edge, not a
389
+ security gap: no run ever exceeds the ceiling, it
390
+ just fails less legibly than it could.
391
+
392
+ `CODING_MAX_CONCURRENT` (default `4`) caps coding runs that hold a slot at
393
+ once, across every control-plane replica: the cap is enforced in Postgres
394
+ inside the provisioning claim, so adding replicas never raises it. A run
395
+ over the cap stays `pending` and `get_run`/`list_runs` show `codingQueuedAt`;
396
+ it starts, oldest first, when a slot frees (immediately in the process whose
397
+ run finished, or on the scheduler leader's next tick). A newly dispatched run
398
+ never takes a free slot ahead of an older queued run. A run still queued after
399
+ `CODING_QUEUE_TIMEOUT_SEC` (default `3600`) fails with `coding_queue_timeout`.
400
+ Slot usage is derived from run state, so a crashed replica cannot leak slots:
401
+ its runs are reconciled to `lost`, which frees them.
402
+
403
+ Operating the queue across replicas:
404
+
405
+ - Every replica must set the same `CODING_MAX_CONCURRENT`. Each claim
406
+ enforces the value of the replica making it, so mismatched values make the
407
+ effective cap depend on which replica dispatched the run.
408
+ - Draining on a timer and applying `CODING_QUEUE_TIMEOUT_SEC` need a
409
+ scheduler process (`wardby serve` or `wardby scheduler`). A process that
410
+ only serves MCP (`wardby mcp`) drains only when one of its own coding runs
411
+ finishes; without a scheduler somewhere, queued runs can wait indefinitely
412
+ and never time out.
413
+ - Upgrade all replicas together. A replica running a version from before
414
+ the queue ignores the cap, and its reconciler reaps queued runs as `lost`. Clones are shallow
415
+ (`--depth 1`); the worker never receives Git history and finalization needs
416
+ only the base commit.
417
+
418
+ Every wardby process also polls the model catalog on
419
+ `WARDBY_MODEL_CATALOG_REFRESH_SECONDS` (default `45`), which decides whether
420
+ a coding run's model is still available and how it is priced at dispatch; see
421
+ [Models and pricing](models.md).
422
+
423
+ The GitHub adapter requires `GITHUB_APP_ID` and `GITHUB_APP_PRIVATE_KEY`; the
424
+ App installation is checked while preparing the workspace, before the
425
+ billable proxy session is created. Upstream keys remain behind
426
+ `CODING_OPENAI_CREDENTIAL_REF` and `CODING_ANTHROPIC_CREDENTIAL_REF` and are
427
+ never written to the database, input
428
+ artifact, Docker arguments, or Git workspace.
429
+
430
+ The embedded Codex SDK runs with its inner sandbox disabled because the
431
+ worker's Docker boundary is authoritative: it has a read-only root filesystem,
432
+ no Linux capabilities, no host mounts or Docker socket, no public network,
433
+ and only isolated workspace/output volumes plus the trusted proxy connection.
434
+ This avoids relying on a nested sandbox that cannot validate Wardby's
435
+ intentionally Git-metadata-free workspace.
436
+
437
+ Coding-agent authoring and execution are MCP-first. `trigger_agent` accepts
438
+ an optional bounded `task` and `baseRef` only for a coding agent owned by the
439
+ caller; the chosen values are copied into the immutable run record. Webhooks
440
+ use the profile default task unless `allowWebhookTaskOverride` is explicitly
441
+ enabled on that coding profile. Coding results returned through `get_run` and
442
+ `tasks/get` are validated, redacted summaries/tests only; job handles and
443
+ execution policy stay internal.
444
+
445
+ The worker receives only its task text, so a coding agent's own `systemPrompt`
446
+ is placed ahead of the request at dispatch ("Standing instructions for this
447
+ coding agent: … Request: …") and stored with the run. A blank prompt leaves
448
+ the task unchanged; a combination over the 16 KiB task limit is refused rather
449
+ than truncated.
450
+
451
+ The worker's result may carry an optional `tag`, shown as `[tag]` in the pull
452
+ request title: at most 32 characters of letters, digits, `.`, `_`, `/`, and
453
+ `-`, starting with a letter or digit. The output schema describes that rule to
454
+ the model, and a tag that breaks it is normalized to a lowercase slug (or
455
+ dropped) instead of failing an otherwise finished run. GitHub finalization
456
+ re-validates the tag independently. Linking an issue is not the tag's job: put
457
+ a closing keyword such as `Resolves #37` in the task so the worker includes it
458
+ in the summary, which becomes the pull request body.
459
+
460
+ Some workspace folders are never collected from a run: `node_modules`, `.venv`,
461
+ `venv`, `__pycache__`, `.pytest_cache`, `.ruff_cache`, `.mypy_cache`, `.tox`,
462
+ `.vite`, and `.cache`, at any depth, plus any repository-relative paths in the
463
+ agent's `collectExclude` (for example `web/dist`). On Kubernetes the keeper's
464
+ `tar` leaves them out, so they never leave the pod; on Docker they are removed
465
+ from the staging copy before it is validated. They therefore never count toward
466
+ the entry, size, symlink, or nested-repository checks, and Git staging excludes
467
+ them too, so a tracked file under an excluded folder is left unchanged.
468
+
469
+ Operator-only checks and cleanup remain available through the CLI:
470
+
471
+ ```sh
472
+ wardby coding preflight
473
+ wardby coding cleanup --run-id <id>
474
+ ```
475
+
476
+ The preflight command requires Docker mode, validates the pinned worker-image
477
+ digest, and confirms the image is available to Docker. Cleanup delegates to
478
+ the configured executor so it resolves and stops the persisted container job.
479
+
480
+ ## Audit And Retention
481
+
482
+ The executor emits metadata-only lifecycle events for queueing, preparation,
483
+ launch, running, budget cutoff, stopping, collection, PR creation, terminal
484
+ outcome, and cleanup. Events carry run ID, opaque job ID, sanitized failure
485
+ category, opaque diagnostic ID, duration, and budget totals only. They never
486
+ carry task text, prompts, repository contents, diffs, worker environment, raw
487
+ Docker logs, or credentials.
488
+
489
+ The production log/metrics collector retains those events for 90 days by
490
+ default policy. `CodingRun` stores only the sanitized failure category and
491
+ diagnostic ID alongside the normal run record; it is not an artifact store.
492
+ Every terminal path removes the worker volume, input artifact, job state, and
493
+ trusted checkout. Restart reconciliation repeats that cleanup from the
494
+ persisted job handle and marks ambiguous provisioning as `lost` instead of
495
+ relaunching it.
496
+
497
+ See [release verification](release-verification.md) for automated gates and
498
+ live-fixture rules.
499
+
500
+ ## Kubernetes launcher (`JOB_LAUNCHER=kubernetes`)
501
+
502
+ `KubernetesJobLauncher` (`src/providers/jobs/kubernetes.ts`) implements the
503
+ same `WorkspaceJobLauncher` contract as `DockerJobLauncher` and is a drop-in
504
+ alternative for deployments with no Docker daemon available to the control
505
+ plane (e.g. a Cloud Run host). Everything above `ContainerExecutor` —
506
+ proxy sessions, the Git finalizer, workspace validation, recovery, cleanup —
507
+ is unchanged; only the container-orchestration seam is replaced. Both Codex
508
+ and Claude Code run on this launcher; Claude Code's pod adds a tool-runner
509
+ sidecar, described in "Pod layout" below.
510
+
511
+ ### Enabling it
512
+
513
+ Set `JOB_LAUNCHER=kubernetes` and `CODING_WORKER_IMAGE` to a **registry
514
+ digest** (`repo@sha256:<64 hex>` — a bare `sha256:` local image ID is
515
+ rejected; a cluster cannot pull it). For Claude Code, also set
516
+ `CODING_CLAUDE_WORKER_IMAGE` and `CODING_CLAUDE_TOOL_RUNNER_IMAGE` (and, for
517
+ Claude agents on the `node-python` toolchain,
518
+ `CODING_CLAUDE_TOOL_RUNNER_IMAGE_NODE_PYTHON_3_12`) to registry digests; the
519
+ control plane refuses to start if any of them is set to anything else.
520
+ Kubernetes-specific settings (`src/config/providers.ts`,
521
+ `loadKubernetesJobConfig`):
522
+
523
+ | Variable | Default | Meaning |
524
+ | ------------------------------- | ----------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
525
+ | `KUBERNETES_NAMESPACE` | `wardby-coding` | The one namespace holding the proxy and every per-run object. |
526
+ | `KUBERNETES_PROXY_SERVICE` | `wardby-coding-proxy` | The proxy's Service name; its ClusterIP is what `hostAliases` points runs at. |
527
+ | `KUBERNETES_CONTEXT` | (unset → in-cluster/default kubeconfig context) | Which kubeconfig context `ClientNodeKubernetesApi` connects with. |
528
+ | `KUBERNETES_RUNTIME_CLASS` | (unset) | e.g. `gvisor` on GKE. Unset means pods run without a sandboxing runtime class — logged once per launch as `kubernetes_runtime_class_unset` and development-only. |
529
+ | `KUBERNETES_RUN_PRIORITY_CLASS` | (unset) | PriorityClass for run pods and the preflight canary, e.g. `wardby-coding-run` in the GKE overlay. Must be an existing class and a DNS-1123 subdomain; `system-` classes are refused. Unset means no class (priority 0). |
530
+
531
+ The `CODING_CPUS` / `CODING_MEMORY_MB` / `CODING_PIDS` / `CODING_DISK_MB` /
532
+ `CODING_MAX_DISK_MB` settings above apply identically; the per-agent
533
+ `workspaceDiskMb` profile field sizes the pod's `storage` `emptyDir` the same
534
+ way it sizes Docker's tmpfs volume.
535
+
536
+ ### Pod layout
537
+
538
+ One pod per run, built by the canonical, deny-by-default policy in
539
+ `kubernetes-isolation.ts`'s `buildRunPod`:
540
+
541
+ - An **init container `storage-init`** runs first and creates
542
+ `/run/wardby/storage/{workspace,input,output}` (mode `0700`, owned by uid 10001) before any regular container starts. This exists because kubelet
543
+ creates a subPath mount's target directory root-owned the first time it
544
+ sets up the worker's volume mounts, and the keeper (uid 10001, no Linux
545
+ capabilities) cannot `chmod` a root-owned directory it doesn't own. This
546
+ is required because otherwise a real-cluster run fails
547
+ `kubernetes_pod_start_timeout`. `keeper.js` itself is unchanged — Docker still
548
+ shares it, and Docker's bind-mount-free volume never had this problem.
549
+ - **Claude Code's `tool-runner`** (only for a Claude Code run): a native
550
+ sidecar (init container with `restartPolicy: Always`) after `storage-init`
551
+ and before the service sidecars, `keeper`, and `worker`. It mounts the
552
+ workspace and the socket directory `/run/wardby/tool`, gets its `WARDBY_TOOL_SETUP`
553
+ environment variable from the run Secret's `tool-setup` key (never the run
554
+ capability), and shares the pod's network and IPC namespaces with the
555
+ worker, like every container in a pod: loopback and abstract Unix sockets
556
+ are common to both, and the worker listens on nothing. The run's one egress
557
+ rule (the proxy) applies to it the same as the worker. Unlike on Docker
558
+ (where it runs under `--init`), it has no init process on Kubernetes; it
559
+ stops when the pod is deleted. Its
560
+ `startupProbe` runs `test -S /run/wardby/tool/runner.sock`, so the keeper
561
+ and worker wait for the socket to exist before they start. The socket
562
+ directory is its own small memory-backed `emptyDir` (`tool-socket`),
563
+ mounted by the tool runner and the worker only, never a subdirectory of the
564
+ disk-backed storage volume: under gVisor a Unix socket bound on a volume
565
+ that isn't shared across the sandbox is invisible to the other container,
566
+ while a memory-backed `emptyDir` that two containers mount is one shared
567
+ tmpfs (GKE Autopilot annotates it `share: pod`). Before it calls the model,
568
+ the Claude worker connects to the socket once and fails the run as
569
+ `worker_tool_runner_unreachable` if it can't, so a run never proceeds
570
+ without its command tool. Of the pod's
571
+ resources, the tool runner gets a fixed 0.25 CPU and a third of the run's
572
+ memory (clamped between 128 and 512 MiB); the agent (`worker`) gets the
573
+ rest of both. The two also split the worker's 1024 MiB ephemeral-storage
574
+ reservation: 256 MiB to the tool runner, 768 MiB to the agent. GKE
575
+ Autopilot may round each container's requests up to its own minimums, so
576
+ there the split is approximate.
577
+ - **Service sidecars** (only for a run with services, see
578
+ [coding-services.md](coding-services.md)): one init container per service,
579
+ `service-<name>`, with `restartPolicy: Always` and a `startupProbe` from its
580
+ catalog entry — this native-sidecar shape (an init container that keeps
581
+ running) needs Kubernetes 1.29 or later — after `storage-init` (and, for a
582
+ Claude Code run, `tool-runner`) and before `keeper` and `worker`, so
583
+ neither starts until every service is ready. Each runs its catalog image,
584
+ pinned by digest, with the same security context as every other container
585
+ (uid 10001, read-only root filesystem, no privilege escalation, all
586
+ capabilities dropped), `emptyDir` volumes at its data and writable paths, and
587
+ requests equal to limits. It shares the pod's network namespace, so the worker
588
+ reaches it on `127.0.0.1` and the run's NetworkPolicy is unchanged: a service
589
+ can reach nothing the worker cannot. For a Claude Code run, the tool runner
590
+ reaches it on `127.0.0.1` too, and gets the same catalog `testEnv` the worker
591
+ would. The worker never receives a service's own
592
+ environment, only the catalog's `testEnv`. Attestation compares sidecars like
593
+ every other container.
594
+ - **`keeper`**: trusted, holds the one `storage` volume (an `emptyDir` sized
595
+ `spec.limits.diskMb` MiB, **disk-backed**, not `medium: Memory`) open for
596
+ the pod's life; the seam streams the workspace and input artifact in and
597
+ the output artifact out through it via `kubectl exec`-style calls
598
+ (`tar` in/out). Readiness probe: the `output` subdirectory exists.
599
+ - **`worker`**: untrusted. Its command is overridden to a small polling gate
600
+ (`WORKER_GATE`) that waits for `/run/wardby/input/.seeded` before
601
+ `import()`-ing the image's real entrypoint — for a Claude Code run, the
602
+ Claude entrypoint, which never mounts `/workspace` (only `/run/wardby/tool`,
603
+ to reach the socket, plus `input` and `output`). A pod's containers all
604
+ start together — Kubernetes has no "start this container later" — so the
605
+ gate is what makes "seed first, run second" possible without a native
606
+ sidecar (which Kubernetes terminates when the main container exits, killing
607
+ the keeper before result collection).
608
+ - `/tmp` and `/home/wardby` are small `medium: Memory` `emptyDir`s (bounded
609
+ `min(64, max(16, memoryMb/8))` MiB), matching Docker's bounded tmpfs mounts;
610
+ a Claude Code run's tool runner gets its own pair, sized the same way from
611
+ its own memory share. Neither is where the agent's commands get `TMPDIR`:
612
+ that points at the workspace's `.cache/tmp` on the disk-backed `storage`
613
+ volume instead, so it scales with `workspaceDiskMb` rather than this tiny
614
+ in-memory scratch.
615
+ - `dnsPolicy: None` with `dnsConfig.nameservers: ["127.0.0.1"]` — **no DNS is
616
+ configured for worker/agent pods at all.** The proxy is reached by name
617
+ (`WARDBY_PROXY_URL=http://wardby-proxy:8787` — `CODING_PROXY_ALIAS` in
618
+ `docker-isolation.ts`, the same alias the Docker launcher uses; this is
619
+ distinct from the `KUBERNETES_PROXY_SERVICE` Kubernetes Service name,
620
+ `wardby-coding-proxy` by default — the proxy checks the `Host` header)
621
+ only because the pod's `hostAliases` maps `wardby-proxy` directly to the
622
+ proxy Service's ClusterIP — closing DNS as an exfiltration channel
623
+ without needing a resolver at all.
624
+ - `activeDeadlineSeconds = spec.timeoutSec + POD_DEADLINE_GRACE_SECONDS`
625
+ (300s) — a **backstop only**. The launcher enforces the real wall-clock
626
+ deadline itself (`observePod` in `kubernetes.ts`); the extra 300s exists so
627
+ the keeper survives long enough after the worker's deadline for result
628
+ collection to still succeed. A worker that exits 0 counts as `succeeded`
629
+ only if all three hold: the pod's status reason isn't `DeadlineExceeded`,
630
+ the pod isn't being deleted (`metadata.deletionTimestamp` unset), and the
631
+ terminated container's `finishedAt` is at or before
632
+ `deadlineAt + 5s` (clock-skew slack). This closes off a SIGTERM-trapping
633
+ worker turning a deadline kill into a fake success, while still accepting
634
+ a run that genuinely finished just before its deadline and was only
635
+ observed after it.
636
+ - Everything else matches the Docker policy's spirit: uid/gid 10001,
637
+ `runAsNonRoot`, all capabilities dropped, `allowPrivilegeEscalation:
638
+ false`, seccomp `RuntimeDefault`, read-only root filesystem, no host
639
+ network/PID/IPC, `automountServiceAccountToken: false`, a dedicated
640
+ no-RBAC service account (`wardby-coding-worker`).
641
+
642
+ Per-run objects, all labeled `app.kubernetes.io/managed-by: wardby`,
643
+ `wardby.io/component: coding-run`, `wardby.io/run-sha256: <sha256(runId)
644
+ prefix>`, named `wardby-run-<token>` (`<token>` = first 20 hex chars of
645
+ `sha256(runId)`):
646
+
647
+ - The **pod** and its **NetworkPolicy** (same name).
648
+ - A **capability Secret** (`wardby-run-<token>-cap`) holding the run's proxy
649
+ capability, injected only via `secretKeyRef` — never in the pod spec,
650
+ command, or arguments. For a Claude Code run, the same Secret also carries
651
+ a `tool-setup` key (the tool runner's `WARDBY_TOOL_SETUP` JSON), likewise
652
+ injected only via `secretKeyRef`.
653
+ - A **record ConfigMap** (also `wardby-run-<token>`) holding all job state —
654
+ phase, deadline, result — updated with optimistic concurrency
655
+ (`resourceVersion`). This replaces local state files and in-process
656
+ timers entirely: any control-plane replica can observe, collect, stop, or
657
+ remove any run, and a restarted process loses nothing.
658
+
659
+ **Record ConfigMaps are retained as tombstones by design.**
660
+ `remove()` deletes the pod, NetworkPolicy, and capability Secret, but
661
+ deliberately _rewrites the record to `phase: "removed"` instead of deleting
662
+ it_ (`kubernetes.ts`, `remove()`) — the contract is that a removed run is
663
+ never relaunched, and the record is what a later `launch()` call for the
664
+ same run ID checks. Wardby does not yet garbage-collect old tombstones, so they
665
+ accumulate until an operator removes them. **A launch that fails during
666
+ provisioning accumulates a record too**,
667
+ not just a `remove()`d run's tombstone: the failure path writes `phase:
668
+ "failed"` and the executor never calls `remove()` for a launch that threw,
669
+ so every failed launch leaves a permanent record as well. On a busy cluster
670
+ this is etcd growth to budget for operationally, not a correctness or
671
+ security problem — see "Known gaps" below.
672
+
673
+ **The executor persists a run's job handle before calling `jobs.launch()`,
674
+ closing the crash window that would otherwise orphan a running pod.** The
675
+ Kubernetes handle (`{ backend: "kubernetes", id: "<namespace>/<token>" }`)
676
+ is fully derivable from the run ID alone, with no cluster call, so it can be
677
+ (and is) written to the run record before the launcher ever creates
678
+ anything. If the control plane crashes or is replaced mid-launch — after the
679
+ pod has been attested and the worker gate has opened, but before the old
680
+ code path would have recorded the handle — the run's handle is already on
681
+ record, so restart reconciliation can find, stop, and clean up the pod,
682
+ NetworkPolicy, and capability Secret through the normal `abandon()` path
683
+ instead of leaking them permanently.
684
+
685
+ ### Attestation — deny-by-default, fail closed
686
+
687
+ Before the worker gate ever opens, the launcher reads back the created pod
688
+ and NetworkPolicy and compares them against the canonical builder's output
689
+ with `assertRunPodMatches` / `assertRunNetworkPolicyMatches`
690
+ (`kubernetes-isolation.ts`). This is a **full, canonical, deep comparison of
691
+ the entire spec, labels, and annotations** — not an allowlist of fields the
692
+ launcher expects to see. An early allowlist-based version of this comparator
693
+ was replaced during implementation review specifically because an allowlist
694
+ silently accepts anything it forgot to check (lifecycle hooks,
695
+ liveness/readiness/startup probes whose `httpGet.host` can reach the node's
696
+ link-local metadata endpoint bypassing the NetworkPolicy, `procMount`,
697
+ `seLinuxOptions`, extra tolerations, `nodeSelector`, stray annotations, ...).
698
+ The only normalization applied before comparing is an explicit, narrow list
699
+ of transformations the Kubernetes API server itself is known to perform on
700
+ write/read — never a loosening of what's compared:
701
+
702
+ - Dropping `schedulerName`, `nodeName`, `priority`, `preemptionPolicy` (server-assigned).
703
+ - Dropping `hostNetwork`/`hostPID`/`hostIPC` when `false`, and an empty
704
+ NetworkPolicy `ingress: []`, because Go's `omitempty` drops a zero-value
705
+ bool or empty slice on serialization — a real read-back never carries
706
+ these fields at their false/empty value, only when true/non-empty.
707
+ - Removing the mirrored `serviceAccount` field when it equals
708
+ `serviceAccountName` (and failing closed if it doesn't).
709
+ - Removing exactly the two well-known `NoExecute` node-health tolerations
710
+ (`node.kubernetes.io/not-ready` / `unreachable`, 300s) every pod gets by
711
+ admission-time default — any other toleration must match exactly.
712
+ - Dropping container `terminationMessagePath`/`terminationMessagePolicy`/`imagePullPolicy`
713
+ and probe threshold/period defaults, and normalizing CPU/memory quantities
714
+ to a canonical millicore/byte count (so `"1"`, `"1.0"`, and `"1000m"`
715
+ compare equal) — with a non-integer-at-that-scale value mapped to a
716
+ sentinel that can never equal a real value, so quantity drift fails closed
717
+ instead of rounding two different resources together.
718
+ - Dropping `mountPropagation: "None"` and an `emptyDir.medium: ""`.
719
+
720
+ Any other difference — anything not on this list — fails the launch closed
721
+ with `kubernetes_isolation_unsupported`. No fallback to a weaker profile.
722
+
723
+ On top of that, a **platform profile** (`KUBERNETES_PLATFORM`) may forgive a
724
+ named, narrow set of mutations its managed admission chain is known to make —
725
+ for `gke-autopilot`: annotations under `autopilot.gke.io/` and `dev.gvisor.`,
726
+ labels under `autopilot.gke.io/` and `topology.kubernetes.io/`, the gVisor
727
+ `nodeSelector`, and two exact tolerations. A profile can only delete named keys
728
+ from both operands before the comparison; it can never disable or short-circuit
729
+ it, and `generic` (the default, and what an omitted argument selects) forgives
730
+ nothing. Most of that list was captured from a real cluster with
731
+ `npm run capture:autopilot`, which submits the pod with `dryRun=All`. That
732
+ capture has a structural blind spot worth knowing: a dry run is admission only
733
+ and never schedules, so anything the platform stamps on **after binding** —
734
+ GKE's `topology.kubernetes.io/{region,zone}`, taken from the node the pod
735
+ landed on — cannot appear in it. That case reached a live Autopilot launch with
736
+ every dry-run-derived check passing, and is now pinned by tests. Treat a
737
+ captured fixture as a lower bound on what a platform mutates.
738
+
739
+ ### The enforcement gate
740
+
741
+ The Kubernetes API can create a NetworkPolicy object without that policy
742
+ being enforced yet — CNIs (including `kind`'s default kindnet) program a new
743
+ pod's policy a few seconds _after_ the pod starts, not atomically with pod
744
+ creation. The delay is observable on real clusters and threatens every run,
745
+ not just the harness: the worker gate could otherwise
746
+ open on a pod whose isolation isn't active yet.
747
+
748
+ The fix, before seeding or opening the worker gate: the launcher execs into
749
+ the keeper (which shares the pod's network namespace with the worker) one
750
+ `node -e` probe that measures **both of the coding proxy's ports against the
751
+ proxy Service's ClusterIP, in the same pass** — `8787` (the proxy itself, which
752
+ the run policy permits) and `8788` (the deny port, which no run policy ever
753
+ permits). Only the outcome **(8787 connected, 8788 blocked)** counts toward the
754
+ streak.
755
+
756
+ Stated exactly, that outcome proves: **the SYN to 8788 was dropped somewhere on
757
+ the path, while the same destination answered on 8787.** That the drop was the
758
+ _run pod's own egress policy_ does not follow from the measurement alone — it
759
+ follows from the proxy admitting run pods on 8788 at its own ingress, so that no
760
+ other hop is left to drop it. That precondition used to be asserted by manifest
761
+ and checked nowhere, and was falsified on a live cluster: with the proxy's policy
762
+ admitting 8787 only, a prober with no policy at all and full internet egress read
763
+ **proven**, because the proxy's own ingress dropped the packet. The `proxy-service`
764
+ preflight check now reads the proxy's NetworkPolicy and verifies that rule, so the
765
+ attribution is checked rather than assumed.
766
+
767
+ Both halves are required, and measuring only one port would be unsound. A
768
+ NetworkPolicy denial **drops** the packet rather than rejecting it — GKE
769
+ Dataplane V2 (Cilium), which Autopilot runs, always drops — so "8788 did not
770
+ answer" on its own is equally consistent with "the proxy is gone and no policy
771
+ exists at all", and a run would be released onto an unpoliced network.
772
+ Requiring 8787 to connect in the _same_ exec turns "something is listening"
773
+ from a control-plane inference (which is stale the moment it is read) into a
774
+ fact this pod just observed, at the instant of the blocked observation. The
775
+ proxy's own policy deliberately **allows** ingress on 8788 from run pods:
776
+ ingress is enforced at the destination, so denying it there would make a run
777
+ pod whose own egress policy was not yet programmed read as "blocked" — which is
778
+ precisely the live falsification above, and why that rule is now verified at
779
+ preflight instead of trusted.
780
+
781
+ **"Blocked" means a timeout specifically, not "did not connect".** Pairing the
782
+ two ports only rules out "the whole proxy pod is dead"; it does not rule out
783
+ the deny _listener_ being unserved while the pod is otherwise healthy. This was
784
+ reproduced on a live cluster: in a namespace with no NetworkPolicy at all,
785
+ against a pod listening on 8787 and serving nothing on 8788, a probe that
786
+ treated any failed connect as "blocked" exited **proven** while it had full
787
+ internet egress. No control-plane read closes this — an Endpoints subset port is
788
+ the Service's numeric `targetPort`, not evidence that anything is bound — so the
789
+ probe distinguishes the two socket outcomes itself: a **timeout** means the
790
+ packet was dropped (a policy), while a **refusal** (RST / `ECONNREFUSED`) proves
791
+ the SYN reached the destination host, on every dataplane, since a drop cannot
792
+ produce an RST. A refused deny port is therefore reachable-but-unserved and is
793
+ reported as `kubernetes_policy_witness_unserved`, never as proven.
794
+
795
+ The intended consequence: on a **reject-style** CNI a genuine policy denial also
796
+ arrives as an RST, so such a cluster now fails closed here rather than passing
797
+ vacuously. That is the correct direction — a witness that cannot tell "denied"
798
+ from "unserved" is not a witness — and such a cluster needs a different one.
799
+
800
+ It requires **3 consecutive proven results, 500ms apart** (anything else
801
+ resets the streak — this guards against a single dropped SYN packet on an
802
+ allowed path being misread as "policy enforced"), bounded by
803
+ `enforcementTimeoutMs` (default 30,000ms — configurable via
804
+ `KubernetesJobLauncherOptions.enforcementTimeoutMs`; a drop-style CNI can
805
+ need close to this whole window). The verdict at the bound comes from the
806
+ _last_ probe — not from whether any probe was ever unavailable, so an early
807
+ blip while the pod's networking came up does not misdirect the operator — and
808
+ the three non-proven outcomes stay distinct, because each sends an operator
809
+ somewhere different:
810
+
811
+ | Last probe | Verdict | Where to look |
812
+ | ---------------- | --------------------------------------- | ---------------------------------------- |
813
+ | 8788 connected | `kubernetes_policy_not_enforced` | the CNI: no policy, or not port-scoped |
814
+ | 8787 unreachable | `kubernetes_policy_witness_unavailable` | the proxy pod / its Service |
815
+ | 8788 refused | `kubernetes_policy_witness_unserved` | the deny listener, or a reject-style CNI |
816
+
817
+ Any other exit code means the probe never ran to completion (a crash, a missing
818
+ interpreter, an OOM-killed keeper), so nothing was measured and nothing about
819
+ the policy can be concluded: that is `kubernetes_policy_probe_unusable`, and it
820
+ carries the observed exit code. Every one of these messages names the probed
821
+ address and the exit code it saw, because an operator reads the error, not this
822
+ page.
823
+
824
+ This wait happens inside the pod's overall ready-timeout window, not on top of
825
+ it.
826
+
827
+ `wardby coding preflight`'s canary pod waits the same way before running its
828
+ probes, for the same reason.
829
+
830
+ ### Preflight
831
+
832
+ `kubernetesPreflight` / `runKubernetesPreflight`
833
+ (`src/providers/jobs/kubernetes-preflight.ts`) run **five checks in order**,
834
+ each producing `kubernetes_isolation_unsupported:<check>` on failure (or
835
+ `:timeout` if the whole preflight — cleanup included — exceeds `timeoutMs`,
836
+ default 90,000ms):
837
+
838
+ 1. `platform` — pure configuration, checked before any cluster API call:
839
+ `assertPlatformConfig` (`src/providers/jobs/kubernetes-platform.ts`) refuses
840
+ a deployment that cannot work under `KUBERNETES_PLATFORM`. Under
841
+ `gke-autopilot` this means `KUBERNETES_RUNTIME_CLASS` must be `gvisor`
842
+ (gVisor is mandatory there — an unset or different runtime class is refused,
843
+ not warned about), and the effective `CODING_MAX_DISK_MB` must leave room
844
+ for the worker container's 1 GiB reservation inside Autopilot's 10 GiB pod
845
+ ephemeral-storage ceiling. The same assertion runs again at process
846
+ start-up in `buildConfiguredExecutor`
847
+ (`src/providers/executor/composition.ts`), so an unrunnable configuration
848
+ fails the process immediately rather than waiting for the first coding run
849
+ or the next `wardby coding preflight` invocation to discover it.
850
+ 2. `namespace` — the configured namespace exists.
851
+ 3. `proxy-service` — the proxy Service exists, has a ClusterIP, **exposes both
852
+ the proxy port (8787) and the deny port (8788)** over TCP, has at least one
853
+ ready endpoint serving both, **and the proxy's own NetworkPolicy admits
854
+ `wardby.io/component: coding-run` on 8788** — failing closed with
855
+ `kubernetes_isolation_unsupported:proxy-service` otherwise. That last clause
856
+ is the attribution precondition: a dropped connect to 8788 only indicts the
857
+ run pod's own egress policy if no other hop would have dropped it, and
858
+ ingress is enforced at the destination, so without it deleting one line from
859
+ an overlay's proxy policy makes every run read as enforced. The deny port is
860
+ the enforcement witness: a second listener on the proxy
861
+ (`src/providers/coding-proxy/deny-port.ts`) that serves nothing and that no
862
+ run's NetworkPolicy ever permits. A run pod that reaches 8787 but not 8788
863
+ has proven its policy is both programmed and port-scoped. It replaces the old
864
+ `cluster-dns` witness, which does not exist on GKE Autopilot (Cloud DNS is the
865
+ only provider there, so no kube-dns pods run) and which made the launcher read
866
+ `kube-system`.
867
+ 4. `worker-image` — `CODING_WORKER_IMAGE` is a registry digest.
868
+ 5. `canary` — creates a real run pod + NetworkPolicy from the same builders
869
+ as a live run, running a script that waits for policy enforcement (as
870
+ above) then attempts DNS resolution, a connect to the proxy's deny port,
871
+ the internet (`1.1.1.1:443`), the metadata server, and the proxy itself —
872
+ requiring every one of the first four to fail and the proxy connect to
873
+ succeed. Any other outcome, or a canary pod that itself fails to
874
+ schedule/run, is `kubernetes_isolation_unsupported:canary`.
875
+
876
+ This whole preflight is **memoized per launcher instance and its failure is
877
+ sticky**: `KubernetesJobLauncher.runPreflight()` caches the first call's
878
+ promise (`this.preflightResult ??= ...`, `kubernetes.ts:472-487`), including
879
+ a rejection — so once a launcher process has seen preflight fail, every
880
+ subsequent `launch()` in that process fails immediately with the same error
881
+ without re-probing the cluster. A fresh preflight requires a new process
882
+ (or, from the CLI, a fresh `wardby coding preflight` invocation, which is
883
+ not memoized).
884
+
885
+ **A hung pod create during preflight can leave a preflight pod and its
886
+ NetworkPolicy behind.** If `createPod` never settles (rather than failing),
887
+ the preflight's own timeout still fires and the caller sees `:timeout`, but
888
+ cleanup for that pod/policy is deferred to whenever the stuck create call
889
+ eventually resolves (`tracked`/`lateCleanup` in `kubernetes-preflight.ts`) —
890
+ if it never does, the objects are never removed automatically. **A canary
891
+ pod and policy are not distinguishable by name from a real run's:** both are
892
+ named `wardby-run-<runId's sha256 prefix>` (`kubernetesRunNames`,
893
+ `kubernetes-isolation.ts:90-96`, used by the preflight at
894
+ `kubernetes-preflight.ts:218,241`) and carry the same
895
+ `wardby.io/component: coding-run` label as a live run
896
+ (`kubernetes-isolation.ts:98-104,249`) — there is no
897
+ `wardby-run-preflight-*` naming pattern. After a `:timeout` failure,
898
+ operators should instead list every object with
899
+ `wardby.io/component=coding-run` in the namespace and cross-reference
900
+ against the run record ConfigMaps that legitimately exist (a stray canary
901
+ object has no corresponding non-tombstoned run record, since preflight
902
+ never creates one). Giving preflight objects a distinguishing label (e.g.
903
+ `wardby.io/component: coding-preflight`) would make this a direct label query;
904
+ see "Known limitations" below.
905
+
906
+ ### RBAC actually required
907
+
908
+ The launcher's `ClientNodeKubernetesApi` issues a narrow, specific set of
909
+ calls, and `deploy/kind-coding/manifests/base/` grants exactly that (no
910
+ `list`/`watch` on pods, no `get` on secrets — the seam never reads one
911
+ back):
912
+
913
+ - Namespace `Role` **`wardby-coding-launcher`** (in the coding namespace):
914
+ `pods` create/get/delete, `pods/exec` create/get, `pods/log` get,
915
+ `secrets` create/delete, `configmaps` create/get/update, `networkpolicies`
916
+ create/get/delete, `services` get, and `endpoints` get scoped by
917
+ `resourceNames: ["wardby-coding-proxy"]` — the one read `readProxyWitness`
918
+ needs to prove the deny port is exposed with a ready backend. **Nothing in
919
+ `kube-system` any more:** the old `wardby-coding-dns-reader` Role is gone
920
+ with the `cluster-dns` check.
921
+ - `ClusterRole` **`wardby-coding-namespace-reader`**: `get` on the
922
+ cluster-scoped `namespaces` resource, `resourceNames: [<the namespace>]`.
923
+ This one has to be cluster-scoped — no namespaced `Role` can grant `get`
924
+ on `namespaces` — but it's still scoped down to the one namespace via
925
+ `resourceNames`, so the launcher identity can't discover any other
926
+ namespace's existence.
927
+
928
+ (`deploy/kind-coding/` binds none of these to a service account — the local
929
+ harness runs every command against your own admin kubeconfig. A production
930
+ overlay, e.g. GKE, binds these two to the control plane's identity.)
931
+
932
+ **Running the real-cluster integration suite needs more than this.**
933
+ `npm run test:kubernetes` (`kubernetes.integration.test.ts`) uses a raw
934
+ `@kubernetes/client-node` client directly, alongside the launcher's own
935
+ `KubernetesApi` seam, for two things outside what the launcher itself ever
936
+ does: it reads `Endpoints` objects (`get endpoints`) to resolve kube-dns's
937
+ and the proxy pod's addresses for its isolation probes, and its cleanup
938
+ deletes each run's record ConfigMap (`delete configmaps`) — the launcher
939
+ intentionally never deletes that object (see "kept as tombstones" above), so
940
+ `deleteConfigMap` isn't even part of the `KubernetesApi` seam; the test goes
941
+ straight to the library. The committed `wardby-coding-launcher` Role grants
942
+ neither verb. On `kind` this gap is invisible because the suite runs against
943
+ the admin kubeconfig; a kubeconfig scoped to only the two launcher roles
944
+ above needs `get endpoints` (coding namespace and `kube-system`) and
945
+ `delete configmaps` (coding namespace) added before the integration suite
946
+ will pass against it.
947
+
948
+ ### Diagnostics
949
+
950
+ On a failed run, the launcher reads only the failed worker container's last
951
+ 8 log lines (bounded to 128 KiB, room for eight full-size
952
+ [debug trace](#debug-trace) lines) through the Kubernetes API
953
+ (`pods/log`), and keeps only a code matching the existing
954
+ `SAFE_WORKER_DIAGNOSTIC` pattern (imported from the Docker launcher) — the
955
+ raw text itself is never stored, logged, or returned
956
+ (`readWorkerDiagnostic`, `kubernetes.ts`). For `coding_output_invalid`, the
957
+ same line may also carry `issues`: up to 8 `path:code` entries (for example
958
+ `tag:invalid_string`) naming which output-schema fields failed, never their
959
+ values. The list is kept only when every entry matches
960
+ `SAFE_CODING_OUTPUT_ISSUE`, and the executor logs it next to the diagnostic id.
961
+
962
+ ### Known limitations
963
+
964
+ - **No gVisor / per-pod process (PID) limit on `kind`.** `kind` has no
965
+ runtime-class sandboxing; `KUBERNETES_RUNTIME_CLASS` is unset in the local
966
+ harness and the launcher logs `kubernetes_runtime_class_unset` once per
967
+ launch as a loud "development cluster" warning. Use the GKE Autopilot
968
+ overlay with its required `gvisor` runtime class for production.
969
+ - **The real-cluster integration coverage is partial.** It covers the contract
970
+ suite against the real API, the enforcement gate, and a full run lifecycle.
971
+ It does not yet cover the isolation acceptance suite's
972
+ OOM/disk-full/wall-clock containment assertions, "canary fails when a
973
+ policy is removed", or Claude Code's tool-runner sidecar (there is no
974
+ separate tool pod: it shares the run pod's network namespace and
975
+ NetworkPolicy with the worker, so the coverage needed is different from a
976
+ second pod's isolation).
977
+ - **No end-to-end backpressure or host-side archive-size cap** once the
978
+ exec WebSocket for a seed/collect transfer is connected — pre-existing,
979
+ documented, not a regression of this milestone.
980
+ - **An over-ceiling `workspaceDiskMb` fails late and generically** — see
981
+ "Control Plane Configuration" above.
982
+ - **Namespace handling must remain explicit in every overlay.**
983
+ `deploy/kind-coding/manifests/base/kustomization.yaml`
984
+ deliberately has **no top-level `namespace:` override**, because
985
+ kustomize's namespace transformer would force `metadata.namespace` onto
986
+ every namespaced resource it lists. Every manifest instead sets its own
987
+ `metadata.namespace` explicitly. Any overlay author copying this harness for
988
+ another cluster must do the same.
989
+ - **Per-run record ConfigMap GC is unimplemented.** Both `remove()`'s
990
+ deliberate tombstones and a failed launch's records (see above)
991
+ accumulate, one small ConfigMap per run, with nothing that automatically
992
+ deletes them. Operators need a retention job for a long-lived deployment;
993
+ this is an operational concern rather than a correctness or security issue.
994
+ - **Autopilot dry-run capture is a lower bound.** Admission dry runs cannot
995
+ observe topology labels added after pod scheduling. The reviewed platform
996
+ profile and tests include those known labels, but every target cluster must
997
+ still pass preflight before accepting work.
998
+ - **A distinguishing label for preflight/canary objects.** Canary pods and
999
+ policies are currently named and labeled identically to real run objects
1000
+ (`wardby.io/component: coding-run`), which is why the troubleshooting
1001
+ guidance above can only recommend cross-referencing against run records
1002
+ rather than a direct label query. Giving preflight objects their own
1003
+ `wardby.io/component: coding-preflight` label would fix this and would
1004
+ also let a future orphan reaper (the ConfigMap GC item above, extended to
1005
+ pods) tell a canary apart from a live run.
1006
+ - **The real-cluster integration suite still uses the deprecated core/v1
1007
+ `Endpoints` API** (`kubernetes.integration.test.ts`) rather than
1008
+ `discovery.k8s.io/v1` `EndpointSlice`. `Endpoints` is deprecated, not yet
1009
+ removed, and this is test-only code, but the migration is a tracked
1010
+ remaining compatibility task.
1011
+
1012
+ ### Troubleshooting: failure codes
1013
+
1014
+ | Code | Meaning |
1015
+ | ---------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
1016
+ | `kubernetes_isolation_unsupported:<check>` | A preflight check failed; `<check>` is one of `platform`, `namespace`, `proxy-service`, `worker-image`, `canary`. Sticky for the launcher's process lifetime once seen (see Preflight above). |
1017
+ | `kubernetes_isolation_unsupported:timeout` | The whole preflight (including cleanup) exceeded its timeout. |
1018
+ | `kubernetes_isolation_unsupported` | (No suffix) Attestation failure: the read-back pod or NetworkPolicy didn't canonically match the builder's output. |
1019
+ | `kubernetes_policy_not_enforced` | The run's NetworkPolicy wasn't observed enforced (8787 reachable, 8788 blocked) within `enforcementTimeoutMs`; the worker gate was never opened. |
1020
+ | `kubernetes_policy_witness_unavailable` | The last enforcement probe could not reach the proxy on 8787 at all, so nothing could be witnessed — the proxy or its Service is the thing to check, not the CNI. The worker gate was never opened. |
1021
+ | `kubernetes_policy_witness_unserved` | The last enforcement probe was **refused** (RST) on 8788 rather than dropped. The packet reached the host, so nothing is blocking the path and the witness proves nothing — the deny listener is not serving, or the CNI rejects instead of dropping. Fails closed; the worker gate was never opened. |
1022
+ | `kubernetes_proxy_witness_unusable: <reason>` | The proxy Service is not a usable enforcement witness (missing, no ClusterIP, a required port not exposed as TCP, or no ready endpoint serving both 8787 and 8788). Seen **unwrapped** like this from `provision`'s per-launch re-read, which runs after an earlier witness check already passed -- a supplied `preflight`, or this launcher's own memoized read, the likelier production sighting being a proxy that degrades after that read succeeded; the launcher's own memoized read reports the same condition wrapped as `kubernetes_isolation_unsupported` (with this string on `cause`), and the preflight reports it as `:proxy-service`. Fails closed before the pod is created. |
1023
+ | `kubernetes_pod_start_timeout` | The keeper didn't become ready within `readyTimeoutMs` (default 120s). Historically caused by the subPath root-ownership issue the `storage-init` init container now fixes; if seen again, check init-container status first. |
1024
+ | `kubernetes_pod_start_failed` | The pod (or its `storage-init` init container) failed outright rather than timing out. |
1025
+ | `kubernetes_tool_runner_failed` | (Claude Code only) The tool runner sidecar restarted, exited, or its image could not be pulled before the keeper started. A malformed `WARDBY_TOOL_SETUP` makes the tool runner exit before it is ready (`tool_setup_invalid` in its log), which surfaces as this code on Kubernetes and as `docker_tool_runner_not_ready` on Docker. |
1026
+ | `kubernetes_tool_runner_unready` | (Claude Code only) The tool runner sidecar never started (its socket startup probe never passed) by the pod-start bound (`readyTimeoutMs`). |
1027
+ | `worker_tool_runner_failed` | (Claude Code only; also seen on the Docker launcher) The tool runner died mid-run, after the pod/containers started successfully. |
1028
+ | `worker_tool_runner_unreachable` | (Claude Code only; both launchers) The Claude worker could not connect to the tool runner's socket before calling the model, or Claude Code reported its command tool as not connected. On Kubernetes, check that the pod has the `tool-socket` volume mounted in both the `worker` and `tool-runner` containers. |
1029
+ | `claude_tool_setup_too_large` | (Claude Code only) The run's tool-runner setup (registry settings plus every service test variable) would exceed the tool runner's bounds (1024 variables, 4096 bytes per value, 96 KiB in total), so the launch fails before the tool runner is created. |
1030
+ | `kubernetes_isolation_unsupported:tool-image-not-registry-digest` | (Claude Code only) `CODING_CLAUDE_TOOL_RUNNER_IMAGE` isn't a registry digest. Checked on every launch, not only preflight. |
1031
+ | `kubernetes_isolation_unsupported:claude-limits-too-small` | (Claude Code only) The run's `cpus`/`memoryMb` are too small to leave the tool runner its fixed floor: Claude Code needs `cpus ≥ 0.35` and `memoryMb ≥ 256`. |
1032
+ | `kubernetes_seed_failed` | Streaming the workspace or input artifact into the keeper failed. |
1033
+ | `kubernetes_workspace_archive_failed` / `kubernetes_result_artifact_invalid` | Collection (workspace or output artifact) failed or was invalid. |
1034
+ | `coding_workspace_disk_exceeds_limit` | The run's `workspaceDiskMb` exceeds `CODING_MAX_DISK_MB`; surfaces on the run record as `coding_failure_workspace:<id>`. |
1035
+ | `coding_service_unready:<name>` | A service sidecar restarted after its startup probe gave up, could not be pulled or started, or had not started when `readyTimeoutMs` ran out. The run fails with category `service_unready`. |
1036
+
1037
+ ### Local harness
1038
+
1039
+ `deploy/kind-coding/` stands up a local `kind` cluster (with a local image
1040
+ registry so images are pulled by digest, as on GKE) that proves this
1041
+ launcher end to end, including real NetworkPolicy enforcement. See
1042
+ [`deploy/kind-coding/README.md`](../deploy/kind-coding/README.md) for
1043
+ prerequisites, the up/down scripts, and what each step does.
1044
+
1045
+ ## Verification
1046
+
1047
+ Build the image and run the destructive, self-cleaning acceptance suite:
1048
+
1049
+ ```sh
1050
+ docker build -f src/coding-worker/Dockerfile -t wardby-coding-worker:task8 .
1051
+ npm run test:docker-isolation
1052
+ npm run verify:claude-code
1053
+ ```
1054
+
1055
+ Set `WARDBY_WORKER_IMAGE` to test another local tag. The runner resolves that
1056
+ tag to an immutable image ID before testing. The suite verifies effective
1057
+ Docker inspection, no default route, proxy-only connectivity, denied
1058
+ Docker-socket/host/metadata/localhost/public access, read-only mounts and
1059
+ rootfs, zero effective capabilities, seccomp and no-new-privileges, private PID
1060
+ 1, PID exhaustion, OOM containment, disk ENOSPC, and wall-clock termination.
1061
+
1062
+ Related Docker references:
1063
+
1064
+ - Internal and isolated bridge networks: https://docs.docker.com/reference/cli/docker/network/create/
1065
+ - CPU, memory, swap, and PID controls: https://docs.docker.com/engine/containers/resource_constraints/
1066
+ - Seccomp and no-new-privileges: https://docs.docker.com/reference/cli/docker/container/run/
1067
+ - Tmpfs behavior and limits: https://docs.docker.com/engine/storage/tmpfs/
1068
+ - Volume subpaths and `nocopy`: https://docs.docker.com/engine/storage/volumes/