@wardby/cli 0.2.1 → 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.env.example +34 -4
- package/README.md +41 -4
- package/dist/claude-coding-worker/driver.d.ts +6 -1
- package/dist/claude-coding-worker/driver.js +27 -1
- package/dist/claude-coding-worker/main.js +2 -0
- package/dist/claude-coding-worker/tool-socket.d.ts +9 -0
- package/dist/claude-coding-worker/tool-socket.js +26 -0
- package/dist/cli-help.d.ts +1 -1
- package/dist/cli-help.js +13 -4
- package/dist/cli.d.ts +1 -1
- package/dist/cli.js +219 -29
- package/dist/coding/collect-exclude.d.ts +25 -0
- package/dist/coding/collect-exclude.js +76 -0
- package/dist/coding/profile.d.ts +67 -22
- package/dist/coding/profile.js +58 -25
- package/dist/coding/protected-path-wording.d.ts +24 -0
- package/dist/coding/protected-path-wording.js +55 -0
- package/dist/coding/protected-paths.d.ts +31 -0
- package/dist/coding/protected-paths.js +48 -0
- package/dist/coding/protocol.d.ts +75 -4
- package/dist/coding/protocol.js +107 -7
- package/dist/coding/registry/adapters.d.ts +2 -0
- package/dist/coding/registry/adapters.js +9 -0
- package/dist/coding/registry/allowlist.d.ts +9 -0
- package/dist/coding/registry/allowlist.js +28 -0
- package/dist/coding/registry/json-scan.d.ts +55 -0
- package/dist/coding/registry/json-scan.js +191 -0
- package/dist/coding/registry/lockfiles.d.ts +6 -0
- package/dist/coding/registry/lockfiles.js +89 -0
- package/dist/coding/registry/npm-lockfile.d.ts +9 -0
- package/dist/coding/registry/npm-lockfile.js +73 -0
- package/dist/coding/registry/npm-plan.d.ts +29 -0
- package/dist/coding/registry/npm-plan.js +289 -0
- package/dist/coding/registry/npm.d.ts +7 -0
- package/dist/coding/registry/npm.js +252 -0
- package/dist/coding/registry/pypi.d.ts +4 -0
- package/dist/coding/registry/pypi.js +210 -0
- package/dist/coding/registry/report.d.ts +37 -0
- package/dist/coding/registry/report.js +31 -0
- package/dist/coding/registry/token.d.ts +4 -0
- package/dist/coding/registry/token.js +7 -0
- package/dist/coding/registry/types.d.ts +286 -0
- package/dist/coding/registry/types.js +19 -0
- package/dist/coding/registry/worker-config.d.ts +14 -0
- package/dist/coding/registry/worker-config.js +31 -0
- package/dist/coding/services/builtins.d.ts +22 -0
- package/dist/coding/services/builtins.js +85 -0
- package/dist/coding/services/catalog.d.ts +443 -0
- package/dist/coding/services/catalog.js +159 -0
- package/dist/coding/services/declaration.d.ts +11 -0
- package/dist/coding/services/declaration.js +114 -0
- package/dist/coding/services/note.d.ts +8 -0
- package/dist/coding/services/note.js +15 -0
- package/dist/coding/services/resolve.d.ts +25 -0
- package/dist/coding/services/resolve.js +50 -0
- package/dist/coding/services/wording.d.ts +25 -0
- package/dist/coding/services/wording.js +70 -0
- package/dist/coding-proxy/main.js +24 -4
- package/dist/coding-worker/artifact.d.ts +6 -0
- package/dist/coding-worker/debug-trace.d.ts +37 -0
- package/dist/coding-worker/debug-trace.js +117 -0
- package/dist/coding-worker/driver.d.ts +13 -1
- package/dist/coding-worker/driver.js +102 -13
- package/dist/coding-worker/errors.js +8 -0
- package/dist/coding-worker/main.js +10 -2
- package/dist/coding-worker/sdk.d.ts +2 -2
- package/dist/coding-worker/sdk.js +7 -2
- package/dist/coding-worker/types.d.ts +3 -0
- package/dist/config/providers.d.ts +41 -0
- package/dist/config/providers.js +66 -0
- package/dist/core/budget-groups.d.ts +108 -5
- package/dist/core/budget-groups.js +125 -19
- package/dist/core/budget-wording.d.ts +22 -0
- package/dist/core/budget-wording.js +55 -0
- package/dist/core/datastores.js +18 -2
- package/dist/core/db.d.ts +5 -1
- package/dist/core/db.js +8 -2
- package/dist/core/dispatch.d.ts +33 -6
- package/dist/core/dispatch.js +260 -43
- package/dist/core/engine-native.js +27 -10
- package/dist/core/grants.d.ts +86 -0
- package/dist/core/grants.js +126 -0
- package/dist/core/host-events.d.ts +49 -0
- package/dist/core/host-events.js +257 -0
- package/dist/core/host-identity-links.d.ts +54 -0
- package/dist/core/host-identity-links.js +189 -0
- package/dist/core/host-status.d.ts +61 -0
- package/dist/core/host-status.js +211 -0
- package/dist/core/in-flight-runs.d.ts +8 -0
- package/dist/core/in-flight-runs.js +56 -0
- package/dist/core/provider-wording.d.ts +12 -0
- package/dist/core/provider-wording.js +44 -0
- package/dist/core/reconciler.d.ts +38 -2
- package/dist/core/reconciler.js +86 -2
- package/dist/core/repo-access.d.ts +99 -0
- package/dist/core/repo-access.js +136 -0
- package/dist/core/review-host-checks.d.ts +11 -0
- package/dist/core/review-host-checks.js +41 -0
- package/dist/core/review-host-tools.d.ts +47 -0
- package/dist/core/review-host-tools.js +347 -0
- package/dist/core/run-heartbeat.d.ts +27 -0
- package/dist/core/run-heartbeat.js +54 -0
- package/dist/core/runner.d.ts +30 -8
- package/dist/core/runner.js +237 -40
- package/dist/core/secrets.d.ts +11 -2
- package/dist/core/secrets.js +25 -6
- package/dist/core/subagent-memory-tools.d.ts +1 -1
- package/dist/core/subagent-memory-tools.js +16 -3
- package/dist/core/tool-admin.d.ts +81 -0
- package/dist/core/tool-admin.js +129 -0
- package/dist/core/tool-names.d.ts +42 -0
- package/dist/core/tool-names.js +64 -0
- package/dist/core/untrusted-content.d.ts +32 -0
- package/dist/core/untrusted-content.js +72 -0
- package/dist/core/webhooks.d.ts +8 -1
- package/dist/core/webhooks.js +23 -2
- package/dist/generated/prisma/browser.d.ts +100 -0
- package/dist/generated/prisma/client.d.ts +100 -0
- package/dist/generated/prisma/commonInputTypes.d.ts +30 -0
- package/dist/generated/prisma/enums.d.ts +6 -0
- package/dist/generated/prisma/enums.js +6 -1
- package/dist/generated/prisma/internal/class.d.ts +143 -0
- package/dist/generated/prisma/internal/class.js +4 -4
- package/dist/generated/prisma/internal/prismaNamespace.d.ts +1746 -577
- package/dist/generated/prisma/internal/prismaNamespace.js +179 -6
- package/dist/generated/prisma/internal/prismaNamespaceBrowser.d.ts +186 -0
- package/dist/generated/prisma/internal/prismaNamespaceBrowser.js +179 -6
- package/dist/generated/prisma/models/Agent.d.ts +294 -1
- package/dist/generated/prisma/models/AgentRepository.d.ts +1425 -0
- package/dist/generated/prisma/models/AgentRepository.js +1 -0
- package/dist/generated/prisma/models/AgentTool.d.ts +95 -1
- package/dist/generated/prisma/models/AuthUser.d.ts +56 -1
- package/dist/generated/prisma/models/CodingAgentProfile.d.ts +258 -7
- package/dist/generated/prisma/models/CodingProxySession.d.ts +123 -2
- package/dist/generated/prisma/models/CodingRun.d.ts +1276 -95
- package/dist/generated/prisma/models/CodingService.d.ts +1348 -0
- package/dist/generated/prisma/models/CodingService.js +1 -0
- package/dist/generated/prisma/models/HostEventDelivery.d.ts +946 -0
- package/dist/generated/prisma/models/HostEventDelivery.js +1 -0
- package/dist/generated/prisma/models/HostIdentity.d.ts +1232 -0
- package/dist/generated/prisma/models/HostIdentity.js +1 -0
- package/dist/generated/prisma/models/HostIdentityLinkRequest.d.ts +1473 -0
- package/dist/generated/prisma/models/HostIdentityLinkRequest.js +1 -0
- package/dist/generated/prisma/models/Principal.d.ts +455 -0
- package/dist/generated/prisma/models/RegistryAllowance.d.ts +1148 -0
- package/dist/generated/prisma/models/RegistryAllowance.js +1 -0
- package/dist/generated/prisma/models/RegistryApprovedVersion.d.ts +1219 -0
- package/dist/generated/prisma/models/RegistryApprovedVersion.js +1 -0
- package/dist/generated/prisma/models/RegistryFetch.d.ts +1428 -0
- package/dist/generated/prisma/models/RegistryFetch.js +1 -0
- package/dist/generated/prisma/models/RegistryPlanRefusal.d.ts +1294 -0
- package/dist/generated/prisma/models/RegistryPlanRefusal.js +1 -0
- package/dist/generated/prisma/models/RegistryVersionFact.d.ts +1085 -0
- package/dist/generated/prisma/models/RegistryVersionFact.js +1 -0
- package/dist/generated/prisma/models/ResourceGrant.d.ts +1437 -0
- package/dist/generated/prisma/models/ResourceGrant.js +1 -0
- package/dist/generated/prisma/models/Run.d.ts +389 -1
- package/dist/generated/prisma/models/RunHostCheck.d.ts +1239 -0
- package/dist/generated/prisma/models/RunHostCheck.js +1 -0
- package/dist/generated/prisma/models/RunHostStatus.d.ts +1315 -0
- package/dist/generated/prisma/models/RunHostStatus.js +1 -0
- package/dist/generated/prisma/models/Tool.d.ts +15 -3
- package/dist/generated/prisma/models.d.ts +13 -0
- package/dist/help/build.d.ts +1 -0
- package/dist/help/build.js +9 -0
- package/dist/help/catalog.d.ts +24 -0
- package/dist/help/catalog.js +160 -0
- package/dist/help/cli.d.ts +2 -0
- package/dist/help/cli.js +65 -0
- package/dist/help/runtime.d.ts +3 -0
- package/dist/help/runtime.js +44 -0
- package/dist/help/search.d.ts +9 -0
- package/dist/help/search.js +104 -0
- package/dist/help-index.json +660 -0
- package/dist/import/cli-args.js +3 -2
- package/dist/import/create.d.ts +5 -0
- package/dist/import/create.js +67 -13
- package/dist/import/index.js +19 -9
- package/dist/import/neutral-schema.d.ts +11 -11
- package/dist/mcp/auth/access.d.ts +70 -0
- package/dist/mcp/auth/access.js +58 -0
- package/dist/mcp/auth/grants-cli.d.ts +149 -0
- package/dist/mcp/auth/grants-cli.js +518 -0
- package/dist/mcp/auth/host-account-cli.d.ts +2 -0
- package/dist/mcp/auth/host-account-cli.js +47 -0
- package/dist/mcp/auth/ownership.d.ts +26 -52
- package/dist/mcp/auth/ownership.js +19 -14
- package/dist/mcp/auth/repo-authorization.d.ts +22 -0
- package/dist/mcp/auth/repo-authorization.js +51 -0
- package/dist/mcp/auth/resource-server.d.ts +31 -2
- package/dist/mcp/auth/resource-server.js +84 -3
- package/dist/mcp/auth/self-hosted/browser.js +2 -2
- package/dist/mcp/auth/self-hosted/cli.js +47 -13
- package/dist/mcp/auth/self-hosted/credentials.d.ts +20 -2
- package/dist/mcp/auth/self-hosted/credentials.js +62 -4
- package/dist/mcp/auth/self-hosted/session.d.ts +2 -1
- package/dist/mcp/context.d.ts +27 -1
- package/dist/mcp/errors.d.ts +25 -6
- package/dist/mcp/errors.js +98 -0
- package/dist/mcp/host-events/github-ingress.d.ts +36 -0
- package/dist/mcp/host-events/github-ingress.js +92 -0
- package/dist/mcp/host-events/github-user-callback.d.ts +18 -0
- package/dist/mcp/host-events/github-user-callback.js +50 -0
- package/dist/mcp/index.d.ts +8 -1
- package/dist/mcp/index.js +106 -16
- package/dist/mcp/server.js +27 -9
- package/dist/mcp/tools/agents.js +446 -48
- package/dist/mcp/tools/budget-groups.js +3 -3
- package/dist/mcp/tools/datastore.js +22 -15
- package/dist/mcp/tools/grants.d.ts +2 -0
- package/dist/mcp/tools/grants.js +239 -0
- package/dist/mcp/tools/help.d.ts +5 -0
- package/dist/mcp/tools/help.js +67 -0
- package/dist/mcp/tools/host-accounts.d.ts +2 -0
- package/dist/mcp/tools/host-accounts.js +114 -0
- package/dist/mcp/tools/memory.d.ts +7 -1
- package/dist/mcp/tools/memory.js +5 -5
- package/dist/mcp/tools/repositories.d.ts +2 -0
- package/dist/mcp/tools/repositories.js +181 -0
- package/dist/mcp/tools/runs.d.ts +6 -0
- package/dist/mcp/tools/runs.js +60 -6
- package/dist/mcp/tools/scheduling.js +3 -6
- package/dist/mcp/tools/secrets.js +18 -6
- package/dist/mcp/tools/services.d.ts +2 -0
- package/dist/mcp/tools/services.js +222 -0
- package/dist/mcp/tools/subagents.js +52 -16
- package/dist/mcp/tools/tools.d.ts +41 -0
- package/dist/mcp/tools/tools.js +201 -40
- package/dist/mcp/tools/trigger.js +36 -6
- package/dist/mcp/tools/webhooks.js +15 -4
- package/dist/mcp/transport/streamable-http.d.ts +10 -0
- package/dist/mcp/transport/streamable-http.js +32 -2
- package/dist/providers/auth/delegating.d.ts +19 -0
- package/dist/providers/auth/delegating.js +71 -0
- package/dist/providers/auth/index.d.ts +2 -0
- package/dist/providers/auth/index.js +11 -0
- package/dist/providers/auth/self-hosted.d.ts +2 -1
- package/dist/providers/auth/self-hosted.js +16 -4
- package/dist/providers/auth/types.d.ts +6 -0
- package/dist/providers/coding-proxy/memory-ledger.d.ts +3 -0
- package/dist/providers/coding-proxy/memory-ledger.js +21 -2
- package/dist/providers/coding-proxy/metering.d.ts +4 -0
- package/dist/providers/coding-proxy/metering.js +28 -2
- package/dist/providers/coding-proxy/prisma-ledger.d.ts +17 -0
- package/dist/providers/coding-proxy/prisma-ledger.js +50 -5
- package/dist/providers/coding-proxy/proxy.d.ts +13 -0
- package/dist/providers/coding-proxy/proxy.js +479 -18
- package/dist/providers/coding-proxy/registry/audit.d.ts +131 -0
- package/dist/providers/coding-proxy/registry/audit.js +380 -0
- package/dist/providers/coding-proxy/registry/bounded-fetch.d.ts +16 -0
- package/dist/providers/coding-proxy/registry/bounded-fetch.js +77 -0
- package/dist/providers/coding-proxy/registry/plan.d.ts +58 -0
- package/dist/providers/coding-proxy/registry/plan.js +304 -0
- package/dist/providers/coding-proxy/registry/prisma-store.d.ts +34 -0
- package/dist/providers/coding-proxy/registry/prisma-store.js +151 -0
- package/dist/providers/coding-proxy/registry/service.d.ts +213 -0
- package/dist/providers/coding-proxy/registry/service.js +1137 -0
- package/dist/providers/coding-proxy/registry/store.d.ts +127 -0
- package/dist/providers/coding-proxy/registry/store.js +70 -0
- package/dist/providers/coding-proxy/runtime.d.ts +17 -0
- package/dist/providers/coding-proxy/runtime.js +63 -0
- package/dist/providers/coding-proxy/secure-fetch.js +0 -1
- package/dist/providers/coding-proxy/server.d.ts +11 -0
- package/dist/providers/coding-proxy/server.js +163 -0
- package/dist/providers/coding-proxy/types.d.ts +15 -0
- package/dist/providers/engine/types.d.ts +10 -0
- package/dist/providers/executor/build.d.ts +2 -2
- package/dist/providers/executor/composition.d.ts +3 -0
- package/dist/providers/executor/composition.js +19 -0
- package/dist/providers/executor/container.d.ts +105 -3
- package/dist/providers/executor/container.js +296 -20
- package/dist/providers/executor/dbos.d.ts +2 -3
- package/dist/providers/executor/in-process.d.ts +2 -3
- package/dist/providers/executor/routing.d.ts +7 -0
- package/dist/providers/executor/routing.js +13 -0
- package/dist/providers/executor/types.d.ts +22 -0
- package/dist/providers/jobs/claude-tool-setup.d.ts +23 -0
- package/dist/providers/jobs/claude-tool-setup.js +50 -0
- package/dist/providers/jobs/collect-prune.d.ts +7 -0
- package/dist/providers/jobs/collect-prune.js +27 -0
- package/dist/providers/jobs/docker-isolation.d.ts +48 -3
- package/dist/providers/jobs/docker-isolation.js +213 -22
- package/dist/providers/jobs/docker-services.d.ts +35 -0
- package/dist/providers/jobs/docker-services.js +191 -0
- package/dist/providers/jobs/docker.d.ts +49 -2
- package/dist/providers/jobs/docker.js +279 -32
- package/dist/providers/jobs/fake-kubernetes-api.d.ts +1 -0
- package/dist/providers/jobs/fake-kubernetes-api.js +12 -3
- package/dist/providers/jobs/kubernetes-isolation.d.ts +17 -2
- package/dist/providers/jobs/kubernetes-isolation.js +238 -57
- package/dist/providers/jobs/kubernetes-platform.d.ts +10 -4
- package/dist/providers/jobs/kubernetes-platform.js +11 -5
- package/dist/providers/jobs/kubernetes-preflight.js +3 -0
- package/dist/providers/jobs/kubernetes.d.ts +21 -2
- package/dist/providers/jobs/kubernetes.js +112 -21
- package/dist/providers/jobs/types.d.ts +18 -0
- package/dist/providers/llm/anthropic.js +6 -2
- package/dist/providers/llm/bedrock.js +2 -1
- package/dist/providers/llm/claude-messages.d.ts +5 -1
- package/dist/providers/llm/claude-messages.js +1 -0
- package/dist/providers/llm/claude-provider.d.ts +3 -1
- package/dist/providers/llm/claude-provider.js +5 -1
- package/dist/providers/llm/index.d.ts +1 -1
- package/dist/providers/llm/index.js +1 -1
- package/dist/providers/llm/pricing-anthropic.d.ts +10 -0
- package/dist/providers/llm/pricing-anthropic.js +32 -14
- package/dist/providers/llm/pricing-bedrock-claude.d.ts +8 -0
- package/dist/providers/llm/pricing-bedrock-claude.js +9 -0
- package/dist/providers/llm/routing.d.ts +10 -1
- package/dist/providers/llm/routing.js +18 -0
- package/dist/providers/llm/types.d.ts +12 -0
- package/dist/providers/llm/types.js +8 -1
- package/dist/providers/review-host/diff-lines.d.ts +16 -0
- package/dist/providers/review-host/diff-lines.js +59 -0
- package/dist/providers/review-host/github-events.d.ts +6 -0
- package/dist/providers/review-host/github-events.js +180 -0
- package/dist/providers/review-host/github-user-auth.d.ts +38 -0
- package/dist/providers/review-host/github-user-auth.js +128 -0
- package/dist/providers/review-host/github.d.ts +57 -0
- package/dist/providers/review-host/github.js +567 -0
- package/dist/providers/review-host/index.d.ts +12 -0
- package/dist/providers/review-host/index.js +30 -0
- package/dist/providers/review-host/review-format.d.ts +19 -0
- package/dist/providers/review-host/review-format.js +59 -0
- package/dist/providers/review-host/types.d.ts +284 -0
- package/dist/providers/review-host/types.js +24 -0
- package/dist/providers/vcs/git.d.ts +16 -6
- package/dist/providers/vcs/git.js +77 -13
- package/dist/providers/vcs/github.d.ts +74 -3
- package/dist/providers/vcs/github.js +195 -15
- package/dist/providers/vcs/types.d.ts +42 -4
- package/dist/quickstart/index.js +13 -2
- package/dist/sandbox/fetch-policy.d.ts +20 -2
- package/dist/sandbox/fetch-policy.js +64 -3
- package/dist/sandbox/host-functions.d.ts +9 -1
- package/dist/sandbox/host-functions.js +12 -5
- package/dist/serve.js +1 -1
- package/dist/wardby-bin.js +6 -0
- package/docs/README.md +30 -0
- package/docs/architecture-runtime.md +90 -0
- package/docs/assets/brand/wardby-icon-512.png +0 -0
- package/docs/assets/brand/wardby-icon.svg +16 -0
- package/docs/assets/brand/wardby-mascot-profile-512.png +0 -0
- package/docs/assets/brand/wardby-mascot.png +0 -0
- package/docs/assets/brand/wardby-mascot.svg +5 -0
- package/docs/assets/wardby-workflow.svg +106 -0
- package/docs/code-review-agents.md +481 -0
- package/docs/coding-agent-setup.md +172 -0
- package/docs/coding-packages.md +455 -0
- package/docs/coding-services.md +300 -0
- package/docs/coding-worker-byo-images.md +98 -0
- package/docs/coding-worker-isolation.md +1061 -0
- package/docs/getting-started-gke.md +615 -0
- package/docs/getting-started-identity-provider.md +308 -0
- package/docs/getting-started.md +133 -0
- package/docs/observability.md +53 -0
- package/docs/release-verification.md +66 -0
- package/docs/security-deployment.md +657 -0
- package/help/code-review-agents.md +31 -0
- package/help/coding-packages.md +31 -0
- package/help/coding-services.md +71 -0
- package/help/creating-agents.md +68 -0
- package/help/deploy-gke.md +39 -0
- package/help/deployment-targets.md +39 -0
- package/help/errors/budget-group-exhausted.md +25 -0
- package/help/errors/docker-isolation-unsupported.md +26 -0
- package/help/errors/protected-path.md +45 -0
- package/help/errors/repo-access.md +27 -0
- package/help/errors/service-declaration-invalid.md +27 -0
- package/help/errors/service-declaration-unavailable.md +25 -0
- package/help/errors/service-launcher-unsupported.md +30 -0
- package/help/errors/service-not-allowed.md +27 -0
- package/help/errors/service-unknown.md +23 -0
- package/help/errors/service-unready.md +49 -0
- package/help/getting-started.md +31 -0
- package/help/github.md +30 -0
- package/help/identity-and-access.md +39 -0
- package/help/mcp.md +30 -0
- package/help/native-capabilities.md +30 -0
- package/help/observability.md +41 -0
- package/help/operating-agents.md +30 -0
- package/help/security.md +27 -0
- package/help/troubleshooting/budgets.md +26 -0
- package/help/troubleshooting/coding-workers.md +47 -0
- package/help/troubleshooting/repository-access.md +27 -0
- package/package.json +12 -3
- package/prisma/migrations/20260925010000_coding_collect_exclude/migration.sql +8 -0
- package/prisma/migrations/20260925015000_allowed_egress_default/migration.sql +8 -0
- package/prisma/migrations/20260925020000_coding_package_registry/migration.sql +55 -0
- package/prisma/migrations/20260925030000_tool_name_per_owner/migration.sql +11 -0
- package/prisma/migrations/20260926010000_run_trigger_host_event/migration.sql +7 -0
- package/prisma/migrations/20260926020000_code_review_hosts/migration.sql +53 -0
- package/prisma/migrations/20260926030000_auth_user_roles/migration.sql +9 -0
- package/prisma/migrations/20260926040000_agent_effort/migration.sql +6 -0
- package/prisma/migrations/20260926050000_registry_lockfile_plan/migration.sql +32 -0
- package/prisma/migrations/20260926050000_repo_access_authorization/migration.sql +82 -0
- package/prisma/migrations/20260926060000_registry_plan_refusal/migration.sql +19 -0
- package/prisma/migrations/20260926100000_registry_plan_refusal_published_at/migration.sql +5 -0
- package/prisma/migrations/20260926190000_run_host_status/migration.sql +18 -0
- package/prisma/migrations/20260926210000_run_host_status_at_dispatch/migration.sql +5 -0
- package/prisma/migrations/20260927010000_resource_grants/migration.sql +72 -0
- package/prisma/migrations/20260927020000_proxy_session_budget_exhausted/migration.sql +3 -0
- package/prisma/migrations/20260927030000_coding_debug_trace/migration.sql +4 -0
- package/prisma/migrations/20260927040000_proxy_session_upstream_failure/migration.sql +3 -0
- package/prisma/migrations/20260927050000_coding_run_services/migration.sql +74 -0
- package/prisma/migrations/20260928000000_coding_run_tool_image/migration.sql +5 -0
- package/prisma/schema.prisma +407 -21
|
@@ -0,0 +1,1061 @@
|
|
|
1
|
+
# Coding Worker Isolation
|
|
2
|
+
|
|
3
|
+
Wardby supports isolated Codex execution with Docker or Kubernetes and
|
|
4
|
+
credential-separated Claude Code execution with Docker. Both launchers fail
|
|
5
|
+
closed when the effective runtime does not match the reviewed policy.
|
|
6
|
+
|
|
7
|
+
## Security Boundary
|
|
8
|
+
|
|
9
|
+
Coding-agent repositories and instructions are untrusted. The Docker daemon,
|
|
10
|
+
host kernel, Wardby control plane, immutable worker image, and dedicated coding
|
|
11
|
+
proxy are trusted. Containers are defense in depth rather than a VM boundary;
|
|
12
|
+
production should run the Docker host on a dedicated worker node or VM with no
|
|
13
|
+
production credentials beyond those required by the proxy.
|
|
14
|
+
|
|
15
|
+
The Codex worker has one network attachment: a unique per-run internal bridge using
|
|
16
|
+
Docker's isolated gateway mode. It has no default external route, published
|
|
17
|
+
port, host mapping, custom DNS server, or direct connection to the control
|
|
18
|
+
plane. A dedicated proxy container is attached to both that internal network
|
|
19
|
+
as `wardby-proxy` and an external network. No other service may join the run
|
|
20
|
+
network.
|
|
21
|
+
|
|
22
|
+
For a coding run with a package allowlist, the same proxy container also
|
|
23
|
+
serves the coding package registry (npm and PyPI) on the same
|
|
24
|
+
`wardby-proxy:8787` port, over a separate, registry-only token derived from
|
|
25
|
+
the run's capability; npm and pip never receive the run's model-API
|
|
26
|
+
capability. See [Installing packages in coding runs](coding-packages.md).
|
|
27
|
+
Registry mode serves both providers: Claude Code's tool runner gets that
|
|
28
|
+
registry-only token's settings from the trusted launcher, delivered as
|
|
29
|
+
`WARDBY_TOOL_SETUP` — never the run capability. Claude Code runs support the
|
|
30
|
+
`node` and `node-python` toolchains; the toolchain selects the tool-runner
|
|
31
|
+
image, which is resolved when the run is dispatched and kept for the whole run.
|
|
32
|
+
|
|
33
|
+
Claude Code uses a credential-separated composite job. Its agent container
|
|
34
|
+
holds the run capability and is attached only to the proxy network; it never
|
|
35
|
+
mounts the repository. Its tool runner container mounts the workspace and
|
|
36
|
+
shares that same run network (for a run with services, through the network
|
|
37
|
+
keeper's namespace; see below), so it can reach the coding proxy too, but it
|
|
38
|
+
holds no model capability or provider credential of its own — only the
|
|
39
|
+
registry-only settings and service variables in `WARDBY_TOOL_SETUP`. The two talk over a private
|
|
40
|
+
Unix socket (`/run/wardby/tool/runner.sock`). The trade is deliberate: the
|
|
41
|
+
tool runner can reach the proxy, but the proxy accepts nothing from it for a
|
|
42
|
+
model call. Both containers, the socket volume, keeper, network, and
|
|
43
|
+
artifacts are attested and cleaned as one persisted handle. On Docker, the tool
|
|
44
|
+
runner gets a fixed 0.25 CPU, a third of the run's memory (clamped between
|
|
45
|
+
128 and 512 MiB), and 64 PIDs, and the agent gets the rest, so a Claude Code
|
|
46
|
+
run needs `cpus ≥ 0.35`, `memoryMb ≥ 256`, and `pids ≥ 96` (the default
|
|
47
|
+
`CODING_PIDS` is 128).
|
|
48
|
+
|
|
49
|
+
**Upgrading with Claude Code runs in flight (Docker launcher).** Let in-flight
|
|
50
|
+
Claude Code runs finish, or stop them, before upgrading the control plane. A
|
|
51
|
+
run launched by a different version fails the new version's container
|
|
52
|
+
attestation and is reported lost.
|
|
53
|
+
|
|
54
|
+
The proxy accepts a run-scoped capability, resolves only exact configured HTTPS
|
|
55
|
+
hostnames, rejects IP literals and every private, loopback, link-local,
|
|
56
|
+
documentation, transition, multicast, and metadata address, rejects mixed DNS
|
|
57
|
+
answers, and pins the vetted address into the socket lookup. Redirects are
|
|
58
|
+
denied. Injected fetch implementations are a test seam and must not be used in
|
|
59
|
+
production composition.
|
|
60
|
+
|
|
61
|
+
The Codex worker retries a failed model request, or a response stream that
|
|
62
|
+
drops, up to three times each before it fails the run. The Claude worker
|
|
63
|
+
likewise retries a failed model request up to three times. Every retry is a
|
|
64
|
+
new request through the proxy, so it is checked against the run's budget
|
|
65
|
+
again.
|
|
66
|
+
If the stream still fails, the run's diagnostic (in the control-plane log,
|
|
67
|
+
under the run's `coding_diag_…` id) names the cause as a fixed code, never the
|
|
68
|
+
error text: `job_coding_stream_proxy_denied` (401/403 from the proxy),
|
|
69
|
+
`job_coding_stream_rate_limited` (429), `job_coding_stream_upstream_error`
|
|
70
|
+
(5xx), `job_coding_stream_timeout`, `job_coding_stream_proxy_unreachable`
|
|
71
|
+
(the worker could not connect), `job_coding_stream_agent_exited` (the agent
|
|
72
|
+
process exited without an HTTP error), or `job_coding_stream_failed` when none
|
|
73
|
+
of these match. A refusal because the run's budget is used up ends the run as
|
|
74
|
+
out of budget instead.
|
|
75
|
+
|
|
76
|
+
The proxy side of the same request is in the coding proxy's log as
|
|
77
|
+
`audit.*` events (`audit.request.reserved`, `audit.response.completed`,
|
|
78
|
+
`audit.request.uncertain`, …), with ids, models, amounts and fixed reason
|
|
79
|
+
codes only. When the proxy has to end a stream early, `audit.request.uncertain`
|
|
80
|
+
says why: `upstream_failed:<code>` when the model API reported a failure (the
|
|
81
|
+
failure event is passed on to the worker unchanged), or a code such as
|
|
82
|
+
`terminal_usage_missing`, with the upstream response's content type and
|
|
83
|
+
encoding. To see the error text behind a code, turn on a
|
|
84
|
+
[debug trace](#debug-trace) for the agent.
|
|
85
|
+
|
|
86
|
+
**Model-provider failures.** The proxy also remembers the first failure the
|
|
87
|
+
model provider reported for a run: the error code of a failed stream, or of a
|
|
88
|
+
rejected request (its HTTP status as `http_<status>` when the response names no
|
|
89
|
+
code). When the run's worker job then fails, the run ends `failed` with the
|
|
90
|
+
error `coding_provider_<class>` and the failure category `provider_<class>`,
|
|
91
|
+
where the class is:
|
|
92
|
+
|
|
93
|
+
| Category | Provider codes | What the host is told |
|
|
94
|
+
| ----------------------- | --------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
95
|
+
| `provider_quota` | codes naming a quota, spend limit, billing, funds or credit (`insufficient_quota`, `project_spend_limit_exceeded`, …) | "The model provider refused the request: its account has reached a spending or quota limit. An operator needs to raise the limit with the provider, then retry." |
|
|
96
|
+
| `provider_rate_limited` | `rate_limit_exceeded` and other rate-limit codes | "The model provider is rate-limiting requests. Try again later." |
|
|
97
|
+
| `provider_unavailable` | `server_error`, `overloaded`, `overloaded_error`, `service_unavailable`, `api_error`, `http_5xx` | "The model provider reported an outage or overload. Try again later." |
|
|
98
|
+
| `provider_rejected` | anything else | "The model provider rejected the request." |
|
|
99
|
+
|
|
100
|
+
The host sees that sentence in the continuation status comment
|
|
101
|
+
(`❌ <agent> could not run: …`) and in an @-mention's status comment
|
|
102
|
+
(`❌ A sub-run could not reach the model: …`); it never sees the provider's
|
|
103
|
+
code or message. `provider_quota` means the **provider account** behind the
|
|
104
|
+
coding credential is out of quota or has hit its spend limit — raise the limit
|
|
105
|
+
with the provider. It is not wardby's budget: a run that uses up its own
|
|
106
|
+
wardby budget still ends `budget_exhausted`, and that wins when both happen.
|
|
107
|
+
|
|
108
|
+
For alerting, the control plane logs one `warn` line per such run with
|
|
109
|
+
`event: "coding.provider_failure"`, the `runId`, the `class`, and the raw
|
|
110
|
+
`upstreamCode`.
|
|
111
|
+
|
|
112
|
+
### Model request allowlist
|
|
113
|
+
|
|
114
|
+
The network policy only helps if the one reachable destination, the proxy,
|
|
115
|
+
cannot be asked to reach somewhere else. Model APIs can do that on the
|
|
116
|
+
caller's behalf: an OpenAI hosted `mcp` tool or a remote `input_image` URL
|
|
117
|
+
makes OpenAI's servers contact any host, and hosted-tool fees and non-default
|
|
118
|
+
service tiers are billed outside the metered tokens. Code in a Codex worker can
|
|
119
|
+
read the run capability, so the proxy validates every request body against an
|
|
120
|
+
allowlist before it resolves a credential, and refuses anything else with a
|
|
121
|
+
`400` (`src/providers/coding-proxy/proxy.ts`). Each refusal is also written to
|
|
122
|
+
the proxy audit log as `request.rejected` with the run id and the refusal
|
|
123
|
+
code, so a smuggling attempt is attributable to its run. JSON nested too
|
|
124
|
+
deeply to process is refused with `request_nesting_too_deep`.
|
|
125
|
+
|
|
126
|
+
- **Anthropic Messages** (`parseAnthropicRequest`): a fixed set of keys,
|
|
127
|
+
text/`tool_use`/`tool_result`/thinking blocks, exactly the two Wardby tool
|
|
128
|
+
names, and the reviewed beta values.
|
|
129
|
+
- **OpenAI Responses** (`parseOpenAiRequest`), built from what the pinned
|
|
130
|
+
Codex CLI actually sends (recorded in
|
|
131
|
+
`src/providers/coding-proxy/fixtures/codex-<version>-responses-requests.json`
|
|
132
|
+
and exercised end to end by `src/coding-worker/codex-compatibility.test.ts`):
|
|
133
|
+
- Top-level keys: `model`, `instructions`, `input`, `tools`, `tool_choice`,
|
|
134
|
+
`parallel_tool_calls`, `reasoning`, `store`, `stream`, `include`,
|
|
135
|
+
`prompt_cache_key`, `text`, `client_metadata`, `max_output_tokens`,
|
|
136
|
+
`background`, `service_tier`. Any other key, including `prompt`,
|
|
137
|
+
`previous_response_id`, `conversation` and `metadata`, is refused with
|
|
138
|
+
`openai_request_key_not_allowed:<key>`.
|
|
139
|
+
- Tools: `function` and `custom` definitions, and `namespace` groups of
|
|
140
|
+
them, either in `tools` or in Codex's `additional_tools` input item. Every
|
|
141
|
+
other tool type, including all hosted tools (`web_search`, `mcp`,
|
|
142
|
+
`code_interpreter`, `image_generation`, `file_search`, `local_shell`, and
|
|
143
|
+
so on), is refused with `openai_tool_not_allowed:<type>`. Tool names are
|
|
144
|
+
checked for shape (`[A-Za-z0-9_-]{1,128}`) but not against a fixed list:
|
|
145
|
+
they change with the model and the Codex version (for example `exec` in
|
|
146
|
+
code mode against `exec_command` otherwise), and a client-side tool runs
|
|
147
|
+
inside the worker, where the container boundary already governs it.
|
|
148
|
+
`tool_choice` may only be `auto`, `none` or `required`.
|
|
149
|
+
- Input items: `message` (`input_text`/`input_image` for user, developer
|
|
150
|
+
and system; `output_text` for assistant), `reasoning`, `function_call`,
|
|
151
|
+
`function_call_output`, `custom_tool_call`, `custom_tool_call_output`,
|
|
152
|
+
`agent_message` and `additional_tools`, each with only the keys Codex
|
|
153
|
+
sends. Anything else (`input_file`, `input_audio`, `item_reference`,
|
|
154
|
+
replayed hosted-tool calls, `compaction`, `local_shell_call`) is refused
|
|
155
|
+
with `openai_input_not_allowed:<type>`. An `input_image` must be an inline
|
|
156
|
+
`data:image/{png,jpeg,gif,webp};base64,` URL, as Codex's `view_image`
|
|
157
|
+
produces; a remote URL or a `file_id` is refused with
|
|
158
|
+
`openai_remote_input_not_allowed`.
|
|
159
|
+
- Pinned values: `include` only `reasoning.encrypted_content`; `reasoning`
|
|
160
|
+
only `effort`, `summary` and `context` with known values; `text` only
|
|
161
|
+
`verbosity` and a `text` or `json_schema` format; `service_tier` only
|
|
162
|
+
unset or `default` (`service_tier_not_allowed`; `auto` is refused because
|
|
163
|
+
it defers to the OpenAI project's own tier, which may be priority);
|
|
164
|
+
`client_metadata` only the string-valued keys Codex sends
|
|
165
|
+
(`openai_client_metadata_key_not_allowed:<key>`). The proxy still
|
|
166
|
+
forces `store: false` and `background: false` and adds a
|
|
167
|
+
`max_output_tokens` ceiling when Codex omits it.
|
|
168
|
+
|
|
169
|
+
Upgrading the pinned `@openai/codex-sdk` means re-recording that fixture
|
|
170
|
+
against a local fake upstream and rerunning the compatibility test. A Codex
|
|
171
|
+
release that sends a new key or item type fails closed at the proxy rather
|
|
172
|
+
than silently widening what reaches OpenAI.
|
|
173
|
+
|
|
174
|
+
The proxy also binds a second listener, the **deny port** (`8788`,
|
|
175
|
+
`CODING_PROXY_DENY_PORT`), which serves nothing: it accepts a connection,
|
|
176
|
+
sends no bytes and closes it immediately
|
|
177
|
+
(`src/providers/coding-proxy/deny-port.ts`). It exists so a run pod can prove
|
|
178
|
+
its own NetworkPolicy is enforced — see "The enforcement gate" below.
|
|
179
|
+
|
|
180
|
+
The deny port ships to **every** deployment, not just Kubernetes:
|
|
181
|
+
`startConfiguredCodingProxy` is the only proxy entry point, so a Docker
|
|
182
|
+
deployment's proxy also binds `0.0.0.0:8788` (and refuses to start if it
|
|
183
|
+
cannot). No host port is published for it, so there is no port conflict. In
|
|
184
|
+
Docker mode a run container can reach it, since a Docker network has no
|
|
185
|
+
port-level policy — that is harmless, because the listener accepts the
|
|
186
|
+
connection, sends nothing and closes it. It is not a leak; it is a fact about
|
|
187
|
+
reachability that the Kubernetes launcher turns into evidence.
|
|
188
|
+
|
|
189
|
+
## Debug trace
|
|
190
|
+
|
|
191
|
+
When a fixed code is not enough to tell why a coding agent's runs fail, an
|
|
192
|
+
admin can turn on a **debug trace** for that agent for a limited time:
|
|
193
|
+
|
|
194
|
+
```
|
|
195
|
+
update_agent { "id": "<agent id>", "codingProfile": { "debugTraceMinutes": 30 } }
|
|
196
|
+
```
|
|
197
|
+
|
|
198
|
+
`debugTraceMinutes` is 1 to 1440 and needs the `agents:admin` scope (the admin
|
|
199
|
+
role), as `workerImageRef` does; `null` turns the trace off early. Every change
|
|
200
|
+
is written to the control-plane log as a `coding.debug_trace.set` audit line.
|
|
201
|
+
`get_agent` shows the expiry as `codingProfile.debugTraceUntil`. The trace
|
|
202
|
+
expires on its own: each coding run dispatched before the expiry is traced for
|
|
203
|
+
its whole life, and `get_run` shows `debugTrace: true` for it. Runs dispatched
|
|
204
|
+
afterwards are not.
|
|
205
|
+
|
|
206
|
+
A traced run's Codex worker writes every Codex stream event, and the full text
|
|
207
|
+
of a stream failure (including the agent process's own error output and cause
|
|
208
|
+
chain), to its **own log**, one JSON line each under a `debugTrace` key. On
|
|
209
|
+
Kubernetes that is the run pod's `worker` container log; on GKE, Cloud Logging
|
|
210
|
+
keeps it after the pod is deleted. With the Docker launcher it is the worker
|
|
211
|
+
container's log (`docker logs`), for as long as the container exists. Nothing from the trace goes to the database,
|
|
212
|
+
GitHub, the run's error, or the control-plane log. Token-shaped values (the run
|
|
213
|
+
capability, bearer tokens, API keys, GitHub tokens) are redacted, each line is
|
|
214
|
+
capped at 16 KiB and each run at 2 MiB.
|
|
215
|
+
|
|
216
|
+
The trace can contain prompts, model output and repository content. Turn it on
|
|
217
|
+
only while diagnosing a failure, for as short a time as you can, and treat the
|
|
218
|
+
pod logs of traced runs as sensitive. Tracing needs a worker driver image that
|
|
219
|
+
supports it; a worker that predates it rejects a traced run's input as
|
|
220
|
+
invalid.
|
|
221
|
+
|
|
222
|
+
## Container Policy
|
|
223
|
+
|
|
224
|
+
`src/providers/jobs/docker-isolation.ts` is the canonical policy builder and
|
|
225
|
+
startup attestation layer. `DockerJobLauncher` executes its argument arrays
|
|
226
|
+
directly with `spawn`; it never invokes a shell.
|
|
227
|
+
|
|
228
|
+
The worker policy requires:
|
|
229
|
+
|
|
230
|
+
- An immutable `sha256:` image ID or repository digest with `--pull never`.
|
|
231
|
+
- UID/GID `10001:10001`, all capabilities dropped, no new privileges, Docker's
|
|
232
|
+
built-in seccomp profile, private cgroup and PID namespaces, and no host IPC.
|
|
233
|
+
- A read-only root filesystem with bounded `noexec,nosuid,nodev` tmpfs mounts
|
|
234
|
+
for `/tmp` and `/home/wardby`. The agent's own temporary files (`TMPDIR`)
|
|
235
|
+
are not among them: they go to the workspace's `.cache/tmp`, which is
|
|
236
|
+
disk-backed and counts toward the run's `workspaceDiskMb`, not this small
|
|
237
|
+
in-memory scratch — a real `pip install` or `npm install` unpacks and builds
|
|
238
|
+
in `TMPDIR`, which the tmpfs is usually too small to hold. `.cache` is
|
|
239
|
+
never part of the collected diff or pull request.
|
|
240
|
+
- Exact CPU, memory, equal memory+swap, PID, shared-memory, disk, and wall-clock
|
|
241
|
+
limits. Equal memory and memory+swap disables additional swap allowance.
|
|
242
|
+
- No devices, device requests, bind mounts, extra groups, custom DNS, extra
|
|
243
|
+
hosts, published ports, or restart policy.
|
|
244
|
+
- At most 2 MiB of local Docker logs and a cooperative SIGTERM grace period
|
|
245
|
+
before forced termination.
|
|
246
|
+
|
|
247
|
+
A run with services (see [coding-services.md](coding-services.md)) adds two
|
|
248
|
+
kinds of container, both labelled and attested like the others before they
|
|
249
|
+
start:
|
|
250
|
+
|
|
251
|
+
- A **network keeper** (`wardby-netns-<token>`): the worker image running an
|
|
252
|
+
idle `node` process as `10001:10001`, read-only, all capabilities dropped,
|
|
253
|
+
no new privileges, built-in seccomp, 64 MiB and 32 PIDs, with no mounts, on
|
|
254
|
+
the run's internal network. It owns the run's network namespace.
|
|
255
|
+
- One container per service (`wardby-svc-<token>-<name>`): the catalog image
|
|
256
|
+
by digest, as `10001:10001`, with a read-only root filesystem, all
|
|
257
|
+
capabilities dropped, no new privileges, built-in seccomp, private cgroup
|
|
258
|
+
namespace, private IPC with 64 MiB of shared memory, the catalog's CPU,
|
|
259
|
+
memory (plus its tmpfs disk and shared memory), 512 PIDs, a bounded tmpfs at
|
|
260
|
+
its data path and each writable path, its catalog `serviceEnv`, and no
|
|
261
|
+
published ports, mounts, devices or restart policy.
|
|
262
|
+
|
|
263
|
+
A Codex run's worker and every service use `--network container:<network
|
|
264
|
+
keeper>`, so a service answers the worker on `127.0.0.1` and the worker still
|
|
265
|
+
reaches the proxy by its alias on the internal network. Nothing else about the
|
|
266
|
+
worker changes.
|
|
267
|
+
|
|
268
|
+
For a Claude Code run it is the tool runner, which runs the agent's shell
|
|
269
|
+
commands, that joins the network keeper's namespace: it reaches every service
|
|
270
|
+
on `127.0.0.1` and the proxy by its alias, and it is created only after every
|
|
271
|
+
service is ready. The agent container stays on the run's internal network,
|
|
272
|
+
unchanged; it reaches the tool runner over the Unix socket in their shared
|
|
273
|
+
storage volume, which does not depend on either container's network. The
|
|
274
|
+
launcher attests the tool runner's network before it starts and on every
|
|
275
|
+
status check: the network keeper's namespace and no network of its own.
|
|
276
|
+
|
|
277
|
+
A service may listen on every interface of the network keeper's namespace, so
|
|
278
|
+
anything on the run's internal network (the proxy, and for a Claude Code run
|
|
279
|
+
the agent container) can address it there, just as every container in a
|
|
280
|
+
Kubernetes run's pod shares the services' namespace. Services are disposable
|
|
281
|
+
test fixtures with catalog credentials; do not treat them as a boundary.
|
|
282
|
+
|
|
283
|
+
The capability value is inherited from the trusted launcher's child-process
|
|
284
|
+
environment with `--env WARDBY_RUN_CAPABILITY`; it is never included in command
|
|
285
|
+
arguments. Docker administrators can still inspect container environment, so
|
|
286
|
+
daemon access remains privileged and must be tightly restricted.
|
|
287
|
+
|
|
288
|
+
## Ephemeral Storage
|
|
289
|
+
|
|
290
|
+
Each run receives one quota-bounded local tmpfs volume. A hardened, no-network
|
|
291
|
+
keeper container holds the volume open from preparation through result
|
|
292
|
+
collection. It creates exactly four private subdirectories and emits
|
|
293
|
+
`wardby_storage_ready` before the launcher may seed them.
|
|
294
|
+
|
|
295
|
+
The worker sees only these volume subpaths:
|
|
296
|
+
|
|
297
|
+
- `/workspace`: read-write checkout files.
|
|
298
|
+
- Git metadata remains in the trusted keeper volume and is not mounted into the worker.
|
|
299
|
+
- `/run/wardby/input`: read-only, validated input artifact.
|
|
300
|
+
- `/run/wardby/output`: read-write result artifact.
|
|
301
|
+
|
|
302
|
+
There are no production host bind mounts. The Docker JobLauncher transfers data
|
|
303
|
+
through the keeper with Docker copy/archive APIs, validates it before launch and
|
|
304
|
+
after collection, and stops the keeper only after collection. Stopping the last
|
|
305
|
+
container that mounts this local tmpfs intentionally destroys the run data.
|
|
306
|
+
The quota is RAM-backed; operators must bound aggregate concurrent `diskMb`
|
|
307
|
+
allocations at the host scheduler as well as per run.
|
|
308
|
+
|
|
309
|
+
## Required Lifecycle
|
|
310
|
+
|
|
311
|
+
1. Validate `JobSpec`, image digest, proxy identity, and host support.
|
|
312
|
+
2. Create and inspect the internal network and tmpfs volume.
|
|
313
|
+
3. Create, inspect, and start the keeper; wait for its readiness marker.
|
|
314
|
+
4. Seed the four fixed storage areas through the keeper, never a bind mount.
|
|
315
|
+
5. Attach the dedicated proxy and attest that it is both internally and
|
|
316
|
+
externally connected.
|
|
317
|
+
6. For a run with services only: create, inspect and start the network
|
|
318
|
+
keeper; then, for each service in turn, use the host's copy of its image or
|
|
319
|
+
pull it by digest, create and inspect its container in the keeper's
|
|
320
|
+
namespace, start it, and run its readiness command until it passes or the
|
|
321
|
+
run fails with `coding_service_unready:<name>`. All of this shares one
|
|
322
|
+
120-second start-up limit (image pulls, each bounded to 5 minutes, are not
|
|
323
|
+
counted) and never runs past the run's deadline.
|
|
324
|
+
A Claude Code run's tool runner is created, inspected and started only
|
|
325
|
+
after this step, in the network keeper's namespace.
|
|
326
|
+
7. Create the worker with the run capability supplied only in the child
|
|
327
|
+
environment; inspect every effective control before start.
|
|
328
|
+
8. Start the worker and enforce `deadlineMs`. Send SIGTERM at expiry, then
|
|
329
|
+
SIGKILL after `stopGraceSeconds` if it remains alive.
|
|
330
|
+
9. Cancel the proxy session, collect and validate bounded output, and re-read
|
|
331
|
+
authoritative usage before any repository publication.
|
|
332
|
+
10. For `changes_ready` only, copy the worker workspace into a new host staging
|
|
333
|
+
directory, reject special files, nested `.git`, escaping symlinks, and size
|
|
334
|
+
or entry-limit violations, then atomically replace the trusted checkout.
|
|
335
|
+
11. Revalidate protected paths, Git configuration, branch ancestry, remotes,
|
|
336
|
+
and budget; create one controlled commit, push one deterministic branch,
|
|
337
|
+
and create or find one draft pull request.
|
|
338
|
+
12. Persist the typed coding result and terminal run status in one transaction,
|
|
339
|
+
then remove the worker, any service containers and network keeper, the
|
|
340
|
+
keeper, network, volume, input artifact, and VCS workspace. `no_changes`
|
|
341
|
+
and `budget_exhausted` never push.
|
|
342
|
+
|
|
343
|
+
`ContainerExecutor` treats a durable proxy session without a durable job handle
|
|
344
|
+
as ambiguous provisioning and never relaunches it. A persisted handle is the
|
|
345
|
+
only recovery path. Duplicate starts, terminal collection, Git finalization,
|
|
346
|
+
and cleanup converge on the same handle, branch, commit, pull request, usage,
|
|
347
|
+
and status.
|
|
348
|
+
|
|
349
|
+
Any missing host feature, unsupported network option, failed inspection,
|
|
350
|
+
unexpected mount/network/environment, or cleanup ambiguity is the fixed
|
|
351
|
+
`docker_isolation_unsupported` failure. Production must not fall back to a
|
|
352
|
+
weaker profile.
|
|
353
|
+
|
|
354
|
+
## Control Plane Configuration
|
|
355
|
+
|
|
356
|
+
Set `JOB_LAUNCHER=docker`, `CODING_WORKER_IMAGE` to an immutable repository
|
|
357
|
+
digest or Docker local image ID, and `CODING_PROXY_CONTAINER` to the dedicated proxy container name.
|
|
358
|
+
For Claude Code, also set `CODING_CLAUDE_WORKER_IMAGE` and
|
|
359
|
+
`CODING_CLAUDE_TOOL_RUNNER_IMAGE` to their immutable IDs.
|
|
360
|
+
`VCS_WORK_ROOT`, `CODING_JOB_STATE_ROOT`, and `CODING_ARTIFACT_ROOT` must be
|
|
361
|
+
trusted host-only directories. Resource limits are controlled by
|
|
362
|
+
`CODING_CPUS`, `CODING_MEMORY_MB`, `CODING_PIDS`, and `CODING_DISK_MB`.
|
|
363
|
+
`CODING_MAX_DISK_MB` (must be an integer between 64 and 32768 and at least
|
|
364
|
+
the effective `CODING_DISK_MB`; **defaults to the effective `CODING_DISK_MB`
|
|
365
|
+
itself**, not a larger number) is the operator ceiling on the per-agent
|
|
366
|
+
`workspaceDiskMb` coding-profile field described below — without it, any
|
|
367
|
+
`agents:write` caller could size a run's workspace disk up to 32 GiB
|
|
368
|
+
(Docker: a RAM-backed tmpfs; Kubernetes: an `emptyDir`), times
|
|
369
|
+
`CODING_MAX_CONCURRENT`, on every run. Defaulting the ceiling to the
|
|
370
|
+
existing disk size means upgrading to this branch changes nothing for a
|
|
371
|
+
deployment that doesn't set `CODING_MAX_DISK_MB`: `workspaceDiskMb` stays
|
|
372
|
+
inert until an operator explicitly raises the ceiling above `CODING_DISK_MB`.
|
|
373
|
+
|
|
374
|
+
A coding agent's profile carries an optional `workspaceDiskMb` (MiB; `null`
|
|
375
|
+
means "use the deployment default `CODING_DISK_MB`"). It is snapshotted onto
|
|
376
|
+
the `CodingRun` at dispatch time, so a later profile edit never changes an
|
|
377
|
+
in-flight run's size, and it is capped by `CODING_MAX_DISK_MB`: a run whose
|
|
378
|
+
snapshotted `workspaceDiskMb` exceeds the ceiling fails with
|
|
379
|
+
`coding_workspace_disk_exceeds_limit`
|
|
380
|
+
(`src/providers/executor/container.ts`, the `jobSpec()` check). This check
|
|
381
|
+
runs after the workspace has already been cloned onto the control plane, so
|
|
382
|
+
an over-ceiling agent still pays for a clone before failing, and the failure
|
|
383
|
+
reaches the run record only as the sanitized `coding_failure_workspace:<id>`
|
|
384
|
+
(the same generic bucketing every workspace/git-related failure gets) — not
|
|
385
|
+
a distinctly labeled "refused" outcome, and not currently logged anywhere
|
|
386
|
+
more diagnosable on the control plane. This is a known rough edge, not a
|
|
387
|
+
security gap: no run ever exceeds the ceiling, it
|
|
388
|
+
just fails less legibly than it could.
|
|
389
|
+
|
|
390
|
+
`CODING_MAX_CONCURRENT` (default `4`) caps coding runs that hold a slot at
|
|
391
|
+
once, across every control-plane replica: the cap is enforced in Postgres
|
|
392
|
+
inside the provisioning claim, so adding replicas never raises it. A run
|
|
393
|
+
over the cap stays `pending` and `get_run`/`list_runs` show `codingQueuedAt`;
|
|
394
|
+
it starts, oldest first, when a slot frees (immediately in the process whose
|
|
395
|
+
run finished, or on the scheduler leader's next tick). A newly dispatched run
|
|
396
|
+
never takes a free slot ahead of an older queued run. A run still queued after
|
|
397
|
+
`CODING_QUEUE_TIMEOUT_SEC` (default `3600`) fails with `coding_queue_timeout`.
|
|
398
|
+
Slot usage is derived from run state, so a crashed replica cannot leak slots:
|
|
399
|
+
its runs are reconciled to `lost`, which frees them.
|
|
400
|
+
|
|
401
|
+
Operating the queue across replicas:
|
|
402
|
+
|
|
403
|
+
- Every replica must set the same `CODING_MAX_CONCURRENT`. Each claim
|
|
404
|
+
enforces the value of the replica making it, so mismatched values make the
|
|
405
|
+
effective cap depend on which replica dispatched the run.
|
|
406
|
+
- Draining on a timer and applying `CODING_QUEUE_TIMEOUT_SEC` need a
|
|
407
|
+
scheduler process (`wardby serve` or `wardby scheduler`). A process that
|
|
408
|
+
only serves MCP (`wardby mcp`) drains only when one of its own coding runs
|
|
409
|
+
finishes; without a scheduler somewhere, queued runs can wait indefinitely
|
|
410
|
+
and never time out.
|
|
411
|
+
- Upgrade all replicas together. A replica running a version from before
|
|
412
|
+
the queue ignores the cap, and its reconciler reaps queued runs as `lost`. Clones are shallow
|
|
413
|
+
(`--depth 1`); the worker never receives Git history and finalization needs
|
|
414
|
+
only the base commit.
|
|
415
|
+
|
|
416
|
+
The GitHub adapter requires `GITHUB_APP_ID` and `GITHUB_APP_PRIVATE_KEY`; the
|
|
417
|
+
App installation is checked while preparing the workspace, before the
|
|
418
|
+
billable proxy session is created. Upstream keys remain behind
|
|
419
|
+
`CODING_OPENAI_CREDENTIAL_REF` and `CODING_ANTHROPIC_CREDENTIAL_REF` and are
|
|
420
|
+
never written to the database, input
|
|
421
|
+
artifact, Docker arguments, or Git workspace.
|
|
422
|
+
|
|
423
|
+
The embedded Codex SDK runs with its inner sandbox disabled because the
|
|
424
|
+
worker's Docker boundary is authoritative: it has a read-only root filesystem,
|
|
425
|
+
no Linux capabilities, no host mounts or Docker socket, no public network,
|
|
426
|
+
and only isolated workspace/output volumes plus the trusted proxy connection.
|
|
427
|
+
This avoids relying on a nested sandbox that cannot validate Wardby's
|
|
428
|
+
intentionally Git-metadata-free workspace.
|
|
429
|
+
|
|
430
|
+
Coding-agent authoring and execution are MCP-first. `trigger_agent` accepts
|
|
431
|
+
an optional bounded `task` and `baseRef` only for a coding agent owned by the
|
|
432
|
+
caller; the chosen values are copied into the immutable run record. Webhooks
|
|
433
|
+
use the profile default task unless `allowWebhookTaskOverride` is explicitly
|
|
434
|
+
enabled on that coding profile. Coding results returned through `get_run` and
|
|
435
|
+
`tasks/get` are validated, redacted summaries/tests only; job handles and
|
|
436
|
+
execution policy stay internal.
|
|
437
|
+
|
|
438
|
+
The worker receives only its task text, so a coding agent's own `systemPrompt`
|
|
439
|
+
is placed ahead of the request at dispatch ("Standing instructions for this
|
|
440
|
+
coding agent: … Request: …") and stored with the run. A blank prompt leaves
|
|
441
|
+
the task unchanged; a combination over the 16 KiB task limit is refused rather
|
|
442
|
+
than truncated.
|
|
443
|
+
|
|
444
|
+
The worker's result may carry an optional `tag`, shown as `[tag]` in the pull
|
|
445
|
+
request title: at most 32 characters of letters, digits, `.`, `_`, `/`, and
|
|
446
|
+
`-`, starting with a letter or digit. The output schema describes that rule to
|
|
447
|
+
the model, and a tag that breaks it is normalized to a lowercase slug (or
|
|
448
|
+
dropped) instead of failing an otherwise finished run. GitHub finalization
|
|
449
|
+
re-validates the tag independently. Linking an issue is not the tag's job: put
|
|
450
|
+
a closing keyword such as `Resolves #37` in the task so the worker includes it
|
|
451
|
+
in the summary, which becomes the pull request body.
|
|
452
|
+
|
|
453
|
+
Some workspace folders are never collected from a run: `node_modules`, `.venv`,
|
|
454
|
+
`venv`, `__pycache__`, `.pytest_cache`, `.ruff_cache`, `.mypy_cache`, `.tox`,
|
|
455
|
+
`.vite`, and `.cache`, at any depth, plus any repository-relative paths in the
|
|
456
|
+
agent's `collectExclude` (for example `web/dist`). On Kubernetes the keeper's
|
|
457
|
+
`tar` leaves them out, so they never leave the pod; on Docker they are removed
|
|
458
|
+
from the staging copy before it is validated. They therefore never count toward
|
|
459
|
+
the entry, size, symlink, or nested-repository checks, and Git staging excludes
|
|
460
|
+
them too, so a tracked file under an excluded folder is left unchanged.
|
|
461
|
+
|
|
462
|
+
Operator-only checks and cleanup remain available through the CLI:
|
|
463
|
+
|
|
464
|
+
```sh
|
|
465
|
+
wardby coding preflight
|
|
466
|
+
wardby coding cleanup --run-id <id>
|
|
467
|
+
```
|
|
468
|
+
|
|
469
|
+
The preflight command requires Docker mode, validates the pinned worker-image
|
|
470
|
+
digest, and confirms the image is available to Docker. Cleanup delegates to
|
|
471
|
+
the configured executor so it resolves and stops the persisted container job.
|
|
472
|
+
|
|
473
|
+
## Audit And Retention
|
|
474
|
+
|
|
475
|
+
The executor emits metadata-only lifecycle events for queueing, preparation,
|
|
476
|
+
launch, running, budget cutoff, stopping, collection, PR creation, terminal
|
|
477
|
+
outcome, and cleanup. Events carry run ID, opaque job ID, sanitized failure
|
|
478
|
+
category, opaque diagnostic ID, duration, and budget totals only. They never
|
|
479
|
+
carry task text, prompts, repository contents, diffs, worker environment, raw
|
|
480
|
+
Docker logs, or credentials.
|
|
481
|
+
|
|
482
|
+
The production log/metrics collector retains those events for 90 days by
|
|
483
|
+
default policy. `CodingRun` stores only the sanitized failure category and
|
|
484
|
+
diagnostic ID alongside the normal run record; it is not an artifact store.
|
|
485
|
+
Every terminal path removes the worker volume, input artifact, job state, and
|
|
486
|
+
trusted checkout. Restart reconciliation repeats that cleanup from the
|
|
487
|
+
persisted job handle and marks ambiguous provisioning as `lost` instead of
|
|
488
|
+
relaunching it.
|
|
489
|
+
|
|
490
|
+
See [release verification](release-verification.md) for automated gates and
|
|
491
|
+
live-fixture rules.
|
|
492
|
+
|
|
493
|
+
## Kubernetes launcher (`JOB_LAUNCHER=kubernetes`)
|
|
494
|
+
|
|
495
|
+
`KubernetesJobLauncher` (`src/providers/jobs/kubernetes.ts`) implements the
|
|
496
|
+
same `WorkspaceJobLauncher` contract as `DockerJobLauncher` and is a drop-in
|
|
497
|
+
alternative for deployments with no Docker daemon available to the control
|
|
498
|
+
plane (e.g. a Cloud Run host). Everything above `ContainerExecutor` —
|
|
499
|
+
proxy sessions, the Git finalizer, workspace validation, recovery, cleanup —
|
|
500
|
+
is unchanged; only the container-orchestration seam is replaced. Both Codex
|
|
501
|
+
and Claude Code run on this launcher; Claude Code's pod adds a tool-runner
|
|
502
|
+
sidecar, described in "Pod layout" below.
|
|
503
|
+
|
|
504
|
+
### Enabling it
|
|
505
|
+
|
|
506
|
+
Set `JOB_LAUNCHER=kubernetes` and `CODING_WORKER_IMAGE` to a **registry
|
|
507
|
+
digest** (`repo@sha256:<64 hex>` — a bare `sha256:` local image ID is
|
|
508
|
+
rejected; a cluster cannot pull it). For Claude Code, also set
|
|
509
|
+
`CODING_CLAUDE_WORKER_IMAGE` and `CODING_CLAUDE_TOOL_RUNNER_IMAGE` (and, for
|
|
510
|
+
Claude agents on the `node-python` toolchain,
|
|
511
|
+
`CODING_CLAUDE_TOOL_RUNNER_IMAGE_NODE_PYTHON_3_12`) to registry digests; the
|
|
512
|
+
control plane refuses to start if any of them is set to anything else.
|
|
513
|
+
Kubernetes-specific settings (`src/config/providers.ts`,
|
|
514
|
+
`loadKubernetesJobConfig`):
|
|
515
|
+
|
|
516
|
+
| Variable | Default | Meaning |
|
|
517
|
+
| ------------------------------- | ----------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
518
|
+
| `KUBERNETES_NAMESPACE` | `wardby-coding` | The one namespace holding the proxy and every per-run object. |
|
|
519
|
+
| `KUBERNETES_PROXY_SERVICE` | `wardby-coding-proxy` | The proxy's Service name; its ClusterIP is what `hostAliases` points runs at. |
|
|
520
|
+
| `KUBERNETES_CONTEXT` | (unset → in-cluster/default kubeconfig context) | Which kubeconfig context `ClientNodeKubernetesApi` connects with. |
|
|
521
|
+
| `KUBERNETES_RUNTIME_CLASS` | (unset) | e.g. `gvisor` on GKE. Unset means pods run without a sandboxing runtime class — logged once per launch as `kubernetes_runtime_class_unset` and development-only. |
|
|
522
|
+
| `KUBERNETES_RUN_PRIORITY_CLASS` | (unset) | PriorityClass for run pods and the preflight canary, e.g. `wardby-coding-run` in the GKE overlay. Must be an existing class and a DNS-1123 subdomain; `system-` classes are refused. Unset means no class (priority 0). |
|
|
523
|
+
|
|
524
|
+
The `CODING_CPUS` / `CODING_MEMORY_MB` / `CODING_PIDS` / `CODING_DISK_MB` /
|
|
525
|
+
`CODING_MAX_DISK_MB` settings above apply identically; the per-agent
|
|
526
|
+
`workspaceDiskMb` profile field sizes the pod's `storage` `emptyDir` the same
|
|
527
|
+
way it sizes Docker's tmpfs volume.
|
|
528
|
+
|
|
529
|
+
### Pod layout
|
|
530
|
+
|
|
531
|
+
One pod per run, built by the canonical, deny-by-default policy in
|
|
532
|
+
`kubernetes-isolation.ts`'s `buildRunPod`:
|
|
533
|
+
|
|
534
|
+
- An **init container `storage-init`** runs first and creates
|
|
535
|
+
`/run/wardby/storage/{workspace,input,output}` (mode `0700`, owned by uid 10001) before any regular container starts. This exists because kubelet
|
|
536
|
+
creates a subPath mount's target directory root-owned the first time it
|
|
537
|
+
sets up the worker's volume mounts, and the keeper (uid 10001, no Linux
|
|
538
|
+
capabilities) cannot `chmod` a root-owned directory it doesn't own. This
|
|
539
|
+
is required because otherwise a real-cluster run fails
|
|
540
|
+
`kubernetes_pod_start_timeout`. `keeper.js` itself is unchanged — Docker still
|
|
541
|
+
shares it, and Docker's bind-mount-free volume never had this problem.
|
|
542
|
+
- **Claude Code's `tool-runner`** (only for a Claude Code run): a native
|
|
543
|
+
sidecar (init container with `restartPolicy: Always`) after `storage-init`
|
|
544
|
+
and before the service sidecars, `keeper`, and `worker`. It mounts the
|
|
545
|
+
workspace and the socket directory `/run/wardby/tool`, gets its `WARDBY_TOOL_SETUP`
|
|
546
|
+
environment variable from the run Secret's `tool-setup` key (never the run
|
|
547
|
+
capability), and shares the pod's network and IPC namespaces with the
|
|
548
|
+
worker, like every container in a pod: loopback and abstract Unix sockets
|
|
549
|
+
are common to both, and the worker listens on nothing. The run's one egress
|
|
550
|
+
rule (the proxy) applies to it the same as the worker. Unlike on Docker
|
|
551
|
+
(where it runs under `--init`), it has no init process on Kubernetes; it
|
|
552
|
+
stops when the pod is deleted. Its
|
|
553
|
+
`startupProbe` runs `test -S /run/wardby/tool/runner.sock`, so the keeper
|
|
554
|
+
and worker wait for the socket to exist before they start. The socket
|
|
555
|
+
directory is its own small memory-backed `emptyDir` (`tool-socket`),
|
|
556
|
+
mounted by the tool runner and the worker only, never a subdirectory of the
|
|
557
|
+
disk-backed storage volume: under gVisor a Unix socket bound on a volume
|
|
558
|
+
that isn't shared across the sandbox is invisible to the other container,
|
|
559
|
+
while a memory-backed `emptyDir` that two containers mount is one shared
|
|
560
|
+
tmpfs (GKE Autopilot annotates it `share: pod`). Before it calls the model,
|
|
561
|
+
the Claude worker connects to the socket once and fails the run as
|
|
562
|
+
`worker_tool_runner_unreachable` if it can't, so a run never proceeds
|
|
563
|
+
without its command tool. Of the pod's
|
|
564
|
+
resources, the tool runner gets a fixed 0.25 CPU and a third of the run's
|
|
565
|
+
memory (clamped between 128 and 512 MiB); the agent (`worker`) gets the
|
|
566
|
+
rest of both. The two also split the worker's 1024 MiB ephemeral-storage
|
|
567
|
+
reservation: 256 MiB to the tool runner, 768 MiB to the agent. GKE
|
|
568
|
+
Autopilot may round each container's requests up to its own minimums, so
|
|
569
|
+
there the split is approximate.
|
|
570
|
+
- **Service sidecars** (only for a run with services, see
|
|
571
|
+
[coding-services.md](coding-services.md)): one init container per service,
|
|
572
|
+
`service-<name>`, with `restartPolicy: Always` and a `startupProbe` from its
|
|
573
|
+
catalog entry — this native-sidecar shape (an init container that keeps
|
|
574
|
+
running) needs Kubernetes 1.29 or later — after `storage-init` (and, for a
|
|
575
|
+
Claude Code run, `tool-runner`) and before `keeper` and `worker`, so
|
|
576
|
+
neither starts until every service is ready. Each runs its catalog image,
|
|
577
|
+
pinned by digest, with the same security context as every other container
|
|
578
|
+
(uid 10001, read-only root filesystem, no privilege escalation, all
|
|
579
|
+
capabilities dropped), `emptyDir` volumes at its data and writable paths, and
|
|
580
|
+
requests equal to limits. It shares the pod's network namespace, so the worker
|
|
581
|
+
reaches it on `127.0.0.1` and the run's NetworkPolicy is unchanged: a service
|
|
582
|
+
can reach nothing the worker cannot. For a Claude Code run, the tool runner
|
|
583
|
+
reaches it on `127.0.0.1` too, and gets the same catalog `testEnv` the worker
|
|
584
|
+
would. The worker never receives a service's own
|
|
585
|
+
environment, only the catalog's `testEnv`. Attestation compares sidecars like
|
|
586
|
+
every other container.
|
|
587
|
+
- **`keeper`**: trusted, holds the one `storage` volume (an `emptyDir` sized
|
|
588
|
+
`spec.limits.diskMb` MiB, **disk-backed**, not `medium: Memory`) open for
|
|
589
|
+
the pod's life; the seam streams the workspace and input artifact in and
|
|
590
|
+
the output artifact out through it via `kubectl exec`-style calls
|
|
591
|
+
(`tar` in/out). Readiness probe: the `output` subdirectory exists.
|
|
592
|
+
- **`worker`**: untrusted. Its command is overridden to a small polling gate
|
|
593
|
+
(`WORKER_GATE`) that waits for `/run/wardby/input/.seeded` before
|
|
594
|
+
`import()`-ing the image's real entrypoint — for a Claude Code run, the
|
|
595
|
+
Claude entrypoint, which never mounts `/workspace` (only `/run/wardby/tool`,
|
|
596
|
+
to reach the socket, plus `input` and `output`). A pod's containers all
|
|
597
|
+
start together — Kubernetes has no "start this container later" — so the
|
|
598
|
+
gate is what makes "seed first, run second" possible without a native
|
|
599
|
+
sidecar (which Kubernetes terminates when the main container exits, killing
|
|
600
|
+
the keeper before result collection).
|
|
601
|
+
- `/tmp` and `/home/wardby` are small `medium: Memory` `emptyDir`s (bounded
|
|
602
|
+
`min(64, max(16, memoryMb/8))` MiB), matching Docker's bounded tmpfs mounts;
|
|
603
|
+
a Claude Code run's tool runner gets its own pair, sized the same way from
|
|
604
|
+
its own memory share. Neither is where the agent's commands get `TMPDIR`:
|
|
605
|
+
that points at the workspace's `.cache/tmp` on the disk-backed `storage`
|
|
606
|
+
volume instead, so it scales with `workspaceDiskMb` rather than this tiny
|
|
607
|
+
in-memory scratch.
|
|
608
|
+
- `dnsPolicy: None` with `dnsConfig.nameservers: ["127.0.0.1"]` — **no DNS is
|
|
609
|
+
configured for worker/agent pods at all.** The proxy is reached by name
|
|
610
|
+
(`WARDBY_PROXY_URL=http://wardby-proxy:8787` — `CODING_PROXY_ALIAS` in
|
|
611
|
+
`docker-isolation.ts`, the same alias the Docker launcher uses; this is
|
|
612
|
+
distinct from the `KUBERNETES_PROXY_SERVICE` Kubernetes Service name,
|
|
613
|
+
`wardby-coding-proxy` by default — the proxy checks the `Host` header)
|
|
614
|
+
only because the pod's `hostAliases` maps `wardby-proxy` directly to the
|
|
615
|
+
proxy Service's ClusterIP — closing DNS as an exfiltration channel
|
|
616
|
+
without needing a resolver at all.
|
|
617
|
+
- `activeDeadlineSeconds = spec.timeoutSec + POD_DEADLINE_GRACE_SECONDS`
|
|
618
|
+
(300s) — a **backstop only**. The launcher enforces the real wall-clock
|
|
619
|
+
deadline itself (`observePod` in `kubernetes.ts`); the extra 300s exists so
|
|
620
|
+
the keeper survives long enough after the worker's deadline for result
|
|
621
|
+
collection to still succeed. A worker that exits 0 counts as `succeeded`
|
|
622
|
+
only if all three hold: the pod's status reason isn't `DeadlineExceeded`,
|
|
623
|
+
the pod isn't being deleted (`metadata.deletionTimestamp` unset), and the
|
|
624
|
+
terminated container's `finishedAt` is at or before
|
|
625
|
+
`deadlineAt + 5s` (clock-skew slack). This closes off a SIGTERM-trapping
|
|
626
|
+
worker turning a deadline kill into a fake success, while still accepting
|
|
627
|
+
a run that genuinely finished just before its deadline and was only
|
|
628
|
+
observed after it.
|
|
629
|
+
- Everything else matches the Docker policy's spirit: uid/gid 10001,
|
|
630
|
+
`runAsNonRoot`, all capabilities dropped, `allowPrivilegeEscalation:
|
|
631
|
+
false`, seccomp `RuntimeDefault`, read-only root filesystem, no host
|
|
632
|
+
network/PID/IPC, `automountServiceAccountToken: false`, a dedicated
|
|
633
|
+
no-RBAC service account (`wardby-coding-worker`).
|
|
634
|
+
|
|
635
|
+
Per-run objects, all labeled `app.kubernetes.io/managed-by: wardby`,
|
|
636
|
+
`wardby.io/component: coding-run`, `wardby.io/run-sha256: <sha256(runId)
|
|
637
|
+
prefix>`, named `wardby-run-<token>` (`<token>` = first 20 hex chars of
|
|
638
|
+
`sha256(runId)`):
|
|
639
|
+
|
|
640
|
+
- The **pod** and its **NetworkPolicy** (same name).
|
|
641
|
+
- A **capability Secret** (`wardby-run-<token>-cap`) holding the run's proxy
|
|
642
|
+
capability, injected only via `secretKeyRef` — never in the pod spec,
|
|
643
|
+
command, or arguments. For a Claude Code run, the same Secret also carries
|
|
644
|
+
a `tool-setup` key (the tool runner's `WARDBY_TOOL_SETUP` JSON), likewise
|
|
645
|
+
injected only via `secretKeyRef`.
|
|
646
|
+
- A **record ConfigMap** (also `wardby-run-<token>`) holding all job state —
|
|
647
|
+
phase, deadline, result — updated with optimistic concurrency
|
|
648
|
+
(`resourceVersion`). This replaces local state files and in-process
|
|
649
|
+
timers entirely: any control-plane replica can observe, collect, stop, or
|
|
650
|
+
remove any run, and a restarted process loses nothing.
|
|
651
|
+
|
|
652
|
+
**Record ConfigMaps are retained as tombstones by design.**
|
|
653
|
+
`remove()` deletes the pod, NetworkPolicy, and capability Secret, but
|
|
654
|
+
deliberately _rewrites the record to `phase: "removed"` instead of deleting
|
|
655
|
+
it_ (`kubernetes.ts`, `remove()`) — the contract is that a removed run is
|
|
656
|
+
never relaunched, and the record is what a later `launch()` call for the
|
|
657
|
+
same run ID checks. Wardby does not yet garbage-collect old tombstones, so they
|
|
658
|
+
accumulate until an operator removes them. **A launch that fails during
|
|
659
|
+
provisioning accumulates a record too**,
|
|
660
|
+
not just a `remove()`d run's tombstone: the failure path writes `phase:
|
|
661
|
+
"failed"` and the executor never calls `remove()` for a launch that threw,
|
|
662
|
+
so every failed launch leaves a permanent record as well. On a busy cluster
|
|
663
|
+
this is etcd growth to budget for operationally, not a correctness or
|
|
664
|
+
security problem — see "Known gaps" below.
|
|
665
|
+
|
|
666
|
+
**The executor persists a run's job handle before calling `jobs.launch()`,
|
|
667
|
+
closing the crash window that would otherwise orphan a running pod.** The
|
|
668
|
+
Kubernetes handle (`{ backend: "kubernetes", id: "<namespace>/<token>" }`)
|
|
669
|
+
is fully derivable from the run ID alone, with no cluster call, so it can be
|
|
670
|
+
(and is) written to the run record before the launcher ever creates
|
|
671
|
+
anything. If the control plane crashes or is replaced mid-launch — after the
|
|
672
|
+
pod has been attested and the worker gate has opened, but before the old
|
|
673
|
+
code path would have recorded the handle — the run's handle is already on
|
|
674
|
+
record, so restart reconciliation can find, stop, and clean up the pod,
|
|
675
|
+
NetworkPolicy, and capability Secret through the normal `abandon()` path
|
|
676
|
+
instead of leaking them permanently.
|
|
677
|
+
|
|
678
|
+
### Attestation — deny-by-default, fail closed
|
|
679
|
+
|
|
680
|
+
Before the worker gate ever opens, the launcher reads back the created pod
|
|
681
|
+
and NetworkPolicy and compares them against the canonical builder's output
|
|
682
|
+
with `assertRunPodMatches` / `assertRunNetworkPolicyMatches`
|
|
683
|
+
(`kubernetes-isolation.ts`). This is a **full, canonical, deep comparison of
|
|
684
|
+
the entire spec, labels, and annotations** — not an allowlist of fields the
|
|
685
|
+
launcher expects to see. An early allowlist-based version of this comparator
|
|
686
|
+
was replaced during implementation review specifically because an allowlist
|
|
687
|
+
silently accepts anything it forgot to check (lifecycle hooks,
|
|
688
|
+
liveness/readiness/startup probes whose `httpGet.host` can reach the node's
|
|
689
|
+
link-local metadata endpoint bypassing the NetworkPolicy, `procMount`,
|
|
690
|
+
`seLinuxOptions`, extra tolerations, `nodeSelector`, stray annotations, ...).
|
|
691
|
+
The only normalization applied before comparing is an explicit, narrow list
|
|
692
|
+
of transformations the Kubernetes API server itself is known to perform on
|
|
693
|
+
write/read — never a loosening of what's compared:
|
|
694
|
+
|
|
695
|
+
- Dropping `schedulerName`, `nodeName`, `priority`, `preemptionPolicy` (server-assigned).
|
|
696
|
+
- Dropping `hostNetwork`/`hostPID`/`hostIPC` when `false`, and an empty
|
|
697
|
+
NetworkPolicy `ingress: []`, because Go's `omitempty` drops a zero-value
|
|
698
|
+
bool or empty slice on serialization — a real read-back never carries
|
|
699
|
+
these fields at their false/empty value, only when true/non-empty.
|
|
700
|
+
- Removing the mirrored `serviceAccount` field when it equals
|
|
701
|
+
`serviceAccountName` (and failing closed if it doesn't).
|
|
702
|
+
- Removing exactly the two well-known `NoExecute` node-health tolerations
|
|
703
|
+
(`node.kubernetes.io/not-ready` / `unreachable`, 300s) every pod gets by
|
|
704
|
+
admission-time default — any other toleration must match exactly.
|
|
705
|
+
- Dropping container `terminationMessagePath`/`terminationMessagePolicy`/`imagePullPolicy`
|
|
706
|
+
and probe threshold/period defaults, and normalizing CPU/memory quantities
|
|
707
|
+
to a canonical millicore/byte count (so `"1"`, `"1.0"`, and `"1000m"`
|
|
708
|
+
compare equal) — with a non-integer-at-that-scale value mapped to a
|
|
709
|
+
sentinel that can never equal a real value, so quantity drift fails closed
|
|
710
|
+
instead of rounding two different resources together.
|
|
711
|
+
- Dropping `mountPropagation: "None"` and an `emptyDir.medium: ""`.
|
|
712
|
+
|
|
713
|
+
Any other difference — anything not on this list — fails the launch closed
|
|
714
|
+
with `kubernetes_isolation_unsupported`. No fallback to a weaker profile.
|
|
715
|
+
|
|
716
|
+
On top of that, a **platform profile** (`KUBERNETES_PLATFORM`) may forgive a
|
|
717
|
+
named, narrow set of mutations its managed admission chain is known to make —
|
|
718
|
+
for `gke-autopilot`: annotations under `autopilot.gke.io/` and `dev.gvisor.`,
|
|
719
|
+
labels under `autopilot.gke.io/` and `topology.kubernetes.io/`, the gVisor
|
|
720
|
+
`nodeSelector`, and two exact tolerations. A profile can only delete named keys
|
|
721
|
+
from both operands before the comparison; it can never disable or short-circuit
|
|
722
|
+
it, and `generic` (the default, and what an omitted argument selects) forgives
|
|
723
|
+
nothing. Most of that list was captured from a real cluster with
|
|
724
|
+
`npm run capture:autopilot`, which submits the pod with `dryRun=All`. That
|
|
725
|
+
capture has a structural blind spot worth knowing: a dry run is admission only
|
|
726
|
+
and never schedules, so anything the platform stamps on **after binding** —
|
|
727
|
+
GKE's `topology.kubernetes.io/{region,zone}`, taken from the node the pod
|
|
728
|
+
landed on — cannot appear in it. That case reached a live Autopilot launch with
|
|
729
|
+
every dry-run-derived check passing, and is now pinned by tests. Treat a
|
|
730
|
+
captured fixture as a lower bound on what a platform mutates.
|
|
731
|
+
|
|
732
|
+
### The enforcement gate
|
|
733
|
+
|
|
734
|
+
The Kubernetes API can create a NetworkPolicy object without that policy
|
|
735
|
+
being enforced yet — CNIs (including `kind`'s default kindnet) program a new
|
|
736
|
+
pod's policy a few seconds _after_ the pod starts, not atomically with pod
|
|
737
|
+
creation. The delay is observable on real clusters and threatens every run,
|
|
738
|
+
not just the harness: the worker gate could otherwise
|
|
739
|
+
open on a pod whose isolation isn't active yet.
|
|
740
|
+
|
|
741
|
+
The fix, before seeding or opening the worker gate: the launcher execs into
|
|
742
|
+
the keeper (which shares the pod's network namespace with the worker) one
|
|
743
|
+
`node -e` probe that measures **both of the coding proxy's ports against the
|
|
744
|
+
proxy Service's ClusterIP, in the same pass** — `8787` (the proxy itself, which
|
|
745
|
+
the run policy permits) and `8788` (the deny port, which no run policy ever
|
|
746
|
+
permits). Only the outcome **(8787 connected, 8788 blocked)** counts toward the
|
|
747
|
+
streak.
|
|
748
|
+
|
|
749
|
+
Stated exactly, that outcome proves: **the SYN to 8788 was dropped somewhere on
|
|
750
|
+
the path, while the same destination answered on 8787.** That the drop was the
|
|
751
|
+
_run pod's own egress policy_ does not follow from the measurement alone — it
|
|
752
|
+
follows from the proxy admitting run pods on 8788 at its own ingress, so that no
|
|
753
|
+
other hop is left to drop it. That precondition used to be asserted by manifest
|
|
754
|
+
and checked nowhere, and was falsified on a live cluster: with the proxy's policy
|
|
755
|
+
admitting 8787 only, a prober with no policy at all and full internet egress read
|
|
756
|
+
**proven**, because the proxy's own ingress dropped the packet. The `proxy-service`
|
|
757
|
+
preflight check now reads the proxy's NetworkPolicy and verifies that rule, so the
|
|
758
|
+
attribution is checked rather than assumed.
|
|
759
|
+
|
|
760
|
+
Both halves are required, and measuring only one port would be unsound. A
|
|
761
|
+
NetworkPolicy denial **drops** the packet rather than rejecting it — GKE
|
|
762
|
+
Dataplane V2 (Cilium), which Autopilot runs, always drops — so "8788 did not
|
|
763
|
+
answer" on its own is equally consistent with "the proxy is gone and no policy
|
|
764
|
+
exists at all", and a run would be released onto an unpoliced network.
|
|
765
|
+
Requiring 8787 to connect in the _same_ exec turns "something is listening"
|
|
766
|
+
from a control-plane inference (which is stale the moment it is read) into a
|
|
767
|
+
fact this pod just observed, at the instant of the blocked observation. The
|
|
768
|
+
proxy's own policy deliberately **allows** ingress on 8788 from run pods:
|
|
769
|
+
ingress is enforced at the destination, so denying it there would make a run
|
|
770
|
+
pod whose own egress policy was not yet programmed read as "blocked" — which is
|
|
771
|
+
precisely the live falsification above, and why that rule is now verified at
|
|
772
|
+
preflight instead of trusted.
|
|
773
|
+
|
|
774
|
+
**"Blocked" means a timeout specifically, not "did not connect".** Pairing the
|
|
775
|
+
two ports only rules out "the whole proxy pod is dead"; it does not rule out
|
|
776
|
+
the deny _listener_ being unserved while the pod is otherwise healthy. This was
|
|
777
|
+
reproduced on a live cluster: in a namespace with no NetworkPolicy at all,
|
|
778
|
+
against a pod listening on 8787 and serving nothing on 8788, a probe that
|
|
779
|
+
treated any failed connect as "blocked" exited **proven** while it had full
|
|
780
|
+
internet egress. No control-plane read closes this — an Endpoints subset port is
|
|
781
|
+
the Service's numeric `targetPort`, not evidence that anything is bound — so the
|
|
782
|
+
probe distinguishes the two socket outcomes itself: a **timeout** means the
|
|
783
|
+
packet was dropped (a policy), while a **refusal** (RST / `ECONNREFUSED`) proves
|
|
784
|
+
the SYN reached the destination host, on every dataplane, since a drop cannot
|
|
785
|
+
produce an RST. A refused deny port is therefore reachable-but-unserved and is
|
|
786
|
+
reported as `kubernetes_policy_witness_unserved`, never as proven.
|
|
787
|
+
|
|
788
|
+
The intended consequence: on a **reject-style** CNI a genuine policy denial also
|
|
789
|
+
arrives as an RST, so such a cluster now fails closed here rather than passing
|
|
790
|
+
vacuously. That is the correct direction — a witness that cannot tell "denied"
|
|
791
|
+
from "unserved" is not a witness — and such a cluster needs a different one.
|
|
792
|
+
|
|
793
|
+
It requires **3 consecutive proven results, 500ms apart** (anything else
|
|
794
|
+
resets the streak — this guards against a single dropped SYN packet on an
|
|
795
|
+
allowed path being misread as "policy enforced"), bounded by
|
|
796
|
+
`enforcementTimeoutMs` (default 30,000ms — configurable via
|
|
797
|
+
`KubernetesJobLauncherOptions.enforcementTimeoutMs`; a drop-style CNI can
|
|
798
|
+
need close to this whole window). The verdict at the bound comes from the
|
|
799
|
+
_last_ probe — not from whether any probe was ever unavailable, so an early
|
|
800
|
+
blip while the pod's networking came up does not misdirect the operator — and
|
|
801
|
+
the three non-proven outcomes stay distinct, because each sends an operator
|
|
802
|
+
somewhere different:
|
|
803
|
+
|
|
804
|
+
| Last probe | Verdict | Where to look |
|
|
805
|
+
| ---------------- | --------------------------------------- | ---------------------------------------- |
|
|
806
|
+
| 8788 connected | `kubernetes_policy_not_enforced` | the CNI: no policy, or not port-scoped |
|
|
807
|
+
| 8787 unreachable | `kubernetes_policy_witness_unavailable` | the proxy pod / its Service |
|
|
808
|
+
| 8788 refused | `kubernetes_policy_witness_unserved` | the deny listener, or a reject-style CNI |
|
|
809
|
+
|
|
810
|
+
Any other exit code means the probe never ran to completion (a crash, a missing
|
|
811
|
+
interpreter, an OOM-killed keeper), so nothing was measured and nothing about
|
|
812
|
+
the policy can be concluded: that is `kubernetes_policy_probe_unusable`, and it
|
|
813
|
+
carries the observed exit code. Every one of these messages names the probed
|
|
814
|
+
address and the exit code it saw, because an operator reads the error, not this
|
|
815
|
+
page.
|
|
816
|
+
|
|
817
|
+
This wait happens inside the pod's overall ready-timeout window, not on top of
|
|
818
|
+
it.
|
|
819
|
+
|
|
820
|
+
`wardby coding preflight`'s canary pod waits the same way before running its
|
|
821
|
+
probes, for the same reason.
|
|
822
|
+
|
|
823
|
+
### Preflight
|
|
824
|
+
|
|
825
|
+
`kubernetesPreflight` / `runKubernetesPreflight`
|
|
826
|
+
(`src/providers/jobs/kubernetes-preflight.ts`) run **five checks in order**,
|
|
827
|
+
each producing `kubernetes_isolation_unsupported:<check>` on failure (or
|
|
828
|
+
`:timeout` if the whole preflight — cleanup included — exceeds `timeoutMs`,
|
|
829
|
+
default 90,000ms):
|
|
830
|
+
|
|
831
|
+
1. `platform` — pure configuration, checked before any cluster API call:
|
|
832
|
+
`assertPlatformConfig` (`src/providers/jobs/kubernetes-platform.ts`) refuses
|
|
833
|
+
a deployment that cannot work under `KUBERNETES_PLATFORM`. Under
|
|
834
|
+
`gke-autopilot` this means `KUBERNETES_RUNTIME_CLASS` must be `gvisor`
|
|
835
|
+
(gVisor is mandatory there — an unset or different runtime class is refused,
|
|
836
|
+
not warned about), and the effective `CODING_MAX_DISK_MB` must leave room
|
|
837
|
+
for the worker container's 1 GiB reservation inside Autopilot's 10 GiB pod
|
|
838
|
+
ephemeral-storage ceiling. The same assertion runs again at process
|
|
839
|
+
start-up in `buildConfiguredExecutor`
|
|
840
|
+
(`src/providers/executor/composition.ts`), so an unrunnable configuration
|
|
841
|
+
fails the process immediately rather than waiting for the first coding run
|
|
842
|
+
or the next `wardby coding preflight` invocation to discover it.
|
|
843
|
+
2. `namespace` — the configured namespace exists.
|
|
844
|
+
3. `proxy-service` — the proxy Service exists, has a ClusterIP, **exposes both
|
|
845
|
+
the proxy port (8787) and the deny port (8788)** over TCP, has at least one
|
|
846
|
+
ready endpoint serving both, **and the proxy's own NetworkPolicy admits
|
|
847
|
+
`wardby.io/component: coding-run` on 8788** — failing closed with
|
|
848
|
+
`kubernetes_isolation_unsupported:proxy-service` otherwise. That last clause
|
|
849
|
+
is the attribution precondition: a dropped connect to 8788 only indicts the
|
|
850
|
+
run pod's own egress policy if no other hop would have dropped it, and
|
|
851
|
+
ingress is enforced at the destination, so without it deleting one line from
|
|
852
|
+
an overlay's proxy policy makes every run read as enforced. The deny port is
|
|
853
|
+
the enforcement witness: a second listener on the proxy
|
|
854
|
+
(`src/providers/coding-proxy/deny-port.ts`) that serves nothing and that no
|
|
855
|
+
run's NetworkPolicy ever permits. A run pod that reaches 8787 but not 8788
|
|
856
|
+
has proven its policy is both programmed and port-scoped. It replaces the old
|
|
857
|
+
`cluster-dns` witness, which does not exist on GKE Autopilot (Cloud DNS is the
|
|
858
|
+
only provider there, so no kube-dns pods run) and which made the launcher read
|
|
859
|
+
`kube-system`.
|
|
860
|
+
4. `worker-image` — `CODING_WORKER_IMAGE` is a registry digest.
|
|
861
|
+
5. `canary` — creates a real run pod + NetworkPolicy from the same builders
|
|
862
|
+
as a live run, running a script that waits for policy enforcement (as
|
|
863
|
+
above) then attempts DNS resolution, a connect to the proxy's deny port,
|
|
864
|
+
the internet (`1.1.1.1:443`), the metadata server, and the proxy itself —
|
|
865
|
+
requiring every one of the first four to fail and the proxy connect to
|
|
866
|
+
succeed. Any other outcome, or a canary pod that itself fails to
|
|
867
|
+
schedule/run, is `kubernetes_isolation_unsupported:canary`.
|
|
868
|
+
|
|
869
|
+
This whole preflight is **memoized per launcher instance and its failure is
|
|
870
|
+
sticky**: `KubernetesJobLauncher.runPreflight()` caches the first call's
|
|
871
|
+
promise (`this.preflightResult ??= ...`, `kubernetes.ts:472-487`), including
|
|
872
|
+
a rejection — so once a launcher process has seen preflight fail, every
|
|
873
|
+
subsequent `launch()` in that process fails immediately with the same error
|
|
874
|
+
without re-probing the cluster. A fresh preflight requires a new process
|
|
875
|
+
(or, from the CLI, a fresh `wardby coding preflight` invocation, which is
|
|
876
|
+
not memoized).
|
|
877
|
+
|
|
878
|
+
**A hung pod create during preflight can leave a preflight pod and its
|
|
879
|
+
NetworkPolicy behind.** If `createPod` never settles (rather than failing),
|
|
880
|
+
the preflight's own timeout still fires and the caller sees `:timeout`, but
|
|
881
|
+
cleanup for that pod/policy is deferred to whenever the stuck create call
|
|
882
|
+
eventually resolves (`tracked`/`lateCleanup` in `kubernetes-preflight.ts`) —
|
|
883
|
+
if it never does, the objects are never removed automatically. **A canary
|
|
884
|
+
pod and policy are not distinguishable by name from a real run's:** both are
|
|
885
|
+
named `wardby-run-<runId's sha256 prefix>` (`kubernetesRunNames`,
|
|
886
|
+
`kubernetes-isolation.ts:90-96`, used by the preflight at
|
|
887
|
+
`kubernetes-preflight.ts:218,241`) and carry the same
|
|
888
|
+
`wardby.io/component: coding-run` label as a live run
|
|
889
|
+
(`kubernetes-isolation.ts:98-104,249`) — there is no
|
|
890
|
+
`wardby-run-preflight-*` naming pattern. After a `:timeout` failure,
|
|
891
|
+
operators should instead list every object with
|
|
892
|
+
`wardby.io/component=coding-run` in the namespace and cross-reference
|
|
893
|
+
against the run record ConfigMaps that legitimately exist (a stray canary
|
|
894
|
+
object has no corresponding non-tombstoned run record, since preflight
|
|
895
|
+
never creates one). Giving preflight objects a distinguishing label (e.g.
|
|
896
|
+
`wardby.io/component: coding-preflight`) would make this a direct label query;
|
|
897
|
+
see "Known limitations" below.
|
|
898
|
+
|
|
899
|
+
### RBAC actually required
|
|
900
|
+
|
|
901
|
+
The launcher's `ClientNodeKubernetesApi` issues a narrow, specific set of
|
|
902
|
+
calls, and `deploy/kind-coding/manifests/base/` grants exactly that (no
|
|
903
|
+
`list`/`watch` on pods, no `get` on secrets — the seam never reads one
|
|
904
|
+
back):
|
|
905
|
+
|
|
906
|
+
- Namespace `Role` **`wardby-coding-launcher`** (in the coding namespace):
|
|
907
|
+
`pods` create/get/delete, `pods/exec` create/get, `pods/log` get,
|
|
908
|
+
`secrets` create/delete, `configmaps` create/get/update, `networkpolicies`
|
|
909
|
+
create/get/delete, `services` get, and `endpoints` get scoped by
|
|
910
|
+
`resourceNames: ["wardby-coding-proxy"]` — the one read `readProxyWitness`
|
|
911
|
+
needs to prove the deny port is exposed with a ready backend. **Nothing in
|
|
912
|
+
`kube-system` any more:** the old `wardby-coding-dns-reader` Role is gone
|
|
913
|
+
with the `cluster-dns` check.
|
|
914
|
+
- `ClusterRole` **`wardby-coding-namespace-reader`**: `get` on the
|
|
915
|
+
cluster-scoped `namespaces` resource, `resourceNames: [<the namespace>]`.
|
|
916
|
+
This one has to be cluster-scoped — no namespaced `Role` can grant `get`
|
|
917
|
+
on `namespaces` — but it's still scoped down to the one namespace via
|
|
918
|
+
`resourceNames`, so the launcher identity can't discover any other
|
|
919
|
+
namespace's existence.
|
|
920
|
+
|
|
921
|
+
(`deploy/kind-coding/` binds none of these to a service account — the local
|
|
922
|
+
harness runs every command against your own admin kubeconfig. A production
|
|
923
|
+
overlay, e.g. GKE, binds these two to the control plane's identity.)
|
|
924
|
+
|
|
925
|
+
**Running the real-cluster integration suite needs more than this.**
|
|
926
|
+
`npm run test:kubernetes` (`kubernetes.integration.test.ts`) uses a raw
|
|
927
|
+
`@kubernetes/client-node` client directly, alongside the launcher's own
|
|
928
|
+
`KubernetesApi` seam, for two things outside what the launcher itself ever
|
|
929
|
+
does: it reads `Endpoints` objects (`get endpoints`) to resolve kube-dns's
|
|
930
|
+
and the proxy pod's addresses for its isolation probes, and its cleanup
|
|
931
|
+
deletes each run's record ConfigMap (`delete configmaps`) — the launcher
|
|
932
|
+
intentionally never deletes that object (see "kept as tombstones" above), so
|
|
933
|
+
`deleteConfigMap` isn't even part of the `KubernetesApi` seam; the test goes
|
|
934
|
+
straight to the library. The committed `wardby-coding-launcher` Role grants
|
|
935
|
+
neither verb. On `kind` this gap is invisible because the suite runs against
|
|
936
|
+
the admin kubeconfig; a kubeconfig scoped to only the two launcher roles
|
|
937
|
+
above needs `get endpoints` (coding namespace and `kube-system`) and
|
|
938
|
+
`delete configmaps` (coding namespace) added before the integration suite
|
|
939
|
+
will pass against it.
|
|
940
|
+
|
|
941
|
+
### Diagnostics
|
|
942
|
+
|
|
943
|
+
On a failed run, the launcher reads only the failed worker container's last
|
|
944
|
+
8 log lines (bounded to 128 KiB, room for eight full-size
|
|
945
|
+
[debug trace](#debug-trace) lines) through the Kubernetes API
|
|
946
|
+
(`pods/log`), and keeps only a code matching the existing
|
|
947
|
+
`SAFE_WORKER_DIAGNOSTIC` pattern (imported from the Docker launcher) — the
|
|
948
|
+
raw text itself is never stored, logged, or returned
|
|
949
|
+
(`readWorkerDiagnostic`, `kubernetes.ts`). For `coding_output_invalid`, the
|
|
950
|
+
same line may also carry `issues`: up to 8 `path:code` entries (for example
|
|
951
|
+
`tag:invalid_string`) naming which output-schema fields failed, never their
|
|
952
|
+
values. The list is kept only when every entry matches
|
|
953
|
+
`SAFE_CODING_OUTPUT_ISSUE`, and the executor logs it next to the diagnostic id.
|
|
954
|
+
|
|
955
|
+
### Known limitations
|
|
956
|
+
|
|
957
|
+
- **No gVisor / per-pod process (PID) limit on `kind`.** `kind` has no
|
|
958
|
+
runtime-class sandboxing; `KUBERNETES_RUNTIME_CLASS` is unset in the local
|
|
959
|
+
harness and the launcher logs `kubernetes_runtime_class_unset` once per
|
|
960
|
+
launch as a loud "development cluster" warning. Use the GKE Autopilot
|
|
961
|
+
overlay with its required `gvisor` runtime class for production.
|
|
962
|
+
- **The real-cluster integration coverage is partial.** It covers the contract
|
|
963
|
+
suite against the real API, the enforcement gate, and a full run lifecycle.
|
|
964
|
+
It does not yet cover the isolation acceptance suite's
|
|
965
|
+
OOM/disk-full/wall-clock containment assertions, "canary fails when a
|
|
966
|
+
policy is removed", or Claude Code's tool-runner sidecar (there is no
|
|
967
|
+
separate tool pod: it shares the run pod's network namespace and
|
|
968
|
+
NetworkPolicy with the worker, so the coverage needed is different from a
|
|
969
|
+
second pod's isolation).
|
|
970
|
+
- **No end-to-end backpressure or host-side archive-size cap** once the
|
|
971
|
+
exec WebSocket for a seed/collect transfer is connected — pre-existing,
|
|
972
|
+
documented, not a regression of this milestone.
|
|
973
|
+
- **An over-ceiling `workspaceDiskMb` fails late and generically** — see
|
|
974
|
+
"Control Plane Configuration" above.
|
|
975
|
+
- **Namespace handling must remain explicit in every overlay.**
|
|
976
|
+
`deploy/kind-coding/manifests/base/kustomization.yaml`
|
|
977
|
+
deliberately has **no top-level `namespace:` override**, because
|
|
978
|
+
kustomize's namespace transformer would force `metadata.namespace` onto
|
|
979
|
+
every namespaced resource it lists. Every manifest instead sets its own
|
|
980
|
+
`metadata.namespace` explicitly. Any overlay author copying this harness for
|
|
981
|
+
another cluster must do the same.
|
|
982
|
+
- **Per-run record ConfigMap GC is unimplemented.** Both `remove()`'s
|
|
983
|
+
deliberate tombstones and a failed launch's records (see above)
|
|
984
|
+
accumulate, one small ConfigMap per run, with nothing that automatically
|
|
985
|
+
deletes them. Operators need a retention job for a long-lived deployment;
|
|
986
|
+
this is an operational concern rather than a correctness or security issue.
|
|
987
|
+
- **Autopilot dry-run capture is a lower bound.** Admission dry runs cannot
|
|
988
|
+
observe topology labels added after pod scheduling. The reviewed platform
|
|
989
|
+
profile and tests include those known labels, but every target cluster must
|
|
990
|
+
still pass preflight before accepting work.
|
|
991
|
+
- **A distinguishing label for preflight/canary objects.** Canary pods and
|
|
992
|
+
policies are currently named and labeled identically to real run objects
|
|
993
|
+
(`wardby.io/component: coding-run`), which is why the troubleshooting
|
|
994
|
+
guidance above can only recommend cross-referencing against run records
|
|
995
|
+
rather than a direct label query. Giving preflight objects their own
|
|
996
|
+
`wardby.io/component: coding-preflight` label would fix this and would
|
|
997
|
+
also let a future orphan reaper (the ConfigMap GC item above, extended to
|
|
998
|
+
pods) tell a canary apart from a live run.
|
|
999
|
+
- **The real-cluster integration suite still uses the deprecated core/v1
|
|
1000
|
+
`Endpoints` API** (`kubernetes.integration.test.ts`) rather than
|
|
1001
|
+
`discovery.k8s.io/v1` `EndpointSlice`. `Endpoints` is deprecated, not yet
|
|
1002
|
+
removed, and this is test-only code, but the migration is a tracked
|
|
1003
|
+
remaining compatibility task.
|
|
1004
|
+
|
|
1005
|
+
### Troubleshooting: failure codes
|
|
1006
|
+
|
|
1007
|
+
| Code | Meaning |
|
|
1008
|
+
| ---------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
1009
|
+
| `kubernetes_isolation_unsupported:<check>` | A preflight check failed; `<check>` is one of `platform`, `namespace`, `proxy-service`, `worker-image`, `canary`. Sticky for the launcher's process lifetime once seen (see Preflight above). |
|
|
1010
|
+
| `kubernetes_isolation_unsupported:timeout` | The whole preflight (including cleanup) exceeded its timeout. |
|
|
1011
|
+
| `kubernetes_isolation_unsupported` | (No suffix) Attestation failure: the read-back pod or NetworkPolicy didn't canonically match the builder's output. |
|
|
1012
|
+
| `kubernetes_policy_not_enforced` | The run's NetworkPolicy wasn't observed enforced (8787 reachable, 8788 blocked) within `enforcementTimeoutMs`; the worker gate was never opened. |
|
|
1013
|
+
| `kubernetes_policy_witness_unavailable` | The last enforcement probe could not reach the proxy on 8787 at all, so nothing could be witnessed — the proxy or its Service is the thing to check, not the CNI. The worker gate was never opened. |
|
|
1014
|
+
| `kubernetes_policy_witness_unserved` | The last enforcement probe was **refused** (RST) on 8788 rather than dropped. The packet reached the host, so nothing is blocking the path and the witness proves nothing — the deny listener is not serving, or the CNI rejects instead of dropping. Fails closed; the worker gate was never opened. |
|
|
1015
|
+
| `kubernetes_proxy_witness_unusable: <reason>` | The proxy Service is not a usable enforcement witness (missing, no ClusterIP, a required port not exposed as TCP, or no ready endpoint serving both 8787 and 8788). Seen **unwrapped** like this from `provision`'s per-launch re-read, which runs after an earlier witness check already passed -- a supplied `preflight`, or this launcher's own memoized read, the likelier production sighting being a proxy that degrades after that read succeeded; the launcher's own memoized read reports the same condition wrapped as `kubernetes_isolation_unsupported` (with this string on `cause`), and the preflight reports it as `:proxy-service`. Fails closed before the pod is created. |
|
|
1016
|
+
| `kubernetes_pod_start_timeout` | The keeper didn't become ready within `readyTimeoutMs` (default 120s). Historically caused by the subPath root-ownership issue the `storage-init` init container now fixes; if seen again, check init-container status first. |
|
|
1017
|
+
| `kubernetes_pod_start_failed` | The pod (or its `storage-init` init container) failed outright rather than timing out. |
|
|
1018
|
+
| `kubernetes_tool_runner_failed` | (Claude Code only) The tool runner sidecar restarted, exited, or its image could not be pulled before the keeper started. A malformed `WARDBY_TOOL_SETUP` makes the tool runner exit before it is ready (`tool_setup_invalid` in its log), which surfaces as this code on Kubernetes and as `docker_tool_runner_not_ready` on Docker. |
|
|
1019
|
+
| `kubernetes_tool_runner_unready` | (Claude Code only) The tool runner sidecar never started (its socket startup probe never passed) by the pod-start bound (`readyTimeoutMs`). |
|
|
1020
|
+
| `worker_tool_runner_failed` | (Claude Code only; also seen on the Docker launcher) The tool runner died mid-run, after the pod/containers started successfully. |
|
|
1021
|
+
| `worker_tool_runner_unreachable` | (Claude Code only; both launchers) The Claude worker could not connect to the tool runner's socket before calling the model, or Claude Code reported its command tool as not connected. On Kubernetes, check that the pod has the `tool-socket` volume mounted in both the `worker` and `tool-runner` containers. |
|
|
1022
|
+
| `claude_tool_setup_too_large` | (Claude Code only) The run's tool-runner setup (registry settings plus every service test variable) would exceed the tool runner's bounds (1024 variables, 4096 bytes per value, 96 KiB in total), so the launch fails before the tool runner is created. |
|
|
1023
|
+
| `kubernetes_isolation_unsupported:tool-image-not-registry-digest` | (Claude Code only) `CODING_CLAUDE_TOOL_RUNNER_IMAGE` isn't a registry digest. Checked on every launch, not only preflight. |
|
|
1024
|
+
| `kubernetes_isolation_unsupported:claude-limits-too-small` | (Claude Code only) The run's `cpus`/`memoryMb` are too small to leave the tool runner its fixed floor: Claude Code needs `cpus ≥ 0.35` and `memoryMb ≥ 256`. |
|
|
1025
|
+
| `kubernetes_seed_failed` | Streaming the workspace or input artifact into the keeper failed. |
|
|
1026
|
+
| `kubernetes_workspace_archive_failed` / `kubernetes_result_artifact_invalid` | Collection (workspace or output artifact) failed or was invalid. |
|
|
1027
|
+
| `coding_workspace_disk_exceeds_limit` | The run's `workspaceDiskMb` exceeds `CODING_MAX_DISK_MB`; surfaces on the run record as `coding_failure_workspace:<id>`. |
|
|
1028
|
+
| `coding_service_unready:<name>` | A service sidecar restarted after its startup probe gave up, could not be pulled or started, or had not started when `readyTimeoutMs` ran out. The run fails with category `service_unready`. |
|
|
1029
|
+
|
|
1030
|
+
### Local harness
|
|
1031
|
+
|
|
1032
|
+
`deploy/kind-coding/` stands up a local `kind` cluster (with a local image
|
|
1033
|
+
registry so images are pulled by digest, as on GKE) that proves this
|
|
1034
|
+
launcher end to end, including real NetworkPolicy enforcement. See
|
|
1035
|
+
[`deploy/kind-coding/README.md`](../deploy/kind-coding/README.md) for
|
|
1036
|
+
prerequisites, the up/down scripts, and what each step does.
|
|
1037
|
+
|
|
1038
|
+
## Verification
|
|
1039
|
+
|
|
1040
|
+
Build the image and run the destructive, self-cleaning acceptance suite:
|
|
1041
|
+
|
|
1042
|
+
```sh
|
|
1043
|
+
docker build -f src/coding-worker/Dockerfile -t wardby-coding-worker:task8 .
|
|
1044
|
+
npm run test:docker-isolation
|
|
1045
|
+
npm run verify:claude-code
|
|
1046
|
+
```
|
|
1047
|
+
|
|
1048
|
+
Set `WARDBY_WORKER_IMAGE` to test another local tag. The runner resolves that
|
|
1049
|
+
tag to an immutable image ID before testing. The suite verifies effective
|
|
1050
|
+
Docker inspection, no default route, proxy-only connectivity, denied
|
|
1051
|
+
Docker-socket/host/metadata/localhost/public access, read-only mounts and
|
|
1052
|
+
rootfs, zero effective capabilities, seccomp and no-new-privileges, private PID
|
|
1053
|
+
1, PID exhaustion, OOM containment, disk ENOSPC, and wall-clock termination.
|
|
1054
|
+
|
|
1055
|
+
Related Docker references:
|
|
1056
|
+
|
|
1057
|
+
- Internal and isolated bridge networks: https://docs.docker.com/reference/cli/docker/network/create/
|
|
1058
|
+
- CPU, memory, swap, and PID controls: https://docs.docker.com/engine/containers/resource_constraints/
|
|
1059
|
+
- Seccomp and no-new-privileges: https://docs.docker.com/reference/cli/docker/container/run/
|
|
1060
|
+
- Tmpfs behavior and limits: https://docs.docker.com/engine/storage/tmpfs/
|
|
1061
|
+
- Volume subpaths and `nocopy`: https://docs.docker.com/engine/storage/volumes/
|