@edgehero/pi-dispatch 1.10.2 → 2.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (97) hide show
  1. package/.env.example +300 -148
  2. package/README.md +50 -0
  3. package/deploy/com.pi-dispatch.worker.plist +9 -3
  4. package/deploy/docker-compose.yml +49 -16
  5. package/deploy/egress-proxy.conf +32 -2
  6. package/deploy/nssm-install.cmd +12 -6
  7. package/deploy/pi-dispatch-egress-out.network +10 -0
  8. package/deploy/pi-dispatch-egress-proxy.container +50 -0
  9. package/deploy/pi-dispatch-netns-keeper.container +80 -0
  10. package/deploy/pi-dispatch-netns-keeper.network +18 -0
  11. package/deploy/pi-dispatch-valkey.container +51 -0
  12. package/deploy/pi-dispatch-valkey.network +16 -0
  13. package/deploy/receiver.service +6 -0
  14. package/deploy/worker-env-wrapper.cmd +11 -0
  15. package/deploy/worker-env-wrapper.sh +60 -34
  16. package/deploy/worker.service +18 -8
  17. package/package.json +14 -4
  18. package/src/azure-host.mjs +19 -0
  19. package/src/azure-identity.mjs +18 -2
  20. package/src/backend-conformance.mjs +71 -18
  21. package/src/backend-local.mjs +637 -21
  22. package/src/backend-podman.mjs +1168 -0
  23. package/src/backend-registry.mjs +86 -3
  24. package/src/backends.mjs +489 -37
  25. package/src/branch.mjs +7 -2
  26. package/src/cancel-cli.mjs +174 -0
  27. package/src/cancel-state.mjs +125 -0
  28. package/src/cli.mjs +188 -90
  29. package/src/config.mjs +503 -43
  30. package/src/connection.mjs +374 -8
  31. package/src/container-spec.mjs +102 -7
  32. package/src/daemon-facts.mjs +167 -0
  33. package/src/deployment-venue.mjs +158 -0
  34. package/src/docker-run.mjs +146 -15
  35. package/src/doctor.mjs +4701 -414
  36. package/src/egress-conf-copy.mjs +166 -0
  37. package/src/egress-proxy-state.mjs +151 -0
  38. package/src/egress.mjs +455 -25
  39. package/src/entry.mjs +27 -0
  40. package/src/env-allowlist.mjs +222 -40
  41. package/src/env-file.mjs +1869 -33
  42. package/src/exit-code.mjs +15 -0
  43. package/src/flow-gate.mjs +5 -3
  44. package/src/forgejo-host.mjs +19 -0
  45. package/src/forgejo-identity.mjs +21 -2
  46. package/src/get-token.mjs +67 -18
  47. package/src/git-dirty.mjs +9 -1
  48. package/src/git-hardening.mjs +33 -0
  49. package/src/github-app-setup.mjs +29 -12
  50. package/src/github-prompt.mjs +4 -1
  51. package/src/gitlab-host.mjs +19 -0
  52. package/src/gitlab-identity.mjs +19 -2
  53. package/src/host-registry.mjs +32 -5
  54. package/src/identity.mjs +29 -4
  55. package/src/image-preflight.mjs +46 -11
  56. package/src/image-ref.mjs +21 -0
  57. package/src/index.mjs +387 -17
  58. package/src/init.mjs +197 -38
  59. package/src/job-user.mjs +252 -0
  60. package/src/json-duplicates.mjs +204 -0
  61. package/src/live-probes.mjs +1020 -0
  62. package/src/materialize.mjs +4 -11
  63. package/src/netns-keeper.mjs +264 -0
  64. package/src/on-failure.mjs +119 -0
  65. package/src/outbox.mjs +7 -0
  66. package/src/podman-stack.mjs +1304 -0
  67. package/src/prepare-github.mjs +6 -6
  68. package/src/prepare-local.mjs +51 -17
  69. package/src/prepare.mjs +27 -6
  70. package/src/processor.mjs +505 -26
  71. package/src/provider-key.mjs +41 -0
  72. package/src/provider-steering.mjs +144 -0
  73. package/src/queue.mjs +35 -8
  74. package/src/redact.mjs +84 -0
  75. package/src/reserved-env.mjs +7 -3
  76. package/src/retention-sweep.mjs +178 -0
  77. package/src/run-container.mjs +181 -14
  78. package/src/run-history.mjs +105 -16
  79. package/src/runtime-observations.mjs +1152 -0
  80. package/src/runtime-settings.mjs +13 -8
  81. package/src/sandbox-cli.mjs +100 -95
  82. package/src/sandbox-store.mjs +612 -45
  83. package/src/sandbox.mjs +1459 -37
  84. package/src/schedules.mjs +16 -3
  85. package/src/secret-profiles.mjs +2 -1
  86. package/src/secrets.mjs +23 -6
  87. package/src/service-env.mjs +247 -0
  88. package/src/service.mjs +618 -28
  89. package/src/session-store.mjs +678 -53
  90. package/src/start.mjs +1395 -268
  91. package/src/transient.mjs +240 -0
  92. package/src/triggers-file.mjs +71 -15
  93. package/src/triggers.mjs +176 -19
  94. package/src/up.mjs +1399 -85
  95. package/src/valkey-auth.mjs +529 -0
  96. package/src/valkey-endpoint.mjs +367 -0
  97. package/src/watch-closer.mjs +158 -0
package/README.md ADDED
@@ -0,0 +1,50 @@
1
+ # @edgehero/pi-dispatch
2
+
3
+ The worker and the `pi-dispatch` command line tool of
4
+ **[pi-dispatch](https://github.com/edgehero/pi-dispatch)**, which runs the
5
+ [pi](https://github.com/earendil-works/pi) coding agent as a self hosted background service.
6
+
7
+ The worker takes jobs from a durable queue (Valkey with BullMQ). It checks the spend caps before
8
+ anything is spent, starts one locked down container per job (Docker or Podman), gives the agent the
9
+ job, records what it did and what it cost, and removes the container. Jobs come from the command line,
10
+ from cron triggers, or from forge events that the
11
+ [receiver](https://www.npmjs.com/package/@edgehero/pi-dispatch-receiver) queues.
12
+
13
+ ## Start
14
+
15
+ You need Docker or Podman, Node 22.19 or newer, and an API key for a model provider that pi supports.
16
+
17
+ ```bash
18
+ mkdir my-dispatch && cd my-dispatch
19
+ npx @edgehero/pi-dispatch up # job image, Valkey with a password, egress proxy, config files, doctor
20
+ # edit .env and set your provider key
21
+ npx @edgehero/pi-dispatch worker # run jobs
22
+ npx @edgehero/pi-dispatch run ./my-project --task "add type hints" --flow tidy
23
+ ```
24
+
25
+ A local job edits your folder **in place**, and there is no undo, so commit first. The worker refuses a
26
+ folder with uncommitted changes unless you pass `--force`.
27
+
28
+ ## Commands
29
+
30
+ | Command | Does |
31
+ |---|---|
32
+ | `init` | write `.env` and the config files into this folder (never overwrites) |
33
+ | `up` | one pass that asks before each container action: image, Valkey, egress proxy, `init`, `doctor` |
34
+ | `doctor [--fix] [--live]` | check the whole setup; `--live` reads the isolation guarantees back from real containers |
35
+ | `worker` | run jobs from the queue |
36
+ | `run <folder> --task "..."` | queue one job against a local folder |
37
+ | `status`, `pause`, `resume`, `cancel <jobId>` | steer the running worker from any terminal |
38
+ | `service render\|install\|status\|restart` | run the worker (or `--receiver`) as a service (user level on macOS and Linux, nssm on Windows) |
39
+ | `setup github` | create a GitHub App in one browser click |
40
+ | `import-pi` | stage your own pi setup (models, skills, persona) for every job, without credentials |
41
+ | `sandbox <jobId>` | reopen a finished run's workspace in a fresh container with no credentials |
42
+
43
+ Before you rely on it, read [`SECURITY.md`](https://github.com/edgehero/pi-dispatch/blob/main/SECURITY.md).
44
+ In short: the agent can read the provider key and the forge token inside its container, and the
45
+ guardrails that tell it not to leak them are prompt text. The default GitHub source (`gh`) forwards your
46
+ whole login, full scope and never expiring. pi-dispatch never merges anything, but the job's token can,
47
+ so branch protection on your default branch is the real control.
48
+
49
+ The full guide, with triggers, providers, the admin panel and every forge, is the
50
+ [main README](https://github.com/edgehero/pi-dispatch#readme). MIT licensed.
@@ -37,9 +37,15 @@
37
37
  One worker per host (DES-CONCURRENCY-3): parallelism is PI_CONCURRENCY inside the one process.
38
38
  Requires the AOF-enabled Valkey from deploy/docker-compose.yml.
39
39
 
40
- PI_LOGS_DIR (run-history records; default OS-temp /pi-dispatch/logs) is created and written by the
41
- worker at boot, so it must be writable by the account the daemon runs as. Set via `.env` (the wrapper),
42
- not a plist change; its default is distinct from the StandardOutPath worker.out.log below.
40
+ PI_LOGS_DIR (run-history records) and PI_SETTINGS_FILE (the settings overlay) default to
41
+ ~/.pi-dispatch, which is DURABLE across a reboot, and are created and written by the worker at boot,
42
+ so they must be writable by the account the daemon runs as. SET BOTH EXPLICITLY via `.env` (the
43
+ wrapper), not a plist change: the default is per USER, so a daemon running as another account writes
44
+ a different directory than the panel reads, and the run list then shows nothing while panel-set caps
45
+ land where the worker never looks. `pi-dispatch up` writes both into `.env` for you, as the account
46
+ default resolved by whoever ran it; if this daemon runs as another account, check it can write there.
47
+ Keep PI_LOGS_DIR
48
+ distinct from the StandardOutPath worker.out.log below: the retention sweep would delete it.
43
49
  -->
44
50
  <plist version="1.0">
45
51
  <dict>
@@ -8,23 +8,52 @@
8
8
  # stays first-class. The `receiver` profile keeps the container OPT-IN: a plain `up` stays Valkey-only,
9
9
  # so worker-on-host deployments that already run the receiver on the host gain no second copy.
10
10
  #
11
- # docker compose -f deploy/docker-compose.yml up -d # start Valkey (default)
12
- # docker compose -f deploy/docker-compose.yml --profile receiver up -d # Valkey + receiver container
13
- # pi-dispatch worker # drain the queue (host)
11
+ # docker compose --env-file .env -f deploy/docker-compose.yml up -d # start Valkey (default)
12
+ # docker compose --env-file .env -f deploy/docker-compose.yml --profile receiver up -d # Valkey + receiver container
13
+ # pi-dispatch worker # drain the queue (host)
14
+ #
15
+ # `--env-file .env` (run from the deployment folder) is how compose reads VALKEY_PASSWORD for the Valkey below (issue
16
+ # #468). Without it compose looks for a .env beside THIS file, finds none, and starts Valkey with no password, which
17
+ # any local account can then read and feed; doctor warns about exactly that.
18
+ #
19
+ # On rootful Podman, run the same commands with the real docker CLI pointed at Podman's socket through a docker
20
+ # context (docs/podman.md). Measured with the egress profile on Podman 5.8.2; podman-docker is not the supported route.
21
+ # Measured again on a real host (issue #355): Podman 5.8.1 on Fedora 44 with SELinux enforcing, netavark on nftables
22
+ # and health checks run by systemd timers. The proxy reached healthy with the `z` labels below and no manual help.
23
+ #
24
+ # Why `z` on the single config files below: with SELinux enforcing, a bind mount carries its host label into the
25
+ # container, and a file under a checkout is not container_file_t, so the container may not read it. Measured on
26
+ # Fedora 44, squid crash-looped with "Unable to open configuration file: /etc/squid/squid.conf: (13) Permission
27
+ # denied". `z` (shared, lower case) relabels that one file container_file_t so any container may read it, and is a
28
+ # no-op where SELinux is off. Never `Z`: that label is private to one container, so it would lock the other
29
+ # services out of the same file. The host's own tools, the worker included, run unconfined and still read it.
14
30
 
15
31
  services:
16
32
  valkey:
17
33
  image: valkey/valkey:8
18
- # AOF on: the wait-list must survive a reboot (REQ-QUEUE-BURST-NO-DROP).
19
- command: valkey-server --appendonly yes
20
- # Bound to localhost only -- the queue is not a public surface.
34
+ # The password (issue #468), from the deployment's .env through `--env-file .env`, into the container's ENVIRONMENT
35
+ # and never onto a command line: the image's PID 1 is tini, which keeps its whole argv for the life of the container,
36
+ # readable by every account on the host in /proc (measured). Empty or unset starts Valkey without one, as before.
37
+ environment:
38
+ VALKEY_PASSWORD: ${VALKEY_PASSWORD:-}
39
+ # AOF on: the wait-list must survive a reboot (REQ-QUEUE-BURST-NO-DROP). The password reaches valkey-server as
40
+ # configuration on stdin (`valkey-server -`), written by the shell's own echo builtin to a 0600 temp file that is
41
+ # opened and deleted before the image's entrypoint runs. `$$` is compose's escape for a literal `$`: the shell in the
42
+ # container expands these, never compose. The same script as the Quadlet unit and `pi-dispatch up`'s docker run.
43
+ command: ["sh", "-c", 'set -eu; if [ -n "$${VALKEY_PASSWORD:-}" ]; then umask 077; f=$$(mktemp); echo "requirepass $$VALKEY_PASSWORD" >"$$f"; exec 0<"$$f"; rm -f "$$f"; set -- -; else set --; fi; unset VALKEY_PASSWORD; exec docker-entrypoint.sh valkey-server "$$@" --appendonly yes']
44
+ # Bound to localhost only -- the queue is not a public surface. On VALKEY_URL's port: PI_VALKEY_PORT, which
45
+ # `pi-dispatch up` and the setup wizard set from VALKEY_URL when they run this file, and which a deployment on a port
46
+ # other than 6379 puts in its .env for a compose command run by hand (PR #475's review: the port was 6379 whatever
47
+ # VALKEY_URL said, so the queue could move away from the worker).
21
48
  ports:
22
- - "127.0.0.1:6379:6379"
49
+ - "127.0.0.1:${PI_VALKEY_PORT:-6379}:6379"
23
50
  volumes:
24
51
  - valkey-data:/data
25
52
  restart: unless-stopped
26
53
  healthcheck:
27
- test: ["CMD", "valkey-cli", "ping"]
54
+ # Healthy only when Valkey answers PONG to this deployment's password (valkey-cli reads REDISCLI_AUTH from its
55
+ # environment); a bare `valkey-cli ping` exits 0 on NOAUTH too.
56
+ test: ["CMD-SHELL", 'if [ -n "$${VALKEY_PASSWORD:-}" ]; then REDISCLI_AUTH="$$VALKEY_PASSWORD"; export REDISCLI_AUTH; fi; valkey-cli ping | grep -q PONG']
28
57
  interval: 10s
29
58
  timeout: 3s
30
59
  retries: 5
@@ -39,7 +68,7 @@ services:
39
68
  profiles: ["receiver"]
40
69
  image: ghcr.io/edgehero/pi-dispatch-receiver:latest
41
70
  # The operator's .env at the repo root, same file the host services read (WEBHOOK_SECRET, github
42
- # auth, forge blocks). Compose refuses to start the profile when it is missing -- correctly: a
71
+ # auth, forge blocks, and VALKEY_PASSWORD, which the receiver sends to the Valkey above). Compose refuses to start the profile when it is missing -- correctly: a
43
72
  # receiver without WEBHOOK_SECRET refuses to boot anyway, so fail at the clearer message.
44
73
  env_file: ../.env
45
74
  # `environment` OVERRIDES env_file, deliberately: a host .env says redis://127.0.0.1:6379, which
@@ -55,7 +84,8 @@ services:
55
84
  # invisible in here until the container restarts, and the worker's own pre-spend check is what
56
85
  # keeps a spent one-shot from running again in the meantime. A restart picks up the current file.
57
86
  volumes:
58
- - ../triggers.json:/config/triggers.json:ro
87
+ # `z` for the SELinux reason at the top of this file: unlabelled, the receiver cannot read it on an enforcing host.
88
+ - ../triggers.json:/config/triggers.json:ro,z
59
89
  # Loopback only, like Valkey's port above: the operator's reverse proxy or tunnel (TLS, public
60
90
  # hostname) is what faces the internet -- never this port raw. Assumes the in-container default
61
91
  # RECEIVER_PORT=3000; if your .env overrides it, adjust the right-hand side to match.
@@ -70,11 +100,12 @@ services:
70
100
  # above: a plain `up` stays Valkey-only, and a deployment that has not turned PI_EGRESS on never needs
71
101
  # this service at all.
72
102
  #
73
- # docker compose -f deploy/docker-compose.yml --profile egress up -d
103
+ # docker compose --env-file .env -f deploy/docker-compose.yml --profile egress up -d
74
104
  #
75
105
  # It sits on ONE network here, the upstream one. The networks that matter are created by the WORKER, one
76
106
  # per job, `--internal`, and this container is attached to each for the life of that job and detached
77
- # after (worker/src/egress.mjs). Per-job rather than one shared network because a shared network is a
107
+ # after (worker/src/egress.mjs), or by the next boot's reaper when a worker died before it could detach
108
+ # (issue #357). Per-job rather than one shared network because a shared network is a
78
109
  # shared L2 segment: at DES-CONCURRENCY-3 that is three mutually-untrusting issue authors who can reach
79
110
  # each other. `enable_icc=false` cannot fix that -- ICC governs every container pair on the bridge and
80
111
  # this proxy is a container, so it would block the very path the design depends on.
@@ -95,12 +126,14 @@ services:
95
126
  # name and the two literals would disagree -- silently, and only on a deployment that renamed nothing.
96
127
  container_name: pi-dispatch-egress-proxy
97
128
  volumes:
98
- # The RULES, shipped and not edited. Relative paths resolve against THIS FILE's directory.
99
- - ./egress-proxy.conf:/etc/squid/squid.conf:ro
129
+ # The RULES, shipped and not edited. Relative paths resolve against THIS FILE's directory. Both mounts
130
+ # here carry `z` for the SELinux reason at the top of this file: without it squid crash-loops on an
131
+ # enforcing host, unable to open its own configuration.
132
+ - ./egress-proxy.conf:/etc/squid/squid.conf:ro,z
100
133
  # The LIST, yours. `pi-dispatch init` scaffolds it next to your .env; it is create-only, so a re-run
101
134
  # never clobbers an edited allowlist. Mounting a path that does not exist makes Docker create it as a
102
135
  # DIRECTORY and squid then fails confusingly, which is why init writes it first.
103
- - ../egress-allowlist.conf:/etc/pi-dispatch/allowlist.conf:ro
136
+ - ../egress-allowlist.conf:/etc/pi-dispatch/allowlist.conf:ro,z
104
137
  networks:
105
138
  - egress-out
106
139
  restart: unless-stopped
@@ -121,7 +154,7 @@ services:
121
154
 
122
155
  networks:
123
156
  # The proxy's route out. Only the proxy is ever on it: job containers live on their own per-job
124
- # `--internal` networks, which have no route anywhere except to this container.
157
+ # `--internal` networks, whose only other member is this container and which have no route off this host.
125
158
  egress-out:
126
159
  name: pi-dispatch-egress-out
127
160
 
@@ -7,7 +7,8 @@
7
7
  #
8
8
  # What this policy is, and is not:
9
9
  # - Hostname filtering on CONNECT, to port 443 only. The provider, the forge and the registry are
10
- # ordinary entries in the allowlist; there is no address-based rule anywhere and nothing is special.
10
+ # ordinary entries in the allowlist, and nothing is special. The one address-based rule is a DENY:
11
+ # a listed name that resolves to a fixed loopback or link-local address of this host is refused.
11
12
  # - TLS is NEVER terminated. This proxy sees the name a client asks for and no byte inside the tunnel,
12
13
  # so it cannot read a credential and cannot account for a token. A proxy that decrypts provider
13
14
  # traffic is OQ-011's mechanism, a materially larger change, and it is not this.
@@ -20,12 +21,41 @@
20
21
  http_port 3128
21
22
 
22
23
  # The operator's list. Bare hostnames, one per line; a leading dot matches subdomains (.github.com).
23
- acl allowed dstdomain "/etc/pi-dispatch/allowlist.conf"
24
+ # `-n` (issue #428): without it squid answers an IP-literal request (CONNECT 93.184.215.14:443, or an
25
+ # IPv6 literal) that matches no entry by a REVERSE lookup of that address, and matches the name it
26
+ # gets back (measured: a PTR query per literal, then 403). A job choosing the address chooses whose
27
+ # reverse zone is asked, which is a DNS channel out. With `-n` no lookup is made: the request's host
28
+ # is compared as written, so a literal is still refused unless the list names that literal itself.
29
+ acl allowed dstdomain -n "/etc/pi-dispatch/allowlist.conf"
24
30
 
25
31
  acl SSL_ports port 443
26
32
  acl CONNECT method CONNECT
27
33
 
34
+ # The host's own services, by address (issue #428). On rootless Podman an account's containers.conf
35
+ # (pasta_options such as --map-host-loopback 169.254.1.2, or slirp4netns's allow_host_loopback=true,
36
+ # which answers on 10.0.2.2) gives the network this proxy sits on the host's 127.0.0.1 services, the
37
+ # job queue's Valkey among them, and a name in the allowlist that resolves (or is rebound) to one of
38
+ # these addresses would then reach them on a job's behalf. 0.0.0.0 and :: are loopback too on Linux.
39
+ # 10.0.2.2 is slirp4netns's host alias and only means the host there; elsewhere it is an ordinary
40
+ # private address, and an allowlisted service that really lives at it is refused (move it, or drop
41
+ # this entry in a copy you maintain).
42
+ # DEFENCE IN DEPTH, not the closure, and only for the FIXED addresses: the host can be mapped to any
43
+ # address (--map-host-loopback <address>, slirp4netns's cidr=, and --map-gw, which uses the default
44
+ # gateway's), and no fixed rule names those. The closure is the worker's: its podman venue refuses
45
+ # every job while the account's containers.conf sets pasta_options, network_cmd_options, annotations
46
+ # or env. That refusal reads the conf, not the live network: after a key is removed, a proxy started
47
+ # under the old conf keeps the old options until the account's containers on a bridge network restart,
48
+ # which the refusal's own text tells the operator to do.
49
+ acl to_host_local dst 127.0.0.0/8 0.0.0.0/32 169.254.0.0/16 10.0.2.2/32 ::1 ::/128 fe80::/10
50
+
28
51
  # Order is the design, not a detail. Read top to bottom, first match wins.
52
+ # The deny names `allowed` FIRST, and that order is load-bearing: squid evaluates an http_access line's
53
+ # ACLs left to right and stops at the first that does not match, and `dst` resolves the requested name.
54
+ # A bare `deny to_host_local` resolved EVERY name a job asked for, listed or not (measured with
55
+ # debug_options 78,3: a CONNECT to an unlisted name sent a DNS query, where without the rule it sent
56
+ # none), which is a DNS channel out of an egress-armed job. With `allowed` first, only a listed name
57
+ # is ever resolved here, and it is the only kind the allows below could let through anyway.
58
+ http_access deny allowed to_host_local
29
59
  http_access deny CONNECT !SSL_ports
30
60
  http_access allow CONNECT allowed
31
61
  http_access allow allowed
@@ -20,13 +20,19 @@ REM multiple services. Requires the AOF-enabled Valkey from deploy/docker-compos
20
20
  REM
21
21
  REM Per-host PLACEHOLDERS: set SERVICE / REPO / LOGDIR below for your host before running.
22
22
  REM
23
- REM PI_LOGS_DIR (run-history records; default OS-temp \pi-dispatch\logs) is created and written by the
24
- REM worker at boot, so it must be writable by the service account. Set via `.env` (the wrapper), not a
25
- REM change here; its default avoids colliding with the nssm LOGDIR worker.out log set below.
23
+ REM PI_LOGS_DIR (run-history records) and PI_SETTINGS_FILE (the settings overlay) default to
24
+ REM %USERPROFILE%\.pi-dispatch, which is DURABLE across a reboot. Both are created and written by the
25
+ REM worker at boot, so they must be writable by the service account.
26
26
  REM
27
- REM PI_SETTINGS_FILE is the runtime-tunable settings overlay (default under OS temp, which may be wiped
28
- REM on reboot) -- point it at a durable path in production. Set via `.env` (the wrapper), not a change
29
- REM here; it is worker-owned and never belongs in the container env allowlist.
27
+ REM SET BOTH EXPLICITLY via `.env` (the wrapper), not a change here. The default is per USER, and a
28
+ REM service running as LocalSystem resolves it under the system profile, not yours -- so the worker and
29
+ REM your /dispatch panel would read different directories, the run list would show nothing, and caps set
30
+ REM from the panel would land where the worker never looks. `pi-dispatch up` writes both into `.env`,
31
+ REM as the account default resolved by whoever ran it; a service running as LocalSystem is very likely
32
+ REM NOT that account, so check it can write there.
33
+ REM Keep PI_LOGS_DIR away from the nssm LOGDIR worker.out log set below: the retention sweep deletes
34
+ REM every .log and .json past its window. Both are worker-owned and never belong in the container env
35
+ REM allowlist.
30
36
 
31
37
  setlocal
32
38
 
@@ -0,0 +1,10 @@
1
+ # Quadlet network for the egress proxy's route out on the native rootless `podman` venue (issue #430,
2
+ # docs/podman.md): the compose file's `egress-out` network, named exactly as there.
3
+ #
4
+ # Installed with pi-dispatch-egress-proxy.container, only while the egress policy is armed. Podman's generator turns
5
+ # it into pi-dispatch-egress-out-network.service, which the proxy's unit requires. NetworkName= gives the network
6
+ # exactly this name (measured on Podman 5.8.1). Only the proxy is ever on it: job containers live on their own
7
+ # per-job `--internal` networks, which the WORKER creates and attaches the proxy to by name.
8
+
9
+ [Network]
10
+ NetworkName=pi-dispatch-egress-out
@@ -0,0 +1,50 @@
1
+ # Quadlet unit for the egress policy's allowlist proxy on the native rootless `podman` venue (REQ-EGRESS-ALLOWLIST,
2
+ # issue #430, docs/podman.md): deploy/docker-compose.yml's `egress-proxy` service, for a host whose jobs run under
3
+ # this account's own rootless Podman.
4
+ #
5
+ # Installed by `pi-dispatch service install` (user scope, with podman in PI_BACKENDS and PI_EGRESS armed) and by
6
+ # `pi-dispatch up`. Unlike the Valkey unit this one is RENDERED: the two /opt/pi-dispatch paths below are
7
+ # placeholders, replaced with an account-owned copy of the package's egress-proxy.conf (written beside this account's
8
+ # config, because `z` cannot relabel a root-owned package file: measured, exit 126) and with the deployment folder's
9
+ # egress-allowlist.conf. Never `systemctl --user enable` it (a generated unit refuses that); the [Install] section is
10
+ # read by the generator itself.
11
+ #
12
+ # A RESTART DROPS EVERY PER-JOB NETWORK (measured on Podman 5.8.1). The generated unit runs `podman run --replace
13
+ # --rm`, so `systemctl --user restart` makes a NEW container, and the per-job `--internal` networks the worker had
14
+ # connected to the old one are gone from it. A job started afterwards connects the new proxy at its start and is
15
+ # fine; a job already running when the proxy restarts has lost its only route out for the rest of that run.
16
+ #
17
+ # A STOP TAKES 10 s: squid does not exit on SIGTERM, so systemd waits out podman's stop timeout, kills it, and the
18
+ # unit ends `failed` with exit 137 (measured). Harmless; `pi-dispatch service uninstall` clears the failed state.
19
+
20
+ [Unit]
21
+ Description=pi-dispatch egress allowlist proxy (rootless Podman, Quadlet)
22
+
23
+ [Container]
24
+ # The compose file's digest, fully qualified. A test holds the two equal: this container IS the allowlist, so a
25
+ # digest bumped in one file only would give the two venues two different policies.
26
+ Image=docker.io/ubuntu/squid@sha256:6a097f68bae708cedbabd6188d68c7e2e7a38cedd05a176e1cc0ba29e3bbe029
27
+ # The name the worker attaches to each job network. A deployment that sets PI_EGRESS_PROXY to another name runs its
28
+ # own proxy, and neither `service install` nor `up` installs this unit for it: `--replace` would otherwise remove
29
+ # whatever container of that name the operator runs.
30
+ ContainerName=pi-dispatch-egress-proxy
31
+ Network=pi-dispatch-egress-out.network
32
+ # The RULES (shipped) and the LIST (yours). `z` for the SELinux reason the compose file records: unlabelled, squid
33
+ # cannot read its own configuration on an enforcing host and crash-loops. Shared `z`, never private `Z`, as there:
34
+ # on a host that also runs the compose proxy the same allowlist file is mounted by both.
35
+ Volume=/opt/pi-dispatch/deploy/egress-proxy.conf:/etc/squid/squid.conf:ro,z
36
+ Volume=/opt/pi-dispatch/egress-allowlist.conf:/etc/pi-dispatch/allowlist.conf:ro,z
37
+ # The exec form, a JSON array, for the reason the compose file gives: a plain string runs under /bin/sh, which on
38
+ # this image is dash, where /dev/tcp does not exist and the check would fail forever while squid serves.
39
+ HealthCmd=["bash", "-c", "exec 3<>/dev/tcp/127.0.0.1/3128"]
40
+ HealthInterval=30s
41
+ HealthTimeout=3s
42
+ HealthRetries=3
43
+ HealthStartPeriod=10s
44
+
45
+ [Service]
46
+ TimeoutStartSec=900
47
+ Restart=on-failure
48
+
49
+ [Install]
50
+ WantedBy=default.target
@@ -0,0 +1,80 @@
1
+ # Quadlet unit for the rootless network keeper on the native rootless `podman` venue (issue #458, docs/podman.md):
2
+ # one idle container that holds this account's shared rootless network helper open, so the egress proxy keeps its
3
+ # route out after a job's network is torn down.
4
+ #
5
+ # WHY IT EXISTS. On Podman 4.9 (Ubuntu 24.04's), `podman network disconnect` of the proxy from a per-job network
6
+ # tears down the account's shared rootless network namespace and its slirp4netns WHILE THE PROXY STILL RUNS, whenever
7
+ # no other running container is on a bridge network: v4.9.3's libpod/networking_linux.go counts the disconnecting
8
+ # container as the caller and cleans up at one. The next bridge container then starts a new namespace, the proxy's
9
+ # eth0 goes with the old one, and every later egress job gets 503 from the proxy until it restarts (measured: job 1
10
+ # 200, jobs 2 to 5 503). Podman 5.0's rootless network rewrite (containers/podman#20772, containers/common#1761)
11
+ # counts attachments instead, and 5.8.1 was measured unaffected. With this container running on a bridge network of
12
+ # its own, 4.9.3 gave 5 sequential egress jobs 200 each, and the two-network teardown `doctor --live` does stayed
13
+ # clean. It is installed on every Podman version, because on 5.x it is one idle container and nothing else
14
+ # (measured harmless on 5.8.1), and one rule is simpler than a version read at install time.
15
+ #
16
+ # Installed by `pi-dispatch service install` (user scope, with podman in PI_BACKENDS and PI_EGRESS armed) and by
17
+ # `pi-dispatch up`, which copy it unchanged into ~/.config/containers/systemd/. Never `systemctl --user enable` it
18
+ # (a generated unit refuses that); the [Install] section is read by the generator itself.
19
+ #
20
+ # IT MUST STAY RUNNING: it protects only while it runs. Stopped, the next job teardown breaks the proxy again (measured).
21
+ #
22
+ # IT WIDENS NOTHING. No published port, no mount, no capability, a read-only root, no new privileges, the nobody
23
+ # uid, and a network with no route out and no DNS: inside it, the proxy and 1.1.1.1 were both unreachable (measured).
24
+
25
+ [Unit]
26
+ Description=pi-dispatch rootless network keeper (holds this account's rootless netns open; rootless Podman, Quadlet)
27
+ # No start limit: with Restart=always a keeper that cannot start (its network pruned, say) would otherwise use up
28
+ # systemd's default 5 starts in 10 s and stay `failed` for good. Measured on 4.9.3: the generator passes this through,
29
+ # 10 quick kills in a row and 15 s of a keeper that could not start never ended `failed`.
30
+ StartLimitIntervalSec=0
31
+
32
+ [Container]
33
+ # The proxy's own image and digest, so this unit pulls nothing the egress stack has not already pulled and adds no
34
+ # pin to keep in step: it only runs `sleep`, which the image has. A test holds the digest equal to the proxy's.
35
+ Image=docker.io/ubuntu/squid@sha256:6a097f68bae708cedbabd6188d68c7e2e7a38cedd05a176e1cc0ba29e3bbe029
36
+ # Outside every prefix a pi-dispatch sweep removes (pi-job-, pi-sandbox-, pi-dispatch-live-, pi-dispatch-egress-doctor-,
37
+ # pi-dispatch-egress-probe-): a test holds that, since a sweep that removed it would bring the defect back unseen.
38
+ ContainerName=pi-dispatch-netns-keeper
39
+ # Its OWN network, never the proxy's or a job's: it has to be a running member of SOME bridge network, and on one of
40
+ # its own it is reachable from nothing.
41
+ Network=pi-dispatch-netns-keeper.network
42
+ ReadOnly=true
43
+ DropCapability=all
44
+ NoNewPrivileges=true
45
+ User=65534
46
+ Group=65534
47
+ # catatonit as PID 1 (`--init`), so a stop ends `sleep` at once rather than after podman's 10 s timeout (measured on
48
+ # 4.9.3: 123 ms); a job's argv carries the same flag, so the venue already needs catatonit.
49
+ RunInit=true
50
+ # Two spellings Quadlet 4.9 has no key for. `--entrypoint=sleep` with `Exec=infinity`: 4.9.3's generator refuses an
51
+ # `Entrypoint=` key ("unsupported key", measured), and a JSON array in PodmanArgs (`--entrypoint=["sleep","infinity"]`)
52
+ # came through as `--entrypoint=[sleep,infinity]`, which catatonit could not exec (measured). `--image-volume=ignore`:
53
+ # the squid image declares VOLUMEs, which would otherwise make two anonymous volumes on every start.
54
+ PodmanArgs=--entrypoint=sleep --image-volume=ignore
55
+ Exec=infinity
56
+
57
+ [Service]
58
+ TimeoutStartSec=900
59
+ # Its network, made again before every start, with exactly the flags its .network unit's own `podman network create`
60
+ # carries (--ignore makes it a no-op while the network exists). That unit runs once and stays `active`, so without
61
+ # this a network pruned (or removed) under a stopped keeper kept every later start failing until someone restarted
62
+ # both units by hand; measured on 4.9.3, with this line a start after `podman network prune` recreated the network
63
+ # (internal, no DNS) and the keeper came up.
64
+ ExecStartPre=/usr/bin/podman network create --ignore --disable-dns --internal pi-dispatch-netns-keeper
65
+ # always, not the proxy's on-failure: a keeper that is not running is the whole defect, and `sleep infinity` never
66
+ # exits on its own, so any exit at all is one to undo. An explicit `systemctl --user stop` still stops it; a `podman
67
+ # stop` or `podman rm -f` behind systemd's back does not (measured: it is running again after RestartSec).
68
+ Restart=always
69
+ # One second, not systemd's 100 ms default, so a keeper that cannot start retries at a calm pace (the start limit is
70
+ # off above). Measured on 4.9.3: back on its bridge network 1.2 to 1.3 s after a kill; 250 ms gave 0.45 s, and was not
71
+ # taken, because a keeper that came back more than the grace (15 s) after the proxy started is already treated as
72
+ # damage by the worker and doctor (a proxy restart is asked for, whatever the gap), so a shorter gap buys nothing there, and a
73
+ # keeper that cannot start would retry four times a second.
74
+ RestartSec=1s
75
+ # A stop ends `sleep` with SIGTERM, which the unit sees as exit 143 and, without this line, as a failure (measured on
76
+ # Podman 4.9.3: every clean `systemctl --user stop` left the unit `failed`).
77
+ SuccessExitStatus=143
78
+
79
+ [Install]
80
+ WantedBy=default.target
@@ -0,0 +1,18 @@
1
+ # Quadlet network for the rootless network keeper on the native rootless `podman` venue (issue #458,
2
+ # docs/podman.md): a network of the keeper's OWN, with nothing else ever on it.
3
+ #
4
+ # Installed with pi-dispatch-netns-keeper.container, while the egress policy is armed. Podman's generator turns it
5
+ # into pi-dispatch-netns-keeper-network.service, which the keeper's unit requires. Internal= gives it no route out
6
+ # and DisableDNS= no aardvark-dns entry, so the one container on it can reach nothing at all, not even the proxy:
7
+ # jobs live on their own per-job `--internal` networks, which have no route to this subnet.
8
+ #
9
+ # NEVER REMOVE IT. Measured on Podman 4.9.3: `podman network rm -f` of this network kills the keeper, and a `podman
10
+ # network prune` or `podman system prune` run while the keeper is stopped removes it (with the keeper running, both
11
+ # leave it). Either way the keeper's unit cannot start again while this network's unit stays `active` (it ran once and
12
+ # exited), until `systemctl --user reset-failed` and a restart of both units. Its name is outside every prefix a
13
+ # pi-dispatch sweep removes.
14
+
15
+ [Network]
16
+ NetworkName=pi-dispatch-netns-keeper
17
+ Internal=true
18
+ DisableDNS=true
@@ -0,0 +1,51 @@
1
+ # Quadlet unit for Valkey on the native rootless `podman` venue (issue #430, docs/podman.md): the same component
2
+ # deploy/docker-compose.yml's `valkey` service and `pi-dispatch up`'s docker path start, for a host that runs jobs
3
+ # under this account's own rootless Podman instead of a docker daemon.
4
+ #
5
+ # Installed by `pi-dispatch service install` (user scope, with podman in PI_BACKENDS) and by `pi-dispatch up`, which
6
+ # copy it into ~/.config/containers/systemd/ unchanged but for a VALKEY_URL port other than 6379, with the password
7
+ # file it reads beside it. Podman's systemd generator turns it into
8
+ # pi-dispatch-valkey.service. Never `systemctl --user enable` it: a generated unit refuses that ("transient or
9
+ # generated"); the [Install] section below is read by the generator itself, which links it into default.target
10
+ # (measured on Podman 5.8.1: with linger on it came up 25 s after a reboot with nobody logged in).
11
+ #
12
+ # Fully qualified image names throughout: under short-name enforcement a unit cannot answer the prompt a short
13
+ # name raises, so `valkey/valkey:8` could fail to start where Docker quietly assumed docker.io.
14
+
15
+ [Unit]
16
+ Description=pi-dispatch Valkey (the job queue; rootless Podman, Quadlet)
17
+
18
+ [Container]
19
+ Image=docker.io/valkey/valkey:8
20
+ ContainerName=pi-dispatch-valkey
21
+ # EXPLICIT, see pi-dispatch-valkey.network for why a container here must never rely on the default network.
22
+ Network=pi-dispatch-valkey.network
23
+ # Bound to localhost only, as in the compose file: the queue is not a public surface. The worker reaches it at
24
+ # redis://127.0.0.1:6379, the default VALKEY_URL.
25
+ PublishPort=127.0.0.1:6379:6379
26
+ # The same named volume `pi-dispatch up` creates on docker, so the wait-list survives a restart and a reboot.
27
+ Volume=pi-dispatch-valkey-data:/data
28
+ # The password (issue #468): VALKEY_PASSWORD from a 0600 file `service install` and `up` write from the deployment's
29
+ # .env, into the container's ENVIRONMENT. Never `--requirepass` on a command line: the image's PID 1 is tini, which keeps
30
+ # its whole argv for the life of the container, and any account on the host reads it in /proc/<pid>/cmdline (measured
31
+ # on Podman 5.8.1 and 4.9.3). Quadlet also loads this file into the unit's own environment, where systemd would expand a
32
+ # bare $VALKEY_PASSWORD in the command below into the password itself: every $ below is written $$ for that reason.
33
+ EnvironmentFile=%h/.config/pi-dispatch/valkey.env
34
+ # AOF on: the wait-list must survive a reboot (REQ-QUEUE-BURST-NO-DROP). The password reaches valkey-server as
35
+ # configuration on stdin (`valkey-server -`), written by the shell's own echo builtin to a 0600 temp file that is opened
36
+ # and deleted before the image's entrypoint runs; with VALKEY_PASSWORD empty it starts with none, as before.
37
+ Exec=sh -c 'set -eu; if [ -n "$${VALKEY_PASSWORD:-}" ]; then umask 077; f=$$(mktemp); echo "requirepass $$VALKEY_PASSWORD" >"$$f"; exec 0<"$$f"; rm -f "$$f"; set -- -; else set --; fi; unset VALKEY_PASSWORD; exec docker-entrypoint.sh valkey-server "$$@" --appendonly yes'
38
+ # Healthy only when Valkey answers PONG to this deployment's password (valkey-cli reads REDISCLI_AUTH from its
39
+ # environment); a bare `valkey-cli ping` exits 0 on NOAUTH too.
40
+ HealthCmd=if [ -n "$${VALKEY_PASSWORD:-}" ]; then REDISCLI_AUTH="$$VALKEY_PASSWORD"; export REDISCLI_AUTH; fi; valkey-cli ping | grep -q PONG
41
+ HealthInterval=10s
42
+ HealthTimeout=3s
43
+ HealthRetries=5
44
+
45
+ [Service]
46
+ # The first start pulls the image, which can outlast systemd's default 90 s start timeout on a slow link.
47
+ TimeoutStartSec=900
48
+ Restart=on-failure
49
+
50
+ [Install]
51
+ WantedBy=default.target
@@ -0,0 +1,16 @@
1
+ # Quadlet network for Valkey on the native rootless `podman` venue (issue #430, docs/podman.md).
2
+ #
3
+ # `pi-dispatch service install` (user scope, with podman in PI_BACKENDS) and `pi-dispatch up` copy this file
4
+ # unchanged into ~/.config/containers/systemd/, where Podman's systemd generator turns it into
5
+ # pi-dispatch-valkey-network.service. NetworkName= gives the network exactly this name, with no `systemd-` prefix
6
+ # (measured on Podman 5.8.1).
7
+ #
8
+ # Why Valkey gets a bridge network of its own rather than no Network= at all: a containers.conf that sets
9
+ # `netns = "host"` puts a container started WITHOUT a network option into the host's network namespace (measured),
10
+ # where the port mapping below means nothing and Valkey would listen on every host interface. An explicit Network=
11
+ # is the only thing that wins over that key. `Network=pasta` would also win, but Podman 4.9 (Ubuntu 24.04) defaults
12
+ # to slirp4netns and ships without passt on some hosts, so a named bridge is the one spelling both 4.9 and 5.x run.
13
+ # Its own network, not the proxy's: nothing on the egress side has any business reaching the queue.
14
+
15
+ [Network]
16
+ NetworkName=pi-dispatch-valkey
@@ -23,6 +23,12 @@ Wants=network-online.target
23
23
  # trying. The worker unit has carried this since it shipped; the receiver did not, and the gap is not
24
24
  # cosmetic -- the receiver is the process that dies on a triggers file it cannot parse, and it dies on
25
25
  # EVERY start. Pairs with Restart=on-failure below.
26
+ #
27
+ # This bounds CRASH loops (a process dying in milliseconds), not transient outages (issue #318): a
28
+ # transient identity failure at boot retries INSIDE the process for RECEIVER_IDENTITY_RETRY_SECONDS
29
+ # (default 600, floored at 60 -- see .env.example) before exiting 1, so a retrying boot exits at most
30
+ # once per window and can never place StartLimitBurst starts inside this interval. Keep these values
31
+ # and widen the env key instead: raising the burst here would only re-open the crash-loop hole.
26
32
  StartLimitIntervalSec=60
27
33
  StartLimitBurst=5
28
34
 
@@ -46,8 +46,18 @@ REM loaded, on purpose: the path is SERVICE configuration, and a `.env` line mus
46
46
  REM a script this wrapper then runs.
47
47
  set "ENV_SETUP=%PI_ENV_SETUP%"
48
48
 
49
+ REM NOTHING THIS WRAPPER READS IS LEFT WHERE .env CAN ASSIGN IT (issue #470). Every .env line is a `set`, and
50
+ REM set ignores case, so an `ENV_SETUP=` (or `env_setup=`) line replaced the copy above and named a script
51
+ REM this wrapper then ran, and an `ERRORLEVEL=` line shadows the dynamic %ERRORLEVEL% (`set /?`: a variable
52
+ REM defined under one of those names overrides it), so the exit-code conversion below would read the file's
53
+ REM number. So after the load the wrapper re-asserts its own: ENV_SETUP and PI_ENV_SETUP from
54
+ REM %PI_ENV_SETUP%, which cmd expands when it READS this parenthesized block, before its first line runs
55
+ REM (`set /?`, the reason delayed expansion exists), so the value is the unit's and no .env line reaches it;
56
+ REM and ERRORLEVEL cleared just before the command runs. RC is assigned after the command, before it is read.
49
57
  if exist ".env" (
50
58
  for /f "usebackq eol=# tokens=1,* delims==" %%A in (".env") do set "%%A=%%B"
59
+ set "ENV_SETUP=%PI_ENV_SETUP%"
60
+ set "PI_ENV_SETUP=%PI_ENV_SETUP%"
51
61
  ) else (
52
62
  if not defined ENV_SETUP (
53
63
  echo worker-env-wrapper: .env not found in "%CD%" -- this wrapper must be started in the deployment folder, the service's nssm AppDirectory; it no longer guesses a location from its own path 1>&2
@@ -82,6 +92,7 @@ REM closed.
82
92
  REM
83
93
  REM The argv runs verbatim -- absolute node, absolute script, composed by `pi-dispatch service` (see
84
94
  REM the .sh twin for the whole contract).
95
+ set "ERRORLEVEL="
85
96
  %*
86
97
  set "RC=%ERRORLEVEL%"
87
98