@orchestraworks/worker 0.0.134 → 0.0.136

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (3) hide show
  1. package/README.md +22 -501
  2. package/docs/REFERENCE.md +520 -0
  3. package/package.json +9 -8
package/README.md CHANGED
@@ -1,518 +1,39 @@
1
- # The worker
1
+ # @orchestraworks/worker
2
2
 
3
- One Go program runs inside every machine and speaks the Factory Worker Protocol to the
4
- platform. This tree is that program, and it is also the `@orchestraworks/worker` npm package: the launcher on
5
- the path, and one per-platform package carrying the binary for the machine it is installed on.
3
+ The worker is the one program that runs on each Orchestra machine: it connects the machine to the platform and runs your agent's harness.
6
4
 
7
- The worker exposes no API of its own. Every route it calls is the platform's.
5
+ ## Install
8
6
 
9
- ## The tree
10
-
11
- | path | what it holds |
12
- |---|---|
13
- | `cmd/worker/` | the one binary: `run`, `install`, `enroll`, `check`, `version` |
14
- | `identity/` | the four identity exchanges — claim, refresh, reattach-challenge, reattach — the holder key, and the loop that keeps the machine token alive |
15
- | `broker/` | the protocol on the message broker: the connection, the register on every connect, the heartbeat and the status cadence, the pull of this machine's own work consumer, and the publisher every other package sends through |
16
- | `env/` | the machine-side environment the platform composes onto a launch, read from the names the contract declares |
17
- | `facts/` | what this build knows about the four harness families: the program to install and the digest that authenticates it, the environment each family is handed, the templates it seeds, and every location its content lands in — embedded in the binary, so one release digest covers the facts and the program that reads them |
18
- | `harness/` | the only place family-specific knowledge lives: the launch plan, the content layout, the store lookup, the mid-turn message each family accepts, and the installer that puts the pinned Node and a family's programs under the worker root |
19
- | `process/` | the agent's setup and the harness process behind each session placement: the start, the stop and the last save a remove asks for, the startup fence over what a previous run left, and the stop of everything when the worker is told to stop |
20
- | `stream/` | the queue: every placement message and every prompt, steer and cancel pulled off this machine's work consumer, the ACP client that opens a session's harness conversation and runs its turns, each turn's events batched to the platform and published live, and the answer every placement message is owed |
21
- | `mount/` | the volumes the launch declares, placed on the lanes that mount them: the home volume first, because anything written under a mount point before the mount is hidden by it |
22
- | `network/` | the transparent redirect on the virtual-machine lanes, as the iptables rules the launch describes: what the worker's own uid reaches directly, and what every uid on the machine reaches through the egress edge |
23
- | `shim/` | the in-guest transparent proxy: the listener the redirect lands on, the recovery of the destination each connection was actually for, the hop to the egress edge under the machine-identity proof, the WinDivert half that is the Windows redirect, and the loopback feed a separate shim reads this machine's token from |
24
- | `egress/` | what puts that proxy in place on the virtual-machine lanes: the edge's certificate authority in this machine's trust store, the inbound firewall rule Windows needs, the redirect itself, and the state line an operator reads off the machine's console |
25
- | `ingress/` | the machine's half of the preview tunnel: the connection this worker holds open to the preview gateway, the challenge and the signed upgrade that authenticate it, and the loopback pipe that carries each viewer request to a port this machine declared and refuses every other |
26
- | `termination/` | the cloud's notice that this machine is being reclaimed, delivered to the broker so the next status report carries it |
27
- | `paths/` | which native place a logical path denotes, and when two logical paths denote one place — pinned against the platform's own answer by committed vectors |
28
- | `service/` | `worker run` under this operating system's service manager, in one of two scopes: a system or user `systemd` unit, a Windows service or logon task, a launchd daemon or launch agent |
29
- | `file/` | everything that moves bytes onto this machine and back off it: the install of the skills, memory, instructions and model settings an agent placement's apply carries, against the install journal, the capture that offers the machine's declared roots back to the platform, the restore that lays a previous machine's capture down before a manifest is installed over it, and the self-update that replaces this worker's own release package |
30
- | `script-harness/` | the first-party `script` harness program, which is the one family the worker package carries rather than installs |
31
- | `bin/worker.js` | the npm launcher: it runs the binary out of the per-platform package installed for this machine, stays in front of it, and forwards the signal a service manager ends it with. Its own tests sit beside it and are not packed |
32
- | `npm/<platform>/` | the five per-platform packages, one per release target: the binary, and the `script` harness copied in beside it at build time |
33
-
34
- `../fwp` is the protocol. Every wire shape comes from `github.com/livingcomputers/fwp/go/wire`
35
- through a relative `replace`, so the program and the conformance suite compile against the exact
36
- shapes the schema ships, and nothing here re-declares one.
37
-
38
- ## The commands
39
-
40
- - **`worker run`** claims this machine's identity — or reattaches, where its row is already bound to
41
- the key this machine holds — and then serves the machine. `run` hands the held identity, with the
42
- broker credential the claim co-issued, to the registrar seam in `cmd/worker/run.go`, which the
43
- `broker` package fills: it dials, registers the machine on every connect, pulls the machine's own
44
- work consumer and reports on the cadence.
45
-
46
- Before any of that, `run` starts what the launch describes: it places the declared volumes, puts
47
- the in-guest egress in place where the launch names an intercept set, opens the egress shim's token
48
- feed, and — on a machine the cloud can reclaim — watches for the notice that it is going away. Each
49
- is started only where it is described: a Kubernetes pod places its own volumes and its egress
50
- sidecar installs its own redirect, so a launch there names neither. If one of those duties is
51
- installed and then dies, `run` ends rather than serving a machine the platform counts as healthy
52
- with no route off it; the service manager starts the worker again.
53
-
54
- `--service` is how the process is hosted rather than what it does: on Windows it hands the body to
55
- the service control manager, which kills a service that does not answer its handshake, and on the
56
- other operating systems it is refused by name, because systemd and launchd supervise a foreground
57
- process from outside.
58
-
59
- On a machine that hosts a harness, `run` also builds the three lanes that answer the platform's
60
- verbs: `process`, which holds the placements, `stream`, which carries their turns, and `file`,
61
- which installs what an agent is made of and captures what its work left behind. The three are
62
- built before the broker, because the broker's configuration carries their handlers, and each
63
- reaches it afterwards through the deferred link in `cmd/worker/serve.go`.
64
-
65
- Once the register has been answered, `run` starts the capture cadence and, where the platform named
66
- a version other than the one this process is running, the self-update: the release package is
67
- fetched from the npm registry when this machine can reach one and from the platform when it cannot,
68
- checked against the digest the register answer named, unpacked over the release package directory,
69
- and the process restarted — but only once no turn has been open for five seconds together, because one
70
- open turn per machine is the platform's rule and killing one would lose its answer. A machine whose
71
- turns never leave that gap keeps running the version it has until some later start picks the unpacked
72
- one up.
73
- - **`worker install --role worker|image|laptop --harness <families> --yes [--root <dir>] [--service=user|none]`**
74
- installs the worker on this machine and starts its service. Every answer comes from a flag: a missing
75
- one is an error and never a default. It installs the harness half itself — the pinned Node, then each
76
- named family's programs from the facts file's own pins, with the tarball checked against the pinned
77
- digest before anything is unpacked — and then registers the service. The per-operating-system service
78
- forms fill the installer seam in `cmd/worker/install.go`; a build with none refuses rather than
79
- reporting an install it did not perform.
80
-
81
- **`--service` says which session the worker is registered in, and neither of its two values needs any
82
- privilege.** `user` registers it in the invoking account's own session: a launch agent under
83
- `~/Library/LaunchAgents` on macOS, a `systemd --user` unit under `~/.config/systemd/user` on Linux, a
84
- scheduled task the account's logon starts on Windows. `none` installs the files and registers nothing,
85
- for the one host that has no service manager — a container image, whose own command is the worker —
86
- and `false` is that same answer under its older spelling. The default is `user`. An install run as
87
- root is refused, because it would put the worker's root under root's home and run every harness as
88
- root; there is no privileged install to redirect it to, and `--service=system` is refused with the
89
- same answer. Nothing on a machine somebody owns needs privilege: network enforcement is outside the
90
- machine, at the egress edge and the directional security group.
91
-
92
- **The machine-wide service is the image role's, and the role is the whole of what asks for it.** An
93
- `image` gets it — a property list under `/Library/LaunchDaemons`, a unit under `/etc/systemd/system`,
94
- a Windows service — because the platform builds the image with nobody logged in to it and nobody
95
- typing the command, and the builder is already root. No flag selects that scope, on any role.
96
-
97
- `--root` follows the scope the same way: `~/.sfdk/worker` (`%LOCALAPPDATA%\sfdk\worker` on Windows) for
98
- a `user` install, `/var/lib/sfdk/worker` (`%ProgramData%\orchestra\worker`) for the other two. A `user`
99
- install writes its log under that root, at `log/worker.log`, because no service manager in that scope
100
- keeps one a person can read.
101
-
102
- What the role decides is when the service runs — a `worker` starts now and at every boot, an `image`
103
- being baked starts at the first boot of a machine made from it and claims no identity now, and a
104
- `laptop` starts when a person enrolls a machine and says so. An `image` install also refuses, before
105
- it installs anything, when a bootstrap file already sits at any path the run would read: an image is
106
- copied onto every machine made from it, so a baked bootstrap would hand one machine's identity to all
107
- of them.
108
-
109
- Each form carries one thing its machine would otherwise be wrong without. A macOS launch DAEMON — the
110
- image role's form, which a person never reaches — names the account that ran the install, because
111
- launchd's system domain runs a daemon as root otherwise, and a harness runs as whoever this worker is;
112
- the worker's own root and the daemon's log file are handed to that account in the same step, so the
113
- service can read the owner-only bootstrap the root install just wrote, write its state file beside it,
114
- and say in its log why it did not start. A launch AGENT names no account: it is
115
- already the session's. The Windows service is registered to restart itself three times at two seconds,
116
- on a clean non-zero exit as well as a crash — the control manager restarts nothing it was not told to,
117
- and this worker exits non-zero deliberately to be started again on a release its self-update put in
118
- place — and the logon task carries the task scheduler's nearest equivalent, three restarts a minute
119
- apart.
120
-
121
- A `worker` install run as root on Linux refuses outright where the `agent` account a harness is
122
- dropped to is not on the machine, and names the command that creates it: a root worker runs its
123
- harness as root without that account, so this machine would register with no harness capability,
124
- report healthy, and refuse every conversation sent to it, having reported an install that succeeded.
125
- Creating a system account on a machine somebody else owns is the operator's to do, so the install
126
- names it rather than making it. One `worker` install can still be root's — `--service=none` on a
127
- machine whose own command is the worker — and every other one is refused as root before it gets this
128
- far. An `image` carries the account already, and an install that registers a service in the invoking
129
- account's session has no second account at all: the harness runs as the person who installed it.
130
-
131
- A family a placement names and this machine lacks is installed on the preparation path instead,
132
- before its harness starts, so a first provision is not slowed by a download of hundreds of megabytes
133
- that nobody asked for. The root is where both paths install. It is deliberately not the release
134
- package's own directory, which a self-update replaces whole. Under it sit the pinned Node, each
135
- family's programs, and `workspaces/<family>`, the directory that family's harness runs in — one
136
- create at install covers all three, and a worker enrolled on somebody's own machine can take no
137
- second one.
138
- - **`worker enroll`** redeems the enrollment token this machine was given, prints the row it was
139
- issued, and then serves that machine. It does not stop at the redemption, because the holder key
140
- the redemption bound lives in this process.
141
- - **`worker check [--harness <families>] [--root <dir>]`** reads the runtime image contract from
142
- inside an image, which is where the rest of that contract cannot be read. An inspection from the
143
- container runtime sees the command, the account it starts as, the platform it was built for and the
144
- version label; whether a family's programs are there at the versions this build's facts pin, whether
145
- the Orchestra CLI is on the path that family's harness is handed, whether a Linux image has the `agent`
146
- account at uid 1000 with home `/home` and glibc's loader, and whether a Windows image has
147
- `WinDivert.dll` and `WinDivert64.sys` under `<root>\bin`, where the redirect loads them, are only
148
- visible from within. Each miss names the line of
149
- `agents/runtime/image/Dockerfile` that fixes it, and every miss is reported in one run, because the
150
- rows are independent and a person rebuilding an image should read the whole list once.
151
-
152
- It starts nothing, claims no identity and speaks to no platform, so it needs neither a bootstrap nor
153
- a daemon: `docker run --rm --entrypoint worker <image> check --harness <family>` is the whole of how
154
- it is used. With no `--harness` it checks the rows that belong to no family and says so — a customer
155
- image installs the families it needs, so holding one to every family this build can host would
156
- refuse it for lacking programs nothing on it was ever going to run.
157
- - **`worker trust-nested [--registry <host>]... [--root <dir>]`** lets the containers this machine
158
- runs trust the egress edge, which nothing does by default. It reads the authority from
159
- `SFDK_EGRESS_CA_FILE`, which the worker sets on every harness process and on the setup script;
160
- writes it into dockerd's `certs.d` for the well-known registries and each one named; writes
161
- `k3d-registries.yaml` under the root, whose `ca_file` is the authority's path inside each k3d node;
162
- and prints the flags `docker run`, compose and `k3d cluster create` need. On a machine that names no
163
- authority, such as one with no edge in front of it, it says so, writes nothing and succeeds.
164
- - **`worker version`** prints this binary's version, the release package it belongs to, and the
165
- protocol versions it speaks.
166
-
167
- The version is stamped at link time by the publishing lane, and both `package.json` files carry
168
- `0.0.0` until that lane stamps them; a build that stamps nothing says so rather than claiming a
169
- version it is not.
170
-
171
- ## The bootstrap a machine is started with
172
-
173
- Everything above assumes the worker already knows where its platform is. That is the one family of
174
- names it reads before it can speak to anything, so they are the worker's own rather than the
175
- contract's: the environment document owns the other `SFDK_` names the platform or the worker sets for a
176
- running machine and says these are outside its scope. They carry the same reserved prefix, so an environment
177
- cannot declare one; the only `SFDK_` names an environment may declare are the release settings
178
- `fwp/schema/environment-v1.json` lists under `customerDeclarable`, which the worker reads from its own
179
- environment in preference to the register answer. The platform composes what is listed here; a value it does not
180
- set is a concern this machine does not have.
181
-
182
- | name | what it carries |
183
- |---|---|
184
- | `SFDK_CLAIM_ENDPOINT_URL` | the platform's origin, not a route: the paths belong to the protocol |
185
- | `SFDK_TRANSPORT_TRUST_ROOTS` | the certificate bundle this machine trusts, inline |
186
- | `SFDK_TRANSPORT_TRUST_ROOTS_FILE` | the same bundle as an absolute path, for a delivery that cannot carry the newlines; setting both is refused |
187
- | `SFDK_CLAIM_LOCATOR` | the machine row this worker claims as |
188
- | `SFDK_ENROLLMENT_TOKEN` | the show-once token an enrolling machine redeems; mutually exclusive with the locator |
189
- | `SFDK_ENROLLMENT_TOKEN_FILE` | the same token as an absolute path; setting both is refused |
190
- | `SFDK_HOLDER_KEY_ALGORITHM` | which of the protocol's two holder-key algorithms this registration declares |
191
- | `SFDK_SUBSTRATE` | which evidence this worker gathers, so an unknown value is refused rather than defaulted |
192
- | `SFDK_CLAIM_AUDIENCE` | who the presented caller-identity request was meant for; required on the virtual-machine lanes |
193
- | `SFDK_STS_ENDPOINT_URL` | the security-token endpoint the platform will send that request to |
194
- | `SFDK_SERVICE_ACCOUNT_TOKEN_FILE` | where the audience-bound projected token is mounted, on the Kubernetes lane |
195
- | `SFDK_TPM_DEVICE_PATH` | the chip's device path where it is not this operating system's default |
196
- | `SFDK_SUBSTRATE_TOKEN` | the `local` substrate's stand-in factor, inline |
197
- | `SFDK_SUBSTRATE_TOKEN_FILE` | the same factor as an absolute path; setting both is refused |
198
- | `SFDK_INTERCEPT_PORTS` | the destination ports whose traffic goes to the egress edge, comma separated; naming none installs no redirect |
199
- | `SFDK_INTERCEPT_BYPASS` | the shared services the worker's own uid reaches directly, as `[{"host": …, "port": …}]` — our broker and our gateway, and nothing broader |
200
- | `SFDK_INTERCEPT_SKIP_RANGES` | the private ranges the redirect never takes, for any uid on any port, as comma-separated CIDR ranges; a range that is not wholly private is refused |
201
- | `SFDK_EDGE_ENDPOINT_URL` | where the redirect sends this machine's traffic; its scheme is the whole TLS statement for that hop |
202
- | `SFDK_EGRESS_LISTEN_PORT` | where the redirect lands, which is the shim's own listen port |
203
- | `SFDK_EGRESS_IPV6` | `true` mirrors the redirect through `ip6tables` |
204
- | `SFDK_EGRESS_PROXY` | `on`, the default, or `off`; `off` tells each harness that no edge of ours stands in front of this machine, as on an enrolled one. The platform writes `off` into every Launcher machine's boot file, because no proxy entrance of ours stands in front of one yet |
205
- | `SFDK_RUNTIME_ENVIRONMENT` | a whole environment folded into one JSON object of string values, for a boot file whose lines cannot carry a line break; each pair is placed unless already set |
206
- | `SFDK_VOLUME_MOUNTS` | the volumes this launch declares, as one JSON document |
207
- | `SFDK_MACHINE_SLUG` | the machine's platform-minted slug: the leading label of its preview hostnames, and what the tunnel's upgrade signs |
208
- | `SFDK_TUNNEL_ENDPOINT_URL` | where the preview gateway serves its tunnel listener; its scheme is the whole TLS statement for that hop |
209
- | `SFDK_INGRESS_PORTS` | the loopback ports this machine projects through that tunnel, comma separated; with no reserved ports below, naming none holds no tunnel |
210
- | `SFDK_INGRESS_ANY_PORT_EXCEPT` | the reserved loopback ports, comma separated; set, the tunnel is held from start, forwards every other port whatever `SFDK_INGRESS_PORTS` says, and the worker registers `ingress-any-port` |
211
-
212
- **The worker reads that file itself, on every operating system, and no service manager is handed it.**
213
- Every `worker run` places each pair into its own environment before it serves, leaving alone any name
214
- that is already set, so a person's own foreground run still wins over the file. systemd's
215
- `EnvironmentFile=` is deliberately not used: it processes backslash escapes in an unquoted value, so
216
- the two characters `\n` that carry a line break inside `SFDK_RUNTIME_ENVIRONMENT` would arrive as a bare
217
- `n` and the certificate folded in there would hold no certificate at all. A machine with no file
218
- starts anyway — that is an enrolled machine's ordinary start — and a line that is not `NAME=VALUE`
219
- stops the start and is named.
220
-
221
- Where the file sits is the one thing that differs, and only because of who may write it. On a Linux
222
- machine the platform launched, cloud-init writes `/etc/sfdk/machine.env` from the user data and the run
223
- reads it; no install by a person writes there. A `worker` install that registers a service writes
224
- `<root>/machine.env` under the worker's own root — `C:\ProgramData\sfdk\worker\machine.env` on a default
225
- machine-wide Windows install — because an install that registers the worker in one person's own session
226
- cannot write under `/etc`. A Linux run reads both, the machine-wide one first. The platform's user data writes that file on a machine
227
- it launched, and a person enrolling one writes it or exports the same names in the shell that runs
228
- `worker enroll`.
229
-
230
- Beside them the worker keeps one file of its own, `<root>/state/machine-identity.json`, owner-only.
231
- It records the row this worker is bound to and — on a machine with no chip — the software holder key
232
- it is bound under, and it decides whether the next start claims or reattaches. A row this
233
- worker has already claimed cannot be claimed again, so a start that finds its own record takes the
234
- next generation under the key instead, and it treats the enrollment token its bootstrap file still
235
- carries as spent rather than refusing to start beside the row it names. The record also carries the
236
- last run that registered with the platform and the epoch it registered under, which is what an
237
- install waits for: everything else in it is written at the claim, before there is a broker link at
238
- all. It is kept only where the row outlives the program: a virtual
239
- machine that parks and wakes with its own disk, and a machine a person enrolled. A pod is replaced
240
- rather than woken, and its row is fresh with it.
241
-
242
- ## The in-guest egress
243
-
244
- On the virtual-machine lanes this worker's own process is the transparent proxy, and it puts three
245
- things in place in one order that is not interchangeable:
246
-
247
- 1. **The edge's certificate authority into this machine's trust store**, from the file the launch
248
- wrote (or, where that file is not there, from the URL it names). Windows reads the machine
249
- certificate stores and nothing else, so a certificate that is not there is every TLS client on
250
- the machine failing its handshake with the edge; and on Windows a chain verdict computed before
251
- the import stays cached machine-wide, so an import that actually wrote also throws that cache
252
- away. It happens first because the redirect is live the moment it lands.
253
- 2. **The listener**, bound before any rule exists, so no connection is ever sent to a port with
254
- nothing behind it. On Windows an inbound firewall rule is written first as well: the rewritten
255
- packet arrives as an inbound connection, which Windows Server blocks by default.
256
- 3. **The redirect** — `iptables` on Ubuntu, WinDivert on Windows. The two WinDivert files are loaded
257
- by absolute path from `<root>\bin`, where the image bake stages them, and never through the
258
- operating system's own search order: opening the driver installs and starts a kernel service.
259
-
260
- Each accepted connection's original destination is recovered (`SO_ORIGINAL_DST` on Linux, the
261
- redirect's own proxy-port map on Windows; a recovery that misses is refused rather than sent to the
262
- peer, which carries the right host and the wrong port) and carried to the edge under the machine
263
- token this worker holds. There is no loopback round trip on these lanes: the process that holds the
264
- token is the process that presents it.
265
-
266
- What the worker's own uid reaches **without** the redirect is the launch's list — our broker and our
267
- gateway — plus three the launch cannot state: the egress edge, because redirecting the hop that
268
- carries a redirected connection is a loop with no exit; the platform's identity endpoint, because
269
- this worker calls it before it holds the token the edge would identify it by; and the metadata
270
- service, which answers the evidence that call presents. Nothing is uid-wide, so the object store
271
- stays intercepted and this worker's own content pulls meet the platform's rule at the edge.
272
-
273
- The redirect takes its Windows form on Windows and its `iptables` form everywhere else, whatever the
274
- substrate. A pod with an egress sidecar runs only the certificate step: the sidecar installs that
275
- lane's redirect and runs its own shim, which reads the token from the loopback feed below. A pod with
276
- no sidecar, as a customer's Launcher creates, names an intercept set and this worker installs the
277
- redirect itself.
278
-
279
- The private ranges `SFDK_INTERCEPT_SKIP_RANGES` names are never redirected, for any uid on any port:
280
- the rules return them ahead of the redirect, a nested container reaches them through its bridge, and
281
- the Windows filter excludes each range beside the single destinations.
282
-
283
- The proxy names a customer declares (`HTTPS_PROXY`, `HTTP_PROXY` and `NO_PROXY`, in either case) serve
284
- this worker's own calls to the platform and the broker. `NO_PROXY` must then hold `169.254.169.254`,
285
- because the metadata service is never reached through a proxy. Where a redirect is in place, the same
286
- names are left out of the environment each harness and the setup script are handed, so the agent's
287
- calls reach the redirect, which carries them to the edge.
288
-
289
- ## The egress shim's token feed
290
-
291
- The shim that intercepts this machine's traffic stamps the machine token on it, and the worker is
292
- what holds that token. On Kubernetes the two are in different containers of one pod, so the only
293
- channel between them is the network: the worker listens on `127.0.0.1:7788` — a constant both halves
294
- carry, not a placed value — and the shim connects with retry.
295
-
296
- Each frame is one JSON line, newline-terminated:
7
+ The Orchestra CLI installs this package for you. To install it alone:
297
8
 
298
- ```
299
- {"generation": 7, "token": "<the machine token>", "expiresAt": "2026-09-12T10:15:00.123456Z"}
9
+ ```bash
10
+ npm install -g @orchestraworks/worker
300
11
  ```
301
12
 
302
- On connect the worker writes the generation this machine holds now, and one line per rotation after
303
- that. It writes nothing else on that socket and never reads from it, and it listens on loopback and
304
- nowhere else. A worker restart drops the connection; the shim reconnects and is fed the current
305
- generation again. The broker credential co-issued beside the token never crosses this channel, and
306
- neither does the holder key.
13
+ ## Quick start
307
14
 
308
- ## Building and testing
15
+ To put one of your own machines to work, use the CLI. It signs you in, asks what this machine is, and installs the worker:
309
16
 
17
+ ```bash
18
+ npm install -g @orchestraworks/cli
19
+ orchestra install
310
20
  ```
311
- make verify # gofmt, then vet for linux, windows and macOS
312
- make vet-all # the vet half on its own, which is what type-checks the per-operating-system files
313
- make test # the script harness's own tests, then go test with the coverage gate
314
- make build # one static binary per release package, into npm/<platform>/, with the script harness beside it
315
- ```
316
-
317
- ### What a release is
318
-
319
- `make build` writes the five per-platform binaries into their package directories and copies the
320
- `script` harness in beside each, and those directories are what the publishing workflow packs. Two
321
- shapes of tarball come out of it, and the layout is stated here because that workflow reads it
322
- rather than deciding it:
323
-
324
- ```
325
- @orchestraworks/worker-<platform> @orchestraworks/worker
326
- package.json package.json
327
- worker (worker.exe on win32) bin/worker.js
328
- script-harness/ README.md
329
- package.json LICENSE
330
- index.js
331
- fetch-failure.js
332
- LICENSE
333
- ```
334
-
335
- Six published packages redistribute this tree, and an Apache-2.0 package that ships without the licence
336
- is redistributed on terms it does not carry — so `LICENSE` is one file here and a copy in every tarball.
337
-
338
- The per-platform tarball is the release artifact: its SHA-256 is the one digest the catalog carries,
339
- the register answer names and both fetch paths check, and it covers the binary and the `script`
340
- harness together, so a self-update carries the program with the binary. The launcher package holds
341
- no binary at all — it depends on the five as optional dependencies and runs whichever one npm
342
- installed for this machine. The build's output is ignored by `.gitignore`: the packages are
343
- published from a build rather than from the tree.
344
21
 
345
- The tests are Go tests plus two sets of Node's own — the `script` harness's, and the npm launcher's: no container runtime, no
346
- database, no stack, and no harness family is ever downloaded — a fake registry serves small archives
347
- that prove the install's own rules. The tests that
348
- need a real trusted platform module skip where there is none, and the broker suite drives a real
349
- `nats-server` named by `WORKER_TEST_NATS_SERVER` and skips when that is unset — so a `make test`
350
- meant to meet the coverage gate sets it:
22
+ To check that a machine image can host a harness, run `worker check` inside it:
351
23
 
24
+ ```bash
25
+ docker run --rm --entrypoint worker <image> check --harness claude
352
26
  ```
353
- WORKER_TEST_NATS_SERVER=/path/to/nats-server-2.14.6 make test
354
- ```
355
-
356
- `WORKER_TEST_NATS_SERVER` is a knob a test process sets for itself, not a variable the platform sets
357
- on a machine, which is why it carries no `SFDK_` prefix and appears in no environment document.
358
-
359
- ## Where a family keeps its state
360
-
361
- Every family keeps its conversation history — the transcripts that let a later session pick up where an
362
- earlier one stopped — in a directory it derives from `$HOME`. That is true of all four, and nothing in the
363
- runtime image baits any of them elsewhere:
364
-
365
- | family | state directory | how it resolves |
366
- | --- | --- | --- |
367
- | claude | `$HOME/.claude/` | Node `os.homedir()` — `$HOME` first, `getpwuid` fallback. |
368
- | codex | `$HOME/.codex/` | The Rust `home` crate, used when `CODEX_HOME` is unset. |
369
- | pi | `$HOME/.pi/agent/` (transcripts under `sessions/`) | `process.env.PI_CODING_AGENT_DIR ? resolve(...) : join(homedir(), ".pi", "agent")` — `pi-acp@0.0.33` `dist/index.js:1381` and `:1725`, and the harness's own `getAgentDir()` at `@earendil-works/pi-coding-agent@0.84.4` `dist/config.js:420-426`, whose variable NAME is assembled at `:405` from `APP_NAME` (`"pi"`). Transcripts: harness `dist/config.js:456-457`, adapter `dist/index.js:1397-1399` — the adapter prefers a `sessionDir` key in `<agentDir>/settings.json` when one is set (`:1383-1396`), and the baked settings file sets none. `pi-acp` also keeps its own session map at `$HOME/.pi/pi-acp/` (`:352`). |
370
- | script | none | The harness holds no conversational state. |
371
-
372
- An environment makes that state durable by declaring a volume mount at `/home`, which is where the
373
- account inside the runtime image already has its home directory. An environment that declares no such
374
- mount behaves exactly as it did before: the state is written to the machine's own disk and goes with it.
375
-
376
- That is why no image bakes `CODEX_HOME` or `PI_CODING_AGENT_DIR`. A pointer to a path outside `$HOME` is
377
- exactly what a durable home cannot survive: it would send codex's and pi's transcripts to the container's
378
- writable layer, which a recycle discards. Leaving both knobs unset lets each family's own `$HOME` fallback
379
- engage. Injecting a literal `CODEX_HOME="$HOME/.codex"` instead was rejected because env-value expansion is
380
- not portable — Kubernetes expands only `$(VAR)` and Docker expands nothing.
381
-
382
- ### First-boot seeding
383
-
384
- codex and pi each need one file to already exist inside their state directory: codex's egress-placeholder
385
- `auth.json`, pi's provider-pinning `settings.json`. Both are image content, but the directory they belong
386
- in lives on a volume that is empty the first time a machine boots. So the image bakes each file as a
387
- template and the family's `seed` list in `facts/families.json` says where it goes:
388
-
389
- ```
390
- codex: /opt/software-factory-codex/auth.json -> .codex/auth.json
391
- pi: /opt/software-factory-pi/settings.json -> .pi/agent/settings.json
392
- ```
393
-
394
- The worker copies each pair before it spawns the harness, resolving every destination against `$HOME`. The
395
- copy is **copy-if-absent**: the first placement lays the templates down, and every later one finds whatever
396
- the harness has written since — a refreshed token, an edited setting — and leaves it alone. Seeded files are
397
- created `0600`, parent directories `0700`, and the step never creates `$HOME` itself. The claude and script
398
- families declare an empty `seed`, so the step is a no-op for them.
399
-
400
- ### Whether a conversation resumes
401
-
402
- Durable storage makes state survive a recycle; whether a family *resumes* a conversation also needs its ACP
403
- adapter to implement `session/load`. All three model-backed adapters do, verified against the shipped
404
- artifacts:
405
-
406
- - **`claude-agent-acp@0.70.0`** — advertises `agentCapabilities.loadSession: true` and implements
407
- `loadSession(params)`, registered on the connection as the `session/load` request handler
408
- (`dist/acp-agent.js:788`, `:6909`).
409
- - **`codex-acp@1.11.0`** — implements `session/load` by resuming the codex thread and replaying that
410
- thread's history as `session/update` notifications before it answers the request. Read in the adapter's
411
- built code at the pinned version; the `0.16.0` adapter it replaced carried the same method inside its own
412
- native binary.
413
- - **`pi-acp@0.0.33`** — advertises `agentCapabilities.loadSession: true` and implements `loadSession(params)`,
414
- which respawns pi against the stored session file and replays its messages (`dist/index.js:1935`, `:2460`;
415
- the helper that binds a session object to an existing pi process is at `:753`). Confirmed on the wire: an
416
- ACP `initialize` against the shipped binary answers `protocolVersion: 1` with
417
- `agentCapabilities.loadSession: true`. That is the *capability*, not the replay itself — pi still has no
418
- captured `session/load` transcript, which is the one resume-evidence gap open on this lane.
419
-
420
- ## The command's environment is the family's pass-through list, and nothing else
421
-
422
- The worker does not *merge* an environment into a harness process: the family's `passEnv` enumeration in
423
- `facts/families.json` **is** that process's complete environment. An environment cannot declare an
424
- `SFDK_` name for it, because `SFDK_` is reserved apart from the four `SFDK_WORKER_*` release settings, which
425
- reach the worker and never a harness process. A derived image that bakes a name the list does not carry, such as `ENV JOB_TARGET=…`, gets **silence**,
426
- not an override: the variable never reaches the command, and a job running under `set -u` aborts on the
427
- first reference to it. Per-job configuration travels the prompt, the agent's instructions, or a file in the
428
- image.
429
-
430
- Every family's harness process also gets `SFDK_SESSION_ID`, which the worker sets from the `sessionId` on
431
- `session_placement.apply`: the worker adds it itself, so it needs no `passEnv` entry, whatever the worker's
432
- own environment holds. The `script` family adds only
433
- `SFDK_ORGANIZATION_ID` to the command it runs, from its stored conversation metadata; its command inherits
434
- the session variable from the harness process. Neither grants egress authority.
435
-
436
- **On Windows that list is shorter than the operating system's own habits assume.** It is
437
- OS-independent, so a Windows job sees `PATH`, `HOME` and `USERPROFILE` and *not* `TEMP`, `APPDATA` or
438
- `ComSpec` — names a PowerShell one-liner reaches for without thinking (`$env:TEMP`). The list is not a
439
- filter over the environment: the worker *builds* the child's environment from it, so a name the platform
440
- set on the service and the family does not declare is simply gone by the time the command runs.
441
-
442
- **Three names are backfilled for Windows and none of them is on any family's list.** A backfill is the
443
- treatment for operating-system *plumbing* — a variable the machine needs to work at all, as opposed to
444
- configuration a customer chose — and it is if-absent, so a declared spelling always wins. Two come from
445
- libraries: Go's `os/exec` adds `SYSTEMROOT` whenever a caller sets an explicit environment, and libuv copies
446
- its own required-variable set into a spawned child from *this* process's environment — which is why `PATH`
447
- reaching the command is what lets `powershell.exe` and `taskkill.exe` resolve at all. The third is ours:
448
- **`PATHEXT`**, the list of extensions PowerShell treats as *executable*. With it unset the effective list
449
- collapses to `.CPL`, so a `curl.exe` PowerShell resolved on `PATH` is a **document** rather than a program
450
- and invoking it throws `CantActivateDocumentInPipeline` — before a process starts, before a packet leaves.
451
- Every network probe on the Windows lane came back `000`, which reads exactly like a blocked connection.
452
-
453
- The neighbours a reader expects beside it were **measured and left out**, each changing nothing on a Windows
454
- Server 2025 runtime AMI: `cmd.exe` and `.bat` files run from CreateProcess's own default without `ComSpec`,
455
- `%SystemRoot%` expands from the `SYSTEMROOT` `os/exec` already backfills, `SystemDrive` and `windir` are
456
- unused by realistic work, and a service's temp directory falls back to `C:\Windows\SystemTemp` without
457
- `TEMP` or `TMP`. Add a fourth name only with that kind of evidence. Machine plumbing is not configuration
458
- and must never be pushed into the prompt: a job cannot set `PATHEXT` for a shell that has already refused to
459
- start its first program.
460
-
461
- Three more bounds a job author should know:
462
-
463
- - **Turn wall clock.** The platform bounds a turn with its own response timeout (default 600 seconds,
464
- configurable). A job that can run longer needs that raised, or it is killed mid-turn and reported as a
465
- runtime failure regardless of what the script was doing.
466
- - **`session/load` re-executes.** The script harness holds no conversational state, so a resumed
467
- conversation's command runs again. That is the safe direction — a fresh, deterministic run — and another
468
- reason to keep commands idempotent.
469
- - **`cwd` is the session's.** `session/new` (and `session/load` on a revival) carries the workspace path, the
470
- worker maps it to this machine's native filesystem — `C:\sf\workspace` on Windows — and the harness spawns
471
- the command there. A session that declares none inherits the harness's own directory; one that does not
472
- exist fails the spawn, which the turn reports as an ordinary exit-127 failure with the operating system's
473
- message attached.
474
27
 
475
- ## What the `script` harness answers
28
+ `worker version` prints the version and the protocol versions this build speaks.
476
29
 
477
- The worker runs one harness process per session placement and speaks ACP to it as its client. The `script`
478
- family's harness is the one the worker carries itself (`script-harness/`), and its dispatch table is:
30
+ ## For coding agents
479
31
 
480
- - `initialize` — the harness identifies itself.
481
- - `session/new` — a fresh session id, with the session's declared working directory anchored on it.
482
- - `session/load` — the session's replay, as the `session/update` notifications emitted between the request
483
- and its response. It executes nothing and calls nothing, and it restores the duplicate-execution guard
484
- below, which the notifications cannot carry.
485
- - `session/prompt` — execute the agent's instructions, stream stdout and stderr as `session/update`
486
- notifications, move platform state from the exit code, then answer. Streaming is bounded twice: each frame
487
- is sliced at 32000 characters (one over-long line would trip the line scanner and kill every session on
488
- the pipe, not just this turn), and a turn's **total** streamed volume is capped at 5 MiB of **wire** bytes
489
- — the JSON-encoded frame text, since that is what the platform records, and JSON escaping inflates
490
- control-heavy output up to about six times (one NUL byte is six wire bytes). Past the ceiling the harness
491
- emits one truncation notice and stops streaming, while the command still runs to completion with its
492
- exit-code contract unchanged.
493
- - `session/cancel` — stop the in-flight command **and everything it started**, then answer. On Linux that is
494
- a `SIGTERM` to the command's whole **process group** (commands run detached, so the group dies with the
495
- command and not just with the `bash -c` parent); on Windows there are no process groups to signal, so it
496
- is `taskkill /T /F /PID <pid>`, which walks the process tree. A cancelled turn made no determination, so
497
- it claims nothing and answers the prompt with a JSON-RPC error. Whether a turn counts as cancelled is read
498
- off an observed **fact about the kill**, never off the arrival of the cancel, and each operating system
499
- supplies the fact it can: on Linux the command's **exit shape** (one the signal actually killed reports
500
- `signal = SIGTERM`, one that had already exited reports its own status), and on Windows — which terminates
501
- a process with an exit *code* and no signal — the kill's own answer, so the harness checks whether the
502
- shell was still running when the cancel arrived and only then kills. Either way a cancel racing a script's
503
- own `exit 0`, including the window after the shell exits but before the harness reaps it, leaves the turn
504
- with the result it earned, claim included. (That liveness check is also a safety requirement on Windows,
505
- which recycles process ids: a `/T` on a stale one would kill a stranger's process tree.)
32
+ - `orchestra worker --help` lists the worker commands. `orchestra worker <command> --help` gives their flags.
33
+ - `worker` with no command prints its usage.
34
+ - `skills/orchestra/SKILL.md` in `@orchestraworks/cli` explains machines and what the platform can do.
35
+ - [docs/REFERENCE.md](docs/REFERENCE.md): every command, the start-up settings, egress, and how a release is built.
506
36
 
507
- A turn that settles while a backgrounded grandchild still holds the pipes signals that process group and
508
- detaches the streams, so nothing leaks into the next turn's transcript. That sweep is **conditional and
509
- best-effort**, and a job author should not read it as a guarantee that background work dies with the turn:
510
- it runs only on the grace-timer path, so a background job whose stdio is redirected (`cmd >/dev/null 2>&1 &`)
511
- lets the pipes close promptly, settles the turn before the timer, and survives it untouched. It is also
512
- `SIGTERM`-only — a process that traps or ignores `TERM` survives either way. Nothing of the sort leaks into
513
- the *transcript*; what survives is the process.
37
+ ## License
514
38
 
515
- A second prompt on a **workflow-step** session whose earlier claim is confirmed landed is answered as a
516
- no-op turn without re-executing the command (warm reuse and steer promotion both make that routine); an
517
- earlier claim that is merely *attempted* is re-confirmed against the run first. The no-op is step-scoped on
518
- purpose: an ordinary message appended to a direct-start session is a real new turn and always executes.
39
+ Apache License 2.0. The full text is in [LICENSE](LICENSE).
@@ -0,0 +1,520 @@
1
+ # The worker: the full reference
2
+
3
+ The short introduction is the package [README](../README.md). Paths below are relative to the `worker/` folder.
4
+
5
+ One Go program runs inside every machine and speaks the Factory Worker Protocol to the
6
+ platform. This tree is that program, and it is also the `@orchestraworks/worker` npm package: the launcher on
7
+ the path, and one per-platform package carrying the binary for the machine it is installed on.
8
+
9
+ The worker exposes no API of its own. Every route it calls is the platform's.
10
+
11
+ ## The tree
12
+
13
+ | path | what it holds |
14
+ |---|---|
15
+ | `cmd/worker/` | the one binary: `run`, `install`, `enroll`, `check`, `version` |
16
+ | `identity/` | the four identity exchanges — claim, refresh, reattach-challenge, reattach — the holder key, and the loop that keeps the machine token alive |
17
+ | `broker/` | the protocol on the message broker: the connection, the register on every connect, the heartbeat and the status cadence, the pull of this machine's own work consumer, and the publisher every other package sends through |
18
+ | `env/` | the machine-side environment the platform composes onto a launch, read from the names the contract declares |
19
+ | `facts/` | what this build knows about the four harness families: the program to install and the digest that authenticates it, the environment each family is handed, the templates it seeds, and every location its content lands in — embedded in the binary, so one release digest covers the facts and the program that reads them |
20
+ | `harness/` | the only place family-specific knowledge lives: the launch plan, the content layout, the store lookup, the mid-turn message each family accepts, and the installer that puts the pinned Node and a family's programs under the worker root |
21
+ | `process/` | the agent's setup and the harness process behind each session placement: the start, the stop and the last save a remove asks for, the startup fence over what a previous run left, and the stop of everything when the worker is told to stop |
22
+ | `stream/` | the queue: every placement message and every prompt, steer and cancel pulled off this machine's work consumer, the ACP client that opens a session's harness conversation and runs its turns, each turn's events batched to the platform and published live, and the answer every placement message is owed |
23
+ | `mount/` | the volumes the launch declares, placed on the lanes that mount them: the home volume first, because anything written under a mount point before the mount is hidden by it |
24
+ | `network/` | the transparent redirect on the virtual-machine lanes, as the iptables rules the launch describes: what the worker's own uid reaches directly, and what every uid on the machine reaches through the egress edge |
25
+ | `shim/` | the in-guest transparent proxy: the listener the redirect lands on, the recovery of the destination each connection was actually for, the hop to the egress edge under the machine-identity proof, the WinDivert half that is the Windows redirect, and the loopback feed a separate shim reads this machine's token from |
26
+ | `egress/` | what puts that proxy in place on the virtual-machine lanes: the edge's certificate authority in this machine's trust store, the inbound firewall rule Windows needs, the redirect itself, and the state line an operator reads off the machine's console |
27
+ | `ingress/` | the machine's half of the preview tunnel: the connection this worker holds open to the preview gateway, the challenge and the signed upgrade that authenticate it, and the loopback pipe that carries each viewer request to a port this machine declared and refuses every other |
28
+ | `termination/` | the cloud's notice that this machine is being reclaimed, delivered to the broker so the next status report carries it |
29
+ | `paths/` | which native place a logical path denotes, and when two logical paths denote one place — pinned against the platform's own answer by committed vectors |
30
+ | `service/` | `worker run` under this operating system's service manager, in one of two scopes: a system or user `systemd` unit, a Windows service or logon task, a launchd daemon or launch agent |
31
+ | `file/` | everything that moves bytes onto this machine and back off it: the install of the skills, memory, instructions and model settings an agent placement's apply carries, against the install journal, the capture that offers the machine's declared roots back to the platform, the restore that lays a previous machine's capture down before a manifest is installed over it, and the self-update that replaces this worker's own release package |
32
+ | `script-harness/` | the first-party `script` harness program, which is the one family the worker package carries rather than installs |
33
+ | `bin/worker.js` | the npm launcher: it runs the binary out of the per-platform package installed for this machine, stays in front of it, and forwards the signal a service manager ends it with. Its own tests sit beside it and are not packed |
34
+ | `npm/<platform>/` | the five per-platform packages, one per release target: the binary, and the `script` harness copied in beside it at build time |
35
+
36
+ `../fwp` is the protocol. Every wire shape comes from `github.com/livingcomputers/fwp/go/wire`
37
+ through a relative `replace`, so the program and the conformance suite compile against the exact
38
+ shapes the schema ships, and nothing here re-declares one.
39
+
40
+ ## The commands
41
+
42
+ - **`worker run`** claims this machine's identity — or reattaches, where its row is already bound to
43
+ the key this machine holds — and then serves the machine. `run` hands the held identity, with the
44
+ broker credential the claim co-issued, to the registrar seam in `cmd/worker/run.go`, which the
45
+ `broker` package fills: it dials, registers the machine on every connect, pulls the machine's own
46
+ work consumer and reports on the cadence.
47
+
48
+ Before any of that, `run` starts what the launch describes: it places the declared volumes, puts
49
+ the in-guest egress in place where the launch names an intercept set, opens the egress shim's token
50
+ feed, and — on a machine the cloud can reclaim — watches for the notice that it is going away. Each
51
+ is started only where it is described: a Kubernetes pod places its own volumes and its egress
52
+ sidecar installs its own redirect, so a launch there names neither. If one of those duties is
53
+ installed and then dies, `run` ends rather than serving a machine the platform counts as healthy
54
+ with no route off it; the service manager starts the worker again.
55
+
56
+ `--service` is how the process is hosted rather than what it does: on Windows it hands the body to
57
+ the service control manager, which kills a service that does not answer its handshake, and on the
58
+ other operating systems it is refused by name, because systemd and launchd supervise a foreground
59
+ process from outside.
60
+
61
+ On a machine that hosts a harness, `run` also builds the three lanes that answer the platform's
62
+ verbs: `process`, which holds the placements, `stream`, which carries their turns, and `file`,
63
+ which installs what an agent is made of and captures what its work left behind. The three are
64
+ built before the broker, because the broker's configuration carries their handlers, and each
65
+ reaches it afterwards through the deferred link in `cmd/worker/serve.go`.
66
+
67
+ Once the register has been answered, `run` starts the capture cadence and, where the platform named
68
+ a version other than the one this process is running, the self-update: the release package is
69
+ fetched from the npm registry when this machine can reach one and from the platform when it cannot,
70
+ checked against the digest the register answer named, unpacked over the release package directory,
71
+ and the process restarted — but only once no turn has been open for five seconds together, because one
72
+ open turn per machine is the platform's rule and killing one would lose its answer. A machine whose
73
+ turns never leave that gap keeps running the version it has until some later start picks the unpacked
74
+ one up.
75
+ - **`worker install --role worker|image|laptop --harness <families> --yes [--root <dir>] [--service=user|none]`**
76
+ installs the worker on this machine and starts its service. Every answer comes from a flag: a missing
77
+ one is an error and never a default. It installs the harness half itself — the pinned Node, then each
78
+ named family's programs from the facts file's own pins, with the tarball checked against the pinned
79
+ digest before anything is unpacked — and then registers the service. The per-operating-system service
80
+ forms fill the installer seam in `cmd/worker/install.go`; a build with none refuses rather than
81
+ reporting an install it did not perform.
82
+
83
+ **`--service` says which session the worker is registered in, and neither of its two values needs any
84
+ privilege.** `user` registers it in the invoking account's own session: a launch agent under
85
+ `~/Library/LaunchAgents` on macOS, a `systemd --user` unit under `~/.config/systemd/user` on Linux, a
86
+ scheduled task the account's logon starts on Windows. `none` installs the files and registers nothing,
87
+ for the one host that has no service manager — a container image, whose own command is the worker —
88
+ and `false` is that same answer under its older spelling. The default is `user`. An install run as
89
+ root is refused, because it would put the worker's root under root's home and run every harness as
90
+ root; there is no privileged install to redirect it to, and `--service=system` is refused with the
91
+ same answer. Nothing on a machine somebody owns needs privilege: network enforcement is outside the
92
+ machine, at the egress edge and the directional security group.
93
+
94
+ **The machine-wide service is the image role's, and the role is the whole of what asks for it.** An
95
+ `image` gets it — a property list under `/Library/LaunchDaemons`, a unit under `/etc/systemd/system`,
96
+ a Windows service — because the platform builds the image with nobody logged in to it and nobody
97
+ typing the command, and the builder is already root. No flag selects that scope, on any role.
98
+
99
+ `--root` follows the scope the same way: `~/.sfdk/worker` (`%LOCALAPPDATA%\sfdk\worker` on Windows) for
100
+ a `user` install, `/var/lib/sfdk/worker` (`%ProgramData%\orchestra\worker`) for the other two. A `user`
101
+ install writes its log under that root, at `log/worker.log`, because no service manager in that scope
102
+ keeps one a person can read.
103
+
104
+ What the role decides is when the service runs — a `worker` starts now and at every boot, an `image`
105
+ being baked starts at the first boot of a machine made from it and claims no identity now, and a
106
+ `laptop` starts when a person enrolls a machine and says so. An `image` install also refuses, before
107
+ it installs anything, when a bootstrap file already sits at any path the run would read: an image is
108
+ copied onto every machine made from it, so a baked bootstrap would hand one machine's identity to all
109
+ of them.
110
+
111
+ Each form carries one thing its machine would otherwise be wrong without. A macOS launch DAEMON — the
112
+ image role's form, which a person never reaches — names the account that ran the install, because
113
+ launchd's system domain runs a daemon as root otherwise, and a harness runs as whoever this worker is;
114
+ the worker's own root and the daemon's log file are handed to that account in the same step, so the
115
+ service can read the owner-only bootstrap the root install just wrote, write its state file beside it,
116
+ and say in its log why it did not start. A launch AGENT names no account: it is
117
+ already the session's. The Windows service is registered to restart itself three times at two seconds,
118
+ on a clean non-zero exit as well as a crash — the control manager restarts nothing it was not told to,
119
+ and this worker exits non-zero deliberately to be started again on a release its self-update put in
120
+ place — and the logon task carries the task scheduler's nearest equivalent, three restarts a minute
121
+ apart.
122
+
123
+ A `worker` install run as root on Linux refuses outright where the `agent` account a harness is
124
+ dropped to is not on the machine, and names the command that creates it: a root worker runs its
125
+ harness as root without that account, so this machine would register with no harness capability,
126
+ report healthy, and refuse every conversation sent to it, having reported an install that succeeded.
127
+ Creating a system account on a machine somebody else owns is the operator's to do, so the install
128
+ names it rather than making it. One `worker` install can still be root's — `--service=none` on a
129
+ machine whose own command is the worker — and every other one is refused as root before it gets this
130
+ far. An `image` carries the account already, and an install that registers a service in the invoking
131
+ account's session has no second account at all: the harness runs as the person who installed it.
132
+
133
+ A family a placement names and this machine lacks is installed on the preparation path instead,
134
+ before its harness starts, so a first provision is not slowed by a download of hundreds of megabytes
135
+ that nobody asked for. The root is where both paths install. It is deliberately not the release
136
+ package's own directory, which a self-update replaces whole. Under it sit the pinned Node, each
137
+ family's programs, and `workspaces/<family>`, the directory that family's harness runs in — one
138
+ create at install covers all three, and a worker enrolled on somebody's own machine can take no
139
+ second one.
140
+ - **`worker enroll`** redeems the enrollment token this machine was given, prints the row it was
141
+ issued, and then serves that machine. It does not stop at the redemption, because the holder key
142
+ the redemption bound lives in this process.
143
+ - **`worker check [--harness <families>] [--root <dir>]`** reads the runtime image contract from
144
+ inside an image, which is where the rest of that contract cannot be read. An inspection from the
145
+ container runtime sees the command, the account it starts as, the platform it was built for and the
146
+ version label; whether a family's programs are there at the versions this build's facts pin, whether
147
+ the Orchestra CLI is on the path that family's harness is handed, whether a Linux image has the `agent`
148
+ account at uid 1000 with home `/home` and glibc's loader, and whether a Windows image has
149
+ `WinDivert.dll` and `WinDivert64.sys` under `<root>\bin`, where the redirect loads them, are only
150
+ visible from within. Each miss names the line of
151
+ `agents/runtime/image/Dockerfile` that fixes it, and every miss is reported in one run, because the
152
+ rows are independent and a person rebuilding an image should read the whole list once.
153
+
154
+ It starts nothing, claims no identity and speaks to no platform, so it needs neither a bootstrap nor
155
+ a daemon: `docker run --rm --entrypoint worker <image> check --harness <family>` is the whole of how
156
+ it is used. With no `--harness` it checks the rows that belong to no family and says so — a customer
157
+ image installs the families it needs, so holding one to every family this build can host would
158
+ refuse it for lacking programs nothing on it was ever going to run.
159
+ - **`worker trust-nested [--registry <host>]... [--root <dir>]`** lets the containers this machine
160
+ runs trust the egress edge, which nothing does by default. It reads the authority from
161
+ `SFDK_EGRESS_CA_FILE`, which the worker sets on every harness process and on the setup script;
162
+ writes it into dockerd's `certs.d` for the well-known registries and each one named; writes
163
+ `k3d-registries.yaml` under the root, whose `ca_file` is the authority's path inside each k3d node;
164
+ and prints the flags `docker run`, compose and `k3d cluster create` need. On a machine that names no
165
+ authority, such as one with no edge in front of it, it says so, writes nothing and succeeds.
166
+ - **`worker version`** prints this binary's version, the release package it belongs to, and the
167
+ protocol versions it speaks.
168
+
169
+ The version is stamped at link time by the publishing lane, and both `package.json` files carry
170
+ `0.0.0` until that lane stamps them; a build that stamps nothing says so rather than claiming a
171
+ version it is not.
172
+
173
+ ## The bootstrap a machine is started with
174
+
175
+ Everything above assumes the worker already knows where its platform is. That is the one family of
176
+ names it reads before it can speak to anything, so they are the worker's own rather than the
177
+ contract's: the environment document owns the other `SFDK_` names the platform or the worker sets for a
178
+ running machine and says these are outside its scope. They carry the same reserved prefix, so an environment
179
+ cannot declare one; the only `SFDK_` names an environment may declare are the release settings
180
+ `fwp/schema/environment-v1.json` lists under `customerDeclarable`, which the worker reads from its own
181
+ environment in preference to the register answer. The platform composes what is listed here; a value it does not
182
+ set is a concern this machine does not have.
183
+
184
+ | name | what it carries |
185
+ |---|---|
186
+ | `SFDK_CLAIM_ENDPOINT_URL` | the platform's origin, not a route: the paths belong to the protocol |
187
+ | `SFDK_TRANSPORT_TRUST_ROOTS` | the certificate bundle this machine trusts, inline |
188
+ | `SFDK_TRANSPORT_TRUST_ROOTS_FILE` | the same bundle as an absolute path, for a delivery that cannot carry the newlines; setting both is refused |
189
+ | `SFDK_CLAIM_LOCATOR` | the machine row this worker claims as |
190
+ | `SFDK_ENROLLMENT_TOKEN` | the show-once token an enrolling machine redeems; mutually exclusive with the locator |
191
+ | `SFDK_ENROLLMENT_TOKEN_FILE` | the same token as an absolute path; setting both is refused |
192
+ | `SFDK_HOLDER_KEY_ALGORITHM` | which of the protocol's two holder-key algorithms this registration declares |
193
+ | `SFDK_SUBSTRATE` | which evidence this worker gathers, so an unknown value is refused rather than defaulted |
194
+ | `SFDK_CLAIM_AUDIENCE` | who the presented caller-identity request was meant for; required on the virtual-machine lanes |
195
+ | `SFDK_STS_ENDPOINT_URL` | the security-token endpoint the platform will send that request to |
196
+ | `SFDK_SERVICE_ACCOUNT_TOKEN_FILE` | where the audience-bound projected token is mounted, on the Kubernetes lane |
197
+ | `SFDK_TPM_DEVICE_PATH` | the chip's device path where it is not this operating system's default |
198
+ | `SFDK_SUBSTRATE_TOKEN` | the `local` substrate's stand-in factor, inline |
199
+ | `SFDK_SUBSTRATE_TOKEN_FILE` | the same factor as an absolute path; setting both is refused |
200
+ | `SFDK_INTERCEPT_PORTS` | the destination ports whose traffic goes to the egress edge, comma separated; naming none installs no redirect |
201
+ | `SFDK_INTERCEPT_BYPASS` | the shared services the worker's own uid reaches directly, as `[{"host": …, "port": …}]` — our broker and our gateway, and nothing broader |
202
+ | `SFDK_INTERCEPT_SKIP_RANGES` | the private ranges the redirect never takes, for any uid on any port, as comma-separated CIDR ranges; a range that is not wholly private is refused |
203
+ | `SFDK_EDGE_ENDPOINT_URL` | where the redirect sends this machine's traffic; its scheme is the whole TLS statement for that hop |
204
+ | `SFDK_EGRESS_LISTEN_PORT` | where the redirect lands, which is the shim's own listen port |
205
+ | `SFDK_EGRESS_IPV6` | `true` mirrors the redirect through `ip6tables` |
206
+ | `SFDK_EGRESS_PROXY` | `on`, the default, or `off`; `off` tells each harness that no edge of ours stands in front of this machine, as on an enrolled one. The platform writes `off` into every Launcher machine's boot file, because no proxy entrance of ours stands in front of one yet |
207
+ | `SFDK_RUNTIME_ENVIRONMENT` | a whole environment folded into one JSON object of string values, for a boot file whose lines cannot carry a line break; each pair is placed unless already set |
208
+ | `SFDK_VOLUME_MOUNTS` | the volumes this launch declares, as one JSON document |
209
+ | `SFDK_MACHINE_SLUG` | the machine's platform-minted slug: the leading label of its preview hostnames, and what the tunnel's upgrade signs |
210
+ | `SFDK_TUNNEL_ENDPOINT_URL` | where the preview gateway serves its tunnel listener; its scheme is the whole TLS statement for that hop |
211
+ | `SFDK_INGRESS_PORTS` | the loopback ports this machine projects through that tunnel, comma separated; with no reserved ports below, naming none holds no tunnel |
212
+ | `SFDK_INGRESS_ANY_PORT_EXCEPT` | the reserved loopback ports, comma separated; set, the tunnel is held from start, forwards every other port whatever `SFDK_INGRESS_PORTS` says, and the worker registers `ingress-any-port` |
213
+
214
+ **The worker reads that file itself, on every operating system, and no service manager is handed it.**
215
+ Every `worker run` places each pair into its own environment before it serves, leaving alone any name
216
+ that is already set, so a person's own foreground run still wins over the file. systemd's
217
+ `EnvironmentFile=` is deliberately not used: it processes backslash escapes in an unquoted value, so
218
+ the two characters `\n` that carry a line break inside `SFDK_RUNTIME_ENVIRONMENT` would arrive as a bare
219
+ `n` and the certificate folded in there would hold no certificate at all. A machine with no file
220
+ starts anyway — that is an enrolled machine's ordinary start — and a line that is not `NAME=VALUE`
221
+ stops the start and is named.
222
+
223
+ Where the file sits is the one thing that differs, and only because of who may write it. On a Linux
224
+ machine the platform launched, cloud-init writes `/etc/sfdk/machine.env` from the user data and the run
225
+ reads it; no install by a person writes there. A `worker` install that registers a service writes
226
+ `<root>/machine.env` under the worker's own root — `C:\ProgramData\sfdk\worker\machine.env` on a default
227
+ machine-wide Windows install — because an install that registers the worker in one person's own session
228
+ cannot write under `/etc`. A Linux run reads both, the machine-wide one first. The platform's user data writes that file on a machine
229
+ it launched, and a person enrolling one writes it or exports the same names in the shell that runs
230
+ `worker enroll`.
231
+
232
+ Beside them the worker keeps one file of its own, `<root>/state/machine-identity.json`, owner-only.
233
+ It records the row this worker is bound to and — on a machine with no chip — the software holder key
234
+ it is bound under, and it decides whether the next start claims or reattaches. A row this
235
+ worker has already claimed cannot be claimed again, so a start that finds its own record takes the
236
+ next generation under the key instead, and it treats the enrollment token its bootstrap file still
237
+ carries as spent rather than refusing to start beside the row it names. The record also carries the
238
+ last run that registered with the platform and the epoch it registered under, which is what an
239
+ install waits for: everything else in it is written at the claim, before there is a broker link at
240
+ all. It is kept only where the row outlives the program: a virtual
241
+ machine that parks and wakes with its own disk, and a machine a person enrolled. A pod is replaced
242
+ rather than woken, and its row is fresh with it.
243
+
244
+ ## The in-guest egress
245
+
246
+ On the virtual-machine lanes this worker's own process is the transparent proxy, and it puts three
247
+ things in place in one order that is not interchangeable:
248
+
249
+ 1. **The edge's certificate authority into this machine's trust store**, from the file the launch
250
+ wrote (or, where that file is not there, from the URL it names). Windows reads the machine
251
+ certificate stores and nothing else, so a certificate that is not there is every TLS client on
252
+ the machine failing its handshake with the edge; and on Windows a chain verdict computed before
253
+ the import stays cached machine-wide, so an import that actually wrote also throws that cache
254
+ away. It happens first because the redirect is live the moment it lands.
255
+ 2. **The listener**, bound before any rule exists, so no connection is ever sent to a port with
256
+ nothing behind it. On Windows an inbound firewall rule is written first as well: the rewritten
257
+ packet arrives as an inbound connection, which Windows Server blocks by default.
258
+ 3. **The redirect** — `iptables` on Ubuntu, WinDivert on Windows. The two WinDivert files are loaded
259
+ by absolute path from `<root>\bin`, where the image bake stages them, and never through the
260
+ operating system's own search order: opening the driver installs and starts a kernel service.
261
+
262
+ Each accepted connection's original destination is recovered (`SO_ORIGINAL_DST` on Linux, the
263
+ redirect's own proxy-port map on Windows; a recovery that misses is refused rather than sent to the
264
+ peer, which carries the right host and the wrong port) and carried to the edge under the machine
265
+ token this worker holds. There is no loopback round trip on these lanes: the process that holds the
266
+ token is the process that presents it.
267
+
268
+ What the worker's own uid reaches **without** the redirect is the launch's list — our broker and our
269
+ gateway — plus three the launch cannot state: the egress edge, because redirecting the hop that
270
+ carries a redirected connection is a loop with no exit; the platform's identity endpoint, because
271
+ this worker calls it before it holds the token the edge would identify it by; and the metadata
272
+ service, which answers the evidence that call presents. Nothing is uid-wide, so the object store
273
+ stays intercepted and this worker's own content pulls meet the platform's rule at the edge.
274
+
275
+ The redirect takes its Windows form on Windows and its `iptables` form everywhere else, whatever the
276
+ substrate. A pod with an egress sidecar runs only the certificate step: the sidecar installs that
277
+ lane's redirect and runs its own shim, which reads the token from the loopback feed below. A pod with
278
+ no sidecar, as a customer's Launcher creates, names an intercept set and this worker installs the
279
+ redirect itself.
280
+
281
+ The private ranges `SFDK_INTERCEPT_SKIP_RANGES` names are never redirected, for any uid on any port:
282
+ the rules return them ahead of the redirect, a nested container reaches them through its bridge, and
283
+ the Windows filter excludes each range beside the single destinations.
284
+
285
+ The proxy names a customer declares (`HTTPS_PROXY`, `HTTP_PROXY` and `NO_PROXY`, in either case) serve
286
+ this worker's own calls to the platform and the broker. `NO_PROXY` must then hold `169.254.169.254`,
287
+ because the metadata service is never reached through a proxy. Where a redirect is in place, the same
288
+ names are left out of the environment each harness and the setup script are handed, so the agent's
289
+ calls reach the redirect, which carries them to the edge.
290
+
291
+ ## The egress shim's token feed
292
+
293
+ The shim that intercepts this machine's traffic stamps the machine token on it, and the worker is
294
+ what holds that token. On Kubernetes the two are in different containers of one pod, so the only
295
+ channel between them is the network: the worker listens on `127.0.0.1:7788` — a constant both halves
296
+ carry, not a placed value — and the shim connects with retry.
297
+
298
+ Each frame is one JSON line, newline-terminated:
299
+
300
+ ```
301
+ {"generation": 7, "token": "<the machine token>", "expiresAt": "2026-09-12T10:15:00.123456Z"}
302
+ ```
303
+
304
+ On connect the worker writes the generation this machine holds now, and one line per rotation after
305
+ that. It writes nothing else on that socket and never reads from it, and it listens on loopback and
306
+ nowhere else. A worker restart drops the connection; the shim reconnects and is fed the current
307
+ generation again. The broker credential co-issued beside the token never crosses this channel, and
308
+ neither does the holder key.
309
+
310
+ ## Building and testing
311
+
312
+ ```
313
+ make verify # gofmt, then vet for linux, windows and macOS
314
+ make vet-all # the vet half on its own, which is what type-checks the per-operating-system files
315
+ make test # the script harness's own tests, then go test with the coverage gate
316
+ make build # one static binary per release package, into npm/<platform>/, with the script harness beside it
317
+ ```
318
+
319
+ ### What a release is
320
+
321
+ `make build` writes the five per-platform binaries into their package directories and copies the
322
+ `script` harness in beside each, and those directories are what the publishing workflow packs. Two
323
+ shapes of tarball come out of it, and the layout is stated here because that workflow reads it
324
+ rather than deciding it:
325
+
326
+ ```
327
+ @orchestraworks/worker-<platform> @orchestraworks/worker
328
+ package.json package.json
329
+ worker (worker.exe on win32) bin/worker.js
330
+ script-harness/ README.md
331
+ package.json LICENSE
332
+ index.js
333
+ fetch-failure.js
334
+ LICENSE
335
+ ```
336
+
337
+ Six published packages redistribute this tree, and an Apache-2.0 package that ships without the licence
338
+ is redistributed on terms it does not carry — so `LICENSE` is one file here and a copy in every tarball.
339
+
340
+ The per-platform tarball is the release artifact: its SHA-256 is the one digest the catalog carries,
341
+ the register answer names and both fetch paths check, and it covers the binary and the `script`
342
+ harness together, so a self-update carries the program with the binary. The launcher package holds
343
+ no binary at all — it depends on the five as optional dependencies and runs whichever one npm
344
+ installed for this machine. The build's output is ignored by `.gitignore`: the packages are
345
+ published from a build rather than from the tree.
346
+
347
+ The tests are Go tests plus two sets of Node's own — the `script` harness's, and the npm launcher's: no container runtime, no
348
+ database, no stack, and no harness family is ever downloaded — a fake registry serves small archives
349
+ that prove the install's own rules. The tests that
350
+ need a real trusted platform module skip where there is none, and the broker suite drives a real
351
+ `nats-server` named by `WORKER_TEST_NATS_SERVER` and skips when that is unset — so a `make test`
352
+ meant to meet the coverage gate sets it:
353
+
354
+ ```
355
+ WORKER_TEST_NATS_SERVER=/path/to/nats-server-2.14.6 make test
356
+ ```
357
+
358
+ `WORKER_TEST_NATS_SERVER` is a knob a test process sets for itself, not a variable the platform sets
359
+ on a machine, which is why it carries no `SFDK_` prefix and appears in no environment document.
360
+
361
+ ## Where a family keeps its state
362
+
363
+ Every family keeps its conversation history — the transcripts that let a later session pick up where an
364
+ earlier one stopped — in a directory it derives from `$HOME`. That is true of all four, and nothing in the
365
+ runtime image baits any of them elsewhere:
366
+
367
+ | family | state directory | how it resolves |
368
+ | --- | --- | --- |
369
+ | claude | `$HOME/.claude/` | Node `os.homedir()` — `$HOME` first, `getpwuid` fallback. |
370
+ | codex | `$HOME/.codex/` | The Rust `home` crate, used when `CODEX_HOME` is unset. |
371
+ | pi | `$HOME/.pi/agent/` (transcripts under `sessions/`) | `process.env.PI_CODING_AGENT_DIR ? resolve(...) : join(homedir(), ".pi", "agent")` — `pi-acp@0.0.33` `dist/index.js:1381` and `:1725`, and the harness's own `getAgentDir()` at `@earendil-works/pi-coding-agent@0.84.4` `dist/config.js:420-426`, whose variable NAME is assembled at `:405` from `APP_NAME` (`"pi"`). Transcripts: harness `dist/config.js:456-457`, adapter `dist/index.js:1397-1399` — the adapter prefers a `sessionDir` key in `<agentDir>/settings.json` when one is set (`:1383-1396`), and the baked settings file sets none. `pi-acp` also keeps its own session map at `$HOME/.pi/pi-acp/` (`:352`). |
372
+ | script | none | The harness holds no conversational state. |
373
+
374
+ An environment makes that state durable by declaring a volume mount at `/home`, which is where the
375
+ account inside the runtime image already has its home directory. An environment that declares no such
376
+ mount behaves exactly as it did before: the state is written to the machine's own disk and goes with it.
377
+
378
+ That is why no image bakes `CODEX_HOME` or `PI_CODING_AGENT_DIR`. A pointer to a path outside `$HOME` is
379
+ exactly what a durable home cannot survive: it would send codex's and pi's transcripts to the container's
380
+ writable layer, which a recycle discards. Leaving both knobs unset lets each family's own `$HOME` fallback
381
+ engage. Injecting a literal `CODEX_HOME="$HOME/.codex"` instead was rejected because env-value expansion is
382
+ not portable — Kubernetes expands only `$(VAR)` and Docker expands nothing.
383
+
384
+ ### First-boot seeding
385
+
386
+ codex and pi each need one file to already exist inside their state directory: codex's egress-placeholder
387
+ `auth.json`, pi's provider-pinning `settings.json`. Both are image content, but the directory they belong
388
+ in lives on a volume that is empty the first time a machine boots. So the image bakes each file as a
389
+ template and the family's `seed` list in `facts/families.json` says where it goes:
390
+
391
+ ```
392
+ codex: /opt/software-factory-codex/auth.json -> .codex/auth.json
393
+ pi: /opt/software-factory-pi/settings.json -> .pi/agent/settings.json
394
+ ```
395
+
396
+ The worker copies each pair before it spawns the harness, resolving every destination against `$HOME`. The
397
+ copy is **copy-if-absent**: the first placement lays the templates down, and every later one finds whatever
398
+ the harness has written since — a refreshed token, an edited setting — and leaves it alone. Seeded files are
399
+ created `0600`, parent directories `0700`, and the step never creates `$HOME` itself. The claude and script
400
+ families declare an empty `seed`, so the step is a no-op for them.
401
+
402
+ ### Whether a conversation resumes
403
+
404
+ Durable storage makes state survive a recycle; whether a family *resumes* a conversation also needs its ACP
405
+ adapter to implement `session/load`. All three model-backed adapters do, verified against the shipped
406
+ artifacts:
407
+
408
+ - **`claude-agent-acp@0.70.0`** — advertises `agentCapabilities.loadSession: true` and implements
409
+ `loadSession(params)`, registered on the connection as the `session/load` request handler
410
+ (`dist/acp-agent.js:788`, `:6909`).
411
+ - **`codex-acp@1.11.0`** — implements `session/load` by resuming the codex thread and replaying that
412
+ thread's history as `session/update` notifications before it answers the request. Read in the adapter's
413
+ built code at the pinned version; the `0.16.0` adapter it replaced carried the same method inside its own
414
+ native binary.
415
+ - **`pi-acp@0.0.33`** — advertises `agentCapabilities.loadSession: true` and implements `loadSession(params)`,
416
+ which respawns pi against the stored session file and replays its messages (`dist/index.js:1935`, `:2460`;
417
+ the helper that binds a session object to an existing pi process is at `:753`). Confirmed on the wire: an
418
+ ACP `initialize` against the shipped binary answers `protocolVersion: 1` with
419
+ `agentCapabilities.loadSession: true`. That is the *capability*, not the replay itself — pi still has no
420
+ captured `session/load` transcript, which is the one resume-evidence gap open on this lane.
421
+
422
+ ## The command's environment is the family's pass-through list, and nothing else
423
+
424
+ The worker does not *merge* an environment into a harness process: the family's `passEnv` enumeration in
425
+ `facts/families.json` **is** that process's complete environment. An environment cannot declare an
426
+ `SFDK_` name for it, because `SFDK_` is reserved apart from the four `SFDK_WORKER_*` release settings, which
427
+ reach the worker and never a harness process. A derived image that bakes a name the list does not carry, such as `ENV JOB_TARGET=…`, gets **silence**,
428
+ not an override: the variable never reaches the command, and a job running under `set -u` aborts on the
429
+ first reference to it. Per-job configuration travels the prompt, the agent's instructions, or a file in the
430
+ image.
431
+
432
+ Every family's harness process also gets `SFDK_SESSION_ID`, which the worker sets from the `sessionId` on
433
+ `session_placement.apply`: the worker adds it itself, so it needs no `passEnv` entry, whatever the worker's
434
+ own environment holds. The `script` family adds only the organization to the command it runs, as
435
+ `ORCHESTRA_ORGANIZATION_ID` and `SFDK_ORGANIZATION_ID`, from its stored conversation metadata; its command inherits
436
+ the session variable from the harness process. Neither grants egress authority.
437
+
438
+ **On Windows that list is shorter than the operating system's own habits assume.** It is
439
+ OS-independent, so a Windows job sees `PATH`, `HOME` and `USERPROFILE` and *not* `TEMP`, `APPDATA` or
440
+ `ComSpec` — names a PowerShell one-liner reaches for without thinking (`$env:TEMP`). The list is not a
441
+ filter over the environment: the worker *builds* the child's environment from it, so a name the platform
442
+ set on the service and the family does not declare is simply gone by the time the command runs.
443
+
444
+ **Three names are backfilled for Windows and none of them is on any family's list.** A backfill is the
445
+ treatment for operating-system *plumbing* — a variable the machine needs to work at all, as opposed to
446
+ configuration a customer chose — and it is if-absent, so a declared spelling always wins. Two come from
447
+ libraries: Go's `os/exec` adds `SYSTEMROOT` whenever a caller sets an explicit environment, and libuv copies
448
+ its own required-variable set into a spawned child from *this* process's environment — which is why `PATH`
449
+ reaching the command is what lets `powershell.exe` and `taskkill.exe` resolve at all. The third is ours:
450
+ **`PATHEXT`**, the list of extensions PowerShell treats as *executable*. With it unset the effective list
451
+ collapses to `.CPL`, so a `curl.exe` PowerShell resolved on `PATH` is a **document** rather than a program
452
+ and invoking it throws `CantActivateDocumentInPipeline` — before a process starts, before a packet leaves.
453
+ Every network probe on the Windows lane came back `000`, which reads exactly like a blocked connection.
454
+
455
+ The neighbours a reader expects beside it were **measured and left out**, each changing nothing on a Windows
456
+ Server 2025 runtime AMI: `cmd.exe` and `.bat` files run from CreateProcess's own default without `ComSpec`,
457
+ `%SystemRoot%` expands from the `SYSTEMROOT` `os/exec` already backfills, `SystemDrive` and `windir` are
458
+ unused by realistic work, and a service's temp directory falls back to `C:\Windows\SystemTemp` without
459
+ `TEMP` or `TMP`. Add a fourth name only with that kind of evidence. Machine plumbing is not configuration
460
+ and must never be pushed into the prompt: a job cannot set `PATHEXT` for a shell that has already refused to
461
+ start its first program.
462
+
463
+ Three more bounds a job author should know:
464
+
465
+ - **Turn wall clock.** The platform bounds a turn with its own response timeout (default 600 seconds,
466
+ configurable). A job that can run longer needs that raised, or it is killed mid-turn and reported as a
467
+ runtime failure regardless of what the script was doing.
468
+ - **`session/load` re-executes.** The script harness holds no conversational state, so a resumed
469
+ conversation's command runs again. That is the safe direction — a fresh, deterministic run — and another
470
+ reason to keep commands idempotent.
471
+ - **`cwd` is the session's.** `session/new` (and `session/load` on a revival) carries the workspace path, the
472
+ worker maps it to this machine's native filesystem — `C:\sf\workspace` on Windows — and the harness spawns
473
+ the command there. A session that declares none inherits the harness's own directory; one that does not
474
+ exist fails the spawn, which the turn reports as an ordinary exit-127 failure with the operating system's
475
+ message attached.
476
+
477
+ ## What the `script` harness answers
478
+
479
+ The worker runs one harness process per session placement and speaks ACP to it as its client. The `script`
480
+ family's harness is the one the worker carries itself (`script-harness/`), and its dispatch table is:
481
+
482
+ - `initialize` — the harness identifies itself.
483
+ - `session/new` — a fresh session id, with the session's declared working directory anchored on it.
484
+ - `session/load` — the session's replay, as the `session/update` notifications emitted between the request
485
+ and its response. It executes nothing and calls nothing, and it restores the duplicate-execution guard
486
+ below, which the notifications cannot carry.
487
+ - `session/prompt` — execute the agent's instructions, stream stdout and stderr as `session/update`
488
+ notifications, move platform state from the exit code, then answer. Streaming is bounded twice: each frame
489
+ is sliced at 32000 characters (one over-long line would trip the line scanner and kill every session on
490
+ the pipe, not just this turn), and a turn's **total** streamed volume is capped at 5 MiB of **wire** bytes
491
+ — the JSON-encoded frame text, since that is what the platform records, and JSON escaping inflates
492
+ control-heavy output up to about six times (one NUL byte is six wire bytes). Past the ceiling the harness
493
+ emits one truncation notice and stops streaming, while the command still runs to completion with its
494
+ exit-code contract unchanged.
495
+ - `session/cancel` — stop the in-flight command **and everything it started**, then answer. On Linux that is
496
+ a `SIGTERM` to the command's whole **process group** (commands run detached, so the group dies with the
497
+ command and not just with the `bash -c` parent); on Windows there are no process groups to signal, so it
498
+ is `taskkill /T /F /PID <pid>`, which walks the process tree. A cancelled turn made no determination, so
499
+ it claims nothing and answers the prompt with a JSON-RPC error. Whether a turn counts as cancelled is read
500
+ off an observed **fact about the kill**, never off the arrival of the cancel, and each operating system
501
+ supplies the fact it can: on Linux the command's **exit shape** (one the signal actually killed reports
502
+ `signal = SIGTERM`, one that had already exited reports its own status), and on Windows — which terminates
503
+ a process with an exit *code* and no signal — the kill's own answer, so the harness checks whether the
504
+ shell was still running when the cancel arrived and only then kills. Either way a cancel racing a script's
505
+ own `exit 0`, including the window after the shell exits but before the harness reaps it, leaves the turn
506
+ with the result it earned, claim included. (That liveness check is also a safety requirement on Windows,
507
+ which recycles process ids: a `/T` on a stale one would kill a stranger's process tree.)
508
+
509
+ A turn that settles while a backgrounded grandchild still holds the pipes signals that process group and
510
+ detaches the streams, so nothing leaks into the next turn's transcript. That sweep is **conditional and
511
+ best-effort**, and a job author should not read it as a guarantee that background work dies with the turn:
512
+ it runs only on the grace-timer path, so a background job whose stdio is redirected (`cmd >/dev/null 2>&1 &`)
513
+ lets the pipes close promptly, settles the turn before the timer, and survives it untouched. It is also
514
+ `SIGTERM`-only — a process that traps or ignores `TERM` survives either way. Nothing of the sort leaks into
515
+ the *transcript*; what survives is the process.
516
+
517
+ A second prompt on a **workflow-step** session whose earlier claim is confirmed landed is answered as a
518
+ no-op turn without re-executing the command (warm reuse and steer promotion both make that routine); an
519
+ earlier claim that is merely *attempted* is re-confirmed against the run first. The no-op is step-scoped on
520
+ purpose: an ordinary message appended to a direct-start session is a real new turn and always executes.
package/package.json CHANGED
@@ -1,8 +1,8 @@
1
1
  {
2
2
  "name": "@orchestraworks/worker",
3
- "version": "0.0.134",
3
+ "version": "0.0.136",
4
4
  "type": "module",
5
- "description": "The Software Factory worker: the one program that runs inside a machine and speaks the worker protocol to the platform.",
5
+ "description": "The Orchestra worker.",
6
6
  "license": "Apache-2.0",
7
7
  "author": "Living Computers",
8
8
  "homepage": "https://orchestraworks.ai",
@@ -18,13 +18,14 @@
18
18
  },
19
19
  "files": [
20
20
  "bin",
21
- "README.md"
21
+ "README.md",
22
+ "docs/REFERENCE.md"
22
23
  ],
23
24
  "optionalDependencies": {
24
- "@orchestraworks/worker-linux-x64": "0.0.134",
25
- "@orchestraworks/worker-linux-arm64": "0.0.134",
26
- "@orchestraworks/worker-darwin-x64": "0.0.134",
27
- "@orchestraworks/worker-darwin-arm64": "0.0.134",
28
- "@orchestraworks/worker-win32-x64": "0.0.134"
25
+ "@orchestraworks/worker-linux-x64": "0.0.136",
26
+ "@orchestraworks/worker-linux-arm64": "0.0.136",
27
+ "@orchestraworks/worker-darwin-x64": "0.0.136",
28
+ "@orchestraworks/worker-darwin-arm64": "0.0.136",
29
+ "@orchestraworks/worker-win32-x64": "0.0.136"
29
30
  }
30
31
  }