@orchestraworks/worker 0.0.134 → 0.0.136
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +22 -501
- package/docs/REFERENCE.md +520 -0
- package/package.json +9 -8
package/README.md
CHANGED
|
@@ -1,518 +1,39 @@
|
|
|
1
|
-
#
|
|
1
|
+
# @orchestraworks/worker
|
|
2
2
|
|
|
3
|
-
|
|
4
|
-
platform. This tree is that program, and it is also the `@orchestraworks/worker` npm package: the launcher on
|
|
5
|
-
the path, and one per-platform package carrying the binary for the machine it is installed on.
|
|
3
|
+
The worker is the one program that runs on each Orchestra machine: it connects the machine to the platform and runs your agent's harness.
|
|
6
4
|
|
|
7
|
-
|
|
5
|
+
## Install
|
|
8
6
|
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
| path | what it holds |
|
|
12
|
-
|---|---|
|
|
13
|
-
| `cmd/worker/` | the one binary: `run`, `install`, `enroll`, `check`, `version` |
|
|
14
|
-
| `identity/` | the four identity exchanges — claim, refresh, reattach-challenge, reattach — the holder key, and the loop that keeps the machine token alive |
|
|
15
|
-
| `broker/` | the protocol on the message broker: the connection, the register on every connect, the heartbeat and the status cadence, the pull of this machine's own work consumer, and the publisher every other package sends through |
|
|
16
|
-
| `env/` | the machine-side environment the platform composes onto a launch, read from the names the contract declares |
|
|
17
|
-
| `facts/` | what this build knows about the four harness families: the program to install and the digest that authenticates it, the environment each family is handed, the templates it seeds, and every location its content lands in — embedded in the binary, so one release digest covers the facts and the program that reads them |
|
|
18
|
-
| `harness/` | the only place family-specific knowledge lives: the launch plan, the content layout, the store lookup, the mid-turn message each family accepts, and the installer that puts the pinned Node and a family's programs under the worker root |
|
|
19
|
-
| `process/` | the agent's setup and the harness process behind each session placement: the start, the stop and the last save a remove asks for, the startup fence over what a previous run left, and the stop of everything when the worker is told to stop |
|
|
20
|
-
| `stream/` | the queue: every placement message and every prompt, steer and cancel pulled off this machine's work consumer, the ACP client that opens a session's harness conversation and runs its turns, each turn's events batched to the platform and published live, and the answer every placement message is owed |
|
|
21
|
-
| `mount/` | the volumes the launch declares, placed on the lanes that mount them: the home volume first, because anything written under a mount point before the mount is hidden by it |
|
|
22
|
-
| `network/` | the transparent redirect on the virtual-machine lanes, as the iptables rules the launch describes: what the worker's own uid reaches directly, and what every uid on the machine reaches through the egress edge |
|
|
23
|
-
| `shim/` | the in-guest transparent proxy: the listener the redirect lands on, the recovery of the destination each connection was actually for, the hop to the egress edge under the machine-identity proof, the WinDivert half that is the Windows redirect, and the loopback feed a separate shim reads this machine's token from |
|
|
24
|
-
| `egress/` | what puts that proxy in place on the virtual-machine lanes: the edge's certificate authority in this machine's trust store, the inbound firewall rule Windows needs, the redirect itself, and the state line an operator reads off the machine's console |
|
|
25
|
-
| `ingress/` | the machine's half of the preview tunnel: the connection this worker holds open to the preview gateway, the challenge and the signed upgrade that authenticate it, and the loopback pipe that carries each viewer request to a port this machine declared and refuses every other |
|
|
26
|
-
| `termination/` | the cloud's notice that this machine is being reclaimed, delivered to the broker so the next status report carries it |
|
|
27
|
-
| `paths/` | which native place a logical path denotes, and when two logical paths denote one place — pinned against the platform's own answer by committed vectors |
|
|
28
|
-
| `service/` | `worker run` under this operating system's service manager, in one of two scopes: a system or user `systemd` unit, a Windows service or logon task, a launchd daemon or launch agent |
|
|
29
|
-
| `file/` | everything that moves bytes onto this machine and back off it: the install of the skills, memory, instructions and model settings an agent placement's apply carries, against the install journal, the capture that offers the machine's declared roots back to the platform, the restore that lays a previous machine's capture down before a manifest is installed over it, and the self-update that replaces this worker's own release package |
|
|
30
|
-
| `script-harness/` | the first-party `script` harness program, which is the one family the worker package carries rather than installs |
|
|
31
|
-
| `bin/worker.js` | the npm launcher: it runs the binary out of the per-platform package installed for this machine, stays in front of it, and forwards the signal a service manager ends it with. Its own tests sit beside it and are not packed |
|
|
32
|
-
| `npm/<platform>/` | the five per-platform packages, one per release target: the binary, and the `script` harness copied in beside it at build time |
|
|
33
|
-
|
|
34
|
-
`../fwp` is the protocol. Every wire shape comes from `github.com/livingcomputers/fwp/go/wire`
|
|
35
|
-
through a relative `replace`, so the program and the conformance suite compile against the exact
|
|
36
|
-
shapes the schema ships, and nothing here re-declares one.
|
|
37
|
-
|
|
38
|
-
## The commands
|
|
39
|
-
|
|
40
|
-
- **`worker run`** claims this machine's identity — or reattaches, where its row is already bound to
|
|
41
|
-
the key this machine holds — and then serves the machine. `run` hands the held identity, with the
|
|
42
|
-
broker credential the claim co-issued, to the registrar seam in `cmd/worker/run.go`, which the
|
|
43
|
-
`broker` package fills: it dials, registers the machine on every connect, pulls the machine's own
|
|
44
|
-
work consumer and reports on the cadence.
|
|
45
|
-
|
|
46
|
-
Before any of that, `run` starts what the launch describes: it places the declared volumes, puts
|
|
47
|
-
the in-guest egress in place where the launch names an intercept set, opens the egress shim's token
|
|
48
|
-
feed, and — on a machine the cloud can reclaim — watches for the notice that it is going away. Each
|
|
49
|
-
is started only where it is described: a Kubernetes pod places its own volumes and its egress
|
|
50
|
-
sidecar installs its own redirect, so a launch there names neither. If one of those duties is
|
|
51
|
-
installed and then dies, `run` ends rather than serving a machine the platform counts as healthy
|
|
52
|
-
with no route off it; the service manager starts the worker again.
|
|
53
|
-
|
|
54
|
-
`--service` is how the process is hosted rather than what it does: on Windows it hands the body to
|
|
55
|
-
the service control manager, which kills a service that does not answer its handshake, and on the
|
|
56
|
-
other operating systems it is refused by name, because systemd and launchd supervise a foreground
|
|
57
|
-
process from outside.
|
|
58
|
-
|
|
59
|
-
On a machine that hosts a harness, `run` also builds the three lanes that answer the platform's
|
|
60
|
-
verbs: `process`, which holds the placements, `stream`, which carries their turns, and `file`,
|
|
61
|
-
which installs what an agent is made of and captures what its work left behind. The three are
|
|
62
|
-
built before the broker, because the broker's configuration carries their handlers, and each
|
|
63
|
-
reaches it afterwards through the deferred link in `cmd/worker/serve.go`.
|
|
64
|
-
|
|
65
|
-
Once the register has been answered, `run` starts the capture cadence and, where the platform named
|
|
66
|
-
a version other than the one this process is running, the self-update: the release package is
|
|
67
|
-
fetched from the npm registry when this machine can reach one and from the platform when it cannot,
|
|
68
|
-
checked against the digest the register answer named, unpacked over the release package directory,
|
|
69
|
-
and the process restarted — but only once no turn has been open for five seconds together, because one
|
|
70
|
-
open turn per machine is the platform's rule and killing one would lose its answer. A machine whose
|
|
71
|
-
turns never leave that gap keeps running the version it has until some later start picks the unpacked
|
|
72
|
-
one up.
|
|
73
|
-
- **`worker install --role worker|image|laptop --harness <families> --yes [--root <dir>] [--service=user|none]`**
|
|
74
|
-
installs the worker on this machine and starts its service. Every answer comes from a flag: a missing
|
|
75
|
-
one is an error and never a default. It installs the harness half itself — the pinned Node, then each
|
|
76
|
-
named family's programs from the facts file's own pins, with the tarball checked against the pinned
|
|
77
|
-
digest before anything is unpacked — and then registers the service. The per-operating-system service
|
|
78
|
-
forms fill the installer seam in `cmd/worker/install.go`; a build with none refuses rather than
|
|
79
|
-
reporting an install it did not perform.
|
|
80
|
-
|
|
81
|
-
**`--service` says which session the worker is registered in, and neither of its two values needs any
|
|
82
|
-
privilege.** `user` registers it in the invoking account's own session: a launch agent under
|
|
83
|
-
`~/Library/LaunchAgents` on macOS, a `systemd --user` unit under `~/.config/systemd/user` on Linux, a
|
|
84
|
-
scheduled task the account's logon starts on Windows. `none` installs the files and registers nothing,
|
|
85
|
-
for the one host that has no service manager — a container image, whose own command is the worker —
|
|
86
|
-
and `false` is that same answer under its older spelling. The default is `user`. An install run as
|
|
87
|
-
root is refused, because it would put the worker's root under root's home and run every harness as
|
|
88
|
-
root; there is no privileged install to redirect it to, and `--service=system` is refused with the
|
|
89
|
-
same answer. Nothing on a machine somebody owns needs privilege: network enforcement is outside the
|
|
90
|
-
machine, at the egress edge and the directional security group.
|
|
91
|
-
|
|
92
|
-
**The machine-wide service is the image role's, and the role is the whole of what asks for it.** An
|
|
93
|
-
`image` gets it — a property list under `/Library/LaunchDaemons`, a unit under `/etc/systemd/system`,
|
|
94
|
-
a Windows service — because the platform builds the image with nobody logged in to it and nobody
|
|
95
|
-
typing the command, and the builder is already root. No flag selects that scope, on any role.
|
|
96
|
-
|
|
97
|
-
`--root` follows the scope the same way: `~/.sfdk/worker` (`%LOCALAPPDATA%\sfdk\worker` on Windows) for
|
|
98
|
-
a `user` install, `/var/lib/sfdk/worker` (`%ProgramData%\orchestra\worker`) for the other two. A `user`
|
|
99
|
-
install writes its log under that root, at `log/worker.log`, because no service manager in that scope
|
|
100
|
-
keeps one a person can read.
|
|
101
|
-
|
|
102
|
-
What the role decides is when the service runs — a `worker` starts now and at every boot, an `image`
|
|
103
|
-
being baked starts at the first boot of a machine made from it and claims no identity now, and a
|
|
104
|
-
`laptop` starts when a person enrolls a machine and says so. An `image` install also refuses, before
|
|
105
|
-
it installs anything, when a bootstrap file already sits at any path the run would read: an image is
|
|
106
|
-
copied onto every machine made from it, so a baked bootstrap would hand one machine's identity to all
|
|
107
|
-
of them.
|
|
108
|
-
|
|
109
|
-
Each form carries one thing its machine would otherwise be wrong without. A macOS launch DAEMON — the
|
|
110
|
-
image role's form, which a person never reaches — names the account that ran the install, because
|
|
111
|
-
launchd's system domain runs a daemon as root otherwise, and a harness runs as whoever this worker is;
|
|
112
|
-
the worker's own root and the daemon's log file are handed to that account in the same step, so the
|
|
113
|
-
service can read the owner-only bootstrap the root install just wrote, write its state file beside it,
|
|
114
|
-
and say in its log why it did not start. A launch AGENT names no account: it is
|
|
115
|
-
already the session's. The Windows service is registered to restart itself three times at two seconds,
|
|
116
|
-
on a clean non-zero exit as well as a crash — the control manager restarts nothing it was not told to,
|
|
117
|
-
and this worker exits non-zero deliberately to be started again on a release its self-update put in
|
|
118
|
-
place — and the logon task carries the task scheduler's nearest equivalent, three restarts a minute
|
|
119
|
-
apart.
|
|
120
|
-
|
|
121
|
-
A `worker` install run as root on Linux refuses outright where the `agent` account a harness is
|
|
122
|
-
dropped to is not on the machine, and names the command that creates it: a root worker runs its
|
|
123
|
-
harness as root without that account, so this machine would register with no harness capability,
|
|
124
|
-
report healthy, and refuse every conversation sent to it, having reported an install that succeeded.
|
|
125
|
-
Creating a system account on a machine somebody else owns is the operator's to do, so the install
|
|
126
|
-
names it rather than making it. One `worker` install can still be root's — `--service=none` on a
|
|
127
|
-
machine whose own command is the worker — and every other one is refused as root before it gets this
|
|
128
|
-
far. An `image` carries the account already, and an install that registers a service in the invoking
|
|
129
|
-
account's session has no second account at all: the harness runs as the person who installed it.
|
|
130
|
-
|
|
131
|
-
A family a placement names and this machine lacks is installed on the preparation path instead,
|
|
132
|
-
before its harness starts, so a first provision is not slowed by a download of hundreds of megabytes
|
|
133
|
-
that nobody asked for. The root is where both paths install. It is deliberately not the release
|
|
134
|
-
package's own directory, which a self-update replaces whole. Under it sit the pinned Node, each
|
|
135
|
-
family's programs, and `workspaces/<family>`, the directory that family's harness runs in — one
|
|
136
|
-
create at install covers all three, and a worker enrolled on somebody's own machine can take no
|
|
137
|
-
second one.
|
|
138
|
-
- **`worker enroll`** redeems the enrollment token this machine was given, prints the row it was
|
|
139
|
-
issued, and then serves that machine. It does not stop at the redemption, because the holder key
|
|
140
|
-
the redemption bound lives in this process.
|
|
141
|
-
- **`worker check [--harness <families>] [--root <dir>]`** reads the runtime image contract from
|
|
142
|
-
inside an image, which is where the rest of that contract cannot be read. An inspection from the
|
|
143
|
-
container runtime sees the command, the account it starts as, the platform it was built for and the
|
|
144
|
-
version label; whether a family's programs are there at the versions this build's facts pin, whether
|
|
145
|
-
the Orchestra CLI is on the path that family's harness is handed, whether a Linux image has the `agent`
|
|
146
|
-
account at uid 1000 with home `/home` and glibc's loader, and whether a Windows image has
|
|
147
|
-
`WinDivert.dll` and `WinDivert64.sys` under `<root>\bin`, where the redirect loads them, are only
|
|
148
|
-
visible from within. Each miss names the line of
|
|
149
|
-
`agents/runtime/image/Dockerfile` that fixes it, and every miss is reported in one run, because the
|
|
150
|
-
rows are independent and a person rebuilding an image should read the whole list once.
|
|
151
|
-
|
|
152
|
-
It starts nothing, claims no identity and speaks to no platform, so it needs neither a bootstrap nor
|
|
153
|
-
a daemon: `docker run --rm --entrypoint worker <image> check --harness <family>` is the whole of how
|
|
154
|
-
it is used. With no `--harness` it checks the rows that belong to no family and says so — a customer
|
|
155
|
-
image installs the families it needs, so holding one to every family this build can host would
|
|
156
|
-
refuse it for lacking programs nothing on it was ever going to run.
|
|
157
|
-
- **`worker trust-nested [--registry <host>]... [--root <dir>]`** lets the containers this machine
|
|
158
|
-
runs trust the egress edge, which nothing does by default. It reads the authority from
|
|
159
|
-
`SFDK_EGRESS_CA_FILE`, which the worker sets on every harness process and on the setup script;
|
|
160
|
-
writes it into dockerd's `certs.d` for the well-known registries and each one named; writes
|
|
161
|
-
`k3d-registries.yaml` under the root, whose `ca_file` is the authority's path inside each k3d node;
|
|
162
|
-
and prints the flags `docker run`, compose and `k3d cluster create` need. On a machine that names no
|
|
163
|
-
authority, such as one with no edge in front of it, it says so, writes nothing and succeeds.
|
|
164
|
-
- **`worker version`** prints this binary's version, the release package it belongs to, and the
|
|
165
|
-
protocol versions it speaks.
|
|
166
|
-
|
|
167
|
-
The version is stamped at link time by the publishing lane, and both `package.json` files carry
|
|
168
|
-
`0.0.0` until that lane stamps them; a build that stamps nothing says so rather than claiming a
|
|
169
|
-
version it is not.
|
|
170
|
-
|
|
171
|
-
## The bootstrap a machine is started with
|
|
172
|
-
|
|
173
|
-
Everything above assumes the worker already knows where its platform is. That is the one family of
|
|
174
|
-
names it reads before it can speak to anything, so they are the worker's own rather than the
|
|
175
|
-
contract's: the environment document owns the other `SFDK_` names the platform or the worker sets for a
|
|
176
|
-
running machine and says these are outside its scope. They carry the same reserved prefix, so an environment
|
|
177
|
-
cannot declare one; the only `SFDK_` names an environment may declare are the release settings
|
|
178
|
-
`fwp/schema/environment-v1.json` lists under `customerDeclarable`, which the worker reads from its own
|
|
179
|
-
environment in preference to the register answer. The platform composes what is listed here; a value it does not
|
|
180
|
-
set is a concern this machine does not have.
|
|
181
|
-
|
|
182
|
-
| name | what it carries |
|
|
183
|
-
|---|---|
|
|
184
|
-
| `SFDK_CLAIM_ENDPOINT_URL` | the platform's origin, not a route: the paths belong to the protocol |
|
|
185
|
-
| `SFDK_TRANSPORT_TRUST_ROOTS` | the certificate bundle this machine trusts, inline |
|
|
186
|
-
| `SFDK_TRANSPORT_TRUST_ROOTS_FILE` | the same bundle as an absolute path, for a delivery that cannot carry the newlines; setting both is refused |
|
|
187
|
-
| `SFDK_CLAIM_LOCATOR` | the machine row this worker claims as |
|
|
188
|
-
| `SFDK_ENROLLMENT_TOKEN` | the show-once token an enrolling machine redeems; mutually exclusive with the locator |
|
|
189
|
-
| `SFDK_ENROLLMENT_TOKEN_FILE` | the same token as an absolute path; setting both is refused |
|
|
190
|
-
| `SFDK_HOLDER_KEY_ALGORITHM` | which of the protocol's two holder-key algorithms this registration declares |
|
|
191
|
-
| `SFDK_SUBSTRATE` | which evidence this worker gathers, so an unknown value is refused rather than defaulted |
|
|
192
|
-
| `SFDK_CLAIM_AUDIENCE` | who the presented caller-identity request was meant for; required on the virtual-machine lanes |
|
|
193
|
-
| `SFDK_STS_ENDPOINT_URL` | the security-token endpoint the platform will send that request to |
|
|
194
|
-
| `SFDK_SERVICE_ACCOUNT_TOKEN_FILE` | where the audience-bound projected token is mounted, on the Kubernetes lane |
|
|
195
|
-
| `SFDK_TPM_DEVICE_PATH` | the chip's device path where it is not this operating system's default |
|
|
196
|
-
| `SFDK_SUBSTRATE_TOKEN` | the `local` substrate's stand-in factor, inline |
|
|
197
|
-
| `SFDK_SUBSTRATE_TOKEN_FILE` | the same factor as an absolute path; setting both is refused |
|
|
198
|
-
| `SFDK_INTERCEPT_PORTS` | the destination ports whose traffic goes to the egress edge, comma separated; naming none installs no redirect |
|
|
199
|
-
| `SFDK_INTERCEPT_BYPASS` | the shared services the worker's own uid reaches directly, as `[{"host": …, "port": …}]` — our broker and our gateway, and nothing broader |
|
|
200
|
-
| `SFDK_INTERCEPT_SKIP_RANGES` | the private ranges the redirect never takes, for any uid on any port, as comma-separated CIDR ranges; a range that is not wholly private is refused |
|
|
201
|
-
| `SFDK_EDGE_ENDPOINT_URL` | where the redirect sends this machine's traffic; its scheme is the whole TLS statement for that hop |
|
|
202
|
-
| `SFDK_EGRESS_LISTEN_PORT` | where the redirect lands, which is the shim's own listen port |
|
|
203
|
-
| `SFDK_EGRESS_IPV6` | `true` mirrors the redirect through `ip6tables` |
|
|
204
|
-
| `SFDK_EGRESS_PROXY` | `on`, the default, or `off`; `off` tells each harness that no edge of ours stands in front of this machine, as on an enrolled one. The platform writes `off` into every Launcher machine's boot file, because no proxy entrance of ours stands in front of one yet |
|
|
205
|
-
| `SFDK_RUNTIME_ENVIRONMENT` | a whole environment folded into one JSON object of string values, for a boot file whose lines cannot carry a line break; each pair is placed unless already set |
|
|
206
|
-
| `SFDK_VOLUME_MOUNTS` | the volumes this launch declares, as one JSON document |
|
|
207
|
-
| `SFDK_MACHINE_SLUG` | the machine's platform-minted slug: the leading label of its preview hostnames, and what the tunnel's upgrade signs |
|
|
208
|
-
| `SFDK_TUNNEL_ENDPOINT_URL` | where the preview gateway serves its tunnel listener; its scheme is the whole TLS statement for that hop |
|
|
209
|
-
| `SFDK_INGRESS_PORTS` | the loopback ports this machine projects through that tunnel, comma separated; with no reserved ports below, naming none holds no tunnel |
|
|
210
|
-
| `SFDK_INGRESS_ANY_PORT_EXCEPT` | the reserved loopback ports, comma separated; set, the tunnel is held from start, forwards every other port whatever `SFDK_INGRESS_PORTS` says, and the worker registers `ingress-any-port` |
|
|
211
|
-
|
|
212
|
-
**The worker reads that file itself, on every operating system, and no service manager is handed it.**
|
|
213
|
-
Every `worker run` places each pair into its own environment before it serves, leaving alone any name
|
|
214
|
-
that is already set, so a person's own foreground run still wins over the file. systemd's
|
|
215
|
-
`EnvironmentFile=` is deliberately not used: it processes backslash escapes in an unquoted value, so
|
|
216
|
-
the two characters `\n` that carry a line break inside `SFDK_RUNTIME_ENVIRONMENT` would arrive as a bare
|
|
217
|
-
`n` and the certificate folded in there would hold no certificate at all. A machine with no file
|
|
218
|
-
starts anyway — that is an enrolled machine's ordinary start — and a line that is not `NAME=VALUE`
|
|
219
|
-
stops the start and is named.
|
|
220
|
-
|
|
221
|
-
Where the file sits is the one thing that differs, and only because of who may write it. On a Linux
|
|
222
|
-
machine the platform launched, cloud-init writes `/etc/sfdk/machine.env` from the user data and the run
|
|
223
|
-
reads it; no install by a person writes there. A `worker` install that registers a service writes
|
|
224
|
-
`<root>/machine.env` under the worker's own root — `C:\ProgramData\sfdk\worker\machine.env` on a default
|
|
225
|
-
machine-wide Windows install — because an install that registers the worker in one person's own session
|
|
226
|
-
cannot write under `/etc`. A Linux run reads both, the machine-wide one first. The platform's user data writes that file on a machine
|
|
227
|
-
it launched, and a person enrolling one writes it or exports the same names in the shell that runs
|
|
228
|
-
`worker enroll`.
|
|
229
|
-
|
|
230
|
-
Beside them the worker keeps one file of its own, `<root>/state/machine-identity.json`, owner-only.
|
|
231
|
-
It records the row this worker is bound to and — on a machine with no chip — the software holder key
|
|
232
|
-
it is bound under, and it decides whether the next start claims or reattaches. A row this
|
|
233
|
-
worker has already claimed cannot be claimed again, so a start that finds its own record takes the
|
|
234
|
-
next generation under the key instead, and it treats the enrollment token its bootstrap file still
|
|
235
|
-
carries as spent rather than refusing to start beside the row it names. The record also carries the
|
|
236
|
-
last run that registered with the platform and the epoch it registered under, which is what an
|
|
237
|
-
install waits for: everything else in it is written at the claim, before there is a broker link at
|
|
238
|
-
all. It is kept only where the row outlives the program: a virtual
|
|
239
|
-
machine that parks and wakes with its own disk, and a machine a person enrolled. A pod is replaced
|
|
240
|
-
rather than woken, and its row is fresh with it.
|
|
241
|
-
|
|
242
|
-
## The in-guest egress
|
|
243
|
-
|
|
244
|
-
On the virtual-machine lanes this worker's own process is the transparent proxy, and it puts three
|
|
245
|
-
things in place in one order that is not interchangeable:
|
|
246
|
-
|
|
247
|
-
1. **The edge's certificate authority into this machine's trust store**, from the file the launch
|
|
248
|
-
wrote (or, where that file is not there, from the URL it names). Windows reads the machine
|
|
249
|
-
certificate stores and nothing else, so a certificate that is not there is every TLS client on
|
|
250
|
-
the machine failing its handshake with the edge; and on Windows a chain verdict computed before
|
|
251
|
-
the import stays cached machine-wide, so an import that actually wrote also throws that cache
|
|
252
|
-
away. It happens first because the redirect is live the moment it lands.
|
|
253
|
-
2. **The listener**, bound before any rule exists, so no connection is ever sent to a port with
|
|
254
|
-
nothing behind it. On Windows an inbound firewall rule is written first as well: the rewritten
|
|
255
|
-
packet arrives as an inbound connection, which Windows Server blocks by default.
|
|
256
|
-
3. **The redirect** — `iptables` on Ubuntu, WinDivert on Windows. The two WinDivert files are loaded
|
|
257
|
-
by absolute path from `<root>\bin`, where the image bake stages them, and never through the
|
|
258
|
-
operating system's own search order: opening the driver installs and starts a kernel service.
|
|
259
|
-
|
|
260
|
-
Each accepted connection's original destination is recovered (`SO_ORIGINAL_DST` on Linux, the
|
|
261
|
-
redirect's own proxy-port map on Windows; a recovery that misses is refused rather than sent to the
|
|
262
|
-
peer, which carries the right host and the wrong port) and carried to the edge under the machine
|
|
263
|
-
token this worker holds. There is no loopback round trip on these lanes: the process that holds the
|
|
264
|
-
token is the process that presents it.
|
|
265
|
-
|
|
266
|
-
What the worker's own uid reaches **without** the redirect is the launch's list — our broker and our
|
|
267
|
-
gateway — plus three the launch cannot state: the egress edge, because redirecting the hop that
|
|
268
|
-
carries a redirected connection is a loop with no exit; the platform's identity endpoint, because
|
|
269
|
-
this worker calls it before it holds the token the edge would identify it by; and the metadata
|
|
270
|
-
service, which answers the evidence that call presents. Nothing is uid-wide, so the object store
|
|
271
|
-
stays intercepted and this worker's own content pulls meet the platform's rule at the edge.
|
|
272
|
-
|
|
273
|
-
The redirect takes its Windows form on Windows and its `iptables` form everywhere else, whatever the
|
|
274
|
-
substrate. A pod with an egress sidecar runs only the certificate step: the sidecar installs that
|
|
275
|
-
lane's redirect and runs its own shim, which reads the token from the loopback feed below. A pod with
|
|
276
|
-
no sidecar, as a customer's Launcher creates, names an intercept set and this worker installs the
|
|
277
|
-
redirect itself.
|
|
278
|
-
|
|
279
|
-
The private ranges `SFDK_INTERCEPT_SKIP_RANGES` names are never redirected, for any uid on any port:
|
|
280
|
-
the rules return them ahead of the redirect, a nested container reaches them through its bridge, and
|
|
281
|
-
the Windows filter excludes each range beside the single destinations.
|
|
282
|
-
|
|
283
|
-
The proxy names a customer declares (`HTTPS_PROXY`, `HTTP_PROXY` and `NO_PROXY`, in either case) serve
|
|
284
|
-
this worker's own calls to the platform and the broker. `NO_PROXY` must then hold `169.254.169.254`,
|
|
285
|
-
because the metadata service is never reached through a proxy. Where a redirect is in place, the same
|
|
286
|
-
names are left out of the environment each harness and the setup script are handed, so the agent's
|
|
287
|
-
calls reach the redirect, which carries them to the edge.
|
|
288
|
-
|
|
289
|
-
## The egress shim's token feed
|
|
290
|
-
|
|
291
|
-
The shim that intercepts this machine's traffic stamps the machine token on it, and the worker is
|
|
292
|
-
what holds that token. On Kubernetes the two are in different containers of one pod, so the only
|
|
293
|
-
channel between them is the network: the worker listens on `127.0.0.1:7788` — a constant both halves
|
|
294
|
-
carry, not a placed value — and the shim connects with retry.
|
|
295
|
-
|
|
296
|
-
Each frame is one JSON line, newline-terminated:
|
|
7
|
+
The Orchestra CLI installs this package for you. To install it alone:
|
|
297
8
|
|
|
298
|
-
```
|
|
299
|
-
|
|
9
|
+
```bash
|
|
10
|
+
npm install -g @orchestraworks/worker
|
|
300
11
|
```
|
|
301
12
|
|
|
302
|
-
|
|
303
|
-
that. It writes nothing else on that socket and never reads from it, and it listens on loopback and
|
|
304
|
-
nowhere else. A worker restart drops the connection; the shim reconnects and is fed the current
|
|
305
|
-
generation again. The broker credential co-issued beside the token never crosses this channel, and
|
|
306
|
-
neither does the holder key.
|
|
13
|
+
## Quick start
|
|
307
14
|
|
|
308
|
-
|
|
15
|
+
To put one of your own machines to work, use the CLI. It signs you in, asks what this machine is, and installs the worker:
|
|
309
16
|
|
|
17
|
+
```bash
|
|
18
|
+
npm install -g @orchestraworks/cli
|
|
19
|
+
orchestra install
|
|
310
20
|
```
|
|
311
|
-
make verify # gofmt, then vet for linux, windows and macOS
|
|
312
|
-
make vet-all # the vet half on its own, which is what type-checks the per-operating-system files
|
|
313
|
-
make test # the script harness's own tests, then go test with the coverage gate
|
|
314
|
-
make build # one static binary per release package, into npm/<platform>/, with the script harness beside it
|
|
315
|
-
```
|
|
316
|
-
|
|
317
|
-
### What a release is
|
|
318
|
-
|
|
319
|
-
`make build` writes the five per-platform binaries into their package directories and copies the
|
|
320
|
-
`script` harness in beside each, and those directories are what the publishing workflow packs. Two
|
|
321
|
-
shapes of tarball come out of it, and the layout is stated here because that workflow reads it
|
|
322
|
-
rather than deciding it:
|
|
323
|
-
|
|
324
|
-
```
|
|
325
|
-
@orchestraworks/worker-<platform> @orchestraworks/worker
|
|
326
|
-
package.json package.json
|
|
327
|
-
worker (worker.exe on win32) bin/worker.js
|
|
328
|
-
script-harness/ README.md
|
|
329
|
-
package.json LICENSE
|
|
330
|
-
index.js
|
|
331
|
-
fetch-failure.js
|
|
332
|
-
LICENSE
|
|
333
|
-
```
|
|
334
|
-
|
|
335
|
-
Six published packages redistribute this tree, and an Apache-2.0 package that ships without the licence
|
|
336
|
-
is redistributed on terms it does not carry — so `LICENSE` is one file here and a copy in every tarball.
|
|
337
|
-
|
|
338
|
-
The per-platform tarball is the release artifact: its SHA-256 is the one digest the catalog carries,
|
|
339
|
-
the register answer names and both fetch paths check, and it covers the binary and the `script`
|
|
340
|
-
harness together, so a self-update carries the program with the binary. The launcher package holds
|
|
341
|
-
no binary at all — it depends on the five as optional dependencies and runs whichever one npm
|
|
342
|
-
installed for this machine. The build's output is ignored by `.gitignore`: the packages are
|
|
343
|
-
published from a build rather than from the tree.
|
|
344
21
|
|
|
345
|
-
|
|
346
|
-
database, no stack, and no harness family is ever downloaded — a fake registry serves small archives
|
|
347
|
-
that prove the install's own rules. The tests that
|
|
348
|
-
need a real trusted platform module skip where there is none, and the broker suite drives a real
|
|
349
|
-
`nats-server` named by `WORKER_TEST_NATS_SERVER` and skips when that is unset — so a `make test`
|
|
350
|
-
meant to meet the coverage gate sets it:
|
|
22
|
+
To check that a machine image can host a harness, run `worker check` inside it:
|
|
351
23
|
|
|
24
|
+
```bash
|
|
25
|
+
docker run --rm --entrypoint worker <image> check --harness claude
|
|
352
26
|
```
|
|
353
|
-
WORKER_TEST_NATS_SERVER=/path/to/nats-server-2.14.6 make test
|
|
354
|
-
```
|
|
355
|
-
|
|
356
|
-
`WORKER_TEST_NATS_SERVER` is a knob a test process sets for itself, not a variable the platform sets
|
|
357
|
-
on a machine, which is why it carries no `SFDK_` prefix and appears in no environment document.
|
|
358
|
-
|
|
359
|
-
## Where a family keeps its state
|
|
360
|
-
|
|
361
|
-
Every family keeps its conversation history — the transcripts that let a later session pick up where an
|
|
362
|
-
earlier one stopped — in a directory it derives from `$HOME`. That is true of all four, and nothing in the
|
|
363
|
-
runtime image baits any of them elsewhere:
|
|
364
|
-
|
|
365
|
-
| family | state directory | how it resolves |
|
|
366
|
-
| --- | --- | --- |
|
|
367
|
-
| claude | `$HOME/.claude/` | Node `os.homedir()` — `$HOME` first, `getpwuid` fallback. |
|
|
368
|
-
| codex | `$HOME/.codex/` | The Rust `home` crate, used when `CODEX_HOME` is unset. |
|
|
369
|
-
| pi | `$HOME/.pi/agent/` (transcripts under `sessions/`) | `process.env.PI_CODING_AGENT_DIR ? resolve(...) : join(homedir(), ".pi", "agent")` — `pi-acp@0.0.33` `dist/index.js:1381` and `:1725`, and the harness's own `getAgentDir()` at `@earendil-works/pi-coding-agent@0.84.4` `dist/config.js:420-426`, whose variable NAME is assembled at `:405` from `APP_NAME` (`"pi"`). Transcripts: harness `dist/config.js:456-457`, adapter `dist/index.js:1397-1399` — the adapter prefers a `sessionDir` key in `<agentDir>/settings.json` when one is set (`:1383-1396`), and the baked settings file sets none. `pi-acp` also keeps its own session map at `$HOME/.pi/pi-acp/` (`:352`). |
|
|
370
|
-
| script | none | The harness holds no conversational state. |
|
|
371
|
-
|
|
372
|
-
An environment makes that state durable by declaring a volume mount at `/home`, which is where the
|
|
373
|
-
account inside the runtime image already has its home directory. An environment that declares no such
|
|
374
|
-
mount behaves exactly as it did before: the state is written to the machine's own disk and goes with it.
|
|
375
|
-
|
|
376
|
-
That is why no image bakes `CODEX_HOME` or `PI_CODING_AGENT_DIR`. A pointer to a path outside `$HOME` is
|
|
377
|
-
exactly what a durable home cannot survive: it would send codex's and pi's transcripts to the container's
|
|
378
|
-
writable layer, which a recycle discards. Leaving both knobs unset lets each family's own `$HOME` fallback
|
|
379
|
-
engage. Injecting a literal `CODEX_HOME="$HOME/.codex"` instead was rejected because env-value expansion is
|
|
380
|
-
not portable — Kubernetes expands only `$(VAR)` and Docker expands nothing.
|
|
381
|
-
|
|
382
|
-
### First-boot seeding
|
|
383
|
-
|
|
384
|
-
codex and pi each need one file to already exist inside their state directory: codex's egress-placeholder
|
|
385
|
-
`auth.json`, pi's provider-pinning `settings.json`. Both are image content, but the directory they belong
|
|
386
|
-
in lives on a volume that is empty the first time a machine boots. So the image bakes each file as a
|
|
387
|
-
template and the family's `seed` list in `facts/families.json` says where it goes:
|
|
388
|
-
|
|
389
|
-
```
|
|
390
|
-
codex: /opt/software-factory-codex/auth.json -> .codex/auth.json
|
|
391
|
-
pi: /opt/software-factory-pi/settings.json -> .pi/agent/settings.json
|
|
392
|
-
```
|
|
393
|
-
|
|
394
|
-
The worker copies each pair before it spawns the harness, resolving every destination against `$HOME`. The
|
|
395
|
-
copy is **copy-if-absent**: the first placement lays the templates down, and every later one finds whatever
|
|
396
|
-
the harness has written since — a refreshed token, an edited setting — and leaves it alone. Seeded files are
|
|
397
|
-
created `0600`, parent directories `0700`, and the step never creates `$HOME` itself. The claude and script
|
|
398
|
-
families declare an empty `seed`, so the step is a no-op for them.
|
|
399
|
-
|
|
400
|
-
### Whether a conversation resumes
|
|
401
|
-
|
|
402
|
-
Durable storage makes state survive a recycle; whether a family *resumes* a conversation also needs its ACP
|
|
403
|
-
adapter to implement `session/load`. All three model-backed adapters do, verified against the shipped
|
|
404
|
-
artifacts:
|
|
405
|
-
|
|
406
|
-
- **`claude-agent-acp@0.70.0`** — advertises `agentCapabilities.loadSession: true` and implements
|
|
407
|
-
`loadSession(params)`, registered on the connection as the `session/load` request handler
|
|
408
|
-
(`dist/acp-agent.js:788`, `:6909`).
|
|
409
|
-
- **`codex-acp@1.11.0`** — implements `session/load` by resuming the codex thread and replaying that
|
|
410
|
-
thread's history as `session/update` notifications before it answers the request. Read in the adapter's
|
|
411
|
-
built code at the pinned version; the `0.16.0` adapter it replaced carried the same method inside its own
|
|
412
|
-
native binary.
|
|
413
|
-
- **`pi-acp@0.0.33`** — advertises `agentCapabilities.loadSession: true` and implements `loadSession(params)`,
|
|
414
|
-
which respawns pi against the stored session file and replays its messages (`dist/index.js:1935`, `:2460`;
|
|
415
|
-
the helper that binds a session object to an existing pi process is at `:753`). Confirmed on the wire: an
|
|
416
|
-
ACP `initialize` against the shipped binary answers `protocolVersion: 1` with
|
|
417
|
-
`agentCapabilities.loadSession: true`. That is the *capability*, not the replay itself — pi still has no
|
|
418
|
-
captured `session/load` transcript, which is the one resume-evidence gap open on this lane.
|
|
419
|
-
|
|
420
|
-
## The command's environment is the family's pass-through list, and nothing else
|
|
421
|
-
|
|
422
|
-
The worker does not *merge* an environment into a harness process: the family's `passEnv` enumeration in
|
|
423
|
-
`facts/families.json` **is** that process's complete environment. An environment cannot declare an
|
|
424
|
-
`SFDK_` name for it, because `SFDK_` is reserved apart from the four `SFDK_WORKER_*` release settings, which
|
|
425
|
-
reach the worker and never a harness process. A derived image that bakes a name the list does not carry, such as `ENV JOB_TARGET=…`, gets **silence**,
|
|
426
|
-
not an override: the variable never reaches the command, and a job running under `set -u` aborts on the
|
|
427
|
-
first reference to it. Per-job configuration travels the prompt, the agent's instructions, or a file in the
|
|
428
|
-
image.
|
|
429
|
-
|
|
430
|
-
Every family's harness process also gets `SFDK_SESSION_ID`, which the worker sets from the `sessionId` on
|
|
431
|
-
`session_placement.apply`: the worker adds it itself, so it needs no `passEnv` entry, whatever the worker's
|
|
432
|
-
own environment holds. The `script` family adds only
|
|
433
|
-
`SFDK_ORGANIZATION_ID` to the command it runs, from its stored conversation metadata; its command inherits
|
|
434
|
-
the session variable from the harness process. Neither grants egress authority.
|
|
435
|
-
|
|
436
|
-
**On Windows that list is shorter than the operating system's own habits assume.** It is
|
|
437
|
-
OS-independent, so a Windows job sees `PATH`, `HOME` and `USERPROFILE` and *not* `TEMP`, `APPDATA` or
|
|
438
|
-
`ComSpec` — names a PowerShell one-liner reaches for without thinking (`$env:TEMP`). The list is not a
|
|
439
|
-
filter over the environment: the worker *builds* the child's environment from it, so a name the platform
|
|
440
|
-
set on the service and the family does not declare is simply gone by the time the command runs.
|
|
441
|
-
|
|
442
|
-
**Three names are backfilled for Windows and none of them is on any family's list.** A backfill is the
|
|
443
|
-
treatment for operating-system *plumbing* — a variable the machine needs to work at all, as opposed to
|
|
444
|
-
configuration a customer chose — and it is if-absent, so a declared spelling always wins. Two come from
|
|
445
|
-
libraries: Go's `os/exec` adds `SYSTEMROOT` whenever a caller sets an explicit environment, and libuv copies
|
|
446
|
-
its own required-variable set into a spawned child from *this* process's environment — which is why `PATH`
|
|
447
|
-
reaching the command is what lets `powershell.exe` and `taskkill.exe` resolve at all. The third is ours:
|
|
448
|
-
**`PATHEXT`**, the list of extensions PowerShell treats as *executable*. With it unset the effective list
|
|
449
|
-
collapses to `.CPL`, so a `curl.exe` PowerShell resolved on `PATH` is a **document** rather than a program
|
|
450
|
-
and invoking it throws `CantActivateDocumentInPipeline` — before a process starts, before a packet leaves.
|
|
451
|
-
Every network probe on the Windows lane came back `000`, which reads exactly like a blocked connection.
|
|
452
|
-
|
|
453
|
-
The neighbours a reader expects beside it were **measured and left out**, each changing nothing on a Windows
|
|
454
|
-
Server 2025 runtime AMI: `cmd.exe` and `.bat` files run from CreateProcess's own default without `ComSpec`,
|
|
455
|
-
`%SystemRoot%` expands from the `SYSTEMROOT` `os/exec` already backfills, `SystemDrive` and `windir` are
|
|
456
|
-
unused by realistic work, and a service's temp directory falls back to `C:\Windows\SystemTemp` without
|
|
457
|
-
`TEMP` or `TMP`. Add a fourth name only with that kind of evidence. Machine plumbing is not configuration
|
|
458
|
-
and must never be pushed into the prompt: a job cannot set `PATHEXT` for a shell that has already refused to
|
|
459
|
-
start its first program.
|
|
460
|
-
|
|
461
|
-
Three more bounds a job author should know:
|
|
462
|
-
|
|
463
|
-
- **Turn wall clock.** The platform bounds a turn with its own response timeout (default 600 seconds,
|
|
464
|
-
configurable). A job that can run longer needs that raised, or it is killed mid-turn and reported as a
|
|
465
|
-
runtime failure regardless of what the script was doing.
|
|
466
|
-
- **`session/load` re-executes.** The script harness holds no conversational state, so a resumed
|
|
467
|
-
conversation's command runs again. That is the safe direction — a fresh, deterministic run — and another
|
|
468
|
-
reason to keep commands idempotent.
|
|
469
|
-
- **`cwd` is the session's.** `session/new` (and `session/load` on a revival) carries the workspace path, the
|
|
470
|
-
worker maps it to this machine's native filesystem — `C:\sf\workspace` on Windows — and the harness spawns
|
|
471
|
-
the command there. A session that declares none inherits the harness's own directory; one that does not
|
|
472
|
-
exist fails the spawn, which the turn reports as an ordinary exit-127 failure with the operating system's
|
|
473
|
-
message attached.
|
|
474
27
|
|
|
475
|
-
|
|
28
|
+
`worker version` prints the version and the protocol versions this build speaks.
|
|
476
29
|
|
|
477
|
-
|
|
478
|
-
family's harness is the one the worker carries itself (`script-harness/`), and its dispatch table is:
|
|
30
|
+
## For coding agents
|
|
479
31
|
|
|
480
|
-
- `
|
|
481
|
-
- `
|
|
482
|
-
- `
|
|
483
|
-
|
|
484
|
-
below, which the notifications cannot carry.
|
|
485
|
-
- `session/prompt` — execute the agent's instructions, stream stdout and stderr as `session/update`
|
|
486
|
-
notifications, move platform state from the exit code, then answer. Streaming is bounded twice: each frame
|
|
487
|
-
is sliced at 32000 characters (one over-long line would trip the line scanner and kill every session on
|
|
488
|
-
the pipe, not just this turn), and a turn's **total** streamed volume is capped at 5 MiB of **wire** bytes
|
|
489
|
-
— the JSON-encoded frame text, since that is what the platform records, and JSON escaping inflates
|
|
490
|
-
control-heavy output up to about six times (one NUL byte is six wire bytes). Past the ceiling the harness
|
|
491
|
-
emits one truncation notice and stops streaming, while the command still runs to completion with its
|
|
492
|
-
exit-code contract unchanged.
|
|
493
|
-
- `session/cancel` — stop the in-flight command **and everything it started**, then answer. On Linux that is
|
|
494
|
-
a `SIGTERM` to the command's whole **process group** (commands run detached, so the group dies with the
|
|
495
|
-
command and not just with the `bash -c` parent); on Windows there are no process groups to signal, so it
|
|
496
|
-
is `taskkill /T /F /PID <pid>`, which walks the process tree. A cancelled turn made no determination, so
|
|
497
|
-
it claims nothing and answers the prompt with a JSON-RPC error. Whether a turn counts as cancelled is read
|
|
498
|
-
off an observed **fact about the kill**, never off the arrival of the cancel, and each operating system
|
|
499
|
-
supplies the fact it can: on Linux the command's **exit shape** (one the signal actually killed reports
|
|
500
|
-
`signal = SIGTERM`, one that had already exited reports its own status), and on Windows — which terminates
|
|
501
|
-
a process with an exit *code* and no signal — the kill's own answer, so the harness checks whether the
|
|
502
|
-
shell was still running when the cancel arrived and only then kills. Either way a cancel racing a script's
|
|
503
|
-
own `exit 0`, including the window after the shell exits but before the harness reaps it, leaves the turn
|
|
504
|
-
with the result it earned, claim included. (That liveness check is also a safety requirement on Windows,
|
|
505
|
-
which recycles process ids: a `/T` on a stale one would kill a stranger's process tree.)
|
|
32
|
+
- `orchestra worker --help` lists the worker commands. `orchestra worker <command> --help` gives their flags.
|
|
33
|
+
- `worker` with no command prints its usage.
|
|
34
|
+
- `skills/orchestra/SKILL.md` in `@orchestraworks/cli` explains machines and what the platform can do.
|
|
35
|
+
- [docs/REFERENCE.md](docs/REFERENCE.md): every command, the start-up settings, egress, and how a release is built.
|
|
506
36
|
|
|
507
|
-
|
|
508
|
-
detaches the streams, so nothing leaks into the next turn's transcript. That sweep is **conditional and
|
|
509
|
-
best-effort**, and a job author should not read it as a guarantee that background work dies with the turn:
|
|
510
|
-
it runs only on the grace-timer path, so a background job whose stdio is redirected (`cmd >/dev/null 2>&1 &`)
|
|
511
|
-
lets the pipes close promptly, settles the turn before the timer, and survives it untouched. It is also
|
|
512
|
-
`SIGTERM`-only — a process that traps or ignores `TERM` survives either way. Nothing of the sort leaks into
|
|
513
|
-
the *transcript*; what survives is the process.
|
|
37
|
+
## License
|
|
514
38
|
|
|
515
|
-
|
|
516
|
-
no-op turn without re-executing the command (warm reuse and steer promotion both make that routine); an
|
|
517
|
-
earlier claim that is merely *attempted* is re-confirmed against the run first. The no-op is step-scoped on
|
|
518
|
-
purpose: an ordinary message appended to a direct-start session is a real new turn and always executes.
|
|
39
|
+
Apache License 2.0. The full text is in [LICENSE](LICENSE).
|
|
@@ -0,0 +1,520 @@
|
|
|
1
|
+
# The worker: the full reference
|
|
2
|
+
|
|
3
|
+
The short introduction is the package [README](../README.md). Paths below are relative to the `worker/` folder.
|
|
4
|
+
|
|
5
|
+
One Go program runs inside every machine and speaks the Factory Worker Protocol to the
|
|
6
|
+
platform. This tree is that program, and it is also the `@orchestraworks/worker` npm package: the launcher on
|
|
7
|
+
the path, and one per-platform package carrying the binary for the machine it is installed on.
|
|
8
|
+
|
|
9
|
+
The worker exposes no API of its own. Every route it calls is the platform's.
|
|
10
|
+
|
|
11
|
+
## The tree
|
|
12
|
+
|
|
13
|
+
| path | what it holds |
|
|
14
|
+
|---|---|
|
|
15
|
+
| `cmd/worker/` | the one binary: `run`, `install`, `enroll`, `check`, `version` |
|
|
16
|
+
| `identity/` | the four identity exchanges — claim, refresh, reattach-challenge, reattach — the holder key, and the loop that keeps the machine token alive |
|
|
17
|
+
| `broker/` | the protocol on the message broker: the connection, the register on every connect, the heartbeat and the status cadence, the pull of this machine's own work consumer, and the publisher every other package sends through |
|
|
18
|
+
| `env/` | the machine-side environment the platform composes onto a launch, read from the names the contract declares |
|
|
19
|
+
| `facts/` | what this build knows about the four harness families: the program to install and the digest that authenticates it, the environment each family is handed, the templates it seeds, and every location its content lands in — embedded in the binary, so one release digest covers the facts and the program that reads them |
|
|
20
|
+
| `harness/` | the only place family-specific knowledge lives: the launch plan, the content layout, the store lookup, the mid-turn message each family accepts, and the installer that puts the pinned Node and a family's programs under the worker root |
|
|
21
|
+
| `process/` | the agent's setup and the harness process behind each session placement: the start, the stop and the last save a remove asks for, the startup fence over what a previous run left, and the stop of everything when the worker is told to stop |
|
|
22
|
+
| `stream/` | the queue: every placement message and every prompt, steer and cancel pulled off this machine's work consumer, the ACP client that opens a session's harness conversation and runs its turns, each turn's events batched to the platform and published live, and the answer every placement message is owed |
|
|
23
|
+
| `mount/` | the volumes the launch declares, placed on the lanes that mount them: the home volume first, because anything written under a mount point before the mount is hidden by it |
|
|
24
|
+
| `network/` | the transparent redirect on the virtual-machine lanes, as the iptables rules the launch describes: what the worker's own uid reaches directly, and what every uid on the machine reaches through the egress edge |
|
|
25
|
+
| `shim/` | the in-guest transparent proxy: the listener the redirect lands on, the recovery of the destination each connection was actually for, the hop to the egress edge under the machine-identity proof, the WinDivert half that is the Windows redirect, and the loopback feed a separate shim reads this machine's token from |
|
|
26
|
+
| `egress/` | what puts that proxy in place on the virtual-machine lanes: the edge's certificate authority in this machine's trust store, the inbound firewall rule Windows needs, the redirect itself, and the state line an operator reads off the machine's console |
|
|
27
|
+
| `ingress/` | the machine's half of the preview tunnel: the connection this worker holds open to the preview gateway, the challenge and the signed upgrade that authenticate it, and the loopback pipe that carries each viewer request to a port this machine declared and refuses every other |
|
|
28
|
+
| `termination/` | the cloud's notice that this machine is being reclaimed, delivered to the broker so the next status report carries it |
|
|
29
|
+
| `paths/` | which native place a logical path denotes, and when two logical paths denote one place — pinned against the platform's own answer by committed vectors |
|
|
30
|
+
| `service/` | `worker run` under this operating system's service manager, in one of two scopes: a system or user `systemd` unit, a Windows service or logon task, a launchd daemon or launch agent |
|
|
31
|
+
| `file/` | everything that moves bytes onto this machine and back off it: the install of the skills, memory, instructions and model settings an agent placement's apply carries, against the install journal, the capture that offers the machine's declared roots back to the platform, the restore that lays a previous machine's capture down before a manifest is installed over it, and the self-update that replaces this worker's own release package |
|
|
32
|
+
| `script-harness/` | the first-party `script` harness program, which is the one family the worker package carries rather than installs |
|
|
33
|
+
| `bin/worker.js` | the npm launcher: it runs the binary out of the per-platform package installed for this machine, stays in front of it, and forwards the signal a service manager ends it with. Its own tests sit beside it and are not packed |
|
|
34
|
+
| `npm/<platform>/` | the five per-platform packages, one per release target: the binary, and the `script` harness copied in beside it at build time |
|
|
35
|
+
|
|
36
|
+
`../fwp` is the protocol. Every wire shape comes from `github.com/livingcomputers/fwp/go/wire`
|
|
37
|
+
through a relative `replace`, so the program and the conformance suite compile against the exact
|
|
38
|
+
shapes the schema ships, and nothing here re-declares one.
|
|
39
|
+
|
|
40
|
+
## The commands
|
|
41
|
+
|
|
42
|
+
- **`worker run`** claims this machine's identity — or reattaches, where its row is already bound to
|
|
43
|
+
the key this machine holds — and then serves the machine. `run` hands the held identity, with the
|
|
44
|
+
broker credential the claim co-issued, to the registrar seam in `cmd/worker/run.go`, which the
|
|
45
|
+
`broker` package fills: it dials, registers the machine on every connect, pulls the machine's own
|
|
46
|
+
work consumer and reports on the cadence.
|
|
47
|
+
|
|
48
|
+
Before any of that, `run` starts what the launch describes: it places the declared volumes, puts
|
|
49
|
+
the in-guest egress in place where the launch names an intercept set, opens the egress shim's token
|
|
50
|
+
feed, and — on a machine the cloud can reclaim — watches for the notice that it is going away. Each
|
|
51
|
+
is started only where it is described: a Kubernetes pod places its own volumes and its egress
|
|
52
|
+
sidecar installs its own redirect, so a launch there names neither. If one of those duties is
|
|
53
|
+
installed and then dies, `run` ends rather than serving a machine the platform counts as healthy
|
|
54
|
+
with no route off it; the service manager starts the worker again.
|
|
55
|
+
|
|
56
|
+
`--service` is how the process is hosted rather than what it does: on Windows it hands the body to
|
|
57
|
+
the service control manager, which kills a service that does not answer its handshake, and on the
|
|
58
|
+
other operating systems it is refused by name, because systemd and launchd supervise a foreground
|
|
59
|
+
process from outside.
|
|
60
|
+
|
|
61
|
+
On a machine that hosts a harness, `run` also builds the three lanes that answer the platform's
|
|
62
|
+
verbs: `process`, which holds the placements, `stream`, which carries their turns, and `file`,
|
|
63
|
+
which installs what an agent is made of and captures what its work left behind. The three are
|
|
64
|
+
built before the broker, because the broker's configuration carries their handlers, and each
|
|
65
|
+
reaches it afterwards through the deferred link in `cmd/worker/serve.go`.
|
|
66
|
+
|
|
67
|
+
Once the register has been answered, `run` starts the capture cadence and, where the platform named
|
|
68
|
+
a version other than the one this process is running, the self-update: the release package is
|
|
69
|
+
fetched from the npm registry when this machine can reach one and from the platform when it cannot,
|
|
70
|
+
checked against the digest the register answer named, unpacked over the release package directory,
|
|
71
|
+
and the process restarted — but only once no turn has been open for five seconds together, because one
|
|
72
|
+
open turn per machine is the platform's rule and killing one would lose its answer. A machine whose
|
|
73
|
+
turns never leave that gap keeps running the version it has until some later start picks the unpacked
|
|
74
|
+
one up.
|
|
75
|
+
- **`worker install --role worker|image|laptop --harness <families> --yes [--root <dir>] [--service=user|none]`**
|
|
76
|
+
installs the worker on this machine and starts its service. Every answer comes from a flag: a missing
|
|
77
|
+
one is an error and never a default. It installs the harness half itself — the pinned Node, then each
|
|
78
|
+
named family's programs from the facts file's own pins, with the tarball checked against the pinned
|
|
79
|
+
digest before anything is unpacked — and then registers the service. The per-operating-system service
|
|
80
|
+
forms fill the installer seam in `cmd/worker/install.go`; a build with none refuses rather than
|
|
81
|
+
reporting an install it did not perform.
|
|
82
|
+
|
|
83
|
+
**`--service` says which session the worker is registered in, and neither of its two values needs any
|
|
84
|
+
privilege.** `user` registers it in the invoking account's own session: a launch agent under
|
|
85
|
+
`~/Library/LaunchAgents` on macOS, a `systemd --user` unit under `~/.config/systemd/user` on Linux, a
|
|
86
|
+
scheduled task the account's logon starts on Windows. `none` installs the files and registers nothing,
|
|
87
|
+
for the one host that has no service manager — a container image, whose own command is the worker —
|
|
88
|
+
and `false` is that same answer under its older spelling. The default is `user`. An install run as
|
|
89
|
+
root is refused, because it would put the worker's root under root's home and run every harness as
|
|
90
|
+
root; there is no privileged install to redirect it to, and `--service=system` is refused with the
|
|
91
|
+
same answer. Nothing on a machine somebody owns needs privilege: network enforcement is outside the
|
|
92
|
+
machine, at the egress edge and the directional security group.
|
|
93
|
+
|
|
94
|
+
**The machine-wide service is the image role's, and the role is the whole of what asks for it.** An
|
|
95
|
+
`image` gets it — a property list under `/Library/LaunchDaemons`, a unit under `/etc/systemd/system`,
|
|
96
|
+
a Windows service — because the platform builds the image with nobody logged in to it and nobody
|
|
97
|
+
typing the command, and the builder is already root. No flag selects that scope, on any role.
|
|
98
|
+
|
|
99
|
+
`--root` follows the scope the same way: `~/.sfdk/worker` (`%LOCALAPPDATA%\sfdk\worker` on Windows) for
|
|
100
|
+
a `user` install, `/var/lib/sfdk/worker` (`%ProgramData%\orchestra\worker`) for the other two. A `user`
|
|
101
|
+
install writes its log under that root, at `log/worker.log`, because no service manager in that scope
|
|
102
|
+
keeps one a person can read.
|
|
103
|
+
|
|
104
|
+
What the role decides is when the service runs — a `worker` starts now and at every boot, an `image`
|
|
105
|
+
being baked starts at the first boot of a machine made from it and claims no identity now, and a
|
|
106
|
+
`laptop` starts when a person enrolls a machine and says so. An `image` install also refuses, before
|
|
107
|
+
it installs anything, when a bootstrap file already sits at any path the run would read: an image is
|
|
108
|
+
copied onto every machine made from it, so a baked bootstrap would hand one machine's identity to all
|
|
109
|
+
of them.
|
|
110
|
+
|
|
111
|
+
Each form carries one thing its machine would otherwise be wrong without. A macOS launch DAEMON — the
|
|
112
|
+
image role's form, which a person never reaches — names the account that ran the install, because
|
|
113
|
+
launchd's system domain runs a daemon as root otherwise, and a harness runs as whoever this worker is;
|
|
114
|
+
the worker's own root and the daemon's log file are handed to that account in the same step, so the
|
|
115
|
+
service can read the owner-only bootstrap the root install just wrote, write its state file beside it,
|
|
116
|
+
and say in its log why it did not start. A launch AGENT names no account: it is
|
|
117
|
+
already the session's. The Windows service is registered to restart itself three times at two seconds,
|
|
118
|
+
on a clean non-zero exit as well as a crash — the control manager restarts nothing it was not told to,
|
|
119
|
+
and this worker exits non-zero deliberately to be started again on a release its self-update put in
|
|
120
|
+
place — and the logon task carries the task scheduler's nearest equivalent, three restarts a minute
|
|
121
|
+
apart.
|
|
122
|
+
|
|
123
|
+
A `worker` install run as root on Linux refuses outright where the `agent` account a harness is
|
|
124
|
+
dropped to is not on the machine, and names the command that creates it: a root worker runs its
|
|
125
|
+
harness as root without that account, so this machine would register with no harness capability,
|
|
126
|
+
report healthy, and refuse every conversation sent to it, having reported an install that succeeded.
|
|
127
|
+
Creating a system account on a machine somebody else owns is the operator's to do, so the install
|
|
128
|
+
names it rather than making it. One `worker` install can still be root's — `--service=none` on a
|
|
129
|
+
machine whose own command is the worker — and every other one is refused as root before it gets this
|
|
130
|
+
far. An `image` carries the account already, and an install that registers a service in the invoking
|
|
131
|
+
account's session has no second account at all: the harness runs as the person who installed it.
|
|
132
|
+
|
|
133
|
+
A family a placement names and this machine lacks is installed on the preparation path instead,
|
|
134
|
+
before its harness starts, so a first provision is not slowed by a download of hundreds of megabytes
|
|
135
|
+
that nobody asked for. The root is where both paths install. It is deliberately not the release
|
|
136
|
+
package's own directory, which a self-update replaces whole. Under it sit the pinned Node, each
|
|
137
|
+
family's programs, and `workspaces/<family>`, the directory that family's harness runs in — one
|
|
138
|
+
create at install covers all three, and a worker enrolled on somebody's own machine can take no
|
|
139
|
+
second one.
|
|
140
|
+
- **`worker enroll`** redeems the enrollment token this machine was given, prints the row it was
|
|
141
|
+
issued, and then serves that machine. It does not stop at the redemption, because the holder key
|
|
142
|
+
the redemption bound lives in this process.
|
|
143
|
+
- **`worker check [--harness <families>] [--root <dir>]`** reads the runtime image contract from
|
|
144
|
+
inside an image, which is where the rest of that contract cannot be read. An inspection from the
|
|
145
|
+
container runtime sees the command, the account it starts as, the platform it was built for and the
|
|
146
|
+
version label; whether a family's programs are there at the versions this build's facts pin, whether
|
|
147
|
+
the Orchestra CLI is on the path that family's harness is handed, whether a Linux image has the `agent`
|
|
148
|
+
account at uid 1000 with home `/home` and glibc's loader, and whether a Windows image has
|
|
149
|
+
`WinDivert.dll` and `WinDivert64.sys` under `<root>\bin`, where the redirect loads them, are only
|
|
150
|
+
visible from within. Each miss names the line of
|
|
151
|
+
`agents/runtime/image/Dockerfile` that fixes it, and every miss is reported in one run, because the
|
|
152
|
+
rows are independent and a person rebuilding an image should read the whole list once.
|
|
153
|
+
|
|
154
|
+
It starts nothing, claims no identity and speaks to no platform, so it needs neither a bootstrap nor
|
|
155
|
+
a daemon: `docker run --rm --entrypoint worker <image> check --harness <family>` is the whole of how
|
|
156
|
+
it is used. With no `--harness` it checks the rows that belong to no family and says so — a customer
|
|
157
|
+
image installs the families it needs, so holding one to every family this build can host would
|
|
158
|
+
refuse it for lacking programs nothing on it was ever going to run.
|
|
159
|
+
- **`worker trust-nested [--registry <host>]... [--root <dir>]`** lets the containers this machine
|
|
160
|
+
runs trust the egress edge, which nothing does by default. It reads the authority from
|
|
161
|
+
`SFDK_EGRESS_CA_FILE`, which the worker sets on every harness process and on the setup script;
|
|
162
|
+
writes it into dockerd's `certs.d` for the well-known registries and each one named; writes
|
|
163
|
+
`k3d-registries.yaml` under the root, whose `ca_file` is the authority's path inside each k3d node;
|
|
164
|
+
and prints the flags `docker run`, compose and `k3d cluster create` need. On a machine that names no
|
|
165
|
+
authority, such as one with no edge in front of it, it says so, writes nothing and succeeds.
|
|
166
|
+
- **`worker version`** prints this binary's version, the release package it belongs to, and the
|
|
167
|
+
protocol versions it speaks.
|
|
168
|
+
|
|
169
|
+
The version is stamped at link time by the publishing lane, and both `package.json` files carry
|
|
170
|
+
`0.0.0` until that lane stamps them; a build that stamps nothing says so rather than claiming a
|
|
171
|
+
version it is not.
|
|
172
|
+
|
|
173
|
+
## The bootstrap a machine is started with
|
|
174
|
+
|
|
175
|
+
Everything above assumes the worker already knows where its platform is. That is the one family of
|
|
176
|
+
names it reads before it can speak to anything, so they are the worker's own rather than the
|
|
177
|
+
contract's: the environment document owns the other `SFDK_` names the platform or the worker sets for a
|
|
178
|
+
running machine and says these are outside its scope. They carry the same reserved prefix, so an environment
|
|
179
|
+
cannot declare one; the only `SFDK_` names an environment may declare are the release settings
|
|
180
|
+
`fwp/schema/environment-v1.json` lists under `customerDeclarable`, which the worker reads from its own
|
|
181
|
+
environment in preference to the register answer. The platform composes what is listed here; a value it does not
|
|
182
|
+
set is a concern this machine does not have.
|
|
183
|
+
|
|
184
|
+
| name | what it carries |
|
|
185
|
+
|---|---|
|
|
186
|
+
| `SFDK_CLAIM_ENDPOINT_URL` | the platform's origin, not a route: the paths belong to the protocol |
|
|
187
|
+
| `SFDK_TRANSPORT_TRUST_ROOTS` | the certificate bundle this machine trusts, inline |
|
|
188
|
+
| `SFDK_TRANSPORT_TRUST_ROOTS_FILE` | the same bundle as an absolute path, for a delivery that cannot carry the newlines; setting both is refused |
|
|
189
|
+
| `SFDK_CLAIM_LOCATOR` | the machine row this worker claims as |
|
|
190
|
+
| `SFDK_ENROLLMENT_TOKEN` | the show-once token an enrolling machine redeems; mutually exclusive with the locator |
|
|
191
|
+
| `SFDK_ENROLLMENT_TOKEN_FILE` | the same token as an absolute path; setting both is refused |
|
|
192
|
+
| `SFDK_HOLDER_KEY_ALGORITHM` | which of the protocol's two holder-key algorithms this registration declares |
|
|
193
|
+
| `SFDK_SUBSTRATE` | which evidence this worker gathers, so an unknown value is refused rather than defaulted |
|
|
194
|
+
| `SFDK_CLAIM_AUDIENCE` | who the presented caller-identity request was meant for; required on the virtual-machine lanes |
|
|
195
|
+
| `SFDK_STS_ENDPOINT_URL` | the security-token endpoint the platform will send that request to |
|
|
196
|
+
| `SFDK_SERVICE_ACCOUNT_TOKEN_FILE` | where the audience-bound projected token is mounted, on the Kubernetes lane |
|
|
197
|
+
| `SFDK_TPM_DEVICE_PATH` | the chip's device path where it is not this operating system's default |
|
|
198
|
+
| `SFDK_SUBSTRATE_TOKEN` | the `local` substrate's stand-in factor, inline |
|
|
199
|
+
| `SFDK_SUBSTRATE_TOKEN_FILE` | the same factor as an absolute path; setting both is refused |
|
|
200
|
+
| `SFDK_INTERCEPT_PORTS` | the destination ports whose traffic goes to the egress edge, comma separated; naming none installs no redirect |
|
|
201
|
+
| `SFDK_INTERCEPT_BYPASS` | the shared services the worker's own uid reaches directly, as `[{"host": …, "port": …}]` — our broker and our gateway, and nothing broader |
|
|
202
|
+
| `SFDK_INTERCEPT_SKIP_RANGES` | the private ranges the redirect never takes, for any uid on any port, as comma-separated CIDR ranges; a range that is not wholly private is refused |
|
|
203
|
+
| `SFDK_EDGE_ENDPOINT_URL` | where the redirect sends this machine's traffic; its scheme is the whole TLS statement for that hop |
|
|
204
|
+
| `SFDK_EGRESS_LISTEN_PORT` | where the redirect lands, which is the shim's own listen port |
|
|
205
|
+
| `SFDK_EGRESS_IPV6` | `true` mirrors the redirect through `ip6tables` |
|
|
206
|
+
| `SFDK_EGRESS_PROXY` | `on`, the default, or `off`; `off` tells each harness that no edge of ours stands in front of this machine, as on an enrolled one. The platform writes `off` into every Launcher machine's boot file, because no proxy entrance of ours stands in front of one yet |
|
|
207
|
+
| `SFDK_RUNTIME_ENVIRONMENT` | a whole environment folded into one JSON object of string values, for a boot file whose lines cannot carry a line break; each pair is placed unless already set |
|
|
208
|
+
| `SFDK_VOLUME_MOUNTS` | the volumes this launch declares, as one JSON document |
|
|
209
|
+
| `SFDK_MACHINE_SLUG` | the machine's platform-minted slug: the leading label of its preview hostnames, and what the tunnel's upgrade signs |
|
|
210
|
+
| `SFDK_TUNNEL_ENDPOINT_URL` | where the preview gateway serves its tunnel listener; its scheme is the whole TLS statement for that hop |
|
|
211
|
+
| `SFDK_INGRESS_PORTS` | the loopback ports this machine projects through that tunnel, comma separated; with no reserved ports below, naming none holds no tunnel |
|
|
212
|
+
| `SFDK_INGRESS_ANY_PORT_EXCEPT` | the reserved loopback ports, comma separated; set, the tunnel is held from start, forwards every other port whatever `SFDK_INGRESS_PORTS` says, and the worker registers `ingress-any-port` |
|
|
213
|
+
|
|
214
|
+
**The worker reads that file itself, on every operating system, and no service manager is handed it.**
|
|
215
|
+
Every `worker run` places each pair into its own environment before it serves, leaving alone any name
|
|
216
|
+
that is already set, so a person's own foreground run still wins over the file. systemd's
|
|
217
|
+
`EnvironmentFile=` is deliberately not used: it processes backslash escapes in an unquoted value, so
|
|
218
|
+
the two characters `\n` that carry a line break inside `SFDK_RUNTIME_ENVIRONMENT` would arrive as a bare
|
|
219
|
+
`n` and the certificate folded in there would hold no certificate at all. A machine with no file
|
|
220
|
+
starts anyway — that is an enrolled machine's ordinary start — and a line that is not `NAME=VALUE`
|
|
221
|
+
stops the start and is named.
|
|
222
|
+
|
|
223
|
+
Where the file sits is the one thing that differs, and only because of who may write it. On a Linux
|
|
224
|
+
machine the platform launched, cloud-init writes `/etc/sfdk/machine.env` from the user data and the run
|
|
225
|
+
reads it; no install by a person writes there. A `worker` install that registers a service writes
|
|
226
|
+
`<root>/machine.env` under the worker's own root — `C:\ProgramData\sfdk\worker\machine.env` on a default
|
|
227
|
+
machine-wide Windows install — because an install that registers the worker in one person's own session
|
|
228
|
+
cannot write under `/etc`. A Linux run reads both, the machine-wide one first. The platform's user data writes that file on a machine
|
|
229
|
+
it launched, and a person enrolling one writes it or exports the same names in the shell that runs
|
|
230
|
+
`worker enroll`.
|
|
231
|
+
|
|
232
|
+
Beside them the worker keeps one file of its own, `<root>/state/machine-identity.json`, owner-only.
|
|
233
|
+
It records the row this worker is bound to and — on a machine with no chip — the software holder key
|
|
234
|
+
it is bound under, and it decides whether the next start claims or reattaches. A row this
|
|
235
|
+
worker has already claimed cannot be claimed again, so a start that finds its own record takes the
|
|
236
|
+
next generation under the key instead, and it treats the enrollment token its bootstrap file still
|
|
237
|
+
carries as spent rather than refusing to start beside the row it names. The record also carries the
|
|
238
|
+
last run that registered with the platform and the epoch it registered under, which is what an
|
|
239
|
+
install waits for: everything else in it is written at the claim, before there is a broker link at
|
|
240
|
+
all. It is kept only where the row outlives the program: a virtual
|
|
241
|
+
machine that parks and wakes with its own disk, and a machine a person enrolled. A pod is replaced
|
|
242
|
+
rather than woken, and its row is fresh with it.
|
|
243
|
+
|
|
244
|
+
## The in-guest egress
|
|
245
|
+
|
|
246
|
+
On the virtual-machine lanes this worker's own process is the transparent proxy, and it puts three
|
|
247
|
+
things in place in one order that is not interchangeable:
|
|
248
|
+
|
|
249
|
+
1. **The edge's certificate authority into this machine's trust store**, from the file the launch
|
|
250
|
+
wrote (or, where that file is not there, from the URL it names). Windows reads the machine
|
|
251
|
+
certificate stores and nothing else, so a certificate that is not there is every TLS client on
|
|
252
|
+
the machine failing its handshake with the edge; and on Windows a chain verdict computed before
|
|
253
|
+
the import stays cached machine-wide, so an import that actually wrote also throws that cache
|
|
254
|
+
away. It happens first because the redirect is live the moment it lands.
|
|
255
|
+
2. **The listener**, bound before any rule exists, so no connection is ever sent to a port with
|
|
256
|
+
nothing behind it. On Windows an inbound firewall rule is written first as well: the rewritten
|
|
257
|
+
packet arrives as an inbound connection, which Windows Server blocks by default.
|
|
258
|
+
3. **The redirect** — `iptables` on Ubuntu, WinDivert on Windows. The two WinDivert files are loaded
|
|
259
|
+
by absolute path from `<root>\bin`, where the image bake stages them, and never through the
|
|
260
|
+
operating system's own search order: opening the driver installs and starts a kernel service.
|
|
261
|
+
|
|
262
|
+
Each accepted connection's original destination is recovered (`SO_ORIGINAL_DST` on Linux, the
|
|
263
|
+
redirect's own proxy-port map on Windows; a recovery that misses is refused rather than sent to the
|
|
264
|
+
peer, which carries the right host and the wrong port) and carried to the edge under the machine
|
|
265
|
+
token this worker holds. There is no loopback round trip on these lanes: the process that holds the
|
|
266
|
+
token is the process that presents it.
|
|
267
|
+
|
|
268
|
+
What the worker's own uid reaches **without** the redirect is the launch's list — our broker and our
|
|
269
|
+
gateway — plus three the launch cannot state: the egress edge, because redirecting the hop that
|
|
270
|
+
carries a redirected connection is a loop with no exit; the platform's identity endpoint, because
|
|
271
|
+
this worker calls it before it holds the token the edge would identify it by; and the metadata
|
|
272
|
+
service, which answers the evidence that call presents. Nothing is uid-wide, so the object store
|
|
273
|
+
stays intercepted and this worker's own content pulls meet the platform's rule at the edge.
|
|
274
|
+
|
|
275
|
+
The redirect takes its Windows form on Windows and its `iptables` form everywhere else, whatever the
|
|
276
|
+
substrate. A pod with an egress sidecar runs only the certificate step: the sidecar installs that
|
|
277
|
+
lane's redirect and runs its own shim, which reads the token from the loopback feed below. A pod with
|
|
278
|
+
no sidecar, as a customer's Launcher creates, names an intercept set and this worker installs the
|
|
279
|
+
redirect itself.
|
|
280
|
+
|
|
281
|
+
The private ranges `SFDK_INTERCEPT_SKIP_RANGES` names are never redirected, for any uid on any port:
|
|
282
|
+
the rules return them ahead of the redirect, a nested container reaches them through its bridge, and
|
|
283
|
+
the Windows filter excludes each range beside the single destinations.
|
|
284
|
+
|
|
285
|
+
The proxy names a customer declares (`HTTPS_PROXY`, `HTTP_PROXY` and `NO_PROXY`, in either case) serve
|
|
286
|
+
this worker's own calls to the platform and the broker. `NO_PROXY` must then hold `169.254.169.254`,
|
|
287
|
+
because the metadata service is never reached through a proxy. Where a redirect is in place, the same
|
|
288
|
+
names are left out of the environment each harness and the setup script are handed, so the agent's
|
|
289
|
+
calls reach the redirect, which carries them to the edge.
|
|
290
|
+
|
|
291
|
+
## The egress shim's token feed
|
|
292
|
+
|
|
293
|
+
The shim that intercepts this machine's traffic stamps the machine token on it, and the worker is
|
|
294
|
+
what holds that token. On Kubernetes the two are in different containers of one pod, so the only
|
|
295
|
+
channel between them is the network: the worker listens on `127.0.0.1:7788` — a constant both halves
|
|
296
|
+
carry, not a placed value — and the shim connects with retry.
|
|
297
|
+
|
|
298
|
+
Each frame is one JSON line, newline-terminated:
|
|
299
|
+
|
|
300
|
+
```
|
|
301
|
+
{"generation": 7, "token": "<the machine token>", "expiresAt": "2026-09-12T10:15:00.123456Z"}
|
|
302
|
+
```
|
|
303
|
+
|
|
304
|
+
On connect the worker writes the generation this machine holds now, and one line per rotation after
|
|
305
|
+
that. It writes nothing else on that socket and never reads from it, and it listens on loopback and
|
|
306
|
+
nowhere else. A worker restart drops the connection; the shim reconnects and is fed the current
|
|
307
|
+
generation again. The broker credential co-issued beside the token never crosses this channel, and
|
|
308
|
+
neither does the holder key.
|
|
309
|
+
|
|
310
|
+
## Building and testing
|
|
311
|
+
|
|
312
|
+
```
|
|
313
|
+
make verify # gofmt, then vet for linux, windows and macOS
|
|
314
|
+
make vet-all # the vet half on its own, which is what type-checks the per-operating-system files
|
|
315
|
+
make test # the script harness's own tests, then go test with the coverage gate
|
|
316
|
+
make build # one static binary per release package, into npm/<platform>/, with the script harness beside it
|
|
317
|
+
```
|
|
318
|
+
|
|
319
|
+
### What a release is
|
|
320
|
+
|
|
321
|
+
`make build` writes the five per-platform binaries into their package directories and copies the
|
|
322
|
+
`script` harness in beside each, and those directories are what the publishing workflow packs. Two
|
|
323
|
+
shapes of tarball come out of it, and the layout is stated here because that workflow reads it
|
|
324
|
+
rather than deciding it:
|
|
325
|
+
|
|
326
|
+
```
|
|
327
|
+
@orchestraworks/worker-<platform> @orchestraworks/worker
|
|
328
|
+
package.json package.json
|
|
329
|
+
worker (worker.exe on win32) bin/worker.js
|
|
330
|
+
script-harness/ README.md
|
|
331
|
+
package.json LICENSE
|
|
332
|
+
index.js
|
|
333
|
+
fetch-failure.js
|
|
334
|
+
LICENSE
|
|
335
|
+
```
|
|
336
|
+
|
|
337
|
+
Six published packages redistribute this tree, and an Apache-2.0 package that ships without the licence
|
|
338
|
+
is redistributed on terms it does not carry — so `LICENSE` is one file here and a copy in every tarball.
|
|
339
|
+
|
|
340
|
+
The per-platform tarball is the release artifact: its SHA-256 is the one digest the catalog carries,
|
|
341
|
+
the register answer names and both fetch paths check, and it covers the binary and the `script`
|
|
342
|
+
harness together, so a self-update carries the program with the binary. The launcher package holds
|
|
343
|
+
no binary at all — it depends on the five as optional dependencies and runs whichever one npm
|
|
344
|
+
installed for this machine. The build's output is ignored by `.gitignore`: the packages are
|
|
345
|
+
published from a build rather than from the tree.
|
|
346
|
+
|
|
347
|
+
The tests are Go tests plus two sets of Node's own — the `script` harness's, and the npm launcher's: no container runtime, no
|
|
348
|
+
database, no stack, and no harness family is ever downloaded — a fake registry serves small archives
|
|
349
|
+
that prove the install's own rules. The tests that
|
|
350
|
+
need a real trusted platform module skip where there is none, and the broker suite drives a real
|
|
351
|
+
`nats-server` named by `WORKER_TEST_NATS_SERVER` and skips when that is unset — so a `make test`
|
|
352
|
+
meant to meet the coverage gate sets it:
|
|
353
|
+
|
|
354
|
+
```
|
|
355
|
+
WORKER_TEST_NATS_SERVER=/path/to/nats-server-2.14.6 make test
|
|
356
|
+
```
|
|
357
|
+
|
|
358
|
+
`WORKER_TEST_NATS_SERVER` is a knob a test process sets for itself, not a variable the platform sets
|
|
359
|
+
on a machine, which is why it carries no `SFDK_` prefix and appears in no environment document.
|
|
360
|
+
|
|
361
|
+
## Where a family keeps its state
|
|
362
|
+
|
|
363
|
+
Every family keeps its conversation history — the transcripts that let a later session pick up where an
|
|
364
|
+
earlier one stopped — in a directory it derives from `$HOME`. That is true of all four, and nothing in the
|
|
365
|
+
runtime image baits any of them elsewhere:
|
|
366
|
+
|
|
367
|
+
| family | state directory | how it resolves |
|
|
368
|
+
| --- | --- | --- |
|
|
369
|
+
| claude | `$HOME/.claude/` | Node `os.homedir()` — `$HOME` first, `getpwuid` fallback. |
|
|
370
|
+
| codex | `$HOME/.codex/` | The Rust `home` crate, used when `CODEX_HOME` is unset. |
|
|
371
|
+
| pi | `$HOME/.pi/agent/` (transcripts under `sessions/`) | `process.env.PI_CODING_AGENT_DIR ? resolve(...) : join(homedir(), ".pi", "agent")` — `pi-acp@0.0.33` `dist/index.js:1381` and `:1725`, and the harness's own `getAgentDir()` at `@earendil-works/pi-coding-agent@0.84.4` `dist/config.js:420-426`, whose variable NAME is assembled at `:405` from `APP_NAME` (`"pi"`). Transcripts: harness `dist/config.js:456-457`, adapter `dist/index.js:1397-1399` — the adapter prefers a `sessionDir` key in `<agentDir>/settings.json` when one is set (`:1383-1396`), and the baked settings file sets none. `pi-acp` also keeps its own session map at `$HOME/.pi/pi-acp/` (`:352`). |
|
|
372
|
+
| script | none | The harness holds no conversational state. |
|
|
373
|
+
|
|
374
|
+
An environment makes that state durable by declaring a volume mount at `/home`, which is where the
|
|
375
|
+
account inside the runtime image already has its home directory. An environment that declares no such
|
|
376
|
+
mount behaves exactly as it did before: the state is written to the machine's own disk and goes with it.
|
|
377
|
+
|
|
378
|
+
That is why no image bakes `CODEX_HOME` or `PI_CODING_AGENT_DIR`. A pointer to a path outside `$HOME` is
|
|
379
|
+
exactly what a durable home cannot survive: it would send codex's and pi's transcripts to the container's
|
|
380
|
+
writable layer, which a recycle discards. Leaving both knobs unset lets each family's own `$HOME` fallback
|
|
381
|
+
engage. Injecting a literal `CODEX_HOME="$HOME/.codex"` instead was rejected because env-value expansion is
|
|
382
|
+
not portable — Kubernetes expands only `$(VAR)` and Docker expands nothing.
|
|
383
|
+
|
|
384
|
+
### First-boot seeding
|
|
385
|
+
|
|
386
|
+
codex and pi each need one file to already exist inside their state directory: codex's egress-placeholder
|
|
387
|
+
`auth.json`, pi's provider-pinning `settings.json`. Both are image content, but the directory they belong
|
|
388
|
+
in lives on a volume that is empty the first time a machine boots. So the image bakes each file as a
|
|
389
|
+
template and the family's `seed` list in `facts/families.json` says where it goes:
|
|
390
|
+
|
|
391
|
+
```
|
|
392
|
+
codex: /opt/software-factory-codex/auth.json -> .codex/auth.json
|
|
393
|
+
pi: /opt/software-factory-pi/settings.json -> .pi/agent/settings.json
|
|
394
|
+
```
|
|
395
|
+
|
|
396
|
+
The worker copies each pair before it spawns the harness, resolving every destination against `$HOME`. The
|
|
397
|
+
copy is **copy-if-absent**: the first placement lays the templates down, and every later one finds whatever
|
|
398
|
+
the harness has written since — a refreshed token, an edited setting — and leaves it alone. Seeded files are
|
|
399
|
+
created `0600`, parent directories `0700`, and the step never creates `$HOME` itself. The claude and script
|
|
400
|
+
families declare an empty `seed`, so the step is a no-op for them.
|
|
401
|
+
|
|
402
|
+
### Whether a conversation resumes
|
|
403
|
+
|
|
404
|
+
Durable storage makes state survive a recycle; whether a family *resumes* a conversation also needs its ACP
|
|
405
|
+
adapter to implement `session/load`. All three model-backed adapters do, verified against the shipped
|
|
406
|
+
artifacts:
|
|
407
|
+
|
|
408
|
+
- **`claude-agent-acp@0.70.0`** — advertises `agentCapabilities.loadSession: true` and implements
|
|
409
|
+
`loadSession(params)`, registered on the connection as the `session/load` request handler
|
|
410
|
+
(`dist/acp-agent.js:788`, `:6909`).
|
|
411
|
+
- **`codex-acp@1.11.0`** — implements `session/load` by resuming the codex thread and replaying that
|
|
412
|
+
thread's history as `session/update` notifications before it answers the request. Read in the adapter's
|
|
413
|
+
built code at the pinned version; the `0.16.0` adapter it replaced carried the same method inside its own
|
|
414
|
+
native binary.
|
|
415
|
+
- **`pi-acp@0.0.33`** — advertises `agentCapabilities.loadSession: true` and implements `loadSession(params)`,
|
|
416
|
+
which respawns pi against the stored session file and replays its messages (`dist/index.js:1935`, `:2460`;
|
|
417
|
+
the helper that binds a session object to an existing pi process is at `:753`). Confirmed on the wire: an
|
|
418
|
+
ACP `initialize` against the shipped binary answers `protocolVersion: 1` with
|
|
419
|
+
`agentCapabilities.loadSession: true`. That is the *capability*, not the replay itself — pi still has no
|
|
420
|
+
captured `session/load` transcript, which is the one resume-evidence gap open on this lane.
|
|
421
|
+
|
|
422
|
+
## The command's environment is the family's pass-through list, and nothing else
|
|
423
|
+
|
|
424
|
+
The worker does not *merge* an environment into a harness process: the family's `passEnv` enumeration in
|
|
425
|
+
`facts/families.json` **is** that process's complete environment. An environment cannot declare an
|
|
426
|
+
`SFDK_` name for it, because `SFDK_` is reserved apart from the four `SFDK_WORKER_*` release settings, which
|
|
427
|
+
reach the worker and never a harness process. A derived image that bakes a name the list does not carry, such as `ENV JOB_TARGET=…`, gets **silence**,
|
|
428
|
+
not an override: the variable never reaches the command, and a job running under `set -u` aborts on the
|
|
429
|
+
first reference to it. Per-job configuration travels the prompt, the agent's instructions, or a file in the
|
|
430
|
+
image.
|
|
431
|
+
|
|
432
|
+
Every family's harness process also gets `SFDK_SESSION_ID`, which the worker sets from the `sessionId` on
|
|
433
|
+
`session_placement.apply`: the worker adds it itself, so it needs no `passEnv` entry, whatever the worker's
|
|
434
|
+
own environment holds. The `script` family adds only the organization to the command it runs, as
|
|
435
|
+
`ORCHESTRA_ORGANIZATION_ID` and `SFDK_ORGANIZATION_ID`, from its stored conversation metadata; its command inherits
|
|
436
|
+
the session variable from the harness process. Neither grants egress authority.
|
|
437
|
+
|
|
438
|
+
**On Windows that list is shorter than the operating system's own habits assume.** It is
|
|
439
|
+
OS-independent, so a Windows job sees `PATH`, `HOME` and `USERPROFILE` and *not* `TEMP`, `APPDATA` or
|
|
440
|
+
`ComSpec` — names a PowerShell one-liner reaches for without thinking (`$env:TEMP`). The list is not a
|
|
441
|
+
filter over the environment: the worker *builds* the child's environment from it, so a name the platform
|
|
442
|
+
set on the service and the family does not declare is simply gone by the time the command runs.
|
|
443
|
+
|
|
444
|
+
**Three names are backfilled for Windows and none of them is on any family's list.** A backfill is the
|
|
445
|
+
treatment for operating-system *plumbing* — a variable the machine needs to work at all, as opposed to
|
|
446
|
+
configuration a customer chose — and it is if-absent, so a declared spelling always wins. Two come from
|
|
447
|
+
libraries: Go's `os/exec` adds `SYSTEMROOT` whenever a caller sets an explicit environment, and libuv copies
|
|
448
|
+
its own required-variable set into a spawned child from *this* process's environment — which is why `PATH`
|
|
449
|
+
reaching the command is what lets `powershell.exe` and `taskkill.exe` resolve at all. The third is ours:
|
|
450
|
+
**`PATHEXT`**, the list of extensions PowerShell treats as *executable*. With it unset the effective list
|
|
451
|
+
collapses to `.CPL`, so a `curl.exe` PowerShell resolved on `PATH` is a **document** rather than a program
|
|
452
|
+
and invoking it throws `CantActivateDocumentInPipeline` — before a process starts, before a packet leaves.
|
|
453
|
+
Every network probe on the Windows lane came back `000`, which reads exactly like a blocked connection.
|
|
454
|
+
|
|
455
|
+
The neighbours a reader expects beside it were **measured and left out**, each changing nothing on a Windows
|
|
456
|
+
Server 2025 runtime AMI: `cmd.exe` and `.bat` files run from CreateProcess's own default without `ComSpec`,
|
|
457
|
+
`%SystemRoot%` expands from the `SYSTEMROOT` `os/exec` already backfills, `SystemDrive` and `windir` are
|
|
458
|
+
unused by realistic work, and a service's temp directory falls back to `C:\Windows\SystemTemp` without
|
|
459
|
+
`TEMP` or `TMP`. Add a fourth name only with that kind of evidence. Machine plumbing is not configuration
|
|
460
|
+
and must never be pushed into the prompt: a job cannot set `PATHEXT` for a shell that has already refused to
|
|
461
|
+
start its first program.
|
|
462
|
+
|
|
463
|
+
Three more bounds a job author should know:
|
|
464
|
+
|
|
465
|
+
- **Turn wall clock.** The platform bounds a turn with its own response timeout (default 600 seconds,
|
|
466
|
+
configurable). A job that can run longer needs that raised, or it is killed mid-turn and reported as a
|
|
467
|
+
runtime failure regardless of what the script was doing.
|
|
468
|
+
- **`session/load` re-executes.** The script harness holds no conversational state, so a resumed
|
|
469
|
+
conversation's command runs again. That is the safe direction — a fresh, deterministic run — and another
|
|
470
|
+
reason to keep commands idempotent.
|
|
471
|
+
- **`cwd` is the session's.** `session/new` (and `session/load` on a revival) carries the workspace path, the
|
|
472
|
+
worker maps it to this machine's native filesystem — `C:\sf\workspace` on Windows — and the harness spawns
|
|
473
|
+
the command there. A session that declares none inherits the harness's own directory; one that does not
|
|
474
|
+
exist fails the spawn, which the turn reports as an ordinary exit-127 failure with the operating system's
|
|
475
|
+
message attached.
|
|
476
|
+
|
|
477
|
+
## What the `script` harness answers
|
|
478
|
+
|
|
479
|
+
The worker runs one harness process per session placement and speaks ACP to it as its client. The `script`
|
|
480
|
+
family's harness is the one the worker carries itself (`script-harness/`), and its dispatch table is:
|
|
481
|
+
|
|
482
|
+
- `initialize` — the harness identifies itself.
|
|
483
|
+
- `session/new` — a fresh session id, with the session's declared working directory anchored on it.
|
|
484
|
+
- `session/load` — the session's replay, as the `session/update` notifications emitted between the request
|
|
485
|
+
and its response. It executes nothing and calls nothing, and it restores the duplicate-execution guard
|
|
486
|
+
below, which the notifications cannot carry.
|
|
487
|
+
- `session/prompt` — execute the agent's instructions, stream stdout and stderr as `session/update`
|
|
488
|
+
notifications, move platform state from the exit code, then answer. Streaming is bounded twice: each frame
|
|
489
|
+
is sliced at 32000 characters (one over-long line would trip the line scanner and kill every session on
|
|
490
|
+
the pipe, not just this turn), and a turn's **total** streamed volume is capped at 5 MiB of **wire** bytes
|
|
491
|
+
— the JSON-encoded frame text, since that is what the platform records, and JSON escaping inflates
|
|
492
|
+
control-heavy output up to about six times (one NUL byte is six wire bytes). Past the ceiling the harness
|
|
493
|
+
emits one truncation notice and stops streaming, while the command still runs to completion with its
|
|
494
|
+
exit-code contract unchanged.
|
|
495
|
+
- `session/cancel` — stop the in-flight command **and everything it started**, then answer. On Linux that is
|
|
496
|
+
a `SIGTERM` to the command's whole **process group** (commands run detached, so the group dies with the
|
|
497
|
+
command and not just with the `bash -c` parent); on Windows there are no process groups to signal, so it
|
|
498
|
+
is `taskkill /T /F /PID <pid>`, which walks the process tree. A cancelled turn made no determination, so
|
|
499
|
+
it claims nothing and answers the prompt with a JSON-RPC error. Whether a turn counts as cancelled is read
|
|
500
|
+
off an observed **fact about the kill**, never off the arrival of the cancel, and each operating system
|
|
501
|
+
supplies the fact it can: on Linux the command's **exit shape** (one the signal actually killed reports
|
|
502
|
+
`signal = SIGTERM`, one that had already exited reports its own status), and on Windows — which terminates
|
|
503
|
+
a process with an exit *code* and no signal — the kill's own answer, so the harness checks whether the
|
|
504
|
+
shell was still running when the cancel arrived and only then kills. Either way a cancel racing a script's
|
|
505
|
+
own `exit 0`, including the window after the shell exits but before the harness reaps it, leaves the turn
|
|
506
|
+
with the result it earned, claim included. (That liveness check is also a safety requirement on Windows,
|
|
507
|
+
which recycles process ids: a `/T` on a stale one would kill a stranger's process tree.)
|
|
508
|
+
|
|
509
|
+
A turn that settles while a backgrounded grandchild still holds the pipes signals that process group and
|
|
510
|
+
detaches the streams, so nothing leaks into the next turn's transcript. That sweep is **conditional and
|
|
511
|
+
best-effort**, and a job author should not read it as a guarantee that background work dies with the turn:
|
|
512
|
+
it runs only on the grace-timer path, so a background job whose stdio is redirected (`cmd >/dev/null 2>&1 &`)
|
|
513
|
+
lets the pipes close promptly, settles the turn before the timer, and survives it untouched. It is also
|
|
514
|
+
`SIGTERM`-only — a process that traps or ignores `TERM` survives either way. Nothing of the sort leaks into
|
|
515
|
+
the *transcript*; what survives is the process.
|
|
516
|
+
|
|
517
|
+
A second prompt on a **workflow-step** session whose earlier claim is confirmed landed is answered as a
|
|
518
|
+
no-op turn without re-executing the command (warm reuse and steer promotion both make that routine); an
|
|
519
|
+
earlier claim that is merely *attempted* is re-confirmed against the run first. The no-op is step-scoped on
|
|
520
|
+
purpose: an ordinary message appended to a direct-start session is a real new turn and always executes.
|
package/package.json
CHANGED
|
@@ -1,8 +1,8 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@orchestraworks/worker",
|
|
3
|
-
"version": "0.0.
|
|
3
|
+
"version": "0.0.136",
|
|
4
4
|
"type": "module",
|
|
5
|
-
"description": "The
|
|
5
|
+
"description": "The Orchestra worker.",
|
|
6
6
|
"license": "Apache-2.0",
|
|
7
7
|
"author": "Living Computers",
|
|
8
8
|
"homepage": "https://orchestraworks.ai",
|
|
@@ -18,13 +18,14 @@
|
|
|
18
18
|
},
|
|
19
19
|
"files": [
|
|
20
20
|
"bin",
|
|
21
|
-
"README.md"
|
|
21
|
+
"README.md",
|
|
22
|
+
"docs/REFERENCE.md"
|
|
22
23
|
],
|
|
23
24
|
"optionalDependencies": {
|
|
24
|
-
"@orchestraworks/worker-linux-x64": "0.0.
|
|
25
|
-
"@orchestraworks/worker-linux-arm64": "0.0.
|
|
26
|
-
"@orchestraworks/worker-darwin-x64": "0.0.
|
|
27
|
-
"@orchestraworks/worker-darwin-arm64": "0.0.
|
|
28
|
-
"@orchestraworks/worker-win32-x64": "0.0.
|
|
25
|
+
"@orchestraworks/worker-linux-x64": "0.0.136",
|
|
26
|
+
"@orchestraworks/worker-linux-arm64": "0.0.136",
|
|
27
|
+
"@orchestraworks/worker-darwin-x64": "0.0.136",
|
|
28
|
+
"@orchestraworks/worker-darwin-arm64": "0.0.136",
|
|
29
|
+
"@orchestraworks/worker-win32-x64": "0.0.136"
|
|
29
30
|
}
|
|
30
31
|
}
|