@nimbus-sh/fabric 0.1.0 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (104) hide show
  1. package/README.md +208 -293
  2. package/dist/bindings.js +5 -5
  3. package/dist/budgets.d.ts +132 -0
  4. package/dist/budgets.d.ts.map +1 -0
  5. package/dist/budgets.js +248 -0
  6. package/dist/composition.d.ts +3 -0
  7. package/dist/composition.d.ts.map +1 -0
  8. package/dist/composition.js +2 -0
  9. package/dist/connections.d.ts +81 -0
  10. package/dist/connections.d.ts.map +1 -0
  11. package/dist/connections.js +114 -0
  12. package/dist/derived.d.ts +65 -0
  13. package/dist/derived.d.ts.map +1 -0
  14. package/dist/derived.js +95 -0
  15. package/dist/do-calls.d.ts +94 -0
  16. package/dist/do-calls.d.ts.map +1 -0
  17. package/dist/do-calls.js +111 -0
  18. package/dist/facet-pool.d.ts +90 -0
  19. package/dist/facet-pool.d.ts.map +1 -0
  20. package/dist/facet-pool.js +113 -0
  21. package/dist/{fanout-pool.d.ts → fanout.d.ts} +20 -20
  22. package/dist/fanout.d.ts.map +1 -0
  23. package/dist/{fanout-pool.js → fanout.js} +20 -20
  24. package/dist/{launch-journal.d.ts → fenced-work.d.ts} +58 -17
  25. package/dist/fenced-work.d.ts.map +1 -0
  26. package/dist/fenced-work.js +241 -0
  27. package/dist/generation.d.ts +69 -0
  28. package/dist/generation.d.ts.map +1 -0
  29. package/dist/generation.js +118 -0
  30. package/dist/{facet-image-store.d.ts → image-store.d.ts} +8 -8
  31. package/dist/image-store.d.ts.map +1 -0
  32. package/dist/{facet-image-store.js → image-store.js} +4 -4
  33. package/dist/index.d.ts +16 -8
  34. package/dist/index.d.ts.map +1 -1
  35. package/dist/index.js +16 -8
  36. package/dist/{loader-pool.d.ts → isolate-pool.d.ts} +19 -19
  37. package/dist/isolate-pool.d.ts.map +1 -0
  38. package/dist/{loader-pool.js → isolate-pool.js} +20 -20
  39. package/dist/journal.d.ts +111 -0
  40. package/dist/journal.d.ts.map +1 -0
  41. package/dist/journal.js +177 -0
  42. package/dist/outbox.d.ts +249 -0
  43. package/dist/outbox.d.ts.map +1 -0
  44. package/dist/outbox.js +355 -0
  45. package/dist/process-fabric.d.ts +33 -15
  46. package/dist/process-fabric.d.ts.map +1 -1
  47. package/dist/process-fabric.js +25 -15
  48. package/dist/process-host.d.ts +1 -1
  49. package/dist/process-host.d.ts.map +1 -1
  50. package/dist/process-host.js +19 -11
  51. package/dist/sealed.d.ts +78 -0
  52. package/dist/sealed.d.ts.map +1 -0
  53. package/dist/sealed.js +145 -0
  54. package/dist/timers.d.ts +138 -0
  55. package/dist/timers.d.ts.map +1 -0
  56. package/dist/timers.js +231 -0
  57. package/dist/{launch-pacer.d.ts → turn-budget.d.ts} +24 -21
  58. package/dist/turn-budget.d.ts.map +1 -0
  59. package/dist/{launch-pacer.js → turn-budget.js} +24 -12
  60. package/dist/workerd-facet-host.d.ts +67 -70
  61. package/dist/workerd-facet-host.d.ts.map +1 -1
  62. package/dist/workerd-facet-host.js +129 -181
  63. package/examples/agent-core-adapter.ts +191 -0
  64. package/package.json +4 -2
  65. package/src/bindings.ts +6 -6
  66. package/src/budgets.ts +308 -0
  67. package/src/composition.ts +16 -0
  68. package/src/connections.ts +140 -0
  69. package/src/derived.ts +135 -0
  70. package/src/do-calls.ts +156 -0
  71. package/src/facet-pool.ts +157 -0
  72. package/src/{fanout-pool.ts → fanout.ts} +35 -35
  73. package/src/{launch-journal.ts → fenced-work.ts} +129 -42
  74. package/src/generation.ts +144 -0
  75. package/src/{facet-image-store.ts → image-store.ts} +9 -9
  76. package/src/index.ts +16 -8
  77. package/src/{loader-pool.ts → isolate-pool.ts} +34 -34
  78. package/src/journal.ts +242 -0
  79. package/src/node-async-hooks.d.ts +14 -0
  80. package/src/outbox.ts +520 -0
  81. package/src/process-fabric.ts +43 -34
  82. package/src/process-host.ts +22 -20
  83. package/src/sealed.ts +150 -0
  84. package/src/timers.ts +294 -0
  85. package/src/{launch-pacer.ts → turn-budget.ts} +34 -27
  86. package/src/workerd-facet-host.ts +159 -208
  87. package/dist/alarms.d.ts +0 -134
  88. package/dist/alarms.d.ts.map +0 -1
  89. package/dist/alarms.js +0 -214
  90. package/dist/ctx-exports.d.ts +0 -47
  91. package/dist/ctx-exports.d.ts.map +0 -1
  92. package/dist/ctx-exports.js +0 -54
  93. package/dist/facet-image-store.d.ts.map +0 -1
  94. package/dist/fanout-pool.d.ts.map +0 -1
  95. package/dist/launch-journal.d.ts.map +0 -1
  96. package/dist/launch-journal.js +0 -154
  97. package/dist/launch-pacer.d.ts.map +0 -1
  98. package/dist/loader-ledger.d.ts +0 -57
  99. package/dist/loader-ledger.d.ts.map +0 -1
  100. package/dist/loader-ledger.js +0 -91
  101. package/dist/loader-pool.d.ts.map +0 -1
  102. package/src/alarms.ts +0 -275
  103. package/src/ctx-exports.ts +0 -77
  104. package/src/loader-ledger.ts +0 -112
package/README.md CHANGED
@@ -4,69 +4,74 @@
4
4
  > cloud OS. This README is edited and maintained with Claude (AI) and
5
5
  > presented as-is.
6
6
 
7
- The Cloudflare-specific half of Nimbus: the machinery for running real
8
- programs on Durable Objects, DO facets, and the Worker Loader. Where
9
- [`@nimbus-sh/core`](https://www.npmjs.com/package/@nimbus-sh/core) is the
10
- backend-agnostic OS (filesystem, shell, process contracts), this package is
11
- what that OS stands on when the host is Cloudflare — and it never imports the
12
- OS's policy, only its shared primitives.
13
-
14
- I extracted it because almost none of it is specific to Nimbus. Anyone who
15
- hosts long-lived processes on Durable Objects meets the same platform
16
- behaviors we did: `await put()` resolving before durability, one alarm per
17
- object, a 65,536-facet lifetime budget, a frozen in-DO clock, RPC stubs that
18
- die with their request context. This package is the machinery we built against
19
- those behaviors, with the measured numbers that justified each mechanism
20
- carried in the doc comments — they are the design record, and they travel with
21
- the code on purpose.
22
-
23
- Everything below was measured on deployed production workerd, not on
24
- `wrangler dev` and not inferred from types, between June and August 2026.
25
- Where a specific date matters it is given.
26
-
27
- ## Importing it
28
-
29
- The root export pulls `cloudflare:workers`, so `import ... from
30
- '@nimbus-sh/fabric'` resolves only inside a Worker. Outside workerd (unit
31
- tests, tooling) import the subpath modules directly —
32
- `@nimbus-sh/fabric/alarms.js`, `@nimbus-sh/fabric/launch-journal.js`, and so
33
- on. Most of the package is structurally typed against plain objects precisely
34
- so it can be tested in bun or node.
35
-
36
- An embedder wires three seams at composition time, each first-write-wins:
7
+ Run long-lived programs on Cloudflare Durable Objects.
8
+
9
+ A Durable Object gives you one alarm, one 128 MiB isolate, a 30-second CPU
10
+ turn, and storage that can reset under you. This package turns those into
11
+ things you can build on: many timers on the one alarm, work that survives a
12
+ reset, CPU work that spans turns, and real processes in their own isolates.
13
+
14
+ Use it if you host something that outlives a request. A dev server, a build,
15
+ an agent, a terminal session.
16
+
17
+ ## Install
18
+
19
+ ```bash
20
+ npm install @nimbus-sh/fabric
21
+ ```
22
+
23
+ ## Requirements
24
+
25
+ Set `compatibility_flags: ["nodejs_compat"]` in your Worker. The timer
26
+ dispatcher needs `AsyncLocalStorage`, which workerd ships only under that
27
+ flag. Without it the module fails to load at deploy time.
28
+
29
+ Import the root inside a Worker. Outside workerd, import subpaths such as
30
+ `@nimbus-sh/fabric/timers.js`, which are typed against plain objects and run
31
+ in bun or node.
32
+
33
+ ## Setup
34
+
35
+ Declare your composition once, in your Worker's entry module.
37
36
 
38
37
  ```ts
39
- import { setCtxExports, setSupervisorEntrypointName, setStagedBootAssembler } from '@nimbus-sh/fabric';
40
-
41
- // In the Worker's fetch handler, once:
42
- setCtxExports(ctx.exports);
43
- // The name of your supervisor WorkerEntrypoint export. The fabric mints one
44
- // binding per hosted program from it (env.SUPERVISOR inside the facet).
45
- setSupervisorEntrypointName('MySupervisorRPC');
46
- // Only if you use 'staged' boot specs; 'code' boots need no assembler.
47
- setStagedBootAssembler(async (env, stage) => assembleLoaderConfig(env, stage));
38
+ import { composeFabric, adoptCtxExports } from '@nimbus-sh/fabric';
39
+
40
+ // Module scope of the Worker entry, once per isolate:
41
+ composeFabric({
42
+ // The name of your supervisor WorkerEntrypoint export. The fabric mints one
43
+ // binding per hosted program from it (env.SUPERVISOR inside the facet).
44
+ supervisorEntrypoint: 'MySupervisorRPC',
45
+ // Only if you use 'staged' boot specs; 'code' boots need no assembler.
46
+ stagedBootAssembler: async (env, stage) => assembleLoaderConfig(env, stage),
47
+ });
48
+
49
+ // ctx.exports is runtime state, not composition. Capture it where the
50
+ // platform hands it over — the first fetch, or the DO constructor:
51
+ adoptCtxExports(ctx.exports);
48
52
  ```
49
53
 
50
- ## One alarm, many reasons
54
+ Both calls take the first value they are given.
51
55
 
52
- A Durable Object has ONE alarm, and a second `setAlarm()` silently overwrites
53
- the first. Every alarm-driven subsystem therefore coordinates through a single
54
- reason→deadline map in storage, with one dispatcher:
56
+ ## Timers
57
+
58
+ A Durable Object has one alarm, and a second `setAlarm()` overwrites the
59
+ first. Route every timer through one dispatcher instead.
55
60
 
56
61
  ```ts
57
62
  import { DurableObject } from 'cloudflare:workers';
58
- import { scheduleAlarm, dispatchAlarm } from '@nimbus-sh/fabric';
63
+ import { timers } from '@nimbus-sh/fabric';
59
64
 
60
65
  export class MySession extends DurableObject {
61
- _alarmChain?: Promise<unknown>; // serializes the map's read-modify-write
66
+ _timerChain?: Promise<unknown>; // serializes the map's read-modify-write
62
67
 
63
68
  async fetch(request: Request): Promise<Response> {
64
- await scheduleAlarm(this, this.ctx, 'janitor', Date.now() + 60_000);
69
+ await timers(this, this.ctx).schedule('janitor', Date.now() + 60_000);
65
70
  return new Response('ok');
66
71
  }
67
72
 
68
73
  async alarm(): Promise<void> {
69
- await dispatchAlarm(this, this.ctx, {
74
+ await timers(this, this.ctx).dispatch({
70
75
  janitor: async (now) => {
71
76
  await this.cleanUp();
72
77
  return { rearmAt: now + 60_000 }; // re-arm through the return value
@@ -76,61 +81,49 @@ export class MySession extends DurableObject {
76
81
  }
77
82
  ```
78
83
 
79
- `scheduleAlarm` keeps the earliest deadline per reason and arms the real alarm
80
- at the minimum across all of them. `dispatchAlarm` snapshots the fireable set
81
- before running any handler (so a handler that re-schedules itself is not
82
- re-fired in the same dispatch), silently drops unknown reasons (a rollback
83
- from a deploy that added reasons must not wedge the alarm), and when no
84
- reasons remain it deletes the map and does not re-arm — which is what lets the
85
- object hibernate.
84
+ `schedule` stores a deadline per reason and arms the alarm at the earliest
85
+ one. `dispatch` runs the reasons that are due, ignores reasons it does not
86
+ know, and stops re-arming when none are left, which lets the object
87
+ hibernate. A handler re-arms itself by returning `{ rearmAt }`.
86
88
 
87
- ## Knowing which incarnation you are
89
+ ## Generations and reset detection
88
90
 
89
- Workerd recycles isolates freely: cold starts, hibernation wakes, and resets
90
- all hand you a fresh module scope over the same storage. The isolate
91
- generation is a persisted counter that increments once per fresh isolate, and
92
- process IDs derive from it (`PID_GEN_STRIDE` = 1,000,000 in core's process
93
- table), which yields the one reset predicate everything else builds on: **a
94
- pid at or below the current generation's base was allocated by a previous
95
- incarnation.**
91
+ Workerd hands you a fresh isolate on cold starts, hibernation wakes, and
92
+ resets. The generation counter tells you which incarnation you are in, and
93
+ process IDs derive from it.
96
94
 
97
95
  ```ts
98
- import { maybeBumpIsolateGen } from '@nimbus-sh/fabric';
96
+ import { adoptGeneration, generation } from '@nimbus-sh/fabric';
99
97
 
100
98
  export class MySession extends DurableObject {
101
- _isolateGen = 0;
102
- _isolateGenPersisted = false;
103
-
104
99
  async fetch(request: Request): Promise<Response> {
105
- await maybeBumpIsolateGen(this, this.ctx); // idempotent per instance
106
- // this._isolateGen is now this incarnation's generation
100
+ await adoptGeneration(this.ctx); // idempotent per instance
101
+ // generation(this.ctx) is now this incarnation's generation
107
102
  }
108
103
  }
109
104
  ```
110
105
 
111
- The ordering inside is deliberate: adopt the persisted value first, bump only
112
- after the `put` resolves. An unpersisted bump would be re-read by the next
113
- boot and re-issued — two instances sharing one generation is exactly the pid
114
- aliasing the counter exists to prevent. Note that `await put()` returning is
115
- not durability; the output gate is what keeps a pid from generation N from
116
- escaping before N is on disk.
106
+ This gives you one reliable test for stale state: **an ID at or below the
107
+ current generation's base came from a previous incarnation.**
117
108
 
118
- ## The launch journal: surviving resets
109
+ Adopt the persisted value before bumping it. `await put()` returning is not
110
+ durability, so a bump that has not landed can be re-issued to the next boot,
111
+ and two instances would share a generation.
119
112
 
120
- The platform resets a Durable Object over what one turn has outstanding in
121
- storage, and the reset destroys every write that turn had in flight. A
122
- long-running launch holds everything in memory, so the process it is building
123
- dies silently with the instance. The journal is what a later instance reads to
124
- know that happened:
113
+ ## Fenced work
114
+
115
+ A reset destroys whatever the current turn had in flight, including a
116
+ half-built process. Journal the work first, and a later instance can finish
117
+ it.
125
118
 
126
119
  ```ts
127
- import { ResidentLaunchJournal, type ResidentLaunchRecord } from '@nimbus-sh/fabric';
120
+ import { FencedWork, type FencedWorkRecord } from '@nimbus-sh/fabric';
128
121
 
129
- interface MyLaunch extends ResidentLaunchRecord {
122
+ interface MyLaunch extends FencedWorkRecord {
130
123
  argv: string[]; // whatever your redrive needs; the journal never reads it
131
124
  }
132
125
 
133
- const journal = new ResidentLaunchJournal<MyLaunch>(this.ctx.storage, {
126
+ const journal = new FencedWork<MyLaunch>(this.ctx.storage, {
134
127
  generationBase: () => this.pidBase,
135
128
  waitUntil: (p) => this.ctx.waitUntil(p),
136
129
  redrive: (record, attempt) => this.launch(record.argv, attempt),
@@ -144,68 +137,47 @@ await journal.release(pid);
144
137
  await journal.recoverInterrupted();
145
138
  ```
146
139
 
147
- Two details here cost us real incidents before they were mechanisms:
148
-
149
- - **`put` then `sync()`.** `await storage.put()` resolves before durability.
150
- Measured live: a launch killed in its first chunks left NO row for the
151
- replacement instance to find, which is how the recovery this feeds sat inert
152
- while its own test stayed green. `sync()` is the storage layer's durability
153
- barrier; the journal writes through it on the way in and on the way out
154
- (delete-then-sync, so a reset moments after release cannot resurrect a
155
- process the user watched end).
156
- - **The row lives for the process's lifetime, not the launch's.** Measured on
157
- staging, 2026-08-13: every observed reset struck seconds AFTER the launch
158
- settled. A launch-scoped row would already have been deleted when recovery
159
- went looking.
160
-
161
- Recovery applies the generation predicate (`pid <= generationBase()`), deletes
162
- each stale row, and re-drives once per record (`RESIDENT_LAUNCH_MAX_ATTEMPT` =
163
- 1) — a reset that recurs is not the transient kind.
164
-
165
- ## Pacing big work across turns
166
-
167
- One DO turn has a CPU budget of about 30 s (we were killed with `exceededCpu`
168
- at 31.8 s and 32.5 s), and yielding inside an invocation buys nothing — CPU
169
- accrues to the invocation, and only genuinely re-entering the object resets
170
- it. Worse, a long turn pins the actor's only thread, so the terminal WebSocket
171
- dies even when the work succeeds. And progress cannot be measured in
172
- milliseconds, because the in-DO clock does not advance without I/O (0 ms
173
- across 200,000 consecutive reads). So the pacer accounts **bytes**:
140
+ Write the row before the work starts and release it when the process ends,
141
+ not when the launch ends. Resets usually arrive after a launch settles, so a
142
+ launch-scoped row is already gone when recovery looks for it.
143
+
144
+ The journal writes through `ctx.storage.sync()`, because `await put()`
145
+ resolves before the write is durable. Recovery re-drives each stale row once.
146
+
147
+ ## Turn pacing
148
+
149
+ One turn gets about 30 seconds of CPU. Yielding inside a turn does not help,
150
+ because CPU accrues to the invocation. Only re-entering the object resets the
151
+ budget, and a long turn also blocks the actor's thread and drops WebSockets.
174
152
 
175
153
  ```ts
176
- import { LaunchPacer, LaunchTurnPump, scheduleAlarm } from '@nimbus-sh/fabric';
154
+ import { TurnBudget, PacedWork, onColdStart, timers } from '@nimbus-sh/fabric';
177
155
 
178
- const pump = new LaunchTurnPump({
179
- requestTurn: () => { void scheduleAlarm(this, this.ctx, 'launch-turn', Date.now()); },
180
- recover: () => journal.recoverInterrupted(),
156
+ const pump = new PacedWork(this.ctx, {
157
+ requestTurn: () => { void timers(this, this.ctx).schedule('launch-turn', Date.now()); },
181
158
  });
182
- const pacer = new LaunchPacer(pump);
159
+ // Deferred reconciliation rides the first pump, off the init gate:
160
+ onColdStart(this.ctx, () => journal.recoverInterrupted());
161
+ const budget = new TurnBudget(pump);
183
162
 
184
163
  // Inside the launch, after each unit of work:
185
- await pacer.spend(bytesJustProcessed); // suspends every LAUNCH_CHUNK_MAX_BYTES (2 MB)
164
+ await budget.spend(bytesJustProcessed); // suspends every TURN_CHUNK_MAX_BYTES (2 MB)
186
165
  // In alarm(), as one of the dispatcher's reasons:
187
166
  'launch-turn': () => pump.pump(),
188
167
  ```
189
168
 
190
- The pump awaits each resumed chunk, so the invocation that granted the turn is
191
- the invocation that pays for the work — nothing runs detached in a handler's
192
- microtask drain. A past-deadline alarm is delivered as soon as the object is
193
- free, which makes `scheduleAlarm(..., Date.now())` a genuine "re-enter now"
194
- primitive. Without an alarm-capable host the pump degrades to a same-context
195
- timer: the single-turn behaviour this path always had, minus the
196
- responsiveness.
169
+ Account in bytes, not milliseconds: the in-DO clock does not advance without
170
+ I/O. `spend()` suspends every 2 MB and resumes on a fresh turn. Scheduling a
171
+ past deadline re-enters the object immediately.
197
172
 
198
- ## Running programs: the loader pool
173
+ ## Isolate pool
199
174
 
200
- `LoaderPool` runs plain functions in warm dynamic-worker isolates over
201
- `env.LOADER`. Functions are serialized with `fn.toString()`, so they must be
202
- self-contained: no captured variables, no `this` (rejected at dispatch), and
203
- their last parameter receives the forwarded bindings.
175
+ Run plain functions in warm dynamic-worker isolates.
204
176
 
205
177
  ```ts
206
- import { LoaderPool } from '@nimbus-sh/fabric';
178
+ import { IsolatePool } from '@nimbus-sh/fabric';
207
179
 
208
- const pool = new LoaderPool(env, this.ctx, {
180
+ const pool = new IsolatePool(env, this.ctx, {
209
181
  concurrency: 4,
210
182
  tag: 'checksum',
211
183
  omitSupervisor: true, // this pool needs no callback into the DO
@@ -220,45 +192,32 @@ try {
220
192
  }
221
193
  ```
222
194
 
223
- Slots are stable (`slot = index % concurrency`) so a batch of 67 tarball
224
- extractions reuses 4 warm isolates instead of paying 67 cold starts. Wasm
225
- rides the loader's modules map as `{ wasm: ArrayBuffer }` — the only path that
226
- works, since request-time `WebAssembly.compile` is CSP-blocked, RPC of a
227
- compiled `Module` is refused by structured clone, and inlining bytes into the
228
- module source OOMs the supervisor.
229
-
230
- The cache key folds the function hash, the preamble hash, a wasm fingerprint,
231
- and **the first 12 characters of the owning DO's id**. That last term is a
232
- security lesson, not an optimization: without it, session B's pool reused
233
- session A's warm isolate — which still carried A's `env.SUPERVISOR` binding —
234
- and B's writes landed silently in A's filesystem while B's install reported
235
- success. Warm isolates are scoped to one session unless a pool explicitly opts
236
- into `cacheScope: 'global'`, which is reserved for stateless compute pools
237
- that take no supervisor binding and retain no user state.
238
-
239
- `FanoutPool` is the tier above: a single DO method can drive at most 4
240
- concurrent Worker Loader fetches, so batches of fewer than 5 tasks run in the
241
- coordinator through a `LoaderPool` and wider batches shard deterministically
242
- across sibling DOs (up to 32, dispatched in phases of 4 to bound simultaneous
243
- cold starts). Transient peer resets retry on a 250/750/1500 ms schedule; an
244
- overloaded peer gets the 1/3/6 s one.
245
-
246
- Every fabric call into the loader lands on a per-DO ledger: distinct ids ever
247
- gotten — each permanently holds one of the ~5–6 dynamic-worker slots, because
248
- a keyed `loader.get(id)` is never released — plus live and peak concurrent
249
- Loader fetches, read via `loaderLedgerStats(ctx)`. A "Too many concurrent
250
- dynamic workers" refusal classifies as `dynamic_worker_cap` and is annotated
251
- with the ids actually holding slots. Measurement and honest failure naming
252
- only — no admission control, because the cap is the platform's and
253
- approximate, and a gate on an approximate number would refuse work the
254
- platform would have run.
255
-
256
- ## Running processes: the resident fabric
257
-
258
- A resident process — a dev server, a socket runner, an attached TUI — is a DO
259
- facet whose class comes from a dynamic worker. `openResidentFacet` is the one
260
- way such a process comes into existence; `ProcessFabric` is the lifecycle
261
- around it:
195
+ Functions are serialized with `fn.toString()`, so they must be
196
+ self-contained: no captured variables, no `this`. Bindings arrive as the last
197
+ parameter. Slots are stable, so a batch of 67 tasks reuses 4 warm isolates
198
+ instead of paying 67 cold starts.
199
+
200
+ Ship wasm through the loader's modules map as `{ wasm: ArrayBuffer }`. It is
201
+ the only path that works. Request-time `WebAssembly.compile` is blocked by
202
+ CSP, structured clone refuses a compiled `Module`, and inlining bytes into
203
+ the source exhausts the supervisor's memory.
204
+
205
+ Warm isolates are scoped to one session. A pool may opt into
206
+ `cacheScope: 'global'` only if it takes no supervisor binding and keeps no
207
+ user state.
208
+
209
+ `Fanout` handles wider batches. One DO method can drive at most 4 concurrent
210
+ loader fetches, so batches under 5 run in the coordinator and larger ones
211
+ shard across up to 32 sibling objects, 4 at a time.
212
+
213
+ Each keyed `loader.get(id)` permanently holds one of roughly 5–6
214
+ dynamic-worker slots. `loaderLedgerStats(ctx)` reports what you have
215
+ consumed, and a cap refusal names the IDs holding slots.
216
+
217
+ ## Process fabric
218
+
219
+ A resident process is a Durable Object facet running a class from a dynamic
220
+ worker.
262
221
 
263
222
  ```ts
264
223
  import { ProcessFabric, createProcessHost } from '@nimbus-sh/fabric';
@@ -283,124 +242,81 @@ handle.kill();
283
242
  await handle.done;
284
243
  ```
285
244
 
286
- The dynamic worker must export a Durable Object class named `NimbusProcess`
287
- (`RESIDENT_PROCESS_CLASS`) with `startProcess(args)` and
288
- `handleHttpRequest(request)`. Its `startProcess` declares one of two contracts:
289
- `'lifetime'` (the call is held open for the process's whole life and settles at
290
- exit — an attached TUI) or `'boot'` (the call returns a payload once the
291
- process is up and the facet stays resident — a server).
292
-
293
- Pieces worth knowing about, each earned the hard way:
294
-
295
- - **The slot book.** A Durable Object admits 65,536 facets over its LIFETIME —
296
- the IDs are append-only and never reclaimed, so the bound is on facets ever
297
- created. Naming facets after pids burned one ID per spawn with no way back.
298
- Reusing a NAME costs no new ID, so facet names come from a per-DO free list
299
- (`proc-slot-<n>`, lowest reused first), and a slot is released only after
300
- `facets.abort` + `facets.delete` — a slot handed out during teardown would
301
- put two processes on one name. The names the book does mint are counted
302
- durably — `facetIdBudget(ctx)` reports `{ consumed, budget }`, first uses
303
- only, adopted across resets — and a creation failure with the budget
304
- consumed names the budget and the count instead of repeating the platform's
305
- opaque message. Exhaustion is permanent for the object, so it is the one
306
- failure worth naming precisely.
307
- - **At-most-once start.** The facet's start callback re-running would
308
- re-execute the user's program, answering a request from a process the user
309
- never started. Both re-entry cases (released, lost) throw instead.
310
- - **Boot specs name large members by VFS path.** A whole structured-clone RPC
311
- value caps at 32 MiB, and one node snapshot alone serialized to 44,252,709
312
- bytes. `vfsWasmModules` and `vfsTextModules` send paths; the hosting actor
313
- reads the bytes through the `ResidentDiskReader` it was given, inside the
314
- loader's cache-miss callback, so they exist only for the duration of the
315
- load. Text images are verified against the digest their own path claims —
316
- a truncated image would otherwise boot as silently-wrong code.
317
- - **The substrate is one deployment-wide value** (`createProcessHost`'s mode,
318
- `'facet'` or `'peer'`), never per-spawn. No program name, mode, or payload
319
- size reaches the choice.
320
-
321
- What each substrate costs, measured on the production shape:
245
+ The worker exports a Durable Object class named `NimbusProcess` with
246
+ `startProcess(args)` and `handleHttpRequest(request)`. `startProcess`
247
+ declares one of two contracts. Use `'lifetime'` when the call should stay
248
+ open for the process's life, as an attached terminal does. Use `'boot'` when
249
+ it should return once the process is up and leave the facet resident, as a
250
+ server does.
251
+
252
+ **Facet names come from a free list.** A Durable Object allows 65,536 facets
253
+ over its lifetime, and IDs are never reclaimed, so the limit counts facets
254
+ ever created. Reusing a name costs no new ID. `facetIdBudget(ctx)` reports
255
+ `{ consumed, budget }`.
256
+
257
+ **Large boot members travel as VFS paths.** A structured-clone RPC value caps
258
+ at 32 MiB. Use `vfsWasmModules` and `vfsTextModules`; the host reads the
259
+ bytes during the load and verifies each image against the digest its path
260
+ claims.
261
+
262
+ **The substrate is one deployment-wide setting**, `'facet'` or `'peer'`,
263
+ never a per-spawn choice.
322
264
 
323
265
  | | spawn | memory | CPU | SQLite |
324
266
  |---|---|---|---|---|
325
- | facet | 8–16 ms | independent (~208 MiB each) | SHARED | own |
267
+ | facet | 8–16 ms | independent (~208 MiB each) | shared | own |
326
268
  | peer | 242–359 ms | independent | independent | own |
327
269
 
328
- Facet CPU is shared because facets are separate isolates inside one actor
329
- thread: awaiting I/O yields it completely, but a deliberate 9,956 ms CPU burn
330
- stalled a sibling for 9,966 ms. A peer pays roughly 20× the spawn cost to buy
331
- that back, and verifies its placement rather than assuming it — a module-scope
332
- UUID token is compared across the hop, up to 4 sibling names tried, because a
333
- peer that co-located shares the CPU it was chosen to escape.
334
-
335
- The substrates also differ in image delivery, stated in the
336
- `ProcessImageDelivery` contract rather than smoothed over: a facet shares its
337
- session's Durable Object, so the session's store is reachable by
338
- copy-on-write (`ctx.facets.clone`: 18–31 ms for a 45.73 MB corpus, 34–54 ms
339
- for 1 GB — flat, because nothing is copied) but also shares the session's
340
- ~10 GiB storage budget. A peer brings its own budget and no reflink: clone is
341
- same-object-only and workerd exposes no `VACUUM INTO`, `ATTACH`, or
342
- `sqlite3_backup` across objects. And a clone hazard we measured rather than
343
- assumed: ANY unresolvable `src` — a typo, a name not created yet — silently
344
- EMPTIES the destination and reports success. `cloneFacetStorage` is the one
345
- way the fabric calls clone: it takes the caller's `populated(name)` probe and
346
- asserts it positively on the source before the clone and on the destination
347
- after, so a typo is refused before the platform call and a wiped destination
348
- is never reported as success. An emptied facet still shows a 4,096-byte
349
- database — one page — which is why the probe must find the caller's own data,
350
- not a non-zero size.
351
-
352
- ## The image store
353
-
354
- `FacetImageStore` materializes generated boot images into a content-addressed
355
- store (`var/lib/nimbus/facet-images/<sha256>.js`) through a small
356
- `FacetImageBlobStore` port — the embedder owns the disk, the store owns the
357
- protocol:
358
-
359
- - **Root before the first byte.** The whole root set is registered
360
- synchronously before any byte lands, so the sweep can never observe a
361
- written-but-unclaimed image, however many turns the write spans.
362
- - **Sliced writes.** One transaction takes `FACET_IMAGE_WRITE_SLICE_BYTES`
363
- (a whole number of VFS chunks under the 1 MiB transaction bound — a slice
364
- ending mid-chunk forces a read-back, and an oversize write falls back to
365
- copy-on-write, which is quadratic). A 22.9 MB map written in one turn took
366
- the session down with it about 25% of the time; sliced and paced, it
367
- doesn't.
368
- - **Size equality is completeness.** A write only ever grows the file from
369
- offset zero, so an interrupted write leaves a strictly shorter file; the
370
- reader verifies the digest before the loader sees the bytes.
371
- - **The sweep roots off the process table.** An image is live for exactly as
372
- long as a process boots from it. No TTL, no eviction heuristic; after a
373
- reset the table is empty and every orphan goes.
374
-
375
- ## Binding shims for inner workers
270
+ Facets are separate isolates in one actor thread, so they scale memory but
271
+ not CPU. Awaiting I/O yields the thread; a 9,956 ms CPU burn stalled a
272
+ sibling for 9,966 ms. A peer costs about 20× the spawn time and buys real CPU
273
+ isolation, and it verifies it did not co-locate.
274
+
275
+ A facet shares its session's object, so it can copy the session's store by
276
+ reflink (`ctx.facets.clone`, 18–31 ms for 45.73 MB, 34–54 ms for 1 GB) and
277
+ shares the session's ~10 GiB budget. A peer brings its own storage and cannot
278
+ reflink.
279
+
280
+ Call clone through `cloneStorage`. An unresolvable `src` empties the
281
+ destination and reports success, so the wrapper checks the source before the
282
+ call and the destination after. An emptied facet still shows a 4,096-byte
283
+ database, so check for your own data rather than a non-zero size.
284
+
285
+ ## Image store
286
+
287
+ `ImageStore` writes generated boot images into a content-addressed store at
288
+ `var/lib/nimbus/facet-images/<sha256>.js`, through an `ImageBlobStore` port
289
+ you implement.
290
+
291
+ It registers the whole root set before writing any byte, so a sweep never
292
+ sees an unclaimed image. Writes are sliced to stay inside the 1 MiB
293
+ transaction bound; a 22.9 MB map written in one turn reset the session about
294
+ a quarter of the time. Because a write only grows the file from offset zero,
295
+ matching size means a complete write, and the reader verifies the digest
296
+ before the loader sees the bytes. Images stay live while a process boots from
297
+ them, rooted in the process table, with no TTL.
298
+
299
+ ## Binding shims
376
300
 
377
301
  `NimbusLoaderRPC`, `NimbusLoadedWorker`, `NimbusLoadedEntrypoint`,
378
302
  `NimbusAssetsRPC`, `NimbusDurableObjectNamespace`, and `NimbusDOStub` give a
379
- dynamically-loaded inner Worker working `env` bindings. They exist because of
380
- three platform behaviors, each of which cost a debugging session:
381
-
382
- - **`WorkerStub` does not serialize**, so each hop a caller makes
383
- (`load → getEntrypoint → fetch`) is its own `WorkerEntrypoint` class.
384
- - **Stubs are I/O objects bound to the request that minted them** ("Cannot
385
- perform I/O on behalf of a different request"), so the shims store CODE,
386
- never stubs, and re-resolve through `LOADER.get(id, cb)` in the current
387
- context — workerd caches by id, so repeated loads are close to free. The
388
- code map is a hard-capped LRU of 32 entries: `wrangler dev`'s
389
- rebuild-on-save loop once grew it without bound to a 128 MiB isolate crash.
390
- - **An RPC stub's method is a wildcard property**: `method.call(ep, request)`
391
- builds the pipelined path `method.call` and serializes `ep` as an argument,
392
- which workerd refuses ("Entrypoints to dynamically-loaded workers cannot be
393
- transferred"). Calls must be written `ep.method(request)`.
394
-
395
- Nesting is capped at depth 4 (`NIMBUS_INNER_LOADER_DEPTH` raises it) —
396
- Nimbus-in-Nimbus is fine, five levels is a runaway.
397
-
398
- ## The platform, measured
399
-
400
- These tables are the part of this package I most wanted to publish. They are
401
- enforced by the code above where code can enforce them; the rest is here so
402
- the next person does not have to measure them again. All figures are from
403
- production workerd, June–August 2026.
303
+ dynamically-loaded inner Worker working `env` bindings.
304
+
305
+ Three platform rules shape them. A `WorkerStub` does not serialize, so each
306
+ hop is its own `WorkerEntrypoint` class. A stub belongs to the request that
307
+ minted it, so the shims store code rather than stubs and re-resolve through
308
+ `LOADER.get(id, cb)`; workerd caches by ID, and the code map is capped at 32
309
+ entries. An RPC method is a wildcard property, so call `ep.method(request)`
310
+ and never `method.call(ep, request)`, which workerd refuses.
311
+
312
+ Nesting is capped at depth 4. Raise it with `NIMBUS_INNER_LOADER_DEPTH`.
313
+
314
+ ## Measured platform limits
315
+
316
+ Figures below come from production workerd, June to August 2026. The code
317
+ above enforces them where it can. [PLATFORM.md](PLATFORM.md) is the full
318
+ catalog: every entry dated, graded by evidence, and marked as enforced here
319
+ or left to you.
404
320
 
405
321
  ### Durable Object storage
406
322
 
@@ -409,7 +325,7 @@ production workerd, June–August 2026.
409
325
  | `await put()` resolves BEFORE durability; `ctx.storage.sync()` is the barrier; the output gate holds the guarantee | a launch killed in its first chunks left NO journal row (staging, 2026-08-13) |
410
326
  | A reset destroys every write its turn had outstanding; an alarm write rolls back with it and the platform re-delivers the alarm to the replacement instance | the first turn after a reset is a recovery turn, for free |
411
327
  | SQLite value cap is 2 MB per ROW, key length included | single-value ceiling 2,199,981 B with a 12-char key; overflow throws clean, catchable `SQLITE_TOOBIG` |
412
- | One alarm per object; a second `setAlarm()` silently overwrites | why `ALARM_REASONS_KEY` is a map |
328
+ | One alarm per object; a second `setAlarm()` silently overwrites | why `TIMER_REASONS_KEY` is a map |
413
329
  | Input gates stay closed across `get`/`put` | set-if-absent is atomic per DO with no CAS loop |
414
330
  | A facet's own SQLite survives a fresh module scope | 7,141 rows / 45.7 MB intact across recycling — keep provenance in rows, never heap |
415
331
  | ~10 GiB storage budget shared by the DO root and every facet and clone under it, with no copy-on-write credit | N clones of X bytes cost X·(N+1); crossing RESETS the object rather than raising an error |
@@ -418,11 +334,11 @@ production workerd, June–August 2026.
418
334
 
419
335
  | Invariant | Evidence |
420
336
  |---|---|
421
- | No pending alarm ⇒ hibernation-eligible after ~10 s idle | why `dispatchAlarm` deletes the map when nothing remains |
337
+ | No pending alarm ⇒ hibernation-eligible after ~10 s idle | why `timers.dispatch` deletes the map when nothing remains |
422
338
  | One-turn CPU budget ~30 s; yielding inside an invocation buys nothing; only genuine re-entry (an alarm) resets it | killed with `exceededCpu` at 31.8 s and 32.5 s |
423
339
  | A long turn drops the object's WebSockets even when the work succeeds | the launch turn finished `outcome=ok` and the terminal died anyway |
424
340
  | The in-DO clock does not advance without I/O | 0 ms across 200,000 consecutive `Time.now` reads — pace in bytes, hand deadlines to the host |
425
- | Isolate generation increments on EVERY fresh isolate: cold start and hibernation wake, not only resets | `maybeBumpIsolateGen` adopts persisted truth first |
341
+ | Isolate generation increments on EVERY fresh isolate: cold start and hibernation wake, not only resets | `adoptGeneration` adopts persisted truth first |
426
342
  | `pid <= generation base` ⇒ previous generation | THE reset predicate; `PID_GEN_STRIDE` = 1,000,000 |
427
343
  | `setTimeout`/`setInterval` prevent hibernation | one-shot self-nulling timers only |
428
344
 
@@ -437,7 +353,7 @@ production workerd, June–August 2026.
437
353
  | Module scope bans I/O; `new Function` succeeds at module scope and throws at request time | code reaches a facet through the module map or not at all |
438
354
  | The facet start callback fires at most once | re-running it would re-execute the user's program |
439
355
  | ~5–6 concurrent dynamic workers per DO; at most 4 concurrent Loader fetches per DO method; loader-cache entries are never released | `IN_DO_THRESHOLD` = 5 sits under the fetch cap; every `loader.get(id)` permanently consumes a slot — counted per DO by the loader ledger, and a cap refusal names the ids holding them |
440
- | `ctx.facets.clone` is same-object only, absent from `@cloudflare/workers-types` and the pinned workerd, present in production | 18–31 ms / 45.7 MB, 34–54 ms / 1 GB; an unresolvable `src` silently EMPTIES the destination and reports success — `cloneFacetStorage` enforces the both-ends validation |
356
+ | `ctx.facets.clone` is same-object only, absent from `@cloudflare/workers-types` and the pinned workerd, present in production | 18–31 ms / 45.7 MB, 34–54 ms / 1 GB; an unresolvable `src` silently EMPTIES the destination and reports success — `cloneStorage` enforces the both-ends validation |
441
357
  | A DO dies at ~200 MiB of live wasm linear memory; reserved and written pages die at the same ceiling | lazy growth buys nothing; bound guest memory by rewriting the memory section |
442
358
  | A wasm stack suspended (JSPI) in one request cannot resume in another | 3 in-context resumes took 6 ms; the first cross-context one hit a 30 s timeout |
443
359
 
@@ -460,7 +376,7 @@ production workerd, June–August 2026.
460
376
  |---|---|
461
377
  | Without `setWebSocketAutoResponse(ping/pong)`, every idle-tab ping wakes the actor | ~2,880 wakes/day per idle tab; the config survives hibernation |
462
378
  | A hibernatable WS owned by a DO cannot be written from a sibling `WorkerEntrypoint` isolate | sends happen in the DO's own context (relay pattern) |
463
- | A resident process never receives a WebSocket | route targets are `handleHttpRequest` only; every socket terminates on the session DO |
379
+ | A WebSocket upgrade cannot ride the RPC hop a resident's HTTP takes | a 101 owns a live socket and RPC reconstructs values rather than handing sockets over; an upgrade takes the separate fetch-semantic entrypoint and stays on `fetch` for every hop (a facet is fetched directly; a peer fetches its own facet), and a target without that entrypoint answers 501 |
464
380
 
465
381
  ### Sharing an isolate
466
382
 
@@ -470,17 +386,16 @@ production workerd, June–August 2026.
470
386
  | `exceededMemory` and `exceededCpu` are both uncatchable inside the dying isolate, observable only across an RPC boundary | absence of the error is not evidence of its absence |
471
387
  | `process.memoryUsage()` returns 0 in DO context | any heap estimate is a lower bound; say so |
472
388
 
473
- ## Relation to the other packages
389
+ ## Related packages
474
390
 
475
- `@nimbus-sh/core` is the OS this machinery hosts — fabric depends on it for
476
- shared primitives (constants, RPC disposal, error classification) and core
477
- never imports fabric.
391
+ [`@nimbus-sh/core`](https://www.npmjs.com/package/@nimbus-sh/core) is the
392
+ backend-agnostic OS: filesystem, shell, process contracts.
393
+ [`@nimbus-sh/platform`](https://www.npmjs.com/package/@nimbus-sh/platform)
394
+ holds the limits tables, the error taxonomy, RPC disposal, and the supervisor
395
+ budget machinery.
478
396
  [`@nimbus-sh/worker`](https://www.npmjs.com/package/@nimbus-sh/worker) is the
479
- canonical embedder: it supplies the seams above, the supervisor entrypoint,
480
- the session protocol, and everything user-facing. If you want the full hosted
481
- product shape, start from `npx create-nimbus-app`; if you are building your
482
- own thing on Durable Objects, this package and its doc comments are the part
483
- of Nimbus you can take without taking Nimbus.
397
+ reference embedder, with the supervisor entrypoint and the session protocol.
398
+ For the full hosted product, start from `npx create-nimbus-app`.
484
399
 
485
400
  ## License
486
401