space-data-module-sdk 0.8.23 → 0.8.25
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/docs/isomorphic-pthreads.html +29 -5
- package/docs/isomorphic-pthreads.md +98 -8
- package/package.json +1 -1
- package/src/compiler/compileModule.js +378 -385
- package/src/compiler/emceptionNode.js +339 -70
- package/src/host/browserModuleHarness.js +13 -0
- package/src/host/hostWorkerBundleSources.js +3 -3
- package/src/host/nodeSyncFsIo.js +3 -1
- package/src/host/wasiThreadBrowserWorker.mjs +41 -21
- package/src/host/wasiThreadHost.js +279 -150
- package/src/host/wasiThreadPool.js +389 -0
- package/src/host/wasiThreadWorker.mjs +62 -20
- package/src/host/wasiThreadWorkerRuntime.js +6 -1
- package/src/index.d.ts +31 -0
- package/src/testing/index.d.ts +2 -0
|
@@ -156,6 +156,27 @@ node --test test/direct-call-constructors.test.js</code></pre></div>
|
|
|
156
156
|
<h3 id="guest-thread-faults"><a class="anchor" href="#guest-thread-faults" aria-hidden="true">#</a>Guest thread faults</h3>
|
|
157
157
|
<p>A guest thread that traps never finishes the pthread exit protocol. WasmEdge cancels the whole command. In the browser and Node harnesses the joining thread stays blocked inside the guest and cannot run the worker's error event, so the worker itself writes <code>[wasi-thread] guest thread N trapped: ...</code> to stderr (Node) or the console (browser). The call still does not return.</p>
|
|
158
158
|
<p>V8 (Node 20 to 25) checks bulk memory operations, and every access when the WebAssembly trap handler is off (Node on Linux arm64), against a per-instance copy of a shared memory's size. That copy is refreshed asynchronously after another thread grows the memory, so a thread that writes into memory another thread has just grown can trap with "memory access out of bounds", even after synchronizing with the growing thread. The guest's allocator uses the heap the artifact was linked with and then grows the memory, so a larger imported initial memory does not prevent it.</p>
|
|
159
|
+
<h3 id="sdk-0825-one-pool-protocol-threads-reused-within-an-invoke-spawns-from-any-thread"><a class="anchor" href="#sdk-0825-one-pool-protocol-threads-reused-within-an-invoke-spawns-from-any-thread" aria-hidden="true">#</a>SDK 0.8.25: one pool protocol; threads reused within an invoke; spawns from any thread</h3>
|
|
160
|
+
<p>A guest runs <code>pthread_create</code> through <code>pthread_join</code> without yielding the spawning thread's event loop, often for a whole invoke that spawns wave after wave of threads (conjunction screening spawns a coarse wave, then a refine wave), and any guest thread may call <code>pthread_create</code>. Through 0.8.24:</p>
|
|
161
|
+
<ul>
|
|
162
|
+
<li>a pooled browser worker was sent each thread by message and went back to idle only when the spawner handled its <code>{t:"exit"}</code> message, which could not happen until the invoke returned. One invoke could spawn at most <code>poolSize</code> threads in total, however few ran at once, and a guest that spawned more got <code>EAGAIN</code> (<code>std::thread</code> aborts with <code>unreachable</code>);</li>
|
|
163
|
+
<li>Node started a worker per spawn. Its live-thread count and the join of a finished worker waited for <code>exit</code> events on the blocked event loop, so an explicit <code>poolSize</code> had the same limit, and a guest that started thousands of threads kept thousands of finished workers;</li>
|
|
164
|
+
<li>a spawn from a guest thread other than the main thread always returned -1.</li>
|
|
165
|
+
</ul>
|
|
166
|
+
<p>From 0.8.25 both runtimes share one pool protocol (<code>src/host/wasiThreadPool.js</code>), a <code>SharedArrayBuffer</code> every thread of the process holds: a tid counter, spawn counters, and per worker a slot (<code>PENDING</code>, <code>IDLE</code>, <code>CLAIMED</code>, <code>ASSIGNED</code>, <code>RUNNING</code>, <code>RETIRED</code>, plus the tid and start argument). A pool worker blocks on its slot between threads. A spawner on any thread claims an <code>IDLE</code> slot with a compare-exchange, writes the tid and argument, marks it <code>ASSIGNED</code> and notifies it. The worker runs <code>wasi_thread_start</code>, then marks the slot <code>IDLE</code> and bumps a release generation that waiting spawners sleep on. No step needs an event loop, so:</p>
|
|
167
|
+
<ul>
|
|
168
|
+
<li>a pool bounds how many threads run <strong>at the same time</strong>, not how many one invoke may start (measured: 14,000 threads in one invoke ran on 7 Node workers);</li>
|
|
169
|
+
<li>every guest thread's own <code>wasi.thread-spawn</code> is the pool's, so a guest thread can start threads of its own;</li>
|
|
170
|
+
<li>Node reuses its workers. With no <code>poolSize</code> it grows the pool to at most 1024 workers; with one, to <code>poolSize</code>. A spawn that finds every worker busy waits up to 2 ms for one to come free, then starts a new worker from the spawning thread, which may itself be a pool worker (a Node worker starts without its parent's event loop). Idle workers do not keep the process alive;</li>
|
|
171
|
+
<li>the browser keeps its pre-started warm pool (a browser worker cannot start while its parent is blocked). Its workers now serve their slots and never return to their event loop after the probe.</li>
|
|
172
|
+
</ul>
|
|
173
|
+
<p>A joined thread's worker is still a few instructions from returning when the joiner wakes, so a wave spawned right after a join can find the pool busy for a moment (measured in headless Chromium: 0 to 30 µs). A spawn that finds every thread busy and cannot grow the pool therefore waits on the generation for up to <code>spawnWaitMs</code> (default 250 ms; harness option <code>wasiThreadSpawnWaitMs</code>) before it is declined. After one wait runs out, later spawns from that thread are declined at once until some thread finishes, so a pool held by long-running threads costs one wait, not one per spawn. The wait uses <code>Atomics.wait</code>, which workers and Node allow; where it is not allowed (a window's main thread) a spawn never waits.</p>
|
|
174
|
+
<p><code>spawnReport()</code> counts spawns from every thread (<code>spawned</code>, <code>waited</code>, and <code>pool-exhausted</code> declines); <code>onSpawnDeclined</code> fires for the owning thread's spawns. A Node worker started by a guest thread reports a guest fault on stderr only; <code>onGuestError</code> covers workers the owning thread started and every browser worker.</p>
|
|
175
|
+
<p>A browser worker that answers the probe without the pool protocol (a worker script from an older SDK) still works: the owning thread sends it its threads by <code>{t:"run"}</code> and frees it on <code>{t:"exit"}</code>, as before, and guest threads never spawn onto it. A worker served from the matching SDK gets the fix.</p>
|
|
176
|
+
<p><code>test/wasi-thread-pool-reuse-guest.test.js</code> compiles a guest that spawns 3 waves of <code>poolSize - 1</code> threads (the main thread runs one stripe), 3 waves of <code>poolSize</code> threads, and 3 waves whose threads a non-main guest thread spawns, in one invoke. It checks the output bytes and spawn counts in Node (explicit pool, no pool, and 14,000 threads on at most 16 workers), in headless Chromium, Firefox and WebKit, and across the three parity lanes:</p>
|
|
177
|
+
<div class="codeblock"><div class="codeblock-head">sh</div><pre><code>SPACE_DATA_MODULE_SDK_ENABLE_BROWSER_THREADS=1 \
|
|
178
|
+
SPACE_DATA_MODULE_SDK_ENABLE_TRI_RUNTIME_PARITY=1 \
|
|
179
|
+
node --test test/wasi-thread-pool-reuse-guest.test.js</code></pre></div>
|
|
159
180
|
<p>The old source path <code>src/testing/browserModuleHarness.js</code> remains a pure compatibility re-export. New browser consumers should use the public <code>space-data-module-sdk/host/browser-module</code> entry point.</p>
|
|
160
181
|
<h2 id="4-integrators-the-browser-worker-anchor-required-when-you-bundle"><a class="anchor" href="#4-integrators-the-browser-worker-anchor-required-when-you-bundle" aria-hidden="true">#</a>4. Integrators: the browser worker anchor (REQUIRED when you bundle)</h2>
|
|
161
182
|
<p>The browser leg of the wasi-threads host runs each guest pthread on a pooled module <code>Worker</code>. That worker is a <strong>served asset</strong>, and the SDK cannot guess where your build published it.</p>
|
|
@@ -177,7 +198,7 @@ setBrowserWasiThreadWorkerBase("js/vendor/space-data-module-sdk/src/host/&q
|
|
|
177
198
|
<div class="table-wrap"><table><thead><tr><th>Source</th><th>Option / API</th></tr></thead><tbody><tr><td>1. Explicit worker file</td><td><code>browserWorkerUrl</code> / <code>wasiThreadWorkerUrl</code></td></tr><tr><td>2. Explicit directory</td><td><code>browserWorkerBaseUrl</code> / <code>wasiThreadWorkerBaseUrl</code></td></tr><tr><td>3. Process-wide base</td><td><code>setBrowserWasiThreadWorkerBase(base)</code></td></tr><tr><td>4. Packaged sibling</td><td><code>new URL("./wasiThreadBrowserWorker.mjs", import.meta.url)</code></td></tr></tbody></table></div>
|
|
178
199
|
<p>Rules that make this contract honest:</p>
|
|
179
200
|
<ul>
|
|
180
|
-
<li><strong>The anchor names a DIRECTORY that serves the WHOLE chain.</strong> <code>wasiThreadBrowserWorker.mjs</code> imports <code>./wasiThreadWorkerRuntime.js</code
|
|
201
|
+
<li><strong>The anchor names a DIRECTORY that serves the WHOLE chain.</strong> <code>wasiThreadBrowserWorker.mjs</code> imports <code>./wasiThreadWorkerRuntime.js</code> and <code>./wasiThreadPool.js</code>, which import their own siblings. Staging the single <code>.mjs</code> next to your bundle does <strong>not</strong> work.</li>
|
|
181
202
|
<li><strong>A relative base resolves against the document</strong>; absolute URLs pass through unchanged.</li>
|
|
182
203
|
<li><strong>No consumer-side <code>location</code> sniffing.</strong> Forking worker resolution per consumer is not the contract; you already know your build's layout, so state it.</li>
|
|
183
204
|
<li><strong>No fetch-and-retry probe.</strong> Resolution never touches the network: Node never fetches, and browser thread count may not become a function of network timing.</li>
|
|
@@ -188,8 +209,8 @@ setBrowserWasiThreadWorkerBase("js/vendor/space-data-module-sdk/src/host/&q
|
|
|
188
209
|
<p>The FlatSQL partition store (stack design <code>docs/architecture/flatsql-partition-store.md</code>, §5.5, §5.6, §18 T9, A7, A36, A38, A39) runs one wasi-threads artifact as a writer instance and reader instances, each with guest threads that call FlatSQL's seven <code>env.flatsql_io_*</code> imports. This section is the SDK half: the thread pool, the I/O channel and workers, the Node provider, link shim v2, the worker bundles and the browser capability probe.</p>
|
|
189
210
|
<h3 id="explicit-pool-size-partial-spawns-supervision-hooks"><a class="anchor" href="#explicit-pool-size-partial-spawns-supervision-hooks" aria-hidden="true">#</a>Explicit pool size, partial spawns, supervision hooks</h3>
|
|
190
211
|
<p><code>createWasiThreadSpawn</code> options added in 0.8.22:</p>
|
|
191
|
-
<div class="table-wrap"><table><thead><tr><th>Option</th><th>Effect</th></tr></thead><tbody><tr><td><code>poolSize</code></td><td>Browser: pre-start exactly this many workers, independent of <code>hardwareConcurrency - 1</code> (pools are sized for isolation: writers + lanes). Node: cap on live guest threads. Arming is partial: workers that fail to start are dropped, the rest serve.</td></tr><tr><td><code>extraImports</code></td><td>Per-worker import objects, as structured-cloneable descriptors (below). Factories run once per worker.</td></tr><tr><td><code>instanceId</code>, <code>onGuestError(instanceId, tid, error)</code></td><td>Called when a guest thread traps or its worker dies (A36). A dead pooled worker leaves the pool.</td></tr><tr><td><code>onSpawnDeclined({ reason, poolSize, declined })</code></td><td>Called for every spawn that returns -1. Reasons: <code>pool-exhausted</code>, <code>pool-not-armed</code>, <code>threads-unavailable</code>, <code>pool-empty</code>, <code>worker-create-failed</code>, <code>dispatch-failed</code>, <code>hostcall-channel-missing</code>, <code>terminated</code>.</td></tr><tr><td><code>probeTimeoutMs</code></td><td>Browser warm-pool probe deadline.</td></tr><tr><td><code>browserWorkerType: "classic"</code></td><td>Spawn classic workers, for the blob bundles below.</td></tr></tbody></table></div>
|
|
192
|
-
<p>The returned host adds <code>spawnReport()</code>: <code>{ poolSize, armed, failedToArm, spawned, declined, declinedByReason, lastDeclineReason, active, idle }</code
|
|
212
|
+
<div class="table-wrap"><table><thead><tr><th>Option</th><th>Effect</th></tr></thead><tbody><tr><td><code>poolSize</code></td><td>Browser: pre-start exactly this many workers, independent of <code>hardwareConcurrency - 1</code> (pools are sized for isolation: writers + lanes). Node: cap on live guest threads (the pool grows to it on demand; without it, to 1024). Arming is partial: workers that fail to start are dropped, the rest serve.</td></tr><tr><td><code>extraImports</code></td><td>Per-worker import objects, as structured-cloneable descriptors (below). Factories run once per worker.</td></tr><tr><td><code>instanceId</code>, <code>onGuestError(instanceId, tid, error)</code></td><td>Called when a guest thread traps or its worker dies (A36). A dead pooled worker leaves the pool. In Node, for workers the owning thread started.</td></tr><tr><td><code>onSpawnDeclined({ reason, poolSize, declined })</code></td><td>Called for every spawn from the owning thread that returns -1. Reasons: <code>pool-exhausted</code>, <code>pool-not-armed</code>, <code>threads-unavailable</code>, <code>pool-empty</code>, <code>worker-create-failed</code>, <code>dispatch-failed</code>, <code>hostcall-channel-missing</code>, <code>terminated</code>.</td></tr><tr><td><code>probeTimeoutMs</code></td><td>Browser warm-pool probe deadline.</td></tr><tr><td><code>spawnWaitMs</code></td><td>How long a spawn that finds every pool thread busy waits for one to finish before it returns -1 (default 250; 0 never waits). Added in 0.8.25; see "one pool protocol" in §3.</td></tr><tr><td><code>browserWorkerType: "classic"</code></td><td>Spawn classic workers, for the blob bundles below.</td></tr></tbody></table></div>
|
|
213
|
+
<p>The returned host adds <code>spawnReport()</code>: <code>{ poolSize, armed, failedToArm, spawned, waited, declined, declinedByReason, lastDeclineReason, active, idle }</code> (Node adds <code>workers</code>). The counts cover spawns from every thread of the process. The implicit (no <code>poolSize</code>) path keeps its all-or-nothing arming.</p>
|
|
193
214
|
<p><code>extraImports</code> entries (the same descriptors work in the engine worker through <code>resolveExtraImports</code>):</p>
|
|
194
215
|
<div class="table-wrap"><table><thead><tr><th>Descriptor</th><th>Imports</th></tr></thead><tbody><tr><td><code>{ provider: "flatsql-io", instanceId, channels, mirror?, trace? }</code></td><td><code>env.flatsql_io_*</code> over one SAB I/O channel per I/O worker. Each worker claims its own request slot.</td></tr><tr><td><code>{ provider: "flatsql-io-node", root, table, instanceId }</code></td><td><code>env.flatsql_io_*</code> over synchronous <code>fs</code> (Node workers).</td></tr><tr><td><code>{ moduleUrl, exportName?, config? }</code></td><td>A factory module (module workers and Node only).</td></tr></tbody></table></div>
|
|
195
216
|
<p>A module that bundles this host into an IIFE or a blob: module worker has no usable <code>import.meta.url</code>. <code>DEFAULT_BROWSER_WORKER_URL</code> is then <code>null</code> instead of a module-evaluation <code>TypeError</code>, and such hosts pass <code>browserWorkerUrl</code>.</p>
|
|
@@ -216,7 +237,7 @@ setBrowserWasiThreadWorkerBase("js/vendor/space-data-module-sdk/src/host/&q
|
|
|
216
237
|
<p>Store lock (A37): with <code>lock: { name }</code> the I/O worker takes that Web Lock before it opens any handle and releases it only after <code>stop()</code> has closed them, so the lock lives exactly as long as the handles. <code>ifAvailable: true</code> fails the start with <code>lockUnavailable</code> instead of waiting; <code>busyRetry</code> retries a handle a previous leader still holds, with backoff. Measured takeover (holder <code>stop()</code> to successor ready): 5 ms Chromium, 52 ms Firefox, 5 ms WebKit. Leadership, follower proxying and heartbeats are the engine's (sdn-js, T10).</p>
|
|
217
238
|
<p>Head mirror (A7): with <code>mirror: { buffer, suffixes: ["/h.fsh"] }</code>, the writer I/O worker copies every write to a matching path into a seqlock mirror in a SharedArrayBuffer (<code>sabIoMirror.js</code>), and reader imports configured with the same mirror serve reads of those paths from it without a round trip.</p>
|
|
218
239
|
<h3 id="node-synchronous-fs-56"><a class="anchor" href="#node-synchronous-fs-56" aria-hidden="true">#</a>Node synchronous fs (§5.6)</h3>
|
|
219
|
-
<p><code>nodeSyncFsIo.js</code>: synchronous <code>fs</code> in each worker over a shared virtual-handle table in a SharedArrayBuffer (<code>createNodeSyncFsIoTable</code>). A handle is <code>(slot << 8) | gen</code>; each worker opens its own fd for a slot on first use and drops stale fds when the generation moves. Paths are confined below <code>root</code>, including through symlinked parents. <code>sync</code> is <code>fdatasync</code> (libuv issues <code>F_FULLFSYNC</code> on darwin). <code>revokeNodeSyncFsIoInstance</code> follows A23: it sets the revoked flag and waits for the instance's in-flight calls to drain. Fault injection (§19, 22.3a-7) belongs to FlatSQL's Node host (T4); <code>interpose</code> wraps every syscall for it.</p>
|
|
240
|
+
<p><code>nodeSyncFsIo.js</code>: synchronous <code>fs</code> in each worker over a shared virtual-handle table in a SharedArrayBuffer (<code>createNodeSyncFsIoTable</code>). A handle is <code>(slot << 8) | gen</code>; each worker opens its own fd for a slot on first use and drops stale fds when the generation moves. Paths are confined below <code>root</code>, including through symlinked parents; a filesystem root (<code>root: "/"</code>) reaches every path under it (through 0.8.24 every such path was <code>ACCESS</code>). <code>sync</code> is <code>fdatasync</code> (libuv issues <code>F_FULLFSYNC</code> on darwin). <code>revokeNodeSyncFsIoInstance</code> follows A23: it sets the revoked flag and waits for the instance's in-flight calls to drain. Fault injection (§19, 22.3a-7) belongs to FlatSQL's Node host (T4); <code>interpose</code> wraps every syscall for it.</p>
|
|
220
241
|
<h3 id="link-shim-v2"><a class="anchor" href="#link-shim-v2" aria-hidden="true">#</a>Link shim v2</h3>
|
|
221
242
|
<p><code>FLATSQL_LINK_SHIM_V2_WASM</code> (<code>src/flow/flatsqlLinkShim.js</code>) is a deterministic module that imports the reader instance's SHARED lane memory as <code>flatsql.memory</code> and gives a linked flow the lane mailbox as direct calls. v1 is unchanged (sha256 <code>8d83e69b…</code>), because today's linked flows use it. v2 sha256: <code>67d5b2d9bc2d1b346a14a253a586fd4d08c8056d54eb62701b48000585ed9613</code>.</p>
|
|
222
243
|
<div class="table-wrap"><table><thead><tr><th>Export</th><th>Result</th></tr></thead><tbody><tr><td><code>mb_submit(mailbox, op, req_ptr, req_len)</code></td><td><code>seq</code>, or -1 when the mailbox holds a request</td></tr><tr><td><code>mb_poll(mailbox, seq)</code></td><td>1 when complete</td></tr><tr><td><code>mb_wait(mailbox, seq, poll_ns: i64, max_polls)</code></td><td>the lane's status, or -110 after <code>max_polls</code> bounded waits (<code>max_polls <= 0</code>: no limit)</td></tr><tr><td><code>mb_release(mailbox, seq)</code></td><td>0, or -1 when not complete</td></tr><tr><td><code>mb_cancel(mailbox, seq)</code></td><td>0; the lane polls the cancel word</td></tr><tr><td><code>load32_acquire</code>, <code>store32_release</code>, <code>peek8/32/64</code>, <code>poke8/32</code>, <code>fnv1a64</code>, <code>count_frames</code></td><td>as in v1, over lane memory</td></tr></tbody></table></div>
|
|
@@ -248,7 +269,7 @@ const pool = await createWasiThreadSpawn({
|
|
|
248
269
|
<p><code>runFlatsqlIoConformance(io)</code> (<code>flatsqlIoConformance.js</code>) is one script for every host (22.3a-6): 15 cases covering statuses, short reads, sparse writes, truncation, EXCL/TRUNC/PROBE/UNLINK/UNLINK_IF_UNUSED/CREATE_PARENTS/ DELETE_ON_CLOSE/OPEN_DEFERRED, access modes, confinement and multi-handle visibility. It passes on the Node sync-fs provider, on the channel over the memory backend (blocking and async clients, both doorbells), on the blob bundle run as a standalone script, and on OPFS and memory in Chromium, Firefox and WebKit. Plain <code>UNLINK</code> of an open path is <code>BUSY</code> in the I/O worker (OPFS cannot remove a file with an open sync handle) and succeeds on POSIX hosts; the script does not test it.</p>
|
|
249
270
|
<h3 id="measured-acceptance-18-t9-a7-a38"><a class="anchor" href="#measured-acceptance-18-t9-a7-a38" aria-hidden="true">#</a>Measured acceptance (§18 T9, A7, A38)</h3>
|
|
250
271
|
<p>Owner's Mac Studio (Apple M3 Ultra, 28 cores, macOS 26.3.1), 2026-09-27, headless, final full run of <code>test/opfs-io-worker.browser.test.js</code>. The machine was shared with other lanes (load average 23-33 on 28 cores). Latencies are guest-observed import calls (<code>trace</code>), in microseconds.</p>
|
|
251
|
-
<div class="table-wrap"><table><thead><tr><th>Item</th><th>Chromium 153</th><th>Firefox 155</th><th>WebKit 26.6</th></tr></thead><tbody><tr><td>#1 8 threads, mixed I/O + 200 async opens + 50 unlinks through 1 I/O worker: errors / lost writes</td><td>0 / 0</td><td>0 / 0</td><td>0 / 0</td></tr><tr><td>same run, 4 KiB read p50 / p99</td><td>50 / 2145</td><td>260 / 8180</td><td>60 / 2660</td></tr><tr><td>same run, message doorbell: errors / lost writes</td><td>0 / 0</td><td>0 / 0</td><td>0 / 0</td></tr><tr><td>A7 lane 4 KiB read p99 during 100 x 4 MiB write+flush, same partition</td><td>790 (active file, shared mode)</td><td>120 (sealed segment)</td><td>40 (sealed segment)</td></tr><tr><td>A7 same, other partition</td><td>45</td><td>120</td><td>40</td></tr><tr><td>A7 writer flush p50 / p99 (4 MiB)</td><td>6365 / 28945</td><td>3540 / 27740</td><td>5260 / 70700</td></tr><tr><td>#2 500 ms open, median of 3 runs: other threads' read p99, baseline / during</td><td>85 / 85</td><td>440 / 400</td><td>60 / 60</td></tr><tr><td>#2 longest other-thread read while the open was pending</td><td>1020</td><td>18500</td><td>320</td></tr><tr><td>A38 OPEN_DEFERRED: open returns in / first write waits (ms)</td><td>3.8 / 504</td><td>28.8 / 511</td><td>5.2 / 755</td></tr><tr><td>#3 <code>poolSize=6</code> on <code>hardwareConcurrency=2</code
|
|
272
|
+
<div class="table-wrap"><table><thead><tr><th>Item</th><th>Chromium 153</th><th>Firefox 155</th><th>WebKit 26.6</th></tr></thead><tbody><tr><td>#1 8 threads, mixed I/O + 200 async opens + 50 unlinks through 1 I/O worker: errors / lost writes</td><td>0 / 0</td><td>0 / 0</td><td>0 / 0</td></tr><tr><td>same run, 4 KiB read p50 / p99</td><td>50 / 2145</td><td>260 / 8180</td><td>60 / 2660</td></tr><tr><td>same run, message doorbell: errors / lost writes</td><td>0 / 0</td><td>0 / 0</td><td>0 / 0</td></tr><tr><td>A7 lane 4 KiB read p99 during 100 x 4 MiB write+flush, same partition</td><td>790 (active file, shared mode)</td><td>120 (sealed segment)</td><td>40 (sealed segment)</td></tr><tr><td>A7 same, other partition</td><td>45</td><td>120</td><td>40</td></tr><tr><td>A7 writer flush p50 / p99 (4 MiB)</td><td>6365 / 28945</td><td>3540 / 27740</td><td>5260 / 70700</td></tr><tr><td>#2 500 ms open, median of 3 runs: other threads' read p99, baseline / during</td><td>85 / 85</td><td>440 / 400</td><td>60 / 60</td></tr><tr><td>#2 longest other-thread read while the open was pending</td><td>1020</td><td>18500</td><td>320</td></tr><tr><td>A38 OPEN_DEFERRED: open returns in / first write waits (ms)</td><td>3.8 / 504</td><td>28.8 / 511</td><td>5.2 / 755</td></tr><tr><td>#3 <code>poolSize=6</code> on <code>hardwareConcurrency=2</code> (<code>spawnWaitMs: 0</code> since 0.8.25)</td><td>6 workers; 7th spawn -1, reported <code>pool-exhausted</code></td><td>same</td><td>same</td></tr><tr><td>#5 link shim v2 (Node 25): 10,000 calls, lost completions</td><td>0 with a notifying lane, 0 with a lane that never notifies, 0 spinning (<code>poll_ns = 0</code>)</td><td></td><td></td></tr></tbody></table></div>
|
|
252
273
|
<p>Chromium's same-partition A7 p99 varies with host load: 90, 110, 295 and 790 over four runs with the adaptive spin (below), and 125-1165 over four runs before it. The original #1 target (4 KiB read p99 <= 200 µs under the mixed load) is not met through one I/O worker; A7 replaced #1 with the <= 1 ms lane target, which is met. Firefox and WebKit have no shared handle modes, so their lanes read a partition's sealed segment; per §22.4-6 those browsers get no local store.</p>
|
|
253
274
|
<p>Adaptive waits: an idle I/O worker polls its doorbell for 50 µs, and a blocking client polls for its answer for 20 µs, before sleeping (<code>spinMicros</code>), so back-to-back requests skip a thread wake-up. In the Node channel test this took the fast-thread read p50 from 30-39 µs to 11-12 µs.</p>
|
|
254
275
|
<h2 id="tests"><a class="anchor" href="#tests" aria-hidden="true">#</a>Tests</h2>
|
|
@@ -259,6 +280,8 @@ const pool = await createWasiThreadSpawn({
|
|
|
259
280
|
<li><code>test/guest-link-symbol-prefix.test.js</code> — the guest-link prefix is the full injective hex of the pluginId; two previously-colliding ids now get distinct prefixes; fresh prefixes match the committed modules-branch artifacts byte-for-byte; and the compose path treats the guest-link metadata's <code>symbolPrefix</code> / <code>methodSymbols</code> as authoritative (never re-derived), keeping a legacy truncated-prefix artifact compatible.</li>
|
|
260
281
|
</ul>
|
|
261
282
|
<ul>
|
|
283
|
+
<li><code>test/wasi-thread-pool-reuse.test.js</code> — the pool protocol with no message and no event-loop turn: 3 waves of <code>poolSize - 1</code> and of <code>poolSize</code> spawns, a spawn that blocks until a thread finishes on another thread, a spawn from a thread that is not the owner, the <code>spawnWaitMs</code> bound and the no-second-wait rule, an older worker script and a late <code>{t:"exit"}</code>, a dead worker, <code>terminateAll</code>.</li>
|
|
284
|
+
<li><code>test/wasi-thread-pool-reuse-guest.test.js</code> — the real wave guest in Node, headless Chromium, Firefox and WebKit, and the three parity lanes (the browser and parity runs are env-gated, see §3).</li>
|
|
262
285
|
<li><code>test/wasi-thread-pool-size.test.js</code> — explicit <code>poolSize</code> (6 workers on <code>hardwareConcurrency=2</code>, the 7th spawn -1 and reported), partial arming, extraImports delivery, <code>onGuestError</code>, classic workers.</li>
|
|
263
286
|
<li><code>test/sab-io-channel.test.js</code> — the I/O channel under a real wasi-threads guest (<code>test/support/flatsql-io/ioGuestWasm.mjs</code>): 8 threads of mixed I/O in both doorbell modes, a 500 ms open blocking only its caller, OPEN_DEFERRED, revocation, a dead I/O worker, supervisor restart, 256 KiB steps, the head mirror, scratch vs direct transfers, pre-open, and path-hash routing.</li>
|
|
264
287
|
<li><code>test/node-sync-fs-io.test.js</code> — shared handles across workers, stale handles, CREATE_PARENTS, UNLINK_IF_UNUSED, confinement, A23 revocation.</li>
|
|
@@ -294,6 +317,7 @@ const pool = await createWasiThreadSpawn({
|
|
|
294
317
|
<li><a class="depth-3" href="#sdk-0820-command-hosts">SDK 0.8.20 command hosts</a></li>
|
|
295
318
|
<li><a class="depth-3" href="#sdk-0821-constructors-on-the-direct-surface">SDK 0.8.21: constructors on the direct surface</a></li>
|
|
296
319
|
<li><a class="depth-3" href="#guest-thread-faults">Guest thread faults</a></li>
|
|
320
|
+
<li><a class="depth-3" href="#sdk-0825-one-pool-protocol-threads-reused-within-an-invoke-spawns-from-any-thread">SDK 0.8.25: one pool protocol; threads reused within an invoke; spawns from any thread</a></li>
|
|
297
321
|
<li><a class="depth-2" href="#4-integrators-the-browser-worker-anchor-required-when-you-bundle">4. Integrators: the browser worker anchor (REQUIRED when you bundle)</a></li>
|
|
298
322
|
<li><a class="depth-2" href="#5-flatsql-partition-store-host-io-sdk-0822">5. FlatSQL partition store host I/O (SDK 0.8.22)</a></li>
|
|
299
323
|
<li><a class="depth-3" href="#explicit-pool-size-partial-spawns-supervision-hooks">Explicit pool size, partial spawns, supervision hooks</a></li>
|
|
@@ -261,6 +261,83 @@ synchronizing with the growing thread. The guest's allocator uses the heap the
|
|
|
261
261
|
artifact was linked with and then grows the memory, so a larger imported initial
|
|
262
262
|
memory does not prevent it.
|
|
263
263
|
|
|
264
|
+
### SDK 0.8.25: one pool protocol; threads reused within an invoke; spawns from any thread
|
|
265
|
+
|
|
266
|
+
A guest runs `pthread_create` through `pthread_join` without yielding the
|
|
267
|
+
spawning thread's event loop, often for a whole invoke that spawns wave after
|
|
268
|
+
wave of threads (conjunction screening spawns a coarse wave, then a refine
|
|
269
|
+
wave), and any guest thread may call `pthread_create`. Through 0.8.24:
|
|
270
|
+
|
|
271
|
+
- a pooled browser worker was sent each thread by message and went back to
|
|
272
|
+
idle only when the spawner handled its `{t:"exit"}` message, which could not
|
|
273
|
+
happen until the invoke returned. One invoke could spawn at most `poolSize`
|
|
274
|
+
threads in total, however few ran at once, and a guest that spawned more got
|
|
275
|
+
`EAGAIN` (`std::thread` aborts with `unreachable`);
|
|
276
|
+
- Node started a worker per spawn. Its live-thread count and the join of a
|
|
277
|
+
finished worker waited for `exit` events on the blocked event loop, so an
|
|
278
|
+
explicit `poolSize` had the same limit, and a guest that started thousands of
|
|
279
|
+
threads kept thousands of finished workers;
|
|
280
|
+
- a spawn from a guest thread other than the main thread always returned -1.
|
|
281
|
+
|
|
282
|
+
From 0.8.25 both runtimes share one pool protocol (`src/host/wasiThreadPool.js`),
|
|
283
|
+
a `SharedArrayBuffer` every thread of the process holds: a tid counter, spawn
|
|
284
|
+
counters, and per worker a slot (`PENDING`, `IDLE`, `CLAIMED`, `ASSIGNED`,
|
|
285
|
+
`RUNNING`, `RETIRED`, plus the tid and start argument). A pool worker blocks on
|
|
286
|
+
its slot between threads. A spawner on any thread claims an `IDLE` slot with a
|
|
287
|
+
compare-exchange, writes the tid and argument, marks it `ASSIGNED` and notifies
|
|
288
|
+
it. The worker runs `wasi_thread_start`, then marks the slot `IDLE` and bumps a
|
|
289
|
+
release generation that waiting spawners sleep on. No step needs an event loop,
|
|
290
|
+
so:
|
|
291
|
+
|
|
292
|
+
- a pool bounds how many threads run **at the same time**, not how many one
|
|
293
|
+
invoke may start (measured: 14,000 threads in one invoke ran on 7 Node
|
|
294
|
+
workers);
|
|
295
|
+
- every guest thread's own `wasi.thread-spawn` is the pool's, so a guest thread
|
|
296
|
+
can start threads of its own;
|
|
297
|
+
- Node reuses its workers. With no `poolSize` it grows the pool to at most 1024
|
|
298
|
+
workers; with one, to `poolSize`. A spawn that finds every worker busy waits
|
|
299
|
+
up to 2 ms for one to come free, then starts a new worker from the spawning
|
|
300
|
+
thread, which may itself be a pool worker (a Node worker starts without its
|
|
301
|
+
parent's event loop). Idle workers do not keep the process alive;
|
|
302
|
+
- the browser keeps its pre-started warm pool (a browser worker cannot start
|
|
303
|
+
while its parent is blocked). Its workers now serve their slots and never
|
|
304
|
+
return to their event loop after the probe.
|
|
305
|
+
|
|
306
|
+
A joined thread's worker is still a few instructions from returning when the
|
|
307
|
+
joiner wakes, so a wave spawned right after a join can find the pool busy for a
|
|
308
|
+
moment (measured in headless Chromium: 0 to 30 µs). A spawn that finds every
|
|
309
|
+
thread busy and cannot grow the pool therefore waits on the generation for up
|
|
310
|
+
to `spawnWaitMs` (default 250 ms; harness option `wasiThreadSpawnWaitMs`) before
|
|
311
|
+
it is declined. After one wait runs out, later spawns from that thread are
|
|
312
|
+
declined at once until some thread finishes, so a pool held by long-running
|
|
313
|
+
threads costs one wait, not one per spawn. The wait uses `Atomics.wait`, which
|
|
314
|
+
workers and Node allow; where it is not allowed (a window's main thread) a spawn
|
|
315
|
+
never waits.
|
|
316
|
+
|
|
317
|
+
`spawnReport()` counts spawns from every thread (`spawned`, `waited`, and
|
|
318
|
+
`pool-exhausted` declines); `onSpawnDeclined` fires for the owning thread's
|
|
319
|
+
spawns. A Node worker started by a guest thread reports a guest fault on stderr
|
|
320
|
+
only; `onGuestError` covers workers the owning thread started and every browser
|
|
321
|
+
worker.
|
|
322
|
+
|
|
323
|
+
A browser worker that answers the probe without the pool protocol (a worker
|
|
324
|
+
script from an older SDK) still works: the owning thread sends it its threads by
|
|
325
|
+
`{t:"run"}` and frees it on `{t:"exit"}`, as before, and guest threads never
|
|
326
|
+
spawn onto it. A worker served from the matching SDK gets the fix.
|
|
327
|
+
|
|
328
|
+
`test/wasi-thread-pool-reuse-guest.test.js` compiles a guest that spawns 3
|
|
329
|
+
waves of `poolSize - 1` threads (the main thread runs one stripe), 3 waves of
|
|
330
|
+
`poolSize` threads, and 3 waves whose threads a non-main guest thread spawns, in
|
|
331
|
+
one invoke. It checks the output bytes and spawn counts in Node (explicit pool,
|
|
332
|
+
no pool, and 14,000 threads on at most 16 workers), in headless Chromium,
|
|
333
|
+
Firefox and WebKit, and across the three parity lanes:
|
|
334
|
+
|
|
335
|
+
```sh
|
|
336
|
+
SPACE_DATA_MODULE_SDK_ENABLE_BROWSER_THREADS=1 \
|
|
337
|
+
SPACE_DATA_MODULE_SDK_ENABLE_TRI_RUNTIME_PARITY=1 \
|
|
338
|
+
node --test test/wasi-thread-pool-reuse-guest.test.js
|
|
339
|
+
```
|
|
340
|
+
|
|
264
341
|
The old source path `src/testing/browserModuleHarness.js` remains a pure
|
|
265
342
|
compatibility re-export. New browser consumers should use the public
|
|
266
343
|
`space-data-module-sdk/host/browser-module` entry point.
|
|
@@ -312,8 +389,9 @@ participates:
|
|
|
312
389
|
Rules that make this contract honest:
|
|
313
390
|
|
|
314
391
|
- **The anchor names a DIRECTORY that serves the WHOLE chain.**
|
|
315
|
-
`wasiThreadBrowserWorker.mjs` imports `./wasiThreadWorkerRuntime.js
|
|
316
|
-
|
|
392
|
+
`wasiThreadBrowserWorker.mjs` imports `./wasiThreadWorkerRuntime.js` and
|
|
393
|
+
`./wasiThreadPool.js`, which import their own siblings. Staging the single
|
|
394
|
+
`.mjs` next to your bundle does **not** work.
|
|
317
395
|
- **A relative base resolves against the document**; absolute URLs pass through
|
|
318
396
|
unchanged.
|
|
319
397
|
- **No consumer-side `location` sniffing.** Forking worker resolution per
|
|
@@ -348,15 +426,17 @@ worker bundles and the browser capability probe.
|
|
|
348
426
|
|
|
349
427
|
| Option | Effect |
|
|
350
428
|
| --- | --- |
|
|
351
|
-
| `poolSize` | Browser: pre-start exactly this many workers, independent of `hardwareConcurrency - 1` (pools are sized for isolation: writers + lanes). Node: cap on live guest threads. Arming is partial: workers that fail to start are dropped, the rest serve. |
|
|
429
|
+
| `poolSize` | Browser: pre-start exactly this many workers, independent of `hardwareConcurrency - 1` (pools are sized for isolation: writers + lanes). Node: cap on live guest threads (the pool grows to it on demand; without it, to 1024). Arming is partial: workers that fail to start are dropped, the rest serve. |
|
|
352
430
|
| `extraImports` | Per-worker import objects, as structured-cloneable descriptors (below). Factories run once per worker. |
|
|
353
|
-
| `instanceId`, `onGuestError(instanceId, tid, error)` | Called when a guest thread traps or its worker dies (A36). A dead pooled worker leaves the pool. |
|
|
354
|
-
| `onSpawnDeclined({ reason, poolSize, declined })` | Called for every spawn that returns -1. Reasons: `pool-exhausted`, `pool-not-armed`, `threads-unavailable`, `pool-empty`, `worker-create-failed`, `dispatch-failed`, `hostcall-channel-missing`, `terminated`. |
|
|
431
|
+
| `instanceId`, `onGuestError(instanceId, tid, error)` | Called when a guest thread traps or its worker dies (A36). A dead pooled worker leaves the pool. In Node, for workers the owning thread started. |
|
|
432
|
+
| `onSpawnDeclined({ reason, poolSize, declined })` | Called for every spawn from the owning thread that returns -1. Reasons: `pool-exhausted`, `pool-not-armed`, `threads-unavailable`, `pool-empty`, `worker-create-failed`, `dispatch-failed`, `hostcall-channel-missing`, `terminated`. |
|
|
355
433
|
| `probeTimeoutMs` | Browser warm-pool probe deadline. |
|
|
434
|
+
| `spawnWaitMs` | How long a spawn that finds every pool thread busy waits for one to finish before it returns -1 (default 250; 0 never waits). Added in 0.8.25; see "one pool protocol" in §3. |
|
|
356
435
|
| `browserWorkerType: "classic"` | Spawn classic workers, for the blob bundles below. |
|
|
357
436
|
|
|
358
437
|
The returned host adds `spawnReport()`: `{ poolSize, armed, failedToArm, spawned,
|
|
359
|
-
declined, declinedByReason, lastDeclineReason, active, idle }
|
|
438
|
+
waited, declined, declinedByReason, lastDeclineReason, active, idle }` (Node adds
|
|
439
|
+
`workers`). The counts cover spawns from every thread of the process. The implicit
|
|
360
440
|
(no `poolSize`) path keeps its all-or-nothing arming.
|
|
361
441
|
|
|
362
442
|
`extraImports` entries (the same descriptors work in the engine worker through
|
|
@@ -461,7 +541,8 @@ same mirror serve reads of those paths from it without a round trip.
|
|
|
461
541
|
table in a SharedArrayBuffer (`createNodeSyncFsIoTable`). A handle is
|
|
462
542
|
`(slot << 8) | gen`; each worker opens its own fd for a slot on first use and
|
|
463
543
|
drops stale fds when the generation moves. Paths are confined below `root`,
|
|
464
|
-
including through symlinked parents
|
|
544
|
+
including through symlinked parents; a filesystem root (`root: "/"`) reaches
|
|
545
|
+
every path under it (through 0.8.24 every such path was `ACCESS`). `sync` is `fdatasync` (libuv issues
|
|
465
546
|
`F_FULLFSYNC` on darwin). `revokeNodeSyncFsIoInstance` follows A23: it sets the
|
|
466
547
|
revoked flag and waits for the instance's in-flight calls to drain. Fault
|
|
467
548
|
injection (§19, 22.3a-7) belongs to FlatSQL's Node host (T4); `interpose` wraps
|
|
@@ -580,7 +661,7 @@ guest-observed import calls (`trace`), in microseconds.
|
|
|
580
661
|
| #2 500 ms open, median of 3 runs: other threads' read p99, baseline / during | 85 / 85 | 440 / 400 | 60 / 60 |
|
|
581
662
|
| #2 longest other-thread read while the open was pending | 1020 | 18500 | 320 |
|
|
582
663
|
| A38 OPEN_DEFERRED: open returns in / first write waits (ms) | 3.8 / 504 | 28.8 / 511 | 5.2 / 755 |
|
|
583
|
-
| #3 `poolSize=6` on `hardwareConcurrency=2` | 6 workers; 7th spawn -1, reported `pool-exhausted` | same | same |
|
|
664
|
+
| #3 `poolSize=6` on `hardwareConcurrency=2` (`spawnWaitMs: 0` since 0.8.25) | 6 workers; 7th spawn -1, reported `pool-exhausted` | same | same |
|
|
584
665
|
| #5 link shim v2 (Node 25): 10,000 calls, lost completions | 0 with a notifying lane, 0 with a lane that never notifies, 0 spinning (`poll_ns = 0`) | | |
|
|
585
666
|
|
|
586
667
|
Chromium's same-partition A7 p99 varies with host load: 90, 110, 295 and 790
|
|
@@ -623,6 +704,15 @@ the fast-thread read p50 from 30-39 µs to 11-12 µs.
|
|
|
623
704
|
`symbolPrefix` / `methodSymbols` as authoritative (never re-derived), keeping a
|
|
624
705
|
legacy truncated-prefix artifact compatible.
|
|
625
706
|
|
|
707
|
+
- `test/wasi-thread-pool-reuse.test.js` — the pool protocol with no message
|
|
708
|
+
and no event-loop turn: 3 waves of `poolSize - 1` and of `poolSize` spawns, a
|
|
709
|
+
spawn that blocks until a thread finishes on another thread, a spawn from a
|
|
710
|
+
thread that is not the owner, the `spawnWaitMs` bound and the no-second-wait
|
|
711
|
+
rule, an older worker script and a late `{t:"exit"}`, a dead worker,
|
|
712
|
+
`terminateAll`.
|
|
713
|
+
- `test/wasi-thread-pool-reuse-guest.test.js` — the real wave guest in Node,
|
|
714
|
+
headless Chromium, Firefox and WebKit, and the three parity lanes (the
|
|
715
|
+
browser and parity runs are env-gated, see §3).
|
|
626
716
|
- `test/wasi-thread-pool-size.test.js` — explicit `poolSize` (6 workers on
|
|
627
717
|
`hardwareConcurrency=2`, the 7th spawn -1 and reported), partial arming,
|
|
628
718
|
extraImports delivery, `onGuestError`, classic workers.
|
package/package.json
CHANGED