workflow 5.0.0 → 5.0.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -140,7 +140,7 @@ Supply a unique, deterministic approval token from the caller. The workflow must
140
140
 
141
141
  `start()` validates the source before it writes anything, so a definition that could never run fails at the call site rather than on a queue delivery:
142
142
 
143
- - It must declare `async function workflow(...)`. Pass `experimental_dynamic.exportName` to use a different name; export names may contain letters, digits, and `_`, and cannot start with a digit.
143
+ - It must declare `async function workflow(...)`. Pass `experimental_dynamic.exportName` to use a different name; export names may contain letters, digits, and `_`, cannot start with a digit, and are at most 64 characters.
144
144
  - The function's first statement must be the `"use workflow"` directive.
145
145
  - No `import` or `export`. Reach steps through `steps`, not through modules.
146
146
  - JavaScript only — no TypeScript syntax, no npm dependencies, no bundling.
@@ -182,6 +182,8 @@ Both paths are transparent — there is nothing to configure. Here, “inline”
182
182
 
183
183
  Alongside the serialized code, the run records small plaintext metadata on `executionContext.dynamicWorkflow`: the source hash, the export name, and the alias-to-step-ID map. That is what lets a run be identified as dynamic without decoding the source. It is plaintext even when the code is encrypted, so anyone who can read the run can see which step IDs it was given and the aliases they were given under.
184
184
 
185
+ In the [observability UI](/docs/observability), a dynamic run's detail view shows its stored code in a **Workflow Code** section, behind the same decrypt action as the run's input and output. **Replay Run** is unavailable for dynamic runs. To run the definition again, call `start()` with the same source.
186
+
185
187
  ## World support
186
188
 
187
189
  Dynamic workflows need a World that can store the run's workflow code.
@@ -219,6 +221,7 @@ Treat dynamic source the way you would treat code in a pull request: written or
219
221
  - Steps must already be registered in the deployment; no runtime step registration.
220
222
  - No inline `"use step"` functions, `createWebhook`, or `getWritable`.
221
223
  - No caller-provided workflow IDs.
224
+ - No **Replay Run** from the observability UI.
222
225
  - Parser-based validation checks JavaScript syntax and the required source/wrapper shape without executing it. It does not validate behavior, determinism, or intent.
223
226
  - On Vercel, roughly 30 step aliases fit the 2,048-byte execution-context limit.
224
227
  - Requires a World with dynamic-source storage.
@@ -0,0 +1,13 @@
1
+ ---
2
+ title: Advanced
3
+ description: Features for workflows that go beyond the build-time execution model.
4
+ type: overview
5
+ summary: Explore features for cases that workflows compiled into your build do not cover.
6
+ related:
7
+ - /docs/foundations
8
+ - /docs/how-it-works/code-transform
9
+ ---
10
+
11
+ These features build on the [foundations](/docs/foundations) for cases that workflows compiled into your build do not cover, such as orchestration whose shape is only known after you deploy.
12
+
13
+ <AutoCards />
@@ -0,0 +1,5 @@
1
+ {
2
+ "title": "Advanced",
3
+ "pages": ["dynamic-workflows"],
4
+ "defaultOpen": false
5
+ }
@@ -154,6 +154,10 @@ export async function POST(request: Request) {
154
154
  }
155
155
  ```
156
156
 
157
+ <Callout type="warn">
158
+ This route resumes whichever hook owns `toolCallId`, so anyone who learns the ID can submit a decision. In production, authenticate the request and check that the signed-in user may approve this booking before calling `resume()`. See [Hook and webhook security](/docs/foundations/hooks#security).
159
+ </Callout>
160
+
157
161
  </Step>
158
162
 
159
163
  <Step>
@@ -334,7 +338,7 @@ export default function ChatPage() {
334
338
 
335
339
  ## Using webhooks directly
336
340
 
337
- For simpler cases where you don't need type-safe validation or programmatic resumption, you can use [`createWebhook()`](/docs/api-reference/workflow/create-webhook) directly. This generates a unique URL that can be called to resume the workflow:
341
+ For simpler cases where you don't need type-safe validation or programmatic resumption, you can use [`createWebhook()`](/docs/api-reference/workflow/create-webhook) directly. This generates a unique URL that can be called to resume the workflow. Anyone with the URL can resume it, so avoid this for approvals that need to know who approved, unless you use the incoming payload for authorization. See [Hook and webhook security](/docs/foundations/hooks#security).
338
342
 
339
343
  ```typescript title="workflows/chat/steps/tools.ts" lineNumbers
340
344
  import { createWebhook } from "workflow";
@@ -8,11 +8,11 @@ The `@workflow/vitest` package provides a Vitest plugin and test helpers for run
8
8
  ## Installation
9
9
 
10
10
  ```package-install
11
- npm i -D @workflow/vitest@beta
11
+ npm i -D @workflow/vitest
12
12
  ```
13
13
 
14
14
  <Callout type="warn">
15
- `@workflow/vitest@latest` is still the 4.x line, so a Workflow 5 app has to install the `beta` tag (or pin the matching beta, for example `@workflow/vitest@5.0.0-beta.53`). The package carries its own copy of `@workflow/core` and runs your workflows against it, so it has to move with `workflow`.
15
+ The package carries its own copy of `@workflow/core` and runs your workflows against it, so it has to move with `workflow` and stay on the same major.
16
16
 
17
17
  `globalSetup` compares the two copies once per run: a different major fails the run with the install command that fixes it, and any other difference logs a warning. Set [`WORKFLOW_VITEST_VERSION_CHECK=off`](/docs/configuration/build-and-diagnostics#workflow_vitest_version_check) to skip the check.
18
18
  </Callout>
@@ -15,6 +15,10 @@ Creates a low-level hook primitive that can be used to resume a workflow run wit
15
15
 
16
16
  Hooks allow external systems to send data to a paused workflow without the HTTP-specific constraints of webhooks. They're identified by a token and can receive any serializable payload.
17
17
 
18
+ <Callout type="warn">
19
+ A hook token routes a payload to the right hook; it does not authorize the sender. Generated tokens are hard to guess but are not secrets, and custom tokens are usually easy to reconstruct. Authorize callers in the route that calls [`resumeHook()`](/docs/api-reference/workflow-api/resume-hook). See [Hook and webhook security](/docs/foundations/hooks#security).
20
+ </Callout>
21
+
18
22
  ```ts lineNumbers
19
23
  import { createHook } from "workflow"
20
24
 
@@ -14,7 +14,7 @@ Creates a webhook that can be used to suspend and resume a workflow run upon rec
14
14
  Webhooks provide a way for external systems to send HTTP requests directly to your workflow. Unlike hooks which accept arbitrary payloads, webhooks work with standard HTTP `Request` objects and can return HTTP `Response` objects.
15
15
 
16
16
  <Callout type="warn">
17
- `createWebhook()` creates a public endpoint at `/.well-known/workflow/v1/webhook/:token`, and the token in that URL is the only authorization performed for incoming requests resuming that webhook. This is convenient for prototypes and basic resume links because it avoids creating another route, but if you need stronger security, prefer [`createHook()`](/docs/api-reference/workflow/create-hook) behind your own route and authorize the request before calling [`resumeHook()`](/docs/api-reference/workflow-api/resume-hook) to avoid unauthenticated workflow resumptions.
17
+ `createWebhook()` creates a public endpoint at `/.well-known/workflow/v1/webhook/:token`, and the token in that URL is the only authorization performed for incoming requests resuming that webhook. Anyone who has the URL can resume the workflow with a request of their choosing, so verify requests before acting on them or use [`createHook()`](/docs/api-reference/workflow/create-hook) behind your own authorized route. See [Hook and webhook security](/docs/foundations/hooks#security).
18
18
  </Callout>
19
19
 
20
20
  ```ts lineNumbers
@@ -17,6 +17,10 @@ This is a lightweight wrapper around [`createHook()`](/docs/api-reference/workfl
17
17
  We recommend using `defineHook()` over `createHook()` in production codebases for better type safety and optional runtime validation.
18
18
  </Callout>
19
19
 
20
+ <Callout type="warn">
21
+ Schema validation checks the payload's shape, not who sent it. Authorize callers before calling `resume()`. See [Hook and webhook security](/docs/foundations/hooks#security).
22
+ </Callout>
23
+
20
24
  ```ts lineNumbers
21
25
  import { defineHook } from "workflow";
22
26
 
@@ -93,13 +93,19 @@ export async function POST(request: Request) {
93
93
 
94
94
  ### Validating hook before resume
95
95
 
96
- Use `getHookByToken` to validate hook ownership or metadata before resuming:
96
+ Use `getHookByToken` to validate hook ownership or metadata before resuming. Take the user's identity from your authentication layer, not from the request body. See [Hook and webhook security](/docs/foundations/hooks#security).
97
97
 
98
98
  ```typescript lineNumbers
99
99
  import { getHookByToken, resumeHook } from "workflow/api";
100
100
 
101
+ declare function getSession(request: Request): Promise<{ userId: string } | null>; // @setup
102
+
101
103
  export async function POST(request: Request) {
102
- const { token, userId, data } = await request.json();
104
+ const session = await getSession(request);
105
+ if (!session) {
106
+ return Response.json({ error: "Unauthorized" }, { status: 401 });
107
+ }
108
+ const { token, data } = await request.json();
103
109
 
104
110
  try {
105
111
  const hook = await getHookByToken(token); // [!code highlight]
@@ -107,7 +113,7 @@ export async function POST(request: Request) {
107
113
  const metadata = (await hook.metadata) as { allowedUserId?: string } | undefined; // [!code highlight]
108
114
 
109
115
  // Validate that the hook metadata matches the user
110
- if (metadata?.allowedUserId !== userId) {
116
+ if (metadata?.allowedUserId !== session.userId) {
111
117
  return Response.json(
112
118
  { error: "Unauthorized to resume this hook" },
113
119
  { status: 403 }
@@ -22,6 +22,10 @@ If `resumeHook()` throws any other error, the outcome is ambiguous only in dispa
22
22
  `resumeHook` is a runtime function that must be called from outside a workflow function.
23
23
  </Callout>
24
24
 
25
+ <Callout type="warn">
26
+ `resumeHook()` does not check who is calling it. Authenticate the caller and confirm they may resume this hook before calling it; knowing the token is not enough. The examples below omit that check for brevity. See [Hook and webhook security](/docs/foundations/hooks#security).
27
+ </Callout>
28
+
25
29
  ```typescript lineNumbers
26
30
  import { resumeHook } from "workflow/api";
27
31
 
@@ -17,6 +17,10 @@ This function publishes a workflow invocation carrying the request; the runtime
17
17
  `resumeWebhook` is a runtime function that must be called from outside a workflow function.
18
18
  </Callout>
19
19
 
20
+ <Callout type="warn">
21
+ The webhook token is the only authorization `resumeWebhook()` and the public webhook route perform. Verify requests before acting on them. See [Hook and webhook security](/docs/foundations/hooks#security).
22
+ </Callout>
23
+
20
24
  ```typescript lineNumbers
21
25
  import { resumeWebhook } from "workflow/api";
22
26
 
@@ -32,6 +32,10 @@ These APIs are available but are **seeded or fixed** to ensure deterministic beh
32
32
  You can safely use `Math.random()`, `Date.now()`, and `crypto.randomUUID()` in workflow functions. The framework ensures these return the same values across replays.
33
33
  </Callout>
34
34
 
35
+ <Callout type="warn">
36
+ `Math.random()`, `crypto.randomUUID()`, and `crypto.getRandomValues()` are derived from the run's seed rather than a secret, so their values are predictable to anyone who knows that seed. Don't use them for secrets, one-time codes, or tokens that must be unguessable; generate those in a step instead. See [Hook and webhook security](/docs/foundations/hooks#security).
37
+ </Callout>
38
+
35
39
  ## Web platform APIs
36
40
 
37
41
  These standard Web APIs are available in workflow functions:
@@ -204,6 +204,35 @@ const result = await world.runs.list({ // [!code highlight]
204
204
 
205
205
  **Returns:** `{ data: WorkflowRun[], cursor?: string }`
206
206
 
207
+ The array form of `status` lets you express set filters without restating the
208
+ status vocabulary. `@workflow/world` exports
209
+ `TERMINAL_WORKFLOW_RUN_STATUSES` (`['completed', 'failed', 'cancelled']`) for
210
+ that purpose:
211
+
212
+ ```typescript lineNumbers
213
+ import { TERMINAL_WORKFLOW_RUN_STATUSES } from "@workflow/world";
214
+
215
+ // Every run that has not reached a terminal status
216
+ const inFlight = await world.runs.list({
217
+ status: ["pending", "running"],
218
+ });
219
+
220
+ // The complement, without hardcoding the list
221
+ const finished = await world.runs.list({
222
+ status: [...TERMINAL_WORKFLOW_RUN_STATUSES],
223
+ });
224
+ ```
225
+
226
+ <Callout type="warn">
227
+ `status: []` matches **no** runs, mirroring SQL `IN ()`. To leave the filter
228
+ unset, omit the field entirely.
229
+
230
+ Support for the array form is per World. The Local and Postgres Worlds accept
231
+ it; the Vercel World's `/v2/runs` endpoint takes a single status today and
232
+ throws a `WorkflowWorldError` (`INVALID_ARGUMENT`) when given an array, rather
233
+ than silently filtering on something else.
234
+ </Callout>
235
+
207
236
  <Callout type="warn">
208
237
  Observability and inspection usage of `world.runs.list()` is deprecated. Use
209
238
  [`world.analytics.runs.list()`](/docs/api-reference/workflow-runtime/world/analytics#runslist)
@@ -76,4 +76,14 @@ Accepted values:
76
76
  - Dev-mode only (`next dev`). Comma-separated list of path fragments the file watcher should never watch, in addition to the built-in ignores and your project's `.gitignore`.
77
77
  - Each entry is matched as a substring of the absolute path (for example, `/fixtures/,/generated/`).
78
78
  - The watcher already respects `.gitignore` (walking from the app directory up to the workspace root). Use this variable only for large directories you cannot or do not want to add to `.gitignore`.
79
- - Useful when a project has thousands of non-ignored directories and `next dev` fails with `EMFILE: too many open files, watch`.
79
+
80
+ #### What the dev watcher tracks
81
+
82
+ In `next dev`, Workflow watches the modules your app actually imports, not your whole project:
83
+
84
+ - Every file the workflow build reached from your Next.js entrypoints. Editing one rebuilds the affected workflow bundles.
85
+ - The directories those files live in, so a module you add under an import you have already written is picked up.
86
+ - The paths behind imports that do not resolve yet. Write `import './billing/workflow'` before creating the file, or delete a workflow module and restore it, and the module is bundled as soon as it exists, wherever it lives.
87
+ - The `app` and `pages` directories (`src/` variants included), so a route you create is noticed even though nothing imports it yet, along with root entrypoints such as `middleware.ts` and `instrumentation.ts`.
88
+
89
+ A directory your app never imports is not watched at all, so files there never produce workflow bundles or rebuilds. Import one from a page and the next rebuild brings it into the watched set. Files are only ever watched through the directory that contains them, so the watcher's cost scales with the number of watched directories rather than the number of source files.
@@ -220,7 +220,7 @@ For example, a workflow can run a 10-minute inline step even with `WORKFLOW_REPL
220
220
  - Values: `node` or `quickjs`
221
221
  - Selects the sandboxed VM engine that executes workflow functions (`"use workflow"`). Step functions are unaffected and always run with full Node.js access.
222
222
  - `node` (default) runs workflow code in a [`node:vm`](https://nodejs.org/api/vm.html) context.
223
- - `quickjs` (experimental) runs workflow code in a [QuickJS](https://github.com/quickjs-ng/quickjs) VM compiled to WebAssembly (via [`quickjs-wasi`](https://github.com/vercel-labs/quickjs-wasi)). Both engines implement the same event-replay execution model (seeded PRNG, deterministic clock, and correlation-ID sequences are identical), but the **global surface is not identical**. Review the differences below before switching an existing deployment. The QuickJS engine is intended for platforms that do not implement `node:vm`, and is the foundation for future VM-memory snapshotting.
223
+ - `quickjs` (experimental) runs workflow code in a [QuickJS](https://github.com/quickjs-ng/quickjs) VM compiled to WebAssembly (via [`quickjs-wasi`](https://github.com/vercel-labs/quickjs-wasi)). Both engines implement the same event-replay execution model (seeded PRNG, deterministic clock, and correlation-ID sequences are identical), but the **global surface is not identical**. Review the differences below before switching an existing deployment. The QuickJS engine is intended for platforms that do not implement `node:vm`, and supports VM-memory snapshotting (see [`WORKFLOW_SNAPSHOT_THRESHOLD`](#workflow_snapshot_threshold)).
224
224
  - Global-surface differences under `quickjs` apply to workflow functions only. Step functions always have full Node.js:
225
225
  - `crypto.getRandomValues()` and `crypto.randomUUID()` are provided and deterministic (seeded like the node engine's). All `crypto.subtle.*` methods, including `digest`, throw with guidance to move to a step function. The node engine supports `digest`.
226
226
  - `Intl` is not available (QuickJS has no ICU). The `Intl.*` constructors throw, and `toLocaleString`-family methods (including `localeCompare`) throw when called **with an explicit locale**. Calling them without arguments keeps the engine default. Perform locale-sensitive formatting in a step function.
@@ -237,6 +237,31 @@ For example, a workflow can run a 10-minute inline step even with `WORKFLOW_REPL
237
237
  - A bundle whose module scope consumes randomness, reads the clock, or replaces a serialization intrinsic cannot be snapshotted safely. The runtime detects these cases when preparing the snapshot and falls back to per-invocation evaluation.
238
238
  - Set `0` or `false` to always evaluate the bundle per invocation.
239
239
 
240
+ ### `WORKFLOW_SNAPSHOT_THRESHOLD`
241
+
242
+ - Default: `0` (disabled)
243
+ - Values: non-negative integer
244
+ - Experimental. Only used by the QuickJS engine (`WORKFLOW_VM=quickjs`), and only with a World that provides snapshot storage (world-local, world-postgres, and world-vercel do).
245
+ - When set above `0`, the runtime persists a **VM-memory snapshot** at a suspension once at least this many events have been processed since the last snapshot. Subsequent invocations restore the VM from the snapshot and replay only the events recorded since, instead of re-executing the workflow from the top against the full event log.
246
+ - Short-lived runs below the threshold never pay the snapshot cost; long-running or unbounded runs stop scaling their resume cost with total event-log length. `1` snapshots at every qualifying suspension.
247
+ - Choosing a value: a restore costs roughly the same regardless of log length (tens of milliseconds to decompress and restore the heap, plus fetching the snapshot), while a full replay grows with the log (tens of milliseconds at a few hundred events, seconds at several thousand). Restoring starts to pay off at a few hundred events, so values in the hundreds suit most long-running workflows. Very low values mostly add snapshot saves to runs that replay quickly anyway.
248
+ - Snapshots are an optimization, not a source of truth: the event log remains authoritative, and a missing, corrupt, oversized, or incompatible snapshot automatically falls back to a full replay. Snapshots over 32 MB uncompressed are not saved.
249
+ - Like `WORKFLOW_VM`, the policy is stamped into the run's `executionContext` at start, so a run keeps the snapshot policy it started with. `start()` rejects an invalid value. An invalid value on the workflow handler disables snapshotting with a warning.
250
+ - A snapshot can only be restored by the exact QuickJS engine build that captured it. On Vercel, runs keep executing on the deployment they started on, so deploys don't affect in-flight snapshots. If your deployment serves in-flight runs from new code (for example, a self-hosted rolling update), an upgrade that changes the embedded engine makes every in-flight run fall back to a full replay at its next resume, all at once, before re-snapshotting at its next qualifying suspension. Plan for that one-time replay load when upgrading.
251
+
252
+ What a snapshot contains, and how it is protected:
253
+
254
+ - A snapshot is the workflow VM's memory: the workflow's code and in-memory state, including step results and hook payloads it holds, and the copy of `process.env` the VM exposes. It is executable state, so write access to snapshot storage is equivalent to running code in the workflow VM.
255
+ - Snapshots are compressed and then encrypted with the run's encryption key. The encryption also authenticates the snapshot: on restore, the runtime rejects any snapshot that isn't encrypted with the run's key, or whose stored metadata doesn't match the metadata sealed inside it.
256
+ - By default, runs without an encryption key are never snapshotted. Only world-vercel provides run encryption keys. On world-local and world-postgres, set `WORKFLOW_SNAPSHOT_ALLOW_UNENCRYPTED=1` on the workflow handler to snapshot anyway. Those snapshots are stored in plaintext (in `.workflow-data/snapshots/` or the `workflow_snapshots` table) and are not authenticated, so only enable this where snapshot storage is as trusted as your deployment.
257
+ - A restored run sees `process.env` as of the current invocation, the same as a full replay.
258
+ - Snapshots are deleted after the run reaches a terminal state. world-local and world-postgres have no separate retention for them, so a snapshot of a run that ends without any invocation observing it (for example, one cancelled externally) stays until you remove it.
259
+
260
+ ### `WORKFLOW_SNAPSHOT_ALLOW_UNENCRYPTED`
261
+
262
+ - Default: disabled
263
+ - Set `1` or `true` on the workflow handler to let the QuickJS engine persist VM snapshots for runs without an encryption key. See [`WORKFLOW_SNAPSHOT_THRESHOLD`](#workflow_snapshot_threshold) for what that stores.
264
+
240
265
  ## Dynamic workflows
241
266
 
242
267
  ### `WORKFLOW_EXPERIMENTAL_DYNAMIC_WORKFLOWS`
@@ -160,6 +160,12 @@ The Postgres World is a self-hosted durable backend for long-running server proc
160
160
  - Number of concurrent workers polling for jobs.
161
161
  - Also bounds concurrent parent-to-child workflow return-value polls.
162
162
 
163
+ ### `pollInterval`
164
+
165
+ - Environment variable: `WORKFLOW_POSTGRES_POLL_INTERVAL_MS`
166
+ - Default: `500`
167
+ - Milliseconds between idle job fetches per worker.
168
+
163
169
  ### `applicationManagedShutdown`
164
170
 
165
171
  - Environment variable: `WORKFLOW_POSTGRES_APPLICATION_MANAGED_SHUTDOWN` (`1` enables)
@@ -290,6 +296,7 @@ Platform-provided values such as `VERCEL_DEPLOYMENT_ID`, `VERCEL_PROJECT_ID`, an
290
296
  - Default: `http`
291
297
  - Experimental stream-write transport capability. Set to exactly `ws` to attempt `workflow-stream-ws/v1`. The server authoritatively accepts or declines each upgrade; a decline uses HTTP directly for that writer lifetime. Stream reads remain HTTP and demand-driven.
292
298
  - This is not tenant rollout policy or a package-version check. HTTP remains the compatibility path. `/websockets/v1` is independent of REST v2/v4 and persisted workflow `specVersion` values.
299
+ - A throttled (429) socket write or close is retried after the server's `Retry-After`, as over HTTP. It moves to HTTP if the connection ends during the wait or the cumulative wait passes 30 seconds.
293
300
 
294
301
  ### `WORKFLOW_DISABLE_ANALYTICS_READS`
295
302
 
@@ -313,3 +320,22 @@ When enabled (the default), a suspension's eager `step_created` and `wait_create
313
320
  - Default: `http`
314
321
  - Set to `ws` to ship workflow run events to the Vercel World over a WebSocket instead of one HTTP request each. Only `ws` (case-insensitive) opts in; any other value, including unset, empty, or `http`, keeps HTTP.
315
322
  - Ignored when the World is configured with `projectConfig` and routes through the `api-workflow` proxy: that endpoint is an HTTP-only REST gateway and does not forward a WebSocket upgrade, so events stay on HTTP.
323
+
324
+ ### `WORKFLOW_EVENTS_TRANSPORT_WS_OVERRIDE_WORKFLOWS`
325
+
326
+ - Factory option: none
327
+ - CLI flag: none
328
+ - Default: none
329
+ - Comma-separated workflows whose runs use the WebSocket events transport even when `WORKFLOW_EVENTS_TRANSPORT` is `http` or unset. Each entry is a function name (`processOrder`) or a full workflow name (`workflow//./src/workflows/order//processOrder`), matched exactly and case-sensitively.
330
+ - Has no effect when `WORKFLOW_EVENTS_TRANSPORT=ws`, which already enables every workflow.
331
+ - See [`WORKFLOW_EVENTS_TRANSPORT_WS_OVERRIDE_WORKFLOWS`](/worlds/vercel#workflow_events_transport_ws_override_workflows) for details.
332
+
333
+ ### `WORKFLOW_WS_MAX_MESSAGE_BYTES`
334
+
335
+ - Factory option: none
336
+ - CLI flag: none
337
+ - Default: `12582912` (12 MiB)
338
+ - Clamp: `2097152` to `16777216` (values outside are clamped, with a warning)
339
+ - Largest WebSocket message the Vercel World sends, header included, for both the events and stream-write transports. The ceiling is the 16 MiB WebSocket message limit.
340
+ - Events: a larger frame is sent as several messages and rebuilt on the other side.
341
+ - Stream writes: a larger write group is split into several ordered write requests; a single chunk too large for one message is written over HTTP, as are the rest of that writer's writes.
@@ -50,7 +50,7 @@ export async function paymentWebhook(orderId: string) {
50
50
 
51
51
  ### Step function for processing
52
52
 
53
- Each webhook request is processed in its own step, giving you full Node.js access for validation, database writes, and responding to the caller:
53
+ Each webhook request is processed in its own step, giving you full Node.js access for validation, database writes, and responding to the caller. In production, verify the provider's signature here before trusting the body; the example omits it for brevity.
54
54
 
55
55
  ```typescript
56
56
  import { type RequestWithResponse } from "workflow";
@@ -176,6 +176,7 @@ export async function POST(request: Request) {
176
176
  - **`respondWith: "manual"`** gives you control over the HTTP response from inside a step. Use this when you need to validate the request before responding.
177
177
  - **`for await` on a webhook** lets you process multiple events from the same URL. Use `break` to stop listening after a terminal event.
178
178
  - **Webhooks auto-generate URLs** at `/.well-known/workflow/v1/webhook/:token`. Pass this URL to external services.
179
+ - **The URL is the only authorization.** Anyone with the webhook URL can send it a request, so verify the provider's signature in the processing step before acting on the payload. See [Hook and webhook security](/docs/foundations/hooks#security).
179
180
  - **Race webhooks against `sleep()`** for deadlines. If the callback doesn't arrive in time, the workflow can take a fallback action.
180
181
  - **For large payloads**, use a hook and reference token instead of passing the data through the workflow. The event log serializes all step inputs and outputs, so large payloads hurt performance.
181
182
 
@@ -168,7 +168,7 @@ Uncaught, the run fails immediately with the `USER_ERROR` code, without retrying
168
168
 
169
169
  On Vercel, backend connection failures and interrupted event streams use the existing retry policies, even when their error codes are unrecognized. This lets workflows recover from network failures instead of immediately failing with `USER_ERROR`. Persistent failures can still exhaust the retry budget.
170
170
 
171
- The SDK also replaces its shared events connection pool after repeated HTTP/2 session failures. Invalid backend URLs, including unsupported protocols and embedded credentials, fail immediately. Fetch requests to blocked ports or with unsupported headers (such as `Expect`) also fail without retrying. Interrupted event writes retain their existing retries; caller cancellations do not trigger another write.
171
+ The SDK also replaces its shared events connection pool after repeated HTTP/2 session failures, and retries event-log reads whose HTTP/2 stream the backend resets. An events request that receives no response data for 60 seconds fails and is retried instead of stalling the invocation. Invalid backend URLs, including unsupported protocols and embedded credentials, fail immediately. Fetch requests to blocked ports or with unsupported headers (such as `Expect`) also fail without retrying. Interrupted event writes retain their existing retries; caller cancellations do not trigger another write.
172
172
 
173
173
  A connection failure does not prove that the backend rejected a write: it may have accepted it before the response was lost. Continue to make step side effects [idempotent](/docs/foundations/idempotency).
174
174
 
@@ -206,9 +206,9 @@ try {
206
206
  | `REPLAY_TIMEOUT` | A workflow replay exceeded the maximum allowed duration |
207
207
  | `REPLAY_DIVERGENCE` | A replay could not consume the event log deterministically, usually because of non-deterministic workflow code. |
208
208
  | `CORRUPTED_EVENT_LOG` | The event log cannot be replayed: it contains orphaned or mismatched events, or one of its stored payloads is no longer readable from the World's storage. If you see this, please [file an issue](https://github.com/vercel/workflow/issues) |
209
- | `STREAM_ERROR` | Workflow stream infrastructure failed while reading or writing data. This is an SDK or backend failure rather than an error in workflow code |
209
+ | `STREAM_ERROR` | Workflow stream infrastructure failed while reading or writing data. This is an SDK or backend failure rather than an error in workflow code; retry the run and report persistent failures with the `runId` |
210
210
  | `WORLD_CONTRACT_ERROR` | A World response violated the SDK contract; points at a World implementation bug |
211
- | `DEPLOYMENT_MISMATCH` | The run was delivered to a deployment other than the one it is pinned to, and automatic re-routing did not recover it |
211
+ | `DEPLOYMENT_MISMATCH` | The run was delivered to a deployment other than the one it is pinned to, and automatic re-routing did not recover it. See [deployment-mismatch](/docs/errors/deployment-mismatch) |
212
212
  | `RUNTIME_ERROR` | An internal runtime error. If you see this, please [file an issue](https://github.com/vercel/workflow/issues) |
213
213
 
214
214
  <Callout type="info">
@@ -82,7 +82,7 @@ export async function POST(request: Request) {
82
82
 
83
83
  The key points:
84
84
  - Hooks allow you to pass **any [serializable data](/docs/foundations/serialization)** as the payload
85
- - You need the hook's `token` to resume it
85
+ - You need the hook's `token` to resume it, but knowing the token does not authorize the caller. Check who is calling before `resumeHook()`; see [Security](#security)
86
86
  - The workflow will resume execution right where it left off
87
87
 
88
88
  ### Checking for token conflicts
@@ -285,7 +285,7 @@ Hooks require you to manually handle HTTP requests and route them to workflows.
285
285
  When using Workflow SDK, webhooks are automatically wired up at `/.well-known/workflow/v1/webhook/:token` without any additional setup.
286
286
 
287
287
  <Callout type="warn">
288
- `createWebhook()` exposes a public route at `/.well-known/workflow/v1/webhook/:token`, and the token in that URL is the only authorization performed for incoming requests. This is convenient for prototypes because you can share the webhook URL (endpoint) without creating another route, but if you need stronger security, prefer [`createHook()`](/docs/api-reference/workflow/create-hook) behind your own route and authorize the request before calling [`resumeHook()`](/docs/api-reference/workflow-api/resume-hook) to avoid unauthenticated workflow resumptions.
288
+ `createWebhook()` exposes a public route at `/.well-known/workflow/v1/webhook/:token`, and the token in that URL is the only authorization performed for incoming requests. Make sure to read up on [security](#security) before using a webhook for calls that need to be authenticated.
289
289
  </Callout>
290
290
 
291
291
  <Callout type="info">
@@ -459,6 +459,94 @@ export async function eventCollectorWorkflow() {
459
459
  - You need to send HTTP responses back to the caller
460
460
  - You want automatic URL routing without writing API handlers
461
461
 
462
+ ## Security
463
+
464
+ A hook token tells the runtime which hook a payload belongs to. It is not an authentication mechanism, and neither `resumeHook()` nor the webhook endpoint checks who is sending the payload.
465
+
466
+ ### Generated tokens are hard to guess, not secret
467
+
468
+ When you don't pass a `token`, the SDK generates one inside the workflow function. Workflow code must produce the same values on every replay, so the generated token comes from the run's deterministic random number generator, the same one that backs [`Math.random()` and `crypto.randomUUID()`](/docs/api-reference/workflow-globals) in workflow functions. That generator is seeded from identifiers of the run, including the run ID, not from a secret key.
469
+
470
+ A generated token is hard to guess without knowing the run, but the values it is derived from are not designed to be kept secret. Treat a generated token like an unlisted link, not like a credential.
471
+
472
+ Custom tokens passed to `createHook({ token })` are usually built from domain data such as an order ID, so they are even easier to reconstruct. That is what makes them useful for routing, and it is why the route that resumes them must do its own authorization. If you can not perform your own authorization on the route that calls `resume` for any reason, and need to generate an unguessable token instead, generate it in a step where `crypto` is not seeded and pass it as `token`:
473
+
474
+ ```typescript lineNumbers
475
+ import { createHook } from "workflow";
476
+
477
+ async function generateToken() {
478
+ "use step";
479
+ // Steps run outside the workflow sandbox, so this uses the platform's
480
+ // cryptographic random source instead of the run's seed.
481
+ return crypto.randomUUID();
482
+ }
483
+
484
+ export async function approvalWorkflow() {
485
+ "use workflow";
486
+
487
+ const token = await generateToken(); // [!code highlight]
488
+ using hook = createHook<{ approved: boolean }>({ token }); // [!code highlight]
489
+
490
+ return (await hook).approved;
491
+ }
492
+ ```
493
+
494
+ The step result is recorded in the run's event log like any other step result. See [Encryption](/docs/how-it-works/encryption) to keep it encrypted at rest.
495
+
496
+ ### Webhook URLs
497
+
498
+ [`createWebhook()`](/docs/api-reference/workflow/create-webhook) serves a public route at `/.well-known/workflow/v1/webhook/:token`, and matching the token is the only check it performs. Anyone who has the URL, or can compute its token, can resume the workflow with a request of their choosing. That is fine for low-stakes callbacks and prototypes. When a webhook request triggers something consequential, either:
499
+
500
+ - Verify each request before acting on it, for example by checking the provider's HMAC signature in a step with [`respondWith: "manual"`](#dynamic-responses-manual-mode), and keep waiting for the next request when verification fails.
501
+ - Use [`createHook()`](/docs/api-reference/workflow/create-hook) behind your own route and authorize the caller before calling [`resumeHook()`](/docs/api-reference/workflow-api/resume-hook), as shown below.
502
+
503
+ ### Authorize before calling `resumeHook()`
504
+
505
+ `resumeHook()` delivers the payload to whichever hook owns the token. The route that calls it has to authenticate the caller and check that they are allowed to resume that specific hook. Knowing the token is not proof of either. One way is to record who may resume the hook in its `metadata`, then compare it to the signed-in user:
506
+
507
+ ```typescript lineNumbers
508
+ import { createHook } from "workflow";
509
+
510
+ export async function expenseWorkflow(approverId: string) {
511
+ "use workflow";
512
+
513
+ using hook = createHook<{ approved: boolean }>({
514
+ metadata: { approverId }, // [!code highlight]
515
+ });
516
+
517
+ return (await hook).approved;
518
+ }
519
+ ```
520
+
521
+ ```typescript lineNumbers
522
+ import { getHookByToken, resumeHook } from "workflow/api";
523
+
524
+ declare function getSession(request: Request): Promise<{ userId: string } | null>; // @setup
525
+
526
+ export async function POST(request: Request) {
527
+ const session = await getSession(request); // [!code highlight]
528
+ if (!session) {
529
+ return Response.json({ error: "Unauthorized" }, { status: 401 });
530
+ }
531
+
532
+ const { token, approved } = await request.json();
533
+ const hook = await getHookByToken(token);
534
+ const metadata = (await hook.metadata) as { approverId?: string } | undefined;
535
+ if (metadata?.approverId !== session.userId) { // [!code highlight]
536
+ return Response.json({ error: "Forbidden" }, { status: 403 });
537
+ }
538
+
539
+ await resumeHook(token, { approved });
540
+ return Response.json({ success: true });
541
+ }
542
+ ```
543
+
544
+ Take the user identity from your authentication layer, never from the request body. Other examples in these docs omit authorization to stay short.
545
+
546
+ ### Randomness in workflow functions
547
+
548
+ The same determinism applies to your own code. `Math.random()`, `crypto.randomUUID()`, and `crypto.getRandomValues()` in a workflow function return values derived from the run's seed, so they are predictable to anyone who knows it. Don't use them for secrets, passwords, one-time codes, or any value that must be unguessable. Generate those in a step, as in the token example above.
549
+
462
550
  ## Advanced patterns
463
551
 
464
552
  ### Type-safe hooks with `defineHook()`
@@ -514,7 +602,7 @@ This pattern is especially valuable in larger applications where the workflow an
514
602
 
515
603
  ### Token design
516
604
 
517
- Custom tokens are available for `createHook()` with server-side `resumeHook()` only. Webhooks (`createWebhook()`) always generate their own unique tokens. A generated token is not trivial to guess, but it is not a strong security contract either, so anyone who obtains the URL can invoke an unintended webhook resumption. To prevent unauthenticated run resumptions entirely, prefer a **hook** over the **webhook** convenience and implement your own authentication on the route that calls `resumeHook()`.
605
+ Custom tokens are available for `createHook()` with server-side `resumeHook()` only. Webhooks (`createWebhook()`) always generate their own unique tokens. Neither kind of token authorizes the sender; see [Security](#security).
518
606
 
519
607
  When using custom tokens with `createHook()`:
520
608
 
@@ -49,7 +49,7 @@ export async function processOrderWorkflow(orderId: string) {
49
49
 
50
50
  Determinism in the workflow is required to resume the workflow from a suspension. Essentially, the workflow code gets re-run multiple times during its lifecycle, each time using the [event log](/docs/how-it-works/event-sourcing) to resume the workflow to the correct spot.
51
51
 
52
- The sandboxed environment that workflows run in already ensures determinism. For instance, `Math.random` and `Date` constructors are fixed in workflow runs, so you are safe to use them, and the framework ensures that the values don't change across replays.
52
+ The sandboxed environment that workflows run in already ensures determinism. For instance, `Math.random` and `Date` constructors are fixed in workflow runs, so you are safe to use them, and the framework ensures that the values don't change across replays. Seeded random values are not secret, so generate secrets and unguessable tokens in a step (see [Hook and webhook security](/docs/foundations/hooks#security)).
53
53
 
54
54
  ## Step functions
55
55
 
@@ -268,7 +268,7 @@ Contains all step functions transformed in **step mode**. The combined flow hand
268
268
  This module must not be exposed as an HTTP endpoint.
269
269
 
270
270
  <Callout type="info">
271
- **Changed in 5.0:** In 4.x, the step bundle was served as its own HTTP route at `POST /.well-known/workflow/v1/step`, with step messages delivered on a separate `__wkf_step_*` queue topic. v5 merged both into the combined flow handler. The step bundle became a registration module imported by `flow.js`, and step messages arrive on the shared workflow queue. Use the version picker to see the old layout on the v4 version of this page.
271
+ **Changed in 5.0:** In 4.x, the step bundle was served as its own HTTP route at `POST /.well-known/workflow/v1/step`, with step messages delivered on a separate `__wkf_step_*` queue topic. v5 merged both into the combined flow handler. The step bundle became a registration module imported by `flow.js`, and step messages arrive on the shared workflow queue. See the [v4 version of this page](/v4/docs/how-it-works/code-transform) for the old layout.
272
272
  </Callout>
273
273
 
274
274
  ### `webhook.js`
package/docs/meta.json CHANGED
@@ -5,7 +5,7 @@
5
5
  "getting-started",
6
6
  "foundations",
7
7
  "how-it-works",
8
- "advanced/dynamic-workflows",
8
+ "advanced",
9
9
  "observability",
10
10
  "ai",
11
11
  "testing",
@@ -91,13 +91,11 @@ For workflows that rely on runtime features like [hooks](/docs/foundations/hooks
91
91
  ### Installation
92
92
 
93
93
  ```package-install
94
- npm i -D @workflow/vitest@beta
94
+ npm i -D @workflow/vitest
95
95
  ```
96
96
 
97
97
  <Callout type="warn">
98
- **Workflow 5 needs the `beta` tag.** `@workflow/vitest@latest` is still the 4.x line, so a plain `npm i -D @workflow/vitest` installs a v4 plugin next to a v5 app. Install `@workflow/vitest@beta`, or pin the beta your app is on (for example `@workflow/vitest@5.0.0-beta.53`), and upgrade it together with `workflow`.
99
-
100
- The plugin builds and runs your workflows against the copy of `@workflow/core` it was installed with, so it checks this for you. At the start of each run it compares that copy with the one your app resolves: a different major fails the run with the install command that fixes it, and any other difference logs a warning. Set [`WORKFLOW_VITEST_VERSION_CHECK=off`](/docs/configuration/build-and-diagnostics#workflow_vitest_version_check) to skip the check.
98
+ **Keep `@workflow/vitest` on the same major as `workflow`.** The plugin builds and runs your workflows against the copy of `@workflow/core` it was installed with, so it checks this for you. At the start of each run it compares that copy with the one your app resolves: a different major fails the run with the install command that fixes it, and any other difference logs a warning. Set [`WORKFLOW_VITEST_VERSION_CHECK=off`](/docs/configuration/build-and-diagnostics#workflow_vitest_version_check) to skip the check.
101
99
  </Callout>
102
100
 
103
101
  ### Vitest configuration
@@ -483,7 +481,7 @@ Integration tests are the right place to verify that your workflows handle error
483
481
 
484
482
  ### Upgrade `@workflow/vitest` with the SDK
485
483
 
486
- `@workflow/vitest` carries its own copy of the Workflow runtime, so treat it as part of the same upgrade as `workflow`. While Workflow 5 is in beta, that means the `beta` tag on both. See [Installation](#installation).
484
+ `@workflow/vitest` carries its own copy of the Workflow runtime, so treat it as part of the same upgrade as `workflow`: bump both together, and keep them on the same major. See [Installation](#installation).
487
485
 
488
486
  ## Further reading
489
487
 
@@ -20,7 +20,7 @@ npx skills add https://github.com/vercel/workflow --skill migrating-workflow-v4-
20
20
 
21
21
  <Callout type="info">
22
22
  Workflow SDK v4 remains installable as `workflow@4` and receives stability
23
- fixes. Switch to its documentation with the version picker in the sidebar.
23
+ fixes. Its documentation lives at [/v4/docs](/v4/docs).
24
24
  </Callout>
25
25
 
26
26
  ## Highlights
@@ -39,6 +39,10 @@ The largest change in v5 has no API surface: the runtime does far less work per
39
39
 
40
40
  **Resuming a hook takes one round trip instead of two.** `resumeHook()` writes the `hook_received` event and dispatches the queue message concurrently, with a `(runId, resumeId)` dedup constraint keeping the two writers converging on exactly one event. See [Resilient hook resumption](/docs/changelog/resilient-resume).
41
41
 
42
+ **A suspension's writes go out as one batch.** The `step_created` and `wait_created` events a suspension produces are folded into a single durable write with a per-event outcome, instead of one request each, and a fan-out's inline step bodies start straight off that commit rather than each claiming its step first. This engages on Worlds that implement the batch API and can be turned off with [`WORKFLOW_BATCH_TRANSITIONS=0`](/docs/configuration/worlds#workflow_batch_transitions). See [Batched event writes](/docs/changelog/batched-event-writes).
43
+
44
+ **Concurrent writers no longer compete for a position in the log.** Each event's position used to be claimed by the write that filled it, so a wide fan-out serialized on that claim. Positions are now handed out ahead of the commit, which is what makes the fan-out above cheap. The cost is a position whose writer dies, which the backend closes with a `noop` event that replay steps over. Nothing about this is visible from workflow code; [`WORKFLOW_SEALED_LOG=0`](/docs/configuration/runtime-tuning#workflow_sealed_log) is the kill switch, and existing runs keep the scheme they were created on.
45
+
42
46
  **Payloads are compressed** before they are encrypted and sent to the API. Repetitive payloads compress heavily; AI token streams average around 80% smaller. That is less stored data and less to move over the network.
43
47
 
44
48
  On [Vercel Workflows](/worlds/vercel) these benefits compound to reduce compute costs by up to 80% for workflows made of many small steps, and storage cost by up to 70%, depending on the workload. Depending on your setup, you may see similar gains using self-hosted or third-party Worlds.
@@ -149,7 +153,9 @@ All three first-party Worlds now implement it: Vercel accepts up to 30 days, and
149
153
  - **A misrouted delivery no longer fails a run.** Runs are pinned to the deployment that created them. A delivery that arrives at a different deployment is now re-routed to the pinned one with backoff instead of failing, and only gives up with the new [`DEPLOYMENT_MISMATCH`](/docs/errors/deployment-mismatch) error once the recovery budget is spent. Nothing executes on the wrong deployment while this happens. In 4.x the same situation surfaced as an unexplained decryption failure.
150
154
  - **A run cannot be forked across environments.** `start()` stamps the environment it was called from onto the queue message, and a deployment refuses a delivery whose run was created in a different environment. Previously a preview client and a production deployment could each hold half of one run ID.
151
155
  - **An experimental QuickJS VM engine.** Set [`WORKFLOW_VM=quickjs`](/docs/configuration/runtime-tuning#workflow_vm) to run workflow functions in a QuickJS VM compiled to WebAssembly instead of `node:vm`, for platforms that do not provide `node:vm`. Replay semantics are identical, but the available globals are not: check the differences before switching an existing deployment.
152
- - **An opt-in WebSocket transport for event writes** on the Vercel World, via [`WORKFLOW_EVENTS_TRANSPORT=ws`](/docs/configuration/worlds). HTTP remains the default.
156
+ - **Experimental dynamic workflows.** Pass workflow source as a string to `start()` to run orchestration assembled after deployment over steps already deployed with your app. The code is stored with the run (encrypted on Vercel), so replays always execute the code the run started with. Off unless the deployment sets `WORKFLOW_EXPERIMENTAL_DYNAMIC_WORKFLOWS=1`. See [Dynamic Workflows](/docs/advanced/dynamic-workflows).
157
+ - **An opt-in WebSocket transport for event writes** on the Vercel World, which ships a run's events over one socket instead of one HTTP request each. Set [`WORKFLOW_EVENTS_TRANSPORT=ws`](/worlds/vercel#workflow_events_transport) to opt in. Any other value, including unset, keeps HTTP. Tracing is unaffected: each write still emits an `http POST` client span, synthesized around the frame, with the transport on `workflow.events.transport`.
158
+ - **`await run.returnValue` waits instead of polling.** Reading a run's result long-polls the World until the run reaches a terminal status, rather than asking again on a fixed interval. A result is observed as soon as it exists, and an idle wait costs one open request instead of a request per tick. Worlds that do not implement the long poll keep the interval.
153
159
  - **An event arriving mid-replay no longer fails the run.** A hook resume or step completion landing while a replay is in flight used to be able to fail it with [`CORRUPTED_EVENT_LOG`](/docs/errors/corrupted-event-log). Writes now come back with the events the replay had not seen, and the event is held for whichever part of the workflow awaits it. A run only fails when the log is genuinely missing a position.
154
160
 
155
161
  ## Breaking changes
@@ -130,7 +130,12 @@ interface Storage {
130
130
 
131
131
  // Create an event for an existing run
132
132
  create(runId: string, data: CreateEventRequest, params?: CreateEventParams): Promise<EventResult>;
133
-
133
+
134
+ // Optional: append an ordered list of events in one durable write, with a
135
+ // per-event outcome for each. Implementing it is the declaration; there is
136
+ // no capability flag.
137
+ createBatch?(runId: string, events: BatchEventRequest[], params?: CreateEventBatchParams): Promise<EventBatchResult>;
138
+
134
139
  list(params: ListEventsParams): Promise<PaginatedResponse<Event>>;
135
140
  listByCorrelationId(params: ListEventsByCorrelationIdParams): Promise<PaginatedResponse<Event>>;
136
141
  };
@@ -151,6 +156,8 @@ interface Storage {
151
156
  2. Atomically update the affected entity (run, step, or hook)
152
157
  3. Return both the created event and the updated entity
153
158
 
159
+ **Batch writes:** `events.createBatch()` is optional, and implementing it is what declares it. The runtime folds a suspension's `step_created` and `wait_created` writes into batches only when the method exists, and otherwise takes the single-event path unchanged. Implement it with real atomicity per attempt, so a lost race leaves nothing behind, or do not implement it at all. The events land in request order at consecutive positions, and a concurrent writer may push the whole batch above the caller's view of the log. No skipped-event report accompanies the result, so a position-tracking caller compares the committed positions against what it expected and reloads. Reject the whole batch, with a request-level error, for `run_created`, `run_started`, `run_cancelled`, `hook_created`, `hook_disposed`, `attr_set`, for more events targeting one entity than a single write can express, and for a batch over your own size caps. The one legal same-entity pair is `step_created` followed by `step_started`, which creates the step born-running, and there the input must ride the `step_created`.
160
+
154
161
  **Run creation:** For `run_created` events, the `runId` parameter may be a client-provided string or `null`. When `null`, your World generates and returns a new `runId`.
155
162
 
156
163
  **Event data resolution:** `events.list()` and `events.create()` accept `resolveData: 'skip-step-inputs'`, and the runtime replays with it. The World may leave `input` out of `step_created` and `step_started` events, and returns everything else as for `'all'`. Treat any value other than `'none'` as `'all'`, because a World that tests `resolveData === 'all'` strips step results and breaks every replay. Map the value with `entityResolveData()` from `@workflow/world` before passing it to an entity read.
@@ -163,6 +170,26 @@ Keep the owning run available for at least as long as its token remains unavaila
163
170
 
164
171
  **Automatic hook cleanup:** When a run ends, remove its live hooks. Make each token available unless its `tokenRetentionUntil` is still in the future. A `hook_disposed` event always makes the token available immediately.
165
172
 
173
+ **Terminal-run start fence:** reject a `step_started` write once the run has
174
+ reached a terminal status, including a redelivery whose step row still reads
175
+ `running` because an earlier delivery claimed it. Starting work on a finished run
176
+ is never valid, and the outcome of such a step cannot be consumed by anything.
177
+ The queued-step path depends on this rejection, because a step message carries
178
+ run identity in place of a run fetch and has no other liveness check. A step that
179
+ was already in flight when the run ended still writes its `step_completed` or
180
+ `step_failed` unchanged: the fence is on starting, not on finishing.
181
+
182
+ **Listing runs:** `runs.list()` takes an optional `workflowName` and an optional
183
+ `status`, and `status` is either one `WorkflowRunStatus` or an array of them. An
184
+ array matches runs in *any* of the listed statuses, and an empty array matches
185
+ nothing, mirroring SQL `IN ()`. An unset filter means "every status", so treat
186
+ `status: []` and an omitted `status` as different requests. If your backend
187
+ cannot express the array form, reject it with an explicit error rather than
188
+ filtering on the first element or ignoring the filter: a silently wrong result
189
+ set is worse than a failed call. `@workflow/world` exports
190
+ `TERMINAL_WORKFLOW_RUN_STATUSES` so neither you nor your callers restate the
191
+ status vocabulary.
192
+
166
193
  ### Optional: waiting for a terminal run status
167
194
 
168
195
  `await run.returnValue` must determine when a run finishes. Without help, it rereads the run every second, so it reports a run that finishes just after a read up to 1s late. Implement `runs.waitForTerminalStatus(id, { timeoutMs, signal, resolveData })` so the runtime asks once and receives an answer when the run ends.
@@ -324,11 +351,22 @@ interface WorkflowInvokePayload {
324
351
  runId: string;
325
352
  stepId?: string;
326
353
  stepName?: string;
354
+ runContext?: RunDispatchContext; // Immutable run identity, step messages only
327
355
  traceCarrier?: Record<string, string>; // OpenTelemetry context
328
356
  requestedAt?: Date;
329
357
  }
330
358
  ```
331
359
 
360
+ **Run identity on step messages:** a step-execution message carries `runContext`
361
+ with the run's `deploymentId`, `specVersion`, `startedAt` and `rootRunId`. Each
362
+ of those is fixed for the life of a run, so a consumer that receives them starts
363
+ the step from the message alone instead of fetching the run first. Run *status*
364
+ is deliberately absent: the liveness check for a queued step is the
365
+ `step_started` write itself, which your World must reject on a terminal run (see
366
+ [Key implementation details](#key-implementation-details)). Messages from older
367
+ producers omit `runContext` and take the older path that fetches the run, so a
368
+ World needs to do nothing to support either.
369
+
332
370
  The SDK also sends an internal `HealthCheckPayload` through the same workflow queue.
333
371
 
334
372
  ### Implementation considerations
@@ -350,6 +388,13 @@ interface Streamer {
350
388
  streamFlushIntervalMs?: number;
351
389
 
352
390
  streams: {
391
+ // Optional. A stateful writer lifetime; see below.
392
+ createWriteSession?(
393
+ runId: string,
394
+ name: string,
395
+ options: { writerId: `wrtr_${string}` }
396
+ ): StreamWriteSession;
397
+
353
398
  write(
354
399
  runId: string,
355
400
  name: string,
@@ -396,6 +441,18 @@ interface Streamer {
396
441
  Streams are identified by a combination of `runId` and `name`. Each workflow run can have multiple named streams.
397
442
  `writeMulti()` is an optional optimization for batching multiple writes.
398
443
 
444
+ `createWriteSession()` is optional. Implement it when your transport can do
445
+ better by holding state across a writer's chunks, for example by keeping one
446
+ connection open instead of reopening it per write. The runtime creates at most
447
+ one session per in-memory `WritableStream` and passes a `writerId` identifying
448
+ that lifetime, so the session's `write(chunkSeq, chunks)` receives a sequence
449
+ number that is writer-local rather than stream-global. That is what lets a
450
+ World order concurrent writers to the same stream. The session's `close()` must
451
+ not resolve before every prior write is durable; `dispose()` is optional and
452
+ releases transport resources without ending the stream. A World that does not
453
+ implement `createWriteSession` is unaffected: the runtime uses
454
+ `write`/`writeMulti`/`close` exactly as before.
455
+
399
456
  `getChunks` returns a paginated snapshot of currently available chunks, unlike `get`, which returns a live `ReadableStream` that waits for new chunks. `getInfo` returns the tail index (last chunk index, 0-based, or `-1` when empty) and whether the stream is complete. Use this information to resolve negative `startIndex` values into absolute positions.
400
457
 
401
458
  ## Analytics interface (optional)
@@ -422,7 +479,11 @@ interface Analytics {
422
479
  };
423
480
  events: {
424
481
  get(runId: string, eventId: string): Promise<AnalyticsEvent>;
482
+ // Bounded batch lookup within one run. Not paginated; rows that analytics
483
+ // has not ingested yet are omitted rather than throwing.
484
+ getMany(runId: string, eventIds: readonly string[]): Promise<AnalyticsEvent[]>;
425
485
  list(params: AnalyticsListEventsParams): Promise<PaginatedResponse<AnalyticsEvent>>;
486
+ /** @deprecated A special case of `list({ runId, correlationId })`. Removed in the next major. */
426
487
  listByCorrelationId(params: AnalyticsListEventsByCorrelationIdParams): Promise<PaginatedResponse<AnalyticsEvent>>;
427
488
  };
428
489
  hooks: {
@@ -441,6 +502,7 @@ If you implement this namespace, observe the following requirements:
441
502
  - **Metadata only.** Analytics responses must not include run inputs or outputs, step data, hook tokens, or other payload data. The `Analytics*` schemas exported by `@workflow/world` define the complete set of permitted fields. Payload retrieval remains exclusively available through the Storage APIs.
442
503
  - **Attribute filters use the latest value.** `runs.list({ attributes })` evaluates each filter against the run's most recently written value for that key. A request may contain up to eight key-value pairs. Reserved `$`-prefixed attributes are valid filters, although users cannot write them directly.
443
504
  - **Time boundaries must be paired.** `startTime` and `endTime` may either both be omitted or both be supplied. Responses may include `pageInfo` describing retention and the available query window. Implementations with retention limits should return this information so tooling can present valid date ranges.
505
+ - **Page limits are part of the contract.** `pagination.limit` defaults to 40 and caps at 1000 for the run-scoped listings (`steps.list`, `events.list`, `waits.list`) and at 100 for the cross-run ones (`runs.list`, `attributes.list`, `hooks.list`); `events.getMany` accepts 1 to 100 ids. Callers hit these bounds before a request leaves their process, as a `RangeError`, so a request that reaches your implementation is already within them.
444
506
  - **Results may be eventually consistent.** Analytics records may lag live workflow state. Consumers use this namespace for discovery and listing; Storage remains the authoritative interface for current workflow state and payload access.
445
507
 
446
508
  See the [Analytics API reference](/docs/api-reference/workflow-runtime/world/analytics) for per-method parameters, row shapes, and `pageInfo` semantics.
@@ -278,6 +278,10 @@ Number of concurrent workers polling for jobs. Default: `50`.
278
278
 
279
279
  This value also bounds how many parent→child workflow polls can be in flight simultaneously. Every `await childRun.returnValue` inside a workflow holds a worker slot until the child run terminates. If you expect recursive or highly-fanned-out parent/child workflows, raise this ceiling above the peak number of concurrent polls. With the default of 50, the included `fibonacciWorkflow` end-to-end test (`fib(6)`, about 24 concurrent polls at peak) passes; deeper recursion or larger fanouts need a correspondingly larger setting.
280
280
 
281
+ ### `WORKFLOW_POSTGRES_POLL_INTERVAL_MS`
282
+
283
+ Milliseconds between idle job fetches per worker. Each worker polls on its own, so an idle process runs about `queueConcurrency × 1000 / pollInterval` fetches per second. Default: `500`.
284
+
281
285
  ### `WORKFLOW_POSTGRES_MAX_POOL_SIZE`
282
286
 
283
287
  Maximum size of the internal `pg.Pool` used when `createWorld()` constructs the pool. Default: the `pg` default (`10`).
@@ -36,27 +36,36 @@ We're working on bringing back World compatibility tests and reporting on the [W
36
36
 
37
37
  ## Spec versions
38
38
 
39
- A World declares the protocol version it speaks on `specVersion`, and that number is stamped on every run it creates. Declare `SPEC_VERSION_CURRENT` from `@workflow/world`, not a literal:
39
+ A World declares the protocol version it speaks on `specVersion`, and that number is stamped on every run it creates. Declare `mintedSpecVersion()` from `@workflow/world`, not a literal:
40
40
 
41
41
  {/* @skip-typecheck - partial World, the other members are elided */}
42
42
  ```typescript
43
- import { SPEC_VERSION_CURRENT } from '@workflow/world';
43
+ import { mintedSpecVersion } from '@workflow/world';
44
44
 
45
45
  export function createWorld(): World {
46
46
  return {
47
- specVersion: SPEC_VERSION_CURRENT,
47
+ specVersion: mintedSpecVersion(),
48
48
  // ...
49
49
  };
50
50
  }
51
51
  ```
52
52
 
53
- In v4 the runtime required that number to equal its own current version exactly. In v5 it checks the declaration against a range, `[SPEC_VERSION_CURRENT, SPEC_VERSION_MAX_SUPPORTED]`, before it creates or replays anything, and refuses a World outside it with an error naming both the range and what your World declared.
53
+ In v4 the runtime required that number to equal its own current version exactly. In v5 it checks the declaration against a range before it creates or replays anything, and refuses a World outside it with an error naming both the range and what your World declared. The floor is the version that introduced [slot-numbered event IDs](#event-id-allocation), because a World below it allocates IDs the runtime cannot read positions out of, and admitting one would only move the failure from startup into the middle of a run. The ceiling is the highest version this runtime can read.
54
54
 
55
- The two bounds are the same version today, so exactly one is accepted. That is a consequence of [event ID allocation](#event-id-allocation) being a requirement rather than an option: a World declaring anything lower allocates IDs the runtime cannot read positions out of, and admitting it would only move the failure from startup into the middle of a run. The check is written as a range because the constants answer different questions and come apart while a version bump is staged. The ceiling rises when the runtime learns to read the next version, and the floor rises when that version becomes the one Worlds stamp.
55
+ `mintedSpecVersion()` is a function rather than a constant because the version a World stamps is a deployment-level choice. It answers with the sealed-log version by default, and with the slot-identity version when [`WORKFLOW_SEALED_LOG=0`](/docs/configuration/runtime-tuning#workflow_sealed_log) opts new runs out. Both sit inside the accepted range, so either answer is a valid declaration. Reading it per `createWorld()` call rather than once at module load is what lets a single process create Worlds in both modes.
56
56
 
57
- Using the constant is what keeps the check passing across upgrades. It moves with the `@workflow/world` version your package resolves, so a bump raises your declaration and the runtime's floor together, while a hard-coded number leaves your World a version behind the next bump and gets it rejected by the runtime it ships alongside. This is worth re-checking if you followed earlier guidance: `SPEC_VERSION_SUPPORTS_SLOT_IDENTITY` names the version that introduced slot-numbered IDs and is equal to `SPEC_VERSION_CURRENT` today, but declaring it pins you to a literal by another name. `@workflow/world-vercel` declared it and now declares the current version instead. Keep `@workflow/world` in the same release channel as the `workflow` version your users install.
57
+ Calling it is also what keeps the check passing across upgrades, since it moves with the `@workflow/world` version your package resolves. A hard-coded number leaves your World a version behind the next bump and gets it rejected by the runtime it ships alongside. That includes the constants: `SPEC_VERSION_CURRENT` and `SPEC_VERSION_SUPPORTS_SLOT_IDENTITY` are literals by another name for this purpose, since neither follows the sealed-log setting. Keep `@workflow/world` in the same release channel as the `workflow` version your users install.
58
58
 
59
- Runs carry a spec version too, and a run keeps the version it was created under for its whole life. Read the stamped version off the run rather than assuming every run matches what your World declares today. Bumping the constant does not reach runs already in your store: their version is persisted, every version test in the runtime is a lower bound, and a run's event ID scheme is resolved from what is stored.
59
+ ### Sealed logs and `noop` events
60
+
61
+ The sealed-log version exists for a World whose store makes allocating a position at the commit a contention bottleneck. Such a World may hand positions out from a per-run counter *before* the commit, so concurrent writers never race for one, and then restore density at read time by writing a `noop` event into any position it can prove was abandoned. A `noop` occupies its position and means nothing: replay steps over it without delivering it and without advancing the deterministic clock.
62
+
63
+ Two consequences for an implementation:
64
+
65
+ - **If you allocate at the commit, you are already compliant** and have nothing to build. No write can leave a position empty, so you have no holes to seal and will never emit a `noop`. `@workflow/world-local` and `@workflow/world-postgres` are in this position.
66
+ - **What the version actually gates is the reader.** A run stamped at the sealed-log version can only be replayed by a reader that knows to skip `noop`. That is every runtime on this release train, but a runtime pinning its own accepted range separately, such as the Python runtime, has to catch up first. `WORKFLOW_SEALED_LOG=0` is the switch for an environment where it has not.
67
+
68
+ Runs carry a spec version too, and a run keeps the version it was created under for its whole life. Read the stamped version off the run rather than assuming every run matches what your World declares today. Changing what you stamp does not reach runs already in your store: their version is persisted, every version test in the runtime is a lower bound, and a run's event ID scheme is resolved from what is stored.
60
69
 
61
70
  ## Interface changes
62
71
 
@@ -84,6 +93,14 @@ These do not change any signature, so an implementation ported by types alone wi
84
93
 
85
94
  **A stale replay no longer has to be refused.** v5 shipped with a `preconditionGuard` capability for a World that rejected an event creation whose snapshot was behind the log. It is gone, and nothing replaced it: allocating positions at the commit means a reader's log is a prefix rather than a prefix with a hole, replay is deterministic on a prefix, and a write reports the events it was pushed past. As a result, a stale replay costs a merge instead of a rejection. If you implemented the guard, you can delete it. `PreconditionFailedError` and the runtime's handling of it remain for a World that allocates positions away from the commit (see [Event ID allocation](#event-id-allocation)); no World in the SDK throws it.
86
95
 
96
+ **Process-wide state has to live on `globalThis`.** A module's top-level `const` or `let` is one instance per *module instance*, not per process, and a host server routinely holds several. Next.js compiles its server output into independent module graphs, and a bundled module is compiled into each one with its own module-scope bindings. Since `@workflow/world-vercel` moved from external to bundled, every module-scope singleton in it quietly became one per layer. The visible casualty was the WebSocket events transport: the queue consumer registered its channel in the route copy's registry while the write path looked it up in the instrumentation copy's empty one, so every event silently fell back to HTTP for the life of the process.
97
+
98
+ This bites rather than merely wasting memory because the runtime caches the *World object* process-wide while any module state that World closes over stays layer-local. Anything your World reaches at request time therefore has to be process-wide too: connection pools, transport registries, ID factories, caches, and log-once latches. Hold them in one object behind [`globalSingleton()`](https://github.com/vercel/workflow/blob/main/packages/utils/src/global-singleton.ts) from `@workflow/utils`, which keys the object off a `Symbol.for` on `globalThis`. A `let` cannot be shared by reference, so a latch becomes a field on that object.
99
+
100
+ **One World per process.** The workflow entrypoint's queue handler is now built from the runtime World that `getWorld()` returns, rather than from `getWorldHandlers()`. A stateful World is no longer instantiated twice in one process, so it stops getting duplicate connection pools and duplicate queue workers. If you added your own de-duplication to work around that, it is now redundant, though harmless if it keys on process-wide state.
101
+
102
+ **Your own transport is your own business, except for the tracing.** How a World ships events to its backend is unconstrained: `@workflow/world-vercel` can opt into a WebSocket and falls back to HTTP. What is constrained is what a reader of a trace sees. A non-HTTP transport still has to emit the per-event client span that an HTTP write would, or the per-event view of a run silently disappears. See [`WORKFLOW_EVENTS_TRANSPORT`](/worlds/vercel#workflow_events_transport) for the span shape and attributes the Vercel World uses, including a separate span for the handshake.
103
+
87
104
  **Replay reads the event log with `resolveData: 'skip-step-inputs'`.** A World may leave `input` out of `step_created` and `step_started` events for this value, and must otherwise treat it as `'all'`. A World that tests `resolveData === 'all'`, or validates against `['none', 'all']`, strips step results or rejects the read, and every replay fails. Test for `'none'` instead, or map with `entityResolveData()` from `@workflow/world`. `@workflow/world-testing` covers this case.
88
105
 
89
106
  **Event creation can return a delta.** `events.create()` may return events alongside the one it created, in `events` with a matching `cursor` and `hasMore`. The runtime uses this to skip a follow-up `events.list` round trip on `run_started`, on step-terminal writes that carried a `sinceCursor`, and on `hook_received` writes that carried `preloadEvents`. All three are advisory: a World that returns only the created event stays correct and pays one more round trip.
@@ -103,6 +120,8 @@ The scheme exists for what a reader can conclude from a log it just fetched: pos
103
120
  - **Bump and report.** `events.create()` params carry `eventCount`, so the expected position is `eventCount + 1`. When it is taken, do not reject the write: commit at the next free position and return the events you skipped on the success response. A stale count is the normal case for a parallel fan-out, and rejecting it would serialize writes the runtime deliberately issues concurrently.
104
121
  - **Allocate at the commit.** Take the position in the same operation that appends the event, not earlier. This is what makes a reader's log a prefix of the run's log rather than a prefix with a hole in it: nothing can land behind a position a reader has already passed. A World that mints a position in a request handler and commits later breaks the property every replay depends on, and is the only kind that still has a use for a stale-write rejection.
105
122
 
123
+ The one sanctioned exception is the sealed log, which is what the [sealed-log spec version](#sealed-logs-and-noop-events) is for: a World may pre-assign positions if it also seals the holes that leaves. Everything below assumes you allocate at the commit, which is the simpler contract and the one both first-party non-Vercel Worlds keep.
124
+
106
125
  [Event ID Allocation](/worlds/building-a-world#event-id-allocation) carries the full rules, and [Event IDs](/docs/how-it-works/event-sourcing#event-ids) covers what the format means for anything that reads an ID back.
107
126
 
108
127
  One consequence is specific to an upgrade, and it is the thing to plan around.
@@ -118,6 +137,8 @@ None of this is required. Each entry is a hook the runtime uses if your World pr
118
137
  | Member | What it buys |
119
138
  | --- | --- |
120
139
  | `capabilities` | Advertises `hookRetention.active`, `hookResumeDedup`, `hookForceClaim`, `deploymentAffinity`, `maxConcurrency`, and `dynamicWorkflowCode`. See the contract note above about failing closed. Event ID allocation is *not* in here: it is a requirement, not a capability. |
140
+ | `events.createBatch` | Appends an ordered list of events in one durable write, with a per-event outcome for each. Implementing the method *is* the declaration: the runtime folds a suspension's `step_created` / `wait_created` writes into batches only when it exists, and otherwise takes the single-event path unchanged. Implement it with real atomicity per attempt, so a lost race leaves nothing behind, or leave it out. See [Batched event writes](/docs/changelog/batched-event-writes). |
141
+ | `runs.waitForTerminalStatus` | Long-polls until a run reaches a terminal status. `await run.returnValue` uses it when present, instead of polling on an interval. |
121
142
  | `analytics` | A metadata-only read namespace for observability surfaces. Payload-bearing reads stay on `runs`, `steps`, `events`, and `hooks`. |
122
143
  | `runs.experimentalSetAttributes` | Backs `setAttributes()` from application code. Without it, run attributes are unavailable. |
123
144
  | `runs.cancelMany` | Bulk cancellation: up to 500 unique run IDs per request (`BULK_CANCEL_MAX_RUN_IDS`), an optional `cancelReason` of at most 512 characters, and a per-run outcome for every ID. Without it, the runtime falls back to bounded-concurrency individual cancels. |
@@ -140,7 +161,7 @@ Compiling workflow files changed independently of the storage contract.
140
161
  | Change | What to do |
141
162
  | --- | --- |
142
163
  | The `client` SWC transform mode was removed | It merged into `step` mode. Integrations passing `mode: 'client'` pass `mode: 'step'`. |
143
- | `stepEntrypoint` removed from `workflow/runtime` | Steps execute through the combined workflow handler the framework integrations generate. Custom hosts use `getWorldHandlers()`. |
164
+ | `stepEntrypoint` removed from `workflow/runtime` | Steps execute through the combined workflow handler the framework integrations generate. A custom host builds that handler from the World `getWorld()` returns. `getWorldHandlers()` still exists for the build-time view of a World, which is what a build integration wants; it is no longer how a request-time handler is assembled. |
144
165
  | Step, workflow and webhook bundles are ESM | Generated output moved from CJS to ESM, with a `createRequire` banner for CJS dependencies. The VM-executed workflow bundle stays CJS. The CLI's standalone output is renamed to match: `flow.mjs`, `webhook.mjs`, and `__step_registrations.mjs` in place of `flow.js`, `webhook.js`, and `step.js`. Consumers import the namespace rather than a default. |
145
166
  | `workflow/internal/private` and `@workflow/core/private` removed | These were never public API. The compiler no longer emits imports from them, so regenerate build output rather than importing them yourself. |
146
167
  | Duplicate step or workflow IDs fail the build | 4.x resolved collisions across non-exported workspace files last-write-wins. A build integration that derived IDs from a partial path may now produce build failures. |
@@ -246,10 +246,12 @@ Experimental stream-write transport capability. Default: `http`. Set `WORKFLOW_S
246
246
 
247
247
  This setting does not own tenant rollout policy and does not infer server support from package versions. HTTP remains the compatibility path. The protocol route (`/websockets/v1`) is versioned independently from REST v2/v4 and persisted workflow `specVersion` values.
248
248
 
249
- Writes and close requests execute serially. The first operation waits up to 250 ms for an opted-in socket to open; this is an implementation-level measurement knob, not protocol semantics. If the budget expires, or the upgrade fails or is declined before acceptance, the writer uses HTTP without racing the same operation over both transports. After the socket accepts a write, a missing acknowledgement has an unknown outcome: the writer fails rather than replaying the write over HTTP and risking a duplicate append.
249
+ Writes and close requests execute serially. The first operation waits up to 250 ms for an opted-in socket to open; this is an implementation-level measurement knob, not protocol semantics. If the budget expires, or the upgrade fails or is declined before acceptance, the writer uses HTTP without racing the same operation over both transports. After the socket accepts a write, a missing acknowledgement has an unknown outcome: the writer fails rather than replaying the write over HTTP and risking a duplicate append. A throttled (429) write or close did not apply, so the writer retries it after the server's `Retry-After`, as the HTTP writer does, on the socket if it stays open (or its replacement, after a drain) and otherwise over HTTP. Past 30 seconds of cumulative wait, the write moves to HTTP, whose own 429 retry applies. A close that fails with a 5xx is retried over HTTP, since close is idempotent.
250
250
 
251
251
  Before routine authentication expiry or server max duration, the server sends a v1 `drain` control. The client stops sending, lets an already-admitted request receive its reply, and reconnects after the server closes with code 1001. Authentication drains re-resolve a fresh bearer. A request that was sent but receives no reply before close still has an unknown outcome and is never replayed.
252
252
 
253
+ Each stream WebSocket message is at most [`WORKFLOW_WS_MAX_MESSAGE_BYTES`](#workflow_ws_max_message_bytes) bytes. A write group larger than the limit is sent as several ordered write requests. A single chunk too large for one message is written over HTTP instead, as are the rest of its writer's writes.
254
+
253
255
  ### `WORKFLOW_EVENTS_TRANSPORT`
254
256
 
255
257
  Opt-in WebSocket transport for workflow run events, which ships them to the Vercel World over one socket per run instead of one HTTP request each. Default: `http`.
@@ -258,6 +260,24 @@ Set `WORKFLOW_EVENTS_TRANSPORT=ws` to opt in. Only that value (case-insensitive)
258
260
 
259
261
  The setting is ignored when the World is configured with `projectConfig` and therefore routes through the `api-workflow` proxy: that endpoint is an HTTP-only REST gateway and does not forward a WebSocket upgrade, so events stay on HTTP. The fallback is silent, because the variable is typically set deployment-wide and a `projectConfig` World (such as the CLI) cannot act on it, and is reported once per process under `DEBUG=workflow:*`. `workflow.events.transport` on the per-write span records which transport actually carried a run.
260
262
 
263
+ Event writes use the WebSocket only while that run's events channel is open. The queue handler opens it for every delivery, so workflows need nothing extra. Code that writes a run's events outside a delivery, such as a custom driver or a long-lived process, opens the channel with `openEventsChannel(runId)`, or those writes go over HTTP. Call the returned release when you're done, because an open socket keeps the process alive:
264
+
265
+ {/*@skip-typecheck: incomplete code sample*/}
266
+
267
+ ```typescript title="write-events.ts" lineNumbers
268
+ import { createWorld, openEventsChannel } from "@workflow/world-vercel";
269
+
270
+ const world = createWorld();
271
+ const release = openEventsChannel(runId);
272
+ try {
273
+ await world.events.create(runId, event);
274
+ } finally {
275
+ release?.();
276
+ }
277
+ ```
278
+
279
+ `openEventsChannel()` returns `undefined`, and writes stay on HTTP, when the transport is disabled or the World cannot hold a socket, such as a `projectConfig` World.
280
+
261
281
  Tracing is unaffected by the choice. Each event write emits an `http POST` client span against the same `url.full`, regardless of which transport carries it. On the WebSocket path, the span is synthesized around the frame because no HTTP request is made. Attributes distinguish the transports:
262
282
 
263
283
  | Attribute | HTTP | WebSocket |
@@ -267,9 +287,33 @@ Tracing is unaffected by the choice. Each event write emits an `http POST` clien
267
287
  | `network.protocol.name` | — | `websocket` |
268
288
  | `workflow.events.ws.url` | — | the socket the frame went over |
269
289
  | `workflow.events.ws.req_id` | — | per-connection request id, matching the server's log line |
290
+ | `workflow.events.ws.request_parts` | — | messages the request was sent as, when it was split |
291
+ | `workflow.events.ws.reply_parts` | — | messages the reply arrived as, when it was split |
270
292
 
271
293
  The WebSocket handshake is itself a span, `workflow.events.ws.connect`, so the cost of opening (or eagerly reopening) a connection is attributable rather than showing up as unexplained time inside the first write.
272
294
 
295
+ ### `WORKFLOW_EVENTS_TRANSPORT_WS_OVERRIDE_WORKFLOWS`
296
+
297
+ Comma-separated workflows whose runs use the WebSocket events transport even when [`WORKFLOW_EVENTS_TRANSPORT`](#workflow_events_transport) is `http` or unset. Use it to move selected workflows onto the WebSocket while the rest of the deployment stays on HTTP. Default: none.
298
+
299
+ Each entry is a workflow's function name, such as `processOrder`, or its full workflow name, such as `workflow//./src/workflows/order//processOrder`. Matching is exact and case-sensitive. A function name matches every workflow with that name, in any file.
300
+
301
+ ```bash
302
+ WORKFLOW_EVENTS_TRANSPORT_WS_OVERRIDE_WORKFLOWS=processOrder,syncInventory
303
+ ```
304
+
305
+ The override applies to the events channel the queue handler opens for each delivery of a listed workflow's run. The workflow name comes from the delivery's queue, so it covers the run's workflow and step invocations alike. `openEventsChannel(runId, config, { workflowName })` applies it to channels you open yourself; without `workflowName`, only `WORKFLOW_EVENTS_TRANSPORT=ws` opens one. When `WORKFLOW_EVENTS_TRANSPORT=ws` is set, every workflow already uses the WebSocket and the override has no effect.
306
+
307
+ Like other environment variables, the setting is fixed for each deployment, and a run stays on the deployment it started on. Changing it takes effect for runs started on the next deployment; runs already in progress keep the transport of the deployment they're pinned to.
308
+
309
+ ### `WORKFLOW_WS_MAX_MESSAGE_BYTES`
310
+
311
+ Largest WebSocket message the Vercel World sends, header included, on both the events and stream-write transports. Default: `12582912` (12 MiB).
312
+
313
+ Some WebSocket paths limit the size of a single message, commonly to 16 MiB. Event payloads can be larger: a step's input or output is carried inside its event. A frame over this limit is sent as several messages, and the other side rebuilds it before handling it. Stream writes split a large write group into several ordered write requests instead (see [`WORKFLOW_STREAMS_TRANSPORT`](#workflow_streams_transport)). Frames at or under the limit are unaffected. Integer values are clamped to `2097152`-`16777216` (2–16 MiB) with a warning; other values use the default.
314
+
315
+ The client splits its own large requests, which needs a Vercel World backend that accepts split frames. It lists `frame-parts` in the `x-workflow-ws-flags` header of the WebSocket upgrade, and the backend splits large replies only for a client that does, so older clients keep receiving whole replies.
316
+
273
317
  ### Programmatic configuration
274
318
 
275
319
  `createWorld()` accepts explicit API configuration. It does not read `WORKFLOW_VERCEL_*` automatically, so pass the environment values yourself when you want a configured World module:
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "workflow",
3
- "version": "5.0.0",
3
+ "version": "5.0.1",
4
4
  "description": "Workflow SDK - Build durable, resilient, and observable workflows",
5
5
  "main": "dist/typescript-plugin.cjs",
6
6
  "type": "module",
@@ -58,19 +58,19 @@
58
58
  }
59
59
  },
60
60
  "dependencies": {
61
- "@workflow/astro": "5.0.0",
62
- "@workflow/cli": "5.0.0",
63
- "@workflow/core": "5.0.0",
64
- "@workflow/errors": "5.0.0",
61
+ "@workflow/astro": "5.0.1",
62
+ "@workflow/cli": "5.0.1",
63
+ "@workflow/core": "5.0.1",
64
+ "@workflow/errors": "5.0.1",
65
65
  "@workflow/typescript-plugin": "5.0.0",
66
66
  "@workflow/utils": "5.0.0",
67
67
  "ms": "2.1.3",
68
- "@workflow/next": "5.0.0",
69
- "@workflow/nest": "5.0.0",
70
- "@workflow/nitro": "5.0.0",
71
- "@workflow/nuxt": "5.0.0",
72
- "@workflow/sveltekit": "5.0.0",
73
- "@workflow/rollup": "5.0.0"
68
+ "@workflow/next": "5.0.1",
69
+ "@workflow/nest": "5.0.1",
70
+ "@workflow/nitro": "5.0.1",
71
+ "@workflow/nuxt": "5.0.1",
72
+ "@workflow/sveltekit": "5.0.1",
73
+ "@workflow/rollup": "5.0.1"
74
74
  },
75
75
  "devDependencies": {
76
76
  "@types/ms": "2.1.0",