workflow 5.0.0-beta.56 → 5.0.0-beta.58

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (38) hide show
  1. package/dist/api.d.ts +1 -1
  2. package/dist/api.d.ts.map +1 -1
  3. package/dist/api.js +1 -1
  4. package/docs/advanced/dynamic-workflows.mdx +224 -0
  5. package/docs/api-reference/workflow/create-hook.mdx +30 -1
  6. package/docs/api-reference/workflow/create-webhook.mdx +1 -1
  7. package/docs/api-reference/workflow/define-hook.mdx +2 -0
  8. package/docs/api-reference/workflow-api/register-lifecycle-hooks.mdx +6 -4
  9. package/docs/api-reference/workflow-api/start.mdx +49 -0
  10. package/docs/api-reference/workflow-errors/hook-conflict-error.mdx +1 -1
  11. package/docs/api-reference/workflow-errors/workflow-run-cancelled-error.mdx +7 -0
  12. package/docs/api-reference/workflow-errors/workflow-run-failed-error.mdx +7 -0
  13. package/docs/api-reference/workflow-runtime/health-check.mdx +1 -0
  14. package/docs/api-reference/workflow-runtime/world/storage.mdx +3 -1
  15. package/docs/changelog/batched-event-writes.mdx +2 -2
  16. package/docs/configuration/runtime-tuning.mdx +13 -3
  17. package/docs/configuration/worlds.mdx +7 -7
  18. package/docs/cookbook/advanced/child-workflows.mdx +3 -1
  19. package/docs/cookbook/advanced/upgrading-workflows.mdx +1 -1
  20. package/docs/cookbook/common-patterns/batching.mdx +2 -0
  21. package/docs/cookbook/common-patterns/sequential-and-parallel.mdx +2 -2
  22. package/docs/errors/hook-conflict.mdx +23 -0
  23. package/docs/errors/index.mdx +21 -0
  24. package/docs/foundations/errors-and-retries.mdx +3 -1
  25. package/docs/foundations/idempotency.mdx +44 -21
  26. package/docs/foundations/serialization.mdx +1 -0
  27. package/docs/foundations/starting-workflows.mdx +2 -0
  28. package/docs/how-it-works/event-sourcing.mdx +7 -1
  29. package/docs/meta.json +1 -0
  30. package/docs/observability/lifecycle-hooks.mdx +6 -4
  31. package/docs/whats-new.mdx +4 -1
  32. package/docs/worlds/building-a-world.mdx +538 -0
  33. package/docs/worlds/local.mdx +129 -0
  34. package/docs/worlds/meta.json +10 -0
  35. package/docs/worlds/postgres.mdx +424 -0
  36. package/docs/worlds/upgrading-to-v5.mdx +162 -0
  37. package/docs/worlds/vercel.mdx +345 -0
  38. package/package.json +11 -11
@@ -91,14 +91,14 @@ The Local World is the default outside Vercel and is intended for development.
91
91
  ### `WORKFLOW_LOCAL_HEADERS_TIMEOUT_MS`
92
92
 
93
93
  - Factory option: none
94
- - Default: `30000`
95
- - Maximum milliseconds to wait for a local queue handler to begin responding before the durable message is redelivered. Set to `0` to disable.
94
+ - Default: `0` (no deadline)
95
+ - Maximum milliseconds to wait for a local queue handler to begin responding before the durable message is redelivered. A value below your longest inline step re-executes that step while it is still running.
96
96
 
97
97
  ### `WORKFLOW_LOCAL_BODY_TIMEOUT_MS`
98
98
 
99
99
  - Factory option: none
100
- - Default: `30000`
101
- - Maximum gap in milliseconds between response body chunks from a local queue handler before the durable message is redelivered. Set to `0` to disable.
100
+ - Default: `0` (no deadline)
101
+ - Maximum gap in milliseconds between response body chunks from a local queue handler before the durable message is redelivered.
102
102
 
103
103
  ### `recoverActiveRuns`
104
104
 
@@ -304,12 +304,12 @@ Platform-provided values such as `VERCEL_DEPLOYMENT_ID`, `VERCEL_PROJECT_ID`, an
304
304
  - Default: on
305
305
  - Set to `0` (or `false`) to **disable** batched event writes, the escape hatch that restores the exact prior one-write-per-event path.
306
306
 
307
- When enabled (the default), a suspension's eager `step_created` and `wait_created` writes fold into batched `events.createBatch` calls (one durable write with per-event outcomes) on Worlds that implement the optional batch API. The fold only engages when the World implements `events.createBatch` (the Vercel World does; Local and Postgres do not), the run's spec version supports slot identity (≥ 6), and the suspension carries no attribute writes, hook writes, or resilient step dispatch. Everything else keeps the single-event path unchanged, so disabling is only needed as an operational escape hatch. Batches are capped at 32 events; larger fan-outs commit in successive batches. See the [batched event writes changelog](/docs/changelog/batched-event-writes) for the World API contract.
307
+ When enabled (the default), a suspension's eager `step_created` and `wait_created` writes fold into batched `events.createBatch` calls (one durable write with per-event outcomes) on Worlds that implement the optional batch API. The fold only engages when the World implements `events.createBatch` (the Vercel World does; Local and Postgres do not), the run's spec version supports slot identity (≥ 6), and the suspension carries no attribute writes or resilient step dispatch. Hook writes in the same suspension go through the single-event path concurrently with the batch. Everything else keeps the single-event path unchanged, so disabling is only needed as an operational escape hatch. Batches are capped at 32 events; larger fan-outs commit in successive batches. See the [batched event writes changelog](/docs/changelog/batched-event-writes) for the World API contract.
308
308
 
309
309
  ### `WORKFLOW_EVENTS_TRANSPORT`
310
310
 
311
311
  - Factory option: none
312
312
  - CLI flag: none
313
- - Default: `ws`
314
- - Ships workflow run events to the Vercel World over a WebSocket instead of one HTTP request each. Set to exactly `http` to opt out; any other value, including unset or empty, uses the WebSocket.
313
+ - Default: `http`
314
+ - Set to `ws` to ship workflow run events to the Vercel World over a WebSocket instead of one HTTP request each. Only `ws` (case-insensitive) opts in; any other value, including unset, empty, or `http`, keeps HTTP.
315
315
  - Ignored when the World is configured with `projectConfig` and routes through the `api-workflow` proxy: that endpoint is an HTTP-only REST gateway and does not forward a WebSocket upgrade, so events stay on HTTP.
@@ -19,8 +19,9 @@ Child workflows are the right choice when:
19
19
 
20
20
  - **Work units are independent**: Each child can run without knowing about the others (for example, when processing individual documents or generating separate reports).
21
21
  - **You need isolated failure boundaries**: A failing child should not abort unrelated work. The parent decides how to handle failures.
22
- - **You want large fan-out**: Spawning 50 or 500 children is practical because each runs on its own infrastructure.
22
+ - **You want large fan-out**: Spawning 50 or 500 children is practical because each runs on its own infrastructure. The parent still pays for each spawn, so start them in chunks rather than all at once (see [Chunked spawning](#fan-out-pattern-chunked-spawning)).
23
23
  - **You need per-item observability**: Each child workflow has its own run ID, status, and event log for monitoring.
24
+ - **One run would otherwise get too big**: A run's event log and step count are both capped, see [Vercel World limits](/worlds/vercel#per-run-limits). Replay reads the whole log, so split before the cap: a run headed for more than a few thousand events belongs in several runs.
24
25
 
25
26
  For simpler cases where steps share a single event log, use [direct await composition](/cookbook/common-patterns/workflow-composition#direct-await-flattening) instead.
26
27
 
@@ -306,6 +307,7 @@ async function startAndWaitWithRetries(
306
307
  - **Export wrapped children at module scope**: The SDK registers `"use workflow"` functions statically, so a runtime higher-order function returned from `withChildCompletionHook()` cannot be passed to `start()`.
307
308
  - **Use stable hook keys**: Document IDs, job IDs, or indexes prevent token collisions between parallel children in one parent run.
308
309
  - **Use chunked spawning for large batches**: Starting 500 children at once can create a large burst of work. Break the work into chunks of 10–50.
310
+ - **Watch the parent's log too**: Each child bounds its own log, but the parent records events for spawning and for collecting every child, so the parent's log grows with the number of children. Count those events per child against the parent's own [run limits](/worlds/vercel#per-run-limits); when the parent alone would exceed them, add a layer, so each parent spawns a modest number of intermediate runs that in turn spawn the leaves.
309
311
  - **Account for each child's retry semantics**: Steps inside child workflows retry independently. The parent sees the final `{ status, value | error }` payload from the hook.
310
312
  - **Use `deploymentId: "latest"` when children should run on the most recent deployment**: See [Versioning](/docs/foundations/versioning) for the full model and the [`start()` API reference](/docs/api-reference/workflow-api/start#using-deploymentid-latest) for compatibility considerations.
311
313
 
@@ -171,7 +171,7 @@ To upgrade a fleet of runs after a deployment, list active runs from a tracking
171
171
 
172
172
  ## How it works
173
173
 
174
- 1. **`deploymentId: "latest"` is the upgrade knob.** Without it, the spawn pins to the current deployment. With it, the new run resolves to whatever deployment is current when the runtime picks it up, so any shipped fix applies starting from that respawn. Both methods rely on this.
174
+ 1. **`deploymentId: "latest"` is the upgrade knob.** Without it, the spawn pins to the current deployment. With it, the new run resolves to whatever deployment is current when the runtime picks it up, so any shipped fix applies starting from that respawn. The respawned run runs at the spec version of the deployment it lands on. Both methods rely on this.
175
175
  2. **`start()` runs directly from the workflow body.** In v5, [`start()`](/docs/api-reference/workflow-api/start) is step-backed, so it can be called from a workflow function and still records a deterministic step boundary in the event log, so no manual `"use step"` wrapper is required.
176
176
  3. **State carries through the function argument.** The accumulating context flows from run N to run N+1 as a serialized argument. No external store is required for the state itself.
177
177
  4. **Per-run hook tokens.** Using `workflowRunId` as the hook token scopes each iteration's wait to its own run, so multiple chains can run concurrently without interfering.
@@ -17,6 +17,7 @@ Use batching when you need to process a large list of items in parallel while co
17
17
  - Processing hundreds or thousands of items against external APIs
18
18
  - Calling rate-limited APIs where you need to control concurrency
19
19
  - Any fan-out where you want failure isolation between groups
20
+ - High-concurrency fan-out, where one flat `Promise.all` over the whole list would put more work in flight than your downstream services or the run's event log should carry at once
20
21
 
21
22
  ## How it works
22
23
 
@@ -98,6 +99,7 @@ async function processRecord(record: Record): Promise<string> {
98
99
 
99
100
  - **Use `Promise.allSettled` instead of `Promise.all`**: Use this pattern when you want to continue even if some items fail. `Promise.all` rejects on the first failure, while `allSettled` waits for everything and identifies failures.
100
101
  - **Tune batch size to your downstream API limits**: If the API allows 10 concurrent requests, use `batchSize: 10`.
102
+ - **Batching bounds concurrency, not the run's total size**: Every batch still appends to the same [event log](/docs/how-it-works/event-sourcing#how-fast-a-log-grows), so a long enough list walks one run toward its [run limits](/worlds/vercel#per-run-limits) no matter how small the batches are. To shrink the run itself, bundle more items per step so one step covers many items, or spawn a [child workflow](/cookbook/advanced/child-workflows) per batch. Reach for one of those once a single run would grow past a few thousand events.
101
103
  - **Add pacing with `sleep()`**: Add a delay between batches to respect rate limits. The sleep is durable and survives cold starts.
102
104
  - **Treat each `processRecord` call as an independent step**: If one call fails, it retries up to three times without affecting other items in the batch.
103
105
 
@@ -10,7 +10,7 @@ related:
10
10
  ---
11
11
 
12
12
  <CopyPrompt
13
- text="Compose these workflow steps with standard async/await patterns. In the exported &quot;use workflow&quot; function, chain dependent &quot;use step&quot; calls with sequential `await`; run independent steps concurrently by starting them without `await` and awaiting `Promise.all([...])`; and use `Promise.race([...])` to act on whichever promise settles first. These compose with durable primitives: race a step or a webhook from `createWebhook()` against `sleep()` from `workflow` for deadlines. Keep every step input and output serializable, and remember `Promise.race` does not cancel the losing branch (it keeps running), so side-effectful losers need idempotency keys. Verify sequential ordering, parallel execution, and both race outcomes."
13
+ text="Compose these workflow steps with standard async/await patterns. In the exported &quot;use workflow&quot; function, chain dependent &quot;use step&quot; calls with sequential `await`; run independent steps concurrently by starting them without `await` and awaiting `Promise.all([...])`; and use `Promise.race([...])` to act on whichever promise settles first. These compose with durable primitives: race a step or a webhook from `createWebhook()` against `sleep()` from `workflow` for deadlines. Keep every step input and output serializable, and remember `Promise.race` does not cancel the losing branch (it keeps running), so side-effectful losers need idempotency keys. When the fan-out is wide, consider keeping concurrency bounded by chunking the array into batches or bundling several items per step rather than one flat `Promise.all` over everything. Verify sequential ordering, parallel execution, and both race outcomes."
14
14
  />
15
15
 
16
16
  Workflows are written in plain async/await: there's no new control-flow API to learn. Sequential awaits chain steps that depend on each other, `Promise.all` runs independent steps in parallel, and `Promise.race` returns whichever finishes first. These compose with workflow primitives like [`sleep()`](/docs/api-reference/workflow/sleep) and [`createWebhook()`](/docs/api-reference/workflow/create-webhook) since those are also promises.
@@ -144,7 +144,7 @@ export async function birthdayWorkflow(
144
144
  ## Adapting to your use case
145
145
 
146
146
  - **Replace `Promise.all` with `Promise.allSettled`**: Use this option when partial failures shouldn't abort the remaining operations. You'll get an array of `{ status, value | reason }` instead of an error on the first rejection.
147
- - **Bound the parallelism**: `Promise.all` over 1,000 items will fan out 1,000 concurrent steps. If your downstream APIs can't handle that, batch the array into chunks (see [Batching](/cookbook/common-patterns/batching)).
147
+ - **Bound the parallelism**: `Promise.all` over a large array fans out one concurrent step per item, all in the same run, and each item's events land in that one log. When concurrency gets high, batch the array into chunks or bundle several items into each step (see [Batching](/cookbook/common-patterns/batching)) so fewer, larger units of work are in flight. Splitting a wide fan-out across [child workflows](/cookbook/advanced/child-workflows) keeps each log shorter and isolates failures, but it does not by itself narrow the fan-out. Both the log and the step count are capped per run: see [Vercel World limits](/worlds/vercel#per-run-limits), and split before a run would grow past a few thousand events.
148
148
  - **Add a deadline to any race**: Pair the operation with `sleep("30s").then(() => "timeout" as const)` and check the discriminated result. See [Timeouts](/cookbook/common-patterns/timeouts).
149
149
  - **Mix steps and hooks in a race**: Wait for an external signal, a deadline, or a step result in the same `Promise.race`. The first promise to resolve wins.
150
150
 
@@ -156,6 +156,28 @@ export async function POST(request: Request) {
156
156
 
157
157
  If the caller needs live output instead of the final result, return `activeRun.getReadable()` from the same branch. If the duplicate request should replace the active work, call `await activeRun.cancel()` after inspecting the run.
158
158
 
159
+ ### Take the token over
160
+
161
+ If the newest run should always own the token, create the hook with [`experimental_force: true`](/docs/api-reference/workflow/create-hook#take-over-a-token-another-run-holds). The new run takes the token instead of getting `HookConflictError`, and the previous owner's `await hook` rejects with [`HookForceClaimedError`](/docs/errors/hook-force-claimed).
162
+
163
+ This enables zero-downtime transfers for hooks so one run can hand off a hook to another without dropping messages.
164
+
165
+ ```typescript lineNumbers
166
+ import { createHook } from "workflow";
167
+
168
+ export async function processPayment(orderId: string) {
169
+ "use workflow";
170
+
171
+ const hook = createHook({
172
+ token: `payment-${orderId}`,
173
+ experimental_force: true, // [!code highlight]
174
+ });
175
+ const payment = await hook;
176
+ }
177
+ ```
178
+
179
+ A run started at a Workflow spec version below 8, including runs started by older SDK releases, can't be taken from, so the forced hook still gets `HookConflictError` in that case.
180
+
159
181
  ## When hook tokens are released
160
182
 
161
183
  Hook tokens are automatically released when:
@@ -172,6 +194,7 @@ After a workflow completes, its hook tokens become available for reuse by other
172
194
  2. **Include unique identifiers** if you need custom tokens (order ID, user ID, etc.)
173
195
  3. **Avoid reusing the same token** across multiple concurrent workflow runs
174
196
  4. **Consider using webhooks** (`createWebhook`) if you need a fixed, predictable URL that can receive multiple payloads
197
+ 5. **Use `experimental_force`** when a newer run should replace the run holding the token
175
198
 
176
199
  ## Related
177
200
 
@@ -9,6 +9,27 @@ related:
9
9
 
10
10
  Fix common mistakes when creating and executing workflows in the **Workflow SDK**.
11
11
 
12
+ ## Error codes
13
+
14
+ When a workflow run fails, its `errorCode` identifies the failure category. You can read it from [`WorkflowRunFailedError`](/docs/api-reference/workflow-errors/workflow-run-failed-error), the Workflow CLI's `error.code` field, or the `workflow.error.code` OpenTelemetry span attribute.
15
+
16
+ | Code | Description |
17
+ | --- | --- |
18
+ | `USER_ERROR` | An error thrown by workflow or step code, including an unhandled step failure or `FatalError`. |
19
+ | `RUNTIME_ERROR` | The Workflow runtime encountered an internal error, such as missing runtime data or an invariant failure. Persistent occurrences should be [reported](https://github.com/vercel/workflow/issues). |
20
+ | [`CORRUPTED_EVENT_LOG`](/docs/errors/corrupted-event-log) | The run's event log cannot be replayed because it contains orphaned or mismatched events, a gap, or an unreadable stored payload. |
21
+ | [`REPLAY_DIVERGENCE`](/docs/errors/replay-divergence) | One replay could not consume the event log deterministically. The runtime automatically retries before treating repeated divergence as a corrupted event log. |
22
+ | `MAX_DELIVERIES_EXCEEDED` | The run exceeded the maximum number of queue deliveries, usually because a persistent failure kept causing redelivery. |
23
+ | `MAX_EVENTS_EXCEEDED` | The run reached the World's per-run event limit. Split unbounded work into child workflows before reaching the limit. |
24
+ | `REPLAY_TIMEOUT` | Workflow replay exceeded the configured duration limit. This measures workflow execution and event-log replay between step boundaries, not time spent inside step functions. |
25
+ | `STREAM_ERROR` | Workflow stream infrastructure failed while reading or writing data. This is an SDK or backend failure rather than an error in workflow code. |
26
+ | `WORLD_CONTRACT_ERROR` | A World returned data that violated the SDK contract and could not be retried safely. This usually indicates a World implementation bug. |
27
+ | [`DEPLOYMENT_MISMATCH`](/docs/errors/deployment-mismatch) | A run was delivered to a deployment other than the one it is pinned to, and automatic re-routing did not recover it. |
28
+
29
+ For guidance on catching failures, retry behavior, and inspecting `WorkflowRunFailedError`, see [Errors and retries](/docs/foundations/errors-and-retries).
30
+
31
+ ## Troubleshooting guides
32
+
12
33
  <AutoCards />
13
34
 
14
35
  ## Learn more
@@ -201,12 +201,14 @@ try {
201
201
  | Code | Meaning |
202
202
  | --- | --- |
203
203
  | `USER_ERROR` | An error thrown in your workflow or step code (including propagated step failures like `FatalError`) |
204
- | `MAX_EVENTS_EXCEEDED` | The run reached the World's per-run event ceiling (25,000 on the Local and Vercel Worlds). Split unbounded loops into [child workflows](/cookbook/advanced/child-workflows); see [Limits](/docs/configuration/runtime-tuning#limits) |
204
+ | `MAX_EVENTS_EXCEEDED` | The run reached the World's per-run event ceiling (for the Vercel World, see [Workflow run limits](https://vercel.com/docs/workflows/pricing#workflow-run-limits)). Split unbounded loops into [child workflows](/cookbook/advanced/child-workflows) well before the ceiling; see [Vercel World limits](/worlds/vercel#per-run-limits) and [Limits](/docs/configuration/runtime-tuning#limits) |
205
205
  | `MAX_DELIVERIES_EXCEEDED` | The run exceeded the maximum number of queue deliveries |
206
206
  | `REPLAY_TIMEOUT` | A workflow replay exceeded the maximum allowed duration |
207
207
  | `REPLAY_DIVERGENCE` | A replay could not consume the event log deterministically, usually because of non-deterministic workflow code. |
208
208
  | `CORRUPTED_EVENT_LOG` | The event log cannot be replayed: it contains orphaned or mismatched events, or one of its stored payloads is no longer readable from the World's storage. If you see this, please [file an issue](https://github.com/vercel/workflow/issues) |
209
+ | `STREAM_ERROR` | Workflow stream infrastructure failed while reading or writing data. This is an SDK or backend failure rather than an error in workflow code |
209
210
  | `WORLD_CONTRACT_ERROR` | A World response violated the SDK contract; points at a World implementation bug |
211
+ | `DEPLOYMENT_MISMATCH` | The run was delivered to a deployment other than the one it is pinned to, and automatic re-routing did not recover it |
210
212
  | `RUNTIME_ERROR` | An internal runtime error. If you see this, please [file an issue](https://github.com/vercel/workflow/issues) |
211
213
 
212
214
  <Callout type="info">
@@ -61,7 +61,7 @@ Because [hooks](/docs/foundations/hooks) already ensure globally unique active t
61
61
 
62
62
  Use a hook token as the idempotency key for an active workflow run. Hook tokens are globally unique while they are active: if another run tries to create a hook with the same token, the runtime records a conflict, `hook.getConflict()` resolves with a `Run` handle for the run that owns the token, and the hook rejects with [`HookConflictError`](/docs/errors/hook-conflict) when the workflow awaits or iterates its payload.
63
63
 
64
- The token should come from your domain, such as an order ID, invoice ID, import ID, or request ID. Create the hook near the beginning of the workflow and check `await hook.getConflict()` before doing duplicate-sensitive work that depends on owning the active token. Calling `createHook()` alone does not register the hook; awaiting `getConflict()` suspends the workflow to commit the registration.
64
+ The token should come from your domain, such as an order ID, invoice ID, import ID, or request ID. Create the hook near the beginning of the workflow and check `await hook.getConflict()` before doing duplicate-sensitive work that depends on owning the active token. Calling `createHook()` alone does not register the hook; awaiting `getConflict()` suspends the workflow to commit the registration. Check it before calling any step: steps the workflow calls before it suspends are started alongside the hook's registration, so a duplicate run that only learns of the conflict later (for example, by awaiting the hook and letting `HookConflictError` end the run) may already have started them. See [Registering a hook before a step uses it](/docs/api-reference/workflow/create-hook#registering-a-hook-before-a-step-uses-it).
65
65
 
66
66
  ```typescript lineNumbers
67
67
  import { createHook } from "workflow";
@@ -242,7 +242,7 @@ export async function processOrder(orderId: string, confirmed: boolean) {
242
242
  }
243
243
  ```
244
244
 
245
- **Supersede the owner.** Without minimum retention, cancel the active run, then claim the released token. The retry loop covers the window where cancellation cleanup has not propagated yet:
245
+ **Supersede the owner.** Create the Hook with [`experimental_force: true`](/docs/api-reference/workflow/create-hook#take-over-a-token-another-run-holds) so the newest run takes the token over from the active owner. This also works for a finished run holding the token under `experimental_minRetention`:
246
246
 
247
247
  ```typescript lineNumbers
248
248
  import { createHook } from "workflow";
@@ -253,31 +253,54 @@ declare function chargeOrder(orderId: string): Promise<void>; // @setup
253
253
  export async function processOrderNewestWins(orderId: string) {
254
254
  "use workflow";
255
255
 
256
- const token = `order:${orderId}`;
257
-
258
- for (let attempt = 0; attempt < 3; attempt++) {
259
- using request = createHook<OrderRequest>({ token });
260
-
261
- const conflict = await request.getConflict();
262
- if (!conflict) {
263
- // Token claimed: this run is now the owner.
264
- const { confirmed } = await request;
265
- if (confirmed) {
266
- await chargeOrder(orderId);
267
- }
268
- return { status: "processed" as const };
269
- }
256
+ using request = createHook<OrderRequest>({
257
+ token: `order:${orderId}`,
258
+ experimental_force: true, // [!code highlight]
259
+ });
270
260
 
271
- await conflict.cancel(); // [!code highlight]
261
+ // This run now owns the token.
262
+ const { confirmed } = await request;
263
+ if (confirmed) {
264
+ await chargeOrder(orderId);
272
265
  }
266
+ return { status: "processed" as const };
267
+ }
268
+ ```
269
+
270
+ The previous owner's `await request` rejects with [`HookForceClaimedError`](/docs/api-reference/workflow-errors/hook-force-claimed-error), and every `resumeHook()` for the token reaches the new run from then on. Any run of this workflow can be superseded by a later one, so catch the error and exit cleanly:
271
+
272
+ ```typescript lineNumbers
273
+ import { createHook } from "workflow";
274
+ import { HookForceClaimedError } from "workflow/errors";
273
275
 
274
- throw new Error(`Could not claim ${token} after canceling the owner`);
276
+ type OrderRequest = { confirmed: boolean };
277
+ declare function chargeOrder(orderId: string): Promise<void>; // @setup
278
+
279
+ export async function processOrderNewestWins(orderId: string) {
280
+ "use workflow";
281
+
282
+ using request = createHook<OrderRequest>({
283
+ token: `order:${orderId}`,
284
+ experimental_force: true,
285
+ });
286
+
287
+ try {
288
+ const { confirmed } = await request;
289
+ if (confirmed) {
290
+ await chargeOrder(orderId);
291
+ }
292
+ return { status: "processed" as const };
293
+ } catch (error) {
294
+ if (HookForceClaimedError.is(error)) { // [!code highlight]
295
+ // A newer run for this order owns the token now. Stop here.
296
+ return { status: "superseded" as const, runId: error.claimedByRunId }; // [!code highlight]
297
+ }
298
+ throw error;
299
+ }
275
300
  }
276
301
  ```
277
302
 
278
- <Callout type="warn">
279
- This pattern does not work with `experimental_minRetention`: canceling the old run does not make its token available early.
280
- </Callout>
303
+ A run started at a Workflow spec version below 8, including runs started by older SDK releases, can't be taken from. In that case the forced Hook rejects with `HookConflictError`, as it would without `experimental_force`.
281
304
 
282
305
  If duplicate requests should only reuse the active run without sending data, use [`getHookByToken()`](/docs/api-reference/workflow-api/get-hook-by-token) as an advisory pre-check before calling `start()`. The workflow should still check `hook.getConflict()`, because the lookup and `start()` are not atomic.
283
306
 
@@ -34,6 +34,7 @@ The following types can be serialized and passed through workflow functions:
34
34
  - `bigint`
35
35
  - `ArrayBuffer`
36
36
  - `BigInt64Array`, `BigUint64Array`
37
+ - `DataView`
37
38
  - `Date`
38
39
  - `Float32Array`, `Float64Array`
39
40
  - `Int8Array`, `Int16Array`, `Int32Array`
@@ -102,6 +102,8 @@ export async function parentWorkflow(inputValue: number) {
102
102
 
103
103
  When you call `start()` inside a workflow function, it automatically executes through an internal step to maintain deterministic replay. The returned `Run` object works as it does outside workflows. Properties such as `.runId`, `.status`, and `.returnValue`, and methods such as `.cancel()`, are all available. Each property access or method call executes as a separate step.
104
104
 
105
+ If the child run fails or is canceled, `await childRun.returnValue` throws [`WorkflowRunFailedError`](/docs/api-reference/workflow-errors/workflow-run-failed-error) or [`WorkflowRunCancelledError`](/docs/api-reference/workflow-errors/workflow-run-cancelled-error) on the first attempt.
106
+
105
107
  <Callout type="info">
106
108
  Inside workflow functions, each `Run` property access (e.g., `run.status`, `run.returnValue`) triggers a workflow step. This means each access is recorded in the event log and replayed deterministically.
107
109
  </Callout>
@@ -188,7 +188,7 @@ Events are categorized by the entity type they affect. Each event contains metad
188
188
  | `hook_created` | Creates a new hook in `active` state. Contains the hook token and optional metadata. |
189
189
  | `hook_conflict` | Records that hook creation failed because another run owns the token. Contains the token and, for current worlds, the owner's run ID. The hook is not created: `hook.getConflict()` resolves with the conflicting run, and awaiting the hook payload rejects with a `HookConflictError`. |
190
190
  | `hook_received` | Records that a payload was delivered to the hook. The hook remains `active` and can receive more payloads. |
191
- | `hook_disposed` | Deletes the hook from storage (conceptually transitioning to `disposed` state). The token is released for reuse by future workflows. |
191
+ | `hook_disposed` | Deletes the hook from storage (conceptually transitioning to `disposed` state). The token is released for reuse by future workflows. When another run takes the token over with [`experimental_force`](/docs/api-reference/workflow/create-hook#take-over-a-token-another-run-holds), the event names the new owner in `forceClaimedBy`, and the token moves to that run instead of being released. |
192
192
 
193
193
  ### Wait events
194
194
 
@@ -203,6 +203,12 @@ Events are categorized by the entity type they affect. Each event contains metad
203
203
  |-------|-------------|
204
204
  | `noop` | Seals an abandoned log position (`specVersion` 7 and above). Only the backend writes this event; the create endpoints reject it. See [Sealed positions](#sealed-positions-noop-events). |
205
205
 
206
+ ### How fast a log grows
207
+
208
+ Counting the events above is how you size a run. A step that succeeds on the first attempt contributes three events; a retry adds another `step_started`, preceded by a `step_retrying` when the World emits one; a sleep contributes two, and a hook at least two.
209
+
210
+ A run's log is capped — the [Vercel World](/worlds/vercel#per-run-limits) fails anything past its ceiling with `MAX_EVENTS_EXCEEDED`. Replay reads the whole log, so a run gets slower as it grows and is worth splitting well before it fails: split into [child workflows](/cookbook/advanced/child-workflows) once a run would grow past a few thousand events.
211
+
206
212
  ## Terminal states
207
213
 
208
214
  Terminal states represent the end of an entity's lifecycle. Once an entity reaches a terminal state, no further events can transition it to another state.
package/docs/meta.json CHANGED
@@ -5,6 +5,7 @@
5
5
  "getting-started",
6
6
  "foundations",
7
7
  "how-it-works",
8
+ "advanced/dynamic-workflows",
8
9
  "observability",
9
10
  "ai",
10
11
  "testing",
@@ -39,21 +39,23 @@ export async function register() {
39
39
 
40
40
  Keep the dynamic `workflow/api` import inside the `NEXT_RUNTIME === "nodejs"` guard. Next.js also compiles `instrumentation.ts` for the Edge runtime. A top-level static import pulls Node.js-only dependencies into that compilation and breaks webpack Edge builds, even if the registration call is guarded.
41
41
 
42
- `registerLifecycleHooks` returns an unregister function. You can register multiple hook sets, and handlers run in registration order.
42
+ `registerLifecycleHooks` returns an unregister function. You can register multiple hook sets, and handlers run in registration order. Registrations are not deduplicated: register each hook set once per process, and call its unregister function before registering it again during hot reload or module re-evaluation. Otherwise, repeated registrations invoke the same handler multiple times for each transition.
43
43
 
44
44
  ## Handler parameters
45
45
 
46
46
  Both handlers receive a `workflowName` string and the [`Run`](/docs/api-reference/workflow-api/get-run) instance for the transitioned run. `workflowName` is the machine-readable workflow identifier, such as `workflow//./src/workflows/order//processOrder`. Use this parameter to filter runs without a backend read. `run.runId` is also available without a read.
47
47
 
48
- The `Run` instance hydrates lazily. Accessors such as `run.workflowName`, `run.status`, and `run.returnValue` still fetch from the backend when used. Lazy access defers those reads; it does not make them free.
48
+ The `Run` instance hydrates lazily. Accessors such as `run.workflowName`, `run.status`, and `run.returnValue` still fetch from the backend when used. In particular, `workflowName` is a string, while `run.workflowName` is a `Promise<string>`. Lazy access defers those reads; it does not make them free.
49
49
 
50
- `onRunFailed` additionally receives the failure as a `WorkflowRunFailedError`, the same shape `run.returnValue` rejects with:
50
+ `onRunFailed` additionally receives a `WorkflowRunFailedError` hydrated for reporting. Unlike `run.returnValue`, it defers readable stream I/O and uses persisted abort snapshots:
51
51
 
52
52
  - `error.errorCode`: the failure classification (`USER_ERROR`, `RUNTIME_ERROR`, `MAX_DELIVERIES_EXCEEDED`, and more). See [error codes](/docs/errors) for the full list.
53
- - `error.cause`: the thrown value hydrated from the persisted error data, with registered Error subclass identity, message, stack, and cause chain preserved. Streamed values load lazily when consumed, and abort signals reflect their persisted state without live subscriptions. If hydration fails, the cause is a generic `Error`, matching `run.returnValue`'s fallback. Any JavaScript value can be thrown, so this is typed `unknown`.
53
+ - `error.cause`: the thrown value hydrated from the persisted error data, with registered Error subclass identity, message, stack, and cause chain preserved. Readable streams load lazily when consumed, and abort signals reflect their persisted state without live subscriptions. Writable streams retain their normal forwarding pipe and lock-polling setup during hydration. If hydration fails, the cause is a generic `Error`, matching `run.returnValue`'s fallback. Any JavaScript value can be thrown, so this is typed `unknown`.
54
54
 
55
55
  In `onRunFailed`, `run.returnValue` rejects because the run failed. Use `error.cause` to inspect or report the thrown value instead of awaiting `run.returnValue`.
56
56
 
57
+ The invocation's `waitUntil` scope includes background stream operations from the hydrated cause, even after a handler returns or throws. Close or release stream reader and writer locks when finished so that work can settle. Await other asynchronous reporting work in your handler, such as `Sentry.flush()` below, to keep it in the same lifetime scope.
58
+
57
59
  ## Reporting failed runs to Sentry
58
60
 
59
61
  This example reports failures for workflows named `processOrder`. Remove the filter to report failures from all workflows.
@@ -35,6 +35,8 @@ The largest change in v5 has no API surface: the runtime does far less work per
35
35
 
36
36
  **The workflow VM is kept alive across inline steps.** Within one invocation, a step- or attribute-driven suspension keeps the live VM and hydrated state, including when hooks or waits are open or created at the same boundary, so the next iteration appends only the newly written events instead of rebuilding the sandbox and replaying the whole log. Step inputs made of plain data or standard built-ins keep this fast path; see [`WORKFLOW_RETAINED_VM`](/docs/configuration/runtime-tuning#workflow_retained_vm).
37
37
 
38
+ **Replay no longer downloads recorded step inputs.** Replay recomputes step arguments by re-running workflow code, so it reads the event log without the `input` of each step. For workflows that pass growing state into their steps, this removes the part of the replay transfer that grew quadratically with the run. Worlds opt in through [`resolveData: 'skip-step-inputs'`](/docs/api-reference/workflow-runtime/world/storage#eventslist), and Vercel Workflows supports it.
39
+
38
40
  **Resuming a hook takes one round trip instead of two.** `resumeHook()` writes the `hook_received` event and dispatches the queue message concurrently, with a `(runId, resumeId)` dedup constraint keeping the two writers converging on exactly one event. See [Resilient hook resumption](/docs/changelog/resilient-resume).
39
41
 
40
42
  **Payloads are compressed** before they are encrypted and sent to the API. Repetitive payloads compress heavily; AI token streams average around 80% smaller. That is less stored data and less to move over the network.
@@ -166,13 +168,14 @@ All three first-party Worlds now implement it: Vercel accepts up to 30 days, and
166
168
  | Default trace mode is `linked` | Update dashboards that assume one trace per run, or set `WORKFLOW_TRACE_MODE=continuous`. |
167
169
  | The event-creation precondition guard is gone | `WORKFLOW_PRECONDITION_GUARD` no longer exists, and no World in the SDK rejects a write for a stale snapshot. Remove the variable if you set it. A replay that is behind now learns what it missed from the write it makes next instead of from a rejection, and [`PreconditionFailedError`](/docs/api-reference/workflow-errors/precondition-failed-error) remains only for a custom World that would still rather refuse. |
168
170
  | Event IDs are slot numbers, not ULIDs | An event ID is now its 1-based position in the run's log (`evnt_00000000000000000000000042`). It is unique only within a run, so pair it with the `runId` as a key, and it carries no timestamp: decoding one yields the Unix epoch rather than a creation time, so read `createdAt` off the event instead. Other entity IDs are unchanged. See [Event IDs](/docs/how-it-works/event-sourcing#event-ids). |
169
- | A per-run event limit is enforced | The World supplies the ceiling, which is 25,000 events on the Local and Vercel Worlds. A run that reaches it fails with `MAX_EVENTS_EXCEEDED`. Split unbounded loops into [child workflows](/cookbook/advanced/child-workflows). As a fallback, you can tune the ceiling. See [Limits](/docs/configuration/runtime-tuning#limits). |
171
+ | A per-run event limit is enforced | The World supplies the ceiling; for the Vercel World see [Workflow run limits](https://vercel.com/docs/workflows/pricing#workflow-run-limits). A run that reaches it fails with `MAX_EVENTS_EXCEEDED`. Split unbounded loops into [child workflows](/cookbook/advanced/child-workflows). [`WORKFLOW_MAX_EVENTS`](/docs/configuration/runtime-tuning#workflow_max_events) tunes the ceiling on the Local and Postgres Worlds; the Vercel World's is service-owned and the variable does not override it. |
170
172
  | Stream writes flush the first chunk immediately | The leading-edge flush window defaults to `0` instead of 10ms. Restore a window with `streamFlushIntervalMs` or `WORKFLOW_STREAM_FLUSH_INTERVAL_MS`. |
171
173
  | The workflow sandbox is stricter about nondeterminism | `WeakRef`, `FinalizationRegistry`, `Atomics.waitAsync`, and async `WebAssembly` compilation are no longer available inside workflow functions, and `crypto.subtle.digest` computes synchronously (same results, deterministic timing). Move code that needs them into a step. |
172
174
  | `Date()` without `new` returns a string inside workflow functions | This matches the language spec, and 4.x returned a `Date` object. Use `new Date()` where you need the object. Subclassing `Date` now works, so libraries like `TZDate` keep their identity across the sandbox boundary. |
173
175
  | `NestLocalBuilder` moved out of `@workflow/nest` root | Import it from `workflow/nest/builder`, so `WorkflowModule` no longer pulls the build toolchain into the runtime bundle. `NestVercelBuilder` lives at `workflow/nest/vercel-builder`. |
174
176
  | `workflow/internal/private` and `@workflow/core/private` removed | These were never public API. The compiler no longer emits imports from them, so regenerate build output rather than importing them yourself. |
175
177
  | The legacy trace viewer is gone from `@workflow/web-shared` | Only affects apps embedding the observability UI. `RunTraceView` and `WorkflowTraceViewer` are removed, and `NewTraceViewer` is now `TraceViewer` (module path `trace-viewer`). `Span`, `SpanEvent`, and `Trace` are still exported from the package root. |
178
+ | `startServer()` from `@workflow/web/server` resolves a `srvx` `Server` | Only affects apps that self-host the observability UI with it. Stop the server with `await server.close()`, and reach the Node `http.Server` at `server.node.server` to listen for its events. The server now also answers conditional and range requests and compresses responses. |
176
179
 
177
180
  Runs created on 4.x keep executing on the deployment that created them, so upgrading a deployment does not migrate in-flight runs. One storage caveat is worth knowing about: failed runs stored by `@workflow/world-postgres` before the upgrade read back with `error: undefined`, because the payload lives in the legacy `error` text column rather than `errorJson`.
178
181