workflow 5.0.0-beta.56 → 5.0.0-beta.57

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (31) hide show
  1. package/docs/api-reference/workflow/create-hook.mdx +30 -1
  2. package/docs/api-reference/workflow/create-webhook.mdx +1 -1
  3. package/docs/api-reference/workflow/define-hook.mdx +2 -0
  4. package/docs/api-reference/workflow-api/register-lifecycle-hooks.mdx +6 -4
  5. package/docs/api-reference/workflow-api/start.mdx +1 -0
  6. package/docs/api-reference/workflow-errors/hook-conflict-error.mdx +1 -1
  7. package/docs/api-reference/workflow-errors/workflow-run-cancelled-error.mdx +7 -0
  8. package/docs/api-reference/workflow-errors/workflow-run-failed-error.mdx +7 -0
  9. package/docs/api-reference/workflow-runtime/world/storage.mdx +3 -1
  10. package/docs/changelog/batched-event-writes.mdx +2 -2
  11. package/docs/configuration/runtime-tuning.mdx +3 -2
  12. package/docs/configuration/worlds.mdx +5 -5
  13. package/docs/cookbook/advanced/child-workflows.mdx +3 -1
  14. package/docs/cookbook/advanced/upgrading-workflows.mdx +1 -1
  15. package/docs/cookbook/common-patterns/batching.mdx +2 -0
  16. package/docs/cookbook/common-patterns/sequential-and-parallel.mdx +2 -2
  17. package/docs/errors/hook-conflict.mdx +23 -0
  18. package/docs/errors/index.mdx +21 -0
  19. package/docs/foundations/errors-and-retries.mdx +3 -1
  20. package/docs/foundations/idempotency.mdx +44 -21
  21. package/docs/foundations/starting-workflows.mdx +2 -0
  22. package/docs/how-it-works/event-sourcing.mdx +7 -1
  23. package/docs/observability/lifecycle-hooks.mdx +6 -4
  24. package/docs/whats-new.mdx +3 -1
  25. package/docs/worlds/building-a-world.mdx +537 -0
  26. package/docs/worlds/local.mdx +129 -0
  27. package/docs/worlds/meta.json +10 -0
  28. package/docs/worlds/postgres.mdx +424 -0
  29. package/docs/worlds/upgrading-to-v5.mdx +162 -0
  30. package/docs/worlds/vercel.mdx +345 -0
  31. package/package.json +11 -11
@@ -143,6 +143,35 @@ async function processOrder(orderId: string) {
143
143
 
144
144
  Because `createHook()` alone does not suspend the workflow, awaiting `hook.getConflict()` is what actually suspends the run and commits the hook registration. It only waits for registration. To receive payload data from a future `resumeHook()` call, await the hook itself or iterate it with `for await...of`.
145
145
 
146
+ ### Registering a hook before a step uses it
147
+
148
+ A hook's registration is committed alongside everything else the workflow started before it suspended, not ahead of it. When a workflow creates a hook and calls a step without awaiting anything in between, the step can start running before the hook is registered, and it can run even if the registration turns out to conflict. That matters in two cases:
149
+
150
+ - The step hands the token to something that may call `resumeHook()` right away, which throws `HookNotFoundError` until the hook exists.
151
+ - The hook guards against duplicate runs. A run that only learns of the conflict after calling the step, for example by awaiting the hook and letting `HookConflictError` end the run, may already have started that step.
152
+
153
+ In either case, await `hook.getConflict()` before calling the step:
154
+
155
+ ```typescript lineNumbers
156
+ import { createHook } from "workflow";
157
+
158
+ declare function requestApproval(token: string): Promise<void>; // @setup
159
+
160
+ async function approvalWorkflow() {
161
+ "use workflow";
162
+
163
+ using hook = createHook<{ approved: boolean }>();
164
+ await hook.getConflict(); // [!code highlight]
165
+
166
+ // The hook is registered, so an approver that resumes it immediately
167
+ // finds it.
168
+ await requestApproval(hook.token);
169
+
170
+ const { approved } = await hook;
171
+ return approved;
172
+ }
173
+ ```
174
+
146
175
  On a conflict, the resolved value is a `Run` handle for the run that owns the token, with durable step-backed accessors. The duplicate run can decide in code how to handle it: return or log `conflict.runId`, inspect `await conflict.status`, wait on `await conflict.returnValue`, or cancel the owner with `await conflict.cancel()` and continue in the current run. See [Run idempotency](/docs/foundations/idempotency#run-idempotency) for these strategies in context.
147
176
 
148
177
  <Callout type="info">
@@ -222,7 +251,7 @@ With `experimental_force`, this run always ends up owning the token:
222
251
  - Any number of runs forcing the same token at the same time converge on a single owner. The takeovers form a chain: each run that loses the token gets `HookForceClaimedError`, exactly one run ends up owning it, and none of them can get stuck. Which run wins among simultaneous claimers is not defined; if the order matters, start them in order.
223
252
  - A finished run that still holds the token under [`experimental_minRetention`](#keep-a-token-unavailable-after-the-run-ends) is taken over silently, since there is nothing left to wake. A run can also take over a token held by its own earlier Hook.
224
253
 
225
- The takeover is durable. If either run's compute fails partway through, the next request for the token completes it, so the token never ends up held by nobody or by both runs. The previous owner's wake is durable too: the new owner republishes it on replay until it records its next event, and the wake is idempotent, so a crash between registering the Hook and waking the previous owner is repaired by the new owner's next invocation.
254
+ The takeover is durable. If either run's compute fails partway through, the next request for the token completes it, so the token never ends up held by nobody or by both runs. The previous owner's wake is durable too: if the new owner's compute fails between registering the Hook and waking the previous owner, the new owner's next invocation republishes the wake, whatever else the new owner has recorded since (a step it started alongside the Hook, for example). Every invocation of the new owner within 24 hours of the takeover republishes it under the same idempotency key, which collapses the repeats into one wake; a repeat that does get through only replays the previous owner, which finds nothing new.
226
255
 
227
256
  <Callout type="info">
228
257
  A token can only be taken from a run whose runtime understands being taken from. Runs started at a Workflow spec version below 8, which includes every run started by an older SDK release, a Python SDK run, or a deployment with `WORKFLOW_SEALED_LOG=0`, would never learn that their Hook was disposed. The World declines to take their token and the forced Hook rejects with the ordinary [`HookConflictError`](/docs/api-reference/workflow-errors/hook-conflict-error) instead, exactly as if `experimental_force` had not been set. Finished runs holding a retained token are taken over at any version.
@@ -55,7 +55,7 @@ The returned `Webhook` object has:
55
55
 
56
56
  - `url`: The HTTP endpoint URL that external systems can call
57
57
  - `token`: The unique token identifying this webhook
58
- - `getConflict()`: A promise that resolves with the conflicting run if another active hook already owns this token, or `null` once the webhook endpoint has been registered
58
+ - `getConflict()`: A promise that resolves with the conflicting run if another active hook already owns this token, or `null` once the webhook endpoint has been registered. The endpoint is registered alongside the steps the workflow starts at the same time, not ahead of them, so await `getConflict()` before a step that hands `url` to a caller who may request it right away. See [Registering a hook before a step uses it](/docs/api-reference/workflow/create-hook#registering-a-hook-before-a-step-uses-it).
59
59
  - Implements `AsyncIterable<T>` for handling multiple requests, where `T` is `Request` (default) or `RequestWithResponse` (manual mode)
60
60
 
61
61
  When using `createWebhook({ respondWith: 'manual' })`, the resolved request type is `RequestWithResponse`, which extends the standard `Request` interface with a `respondWith(response: Response): Promise<void>` method for sending custom responses back to the caller.
@@ -211,6 +211,8 @@ export async function slackBotWorkflow(channelId: string) {
211
211
  }
212
212
  ```
213
213
 
214
+ `create()` accepts the same options as `createHook()`. If a newer run should replace one that still holds the token, pass [`experimental_force: true`](/docs/api-reference/workflow/create-hook#take-over-a-token-another-run-holds) to take the token over instead of getting [`HookConflictError`](/docs/api-reference/workflow-errors/hook-conflict-error).
215
+
214
216
  ## Related functions
215
217
 
216
218
  - [`createHook()`](/docs/api-reference/workflow/create-hook): Create a hook in a workflow.
@@ -49,11 +49,11 @@ showSections={["parameters"]}
49
49
 
50
50
  ### Returns
51
51
 
52
- Returns a function that unregisters these hooks.
52
+ Returns a function that unregisters these hooks. Registrations are not deduplicated. Register each hook set once per process, and unregister the previous hooks before registering again during hot reload or module re-evaluation.
53
53
 
54
54
  ## Handlers
55
55
 
56
- Both handlers receive a `workflowName` string and a lazily hydrated [`Run`](/docs/api-reference/workflow-api/get-run) instance. Use the `workflowName` parameter to filter without a backend read; `run.runId` also requires no read. Accessors such as `run.workflowName`, `run.status`, and `run.returnValue` still fetch from the backend when used. Lazy access defers those reads rather than eliminating them.
56
+ Both handlers receive a `workflowName` string and a lazily hydrated [`Run`](/docs/api-reference/workflow-api/get-run) instance. Use the `workflowName` parameter to filter without a backend read; `run.runId` also requires no read. Accessors such as `run.workflowName`, `run.status`, and `run.returnValue` still fetch from the backend when used. In particular, `workflowName` is a string, while `run.workflowName` is a `Promise<string>`. Lazy access defers those reads rather than eliminating them.
57
57
 
58
58
  ### `onRunCompleted`
59
59
 
@@ -72,9 +72,11 @@ Invoked when a workflow run fails terminally (after any retries).
72
72
  | --- | --- | --- |
73
73
  | `params.run` | `Run` | The failed run. |
74
74
  | `params.workflowName` | `string` | The machine-readable workflow identifier, such as `workflow//./src/workflows/order//processOrder`. Available without a backend read. |
75
- | `params.error` | `WorkflowRunFailedError` | The failure, in the same shape `run.returnValue` rejects with: `error.errorCode` carries the classification (e.g. `USER_ERROR`) and `error.cause` is the hydrated thrown value. |
75
+ | `params.error` | `WorkflowRunFailedError` | The persisted failure hydrated for reporting: `error.errorCode` carries the classification (e.g. `USER_ERROR`) and `error.cause` is the hydrated thrown value. |
76
76
 
77
- `error.cause` is hydrated from the persisted error data, with streamed values loaded lazily when consumed. Abort signals reflect their persisted state without live subscriptions. If hydration fails, the cause is a generic `Error`, matching `run.returnValue`'s fallback. In `onRunFailed`, `run.returnValue` rejects because the run failed. Use `error.cause` to inspect or report the thrown value instead.
77
+ Unlike `run.returnValue`, `error.cause` defers readable stream I/O until consumption and revives abort signals as persisted snapshots without live subscriptions. Writable streams retain their normal forwarding pipe and lock-polling setup during hydration. If hydration fails, the cause is a generic `Error`, matching `run.returnValue`'s fallback. In `onRunFailed`, `run.returnValue` rejects because the run failed. Use `error.cause` to inspect or report the thrown value instead.
78
+
79
+ The invocation's `waitUntil` scope includes background stream operations from the hydrated cause, even after a handler returns or throws. Close or release stream reader and writer locks when finished so that work can settle. Await other asynchronous reporting work in your handler to keep it in the same lifetime scope.
78
80
 
79
81
  ## Behavior
80
82
 
@@ -59,6 +59,7 @@ Learn more about [`WorkflowReadableStreamOptions`](/docs/api-reference/workflow-
59
59
  * Each call to `start()` creates a new workflow run. If retried requests must route to one active workflow, have the workflow create a deterministic hook token and use [`getHookByToken()`](/docs/api-reference/workflow-api/get-hook-by-token) to reuse an already-registered active hook. The lookup is not atomic with `start()`, so concurrent callers can still create extra runs before the hook is registered. Handle that race inside the workflow by checking `await hook.getConflict()` before duplicate-sensitive work. On a conflict, it resolves with the run that owns the token, so the duplicate can return the active owner to the caller. If duplicates must be rejected before a workflow body runs, keep a durable request record until native atomic start-and-hook registration exists. See [Idempotency](/docs/foundations/idempotency#run-idempotency).
60
60
  * All arguments must be [serializable](/docs/foundations/serialization).
61
61
  * When you provide `deploymentId`, the argument types and return type become `unknown` because the workflow function's types may differ across deployments.
62
+ * When `deploymentId` names a deployment other than the caller's, the run is stamped with the spec version the *target* deployment reports on its capability probe (capped at the caller's own), since the target is what executes it. `start()` waits up to 10 seconds for the first probe to a deployment and returns as soon as the target answers; later starts to the same deployment reuse the answer. If the target does not answer in time, the run is stamped with spec version 6, the lowest version a v5 runtime executes; a target on an older major version (such as `stable`) cannot execute such a run, so it logs a warning. In either case, the `attributes` and `experimental_retention` checks below apply to the stamped version, and fail naming the target deployment.
62
63
  * `attributes` seeds plaintext run metadata as part of creation and requires a World implementing spec version 4 or later. Keys that start with `$` are reserved for framework and library code; framework-level callers can pass `allowReservedAttributes: true` to seed reserved keys, with the same semantics as the [`setAttributes`](/docs/api-reference/workflow/set-attributes) option of the same name.
63
64
  * `region` pins the new run to a specific region on Worlds with a regional dimension. The [Vercel World](/worlds/vercel#explicit-region-selection) then serves the run's storage, queue dispatch, and streams from that region. When you omit `region`, the run is pinned to the region where it was created. Worlds without regions ignore the option.
64
65
  * `experimental_retention` asks the World to delete the run's user data as soon as the run completes or fails, instead of keeping it for the World's default window. `0` requests immediate deletion; `'default'` is identical to omitting the option. These are the only two values accepted — the value is a duration and zero is the only one implemented, and its unit is not yet decided. Recorded as the reserved `$retention` attribute, so it needs a World implementing spec version 4 or later. Retention is enforced by the World, not the SDK: the first-party Worlds implement it and a World that does not keeps the data. Note that `await run.returnValue` on a run started with `experimental_retention: 0` usually throws [`RunExpiredError`](/docs/errors/run-expired) rather than resolving, because the deletion races the read. See [Data retention](/docs/observability/retention).
@@ -9,7 +9,7 @@ related:
9
9
  - /docs/errors/hook-conflict
10
10
  ---
11
11
 
12
- `HookConflictError` is thrown when creating a hook with a token that is already in use by another active workflow run. Hook tokens must be unique across all running workflows. See the [hook-conflict](/docs/errors/hook-conflict) error guide for resolution strategies.
12
+ `HookConflictError` is thrown when creating a hook with a token that is already in use by another active workflow run. Hook tokens must be unique across all running workflows. See the [hook-conflict](/docs/errors/hook-conflict) error guide for resolution strategies. To take the token over instead, create the hook with [`experimental_force: true`](/docs/api-reference/workflow/create-hook#take-over-a-token-another-run-holds).
13
13
 
14
14
  ```typescript lineNumbers
15
15
  import { HookConflictError } from "workflow/errors"
@@ -12,6 +12,8 @@ related:
12
12
 
13
13
  You can check for cancellation before awaiting by inspecting `run.status`.
14
14
 
15
+ A canceled run is terminal, so this error is non-retryable (`fatal: true`). Inside a workflow, `await run.returnValue` runs as a step, and that step fails on its first attempt instead of spending its retry budget re-reading a run that cannot change. Errors from *failing to read* the run, such as a transport blip, stay retryable.
16
+
15
17
  ```typescript lineNumbers
16
18
  import { WorkflowRunCancelledError } from "workflow/errors"
17
19
  declare const run: { status: Promise<string>; returnValue: Promise<any> }; // @setup
@@ -34,6 +36,11 @@ definition={`
34
36
  interface WorkflowRunCancelledError {
35
37
  /** The ID of the canceled run. */
36
38
  runId: string;
39
+ /**
40
+ * Always \`true\`. A canceled run is terminal, so a step that reads one is
41
+ * not retried.
42
+ */
43
+ fatal: true;
37
44
  /** The error message. */
38
45
  message: string;
39
46
  }
@@ -13,6 +13,8 @@ related:
13
13
 
14
14
  The `cause` property holds the original thrown value, hydrated through the workflow serialization pipeline so its type identity (e.g. `FatalError`, `RetryableError`, custom `Error` subclasses), `cause` chain, and custom properties are preserved. Because any JavaScript value can be thrown, `cause` is typed as `unknown`, so narrow it with `instanceof Error` (or a more specific check) before accessing fields like `message`. The high-level error classification is exposed as the top-level `errorCode` property.
15
15
 
16
+ A failed run is terminal, so this error is non-retryable (`fatal: true`). Inside a workflow, `await run.returnValue` runs as a step, and that step fails on its first attempt instead of spending its retry budget re-reading a run that cannot change: the remote failure reaches the caller immediately, and the caller catches a `WorkflowRunFailedError` rather than a retry-exhaustion wrapper. Errors from *failing to read* the run, such as a transport blip, stay retryable.
17
+
16
18
  ```typescript lineNumbers
17
19
  import { WorkflowRunFailedError } from "workflow/errors"
18
20
  declare const run: { status: Promise<string>; returnValue: Promise<any> }; // @setup
@@ -50,6 +52,11 @@ interface WorkflowRunFailedError {
50
52
  cause: unknown;
51
53
  /** The high-level error category (e.g. \`USER_ERROR\`, \`RUNTIME_ERROR\`). */
52
54
  errorCode?: string;
55
+ /**
56
+ * Always \`true\`. A failed run is terminal, so a step that reads one is not
57
+ * retried.
58
+ */
59
+ fatal: true;
53
60
  /** The error message. */
54
61
  message: string;
55
62
  }
@@ -23,6 +23,7 @@ keywords:
23
23
  - Event
24
24
  - cursor pagination
25
25
  - resolveData
26
+ - skip-step-inputs
26
27
  - run_cancelled
27
28
  - correlation ID
28
29
  - parseStepName
@@ -94,7 +95,7 @@ const result = await world.events.list({ runId, pagination: { cursor } }); // [!
94
95
  | `params.pagination.cursor` | `string` | Cursor for the next page |
95
96
  | `params.pagination.limit` | `number` | Maximum events to return. When omitted, returns every remaining event up to the World's event ceiling. |
96
97
  | `params.pagination.sortOrder` | `"asc" \| "desc"` | Event order |
97
- | `params.resolveData` | `"all" \| "none"` | Include or omit event payload data |
98
+ | `params.resolveData` | `"all" \| "none" \| "skip-step-inputs"` | Include or omit event payload data. `"skip-step-inputs"` is `"all"` without the `input` of `step_created` and `step_started` events, which replay does not read. |
98
99
 
99
100
  **Returns:** `{ data: Event[], cursor: string | null, hasMore: boolean }`
100
101
 
@@ -359,6 +360,7 @@ const result = await world.hooks.list({ // [!code highlight]
359
360
  | `environment` | `string` | Deployment environment |
360
361
  | `metadata` | `object` | Custom metadata attached to the hook |
361
362
  | `isWebhook` | `boolean` | Whether this is a webhook-style hook |
363
+ | `claimedFrom` | `{ runId: string; hookId: string } \| undefined` | Set when this hook took its token from another run with [`experimental_force`](/docs/api-reference/workflow/create-hook#take-over-a-token-another-run-holds). Names that run and its hook |
362
364
 
363
365
  ---
364
366
 
@@ -62,11 +62,11 @@ The contract:
62
62
 
63
63
  ## The runtime integration (suspension fan-out fold)
64
64
 
65
- **On by default.** The suspension handler folds a **clean fan-out** (the suspension's eager `step_created` and `wait_created` writes) into `createBatch` calls of at most 32 events (mirroring the server's transaction budgets). Chunks of a larger fan-out commit **concurrently**: slot assignment is the World's, so parallel chunks race for slot ranges exactly like the pre-fold path's parallel single writes did, and per-entity conditions, not commit order, carry correctness. The fold only engages when the World implements `createBatch`, the run is on slot identity, and the suspension carries no attribute writes, no hook writes, and no resilient step dispatch; everything else keeps the single-event path byte-for-byte.
65
+ **On by default.** The suspension handler folds a **clean fan-out** (the suspension's eager `step_created` and `wait_created` writes) into `createBatch` calls of at most 32 events (mirroring the server's transaction budgets). Chunks of a larger fan-out commit **concurrently**: slot assignment is the World's, so parallel chunks race for slot ranges exactly like the pre-fold path's parallel single writes did, and per-entity conditions, not commit order, carry correctness. The fold only engages when the World implements `createBatch`, the run is on slot identity, and the suspension carries no attribute writes and no resilient step dispatch; everything else keeps the single-event path byte-for-byte. A suspension that also creates or disposes hooks still folds: hook writes are not batchable, so they go through the single-event path **concurrently** with the fold rather than ahead of it.
66
66
 
67
67
  **Per-chunk continuation.** Each chunk's follow-on work starts the moment **that chunk** commits, not when the whole fold does: a chunk's step-execution queue messages publish right off its own commit (publish-after-create holds per step), and only the chunk carrying the inline pairs gates the replay's continuation: trailing chunks' commits and publishes are joined before the invocation can acknowledge its message, so the durability contract ("every create durable before ack") is unchanged.
68
68
 
69
- **Pre-claimed inline pairs.** When the fold engages with at least two inline steps, the steps the runtime is about to execute inline join the batch as adjacent `[step_created, step_started]` pairs: the created row carrying the input, the started row a bare ownership-stamped claim the World folds into a born-running create. The pairs commit in a chunk of their own, ahead of the plain `step_created` and `wait_created` chunks, so the write the inline bodies wait for carries only two rows per inline step (a small transaction that commits faster than a full 32-event chunk) while the plain creates commit concurrently beside it. The inline bodies start straight off the pair chunk's commit (in parallel with the queue publishes and the sibling chunks) with no per-step claim POST at all, and a pair that loses its atomic create-claim to a concurrent delivery skips its body exactly as a lost lazy claim does. A lone inline step keeps the optimistic lazy-start path (one row, whose claim overlaps the body) even when eager creates batch beside it: the pairs share no round trip with those creates, so only two or more inline steps make a pair chunk worth the trade. A plain partition of exactly one `step_created` or `wait_created` beside the pairs is written through the ordinary single path rather than a one-row batch, and its queue message still waits for that write.
69
+ **Pre-claimed inline pairs.** When the fold engages with at least two inline steps, the steps the runtime is about to execute inline join the batch as adjacent `[step_created, step_started]` pairs: the created row carrying the input, the started row a bare ownership-stamped claim the World folds into a born-running create. The pairs commit in a chunk of their own, ahead of the plain `step_created` and `wait_created` chunks, so the write the inline bodies wait for carries only two rows per inline step (a small transaction that commits faster than a full 32-event chunk) while the plain creates commit concurrently beside it. The inline bodies start straight off the pair chunk's commit (in parallel with the queue publishes and the sibling chunks) with no per-step claim POST at all, and a pair that loses its atomic create-claim to a concurrent delivery skips its body exactly as a lost lazy claim does. A lone inline step keeps the optimistic lazy-start path (one row, whose claim overlaps the body) even when eager creates batch beside it: the pairs share no round trip with those creates, so only two or more inline steps make a pair chunk worth the trade. The exception is a lone inline step in a suspension that creates a hook: the runtime never starts a body before its claim settles while a hook is being created, and a lazy claim could only be sent after the hook write committed, so the step's pair is folded instead and its claim commits concurrently with the hook write. A plain partition of exactly one `step_created` or `wait_created` beside the pairs is written through the ordinary single path rather than a one-row batch, and its queue message still waits for that write.
70
70
 
71
71
  Per-event `409`s are tolerated the same way the single path tolerates `EntityConflictError` (a concurrent delivery already created the entity); any other per-event failure fails the suspension write the way a single-path rejection would. A batch carrying a `step_started` (that is, any batch with inline pairs) is **not** retried in-process on a transport blip: a pair's `409` cannot be told apart from the caller's own earlier attempt having committed it, so recovery goes through queue redelivery instead, where the step's ownership stamp routes it back to the same invocation.
72
72
 
@@ -73,7 +73,7 @@ For example, a workflow can run a 10-minute inline step even with `WORKFLOW_REPL
73
73
  - Default: `25000`
74
74
  - Positive-integer event limit reported by the Local World and enforced by the runtime as `MAX_EVENTS_EXCEEDED`.
75
75
  - The Local and Postgres Worlds also use it as the maximum number of events returned when `events.list()` is called without a limit. If more events exist, the response includes `hasMore: true` and a continuation cursor.
76
- - The Vercel World receives its event limit from the service; this environment variable does not override that service-owned value.
76
+ - The Vercel World receives its event limit from the service; this environment variable does not override that service-owned value. See [Vercel World limits](/worlds/vercel#per-run-limits).
77
77
  - Invalid or non-positive values fall back to the default.
78
78
 
79
79
  ### `WORKFLOW_REPLAY_DIVERGENCE_MAX_RETRIES`
@@ -129,6 +129,7 @@ For example, a workflow can run a 10-minute inline step even with `WORKFLOW_REPL
129
129
  - New runs are created at the sealed-log spec version, in which the World's backend assigns each event its position *before* the write commits rather than letting concurrent writers race for one. Concurrent writes then never contend for a position, which is what makes a wide fan-out cheap.
130
130
  - The price of assigning positions in advance is that a writer which claims one and then dies leaves a position no writer will ever fill. The backend closes such a position by writing a `noop` event into it once it can prove the position was abandoned, so a reader still sees the dense log it needs. Replay steps over a `noop` without delivering it to the workflow or advancing the deterministic clock. Its timestamp belongs to whichever reader sealed it, not to the run.
131
131
  - Set `0` to put a deployment back on the previous scheme, where each position is allocated by the write that occupies it. Use this as the kill switch if position assignment turns out to be at fault for event-log problems.
132
+ - Setting `0` also stamps new runs below spec version 8, so another run can't take their hook tokens over with [`experimental_force`](/docs/api-reference/workflow/create-hook#take-over-a-token-another-run-holds). A forced hook in another run gets `HookConflictError` instead. Runs created with `0` can still take tokens over from others.
132
133
  - Existing runs are unaffected either way. A run's spec version is stamped once, at creation, and read from the run for the rest of its life, so flipping this changes only what *new* runs get, and a run in flight keeps the scheme it started on. Every build reads sealed logs regardless of the setting.
133
134
  - A run created at the sealed-log version can only be replayed by a reader that knows to skip `noop` events. That includes every runtime on this release train, but a runtime that pins its own accepted spec range separately, such as the Python runtime, has to catch up before it can read these runs. Switch this off in an environment where it has not.
134
135
  - Only the Vercel World seals. The Local and Postgres Worlds allocate each position at the commit that occupies it, so they cannot leave a hole and never write a `noop`; the setting still moves the version they stamp, so the fleet stays on one spec.
@@ -386,4 +387,4 @@ These variables are primarily for tests, debugging, or unusual deployments.
386
387
  - Default: unset
387
388
  - Lowers the per-run event ceiling supplied by the World. A run whose event log reaches the ceiling fails with `MAX_EVENTS_EXCEEDED`, which stops a runaway loop from growing its log without bound.
388
389
  - Clamp-down only: it never raises the World's limit, and it applies even when the World supplies none. With no World limit and no override, nothing is enforced.
389
- - The Local and Vercel Worlds both supply a limit; the Local World defaults to 25,000 and is configurable with [`WORKFLOW_MAX_EVENTS`](/docs/configuration/worlds#workflow_max_events).
390
+ - The Local and Vercel Worlds both supply a limit; the Local World defaults to 25,000 and is configurable with [`WORKFLOW_MAX_EVENTS`](/docs/configuration/worlds#workflow_max_events). The Vercel World's ceiling is documented under [Workflow run limits](https://vercel.com/docs/workflows/pricing#workflow-run-limits).
@@ -91,14 +91,14 @@ The Local World is the default outside Vercel and is intended for development.
91
91
  ### `WORKFLOW_LOCAL_HEADERS_TIMEOUT_MS`
92
92
 
93
93
  - Factory option: none
94
- - Default: `30000`
95
- - Maximum milliseconds to wait for a local queue handler to begin responding before the durable message is redelivered. Set to `0` to disable.
94
+ - Default: `0` (no deadline)
95
+ - Maximum milliseconds to wait for a local queue handler to begin responding before the durable message is redelivered. A value below your longest inline step re-executes that step while it is still running.
96
96
 
97
97
  ### `WORKFLOW_LOCAL_BODY_TIMEOUT_MS`
98
98
 
99
99
  - Factory option: none
100
- - Default: `30000`
101
- - Maximum gap in milliseconds between response body chunks from a local queue handler before the durable message is redelivered. Set to `0` to disable.
100
+ - Default: `0` (no deadline)
101
+ - Maximum gap in milliseconds between response body chunks from a local queue handler before the durable message is redelivered.
102
102
 
103
103
  ### `recoverActiveRuns`
104
104
 
@@ -304,7 +304,7 @@ Platform-provided values such as `VERCEL_DEPLOYMENT_ID`, `VERCEL_PROJECT_ID`, an
304
304
  - Default: on
305
305
  - Set to `0` (or `false`) to **disable** batched event writes, the escape hatch that restores the exact prior one-write-per-event path.
306
306
 
307
- When enabled (the default), a suspension's eager `step_created` and `wait_created` writes fold into batched `events.createBatch` calls (one durable write with per-event outcomes) on Worlds that implement the optional batch API. The fold only engages when the World implements `events.createBatch` (the Vercel World does; Local and Postgres do not), the run's spec version supports slot identity (≥ 6), and the suspension carries no attribute writes, hook writes, or resilient step dispatch. Everything else keeps the single-event path unchanged, so disabling is only needed as an operational escape hatch. Batches are capped at 32 events; larger fan-outs commit in successive batches. See the [batched event writes changelog](/docs/changelog/batched-event-writes) for the World API contract.
307
+ When enabled (the default), a suspension's eager `step_created` and `wait_created` writes fold into batched `events.createBatch` calls (one durable write with per-event outcomes) on Worlds that implement the optional batch API. The fold only engages when the World implements `events.createBatch` (the Vercel World does; Local and Postgres do not), the run's spec version supports slot identity (≥ 6), and the suspension carries no attribute writes or resilient step dispatch. Hook writes in the same suspension go through the single-event path concurrently with the batch. Everything else keeps the single-event path unchanged, so disabling is only needed as an operational escape hatch. Batches are capped at 32 events; larger fan-outs commit in successive batches. See the [batched event writes changelog](/docs/changelog/batched-event-writes) for the World API contract.
308
308
 
309
309
  ### `WORKFLOW_EVENTS_TRANSPORT`
310
310
 
@@ -19,8 +19,9 @@ Child workflows are the right choice when:
19
19
 
20
20
  - **Work units are independent**: Each child can run without knowing about the others (for example, when processing individual documents or generating separate reports).
21
21
  - **You need isolated failure boundaries**: A failing child should not abort unrelated work. The parent decides how to handle failures.
22
- - **You want large fan-out**: Spawning 50 or 500 children is practical because each runs on its own infrastructure.
22
+ - **You want large fan-out**: Spawning 50 or 500 children is practical because each runs on its own infrastructure. The parent still pays for each spawn, so start them in chunks rather than all at once (see [Chunked spawning](#fan-out-pattern-chunked-spawning)).
23
23
  - **You need per-item observability**: Each child workflow has its own run ID, status, and event log for monitoring.
24
+ - **One run would otherwise get too big**: A run's event log and step count are both capped, see [Vercel World limits](/worlds/vercel#per-run-limits). Replay reads the whole log, so split before the cap: a run headed for more than a few thousand events belongs in several runs.
24
25
 
25
26
  For simpler cases where steps share a single event log, use [direct await composition](/cookbook/common-patterns/workflow-composition#direct-await-flattening) instead.
26
27
 
@@ -306,6 +307,7 @@ async function startAndWaitWithRetries(
306
307
  - **Export wrapped children at module scope**: The SDK registers `"use workflow"` functions statically, so a runtime higher-order function returned from `withChildCompletionHook()` cannot be passed to `start()`.
307
308
  - **Use stable hook keys**: Document IDs, job IDs, or indexes prevent token collisions between parallel children in one parent run.
308
309
  - **Use chunked spawning for large batches**: Starting 500 children at once can create a large burst of work. Break the work into chunks of 10–50.
310
+ - **Watch the parent's log too**: Each child bounds its own log, but the parent records events for spawning and for collecting every child, so the parent's log grows with the number of children. Count those events per child against the parent's own [run limits](/worlds/vercel#per-run-limits); when the parent alone would exceed them, add a layer, so each parent spawns a modest number of intermediate runs that in turn spawn the leaves.
309
311
  - **Account for each child's retry semantics**: Steps inside child workflows retry independently. The parent sees the final `{ status, value | error }` payload from the hook.
310
312
  - **Use `deploymentId: "latest"` when children should run on the most recent deployment**: See [Versioning](/docs/foundations/versioning) for the full model and the [`start()` API reference](/docs/api-reference/workflow-api/start#using-deploymentid-latest) for compatibility considerations.
311
313
 
@@ -171,7 +171,7 @@ To upgrade a fleet of runs after a deployment, list active runs from a tracking
171
171
 
172
172
  ## How it works
173
173
 
174
- 1. **`deploymentId: "latest"` is the upgrade knob.** Without it, the spawn pins to the current deployment. With it, the new run resolves to whatever deployment is current when the runtime picks it up, so any shipped fix applies starting from that respawn. Both methods rely on this.
174
+ 1. **`deploymentId: "latest"` is the upgrade knob.** Without it, the spawn pins to the current deployment. With it, the new run resolves to whatever deployment is current when the runtime picks it up, so any shipped fix applies starting from that respawn. The respawned run runs at the spec version of the deployment it lands on. Both methods rely on this.
175
175
  2. **`start()` runs directly from the workflow body.** In v5, [`start()`](/docs/api-reference/workflow-api/start) is step-backed, so it can be called from a workflow function and still records a deterministic step boundary in the event log, so no manual `"use step"` wrapper is required.
176
176
  3. **State carries through the function argument.** The accumulating context flows from run N to run N+1 as a serialized argument. No external store is required for the state itself.
177
177
  4. **Per-run hook tokens.** Using `workflowRunId` as the hook token scopes each iteration's wait to its own run, so multiple chains can run concurrently without interfering.
@@ -17,6 +17,7 @@ Use batching when you need to process a large list of items in parallel while co
17
17
  - Processing hundreds or thousands of items against external APIs
18
18
  - Calling rate-limited APIs where you need to control concurrency
19
19
  - Any fan-out where you want failure isolation between groups
20
+ - High-concurrency fan-out, where one flat `Promise.all` over the whole list would put more work in flight than your downstream services or the run's event log should carry at once
20
21
 
21
22
  ## How it works
22
23
 
@@ -98,6 +99,7 @@ async function processRecord(record: Record): Promise<string> {
98
99
 
99
100
  - **Use `Promise.allSettled` instead of `Promise.all`**: Use this pattern when you want to continue even if some items fail. `Promise.all` rejects on the first failure, while `allSettled` waits for everything and identifies failures.
100
101
  - **Tune batch size to your downstream API limits**: If the API allows 10 concurrent requests, use `batchSize: 10`.
102
+ - **Batching bounds concurrency, not the run's total size**: Every batch still appends to the same [event log](/docs/how-it-works/event-sourcing#how-fast-a-log-grows), so a long enough list walks one run toward its [run limits](/worlds/vercel#per-run-limits) no matter how small the batches are. To shrink the run itself, bundle more items per step so one step covers many items, or spawn a [child workflow](/cookbook/advanced/child-workflows) per batch. Reach for one of those once a single run would grow past a few thousand events.
101
103
  - **Add pacing with `sleep()`**: Add a delay between batches to respect rate limits. The sleep is durable and survives cold starts.
102
104
  - **Treat each `processRecord` call as an independent step**: If one call fails, it retries up to three times without affecting other items in the batch.
103
105
 
@@ -10,7 +10,7 @@ related:
10
10
  ---
11
11
 
12
12
  <CopyPrompt
13
- text="Compose these workflow steps with standard async/await patterns. In the exported &quot;use workflow&quot; function, chain dependent &quot;use step&quot; calls with sequential `await`; run independent steps concurrently by starting them without `await` and awaiting `Promise.all([...])`; and use `Promise.race([...])` to act on whichever promise settles first. These compose with durable primitives: race a step or a webhook from `createWebhook()` against `sleep()` from `workflow` for deadlines. Keep every step input and output serializable, and remember `Promise.race` does not cancel the losing branch (it keeps running), so side-effectful losers need idempotency keys. Verify sequential ordering, parallel execution, and both race outcomes."
13
+ text="Compose these workflow steps with standard async/await patterns. In the exported &quot;use workflow&quot; function, chain dependent &quot;use step&quot; calls with sequential `await`; run independent steps concurrently by starting them without `await` and awaiting `Promise.all([...])`; and use `Promise.race([...])` to act on whichever promise settles first. These compose with durable primitives: race a step or a webhook from `createWebhook()` against `sleep()` from `workflow` for deadlines. Keep every step input and output serializable, and remember `Promise.race` does not cancel the losing branch (it keeps running), so side-effectful losers need idempotency keys. When the fan-out is wide, consider keeping concurrency bounded by chunking the array into batches or bundling several items per step rather than one flat `Promise.all` over everything. Verify sequential ordering, parallel execution, and both race outcomes."
14
14
  />
15
15
 
16
16
  Workflows are written in plain async/await: there's no new control-flow API to learn. Sequential awaits chain steps that depend on each other, `Promise.all` runs independent steps in parallel, and `Promise.race` returns whichever finishes first. These compose with workflow primitives like [`sleep()`](/docs/api-reference/workflow/sleep) and [`createWebhook()`](/docs/api-reference/workflow/create-webhook) since those are also promises.
@@ -144,7 +144,7 @@ export async function birthdayWorkflow(
144
144
  ## Adapting to your use case
145
145
 
146
146
  - **Replace `Promise.all` with `Promise.allSettled`**: Use this option when partial failures shouldn't abort the remaining operations. You'll get an array of `{ status, value | reason }` instead of an error on the first rejection.
147
- - **Bound the parallelism**: `Promise.all` over 1,000 items will fan out 1,000 concurrent steps. If your downstream APIs can't handle that, batch the array into chunks (see [Batching](/cookbook/common-patterns/batching)).
147
+ - **Bound the parallelism**: `Promise.all` over a large array fans out one concurrent step per item, all in the same run, and each item's events land in that one log. When concurrency gets high, batch the array into chunks or bundle several items into each step (see [Batching](/cookbook/common-patterns/batching)) so fewer, larger units of work are in flight. Splitting a wide fan-out across [child workflows](/cookbook/advanced/child-workflows) keeps each log shorter and isolates failures, but it does not by itself narrow the fan-out. Both the log and the step count are capped per run: see [Vercel World limits](/worlds/vercel#per-run-limits), and split before a run would grow past a few thousand events.
148
148
  - **Add a deadline to any race**: Pair the operation with `sleep("30s").then(() => "timeout" as const)` and check the discriminated result. See [Timeouts](/cookbook/common-patterns/timeouts).
149
149
  - **Mix steps and hooks in a race**: Wait for an external signal, a deadline, or a step result in the same `Promise.race`. The first promise to resolve wins.
150
150
 
@@ -156,6 +156,28 @@ export async function POST(request: Request) {
156
156
 
157
157
  If the caller needs live output instead of the final result, return `activeRun.getReadable()` from the same branch. If the duplicate request should replace the active work, call `await activeRun.cancel()` after inspecting the run.
158
158
 
159
+ ### Take the token over
160
+
161
+ If the newest run should always own the token, create the hook with [`experimental_force: true`](/docs/api-reference/workflow/create-hook#take-over-a-token-another-run-holds). The new run takes the token instead of getting `HookConflictError`, and the previous owner's `await hook` rejects with [`HookForceClaimedError`](/docs/errors/hook-force-claimed).
162
+
163
+ This enables zero-downtime transfers for hooks so one run can hand off a hook to another without dropping messages.
164
+
165
+ ```typescript lineNumbers
166
+ import { createHook } from "workflow";
167
+
168
+ export async function processPayment(orderId: string) {
169
+ "use workflow";
170
+
171
+ const hook = createHook({
172
+ token: `payment-${orderId}`,
173
+ experimental_force: true, // [!code highlight]
174
+ });
175
+ const payment = await hook;
176
+ }
177
+ ```
178
+
179
+ A run started at a Workflow spec version below 8, including runs started by older SDK releases, can't be taken from, so the forced hook still gets `HookConflictError` in that case.
180
+
159
181
  ## When hook tokens are released
160
182
 
161
183
  Hook tokens are automatically released when:
@@ -172,6 +194,7 @@ After a workflow completes, its hook tokens become available for reuse by other
172
194
  2. **Include unique identifiers** if you need custom tokens (order ID, user ID, etc.)
173
195
  3. **Avoid reusing the same token** across multiple concurrent workflow runs
174
196
  4. **Consider using webhooks** (`createWebhook`) if you need a fixed, predictable URL that can receive multiple payloads
197
+ 5. **Use `experimental_force`** when a newer run should replace the run holding the token
175
198
 
176
199
  ## Related
177
200
 
@@ -9,6 +9,27 @@ related:
9
9
 
10
10
  Fix common mistakes when creating and executing workflows in the **Workflow SDK**.
11
11
 
12
+ ## Error codes
13
+
14
+ When a workflow run fails, its `errorCode` identifies the failure category. You can read it from [`WorkflowRunFailedError`](/docs/api-reference/workflow-errors/workflow-run-failed-error), the Workflow CLI's `error.code` field, or the `workflow.error.code` OpenTelemetry span attribute.
15
+
16
+ | Code | Description |
17
+ | --- | --- |
18
+ | `USER_ERROR` | An error thrown by workflow or step code, including an unhandled step failure or `FatalError`. |
19
+ | `RUNTIME_ERROR` | The Workflow runtime encountered an internal error, such as missing runtime data or an invariant failure. Persistent occurrences should be [reported](https://github.com/vercel/workflow/issues). |
20
+ | [`CORRUPTED_EVENT_LOG`](/docs/errors/corrupted-event-log) | The run's event log cannot be replayed because it contains orphaned or mismatched events, a gap, or an unreadable stored payload. |
21
+ | [`REPLAY_DIVERGENCE`](/docs/errors/replay-divergence) | One replay could not consume the event log deterministically. The runtime automatically retries before treating repeated divergence as a corrupted event log. |
22
+ | `MAX_DELIVERIES_EXCEEDED` | The run exceeded the maximum number of queue deliveries, usually because a persistent failure kept causing redelivery. |
23
+ | `MAX_EVENTS_EXCEEDED` | The run reached the World's per-run event limit. Split unbounded work into child workflows before reaching the limit. |
24
+ | `REPLAY_TIMEOUT` | Workflow replay exceeded the configured duration limit. This measures workflow execution and event-log replay between step boundaries, not time spent inside step functions. |
25
+ | `STREAM_ERROR` | Workflow stream infrastructure failed while reading or writing data. This is an SDK or backend failure rather than an error in workflow code. |
26
+ | `WORLD_CONTRACT_ERROR` | A World returned data that violated the SDK contract and could not be retried safely. This usually indicates a World implementation bug. |
27
+ | [`DEPLOYMENT_MISMATCH`](/docs/errors/deployment-mismatch) | A run was delivered to a deployment other than the one it is pinned to, and automatic re-routing did not recover it. |
28
+
29
+ For guidance on catching failures, retry behavior, and inspecting `WorkflowRunFailedError`, see [Errors and retries](/docs/foundations/errors-and-retries).
30
+
31
+ ## Troubleshooting guides
32
+
12
33
  <AutoCards />
13
34
 
14
35
  ## Learn more
@@ -201,12 +201,14 @@ try {
201
201
  | Code | Meaning |
202
202
  | --- | --- |
203
203
  | `USER_ERROR` | An error thrown in your workflow or step code (including propagated step failures like `FatalError`) |
204
- | `MAX_EVENTS_EXCEEDED` | The run reached the World's per-run event ceiling (25,000 on the Local and Vercel Worlds). Split unbounded loops into [child workflows](/cookbook/advanced/child-workflows); see [Limits](/docs/configuration/runtime-tuning#limits) |
204
+ | `MAX_EVENTS_EXCEEDED` | The run reached the World's per-run event ceiling (for the Vercel World, see [Workflow run limits](https://vercel.com/docs/workflows/pricing#workflow-run-limits)). Split unbounded loops into [child workflows](/cookbook/advanced/child-workflows) well before the ceiling; see [Vercel World limits](/worlds/vercel#per-run-limits) and [Limits](/docs/configuration/runtime-tuning#limits) |
205
205
  | `MAX_DELIVERIES_EXCEEDED` | The run exceeded the maximum number of queue deliveries |
206
206
  | `REPLAY_TIMEOUT` | A workflow replay exceeded the maximum allowed duration |
207
207
  | `REPLAY_DIVERGENCE` | A replay could not consume the event log deterministically, usually because of non-deterministic workflow code. |
208
208
  | `CORRUPTED_EVENT_LOG` | The event log cannot be replayed: it contains orphaned or mismatched events, or one of its stored payloads is no longer readable from the World's storage. If you see this, please [file an issue](https://github.com/vercel/workflow/issues) |
209
+ | `STREAM_ERROR` | Workflow stream infrastructure failed while reading or writing data. This is an SDK or backend failure rather than an error in workflow code |
209
210
  | `WORLD_CONTRACT_ERROR` | A World response violated the SDK contract; points at a World implementation bug |
211
+ | `DEPLOYMENT_MISMATCH` | The run was delivered to a deployment other than the one it is pinned to, and automatic re-routing did not recover it |
210
212
  | `RUNTIME_ERROR` | An internal runtime error. If you see this, please [file an issue](https://github.com/vercel/workflow/issues) |
211
213
 
212
214
  <Callout type="info">
@@ -61,7 +61,7 @@ Because [hooks](/docs/foundations/hooks) already ensure globally unique active t
61
61
 
62
62
  Use a hook token as the idempotency key for an active workflow run. Hook tokens are globally unique while they are active: if another run tries to create a hook with the same token, the runtime records a conflict, `hook.getConflict()` resolves with a `Run` handle for the run that owns the token, and the hook rejects with [`HookConflictError`](/docs/errors/hook-conflict) when the workflow awaits or iterates its payload.
63
63
 
64
- The token should come from your domain, such as an order ID, invoice ID, import ID, or request ID. Create the hook near the beginning of the workflow and check `await hook.getConflict()` before doing duplicate-sensitive work that depends on owning the active token. Calling `createHook()` alone does not register the hook; awaiting `getConflict()` suspends the workflow to commit the registration.
64
+ The token should come from your domain, such as an order ID, invoice ID, import ID, or request ID. Create the hook near the beginning of the workflow and check `await hook.getConflict()` before doing duplicate-sensitive work that depends on owning the active token. Calling `createHook()` alone does not register the hook; awaiting `getConflict()` suspends the workflow to commit the registration. Check it before calling any step: steps the workflow calls before it suspends are started alongside the hook's registration, so a duplicate run that only learns of the conflict later (for example, by awaiting the hook and letting `HookConflictError` end the run) may already have started them. See [Registering a hook before a step uses it](/docs/api-reference/workflow/create-hook#registering-a-hook-before-a-step-uses-it).
65
65
 
66
66
  ```typescript lineNumbers
67
67
  import { createHook } from "workflow";
@@ -242,7 +242,7 @@ export async function processOrder(orderId: string, confirmed: boolean) {
242
242
  }
243
243
  ```
244
244
 
245
- **Supersede the owner.** Without minimum retention, cancel the active run, then claim the released token. The retry loop covers the window where cancellation cleanup has not propagated yet:
245
+ **Supersede the owner.** Create the Hook with [`experimental_force: true`](/docs/api-reference/workflow/create-hook#take-over-a-token-another-run-holds) so the newest run takes the token over from the active owner. This also works for a finished run holding the token under `experimental_minRetention`:
246
246
 
247
247
  ```typescript lineNumbers
248
248
  import { createHook } from "workflow";
@@ -253,31 +253,54 @@ declare function chargeOrder(orderId: string): Promise<void>; // @setup
253
253
  export async function processOrderNewestWins(orderId: string) {
254
254
  "use workflow";
255
255
 
256
- const token = `order:${orderId}`;
257
-
258
- for (let attempt = 0; attempt < 3; attempt++) {
259
- using request = createHook<OrderRequest>({ token });
260
-
261
- const conflict = await request.getConflict();
262
- if (!conflict) {
263
- // Token claimed: this run is now the owner.
264
- const { confirmed } = await request;
265
- if (confirmed) {
266
- await chargeOrder(orderId);
267
- }
268
- return { status: "processed" as const };
269
- }
256
+ using request = createHook<OrderRequest>({
257
+ token: `order:${orderId}`,
258
+ experimental_force: true, // [!code highlight]
259
+ });
270
260
 
271
- await conflict.cancel(); // [!code highlight]
261
+ // This run now owns the token.
262
+ const { confirmed } = await request;
263
+ if (confirmed) {
264
+ await chargeOrder(orderId);
272
265
  }
266
+ return { status: "processed" as const };
267
+ }
268
+ ```
269
+
270
+ The previous owner's `await request` rejects with [`HookForceClaimedError`](/docs/api-reference/workflow-errors/hook-force-claimed-error), and every `resumeHook()` for the token reaches the new run from then on. Any run of this workflow can be superseded by a later one, so catch the error and exit cleanly:
271
+
272
+ ```typescript lineNumbers
273
+ import { createHook } from "workflow";
274
+ import { HookForceClaimedError } from "workflow/errors";
273
275
 
274
- throw new Error(`Could not claim ${token} after canceling the owner`);
276
+ type OrderRequest = { confirmed: boolean };
277
+ declare function chargeOrder(orderId: string): Promise<void>; // @setup
278
+
279
+ export async function processOrderNewestWins(orderId: string) {
280
+ "use workflow";
281
+
282
+ using request = createHook<OrderRequest>({
283
+ token: `order:${orderId}`,
284
+ experimental_force: true,
285
+ });
286
+
287
+ try {
288
+ const { confirmed } = await request;
289
+ if (confirmed) {
290
+ await chargeOrder(orderId);
291
+ }
292
+ return { status: "processed" as const };
293
+ } catch (error) {
294
+ if (HookForceClaimedError.is(error)) { // [!code highlight]
295
+ // A newer run for this order owns the token now. Stop here.
296
+ return { status: "superseded" as const, runId: error.claimedByRunId }; // [!code highlight]
297
+ }
298
+ throw error;
299
+ }
275
300
  }
276
301
  ```
277
302
 
278
- <Callout type="warn">
279
- This pattern does not work with `experimental_minRetention`: canceling the old run does not make its token available early.
280
- </Callout>
303
+ A run started at a Workflow spec version below 8, including runs started by older SDK releases, can't be taken from. In that case the forced Hook rejects with `HookConflictError`, as it would without `experimental_force`.
281
304
 
282
305
  If duplicate requests should only reuse the active run without sending data, use [`getHookByToken()`](/docs/api-reference/workflow-api/get-hook-by-token) as an advisory pre-check before calling `start()`. The workflow should still check `hook.getConflict()`, because the lookup and `start()` are not atomic.
283
306
 
@@ -102,6 +102,8 @@ export async function parentWorkflow(inputValue: number) {
102
102
 
103
103
  When you call `start()` inside a workflow function, it automatically executes through an internal step to maintain deterministic replay. The returned `Run` object works as it does outside workflows. Properties such as `.runId`, `.status`, and `.returnValue`, and methods such as `.cancel()`, are all available. Each property access or method call executes as a separate step.
104
104
 
105
+ If the child run fails or is canceled, `await childRun.returnValue` throws [`WorkflowRunFailedError`](/docs/api-reference/workflow-errors/workflow-run-failed-error) or [`WorkflowRunCancelledError`](/docs/api-reference/workflow-errors/workflow-run-cancelled-error) on the first attempt.
106
+
105
107
  <Callout type="info">
106
108
  Inside workflow functions, each `Run` property access (e.g., `run.status`, `run.returnValue`) triggers a workflow step. This means each access is recorded in the event log and replayed deterministically.
107
109
  </Callout>
@@ -188,7 +188,7 @@ Events are categorized by the entity type they affect. Each event contains metad
188
188
  | `hook_created` | Creates a new hook in `active` state. Contains the hook token and optional metadata. |
189
189
  | `hook_conflict` | Records that hook creation failed because another run owns the token. Contains the token and, for current worlds, the owner's run ID. The hook is not created: `hook.getConflict()` resolves with the conflicting run, and awaiting the hook payload rejects with a `HookConflictError`. |
190
190
  | `hook_received` | Records that a payload was delivered to the hook. The hook remains `active` and can receive more payloads. |
191
- | `hook_disposed` | Deletes the hook from storage (conceptually transitioning to `disposed` state). The token is released for reuse by future workflows. |
191
+ | `hook_disposed` | Deletes the hook from storage (conceptually transitioning to `disposed` state). The token is released for reuse by future workflows. When another run takes the token over with [`experimental_force`](/docs/api-reference/workflow/create-hook#take-over-a-token-another-run-holds), the event names the new owner in `forceClaimedBy`, and the token moves to that run instead of being released. |
192
192
 
193
193
  ### Wait events
194
194
 
@@ -203,6 +203,12 @@ Events are categorized by the entity type they affect. Each event contains metad
203
203
  |-------|-------------|
204
204
  | `noop` | Seals an abandoned log position (`specVersion` 7 and above). Only the backend writes this event; the create endpoints reject it. See [Sealed positions](#sealed-positions-noop-events). |
205
205
 
206
+ ### How fast a log grows
207
+
208
+ Counting the events above is how you size a run. A step that succeeds on the first attempt contributes three events; a retry adds another `step_started`, preceded by a `step_retrying` when the World emits one; a sleep contributes two, and a hook at least two.
209
+
210
+ A run's log is capped — the [Vercel World](/worlds/vercel#per-run-limits) fails anything past its ceiling with `MAX_EVENTS_EXCEEDED`. Replay reads the whole log, so a run gets slower as it grows and is worth splitting well before it fails: split into [child workflows](/cookbook/advanced/child-workflows) once a run would grow past a few thousand events.
211
+
206
212
  ## Terminal states
207
213
 
208
214
  Terminal states represent the end of an entity's lifecycle. Once an entity reaches a terminal state, no further events can transition it to another state.
@@ -39,21 +39,23 @@ export async function register() {
39
39
 
40
40
  Keep the dynamic `workflow/api` import inside the `NEXT_RUNTIME === "nodejs"` guard. Next.js also compiles `instrumentation.ts` for the Edge runtime. A top-level static import pulls Node.js-only dependencies into that compilation and breaks webpack Edge builds, even if the registration call is guarded.
41
41
 
42
- `registerLifecycleHooks` returns an unregister function. You can register multiple hook sets, and handlers run in registration order.
42
+ `registerLifecycleHooks` returns an unregister function. You can register multiple hook sets, and handlers run in registration order. Registrations are not deduplicated: register each hook set once per process, and call its unregister function before registering it again during hot reload or module re-evaluation. Otherwise, repeated registrations invoke the same handler multiple times for each transition.
43
43
 
44
44
  ## Handler parameters
45
45
 
46
46
  Both handlers receive a `workflowName` string and the [`Run`](/docs/api-reference/workflow-api/get-run) instance for the transitioned run. `workflowName` is the machine-readable workflow identifier, such as `workflow//./src/workflows/order//processOrder`. Use this parameter to filter runs without a backend read. `run.runId` is also available without a read.
47
47
 
48
- The `Run` instance hydrates lazily. Accessors such as `run.workflowName`, `run.status`, and `run.returnValue` still fetch from the backend when used. Lazy access defers those reads; it does not make them free.
48
+ The `Run` instance hydrates lazily. Accessors such as `run.workflowName`, `run.status`, and `run.returnValue` still fetch from the backend when used. In particular, `workflowName` is a string, while `run.workflowName` is a `Promise<string>`. Lazy access defers those reads; it does not make them free.
49
49
 
50
- `onRunFailed` additionally receives the failure as a `WorkflowRunFailedError`, the same shape `run.returnValue` rejects with:
50
+ `onRunFailed` additionally receives a `WorkflowRunFailedError` hydrated for reporting. Unlike `run.returnValue`, it defers readable stream I/O and uses persisted abort snapshots:
51
51
 
52
52
  - `error.errorCode`: the failure classification (`USER_ERROR`, `RUNTIME_ERROR`, `MAX_DELIVERIES_EXCEEDED`, and more). See [error codes](/docs/errors) for the full list.
53
- - `error.cause`: the thrown value hydrated from the persisted error data, with registered Error subclass identity, message, stack, and cause chain preserved. Streamed values load lazily when consumed, and abort signals reflect their persisted state without live subscriptions. If hydration fails, the cause is a generic `Error`, matching `run.returnValue`'s fallback. Any JavaScript value can be thrown, so this is typed `unknown`.
53
+ - `error.cause`: the thrown value hydrated from the persisted error data, with registered Error subclass identity, message, stack, and cause chain preserved. Readable streams load lazily when consumed, and abort signals reflect their persisted state without live subscriptions. Writable streams retain their normal forwarding pipe and lock-polling setup during hydration. If hydration fails, the cause is a generic `Error`, matching `run.returnValue`'s fallback. Any JavaScript value can be thrown, so this is typed `unknown`.
54
54
 
55
55
  In `onRunFailed`, `run.returnValue` rejects because the run failed. Use `error.cause` to inspect or report the thrown value instead of awaiting `run.returnValue`.
56
56
 
57
+ The invocation's `waitUntil` scope includes background stream operations from the hydrated cause, even after a handler returns or throws. Close or release stream reader and writer locks when finished so that work can settle. Await other asynchronous reporting work in your handler, such as `Sentry.flush()` below, to keep it in the same lifetime scope.
58
+
57
59
  ## Reporting failed runs to Sentry
58
60
 
59
61
  This example reports failures for workflows named `processOrder`. Remove the filter to report failures from all workflows.