workflow 5.0.0-beta.10 → 5.0.0-beta.12
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
|
@@ -8,12 +8,11 @@ type: overview
|
|
|
8
8
|
|
|
9
9
|
**Date**: March 2026
|
|
10
10
|
|
|
11
|
-
This is a major internal architecture change to how Workflow DevKit executes workflows and steps
|
|
11
|
+
This is a major internal architecture change to how Workflow DevKit executes workflows and steps. It reduces function invocations and queue overhead by executing steps _inline_ within the same function invocation as the workflow replay, rather than dispatching every step to a separate function via the queue.
|
|
12
12
|
|
|
13
13
|
## Previous Architecture
|
|
14
14
|
|
|
15
|
-
The previous architecture used two separate routes,
|
|
16
|
-
each backed by its own queue trigger:
|
|
15
|
+
The previous architecture used two separate routes, each backed by its own queue trigger:
|
|
17
16
|
|
|
18
17
|
```
|
|
19
18
|
Queue: __wkf_workflow_* --> /.well-known/workflow/v1/flow (workflow replay in VM)
|
|
@@ -69,118 +68,64 @@ suspension with pending operations
|
|
|
69
68
|
|
|
70
69
|
A serial workflow with 10 steps now completes in **1 function invocation**.
|
|
71
70
|
|
|
72
|
-
## Inline Step Execution
|
|
73
|
-
|
|
74
|
-
After the workflow suspends with pending steps, the handler executes one step inline:
|
|
75
|
-
|
|
76
|
-
1. Create `step_started` event
|
|
77
|
-
2. Hydrate step input from the event log
|
|
78
|
-
3. Look up the step function via `getStepFunction(stepName)`
|
|
79
|
-
4. Execute the step function
|
|
80
|
-
5. Create `step_completed` or `step_failed` event
|
|
81
|
-
6. Loop back to workflow replay
|
|
82
|
-
|
|
83
|
-
This logic lives in `executeStep()` in `packages/core/src/runtime/step-executor.ts`.
|
|
84
|
-
|
|
85
71
|
## Background Steps (Parallel Execution)
|
|
86
72
|
|
|
87
|
-
When a workflow suspends with multiple pending steps (e.g., from `Promise.all`), the handler
|
|
88
|
-
|
|
89
|
-
1. Creates `step_created` events for all pending steps
|
|
90
|
-
2. Queues N-1 steps back to `__wkf_workflow_*` with `stepId` and `stepName` in the message payload
|
|
91
|
-
3. Executes 1 step inline
|
|
92
|
-
4. Loops back to replay
|
|
73
|
+
When a workflow suspends with multiple pending steps (e.g., from `Promise.all`), the handler creates `step_created` events for all of them, queues N-1 back to `__wkf_workflow_*` with `stepId` and `stepName` in the message payload, executes 1 step inline, and loops back to replay.
|
|
93
74
|
|
|
94
|
-
Each background step message is handled by a separate function invocation of the same handler. When a message arrives with `stepId` and `stepName`, the handler executes that specific step, then checks if all parallel steps from the batch are done by
|
|
75
|
+
Each background step message is handled by a separate function invocation of the same handler. When a message arrives with `stepId` and `stepName`, the handler executes that specific step, then checks if all parallel steps from the batch are done by comparing `step_created` events against terminal events (`step_completed`/`step_failed`):
|
|
95
76
|
|
|
96
|
-
- **All steps done**: The handler replays the workflow inline, continuing the execution loop without a queue roundtrip.
|
|
77
|
+
- **All steps done**: The handler replays the workflow inline, continuing the execution loop without a queue roundtrip.
|
|
97
78
|
- **Steps still pending**: The handler returns without queuing a continuation. The last handler to complete its step will see all steps done and replay inline.
|
|
98
|
-
- **Pending ops (stream writes)**: The handler queues a continuation and returns, so `waitUntil` can flush the pending stream data
|
|
79
|
+
- **Pending ops (stream writes)**: The handler queues a continuation and returns, so `waitUntil` can flush the pending stream data.
|
|
99
80
|
|
|
100
81
|
### Convergence After Parallel Steps
|
|
101
82
|
|
|
102
|
-
When multiple background steps complete near-simultaneously, multiple handlers may observe "all steps done" and attempt to advance the workflow concurrently. The event-sourced architecture plus the invariants
|
|
83
|
+
When multiple background steps complete near-simultaneously, multiple handlers may observe "all steps done" and attempt to advance the workflow concurrently. The event-sourced architecture plus the invariants below ensure safe convergence:
|
|
103
84
|
|
|
104
|
-
- **`step_created` idempotency**
|
|
105
|
-
- **`step_completed` / `step_failed` idempotency**
|
|
106
|
-
- **Queue idempotency keys**
|
|
107
|
-
- **Deterministic replay**
|
|
85
|
+
- **`step_created` idempotency** — duplicate creates return 409; exactly one handler owns each step
|
|
86
|
+
- **`step_completed` / `step_failed` idempotency** — only the first invocation to record a terminal result wins
|
|
87
|
+
- **Queue idempotency keys** — background step messages use `correlationId` as idempotency key
|
|
88
|
+
- **Deterministic replay** — all invocations produce the same result given the same event log
|
|
108
89
|
|
|
109
90
|
### Single Inline Executor Per Step
|
|
110
91
|
|
|
111
|
-
Inline step execution combined with background-step dispatch introduces a new coordination requirement
|
|
92
|
+
Inline step execution combined with background-step dispatch introduces a new coordination requirement: when multiple handlers reach the same `Promise.all` batch concurrently, we need to guarantee that each step body runs at most once via the inline path. Without that guarantee, the event log accumulates duplicate `step_started` events (including some written *after* `step_completed`, which orphans them on replay) and step bodies run redundantly.
|
|
112
93
|
|
|
113
94
|
The design enforces a simple invariant: **exactly one handler owns each step, and only the owner may execute it inline**. Ownership is established by the atomicity of `step_created`:
|
|
114
95
|
|
|
115
|
-
1. **Atomic `step_created`**
|
|
116
|
-
2. **Suspension handler reports ownership**
|
|
117
|
-
3. **Inline execution is gated on ownership**
|
|
118
|
-
4. **Queueing is unconditional**
|
|
96
|
+
1. **Atomic `step_created`** — `events.create('step_created', correlationId=X)` is serialized per-correlationId in every world. Exactly one concurrent caller succeeds; the rest receive `EntityConflictError`.
|
|
97
|
+
2. **Suspension handler reports ownership** — only `step_created` writes that actually succeeded (not those that caught 409) count toward ownership.
|
|
98
|
+
3. **Inline execution is gated on ownership** — a handler that didn't win any `step_created` race performs no inline execution.
|
|
99
|
+
4. **Queueing is unconditional** — for every pending step except the one being inline-executed, the handler enqueues a background step message with `idempotencyKey: correlationId`. This is what makes crash recovery work: if a prior handler wrote `step_created` but crashed before enqueueing, a later handler will enqueue the orphaned step. Concurrent handlers' redundant enqueues dedupe on the idempotency key.
|
|
119
100
|
|
|
120
|
-
Together these give: every `step_created` event has exactly one inline executor
|
|
101
|
+
Together these give: every `step_created` event has exactly one inline executor **and** at least one queued dispatch. Step bodies are never executed concurrently, and `step_started` events never land in the log after `step_completed` for the same step.
|
|
121
102
|
|
|
122
|
-
**Retry semantics are preserved**:
|
|
103
|
+
**Retry semantics are preserved**: a step that is currently running (status=`running`) still accepts a second `step_started` write with an incremented attempt counter — this is how queue redelivery after a SIGKILL mid-execution legitimately re-runs the step.
|
|
123
104
|
|
|
124
105
|
## Incremental Event Loading
|
|
125
106
|
|
|
126
107
|
The handler caches the event log in memory across loop iterations. Instead of re-fetching the entire event log on each replay:
|
|
127
108
|
|
|
128
|
-
1. **First iteration**: full load
|
|
129
|
-
2. **Subsequent iterations**:
|
|
109
|
+
1. **First iteration**: full load, returning both the events and the final pagination cursor
|
|
110
|
+
2. **Subsequent iterations**: fetch only events created after the saved cursor and append them to the cached array
|
|
130
111
|
|
|
131
112
|
For a 10-step serial workflow completing in one invocation, the 10th replay loads ~2 new events instead of re-fetching all ~30.
|
|
132
113
|
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
The incremental loading depends on the server returning a cursor even on the final page of results (`hasMore: false`). Previously, `workflow-server` returned `cursor: null` when there were no more pages. This was fixed in the `peter/fix-end-cursor` branch to always return an `eid:<eventId>` cursor when there are events, aligning with `world-local` and `world-postgres` behavior.
|
|
136
|
-
|
|
137
|
-
If a World implementation does not return a cursor after the initial load, the handler logs an error and falls back to a full reload.
|
|
114
|
+
Incremental loading depends on the World returning a cursor even on the final page of results. If a World implementation does not return a cursor after the initial load, the handler logs an error and falls back to a full reload.
|
|
138
115
|
|
|
139
116
|
## Timeout Handling
|
|
140
117
|
|
|
141
|
-
The inline execution loop checks wall-clock time before each replay iteration. If the elapsed time exceeds a configurable threshold (default: 110 seconds, for a 120-second function limit), the handler re-schedules itself via the queue and returns.
|
|
142
|
-
|
|
143
|
-
The threshold is configurable via the `WORKFLOW_V2_TIMEOUT_MS` environment variable.
|
|
118
|
+
The inline execution loop checks wall-clock time before each replay iteration. If the elapsed time exceeds a configurable threshold (default: 110 seconds, for a 120-second function limit), the handler re-schedules itself via the queue and returns. Configurable via `WORKFLOW_V2_TIMEOUT_MS`.
|
|
144
119
|
|
|
145
120
|
If a single step takes longer than the timeout threshold, the step runs to completion (or SIGKILL) — there is no interruption mechanism for in-progress step execution. This is the same behavior as the previous architecture.
|
|
146
121
|
|
|
147
122
|
## Queue Message Changes
|
|
148
123
|
|
|
149
|
-
The `WorkflowInvokePayload` schema has two new optional fields:
|
|
150
|
-
|
|
151
|
-
{/*@skip-typecheck - snippet, not runnable code*/}
|
|
152
|
-
|
|
153
|
-
```typescript
|
|
154
|
-
stepId: z.string().optional()
|
|
155
|
-
stepName: z.string().optional()
|
|
156
|
-
```
|
|
157
|
-
|
|
158
|
-
When `stepId` is present, the handler executes that specific step before (or instead of) replaying the workflow. Background steps are queued with both `stepId` and `stepName` set, so the handler knows which step function to call without loading the event log. Previously, `stepName` was resolved by loading all events and searching for the `step_created` event matching the `stepId` — an O(N) operation on the full event history for every background step arrival.
|
|
124
|
+
The `WorkflowInvokePayload` schema has two new optional fields: `stepId` and `stepName`. When `stepId` is present, the handler executes that specific step before (or instead of) replaying the workflow. Background steps are queued with both set, so the handler knows which step function to call without loading the event log. Previously, `stepName` was resolved by loading all events and searching for the `step_created` event matching the `stepId` — an O(N) operation on the full event history for every background step arrival.
|
|
159
125
|
|
|
160
126
|
The queue trigger configuration uses `WORKFLOW_QUEUE_TRIGGER` on the `__wkf_workflow_*` topic. The `__wkf_step_*` topic and its separate trigger are no longer generated.
|
|
161
127
|
|
|
162
|
-
##
|
|
163
|
-
|
|
164
|
-
### Base Builder
|
|
165
|
-
|
|
166
|
-
New method `createCombinedBundle()` in `packages/builders/src/base-builder.ts`:
|
|
167
|
-
|
|
168
|
-
1. Builds the step registrations bundle (same esbuild + SWC step mode as before)
|
|
169
|
-
2. Builds the workflow VM code string (same esbuild + SWC workflow mode as before)
|
|
170
|
-
3. Generates a combined route file that imports the step registrations and uses `workflowEntrypoint(workflowCode)`
|
|
171
|
-
|
|
172
|
-
No changes to the SWC plugin were needed. The two-pass build approach (separate step and workflow SWC modes) still applies.
|
|
173
|
-
|
|
174
|
-
### Framework Builders
|
|
175
|
-
|
|
176
|
-
All framework builders were updated to use `createCombinedBundle()`:
|
|
177
|
-
|
|
178
|
-
- **Next.js** (eager and deferred/lazyDiscovery): replaces separate step + flow route generation
|
|
179
|
-
- **NestJS, Nitro, Standalone**: replaces separate `createStepsBundle()` + `createWorkflowsBundle()` calls
|
|
180
|
-
- **SvelteKit, Astro**: same, plus post-processing regex updated to match `workflowEntrypoint`
|
|
181
|
-
- **Vercel Build Output API** (used by Nitro/Astro production): single `flow.func/` with `WORKFLOW_QUEUE_TRIGGER`
|
|
182
|
-
|
|
183
|
-
### Generated File Layout
|
|
128
|
+
## Generated File Layout
|
|
184
129
|
|
|
185
130
|
```
|
|
186
131
|
.well-known/workflow/v1/
|
|
@@ -196,37 +141,19 @@ All framework builders were updated to use `createCombinedBundle()`:
|
|
|
196
141
|
|
|
197
142
|
The `step/` directory is no longer generated.
|
|
198
143
|
|
|
199
|
-
##
|
|
200
|
-
|
|
201
|
-
`handleSuspension()` in `packages/core/src/runtime/suspension-handler.ts` creates events for all pending operations (hooks, step events, wait events) but does **not** queue step messages. It returns the pending step items so the handler can decide which to execute inline vs. queue to background.
|
|
202
|
-
|
|
203
|
-
## Concerns and Edge Cases
|
|
144
|
+
## Design Notes and Tradeoffs
|
|
204
145
|
|
|
205
146
|
### Parent→Child Polling Holds Worker Slots
|
|
206
147
|
|
|
207
148
|
`Run#returnValue` is implemented as a polling step: the workflow awaits the child run's terminal status inside a step body. In worker-based worlds (notably `world-postgres`), each such poll occupies a queue worker slot until the child run finishes. Parent workflows that fan out to many child runs — recursive workflows like `fibonacciWorkflow` are the obvious case — can therefore consume a large fraction of available workers just holding positions in `Promise.all([...children.map(c => c.returnValue)])`.
|
|
208
149
|
|
|
209
|
-
If `queueConcurrency` is smaller than the peak number of concurrent parent polls plus the workers needed for any in-flight children, the system deadlocks: every slot is held by a parent waiting on a child, but no child can acquire a slot to start.
|
|
150
|
+
If `queueConcurrency` is smaller than the peak number of concurrent parent polls plus the workers needed for any in-flight children, the system deadlocks: every slot is held by a parent waiting on a child, but no child can acquire a slot to start.
|
|
210
151
|
|
|
211
|
-
For `world-postgres`, the default `queueConcurrency` is set to **50
|
|
152
|
+
For `world-postgres`, the default `queueConcurrency` is set to **50**. Workflows that fan out more aggressively must raise this ceiling.
|
|
212
153
|
|
|
213
|
-
|
|
154
|
+
To prevent deadlock when polling is executed inline by the step executor, `Run#pollReturnValue()` detects when it's running inside a step executor and throws `TooEarlyError` instead of polling in a blocking loop. The step executor handles `TooEarlyError` by re-queueing the step with a 1-second delay, freeing the worker. Unlike `RetryableError`, `TooEarlyError` does NOT count against `maxRetries`, so polling steps can retry indefinitely until the child completes.
|
|
214
155
|
|
|
215
|
-
|
|
216
|
-
|
|
217
|
-
Workflow code still runs in a Node.js VM for determinism and sandboxing. Step code runs in the Node.js host context. The only change is that both happen within the same function invocation.
|
|
218
|
-
|
|
219
|
-
### Bundle Size and Cold Start
|
|
220
|
-
|
|
221
|
-
The combined bundle is larger (contains both step code and workflow VM code). Cold start time increases slightly. The reduction in total function invocations more than compensates.
|
|
222
|
-
|
|
223
|
-
### Step Retries
|
|
224
|
-
|
|
225
|
-
When an inline step fails with retries remaining:
|
|
226
|
-
|
|
227
|
-
- `RetryableError` with explicit `retryAfter` delay: re-queue to self with `stepId` and delay
|
|
228
|
-
- Transient errors with immediate retry: re-queue to self with `stepId` (delay = 1s)
|
|
229
|
-
- `FatalError`: fail immediately
|
|
156
|
+
**Follow-up**: Replace the worker-pool sizing requirement with a polling design that does not occupy a worker slot. Options under consideration: moving child-completion polling out of the step body into the suspension layer, or emitting a `run_completed` notification on the parent's stream/queue so the parent only resumes when the child actually finishes.
|
|
230
157
|
|
|
231
158
|
### Mixed Suspensions
|
|
232
159
|
|
|
@@ -236,360 +163,87 @@ A suspension may contain steps, hooks, and waits simultaneously. The handler cre
|
|
|
236
163
|
- **Steps + at least one wait**: every step is queued (no inline execution). The handler returns with the wait timeout. Whichever lands first — a step's continuation or the wait timer — drives the next replay.
|
|
237
164
|
- **Hooks / waits only**: handler returns with the wait timeout (or no timeout, for hook-only suspensions). The next continuation is driven by external resume or the wait timer.
|
|
238
165
|
|
|
239
|
-
The "no inline when there's a wait" carve-out is necessary to preserve `Promise.race(step, sleep)` semantics. Inline `await executeStep(...)` blocks the handler for the full step duration, and `wait_completed` events are only created on the *next* loop iteration's "complete elapsed waits" pass — so a longer-running step would always swallow the shorter sleep and `Promise.race` would resolve incorrectly. Queueing the step in this case lets the wait timer drive a continuation in parallel
|
|
166
|
+
The "no inline when there's a wait" carve-out is necessary to preserve `Promise.race(step, sleep)` semantics. Inline `await executeStep(...)` blocks the handler for the full step duration, and `wait_completed` events are only created on the *next* loop iteration's "complete elapsed waits" pass — so a longer-running step would always swallow the shorter sleep and `Promise.race` would resolve incorrectly. Queueing the step in this case lets the wait timer drive a continuation in parallel.
|
|
240
167
|
|
|
241
168
|
Pure step suspensions (without waits) still benefit from inline execution; the carve-out only costs an extra queue roundtrip when a step and a sleep coexist.
|
|
242
169
|
|
|
243
|
-
###
|
|
244
|
-
|
|
245
|
-
If a hook conflict is detected during suspension handling, the handler breaks the loop and returns `{ timeoutSeconds: 0 }` for immediate re-invocation, same as the previous behavior.
|
|
246
|
-
|
|
247
|
-
### Encryption Key Resolution
|
|
248
|
-
|
|
249
|
-
Encryption keys are resolved once before the inline execution loop starts (after the run status is confirmed as `running`) and reused across all iterations. Background step executions resolve the key independently. The key does not change within a run.
|
|
250
|
-
|
|
251
|
-
## Framework Support
|
|
252
|
-
|
|
253
|
-
All framework integrations have been updated: Next.js (eager and deferred/lazyDiscovery), NestJS, SvelteKit, Astro, Nitro/Nuxt/Hono/Express/Vite, and CLI standalone. The Vercel Build Output API builder (used by Nitro and Astro for production deploys) also uses the combined bundle with `WORKFLOW_QUEUE_TRIGGER`.
|
|
254
|
-
|
|
255
|
-
## Non-Next.js Integration Challenges
|
|
256
|
-
|
|
257
|
-
### Module Scope Duplication in Re-Bundled Output
|
|
258
|
-
|
|
259
|
-
Builders that use `bundleFinalOutput: true` (standalone CLI, Vercel Build Output API, NestJS) produce a single file where esbuild re-bundles the step registrations and the workflow runtime together. esbuild creates isolated module scopes for each source module, even within the same output file. This meant `registerStepFunction` and `getStepFunction` operated on different `Map` instances — steps were registered into one Map but looked up from another.
|
|
260
|
-
|
|
261
|
-
**Fix**: The step function registry (`registeredSteps` Map in `@workflow/core/private`) and the step context storage (`contextStorage` AsyncLocalStorage in `@workflow/core/step/context-storage`) were changed from module-scoped variables to `globalThis` singletons using `Symbol.for`. This ensures all esbuild module scopes share the same instances. The pattern was already used in the codebase for the World singleton and the class serialization registry.
|
|
262
|
-
|
|
263
|
-
### Workflow Package CJS Export Condition
|
|
264
|
-
|
|
265
|
-
The `workflow` package's root export has `"require": "./dist/typescript-plugin.cjs"` for TypeScript editor plugin loading. When esbuild bundles with CJS format, it resolves `import { defineHook } from 'workflow'` via the `require` condition, getting the TS plugin instead of the API.
|
|
266
|
-
|
|
267
|
-
**Fix**: Added a `"node"` condition (`"node": "./dist/index.js"`) before the `"require"` condition in the workflow package's exports. esbuild with `conditions: ['node']` matches `"node"` first and uses the correct API entry. TypeScript's plugin loader doesn't use `conditions: ['node']`, so it still falls through to `"require"` for the TS plugin.
|
|
268
|
-
|
|
269
|
-
### Local World Concurrent Replay Interference
|
|
270
|
-
|
|
271
|
-
The local development world (`world-local`) processes queue messages with high concurrency (default: 1000). With the V2 combined handler, parallel steps generate multiple workflow continuation messages. When these are processed concurrently, each triggers a replay that sees in-flight events from other concurrent replays. This causes "unconsumed event" errors because the event consumer encounters events that don't match any subscriber in the current replay state.
|
|
272
|
-
|
|
273
|
-
In production (Vercel), this doesn't happen — each function invocation is isolated with its own event loading.
|
|
274
|
-
|
|
275
|
-
**Fix**: The `EventsConsumer`'s `onUnconsumedEvent` callback (see "Concurrent Replay Interference with Multi-Batch Workflows" below) handles the concurrent event visibility issue. The V2 inline replay optimization (where the last background step to complete replays inline instead of queuing) further reduces concurrent replays. Redundant step executions from concurrent handlers are harmless due to `step_completed` idempotency — only the first completion wins.
|
|
276
|
-
|
|
277
|
-
### ESM `bundleFinalOutput` and Dynamic Require Errors
|
|
278
|
-
|
|
279
|
-
When `bundleFinalOutput: true` is used with ESM format, esbuild bundles CJS dependencies (like `debug`) into the output. CJS `require()` calls are wrapped in esbuild's `__require` polyfill, which throws "Dynamic require of X is not supported" in ESM contexts where `require` is undefined. This affected all ESM-based framework builders (Nitro, NestJS, SvelteKit, Astro) that were switched to `bundleFinalOutput: true` during the V2 migration.
|
|
280
|
-
|
|
281
|
-
**Fix**: ESM builders use `bundleFinalOutput: false` with `externalizeNonSteps: true`, matching the pre-V2 behavior. The framework's own bundler (Vite, Rollup, Turbopack) handles dependency resolution. The standalone CLI and Vercel Build Output API builders use `bundleFinalOutput: true` with ESM output plus a `createRequire(import.meta.url)` banner (see "V2 Combined Bundle Switched from CJS to ESM" below) so CJS dependencies can still call `require()` for Node.js builtins.
|
|
282
|
-
|
|
283
|
-
### Rollup Tree-Shaking of Step Registrations
|
|
284
|
-
|
|
285
|
-
When `bundleFinalOutput: false` is used with Nitro's rollup pipeline, the step registrations bundle (`steps.mjs`) only contains side-effect code (`registerStepFunction` calls) with no exports. Rollup tree-shakes the entire module because it has no used exports, removing all step registrations from the production bundle. This causes "Step not found" errors at runtime.
|
|
286
|
-
|
|
287
|
-
**Fix**: The steps bundle now exports a sentinel value (`export const __steps_registered = true`), and the combined route file imports it (`import { __steps_registered } from './steps.mjs'`). This gives rollup a used binding to track, preventing it from dropping the module and its side effects.
|
|
288
|
-
|
|
289
|
-
### Concurrent Replay Interference with Multi-Batch Workflows (historic)
|
|
290
|
-
|
|
291
|
-
An earlier iteration of the V2 work hit "Unconsumed event in event log" errors when multiple concurrent handlers raced into the same batch boundary. The diagnosis at the time was that concurrent handlers could see events the current replay hadn't reached yet, and the mitigation was a skip path in `onUnconsumedEvent` that tolerated step/hook/wait lifecycle events whose correlationId had a matching `step_created` / `hook_created` / earlier `wait_completed` in the log.
|
|
292
|
-
|
|
293
|
-
Later work on the "Single Inline Executor Per Step" invariant (described above) identified the actual root cause: duplicate `step_started` events were being written *after* `step_completed` on the same step, because the local world's `step_started` was not atomic w.r.t. terminal state and the main loop was re-picking already-queued steps for inline execution. Fixing those at the source (per-step mutex in `world-local` + ownership-gated inline dispatch + unconditional queueing with idempotency keys) eliminated the unconsumed-step-event path entirely, and fixing the `wait_completed` cursor bug (the main loop manually pushed `wait_completed` events without advancing `eventsCursor`, so the next incremental fetch re-returned them as local-array duplicates) eliminated the wait case.
|
|
170
|
+
### VM Sandboxing
|
|
294
171
|
|
|
295
|
-
|
|
172
|
+
Workflow code still runs in a Node.js VM for determinism and sandboxing. Step code runs in the Node.js host context. The only change is that both happen within the same function invocation.
|
|
296
173
|
|
|
297
|
-
###
|
|
174
|
+
### Bundle Size and Cold Start
|
|
298
175
|
|
|
299
|
-
|
|
176
|
+
The combined bundle is larger (contains both step code and workflow VM code). Cold start time increases slightly. The reduction in total function invocations more than compensates.
|
|
300
177
|
|
|
301
|
-
|
|
178
|
+
### Step Retries
|
|
302
179
|
|
|
303
|
-
|
|
180
|
+
When an inline step fails with retries remaining:
|
|
304
181
|
|
|
305
|
-
|
|
182
|
+
- `RetryableError` with explicit `retryAfter` delay: re-queue to self with `stepId` and delay
|
|
183
|
+
- Transient errors with immediate retry: re-queue to self with `stepId` (delay = 1s)
|
|
184
|
+
- `FatalError`: fail immediately
|
|
306
185
|
|
|
307
|
-
|
|
186
|
+
### Encryption Key Resolution
|
|
308
187
|
|
|
309
|
-
|
|
188
|
+
Encryption keys are resolved once before the inline execution loop starts (after the run status is confirmed as `running`) and reused across all iterations. Background step executions resolve the key independently. The key does not change within a run.
|
|
310
189
|
|
|
311
|
-
|
|
190
|
+
### Module Scope Duplication in Re-Bundled Output
|
|
312
191
|
|
|
313
|
-
|
|
192
|
+
Builders that re-bundle the combined output into a single file (standalone CLI, Vercel Build Output API, NestJS) produce a layout where esbuild creates isolated module scopes for each source module, even within the same output file. Without intervention this means `registerStepFunction` and `getStepFunction` operate on different `Map` instances — steps are registered into one Map but looked up from another.
|
|
314
193
|
|
|
315
|
-
|
|
194
|
+
The step function registry and the step context storage are `globalThis` singletons (via `Symbol.for`) to ensure all module scopes share the same instances. The same pattern is used for the World singleton and the class serialization registry.
|
|
316
195
|
|
|
317
196
|
### Inline Step Execution with Pending Stream Operations
|
|
318
197
|
|
|
319
198
|
When a step's arguments or return value include serialized streams (e.g., `WritableStream` from `getWritable()`, or AI SDK streaming steps), the serialization layer creates background `flushablePipe` operations that pipe data to S3. These ops are tracked in an `ops` array and need to complete before the stream data is readable by external consumers.
|
|
320
199
|
|
|
321
|
-
In V1, each step ran in a separate function invocation. After the step completed, `waitUntil(ops)` kept the function alive to flush the ops.
|
|
322
|
-
|
|
323
|
-
**Current state**: `executeStep()` attempts a 500ms `Promise.race` between the ops settling and a timeout. If ops settle in time (data confirmed on server), it returns `hasPendingOps: false` and the V2 handler continues the inline loop. If ops don't settle in 500ms (e.g., `WritableStream` kept open across steps), it returns `hasPendingOps: true` and the V2 handler breaks the loop and queues a continuation so `waitUntil` can flush them.
|
|
324
|
-
|
|
325
|
-
**Earlier attempts that failed** (before the flush waiter fix below):
|
|
326
|
-
|
|
327
|
-
1. **500ms inline ops await without flush waiters** — The same 500ms race, but `WorkflowServerWritableStream` used a buffered 10ms flush timer: the `flushablePipe`'s `pendingOps` reached 0 when the buffered `write()` returned (instant), but the actual S3 HTTP write hadn't started yet. The ops appeared settled but data wasn't on S3. Multiple approaches to fix the timing (delaying `pollWritableLock`, closing the writable to trigger flush, adding a post-settle delay) all failed or caused other issues (deadlocks, premature stream closure).
|
|
328
|
-
|
|
329
|
-
2. **Root cause of the buffered write issue**: `WorkflowServerWritableStream.write()` buffers chunks and schedules a flush via `setTimeout(flush, 10ms)`. The `flushablePipe` calls `await writer.write(chunk)` which returns immediately (data buffered). `pendingOps--` fires before the 10ms timer. The `pollWritableLock` sees `pendingOps === 0` and resolves `state.promise`. The ops appear settled, but data is still in the buffer.
|
|
330
|
-
|
|
331
|
-
3. **Why this only affects Vercel Prod**: On local (world-local), stream writes go to the filesystem — effectively instant. On Vercel (world-vercel), writes go through HTTP to workflow-server → S3, adding 50-100ms latency. The buffered write returns instantly but the HTTP round-trip is deferred. When the V2 loop continues and the function eventually returns, `waitUntil` may not have enough time to flush.
|
|
332
|
-
|
|
333
|
-
**Follow-up**: The flush-waiter design described under "Buffered Stream Flush with Waiter Promises" below is the landed fix and resolves the buffered-write race. The remaining work is to shrink the 500ms inline-ops budget once we have confidence that the flush-waiter path settles deterministically across all worlds (today the budget is a defensive ceiling, not a tuned latency target), and to surface a stronger contract for "ops settled" — currently a 500ms timeout means "probably settled, give up and queue a continuation", which is correct but coarse. A signaled "ops drained" event from the world layer would let `executeStep()` proceed without the timeout in the common case, removing latency for streaming workflows whose ops settle in well under 500ms.
|
|
334
|
-
|
|
335
|
-
### CJS `module.exports` Collision in BOA Bundles (RESOLVED)
|
|
336
|
-
|
|
337
|
-
The Vercel Build Output API (BOA) builder creates a single CJS bundle via `createCombinedBundle` with `bundleFinalOutput: true`. The combined route file imports the steps bundle:
|
|
338
|
-
|
|
339
|
-
```js
|
|
340
|
-
import { __steps_registered } from './__step_registrations.js';
|
|
341
|
-
import { workflowEntrypoint } from 'workflow/runtime';
|
|
342
|
-
export const POST = workflowEntrypoint(workflowCode);
|
|
343
|
-
```
|
|
344
|
-
|
|
345
|
-
When esbuild re-bundles this into CJS, the steps bundle's code is inlined. If the steps bundle is also CJS format, it contains its own `module.exports = __toCommonJS(...)` at the top level. esbuild sometimes inlines CJS modules **without** a `__commonJS()` wrapper (the heuristic depends on the module's detected format). When unwrapped, the steps bundle's `module.exports` assignment executes at the top level and **overwrites** the combined route's `module.exports`, removing the `POST` handler export.
|
|
346
|
-
|
|
347
|
-
**Symptoms**: The Vercel deployment builds and starts successfully, but the `POST` handler is missing from the function's exports. Queue messages are delivered to the function but nothing processes them. All e2e tests hang indefinitely.
|
|
200
|
+
In V1, each step ran in a separate function invocation. After the step completed, `waitUntil(ops)` kept the function alive to flush the ops. In V2, the inline execution loop continues immediately after the step body returns — so we need to know whether to keep looping or break out and let `waitUntil` flush.
|
|
348
201
|
|
|
349
|
-
|
|
202
|
+
`executeStep()` attempts a 500ms `Promise.race` between the ops settling and a timeout. If ops settle in time (data confirmed on server), it returns `hasPendingOps: false` and the V2 handler continues the inline loop. If ops don't settle in 500ms (e.g., `WritableStream` kept open across steps), it returns `hasPendingOps: true` and the V2 handler breaks the loop and queues a continuation so `waitUntil` can flush them.
|
|
350
203
|
|
|
351
|
-
|
|
352
|
-
2. Found two `module.exports` assignments in the bundle: line ~45K (from the combined route, exporting `POST`) and line ~95K (from the inlined steps bundle, exporting `__steps_registered`). The second overwrites the first.
|
|
353
|
-
3. Compared with the standalone builder's bundle which had the same steps code wrapped in `__commonJS()` — esbuild's wrapper prevents the inner `module.exports` from leaking.
|
|
354
|
-
|
|
355
|
-
**Fix**: When `bundleFinalOutput` is true, build the steps bundle in **ESM format** regardless of the final output format. The final esbuild pass converts everything to CJS correctly. ESM steps don't have `module.exports`, so there's no collision. The combined route's `export const POST` becomes the sole `module.exports` entry.
|
|
356
|
-
|
|
357
|
-
### Step Error Source Maps on BOA Deployments
|
|
358
|
-
|
|
359
|
-
The V2 combined CJS bundle (`bundleFinalOutput: true`) loses original source file names during re-bundling. Error stack traces show `/var/task/index.js` instead of `99_e2e.ts`. The `hasStepSourceMaps()` utility was updated to return `false` for BOA-builder frameworks (Express, Fastify, Hono, Nitro, Nuxt, Vite, Astro, Example) on Vercel preview, aligning test expectations with the actual bundle behavior.
|
|
360
|
-
|
|
361
|
-
### CLI Health Check Port Mismatch
|
|
362
|
-
|
|
363
|
-
The CLI `health` command defaults to `http://localhost:3000` when `WORKFLOW_LOCAL_BASE_URL` is not set. Different frameworks use different ports (Astro: 4321, SvelteKit: 5173). The e2e test passed `WORKFLOW_LOCAL_BASE_URL` via the spawn env, but the CLI's `getEnvVars()` function had a fixed list of env vars that didn't include `WORKFLOW_LOCAL_BASE_URL`. The env var was set but never read.
|
|
364
|
-
|
|
365
|
-
**Fix**: Added `WORKFLOW_LOCAL_BASE_URL` to the CLI's `getEnvVars()` return object.
|
|
204
|
+
**Follow-up**: Shrink the 500ms inline-ops budget once we have confidence that the flush-waiter path settles deterministically across all worlds. A signaled "ops drained" event from the world layer would let `executeStep()` proceed without the timeout in the common case.
|
|
366
205
|
|
|
367
206
|
### Buffered Stream Flush with Waiter Promises
|
|
368
207
|
|
|
369
|
-
`WorkflowServerWritableStream` buffers writes and flushes via a 10ms `setTimeout` for batching.
|
|
208
|
+
`WorkflowServerWritableStream` buffers writes and flushes via a 10ms `setTimeout` for batching. Naively, `write()` could return immediately after buffering, but that would cause the `flushablePipe`'s `pendingOps` counter to reach 0 before data actually reached the server — the V2 inline loop would see ops as settled prematurely and produce data-loss races on every step with `WritableStream` serialization.
|
|
370
209
|
|
|
371
|
-
|
|
210
|
+
`write()` returns a promise that resolves only after the scheduled flush completes. Multiple writes within the 10ms window still share a single batched HTTP request (the batching optimization is preserved). Each write registers a `{resolve, reject}` pair in a `flushWaiters` array. When the `setTimeout` fires and `flush()` completes the HTTP round-trip, all waiters are resolved (or rejected on error). This makes `pendingOps` accurately reflect server-side data state while keeping network-efficient batching.
|
|
372
211
|
|
|
373
|
-
|
|
374
|
-
|
|
375
|
-
- **Steps where ops settle** (data on server, ~200ms after lock release + flush) → continue loop inline
|
|
376
|
-
- **Steps where ops don't settle** (WritableStream kept open across steps) → break loop
|
|
377
|
-
|
|
378
|
-
### Lock-Release Polling Interval Lowered to 10ms
|
|
212
|
+
### Lock-Release Polling Interval
|
|
379
213
|
|
|
380
214
|
`flushablePipe`'s `pollWritableLock` / `pollReadableLock` use `setInterval` to detect when a user releases their stream lock without closing the stream — the Web Streams API has no event for that state. The V2 step executor's `opsSettled` race waits for this poll to resolve after each writable-bearing step body returns, so the polling interval sits on the critical path of every streaming step.
|
|
381
215
|
|
|
382
|
-
The interval was
|
|
383
|
-
|
|
384
|
-
**Fix**: dropped the polling interval from 100ms to 10ms in `packages/core/src/flushable-stream.ts`. Per-step wait drops from ~50ms average to ~5ms (a 10× improvement, expected to scale linearly with the number of writable-bearing steps in a workflow). For `DurableAgent.chat` with one tool call (4 writable-bearing steps), this removes ~180ms from the streaming chat response's critical path. Per-tick work is just `writable.locked` plus a `getWriter()`/`releaseLock()` probe — both microsecond-scale, so 10× more ticks during a stream's lifetime is not measurable in practice.
|
|
385
|
-
|
|
386
|
-
**Follow-up**: Replace the polling entirely with an event-driven release signal — wrap the writable returned from the `WritableStream` reviver with a writer that fires on `releaseLock()` — bringing the wait to ~0ms. The 10ms polling interval is the cheap path that captures most of the available win without the structural change, but every writable-bearing step still pays a ~5ms tax that the event-driven design would eliminate. The structural change is also worth pursuing because it removes a source of timing drift between `world-local` (filesystem-instant) and `world-vercel` (HTTP-deferred) — both would see truly synchronous lock-release detection rather than periodic-poll detection.
|
|
387
|
-
|
|
388
|
-
### Event Consumer Skip Logic Was Too Broad For Wait Replays
|
|
389
|
-
|
|
390
|
-
The V2 handler needs some tolerance for out-of-order replay, especially around step events created by concurrent continuations. An early follow-up broadened that fallback to all wait lifecycle events too, so `onUnconsumedEvent` would skip `wait_created` and the first `wait_completed` whenever they matched a known wait. In the BOA-backed previews that broke `hookDisposeTestWorkflow`: once the first run disposed its hook and went into `sleep('5s')`, a replay could skip the live wait event before `sleep()` registered its subscriber, leaving the run stuck forever at `wait_created`.
|
|
391
|
-
|
|
392
|
-
**Fix**: Keep the step/hook replay tolerance, but narrow the wait fallback to the one case we actually need: duplicate `wait_completed` events that appear *after* an earlier completion for the same wait. The `hookDispose` e2e was also updated to poll for hook registration/disposal instead of relying on fixed 3-5 second sleeps, which made the Vercel preview timing less brittle.
|
|
393
|
-
|
|
394
|
-
### TooEarlyError Retry Delay in Step Executor
|
|
395
|
-
|
|
396
|
-
The `executeStep()` function handles `TooEarlyError` (thrown when a step's `retryAfter` timestamp hasn't been reached yet) by returning a `retry` result with a timeout. The original implementation used a stale access pattern `(err as any).meta?.retryAfter` copied from an older error shape. The `TooEarlyError` class (from `@workflow/errors`) has `retryAfter` as a direct property (number of seconds), not nested under `.meta`. The stale pattern always evaluated to `undefined`, falling back to a 1-second delay regardless of the server's actual retry-after value.
|
|
397
|
-
|
|
398
|
-
**Fix**: Changed to `err.retryAfter ?? 1`, matching the correct pattern used in `step-handler.ts`.
|
|
399
|
-
|
|
400
|
-
### Health Check Endpoint JSON Response
|
|
401
|
-
|
|
402
|
-
The `withHealthCheck()` wrapper in `helpers.ts` was updated (on main) to return a JSON response with `{ healthy, endpoint, specVersion, workflowCoreVersion }` instead of a plain text string. The V2 branch's e2e test still expected `Content-Type: text/plain` and a text body after merging main, causing the "health check endpoint (HTTP)" test to fail across all frameworks and environments.
|
|
403
|
-
|
|
404
|
-
**Fix**: Updated the e2e test to expect `Content-Type: application/json` and validate the JSON body structure, including a `specVersion >= SPEC_VERSION_CURRENT` range assertion.
|
|
405
|
-
|
|
406
|
-
### V2 Combined Bundle Switched from CJS to ESM
|
|
407
|
-
|
|
408
|
-
The V2 combined bundle was initially emitted as CJS by the standalone CLI and Vercel Build Output API builders, while `main` had already moved those outputs to ESM in [#1562](https://github.com/vercel/workflow/pull/1562). Staying on CJS meant `import.meta.url` was polyfilled (often producing the wrong path in re-bundled contexts), and the `world-testing` server had to import from `flow.js` via `createRequire` to force CJS semantics on what was really a CJS bundle.
|
|
409
|
-
|
|
410
|
-
**Fix**: Align V2 with `main`'s ESM defaults:
|
|
411
|
-
|
|
412
|
-
1. The BOA builder emits `__step_registrations.mjs` and `index.mjs`, writes `"type": "module"` in `package.json`, and sets `handler: "index.mjs"` in `.vc-config.json`.
|
|
413
|
-
2. The standalone builder no longer overrides `format`; it inherits the base builder's `'esm'` default.
|
|
414
|
-
3. The standalone config outputs `step.mjs` / `flow.mjs` instead of `.js`.
|
|
415
|
-
4. The `world-testing` server uses a native `import { POST } from '../.well-known/workflow/v1/flow.mjs'` instead of `createRequire`.
|
|
416
|
-
5. `createCombinedBundle`'s final esbuild pass (for `bundleFinalOutput: true`) now prepends the same `createRequire(import.meta.url)` banner used by the workflow/webhook bundles so CJS dependencies that call `require()` for Node.js builtins (for example the `events` module referenced by bundled libraries) still resolve at runtime.
|
|
417
|
-
6. To avoid a duplicate `__createRequire` declaration, the inner steps bundle that gets inlined by the final pass skips the banner — only the outer bundle emits it. This is threaded through via a new `skipEsmRequireBanner` option on `createStepsBundle`.
|
|
418
|
-
|
|
419
|
-
### World specVersion in Health Check Responses
|
|
420
|
-
|
|
421
|
-
The `getWorldHandlers()` return value was updated on main to include `specVersion` (the World's declared spec version). The V2 handler destructures this as `worldSpecVersion` and passes it to `handleHealthCheckMessage()` for inclusion in queue-based health check responses. This was merged alongside the V2 timeout configuration.
|
|
422
|
-
|
|
423
|
-
### Async World Singleton Drift After Merge
|
|
424
|
-
|
|
425
|
-
The later `main` merge changed `getWorld()` and `getWorldHandlers()` to be asynchronous promise-backed singletons, but the eager-processing branch still had synchronous call sites in the V2 runtime path. That left `packages/core/src/runtime.ts` and `packages/core/src/runtime/helpers.ts` trying to access `.events` on a `Promise<World>`, which failed typecheck immediately after the merge.
|
|
426
|
-
|
|
427
|
-
**Fix**: Rebases the V2 workflow entrypoint onto the async world API by lazily awaiting `getWorldHandlers()` when wiring the queue handler and awaiting `getWorld()` at the remaining runtime/helper call sites. This preserves the inline replay loop while matching `main`'s new world initialization contract.
|
|
428
|
-
|
|
429
|
-
### Lazy World Loading for Next.js Production Builds
|
|
430
|
-
|
|
431
|
-
After the async world merge, `packages/core/src/runtime/world.ts` still eagerly imported both `@workflow/world-local` and `@workflow/world-vercel`, and it initialized `createRequire()` from `process.cwd() + '/package.json'` at module load time. In the Next.js production build jobs that caused the generated flow route to pull `@workflow/world-vercel` and its `debug` dependency into local builds, then fail during page-data collection with `module.createRequire failed parsing argument` and `Dynamic require of "tty" is not supported`.
|
|
432
|
-
|
|
433
|
-
**Fix**: Switched the runtime world loader to use `createRequire(import.meta.url)` and moved the local/Vercel world imports behind the existing async `createWorld()` branches. Local Next.js builds now only load the selected world implementation at runtime instead of bundling both worlds eagerly into the route module.
|
|
434
|
-
|
|
435
|
-
### Deferred Next.js Builds Re-Ran Eager Discovery
|
|
436
|
-
|
|
437
|
-
The later merge also pulled `BaseBuilder.createCombinedBundle()` into the deferred Next.js path without a way to pass the already-discovered workflow/step/serde entry sets. As a result, `packages/next/src/builder-deferred.ts` quietly fell back to `discoverEntries()` during production builds, re-emitting `Discovering workflow directives ...` and failing the local build tests that assert deferred mode avoids eager input-graph scans.
|
|
438
|
-
|
|
439
|
-
**Fix**: Threaded explicit `discoveredEntries` through `createCombinedBundle()` and passed the deferred builder's tracked workflow/step/serde file sets into that call. Deferred Next.js builds now reuse the socket/cache-driven discovery state instead of re-running the base eager discovery pass.
|
|
440
|
-
|
|
441
|
-
### Deferred Package Steps Fell Back to Compiled `dist/` Files
|
|
442
|
-
|
|
443
|
-
Once deferred discovery stopped re-running the base eager scan, some package-provided steps were only being rediscovered from built artifacts such as `packages/ai/dist/agent/durable-agent.js`. Those compiled files no longer carried every nested `'use step'` directive, so local production Next.js builds could miss registrations like `@workflow/ai/agent`'s `closeStream` helper and fail at runtime with "step is not registered in the current deployment".
|
|
444
|
-
|
|
445
|
-
**Fix**: The deferred Next.js builder now rewrites discovered workspace package paths from `dist/` back to their matching `src/` files when those sources exist. That keeps deferred bundling pointed at the directive-bearing source modules instead of their compiled output.
|
|
446
|
-
|
|
447
|
-
### Workspace Source Step IDs Lost Export Subpaths
|
|
448
|
-
|
|
449
|
-
Switching deferred builds over to workspace `src/` files fixed the missing nested directives, but it exposed a second mismatch in the SWC manifest path logic. `packages/builders/src/module-specifier.ts` only matched package exports against the on-disk file being transformed, so `packages/ai/src/agent/durable-agent.ts` was assigned `@workflow/ai@...` while the runtime still referenced the exported subpath id `@workflow/ai/agent@...`. Local Next.js agent runs then failed with "Step `step//@workflow/ai/agent@...//closeStream` is not registered" even though the source file was finally back in the bundle.
|
|
450
|
-
|
|
451
|
-
**Fix**: `resolveModuleSpecifier()` now treats workspace source files as the source-backed form of their exported `dist/` targets when deriving step ids. That preserves package export subpaths like `@workflow/ai/agent` for id generation while still bundling the directive-bearing `src/` modules.
|
|
452
|
-
|
|
453
|
-
### Tarball-Staged Next.js Builds Still Lost Package Step Sources
|
|
454
|
-
|
|
455
|
-
The local production and Postgres Next.js jobs stage the workbenches by packing workspace packages into tarballs and installing those tarballs into a temporary `node_modules` tree. Deferred discovery was already willing to rewrite workspace `packages/*/dist/*` files back to `src/*`, but the tarballed `@workflow/ai` package did not publish its `src/` tree and the base builder still treated `node_modules/@workflow/*/src/*` as ordinary package imports. That meant the staged CI path fell back to `dist/` again and dropped nested steps like `@workflow/ai/agent`'s `closeStream`, even after the workspace build path had been fixed.
|
|
456
|
-
|
|
457
|
-
**Fix**: Publish `packages/ai/src` in the tarball, treat source-backed `node_modules/@workflow/*/src/*` files like external workspace source files when generating bundle imports, and extend deferred transitive step discovery to follow bare `workflow` / `@workflow/*` package imports during non-watch builds.
|
|
458
|
-
|
|
459
|
-
### Vercel Step Source Map Expectations Were Too Optimistic
|
|
216
|
+
The interval was lowered from 100ms to 10ms. Per-step wait drops from ~50ms average to ~5ms (scaling linearly with the number of writable-bearing steps in a workflow). Per-tick work is just `writable.locked` plus a `getWriter()`/`releaseLock()` probe — microsecond-scale, so 10× more ticks is not measurable in practice.
|
|
460
217
|
|
|
461
|
-
|
|
218
|
+
**Follow-up**: Replace polling entirely with an event-driven release signal — wrap the writable returned from the `WritableStream` reviver with a writer that fires on `releaseLock()` — bringing the wait to ~0ms. This would also remove a source of timing drift between worlds with synchronous storage (`world-local`) and worlds with HTTP-deferred storage (`world-vercel`).
|
|
462
219
|
|
|
463
|
-
|
|
220
|
+
### Concurrent `step_started` and Attempt Counter
|
|
464
221
|
|
|
465
|
-
|
|
222
|
+
When the V2 handler dispatches N parallel steps as background messages, each background step completion queues a workflow continuation. Up to N continuations may replay concurrently, and each may attempt to start the same not-yet-completed step (since `step_started` succeeds for already-running steps). Each call atomically increments the `attempt` counter — so with N=5 parallel steps, the counter can reach 5 on the first genuine execution.
|
|
466
223
|
|
|
467
|
-
|
|
468
|
-
|
|
469
|
-
The Redis community-world benchmark still loads an external world package that has not adopted the newer `world.streams.*` interface yet. Once the eager-processing changes exercised stream writes through the modern namespace consistently, that adapter started failing with `Cannot read properties of undefined (reading 'writeMulti')` before the benchmark could even start.
|
|
470
|
-
|
|
471
|
-
**Decision**: Community world adapters must implement the `world.streams.*` interface. The runtime legacy stream normalization (`normalizeLegacyWorld`) was removed. Community world e2e tests are skipped until the adapters are updated.
|
|
472
|
-
|
|
473
|
-
### Build Output API Flow Handler Drift
|
|
474
|
-
|
|
475
|
-
The Vercel Build Output API builder still emitted the combined flow function as `index.js`, but the surrounding metadata kept pointing at `index.mjs`. That mismatch meant BOA-based preview deployments published neither `/.well-known/workflow/v1/flow` nor the public manifest, so the Vercel production e2e suite collapsed into manifest `404` errors immediately after deployment.
|
|
476
|
-
|
|
477
|
-
**Fix**: Updated `packages/builders/src/vercel-build-output-api.ts` to point both `.vc-config.json` and manifest extraction at `flow.func/index.js`, which matches the CommonJS file the builder actually writes.
|
|
478
|
-
|
|
479
|
-
### Async World Loading Broke Custom Target Worlds
|
|
480
|
-
|
|
481
|
-
The first lazy-world-loading fix switched package resolution over to `createRequire(import.meta.url)` globally. That solved the Next.js bundling problem for built-in worlds, but it also made custom targets like `@workflow/world-postgres` resolve relative to `@workflow/core` instead of the consuming app. Local Postgres tests then failed at startup with `Cannot find module '@workflow/world-postgres'`.
|
|
482
|
-
|
|
483
|
-
**Fix**: The runtime now creates the package resolver lazily from `process.cwd()/package.json` when possible, falling back to `import.meta.url` only when the app root cannot be resolved. That keeps custom world modules app-relative without reintroducing the eager module-load failure in Next.js builds.
|
|
484
|
-
|
|
485
|
-
### Core Logger Still Pulled `debug` Into Webpack Flow Routes
|
|
486
|
-
|
|
487
|
-
Even after lazy world loading stopped eagerly importing `@workflow/world-vercel`, the generated Next.js webpack flow route still evaluated `packages/core/src/logger.ts` at module load. That file had a top-level `import debug from 'debug'`, which in turn pulled `debug/src/node` and its `tty` dynamic require into `/.well-known/workflow/v1/flow`. Webpack then failed during page-data collection with `Dynamic require of "tty" is not supported`.
|
|
488
|
-
|
|
489
|
-
**Fix**: Replace the static `debug` dependency in the core logger with lightweight `process.env.DEBUG` matching plus `console.debug`. That keeps verbose opt-in logging for local debugging without forcing webpack to bundle `debug` and its Node-only terminal helpers into the flow route.
|
|
490
|
-
|
|
491
|
-
### Deferred Next.js Builder Helper Drift After Merge
|
|
492
|
-
|
|
493
|
-
Merging `main` into the eager-processing branch pulled in a set of helper methods for copied-step import rewriting in `packages/next/src/builder-deferred.ts`, but the corresponding call sites were not present on this branch yet. That left `getRelativeImportSpecifier`, `getStepCopyFileName`, and `rewriteRelativeImportsForCopiedStep` orphaned, and `@workflow/next` failed to build with `TS6133` unused-private-member errors immediately after the merge.
|
|
494
|
-
|
|
495
|
-
**Fix**: Removed the orphaned helper methods during merge resolution and kept the existing deferred-builder behavior unchanged. The copied-step import-rewrite work should land as a complete change set rather than a partial backport from `main`.
|
|
496
|
-
|
|
497
|
-
### Next.js React Step Fixture and `eval('require(...)')`
|
|
498
|
-
|
|
499
|
-
The `nextjs-webpack` e2e suite still failed after the merge in `workflows/8_react_render.tsx`, where the step intentionally did `eval('require("react-dom/server")')` to avoid Next.js linting rules around importing `react-dom/server` directly. That pattern was brittle under webpack rebundling: even though the intermediate step bundle had a `createRequire(import.meta.url)` banner, the rebundled route still failed at runtime with `TypeError: require is not a function`.
|
|
500
|
-
|
|
501
|
-
**Fix**: Updated the React-rendering step fixture in both Next.js workbenches to use `await import('react-dom/server')` instead. The test still exercises server-side React rendering inside a step, but no longer depends on bundler-specific `eval('require(...)')` behavior.
|
|
502
|
-
|
|
503
|
-
## Inline Execution Verification Tests
|
|
504
|
-
|
|
505
|
-
The `@workflow/world-testing` package includes invocation-counting tests that verify the V2 inline loop behavior for each workflow pattern:
|
|
506
|
-
|
|
507
|
-
| Workflow Pattern | Expected Invocations | Why |
|
|
508
|
-
|-----------------|---------------------|-----|
|
|
509
|
-
| Sequential steps (3 adds) | **1** | All steps execute inline |
|
|
510
|
-
| Sequential steps + WritableStream | **1** | Ops settle via flush waiter promises (500ms race) |
|
|
511
|
-
| Sleep (1s) + step | **2** | Sleep requires queue round-trip |
|
|
512
|
-
| Promise.all (2 steps) | **2-3** | Background step + inline replay after all steps done |
|
|
513
|
-
|
|
514
|
-
The test server tracks flow handler invocations per `runId` via an internal counter. Each test asserts the exact invocation count after the workflow completes.
|
|
515
|
-
|
|
516
|
-
### world-testing Flow Invocation Counting Missed Wrapped Queue Payloads
|
|
517
|
-
|
|
518
|
-
The inline-execution assertions in `packages/world-testing` count how many times the flow handler runs by inspecting the queue callback body and extracting `runId`. After the queue callback shape drifted, some worlds were only exposing the workflow payload under `body.payload.runId`, so the helper recorded `0` invocations even when the workflow completed correctly. That showed up in CI as the Postgres inline-execution spec failing its "single flow invocation" assertion.
|
|
519
|
-
|
|
520
|
-
**Fix**: Accept both top-level `runId` and nested `payload.runId` when tracking flow invocations in the embedded test server.
|
|
521
|
-
|
|
522
|
-
### Turbopack NFT Tracing Errors in V2 Combined Flow Route
|
|
523
|
-
|
|
524
|
-
The V2 combined flow route imports the step registrations bundle (`__step_registrations.js`), which esbuild produces as a monolithic file. On `main`, step registrations live in a separate route (`step/route.js`), so Turbopack traces them independently. In V2, Turbopack traces the step registrations through the flow route's import graph, encountering `world.ts` code with `process.cwd()`, dynamic `import()` calls to `@workflow/world-local`/`@workflow/world-vercel`, and `createRequire()` patterns — all of which trigger fatal NFT (Node File Trace) errors.
|
|
525
|
-
|
|
526
|
-
**Fix**: Introduced `get-world-lazy.ts`, a globalThis `Symbol.for`-based accessor that replaces the static `import { getWorld } from './runtime/world.js'` in all step-side modules (`serialization.ts`, `run.ts`, `helpers.ts`, `start.ts`, `resume-hook.ts`). This breaks the static import chain from step code to `world.ts`, preventing esbuild from bundling `world.ts` (and its transitive deps) into the step registrations. The step registrations bundle dropped from ~37k lines to ~6.6k lines (matching `main`), with zero `process.cwd()` or world package references.
|
|
527
|
-
|
|
528
|
-
The `getWorldLazy()` function reads from the globalThis world singleton cache (populated by the runtime's `getWorld()` on first call). When the cache is empty (e.g., `start()` called from application code before any workflow runs), it falls back to a dynamic `import()` of `world.js` to initialize the world.
|
|
529
|
-
|
|
530
|
-
Additional changes for Turbopack compatibility:
|
|
531
|
-
- Removed `stepEntrypoint` re-export from `runtime.ts` (V2 doesn't use separate step routes)
|
|
532
|
-
- Lazy-loaded `getPort` via `createRequire` with opaque specifier to prevent `@workflow/utils/get-port` filesystem operations from being traced
|
|
533
|
-
- `getRuntimeRequire()` uses `process.cwd()` as primary resolution base (for custom world packages like `@workflow/world-postgres` that are app-level deps, not `@workflow/core` deps), with `import.meta.url` fallback
|
|
534
|
-
|
|
535
|
-
### Cold-Start `MODULE_NOT_FOUND: './world.js'` From `getWorldLazy` Fallback
|
|
536
|
-
|
|
537
|
-
The `getWorldLazy()` design assumed one of two paths would always succeed: either `globalThis[GetWorldFnKey]` is populated (because some prior code reached `world.ts`'s module body), or the dynamic `import('./world.js')` fallback resolves at runtime.
|
|
538
|
-
|
|
539
|
-
Both assumptions break for routes that consume `start` (or any other `getWorldLazy` consumer) without going through the queue-driven flow handler first:
|
|
540
|
-
|
|
541
|
-
1. Webpack and Turbopack tree-shake the named import `{ getWorld } from './runtime/world.js'` out of `runtime.ts` once a consumer only uses `start`. `world.ts` is dropped from the bundle entirely, so its module-load `globalThis[GetWorldFnKey] ??= getWorld` registration never fires.
|
|
542
|
-
2. The dynamic-import fallback inside `get-world-lazy.ts` builds the specifier `./world.js` at runtime to evade bundler tracing — but webpack inlines `get-world-lazy.js` into the route bundle, so the relative specifier resolves against `/var/task/<app>/.next/server/app/<route>/route.js` where no sibling `world.js` exists. Node throws `MODULE_NOT_FOUND`.
|
|
543
|
-
|
|
544
|
-
The symptom: the very first request that goes through `start()` on a cold serverless invocation fails. Once any other code path (typically the queue-driven `/.well-known/workflow/v1/flow` route, which uses `getWorld` directly via `workflowEntrypoint`) has loaded `world.ts`, subsequent `start()` calls succeed for the rest of the process lifetime — making the failure flake-shaped: hard to reproduce in dev where everything tends to be warmed, but reliable on first user traffic into a fresh function instance.
|
|
545
|
-
|
|
546
|
-
**Fix**: Added `@workflow/core/runtime/world-init`, a server-only side-effect module that imports `./world.js` purely for its module-load side effect (the globalThis registration). It's exported via package conditions:
|
|
547
|
-
|
|
548
|
-
- `default` → `./dist/runtime/world-init.js` (real, loads `world.ts`)
|
|
549
|
-
- `workflow` → `./dist/workflow/world-init-stub.js` (empty, used by VM/step bundles)
|
|
550
|
-
|
|
551
|
-
`packages/workflow/src/api.ts` (the host file behind `workflow/api`'s `default` condition) imports it for its side effect. The matching VM/step entry `api-workflow.ts` does not, so `world.ts` and its server-only deps (`@workflow/world-local`, `@workflow/world-vercel`, `cbor-x`, …) stay out of the workflow sandbox bundle.
|
|
552
|
-
|
|
553
|
-
Reverification: built bundles for `vade-review` (Next.js webpack) show `createLocalWorld`/`createVercelWorld`/`GetWorldFnKey` present in the route's vendor chunk for `@workflow/core` (zero before the fix), and the workflow VM bundle's `flow/route.js` and `__step_registrations.js` continue to have zero references to either the world-init module or `world.ts`. Cold-start `POST /api/review/submit` succeeds on the first request after a fresh server boot — the regression case.
|
|
554
|
-
|
|
555
|
-
The dynamic-import fallback in `get-world-lazy.ts` is preserved as defense-in-depth for environments outside the documented configurations (CJS test runners, scripts that import deeply into `@workflow/core` without going through `workflow/api`).
|
|
556
|
-
|
|
557
|
-
### Run#returnValue Worker Deadlock in V2 Inline Execution
|
|
558
|
-
|
|
559
|
-
When a workflow calls `start()` to spawn child workflows (e.g., `fibonacciWorkflow`), the parent's `Run#returnValue` step polls the child's completion status in a blocking loop (`while (true) { ... sleep(1000) ... }`). In V2, this step is executed inline by the step executor, holding a worker thread slot. If the child workflow's queue message is waiting for the same worker pool, the parent blocks the child from starting — a classic deadlock.
|
|
560
|
-
|
|
561
|
-
**Fix**: `Run#pollReturnValue()` detects whether it's running inside a step executor (via `contextStorage.getStore()`) and, if so, throws `TooEarlyError` instead of polling in a blocking loop. `TooEarlyError` is handled specially by the step executor — it returns `{ type: 'retry', timeoutSeconds }` which re-queues the step with a 1-second delay, freeing the worker to process child workflows. Unlike `RetryableError`, `TooEarlyError` does NOT count against `maxRetries`, so polling steps can retry indefinitely until the child completes.
|
|
562
|
-
|
|
563
|
-
When called from outside a step (e.g., test code, API routes), `pollReturnValue()` retains the original blocking loop behavior for backward compatibility.
|
|
224
|
+
The max retries check in `executeStep()` only enforces when `step.error` exists — distinguishing actual retries (failed → retry with error) from concurrent first-attempt races (multiple handlers start the same step simultaneously without any prior failure). Concurrent starts are harmless since `step_completed` idempotency ensures only the first completion wins.
|
|
564
225
|
|
|
565
226
|
### Unconsumed Event Check Two-Phase Drain
|
|
566
227
|
|
|
567
|
-
|
|
228
|
+
The `EventsConsumer`'s unconsumed event check uses a two-phase promise queue drain: yield once after the first drain (via `setTimeout(0)`) so cross-VM promise chains can append follow-up async work, then re-drain before checking. This improves timing for scenarios like `step_completed` → for-await loop resume → next hook hydration.
|
|
568
229
|
|
|
569
230
|
The check additionally arms a `DEFERRED_CHECK_DELAY_MS = 100` `setTimeout` after the second drain, since Node.js does not guarantee that `setTimeout(0)` fires after all cross-context microtasks settle. Any `subscribe()` call arriving during that 100ms window cancels the check via version invalidation + `clearTimeout`, so the delay only adds latency to genuine corruption — never to the happy path.
|
|
570
231
|
|
|
571
|
-
**Follow-up**: 100ms is a heuristic chosen empirically
|
|
232
|
+
**Follow-up**: 100ms is a heuristic chosen empirically. A deterministic settlement signal — for example, a "VM idle" callback exposed by the workflow VM bridge that fires only after all pending cross-context promise chains have resolved — would let the consumer fire the unconsumed-event check immediately on quiescence instead of waiting for a wall-clock timeout.
|
|
233
|
+
|
|
234
|
+
### Lazy World Loading
|
|
572
235
|
|
|
573
|
-
|
|
236
|
+
Static imports of `world-local` and `world-vercel` from the runtime caused two distinct build-time issues: Next.js production builds pulled both worlds (including their Node-only deps like `debug`'s `tty` requires) into the route module, and Turbopack's NFT (Node File Trace) errored on `process.cwd()` and dynamic `import()` patterns it couldn't statically analyze.
|
|
574
237
|
|
|
575
|
-
|
|
238
|
+
A `getWorldLazy()` accessor (backed by a `globalThis` `Symbol.for` cache) replaces the static import in step-side modules. This breaks the static import chain from step code to `world.ts`, preventing both worlds from being bundled into the step registrations.
|
|
576
239
|
|
|
577
|
-
|
|
240
|
+
Because tree-shaking can otherwise drop `world.ts`'s module-load registration entirely, a server-only side-effect module (`@workflow/core/runtime/world-init`) imports `./world.js` purely for its module-load side effect. It's wired via package conditions:
|
|
578
241
|
|
|
579
|
-
|
|
242
|
+
- `default` → real, loads `world.ts`
|
|
243
|
+
- `workflow` → empty stub, used by VM/step bundles
|
|
580
244
|
|
|
581
|
-
|
|
582
|
-
|-----------|-----------|--------|
|
|
583
|
-
| Unit Tests | core | 581/581 |
|
|
584
|
-
| Embedded Tests | world-testing | 9/9 (including inline execution) |
|
|
585
|
-
| Local Dev | 14 frameworks | All pass |
|
|
586
|
-
| Local Prod | 14 configurations | All pass |
|
|
587
|
-
| Postgres | 14 frameworks | All pass |
|
|
588
|
-
| Vercel Prod | 11 frameworks | All pass |
|
|
589
|
-
| Vercel Deployments | 15 projects | All succeed |
|
|
590
|
-
| Community Worlds | Turso, MongoDB, Redis | All pass |
|
|
591
|
-
| Windows | e2e | Pass |
|
|
245
|
+
This guarantees the world is loaded for routes that consume `start()` without going through the queue-driven flow handler first, while keeping `world.ts` and its server-only deps out of the workflow sandbox bundle.
|
|
592
246
|
|
|
593
|
-
|
|
247
|
+
### Community Worlds and the `world.streams` API
|
|
594
248
|
|
|
595
|
-
|
|
249
|
+
Community world adapters must implement the `world.streams.*` interface. The runtime legacy stream normalization was removed as part of this work.
|
|
@@ -7,321 +7,69 @@ description: Overhaul run start logic to tolerate world storage unavailability,
|
|
|
7
7
|
|
|
8
8
|
## Motivation
|
|
9
9
|
|
|
10
|
-
When `world` storage is unavailable but the queue is up,
|
|
11
|
-
`start()` previously failed entirely because `world.events.create(run_created)`
|
|
12
|
-
is called before `world.queue()`. This change decouples run creation from queue
|
|
13
|
-
dispatch so that runs can still be accepted when storage is degraded.
|
|
10
|
+
When `world` storage is unavailable but the queue is up, `start()` previously failed entirely because `world.events.create(run_created)` is called before `world.queue()`. This change decouples run creation from queue dispatch so that runs can still be accepted when storage is degraded.
|
|
14
11
|
|
|
15
|
-
Additionally, the runtime previously called `world.runs.get(runId)` before
|
|
16
|
-
`run_started`, adding an extra round-trip. By always calling `run_started`
|
|
17
|
-
directly, we save that round-trip and can return pre-loaded events in the
|
|
18
|
-
response to skip the initial `events.list` call, reducing TTFB.
|
|
12
|
+
Additionally, the runtime previously called `world.runs.get(runId)` before `run_started`, adding an extra round-trip. By always calling `run_started` directly, we save that round-trip and can return pre-loaded events in the response to skip the initial `events.list` call, reducing TTFB.
|
|
19
13
|
|
|
20
14
|
## Design
|
|
21
15
|
|
|
22
|
-
### `start()` changes
|
|
16
|
+
### `start()` changes
|
|
23
17
|
|
|
24
|
-
- `world.events.create` (run_created) and `world.queue` are now called **in parallel**
|
|
25
|
-
|
|
26
|
-
- If `events.create` errors with **
|
|
27
|
-
creation failed but the run was accepted — creation will be re-tried async by the
|
|
28
|
-
runtime when it processes the queue message. The returned `Run` instance is marked
|
|
29
|
-
with `resilientStart = true`.
|
|
30
|
-
- If `events.create` errors with **409** (EntityConflictError), the run already exists
|
|
31
|
-
(e.g., the queue handler's resilient start path created it first due to a cold-start
|
|
32
|
-
race). This is treated as success.
|
|
18
|
+
- `world.events.create` (run_created) and `world.queue` are now called **in parallel** via `Promise.allSettled`.
|
|
19
|
+
- If `events.create` errors with **429 or 5xx**, we log a warning saying that run creation failed but the run was accepted — creation will be re-tried async by the runtime when it processes the queue message. The returned `Run` instance is marked with `resilientStart = true`.
|
|
20
|
+
- If `events.create` errors with **409** (EntityConflictError), the run already exists (e.g., the queue handler's resilient start path created it first due to a cold-start race). This is treated as success.
|
|
33
21
|
- If `world.queue` fails, we still throw — the run truly failed and was not enqueued.
|
|
34
|
-
- The queue invocation now receives all the run inputs (`input`, `deploymentId`,
|
|
35
|
-
|
|
36
|
-
create the run later if needed.
|
|
37
|
-
- When the runtime re-enqueues itself, it does **not** pass these inputs — only the
|
|
38
|
-
first queue cycle carries them.
|
|
22
|
+
- The queue invocation now receives all the run inputs (`input`, `deploymentId`, `workflowName`, `specVersion`, `executionContext`) via `runInput` so the runtime can create the run later if needed.
|
|
23
|
+
- When the runtime re-enqueues itself, it does **not** pass these inputs — only the first queue cycle carries them.
|
|
39
24
|
|
|
40
|
-
### `workflowEntrypoint` changes
|
|
25
|
+
### `workflowEntrypoint` changes
|
|
41
26
|
|
|
42
|
-
- When calling `world.events.create` with `run_started`, we now also always pass the
|
|
43
|
-
run input that was sent through the queue, if available. The response will still be on off:
|
|
44
|
-
- **200 with event (now running)**: As usual, but the server could have used the run input to create the run if it didn't exist yet. The response will be opaque to the runtime.
|
|
45
|
-
- **200 without event (already running)**: As usual
|
|
46
|
-
- **409 or 410 (already finished)**: As usual
|
|
27
|
+
- When calling `world.events.create` with `run_started`, we now also always pass the run input that was sent through the queue, if available. The world is responsible for creating the run if it doesn't already exist.
|
|
47
28
|
|
|
48
|
-
### `Run.returnValue` polling
|
|
29
|
+
### `Run.returnValue` polling
|
|
49
30
|
|
|
50
|
-
- When `resilientStart` is true on the Run instance (run_created failed), the
|
|
51
|
-
|
|
52
|
-
(1s + 3s + 6s = 10s total) to give the queue time to deliver and the runtime
|
|
53
|
-
to create the run via `run_started`.
|
|
54
|
-
- When `resilientStart` is false (normal path), 404 fails immediately — no delay
|
|
55
|
-
for the common case of a wrong run ID.
|
|
31
|
+
- When `resilientStart` is true on the Run instance (run_created failed), the `pollReturnValue` loop retries on `WorkflowRunNotFoundError` up to 3 times (1s + 3s + 6s = 10s total) to give the queue time to deliver and the runtime to create the run via `run_started`.
|
|
32
|
+
- When `resilientStart` is false (normal path), 404 fails immediately — no delay for the common case of a wrong run ID.
|
|
56
33
|
|
|
57
|
-
### World
|
|
34
|
+
### World contract changes
|
|
58
35
|
|
|
59
|
-
- Posting `run_started` to a **non-existent** run is now allowed when the run input is
|
|
60
|
-
|
|
61
|
-
1. Creates a `run_created` event first (so the event log is consistent).
|
|
62
|
-
2. Strips the input from the `run_started` event data (it lives on `run_created`).
|
|
63
|
-
3. Then creates the `run_started` event normally.
|
|
64
|
-
4. Emits a log and a Datadog metric (`workflow_server.resilient_start.run_created_via_run_started`)
|
|
65
|
-
to track when this fallback path is hit.
|
|
66
|
-
- When `run_started` encounters an **already-running** run, all worlds return `{ run }`
|
|
67
|
-
with `event: undefined` instead of throwing. No duplicate event is created.
|
|
36
|
+
- Posting `run_started` to a **non-existent** run is now allowed when the run input is sent along with the payload. The world creates a `run_created` event first (so the event log is consistent), then creates the `run_started` event normally.
|
|
37
|
+
- When `run_started` encounters an **already-running** run, all worlds return `{ run }` with `event: undefined` instead of throwing. No duplicate event is created.
|
|
68
38
|
|
|
69
39
|
### Queue transport changes
|
|
70
40
|
|
|
71
|
-
`Uint8Array` values (the serialized workflow input in `runInput`) don't survive plain
|
|
72
|
-
JSON serialization. Each world uses a transport that preserves binary data:
|
|
41
|
+
`Uint8Array` values (the serialized workflow input in `runInput`) don't survive plain JSON serialization. Each world uses a transport that preserves binary data:
|
|
73
42
|
|
|
74
|
-
- **world-vercel**: CBOR transport — CBOR-encodes the entire queue payload into a
|
|
75
|
-
|
|
76
|
-
- **world-
|
|
77
|
-
from `fs.ts` that encode Uint8Array as `{ __type: 'Uint8Array', data: '<base64>' }`.
|
|
78
|
-
- **world-postgres**: Inline typed JSON transport — same tagged-envelope approach as
|
|
79
|
-
world-local, inlined since world-postgres doesn't import from world-local.
|
|
43
|
+
- **world-vercel**: CBOR transport — CBOR-encodes the entire queue payload into a `Buffer` and uses `BufferTransport` from `@vercel/queue`. Uint8Array survives natively.
|
|
44
|
+
- **world-local**: `TypedJsonTransport` — encodes Uint8Array as `{ __type: 'Uint8Array', data: '<base64>' }`.
|
|
45
|
+
- **world-postgres**: Inline typed JSON transport — same tagged-envelope approach as world-local.
|
|
80
46
|
|
|
81
47
|
## Decisions
|
|
82
48
|
|
|
83
|
-
1. **Parallel not sequential**: We chose `Promise.allSettled` over sequential calls to
|
|
84
|
-
minimize latency in the happy path.
|
|
49
|
+
1. **Parallel not sequential**: We chose `Promise.allSettled` over sequential calls to minimize latency in the happy path.
|
|
85
50
|
|
|
86
|
-
2. **Already-running returns run without event**: When `run_started` encounters an
|
|
87
|
-
already-running run, all worlds return `{ run }` with `event: undefined` (no
|
|
88
|
-
`events` array) instead of throwing. The runtime detects this by checking for
|
|
89
|
-
`result.event === undefined`. This avoids an extra `world.runs.get` round-trip.
|
|
51
|
+
2. **Already-running returns run without event**: When `run_started` encounters an already-running run, all worlds return `{ run }` with `event: undefined` (no `events` array) instead of throwing. The runtime detects this by checking for `result.event === undefined`. This avoids an extra `world.runs.get` round-trip.
|
|
90
52
|
|
|
91
|
-
3. **Events in 200 response**: We only return events on the 200 path (first caller).
|
|
92
|
-
On the already-running path, we fall back to the normal `events.list` call. This is
|
|
93
|
-
correct because only on 200 can we be certain we know the full event history.
|
|
53
|
+
3. **Events in 200 response**: We only return events on the 200 path (first caller). On the already-running path, we fall back to the normal `events.list` call. This is correct because only on 200 can we be certain we know the full event history.
|
|
94
54
|
|
|
95
|
-
4. **Conditional 404 retry on Run.returnValue**: Only when `resilientStart = true`
|
|
96
|
-
(run_created failed). Normal runs fail fast on 404.
|
|
55
|
+
4. **Conditional 404 retry on Run.returnValue**: Only when `resilientStart = true` (run_created failed). Normal runs fail fast on 404.
|
|
97
56
|
|
|
98
57
|
## Known concerns
|
|
99
58
|
|
|
100
|
-
### Cold-start race on Vercel
|
|
59
|
+
### Cold-start race on Vercel
|
|
101
60
|
|
|
102
|
-
On Vercel, the parallel dispatch can cause the queue message to be processed before
|
|
103
|
-
`run_created` completes, if `run_created` hits a cold-start lambda. Confirmed via
|
|
104
|
-
Datadog: the `run_started` request hit a warm lambda (23ms) while `run_created` hit
|
|
105
|
-
a cold lambda (727ms), even though `run_created` arrived at the edge 116ms earlier.
|
|
106
|
-
When this happens:
|
|
61
|
+
On Vercel, the parallel dispatch can cause the queue message to be processed before `run_created` completes, if `run_created` hits a cold-start lambda. When this happens:
|
|
107
62
|
|
|
108
63
|
1. The runtime's resilient start path creates the run from `run_started`.
|
|
109
64
|
2. The original `run_created` arrives and gets 409 (EntityConflictError).
|
|
110
65
|
3. `start()` treats the 409 as success (the run exists).
|
|
111
66
|
|
|
112
|
-
|
|
113
|
-
in this case (409 is not a retryable error), so `returnValue` fails fast on 404.
|
|
67
|
+
The `resilientStart` flag is NOT set on the Run instance in this case (409 is not a retryable error), so `returnValue` fails fast on 404.
|
|
114
68
|
|
|
115
|
-
###
|
|
69
|
+
### Atomicity of run entity creation
|
|
116
70
|
|
|
117
|
-
|
|
118
|
-
`events.create(run_created)` finishes writing to the shared filesystem. The
|
|
119
|
-
resilient start path should handle this, but Local Prod tests showed occasional
|
|
120
|
-
runs stuck at `pending` (no `run_started` event), and Windows CI showed
|
|
121
|
-
"Unconsumed event in event log" errors from duplicate `run_created` events.
|
|
71
|
+
The normal `run_created` path and the resilient start path can race on creating the run entity. In `world-local`, both paths use `writeExclusive` (O_CREAT|O_EXCL) — atomic at the OS level, so exactly one writer wins and the other gets EEXIST. The normal path throws `EntityConflictError` on conflict (handled by `start()` as 409); the resilient start path re-reads the run from disk on conflict.
|
|
122
72
|
|
|
123
|
-
|
|
124
|
-
resilient start path. Both used `writeJSON` which checks existence with
|
|
125
|
-
`fs.access()` (non-atomic), so both could pass the check and write separate
|
|
126
|
-
`run_created` events with different event IDs. Fixed by switching both paths to
|
|
127
|
-
`writeExclusive` (O_CREAT|O_EXCL) — see retrospective items 12 and 16.
|
|
73
|
+
In `world-postgres`, the resilient start path uses `onConflictDoNothing` plus a re-read on conflict for the same effect, with the same outcome on either side of the race.
|
|
128
74
|
|
|
129
|
-
|
|
130
|
-
|
|
131
|
-
- [x] ~~Investigate Local Prod test flakiness~~ — resolved via `writeExclusive`
|
|
132
|
-
for run entity creation (retrospective items 12, 16).
|
|
133
|
-
- [ ] Monitor the Datadog metric in production to understand how often the fallback is hit.
|
|
134
|
-
- [x] ~~Events optimization for re-enqueue cycles~~ — decided against. The
|
|
135
|
-
already-running path returns early without writing an event, so preloading
|
|
136
|
-
events there would require an extra filesystem/DB query on every re-enqueue.
|
|
137
|
-
More importantly, on Vercel with at-least-once delivery, multiple lambdas can
|
|
138
|
-
process the same run concurrently — the event snapshot could be stale or
|
|
139
|
-
incomplete. The runtime's fallback to `events.list` is the correct behavior
|
|
140
|
-
for re-enqueue cycles.
|
|
141
|
-
- [x] ~~CborTransport pass-through~~ — refactored. `encode()`/`decode()` now
|
|
142
|
-
live inside `CborTransport.serialize()`/`deserialize()`, matching the pattern
|
|
143
|
-
used by TypedJsonTransport (world-local) and the inline transport
|
|
144
|
-
(world-postgres). Call sites pass plain objects instead of pre-encoded buffers.
|
|
145
|
-
|
|
146
|
-
## Development retrospective
|
|
147
|
-
|
|
148
|
-
Chronological log of mistakes, misunderstandings, and reverted approaches during
|
|
149
|
-
development. Included for future reference when working on similar cross-cutting
|
|
150
|
-
runtime changes.
|
|
151
|
-
|
|
152
|
-
### 1. Uint8Array corruption through JSON queue transport
|
|
153
|
-
|
|
154
|
-
The initial implementation passed `runInput.input` (a `Uint8Array`) directly through
|
|
155
|
-
the queue payload. `Uint8Array` doesn't survive `JSON.stringify` — it becomes
|
|
156
|
-
`{"0":72,"1":101,...}`. This corrupted the workflow input when the resilient start
|
|
157
|
-
path tried to recreate the run from the queue-delivered data.
|
|
158
|
-
|
|
159
|
-
Caught by the `spawnWorkflowFromStepWorkflow` e2e test and the `world-testing`
|
|
160
|
-
embedded tests, which failed with "Invalid input" from devalue's `unflatten()`.
|
|
161
|
-
|
|
162
|
-
Three approaches were tried before landing on the final solution:
|
|
163
|
-
|
|
164
|
-
1. **Base64 encoding** (`btoa`/`atob`) — worked but fragile. The decode side used
|
|
165
|
-
`typeof runInput.input === 'string'` as a discriminant, which was flagged as
|
|
166
|
-
dangerous since non-binary inputs could also be strings.
|
|
167
|
-
2. **`Array.from()`/`new Uint8Array()`** — replaced base64 with a plain number array.
|
|
168
|
-
Two problems: (a) 3x JSON size regression vs base64, and (b) `Array.isArray()`
|
|
169
|
-
false-positives on v1Compat runs where `dehydrateWorkflowArguments` returns
|
|
170
|
-
devalue's flat Array format.
|
|
171
|
-
3. **CBOR + BufferTransport** (final) — world-vercel CBOR-encodes the queue payload;
|
|
172
|
-
world-local and world-postgres use a `TypedJsonTransport` with a tagged envelope.
|
|
173
|
-
|
|
174
|
-
### 2. Forgot to commit world-postgres transport fix (twice)
|
|
175
|
-
|
|
176
|
-
After fixing world-local and world-vercel queue transports, the same `JsonTransport`
|
|
177
|
-
corruption bug existed in world-postgres. The fix was written during a session but
|
|
178
|
-
never committed — lost when the working directory was reset via stash/checkout. This
|
|
179
|
-
happened twice. The fix only landed on the third attempt when it was committed and
|
|
180
|
-
pushed immediately. All 14 Postgres e2e jobs failed each time.
|
|
181
|
-
|
|
182
|
-
### 3. Incorrect diagnosis of Vercel Prod 409 errors
|
|
183
|
-
|
|
184
|
-
Multiple Vercel Prod e2e tests failed with `EntityConflictError: Workflow run with
|
|
185
|
-
ID wrun_... already exists` on `run_created`. The initial assumption was that VQS
|
|
186
|
-
couldn't deliver the queue message fast enough to beat the `run_created` call.
|
|
187
|
-
|
|
188
|
-
Datadog logs showed otherwise: the `run_created` request arrived at Vercel's edge
|
|
189
|
-
116ms before `run_started`, but `run_created` hit a cold-start lambda (727ms) while
|
|
190
|
-
`run_started` hit a warm one (23ms). Cold starts can invert expected execution order.
|
|
191
|
-
|
|
192
|
-
### 4. Removed EntityConflictError catch, then had to restore it
|
|
193
|
-
|
|
194
|
-
The `workflowEntrypoint` error handler originally caught both `EntityConflictError`
|
|
195
|
-
and `RunExpiredError`. When adding the "already-running returns run without event"
|
|
196
|
-
behavior, `EntityConflictError` was removed from the catch since the new worlds
|
|
197
|
-
wouldn't throw it. Reviewer flagged this: old worlds or world-vercel hitting an
|
|
198
|
-
older workflow-server could still throw it. The catch was restored.
|
|
199
|
-
|
|
200
|
-
### 5. Duplicate `startedAt` check
|
|
201
|
-
|
|
202
|
-
After refactoring the `run_started` flow, a `workflowRun.startedAt` null check
|
|
203
|
-
existed both inside the `try` block and after the `catch` block. The second was
|
|
204
|
-
unreachable. Removed after review.
|
|
205
|
-
|
|
206
|
-
### 6. WORKFLOW_SERVER_URL_OVERRIDE left set
|
|
207
|
-
|
|
208
|
-
During development, `WORKFLOW_SERVER_URL_OVERRIDE` was set to a test URL pointing
|
|
209
|
-
at the workflow-server preview deployment and accidentally committed. The Vercel
|
|
210
|
-
bot flagged this. Reset to empty string.
|
|
211
|
-
|
|
212
|
-
### 7. e2e test assertion was wrong
|
|
213
|
-
|
|
214
|
-
The resilient start e2e test stubbed `world.events.create` and asserted
|
|
215
|
-
`createCallCount >= 2`. But the stub only intercepts calls from the test runner
|
|
216
|
-
process — the server uses its own world. `createCallCount` was always 1. Changed
|
|
217
|
-
to `expect(createCallCount).toBe(1)`.
|
|
218
|
-
|
|
219
|
-
### 8. Misattributed Local Prod timeouts as "pre-existing"
|
|
220
|
-
|
|
221
|
-
Local Prod tests showed 60-second timeouts across various tests. Initially dismissed
|
|
222
|
-
as CI flakes. Checking main's CI showed all Local Prod tests pass on main — the
|
|
223
|
-
timeouts are caused by our changes. Should have compared against main immediately.
|
|
224
|
-
|
|
225
|
-
### 9. Attempted to revert parallel dispatch
|
|
226
|
-
|
|
227
|
-
After identifying Local Prod timeouts, `start()` was partially reverted back to
|
|
228
|
-
sequential dispatch. The user pointed out that parallel dispatch is the core value
|
|
229
|
-
proposition of the PR. The revert was undone.
|
|
230
|
-
|
|
231
|
-
### 10. WorkflowRunNotFoundError retry was unconditional
|
|
232
|
-
|
|
233
|
-
The initial `pollReturnValue` retry on `WorkflowRunNotFoundError` applied to all
|
|
234
|
-
`Run` instances. A user calling `getRun()` with a wrong ID would wait 10 seconds
|
|
235
|
-
before getting a 404. Fixed by adding a `resilientStart` flag: only retries when
|
|
236
|
-
`run_created` actually failed.
|
|
237
|
-
|
|
238
|
-
### 11. Changeset `minor` vs `patch`
|
|
239
|
-
|
|
240
|
-
The changeset was created with `"@workflow/core": minor`. Reviewer flagged this as
|
|
241
|
-
violating repo rules ("all changes should be patch"). Changed after discussion.
|
|
242
|
-
|
|
243
|
-
### 12. world-local TOCTOU race causing duplicate `run_created` events (Windows CI)
|
|
244
|
-
|
|
245
|
-
The resilient start path AND the normal `run_created` path in `world-local/events-storage.ts`
|
|
246
|
-
both used `writeJSON` to create the run entity. `writeJSON` checks file existence with
|
|
247
|
-
`fs.access()` then writes via temp+rename — a classic TOCTOU race. On the local world,
|
|
248
|
-
the queue delivers via an async IIFE in the same event loop, so `events.create(run_created)`
|
|
249
|
-
and `events.create(run_started)` (with resilient start) run concurrently:
|
|
250
|
-
|
|
251
|
-
1. Both paths call `fs.access(runPath)` → ENOENT (file doesn't exist yet)
|
|
252
|
-
2. Both proceed to write → the last `fs.rename` wins
|
|
253
|
-
3. Both succeed → both write their own `run_created` event with different event IDs
|
|
254
|
-
4. During replay, the consumer sees two `run_created` events → "Unconsumed event" error
|
|
255
|
-
|
|
256
|
-
This caused consistent failures in `world-testing` embedded tests on Windows CI (`hooks`,
|
|
257
|
-
`supports null bytes in step results`, `retriable and fatal errors` — all timing out at
|
|
258
|
-
60s with "Unconsumed event in event log" errors). Linux CI was not affected because the
|
|
259
|
-
timing was different enough that the race window was rarely hit.
|
|
260
|
-
|
|
261
|
-
Fixed by switching BOTH paths to `writeExclusive` (O_CREAT|O_EXCL), which is atomic at
|
|
262
|
-
the OS level — exactly one writer wins, the other gets EEXIST. The normal `run_created`
|
|
263
|
-
path throws `EntityConflictError` on conflict (handled by `start()` as 409). The resilient
|
|
264
|
-
start path re-reads the run from disk on conflict. Either way, only one `run_created`
|
|
265
|
-
event is written.
|
|
266
|
-
|
|
267
|
-
### 13. Non-atomic run + run_created event in world-postgres resilient path
|
|
268
|
-
|
|
269
|
-
The resilient start path in `world-postgres/storage.ts` did two separate writes (run
|
|
270
|
-
insert, then event insert) without a transaction. If the process crashed between them,
|
|
271
|
-
the run would exist without a `run_created` event — an inconsistent event log.
|
|
272
|
-
|
|
273
|
-
A `drizzle.transaction()` wrapper was attempted but dropped due to TypeScript inference
|
|
274
|
-
issues with drizzle's transaction callback and the insert builder's overloads. The current
|
|
275
|
-
fix keeps the two writes sequential but adds the same conflict-aware re-read pattern as
|
|
276
|
-
world-local: when `onConflictDoNothing` produces no result (run already existed), the run
|
|
277
|
-
is re-read so downstream logic sees the real state. The narrow crash window between the
|
|
278
|
-
two writes is acceptable — if the run insert succeeds but the event insert crashes, the
|
|
279
|
-
run exists and `run_started` will still proceed normally (the event log will be missing a
|
|
280
|
-
`run_created` entry, but the run itself is functional).
|
|
281
|
-
|
|
282
|
-
### 14. Missing `WorkflowRunStatus` span attribute after parallel refactor
|
|
283
|
-
|
|
284
|
-
The `start()` span previously set `Attribute.WorkflowRunStatus(result.run.status)`, but
|
|
285
|
-
this was dropped in the parallel refactor because `result.run` is only available when
|
|
286
|
-
`runCreatedResult` fulfilled. The attribute is now conditionally set when the result is
|
|
287
|
-
available. In the resilient start case (run_created failed), the attribute is omitted
|
|
288
|
-
rather than erroring.
|
|
289
|
-
|
|
290
|
-
### 15. `run_started` eventData leak in world-postgres result
|
|
291
|
-
|
|
292
|
-
The `...data` spread in the result construction leaked `eventData` from `run_started`
|
|
293
|
-
into the returned event object. Storage was already correct (`storedEventData` is
|
|
294
|
-
`undefined` for `run_started`), but the returned result carried the input data. While
|
|
295
|
-
harmless (the runtime doesn't use `result.event.eventData`), it was restored to match
|
|
296
|
-
the pre-refactor behavior where eventData was explicitly stripped from the result.
|
|
297
|
-
|
|
298
|
-
### 16. Normal `run_created` path also needed `writeExclusive` (Windows CI)
|
|
299
|
-
|
|
300
|
-
The initial TOCTOU fix (item 12) only changed the resilient start path to use
|
|
301
|
-
`writeExclusive`. The normal `run_created` entity write still used `writeJSON` which
|
|
302
|
-
checks existence with `fs.access()` then writes via temp+rename — not atomic. On
|
|
303
|
-
Windows CI, the local queue's async IIFE delivered fast enough for both paths to pass
|
|
304
|
-
their existence checks simultaneously, producing two `run_created` events with different
|
|
305
|
-
event IDs. The events consumer saw the duplicate as "Unconsumed event in event log,"
|
|
306
|
-
causing `hooks`, `supports null bytes in step results`, and `retriable and fatal errors`
|
|
307
|
-
tests to time out at 60s. Fixed by also switching the normal `run_created` entity write to
|
|
308
|
-
`writeExclusive`, making both paths use the same atomic gate.
|
|
309
|
-
|
|
310
|
-
### 17. CborTransport was a pass-through wrapper
|
|
311
|
-
|
|
312
|
-
`world-vercel/queue.ts` had `CborTransport` implementing `Transport<Buffer>` with a
|
|
313
|
-
no-op `serialize` (identity function) and a `deserialize` that reassembled chunks into
|
|
314
|
-
a Buffer without decoding. The actual CBOR `encode()`/`decode()` calls happened at the
|
|
315
|
-
call sites — `queue()` pre-encoded before calling `client.send()`, and the handler
|
|
316
|
-
post-decoded after receiving from `client.handleCallback()`. This violated the transport
|
|
317
|
-
abstraction (every other transport does its encoding inside serialize/deserialize) and
|
|
318
|
-
meant the call site had to remember to pre-encode. Refactored to move `encode()`/`decode()`
|
|
319
|
-
into the transport methods and changed the type from `Transport<Buffer>` to
|
|
320
|
-
`Transport<unknown>`.
|
|
321
|
-
|
|
322
|
-
## Follow-up work (additional)
|
|
323
|
-
|
|
324
|
-
- [x] ~~**CborTransport is a pass-through**~~ — Resolved. Moved `encode()`/`decode()`
|
|
325
|
-
into `CborTransport.serialize()`/`CborTransport.deserialize()`. The transport is now
|
|
326
|
-
self-contained: call sites pass plain objects, and the handler receives decoded objects.
|
|
327
|
-
See retrospective item 17.
|
|
75
|
+
The narrow crash window in `world-postgres` between the run insert and the event insert is acceptable — if the run insert succeeds but the event insert crashes, the run exists and `run_started` will still proceed normally (the event log will be missing a `run_created` entry, but the run itself is functional).
|
|
@@ -9,21 +9,21 @@ related:
|
|
|
9
9
|
- /docs/foundations/errors-and-retries
|
|
10
10
|
---
|
|
11
11
|
|
|
12
|
-
This error occurs when the Workflow runtime
|
|
12
|
+
This error occurs when the Workflow runtime repeatedly cannot replay events in the event log. This usually means the event log is in an invalid state, such as duplicate or orphaned events, or that a runtime determinism bug persists across retry attempts.
|
|
13
13
|
|
|
14
|
-
This is a **workflow-level fatal error**. It cannot be caught or handled inside your workflow code.
|
|
14
|
+
This is a **workflow-level fatal error**. It cannot be caught or handled inside your workflow code. The runtime first retries transient replay divergence automatically; it marks the run as failed with this error only after replay still cannot recover.
|
|
15
15
|
|
|
16
16
|
## Error Message
|
|
17
17
|
|
|
18
18
|
```
|
|
19
|
-
|
|
19
|
+
Workflow replay diverged <divergenceCount> times after <maxRecoveryReplays> recovery replays; latest divergent event was <eventId>. Last divergence: <details>
|
|
20
20
|
```
|
|
21
21
|
|
|
22
22
|
## Why This Happens
|
|
23
23
|
|
|
24
24
|
Workflows persist their progress as an ordered event log. During replay, the runtime processes each event in sequence — every event must be consumed by a matching callback (e.g., a step or sleep waiting for its result). When an event has no matching consumer, the runtime cannot advance past it, which would block all subsequent events and hang the workflow indefinitely.
|
|
25
25
|
|
|
26
|
-
Instead of silently hanging, the runtime
|
|
26
|
+
Instead of silently hanging, the runtime retries a divergent replay before failing the workflow and surfacing this terminal error.
|
|
27
27
|
|
|
28
28
|
Common scenarios that produce this error:
|
|
29
29
|
|
|
@@ -45,7 +45,7 @@ npm install workflow@latest
|
|
|
45
45
|
|
|
46
46
|
### 2. Retry the failed run
|
|
47
47
|
|
|
48
|
-
|
|
48
|
+
If this error is displayed, automatic replay recovery has already been exhausted and the run has been marked as `failed`. You can re-run it using the **Re-run** button in the Workflow Dashboard.
|
|
49
49
|
|
|
50
50
|
### 3. Report the issue
|
|
51
51
|
|
package/docs/errors/index.mdx
CHANGED
|
@@ -37,6 +37,9 @@ Fix common mistakes when creating and executing workflows in the **Workflow SDK*
|
|
|
37
37
|
<Card href="/docs/errors/corrupted-event-log" title="corrupted-event-log">
|
|
38
38
|
Learn how to handle corrupted or invalid event logs.
|
|
39
39
|
</Card>
|
|
40
|
+
<Card href="/docs/errors/replay-divergence" title="replay-divergence">
|
|
41
|
+
Learn how workflow replay divergence is recovered automatically.
|
|
42
|
+
</Card>
|
|
40
43
|
<Card href="/docs/errors/step-not-registered" title="step-not-registered">
|
|
41
44
|
Resolve step not registered errors caused by deployment mismatches.
|
|
42
45
|
</Card>
|
|
@@ -0,0 +1,27 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: replay-divergence
|
|
3
|
+
description: A workflow replay temporarily followed a path that did not match its recorded events.
|
|
4
|
+
type: troubleshooting
|
|
5
|
+
summary: Understand automatic recovery when a workflow replay diverges from its event history.
|
|
6
|
+
prerequisites:
|
|
7
|
+
- /docs/foundations/workflows-and-steps
|
|
8
|
+
related:
|
|
9
|
+
- /docs/errors/corrupted-event-log
|
|
10
|
+
- /docs/foundations/errors-and-retries
|
|
11
|
+
---
|
|
12
|
+
|
|
13
|
+
A replay divergence occurs when one invocation of a workflow cannot consume the durable event history using the promises, hooks, sleeps, or steps it created during replay.
|
|
14
|
+
|
|
15
|
+
This is an SDK/runtime signal, not an error thrown by your workflow code. It is not catchable inside a workflow function.
|
|
16
|
+
|
|
17
|
+
## Automatic Recovery
|
|
18
|
+
|
|
19
|
+
A single divergent replay does not prove that persisted history is corrupted. For example, asynchronous delivery ordering may cause one invocation to follow the wrong side of a race while another replay can follow the recorded history correctly.
|
|
20
|
+
|
|
21
|
+
The runtime automatically queues another replay when an invocation reports `REPLAY_DIVERGENCE`. No terminal `run_failed` event is written during these recovery attempts.
|
|
22
|
+
|
|
23
|
+
If recovery replays continue to diverge after the retry budget is exhausted, the runtime marks the run as failed with `CORRUPTED_EVENT_LOG` and records the latest divergent event for diagnosis.
|
|
24
|
+
|
|
25
|
+
## What To Do
|
|
26
|
+
|
|
27
|
+
Most replay divergence signals recover without action. If a run ultimately fails with `CORRUPTED_EVENT_LOG`, update to the latest `workflow` package and report the run ID and error details if the failure persists.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "workflow",
|
|
3
|
-
"version": "5.0.0-beta.
|
|
3
|
+
"version": "5.0.0-beta.12",
|
|
4
4
|
"description": "Workflow SDK - Build durable, resilient, and observable workflows",
|
|
5
5
|
"main": "dist/typescript-plugin.cjs",
|
|
6
6
|
"type": "module",
|
|
@@ -57,18 +57,18 @@
|
|
|
57
57
|
},
|
|
58
58
|
"dependencies": {
|
|
59
59
|
"ms": "2.1.3",
|
|
60
|
-
"@workflow/astro": "5.0.0-beta.
|
|
61
|
-
"@workflow/cli": "5.0.0-beta.
|
|
62
|
-
"@workflow/core": "5.0.0-beta.
|
|
63
|
-
"@workflow/errors": "5.0.0-beta.
|
|
60
|
+
"@workflow/astro": "5.0.0-beta.12",
|
|
61
|
+
"@workflow/cli": "5.0.0-beta.12",
|
|
62
|
+
"@workflow/core": "5.0.0-beta.12",
|
|
63
|
+
"@workflow/errors": "5.0.0-beta.7",
|
|
64
64
|
"@workflow/typescript-plugin": "5.0.0-beta.4",
|
|
65
65
|
"@workflow/utils": "5.0.0-beta.3",
|
|
66
|
-
"@workflow/next": "5.0.0-beta.
|
|
67
|
-
"@workflow/nest": "5.0.0-beta.
|
|
68
|
-
"@workflow/nitro": "5.0.0-beta.
|
|
69
|
-
"@workflow/nuxt": "5.0.0-beta.
|
|
70
|
-
"@workflow/sveltekit": "5.0.0-beta.
|
|
71
|
-
"@workflow/rollup": "5.0.0-beta.
|
|
66
|
+
"@workflow/next": "5.0.0-beta.12",
|
|
67
|
+
"@workflow/nest": "5.0.0-beta.12",
|
|
68
|
+
"@workflow/nitro": "5.0.0-beta.12",
|
|
69
|
+
"@workflow/nuxt": "5.0.0-beta.12",
|
|
70
|
+
"@workflow/sveltekit": "5.0.0-beta.12",
|
|
71
|
+
"@workflow/rollup": "5.0.0-beta.12"
|
|
72
72
|
},
|
|
73
73
|
"devDependencies": {
|
|
74
74
|
"@types/ms": "2.1.0",
|