workflow 5.0.0-beta.11 → 5.0.0-beta.13
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/docs/ai/index.mdx +21 -18
- package/docs/api-reference/workflow-ai/durable-agent.mdx +7 -45
- package/docs/changelog/eager-processing.mdx +67 -413
- package/docs/changelog/resilient-start.mdx +31 -283
- package/docs/cookbook/agent-patterns/durable-agent.mdx +10 -146
- package/docs/cookbook/index.mdx +1 -1
- package/docs/errors/corrupted-event-log.mdx +5 -5
- package/docs/errors/index.mdx +3 -0
- package/docs/errors/replay-divergence.mdx +27 -0
- package/docs/foundations/streaming.mdx +13 -22
- package/docs/internal/index.mdx +2 -0
- package/docs/internal/meta.json +6 -1
- package/docs/internal/nitro-native-build.mdx +38 -0
- package/docs/internal/nitro-web-ui.mdx +24 -0
- package/package.json +11 -11
|
@@ -8,12 +8,11 @@ type: overview
|
|
|
8
8
|
|
|
9
9
|
**Date**: March 2026
|
|
10
10
|
|
|
11
|
-
This is a major internal architecture change to how Workflow DevKit executes workflows and steps
|
|
11
|
+
This is a major internal architecture change to how Workflow DevKit executes workflows and steps. It reduces function invocations and queue overhead by executing steps _inline_ within the same function invocation as the workflow replay, rather than dispatching every step to a separate function via the queue.
|
|
12
12
|
|
|
13
13
|
## Previous Architecture
|
|
14
14
|
|
|
15
|
-
The previous architecture used two separate routes,
|
|
16
|
-
each backed by its own queue trigger:
|
|
15
|
+
The previous architecture used two separate routes, each backed by its own queue trigger:
|
|
17
16
|
|
|
18
17
|
```
|
|
19
18
|
Queue: __wkf_workflow_* --> /.well-known/workflow/v1/flow (workflow replay in VM)
|
|
@@ -69,118 +68,64 @@ suspension with pending operations
|
|
|
69
68
|
|
|
70
69
|
A serial workflow with 10 steps now completes in **1 function invocation**.
|
|
71
70
|
|
|
72
|
-
## Inline Step Execution
|
|
73
|
-
|
|
74
|
-
After the workflow suspends with pending steps, the handler executes one step inline:
|
|
75
|
-
|
|
76
|
-
1. Create `step_started` event
|
|
77
|
-
2. Hydrate step input from the event log
|
|
78
|
-
3. Look up the step function via `getStepFunction(stepName)`
|
|
79
|
-
4. Execute the step function
|
|
80
|
-
5. Create `step_completed` or `step_failed` event
|
|
81
|
-
6. Loop back to workflow replay
|
|
82
|
-
|
|
83
|
-
This logic lives in `executeStep()` in `packages/core/src/runtime/step-executor.ts`.
|
|
84
|
-
|
|
85
71
|
## Background Steps (Parallel Execution)
|
|
86
72
|
|
|
87
|
-
When a workflow suspends with multiple pending steps (e.g., from `Promise.all`), the handler
|
|
88
|
-
|
|
89
|
-
1. Creates `step_created` events for all pending steps
|
|
90
|
-
2. Queues N-1 steps back to `__wkf_workflow_*` with `stepId` and `stepName` in the message payload
|
|
91
|
-
3. Executes 1 step inline
|
|
92
|
-
4. Loops back to replay
|
|
73
|
+
When a workflow suspends with multiple pending steps (e.g., from `Promise.all`), the handler creates `step_created` events for all of them, queues N-1 back to `__wkf_workflow_*` with `stepId` and `stepName` in the message payload, executes 1 step inline, and loops back to replay.
|
|
93
74
|
|
|
94
|
-
Each background step message is handled by a separate function invocation of the same handler. When a message arrives with `stepId` and `stepName`, the handler executes that specific step, then checks if all parallel steps from the batch are done by
|
|
75
|
+
Each background step message is handled by a separate function invocation of the same handler. When a message arrives with `stepId` and `stepName`, the handler executes that specific step, then checks if all parallel steps from the batch are done by comparing `step_created` events against terminal events (`step_completed`/`step_failed`):
|
|
95
76
|
|
|
96
|
-
- **All steps done**: The handler replays the workflow inline, continuing the execution loop without a queue roundtrip.
|
|
77
|
+
- **All steps done**: The handler replays the workflow inline, continuing the execution loop without a queue roundtrip.
|
|
97
78
|
- **Steps still pending**: The handler returns without queuing a continuation. The last handler to complete its step will see all steps done and replay inline.
|
|
98
|
-
- **Pending ops (stream writes)**: The handler queues a continuation and returns, so `waitUntil` can flush the pending stream data
|
|
79
|
+
- **Pending ops (stream writes)**: The handler queues a continuation and returns, so `waitUntil` can flush the pending stream data.
|
|
99
80
|
|
|
100
81
|
### Convergence After Parallel Steps
|
|
101
82
|
|
|
102
|
-
When multiple background steps complete near-simultaneously, multiple handlers may observe "all steps done" and attempt to advance the workflow concurrently. The event-sourced architecture plus the invariants
|
|
83
|
+
When multiple background steps complete near-simultaneously, multiple handlers may observe "all steps done" and attempt to advance the workflow concurrently. The event-sourced architecture plus the invariants below ensure safe convergence:
|
|
103
84
|
|
|
104
|
-
- **`step_created` idempotency**
|
|
105
|
-
- **`step_completed` / `step_failed` idempotency**
|
|
106
|
-
- **Queue idempotency keys**
|
|
107
|
-
- **Deterministic replay**
|
|
85
|
+
- **`step_created` idempotency** — duplicate creates return 409; exactly one handler owns each step
|
|
86
|
+
- **`step_completed` / `step_failed` idempotency** — only the first invocation to record a terminal result wins
|
|
87
|
+
- **Queue idempotency keys** — background step messages use `correlationId` as idempotency key
|
|
88
|
+
- **Deterministic replay** — all invocations produce the same result given the same event log
|
|
108
89
|
|
|
109
90
|
### Single Inline Executor Per Step
|
|
110
91
|
|
|
111
|
-
Inline step execution combined with background-step dispatch introduces a new coordination requirement
|
|
92
|
+
Inline step execution combined with background-step dispatch introduces a new coordination requirement: when multiple handlers reach the same `Promise.all` batch concurrently, we need to guarantee that each step body runs at most once via the inline path. Without that guarantee, the event log accumulates duplicate `step_started` events (including some written *after* `step_completed`, which orphans them on replay) and step bodies run redundantly.
|
|
112
93
|
|
|
113
94
|
The design enforces a simple invariant: **exactly one handler owns each step, and only the owner may execute it inline**. Ownership is established by the atomicity of `step_created`:
|
|
114
95
|
|
|
115
|
-
1. **Atomic `step_created`**
|
|
116
|
-
2. **Suspension handler reports ownership**
|
|
117
|
-
3. **Inline execution is gated on ownership**
|
|
118
|
-
4. **Queueing is unconditional**
|
|
96
|
+
1. **Atomic `step_created`** — `events.create('step_created', correlationId=X)` is serialized per-correlationId in every world. Exactly one concurrent caller succeeds; the rest receive `EntityConflictError`.
|
|
97
|
+
2. **Suspension handler reports ownership** — only `step_created` writes that actually succeeded (not those that caught 409) count toward ownership.
|
|
98
|
+
3. **Inline execution is gated on ownership** — a handler that didn't win any `step_created` race performs no inline execution.
|
|
99
|
+
4. **Queueing is unconditional** — for every pending step except the one being inline-executed, the handler enqueues a background step message with `idempotencyKey: correlationId`. This is what makes crash recovery work: if a prior handler wrote `step_created` but crashed before enqueueing, a later handler will enqueue the orphaned step. Concurrent handlers' redundant enqueues dedupe on the idempotency key.
|
|
119
100
|
|
|
120
|
-
Together these give: every `step_created` event has exactly one inline executor
|
|
101
|
+
Together these give: every `step_created` event has exactly one inline executor **and** at least one queued dispatch. Step bodies are never executed concurrently, and `step_started` events never land in the log after `step_completed` for the same step.
|
|
121
102
|
|
|
122
|
-
**Retry semantics are preserved**:
|
|
103
|
+
**Retry semantics are preserved**: a step that is currently running (status=`running`) still accepts a second `step_started` write with an incremented attempt counter — this is how queue redelivery after a SIGKILL mid-execution legitimately re-runs the step.
|
|
123
104
|
|
|
124
105
|
## Incremental Event Loading
|
|
125
106
|
|
|
126
107
|
The handler caches the event log in memory across loop iterations. Instead of re-fetching the entire event log on each replay:
|
|
127
108
|
|
|
128
|
-
1. **First iteration**: full load
|
|
129
|
-
2. **Subsequent iterations**:
|
|
109
|
+
1. **First iteration**: full load, returning both the events and the final pagination cursor
|
|
110
|
+
2. **Subsequent iterations**: fetch only events created after the saved cursor and append them to the cached array
|
|
130
111
|
|
|
131
112
|
For a 10-step serial workflow completing in one invocation, the 10th replay loads ~2 new events instead of re-fetching all ~30.
|
|
132
113
|
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
The incremental loading depends on the server returning a cursor even on the final page of results (`hasMore: false`). Previously, `workflow-server` returned `cursor: null` when there were no more pages. This was fixed in the `peter/fix-end-cursor` branch to always return an `eid:<eventId>` cursor when there are events, aligning with `world-local` and `world-postgres` behavior.
|
|
136
|
-
|
|
137
|
-
If a World implementation does not return a cursor after the initial load, the handler logs an error and falls back to a full reload.
|
|
114
|
+
Incremental loading depends on the World returning a cursor even on the final page of results. If a World implementation does not return a cursor after the initial load, the handler logs an error and falls back to a full reload.
|
|
138
115
|
|
|
139
116
|
## Timeout Handling
|
|
140
117
|
|
|
141
|
-
The inline execution loop checks wall-clock time before each replay iteration. If the elapsed time exceeds a configurable threshold (default: 110 seconds, for a 120-second function limit), the handler re-schedules itself via the queue and returns.
|
|
142
|
-
|
|
143
|
-
The threshold is configurable via the `WORKFLOW_V2_TIMEOUT_MS` environment variable.
|
|
118
|
+
The inline execution loop checks wall-clock time before each replay iteration. If the elapsed time exceeds a configurable threshold (default: 110 seconds, for a 120-second function limit), the handler re-schedules itself via the queue and returns. Configurable via `WORKFLOW_V2_TIMEOUT_MS`.
|
|
144
119
|
|
|
145
120
|
If a single step takes longer than the timeout threshold, the step runs to completion (or SIGKILL) — there is no interruption mechanism for in-progress step execution. This is the same behavior as the previous architecture.
|
|
146
121
|
|
|
147
122
|
## Queue Message Changes
|
|
148
123
|
|
|
149
|
-
The `WorkflowInvokePayload` schema has two new optional fields:
|
|
150
|
-
|
|
151
|
-
{/*@skip-typecheck - snippet, not runnable code*/}
|
|
152
|
-
|
|
153
|
-
```typescript
|
|
154
|
-
stepId: z.string().optional()
|
|
155
|
-
stepName: z.string().optional()
|
|
156
|
-
```
|
|
157
|
-
|
|
158
|
-
When `stepId` is present, the handler executes that specific step before (or instead of) replaying the workflow. Background steps are queued with both `stepId` and `stepName` set, so the handler knows which step function to call without loading the event log. Previously, `stepName` was resolved by loading all events and searching for the `step_created` event matching the `stepId` — an O(N) operation on the full event history for every background step arrival.
|
|
124
|
+
The `WorkflowInvokePayload` schema has two new optional fields: `stepId` and `stepName`. When `stepId` is present, the handler executes that specific step before (or instead of) replaying the workflow. Background steps are queued with both set, so the handler knows which step function to call without loading the event log. Previously, `stepName` was resolved by loading all events and searching for the `step_created` event matching the `stepId` — an O(N) operation on the full event history for every background step arrival.
|
|
159
125
|
|
|
160
126
|
The queue trigger configuration uses `WORKFLOW_QUEUE_TRIGGER` on the `__wkf_workflow_*` topic. The `__wkf_step_*` topic and its separate trigger are no longer generated.
|
|
161
127
|
|
|
162
|
-
##
|
|
163
|
-
|
|
164
|
-
### Base Builder
|
|
165
|
-
|
|
166
|
-
New method `createCombinedBundle()` in `packages/builders/src/base-builder.ts`:
|
|
167
|
-
|
|
168
|
-
1. Builds the step registrations bundle (same esbuild + SWC step mode as before)
|
|
169
|
-
2. Builds the workflow VM code string (same esbuild + SWC workflow mode as before)
|
|
170
|
-
3. Generates a combined route file that imports the step registrations and uses `workflowEntrypoint(workflowCode)`
|
|
171
|
-
|
|
172
|
-
No changes to the SWC plugin were needed. The two-pass build approach (separate step and workflow SWC modes) still applies.
|
|
173
|
-
|
|
174
|
-
### Framework Builders
|
|
175
|
-
|
|
176
|
-
All framework builders were updated to use `createCombinedBundle()`:
|
|
177
|
-
|
|
178
|
-
- **Next.js** (eager and deferred/lazyDiscovery): replaces separate step + flow route generation
|
|
179
|
-
- **NestJS, Nitro, Standalone**: replaces separate `createStepsBundle()` + `createWorkflowsBundle()` calls
|
|
180
|
-
- **SvelteKit, Astro**: same, plus post-processing regex updated to match `workflowEntrypoint`
|
|
181
|
-
- **Vercel Build Output API** (used by Nitro/Astro production): single `flow.func/` with `WORKFLOW_QUEUE_TRIGGER`
|
|
182
|
-
|
|
183
|
-
### Generated File Layout
|
|
128
|
+
## Generated File Layout
|
|
184
129
|
|
|
185
130
|
```
|
|
186
131
|
.well-known/workflow/v1/
|
|
@@ -196,37 +141,19 @@ All framework builders were updated to use `createCombinedBundle()`:
|
|
|
196
141
|
|
|
197
142
|
The `step/` directory is no longer generated.
|
|
198
143
|
|
|
199
|
-
##
|
|
200
|
-
|
|
201
|
-
`handleSuspension()` in `packages/core/src/runtime/suspension-handler.ts` creates events for all pending operations (hooks, step events, wait events) but does **not** queue step messages. It returns the pending step items so the handler can decide which to execute inline vs. queue to background.
|
|
202
|
-
|
|
203
|
-
## Concerns and Edge Cases
|
|
144
|
+
## Design Notes and Tradeoffs
|
|
204
145
|
|
|
205
146
|
### Parent→Child Polling Holds Worker Slots
|
|
206
147
|
|
|
207
148
|
`Run#returnValue` is implemented as a polling step: the workflow awaits the child run's terminal status inside a step body. In worker-based worlds (notably `world-postgres`), each such poll occupies a queue worker slot until the child run finishes. Parent workflows that fan out to many child runs — recursive workflows like `fibonacciWorkflow` are the obvious case — can therefore consume a large fraction of available workers just holding positions in `Promise.all([...children.map(c => c.returnValue)])`.
|
|
208
149
|
|
|
209
|
-
If `queueConcurrency` is smaller than the peak number of concurrent parent polls plus the workers needed for any in-flight children, the system deadlocks: every slot is held by a parent waiting on a child, but no child can acquire a slot to start.
|
|
150
|
+
If `queueConcurrency` is smaller than the peak number of concurrent parent polls plus the workers needed for any in-flight children, the system deadlocks: every slot is held by a parent waiting on a child, but no child can acquire a slot to start.
|
|
210
151
|
|
|
211
|
-
For `world-postgres`, the default `queueConcurrency` is set to **50
|
|
152
|
+
For `world-postgres`, the default `queueConcurrency` is set to **50**. Workflows that fan out more aggressively must raise this ceiling.
|
|
212
153
|
|
|
213
|
-
|
|
154
|
+
To prevent deadlock when polling is executed inline by the step executor, `Run#pollReturnValue()` detects when it's running inside a step executor and throws `TooEarlyError` instead of polling in a blocking loop. The step executor handles `TooEarlyError` by re-queueing the step with a 1-second delay, freeing the worker. Unlike `RetryableError`, `TooEarlyError` does NOT count against `maxRetries`, so polling steps can retry indefinitely until the child completes.
|
|
214
155
|
|
|
215
|
-
|
|
216
|
-
|
|
217
|
-
Workflow code still runs in a Node.js VM for determinism and sandboxing. Step code runs in the Node.js host context. The only change is that both happen within the same function invocation.
|
|
218
|
-
|
|
219
|
-
### Bundle Size and Cold Start
|
|
220
|
-
|
|
221
|
-
The combined bundle is larger (contains both step code and workflow VM code). Cold start time increases slightly. The reduction in total function invocations more than compensates.
|
|
222
|
-
|
|
223
|
-
### Step Retries
|
|
224
|
-
|
|
225
|
-
When an inline step fails with retries remaining:
|
|
226
|
-
|
|
227
|
-
- `RetryableError` with explicit `retryAfter` delay: re-queue to self with `stepId` and delay
|
|
228
|
-
- Transient errors with immediate retry: re-queue to self with `stepId` (delay = 1s)
|
|
229
|
-
- `FatalError`: fail immediately
|
|
156
|
+
**Follow-up**: Replace the worker-pool sizing requirement with a polling design that does not occupy a worker slot. Options under consideration: moving child-completion polling out of the step body into the suspension layer, or emitting a `run_completed` notification on the parent's stream/queue so the parent only resumes when the child actually finishes.
|
|
230
157
|
|
|
231
158
|
### Mixed Suspensions
|
|
232
159
|
|
|
@@ -236,360 +163,87 @@ A suspension may contain steps, hooks, and waits simultaneously. The handler cre
|
|
|
236
163
|
- **Steps + at least one wait**: every step is queued (no inline execution). The handler returns with the wait timeout. Whichever lands first — a step's continuation or the wait timer — drives the next replay.
|
|
237
164
|
- **Hooks / waits only**: handler returns with the wait timeout (or no timeout, for hook-only suspensions). The next continuation is driven by external resume or the wait timer.
|
|
238
165
|
|
|
239
|
-
The "no inline when there's a wait" carve-out is necessary to preserve `Promise.race(step, sleep)` semantics. Inline `await executeStep(...)` blocks the handler for the full step duration, and `wait_completed` events are only created on the *next* loop iteration's "complete elapsed waits" pass — so a longer-running step would always swallow the shorter sleep and `Promise.race` would resolve incorrectly. Queueing the step in this case lets the wait timer drive a continuation in parallel
|
|
166
|
+
The "no inline when there's a wait" carve-out is necessary to preserve `Promise.race(step, sleep)` semantics. Inline `await executeStep(...)` blocks the handler for the full step duration, and `wait_completed` events are only created on the *next* loop iteration's "complete elapsed waits" pass — so a longer-running step would always swallow the shorter sleep and `Promise.race` would resolve incorrectly. Queueing the step in this case lets the wait timer drive a continuation in parallel.
|
|
240
167
|
|
|
241
168
|
Pure step suspensions (without waits) still benefit from inline execution; the carve-out only costs an extra queue roundtrip when a step and a sleep coexist.
|
|
242
169
|
|
|
243
|
-
###
|
|
244
|
-
|
|
245
|
-
If a hook conflict is detected during suspension handling, the handler breaks the loop and returns `{ timeoutSeconds: 0 }` for immediate re-invocation, same as the previous behavior.
|
|
246
|
-
|
|
247
|
-
### Encryption Key Resolution
|
|
248
|
-
|
|
249
|
-
Encryption keys are resolved once before the inline execution loop starts (after the run status is confirmed as `running`) and reused across all iterations. Background step executions resolve the key independently. The key does not change within a run.
|
|
250
|
-
|
|
251
|
-
## Framework Support
|
|
252
|
-
|
|
253
|
-
All framework integrations have been updated: Next.js (eager and deferred/lazyDiscovery), NestJS, SvelteKit, Astro, Nitro/Nuxt/Hono/Express/Vite, and CLI standalone. The Vercel Build Output API builder (used by Nitro and Astro for production deploys) also uses the combined bundle with `WORKFLOW_QUEUE_TRIGGER`.
|
|
254
|
-
|
|
255
|
-
## Non-Next.js Integration Challenges
|
|
256
|
-
|
|
257
|
-
### Module Scope Duplication in Re-Bundled Output
|
|
258
|
-
|
|
259
|
-
Builders that use `bundleFinalOutput: true` (standalone CLI, Vercel Build Output API, NestJS) produce a single file where esbuild re-bundles the step registrations and the workflow runtime together. esbuild creates isolated module scopes for each source module, even within the same output file. This meant `registerStepFunction` and `getStepFunction` operated on different `Map` instances — steps were registered into one Map but looked up from another.
|
|
260
|
-
|
|
261
|
-
**Fix**: The step function registry (`registeredSteps` Map in `@workflow/core/private`) and the step context storage (`contextStorage` AsyncLocalStorage in `@workflow/core/step/context-storage`) were changed from module-scoped variables to `globalThis` singletons using `Symbol.for`. This ensures all esbuild module scopes share the same instances. The pattern was already used in the codebase for the World singleton and the class serialization registry.
|
|
262
|
-
|
|
263
|
-
### Workflow Package CJS Export Condition
|
|
264
|
-
|
|
265
|
-
The `workflow` package's root export has `"require": "./dist/typescript-plugin.cjs"` for TypeScript editor plugin loading. When esbuild bundles with CJS format, it resolves `import { defineHook } from 'workflow'` via the `require` condition, getting the TS plugin instead of the API.
|
|
266
|
-
|
|
267
|
-
**Fix**: Added a `"node"` condition (`"node": "./dist/index.js"`) before the `"require"` condition in the workflow package's exports. esbuild with `conditions: ['node']` matches `"node"` first and uses the correct API entry. TypeScript's plugin loader doesn't use `conditions: ['node']`, so it still falls through to `"require"` for the TS plugin.
|
|
268
|
-
|
|
269
|
-
### Local World Concurrent Replay Interference
|
|
270
|
-
|
|
271
|
-
The local development world (`world-local`) processes queue messages with high concurrency (default: 1000). With the V2 combined handler, parallel steps generate multiple workflow continuation messages. When these are processed concurrently, each triggers a replay that sees in-flight events from other concurrent replays. This causes "unconsumed event" errors because the event consumer encounters events that don't match any subscriber in the current replay state.
|
|
272
|
-
|
|
273
|
-
In production (Vercel), this doesn't happen — each function invocation is isolated with its own event loading.
|
|
274
|
-
|
|
275
|
-
**Fix**: The `EventsConsumer`'s `onUnconsumedEvent` callback (see "Concurrent Replay Interference with Multi-Batch Workflows" below) handles the concurrent event visibility issue. The V2 inline replay optimization (where the last background step to complete replays inline instead of queuing) further reduces concurrent replays. Redundant step executions from concurrent handlers are harmless due to `step_completed` idempotency — only the first completion wins.
|
|
276
|
-
|
|
277
|
-
### ESM `bundleFinalOutput` and Dynamic Require Errors
|
|
278
|
-
|
|
279
|
-
When `bundleFinalOutput: true` is used with ESM format, esbuild bundles CJS dependencies (like `debug`) into the output. CJS `require()` calls are wrapped in esbuild's `__require` polyfill, which throws "Dynamic require of X is not supported" in ESM contexts where `require` is undefined. This affected all ESM-based framework builders (Nitro, NestJS, SvelteKit, Astro) that were switched to `bundleFinalOutput: true` during the V2 migration.
|
|
280
|
-
|
|
281
|
-
**Fix**: ESM builders use `bundleFinalOutput: false` with `externalizeNonSteps: true`, matching the pre-V2 behavior. The framework's own bundler (Vite, Rollup, Turbopack) handles dependency resolution. The standalone CLI and Vercel Build Output API builders use `bundleFinalOutput: true` with ESM output plus a `createRequire(import.meta.url)` banner (see "V2 Combined Bundle Switched from CJS to ESM" below) so CJS dependencies can still call `require()` for Node.js builtins.
|
|
282
|
-
|
|
283
|
-
### Rollup Tree-Shaking of Step Registrations
|
|
284
|
-
|
|
285
|
-
When `bundleFinalOutput: false` is used with Nitro's rollup pipeline, the step registrations bundle (`steps.mjs`) only contains side-effect code (`registerStepFunction` calls) with no exports. Rollup tree-shakes the entire module because it has no used exports, removing all step registrations from the production bundle. This causes "Step not found" errors at runtime.
|
|
286
|
-
|
|
287
|
-
**Fix**: The steps bundle now exports a sentinel value (`export const __steps_registered = true`), and the combined route file imports it (`import { __steps_registered } from './steps.mjs'`). This gives rollup a used binding to track, preventing it from dropping the module and its side effects.
|
|
288
|
-
|
|
289
|
-
### Concurrent Replay Interference with Multi-Batch Workflows (historic)
|
|
290
|
-
|
|
291
|
-
An earlier iteration of the V2 work hit "Unconsumed event in event log" errors when multiple concurrent handlers raced into the same batch boundary. The diagnosis at the time was that concurrent handlers could see events the current replay hadn't reached yet, and the mitigation was a skip path in `onUnconsumedEvent` that tolerated step/hook/wait lifecycle events whose correlationId had a matching `step_created` / `hook_created` / earlier `wait_completed` in the log.
|
|
292
|
-
|
|
293
|
-
Later work on the "Single Inline Executor Per Step" invariant (described above) identified the actual root cause: duplicate `step_started` events were being written *after* `step_completed` on the same step, because the local world's `step_started` was not atomic w.r.t. terminal state and the main loop was re-picking already-queued steps for inline execution. Fixing those at the source (per-step mutex in `world-local` + ownership-gated inline dispatch + unconditional queueing with idempotency keys) eliminated the unconsumed-step-event path entirely, and fixing the `wait_completed` cursor bug (the main loop manually pushed `wait_completed` events without advancing `eventsCursor`, so the next incremental fetch re-returned them as local-array duplicates) eliminated the wait case.
|
|
170
|
+
### VM Sandboxing
|
|
294
171
|
|
|
295
|
-
|
|
172
|
+
Workflow code still runs in a Node.js VM for determinism and sandboxing. Step code runs in the Node.js host context. The only change is that both happen within the same function invocation.
|
|
296
173
|
|
|
297
|
-
###
|
|
174
|
+
### Bundle Size and Cold Start
|
|
298
175
|
|
|
299
|
-
|
|
176
|
+
The combined bundle is larger (contains both step code and workflow VM code). Cold start time increases slightly. The reduction in total function invocations more than compensates.
|
|
300
177
|
|
|
301
|
-
|
|
178
|
+
### Step Retries
|
|
302
179
|
|
|
303
|
-
|
|
180
|
+
When an inline step fails with retries remaining:
|
|
304
181
|
|
|
305
|
-
|
|
182
|
+
- `RetryableError` with explicit `retryAfter` delay: re-queue to self with `stepId` and delay
|
|
183
|
+
- Transient errors with immediate retry: re-queue to self with `stepId` (delay = 1s)
|
|
184
|
+
- `FatalError`: fail immediately
|
|
306
185
|
|
|
307
|
-
|
|
186
|
+
### Encryption Key Resolution
|
|
308
187
|
|
|
309
|
-
|
|
188
|
+
Encryption keys are resolved once before the inline execution loop starts (after the run status is confirmed as `running`) and reused across all iterations. Background step executions resolve the key independently. The key does not change within a run.
|
|
310
189
|
|
|
311
|
-
|
|
190
|
+
### Module Scope Duplication in Re-Bundled Output
|
|
312
191
|
|
|
313
|
-
|
|
192
|
+
Builders that re-bundle the combined output into a single file (standalone CLI, Vercel Build Output API, NestJS) produce a layout where esbuild creates isolated module scopes for each source module, even within the same output file. Without intervention this means `registerStepFunction` and `getStepFunction` operate on different `Map` instances — steps are registered into one Map but looked up from another.
|
|
314
193
|
|
|
315
|
-
|
|
194
|
+
The step function registry and the step context storage are `globalThis` singletons (via `Symbol.for`) to ensure all module scopes share the same instances. The same pattern is used for the World singleton and the class serialization registry.
|
|
316
195
|
|
|
317
196
|
### Inline Step Execution with Pending Stream Operations
|
|
318
197
|
|
|
319
198
|
When a step's arguments or return value include serialized streams (e.g., `WritableStream` from `getWritable()`, or AI SDK streaming steps), the serialization layer creates background `flushablePipe` operations that pipe data to S3. These ops are tracked in an `ops` array and need to complete before the stream data is readable by external consumers.
|
|
320
199
|
|
|
321
|
-
In V1, each step ran in a separate function invocation. After the step completed, `waitUntil(ops)` kept the function alive to flush the ops.
|
|
322
|
-
|
|
323
|
-
**Current state**: `executeStep()` attempts a 500ms `Promise.race` between the ops settling and a timeout. If ops settle in time (data confirmed on server), it returns `hasPendingOps: false` and the V2 handler continues the inline loop. If ops don't settle in 500ms (e.g., `WritableStream` kept open across steps), it returns `hasPendingOps: true` and the V2 handler breaks the loop and queues a continuation so `waitUntil` can flush them.
|
|
324
|
-
|
|
325
|
-
**Earlier attempts that failed** (before the flush waiter fix below):
|
|
326
|
-
|
|
327
|
-
1. **500ms inline ops await without flush waiters** — The same 500ms race, but `WorkflowServerWritableStream` used a buffered 10ms flush timer: the `flushablePipe`'s `pendingOps` reached 0 when the buffered `write()` returned (instant), but the actual S3 HTTP write hadn't started yet. The ops appeared settled but data wasn't on S3. Multiple approaches to fix the timing (delaying `pollWritableLock`, closing the writable to trigger flush, adding a post-settle delay) all failed or caused other issues (deadlocks, premature stream closure).
|
|
328
|
-
|
|
329
|
-
2. **Root cause of the buffered write issue**: `WorkflowServerWritableStream.write()` buffers chunks and schedules a flush via `setTimeout(flush, 10ms)`. The `flushablePipe` calls `await writer.write(chunk)` which returns immediately (data buffered). `pendingOps--` fires before the 10ms timer. The `pollWritableLock` sees `pendingOps === 0` and resolves `state.promise`. The ops appear settled, but data is still in the buffer.
|
|
330
|
-
|
|
331
|
-
3. **Why this only affects Vercel Prod**: On local (world-local), stream writes go to the filesystem — effectively instant. On Vercel (world-vercel), writes go through HTTP to workflow-server → S3, adding 50-100ms latency. The buffered write returns instantly but the HTTP round-trip is deferred. When the V2 loop continues and the function eventually returns, `waitUntil` may not have enough time to flush.
|
|
332
|
-
|
|
333
|
-
**Follow-up**: The flush-waiter design described under "Buffered Stream Flush with Waiter Promises" below is the landed fix and resolves the buffered-write race. The remaining work is to shrink the 500ms inline-ops budget once we have confidence that the flush-waiter path settles deterministically across all worlds (today the budget is a defensive ceiling, not a tuned latency target), and to surface a stronger contract for "ops settled" — currently a 500ms timeout means "probably settled, give up and queue a continuation", which is correct but coarse. A signaled "ops drained" event from the world layer would let `executeStep()` proceed without the timeout in the common case, removing latency for streaming workflows whose ops settle in well under 500ms.
|
|
334
|
-
|
|
335
|
-
### CJS `module.exports` Collision in BOA Bundles (RESOLVED)
|
|
336
|
-
|
|
337
|
-
The Vercel Build Output API (BOA) builder creates a single CJS bundle via `createCombinedBundle` with `bundleFinalOutput: true`. The combined route file imports the steps bundle:
|
|
338
|
-
|
|
339
|
-
```js
|
|
340
|
-
import { __steps_registered } from './__step_registrations.js';
|
|
341
|
-
import { workflowEntrypoint } from 'workflow/runtime';
|
|
342
|
-
export const POST = workflowEntrypoint(workflowCode);
|
|
343
|
-
```
|
|
344
|
-
|
|
345
|
-
When esbuild re-bundles this into CJS, the steps bundle's code is inlined. If the steps bundle is also CJS format, it contains its own `module.exports = __toCommonJS(...)` at the top level. esbuild sometimes inlines CJS modules **without** a `__commonJS()` wrapper (the heuristic depends on the module's detected format). When unwrapped, the steps bundle's `module.exports` assignment executes at the top level and **overwrites** the combined route's `module.exports`, removing the `POST` handler export.
|
|
346
|
-
|
|
347
|
-
**Symptoms**: The Vercel deployment builds and starts successfully, but the `POST` handler is missing from the function's exports. Queue messages are delivered to the function but nothing processes them. All e2e tests hang indefinitely.
|
|
200
|
+
In V1, each step ran in a separate function invocation. After the step completed, `waitUntil(ops)` kept the function alive to flush the ops. In V2, the inline execution loop continues immediately after the step body returns — so we need to know whether to keep looping or break out and let `waitUntil` flush.
|
|
348
201
|
|
|
349
|
-
|
|
202
|
+
`executeStep()` attempts a 500ms `Promise.race` between the ops settling and a timeout. If ops settle in time (data confirmed on server), it returns `hasPendingOps: false` and the V2 handler continues the inline loop. If ops don't settle in 500ms (e.g., `WritableStream` kept open across steps), it returns `hasPendingOps: true` and the V2 handler breaks the loop and queues a continuation so `waitUntil` can flush them.
|
|
350
203
|
|
|
351
|
-
|
|
352
|
-
2. Found two `module.exports` assignments in the bundle: line ~45K (from the combined route, exporting `POST`) and line ~95K (from the inlined steps bundle, exporting `__steps_registered`). The second overwrites the first.
|
|
353
|
-
3. Compared with the standalone builder's bundle which had the same steps code wrapped in `__commonJS()` — esbuild's wrapper prevents the inner `module.exports` from leaking.
|
|
354
|
-
|
|
355
|
-
**Fix**: When `bundleFinalOutput` is true, build the steps bundle in **ESM format** regardless of the final output format. The final esbuild pass converts everything to CJS correctly. ESM steps don't have `module.exports`, so there's no collision. The combined route's `export const POST` becomes the sole `module.exports` entry.
|
|
356
|
-
|
|
357
|
-
### Step Error Source Maps on BOA Deployments
|
|
358
|
-
|
|
359
|
-
The V2 combined CJS bundle (`bundleFinalOutput: true`) loses original source file names during re-bundling. Error stack traces show `/var/task/index.js` instead of `99_e2e.ts`. The `hasStepSourceMaps()` utility was updated to return `false` for BOA-builder frameworks (Express, Fastify, Hono, Nitro, Nuxt, Vite, Astro, Example) on Vercel preview, aligning test expectations with the actual bundle behavior.
|
|
360
|
-
|
|
361
|
-
### CLI Health Check Port Mismatch
|
|
362
|
-
|
|
363
|
-
The CLI `health` command defaults to `http://localhost:3000` when `WORKFLOW_LOCAL_BASE_URL` is not set. Different frameworks use different ports (Astro: 4321, SvelteKit: 5173). The e2e test passed `WORKFLOW_LOCAL_BASE_URL` via the spawn env, but the CLI's `getEnvVars()` function had a fixed list of env vars that didn't include `WORKFLOW_LOCAL_BASE_URL`. The env var was set but never read.
|
|
364
|
-
|
|
365
|
-
**Fix**: Added `WORKFLOW_LOCAL_BASE_URL` to the CLI's `getEnvVars()` return object.
|
|
204
|
+
**Follow-up**: Shrink the 500ms inline-ops budget once we have confidence that the flush-waiter path settles deterministically across all worlds. A signaled "ops drained" event from the world layer would let `executeStep()` proceed without the timeout in the common case.
|
|
366
205
|
|
|
367
206
|
### Buffered Stream Flush with Waiter Promises
|
|
368
207
|
|
|
369
|
-
`WorkflowServerWritableStream` buffers writes and flushes via a 10ms `setTimeout` for batching.
|
|
208
|
+
`WorkflowServerWritableStream` buffers writes and flushes via a 10ms `setTimeout` for batching. Naively, `write()` could return immediately after buffering, but that would cause the `flushablePipe`'s `pendingOps` counter to reach 0 before data actually reached the server — the V2 inline loop would see ops as settled prematurely and produce data-loss races on every step with `WritableStream` serialization.
|
|
370
209
|
|
|
371
|
-
|
|
210
|
+
`write()` returns a promise that resolves only after the scheduled flush completes. Multiple writes within the 10ms window still share a single batched HTTP request (the batching optimization is preserved). Each write registers a `{resolve, reject}` pair in a `flushWaiters` array. When the `setTimeout` fires and `flush()` completes the HTTP round-trip, all waiters are resolved (or rejected on error). This makes `pendingOps` accurately reflect server-side data state while keeping network-efficient batching.
|
|
372
211
|
|
|
373
|
-
|
|
374
|
-
|
|
375
|
-
- **Steps where ops settle** (data on server, ~200ms after lock release + flush) → continue loop inline
|
|
376
|
-
- **Steps where ops don't settle** (WritableStream kept open across steps) → break loop
|
|
377
|
-
|
|
378
|
-
### Lock-Release Polling Interval Lowered to 10ms
|
|
212
|
+
### Lock-Release Polling Interval
|
|
379
213
|
|
|
380
214
|
`flushablePipe`'s `pollWritableLock` / `pollReadableLock` use `setInterval` to detect when a user releases their stream lock without closing the stream — the Web Streams API has no event for that state. The V2 step executor's `opsSettled` race waits for this poll to resolve after each writable-bearing step body returns, so the polling interval sits on the critical path of every streaming step.
|
|
381
215
|
|
|
382
|
-
The interval was
|
|
383
|
-
|
|
384
|
-
**Fix**: dropped the polling interval from 100ms to 10ms in `packages/core/src/flushable-stream.ts`. Per-step wait drops from ~50ms average to ~5ms (a 10× improvement, expected to scale linearly with the number of writable-bearing steps in a workflow). For `DurableAgent.chat` with one tool call (4 writable-bearing steps), this removes ~180ms from the streaming chat response's critical path. Per-tick work is just `writable.locked` plus a `getWriter()`/`releaseLock()` probe — both microsecond-scale, so 10× more ticks during a stream's lifetime is not measurable in practice.
|
|
385
|
-
|
|
386
|
-
**Follow-up**: Replace the polling entirely with an event-driven release signal — wrap the writable returned from the `WritableStream` reviver with a writer that fires on `releaseLock()` — bringing the wait to ~0ms. The 10ms polling interval is the cheap path that captures most of the available win without the structural change, but every writable-bearing step still pays a ~5ms tax that the event-driven design would eliminate. The structural change is also worth pursuing because it removes a source of timing drift between `world-local` (filesystem-instant) and `world-vercel` (HTTP-deferred) — both would see truly synchronous lock-release detection rather than periodic-poll detection.
|
|
387
|
-
|
|
388
|
-
### Event Consumer Skip Logic Was Too Broad For Wait Replays
|
|
389
|
-
|
|
390
|
-
The V2 handler needs some tolerance for out-of-order replay, especially around step events created by concurrent continuations. An early follow-up broadened that fallback to all wait lifecycle events too, so `onUnconsumedEvent` would skip `wait_created` and the first `wait_completed` whenever they matched a known wait. In the BOA-backed previews that broke `hookDisposeTestWorkflow`: once the first run disposed its hook and went into `sleep('5s')`, a replay could skip the live wait event before `sleep()` registered its subscriber, leaving the run stuck forever at `wait_created`.
|
|
391
|
-
|
|
392
|
-
**Fix**: Keep the step/hook replay tolerance, but narrow the wait fallback to the one case we actually need: duplicate `wait_completed` events that appear *after* an earlier completion for the same wait. The `hookDispose` e2e was also updated to poll for hook registration/disposal instead of relying on fixed 3-5 second sleeps, which made the Vercel preview timing less brittle.
|
|
393
|
-
|
|
394
|
-
### TooEarlyError Retry Delay in Step Executor
|
|
395
|
-
|
|
396
|
-
The `executeStep()` function handles `TooEarlyError` (thrown when a step's `retryAfter` timestamp hasn't been reached yet) by returning a `retry` result with a timeout. The original implementation used a stale access pattern `(err as any).meta?.retryAfter` copied from an older error shape. The `TooEarlyError` class (from `@workflow/errors`) has `retryAfter` as a direct property (number of seconds), not nested under `.meta`. The stale pattern always evaluated to `undefined`, falling back to a 1-second delay regardless of the server's actual retry-after value.
|
|
397
|
-
|
|
398
|
-
**Fix**: Changed to `err.retryAfter ?? 1`, matching the correct pattern used in `step-handler.ts`.
|
|
399
|
-
|
|
400
|
-
### Health Check Endpoint JSON Response
|
|
401
|
-
|
|
402
|
-
The `withHealthCheck()` wrapper in `helpers.ts` was updated (on main) to return a JSON response with `{ healthy, endpoint, specVersion, workflowCoreVersion }` instead of a plain text string. The V2 branch's e2e test still expected `Content-Type: text/plain` and a text body after merging main, causing the "health check endpoint (HTTP)" test to fail across all frameworks and environments.
|
|
403
|
-
|
|
404
|
-
**Fix**: Updated the e2e test to expect `Content-Type: application/json` and validate the JSON body structure, including a `specVersion >= SPEC_VERSION_CURRENT` range assertion.
|
|
405
|
-
|
|
406
|
-
### V2 Combined Bundle Switched from CJS to ESM
|
|
407
|
-
|
|
408
|
-
The V2 combined bundle was initially emitted as CJS by the standalone CLI and Vercel Build Output API builders, while `main` had already moved those outputs to ESM in [#1562](https://github.com/vercel/workflow/pull/1562). Staying on CJS meant `import.meta.url` was polyfilled (often producing the wrong path in re-bundled contexts), and the `world-testing` server had to import from `flow.js` via `createRequire` to force CJS semantics on what was really a CJS bundle.
|
|
409
|
-
|
|
410
|
-
**Fix**: Align V2 with `main`'s ESM defaults:
|
|
411
|
-
|
|
412
|
-
1. The BOA builder emits `__step_registrations.mjs` and `index.mjs`, writes `"type": "module"` in `package.json`, and sets `handler: "index.mjs"` in `.vc-config.json`.
|
|
413
|
-
2. The standalone builder no longer overrides `format`; it inherits the base builder's `'esm'` default.
|
|
414
|
-
3. The standalone config outputs `step.mjs` / `flow.mjs` instead of `.js`.
|
|
415
|
-
4. The `world-testing` server uses a native `import { POST } from '../.well-known/workflow/v1/flow.mjs'` instead of `createRequire`.
|
|
416
|
-
5. `createCombinedBundle`'s final esbuild pass (for `bundleFinalOutput: true`) now prepends the same `createRequire(import.meta.url)` banner used by the workflow/webhook bundles so CJS dependencies that call `require()` for Node.js builtins (for example the `events` module referenced by bundled libraries) still resolve at runtime.
|
|
417
|
-
6. To avoid a duplicate `__createRequire` declaration, the inner steps bundle that gets inlined by the final pass skips the banner — only the outer bundle emits it. This is threaded through via a new `skipEsmRequireBanner` option on `createStepsBundle`.
|
|
418
|
-
|
|
419
|
-
### World specVersion in Health Check Responses
|
|
420
|
-
|
|
421
|
-
The `getWorldHandlers()` return value was updated on main to include `specVersion` (the World's declared spec version). The V2 handler destructures this as `worldSpecVersion` and passes it to `handleHealthCheckMessage()` for inclusion in queue-based health check responses. This was merged alongside the V2 timeout configuration.
|
|
422
|
-
|
|
423
|
-
### Async World Singleton Drift After Merge
|
|
424
|
-
|
|
425
|
-
The later `main` merge changed `getWorld()` and `getWorldHandlers()` to be asynchronous promise-backed singletons, but the eager-processing branch still had synchronous call sites in the V2 runtime path. That left `packages/core/src/runtime.ts` and `packages/core/src/runtime/helpers.ts` trying to access `.events` on a `Promise<World>`, which failed typecheck immediately after the merge.
|
|
426
|
-
|
|
427
|
-
**Fix**: Rebases the V2 workflow entrypoint onto the async world API by lazily awaiting `getWorldHandlers()` when wiring the queue handler and awaiting `getWorld()` at the remaining runtime/helper call sites. This preserves the inline replay loop while matching `main`'s new world initialization contract.
|
|
428
|
-
|
|
429
|
-
### Lazy World Loading for Next.js Production Builds
|
|
430
|
-
|
|
431
|
-
After the async world merge, `packages/core/src/runtime/world.ts` still eagerly imported both `@workflow/world-local` and `@workflow/world-vercel`, and it initialized `createRequire()` from `process.cwd() + '/package.json'` at module load time. In the Next.js production build jobs that caused the generated flow route to pull `@workflow/world-vercel` and its `debug` dependency into local builds, then fail during page-data collection with `module.createRequire failed parsing argument` and `Dynamic require of "tty" is not supported`.
|
|
432
|
-
|
|
433
|
-
**Fix**: Switched the runtime world loader to use `createRequire(import.meta.url)` and moved the local/Vercel world imports behind the existing async `createWorld()` branches. Local Next.js builds now only load the selected world implementation at runtime instead of bundling both worlds eagerly into the route module.
|
|
434
|
-
|
|
435
|
-
### Deferred Next.js Builds Re-Ran Eager Discovery
|
|
436
|
-
|
|
437
|
-
The later merge also pulled `BaseBuilder.createCombinedBundle()` into the deferred Next.js path without a way to pass the already-discovered workflow/step/serde entry sets. As a result, `packages/next/src/builder-deferred.ts` quietly fell back to `discoverEntries()` during production builds, re-emitting `Discovering workflow directives ...` and failing the local build tests that assert deferred mode avoids eager input-graph scans.
|
|
438
|
-
|
|
439
|
-
**Fix**: Threaded explicit `discoveredEntries` through `createCombinedBundle()` and passed the deferred builder's tracked workflow/step/serde file sets into that call. Deferred Next.js builds now reuse the socket/cache-driven discovery state instead of re-running the base eager discovery pass.
|
|
440
|
-
|
|
441
|
-
### Deferred Package Steps Fell Back to Compiled `dist/` Files
|
|
442
|
-
|
|
443
|
-
Once deferred discovery stopped re-running the base eager scan, some package-provided steps were only being rediscovered from built artifacts such as `packages/ai/dist/agent/durable-agent.js`. Those compiled files no longer carried every nested `'use step'` directive, so local production Next.js builds could miss registrations like `@workflow/ai/agent`'s `closeStream` helper and fail at runtime with "step is not registered in the current deployment".
|
|
444
|
-
|
|
445
|
-
**Fix**: The deferred Next.js builder now rewrites discovered workspace package paths from `dist/` back to their matching `src/` files when those sources exist. That keeps deferred bundling pointed at the directive-bearing source modules instead of their compiled output.
|
|
446
|
-
|
|
447
|
-
### Workspace Source Step IDs Lost Export Subpaths
|
|
448
|
-
|
|
449
|
-
Switching deferred builds over to workspace `src/` files fixed the missing nested directives, but it exposed a second mismatch in the SWC manifest path logic. `packages/builders/src/module-specifier.ts` only matched package exports against the on-disk file being transformed, so `packages/ai/src/agent/durable-agent.ts` was assigned `@workflow/ai@...` while the runtime still referenced the exported subpath id `@workflow/ai/agent@...`. Local Next.js agent runs then failed with "Step `step//@workflow/ai/agent@...//closeStream` is not registered" even though the source file was finally back in the bundle.
|
|
450
|
-
|
|
451
|
-
**Fix**: `resolveModuleSpecifier()` now treats workspace source files as the source-backed form of their exported `dist/` targets when deriving step ids. That preserves package export subpaths like `@workflow/ai/agent` for id generation while still bundling the directive-bearing `src/` modules.
|
|
452
|
-
|
|
453
|
-
### Tarball-Staged Next.js Builds Still Lost Package Step Sources
|
|
454
|
-
|
|
455
|
-
The local production and Postgres Next.js jobs stage the workbenches by packing workspace packages into tarballs and installing those tarballs into a temporary `node_modules` tree. Deferred discovery was already willing to rewrite workspace `packages/*/dist/*` files back to `src/*`, but the tarballed `@workflow/ai` package did not publish its `src/` tree and the base builder still treated `node_modules/@workflow/*/src/*` as ordinary package imports. That meant the staged CI path fell back to `dist/` again and dropped nested steps like `@workflow/ai/agent`'s `closeStream`, even after the workspace build path had been fixed.
|
|
456
|
-
|
|
457
|
-
**Fix**: Publish `packages/ai/src` in the tarball, treat source-backed `node_modules/@workflow/*/src/*` files like external workspace source files when generating bundle imports, and extend deferred transitive step discovery to follow bare `workflow` / `@workflow/*` package imports during non-watch builds.
|
|
458
|
-
|
|
459
|
-
### Vercel Step Source Map Expectations Were Too Optimistic
|
|
216
|
+
The interval was lowered from 100ms to 10ms. Per-step wait drops from ~50ms average to ~5ms (scaling linearly with the number of writable-bearing steps in a workflow). Per-tick work is just `writable.locked` plus a `getWriter()`/`releaseLock()` probe — microsecond-scale, so 10× more ticks is not measurable in practice.
|
|
460
217
|
|
|
461
|
-
|
|
218
|
+
**Follow-up**: Replace polling entirely with an event-driven release signal — wrap the writable returned from the `WritableStream` reviver with a writer that fires on `releaseLock()` — bringing the wait to ~0ms. This would also remove a source of timing drift between worlds with synchronous storage (`world-local`) and worlds with HTTP-deferred storage (`world-vercel`).
|
|
462
219
|
|
|
463
|
-
|
|
220
|
+
### Concurrent `step_started` and Attempt Counter
|
|
464
221
|
|
|
465
|
-
|
|
222
|
+
When the V2 handler dispatches N parallel steps as background messages, each background step completion queues a workflow continuation. Up to N continuations may replay concurrently, and each may attempt to start the same not-yet-completed step (since `step_started` succeeds for already-running steps). Each call atomically increments the `attempt` counter — so with N=5 parallel steps, the counter can reach 5 on the first genuine execution.
|
|
466
223
|
|
|
467
|
-
|
|
468
|
-
|
|
469
|
-
The Redis community-world benchmark still loads an external world package that has not adopted the newer `world.streams.*` interface yet. Once the eager-processing changes exercised stream writes through the modern namespace consistently, that adapter started failing with `Cannot read properties of undefined (reading 'writeMulti')` before the benchmark could even start.
|
|
470
|
-
|
|
471
|
-
**Decision**: Community world adapters must implement the `world.streams.*` interface. The runtime legacy stream normalization (`normalizeLegacyWorld`) was removed. Community world e2e tests are skipped until the adapters are updated.
|
|
472
|
-
|
|
473
|
-
### Build Output API Flow Handler Drift
|
|
474
|
-
|
|
475
|
-
The Vercel Build Output API builder still emitted the combined flow function as `index.js`, but the surrounding metadata kept pointing at `index.mjs`. That mismatch meant BOA-based preview deployments published neither `/.well-known/workflow/v1/flow` nor the public manifest, so the Vercel production e2e suite collapsed into manifest `404` errors immediately after deployment.
|
|
476
|
-
|
|
477
|
-
**Fix**: Updated `packages/builders/src/vercel-build-output-api.ts` to point both `.vc-config.json` and manifest extraction at `flow.func/index.js`, which matches the CommonJS file the builder actually writes.
|
|
478
|
-
|
|
479
|
-
### Async World Loading Broke Custom Target Worlds
|
|
480
|
-
|
|
481
|
-
The first lazy-world-loading fix switched package resolution over to `createRequire(import.meta.url)` globally. That solved the Next.js bundling problem for built-in worlds, but it also made custom targets like `@workflow/world-postgres` resolve relative to `@workflow/core` instead of the consuming app. Local Postgres tests then failed at startup with `Cannot find module '@workflow/world-postgres'`.
|
|
482
|
-
|
|
483
|
-
**Fix**: The runtime now creates the package resolver lazily from `process.cwd()/package.json` when possible, falling back to `import.meta.url` only when the app root cannot be resolved. That keeps custom world modules app-relative without reintroducing the eager module-load failure in Next.js builds.
|
|
484
|
-
|
|
485
|
-
### Core Logger Still Pulled `debug` Into Webpack Flow Routes
|
|
486
|
-
|
|
487
|
-
Even after lazy world loading stopped eagerly importing `@workflow/world-vercel`, the generated Next.js webpack flow route still evaluated `packages/core/src/logger.ts` at module load. That file had a top-level `import debug from 'debug'`, which in turn pulled `debug/src/node` and its `tty` dynamic require into `/.well-known/workflow/v1/flow`. Webpack then failed during page-data collection with `Dynamic require of "tty" is not supported`.
|
|
488
|
-
|
|
489
|
-
**Fix**: Replace the static `debug` dependency in the core logger with lightweight `process.env.DEBUG` matching plus `console.debug`. That keeps verbose opt-in logging for local debugging without forcing webpack to bundle `debug` and its Node-only terminal helpers into the flow route.
|
|
490
|
-
|
|
491
|
-
### Deferred Next.js Builder Helper Drift After Merge
|
|
492
|
-
|
|
493
|
-
Merging `main` into the eager-processing branch pulled in a set of helper methods for copied-step import rewriting in `packages/next/src/builder-deferred.ts`, but the corresponding call sites were not present on this branch yet. That left `getRelativeImportSpecifier`, `getStepCopyFileName`, and `rewriteRelativeImportsForCopiedStep` orphaned, and `@workflow/next` failed to build with `TS6133` unused-private-member errors immediately after the merge.
|
|
494
|
-
|
|
495
|
-
**Fix**: Removed the orphaned helper methods during merge resolution and kept the existing deferred-builder behavior unchanged. The copied-step import-rewrite work should land as a complete change set rather than a partial backport from `main`.
|
|
496
|
-
|
|
497
|
-
### Next.js React Step Fixture and `eval('require(...)')`
|
|
498
|
-
|
|
499
|
-
The `nextjs-webpack` e2e suite still failed after the merge in `workflows/8_react_render.tsx`, where the step intentionally did `eval('require("react-dom/server")')` to avoid Next.js linting rules around importing `react-dom/server` directly. That pattern was brittle under webpack rebundling: even though the intermediate step bundle had a `createRequire(import.meta.url)` banner, the rebundled route still failed at runtime with `TypeError: require is not a function`.
|
|
500
|
-
|
|
501
|
-
**Fix**: Updated the React-rendering step fixture in both Next.js workbenches to use `await import('react-dom/server')` instead. The test still exercises server-side React rendering inside a step, but no longer depends on bundler-specific `eval('require(...)')` behavior.
|
|
502
|
-
|
|
503
|
-
## Inline Execution Verification Tests
|
|
504
|
-
|
|
505
|
-
The `@workflow/world-testing` package includes invocation-counting tests that verify the V2 inline loop behavior for each workflow pattern:
|
|
506
|
-
|
|
507
|
-
| Workflow Pattern | Expected Invocations | Why |
|
|
508
|
-
|-----------------|---------------------|-----|
|
|
509
|
-
| Sequential steps (3 adds) | **1** | All steps execute inline |
|
|
510
|
-
| Sequential steps + WritableStream | **1** | Ops settle via flush waiter promises (500ms race) |
|
|
511
|
-
| Sleep (1s) + step | **2** | Sleep requires queue round-trip |
|
|
512
|
-
| Promise.all (2 steps) | **2-3** | Background step + inline replay after all steps done |
|
|
513
|
-
|
|
514
|
-
The test server tracks flow handler invocations per `runId` via an internal counter. Each test asserts the exact invocation count after the workflow completes.
|
|
515
|
-
|
|
516
|
-
### world-testing Flow Invocation Counting Missed Wrapped Queue Payloads
|
|
517
|
-
|
|
518
|
-
The inline-execution assertions in `packages/world-testing` count how many times the flow handler runs by inspecting the queue callback body and extracting `runId`. After the queue callback shape drifted, some worlds were only exposing the workflow payload under `body.payload.runId`, so the helper recorded `0` invocations even when the workflow completed correctly. That showed up in CI as the Postgres inline-execution spec failing its "single flow invocation" assertion.
|
|
519
|
-
|
|
520
|
-
**Fix**: Accept both top-level `runId` and nested `payload.runId` when tracking flow invocations in the embedded test server.
|
|
521
|
-
|
|
522
|
-
### Turbopack NFT Tracing Errors in V2 Combined Flow Route
|
|
523
|
-
|
|
524
|
-
The V2 combined flow route imports the step registrations bundle (`__step_registrations.js`), which esbuild produces as a monolithic file. On `main`, step registrations live in a separate route (`step/route.js`), so Turbopack traces them independently. In V2, Turbopack traces the step registrations through the flow route's import graph, encountering `world.ts` code with `process.cwd()`, dynamic `import()` calls to `@workflow/world-local`/`@workflow/world-vercel`, and `createRequire()` patterns — all of which trigger fatal NFT (Node File Trace) errors.
|
|
525
|
-
|
|
526
|
-
**Fix**: Introduced `get-world-lazy.ts`, a globalThis `Symbol.for`-based accessor that replaces the static `import { getWorld } from './runtime/world.js'` in all step-side modules (`serialization.ts`, `run.ts`, `helpers.ts`, `start.ts`, `resume-hook.ts`). This breaks the static import chain from step code to `world.ts`, preventing esbuild from bundling `world.ts` (and its transitive deps) into the step registrations. The step registrations bundle dropped from ~37k lines to ~6.6k lines (matching `main`), with zero `process.cwd()` or world package references.
|
|
527
|
-
|
|
528
|
-
The `getWorldLazy()` function reads from the globalThis world singleton cache (populated by the runtime's `getWorld()` on first call). When the cache is empty (e.g., `start()` called from application code before any workflow runs), it falls back to a dynamic `import()` of `world.js` to initialize the world.
|
|
529
|
-
|
|
530
|
-
Additional changes for Turbopack compatibility:
|
|
531
|
-
- Removed `stepEntrypoint` re-export from `runtime.ts` (V2 doesn't use separate step routes)
|
|
532
|
-
- Lazy-loaded `getPort` via `createRequire` with opaque specifier to prevent `@workflow/utils/get-port` filesystem operations from being traced
|
|
533
|
-
- `getRuntimeRequire()` uses `process.cwd()` as primary resolution base (for custom world packages like `@workflow/world-postgres` that are app-level deps, not `@workflow/core` deps), with `import.meta.url` fallback
|
|
534
|
-
|
|
535
|
-
### Cold-Start `MODULE_NOT_FOUND: './world.js'` From `getWorldLazy` Fallback
|
|
536
|
-
|
|
537
|
-
The `getWorldLazy()` design assumed one of two paths would always succeed: either `globalThis[GetWorldFnKey]` is populated (because some prior code reached `world.ts`'s module body), or the dynamic `import('./world.js')` fallback resolves at runtime.
|
|
538
|
-
|
|
539
|
-
Both assumptions break for routes that consume `start` (or any other `getWorldLazy` consumer) without going through the queue-driven flow handler first:
|
|
540
|
-
|
|
541
|
-
1. Webpack and Turbopack tree-shake the named import `{ getWorld } from './runtime/world.js'` out of `runtime.ts` once a consumer only uses `start`. `world.ts` is dropped from the bundle entirely, so its module-load `globalThis[GetWorldFnKey] ??= getWorld` registration never fires.
|
|
542
|
-
2. The dynamic-import fallback inside `get-world-lazy.ts` builds the specifier `./world.js` at runtime to evade bundler tracing — but webpack inlines `get-world-lazy.js` into the route bundle, so the relative specifier resolves against `/var/task/<app>/.next/server/app/<route>/route.js` where no sibling `world.js` exists. Node throws `MODULE_NOT_FOUND`.
|
|
543
|
-
|
|
544
|
-
The symptom: the very first request that goes through `start()` on a cold serverless invocation fails. Once any other code path (typically the queue-driven `/.well-known/workflow/v1/flow` route, which uses `getWorld` directly via `workflowEntrypoint`) has loaded `world.ts`, subsequent `start()` calls succeed for the rest of the process lifetime — making the failure flake-shaped: hard to reproduce in dev where everything tends to be warmed, but reliable on first user traffic into a fresh function instance.
|
|
545
|
-
|
|
546
|
-
**Fix**: Added `@workflow/core/runtime/world-init`, a server-only side-effect module that imports `./world.js` purely for its module-load side effect (the globalThis registration). It's exported via package conditions:
|
|
547
|
-
|
|
548
|
-
- `default` → `./dist/runtime/world-init.js` (real, loads `world.ts`)
|
|
549
|
-
- `workflow` → `./dist/workflow/world-init-stub.js` (empty, used by VM/step bundles)
|
|
550
|
-
|
|
551
|
-
`packages/workflow/src/api.ts` (the host file behind `workflow/api`'s `default` condition) imports it for its side effect. The matching VM/step entry `api-workflow.ts` does not, so `world.ts` and its server-only deps (`@workflow/world-local`, `@workflow/world-vercel`, `cbor-x`, …) stay out of the workflow sandbox bundle.
|
|
552
|
-
|
|
553
|
-
Reverification: built bundles for `vade-review` (Next.js webpack) show `createLocalWorld`/`createVercelWorld`/`GetWorldFnKey` present in the route's vendor chunk for `@workflow/core` (zero before the fix), and the workflow VM bundle's `flow/route.js` and `__step_registrations.js` continue to have zero references to either the world-init module or `world.ts`. Cold-start `POST /api/review/submit` succeeds on the first request after a fresh server boot — the regression case.
|
|
554
|
-
|
|
555
|
-
The dynamic-import fallback in `get-world-lazy.ts` is preserved as defense-in-depth for environments outside the documented configurations (CJS test runners, scripts that import deeply into `@workflow/core` without going through `workflow/api`).
|
|
556
|
-
|
|
557
|
-
### Run#returnValue Worker Deadlock in V2 Inline Execution
|
|
558
|
-
|
|
559
|
-
When a workflow calls `start()` to spawn child workflows (e.g., `fibonacciWorkflow`), the parent's `Run#returnValue` step polls the child's completion status in a blocking loop (`while (true) { ... sleep(1000) ... }`). In V2, this step is executed inline by the step executor, holding a worker thread slot. If the child workflow's queue message is waiting for the same worker pool, the parent blocks the child from starting — a classic deadlock.
|
|
560
|
-
|
|
561
|
-
**Fix**: `Run#pollReturnValue()` detects whether it's running inside a step executor (via `contextStorage.getStore()`) and, if so, throws `TooEarlyError` instead of polling in a blocking loop. `TooEarlyError` is handled specially by the step executor — it returns `{ type: 'retry', timeoutSeconds }` which re-queues the step with a 1-second delay, freeing the worker to process child workflows. Unlike `RetryableError`, `TooEarlyError` does NOT count against `maxRetries`, so polling steps can retry indefinitely until the child completes.
|
|
562
|
-
|
|
563
|
-
When called from outside a step (e.g., test code, API routes), `pollReturnValue()` retains the original blocking loop behavior for backward compatibility.
|
|
224
|
+
The max retries check in `executeStep()` only enforces when `step.error` exists — distinguishing actual retries (failed → retry with error) from concurrent first-attempt races (multiple handlers start the same step simultaneously without any prior failure). Concurrent starts are harmless since `step_completed` idempotency ensures only the first completion wins.
|
|
564
225
|
|
|
565
226
|
### Unconsumed Event Check Two-Phase Drain
|
|
566
227
|
|
|
567
|
-
|
|
228
|
+
The `EventsConsumer`'s unconsumed event check uses a two-phase promise queue drain: yield once after the first drain (via `setTimeout(0)`) so cross-VM promise chains can append follow-up async work, then re-drain before checking. This improves timing for scenarios like `step_completed` → for-await loop resume → next hook hydration.
|
|
568
229
|
|
|
569
230
|
The check additionally arms a `DEFERRED_CHECK_DELAY_MS = 100` `setTimeout` after the second drain, since Node.js does not guarantee that `setTimeout(0)` fires after all cross-context microtasks settle. Any `subscribe()` call arriving during that 100ms window cancels the check via version invalidation + `clearTimeout`, so the delay only adds latency to genuine corruption — never to the happy path.
|
|
570
231
|
|
|
571
|
-
**Follow-up**: 100ms is a heuristic chosen empirically
|
|
232
|
+
**Follow-up**: 100ms is a heuristic chosen empirically. A deterministic settlement signal — for example, a "VM idle" callback exposed by the workflow VM bridge that fires only after all pending cross-context promise chains have resolved — would let the consumer fire the unconsumed-event check immediately on quiescence instead of waiting for a wall-clock timeout.
|
|
233
|
+
|
|
234
|
+
### Lazy World Loading
|
|
572
235
|
|
|
573
|
-
|
|
236
|
+
Static imports of `world-local` and `world-vercel` from the runtime caused two distinct build-time issues: Next.js production builds pulled both worlds (including their Node-only deps like `debug`'s `tty` requires) into the route module, and Turbopack's NFT (Node File Trace) errored on `process.cwd()` and dynamic `import()` patterns it couldn't statically analyze.
|
|
574
237
|
|
|
575
|
-
|
|
238
|
+
A `getWorldLazy()` accessor (backed by a `globalThis` `Symbol.for` cache) replaces the static import in step-side modules. This breaks the static import chain from step code to `world.ts`, preventing both worlds from being bundled into the step registrations.
|
|
576
239
|
|
|
577
|
-
|
|
240
|
+
Because tree-shaking can otherwise drop `world.ts`'s module-load registration entirely, a server-only side-effect module (`@workflow/core/runtime/world-init`) imports `./world.js` purely for its module-load side effect. It's wired via package conditions:
|
|
578
241
|
|
|
579
|
-
|
|
242
|
+
- `default` → real, loads `world.ts`
|
|
243
|
+
- `workflow` → empty stub, used by VM/step bundles
|
|
580
244
|
|
|
581
|
-
|
|
582
|
-
|-----------|-----------|--------|
|
|
583
|
-
| Unit Tests | core | 581/581 |
|
|
584
|
-
| Embedded Tests | world-testing | 9/9 (including inline execution) |
|
|
585
|
-
| Local Dev | 14 frameworks | All pass |
|
|
586
|
-
| Local Prod | 14 configurations | All pass |
|
|
587
|
-
| Postgres | 14 frameworks | All pass |
|
|
588
|
-
| Vercel Prod | 11 frameworks | All pass |
|
|
589
|
-
| Vercel Deployments | 15 projects | All succeed |
|
|
590
|
-
| Community Worlds | Turso, MongoDB, Redis | All pass |
|
|
591
|
-
| Windows | e2e | Pass |
|
|
245
|
+
This guarantees the world is loaded for routes that consume `start()` without going through the queue-driven flow handler first, while keeping `world.ts` and its server-only deps out of the workflow sandbox bundle.
|
|
592
246
|
|
|
593
|
-
|
|
247
|
+
### Community Worlds and the `world.streams` API
|
|
594
248
|
|
|
595
|
-
|
|
249
|
+
Community world adapters must implement the `world.streams.*` interface. The runtime legacy stream normalization was removed as part of this work.
|