workflow 5.0.0-beta.2 → 5.0.0-beta.21
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/api-workflow.d.ts +1 -1
- package/dist/api-workflow.d.ts.map +1 -1
- package/dist/api-workflow.js +2 -2
- package/dist/api.d.ts +5 -1
- package/dist/api.d.ts.map +1 -1
- package/dist/api.js +14 -2
- package/dist/internal/builtins.d.ts +17 -0
- package/dist/internal/builtins.d.ts.map +1 -1
- package/dist/internal/builtins.js +65 -1
- package/dist/observability.d.ts +1 -1
- package/dist/observability.js +2 -2
- package/dist/runtime.d.ts +1 -1
- package/dist/runtime.d.ts.map +1 -1
- package/dist/runtime.js +2 -2
- package/docs/ai/index.mdx +27 -23
- package/docs/api-reference/index.mdx +24 -0
- package/docs/api-reference/meta.json +8 -0
- package/docs/api-reference/vitest/index.mdx +28 -7
- package/docs/api-reference/workflow/create-hook.mdx +38 -0
- package/docs/api-reference/workflow/create-webhook.mdx +1 -0
- package/docs/api-reference/workflow/experimental-set-attributes.mdx +65 -0
- package/docs/api-reference/workflow/fetch.mdx +5 -0
- package/docs/api-reference/workflow/index.mdx +3 -0
- package/docs/api-reference/workflow-ai/durable-agent.mdx +7 -45
- package/docs/api-reference/workflow-api/get-hook-by-token.mdx +7 -0
- package/docs/api-reference/workflow-api/get-run.mdx +6 -0
- package/docs/api-reference/workflow-api/index.mdx +6 -8
- package/docs/api-reference/workflow-api/resume-hook.mdx +57 -0
- package/docs/api-reference/workflow-api/start.mdx +13 -5
- package/docs/api-reference/workflow-astro/index.mdx +18 -0
- package/docs/api-reference/workflow-astro/meta.json +4 -0
- package/docs/api-reference/workflow-astro/workflow.mdx +45 -0
- package/docs/api-reference/workflow-errors/hook-conflict-error.mdx +60 -0
- package/docs/api-reference/workflow-errors/index.mdx +85 -0
- package/docs/api-reference/workflow-errors/meta.json +5 -0
- package/docs/api-reference/workflow-errors/run-not-supported-error.mdx +58 -0
- package/docs/api-reference/workflow-errors/workflow-error.mdx +52 -0
- package/docs/api-reference/workflow-errors/workflow-run-failed-error.mdx +16 -6
- package/docs/api-reference/workflow-errors/workflow-run-not-completed-error.mdx +58 -0
- package/docs/api-reference/workflow-errors/workflow-runtime-error.mdx +58 -0
- package/docs/api-reference/workflow-nest/configure-workflow-controller.mdx +33 -0
- package/docs/api-reference/workflow-nest/index.mdx +31 -0
- package/docs/api-reference/workflow-nest/meta.json +9 -0
- package/docs/api-reference/workflow-nest/nest-local-builder.mdx +64 -0
- package/docs/api-reference/workflow-nest/workflow-controller.mdx +40 -0
- package/docs/api-reference/workflow-nest/workflow-module.mdx +74 -0
- package/docs/api-reference/workflow-next/with-workflow.mdx +56 -2
- package/docs/api-reference/workflow-nitro/index.mdx +59 -0
- package/docs/api-reference/workflow-nuxt/index.mdx +47 -0
- package/docs/api-reference/workflow-observability/hydrate-data.mdx +35 -0
- package/docs/api-reference/workflow-observability/hydrate-resource-io.mdx +62 -0
- package/docs/api-reference/workflow-observability/index.mdx +64 -0
- package/docs/api-reference/workflow-observability/meta.json +11 -0
- package/docs/api-reference/workflow-observability/observability-revivers.mdx +50 -0
- package/docs/api-reference/workflow-observability/parse-class-name.mdx +41 -0
- package/docs/api-reference/workflow-observability/parse-step-name.mdx +40 -0
- package/docs/api-reference/workflow-observability/parse-workflow-name.mdx +55 -0
- package/docs/api-reference/workflow-runtime/create-world.mdx +39 -0
- package/docs/api-reference/workflow-runtime/get-world-handlers.mdx +44 -0
- package/docs/api-reference/{workflow-api → workflow-runtime}/get-world.mdx +7 -10
- package/docs/api-reference/workflow-runtime/health-check.mdx +50 -0
- package/docs/api-reference/workflow-runtime/index.mdx +43 -0
- package/docs/api-reference/workflow-runtime/meta.json +12 -0
- package/docs/api-reference/workflow-runtime/set-world.mdx +49 -0
- package/docs/api-reference/workflow-runtime/workflow-entrypoint.mdx +42 -0
- package/docs/api-reference/{workflow-api → workflow-runtime}/world/index.mdx +5 -8
- package/docs/api-reference/workflow-runtime/world/meta.json +4 -0
- package/docs/api-reference/{workflow-api → workflow-runtime}/world/queue.mdx +2 -2
- package/docs/api-reference/{workflow-api → workflow-runtime}/world/storage.mdx +11 -4
- package/docs/api-reference/{workflow-api → workflow-runtime}/world/streams.mdx +2 -2
- package/docs/api-reference/workflow-serde/index.mdx +0 -1
- package/docs/api-reference/workflow-serde/workflow-deserialize.mdx +1 -2
- package/docs/api-reference/workflow-serde/workflow-serialize.mdx +1 -2
- package/docs/api-reference/workflow-sveltekit/index.mdx +18 -0
- package/docs/api-reference/workflow-sveltekit/meta.json +4 -0
- package/docs/api-reference/workflow-sveltekit/workflow-plugin.mdx +42 -0
- package/docs/api-reference/workflow-vite/index.mdx +18 -0
- package/docs/api-reference/workflow-vite/meta.json +4 -0
- package/docs/api-reference/workflow-vite/workflow.mdx +48 -0
- package/docs/changelog/attributes-mvp.mdx +380 -0
- package/docs/changelog/eager-processing.mdx +269 -0
- package/docs/changelog/index.mdx +2 -1
- package/docs/changelog/lazy-event-creation.md +127 -0
- package/docs/changelog/meta.json +7 -1
- package/docs/changelog/resilient-start.mdx +31 -283
- package/docs/changelog/turbo-mode.md +87 -0
- package/docs/cookbook/advanced/child-workflows.mdx +315 -0
- package/docs/cookbook/advanced/meta.json +2 -3
- package/docs/cookbook/advanced/publishing-libraries.mdx +87 -29
- package/docs/cookbook/advanced/serializable-steps.mdx +17 -5
- package/docs/cookbook/advanced/upgrading-workflows.mdx +195 -0
- package/docs/cookbook/agent-patterns/agent-cancellation.mdx +156 -0
- package/docs/cookbook/agent-patterns/durable-agent.mdx +11 -184
- package/docs/cookbook/agent-patterns/human-in-the-loop.mdx +150 -173
- package/docs/cookbook/agent-patterns/meta.json +1 -7
- package/docs/cookbook/common-patterns/batching.mdx +44 -118
- package/docs/cookbook/common-patterns/idempotency.mdx +36 -52
- package/docs/cookbook/common-patterns/meta.json +4 -4
- package/docs/cookbook/common-patterns/rate-limiting.mdx +1 -1
- package/docs/cookbook/common-patterns/saga.mdx +128 -33
- package/docs/cookbook/common-patterns/scheduling.mdx +77 -193
- package/docs/cookbook/common-patterns/sequential-and-parallel.mdx +155 -0
- package/docs/cookbook/common-patterns/timeouts.mdx +100 -0
- package/docs/cookbook/common-patterns/workflow-composition.mdx +117 -0
- package/docs/cookbook/index.mdx +14 -17
- package/docs/cookbook/integrations/ai-sdk.mdx +330 -142
- package/docs/cookbook/integrations/chat-sdk.mdx +264 -151
- package/docs/cookbook/integrations/sandbox.mdx +482 -81
- package/docs/cookbook/meta.json +1 -1
- package/docs/deploying/building-a-world.mdx +1 -1
- package/docs/deploying/world/postgres-world.mdx +5 -3
- package/docs/deploying/world/vercel-world.mdx +2 -0
- package/docs/errors/abort-signal-timeout-in-workflow.mdx +80 -0
- package/docs/errors/corrupted-event-log.mdx +5 -5
- package/docs/errors/hook-conflict.mdx +56 -4
- package/docs/errors/index.mdx +9 -0
- package/docs/errors/replay-divergence.mdx +27 -0
- package/docs/errors/runtime-decryption-failed.mdx +77 -0
- package/docs/errors/step-executed-multiple-times.mdx +23 -0
- package/docs/errors/step-not-registered.mdx +1 -1
- package/docs/foundations/cancellation.mdx +459 -0
- package/docs/foundations/errors-and-retries.mdx +7 -3
- package/docs/foundations/hooks.mdx +29 -0
- package/docs/foundations/idempotency.mdx +236 -11
- package/docs/foundations/index.mdx +3 -3
- package/docs/foundations/meta.json +3 -2
- package/docs/foundations/serialization.mdx +78 -42
- package/docs/foundations/starting-workflows.mdx +6 -2
- package/docs/foundations/streaming.mdx +14 -23
- package/docs/foundations/versioning.mdx +263 -0
- package/docs/getting-started/astro.mdx +6 -0
- package/docs/getting-started/index.mdx +6 -7
- package/docs/getting-started/meta.json +1 -0
- package/docs/getting-started/nestjs.mdx +9 -0
- package/docs/getting-started/next.mdx +5 -3
- package/docs/getting-started/nitro.mdx +22 -0
- package/docs/getting-started/sveltekit.mdx +6 -0
- package/docs/getting-started/tanstack-start.mdx +241 -0
- package/docs/how-it-works/cancellation.mdx +287 -0
- package/docs/how-it-works/code-transform.mdx +2 -2
- package/docs/how-it-works/encryption.mdx +2 -2
- package/docs/how-it-works/event-sourcing.mdx +2 -2
- package/docs/how-it-works/meta.json +2 -1
- package/docs/internal/index.mdx +21 -0
- package/docs/internal/meta.json +10 -0
- package/docs/internal/nitro-native-build.mdx +38 -0
- package/docs/internal/nitro-web-ui.mdx +24 -0
- package/docs/internal/serializable-abort-controller.mdx +148 -0
- package/docs/migration-guides/migrating-from-aws-step-functions.mdx +63 -16
- package/docs/migration-guides/migrating-from-inngest.mdx +44 -22
- package/docs/migration-guides/migrating-from-temporal.mdx +43 -14
- package/docs/migration-guides/migrating-from-trigger-dev.mdx +59 -27
- package/docs/observability/attributes.mdx +87 -0
- package/docs/observability/index.mdx +25 -1
- package/docs/observability/meta.json +1 -1
- package/docs/observability/tracing.mdx +106 -0
- package/docs/testing/index.mdx +2 -2
- package/package.json +14 -13
- package/docs/api-reference/workflow-api/world/meta.json +0 -4
- package/docs/api-reference/workflow-api/world/observability.mdx +0 -164
- package/docs/cookbook/advanced/custom-serialization.mdx +0 -168
- package/docs/cookbook/advanced/durable-objects.mdx +0 -148
- package/docs/cookbook/advanced/isomorphic-packages.mdx +0 -145
- package/docs/cookbook/agent-patterns/stop-workflow.mdx +0 -216
- package/docs/cookbook/agent-patterns/tool-orchestration.mdx +0 -255
- package/docs/cookbook/agent-patterns/tool-streaming.mdx +0 -181
- package/docs/cookbook/common-patterns/child-workflows.mdx +0 -372
- package/docs/cookbook/common-patterns/content-router.mdx +0 -207
- package/docs/cookbook/common-patterns/fan-out.mdx +0 -208
- package/docs/foundations/common-patterns.mdx +0 -265
|
@@ -7,321 +7,69 @@ description: Overhaul run start logic to tolerate world storage unavailability,
|
|
|
7
7
|
|
|
8
8
|
## Motivation
|
|
9
9
|
|
|
10
|
-
When `world` storage is unavailable but the queue is up,
|
|
11
|
-
`start()` previously failed entirely because `world.events.create(run_created)`
|
|
12
|
-
is called before `world.queue()`. This change decouples run creation from queue
|
|
13
|
-
dispatch so that runs can still be accepted when storage is degraded.
|
|
10
|
+
When `world` storage is unavailable but the queue is up, `start()` previously failed entirely because `world.events.create(run_created)` is called before `world.queue()`. This change decouples run creation from queue dispatch so that runs can still be accepted when storage is degraded.
|
|
14
11
|
|
|
15
|
-
Additionally, the runtime previously called `world.runs.get(runId)` before
|
|
16
|
-
`run_started`, adding an extra round-trip. By always calling `run_started`
|
|
17
|
-
directly, we save that round-trip and can return pre-loaded events in the
|
|
18
|
-
response to skip the initial `events.list` call, reducing TTFB.
|
|
12
|
+
Additionally, the runtime previously called `world.runs.get(runId)` before `run_started`, adding an extra round-trip. By always calling `run_started` directly, we save that round-trip and can return pre-loaded events in the response to skip the initial `events.list` call, reducing TTFB.
|
|
19
13
|
|
|
20
14
|
## Design
|
|
21
15
|
|
|
22
|
-
### `start()` changes
|
|
16
|
+
### `start()` changes
|
|
23
17
|
|
|
24
|
-
- `world.events.create` (run_created) and `world.queue` are now called **in parallel**
|
|
25
|
-
|
|
26
|
-
- If `events.create` errors with **
|
|
27
|
-
creation failed but the run was accepted — creation will be re-tried async by the
|
|
28
|
-
runtime when it processes the queue message. The returned `Run` instance is marked
|
|
29
|
-
with `resilientStart = true`.
|
|
30
|
-
- If `events.create` errors with **409** (EntityConflictError), the run already exists
|
|
31
|
-
(e.g., the queue handler's resilient start path created it first due to a cold-start
|
|
32
|
-
race). This is treated as success.
|
|
18
|
+
- `world.events.create` (run_created) and `world.queue` are now called **in parallel** via `Promise.allSettled`.
|
|
19
|
+
- If `events.create` errors with **429 or 5xx**, we log a warning saying that run creation failed but the run was accepted — creation will be re-tried async by the runtime when it processes the queue message. The returned `Run` instance is marked with `resilientStart = true`.
|
|
20
|
+
- If `events.create` errors with **409** (EntityConflictError), the run already exists (e.g., the queue handler's resilient start path created it first due to a cold-start race). This is treated as success.
|
|
33
21
|
- If `world.queue` fails, we still throw — the run truly failed and was not enqueued.
|
|
34
|
-
- The queue invocation now receives all the run inputs (`input`, `deploymentId`,
|
|
35
|
-
|
|
36
|
-
create the run later if needed.
|
|
37
|
-
- When the runtime re-enqueues itself, it does **not** pass these inputs — only the
|
|
38
|
-
first queue cycle carries them.
|
|
22
|
+
- The queue invocation now receives all the run inputs (`input`, `deploymentId`, `workflowName`, `specVersion`, `executionContext`) via `runInput` so the runtime can create the run later if needed.
|
|
23
|
+
- When the runtime re-enqueues itself, it does **not** pass these inputs — only the first queue cycle carries them.
|
|
39
24
|
|
|
40
|
-
### `workflowEntrypoint` changes
|
|
25
|
+
### `workflowEntrypoint` changes
|
|
41
26
|
|
|
42
|
-
- When calling `world.events.create` with `run_started`, we now also always pass the
|
|
43
|
-
run input that was sent through the queue, if available. The response will still be on off:
|
|
44
|
-
- **200 with event (now running)**: As usual, but the server could have used the run input to create the run if it didn't exist yet. The response will be opaque to the runtime.
|
|
45
|
-
- **200 without event (already running)**: As usual
|
|
46
|
-
- **409 or 410 (already finished)**: As usual
|
|
27
|
+
- When calling `world.events.create` with `run_started`, we now also always pass the run input that was sent through the queue, if available. The world is responsible for creating the run if it doesn't already exist.
|
|
47
28
|
|
|
48
|
-
### `Run.returnValue` polling
|
|
29
|
+
### `Run.returnValue` polling
|
|
49
30
|
|
|
50
|
-
- When `resilientStart` is true on the Run instance (run_created failed), the
|
|
51
|
-
|
|
52
|
-
(1s + 3s + 6s = 10s total) to give the queue time to deliver and the runtime
|
|
53
|
-
to create the run via `run_started`.
|
|
54
|
-
- When `resilientStart` is false (normal path), 404 fails immediately — no delay
|
|
55
|
-
for the common case of a wrong run ID.
|
|
31
|
+
- When `resilientStart` is true on the Run instance (run_created failed), the `pollReturnValue` loop retries on `WorkflowRunNotFoundError` up to 3 times (1s + 3s + 6s = 10s total) to give the queue time to deliver and the runtime to create the run via `run_started`.
|
|
32
|
+
- When `resilientStart` is false (normal path), 404 fails immediately — no delay for the common case of a wrong run ID.
|
|
56
33
|
|
|
57
|
-
### World
|
|
34
|
+
### World contract changes
|
|
58
35
|
|
|
59
|
-
- Posting `run_started` to a **non-existent** run is now allowed when the run input is
|
|
60
|
-
|
|
61
|
-
1. Creates a `run_created` event first (so the event log is consistent).
|
|
62
|
-
2. Strips the input from the `run_started` event data (it lives on `run_created`).
|
|
63
|
-
3. Then creates the `run_started` event normally.
|
|
64
|
-
4. Emits a log and a Datadog metric (`workflow_server.resilient_start.run_created_via_run_started`)
|
|
65
|
-
to track when this fallback path is hit.
|
|
66
|
-
- When `run_started` encounters an **already-running** run, all worlds return `{ run }`
|
|
67
|
-
with `event: undefined` instead of throwing. No duplicate event is created.
|
|
36
|
+
- Posting `run_started` to a **non-existent** run is now allowed when the run input is sent along with the payload. The world creates a `run_created` event first (so the event log is consistent), then creates the `run_started` event normally.
|
|
37
|
+
- When `run_started` encounters an **already-running** run, all worlds return `{ run }` with `event: undefined` instead of throwing. No duplicate event is created.
|
|
68
38
|
|
|
69
39
|
### Queue transport changes
|
|
70
40
|
|
|
71
|
-
`Uint8Array` values (the serialized workflow input in `runInput`) don't survive plain
|
|
72
|
-
JSON serialization. Each world uses a transport that preserves binary data:
|
|
41
|
+
`Uint8Array` values (the serialized workflow input in `runInput`) don't survive plain JSON serialization. Each world uses a transport that preserves binary data:
|
|
73
42
|
|
|
74
|
-
- **world-vercel**: CBOR transport — CBOR-encodes the entire queue payload into a
|
|
75
|
-
|
|
76
|
-
- **world-
|
|
77
|
-
from `fs.ts` that encode Uint8Array as `{ __type: 'Uint8Array', data: '<base64>' }`.
|
|
78
|
-
- **world-postgres**: Inline typed JSON transport — same tagged-envelope approach as
|
|
79
|
-
world-local, inlined since world-postgres doesn't import from world-local.
|
|
43
|
+
- **world-vercel**: CBOR transport — CBOR-encodes the entire queue payload into a `Buffer` and uses `BufferTransport` from `@vercel/queue`. Uint8Array survives natively.
|
|
44
|
+
- **world-local**: `TypedJsonTransport` — encodes Uint8Array as `{ __type: 'Uint8Array', data: '<base64>' }`.
|
|
45
|
+
- **world-postgres**: Inline typed JSON transport — same tagged-envelope approach as world-local.
|
|
80
46
|
|
|
81
47
|
## Decisions
|
|
82
48
|
|
|
83
|
-
1. **Parallel not sequential**: We chose `Promise.allSettled` over sequential calls to
|
|
84
|
-
minimize latency in the happy path.
|
|
49
|
+
1. **Parallel not sequential**: We chose `Promise.allSettled` over sequential calls to minimize latency in the happy path.
|
|
85
50
|
|
|
86
|
-
2. **Already-running returns run without event**: When `run_started` encounters an
|
|
87
|
-
already-running run, all worlds return `{ run }` with `event: undefined` (no
|
|
88
|
-
`events` array) instead of throwing. The runtime detects this by checking for
|
|
89
|
-
`result.event === undefined`. This avoids an extra `world.runs.get` round-trip.
|
|
51
|
+
2. **Already-running returns run without event**: When `run_started` encounters an already-running run, all worlds return `{ run }` with `event: undefined` (no `events` array) instead of throwing. The runtime detects this by checking for `result.event === undefined`. This avoids an extra `world.runs.get` round-trip.
|
|
90
52
|
|
|
91
|
-
3. **Events in 200 response**: We only return events on the 200 path (first caller).
|
|
92
|
-
On the already-running path, we fall back to the normal `events.list` call. This is
|
|
93
|
-
correct because only on 200 can we be certain we know the full event history.
|
|
53
|
+
3. **Events in 200 response**: We only return events on the 200 path (first caller). On the already-running path, we fall back to the normal `events.list` call. This is correct because only on 200 can we be certain we know the full event history.
|
|
94
54
|
|
|
95
|
-
4. **Conditional 404 retry on Run.returnValue**: Only when `resilientStart = true`
|
|
96
|
-
(run_created failed). Normal runs fail fast on 404.
|
|
55
|
+
4. **Conditional 404 retry on Run.returnValue**: Only when `resilientStart = true` (run_created failed). Normal runs fail fast on 404.
|
|
97
56
|
|
|
98
57
|
## Known concerns
|
|
99
58
|
|
|
100
|
-
### Cold-start race on Vercel
|
|
59
|
+
### Cold-start race on Vercel
|
|
101
60
|
|
|
102
|
-
On Vercel, the parallel dispatch can cause the queue message to be processed before
|
|
103
|
-
`run_created` completes, if `run_created` hits a cold-start lambda. Confirmed via
|
|
104
|
-
Datadog: the `run_started` request hit a warm lambda (23ms) while `run_created` hit
|
|
105
|
-
a cold lambda (727ms), even though `run_created` arrived at the edge 116ms earlier.
|
|
106
|
-
When this happens:
|
|
61
|
+
On Vercel, the parallel dispatch can cause the queue message to be processed before `run_created` completes, if `run_created` hits a cold-start lambda. When this happens:
|
|
107
62
|
|
|
108
63
|
1. The runtime's resilient start path creates the run from `run_started`.
|
|
109
64
|
2. The original `run_created` arrives and gets 409 (EntityConflictError).
|
|
110
65
|
3. `start()` treats the 409 as success (the run exists).
|
|
111
66
|
|
|
112
|
-
|
|
113
|
-
in this case (409 is not a retryable error), so `returnValue` fails fast on 404.
|
|
67
|
+
The `resilientStart` flag is NOT set on the Run instance in this case (409 is not a retryable error), so `returnValue` fails fast on 404.
|
|
114
68
|
|
|
115
|
-
###
|
|
69
|
+
### Atomicity of run entity creation
|
|
116
70
|
|
|
117
|
-
|
|
118
|
-
`events.create(run_created)` finishes writing to the shared filesystem. The
|
|
119
|
-
resilient start path should handle this, but Local Prod tests showed occasional
|
|
120
|
-
runs stuck at `pending` (no `run_started` event), and Windows CI showed
|
|
121
|
-
"Unconsumed event in event log" errors from duplicate `run_created` events.
|
|
71
|
+
The normal `run_created` path and the resilient start path can race on creating the run entity. In `world-local`, both paths use `writeExclusive` (O_CREAT|O_EXCL) — atomic at the OS level, so exactly one writer wins and the other gets EEXIST. The normal path throws `EntityConflictError` on conflict (handled by `start()` as 409); the resilient start path re-reads the run from disk on conflict.
|
|
122
72
|
|
|
123
|
-
|
|
124
|
-
resilient start path. Both used `writeJSON` which checks existence with
|
|
125
|
-
`fs.access()` (non-atomic), so both could pass the check and write separate
|
|
126
|
-
`run_created` events with different event IDs. Fixed by switching both paths to
|
|
127
|
-
`writeExclusive` (O_CREAT|O_EXCL) — see retrospective items 12 and 16.
|
|
73
|
+
In `world-postgres`, the resilient start path uses `onConflictDoNothing` plus a re-read on conflict for the same effect, with the same outcome on either side of the race.
|
|
128
74
|
|
|
129
|
-
|
|
130
|
-
|
|
131
|
-
- [x] ~~Investigate Local Prod test flakiness~~ — resolved via `writeExclusive`
|
|
132
|
-
for run entity creation (retrospective items 12, 16).
|
|
133
|
-
- [ ] Monitor the Datadog metric in production to understand how often the fallback is hit.
|
|
134
|
-
- [x] ~~Events optimization for re-enqueue cycles~~ — decided against. The
|
|
135
|
-
already-running path returns early without writing an event, so preloading
|
|
136
|
-
events there would require an extra filesystem/DB query on every re-enqueue.
|
|
137
|
-
More importantly, on Vercel with at-least-once delivery, multiple lambdas can
|
|
138
|
-
process the same run concurrently — the event snapshot could be stale or
|
|
139
|
-
incomplete. The runtime's fallback to `events.list` is the correct behavior
|
|
140
|
-
for re-enqueue cycles.
|
|
141
|
-
- [x] ~~CborTransport pass-through~~ — refactored. `encode()`/`decode()` now
|
|
142
|
-
live inside `CborTransport.serialize()`/`deserialize()`, matching the pattern
|
|
143
|
-
used by TypedJsonTransport (world-local) and the inline transport
|
|
144
|
-
(world-postgres). Call sites pass plain objects instead of pre-encoded buffers.
|
|
145
|
-
|
|
146
|
-
## Development retrospective
|
|
147
|
-
|
|
148
|
-
Chronological log of mistakes, misunderstandings, and reverted approaches during
|
|
149
|
-
development. Included for future reference when working on similar cross-cutting
|
|
150
|
-
runtime changes.
|
|
151
|
-
|
|
152
|
-
### 1. Uint8Array corruption through JSON queue transport
|
|
153
|
-
|
|
154
|
-
The initial implementation passed `runInput.input` (a `Uint8Array`) directly through
|
|
155
|
-
the queue payload. `Uint8Array` doesn't survive `JSON.stringify` — it becomes
|
|
156
|
-
`{"0":72,"1":101,...}`. This corrupted the workflow input when the resilient start
|
|
157
|
-
path tried to recreate the run from the queue-delivered data.
|
|
158
|
-
|
|
159
|
-
Caught by the `spawnWorkflowFromStepWorkflow` e2e test and the `world-testing`
|
|
160
|
-
embedded tests, which failed with "Invalid input" from devalue's `unflatten()`.
|
|
161
|
-
|
|
162
|
-
Three approaches were tried before landing on the final solution:
|
|
163
|
-
|
|
164
|
-
1. **Base64 encoding** (`btoa`/`atob`) — worked but fragile. The decode side used
|
|
165
|
-
`typeof runInput.input === 'string'` as a discriminant, which was flagged as
|
|
166
|
-
dangerous since non-binary inputs could also be strings.
|
|
167
|
-
2. **`Array.from()`/`new Uint8Array()`** — replaced base64 with a plain number array.
|
|
168
|
-
Two problems: (a) 3x JSON size regression vs base64, and (b) `Array.isArray()`
|
|
169
|
-
false-positives on v1Compat runs where `dehydrateWorkflowArguments` returns
|
|
170
|
-
devalue's flat Array format.
|
|
171
|
-
3. **CBOR + BufferTransport** (final) — world-vercel CBOR-encodes the queue payload;
|
|
172
|
-
world-local and world-postgres use a `TypedJsonTransport` with a tagged envelope.
|
|
173
|
-
|
|
174
|
-
### 2. Forgot to commit world-postgres transport fix (twice)
|
|
175
|
-
|
|
176
|
-
After fixing world-local and world-vercel queue transports, the same `JsonTransport`
|
|
177
|
-
corruption bug existed in world-postgres. The fix was written during a session but
|
|
178
|
-
never committed — lost when the working directory was reset via stash/checkout. This
|
|
179
|
-
happened twice. The fix only landed on the third attempt when it was committed and
|
|
180
|
-
pushed immediately. All 14 Postgres e2e jobs failed each time.
|
|
181
|
-
|
|
182
|
-
### 3. Incorrect diagnosis of Vercel Prod 409 errors
|
|
183
|
-
|
|
184
|
-
Multiple Vercel Prod e2e tests failed with `EntityConflictError: Workflow run with
|
|
185
|
-
ID wrun_... already exists` on `run_created`. The initial assumption was that VQS
|
|
186
|
-
couldn't deliver the queue message fast enough to beat the `run_created` call.
|
|
187
|
-
|
|
188
|
-
Datadog logs showed otherwise: the `run_created` request arrived at Vercel's edge
|
|
189
|
-
116ms before `run_started`, but `run_created` hit a cold-start lambda (727ms) while
|
|
190
|
-
`run_started` hit a warm one (23ms). Cold starts can invert expected execution order.
|
|
191
|
-
|
|
192
|
-
### 4. Removed EntityConflictError catch, then had to restore it
|
|
193
|
-
|
|
194
|
-
The `workflowEntrypoint` error handler originally caught both `EntityConflictError`
|
|
195
|
-
and `RunExpiredError`. When adding the "already-running returns run without event"
|
|
196
|
-
behavior, `EntityConflictError` was removed from the catch since the new worlds
|
|
197
|
-
wouldn't throw it. Reviewer flagged this: old worlds or world-vercel hitting an
|
|
198
|
-
older workflow-server could still throw it. The catch was restored.
|
|
199
|
-
|
|
200
|
-
### 5. Duplicate `startedAt` check
|
|
201
|
-
|
|
202
|
-
After refactoring the `run_started` flow, a `workflowRun.startedAt` null check
|
|
203
|
-
existed both inside the `try` block and after the `catch` block. The second was
|
|
204
|
-
unreachable. Removed after review.
|
|
205
|
-
|
|
206
|
-
### 6. WORKFLOW_SERVER_URL_OVERRIDE left set
|
|
207
|
-
|
|
208
|
-
During development, `WORKFLOW_SERVER_URL_OVERRIDE` was set to a test URL pointing
|
|
209
|
-
at the workflow-server preview deployment and accidentally committed. The Vercel
|
|
210
|
-
bot flagged this. Reset to empty string.
|
|
211
|
-
|
|
212
|
-
### 7. e2e test assertion was wrong
|
|
213
|
-
|
|
214
|
-
The resilient start e2e test stubbed `world.events.create` and asserted
|
|
215
|
-
`createCallCount >= 2`. But the stub only intercepts calls from the test runner
|
|
216
|
-
process — the server uses its own world. `createCallCount` was always 1. Changed
|
|
217
|
-
to `expect(createCallCount).toBe(1)`.
|
|
218
|
-
|
|
219
|
-
### 8. Misattributed Local Prod timeouts as "pre-existing"
|
|
220
|
-
|
|
221
|
-
Local Prod tests showed 60-second timeouts across various tests. Initially dismissed
|
|
222
|
-
as CI flakes. Checking main's CI showed all Local Prod tests pass on main — the
|
|
223
|
-
timeouts are caused by our changes. Should have compared against main immediately.
|
|
224
|
-
|
|
225
|
-
### 9. Attempted to revert parallel dispatch
|
|
226
|
-
|
|
227
|
-
After identifying Local Prod timeouts, `start()` was partially reverted back to
|
|
228
|
-
sequential dispatch. The user pointed out that parallel dispatch is the core value
|
|
229
|
-
proposition of the PR. The revert was undone.
|
|
230
|
-
|
|
231
|
-
### 10. WorkflowRunNotFoundError retry was unconditional
|
|
232
|
-
|
|
233
|
-
The initial `pollReturnValue` retry on `WorkflowRunNotFoundError` applied to all
|
|
234
|
-
`Run` instances. A user calling `getRun()` with a wrong ID would wait 10 seconds
|
|
235
|
-
before getting a 404. Fixed by adding a `resilientStart` flag: only retries when
|
|
236
|
-
`run_created` actually failed.
|
|
237
|
-
|
|
238
|
-
### 11. Changeset `minor` vs `patch`
|
|
239
|
-
|
|
240
|
-
The changeset was created with `"@workflow/core": minor`. Reviewer flagged this as
|
|
241
|
-
violating repo rules ("all changes should be patch"). Changed after discussion.
|
|
242
|
-
|
|
243
|
-
### 12. world-local TOCTOU race causing duplicate `run_created` events (Windows CI)
|
|
244
|
-
|
|
245
|
-
The resilient start path AND the normal `run_created` path in `world-local/events-storage.ts`
|
|
246
|
-
both used `writeJSON` to create the run entity. `writeJSON` checks file existence with
|
|
247
|
-
`fs.access()` then writes via temp+rename — a classic TOCTOU race. On the local world,
|
|
248
|
-
the queue delivers via an async IIFE in the same event loop, so `events.create(run_created)`
|
|
249
|
-
and `events.create(run_started)` (with resilient start) run concurrently:
|
|
250
|
-
|
|
251
|
-
1. Both paths call `fs.access(runPath)` → ENOENT (file doesn't exist yet)
|
|
252
|
-
2. Both proceed to write → the last `fs.rename` wins
|
|
253
|
-
3. Both succeed → both write their own `run_created` event with different event IDs
|
|
254
|
-
4. During replay, the consumer sees two `run_created` events → "Unconsumed event" error
|
|
255
|
-
|
|
256
|
-
This caused consistent failures in `world-testing` embedded tests on Windows CI (`hooks`,
|
|
257
|
-
`supports null bytes in step results`, `retriable and fatal errors` — all timing out at
|
|
258
|
-
60s with "Unconsumed event in event log" errors). Linux CI was not affected because the
|
|
259
|
-
timing was different enough that the race window was rarely hit.
|
|
260
|
-
|
|
261
|
-
Fixed by switching BOTH paths to `writeExclusive` (O_CREAT|O_EXCL), which is atomic at
|
|
262
|
-
the OS level — exactly one writer wins, the other gets EEXIST. The normal `run_created`
|
|
263
|
-
path throws `EntityConflictError` on conflict (handled by `start()` as 409). The resilient
|
|
264
|
-
start path re-reads the run from disk on conflict. Either way, only one `run_created`
|
|
265
|
-
event is written.
|
|
266
|
-
|
|
267
|
-
### 13. Non-atomic run + run_created event in world-postgres resilient path
|
|
268
|
-
|
|
269
|
-
The resilient start path in `world-postgres/storage.ts` did two separate writes (run
|
|
270
|
-
insert, then event insert) without a transaction. If the process crashed between them,
|
|
271
|
-
the run would exist without a `run_created` event — an inconsistent event log.
|
|
272
|
-
|
|
273
|
-
A `drizzle.transaction()` wrapper was attempted but dropped due to TypeScript inference
|
|
274
|
-
issues with drizzle's transaction callback and the insert builder's overloads. The current
|
|
275
|
-
fix keeps the two writes sequential but adds the same conflict-aware re-read pattern as
|
|
276
|
-
world-local: when `onConflictDoNothing` produces no result (run already existed), the run
|
|
277
|
-
is re-read so downstream logic sees the real state. The narrow crash window between the
|
|
278
|
-
two writes is acceptable — if the run insert succeeds but the event insert crashes, the
|
|
279
|
-
run exists and `run_started` will still proceed normally (the event log will be missing a
|
|
280
|
-
`run_created` entry, but the run itself is functional).
|
|
281
|
-
|
|
282
|
-
### 14. Missing `WorkflowRunStatus` span attribute after parallel refactor
|
|
283
|
-
|
|
284
|
-
The `start()` span previously set `Attribute.WorkflowRunStatus(result.run.status)`, but
|
|
285
|
-
this was dropped in the parallel refactor because `result.run` is only available when
|
|
286
|
-
`runCreatedResult` fulfilled. The attribute is now conditionally set when the result is
|
|
287
|
-
available. In the resilient start case (run_created failed), the attribute is omitted
|
|
288
|
-
rather than erroring.
|
|
289
|
-
|
|
290
|
-
### 15. `run_started` eventData leak in world-postgres result
|
|
291
|
-
|
|
292
|
-
The `...data` spread in the result construction leaked `eventData` from `run_started`
|
|
293
|
-
into the returned event object. Storage was already correct (`storedEventData` is
|
|
294
|
-
`undefined` for `run_started`), but the returned result carried the input data. While
|
|
295
|
-
harmless (the runtime doesn't use `result.event.eventData`), it was restored to match
|
|
296
|
-
the pre-refactor behavior where eventData was explicitly stripped from the result.
|
|
297
|
-
|
|
298
|
-
### 16. Normal `run_created` path also needed `writeExclusive` (Windows CI)
|
|
299
|
-
|
|
300
|
-
The initial TOCTOU fix (item 12) only changed the resilient start path to use
|
|
301
|
-
`writeExclusive`. The normal `run_created` entity write still used `writeJSON` which
|
|
302
|
-
checks existence with `fs.access()` then writes via temp+rename — not atomic. On
|
|
303
|
-
Windows CI, the local queue's async IIFE delivered fast enough for both paths to pass
|
|
304
|
-
their existence checks simultaneously, producing two `run_created` events with different
|
|
305
|
-
event IDs. The events consumer saw the duplicate as "Unconsumed event in event log,"
|
|
306
|
-
causing `hooks`, `supports null bytes in step results`, and `retriable and fatal errors`
|
|
307
|
-
tests to time out at 60s. Fixed by also switching the normal `run_created` entity write to
|
|
308
|
-
`writeExclusive`, making both paths use the same atomic gate.
|
|
309
|
-
|
|
310
|
-
### 17. CborTransport was a pass-through wrapper
|
|
311
|
-
|
|
312
|
-
`world-vercel/queue.ts` had `CborTransport` implementing `Transport<Buffer>` with a
|
|
313
|
-
no-op `serialize` (identity function) and a `deserialize` that reassembled chunks into
|
|
314
|
-
a Buffer without decoding. The actual CBOR `encode()`/`decode()` calls happened at the
|
|
315
|
-
call sites — `queue()` pre-encoded before calling `client.send()`, and the handler
|
|
316
|
-
post-decoded after receiving from `client.handleCallback()`. This violated the transport
|
|
317
|
-
abstraction (every other transport does its encoding inside serialize/deserialize) and
|
|
318
|
-
meant the call site had to remember to pre-encode. Refactored to move `encode()`/`decode()`
|
|
319
|
-
into the transport methods and changed the type from `Transport<Buffer>` to
|
|
320
|
-
`Transport<unknown>`.
|
|
321
|
-
|
|
322
|
-
## Follow-up work (additional)
|
|
323
|
-
|
|
324
|
-
- [x] ~~**CborTransport is a pass-through**~~ — Resolved. Moved `encode()`/`decode()`
|
|
325
|
-
into `CborTransport.serialize()`/`CborTransport.deserialize()`. The transport is now
|
|
326
|
-
self-contained: call sites pass plain objects, and the handler receives decoded objects.
|
|
327
|
-
See retrospective item 17.
|
|
75
|
+
The narrow crash window in `world-postgres` between the run insert and the event insert is acceptable — if the run insert succeeds but the event insert crashes, the run exists and `run_started` will still proceed normally (the event log will be missing a `run_created` entry, but the run itself is functional).
|
|
@@ -0,0 +1,87 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: Turbo mode (fast first invocation)
|
|
3
|
+
description: Fast-path the very first delivery of the first invocation — background run_started, skip the initial event-log load, and force optimistic inline start — so a run blazes through its first steps. A no-op for everything else.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Turbo mode
|
|
7
|
+
|
|
8
|
+
## Motivation
|
|
9
|
+
|
|
10
|
+
The first invocation of a workflow run is where time-to-first-step matters most, yet it pays the most fixed network latency before any user code runs. Three round-trips sit on that critical path today:
|
|
11
|
+
|
|
12
|
+
1. **`run_started` is awaited.** The handler writes `run_started` and waits for it to return the run entity before doing anything else.
|
|
13
|
+
2. **The event log is loaded.** A full `events.list` runs before the first replay — even though on the very first delivery nothing has written any events yet.
|
|
14
|
+
3. **Optimistic inline start is off by default.** The [optimistic inline start](./lazy-event-creation#optimistic-inline-start-opt-in-off-by-default) optimization (running a step body before its `step_started` is confirmed) is off by default because under contention two handlers can both run a body and corrupt non-idempotent side effects.
|
|
15
|
+
|
|
16
|
+
Turbo mode removes all three costs **for the first delivery of the first invocation only**, where each is provably safe to remove, then gets out of the way. For every subsequent invocation it is a complete no-op.
|
|
17
|
+
|
|
18
|
+
## What turbo mode does
|
|
19
|
+
|
|
20
|
+
When the handler detects the first delivery of the first invocation, it:
|
|
21
|
+
|
|
22
|
+
1. **Backgrounds `run_started`.** The event is written without awaiting; the run entity is synthesized locally from the queued run input (status `running`, `startedAt` now) so replay can begin immediately. The `run_started` round-trip overlaps replay instead of blocking it. This reuses the [resilient start](./resilient-start) contract — `run_started` carrying the run input creates the run on the fly (synthetic `run_created`) if it doesn't exist yet.
|
|
23
|
+
2. **Skips the initial event-log load.** Nothing has been written, so the first replay runs against an empty log. The second loop iteration does a normal incremental load once the first step's events exist.
|
|
24
|
+
3. **Forces optimistic inline start** for that invocation, independent of `WORKFLOW_OPTIMISTIC_INLINE_START`. The step body runs immediately against locally-synthesized state; only the `step_started` network write waits for the backgrounded `run_started`.
|
|
25
|
+
|
|
26
|
+
The net effect: the first step body starts after just the in-process replay, with `run_started` and `step_started` happening in the background around it, and no `events.list` before it.
|
|
27
|
+
|
|
28
|
+
## Why this is safe (and where it stops)
|
|
29
|
+
|
|
30
|
+
### Detection
|
|
31
|
+
|
|
32
|
+
The first-invocation message is the only one that carries the queued **run input**, and the queue delivery **attempt is 1** (a redelivery is attempt ≥ 2). Together with "not a background-step invocation" and "not a divergence recovery", that uniquely identifies the first delivery of the first invocation — with no new message field and no world/backend change.
|
|
33
|
+
|
|
34
|
+
### The single-handler guarantee
|
|
35
|
+
|
|
36
|
+
Forcing optimistic start is unsafe *in general* because two handlers racing the same step's create-claim can both run the body before one wins. On the first delivery of the first invocation there is **no concurrent peer handler** — the run was created moments ago by `start()` and only this one message is in flight. So the body runs exactly once, and forcing optimistic start is safe here even though the global flag is off.
|
|
37
|
+
|
|
38
|
+
### Turbo exits on the first hook or wait
|
|
39
|
+
|
|
40
|
+
That single-handler guarantee ends the moment the run creates a **hook** or **wait** (or writes attributes): those introduce later resume/parallel invocations that *can* race. So turbo stops forcing optimistic start as soon as a suspension creates any of them — the inline steps of that suspension fall back to the normal await-then-run path, and the rest of the run behaves exactly as it does today. A pure-step suspension (the common hot path) stays on the fast path.
|
|
41
|
+
|
|
42
|
+
### Write ordering is preserved
|
|
43
|
+
|
|
44
|
+
Because `run_started` is backgrounded, every event write is gated on a run-ready barrier so nothing is written before the run exists:
|
|
45
|
+
|
|
46
|
+
- The optimistic `step_started` is **chained** on the barrier — the body still runs immediately, only the network write waits.
|
|
47
|
+
- The suspension handler **awaits** the barrier before any eager write (`hook_created`, `wait_created`, overflow `step_created`). The pure inline hot path defers all its steps and writes nothing here, so it never blocks on the barrier.
|
|
48
|
+
- Terminal run writes (`run_completed` / `run_failed`) await the barrier too, so a workflow that finishes with no steps still orders its completion after `run_started`.
|
|
49
|
+
|
|
50
|
+
The event log therefore still reads `run_created → run_started → step_created → step_started → step_completed`. If the backgrounded `run_started` genuinely fails (e.g. the run was cancelled in the meantime), the chained writes surface the real error (`gone` / run-not-found) and the message redelivers as a normal, non-turbo attempt.
|
|
51
|
+
|
|
52
|
+
The barrier orders **event** writes. The forced-optimistic first step **body** runs immediately, so any side effects it performs *before* the terminal write — stream writes via `getWritable()` and the per-step ops flush — are **not** gated on the barrier and can reach the world before the backgrounded `run_started` lands (and are orphaned if it ultimately fails). This is the same exposure as optimistic inline start and is covered by the stream-safety caveat below; deployments whose first step writes to the workflow stream and require strict `run_created → run_started` ordering of stream data should set `WORKFLOW_TURBO=0`.
|
|
53
|
+
|
|
54
|
+
### A run cancelled before its first delivery still runs the first step body
|
|
55
|
+
|
|
56
|
+
The non-turbo path awaits `run_started` up front and, if the run was cancelled or expired between `start()` and this delivery, returns before any workflow/step code runs. Turbo synthesizes `status: 'running'` and runs the first step body optimistically, so such a cancellation is only observed when the backgrounded `run_started` (and the barrier-chained `step_started`) rejects — *after* the body's side effects have executed (they are then discarded via reconciliation). For non-idempotent first steps this is the same "body runs before ownership is confirmed" tradeoff as optimistic inline start; `WORKFLOW_TURBO=0` restores the up-front skip.
|
|
57
|
+
|
|
58
|
+
### `workflowStartedAt` reflects the first delivery's clock
|
|
59
|
+
|
|
60
|
+
Replay matching — step/wait/hook correlation IDs, the VM seed, and the in-VM `Date.now()` — is derived from a replay-stable timestamp recovered from the run ID, so it does **not** depend on `startedAt` and is identical on every delivery. The one value that still tracks `startedAt` is the user-facing `getWorkflowMetadata().workflowStartedAt`: under turbo the first delivery synthesizes it from the local clock, while a later (non-turbo) delivery loads the server-canonical `startedAt`, so the two can differ by the start→first-delivery latency. Treat `workflowStartedAt` as an approximate, human-facing timestamp — do **not** branch workflow control flow on it (e.g. `Date.now() - +workflowStartedAt > threshold`), since that can take different paths across deliveries and diverge on replay. For timing logic that must survive replay, use the in-VM `Date.now()` / `new Date()`, which is replay-stable.
|
|
61
|
+
|
|
62
|
+
### Attributes seeded at `start()` survive the skipped event load
|
|
63
|
+
|
|
64
|
+
`start({ attributes })` does **not** disable turbo, and it needs no synthetic event in the empty log. Seed attributes are folded into the `run_created` event's data (not separate `attr_set` events) and ride along in the queued run input, so the locally-synthesized run snapshot carries them — turbo skipping the initial `events.list` loses nothing.
|
|
65
|
+
|
|
66
|
+
This is safe specifically because **attributes are write-only inside a workflow**: there is no in-workflow read API today, and `run_created` is consumed structurally during replay without inspecting its attributes. So an empty initial event log replays identically whether or not the run was seeded with attributes.
|
|
67
|
+
|
|
68
|
+
That safety is a standing invariant for any future change: if an in-workflow attribute *read* API is ever added, it MUST read from the run snapshot (which turbo populates from the run input) and **not** by replaying `run_created` / `attr_set` events. Reading from the event log would surface seed attributes as empty on the first turbo delivery only — a turbo-exclusive divergence from the non-turbo path. `start()` cannot seed hooks or waits, so there is no start-seeded suspension state for the skipped load to miss.
|
|
69
|
+
|
|
70
|
+
## Configuration
|
|
71
|
+
|
|
72
|
+
Turbo mode is **on by default**. Set `WORKFLOW_TURBO=0` (or `false`) to disable it — every invocation then takes the existing awaited path. This is a useful kill-switch for deployments whose first-step bodies are not idempotent and stream-safe (the same caveat as optimistic inline start), or for isolating behavior while debugging.
|
|
73
|
+
|
|
74
|
+
Turbo forces optimistic inline start on the first invocation regardless of `WORKFLOW_OPTIMISTIC_INLINE_START` (its single-handler guarantee removes the double-execution race that flag guards against). It does, however, **honor an explicit `WORKFLOW_OPTIMISTIC_INLINE_START=0`**: because forced optimistic start still runs the body before `step_started`/`run_started` is confirmed, an operator who has explicitly disabled optimistic start keeps the await-then-run path even under turbo (the rest of turbo — backgrounded `run_started`, skipped initial load — still applies). With the flag unset (the default), turbo forces it on.
|
|
75
|
+
|
|
76
|
+
Turbo mode is purely client-side and builds on the lazy/optimistic inline start support already shipped — it requires no world or backend changes.
|
|
77
|
+
|
|
78
|
+
## Considered: running ahead of durable writes (not implemented)
|
|
79
|
+
|
|
80
|
+
Turbo overlaps the *start* round-trips with a step's body, but it still **awaits each `step_completed` before advancing** to the next step. We explored going further — "run-ahead": within a single invocation, execute the workflow forward across a sequential chain *without* awaiting each step's event writes, draining `step_started`/`step_completed` through a background FIFO queue and only blocking on a full drain before acking. A run of three sub-millisecond steps would then fire all the bodies back-to-back while the six event posts caught up in the background, turning per-step latency into `max(Σ body, Σ post)` instead of `Σ(body + post)`.
|
|
81
|
+
|
|
82
|
+
We decided **not** to ship it, for two reasons:
|
|
83
|
+
|
|
84
|
+
1. **Re-execution blast radius on failure.** Awaiting each completion means a crash re-runs essentially one in-flight step. Running ahead leaves many completions undrained at once, so a crash or `maxDuration` SIGTERM re-runs *all* of them on redelivery — a much larger at-least-once blast radius, precisely on the latency-sensitive runs most likely to pack many steps into one invocation.
|
|
85
|
+
2. **Divergent branches from non-durable results.** Advancing past a step before its result is durable lets the workflow commit to a forward path that a crash-and-redeliver can re-decide differently. A `Promise.race([B, C])` resolved by local timing can pick `B`, run `D(B)`, then crash before `step_completed_B` is durable — and the redelivery may re-resolve to `C`, so `D` executed against a winner the durable history never records. The same shape appears for a branch on a non-deterministic step output (`B(v1)` runs, crash, redelivery commits `B(v2)`). Idempotency doesn't cover these — `D(B)`/`D(C)` and `B(v1)`/`B(v2)` are *different* operations, not retries of one. A "run ahead only while at most one result is undurable" gate would contain the race case (a race needs ≥2 concurrent undurable steps) but not the non-deterministic-output case, and that residual hazard plus the re-execution blast radius outweighed the gain.
|
|
86
|
+
|
|
87
|
+
So turbo deliberately stops at forced-optimistic *start* and awaits each `step_completed` before moving on: re-execution after a crash stays deterministic (each step re-runs against the same durable inputs) and bounded (roughly one step, not the whole chain). The idea is recorded here in case a future change (e.g. a determinism signal on steps, or deterministic race resolution) makes run-ahead safe enough to revisit.
|