@agentproto/workflow-runtime 0.12.0 → 0.13.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -55,13 +55,16 @@ const { output, bindings } = await runWorkflow({ workflow: wf })
55
55
 
56
56
  ## Step kinds
57
57
 
58
- Every step reads `Bindings` (`{ input, steps, item?, index? }`) via selectors
59
- and binds its output under `id`.
58
+ Every step reads `Bindings` (`{ input, steps, item?, index?, run? }`) via
59
+ selectors and binds its output under `id`. `run` (`{ workspace }`) is present
60
+ whenever the host wires `RunWorkflowArgs.workspace` — see **Run workspace
61
+ (AIP-58 §4)** below.
60
62
 
61
63
  | kind | what it does |
62
64
  | --- | --- |
63
65
  | `tool` | Resolve a TOOL contract against `candidates` and run it. Cacheable. |
64
66
  | `agent` | Spawn/reuse an agent session, send a prompt, wait for the turn. Cacheable. |
67
+ | `artifact` | Hash + copy a file out of `$run.workspace` into `artifactsDir`, recorded as an `ArtifactEntry`. Always cache-aware — see **Run workspace**. |
65
68
  | `transform` | Pure in-run compute — combine / filter / shape. No dispatch. |
66
69
  | `map` | Run a body once per array element (`bindings.item` / `index`); collect outputs. Fail-fast by default; `onError: "collect"` opts into per-item tolerance. |
67
70
  | `pipeline` | Run N items through K sequential stages with **no cross-item barrier**. Same `onError: "collect"` opt-in as `map`. |
@@ -73,6 +76,49 @@ and binds its output under `id`.
73
76
  | `group` | Run a list of steps as one unit; output is the last step's. |
74
77
  | `subworkflow` | Run a nested workflow with its own isolated bindings. |
75
78
 
79
+ ### WORKFLOW.md `kind: branch` — exclusive arms + join
80
+
81
+ `compileWorkflow` compiles a declarative AIP-15 `kind: branch` step (goto-style
82
+ `branches[].next` / `default`, forward-only: every target is a LATER sibling in
83
+ the same step list) into nested runtime `branch` nodes. Arms are **exclusive**:
84
+
85
+ - sorted by position, the arm targets split the siblings after the branch into
86
+ arm bodies — an arm's body is its target step up to (not including) the next
87
+ arm's target; the last arm's body runs up to the **join**;
88
+ - the join is the `join:` sibling if declared, else the step right after the
89
+ last arm's target (so without `join:` the last arm is exactly one step);
90
+ - exactly one body runs (first truthy `when`, else `default`'s), then
91
+ execution continues at the join — every step from the join on runs once,
92
+ whichever arm was taken;
93
+ - no `default` ⇒ a no-match runs the steps between the branch and its first
94
+ arm target (usually none) and continues at the join;
95
+ - the untaken arms' steps are reported through `onStepSkipped` (the daemon
96
+ surfaces them as `skipped` in `workflow_status` and as AIP-58 `step.skipped`
97
+ events).
98
+
99
+ ```yaml
100
+ steps:
101
+ - id: maybe-render-pdf
102
+ kind: branch
103
+ branches:
104
+ - when: $input.exportPdf
105
+ next: pdf-render # arm 1 body: pdf-render, pdf-upload
106
+ join: publish # optional — omit and the join is the step after the last arm
107
+ - id: pdf-render
108
+ kind: tool
109
+ tool: pdf.render
110
+ - id: pdf-upload
111
+ kind: tool
112
+ tool: pdf.upload
113
+ - id: publish # runs once, whether or not the PDF arm ran
114
+ kind: tool
115
+ tool: site.publish
116
+ ```
117
+
118
+ `fallthrough: true` restores the legacy (pre-exclusive) semantics — the chosen
119
+ target and EVERY sibling after it run, so an earlier arm also runs every later
120
+ arm. It is incompatible with `join`; prefer the exclusive form.
121
+
76
122
  ## Harness-parity capabilities
77
123
 
78
124
  The engine reaches parity with a code-first agent harness across five axes.
@@ -224,18 +270,167 @@ await runWorkflow({ workflow: wf, cache, cacheKey: "nightly-review" }) // run 2:
224
270
  Both `cache` and `cacheKey` must be set for any caching to happen. Supply your
225
271
  own `StepCache` (`{ get, set }`) for an in-memory or custom-backed journal.
226
272
 
273
+ Both a declarative `kind: "tool"` step and a `kind: "agent"` step accept
274
+ `cacheable: true` — in WORKFLOW.md frontmatter as well as a TS-authored step.
275
+
276
+ Inside a `map`, each item caches **independently**: the journal key includes
277
+ the item's `[<index>]` path (nested maps append their own index), so item 1
278
+ changing doesn't invalidate items 0 and 2, and every item of a >1-item map
279
+ gets its own cache entry instead of all items sharing (and stomping) one key.
280
+
281
+ #### Re-running a failed run from the same cacheKey
282
+
283
+ Because the journal is written per step as the run progresses — not only at
284
+ the end — a run that throws partway through still leaves every already-
285
+ succeeded cacheable step's output in the journal. Re-invoking `runWorkflow`
286
+ with the SAME `workflow`, `input`, `cache`, and `cacheKey` after a failure
287
+ replays every step whose resolved inputs are unchanged and only re-executes
288
+ the step(s) that failed (or whose resolved inputs changed since the failed
289
+ run):
290
+
291
+ ```ts
292
+ try {
293
+ await runWorkflow({ workflow: wf, input, cache, cacheKey: "run-42" })
294
+ } catch {
295
+ // fix the underlying issue, then re-run with the SAME cacheKey —
296
+ // every cacheable step that already succeeded replays; only the
297
+ // step that failed (and anything downstream of it) re-executes.
298
+ await runWorkflow({ workflow: wf, input, cache, cacheKey: "run-42" })
299
+ }
300
+ ```
301
+
302
+ This is a manual retry, not a resumable run object — `runWorkflow` has no
303
+ notion of "the run that failed"; the cacheKey is just a namespace the caller
304
+ re-supplies. A first-class `run.retry`/`run.replay` verb that resumes a named
305
+ run without the caller re-threading `workflow`/`input`/`cacheKey` by hand is
306
+ AIP-58 P5, not implemented here.
307
+
308
+ ### Run workspace (AIP-58 §4)
309
+
310
+ A host wires `workspace` (an absolute directory) and `artifactsDir` (a
311
+ sibling directory) onto `RunWorkflowArgs` — one allocation per run,
312
+ `<runsRoot>/<runId>/{scratch,artifacts}` in `@agentproto/runtime`'s
313
+ `WorkflowRunner`. Steps read `workspace` two ways:
314
+
315
+ - **`$run.workspace`** / **`{{run.workspace}}`** — the same `$…`-ref and
316
+ mustache grammar `$input`/`$steps.<id>` already use, so it drops straight
317
+ into a `tool` step's `inputs`, a `gate` step's `args`/`cwd`, or an `agent`
318
+ step's `prompt`.
319
+ - **`_workflowFsRoot`** — the AIP-16 file-contract convention (`inputs
320
+ ._workflowFsRoot`, injected by the host); use this form when the same
321
+ manifest needs to stay AIP-16-conformant independent of this runtime.
322
+
323
+ Both name the identical path — pick whichever reads better at each call
324
+ site. An `agent` step's own **cwd is unaffected** (AIP-15 F25: it stays the
325
+ owning app's root) — the workspace path only reaches it as *text*, via the
326
+ binding in its prompt.
327
+
328
+ ```yaml
329
+ steps:
330
+ - id: fetch
331
+ kind: tool
332
+ tool: yt.fetch-captions
333
+ inputs:
334
+ url: $input.url
335
+ outDir: $run.workspace # was: $input.outDir (one fixed, reused dir)
336
+ - id: clean
337
+ kind: agent
338
+ prompt: >-
339
+ Clean {{item}} and write the result under {{run.workspace}}/cleaned/.
340
+ ```
341
+
342
+ #### `kind: "artifact"` — declaring a run artifact
343
+
344
+ ```ts
345
+ { kind: "artifact", id: "save", key: "brief", path: "briefs/latest.md", contentType: "text/markdown" }
346
+ ```
347
+
348
+ `path` is read relative to `$run.workspace` (absolute paths, and any path
349
+ that would resolve OUTSIDE the workspace, throw). The step hashes
350
+ (`sha256`) and sizes the file, copies it to `artifactsDir/<sanitized key>`,
351
+ and binds/report an `ArtifactEntry` — `{ key, path: "artifacts/<key>", sha256,
352
+ size, stepId, contentType? }` (`path` here is relative to the RUN WORKSPACE
353
+ ROOT, the parent of `$run.workspace` itself — not the source location).
354
+ Pass `onArtifact` to `runWorkflow` to observe every one recorded, cache hit
355
+ or fresh.
356
+
357
+ A declarative WORKFLOW.md manifest may instead declare **`outputsFiles`**
358
+ (AIP-16, amended by AIP-58 §4 with `required`) at the top level — checked
359
+ ONCE, after every top-level step finishes:
360
+
361
+ ```yaml
362
+ outputsFiles:
363
+ brief:
364
+ path: "./briefs/<runId>.md" # <runId> / <workflowId> / <isoDate> interpolate
365
+ required: true
366
+ ```
367
+
368
+ Present ⇒ copied into `artifactsDir` + reported, same as a `kind: "artifact"`
369
+ step. Absent with `required: true` ⇒ the run throws `MissingArtifactError`
370
+ (`code: "missing-artifact"`, attributed to the last step that ran). Absent
371
+ with `required: false` **or omitted (the default — `required` is opt-in)**
372
+ ⇒ a `console.warn`, the run still succeeds.
373
+
374
+ **This is where a shared/fixed destination stops racing.** The
375
+ Motivation this AIP exists for — two concurrent runs of the same workflow
376
+ overwriting each other's output — is closed by `artifactsDir` being
377
+ per-run: `outputsFiles.<key>.path`'s `<runId>`-interpolated form (or a bare
378
+ fixed path) is *never* written to directly; it becomes the **default
379
+ destination for an explicit `run.publish`** (`@agentproto/runtime`'s
380
+ `WorkflowRunner.publish` / the `workflow_publish` MCP tool), which a host
381
+ MUST refuse unless the run has already `"succeeded"`. See
382
+ `@agentproto/runtime`'s own docs for `publish`/`readArtifact` (the "fetchable"
383
+ half) and the compact `{key, path, size}` projection `workflow_status` shows.
384
+
385
+ #### Cache/replay interplay — a design decision
386
+
387
+ A cache hit (`cacheable: true` / journaled) replays a step's **recorded
388
+ output** without re-executing it — but if that output only meant something
389
+ because a FILE existed at some workspace-relative path, and this run's own
390
+ workspace is a **different directory** than the run that originally wrote
391
+ it (AIP-58 §4 "two runs MUST NEVER share a workspace" — replay/re-invocation
392
+ under the same `cacheKey` is always a fresh, disjoint `runId`), a naive
393
+ cache hit would report an output pointing at a file that was never actually
394
+ written into THIS run's workspace.
395
+
396
+ `kind: "artifact"` (and the `outputsFiles` check, which is really the same
397
+ mechanism unrolled) closes this the way the AIP steers hosts to think about
398
+ runs generally: **copy, never share**. A cache hit for an `artifact` step
399
+ relocates (copies) the file from the ORIGINAL run's `artifactsDir` into the
400
+ CURRENT run's own before returning — the two runs' workspaces stay fully
401
+ disjoint on disk; only the *bytes* are reused, exactly the same way
402
+ `run.replay`'s journal-sourced step reuse is specified to work (AIP-58 §6).
403
+ This is a deliberate, narrower scope than "any cacheable step's output might
404
+ reference a workspace file" — a plain cacheable `tool`/`agent` step whose
405
+ output happens to name a path is NOT relocated automatically; route a
406
+ step's file-shaped output through `kind: "artifact"` (or `outputsFiles`) to
407
+ get cache/replay continuity for it.
408
+
227
409
  ### Step lifecycle callbacks
228
410
 
229
- Pass `onStepStart` and `onStepComplete` to `runWorkflow` to observe progress in
411
+ Pass `onStepStart`, `onStepComplete` and `onStepSkipped` to `runWorkflow` to observe progress in
230
412
  real time. The callbacks fire for every step kind; for `agent` steps, start fires
231
413
  before spawn and complete fires after the turn (and any output-schema retry loop)
232
414
  finishes.
233
415
 
416
+ A cacheable step replayed from the journal still fires both callbacks (it
417
+ doesn't vanish from progress), with a third `info` argument of
418
+ `{ cached: true }`; an executed step gets `info === undefined`. The daemon's
419
+ `workflow_status` surfaces this as `cached: true` on the step row, and on the
420
+ AIP-58 `step.started`/`step.succeeded` events' `data`.
421
+
422
+ `onStepSkipped(stepId, { reason: "branch-not-taken", branchId })` fires, once a
423
+ `branch` decides, for every statically-known step in the arms it did NOT take
424
+ (a `map`/`pipeline`/`subworkflow` step reports its own id; a step id that also
425
+ sits on the taken path is never reported). Inside a `map` item the id is
426
+ indexed (`<id>[<index>]`) like the other two callbacks.
427
+
234
428
  ```ts
235
429
  await runWorkflow({
236
430
  workflow: wf,
237
- onStepStart: (stepId) => console.log("starting", stepId),
238
- onStepComplete: (stepId, output) => console.log("done", stepId, output),
431
+ onStepStart: (stepId, info) => console.log("starting", stepId, info?.cached ? "(cached)" : ""),
432
+ onStepComplete: (stepId, output, info) => console.log("done", stepId, output, info?.cached ? "(cached)" : ""),
433
+ onStepSkipped: (stepId, info) => console.log("skipped", stepId, "by", info.branchId),
239
434
  })
240
435
  ```
241
436