@agentproto/workflow-runtime 0.12.0 → 0.13.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -55,13 +55,16 @@ const { output, bindings } = await runWorkflow({ workflow: wf })
55
55
 
56
56
  ## Step kinds
57
57
 
58
- Every step reads `Bindings` (`{ input, steps, item?, index? }`) via selectors
59
- and binds its output under `id`.
58
+ Every step reads `Bindings` (`{ input, steps, item?, index?, run? }`) via
59
+ selectors and binds its output under `id`. `run` (`{ workspace }`) is present
60
+ whenever the host wires `RunWorkflowArgs.workspace` — see **Run workspace
61
+ (AIP-58 §4)** below.
60
62
 
61
63
  | kind | what it does |
62
64
  | --- | --- |
63
65
  | `tool` | Resolve a TOOL contract against `candidates` and run it. Cacheable. |
64
66
  | `agent` | Spawn/reuse an agent session, send a prompt, wait for the turn. Cacheable. |
67
+ | `artifact` | Hash + copy a file out of `$run.workspace` into `artifactsDir`, recorded as an `ArtifactEntry`. Always cache-aware — see **Run workspace**. |
65
68
  | `transform` | Pure in-run compute — combine / filter / shape. No dispatch. |
66
69
  | `map` | Run a body once per array element (`bindings.item` / `index`); collect outputs. Fail-fast by default; `onError: "collect"` opts into per-item tolerance. |
67
70
  | `pipeline` | Run N items through K sequential stages with **no cross-item barrier**. Same `onError: "collect"` opt-in as `map`. |
@@ -73,6 +76,49 @@ and binds its output under `id`.
73
76
  | `group` | Run a list of steps as one unit; output is the last step's. |
74
77
  | `subworkflow` | Run a nested workflow with its own isolated bindings. |
75
78
 
79
+ ### WORKFLOW.md `kind: branch` — exclusive arms + join
80
+
81
+ `compileWorkflow` compiles a declarative AIP-15 `kind: branch` step (goto-style
82
+ `branches[].next` / `default`, forward-only: every target is a LATER sibling in
83
+ the same step list) into nested runtime `branch` nodes. Arms are **exclusive**:
84
+
85
+ - sorted by position, the arm targets split the siblings after the branch into
86
+ arm bodies — an arm's body is its target step up to (not including) the next
87
+ arm's target; the last arm's body runs up to the **join**;
88
+ - the join is the `join:` sibling if declared, else the step right after the
89
+ last arm's target (so without `join:` the last arm is exactly one step);
90
+ - exactly one body runs (first truthy `when`, else `default`'s), then
91
+ execution continues at the join — every step from the join on runs once,
92
+ whichever arm was taken;
93
+ - no `default` ⇒ a no-match runs the steps between the branch and its first
94
+ arm target (usually none) and continues at the join;
95
+ - the untaken arms' steps are reported through `onStepSkipped` (the daemon
96
+ surfaces them as `skipped` in `workflow_status` and as AIP-58 `step.skipped`
97
+ events).
98
+
99
+ ```yaml
100
+ steps:
101
+ - id: maybe-render-pdf
102
+ kind: branch
103
+ branches:
104
+ - when: $input.exportPdf
105
+ next: pdf-render # arm 1 body: pdf-render, pdf-upload
106
+ join: publish # optional — omit and the join is the step after the last arm
107
+ - id: pdf-render
108
+ kind: tool
109
+ tool: pdf.render
110
+ - id: pdf-upload
111
+ kind: tool
112
+ tool: pdf.upload
113
+ - id: publish # runs once, whether or not the PDF arm ran
114
+ kind: tool
115
+ tool: site.publish
116
+ ```
117
+
118
+ `fallthrough: true` restores the legacy (pre-exclusive) semantics — the chosen
119
+ target and EVERY sibling after it run, so an earlier arm also runs every later
120
+ arm. It is incompatible with `join`; prefer the exclusive form.
121
+
76
122
  ## Harness-parity capabilities
77
123
 
78
124
  The engine reaches parity with a code-first agent harness across five axes.
@@ -224,18 +270,204 @@ await runWorkflow({ workflow: wf, cache, cacheKey: "nightly-review" }) // run 2:
224
270
  Both `cache` and `cacheKey` must be set for any caching to happen. Supply your
225
271
  own `StepCache` (`{ get, set }`) for an in-memory or custom-backed journal.
226
272
 
273
+ Both a declarative `kind: "tool"` step and a `kind: "agent"` step accept
274
+ `cacheable: true` — in WORKFLOW.md frontmatter as well as a TS-authored step.
275
+
276
+ Inside a `map`, each item caches **independently**: the journal key includes
277
+ the item's `[<index>]` path (nested maps append their own index), so item 1
278
+ changing doesn't invalidate items 0 and 2, and every item of a >1-item map
279
+ gets its own cache entry instead of all items sharing (and stomping) one key.
280
+
281
+ #### Re-running a failed run from the same cacheKey
282
+
283
+ Because the journal is written per step as the run progresses — not only at
284
+ the end — a run that throws partway through still leaves every already-
285
+ succeeded cacheable step's output in the journal. Re-invoking `runWorkflow`
286
+ with the SAME `workflow`, `input`, `cache`, and `cacheKey` after a failure
287
+ replays every step whose resolved inputs are unchanged and only re-executes
288
+ the step(s) that failed (or whose resolved inputs changed since the failed
289
+ run):
290
+
291
+ ```ts
292
+ try {
293
+ await runWorkflow({ workflow: wf, input, cache, cacheKey: "run-42" })
294
+ } catch {
295
+ // fix the underlying issue, then re-run with the SAME cacheKey —
296
+ // every cacheable step that already succeeded replays; only the
297
+ // step that failed (and anything downstream of it) re-executes.
298
+ await runWorkflow({ workflow: wf, input, cache, cacheKey: "run-42" })
299
+ }
300
+ ```
301
+
302
+ This is a manual retry, not a resumable run object — `runWorkflow` has no
303
+ notion of "the run that failed"; the cacheKey is just a namespace the caller
304
+ re-supplies. A first-class `run.retry` verb that resumes a named run without
305
+ the caller re-threading `workflow`/`input`/`cacheKey` by hand — AIP-58 P5 —
306
+ is a HOST concern, not something this transport-agnostic package implements
307
+ itself: see `@agentproto/runtime`'s `WorkflowRunner.retry()` / the
308
+ `workflow_retry` MCP tool, which owns runId allocation and an always-on
309
+ internal journal (so it works even when the original run never passed a
310
+ `cacheKey` at all — every step it runs is journaled internally either way).
311
+
312
+ ### Run workspace (AIP-58 §4)
313
+
314
+ A host wires `workspace` (an absolute directory) and `artifactsDir` (a
315
+ sibling directory) onto `RunWorkflowArgs` — one allocation per run,
316
+ `<runsRoot>/<runId>/{scratch,artifacts}` in `@agentproto/runtime`'s
317
+ `WorkflowRunner`. Steps read `workspace` two ways:
318
+
319
+ - **`$run.workspace`** / **`{{run.workspace}}`** — the same `$…`-ref and
320
+ mustache grammar `$input`/`$steps.<id>` already use, so it drops straight
321
+ into a `tool` step's `inputs`, a `gate` step's `args`/`cwd`, or an `agent`
322
+ step's `prompt`.
323
+ - **`_workflowFsRoot`** — the AIP-16 file-contract convention (`inputs
324
+ ._workflowFsRoot`, injected by the host); use this form when the same
325
+ manifest needs to stay AIP-16-conformant independent of this runtime.
326
+
327
+ Both name the identical path — pick whichever reads better at each call
328
+ site. An `agent` step's own **cwd is unaffected** (AIP-15 F25: it stays the
329
+ owning app's root) — the workspace path only reaches it as *text*, via the
330
+ binding in its prompt.
331
+
332
+ ```yaml
333
+ steps:
334
+ - id: fetch
335
+ kind: tool
336
+ tool: yt.fetch-captions
337
+ inputs:
338
+ url: $input.url
339
+ outDir: $run.workspace # was: $input.outDir (one fixed, reused dir)
340
+ - id: clean
341
+ kind: agent
342
+ prompt: >-
343
+ Clean {{item}} and write the result under {{run.workspace}}/cleaned/.
344
+ ```
345
+
346
+ #### `kind: "artifact"` — declaring a run artifact
347
+
348
+ ```ts
349
+ { kind: "artifact", id: "save", key: "brief", path: "briefs/latest.md", contentType: "text/markdown" }
350
+ ```
351
+
352
+ `path` is read relative to `$run.workspace` (absolute paths, and any path
353
+ that would resolve OUTSIDE the workspace, throw). The step hashes
354
+ (`sha256`) and sizes the file, copies it to `artifactsDir/<basename of path,
355
+ sanitized>` — `path: "briefs/latest.md"` above lands at
356
+ `artifacts/latest.md`, keeping the extension, NOT the bare key
357
+ (`artifacts/brief`) — and binds/reports an `ArtifactEntry` — `{ key, path:
358
+ "artifacts/<basename>", sha256, size, stepId, contentType? }` (`path` here is
359
+ relative to the RUN WORKSPACE ROOT, the parent of `$run.workspace` itself —
360
+ not the source location; always read it back rather than assuming a name).
361
+ If two keys' files share a basename, the second one claimed is prefixed with
362
+ its own sanitized key (`artifacts/<key>-<basename>`) instead of silently
363
+ overwriting the first — deterministic, same result run to run. Pass
364
+ `onArtifact` to `runWorkflow` to observe every one recorded, cache hit or
365
+ fresh.
366
+
367
+ A declarative WORKFLOW.md manifest may instead declare **`outputsFiles`**
368
+ (AIP-16, amended by AIP-58 §4 with `required`) at the top level — checked
369
+ ONCE, after every top-level step finishes:
370
+
371
+ ```yaml
372
+ outputsFiles:
373
+ brief:
374
+ path: "./briefs/<runId>.md" # <runId> / <workflowId> / <isoDate> interpolate
375
+ required: true
376
+ ```
377
+
378
+ Present ⇒ copied into `artifactsDir` + reported, same as a `kind: "artifact"`
379
+ step. Absent with `required: true` ⇒ the run throws `MissingArtifactError`
380
+ (`code: "missing-artifact"`, attributed to the last step that ran). Absent
381
+ with `required: false` **or omitted (the default — `required` is opt-in)**
382
+ ⇒ a `console.warn`, the run still succeeds.
383
+
384
+ **This is where a shared/fixed destination stops racing.** The
385
+ Motivation this AIP exists for — two concurrent runs of the same workflow
386
+ overwriting each other's output — is closed by `artifactsDir` being
387
+ per-run: `outputsFiles.<key>.path`'s `<runId>`-interpolated form (or a bare
388
+ fixed path) is *never* written to directly; it becomes the **default
389
+ destination for an explicit `run.publish`** (`@agentproto/runtime`'s
390
+ `WorkflowRunner.publish` / the `workflow_publish` MCP tool), which a host
391
+ MUST refuse unless the run has already `"succeeded"`. See
392
+ `@agentproto/runtime`'s own docs for `publish`/`readArtifact` (the "fetchable"
393
+ half) and the compact `{key, path, size}` projection `workflow_status` shows.
394
+
395
+ #### Cache/replay interplay — a design decision
396
+
397
+ A cache hit (`cacheable: true` / journaled) replays a step's **recorded
398
+ output** without re-executing it — but if that output only meant something
399
+ because a FILE existed at some workspace-relative path, and this run's own
400
+ workspace is a **different directory** than the run that originally wrote
401
+ it (AIP-58 §4 "two runs MUST NEVER share a workspace" — replay/re-invocation
402
+ under the same `cacheKey` is always a fresh, disjoint `runId`), a naive
403
+ cache hit would report an output pointing at a file that was never actually
404
+ written into THIS run's workspace.
405
+
406
+ `kind: "artifact"` (and the `outputsFiles` check, which is really the same
407
+ mechanism unrolled) closes this the way the AIP steers hosts to think about
408
+ runs generally: **copy, never share**. A cache hit for an `artifact` step
409
+ relocates (copies) the file from the ORIGINAL run's `artifactsDir` into the
410
+ CURRENT run's own before returning — the two runs' workspaces stay fully
411
+ disjoint on disk; only the *bytes* are reused, exactly the same way
412
+ `run.replay`'s journal-sourced step reuse is specified to work (AIP-58 §6).
413
+
414
+ A plain cacheable `tool`/`agent` step gets the same treatment, not a narrower
415
+ one (this used to be a documented gap — see "Cache key vs. the run
416
+ workspace" below for why it had to stop being one):
417
+
418
+ - **Hashing ignores the workspace's identity.** `hashResolvedInputs`
419
+ replaces every occurrence of `ctx.workspace` inside the serialized
420
+ resolved input/prompt (including trailing subpaths, e.g.
421
+ `<workspace>/cleaned/out.txt`) with a stable placeholder before hashing.
422
+ Two runs of the same logical step under the same `cacheKey` hash
423
+ identically even though AIP-58 §4 gives each one a fresh, disjoint
424
+ workspace directory — without this, ANY step whose input/prompt names
425
+ `$run.workspace` / `{{run.workspace}}` / `_workflowFsRoot` could never hit
426
+ the journal at all.
427
+ - **A hit relocates forward.** On a hit, every workspace-relative file/
428
+ directory the entry recorded (`StepCacheEntry.workspaceFiles` — collected,
429
+ best-effort, from every string in the step's output that resolved to a
430
+ real path under the ORIGINAL run's workspace) is copied into the matching
431
+ path under the CURRENT run's own; the recorded output's path strings are
432
+ then rewritten from the original workspace onto the current one. A
433
+ downstream step whose input reads one of those paths (`$steps.<id>.path`)
434
+ finds the bytes there, not just the (by-then-gone) original run's.
435
+ Best-effort, same posture as the `artifact` step's own relocation: a
436
+ source already swept by a host's `scratch/` retention policy is not this
437
+ run's problem to recover.
438
+ - **Still no cross-step content hashing.** This closes the *workspace-path*
439
+ gap only — the journal still hashes each step's own resolved inputs, not
440
+ a transitive hash of everything upstream. A step whose resolved input is
441
+ an opaque reference that stays textually identical across runs (a fixed
442
+ relative filename, an id) replays from cache on that basis alone,
443
+ regardless of whether the value behind that reference changed upstream —
444
+ same limitation any resolved-input-hash cache has, workspace or not.
445
+
227
446
  ### Step lifecycle callbacks
228
447
 
229
- Pass `onStepStart` and `onStepComplete` to `runWorkflow` to observe progress in
448
+ Pass `onStepStart`, `onStepComplete` and `onStepSkipped` to `runWorkflow` to observe progress in
230
449
  real time. The callbacks fire for every step kind; for `agent` steps, start fires
231
450
  before spawn and complete fires after the turn (and any output-schema retry loop)
232
451
  finishes.
233
452
 
453
+ A cacheable step replayed from the journal still fires both callbacks (it
454
+ doesn't vanish from progress), with a third `info` argument of
455
+ `{ cached: true }`; an executed step gets `info === undefined`. The daemon's
456
+ `workflow_status` surfaces this as `cached: true` on the step row, and on the
457
+ AIP-58 `step.started`/`step.succeeded` events' `data`.
458
+
459
+ `onStepSkipped(stepId, { reason: "branch-not-taken", branchId })` fires, once a
460
+ `branch` decides, for every statically-known step in the arms it did NOT take
461
+ (a `map`/`pipeline`/`subworkflow` step reports its own id; a step id that also
462
+ sits on the taken path is never reported). Inside a `map` item the id is
463
+ indexed (`<id>[<index>]`) like the other two callbacks.
464
+
234
465
  ```ts
235
466
  await runWorkflow({
236
467
  workflow: wf,
237
- onStepStart: (stepId) => console.log("starting", stepId),
238
- onStepComplete: (stepId, output) => console.log("done", stepId, output),
468
+ onStepStart: (stepId, info) => console.log("starting", stepId, info?.cached ? "(cached)" : ""),
469
+ onStepComplete: (stepId, output, info) => console.log("done", stepId, output, info?.cached ? "(cached)" : ""),
470
+ onStepSkipped: (stepId, info) => console.log("skipped", stepId, "by", info.branchId),
239
471
  })
240
472
  ```
241
473