hermes-taskflow 0.2.9

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,692 @@
1
+ ---
2
+ name: taskflow
3
+ description: Orchestrate multi-phase subagent workflows with Taskflow. Use whenever a request spans a whole project or many items — deeply exploring / 探索 / auditing / 审计 / analyzing a codebase, reviewing or migrating many files or modules in parallel, cross-checked/adversarial review, codebase-wide research, or any repeatable orchestration you want to save and rerun. Prefer this over ad-hoc parallel work when the task has multiple phases (discover → work → review → report) or dynamic fan-out over a discovered list. Drives the taskflow_* MCP tools (Hermes registers them as mcp_taskflow_*).
4
+ ---
5
+
6
+ <!-- GENERATED FILE — do not edit. Source: skills-src/taskflow/entry.hermes.md + core.md (npm run build:skills) -->
7
+
8
+ # Taskflow (Hermes Agent)
9
+
10
+ **Host binding (Hermes):** everything below is driven through the `taskflow_*`
11
+ MCP tools. Hermes prefixes MCP tools as `mcp_taskflow_*` (e.g.
12
+ `mcp_taskflow_verify`). Where an example shows a host-neutral invocation like
13
+ `verify`, use the Hermes form (`mcp_taskflow_verify` / `taskflow_verify`).
14
+ Each phase's subagent runs as an isolated `hermes chat -q -Q --source tool`
15
+ session.
16
+
17
+ | Tool | What it does |
18
+ |------|--------------|
19
+ | `taskflow_run` | Run a saved or inline flow. Optional `args`, `incremental`; `mode: "background"` returns a durable `runId` immediately. |
20
+ | `taskflow_runs` | List background runs or `status` / `wait` / `cancel` one by `runId`. |
21
+ | `taskflow_resume` | Fork a failed/paused run into a new immutable child run, optionally overriding one phase's task/model/timeouts. |
22
+ | `taskflow_version` | Report the executing package version, build commit, schema version, build time, and host identity. |
23
+ | `taskflow_list` | List saved flows discoverable from the current working directory. |
24
+ | `taskflow_show` | Show a saved flow's full definition as JSON. |
25
+ | `taskflow_plan` | Preflight plan: bind args, phase order, dynamic bindings, worst-case agent-call bound — zero tokens, no execution. |
26
+ | `taskflow_analytics` | Aggregate last-N runs for a flow (status histogram, durations, per-phase fail/cache rates). Read-only. |
27
+ | `taskflow_verify` | Statically verify a flow (cycles, missing deps, undefined refs, contract typos) — no execution, zero tokens. |
28
+ | `taskflow_compile` | Render a flow's DAG as an inline SVG **and** text outline + a verification report — no execution. |
29
+ | `taskflow_peek` | Inspect one phase's intermediate output from a stored run (post-hoc debugging). Omit `phaseId` to list phases; `json`/`item`/`limit` refine the slice. Hard-truncated, read-only. |
30
+ | `taskflow_trace` | Read a run's append-only event timeline. |
31
+ | `taskflow_replay` | Replay recorded decisions offline with optional overrides — zero model calls. |
32
+ | `taskflow_why_stale` | Explain why phases are stale from observed and declared dependencies — zero tokens. |
33
+ | `taskflow_recompute` | Compute the stale frontier (**dry-run only** over MCP; never executes phases). |
34
+ | `taskflow_reconcile_workspace` | After inspection/repair, accept a failed resolve-only workspace. Requires host `TASKFLOW_WORKSPACE_RECONCILE_MODE=explicit`; never restores files. |
35
+ | `taskflow_save` | Save a reusable flow and optional library metadata. |
36
+ | `taskflow_search` | Search and rank reusable flows before authoring another one. |
37
+
38
+ **Always `taskflow_plan` (or at least `taskflow_verify`) a non-trivial flow before `taskflow_run`** — free, binds args, and catches most authoring mistakes.
39
+
40
+ **Mutating runs:** set `PI_TASKFLOW_HERMES_UNSAFE_YOLO=1` on the MCP server env so agent phases may use Hermes `--yolo` (see package README / `plugin/hermes.config.snippet.yaml`).
41
+
42
+ Build and run **declarative, multi-phase workflows** of subagents. The runtime
43
+ holds intermediate results and the phase DAG, so your main context only receives
44
+ the final answer — not every step's transcript.
45
+
46
+ ## Documentation map (progressive loading)
47
+
48
+ This file teaches the core: phase types, control flow, interpolation, and the
49
+ mistakes that break flows. Load the companion files **only when needed**:
50
+
51
+ | File | Load when you need |
52
+ |------|--------------------|
53
+ | `patterns.md` | **Designing a non-trivial flow.** Proven flow archetypes (audit fan-out, self-healing rework, plan→approve→execute, dynamic replanning, tournament synthesis, incremental audit), anti-patterns, and the production-flow quality checklist. |
54
+ | `advanced.md` | Dynamic sub-flow (`flow{def}`) contracts & security caps, workspace isolation (`cwd: temp/dedicated/worktree`), immutable resume (`taskflow_resume`), and build/host identity (`taskflow_version`). |
55
+ | `configuration.md` | Every knob: per-phase `model`/`thinking`/`tools`/`cwd`, concurrency model, agent discovery, `settings.json`, cross-run caching (`cache`, `fingerprint`, per-item map caching), args, storage paths. **TypeScript DSL CLI** (`taskflow-dsl` / S4). |
56
+ | `library.md` | **Before authoring a non-trivial flow — SEARCH the reusable-flow library.** Save reusable flows with `purpose`+`tags` so future search finds them; reuse + generalize instead of rewriting from scratch. The compounding flywheel. |
57
+
58
+ > Rule of thumb: writing a flow with ≥ 4 phases, a gate, or any fan-out?
59
+ > **Read `patterns.md` first** — it will make the flow better, not just valid.
60
+
61
+ ## When to use
62
+
63
+ - A task needs **several coordinated steps** (discover → work → review → report).
64
+ - You need to **fan out over many items** (audit every endpoint, summarize every file).
65
+ - You want **cross-checked / adversarial review** before reporting.
66
+ - You want a **repeatable** orchestration you can save and rerun by name.
67
+ - The same expensive analysis will be **re-run as the repo evolves** (use
68
+ `incremental: true` + fingerprints — see `configuration.md` §8).
69
+
70
+ ## When NOT to use
71
+
72
+ - A **single-file, single-step** change you can do directly — just do it.
73
+ - **Interactive debugging** where each step depends on watching live output.
74
+ - Work that is **one bash command** — run it yourself, don't wrap it in a flow.
75
+
76
+ ## Flow design ladder
77
+
78
+ Match the flow's sophistication to the task. Don't stop at level 1 when the
79
+ task deserves level 3 — the higher levels are where taskflow pays for itself.
80
+
81
+ | Level | Shape | Reach for it when |
82
+ |-------|-------|-------------------|
83
+ | 0 | shorthand `task` / `tasks` / `chain` | one-off delegation, simple sequence |
84
+ | 1 | linear DAG with `dependsOn` | fixed steps, each consuming the last |
85
+ | 2 | discover → `map` fan-out → `gate` → `reduce` | many items, needs review before reporting |
86
+ | 3 | + `eval` zero-token gates, `expect` contracts, `retry`, `onBlock: "retry"`, `budget`, `optional` fallbacks | production-grade: self-healing, cost stop-loss, fails precisely |
87
+ | 4 | + `loop`, `tournament`, `flow{def}` / `expand`, `race` | the work itself is discovered at runtime; one shot is unreliable; try parallel approaches and keep the first win |
88
+ | 5 | + `incremental: true`, `cache.fingerprint` | the flow re-runs as the repo changes; only re-pay for what changed |
89
+
90
+ **A production-grade flow (level 3+) usually has:** machine checks before LLM
91
+ checks (`eval`, `script`), an `expect` contract on every JSON-emitting phase,
92
+ `retry` on contract-checked phases, a `budget`, `optional: true` on
93
+ degradable phases with a downstream fallback, and exactly one `final` phase.
94
+ `patterns.md` shows each of these composed into full archetypes.
95
+
96
+ ## Shorthand (non-DAG)
97
+
98
+ Skip the DSL entirely for simple delegations. The runtime desugars these into a
99
+ proper flow, so you still get progress, persistence, and resume.
100
+
101
+ ```jsonc
102
+ // single — one agent, one task
103
+ { "task": "Summarize the architecture of src/", "agent": "explorer" }
104
+
105
+ // parallel — run several tasks at once, outputs merged
106
+ { "tasks": [
107
+ { "task": "Audit auth in src/api", "agent": "analyst" },
108
+ { "task": "Audit input validation in src/api", "agent": "analyst" }
109
+ ] }
110
+
111
+ // chain — run sequentially; reference the prior step with {previous.output}
112
+ { "chain": [
113
+ { "task": "List the public API of src/lib", "agent": "scout" },
114
+ { "task": "Write docs for:\n{previous.output}", "agent": "writer" }
115
+ ] }
116
+ ```
117
+
118
+ - `agent` is optional (defaults to the first available agent).
119
+ - `context` (optional, per step or top-level in single mode): file paths to
120
+ pre-read and inject before the task — same as the full-DSL `Phase.context`
121
+ (per-file `contextLimit`, default 8000 chars). In **parallel `tasks` mode**
122
+ all branches SHARE the union of step contexts. In **chain mode** declare
123
+ `context` on individual steps; a top-level `context` is ignored (with a warning).
124
+ - `cwd` (optional, top-level or per-step): working directory for the subagent —
125
+ same as the full-DSL `Phase.cwd`. A top-level `cwd` is the default for every
126
+ step; a per-step `cwd` overrides it. For **single** and **chain** it lands on
127
+ each `Phase.cwd` (full workspace-keyword lifecycle: `temp`/`dedicated`/
128
+ `worktree`). For **parallel `tasks`**, the top-level `cwd` is the shared phase
129
+ cwd, and each branch may set its own **literal-path** `cwd` (mixed branch cwds
130
+ are honored independently). Per-branch workspace keywords are **rejected**
131
+ (the workspace lifecycle is per-phase — use the top-level `cwd` for isolation).
132
+ - Add `name` to label the run.
133
+ - Precedence if several are given: `chain` > `tasks` > `task`.
134
+ - Pass these as the `define` argument to `taskflow_run`.
135
+
136
+ ## How to author a taskflow
137
+
138
+ Call `taskflow_run` with an inline `define` object, or `name` for a saved flow.
139
+ **Before running a non-trivial flow, `taskflow_plan` it (or at least
140
+ `taskflow_verify`) — zero tokens: binds args, projects the phase plan + budget
141
+ bound, and catches cycles / missing deps / undefined refs / contract typos.**
142
+
143
+ ### Iterating on a big flow? Use `defineFile` (write once, verify / edit / run by path)
144
+
145
+ For a non-trivial flow you'll iterate on, **write the definition to a file**
146
+ (typically in the OS tmp dir) and point every call at it with `defineFile`:
147
+
148
+ ```jsonc
149
+ // 1. write /tmp/audit.json with the `write` tool (a full {name, phases:[…]} object)
150
+ // 2. verify, iterate, run — all reference the SAME file by path:
151
+ { "name": "taskflow_plan", "arguments": { "defineFile": "/tmp/audit.json", "args": { … } } } // zero tokens: bind + plan + budget bound
152
+ { "name": "taskflow_verify", "arguments": { "defineFile": "/tmp/audit.json" } } // zero tokens
153
+ { "name": "taskflow_compile", "arguments": { "defineFile": "/tmp/audit.json" } } // diagram
154
+ { "name": "taskflow_lint", "arguments": { "defineFile": "/tmp/audit.json" } } // script-lint + custom verifiers
155
+ { "name": "taskflow_run", "arguments": { "defineFile": "/tmp/audit.json" } }
156
+ ```
157
+
158
+ The file can be raw JSON **or** a Markdown doc with a fenced ```json block
159
+ (`write` the JSON form, or paste the flow into a note and fence it). Between
160
+ calls, edit the file (not the call) and re-`verify`. This avoids re-sending a
161
+ large definition on every call and keeps a durable draft you can diff. Falls
162
+ back cleanly: precedence is `define` (inline) > `defineFile` (disk) > `name`
163
+ (saved flow).
164
+
165
+ ### DSL shape
166
+
167
+ ```jsonc
168
+ {
169
+ "name": "audit-endpoints",
170
+ "description": "Audit API endpoints for missing auth",
171
+ "args": { "dir": { "default": "src/routes" } },
172
+ "concurrency": 8,
173
+ "budget": { "maxUSD": 2.00 },
174
+ "agentScope": "user", // user | project | both
175
+ "phases": [
176
+ { "id": "discover", "type": "agent", "agent": "scout",
177
+ "task": "List endpoints under {args.dir}. Output ONLY a JSON array [{\"route\":\"\",\"file\":\"\"}].",
178
+ "output": "json",
179
+ "expect": { "type": "array", "items": { "type": "object", "required": ["route", "file"] } },
180
+ "retry": { "max": 2, "backoffMs": 0 } },
181
+ { "id": "audit", "type": "map", "over": "{steps.discover.json}", "as": "item",
182
+ "agent": "analyst", "task": "Audit {item.route} ({item.file}) for missing auth.",
183
+ "dependsOn": ["discover"] },
184
+ { "id": "review", "type": "gate", "agent": "reviewer",
185
+ "task": "Remove false positives from:\n{steps.audit.output}\nVERDICT: PASS or BLOCK.",
186
+ "dependsOn": ["audit"] },
187
+ { "id": "report", "type": "reduce", "from": ["review"], "agent": "writer",
188
+ "task": "Write a final report:\n{steps.review.output}", "dependsOn": ["review"],
189
+ "final": true }
190
+ ]
191
+ }
192
+ ```
193
+
194
+ ### Phase types (12)
195
+
196
+ | type | meaning | details |
197
+ |------|---------|---------|
198
+ | `agent` | one subagent runs `task` | this file |
199
+ | `parallel` | run static `branches[]` concurrently (all complete) | this file |
200
+ | `map` | fan out over `over` (an array) — one subagent per item, `{item}` bound | this file |
201
+ | `gate` | quality/review step that can **halt the flow** | Gate phases below |
202
+ | `reduce` | aggregate `from[]` phases into one output | this file |
203
+ | `approval` | **human-in-the-loop** pause: approve / reject / edit | Approval phases below |
204
+ | `flow` | run a **sub-flow** as one phase — saved (`use`) or runtime-generated (`def`) | summary below; deep contract in `advanced.md` |
205
+ | `loop` | repeat a body until a condition / convergence / `maxIterations` | Loop phases below |
206
+ | `tournament` | run N competing `variants`, a `judge` picks best or aggregates | Tournament phases below |
207
+ | `script` | run a **shell command** (no LLM, zero tokens) — stdout is the output | Script phases below |
208
+ | `race` | run `branches[]` concurrently; **first success wins** (unlike parallel) | Race phases below |
209
+ | `expand` | run a dynamic fragment (`def`); `nested` (isolated) or `graft` (promote onto parent) | Expand phases below |
210
+
211
+ ### Control-flow fields (any phase)
212
+
213
+ | field | meaning |
214
+ |-------|---------|
215
+ | `when` | conditional guard — skip the phase unless the expression is truthy. Supports `{refs}`, `== != < > <= >=`, `&& \|\| !`, parentheses, quoted strings/numbers. Parse errors fail **open** (phase runs). |
216
+ | `join` | dependency join: `"all"` (default — wait for every dep) or `"any"` (OR-join — run as soon as one dep completes). |
217
+ | `retry` | `{ "max": N, "backoffMs": ms, "factor": k }` — retry a failing subagent up to N times; delay is `backoffMs * factor^attempt` (`factor:1`=fixed, `2`=exponential). |
218
+ | `timeout` | max ms per subagent call (>= 1000). On expiry the subagent is aborted and the phase fails with a `timedOut` marker — deterministic, **never retried**. Caps EACH call, so a map/parallel/race/loop/tournament phase's wall time is per item/iteration/variant (a tournament's judge call gets its own cap too). Script phases keep their own child-process timeout (default 60s, max 300s). Not supported on approval/flow/expand. Pair with `optional: true` + a downstream fallback phase to degrade instead of failing the run. |
219
+ | `expect` | output contract for `output: "json"` phases (agent/gate/reduce/loop): a JSON-Schema-like shape `{type, properties, required, items, enum}` validated the moment the subagent finishes. A violation fails the phase with per-path diagnostics (e.g. `$.score: required key is missing`) and is retryable under the phase's explicit `retry`. `verify`/`compile` also statically warn when a `{steps.X.json.field}` ref names a field absent from X's declared contract. |
220
+ | `idempotent` | side-effect classification. Default `true` (safe to cache + auto-retry). Set `false` on phases with **irreversible side effects** (webhook POSTs, deploys, DB writes, file mutations): transient provider errors are **not** auto-retried (an explicit `retry{}` IS still honored — it's your declaration that repeats are acceptable) and the result is **never cached** in any scope (within-run resume, cross-run, `incremental` — the phase re-runs every time). The phase state records `sideEffect: true` (rendered as ⚡). |
221
+ | `optional` | fail-soft — a failed/blocked phase won't abort the run; downstream sees empty output. Pair with a fallback phase guarded by `when`. |
222
+ | `cache` | per-phase reuse policy (`run-only` default / `cross-run` / `off`). See `configuration.md` §8. |
223
+
224
+ ### Conditional routing (when + gate/branches)
225
+
226
+ Pair `when` with an upstream phase that emits a decision to build real if/else
227
+ routing. Use `join: "any"` on the merge phase so it runs whichever branch fired.
228
+ For static (non-conditional) concurrency, a `parallel` phase runs fixed
229
+ `branches[]` instead — `{ "type": "parallel", "branches": [{"task":"..."}, {"task":"...","agent":"reviewer"}] }`.
230
+
231
+ ```jsonc
232
+ { "id": "triage", "type": "agent", "agent": "analyst", "output": "json",
233
+ "task": "Classify the task. Output ONLY {\"route\":\"deep\"} or {\"route\":\"quick\"}.",
234
+ "expect": { "type": "object", "required": ["route"], "properties": { "route": { "enum": ["deep", "quick"] } } } },
235
+ { "id": "deep", "when": "{steps.triage.json.route} == deep", "dependsOn": ["triage"], "agent": "analyst", "task": "..." },
236
+ { "id": "quick", "when": "{steps.triage.json.route} == quick", "dependsOn": ["triage"], "agent": "executor-fast", "task": "..." },
237
+ { "id": "report", "type": "reduce", "from": ["deep","quick"], "join": "any",
238
+ "dependsOn": ["deep","quick"], "agent": "writer", "task": "...", "final": true }
239
+ ```
240
+
241
+ > **⚠️ Breaking change (0.2.0 dogfood fix):** a `reduce` phase's `{previous.output}` now aggregates **all** completed `from[]` sources (in from-array order), not just the last completed dependency. If your reduce task referenced `{previous.output}` expecting only the last dep, it now receives every `from[]` output. Use explicit `{steps.ID.output}` refs to address individual sources. For large aggregations, set `reduceStrategy: "tree"` + `batchSize` to run batched intermediate reducer rounds (forces the imperative runtime).
242
+
243
+ > `when` should reference **upstream** (`dependsOn`) phases — a ref to a phase
244
+ > that hasn't completed resolves empty and the guard is treated as false. Note
245
+ > the `expect` enum on the router: it converts "the router said `Deep` with a
246
+ > capital D and both branches silently skipped" into an immediate retryable
247
+ > failure at the router.
248
+
249
+ ### Gate phases (quality control)
250
+
251
+ A `gate` phase runs an agent to review upstream output and can **block the rest
252
+ of the workflow**. The runtime needs to read a verdict from the agent's output.
253
+ There are three ways to provide one, in order of robustness:
254
+
255
+ **1. JSON contract (most robust — preferred).** Set `output: "json"` + an `expect`
256
+ enum so the output is machine-validated. A verdict that isn't exactly `"pass"` or
257
+ `"block"` (wrong case, extra formatting, a synonym) fails the `expect` contract and
258
+ is retried — the verdict can never be silently misread.
259
+
260
+ ```jsonc
261
+ { "id": "review", "type": "gate", "agent": "reviewer", "dependsOn": ["impl"],
262
+ "output": "json",
263
+ "expect": { "type": "object",
264
+ "properties": { "verdict": { "enum": ["pass", "block"] }, "reason": { "type": "string" } },
265
+ "required": ["verdict", "reason"] },
266
+ "task": "Review the diff. Respond ONLY with JSON: {\"verdict\":\"pass\"|\"block\",\"reason\":\"...\"}" }
267
+ ```
268
+
269
+ **2. Explicit text marker.** End the task by asking the agent to emit a final line
270
+ `VERDICT: PASS` or `VERDICT: BLOCK` (also accepts OK/FAIL/STOP/REJECT/HALT; common
271
+ Markdown emphasis like `VERDICT: **BLOCK**` is tolerated). JSON objects such as
272
+ `{"continue": false, "reason": "missing auth checks"}` / `{"verdict": "block"}` also work.
273
+
274
+ **3. Auto-appended format suffix.** If a free-text gate's task does **not** already
275
+ ask for a `VERDICT:` marker (and has no JSON contract), the runtime automatically
276
+ appends the exact format instruction. You don't need to remember to add it — but
277
+ writing it yourself (option 2) makes the intent explicit in your flow.
278
+
279
+ On **BLOCK**, downstream phases are skipped and the run ends as `blocked` with the
280
+ reason surfaced. Unparseable gate **model output fails closed** (treated as BLOCK):
281
+ a gate that cannot reach a verdict cannot be trusted to pass (issue #54). Note
282
+ that *config* slips (an unresolved `score.target`, malformed `scorers`) are
283
+ different and still fail **open** with a warning — those are authoring errors that
284
+ degrade to the historical behavior, not a judge that couldn't decide. An explicit
285
+ non-blocking JSON verdict (e.g. `{"verdict":"No issues found"}`) is a semantic PASS,
286
+ not ambiguity.
287
+
288
+ **Zero-token machine checks (`eval`) — use these before spending tokens.**
289
+ List machine-checkable assertions in `eval`. If **all** pass, the gate
290
+ auto-passes with **no LLM call**; if any fails, it falls through to the LLM
291
+ `task` (the qualitative residue). Each entry supports the `when` operators plus
292
+ `X contains Y` (substring). A parse error fails **open**.
293
+
294
+ ```jsonc
295
+ { "id": "quality", "type": "gate", "dependsOn": ["build","test"],
296
+ "eval": ["{steps.build.output} contains BUILD SUCCESS", "{steps.test.json.failures} == 0"],
297
+ "task": "Review the diff for subtle logic errors a linter can't catch. VERDICT: PASS or BLOCK." }
298
+ ```
299
+
300
+ **Self-healing (`onBlock: "retry"`).** By default a blocking gate halts the run
301
+ (`onBlock: "halt"`). With `onBlock: "retry"` the gate instead **re-runs its
302
+ upstream `dependsOn` phases and re-evaluates**, up to `retry.max` rounds (or
303
+ until PASS / budget / abort) — a generate→critique→regenerate rework loop. See
304
+ `patterns.md` for the full archetype.
305
+
306
+ ```jsonc
307
+ { "id": "spec-gate", "type": "gate", "onBlock": "retry", "retry": { "max": 3 },
308
+ "dependsOn": ["implement"],
309
+ "task": "Does the implementation satisfy ALL acceptance criteria? VERDICT: PASS or BLOCK with reasons." }
310
+ ```
311
+
312
+ **Scoring gates (`score`) — graded, composable, auditable quality checks.**
313
+ Where `eval` gives boolean assertions, `score` runs deterministic scorers
314
+ against a target string at **zero tokens**, combines them into a [0,1] score,
315
+ and only escalates to an LLM when they can't decide. The structured result is
316
+ the gate's `.json` — downstream phases read `{steps.<gate>.json.combined}` /
317
+ `.json.results.0.passed` and route on quality, not just pass/fail.
318
+
319
+ | field | meaning |
320
+ |-------|---------|
321
+ | `target` | interpolation ref for the scored string (default `{previous.output}`) |
322
+ | `scorers` | array of checks: `exact-match` (`value`), `contains` (`value`), `regex` (`pattern`, optional `negate`), `json-schema` (`schema`, an `expect`-style contract), `length-range` (`min`/`max`), `code-compiles` (`language`: javascript\|typescript) |
323
+ | `combine` | `all` (default) / `any` / `weighted` |
324
+ | `weights` | weighted only — one entry per scorer, **+1 trailing entry for the judge** when present |
325
+ | `threshold` | weighted only — combined-score cutoff in (0,1], default 0.5 |
326
+ | `judge` | optional LLM-as-judge fallback `{agent?, task}` — runs when the deterministics fail (and, for `all`/`any`, whenever configured); sees the target + scorer report; returns `{"score": 0-1, "verdict": "pass"\|"block", "reason"}` |
327
+
328
+ Decision order: (1) deterministics pass **and the judge cannot veto** →
329
+ **auto-PASS, zero LLM tokens** — that means: no judge configured, or `weighted`
330
+ where the deterministic score is a lower bound already clearing the threshold
331
+ (the judge could not drop it). With `all`/`any` + a judge the judge **always
332
+ runs** — its verdict is authoritative (it may check what scorers cannot, e.g.
333
+ factuality); (2) fail + `judge` → judge decides; (3) fail + `task` → the gate
334
+ task runs with the scorer report appended; (4) fail + no fallback → **explicit
335
+ BLOCK** (a deterministic failure is not ambiguity). Fail-closed: an unparseable
336
+ judge → BLOCK (issue #54); unresolved `target` with no fallback → PASS +
337
+ warning (config slip, not a judge verdict); malformed `score` → the plain LLM
338
+ gate. **Security:** LLM-generated dynamic sub-flows
339
+ (`flow{def}`) may not use `code-compiles` (compiler execution) or `regex`
340
+ (ReDoS) scorers — same hardening class as the `script` block.
341
+
342
+ ```jsonc
343
+ { "id": "quality", "type": "gate", "dependsOn": ["gen"],
344
+ "score": {
345
+ "target": "{steps.gen.output}",
346
+ "scorers": [
347
+ { "type": "json-schema", "name": "shape", "schema": { "type": "object", "required": ["summary", "risks"] } },
348
+ { "type": "regex", "name": "no-placeholders", "pattern": "TODO|TBD", "negate": true },
349
+ { "type": "length-range", "name": "substantive", "min": 200 }
350
+ ],
351
+ "combine": "weighted", "weights": [3, 2, 1, 2], "threshold": 0.8,
352
+ "judge": { "agent": "reviewer", "task": "Score the analysis quality 0-1: depth, evidence, actionability." }
353
+ } }
354
+ // downstream: { "when": "{steps.quality.json.combined} >= 0.9", ... }
355
+ ```
356
+
357
+ ### Approval phases (human-in-the-loop)
358
+
359
+ An `approval` phase pauses the run and asks the operator to **Approve / Reject /
360
+ Edit**. Distinct from `gate` (an *agent* reviewing): this is a *human* deciding.
361
+ The (interpolated) `task` is the prompt shown.
362
+
363
+ - **Approve** → continue; the phase output is `(approve)`.
364
+ - **Reject** → halt the flow (same mechanism as a blocking gate).
365
+ - **Edit** → the typed note becomes this phase's `output` — inject guidance
366
+ mid-run and reference it downstream with `{steps.<id>.output}`.
367
+ - **Non-interactive** runs (headless/CI/print mode) **auto-reject** and record it.
368
+ - **Background (detached)** runs **auto-reject** (no interactive approver);
369
+ downstream sees the rejection; the flow continues (fail-open).
370
+
371
+ > **MCP-host caveat (Codex / Claude Code / OpenCode / Grok / Hermes):** MCP-driven runs are
372
+ > non-interactive, so an `approval` phase **auto-rejects**. Prefer a `gate`
373
+ > (agent review) in flows you run through the `taskflow_*` tools; use `approval`
374
+ > only in flows a human runs interactively.
375
+
376
+ ### Sub-flows (composition) — summary
377
+
378
+ A `flow` phase runs another taskflow as a single phase and bubbles up its final
379
+ output. Two mutually-exclusive sources:
380
+
381
+ - **Saved** (`use`): `{ "type": "flow", "use": "deep-research", "with": { "topic": "{item}" } }`
382
+ — args via `with` (string values interpolate); recursion is detected and rejected.
383
+ - **Runtime-generated** (`def`): `{ "type": "flow", "def": "{steps.plan.json}" }`
384
+ — an upstream planner emits a whole flow as JSON; the runtime validates it
385
+ (cycles / dangling refs / security caps) then runs it nested. This is how a
386
+ planner decides *at runtime* what work to spawn — the declarative answer to a
387
+ code-mode `for`/`if` loop.
388
+
389
+ The `def` output contract, fail-open semantics (`defError`), nesting/breadth
390
+ caps, and the iterative-replanning pattern (`loop` + `flow{def}`) are in
391
+ `advanced.md`. The plan→execute and replan archetypes are in `patterns.md`.
392
+
393
+ ### Loop phases (iterate until done)
394
+
395
+ A `loop` phase runs its body repeatedly, exposing each iteration's output as
396
+ `{steps.<thisId>.output}` / `.json` so the next round can react to the last. It
397
+ stops on the first of: `until` truthy, **convergence** (output stops changing),
398
+ or `maxIterations` (hard cap, required). The runtime always terminates.
399
+
400
+ - `until` — stop condition, same operators as `when` (a parse error stops the loop, fail-safe).
401
+ - `maxIterations` — hard iteration cap (required).
402
+ - `convergence` — `true` to stop early when an iteration's output equals the previous one.
403
+ - `reflexion` — `true` to feed each iteration a structured summary of the prior one (see below).
404
+
405
+ ```jsonc
406
+ {
407
+ "id": "refine", "type": "loop", "agent": "executor",
408
+ "maxIterations": 5,
409
+ "until": "{steps.refine.json.done} == true",
410
+ "convergence": true,
411
+ "task": "Improve the draft. When nothing else needs fixing, output JSON {\"done\":true,\"draft\":\"...\"}; otherwise {\"done\":false,\"draft\":\"...\"}.",
412
+ "output": "json",
413
+ "expect": { "type": "object", "required": ["done", "draft"] },
414
+ "final": true
415
+ }
416
+ ```
417
+
418
+ **Reflexion memory (`reflexion: true`).** By default each iteration sees only
419
+ the prior *output* — the *reason* it wasn't good enough (an `expect` contract
420
+ violation, an error, the unmet `until`) is discarded, so models repeat mistakes.
421
+ With `reflexion: true`, every iteration after the first receives a structured
422
+ failure summary of the prior one via the `{reflexion}` placeholder
423
+ (auto-appended if the task omits it, with a one-time warning; capped at 2000
424
+ chars): contract diagnostics like `$.done: required key is missing`, the
425
+ (sanitized) error, or the unmet stop condition, plus a truncated output
426
+ snippet. Iteration 1 sees a sentinel.
427
+
428
+ Semantics shift to enable self-correction: **body failures become feedback
429
+ instead of terminating the loop**. Timeout/abort/over-budget still hard-stop,
430
+ and if `maxIterations` exhausts with the last iteration failed, the phase fails
431
+ (reflexion defers failure, never erases it). Cost is bounded by `maxIterations`
432
+ + the run `budget`.
433
+
434
+ ```jsonc
435
+ { "id": "emit-plan", "type": "loop", "reflexion": true, "maxIterations": 4,
436
+ "output": "json", "expect": { "type": "object", "required": ["steps", "done"] },
437
+ "until": "{steps.emit-plan.json.done} == true",
438
+ "task": "Emit the migration plan as JSON {steps:[...], done:bool}.\n{reflexion}" }
439
+ ```
440
+
441
+ ### Tournament phases (N variants, judge picks best)
442
+
443
+ A `tournament` phase runs `variants` competing attempts in parallel, then a
444
+ **judge** sub-phase selects the winner (`mode: "best"`) or merges them
445
+ (`mode: "aggregate"`). Use it when one shot is unreliable and you want the best
446
+ of several drafts, or a synthesis of diverse approaches.
447
+
448
+ - `variants` — number of competing variants spawned from `task` (default 3, max 20).
449
+ For genuinely different *approaches*, use `branches` instead — an explicit
450
+ array of `{task, agent?}` definitions (e.g. one conservative, one aggressive).
451
+ - `mode` — `"best"` (judge picks one winner, default) or `"aggregate"` (judge merges all).
452
+ - `judge` — the judge's rubric/instructions. `judgeAgent` — optional judge agent
453
+ (defaults to the phase `agent`; use a stronger model here).
454
+ - **Winner format — prefer JSON.** Have the judge return `{"winner": <n>}` (and an
455
+ optional `"reason"`); the runtime also reads a `WINNER: <n>` line (`#3` and
456
+ common Markdown emphasis like `WINNER: **3**` are tolerated — issue #54).
457
+ JSON is more robust than a text marker: there's no formatting the model can
458
+ get subtly wrong.
459
+ - Fail-open: if the judge's pick is still unparseable, variant 1 is returned
460
+ (work is never lost — the variants are already computed, so blocking would be
461
+ worse than picking a safe default).
462
+
463
+ ```jsonc
464
+ {
465
+ "id": "headline", "type": "tournament", "agent": "executor",
466
+ "variants": 3, "mode": "best",
467
+ "judge": "Pick the clearest, most accurate headline. Return JSON {\"winner\": <n>, \"reason\": \"...\"}.",
468
+ "task": "Write one headline for the article below.\n\n{steps.draft.output}",
469
+ "dependsOn": ["draft"], "final": true
470
+ }
471
+ ```
472
+
473
+ ### Script phases (shell commands, zero tokens)
474
+
475
+ A `script` phase runs a **shell command** directly — no subagent, no tokens — and
476
+ captures its stdout as the phase output. Use it to anchor LLM phases to ground
477
+ truth: builds, tests, `git`, formatters, scoring scripts. **Prefer a `script`
478
+ phase over asking an agent to run a command** — it is cheaper, faster, and the
479
+ output is exact.
480
+
481
+ - `run` — **required**. A **string** runs through a shell; an **array** is
482
+ spawned directly (execvp, no shell). A string `run` containing an
483
+ interpolation placeholder is **rejected at validation** (shell-injection
484
+ guard) — use the array form or `input` for dynamic values.
485
+ - `input` — optional text piped to stdin (supports interpolation).
486
+ - `timeout` — optional ms cap (1000–300000, default 60000); SIGTERM → SIGKILL on expiry.
487
+ - A non-zero exit fails the phase (stderr captured); stdout capped at 1 MB.
488
+ No `retry`, no `output: "json"`; **excluded from cross-run cache** (may have
489
+ side effects). Not allowed inside LLM-generated dynamic sub-flows (RCE guard).
490
+
491
+ ```jsonc
492
+ { "id": "build", "type": "script", "run": "pnpm run build", "timeout": 120000 },
493
+ { "id": "score", "type": "script", "run": ["python", "score.py"],
494
+ "input": "{steps.analyze.output}", "dependsOn": ["analyze"], "final": true }
495
+ ```
496
+
497
+ ### Race phases (first success wins)
498
+
499
+ A `race` phase runs static `branches[]` concurrently and **returns the first
500
+ branch that finishes successfully** (failed settles do **not** win — a slower
501
+ success still wins over a fast hard-fail). Unlike `parallel` (waits for all) or
502
+ `tournament` (judges quality after all variants), use race when latency matters
503
+ more than comparing every approach.
504
+
505
+ - `branches` — **required**, at least two `{task, agent?}`.
506
+ - `cancelLosers` — optional boolean (default `true`). After the first **success**,
507
+ abort other branches via `AbortSignal` (best-effort — host must honor the
508
+ signal). Set `false` to let losers finish naturally.
509
+ - Phase `usage` **aggregates all branches** (including aborted partials) so
510
+ budgets stay honest.
511
+ - Output of the winning branch becomes the race phase output; a warning records
512
+ which branch won.
513
+
514
+ ```jsonc
515
+ {
516
+ "id": "quick", "type": "race",
517
+ "branches": [
518
+ { "task": "Answer with a short heuristic…", "agent": "executor" },
519
+ { "task": "Answer with a thorough search…", "agent": "researcher" }
520
+ ],
521
+ "final": true
522
+ }
523
+ ```
524
+
525
+ ### Expand phases (dynamic fragment: nested or graft)
526
+
527
+ An `expand` phase runs a **fragment Taskflow** from `def` (inline object,
528
+ phases array, or interpolated `{steps.plan.json}`). Two modes:
529
+
530
+ | `expandMode` | Behavior |
531
+ |--------------|----------|
532
+ | `nested` (default) | Run as an isolated sub-flow (like `flow{def}`); child phase ids stay **off** the parent. |
533
+ | `graft` | After success, **promote** child phase states onto the parent as `<expandId>-<childId>` so later phases can read `{steps.grow-leaf.output}`. |
534
+
535
+ - `def` — **required** for expand.
536
+ - `maxNodes` — optional cap on fragment phase count (default 50, hard max 100).
537
+ - Dynamic validation + nesting caps match `flow{def}` (see `advanced.md`).
538
+ - Prefer `expand` when the planner fragment is a first-class kind; prefer
539
+ `flow` + `use` for saved reusable flows; prefer `flow` + `def` when you want
540
+ the classic nested sub-flow without graft promote.
541
+
542
+ ```jsonc
543
+ {
544
+ "id": "grow", "type": "expand", "expandMode": "graft",
545
+ "def": "{steps.plan.json}",
546
+ "dependsOn": ["plan"], "final": true
547
+ }
548
+ ```
549
+
550
+ ### Budget (observed-usage stop-loss)
551
+
552
+ Add a run-wide stop-loss at the top level. Ordinary budgeted DAG layers and
553
+ `map`/`parallel`/`tournament` fan-out use serial call admission. Once reported
554
+ cost/tokens exceed the threshold, no new model call is started; the run ends as
555
+ `blocked` with partial outputs preserved. An admitted call may cross the
556
+ threshold. A `race` necessarily starts competing branches together, so all
557
+ already-active race branches may contribute overshoot. This is never a
558
+ zero-overshoot guarantee.
559
+
560
+ ```jsonc
561
+ { "name": "...", "budget": { "maxUSD": 1.50, "maxTokens": 2000000 }, "phases": [ ... ] }
562
+ ```
563
+
564
+ **Any flow with a fan-out should have a `budget`** — a map over a
565
+ mis-discovered 500-item array is otherwise unbounded spend.
566
+
567
+ Host accounting matters: Codex reports tokens but not cost, so Codex accepts
568
+ `maxTokens` and rejects `maxUSD`. Grok 0.2.93 and Hermes quiet mode report
569
+ neither, so both reject every flow declaring `budget`. Pi, Claude Code, and
570
+ OpenCode accept both dimensions.
571
+
572
+ ### Strict interpolation
573
+
574
+ By default an unresolved placeholder (typo'd `{steps.X.output}`, missing
575
+ `{args.Y}`) resolves to an empty string and validation issues a *warning* —
576
+ the flow still runs, possibly doing subtly wrong work. Set
577
+ `"strictInterpolation": true` at the flow level to promote unresolved
578
+ placeholders and missing-dep/arg warnings to **hard errors**. Recommended for
579
+ any flow you save — a saved flow will be run later with args you're not
580
+ watching.
581
+
582
+ ## Interpolation
583
+
584
+ - `{args.X}` — invocation argument
585
+ - `{steps.ID.output}` — a prior phase's text output
586
+ - `{steps.ID.json}` / `{steps.ID.json.field}` — prior output parsed as JSON
587
+ - `{item}` / `{item.field}` — current item inside a `map` phase
588
+ - `{previous.output}` — the immediately-upstream phase output. For `reduce` phases, this resolves to **all completed `from[]` outputs** in from-array order: one completed input → its raw output; many → `### <id>\n\n<output>` sections joined by `\n\n---\n\n`. `join: "any"` includes only completed branches (skipped/failed are omitted). Explicit `{steps.ID.output}` refs are unaffected.
589
+ - `{loop.iteration}` / `{loop.lastOutput}` / `{loop.maxIterations}` — inside a `loop` body: the 1-based round, the prior iteration's output, and the cap
590
+ - `{reflexion}` — inside a `loop` body with `reflexion: true`: the structured failure summary of the prior iteration (sentinel on iteration 1)
591
+
592
+ Interpolation also runs on a scoring gate's `score.target` and `score.judge.task`
593
+ — refs there need `dependsOn` like any other `{steps.X}` use.
594
+
595
+ ## Rules that make flows work
596
+
597
+ 1. For a `map` phase, make the upstream phase **emit a JSON array** and set
598
+ `output: "json"` on it. Tell that agent to output **only** JSON, and pin the
599
+ shape with an `expect` contract + `retry`.
600
+ 2. Give each phase a clear, single responsibility.
601
+ 3. Reference upstream results explicitly with `{steps.ID...}` and set `dependsOn`.
602
+ 4. Mark the result-bearing phase with `"final": true` (else the last phase wins).
603
+ 5. Machine checks before LLM checks: `script` for ground truth, gate `eval`
604
+ before gate `task`, `expect` before a downstream "did it parse?" phase.
605
+ 6. **Decision phases should emit structured output, not free text.** Any phase
606
+ whose output is a *decision* a downstream phase (or the runtime) acts on — a
607
+ gate verdict, a router's branch, a tournament winner, a judge's score — should
608
+ use `output: "json"` + an `expect` enum/contract so the decision is
609
+ machine-validated. Free-text markers (`VERDICT:`, `WINNER:`, `SCORE:`) are
610
+ tolerated and Markdown-emphasis-tolerant (issue #54), but a JSON contract is
611
+ strictly more robust: there's no formatting the model can get subtly wrong, and
612
+ a malformed decision fails the contract (retryable) instead of being silently
613
+ mis-read.
614
+ 7. `verify` before `run` for anything non-trivial (zero tokens).
615
+
616
+ ## Common mistakes (the runtime rejects these at validation time)
617
+
618
+ ### 1. Referencing `{steps.X}` without `dependsOn: ["X"]`
619
+
620
+ ```jsonc
621
+ // ❌ WRONG — 'fix-issues' runs in parallel with 'code-review-1' and sees the
622
+ // literal string "{steps.code-review-1.output}" instead of the review text.
623
+ { "id": "code-review-1", "type": "agent", "task": "review code" },
624
+ { "id": "fix-issues", "type": "agent",
625
+ "task": "fix {steps.code-review-1.output}" } // ← no dependsOn!
626
+ ```
627
+
628
+ Validation rejects this: `Phase 'fix-issues': task references
629
+ {steps.code-review-1.*} but 'code-review-1' is not in dependsOn. ...`
630
+ **Always declare the chain:**
631
+
632
+ ```jsonc
633
+ // ✅ RIGHT
634
+ { "id": "code-review-1", "type": "agent", "task": "review code" },
635
+ { "id": "fix-issues", "type": "agent",
636
+ "task": "fix {steps.code-review-1.output}",
637
+ "dependsOn": ["code-review-1"] }
638
+ ```
639
+
640
+ Tip: write the `task` first (it tells you what each phase needs), then scan for
641
+ `{steps.*}` references and add the matching `dependsOn`.
642
+ Exception: phases with `join: "any"` are exempt (they deliberately wait for only
643
+ one dep and may reference others as informational context).
644
+
645
+ ### 2. Assuming the runtime knows "this is a chain"
646
+
647
+ Phase order in the `phases` array is **documentation, not execution order**.
648
+ The DAG comes from `dependsOn`. Four phases listed in order with no `dependsOn`
649
+ are four **parallel** phases, all racing in layer 0. Use the shorthand `chain`
650
+ if you literally want `a → b → c → d`, or write explicit `dependsOn`.
651
+
652
+ ### 3. Underscores in ids / invented agent names
653
+
654
+ Phase ids and agent names use **hyphens** (`audit-each`, `risk-reviewer`).
655
+ An unknown agent name fails the phase with the list of available agents.
656
+ Built-in agents: `executor`, `executor-code` (complex, multi-file),
657
+ `executor-fast` (trivial), `executor-ui`, `scout` (cheap recon), `planner`,
658
+ `analyst`, `critic`, `reviewer`, `risk-reviewer`, `security-reviewer`,
659
+ `plan-arbiter`, `final-arbiter`, `test-engineer`, `doc-writer`, `verifier`,
660
+ `recover`, `visual-explorer`. **Do not invent agent names** — omit `agent` to
661
+ use the default. Use cheap agents (`scout`) for discovery and strong agents
662
+ (`critic`, `final-arbiter`) for gates/judging.
663
+
664
+ ## Operating a run (lifecycle & inspection)
665
+
666
+ A run moves through: **running →** `completed` (a `final` phase produced output)
667
+ **/** `blocked` (gate BLOCK, approval rejected, or `budget` hit) **/** `failed`
668
+ (a non-`optional` phase errored) **/** `paused` (aborted).
669
+
670
+ `taskflow_run` reports a `runId`. If the final output looks wrong, don't
671
+ re-run blind — `taskflow_peek` the run: omit `phaseId` to list phase statuses
672
+ and output sizes, then peek the suspicious phase (`json: true` for parsed
673
+ output, `item: n` for one fan-out section). Output is hard-truncated
674
+ (default 4000 chars, max 32000) so a peek never floods your context.
675
+
676
+ For a flow that may outlive one MCP tool call, set `mode: "background"` on
677
+ `taskflow_run`. It returns immediately; use `taskflow_runs` with `action:
678
+ "status"`, `"wait"`, or `"cancel"` and the returned `runId`. A bounded `wait`
679
+ can be called repeatedly, and completion returns the persisted final output.
680
+ Use `action: "list"` with optional `status: "running" | "terminal"` to see
681
+ active concurrency. Starting a sixth active run warns that Taskflow has no
682
+ hidden global cross-host concurrency or budget coordinator.
683
+
684
+ Use `taskflow_trace` to inspect the append-only event log for a finished run,
685
+ then `taskflow_replay` to re-judge it under alternate thresholds/budget **offline
686
+ (zero tokens)** — e.g. "would a 0.9 gate threshold have blocked this run?"
687
+
688
+ For flows re-run as the repo evolves, pass `incremental: true` to
689
+ `taskflow_run` — every phase defaults to **cross-run cache reuse**: identical
690
+ input → $0 instant hit. Per-phase `cache.fingerprint` entries
691
+ (`git:HEAD`, `glob!:src/**/*.ts`, `file:package.json`) invalidate on world
692
+ changes; a cached `map` re-executes only changed items. See `configuration.md` §8.