hermes-taskflow 0.2.9

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,267 @@
1
+ <!-- GENERATED FILE — do not edit. Source: skills-src/taskflow/advanced.md (npm run build:skills) -->
2
+
3
+ # Taskflow Advanced — dynamic sub-flows & workspace isolation
4
+
5
+ Load this when a flow needs: runtime-generated work (`flow{def}` / `expand`) or
6
+ isolated working directories (`cwd: temp/dedicated/worktree`).
7
+
8
+ ---
9
+
10
+ ## `flow{def}` vs `expand` (when to use which)
11
+
12
+ | Need | Prefer |
13
+ |------|--------|
14
+ | Saved reusable flow by name | `flow` + `use` |
15
+ | Planner JSON as isolated nested sub-flow (classic) | `flow` + `def` **or** `expand` + `expandMode: "nested"` |
16
+ | Fragment phases must appear on the **parent** run as `<expandId>-<childId>` | `expand` + `expandMode: "graft"` |
17
+ | First of several static approaches (latency) | `race` (not tournament) |
18
+
19
+ `expand` is a first-class phase type (Horizon B). Dynamic validation / nesting /
20
+ breadth caps match `flow{def}`. **Event kernel** still excludes `race`/`expand`
21
+ (imperative path only until step handlers exist).
22
+
23
+ ---
24
+
25
+ ## Dynamic sub-flows (`flow{def}`) — the full contract
26
+
27
+ A `flow` phase with `def` resolves a sub-flow **at runtime**, usually from an
28
+ upstream phase's JSON output. The runtime interpolates + JSON-parses the `def`,
29
+ validates it, then runs it nested. This is how a planner decides at runtime
30
+ what work to spawn — with each generated plan checked before it spends a token.
31
+
32
+ ```jsonc
33
+ { "id": "plan", "type": "agent", "agent": "planner", "output": "json",
34
+ "task": "Scan the repo. Output ONLY JSON {\"name\":\"audit\",\"phases\":[...]} — one audit phase per file." },
35
+ { "id": "run", "type": "flow", "def": "{steps.plan.json}", "optional": true,
36
+ "dependsOn": ["plan"], "final": true }
37
+ ```
38
+
39
+ **LLM output contract for `def`** (put this in the planner's task):
40
+ - A *full* Taskflow `{"name":"...","phases":[...]}`, a bare `phases` array, or
41
+ `{"phases":[...]}` — pure JSON (a ```json fence is tolerated and stripped).
42
+ - Hyphens in ids, never underscores.
43
+ - Sub-flow phases reference each other in their **own** `{steps.x.output}`
44
+ namespace (no parent-id prefixing).
45
+ - An **empty** `phases` array is a valid no-op (the planner decided there's
46
+ nothing to do).
47
+
48
+ **Security caps on generated flows** (validation rejects; tell the planner not
49
+ to emit these so a retry isn't wasted):
50
+ - **No `script` phases** — shell execution from an LLM-authored plan is an RCE
51
+ vector; only author-written flows may use `script`.
52
+ - **No workspace `cwd` keywords** (`temp`/`dedicated`/`worktree`) and no `cwd`
53
+ escaping the run directory.
54
+ - Breadth caps: ≤100 phases, concurrency ≤16 (flow and per-phase).
55
+ - Depth: inline nesting capped at 5 (shared with `ctx_spawn` subflows).
56
+
57
+ **Fail-open semantics:** if the `def` doesn't parse, has the wrong shape, or
58
+ fails validation, the phase completes with `status: "done"` and a `defError`
59
+ diagnostic field; downstream phases receive empty output and the run continues.
60
+ Design for it:
61
+ - Add `optional: true` on the flow phase so a bad plan never aborts the run.
62
+ - Want a hard stop instead? Add a downstream gate:
63
+ `{ "type": "gate", "eval": ["{steps.run.output} != "], "task": "…VERDICT: BLOCK if the plan failed." }`
64
+
65
+ **Iterative replanning** — pair `flow{def}` with a `loop` whose body emits the
66
+ next plan from the previous round's result: the declarative equivalent of
67
+ `for (...) { read result; decide next }`. See `examples/dynamic-plan-execute.json`
68
+ and `examples/iterative-replan.json`.
69
+
70
+ ---
71
+
72
+ ## Workspace isolation (`cwd` keywords)
73
+
74
+ A phase's `cwd` is normally a literal path (or inherited from the run). Three
75
+ **reserved keywords** ask the runtime to allocate an isolated working directory
76
+ for the phase's subagent and tear it down afterwards — scratch work or file
77
+ mutation without touching the main tree:
78
+
79
+ | `cwd` value | what the runtime does | lifecycle |
80
+ |-------------|-----------------------|-----------|
81
+ | `"temp"` | ephemeral dir under the OS tmpdir | removed when the phase finishes |
82
+ | `"dedicated"` | persistent dir under the run state (`runs/ws/<runId>/<phaseId>`) | **kept** for inspection; deterministic per phase (resume reuses it) |
83
+ | `"worktree"` | `git worktree add` on a throwaway branch off `HEAD` | `git worktree remove` + branch delete when the phase finishes |
84
+
85
+ ```jsonc
86
+ { "id": "experiment", "type": "agent", "agent": "executor", "cwd": "worktree",
87
+ "task": "Try the risky refactor and run the tests. Your edits are isolated in a git worktree." }
88
+ ```
89
+
90
+ - **Fail-open.** If allocation fails (e.g. `worktree` outside a git tree), the
91
+ phase degrades — `worktree`→`temp`, any other failure → the base cwd — with a
92
+ `warnings` diagnostic. A phase never fails to run because of isolation.
93
+ - **Security.** Keywords are honoured only in **author-written** flows; a
94
+ generated plan (`flow{def}` / `ctx_spawn` subflow) requesting one is rejected
95
+ at validation.
96
+ - A literal path passes through unchanged.
97
+
98
+ ### Argument-selected cwd (0.2.1 compatibility bridge)
99
+
100
+ An author-written flow may set `cwd: "{args.package}"` only when `package` is
101
+ declared as `{ "type": "relative-path" }`. The placeholder must occupy the
102
+ whole field. The value is resolved below the invocation cwd, must be an existing
103
+ directory, and cannot escape through `..`, absolute paths, or symlinks at bind time.
104
+
105
+ The bridge is fail-closed and disabled unless the host operator explicitly sets
106
+ `TASKFLOW_CWD_BRIDGE_MODE=resolve-only`. That mode performs a time-of-check path
107
+ validation but has no no-follow filesystem handle and is not an OS filesystem sandbox; each phase emits a warning stating the lower
108
+ guarantee. Cwd-bridge flow trees do not reuse output-only cache/resume entries,
109
+ because the current 0.2.x runtime cannot restore filesystem mutations on a cache hit. Generated
110
+ sub-flows cannot use the bridge. Saved-flow definitions are frozen for one
111
+ top-level execution, the invocation root identity is persisted for resume, and
112
+ a selected sub-flow inherits a non-expanding canonical boundary: nested literal
113
+ cwd and context files may narrow it but cannot escape it or allocate a workspace
114
+ provider.
115
+
116
+ Do not combine this bridge with `retry.max > 0`; validation rejects that
117
+ combination. After a failed writer, the filesystem outcome is unknown and an
118
+ operator must explicitly reconcile the invocation workspace before another
119
+ write can start.
120
+
121
+ **Pattern — competing experiments in worktrees:** run two `parallel` branches,
122
+ each `cwd: "worktree"`, each attempting a different refactor strategy and
123
+ reporting its test results; a downstream gate/judge picks which diff to apply
124
+ for real. The main tree is never touched by the losers.
125
+
126
+ ---
127
+
128
+ ## Trace & offline replay (`trace` / `replay`) — vs resume / recompute
129
+
130
+ Three **different** reuse tools; do not conflate them:
131
+
132
+ | Tool | Spends tokens? | Mutates the run? | Answers |
133
+ |------|----------------|------------------|---------|
134
+ | **`resume`** | Only unfinished / cache-miss phases | **Forks a new run** (parent untouched; child carries `parentRunId`) | "Pick up where we stopped" |
135
+ | **`why-stale` → `recompute`** | Dry-run free; `--apply` / `dryRun:false` spends | Optional write of recompute result | "World/input changed — which phases re-run?" |
136
+ | **`trace` → `replay`** | **Never** | Never | "If the gate threshold / budget had been different, would we have blocked?" |
137
+
138
+ ### Trace (read the evidence)
139
+
140
+ Every instrumented run may write an append-only **event log**
141
+ (`runs/<flow>/<runId>.trace.jsonl`): phase lifecycle, each subagent
142
+ input/output, and runtime **decisions** (gate verdict/score, when-guard,
143
+ cache-hit, budget-hit, tournament-winner, unreplayable).
144
+
145
+ ```
146
+ taskflow_trace { runId: "<id>" }
147
+ taskflow_trace { runId: "<id>", json: true }
148
+ ```
149
+
150
+ MCP trace responses are bounded. JSON mode returns an envelope with
151
+ `total`/`returned`/`truncated`; use `limit` (default 200, max 1000) to select the
152
+ newest events without flooding the host context.
153
+
154
+ If there is no log (pre-trace run, or no sink injected), the tool reports that
155
+ clearly — it never invents events.
156
+
157
+ ### Offline replay (what-if, zero tokens)
158
+
159
+ `replay` **re-folds** the recorded log under alternate **decision knobs** without
160
+ calling any model:
161
+
162
+ - `thresholds` — map of `phaseId → new score threshold` (gate-score events)
163
+ - `budgetMaxUSD` / `budgetMaxTokens` — would later phases have been skipped?
164
+ - `models` / `args` — currently report `needs-live-rerun` (quality cannot be
165
+ re-judged offline without re-execution)
166
+
167
+ Outcomes per phase: `reused`, `would-block`, `verdict-flipped`,
168
+ `would-exceed-budget`, `threshold-changed`, `needs-live-rerun`, `failed`.
169
+
170
+ ```
171
+ taskflow_replay { runId: "<id>", thresholds: { review: 0.9 } }
172
+ taskflow_replay { runId: "<id>", budgetMaxUSD: 0.05, json: true }
173
+ ```
174
+
175
+ **Import-graph guarantee:** `replayRun` never imports the process-spawning
176
+ runtime or event kernel — offline replay cannot accidentally spend tokens.
177
+
178
+ ### When to use which
179
+
180
+ | Situation | Use |
181
+ |-----------|-----|
182
+ | Rate-limit mid-run; inputs unchanged | `resume` |
183
+ | Repo file changed; re-pay only affected phases | `why-stale` → `recompute` |
184
+ | "Would a stricter gate have blocked last night's run?" | `trace` → `replay` with new `thresholds` |
185
+ | "Would a $0.10 cap have stopped the fan-out?" | `replay` with `budgetMaxUSD` |
186
+ | Need fresh model judgment under a new model id | `replay` will say `needs-live-rerun` → live `recompute`/`run` |
187
+
188
+ ---
189
+
190
+ ## Resume overrides (re-run one phase with a patch)
191
+
192
+ `taskflow_resume` accepts a `failed` or `paused` run and **forks a new
193
+ run** — the original run file is never
194
+ modified (the child carries `parentRunId`). To re-run exactly one phase with a
195
+ patched task/model/timeout/idleTimeout, pass override fields alongside
196
+ `phaseId`:
197
+
198
+ ```
199
+ taskflow_resume { runId: "<id>", phaseId: "audit",
200
+ task: "re-audit src/api with the new checklist",
201
+ model: "gpt-5" }
202
+ ```
203
+
204
+ The overrides apply to the child's def only; the parent is untouched. Without
205
+ overrides, ordinary resume re-runs the non-done phases.
206
+
207
+ ---
208
+
209
+ ## Pluggable verifiers — zero-token custom static checks
210
+
211
+ Beyond the built-in structural detectors (dead-ends, unreachable, gate-exhaustion,
212
+ budget-overflow, concurrency, ref-integrity, guard-contradictions, contracts),
213
+ Taskflow supports **pluggable verifiers**: pure functions that lint a flow's
214
+ declarations at compile time, before any model is spawned.
215
+
216
+ ### Built-in: script-lint
217
+
218
+ `compileTaskflow` auto-includes the **script-lint** verifier (opt out with
219
+ `lint: false`). It catches common shell mistakes in `script` phase `run`
220
+ commands:
221
+
222
+ - `grep` pattern starting with `-` without a `--` separator (exit 2, false RED)
223
+ - Unbalanced `[` or `(` in `grep`/`sed` regex (exit 2)
224
+ - Pipeline ending with a filter (`grep`/`awk`/`head`/`tail`/`wc`/`sort`)
225
+ without `set -o pipefail` or `PIPESTATUS` (failing upstream masked)
226
+
227
+ ### Custom verifiers (convention dir)
228
+
229
+ Drop a `.ts`/`.js`/`.mjs` file in `.pi/taskflows/verifiers/` (project) or
230
+ `~/.pi/taskflows/verifiers/` (user). Export a `TaskflowVerifier`:
231
+
232
+ ```ts
233
+ export default {
234
+ name: "my-check",
235
+ verify(flow) {
236
+ // flow.phases, flow.budget, flow.name are available.
237
+ // Return VerifierIssue[]: { phaseId?, message, severity: "error"|"warning" }.
238
+ return [];
239
+ },
240
+ };
241
+ ```
242
+
243
+ Project-scope verifiers shadow user-scope by `name`. Broken modules are
244
+ skipped with a warning (fail-open). Use `taskflow_lint` (MCP) or
245
+ `verifyTaskflow(flow, { verifiers })` (programmatic) to run them.
246
+
247
+ ### MCP: `taskflow_lint`
248
+
249
+ ```
250
+ taskflow_lint { "defineFile": "/tmp/flow.json" }
251
+ ```
252
+
253
+ Runs built-in + discovered verifiers. Plugin issues are stamped
254
+ `category: "plugin"` with `source: <verifier-name>`. Structural issues
255
+ are covered by `taskflow_verify`; `taskflow_lint` reports only plugin findings.
256
+
257
+ ---
258
+
259
+ ## `taskflow_version` — build/host identity
260
+
261
+ `taskflow_version` reports the engine package version, the git commit the dist
262
+ was built from, the run-state schema version, and the bound host
263
+ (`codex`/`claude`/`opencode`/`grok`/`hermes`). The git commit is stamped at build time —
264
+ `git` is never run at runtime.
265
+ ```
266
+ taskflow_version {}
267
+ ```