hermes-taskflow 0.2.9
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +447 -0
- package/dist/index.d.ts +13 -0
- package/dist/index.d.ts.map +1 -0
- package/dist/index.js +13 -0
- package/dist/index.js.map +1 -0
- package/dist/mcp/bin.d.ts +31 -0
- package/dist/mcp/bin.d.ts.map +1 -0
- package/dist/mcp/bin.js +39 -0
- package/dist/mcp/bin.js.map +1 -0
- package/dist/mcp/server.d.ts +16 -0
- package/dist/mcp/server.d.ts.map +1 -0
- package/dist/mcp/server.js +30 -0
- package/dist/mcp/server.js.map +1 -0
- package/package.json +60 -0
- package/plugin/assets/taskflow-small.svg +14 -0
- package/plugin/assets/taskflow.svg +17 -0
- package/plugin/hermes.config.snippet.yaml +27 -0
- package/plugin/skills/taskflow/SKILL.md +692 -0
- package/plugin/skills/taskflow/advanced.md +267 -0
- package/plugin/skills/taskflow/configuration.md +598 -0
- package/plugin/skills/taskflow/library.md +105 -0
- package/plugin/skills/taskflow/patterns.md +348 -0
|
@@ -0,0 +1,267 @@
|
|
|
1
|
+
<!-- GENERATED FILE — do not edit. Source: skills-src/taskflow/advanced.md (npm run build:skills) -->
|
|
2
|
+
|
|
3
|
+
# Taskflow Advanced — dynamic sub-flows & workspace isolation
|
|
4
|
+
|
|
5
|
+
Load this when a flow needs: runtime-generated work (`flow{def}` / `expand`) or
|
|
6
|
+
isolated working directories (`cwd: temp/dedicated/worktree`).
|
|
7
|
+
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
## `flow{def}` vs `expand` (when to use which)
|
|
11
|
+
|
|
12
|
+
| Need | Prefer |
|
|
13
|
+
|------|--------|
|
|
14
|
+
| Saved reusable flow by name | `flow` + `use` |
|
|
15
|
+
| Planner JSON as isolated nested sub-flow (classic) | `flow` + `def` **or** `expand` + `expandMode: "nested"` |
|
|
16
|
+
| Fragment phases must appear on the **parent** run as `<expandId>-<childId>` | `expand` + `expandMode: "graft"` |
|
|
17
|
+
| First of several static approaches (latency) | `race` (not tournament) |
|
|
18
|
+
|
|
19
|
+
`expand` is a first-class phase type (Horizon B). Dynamic validation / nesting /
|
|
20
|
+
breadth caps match `flow{def}`. **Event kernel** still excludes `race`/`expand`
|
|
21
|
+
(imperative path only until step handlers exist).
|
|
22
|
+
|
|
23
|
+
---
|
|
24
|
+
|
|
25
|
+
## Dynamic sub-flows (`flow{def}`) — the full contract
|
|
26
|
+
|
|
27
|
+
A `flow` phase with `def` resolves a sub-flow **at runtime**, usually from an
|
|
28
|
+
upstream phase's JSON output. The runtime interpolates + JSON-parses the `def`,
|
|
29
|
+
validates it, then runs it nested. This is how a planner decides at runtime
|
|
30
|
+
what work to spawn — with each generated plan checked before it spends a token.
|
|
31
|
+
|
|
32
|
+
```jsonc
|
|
33
|
+
{ "id": "plan", "type": "agent", "agent": "planner", "output": "json",
|
|
34
|
+
"task": "Scan the repo. Output ONLY JSON {\"name\":\"audit\",\"phases\":[...]} — one audit phase per file." },
|
|
35
|
+
{ "id": "run", "type": "flow", "def": "{steps.plan.json}", "optional": true,
|
|
36
|
+
"dependsOn": ["plan"], "final": true }
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
**LLM output contract for `def`** (put this in the planner's task):
|
|
40
|
+
- A *full* Taskflow `{"name":"...","phases":[...]}`, a bare `phases` array, or
|
|
41
|
+
`{"phases":[...]}` — pure JSON (a ```json fence is tolerated and stripped).
|
|
42
|
+
- Hyphens in ids, never underscores.
|
|
43
|
+
- Sub-flow phases reference each other in their **own** `{steps.x.output}`
|
|
44
|
+
namespace (no parent-id prefixing).
|
|
45
|
+
- An **empty** `phases` array is a valid no-op (the planner decided there's
|
|
46
|
+
nothing to do).
|
|
47
|
+
|
|
48
|
+
**Security caps on generated flows** (validation rejects; tell the planner not
|
|
49
|
+
to emit these so a retry isn't wasted):
|
|
50
|
+
- **No `script` phases** — shell execution from an LLM-authored plan is an RCE
|
|
51
|
+
vector; only author-written flows may use `script`.
|
|
52
|
+
- **No workspace `cwd` keywords** (`temp`/`dedicated`/`worktree`) and no `cwd`
|
|
53
|
+
escaping the run directory.
|
|
54
|
+
- Breadth caps: ≤100 phases, concurrency ≤16 (flow and per-phase).
|
|
55
|
+
- Depth: inline nesting capped at 5 (shared with `ctx_spawn` subflows).
|
|
56
|
+
|
|
57
|
+
**Fail-open semantics:** if the `def` doesn't parse, has the wrong shape, or
|
|
58
|
+
fails validation, the phase completes with `status: "done"` and a `defError`
|
|
59
|
+
diagnostic field; downstream phases receive empty output and the run continues.
|
|
60
|
+
Design for it:
|
|
61
|
+
- Add `optional: true` on the flow phase so a bad plan never aborts the run.
|
|
62
|
+
- Want a hard stop instead? Add a downstream gate:
|
|
63
|
+
`{ "type": "gate", "eval": ["{steps.run.output} != "], "task": "…VERDICT: BLOCK if the plan failed." }`
|
|
64
|
+
|
|
65
|
+
**Iterative replanning** — pair `flow{def}` with a `loop` whose body emits the
|
|
66
|
+
next plan from the previous round's result: the declarative equivalent of
|
|
67
|
+
`for (...) { read result; decide next }`. See `examples/dynamic-plan-execute.json`
|
|
68
|
+
and `examples/iterative-replan.json`.
|
|
69
|
+
|
|
70
|
+
---
|
|
71
|
+
|
|
72
|
+
## Workspace isolation (`cwd` keywords)
|
|
73
|
+
|
|
74
|
+
A phase's `cwd` is normally a literal path (or inherited from the run). Three
|
|
75
|
+
**reserved keywords** ask the runtime to allocate an isolated working directory
|
|
76
|
+
for the phase's subagent and tear it down afterwards — scratch work or file
|
|
77
|
+
mutation without touching the main tree:
|
|
78
|
+
|
|
79
|
+
| `cwd` value | what the runtime does | lifecycle |
|
|
80
|
+
|-------------|-----------------------|-----------|
|
|
81
|
+
| `"temp"` | ephemeral dir under the OS tmpdir | removed when the phase finishes |
|
|
82
|
+
| `"dedicated"` | persistent dir under the run state (`runs/ws/<runId>/<phaseId>`) | **kept** for inspection; deterministic per phase (resume reuses it) |
|
|
83
|
+
| `"worktree"` | `git worktree add` on a throwaway branch off `HEAD` | `git worktree remove` + branch delete when the phase finishes |
|
|
84
|
+
|
|
85
|
+
```jsonc
|
|
86
|
+
{ "id": "experiment", "type": "agent", "agent": "executor", "cwd": "worktree",
|
|
87
|
+
"task": "Try the risky refactor and run the tests. Your edits are isolated in a git worktree." }
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
- **Fail-open.** If allocation fails (e.g. `worktree` outside a git tree), the
|
|
91
|
+
phase degrades — `worktree`→`temp`, any other failure → the base cwd — with a
|
|
92
|
+
`warnings` diagnostic. A phase never fails to run because of isolation.
|
|
93
|
+
- **Security.** Keywords are honoured only in **author-written** flows; a
|
|
94
|
+
generated plan (`flow{def}` / `ctx_spawn` subflow) requesting one is rejected
|
|
95
|
+
at validation.
|
|
96
|
+
- A literal path passes through unchanged.
|
|
97
|
+
|
|
98
|
+
### Argument-selected cwd (0.2.1 compatibility bridge)
|
|
99
|
+
|
|
100
|
+
An author-written flow may set `cwd: "{args.package}"` only when `package` is
|
|
101
|
+
declared as `{ "type": "relative-path" }`. The placeholder must occupy the
|
|
102
|
+
whole field. The value is resolved below the invocation cwd, must be an existing
|
|
103
|
+
directory, and cannot escape through `..`, absolute paths, or symlinks at bind time.
|
|
104
|
+
|
|
105
|
+
The bridge is fail-closed and disabled unless the host operator explicitly sets
|
|
106
|
+
`TASKFLOW_CWD_BRIDGE_MODE=resolve-only`. That mode performs a time-of-check path
|
|
107
|
+
validation but has no no-follow filesystem handle and is not an OS filesystem sandbox; each phase emits a warning stating the lower
|
|
108
|
+
guarantee. Cwd-bridge flow trees do not reuse output-only cache/resume entries,
|
|
109
|
+
because the current 0.2.x runtime cannot restore filesystem mutations on a cache hit. Generated
|
|
110
|
+
sub-flows cannot use the bridge. Saved-flow definitions are frozen for one
|
|
111
|
+
top-level execution, the invocation root identity is persisted for resume, and
|
|
112
|
+
a selected sub-flow inherits a non-expanding canonical boundary: nested literal
|
|
113
|
+
cwd and context files may narrow it but cannot escape it or allocate a workspace
|
|
114
|
+
provider.
|
|
115
|
+
|
|
116
|
+
Do not combine this bridge with `retry.max > 0`; validation rejects that
|
|
117
|
+
combination. After a failed writer, the filesystem outcome is unknown and an
|
|
118
|
+
operator must explicitly reconcile the invocation workspace before another
|
|
119
|
+
write can start.
|
|
120
|
+
|
|
121
|
+
**Pattern — competing experiments in worktrees:** run two `parallel` branches,
|
|
122
|
+
each `cwd: "worktree"`, each attempting a different refactor strategy and
|
|
123
|
+
reporting its test results; a downstream gate/judge picks which diff to apply
|
|
124
|
+
for real. The main tree is never touched by the losers.
|
|
125
|
+
|
|
126
|
+
---
|
|
127
|
+
|
|
128
|
+
## Trace & offline replay (`trace` / `replay`) — vs resume / recompute
|
|
129
|
+
|
|
130
|
+
Three **different** reuse tools; do not conflate them:
|
|
131
|
+
|
|
132
|
+
| Tool | Spends tokens? | Mutates the run? | Answers |
|
|
133
|
+
|------|----------------|------------------|---------|
|
|
134
|
+
| **`resume`** | Only unfinished / cache-miss phases | **Forks a new run** (parent untouched; child carries `parentRunId`) | "Pick up where we stopped" |
|
|
135
|
+
| **`why-stale` → `recompute`** | Dry-run free; `--apply` / `dryRun:false` spends | Optional write of recompute result | "World/input changed — which phases re-run?" |
|
|
136
|
+
| **`trace` → `replay`** | **Never** | Never | "If the gate threshold / budget had been different, would we have blocked?" |
|
|
137
|
+
|
|
138
|
+
### Trace (read the evidence)
|
|
139
|
+
|
|
140
|
+
Every instrumented run may write an append-only **event log**
|
|
141
|
+
(`runs/<flow>/<runId>.trace.jsonl`): phase lifecycle, each subagent
|
|
142
|
+
input/output, and runtime **decisions** (gate verdict/score, when-guard,
|
|
143
|
+
cache-hit, budget-hit, tournament-winner, unreplayable).
|
|
144
|
+
|
|
145
|
+
```
|
|
146
|
+
taskflow_trace { runId: "<id>" }
|
|
147
|
+
taskflow_trace { runId: "<id>", json: true }
|
|
148
|
+
```
|
|
149
|
+
|
|
150
|
+
MCP trace responses are bounded. JSON mode returns an envelope with
|
|
151
|
+
`total`/`returned`/`truncated`; use `limit` (default 200, max 1000) to select the
|
|
152
|
+
newest events without flooding the host context.
|
|
153
|
+
|
|
154
|
+
If there is no log (pre-trace run, or no sink injected), the tool reports that
|
|
155
|
+
clearly — it never invents events.
|
|
156
|
+
|
|
157
|
+
### Offline replay (what-if, zero tokens)
|
|
158
|
+
|
|
159
|
+
`replay` **re-folds** the recorded log under alternate **decision knobs** without
|
|
160
|
+
calling any model:
|
|
161
|
+
|
|
162
|
+
- `thresholds` — map of `phaseId → new score threshold` (gate-score events)
|
|
163
|
+
- `budgetMaxUSD` / `budgetMaxTokens` — would later phases have been skipped?
|
|
164
|
+
- `models` / `args` — currently report `needs-live-rerun` (quality cannot be
|
|
165
|
+
re-judged offline without re-execution)
|
|
166
|
+
|
|
167
|
+
Outcomes per phase: `reused`, `would-block`, `verdict-flipped`,
|
|
168
|
+
`would-exceed-budget`, `threshold-changed`, `needs-live-rerun`, `failed`.
|
|
169
|
+
|
|
170
|
+
```
|
|
171
|
+
taskflow_replay { runId: "<id>", thresholds: { review: 0.9 } }
|
|
172
|
+
taskflow_replay { runId: "<id>", budgetMaxUSD: 0.05, json: true }
|
|
173
|
+
```
|
|
174
|
+
|
|
175
|
+
**Import-graph guarantee:** `replayRun` never imports the process-spawning
|
|
176
|
+
runtime or event kernel — offline replay cannot accidentally spend tokens.
|
|
177
|
+
|
|
178
|
+
### When to use which
|
|
179
|
+
|
|
180
|
+
| Situation | Use |
|
|
181
|
+
|-----------|-----|
|
|
182
|
+
| Rate-limit mid-run; inputs unchanged | `resume` |
|
|
183
|
+
| Repo file changed; re-pay only affected phases | `why-stale` → `recompute` |
|
|
184
|
+
| "Would a stricter gate have blocked last night's run?" | `trace` → `replay` with new `thresholds` |
|
|
185
|
+
| "Would a $0.10 cap have stopped the fan-out?" | `replay` with `budgetMaxUSD` |
|
|
186
|
+
| Need fresh model judgment under a new model id | `replay` will say `needs-live-rerun` → live `recompute`/`run` |
|
|
187
|
+
|
|
188
|
+
---
|
|
189
|
+
|
|
190
|
+
## Resume overrides (re-run one phase with a patch)
|
|
191
|
+
|
|
192
|
+
`taskflow_resume` accepts a `failed` or `paused` run and **forks a new
|
|
193
|
+
run** — the original run file is never
|
|
194
|
+
modified (the child carries `parentRunId`). To re-run exactly one phase with a
|
|
195
|
+
patched task/model/timeout/idleTimeout, pass override fields alongside
|
|
196
|
+
`phaseId`:
|
|
197
|
+
|
|
198
|
+
```
|
|
199
|
+
taskflow_resume { runId: "<id>", phaseId: "audit",
|
|
200
|
+
task: "re-audit src/api with the new checklist",
|
|
201
|
+
model: "gpt-5" }
|
|
202
|
+
```
|
|
203
|
+
|
|
204
|
+
The overrides apply to the child's def only; the parent is untouched. Without
|
|
205
|
+
overrides, ordinary resume re-runs the non-done phases.
|
|
206
|
+
|
|
207
|
+
---
|
|
208
|
+
|
|
209
|
+
## Pluggable verifiers — zero-token custom static checks
|
|
210
|
+
|
|
211
|
+
Beyond the built-in structural detectors (dead-ends, unreachable, gate-exhaustion,
|
|
212
|
+
budget-overflow, concurrency, ref-integrity, guard-contradictions, contracts),
|
|
213
|
+
Taskflow supports **pluggable verifiers**: pure functions that lint a flow's
|
|
214
|
+
declarations at compile time, before any model is spawned.
|
|
215
|
+
|
|
216
|
+
### Built-in: script-lint
|
|
217
|
+
|
|
218
|
+
`compileTaskflow` auto-includes the **script-lint** verifier (opt out with
|
|
219
|
+
`lint: false`). It catches common shell mistakes in `script` phase `run`
|
|
220
|
+
commands:
|
|
221
|
+
|
|
222
|
+
- `grep` pattern starting with `-` without a `--` separator (exit 2, false RED)
|
|
223
|
+
- Unbalanced `[` or `(` in `grep`/`sed` regex (exit 2)
|
|
224
|
+
- Pipeline ending with a filter (`grep`/`awk`/`head`/`tail`/`wc`/`sort`)
|
|
225
|
+
without `set -o pipefail` or `PIPESTATUS` (failing upstream masked)
|
|
226
|
+
|
|
227
|
+
### Custom verifiers (convention dir)
|
|
228
|
+
|
|
229
|
+
Drop a `.ts`/`.js`/`.mjs` file in `.pi/taskflows/verifiers/` (project) or
|
|
230
|
+
`~/.pi/taskflows/verifiers/` (user). Export a `TaskflowVerifier`:
|
|
231
|
+
|
|
232
|
+
```ts
|
|
233
|
+
export default {
|
|
234
|
+
name: "my-check",
|
|
235
|
+
verify(flow) {
|
|
236
|
+
// flow.phases, flow.budget, flow.name are available.
|
|
237
|
+
// Return VerifierIssue[]: { phaseId?, message, severity: "error"|"warning" }.
|
|
238
|
+
return [];
|
|
239
|
+
},
|
|
240
|
+
};
|
|
241
|
+
```
|
|
242
|
+
|
|
243
|
+
Project-scope verifiers shadow user-scope by `name`. Broken modules are
|
|
244
|
+
skipped with a warning (fail-open). Use `taskflow_lint` (MCP) or
|
|
245
|
+
`verifyTaskflow(flow, { verifiers })` (programmatic) to run them.
|
|
246
|
+
|
|
247
|
+
### MCP: `taskflow_lint`
|
|
248
|
+
|
|
249
|
+
```
|
|
250
|
+
taskflow_lint { "defineFile": "/tmp/flow.json" }
|
|
251
|
+
```
|
|
252
|
+
|
|
253
|
+
Runs built-in + discovered verifiers. Plugin issues are stamped
|
|
254
|
+
`category: "plugin"` with `source: <verifier-name>`. Structural issues
|
|
255
|
+
are covered by `taskflow_verify`; `taskflow_lint` reports only plugin findings.
|
|
256
|
+
|
|
257
|
+
---
|
|
258
|
+
|
|
259
|
+
## `taskflow_version` — build/host identity
|
|
260
|
+
|
|
261
|
+
`taskflow_version` reports the engine package version, the git commit the dist
|
|
262
|
+
was built from, the run-state schema version, and the bound host
|
|
263
|
+
(`codex`/`claude`/`opencode`/`grok`/`hermes`). The git commit is stamped at build time —
|
|
264
|
+
`git` is never run at runtime.
|
|
265
|
+
```
|
|
266
|
+
taskflow_version {}
|
|
267
|
+
```
|