grok-taskflow 0.2.0 → 0.2.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +247 -862
- package/dist/mcp/server.js +3 -3
- package/dist/mcp/server.js.map +1 -1
- package/package.json +5 -5
package/README.md
CHANGED
|
@@ -1,1008 +1,393 @@
|
|
|
1
1
|
<div align="center">
|
|
2
2
|
|
|
3
|
-
<img src="https://raw.githubusercontent.com/heggria/taskflow/main/assets/hero.png" alt="taskflow
|
|
3
|
+
<img src="https://raw.githubusercontent.com/heggria/taskflow/main/assets/hero.png" alt="taskflow: compile, verify, and run multi-agent DAGs across five coding-agent hosts" width="100%">
|
|
4
4
|
|
|
5
|
-
<
|
|
6
|
-
<a href="https://www.npmjs.com/package/pi-taskflow"><img src="https://img.shields.io/npm/v/pi-taskflow?style=flat-square&color=4B4ACF&label=npm" alt="npm version"></a>
|
|
7
|
-
<a href="https://www.npmjs.com/package/pi-taskflow"><img src="https://img.shields.io/npm/dm/pi-taskflow?style=flat-square&color=5A5D63&label=downloads" alt="npm downloads"></a>
|
|
8
|
-
<a href="https://github.com/heggria/taskflow/blob/main/LICENSE"><img src="https://img.shields.io/badge/license-MIT-0E8A66?style=flat-square" alt="MIT license"></a>
|
|
9
|
-
<a href="https://github.com/heggria/taskflow/actions/workflows/ci.yml"><img src="https://img.shields.io/github/actions/workflow/status/heggria/taskflow/ci.yml?branch=main&style=flat-square&label=CI" alt="CI status"></a>
|
|
10
|
-
<a href="#whats-inside"><img src="https://img.shields.io/badge/tests-1500+-4B4ACF?style=flat-square" alt="1500+ tests"></a>
|
|
11
|
-
<a href="#whats-inside"><img src="https://img.shields.io/badge/dogfooded-%E2%9C%93-0E8A66?style=flat-square" alt="dogfooded"></a>
|
|
12
|
-
<a href="#run-it-on-your-agent"><img src="https://img.shields.io/badge/runs%20on-Pi%20%2B%20Codex%20%2B%20Claude%20Code%20%2B%20OpenCode%20%2B%20Grok-4B4ACF?style=flat-square" alt="runs on Pi, Codex, Claude Code, OpenCode, and Grok Build"></a>
|
|
13
|
-
</p>
|
|
5
|
+
<br />
|
|
14
6
|
|
|
15
|
-
|
|
7
|
+
[](https://www.npmjs.com/package/pi-taskflow)
|
|
8
|
+
[](https://github.com/heggria/taskflow/actions/workflows/ci.yml)
|
|
9
|
+
[](https://nodejs.org)
|
|
10
|
+
[](https://github.com/heggria/taskflow/blob/main/LICENSE)
|
|
11
|
+
[](#install-on-your-host)
|
|
12
|
+
[](#built-to-survive-real-work)
|
|
16
13
|
|
|
17
|
-
|
|
18
|
-
<b>English</b> ·
|
|
19
|
-
<a href="https://github.com/heggria/taskflow/blob/main/README.zh-CN.md">简体中文</a>
|
|
20
|
-
</p>
|
|
14
|
+
**English** · [简体中文](https://github.com/heggria/taskflow/blob/main/README.zh-CN.md)
|
|
21
15
|
|
|
22
|
-
|
|
23
|
-
<a href="https://heggria.github.io/taskflow/en"><img src="https://img.shields.io/badge/📖_Read_the_docs-heggria.github.io%2Ftaskflow-4B4ACF?style=for-the-badge&labelColor=2D2F5A" alt="Read the docs — heggria.github.io/taskflow"></a>
|
|
24
|
-
</p>
|
|
25
|
-
|
|
26
|
-
<p><strong>A declarative, verifiable <em>graph of tasks</em> for coding-agent subagents.</strong><br/>
|
|
27
|
-
Not a workflow you script — a DAG you declare. Fan out · gate · loop · tournament · resume · save as a command — intermediate results stay out of your context.<br/>
|
|
28
|
-
Runs on the <a href="https://pi.dev">Pi</a> coding agent, on <a href="https://github.com/openai/codex">OpenAI Codex</a>, on <a href="https://claude.com/product/claude-code">Claude Code</a>, on <a href="https://opencode.ai">OpenCode</a>, and on <a href="https://docs.x.ai/build/overview">Grok Build</a>.</p>
|
|
16
|
+
[Install](#install-on-your-host) · [Quickstart](#60-second-start) · [What's new in 0.2](#02-is-the-compiler-turn) · [Docs](https://heggria.github.io/taskflow/en/docs) · [Examples](https://github.com/heggria/taskflow/blob/main/examples)
|
|
29
17
|
|
|
30
18
|
</div>
|
|
31
19
|
|
|
32
|
-
```bash
|
|
33
|
-
# Pi
|
|
34
|
-
pi install npm:pi-taskflow
|
|
35
|
-
|
|
36
|
-
# Codex
|
|
37
|
-
codex plugin marketplace add heggria/taskflow
|
|
38
|
-
codex plugin add taskflow@taskflow
|
|
39
|
-
|
|
40
|
-
# Claude Code
|
|
41
|
-
claude plugin marketplace add heggria/taskflow
|
|
42
|
-
claude plugin install claude-taskflow@taskflow
|
|
43
|
-
|
|
44
|
-
# OpenCode — add the MCP server to opencode.json (see the OpenCode guide)
|
|
45
|
-
opencode mcp add taskflow -- npx -y -p opencode-taskflow opencode-taskflow-mcp
|
|
46
|
-
|
|
47
|
-
# Grok Build (published MCP package)
|
|
48
|
-
# First define custom taskflow-workspace/taskflow-readonly profiles extending
|
|
49
|
-
# workspace/read-only respectively in ~/.grok/sandbox.toml, then:
|
|
50
|
-
export PI_TASKFLOW_GROK_MUTATING_SANDBOX_PROFILE=taskflow-workspace
|
|
51
|
-
export PI_TASKFLOW_GROK_READONLY_SANDBOX_PROFILE=taskflow-readonly
|
|
52
|
-
grok mcp add taskflow -- npx -y -p grok-taskflow@0.2.0 grok-taskflow-mcp
|
|
53
|
-
```
|
|
54
|
-
|
|
55
20
|
---
|
|
56
21
|
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
You already know your agent's built-in subagent shorthand — `task` / `tasks` / `chain`. `taskflow` speaks the *same* shorthand — so your existing delegations instantly become **tracked, resumable, and saveable by name** (on Pi, a saved flow becomes a one-word `/tf:<name>` command; on Codex, Claude Code, OpenCode, and Grok Build you run it by name through `taskflow_run`). When you outgrow the shorthand, the full DSL gives you a real DAG: dynamic fan-out over dozens of items, conditional routing, quality gates, human approvals, retries, loops, tournaments, and an observed-usage budget stop-loss.
|
|
22
|
+
# Build multi-agent systems you can inspect before they run.
|
|
60
23
|
|
|
61
|
-
|
|
24
|
+
**taskflow turns agent plans into compiled task graphs**: declared once, verified before model spend, executed in isolated subagents, resumed across sessions, replayed without tokens, and recomputed from the smallest stale frontier.
|
|
62
25
|
|
|
63
|
-
|
|
26
|
+
It runs on the coding agent you already use:
|
|
64
27
|
|
|
65
|
-
|
|
28
|
+
**Pi · Codex · Claude Code · OpenCode · Grok Build**
|
|
66
29
|
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
Here's the wall you hit with raw subagents: you describe a multi-step plan in prose, the model re-derives it every single run, the intermediate transcripts flood your context, and the moment one model call fails you start over from zero. There's no reuse, no recovery, no structure — and no way to *check* the plan before it burns tokens.
|
|
30
|
+
```text
|
|
31
|
+
JSON or .tf.ts
|
|
32
|
+
│
|
|
33
|
+
▼
|
|
34
|
+
validate ──► Taskflow JSON ──► FlowIR + content hash
|
|
35
|
+
│
|
|
36
|
+
▼
|
|
37
|
+
isolated DAG runtime
|
|
38
|
+
│
|
|
39
|
+
┌────────────┼────────────┐
|
|
40
|
+
▼ ▼ ▼
|
|
41
|
+
resume replay recompute
|
|
42
|
+
```
|
|
81
43
|
|
|
82
|
-
|
|
44
|
+
> Your host receives the final result. Intermediate transcripts stay inside the runtime unless you explicitly inspect them.
|
|
83
45
|
|
|
84
|
-
|
|
85
|
-
<img src="https://raw.githubusercontent.com/heggria/taskflow/main/assets/context-isolation.png" alt="With raw subagents every transcript floods your context; with taskflow transcripts stay in the runtime and only the final result returns" width="900">
|
|
86
|
-
</div>
|
|
46
|
+
## Why taskflow?
|
|
87
47
|
|
|
88
|
-
|
|
48
|
+
Built-in subagent tools are excellent for one turn. The moment the work branches, retries, crosses sessions, or needs a quality gate, the plan becomes infrastructure.
|
|
89
49
|
|
|
90
|
-
| |
|
|
91
|
-
|---|---|---|
|
|
92
|
-
| **Who drives** | the model, turn by turn | the runtime, from a definition |
|
|
93
|
-
| **Topology** | chain / flat parallel | **DAG with layered concurrency + routing** |
|
|
94
|
-
| **Intermediate results** | in your context window | **in the runtime — not your context** |
|
|
95
|
-
| **Scale** | a handful of tasks | **dynamic `map` fan-out over dozens of items** |
|
|
96
|
-
| **Reusable** | re-described every time | **saved by name (`/tf:<name>` on Pi; `taskflow_run` by name on Codex)** |
|
|
97
|
-
| **Resumable** | ✗ | **✓ cross-session — cached phases auto-skip** |
|
|
98
|
-
| **Quality gates** | ✗ | **`gate` phases that halt on `VERDICT: BLOCK`** |
|
|
99
|
-
| **Conditional routing** | ✗ | **`when` guards + `join: any` OR-joins** |
|
|
100
|
-
| **Fault tolerance** | ✗ | **per-phase `retry` + auto-retry on transient errors** |
|
|
101
|
-
| **Human-in-the-loop** | ✗ | **`approval` phases (approve / reject / edit)** |
|
|
102
|
-
| **Cost control** | ✗ | **run-wide observed-usage `budget` stop-losses (USD / tokens)** |
|
|
103
|
-
| **Composition** | ✗ | **`flow` phases run saved *or runtime-generated* sub-flows** |
|
|
104
|
-
| **Iterative loops** | ✗ | **`loop` phases — repeat until condition, convergence, or cap** |
|
|
105
|
-
| **Competitive selection** | ✗ | **`tournament` phases — N variants + judge** |
|
|
106
|
-
| **Live progress** | opaque while running | **live DAG render with timing + cost (Pi `/tf`); one streaming tool call on Codex** |
|
|
107
|
-
| **Ergonomics** | inline JSON each time | **shorthand (`task`/`tasks`/`chain`) *or* DSL** |
|
|
108
|
-
|
|
109
|
-
It doesn't replace the subagent tool. It gives your subagents a **graph**, a memory, and a name.
|
|
110
|
-
|
|
111
|
-
## Declarative graph vs. imperative script
|
|
112
|
-
|
|
113
|
-
The closest thing to `taskflow` in spirit is the **dynamic / code-mode workflow** — where the model writes a JavaScript orchestration script. It's powerful and genuinely expressive. But it sits at the *opposite* end of one fundamental axis: **expressivity vs. verifiability.**
|
|
114
|
-
|
|
115
|
-
| | dynamic `workflow` (code-mode) | **`taskflow`** (declarative graph) |
|
|
50
|
+
| | Ad-hoc agents / scripts | **taskflow** |
|
|
116
51
|
|---|---|---|
|
|
117
|
-
| **
|
|
118
|
-
| **
|
|
119
|
-
| **
|
|
120
|
-
| **
|
|
121
|
-
| **
|
|
122
|
-
| **
|
|
123
|
-
| **Expressivity ceiling** | **higher** — arbitrary control flow | bounded by the DSL, but `map`/`when`/`loop`/`gate`/`tournament` — plus **runtime-generated sub-flows (`flow {def}`)** for plan-then-execute and iterative replanning — cover most jobs |
|
|
124
|
-
|
|
125
|
-
We chose the **verifiable** side on purpose. The expressivity you give up is real; what you get back — a plan you can check, watch, replay, and safely let a model author — is what turns one-off prompting into durable orchestration.
|
|
126
|
-
|
|
127
|
-
## Compared to other Pi extensions
|
|
128
|
-
|
|
129
|
-
> This section is **Pi-specific** — it maps `pi-taskflow` against other packages in the Pi ecosystem. If you're on Codex, skip to [Phase types](#phase-types); the engine and DSL are identical.
|
|
52
|
+
| **Plan** | Re-derived from prose or hidden in a script | **An explicit, versionable DAG** |
|
|
53
|
+
| **Before execution** | Discover mistakes while spending | **Verify structure at zero model calls** |
|
|
54
|
+
| **Intermediate output** | Floods the host context | **Stays isolated in the runtime** |
|
|
55
|
+
| **Failure** | Start over or reconstruct state | **Resume from persisted phase state** |
|
|
56
|
+
| **Changed input** | Re-run broadly | **Explain staleness and re-run the affected frontier** |
|
|
57
|
+
| **Portability** | Coupled to one agent | **One JSON contract across five hosts** |
|
|
130
58
|
|
|
131
|
-
The
|
|
59
|
+
The trade is deliberate: less arbitrary orchestration code, more **verifiability, observability, recovery, and reuse**.
|
|
132
60
|
|
|
133
|
-
|
|
134
|
-
|---|---|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|
|
|
135
|
-
| **taskflow** | **declarative multi-phase taskflows** | **✓** | **✓** | **✓ `map`** | **✓ phase-hash** | **✓** | **✓** | **✓ `/tf:<name>`** | **✕ (1 + peers)** |
|
|
136
|
-
| [`@pi-agents/orchid`](https://www.npmjs.com/package/@pi-agents/orchid) | opinionated 9-phase pipeline + Ralph loop | fixed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✕ (2) |
|
|
137
|
-
| [`pi-crew`](https://www.npmjs.com/package/pi-crew) | role teams + git worktrees + async | partial | ✓ | ✓ | ✓ | ✓ | ✓ | – | ✕ (7) |
|
|
138
|
-
| [`ultimate-pi`](https://www.npmjs.com/package/ultimate-pi) | governed plan→execute→review harness | YAML contracts | ✓ (plan-time) | ✕ | ✓ | ✓ (3-tier) | ✓ | ✓ | ✕ (16) |
|
|
139
|
-
| [`@zhushanwen/pi-workflow`](https://www.npmjs.com/package/@zhushanwen/pi-workflow) | JS scripts (`agent`/`parallel`/`pipeline`) | yes (JS) | ✕ (linear) | ✓ | ✓ | ✕ | ✕ | ✓ (call cache) | ✓ |
|
|
140
|
-
| [`@fiale-plus/pi-rogue-orchestration`](https://www.npmjs.com/package/@fiale-plus/pi-rogue-orchestration) | timer loop + goal resolution | ✕ | ✕ | ✕ | ✓ | ✓ (goal-check) | ✕ | ✕ | ✓ |
|
|
141
|
-
| [`pi-subagents`](https://www.npmjs.com/package/pi-subagents) | single / parallel / chain delegation | ✕ | ✕ | static | – | ✕ | clarify | named workflows | ✕ (3) |
|
|
142
|
-
| [`@gotgenes/pi-subagents`](https://www.npmjs.com/package/@gotgenes/pi-subagents) | Claude-Code-style subagents + worktrees | ✕ | ✕ | ✕ | ✓ (by id) | ✕ | per-agent | ✕ | ✕ (1) |
|
|
143
|
-
| [`pi-pipeline`](https://www.npmjs.com/package/pi-pipeline) | fixed SPEC→PLAN→TASKS→VERIFY | ✕ | fixed | ✕ | session planning | ✓ | clarify | ✕ | ✕ (2) |
|
|
144
|
-
| [`pi-agent-flow`](https://www.npmjs.com/package/pi-agent-flow) | one-shot parallel specialist `fork` | yes | ✕ | ✕ | – | ✕ | ✕ | – | ✕ (2) |
|
|
61
|
+
## 60-second start
|
|
145
62
|
|
|
146
|
-
|
|
147
|
-
|
|
148
|
-
**How to choose:**
|
|
149
|
-
|
|
150
|
-
- **`@pi-agents/orchid`** is the most feature-complete orchestrator in the ecosystem (DAG + worktrees + Ralph loop + agent mailbox) — but its DSL is a *fixed* 9-phase pipeline, it carries runtime deps + jiti, and it's beta. Reach for `taskflow` when you want to **define your own graph** (not adopt an opinionated one) with **no host-SDK coupling** and a one-command install.
|
|
151
|
-
- **`pi-crew` / `ultimate-pi`** go heavier — worktree isolation, durable async teams, multi-tier governance. If you want a lightweight declarative engine with no host-SDK coupling, that's this project.
|
|
152
|
-
- **`@zhushanwen/pi-workflow`** is the closest in spirit and also zero-dep, but it's the **imperative** side of the split above: you author workflows as **JavaScript scripts** the model writes and runs. `taskflow`'s **declarative JSON DAG** is the verifiable side — statically checkable, visualizable, safe to LLM-generate, and resumable at phase granularity rather than call-cache dedup.
|
|
153
|
-
- **`@fiale-plus/pi-rogue-orchestration`** has a real **loop-until-done** (goal-driven iteration). `taskflow` now ships its own `loop` phase (v0.0.13+) plus `tournament` for competitive selection — and unlike rogue-orchestration, `taskflow` has a full DAG with gates, compositional sub-flows, and cross-session resume. For raw "keep going until the goal is met" with minimal structure, rogue-orchestration is still lighter; for structured, branching pipelines, `taskflow` covers the same ground and more.
|
|
154
|
-
- **`pi-subagents` / `@gotgenes/pi-subagents`** are the mature picks for ad-hoc "use reviewer on this diff" delegation and background jobs. `taskflow` is for when those delegations need to become a *repeatable, resumable pipeline*.
|
|
155
|
-
- **`pi-pipeline` / `pi-agent-flow`** ship *opinionated, fixed* flows. `taskflow` ships an *empty canvas*: you (or the model) declare the graph that fits the job.
|
|
156
|
-
|
|
157
|
-
> The honest one-liner: **`pi-taskflow` gives you a *declarative, verifiable, resumable* DAG of task nodes — saved as a one-word `/tf:<name>` command, with context isolation by design** (and the same engine runs on MCP hosts). The engine avoids host-SDK coupling; `typebox` is a peer dependency, the TypeScript DSL includes the compiler, and delivery packages depend on the internal taskflow packages.
|
|
158
|
-
|
|
159
|
-
## 30-second start
|
|
160
|
-
|
|
161
|
-
### On Pi
|
|
162
|
-
|
|
163
|
-
**1. Install** — one command:
|
|
63
|
+
Install taskflow on [Pi](https://pi.dev):
|
|
164
64
|
|
|
165
65
|
```bash
|
|
166
66
|
pi install npm:pi-taskflow
|
|
167
67
|
```
|
|
168
68
|
|
|
169
|
-
|
|
170
|
-
> (`fast`, `strong`, `thinker`, …) to your own models — an interactive picker.
|
|
171
|
-
> Skip it and agents just use Pi's default model. See [Model roles](#model-roles).
|
|
172
|
-
|
|
173
|
-
**2. Run** — just ask the model in a Pi session:
|
|
174
|
-
|
|
175
|
-
> *Run a chain: first explore the auth flow, then summarize the findings.*
|
|
176
|
-
|
|
177
|
-
The model calls the `taskflow` tool automatically. You get live progress, per-step timing, token cost, and a saved run record — **same effort as the built-in tool, now tracked and resumable.**
|
|
178
|
-
|
|
179
|
-
**3. Save** — say *"save it"* and you have `/tf:<name>` forever.
|
|
180
|
-
|
|
181
|
-
That's it. You can be running your first workflow before your coffee cools — without writing a single phase definition.
|
|
182
|
-
|
|
183
|
-
<a id="run-it-on-your-agent"></a>
|
|
184
|
-
### On Codex
|
|
185
|
-
|
|
186
|
-
taskflow ships as a Codex **plugin** — install it once and the `taskflow_*` MCP tools plus a routing skill light up automatically, no manual `mcp add` and no config editing:
|
|
187
|
-
|
|
188
|
-
```bash
|
|
189
|
-
codex plugin marketplace add heggria/taskflow
|
|
190
|
-
codex plugin add taskflow@taskflow
|
|
191
|
-
```
|
|
192
|
-
|
|
193
|
-
The plugin's MCP server runs via `npx` (a version-pinned `codex-taskflow`), so there's nothing else to install globally and the plugin version binds the exact code that runs. Then just ask Codex to run a multi-phase or fan-out job and it calls the tools. See the [Codex guide](https://github.com/heggria/taskflow/blob/main/docs/codex-mcp.md).
|
|
194
|
-
|
|
195
|
-
### On Claude Code
|
|
196
|
-
|
|
197
|
-
taskflow ships as a Claude Code **plugin** too — install it once and the `taskflow_*` MCP tools plus a routing skill light up automatically, no manual `mcp add` and no config editing:
|
|
198
|
-
|
|
199
|
-
```bash
|
|
200
|
-
claude plugin marketplace add heggria/taskflow
|
|
201
|
-
claude plugin install claude-taskflow@taskflow
|
|
202
|
-
```
|
|
203
|
-
|
|
204
|
-
The plugin's MCP server runs via `npx` (a version-pinned `claude-taskflow`), so there's nothing else to install globally and the plugin version binds the exact code that runs. Each phase's subagent then runs as an isolated `claude -p` session. **Claude Code 2.1.169+ is required** for the safe-mode isolation contract. Just ask Claude Code to run a multi-phase or fan-out job and it calls the tools. See the [Claude Code guide](https://github.com/heggria/taskflow/blob/main/docs/claude-mcp.md).
|
|
69
|
+
Then ask naturally:
|
|
205
70
|
|
|
206
|
-
|
|
71
|
+
> Use taskflow to audit `src/api` in parallel and return one prioritized report.
|
|
207
72
|
|
|
208
|
-
|
|
73
|
+
The routing skill uses the same familiar `task` / `tasks` / `chain` shape:
|
|
209
74
|
|
|
210
|
-
```
|
|
211
|
-
opencode mcp add taskflow -- npx -y -p opencode-taskflow opencode-taskflow-mcp
|
|
212
|
-
```
|
|
213
|
-
|
|
214
|
-
```jsonc
|
|
215
|
-
// opencode.json
|
|
75
|
+
```json
|
|
216
76
|
{
|
|
217
|
-
"
|
|
218
|
-
|
|
219
|
-
|
|
220
|
-
"
|
|
221
|
-
"
|
|
222
|
-
|
|
77
|
+
"chain": [
|
|
78
|
+
{ "agent": "scout", "task": "Map the public API under src/api." },
|
|
79
|
+
{
|
|
80
|
+
"agent": "security-reviewer",
|
|
81
|
+
"task": "Audit this surface for missing auth and unsafe input boundaries:\n{previous.output}"
|
|
82
|
+
},
|
|
83
|
+
{
|
|
84
|
+
"agent": "reviewer",
|
|
85
|
+
"task": "Turn these findings into one prioritized report:\n{previous.output}"
|
|
223
86
|
}
|
|
224
|
-
|
|
87
|
+
]
|
|
225
88
|
}
|
|
226
89
|
```
|
|
227
90
|
|
|
228
|
-
|
|
229
|
-
|
|
230
|
-
### On Grok Build
|
|
231
|
-
|
|
232
|
-
The published path is the MCP package:
|
|
233
|
-
|
|
234
|
-
```toml
|
|
235
|
-
# ~/.grok/sandbox.toml
|
|
236
|
-
[profiles.taskflow-workspace]
|
|
237
|
-
extends = "workspace"
|
|
238
|
-
|
|
239
|
-
[profiles.taskflow-readonly]
|
|
240
|
-
extends = "read-only"
|
|
241
|
-
```
|
|
242
|
-
|
|
243
|
-
```bash
|
|
244
|
-
export PI_TASKFLOW_GROK_MUTATING_SANDBOX_PROFILE=taskflow-workspace
|
|
245
|
-
export PI_TASKFLOW_GROK_READONLY_SANDBOX_PROFILE=taskflow-readonly
|
|
246
|
-
grok mcp add taskflow -- npx -y -p grok-taskflow@0.2.0 grok-taskflow-mcp
|
|
247
|
-
```
|
|
248
|
-
|
|
249
|
-
A plugin scaffold is also available from a monorepo checkout:
|
|
91
|
+
That already gives you an isolated, tracked run. When the job needs real topology, declare the graph:
|
|
250
92
|
|
|
251
|
-
```
|
|
252
|
-
|
|
253
|
-
|
|
254
|
-
|
|
255
|
-
|
|
93
|
+
```json
|
|
94
|
+
{
|
|
95
|
+
"name": "audit-api",
|
|
96
|
+
"args": { "dir": { "default": "src/api" } },
|
|
97
|
+
"concurrency": 4,
|
|
98
|
+
"phases": [
|
|
99
|
+
{
|
|
100
|
+
"id": "discover",
|
|
101
|
+
"type": "agent",
|
|
102
|
+
"agent": "scout",
|
|
103
|
+
"task": "List source files under {args.dir}. Output ONLY a JSON array of {\"path\":\"...\"} objects.",
|
|
104
|
+
"output": "json"
|
|
105
|
+
},
|
|
106
|
+
{
|
|
107
|
+
"id": "audit-each",
|
|
108
|
+
"type": "map",
|
|
109
|
+
"over": "{steps.discover.json}",
|
|
110
|
+
"as": "file",
|
|
111
|
+
"agent": "security-reviewer",
|
|
112
|
+
"task": "Audit {file.path}. Cite evidence and assign severity.",
|
|
113
|
+
"dependsOn": ["discover"]
|
|
114
|
+
},
|
|
115
|
+
{
|
|
116
|
+
"id": "report",
|
|
117
|
+
"type": "reduce",
|
|
118
|
+
"from": ["audit-each"],
|
|
119
|
+
"agent": "reviewer",
|
|
120
|
+
"task": "Synthesize one prioritized report:\n{steps.audit-each.output}",
|
|
121
|
+
"dependsOn": ["audit-each"],
|
|
122
|
+
"final": true
|
|
123
|
+
}
|
|
124
|
+
]
|
|
125
|
+
}
|
|
256
126
|
```
|
|
257
127
|
|
|
258
|
-
|
|
259
|
-
|
|
260
|
-
### The shorthand (same shape as the built-in tool)
|
|
261
|
-
|
|
262
|
-
```jsonc
|
|
263
|
-
// Single — one agent, one job
|
|
264
|
-
{ "task": "Summarize the architecture of src/", "agent": "explorer" }
|
|
265
|
-
|
|
266
|
-
// Parallel — fire several at once, outputs merge
|
|
267
|
-
{ "tasks": [
|
|
268
|
-
{ "task": "Audit auth in src/api", "agent": "analyst" },
|
|
269
|
-
{ "task": "Audit input validation in src/api", "agent": "analyst" }
|
|
270
|
-
] }
|
|
128
|
+
Save it as `.pi/taskflows/audit-api.json`, then run:
|
|
271
129
|
|
|
272
|
-
|
|
273
|
-
|
|
274
|
-
{ "task": "List the public API of src/lib", "agent": "scout" },
|
|
275
|
-
{ "task": "Write docs for:\n{previous.output}", "agent": "writer" }
|
|
276
|
-
] }
|
|
130
|
+
```text
|
|
131
|
+
/tf:audit-api dir=src/api
|
|
277
132
|
```
|
|
278
133
|
|
|
279
|
-
|
|
134
|
+
On Codex, Claude Code, OpenCode, and Grok Build, run the same saved definition by name through `taskflow_run`.
|
|
280
135
|
|
|
281
|
-
|
|
136
|
+
[Follow the full quickstart →](https://heggria.github.io/taskflow/en/docs/getting-started)
|
|
282
137
|
|
|
283
|
-
|
|
284
|
-
// Chain with context files injected into each step
|
|
285
|
-
{ "chain": [
|
|
286
|
-
{ "task": "List the public API", "agent": "scout", "context": ["src/lib/**/*.ts"] },
|
|
287
|
-
{ "task": "Write docs for:\n{previous.output}", "agent": "writer" }
|
|
288
|
-
] }
|
|
289
|
-
```
|
|
290
|
-
|
|
291
|
-
## Watch it run
|
|
138
|
+
## See the graph run
|
|
292
139
|
|
|
293
|
-
This is
|
|
140
|
+
This is real output from a Pi run—not a mock dashboard:
|
|
294
141
|
|
|
295
|
-
```
|
|
142
|
+
```text
|
|
296
143
|
⊗ taskflow self-improve 6/7 · blocked · $0.095
|
|
297
144
|
✓ discover agent deepseek-v4-flash 10t ↑38k ↓6.7k $0.011
|
|
298
145
|
┌ ✓ write-runner-tests agent claude-sonnet-4-6 10t ↑13 ↓6.6k $0.020
|
|
299
146
|
├ ✓ write-store-tests agent claude-sonnet-4-6 10t ↑11 ↓10k $0.018
|
|
300
147
|
├ ✓ write-agents-tests agent claude-sonnet-4-6 10t ↑28 ↓13k $0.030
|
|
301
148
|
└ ✓ fix-stability agent claude-sonnet-4-6 10t ↑13 ↓3.9k $0.012
|
|
302
|
-
✓ verify gate BLOCK 3 type errors in test files
|
|
149
|
+
✓ verify gate BLOCK 3 type errors in test files
|
|
303
150
|
⊘ report reduce skipped · Gate blocked ↳ fix-stability
|
|
304
151
|
```
|
|
305
152
|
|
|
306
|
-
|
|
153
|
+
The layout **is** the DAG. Parallel rails expose concurrency; long edges expose dependencies; the gate explains why downstream work stopped. No separate control plane is required to understand the run.
|
|
307
154
|
|
|
308
|
-
|
|
309
|
-
- **Status icons** — `✓` done · `◐` running · `✗` failed · `⊘` skipped · `○` pending.
|
|
310
|
-
- **Rail `┌ ├ └`** — phases in the same DAG layer, running concurrently. The four `write-*`/`fix-stability` tasks fan out from `discover`. A blank gutter = a single-phase layer.
|
|
311
|
-
- **`↳`** — a long, layer-skipping dependency. `report` depends on the adjacent `verify` *and* on `fix-stability` two layers back, so only that skip edge is annotated.
|
|
312
|
-
- **Gate** — `verify` emitted `VERDICT: BLOCK`, so the runtime skipped `report` and ended the run as `blocked`, surfacing the reason inline.
|
|
313
|
-
- **Detail** — per phase: model, token counts (`↑`in `↓`out), cost, timing. Fan-out phases also show sub-task progress (`3/15 2✗ 8▸`).
|
|
155
|
+
## 0.2 is the compiler turn
|
|
314
156
|
|
|
315
|
-
|
|
157
|
+
Before 0.2, taskflow executed declarative graphs. Now the graph also has a compile-time frontend, a canonical intermediate representation, an append-only decision trace, offline replay, and incremental recompute.
|
|
316
158
|
|
|
317
|
-
|
|
159
|
+
### Author in JSON or TypeScript
|
|
318
160
|
|
|
319
|
-
|
|
161
|
+
JSON remains the portable runtime contract. For larger flows, `taskflow-dsl` adds a compile-time TypeScript authoring layer:
|
|
320
162
|
|
|
321
|
-
```
|
|
322
|
-
{
|
|
323
|
-
"name": "summarize-files",
|
|
324
|
-
"description": "Discover files, summarize each, produce one report",
|
|
325
|
-
"args": { "dir": { "default": "." } },
|
|
326
|
-
"concurrency": 8,
|
|
327
|
-
"phases": [
|
|
328
|
-
{ "id": "discover", "type": "agent", "agent": "scout",
|
|
329
|
-
"task": "List source files under {args.dir} (non-recursive).\nOutput ONLY a JSON array [{\"file\":\"\"}]. No prose.",
|
|
330
|
-
"output": "json" },
|
|
331
|
-
{ "id": "summarize", "type": "map",
|
|
332
|
-
"over": "{steps.discover.json}", "as": "item", "agent": "scout",
|
|
333
|
-
"task": "Read {item.file} and give a one-sentence summary.",
|
|
334
|
-
"dependsOn": ["discover"] },
|
|
335
|
-
{ "id": "report", "type": "reduce", "from": ["summarize"], "agent": "writer",
|
|
336
|
-
"task": "Combine into a short overview:\n{steps.summarize.output}",
|
|
337
|
-
"dependsOn": ["summarize"], "final": true }
|
|
338
|
-
]
|
|
339
|
-
}
|
|
340
|
-
```
|
|
163
|
+
```ts
|
|
164
|
+
import { agent, flow, json, map, reduce } from "taskflow-dsl";
|
|
341
165
|
|
|
342
|
-
|
|
343
|
-
|
|
344
|
-
3. **`report`** is a `reduce` — it merges every summary into one clean overview.
|
|
166
|
+
export default flow("audit", (ctx) => {
|
|
167
|
+
ctx.budget({ maxUSD: 2 });
|
|
345
168
|
|
|
346
|
-
|
|
169
|
+
const files = agent("List files under {args.dir}", {
|
|
170
|
+
agent: "scout",
|
|
171
|
+
output: json<{ path: string }[]>(),
|
|
172
|
+
});
|
|
347
173
|
|
|
348
|
-
|
|
174
|
+
const audits = map(files, (file) =>
|
|
175
|
+
agent(`Audit ${file.path}`, { agent: "security-reviewer" }),
|
|
176
|
+
);
|
|
349
177
|
|
|
350
|
-
|
|
351
|
-
|
|
352
|
-
|
|
353
|
-
|
|
354
|
-
|
|
355
|
-
|
|
356
|
-
"task": "Classify the bug. Output ONLY {\"severity\":\"high\"} or {\"severity\":\"low\"}." },
|
|
357
|
-
{ "id": "deep", "when": "{steps.triage.json.severity} == high", "dependsOn": ["triage"],
|
|
358
|
-
"agent": "executor-code", "task": "Root-cause and patch it.",
|
|
359
|
-
"retry": { "max": 2, "backoffMs": 500 } },
|
|
360
|
-
{ "id": "quick", "when": "{steps.triage.json.severity} == low", "dependsOn": ["triage"],
|
|
361
|
-
"agent": "executor-fast", "task": "Apply the quick fix." },
|
|
362
|
-
{ "id": "approve", "type": "approval", "join": "any", "dependsOn": ["deep", "quick"],
|
|
363
|
-
"task": "Review the fix before it ships." },
|
|
364
|
-
{ "id": "ship", "type": "agent", "dependsOn": ["approve"],
|
|
365
|
-
"task": "Open a PR with the change.", "final": true }
|
|
366
|
-
]
|
|
367
|
-
}
|
|
178
|
+
return reduce(
|
|
179
|
+
[audits],
|
|
180
|
+
(parts) => agent(`Write one report:\n${parts.audits.output}`),
|
|
181
|
+
{ final: true },
|
|
182
|
+
);
|
|
183
|
+
});
|
|
368
184
|
```
|
|
369
185
|
|
|
370
|
-
|
|
371
|
-
|
|
372
|
-
-
|
|
373
|
-
-
|
|
374
|
-
|
|
375
|
-
No scripting. No JavaScript `eval`. Just data the runtime executes — safe enough to run LLM-generated definitions directly.
|
|
376
|
-
|
|
377
|
-
### Loop until done
|
|
378
|
-
|
|
379
|
-
Some work is inherently iterative — refine a draft until a reviewer is satisfied, retry-and-improve until tests pass, converge on an answer:
|
|
380
|
-
|
|
381
|
-
```jsonc
|
|
382
|
-
{
|
|
383
|
-
"id": "refine",
|
|
384
|
-
"type": "loop",
|
|
385
|
-
"task": "Improve this draft (iteration {loop.iteration}). Previous attempt:\n{loop.lastOutput}\n\nReturn JSON {\"draft\":\"…\",\"done\":true|false}.",
|
|
386
|
-
"until": "{steps.refine.json.done} == true",
|
|
387
|
-
"output": "json",
|
|
388
|
-
"maxIterations": 6,
|
|
389
|
-
"convergence": true
|
|
390
|
-
}
|
|
391
|
-
```
|
|
392
|
-
|
|
393
|
-
See [Loop phases](#loop-until-done-loop) for the full reference.
|
|
394
|
-
|
|
395
|
-
### Plan, then execute (runtime sub-flows)
|
|
396
|
-
|
|
397
|
-
A planner decides *at runtime* what work to spawn — each iteration's plan depends on the previous result:
|
|
398
|
-
|
|
399
|
-
```jsonc
|
|
400
|
-
{
|
|
401
|
-
"name": "iterative-replan",
|
|
402
|
-
"phases": [
|
|
403
|
-
{ "id": "plan", "type": "agent", "agent": "planner",
|
|
404
|
-
"task": "Given the current state, output a JSON taskflow definition (with phases[]).",
|
|
405
|
-
"output": "json" },
|
|
406
|
-
{ "id": "execute", "type": "flow", "def": "{steps.plan.json}",
|
|
407
|
-
"dependsOn": ["plan"] }
|
|
408
|
-
]
|
|
409
|
-
}
|
|
410
|
-
```
|
|
411
|
-
|
|
412
|
-
The generated sub-flow is **validated** (no cycles, no dangling refs, no duplicate IDs) before a single token is spent. See [`examples/dynamic-plan-execute.json`](https://github.com/heggria/taskflow/blob/main/examples/dynamic-plan-execute.json) and [`examples/iterative-replan.json`](https://github.com/heggria/taskflow/blob/main/examples/iterative-replan.json).
|
|
413
|
-
|
|
414
|
-
### Tournament (compete and judge)
|
|
415
|
-
|
|
416
|
-
For open-ended creative or subjective work, spawn several competing variants and let a judge pick the best:
|
|
417
|
-
|
|
418
|
-
```jsonc
|
|
419
|
-
{
|
|
420
|
-
"id": "headline",
|
|
421
|
-
"type": "tournament",
|
|
422
|
-
"task": "Write a punchy headline for this launch post.",
|
|
423
|
-
"variants": 4,
|
|
424
|
-
"judge": "Pick the headline with the strongest hook and clearest promise.",
|
|
425
|
-
"mode": "best"
|
|
426
|
-
}
|
|
186
|
+
```bash
|
|
187
|
+
pnpm add -D taskflow-dsl
|
|
188
|
+
taskflow-dsl check audit.tf.ts
|
|
189
|
+
taskflow-dsl build audit.tf.ts --emit both
|
|
190
|
+
# → audit.taskflow.json + audit.flowir.json
|
|
427
191
|
```
|
|
428
192
|
|
|
429
|
-
|
|
430
|
-
|
|
431
|
-
## Phase types
|
|
432
|
-
|
|
433
|
-
| type | what it does | required fields |
|
|
434
|
-
|------|--------------|-----------------|
|
|
435
|
-
| `agent` | one subagent runs a single task | `task` |
|
|
436
|
-
| `parallel` | run `branches[]` concurrently | `branches` (array of `{task, agent?}`) |
|
|
437
|
-
| `map` | **fan out** over an array — one subagent per item, `{item}` bound | `over`, `task` |
|
|
438
|
-
| `gate` | quality/review step that can **halt the flow** | `task` |
|
|
439
|
-
| `reduce` | aggregate `from[]` phase outputs into one | `from`, `task` |
|
|
440
|
-
| `approval` | **human-in-the-loop** pause — approve / reject / edit | — |
|
|
441
|
-
| `flow` | run a **sub-flow** as one phase — a **saved** flow (`use`) or a **runtime-generated** one (`def`) | `use` \| `def` |
|
|
442
|
-
| `loop` | **iterate a task until done** — re-run a body until a condition, convergence, or a cap | `task`, `until` |
|
|
443
|
-
| `tournament` | **N variants compete**, a judge picks the best (or aggregates) | `task` \| `branches` |
|
|
444
|
-
| `script` | run a **shell command** — no LLM, zero tokens — capturing stdout as the phase output | `run` |
|
|
445
|
-
| `race` | **first successful** branch wins (optional `cancelLosers` abort) | `branches` (≥2) |
|
|
446
|
-
| `expand` | run a dynamic fragment (`nested` or `graft` promote) | `def` (+ `expandMode?`) |
|
|
193
|
+
`.tf.ts` is **compile-time only**. Hosts execute the emitted Taskflow JSON; they never interpret TypeScript.
|
|
447
194
|
|
|
448
|
-
###
|
|
195
|
+
### Compile to a contract you can reason about
|
|
449
196
|
|
|
450
|
-
|
|
197
|
+
FlowIR canonicalizes the graph and gives it a content hash. That compiled identity makes provenance and stale analysis inspectable, while the runtime adds content-addressed caching and deterministic tools:
|
|
451
198
|
|
|
452
|
-
|
|
|
453
|
-
|
|
454
|
-
| `
|
|
455
|
-
| `
|
|
456
|
-
| `
|
|
457
|
-
| `
|
|
458
|
-
| `
|
|
459
|
-
| `
|
|
460
|
-
| `
|
|
461
|
-
| `cwd` | Working directory for the subagent. A literal path, or a reserved keyword for **workspace isolation** — `"temp"` (ephemeral dir, removed after), `"dedicated"` (persistent dir under the run state, kept), `"worktree"` (a git worktree on a throwaway branch, removed after). Fail-open; rejected in LLM-authored sub-flows. |
|
|
462
|
-
| `context` | File paths to pre-read and inject into the agent prompt |
|
|
463
|
-
| `contextLimit` | Max chars per context file (default 8000) |
|
|
464
|
-
| `concurrency` | Fan-out cap for `map` / `parallel` (overrides the flow default) |
|
|
465
|
-
| `final` | Marks the result-bearing phase (else the last phase wins) |
|
|
466
|
-
| `optional` | A failure here does **not** abort the run |
|
|
467
|
-
| `shareContext` | Opt this phase's subagent into the **Shared Context Tree** (see below). Set `contextSharing: true` at the flow level to enable it for every phase |
|
|
468
|
-
| `cache` | `{ scope, ttl?, fingerprint? }` — cross-run memoization (see below) |
|
|
469
|
-
| `onBlock` | `"halt"` (default) or `"retry"` — what happens when a gate blocks |
|
|
470
|
-
| `eval` | Zero-token machine-checkable criteria that run *before* the LLM gate |
|
|
471
|
-
|
|
472
|
-
Flow-level keys: `name`, `description`, `args`, `concurrency` (default 8), `agentScope`, `contextSharing`, `strictInterpolation`, and `budget: { maxUSD?, maxTokens? }`.
|
|
473
|
-
|
|
474
|
-
### Shared Context Tree (blackboard + supervision)
|
|
475
|
-
|
|
476
|
-
> **Host scope in 0.2.0:** `ctx_read` / `ctx_write` / `ctx_report` /
|
|
477
|
-
> `ctx_spawn` tool injection is available through `pi-taskflow`. The Codex,
|
|
478
|
-
> Claude, OpenCode, and Grok runners do not inject these tools yet.
|
|
479
|
-
|
|
480
|
-
By default subagents are fully isolated — they share nothing and only return a
|
|
481
|
-
final string. Opt a phase in with `shareContext: true` (or `contextSharing: true`
|
|
482
|
-
flow-wide) to give its subagent four extra tools backed by a per-run, file-based
|
|
483
|
-
blackboard:
|
|
484
|
-
|
|
485
|
-
| tool | direction | use |
|
|
486
|
-
|------|-----------|-----|
|
|
487
|
-
| `ctx_write(key, value)` | horizontal | publish a finding so siblings/descendants reuse it (stop re-reading the same files) |
|
|
488
|
-
| `ctx_read(key?)` | horizontal | read findings visible to this node: its own + ancestors' + **completed** others' |
|
|
489
|
-
| `ctx_report(summary, structured?)` | vertical ↑ | report a result up to the parent |
|
|
490
|
-
| `ctx_spawn(assignments[])` | vertical ↓ | delegate child work at runtime; each assignment is a flat `{task}` **or** a `{subflow}` (a dependency-bearing DAG the runtime validates and runs nested). Child reports fold back into this phase's output |
|
|
491
|
-
|
|
492
|
-
The first two are a **horizontal blackboard** (siblings reuse expensive context);
|
|
493
|
-
the last two are a **vertical supervision tree** (a node delegates work and its
|
|
494
|
-
children report up). Everything is opt-in, fail-open, depth-capped (5 levels), size-bounded
|
|
495
|
-
(256KB per value, 256 keys per node, 16 spawn assignments max), and cleaned up
|
|
496
|
-
with the run — flows that don't opt in behave exactly as before.
|
|
497
|
-
|
|
498
|
-
```jsonc
|
|
499
|
-
{ "id": "survey", "type": "agent", "agent": "scout", "shareContext": true,
|
|
500
|
-
"task": "Map the API surface. ctx_write key 'endpoints' so the auditors don't re-scan." },
|
|
501
|
-
{ "id": "audit", "type": "map", "over": "{steps.survey.json}", "shareContext": true,
|
|
502
|
-
"dependsOn": ["survey"], "agent": "analyst",
|
|
503
|
-
"task": "ctx_read 'endpoints' for shared context, then audit {item} for missing auth." }
|
|
504
|
-
```
|
|
199
|
+
| Operation | What it answers | Model calls |
|
|
200
|
+
|---|---|---:|
|
|
201
|
+
| `verify` / `compile` | Is the graph structurally safe to run? | **0** |
|
|
202
|
+
| `ir` | What is the canonical graph and content hash? | **0** |
|
|
203
|
+
| `resume` | What unfinished work remains? (forks a new run; original untouched) | Only unfinished phases |
|
|
204
|
+
| `trace` | What calls and runtime decisions actually happened? | **0** to inspect |
|
|
205
|
+
| `replay` | What if thresholds or budgets had been different? | **0** |
|
|
206
|
+
| `why-stale` | What changed, and what depends on it? | **0** |
|
|
207
|
+
| `recompute` | What is the smallest observable affected frontier? | Only affected phases |
|
|
505
208
|
|
|
506
|
-
|
|
209
|
+
[Explore the compiler and runtime →](https://heggria.github.io/taskflow/en/docs/compiler-runtime/)
|
|
507
210
|
|
|
508
|
-
|
|
509
|
-
again. Phase 1 of the **taskflow library** adds sidecar `.meta.json` metadata
|
|
510
|
-
(`purpose`, `tags`, `phaseSignature`, `generality`, `reuseCount`) to saved
|
|
511
|
-
flows, plus a search tool that surfaces reusable flows before you write a new
|
|
512
|
-
one:
|
|
513
|
-
|
|
514
|
-
```jsonc
|
|
515
|
-
// MCP
|
|
516
|
-
{ "name": "taskflow_search", "arguments": { "query": "audit API endpoints for missing auth" } }
|
|
517
|
-
// Pi tool
|
|
518
|
-
{ "action": "search", "query": "审计接口鉴权" }
|
|
519
|
-
```
|
|
211
|
+
## One runtime, 12 phase types
|
|
520
212
|
|
|
521
|
-
|
|
522
|
-
|
|
523
|
-
|
|
524
|
-
`
|
|
525
|
-
`
|
|
526
|
-
|
|
527
|
-
is not configured.
|
|
528
|
-
|
|
529
|
-
When you run a flow you discovered via search, pass `reusedFromSearch: true`
|
|
530
|
-
(`taskflow_run` MCP, or `action=run` in Pi). That bumps the flow's
|
|
531
|
-
`reuseCount`, so high-quality reusable patterns rank higher over time. When you
|
|
532
|
-
write a new reusable flow, save it with metadata:
|
|
533
|
-
|
|
534
|
-
```jsonc
|
|
535
|
-
{ "name": "taskflow_save",
|
|
536
|
-
"arguments": {
|
|
537
|
-
"name": "audit-endpoints",
|
|
538
|
-
"definition": { ... },
|
|
539
|
-
"purpose": "Audit API endpoints for missing auth checks",
|
|
540
|
-
"tags": ["audit", "security", "auth"] } }
|
|
541
|
-
```
|
|
542
|
-
|
|
543
|
-
See `docs/rfc-library-reuse.md` for the design and `skills-src/taskflow/library.md`
|
|
544
|
-
for the agent-facing search→reuse→generalize→re-save workflow.
|
|
213
|
+
| Family | Phases | Use them for |
|
|
214
|
+
|---|---|---|
|
|
215
|
+
| **Work** | `agent` · `parallel` · `map` · `reduce` · `script` | Single tasks, static fan-out, dynamic fan-out, aggregation, zero-token shell steps |
|
|
216
|
+
| **Control** | `gate` · `approval` · `flow` · `loop` | Quality decisions, human checkpoints, composition, iterative refinement |
|
|
217
|
+
| **Selection** | `tournament` · `race` | Best-of-N quality or first-success latency |
|
|
218
|
+
| **Dynamic graph** | `expand` | Validate and execute a runtime-produced fragment, nested or grafted |
|
|
545
219
|
|
|
546
|
-
|
|
220
|
+
Across those phase types, the DSL provides dependencies, conditions, retries, timeouts, output contracts, budgets, workspace isolation, and explicit final-output selection. Each kind accepts only the fields that are safe and meaningful for it; freshness-sensitive phases are excluded from cross-run caching.
|
|
547
221
|
|
|
548
|
-
|
|
549
|
-
- **`join: "any"`** — an OR-join: the phase runs as soon as *one* dependency completes (default `"all"` waits for all).
|
|
550
|
-
- **`retry`** — `{ "max": 2, "backoffMs": 500, "factor": 2 }` retries a failing subagent with fixed or exponential backoff; usage is summed and the attempt count shows as `↻N` in the TUI. Transient provider errors (rate-limit / 5xx / timeout) **auto-retry even without an explicit policy**; hard errors don't.
|
|
551
|
-
- **`onBlock`** — `"halt"` (default) stops the run when a gate blocks. `"retry"` retries upstream phases when a gate blocks, instead of halting — a self-healing rework loop with budget and idle-watchdog guards and a nested recursion depth cap.
|
|
552
|
-
- **`eval`** — zero-token machine-checkable criteria that run *before* the LLM gate. If the eval check fails, the gate blocks without spawning an agent.
|
|
553
|
-
- **`score`** — graded, composable quality gates: deterministic scorers (`exact-match`, `contains`, `regex`, `json-schema`, `length-range`, `code-compiles`) run against a target string at **zero tokens** and combine via `all`/`any`/`weighted` against a `threshold`. Deterministic pass → auto-PASS with no LLM call **when the judge cannot veto** — no judge configured, or `weighted` where the deterministic score is a *lower bound* already clearing the threshold. With `all`/`any` + a judge, the judge always runs (its verdict is authoritative — it may check what scorers cannot, e.g. factuality). Deterministic fail → the optional LLM `judge` decides (**fail-closed** on unparseable output — issue #54), or the gate `task` runs with the scorer report appended, or — with no fallback — the gate **blocks explicitly**. The structured result is the gate's `.json` (`{steps.<gate>.json.combined}`, `.json.results`), so downstream phases can route on quality, not just pass/fail. LLM-generated dynamic sub-flows may not use `code-compiles` (compiler execution) or `regex` (ReDoS) scorers — same hardening class as the `script` block.
|
|
554
|
-
- **`idempotent: false`** — side-effect classification for phases with **irreversible effects** (webhook POSTs, deploys, DB writes): the implicit transient auto-retry is suppressed (an explicit `retry{}` is still honored — it's the author's declaration that repeats are acceptable) and the result is **never cached** in any scope (within-run resume, cross-run, `incremental`) — the phase re-runs every time. The phase state records `sideEffect: true` (rendered as ⚡). Default `true` — existing flows are unchanged.
|
|
555
|
-
- **`approval`** — pause for a human (Approve / Reject / Edit). Reject halts the flow; Edit injects the typed note as the phase output for downstream steps. Non-interactive runs (detached / CI) **auto-reject** (safety: approval gates are never bypassed).
|
|
556
|
-
- **`flow`** — `{ "type": "flow", "use": "deep-research", "with": { "topic": "{item}" } }` runs a **saved** flow as a phase (recursion is detected and rejected). Or **generate the sub-flow at runtime**: `{ "type": "flow", "def": "{steps.plan.json}" }` resolves an upstream phase's JSON output into a sub-flow, **validates it (cycles / dangling refs / duplicate ids / dead-ends), then runs it** — the number and shape of the generated phases is decided at runtime, not authored in advance. A malformed plan fails *open* (the phase is skipped with a `defError`, the run continues). This is how a planner decides *at runtime* what work to spawn — the declarative answer to a code-mode `for` loop, with each generated plan checked before it spends a token. Security hardening for LLM-generated sub-flows: breadth caps (100 phases, 200 map items, 16 concurrency), `cwd` containment, budget clamped to `min(child, parent)`, nesting cap (5 levels), and prototype-pollution defense (deep-cloned, `__proto__`/`constructor`/`prototype` stripped). Pair it with `loop` for **data-dependent iterative replanning** (round N's plan depends on round N-1's result). See [`examples/dynamic-plan-execute.json`](https://github.com/heggria/taskflow/blob/main/examples/dynamic-plan-execute.json) and [`examples/iterative-replan.json`](https://github.com/heggria/taskflow/blob/main/examples/iterative-replan.json).
|
|
222
|
+
[Read the phase reference →](https://heggria.github.io/taskflow/en/docs/syntax/phase-types)
|
|
557
223
|
|
|
558
|
-
|
|
224
|
+
## Runtime guarantees, not prompt conventions
|
|
559
225
|
|
|
560
|
-
|
|
226
|
+
### Verify before spend
|
|
561
227
|
|
|
562
|
-
|
|
563
|
-
{
|
|
564
|
-
"id": "refine",
|
|
565
|
-
"type": "loop",
|
|
566
|
-
"task": "Improve this draft (iteration {loop.iteration}). Previous attempt:\n{loop.lastOutput}\n\nReturn JSON {\"draft\":\"…\",\"done\":true|false}.",
|
|
567
|
-
"until": "{steps.refine.json.done} == true", // the iteration's own output is exposed here
|
|
568
|
-
"output": "json",
|
|
569
|
-
"maxIterations": 6, // default 10, hard cap 100 — the loop ALWAYS terminates
|
|
570
|
-
"convergence": true // default: stop early if an iteration's output is identical to the last
|
|
571
|
-
}
|
|
572
|
-
```
|
|
228
|
+
Cycles, dangling dependencies, invalid references, impossible joins, unsafe dynamic fragments, and configuration hazards are rejected or surfaced before the expensive work starts.
|
|
573
229
|
|
|
574
|
-
|
|
575
|
-
- **`until`** — evaluated after each iteration with the iteration's output exposed as `{steps.<thisId>.output}` / `.json`. Same operators as `when`. The loop stops the moment it's truthy.
|
|
576
|
-
- **Always terminates.** Four independent stops: `until` truthy, **convergence** (a fixed point — output identical to the previous iteration), **`maxIterations`** (hard-capped at 100), or a **failing iteration** (the phase fails with the partial output preserved). A malformed `until` **stops** the loop rather than spinning forever (fail-safe) and surfaces a warning on the phase.
|
|
577
|
-
- **Reflexion memory (`reflexion: true`)** — by default each iteration only sees the prior *output*; the **reason** it wasn't good enough (an `expect` contract violation, an error, the unmet `until`) is discarded, so models repeat mistakes. With `reflexion: true` every iteration after the first receives a structured failure summary of the prior one via the `{reflexion}` placeholder (auto-appended when absent, capped at 2000 chars): contract diagnostics like `$.done: required key is missing`, the error message, or the unmet stop condition, plus a truncated output snippet. Semantics shift to enable self-correction: **body failures become feedback instead of terminating the loop** — timeout/abort still hard-stop, and exhausting `maxIterations` on a failure still fails the phase (reflexion defers failure, never erases it). The last injected summary is persisted on the phase state for audit.
|
|
578
|
-
- The TUI shows `↻N` with the stop reason (`done` / `converged` / `max` / `failed`); usage is summed across iterations. Like `gate`/`approval`, `loop` is **excluded from `cross-run` cache** (each run must iterate fresh).
|
|
230
|
+
### Keep intermediate work out of the host context
|
|
579
231
|
|
|
580
|
-
|
|
232
|
+
Agent-running phases execute in isolated subagent processes; control and script phases stay inside the runtime. Upstream outputs are wired into downstream inputs internally. Only `finalOutput` returns to the host unless you explicitly use `peek` or `trace`.
|
|
581
233
|
|
|
582
|
-
|
|
234
|
+
### Survive sessions and failures
|
|
583
235
|
|
|
584
|
-
|
|
585
|
-
{
|
|
586
|
-
"id": "headline",
|
|
587
|
-
"type": "tournament",
|
|
588
|
-
"task": "Write a punchy headline for this launch post.",
|
|
589
|
-
"variants": 4, // spawn 4 competitors of the SAME task (default 3, max 20)
|
|
590
|
-
"judge": "Pick the headline with the strongest hook and clearest promise.",
|
|
591
|
-
"judgeAgent": "reviewer", // optional; defaults to the phase agent
|
|
592
|
-
"mode": "best" // "best" (default) | "aggregate"
|
|
593
|
-
}
|
|
594
|
-
```
|
|
236
|
+
Phase state is persisted atomically. Resume skips unchanged completed work; detached Pi runs can outlive the initiating session; an idle watchdog terminates stalled subagents.
|
|
595
237
|
|
|
596
|
-
|
|
597
|
-
- **Judge** — after the fan-out, one judge agent sees every variant (numbered) plus your `judge` rubric and picks a winner. Prefer a JSON pick `{"winner": n}` (most robust); the runtime also reads a `WINNER: <n>` line (`#n` and Markdown emphasis like `WINNER: **n**` are tolerated — issue #54). An unreadable verdict **fails open** to variant 1; a failed judge falls back too — the work is never lost.
|
|
598
|
-
- **`mode`** — `best` returns the winning variant **verbatim**; `aggregate` returns the judge's **synthesized** answer combining the strongest parts.
|
|
599
|
-
- **Short-circuits:** if only one competitor survives, it wins with no judge call; if all fail, the phase fails. The TUI shows `⚑ N→#k`; usage sums variants + judge. Like `gate`, it's **excluded from `cross-run` cache**.
|
|
238
|
+
### Reuse work honestly
|
|
600
239
|
|
|
601
|
-
|
|
240
|
+
Within-run resume is content-addressed. Cross-run caching is opt-in and can fingerprint Git commits, files, globs, environment variables, and TTLs. Change one declared input and only its dependents become stale.
|
|
602
241
|
|
|
603
|
-
|
|
242
|
+
### Bound the blast radius
|
|
604
243
|
|
|
605
|
-
|
|
606
|
-
{
|
|
607
|
-
"id": "build",
|
|
608
|
-
"type": "script",
|
|
609
|
-
"run": "pnpm run build", // string → runs in a shell
|
|
610
|
-
"timeout": 120000 // optional ms cap (1000–300000, default 60000)
|
|
611
|
-
},
|
|
612
|
-
{
|
|
613
|
-
"id": "score",
|
|
614
|
-
"type": "script",
|
|
615
|
-
"run": ["python", "score.py"], // array → direct exec, no shell (injection-safe)
|
|
616
|
-
"input": "{steps.analyze.output}", // optional — piped to stdin (interpolation-enabled)
|
|
617
|
-
"dependsOn": ["analyze"]
|
|
618
|
-
}
|
|
619
|
-
```
|
|
244
|
+
Budgets, concurrency caps, retries, timeouts, nesting limits, dynamic-graph breadth caps, path containment, non-idempotent phase classification, and fail-closed approval behavior are runtime semantics—not suggestions in a prompt.
|
|
620
245
|
|
|
621
|
-
|
|
622
|
-
- **`input`** — optional text piped to the command's stdin; supports interpolation (`{steps.X.output}`, `{args.X}`). If omitted, stdin is closed.
|
|
623
|
-
- **`timeout`** — optional millisecond cap (1000–300000, default 60000). On timeout the child gets `SIGTERM`, then `SIGKILL` after a grace period, and the phase fails.
|
|
624
|
-
- A non-zero exit **fails** the phase (stderr is captured); stdout is capped at 1 MB. `script` phases spend **zero tokens**, do not support `retry` or `output: "json"`, and are **excluded from `cross-run` cache** (a shell step may have side effects). The `compile` diagram renders them as `⚡ script`.
|
|
246
|
+
### 0.2.1: safe dynamic cwd and Pi terminal reaping
|
|
625
247
|
|
|
626
|
-
|
|
248
|
+
An invocation argument declared as `type: "relative-path"` may select a phase
|
|
249
|
+
working directory with the exact form `cwd: "{args.package}"`. The bridge is
|
|
250
|
+
default-off, requires host `resolve-only` authorization, and confines the
|
|
251
|
+
canonical directory to the invocation root. Absolute paths, concatenation, and
|
|
252
|
+
`{steps.*}` remain rejected; this compatibility bridge is not an OS sandbox.
|
|
253
|
+
Resolve-only writer phases within one invocation are serialized before durable
|
|
254
|
+
lease acquisition, so fan-out cannot self-timeout while separate processes
|
|
255
|
+
remain protected by cross-process leases.
|
|
627
256
|
|
|
628
|
-
|
|
257
|
+
Pi child agents no longer inherit ambient extensions by default. Trusted host
|
|
258
|
+
settings can use an explicit extension allowlist or opt back into legacy
|
|
259
|
+
inheritance. If a Pi child produces a validated final answer and terminal event
|
|
260
|
+
but an extension keeps the process alive, Taskflow waits a bounded grace window,
|
|
261
|
+
reaps the process group, and records `completionSource: "terminal-reap"` instead
|
|
262
|
+
of reporting a false timeout.
|
|
629
263
|
|
|
630
|
-
```
|
|
264
|
+
```json
|
|
631
265
|
{
|
|
632
|
-
"
|
|
633
|
-
|
|
634
|
-
|
|
635
|
-
|
|
636
|
-
|
|
637
|
-
|
|
638
|
-
"fingerprint": ["git:HEAD", "glob:src/auth/**/*.ts"] // fold world-state into the key
|
|
266
|
+
"taskflow": {
|
|
267
|
+
"piChild": {
|
|
268
|
+
"resourceProfile": "isolated",
|
|
269
|
+
"extensions": [],
|
|
270
|
+
"terminalGraceMs": 1500
|
|
271
|
+
}
|
|
639
272
|
}
|
|
640
273
|
}
|
|
641
274
|
```
|
|
642
275
|
|
|
643
|
-
|
|
644
|
-
|
|
645
|
-
|
|
646
|
-
- Cache lives in `.pi/taskflows/cache/` (gitignored). Clear it with `action: "cache-clear"` on the tool. Full rationale: [`docs/internal/rfc-cross-run-memoization.md`](https://github.com/heggria/taskflow/blob/main/docs/internal/rfc-cross-run-memoization.md).
|
|
647
|
-
|
|
648
|
-
### Gate phases (quality control)
|
|
276
|
+
`allowlist` accepts explicit trusted extension files; `inherit` restores ambient
|
|
277
|
+
Pi extension discovery as a compatibility mode. Flows cannot widen this host
|
|
278
|
+
authority.
|
|
649
279
|
|
|
650
|
-
|
|
280
|
+
[Read the core concepts →](https://heggria.github.io/taskflow/en/docs/concepts/)
|
|
651
281
|
|
|
652
|
-
|
|
653
|
-
- JSON like `{"continue": false, "reason": "missing auth checks"}` / `{"verdict": "block", "reason": "..."}`.
|
|
282
|
+
## Install on your host
|
|
654
283
|
|
|
655
|
-
|
|
284
|
+
All packages require **Node.js ≥ 22.19.0**.
|
|
656
285
|
|
|
657
|
-
|
|
658
|
-
Review the audit below. If any endpoint is missing auth, end with
|
|
659
|
-
"VERDICT: BLOCK" and a one-line reason; otherwise end with "VERDICT: PASS".
|
|
286
|
+
### Pi
|
|
660
287
|
|
|
661
|
-
|
|
662
|
-
|
|
663
|
-
|
|
664
|
-
## Interpolation & expressions
|
|
665
|
-
|
|
666
|
-
| placeholder | resolves to |
|
|
667
|
-
|---|---|
|
|
668
|
-
| `{args.X}` | invocation argument |
|
|
669
|
-
| `{steps.ID.output}` | a prior phase's text output |
|
|
670
|
-
| `{steps.ID.json}` | prior output parsed as JSON (or `{steps.ID.json.field}`) |
|
|
671
|
-
| `{item}` / `{item.field}` | current item inside a `map` phase |
|
|
672
|
-
| `{previous.output}` | the immediately-upstream phase output |
|
|
673
|
-
| `{loop.iteration}` | current iteration number inside a `loop` phase |
|
|
674
|
-
| `{loop.lastOutput}` | previous iteration's output inside a `loop` phase |
|
|
675
|
-
| `{loop.maxIterations}` | the iteration cap inside a `loop` phase |
|
|
676
|
-
|
|
677
|
-
Condition grammar (for `when`): `== != < > <= >=`, `&& || !`, parentheses, quoted strings/numbers, and any `{...}` reference — e.g. `"when": "{steps.triage.json.route} == deep && {args.force} != true"`.
|
|
678
|
-
|
|
679
|
-
> Referencing `{steps.X}` that isn't declared in `dependsOn` is a **hard validation error** — the runtime catches the most common pipeline bug before a single agent runs.
|
|
680
|
-
|
|
681
|
-
> Unresolved interpolation refs (e.g. `{args.typo}` or a missing `dependsOn`) are surfaced as **phase warnings** (`PhaseState.warnings`) in the run record and `/tf runs` — no more silent intact placeholders.
|
|
682
|
-
|
|
683
|
-
## Commands
|
|
684
|
-
|
|
685
|
-
Saved flows become CLI shortcuts. **These `/tf` commands are Pi-only** (they run in the Pi session). On Codex, Claude Code, OpenCode, and Grok Build, use the `taskflow_*` MCP tools instead — full set: `taskflow_run` / `list` / `show` / `verify` / `compile` / `peek` / `trace` / `replay` / `why_stale` / `recompute` (dry-run) / `save` / `search`.
|
|
686
|
-
|
|
687
|
-
| Command | What it does |
|
|
688
|
-
|---|---|
|
|
689
|
-
| `/tf list` | List all saved flows |
|
|
690
|
-
| `/tf run <name> [args]` | Run a saved flow (e.g. `/tf run summarize-files dir=src`) |
|
|
691
|
-
| `/tf show <name>` | Print a flow's definition |
|
|
692
|
-
| `/tf compile <name> [lr\|td]` | **Render the flow as a Mermaid diagram + verification overlay** — 0 tokens, no LLM; paste into a README/issue/PR |
|
|
693
|
-
| `/tf ir <name>` | Compile to **FlowIR** + content hash (`ir:<64-hex>`) — 0 tokens |
|
|
694
|
-
| `/tf runs` | Browse recent run history (interactive TUI — **live auto-refreshes** while any run is active) |
|
|
695
|
-
| `/tf resume <runId>` | Continue a paused/failed run — cached phases skip automatically |
|
|
696
|
-
| `/tf peek <runId> [phaseId]` | Inspect a phase's intermediate output (the debugging escape hatch) |
|
|
697
|
-
| `/tf provenance <runId>` | Show observed read-sets for a completed run |
|
|
698
|
-
| `/tf trace <runId> [--json]` | Show a run's **deterministic-replay event trace** (each subagent call + runtime decisions) |
|
|
699
|
-
| `/tf replay <runId> [--threshold phase=n] [--budget-usd n] [--json]` | **Offline what-if** re-judge of thresholds/budget from a recorded trace (zero tokens) |
|
|
700
|
-
| `/tf why-stale <runId> [phaseId]` | Explain the stale frontier (observed ∪ declared deps) |
|
|
701
|
-
| `/tf recompute <runId> <phaseId> [--apply]` | Dry-run (default) or apply minimal recompute of the stale frontier |
|
|
702
|
-
| `/tf init` | **Interactively map model roles** to your enabled models (writes `~/.pi/agent/settings.json`) |
|
|
703
|
-
| `/tf:<name> [args]` | Shortcut — runs the flow in one tap |
|
|
704
|
-
|
|
705
|
-
Tool actions (used by the model on Pi): `run` (inline `define` or saved `name`), `save`, `resume`, `list`, `agents`, `init`, `verify`, `compile`, `ir`, `provenance`, `trace`, `replay`, `why-stale`, `recompute`, `cache-clear`, `search`. On Codex, Claude Code, OpenCode, and Grok Build the exposed MCP tools are `taskflow_run` / `taskflow_list` / `taskflow_show` / `taskflow_verify` / `taskflow_compile` / `taskflow_peek` / `taskflow_trace` / `taskflow_replay` / `taskflow_why_stale` / `taskflow_recompute` (dry-run only) / `taskflow_save` / `taskflow_search`.
|
|
706
|
-
|
|
707
|
-
## Background (detached) execution
|
|
708
|
-
|
|
709
|
-
Pass `detach: true` to run a taskflow in a detached child process — the tool returns immediately with the `runId` and the flow continues running even if the host session exits:
|
|
710
|
-
|
|
711
|
-
```jsonc
|
|
712
|
-
{
|
|
713
|
-
"action": "run",
|
|
714
|
-
"name": "nightly-audit",
|
|
715
|
-
"detach": true
|
|
716
|
-
}
|
|
717
|
-
```
|
|
718
|
-
|
|
719
|
-
- The child process reads serialized context, calls the orchestration engine, and persists terminal state to the store.
|
|
720
|
-
- Status is polled via `/tf runs` (which now **auto-refreshes live** when any run is running) or `action: "resume"`.
|
|
721
|
-
- Stale PID detection via signal-0 probe; the idle watchdog kills stalled children.
|
|
722
|
-
- **Approval phases auto-reject** in detached mode — human gates are never silently bypassed.
|
|
723
|
-
- `resume` works normally after a detached run completes or fails.
|
|
724
|
-
|
|
725
|
-
## Resume across sessions
|
|
726
|
-
|
|
727
|
-
A taskflow run isn't tied to your session. Every completed phase is written to disk, so a run that fails (or that you stop) can be continued later with `/tf resume <runId>` — **cached phases skip automatically** and only the remaining work spends tokens.
|
|
728
|
-
|
|
729
|
-
<div align="center">
|
|
730
|
-
<img src="https://raw.githubusercontent.com/heggria/taskflow/main/assets/resume.png" alt="A run fails midway in session 1; in session 2 /tf resume skips the cached phases and only re-runs the failed phase and what follows" width="900">
|
|
731
|
-
</div>
|
|
732
|
-
|
|
733
|
-
Resume is keyed on each phase's input hash — if an upstream output changed, dependent phases re-run; if nothing changed, they're reused. No competing Pi extension does this across sessions.
|
|
734
|
-
|
|
735
|
-
## Storage
|
|
736
|
-
|
|
737
|
-
```
|
|
738
|
-
.pi/taskflows/<name>.json # project-scope definitions (commit to share)
|
|
739
|
-
~/.pi/agent/taskflows/<name>.json # user-scope definitions
|
|
740
|
-
.pi/taskflows/runs/<flowName>/<runId>.json # run state for resume (gitignore this)
|
|
741
|
-
.pi/taskflows/cache/ # cross-run memoization cache (gitignored)
|
|
288
|
+
```bash
|
|
289
|
+
pi install npm:pi-taskflow
|
|
742
290
|
```
|
|
743
291
|
|
|
744
|
-
|
|
745
|
-
|
|
746
|
-
Agent discovery scope (via `agentScope` in the flow definition):
|
|
292
|
+
Pi provides the richest local experience: the `taskflow` tool, `/tf` commands, live DAG rendering, interactive approvals, background runs, and model-role setup.
|
|
747
293
|
|
|
748
|
-
|
|
749
|
-
|---|---|
|
|
750
|
-
| `"user"` (default) | `~/.pi/agent/agents/*.md` |
|
|
751
|
-
| `"project"` | `.pi/agents/*.md` (walks up the tree) |
|
|
752
|
-
| `"both"` | user + project; project wins on name collision |
|
|
753
|
-
|
|
754
|
-
Run cleanup is configurable via `maxKeptRuns` and `maxRunAgeDays` in settings.
|
|
755
|
-
|
|
756
|
-
## Agents
|
|
757
|
-
|
|
758
|
-
Taskflow ships **18 built-in agents** — each a `.md` file with a tuned system prompt, thinking level, and tool set. You can reference them by `name` in any phase or shorthand, right after install. No setup required.
|
|
759
|
-
|
|
760
|
-
### Built-in agent roster
|
|
761
|
-
|
|
762
|
-
| Agent | Role | Thinking | Default role |
|
|
763
|
-
|---|---|---:|---|
|
|
764
|
-
| `executor` | Implement planned code changes | high | `{{fast}}` |
|
|
765
|
-
| `executor-fast` | Trivial fixes (≤2 files, ≤50 lines) | off | `{{fast}}` |
|
|
766
|
-
| `executor-code` | Complex multi-file implementation | high | `{{strong}}` |
|
|
767
|
-
| `executor-ui` | Frontend / styling / visual changes | high | `{{vision}}` |
|
|
768
|
-
| `scout` | Fast codebase recon & file mapping | off | `{{fast}}` |
|
|
769
|
-
| `planner` | Implementation plan creation | high | `{{strong}}` |
|
|
770
|
-
| `analyst` | Requirements analysis, ambiguity detection | high | `{{thinker}}` |
|
|
771
|
-
| `critic` | Inline self-doubt during reasoning | xhigh | `{{thinker}}` |
|
|
772
|
-
| `reviewer` | General code / architecture review | high | `{{strong}}` |
|
|
773
|
-
| `risk-reviewer` | Backend / infra / DB / API risk | high | `{{reasoner}}` |
|
|
774
|
-
| `security-reviewer` | Security vulns, auth/crypto | xhigh | `{{reasoner}}` |
|
|
775
|
-
| `plan-arbiter` | Plan quality gate (complex tasks) | high | `{{arbiter}}` |
|
|
776
|
-
| `final-arbiter` | Tiebreaker when critics disagree | xhigh | `{{arbiter}}` |
|
|
777
|
-
| `test-engineer` | Design & implement tests | high | `{{fast}}` |
|
|
778
|
-
| `doc-writer` | Documentation authoring | off | `{{fast}}` |
|
|
779
|
-
| `recover` | Session recovery after compaction | low | `{{fast}}` |
|
|
780
|
-
| `verifier` | Run tests, validate outcomes | off | `{{fast}}` |
|
|
781
|
-
| `visual-explorer` | Figma design metadata analysis | high | `{{vision}}` |
|
|
782
|
-
|
|
783
|
-
Agents are layered: **built-in → user (`~/.pi/agent/agents/`) → project (`.pi/agents/`)**. A user or project agent with the same `name` overrides the built-in — so you can customize any agent without touching the package.
|
|
784
|
-
|
|
785
|
-
### Model roles
|
|
786
|
-
|
|
787
|
-
Each built-in agent's `model` field uses a **role placeholder** (e.g. `{{fast}}`) instead of a hardcoded provider string. This decouples *intent* from *implementation* — you map roles to models once, and every agent adapts.
|
|
788
|
-
|
|
789
|
-
| Role | Intent | Typical model |
|
|
790
|
-
|---|---|---|
|
|
791
|
-
| `{{fast}}` | Cheap & quick — high-volume, low-stakes | DeepSeek V4 Flash |
|
|
792
|
-
| `{{strong}}` | Balanced — planning, review, moderate complexity | MiMo v2.5 Pro |
|
|
793
|
-
| `{{thinker}}` | Deep analysis — requirements, critique | DeepSeek V4 Pro |
|
|
794
|
-
| `{{arbiter}}` | Final judgment — tiebreak, plan quality gates | Qwen 3.7 Max |
|
|
795
|
-
| `{{vision}}` | Multimodal — UI work, design reading | MiniMax M3 |
|
|
796
|
-
| `{{reasoner}}` | Cautious reasoning — security, risk | GLM 5.1 |
|
|
294
|
+
[Pi guide →](https://heggria.github.io/taskflow/en/docs/guides/pi)
|
|
797
295
|
|
|
798
|
-
|
|
296
|
+
### OpenAI Codex
|
|
799
297
|
|
|
800
298
|
```bash
|
|
801
|
-
/
|
|
299
|
+
codex plugin marketplace add heggria/taskflow
|
|
300
|
+
codex plugin add taskflow@taskflow
|
|
802
301
|
```
|
|
803
302
|
|
|
804
|
-
|
|
805
|
-
|
|
806
|
-
```
|
|
807
|
-
? What do you want to do with model roles?
|
|
808
|
-
❯ Use recommended defaults
|
|
809
|
-
Configure each role
|
|
810
|
-
Edit one role
|
|
811
|
-
Show current roles
|
|
812
|
-
Cancel
|
|
813
|
-
```
|
|
303
|
+
[Codex guide →](https://heggria.github.io/taskflow/en/docs/guides/codex)
|
|
814
304
|
|
|
815
|
-
|
|
305
|
+
### Claude Code
|
|
816
306
|
|
|
817
|
-
```
|
|
818
|
-
|
|
819
|
-
|
|
820
|
-
Recommended: minimax/MiniMax-M3
|
|
821
|
-
───────────────
|
|
822
|
-
❯ MiniMax M3 (minimax/MiniMax-M3) · image ✓ · reasoning ✓ · (recommended)
|
|
823
|
-
Claude Sonnet 4.6 (openrouter/anthropic/...) · image ✓ · reasoning ✓ · (current)
|
|
824
|
-
GPT-5 (openrouter/openai/gpt-5) · image ✓
|
|
825
|
-
DeepSeek V4 Flash (openrouter/deepseek/v4-flash)
|
|
826
|
-
───────────────
|
|
827
|
-
Custom (type your own)
|
|
828
|
-
Keep current
|
|
829
|
-
Back to action menu
|
|
307
|
+
```bash
|
|
308
|
+
claude plugin marketplace add heggria/taskflow
|
|
309
|
+
claude plugin install claude-taskflow@taskflow
|
|
830
310
|
```
|
|
831
311
|
|
|
832
|
-
|
|
833
|
-
|
|
834
|
-
```
|
|
835
|
-
? Review changes:
|
|
836
|
-
fast openrouter/deepseek/deepseek-v4-flash (unchanged)
|
|
837
|
-
strong openrouter/xiaomi/mimo-v2.5-pro (unchanged)
|
|
838
|
-
thinker openrouter/qwen/qwen3.7-max (changed ← was: openrouter/deepseek/v4-pro)
|
|
839
|
-
arbiter openrouter/qwen/qwen3.7-max (unchanged)
|
|
840
|
-
vision minimax/MiniMax-M3 (unchanged)
|
|
841
|
-
reasoner z-ai/glm-5.1 (unchanged)
|
|
842
|
-
───────────────
|
|
843
|
-
❯ Save these changes
|
|
844
|
-
Edit a role
|
|
845
|
-
Cancel
|
|
846
|
-
```
|
|
312
|
+
[Claude Code guide →](https://heggria.github.io/taskflow/en/docs/guides/claude-code)
|
|
847
313
|
|
|
848
|
-
|
|
314
|
+
### OpenCode
|
|
849
315
|
|
|
850
|
-
```
|
|
851
|
-
|
|
852
|
-
|
|
853
|
-
"fast": "openrouter/deepseek/deepseek-v4-flash",
|
|
854
|
-
"strong": "openrouter/xiaomi/mimo-v2.5-pro",
|
|
855
|
-
"thinker": "openrouter/deepseek/deepseek-v4-pro",
|
|
856
|
-
"arbiter": "openrouter/qwen/qwen3.7-max",
|
|
857
|
-
"vision": "minimax/MiniMax-M3",
|
|
858
|
-
"reasoner": "z-ai/glm-5.1"
|
|
859
|
-
}
|
|
860
|
-
}
|
|
316
|
+
```bash
|
|
317
|
+
opencode mcp add taskflow -- \
|
|
318
|
+
npx -y -p opencode-taskflow@0.2.2 opencode-taskflow-mcp
|
|
861
319
|
```
|
|
862
320
|
|
|
863
|
-
|
|
864
|
-
|
|
865
|
-
To customize a specific agent's model or thinking without changing `modelRoles`, create an agent file at `~/.pi/agent/agents/<name>.md` with the desired overrides in the YAML frontmatter.
|
|
866
|
-
|
|
867
|
-
### Tool path (`action="init"`)
|
|
321
|
+
[OpenCode guide →](https://heggria.github.io/taskflow/en/docs/guides/opencode)
|
|
868
322
|
|
|
869
|
-
|
|
870
|
-
|
|
871
|
-
| Mode | Behavior |
|
|
872
|
-
|---|---|
|
|
873
|
-
| `mode: "show"` (default) | Read-only report of current `modelRoles`. Never overwrites. |
|
|
874
|
-
| `mode: "apply-defaults"` + `force: true` | Writes `RECOMMENDED_DEFAULTS` to `settings.json`, preserving stale keys. |
|
|
875
|
-
| `mode: "interactive"` | Launches the full action menu + picker flow (requires a UI session). |
|
|
876
|
-
|
|
877
|
-
|
|
878
|
-
### Custom agents
|
|
879
|
-
|
|
880
|
-
Drop a `.md` file into `~/.pi/agent/agents/` (user-level) or `.pi/agents/` (project-level, commit it) to add your own:
|
|
881
|
-
|
|
882
|
-
```markdown
|
|
883
|
-
---
|
|
884
|
-
name: my-linter
|
|
323
|
+
### Grok Build
|
|
885
324
|
|
|
886
|
-
|
|
887
|
-
|
|
888
|
-
|
|
889
|
-
|
|
890
|
-
model: "{{fast}}"
|
|
891
|
-
|
|
892
|
-
thinking: off
|
|
893
|
-
---
|
|
894
|
-
|
|
895
|
-
You are a linting agent. Run `npx eslint --format json` on the
|
|
896
|
-
provided files. Report violations grouped by file. No fixes.
|
|
325
|
+
```bash
|
|
326
|
+
grok mcp add taskflow -- \
|
|
327
|
+
npx -y -p grok-taskflow@0.2.2 grok-taskflow-mcp
|
|
897
328
|
```
|
|
898
329
|
|
|
899
|
-
|
|
900
|
-
|
|
901
|
-
## Examples
|
|
330
|
+
Grok Build support is new in 0.2. Its CLI stream does not report token/cost usage, so budget-declaring flows are rejected rather than silently running without enforcement.
|
|
902
331
|
|
|
903
|
-
|
|
332
|
+
[Grok Build guide →](https://heggria.github.io/taskflow/en/docs/guides/grok-build)
|
|
904
333
|
|
|
905
|
-
|
|
906
|
-
|---|---|
|
|
907
|
-
| [`summarize-files.json`](https://github.com/heggria/taskflow/blob/main/examples/summarize-files.json) | discover → `map` fan-out → `reduce` |
|
|
908
|
-
| [`conditional-research.json`](https://github.com/heggria/taskflow/blob/main/examples/conditional-research.json) | `when` routing + `join: any` + `gate` + `budget` |
|
|
909
|
-
| [`guarded-refactor.json`](https://github.com/heggria/taskflow/blob/main/examples/guarded-refactor.json) | `approval` (human-in-the-loop) + `retry` + `gate` |
|
|
910
|
-
| [`dynamic-plan-execute.json`](https://github.com/heggria/taskflow/blob/main/examples/dynamic-plan-execute.json) | `flow { def }` — plan then execute at runtime |
|
|
911
|
-
| [`iterative-replan.json`](https://github.com/heggria/taskflow/blob/main/examples/iterative-replan.json) | `loop` + `flow { def }` — iterative replanning |
|
|
912
|
-
|
|
913
|
-
Copy one into `.pi/taskflows/<name>.json` (or `~/.pi/agent/taskflows/`) and it registers as `/tf:<name>` — or just point the model at it.
|
|
914
|
-
|
|
915
|
-
## What's inside
|
|
334
|
+
## Built to survive real work
|
|
916
335
|
|
|
917
336
|
<div align="center">
|
|
918
337
|
|
|
919
|
-
**
|
|
338
|
+
**9 packages** · **5 hosts** · **12 phase types** · **18 built-in agents** · **1,500+ tests** · **MIT**
|
|
920
339
|
|
|
921
340
|
</div>
|
|
922
341
|
|
|
923
|
-
|
|
924
|
-
-
|
|
925
|
-
|
|
926
|
-
|
|
927
|
-
|
|
928
|
-
|
|
929
|
-
|
|
930
|
-
|
|
931
|
-
|
|
932
|
-
|
|
933
|
-
|
|
934
|
-
| Campaign | Scale | Phases | Outcome |
|
|
935
|
-
|----------|-------|--------|---------|
|
|
936
|
-
| [v0.0.8 dogfood](https://github.com/heggria/taskflow/blob/main/docs/internal/dogfooding-v0.0.8-report.md) | Full codebase audit → triage → fix → verify | 10 phases, 234 tests | 13 fixes, all pass |
|
|
937
|
-
| [v0.0.6 self-audit](https://github.com/heggria/taskflow/blob/main/docs/internal/self-audit-report.md) | inventory → map audit → gate → approval → map fix → reduce | 9 phases | 11 critical defects fixed |
|
|
938
|
-
| [Cross-run cache dogfood](https://github.com/heggria/taskflow/blob/main/docs/internal/rfc-cross-run-memoization.md) | Real runtime + on-disk store | Dedicated test harness | Cache correctness under adversarial fingerprints |
|
|
939
|
-
| [Adversarial cross-review](https://github.com/heggria/taskflow/blob/main/docs/internal/brainstorm-adversarial-review-report.md) | Multi-agent adversarial review | `tournament` + `gate` | P0 cache-key fix shipped |
|
|
940
|
-
| [Init redesign review](https://github.com/heggria/taskflow/blob/main/docs/internal/issue-necessity-review-report.md) | Necessity audit → parallel checks → verdict | 7 phases | Full redesign plan validated |
|
|
941
|
-
| [Round 2 adversarial audit](https://github.com/heggria/taskflow/blob/main/docs/internal/dogfooding-report.md) | Integration layer + cross-module — 12 findings across runner/runtime/interpolate/verify | 14 phases | 10 fixes applied, 0 regressions |
|
|
942
|
-
| [Round 3 adversarial audit](https://github.com/heggria/taskflow/blob/main/docs/internal/dogfooding-report.md) | Integration layer + cross-module — 10 findings across index/agents/cache/render/runs-view | 9 phases | 10 fixes applied, 0 regressions |
|
|
943
|
-
| [v0.0.23 Shared Context Tree](https://github.com/heggria/taskflow/blob/main/docs/internal/dogfooding-report.md) | End-to-end validation: org-tree spawn, 5-way audit via loop+gate | 6 e2e runs | Spawn-drain bug fixed, 50 new tests |
|
|
944
|
-
|
|
945
|
-
> **Meta:** we used `taskflow`'s `map` fan-out, `gate` verdicts, `approval` human-in-the-loop, `tournament` best-of-N, `loop` until-done, and `cross-run` cache — to build `taskflow`.
|
|
946
|
-
|
|
947
|
-
## Status & limits
|
|
948
|
-
|
|
949
|
-
**Compatibility baseline from v0.1.8:** interpolation placeholders in phase
|
|
950
|
-
`cwd` are rejected; the release dependency/security sweep is also retained.
|
|
342
|
+
```text
|
|
343
|
+
taskflow-core
|
|
344
|
+
┌──────────────┼───────────────┐
|
|
345
|
+
│ │ │
|
|
346
|
+
taskflow-dsl pi-taskflow taskflow-mcp-core ─┐
|
|
347
|
+
taskflow-hosts ─────┼─ codex-taskflow
|
|
348
|
+
├─ claude-taskflow
|
|
349
|
+
├─ opencode-taskflow
|
|
350
|
+
└─ grok-taskflow
|
|
351
|
+
```
|
|
951
352
|
|
|
952
|
-
|
|
353
|
+
`taskflow-core` is host-neutral and imports no host SDK. `taskflow-mcp-core` implements stdio JSON-RPC without an MCP SDK dependency; `taskflow-hosts` owns the shared host process runners. The four MCP delivery packages bind both layers (and core), while Pi keeps its native adapter.
|
|
953
354
|
|
|
954
|
-
|
|
355
|
+
The test suite covers orchestration semantics, persistence and file-lock races, cache freshness, path traversal, dynamic graph hardening, cancellation, budgets, all 12 phase kinds, FlowIR/replay/recompute, TypeScript DSL erasure, host argv contracts, MCP servers, and packed consumer imports.
|
|
955
356
|
|
|
956
|
-
|
|
957
|
-
- **Workspace isolation is fail-open.** `cwd: "worktree"` requires the base cwd to be a git work tree; otherwise it degrades to a `temp` dir (with a warning). `temp`/`worktree` dirs are removed when the phase ends — a hard crash mid-phase may leave a stray dir (cleaned on the next run for `dedicated`; `temp`/`worktree` are under the OS tmpdir). The reserved keywords are honoured only in author-written flows.
|
|
958
|
-
- **No `output: "file"`.** Outputs are text/JSON only — write files via an agent's `write` tool call.
|
|
959
|
-
- **`map` fans out over a JSON array from a string `over`.** The `over` field is a string that either interpolates to a JSON array (e.g. `{steps.ID.json}`) or is a literal JSON-array string. Wrap a plain text list in a single-agent `output: "json"` phase first, or pass `JSON.stringify([...])` for a fixed list. (A raw literal array is rejected — emit it from a phase and reference that.)
|
|
960
|
-
- **The DAG must be acyclic.** Cycles are rejected at validation.
|
|
961
|
-
- **Cross-run cache excludes `gate`, `approval`, `loop`, `tournament`, `script`, `race`, and `expand`.** These must produce a fresh result each run.
|
|
962
|
-
- **Approval auto-rejects in detached mode.** This is a safety invariant — approval gates are never silently bypassed.
|
|
357
|
+
## Documentation
|
|
963
358
|
|
|
964
|
-
|
|
359
|
+
| Start here | When you need |
|
|
360
|
+
|---|---|
|
|
361
|
+
| [Getting Started](https://heggria.github.io/taskflow/en/docs/getting-started) | Your first successful run |
|
|
362
|
+
| [Concepts](https://heggria.github.io/taskflow/en/docs/concepts/) | DAGs, isolation, verification, resume, shared context |
|
|
363
|
+
| [Syntax](https://heggria.github.io/taskflow/en/docs/syntax/) | Phase fields, control flow, budgets, caching, scorers |
|
|
364
|
+
| [Compiler & Runtime](https://heggria.github.io/taskflow/en/docs/compiler-runtime/) | TypeScript DSL, FlowIR, replay, recompute, background runs |
|
|
365
|
+
| [Host Guides](https://heggria.github.io/taskflow/en/docs/guides/) | Pi, Codex, Claude Code, OpenCode, and Grok setup |
|
|
366
|
+
| [Reference](https://heggria.github.io/taskflow/en/docs/reference/) | Commands, shorthand, and exact tool surfaces |
|
|
367
|
+
| [Showcase](https://heggria.github.io/taskflow/en/docs/showcase/) | Real flows and case studies |
|
|
965
368
|
|
|
966
|
-
`taskflow
|
|
369
|
+
Also see [`examples/`](https://github.com/heggria/taskflow/blob/main/examples), the [changelog](https://github.com/heggria/taskflow/blob/main/CHANGELOG.md), and the [release guide](https://github.com/heggria/taskflow/blob/main/RELEASE.md).
|
|
967
370
|
|
|
968
|
-
|
|
969
|
-
|---------|------|
|
|
970
|
-
| [`taskflow-core`](https://github.com/heggria/taskflow/blob/main/packages/taskflow-core) | Host-neutral orchestration engine (zero host-SDK deps; only `typebox`) — runtime, DSL, cache, verify |
|
|
971
|
-
| [`taskflow-mcp-core`](https://github.com/heggria/taskflow/blob/main/packages/taskflow-mcp-core) | Host-neutral MCP server (stdio JSON-RPC + `taskflow_*` tools + DAG renderer); depends on core |
|
|
972
|
-
| [`taskflow-hosts`](https://github.com/heggria/taskflow/blob/main/packages/taskflow-hosts) | Shared host-runner collection — the codex/claude/opencode/grok `SubagentRunner` impls + their argv builders + event-stream parsers; depends on core |
|
|
973
|
-
| [`taskflow-dsl`](https://github.com/heggria/taskflow/blob/main/packages/taskflow-dsl) | TypeScript DSL CLI/package — erases `.tf.ts` to Taskflow JSON and optional FlowIR; depends on core |
|
|
974
|
-
| [`pi-taskflow`](https://github.com/heggria/taskflow/blob/main/packages/pi-taskflow) | Pi extension adapter — `taskflow` tool + `/tf` commands (what `pi install npm:pi-taskflow` gives you) |
|
|
975
|
-
| [`codex-taskflow`](https://github.com/heggria/taskflow/blob/main/packages/codex-taskflow) | Codex MCP server + bin + [Codex plugin](https://github.com/heggria/taskflow/blob/main/packages/codex-taskflow/plugin) (re-exports the runner from `taskflow-hosts`) ([guide](https://github.com/heggria/taskflow/blob/main/docs/codex-mcp.md)) |
|
|
976
|
-
| [`claude-taskflow`](https://github.com/heggria/taskflow/blob/main/packages/claude-taskflow) | Claude Code MCP server + bin + [Claude Code plugin](https://github.com/heggria/taskflow/blob/main/packages/claude-taskflow/plugin) (re-exports the runner from `taskflow-hosts`) ([guide](https://github.com/heggria/taskflow/blob/main/docs/claude-mcp.md)) |
|
|
977
|
-
| [`opencode-taskflow`](https://github.com/heggria/taskflow/blob/main/packages/opencode-taskflow) | OpenCode MCP server + bin + [OpenCode config scaffold](https://github.com/heggria/taskflow/blob/main/packages/opencode-taskflow/plugin) (re-exports the runner from `taskflow-hosts`) ([guide](https://github.com/heggria/taskflow/blob/main/docs/opencode-mcp.md)) |
|
|
978
|
-
| [`grok-taskflow`](https://github.com/heggria/taskflow/blob/main/packages/grok-taskflow) | Grok Build MCP server + bin + [Grok plugin](https://github.com/heggria/taskflow/blob/main/packages/grok-taskflow/plugin) (re-exports the runner from `taskflow-hosts`) ([guide](https://github.com/heggria/taskflow/blob/main/docs/grok-mcp.md)) |
|
|
371
|
+
## Contributing
|
|
979
372
|
|
|
980
373
|
```bash
|
|
981
374
|
pnpm install
|
|
982
|
-
pnpm run typecheck
|
|
983
|
-
pnpm test
|
|
984
|
-
pnpm run
|
|
985
|
-
pnpm run
|
|
986
|
-
pnpm run test:e2e-codex # codex executor e2e (needs `codex` + model access)
|
|
987
|
-
pnpm run test:e2e-codex-mcp # codex MCP server e2e
|
|
988
|
-
pnpm run test:e2e-grok-mcp # grok MCP server e2e (no live model required)
|
|
375
|
+
pnpm run typecheck
|
|
376
|
+
pnpm test
|
|
377
|
+
pnpm run build
|
|
378
|
+
pnpm run test:pack
|
|
989
379
|
```
|
|
990
380
|
|
|
991
|
-
|
|
992
|
-
the `.mts` extension so the unit-test glob skips them), e.g.:
|
|
381
|
+
Contributions are welcome. Start with [`CONTRIBUTING.md`](https://github.com/heggria/taskflow/blob/main/CONTRIBUTING.md) for the workflow and [`AGENTS.md`](https://github.com/heggria/taskflow/blob/main/AGENTS.md) for architecture and coding conventions.
|
|
993
382
|
|
|
994
|
-
|
|
995
|
-
node --conditions=development --experimental-strip-types packages/pi-taskflow/test/e2e.mts
|
|
996
|
-
# others: e2e-team, e2e-context, e2e-context-value, e2e-spawn-subflow,
|
|
997
|
-
# e2e-flowir, e2e-incremental-suite, dogfood-cache
|
|
998
|
-
```
|
|
383
|
+
## License
|
|
999
384
|
|
|
1000
|
-
|
|
385
|
+
[MIT](https://github.com/heggria/taskflow/blob/main/LICENSE) © [heggria](https://github.com/heggria)
|
|
1001
386
|
|
|
1002
|
-
|
|
387
|
+
<div align="center">
|
|
1003
388
|
|
|
1004
|
-
|
|
389
|
+
**Declare once. Verify first. Recompute only what changed.**
|
|
1005
390
|
|
|
1006
|
-
|
|
391
|
+
[Read the docs](https://heggria.github.io/taskflow/en/docs) · [Try an example](https://github.com/heggria/taskflow/blob/main/examples) · [View releases](https://github.com/heggria/taskflow/releases)
|
|
1007
392
|
|
|
1008
|
-
|
|
393
|
+
</div>
|