taskflow-mcp-core 0.1.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +951 -0
- package/dist/mcp/jsonrpc.d.ts +54 -0
- package/dist/mcp/jsonrpc.d.ts.map +1 -0
- package/dist/mcp/jsonrpc.js +118 -0
- package/dist/mcp/jsonrpc.js.map +1 -0
- package/dist/mcp/server.d.ts +38 -0
- package/dist/mcp/server.d.ts.map +1 -0
- package/dist/mcp/server.js +538 -0
- package/dist/mcp/server.js.map +1 -0
- package/dist/mcp/svg.d.ts +52 -0
- package/dist/mcp/svg.d.ts.map +1 -0
- package/dist/mcp/svg.js +312 -0
- package/dist/mcp/svg.js.map +1 -0
- package/package.json +64 -0
package/README.md
ADDED
|
@@ -0,0 +1,951 @@
|
|
|
1
|
+
<div align="center">
|
|
2
|
+
|
|
3
|
+
<img src="https://raw.githubusercontent.com/heggria/taskflow/main/assets/hero.png" alt="taskflow — a declarative, verifiable graph of task nodes for coding-agent subagents: stateful, resumable, context-isolated" width="900">
|
|
4
|
+
|
|
5
|
+
<p>
|
|
6
|
+
<a href="https://www.npmjs.com/package/pi-taskflow"><img src="https://img.shields.io/npm/v/pi-taskflow?style=flat-square&color=4B4ACF&label=npm" alt="npm version"></a>
|
|
7
|
+
<a href="https://www.npmjs.com/package/pi-taskflow"><img src="https://img.shields.io/npm/dm/pi-taskflow?style=flat-square&color=5A5D63&label=downloads" alt="npm downloads"></a>
|
|
8
|
+
<a href="https://github.com/heggria/taskflow/blob/main/LICENSE"><img src="https://img.shields.io/badge/license-MIT-0E8A66?style=flat-square" alt="MIT license"></a>
|
|
9
|
+
<a href="#whats-inside"><img src="https://img.shields.io/badge/runtime%20deps-0-0E8A66?style=flat-square" alt="zero runtime dependencies"></a>
|
|
10
|
+
<a href="https://github.com/heggria/taskflow/actions/workflows/ci.yml"><img src="https://img.shields.io/github/actions/workflow/status/heggria/taskflow/ci.yml?branch=main&style=flat-square&label=CI" alt="CI status"></a>
|
|
11
|
+
<a href="#whats-inside"><img src="https://img.shields.io/badge/tests-1140-4B4ACF?style=flat-square" alt="1140 tests"></a>
|
|
12
|
+
<a href="#whats-inside"><img src="https://img.shields.io/badge/dogfooded-%E2%9C%93-0E8A66?style=flat-square" alt="dogfooded"></a>
|
|
13
|
+
<a href="#run-it-on-your-agent"><img src="https://img.shields.io/badge/runs%20on-Pi%20%2B%20Codex%20%2B%20Claude%20Code%20%2B%20OpenCode-4B4ACF?style=flat-square" alt="runs on Pi, Codex, Claude Code, and OpenCode"></a>
|
|
14
|
+
</p>
|
|
15
|
+
|
|
16
|
+
<p align="center">
|
|
17
|
+
<b>English</b> ·
|
|
18
|
+
<a href="https://github.com/heggria/taskflow/blob/main/README.zh-CN.md">简体中文</a>
|
|
19
|
+
</p>
|
|
20
|
+
|
|
21
|
+
<p align="center">
|
|
22
|
+
<a href="https://heggria.github.io/taskflow/en"><img src="https://img.shields.io/badge/📖_Read_the_docs-heggria.github.io%2Ftaskflow-4B4ACF?style=for-the-badge&labelColor=2D2F5A" alt="Read the docs — heggria.github.io/taskflow"></a>
|
|
23
|
+
</p>
|
|
24
|
+
|
|
25
|
+
<p><strong>A declarative, verifiable <em>graph of tasks</em> for coding-agent subagents.</strong><br/>
|
|
26
|
+
Not a workflow you script — a DAG you declare. Fan out · gate · loop · tournament · resume · save as a command — intermediate results stay out of your context.<br/>
|
|
27
|
+
Runs on the <a href="https://pi.dev">Pi</a> coding agent, on <a href="https://github.com/openai/codex">OpenAI Codex</a>, on <a href="https://claude.com/product/claude-code">Claude Code</a>, and on <a href="https://opencode.ai">OpenCode</a>.</p>
|
|
28
|
+
|
|
29
|
+
</div>
|
|
30
|
+
|
|
31
|
+
```bash
|
|
32
|
+
# Pi
|
|
33
|
+
pi install npm:pi-taskflow
|
|
34
|
+
|
|
35
|
+
# Codex
|
|
36
|
+
codex plugin marketplace add heggria/taskflow
|
|
37
|
+
codex plugin add taskflow@taskflow
|
|
38
|
+
|
|
39
|
+
# Claude Code
|
|
40
|
+
claude plugin marketplace add heggria/taskflow
|
|
41
|
+
claude plugin install claude-taskflow@taskflow
|
|
42
|
+
|
|
43
|
+
# OpenCode — add the MCP server to opencode.json (see the OpenCode guide)
|
|
44
|
+
opencode mcp add taskflow -- npx -y -p opencode-taskflow opencode-taskflow-mcp
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
---
|
|
48
|
+
|
|
49
|
+
**A `workflow` flows. A `taskflow` is a *graph*.** Other orchestrators let the model *script* the work — imperative code that flows step by step, with the graph hidden inside control flow. `taskflow` does the opposite: you **declare** the work as a graph of discrete, named **task** nodes connected by `dependsOn` edges — and the runtime *verifies that graph before it spends a single token.*
|
|
50
|
+
|
|
51
|
+
You already know your agent's built-in subagent shorthand — `task` / `tasks` / `chain`. `taskflow` speaks the *same* shorthand — so your existing delegations instantly become **tracked, resumable, and saveable by name** (on Pi, a saved flow becomes a one-word `/tf:<name>` command; on Codex, Claude Code, and OpenCode you run it by name through `taskflow_run`). When you outgrow the shorthand, the full DSL gives you a real DAG: dynamic fan-out over dozens of items, conditional routing, quality gates, human approvals, retries, loops, tournaments, and a hard spend ceiling.
|
|
52
|
+
|
|
53
|
+
And the whole time, **only the final phase reaches your conversation.** Every intermediate transcript stays in the runtime, never your context window.
|
|
54
|
+
|
|
55
|
+
## Why "taskflow" and not "workflow"?
|
|
56
|
+
|
|
57
|
+
The name is the thesis. In engineering, a **task** is a *discrete, declared unit of work* — the node of a task graph (the same `task` a build system, scheduler, or compiler wires into a DAG). **Work**, by contrast, is *fluid and unbounded* — the continuous, imperative act of doing.
|
|
58
|
+
|
|
59
|
+
That distinction is exactly the design split playing out across coding-agent ecosystems:
|
|
60
|
+
|
|
61
|
+
<div align="center">
|
|
62
|
+
<img src="https://raw.githubusercontent.com/heggria/taskflow/main/assets/task-vs-work.png" alt="work is a fluid imperative script whose graph hides in control flow and can't be verified before it runs; a taskflow is a declarative graph of discrete task nodes that is statically verified before any token is spent" width="900">
|
|
63
|
+
</div>
|
|
64
|
+
|
|
65
|
+
- A **`workflow`** (the dynamic, code-mode kind) is the model writing an **imperative script** that *flows*: `await agent(...)`, an `if`, a `for`, another `await`. Expressive — it's Turing-complete — but the graph only exists *as the code runs*. You can't see it, diff it, or prove it terminates before you pay for it.
|
|
66
|
+
- A **`taskflow`** moves the plan **out of code and into a declarative graph of `task` nodes.** Because the graph is *data*, the runtime can do what an imperative script structurally cannot: **statically verify it** (no cycles, no dead ends, no budget overflow, no dangling refs) before a single subagent spawns, **render it** (the live progress *is* the DAG), **resume it** phase-by-phase, and **save it** as a one-word command.
|
|
67
|
+
|
|
68
|
+
> **The trade we make on purpose:** we give up the raw expressivity of arbitrary code to gain something an imperative script can't have — a graph that is **verifiable, observable, replayable, and safe to generate with an LLM.** When a job needs twelve steps with branching fan-out and a review gate, you want a graph you can *check* — not a script you *hope* runs right.
|
|
69
|
+
|
|
70
|
+
## Why this exists
|
|
71
|
+
|
|
72
|
+
Here's the wall you hit with raw subagents: you describe a multi-step plan in prose, the model re-derives it every single run, the intermediate transcripts flood your context, and the moment one model call fails you start over from zero. There's no reuse, no recovery, no structure — and no way to *check* the plan before it burns tokens.
|
|
73
|
+
|
|
74
|
+
`taskflow` moves the plan **out of the prompt and into a declarative graph of task nodes.** The runtime owns the DAG, the loops, the retries, and the intermediate state. You declare a pipeline once and run it a hundred times — by name. Because the plan is data, not prose and not code, it can be **validated, visualized, and replayed.**
|
|
75
|
+
|
|
76
|
+
<div align="center">
|
|
77
|
+
<img src="https://raw.githubusercontent.com/heggria/taskflow/main/assets/context-isolation.png" alt="With raw subagents every transcript floods your context; with taskflow transcripts stay in the runtime and only the final result returns" width="900">
|
|
78
|
+
</div>
|
|
79
|
+
|
|
80
|
+
> Twelve steps, branching fan-out, a review gate, a spend cap — that's a graph, and you want to *see and check* it, not re-prompt it every run.
|
|
81
|
+
|
|
82
|
+
| | subagent (built-in) | **taskflow** |
|
|
83
|
+
|---|---|---|
|
|
84
|
+
| **Who drives** | the model, turn by turn | the runtime, from a definition |
|
|
85
|
+
| **Topology** | chain / flat parallel | **DAG with layered concurrency + routing** |
|
|
86
|
+
| **Intermediate results** | in your context window | **in the runtime — not your context** |
|
|
87
|
+
| **Scale** | a handful of tasks | **dynamic `map` fan-out over dozens of items** |
|
|
88
|
+
| **Reusable** | re-described every time | **saved by name (`/tf:<name>` on Pi; `taskflow_run` by name on Codex)** |
|
|
89
|
+
| **Resumable** | ✗ | **✓ cross-session — cached phases auto-skip** |
|
|
90
|
+
| **Quality gates** | ✗ | **`gate` phases that halt on `VERDICT: BLOCK`** |
|
|
91
|
+
| **Conditional routing** | ✗ | **`when` guards + `join: any` OR-joins** |
|
|
92
|
+
| **Fault tolerance** | ✗ | **per-phase `retry` + auto-retry on transient errors** |
|
|
93
|
+
| **Human-in-the-loop** | ✗ | **`approval` phases (approve / reject / edit)** |
|
|
94
|
+
| **Cost control** | ✗ | **run-wide `budget` (USD / token caps)** |
|
|
95
|
+
| **Composition** | ✗ | **`flow` phases run saved *or runtime-generated* sub-flows** |
|
|
96
|
+
| **Iterative loops** | ✗ | **`loop` phases — repeat until condition, convergence, or cap** |
|
|
97
|
+
| **Competitive selection** | ✗ | **`tournament` phases — N variants + judge** |
|
|
98
|
+
| **Live progress** | opaque while running | **live DAG render with timing + cost (Pi `/tf`); one streaming tool call on Codex** |
|
|
99
|
+
| **Ergonomics** | inline JSON each time | **shorthand (`task`/`tasks`/`chain`) *or* DSL** |
|
|
100
|
+
|
|
101
|
+
It doesn't replace the subagent tool. It gives your subagents a **graph**, a memory, and a name.
|
|
102
|
+
|
|
103
|
+
## Declarative graph vs. imperative script
|
|
104
|
+
|
|
105
|
+
The closest thing to `taskflow` in spirit is the **dynamic / code-mode workflow** — where the model writes a JavaScript orchestration script. It's powerful and genuinely expressive. But it sits at the *opposite* end of one fundamental axis: **expressivity vs. verifiability.**
|
|
106
|
+
|
|
107
|
+
| | dynamic `workflow` (code-mode) | **`taskflow`** (declarative graph) |
|
|
108
|
+
|---|---|---|
|
|
109
|
+
| **The plan is** | imperative JS the model writes & runs | **declarative JSON data the runtime executes** |
|
|
110
|
+
| **The graph** | implicit — hidden in `if`/`for`/`await` control flow | **explicit — `phases[]` + `dependsOn` edges, a first-class object** |
|
|
111
|
+
| **Verify before running** | ✗ Turing-complete; can't prove it terminates | **✓ static checks: no cycles, dead-ends, budget overflow, dangling refs** |
|
|
112
|
+
| **See it** | ✗ the graph only exists as the code runs | **✓ the live progress render *is* the DAG** |
|
|
113
|
+
| **Resume** | coarse (call-cache dedup) | **✓ phase-by-phase input-hash resume, cross-session** |
|
|
114
|
+
| **Safe to LLM-generate** | risky — it's executable code | **✓ it's just data — no JavaScript `eval`; and a runtime-generated sub-flow is *structurally validated* (cycles / dangling refs / duplicate ids) before it runs** |
|
|
115
|
+
| **Expressivity ceiling** | **higher** — arbitrary control flow | bounded by the DSL, but `map`/`when`/`loop`/`gate`/`tournament` — plus **runtime-generated sub-flows (`flow {def}`)** for plan-then-execute and iterative replanning — cover most jobs |
|
|
116
|
+
|
|
117
|
+
We chose the **verifiable** side on purpose. The expressivity you give up is real; what you get back — a plan you can check, watch, replay, and safely let a model author — is what turns one-off prompting into durable orchestration.
|
|
118
|
+
|
|
119
|
+
## Compared to other Pi extensions
|
|
120
|
+
|
|
121
|
+
> This section is **Pi-specific** — it maps `pi-taskflow` against other packages in the Pi ecosystem. If you're on Codex, skip to [Phase types](#phase-types); the engine and DSL are identical.
|
|
122
|
+
|
|
123
|
+
The Pi ecosystem now has **20+ delegation, workflow, and orchestration extensions** — each great at what it's for. Here's an honest map of where `pi-taskflow` sits (verified against each package's latest npm release, June 2026). For the full breakdown — every package, strengths *and* weaknesses — see [`docs/internal/PI-ECOSYSTEM.md`](https://github.com/heggria/taskflow/blob/main/docs/internal/PI-ECOSYSTEM.md). For the broader, non-Pi landscape (LangGraph, Temporal, CrewAI, Mastra…) see [`docs/internal/COMPETITORS.md`](https://github.com/heggria/taskflow/blob/main/docs/internal/COMPETITORS.md).
|
|
124
|
+
|
|
125
|
+
| Extension | Model | Custom DSL | DAG | Dynamic fan-out | Cross-session resume | Quality gate | Human approval | Save as command | Zero deps |
|
|
126
|
+
|---|---|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|
|
|
127
|
+
| **taskflow** | **declarative multi-phase taskflows** | **✓** | **✓** | **✓ `map`** | **✓ phase-hash** | **✓** | **✓** | **✓ `/tf:<name>`** | **✓** |
|
|
128
|
+
| [`@pi-agents/orchid`](https://www.npmjs.com/package/@pi-agents/orchid) | opinionated 9-phase pipeline + Ralph loop | fixed | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✕ (2) |
|
|
129
|
+
| [`pi-crew`](https://www.npmjs.com/package/pi-crew) | role teams + git worktrees + async | partial | ✓ | ✓ | ✓ | ✓ | ✓ | – | ✕ (7) |
|
|
130
|
+
| [`ultimate-pi`](https://www.npmjs.com/package/ultimate-pi) | governed plan→execute→review harness | YAML contracts | ✓ (plan-time) | ✕ | ✓ | ✓ (3-tier) | ✓ | ✓ | ✕ (16) |
|
|
131
|
+
| [`@zhushanwen/pi-workflow`](https://www.npmjs.com/package/@zhushanwen/pi-workflow) | JS scripts (`agent`/`parallel`/`pipeline`) | yes (JS) | ✕ (linear) | ✓ | ✓ | ✕ | ✕ | ✓ (call cache) | ✓ |
|
|
132
|
+
| [`@fiale-plus/pi-rogue-orchestration`](https://www.npmjs.com/package/@fiale-plus/pi-rogue-orchestration) | timer loop + goal resolution | ✕ | ✕ | ✕ | ✓ | ✓ (goal-check) | ✕ | ✕ | ✓ |
|
|
133
|
+
| [`pi-subagents`](https://www.npmjs.com/package/pi-subagents) | single / parallel / chain delegation | ✕ | ✕ | static | – | ✕ | clarify | named workflows | ✕ (3) |
|
|
134
|
+
| [`@gotgenes/pi-subagents`](https://www.npmjs.com/package/@gotgenes/pi-subagents) | Claude-Code-style subagents + worktrees | ✕ | ✕ | ✕ | ✓ (by id) | ✕ | per-agent | ✕ | ✕ (1) |
|
|
135
|
+
| [`pi-pipeline`](https://www.npmjs.com/package/pi-pipeline) | fixed SPEC→PLAN→TASKS→VERIFY | ✕ | fixed | ✕ | session planning | ✓ | clarify | ✕ | ✕ (2) |
|
|
136
|
+
| [`pi-agent-flow`](https://www.npmjs.com/package/pi-agent-flow) | one-shot parallel specialist `fork` | yes | ✕ | ✕ | – | ✕ | ✕ | – | ✕ (2) |
|
|
137
|
+
|
|
138
|
+
*(Representative slice of the 20+ — see [`docs/internal/PI-ECOSYSTEM.md`](https://github.com/heggria/taskflow/blob/main/docs/internal/PI-ECOSYSTEM.md) for all of them, plus `@0xkobold/pi-orchestration`, `@melihmucuk/pi-crew`, `@mediadatafusion/pi-workflow-suite`, `gentle-pi`, `@dreki-gg/pi-subagent`, and more.)*
|
|
139
|
+
|
|
140
|
+
**How to choose:**
|
|
141
|
+
|
|
142
|
+
- **`@pi-agents/orchid`** is the most feature-complete orchestrator in the ecosystem (DAG + worktrees + Ralph loop + agent mailbox) — but its DSL is a *fixed* 9-phase pipeline, it carries runtime deps + jiti, and it's beta. Reach for `taskflow` when you want to **define your own graph** (not adopt an opinionated one) with **zero dependencies** and a one-command install.
|
|
143
|
+
- **`pi-crew` / `ultimate-pi`** go heavier — worktree isolation, durable async teams, multi-tier governance. If you want lightweight, declarative, and zero-dependency, that's this project.
|
|
144
|
+
- **`@zhushanwen/pi-workflow`** is the closest in spirit and also zero-dep, but it's the **imperative** side of the split above: you author workflows as **JavaScript scripts** the model writes and runs. `taskflow`'s **declarative JSON DAG** is the verifiable side — statically checkable, visualizable, safe to LLM-generate, and resumable at phase granularity rather than call-cache dedup.
|
|
145
|
+
- **`@fiale-plus/pi-rogue-orchestration`** has a real **loop-until-done** (goal-driven iteration). `taskflow` now ships its own `loop` phase (v0.0.13+) plus `tournament` for competitive selection — and unlike rogue-orchestration, `taskflow` has a full DAG with gates, compositional sub-flows, and cross-session resume. For raw "keep going until the goal is met" with minimal structure, rogue-orchestration is still lighter; for structured, branching pipelines, `taskflow` covers the same ground and more.
|
|
146
|
+
- **`pi-subagents` / `@gotgenes/pi-subagents`** are the mature picks for ad-hoc "use reviewer on this diff" delegation and background jobs. `taskflow` is for when those delegations need to become a *repeatable, resumable pipeline*.
|
|
147
|
+
- **`pi-pipeline` / `pi-agent-flow`** ship *opinionated, fixed* flows. `taskflow` ships an *empty canvas*: you (or the model) declare the graph that fits the job.
|
|
148
|
+
|
|
149
|
+
> The honest one-liner: **`pi-taskflow` is the only Pi extension that gives you a *declarative, verifiable, resumable* DAG of task nodes — saved as a one-word `/tf:<name>` command, with zero runtime dependencies and context isolation by design** (and the same engine runs on Codex via the `taskflow_*` MCP tools). Where code-mode workflows let the model *script* the work, `taskflow` lets it *declare a graph the runtime can prove correct before running.* Recently shipped from the roadmap: the Shared Context Tree (blackboard + supervision) and worktree isolation (see [`docs/internal/STRATEGY.md`](https://github.com/heggria/taskflow/blob/main/docs/internal/STRATEGY.md)).
|
|
150
|
+
|
|
151
|
+
## 30-second start
|
|
152
|
+
|
|
153
|
+
### On Pi
|
|
154
|
+
|
|
155
|
+
**1. Install** — one command:
|
|
156
|
+
|
|
157
|
+
```bash
|
|
158
|
+
pi install npm:pi-taskflow
|
|
159
|
+
```
|
|
160
|
+
|
|
161
|
+
> **Optional:** run `/tf init` once to map the 18 built-in agents' model roles
|
|
162
|
+
> (`fast`, `strong`, `thinker`, …) to your own models — an interactive picker.
|
|
163
|
+
> Skip it and agents just use Pi's default model. See [Model roles](#model-roles).
|
|
164
|
+
|
|
165
|
+
**2. Run** — just ask the model in a Pi session:
|
|
166
|
+
|
|
167
|
+
> *Run a chain: first explore the auth flow, then summarize the findings.*
|
|
168
|
+
|
|
169
|
+
The model calls the `taskflow` tool automatically. You get live progress, per-step timing, token cost, and a saved run record — **same effort as the built-in tool, now tracked and resumable.**
|
|
170
|
+
|
|
171
|
+
**3. Save** — say *"save it"* and you have `/tf:<name>` forever.
|
|
172
|
+
|
|
173
|
+
That's it. You can be running your first workflow before your coffee cools — without writing a single phase definition.
|
|
174
|
+
|
|
175
|
+
<a id="run-it-on-your-agent"></a>
|
|
176
|
+
### On Codex
|
|
177
|
+
|
|
178
|
+
taskflow ships as a Codex **plugin** — install it once and the `taskflow_*` MCP tools plus a routing skill light up automatically, no manual `mcp add` and no config editing:
|
|
179
|
+
|
|
180
|
+
```bash
|
|
181
|
+
codex plugin marketplace add heggria/taskflow
|
|
182
|
+
codex plugin add taskflow@taskflow
|
|
183
|
+
```
|
|
184
|
+
|
|
185
|
+
The plugin's MCP server runs via `npx` (a version-pinned `codex-taskflow`), so there's nothing else to install globally and the plugin version binds the exact code that runs. Then just ask Codex to run a multi-phase or fan-out job and it calls the tools. See the [Codex guide](https://github.com/heggria/taskflow/blob/main/docs/codex-mcp.md).
|
|
186
|
+
|
|
187
|
+
### On Claude Code
|
|
188
|
+
|
|
189
|
+
taskflow ships as a Claude Code **plugin** too — install it once and the `taskflow_*` MCP tools plus a routing skill light up automatically, no manual `mcp add` and no config editing:
|
|
190
|
+
|
|
191
|
+
```bash
|
|
192
|
+
claude plugin marketplace add heggria/taskflow
|
|
193
|
+
claude plugin install claude-taskflow@taskflow
|
|
194
|
+
```
|
|
195
|
+
|
|
196
|
+
The plugin's MCP server runs via `npx` (a version-pinned `claude-taskflow`), so there's nothing else to install globally and the plugin version binds the exact code that runs. Each phase's subagent then runs as an isolated `claude -p` session. Just ask Claude Code to run a multi-phase or fan-out job and it calls the tools. See the [Claude Code guide](https://github.com/heggria/taskflow/blob/main/docs/claude-mcp.md).
|
|
197
|
+
|
|
198
|
+
### On OpenCode
|
|
199
|
+
|
|
200
|
+
OpenCode reaches taskflow through the same MCP server. Register it once — either with the CLI or by adding an `mcp` entry to your `opencode.json`:
|
|
201
|
+
|
|
202
|
+
```bash
|
|
203
|
+
opencode mcp add taskflow -- npx -y -p opencode-taskflow opencode-taskflow-mcp
|
|
204
|
+
```
|
|
205
|
+
|
|
206
|
+
```jsonc
|
|
207
|
+
// opencode.json
|
|
208
|
+
{
|
|
209
|
+
"$schema": "https://opencode.ai/config.json",
|
|
210
|
+
"mcp": {
|
|
211
|
+
"taskflow": {
|
|
212
|
+
"type": "local",
|
|
213
|
+
"command": ["npx", "-y", "-p", "opencode-taskflow", "opencode-taskflow-mcp"],
|
|
214
|
+
"enabled": true
|
|
215
|
+
}
|
|
216
|
+
}
|
|
217
|
+
}
|
|
218
|
+
```
|
|
219
|
+
|
|
220
|
+
The server runs via `npx` (a version-pinned `opencode-taskflow`), and each phase's subagent runs as an isolated `opencode run` session. OpenCode also auto-discovers the bundled routing skill (`**/SKILL.md`). Then just ask OpenCode to run a multi-phase or fan-out job and it calls the tools. See the [OpenCode guide](https://github.com/heggria/taskflow/blob/main/docs/opencode-mcp.md).
|
|
221
|
+
|
|
222
|
+
### The shorthand (same shape as the built-in tool)
|
|
223
|
+
|
|
224
|
+
```jsonc
|
|
225
|
+
// Single — one agent, one job
|
|
226
|
+
{ "task": "Summarize the architecture of src/", "agent": "explorer" }
|
|
227
|
+
|
|
228
|
+
// Parallel — fire several at once, outputs merge
|
|
229
|
+
{ "tasks": [
|
|
230
|
+
{ "task": "Audit auth in src/api", "agent": "analyst" },
|
|
231
|
+
{ "task": "Audit input validation in src/api", "agent": "analyst" }
|
|
232
|
+
] }
|
|
233
|
+
|
|
234
|
+
// Chain — sequential; each step sees the previous output
|
|
235
|
+
{ "chain": [
|
|
236
|
+
{ "task": "List the public API of src/lib", "agent": "scout" },
|
|
237
|
+
{ "task": "Write docs for:\n{previous.output}", "agent": "writer" }
|
|
238
|
+
] }
|
|
239
|
+
```
|
|
240
|
+
|
|
241
|
+
`agent` is optional (defaults to the first discovered agent). Add a `name` to label the run and unlock saving it as a command.
|
|
242
|
+
|
|
243
|
+
Shorthand modes also support per-step **context pre-reading** — pass `context` (file paths) and optionally `contextLimit` (max chars per file, default 8000) at the step level:
|
|
244
|
+
|
|
245
|
+
```jsonc
|
|
246
|
+
// Chain with context files injected into each step
|
|
247
|
+
{ "chain": [
|
|
248
|
+
{ "task": "List the public API", "agent": "scout", "context": ["src/lib/**/*.ts"] },
|
|
249
|
+
{ "task": "Write docs for:\n{previous.output}", "agent": "writer" }
|
|
250
|
+
] }
|
|
251
|
+
```
|
|
252
|
+
|
|
253
|
+
## Watch it run
|
|
254
|
+
|
|
255
|
+
This is not a mockup. **This is stdout from a real run** (the Pi TUI) — the `self-improve` flow that writes and verifies its own test suites, caught mid-flight by a quality gate:
|
|
256
|
+
|
|
257
|
+
```
|
|
258
|
+
⊗ taskflow self-improve 6/7 · blocked · $0.095
|
|
259
|
+
✓ discover agent deepseek-v4-flash 10t ↑38k ↓6.7k $0.011
|
|
260
|
+
┌ ✓ write-runner-tests agent claude-sonnet-4-6 10t ↑13 ↓6.6k $0.020
|
|
261
|
+
├ ✓ write-store-tests agent claude-sonnet-4-6 10t ↑11 ↓10k $0.018
|
|
262
|
+
├ ✓ write-agents-tests agent claude-sonnet-4-6 10t ↑28 ↓13k $0.030
|
|
263
|
+
└ ✓ fix-stability agent claude-sonnet-4-6 10t ↑13 ↓3.9k $0.012
|
|
264
|
+
✓ verify gate BLOCK 3 type errors in test files deepseek-v4-flash
|
|
265
|
+
⊘ report reduce skipped · Gate blocked ↳ fix-stability
|
|
266
|
+
```
|
|
267
|
+
|
|
268
|
+
**The layout *is* the DAG.** No dashboard, no logs to grep — you read the progress bar and you understand the whole pipeline:
|
|
269
|
+
|
|
270
|
+
- **Header** — `⊗` = blocked (a gate halted it); `6/7` phases processed; aggregate cost `$0.095`.
|
|
271
|
+
- **Status icons** — `✓` done · `◐` running · `✗` failed · `⊘` skipped · `○` pending.
|
|
272
|
+
- **Rail `┌ ├ └`** — phases in the same DAG layer, running concurrently. The four `write-*`/`fix-stability` tasks fan out from `discover`. A blank gutter = a single-phase layer.
|
|
273
|
+
- **`↳`** — a long, layer-skipping dependency. `report` depends on the adjacent `verify` *and* on `fix-stability` two layers back, so only that skip edge is annotated.
|
|
274
|
+
- **Gate** — `verify` emitted `VERDICT: BLOCK`, so the runtime skipped `report` and ended the run as `blocked`, surfacing the reason inline.
|
|
275
|
+
- **Detail** — per phase: model, token counts (`↑`in `↓`out), cost, timing. Fan-out phases also show sub-task progress (`3/15 2✗ 8▸`).
|
|
276
|
+
|
|
277
|
+
## Go declarative
|
|
278
|
+
|
|
279
|
+
The shorthand is your onramp. The DSL is where `taskflow` earns its keep — dynamic fan-out, structured routing, and quality gates.
|
|
280
|
+
|
|
281
|
+
### Fan out and reduce
|
|
282
|
+
|
|
283
|
+
```jsonc
|
|
284
|
+
{
|
|
285
|
+
"name": "summarize-files",
|
|
286
|
+
"description": "Discover files, summarize each, produce one report",
|
|
287
|
+
"args": { "dir": { "default": "." } },
|
|
288
|
+
"concurrency": 8,
|
|
289
|
+
"phases": [
|
|
290
|
+
{ "id": "discover", "type": "agent", "agent": "scout",
|
|
291
|
+
"task": "List source files under {args.dir} (non-recursive).\nOutput ONLY a JSON array [{\"file\":\"\"}]. No prose.",
|
|
292
|
+
"output": "json" },
|
|
293
|
+
{ "id": "summarize", "type": "map",
|
|
294
|
+
"over": "{steps.discover.json}", "as": "item", "agent": "scout",
|
|
295
|
+
"task": "Read {item.file} and give a one-sentence summary.",
|
|
296
|
+
"dependsOn": ["discover"] },
|
|
297
|
+
{ "id": "report", "type": "reduce", "from": ["summarize"], "agent": "writer",
|
|
298
|
+
"task": "Combine into a short overview:\n{steps.summarize.output}",
|
|
299
|
+
"dependsOn": ["summarize"], "final": true }
|
|
300
|
+
]
|
|
301
|
+
}
|
|
302
|
+
```
|
|
303
|
+
|
|
304
|
+
1. **`discover`** lists every file and emits a JSON array.
|
|
305
|
+
2. **`summarize`** is a `map` — it fans out one subagent per file, throttled to 8 concurrent, with `{item.file}` bound to each path.
|
|
306
|
+
3. **`report`** is a `reduce` — it merges every summary into one clean overview.
|
|
307
|
+
|
|
308
|
+
The intermediate summaries never enter your context. The runtime owns them; you get the report. **Save it once → `/tf:summarize-files dir=src` forever.**
|
|
309
|
+
|
|
310
|
+
### Route, gate, retry, approve, and cap the spend
|
|
311
|
+
|
|
312
|
+
```jsonc
|
|
313
|
+
{
|
|
314
|
+
"name": "triage-and-fix",
|
|
315
|
+
"budget": { "maxUSD": 1.5 },
|
|
316
|
+
"phases": [
|
|
317
|
+
{ "id": "triage", "type": "agent", "agent": "analyst", "output": "json",
|
|
318
|
+
"task": "Classify the bug. Output ONLY {\"severity\":\"high\"} or {\"severity\":\"low\"}." },
|
|
319
|
+
{ "id": "deep", "when": "{steps.triage.json.severity} == high", "dependsOn": ["triage"],
|
|
320
|
+
"agent": "executor-code", "task": "Root-cause and patch it.",
|
|
321
|
+
"retry": { "max": 2, "backoffMs": 500 } },
|
|
322
|
+
{ "id": "quick", "when": "{steps.triage.json.severity} == low", "dependsOn": ["triage"],
|
|
323
|
+
"agent": "executor-fast", "task": "Apply the quick fix." },
|
|
324
|
+
{ "id": "approve", "type": "approval", "join": "any", "dependsOn": ["deep", "quick"],
|
|
325
|
+
"task": "Review the fix before it ships." },
|
|
326
|
+
{ "id": "ship", "type": "agent", "dependsOn": ["approve"],
|
|
327
|
+
"task": "Open a PR with the change.", "final": true }
|
|
328
|
+
]
|
|
329
|
+
}
|
|
330
|
+
```
|
|
331
|
+
|
|
332
|
+
- **`when`** routes to `deep` *or* `quick` from the triage JSON — the other branch is skipped.
|
|
333
|
+
- **`join: "any"`** lets `approve` fire the moment whichever branch ran completes (an OR-join).
|
|
334
|
+
- **`retry`** re-runs a flaky patch with backoff; **`budget`** halts the whole run if it gets too expensive.
|
|
335
|
+
- **`approval`** pauses for a human (approve / reject / edit) before the final `ship`.
|
|
336
|
+
|
|
337
|
+
No scripting. No JavaScript `eval`. Just data the runtime executes — safe enough to run LLM-generated definitions directly.
|
|
338
|
+
|
|
339
|
+
### Loop until done
|
|
340
|
+
|
|
341
|
+
Some work is inherently iterative — refine a draft until a reviewer is satisfied, retry-and-improve until tests pass, converge on an answer:
|
|
342
|
+
|
|
343
|
+
```jsonc
|
|
344
|
+
{
|
|
345
|
+
"id": "refine",
|
|
346
|
+
"type": "loop",
|
|
347
|
+
"task": "Improve this draft (iteration {loop.iteration}). Previous attempt:\n{loop.lastOutput}\n\nReturn JSON {\"draft\":\"…\",\"done\":true|false}.",
|
|
348
|
+
"until": "{steps.refine.json.done} == true",
|
|
349
|
+
"output": "json",
|
|
350
|
+
"maxIterations": 6,
|
|
351
|
+
"convergence": true
|
|
352
|
+
}
|
|
353
|
+
```
|
|
354
|
+
|
|
355
|
+
See [Loop phases](#loop-until-done-loop) for the full reference.
|
|
356
|
+
|
|
357
|
+
### Plan, then execute (runtime sub-flows)
|
|
358
|
+
|
|
359
|
+
A planner decides *at runtime* what work to spawn — each iteration's plan depends on the previous result:
|
|
360
|
+
|
|
361
|
+
```jsonc
|
|
362
|
+
{
|
|
363
|
+
"name": "iterative-replan",
|
|
364
|
+
"phases": [
|
|
365
|
+
{ "id": "plan", "type": "agent", "agent": "planner",
|
|
366
|
+
"task": "Given the current state, output a JSON taskflow definition (with phases[]).",
|
|
367
|
+
"output": "json" },
|
|
368
|
+
{ "id": "execute", "type": "flow", "def": "{steps.plan.json}",
|
|
369
|
+
"dependsOn": ["plan"] }
|
|
370
|
+
]
|
|
371
|
+
}
|
|
372
|
+
```
|
|
373
|
+
|
|
374
|
+
The generated sub-flow is **validated** (no cycles, no dangling refs, no duplicate IDs) before a single token is spent. See [`examples/dynamic-plan-execute.json`](https://github.com/heggria/taskflow/blob/main/examples/dynamic-plan-execute.json) and [`examples/iterative-replan.json`](https://github.com/heggria/taskflow/blob/main/examples/iterative-replan.json).
|
|
375
|
+
|
|
376
|
+
### Tournament (compete and judge)
|
|
377
|
+
|
|
378
|
+
For open-ended creative or subjective work, spawn several competing variants and let a judge pick the best:
|
|
379
|
+
|
|
380
|
+
```jsonc
|
|
381
|
+
{
|
|
382
|
+
"id": "headline",
|
|
383
|
+
"type": "tournament",
|
|
384
|
+
"task": "Write a punchy headline for this launch post.",
|
|
385
|
+
"variants": 4,
|
|
386
|
+
"judge": "Pick the headline with the strongest hook and clearest promise.",
|
|
387
|
+
"mode": "best"
|
|
388
|
+
}
|
|
389
|
+
```
|
|
390
|
+
|
|
391
|
+
See [Tournament phases](#tournament-tournament) for the full reference.
|
|
392
|
+
|
|
393
|
+
## Phase types
|
|
394
|
+
|
|
395
|
+
| type | what it does | required fields |
|
|
396
|
+
|------|--------------|-----------------|
|
|
397
|
+
| `agent` | one subagent runs a single task | `task` |
|
|
398
|
+
| `parallel` | run `branches[]` concurrently | `branches` (array of `{task, agent?}`) |
|
|
399
|
+
| `map` | **fan out** over an array — one subagent per item, `{item}` bound | `over`, `task` |
|
|
400
|
+
| `gate` | quality/review step that can **halt the flow** | `task` |
|
|
401
|
+
| `reduce` | aggregate `from[]` phase outputs into one | `from`, `task` |
|
|
402
|
+
| `approval` | **human-in-the-loop** pause — approve / reject / edit | — |
|
|
403
|
+
| `flow` | run a **sub-flow** as one phase — a **saved** flow (`use`) or a **runtime-generated** one (`def`) | `use` \| `def` |
|
|
404
|
+
| `loop` | **iterate a task until done** — re-run a body until a condition, convergence, or a cap | `task`, `until` |
|
|
405
|
+
| `tournament` | **N variants compete**, a judge picks the best (or aggregates) | `task` \| `branches` |
|
|
406
|
+
| `script` | run a **shell command** — no LLM, zero tokens — capturing stdout as the phase output | `run` |
|
|
407
|
+
|
|
408
|
+
### Common phase fields
|
|
409
|
+
|
|
410
|
+
Every phase needs a unique `id` and a `type` (defaults to `agent`). On top of the per-type fields:
|
|
411
|
+
|
|
412
|
+
| Field | Meaning |
|
|
413
|
+
|---|---|
|
|
414
|
+
| `agent` | Agent to run (defaults to the first discovered agent) |
|
|
415
|
+
| `dependsOn` | Phase ids this phase waits for — builds the DAG |
|
|
416
|
+
| `join` | `"all"` (default) waits for every dep; `"any"` is an OR-join |
|
|
417
|
+
| `when` | Conditional guard — skip unless the expression is truthy |
|
|
418
|
+
| `retry` | `{ max, backoffMs?, factor? }` — retry a failing subagent |
|
|
419
|
+
| `output` | `"text"` (default) or `"json"` (exposes `{steps.ID.json}`) |
|
|
420
|
+
| `model` / `thinking` / `tools` | Per-phase overrides for the subagent |
|
|
421
|
+
| `cwd` | Working directory for the subagent. A literal path, or a reserved keyword for **workspace isolation** — `"temp"` (ephemeral dir, removed after), `"dedicated"` (persistent dir under the run state, kept), `"worktree"` (a git worktree on a throwaway branch, removed after). Fail-open; rejected in LLM-authored sub-flows. |
|
|
422
|
+
| `context` | File paths to pre-read and inject into the agent prompt |
|
|
423
|
+
| `contextLimit` | Max chars per context file (default 8000) |
|
|
424
|
+
| `concurrency` | Fan-out cap for `map` / `parallel` (overrides the flow default) |
|
|
425
|
+
| `final` | Marks the result-bearing phase (else the last phase wins) |
|
|
426
|
+
| `optional` | A failure here does **not** abort the run |
|
|
427
|
+
| `shareContext` | Opt this phase's subagent into the **Shared Context Tree** (see below). Set `contextSharing: true` at the flow level to enable it for every phase |
|
|
428
|
+
| `cache` | `{ scope, ttl?, fingerprint? }` — cross-run memoization (see below) |
|
|
429
|
+
| `onBlock` | `"halt"` (default) or `"retry"` — what happens when a gate blocks |
|
|
430
|
+
| `eval` | Zero-token machine-checkable criteria that run *before* the LLM gate |
|
|
431
|
+
|
|
432
|
+
Flow-level keys: `name`, `description`, `args`, `concurrency` (default 8), `agentScope`, `contextSharing`, `strictInterpolation`, and `budget: { maxUSD?, maxTokens? }`.
|
|
433
|
+
|
|
434
|
+
### Shared Context Tree (blackboard + supervision)
|
|
435
|
+
|
|
436
|
+
By default subagents are fully isolated — they share nothing and only return a
|
|
437
|
+
final string. Opt a phase in with `shareContext: true` (or `contextSharing: true`
|
|
438
|
+
flow-wide) to give its subagent four extra tools backed by a per-run, file-based
|
|
439
|
+
blackboard:
|
|
440
|
+
|
|
441
|
+
| tool | direction | use |
|
|
442
|
+
|------|-----------|-----|
|
|
443
|
+
| `ctx_write(key, value)` | horizontal | publish a finding so siblings/descendants reuse it (stop re-reading the same files) |
|
|
444
|
+
| `ctx_read(key?)` | horizontal | read findings visible to this node: its own + ancestors' + **completed** others' |
|
|
445
|
+
| `ctx_report(summary, structured?)` | vertical ↑ | report a result up to the parent |
|
|
446
|
+
| `ctx_spawn(assignments[])` | vertical ↓ | delegate child work at runtime; each assignment is a flat `{task}` **or** a `{subflow}` (a dependency-bearing DAG the runtime validates and runs nested). Child reports fold back into this phase's output |
|
|
447
|
+
|
|
448
|
+
The first two are a **horizontal blackboard** (siblings reuse expensive context);
|
|
449
|
+
the last two are a **vertical supervision tree** (a node delegates work and its
|
|
450
|
+
children report up). Everything is opt-in, fail-open, depth-capped (5 levels), size-bounded
|
|
451
|
+
(256KB per value, 256 keys per node, 16 spawn assignments max), and cleaned up
|
|
452
|
+
with the run — flows that don't opt in behave exactly as before.
|
|
453
|
+
|
|
454
|
+
```jsonc
|
|
455
|
+
{ "id": "survey", "type": "agent", "agent": "scout", "shareContext": true,
|
|
456
|
+
"task": "Map the API surface. ctx_write key 'endpoints' so the auditors don't re-scan." },
|
|
457
|
+
{ "id": "audit", "type": "map", "over": "{steps.survey.json}", "shareContext": true,
|
|
458
|
+
"dependsOn": ["survey"], "agent": "analyst",
|
|
459
|
+
"task": "ctx_read 'endpoints' for shared context, then audit {item} for missing auth." }
|
|
460
|
+
```
|
|
461
|
+
|
|
462
|
+
### Library: search before author
|
|
463
|
+
|
|
464
|
+
A flow that took work to generalize is an asset — but only if you can find it
|
|
465
|
+
again. Phase 1 of the **taskflow library** adds sidecar `.meta.json` metadata
|
|
466
|
+
(`purpose`, `tags`, `phaseSignature`, `generality`, `reuseCount`) to saved
|
|
467
|
+
flows, plus a search tool that surfaces reusable flows before you write a new
|
|
468
|
+
one:
|
|
469
|
+
|
|
470
|
+
```jsonc
|
|
471
|
+
// MCP
|
|
472
|
+
{ "name": "taskflow_search", "arguments": { "query": "audit API endpoints for missing auth" } }
|
|
473
|
+
// Pi tool
|
|
474
|
+
{ "action": "search", "query": "审计接口鉴权" }
|
|
475
|
+
```
|
|
476
|
+
|
|
477
|
+
Search is **structural + keyword** by default (zero dependencies, zero tokens):
|
|
478
|
+
`phaseSignature` (`agent→map→reduce`) and phase-count similarity are blended
|
|
479
|
+
with keyword overlap. It is **CJK-aware**, so a Chinese query like
|
|
480
|
+
`"检查接口安全性 鉴权 缺失"` still matches a purpose containing
|
|
481
|
+
`"审计...是否缺少鉴权检查"`. Embedding-based semantic search is Phase 2 —
|
|
482
|
+
the `embedder` seam is already in place and search degrades gracefully when it
|
|
483
|
+
is not configured.
|
|
484
|
+
|
|
485
|
+
When you run a flow you discovered via search, pass `reusedFromSearch: true`
|
|
486
|
+
(`taskflow_run` MCP, or `action=run` in Pi). That bumps the flow's
|
|
487
|
+
`reuseCount`, so high-quality reusable patterns rank higher over time. When you
|
|
488
|
+
write a new reusable flow, save it with metadata:
|
|
489
|
+
|
|
490
|
+
```jsonc
|
|
491
|
+
{ "name": "taskflow_save",
|
|
492
|
+
"arguments": {
|
|
493
|
+
"name": "audit-endpoints",
|
|
494
|
+
"definition": { ... },
|
|
495
|
+
"purpose": "Audit API endpoints for missing auth checks",
|
|
496
|
+
"tags": ["audit", "security", "auth"] } }
|
|
497
|
+
```
|
|
498
|
+
|
|
499
|
+
See `docs/rfc-library-reuse.md` for the design and `skills-src/taskflow/library.md`
|
|
500
|
+
for the agent-facing search→reuse→generalize→re-save workflow.
|
|
501
|
+
|
|
502
|
+
### Control flow & reliability
|
|
503
|
+
|
|
504
|
+
- **`when`** — skip a phase unless an expression is truthy. Supports `{refs}`, `== != < > <= >=`, `&& || !`, parentheses, and quoted strings/numbers. Pair with `join: "any"` on the merge phase for real if/else routing. Parse errors **fail open** (the phase runs — never silently dropped).
|
|
505
|
+
- **`join: "any"`** — an OR-join: the phase runs as soon as *one* dependency completes (default `"all"` waits for all).
|
|
506
|
+
- **`retry`** — `{ "max": 2, "backoffMs": 500, "factor": 2 }` retries a failing subagent with fixed or exponential backoff; usage is summed and the attempt count shows as `↻N` in the TUI. Transient provider errors (rate-limit / 5xx / timeout) **auto-retry even without an explicit policy**; hard errors don't.
|
|
507
|
+
- **`onBlock`** — `"halt"` (default) stops the run when a gate blocks. `"retry"` retries upstream phases when a gate blocks, instead of halting — a self-healing rework loop with budget and idle-watchdog guards and a nested recursion depth cap.
|
|
508
|
+
- **`eval`** — zero-token machine-checkable criteria that run *before* the LLM gate. If the eval check fails, the gate blocks without spawning an agent.
|
|
509
|
+
- **`score`** — graded, composable quality gates: deterministic scorers (`exact-match`, `contains`, `regex`, `json-schema`, `length-range`, `code-compiles`) run against a target string at **zero tokens** and combine via `all`/`any`/`weighted` against a `threshold`. Deterministic pass → auto-PASS with no LLM call **when the judge cannot veto** — no judge configured, or `weighted` where the deterministic score is a *lower bound* already clearing the threshold. With `all`/`any` + a judge, the judge always runs (its verdict is authoritative — it may check what scorers cannot, e.g. factuality). Deterministic fail → the optional LLM `judge` decides (fail-open on unparseable output), or the gate `task` runs with the scorer report appended, or — with no fallback — the gate **blocks explicitly**. The structured result is the gate's `.json` (`{steps.<gate>.json.combined}`, `.json.results`), so downstream phases can route on quality, not just pass/fail. LLM-generated dynamic sub-flows may not use `code-compiles` (compiler execution) or `regex` (ReDoS) scorers — same hardening class as the `script` block.
|
|
510
|
+
- **`idempotent: false`** — side-effect classification for phases with **irreversible effects** (webhook POSTs, deploys, DB writes): the implicit transient auto-retry is suppressed (an explicit `retry{}` is still honored — it's the author's declaration that repeats are acceptable) and the result is **never cached** in any scope (within-run resume, cross-run, `incremental`) — the phase re-runs every time. The phase state records `sideEffect: true` (rendered as ⚡). Default `true` — existing flows are unchanged.
|
|
511
|
+
- **`approval`** — pause for a human (Approve / Reject / Edit). Reject halts the flow; Edit injects the typed note as the phase output for downstream steps. Non-interactive runs (detached / CI) **auto-reject** (safety: approval gates are never bypassed).
|
|
512
|
+
- **`flow`** — `{ "type": "flow", "use": "deep-research", "with": { "topic": "{item}" } }` runs a **saved** flow as a phase (recursion is detected and rejected). Or **generate the sub-flow at runtime**: `{ "type": "flow", "def": "{steps.plan.json}" }` resolves an upstream phase's JSON output into a sub-flow, **validates it (cycles / dangling refs / duplicate ids / dead-ends), then runs it** — the number and shape of the generated phases is decided at runtime, not authored in advance. A malformed plan fails *open* (the phase is skipped with a `defError`, the run continues). This is how a planner decides *at runtime* what work to spawn — the declarative answer to a code-mode `for` loop, with each generated plan checked before it spends a token. Security hardening for LLM-generated sub-flows: breadth caps (100 phases, 200 map items, 16 concurrency), `cwd` containment, budget clamped to `min(child, parent)`, nesting cap (5 levels), and prototype-pollution defense (deep-cloned, `__proto__`/`constructor`/`prototype` stripped). Pair it with `loop` for **data-dependent iterative replanning** (round N's plan depends on round N-1's result). See [`examples/dynamic-plan-execute.json`](https://github.com/heggria/taskflow/blob/main/examples/dynamic-plan-execute.json) and [`examples/iterative-replan.json`](https://github.com/heggria/taskflow/blob/main/examples/iterative-replan.json).
|
|
513
|
+
|
|
514
|
+
### Loop-until-done (`loop`)
|
|
515
|
+
|
|
516
|
+
Some work is inherently iterative — refine a draft until a reviewer is satisfied, retry-and-improve until tests pass, converge on an answer. A `loop` phase re-runs one task body until a stop condition holds:
|
|
517
|
+
|
|
518
|
+
```jsonc
|
|
519
|
+
{
|
|
520
|
+
"id": "refine",
|
|
521
|
+
"type": "loop",
|
|
522
|
+
"task": "Improve this draft (iteration {loop.iteration}). Previous attempt:\n{loop.lastOutput}\n\nReturn JSON {\"draft\":\"…\",\"done\":true|false}.",
|
|
523
|
+
"until": "{steps.refine.json.done} == true", // the iteration's own output is exposed here
|
|
524
|
+
"output": "json",
|
|
525
|
+
"maxIterations": 6, // default 10, hard cap 100 — the loop ALWAYS terminates
|
|
526
|
+
"convergence": true // default: stop early if an iteration's output is identical to the last
|
|
527
|
+
}
|
|
528
|
+
```
|
|
529
|
+
|
|
530
|
+
- **Body locals** — the task can read `{loop.iteration}` (1-based), `{loop.lastOutput}` (the prior iteration's output), and `{loop.maxIterations}` to build on its own previous work; all three are also available to the `until` condition.
|
|
531
|
+
- **`until`** — evaluated after each iteration with the iteration's output exposed as `{steps.<thisId>.output}` / `.json`. Same operators as `when`. The loop stops the moment it's truthy.
|
|
532
|
+
- **Always terminates.** Four independent stops: `until` truthy, **convergence** (a fixed point — output identical to the previous iteration), **`maxIterations`** (hard-capped at 100), or a **failing iteration** (the phase fails with the partial output preserved). A malformed `until` **stops** the loop rather than spinning forever (fail-safe) and surfaces a warning on the phase.
|
|
533
|
+
- **Reflexion memory (`reflexion: true`)** — by default each iteration only sees the prior *output*; the **reason** it wasn't good enough (an `expect` contract violation, an error, the unmet `until`) is discarded, so models repeat mistakes. With `reflexion: true` every iteration after the first receives a structured failure summary of the prior one via the `{reflexion}` placeholder (auto-appended when absent, capped at 2000 chars): contract diagnostics like `$.done: required key is missing`, the error message, or the unmet stop condition, plus a truncated output snippet. Semantics shift to enable self-correction: **body failures become feedback instead of terminating the loop** — timeout/abort still hard-stop, and exhausting `maxIterations` on a failure still fails the phase (reflexion defers failure, never erases it). The last injected summary is persisted on the phase state for audit.
|
|
534
|
+
- The TUI shows `↻N` with the stop reason (`done` / `converged` / `max` / `failed`); usage is summed across iterations. Like `gate`/`approval`, `loop` is **excluded from `cross-run` cache** (each run must iterate fresh).
|
|
535
|
+
|
|
536
|
+
### Tournament (`tournament`)
|
|
537
|
+
|
|
538
|
+
For open-ended work, the best result often comes from generating several candidates and picking the strongest — best-of-N with a judge, in one declarative phase:
|
|
539
|
+
|
|
540
|
+
```jsonc
|
|
541
|
+
{
|
|
542
|
+
"id": "headline",
|
|
543
|
+
"type": "tournament",
|
|
544
|
+
"task": "Write a punchy headline for this launch post.",
|
|
545
|
+
"variants": 4, // spawn 4 competitors of the SAME task (default 3, max 20)
|
|
546
|
+
"judge": "Pick the headline with the strongest hook and clearest promise.",
|
|
547
|
+
"judgeAgent": "reviewer", // optional; defaults to the phase agent
|
|
548
|
+
"mode": "best" // "best" (default) | "aggregate"
|
|
549
|
+
}
|
|
550
|
+
```
|
|
551
|
+
|
|
552
|
+
- **Competitors** — either `variants: N` copies of one `task` (diversity comes from model nondeterminism), or distinct `branches: [{task, agent?}, …]` when you want to pit *different approaches* against each other.
|
|
553
|
+
- **Judge** — after the fan-out, one judge agent sees every variant (numbered) plus your `judge` rubric and picks a winner via a `WINNER: <n>` line or `{"winner": n}`. An unreadable verdict **fails open** to variant 1; a failed judge falls back too — the work is never lost.
|
|
554
|
+
- **`mode`** — `best` returns the winning variant **verbatim**; `aggregate` returns the judge's **synthesized** answer combining the strongest parts.
|
|
555
|
+
- **Short-circuits:** if only one competitor survives, it wins with no judge call; if all fail, the phase fails. The TUI shows `⚑ N→#k`; usage sums variants + judge. Like `gate`, it's **excluded from `cross-run` cache**.
|
|
556
|
+
|
|
557
|
+
### Shell steps (`script`)
|
|
558
|
+
|
|
559
|
+
Not every step needs a model. A `script` phase runs a **shell command** directly — zero tokens, no subagent — and captures its stdout as the phase output. Use it to glue LLM work to real tools: run a build or test suite, a formatter, `git`, `curl` a webhook, or pipe a previous phase's output through a script.
|
|
560
|
+
|
|
561
|
+
```jsonc
|
|
562
|
+
{
|
|
563
|
+
"id": "build",
|
|
564
|
+
"type": "script",
|
|
565
|
+
"run": "pnpm run build", // string → runs in a shell
|
|
566
|
+
"timeout": 120000 // optional ms cap (1000–300000, default 60000)
|
|
567
|
+
},
|
|
568
|
+
{
|
|
569
|
+
"id": "score",
|
|
570
|
+
"type": "script",
|
|
571
|
+
"run": ["python", "score.py"], // array → direct exec, no shell (injection-safe)
|
|
572
|
+
"input": "{steps.analyze.output}", // optional — piped to stdin (interpolation-enabled)
|
|
573
|
+
"dependsOn": ["analyze"]
|
|
574
|
+
}
|
|
575
|
+
```
|
|
576
|
+
|
|
577
|
+
- **`run`** — the command. A **string** runs through a shell (`sh -c` / `cmd`); an **array** is spawned directly (execvp-style, no shell). Prefer the array form for anything containing interpolated values: a string `run` that contains an interpolation placeholder is **rejected at validation** (a shell-injection guard) — pass dynamic values via the array form or `input` instead.
|
|
578
|
+
- **`input`** — optional text piped to the command's stdin; supports interpolation (`{steps.X.output}`, `{args.X}`). If omitted, stdin is closed.
|
|
579
|
+
- **`timeout`** — optional millisecond cap (1000–300000, default 60000). On timeout the child gets `SIGTERM`, then `SIGKILL` after a grace period, and the phase fails.
|
|
580
|
+
- A non-zero exit **fails** the phase (stderr is captured); stdout is capped at 1 MB. `script` phases spend **zero tokens**, do not support `retry` or `output: "json"`, and are **excluded from `cross-run` cache** (a shell step may have side effects). The `compile` diagram renders them as `⚡ script`.
|
|
581
|
+
|
|
582
|
+
### Cross-run memoization (`cache`)
|
|
583
|
+
|
|
584
|
+
Every phase is already content-addressed: within a single run's **resume**, a phase whose resolved inputs are unchanged is skipped. `cache` extends that reuse **across independent runs** — if any prior run computed a phase with an identical input hash, its result is reused for **$0.00**.
|
|
585
|
+
|
|
586
|
+
```jsonc
|
|
587
|
+
{
|
|
588
|
+
"id": "analyze-auth",
|
|
589
|
+
"task": "Summarize how the auth module works.",
|
|
590
|
+
"context": ["src/auth/**/*.ts"],
|
|
591
|
+
"cache": {
|
|
592
|
+
"scope": "cross-run", // "run-only" (default) | "cross-run" | "off"
|
|
593
|
+
"ttl": "6h", // optional max age before a hit is treated as a miss
|
|
594
|
+
"fingerprint": ["git:HEAD", "glob:src/auth/**/*.ts"] // fold world-state into the key
|
|
595
|
+
}
|
|
596
|
+
}
|
|
597
|
+
```
|
|
598
|
+
|
|
599
|
+
- **`scope`** — `"run-only"` (default) is exactly the historical behavior (within-run resume only). `"cross-run"` opts the phase into the persistent store. `"off"` disables reuse entirely (even within a run), for debugging.
|
|
600
|
+
- **Freshness is the whole game.** The cache key already includes the prompt, the `over` items, and any `context` files (pre-read into the task). `fingerprint` folds *implicit* inputs into the key so "the world changed" becomes a cache miss: `git:HEAD`, `glob:<pat>` (size+mtime), `glob!:<pat>` (content hash), `file:<path>`, `env:<NAME>`. `ttl` (`30m`/`6h`/`7d`) is a time backstop.
|
|
601
|
+
- **Honest limit:** a subagent that reads a file it didn't declare in `context`/`fingerprint` can still serve a stale `cross-run` hit. That's why the default is `run-only` and why `gate`/`approval` phases are **forbidden** from `cross-run` (they must produce a fresh result each run). Opt in only for phases whose output is a function of declared inputs.
|
|
602
|
+
- Cache lives in `.pi/taskflows/cache/` (gitignored). Clear it with `action: "cache-clear"` on the tool. Full rationale: [`docs/internal/rfc-cross-run-memoization.md`](https://github.com/heggria/taskflow/blob/main/docs/internal/rfc-cross-run-memoization.md).
|
|
603
|
+
|
|
604
|
+
### Gate phases (quality control)
|
|
605
|
+
|
|
606
|
+
A `gate` runs an agent to review upstream output and can **block the rest of the workflow.** End the gate task by asking for a verdict the runtime can read:
|
|
607
|
+
|
|
608
|
+
- a final line `VERDICT: PASS` or `VERDICT: BLOCK` (also accepts `OK`, `FAIL`, `STOP`, `REJECT`, `HALT` — last occurrence wins), or
|
|
609
|
+
- JSON like `{"continue": false, "reason": "missing auth checks"}` / `{"verdict": "block", "reason": "..."}`.
|
|
610
|
+
|
|
611
|
+
On **BLOCK**, downstream phases skip and the run ends as `blocked` with the reason surfaced. **Ambiguous output fails open** (treated as PASS) — a gate never halts your flow by accident.
|
|
612
|
+
|
|
613
|
+
```
|
|
614
|
+
Review the audit below. If any endpoint is missing auth, end with
|
|
615
|
+
"VERDICT: BLOCK" and a one-line reason; otherwise end with "VERDICT: PASS".
|
|
616
|
+
|
|
617
|
+
{steps.audit.output}
|
|
618
|
+
```
|
|
619
|
+
|
|
620
|
+
## Interpolation & expressions
|
|
621
|
+
|
|
622
|
+
| placeholder | resolves to |
|
|
623
|
+
|---|---|
|
|
624
|
+
| `{args.X}` | invocation argument |
|
|
625
|
+
| `{steps.ID.output}` | a prior phase's text output |
|
|
626
|
+
| `{steps.ID.json}` | prior output parsed as JSON (or `{steps.ID.json.field}`) |
|
|
627
|
+
| `{item}` / `{item.field}` | current item inside a `map` phase |
|
|
628
|
+
| `{previous.output}` | the immediately-upstream phase output |
|
|
629
|
+
| `{loop.iteration}` | current iteration number inside a `loop` phase |
|
|
630
|
+
| `{loop.lastOutput}` | previous iteration's output inside a `loop` phase |
|
|
631
|
+
| `{loop.maxIterations}` | the iteration cap inside a `loop` phase |
|
|
632
|
+
|
|
633
|
+
Condition grammar (for `when`): `== != < > <= >=`, `&& || !`, parentheses, quoted strings/numbers, and any `{...}` reference — e.g. `"when": "{steps.triage.json.route} == deep && {args.force} != true"`.
|
|
634
|
+
|
|
635
|
+
> Referencing `{steps.X}` that isn't declared in `dependsOn` is a **hard validation error** — the runtime catches the most common pipeline bug before a single agent runs.
|
|
636
|
+
|
|
637
|
+
> Unresolved interpolation refs (e.g. `{args.typo}` or a missing `dependsOn`) are surfaced as **phase warnings** (`PhaseState.warnings`) in the run record and `/tf runs` — no more silent intact placeholders.
|
|
638
|
+
|
|
639
|
+
## Commands
|
|
640
|
+
|
|
641
|
+
Saved flows become CLI shortcuts. **These `/tf` commands are Pi-only** (they run in the Pi session). On Codex, Claude Code, and OpenCode, use the `taskflow_*` MCP tools instead — `taskflow_list` / `taskflow_show` / `taskflow_run` (by `name`) / `taskflow_verify` / `taskflow_compile` / `taskflow_peek`.
|
|
642
|
+
|
|
643
|
+
| Command | What it does |
|
|
644
|
+
|---|---|
|
|
645
|
+
| `/tf list` | List all saved flows |
|
|
646
|
+
| `/tf run <name> [args]` | Run a saved flow (e.g. `/tf run summarize-files dir=src`) |
|
|
647
|
+
| `/tf show <name>` | Print a flow's definition |
|
|
648
|
+
| `/tf compile <name> [lr\|td]` | **Render the flow as a Mermaid diagram + verification overlay** — 0 tokens, no LLM; paste into a README/issue/PR |
|
|
649
|
+
| `/tf runs` | Browse recent run history (interactive TUI — **live auto-refreshes** while any run is active) |
|
|
650
|
+
| `/tf resume <runId>` | Continue a paused/failed run — cached phases skip automatically |
|
|
651
|
+
| `/tf init` | **Interactively map model roles** to your enabled models (writes `~/.pi/agent/settings.json`) |
|
|
652
|
+
| `/tf:<name> [args]` | Shortcut — runs the flow in one tap |
|
|
653
|
+
|
|
654
|
+
Tool actions (used by the model on Pi): `run` (inline `define` or saved `name`), `save`, `resume`, `list`, `agents`, `init`, `verify`, `compile`, `ir`, `provenance`, `why-stale`, `recompute`, `cache-clear`. On Codex, Claude Code, and OpenCode the exposed MCP tools are `taskflow_run` / `taskflow_list` / `taskflow_show` / `taskflow_verify` / `taskflow_compile` / `taskflow_peek`.
|
|
655
|
+
|
|
656
|
+
## Background (detached) execution
|
|
657
|
+
|
|
658
|
+
Pass `detach: true` to run a taskflow in a detached child process — the tool returns immediately with the `runId` and the flow continues running even if the host session exits:
|
|
659
|
+
|
|
660
|
+
```jsonc
|
|
661
|
+
{
|
|
662
|
+
"action": "run",
|
|
663
|
+
"name": "nightly-audit",
|
|
664
|
+
"detach": true
|
|
665
|
+
}
|
|
666
|
+
```
|
|
667
|
+
|
|
668
|
+
- The child process reads serialized context, calls the orchestration engine, and persists terminal state to the store.
|
|
669
|
+
- Status is polled via `/tf runs` (which now **auto-refreshes live** when any run is running) or `action: "resume"`.
|
|
670
|
+
- Stale PID detection via signal-0 probe; the idle watchdog kills stalled children.
|
|
671
|
+
- **Approval phases auto-reject** in detached mode — human gates are never silently bypassed.
|
|
672
|
+
- `resume` works normally after a detached run completes or fails.
|
|
673
|
+
|
|
674
|
+
## Resume across sessions
|
|
675
|
+
|
|
676
|
+
A taskflow run isn't tied to your session. Every completed phase is written to disk, so a run that fails (or that you stop) can be continued later with `/tf resume <runId>` — **cached phases skip automatically** and only the remaining work spends tokens.
|
|
677
|
+
|
|
678
|
+
<div align="center">
|
|
679
|
+
<img src="https://raw.githubusercontent.com/heggria/taskflow/main/assets/resume.png" alt="A run fails midway in session 1; in session 2 /tf resume skips the cached phases and only re-runs the failed phase and what follows" width="900">
|
|
680
|
+
</div>
|
|
681
|
+
|
|
682
|
+
Resume is keyed on each phase's input hash — if an upstream output changed, dependent phases re-run; if nothing changed, they're reused. No competing Pi extension does this across sessions.
|
|
683
|
+
|
|
684
|
+
## Storage
|
|
685
|
+
|
|
686
|
+
```
|
|
687
|
+
.pi/taskflows/<name>.json # project-scope definitions (commit to share)
|
|
688
|
+
~/.pi/agent/taskflows/<name>.json # user-scope definitions
|
|
689
|
+
.pi/taskflows/runs/<flowName>/<runId>.json # run state for resume (gitignore this)
|
|
690
|
+
.pi/taskflows/cache/ # cross-run memoization cache (gitignored)
|
|
691
|
+
```
|
|
692
|
+
|
|
693
|
+
> Commit `.pi/taskflows/` and your whole team shares the pipelines — no config sync, no onboarding doc. Run state is written atomically via `writeFileAtomic()` (temp file + `renameSync`) and guarded by a zero-dependency file lock (`O_CREAT|O_EXCL` with stale-lock steal via atomic rename), so concurrent runs never corrupt the index.
|
|
694
|
+
|
|
695
|
+
Agent discovery scope (via `agentScope` in the flow definition):
|
|
696
|
+
|
|
697
|
+
| value | discovers agents from |
|
|
698
|
+
|---|---|
|
|
699
|
+
| `"user"` (default) | `~/.pi/agent/agents/*.md` |
|
|
700
|
+
| `"project"` | `.pi/agents/*.md` (walks up the tree) |
|
|
701
|
+
| `"both"` | user + project; project wins on name collision |
|
|
702
|
+
|
|
703
|
+
Run cleanup is configurable via `maxKeptRuns` and `maxRunAgeDays` in settings.
|
|
704
|
+
|
|
705
|
+
## Agents
|
|
706
|
+
|
|
707
|
+
Taskflow ships **18 built-in agents** — each a `.md` file with a tuned system prompt, thinking level, and tool set. You can reference them by `name` in any phase or shorthand, right after install. No setup required.
|
|
708
|
+
|
|
709
|
+
### Built-in agent roster
|
|
710
|
+
|
|
711
|
+
| Agent | Role | Thinking | Default role |
|
|
712
|
+
|---|---|---:|---|
|
|
713
|
+
| `executor` | Implement planned code changes | high | `{{fast}}` |
|
|
714
|
+
| `executor-fast` | Trivial fixes (≤2 files, ≤50 lines) | off | `{{fast}}` |
|
|
715
|
+
| `executor-code` | Complex multi-file implementation | high | `{{strong}}` |
|
|
716
|
+
| `executor-ui` | Frontend / styling / visual changes | high | `{{vision}}` |
|
|
717
|
+
| `scout` | Fast codebase recon & file mapping | off | `{{fast}}` |
|
|
718
|
+
| `planner` | Implementation plan creation | high | `{{strong}}` |
|
|
719
|
+
| `analyst` | Requirements analysis, ambiguity detection | high | `{{thinker}}` |
|
|
720
|
+
| `critic` | Inline self-doubt during reasoning | xhigh | `{{thinker}}` |
|
|
721
|
+
| `reviewer` | General code / architecture review | high | `{{strong}}` |
|
|
722
|
+
| `risk-reviewer` | Backend / infra / DB / API risk | high | `{{reasoner}}` |
|
|
723
|
+
| `security-reviewer` | Security vulns, auth/crypto | xhigh | `{{reasoner}}` |
|
|
724
|
+
| `plan-arbiter` | Plan quality gate (complex tasks) | high | `{{arbiter}}` |
|
|
725
|
+
| `final-arbiter` | Tiebreaker when critics disagree | xhigh | `{{arbiter}}` |
|
|
726
|
+
| `test-engineer` | Design & implement tests | high | `{{fast}}` |
|
|
727
|
+
| `doc-writer` | Documentation authoring | off | `{{fast}}` |
|
|
728
|
+
| `recover` | Session recovery after compaction | low | `{{fast}}` |
|
|
729
|
+
| `verifier` | Run tests, validate outcomes | off | `{{fast}}` |
|
|
730
|
+
| `visual-explorer` | Figma design metadata analysis | high | `{{vision}}` |
|
|
731
|
+
|
|
732
|
+
Agents are layered: **built-in → user (`~/.pi/agent/agents/`) → project (`.pi/agents/`)**. A user or project agent with the same `name` overrides the built-in — so you can customize any agent without touching the package.
|
|
733
|
+
|
|
734
|
+
### Model roles
|
|
735
|
+
|
|
736
|
+
Each built-in agent's `model` field uses a **role placeholder** (e.g. `{{fast}}`) instead of a hardcoded provider string. This decouples *intent* from *implementation* — you map roles to models once, and every agent adapts.
|
|
737
|
+
|
|
738
|
+
| Role | Intent | Typical model |
|
|
739
|
+
|---|---|---|
|
|
740
|
+
| `{{fast}}` | Cheap & quick — high-volume, low-stakes | DeepSeek V4 Flash |
|
|
741
|
+
| `{{strong}}` | Balanced — planning, review, moderate complexity | MiMo v2.5 Pro |
|
|
742
|
+
| `{{thinker}}` | Deep analysis — requirements, critique | DeepSeek V4 Pro |
|
|
743
|
+
| `{{arbiter}}` | Final judgment — tiebreak, plan quality gates | Qwen 3.7 Max |
|
|
744
|
+
| `{{vision}}` | Multimodal — UI work, design reading | MiniMax M3 |
|
|
745
|
+
| `{{reasoner}}` | Cautious reasoning — security, risk | GLM 5.1 |
|
|
746
|
+
|
|
747
|
+
Without configuration, agents fall back to Pi's default model. To map roles to real models, run the interactive setup:
|
|
748
|
+
|
|
749
|
+
```bash
|
|
750
|
+
/tf init
|
|
751
|
+
```
|
|
752
|
+
|
|
753
|
+
`/tf init` starts with an **action menu**. First-time users get a 2-option shortcut ("Use recommended defaults" / "Configure each role"). Returning users see the full 5-option menu:
|
|
754
|
+
|
|
755
|
+
```
|
|
756
|
+
? What do you want to do with model roles?
|
|
757
|
+
❯ Use recommended defaults
|
|
758
|
+
Configure each role
|
|
759
|
+
Edit one role
|
|
760
|
+
Show current roles
|
|
761
|
+
Cancel
|
|
762
|
+
```
|
|
763
|
+
|
|
764
|
+
The picker shows model **display names** with capability flags and current/recommended markers:
|
|
765
|
+
|
|
766
|
+
```
|
|
767
|
+
? Model for 'vision' — Multimodal (executor-ui, visual-explorer)
|
|
768
|
+
Current: openrouter/anthropic/claude-sonnet-4-6
|
|
769
|
+
Recommended: minimax/MiniMax-M3
|
|
770
|
+
───────────────
|
|
771
|
+
❯ MiniMax M3 (minimax/MiniMax-M3) · image ✓ · reasoning ✓ · (recommended)
|
|
772
|
+
Claude Sonnet 4.6 (openrouter/anthropic/...) · image ✓ · reasoning ✓ · (current)
|
|
773
|
+
GPT-5 (openrouter/openai/gpt-5) · image ✓
|
|
774
|
+
DeepSeek V4 Flash (openrouter/deepseek/v4-flash)
|
|
775
|
+
───────────────
|
|
776
|
+
Custom (type your own)
|
|
777
|
+
Keep current
|
|
778
|
+
Back to action menu
|
|
779
|
+
```
|
|
780
|
+
|
|
781
|
+
Before saving, a **preview screen** shows the diff of your changes:
|
|
782
|
+
|
|
783
|
+
```
|
|
784
|
+
? Review changes:
|
|
785
|
+
fast openrouter/deepseek/deepseek-v4-flash (unchanged)
|
|
786
|
+
strong openrouter/xiaomi/mimo-v2.5-pro (unchanged)
|
|
787
|
+
thinker openrouter/qwen/qwen3.7-max (changed ← was: openrouter/deepseek/v4-pro)
|
|
788
|
+
arbiter openrouter/qwen/qwen3.7-max (unchanged)
|
|
789
|
+
vision minimax/MiniMax-M3 (unchanged)
|
|
790
|
+
reasoner z-ai/glm-5.1 (unchanged)
|
|
791
|
+
───────────────
|
|
792
|
+
❯ Save these changes
|
|
793
|
+
Edit a role
|
|
794
|
+
Cancel
|
|
795
|
+
```
|
|
796
|
+
|
|
797
|
+
Your choices are written to `~/.pi/agent/settings.json`:
|
|
798
|
+
|
|
799
|
+
```json
|
|
800
|
+
{
|
|
801
|
+
"modelRoles": {
|
|
802
|
+
"fast": "openrouter/deepseek/deepseek-v4-flash",
|
|
803
|
+
"strong": "openrouter/xiaomi/mimo-v2.5-pro",
|
|
804
|
+
"thinker": "openrouter/deepseek/deepseek-v4-pro",
|
|
805
|
+
"arbiter": "openrouter/qwen/qwen3.7-max",
|
|
806
|
+
"vision": "minimax/MiniMax-M3",
|
|
807
|
+
"reasoner": "z-ai/glm-5.1"
|
|
808
|
+
}
|
|
809
|
+
}
|
|
810
|
+
```
|
|
811
|
+
|
|
812
|
+
Edit the values manually any time, or just re-run `/tf init`.
|
|
813
|
+
|
|
814
|
+
To customize a specific agent's model or thinking without changing `modelRoles`, create an agent file at `~/.pi/agent/agents/<name>.md` with the desired overrides in the YAML frontmatter.
|
|
815
|
+
|
|
816
|
+
### Tool path (`action="init"`)
|
|
817
|
+
|
|
818
|
+
The model can also configure roles via the `taskflow` tool:
|
|
819
|
+
|
|
820
|
+
| Mode | Behavior |
|
|
821
|
+
|---|---|
|
|
822
|
+
| `mode: "show"` (default) | Read-only report of current `modelRoles`. Never overwrites. |
|
|
823
|
+
| `mode: "apply-defaults"` + `force: true` | Writes `RECOMMENDED_DEFAULTS` to `settings.json`, preserving stale keys. |
|
|
824
|
+
| `mode: "interactive"` | Launches the full action menu + picker flow (requires a UI session). |
|
|
825
|
+
|
|
826
|
+
|
|
827
|
+
### Custom agents
|
|
828
|
+
|
|
829
|
+
Drop a `.md` file into `~/.pi/agent/agents/` (user-level) or `.pi/agents/` (project-level, commit it) to add your own:
|
|
830
|
+
|
|
831
|
+
```markdown
|
|
832
|
+
---
|
|
833
|
+
name: my-linter
|
|
834
|
+
|
|
835
|
+
description: Run ESLint and report violations
|
|
836
|
+
|
|
837
|
+
tools: read, bash
|
|
838
|
+
|
|
839
|
+
model: "{{fast}}"
|
|
840
|
+
|
|
841
|
+
thinking: off
|
|
842
|
+
---
|
|
843
|
+
|
|
844
|
+
You are a linting agent. Run `npx eslint --format json` on the
|
|
845
|
+
provided files. Report violations grouped by file. No fixes.
|
|
846
|
+
```
|
|
847
|
+
|
|
848
|
+
Then reference it in any phase: `{ "agent": "my-linter", "task": "Lint src/" }`.
|
|
849
|
+
|
|
850
|
+
## Examples
|
|
851
|
+
|
|
852
|
+
Ready-to-read definitions in [`examples/`](https://github.com/heggria/taskflow/blob/main/examples):
|
|
853
|
+
|
|
854
|
+
| File | Demonstrates |
|
|
855
|
+
|---|---|
|
|
856
|
+
| [`summarize-files.json`](https://github.com/heggria/taskflow/blob/main/examples/summarize-files.json) | discover → `map` fan-out → `reduce` |
|
|
857
|
+
| [`conditional-research.json`](https://github.com/heggria/taskflow/blob/main/examples/conditional-research.json) | `when` routing + `join: any` + `gate` + `budget` |
|
|
858
|
+
| [`guarded-refactor.json`](https://github.com/heggria/taskflow/blob/main/examples/guarded-refactor.json) | `approval` (human-in-the-loop) + `retry` + `gate` |
|
|
859
|
+
| [`dynamic-plan-execute.json`](https://github.com/heggria/taskflow/blob/main/examples/dynamic-plan-execute.json) | `flow { def }` — plan then execute at runtime |
|
|
860
|
+
| [`iterative-replan.json`](https://github.com/heggria/taskflow/blob/main/examples/iterative-replan.json) | `loop` + `flow { def }` — iterative replanning |
|
|
861
|
+
|
|
862
|
+
Copy one into `.pi/taskflows/<name>.json` (or `~/.pi/agent/taskflows/`) and it registers as `/tf:<name>` — or just point the model at it.
|
|
863
|
+
|
|
864
|
+
## What's inside
|
|
865
|
+
|
|
866
|
+
<div align="center">
|
|
867
|
+
|
|
868
|
+
**0 runtime dependencies** · **1140 tests** · **10 phase types** · **shared context tree** · **cross-session resume** · **cross-run memoization** · **per-item map caching** · **incremental recompute** · **FlowIR compile seam** · **detached execution** · **`compile` Mermaid renderer** · **~9k LOC runtime**
|
|
869
|
+
|
|
870
|
+
</div>
|
|
871
|
+
|
|
872
|
+
- **Zero runtime dependencies.** No `dependencies` field — the runtime is built entirely on Node built-ins (`fs` / `path` / `os` / `child_process` / `crypto`). The file lock is `fs.openSync("wx")`, not a third-party library.
|
|
873
|
+
- **1140 tests across 70 test files** covering concurrency, atomic file locking (8-process race regressions), path-traversal hardening, cross-session resume, cross-run cache freshness (flow/thinking/tools key isolation, fingerprint invalidation, TTL/LRU eviction), backward-compatible cache-key migration (4-tier legacy fallback), per-phase structural sub-fingerprint (v3:phasefp — editing one phase invalidates only it and its dependents), per-item map caching (one changed item re-executes, N−1 cache hits), the `incremental` flag (run-wide cross-run default), reuse reporting, the FlowIR compile seam (determinism, declared-plane synthesis), incremental recompute (early-cutoff propagation, partial cascade strictly < full, observed ∪ declared union frontier), gate verdicts, budget caps, retry/backoff, approval flows, loop termination, tournament judging, sub-flow composition, the shared context tree (blackboard reuse, supervision spawn, subflow validation/nesting), workspace isolation (temp/dedicated/worktree lifecycle, fail-open degrade, dynamic-flow rejection), dynamic sub-flow security hardening, detached execution (PID persistence, stale detection, crash→failed, resume after failure), live run-history refresh, callback isolation, the idle watchdog, model-role init config, parseModelFromLabel with parenthesized-model-name regression, multi-fence `safeParse` recovery, host argv-contract locking (codex/claude/opencode `buildXxxArgs`), the `compile` Mermaid renderer (id-collision disambiguation, markdown-injection hardening, and full verify-overlay category coverage), plus the library Phase 1 metadata/search/store layer (phaseSignature, generality, CJK text scoring, staleness detection, sidecar persistence, A1 ghost-flow guard).
|
|
874
|
+
- **Hardened by design.** Path-traversal defense (lexical + `realpath` containment check), runId validation, HTML/error sanitization, atomic writes, stale-lock stealing via `rename`, and an idle watchdog that kills wedged subagents (SIGTERM → SIGKILL after 5 minutes of silence). Dynamic sub-flows additionally get breadth caps, `cwd` containment, budget clamping, nesting depth caps, and prototype-pollution defense.
|
|
875
|
+
- **Dogfooded.** Every new feature has to survive the project's own `self-improve` taskflow before it ships.
|
|
876
|
+
|
|
877
|
+
## 🍽️ We eat our own dog food
|
|
878
|
+
|
|
879
|
+
Every feature in `taskflow` ships **through `taskflow`.**
|
|
880
|
+
|
|
881
|
+
Our `self-improve` flow is a 10-phase DAG — it audits the codebase, patches defects, verifies correctness, gates on quality, and surfaces the report — all declaratively. We run it (as a user-scope `/tf:self-improve` flow) before releases. No other agent orchestrator in the Pi ecosystem builds itself with itself.
|
|
882
|
+
|
|
883
|
+
| Campaign | Scale | Phases | Outcome |
|
|
884
|
+
|----------|-------|--------|---------|
|
|
885
|
+
| [v0.0.8 dogfood](https://github.com/heggria/taskflow/blob/main/docs/internal/dogfooding-v0.0.8-report.md) | Full codebase audit → triage → fix → verify | 10 phases, 234 tests | 13 fixes, all pass |
|
|
886
|
+
| [v0.0.6 self-audit](https://github.com/heggria/taskflow/blob/main/docs/internal/self-audit-report.md) | inventory → map audit → gate → approval → map fix → reduce | 9 phases | 11 critical defects fixed |
|
|
887
|
+
| [Cross-run cache dogfood](https://github.com/heggria/taskflow/blob/main/docs/internal/rfc-cross-run-memoization.md) | Real runtime + on-disk store | Dedicated test harness | Cache correctness under adversarial fingerprints |
|
|
888
|
+
| [Adversarial cross-review](https://github.com/heggria/taskflow/blob/main/docs/internal/brainstorm-adversarial-review-report.md) | Multi-agent adversarial review | `tournament` + `gate` | P0 cache-key fix shipped |
|
|
889
|
+
| [Init redesign review](https://github.com/heggria/taskflow/blob/main/docs/internal/issue-necessity-review-report.md) | Necessity audit → parallel checks → verdict | 7 phases | Full redesign plan validated |
|
|
890
|
+
| [Round 2 adversarial audit](https://github.com/heggria/taskflow/blob/main/docs/internal/dogfooding-report.md) | Integration layer + cross-module — 12 findings across runner/runtime/interpolate/verify | 14 phases | 10 fixes applied, 0 regressions |
|
|
891
|
+
| [Round 3 adversarial audit](https://github.com/heggria/taskflow/blob/main/docs/internal/dogfooding-report.md) | Integration layer + cross-module — 10 findings across index/agents/cache/render/runs-view | 9 phases | 10 fixes applied, 0 regressions |
|
|
892
|
+
| [v0.0.23 Shared Context Tree](https://github.com/heggria/taskflow/blob/main/docs/internal/dogfooding-report.md) | End-to-end validation: org-tree spawn, 5-way audit via loop+gate | 6 e2e runs | Spawn-drain bug fixed, 50 new tests |
|
|
893
|
+
|
|
894
|
+
> **Meta:** we used `taskflow`'s `map` fan-out, `gate` verdicts, `approval` human-in-the-loop, `tournament` best-of-N, `loop` until-done, and `cross-run` cache — to build `taskflow`.
|
|
895
|
+
|
|
896
|
+
## Status & limits
|
|
897
|
+
|
|
898
|
+
**v0.1.6** (current release) — adds **library Phase 1** (search-before-author + reusable-flow sidecar metadata), the **`defineFile`** parameter (verify/compile/run a flow from a path on disk), and **JSONC comment support** in flow definition files (`//` and `/* */` comments + trailing commas, parsed by the new zero-dependency `parseJsonc`). **v0.1.5** added **Claude Code and OpenCode as hosts**, **extracted the MCP server into its own `taskflow-mcp-core` package**, and **de-duplicated the three host runners** into a shared `runSubagentProcess`. See [CHANGELOG](https://github.com/heggria/taskflow/blob/main/CHANGELOG.md) for the full history. Baseline: **multi-host monorepo of seven packages** — the host-neutral `taskflow-core` engine, the host-neutral `taskflow-mcp-core` MCP server, the shared host-runner `taskflow-hosts`, plus `pi-taskflow` (Pi adapter), `codex-taskflow`, `claude-taskflow`, and `opencode-taskflow` (the three delivery packages re-export their runners from `taskflow-hosts` and each ships an MCP bin + plugin/config), all sharing the host-neutral MCP server in `taskflow-mcp-core`. **Library Phase 1**: save flows with `purpose`+`tags` via `taskflow_save` (MCP) or `action=save` (Pi), search them with structural + CJK-aware keyword scoring via `taskflow_search`/`action=search`, and track `reuseCount` via `reusedFromSearch`. **`defineFile`**: pass a `defineFile` path (or `{defineFile, name}`) to `action=run` (Pi) or `taskflow_run`/`taskflow_verify`/`taskflow_compile` (MCP) instead of an inline `define`, and the engine reads the flow from disk — pair it with JSONC comments to annotate saved flows. **JSONC**: flow-definition `.json` files may now carry `//` and `/* */` comments and trailing commas (parsed by `parseJsonc`, re-exported from the `taskflow-core` barrel); LLM-output parsing via `safeParse` stays strict. **Shared Context Tree**: opt-in (`shareContext` / `contextSharing`) blackboard + supervision tools (`ctx_read`/`ctx_write` horizontal reuse, `ctx_report`/`ctx_spawn` vertical supervision); `ctx_spawn` accepts a flat task **or** a dependency-bearing `subflow` (a runtime-validated nested DAG), depth-capped on a unified nesting counter with budget accounting. **Workspace isolation**: a phase's `cwd` accepts reserved keywords `temp`/`dedicated`/`worktree` — the runtime allocates an isolated dir (or a git worktree on a throwaway branch) and tears it down after the phase, fail-open, rejected in LLM-authored sub-flows. **Detached execution**: runs can execute in the background, detached from the Pi session. Prior: loop-until-done (`loop`), tournament (best-of-N with a judge), cross-run memoization (content-addressed cache with git/file/glob/env fingerprints and TTL), interactive `/tf init`, configurable built-in agents, 18 built-in agents with 6 model roles. Full control-flow & reliability layer (`when` guards, `join: any`, `retry`/backoff, `approval`, `flow` composition, `budget` caps, `onBlock: "retry"`, `eval` machine gates, idle watchdog) on top of the DSL + DAG runtime (`agent`/`parallel`/`map`/`gate`/`reduce`). Inline + saved flows, cross-session resume, live progress, and isolated context. A run executes as one streaming tool call.
|
|
899
|
+
|
|
900
|
+
Known boundaries (tracked, bounded — no surprises mid-flow):
|
|
901
|
+
|
|
902
|
+
- **Shared context is opt-in.** Subagents share nothing unless a phase sets `shareContext` (or the flow sets `contextSharing`). The blackboard is per-run, file-based, size-bounded, and cleaned up with the run. Spawn nesting is capped at `MAX_DYNAMIC_NESTING` (5). A spawned flat task is not individually checkpointed — on crash it re-runs on resume (spawned *subflows* resume their completed inner phases via the cache).
|
|
903
|
+
- **Workspace isolation is fail-open.** `cwd: "worktree"` requires the base cwd to be a git work tree; otherwise it degrades to a `temp` dir (with a warning). `temp`/`worktree` dirs are removed when the phase ends — a hard crash mid-phase may leave a stray dir (cleaned on the next run for `dedicated`; `temp`/`worktree` are under the OS tmpdir). The reserved keywords are honoured only in author-written flows.
|
|
904
|
+
- **No `output: "file"`.** Outputs are text/JSON only — write files via an agent's `write` tool call.
|
|
905
|
+
- **`map` fans out over a JSON array from a string `over`.** The `over` field is a string that either interpolates to a JSON array (e.g. `{steps.ID.json}`) or is a literal JSON-array string. Wrap a plain text list in a single-agent `output: "json"` phase first, or pass `JSON.stringify([...])` for a fixed list. (A raw literal array is rejected — emit it from a phase and reference that.)
|
|
906
|
+
- **The DAG must be acyclic.** Cycles are rejected at validation.
|
|
907
|
+
- **Cross-run cache excludes `gate`, `approval`, `loop`, `tournament`, and `script`.** These must produce a fresh result each run (a `script` phase may also have side effects).
|
|
908
|
+
- **Approval auto-rejects in detached mode.** This is a safety invariant — approval gates are never silently bypassed.
|
|
909
|
+
|
|
910
|
+
## Development
|
|
911
|
+
|
|
912
|
+
`taskflow` is a pnpm-workspace monorepo of seven published packages:
|
|
913
|
+
|
|
914
|
+
| Package | Role |
|
|
915
|
+
|---------|------|
|
|
916
|
+
| [`taskflow-core`](https://github.com/heggria/taskflow/blob/main/packages/taskflow-core) | Host-neutral orchestration engine (zero host-SDK deps; only `typebox`) — runtime, DSL, cache, verify |
|
|
917
|
+
| [`taskflow-mcp-core`](https://github.com/heggria/taskflow/blob/main/packages/taskflow-mcp-core) | Host-neutral MCP server (stdio JSON-RPC + `taskflow_*` tools + DAG renderer); depends on core |
|
|
918
|
+
| [`taskflow-hosts`](https://github.com/heggria/taskflow/blob/main/packages/taskflow-hosts) | Shared host-runner collection — the codex/claude/opencode `SubagentRunner` impls + their argv builders + event-stream parsers; depends on core |
|
|
919
|
+
| [`pi-taskflow`](https://github.com/heggria/taskflow/blob/main/packages/pi-taskflow) | Pi extension adapter — `taskflow` tool + `/tf` commands (what `pi install npm:pi-taskflow` gives you) |
|
|
920
|
+
| [`codex-taskflow`](https://github.com/heggria/taskflow/blob/main/packages/codex-taskflow) | Codex MCP server + bin + [Codex plugin](https://github.com/heggria/taskflow/blob/main/packages/codex-taskflow/plugin) (re-exports the runner from `taskflow-hosts`) ([guide](https://github.com/heggria/taskflow/blob/main/docs/codex-mcp.md)) |
|
|
921
|
+
| [`claude-taskflow`](https://github.com/heggria/taskflow/blob/main/packages/claude-taskflow) | Claude Code MCP server + bin + [Claude Code plugin](https://github.com/heggria/taskflow/blob/main/packages/claude-taskflow/plugin) (re-exports the runner from `taskflow-hosts`) ([guide](https://github.com/heggria/taskflow/blob/main/docs/claude-mcp.md)) |
|
|
922
|
+
| [`opencode-taskflow`](https://github.com/heggria/taskflow/blob/main/packages/opencode-taskflow) | OpenCode MCP server + bin + [OpenCode config scaffold](https://github.com/heggria/taskflow/blob/main/packages/opencode-taskflow/plugin) (re-exports the runner from `taskflow-hosts`) ([guide](https://github.com/heggria/taskflow/blob/main/docs/opencode-mcp.md)) |
|
|
923
|
+
|
|
924
|
+
```bash
|
|
925
|
+
pnpm install
|
|
926
|
+
pnpm run typecheck # tsc --noEmit across all packages (no build needed)
|
|
927
|
+
pnpm test # unit tests — no network, no process spawning
|
|
928
|
+
pnpm run test:hosts # host-runner tests only (also: test:pi, test:codex, test:claude, test:opencode)
|
|
929
|
+
pnpm run build # emit dist/*.js + .d.ts for all seven packages
|
|
930
|
+
pnpm run test:e2e-codex # codex executor e2e (needs `codex` + model access)
|
|
931
|
+
pnpm run test:e2e-codex-mcp # codex MCP server e2e
|
|
932
|
+
```
|
|
933
|
+
|
|
934
|
+
The pi end-to-end suites spawn live `pi` subagents and are run directly (they use
|
|
935
|
+
the `.mts` extension so the unit-test glob skips them), e.g.:
|
|
936
|
+
|
|
937
|
+
```bash
|
|
938
|
+
node --conditions=development --experimental-strip-types packages/pi-taskflow/test/e2e.mts
|
|
939
|
+
# others: e2e-team, e2e-context, e2e-context-value, e2e-spawn-subflow,
|
|
940
|
+
# e2e-flowir, e2e-incremental-suite, dogfood-cache
|
|
941
|
+
```
|
|
942
|
+
|
|
943
|
+
Engine code lives in `packages/taskflow-core/src/`, the Pi adapter in `packages/pi-taskflow/src/`, tests in each package's `test/`, and runnable examples in `examples/`. Published packages ship compiled `dist/`; dev resolves the TypeScript sources directly via a `development` export condition — no build step needed to typecheck or test.
|
|
944
|
+
|
|
945
|
+
## Contributing
|
|
946
|
+
|
|
947
|
+
Contributions welcome — this is a young, fast-moving project. Open an issue or PR on [GitHub](https://github.com/heggria/taskflow). Good first contributions: new example flows, phase-type ideas, and TUI polish. See [`CONTRIBUTING.md`](https://github.com/heggria/taskflow/blob/main/CONTRIBUTING.md) and [`AGENTS.md`](https://github.com/heggria/taskflow/blob/main/AGENTS.md) for coding conventions and common task recipes.
|
|
948
|
+
|
|
949
|
+
## License
|
|
950
|
+
|
|
951
|
+
MIT
|