hermes-taskflow 0.2.9
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +447 -0
- package/dist/index.d.ts +13 -0
- package/dist/index.d.ts.map +1 -0
- package/dist/index.js +13 -0
- package/dist/index.js.map +1 -0
- package/dist/mcp/bin.d.ts +31 -0
- package/dist/mcp/bin.d.ts.map +1 -0
- package/dist/mcp/bin.js +39 -0
- package/dist/mcp/bin.js.map +1 -0
- package/dist/mcp/server.d.ts +16 -0
- package/dist/mcp/server.d.ts.map +1 -0
- package/dist/mcp/server.js +30 -0
- package/dist/mcp/server.js.map +1 -0
- package/package.json +60 -0
- package/plugin/assets/taskflow-small.svg +14 -0
- package/plugin/assets/taskflow.svg +17 -0
- package/plugin/hermes.config.snippet.yaml +27 -0
- package/plugin/skills/taskflow/SKILL.md +692 -0
- package/plugin/skills/taskflow/advanced.md +267 -0
- package/plugin/skills/taskflow/configuration.md +598 -0
- package/plugin/skills/taskflow/library.md +105 -0
- package/plugin/skills/taskflow/patterns.md +348 -0
|
@@ -0,0 +1,105 @@
|
|
|
1
|
+
<!-- GENERATED FILE — do not edit. Source: skills-src/taskflow/library.md (npm run build:skills) -->
|
|
2
|
+
|
|
3
|
+
# Library: reusable flows & the search-before-author loop
|
|
4
|
+
|
|
5
|
+
taskflow saves flows you'll reuse into a **library** with metadata, so you can
|
|
6
|
+
`search` it before authoring a new flow — the **reuse flywheel**. The more you
|
|
7
|
+
save (with a good `purpose` + `tags`), the better future search gets.
|
|
8
|
+
|
|
9
|
+
> Full design: `docs/rfc-library-reuse.md`. This file is the agent-facing
|
|
10
|
+
> "when and how" guide.
|
|
11
|
+
|
|
12
|
+
**Host binding:** use the `taskflow_search`, `taskflow_save`, `taskflow_list`,
|
|
13
|
+
`taskflow_show` tools.
|
|
14
|
+
|
|
15
|
+
## Before authoring a non-trivial flow: SEARCH first
|
|
16
|
+
|
|
17
|
+
Any time you're about to write a flow with ≥3 phases, fan-out, or a gate,
|
|
18
|
+
**search the library first**. It costs nothing and often finds a starter you
|
|
19
|
+
can adapt.
|
|
20
|
+
|
|
21
|
+
```jsonc
|
|
22
|
+
{ "name": "taskflow_search", "arguments": { "query": "audit API endpoints for missing auth", "limit": 5 } }
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
Read the results and the `→ reuseHint`:
|
|
26
|
+
|
|
27
|
+
| score | reuseHint | what to do |
|
|
28
|
+
|-------|-----------|------------|
|
|
29
|
+
| **≥ 0.8** | "直接复用" / direct reuse | Run by name (skip authoring). |
|
|
30
|
+
| **0.5 – 0.8** | "copy + 泛化" / copy & generalize | `show` it, copy, **generalize** (see checklist), save as a new version. |
|
|
31
|
+
| **< 0.5** or no matches | "从头编写" / write fresh | Author from scratch — then **save** it if reusable. |
|
|
32
|
+
|
|
33
|
+
`searchMode` tells you how the ranking was produced: `structural` (keyword +
|
|
34
|
+
phase-signature; no embedding backend configured), `semantic` (cosine over
|
|
35
|
+
embeddings), or `mixed` (some flows had vectors, some didn't). `structural` is
|
|
36
|
+
weaker on paraphrase — if it missed something obvious, try `structureOnly:
|
|
37
|
+
false` or a different phrasing.
|
|
38
|
+
|
|
39
|
+
## After a successful novel flow: SAVE it (if reusable)
|
|
40
|
+
|
|
41
|
+
When you finish a flow you expect to use again, save it **with a `purpose` and
|
|
42
|
+
2–4 `tags`**. These two fields are what search matches on — a flow saved
|
|
43
|
+
without them is nearly invisible to future search.
|
|
44
|
+
|
|
45
|
+
```jsonc
|
|
46
|
+
{ "name": "taskflow_save",
|
|
47
|
+
"arguments": { "name": "audit-endpoints", "definition": { "phases": [ ... ] },
|
|
48
|
+
"purpose": "Audit a directory of API endpoints for missing auth checks",
|
|
49
|
+
"tags": ["audit", "security", "auth", "fan-out"] } }
|
|
50
|
+
```
|
|
51
|
+
|
|
52
|
+
`save` auto-derives structural metadata (phase signature, a `generality` score
|
|
53
|
+
in 0–1) and writes a sidecar `.meta.json` next to the flow file. You don't
|
|
54
|
+
compute any of that — just give `purpose` + `tags`.
|
|
55
|
+
|
|
56
|
+
## The generalization checklist (apply on every reuse)
|
|
57
|
+
|
|
58
|
+
When you copy + generalize a flow, make it **more reusable than the version you
|
|
59
|
+
found**. Each item raises the auto-derived `generality` score and broadens
|
|
60
|
+
future search recall:
|
|
61
|
+
|
|
62
|
+
- [ ] Hardcoded file/dir paths → `{args.X}` (with a `default`).
|
|
63
|
+
- [ ] Specific entity words ("endpoint", "route") → broaden in the discover
|
|
64
|
+
prompt so the flow works for the whole class.
|
|
65
|
+
- [ ] Thresholds / counts → `{args.X}` with sensible defaults.
|
|
66
|
+
- [ ] Add `budget` / `retry` / `expect` if missing (production-grade knobs).
|
|
67
|
+
- [ ] Update `purpose` to reflect the wider scope.
|
|
68
|
+
|
|
69
|
+
Then save it back (version auto-bumps). Over time the library compounds: every
|
|
70
|
+
reuse leaves a more general flow behind.
|
|
71
|
+
|
|
72
|
+
## reuseCount & `reusedFromSearch`
|
|
73
|
+
|
|
74
|
+
Each saved flow has a `reuseCount`. It goes up by 1 **only when** a run was
|
|
75
|
+
chosen because of a prior search — set the `reusedFromSearch: true` flag on the
|
|
76
|
+
run. Direct run-by-name does **not** bump it (that's intentional: `reuseCount`
|
|
77
|
+
measures "found-via-search reuse", the high-quality signal for later auto-prune).
|
|
78
|
+
|
|
79
|
+
```jsonc
|
|
80
|
+
{ "name": "taskflow_run",
|
|
81
|
+
"arguments": { "name": "audit-endpoints", "args": { "dir": "src/api" }, "reusedFromSearch": true } }
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
## Judicious reuse — not every task needs the library
|
|
85
|
+
|
|
86
|
+
Skip search + save for:
|
|
87
|
+
- **One-off tasks** (a quick fix, a throwaway analysis) — `generality < 0.3`
|
|
88
|
+
and obviously won't recur.
|
|
89
|
+
- **Trivial flows** (1–2 phases, no fan-out) — overhead isn't worth it.
|
|
90
|
+
|
|
91
|
+
The library pays off for *patterns* that recur across projects or sessions:
|
|
92
|
+
audits, migrations, reviews, fan-out summarization, plan→approve→execute, etc.
|
|
93
|
+
|
|
94
|
+
## Configuration (embedding backend — optional, Phase 2)
|
|
95
|
+
|
|
96
|
+
Search works with **zero config** (structural mode). For smarter paraphrase
|
|
97
|
+
recall, configure an embedding backend in `~/.pi/agent/settings.json`:
|
|
98
|
+
|
|
99
|
+
```jsonc
|
|
100
|
+
{ "taskflow": { "library": { "enabled": true, "scope": "both" },
|
|
101
|
+
"embedder": { "kind": "http", "url": "http://127.0.0.1:8123/v1/embeddings", "model": "qwen3-embedding-0.6b" } } }
|
|
102
|
+
```
|
|
103
|
+
|
|
104
|
+
Without `embedder`, or if the embedder fails, search **degrades gracefully** to
|
|
105
|
+
structural mode — it never breaks. (See `docs/rfc-library-reuse.md` §4.)
|
|
@@ -0,0 +1,348 @@
|
|
|
1
|
+
<!-- GENERATED FILE — do not edit. Source: skills-src/taskflow/patterns.md (npm run build:skills) -->
|
|
2
|
+
|
|
3
|
+
# Taskflow Patterns — proven archetypes & the production checklist
|
|
4
|
+
|
|
5
|
+
Read this when designing a flow with ≥ 4 phases, a gate, or any fan-out.
|
|
6
|
+
Each archetype below is a complete, runnable shape distilled from real runs.
|
|
7
|
+
Copy the closest one and adapt — don't design from a blank page.
|
|
8
|
+
|
|
9
|
+
---
|
|
10
|
+
|
|
11
|
+
## The production-flow checklist
|
|
12
|
+
|
|
13
|
+
Before running a flow you designed, check it against this list. Every item is
|
|
14
|
+
cheap to add and each one has prevented a real failure class:
|
|
15
|
+
|
|
16
|
+
- [ ] **`verify` first.** Run the static verifier (pi: `action: "verify"` /
|
|
17
|
+
Codex / Claude Code / OpenCode: `taskflow_verify`) — zero tokens,
|
|
18
|
+
catches cycles / missing deps / ref typos / contract mismatches.
|
|
19
|
+
- [ ] **Every JSON-emitting phase has `expect` + `retry`.** "Output ONLY JSON"
|
|
20
|
+
in the task text is a request; `expect` is enforcement. Without it, a
|
|
21
|
+
malformed router output silently skips both branches.
|
|
22
|
+
- [ ] **Machine checks before LLM checks.** `script` phases for build/test
|
|
23
|
+
ground truth; gate `eval` assertions before the gate's LLM `task`. A
|
|
24
|
+
token spent verifying what a shell command can verify is a wasted token.
|
|
25
|
+
- [ ] **`budget` on any flow with a fan-out.** A map over a mis-discovered
|
|
26
|
+
500-item array is unbounded spend without one.
|
|
27
|
+
- [ ] **`optional: true` + a fallback for degradable phases.** An enrichment
|
|
28
|
+
phase that times out shouldn't sink the run — pair `timeout` +
|
|
29
|
+
`optional` with a downstream `when`-guarded fallback.
|
|
30
|
+
- [ ] **Exactly one `final: true`** on the result-bearing phase.
|
|
31
|
+
- [ ] **`strictInterpolation: true` on any flow you save.** Saved flows run
|
|
32
|
+
later with args you're not watching; unresolved placeholders must be
|
|
33
|
+
errors, not empty strings.
|
|
34
|
+
- [ ] **Discovery phases are cheap and tool-restricted.** `agent: "scout"`,
|
|
35
|
+
`tools: ["read","grep","ls"]`, low thinking. Filesystem isolation still
|
|
36
|
+
depends on enforcement by the selected host; don't pay executor prices for `ls`.
|
|
37
|
+
- [ ] **The reviewer is not the producer.** A gate reviewing phase X uses a
|
|
38
|
+
different agent (ideally a different model) than X. Self-review passes
|
|
39
|
+
~everything.
|
|
40
|
+
- [ ] **Re-runnable flows are `incremental: true`** with `fingerprint` entries
|
|
41
|
+
on phases that read the world (`git:HEAD`, `glob!:src/**/*.ts`). See
|
|
42
|
+
`advanced.md`.
|
|
43
|
+
|
|
44
|
+
---
|
|
45
|
+
|
|
46
|
+
## Archetype 1: Audit fan-out (discover → map → gate → reduce)
|
|
47
|
+
|
|
48
|
+
The workhorse. Use for: audit every endpoint / migrate every file /
|
|
49
|
+
summarize every module.
|
|
50
|
+
|
|
51
|
+
```jsonc
|
|
52
|
+
{
|
|
53
|
+
"name": "audit-endpoints",
|
|
54
|
+
"strictInterpolation": true,
|
|
55
|
+
"budget": { "maxUSD": 3.00 },
|
|
56
|
+
"phases": [
|
|
57
|
+
{ "id": "discover", "type": "agent", "agent": "scout",
|
|
58
|
+
"tools": ["read", "grep", "ls"],
|
|
59
|
+
"task": "List every HTTP endpoint under src/routes. Output ONLY a JSON array [{\"route\":\"...\",\"file\":\"...\"}]. No prose.",
|
|
60
|
+
"output": "json",
|
|
61
|
+
"expect": { "type": "array", "items": { "type": "object", "required": ["route", "file"] } },
|
|
62
|
+
"retry": { "max": 2, "backoffMs": 0 } },
|
|
63
|
+
|
|
64
|
+
{ "id": "audit", "type": "map", "over": "{steps.discover.json}", "as": "item",
|
|
65
|
+
"agent": "analyst", "concurrency": 4,
|
|
66
|
+
"task": "Audit {item.route} in {item.file} for missing auth. Report: SEVERITY (high/med/low/none), evidence (file:line), fix.",
|
|
67
|
+
"dependsOn": ["discover"] },
|
|
68
|
+
|
|
69
|
+
{ "id": "screen", "type": "gate", "agent": "reviewer",
|
|
70
|
+
"task": "Cross-check the findings below. Delete false positives (cite why). If ANY confirmed HIGH remains, end with VERDICT: BLOCK and list them; else VERDICT: PASS.\n\n{steps.audit.output}",
|
|
71
|
+
"dependsOn": ["audit"] },
|
|
72
|
+
|
|
73
|
+
{ "id": "report", "type": "reduce", "from": ["screen"], "agent": "doc-writer",
|
|
74
|
+
"task": "Write a prioritized remediation report from:\n{steps.screen.output}",
|
|
75
|
+
"dependsOn": ["screen"], "final": true }
|
|
76
|
+
]
|
|
77
|
+
}
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
Why each piece: `expect`+`retry` on discover means a chatty scout gets a second
|
|
81
|
+
chance instead of feeding garbage to the map; `concurrency: 4` protects rate
|
|
82
|
+
limits; the gate is a *different* agent than the auditor; `budget` limits the
|
|
83
|
+
fan-out.
|
|
84
|
+
|
|
85
|
+
**Variant — per-item caching for repeated audits:** add
|
|
86
|
+
`"cache": { "scope": "cross-run" }` to the map phase. On the next run, only
|
|
87
|
+
items whose task text changed re-execute; the rest are $0 cache hits.
|
|
88
|
+
(Details: `configuration.md` §8.)
|
|
89
|
+
|
|
90
|
+
---
|
|
91
|
+
|
|
92
|
+
## Archetype 2: Self-healing implement→verify→rework
|
|
93
|
+
|
|
94
|
+
Use for: implement against acceptance criteria, fix-until-green.
|
|
95
|
+
The gate re-runs its upstream on BLOCK — a generate→critique→regenerate loop
|
|
96
|
+
without you writing a loop.
|
|
97
|
+
|
|
98
|
+
```jsonc
|
|
99
|
+
{
|
|
100
|
+
"name": "implement-verified",
|
|
101
|
+
"budget": { "maxUSD": 5.00 },
|
|
102
|
+
"phases": [
|
|
103
|
+
{ "id": "implement", "type": "agent", "agent": "executor-code",
|
|
104
|
+
"task": "Implement the feature per the spec in docs/spec.md. Run nothing; just edit." },
|
|
105
|
+
|
|
106
|
+
{ "id": "build-test", "type": "script",
|
|
107
|
+
"run": "npx tsc --noEmit && pnpm test 2>&1 | tail -20",
|
|
108
|
+
"timeout": 180000, "dependsOn": ["implement"] },
|
|
109
|
+
|
|
110
|
+
{ "id": "spec-gate", "type": "gate", "agent": "reviewer",
|
|
111
|
+
"onBlock": "retry", "retry": { "max": 3 },
|
|
112
|
+
"eval": ["{steps.build-test.output} contains pass"],
|
|
113
|
+
"task": "Build/test output:\n{steps.build-test.output}\n\nDoes the implementation satisfy ALL acceptance criteria in docs/spec.md? VERDICT: PASS, or VERDICT: BLOCK with a precise list of what to fix.",
|
|
114
|
+
"dependsOn": ["implement", "build-test"] },
|
|
115
|
+
|
|
116
|
+
{ "id": "summary", "type": "agent", "agent": "doc-writer",
|
|
117
|
+
"task": "Summarize what was implemented and the verification result:\n{steps.spec-gate.output}",
|
|
118
|
+
"dependsOn": ["spec-gate"], "final": true }
|
|
119
|
+
]
|
|
120
|
+
}
|
|
121
|
+
```
|
|
122
|
+
|
|
123
|
+
Key mechanics: on BLOCK, `spec-gate` re-runs **both** its `dependsOn` upstreams
|
|
124
|
+
(`implement` gets the blocker's reasons via re-interpolation, `build-test`
|
|
125
|
+
re-verifies), up to 3 rounds. The `eval` line means a green build+test skips
|
|
126
|
+
the LLM review entirely on the happy path.
|
|
127
|
+
|
|
128
|
+
**Verification phases: force structured output.** LLMs are bad at
|
|
129
|
+
*summarizing* shell output (234 tests read as 230) but good at *copying*
|
|
130
|
+
structured data. If a verification step must go through an agent (not a
|
|
131
|
+
`script`), demand `key=value` lines:
|
|
132
|
+
|
|
133
|
+
```
|
|
134
|
+
Report EXACTLY in this format (one key=value per line, no prose):
|
|
135
|
+
typecheck=PASS|FAIL
|
|
136
|
+
tests_total=N
|
|
137
|
+
tests_fail=N
|
|
138
|
+
If any field is missing, you failed the task — re-run and re-read.
|
|
139
|
+
```
|
|
140
|
+
|
|
141
|
+
Prefer a `script` phase whenever the check is a command — exact, free, fast.
|
|
142
|
+
|
|
143
|
+
---
|
|
144
|
+
|
|
145
|
+
## Archetype 3: Plan → human approval → execute
|
|
146
|
+
|
|
147
|
+
Use for: anything expensive or destructive where a human should see the plan
|
|
148
|
+
before the spend. The approval's **Edit** option injects mid-run guidance.
|
|
149
|
+
|
|
150
|
+
> **MCP-host caveat (Codex / Claude Code / OpenCode / Grok / Hermes):** approval phases auto-reject in
|
|
151
|
+
> MCP-driven (non-interactive) runs. This archetype only works when a human
|
|
152
|
+
> runs the flow interactively; for tool-driven runs, replace the approval with
|
|
153
|
+
> a strict `gate`.
|
|
154
|
+
|
|
155
|
+
```jsonc
|
|
156
|
+
{
|
|
157
|
+
"name": "guarded-migration",
|
|
158
|
+
"phases": [
|
|
159
|
+
{ "id": "plan", "type": "agent", "agent": "planner",
|
|
160
|
+
"task": "Plan the migration of src/legacy/* to the new API. List each file, the change, and the risk." },
|
|
161
|
+
|
|
162
|
+
{ "id": "checkpoint", "type": "approval",
|
|
163
|
+
"task": "Migration plan:\n\n{steps.plan.output}\n\nApprove to execute, reject to abort, or edit to add constraints.",
|
|
164
|
+
"dependsOn": ["plan"] },
|
|
165
|
+
|
|
166
|
+
{ "id": "execute", "type": "agent", "agent": "executor-code",
|
|
167
|
+
"task": "Execute the migration plan:\n{steps.plan.output}\n\nOperator guidance (if any): {steps.checkpoint.output}",
|
|
168
|
+
"dependsOn": ["checkpoint"] },
|
|
169
|
+
|
|
170
|
+
{ "id": "verify", "type": "script", "run": "npx tsc --noEmit && pnpm test 2>&1 | tail -5",
|
|
171
|
+
"timeout": 180000, "dependsOn": ["execute"], "final": true }
|
|
172
|
+
]
|
|
173
|
+
}
|
|
174
|
+
```
|
|
175
|
+
|
|
176
|
+
Note `{steps.checkpoint.output}` — on Edit it carries the operator's note; on
|
|
177
|
+
Approve it's `(approve)`. Don't use approval in detached/headless runs (it
|
|
178
|
+
auto-rejects there — by design).
|
|
179
|
+
|
|
180
|
+
---
|
|
181
|
+
|
|
182
|
+
## Archetype 4: Dynamic plan → execute (`flow{def}`)
|
|
183
|
+
|
|
184
|
+
Use when the *work itself* must be discovered at runtime — the planner emits a
|
|
185
|
+
whole sub-flow as JSON and the runtime validates + runs it. The declarative
|
|
186
|
+
answer to "loop over whatever we find".
|
|
187
|
+
|
|
188
|
+
```jsonc
|
|
189
|
+
{
|
|
190
|
+
"name": "dynamic-audit",
|
|
191
|
+
"budget": { "maxUSD": 4.00 },
|
|
192
|
+
"phases": [
|
|
193
|
+
{ "id": "plan", "type": "agent", "agent": "planner", "output": "json",
|
|
194
|
+
"task": "Scan this repo. Output ONLY a JSON taskflow {\"name\":\"sub\",\"phases\":[...]} with one 'agent' phase per module that needs auditing (agent: \"analyst\"), plus a final 'reduce' phase (agent: \"doc-writer\", from: [all audit ids], final: true). Use hyphens in ids. No script phases, no cwd fields.",
|
|
195
|
+
"expect": { "type": "object", "required": ["name", "phases"] },
|
|
196
|
+
"retry": { "max": 2, "backoffMs": 0 } },
|
|
197
|
+
|
|
198
|
+
{ "id": "run-plan", "type": "flow", "def": "{steps.plan.json}",
|
|
199
|
+
"optional": true, "dependsOn": ["plan"] },
|
|
200
|
+
|
|
201
|
+
{ "id": "deliver", "type": "agent", "agent": "doc-writer",
|
|
202
|
+
"task": "Final result (empty means the plan failed validation — say so):\n{steps.run-plan.output}",
|
|
203
|
+
"dependsOn": ["run-plan"], "final": true }
|
|
204
|
+
]
|
|
205
|
+
}
|
|
206
|
+
```
|
|
207
|
+
|
|
208
|
+
Critical details (full contract in `advanced.md`): a bad plan **fails open**
|
|
209
|
+
(`defError` diagnostic, empty output downstream) — `optional: true` + the
|
|
210
|
+
`deliver` phase turn that into a graceful report instead of a dead run. The
|
|
211
|
+
planner prompt must forbid what validation will reject anyway (`script`
|
|
212
|
+
phases, workspace `cwd` keywords) so the plan doesn't waste a retry.
|
|
213
|
+
|
|
214
|
+
**Iterative replanning:** wrap a plan-emitting body in a `loop` so round N's
|
|
215
|
+
plan reacts to round N−1's result — see `examples/iterative-replan.json`.
|
|
216
|
+
|
|
217
|
+
---
|
|
218
|
+
|
|
219
|
+
## Archetype 5: Tournament for one-shot-unreliable work
|
|
220
|
+
|
|
221
|
+
Use when quality varies run-to-run (naming, copywriting, tricky refactor
|
|
222
|
+
strategy, root-cause hypotheses). Branches > variants when you want genuinely
|
|
223
|
+
different *approaches* judged against each other.
|
|
224
|
+
|
|
225
|
+
```jsonc
|
|
226
|
+
{
|
|
227
|
+
"id": "strategy", "type": "tournament", "mode": "best",
|
|
228
|
+
"judgeAgent": "final-arbiter",
|
|
229
|
+
"judge": "Judge on: correctness under concurrent access, blast radius, migration cost. Quote evidence. Return JSON {\"winner\": <n>, \"reason\": \"...\"}.",
|
|
230
|
+
"branches": [
|
|
231
|
+
{ "task": "Design the cache-invalidation fix with a conservative approach: minimal diff, no schema change.", "agent": "analyst" },
|
|
232
|
+
{ "task": "Design the fix assuming we can change the schema: optimal correctness.", "agent": "analyst" },
|
|
233
|
+
{ "task": "Design the fix as an adversary: what will break each obvious approach? Then propose the one that survives.", "agent": "critic" }
|
|
234
|
+
],
|
|
235
|
+
"dependsOn": ["context"], "final": true
|
|
236
|
+
}
|
|
237
|
+
```
|
|
238
|
+
|
|
239
|
+
Give the judge a **rubric with named criteria**, a stronger model
|
|
240
|
+
(`judgeAgent`), and a structured winner output (`{"winner": <n>}` JSON,
|
|
241
|
+
or an exact `WINNER: <n>` terminator). `mode: "aggregate"`
|
|
242
|
+
instead merges all variants — good for research synthesis, bad for decisions.
|
|
243
|
+
|
|
244
|
+
---
|
|
245
|
+
|
|
246
|
+
## Archetype 7: Race for latency (first approach wins)
|
|
247
|
+
|
|
248
|
+
When several strategies can answer the same question and you care about **time
|
|
249
|
+
to first good answer** more than comparing quality, use `race` (not
|
|
250
|
+
`tournament`). Branches start together; the first successful completion becomes
|
|
251
|
+
the phase output.
|
|
252
|
+
|
|
253
|
+
```jsonc
|
|
254
|
+
{
|
|
255
|
+
"name": "quick-answer",
|
|
256
|
+
"budget": { "maxUSD": 0.5 },
|
|
257
|
+
"phases": [
|
|
258
|
+
{
|
|
259
|
+
"id": "answer", "type": "race",
|
|
260
|
+
"branches": [
|
|
261
|
+
{ "task": "Answer from local heuristics only: {args.q}", "agent": "executor" },
|
|
262
|
+
{ "task": "Answer after a short web/docs look: {args.q}", "agent": "researcher" }
|
|
263
|
+
],
|
|
264
|
+
"final": true
|
|
265
|
+
}
|
|
266
|
+
]
|
|
267
|
+
}
|
|
268
|
+
```
|
|
269
|
+
|
|
270
|
+
**Prefer tournament when** you need a judge to pick the *best* draft after all
|
|
271
|
+
variants finish. **Prefer parallel when** you need *every* branch's output
|
|
272
|
+
downstream (then `reduce`).
|
|
273
|
+
|
|
274
|
+
## Archetype 8: Expand graft (planner fragment on the parent DAG)
|
|
275
|
+
|
|
276
|
+
Planner emits a fragment; `expand` with `expandMode: "graft"` runs it and
|
|
277
|
+
promotes child phase states as `<expandId>-<childId>` so a later phase can read
|
|
278
|
+
them. Use `nested` (or classic `flow{def}`) when you only need the fragment's
|
|
279
|
+
**final** output and do not want child ids on the parent.
|
|
280
|
+
|
|
281
|
+
```jsonc
|
|
282
|
+
{
|
|
283
|
+
"name": "plan-graft",
|
|
284
|
+
"phases": [
|
|
285
|
+
{
|
|
286
|
+
"id": "plan", "type": "agent", "agent": "planner", "output": "json",
|
|
287
|
+
"task": "Emit a mini-flow JSON: {name, phases:[{id,type,agent,task,final?}…]} for the audit.",
|
|
288
|
+
"expect": { "type": "object", "required": ["phases"] }
|
|
289
|
+
},
|
|
290
|
+
{
|
|
291
|
+
"id": "grow", "type": "expand", "expandMode": "graft",
|
|
292
|
+
"def": "{steps.plan.json}", "dependsOn": ["plan"]
|
|
293
|
+
},
|
|
294
|
+
{
|
|
295
|
+
"id": "wrap", "type": "agent", "agent": "writer",
|
|
296
|
+
"task": "Summarize grafted work. Child outputs may appear as steps.grow-* in the run state.",
|
|
297
|
+
"dependsOn": ["grow"], "final": true
|
|
298
|
+
}
|
|
299
|
+
]
|
|
300
|
+
}
|
|
301
|
+
```
|
|
302
|
+
|
|
303
|
+
## Archetype 6: Incremental repo-watch audit (cross-run)
|
|
304
|
+
|
|
305
|
+
Use for flows you'll re-run as the repo evolves. First run pays full price;
|
|
306
|
+
subsequent runs re-pay only for what changed.
|
|
307
|
+
|
|
308
|
+
```jsonc
|
|
309
|
+
{
|
|
310
|
+
"name": "security-sweep",
|
|
311
|
+
"incremental": true,
|
|
312
|
+
"budget": { "maxUSD": 3.00 },
|
|
313
|
+
"phases": [
|
|
314
|
+
{ "id": "discover", "type": "agent", "agent": "scout", "output": "json",
|
|
315
|
+
"task": "List all files handling user input. Output ONLY a JSON array of paths.",
|
|
316
|
+
"expect": { "type": "array" },
|
|
317
|
+
"cache": { "scope": "cross-run", "fingerprint": ["glob!:src/**/*.ts"] } },
|
|
318
|
+
{ "id": "audit", "type": "map", "over": "{steps.discover.json}",
|
|
319
|
+
"agent": "security-reviewer", "task": "Audit {item} for injection/authz issues.",
|
|
320
|
+
"cache": { "scope": "cross-run" },
|
|
321
|
+
"dependsOn": ["discover"] },
|
|
322
|
+
{ "id": "report", "type": "reduce", "from": ["audit"], "agent": "doc-writer",
|
|
323
|
+
"task": "Prioritized findings report:\n{steps.audit.output}",
|
|
324
|
+
"dependsOn": ["audit"], "final": true }
|
|
325
|
+
]
|
|
326
|
+
}
|
|
327
|
+
```
|
|
328
|
+
|
|
329
|
+
The `glob!:` fingerprint makes `discover` a cache miss only when file
|
|
330
|
+
*contents* change; the map's per-item cache means one changed file re-audits
|
|
331
|
+
one item.
|
|
332
|
+
|
|
333
|
+
---
|
|
334
|
+
|
|
335
|
+
## Anti-patterns (seen in real flows)
|
|
336
|
+
|
|
337
|
+
| Anti-pattern | Why it fails | Fix |
|
|
338
|
+
|--------------|--------------|-----|
|
|
339
|
+
| One mega-phase doing discover+audit+report | No parallelism, no caching granularity, one failure loses everything | Split along the archetype-1 shape |
|
|
340
|
+
| Gate whose task doesn't demand a `VERDICT:` terminator | Ambiguous model output fails closed → gate blocks on a model that forgot the verdict | Use `output:"json"` + `expect` enum (preferred), or end the task with the exact `VERDICT: PASS\|BLOCK` instruction (auto-appended if you omit it) |
|
|
341
|
+
| Router phase without `expect` enum | `"Deep"` vs `"deep"` → both `when` branches skip, `join:"any"` reduce gets nothing | `expect: { properties: { route: { enum: [...] } } }` + `retry` |
|
|
342
|
+
| Agent phase that just runs a shell command | Tokens spent, output paraphrased inaccurately | `script` phase |
|
|
343
|
+
| Same agent produces and reviews | Self-review passes everything | Different agent (ideally model) for the gate |
|
|
344
|
+
| Fan-out with no `budget` and no `concurrency` cap | Unbounded spend + rate-limit storms | `budget` + `phase.concurrency` |
|
|
345
|
+
| `dependsOn` declared but output never referenced | The downstream agent doesn't see the upstream's work — dependency ≠ data flow | Interpolate `{steps.X.output}` into the task (or `context`) |
|
|
346
|
+
| Saving a flow without `strictInterpolation` | Later invocations with wrong args silently run on empty strings | `strictInterpolation: true` before saving |
|
|
347
|
+
| `map` over `{steps.X.output}` (text, not json) | `over` must resolve to an array | `output: "json"` upstream + `over: "{steps.X.json}"` |
|
|
348
|
+
| Deep `chain` where steps don't need each other's output | Serialized latency for no reason | `tasks` (parallel) or a DAG with real edges only |
|