hermes-taskflow 0.2.9

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,105 @@
1
+ <!-- GENERATED FILE — do not edit. Source: skills-src/taskflow/library.md (npm run build:skills) -->
2
+
3
+ # Library: reusable flows & the search-before-author loop
4
+
5
+ taskflow saves flows you'll reuse into a **library** with metadata, so you can
6
+ `search` it before authoring a new flow — the **reuse flywheel**. The more you
7
+ save (with a good `purpose` + `tags`), the better future search gets.
8
+
9
+ > Full design: `docs/rfc-library-reuse.md`. This file is the agent-facing
10
+ > "when and how" guide.
11
+
12
+ **Host binding:** use the `taskflow_search`, `taskflow_save`, `taskflow_list`,
13
+ `taskflow_show` tools.
14
+
15
+ ## Before authoring a non-trivial flow: SEARCH first
16
+
17
+ Any time you're about to write a flow with ≥3 phases, fan-out, or a gate,
18
+ **search the library first**. It costs nothing and often finds a starter you
19
+ can adapt.
20
+
21
+ ```jsonc
22
+ { "name": "taskflow_search", "arguments": { "query": "audit API endpoints for missing auth", "limit": 5 } }
23
+ ```
24
+
25
+ Read the results and the `→ reuseHint`:
26
+
27
+ | score | reuseHint | what to do |
28
+ |-------|-----------|------------|
29
+ | **≥ 0.8** | "直接复用" / direct reuse | Run by name (skip authoring). |
30
+ | **0.5 – 0.8** | "copy + 泛化" / copy & generalize | `show` it, copy, **generalize** (see checklist), save as a new version. |
31
+ | **< 0.5** or no matches | "从头编写" / write fresh | Author from scratch — then **save** it if reusable. |
32
+
33
+ `searchMode` tells you how the ranking was produced: `structural` (keyword +
34
+ phase-signature; no embedding backend configured), `semantic` (cosine over
35
+ embeddings), or `mixed` (some flows had vectors, some didn't). `structural` is
36
+ weaker on paraphrase — if it missed something obvious, try `structureOnly:
37
+ false` or a different phrasing.
38
+
39
+ ## After a successful novel flow: SAVE it (if reusable)
40
+
41
+ When you finish a flow you expect to use again, save it **with a `purpose` and
42
+ 2–4 `tags`**. These two fields are what search matches on — a flow saved
43
+ without them is nearly invisible to future search.
44
+
45
+ ```jsonc
46
+ { "name": "taskflow_save",
47
+ "arguments": { "name": "audit-endpoints", "definition": { "phases": [ ... ] },
48
+ "purpose": "Audit a directory of API endpoints for missing auth checks",
49
+ "tags": ["audit", "security", "auth", "fan-out"] } }
50
+ ```
51
+
52
+ `save` auto-derives structural metadata (phase signature, a `generality` score
53
+ in 0–1) and writes a sidecar `.meta.json` next to the flow file. You don't
54
+ compute any of that — just give `purpose` + `tags`.
55
+
56
+ ## The generalization checklist (apply on every reuse)
57
+
58
+ When you copy + generalize a flow, make it **more reusable than the version you
59
+ found**. Each item raises the auto-derived `generality` score and broadens
60
+ future search recall:
61
+
62
+ - [ ] Hardcoded file/dir paths → `{args.X}` (with a `default`).
63
+ - [ ] Specific entity words ("endpoint", "route") → broaden in the discover
64
+ prompt so the flow works for the whole class.
65
+ - [ ] Thresholds / counts → `{args.X}` with sensible defaults.
66
+ - [ ] Add `budget` / `retry` / `expect` if missing (production-grade knobs).
67
+ - [ ] Update `purpose` to reflect the wider scope.
68
+
69
+ Then save it back (version auto-bumps). Over time the library compounds: every
70
+ reuse leaves a more general flow behind.
71
+
72
+ ## reuseCount & `reusedFromSearch`
73
+
74
+ Each saved flow has a `reuseCount`. It goes up by 1 **only when** a run was
75
+ chosen because of a prior search — set the `reusedFromSearch: true` flag on the
76
+ run. Direct run-by-name does **not** bump it (that's intentional: `reuseCount`
77
+ measures "found-via-search reuse", the high-quality signal for later auto-prune).
78
+
79
+ ```jsonc
80
+ { "name": "taskflow_run",
81
+ "arguments": { "name": "audit-endpoints", "args": { "dir": "src/api" }, "reusedFromSearch": true } }
82
+ ```
83
+
84
+ ## Judicious reuse — not every task needs the library
85
+
86
+ Skip search + save for:
87
+ - **One-off tasks** (a quick fix, a throwaway analysis) — `generality < 0.3`
88
+ and obviously won't recur.
89
+ - **Trivial flows** (1–2 phases, no fan-out) — overhead isn't worth it.
90
+
91
+ The library pays off for *patterns* that recur across projects or sessions:
92
+ audits, migrations, reviews, fan-out summarization, plan→approve→execute, etc.
93
+
94
+ ## Configuration (embedding backend — optional, Phase 2)
95
+
96
+ Search works with **zero config** (structural mode). For smarter paraphrase
97
+ recall, configure an embedding backend in `~/.pi/agent/settings.json`:
98
+
99
+ ```jsonc
100
+ { "taskflow": { "library": { "enabled": true, "scope": "both" },
101
+ "embedder": { "kind": "http", "url": "http://127.0.0.1:8123/v1/embeddings", "model": "qwen3-embedding-0.6b" } } }
102
+ ```
103
+
104
+ Without `embedder`, or if the embedder fails, search **degrades gracefully** to
105
+ structural mode — it never breaks. (See `docs/rfc-library-reuse.md` §4.)
@@ -0,0 +1,348 @@
1
+ <!-- GENERATED FILE — do not edit. Source: skills-src/taskflow/patterns.md (npm run build:skills) -->
2
+
3
+ # Taskflow Patterns — proven archetypes & the production checklist
4
+
5
+ Read this when designing a flow with ≥ 4 phases, a gate, or any fan-out.
6
+ Each archetype below is a complete, runnable shape distilled from real runs.
7
+ Copy the closest one and adapt — don't design from a blank page.
8
+
9
+ ---
10
+
11
+ ## The production-flow checklist
12
+
13
+ Before running a flow you designed, check it against this list. Every item is
14
+ cheap to add and each one has prevented a real failure class:
15
+
16
+ - [ ] **`verify` first.** Run the static verifier (pi: `action: "verify"` /
17
+ Codex / Claude Code / OpenCode: `taskflow_verify`) — zero tokens,
18
+ catches cycles / missing deps / ref typos / contract mismatches.
19
+ - [ ] **Every JSON-emitting phase has `expect` + `retry`.** "Output ONLY JSON"
20
+ in the task text is a request; `expect` is enforcement. Without it, a
21
+ malformed router output silently skips both branches.
22
+ - [ ] **Machine checks before LLM checks.** `script` phases for build/test
23
+ ground truth; gate `eval` assertions before the gate's LLM `task`. A
24
+ token spent verifying what a shell command can verify is a wasted token.
25
+ - [ ] **`budget` on any flow with a fan-out.** A map over a mis-discovered
26
+ 500-item array is unbounded spend without one.
27
+ - [ ] **`optional: true` + a fallback for degradable phases.** An enrichment
28
+ phase that times out shouldn't sink the run — pair `timeout` +
29
+ `optional` with a downstream `when`-guarded fallback.
30
+ - [ ] **Exactly one `final: true`** on the result-bearing phase.
31
+ - [ ] **`strictInterpolation: true` on any flow you save.** Saved flows run
32
+ later with args you're not watching; unresolved placeholders must be
33
+ errors, not empty strings.
34
+ - [ ] **Discovery phases are cheap and tool-restricted.** `agent: "scout"`,
35
+ `tools: ["read","grep","ls"]`, low thinking. Filesystem isolation still
36
+ depends on enforcement by the selected host; don't pay executor prices for `ls`.
37
+ - [ ] **The reviewer is not the producer.** A gate reviewing phase X uses a
38
+ different agent (ideally a different model) than X. Self-review passes
39
+ ~everything.
40
+ - [ ] **Re-runnable flows are `incremental: true`** with `fingerprint` entries
41
+ on phases that read the world (`git:HEAD`, `glob!:src/**/*.ts`). See
42
+ `advanced.md`.
43
+
44
+ ---
45
+
46
+ ## Archetype 1: Audit fan-out (discover → map → gate → reduce)
47
+
48
+ The workhorse. Use for: audit every endpoint / migrate every file /
49
+ summarize every module.
50
+
51
+ ```jsonc
52
+ {
53
+ "name": "audit-endpoints",
54
+ "strictInterpolation": true,
55
+ "budget": { "maxUSD": 3.00 },
56
+ "phases": [
57
+ { "id": "discover", "type": "agent", "agent": "scout",
58
+ "tools": ["read", "grep", "ls"],
59
+ "task": "List every HTTP endpoint under src/routes. Output ONLY a JSON array [{\"route\":\"...\",\"file\":\"...\"}]. No prose.",
60
+ "output": "json",
61
+ "expect": { "type": "array", "items": { "type": "object", "required": ["route", "file"] } },
62
+ "retry": { "max": 2, "backoffMs": 0 } },
63
+
64
+ { "id": "audit", "type": "map", "over": "{steps.discover.json}", "as": "item",
65
+ "agent": "analyst", "concurrency": 4,
66
+ "task": "Audit {item.route} in {item.file} for missing auth. Report: SEVERITY (high/med/low/none), evidence (file:line), fix.",
67
+ "dependsOn": ["discover"] },
68
+
69
+ { "id": "screen", "type": "gate", "agent": "reviewer",
70
+ "task": "Cross-check the findings below. Delete false positives (cite why). If ANY confirmed HIGH remains, end with VERDICT: BLOCK and list them; else VERDICT: PASS.\n\n{steps.audit.output}",
71
+ "dependsOn": ["audit"] },
72
+
73
+ { "id": "report", "type": "reduce", "from": ["screen"], "agent": "doc-writer",
74
+ "task": "Write a prioritized remediation report from:\n{steps.screen.output}",
75
+ "dependsOn": ["screen"], "final": true }
76
+ ]
77
+ }
78
+ ```
79
+
80
+ Why each piece: `expect`+`retry` on discover means a chatty scout gets a second
81
+ chance instead of feeding garbage to the map; `concurrency: 4` protects rate
82
+ limits; the gate is a *different* agent than the auditor; `budget` limits the
83
+ fan-out.
84
+
85
+ **Variant — per-item caching for repeated audits:** add
86
+ `"cache": { "scope": "cross-run" }` to the map phase. On the next run, only
87
+ items whose task text changed re-execute; the rest are $0 cache hits.
88
+ (Details: `configuration.md` §8.)
89
+
90
+ ---
91
+
92
+ ## Archetype 2: Self-healing implement→verify→rework
93
+
94
+ Use for: implement against acceptance criteria, fix-until-green.
95
+ The gate re-runs its upstream on BLOCK — a generate→critique→regenerate loop
96
+ without you writing a loop.
97
+
98
+ ```jsonc
99
+ {
100
+ "name": "implement-verified",
101
+ "budget": { "maxUSD": 5.00 },
102
+ "phases": [
103
+ { "id": "implement", "type": "agent", "agent": "executor-code",
104
+ "task": "Implement the feature per the spec in docs/spec.md. Run nothing; just edit." },
105
+
106
+ { "id": "build-test", "type": "script",
107
+ "run": "npx tsc --noEmit && pnpm test 2>&1 | tail -20",
108
+ "timeout": 180000, "dependsOn": ["implement"] },
109
+
110
+ { "id": "spec-gate", "type": "gate", "agent": "reviewer",
111
+ "onBlock": "retry", "retry": { "max": 3 },
112
+ "eval": ["{steps.build-test.output} contains pass"],
113
+ "task": "Build/test output:\n{steps.build-test.output}\n\nDoes the implementation satisfy ALL acceptance criteria in docs/spec.md? VERDICT: PASS, or VERDICT: BLOCK with a precise list of what to fix.",
114
+ "dependsOn": ["implement", "build-test"] },
115
+
116
+ { "id": "summary", "type": "agent", "agent": "doc-writer",
117
+ "task": "Summarize what was implemented and the verification result:\n{steps.spec-gate.output}",
118
+ "dependsOn": ["spec-gate"], "final": true }
119
+ ]
120
+ }
121
+ ```
122
+
123
+ Key mechanics: on BLOCK, `spec-gate` re-runs **both** its `dependsOn` upstreams
124
+ (`implement` gets the blocker's reasons via re-interpolation, `build-test`
125
+ re-verifies), up to 3 rounds. The `eval` line means a green build+test skips
126
+ the LLM review entirely on the happy path.
127
+
128
+ **Verification phases: force structured output.** LLMs are bad at
129
+ *summarizing* shell output (234 tests read as 230) but good at *copying*
130
+ structured data. If a verification step must go through an agent (not a
131
+ `script`), demand `key=value` lines:
132
+
133
+ ```
134
+ Report EXACTLY in this format (one key=value per line, no prose):
135
+ typecheck=PASS|FAIL
136
+ tests_total=N
137
+ tests_fail=N
138
+ If any field is missing, you failed the task — re-run and re-read.
139
+ ```
140
+
141
+ Prefer a `script` phase whenever the check is a command — exact, free, fast.
142
+
143
+ ---
144
+
145
+ ## Archetype 3: Plan → human approval → execute
146
+
147
+ Use for: anything expensive or destructive where a human should see the plan
148
+ before the spend. The approval's **Edit** option injects mid-run guidance.
149
+
150
+ > **MCP-host caveat (Codex / Claude Code / OpenCode / Grok / Hermes):** approval phases auto-reject in
151
+ > MCP-driven (non-interactive) runs. This archetype only works when a human
152
+ > runs the flow interactively; for tool-driven runs, replace the approval with
153
+ > a strict `gate`.
154
+
155
+ ```jsonc
156
+ {
157
+ "name": "guarded-migration",
158
+ "phases": [
159
+ { "id": "plan", "type": "agent", "agent": "planner",
160
+ "task": "Plan the migration of src/legacy/* to the new API. List each file, the change, and the risk." },
161
+
162
+ { "id": "checkpoint", "type": "approval",
163
+ "task": "Migration plan:\n\n{steps.plan.output}\n\nApprove to execute, reject to abort, or edit to add constraints.",
164
+ "dependsOn": ["plan"] },
165
+
166
+ { "id": "execute", "type": "agent", "agent": "executor-code",
167
+ "task": "Execute the migration plan:\n{steps.plan.output}\n\nOperator guidance (if any): {steps.checkpoint.output}",
168
+ "dependsOn": ["checkpoint"] },
169
+
170
+ { "id": "verify", "type": "script", "run": "npx tsc --noEmit && pnpm test 2>&1 | tail -5",
171
+ "timeout": 180000, "dependsOn": ["execute"], "final": true }
172
+ ]
173
+ }
174
+ ```
175
+
176
+ Note `{steps.checkpoint.output}` — on Edit it carries the operator's note; on
177
+ Approve it's `(approve)`. Don't use approval in detached/headless runs (it
178
+ auto-rejects there — by design).
179
+
180
+ ---
181
+
182
+ ## Archetype 4: Dynamic plan → execute (`flow{def}`)
183
+
184
+ Use when the *work itself* must be discovered at runtime — the planner emits a
185
+ whole sub-flow as JSON and the runtime validates + runs it. The declarative
186
+ answer to "loop over whatever we find".
187
+
188
+ ```jsonc
189
+ {
190
+ "name": "dynamic-audit",
191
+ "budget": { "maxUSD": 4.00 },
192
+ "phases": [
193
+ { "id": "plan", "type": "agent", "agent": "planner", "output": "json",
194
+ "task": "Scan this repo. Output ONLY a JSON taskflow {\"name\":\"sub\",\"phases\":[...]} with one 'agent' phase per module that needs auditing (agent: \"analyst\"), plus a final 'reduce' phase (agent: \"doc-writer\", from: [all audit ids], final: true). Use hyphens in ids. No script phases, no cwd fields.",
195
+ "expect": { "type": "object", "required": ["name", "phases"] },
196
+ "retry": { "max": 2, "backoffMs": 0 } },
197
+
198
+ { "id": "run-plan", "type": "flow", "def": "{steps.plan.json}",
199
+ "optional": true, "dependsOn": ["plan"] },
200
+
201
+ { "id": "deliver", "type": "agent", "agent": "doc-writer",
202
+ "task": "Final result (empty means the plan failed validation — say so):\n{steps.run-plan.output}",
203
+ "dependsOn": ["run-plan"], "final": true }
204
+ ]
205
+ }
206
+ ```
207
+
208
+ Critical details (full contract in `advanced.md`): a bad plan **fails open**
209
+ (`defError` diagnostic, empty output downstream) — `optional: true` + the
210
+ `deliver` phase turn that into a graceful report instead of a dead run. The
211
+ planner prompt must forbid what validation will reject anyway (`script`
212
+ phases, workspace `cwd` keywords) so the plan doesn't waste a retry.
213
+
214
+ **Iterative replanning:** wrap a plan-emitting body in a `loop` so round N's
215
+ plan reacts to round N−1's result — see `examples/iterative-replan.json`.
216
+
217
+ ---
218
+
219
+ ## Archetype 5: Tournament for one-shot-unreliable work
220
+
221
+ Use when quality varies run-to-run (naming, copywriting, tricky refactor
222
+ strategy, root-cause hypotheses). Branches > variants when you want genuinely
223
+ different *approaches* judged against each other.
224
+
225
+ ```jsonc
226
+ {
227
+ "id": "strategy", "type": "tournament", "mode": "best",
228
+ "judgeAgent": "final-arbiter",
229
+ "judge": "Judge on: correctness under concurrent access, blast radius, migration cost. Quote evidence. Return JSON {\"winner\": <n>, \"reason\": \"...\"}.",
230
+ "branches": [
231
+ { "task": "Design the cache-invalidation fix with a conservative approach: minimal diff, no schema change.", "agent": "analyst" },
232
+ { "task": "Design the fix assuming we can change the schema: optimal correctness.", "agent": "analyst" },
233
+ { "task": "Design the fix as an adversary: what will break each obvious approach? Then propose the one that survives.", "agent": "critic" }
234
+ ],
235
+ "dependsOn": ["context"], "final": true
236
+ }
237
+ ```
238
+
239
+ Give the judge a **rubric with named criteria**, a stronger model
240
+ (`judgeAgent`), and a structured winner output (`{"winner": <n>}` JSON,
241
+ or an exact `WINNER: <n>` terminator). `mode: "aggregate"`
242
+ instead merges all variants — good for research synthesis, bad for decisions.
243
+
244
+ ---
245
+
246
+ ## Archetype 7: Race for latency (first approach wins)
247
+
248
+ When several strategies can answer the same question and you care about **time
249
+ to first good answer** more than comparing quality, use `race` (not
250
+ `tournament`). Branches start together; the first successful completion becomes
251
+ the phase output.
252
+
253
+ ```jsonc
254
+ {
255
+ "name": "quick-answer",
256
+ "budget": { "maxUSD": 0.5 },
257
+ "phases": [
258
+ {
259
+ "id": "answer", "type": "race",
260
+ "branches": [
261
+ { "task": "Answer from local heuristics only: {args.q}", "agent": "executor" },
262
+ { "task": "Answer after a short web/docs look: {args.q}", "agent": "researcher" }
263
+ ],
264
+ "final": true
265
+ }
266
+ ]
267
+ }
268
+ ```
269
+
270
+ **Prefer tournament when** you need a judge to pick the *best* draft after all
271
+ variants finish. **Prefer parallel when** you need *every* branch's output
272
+ downstream (then `reduce`).
273
+
274
+ ## Archetype 8: Expand graft (planner fragment on the parent DAG)
275
+
276
+ Planner emits a fragment; `expand` with `expandMode: "graft"` runs it and
277
+ promotes child phase states as `<expandId>-<childId>` so a later phase can read
278
+ them. Use `nested` (or classic `flow{def}`) when you only need the fragment's
279
+ **final** output and do not want child ids on the parent.
280
+
281
+ ```jsonc
282
+ {
283
+ "name": "plan-graft",
284
+ "phases": [
285
+ {
286
+ "id": "plan", "type": "agent", "agent": "planner", "output": "json",
287
+ "task": "Emit a mini-flow JSON: {name, phases:[{id,type,agent,task,final?}…]} for the audit.",
288
+ "expect": { "type": "object", "required": ["phases"] }
289
+ },
290
+ {
291
+ "id": "grow", "type": "expand", "expandMode": "graft",
292
+ "def": "{steps.plan.json}", "dependsOn": ["plan"]
293
+ },
294
+ {
295
+ "id": "wrap", "type": "agent", "agent": "writer",
296
+ "task": "Summarize grafted work. Child outputs may appear as steps.grow-* in the run state.",
297
+ "dependsOn": ["grow"], "final": true
298
+ }
299
+ ]
300
+ }
301
+ ```
302
+
303
+ ## Archetype 6: Incremental repo-watch audit (cross-run)
304
+
305
+ Use for flows you'll re-run as the repo evolves. First run pays full price;
306
+ subsequent runs re-pay only for what changed.
307
+
308
+ ```jsonc
309
+ {
310
+ "name": "security-sweep",
311
+ "incremental": true,
312
+ "budget": { "maxUSD": 3.00 },
313
+ "phases": [
314
+ { "id": "discover", "type": "agent", "agent": "scout", "output": "json",
315
+ "task": "List all files handling user input. Output ONLY a JSON array of paths.",
316
+ "expect": { "type": "array" },
317
+ "cache": { "scope": "cross-run", "fingerprint": ["glob!:src/**/*.ts"] } },
318
+ { "id": "audit", "type": "map", "over": "{steps.discover.json}",
319
+ "agent": "security-reviewer", "task": "Audit {item} for injection/authz issues.",
320
+ "cache": { "scope": "cross-run" },
321
+ "dependsOn": ["discover"] },
322
+ { "id": "report", "type": "reduce", "from": ["audit"], "agent": "doc-writer",
323
+ "task": "Prioritized findings report:\n{steps.audit.output}",
324
+ "dependsOn": ["audit"], "final": true }
325
+ ]
326
+ }
327
+ ```
328
+
329
+ The `glob!:` fingerprint makes `discover` a cache miss only when file
330
+ *contents* change; the map's per-item cache means one changed file re-audits
331
+ one item.
332
+
333
+ ---
334
+
335
+ ## Anti-patterns (seen in real flows)
336
+
337
+ | Anti-pattern | Why it fails | Fix |
338
+ |--------------|--------------|-----|
339
+ | One mega-phase doing discover+audit+report | No parallelism, no caching granularity, one failure loses everything | Split along the archetype-1 shape |
340
+ | Gate whose task doesn't demand a `VERDICT:` terminator | Ambiguous model output fails closed → gate blocks on a model that forgot the verdict | Use `output:"json"` + `expect` enum (preferred), or end the task with the exact `VERDICT: PASS\|BLOCK` instruction (auto-appended if you omit it) |
341
+ | Router phase without `expect` enum | `"Deep"` vs `"deep"` → both `when` branches skip, `join:"any"` reduce gets nothing | `expect: { properties: { route: { enum: [...] } } }` + `retry` |
342
+ | Agent phase that just runs a shell command | Tokens spent, output paraphrased inaccurately | `script` phase |
343
+ | Same agent produces and reviews | Self-review passes everything | Different agent (ideally model) for the gate |
344
+ | Fan-out with no `budget` and no `concurrency` cap | Unbounded spend + rate-limit storms | `budget` + `phase.concurrency` |
345
+ | `dependsOn` declared but output never referenced | The downstream agent doesn't see the upstream's work — dependency ≠ data flow | Interpolate `{steps.X.output}` into the task (or `context`) |
346
+ | Saving a flow without `strictInterpolation` | Later invocations with wrong args silently run on empty strings | `strictInterpolation: true` before saving |
347
+ | `map` over `{steps.X.output}` (text, not json) | `over` must resolve to an array | `output: "json"` upstream + `over: "{steps.X.json}"` |
348
+ | Deep `chain` where steps don't need each other's output | Serialized latency for no reason | `tasks` (parallel) or a DAG with real edges only |