codex-taskflow 0.1.6 → 0.1.8
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +9 -7
- package/dist/mcp/server.js.map +1 -1
- package/package.json +4 -4
package/README.md
CHANGED
|
@@ -506,7 +506,7 @@ for the agent-facing search→reuse→generalize→re-save workflow.
|
|
|
506
506
|
- **`retry`** — `{ "max": 2, "backoffMs": 500, "factor": 2 }` retries a failing subagent with fixed or exponential backoff; usage is summed and the attempt count shows as `↻N` in the TUI. Transient provider errors (rate-limit / 5xx / timeout) **auto-retry even without an explicit policy**; hard errors don't.
|
|
507
507
|
- **`onBlock`** — `"halt"` (default) stops the run when a gate blocks. `"retry"` retries upstream phases when a gate blocks, instead of halting — a self-healing rework loop with budget and idle-watchdog guards and a nested recursion depth cap.
|
|
508
508
|
- **`eval`** — zero-token machine-checkable criteria that run *before* the LLM gate. If the eval check fails, the gate blocks without spawning an agent.
|
|
509
|
-
- **`score`** — graded, composable quality gates: deterministic scorers (`exact-match`, `contains`, `regex`, `json-schema`, `length-range`, `code-compiles`) run against a target string at **zero tokens** and combine via `all`/`any`/`weighted` against a `threshold`. Deterministic pass → auto-PASS with no LLM call **when the judge cannot veto** — no judge configured, or `weighted` where the deterministic score is a *lower bound* already clearing the threshold. With `all`/`any` + a judge, the judge always runs (its verdict is authoritative — it may check what scorers cannot, e.g. factuality). Deterministic fail → the optional LLM `judge` decides (fail-
|
|
509
|
+
- **`score`** — graded, composable quality gates: deterministic scorers (`exact-match`, `contains`, `regex`, `json-schema`, `length-range`, `code-compiles`) run against a target string at **zero tokens** and combine via `all`/`any`/`weighted` against a `threshold`. Deterministic pass → auto-PASS with no LLM call **when the judge cannot veto** — no judge configured, or `weighted` where the deterministic score is a *lower bound* already clearing the threshold. With `all`/`any` + a judge, the judge always runs (its verdict is authoritative — it may check what scorers cannot, e.g. factuality). Deterministic fail → the optional LLM `judge` decides (**fail-closed** on unparseable output — issue #54), or the gate `task` runs with the scorer report appended, or — with no fallback — the gate **blocks explicitly**. The structured result is the gate's `.json` (`{steps.<gate>.json.combined}`, `.json.results`), so downstream phases can route on quality, not just pass/fail. LLM-generated dynamic sub-flows may not use `code-compiles` (compiler execution) or `regex` (ReDoS) scorers — same hardening class as the `script` block.
|
|
510
510
|
- **`idempotent: false`** — side-effect classification for phases with **irreversible effects** (webhook POSTs, deploys, DB writes): the implicit transient auto-retry is suppressed (an explicit `retry{}` is still honored — it's the author's declaration that repeats are acceptable) and the result is **never cached** in any scope (within-run resume, cross-run, `incremental`) — the phase re-runs every time. The phase state records `sideEffect: true` (rendered as ⚡). Default `true` — existing flows are unchanged.
|
|
511
511
|
- **`approval`** — pause for a human (Approve / Reject / Edit). Reject halts the flow; Edit injects the typed note as the phase output for downstream steps. Non-interactive runs (detached / CI) **auto-reject** (safety: approval gates are never bypassed).
|
|
512
512
|
- **`flow`** — `{ "type": "flow", "use": "deep-research", "with": { "topic": "{item}" } }` runs a **saved** flow as a phase (recursion is detected and rejected). Or **generate the sub-flow at runtime**: `{ "type": "flow", "def": "{steps.plan.json}" }` resolves an upstream phase's JSON output into a sub-flow, **validates it (cycles / dangling refs / duplicate ids / dead-ends), then runs it** — the number and shape of the generated phases is decided at runtime, not authored in advance. A malformed plan fails *open* (the phase is skipped with a `defError`, the run continues). This is how a planner decides *at runtime* what work to spawn — the declarative answer to a code-mode `for` loop, with each generated plan checked before it spends a token. Security hardening for LLM-generated sub-flows: breadth caps (100 phases, 200 map items, 16 concurrency), `cwd` containment, budget clamped to `min(child, parent)`, nesting cap (5 levels), and prototype-pollution defense (deep-cloned, `__proto__`/`constructor`/`prototype` stripped). Pair it with `loop` for **data-dependent iterative replanning** (round N's plan depends on round N-1's result). See [`examples/dynamic-plan-execute.json`](https://github.com/heggria/taskflow/blob/main/examples/dynamic-plan-execute.json) and [`examples/iterative-replan.json`](https://github.com/heggria/taskflow/blob/main/examples/iterative-replan.json).
|
|
@@ -550,7 +550,7 @@ For open-ended work, the best result often comes from generating several candida
|
|
|
550
550
|
```
|
|
551
551
|
|
|
552
552
|
- **Competitors** — either `variants: N` copies of one `task` (diversity comes from model nondeterminism), or distinct `branches: [{task, agent?}, …]` when you want to pit *different approaches* against each other.
|
|
553
|
-
- **Judge** — after the fan-out, one judge agent sees every variant (numbered) plus your `judge` rubric and picks a winner
|
|
553
|
+
- **Judge** — after the fan-out, one judge agent sees every variant (numbered) plus your `judge` rubric and picks a winner. Prefer a JSON pick `{"winner": n}` (most robust); the runtime also reads a `WINNER: <n>` line (`#n` and Markdown emphasis like `WINNER: **n**` are tolerated — issue #54). An unreadable verdict **fails open** to variant 1; a failed judge falls back too — the work is never lost.
|
|
554
554
|
- **`mode`** — `best` returns the winning variant **verbatim**; `aggregate` returns the judge's **synthesized** answer combining the strongest parts.
|
|
555
555
|
- **Short-circuits:** if only one competitor survives, it wins with no judge call; if all fail, the phase fails. The TUI shows `⚑ N→#k`; usage sums variants + judge. Like `gate`, it's **excluded from `cross-run` cache**.
|
|
556
556
|
|
|
@@ -603,12 +603,12 @@ Every phase is already content-addressed: within a single run's **resume**, a ph
|
|
|
603
603
|
|
|
604
604
|
### Gate phases (quality control)
|
|
605
605
|
|
|
606
|
-
A `gate` runs an agent to review upstream output and can **block the rest of the workflow.**
|
|
606
|
+
A `gate` runs an agent to review upstream output and can **block the rest of the workflow.** Provide a verdict the runtime can read — **preferably via a JSON contract** (`output: "json"` + `expect: { properties: { verdict: { enum: ["pass","block"] } } }`), which machine-validates the output so a verdict can never be silently misread. Otherwise:
|
|
607
607
|
|
|
608
|
-
- a final line `VERDICT: PASS` or `VERDICT: BLOCK` (also accepts `OK`, `FAIL`, `STOP`, `REJECT`, `HALT` — last occurrence wins), or
|
|
608
|
+
- a final line `VERDICT: PASS` or `VERDICT: BLOCK` (also accepts `OK`, `FAIL`, `STOP`, `REJECT`, `HALT` — last occurrence wins; common Markdown emphasis like `VERDICT: **BLOCK**` is tolerated), or
|
|
609
609
|
- JSON like `{"continue": false, "reason": "missing auth checks"}` / `{"verdict": "block", "reason": "..."}`.
|
|
610
610
|
|
|
611
|
-
On **BLOCK**, downstream phases skip and the run ends as `blocked` with the reason surfaced. **
|
|
611
|
+
If a free-text gate's task doesn't already ask for a `VERDICT:` marker, the runtime **auto-appends** the exact format instruction. On **BLOCK**, downstream phases skip and the run ends as `blocked` with the reason surfaced. **Unparseable gate model output fails closed** (treated as BLOCK) — a gate that cannot reach a verdict cannot be trusted to pass (issue #54). Config slips (unresolved `score.target`, malformed `scorers`) still fail **open** with a warning.
|
|
612
612
|
|
|
613
613
|
```
|
|
614
614
|
Review the audit below. If any endpoint is missing auth, end with
|
|
@@ -648,10 +648,12 @@ Saved flows become CLI shortcuts. **These `/tf` commands are Pi-only** (they run
|
|
|
648
648
|
| `/tf compile <name> [lr\|td]` | **Render the flow as a Mermaid diagram + verification overlay** — 0 tokens, no LLM; paste into a README/issue/PR |
|
|
649
649
|
| `/tf runs` | Browse recent run history (interactive TUI — **live auto-refreshes** while any run is active) |
|
|
650
650
|
| `/tf resume <runId>` | Continue a paused/failed run — cached phases skip automatically |
|
|
651
|
+
| `/tf peek <runId> [phaseId]` | Inspect a phase's intermediate output (the debugging escape hatch) |
|
|
652
|
+
| `/tf trace <runId> [--json]` | Show a run's **deterministic-replay event trace** (each subagent call + runtime decisions) |
|
|
651
653
|
| `/tf init` | **Interactively map model roles** to your enabled models (writes `~/.pi/agent/settings.json`) |
|
|
652
654
|
| `/tf:<name> [args]` | Shortcut — runs the flow in one tap |
|
|
653
655
|
|
|
654
|
-
Tool actions (used by the model on Pi): `run` (inline `define` or saved `name`), `save`, `resume`, `list`, `agents`, `init`, `verify`, `compile`, `ir`, `provenance`, `why-stale`, `recompute`, `cache-clear`. On Codex, Claude Code, and OpenCode the exposed MCP tools are `taskflow_run` / `taskflow_list` / `taskflow_show` / `taskflow_verify` / `taskflow_compile` / `taskflow_peek`.
|
|
656
|
+
Tool actions (used by the model on Pi): `run` (inline `define` or saved `name`), `save`, `resume`, `list`, `agents`, `init`, `verify`, `compile`, `ir`, `provenance`, `trace`, `why-stale`, `recompute`, `cache-clear`, `search`. On Codex, Claude Code, and OpenCode the exposed MCP tools are `taskflow_run` / `taskflow_list` / `taskflow_show` / `taskflow_verify` / `taskflow_compile` / `taskflow_peek` / `taskflow_trace` / `taskflow_why_stale` / `taskflow_recompute` (dry-run only) / `taskflow_save` / `taskflow_search`.
|
|
655
657
|
|
|
656
658
|
## Background (detached) execution
|
|
657
659
|
|
|
@@ -895,7 +897,7 @@ Our `self-improve` flow is a 10-phase DAG — it audits the codebase, patches de
|
|
|
895
897
|
|
|
896
898
|
## Status & limits
|
|
897
899
|
|
|
898
|
-
**v0.1.
|
|
900
|
+
**v0.1.7** (current release) — **file loaders now report *why* a file failed with the parse position** (line/column) instead of a merged "not found or unparseable" message — `defineFile`, saved flows, run records, and library sidecars all distinguish *missing* from *malformed*, so a stray bare newline in a hand-authored flow is diagnosable in seconds; `safeParse` stays lenient for LLM output. Also fixes a pi-taskflow hint that re-printed every session. **Gate safety hardening (issue #54)**: a shared emphasis-tolerant marker factory now covers **all three decision markers** — `VERDICT`, `WINNER`, and `SCORE` — so Markdown-wrapped tokens (`VERDICT: **BLOCK**`, `WINNER: __3__`, `SCORE: `0.8``) are never silently mis-read (a genuine BLOCK no longer becomes PASS; a judge's pick no longer silently reverts to variant 1); **unparseable gate *model output now fails closed* (BLOCK)** instead of rubber-stamping PASS — a gate that cannot reach a verdict cannot be trusted to pass, while *config* slips (unresolved `score.target`, malformed `scorers`) stay fail-open with a warning; and free-text gates whose task omits a `VERDICT:` instruction now get the exact format suffix **auto-appended**. For the most robust decision phases, use `output: "json"` + `expect` to machine-validate the output (now the documented default for gate verdicts, tournament winners, and router branches). **v0.1.6** added **library Phase 1** (search-before-author + reusable-flow sidecar metadata), the **`defineFile`** parameter (verify/compile/run a flow from a path on disk), and **JSONC comment support** in flow definition files (`//` and `/* */` comments + trailing commas, parsed by the new zero-dependency `parseJsonc`). **v0.1.5** added **Claude Code and OpenCode as hosts**, **extracted the MCP server into its own `taskflow-mcp-core` package**, and **de-duplicated the three host runners** into a shared `runSubagentProcess`. See [CHANGELOG](https://github.com/heggria/taskflow/blob/main/CHANGELOG.md) for the full history. Baseline: **multi-host monorepo of seven packages** — the host-neutral `taskflow-core` engine, the host-neutral `taskflow-mcp-core` MCP server, the shared host-runner `taskflow-hosts`, plus `pi-taskflow` (Pi adapter), `codex-taskflow`, `claude-taskflow`, and `opencode-taskflow` (the three delivery packages re-export their runners from `taskflow-hosts` and each ships an MCP bin + plugin/config), all sharing the host-neutral MCP server in `taskflow-mcp-core`. **Library Phase 1**: save flows with `purpose`+`tags` via `taskflow_save` (MCP) or `action=save` (Pi), search them with structural + CJK-aware keyword scoring via `taskflow_search`/`action=search`, and track `reuseCount` via `reusedFromSearch`. **`defineFile`**: pass a `defineFile` path (or `{defineFile, name}`) to `action=run` (Pi) or `taskflow_run`/`taskflow_verify`/`taskflow_compile` (MCP) instead of an inline `define`, and the engine reads the flow from disk — pair it with JSONC comments to annotate saved flows. **JSONC**: flow-definition `.json` files may now carry `//` and `/* */` comments and trailing commas (parsed by `parseJsonc`, re-exported from the `taskflow-core` barrel); LLM-output parsing via `safeParse` stays strict. **Shared Context Tree**: opt-in (`shareContext` / `contextSharing`) blackboard + supervision tools (`ctx_read`/`ctx_write` horizontal reuse, `ctx_report`/`ctx_spawn` vertical supervision); `ctx_spawn` accepts a flat task **or** a dependency-bearing `subflow` (a runtime-validated nested DAG), depth-capped on a unified nesting counter with budget accounting. **Workspace isolation**: a phase's `cwd` accepts reserved keywords `temp`/`dedicated`/`worktree` — the runtime allocates an isolated dir (or a git worktree on a throwaway branch) and tears it down after the phase, fail-open, rejected in LLM-authored sub-flows. **Detached execution**: runs can execute in the background, detached from the Pi session. Prior: loop-until-done (`loop`), tournament (best-of-N with a judge), cross-run memoization (content-addressed cache with git/file/glob/env fingerprints and TTL), interactive `/tf init`, configurable built-in agents, 18 built-in agents with 6 model roles. Full control-flow & reliability layer (`when` guards, `join: any`, `retry`/backoff, `approval`, `flow` composition, `budget` caps, `onBlock: "retry"`, `eval` machine gates, idle watchdog) on top of the DSL + DAG runtime (`agent`/`parallel`/`map`/`gate`/`reduce`). Inline + saved flows, cross-session resume, live progress, and isolated context. A run executes as one streaming tool call.
|
|
899
901
|
|
|
900
902
|
Known boundaries (tracked, bounded — no surprises mid-flow):
|
|
901
903
|
|
package/dist/mcp/server.js.map
CHANGED
|
@@ -1 +1 @@
|
|
|
1
|
-
{"version":3,"file":"server.js","sourceRoot":"","sources":["../../src/mcp/server.ts"],"names":[],"mappings":"AAAA;;;;;;;GAOG;AAEH,OAAO,EACN,eAAe,IAAI,mBAAmB,EACtC,gBAAgB,IAAI,oBAAoB,EACxC,cAAc,IAAI,kBAAkB,GACpC,MAAM,0BAA0B,CAAC;AAElC,OAAO,EAAE,mBAAmB,EAAE,MAAM,gBAAgB,CAAC;AAErD,qEAAqE;AACrE,MAAM,UAAU,gBAAgB,CAAC,GAAW;IAC3C,OAAO,oBAAoB,CAAC,GAAG,EAAE,mBAAmB,CAAC,CAAC;AACvD,CAAC;AAED,sEAAsE;AACtE,MAAM,UAAU,eAAe,CAAC,GAAW;IAC1C,OAAO,mBAAmB,CAAC,GAAG,EAAE,mBAAmB,CAAC,CAAC;AACtD,CAAC;AAED,wEAAwE;AACxE,MAAM,UAAU,cAAc,CAAC,
|
|
1
|
+
{"version":3,"file":"server.js","sourceRoot":"","sources":["../../src/mcp/server.ts"],"names":[],"mappings":"AAAA;;;;;;;GAOG;AAEH,OAAO,EACN,eAAe,IAAI,mBAAmB,EACtC,gBAAgB,IAAI,oBAAoB,EACxC,cAAc,IAAI,kBAAkB,GACpC,MAAM,0BAA0B,CAAC;AAElC,OAAO,EAAE,mBAAmB,EAAE,MAAM,gBAAgB,CAAC;AAErD,qEAAqE;AACrE,MAAM,UAAU,gBAAgB,CAAC,GAAW;IAC3C,OAAO,oBAAoB,CAAC,GAAG,EAAE,mBAAmB,CAAC,CAAC;AACvD,CAAC;AAED,sEAAsE;AACtE,MAAM,UAAU,eAAe,CAAC,GAAW;IAC1C,OAAO,mBAAmB,CAAC,GAAG,EAAE,mBAAmB,CAAC,CAAC;AACtD,CAAC;AAED,wEAAwE;AACxE,MAAM,UAAU,cAAc,CAAC,GAAG,GAAW,OAAO,CAAC,GAAG,EAAE;IACzD,OAAO,kBAAkB,CAAC,mBAAmB,EAAE,GAAG,CAAC,CAAC;AACrD,CAAC"}
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "codex-taskflow",
|
|
3
|
-
"version": "0.1.
|
|
3
|
+
"version": "0.1.8",
|
|
4
4
|
"description": "Run taskflow on OpenAI Codex: a Codex subagent runner plus an MCP server (and a plug-and-play Codex plugin) that exposes the taskflow_* tools to Codex users.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"codex",
|
|
@@ -50,9 +50,9 @@
|
|
|
50
50
|
"access": "public"
|
|
51
51
|
},
|
|
52
52
|
"dependencies": {
|
|
53
|
-
"taskflow-core": "0.1.
|
|
54
|
-
"taskflow-hosts": "0.1.
|
|
55
|
-
"taskflow-mcp-core": "0.1.
|
|
53
|
+
"taskflow-core": "0.1.8",
|
|
54
|
+
"taskflow-hosts": "0.1.8",
|
|
55
|
+
"taskflow-mcp-core": "0.1.8"
|
|
56
56
|
},
|
|
57
57
|
"scripts": {
|
|
58
58
|
"build": "rm -rf dist && tsc -p tsconfig.build.json && node ../../scripts/copy-readme.mjs codex-taskflow"
|