@fyeeme/pi-dynamic-workflows 0.1.1 → 2.0.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,5 +1,12 @@
1
1
  # @fyeeme/pi-dynamic-workflows
2
2
 
3
+ **2.0.0 — major release** (from 1.1.0 on npm), part of the 2.0 extensions family wave:
4
+
5
+ - **Composed sub-agent stack** — the extension factory composes [`@fyeeme/pi-subagents`](https://www.npmjs.com/package/@fyeeme/pi-subagents) 2.1.1 from its npm dependency: the live agent fleet surface and the `subagent` tool work out of the box; no manifest path wiring, no separate install.
6
+ - **Run persistence** — journaled runs under `.pi/workflows/` (journal + manifest): a re-run resumes cached agent calls keyed by `sha256(workflow + prompt + signature)` with zero re-dispatch, and the manifest tracks live state for inspection.
7
+ - **Per-call budget override** — `run_workflow` accepts a `budget` object that merges over the workflow definition's own `maxAgents`/`maxTokens`, making "halve the fan-out for large diffs" an actual parameter.
8
+ - **Authoring aids shipped** — `workflow-author` skill, `/wf-*` prompts, and seed workflows (`review-local-diff`, `review-extension`) using `import type` for load determinism.
9
+
3
10
  **Deterministic TypeScript workflow orchestration for [pi](https://github.com/earendil-works/pi-mono).**
4
11
 
5
12
  Define a workflow as a declarative list of typed steps, run it, and get resumable, budget-bounded, abortable execution. Fuses the pi-dynamic-workflows design (10 step primitives + heuristic planner + outcome collectors) with Claude Code's workflow-engine coordination mechanisms (deterministic sandbox, cache-key resume, per-agent abort, dynamic budget, runaway caps).
@@ -20,16 +27,20 @@ A workflow run is a list of steps (`agent` / `code` / `log` / `fan_out` / `loop_
20
27
 
21
28
  ---
22
29
 
23
- ## Install
30
+ **Batteries included**: this package's extension factory composes
31
+ [`@fyeeme/pi-subagents`](https://www.npmjs.com/package/@fyeeme/pi-subagents)
32
+ (the single below-editor fleet surface with its conversation viewer, and the `subagent` tool)
33
+ from its version-pinned dependency copy, so workflow runs are observable out
34
+ of the box. A standalone pi-subagents install is optional and coexists
35
+ (idempotent composition).
24
36
 
25
- This is a pi extension package (workspace / local), not yet published to npm. From a pi workspace:
37
+ ## Install
26
38
 
27
39
  ```bash
28
- npm install --ignore-scripts # hydrate (the package is a workspace dep)
40
+ pi install npm:@fyeeme/pi-dynamic-workflows
29
41
  ```
30
42
 
31
- This resolves [`@fyeeme/pi-subagent-core`](https://www.npmjs.com/package/@fyeeme/pi-subagent-core)
32
- (`^0.3.0`, from the npm registry — no sibling-repo layout requirement).
43
+ This resolves [`@fyeeme/pi-subagents`](https://www.npmjs.com/package/@fyeeme/pi-subagents) 2.1.1 from the npm registry (the composed sub-agent stack lights up with no separate install).
33
44
 
34
45
  Then import the public API from the package root module (a TypeScript barrel; the package ships `.ts` source):
35
46
 
@@ -37,9 +48,11 @@ Then import the public API from the package root module (a TypeScript barrel; th
37
48
  import { defineWorkflow, runWorkflow } from "@fyeeme/pi-dynamic-workflows/src/index.ts";
38
49
  ```
39
50
 
40
- > The package's `pi.extensions` entry (`./index.ts`) registers the `run_workflow`
41
- > tool and the `/wf-inspect` command. The engine is also fully usable via the
42
- > imports shown here.
51
+ > The package's `pi.extensions` entry registers the `run_workflow` tool plus
52
+ > the shared sub-agent UI from pi-subagents (the below-editor fleet roster —
53
+ > every spawned workflow agent appears there under its stable id; ↓ at an
54
+ > empty editor opens the list, Enter shows an agent's live transcript). The
55
+ > engine is also fully usable via the imports shown here.
43
56
 
44
57
  ---
45
58
 
package/README.zh-CN.md CHANGED
@@ -2,7 +2,7 @@
2
2
 
3
3
  **为 [pi](https://github.com/earendil-works/pi-mono) 打造的确定性 TypeScript 工作流编排。**
4
4
 
5
- 把工作流定义成一份类型化的声明式步骤列表,运行后即可获得**可恢复、受预算约束、可中止**的执行。融合 pi-dynamic-workflows 设计(9 个步骤原语 + 启发式 planner + outcome 收集器)与 Claude Code 工作流引擎的协调机制(确定性沙箱、缓存键恢复、按 agent 中止、动态预算、失控上限)。
5
+ 把工作流定义成一份类型化的声明式步骤列表,运行后即可获得**可恢复、受预算约束、可中止**的执行。融合 pi-dynamic-workflows 设计(10 个步骤原语 + 启发式 planner + outcome 收集器)与 Claude Code 工作流引擎的协调机制(确定性沙箱、缓存键恢复、按 agent 中止、动态预算、失控上限)。
6
6
 
7
7
  语言:[English](README.md) | **中文**
8
8
 
@@ -10,7 +10,7 @@
10
10
 
11
11
  ## 为什么需要它
12
12
 
13
- 一次运行 = 一份步骤列表(`agent` / `code` / `fan_out` / `loop_until` / `adversarial` / `tournament` / `classify_route`)。引擎保证:
13
+ 一次运行 = 一份步骤列表(`agent` / `code` / `log` / `fan_out` / `loop_until` / `loop_until_dry` / `adversarial` / `tournament` / `classify_route` / `sub_workflow`)。引擎保证:
14
14
 
15
15
  - **确定性** —— workflow `.ts` 文件经 AST 守卫,禁止 `Date.now()` / `Math.random()` / `new Date()`;run id 是 `(timestamp, sequence)` 的纯函数。
16
16
  - **恢复即不重派** —— 每个 agent 调用以 `sha256(workflow + prompt + signature)` 为键写入 journal;重跑同一 workflow 会回放缓存的 agent(零子进程派发)。
@@ -20,6 +20,12 @@
20
20
 
21
21
  ---
22
22
 
23
+ **开箱即用**:本包的扩展工厂会组合
24
+ [`@fyeeme/pi-subagents`](https://www.npmjs.com/package/@fyeeme/pi-subagents)
25
+ (实时代理 UI——编辑器上方 widget、FleetView、`/agents`——以及 `subagent`
26
+ 工具),且版本钉定在自身依赖副本上,工作流运行的可观测性开箱即得。
27
+ 单独安装 pi-subagents 是可选的,二者可共存(组合幂等)。
28
+
23
29
  ## 安装
24
30
 
25
31
  这是一个 pi 扩展包(workspace / 本地),尚未发布到 npm。在 pi workspace 中:
@@ -28,7 +34,7 @@
28
34
  npm install --ignore-scripts # 水合(本包是 workspace 依赖)
29
35
  ```
30
36
 
31
- 这会解析 npm registry 上的 [`@fyeeme/pi-subagent-core`](https://www.npmjs.com/package/@fyeeme/pi-subagent-core)(`^0.3.0`,无需保持同级仓库目录结构)。
37
+ 这会解析 npm registry 上的 [`@fyeeme/pi-subagents`](https://www.npmjs.com/package/@fyeeme/pi-subagents)(无需保持同级仓库目录结构)。
32
38
 
33
39
  随后从包根模块导入公共 API(TypeScript barrel,包直接以 `.ts` 源码分发):
34
40
 
@@ -36,7 +42,7 @@ npm install --ignore-scripts # 水合(本包是 workspace 依赖)
36
42
  import { defineWorkflow, runWorkflow } from "@fyeeme/pi-dynamic-workflows/src/index.ts";
37
43
  ```
38
44
 
39
- > 包的 `pi.extensions` 入口(`./index.ts`)目前仍是脚手架——把 `run_workflow` 工具接进 pi 是后续工作。引擎本身已可经上述导入直接使用。
45
+ > 包的 `pi.extensions` 入口注册 `run_workflow` 工具,并接入 pi-subagents 的共享子代理 UI(编辑器上方实时 agent widget、下方 FleetView、`/agents` 转录查看器——每个工作流 agent 以其 step id 出现在其中;需先安装 @fyeeme/pi-subagents,见上方提示)。引擎本身也可经上述导入直接使用。
40
46
 
41
47
  ---
42
48
 
@@ -298,6 +304,8 @@ const result = await runWorkflow({ workflow: wf, cwd: tempDir, now: 1000, dispat
298
304
  | `adversarial` | `produce`、`rubric[]`、`judges?`、`minPass?` | `{ candidate, passed, passCount, judges }` |
299
305
  | `tournament` | `candidates`、`judges`、`produce` | `{ candidates, winner, judges }` |
300
306
  | `classify_route` | `classifier`、`routes: Record<cat, Step[]>`、`fallback?` | `{ category, matched, route, routeStatus }` |
307
+ | `sub_workflow` | `workflow: WorkflowDefinition`、`input?`、`inheritBudget?` | `{ steps, status, workflowName, error }` |
308
+ | `loop_until_dry` | `agent(item, i)`、`keyOf?`、`merge?`、`maxRounds?`、`dryThreshold?` | 发现项组成的数组 |
301
309
 
302
310
  每个步骤都接受 `id`、`retry?: { maxRetries }` 与
303
311
  `onBudgetExhaust?: "throw" | "null"`——`"null"` 下预算耗尽时该步骤返回 `null`
package/index.ts CHANGED
@@ -5,34 +5,31 @@
5
5
  * workflow from within pi. The engine (src/runner) does the work; this entry
6
6
  * only adapts the agent's JSON args into the code-form WorkflowDefinition and
7
7
  * runs it with the default dispatch (real `pi --mode json` subprocesses).
8
+ *
9
+ * Live progress UI is delegated to the shared `@fyeeme/pi-subagents`
10
+ * extension (registered via the `pi.extensions` manifest alongside this
11
+ * entry): every spawned workflow agent notifies the process-global monitor
12
+ * through spawnAgent, rendering in the shared above-editor agent widget, the
13
+ * below-editor FleetView, and the `/agents` transcript viewer. This package
14
+ * ships no widget/command of its own — a former `wf:progress` widget and
15
+ * `/wf-inspect` command were removed in favor of that shared surface.
8
16
  */
9
17
  import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
18
+ import { truncateHead } from "@earendil-works/pi-coding-agent";
19
+ import { StringEnum } from "@earendil-works/pi-ai";
10
20
  import { Type } from "typebox";
11
- import { buildRenderGroups, type PhaseDef } from "./src/ui-groups.ts";
12
- import { BOLD, CYAN, DIM, GREEN, RED, YELLOW, fmtTokens, stepIdOf } from "./src/format.ts";
21
+ import piSubagents from "@fyeeme/pi-subagents";
13
22
  import { defineWorkflow, runWorkflow } from "./src/index.ts";
14
- import type { AgentCallId, Budget, RunResult, StageType, StepContext, StepDefinition, StepResult, StepStats, WorkflowDefinition } from "./src/types.ts";
23
+ import type { Budget, StepContext, StepDefinition, WorkflowDefinition } from "./src/types.ts";
15
24
  import { WorkflowError } from "./src/errors.ts";
16
- import type { AgentLifecycleListeners } from "./src/lifecycle.ts";
17
- import { WorkflowInspect } from "./src/inspect.ts";
18
-
19
- /** Last completed run, exposed to /wf-inspect for interactive review. */
20
- let lastRunResult: RunResult | null = null;
21
-
22
- /** Phases of the most recent run — handed to /wf-inspect so the post-run view
23
- * groups steps the same way the live widget did (C1 consistency). */
24
- let lastPhases: readonly PhaseDef[] | undefined;
25
-
26
- /** Active widget during a run — lets /wf-inspect show a live snapshot
27
- * before the run completes (lastRunResult is only set post-run). */
28
- let activeWidget: { snapshot(): RunResult } | null = null;
25
+ import { discoverWorkflowLibrary, loadLibraryWorkflow } from "./src/library.ts";
29
26
 
30
27
  // ---------------------------------------------------------------------------
31
28
  // Parameter schema (the JSON-serializable workflow subset)
32
29
  // ---------------------------------------------------------------------------
33
30
 
34
31
  const BudgetExhaustPolicy = Type.Optional(
35
- Type.Union([Type.Literal("throw"), Type.Literal("null")], {
32
+ StringEnum(["throw", "null"] as const, {
36
33
  description: "Budget-exhaustion policy for this step: \"throw\" (default) aborts the run; \"null\" degrades this step to a null result so siblings/downstream continue (the run result records degraded steps).",
37
34
  }),
38
35
  );
@@ -49,7 +46,7 @@ const StepSchema = Type.Union([
49
46
  Type.Object({
50
47
  id: Type.String(),
51
48
  type: Type.Literal("log"),
52
- message: Type.String({ description: "Narrative line emitted into the progress widget (zero dispatch / zero tokens)" }),
49
+ message: Type.String({ description: "Narrative line fired via the onLog lifecycle listener (zero dispatch / zero tokens)" }),
53
50
  onBudgetExhaust: BudgetExhaustPolicy,
54
51
  }),
55
52
  Type.Object({
@@ -102,22 +99,38 @@ const BudgetSchema = Type.Object({
102
99
  maxDurationMs: Type.Optional(Type.Number()),
103
100
  });
104
101
 
105
- const PhaseSchema = Type.Object({
106
- title: Type.String({ description: "Phase display name" }),
107
- detail: Type.Optional(Type.String({ description: "Short detail shown next to the phase title" })),
108
- stepIds: Type.Array(Type.String(), { description: "Step ids belonging to this phase" }),
109
- });
110
-
111
102
  const WorkflowSchema = Type.Object({
112
103
  name: Type.String(),
113
104
  description: Type.Optional(Type.String()),
114
105
  steps: Type.Array(StepSchema),
115
106
  budget: Type.Optional(BudgetSchema),
116
- phases: Type.Optional(Type.Array(PhaseSchema, { description: "Group steps into phases for progress-tree UI" })),
117
107
  });
118
108
 
119
109
  const RunWorkflowParams = Type.Object({
120
- workflow: WorkflowSchema,
110
+ source: Type.Optional(
111
+ StringEnum(["inline", "library"] as const, {
112
+ description:
113
+ 'Where the workflow comes from. "inline" (default): the `workflow` parameter (JSON subset). "library": a named workflow from the discovered library (bundled workflows/ + project .pi/workflows/lib/, project overrides bundled) — full step set incl. loop_until, loaded through the determinism guard.',
114
+ default: "inline",
115
+ }),
116
+ ),
117
+ name: Type.Optional(
118
+ Type.String({ description: 'Library workflow name (required for source: "library").' }),
119
+ ),
120
+ workflow: Type.Optional(WorkflowSchema),
121
+ budget: Type.Optional(
122
+ Type.Object(
123
+ {
124
+ maxAgents: Type.Optional(Type.Number()),
125
+ maxTokens: Type.Optional(Type.Number()),
126
+ maxDurationMs: Type.Optional(Type.Number()),
127
+ },
128
+ {
129
+ description:
130
+ "Per-call budget override (library mode only — inline workflows declare budget inside `workflow.budget`). Fields merge over the workflow definition's own budget; use it to tighten a library workflow for a large input without editing the definition. Omitted fields keep the definition's values.",
131
+ },
132
+ ),
133
+ ),
121
134
  input: Type.Optional(Type.String({ description: "Initial ctx.input (also {{input}} in prompts)" })),
122
135
  cwd: Type.Optional(Type.String({ description: "Working dir + journal base. Default: session cwd" })),
123
136
  now: Type.Optional(Type.Number({ description: "Deterministic inception ms (resume seed). Default: Date.now()" })),
@@ -287,234 +300,6 @@ type StepData =
287
300
  onBudgetExhaust?: "throw" | "null";
288
301
  };
289
302
 
290
- // ---------------------------------------------------------------------------
291
- // Progress widget — bridges lifecycle events → TUI setWidget
292
- // ---------------------------------------------------------------------------
293
-
294
- type CallStatus = "running" | "done" | "failed" | "skipped" | "retried" | "cached";
295
-
296
- interface CallInfo {
297
- readonly stepId: string;
298
- readonly status: CallStatus;
299
- readonly tokens: number;
300
- readonly model?: string;
301
- }
302
-
303
- export function buildProgressWidget(
304
- steps: readonly { id: string; type: string }[],
305
- setWidget: (lines: string[] | undefined) => void,
306
- setStatus: (text: string | undefined) => void,
307
- phases?: readonly PhaseDef[],
308
- ): AgentLifecycleListeners & { cleanup(): void; snapshot(): RunResult } {
309
- const start = Date.now();
310
- let lastStreamRender = 0;
311
- // Once cleanup() runs (run finished), late events from the abort window — the
312
- // SIGTERM→SIGKILL grace period during which a child can still emit streamed
313
- // deltas — must not re-create the panel via render()/setWidget.
314
- let disposed = false;
315
- const calls = new Map<AgentCallId, CallInfo>();
316
- // C2: narrative lines emitted by `log` steps, keyed by step id.
317
- const logLines = new Map<string, string>();
318
- // C3: accumulated streaming text per in-flight call (delta chunks from onUpdate).
319
- const streamText = new Map<string, string>();
320
- // Live output capture: callId → the settled agent's final text (from onAgentEnd
321
- // `output`), so a live snapshot shows REAL results, not fabricated progress text.
322
- const callOutputs = new Map<string, string>();
323
- // Every step starts at 0 expected agents; onAgentStart/onAgentCacheHit
324
- // increment as calls fire (fan_out totals emerge at runtime). Pre-seeding
325
- // non-fan_out steps to 1 double-counted (1→2 on start), leaving completed
326
- // single-agent steps stuck showing [1/2].
327
- const expected = new Map(steps.map((s) => [s.id, 0]));
328
-
329
- // C1+D1: phase grouping — shared pure builder (also used by /wf-inspect).
330
- const renderGroups = buildRenderGroups(steps, (s) => s.id, phases);
331
-
332
- // callId format: `${stepId}#${n}` (e.g. "fan#2", "adv#produce") — see src/format.ts stepIdOf.
333
-
334
- function render(): void {
335
- const picons: Record<string, string> = {
336
- done: GREEN("✓"),
337
- failed: RED("✗"),
338
- skipped: YELLOW("⏭"),
339
- running: YELLOW("⏳"),
340
- cached: GREEN("↻"),
341
- };
342
-
343
- const lines: string[] = [];
344
- const renderStep = (s: { id: string; type: string }, indent: boolean): void => {
345
- // C2: a `log` step renders as a distinct narrative line, not an agent row.
346
- if (s.type === "log") {
347
- const msg = logLines.get(s.id);
348
- if (msg !== undefined) lines.push(`${indent ? " " : " "}${DIM(msg)}`);
349
- return;
350
- }
351
- const total = expected.get(s.id) ?? 0;
352
- const entries = [...calls.values()].filter((c) => c.stepId === s.id);
353
- const running = entries.some((c) => c.status === "running");
354
- const failed = entries.filter((c) => c.status === "failed").length;
355
- const skipped = entries.filter((c) => c.status === "skipped").length;
356
- const done = entries.filter((c) => c.status === "done" || c.status === "cached").length;
357
- const cached = entries.filter((c) => c.status === "cached").length;
358
- const tokens = entries.reduce((sum, c) => sum + c.tokens, 0);
359
- const models = new Set(entries.map((c) => c.model).filter((m): m is string => Boolean(m)));
360
-
361
- const icon = failed > 0 ? picons.failed
362
- : skipped > 0 && done === 0 ? picons.skipped
363
- : total > 0 && done >= total ? (cached > 0 ? picons.cached : picons.done)
364
- : running ? picons.running
365
- : DIM("○");
366
-
367
- const progress = total > 1 ? ` [${done}/${total}]` : "";
368
- const tok = tokens > 0 ? ` · ${fmtTokens(tokens)} tok` : "";
369
- // A8/C4: show the serving model when known (single model shown; mixed
370
- // fan_out models collapse to a count to avoid a noisy line).
371
- const modelTag = models.size === 1 ? ` · ${[...models][0]}` : models.size > 1 ? ` · ${models.size} models` : "";
372
- const pad = indent ? " " : " ";
373
- lines.push(`${pad}${icon} ${s.id}${progress}${tok}${modelTag}`);
374
- };
375
-
376
- // Every phase group renders its header — a phase interrupted by ungrouped
377
- // items produces two groups; deduplicating the second header would leave
378
- // an indented, header-less orphan row (review L2).
379
- for (const g of renderGroups) {
380
- if (g.kind === "phase" && g.title) {
381
- lines.push(` ${BOLD(g.title)}${g.detail ? DIM(` — ${g.detail}`) : ""}`);
382
- }
383
- for (const s of g.items) renderStep(s, g.kind === "phase");
384
- }
385
-
386
- // C3: streaming tail — the most-recently-started running call's accumulated
387
- // text, truncated to the last 3 lines, so concurrent fan-out previews only
388
- // the active call instead of flooding the widget.
389
- const running = [...calls.entries()].reverse().find(([id, c]) => c.status === "running" && streamText.has(id));
390
- if (running) {
391
- const [id] = running;
392
- const text = streamText.get(id) ?? "";
393
- const tail = text.split("\n").slice(-3);
394
- for (const ln of tail) {
395
- const clipped = ln.length > 100 ? `${ln.slice(0, 99)}…` : ln;
396
- if (clipped) lines.push(DIM(` ↳ ${clipped}`));
397
- }
398
- }
399
-
400
-
401
- setWidget(lines.length > 0 ? lines : void 0);
402
-
403
- const all = [...calls.values()];
404
- const allDone = all.filter((c) => c.status !== "running").length;
405
- const totalTokens = all.reduce((sum, c) => sum + c.tokens, 0);
406
- const elapsed = ((Date.now() - start) / 1000).toFixed(0);
407
- setStatus(`wf ${allDone}/${calls.size} agents · ${fmtTokens(totalTokens)} tok · ${elapsed}s`);
408
- }
409
-
410
- const record = (callId: string, status: CallStatus, tokens = 0, model?: string): void => {
411
- if (disposed) return; // late settle after cleanup — do not re-create the panel
412
- calls.set(callId, { stepId: stepIdOf(callId), status, tokens, model });
413
- render();
414
- };
415
-
416
- return {
417
- cleanup() {
418
- disposed = true;
419
- calls.clear();
420
- streamText.clear();
421
- setWidget(void 0);
422
- setStatus(void 0);
423
- },
424
- /** Build a synthetic RunResult from the live calls map, so /wf-inspect
425
- * can show in-progress agents before the run finishes. */
426
- snapshot(): RunResult {
427
- const snapSteps: StepResult[] = steps.map((s) => {
428
- const entries = [...calls.values()].filter((c) => c.stepId === s.id);
429
- const total = expected.get(s.id) ?? 0;
430
- const done = entries.filter((c) => c.status === "done" || c.status === "cached").length;
431
- const failed = entries.filter((c) => c.status === "failed").length;
432
- const running = entries.filter((c) => c.status === "running").length;
433
- const cached = entries.filter((c) => c.status === "cached").length;
434
- const tokens = entries.reduce((sum, c) => sum + c.tokens, 0);
435
- const status: StepResult["status"] = s.type === "log"
436
- ? "done" // narrative line, no agent call — never "skipped"
437
- : failed > 0 ? "failed" : done >= total && total > 0 ? "done" : running > 0 ? "running" : "skipped";
438
- // Real outputs from settled calls — NOT the fabricated `[n/m] running`
439
- // progress string (that leaked into the detail pane as fake results).
440
- const settled = [...calls.entries()]
441
- .filter(([callId, c]) => c.stepId === s.id)
442
- .map(([callId]) => callOutputs.get(callId))
443
- .filter((o): o is string => Boolean(o));
444
- const results = settled.length > 0
445
- ? (s.type === "fan_out" ? settled : settled[0])
446
- : undefined;
447
- return {
448
- id: s.id,
449
- type: s.type as StageType,
450
- status,
451
- results,
452
- stats: { tokens, cost: 0, durationMs: 0, agents: entries.length, failures: failed },
453
- };
454
- });
455
- const all = [...calls.values()];
456
- const stats: StepStats = {
457
- tokens: all.reduce((sum, c) => sum + c.tokens, 0),
458
- cost: 0,
459
- durationMs: Date.now() - start,
460
- agents: all.length,
461
- failures: all.filter((c) => c.status === "failed").length,
462
- };
463
- return { runId: "live", status: "completed", steps: snapSteps, stats };
464
- },
465
- onAgentStart(callId) {
466
- const stepId = stepIdOf(callId);
467
- expected.set(stepId, (expected.get(stepId) ?? 0) + 1);
468
- record(callId, "running");
469
- },
470
- onAgentEnd(callId, ok, stats, model, output) {
471
- // A skipped/retried/cached call's subprocess still settles (abort →
472
- // notifyEnd(false)); do not overwrite the already-recorded terminal
473
- // state with a "failed" stamp — the run's own bookkeeping marks the
474
- // step skipped/retried, so the widget must show the same.
475
- const cur = calls.get(callId);
476
- if (cur && cur.status !== "running") {
477
- streamText.delete(callId);
478
- return;
479
- }
480
- record(callId, ok ? "done" : "failed", stats?.tokens ?? 0, model);
481
- streamText.delete(callId); // free the accumulated tail once the call settles
482
- if (output) callOutputs.set(callId, output);
483
- },
484
- onAgentSkip(callId) {
485
- record(callId, "skipped");
486
- },
487
- onAgentRetry(callId) {
488
- record(callId, "retried");
489
- },
490
- onAgentCacheHit(callId) {
491
- const stepId = stepIdOf(callId);
492
- expected.set(stepId, (expected.get(stepId) ?? 0) + 1);
493
- record(callId, "cached");
494
- },
495
- onLog(stepId, message) {
496
- if (disposed) return;
497
- logLines.set(stepId, message);
498
- render();
499
- },
500
- onUpdate(callId, partial) {
501
- if (disposed) return;
502
- // Bound the accumulated tail: the widget only ever renders the last 3
503
- // lines (each clipped to ~100 chars), so keeping the full stream alive
504
- // for the call's duration is pure memory growth on long generations.
505
- streamText.set(callId, ((streamText.get(callId) ?? "") + partial).slice(-4096));
506
- // Throttle: a high-frequency stream (fan_out × many deltas) would otherwise
507
- // trigger a full O(steps×calls) render() per chunk. Bound to ~20fps; the
508
- // final onAgentEnd render always fires, so the settled state is exact.
509
- const nowMs = Date.now();
510
- if (nowMs - lastStreamRender >= 50) {
511
- lastStreamRender = nowMs;
512
- render();
513
- }
514
- },
515
- };
516
- }
517
-
518
303
  // ---------------------------------------------------------------------------
519
304
  // Extension
520
305
  // ---------------------------------------------------------------------------
@@ -529,6 +314,13 @@ function dropInvalidModel(id: string, model: string | undefined, validIds: Set<s
529
314
  }
530
315
 
531
316
  export default function (pi: ExtensionAPI): void {
317
+ // Compose the subagent stack (live agent UI — widget / FleetView /agents —
318
+ // and the `subagent` tool) from the pinned dependency copy. The tool
319
+ // registers exactly once per process (guard in pi-subagents' index.ts):
320
+ // coexists with a standalone pi-subagents install and with other consumers
321
+ // (pi-review) composing it.
322
+ piSubagents(pi);
323
+
532
324
  pi.registerTool({
533
325
  name: "run_workflow",
534
326
  label: "Run workflow",
@@ -543,64 +335,85 @@ export default function (pi: ExtensionAPI): void {
543
335
  parameters: RunWorkflowParams,
544
336
 
545
337
  async execute(_toolCallId, params, signal, _onUpdate, ctx) {
546
- // Scoped-models: when the session restricts models (--models /
547
- // enabledModels), validate step models against that scope; otherwise
548
- // fall back to the full registry. Invalid ids (e.g. "sonnet") are
549
- // dropped so the subprocess uses the default session model.
550
- const scoped = ctx.scopedModels;
551
- const validIds = scoped && scoped.length > 0
552
- ? new Set(scoped.map((s) => s.model.id))
553
- : new Set(ctx.modelRegistry.getAll().map((m) => m.id));
554
- const dropped: string[] = [];
555
- const sanitizedSteps = params.workflow.steps.map((s) => {
556
- const model = "model" in s ? dropInvalidModel(s.id, s.model, validIds, dropped) : undefined;
557
- if (s.type === "classify_route") {
558
- // Route/fallback sub-step models must be sanitized too — the schema
559
- // promises "Invalid ids are dropped", which routeStepToCode otherwise
560
- // passes straight through to the subprocess.
561
- const routes = Object.fromEntries(
562
- Object.entries(s.routes).map(([cat, rs]) => [
563
- cat,
564
- rs.map((r) => ({ ...r, model: dropInvalidModel(`${s.id}.${r.id}`, r.model, validIds, dropped) })),
565
- ]),
566
- ) as typeof s.routes;
567
- const fallback = s.fallback?.map((r) => ({ ...r, model: dropInvalidModel(`${s.id}.${r.id}`, r.model, validIds, dropped) }));
568
- return { ...s, model, routes, fallback };
338
+ // Library mode: resolve the named workflow from the discovered
339
+ // library (ast-guard + jiti via the existing loader), then run it
340
+ // through the same runner as inline mode. Unknown name → list the
341
+ // available workflows instead of spawning anything.
342
+ let workflowDef: WorkflowDefinition | undefined;
343
+ if (params.source === "library") {
344
+ if (!params.name)
345
+ throw new Error('run_workflow: source "library" requires a workflow `name`.');
346
+ try {
347
+ const entry = await loadLibraryWorkflow(params.name, params.cwd ?? ctx.cwd);
348
+ if (!entry) {
349
+ const lib = await discoverWorkflowLibrary(params.cwd ?? ctx.cwd);
350
+ const available = [...lib.values()]
351
+ .map((e) => `- ${e.name}${e.description ? ` — ${e.description}` : ""} (${e.filePath})`)
352
+ .join("\n");
353
+ throw new Error(`run_workflow: no library workflow named "${params.name}". Available:\n${available || "(none)"}`);
354
+ }
355
+ workflowDef = entry.workflow;
356
+ } catch (e) {
357
+ if (e instanceof Error && e.message.startsWith("run_workflow:")) throw e;
358
+ const msg = e instanceof Error ? e.message : String(e);
359
+ throw new Error(`run_workflow failed: ${msg}`);
569
360
  }
570
- return { ...s, model };
571
- });
572
- if (dropped.length > 0) {
573
- ctx.ui.notify(`Invalid model(s) dropped, using default: ${dropped.join(", ")}`, "warning");
361
+ } else {
362
+ if (!params.workflow)
363
+ throw new Error('run_workflow: provide either a `workflow` (inline) or `name` with source "library".');
574
364
  }
575
- const sanitizedWorkflow = { ...params.workflow, steps: sanitizedSteps };
576
- const widget = buildProgressWidget(
577
- sanitizedWorkflow.steps,
578
- (lines) => ctx.ui.setWidget("wf:progress", lines),
579
- (text) => ctx.ui.setStatus("wf:summary", text),
580
- sanitizedWorkflow.phases,
581
- );
582
- activeWidget = widget;
583
- lastPhases = sanitizedWorkflow.phases;
584
- try {
585
- const workflow = buildWorkflow(sanitizedWorkflow);
586
- const listeners: AgentLifecycleListeners = {
587
- onAgentStart: widget.onAgentStart,
588
- onAgentEnd: widget.onAgentEnd,
589
- onAgentSkip: widget.onAgentSkip,
590
- onAgentRetry: widget.onAgentRetry,
591
- onAgentCacheHit: widget.onAgentCacheHit,
592
- onLog: widget.onLog,
593
- onUpdate: widget.onUpdate,
594
- };
365
+
366
+ try {
367
+ let workflow: WorkflowDefinition;
368
+ if (params.source === "library") {
369
+ // Library workflows are authored TS (models already validated at
370
+ // authoring time; the determinism guard ran at load). A per-call
371
+ // `budget` merges over the definition's own budget.
372
+ workflow =
373
+ workflowDef && params.budget
374
+ ? { ...workflowDef, budget: { ...workflowDef.budget, ...params.budget } }
375
+ : workflowDef!;
376
+ } else {
377
+ // Scoped-models: when the session restricts models (--models /
378
+ // enabledModels), validate step models against that scope; otherwise
379
+ // fall back to the full registry. Invalid ids (e.g. "sonnet") are
380
+ // dropped so the subprocess uses the default session model.
381
+ const scoped = ctx.scopedModels;
382
+ const validIds =
383
+ scoped && scoped.length > 0
384
+ ? new Set(scoped.map((s) => s.model.id))
385
+ : new Set(ctx.modelRegistry.getAll().map((m) => m.id));
386
+ const dropped: string[] = [];
387
+ const sanitizedSteps = params.workflow!.steps.map((s) => {
388
+ const model = "model" in s ? dropInvalidModel(s.id, s.model, validIds, dropped) : undefined;
389
+ if (s.type === "classify_route") {
390
+ // Route/fallback sub-step models must be sanitized too — the schema
391
+ // promises "Invalid ids are dropped", which routeStepToCode otherwise
392
+ // passes straight through to the subprocess.
393
+ const routes = Object.fromEntries(
394
+ Object.entries(s.routes).map(([cat, rs]) => [
395
+ cat,
396
+ rs.map((r) => ({ ...r, model: dropInvalidModel(`${s.id}.${r.id}`, r.model, validIds, dropped) })),
397
+ ]),
398
+ ) as typeof s.routes;
399
+ const fallback = s.fallback?.map((r) => ({ ...r, model: dropInvalidModel(`${s.id}.${r.id}`, r.model, validIds, dropped) }));
400
+ return { ...s, model, routes, fallback };
401
+ }
402
+ return { ...s, model };
403
+ });
404
+ if (dropped.length > 0) {
405
+ ctx.ui.notify(`Invalid model(s) dropped, using default: ${dropped.join(", ")}`, "warning");
406
+ }
407
+ const sanitizedWorkflow = { ...params.workflow!, steps: sanitizedSteps };
408
+ workflow = buildWorkflow(sanitizedWorkflow);
409
+ }
595
410
  const result = await runWorkflow({
596
411
  workflow,
597
412
  input: params.input,
598
413
  cwd: params.cwd ?? ctx.cwd,
599
414
  now: params.now ?? Date.now(),
600
415
  signal,
601
- listeners,
602
416
  });
603
- lastRunResult = result;
604
417
 
605
418
  const lines = [
606
419
  `workflow "${workflow.name}" → ${result.status} (run ${result.runId})`,
@@ -608,37 +421,35 @@ export default function (pi: ExtensionAPI): void {
608
421
  `stats: ${result.stats.agents} agent(s), ${result.stats.tokens} tokens, $${result.stats.cost.toFixed(4)}`,
609
422
  ];
610
423
  if (result.error) lines.push(`error: ${result.error}`);
424
+ if (result.journalFile) lines.push(`full run details: ${result.journalFile}`);
425
+
426
+ // Failed/aborted runs are tool errors: throw so pi sets isError and
427
+ // reports the summary to the model (returning a value never sets the
428
+ // error flag). Completed runs with degraded steps stay a normal result.
429
+ const text = truncateHead(lines.join("\n"), { maxLines: 2000, maxBytes: 50_000 }).content;
430
+ if (result.status !== "completed") {
431
+ throw new Error(text);
432
+ }
433
+ const u = result.stats.usage;
611
434
  return {
612
- content: [{ type: "text" as const, text: lines.join("\n") }],
435
+ content: [{ type: "text" as const, text }],
613
436
  details: result,
614
- isError: result.status !== "completed",
437
+ // Surface nested-agent usage so pi's footer //session totals include it.
438
+ usage: u
439
+ ? {
440
+ input: u.input,
441
+ output: u.output,
442
+ cacheRead: u.cacheRead,
443
+ cacheWrite: u.cacheWrite,
444
+ totalTokens: result.stats.tokens,
445
+ cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0, total: result.stats.cost },
446
+ }
447
+ : undefined,
615
448
  };
616
449
  } catch (e) {
617
450
  const msg = e instanceof Error ? e.message : String(e);
618
- return { content: [{ type: "text" as const, text: `run_workflow failed: ${msg}` }], details: { error: msg }, isError: true };
619
- } finally {
620
- activeWidget = null;
621
- ctx.ui.setWidget("wf:progress", void 0);
622
- widget.cleanup();
623
- }
624
- },
625
- });
626
-
627
- pi.registerCommand("wf-inspect", {
628
- description: "Inspect the current/last workflow run (↑↓ select, enter detail, esc exit)",
629
- handler: async (_args, ctx) => {
630
- // Prefer a live snapshot while a run is in progress; fall back to
631
- // the last completed result once the run has finished.
632
- const r = activeWidget?.snapshot() ?? lastRunResult;
633
- if (!r) {
634
- ctx.ui.notify("No workflow run yet — run run_workflow first", "warning");
635
- return;
451
+ throw new Error(`run_workflow failed: ${msg}`);
636
452
  }
637
- await ctx.ui.custom(
638
- (tui, _theme, _kb, done) =>
639
- new WorkflowInspect(r, tui, () => done(undefined), lastPhases),
640
- { overlay: true, overlayOptions: { anchor: "center", width: "90%", maxHeight: "80%" } },
641
- );
642
453
  },
643
454
  });
644
455
  }