pi-harness-delegate 0.4.1 → 0.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -2,7 +2,7 @@
2
2
 
3
3
  [![npm version](https://img.shields.io/npm/v/pi-harness-delegate?logo=npm&color=CB3837)](https://www.npmjs.com/package/pi-harness-delegate) [![CI](https://github.com/yorch/pi-harness-delegate/actions/workflows/ci.yml/badge.svg)](https://github.com/yorch/pi-harness-delegate/actions/workflows/ci.yml) [![Release](https://github.com/yorch/pi-harness-delegate/actions/workflows/release.yml/badge.svg)](https://github.com/yorch/pi-harness-delegate/actions/workflows/release.yml) [![Node](https://img.shields.io/badge/node-26.x-brightgreen?logo=node.js)](https://nodejs.org) [![Bun](https://img.shields.io/badge/bun-1.3.14-black?logo=bun)](https://bun.sh) [![Biome](https://img.shields.io/badge/Biome-2.5.10-60a5fa)](https://biomejs.dev) [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)
4
4
 
5
- Delegate work to **any harness** ([Claude Code](https://github.com/anthropics/claude-code), [Muse](https://github.com/openai/codex), [OpenCode](https://opencode.ai), [Amp](https://ampcode.com)) from the [pi coding agent](https://github.com/badlogic/pi-mono): code reviews, detailed plans, implementation, security audits, docs — or your own custom templates.
5
+ Delegate work to **any harness** ([Claude Code](https://github.com/anthropics/claude-code), [Muse](https://github.com/openai/codex), [OpenCode](https://opencode.ai), [Amp](https://ampcode.com), [Devin](https://devin.ai)) from the [pi coding agent](https://github.com/badlogic/pi-mono): code reviews, detailed plans, implementation, security audits, docs — or your own custom templates.
6
6
 
7
7
  Each harness runs headless in your repo with a normalized permission (`readonly` / `edit` / `danger`). Results stream back live, and token/cost usage feeds into pi's footer stats. Templates are portable — prompt bodies live in `templates/shared/`, harness-specific frontmatter selects the native permission.
8
8
 
@@ -16,7 +16,9 @@ pi install npm:pi-harness-delegate
16
16
  pi install git:github.com/yorch/pi-harness-delegate
17
17
  ```
18
18
 
19
- Requires at least one harness binary on PATH (`claude --version`, `codex --version`, `opencode --version`, `amp --version`). Restart pi (or `/reload`) to activate.
19
+ Requires at least one harness binary on PATH (`claude --version`, `codex --version`, `opencode --version`, `amp --version`, `devin --version`). Restart pi (or `/reload`) to activate.
20
+
21
+ **Devin setup note:** `devin` refuses to run interactively (`devin`, `devin -p`) in a directory you haven't trusted yet — but this extension runs Devin over `devin acp` (see below), and live testing found that transport is **not** gated by workspace trust in the tested version (`3000.6.7`): a fresh, never-touched directory worked over ACP with no refusal and no prompt. This extension never sets Devin's `skip_workspace_trust` config key on your behalf either way — that stays a decision you make interactively, if you ever need it for `devin` itself.
20
22
 
21
23
  ## Usage
22
24
 
@@ -32,6 +34,7 @@ Manual delegation:
32
34
  /claude --mode=security-audit --scope=auth/ … # alias → delegate --harness=claude
33
35
  /opencode plan the cache migration
34
36
  /amp implement the caching layer
37
+ /devin review the new auth flow
35
38
  ```
36
39
 
37
40
  Only the prompt is required. A **harness as first word** and/or **mode as next word** selects them; every `--flag` is optional (harness defaults to `delegate.defaultHarness`, mode to `delegate.defaultMode`, scope to whole repo).
@@ -53,6 +56,8 @@ The `delegate` tool takes: `harness`, `task`, `mode`, `scope` (`diff` = git diff
53
56
  delegate({ harness: "all", mode: "review", scope: "diff" }) # tool call form
54
57
  ```
55
58
 
59
+ - `all` resolves against *detected* harnesses, so Devin joins a fan-out automatically once `devin` is installed — a 5-harness fan-out costs more (and runs one more concurrent process) than the 4-harness one did, budget accordingly.
60
+
56
61
  - `all` resolves to whatever's actually installed (`detectAll()`) — an uninstalled harness is skipped and named in the report, it doesn't fail the run. An explicit list is validated the same way; an unknown name is also reported rather than aborting the rest.
57
62
  - Each harness's run goes through the same `delegate()` engine as a single-harness call and writes its own transcript to its own `~/.pi/agent/delegate/outputs/<harness>/`. Runs are launched together and execute in parallel, bounded by `maxConcurrent` (default `4`, one slot per supported harness) — a run beyond the cap queues for a free slot instead of failing, and the cap is enforced across pi processes, not just this one. **This means fan-out spend is genuinely simultaneous**: with the default cap, a 4-harness fan-out can bill all four at once instead of one after another — budget accordingly (`maxBudgetUsd` still applies per run).
58
63
  - The synthesized report is always ordered by the resolved harness list (e.g. `claude, codex, opencode`), regardless of which harness actually finishes first — it groups each harness's metrics + output and a total spend line (unknown-cost runs called out separately, same as `/delegate status`), assembled mechanically, not by asking a model to summarize.
@@ -78,6 +83,7 @@ delegate({ harness: "all", mode: "review", scope: "diff" }) # tool call form
78
83
  | `codex` | `codex` | `readonly→read-only`, `edit→workspace-write`, `danger→danger-full-access` | `codex exec --json`. Schema-verified against codex-cli 0.149.1; cost is always unmeasured (`null`) on ChatGPT-plan auth. |
79
84
  | `opencode` | `opencode` | `readonly→read-only`, `edit→allow-edit`, `danger→danger` | `opencode run --format json`. Schema-verified against opencode 1.18.16. |
80
85
  | `amp` | `amp` (`omp` alias) | `readonly→read-only`, `edit→workspace`, `danger→danger` | `<binary> -p --mode json`, resolves whichever of `amp`/`omp` is actually on `PATH`. Schema-verified against omp 17.2.9 (Sourcegraph's real Amp CLI is unverified). |
86
+ | `devin` | `devin` | `readonly→plan`, `edit→accept-edits`, `danger→bypass` | Runs `devin acp` — [Agent Client Protocol](https://agentclientprotocol.com) over stdio, not stdout JSONL (see `acp-runner.ts`). Real tool-call ids, a genuine context-window %, and a working `sessionId`/resume via `session/load`. Reports no `$` cost (stays `null`) and no turn count. `model` is wired via `devin acp --model <MODEL>` (fuzzy names, e.g. `opus`); the reported `model` is read back from Devin's own `_cognition.ai/agent_stopped` event rather than echoed from the request, so it reflects what actually ran. Schema-verified against `devin 3000.6.7 (260a97c8)`. |
81
87
 
82
88
  Detect availability: `delegate` checks `harness --version` at startup; missing harnesses hint install instructions.
83
89
 
@@ -0,0 +1,310 @@
1
+ /**
2
+ * Sibling to runner.ts for harnesses whose `transport` is `'acp'` (Agent Client Protocol,
3
+ * https://agentclientprotocol.com — JSON-RPC 2.0, newline-delimited, over stdio). Unlike the
4
+ * stdout harnesses' one-way JSONL stream, ACP is bidirectional and stateful: the runner must
5
+ * drive a handshake (`initialize` -> `session/new` -> `session/set_mode` -> `session/prompt`)
6
+ * and hold stdin open for the session's lifetime — the agent exits on stdin EOF. It must also
7
+ * answer requests the agent sends back to us (permission prompts, fs reads) so the session
8
+ * doesn't hang, since we run non-interactively with a permission mode already negotiated.
9
+ *
10
+ * Exposes the exact `RunHarnessOptions`/`HarnessResult` shape as runner.ts, so `delegate()`
11
+ * can pick either runner from `harness.transport` and everything downstream (transcripts,
12
+ * `ToolCallIndex`, progress overlays, fan-out, spend rollup) is unchanged.
13
+ *
14
+ * Deliberately general: an agent's mode ids and result shape live in its `Harness` (`buildArgs`,
15
+ * `permissionMap`, `parseLine`, `extractResult`) — this file only knows the ACP wire protocol.
16
+ */
17
+ import { spawn } from 'node:child_process';
18
+ import { createInterface } from 'node:readline';
19
+ import { DEFAULT_TIMEOUT_MS, type ParseState, type StreamedResult } from './harnesses/types.ts';
20
+ import type { HarnessResult, RunHarnessOptions } from './runner.ts';
21
+
22
+ /** Bound on the initial handshake (initialize / session/new / session/set_mode) so a hung agent
23
+ * doesn't wedge the whole `timeoutMs` budget before `session/prompt` — the actual work — even starts. */
24
+ const HANDSHAKE_TIMEOUT_MS = 30_000;
25
+
26
+ const PROTOCOL_VERSION = 1;
27
+
28
+ function isRecord(v: unknown): v is Record<string, unknown> {
29
+ return typeof v === 'object' && v !== null && !Array.isArray(v);
30
+ }
31
+
32
+ interface PendingRequest {
33
+ resolve: (result: unknown) => void;
34
+ reject: (err: Error) => void;
35
+ timer?: ReturnType<typeof setTimeout>;
36
+ }
37
+
38
+ export function runAcpHarness(opts: RunHarnessOptions): Promise<HarnessResult> {
39
+ return new Promise((resolve, reject) => {
40
+ const args = opts.harness.buildArgs({
41
+ prompt: opts.prompt,
42
+ cwd: opts.cwd,
43
+ permission: opts.permission,
44
+ nativePermission: opts.nativePermission,
45
+ model: opts.model,
46
+ maxBudgetUsd: opts.maxBudgetUsd,
47
+ addDirs: opts.addDirs,
48
+ resumeSessionId: opts.resumeSessionId,
49
+ });
50
+
51
+ const proc = spawn(opts.harness.binary, args, { cwd: opts.cwd, stdio: ['pipe', 'pipe', 'pipe'] });
52
+
53
+ const state: ParseState = { streamedText: '', activities: [], result: null, _harness: {} };
54
+ let stderr = '';
55
+ let settled = false;
56
+ let firstTokenAt: number | null = null;
57
+ // Resumed sessions replay every prior turn as session/update notifications before the new
58
+ // prompt's — set once session/prompt is actually sent, so replayed text/activities (and a
59
+ // replay-skewed TTFT) never reach the caller. See the handshake IIFE below.
60
+ let promptSent = false;
61
+ const startAt = Date.now();
62
+ const MAX_STREAMED = 5 * 1024 * 1024; // 5MB cap to prevent OOM on compromised harness
63
+ const MAX_ACTIVITIES = 5000;
64
+
65
+ let nextId = 1;
66
+ const pending = new Map<number, PendingRequest>();
67
+
68
+ const finish = (r: StreamedResult) => {
69
+ if (settled) return;
70
+ settled = true;
71
+ clearTimeout(timer);
72
+ rejectAllPending(new Error('session ended'));
73
+ const ttft = firstTokenAt !== null ? firstTokenAt - startAt : r.ttftMs;
74
+ resolve({ ...r, ttftMs: ttft, streamedText: state.streamedText, harness: opts.harness.name });
75
+ };
76
+ const fail = (err: Error) => {
77
+ if (settled) return;
78
+ settled = true;
79
+ clearTimeout(timer);
80
+ rejectAllPending(err);
81
+ reject(err);
82
+ };
83
+ function rejectAllPending(err: Error): void {
84
+ for (const p of pending.values()) {
85
+ clearTimeout(p.timer);
86
+ p.reject(err);
87
+ }
88
+ pending.clear();
89
+ }
90
+
91
+ const writeLine = (msg: Record<string, unknown>): void => {
92
+ try {
93
+ proc.stdin.write(`${JSON.stringify(msg)}\n`);
94
+ } catch {
95
+ // stdin already closed (process exiting) — the pending request(s) time out/reject normally.
96
+ }
97
+ };
98
+
99
+ /** Send a JSON-RPC request and await its response. `timeoutMs` bounds only this request —
100
+ * distinct from the overall run timeout — so a hung handshake step fails fast and clearly.
101
+ * Omitted for `session/prompt`: that's the actual work, already bounded by the overall
102
+ * `timer` below, which kills the process and rejects every pending request on fire. */
103
+ const sendRequest = (method: string, params: unknown, timeoutMs?: number): Promise<unknown> => {
104
+ const id = nextId++;
105
+ return new Promise((res, rej) => {
106
+ const entry: PendingRequest = { resolve: res, reject: rej };
107
+ if (timeoutMs !== undefined) {
108
+ entry.timer = setTimeout(() => {
109
+ pending.delete(id);
110
+ rej(new Error(`${method} timed out after ${timeoutMs}ms`));
111
+ }, timeoutMs);
112
+ }
113
+ pending.set(id, entry);
114
+ writeLine({ jsonrpc: '2.0', id, method, params });
115
+ });
116
+ };
117
+
118
+ /** Respond to a request the agent sent to us. Every request needs a reply or the agent's
119
+ * session hangs waiting for it. */
120
+ const respond = (id: unknown, result: unknown): void => writeLine({ jsonrpc: '2.0', id, result });
121
+ const respondError = (id: unknown, message: string): void =>
122
+ writeLine({ jsonrpc: '2.0', id, error: { code: -32601, message } });
123
+
124
+ /** Handle a request FROM the agent (has both `method` and `id`). We run non-interactively
125
+ * with a permission mode already negotiated, so the safe default is to decline anything
126
+ * not already covered by that mode rather than auto-approve — never observed in the captured
127
+ * fixture this harness was built from, but handled defensively since the spec allows it. */
128
+ const handleServerRequest = (msg: Record<string, unknown>): void => {
129
+ const { id, method, params } = msg;
130
+ if (method === 'session/request_permission' && isRecord(params) && Array.isArray(params.options)) {
131
+ const options = params.options as Array<{ optionId?: unknown; kind?: unknown }>;
132
+ const reject =
133
+ options.find(o => o.kind === 'reject_once') ??
134
+ options.find(o => o.kind === 'reject_always') ??
135
+ options.find(o => typeof o.kind === 'string' && o.kind.startsWith('reject'));
136
+ if (reject && typeof reject.optionId === 'string') {
137
+ respond(id, { outcome: { outcome: 'selected', optionId: reject.optionId } });
138
+ } else {
139
+ respond(id, { outcome: { outcome: 'cancelled' } });
140
+ }
141
+ return;
142
+ }
143
+ // fs/read_text_file, fs/write_text_file, terminal/* etc. — we declare no client capabilities
144
+ // for these in `initialize`, so a well-behaved agent shouldn't ask; decline defensively if one does.
145
+ respondError(id, `${String(method)} not supported by this client`);
146
+ };
147
+
148
+ const rl = createInterface({ input: proc.stdout });
149
+ rl.on('line', line => {
150
+ let msg: unknown;
151
+ try {
152
+ msg = JSON.parse(line);
153
+ } catch {
154
+ return;
155
+ }
156
+ if (!isRecord(msg)) return;
157
+
158
+ if (typeof msg.method === 'string' && msg.id !== undefined) {
159
+ handleServerRequest(msg);
160
+ return;
161
+ }
162
+ if (msg.id !== undefined && 'result' in msg) {
163
+ const entry = pending.get(msg.id as number);
164
+ if (entry) {
165
+ pending.delete(msg.id as number);
166
+ clearTimeout(entry.timer);
167
+ entry.resolve(msg.result);
168
+ }
169
+ // still fall through: `harness.parseLine` may also want to extract activity/result data.
170
+ } else if (msg.id !== undefined && 'error' in msg) {
171
+ const entry = pending.get(msg.id as number);
172
+ if (entry) {
173
+ pending.delete(msg.id as number);
174
+ clearTimeout(entry.timer);
175
+ const err = isRecord(msg.error) ? msg.error : {};
176
+ entry.reject(new Error(typeof err.message === 'string' ? err.message : `${msg.id} failed`));
177
+ }
178
+ return;
179
+ }
180
+
181
+ const outcome = opts.harness.parseLine(line, state);
182
+ // Discard streamed text/activities from anything that arrives before the new session/prompt
183
+ // is sent — on a resume that's the replayed prior conversation, not the new turn's own output.
184
+ if (promptSent && outcome.streamedText) {
185
+ if (firstTokenAt === null) firstTokenAt = Date.now();
186
+ if (state.streamedText.length < MAX_STREAMED) {
187
+ const remaining = MAX_STREAMED - state.streamedText.length;
188
+ const chunk =
189
+ outcome.streamedText.length > remaining
190
+ ? `${outcome.streamedText.slice(0, remaining)} [truncated ${outcome.streamedText.length - remaining} chars]`
191
+ : outcome.streamedText;
192
+ state.streamedText += chunk;
193
+ opts.onStream?.(chunk);
194
+ }
195
+ }
196
+ if (promptSent && outcome.activities) {
197
+ for (const a of outcome.activities) {
198
+ if (state.activities.length < MAX_ACTIVITIES) {
199
+ state.activities.push(a);
200
+ opts.onActivity?.(a);
201
+ }
202
+ }
203
+ }
204
+ if (outcome.result) {
205
+ if (!outcome.result.result) outcome.result.result = state.streamedText;
206
+ state.result = outcome.result;
207
+ }
208
+ });
209
+
210
+ proc.stderr.on('data', (d: Buffer) => (stderr += d.toString()));
211
+ proc.on('close', code => {
212
+ rejectAllPending(new Error(`${opts.harness.binary} exited`));
213
+ if (code !== 0 && !state.result) {
214
+ fail(new Error(stderr.trim() || `${opts.harness.binary} exited with code ${code}`));
215
+ return;
216
+ }
217
+ const final = state.result ?? opts.harness.extractResult(state);
218
+ if (final) {
219
+ if (!final.result) final.result = state.streamedText;
220
+ finish(final);
221
+ } else if (code !== 0) {
222
+ fail(new Error(stderr.trim() || `${opts.harness.binary} exited with code ${code}`));
223
+ } else {
224
+ fail(new Error(`${opts.harness.binary} finished without emitting a result`));
225
+ }
226
+ });
227
+ proc.on('error', err => {
228
+ fail(new Error(`failed to start ${opts.harness.binary}: ${err.message}`));
229
+ });
230
+
231
+ const timer = setTimeout(() => {
232
+ proc.kill('SIGKILL');
233
+ fail(new Error(`${opts.harness.binary} timed out after ${opts.timeoutMs ?? DEFAULT_TIMEOUT_MS}ms`));
234
+ }, opts.timeoutMs ?? DEFAULT_TIMEOUT_MS);
235
+ timer.unref?.();
236
+
237
+ opts.signal?.addEventListener(
238
+ 'abort',
239
+ () => {
240
+ proc.kill('SIGKILL');
241
+ fail(new Error('cancelled'));
242
+ },
243
+ { once: true },
244
+ );
245
+
246
+ // Drive the handshake. `session/prompt` has no separate timeout — it's the actual work,
247
+ // bounded by the overall `timer` above like everything else.
248
+ (async () => {
249
+ const modeId = opts.nativePermission ?? opts.harness.permissionMap?.[opts.permission]?.[0] ?? opts.permission;
250
+ await sendRequest(
251
+ 'initialize',
252
+ {
253
+ protocolVersion: PROTOCOL_VERSION,
254
+ clientCapabilities: {}, // no fs/terminal proxying — decline those requests if asked (see handleServerRequest)
255
+ },
256
+ HANDSHAKE_TIMEOUT_MS,
257
+ );
258
+ if (settled) return;
259
+ const sessionParams = {
260
+ cwd: opts.cwd,
261
+ mcpServers: [],
262
+ ...(opts.addDirs && opts.addDirs.length > 0 ? { additionalDirectories: opts.addDirs } : {}),
263
+ };
264
+ // `session/load` resumes a prior session by id (its response carries no sessionId of its
265
+ // own — the client already has it) and replays prior turns as session/update notifications
266
+ // before the new prompt's; `session/new` mints a fresh one. Verified live: loadSession is
267
+ // advertised in agentCapabilities and a real session/load + follow-up prompt round-trips
268
+ // cleanly, replaying history and continuing the same token-usage accounting.
269
+ let sessionId: string | null;
270
+ if (opts.resumeSessionId) {
271
+ await sendRequest('session/load', { sessionId: opts.resumeSessionId, ...sessionParams }, HANDSHAKE_TIMEOUT_MS);
272
+ sessionId = opts.resumeSessionId;
273
+ } else {
274
+ const newSession = await sendRequest('session/new', sessionParams, HANDSHAKE_TIMEOUT_MS);
275
+ sessionId = isRecord(newSession) && typeof newSession.sessionId === 'string' ? newSession.sessionId : null;
276
+ }
277
+ if (settled) return;
278
+ if (!sessionId) throw new Error('session/new did not return a sessionId');
279
+ // session/load's response carries no sessionId of its own (unlike session/new's) — stash it
280
+ // so the harness's parseLine can still report the real session id on the final result.
281
+ if (opts.resumeSessionId) {
282
+ state._harness ??= {};
283
+ state._harness.sessionId = sessionId;
284
+ }
285
+ await sendRequest('session/set_mode', { sessionId, modeId }, HANDSHAKE_TIMEOUT_MS);
286
+ if (settled) return;
287
+ promptSent = true;
288
+ await sendRequest('session/prompt', { sessionId, prompt: [{ type: 'text', text: opts.prompt }] });
289
+ if (settled) return;
290
+ // The agent doesn't exit on its own once the turn is done — an ACP session can outlive a
291
+ // single prompt (resume, follow-up turns). `delegate()` is one-shot per process, so finish
292
+ // as soon as the prompt response resolves (parseLine already turned it into state.result,
293
+ // synchronously, before this await's continuation runs) and tear the process down ourselves.
294
+ const final = state.result ?? opts.harness.extractResult(state);
295
+ if (final) {
296
+ if (!final.result) final.result = state.streamedText;
297
+ finish(final);
298
+ } else {
299
+ fail(new Error(`${opts.harness.binary} session/prompt completed without emitting a result`));
300
+ }
301
+ proc.kill('SIGKILL');
302
+ })().catch(err => {
303
+ // Every other exit path (timeout, abort, success) kills the child — a rejected handshake
304
+ // step (bad modeId, a JSON-RPC error, a HANDSHAKE_TIMEOUT_MS expiry) must too, or the
305
+ // process leaks: ACP agents only exit on stdin EOF, which nothing else here sends.
306
+ proc.kill('SIGKILL');
307
+ fail(err instanceof Error ? err : new Error(String(err)));
308
+ });
309
+ });
310
+ }
@@ -0,0 +1,206 @@
1
+ /**
2
+ * Devin — runs over the Agent Client Protocol (`devin acp`), not stdout JSONL. `transport: 'acp'`
3
+ * routes it through extensions/acp-runner.ts instead of runner.ts; `buildArgs` here only needs to
4
+ * spawn the ACP server (`devin acp`) — the prompt, permission mode, and session lifecycle are all
5
+ * negotiated over the wire by acp-runner.ts, not passed as CLI flags.
6
+ *
7
+ * `parseLine`/`extractResult` translate raw ACP JSON-RPC lines (one per line, same as any other
8
+ * harness's JSONL) into `ParseOutcome`/`StreamedResult` — this is what makes the JSON-RPC plumbing
9
+ * testable via the same fixture-replay pattern as the stdout harnesses (tests/fixtures.test.ts),
10
+ * without spawning a process: `tests/fixtures/devin-acp.jsonl` is a real captured session.
11
+ *
12
+ * Schema verified against `devin 3000.6.7 (260a97c8)` — see docs/devin-acp-harness-design.md.
13
+ *
14
+ * Workspace trust: the design note this was built from flagged `devin`'s interactive workspace-trust
15
+ * gate as a hazard needing a `detect()` hint. Live verification found that gate applies to `devin -p`
16
+ * / interactive `devin`, but NOT to `devin acp` — confirmed by running the raw `initialize`/`session/new`
17
+ * handshake against a directory never seen by devin before (no `--config` bypass), which succeeded with
18
+ * no refusal, while `devin -p` in the same directory refused. So this harness's actual code path was
19
+ * never gated in the first place; no hint or bypass is needed. If a future devin version starts
20
+ * enforcing trust over ACP too, the generic non-zero-exit path in acp-runner.ts already surfaces
21
+ * whatever refusal message devin prints, same as any other process failure.
22
+ */
23
+ import { execFile } from 'node:child_process';
24
+ import { promisify } from 'node:util';
25
+ import type {
26
+ ActivityEvent,
27
+ BuildArgsOpts,
28
+ Harness,
29
+ NormalizedPermission,
30
+ ParseOutcome,
31
+ ParseState,
32
+ StreamedResult,
33
+ } from './types.ts';
34
+
35
+ const execFileAsync = promisify(execFile);
36
+
37
+ function isRecord(v: unknown): v is Record<string, unknown> {
38
+ return typeof v === 'object' && v !== null && !Array.isArray(v);
39
+ }
40
+
41
+ /** Exact structural match to Claude's tiers (session/new's captured `availableModes`: plan, accept-edits,
42
+ * smart, ask, bypass). `smart`/`ask` stay reachable via the existing `nativePermission` escape hatch. */
43
+ const PERMISSION_MAP: Record<NormalizedPermission, string> = {
44
+ readonly: 'plan',
45
+ edit: 'accept-edits',
46
+ danger: 'bypass',
47
+ };
48
+
49
+ /** Translate one `session/update` notification's `params.update` payload into ParseOutcome deltas. */
50
+ function translateUpdate(update: Record<string, unknown>, state: ParseState): ParseOutcome {
51
+ const activities: ActivityEvent[] = [];
52
+ let streamedText: string | undefined;
53
+ const content = isRecord(update.content) ? update.content : undefined;
54
+ const text = content?.type === 'text' && typeof content.text === 'string' ? content.text : undefined;
55
+
56
+ switch (update.sessionUpdate) {
57
+ case 'agent_message_chunk':
58
+ if (text !== undefined) streamedText = text;
59
+ break;
60
+ case 'agent_thought_chunk':
61
+ if (text !== undefined) activities.push({ kind: 'thinking', chars: text.length });
62
+ break;
63
+ case 'tool_call': {
64
+ if (typeof update.toolCallId !== 'string') break;
65
+ const meta = isRecord(update._meta) ? update._meta : {};
66
+ const inferenceName = meta['cognition.ai/inferenceToolName'];
67
+ const name =
68
+ typeof inferenceName === 'string' ? inferenceName : typeof update.kind === 'string' ? update.kind : 'tool';
69
+ activities.push({
70
+ kind: 'tool_input',
71
+ name,
72
+ input: isRecord(update.rawInput) ? update.rawInput : {},
73
+ id: update.toolCallId,
74
+ });
75
+ break;
76
+ }
77
+ case 'tool_call_update': {
78
+ // Only a terminal status produces a tool_result. A call fires multiple `in_progress`
79
+ // updates for the same toolCallId before its `completed`/`failed` (real fixture: 2-3 per
80
+ // id) — ToolCallIndex.resolve() consumes the pending entry on first match, so an earlier
81
+ // in_progress "result" would eat the id and strand the real completion unattributed.
82
+ if (typeof update.toolCallId !== 'string') break;
83
+ if (update.status === 'completed' || update.status === 'failed') {
84
+ activities.push({ kind: 'tool_result', isError: update.status === 'failed', id: update.toolCallId });
85
+ }
86
+ break;
87
+ }
88
+ case 'usage_update':
89
+ state._harness ??= {};
90
+ if (typeof update.size === 'number') state._harness.contextWindow = update.size;
91
+ break;
92
+ default:
93
+ break;
94
+ }
95
+ return { streamedText, activities };
96
+ }
97
+
98
+ export function parseDevinLine(line: string, state: ParseState): ParseOutcome {
99
+ let o: unknown;
100
+ try {
101
+ o = JSON.parse(line);
102
+ } catch {
103
+ return {};
104
+ }
105
+ if (!isRecord(o)) return {};
106
+
107
+ if (o.method === 'session/update' && isRecord(o.params) && isRecord(o.params.update)) {
108
+ return translateUpdate(o.params.update, state);
109
+ }
110
+
111
+ if (o.method === '_cognition.ai/agent_stopped' && isRecord(o.params) && isRecord(o.params.stats)) {
112
+ // The real model Devin ran, independent of whatever --model was requested — the honest
113
+ // value to report, since `--model` accepts fuzzy names and enterprise config can override it.
114
+ const label = o.params.stats.modelLabel;
115
+ if (typeof label === 'string') {
116
+ state._harness ??= {};
117
+ state._harness.model = label;
118
+ }
119
+ return {};
120
+ }
121
+
122
+ if (isRecord(o.result)) {
123
+ const r = o.result;
124
+ if (typeof r.stopReason === 'string') {
125
+ // session/prompt response — the turn is over, build the final result.
126
+ const u = isRecord(r.usage) ? r.usage : null;
127
+ const harnessState = isRecord(state._harness) ? state._harness : {};
128
+ // Devin's inputTokens already includes cachedReadTokens as a subset (fixture: on every
129
+ // usage_update/prompt result, used === inputTokens + outputTokens exactly, and
130
+ // cachedReadTokens < inputTokens). StreamedUsage follows Claude's convention where
131
+ // inputTokens EXCLUDES cache reads (index.ts sums inputTokens + cacheReadInputTokens into
132
+ // promptTokens) — so subtract the cache-read subset back out here, or promptTokens/context%
133
+ // double-counts it. Math.max guards against a negative if that invariant ever breaks.
134
+ const cacheReadInputTokens = typeof u?.cachedReadTokens === 'number' ? u.cachedReadTokens : 0;
135
+ const rawInputTokens = typeof u?.inputTokens === 'number' ? u.inputTokens : 0;
136
+ const result: StreamedResult = {
137
+ result: state.streamedText,
138
+ // Devin's ACP prompt response carries no error flag of its own — a genuine failure
139
+ // (JSON-RPC `error`, non-zero exit, workspace-trust refusal) surfaces via acp-runner.ts's
140
+ // process-level fail() path instead, same as every other harness's non-zero-exit case.
141
+ isError: false,
142
+ numTurns: null, // not reported for a single prompt turn
143
+ totalCostUsd: null, // Devin reports no $ cost over ACP — honest-metrics convention (#11)
144
+ sessionId: typeof harnessState.sessionId === 'string' ? harnessState.sessionId : null,
145
+ stopReason: r.stopReason,
146
+ permissionDenials: [],
147
+ durationMs: null,
148
+ durationApiMs: null,
149
+ ttftMs: null,
150
+ model: typeof harnessState.model === 'string' ? harnessState.model : null,
151
+ contextWindow: typeof harnessState.contextWindow === 'number' ? harnessState.contextWindow : null,
152
+ maxOutputTokens: null,
153
+ usage: u
154
+ ? {
155
+ inputTokens: Math.max(0, rawInputTokens - cacheReadInputTokens),
156
+ outputTokens: typeof u.outputTokens === 'number' ? u.outputTokens : 0,
157
+ cacheCreationInputTokens: 0, // not reported over ACP
158
+ cacheReadInputTokens,
159
+ }
160
+ : null,
161
+ };
162
+ return { result };
163
+ }
164
+ if (typeof r.sessionId === 'string') {
165
+ // session/new response — stash the id; the prompt response above has no sessionId of its own.
166
+ state._harness ??= {};
167
+ state._harness.sessionId = r.sessionId;
168
+ }
169
+ }
170
+ return {};
171
+ }
172
+
173
+ export const devinHarness: Harness = {
174
+ name: 'devin',
175
+ displayName: 'Devin',
176
+ binary: 'devin',
177
+ transport: 'acp',
178
+ async detect() {
179
+ try {
180
+ const { stdout } = await execFileAsync('devin', ['--version'], { timeout: 5000 });
181
+ return { ok: true, version: stdout.trim() };
182
+ } catch {
183
+ return { ok: false, hint: 'Install the Devin CLI: https://docs.devin.ai/' };
184
+ }
185
+ },
186
+ buildArgs(opts: BuildArgsOpts): string[] {
187
+ // The prompt, permission mode, and session lifecycle are all negotiated over the ACP wire
188
+ // (see acp-runner.ts) — this just launches the ACP server. `--model` is the one real CLI flag
189
+ // `devin acp` accepts (verified: `devin acp --help`); it sets the default model for every new
190
+ // ACP session on this server, and accepts fuzzy names (family slug, alias, or partial name).
191
+ const args = ['acp'];
192
+ if (opts.model) args.push('--model', opts.model);
193
+ return args;
194
+ },
195
+ parseLine(line: string, state: ParseState): ParseOutcome {
196
+ return parseDevinLine(line, state);
197
+ },
198
+ extractResult(state: ParseState): StreamedResult | null {
199
+ return state.result;
200
+ },
201
+ permissionMap: {
202
+ readonly: [PERMISSION_MAP.readonly],
203
+ edit: [PERMISSION_MAP.edit],
204
+ danger: [PERMISSION_MAP.danger],
205
+ },
206
+ };
@@ -1,6 +1,7 @@
1
1
  import { ampHarness } from './amp.ts';
2
2
  import { claudeHarness } from './claude.ts';
3
3
  import { codexHarness } from './codex.ts';
4
+ import { devinHarness } from './devin.ts';
4
5
  import { opencodeHarness } from './opencode.ts';
5
6
  import type { Harness } from './types.ts';
6
7
 
@@ -9,6 +10,7 @@ export const HARNESSES: Record<string, Harness> = {
9
10
  codex: codexHarness,
10
11
  opencode: opencodeHarness,
11
12
  amp: ampHarness,
13
+ devin: devinHarness,
12
14
  };
13
15
 
14
16
  export const ALIASES: Record<string, string> = {
@@ -49,3 +51,30 @@ export function isKnownHarness(name: string): boolean {
49
51
  }
50
52
 
51
53
  export const normalizeHarnessName = resolveHarnessName;
54
+
55
+ /**
56
+ * Legacy danger spellings, kept for templates written before harnesses were partitioned.
57
+ */
58
+ const LEGACY_DANGER_TOKENS = new Set(['bypassPermissions', 'danger-full-access', 'danger']);
59
+
60
+ /**
61
+ * Does this native permission string mean "unrestricted" for this harness?
62
+ *
63
+ * A template can declare any native mode via the escape hatch (`permission: <native>`), and
64
+ * `normalizePermission` files anything unrecognised under `nativePermission` with a normalized
65
+ * tier of `edit`. Without this check, a template declaring `yolo` (amp) or `bypass` (devin) would
66
+ * skip the `allowDangerous` gate entirely and run the harness unsandboxed while `delegate()`
67
+ * recorded the run as `edit` — breaking the invariant that danger is only ever reachable through
68
+ * an explicit per-call `allowDangerous: true`.
69
+ *
70
+ * Matches the harness's own `permissionMap.danger`, joined, so a multi-token danger mode is
71
+ * compared as a whole: opencode's danger is `['build', '--auto']`, and bare `build` is its *edit*
72
+ * token — treating each token separately would wrongly gate legitimate `edit` templates.
73
+ */
74
+ export function isNativeDangerPermission(harness: Harness | undefined, nativePermission: string | undefined): boolean {
75
+ if (!nativePermission) return false;
76
+ const native = nativePermission.trim();
77
+ if (LEGACY_DANGER_TOKENS.has(native)) return true;
78
+ const danger = harness?.permissionMap?.danger;
79
+ return Array.isArray(danger) && danger.length > 0 && native === danger.join(' ');
80
+ }
@@ -74,6 +74,11 @@ export interface DetectResult {
74
74
  hint?: string;
75
75
  }
76
76
 
77
+ /** Which runner drives a harness. Defaults to 'stdout' (extensions/runner.ts) when omitted.
78
+ * 'acp' (extensions/acp-runner.ts) is for Agent Client Protocol agents — a bidirectional
79
+ * JSON-RPC session over stdio rather than a one-way JSONL stream. See AGENTS.md. */
80
+ export type Transport = 'stdout' | 'acp';
81
+
77
82
  export interface Harness {
78
83
  name: string;
79
84
  displayName: string;
@@ -90,6 +95,10 @@ export interface Harness {
90
95
  /** Normalized -> native arg fragments. */
91
96
  permissionMap?: Record<NormalizedPermission, string[]>;
92
97
  permissionHint?: (permission: NormalizedPermission) => string[];
98
+ /** Runner selection. Omitted/'stdout' -> runner.ts (the four existing harnesses); 'acp' -> acp-runner.ts.
99
+ * For 'acp', `permissionMap`'s first element per tier is the ACP session mode id (session/set_mode) —
100
+ * the same field the stdout harnesses use for CLI arg fragments, reused rather than duplicated. */
101
+ transport?: Transport;
93
102
  }
94
103
 
95
104
  export const DEFAULT_TIMEOUT_MS = 600_000;
@@ -3,7 +3,7 @@
3
3
  *
4
4
  * Registers:
5
5
  * - `delegate` tool (primary) + `claude_delegate` alias
6
- * - `/delegate` command (primary) + `/claude`, `/codex`, `/opencode`, `/amp`, `/omp` aliases
6
+ * - `/delegate` command (primary) + `/claude`, `/codex`, `/opencode`, `/amp`, `/omp`, `/devin` aliases
7
7
  *
8
8
  * Templates ship in ../templates/shared + ../templates/<harness>; users add custom ones in
9
9
  * ~/.pi/agent/delegate/templates/<harness>/ (global)
@@ -29,6 +29,7 @@ import {
29
29
  truncateToWidth,
30
30
  } from '@earendil-works/pi-tui';
31
31
  import { Type } from 'typebox';
32
+ import { runAcpHarness } from './acp-runner.ts';
32
33
  import {
33
34
  aggregateSpend,
34
35
  buildFanoutReport,
@@ -64,16 +65,16 @@ import {
64
65
  getHarness,
65
66
  HARNESS_NAMES,
66
67
  isKnownHarness,
68
+ isNativeDangerPermission,
67
69
  resolveHarnessName,
68
70
  } from './harnesses/registry.ts';
69
71
  import type { ActivityEvent, NormalizedPermission } from './harnesses/types.ts';
70
-
71
72
  import { delegationHint, stripMarker } from './hint.ts';
72
73
  import { NotifyBatcher } from './notify.ts';
73
74
  import { type FeedEntry, progressWindow } from './progress.ts';
74
75
  import { formatFanoutChip, multiProgressWindow, type RunRow } from './progress-multi.ts';
75
76
  import { runHarness } from './runner.ts';
76
- import { type DelegateTemplate, loadTemplates } from './templates.ts';
77
+ import { type DelegateTemplate, loadTemplates, resolveNativePermission } from './templates.ts';
77
78
  import { mapClaudeUsage } from './usage.ts';
78
79
 
79
80
  /** Render a possibly-unknown cost — `null` means the harness didn't report one, not a measured $0. */
@@ -551,7 +552,7 @@ async function delegate(
551
552
  // permission: normalized, danger requires explicit per-call allowDangerous:true
552
553
  let permission: NormalizedPermission = template.permission;
553
554
  const nativePerm = template.nativePermission;
554
- const isNativeDanger = !!nativePerm && ['bypassPermissions', 'danger-full-access', 'danger'].includes(nativePerm);
555
+ const isNativeDanger = isNativeDangerPermission(harness, nativePerm);
555
556
  if (template.permission === 'danger' || isNativeDanger) {
556
557
  if (opts.allowDangerous !== true) {
557
558
  throw new Error(
@@ -564,6 +565,9 @@ async function delegate(
564
565
  permission = 'danger';
565
566
  }
566
567
  const permissionForDisplay = nativePerm ?? permission;
568
+ // Dropped when an explicit escalation moved us off the template's own tier — see
569
+ // resolveNativePermission(). Applies to both transports.
570
+ const nativePermissionForRun = resolveNativePermission(template.permission, permission, nativePerm);
567
571
 
568
572
  const model = resolveModelForHarness(config, harnessName, opts.model, template.model);
569
573
  const prompt = buildPrompt(template, task, scopeText, ctx.cwd, harnessName);
@@ -572,7 +576,7 @@ async function delegate(
572
576
  let streamedFull = '';
573
577
  let result: import('./runner.ts').HarnessResult;
574
578
  try {
575
- result = await runHarness({
579
+ const baseRunOpts = {
576
580
  harness,
577
581
  prompt,
578
582
  cwd: ctx.cwd,
@@ -586,15 +590,17 @@ async function delegate(
586
590
  signal: opts.signal,
587
591
  timeoutMs: config.harnesses[harnessName]?.timeoutMs ?? config.timeoutMs,
588
592
  resumeSessionId: opts.sessionId,
589
- onStream: t => {
593
+ onStream: (t: string) => {
590
594
  streamedFull += t;
591
595
  opts.onStream?.(t);
592
596
  },
593
- onActivity: ev => {
597
+ onActivity: (ev: ActivityEvent) => {
594
598
  activityEvents.push(ev);
595
599
  opts.onActivity?.(ev);
596
600
  },
597
- });
601
+ nativePermission: nativePermissionForRun,
602
+ };
603
+ result = harness.transport === 'acp' ? await runAcpHarness(baseRunOpts) : await runHarness(baseRunOpts);
598
604
  } catch (err) {
599
605
  release();
600
606
  if (streamedFull.length > 0) {
@@ -960,11 +966,11 @@ export default function (pi: ExtensionAPI) {
960
966
  name: 'delegate',
961
967
  label: 'Delegate',
962
968
  description:
963
- 'Delegate a task to any harness (claude, codex, opencode, amp) running headless in the repo and return its streamed report (cost, token usage, context %, session id). harness selects the backend (default from config, fallback claude) — pass "all" or a comma list (e.g. "claude,codex") to fan out the same task to several harnesses and get back one comparison report. mode selects a template: review, plan, implement, security-audit, docs, general, or custom — some templates run a host-side check (e.g. "bun test") after the harness exits and report pass/fail as separate evidence; that is configured on the template, not a parameter here. scope restricts work: diff for current git diff, pr for PR diff, path list, or whole repo. sessionId continues a prior session.',
969
+ 'Delegate a task to any harness (claude, codex, opencode, amp, devin) running headless in the repo and return its streamed report (cost, token usage, context %, session id). harness selects the backend (default from config, fallback claude) — pass "all" or a comma list (e.g. "claude,codex") to fan out the same task to several harnesses and get back one comparison report. mode selects a template: review, plan, implement, security-audit, docs, general, or custom — some templates run a host-side check (e.g. "bun test") after the harness exits and report pass/fail as separate evidence; that is configured on the template, not a parameter here. scope restricts work: diff for current git diff, pr for PR diff, path list, or whole repo. sessionId continues a prior session.',
964
970
  promptSnippet: 'Delegate a subtask to a harness and return its report',
965
971
  promptGuidelines: [
966
972
  'delegate runs a harness headless in the working directory and returns a streamed report with cost, token usage, and a session id for follow-ups.',
967
- 'Pass harness (claude|codex|opencode|amp) + focused task string + intent and constraints. Use scope: diff for current git diff, pr for PR diff, path list, or omit for whole repo.',
973
+ 'Pass harness (claude|codex|opencode|amp|devin) + focused task string + intent and constraints. Use scope: diff for current git diff, pr for PR diff, path list, or omit for whole repo.',
968
974
  'mode selects the template and its permission level: review/plan/security-audit are readonly; implement/docs/general are edit. Custom template names also work. Some templates verify their own work (e.g. running tests) automatically after the harness finishes — that is not something you configure here.',
969
975
  'harness: "all" or a comma list (e.g. "codex,opencode") fans the same task out to each detected harness and returns one synthesized comparison report — costs multiply, so only use it when the user actually wants a multi-harness comparison.',
970
976
  'sessionId resumes a previous delegated session instead of starting fresh.',
@@ -974,7 +980,7 @@ export default function (pi: ExtensionAPI) {
974
980
  harness: Type.Optional(
975
981
  Type.String({
976
982
  description:
977
- 'Harness to use: claude, codex, opencode, amp (aliases: omp). "all" or a comma list (e.g. "claude,codex") fans out to each detected harness. Defaults to config defaultHarness.',
983
+ 'Harness to use: claude, codex, opencode, amp (aliases: omp), devin. "all" or a comma list (e.g. "claude,codex") fans out to each detected harness. Defaults to config defaultHarness.',
978
984
  }),
979
985
  ),
980
986
  task: Type.String({ description: 'The task/intent to delegate. Be specific.' }),
@@ -1682,7 +1688,7 @@ export default function (pi: ExtensionAPI) {
1682
1688
  );
1683
1689
  else
1684
1690
  ctx.ui.notify?.(
1685
- 'Usage: /delegate [--harness=claude|codex|opencode|amp|all] [--mode=…] [--model=…] [--scope=…] [--verify=…] <prompt>',
1691
+ 'Usage: /delegate [--harness=claude|codex|opencode|amp|devin|all] [--mode=…] [--model=…] [--scope=…] [--verify=…] <prompt>',
1686
1692
  'warning',
1687
1693
  );
1688
1694
  return;
@@ -1752,7 +1758,7 @@ export default function (pi: ExtensionAPI) {
1752
1758
 
1753
1759
  pi.registerCommand('delegate', {
1754
1760
  description:
1755
- 'Delegate a task to any harness. Usage: /delegate [--harness=claude|codex|opencode|amp|all] [--mode=review|plan|implement|security-audit|docs|general] [--model=...] [--scope=diff|pr|paths] [--verify=<cmd>] [--resume=<id>] <prompt> — or use harness as first word: /delegate codex review <prompt>. harness=all or a comma list (e.g. claude,codex) fans out to every detected harness and returns one comparison report.',
1761
+ 'Delegate a task to any harness. Usage: /delegate [--harness=claude|codex|opencode|amp|devin|all] [--mode=review|plan|implement|security-audit|docs|general] [--model=...] [--scope=diff|pr|paths] [--verify=<cmd>] [--resume=<id>] <prompt> — or use harness as first word: /delegate codex review <prompt>. harness=all or a comma list (e.g. claude,codex) fans out to every detected harness and returns one comparison report.',
1756
1762
  handler: makeHandler(),
1757
1763
  });
1758
1764
  pi.registerCommand('claude', {
@@ -1775,6 +1781,10 @@ export default function (pi: ExtensionAPI) {
1775
1781
  description: 'Alias for /delegate --harness=amp (omp compat). Usage: /omp [--mode=...] <prompt>',
1776
1782
  handler: makeHandler('amp'),
1777
1783
  });
1784
+ pi.registerCommand('devin', {
1785
+ description: 'Alias for /delegate --harness=devin. Usage: /devin [--mode=...] <prompt>',
1786
+ handler: makeHandler('devin'),
1787
+ });
1778
1788
 
1779
1789
  pi.on('input', async (event, _ctx) => {
1780
1790
  if (event.source === 'extension') return { action: 'continue' };
@@ -43,11 +43,11 @@ export function normalizePermission(
43
43
  const lower = raw.trim().toLowerCase();
44
44
  if (lower === 'readonly' || lower === 'read-only' || lower === 'read_only')
45
45
  return { permission: 'readonly', permissionMode: 'plan' };
46
- if (lower === 'edit' || lower === 'acceptEdits' || lower === 'accept-edits')
46
+ if (lower === 'edit' || lower === 'acceptedits' || lower === 'accept-edits')
47
47
  return { permission: 'edit', permissionMode: 'acceptEdits' };
48
48
  if (
49
49
  lower === 'danger' ||
50
- lower === 'bypassPermissions' ||
50
+ lower === 'bypasspermissions' ||
51
51
  lower === 'danger-full-access' ||
52
52
  lower === 'danger_full_access'
53
53
  )
@@ -184,3 +184,21 @@ export function loadTemplates(cwd: string, harnessName?: string): Map<string, De
184
184
  export function loadAllTemplates(cwd: string): Map<string, DelegateTemplate> {
185
185
  return loadTemplates(cwd);
186
186
  }
187
+
188
+ /**
189
+ * Which native permission string (if any) to hand the harness for this run.
190
+ *
191
+ * A template's native escape hatch (`permissionMode`/`sandbox`) applies only while the effective
192
+ * permission is still the template's own. Every harness's `buildArgs` prefers `nativePermission`
193
+ * over the normalized map, so passing it unconditionally would let a template's native mode
194
+ * silently override an explicit `allowDangerous` escalation — a per-call escalation would be
195
+ * quietly downgraded back to whatever the template declared.
196
+ */
197
+ export function resolveNativePermission(
198
+ templatePermission: NormalizedPermission,
199
+ effectivePermission: NormalizedPermission,
200
+ nativePermission: string | undefined,
201
+ ): string | undefined {
202
+ if (!nativePermission) return undefined;
203
+ return effectivePermission === templatePermission ? nativePermission : undefined;
204
+ }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-harness-delegate",
3
- "version": "0.4.1",
3
+ "version": "0.5.0",
4
4
  "description": "Delegate work to any harness (Claude Code, Muse, OpenCode, Amp) from the pi coding agent \u2014 code reviews, plans, implementation, security audits, docs, or your own custom templates.",
5
5
  "type": "module",
6
6
  "packageManager": "bun@1.3.14",
@@ -0,0 +1,10 @@
1
+ ---
2
+ name: docs
3
+ description: Generate or update documentation. Writes files.
4
+ permission: edit
5
+ ---
6
+ You are a technical writer delegated by the pi coding agent.
7
+
8
+ Write clear, accurate documentation matching the project's existing doc style
9
+ (check for README/ADRs/docs conventions first). Cover usage, gotchas, and
10
+ worked examples. Report which files you created or changed.
@@ -0,0 +1,8 @@
1
+ ---
2
+ name: general
3
+ description: General delegation — any task. Bounded to file edits unless allowDangerous is set.
4
+ permission: edit
5
+ ---
6
+ You are a capable engineer delegated by the pi coding agent to handle the
7
+ following task. Work in the current repository. Follow its conventions, run
8
+ available checks, and report what you did and why.
@@ -0,0 +1,12 @@
1
+ ---
2
+ name: implement
3
+ description: Implement a task with file edits (auto-accepted). Runs checks.
4
+ permission: edit
5
+ ---
6
+ You are a senior engineer delegated by the pi coding agent to implement a
7
+ task.
8
+
9
+ Implement the task described below. Follow the repo's existing conventions.
10
+ Run the relevant checks when present (tests, typecheck, lint) and fix what
11
+ breaks. Keep changes minimal and focused on the task. Report what you changed
12
+ and the verification you ran.
@@ -0,0 +1,17 @@
1
+ ---
2
+ name: plan
3
+ description: Produce a detailed implementation plan from an intent. Read-only.
4
+ permission: readonly
5
+ ---
6
+ You are a staff engineer delegated by the pi coding agent to produce a
7
+ detailed implementation plan.
8
+
9
+ Understand the current codebase first (read the relevant files). Then produce:
10
+ 1. Goal and non-goals
11
+ 2. Proposed approach, with alternatives considered and why rejected
12
+ 3. Step-by-step implementation plan: ordered steps, each naming the files to
13
+ touch and what changes in them
14
+ 4. Risks, edge cases, and testing strategy
15
+
16
+ Be concrete and reference actual files/functions in the repo. Do not edit
17
+ files — this is a plan.
@@ -0,0 +1,18 @@
1
+ ---
2
+ name: review
3
+ description: Code review of a scope (git diff, files, or the whole repo). Read-only.
4
+ permission: readonly
5
+ defaultTask: Review the current git diff (staged + unstaged)
6
+ defaultScope: diff
7
+ ---
8
+ You are a senior code reviewer delegated by the pi coding agent.
9
+
10
+ Review the provided scope for:
11
+ - Correctness bugs and edge cases
12
+ - Security issues (injection, auth, secrets, unsafe deserialization)
13
+ - Performance problems
14
+ - Code style and maintainability
15
+
16
+ Be specific: cite `file:line` for every finding. Classify each finding as
17
+ Critical / Major / Minor / Nit. End with a prioritized list of the top
18
+ actions. Do not edit files — this is a review.
@@ -0,0 +1,17 @@
1
+ ---
2
+ name: security-audit
3
+ description: Security audit of a scope. Read-only.
4
+ permission: readonly
5
+ defaultTask: Audit this repository for security vulnerabilities
6
+ ---
7
+ You are a security auditor delegated by the pi coding agent.
8
+
9
+ Audit the provided scope for:
10
+ - Injection (SQL, command, path, template)
11
+ - Authentication / authorization gaps
12
+ - Secrets and credential handling
13
+ - Insecure defaults, unsafe deserialization, XXE, SSRF
14
+ - Dependency and supply-chain risks
15
+
16
+ Cite `file:line` for every finding. Rank findings by severity and
17
+ exploitability. Do not edit files — this is an audit.