pi-durable-subagents 1.0.17 → 1.0.19

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,46 @@
1
1
  # Changelog
2
2
 
3
+ ## 1.0.19
4
+
5
+ - The orchestrator no longer keeps every workflow's pinned origin branch (the
6
+ parent session copied for `context: "fork"`) and input bytes in memory; they
7
+ are read from the pinned files when needed. With about 110 workflows of
8
+ history, the orchestrator heap had grown to 3.7 GB; on a copy of that
9
+ history it now stays near 120 MB.
10
+
11
+ ## 1.0.18
12
+
13
+ - The orchestrator uses far less CPU. Running calls share one process-table
14
+ scan instead of each rescanning `/proc` whenever its call directory
15
+ changed, streamed output no longer triggers whole-journal checks for every
16
+ event, and status/idle checks reuse cached journal views. In a benchmark
17
+ with four streaming calls and 200 other processes, CPU fell from about
18
+ 1.9 cores to 0.07; with 100 workflows of history, from about 1.1 cores to
19
+ 0.03.
20
+ - `pi-durable-subagents restart [--force]` (and the `restart` tool action)
21
+ replaces the orchestrator with the installed version. It refuses while an
22
+ execution is running and lists them (`<wid>/<key>`, age, origin); calls
23
+ waiting for a slot or an answer do not block it. `--force` fences the
24
+ running executions, which resume on the new orchestrator. No execution
25
+ starts between the check and the restart. Against an orchestrator from
26
+ 1.0.17 or earlier, the CLI checks the journals itself and then stops the
27
+ old process. Use it instead of killing the orchestrator.
28
+ - One writer call per worktree is now enforced. A call whose tools include
29
+ `edit` or `write` holds its git worktree from its first execution until it
30
+ ends, also while waiting for an answer and across orchestrator restarts.
31
+ Another writer in that worktree waits in order; its status line shows
32
+ `(waiting for writer lock: <root> held by <wid>/<key>)` and its origin gets
33
+ one `conflict` notice. `writer:false` on a call, `isolation:"worktree"`, or
34
+ `writerLock: "off"` in config.json opt out.
35
+ - `pi-durable-subagents hold <resource> [--shared] [--max-wait s] -- <cmd>`
36
+ runs a command under a resource lease, such as `machine` for benchmarks:
37
+ exclusive holders run one at a time, shared holders together, strictly in
38
+ request order. The lease lasts as long as the command, even if the `hold`
39
+ process is killed, and the command's leftover processes end before it is
40
+ released. Subagents find the command on their PATH; it also works from
41
+ your own shell without an orchestrator. `leases` and status show holders
42
+ and waiters, and a call's status line shows the lease it holds or waits for.
43
+
3
44
  ## 1.0.17
4
45
 
5
46
  - A silent subagent shows one warning instead of separate activity and
package/README.md CHANGED
@@ -56,7 +56,7 @@ the npx cache, so `install-service` refuses to run from there.
56
56
  | Two steers arrive out of order and the second replaces the first | Only the second one applies. |
57
57
  | A step is refused, or a dependency fails | The workflow stops that branch cleanly. Nothing is retried in vain. |
58
58
  | A provider's usage window runs out (`No available accounts`, usage limit, quota exceeded) | Found at the second refusal in a row, while pi is still retrying. A call in a pool continues **in the same session** on the pool's next model (within pi's next retry or two); new calls skip that provider. After 15 minutes the next call that wants it tries it once; when it answers, new calls and new generations use it again. A call with a single model waits for it instead of failing. Billing errors (402, insufficient balance) still fail at once. |
59
- | Two subagents edit the same worktree | A reminder names both calls; neither is blocked or locked. Only observed `edit`/`write` calls count (bash-only writes are not seen). Calls with `isolation: "worktree"` have their own worktrees. |
59
+ | Two subagents would write in the same worktree | Only one runs there at a time. A call that can write (its tools include `edit` or `write`, which pi's default tools do) holds its git worktree's writer lock from its launch until it ends, also while it waits for an answer. Another writer for that worktree waits in order, and status shows `waiting for writer lock: <root> held by <wid>/<key>`. `writer: false` (a call that does not write there), `isolation: "worktree"` and `"writerLock": "off"` opt out. |
60
60
  | A subagent waits for an answer for a long time | It releases its model slot and memory, then resumes exactly once when you answer. |
61
61
 
62
62
  ## Use it
@@ -126,7 +126,9 @@ single file: it cannot `import` or `require` other modules.
126
126
  - `timeoutMs` (active time), `output` (a relative path becomes an artifact);
127
127
  - `schema` (a structured `report`);
128
128
  - `gate` (a command, or `{command, output: "json", schema, timeoutMs}`);
129
- - `isolation: "worktree"`, `context: "fork"`, `budget`.
129
+ - `isolation: "worktree"`, `context: "fork"`, `budget`;
130
+ - `writer` (`false`: the call does not write in its cwd's worktree, so it does
131
+ not take that worktree's writer lock; `true`: it does, whatever its tools).
130
132
 
131
133
  The result has `ok`, `status`, the full `output` text and the structured
132
134
  `data`.
@@ -177,7 +179,8 @@ The main agent is interrupted only when there is something to decide:
177
179
  - a finished workflow;
178
180
  - a stalled subagent (the alert names the command it is running and for how long, so a long silent command reads differently from a stuck call);
179
181
  - an unknown outcome;
180
- - two unfinished calls observed editing the same worktree (a reminder, never a block);
182
+ - a call waiting for another call's writer lock on its worktree (once, with the holder);
183
+ - two unfinished calls observed editing the same worktree when one of them does not take the writer lock (a reminder);
181
184
  - a reached budget.
182
185
 
183
186
  Each one arrives once. A reminder that was already resolved is shown as
@@ -239,6 +242,10 @@ pi-durable-subagents stop-all pause every existing workflow now; journ
239
242
  (runs you start afterwards are not held)
240
243
  pi-durable-subagents prune [wid] [--older-than <days>]
241
244
  delete finished workflows (done, failed, stopped); prints count and bytes freed
245
+ pi-durable-subagents restart [--force] switch to the installed version (see "Updating Durable Subagents")
246
+ pi-durable-subagents hold <resource> [--shared] [--max-wait <s>] [--note <text>] -- <command…>
247
+ run one command while holding a resource lease (see below)
248
+ pi-durable-subagents leases [--json] who holds and who waits for each resource
242
249
  pi-durable-subagents doctor [--json] read-only health check; exits 1 when something needs you
243
250
  pi-durable-subagents install-service optional: run `start` at login and every 30 s (systemd / launchd)
244
251
  pi-durable-subagents uninstall-service
@@ -259,6 +266,45 @@ never comes back. `doctor` shows disk use, workflows by status, the largest
259
266
  journals, parked work, old open questions, and leftovers; each finding
260
267
  comes with one command to fix it.
261
268
 
269
+ ### Resource leases
270
+
271
+ Benchmarks, timing measurements and big builds need the machine to
272
+ themselves. Instead of each subagent polling for an idle machine, wrap the
273
+ command:
274
+
275
+ ```sh
276
+ pi-durable-subagents hold machine -- make bench # exclusive
277
+ pi-durable-subagents hold machine --shared -- npm test # with other shared holders, never with an exclusive one
278
+ pi-durable-subagents hold machine --max-wait 600 --note "frame phase" -- ./measure.sh
279
+ ```
280
+
281
+ - The lease covers one command, not a whole call: a subagent that thinks
282
+ or waits for an answer holds nothing.
283
+ - Requests are served strictly in order. An exclusive request waits for
284
+ everything before it, and keeps later shared requests out (no starvation).
285
+ A waiting `hold` prints who holds the resource; `--max-wait` gives up with
286
+ exit 75 without running the command.
287
+ - The command runs without a shell (write `-- sh -c '…'` for one) in its
288
+ own process group; signals to `hold` go to it and its exit status is
289
+ returned. When it exits, whatever it left in its process group is ended
290
+ before the lease passes on.
291
+ - The lease lives as long as the `hold` process or its command lives, so a
292
+ killed `hold` does not hand the machine over while the command still
293
+ runs. If `hold` is killed and its command has exited, processes the
294
+ command left in its group still hold the lease; the next waiter ends them
295
+ (on macOS, a process that took over the command's pid is waited for, not
296
+ ended). State is one small file per request under
297
+ `$DSA_HOME/leases/<resource>/`; no orchestrator is needed, and the user's
298
+ own shell can take part.
299
+ - Subagents find the command on their `PATH` (the orchestrator puts a shim
300
+ in `$DSA_HOME/bin`), and their leases are tagged with their call:
301
+ `status` shows `lease: machine held by <wid>/<key> …; waiting: …` and
302
+ `(holds lease machine)` / `(waiting for lease machine 3m)` on call lines.
303
+ Tell a subagent in its task to run measurements under
304
+ `pi-durable-subagents hold machine -- …`.
305
+ - Leases are cooperative: processes started without `hold` are not held
306
+ back, and a daemon that leaves the process group is not covered.
307
+
262
308
  ## Configuration
263
309
 
264
310
  State lives in `~/.pi/durable-subagents`; set `DSA_HOME` to move it.
@@ -270,7 +316,8 @@ State lives in `~/.pi/durable-subagents`; set `DSA_HOME` to move it.
270
316
  "onQuit": "pause",
271
317
  "pools": { "fast": ["anthropic/claude-haiku-4-5", "openai/gpt-5-mini"] },
272
318
  "providers": { "anthropic": { "slots": 4 } },
273
- "memory": { "reserveMb": 2048, "perChildMb": 300 }
319
+ "memory": { "reserveMb": 2048, "perChildMb": 300 },
320
+ "writerLock": "queue"
274
321
  }
275
322
  ```
276
323
 
@@ -287,6 +334,11 @@ State lives in `~/.pi/durable-subagents`; set `DSA_HOME` to move it.
287
334
  progress.
288
335
  - **Memory:** new subagents wait while memory is short. Running ones are
289
336
  never stopped for memory.
337
+ - **Writer lock:** `"queue"` (default) runs one writing call per git
338
+ worktree (outside git: per directory) at a time; the others wait in order.
339
+ `"off"` lets them run together and only reminds you of edits seen in the same
340
+ worktree. Writes that do not go through a writing call (your own, or a
341
+ `writer: false` call's bash) are not constrained.
290
342
  - **onQuit:** `"pause"` (default) pauses a session's running workflows when
291
343
  you quit that pi; `"continue"` lets them run on in the background.
292
344
  - **Changes apply without a restart:** the orchestrator re-reads the file
@@ -324,23 +376,35 @@ On load, it checks the pi exports and API methods it uses.
324
376
 
325
377
  ## Updating Durable Subagents
326
378
 
327
- Running work stays on the version it started with: the orchestrator is not
328
- restarted under it. When the orchestrator runs another version than the one a
329
- pi session loaded, that pi says so once, and `status` shows the running
330
- version with a note. The orchestrator exits about 10 s after all work ends,
331
- and the next start runs the new version. To switch sooner without stopping
332
- running calls:
333
-
334
- 1. `drain`: running calls finish and nothing new starts in existing
335
- workflows. A call waiting for your answer still counts as running, and a
336
- workflow started after the drain is not held.
337
- 2. Wait until `status` no longer shows the old version (about 10 s after the
338
- last call ends).
339
- 3. `resume` from a pi session started after the update.
340
-
341
- A `resume` before the old orchestrator exits keeps it running the old
342
- version, and a pi session started before the update still starts the old
343
- version; start a new one.
379
+ Running work stays on the version it started with until you restart the
380
+ orchestrator. When the orchestrator runs another version than the one a pi
381
+ session loaded, that pi says so once, and `status` shows the running version
382
+ with a note. The orchestrator exits about 10 s after all work ends, and the
383
+ next start runs the new version. To switch sooner:
384
+
385
+ ```sh
386
+ pi-durable-subagents restart # or the subagents tool: action "restart"
387
+ ```
388
+
389
+ The orchestrator refuses while any execution runs (a subagent process, or a
390
+ gate before a call's seal) and names each one with its session and age; no new
391
+ execution starts while it decides, so nothing slips in between. Calls waiting
392
+ for your answer (hibernated), waiting for a provider slot, or held by a drain
393
+ do not block it. Otherwise it exits and its successor starts at once from the
394
+ installed files and resumes every workflow: an asker keeps its question, a
395
+ queued call launches on the new version. `restart --force` (tool:
396
+ `force: true`) fences running executions instead of refusing; they resume on
397
+ the new version from their sessions, like after a crash, so a tool call that
398
+ was running is repeated or reported as interrupted.
399
+
400
+ To restart only when the machine is quiet, `drain` first (running calls finish
401
+ and nothing new starts in existing workflows), retry `restart` until it is
402
+ accepted, then `resume`. Never kill the orchestrator process: other sessions'
403
+ running calls would be interrupted without a check. An orchestrator from 1.0.17
404
+ or earlier does not know the restart request; `restart` then checks the
405
+ journals itself and ends it with SIGTERM, which is not atomic: a call launched
406
+ in between is fenced and resumes. A pi session started before the update still
407
+ loads the old extension; start a new one.
344
408
 
345
409
  ## What we do not promise
346
410
 
@@ -2,10 +2,10 @@ import { resolve } from "node:path";
2
2
  import { Type } from "@earendil-works/pi-ai";
3
3
  import { validateCallSpec } from "../../compat/spec.js";
4
4
  import { compileFanout } from "../../compat/fanout.js";
5
- const stepsDoc = "Call specs {agent, task, model?, cwd?, timeoutMs?, output?, schema?, gate?, isolation?, context?, budget?, once?, tools?, skills?, key?}; " +
5
+ const stepsDoc = "Call specs {agent, task, model?, cwd?, timeoutMs?, output?, schema?, gate?, isolation?, context?, budget?, once?, tools?, skills?, writer?, key?}; " +
6
6
  "each call is addressed as '<wid>/<key>', where key is the step's own unique key or else 'tasks:<i>' / 'chain:<i>'.";
7
7
  export const parameters = Type.Object({
8
- action: Type.Optional(Type.Union(["run", "agents", "send", "stop", "revise", "status", "resume", "drain"].map(v => Type.Literal(v)))),
8
+ action: Type.Optional(Type.Union(["run", "agents", "send", "stop", "revise", "status", "resume", "drain", "restart"].map(v => Type.Literal(v)))),
9
9
  workflow: Type.Optional(Type.String()), source: Type.Optional(Type.String()), args: Type.Optional(Type.Unknown()),
10
10
  tasks: Type.Optional(Type.Array(Type.Any(), { description: `Parallel calls. ${stepsDoc}` })),
11
11
  chain: Type.Optional(Type.Array(Type.Any(), { description: `Sequential calls ({previous} = previous output). ${stepsDoc}` })),
@@ -20,9 +20,10 @@ export const parameters = Type.Object({
20
20
  timeoutMs: Type.Optional(Type.Number({ description: "Per-call limit on active time in milliseconds (a number). Omit unless a hard limit is needed; prefer budgets." })),
21
21
  key: Type.Optional(Type.String({ description: "A single agent/task run: the call's key. status with wid: that call's full result." })),
22
22
  full: Type.Optional(Type.Boolean({ description: "status: with wid, the complete workflow detail including every output." })),
23
+ force: Type.Optional(Type.Boolean({ description: "restart: fence running executions instead of refusing (they resume on the new orchestrator)." })),
23
24
  }, { additionalProperties: true });
24
25
  /** Call fields a tasks/chain run applies to every step that does not set its own. */
25
- export const stepDefaults = ["model", "timeoutMs", "budget", "isolation", "context", "tools", "skills", "once"];
26
+ export const stepDefaults = ["model", "timeoutMs", "budget", "isolation", "context", "tools", "skills", "once", "writer"];
26
27
  function string(args, name) {
27
28
  if (typeof args[name] !== "string" || !args[name])
28
29
  throw new Error(`${name} is required`);
@@ -52,7 +53,7 @@ export function request(args, cwd) {
52
53
  args.chain !== undefined, args.workflow !== undefined, args.source !== undefined];
53
54
  const action = args.action === undefined && launchForms.filter(Boolean).length === 1 ? "run" : args.action;
54
55
  if (typeof action !== "string" || !action)
55
- throw new Error("action is required: run, agents, send, stop, revise, status, resume, drain");
56
+ throw new Error("action is required: run, agents, send, stop, revise, status, resume, drain, restart");
56
57
  if (action === "run") {
57
58
  const { action: _, workflow, source, tasks, chain, args: inputs, name, usageBudget, maxCalls, inputs: files, by: _by, ...spec } = args;
58
59
  const choices = [workflow, source, tasks, chain, spec.agent === undefined && spec.task === undefined ? undefined : spec];
@@ -141,5 +142,7 @@ export function request(args, cwd) {
141
142
  return { kind: "resume", body: args.wid !== undefined ? { wid: string(args, "wid") } : typeof args.origin === "string" ? { origin: args.origin } : {} };
142
143
  if (action === "drain")
143
144
  return { kind: "drain", body: {} };
144
- throw new Error(`Unsupported action: ${action}; use run, agents, send, stop, revise, status, resume, or drain`);
145
+ if (action === "restart")
146
+ return { kind: "restart", body: args.force === true ? { force: true } : {} };
147
+ throw new Error(`Unsupported action: ${action}; use run, agents, send, stop, revise, status, resume, drain, or restart`);
145
148
  }
@@ -15,6 +15,8 @@ import { attention, presentText, presented, resolved, unfinishedWorkflow } from
15
15
  import { isLive, pausedElsewhere, runningOrchestrator, statusBrief, statusCallDetail, statusCompactDetail, statusDetail, statusView, widOfRid } from "../orchestrator/snapshot.js";
16
16
  import { parameters, request, sendReceipt } from "./main/tool.js";
17
17
  import { discoverAgents } from "../compat/agents.js";
18
+ import { currentOrchestrator, legacyRestart, waitExit } from "../cli/restart.js";
19
+ import { packageVersion } from "../version.js";
18
20
  let noteSink;
19
21
  /** P16: Queue a UI note for the next boundary without waking the model. */
20
22
  export function presentNote(text) { noteSink?.(text); }
@@ -297,6 +299,20 @@ export function registerMain(pi, ui) {
297
299
  const sessionFile = ctx?.sessionManager.getSessionFile();
298
300
  if (normalized.kind === "run" && sessionFile)
299
301
  normalized.body.origin = { sessionFile, leafId: ctx.sessionManager.getLeafId() };
302
+ if (normalized.kind === "restart") {
303
+ // Only an orchestrator that decides restarts is sent one (an older one would keep it as an invalid inbox file).
304
+ const previous = currentOrchestrator(home), force = args.force === true;
305
+ if (!previous)
306
+ return { applied: true, note: "no orchestrator is running; the next one starts on the installed version when work is submitted" };
307
+ if (!previous.restart) {
308
+ const legacy = legacyRestart(home, previous, force);
309
+ if (!legacy.applied)
310
+ return { applied: false, reason: legacy.reason };
311
+ // It does not start its successor; this session does once it has exited (or its next periodic check would).
312
+ void waitExit(previous, 60_000).then(exited => exited ? serial(() => starter()) : undefined).catch(() => { });
313
+ return { applied: true, note: restartNote(previous) };
314
+ }
315
+ }
300
316
  const sent = await serial(async () => {
301
317
  if (!outbox || stopped)
302
318
  throw new Error("Main session is not active");
@@ -330,6 +346,7 @@ export function registerMain(pi, ui) {
330
346
  }
331
347
  return { submitted: { rid: sent.rid } };
332
348
  }
349
+ const restartNote = (previous) => `orchestrator ${previous.version} (pid ${previous.pid}) exits; the installed version (this pi loaded ${packageVersion()}) starts in its place and resumes every workflow`;
333
350
  ui?.(pi, { home, presentNote, submit: args => submit({ ...args, by: "user" }, ctx?.cwd ?? process.cwd(), undefined, false) });
334
351
  // The model must name a real agent; list the ones this project can use (names are checked again per run).
335
352
  let agents = "";
@@ -343,7 +360,7 @@ export function registerMain(pi, ui) {
343
360
  "run (action optional for exactly one launch form): agent+task; tasks:[call specs] parallel; chain:[call specs] sequential ({previous}); workflow:'./script.js' or source (runs.run(key,spec), runs.all([...]), emit(value), args, runs.input(name)). Optional name, cwd, usageBudget, maxCalls, inputs. With tasks/chain, top-level model, timeoutMs, budget, isolation, context, tools, skills, once are defaults for every step (a step's own value wins); a workflow/source script sets them per runs.run call. timeoutMs is milliseconds of active time (a number); omit it unless a hard limit is needed. Explicit unknown agents are rejected BEFORE creation, with available names; unknown script agents fail only their call.",
344
361
  "agents: list names, descriptions, default models and source for this cwd; use these names for run.",
345
362
  "send to:'<wid>/<key>' (bare '<wid>' only for a single-call workflow): steer on a running call delivers at the next safe point (receipt in status/UI); a steer to a call waiting on its question interrupts the question and the subagent usually asks again — use answer to answer it; sealed → finished:<status> — use kind 'follow-up'. follow-up continues a sealed call as generation g+1 or queues after a running turn; follow-up model:'provider/id' or a pool name runs that generation on it. answer: give the qid (or just the call, or nothing when one question is open); to and rev are filled in. A question that needs the user's decision goes to the user; if you answer one yourself, tell the user what you chose. model ('provider/id' or a pool name — its first model not used up): a running call switches at its next provider request; an asking, hibernated or queued call launches on it when it runs again; the reply's model/effect (next-request|next-execution|next-generation) says which. status model = model actually used by the last request; switching = requested, not used yet; switchFailed = refused. A provider content refusal (ToS/usage policy) fails the call at once, not retried. Unknown targets list valid addresses. replaces:[rid] supersedes an earlier send.",
346
- "stop target:<wid|<wid>/<key>> is terminal stopped (usage and partial edits kept); a sealed call → already-sealed:<status>, a finished workflow → terminal:<status>. drain holds existing workflows reversibly (new runs unaffected); resume [wid] releases held workflows. status: without wid, what runs, asks (with its answer address; hibernated:true holds no slot) or failed, sharedWorktree names calls sharing observed edit/write roots (reminder only), finished workflows one line each, provider slots held/limit, the config in effect and providers whose usage window is used up (avoided until a probe finds them answering again), and the orchestrator version (versionNote when it differs from the loaded one); wid: one workflow, outputs clipped; wid+key: one call's full result; full:true: everything. A run's rid from {submitted:{rid}} works wherever a wid is expected. revise wid + workflow/source/args starts a revision.",
363
+ "stop target:<wid|<wid>/<key>> is terminal stopped (usage and partial edits kept); a sealed call → already-sealed:<status>, a finished workflow → terminal:<status>. drain holds existing workflows reversibly (new runs unaffected); resume [wid] releases held workflows. restart (after an update) replaces the orchestrator with the installed version: refused with busy:<running executions> while any runs, unless force:true (they are fenced and resume); hibernated askers and queued calls do not block it. Never kill the orchestrator process. Commands that need the machine (benchmarks, timing) take a lease: tell the subagent to run them as `pi-durable-subagents hold machine [--shared] -- <command>` (FIFO; status lists lease holders and waiters). status: without wid, what runs, asks (with its answer address; hibernated:true holds no slot) or failed, writerWait: a call queued for its git worktree's writer lock (one call whose tools include edit/write runs per worktree; spec writer:false or isolation:'worktree' opts out), sharedWorktree names calls sharing observed edit/write roots (reminder), lease: a call holding or waiting for a resource lease, finished workflows one line each, provider slots held/limit, the config in effect and providers whose usage window is used up (avoided until a probe finds them answering again), and the orchestrator version (versionNote when it differs from the loaded one); wid: one workflow, outputs clipped; wid+key: one call's full result; full:true: everything. A run's rid from {submitted:{rid}} works wherever a wid is expected. revise wid + workflow/source/args starts a revision.",
347
364
  "Control replies are {applied:true,rid} or {applied:false,reason,rid} when decided; otherwise {submitted:{rid}} after 10s.",
348
365
  ...(agents ? [`Available agents: ${agents}.`] : []),
349
366
  "User sees a summary line above the editor; ↓ on an empty editor (or /subagents) opens the list, Enter watches live OR finished calls (finished transcripts remain on disk) and expands finished workflows. List keys: s steer (paste-capable input), x stop (confirm y), m model, a answer when asked, f follow-up on finished calls; action feedback appears in footer.",
@@ -103,6 +103,10 @@ export async function submit(home, command, target, env = process.env, options =
103
103
  const body = { ...(target ? { wid: target } : {}), ...(options.olderThanDays !== undefined ? { olderThanDays: options.olderThanDays } : {}) };
104
104
  requests.push(await outbox.send("orch", "prune", body));
105
105
  }
106
+ else if (command === "restart") {
107
+ const body = options.force ? { force: true } : {};
108
+ requests.push(await outbox.send("orch", "restart", body));
109
+ }
106
110
  else
107
111
  requests.push(await outbox.send("orch", command, command === "stop" ? { target } : command === "resume" && target ? { wid: target } : {}));
108
112
  // Publish first: even a starter failure leaves a recoverable request and no idle-exit race.
@@ -0,0 +1,184 @@
1
+ // `pi-durable-subagents hold <resource> [--shared] [--max-wait <s>] [--note <text>] -- <command> [args…]`
2
+ // Waits for the lease (strict FIFO), runs the command in its own process group, ends what the command left in that
3
+ // group, then releases. See src/platform/lease.ts for the ticket protocol.
4
+ import { spawn } from "node:child_process";
5
+ import { watch } from "node:fs";
6
+ import { constants } from "node:os";
7
+ import { dsaHome } from "../paths.js";
8
+ import { captureStart } from "../platform/proctable.js";
9
+ import { blockers, enqueue, groupAlive, leaseDir, liveTickets, orphaned, removeTicket, RESOURCE, who, writeTicket } from "../platform/lease.js";
10
+ export const HOLD_USAGE = "usage: pi-durable-subagents hold <resource> [--shared] [--max-wait <seconds>] [--note <text>] -- <command> [args…]";
11
+ /** Exit status when --max-wait expires before the lease is granted (EX_TEMPFAIL). */
12
+ export const WAIT_EXPIRED = 75;
13
+ export function parseHold(args) {
14
+ const split = args.indexOf("--");
15
+ if (split < 0)
16
+ throw new Error(`hold needs -- before the command. ${HOLD_USAGE}`);
17
+ const opts = args.slice(0, split), argv = args.slice(split + 1);
18
+ let resource, mode = "exclusive", maxWaitMs, note;
19
+ for (let i = 0; i < opts.length; i++) {
20
+ const a = opts[i];
21
+ if (a === "--shared")
22
+ mode = "shared";
23
+ else if (a === "--max-wait" || a === "--note") {
24
+ const v = opts[++i];
25
+ if (v === undefined)
26
+ throw new Error(`${a} needs a value. ${HOLD_USAGE}`);
27
+ if (a === "--note")
28
+ note = v;
29
+ else {
30
+ const s = Number(v);
31
+ if (!Number.isFinite(s) || s < 0)
32
+ throw new Error(`--max-wait must be a number of seconds ≥ 0. ${HOLD_USAGE}`);
33
+ maxWaitMs = s * 1000;
34
+ }
35
+ }
36
+ else if (a.startsWith("-"))
37
+ throw new Error(`Unknown option ${a}. ${HOLD_USAGE}`);
38
+ else if (resource === undefined)
39
+ resource = a;
40
+ else
41
+ throw new Error(`Unexpected argument ${a}. ${HOLD_USAGE}`);
42
+ }
43
+ if (!resource)
44
+ throw new Error(`hold needs a resource name. ${HOLD_USAGE}`);
45
+ if (!RESOURCE.test(resource))
46
+ throw new Error(`Invalid resource ${JSON.stringify(resource)}: use 1–64 of A-Z a-z 0-9 . _ -`);
47
+ if (!argv.length)
48
+ throw new Error(`hold needs a command after --. ${HOLD_USAGE}`);
49
+ return { resource, mode, ...(maxWaitMs !== undefined ? { maxWaitMs } : {}), ...(note !== undefined ? { note } : {}), argv };
50
+ }
51
+ const age = (ms) => ms < 60_000 ? `${Math.max(0, Math.round(ms / 1000))}s` : ms < 3_600_000 ? `${Math.round(ms / 60_000)}m` : `${(ms / 3_600_000).toFixed(1)}h`;
52
+ const SIGNALS = ["SIGINT", "SIGTERM", "SIGHUP"];
53
+ /** Run one held command; resolves to the exit status for the process. */
54
+ export async function hold(args, options = {}) {
55
+ const env = options.env ?? process.env, home = dsaHome(env), say = options.stderr ?? (l => process.stderr.write(`${l}\n`));
56
+ const pollMs = options.pollMs ?? 250, graceMs = options.graceMs ?? 2000;
57
+ const ticket = await enqueue(home, {
58
+ resource: args.resource, mode: args.mode,
59
+ wrapper: { pid: process.pid, start: (await captureStart(process.pid)) || undefined },
60
+ argv: args.argv.map(a => a.length > 200 ? `${a.slice(0, 199)}…` : a).slice(0, 32), cwd: process.cwd(),
61
+ ...(args.note ? { note: args.note.slice(0, 300) } : {}), since: Date.now(),
62
+ ...(env.DSA_EXEC ? { exec: env.DSA_EXEC } : {}), ...(env.DSA_CALL ? { call: env.DSA_CALL } : {}),
63
+ });
64
+ // Signals: while waiting they withdraw the request; while the command runs they go to its process group.
65
+ let child, interrupted, wake = () => { };
66
+ const onSignal = (signal) => {
67
+ if (child?.pid) {
68
+ try {
69
+ process.kill(-child.pid, signal);
70
+ }
71
+ catch { /* gone */ }
72
+ }
73
+ else {
74
+ interrupted = signal;
75
+ wake();
76
+ }
77
+ };
78
+ for (const s of SIGNALS)
79
+ process.on(s, onSignal);
80
+ let watcher;
81
+ try {
82
+ // Wait in order.
83
+ const deadline = args.maxWaitMs !== undefined ? ticket.since + args.maxWaitMs : undefined;
84
+ try {
85
+ watcher = watch(leaseDir(home, args.resource), () => wake());
86
+ watcher.on("error", () => { });
87
+ }
88
+ catch { /* polling suffices */ }
89
+ let shown;
90
+ const ended = new Map();
91
+ for (;;) {
92
+ if (interrupted) {
93
+ removeTicket(home, ticket);
94
+ return 128 + (constants.signals[interrupted] ?? 1);
95
+ }
96
+ const ahead = blockers(ticket, liveTickets(home, args.resource));
97
+ if (!ahead.length)
98
+ break;
99
+ const now = Date.now();
100
+ // A holder whose wrapper was killed after its command ended can leave processes in the command's group; the
101
+ // wrapper would have ended them before releasing, so a waiter does: TERM, then KILL after the grace period.
102
+ for (const t of ahead.filter(orphaned)) {
103
+ const first = ended.get(t.seq);
104
+ if (first === undefined)
105
+ say(`hold: ending processes left by ${who(t)} (its hold process is gone)`);
106
+ try {
107
+ process.kill(-t.command.pid, first !== undefined && now - first >= graceMs ? "SIGKILL" : "SIGTERM");
108
+ }
109
+ catch { /* gone */ }
110
+ if (first === undefined)
111
+ ended.set(t.seq, now);
112
+ }
113
+ if (deadline !== undefined && now >= deadline) {
114
+ removeTicket(home, ticket);
115
+ say(`hold: ${args.resource} still held by ${ahead.filter(t => t.grantedAt !== undefined).map(who).join(", ") || who(ahead[0])} after ${age(now - ticket.since)}; not running the command (exit ${WAIT_EXPIRED})`);
116
+ return WAIT_EXPIRED;
117
+ }
118
+ const holders = ahead.filter(t => t.grantedAt !== undefined), key = holders.map(t => t.seq).join(",") || `q${ahead[0].seq}`;
119
+ if (key !== shown) {
120
+ shown = key;
121
+ const by = holders.length ? holders.map(t => `${who(t)} (${t.mode}, ${age(now - (t.grantedAt ?? t.since))})`).join(", ") : `${who(ahead[0])} (waiting)`;
122
+ say(`hold: waiting for ${args.resource} (${args.mode}) — held by ${by}; ${ahead.length} ahead`);
123
+ }
124
+ await new Promise(resolve => { const timer = setTimeout(resolve, pollMs); wake = () => { clearTimeout(timer); resolve(); }; });
125
+ }
126
+ watcher?.close();
127
+ watcher = undefined;
128
+ ticket.grantedAt = Date.now();
129
+ await writeTicket(home, ticket);
130
+ if (interrupted) {
131
+ removeTicket(home, ticket);
132
+ return 128 + (constants.signals[interrupted] ?? 1);
133
+ }
134
+ if (shown)
135
+ say(`hold: ${args.resource} granted after ${age(ticket.grantedAt - ticket.since)}`);
136
+ // Run. stdin stays attached only when it is not a terminal: a background process group reading the terminal stops.
137
+ const c = spawn(args.argv[0], args.argv.slice(1), { stdio: [process.stdin.isTTY ? "ignore" : "inherit", "inherit", "inherit"], detached: true, env });
138
+ child = c;
139
+ const exit = new Promise(resolve => {
140
+ c.once("error", error => { say(`hold: cannot run ${args.argv[0]}: ${error.message}`); resolve(127); });
141
+ c.once("exit", (code, signal) => resolve(code ?? 128 + (signal ? constants.signals[signal] ?? 1 : 1)));
142
+ });
143
+ if (c.pid) {
144
+ ticket.command = { pid: c.pid, start: (await captureStart(c.pid)) || undefined };
145
+ await writeTicket(home, ticket).catch(() => { });
146
+ }
147
+ const status = await exit;
148
+ // End what the command left behind in its group before the lease goes to the next holder.
149
+ if (c.pid && groupAlive(c.pid)) {
150
+ try {
151
+ process.kill(-c.pid, "SIGTERM");
152
+ }
153
+ catch { /* gone */ }
154
+ const until = Date.now() + graceMs;
155
+ while (groupAlive(c.pid) && Date.now() < until)
156
+ await new Promise(r => setTimeout(r, 50));
157
+ if (groupAlive(c.pid)) {
158
+ try {
159
+ process.kill(-c.pid, "SIGKILL");
160
+ }
161
+ catch { /* gone */ }
162
+ }
163
+ for (let i = 0; i < 40 && groupAlive(c.pid); i++)
164
+ await new Promise(r => setTimeout(r, 25));
165
+ }
166
+ // A /proc listing is not atomic (a member can fork and exit while it is read): a last group signal ends any member
167
+ // the checks missed, a newborn included. The group id is still ours while it has members.
168
+ if (c.pid) {
169
+ try {
170
+ process.kill(-c.pid, "SIGKILL");
171
+ }
172
+ catch { /* gone */ }
173
+ }
174
+ removeTicket(home, ticket);
175
+ return status;
176
+ }
177
+ finally {
178
+ watcher?.close();
179
+ for (const s of SIGNALS)
180
+ process.off(s, onSignal);
181
+ }
182
+ }
183
+ /** `leases [--json]`. */
184
+ export { leaseLines, leaseState } from "../platform/lease.js";