pi-durable-subagents 1.0.19 → 1.0.21

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,45 @@
1
1
  # Changelog
2
2
 
3
+ ## 1.0.21
4
+
5
+ 1.0.20 was not published; its changes ship in this release.
6
+
7
+ - Programs can name requests: `run`, `send` and `stop` take `--request <id>`
8
+ (CLI) or `request` (the `subagents` tool). A retry with the same id and the
9
+ same content gets the first outcome and never starts a second workflow or
10
+ follow-up; other content under the id is a `request-conflict` and sends
11
+ nothing. Exit codes: 0 applied, 1 rejected, 3 conflict, 75 not decided yet
12
+ (retry with the same id). `describe --key <id>` (or a wid) reports the
13
+ state, open questions and every call's output in full, what live calls
14
+ wait for and why their last execution was fenced. A pruned workflow leaves
15
+ a tombstone with its final status, request id and digest. See "Driving dsa
16
+ from a program" in the README.
17
+ - A forced restart needs the user's approval in a checkable form: a refused
18
+ `restart` lists the running executions grouped by session (with held
19
+ leases) and a token; `restart --force <token> --reason <text>` fences
20
+ exactly that set and is refused again if it changed. A subagent cannot
21
+ force a restart. The ledger records who forced it and why, and `status`
22
+ shows it for 24 hours.
23
+ - An asker cut off by a restart or crash before its planned hibernation no
24
+ longer loses its question: it hibernates and resumes with the answer
25
+ (also an answer given while dsa restarted), and a `once` step waiting for
26
+ an answer no longer ends as `unknown`.
27
+ - A workflow's done notice names a follow-up still going (queued or running)
28
+ instead of calling it `unknown`.
29
+ - Orchestrator starts are fast again. Each start re-parsed every workflow's
30
+ staged snapshot (tens of megabytes each when it holds a forked parent
31
+ session) before handling any request, so a restart stalled all work for
32
+ over a minute. A verified `pins.json` record per revision now replaces that
33
+ work: on a copy of a home with 117 workflows, recovery takes about 6 s
34
+ instead of 80 s. The first start after the update builds the records once.
35
+ A record that does not verify falls back to the snapshot, so a changed
36
+ pinned file is still reported as a conflict.
37
+ - `restart` no longer reports a failure while the orchestrator is still
38
+ recovering: it says the request is submitted and keeps waiting (up to 10
39
+ minutes), then exits 75 with "still pending — do not resubmit" if no
40
+ orchestrator reached it. After the restart it waits for the orchestrator
41
+ that actually decided it.
42
+
3
43
  ## 1.0.19
4
44
 
5
45
  - The orchestrator no longer keeps every workflow's pinned origin branch (the
package/README.md CHANGED
@@ -57,7 +57,9 @@ the npx cache, so `install-service` refuses to run from there.
57
57
  | A step is refused, or a dependency fails | The workflow stops that branch cleanly. Nothing is retried in vain. |
58
58
  | A provider's usage window runs out (`No available accounts`, usage limit, quota exceeded) | Found at the second refusal in a row, while pi is still retrying. A call in a pool continues **in the same session** on the pool's next model (within pi's next retry or two); new calls skip that provider. After 15 minutes the next call that wants it tries it once; when it answers, new calls and new generations use it again. A call with a single model waits for it instead of failing. Billing errors (402, insufficient balance) still fail at once. |
59
59
  | Two subagents would write in the same worktree | Only one runs there at a time. A call that can write (its tools include `edit` or `write`, which pi's default tools do) holds its git worktree's writer lock from its launch until it ends, also while it waits for an answer. Another writer for that worktree waits in order, and status shows `waiting for writer lock: <root> held by <wid>/<key>`. `writer: false` (a call that does not write there), `isolation: "worktree"` and `"writerLock": "off"` opt out. |
60
- | A subagent waits for an answer for a long time | It releases its model slot and memory, then resumes exactly once when you answer. |
60
+ | A subagent waits for an answer for a long time | It releases its model slot and memory, then resumes exactly once when you answer. The question survives orchestrator restarts (also forced ones) and crashes, including one that hits before the subagent released its slot. |
61
+ | A subagent's work ends (finished, stopped, or cut off) | Every process its tools started ends with that execution, also ones started with `nohup`, `setsid` or `&`: they carry the execution's tag (see the limit below). Anything that must outlive the subagent has to be started by you or the parent session. A command run under `hold` is no exception: a forced restart stops it and its lease is released. |
62
+ | A subagent runs in a worktree you made (`isolation: "none"`, the default, with `cwd`) | Durable Subagents never creates, cleans, moves or deletes that directory or its branch, also not on `prune`; `prune` deletes only its own state and the worktrees it created for `isolation: "worktree"`. |
61
63
 
62
64
  ## Use it
63
65
 
@@ -238,11 +240,18 @@ pi-durable-subagents start start the orchestrator if work is pendin
238
240
  pi-durable-subagents resume [wid] continue unfinished or parked work (undoes drain / stop-all)
239
241
  pi-durable-subagents drain hold existing workflows: running calls finish, nothing new starts in them
240
242
  pi-durable-subagents stop <wid|call>
243
+ pi-durable-subagents run --request <id> --spec <file|-> [--cwd <dir>] [--json] [--wait-ms <n>]
244
+ start a run under a caller-chosen id; safe to retry (see below)
245
+ pi-durable-subagents send --request <id> --to <run-id|wid/key> --kind follow-up|answer|steer|model
246
+ [--call <key>] [--qid <qid> --rev <n>] --message <text|@file> [--model <m>] [--json]
247
+ pi-durable-subagents stop --request <id> <run-id|wid|wid/key|call> [--json]
248
+ pi-durable-subagents describe --key <run-id> | <wid> [--json]
249
+ full read-only state of one run (never starts the orchestrator)
241
250
  pi-durable-subagents stop-all pause every existing workflow now; journals stay resumable
242
251
  (runs you start afterwards are not held)
243
252
  pi-durable-subagents prune [wid] [--older-than <days>]
244
253
  delete finished workflows (done, failed, stopped); prints count and bytes freed
245
- pi-durable-subagents restart [--force] switch to the installed version (see "Updating Durable Subagents")
254
+ pi-durable-subagents restart [--force <token> --reason <text>] switch to the installed version (see "Updating Durable Subagents")
246
255
  pi-durable-subagents hold <resource> [--shared] [--max-wait <s>] [--note <text>] -- <command…>
247
256
  run one command while holding a resource lease (see below)
248
257
  pi-durable-subagents leases [--json] who holds and who waits for each resource
@@ -255,6 +264,67 @@ The service only runs `start`: it never resumes work you drained or
255
264
  stopped. Install the CLI globally (`npm i -g pi-durable-subagents`) before
256
265
  `install-service`.
257
266
 
267
+ ### Driving dsa from a program
268
+
269
+ A program (a CI job, a script, another agent) should name its requests with
270
+ `--request <id>`: 1–124 characters `[A-Za-z0-9][A-Za-z0-9._:-]*`, unique per
271
+ `DSA_HOME` across `run`, `send` and `stop`. The id is the identity: the first
272
+ submission under an id is decided once and every retry with **the same
273
+ content** gets that first outcome; a retry never starts a second workflow or
274
+ a second follow-up. The run id is also the workflow key: `send --to <run-id>`
275
+ and `describe --key <run-id>` find the workflow without knowing its wid.
276
+
277
+ The content is hashed into `spec_digest` (the request kind, its body and, for
278
+ answers, the question id and revision). Persist the exact spec bytes you
279
+ submit and retry with those bytes: the run's `cwd` is part of the content
280
+ (`spec.cwd`, else `--cwd`, else the current directory, made absolute), so
281
+ retry from the same place or pass it explicitly. The `spec` file is the
282
+ `subagents` tool's run form (`{agent, task, model?, schema?, …}` or
283
+ `{tasks: […], name?, usageBudget?, maxCalls?}`); `context: "fork"` needs a pi
284
+ session and is rejected.
285
+
286
+ | exit | meaning | `--json` reply |
287
+ | --- | --- | --- |
288
+ | 0 | decided and applied (a retry gets the same answer) | run: `{request, wid, created, spec_digest}`; send/stop: `{request, applied, generation?, call?, spec_digest}` |
289
+ | 1 | decided and rejected (`reason`), or a usage error | `{request, applied: false, reason, spec_digest}` |
290
+ | 3 | `request-conflict`: the id already names other content; nothing was sent | `{request, error, wid?, spec_digest, state}` (the original's digest and state) |
291
+ | 75 | not decided yet: submitted but undecided in `--wait-ms` (default 60 s), submitted and then a later step failed (`reason` says which), or a lock was busy (`reason: "busy"`); retry with the same id | `{request, pending: true, reason?}` |
292
+
293
+ `created` is false when the id had already been decided before the command
294
+ ran. A conflicting id stays conflicting forever, also after `prune`: the
295
+ tombstone keeps the final status, the id and the digest. Any sender's retry
296
+ completes a submission another one recorded but did not publish (it died in
297
+ between), so a retry with the same content always converges.
298
+
299
+ From the `subagents` tool, `request` works the same, with one difference: a
300
+ run's origin session (what `context: "fork"` copies, and where notices go) is
301
+ not part of the digest. A retry of the same id from another session therefore
302
+ gets the first session's workflow, its notices and its forked context. A
303
+ retried answer may leave out `to`, `qid` and `rev`: it addresses the question
304
+ the first attempt answered.
305
+
306
+ `describe` reports one of `absent` (never seen), `pending` (submitted, not
307
+ decided — typically no orchestrator is running; `start` or a retry starts it),
308
+ `rejected` (with `reason`), `running`, `asking` (open questions with their
309
+ full text, `qid`, `rev` and the `to` address to answer), `sealed` (finished:
310
+ `status` plus every call's unclipped `output`, `error` and schema `data`) or
311
+ `pruned` (`pruned: {status, endedAt}`; workflows pruned before 1.0.21 have
312
+ only `endedAt`), with `wid`, `request` and
313
+ `spec_digest`. Live calls also show what they wait for (slot, writer lock,
314
+ lease, exhausted provider) and the last fence of their execution
315
+ (`lastFence: {at, exec, reason}`); the reason is a best-effort reading of the
316
+ journals: `restart-force` when a forced restart listed the execution,
317
+ `orchestrator-crash` when the orchestrator died uncleanly while it ran, else
318
+ `process-died`.
319
+
320
+ ```sh
321
+ pi-durable-subagents run --request build-42 --spec build-42.json --json
322
+ # exit 75: retry later with the same id and bytes
323
+ pi-durable-subagents describe --key build-42 --json
324
+ pi-durable-subagents send --request build-42-a1 --to build-42 --kind answer \
325
+ --qid <qid> --rev <rev> --message "yes"
326
+ ```
327
+
258
328
  ### Housekeeping
259
329
 
260
330
  Journals are never compacted, so state only grows. `prune` removes finished
@@ -387,23 +457,45 @@ pi-durable-subagents restart # or the subagents tool: action "restart"
387
457
  ```
388
458
 
389
459
  The orchestrator refuses while any execution runs (a subagent process, or a
390
- gate before a call's seal) and names each one with its session and age; no new
391
- execution starts while it decides, so nothing slips in between. Calls waiting
460
+ gate before a call's seal). The refusal groups executions by session with ages,
461
+ lease annotations and a token for that exact set; no new execution starts while
462
+ it decides, so nothing slips in between. Calls waiting
392
463
  for your answer (hibernated), waiting for a provider slot, or held by a drain
393
464
  do not block it. Otherwise it exits and its successor starts at once from the
394
465
  installed files and resumes every workflow: an asker keeps its question, a
395
- queued call launches on the new version. `restart --force` (tool:
396
- `force: true`) fences running executions instead of refusing; they resume on
397
- the new version from their sessions, like after a crash, so a tool call that
398
- was running is repeated or reported as interrupted.
466
+ queued call launches on the new version.
467
+
468
+ To interrupt those executions deliberately, first show the user the refusal's
469
+ list and obtain their explicit approval, then use its token and a non-empty
470
+ reason (at most 500 characters):
471
+
472
+ ```sh
473
+ pi-durable-subagents restart --force <token> --reason "<why>"
474
+ # tool: {action:"restart", force:"<token>", reason:"<why>"}
475
+ ```
476
+
477
+ A changed execution set is refused with a fresh list and token. With no live
478
+ executions no token is needed. Bare force cannot fence live executions, and the
479
+ tool rejects `force:true`. Subagents cannot force a restart, even from bash:
480
+ it would fence themselves and other sessions' work. Force fences running
481
+ executions; they resume on the new version from their sessions, like after a
482
+ crash, so a tool call that was running is repeated or reported as interrupted.
483
+ The restart ledger records the reason and initiator; after the next start,
484
+ `status` shows who forced it and why for 24 hours.
485
+
486
+ An orchestrator reads requests only after it has recovered its workflows, so
487
+ a `restart` sent to one that just started waits for it (and says so); if no
488
+ orchestrator reaches the request in time, `restart` exits 75 and the request
489
+ stays pending — it is decided later, so do not send another.
399
490
 
400
491
  To restart only when the machine is quiet, `drain` first (running calls finish
401
492
  and nothing new starts in existing workflows), retry `restart` until it is
402
493
  accepted, then `resume`. Never kill the orchestrator process: other sessions'
403
494
  running calls would be interrupted without a check. An orchestrator from 1.0.17
404
495
  or earlier does not know the restart request; `restart` then checks the
405
- journals itself and ends it with SIGTERM, which is not atomic: a call launched
406
- in between is fenced and resumes. A pi session started before the update still
496
+ journals itself with the same token, reason and subagent checks and ends it with
497
+ SIGTERM, which is not atomic: a call launched in between is fenced and resumes.
498
+ The old orchestrator cannot record the new audit fields. A pi session started before the update still
407
499
  loads the old extension; start a new one.
408
500
 
409
501
  ## What we do not promise
@@ -419,8 +511,15 @@ loads the old extension; start a new one.
419
511
  "outcome unknown" notice; other work keeps running.
420
512
  - A tool that already ran inside a subagent may run again after a crash, if
421
513
  its result never reached the session. Make external side effects
422
- idempotent, or mark the step `once: true` (it then stops as `unknown`
423
- instead of repeating).
514
+ idempotent, or mark the step `once: true`: when an execution is cut off
515
+ while a tool call is running (its result never arrived), the step then
516
+ ends as `unknown` instead of repeating it. Cut off between tool calls, a
517
+ `once` step continues like any other (nothing was left half done). A
518
+ step cut off while its only unfinished tool call is the question it asked
519
+ is not `unknown` either: it keeps waiting and resumes with the answer, also
520
+ an answer given while dsa restarted. If it was cut off while another tool
521
+ call ran beside the question, a `once` step still ends as `unknown`. A follow-up on an `unknown` step continues the
522
+ same session as its next generation.
424
523
  - After a crash, the model call that was in flight is paid for again.
425
524
  - Process containment uses process tags plus a 1-second tracker. A process
426
525
  that clears its tag and leaves the process tree within its first second
@@ -1,3 +1,4 @@
1
+ import { restartInputError } from "../../orchestrator/restart.js";
1
2
  import { resolve } from "node:path";
2
3
  import { Type } from "@earendil-works/pi-ai";
3
4
  import { validateCallSpec } from "../../compat/spec.js";
@@ -20,7 +21,9 @@ export const parameters = Type.Object({
20
21
  timeoutMs: Type.Optional(Type.Number({ description: "Per-call limit on active time in milliseconds (a number). Omit unless a hard limit is needed; prefer budgets." })),
21
22
  key: Type.Optional(Type.String({ description: "A single agent/task run: the call's key. status with wid: that call's full result." })),
22
23
  full: Type.Optional(Type.Boolean({ description: "status: with wid, the complete workflow detail including every output." })),
23
- force: Type.Optional(Type.Boolean({ description: "restart: fence running executions instead of refusing (they resume on the new orchestrator)." })),
24
+ force: Type.Optional(Type.Union([Type.String(), Type.Boolean()], { description: "restart: the token shown by a refusal. Show the user the list and obtain explicit approval first; boolean true is refused. Subagents cannot force a restart." })),
25
+ reason: Type.Optional(Type.String({ description: "restart: non-empty reason, at most 500 characters; required with force." })),
26
+ request: Type.Optional(Type.String({ description: "run/send/stop: your own request id (1-124 chars [A-Za-z0-9][A-Za-z0-9._:-]*) making a retry safe: the same id with the same content gets the first outcome; other content is refused (request-conflict)." })),
24
27
  }, { additionalProperties: true });
25
28
  /** Call fields a tasks/chain run applies to every step that does not set its own. */
26
29
  export const stepDefaults = ["model", "timeoutMs", "budget", "isolation", "context", "tools", "skills", "once", "writer"];
@@ -39,6 +42,16 @@ function call(value, cwd, where) {
39
42
  spec.cwd = resolve(cwd, spec.cwd);
40
43
  return spec;
41
44
  }
45
+ /** v12 §2: Reject unknown explicit call agents before starter or outbox publication; scripts remain call-local. Shared by
46
+ * the tool and the CLI `run --request` (R2). */
47
+ export function checkAgents(body, available) {
48
+ const names = [...(body.call ? [body.call] : []), ...(body.tasks ?? []), ...(body.chain ?? [])].map(call => call.agent);
49
+ if (!names.length)
50
+ return;
51
+ const known = available(), unknown = [...new Set(names.filter(name => !known.includes(name)))];
52
+ if (unknown.length)
53
+ throw new Error(`Unknown agent${unknown.length === 1 ? "" : "s"}: ${unknown.join(", ")}. Available agents: ${known.join(", ") || "(none)"}`);
54
+ }
42
55
  /** v12 §2: Infer unambiguous runs and normalize controls into unchanged wire bodies. */
43
56
  /** P12: a send naming a model is answered with that model and when it applies — `next-request` (a running call switches
44
57
  * at its next provider request), `next-execution` (a call with no live execution launches on it) or `next-generation`
@@ -55,7 +68,7 @@ export function request(args, cwd) {
55
68
  if (typeof action !== "string" || !action)
56
69
  throw new Error("action is required: run, agents, send, stop, revise, status, resume, drain, restart");
57
70
  if (action === "run") {
58
- const { action: _, workflow, source, tasks, chain, args: inputs, name, usageBudget, maxCalls, inputs: files, by: _by, ...spec } = args;
71
+ const { action: _, workflow, source, tasks, chain, args: inputs, name, usageBudget, maxCalls, inputs: files, by: _by, request: _request, ...spec } = args;
59
72
  const choices = [workflow, source, tasks, chain, spec.agent === undefined && spec.task === undefined ? undefined : spec];
60
73
  if (choices.filter(v => v !== undefined).length !== 1)
61
74
  throw new Error("run requires exactly one of workflow, source, tasks, chain, or agent/task");
@@ -142,7 +155,16 @@ export function request(args, cwd) {
142
155
  return { kind: "resume", body: args.wid !== undefined ? { wid: string(args, "wid") } : typeof args.origin === "string" ? { origin: args.origin } : {} };
143
156
  if (action === "drain")
144
157
  return { kind: "drain", body: {} };
145
- if (action === "restart")
146
- return { kind: "restart", body: args.force === true ? { force: true } : {} };
158
+ if (action === "restart") {
159
+ if (args.force === true)
160
+ throw new Error('force:true is refused; show the user the running executions from a restart refusal, then use force:"<token>" and reason:"<why>" only with explicit user approval');
161
+ if (args.force !== undefined && args.force !== false && typeof args.force !== "string")
162
+ throw new Error("force must be the token from a refused restart");
163
+ const body = { ...(typeof args.force === "string" ? { token: args.force } : {}), ...(args.reason !== undefined ? { reason: args.reason } : {}) };
164
+ const invalid = restartInputError(body);
165
+ if (invalid)
166
+ throw new Error(invalid);
167
+ return { kind: "restart", body };
168
+ }
147
169
  throw new Error(`Unsupported action: ${action}; use run, agents, send, stop, revise, status, resume, drain, or restart`);
148
170
  }
@@ -13,8 +13,10 @@ import { dsaHome, orchInbox, orchLedger, orchLock, outboxRoot } from "../paths.j
13
13
  import { CT, JT } from "../types.js";
14
14
  import { attention, presentText, presented, resolved, unfinishedWorkflow } from "./main/snapshots.js";
15
15
  import { isLive, pausedElsewhere, runningOrchestrator, statusBrief, statusCallDetail, statusCompactDetail, statusDetail, statusView, widOfRid } from "../orchestrator/snapshot.js";
16
- import { parameters, request, sendReceipt } from "./main/tool.js";
16
+ import { checkAgents, parameters, request, sendReceipt } from "./main/tool.js";
17
+ import { findRequest, requestRid, sendIdentified } from "../requests.js";
17
18
  import { discoverAgents } from "../compat/agents.js";
19
+ import { restartInputError } from "../orchestrator/restart.js";
18
20
  import { currentOrchestrator, legacyRestart, waitExit } from "../cli/restart.js";
19
21
  import { packageVersion } from "../version.js";
20
22
  let noteSink;
@@ -40,6 +42,7 @@ function quitPolicy(home) {
40
42
  return "pause";
41
43
  }
42
44
  }
45
+ const REQUEST_USE = "request is a string id for run, send or stop (not combined with replaces)";
43
46
  export function registerMain(pi, ui) {
44
47
  const home = dsaHome();
45
48
  let ctx, sender = "", outbox;
@@ -261,6 +264,8 @@ export function registerMain(pi, ui) {
261
264
  for (const field of ["wid", "to", "target"])
262
265
  if (typeof args[field] === "string")
263
266
  args = { ...args, [field]: ridToWid(args[field]) };
267
+ if (args.request !== undefined && (args.action === "status" || args.action === "agents"))
268
+ throw new Error(REQUEST_USE);
264
269
  if (args.action === "status") {
265
270
  if (typeof args.wid !== "string" || !args.wid)
266
271
  return statusBrief(home, { origin: sender });
@@ -278,34 +283,41 @@ export function registerMain(pi, ui) {
278
283
  const { target, ...rest } = args;
279
284
  args = { ...rest, to: target };
280
285
  }
286
+ // A retried answer addresses the question the first attempt resolved (it may be closed by now), like the CLI.
287
+ if (args.action === "send" && args.kind === "answer" && typeof args.request === "string" && (args.qid === undefined || args.rev === undefined)) {
288
+ const prior = (await findRequest(home, requestRid(args.request)))?.request, body = prior?.body;
289
+ if (prior?.kind === "send" && body?.kind === "answer" && typeof body.to === "string" && prior.cond?.qid !== undefined &&
290
+ (args.to === undefined || args.to === body.to) && (args.qid === undefined || args.qid === prior.cond.qid))
291
+ args = { ...args, to: body.to, qid: prior.cond.qid, rev: args.rev ?? prior.cond.rev };
292
+ }
281
293
  if (args.action === "send")
282
294
  args = completeSend(args);
283
295
  // A session resumes its own held work (what its quit paused); the CLI `resume` remains the global one.
284
296
  if (args.action === "resume" && args.wid === undefined)
285
297
  args = { ...args, origin: sender };
286
298
  const normalized = request(args, cwd);
287
- if (normalized.kind === "run") {
288
- const body = normalized.body;
289
- // v12 §2: Reject unknown explicit call agents before starter or outbox publication; scripts remain call-local.
290
- const names = [...(body.call ? [body.call] : []), ...(body.tasks ?? []), ...(body.chain ?? [])].map(call => call.agent);
291
- if (names.length) {
292
- const available = agentsAt(cwd).map(agent => agent.name);
293
- const unknown = [...new Set(names.filter(name => !available.includes(name)))];
294
- if (unknown.length)
295
- throw new Error(`Unknown agent${unknown.length === 1 ? "" : "s"}: ${unknown.join(", ")}. Available agents: ${available.join(", ") || "(none)"}`);
296
- }
297
- }
299
+ // R1: a caller-chosen request id names a run, send or stop; a retry with the same content gets the first outcome.
300
+ if (args.request !== undefined && (typeof args.request !== "string" || !["run", "send", "stop"].includes(normalized.kind) || normalized.replaces?.length))
301
+ throw new Error(REQUEST_USE);
302
+ const rid = typeof args.request === "string" ? requestRid(args.request) : undefined;
303
+ if (normalized.kind === "run")
304
+ checkAgents(normalized.body, () => agentsAt(cwd).map(agent => agent.name));
298
305
  // P33: any call of the run may fork the origin context, so the origin branch is always offered for pinning.
299
306
  const sessionFile = ctx?.sessionManager.getSessionFile();
300
307
  if (normalized.kind === "run" && sessionFile)
301
308
  normalized.body.origin = { sessionFile, leafId: ctx.sessionManager.getLeafId() };
302
309
  if (normalized.kind === "restart") {
303
310
  // Only an orchestrator that decides restarts is sent one (an older one would keep it as an invalid inbox file).
304
- const previous = currentOrchestrator(home), force = args.force === true;
311
+ const body = normalized.body;
312
+ body.initiator = process.env.DSA_CALL ? { call: process.env.DSA_CALL } : { origin: sender };
313
+ const invalid = restartInputError(body, process.env.DSA_EXEC !== undefined);
314
+ if (invalid)
315
+ return { applied: false, reason: invalid };
316
+ const previous = currentOrchestrator(home);
305
317
  if (!previous)
306
318
  return { applied: true, note: "no orchestrator is running; the next one starts on the installed version when work is submitted" };
307
319
  if (!previous.restart) {
308
- const legacy = legacyRestart(home, previous, force);
320
+ const legacy = legacyRestart(home, previous, body, { subagent: process.env.DSA_EXEC !== undefined, tool: true });
309
321
  if (!legacy.applied)
310
322
  return { applied: false, reason: legacy.reason };
311
323
  // It does not start its successor; this session does once it has exited (or its next periodic check would).
@@ -313,7 +325,7 @@ export function registerMain(pi, ui) {
313
325
  return { applied: true, note: restartNote(previous) };
314
326
  }
315
327
  }
316
- const sent = await serial(async () => {
328
+ const outcome = await serial(async () => {
317
329
  if (!outbox || stopped)
318
330
  throw new Error("Main session is not active");
319
331
  signal?.throwIfAborted();
@@ -322,8 +334,16 @@ export function registerMain(pi, ui) {
322
334
  const withdrawn = await outbox.send("orch", "withdraw", { rids: normalized.replaces });
323
335
  normalized.cond = { ...normalized.cond, after: withdrawn.rid };
324
336
  }
337
+ if (rid)
338
+ return sendIdentified(home, outbox, sender, rid, normalized.kind, normalized.body, normalized.cond);
325
339
  return outbox.send("orch", normalized.kind, normalized.body, normalized.cond);
326
340
  });
341
+ if ("conflict" in outcome) {
342
+ const created = ledger().find(e => e.type === JT.created && e.rid === rid);
343
+ return { applied: false, reason: "request-conflict", request: args.request, spec_digest: outcome.digest, ...(created ? { wid: created.wid } : {}),
344
+ note: "this request id was used for different content; use a new id" };
345
+ }
346
+ const sent = "digest" in outcome ? outcome.request : outcome;
327
347
  // P25: run waits for `created`; control requests wait for their terminal lifecycle record (applied or rejected+reason).
328
348
  const deadline = performance.now() + 10_000;
329
349
  while (wait || sent.kind === "run") {
@@ -360,7 +380,7 @@ export function registerMain(pi, ui) {
360
380
  "run (action optional for exactly one launch form): agent+task; tasks:[call specs] parallel; chain:[call specs] sequential ({previous}); workflow:'./script.js' or source (runs.run(key,spec), runs.all([...]), emit(value), args, runs.input(name)). Optional name, cwd, usageBudget, maxCalls, inputs. With tasks/chain, top-level model, timeoutMs, budget, isolation, context, tools, skills, once are defaults for every step (a step's own value wins); a workflow/source script sets them per runs.run call. timeoutMs is milliseconds of active time (a number); omit it unless a hard limit is needed. Explicit unknown agents are rejected BEFORE creation, with available names; unknown script agents fail only their call.",
361
381
  "agents: list names, descriptions, default models and source for this cwd; use these names for run.",
362
382
  "send to:'<wid>/<key>' (bare '<wid>' only for a single-call workflow): steer on a running call delivers at the next safe point (receipt in status/UI); a steer to a call waiting on its question interrupts the question and the subagent usually asks again — use answer to answer it; sealed → finished:<status> — use kind 'follow-up'. follow-up continues a sealed call as generation g+1 or queues after a running turn; follow-up model:'provider/id' or a pool name runs that generation on it. answer: give the qid (or just the call, or nothing when one question is open); to and rev are filled in. A question that needs the user's decision goes to the user; if you answer one yourself, tell the user what you chose. model ('provider/id' or a pool name — its first model not used up): a running call switches at its next provider request; an asking, hibernated or queued call launches on it when it runs again; the reply's model/effect (next-request|next-execution|next-generation) says which. status model = model actually used by the last request; switching = requested, not used yet; switchFailed = refused. A provider content refusal (ToS/usage policy) fails the call at once, not retried. Unknown targets list valid addresses. replaces:[rid] supersedes an earlier send.",
363
- "stop target:<wid|<wid>/<key>> is terminal stopped (usage and partial edits kept); a sealed call → already-sealed:<status>, a finished workflow → terminal:<status>. drain holds existing workflows reversibly (new runs unaffected); resume [wid] releases held workflows. restart (after an update) replaces the orchestrator with the installed version: refused with busy:<running executions> while any runs, unless force:true (they are fenced and resume); hibernated askers and queued calls do not block it. Never kill the orchestrator process. Commands that need the machine (benchmarks, timing) take a lease: tell the subagent to run them as `pi-durable-subagents hold machine [--shared] -- <command>` (FIFO; status lists lease holders and waiters). status: without wid, what runs, asks (with its answer address; hibernated:true holds no slot) or failed, writerWait: a call queued for its git worktree's writer lock (one call whose tools include edit/write runs per worktree; spec writer:false or isolation:'worktree' opts out), sharedWorktree names calls sharing observed edit/write roots (reminder), lease: a call holding or waiting for a resource lease, finished workflows one line each, provider slots held/limit, the config in effect and providers whose usage window is used up (avoided until a probe finds them answering again), and the orchestrator version (versionNote when it differs from the loaded one); wid: one workflow, outputs clipped; wid+key: one call's full result; full:true: everything. A run's rid from {submitted:{rid}} works wherever a wid is expected. revise wid + workflow/source/args starts a revision.",
383
+ "stop target:<wid|<wid>/<key>> is terminal stopped (usage and partial edits kept); a sealed call → already-sealed:<status>, a finished workflow → terminal:<status>. drain holds existing workflows reversibly (new runs unaffected); resume [wid] releases held workflows. restart (after an update) replaces the orchestrator with the installed version: refused with busy:<running executions> while any runs. Never force without the user's explicit approval: show the user the refusal's list first, then supply force:'<token>' and reason. Subagents cannot force; hibernated askers and queued calls do not block it. Never kill the orchestrator process. Commands that need the machine (benchmarks, timing) take a lease: tell the subagent to run them as `pi-durable-subagents hold machine [--shared] -- <command>` (FIFO; status lists lease holders and waiters). status: without wid, what runs, asks (with its answer address; hibernated:true holds no slot) or failed, writerWait: a call queued for its git worktree's writer lock (one call whose tools include edit/write runs per worktree; spec writer:false or isolation:'worktree' opts out), sharedWorktree names calls sharing observed edit/write roots (reminder), lease: a call holding or waiting for a resource lease, finished workflows one line each, provider slots held/limit, the config in effect and providers whose usage window is used up (avoided until a probe finds them answering again), and the orchestrator version (versionNote when it differs from the loaded one); wid: one workflow, outputs clipped; wid+key: one call's full result; full:true: everything. A run's rid from {submitted:{rid}} works wherever a wid is expected. revise wid + workflow/source/args starts a revision.",
364
384
  "Control replies are {applied:true,rid} or {applied:false,reason,rid} when decided; otherwise {submitted:{rid}} after 10s.",
365
385
  ...(agents ? [`Available agents: ${agents}.`] : []),
366
386
  "User sees a summary line above the editor; ↓ on an empty editor (or /subagents) opens the list, Enter watches live OR finished calls (finished transcripts remain on disk) and expands finished workflows. List keys: s steer (paste-capable input), x stop (confirm y), m model, a answer when asked, f follow-up on finished calls; action feedback appears in footer.",
@@ -10,7 +10,10 @@ import { reduceLifecycle } from "../kernel/lifecycle.js";
10
10
  import { OsLock } from "../platform/lock.js";
11
11
  import { orchInbox, orchLedger, orchLock, outboxRoot } from "../paths.js";
12
12
  import { unfinishedWorkflow } from "../agent/main/snapshots.js";
13
+ import { cliInitiator } from "./restart.js";
14
+ import { restartInputError } from "../orchestrator/restart.js";
13
15
  import { JT } from "../types.js";
16
+ import { RequestsBusy, sendIdentified } from "../requests.js";
14
17
  /** P1: Start the detached orchestrator only after probing its OS lock. */
15
18
  export async function startOrchestrator(home, env) {
16
19
  const lock = await new OsLock().tryAcquire(orchLock(home));
@@ -73,10 +76,47 @@ export async function resolution(home, rid, timeoutMs, interval = 100) {
73
76
  await delay(interval);
74
77
  }
75
78
  }
76
- /** P5, P38: Serialize the stable CLI sender across processes and recover its durable outbox. */
79
+ /** P5, P38: Send one control request through the CLI sender, then start the orchestrator. */
77
80
  export async function submit(home, command, target, env = process.env, options = {}) {
78
81
  if (command === "stop" && !target)
79
82
  throw new Error("stop requires a workflow or call id");
83
+ return withSender(home, async (outbox) => {
84
+ const requests = [];
85
+ if (command === "stop-all") {
86
+ const body = { fence: true };
87
+ requests.push(await outbox.send("orch", "drain", body));
88
+ }
89
+ else if (command === "prune") {
90
+ const body = { ...(target ? { wid: target } : {}), ...(options.olderThanDays !== undefined ? { olderThanDays: options.olderThanDays } : {}) };
91
+ requests.push(await outbox.send("orch", "prune", body));
92
+ }
93
+ else if (command === "restart") {
94
+ const body = { ...options.restart, initiator: cliInitiator(env) };
95
+ const invalid = restartInputError(body, env.DSA_EXEC !== undefined);
96
+ if (invalid)
97
+ throw new Error(invalid);
98
+ requests.push(await outbox.send("orch", "restart", body));
99
+ }
100
+ else
101
+ requests.push(await outbox.send("orch", command, command === "stop" ? { target } : command === "resume" && target ? { wid: target } : {}));
102
+ // Publish first: even a starter failure leaves a recoverable request and no idle-exit race.
103
+ await startOrchestrator(home, env);
104
+ return requests;
105
+ });
106
+ }
107
+ /** R1: Submit a request named by a caller-chosen id through the CLI sender: a retry with the same content republishes
108
+ * (or reuses) the recorded envelope, other content is a conflict and publishes nothing. Starts the orchestrator
109
+ * unless the request conflicts. */
110
+ export async function submitIdentified(home, rid, kind, body, cond, env = process.env, starter = startOrchestrator) {
111
+ return withSender(home, async (outbox, sender) => {
112
+ const result = await sendIdentified(home, outbox, sender, rid, kind, body, cond);
113
+ if ("request" in result)
114
+ await starter(home, env);
115
+ return result;
116
+ });
117
+ }
118
+ /** P5, P38: Serialize the stable CLI sender across processes and recover its durable outbox. */
119
+ async function withSender(home, fn) {
80
120
  await mkdir(home, { recursive: true });
81
121
  const sender = `cli:${userInfo().username}@${hostname()}`;
82
122
  const locker = new OsLock(), deadline = performance.now() + 10_000;
@@ -86,7 +126,7 @@ export async function submit(home, command, target, env = process.env, options =
86
126
  lock = await locker.tryAcquire(join(home, `${sender}.lock`));
87
127
  }
88
128
  if (!lock)
89
- throw new Error("CLI sender is busy; retry the command");
129
+ throw new RequestsBusy("CLI sender is busy; retry the command");
90
130
  try {
91
131
  const outbox = await Outbox.open(outboxRoot(home), sender, () => orchInbox(home));
92
132
  try {
@@ -94,24 +134,7 @@ export async function submit(home, command, target, env = process.env, options =
94
134
  for (const rid of reduceLifecycle(records).resolved.keys())
95
135
  await outbox.markResolved(rid);
96
136
  await outbox.republishPending();
97
- const requests = [];
98
- if (command === "stop-all") {
99
- const body = { fence: true };
100
- requests.push(await outbox.send("orch", "drain", body));
101
- }
102
- else if (command === "prune") {
103
- const body = { ...(target ? { wid: target } : {}), ...(options.olderThanDays !== undefined ? { olderThanDays: options.olderThanDays } : {}) };
104
- requests.push(await outbox.send("orch", "prune", body));
105
- }
106
- else if (command === "restart") {
107
- const body = options.force ? { force: true } : {};
108
- requests.push(await outbox.send("orch", "restart", body));
109
- }
110
- else
111
- requests.push(await outbox.send("orch", command, command === "stop" ? { target } : command === "resume" && target ? { wid: target } : {}));
112
- // Publish first: even a starter failure leaves a recoverable request and no idle-exit race.
113
- await startOrchestrator(home, env);
114
- return requests;
137
+ return await fn(outbox, sender);
115
138
  }
116
139
  finally {
117
140
  await outbox.close();