pi-durable-subagents 1.0.19 → 1.0.21
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +40 -0
- package/README.md +111 -12
- package/dist/agent/main/tool.js +26 -4
- package/dist/agent/main.js +36 -16
- package/dist/cli/control.js +43 -20
- package/dist/cli/main.js +54 -17
- package/dist/cli/requests.js +348 -0
- package/dist/cli/restart.js +34 -16
- package/dist/kernel/lifecycle.js +2 -1
- package/dist/orchestrator/engine.js +33 -19
- package/dist/orchestrator/executor/index.js +42 -0
- package/dist/orchestrator/ledger.js +9 -3
- package/dist/orchestrator/restart.js +57 -0
- package/dist/orchestrator/snapshot.js +7 -4
- package/dist/orchestrator/store.js +83 -8
- package/dist/requests.js +87 -0
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,45 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 1.0.21
|
|
4
|
+
|
|
5
|
+
1.0.20 was not published; its changes ship in this release.
|
|
6
|
+
|
|
7
|
+
- Programs can name requests: `run`, `send` and `stop` take `--request <id>`
|
|
8
|
+
(CLI) or `request` (the `subagents` tool). A retry with the same id and the
|
|
9
|
+
same content gets the first outcome and never starts a second workflow or
|
|
10
|
+
follow-up; other content under the id is a `request-conflict` and sends
|
|
11
|
+
nothing. Exit codes: 0 applied, 1 rejected, 3 conflict, 75 not decided yet
|
|
12
|
+
(retry with the same id). `describe --key <id>` (or a wid) reports the
|
|
13
|
+
state, open questions and every call's output in full, what live calls
|
|
14
|
+
wait for and why their last execution was fenced. A pruned workflow leaves
|
|
15
|
+
a tombstone with its final status, request id and digest. See "Driving dsa
|
|
16
|
+
from a program" in the README.
|
|
17
|
+
- A forced restart needs the user's approval in a checkable form: a refused
|
|
18
|
+
`restart` lists the running executions grouped by session (with held
|
|
19
|
+
leases) and a token; `restart --force <token> --reason <text>` fences
|
|
20
|
+
exactly that set and is refused again if it changed. A subagent cannot
|
|
21
|
+
force a restart. The ledger records who forced it and why, and `status`
|
|
22
|
+
shows it for 24 hours.
|
|
23
|
+
- An asker cut off by a restart or crash before its planned hibernation no
|
|
24
|
+
longer loses its question: it hibernates and resumes with the answer
|
|
25
|
+
(also an answer given while dsa restarted), and a `once` step waiting for
|
|
26
|
+
an answer no longer ends as `unknown`.
|
|
27
|
+
- A workflow's done notice names a follow-up still going (queued or running)
|
|
28
|
+
instead of calling it `unknown`.
|
|
29
|
+
- Orchestrator starts are fast again. Each start re-parsed every workflow's
|
|
30
|
+
staged snapshot (tens of megabytes each when it holds a forked parent
|
|
31
|
+
session) before handling any request, so a restart stalled all work for
|
|
32
|
+
over a minute. A verified `pins.json` record per revision now replaces that
|
|
33
|
+
work: on a copy of a home with 117 workflows, recovery takes about 6 s
|
|
34
|
+
instead of 80 s. The first start after the update builds the records once.
|
|
35
|
+
A record that does not verify falls back to the snapshot, so a changed
|
|
36
|
+
pinned file is still reported as a conflict.
|
|
37
|
+
- `restart` no longer reports a failure while the orchestrator is still
|
|
38
|
+
recovering: it says the request is submitted and keeps waiting (up to 10
|
|
39
|
+
minutes), then exits 75 with "still pending — do not resubmit" if no
|
|
40
|
+
orchestrator reached it. After the restart it waits for the orchestrator
|
|
41
|
+
that actually decided it.
|
|
42
|
+
|
|
3
43
|
## 1.0.19
|
|
4
44
|
|
|
5
45
|
- The orchestrator no longer keeps every workflow's pinned origin branch (the
|
package/README.md
CHANGED
|
@@ -57,7 +57,9 @@ the npx cache, so `install-service` refuses to run from there.
|
|
|
57
57
|
| A step is refused, or a dependency fails | The workflow stops that branch cleanly. Nothing is retried in vain. |
|
|
58
58
|
| A provider's usage window runs out (`No available accounts`, usage limit, quota exceeded) | Found at the second refusal in a row, while pi is still retrying. A call in a pool continues **in the same session** on the pool's next model (within pi's next retry or two); new calls skip that provider. After 15 minutes the next call that wants it tries it once; when it answers, new calls and new generations use it again. A call with a single model waits for it instead of failing. Billing errors (402, insufficient balance) still fail at once. |
|
|
59
59
|
| Two subagents would write in the same worktree | Only one runs there at a time. A call that can write (its tools include `edit` or `write`, which pi's default tools do) holds its git worktree's writer lock from its launch until it ends, also while it waits for an answer. Another writer for that worktree waits in order, and status shows `waiting for writer lock: <root> held by <wid>/<key>`. `writer: false` (a call that does not write there), `isolation: "worktree"` and `"writerLock": "off"` opt out. |
|
|
60
|
-
| A subagent waits for an answer for a long time | It releases its model slot and memory, then resumes exactly once when you answer. |
|
|
60
|
+
| A subagent waits for an answer for a long time | It releases its model slot and memory, then resumes exactly once when you answer. The question survives orchestrator restarts (also forced ones) and crashes, including one that hits before the subagent released its slot. |
|
|
61
|
+
| A subagent's work ends (finished, stopped, or cut off) | Every process its tools started ends with that execution, also ones started with `nohup`, `setsid` or `&`: they carry the execution's tag (see the limit below). Anything that must outlive the subagent has to be started by you or the parent session. A command run under `hold` is no exception: a forced restart stops it and its lease is released. |
|
|
62
|
+
| A subagent runs in a worktree you made (`isolation: "none"`, the default, with `cwd`) | Durable Subagents never creates, cleans, moves or deletes that directory or its branch, also not on `prune`; `prune` deletes only its own state and the worktrees it created for `isolation: "worktree"`. |
|
|
61
63
|
|
|
62
64
|
## Use it
|
|
63
65
|
|
|
@@ -238,11 +240,18 @@ pi-durable-subagents start start the orchestrator if work is pendin
|
|
|
238
240
|
pi-durable-subagents resume [wid] continue unfinished or parked work (undoes drain / stop-all)
|
|
239
241
|
pi-durable-subagents drain hold existing workflows: running calls finish, nothing new starts in them
|
|
240
242
|
pi-durable-subagents stop <wid|call>
|
|
243
|
+
pi-durable-subagents run --request <id> --spec <file|-> [--cwd <dir>] [--json] [--wait-ms <n>]
|
|
244
|
+
start a run under a caller-chosen id; safe to retry (see below)
|
|
245
|
+
pi-durable-subagents send --request <id> --to <run-id|wid/key> --kind follow-up|answer|steer|model
|
|
246
|
+
[--call <key>] [--qid <qid> --rev <n>] --message <text|@file> [--model <m>] [--json]
|
|
247
|
+
pi-durable-subagents stop --request <id> <run-id|wid|wid/key|call> [--json]
|
|
248
|
+
pi-durable-subagents describe --key <run-id> | <wid> [--json]
|
|
249
|
+
full read-only state of one run (never starts the orchestrator)
|
|
241
250
|
pi-durable-subagents stop-all pause every existing workflow now; journals stay resumable
|
|
242
251
|
(runs you start afterwards are not held)
|
|
243
252
|
pi-durable-subagents prune [wid] [--older-than <days>]
|
|
244
253
|
delete finished workflows (done, failed, stopped); prints count and bytes freed
|
|
245
|
-
pi-durable-subagents restart [--force] switch to the installed version (see "Updating Durable Subagents")
|
|
254
|
+
pi-durable-subagents restart [--force <token> --reason <text>] switch to the installed version (see "Updating Durable Subagents")
|
|
246
255
|
pi-durable-subagents hold <resource> [--shared] [--max-wait <s>] [--note <text>] -- <command…>
|
|
247
256
|
run one command while holding a resource lease (see below)
|
|
248
257
|
pi-durable-subagents leases [--json] who holds and who waits for each resource
|
|
@@ -255,6 +264,67 @@ The service only runs `start`: it never resumes work you drained or
|
|
|
255
264
|
stopped. Install the CLI globally (`npm i -g pi-durable-subagents`) before
|
|
256
265
|
`install-service`.
|
|
257
266
|
|
|
267
|
+
### Driving dsa from a program
|
|
268
|
+
|
|
269
|
+
A program (a CI job, a script, another agent) should name its requests with
|
|
270
|
+
`--request <id>`: 1–124 characters `[A-Za-z0-9][A-Za-z0-9._:-]*`, unique per
|
|
271
|
+
`DSA_HOME` across `run`, `send` and `stop`. The id is the identity: the first
|
|
272
|
+
submission under an id is decided once and every retry with **the same
|
|
273
|
+
content** gets that first outcome; a retry never starts a second workflow or
|
|
274
|
+
a second follow-up. The run id is also the workflow key: `send --to <run-id>`
|
|
275
|
+
and `describe --key <run-id>` find the workflow without knowing its wid.
|
|
276
|
+
|
|
277
|
+
The content is hashed into `spec_digest` (the request kind, its body and, for
|
|
278
|
+
answers, the question id and revision). Persist the exact spec bytes you
|
|
279
|
+
submit and retry with those bytes: the run's `cwd` is part of the content
|
|
280
|
+
(`spec.cwd`, else `--cwd`, else the current directory, made absolute), so
|
|
281
|
+
retry from the same place or pass it explicitly. The `spec` file is the
|
|
282
|
+
`subagents` tool's run form (`{agent, task, model?, schema?, …}` or
|
|
283
|
+
`{tasks: […], name?, usageBudget?, maxCalls?}`); `context: "fork"` needs a pi
|
|
284
|
+
session and is rejected.
|
|
285
|
+
|
|
286
|
+
| exit | meaning | `--json` reply |
|
|
287
|
+
| --- | --- | --- |
|
|
288
|
+
| 0 | decided and applied (a retry gets the same answer) | run: `{request, wid, created, spec_digest}`; send/stop: `{request, applied, generation?, call?, spec_digest}` |
|
|
289
|
+
| 1 | decided and rejected (`reason`), or a usage error | `{request, applied: false, reason, spec_digest}` |
|
|
290
|
+
| 3 | `request-conflict`: the id already names other content; nothing was sent | `{request, error, wid?, spec_digest, state}` (the original's digest and state) |
|
|
291
|
+
| 75 | not decided yet: submitted but undecided in `--wait-ms` (default 60 s), submitted and then a later step failed (`reason` says which), or a lock was busy (`reason: "busy"`); retry with the same id | `{request, pending: true, reason?}` |
|
|
292
|
+
|
|
293
|
+
`created` is false when the id had already been decided before the command
|
|
294
|
+
ran. A conflicting id stays conflicting forever, also after `prune`: the
|
|
295
|
+
tombstone keeps the final status, the id and the digest. Any sender's retry
|
|
296
|
+
completes a submission another one recorded but did not publish (it died in
|
|
297
|
+
between), so a retry with the same content always converges.
|
|
298
|
+
|
|
299
|
+
From the `subagents` tool, `request` works the same, with one difference: a
|
|
300
|
+
run's origin session (what `context: "fork"` copies, and where notices go) is
|
|
301
|
+
not part of the digest. A retry of the same id from another session therefore
|
|
302
|
+
gets the first session's workflow, its notices and its forked context. A
|
|
303
|
+
retried answer may leave out `to`, `qid` and `rev`: it addresses the question
|
|
304
|
+
the first attempt answered.
|
|
305
|
+
|
|
306
|
+
`describe` reports one of `absent` (never seen), `pending` (submitted, not
|
|
307
|
+
decided — typically no orchestrator is running; `start` or a retry starts it),
|
|
308
|
+
`rejected` (with `reason`), `running`, `asking` (open questions with their
|
|
309
|
+
full text, `qid`, `rev` and the `to` address to answer), `sealed` (finished:
|
|
310
|
+
`status` plus every call's unclipped `output`, `error` and schema `data`) or
|
|
311
|
+
`pruned` (`pruned: {status, endedAt}`; workflows pruned before 1.0.21 have
|
|
312
|
+
only `endedAt`), with `wid`, `request` and
|
|
313
|
+
`spec_digest`. Live calls also show what they wait for (slot, writer lock,
|
|
314
|
+
lease, exhausted provider) and the last fence of their execution
|
|
315
|
+
(`lastFence: {at, exec, reason}`); the reason is a best-effort reading of the
|
|
316
|
+
journals: `restart-force` when a forced restart listed the execution,
|
|
317
|
+
`orchestrator-crash` when the orchestrator died uncleanly while it ran, else
|
|
318
|
+
`process-died`.
|
|
319
|
+
|
|
320
|
+
```sh
|
|
321
|
+
pi-durable-subagents run --request build-42 --spec build-42.json --json
|
|
322
|
+
# exit 75: retry later with the same id and bytes
|
|
323
|
+
pi-durable-subagents describe --key build-42 --json
|
|
324
|
+
pi-durable-subagents send --request build-42-a1 --to build-42 --kind answer \
|
|
325
|
+
--qid <qid> --rev <rev> --message "yes"
|
|
326
|
+
```
|
|
327
|
+
|
|
258
328
|
### Housekeeping
|
|
259
329
|
|
|
260
330
|
Journals are never compacted, so state only grows. `prune` removes finished
|
|
@@ -387,23 +457,45 @@ pi-durable-subagents restart # or the subagents tool: action "restart"
|
|
|
387
457
|
```
|
|
388
458
|
|
|
389
459
|
The orchestrator refuses while any execution runs (a subagent process, or a
|
|
390
|
-
gate before a call's seal)
|
|
391
|
-
|
|
460
|
+
gate before a call's seal). The refusal groups executions by session with ages,
|
|
461
|
+
lease annotations and a token for that exact set; no new execution starts while
|
|
462
|
+
it decides, so nothing slips in between. Calls waiting
|
|
392
463
|
for your answer (hibernated), waiting for a provider slot, or held by a drain
|
|
393
464
|
do not block it. Otherwise it exits and its successor starts at once from the
|
|
394
465
|
installed files and resumes every workflow: an asker keeps its question, a
|
|
395
|
-
queued call launches on the new version.
|
|
396
|
-
|
|
397
|
-
|
|
398
|
-
|
|
466
|
+
queued call launches on the new version.
|
|
467
|
+
|
|
468
|
+
To interrupt those executions deliberately, first show the user the refusal's
|
|
469
|
+
list and obtain their explicit approval, then use its token and a non-empty
|
|
470
|
+
reason (at most 500 characters):
|
|
471
|
+
|
|
472
|
+
```sh
|
|
473
|
+
pi-durable-subagents restart --force <token> --reason "<why>"
|
|
474
|
+
# tool: {action:"restart", force:"<token>", reason:"<why>"}
|
|
475
|
+
```
|
|
476
|
+
|
|
477
|
+
A changed execution set is refused with a fresh list and token. With no live
|
|
478
|
+
executions no token is needed. Bare force cannot fence live executions, and the
|
|
479
|
+
tool rejects `force:true`. Subagents cannot force a restart, even from bash:
|
|
480
|
+
it would fence themselves and other sessions' work. Force fences running
|
|
481
|
+
executions; they resume on the new version from their sessions, like after a
|
|
482
|
+
crash, so a tool call that was running is repeated or reported as interrupted.
|
|
483
|
+
The restart ledger records the reason and initiator; after the next start,
|
|
484
|
+
`status` shows who forced it and why for 24 hours.
|
|
485
|
+
|
|
486
|
+
An orchestrator reads requests only after it has recovered its workflows, so
|
|
487
|
+
a `restart` sent to one that just started waits for it (and says so); if no
|
|
488
|
+
orchestrator reaches the request in time, `restart` exits 75 and the request
|
|
489
|
+
stays pending — it is decided later, so do not send another.
|
|
399
490
|
|
|
400
491
|
To restart only when the machine is quiet, `drain` first (running calls finish
|
|
401
492
|
and nothing new starts in existing workflows), retry `restart` until it is
|
|
402
493
|
accepted, then `resume`. Never kill the orchestrator process: other sessions'
|
|
403
494
|
running calls would be interrupted without a check. An orchestrator from 1.0.17
|
|
404
495
|
or earlier does not know the restart request; `restart` then checks the
|
|
405
|
-
journals itself
|
|
406
|
-
|
|
496
|
+
journals itself with the same token, reason and subagent checks and ends it with
|
|
497
|
+
SIGTERM, which is not atomic: a call launched in between is fenced and resumes.
|
|
498
|
+
The old orchestrator cannot record the new audit fields. A pi session started before the update still
|
|
407
499
|
loads the old extension; start a new one.
|
|
408
500
|
|
|
409
501
|
## What we do not promise
|
|
@@ -419,8 +511,15 @@ loads the old extension; start a new one.
|
|
|
419
511
|
"outcome unknown" notice; other work keeps running.
|
|
420
512
|
- A tool that already ran inside a subagent may run again after a crash, if
|
|
421
513
|
its result never reached the session. Make external side effects
|
|
422
|
-
idempotent, or mark the step `once: true
|
|
423
|
-
|
|
514
|
+
idempotent, or mark the step `once: true`: when an execution is cut off
|
|
515
|
+
while a tool call is running (its result never arrived), the step then
|
|
516
|
+
ends as `unknown` instead of repeating it. Cut off between tool calls, a
|
|
517
|
+
`once` step continues like any other (nothing was left half done). A
|
|
518
|
+
step cut off while its only unfinished tool call is the question it asked
|
|
519
|
+
is not `unknown` either: it keeps waiting and resumes with the answer, also
|
|
520
|
+
an answer given while dsa restarted. If it was cut off while another tool
|
|
521
|
+
call ran beside the question, a `once` step still ends as `unknown`. A follow-up on an `unknown` step continues the
|
|
522
|
+
same session as its next generation.
|
|
424
523
|
- After a crash, the model call that was in flight is paid for again.
|
|
425
524
|
- Process containment uses process tags plus a 1-second tracker. A process
|
|
426
525
|
that clears its tag and leaves the process tree within its first second
|
package/dist/agent/main/tool.js
CHANGED
|
@@ -1,3 +1,4 @@
|
|
|
1
|
+
import { restartInputError } from "../../orchestrator/restart.js";
|
|
1
2
|
import { resolve } from "node:path";
|
|
2
3
|
import { Type } from "@earendil-works/pi-ai";
|
|
3
4
|
import { validateCallSpec } from "../../compat/spec.js";
|
|
@@ -20,7 +21,9 @@ export const parameters = Type.Object({
|
|
|
20
21
|
timeoutMs: Type.Optional(Type.Number({ description: "Per-call limit on active time in milliseconds (a number). Omit unless a hard limit is needed; prefer budgets." })),
|
|
21
22
|
key: Type.Optional(Type.String({ description: "A single agent/task run: the call's key. status with wid: that call's full result." })),
|
|
22
23
|
full: Type.Optional(Type.Boolean({ description: "status: with wid, the complete workflow detail including every output." })),
|
|
23
|
-
force: Type.Optional(Type.Boolean({ description: "restart:
|
|
24
|
+
force: Type.Optional(Type.Union([Type.String(), Type.Boolean()], { description: "restart: the token shown by a refusal. Show the user the list and obtain explicit approval first; boolean true is refused. Subagents cannot force a restart." })),
|
|
25
|
+
reason: Type.Optional(Type.String({ description: "restart: non-empty reason, at most 500 characters; required with force." })),
|
|
26
|
+
request: Type.Optional(Type.String({ description: "run/send/stop: your own request id (1-124 chars [A-Za-z0-9][A-Za-z0-9._:-]*) making a retry safe: the same id with the same content gets the first outcome; other content is refused (request-conflict)." })),
|
|
24
27
|
}, { additionalProperties: true });
|
|
25
28
|
/** Call fields a tasks/chain run applies to every step that does not set its own. */
|
|
26
29
|
export const stepDefaults = ["model", "timeoutMs", "budget", "isolation", "context", "tools", "skills", "once", "writer"];
|
|
@@ -39,6 +42,16 @@ function call(value, cwd, where) {
|
|
|
39
42
|
spec.cwd = resolve(cwd, spec.cwd);
|
|
40
43
|
return spec;
|
|
41
44
|
}
|
|
45
|
+
/** v12 §2: Reject unknown explicit call agents before starter or outbox publication; scripts remain call-local. Shared by
|
|
46
|
+
* the tool and the CLI `run --request` (R2). */
|
|
47
|
+
export function checkAgents(body, available) {
|
|
48
|
+
const names = [...(body.call ? [body.call] : []), ...(body.tasks ?? []), ...(body.chain ?? [])].map(call => call.agent);
|
|
49
|
+
if (!names.length)
|
|
50
|
+
return;
|
|
51
|
+
const known = available(), unknown = [...new Set(names.filter(name => !known.includes(name)))];
|
|
52
|
+
if (unknown.length)
|
|
53
|
+
throw new Error(`Unknown agent${unknown.length === 1 ? "" : "s"}: ${unknown.join(", ")}. Available agents: ${known.join(", ") || "(none)"}`);
|
|
54
|
+
}
|
|
42
55
|
/** v12 §2: Infer unambiguous runs and normalize controls into unchanged wire bodies. */
|
|
43
56
|
/** P12: a send naming a model is answered with that model and when it applies — `next-request` (a running call switches
|
|
44
57
|
* at its next provider request), `next-execution` (a call with no live execution launches on it) or `next-generation`
|
|
@@ -55,7 +68,7 @@ export function request(args, cwd) {
|
|
|
55
68
|
if (typeof action !== "string" || !action)
|
|
56
69
|
throw new Error("action is required: run, agents, send, stop, revise, status, resume, drain, restart");
|
|
57
70
|
if (action === "run") {
|
|
58
|
-
const { action: _, workflow, source, tasks, chain, args: inputs, name, usageBudget, maxCalls, inputs: files, by: _by, ...spec } = args;
|
|
71
|
+
const { action: _, workflow, source, tasks, chain, args: inputs, name, usageBudget, maxCalls, inputs: files, by: _by, request: _request, ...spec } = args;
|
|
59
72
|
const choices = [workflow, source, tasks, chain, spec.agent === undefined && spec.task === undefined ? undefined : spec];
|
|
60
73
|
if (choices.filter(v => v !== undefined).length !== 1)
|
|
61
74
|
throw new Error("run requires exactly one of workflow, source, tasks, chain, or agent/task");
|
|
@@ -142,7 +155,16 @@ export function request(args, cwd) {
|
|
|
142
155
|
return { kind: "resume", body: args.wid !== undefined ? { wid: string(args, "wid") } : typeof args.origin === "string" ? { origin: args.origin } : {} };
|
|
143
156
|
if (action === "drain")
|
|
144
157
|
return { kind: "drain", body: {} };
|
|
145
|
-
if (action === "restart")
|
|
146
|
-
|
|
158
|
+
if (action === "restart") {
|
|
159
|
+
if (args.force === true)
|
|
160
|
+
throw new Error('force:true is refused; show the user the running executions from a restart refusal, then use force:"<token>" and reason:"<why>" only with explicit user approval');
|
|
161
|
+
if (args.force !== undefined && args.force !== false && typeof args.force !== "string")
|
|
162
|
+
throw new Error("force must be the token from a refused restart");
|
|
163
|
+
const body = { ...(typeof args.force === "string" ? { token: args.force } : {}), ...(args.reason !== undefined ? { reason: args.reason } : {}) };
|
|
164
|
+
const invalid = restartInputError(body);
|
|
165
|
+
if (invalid)
|
|
166
|
+
throw new Error(invalid);
|
|
167
|
+
return { kind: "restart", body };
|
|
168
|
+
}
|
|
147
169
|
throw new Error(`Unsupported action: ${action}; use run, agents, send, stop, revise, status, resume, drain, or restart`);
|
|
148
170
|
}
|
package/dist/agent/main.js
CHANGED
|
@@ -13,8 +13,10 @@ import { dsaHome, orchInbox, orchLedger, orchLock, outboxRoot } from "../paths.j
|
|
|
13
13
|
import { CT, JT } from "../types.js";
|
|
14
14
|
import { attention, presentText, presented, resolved, unfinishedWorkflow } from "./main/snapshots.js";
|
|
15
15
|
import { isLive, pausedElsewhere, runningOrchestrator, statusBrief, statusCallDetail, statusCompactDetail, statusDetail, statusView, widOfRid } from "../orchestrator/snapshot.js";
|
|
16
|
-
import { parameters, request, sendReceipt } from "./main/tool.js";
|
|
16
|
+
import { checkAgents, parameters, request, sendReceipt } from "./main/tool.js";
|
|
17
|
+
import { findRequest, requestRid, sendIdentified } from "../requests.js";
|
|
17
18
|
import { discoverAgents } from "../compat/agents.js";
|
|
19
|
+
import { restartInputError } from "../orchestrator/restart.js";
|
|
18
20
|
import { currentOrchestrator, legacyRestart, waitExit } from "../cli/restart.js";
|
|
19
21
|
import { packageVersion } from "../version.js";
|
|
20
22
|
let noteSink;
|
|
@@ -40,6 +42,7 @@ function quitPolicy(home) {
|
|
|
40
42
|
return "pause";
|
|
41
43
|
}
|
|
42
44
|
}
|
|
45
|
+
const REQUEST_USE = "request is a string id for run, send or stop (not combined with replaces)";
|
|
43
46
|
export function registerMain(pi, ui) {
|
|
44
47
|
const home = dsaHome();
|
|
45
48
|
let ctx, sender = "", outbox;
|
|
@@ -261,6 +264,8 @@ export function registerMain(pi, ui) {
|
|
|
261
264
|
for (const field of ["wid", "to", "target"])
|
|
262
265
|
if (typeof args[field] === "string")
|
|
263
266
|
args = { ...args, [field]: ridToWid(args[field]) };
|
|
267
|
+
if (args.request !== undefined && (args.action === "status" || args.action === "agents"))
|
|
268
|
+
throw new Error(REQUEST_USE);
|
|
264
269
|
if (args.action === "status") {
|
|
265
270
|
if (typeof args.wid !== "string" || !args.wid)
|
|
266
271
|
return statusBrief(home, { origin: sender });
|
|
@@ -278,34 +283,41 @@ export function registerMain(pi, ui) {
|
|
|
278
283
|
const { target, ...rest } = args;
|
|
279
284
|
args = { ...rest, to: target };
|
|
280
285
|
}
|
|
286
|
+
// A retried answer addresses the question the first attempt resolved (it may be closed by now), like the CLI.
|
|
287
|
+
if (args.action === "send" && args.kind === "answer" && typeof args.request === "string" && (args.qid === undefined || args.rev === undefined)) {
|
|
288
|
+
const prior = (await findRequest(home, requestRid(args.request)))?.request, body = prior?.body;
|
|
289
|
+
if (prior?.kind === "send" && body?.kind === "answer" && typeof body.to === "string" && prior.cond?.qid !== undefined &&
|
|
290
|
+
(args.to === undefined || args.to === body.to) && (args.qid === undefined || args.qid === prior.cond.qid))
|
|
291
|
+
args = { ...args, to: body.to, qid: prior.cond.qid, rev: args.rev ?? prior.cond.rev };
|
|
292
|
+
}
|
|
281
293
|
if (args.action === "send")
|
|
282
294
|
args = completeSend(args);
|
|
283
295
|
// A session resumes its own held work (what its quit paused); the CLI `resume` remains the global one.
|
|
284
296
|
if (args.action === "resume" && args.wid === undefined)
|
|
285
297
|
args = { ...args, origin: sender };
|
|
286
298
|
const normalized = request(args, cwd);
|
|
287
|
-
|
|
288
|
-
|
|
289
|
-
|
|
290
|
-
|
|
291
|
-
|
|
292
|
-
|
|
293
|
-
const unknown = [...new Set(names.filter(name => !available.includes(name)))];
|
|
294
|
-
if (unknown.length)
|
|
295
|
-
throw new Error(`Unknown agent${unknown.length === 1 ? "" : "s"}: ${unknown.join(", ")}. Available agents: ${available.join(", ") || "(none)"}`);
|
|
296
|
-
}
|
|
297
|
-
}
|
|
299
|
+
// R1: a caller-chosen request id names a run, send or stop; a retry with the same content gets the first outcome.
|
|
300
|
+
if (args.request !== undefined && (typeof args.request !== "string" || !["run", "send", "stop"].includes(normalized.kind) || normalized.replaces?.length))
|
|
301
|
+
throw new Error(REQUEST_USE);
|
|
302
|
+
const rid = typeof args.request === "string" ? requestRid(args.request) : undefined;
|
|
303
|
+
if (normalized.kind === "run")
|
|
304
|
+
checkAgents(normalized.body, () => agentsAt(cwd).map(agent => agent.name));
|
|
298
305
|
// P33: any call of the run may fork the origin context, so the origin branch is always offered for pinning.
|
|
299
306
|
const sessionFile = ctx?.sessionManager.getSessionFile();
|
|
300
307
|
if (normalized.kind === "run" && sessionFile)
|
|
301
308
|
normalized.body.origin = { sessionFile, leafId: ctx.sessionManager.getLeafId() };
|
|
302
309
|
if (normalized.kind === "restart") {
|
|
303
310
|
// Only an orchestrator that decides restarts is sent one (an older one would keep it as an invalid inbox file).
|
|
304
|
-
const
|
|
311
|
+
const body = normalized.body;
|
|
312
|
+
body.initiator = process.env.DSA_CALL ? { call: process.env.DSA_CALL } : { origin: sender };
|
|
313
|
+
const invalid = restartInputError(body, process.env.DSA_EXEC !== undefined);
|
|
314
|
+
if (invalid)
|
|
315
|
+
return { applied: false, reason: invalid };
|
|
316
|
+
const previous = currentOrchestrator(home);
|
|
305
317
|
if (!previous)
|
|
306
318
|
return { applied: true, note: "no orchestrator is running; the next one starts on the installed version when work is submitted" };
|
|
307
319
|
if (!previous.restart) {
|
|
308
|
-
const legacy = legacyRestart(home, previous,
|
|
320
|
+
const legacy = legacyRestart(home, previous, body, { subagent: process.env.DSA_EXEC !== undefined, tool: true });
|
|
309
321
|
if (!legacy.applied)
|
|
310
322
|
return { applied: false, reason: legacy.reason };
|
|
311
323
|
// It does not start its successor; this session does once it has exited (or its next periodic check would).
|
|
@@ -313,7 +325,7 @@ export function registerMain(pi, ui) {
|
|
|
313
325
|
return { applied: true, note: restartNote(previous) };
|
|
314
326
|
}
|
|
315
327
|
}
|
|
316
|
-
const
|
|
328
|
+
const outcome = await serial(async () => {
|
|
317
329
|
if (!outbox || stopped)
|
|
318
330
|
throw new Error("Main session is not active");
|
|
319
331
|
signal?.throwIfAborted();
|
|
@@ -322,8 +334,16 @@ export function registerMain(pi, ui) {
|
|
|
322
334
|
const withdrawn = await outbox.send("orch", "withdraw", { rids: normalized.replaces });
|
|
323
335
|
normalized.cond = { ...normalized.cond, after: withdrawn.rid };
|
|
324
336
|
}
|
|
337
|
+
if (rid)
|
|
338
|
+
return sendIdentified(home, outbox, sender, rid, normalized.kind, normalized.body, normalized.cond);
|
|
325
339
|
return outbox.send("orch", normalized.kind, normalized.body, normalized.cond);
|
|
326
340
|
});
|
|
341
|
+
if ("conflict" in outcome) {
|
|
342
|
+
const created = ledger().find(e => e.type === JT.created && e.rid === rid);
|
|
343
|
+
return { applied: false, reason: "request-conflict", request: args.request, spec_digest: outcome.digest, ...(created ? { wid: created.wid } : {}),
|
|
344
|
+
note: "this request id was used for different content; use a new id" };
|
|
345
|
+
}
|
|
346
|
+
const sent = "digest" in outcome ? outcome.request : outcome;
|
|
327
347
|
// P25: run waits for `created`; control requests wait for their terminal lifecycle record (applied or rejected+reason).
|
|
328
348
|
const deadline = performance.now() + 10_000;
|
|
329
349
|
while (wait || sent.kind === "run") {
|
|
@@ -360,7 +380,7 @@ export function registerMain(pi, ui) {
|
|
|
360
380
|
"run (action optional for exactly one launch form): agent+task; tasks:[call specs] parallel; chain:[call specs] sequential ({previous}); workflow:'./script.js' or source (runs.run(key,spec), runs.all([...]), emit(value), args, runs.input(name)). Optional name, cwd, usageBudget, maxCalls, inputs. With tasks/chain, top-level model, timeoutMs, budget, isolation, context, tools, skills, once are defaults for every step (a step's own value wins); a workflow/source script sets them per runs.run call. timeoutMs is milliseconds of active time (a number); omit it unless a hard limit is needed. Explicit unknown agents are rejected BEFORE creation, with available names; unknown script agents fail only their call.",
|
|
361
381
|
"agents: list names, descriptions, default models and source for this cwd; use these names for run.",
|
|
362
382
|
"send to:'<wid>/<key>' (bare '<wid>' only for a single-call workflow): steer on a running call delivers at the next safe point (receipt in status/UI); a steer to a call waiting on its question interrupts the question and the subagent usually asks again — use answer to answer it; sealed → finished:<status> — use kind 'follow-up'. follow-up continues a sealed call as generation g+1 or queues after a running turn; follow-up model:'provider/id' or a pool name runs that generation on it. answer: give the qid (or just the call, or nothing when one question is open); to and rev are filled in. A question that needs the user's decision goes to the user; if you answer one yourself, tell the user what you chose. model ('provider/id' or a pool name — its first model not used up): a running call switches at its next provider request; an asking, hibernated or queued call launches on it when it runs again; the reply's model/effect (next-request|next-execution|next-generation) says which. status model = model actually used by the last request; switching = requested, not used yet; switchFailed = refused. A provider content refusal (ToS/usage policy) fails the call at once, not retried. Unknown targets list valid addresses. replaces:[rid] supersedes an earlier send.",
|
|
363
|
-
"stop target:<wid|<wid>/<key>> is terminal stopped (usage and partial edits kept); a sealed call → already-sealed:<status>, a finished workflow → terminal:<status>. drain holds existing workflows reversibly (new runs unaffected); resume [wid] releases held workflows. restart (after an update) replaces the orchestrator with the installed version: refused with busy:<running executions> while any runs
|
|
383
|
+
"stop target:<wid|<wid>/<key>> is terminal stopped (usage and partial edits kept); a sealed call → already-sealed:<status>, a finished workflow → terminal:<status>. drain holds existing workflows reversibly (new runs unaffected); resume [wid] releases held workflows. restart (after an update) replaces the orchestrator with the installed version: refused with busy:<running executions> while any runs. Never force without the user's explicit approval: show the user the refusal's list first, then supply force:'<token>' and reason. Subagents cannot force; hibernated askers and queued calls do not block it. Never kill the orchestrator process. Commands that need the machine (benchmarks, timing) take a lease: tell the subagent to run them as `pi-durable-subagents hold machine [--shared] -- <command>` (FIFO; status lists lease holders and waiters). status: without wid, what runs, asks (with its answer address; hibernated:true holds no slot) or failed, writerWait: a call queued for its git worktree's writer lock (one call whose tools include edit/write runs per worktree; spec writer:false or isolation:'worktree' opts out), sharedWorktree names calls sharing observed edit/write roots (reminder), lease: a call holding or waiting for a resource lease, finished workflows one line each, provider slots held/limit, the config in effect and providers whose usage window is used up (avoided until a probe finds them answering again), and the orchestrator version (versionNote when it differs from the loaded one); wid: one workflow, outputs clipped; wid+key: one call's full result; full:true: everything. A run's rid from {submitted:{rid}} works wherever a wid is expected. revise wid + workflow/source/args starts a revision.",
|
|
364
384
|
"Control replies are {applied:true,rid} or {applied:false,reason,rid} when decided; otherwise {submitted:{rid}} after 10s.",
|
|
365
385
|
...(agents ? [`Available agents: ${agents}.`] : []),
|
|
366
386
|
"User sees a summary line above the editor; ↓ on an empty editor (or /subagents) opens the list, Enter watches live OR finished calls (finished transcripts remain on disk) and expands finished workflows. List keys: s steer (paste-capable input), x stop (confirm y), m model, a answer when asked, f follow-up on finished calls; action feedback appears in footer.",
|
package/dist/cli/control.js
CHANGED
|
@@ -10,7 +10,10 @@ import { reduceLifecycle } from "../kernel/lifecycle.js";
|
|
|
10
10
|
import { OsLock } from "../platform/lock.js";
|
|
11
11
|
import { orchInbox, orchLedger, orchLock, outboxRoot } from "../paths.js";
|
|
12
12
|
import { unfinishedWorkflow } from "../agent/main/snapshots.js";
|
|
13
|
+
import { cliInitiator } from "./restart.js";
|
|
14
|
+
import { restartInputError } from "../orchestrator/restart.js";
|
|
13
15
|
import { JT } from "../types.js";
|
|
16
|
+
import { RequestsBusy, sendIdentified } from "../requests.js";
|
|
14
17
|
/** P1: Start the detached orchestrator only after probing its OS lock. */
|
|
15
18
|
export async function startOrchestrator(home, env) {
|
|
16
19
|
const lock = await new OsLock().tryAcquire(orchLock(home));
|
|
@@ -73,10 +76,47 @@ export async function resolution(home, rid, timeoutMs, interval = 100) {
|
|
|
73
76
|
await delay(interval);
|
|
74
77
|
}
|
|
75
78
|
}
|
|
76
|
-
/** P5, P38:
|
|
79
|
+
/** P5, P38: Send one control request through the CLI sender, then start the orchestrator. */
|
|
77
80
|
export async function submit(home, command, target, env = process.env, options = {}) {
|
|
78
81
|
if (command === "stop" && !target)
|
|
79
82
|
throw new Error("stop requires a workflow or call id");
|
|
83
|
+
return withSender(home, async (outbox) => {
|
|
84
|
+
const requests = [];
|
|
85
|
+
if (command === "stop-all") {
|
|
86
|
+
const body = { fence: true };
|
|
87
|
+
requests.push(await outbox.send("orch", "drain", body));
|
|
88
|
+
}
|
|
89
|
+
else if (command === "prune") {
|
|
90
|
+
const body = { ...(target ? { wid: target } : {}), ...(options.olderThanDays !== undefined ? { olderThanDays: options.olderThanDays } : {}) };
|
|
91
|
+
requests.push(await outbox.send("orch", "prune", body));
|
|
92
|
+
}
|
|
93
|
+
else if (command === "restart") {
|
|
94
|
+
const body = { ...options.restart, initiator: cliInitiator(env) };
|
|
95
|
+
const invalid = restartInputError(body, env.DSA_EXEC !== undefined);
|
|
96
|
+
if (invalid)
|
|
97
|
+
throw new Error(invalid);
|
|
98
|
+
requests.push(await outbox.send("orch", "restart", body));
|
|
99
|
+
}
|
|
100
|
+
else
|
|
101
|
+
requests.push(await outbox.send("orch", command, command === "stop" ? { target } : command === "resume" && target ? { wid: target } : {}));
|
|
102
|
+
// Publish first: even a starter failure leaves a recoverable request and no idle-exit race.
|
|
103
|
+
await startOrchestrator(home, env);
|
|
104
|
+
return requests;
|
|
105
|
+
});
|
|
106
|
+
}
|
|
107
|
+
/** R1: Submit a request named by a caller-chosen id through the CLI sender: a retry with the same content republishes
|
|
108
|
+
* (or reuses) the recorded envelope, other content is a conflict and publishes nothing. Starts the orchestrator
|
|
109
|
+
* unless the request conflicts. */
|
|
110
|
+
export async function submitIdentified(home, rid, kind, body, cond, env = process.env, starter = startOrchestrator) {
|
|
111
|
+
return withSender(home, async (outbox, sender) => {
|
|
112
|
+
const result = await sendIdentified(home, outbox, sender, rid, kind, body, cond);
|
|
113
|
+
if ("request" in result)
|
|
114
|
+
await starter(home, env);
|
|
115
|
+
return result;
|
|
116
|
+
});
|
|
117
|
+
}
|
|
118
|
+
/** P5, P38: Serialize the stable CLI sender across processes and recover its durable outbox. */
|
|
119
|
+
async function withSender(home, fn) {
|
|
80
120
|
await mkdir(home, { recursive: true });
|
|
81
121
|
const sender = `cli:${userInfo().username}@${hostname()}`;
|
|
82
122
|
const locker = new OsLock(), deadline = performance.now() + 10_000;
|
|
@@ -86,7 +126,7 @@ export async function submit(home, command, target, env = process.env, options =
|
|
|
86
126
|
lock = await locker.tryAcquire(join(home, `${sender}.lock`));
|
|
87
127
|
}
|
|
88
128
|
if (!lock)
|
|
89
|
-
throw new
|
|
129
|
+
throw new RequestsBusy("CLI sender is busy; retry the command");
|
|
90
130
|
try {
|
|
91
131
|
const outbox = await Outbox.open(outboxRoot(home), sender, () => orchInbox(home));
|
|
92
132
|
try {
|
|
@@ -94,24 +134,7 @@ export async function submit(home, command, target, env = process.env, options =
|
|
|
94
134
|
for (const rid of reduceLifecycle(records).resolved.keys())
|
|
95
135
|
await outbox.markResolved(rid);
|
|
96
136
|
await outbox.republishPending();
|
|
97
|
-
|
|
98
|
-
if (command === "stop-all") {
|
|
99
|
-
const body = { fence: true };
|
|
100
|
-
requests.push(await outbox.send("orch", "drain", body));
|
|
101
|
-
}
|
|
102
|
-
else if (command === "prune") {
|
|
103
|
-
const body = { ...(target ? { wid: target } : {}), ...(options.olderThanDays !== undefined ? { olderThanDays: options.olderThanDays } : {}) };
|
|
104
|
-
requests.push(await outbox.send("orch", "prune", body));
|
|
105
|
-
}
|
|
106
|
-
else if (command === "restart") {
|
|
107
|
-
const body = options.force ? { force: true } : {};
|
|
108
|
-
requests.push(await outbox.send("orch", "restart", body));
|
|
109
|
-
}
|
|
110
|
-
else
|
|
111
|
-
requests.push(await outbox.send("orch", command, command === "stop" ? { target } : command === "resume" && target ? { wid: target } : {}));
|
|
112
|
-
// Publish first: even a starter failure leaves a recoverable request and no idle-exit race.
|
|
113
|
-
await startOrchestrator(home, env);
|
|
114
|
-
return requests;
|
|
137
|
+
return await fn(outbox, sender);
|
|
115
138
|
}
|
|
116
139
|
finally {
|
|
117
140
|
await outbox.close();
|