pi-durable-subagents 1.0.23 → 1.0.25

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,39 @@
1
1
  # Changelog
2
2
 
3
+ ## 1.0.25
4
+
5
+ - `restart` refusals show what a fence would cut short: each lease a running
6
+ call holds, with its mode, how long it has been held, its command and note.
7
+ Leases held outside those executions (a shell, a `systemd-run` unit) are
8
+ listed apart, since a restart leaves them held.
9
+ - The waiting/moving check period is configured as `k.waitCheckMs` (was
10
+ `k.r7Ms` in 1.0.24; the old key is not accepted).
11
+ - README: deduplicate events by `id`, not cursor; a call with no next
12
+ execution has no `fenced`; an open question keeps the orchestrator from its
13
+ idle exit (compact the event log now with a non-force `restart`); leases
14
+ outside calls survive restarts; why a restart cannot hand running
15
+ executions to the new orchestrator.
16
+
17
+ ## 1.0.24
18
+
19
+ - `events --all [--since <cursor>] [--limit <n>]`: one durable log of
20
+ milestones across all workflows (`submitted`, `started`, `asking` with the
21
+ full question, `answered` without the text, `sealed` with status and data,
22
+ `fenced` with the reason of an interruption, `workflow-done`), read with an
23
+ `<epoch>:<seq>` cursor in pages of at most 1000. Delivery is at least once
24
+ without gaps, also across `kill -9` and restarts (re-derived events keep
25
+ their `id`); events are kept at least 7 days and never while their workflow
26
+ is unfinished or asking; an older cursor gets `cursor-expired` (exit 4). A
27
+ prune whose events cannot be logged first is rejected (`event-log: …`).
28
+ - `run --labels <json>` (tool: `labels`): caller labels, part of the spec
29
+ digest, returned by `describe` and echoed on every event of the run.
30
+ - Why a call does not move: `describe` adds `reason`, `detail` and
31
+ `since` to each waiting call, and the event log gets `waiting`/`moving`
32
+ when the reason changes: `unconfirmed-stop`, `provider-exhausted`,
33
+ `writer-lock`, `lease`, `slot`, `silent`.
34
+ - README: the subagent force-restart guard is a rail against accidents, not a
35
+ security boundary.
36
+
3
37
  ## 1.0.23
4
38
 
5
39
  - An asker cut off by a restart (also `restart --force`) now hibernates and
package/README.md CHANGED
@@ -235,12 +235,14 @@ pi-durable-subagents smoke check this machine and this pi (offline,
235
235
  pi-durable-subagents chaos run the fault suite (offline, about 2 minutes)
236
236
  pi-durable-subagents status [wid] [--json]
237
237
  pi-durable-subagents events <wid> [--json] the meaningful timeline of one workflow
238
+ pi-durable-subagents events --all [--since <cursor>] [--limit <n>] [--json]
239
+ milestones of every workflow, read with a cursor (see below)
238
240
  pi-durable-subagents tail [wid] [--json]
239
241
  pi-durable-subagents start start the orchestrator if work is pending; sends nothing
240
242
  pi-durable-subagents resume [wid] continue unfinished or parked work (undoes drain / stop-all)
241
243
  pi-durable-subagents drain hold existing workflows: running calls finish, nothing new starts in them
242
244
  pi-durable-subagents stop <wid|call>
243
- pi-durable-subagents run --request <id> --spec <file|-> [--cwd <dir>] [--json] [--wait-ms <n>]
245
+ pi-durable-subagents run --request <id> --spec <file|-> [--labels <json>] [--cwd <dir>] [--json] [--wait-ms <n>]
244
246
  start a run under a caller-chosen id; safe to retry (see below)
245
247
  pi-durable-subagents send --request <id> --to <run-id|wid/key> --kind follow-up|answer|steer|model
246
248
  [--call <key>] [--qid <qid> --rev <n>] --message <text|@file> [--model <m>] [--json]
@@ -310,8 +312,9 @@ full text, `qid`, `rev` and the `to` address to answer), `sealed` (finished:
310
312
  `status` plus every call's unclipped `output`, `error` and schema `data`) or
311
313
  `pruned` (`pruned: {status, endedAt}`; workflows pruned before 1.0.21 have
312
314
  only `endedAt`), with `wid`, `request` and
313
- `spec_digest`. Live calls also show what they wait for (slot, writer lock,
314
- lease, exhausted provider). `lastFence: {at, exec, reason}` appears only
315
+ `spec_digest`, and `labels` when the run has any. Live calls also show what
316
+ they wait for (slot, writer lock, lease, exhausted provider; see "Labels and
317
+ why a call does not move"). `lastFence: {at, exec, reason}` appears only
315
318
  when an execution was cut off: it had not ended its turn when it was fenced,
316
319
  did not hibernate on a question (also when recovery finds it cut off while only
317
320
  its question's `ask` ran), and was not ended on purpose (stop, timeout,
@@ -329,6 +332,118 @@ pi-durable-subagents send --request build-42-a1 --to build-42 --kind answer \
329
332
  --qid <qid> --rev <rev> --message "yes"
330
333
  ```
331
334
 
335
+ #### Events across workflows
336
+
337
+ `events --all` reads one durable log of milestones of every workflow
338
+ (`$DSA_HOME/events.jsonl`, written only by the orchestrator), so a program
339
+ without a daemon can poll it and react to completions and questions without
340
+ reading every workflow. Output is JSON lines (with or without `--json`).
341
+
342
+ ```sh
343
+ pi-durable-subagents events --all # {"head":"<epoch>:<seq>","more":false}
344
+ pi-durable-subagents events --all --since <cursor> --limit 500
345
+ ```
346
+
347
+ Every event has `id`, `cursor`, `ts` (when the milestone happened), `type`,
348
+ `wid`, `request` (the run id, when the run was created by `run --request`),
349
+ `labels` (the run's labels, when it has any), and for call events `key`,
350
+ `gen` and `call` (`<wid>@<rev>/<key>@<gen>`):
351
+
352
+ | type | fields |
353
+ | --- | --- |
354
+ | `submitted` | `name?` — the workflow was created |
355
+ | `started` | `exec` — the first execution of a call (generation) began |
356
+ | `asking` | `qid`, `rev`, `question` (full text), `to` (`<wid>/<key>`, the answer address) |
357
+ | `answered` | `qid`, `rev`, `by`, `via?`, `digest` (sha256 hex of the UTF-8 answer), `length` (its length in UTF-16 code units, as JavaScript counts) — never the text; `describe` has it |
358
+ | `sealed` | `status` (`ok`, `failed`, `gate-failed`, `stopped`, `timeout`, `budget`, `unknown`, …), `error?` (unclipped), `data` when its JSON is at most 16 KiB, else `data_omitted: <bytes>` (read it with `describe`) |
359
+ | `fenced` | `exec` (the execution cut off), `reason` (`restart-force`, `orchestrator-crash`, `process-died`), `at` — an execution was interrupted and the call resumed in a new one: the processes its tools had started are gone |
360
+ | `workflow-done` | `status`, `error?` |
361
+
362
+ Readers must ignore types they do not know (`waiting`/`moving` follow).
363
+ `by` is the sender of the answer: `session:<id>` for a pi session (with
364
+ `via: "ui"` when it came from the subagent list), `cli:<user>@<host>` for the
365
+ CLI (a subagent answering through the CLI also shows as `cli:…`), else
366
+ `unknown`. `fenced` is emitted when the call's next execution begins (right
367
+ after recovery, before it waits for a slot) and only when the fence
368
+ interrupted work, exactly as `describe`'s `lastFence`: a turn that had ended,
369
+ a hibernated question, an answer's resume or a seal are no `fenced`. A `once`
370
+ call cut off in a tool is never resumed: it gets `sealed` with status
371
+ `unknown` and no `fenced`. A call with no next execution has no `fenced`; its
372
+ `sealed` carries the outcome (`describe`'s `lastFence` still names the fence).
373
+
374
+ Cursors are `<epoch>:<seq>`; `--since c` returns the events after `c` in log
375
+ order, at most `--limit` (default and maximum 1000), then
376
+ `{"head": …, "more": …}`. With `more: true`, `head` is the cursor of the last
377
+ event printed: pass it as the next `--since`. With `more: false`, `head` is
378
+ the log's head; it may name a seq that no event has (every orchestrator start
379
+ skips 1000 seqs, so a seq you saw in a write that a power cut undid is never
380
+ reused), and it is still a valid cursor. Without `--since` only the head is
381
+ printed. When no log exists yet, the command starts the orchestrator (which
382
+ creates it from everything still on disk) and waits up to `--wait-ms`
383
+ (default 60 s), else prints `{"pending": true}` and exits 75; when the log
384
+ exists it never starts anything.
385
+
386
+ Delivery is at least once, without gaps: after a crash the orchestrator
387
+ derives again from its last durable watermark, and an event derived again has
388
+ the same `id` (a new cursor). Deduplicate by `id`, not by cursor (keeping
389
+ ids for the retention window is enough), and persist your cursor only after
390
+ you applied the events of a page.
391
+
392
+ Retention: an event is dropped only when it was logged more than 7 days ago
393
+ (`"k": { "eventRetentionMs": … }` in `$DSA_HOME/config.json`) and its workflow is
394
+ finished in its current revision (done, failed or stopped — not parked) with
395
+ no open question and no unsealed call, or was pruned. The log is compacted at
396
+ orchestrator start and at most hourly; to compact now while a question is open
397
+ (which keeps the orchestrator from idle exit), run `restart` without `--force`. A cursor of another epoch (the log was
398
+ replaced: a corrupt log is kept aside as `events.jsonl.corrupt-<ms>` and a new
399
+ one starts), below the highest dropped seq, or beyond the head gets exit 4
400
+ and one line `{"error":"cursor-expired","head":"…","oldest":"…"}` (`oldest`
401
+ is the smallest cursor still accepted). To recover, run `describe --key` for
402
+ every run you have not closed (it reports `sealed`, `asking` with the full
403
+ question, `pruned`, …), rebuild your state from those answers, then continue
404
+ with `--since <head>` from that reply. A malformed cursor or option exits 1
405
+ with `{"error":"invalid-arguments","message":…}`.
406
+
407
+ #### Labels and why a call does not move
408
+
409
+ `run --request <id> --spec <file> --labels '{"node":"n1","attempt":"2"}'`
410
+ (tool: `labels: {…}`) attaches your own labels to a run: a flat JSON object
411
+ of at most 32 keys `[A-Za-z0-9_.:-]{1,64}` with string values of at most 256
412
+ characters, at most 4096 bytes of JSON. They are part of the content: the
413
+ same id with other labels (or none) is a `request-conflict` (exit 3). Give
414
+ them only with `--labels`; a `labels` field in the spec file is refused.
415
+ Invalid labels exit 1 and submit nothing, and the orchestrator rejects a
416
+ request that carries invalid ones (`invalid-labels: …`). `describe` returns
417
+ them as `labels` (also after `prune`), and every event of the run carries
418
+ them.
419
+
420
+ When an unsealed call does not move, its `waiting` in `describe` adds
421
+ `reason`, `detail` (the status line for that cause, e.g. `waiting for a slot:
422
+ probe 1/1`) and `since` (ms: when that cause started). The event log has the
423
+ same: `waiting {reason, detail, since}` when the reason appears or changes,
424
+ `moving {after}` when it clears (also when the call ends), checked every
425
+ `k.waitCheckMs` (default 5 s; read when the orchestrator starts, unlike the other
426
+ `k` settings a `config.json` change does not apply it until a restart); a
427
+ change of detail alone is no event. The first reason
428
+ that applies wins:
429
+
430
+ | reason | the call … |
431
+ | --- | --- |
432
+ | `unconfirmed-stop` | had processes that did not exit after SIGKILL; it starts nothing until they are gone (look at them) |
433
+ | `provider-exhausted` | runs on, or can only be admitted to, providers whose usage window is used up |
434
+ | `writer-lock` | waits for another call that writes in the same worktree |
435
+ | `lease` | waits for a resource lease (`hold`, below) |
436
+ | `slot` | is queued for a provider slot or memory headroom (at once when its providers are full, else after 3 s) |
437
+ | `silent` | is running (launched, not stopped) without visible activity: the stall notice of that execution, with the command running and for how long |
438
+
439
+ A call whose current execution asked a question (also while it hibernates
440
+ until the answer) is `asking`, not waiting. Once the answer arrives the call
441
+ launches again, and from then on it waits like any call (for a slot, the
442
+ writer lock, …) even though `describe` still lists the question as open until
443
+ the new execution reads it (its `state` stays `asking`). A drained call
444
+ (no execution running) is never `silent`; a sealed call never waits. `provider-exhausted`, `writer-lock`, `lease` and `slot` are queues
445
+ that clear by themselves; `silent` and `unconfirmed-stop` may need a look.
446
+
332
447
  ### Housekeeping
333
448
 
334
449
  Journals are never compacted, so state only grows. `prune` removes finished
@@ -349,7 +464,7 @@ command:
349
464
  ```sh
350
465
  pi-durable-subagents hold machine -- make bench # exclusive
351
466
  pi-durable-subagents hold machine --shared -- npm test # with other shared holders, never with an exclusive one
352
- pi-durable-subagents hold machine --max-wait 600 --note "frame phase" -- ./measure.sh
467
+ pi-durable-subagents hold machine --max-wait 600 --note "profile" -- ./measure.sh
353
468
  ```
354
469
 
355
470
  - The lease covers one command, not a whole call: a subagent that thinks
@@ -370,6 +485,10 @@ pi-durable-subagents hold machine --max-wait 600 --note "frame phase" -- ./measu
370
485
  ended). State is one small file per request under
371
486
  `$DSA_HOME/leases/<resource>/`; no orchestrator is needed, and the user's
372
487
  own shell can take part.
488
+ - A lease taken outside any call (your shell, a `systemd-run --user` unit)
489
+ does not depend on the orchestrator: a restart, forced or not, leaves it
490
+ held, and it is released when its `hold` and command end. A lease taken
491
+ inside a call ends with that call's processes when the call is fenced.
373
492
  - Subagents find the command on their `PATH` (the orchestrator puts a shim
374
493
  in `$DSA_HOME/bin`), and their leases are tagged with their call:
375
494
  `status` shows `lease: machine held by <wid>/<key> …; waiting: …` and
@@ -453,8 +572,11 @@ On load, it checks the pi exports and API methods it uses.
453
572
  Running work stays on the version it started with until you restart the
454
573
  orchestrator. When the orchestrator runs another version than the one a pi
455
574
  session loaded, that pi says so once, and `status` shows the running version
456
- with a note. The orchestrator exits about 10 s after all work ends, and the
457
- next start runs the new version. To switch sooner:
575
+ with a note. The orchestrator exits about 10 s (`k.idleExitMs`) after all
576
+ work ends: every workflow is finished in its current revision (done, failed,
577
+ stopped or parked) or held by `drain`. A workflow with an open question is not
578
+ finished, so a call hibernated on its question keeps the orchestrator running
579
+ (it holds no slot and costs little). The next start runs the new version. To switch sooner:
458
580
 
459
581
  ```sh
460
582
  pi-durable-subagents restart # or the subagents tool: action "restart"
@@ -462,7 +584,9 @@ pi-durable-subagents restart # or the subagents tool: action "restart"
462
584
 
463
585
  The orchestrator refuses while any execution runs (a subagent process, or a
464
586
  gate before a call's seal). The refusal groups executions by session with ages,
465
- lease annotations and a token for that exact set; no new execution starts while
587
+ the leases each one holds (mode, how long, command, note: what a fence would cut
588
+ short) and a token for that exact set. Leases held outside those executions are
589
+ listed apart, since the restart leaves them held; no new execution starts while
466
590
  it decides, so nothing slips in between. Calls waiting
467
591
  for your answer (hibernated), waiting for a provider slot, or held by a drain
468
592
  do not block it. Otherwise it exits and its successor starts at once from the
@@ -481,9 +605,21 @@ pi-durable-subagents restart --force <token> --reason "<why>"
481
605
  A changed execution set is refused with a fresh list and token. With no live
482
606
  executions no token is needed. Bare force cannot fence live executions, and the
483
607
  tool rejects `force:true`. Subagents cannot force a restart, even from bash:
484
- it would fence themselves and other sessions' work. Force fences running
608
+ it would fence themselves and other sessions' work. That guard reads the
609
+ environment on purpose (`DSA_EXEC`/`DSA_CALL` and the request's initiator
610
+ call): it is a rail against accidents and instructions, not a security
611
+ boundary — a subagent runs as the same OS user and could signal the
612
+ orchestrator anyway. Force fences running
485
613
  executions; they resume on the new version from their sessions, like after a
486
614
  crash, so a tool call that was running is repeated or reported as interrupted.
615
+ A running execution cannot be handed over to the new orchestrator: each
616
+ subagent is a pi process the orchestrator drives over its stdin and stdout, and
617
+ those pipes end with the old process. A crash is no different: the successor
618
+ fences every execution that still runs (an execution that had already ended is
619
+ not counted as interrupted). Work that must survive a forced restart, such as a
620
+ long measurement, belongs outside the subagent's processes (for example
621
+ `systemd-run --user … pi-durable-subagents hold machine -- …`), with the
622
+ subagent only watching it.
487
623
  The restart ledger records the reason and initiator; after the next start,
488
624
  `status` shows who forced it and why for 24 hours.
489
625
 
@@ -16,10 +16,11 @@ export const parameters = Type.Object({
16
16
  usageBudget: Type.Optional(Type.Object({ tokens: Type.Optional(Type.Number()), costUsd: Type.Optional(Type.Number()) })),
17
17
  maxCalls: Type.Optional(Type.Integer({ minimum: 1 })), inputs: Type.Optional(Type.Record(Type.String(), Type.String())),
18
18
  name: Type.Optional(Type.String()),
19
+ labels: Type.Optional(Type.Record(Type.String(), Type.String(), { description: "run: your labels, e.g. {node, attempt}: at most 32 keys [A-Za-z0-9_.:-]{1,64}, string values of at most 256 characters, 4096 bytes of JSON; part of the request's content (spec_digest), shown by describe and on its events." })),
19
20
  timeoutMs: Type.Optional(Type.Number({ description: "Per-call limit on active time in milliseconds (a number). Omit unless a hard limit is needed; prefer budgets." })),
20
21
  key: Type.Optional(Type.String({ description: "A single agent/task run: the call's key. status with wid: that call's full result." })),
21
22
  full: Type.Optional(Type.Boolean({ description: "status: with wid, the complete workflow detail including every output." })),
22
- force: Type.Optional(Type.Union([Type.String(), Type.Boolean()], { description: "restart: the token shown by a refusal. Show the user the list and obtain explicit approval first; boolean true is refused. Subagents cannot force a restart." })),
23
+ force: Type.Optional(Type.Union([Type.String(), Type.Boolean()], { description: "restart: the token shown by a refusal. Show the user the list and obtain explicit approval first; boolean true is refused. Subagents cannot force a restart (an environment-based rail against accidents, not a security boundary)." })),
23
24
  reason: Type.Optional(Type.String({ description: "restart: non-empty reason, at most 500 characters; required with force." })),
24
25
  request: Type.Optional(Type.String({ description: "run/send/stop: your own request id (1-124 chars [A-Za-z0-9][A-Za-z0-9._:-]*) making a retry safe: the same id with the same content gets the first outcome; other content is refused (request-conflict)." })),
25
26
  }, { additionalProperties: true });
@@ -2,6 +2,7 @@ import { restartInputError } from "../../orchestrator/restart.js";
2
2
  import { resolve } from "node:path";
3
3
  import { validateCallSpec } from "../../compat/spec.js";
4
4
  import { compileFanout } from "../../compat/fanout.js";
5
+ import { checkLabels } from "../../events/labels.js";
5
6
  /** Call fields a tasks/chain run applies to every step that does not set its own. */
6
7
  export const stepDefaults = ["model", "timeoutMs", "budget", "isolation", "context", "tools", "skills", "once", "writer"];
7
8
  function string(args, name) {
@@ -20,7 +21,7 @@ function call(value, cwd, where) {
20
21
  return spec;
21
22
  }
22
23
  /** v12 §2: Reject unknown explicit call agents before starter or outbox publication; scripts remain call-local. Shared by
23
- * the tool and the CLI `run --request` (R2). */
24
+ * the tool and the CLI `run --request`. */
24
25
  export function checkAgents(body, available) {
25
26
  const names = [...(body.call ? [body.call] : []), ...(body.tasks ?? []), ...(body.chain ?? [])].map(call => call.agent);
26
27
  if (!names.length)
@@ -45,7 +46,7 @@ export function request(args, cwd) {
45
46
  if (typeof action !== "string" || !action)
46
47
  throw new Error("action is required: run, agents, send, stop, revise, status, resume, drain, restart");
47
48
  if (action === "run") {
48
- const { action: _, workflow, source, tasks, chain, args: inputs, name, usageBudget, maxCalls, inputs: files, by: _by, request: _request, ...spec } = args;
49
+ const { action: _, workflow, source, tasks, chain, args: inputs, name, usageBudget, maxCalls, inputs: files, labels, by: _by, request: _request, ...spec } = args;
49
50
  const choices = [workflow, source, tasks, chain, spec.agent === undefined && spec.task === undefined ? undefined : spec];
50
51
  if (choices.filter(v => v !== undefined).length !== 1)
51
52
  throw new Error("run requires exactly one of workflow, source, tasks, chain, or agent/task");
@@ -84,6 +85,9 @@ export function request(args, cwd) {
84
85
  body.args = inputs;
85
86
  if (name !== undefined)
86
87
  body.name = string(args, "name");
88
+ // Part of the spec digest; an empty object is the same as none.
89
+ if (labels !== undefined && Object.keys(checkLabels(labels)).length)
90
+ body.labels = labels;
87
91
  // P31a, P36, P11: workflow-level limits and declared input files (absolute paths, pinned at admission).
88
92
  if (usageBudget !== undefined) {
89
93
  const b = usageBudget;
@@ -297,7 +297,7 @@ export function registerMain(pi, ui) {
297
297
  if (args.action === "resume" && args.wid === undefined)
298
298
  args = { ...args, origin: sender };
299
299
  const normalized = request(args, cwd);
300
- // R1: a caller-chosen request id names a run, send or stop; a retry with the same content gets the first outcome.
300
+ // A caller-chosen request id names a run, send or stop; a retry with the same content gets the first outcome.
301
301
  if (args.request !== undefined && (typeof args.request !== "string" || !["run", "send", "stop"].includes(normalized.kind) || normalized.replaces?.length))
302
302
  throw new Error(REQUEST_USE);
303
303
  const rid = typeof args.request === "string" ? requestRid(args.request) : undefined;
@@ -378,7 +378,7 @@ export function registerMain(pi, ui) {
378
378
  pi.registerTool(defineTool({
379
379
  name: "subagents", label: "Subagents", description: [
380
380
  "Durable asynchronous subagents; run returns {wid} when created (or {submitted:{rid}} while pending). A finished workflow (its notice carries every agent's result) or a question wakes you, so after starting work end your turn: never poll with sleep or repeated status. Crash recovery resumes sessions, not external side effects. Background helper processes (orchestrator, evaluator) exit by themselves about 10 s after all work ends: never kill processes or delete files to 'clean up'. When the user quits pi, this session's running workflows pause (nothing is spent); resume continues them.",
381
- "run (action optional for exactly one launch form): agent+task; tasks:[call specs] parallel; chain:[call specs] sequential ({previous}); workflow:'./script.js' or source (runs.run(key,spec), runs.all([...]), emit(value), args, runs.input(name)). Optional name, cwd, usageBudget, maxCalls, inputs. With tasks/chain, top-level model, timeoutMs, budget, isolation, context, tools, skills, once are defaults for every step (a step's own value wins); a workflow/source script sets them per runs.run call. timeoutMs is milliseconds of active time (a number); omit it unless a hard limit is needed. Explicit unknown agents are rejected BEFORE creation, with available names; unknown script agents fail only their call.",
381
+ "run (action optional for exactly one launch form): agent+task; tasks:[call specs] parallel; chain:[call specs] sequential ({previous}); workflow:'./script.js' or source (runs.run(key,spec), runs.all([...]), emit(value), args, runs.input(name)). Optional name, cwd, usageBudget, maxCalls, inputs, labels. With tasks/chain, top-level model, timeoutMs, budget, isolation, context, tools, skills, once are defaults for every step (a step's own value wins); a workflow/source script sets them per runs.run call. timeoutMs is milliseconds of active time (a number); omit it unless a hard limit is needed. Explicit unknown agents are rejected BEFORE creation, with available names; unknown script agents fail only their call.",
382
382
  "agents: list names, descriptions, default models and source for this cwd; use these names for run.",
383
383
  "send to:'<wid>/<key>' (bare '<wid>' only for a single-call workflow): steer on a running call delivers at the next safe point (receipt in status/UI); a steer to a call waiting on its question interrupts the question and the subagent usually asks again — use answer to answer it; sealed → finished:<status> — use kind 'follow-up'. follow-up continues a sealed call as generation g+1 or queues after a running turn; follow-up model:'provider/id' or a pool name runs that generation on it. answer: give the qid (or just the call, or nothing when one question is open); to and rev are filled in. A question that needs the user's decision goes to the user; if you answer one yourself, tell the user what you chose. model ('provider/id' or a pool name — its first model not used up): a running call switches at its next provider request; an asking, hibernated or queued call launches on it when it runs again; the reply's model/effect (next-request|next-execution|next-generation) says which. status model = model actually used by the last request; switching = requested, not used yet; switchFailed = refused. A provider content refusal (ToS/usage policy) fails the call at once, not retried. Unknown targets list valid addresses. replaces:[rid] supersedes an earlier send.",
384
384
  "stop target:<wid|<wid>/<key>> is terminal stopped (usage and partial edits kept); a sealed call → already-sealed:<status>, a finished workflow → terminal:<status>. drain holds existing workflows reversibly (new runs unaffected); resume [wid] releases held workflows. restart (after an update) replaces the orchestrator with the installed version: refused with busy:<running executions> while any runs. Never force without the user's explicit approval: show the user the refusal's list first, then supply force:'<token>' and reason. Subagents cannot force; hibernated askers and queued calls do not block it. Never kill the orchestrator process. Commands that need the machine (benchmarks, timing) take a lease: tell the subagent to run them as `pi-durable-subagents hold machine [--shared] -- <command>` (FIFO; status lists lease holders and waiters). status: without wid, what runs, asks (with its answer address; hibernated:true holds no slot) or failed, writerWait: a call queued for its git worktree's writer lock (one call whose tools include edit/write runs per worktree; spec writer:false or isolation:'worktree' opts out), sharedWorktree names calls sharing observed edit/write roots (reminder), lease: a call holding or waiting for a resource lease, finished workflows one line each, provider slots held/limit, the config in effect and providers whose usage window is used up (avoided until a probe finds them answering again), and the orchestrator version (versionNote when it differs from the loaded one); wid: one workflow, outputs clipped; wid+key: one call's full result; full:true: everything. A run's rid from {submitted:{rid}} works wherever a wid is expected. revise wid + workflow/source/args starts a revision.",
@@ -104,7 +104,7 @@ export async function submit(home, command, target, env = process.env, options =
104
104
  return requests;
105
105
  });
106
106
  }
107
- /** R1: Submit a request named by a caller-chosen id through the CLI sender: a retry with the same content republishes
107
+ /** Submit a request named by a caller-chosen id through the CLI sender: a retry with the same content republishes
108
108
  * (or reuses) the recorded envelope, other content is a conflict and publishes nothing. Starts the orchestrator
109
109
  * unless the request conflicts. */
110
110
  export async function submitIdentified(home, rid, kind, body, cond, env = process.env, starter = startOrchestrator) {
package/dist/cli/main.js CHANGED
@@ -132,7 +132,7 @@ export function serviceEntryError(entry) {
132
132
  return `install-service refuses to run from an npx cache (${entry}); the cache can be pruned and the service would break. Install the CLI with \`npm i -g pi-durable-subagents\` and run \`pi-durable-subagents install-service\` again.`;
133
133
  return undefined;
134
134
  }
135
- export const HELP = "pi-durable-subagents: smoke | status [wid] [--json] | events <wid> [--json] | tail [wid] [--json] | start | resume [wid] | drain | stop <wid|callId> | stop-all | run --request <id> --spec <file|-> [--cwd <dir>] [--json] [--wait-ms <n>] | send --request <id> --to <run-id|wid/key> [--call <key>] --kind follow-up|answer|steer|model [--qid <qid> --rev <n>] [--message <text|@file>] [--model <m>] [--json] [--wait-ms <n>] | stop --request <id> <run-id|wid|wid/key> [--json] [--wait-ms <n>] | describe --key <id> | describe <wid> [--json] | prune [wid] [--older-than <days>] | restart [--force <token> --reason <text>] | hold <resource> [--shared] [--max-wait <s>] [--note <text>] -- <command…> | leases [--json] | doctor [--json] | install-service [--dry-run] | uninstall-service [--dry-run] | chaos [--scenario <1-9>] [--keep] [--json]";
135
+ export const HELP = "pi-durable-subagents: smoke | status [wid] [--json] | events <wid> [--json] | events --all [--since <cursor>] [--limit <n>] [--json] [--wait-ms <n>] | tail [wid] [--json] | start | resume [wid] | drain | stop <wid|callId> | stop-all | run --request <id> --spec <file|-> [--labels <json>] [--cwd <dir>] [--json] [--wait-ms <n>] | send --request <id> --to <run-id|wid/key> [--call <key>] --kind follow-up|answer|steer|model [--qid <qid> --rev <n>] [--message <text|@file>] [--model <m>] [--json] [--wait-ms <n>] | stop --request <id> <run-id|wid|wid/key> [--json] [--wait-ms <n>] | describe --key <id> | describe <wid> [--json] | prune [wid] [--older-than <days>] | restart [--force <token> --reason <text>] | hold <resource> [--shared] [--max-wait <s>] [--note <text>] -- <command…> | leases [--json] | doctor [--json] | install-service [--dry-run] | uninstall-service [--dry-run] | chaos [--scenario <1-9>] [--keep] [--json]";
136
136
  /** Restart: the orchestrator exits when no execution runs (or `force`) and the installed version takes over. */
137
137
  async function restartCommand(home, env, write, options) {
138
138
  const body = { ...(typeof options.force === "string" ? { token: options.force } : options.force === true ? { force: true } : {}), ...(options.reason !== undefined ? { reason: options.reason } : {}), initiator: cliInitiator(env) };
@@ -194,7 +194,12 @@ export async function main(args = process.argv.slice(2), options = {}) {
194
194
  }
195
195
  if (args[0] === "chaos")
196
196
  return (await import("./chaos/index.js")).chaos(args.slice(1), options.env ?? process.env, options.write);
197
- // R1–R3: program-facing commands named by request ids (strict flags of their own).
197
+ // The cross-workflow event log (strict flags of its own); `events <wid>` stays below.
198
+ if (args[0] === "events" && args.includes("--all")) {
199
+ const env = options.env ?? process.env;
200
+ return (await import("../events/cli.js")).eventsAll(args.slice(1), { home: dsaHome(env), env, write: options.write ?? ((line) => console.log(line)), starter: options.starter ?? startOrchestrator, waitMs: options.waitMs });
201
+ }
202
+ // Program-facing commands named by request ids (strict flags of their own).
198
203
  if (["run", "send", "describe"].includes(args[0]) || (args[0] === "stop" && args.includes("--request"))) {
199
204
  const env = options.env ?? process.env, requests = await import("./requests.js");
200
205
  const ctx = { home: dsaHome(env), env, write: options.write ?? ((line) => console.log(line)), starter: options.starter, waitMs: options.waitMs, cwd: options.cwd, stdin: options.stdin };
@@ -1,12 +1,12 @@
1
- // R1–R4: Program-facing commands named by caller-chosen request ids — `run|send|stop --request <id>` and `describe`.
1
+ // Program-facing commands named by caller-chosen request ids — `run|send|stop --request <id>` and `describe`.
2
2
  // Exit codes: 0 decided (applied/created), 1 rejected or invalid, 3 request-conflict (the id names other content),
3
3
  // 75 not decided within --wait-ms (retry with the same id and content: safe).
4
4
  import { existsSync, readFileSync } from "node:fs";
5
- import { resolve } from "node:path";
5
+ import { join, resolve } from "node:path";
6
6
  import { setTimeout as delay } from "node:timers/promises";
7
7
  import { readJournalSnapshot } from "../kernel/journal.js";
8
8
  import { reduceLifecycle } from "../kernel/lifecycle.js";
9
- import { journalPath, orchLedger } from "../paths.js";
9
+ import { journalPath, orchLedger, pinnedDir } from "../paths.js";
10
10
  import { findRequest, REQUEST_ID, requestId, requestRid, RequestsBusy, specDigest } from "../requests.js";
11
11
  import { isLive, slotsView, workflowSnapshot } from "../orchestrator/snapshot.js";
12
12
  import { leaseCalls, leaseState } from "../platform/lease.js";
@@ -14,6 +14,10 @@ import { checkAgents, request } from "../agent/main/tool.js";
14
14
  import { discoverAgents } from "../compat/agents.js";
15
15
  import { JT } from "../types.js";
16
16
  import { startOrchestrator, submitIdentified } from "./control.js";
17
+ import { endedExecs, fenceReason } from "../events/fence.js";
18
+ import { parseLabels } from "../events/labels.js";
19
+ import { foldWaits, leaseWaits, waitsOf } from "../events/waiting.js";
20
+ import { emptyLedger, foldLedger } from "../orchestrator/ledger.js";
17
21
  export const EXIT = { ok: 0, rejected: 1, conflict: 3, pending: 75 };
18
22
  /** Strict `--name value` / `--flag` parsing; unknown or repeated options are errors. */
19
23
  function flags(args, spec) {
@@ -53,7 +57,7 @@ function decision(entries, rid) {
53
57
  const createdBy = (entries, rid) => entries.find(e => e.type === JT.created && e.rid === rid);
54
58
  async function stdin() { const chunks = []; for await (const chunk of process.stdin)
55
59
  chunks.push(chunk); return Buffer.concat(chunks).toString("utf8"); }
56
- /** R3: The state of a request id (`{request}`) or a workflow (`{wid}`), with full texts (no clipping). */
60
+ /** The state of a request id (`{request}`) or a workflow (`{wid}`), with full texts (no clipping). */
57
61
  export async function describe(home, key, now = Date.now()) {
58
62
  const entries = ledger(home);
59
63
  if ("wid" in key)
@@ -61,7 +65,7 @@ export async function describe(home, key, now = Date.now()) {
61
65
  const rid = requestRid(key.request), found = await findRequest(home, rid);
62
66
  if (!found)
63
67
  return { state: "absent", request: key.request };
64
- const head = { request: key.request, kind: found.request.kind, spec_digest: specDigest(found.request) };
68
+ const head = { request: key.request, kind: found.request.kind, spec_digest: specDigest(found.request), ...labelsOf(found.request) };
65
69
  const decided = decision(entries, rid), created = createdBy(entries, rid);
66
70
  if (found.request.kind === "run" && created)
67
71
  return { ...await describeWorkflow(home, String(created.wid), entries, now), ...head };
@@ -72,24 +76,57 @@ export async function describe(home, key, now = Date.now()) {
72
76
  return { state: "applied", ...head };
73
77
  return { state: "pending", ...head };
74
78
  }
79
+ /** `{labels}` of a run request that has any. */
80
+ function labelsOf(request) {
81
+ const labels = request?.kind === "run" ? request.body?.labels : undefined;
82
+ return labels && typeof labels === "object" && Object.keys(labels).length ? { labels } : {};
83
+ }
84
+ /** The admitted run request that created `wid` (the ledger keeps it after a prune). */
85
+ function runOf(entries, wid) {
86
+ const created = entries.find(e => e.type === JT.created && e.wid === wid);
87
+ return { created, run: created ? entries.find(e => e.type === "request" && e.request.rid === created.rid)?.request : undefined };
88
+ }
89
+ /** The pinned agents of a workflow revision, read once on demand (the waiting check needs the model of a call that names none). */
90
+ function pinnedAgentModel(home, wid, rev) {
91
+ let agents;
92
+ return name => {
93
+ if (!agents) {
94
+ try {
95
+ agents = JSON.parse(readFileSync(join(pinnedDir(home, wid), rev === 1 ? "" : `r${rev}`, "agents.json"), "utf8"));
96
+ }
97
+ catch {
98
+ agents = [];
99
+ }
100
+ }
101
+ return Array.isArray(agents) ? agents.find(a => a?.name === name)?.model : undefined;
102
+ };
103
+ }
75
104
  function describeWorkflow(home, wid, entries, now) {
76
105
  const pruned = entries.find(e => e.type === "pruned" && e.wid === wid);
77
106
  if (pruned)
78
107
  return { state: "pruned", wid, pruned: { ...(pruned.status !== undefined ? { status: String(pruned.status) } : {}), endedAt: Number(pruned.endedAt) },
79
- ...(typeof pruned.request === "string" ? { request: pruned.request } : {}), ...(typeof pruned.spec_digest === "string" ? { spec_digest: pruned.spec_digest } : {}) };
108
+ ...(typeof pruned.request === "string" ? { request: pruned.request } : {}), ...(typeof pruned.spec_digest === "string" ? { spec_digest: pruned.spec_digest } : {}), ...labelsOf(runOf(entries, wid).run) };
80
109
  if (!/^[^/\\\0]+$/.test(wid) || wid === "." || wid === ".." || !existsSync(journalPath(home, wid)))
81
110
  return { state: "absent", wid };
82
- const wf = workflowSnapshot(home, wid), journal = readJournalSnapshot(journalPath(home, wid));
83
- const created = entries.find(e => e.type === JT.created && e.wid === wid), id = created ? requestId(String(created.rid)) : undefined;
84
- const admitted = id ? entries.find(e => e.type === "request" && e.request.rid === created.rid)?.request : undefined;
85
- const slots = slotsView(home, now), leases = leaseCalls(leaseState(home), now);
111
+ // The snapshot and the journal the wait fold reads must be the same bytes (a writer-wait appended between the two reads
112
+ // would give a reason without its writerWait): read again until the journal did not move around the snapshot.
113
+ let journal = readJournalSnapshot(journalPath(home, wid)), wf = workflowSnapshot(home, wid);
114
+ for (let i = 0, again = readJournalSnapshot(journalPath(home, wid)); again !== journal && i < 5; i++, again = readJournalSnapshot(journalPath(home, wid))) {
115
+ journal = again;
116
+ wf = workflowSnapshot(home, wid);
117
+ }
118
+ const { created, run } = runOf(entries, wid), id = created ? requestId(String(created.rid)) : undefined;
119
+ const admitted = id ? run : undefined;
120
+ const lstate = leaseState(home), slots = slotsView(home, now), leases = leaseCalls(lstate, now);
121
+ // The same fold and decision the orchestrator's collector uses, from the disk snapshots.
122
+ const waits = waitsOf(foldWaits(wid, journal), { now, ledger: foldLedger(emptyLedger(), entries), leases: leaseWaits(lstate, now) }, pinnedAgentModel(home, wid, wf.rev));
86
123
  const line = (lines, model) => { const provider = model?.split("/")[0]; return provider ? lines?.find(l => l.startsWith(`${provider} `)) : undefined; };
87
124
  const latest = [...new Map(wf.calls.map(c => [c.key, c])).values()];
88
125
  const calls = latest.map((c) => {
89
126
  const r = c.result, waiting = r ? undefined : {
90
127
  ...(c.writerWait ? { writerWait: c.writerWait } : {}), ...(leases.get(c.callId) ? { lease: leases.get(c.callId) } : {}),
91
128
  ...(line(slots.slots, c.model) ? { slot: line(slots.slots, c.model) } : {}), ...(line(slots.exhausted, c.model) ? { exhausted: line(slots.exhausted, c.model) } : {}),
92
- ...(c.hibernated ? { hibernated: true } : {})
129
+ ...(c.hibernated ? { hibernated: true } : {}), ...waits.get(c.callId)
93
130
  };
94
131
  return { key: c.key, gen: c.gen, phase: c.phase, agent: c.agent, ...(c.model ? { model: c.model } : {}),
95
132
  ...(r ? { status: r.status, ok: r.ok, ...(r.error ? { error: r.error } : {}), output: r.output, ...(r.data !== undefined ? { data: r.data } : {}) } : {}),
@@ -100,48 +137,20 @@ function describeWorkflow(home, wid, entries, now) {
100
137
  const attention = wf.attention.filter(a => a.kind !== "question").map(a => ({ id: a.id, rev: a.rev, kind: a.kind, ...(a.call ? { call: a.call } : {}), text: a.text }));
101
138
  const state = questions.length ? "asking" : isLive(wf) || wf.status === "parked" ? "running" : "sealed";
102
139
  const fence = lastFence(journal, entries);
103
- return { state, wid, ...(id ? { request: id } : {}), ...(admitted ? { spec_digest: specDigest(admitted) } : {}), status: wf.status, ...(wf.error ? { error: wf.error } : {}),
140
+ return { state, wid, ...(id ? { request: id } : {}), ...(admitted ? { spec_digest: specDigest(admitted) } : {}), ...labelsOf(run), status: wf.status, ...(wf.error ? { error: wf.error } : {}),
104
141
  calls, ...(questions.length ? { questions } : {}), ...(attention.length ? { attention } : {}), ...(fence ? { lastFence: fence } : {}) };
105
142
  }
106
- /** R3, best effort: why the latest fence that interrupted work happened. Every execution ends with a fence; one interrupted
107
- * work only when the execution neither settled (its turn ended) before it nor hibernated (it waits for an answer), and
108
- * was not sealed on purpose: a seal ends an execution on purpose unless its outcome is `unknown` (a `once` call cut off
109
- * in a tool) or the execution was recorded as lost (the loss bound sealed it), which are interruptions themselves.
110
- * A seal for an execution that never ran (a launch failure) or that the call's stop, timeout or budget ended is on
111
- * purpose. restart-force: a forced restart listed the execution as live;
112
- * orchestrator-crash: the execution was launched before an orchestrator start that is not preceded by a clean exit and
113
- * fenced after it (startup recovery); otherwise process-died (the child or its host went away, or a drain fenced it). */
143
+ /** Best effort: why the latest fence that interrupted work happened. The per-execution classification and the
144
+ * reason are shared with the event log's `fenced` events (src/events/fence.ts). */
114
145
  export function lastFence(journal, orch) {
115
- const lost = new Set(journal.filter(e => e.type === "loss").map(e => String(e.exec)));
116
- const fencedAt = new Map(journal.filter(e => e.type === JT.fenced).map(e => [String(e.exec), Number(e.seq)]));
117
- const ended = new Set(journal.filter(e => {
118
- const exec = String(e.exec);
119
- // Recovery records `hibernated` after the fence for an execution cut off while only its question's ask ran (P28):
120
- // it was waiting, not working, so that is no interruption either.
121
- if (e.type === "hibernated")
122
- return true;
123
- if (e.type === "settled")
124
- return Number(e.seq) < (fencedAt.get(exec) ?? Infinity);
125
- return e.type === JT.sealed && e.result?.status !== "unknown" && !lost.has(exec);
126
- }).map(e => String(e.exec)));
146
+ const ended = endedExecs(journal);
127
147
  const fence = journal.findLast(e => e.type === JT.fenced && !ended.has(String(e.exec)));
128
- if (!fence)
129
- return undefined;
130
- const exec = String(fence.exec), at = Number(fence.ts);
131
- if (orch.some(e => e.type === "restart" && e.force === true && Array.isArray(e.live) && e.live.includes(exec)))
132
- return { at, exec, reason: "restart-force" };
133
- const launched = journal.find(e => e.type === JT.exec && e.exec === exec), starts = orch.filter(e => e.type === "orchestrator");
134
- const recovery = starts.findLast(s => Number(s.ts) <= at);
135
- if (launched && recovery && Number(launched.ts) < Number(recovery.ts)) {
136
- const prior = orch.filter(e => Number(e.seq) < Number(recovery.seq));
137
- const lastStart = prior.findLast(e => e.type === "orchestrator"), cleanExit = lastStart && prior.some(e => e.type === "orchestrator-exit" && Number(e.seq) > Number(lastStart.seq));
138
- if (!cleanExit)
139
- return { at, exec, reason: "orchestrator-crash" };
140
- }
141
- return { at, exec, reason: "process-died" };
148
+ return fence ? { at: Number(fence.ts), exec: String(fence.exec), reason: fenceReason(journal, orch, fence) } : undefined;
142
149
  }
143
150
  export function renderDescription(d) {
144
151
  const lines = [`${d.request ?? d.wid}: ${d.state}${d.reason ? ` (${d.reason})` : ""}${d.wid && d.request ? ` — ${d.wid}` : ""}${d.status && d.state !== d.status ? ` · ${d.status}` : ""}`];
152
+ if (d.labels)
153
+ lines.push(` labels: ${Object.entries(d.labels).map(([k, v]) => `${k}=${v}`).join(" ")}`);
145
154
  if (d.pruned)
146
155
  lines.push(` pruned: ${d.pruned.status ?? "?"} at ${new Date(d.pruned.endedAt ?? 0).toISOString()}`);
147
156
  if (d.error)
@@ -201,7 +210,7 @@ function pending(ctx, id, json, why = "") {
201
210
  }
202
211
  /** Submit, then wait; a decision whose admitted envelope has other content (another sender won the id) is a conflict. */
203
212
  async function submitAndWait(ctx, seen, id, kind, body, cond, wait, json) {
204
- // R2: `created` is false only when the id was already decided before this invocation submitted (decisions are
213
+ // `created` is false only when the id was already decided before this invocation submitted (decisions are
205
214
  // monotonic); racing first attempts may all report created — the wid is what identifies the run.
206
215
  const rid = requestRid(id), earlier = Boolean(await outcome(ctx.home, rid, kind === "run", 0));
207
216
  let sent;
@@ -227,7 +236,7 @@ async function submitAndWait(ctx, seen, id, kind, body, cond, wait, json) {
227
236
  return { code: await conflict(ctx, id, specDigest(admitted), json) };
228
237
  return { sent, outcome: result, earlier };
229
238
  }
230
- /** R2: a failure once this invocation submitted, or once the id is found recorded with this content, is "not decided
239
+ /** A failure once this invocation submitted, or once the id is found recorded with this content, is "not decided
231
240
  * yet" (75), never a refusal. Otherwise, with --json, a request refused before submission (a usage error, an invalid
232
241
  * spec, an unknown agent, no open question) answers `{request, applied:false, reason, spec_digest?}` with exit 1;
233
242
  * spec_digest is present once the content was complete enough to hash. This invocation submitted nothing. */
@@ -250,17 +259,18 @@ async function refusable(args, ctx, run) {
250
259
  }
251
260
  const forks = (spec) => [spec, ...["tasks", "chain"].flatMap(k => Array.isArray(spec[k]) ? spec[k] : [])]
252
261
  .some(s => s && typeof s === "object" && s.context === "fork");
253
- /** R2: `run --request <id> --spec <file|-> [--cwd <dir>] [--json] [--wait-ms <n>]`. The spec is the `subagents` run
262
+ /** `run --request <id> --spec <file|-> [--cwd <dir>] [--json] [--wait-ms <n>]`. The spec is the `subagents` run
254
263
  * form ({agent,task,…} or {tasks|chain:[…],…}); it is validated by the tool's own normalizer and agent check. */
255
264
  export const runCommand = (args, ctx) => refusable(args, ctx, runRequest);
256
265
  async function runRequest(args, ctx, seen) {
257
- const { values, positionals } = flags(args, { request: "value", spec: "value", cwd: "value", json: "flag", "wait-ms": "value" });
266
+ const { values, positionals } = flags(args, { request: "value", spec: "value", cwd: "value", labels: "value", json: "flag", "wait-ms": "value" });
258
267
  const id = text(values, "request"), file = text(values, "spec"), json = values.json === true, wait = waitMs(values, ctx);
259
268
  seen.id = id;
260
269
  seen.json = json;
261
270
  if (!id || !file || positionals.length)
262
- throw new Error("usage: run --request <id> --spec <file|-> [--cwd <dir>] [--json] [--wait-ms <n>]");
271
+ throw new Error("usage: run --request <id> --spec <file|-> [--labels <json>] [--cwd <dir>] [--json] [--wait-ms <n>]");
263
272
  requestRid(id);
273
+ const raw = text(values, "labels"), labels = raw !== undefined ? parseLabels(raw) : undefined;
264
274
  const bytes = file === "-" ? await (ctx.stdin ?? stdin)() : readFileSync(resolve(ctx.cwd ?? process.cwd(), file), "utf8");
265
275
  let spec;
266
276
  try {
@@ -275,11 +285,13 @@ async function runRequest(args, ctx, seen) {
275
285
  throw new Error("--spec describes a run; action must be absent or \"run\"");
276
286
  if (spec.request !== undefined)
277
287
  throw new Error("the request id is --request, not a spec field");
288
+ if (spec.labels !== undefined)
289
+ throw new Error("labels are given with --labels <json>, not in the spec");
278
290
  if (forks(spec))
279
291
  throw new Error("context \"fork\" needs a pi session to fork; it is not available to run --request");
280
292
  // RunBody.cwd = spec.cwd ?? --cwd ?? the current directory, absolute before the digest.
281
293
  const base = resolve(ctx.cwd ?? process.cwd(), text(values, "cwd") ?? "."), dir = typeof spec.cwd === "string" && spec.cwd ? resolve(base, spec.cwd) : base;
282
- const normalized = request({ ...spec, action: "run", ...(typeof spec.cwd === "string" && spec.cwd ? { cwd: dir } : {}) }, dir);
294
+ const normalized = request({ ...spec, action: "run", ...(typeof spec.cwd === "string" && spec.cwd ? { cwd: dir } : {}), ...(labels ? { labels } : {}) }, dir);
283
295
  const body = normalized.body;
284
296
  seen.digest = specDigest({ kind: "run", body });
285
297
  // An id already recorded is decided by that record: the same content gets its first outcome (the agents it named may
@@ -308,7 +320,7 @@ async function widOf(home, head) {
308
320
  const found = await findRequest(home, rid);
309
321
  return found?.request.kind === "run" && decision(ledger(home), rid)?.type !== "rejected" ? { pending: true } : { wid: head };
310
322
  }
311
- /** R2: `--to <run-id>[/<key>] | <wid>/<key>` (+ `--call <key>`) → `<wid>/<key>`; a run of one call implies its key. */
323
+ /** `--to <run-id>[/<key>] | <wid>/<key>` (+ `--call <key>`) → `<wid>/<key>`; a run of one call implies its key. */
312
324
  async function target(home, to, call, prior) {
313
325
  const cut = to.indexOf("/"), head = cut < 0 ? to : to.slice(0, cut), key = cut < 0 ? call : to.slice(cut + 1);
314
326
  if (cut >= 0 && call !== undefined)
@@ -329,7 +341,7 @@ async function target(home, to, call, prior) {
329
341
  return { to: `${resolved.wid}/${keys[0]}` };
330
342
  throw new Error(`${to} has ${keys.length ? `calls ${keys.join(", ")}` : "no calls yet"}; name one with --call <key> or --to <wid>/<key>`);
331
343
  }
332
- /** R2: `send --request <id> --to <…> --kind follow-up|answer|steer|model [--qid <qid> --rev <n>] --message <text|@file> [--model <m>]`. */
344
+ /** `send --request <id> --to <…> --kind follow-up|answer|steer|model [--qid <qid> --rev <n>] --message <text|@file> [--model <m>]`. */
333
345
  export const sendCommand = (args, ctx) => refusable(args, ctx, sendRequest);
334
346
  async function sendRequest(args, ctx, seen) {
335
347
  const { values, positionals } = flags(args, { request: "value", to: "value", call: "value", kind: "value", qid: "value", rev: "value", message: "value", model: "value", json: "flag", "wait-ms": "value" });
@@ -366,7 +378,7 @@ async function sendRequest(args, ctx, seen) {
366
378
  return done.code;
367
379
  return decided(ctx, id, done, json);
368
380
  }
369
- /** R2: `stop --request <id> <run-id|wid|wid/key|callId>`. */
381
+ /** `stop --request <id> <run-id|wid|wid/key|callId>`. */
370
382
  export const stopCommand = (args, ctx) => refusable(args, ctx, stopRequest);
371
383
  async function stopRequest(args, ctx, seen) {
372
384
  const { values, positionals } = flags(args, { request: "value", json: "flag", "wait-ms": "value" });