pi-durable-subagents 1.0.13 → 1.0.16

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,44 @@
1
1
  # Changelog
2
2
 
3
+ ## 1.0.16
4
+
5
+ - A used-up usage window is found while pi is still retrying: at the second
6
+ quota refusal in a row (`No available accounts`, usage limit, quota
7
+ exceeded), not after pi's retries end. A call from a pool moves to the
8
+ pool's next model that is not used up and has a free slot, in the same
9
+ execution and session, within pi's next retry or two (refused requests use
10
+ no quota); one with a single model waits for its provider. Before, a call kept retrying the used-up provider for as
11
+ long as pi's retry settings allowed (over ten minutes with ten retries).
12
+ - `send kind:"model"` and a follow-up's `model` accept a pool's name: the
13
+ first model of the pool that is not used up (for a running call, also with
14
+ a free slot). The reply names the model picked; a call from that pool stays
15
+ in it, so a later used-up window still moves it on. A follow-up naming a
16
+ pool starts its generation from the pool's order.
17
+
18
+ ## 1.0.15
19
+
20
+ - The changes listed under 1.0.14, which was tagged but never published:
21
+ its macOS CI failed because two test files compared worktree roots (real
22
+ paths) with temporary paths under the `/var` symlink. Only those tests
23
+ changed.
24
+
25
+ ## 1.0.14
26
+
27
+ - The running orchestrator's version is visible. `status` shows it
28
+ (`orchestrator: 1.0.14 (pid …)`), and when it is not the version a pi
29
+ session loaded (after an update, running work stays on the old one), that
30
+ pi says so once and `status` adds a note: the orchestrator switches by
31
+ itself about 10 s after all work ends. To switch sooner without stopping
32
+ running calls: `drain`, wait until `status` no longer shows the old
33
+ version, then `resume` from a pi session started after the update.
34
+ - Two subagents editing the same worktree are pointed out. When two calls
35
+ that have not finished both used `edit` or `write` under the same git
36
+ worktree, the main session that started them gets one reminder per pair,
37
+ and `status` names the other call in `sharedWorktree`. A paused call that
38
+ wrote still counts until it ends. Nothing is blocked; the reminder closes
39
+ when either call ends. Writes made only through `bash`
40
+ are not seen; calls with `isolation: "worktree"` have their own worktree.
41
+
3
42
  ## 1.0.13
4
43
 
5
44
  - A subagent resumed after its execution was interrupted (the orchestrator
package/README.md CHANGED
@@ -55,7 +55,8 @@ the npx cache, so `install-service` refuses to run from there.
55
55
  | You steer a subagent while it is asking you a question | Your message reaches it, in order. Nothing is rejected or lost. |
56
56
  | Two steers arrive out of order and the second replaces the first | Only the second one applies. |
57
57
  | A step is refused, or a dependency fails | The workflow stops that branch cleanly. Nothing is retried in vain. |
58
- | A provider's usage window runs out (`No available accounts`, usage limit, quota exceeded) | A call in a pool continues **in the same session** on the pool's next model; new calls skip that provider. After 15 minutes the next call that wants it tries it once; when it answers, new calls and new generations use it again. A call with a single model waits for it instead of failing. Billing errors (402, insufficient balance) still fail at once. |
58
+ | A provider's usage window runs out (`No available accounts`, usage limit, quota exceeded) | Found at the second refusal in a row, while pi is still retrying. A call in a pool continues **in the same session** on the pool's next model (within pi's next retry or two); new calls skip that provider. After 15 minutes the next call that wants it tries it once; when it answers, new calls and new generations use it again. A call with a single model waits for it instead of failing. Billing errors (402, insufficient balance) still fail at once. |
59
+ | Two subagents edit the same worktree | A reminder names both calls; neither is blocked or locked. Only observed `edit`/`write` calls count (bash-only writes are not seen). Calls with `isolation: "worktree"` have their own worktrees. |
59
60
  | A subagent waits for an answer for a long time | It releases its model slot and memory, then resumes exactly once when you answer. |
60
61
 
61
62
  ## Use it
@@ -97,9 +98,9 @@ it something. Each verb means one thing, and a refusal says what would work:
97
98
  |---|---|---|
98
99
  | `run` | — | Start one subagent, `tasks` in parallel, a `chain`, or a workflow script. An unknown agent name is refused before anything starts, with the list of agents. |
99
100
  | `send steer` | a running subagent | Reaches it at its next safe point. To a finished one: refused, use `follow-up`; To one waiting on its question: it interrupts the question, and the subagent usually asks again; `answer` answers it. |
100
- | `send follow-up` | a finished subagent | Continues the same session as a new generation (`key@2`). With `model`, that generation runs on it. |
101
+ | `send follow-up` | a finished subagent | Continues the same session as a new generation (`key@2`). With `model` (a model or a pool's name), that generation runs on it. |
101
102
  | `send answer` | an open question | Answers it once. |
102
- | `send model` | any subagent | A running one switches at its next request; one asking, hibernated or waiting for a slot launches on it when it runs again. |
103
+ | `send model` | any subagent | A running one switches at its next request; one asking, hibernated or waiting for a slot launches on it when it runs again. A pool's name picks its first model that is not used up (and, for a running call, has a free slot); the reply names the model picked, and a call from that pool stays in it. |
103
104
  | `stop` | a subagent or a workflow | Final: `stopped`, usage kept, edits left as they are. |
104
105
  | `drain` / `resume` | existing workflows | A reversible hold; runs started later are not held. |
105
106
 
@@ -176,6 +177,7 @@ The main agent is interrupted only when there is something to decide:
176
177
  - a finished workflow;
177
178
  - a stalled subagent (the alert names the command it is running and for how long, so a long silent command reads differently from a stuck call);
178
179
  - an unknown outcome;
180
+ - two unfinished calls observed editing the same worktree (a reminder, never a block);
179
181
  - a reached budget.
180
182
 
181
183
  Each one arrives once. A reminder that was already resolved is shown as
@@ -320,6 +322,26 @@ On load, it checks the pi exports and API methods it uses.
320
322
 
321
323
  `smoke` runs the same checks inside your pi.
322
324
 
325
+ ## Updating Durable Subagents
326
+
327
+ Running work stays on the version it started with: the orchestrator is not
328
+ restarted under it. When the orchestrator runs another version than the one a
329
+ pi session loaded, that pi says so once, and `status` shows the running
330
+ version with a note. The orchestrator exits about 10 s after all work ends,
331
+ and the next start runs the new version. To switch sooner without stopping
332
+ running calls:
333
+
334
+ 1. `drain`: running calls finish and nothing new starts in existing
335
+ workflows. A call waiting for your answer still counts as running, and a
336
+ workflow started after the drain is not held.
337
+ 2. Wait until `status` no longer shows the old version (about 10 s after the
338
+ last call ends).
339
+ 3. `resume` from a pi session started after the update.
340
+
341
+ A `resume` before the old orchestrator exits keeps it running the old
342
+ version, and a pi session started before the update still starts the old
343
+ version; start a new one.
344
+
323
345
  ## What we do not promise
324
346
 
325
347
  - Call specs are checked strictly when a call is first proposed: an unknown or
@@ -98,6 +98,8 @@ export function registerChild(pi) {
98
98
  return { action: 'defer' };
99
99
  if (req.kind === 'model') {
100
100
  const body = req.body;
101
+ if (body?.exec !== undefined && body.exec !== exec)
102
+ return { action: 'reject', reason: 'stale-execution' };
101
103
  return body && ctx.modelRegistry.find(body.provider, body.model) ? { action: 'apply' } : { action: 'reject', reason: 'unknown-model' };
102
104
  }
103
105
  if (MESSAGES.includes(req.kind)) {
@@ -44,7 +44,7 @@ function call(value, cwd, where) {
44
44
  * (a follow-up's new generation runs on it). From the orchestrator ledger's `send-note`. */
45
45
  export function sendReceipt(ledger, rid) {
46
46
  const note = ledger.find(e => e.type === "send-note" && e.rid === rid);
47
- return note ? { model: String(note.model), effect: String(note.effect) } : {};
47
+ return note ? { model: String(note.model), effect: String(note.effect), ...(note.pool ? { pool: String(note.pool) } : {}) } : {};
48
48
  }
49
49
  export function request(args, cwd) {
50
50
  // v12 §2: Infer run only when one launch form is present; never guess a control verb.
@@ -12,7 +12,7 @@ import { OsLock } from "../platform/lock.js";
12
12
  import { dsaHome, orchInbox, orchLedger, orchLock, outboxRoot } from "../paths.js";
13
13
  import { CT, JT } from "../types.js";
14
14
  import { attention, presentText, presented, resolved, unfinishedWorkflow } from "./main/snapshots.js";
15
- import { isLive, pausedElsewhere, statusBrief, statusCallDetail, statusCompactDetail, statusDetail, statusView, widOfRid } from "../orchestrator/snapshot.js";
15
+ import { isLive, pausedElsewhere, runningOrchestrator, statusBrief, statusCallDetail, statusCompactDetail, statusDetail, statusView, widOfRid } from "../orchestrator/snapshot.js";
16
16
  import { parameters, request, sendReceipt } from "./main/tool.js";
17
17
  import { discoverAgents } from "../compat/agents.js";
18
18
  let noteSink;
@@ -89,6 +89,16 @@ export function registerMain(pi, ui) {
89
89
  await new Promise((resolve, reject) => { child.once("spawn", resolve); child.once("error", reject); });
90
90
  child.unref();
91
91
  }
92
+ // Version visibility: say once per orchestrator process when it runs another version than this pi loaded (after an
93
+ // update its running work stays on the old one); status shows the same note.
94
+ let versionTold;
95
+ function tellVersion() {
96
+ const { orchestrator, versionNote } = runningOrchestrator(home);
97
+ if (!versionNote || orchestrator === versionTold || !ctx?.hasUI)
98
+ return;
99
+ versionTold = orchestrator;
100
+ ctx.ui.notify(`Durable Subagents: ${versionNote}.`, "warning");
101
+ }
92
102
  function collect() { return ctx ? attention(home, sender, [...presented(ctx), ...reserved]) : []; }
93
103
  function message(items) {
94
104
  return { type: "custom_message", customType: CT.attention, content: items.map(presentText).join("\n"), display: true, details: { items } };
@@ -146,6 +156,12 @@ export function registerMain(pi, ui) {
146
156
  await outbox.republishPending();
147
157
  await starter();
148
158
  lastStarter = Date.now();
159
+ try {
160
+ tellVersion();
161
+ }
162
+ catch (error) {
163
+ console.error("durable-subagents:", error);
164
+ }
149
165
  timer = setInterval(() => {
150
166
  if (polling || stopped)
151
167
  return;
@@ -156,6 +172,7 @@ export function registerMain(pi, ui) {
156
172
  if (Date.now() - lastStarter >= 30_000) {
157
173
  lastStarter = Date.now();
158
174
  await starter();
175
+ tellVersion();
159
176
  }
160
177
  await idle();
161
178
  }).catch(error => console.error("durable-subagents:", error)).finally(() => { polling = false; });
@@ -325,8 +342,8 @@ export function registerMain(pi, ui) {
325
342
  "Durable asynchronous subagents; run returns {wid} when created (or {submitted:{rid}} while pending). A finished workflow (its notice carries every agent's result) or a question wakes you, so after starting work end your turn: never poll with sleep or repeated status. Crash recovery resumes sessions, not external side effects. Background helper processes (orchestrator, evaluator) exit by themselves about 10 s after all work ends: never kill processes or delete files to 'clean up'. When the user quits pi, this session's running workflows pause (nothing is spent); resume continues them.",
326
343
  "run (action optional for exactly one launch form): agent+task; tasks:[call specs] parallel; chain:[call specs] sequential ({previous}); workflow:'./script.js' or source (runs.run(key,spec), runs.all([...]), emit(value), args, runs.input(name)). Optional name, cwd, usageBudget, maxCalls, inputs. With tasks/chain, top-level model, timeoutMs, budget, isolation, context, tools, skills, once are defaults for every step (a step's own value wins); a workflow/source script sets them per runs.run call. timeoutMs is milliseconds of active time (a number); omit it unless a hard limit is needed. Explicit unknown agents are rejected BEFORE creation, with available names; unknown script agents fail only their call.",
327
344
  "agents: list names, descriptions, default models and source for this cwd; use these names for run.",
328
- "send to:'<wid>/<key>' (bare '<wid>' only for a single-call workflow): steer on a running call delivers at the next safe point (receipt in status/UI); a steer to a call waiting on its question interrupts the question and the subagent usually asks again — use answer to answer it; sealed → finished:<status> — use kind 'follow-up'. follow-up continues a sealed call as generation g+1 or queues after a running turn; follow-up model:'provider/id' runs that generation on it. answer: give the qid (or just the call, or nothing when one question is open); to and rev are filled in. A question that needs the user's decision goes to the user; if you answer one yourself, tell the user what you chose. model: a running call switches at its next provider request; an asking, hibernated or queued call launches on it when it runs again; the reply's model/effect (next-request|next-execution|next-generation) says which. status model = model actually used by the last request; switching = requested, not used yet; switchFailed = refused. A provider content refusal (ToS/usage policy) fails the call at once, not retried. Unknown targets list valid addresses. replaces:[rid] supersedes an earlier send.",
329
- "stop target:<wid|<wid>/<key>> is terminal stopped (usage and partial edits kept); a sealed call → already-sealed:<status>, a finished workflow → terminal:<status>. drain holds existing workflows reversibly (new runs unaffected); resume [wid] releases held workflows. status: without wid, what runs, asks (with its answer address; hibernated:true holds no slot) or failed, finished workflows one line each, provider slots held/limit, the config in effect and providers whose usage window is used up (avoided until a probe finds them answering again); wid: one workflow, outputs clipped; wid+key: one call's full result; full:true: everything. A run's rid from {submitted:{rid}} works wherever a wid is expected. revise wid + workflow/source/args starts a revision.",
345
+ "send to:'<wid>/<key>' (bare '<wid>' only for a single-call workflow): steer on a running call delivers at the next safe point (receipt in status/UI); a steer to a call waiting on its question interrupts the question and the subagent usually asks again — use answer to answer it; sealed → finished:<status> — use kind 'follow-up'. follow-up continues a sealed call as generation g+1 or queues after a running turn; follow-up model:'provider/id' or a pool name runs that generation on it. answer: give the qid (or just the call, or nothing when one question is open); to and rev are filled in. A question that needs the user's decision goes to the user; if you answer one yourself, tell the user what you chose. model ('provider/id' or a pool name — its first model not used up): a running call switches at its next provider request; an asking, hibernated or queued call launches on it when it runs again; the reply's model/effect (next-request|next-execution|next-generation) says which. status model = model actually used by the last request; switching = requested, not used yet; switchFailed = refused. A provider content refusal (ToS/usage policy) fails the call at once, not retried. Unknown targets list valid addresses. replaces:[rid] supersedes an earlier send.",
346
+ "stop target:<wid|<wid>/<key>> is terminal stopped (usage and partial edits kept); a sealed call → already-sealed:<status>, a finished workflow → terminal:<status>. drain holds existing workflows reversibly (new runs unaffected); resume [wid] releases held workflows. status: without wid, what runs, asks (with its answer address; hibernated:true holds no slot) or failed, sharedWorktree names calls sharing observed edit/write roots (reminder only), finished workflows one line each, provider slots held/limit, the config in effect and providers whose usage window is used up (avoided until a probe finds them answering again), and the orchestrator version (versionNote when it differs from the loaded one); wid: one workflow, outputs clipped; wid+key: one call's full result; full:true: everything. A run's rid from {submitted:{rid}} works wherever a wid is expected. revise wid + workflow/source/args starts a revision.",
330
347
  "Control replies are {applied:true,rid} or {applied:false,reason,rid} when decided; otherwise {submitted:{rid}} after 10s.",
331
348
  ...(agents ? [`Available agents: ${agents}.`] : []),
332
349
  "User sees a summary line above the editor; ↓ on an empty editor (or /subagents) opens the list, Enter watches live OR finished calls (finished transcripts remain on disk) and expands finished workflows. List keys: s steer (paste-capable input), x stop (confirm y), m model, a answer when asked, f follow-up on finished calls; action feedback appears in footer.",
package/dist/cli/main.js CHANGED
@@ -57,7 +57,7 @@ export function renderStatus(wf) {
57
57
  return [`${wf.wid}@${wf.rev}${wf.name ? ` ${wf.name}` : ""}: ${wf.status}${wf.error ? ` (${clip(wf.error, 300)})` : ""} · ${sealed}/${total} sealed${used ? ` · ${formatUsage(wf.usage)}` : ""}`,
58
58
  ...wf.calls.map(c => {
59
59
  const last = c.result?.output?.split("\n").map(l => l.trim()).filter(Boolean).at(-1), err = c.result?.error;
60
- return ` ${c.key}@${c.gen} ${c.result?.status ?? c.phase}${c.model ? ` ${c.model}` : ""}${c.tools ? ` tools:${c.tools}` : ""}${c.usage && (c.usage.input || c.usage.output) ? ` ${formatUsage(c.usage)}` : ""}${last ? ` ${JSON.stringify(clip(last, 160))}` : err ? ` (${clip(err, 160)})` : ""}`;
60
+ return ` ${c.key}@${c.gen} ${c.result?.status ?? c.phase}${c.model ? ` ${c.model}` : ""}${c.tools ? ` tools:${c.tools}` : ""}${c.usage && (c.usage.input || c.usage.output) ? ` ${formatUsage(c.usage)}` : ""}${last ? ` ${JSON.stringify(clip(last, 160))}` : err ? ` (${clip(err, 160)})` : ""}${c.sharedWorktree ? ` (shares worktree with ${c.sharedWorktree.join(", ")})` : ""}`;
61
61
  }),
62
62
  ...wf.attention.map(a => ` ${a.kind}: ${JSON.stringify(clip(a.text, 300))}`),
63
63
  ...(wf.scriptLog ? [` script log: ${wf.scriptLog}`] : [])].join("\n");
@@ -65,13 +65,13 @@ export function renderStatus(wf) {
65
65
  /** P25, T10: Render the compact status projection shared with the `subagents` tool. */
66
66
  export function renderView(view) {
67
67
  const lines = view.workflows.map(w => [`${w.wid}@${w.rev}${w.name ? ` ${w.name}` : ""}: ${w.status}${w.followUps ? " (follow-up running)" : ""}${w.error ? ` (${clip(w.error, 200)})` : ""} · ${w.done}/${w.planned ?? w.calls.length}${w.planned === undefined && w.status === "running" ? "+" : ""} done${w.usage.input || w.usage.output || w.usage.costUsd ? ` · ${formatUsage(w.usage)}` : ""}`,
68
- ...w.calls.map(c => ` ${c.key}@${c.gen} ${c.status ?? c.phase}${c.hibernated ? " (hibernated, no slot)" : ""}${c.model ? ` ${c.model}` : ""}${c.switching ? ` → ${c.switching} (requested)` : ""}${c.switchFailed ? ` (switch refused: ${c.switchFailed})` : ""}${c.tools ? ` tools:${c.tools}` : ""}${c.usage ? ` ${formatUsage(c.usage)}` : ""}${c.lastLine ? ` ${JSON.stringify(c.lastLine)}` : c.error ? ` (${c.error})` : ""}`),
68
+ ...w.calls.map(c => ` ${c.key}@${c.gen} ${c.status ?? c.phase}${c.hibernated ? " (hibernated, no slot)" : ""}${c.model ? ` ${c.model}` : ""}${c.switching ? ` → ${c.switching} (requested)` : ""}${c.switchFailed ? ` (switch refused: ${c.switchFailed})` : ""}${c.tools ? ` tools:${c.tools}` : ""}${c.usage ? ` ${formatUsage(c.usage)}` : ""}${c.lastLine ? ` ${JSON.stringify(c.lastLine)}` : c.error ? ` (${c.error})` : ""}${c.sharedWorktree ? ` (shares worktree with ${c.sharedWorktree.join(", ")})` : ""}`),
69
69
  ...w.attention.map(a => ` ${a.kind}: ${JSON.stringify(a.text.split("\n")[0])}`)].join("\n"));
70
70
  if (view.paused)
71
71
  lines.unshift(`${view.paused} (pi-durable-subagents resume)`);
72
72
  if (view.olderFinished)
73
73
  lines.push(`(+${view.olderFinished} older finished workflows; status <wid> shows one in detail)`);
74
- const footer = [view.slots?.length ? `slots: ${view.slots.join(", ")}` : "", view.config ? `config: ${view.config}` : "", view.configRejected ? `config.json rejected: ${view.configRejected}` : "", ...(view.exhausted ?? [])].filter(Boolean);
74
+ const footer = [view.slots?.length ? `slots: ${view.slots.join(", ")}` : "", view.config ? `config: ${view.config}` : "", view.configRejected ? `config.json rejected: ${view.configRejected}` : "", view.orchestrator ? `orchestrator: ${view.orchestrator}` : "", ...(view.exhausted ?? []), view.versionNote ? `note: ${view.versionNote}` : ""].filter(Boolean);
75
75
  if (!lines.length)
76
76
  lines.push("No workflows");
77
77
  return [...lines, ...footer].join("\n");
@@ -331,7 +331,9 @@ export class Engine {
331
331
  const seal = wf.journal.entries().find(e => e.type === JT.sealed && e.call === from);
332
332
  if (seal && send.kind === 'steer')
333
333
  return { action: 'reject', reason: `finished:${seal.result.status} — use kind "follow-up" to continue it` };
334
- if (send.kind === 'follow-up' && send.model !== undefined) {
334
+ // A pool's name is a model too: the call keeps the pool, and its order and failover apply to the new generation.
335
+ const pools = this.ledgers.config.pools, pool = send.model !== undefined && pools && Object.hasOwn(pools, send.model);
336
+ if (send.kind === 'follow-up' && send.model !== undefined && !pool) {
335
337
  try {
336
338
  if (!parseModel(send.model).provider)
337
339
  throw new Error('missing provider');
@@ -346,7 +348,7 @@ export class Engine {
346
348
  const spec = send.model !== undefined ? { ...entry.spec, model: send.model } : entry.spec;
347
349
  if (send.model !== undefined)
348
350
  await this.note(req.rid, send.model, 'next-generation');
349
- const opened = await wf.journal.append('generation', { rid: req.rid, key: entry.key, gen, from, spec, revision: wf.revision, opening: { rid: req.rid, kind: send.kind, message: send.message ?? '' }, ...(send.model !== undefined ? { model: send.model } : {}) });
351
+ const opened = await wf.journal.append('generation', { rid: req.rid, key: entry.key, gen, from, spec, revision: wf.revision, opening: { rid: req.rid, kind: send.kind, message: send.message ?? '' }, ...(send.model !== undefined && !pool ? { model: send.model } : {}) });
350
352
  this.dispatchGeneration(wf, opened);
351
353
  return { action: 'apply' };
352
354
  }
@@ -30,9 +30,12 @@ import { availableMemory } from "./memory.js";
30
30
  import { reached, sessionUsage, totalUsage } from "./usage.js";
31
31
  import { recordFenceFailure, resolveFenceAttention, serialContainment, skipLostCandidate, sweepExecutions } from "./sweep.js";
32
32
  import { gateRetired } from "./effects/gate.js";
33
+ import { WorktreeIndex, worktreeCalls, worktreeLabel, worktreePair, worktreeRoots } from "./worktree.js";
33
34
  const MEM_RECORD_MS = 30000;
34
35
  /** A4, P29: A child admitted within this window may not show in MemAvailable yet; its share is reserved explicitly. */
35
36
  const MEM_WARMUP_MS = 30000;
37
+ /** Quota refusals in a row that find a provider's usage window used up while pi is still retrying. */
38
+ const QUOTA_REFUSALS = 2;
36
39
  const ignoreMissing = (error) => { if (error.code !== "ENOENT")
37
40
  throw error; };
38
41
  const callOf = (exec) => exec.slice(0, exec.lastIndexOf("#"));
@@ -60,7 +63,8 @@ export function requestedModel(journal, call, followUp) {
60
63
  wanted = m;
61
64
  }
62
65
  for (const [index, e] of all.entries()) {
63
- if (e.type !== "forward" || e.dest !== call || e.envelope?.kind !== "model")
66
+ // A failover's switch is not a request: an execution that ends before applying it leaves the choice to the pool.
67
+ if (e.type !== "forward" || e.dest !== call || e.envelope?.kind !== "model" || e.failover)
64
68
  continue;
65
69
  const delivered = all.find(r => r.type === "forward-delivered" && r.call === call && r.rid2 === e.rid2);
66
70
  if (delivered) {
@@ -299,9 +303,9 @@ export default function createExecutor(ledgers, options = {}) {
299
303
  }
300
304
  /** P7, P27: Record forward-delivered once when a forward's child receipt is first observed; serial sections only. */
301
305
  /** The reply to a model send says which model and when it applies (orchestrator ledger `send-note`, once per rid). */
302
- async function note(rid, model, effect) {
306
+ async function note(rid, model, effect, pool) {
303
307
  if (!orch.entries().some(e => e.type === "send-note" && e.rid === rid))
304
- await orch.append("send-note", { rid, model, effect });
308
+ await orch.append("send-note", { rid, model, effect, ...(pool ? { pool } : {}) });
305
309
  }
306
310
  async function forwardsDelivered(journal, call, entries) {
307
311
  const all = journal.entries();
@@ -350,7 +354,54 @@ export default function createExecutor(ledgers, options = {}) {
350
354
  await t.journal.append(JT.attention, { item: { id, rev: 1, kind: "budget", text: "Workflow budget reached", wid: t.wid } });
351
355
  return hit;
352
356
  }
357
+ const roots = worktreeRoots(), writes = new WorktreeIndex();
358
+ const indexed = () => writes.scan(journals.values());
359
+ const origin = (w) => writes.origin(w.journal) ?? orch.entries().find(e => e.type === JT.created && e.wid === address(w.call).wid)?.origin;
360
+ // Serial sections only: wrote and attention are durable before another writer or seal can interleave.
361
+ /** Remind of every pair of calls that wrote in `root` and have not ended, once per pair and journal. The index must be
362
+ * current; each append here is indexed at once (only that journal is read). */
363
+ async function remind(root) {
364
+ const live = writes.live(root);
365
+ for (let i = 0; i < live.length; i++)
366
+ for (let k = i + 1; k < live.length; k++) {
367
+ const [first, second] = WorktreeIndex.order(live[i], live[k]), id = worktreePair(first.call, second.call);
368
+ // A retry after a partial cross-workflow append keeps the reminder (and its writer) as first written.
369
+ const prior = writes.reminder(id), secondWid = address(second.call).wid;
370
+ const targets = origin(first) === origin(second) ? [prior && prior.wid !== secondWid ? first : second] : [second, first];
371
+ const item = prior ?? { id, rev: 1, kind: "conflict", call: second.call, wid: secondWid,
372
+ text: `${worktreeLabel(first.call)} and ${worktreeLabel(second.call)} both write in ${root} (edit/write seen); assign one owner or move one to its own worktree` };
373
+ for (const target of targets)
374
+ if (!writes.remindedIn(id, target.journal)) {
375
+ await target.journal.append(JT.attention, { item: { ...item, wid: address(target.call).wid } });
376
+ writes.scan([target.journal]);
377
+ }
378
+ }
379
+ }
380
+ async function wrote(t, exec, cwd, path) {
381
+ const root = await roots(cwd, path);
382
+ if (!root)
383
+ return;
384
+ await serial(async () => {
385
+ if (sealed(t.journal, t.callId) || has(t.journal, JT.fenced, exec) || current(t.journal, t.callId) !== exec)
386
+ return;
387
+ if (indexed().has(exec, root))
388
+ return;
389
+ const after = writes.live(root).map(w => w.call).filter(c => c !== t.callId);
390
+ await t.journal.append("wrote", { exec, root, ...(after.length ? { after } : {}) });
391
+ writes.scan([t.journal]);
392
+ await remind(root);
393
+ });
394
+ }
395
+ /** Close the reminders of pairs where `call` (or, without one, any participant) ended. */
396
+ async function resolveWorktrees(call) {
397
+ for (const { id, journal, item } of indexed().open()) {
398
+ const pair = worktreeCalls(id);
399
+ if (call ? pair.includes(call) : pair.some(c => writes.ended(c)))
400
+ await journal.append(JT.attentionResolved, { id, rev: item.rev, resolution: "ended" });
401
+ }
402
+ }
353
403
  async function retireAttention(journal, call) {
404
+ await resolveWorktrees(call);
354
405
  for (const { item } of attentionEntries(journal.entries())) {
355
406
  if (item.kind === "finished" || item.id === `unknown:${call}` || item.id.startsWith("fence:"))
356
407
  continue;
@@ -443,7 +494,7 @@ export default function createExecutor(ledgers, options = {}) {
443
494
  const decision = await decide(), models = decision.candidates, pool = decision.pool;
444
495
  continuation = decision.continuation;
445
496
  for (const model of models) {
446
- if (!continuation && pool && skipped(pool, model))
497
+ if (!continuation && pool && models.length > 1 && skipped(pool, model))
447
498
  continue;
448
499
  const provider = model.provider;
449
500
  if (unavailable(provider))
@@ -511,14 +562,19 @@ export default function createExecutor(ledgers, options = {}) {
511
562
  const candidate = recorded && candidates.some(m => m.provider === recorded.provider && m.id === recorded.id);
512
563
  // Leave the session's model for the pool's others when its pool skips it after losses, or its provider's usage
513
564
  // window is used up; and at a new generation, go back to the pool's order of preference.
514
- const skip = pool && candidate && (previous && ownSegment && skipped(pool, recorded) || unavailable(recorded.provider) || !ownSegment && !!t.continueFrom);
565
+ // A new generation of a pool call starts from the pool also when the session's model is not one of its models
566
+ // (switched outside it, or the follow-up named the pool).
567
+ const skip = pool && (candidate ? previous && ownSegment && skipped(pool, recorded) || unavailable(recorded.provider) || !ownSegment && !!t.continueFrom
568
+ : !ownSegment && !!t.continueFrom);
515
569
  // A model the call was asked to use replaces the session's: launched with it, and holding its provider's slot.
516
570
  const wanted = requestedModel(t.journal, t.callId, t.model);
517
571
  // It outranks the pool's order at a new generation too, also when it names the model the session already has.
572
+ // A requested model of the call's own pool keeps the pool: a used-up window still moves the call on.
573
+ const keep = pool && candidates.some(m => m.provider === wanted?.provider && m.id === wanted?.id) ? pool : undefined;
518
574
  if (wanted)
519
575
  return recorded && !freshFork && recorded.provider === wanted.provider && recorded.id === wanted.id
520
- ? { candidates: [recorded], continuation: true, pool: undefined }
521
- : { candidates: [wanted], continuation: false, pool: undefined };
576
+ ? { candidates: [recorded], continuation: true, pool: keep }
577
+ : { candidates: [wanted], continuation: false, pool: keep };
522
578
  if (recorded && !freshFork && !skip)
523
579
  return { candidates: [recorded], continuation: true, pool: candidate ? pool : undefined };
524
580
  return { candidates, continuation: false, pool };
@@ -586,15 +642,87 @@ export default function createExecutor(ledgers, options = {}) {
586
642
  // While a probe runs, its outcome alone decides: a late refusal of an execution admitted earlier changes nothing.
587
643
  if (x && (x.probe ? x.probe !== exec : now < x.nextTry))
588
644
  return;
589
- if (orch.entries().some(e => e.type === "provider-exhausted" && e.exec === exec))
645
+ // Once per execution and provider: an execution moved on by failover can find a second provider used up too.
646
+ if (orch.entries().some(e => e.type === "provider-exhausted" && e.exec === exec && e.provider === provider))
590
647
  return;
591
648
  await orch.append("provider-exhausted", { provider, exec, since: x?.since ?? now, nextTry: now + (config.k?.probeMs ?? 900_000), error: error.slice(0, 300) });
592
649
  }
650
+ /** Quota refusals in a row per execution, from one provider (pi retries a refused request on its own). */
651
+ const refusals = new Map();
652
+ /** A used-up window shows while pi still retries: the second refusal in a row (the first for a probe) finds the
653
+ * provider used up, and a call launched from a pool switches to the pool's next model at its next request. */
654
+ async function refused(t, exec, provider, error) {
655
+ if (!quotaExhausted(error))
656
+ return;
657
+ const last = refusals.get(exec), count = last?.provider === provider ? last.count + 1 : 1;
658
+ refusals.set(exec, { provider, count });
659
+ const probe = folded().exhausted.get(provider)?.probe === exec;
660
+ if (count < (probe ? 1 : QUOTA_REFUSALS))
661
+ return;
662
+ await serial(async () => {
663
+ if (has(t.journal, JT.fenced, exec) || current(t.journal, t.callId) !== exec)
664
+ return;
665
+ await recordExhausted(provider, exec, error);
666
+ await failover(t, exec, provider);
667
+ });
668
+ wake();
669
+ }
670
+ /** Switch a running execution off a used-up provider: to the first model of its pool on another provider that is
671
+ * neither used up nor full, reserving that slot as a requested switch does. Without one, pi's retries go on and
672
+ * the call waits for the provider once they end. */
673
+ async function failover(t, exec, provider) {
674
+ if (pendingSwitch(exec))
675
+ return;
676
+ const pool = t.journal.entries().findLast(e => e.type === "selected" && e.exec === exec)?.pool, pools = settings().pools;
677
+ if (!pool || !pools?.[pool])
678
+ return;
679
+ let models;
680
+ try {
681
+ models = resolveModel(pool, pools);
682
+ }
683
+ catch {
684
+ return;
685
+ }
686
+ for (const m of models) {
687
+ if (!m.provider || m.provider === provider || unavailable(m.provider) || skipped(pool, m))
688
+ continue;
689
+ const rid = contentHash([exec, "failover", provider]);
690
+ if (t.journal.entries().some(e => e.type === "forward" && e.rid === rid))
691
+ return;
692
+ const probe = folded().exhausted.has(m.provider); // its next try is due (`unavailable` said so): this is its probe
693
+ if (!await reserveSwitch(exec, m.provider, rid))
694
+ continue;
695
+ if (probe)
696
+ await orch.append("provider-probe", { provider: m.provider, exec });
697
+ // Bound to this execution: replayed after it ended, a later execution (which chose its model at launch) refuses it.
698
+ const body = { provider: m.provider, model: m.id, ...(m.thinking ? { thinking: m.thinking } : {}), exec };
699
+ const envelope = { to: t.callId, kind: "model", body }, hash = contentHash(envelope);
700
+ const entry = await t.journal.append("forward", { rid, rid2: forwardRid(rid, t.callId.slice(0, t.callId.indexOf("/")), t.key, hash), dest: t.callId, hash, envelope, failover: provider });
701
+ await replayForward(entry);
702
+ return;
703
+ }
704
+ }
705
+ /** Hold a slot of `provider` for a running execution's switch, unless it holds one; false when it is full. */
706
+ async function reserveSwitch(exec, provider, rid) {
707
+ if (holdings().some(h => h.exec === exec && h.pool === provider))
708
+ return true;
709
+ const target = holdings().filter(h => h.pool === provider);
710
+ if (!capacity({ kind: "provider", holders: target.length, capacity: settings().providers?.[provider]?.slots ?? Infinity }))
711
+ return false;
712
+ let slot = 0;
713
+ while (target.some(h => h.slot === slot))
714
+ slot++;
715
+ await orch.append("hold", { pool: provider, slot, exec, reserved: true, rid });
716
+ return true;
717
+ }
593
718
  /** An answer from a used-up provider, requested after it was found used up: available again. */
594
- async function answered(exec, event) {
719
+ async function answered(t, exec, event) {
595
720
  const message = event.message, provider = message?.provider;
596
- if (message?.role !== "assistant" || !provider || message.stopReason === "error")
721
+ if (message?.role !== "assistant" || !provider)
597
722
  return;
723
+ if (message.stopReason === "error")
724
+ return refused(t, exec, provider, String(message.errorMessage ?? ""));
725
+ refusals.delete(exec);
598
726
  await serial(async () => {
599
727
  const x = folded().exhausted.get(provider);
600
728
  if (x && (x.probe === exec || Number(message.timestamp) > x.since))
@@ -787,7 +915,8 @@ export default function createExecutor(ledgers, options = {}) {
787
915
  }
788
916
  });
789
917
  }, recordUsage: values => recordUsage(t, values),
790
- switched: event => switched(exec, journal, event), answered: event => answered(exec, event), pendingSwitch: () => pendingSwitch(exec),
918
+ wrote: path => wrote(t, exec, cwd, path),
919
+ switched: event => switched(exec, journal, event), answered: event => answered(t, exec, event), pendingSwitch: () => pendingSwitch(exec),
791
920
  });
792
921
  }
793
922
  finally {
@@ -810,7 +939,12 @@ export default function createExecutor(ledgers, options = {}) {
810
939
  journals.set(ticket.wid, ticket.journal);
811
940
  const a = { ticket, controller: new AbortController(), stopped: false, wake: () => { }, promise: undefined, onPark: new Set() };
812
941
  active.set(ticket.callId, a);
813
- a.promise = execute(a).catch(error => {
942
+ a.promise = serial(async () => {
943
+ // Repair a crash between wrote and attention (or between the two origin journals).
944
+ for (const root of indexed().contested())
945
+ await remind(root);
946
+ await resolveWorktrees();
947
+ }).then(() => execute(a)).catch(error => {
814
948
  completed.delete(ticket.callId);
815
949
  if (a.retired || ticket.journal.entries().some(e => e.type === "retired" && e.call === ticket.callId))
816
950
  return buildCallResult({ key: ticket.key, gen: ticket.gen, status: "stopped", output: "", error: "retired" });
@@ -827,6 +961,9 @@ export default function createExecutor(ledgers, options = {}) {
827
961
  for (const e of inUse.keys())
828
962
  if (callOf(e) === ticket.callId)
829
963
  inUse.delete(e);
964
+ for (const e of refusals.keys())
965
+ if (callOf(e) === ticket.callId)
966
+ refusals.delete(e);
830
967
  forgetSession(callSession(home, ticket.wid, ticket.key, ticket.gen));
831
968
  wake();
832
969
  }
@@ -877,6 +1014,28 @@ export default function createExecutor(ledgers, options = {}) {
877
1014
  /** P12: a model request to this call, recorded with the rid given; a reject has no effect. */
878
1015
  const requestModel = async (rid, model, hash, cond) => {
879
1016
  let body;
1017
+ const exec = current(ctx.journal, dest);
1018
+ // P28: with no live execution (not started yet, between executions, hibernated while asking) the model is
1019
+ // recorded and the next execution launches on it (`requestedModel`); its slot is acquired then, as for any launch.
1020
+ const idle = !exec || has(ctx.journal, JT.fenced, exec) || !has(ctx.journal, "selected", exec);
1021
+ // A pool's name asks for its first model that can take the call now: provider not used up, and (for a running
1022
+ // call) a free slot. The call's own pool stays, so a used-up window later moves it on as before.
1023
+ const pools = settings().pools, pool = pools && Object.hasOwn(pools, model) ? model : undefined;
1024
+ if (pool) {
1025
+ let models;
1026
+ try {
1027
+ models = resolveModel(pool, pools);
1028
+ }
1029
+ catch {
1030
+ return { action: "reject", reason: "unknown-model" };
1031
+ }
1032
+ const free = (p) => idle || holdings().some(h => h.exec === exec && h.pool === p) ||
1033
+ capacity({ kind: "provider", holders: holdings().filter(h => h.pool === p).length, capacity: settings().providers?.[p]?.slots ?? Infinity });
1034
+ const m = models.find(m => m.provider && !unavailable(m.provider) && free(m.provider));
1035
+ if (!m)
1036
+ return { action: "reject", reason: "pool-unavailable" };
1037
+ model = `${m.provider}/${m.id}${m.thinking ? `:${m.thinking}` : ""}`;
1038
+ }
880
1039
  try {
881
1040
  const m = parseModel(model);
882
1041
  if (!m.provider)
@@ -888,28 +1047,14 @@ export default function createExecutor(ledgers, options = {}) {
888
1047
  }
889
1048
  const envelope = { to: dest, kind: "model", body, ...(cond && Object.keys(cond).length ? { cond } : {}) };
890
1049
  const rid2 = forwardRid(rid, ctx.widRev, ctx.key, hash);
891
- const exec = current(ctx.journal, dest), provider = body.provider;
892
- // P28: with no live execution (not started yet, between executions, hibernated while asking) the model is
893
- // recorded and the next execution launches on it (`requestedModel`); its slot is acquired then, as for any launch.
894
1050
  // Launching (`selected`, not `tracked` yet): the child may start on the old model; ask again in a moment.
895
- const idle = !exec || has(ctx.journal, JT.fenced, exec) || !has(ctx.journal, "selected", exec);
896
1051
  if (!idle && !has(ctx.journal, "tracked", exec))
897
1052
  return { action: "reject", reason: "call-starting" };
898
1053
  if (!idle && pendingSwitch(exec))
899
1054
  return { action: "reject", reason: "switch-pending" };
900
- if (!idle) {
901
- const held = holdings().filter(h => h.exec === exec);
902
- if (!held.some(h => h.pool === provider)) {
903
- const target = holdings().filter(h => h.pool === provider);
904
- if (!capacity({ kind: "provider", holders: target.length, capacity: settings().providers?.[provider]?.slots ?? Infinity }))
905
- return { action: "reject", reason: "provider-full" };
906
- let slot = 0;
907
- while (target.some(h => h.slot === slot))
908
- slot++;
909
- await orch.append("hold", { pool: provider, slot, exec, reserved: true, rid });
910
- }
911
- }
912
- await note(req.rid, model, idle ? "next-execution" : "next-request");
1055
+ if (!idle && !await reserveSwitch(exec, body.provider, rid))
1056
+ return { action: "reject", reason: "provider-full" };
1057
+ await note(req.rid, model, idle ? "next-execution" : "next-request", pool);
913
1058
  const entry = await ctx.journal.append("forward", { rid, rid2, dest, hash, envelope });
914
1059
  await replayForward(entry);
915
1060
  if (idle) {
@@ -1005,6 +1150,7 @@ export default function createExecutor(ledgers, options = {}) {
1005
1150
  },
1006
1151
  async recover(wid, journal) {
1007
1152
  journals.set(wid, journal);
1153
+ await serial(() => resolveWorktrees());
1008
1154
  await effects.recover(journal);
1009
1155
  for (const e of journal.entries().filter(e => e.type === JT.exec)) {
1010
1156
  // F1: an exec whose fence fails stays unfenced and holding; its call parks until a sweep retires it.
@@ -123,6 +123,9 @@ export async function observeExecution(d) {
123
123
  event.type === "message_end" && message?.stopReason !== "error" && (message?.usage?.output ?? 0) > 0)
124
124
  progress = performance.now();
125
125
  enqueue(async () => {
126
+ const args = event.args;
127
+ if (event.type === "tool_execution_start" && (event.toolName === "edit" || event.toolName === "write") && typeof args?.path === "string")
128
+ await d.wrote?.(args.path);
126
129
  const slim = observation(event);
127
130
  if (slim) {
128
131
  await serial(() => t.journal.append("observation", { exec, event: slim }));
@@ -0,0 +1,152 @@
1
+ import { lstat, realpath } from "node:fs/promises";
2
+ import { homedir } from "node:os";
3
+ import { dirname, join, resolve } from "node:path";
4
+ import { fileURLToPath } from "node:url";
5
+ import { JT, isEntry } from "../../types.js";
6
+ const UNICODE_SPACES = /[\u00A0\u2000-\u200A\u202F\u205F\u3000]/g;
7
+ /** The file an edit/write tool call names, resolved as pi's tools resolve it (`resolveToCwd`): unicode spaces, an `@`
8
+ * prefix, `~` and file URLs. */
9
+ export function toolPath(cwd, input) {
10
+ let path = input.replace(UNICODE_SPACES, " ");
11
+ if (path.startsWith("@"))
12
+ path = path.slice(1);
13
+ if (path === "~")
14
+ path = homedir();
15
+ else if (path.startsWith("~/"))
16
+ path = join(homedir(), path.slice(2));
17
+ else if (/^file:\/\//.test(path))
18
+ path = fileURLToPath(path);
19
+ return resolve(cwd, path);
20
+ }
21
+ /** The worktree an observed edit/write path is in: the nearest real ancestor directory holding `.git` (a directory, or
22
+ * a file in a linked worktree), for files and parent directories not created yet too. No git process and no cache: a
23
+ * repository made inside another while calls run is seen at once; each lookup is a few `lstat`s. */
24
+ export function worktreeRoots() {
25
+ return async (cwd, path) => {
26
+ // A path pi's tool cannot use either (e.g. a file URL with a host) is no evidence; the tool reports it to the agent.
27
+ let dir;
28
+ try {
29
+ dir = dirname(toolPath(cwd, path));
30
+ }
31
+ catch {
32
+ return;
33
+ }
34
+ for (;;) {
35
+ try {
36
+ dir = await realpath(dir);
37
+ break;
38
+ }
39
+ catch (error) {
40
+ if (!["ENOENT", "ENOTDIR"].includes(error.code ?? ""))
41
+ return;
42
+ const parent = dirname(dir);
43
+ if (parent === dir)
44
+ return;
45
+ dir = parent;
46
+ }
47
+ }
48
+ for (;;) {
49
+ try {
50
+ await lstat(join(dir, ".git"));
51
+ return dir;
52
+ }
53
+ catch (error) {
54
+ if (error.code !== "ENOENT")
55
+ return;
56
+ }
57
+ const parent = dirname(dir);
58
+ if (parent === dir)
59
+ return;
60
+ dir = parent;
61
+ }
62
+ };
63
+ }
64
+ /** Pair identity is durable; it also lets readers and seal recovery identify both participants. */
65
+ export const worktreePair = (a, b) => `worktree:${[a, b].sort().join("|")}`;
66
+ export const worktreeCalls = (id) => id.startsWith("worktree:") ? id.slice(9).split("|") : [];
67
+ export const worktreeLabel = (call) => call.replace(/@\d+\//, "/").replace(/@\d+$/, "");
68
+ const callOf = (exec) => exec.slice(0, exec.lastIndexOf("#"));
69
+ /** What the journals say about observed writes, kept up to date incrementally (each journal entry is read once):
70
+ * the first `wrote` of each call per worktree, calls that ended (sealed or retired), the conflict reminders written and
71
+ * whether each is still open, and each workflow's origin. Participants come from the journals, not from the calls
72
+ * running now: a paused call that wrote still counts until it ends. */
73
+ export class WorktreeIndex {
74
+ scanned = new Map();
75
+ /** Writers per root that have not ended; a root is dropped when its last writer ends. */
76
+ byRoot = new Map();
77
+ callRoots = new Map();
78
+ execRoots = new Set();
79
+ endedCalls = new Set();
80
+ reminders = new Map();
81
+ origins = new Map();
82
+ scan(journals) {
83
+ for (const journal of journals) {
84
+ const entries = journal.entries();
85
+ for (let i = this.scanned.get(journal) ?? 0; i < entries.length; i++)
86
+ this.apply(journal, entries[i]);
87
+ this.scanned.set(journal, entries.length);
88
+ }
89
+ return this;
90
+ }
91
+ apply(journal, e) {
92
+ if (e.type === "wrote" && typeof e.exec === "string" && typeof e.root === "string") {
93
+ const call = callOf(e.exec);
94
+ this.execRoots.add(`${e.exec}\0${e.root}`);
95
+ if (this.endedCalls.has(call))
96
+ return;
97
+ const calls = this.byRoot.get(e.root) ?? new Map();
98
+ this.byRoot.set(e.root, calls);
99
+ if (!calls.has(call))
100
+ calls.set(call, { call, journal, ...(Array.isArray(e.after) ? { after: e.after.map(String) } : {}) });
101
+ const roots = this.callRoots.get(call) ?? new Set();
102
+ roots.add(e.root);
103
+ this.callRoots.set(call, roots);
104
+ }
105
+ else if (e.type === JT.sealed || e.type === "retired") {
106
+ const call = String(e.call);
107
+ this.endedCalls.add(call);
108
+ for (const root of this.callRoots.get(call) ?? []) {
109
+ const calls = this.byRoot.get(root);
110
+ calls?.delete(call);
111
+ if (calls && !calls.size)
112
+ this.byRoot.delete(root);
113
+ }
114
+ this.callRoots.delete(call);
115
+ }
116
+ else if (e.type === "wf-created")
117
+ this.origins.set(journal, e.origin);
118
+ else if (isEntry(e, JT.attention) && e.item?.kind === "conflict") {
119
+ const held = this.reminders.get(e.item.id) ?? new Map();
120
+ this.reminders.set(e.item.id, held);
121
+ if (!held.has(journal))
122
+ held.set(journal, { item: e.item, open: true });
123
+ }
124
+ else if (e.type === JT.attentionResolved && typeof e.id === "string" && e.id.startsWith("worktree:")) {
125
+ const held = this.reminders.get(e.id)?.get(journal);
126
+ if (held)
127
+ held.open = false;
128
+ }
129
+ }
130
+ has(exec, root) { return this.execRoots.has(`${exec}\0${root}`); }
131
+ ended(call) { return this.endedCalls.has(call); }
132
+ /** Roots where two or more calls that have not ended wrote: the only ones a reminder can be due for. */
133
+ contested() { return [...this.byRoot].filter(([, calls]) => calls.size > 1).map(([root]) => root); }
134
+ /** Calls that wrote in `root` and have not ended. */
135
+ live(root) { return [...this.byRoot.get(root)?.values() ?? []]; }
136
+ origin(journal) { return this.origins.get(journal); }
137
+ /** The reminder of a pair as first written (another journal may still lack its copy after a crash). */
138
+ reminder(id) { return this.reminders.get(id)?.values().next().value?.item; }
139
+ remindedIn(id, journal) { return this.reminders.get(id)?.has(journal) ?? false; }
140
+ open() {
141
+ return [...this.reminders].flatMap(([id, held]) => [...held].filter(([, r]) => r.open).map(([journal, r]) => ({ id, journal, item: r.item })));
142
+ }
143
+ /** First and second writer of a pair, from the journals: a write records the calls that had written there before it
144
+ * (`after`), so the order survives a crash before the reminder; writes without it are ordered by call id. */
145
+ static order(x, y) {
146
+ if (y.after?.includes(x.call))
147
+ return [x, y];
148
+ if (x.after?.includes(y.call))
149
+ return [y, x];
150
+ return x.call < y.call ? [x, y] : [y, x];
151
+ }
152
+ }
@@ -5,6 +5,7 @@
5
5
  // provider-*: used-up usage windows (providers.ts).
6
6
  // config{hash,config}: the orchestrator settings in effect from here; config-rejected{error,hash?}: a change of
7
7
  // config.json refused, while the earlier settings stay (cleared by the next config record).
8
+ // orchestrator{version,pid} / orchestrator-exit{pid}: the orchestrator running (its package version) and its exit.
8
9
  import { foldExhaustion } from "./providers.js";
9
10
  import { isEntry } from "../types.js";
10
11
  export function emptyLedger() {
@@ -29,6 +30,12 @@ export function applyLedger(state, e) {
29
30
  }
30
31
  else if (isEntry(e, "config-rejected"))
31
32
  state.rejected = { error: String(e.error), ts: e.ts };
33
+ else if (isEntry(e, "orchestrator"))
34
+ state.orchestrator = { version: String(e.version), pid: Number(e.pid), ...(e.start ? { start: String(e.start) } : {}), ts: e.ts };
35
+ else if (isEntry(e, "orchestrator-exit")) {
36
+ if (state.orchestrator?.pid === e.pid)
37
+ state.orchestrator.exited = true;
38
+ }
32
39
  foldExhaustion(state.exhausted, e);
33
40
  }
34
41
  /** Fold the entries appended since `state` was last folded (the ledger only grows); a ledger that is not the one
@@ -36,7 +43,7 @@ export function applyLedger(state, e) {
36
43
  export function foldLedger(state, entries) {
37
44
  // Journal readers keep the entry objects of a ledger as it grows (kernel/journal.ts), so identity tells them apart.
38
45
  if (state.seen && entries[state.seen - 1] !== state.last)
39
- Object.assign(state, emptyLedger(), { config: undefined, rejected: undefined, last: undefined });
46
+ Object.assign(state, emptyLedger(), { config: undefined, rejected: undefined, orchestrator: undefined, last: undefined });
40
47
  for (; state.seen < entries.length; state.seen++)
41
48
  applyLedger(state, entries[state.seen]);
42
49
  state.last = entries[state.seen - 1];
@@ -13,8 +13,10 @@ import { join, resolve } from 'node:path';
13
13
  import { dsaHome, orchLedger, orchLock } from "../paths.js";
14
14
  import { openJournal } from "../kernel/journal.js";
15
15
  import { OsLock } from "../platform/lock.js";
16
+ import { captureStart } from "../platform/proctable.js";
16
17
  import { Engine } from "./engine.js";
17
18
  import { configPath, configProblem, configStamp, recordConfig, stampedConfig, watchConfig } from "./config.js";
19
+ import { packageVersion } from "../version.js";
18
20
  /** P2, C4, K6: Acquire the single authority before opening ledgers, recover, and idle-exit. */
19
21
  export async function main(options = {}) {
20
22
  const home = options.home ?? dsaHome();
@@ -36,6 +38,10 @@ export async function main(options = {}) {
36
38
  config = raw;
37
39
  }
38
40
  ledgers = { home, config, orch: await openJournal(orchLedger(home)) };
41
+ // Which version runs is visible to every pi session (status; a notice when it differs from the one pi loaded).
42
+ // `start` (Linux) tells this process from a later one given the same pid after a crash.
43
+ const start = await captureStart(process.pid).catch(() => '');
44
+ await ledgers.orch.append('orchestrator', { version: packageVersion(), pid: process.pid, ...(start ? { start } : {}) });
39
45
  const factory = options.executor ?? (await import(__rewriteRelativeImportExtension(new URL(import.meta.url.endsWith('.ts') ? './executor/index.ts' : './executor/index.js', import.meta.url).href))).default;
40
46
  const executor = factory(ledgers);
41
47
  engine = new Engine(ledgers, executor, options);
@@ -56,6 +62,7 @@ export async function main(options = {}) {
56
62
  }
57
63
  finally {
58
64
  try {
65
+ await ledgers?.orch.append('orchestrator-exit', { pid: process.pid }).catch(() => { });
59
66
  await ledgers?.orch.close();
60
67
  }
61
68
  finally {
@@ -7,6 +7,8 @@ import { readJournalSnapshot } from "../kernel/journal.js";
7
7
  import { journalPath, orchLedger, pinnedDir, workflowDir } from "../paths.js";
8
8
  import { JT, isEntry } from "../types.js";
9
9
  import { emptyLedger, foldLedger } from "./ledger.js";
10
+ import { packageVersion } from "../version.js";
11
+ import { worktreeCalls, worktreeLabel } from "./executor/worktree.js";
10
12
  /** A workflow has live work: it runs, or follow-ups opened on it after it finished have not ended yet. */
11
13
  export function isLive(wf) {
12
14
  return wf.status === "running" || (wf.followUps ?? 0) > 0;
@@ -230,6 +232,12 @@ function snapshotReducer(wid, entries) {
230
232
  const list = structuredClone([...calls.values()]);
231
233
  const openAttention = structuredClone(attention.filter(item => !resolved.has(`${item.id}@${item.rev}`)));
232
234
  for (const item of openAttention) {
235
+ if (item.kind === "conflict") {
236
+ const pair = worktreeCalls(item.id);
237
+ for (const c of list)
238
+ if (pair.includes(c.callId))
239
+ c.sharedWorktree = [...new Set([...(c.sharedWorktree ?? []), ...pair.filter(id => id !== c.callId).map(worktreeLabel)])].sort();
240
+ }
233
241
  const call = item.kind === "question" ? list.find(c => item.id.startsWith(`q:${c.callId}:`)) : undefined;
234
242
  if (call && (call.phase === "running" || call.phase === "queued"))
235
243
  call.phase = "asking"; // a hibernated asker is fenced
@@ -253,7 +261,7 @@ function snapshotReducer(wid, entries) {
253
261
  const pending = forwarded.filter(pendingMessage).length;
254
262
  if (pending)
255
263
  c.pending = pending;
256
- const last = c.phase !== "sealed" && !retired.has(c.callId) ? forwarded.findLast(s => s.kind === "model" && s.reason !== "withdrawn") : undefined;
264
+ const last = c.phase !== "sealed" && !retired.has(c.callId) ? forwarded.findLast(s => s.kind === "model" && s.reason !== "withdrawn" && s.reason !== "stale-execution") : undefined;
257
265
  if (last?.reason !== undefined)
258
266
  c.switchFailed = `${last.model} (${last.reason})`;
259
267
  else if (last && last.state !== "retired" && last.model !== c.model?.replace(/:(off|minimal|low|medium|high|xhigh|max)$/, ""))
@@ -420,6 +428,7 @@ export function compactWorkflow(wf) {
420
428
  calls: wf.calls.map(c => {
421
429
  const r = c.result, last = r?.output?.split("\n").map(l => l.trim()).filter(Boolean).at(-1);
422
430
  return { key: c.key, gen: c.gen, callId: c.callId, phase: c.phase, ...(r ? { status: r.status, ok: r.ok } : {}),
431
+ ...(c.sharedWorktree ? { sharedWorktree: c.sharedWorktree } : {}),
423
432
  ...(c.model ? { model: c.model } : {}), ...(c.tools ? { tools: c.tools } : {}), ...(c.pending ? { pending: c.pending } : {}), ...(c.switching ? { switching: c.switching } : {}), ...(c.switchFailed ? { switchFailed: c.switchFailed } : {}), ...(nonzero(c.usage) ? { usage: c.usage } : {}),
424
433
  ...(last ? { lastLine: clip(last, 200) } : {}), ...(r?.error ? { error: clip(r.error, 300) } : {}), ...(c.hibernated ? { hibernated: true } : {}) };
425
434
  }),
@@ -471,6 +480,57 @@ const latestCalls = (wf) => [...new Map(wf.calls.map(c => [c.key, c])).values()]
471
480
  * the settings in effect (the latest config{hash,config}); config-rejected after it is reported too. */
472
481
  // One fold per orchestrator ledger, extended as it grows (ledger.ts; the executor folds the same entries the same way).
473
482
  const ledgerStates = new Map();
483
+ /** The orchestrator running now (its last `orchestrator` record, without an exit, whose process lives), and a note when
484
+ * its version is not the one this process loaded: running work stays on the version it started with. */
485
+ export function orchestratorView(state, loaded = packageVersion()) {
486
+ const o = state.orchestrator;
487
+ if (!o || o.exited || !alive(o.pid, o.start))
488
+ return {};
489
+ return { orchestrator: `${o.version} (pid ${o.pid})`, ...(o.version === loaded ? {} : { versionNote: versionNote(o.version, loaded) }) };
490
+ }
491
+ /** Whether the recorded orchestrator still runs. On Linux its start time also tells it from a later process given the
492
+ * same pid after a crash (elsewhere the pid alone is checked). */
493
+ function alive(pid, start) {
494
+ try {
495
+ process.kill(pid, 0);
496
+ }
497
+ catch (error) {
498
+ if (error.code !== "EPERM")
499
+ return false;
500
+ }
501
+ if (!start || process.platform !== "linux")
502
+ return true;
503
+ try {
504
+ const stat = readFileSync(`/proc/${pid}/stat`, "utf8");
505
+ return stat.slice(stat.lastIndexOf(")") + 2).split(" ")[19] === start;
506
+ }
507
+ catch {
508
+ return false;
509
+ }
510
+ }
511
+ const parts = (v) => v.split(/[.-]/).map(n => Number.parseInt(n, 10) || 0);
512
+ function newer(a, b) {
513
+ const x = parts(a), y = parts(b);
514
+ for (let i = 0; i < Math.max(x.length, y.length); i++)
515
+ if ((x[i] ?? 0) !== (y[i] ?? 0))
516
+ return (x[i] ?? 0) > (y[i] ?? 0);
517
+ return false;
518
+ }
519
+ export function versionNote(running, loaded) {
520
+ return newer(running, loaded)
521
+ ? `this pi session loaded durable-subagents ${loaded}, older than the running orchestrator ${running}; start a new pi session to use ${running}`
522
+ : `the orchestrator runs durable-subagents ${running}, this pi loaded ${loaded}: running work stays on ${running}. ` +
523
+ `It exits about 10 s after all work ends and starts again on ${loaded}. To switch sooner without stopping running calls: ` +
524
+ `drain (running calls finish, nothing new starts in existing workflows; calls waiting for an answer keep it running), ` +
525
+ `wait until status no longer shows ${running}, then resume from a pi session started after the update. ` +
526
+ `A resume before it exits keeps ${running}, and pi sessions started before the update start ${running} again`;
527
+ }
528
+ /** orchestratorView of the home's orchestrator ledger. */
529
+ export function runningOrchestrator(home) {
530
+ const path = orchLedger(home), state = foldLedger(ledgerStates.get(path) ?? emptyLedger(), readJournalSnapshot(path));
531
+ ledgerStates.set(path, state);
532
+ return orchestratorView(state);
533
+ }
474
534
  export function slotsView(home, now = Date.now()) {
475
535
  const path = orchLedger(home), state = foldLedger(ledgerStates.get(path) ?? emptyLedger(), readJournalSnapshot(path));
476
536
  ledgerStates.set(path, state);
@@ -484,7 +544,7 @@ export function slotsView(home, now = Date.now()) {
484
544
  holders.set(e.pool, (holders.get(e.pool) ?? 0) + 1);
485
545
  const names = [...new Set([...Object.keys(limits), ...holders.keys()])].sort();
486
546
  const slots = names.map(p => { const n = holders.get(p) ?? 0, limit = limits[p]?.slots; return typeof limit === "number" ? `${p} ${n}/${limit}` : `${p} ${n} (no limit)`; });
487
- return { ...(slots.length ? { slots } : {}), ...(config ? { config: `${config.hash} since ${age(now - config.ts)} ago` } : {}),
547
+ return { ...orchestratorView(state), ...(slots.length ? { slots } : {}), ...(config ? { config: `${config.hash} since ${age(now - config.ts)} ago` } : {}),
488
548
  ...(exhausted.length ? { exhausted } : {}),
489
549
  ...(rejected ? { configRejected: `${clip(rejected.error, 200)} (${age(now - rejected.ts)} ago); ${config ? config.hash : "the start settings"} stay in effect` } : {}) };
490
550
  }
@@ -504,6 +564,7 @@ export function statusBrief(home, options = {}) {
504
564
  const calls = latestCalls(w).filter(c => c.phase !== "sealed" || (c.result && !c.result.ok)).map((c) => {
505
565
  const live = c.phase !== "sealed", quiet = c.lastActivity !== undefined ? now - c.lastActivity : undefined;
506
566
  return { key: c.key, agent: c.agent, phase: c.phase, ...(c.model ? { model: c.model } : {}),
567
+ ...(c.sharedWorktree ? { sharedWorktree: c.sharedWorktree } : {}),
507
568
  ...(live && c.startedAt !== undefined ? { for: age(now - c.startedAt) } : {}),
508
569
  ...(live && c.phase !== "asking" && quiet !== undefined && quiet >= 60_000 ? { quiet: age(quiet) } : {}),
509
570
  ...(live && c.startedAt !== undefined ? { tokens: tokens(c.usage) } : {}),
package/dist/ui/cards.js CHANGED
@@ -6,6 +6,7 @@ const HEAD = {
6
6
  finished: { icon: "✓", title: "finished", tone: "success" },
7
7
  stall: { icon: "…", title: "no activity", tone: "warning" },
8
8
  unknown: { icon: "!", title: "outcome unknown", tone: "warning" },
9
+ conflict: { icon: "!", title: "shares a worktree", tone: "warning" },
9
10
  budget: { icon: "$", title: "budget reached", tone: "warning" },
10
11
  };
11
12
  const keyOf = (item) => item.call ? item.call.split("/").at(-1).replace(/@1$/, "") : item.wid;
@@ -0,0 +1,14 @@
1
+ // The version of this package, read from its package.json (src/ and dist/ both sit one level below it).
2
+ import { readFileSync } from "node:fs";
3
+ let cached;
4
+ export function packageVersion() {
5
+ if (cached === undefined) {
6
+ try {
7
+ cached = String(JSON.parse(readFileSync(new URL("../package.json", import.meta.url), "utf8")).version ?? "unknown");
8
+ }
9
+ catch {
10
+ cached = "unknown";
11
+ }
12
+ }
13
+ return cached;
14
+ }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-durable-subagents",
3
- "version": "1.0.13",
3
+ "version": "1.0.16",
4
4
  "description": "Subagents for pi that never lose work and never do it twice. Crash-safe workflows, automatic recovery, and a live view just like the main agent.",
5
5
  "type": "module",
6
6
  "license": "MIT",