pi-durable-subagents 1.0.8 → 1.0.10

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,44 @@
1
1
  # Changelog
2
2
 
3
+ ## 1.0.10
4
+
5
+ - Provider failover for a used-up usage window. An error such as `503 No
6
+ available accounts`, "usage limit" or "quota exceeded" (after pi's own
7
+ retries) marks the provider used up instead of counting as a lost
8
+ execution: a call in a pool continues in the same session on the pool's next
9
+ model, and new calls skip the provider. After `k.probeMs` (15 minutes) the
10
+ next call that wants it is admitted to it alone; when it answers, new calls
11
+ and new generations go back to it. A call with a single model waits for the
12
+ provider instead of failing. `status` lists used-up providers with their
13
+ next try. Billing errors (402, insufficient balance) still fail at once.
14
+ A short request rate limit ("429 … resets in 1 second") is not a used-up
15
+ window and is retried as before. A follow-up naming a model runs on it even
16
+ where the pool would start over.
17
+ - After a call moved to another model by a relaunch, the orchestrator now
18
+ reads the session's model as pi restores it (from the last answer), so a
19
+ later relaunch holds the slot of the provider it actually uses.
20
+
21
+ ## 1.0.9
22
+
23
+ - `status` shows the model a call actually uses: the model of its last
24
+ provider request. A requested switch not used yet shows as `switching`
25
+ (also in the brief status), a switch the subagent refused as
26
+ `switchFailed`. Before, `status` kept the model the execution started with,
27
+ so a switch that had worked looked as if it had not.
28
+ - `send follow-up` with `model` runs the new generation on that model; it was
29
+ ignored, and the generation continued on the session's model. A send that
30
+ names a model replies with `model` and `effect`: `next-request` (a running
31
+ call switches at its next request), `next-execution` or `next-generation`.
32
+ - `send model` to a call that is asking (hibernated) or still waiting for a
33
+ slot is recorded and applied when it runs again, instead of being refused
34
+ with `call-not-running`.
35
+ A requested model applies once: afterwards the session and the pool choose
36
+ as before. Withdrawing a follow-up that names a model withdraws the switch
37
+ too.
38
+ - A provider's refusal of the content (terms of service, usage policy) fails
39
+ the call at once with that error. It was retried as a lost execution five
40
+ times and reported as `lost ×5`.
41
+
3
42
  ## 1.0.8
4
43
 
5
44
  - pi stays responsive with a long history: the extension's polling in pi's
package/README.md CHANGED
@@ -55,6 +55,7 @@ the npx cache, so `install-service` refuses to run from there.
55
55
  | You steer a subagent while it is asking you a question | Your message reaches it, in order. Nothing is rejected or lost. |
56
56
  | Two steers arrive out of order and the second replaces the first | Only the second one applies. |
57
57
  | A step is refused, or a dependency fails | The workflow stops that branch cleanly. Nothing is retried in vain. |
58
+ | A provider's usage window runs out (`No available accounts`, usage limit, quota exceeded) | A call in a pool continues **in the same session** on the pool's next model; new calls skip that provider. After 15 minutes the next call that wants it tries it once; when it answers, new calls and new generations use it again. A call with a single model waits for it instead of failing. Billing errors (402, insufficient balance) still fail at once. |
58
59
  | A subagent waits for an answer for a long time | It releases its model slot and memory, then resumes exactly once when you answer. |
59
60
 
60
61
  ## Use it
@@ -77,6 +78,12 @@ answer) or failed, and one line per finished workflow. `status` with a wid
77
78
  shows one workflow with outputs clipped; add `key` for one call's full result,
78
79
  or `full: true` for everything. When a run replies `{submitted: {rid}}`
79
80
  (its workflow was not created within 10 s), the rid works wherever a wid does.
81
+ A call's `model` in `status` is the model its last provider request used; a
82
+ requested switch not used yet shows as `switching`, a refused one as
83
+ `switchFailed`. A send naming a model replies with `model` and `effect`
84
+ (`next-request`, `next-execution` or `next-generation`).
85
+ A provider's refusal of the content (terms of service, usage policy) fails the
86
+ call at once with that error instead of retrying it.
80
87
  With `tasks` or `chain`, top-level `model`, `timeoutMs`, `budget`, `isolation`,
81
88
  `context`, `tools`, `skills` and `once` apply to every step that does not set
82
89
  its own; other call fields there, and any of them beside a workflow script,
@@ -90,9 +97,9 @@ it something. Each verb means one thing, and a refusal says what would work:
90
97
  |---|---|---|
91
98
  | `run` | — | Start one subagent, `tasks` in parallel, a `chain`, or a workflow script. An unknown agent name is refused before anything starts, with the list of agents. |
92
99
  | `send steer` | a running subagent | Reaches it at its next safe point. To a finished one: refused, use `follow-up`; To one waiting on its question: it interrupts the question, and the subagent usually asks again; `answer` answers it. |
93
- | `send follow-up` | a finished subagent | Continues the same session as a new generation (`key@2`). |
100
+ | `send follow-up` | a finished subagent | Continues the same session as a new generation (`key@2`). With `model`, that generation runs on it. |
94
101
  | `send answer` | an open question | Answers it once. |
95
- | `send model` | any subagent | Switches its model at the next request. |
102
+ | `send model` | any subagent | A running one switches at its next request; one asking, hibernated or waiting for a slot launches on it when it runs again. |
96
103
  | `stop` | a subagent or a workflow | Final: `stopped`, usage kept, edits left as they are. |
97
104
  | `drain` / `resume` | existing workflows | A reversible hold; runs started later are not held. |
98
105
 
@@ -266,7 +273,14 @@ State lives in `~/.pi/durable-subagents`; set `DSA_HOME` to move it.
266
273
  ```
267
274
 
268
275
  - **Pools:** a model can name a pool. The first candidate with a free slot is
269
- used, and a candidate that keeps failing is skipped for 10 minutes.
276
+ used, and a candidate that keeps failing is skipped for 10 minutes. The
277
+ order is the preference: list the provider you want to use first.
278
+ - **A used-up provider** is not sent new calls until its next try, 15 minutes
279
+ after it last refused (`"k": { "probeMs": 900000 }`). Then one call at a
280
+ time goes to it, so finding out costs no extra request. A call that moved to
281
+ another provider stays there for the rest of its generation (switching back
282
+ mid-task would lose the prompt cache); a follow-up starts on the first
283
+ candidate again. `status` lists each used-up provider with its next try.
270
284
  - **Provider slots:** never exceeded, including while a model switch is in
271
285
  progress.
272
286
  - **Memory:** new subagents wait while memory is short. Running ones are
@@ -39,6 +39,13 @@ function call(value, cwd, where) {
39
39
  return spec;
40
40
  }
41
41
  /** v12 §2: Infer unambiguous runs and normalize controls into unchanged wire bodies. */
42
+ /** P12: a send naming a model is answered with that model and when it applies — `next-request` (a running call switches
43
+ * at its next provider request), `next-execution` (a call with no live execution launches on it) or `next-generation`
44
+ * (a follow-up's new generation runs on it). From the orchestrator ledger's `send-note`. */
45
+ export function sendReceipt(ledger, rid) {
46
+ const note = ledger.find(e => e.type === "send-note" && e.rid === rid);
47
+ return note ? { model: String(note.model), effect: String(note.effect) } : {};
48
+ }
42
49
  export function request(args, cwd) {
43
50
  // v12 §2: Infer run only when one launch form is present; never guess a control verb.
44
51
  const launchForms = [args.agent !== undefined || args.task !== undefined, args.tasks !== undefined,
@@ -110,7 +117,9 @@ export function request(args, cwd) {
110
117
  const kind = string(args, "kind");
111
118
  if (!["steer", "follow-up", "answer", "model"].includes(kind))
112
119
  throw new Error("Unsupported send kind");
113
- const body = { to: string(args, "to"), kind, ...(kind === "model" ? { model: string(args, "model") } : { message: string(args, "message") }), ...(args.by === "user" ? { by: "user" } : {}) };
120
+ // follow-up may name the model its continuation runs on (P37); other kinds ignore one.
121
+ const model = kind === "model" || kind === "follow-up" && args.model !== undefined ? { model: string(args, "model") } : {};
122
+ const body = { to: string(args, "to"), kind, ...model, ...(kind === "model" ? {} : { message: string(args, "message") }), ...(args.by === "user" ? { by: "user" } : {}) };
114
123
  const cond = {};
115
124
  if (kind === "answer") {
116
125
  cond.qid = string(args, "qid");
@@ -13,7 +13,7 @@ import { dsaHome, orchInbox, orchLedger, orchLock, outboxRoot } from "../paths.j
13
13
  import { CT, JT } from "../types.js";
14
14
  import { attention, presentText, presented, resolved, unfinishedWorkflow } from "./main/snapshots.js";
15
15
  import { isLive, pausedElsewhere, statusBrief, statusCallDetail, statusCompactDetail, statusDetail, statusView, widOfRid } from "../orchestrator/snapshot.js";
16
- import { parameters, request } from "./main/tool.js";
16
+ import { parameters, request, sendReceipt } from "./main/tool.js";
17
17
  import { discoverAgents } from "../compat/agents.js";
18
18
  let noteSink;
19
19
  /** P16: Queue a UI note for the next boundary without waking the model. */
@@ -303,8 +303,9 @@ export function registerMain(pi, ui) {
303
303
  // The rid is returned so a later send can supersede this one (replaces: [rid]).
304
304
  if (decision.type === "rejected")
305
305
  return { applied: false, reason: sent.kind === "resume" && args.wid === undefined ? resumeElsewhere(String(decision.reason)) : decision.reason, rid: sent.rid };
306
- if (sent.kind !== "run")
307
- return { applied: true, rid: sent.rid };
306
+ if (sent.kind !== "run") {
307
+ return { applied: true, rid: sent.rid, ...sendReceipt(ledger(), sent.rid) };
308
+ }
308
309
  }
309
310
  if (performance.now() >= deadline || signal?.aborted)
310
311
  break;
@@ -324,8 +325,8 @@ export function registerMain(pi, ui) {
324
325
  "Durable asynchronous subagents; run returns {wid} when created (or {submitted:{rid}} while pending). A finished workflow (its notice carries every agent's result) or a question wakes you, so after starting work end your turn: never poll with sleep or repeated status. Crash recovery resumes sessions, not external side effects. Background helper processes (orchestrator, evaluator) exit by themselves about 10 s after all work ends: never kill processes or delete files to 'clean up'. When the user quits pi, this session's running workflows pause (nothing is spent); resume continues them.",
325
326
  "run (action optional for exactly one launch form): agent+task; tasks:[call specs] parallel; chain:[call specs] sequential ({previous}); workflow:'./script.js' or source (runs.run(key,spec), runs.all([...]), emit(value), args, runs.input(name)). Optional name, cwd, usageBudget, maxCalls, inputs. With tasks/chain, top-level model, timeoutMs, budget, isolation, context, tools, skills, once are defaults for every step (a step's own value wins); a workflow/source script sets them per runs.run call. timeoutMs is milliseconds of active time (a number); omit it unless a hard limit is needed. Explicit unknown agents are rejected BEFORE creation, with available names; unknown script agents fail only their call.",
326
327
  "agents: list names, descriptions, default models and source for this cwd; use these names for run.",
327
- "send to:'<wid>/<key>' (bare '<wid>' only for a single-call workflow): steer on a running call delivers at the next safe point (receipt in status/UI); a steer to a call waiting on its question interrupts the question and the subagent usually asks again — use answer to answer it; sealed → finished:<status> — use kind 'follow-up'. follow-up continues a sealed call as generation g+1 or queues after a running turn. answer: give the qid (or just the call, or nothing when one question is open); to and rev are filled in. A question that needs the user's decision goes to the user; if you answer one yourself, tell the user what you chose. model switches at next provider request. Unknown targets list valid addresses. replaces:[rid] supersedes an earlier send.",
328
- "stop target:<wid|<wid>/<key>> is terminal stopped (usage and partial edits kept); a sealed call → already-sealed:<status>, a finished workflow → terminal:<status>. drain holds existing workflows reversibly (new runs unaffected); resume [wid] releases held workflows. status: without wid, what runs, asks (with its answer address; hibernated:true holds no slot) or failed, finished workflows one line each, provider slots held/limit and the config in effect; wid: one workflow, outputs clipped; wid+key: one call's full result; full:true: everything. A run's rid from {submitted:{rid}} works wherever a wid is expected. revise wid + workflow/source/args starts a revision.",
328
+ "send to:'<wid>/<key>' (bare '<wid>' only for a single-call workflow): steer on a running call delivers at the next safe point (receipt in status/UI); a steer to a call waiting on its question interrupts the question and the subagent usually asks again — use answer to answer it; sealed → finished:<status> — use kind 'follow-up'. follow-up continues a sealed call as generation g+1 or queues after a running turn; follow-up model:'provider/id' runs that generation on it. answer: give the qid (or just the call, or nothing when one question is open); to and rev are filled in. A question that needs the user's decision goes to the user; if you answer one yourself, tell the user what you chose. model: a running call switches at its next provider request; an asking, hibernated or queued call launches on it when it runs again; the reply's model/effect (next-request|next-execution|next-generation) says which. status model = model actually used by the last request; switching = requested, not used yet; switchFailed = refused. A provider content refusal (ToS/usage policy) fails the call at once, not retried. Unknown targets list valid addresses. replaces:[rid] supersedes an earlier send.",
329
+ "stop target:<wid|<wid>/<key>> is terminal stopped (usage and partial edits kept); a sealed call → already-sealed:<status>, a finished workflow → terminal:<status>. drain holds existing workflows reversibly (new runs unaffected); resume [wid] releases held workflows. status: without wid, what runs, asks (with its answer address; hibernated:true holds no slot) or failed, finished workflows one line each, provider slots held/limit, the config in effect and providers whose usage window is used up (avoided until a probe finds them answering again); wid: one workflow, outputs clipped; wid+key: one call's full result; full:true: everything. A run's rid from {submitted:{rid}} works wherever a wid is expected. revise wid + workflow/source/args starts a revision.",
329
330
  "Control replies are {applied:true,rid} or {applied:false,reason,rid} when decided; otherwise {submitted:{rid}} after 10s.",
330
331
  ...(agents ? [`Available agents: ${agents}.`] : []),
331
332
  "User sees a summary line above the editor; ↓ on an empty editor (or /subagents) opens the list, Enter watches live OR finished calls (finished transcripts remain on disk) and expands finished workflows. List keys: s steer (paste-capable input), x stop (confirm y), m model, a answer when asked, f follow-up on finished calls; action feedback appears in footer.",
package/dist/cli/main.js CHANGED
@@ -65,13 +65,13 @@ export function renderStatus(wf) {
65
65
  /** P25, T10: Render the compact status projection shared with the `subagents` tool. */
66
66
  export function renderView(view) {
67
67
  const lines = view.workflows.map(w => [`${w.wid}@${w.rev}${w.name ? ` ${w.name}` : ""}: ${w.status}${w.followUps ? " (follow-up running)" : ""}${w.error ? ` (${clip(w.error, 200)})` : ""} · ${w.done}/${w.planned ?? w.calls.length}${w.planned === undefined && w.status === "running" ? "+" : ""} done${w.usage.input || w.usage.output || w.usage.costUsd ? ` · ${formatUsage(w.usage)}` : ""}`,
68
- ...w.calls.map(c => ` ${c.key}@${c.gen} ${c.status ?? c.phase}${c.hibernated ? " (hibernated, no slot)" : ""}${c.model ? ` ${c.model}` : ""}${c.tools ? ` tools:${c.tools}` : ""}${c.usage ? ` ${formatUsage(c.usage)}` : ""}${c.lastLine ? ` ${JSON.stringify(c.lastLine)}` : c.error ? ` (${c.error})` : ""}`),
68
+ ...w.calls.map(c => ` ${c.key}@${c.gen} ${c.status ?? c.phase}${c.hibernated ? " (hibernated, no slot)" : ""}${c.model ? ` ${c.model}` : ""}${c.switching ? ` → ${c.switching} (requested)` : ""}${c.switchFailed ? ` (switch refused: ${c.switchFailed})` : ""}${c.tools ? ` tools:${c.tools}` : ""}${c.usage ? ` ${formatUsage(c.usage)}` : ""}${c.lastLine ? ` ${JSON.stringify(c.lastLine)}` : c.error ? ` (${c.error})` : ""}`),
69
69
  ...w.attention.map(a => ` ${a.kind}: ${JSON.stringify(a.text.split("\n")[0])}`)].join("\n"));
70
70
  if (view.paused)
71
71
  lines.unshift(`${view.paused} (pi-durable-subagents resume)`);
72
72
  if (view.olderFinished)
73
73
  lines.push(`(+${view.olderFinished} older finished workflows; status <wid> shows one in detail)`);
74
- const footer = [view.slots?.length ? `slots: ${view.slots.join(", ")}` : "", view.config ? `config: ${view.config}` : "", view.configRejected ? `config.json rejected: ${view.configRejected}` : ""].filter(Boolean);
74
+ const footer = [view.slots?.length ? `slots: ${view.slots.join(", ")}` : "", view.config ? `config: ${view.config}` : "", view.configRejected ? `config.json rejected: ${view.configRejected}` : "", ...(view.exhausted ?? [])].filter(Boolean);
75
75
  if (!lines.length)
76
76
  lines.push("No workflows");
77
77
  return [...lines, ...footer].join("\n");
@@ -9,7 +9,7 @@ import { join } from "node:path";
9
9
  import { contentHash } from "../kernel/ids.js";
10
10
  /** The keys the orchestrator reads; config.json also holds pi-side settings (ui, onQuit) that it ignores. */
11
11
  const KEYS = ["defaultModel", "pools", "providers", "memory", "k"];
12
- const K = ["lossBound", "checkpointMs", "stallMs", "progressMs", "switchTimeoutMs", "idleExitMs", "trackerMs", "hibernateMs", "spawnBudget"];
12
+ const K = ["lossBound", "checkpointMs", "stallMs", "progressMs", "switchTimeoutMs", "idleExitMs", "trackerMs", "hibernateMs", "spawnBudget", "probeMs"];
13
13
  export const configPath = (home) => join(home, "config.json");
14
14
  /** The orchestrator's part of a parsed config.json. */
15
15
  export function orchestratorSettings(raw) {
@@ -22,6 +22,7 @@ import { EvaluatorClient } from "./evaluator-client.js";
22
22
  import { Store, revisionEntries, terminalEntry } from "./store.js";
23
23
  import { formatUsage, holdOf, refusedResult, snapshotFromEntries } from "./snapshot.js";
24
24
  import { validateCallSpec } from "../compat/spec.js";
25
+ import { parseModel } from "../compat/model.js";
25
26
  const tail = (text, n) => text.length > n ? `…${text.slice(-(n - 1))}` : text;
26
27
  const charged = (u) => u && (u.input || u.output || u.costUsd) ? formatUsage(u) : undefined;
27
28
  const wakeStatus = (status) => {
@@ -330,12 +331,26 @@ export class Engine {
330
331
  const seal = wf.journal.entries().find(e => e.type === JT.sealed && e.call === from);
331
332
  if (seal && send.kind === 'steer')
332
333
  return { action: 'reject', reason: `finished:${seal.result.status} — use kind "follow-up" to continue it` };
334
+ if (send.kind === 'follow-up' && send.model !== undefined) {
335
+ try {
336
+ if (!parseModel(send.model).provider)
337
+ throw new Error('missing provider');
338
+ }
339
+ catch {
340
+ return { action: 'reject', reason: 'unknown-model' };
341
+ }
342
+ }
333
343
  if (seal && send.kind === 'follow-up') {
334
344
  const gen = Math.max(0, ...wf.journal.entries().filter(e => ['call', 'generation'].includes(e.type) && e.key === entry.key).map(e => Number(e.gen))) + 1;
335
- const opened = await wf.journal.append('generation', { rid: req.rid, key: entry.key, gen, from, spec: entry.spec, revision: wf.revision, opening: { rid: req.rid, kind: send.kind, message: send.message ?? '' } });
345
+ // A follow-up's model replaces the continued session's for this generation and those continuing it.
346
+ const spec = send.model !== undefined ? { ...entry.spec, model: send.model } : entry.spec;
347
+ if (send.model !== undefined)
348
+ await this.note(req.rid, send.model, 'next-generation');
349
+ const opened = await wf.journal.append('generation', { rid: req.rid, key: entry.key, gen, from, spec, revision: wf.revision, opening: { rid: req.rid, kind: send.kind, message: send.message ?? '' }, ...(send.model !== undefined ? { model: send.model } : {}) });
336
350
  this.dispatchGeneration(wf, opened);
337
351
  return { action: 'apply' };
338
352
  }
353
+ // A follow-up naming a model, queued on unfinished work: the executor records its model request with the message.
339
354
  return this.executor.forward(req, this.context(wf, entry));
340
355
  }
341
356
  else if (req.kind === 'stop') {
@@ -504,6 +519,11 @@ export class Engine {
504
519
  this.background(async () => { throw error; });
505
520
  });
506
521
  }
522
+ /** The reply to a send that names a model says which model and when it applies (orchestrator ledger `send-note`). */
523
+ async note(rid, model, effect) {
524
+ if (!this.ledgers.orch.entries().some(e => e.type === 'send-note' && e.rid === rid))
525
+ await this.ledgers.orch.append('send-note', { rid, model, effect });
526
+ }
507
527
  ticket(st, entry) {
508
528
  const spec = entry.spec, agent = st.wf.pins.agents.find(a => a.name === spec.agent);
509
529
  if (!agent)
@@ -511,7 +531,8 @@ export class Engine {
511
531
  return { wid: st.wf.wid, widRev: `${st.wf.wid}@${st.wf.revision}`, key: entry.key, gen: entry.gen,
512
532
  callId: `${st.wf.wid}@${st.wf.revision}/${entry.key}@${entry.gen}`, spec, agent, workflowBudget: st.wf.pins.usageBudget, cwd: resolve(st.wf.cwd, spec.cwd ?? '.'), journal: st.wf.journal,
513
533
  ...(st.wf.pins.origin !== undefined ? { originSession: join(pinnedDir(this.ledgers.home, st.wf.wid), ...(st.wf.revision === 1 ? [] : [`r${st.wf.revision}`]), 'origin.jsonl') } : {}),
514
- ...(entry.type === 'generation' ? { continueFrom: entry.from, opening: entry.opening } : {}) };
534
+ ...(entry.type === 'generation' ? { continueFrom: entry.from, opening: entry.opening } : {}),
535
+ ...(entry.type === 'generation' && typeof entry.model === 'string' ? { model: entry.model } : {}) };
515
536
  }
516
537
  sealed(st, entry) {
517
538
  if (entry.type === 'refused')
@@ -22,7 +22,8 @@ import { buildCallResult } from "../../compat/result.js";
22
22
  import createEffects from "./effects/index.js";
23
23
  import { continueSession } from "./generation.js";
24
24
  import { hibernation, openQuestion } from "./hibernate.js";
25
- import { evidence, fatalProviderError, forgetSession, readSessionState, receiptId, sessionModel } from "./session.js";
25
+ import { foldExhaustion } from "../providers.js";
26
+ import { evidence, fatalProviderError, quotaExhausted, refusedByProvider, forgetSession, readSessionState, receiptId, sessionModel } from "./session.js";
26
27
  import { activeTotal } from "./time.js";
27
28
  import { observeExecution } from "./observe.js";
28
29
  import { availableMemory } from "./memory.js";
@@ -41,6 +42,39 @@ const entriesFor = (journal, call) => journal.entries().filter(e => e.call === c
41
42
  const current = (journal, call) => entriesFor(journal, call).findLast(e => e.type === JT.exec)?.exec;
42
43
  const sealed = (journal, call) => entriesFor(journal, call).find(e => e.type === JT.sealed)?.result;
43
44
  const has = (journal, type, exec) => journal.entries().some(e => e.type === type && e.exec === exec);
45
+ /** The rid of the model request a follow-up naming a model makes (P12): derived, so withdrawing or replacing the
46
+ * follow-up withdraws its model request too. */
47
+ export const modelRid = (rid) => contentHash([rid, "model"]);
48
+ /** The model a call was asked to use and has not used yet: a follow-up's `model`, then each model send in order. One the
49
+ * child rejected or that was withdrawn does not count; one already used does not either — the child applied it, or an
50
+ * execution of the call answered with it — so the session's model and the pool's fallback rule again after that.
51
+ * Its next execution launches on it (P12, P37). */
52
+ export function requestedModel(journal, call, followUp) {
53
+ const all = journal.entries(), execs = new Set(all.filter(e => e.type === JT.exec && e.call === call).map(e => String(e.exec)));
54
+ const same = (a, b) => b?.provider === a.provider && b?.id === a.id;
55
+ const usedAfter = (m, index) => all.some((e, i) => i > index && (e.type === "selected" || e.type === "model-used") && execs.has(String(e.exec)) && same(m, e.model));
56
+ let wanted;
57
+ if (followUp) {
58
+ const m = parseModel(followUp);
59
+ if (!usedAfter(m, -1))
60
+ wanted = m;
61
+ }
62
+ for (const [index, e] of all.entries()) {
63
+ if (e.type !== "forward" || e.dest !== call || e.envelope?.kind !== "model")
64
+ continue;
65
+ const delivered = all.find(r => r.type === "forward-delivered" && r.call === call && r.rid2 === e.rid2);
66
+ if (delivered) {
67
+ wanted = undefined;
68
+ continue;
69
+ } // applied by the child (now the session's model) or refused by it
70
+ if (all.some(r => r.type === "forward" && r.dest === call && r.envelope.kind === "withdraw" && r.envelope.body.rids?.includes(String(e.rid2))))
71
+ continue;
72
+ const body = e.envelope.body;
73
+ const m = { provider: body.provider, id: body.model, ...(body.thinking ? { thinking: body.thinking } : {}) };
74
+ wanted = usedAfter(m, index) ? undefined : m;
75
+ }
76
+ return wanted;
77
+ }
44
78
  function address(call) {
45
79
  const match = /^(.*)@(\d+)\/(.*)@(\d+)$/.exec(call);
46
80
  if (!match)
@@ -82,7 +116,7 @@ export default function createExecutor(ledgers, options = {}) {
82
116
  const wake = () => { for (const fn of waiters)
83
117
  fn(); waiters.clear(); };
84
118
  // F3: fold only orchestrator entries appended since the last fold (holdings, observed switches, K7 skips).
85
- const ledger = { seen: 0, held: new Map(), observed: new Set(), skips: new Map() };
119
+ const ledger = { seen: 0, held: new Map(), observed: new Set(), skips: new Map(), exhausted: new Map() };
86
120
  const folded = () => {
87
121
  const entries = orch.entries();
88
122
  for (; ledger.seen < entries.length; ledger.seen++) {
@@ -95,11 +129,17 @@ export default function createExecutor(ledgers, options = {}) {
95
129
  ledger.observed.add(`${e.exec}\n${e.rid}`);
96
130
  else if (e.type === "skip")
97
131
  ledger.skips.set(`${e.pool}\n${e.model}`, Math.max(Number(e.until), ledger.skips.get(`${e.pool}\n${e.model}`) ?? 0));
132
+ foldExhaustion(ledger.exhausted, e);
98
133
  }
99
134
  return ledger;
100
135
  };
101
136
  const holdings = () => [...folded().held.values()];
102
137
  const skipped = (pool, model) => (folded().skips.get(`${pool}\n${model.provider}/${model.id}`) ?? 0) > Date.now();
138
+ /** A provider whose usage window is used up admits no call until its next try, and then one probe at a time. */
139
+ const unavailable = (provider) => {
140
+ const x = provider ? folded().exhausted.get(provider) : undefined;
141
+ return !!x && (Date.now() < x.nextTry || x.probe !== undefined);
142
+ };
103
143
  async function release(exec) {
104
144
  await serial(async () => { for (const h of holdings().filter(e => e.exec === exec))
105
145
  await orch.append("release", { pool: h.pool, slot: h.slot, exec }); });
@@ -261,6 +301,11 @@ export default function createExecutor(ledgers, options = {}) {
261
301
  }
262
302
  }
263
303
  /** P7, P27: Record forward-delivered once when a forward's child receipt is first observed; serial sections only. */
304
+ /** The reply to a model send says which model and when it applies (orchestrator ledger `send-note`, once per rid). */
305
+ async function note(rid, model, effect) {
306
+ if (!orch.entries().some(e => e.type === "send-note" && e.rid === rid))
307
+ await orch.append("send-note", { rid, model, effect });
308
+ }
264
309
  async function forwardsDelivered(journal, call, entries) {
265
310
  const all = journal.entries();
266
311
  const open = all.filter(e => e.type === "forward" && e.dest === call &&
@@ -405,6 +450,9 @@ export default function createExecutor(ledgers, options = {}) {
405
450
  if (!continuation && pool && skipped(pool, model))
406
451
  continue;
407
452
  const provider = model.provider;
453
+ if (unavailable(provider))
454
+ continue;
455
+ const probe = provider !== undefined && folded().exhausted.has(provider);
408
456
  const holders = holdings().filter(e => e.pool === provider);
409
457
  const limit = config.providers?.[provider ?? ""]?.slots ?? Infinity;
410
458
  if (!capacity({ kind: "provider", holders: holders.length, capacity: limit }))
@@ -432,6 +480,10 @@ export default function createExecutor(ledgers, options = {}) {
432
480
  while (holders.some(e => e.slot === slot))
433
481
  slot++;
434
482
  await orch.append("hold", { pool: provider, slot, exec });
483
+ // After its next try, the first call admitted to a used-up provider is its probe: its first request
484
+ // either goes through (the provider is available again) or is refused, which uses no quota.
485
+ if (probe)
486
+ await orch.append("provider-probe", { provider, exec });
435
487
  }
436
488
  await a.ticket.journal.append("selected", { exec, model, ...(pool ? { pool } : {}) });
437
489
  return model;
@@ -460,7 +512,16 @@ export default function createExecutor(ledgers, options = {}) {
460
512
  const pool = raw && config.pools?.[raw] ? raw : undefined;
461
513
  const candidates = raw ? resolveModel(raw, config.pools) : [{ id: "" }];
462
514
  const candidate = recorded && candidates.some(m => m.provider === recorded.provider && m.id === recorded.id);
463
- const skip = previous && ownSegment && pool && candidate && skipped(pool, recorded);
515
+ // Leave the session's model for the pool's others when its pool skips it after losses, or its provider's usage
516
+ // window is used up; and at a new generation, go back to the pool's order of preference.
517
+ const skip = pool && candidate && (previous && ownSegment && skipped(pool, recorded) || unavailable(recorded.provider) || !ownSegment && !!t.continueFrom);
518
+ // A model the call was asked to use replaces the session's: launched with it, and holding its provider's slot.
519
+ const wanted = requestedModel(t.journal, t.callId, t.model);
520
+ // It outranks the pool's order at a new generation too, also when it names the model the session already has.
521
+ if (wanted)
522
+ return recorded && !freshFork && recorded.provider === wanted.provider && recorded.id === wanted.id
523
+ ? { candidates: [recorded], continuation: true, pool: undefined }
524
+ : { candidates: [wanted], continuation: false, pool: undefined };
464
525
  if (recorded && !freshFork && !skip)
465
526
  return { candidates: [recorded], continuation: true, pool: candidate ? pool : undefined };
466
527
  return { candidates, continuation: false, pool };
@@ -487,11 +548,26 @@ export default function createExecutor(ledgers, options = {}) {
487
548
  }
488
549
  });
489
550
  }
490
- async function switched(exec, event) {
491
- const provider = event.message?.provider;
551
+ /** The model each execution last answered with (`selected`, then `model-used`), cached per execution. */
552
+ const inUse = new Map();
553
+ async function switched(exec, journal, event) {
554
+ const message = event.message, provider = message?.provider;
492
555
  if (!provider)
493
556
  return;
494
557
  await serial(async () => {
558
+ // Evidence of the model in use: the provider and model of each assistant message, recorded when it changes.
559
+ if (message.role === "assistant" && message.model) {
560
+ const name = `${provider}/${message.model}`;
561
+ if (!inUse.has(exec)) {
562
+ const last = journal.entries().findLast(e => (e.type === "selected" || e.type === "model-used") && e.exec === exec)?.model;
563
+ if (last)
564
+ inUse.set(exec, `${last.provider}/${last.id}`);
565
+ }
566
+ if (inUse.get(exec) !== name) {
567
+ await journal.append("model-used", { exec, model: { provider, id: message.model } });
568
+ inUse.set(exec, name);
569
+ }
570
+ }
495
571
  const target = holdings().find(h => h.exec === exec && h.pool === provider && h.reserved);
496
572
  if (!target)
497
573
  return;
@@ -503,6 +579,32 @@ export default function createExecutor(ledgers, options = {}) {
503
579
  });
504
580
  wake();
505
581
  }
582
+ /** The model an execution last answered with, or was launched on. */
583
+ function modelOf(journal, exec) {
584
+ return journal.entries().findLast(e => (e.type === "selected" || e.type === "model-used") && e.exec === exec)?.model;
585
+ }
586
+ /** Record a used-up provider once per window: again only when its probe (or any call after the next try) is refused. */
587
+ async function recordExhausted(provider, exec, error) {
588
+ const x = folded().exhausted.get(provider), now = Date.now();
589
+ // While a probe runs, its outcome alone decides: a late refusal of an execution admitted earlier changes nothing.
590
+ if (x && (x.probe ? x.probe !== exec : now < x.nextTry))
591
+ return;
592
+ if (orch.entries().some(e => e.type === "provider-exhausted" && e.exec === exec))
593
+ return;
594
+ await orch.append("provider-exhausted", { provider, exec, since: x?.since ?? now, nextTry: now + (config.k?.probeMs ?? 900_000), error: error.slice(0, 300) });
595
+ }
596
+ /** An answer from a used-up provider, requested after it was found used up: available again. */
597
+ async function answered(exec, event) {
598
+ const message = event.message, provider = message?.provider;
599
+ if (message?.role !== "assistant" || !provider || message.stopReason === "error")
600
+ return;
601
+ await serial(async () => {
602
+ const x = folded().exhausted.get(provider);
603
+ if (x && (x.probe === exec || Number(message.timestamp) > x.since))
604
+ await orch.append("provider-available", { provider, exec });
605
+ });
606
+ wake();
607
+ }
506
608
  function pendingSwitch(exec) {
507
609
  return holdings().find(h => h.exec === exec && h.reserved && !folded().observed.has(`${exec}\n${h.rid}`));
508
610
  }
@@ -579,11 +681,20 @@ export default function createExecutor(ledgers, options = {}) {
579
681
  return finish(journal, t.callId, exec, makeResult("unknown", "", `Unknown tool outcomes: ${dangling.join(", ")}`));
580
682
  if (has(journal, "settled", exec) && !ev.text && ev.error && fatalProviderError(ev.error))
581
683
  return finish(journal, t.callId, exec, makeResult("failed", "", `Provider error: ${ev.error}`));
582
- await serial(async () => {
583
- if (!has(journal, "loss", exec))
584
- await journal.append("loss", { exec });
585
- await skipLostCandidate(journal, orch, exec);
586
- });
684
+ // A refusal of the content is deterministic: the same request is refused again, so it is reported, not retried.
685
+ if (has(journal, "settled", exec) && !ev.text && ev.error && refusedByProvider(ev.error))
686
+ return finish(journal, t.callId, exec, makeResult("failed", "", `Refused by the provider (not retried): ${ev.error.slice(0, 500)}`));
687
+ // A used-up usage window is no loss: the provider is avoided until a probe finds it accepting requests again,
688
+ // and the call goes on with the pool's next model, or waits for that provider.
689
+ const exhausted = has(journal, "settled", exec) && !ev.text && ev.error && quotaExhausted(ev.error) ? modelOf(journal, exec)?.provider : undefined;
690
+ if (exhausted)
691
+ await serial(() => recordExhausted(exhausted, exec, ev.error));
692
+ else
693
+ await serial(async () => {
694
+ if (!has(journal, "loss", exec))
695
+ await journal.append("loss", { exec });
696
+ await skipLostCandidate(journal, orch, exec);
697
+ });
587
698
  const losses = journal.entries().filter(e => e.type === "loss" && String(e.exec).startsWith(`${t.callId}#`)).length;
588
699
  if (losses >= (config.k?.lossBound ?? 5))
589
700
  return finish(journal, t.callId, exec, makeResult("failed", "", `lost ×${losses}${ev.error ? `; last error: ${ev.error.slice(0, 300)}` : ""}`));
@@ -679,7 +790,7 @@ export default function createExecutor(ledgers, options = {}) {
679
790
  }
680
791
  });
681
792
  }, recordUsage: values => recordUsage(t, values),
682
- switched: event => switched(exec, event), pendingSwitch: () => pendingSwitch(exec),
793
+ switched: event => switched(exec, journal, event), answered: event => answered(exec, event), pendingSwitch: () => pendingSwitch(exec),
683
794
  });
684
795
  }
685
796
  finally {
@@ -716,6 +827,9 @@ export default function createExecutor(ledgers, options = {}) {
716
827
  finally {
717
828
  active.delete(ticket.callId);
718
829
  collected.delete(ticket.callId);
830
+ for (const e of inUse.keys())
831
+ if (callOf(e) === ticket.callId)
832
+ inUse.delete(e);
719
833
  forgetSession(callSession(home, ticket.wid, ticket.key, ticket.gen));
720
834
  wake();
721
835
  }
@@ -763,28 +877,70 @@ export default function createExecutor(ledgers, options = {}) {
763
877
  return { action: "apply" };
764
878
  }
765
879
  }
880
+ /** P12: a model request to this call, recorded with the rid given; a reject has no effect. */
881
+ const requestModel = async (rid, model, hash, cond) => {
882
+ let body;
883
+ try {
884
+ const m = parseModel(model);
885
+ if (!m.provider)
886
+ throw new Error("Missing provider");
887
+ body = { provider: m.provider, model: m.id, ...(m.thinking ? { thinking: m.thinking } : {}) };
888
+ }
889
+ catch {
890
+ return { action: "reject", reason: "unknown-model" };
891
+ }
892
+ const envelope = { to: dest, kind: "model", body, ...(cond && Object.keys(cond).length ? { cond } : {}) };
893
+ const rid2 = forwardRid(rid, ctx.widRev, ctx.key, hash);
894
+ const exec = current(ctx.journal, dest), provider = body.provider;
895
+ // P28: with no live execution (not started yet, between executions, hibernated while asking) the model is
896
+ // recorded and the next execution launches on it (`requestedModel`); its slot is acquired then, as for any launch.
897
+ // Launching (`selected`, not `tracked` yet): the child may start on the old model; ask again in a moment.
898
+ const idle = !exec || has(ctx.journal, JT.fenced, exec) || !has(ctx.journal, "selected", exec);
899
+ if (!idle && !has(ctx.journal, "tracked", exec))
900
+ return { action: "reject", reason: "call-starting" };
901
+ if (!idle && pendingSwitch(exec))
902
+ return { action: "reject", reason: "switch-pending" };
903
+ if (!idle) {
904
+ const held = holdings().filter(h => h.exec === exec);
905
+ if (!held.some(h => h.pool === provider)) {
906
+ const target = holdings().filter(h => h.pool === provider);
907
+ if (!capacity({ kind: "provider", holders: target.length, capacity: config.providers?.[provider]?.slots ?? Infinity }))
908
+ return { action: "reject", reason: "provider-full" };
909
+ let slot = 0;
910
+ while (target.some(h => h.slot === slot))
911
+ slot++;
912
+ await orch.append("hold", { pool: provider, slot, exec, reserved: true, rid });
913
+ }
914
+ }
915
+ await note(req.rid, model, idle ? "next-execution" : "next-request");
916
+ const entry = await ctx.journal.append("forward", { rid, rid2, dest, hash, envelope });
917
+ await replayForward(entry);
918
+ if (idle) {
919
+ active.get(dest)?.wake();
920
+ wake();
921
+ }
922
+ return undefined;
923
+ };
766
924
  let kind, body;
925
+ if (req.kind === "send" && req.body.kind === "follow-up" && req.body.model !== undefined) {
926
+ // A follow-up naming a model, queued on unfinished work: the model request and the message are recorded in one
927
+ // section, so a seal cannot come between them (both or neither). A replay finds the model request recorded.
928
+ const mrid = modelRid(req.rid);
929
+ if (!ctx.journal.entries().some(e => e.type === "forward" && e.rid === mrid && e.dest === dest)) {
930
+ const refused = await requestModel(mrid, req.body.model, contentHash([hash, "model"]));
931
+ if (refused)
932
+ return { action: "reject", reason: `model: ${refused.reason}` };
933
+ }
934
+ }
767
935
  if (req.kind === "withdraw") {
768
936
  kind = "withdraw";
769
- const targets = req.body.rids;
937
+ const targets = req.body.rids.flatMap(rid => [rid, modelRid(rid)]);
770
938
  body = { rids: ctx.journal.entries().filter(e => e.type === "forward" && e.dest === dest && targets.includes(String(e.rid))).map(e => String(e.rid2)) };
771
939
  }
772
940
  else if (req.kind === "send") {
773
941
  const send = req.body;
774
942
  kind = send.kind;
775
- if (kind === "model") {
776
- try {
777
- const m = parseModel(send.model ?? "");
778
- if (!m.provider)
779
- throw new Error("Missing provider");
780
- body = { provider: m.provider, model: m.id, ...(m.thinking ? { thinking: m.thinking } : {}) };
781
- }
782
- catch {
783
- return { action: "reject", reason: "unknown-model" };
784
- }
785
- }
786
- else
787
- body = { message: send.message ?? "" };
943
+ body = kind === "model" ? undefined : { message: send.message ?? "" };
788
944
  }
789
945
  else
790
946
  return { action: "reject", reason: "unsupported" };
@@ -797,30 +953,15 @@ export default function createExecutor(ledgers, options = {}) {
797
953
  else
798
954
  delete cond.after;
799
955
  }
956
+ if (kind === "model")
957
+ return (await requestModel(req.rid, req.body.model ?? "", hash, cond)) ?? { action: "apply" };
800
958
  const envelope = { to: dest, kind, body, ...(Object.keys(cond).length ? { cond } : {}) };
801
959
  const rid2 = forwardRid(req.rid, ctx.widRev, ctx.key, hash);
802
- if (kind === "model") {
803
- const exec = current(ctx.journal, dest), provider = body.provider;
804
- if (!exec || has(ctx.journal, JT.fenced, exec) || !has(ctx.journal, "tracked", exec))
805
- return { action: "reject", reason: "call-not-running" };
806
- if (pendingSwitch(exec))
807
- return { action: "reject", reason: "switch-pending" };
808
- const held = holdings().filter(h => h.exec === exec);
809
- if (!held.some(h => h.pool === provider)) {
810
- const target = holdings().filter(h => h.pool === provider);
811
- if (!capacity({ kind: "provider", holders: target.length, capacity: config.providers?.[provider]?.slots ?? Infinity }))
812
- return { action: "reject", reason: "provider-full" };
813
- let slot = 0;
814
- while (target.some(h => h.slot === slot))
815
- slot++;
816
- await orch.append("hold", { pool: provider, slot, exec, reserved: true, rid: req.rid });
817
- }
818
- }
819
960
  const entry = await ctx.journal.append("forward", { rid: req.rid, rid2, dest, hash, envelope });
820
961
  await replayForward(entry);
821
962
  if (kind === "withdraw") {
822
963
  const exec = current(ctx.journal, dest), reservation = exec && pendingSwitch(exec);
823
- if (reservation && req.body.rids.includes(String(reservation.rid)))
964
+ if (reservation && req.body.rids.some(rid => rid === reservation.rid || modelRid(rid) === reservation.rid))
824
965
  active.get(dest)?.wake();
825
966
  }
826
967
  return { action: "apply" };
@@ -110,8 +110,10 @@ export async function observeExecution(d) {
110
110
  await serial(() => t.journal.append("observation", { exec, event: slim }));
111
111
  if (event.type === "message_start")
112
112
  await d.switched(event);
113
- if (event.type === "message_end")
113
+ if (event.type === "message_end") {
114
114
  await d.recordUsage([{ id: String(slim.id), usage: slim.usage }]);
115
+ await d.answered?.(event);
116
+ }
115
117
  }
116
118
  await limits();
117
119
  await stall();
@@ -107,12 +107,43 @@ export function evidence(entries, exec) {
107
107
  const error = last?.stopReason === "error" ? last.errorMessage : undefined;
108
108
  return { report, budget, text, error, dangling: [...tools].map(([id, name]) => `${name} (${id})`), usage };
109
109
  }
110
- /** Only explicit quota/payment failures are terminal; rate limits, overload and transport errors still retry. */
110
+ /** Only explicit payment failures are terminal; rate limits, overload and transport errors still retry, and a used-up
111
+ * usage window (`quotaExhausted`) waits for the provider or moves to another one. */
111
112
  export function fatalProviderError(text) {
112
- return /\b402\b|insufficient[_ ]?(quota|balance|funds)|quota (exceeded|exhausted)|billing|credit balance|额度|余额|usage limit/i.test(text);
113
+ return /\b402\b|insufficient[_ ]?(quota|balance|funds)|billing|credit balance|余额/i.test(text);
113
114
  }
114
- /** P13, C8: Restore the effective provider from the native session's model changes. */
115
+ /** A provider's usage window is used up: its requests are refused (and not counted) until the window resets, hours
116
+ * later. Seen as a gateway's `503 No available accounts` once pi's own retries are spent, or a usage-limit message.
117
+ * The provider is then avoided until a probe finds it accepting requests again. */
118
+ export function quotaExhausted(text) {
119
+ if (fatalProviderError(text))
120
+ return false;
121
+ if (/no available accounts?/i.test(text))
122
+ return true;
123
+ // A request rate limit clears in seconds ("rate limit exceeded; resets in 1 second", "quota exceeded for requests
124
+ // per minute"): pi's retries and the lost-execution path handle it; it must not take the provider out for minutes.
125
+ if (/rate.?limit|too many requests|request limit|per (second|minute)|\b[RT]PM\b|resets? in \d+ ?(ms|s|secs?|seconds?|minutes?)\b/i.test(text))
126
+ return false;
127
+ return /usage limit|quota (exceeded|exhausted)|exceeded your (current )?(usage|quota)|limit (reached|exceeded)[^.]*resets?\b|额度/i.test(text);
128
+ }
129
+ /** A refusal of the request's content (terms of service, usage or content policy): the same request is refused again,
130
+ * on this provider and usually on another, so it is reported at once instead of retried as a lost execution. */
131
+ export function refusedByProvider(text) {
132
+ // A content filter that is down ("temporarily unavailable, please retry") is a transient failure, not a refusal.
133
+ if (/temporar|unavailable|try again|retry|timed? ?out|overloaded/i.test(text))
134
+ return false;
135
+ return /terms of service|usage polic(y|ies)|acceptable use|content[_ ]?(policy|filter|management policy)|safety (system|filter)|flagged as (unsafe|harmful)/i.test(text);
136
+ }
137
+ /** P13, C8: Restore the effective provider as pi does: from the last model change or assistant message. */
115
138
  export function sessionModel(entries) {
116
- const last = entries.findLast(e => e.type === "model_change" && e.provider && e.modelId);
117
- return last ? { provider: last.provider, id: last.modelId } : undefined;
139
+ // pi restores the model of the last model change or assistant message: a relaunch with `--model` on an existing
140
+ // session records no model change, so only the answer tells which model the session went on with.
141
+ const last = entries.findLast(e => e.type === "model_change" && e.provider && e.modelId
142
+ || e.type === "message" && e.message?.role === "assistant" && !!e.message.provider && !!e.message.model);
143
+ if (!last)
144
+ return undefined;
145
+ if (last.type === "model_change")
146
+ return { provider: last.provider, id: last.modelId };
147
+ const m = last.message;
148
+ return { provider: m.provider, id: m.model };
118
149
  }
@@ -0,0 +1,21 @@
1
+ /** Apply one orchestrator ledger entry to the map of used-up providers. */
2
+ export function foldExhaustion(exhausted, e) {
3
+ const provider = String(e.provider ?? e.pool ?? "");
4
+ if (e.type === "provider-exhausted") {
5
+ // A refusal of the probe ends it; any other execution's (the executor records none while a probe runs) keeps it.
6
+ const probe = exhausted.get(provider)?.probe;
7
+ exhausted.set(provider, { since: Number(e.since), nextTry: Number(e.nextTry), error: String(e.error ?? ""), ...(probe && probe !== e.exec ? { probe } : {}) });
8
+ }
9
+ else if (e.type === "provider-available")
10
+ exhausted.delete(provider);
11
+ else if (e.type === "provider-probe") {
12
+ const x = exhausted.get(provider);
13
+ if (x)
14
+ x.probe = String(e.exec);
15
+ }
16
+ else if (e.type === "release") {
17
+ const x = exhausted.get(provider);
18
+ if (x && x.probe === e.exec)
19
+ delete x.probe;
20
+ }
21
+ }
@@ -6,6 +6,7 @@ import { compileFanout } from "../compat/fanout.js";
6
6
  import { readJournalSnapshot } from "../kernel/journal.js";
7
7
  import { journalPath, orchLedger, pinnedDir, workflowDir } from "../paths.js";
8
8
  import { JT } from "../types.js";
9
+ import { foldExhaustion } from "./providers.js";
9
10
  /** A workflow has live work: it runs, or follow-ups opened on it after it finished have not ended yet. */
10
11
  export function isLive(wf) {
11
12
  return wf.status === "running" || (wf.followUps ?? 0) > 0;
@@ -152,6 +153,12 @@ function snapshotReducer(wid, entries) {
152
153
  if (call && call.phase === "queued")
153
154
  call.phase = "running";
154
155
  }
156
+ else if (e.type === "model-used") {
157
+ // Evidence of a switch: the execution answered with another model than it launched with.
158
+ const call = byExec.get(String(e.exec)), m = e.model;
159
+ if (call && m)
160
+ call.model = m.provider ? `${m.provider}/${m.id}` : m.id;
161
+ }
155
162
  else if (e.type === "observation") {
156
163
  // Only what the agent did counts as activity; tracker scans and time checkpoints are bookkeeping.
157
164
  const call = byExec.get(String(e.exec));
@@ -195,6 +202,13 @@ function snapshotReducer(wid, entries) {
195
202
  const list = sends.get(String(e.dest)) ?? [];
196
203
  list.push(send);
197
204
  sends.set(String(e.dest), list);
205
+ // A withdrawn model request no longer stands (the executor ignores it for the next launch as well).
206
+ if (envelope?.kind === "withdraw")
207
+ for (const rid2 of envelope.body?.rids ?? []) {
208
+ const target = byRid2.get(`${e.dest}\n${rid2}`);
209
+ if (target?.kind === "model" && target.state === "pending")
210
+ target.reason = "withdrawn";
211
+ }
198
212
  byRid2.set(`${e.dest}\n${e.rid2}`, send);
199
213
  byRid2.set(String(e.rid2), send);
200
214
  }
@@ -239,9 +253,11 @@ function snapshotReducer(wid, entries) {
239
253
  const pending = forwarded.filter(pendingMessage).length;
240
254
  if (pending)
241
255
  c.pending = pending;
242
- const switching = c.phase !== "sealed" ? forwarded.findLast(s => s.kind === "model" && s.state === "pending")?.model : undefined;
243
- if (switching && switching !== c.model?.replace(/:(off|minimal|low|medium|high|xhigh|max)$/, ""))
244
- c.switching = switching;
256
+ const last = c.phase !== "sealed" && !retired.has(c.callId) ? forwarded.findLast(s => s.kind === "model" && s.reason !== "withdrawn") : undefined;
257
+ if (last?.reason !== undefined)
258
+ c.switchFailed = `${last.model} (${last.reason})`;
259
+ else if (last && last.state !== "retired" && last.model !== c.model?.replace(/:(off|minimal|low|medium|high|xhigh|max)$/, ""))
260
+ c.switching = last.model;
245
261
  }
246
262
  }
247
263
  const after = done ? list.filter(c => generations.has(c.callId) && c.phase !== "sealed" && !retired.has(c.callId)) : [];
@@ -404,7 +420,7 @@ export function compactWorkflow(wf) {
404
420
  calls: wf.calls.map(c => {
405
421
  const r = c.result, last = r?.output?.split("\n").map(l => l.trim()).filter(Boolean).at(-1);
406
422
  return { key: c.key, gen: c.gen, callId: c.callId, phase: c.phase, ...(r ? { status: r.status, ok: r.ok } : {}),
407
- ...(c.model ? { model: c.model } : {}), ...(c.tools ? { tools: c.tools } : {}), ...(c.pending ? { pending: c.pending } : {}), ...(c.switching ? { switching: c.switching } : {}), ...(nonzero(c.usage) ? { usage: c.usage } : {}),
423
+ ...(c.model ? { model: c.model } : {}), ...(c.tools ? { tools: c.tools } : {}), ...(c.pending ? { pending: c.pending } : {}), ...(c.switching ? { switching: c.switching } : {}), ...(c.switchFailed ? { switchFailed: c.switchFailed } : {}), ...(nonzero(c.usage) ? { usage: c.usage } : {}),
408
424
  ...(last ? { lastLine: clip(last, 200) } : {}), ...(r?.error ? { error: clip(r.error, 300) } : {}), ...(c.hibernated ? { hibernated: true } : {}) };
409
425
  }),
410
426
  attention: wf.attention.map(a => ({ id: a.id, rev: a.rev, kind: a.kind, text: clip(a.text, 300), ...(a.call ? { call: a.call } : {}), ...(a.qid ? { qid: a.qid } : {}) })),
@@ -454,7 +470,7 @@ const latestCalls = (wf) => [...new Map(wf.calls.map(c => [c.key, c])).values()]
454
470
  /** Provider slots from the orchestrator ledger: holders per provider (hold/release{pool,slot,exec}) and the limits of
455
471
  * the settings in effect (the latest config{hash,config}); config-rejected after it is reported too. */
456
472
  export function slotsView(home, now = Date.now()) {
457
- const held = new Map();
473
+ const held = new Map(), used = new Map();
458
474
  let config, rejected;
459
475
  for (const e of readJournalSnapshot(orchLedger(home))) {
460
476
  if (e.type === "hold")
@@ -467,7 +483,10 @@ export function slotsView(home, now = Date.now()) {
467
483
  }
468
484
  else if (e.type === "config-rejected")
469
485
  rejected = e;
486
+ foldExhaustion(used, e);
470
487
  }
488
+ const exhausted = [...used].sort(([a], [b]) => a.localeCompare(b)).map(([p, x]) => `${p} exhausted since ${age(now - x.since)} ago (${clip(x.error, 80)}), ` +
489
+ (x.probe ? `probing with ${x.probe.split("#")[0]}` : x.nextTry > now ? `next try in ${age(x.nextTry - now)}` : "next call probes it"));
471
490
  const limits = (config?.config?.providers) ?? {};
472
491
  const holders = new Map();
473
492
  for (const e of held.values())
@@ -476,6 +495,7 @@ export function slotsView(home, now = Date.now()) {
476
495
  const names = [...new Set([...Object.keys(limits), ...holders.keys()])].sort();
477
496
  const slots = names.map(p => { const n = holders.get(p) ?? 0, limit = limits[p]?.slots; return typeof limit === "number" ? `${p} ${n}/${limit}` : `${p} ${n} (no limit)`; });
478
497
  return { ...(slots.length ? { slots } : {}), ...(config ? { config: `${String(config.hash)} since ${age(now - config.ts)} ago` } : {}),
498
+ ...(exhausted.length ? { exhausted } : {}),
479
499
  ...(rejected ? { configRejected: `${clip(String(rejected.error), 200)} (${age(now - rejected.ts)} ago); ${config ? String(config.hash) : "the start settings"} stay in effect` } : {}) };
480
500
  }
481
501
  /** Tool status without a wid: what runs, what waits for an answer and what failed, with finished workflows one line each.
@@ -498,7 +518,8 @@ export function statusBrief(home, options = {}) {
498
518
  ...(live && c.phase !== "asking" && quiet !== undefined && quiet >= 60_000 ? { quiet: age(quiet) } : {}),
499
519
  ...(live && c.startedAt !== undefined ? { tokens: tokens(c.usage) } : {}),
500
520
  ...(c.result ? { status: c.result.status, ...(c.result.error ? { error: clip(c.result.error, 200) } : {}) } : {}),
501
- ...(c.hibernated ? { hibernated: true } : {}) };
521
+ ...(c.hibernated ? { hibernated: true } : {}),
522
+ ...(live && c.switching ? { switching: c.switching } : {}), ...(live && c.switchFailed ? { switchFailed: c.switchFailed } : {}) };
502
523
  });
503
524
  const asking = open.filter(a => a.kind === "question" && a.call).map(a => ({ to: `${w.wid}/${callKey(a.call)}`, ...(a.qid ? { qid: a.qid } : {}),
504
525
  ...(w.calls.some(c => c.callId === a.call && c.hibernated) ? { hibernated: true } : {}), question: clip(a.text, 300) }));
package/dist/ui/screen.js CHANGED
@@ -565,7 +565,7 @@ export class SubagentScreen {
565
565
  const facts = this.data.facts.get(c.callId), active = w.calls.filter(c => c.phase !== "sealed"), done = w.calls.length - active.length;
566
566
  const tabs = size.width < 60 ? `${c.key} ${w.calls.indexOf(c) + 1}/${w.calls.length}` : `${[...active.map(c => c.key), ...(done ? [`${done} done`] : [])].join(" · ")} ← → switch`;
567
567
  const tools = toolCount(facts?.tools), pending = pendingText(c.pending), rule = this.theme.fg("borderMuted", "─".repeat(size.width));
568
- const switching = c.switching ? ` → ${this.name(c.switching)} (at the end of this step)` : "";
568
+ const switching = c.switching ? ` → ${this.name(c.switching)} (requested)` : "";
569
569
  const head = [tabs, `${label(c)} · ${this.name(facts?.model ?? c.model)} ▾${switching} · ${facts?.thinking ?? "off"} ▾${tools ? ` · ${tools}` : ""}${pending ? ` · ${pending}` : ""}`,
570
570
  this.theme.fg("dim", this.spend(c, facts)), rule];
571
571
  const asking = w.attention.some(a => a.kind === "question" && a.call === c.callId);
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-durable-subagents",
3
- "version": "1.0.8",
3
+ "version": "1.0.10",
4
4
  "description": "Subagents for pi that never lose work and never do it twice. Crash-safe workflows, automatic recovery, and a live view just like the main agent.",
5
5
  "type": "module",
6
6
  "license": "MIT",