pi-durable-subagents 1.0.8 → 1.0.10
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +39 -0
- package/README.md +17 -3
- package/dist/agent/main/tool.js +10 -1
- package/dist/agent/main.js +6 -5
- package/dist/cli/main.js +2 -2
- package/dist/orchestrator/config.js +1 -1
- package/dist/orchestrator/engine.js +23 -2
- package/dist/orchestrator/executor/index.js +184 -43
- package/dist/orchestrator/executor/observe.js +3 -1
- package/dist/orchestrator/executor/session.js +36 -5
- package/dist/orchestrator/providers.js +21 -0
- package/dist/orchestrator/snapshot.js +27 -6
- package/dist/ui/screen.js +1 -1
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,44 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 1.0.10
|
|
4
|
+
|
|
5
|
+
- Provider failover for a used-up usage window. An error such as `503 No
|
|
6
|
+
available accounts`, "usage limit" or "quota exceeded" (after pi's own
|
|
7
|
+
retries) marks the provider used up instead of counting as a lost
|
|
8
|
+
execution: a call in a pool continues in the same session on the pool's next
|
|
9
|
+
model, and new calls skip the provider. After `k.probeMs` (15 minutes) the
|
|
10
|
+
next call that wants it is admitted to it alone; when it answers, new calls
|
|
11
|
+
and new generations go back to it. A call with a single model waits for the
|
|
12
|
+
provider instead of failing. `status` lists used-up providers with their
|
|
13
|
+
next try. Billing errors (402, insufficient balance) still fail at once.
|
|
14
|
+
A short request rate limit ("429 … resets in 1 second") is not a used-up
|
|
15
|
+
window and is retried as before. A follow-up naming a model runs on it even
|
|
16
|
+
where the pool would start over.
|
|
17
|
+
- After a call moved to another model by a relaunch, the orchestrator now
|
|
18
|
+
reads the session's model as pi restores it (from the last answer), so a
|
|
19
|
+
later relaunch holds the slot of the provider it actually uses.
|
|
20
|
+
|
|
21
|
+
## 1.0.9
|
|
22
|
+
|
|
23
|
+
- `status` shows the model a call actually uses: the model of its last
|
|
24
|
+
provider request. A requested switch not used yet shows as `switching`
|
|
25
|
+
(also in the brief status), a switch the subagent refused as
|
|
26
|
+
`switchFailed`. Before, `status` kept the model the execution started with,
|
|
27
|
+
so a switch that had worked looked as if it had not.
|
|
28
|
+
- `send follow-up` with `model` runs the new generation on that model; it was
|
|
29
|
+
ignored, and the generation continued on the session's model. A send that
|
|
30
|
+
names a model replies with `model` and `effect`: `next-request` (a running
|
|
31
|
+
call switches at its next request), `next-execution` or `next-generation`.
|
|
32
|
+
- `send model` to a call that is asking (hibernated) or still waiting for a
|
|
33
|
+
slot is recorded and applied when it runs again, instead of being refused
|
|
34
|
+
with `call-not-running`.
|
|
35
|
+
A requested model applies once: afterwards the session and the pool choose
|
|
36
|
+
as before. Withdrawing a follow-up that names a model withdraws the switch
|
|
37
|
+
too.
|
|
38
|
+
- A provider's refusal of the content (terms of service, usage policy) fails
|
|
39
|
+
the call at once with that error. It was retried as a lost execution five
|
|
40
|
+
times and reported as `lost ×5`.
|
|
41
|
+
|
|
3
42
|
## 1.0.8
|
|
4
43
|
|
|
5
44
|
- pi stays responsive with a long history: the extension's polling in pi's
|
package/README.md
CHANGED
|
@@ -55,6 +55,7 @@ the npx cache, so `install-service` refuses to run from there.
|
|
|
55
55
|
| You steer a subagent while it is asking you a question | Your message reaches it, in order. Nothing is rejected or lost. |
|
|
56
56
|
| Two steers arrive out of order and the second replaces the first | Only the second one applies. |
|
|
57
57
|
| A step is refused, or a dependency fails | The workflow stops that branch cleanly. Nothing is retried in vain. |
|
|
58
|
+
| A provider's usage window runs out (`No available accounts`, usage limit, quota exceeded) | A call in a pool continues **in the same session** on the pool's next model; new calls skip that provider. After 15 minutes the next call that wants it tries it once; when it answers, new calls and new generations use it again. A call with a single model waits for it instead of failing. Billing errors (402, insufficient balance) still fail at once. |
|
|
58
59
|
| A subagent waits for an answer for a long time | It releases its model slot and memory, then resumes exactly once when you answer. |
|
|
59
60
|
|
|
60
61
|
## Use it
|
|
@@ -77,6 +78,12 @@ answer) or failed, and one line per finished workflow. `status` with a wid
|
|
|
77
78
|
shows one workflow with outputs clipped; add `key` for one call's full result,
|
|
78
79
|
or `full: true` for everything. When a run replies `{submitted: {rid}}`
|
|
79
80
|
(its workflow was not created within 10 s), the rid works wherever a wid does.
|
|
81
|
+
A call's `model` in `status` is the model its last provider request used; a
|
|
82
|
+
requested switch not used yet shows as `switching`, a refused one as
|
|
83
|
+
`switchFailed`. A send naming a model replies with `model` and `effect`
|
|
84
|
+
(`next-request`, `next-execution` or `next-generation`).
|
|
85
|
+
A provider's refusal of the content (terms of service, usage policy) fails the
|
|
86
|
+
call at once with that error instead of retrying it.
|
|
80
87
|
With `tasks` or `chain`, top-level `model`, `timeoutMs`, `budget`, `isolation`,
|
|
81
88
|
`context`, `tools`, `skills` and `once` apply to every step that does not set
|
|
82
89
|
its own; other call fields there, and any of them beside a workflow script,
|
|
@@ -90,9 +97,9 @@ it something. Each verb means one thing, and a refusal says what would work:
|
|
|
90
97
|
|---|---|---|
|
|
91
98
|
| `run` | — | Start one subagent, `tasks` in parallel, a `chain`, or a workflow script. An unknown agent name is refused before anything starts, with the list of agents. |
|
|
92
99
|
| `send steer` | a running subagent | Reaches it at its next safe point. To a finished one: refused, use `follow-up`; To one waiting on its question: it interrupts the question, and the subagent usually asks again; `answer` answers it. |
|
|
93
|
-
| `send follow-up` | a finished subagent | Continues the same session as a new generation (`key@2`). |
|
|
100
|
+
| `send follow-up` | a finished subagent | Continues the same session as a new generation (`key@2`). With `model`, that generation runs on it. |
|
|
94
101
|
| `send answer` | an open question | Answers it once. |
|
|
95
|
-
| `send model` | any subagent |
|
|
102
|
+
| `send model` | any subagent | A running one switches at its next request; one asking, hibernated or waiting for a slot launches on it when it runs again. |
|
|
96
103
|
| `stop` | a subagent or a workflow | Final: `stopped`, usage kept, edits left as they are. |
|
|
97
104
|
| `drain` / `resume` | existing workflows | A reversible hold; runs started later are not held. |
|
|
98
105
|
|
|
@@ -266,7 +273,14 @@ State lives in `~/.pi/durable-subagents`; set `DSA_HOME` to move it.
|
|
|
266
273
|
```
|
|
267
274
|
|
|
268
275
|
- **Pools:** a model can name a pool. The first candidate with a free slot is
|
|
269
|
-
used, and a candidate that keeps failing is skipped for 10 minutes.
|
|
276
|
+
used, and a candidate that keeps failing is skipped for 10 minutes. The
|
|
277
|
+
order is the preference: list the provider you want to use first.
|
|
278
|
+
- **A used-up provider** is not sent new calls until its next try, 15 minutes
|
|
279
|
+
after it last refused (`"k": { "probeMs": 900000 }`). Then one call at a
|
|
280
|
+
time goes to it, so finding out costs no extra request. A call that moved to
|
|
281
|
+
another provider stays there for the rest of its generation (switching back
|
|
282
|
+
mid-task would lose the prompt cache); a follow-up starts on the first
|
|
283
|
+
candidate again. `status` lists each used-up provider with its next try.
|
|
270
284
|
- **Provider slots:** never exceeded, including while a model switch is in
|
|
271
285
|
progress.
|
|
272
286
|
- **Memory:** new subagents wait while memory is short. Running ones are
|
package/dist/agent/main/tool.js
CHANGED
|
@@ -39,6 +39,13 @@ function call(value, cwd, where) {
|
|
|
39
39
|
return spec;
|
|
40
40
|
}
|
|
41
41
|
/** v12 §2: Infer unambiguous runs and normalize controls into unchanged wire bodies. */
|
|
42
|
+
/** P12: a send naming a model is answered with that model and when it applies — `next-request` (a running call switches
|
|
43
|
+
* at its next provider request), `next-execution` (a call with no live execution launches on it) or `next-generation`
|
|
44
|
+
* (a follow-up's new generation runs on it). From the orchestrator ledger's `send-note`. */
|
|
45
|
+
export function sendReceipt(ledger, rid) {
|
|
46
|
+
const note = ledger.find(e => e.type === "send-note" && e.rid === rid);
|
|
47
|
+
return note ? { model: String(note.model), effect: String(note.effect) } : {};
|
|
48
|
+
}
|
|
42
49
|
export function request(args, cwd) {
|
|
43
50
|
// v12 §2: Infer run only when one launch form is present; never guess a control verb.
|
|
44
51
|
const launchForms = [args.agent !== undefined || args.task !== undefined, args.tasks !== undefined,
|
|
@@ -110,7 +117,9 @@ export function request(args, cwd) {
|
|
|
110
117
|
const kind = string(args, "kind");
|
|
111
118
|
if (!["steer", "follow-up", "answer", "model"].includes(kind))
|
|
112
119
|
throw new Error("Unsupported send kind");
|
|
113
|
-
|
|
120
|
+
// follow-up may name the model its continuation runs on (P37); other kinds ignore one.
|
|
121
|
+
const model = kind === "model" || kind === "follow-up" && args.model !== undefined ? { model: string(args, "model") } : {};
|
|
122
|
+
const body = { to: string(args, "to"), kind, ...model, ...(kind === "model" ? {} : { message: string(args, "message") }), ...(args.by === "user" ? { by: "user" } : {}) };
|
|
114
123
|
const cond = {};
|
|
115
124
|
if (kind === "answer") {
|
|
116
125
|
cond.qid = string(args, "qid");
|
package/dist/agent/main.js
CHANGED
|
@@ -13,7 +13,7 @@ import { dsaHome, orchInbox, orchLedger, orchLock, outboxRoot } from "../paths.j
|
|
|
13
13
|
import { CT, JT } from "../types.js";
|
|
14
14
|
import { attention, presentText, presented, resolved, unfinishedWorkflow } from "./main/snapshots.js";
|
|
15
15
|
import { isLive, pausedElsewhere, statusBrief, statusCallDetail, statusCompactDetail, statusDetail, statusView, widOfRid } from "../orchestrator/snapshot.js";
|
|
16
|
-
import { parameters, request } from "./main/tool.js";
|
|
16
|
+
import { parameters, request, sendReceipt } from "./main/tool.js";
|
|
17
17
|
import { discoverAgents } from "../compat/agents.js";
|
|
18
18
|
let noteSink;
|
|
19
19
|
/** P16: Queue a UI note for the next boundary without waking the model. */
|
|
@@ -303,8 +303,9 @@ export function registerMain(pi, ui) {
|
|
|
303
303
|
// The rid is returned so a later send can supersede this one (replaces: [rid]).
|
|
304
304
|
if (decision.type === "rejected")
|
|
305
305
|
return { applied: false, reason: sent.kind === "resume" && args.wid === undefined ? resumeElsewhere(String(decision.reason)) : decision.reason, rid: sent.rid };
|
|
306
|
-
if (sent.kind !== "run")
|
|
307
|
-
return { applied: true, rid: sent.rid };
|
|
306
|
+
if (sent.kind !== "run") {
|
|
307
|
+
return { applied: true, rid: sent.rid, ...sendReceipt(ledger(), sent.rid) };
|
|
308
|
+
}
|
|
308
309
|
}
|
|
309
310
|
if (performance.now() >= deadline || signal?.aborted)
|
|
310
311
|
break;
|
|
@@ -324,8 +325,8 @@ export function registerMain(pi, ui) {
|
|
|
324
325
|
"Durable asynchronous subagents; run returns {wid} when created (or {submitted:{rid}} while pending). A finished workflow (its notice carries every agent's result) or a question wakes you, so after starting work end your turn: never poll with sleep or repeated status. Crash recovery resumes sessions, not external side effects. Background helper processes (orchestrator, evaluator) exit by themselves about 10 s after all work ends: never kill processes or delete files to 'clean up'. When the user quits pi, this session's running workflows pause (nothing is spent); resume continues them.",
|
|
325
326
|
"run (action optional for exactly one launch form): agent+task; tasks:[call specs] parallel; chain:[call specs] sequential ({previous}); workflow:'./script.js' or source (runs.run(key,spec), runs.all([...]), emit(value), args, runs.input(name)). Optional name, cwd, usageBudget, maxCalls, inputs. With tasks/chain, top-level model, timeoutMs, budget, isolation, context, tools, skills, once are defaults for every step (a step's own value wins); a workflow/source script sets them per runs.run call. timeoutMs is milliseconds of active time (a number); omit it unless a hard limit is needed. Explicit unknown agents are rejected BEFORE creation, with available names; unknown script agents fail only their call.",
|
|
326
327
|
"agents: list names, descriptions, default models and source for this cwd; use these names for run.",
|
|
327
|
-
"send to:'<wid>/<key>' (bare '<wid>' only for a single-call workflow): steer on a running call delivers at the next safe point (receipt in status/UI); a steer to a call waiting on its question interrupts the question and the subagent usually asks again — use answer to answer it; sealed → finished:<status> — use kind 'follow-up'. follow-up continues a sealed call as generation g+1 or queues after a running turn. answer: give the qid (or just the call, or nothing when one question is open); to and rev are filled in. A question that needs the user's decision goes to the user; if you answer one yourself, tell the user what you chose. model switches at next provider request. Unknown targets list valid addresses. replaces:[rid] supersedes an earlier send.",
|
|
328
|
-
"stop target:<wid|<wid>/<key>> is terminal stopped (usage and partial edits kept); a sealed call → already-sealed:<status>, a finished workflow → terminal:<status>. drain holds existing workflows reversibly (new runs unaffected); resume [wid] releases held workflows. status: without wid, what runs, asks (with its answer address; hibernated:true holds no slot) or failed, finished workflows one line each, provider slots held/limit
|
|
328
|
+
"send to:'<wid>/<key>' (bare '<wid>' only for a single-call workflow): steer on a running call delivers at the next safe point (receipt in status/UI); a steer to a call waiting on its question interrupts the question and the subagent usually asks again — use answer to answer it; sealed → finished:<status> — use kind 'follow-up'. follow-up continues a sealed call as generation g+1 or queues after a running turn; follow-up model:'provider/id' runs that generation on it. answer: give the qid (or just the call, or nothing when one question is open); to and rev are filled in. A question that needs the user's decision goes to the user; if you answer one yourself, tell the user what you chose. model: a running call switches at its next provider request; an asking, hibernated or queued call launches on it when it runs again; the reply's model/effect (next-request|next-execution|next-generation) says which. status model = model actually used by the last request; switching = requested, not used yet; switchFailed = refused. A provider content refusal (ToS/usage policy) fails the call at once, not retried. Unknown targets list valid addresses. replaces:[rid] supersedes an earlier send.",
|
|
329
|
+
"stop target:<wid|<wid>/<key>> is terminal stopped (usage and partial edits kept); a sealed call → already-sealed:<status>, a finished workflow → terminal:<status>. drain holds existing workflows reversibly (new runs unaffected); resume [wid] releases held workflows. status: without wid, what runs, asks (with its answer address; hibernated:true holds no slot) or failed, finished workflows one line each, provider slots held/limit, the config in effect and providers whose usage window is used up (avoided until a probe finds them answering again); wid: one workflow, outputs clipped; wid+key: one call's full result; full:true: everything. A run's rid from {submitted:{rid}} works wherever a wid is expected. revise wid + workflow/source/args starts a revision.",
|
|
329
330
|
"Control replies are {applied:true,rid} or {applied:false,reason,rid} when decided; otherwise {submitted:{rid}} after 10s.",
|
|
330
331
|
...(agents ? [`Available agents: ${agents}.`] : []),
|
|
331
332
|
"User sees a summary line above the editor; ↓ on an empty editor (or /subagents) opens the list, Enter watches live OR finished calls (finished transcripts remain on disk) and expands finished workflows. List keys: s steer (paste-capable input), x stop (confirm y), m model, a answer when asked, f follow-up on finished calls; action feedback appears in footer.",
|
package/dist/cli/main.js
CHANGED
|
@@ -65,13 +65,13 @@ export function renderStatus(wf) {
|
|
|
65
65
|
/** P25, T10: Render the compact status projection shared with the `subagents` tool. */
|
|
66
66
|
export function renderView(view) {
|
|
67
67
|
const lines = view.workflows.map(w => [`${w.wid}@${w.rev}${w.name ? ` ${w.name}` : ""}: ${w.status}${w.followUps ? " (follow-up running)" : ""}${w.error ? ` (${clip(w.error, 200)})` : ""} · ${w.done}/${w.planned ?? w.calls.length}${w.planned === undefined && w.status === "running" ? "+" : ""} done${w.usage.input || w.usage.output || w.usage.costUsd ? ` · ${formatUsage(w.usage)}` : ""}`,
|
|
68
|
-
...w.calls.map(c => ` ${c.key}@${c.gen} ${c.status ?? c.phase}${c.hibernated ? " (hibernated, no slot)" : ""}${c.model ? ` ${c.model}` : ""}${c.tools ? ` tools:${c.tools}` : ""}${c.usage ? ` ${formatUsage(c.usage)}` : ""}${c.lastLine ? ` ${JSON.stringify(c.lastLine)}` : c.error ? ` (${c.error})` : ""}`),
|
|
68
|
+
...w.calls.map(c => ` ${c.key}@${c.gen} ${c.status ?? c.phase}${c.hibernated ? " (hibernated, no slot)" : ""}${c.model ? ` ${c.model}` : ""}${c.switching ? ` → ${c.switching} (requested)` : ""}${c.switchFailed ? ` (switch refused: ${c.switchFailed})` : ""}${c.tools ? ` tools:${c.tools}` : ""}${c.usage ? ` ${formatUsage(c.usage)}` : ""}${c.lastLine ? ` ${JSON.stringify(c.lastLine)}` : c.error ? ` (${c.error})` : ""}`),
|
|
69
69
|
...w.attention.map(a => ` ${a.kind}: ${JSON.stringify(a.text.split("\n")[0])}`)].join("\n"));
|
|
70
70
|
if (view.paused)
|
|
71
71
|
lines.unshift(`${view.paused} (pi-durable-subagents resume)`);
|
|
72
72
|
if (view.olderFinished)
|
|
73
73
|
lines.push(`(+${view.olderFinished} older finished workflows; status <wid> shows one in detail)`);
|
|
74
|
-
const footer = [view.slots?.length ? `slots: ${view.slots.join(", ")}` : "", view.config ? `config: ${view.config}` : "", view.configRejected ? `config.json rejected: ${view.configRejected}` : ""].filter(Boolean);
|
|
74
|
+
const footer = [view.slots?.length ? `slots: ${view.slots.join(", ")}` : "", view.config ? `config: ${view.config}` : "", view.configRejected ? `config.json rejected: ${view.configRejected}` : "", ...(view.exhausted ?? [])].filter(Boolean);
|
|
75
75
|
if (!lines.length)
|
|
76
76
|
lines.push("No workflows");
|
|
77
77
|
return [...lines, ...footer].join("\n");
|
|
@@ -9,7 +9,7 @@ import { join } from "node:path";
|
|
|
9
9
|
import { contentHash } from "../kernel/ids.js";
|
|
10
10
|
/** The keys the orchestrator reads; config.json also holds pi-side settings (ui, onQuit) that it ignores. */
|
|
11
11
|
const KEYS = ["defaultModel", "pools", "providers", "memory", "k"];
|
|
12
|
-
const K = ["lossBound", "checkpointMs", "stallMs", "progressMs", "switchTimeoutMs", "idleExitMs", "trackerMs", "hibernateMs", "spawnBudget"];
|
|
12
|
+
const K = ["lossBound", "checkpointMs", "stallMs", "progressMs", "switchTimeoutMs", "idleExitMs", "trackerMs", "hibernateMs", "spawnBudget", "probeMs"];
|
|
13
13
|
export const configPath = (home) => join(home, "config.json");
|
|
14
14
|
/** The orchestrator's part of a parsed config.json. */
|
|
15
15
|
export function orchestratorSettings(raw) {
|
|
@@ -22,6 +22,7 @@ import { EvaluatorClient } from "./evaluator-client.js";
|
|
|
22
22
|
import { Store, revisionEntries, terminalEntry } from "./store.js";
|
|
23
23
|
import { formatUsage, holdOf, refusedResult, snapshotFromEntries } from "./snapshot.js";
|
|
24
24
|
import { validateCallSpec } from "../compat/spec.js";
|
|
25
|
+
import { parseModel } from "../compat/model.js";
|
|
25
26
|
const tail = (text, n) => text.length > n ? `…${text.slice(-(n - 1))}` : text;
|
|
26
27
|
const charged = (u) => u && (u.input || u.output || u.costUsd) ? formatUsage(u) : undefined;
|
|
27
28
|
const wakeStatus = (status) => {
|
|
@@ -330,12 +331,26 @@ export class Engine {
|
|
|
330
331
|
const seal = wf.journal.entries().find(e => e.type === JT.sealed && e.call === from);
|
|
331
332
|
if (seal && send.kind === 'steer')
|
|
332
333
|
return { action: 'reject', reason: `finished:${seal.result.status} — use kind "follow-up" to continue it` };
|
|
334
|
+
if (send.kind === 'follow-up' && send.model !== undefined) {
|
|
335
|
+
try {
|
|
336
|
+
if (!parseModel(send.model).provider)
|
|
337
|
+
throw new Error('missing provider');
|
|
338
|
+
}
|
|
339
|
+
catch {
|
|
340
|
+
return { action: 'reject', reason: 'unknown-model' };
|
|
341
|
+
}
|
|
342
|
+
}
|
|
333
343
|
if (seal && send.kind === 'follow-up') {
|
|
334
344
|
const gen = Math.max(0, ...wf.journal.entries().filter(e => ['call', 'generation'].includes(e.type) && e.key === entry.key).map(e => Number(e.gen))) + 1;
|
|
335
|
-
|
|
345
|
+
// A follow-up's model replaces the continued session's for this generation and those continuing it.
|
|
346
|
+
const spec = send.model !== undefined ? { ...entry.spec, model: send.model } : entry.spec;
|
|
347
|
+
if (send.model !== undefined)
|
|
348
|
+
await this.note(req.rid, send.model, 'next-generation');
|
|
349
|
+
const opened = await wf.journal.append('generation', { rid: req.rid, key: entry.key, gen, from, spec, revision: wf.revision, opening: { rid: req.rid, kind: send.kind, message: send.message ?? '' }, ...(send.model !== undefined ? { model: send.model } : {}) });
|
|
336
350
|
this.dispatchGeneration(wf, opened);
|
|
337
351
|
return { action: 'apply' };
|
|
338
352
|
}
|
|
353
|
+
// A follow-up naming a model, queued on unfinished work: the executor records its model request with the message.
|
|
339
354
|
return this.executor.forward(req, this.context(wf, entry));
|
|
340
355
|
}
|
|
341
356
|
else if (req.kind === 'stop') {
|
|
@@ -504,6 +519,11 @@ export class Engine {
|
|
|
504
519
|
this.background(async () => { throw error; });
|
|
505
520
|
});
|
|
506
521
|
}
|
|
522
|
+
/** The reply to a send that names a model says which model and when it applies (orchestrator ledger `send-note`). */
|
|
523
|
+
async note(rid, model, effect) {
|
|
524
|
+
if (!this.ledgers.orch.entries().some(e => e.type === 'send-note' && e.rid === rid))
|
|
525
|
+
await this.ledgers.orch.append('send-note', { rid, model, effect });
|
|
526
|
+
}
|
|
507
527
|
ticket(st, entry) {
|
|
508
528
|
const spec = entry.spec, agent = st.wf.pins.agents.find(a => a.name === spec.agent);
|
|
509
529
|
if (!agent)
|
|
@@ -511,7 +531,8 @@ export class Engine {
|
|
|
511
531
|
return { wid: st.wf.wid, widRev: `${st.wf.wid}@${st.wf.revision}`, key: entry.key, gen: entry.gen,
|
|
512
532
|
callId: `${st.wf.wid}@${st.wf.revision}/${entry.key}@${entry.gen}`, spec, agent, workflowBudget: st.wf.pins.usageBudget, cwd: resolve(st.wf.cwd, spec.cwd ?? '.'), journal: st.wf.journal,
|
|
513
533
|
...(st.wf.pins.origin !== undefined ? { originSession: join(pinnedDir(this.ledgers.home, st.wf.wid), ...(st.wf.revision === 1 ? [] : [`r${st.wf.revision}`]), 'origin.jsonl') } : {}),
|
|
514
|
-
...(entry.type === 'generation' ? { continueFrom: entry.from, opening: entry.opening } : {})
|
|
534
|
+
...(entry.type === 'generation' ? { continueFrom: entry.from, opening: entry.opening } : {}),
|
|
535
|
+
...(entry.type === 'generation' && typeof entry.model === 'string' ? { model: entry.model } : {}) };
|
|
515
536
|
}
|
|
516
537
|
sealed(st, entry) {
|
|
517
538
|
if (entry.type === 'refused')
|
|
@@ -22,7 +22,8 @@ import { buildCallResult } from "../../compat/result.js";
|
|
|
22
22
|
import createEffects from "./effects/index.js";
|
|
23
23
|
import { continueSession } from "./generation.js";
|
|
24
24
|
import { hibernation, openQuestion } from "./hibernate.js";
|
|
25
|
-
import {
|
|
25
|
+
import { foldExhaustion } from "../providers.js";
|
|
26
|
+
import { evidence, fatalProviderError, quotaExhausted, refusedByProvider, forgetSession, readSessionState, receiptId, sessionModel } from "./session.js";
|
|
26
27
|
import { activeTotal } from "./time.js";
|
|
27
28
|
import { observeExecution } from "./observe.js";
|
|
28
29
|
import { availableMemory } from "./memory.js";
|
|
@@ -41,6 +42,39 @@ const entriesFor = (journal, call) => journal.entries().filter(e => e.call === c
|
|
|
41
42
|
const current = (journal, call) => entriesFor(journal, call).findLast(e => e.type === JT.exec)?.exec;
|
|
42
43
|
const sealed = (journal, call) => entriesFor(journal, call).find(e => e.type === JT.sealed)?.result;
|
|
43
44
|
const has = (journal, type, exec) => journal.entries().some(e => e.type === type && e.exec === exec);
|
|
45
|
+
/** The rid of the model request a follow-up naming a model makes (P12): derived, so withdrawing or replacing the
|
|
46
|
+
* follow-up withdraws its model request too. */
|
|
47
|
+
export const modelRid = (rid) => contentHash([rid, "model"]);
|
|
48
|
+
/** The model a call was asked to use and has not used yet: a follow-up's `model`, then each model send in order. One the
|
|
49
|
+
* child rejected or that was withdrawn does not count; one already used does not either — the child applied it, or an
|
|
50
|
+
* execution of the call answered with it — so the session's model and the pool's fallback rule again after that.
|
|
51
|
+
* Its next execution launches on it (P12, P37). */
|
|
52
|
+
export function requestedModel(journal, call, followUp) {
|
|
53
|
+
const all = journal.entries(), execs = new Set(all.filter(e => e.type === JT.exec && e.call === call).map(e => String(e.exec)));
|
|
54
|
+
const same = (a, b) => b?.provider === a.provider && b?.id === a.id;
|
|
55
|
+
const usedAfter = (m, index) => all.some((e, i) => i > index && (e.type === "selected" || e.type === "model-used") && execs.has(String(e.exec)) && same(m, e.model));
|
|
56
|
+
let wanted;
|
|
57
|
+
if (followUp) {
|
|
58
|
+
const m = parseModel(followUp);
|
|
59
|
+
if (!usedAfter(m, -1))
|
|
60
|
+
wanted = m;
|
|
61
|
+
}
|
|
62
|
+
for (const [index, e] of all.entries()) {
|
|
63
|
+
if (e.type !== "forward" || e.dest !== call || e.envelope?.kind !== "model")
|
|
64
|
+
continue;
|
|
65
|
+
const delivered = all.find(r => r.type === "forward-delivered" && r.call === call && r.rid2 === e.rid2);
|
|
66
|
+
if (delivered) {
|
|
67
|
+
wanted = undefined;
|
|
68
|
+
continue;
|
|
69
|
+
} // applied by the child (now the session's model) or refused by it
|
|
70
|
+
if (all.some(r => r.type === "forward" && r.dest === call && r.envelope.kind === "withdraw" && r.envelope.body.rids?.includes(String(e.rid2))))
|
|
71
|
+
continue;
|
|
72
|
+
const body = e.envelope.body;
|
|
73
|
+
const m = { provider: body.provider, id: body.model, ...(body.thinking ? { thinking: body.thinking } : {}) };
|
|
74
|
+
wanted = usedAfter(m, index) ? undefined : m;
|
|
75
|
+
}
|
|
76
|
+
return wanted;
|
|
77
|
+
}
|
|
44
78
|
function address(call) {
|
|
45
79
|
const match = /^(.*)@(\d+)\/(.*)@(\d+)$/.exec(call);
|
|
46
80
|
if (!match)
|
|
@@ -82,7 +116,7 @@ export default function createExecutor(ledgers, options = {}) {
|
|
|
82
116
|
const wake = () => { for (const fn of waiters)
|
|
83
117
|
fn(); waiters.clear(); };
|
|
84
118
|
// F3: fold only orchestrator entries appended since the last fold (holdings, observed switches, K7 skips).
|
|
85
|
-
const ledger = { seen: 0, held: new Map(), observed: new Set(), skips: new Map() };
|
|
119
|
+
const ledger = { seen: 0, held: new Map(), observed: new Set(), skips: new Map(), exhausted: new Map() };
|
|
86
120
|
const folded = () => {
|
|
87
121
|
const entries = orch.entries();
|
|
88
122
|
for (; ledger.seen < entries.length; ledger.seen++) {
|
|
@@ -95,11 +129,17 @@ export default function createExecutor(ledgers, options = {}) {
|
|
|
95
129
|
ledger.observed.add(`${e.exec}\n${e.rid}`);
|
|
96
130
|
else if (e.type === "skip")
|
|
97
131
|
ledger.skips.set(`${e.pool}\n${e.model}`, Math.max(Number(e.until), ledger.skips.get(`${e.pool}\n${e.model}`) ?? 0));
|
|
132
|
+
foldExhaustion(ledger.exhausted, e);
|
|
98
133
|
}
|
|
99
134
|
return ledger;
|
|
100
135
|
};
|
|
101
136
|
const holdings = () => [...folded().held.values()];
|
|
102
137
|
const skipped = (pool, model) => (folded().skips.get(`${pool}\n${model.provider}/${model.id}`) ?? 0) > Date.now();
|
|
138
|
+
/** A provider whose usage window is used up admits no call until its next try, and then one probe at a time. */
|
|
139
|
+
const unavailable = (provider) => {
|
|
140
|
+
const x = provider ? folded().exhausted.get(provider) : undefined;
|
|
141
|
+
return !!x && (Date.now() < x.nextTry || x.probe !== undefined);
|
|
142
|
+
};
|
|
103
143
|
async function release(exec) {
|
|
104
144
|
await serial(async () => { for (const h of holdings().filter(e => e.exec === exec))
|
|
105
145
|
await orch.append("release", { pool: h.pool, slot: h.slot, exec }); });
|
|
@@ -261,6 +301,11 @@ export default function createExecutor(ledgers, options = {}) {
|
|
|
261
301
|
}
|
|
262
302
|
}
|
|
263
303
|
/** P7, P27: Record forward-delivered once when a forward's child receipt is first observed; serial sections only. */
|
|
304
|
+
/** The reply to a model send says which model and when it applies (orchestrator ledger `send-note`, once per rid). */
|
|
305
|
+
async function note(rid, model, effect) {
|
|
306
|
+
if (!orch.entries().some(e => e.type === "send-note" && e.rid === rid))
|
|
307
|
+
await orch.append("send-note", { rid, model, effect });
|
|
308
|
+
}
|
|
264
309
|
async function forwardsDelivered(journal, call, entries) {
|
|
265
310
|
const all = journal.entries();
|
|
266
311
|
const open = all.filter(e => e.type === "forward" && e.dest === call &&
|
|
@@ -405,6 +450,9 @@ export default function createExecutor(ledgers, options = {}) {
|
|
|
405
450
|
if (!continuation && pool && skipped(pool, model))
|
|
406
451
|
continue;
|
|
407
452
|
const provider = model.provider;
|
|
453
|
+
if (unavailable(provider))
|
|
454
|
+
continue;
|
|
455
|
+
const probe = provider !== undefined && folded().exhausted.has(provider);
|
|
408
456
|
const holders = holdings().filter(e => e.pool === provider);
|
|
409
457
|
const limit = config.providers?.[provider ?? ""]?.slots ?? Infinity;
|
|
410
458
|
if (!capacity({ kind: "provider", holders: holders.length, capacity: limit }))
|
|
@@ -432,6 +480,10 @@ export default function createExecutor(ledgers, options = {}) {
|
|
|
432
480
|
while (holders.some(e => e.slot === slot))
|
|
433
481
|
slot++;
|
|
434
482
|
await orch.append("hold", { pool: provider, slot, exec });
|
|
483
|
+
// After its next try, the first call admitted to a used-up provider is its probe: its first request
|
|
484
|
+
// either goes through (the provider is available again) or is refused, which uses no quota.
|
|
485
|
+
if (probe)
|
|
486
|
+
await orch.append("provider-probe", { provider, exec });
|
|
435
487
|
}
|
|
436
488
|
await a.ticket.journal.append("selected", { exec, model, ...(pool ? { pool } : {}) });
|
|
437
489
|
return model;
|
|
@@ -460,7 +512,16 @@ export default function createExecutor(ledgers, options = {}) {
|
|
|
460
512
|
const pool = raw && config.pools?.[raw] ? raw : undefined;
|
|
461
513
|
const candidates = raw ? resolveModel(raw, config.pools) : [{ id: "" }];
|
|
462
514
|
const candidate = recorded && candidates.some(m => m.provider === recorded.provider && m.id === recorded.id);
|
|
463
|
-
|
|
515
|
+
// Leave the session's model for the pool's others when its pool skips it after losses, or its provider's usage
|
|
516
|
+
// window is used up; and at a new generation, go back to the pool's order of preference.
|
|
517
|
+
const skip = pool && candidate && (previous && ownSegment && skipped(pool, recorded) || unavailable(recorded.provider) || !ownSegment && !!t.continueFrom);
|
|
518
|
+
// A model the call was asked to use replaces the session's: launched with it, and holding its provider's slot.
|
|
519
|
+
const wanted = requestedModel(t.journal, t.callId, t.model);
|
|
520
|
+
// It outranks the pool's order at a new generation too, also when it names the model the session already has.
|
|
521
|
+
if (wanted)
|
|
522
|
+
return recorded && !freshFork && recorded.provider === wanted.provider && recorded.id === wanted.id
|
|
523
|
+
? { candidates: [recorded], continuation: true, pool: undefined }
|
|
524
|
+
: { candidates: [wanted], continuation: false, pool: undefined };
|
|
464
525
|
if (recorded && !freshFork && !skip)
|
|
465
526
|
return { candidates: [recorded], continuation: true, pool: candidate ? pool : undefined };
|
|
466
527
|
return { candidates, continuation: false, pool };
|
|
@@ -487,11 +548,26 @@ export default function createExecutor(ledgers, options = {}) {
|
|
|
487
548
|
}
|
|
488
549
|
});
|
|
489
550
|
}
|
|
490
|
-
|
|
491
|
-
|
|
551
|
+
/** The model each execution last answered with (`selected`, then `model-used`), cached per execution. */
|
|
552
|
+
const inUse = new Map();
|
|
553
|
+
async function switched(exec, journal, event) {
|
|
554
|
+
const message = event.message, provider = message?.provider;
|
|
492
555
|
if (!provider)
|
|
493
556
|
return;
|
|
494
557
|
await serial(async () => {
|
|
558
|
+
// Evidence of the model in use: the provider and model of each assistant message, recorded when it changes.
|
|
559
|
+
if (message.role === "assistant" && message.model) {
|
|
560
|
+
const name = `${provider}/${message.model}`;
|
|
561
|
+
if (!inUse.has(exec)) {
|
|
562
|
+
const last = journal.entries().findLast(e => (e.type === "selected" || e.type === "model-used") && e.exec === exec)?.model;
|
|
563
|
+
if (last)
|
|
564
|
+
inUse.set(exec, `${last.provider}/${last.id}`);
|
|
565
|
+
}
|
|
566
|
+
if (inUse.get(exec) !== name) {
|
|
567
|
+
await journal.append("model-used", { exec, model: { provider, id: message.model } });
|
|
568
|
+
inUse.set(exec, name);
|
|
569
|
+
}
|
|
570
|
+
}
|
|
495
571
|
const target = holdings().find(h => h.exec === exec && h.pool === provider && h.reserved);
|
|
496
572
|
if (!target)
|
|
497
573
|
return;
|
|
@@ -503,6 +579,32 @@ export default function createExecutor(ledgers, options = {}) {
|
|
|
503
579
|
});
|
|
504
580
|
wake();
|
|
505
581
|
}
|
|
582
|
+
/** The model an execution last answered with, or was launched on. */
|
|
583
|
+
function modelOf(journal, exec) {
|
|
584
|
+
return journal.entries().findLast(e => (e.type === "selected" || e.type === "model-used") && e.exec === exec)?.model;
|
|
585
|
+
}
|
|
586
|
+
/** Record a used-up provider once per window: again only when its probe (or any call after the next try) is refused. */
|
|
587
|
+
async function recordExhausted(provider, exec, error) {
|
|
588
|
+
const x = folded().exhausted.get(provider), now = Date.now();
|
|
589
|
+
// While a probe runs, its outcome alone decides: a late refusal of an execution admitted earlier changes nothing.
|
|
590
|
+
if (x && (x.probe ? x.probe !== exec : now < x.nextTry))
|
|
591
|
+
return;
|
|
592
|
+
if (orch.entries().some(e => e.type === "provider-exhausted" && e.exec === exec))
|
|
593
|
+
return;
|
|
594
|
+
await orch.append("provider-exhausted", { provider, exec, since: x?.since ?? now, nextTry: now + (config.k?.probeMs ?? 900_000), error: error.slice(0, 300) });
|
|
595
|
+
}
|
|
596
|
+
/** An answer from a used-up provider, requested after it was found used up: available again. */
|
|
597
|
+
async function answered(exec, event) {
|
|
598
|
+
const message = event.message, provider = message?.provider;
|
|
599
|
+
if (message?.role !== "assistant" || !provider || message.stopReason === "error")
|
|
600
|
+
return;
|
|
601
|
+
await serial(async () => {
|
|
602
|
+
const x = folded().exhausted.get(provider);
|
|
603
|
+
if (x && (x.probe === exec || Number(message.timestamp) > x.since))
|
|
604
|
+
await orch.append("provider-available", { provider, exec });
|
|
605
|
+
});
|
|
606
|
+
wake();
|
|
607
|
+
}
|
|
506
608
|
function pendingSwitch(exec) {
|
|
507
609
|
return holdings().find(h => h.exec === exec && h.reserved && !folded().observed.has(`${exec}\n${h.rid}`));
|
|
508
610
|
}
|
|
@@ -579,11 +681,20 @@ export default function createExecutor(ledgers, options = {}) {
|
|
|
579
681
|
return finish(journal, t.callId, exec, makeResult("unknown", "", `Unknown tool outcomes: ${dangling.join(", ")}`));
|
|
580
682
|
if (has(journal, "settled", exec) && !ev.text && ev.error && fatalProviderError(ev.error))
|
|
581
683
|
return finish(journal, t.callId, exec, makeResult("failed", "", `Provider error: ${ev.error}`));
|
|
582
|
-
|
|
583
|
-
|
|
584
|
-
|
|
585
|
-
|
|
586
|
-
|
|
684
|
+
// A refusal of the content is deterministic: the same request is refused again, so it is reported, not retried.
|
|
685
|
+
if (has(journal, "settled", exec) && !ev.text && ev.error && refusedByProvider(ev.error))
|
|
686
|
+
return finish(journal, t.callId, exec, makeResult("failed", "", `Refused by the provider (not retried): ${ev.error.slice(0, 500)}`));
|
|
687
|
+
// A used-up usage window is no loss: the provider is avoided until a probe finds it accepting requests again,
|
|
688
|
+
// and the call goes on with the pool's next model, or waits for that provider.
|
|
689
|
+
const exhausted = has(journal, "settled", exec) && !ev.text && ev.error && quotaExhausted(ev.error) ? modelOf(journal, exec)?.provider : undefined;
|
|
690
|
+
if (exhausted)
|
|
691
|
+
await serial(() => recordExhausted(exhausted, exec, ev.error));
|
|
692
|
+
else
|
|
693
|
+
await serial(async () => {
|
|
694
|
+
if (!has(journal, "loss", exec))
|
|
695
|
+
await journal.append("loss", { exec });
|
|
696
|
+
await skipLostCandidate(journal, orch, exec);
|
|
697
|
+
});
|
|
587
698
|
const losses = journal.entries().filter(e => e.type === "loss" && String(e.exec).startsWith(`${t.callId}#`)).length;
|
|
588
699
|
if (losses >= (config.k?.lossBound ?? 5))
|
|
589
700
|
return finish(journal, t.callId, exec, makeResult("failed", "", `lost ×${losses}${ev.error ? `; last error: ${ev.error.slice(0, 300)}` : ""}`));
|
|
@@ -679,7 +790,7 @@ export default function createExecutor(ledgers, options = {}) {
|
|
|
679
790
|
}
|
|
680
791
|
});
|
|
681
792
|
}, recordUsage: values => recordUsage(t, values),
|
|
682
|
-
switched: event => switched(exec, event), pendingSwitch: () => pendingSwitch(exec),
|
|
793
|
+
switched: event => switched(exec, journal, event), answered: event => answered(exec, event), pendingSwitch: () => pendingSwitch(exec),
|
|
683
794
|
});
|
|
684
795
|
}
|
|
685
796
|
finally {
|
|
@@ -716,6 +827,9 @@ export default function createExecutor(ledgers, options = {}) {
|
|
|
716
827
|
finally {
|
|
717
828
|
active.delete(ticket.callId);
|
|
718
829
|
collected.delete(ticket.callId);
|
|
830
|
+
for (const e of inUse.keys())
|
|
831
|
+
if (callOf(e) === ticket.callId)
|
|
832
|
+
inUse.delete(e);
|
|
719
833
|
forgetSession(callSession(home, ticket.wid, ticket.key, ticket.gen));
|
|
720
834
|
wake();
|
|
721
835
|
}
|
|
@@ -763,28 +877,70 @@ export default function createExecutor(ledgers, options = {}) {
|
|
|
763
877
|
return { action: "apply" };
|
|
764
878
|
}
|
|
765
879
|
}
|
|
880
|
+
/** P12: a model request to this call, recorded with the rid given; a reject has no effect. */
|
|
881
|
+
const requestModel = async (rid, model, hash, cond) => {
|
|
882
|
+
let body;
|
|
883
|
+
try {
|
|
884
|
+
const m = parseModel(model);
|
|
885
|
+
if (!m.provider)
|
|
886
|
+
throw new Error("Missing provider");
|
|
887
|
+
body = { provider: m.provider, model: m.id, ...(m.thinking ? { thinking: m.thinking } : {}) };
|
|
888
|
+
}
|
|
889
|
+
catch {
|
|
890
|
+
return { action: "reject", reason: "unknown-model" };
|
|
891
|
+
}
|
|
892
|
+
const envelope = { to: dest, kind: "model", body, ...(cond && Object.keys(cond).length ? { cond } : {}) };
|
|
893
|
+
const rid2 = forwardRid(rid, ctx.widRev, ctx.key, hash);
|
|
894
|
+
const exec = current(ctx.journal, dest), provider = body.provider;
|
|
895
|
+
// P28: with no live execution (not started yet, between executions, hibernated while asking) the model is
|
|
896
|
+
// recorded and the next execution launches on it (`requestedModel`); its slot is acquired then, as for any launch.
|
|
897
|
+
// Launching (`selected`, not `tracked` yet): the child may start on the old model; ask again in a moment.
|
|
898
|
+
const idle = !exec || has(ctx.journal, JT.fenced, exec) || !has(ctx.journal, "selected", exec);
|
|
899
|
+
if (!idle && !has(ctx.journal, "tracked", exec))
|
|
900
|
+
return { action: "reject", reason: "call-starting" };
|
|
901
|
+
if (!idle && pendingSwitch(exec))
|
|
902
|
+
return { action: "reject", reason: "switch-pending" };
|
|
903
|
+
if (!idle) {
|
|
904
|
+
const held = holdings().filter(h => h.exec === exec);
|
|
905
|
+
if (!held.some(h => h.pool === provider)) {
|
|
906
|
+
const target = holdings().filter(h => h.pool === provider);
|
|
907
|
+
if (!capacity({ kind: "provider", holders: target.length, capacity: config.providers?.[provider]?.slots ?? Infinity }))
|
|
908
|
+
return { action: "reject", reason: "provider-full" };
|
|
909
|
+
let slot = 0;
|
|
910
|
+
while (target.some(h => h.slot === slot))
|
|
911
|
+
slot++;
|
|
912
|
+
await orch.append("hold", { pool: provider, slot, exec, reserved: true, rid });
|
|
913
|
+
}
|
|
914
|
+
}
|
|
915
|
+
await note(req.rid, model, idle ? "next-execution" : "next-request");
|
|
916
|
+
const entry = await ctx.journal.append("forward", { rid, rid2, dest, hash, envelope });
|
|
917
|
+
await replayForward(entry);
|
|
918
|
+
if (idle) {
|
|
919
|
+
active.get(dest)?.wake();
|
|
920
|
+
wake();
|
|
921
|
+
}
|
|
922
|
+
return undefined;
|
|
923
|
+
};
|
|
766
924
|
let kind, body;
|
|
925
|
+
if (req.kind === "send" && req.body.kind === "follow-up" && req.body.model !== undefined) {
|
|
926
|
+
// A follow-up naming a model, queued on unfinished work: the model request and the message are recorded in one
|
|
927
|
+
// section, so a seal cannot come between them (both or neither). A replay finds the model request recorded.
|
|
928
|
+
const mrid = modelRid(req.rid);
|
|
929
|
+
if (!ctx.journal.entries().some(e => e.type === "forward" && e.rid === mrid && e.dest === dest)) {
|
|
930
|
+
const refused = await requestModel(mrid, req.body.model, contentHash([hash, "model"]));
|
|
931
|
+
if (refused)
|
|
932
|
+
return { action: "reject", reason: `model: ${refused.reason}` };
|
|
933
|
+
}
|
|
934
|
+
}
|
|
767
935
|
if (req.kind === "withdraw") {
|
|
768
936
|
kind = "withdraw";
|
|
769
|
-
const targets = req.body.rids;
|
|
937
|
+
const targets = req.body.rids.flatMap(rid => [rid, modelRid(rid)]);
|
|
770
938
|
body = { rids: ctx.journal.entries().filter(e => e.type === "forward" && e.dest === dest && targets.includes(String(e.rid))).map(e => String(e.rid2)) };
|
|
771
939
|
}
|
|
772
940
|
else if (req.kind === "send") {
|
|
773
941
|
const send = req.body;
|
|
774
942
|
kind = send.kind;
|
|
775
|
-
|
|
776
|
-
try {
|
|
777
|
-
const m = parseModel(send.model ?? "");
|
|
778
|
-
if (!m.provider)
|
|
779
|
-
throw new Error("Missing provider");
|
|
780
|
-
body = { provider: m.provider, model: m.id, ...(m.thinking ? { thinking: m.thinking } : {}) };
|
|
781
|
-
}
|
|
782
|
-
catch {
|
|
783
|
-
return { action: "reject", reason: "unknown-model" };
|
|
784
|
-
}
|
|
785
|
-
}
|
|
786
|
-
else
|
|
787
|
-
body = { message: send.message ?? "" };
|
|
943
|
+
body = kind === "model" ? undefined : { message: send.message ?? "" };
|
|
788
944
|
}
|
|
789
945
|
else
|
|
790
946
|
return { action: "reject", reason: "unsupported" };
|
|
@@ -797,30 +953,15 @@ export default function createExecutor(ledgers, options = {}) {
|
|
|
797
953
|
else
|
|
798
954
|
delete cond.after;
|
|
799
955
|
}
|
|
956
|
+
if (kind === "model")
|
|
957
|
+
return (await requestModel(req.rid, req.body.model ?? "", hash, cond)) ?? { action: "apply" };
|
|
800
958
|
const envelope = { to: dest, kind, body, ...(Object.keys(cond).length ? { cond } : {}) };
|
|
801
959
|
const rid2 = forwardRid(req.rid, ctx.widRev, ctx.key, hash);
|
|
802
|
-
if (kind === "model") {
|
|
803
|
-
const exec = current(ctx.journal, dest), provider = body.provider;
|
|
804
|
-
if (!exec || has(ctx.journal, JT.fenced, exec) || !has(ctx.journal, "tracked", exec))
|
|
805
|
-
return { action: "reject", reason: "call-not-running" };
|
|
806
|
-
if (pendingSwitch(exec))
|
|
807
|
-
return { action: "reject", reason: "switch-pending" };
|
|
808
|
-
const held = holdings().filter(h => h.exec === exec);
|
|
809
|
-
if (!held.some(h => h.pool === provider)) {
|
|
810
|
-
const target = holdings().filter(h => h.pool === provider);
|
|
811
|
-
if (!capacity({ kind: "provider", holders: target.length, capacity: config.providers?.[provider]?.slots ?? Infinity }))
|
|
812
|
-
return { action: "reject", reason: "provider-full" };
|
|
813
|
-
let slot = 0;
|
|
814
|
-
while (target.some(h => h.slot === slot))
|
|
815
|
-
slot++;
|
|
816
|
-
await orch.append("hold", { pool: provider, slot, exec, reserved: true, rid: req.rid });
|
|
817
|
-
}
|
|
818
|
-
}
|
|
819
960
|
const entry = await ctx.journal.append("forward", { rid: req.rid, rid2, dest, hash, envelope });
|
|
820
961
|
await replayForward(entry);
|
|
821
962
|
if (kind === "withdraw") {
|
|
822
963
|
const exec = current(ctx.journal, dest), reservation = exec && pendingSwitch(exec);
|
|
823
|
-
if (reservation && req.body.rids.
|
|
964
|
+
if (reservation && req.body.rids.some(rid => rid === reservation.rid || modelRid(rid) === reservation.rid))
|
|
824
965
|
active.get(dest)?.wake();
|
|
825
966
|
}
|
|
826
967
|
return { action: "apply" };
|
|
@@ -110,8 +110,10 @@ export async function observeExecution(d) {
|
|
|
110
110
|
await serial(() => t.journal.append("observation", { exec, event: slim }));
|
|
111
111
|
if (event.type === "message_start")
|
|
112
112
|
await d.switched(event);
|
|
113
|
-
if (event.type === "message_end")
|
|
113
|
+
if (event.type === "message_end") {
|
|
114
114
|
await d.recordUsage([{ id: String(slim.id), usage: slim.usage }]);
|
|
115
|
+
await d.answered?.(event);
|
|
116
|
+
}
|
|
115
117
|
}
|
|
116
118
|
await limits();
|
|
117
119
|
await stall();
|
|
@@ -107,12 +107,43 @@ export function evidence(entries, exec) {
|
|
|
107
107
|
const error = last?.stopReason === "error" ? last.errorMessage : undefined;
|
|
108
108
|
return { report, budget, text, error, dangling: [...tools].map(([id, name]) => `${name} (${id})`), usage };
|
|
109
109
|
}
|
|
110
|
-
/** Only explicit
|
|
110
|
+
/** Only explicit payment failures are terminal; rate limits, overload and transport errors still retry, and a used-up
|
|
111
|
+
* usage window (`quotaExhausted`) waits for the provider or moves to another one. */
|
|
111
112
|
export function fatalProviderError(text) {
|
|
112
|
-
return /\b402\b|insufficient[_ ]?(quota|balance|funds)|
|
|
113
|
+
return /\b402\b|insufficient[_ ]?(quota|balance|funds)|billing|credit balance|余额/i.test(text);
|
|
113
114
|
}
|
|
114
|
-
/**
|
|
115
|
+
/** A provider's usage window is used up: its requests are refused (and not counted) until the window resets, hours
|
|
116
|
+
* later. Seen as a gateway's `503 No available accounts` once pi's own retries are spent, or a usage-limit message.
|
|
117
|
+
* The provider is then avoided until a probe finds it accepting requests again. */
|
|
118
|
+
export function quotaExhausted(text) {
|
|
119
|
+
if (fatalProviderError(text))
|
|
120
|
+
return false;
|
|
121
|
+
if (/no available accounts?/i.test(text))
|
|
122
|
+
return true;
|
|
123
|
+
// A request rate limit clears in seconds ("rate limit exceeded; resets in 1 second", "quota exceeded for requests
|
|
124
|
+
// per minute"): pi's retries and the lost-execution path handle it; it must not take the provider out for minutes.
|
|
125
|
+
if (/rate.?limit|too many requests|request limit|per (second|minute)|\b[RT]PM\b|resets? in \d+ ?(ms|s|secs?|seconds?|minutes?)\b/i.test(text))
|
|
126
|
+
return false;
|
|
127
|
+
return /usage limit|quota (exceeded|exhausted)|exceeded your (current )?(usage|quota)|limit (reached|exceeded)[^.]*resets?\b|额度/i.test(text);
|
|
128
|
+
}
|
|
129
|
+
/** A refusal of the request's content (terms of service, usage or content policy): the same request is refused again,
|
|
130
|
+
* on this provider and usually on another, so it is reported at once instead of retried as a lost execution. */
|
|
131
|
+
export function refusedByProvider(text) {
|
|
132
|
+
// A content filter that is down ("temporarily unavailable, please retry") is a transient failure, not a refusal.
|
|
133
|
+
if (/temporar|unavailable|try again|retry|timed? ?out|overloaded/i.test(text))
|
|
134
|
+
return false;
|
|
135
|
+
return /terms of service|usage polic(y|ies)|acceptable use|content[_ ]?(policy|filter|management policy)|safety (system|filter)|flagged as (unsafe|harmful)/i.test(text);
|
|
136
|
+
}
|
|
137
|
+
/** P13, C8: Restore the effective provider as pi does: from the last model change or assistant message. */
|
|
115
138
|
export function sessionModel(entries) {
|
|
116
|
-
|
|
117
|
-
|
|
139
|
+
// pi restores the model of the last model change or assistant message: a relaunch with `--model` on an existing
|
|
140
|
+
// session records no model change, so only the answer tells which model the session went on with.
|
|
141
|
+
const last = entries.findLast(e => e.type === "model_change" && e.provider && e.modelId
|
|
142
|
+
|| e.type === "message" && e.message?.role === "assistant" && !!e.message.provider && !!e.message.model);
|
|
143
|
+
if (!last)
|
|
144
|
+
return undefined;
|
|
145
|
+
if (last.type === "model_change")
|
|
146
|
+
return { provider: last.provider, id: last.modelId };
|
|
147
|
+
const m = last.message;
|
|
148
|
+
return { provider: m.provider, id: m.model };
|
|
118
149
|
}
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
/** Apply one orchestrator ledger entry to the map of used-up providers. */
|
|
2
|
+
export function foldExhaustion(exhausted, e) {
|
|
3
|
+
const provider = String(e.provider ?? e.pool ?? "");
|
|
4
|
+
if (e.type === "provider-exhausted") {
|
|
5
|
+
// A refusal of the probe ends it; any other execution's (the executor records none while a probe runs) keeps it.
|
|
6
|
+
const probe = exhausted.get(provider)?.probe;
|
|
7
|
+
exhausted.set(provider, { since: Number(e.since), nextTry: Number(e.nextTry), error: String(e.error ?? ""), ...(probe && probe !== e.exec ? { probe } : {}) });
|
|
8
|
+
}
|
|
9
|
+
else if (e.type === "provider-available")
|
|
10
|
+
exhausted.delete(provider);
|
|
11
|
+
else if (e.type === "provider-probe") {
|
|
12
|
+
const x = exhausted.get(provider);
|
|
13
|
+
if (x)
|
|
14
|
+
x.probe = String(e.exec);
|
|
15
|
+
}
|
|
16
|
+
else if (e.type === "release") {
|
|
17
|
+
const x = exhausted.get(provider);
|
|
18
|
+
if (x && x.probe === e.exec)
|
|
19
|
+
delete x.probe;
|
|
20
|
+
}
|
|
21
|
+
}
|
|
@@ -6,6 +6,7 @@ import { compileFanout } from "../compat/fanout.js";
|
|
|
6
6
|
import { readJournalSnapshot } from "../kernel/journal.js";
|
|
7
7
|
import { journalPath, orchLedger, pinnedDir, workflowDir } from "../paths.js";
|
|
8
8
|
import { JT } from "../types.js";
|
|
9
|
+
import { foldExhaustion } from "./providers.js";
|
|
9
10
|
/** A workflow has live work: it runs, or follow-ups opened on it after it finished have not ended yet. */
|
|
10
11
|
export function isLive(wf) {
|
|
11
12
|
return wf.status === "running" || (wf.followUps ?? 0) > 0;
|
|
@@ -152,6 +153,12 @@ function snapshotReducer(wid, entries) {
|
|
|
152
153
|
if (call && call.phase === "queued")
|
|
153
154
|
call.phase = "running";
|
|
154
155
|
}
|
|
156
|
+
else if (e.type === "model-used") {
|
|
157
|
+
// Evidence of a switch: the execution answered with another model than it launched with.
|
|
158
|
+
const call = byExec.get(String(e.exec)), m = e.model;
|
|
159
|
+
if (call && m)
|
|
160
|
+
call.model = m.provider ? `${m.provider}/${m.id}` : m.id;
|
|
161
|
+
}
|
|
155
162
|
else if (e.type === "observation") {
|
|
156
163
|
// Only what the agent did counts as activity; tracker scans and time checkpoints are bookkeeping.
|
|
157
164
|
const call = byExec.get(String(e.exec));
|
|
@@ -195,6 +202,13 @@ function snapshotReducer(wid, entries) {
|
|
|
195
202
|
const list = sends.get(String(e.dest)) ?? [];
|
|
196
203
|
list.push(send);
|
|
197
204
|
sends.set(String(e.dest), list);
|
|
205
|
+
// A withdrawn model request no longer stands (the executor ignores it for the next launch as well).
|
|
206
|
+
if (envelope?.kind === "withdraw")
|
|
207
|
+
for (const rid2 of envelope.body?.rids ?? []) {
|
|
208
|
+
const target = byRid2.get(`${e.dest}\n${rid2}`);
|
|
209
|
+
if (target?.kind === "model" && target.state === "pending")
|
|
210
|
+
target.reason = "withdrawn";
|
|
211
|
+
}
|
|
198
212
|
byRid2.set(`${e.dest}\n${e.rid2}`, send);
|
|
199
213
|
byRid2.set(String(e.rid2), send);
|
|
200
214
|
}
|
|
@@ -239,9 +253,11 @@ function snapshotReducer(wid, entries) {
|
|
|
239
253
|
const pending = forwarded.filter(pendingMessage).length;
|
|
240
254
|
if (pending)
|
|
241
255
|
c.pending = pending;
|
|
242
|
-
const
|
|
243
|
-
if (
|
|
244
|
-
c.
|
|
256
|
+
const last = c.phase !== "sealed" && !retired.has(c.callId) ? forwarded.findLast(s => s.kind === "model" && s.reason !== "withdrawn") : undefined;
|
|
257
|
+
if (last?.reason !== undefined)
|
|
258
|
+
c.switchFailed = `${last.model} (${last.reason})`;
|
|
259
|
+
else if (last && last.state !== "retired" && last.model !== c.model?.replace(/:(off|minimal|low|medium|high|xhigh|max)$/, ""))
|
|
260
|
+
c.switching = last.model;
|
|
245
261
|
}
|
|
246
262
|
}
|
|
247
263
|
const after = done ? list.filter(c => generations.has(c.callId) && c.phase !== "sealed" && !retired.has(c.callId)) : [];
|
|
@@ -404,7 +420,7 @@ export function compactWorkflow(wf) {
|
|
|
404
420
|
calls: wf.calls.map(c => {
|
|
405
421
|
const r = c.result, last = r?.output?.split("\n").map(l => l.trim()).filter(Boolean).at(-1);
|
|
406
422
|
return { key: c.key, gen: c.gen, callId: c.callId, phase: c.phase, ...(r ? { status: r.status, ok: r.ok } : {}),
|
|
407
|
-
...(c.model ? { model: c.model } : {}), ...(c.tools ? { tools: c.tools } : {}), ...(c.pending ? { pending: c.pending } : {}), ...(c.switching ? { switching: c.switching } : {}), ...(nonzero(c.usage) ? { usage: c.usage } : {}),
|
|
423
|
+
...(c.model ? { model: c.model } : {}), ...(c.tools ? { tools: c.tools } : {}), ...(c.pending ? { pending: c.pending } : {}), ...(c.switching ? { switching: c.switching } : {}), ...(c.switchFailed ? { switchFailed: c.switchFailed } : {}), ...(nonzero(c.usage) ? { usage: c.usage } : {}),
|
|
408
424
|
...(last ? { lastLine: clip(last, 200) } : {}), ...(r?.error ? { error: clip(r.error, 300) } : {}), ...(c.hibernated ? { hibernated: true } : {}) };
|
|
409
425
|
}),
|
|
410
426
|
attention: wf.attention.map(a => ({ id: a.id, rev: a.rev, kind: a.kind, text: clip(a.text, 300), ...(a.call ? { call: a.call } : {}), ...(a.qid ? { qid: a.qid } : {}) })),
|
|
@@ -454,7 +470,7 @@ const latestCalls = (wf) => [...new Map(wf.calls.map(c => [c.key, c])).values()]
|
|
|
454
470
|
/** Provider slots from the orchestrator ledger: holders per provider (hold/release{pool,slot,exec}) and the limits of
|
|
455
471
|
* the settings in effect (the latest config{hash,config}); config-rejected after it is reported too. */
|
|
456
472
|
export function slotsView(home, now = Date.now()) {
|
|
457
|
-
const held = new Map();
|
|
473
|
+
const held = new Map(), used = new Map();
|
|
458
474
|
let config, rejected;
|
|
459
475
|
for (const e of readJournalSnapshot(orchLedger(home))) {
|
|
460
476
|
if (e.type === "hold")
|
|
@@ -467,7 +483,10 @@ export function slotsView(home, now = Date.now()) {
|
|
|
467
483
|
}
|
|
468
484
|
else if (e.type === "config-rejected")
|
|
469
485
|
rejected = e;
|
|
486
|
+
foldExhaustion(used, e);
|
|
470
487
|
}
|
|
488
|
+
const exhausted = [...used].sort(([a], [b]) => a.localeCompare(b)).map(([p, x]) => `${p} exhausted since ${age(now - x.since)} ago (${clip(x.error, 80)}), ` +
|
|
489
|
+
(x.probe ? `probing with ${x.probe.split("#")[0]}` : x.nextTry > now ? `next try in ${age(x.nextTry - now)}` : "next call probes it"));
|
|
471
490
|
const limits = (config?.config?.providers) ?? {};
|
|
472
491
|
const holders = new Map();
|
|
473
492
|
for (const e of held.values())
|
|
@@ -476,6 +495,7 @@ export function slotsView(home, now = Date.now()) {
|
|
|
476
495
|
const names = [...new Set([...Object.keys(limits), ...holders.keys()])].sort();
|
|
477
496
|
const slots = names.map(p => { const n = holders.get(p) ?? 0, limit = limits[p]?.slots; return typeof limit === "number" ? `${p} ${n}/${limit}` : `${p} ${n} (no limit)`; });
|
|
478
497
|
return { ...(slots.length ? { slots } : {}), ...(config ? { config: `${String(config.hash)} since ${age(now - config.ts)} ago` } : {}),
|
|
498
|
+
...(exhausted.length ? { exhausted } : {}),
|
|
479
499
|
...(rejected ? { configRejected: `${clip(String(rejected.error), 200)} (${age(now - rejected.ts)} ago); ${config ? String(config.hash) : "the start settings"} stay in effect` } : {}) };
|
|
480
500
|
}
|
|
481
501
|
/** Tool status without a wid: what runs, what waits for an answer and what failed, with finished workflows one line each.
|
|
@@ -498,7 +518,8 @@ export function statusBrief(home, options = {}) {
|
|
|
498
518
|
...(live && c.phase !== "asking" && quiet !== undefined && quiet >= 60_000 ? { quiet: age(quiet) } : {}),
|
|
499
519
|
...(live && c.startedAt !== undefined ? { tokens: tokens(c.usage) } : {}),
|
|
500
520
|
...(c.result ? { status: c.result.status, ...(c.result.error ? { error: clip(c.result.error, 200) } : {}) } : {}),
|
|
501
|
-
...(c.hibernated ? { hibernated: true } : {})
|
|
521
|
+
...(c.hibernated ? { hibernated: true } : {}),
|
|
522
|
+
...(live && c.switching ? { switching: c.switching } : {}), ...(live && c.switchFailed ? { switchFailed: c.switchFailed } : {}) };
|
|
502
523
|
});
|
|
503
524
|
const asking = open.filter(a => a.kind === "question" && a.call).map(a => ({ to: `${w.wid}/${callKey(a.call)}`, ...(a.qid ? { qid: a.qid } : {}),
|
|
504
525
|
...(w.calls.some(c => c.callId === a.call && c.hibernated) ? { hibernated: true } : {}), question: clip(a.text, 300) }));
|
package/dist/ui/screen.js
CHANGED
|
@@ -565,7 +565,7 @@ export class SubagentScreen {
|
|
|
565
565
|
const facts = this.data.facts.get(c.callId), active = w.calls.filter(c => c.phase !== "sealed"), done = w.calls.length - active.length;
|
|
566
566
|
const tabs = size.width < 60 ? `${c.key} ${w.calls.indexOf(c) + 1}/${w.calls.length}` : `${[...active.map(c => c.key), ...(done ? [`${done} done`] : [])].join(" · ")} ← → switch`;
|
|
567
567
|
const tools = toolCount(facts?.tools), pending = pendingText(c.pending), rule = this.theme.fg("borderMuted", "─".repeat(size.width));
|
|
568
|
-
const switching = c.switching ? ` → ${this.name(c.switching)} (
|
|
568
|
+
const switching = c.switching ? ` → ${this.name(c.switching)} (requested)` : "";
|
|
569
569
|
const head = [tabs, `${label(c)} · ${this.name(facts?.model ?? c.model)} ▾${switching} · ${facts?.thinking ?? "off"} ▾${tools ? ` · ${tools}` : ""}${pending ? ` · ${pending}` : ""}`,
|
|
570
570
|
this.theme.fg("dim", this.spend(c, facts)), rule];
|
|
571
571
|
const asking = w.attention.some(a => a.kind === "question" && a.call === c.callId);
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pi-durable-subagents",
|
|
3
|
-
"version": "1.0.
|
|
3
|
+
"version": "1.0.10",
|
|
4
4
|
"description": "Subagents for pi that never lose work and never do it twice. Crash-safe workflows, automatic recovery, and a live view just like the main agent.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"license": "MIT",
|