pi-durable-subagents 1.0.6 → 1.0.7
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +14 -0
- package/README.md +5 -0
- package/dist/agent/main.js +1 -1
- package/dist/cli/main.js +5 -2
- package/dist/orchestrator/config.js +147 -0
- package/dist/orchestrator/executor/index.js +10 -4
- package/dist/orchestrator/executor/observe.js +5 -3
- package/dist/orchestrator/main.js +20 -12
- package/dist/orchestrator/snapshot.js +58 -5
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,19 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 1.0.7
|
|
4
|
+
|
|
5
|
+
- Changes to `config.json` (`defaultModel`, `pools`, `providers`, `memory`,
|
|
6
|
+
`k`) apply without restarting the orchestrator: it checks the file about
|
|
7
|
+
once a second and applies a valid change between slot admissions, so a call
|
|
8
|
+
still waiting for a slot follows the new model, pool or limit. Slots already
|
|
9
|
+
held are kept when a limit drops. An invalid change is refused and the
|
|
10
|
+
settings in effect stay; `status` names the error until the file is valid
|
|
11
|
+
again.
|
|
12
|
+
- `status` marks an asker whose execution hibernated with `hibernated: true`
|
|
13
|
+
(CLI: "hibernated, no slot"): it holds no provider slot until it is
|
|
14
|
+
answered. Both `status` forms list provider slots in use against their
|
|
15
|
+
limits, and the config in effect.
|
|
16
|
+
|
|
3
17
|
## 1.0.6
|
|
4
18
|
|
|
5
19
|
- A follow-up on a finished workflow shows as running work: in the dock, the
|
package/README.md
CHANGED
|
@@ -273,6 +273,11 @@ State lives in `~/.pi/durable-subagents`; set `DSA_HOME` to move it.
|
|
|
273
273
|
never stopped for memory.
|
|
274
274
|
- **onQuit:** `"pause"` (default) pauses a session's running workflows when
|
|
275
275
|
you quit that pi; `"continue"` lets them run on in the background.
|
|
276
|
+
- **Changes apply without a restart:** the orchestrator re-reads the file
|
|
277
|
+
when it changes. A new slot limit, pool or default model applies to the next
|
|
278
|
+
slot acquisition; slots already held are kept when a limit drops. An
|
|
279
|
+
invalid change is not applied, and `status` reports it next to the settings
|
|
280
|
+
still in effect (`config: <hash> since …`) and the slots held per provider.
|
|
276
281
|
|
|
277
282
|
## Switching back
|
|
278
283
|
|
package/dist/agent/main.js
CHANGED
|
@@ -325,7 +325,7 @@ export function registerMain(pi, ui) {
|
|
|
325
325
|
"run (action optional for exactly one launch form): agent+task; tasks:[call specs] parallel; chain:[call specs] sequential ({previous}); workflow:'./script.js' or source (runs.run(key,spec), runs.all([...]), emit(value), args, runs.input(name)). Optional name, cwd, usageBudget, maxCalls, inputs. With tasks/chain, top-level model, timeoutMs, budget, isolation, context, tools, skills, once are defaults for every step (a step's own value wins); a workflow/source script sets them per runs.run call. timeoutMs is milliseconds of active time (a number); omit it unless a hard limit is needed. Explicit unknown agents are rejected BEFORE creation, with available names; unknown script agents fail only their call.",
|
|
326
326
|
"agents: list names, descriptions, default models and source for this cwd; use these names for run.",
|
|
327
327
|
"send to:'<wid>/<key>' (bare '<wid>' only for a single-call workflow): steer on a running call delivers at the next safe point (receipt in status/UI); a steer to a call waiting on its question interrupts the question and the subagent usually asks again — use answer to answer it; sealed → finished:<status> — use kind 'follow-up'. follow-up continues a sealed call as generation g+1 or queues after a running turn. answer: give the qid (or just the call, or nothing when one question is open); to and rev are filled in. A question that needs the user's decision goes to the user; if you answer one yourself, tell the user what you chose. model switches at next provider request. Unknown targets list valid addresses. replaces:[rid] supersedes an earlier send.",
|
|
328
|
-
"stop target:<wid|<wid>/<key>> is terminal stopped (usage and partial edits kept); a sealed call → already-sealed:<status>, a finished workflow → terminal:<status>. drain holds existing workflows reversibly (new runs unaffected); resume [wid] releases held workflows. status: without wid, what runs, asks (with its answer address) or failed, finished workflows one line each; wid: one workflow, outputs clipped; wid+key: one call's full result; full:true: everything. A run's rid from {submitted:{rid}} works wherever a wid is expected. revise wid + workflow/source/args starts a revision.",
|
|
328
|
+
"stop target:<wid|<wid>/<key>> is terminal stopped (usage and partial edits kept); a sealed call → already-sealed:<status>, a finished workflow → terminal:<status>. drain holds existing workflows reversibly (new runs unaffected); resume [wid] releases held workflows. status: without wid, what runs, asks (with its answer address; hibernated:true holds no slot) or failed, finished workflows one line each, provider slots held/limit and the config in effect; wid: one workflow, outputs clipped; wid+key: one call's full result; full:true: everything. A run's rid from {submitted:{rid}} works wherever a wid is expected. revise wid + workflow/source/args starts a revision.",
|
|
329
329
|
"Control replies are {applied:true,rid} or {applied:false,reason,rid} when decided; otherwise {submitted:{rid}} after 10s.",
|
|
330
330
|
...(agents ? [`Available agents: ${agents}.`] : []),
|
|
331
331
|
"User sees a summary line above the editor; ↓ on an empty editor (or /subagents) opens the list, Enter watches live OR finished calls (finished transcripts remain on disk) and expands finished workflows. List keys: s steer (paste-capable input), x stop (confirm y), m model, a answer when asked, f follow-up on finished calls; action feedback appears in footer.",
|
package/dist/cli/main.js
CHANGED
|
@@ -65,13 +65,16 @@ export function renderStatus(wf) {
|
|
|
65
65
|
/** P25, T10: Render the compact status projection shared with the `subagents` tool. */
|
|
66
66
|
export function renderView(view) {
|
|
67
67
|
const lines = view.workflows.map(w => [`${w.wid}@${w.rev}${w.name ? ` ${w.name}` : ""}: ${w.status}${w.followUps ? " (follow-up running)" : ""}${w.error ? ` (${clip(w.error, 200)})` : ""} · ${w.done}/${w.planned ?? w.calls.length}${w.planned === undefined && w.status === "running" ? "+" : ""} done${w.usage.input || w.usage.output || w.usage.costUsd ? ` · ${formatUsage(w.usage)}` : ""}`,
|
|
68
|
-
...w.calls.map(c => ` ${c.key}@${c.gen} ${c.status ?? c.phase}${c.model ? ` ${c.model}` : ""}${c.tools ? ` tools:${c.tools}` : ""}${c.usage ? ` ${formatUsage(c.usage)}` : ""}${c.lastLine ? ` ${JSON.stringify(c.lastLine)}` : c.error ? ` (${c.error})` : ""}`),
|
|
68
|
+
...w.calls.map(c => ` ${c.key}@${c.gen} ${c.status ?? c.phase}${c.hibernated ? " (hibernated, no slot)" : ""}${c.model ? ` ${c.model}` : ""}${c.tools ? ` tools:${c.tools}` : ""}${c.usage ? ` ${formatUsage(c.usage)}` : ""}${c.lastLine ? ` ${JSON.stringify(c.lastLine)}` : c.error ? ` (${c.error})` : ""}`),
|
|
69
69
|
...w.attention.map(a => ` ${a.kind}: ${JSON.stringify(a.text.split("\n")[0])}`)].join("\n"));
|
|
70
70
|
if (view.paused)
|
|
71
71
|
lines.unshift(`${view.paused} (pi-durable-subagents resume)`);
|
|
72
72
|
if (view.olderFinished)
|
|
73
73
|
lines.push(`(+${view.olderFinished} older finished workflows; status <wid> shows one in detail)`);
|
|
74
|
-
|
|
74
|
+
const footer = [view.slots?.length ? `slots: ${view.slots.join(", ")}` : "", view.config ? `config: ${view.config}` : "", view.configRejected ? `config.json rejected: ${view.configRejected}` : ""].filter(Boolean);
|
|
75
|
+
if (!lines.length)
|
|
76
|
+
lines.push("No workflows");
|
|
77
|
+
return [...lines, ...footer].join("\n");
|
|
75
78
|
}
|
|
76
79
|
function snapshots(home, wid) {
|
|
77
80
|
if (!wid)
|
|
@@ -0,0 +1,147 @@
|
|
|
1
|
+
// $DSA_HOME/config.json for the orchestrator: read at start and again whenever the file changes, so a new slot limit,
|
|
2
|
+
// pool order or K-parameter needs no orchestrator restart (a restart at the wrong moment used to cost a dispatch).
|
|
3
|
+
// Orchestrator ledger entries: config{hash,config} — the settings in effect from then on (at start, or after a change);
|
|
4
|
+
// config-rejected{hash,error} — a changed file that was not applied; the settings before it stay in effect.
|
|
5
|
+
// A reload changes the shared config object in place: every later read sees it (the next slot acquisition, model
|
|
6
|
+
// resolution or check). Slots already held are kept when a limit drops; timers of running executions keep their period.
|
|
7
|
+
import { readFile, stat } from "node:fs/promises";
|
|
8
|
+
import { join } from "node:path";
|
|
9
|
+
import { contentHash } from "../kernel/ids.js";
|
|
10
|
+
/** The keys the orchestrator reads; config.json also holds pi-side settings (ui, onQuit) that it ignores. */
|
|
11
|
+
const KEYS = ["defaultModel", "pools", "providers", "memory", "k"];
|
|
12
|
+
const K = ["lossBound", "checkpointMs", "stallMs", "progressMs", "switchTimeoutMs", "idleExitMs", "trackerMs", "hibernateMs", "spawnBudget"];
|
|
13
|
+
export const configPath = (home) => join(home, "config.json");
|
|
14
|
+
/** The orchestrator's part of a parsed config.json. */
|
|
15
|
+
export function orchestratorSettings(raw) {
|
|
16
|
+
const value = raw && typeof raw === "object" && !Array.isArray(raw) ? raw : {};
|
|
17
|
+
return Object.fromEntries(KEYS.filter(k => value[k] !== undefined).map(k => [k, value[k]]));
|
|
18
|
+
}
|
|
19
|
+
export const configHash = (config) => contentHash(orchestratorSettings(config)).slice(0, 12);
|
|
20
|
+
const record = (v) => !!v && typeof v === "object" && !Array.isArray(v);
|
|
21
|
+
const count = (v) => typeof v === "number" && Number.isFinite(v) && v >= 0;
|
|
22
|
+
/** What is wrong with a parsed config.json for the orchestrator, or undefined. Unknown keys are allowed. */
|
|
23
|
+
export function configProblem(raw) {
|
|
24
|
+
if (!record(raw))
|
|
25
|
+
return "config.json must be a JSON object";
|
|
26
|
+
const { defaultModel, pools, providers, memory, k } = raw;
|
|
27
|
+
if (defaultModel !== undefined && typeof defaultModel !== "string")
|
|
28
|
+
return "defaultModel must be a string";
|
|
29
|
+
if (pools !== undefined) {
|
|
30
|
+
if (!record(pools))
|
|
31
|
+
return "pools must map names to model lists";
|
|
32
|
+
for (const [name, list] of Object.entries(pools))
|
|
33
|
+
if (!Array.isArray(list) || !list.length || !list.every(m => typeof m === "string" && m))
|
|
34
|
+
return `pools.${name} must be a nonempty list of "provider/id[:thinking]"`;
|
|
35
|
+
}
|
|
36
|
+
if (providers !== undefined) {
|
|
37
|
+
if (!record(providers))
|
|
38
|
+
return "providers must map provider names to { slots }";
|
|
39
|
+
for (const [name, p] of Object.entries(providers))
|
|
40
|
+
if (!record(p) || !Number.isSafeInteger(p.slots) || Number(p.slots) < 0)
|
|
41
|
+
return `providers.${name}.slots must be a nonnegative integer`;
|
|
42
|
+
}
|
|
43
|
+
if (memory !== undefined) {
|
|
44
|
+
if (!record(memory))
|
|
45
|
+
return "memory must be { reserveMb?, perChildMb? }";
|
|
46
|
+
for (const key of ["reserveMb", "perChildMb"])
|
|
47
|
+
if (memory[key] !== undefined && !count(memory[key]))
|
|
48
|
+
return `memory.${key} must be a nonnegative number`;
|
|
49
|
+
}
|
|
50
|
+
if (k !== undefined) {
|
|
51
|
+
if (!record(k))
|
|
52
|
+
return "k must map K-parameters to numbers";
|
|
53
|
+
for (const [key, v] of Object.entries(k)) {
|
|
54
|
+
if (!K.includes(key))
|
|
55
|
+
return `k.${key} is not a K-parameter (${K.join(", ")})`;
|
|
56
|
+
if (!count(v) || (key !== "lossBound" && key !== "spawnBudget" && v === 0))
|
|
57
|
+
return `k.${key} must be a positive number`;
|
|
58
|
+
}
|
|
59
|
+
}
|
|
60
|
+
}
|
|
61
|
+
/** Parse config.json; a missing file is the empty config. A JSON error throws. */
|
|
62
|
+
export async function readConfig(path) {
|
|
63
|
+
try {
|
|
64
|
+
return JSON.parse(await readFile(path, "utf8"));
|
|
65
|
+
}
|
|
66
|
+
catch (error) {
|
|
67
|
+
if (error.code === "ENOENT")
|
|
68
|
+
return {};
|
|
69
|
+
throw error;
|
|
70
|
+
}
|
|
71
|
+
}
|
|
72
|
+
/** Record the settings in effect unless the ledger already ends on the same ones (a restart with an unchanged file).
|
|
73
|
+
* After a rejection they are recorded again: that confirms the file is back to settings in effect. */
|
|
74
|
+
export async function recordConfig(orch, config) {
|
|
75
|
+
const hash = configHash(config), last = orch.entries().findLast(e => e.type === "config" || e.type === "config-rejected");
|
|
76
|
+
if (last?.type !== "config" || last.hash !== hash)
|
|
77
|
+
await orch.append("config", { hash, config: orchestratorSettings(config) });
|
|
78
|
+
return hash;
|
|
79
|
+
}
|
|
80
|
+
/** Replace the contents of the shared config object; readers hold the object, not a copy. */
|
|
81
|
+
function applyInPlace(target, next) {
|
|
82
|
+
const t = target;
|
|
83
|
+
for (const key of KEYS)
|
|
84
|
+
delete t[key];
|
|
85
|
+
Object.assign(t, orchestratorSettings(next));
|
|
86
|
+
}
|
|
87
|
+
/** Identity of the file's current version (taken before a read, so a change during the read is seen next time). */
|
|
88
|
+
export async function configStamp(path) {
|
|
89
|
+
try {
|
|
90
|
+
const s = await stat(path);
|
|
91
|
+
return `${s.ino}:${s.size}:${s.mtimeMs}`;
|
|
92
|
+
}
|
|
93
|
+
catch (error) {
|
|
94
|
+
if (error.code === "ENOENT")
|
|
95
|
+
return "missing";
|
|
96
|
+
throw error;
|
|
97
|
+
}
|
|
98
|
+
}
|
|
99
|
+
/** Watch config.json from the version `stamp` (read at start) and apply each valid change in place. */
|
|
100
|
+
/** `apply` runs the change where readers cannot observe half of it (the executor's admission section). */
|
|
101
|
+
export function watchConfig(options) {
|
|
102
|
+
const { path, config, orch } = options;
|
|
103
|
+
let stamp = options.stamp, running, rejected, stopped = false;
|
|
104
|
+
const once = async () => {
|
|
105
|
+
const now = await configStamp(path);
|
|
106
|
+
if (now === stamp || stopped)
|
|
107
|
+
return;
|
|
108
|
+
let raw;
|
|
109
|
+
try {
|
|
110
|
+
raw = await readConfig(path);
|
|
111
|
+
}
|
|
112
|
+
catch (error) {
|
|
113
|
+
// A half-written file reads as invalid JSON: retry on the next change of the file, report it once per version.
|
|
114
|
+
stamp = now;
|
|
115
|
+
const hash = `invalid:${now}`;
|
|
116
|
+
if (rejected !== hash) {
|
|
117
|
+
rejected = hash;
|
|
118
|
+
await orch.append("config-rejected", { error: `config.json is not valid JSON: ${error.message}` });
|
|
119
|
+
}
|
|
120
|
+
return;
|
|
121
|
+
}
|
|
122
|
+
stamp = now;
|
|
123
|
+
// Validate the content first: `[]` or `null` project to no settings and must not pass as "unchanged".
|
|
124
|
+
const problem = configProblem(raw), next = orchestratorSettings(raw), hash = configHash(next);
|
|
125
|
+
if (problem) {
|
|
126
|
+
if (rejected !== hash) {
|
|
127
|
+
rejected = hash;
|
|
128
|
+
await orch.append("config-rejected", { hash, error: problem });
|
|
129
|
+
}
|
|
130
|
+
return;
|
|
131
|
+
}
|
|
132
|
+
rejected = undefined;
|
|
133
|
+
let changed = false;
|
|
134
|
+
await (options.apply ?? (change => change()))(async () => {
|
|
135
|
+
changed = hash !== configHash(config);
|
|
136
|
+
if (changed)
|
|
137
|
+
applyInPlace(config, next);
|
|
138
|
+
await recordConfig(orch, config); // also confirms a return to the settings in effect after a rejection
|
|
139
|
+
});
|
|
140
|
+
if (changed)
|
|
141
|
+
options.onApplied?.();
|
|
142
|
+
};
|
|
143
|
+
const check = () => running ??= once().catch(error => console.error(`durable-subagents: config.json check failed: ${String(error)}`)).finally(() => { running = undefined; });
|
|
144
|
+
const timer = setInterval(check, options.intervalMs ?? 1000);
|
|
145
|
+
timer.unref();
|
|
146
|
+
return { stop: async () => { stopped = true; clearInterval(timer); await running; }, check };
|
|
147
|
+
}
|
|
@@ -389,13 +389,18 @@ export default function createExecutor(ledgers, options = {}) {
|
|
|
389
389
|
}
|
|
390
390
|
const stopped = (a) => a.stopped || entriesFor(a.ticket.journal, a.ticket.callId).some(e => e.type === "stop-intent");
|
|
391
391
|
const interrupted = (a) => closed || a.suspended || a.retired || stopped(a);
|
|
392
|
-
|
|
392
|
+
/** Wait for a slot. The candidates are decided again on every attempt, inside the serial section that also applies
|
|
393
|
+
* config.json changes: a queued call follows a reloaded defaultModel, pool or limit, never a mix of two versions. */
|
|
394
|
+
async function acquire(a, exec, decide) {
|
|
395
|
+
let continuation = false;
|
|
393
396
|
for (;;) {
|
|
394
397
|
let signal;
|
|
395
398
|
const changed = new Promise(resolve => { signal = resolve; waiters.add(resolve); });
|
|
396
399
|
const chosen = await serial(async () => {
|
|
397
400
|
if (interrupted(a) || await workflowReached(a.ticket))
|
|
398
401
|
return;
|
|
402
|
+
const decision = await decide(), models = decision.candidates, pool = decision.pool;
|
|
403
|
+
continuation = decision.continuation;
|
|
399
404
|
for (const model of models) {
|
|
400
405
|
if (!continuation && pool && skipped(pool, model))
|
|
401
406
|
continue;
|
|
@@ -434,7 +439,7 @@ export default function createExecutor(ledgers, options = {}) {
|
|
|
434
439
|
});
|
|
435
440
|
if (chosen || interrupted(a) || reached(totalUsage(a.ticket.journal.entries()), a.ticket.workflowBudget)) {
|
|
436
441
|
waiters.delete(signal);
|
|
437
|
-
return chosen;
|
|
442
|
+
return { model: chosen, continuation };
|
|
438
443
|
}
|
|
439
444
|
const timer = setTimeout(signal, config.k?.trackerMs ?? 1000);
|
|
440
445
|
try {
|
|
@@ -599,13 +604,13 @@ export default function createExecutor(ledgers, options = {}) {
|
|
|
599
604
|
}
|
|
600
605
|
let decision;
|
|
601
606
|
try {
|
|
602
|
-
decision = await launchModel(t, entries, previous);
|
|
607
|
+
decision = await acquire(a, exec, () => launchModel(t, entries, previous));
|
|
603
608
|
}
|
|
604
609
|
catch (error) {
|
|
605
610
|
await fence(journal, exec, { park: a });
|
|
606
611
|
return finish(journal, t.callId, exec, makeResult("failed", "", String(error)));
|
|
607
612
|
}
|
|
608
|
-
const model =
|
|
613
|
+
const model = decision.model;
|
|
609
614
|
if (!model || interrupted(a)) {
|
|
610
615
|
await fence(journal, exec, { park: a });
|
|
611
616
|
return finish(journal, t.callId, exec, !interrupted(a) ? makeResult("failed", "", "workflow budget reached") : makeResult("stopped"));
|
|
@@ -687,6 +692,7 @@ export default function createExecutor(ledgers, options = {}) {
|
|
|
687
692
|
await (await outbox).send(envelope.to, envelope.kind, envelope.body, envelope.cond, { rid: String(e.rid2) });
|
|
688
693
|
}
|
|
689
694
|
return {
|
|
695
|
+
async reconfigure(apply) { await serial(apply); wake(); },
|
|
690
696
|
run(ticket) {
|
|
691
697
|
const existing = completed.get(ticket.callId);
|
|
692
698
|
if (existing)
|
|
@@ -123,7 +123,7 @@ export async function observeExecution(d) {
|
|
|
123
123
|
}
|
|
124
124
|
});
|
|
125
125
|
});
|
|
126
|
-
let scanning = false,
|
|
126
|
+
let scanning = false, failingSince;
|
|
127
127
|
const scan = () => {
|
|
128
128
|
if (scanning)
|
|
129
129
|
return;
|
|
@@ -133,12 +133,14 @@ export async function observeExecution(d) {
|
|
|
133
133
|
// One failed process-table scan is no evidence, not a reason to fence a live child: on macOS `ps` can miss its
|
|
134
134
|
// 2 s deadline right after the orchestrator was stopped or the machine slept (seen on CI as a spurious loss and
|
|
135
135
|
// a rerun). Only a scan that keeps failing for a minute ends the observation.
|
|
136
|
+
// Measured in elapsed time: the scan period is fixed at start while k.trackerMs can be reloaded.
|
|
136
137
|
try {
|
|
137
138
|
clock.scan(await d.track());
|
|
138
|
-
|
|
139
|
+
failingSince = undefined;
|
|
139
140
|
}
|
|
140
141
|
catch (error) {
|
|
141
|
-
|
|
142
|
+
failingSince ??= performance.now();
|
|
143
|
+
if (performance.now() - failingSince >= 60_000)
|
|
142
144
|
throw error;
|
|
143
145
|
}
|
|
144
146
|
const nextSize = (await fileStat(session).catch(() => ({ size: 0 }))).size;
|
|
@@ -6,7 +6,7 @@ var __rewriteRelativeImportExtension = (this && this.__rewriteRelativeImportExte
|
|
|
6
6
|
}
|
|
7
7
|
return path;
|
|
8
8
|
};
|
|
9
|
-
import { mkdir
|
|
9
|
+
import { mkdir } from 'node:fs/promises';
|
|
10
10
|
import { fileURLToPath } from 'node:url';
|
|
11
11
|
import { realpathSync } from 'node:fs';
|
|
12
12
|
import { join, resolve } from 'node:path';
|
|
@@ -14,6 +14,7 @@ import { dsaHome, orchLedger, orchLock } from "../paths.js";
|
|
|
14
14
|
import { openJournal } from "../kernel/journal.js";
|
|
15
15
|
import { OsLock } from "../platform/lock.js";
|
|
16
16
|
import { Engine } from "./engine.js";
|
|
17
|
+
import { configPath, configProblem, configStamp, readConfig, recordConfig, watchConfig } from "./config.js";
|
|
17
18
|
/** P2, C4, K6: Acquire the single authority before opening ledgers, recover, and idle-exit. */
|
|
18
19
|
export async function main(options = {}) {
|
|
19
20
|
const home = options.home ?? dsaHome();
|
|
@@ -21,27 +22,34 @@ export async function main(options = {}) {
|
|
|
21
22
|
const lock = await new OsLock().tryAcquire(orchLock(home));
|
|
22
23
|
if (!lock)
|
|
23
24
|
return;
|
|
24
|
-
let engine, ledgers;
|
|
25
|
+
let engine, ledgers, watcher;
|
|
25
26
|
try {
|
|
26
|
-
let config = options.config;
|
|
27
|
+
let config = options.config, stamp;
|
|
27
28
|
if (!config) {
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
}
|
|
29
|
+
// The stamp is taken before the read: a change made while reading is applied by the first check.
|
|
30
|
+
stamp = await configStamp(configPath(home));
|
|
31
|
+
const raw = await readConfig(configPath(home));
|
|
32
|
+
const problem = configProblem(raw);
|
|
33
|
+
if (problem)
|
|
34
|
+
console.error(`durable-subagents: config.json: ${problem}`);
|
|
35
|
+
config = raw;
|
|
36
36
|
}
|
|
37
37
|
ledgers = { home, config, orch: await openJournal(orchLedger(home)) };
|
|
38
38
|
const factory = options.executor ?? (await import(__rewriteRelativeImportExtension(new URL(import.meta.url.endsWith('.ts') ? './executor/index.ts' : './executor/index.js', import.meta.url).href))).default;
|
|
39
|
-
|
|
39
|
+
const executor = factory(ledgers);
|
|
40
|
+
engine = new Engine(ledgers, executor, options);
|
|
41
|
+
if (stamp !== undefined) {
|
|
42
|
+
await recordConfig(ledgers.orch, config);
|
|
43
|
+
// A change applies between slot admissions, so one admission never mixes two versions of the limits.
|
|
44
|
+
watcher = watchConfig({ path: configPath(home), stamp, config, orch: ledgers.orch, intervalMs: Math.min(1000, config.k?.trackerMs ?? 1000),
|
|
45
|
+
apply: change => executor.reconfigure ? executor.reconfigure(change) : change() });
|
|
46
|
+
}
|
|
40
47
|
await engine.recover();
|
|
41
48
|
await engine.loop(options.signal);
|
|
42
49
|
}
|
|
43
50
|
finally {
|
|
44
51
|
try {
|
|
52
|
+
await watcher?.stop();
|
|
45
53
|
await engine?.close();
|
|
46
54
|
}
|
|
47
55
|
finally {
|
|
@@ -76,6 +76,7 @@ export function snapshotFromEntries(wid, entries) {
|
|
|
76
76
|
const done = terminal?.type === JT.done ? terminal : undefined;
|
|
77
77
|
const calls = new Map();
|
|
78
78
|
const byExec = new Map();
|
|
79
|
+
const hibernating = new Map(); // callId → exec that decided to hibernate and is not fenced yet
|
|
79
80
|
const resolved = new Set(entries.filter(e => e.type === JT.attentionResolved).map(e => `${e.id}@${e.rev}`));
|
|
80
81
|
const generations = new Set(), retired = new Set();
|
|
81
82
|
const attention = [];
|
|
@@ -115,12 +116,29 @@ export function snapshotFromEntries(wid, entries) {
|
|
|
115
116
|
call.startedAt ??= e.ts;
|
|
116
117
|
call.lastActivity = e.ts;
|
|
117
118
|
byExec.set(call.exec, call);
|
|
119
|
+
delete call.hibernated;
|
|
120
|
+
hibernating.delete(call.callId);
|
|
121
|
+
}
|
|
122
|
+
else if (e.type === "hibernated") {
|
|
123
|
+
// The decision to hibernate precedes the fence; the slot is released only once the execution is fenced.
|
|
124
|
+
const call = calls.get(String(e.call));
|
|
125
|
+
if (call?.exec && call.exec === e.exec)
|
|
126
|
+
hibernating.set(call.callId, call.exec);
|
|
127
|
+
}
|
|
128
|
+
else if (e.type === "answer-bound" || (e.type === "resumed" && e.call)) {
|
|
129
|
+
const call = calls.get(String(e.call));
|
|
130
|
+
if (call) {
|
|
131
|
+
delete call.hibernated;
|
|
132
|
+
hibernating.delete(call.callId);
|
|
133
|
+
}
|
|
118
134
|
}
|
|
119
135
|
else if (e.type === JT.fenced) {
|
|
120
136
|
// A fenced execution without a seal (stop-all, a quit pi, a loss before its continuation) waits to run again.
|
|
121
137
|
const call = byExec.get(String(e.exec));
|
|
122
138
|
if (call && call.phase === "running")
|
|
123
139
|
call.phase = "queued";
|
|
140
|
+
if (call && hibernating.get(call.callId) === e.exec)
|
|
141
|
+
call.hibernated = true;
|
|
124
142
|
}
|
|
125
143
|
else if (e.type === "selected") {
|
|
126
144
|
const call = byExec.get(String(e.exec)), m = e.model;
|
|
@@ -158,6 +176,8 @@ export function snapshotFromEntries(wid, entries) {
|
|
|
158
176
|
call.phase = "sealed";
|
|
159
177
|
call.result = e.result;
|
|
160
178
|
call.endedAt = e.ts;
|
|
179
|
+
delete call.hibernated;
|
|
180
|
+
hibernating.delete(call.callId);
|
|
161
181
|
}
|
|
162
182
|
else if (e.type === JT.attention) {
|
|
163
183
|
const item = e.item;
|
|
@@ -189,6 +209,9 @@ export function snapshotFromEntries(wid, entries) {
|
|
|
189
209
|
if (call && (call.phase === "running" || call.phase === "queued"))
|
|
190
210
|
call.phase = "asking"; // a hibernated asker is fenced
|
|
191
211
|
}
|
|
212
|
+
for (const c of calls.values())
|
|
213
|
+
if (c.hibernated && c.phase !== "asking")
|
|
214
|
+
delete c.hibernated;
|
|
192
215
|
const usageOf = (callId) => sealedUsage.get(callId) ?? live.get(callId);
|
|
193
216
|
const list = [...calls.values()];
|
|
194
217
|
const counts = { queued: 0, running: 0, asking: 0, sealed: 0 };
|
|
@@ -302,7 +325,7 @@ export function compactWorkflow(wf) {
|
|
|
302
325
|
const r = c.result, last = r?.output?.split("\n").map(l => l.trim()).filter(Boolean).at(-1);
|
|
303
326
|
return { key: c.key, gen: c.gen, callId: c.callId, phase: c.phase, ...(r ? { status: r.status, ok: r.ok } : {}),
|
|
304
327
|
...(c.model ? { model: c.model } : {}), ...(c.tools ? { tools: c.tools } : {}), ...(c.pending ? { pending: c.pending } : {}), ...(c.switching ? { switching: c.switching } : {}), ...(nonzero(c.usage) ? { usage: c.usage } : {}),
|
|
305
|
-
...(last ? { lastLine: clip(last, 200) } : {}), ...(r?.error ? { error: clip(r.error, 300) } : {}) };
|
|
328
|
+
...(last ? { lastLine: clip(last, 200) } : {}), ...(r?.error ? { error: clip(r.error, 300) } : {}), ...(c.hibernated ? { hibernated: true } : {}) };
|
|
306
329
|
}),
|
|
307
330
|
attention: wf.attention.map(a => ({ id: a.id, rev: a.rev, kind: a.kind, text: clip(a.text, 300), ...(a.call ? { call: a.call } : {}), ...(a.qid ? { qid: a.qid } : {}) })),
|
|
308
331
|
...(wf.paused ? { paused: true } : {}), ...(wf.followUps ? { followUps: wf.followUps } : {}),
|
|
@@ -338,7 +361,8 @@ export function statusView(home, options = {}) {
|
|
|
338
361
|
const hidden = all.length - shown.length;
|
|
339
362
|
const paused = all.filter(w => w.paused).length;
|
|
340
363
|
return { workflows: shown.map(compactWorkflow), ...(hidden ? { olderFinished: hidden, hint: "status wid=<wid> shows any workflow in detail" } : {}),
|
|
341
|
-
...(paused ? { paused: `${paused} workflow${paused > 1 ? "s" : ""} paused (stop-all, drain or a quit pi) since ${new Date(since).toISOString()}; resume continues them (new runs are not affected)` } : {})
|
|
364
|
+
...(paused ? { paused: `${paused} workflow${paused > 1 ? "s" : ""} paused (stop-all, drain or a quit pi) since ${new Date(since).toISOString()}; resume continues them (new runs are not affected)` } : {}),
|
|
365
|
+
...slotsView(home) };
|
|
342
366
|
}
|
|
343
367
|
const FINAL = ["done", "failed", "stopped"];
|
|
344
368
|
/** Finished with nothing left running (a follow-up on a finished workflow is live work). */
|
|
@@ -347,6 +371,33 @@ const age = (ms) => ms < 60_000 ? `${Math.max(0, Math.round(ms / 1000))}s` : ms
|
|
|
347
371
|
const tokens = (u) => { const n = u ? u.input + u.output : 0; return n >= 1e6 ? `${(n / 1e6).toFixed(1)}M` : n >= 1e3 ? `${(n / 1e3).toFixed(1)}K` : String(n); };
|
|
348
372
|
/** The latest generation of every key, in first-call order. */
|
|
349
373
|
const latestCalls = (wf) => [...new Map(wf.calls.map(c => [c.key, c])).values()];
|
|
374
|
+
/** Provider slots from the orchestrator ledger: holders per provider (hold/release{pool,slot,exec}) and the limits of
|
|
375
|
+
* the settings in effect (the latest config{hash,config}); config-rejected after it is reported too. */
|
|
376
|
+
export function slotsView(home, now = Date.now()) {
|
|
377
|
+
const held = new Map();
|
|
378
|
+
let config, rejected;
|
|
379
|
+
for (const e of readJournalSnapshot(orchLedger(home))) {
|
|
380
|
+
if (e.type === "hold")
|
|
381
|
+
held.set(`${e.pool}:${e.slot}`, e);
|
|
382
|
+
else if (e.type === "release" && held.get(`${e.pool}:${e.slot}`)?.exec === e.exec)
|
|
383
|
+
held.delete(`${e.pool}:${e.slot}`);
|
|
384
|
+
else if (e.type === "config") {
|
|
385
|
+
config = e;
|
|
386
|
+
rejected = undefined;
|
|
387
|
+
}
|
|
388
|
+
else if (e.type === "config-rejected")
|
|
389
|
+
rejected = e;
|
|
390
|
+
}
|
|
391
|
+
const limits = (config?.config?.providers) ?? {};
|
|
392
|
+
const holders = new Map();
|
|
393
|
+
for (const e of held.values())
|
|
394
|
+
if (e.pool !== "memory")
|
|
395
|
+
holders.set(String(e.pool), (holders.get(String(e.pool)) ?? 0) + 1);
|
|
396
|
+
const names = [...new Set([...Object.keys(limits), ...holders.keys()])].sort();
|
|
397
|
+
const slots = names.map(p => { const n = holders.get(p) ?? 0, limit = limits[p]?.slots; return typeof limit === "number" ? `${p} ${n}/${limit}` : `${p} ${n} (no limit)`; });
|
|
398
|
+
return { ...(slots.length ? { slots } : {}), ...(config ? { config: `${String(config.hash)} since ${age(now - config.ts)} ago` } : {}),
|
|
399
|
+
...(rejected ? { configRejected: `${clip(String(rejected.error), 200)} (${age(now - rejected.ts)} ago); ${config ? String(config.hash) : "the start settings"} stay in effect` } : {}) };
|
|
400
|
+
}
|
|
350
401
|
/** Tool status without a wid: what runs, what waits for an answer and what failed, with finished workflows one line each.
|
|
351
402
|
* The full view (every call's last line and usage) ran to tens of thousands of tokens on a busy home. */
|
|
352
403
|
export function statusBrief(home, options = {}) {
|
|
@@ -366,9 +417,11 @@ export function statusBrief(home, options = {}) {
|
|
|
366
417
|
...(live && c.startedAt !== undefined ? { for: age(now - c.startedAt) } : {}),
|
|
367
418
|
...(live && c.phase !== "asking" && quiet !== undefined && quiet >= 60_000 ? { quiet: age(quiet) } : {}),
|
|
368
419
|
...(live && c.startedAt !== undefined ? { tokens: tokens(c.usage) } : {}),
|
|
369
|
-
...(c.result ? { status: c.result.status, ...(c.result.error ? { error: clip(c.result.error, 200) } : {}) } : {})
|
|
420
|
+
...(c.result ? { status: c.result.status, ...(c.result.error ? { error: clip(c.result.error, 200) } : {}) } : {}),
|
|
421
|
+
...(c.hibernated ? { hibernated: true } : {}) };
|
|
370
422
|
});
|
|
371
|
-
const asking = open.filter(a => a.kind === "question" && a.call).map(a => ({ to: `${w.wid}/${callKey(a.call)}`, ...(a.qid ? { qid: a.qid } : {}),
|
|
423
|
+
const asking = open.filter(a => a.kind === "question" && a.call).map(a => ({ to: `${w.wid}/${callKey(a.call)}`, ...(a.qid ? { qid: a.qid } : {}),
|
|
424
|
+
...(w.calls.some(c => c.callId === a.call && c.hibernated) ? { hibernated: true } : {}), question: clip(a.text, 300) }));
|
|
372
425
|
const alerts = open.filter(a => a.kind !== "question").map(a => `${a.kind}${a.call ? ` ${w.wid}/${callKey(a.call)}` : ""}: ${clip(a.text, 200)}`);
|
|
373
426
|
return { wid: w.wid, ...(w.name ? { name: w.name } : {}), status: w.status, ...(w.paused ? { paused: true } : {}), ...(w.followUps ? { followUps: w.followUps } : {}),
|
|
374
427
|
progress: `${p.done}/${p.total}${p.plus ? "+" : ""}`, tokens: tokens(w.usage), calls,
|
|
@@ -383,7 +436,7 @@ export function statusBrief(home, options = {}) {
|
|
|
383
436
|
return {
|
|
384
437
|
active: active.filter(mine).map(brief), ...(others.length ? { otherSessions: others.slice(0, 10).map(line) } : {}),
|
|
385
438
|
finished: finished.slice(0, keep).map(line), ...(finished.length > keep ? { olderFinished: finished.length - keep } : {}),
|
|
386
|
-
...(paused ? { paused } : {}),
|
|
439
|
+
...(paused ? { paused } : {}), ...slotsView(home, now),
|
|
387
440
|
hint: "status wid=<wid> shows one workflow (outputs clipped); add key=<key> for one call's full result, or full:true for everything",
|
|
388
441
|
};
|
|
389
442
|
}
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pi-durable-subagents",
|
|
3
|
-
"version": "1.0.
|
|
3
|
+
"version": "1.0.7",
|
|
4
4
|
"description": "Subagents for pi that never lose work and never do it twice. Crash-safe workflows, automatic recovery, and a live view just like the main agent.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"license": "MIT",
|