@awebai/oats 0.40.1 → 0.40.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/bin/oats.mjs +6 -6
- package/docs/capabilities.md +35 -2
- package/docs/desktop-cli-api.md +89 -7
- package/docs/execution-targets.md +30 -29
- package/docs/implementation.md +48 -5
- package/docs/oats-local.schema.json +2 -2
- package/docs/official-catalog.md +1 -1
- package/docs/packages.md +5 -5
- package/docs/release-lane.md +8 -4
- package/docs/release-notes/v0.40.2.md +120 -0
- package/docs/schedules.md +40 -11
- package/docs/servers.md +3 -1
- package/docs/workspaces.md +1 -1
- package/lib/core.mjs +27 -12
- package/lib/dir-lock.mjs +7 -4
- package/lib/instance-events.mjs +130 -46
- package/lib/schedule-command-child.mjs +39 -9
- package/lib/schedule.mjs +57 -28
- package/lib/servers.mjs +7 -1
- package/lib/session-input.mjs +42 -47
- package/package-catalog.json +1 -1
- package/package.json +1 -1
- package/skills/oats-getting-started/SKILL.md +1 -1
package/docs/schedules.md
CHANGED
|
@@ -86,8 +86,19 @@ jobs) sets it so that its command jobs can be told apart.
|
|
|
86
86
|
sends SIGTERM to the child's process group, allows two seconds for cleanup,
|
|
87
87
|
then sends SIGKILL if the group remains. It observes the direct child's exit
|
|
88
88
|
before returning; inherited output pipes cannot hold the tick indefinitely.
|
|
89
|
-
|
|
90
|
-
|
|
89
|
+
SIGINT, SIGTERM or SIGHUP received by the supervisor enters that same cleanup
|
|
90
|
+
once; repeated signals do not bypass it. A timeout or interrupted supervisor
|
|
91
|
+
leaves effects unconfirmed even if the child printed an envelope. Spawn
|
|
92
|
+
previews use the same bounded runner.
|
|
93
|
+
|
|
94
|
+
Cleanup covers the owned process group. A descendant that creates its own
|
|
95
|
+
session can escape it; inherited pipes are bounded but that escaped process
|
|
96
|
+
is not terminated by this group cleanup. The supervisor checks the group when
|
|
97
|
+
the leader exits and never signals it after observing it empty. This reduces
|
|
98
|
+
the group-ID reuse window; it does not eliminate PID reuse races. SIGKILL,
|
|
99
|
+
OOM and other unrecoverable supervisor deaths cannot run JavaScript handlers:
|
|
100
|
+
cleanup is not guaranteed then. Without a private supervisor receipt, the
|
|
101
|
+
scheduler keeps the unknown attempt and its slot until reconciliation.
|
|
91
102
|
- **wake** `{…, home, message}` — every due minute inspects the instance at
|
|
92
103
|
`home`. Running: `message` is delivered once as terminal input (bracketed
|
|
93
104
|
paste plus Enter), never an interrupt. Not running: the home is started with
|
|
@@ -273,9 +284,12 @@ add` / `oats schedule add` definitions need no trust.
|
|
|
273
284
|
|
|
274
285
|
**Opting out on one host.** `oats trigger disable <member>/<id>` writes
|
|
275
286
|
`triggers.disabled`, and `oats schedule disable <member>/<id>` writes
|
|
276
|
-
`schedules.disabled`, in `oats-local.yaml`; `enable` removes the entry.
|
|
277
|
-
|
|
278
|
-
`
|
|
287
|
+
`schedules.disabled`, in `oats-local.yaml`; `enable` removes the entry. The
|
|
288
|
+
schedule part of a qualified ID accepts up to 100 characters in both
|
|
289
|
+
`schedules.disabled` and named `automations.trust` entries. Trigger definitions
|
|
290
|
+
and `triggers.disabled` keep their 40-character limit; the shared trust list
|
|
291
|
+
does not widen trigger IDs. A workspace definition is never edited or removed
|
|
292
|
+
from the CLI (`update` and `remove` answer `E_AUTOMATION_WORKSPACE`): change the file in Git.
|
|
279
293
|
|
|
280
294
|
**Refresh.**
|
|
281
295
|
|
|
@@ -329,12 +343,27 @@ current choices. Invalid values fail with `E_BAD_ARGS` before registration or
|
|
|
329
343
|
timer changes. `host status` reports the effective `maxConcurrent` and
|
|
330
344
|
`triggersMaxConcurrent` (`null` when uncapped).
|
|
331
345
|
|
|
332
|
-
The registry stores explicit choices only.
|
|
333
|
-
|
|
334
|
-
|
|
335
|
-
|
|
336
|
-
|
|
337
|
-
|
|
346
|
+
The registry stores explicit choices only. Reading status or the registry takes
|
|
347
|
+
no registry lock and creates or rewrites no files or directories. A reader
|
|
348
|
+
interprets a pre-migration stored `maxConcurrent: 1` as the default of five in
|
|
349
|
+
memory. Registration, unregistration and cap updates persist that migration
|
|
350
|
+
once, under the registry lock, even if the workspace membership is unchanged.
|
|
351
|
+
Other explicit values and the independent trigger cap survive.
|
|
352
|
+
|
|
353
|
+
A pre-migration hand-set one is indistinguishable from the old implicit one;
|
|
354
|
+
both follow that compatibility choice. When a write migrates one to the default,
|
|
355
|
+
it prints a notice on stderr with `oats schedule host install --max-concurrent 1`
|
|
356
|
+
to restore one if needed. Pure reads stay silent and JSON stdout is unchanged.
|
|
357
|
+
A later explicit one survives writes by current kernels. Older binaries sharing
|
|
358
|
+
the registry can write one back while preserving `capsVersion: 2`; there is no
|
|
359
|
+
provenance to distinguish that from a new explicit one. Stop mixed-version
|
|
360
|
+
writes and explicitly set the desired cap with the supported CLI.
|
|
361
|
+
|
|
362
|
+
A present invalid `maxConcurrent` is `E_SCHEDULE_INVALID`, not a fallback to
|
|
363
|
+
five. Correct it with `oats schedule host install --max-concurrent N` (or
|
|
364
|
+
`--max-concurrent default`) from the deployment, or with `--dir <deployment>`.
|
|
365
|
+
An absent value still means five. Set caps through the CLI; do not edit the
|
|
366
|
+
registry by hand.
|
|
338
367
|
|
|
339
368
|
`oats schedule list --json` answers:
|
|
340
369
|
|
package/docs/servers.md
CHANGED
|
@@ -298,7 +298,9 @@ the host's own facts from its `status --json`: `identity`,
|
|
|
298
298
|
`identityAddress`, `teams`, `startedAt`, `createdAt`, `model`,
|
|
299
299
|
`runtimeState`, `parentInstance`, `siblingInstance`, `relation`,
|
|
300
300
|
`relativeTo` and `spawnOrigin`. A fact the host does not supply is `null`
|
|
301
|
-
(an older host, or a saved route the host no longer lists). A
|
|
301
|
+
(an older host, or a saved route the host no longer lists). A row also
|
|
302
|
+
carries `waitingOnYou` (needs input) when the host's kernel reports it, and
|
|
303
|
+
only then: an older host's row has no such key. A removed or edited registration keeps
|
|
302
304
|
its group from the saved routes. State is pulled on every call within
|
|
303
305
|
`--per-target` (default 20 s) of a total `--budget` (default 45 s); a group
|
|
304
306
|
not reached is reported with `E_ROSTER_BUDGET`. `--server <id>` narrows it.
|
package/docs/workspaces.md
CHANGED
|
@@ -54,7 +54,7 @@ members: # repo refs, NO @revision (E_WORKSPAC
|
|
|
54
54
|
- git:github.com/acme/tools # a member that ALSO publishes a package (see below)
|
|
55
55
|
|
|
56
56
|
packages: # the ONLY versioned things
|
|
57
|
-
oats.framework: v1.6.
|
|
57
|
+
oats.framework: v1.6.1 # bare version → resolves through the official catalog
|
|
58
58
|
oats.okf: v4.1.1
|
|
59
59
|
acme.tools: git:github.com/acme/tools@v0.4.0 # outside the catalog → git:<repo>@<tag|OID>; still a package
|
|
60
60
|
|
package/lib/core.mjs
CHANGED
|
@@ -2668,6 +2668,12 @@ export async function spawnInstanceAsync(root, agent, o = {}) {
|
|
|
2668
2668
|
}
|
|
2669
2669
|
return step.value;
|
|
2670
2670
|
}
|
|
2671
|
+
/** The producer sets this only when dispatched effects or their compensation
|
|
2672
|
+
* cannot be confirmed. A diagnostic's wording or home path is not evidence. */
|
|
2673
|
+
function unconfirmedSpawn(error) {
|
|
2674
|
+
error.details = { ...error.details, unconfirmed: true };
|
|
2675
|
+
return error;
|
|
2676
|
+
}
|
|
2671
2677
|
function* spawnBody(root, agent, o = {}) {
|
|
2672
2678
|
if (!o.prepared) throw localMissingForSpawn(agent);
|
|
2673
2679
|
const deliver = (r) => r;
|
|
@@ -2756,7 +2762,7 @@ function* spawnBody(root, agent, o = {}) {
|
|
|
2756
2762
|
// Completion custody: the home exists from the first metadata write, but
|
|
2757
2763
|
// the launch/lineage/events that make it a finished spawn may not have
|
|
2758
2764
|
// happened (crash in the interval). Say which, never replay a half-spawn.
|
|
2759
|
-
if (prior.spawnCompleted !== true) throw Object.assign(oatsError("E_SPAWN_INCOMPLETE", `${prior.instance} was created for this key but its spawn did not complete (launch or lineage unfinished); inspect it with oats session inspect --home ${prior.home} — do not spawn again`), { instance: prior.instance, home: prior.home, launched: prior.launched === true ? "unknown" : false });
|
|
2765
|
+
if (prior.spawnCompleted !== true) throw unconfirmedSpawn(Object.assign(oatsError("E_SPAWN_INCOMPLETE", `${prior.instance} was created for this key but its spawn did not complete (launch or lineage unfinished); inspect it with oats session inspect --home ${prior.home} — do not spawn again`), { instance: prior.instance, home: prior.home, launched: prior.launched === true ? "unknown" : false }));
|
|
2760
2766
|
return { ...prior, replayed: true, launch: undefined, command: undefined, wake: prior.wake ?? { requested: null, saved: null, error: null } };
|
|
2761
2767
|
}
|
|
2762
2768
|
}
|
|
@@ -3218,7 +3224,7 @@ function* spawnBody(root, agent, o = {}) {
|
|
|
3218
3224
|
// name taken (M1). Nothing of a classic home exists at this point either.
|
|
3219
3225
|
const rollbackEmptyOrPreparedHome = (e) => {
|
|
3220
3226
|
try { rmSync(home, { recursive: true, force: true }); }
|
|
3221
|
-
catch (x) { e.message += ` — rollback INCOMPLETE, remove ${home} manually: ${x.message}`; }
|
|
3227
|
+
catch (x) { e.message += ` — rollback INCOMPLETE, remove ${home} manually: ${x.message}`; unconfirmedSpawn(e); }
|
|
3222
3228
|
return e;
|
|
3223
3229
|
};
|
|
3224
3230
|
// Workspace model: copy every resolved capability WHOLE into the new home
|
|
@@ -3386,9 +3392,10 @@ function* spawnBody(root, agent, o = {}) {
|
|
|
3386
3392
|
if (incomplete.length) {
|
|
3387
3393
|
// Nothing outside the home exists yet (no worktree, no hooks, no window), so
|
|
3388
3394
|
// removing the scaffold is the whole rollback.
|
|
3389
|
-
let removal = "";
|
|
3390
|
-
try { rmSync(home, { recursive: true, force: true }); } catch (e) { removal = ` — rollback INCOMPLETE, remove ${home} manually: ${e.message}`; }
|
|
3391
|
-
|
|
3395
|
+
let removal = "", removalFailed = false;
|
|
3396
|
+
try { rmSync(home, { recursive: true, force: true }); } catch (e) { removalFailed = true; removal = ` — rollback INCOMPLETE, remove ${home} manually: ${e.message}`; }
|
|
3397
|
+
const error = oatsError("E_COMPOSITION_INCOMPLETE", `the instance composition did not materialize completely:\n${incomplete.map((m) => ` ${m}`).join("\n")}${removal}`);
|
|
3398
|
+
throw removalFailed ? unconfirmedSpawn(error) : error;
|
|
3392
3399
|
}
|
|
3393
3400
|
|
|
3394
3401
|
// Work tree.
|
|
@@ -3433,7 +3440,8 @@ function* spawnBody(root, agent, o = {}) {
|
|
|
3433
3440
|
}
|
|
3434
3441
|
try { rmSync(home, { recursive: true, force: true }); } catch (e2) { incomplete.push(`instance home ${home}: ${e2.message}`); }
|
|
3435
3442
|
const note = incomplete.length ? ` — rollback INCOMPLETE — clean up manually: ${incomplete.join("; ")}` : "";
|
|
3436
|
-
|
|
3443
|
+
const error = new Error(`git worktree add/canonicalization failed: ${original}${note}`);
|
|
3444
|
+
throw incomplete.length ? unconfirmedSpawn(error) : error;
|
|
3437
3445
|
}
|
|
3438
3446
|
} else if (work === "directory") {
|
|
3439
3447
|
// An owned execution directory, not a link to the source or a fake Git repo.
|
|
@@ -3529,11 +3537,11 @@ function* spawnBody(root, agent, o = {}) {
|
|
|
3529
3537
|
outstandingGit.add("worktree");
|
|
3530
3538
|
if (branch) outstandingGit.add("branch");
|
|
3531
3539
|
}
|
|
3532
|
-
return quarantineInstanceHome({
|
|
3540
|
+
return { unconfirmed: true, note: quarantineInstanceHome({
|
|
3533
3541
|
home, instance, agent, soulDir: homeSoulTarget, soulId: preparedSoulId, incomplete, failed, outstandingHooks, outstandingGit,
|
|
3534
3542
|
repoAbs, work, branch, resolvedCfg, hookMeta: hookRes.meta || {},
|
|
3535
3543
|
launched: true, tmux: spawnTmux, directoryHome: homeReal, recordRetirementBaseline: true,
|
|
3536
|
-
});
|
|
3544
|
+
}) };
|
|
3537
3545
|
}
|
|
3538
3546
|
}
|
|
3539
3547
|
// Once hooks ran, directory execution may already hold authored results.
|
|
@@ -3555,7 +3563,7 @@ function* spawnBody(root, agent, o = {}) {
|
|
|
3555
3563
|
launched: false, directoryPreservation: true, directoryHome: homeReal,
|
|
3556
3564
|
recordRetirementBaseline: true,
|
|
3557
3565
|
});
|
|
3558
|
-
return `${note}${directoryRecoveries.length ? `; prior work recovery: ${directoryRecoveries.join(", ")}` : ""}
|
|
3566
|
+
return { unconfirmed: true, note: `${note}${directoryRecoveries.length ? `; prior work recovery: ${directoryRecoveries.join(", ")}` : ""}` };
|
|
3559
3567
|
};
|
|
3560
3568
|
const preserveDirectory = () => {
|
|
3561
3569
|
if (work !== "directory") return;
|
|
@@ -3639,7 +3647,7 @@ function* spawnBody(root, agent, o = {}) {
|
|
|
3639
3647
|
if (existsSync(home) && !incomplete.some((m) => m.startsWith("instance home"))) incomplete.push(`instance home ${home}: still present`);
|
|
3640
3648
|
note = incomplete.length ? ` — rollback INCOMPLETE, clean up manually: ${incomplete.join("; ")}` : " — spawn rolled back";
|
|
3641
3649
|
}
|
|
3642
|
-
return `${note}${directoryRecoveries.length ? `; directory work preserved at ${directoryRecoveries.join(", ")}` : ""}
|
|
3650
|
+
return { unconfirmed: incomplete.length > 0, note: `${note}${directoryRecoveries.length ? `; directory work preserved at ${directoryRecoveries.join(", ")}` : ""}` };
|
|
3643
3651
|
};
|
|
3644
3652
|
|
|
3645
3653
|
try {
|
|
@@ -3858,8 +3866,15 @@ ${task.trim() ? `\n## Task\n\n${task.trim()}\n` : "\nNo task was provided at spa
|
|
|
3858
3866
|
}
|
|
3859
3867
|
return deliver({ ...meta, ...(o.expectDecision !== undefined ? { replayed: false } : {}), launch: redactLaunchRecipe(recipe), command: redactLaunchCommand(cmdline), attach: `tmux attach -t ${session}`, warnings: spawnWarnings.length ? spawnWarnings : undefined });
|
|
3860
3868
|
} catch (error) {
|
|
3861
|
-
|
|
3862
|
-
|
|
3869
|
+
try {
|
|
3870
|
+
const compensation = compensateSpawn();
|
|
3871
|
+
error.message += compensation.note;
|
|
3872
|
+
if (compensation.unconfirmed) unconfirmedSpawn(error);
|
|
3873
|
+
} catch (cleanupError) {
|
|
3874
|
+
// An interrupted compensation pass cannot establish that effects ended.
|
|
3875
|
+
error.message += ` — rollback INCOMPLETE: cleanup could not be completed: ${cleanupError.message}`;
|
|
3876
|
+
unconfirmedSpawn(error);
|
|
3877
|
+
}
|
|
3863
3878
|
throw error;
|
|
3864
3879
|
}
|
|
3865
3880
|
}
|
package/lib/dir-lock.mjs
CHANGED
|
@@ -10,9 +10,9 @@ function pidAlive(pid) { try { process.kill(pid, 0); return true; } catch (e) {
|
|
|
10
10
|
const pause = (ms) => Atomics.wait(new Int32Array(new SharedArrayBuffer(4)), 0, 0, ms);
|
|
11
11
|
|
|
12
12
|
/** A mkdir lock that is never reclaimed by another process: an existing
|
|
13
|
-
* lock whose owner is unreadable or dead is refused
|
|
14
|
-
* remove, because the gap between mkdir and owner.json
|
|
15
|
-
* acquirer and a dead-owner reclaim races every other acquirer. The
|
|
13
|
+
* lock whose owner is unreadable or dead is refused after bounded waiting,
|
|
14
|
+
* with the directory to remove, because the gap between mkdir and owner.json
|
|
15
|
+
* belongs to a live acquirer and a dead-owner reclaim races every other acquirer. The
|
|
16
16
|
* holder removes its own lock in finally and on SIGINT/SIGTERM. A lock still
|
|
17
17
|
* held after `retryMs` is refused with `busy(why)`, the caller's own error. */
|
|
18
18
|
export function withDirLock(dir, what, fn, { retryMs = 0, busy } = {}) {
|
|
@@ -23,7 +23,10 @@ export function withDirLock(dir, what, fn, { retryMs = 0, busy } = {}) {
|
|
|
23
23
|
catch (e) {
|
|
24
24
|
if (e.code !== "EEXIST") throw e;
|
|
25
25
|
let owner; try { owner = JSON.parse(readFileSync(join(dir, "owner.json"), "utf8")); } catch { owner = undefined; }
|
|
26
|
-
|
|
26
|
+
// Missing/unreadable owners also occur between mkdir and publication,
|
|
27
|
+
// and while the holder removes its directory. Wait without touching it.
|
|
28
|
+
const remaining = deadline - Date.now();
|
|
29
|
+
if (remaining > 0) { pause(Math.min(50, remaining)); continue; }
|
|
27
30
|
throw busy(!owner ? "its owner is not readable yet or the file is missing" : pidAlive(owner.pid) ? `pid ${owner.pid} holds it` : `its owner pid ${owner.pid} is gone`);
|
|
28
31
|
}
|
|
29
32
|
}
|
package/lib/instance-events.mjs
CHANGED
|
@@ -26,9 +26,20 @@
|
|
|
26
26
|
* Producers write claims as `waiting` rows through `setWaiting` (the CLI's
|
|
27
27
|
* `oats instance waiting` and `oats instance attention`): data
|
|
28
28
|
* `{waitingOnYou, reason?, message?}`, appended only on a change. A claim is
|
|
29
|
-
* evidence for display, never authority: nothing in the kernel acts on it.
|
|
30
|
-
|
|
31
|
-
|
|
29
|
+
* evidence for display, never authority: nothing in the kernel acts on it.
|
|
30
|
+
*
|
|
31
|
+
* ONE ADDRESS PER HOME (awebai/oats#583). A home has several spellings when its
|
|
32
|
+
* deployment is reached through a symlink or its agents root is one: status
|
|
33
|
+
* addresses it lexically, a session carries the real path. Rows are keyed by
|
|
34
|
+
* the home string, so every exported function here resolves the home it is
|
|
35
|
+
* given to its REAL path first (`addressOf`): rows record it, readers compare
|
|
36
|
+
* against it, and both logs are found from any spelling. `admit` stays strict.
|
|
37
|
+
* That is storage and matching. An ANSWER (readEvents, setWaiting) names the
|
|
38
|
+
* home the way the caller did: `admit` has proved every returned row is this
|
|
39
|
+
* one home's, so its `home` is the address that was asked about, and a consumer
|
|
40
|
+
* that checks "the rows' home is the home I asked for" keeps holding. */
|
|
41
|
+
import { appendFileSync, closeSync, constants as fsConstants, existsSync, fstatSync, lstatSync, mkdirSync, openSync, readFileSync, readSync, realpathSync } from "node:fs";
|
|
42
|
+
import { basename, dirname, isAbsolute, join, resolve } from "node:path";
|
|
32
43
|
|
|
33
44
|
export const EVENTS_API = 2;
|
|
34
45
|
const HOME_LOG = ".oats-events.jsonl";
|
|
@@ -61,36 +72,104 @@ export function validWaitingMessage(m) {
|
|
|
61
72
|
// closed set; anything else reads as null, and the claim still counts.
|
|
62
73
|
/** A stored reason outside the closed set (a hand-edited or foreign row) reads as null. */
|
|
63
74
|
const readReason = (r) => (WAITING_REASONS.includes(r) ? r : null);
|
|
75
|
+
/** An event time: a string that parses as a date. The writer refuses anything else. */
|
|
76
|
+
const validTime = (at) => typeof at === "string" && Number.isFinite(Date.parse(at));
|
|
64
77
|
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
78
|
+
/** The READ rule for one claim `{since, producer, reason, message}`, the same for a
|
|
79
|
+
* row of this machine's log and for a claim another kernel reports (a remote roster
|
|
80
|
+
* row): `since` is a valid date and `producer` is `kernel` or a producer id, or there
|
|
81
|
+
* is no claim (null); a reason outside the closed set and a message that fails
|
|
82
|
+
* validWaitingMessage read as null, and the claim still counts. Anything that is not
|
|
83
|
+
* such an object is null. */
|
|
84
|
+
export function readWaitingClaim(claim) {
|
|
85
|
+
if (!claim || typeof claim !== "object" || Array.isArray(claim)) return null;
|
|
86
|
+
const { since, producer } = claim;
|
|
87
|
+
if (!validTime(since)) return null;
|
|
88
|
+
// A producer the writer would refuse (a hand-edited log) is never shown as one.
|
|
89
|
+
if (producer !== "kernel" && !(typeof producer === "string" && WAITING_PRODUCER_RE.test(producer))) return null;
|
|
90
|
+
return { since, producer, reason: readReason(claim.reason), message: validWaitingMessage(claim.message) ? claim.message : null };
|
|
71
91
|
}
|
|
72
92
|
|
|
73
|
-
/**
|
|
74
|
-
|
|
75
|
-
|
|
93
|
+
/** realpath of `p`, or, when `p` is gone, the realpath of its nearest existing
|
|
94
|
+
* ancestor with the rest re-appended (retirement writes its rows after the home
|
|
95
|
+
* is removed). */
|
|
96
|
+
function realOrNearest(p) {
|
|
97
|
+
try { return realpathSync(p); } catch { /* absent: resolve what exists */ }
|
|
98
|
+
let d = resolve(p); const tail = [];
|
|
99
|
+
while (!existsSync(d) && dirname(d) !== d) { tail.unshift(basename(d)); d = dirname(d); }
|
|
100
|
+
try { return join(realpathSync(d), ...tail); } catch { return resolve(p); }
|
|
101
|
+
}
|
|
102
|
+
function metaOf(home) {
|
|
103
|
+
try { const m = JSON.parse(readFileSync(join(home, "instance.json"), "utf8")); return m && typeof m === "object" ? m : null; } catch { return null; }
|
|
104
|
+
}
|
|
105
|
+
const incarnationIn = (meta) => (typeof meta?.createdAt === "string" ? meta.createdAt : null);
|
|
106
|
+
|
|
107
|
+
/** <deployment>/.agents/events/<agent>--<instance>.jsonl: deployment-private state
|
|
108
|
+
* beside installed capabilities and schedules; survives the home's removal.
|
|
109
|
+
*
|
|
110
|
+
* The DEPLOYMENT decides where it is, not the real home string: under a symlinked
|
|
111
|
+
* agents root the real home's ancestors are outside the deployment. The rule:
|
|
112
|
+
* 1. the deployment the spawn recorded (instance.json `workspace.deployment`), used
|
|
113
|
+
* ONLY when it verifies: the real path of its agents/<agent>/instances/<instance>
|
|
114
|
+
* is this home, with agent and instance taken from the home's path, never from
|
|
115
|
+
* instance.json;
|
|
116
|
+
* 2. otherwise the fourth ancestor of the home AS THE CALLER SPELLED IT, which is how
|
|
117
|
+
* the kernel addresses a home with no record, or one already removed (retirement).
|
|
118
|
+
* The result is resolved, so every spelling of the home gives one path.
|
|
119
|
+
*
|
|
120
|
+
* Two limits. A home with no recorded deployment (or already removed) that is named
|
|
121
|
+
* by its REAL path under a symlinked agents root still derives a directory outside
|
|
122
|
+
* the deployment; no kernel caller does that. And instance.json is home content: a
|
|
123
|
+
* record edited to name another directory that links back to this home moves the log
|
|
124
|
+
* there. That is the class the home log is already in (a path the agent can write,
|
|
125
|
+
* which the kernel appends to under the same uid), so nothing stricter is checked. */
|
|
126
|
+
function workspaceLogPath(given, home, meta) {
|
|
127
|
+
const instance = basename(home), agent = basename(dirname(dirname(home))); // <root>/<agent>/instances/<instance>
|
|
128
|
+
const recorded = meta?.workspace?.deployment;
|
|
129
|
+
const verified = typeof recorded === "string" && isAbsolute(recorded) && realOrNearest(join(recorded, "agents", agent, "instances", instance)) === home;
|
|
130
|
+
const deployment = verified ? recorded : dirname(dirname(dirname(dirname(resolve(given)))));
|
|
131
|
+
return join(realOrNearest(deployment), ".agents", "events", `${agent}--${instance}.jsonl`);
|
|
76
132
|
}
|
|
77
133
|
|
|
134
|
+
/** The one address of the home a caller named, in any spelling: the real home, its
|
|
135
|
+
* instance name and incarnation, its two logs, and `asked`, the caller's spelling. */
|
|
136
|
+
function addressOf(given) {
|
|
137
|
+
const home = realOrNearest(given);
|
|
138
|
+
const meta = metaOf(home);
|
|
139
|
+
let workspaceLog;
|
|
140
|
+
return {
|
|
141
|
+
home, asked: given, instance: basename(home), incarnation: incarnationIn(meta), homeLog: join(home, HOME_LOG),
|
|
142
|
+
get workspaceLog() { return (workspaceLog ??= workspaceLogPath(given, home, meta)); },
|
|
143
|
+
};
|
|
144
|
+
}
|
|
145
|
+
|
|
146
|
+
/** THE OUTPUT EDGE, and the only one: canonical inside, the caller's spelling outside.
|
|
147
|
+
* Whatever this module ANSWERS about an address (an events read, each row in it, a
|
|
148
|
+
* waiting answer) passes through here, which names the home as the caller spelled
|
|
149
|
+
* it. Rows on disk and every comparison keep the real path. A new answer that
|
|
150
|
+
* carries a `home` must go through this too, never write `a.home` out itself. */
|
|
151
|
+
const answered = (a, body) => ({ ...body, home: a.asked });
|
|
152
|
+
|
|
153
|
+
/** The incarnation of the home at `home`: its instance.json `createdAt`, or null. */
|
|
154
|
+
export function incarnationOf(home) { return incarnationIn(metaOf(home)); }
|
|
155
|
+
|
|
78
156
|
/** Append one typed event. Never throws into the caller's action: an event is
|
|
79
157
|
* evidence, not authority — a failed write is reported in the return value. */
|
|
80
|
-
export function appendEvent(home, event, {
|
|
158
|
+
export function appendEvent(home, event, options) { return append(addressOf(home), event, options); }
|
|
159
|
+
function append(a, event, { workspaceOnly = false, incarnation, at } = {}) {
|
|
81
160
|
if (!EVENT_KINDS.includes(event.kind)) return { ok: false, reason: `unknown event kind ${event.kind}` };
|
|
82
161
|
// `at`: the time the fact became true, when it is recorded later (a start
|
|
83
162
|
// boundary reconciled from its receipt); readers order rows by it.
|
|
84
|
-
if (at !== undefined &&
|
|
85
|
-
const row = { eventsApi: EVENTS_API, at: at ?? new Date().toISOString(), instance:
|
|
163
|
+
if (at !== undefined && !validTime(at)) return { ok: false, reason: `invalid event time ${at}` };
|
|
164
|
+
const row = { eventsApi: EVENTS_API, at: at ?? new Date().toISOString(), instance: a.instance, home: a.home, incarnation: incarnation === undefined ? a.incarnation : incarnation, producer: event.producer || "kernel", kind: event.kind, ...(event.data !== undefined ? { data: event.data } : {}) };
|
|
86
165
|
const line = JSON.stringify(row) + "\n";
|
|
87
166
|
const results = [];
|
|
88
167
|
// Retirement fingerprints the home against its baseline and preserves any
|
|
89
168
|
// changed bytes as "unknown work": events written DURING retirement go only
|
|
90
169
|
// to the workspace log, never into the home being inspected.
|
|
91
|
-
for (const path of [...(workspaceOnly ? [] : [
|
|
170
|
+
for (const path of [...(workspaceOnly ? [] : [a.homeLog]), a.workspaceLog]) {
|
|
92
171
|
try {
|
|
93
|
-
if (path.
|
|
172
|
+
if (path === a.homeLog && !existsSync(a.home)) { results.push({ path, ok: false, reason: "home absent" }); continue; }
|
|
94
173
|
mkdirSync(dirname(path), { recursive: true });
|
|
95
174
|
appendFileSync(path, line);
|
|
96
175
|
results.push({ path, ok: true });
|
|
@@ -174,11 +253,12 @@ function claimsOf(rows, incarnation) {
|
|
|
174
253
|
for (const r of current.slice(boundary + 1)) {
|
|
175
254
|
if (!r.data || typeof r.data.waitingOnYou !== "boolean") continue;
|
|
176
255
|
const p = r.producer ?? "kernel";
|
|
177
|
-
// A
|
|
178
|
-
|
|
256
|
+
// A row the read rule refuses (a hand-edited log: its producer, its time) says nothing.
|
|
257
|
+
const claim = readWaitingClaim({ since: r.at, producer: p, reason: r.data.reason, message: r.data.message });
|
|
258
|
+
if (!claim) continue;
|
|
179
259
|
claims.set(p, r.data.waitingOnYou
|
|
180
|
-
? { producer: p, waiting: true, since:
|
|
181
|
-
: { producer: p, waiting: false, since:
|
|
260
|
+
? { producer: p, waiting: true, since: claim.since, reason: claim.reason, message: claim.message }
|
|
261
|
+
: { producer: p, waiting: false, since: claim.since, reason: null, message: null });
|
|
182
262
|
}
|
|
183
263
|
return claims;
|
|
184
264
|
}
|
|
@@ -186,19 +266,20 @@ const positiveOf = (claims) => [...claims.values()].filter((c) => c.waiting).sor
|
|
|
186
266
|
const waitingShape = (c) => (c ? { since: c.since, producer: c.producer, reason: c.reason, message: c.message } : null);
|
|
187
267
|
|
|
188
268
|
/** Events for one instance ADDRESS, newest last, from both logs, with a bounded
|
|
189
|
-
* window; integrity is reported independently of the selected rows.
|
|
190
|
-
|
|
191
|
-
|
|
192
|
-
const
|
|
193
|
-
const
|
|
269
|
+
* window; integrity is reported independently of the selected rows. The answer's
|
|
270
|
+
* `home`, and each returned row's, is the home as the caller spelled it. */
|
|
271
|
+
export function readEvents(given, { limit = 200, since = null } = {}) {
|
|
272
|
+
const address = addressOf(given);
|
|
273
|
+
const { home, instance, incarnation, homeLog, workspaceLog } = address;
|
|
274
|
+
const a = readLog(homeLog, "home"), b = readLog(workspaceLog, "workspace");
|
|
194
275
|
const { rows, foreign } = admit(home, [a, b]);
|
|
195
276
|
const claims = claimsOf(rows, incarnation);
|
|
196
277
|
const waitingClaims = [...claims.values()];
|
|
197
278
|
const positive = positiveOf(claims);
|
|
198
279
|
const filtered = since ? rows.filter((r) => r.at > since) : rows;
|
|
199
|
-
const window = filtered.slice(-limit);
|
|
280
|
+
const window = filtered.slice(-limit).map((r) => answered(address, r));
|
|
200
281
|
const last = window.at(-1) ?? null;
|
|
201
|
-
return {
|
|
282
|
+
return answered(address, {
|
|
202
283
|
eventsApi: EVENTS_API, instance, home, incarnation, count: filtered.length, returned: window.length,
|
|
203
284
|
truncated: filtered.length > window.length || a.source.status === "tail" || b.source.status === "tail",
|
|
204
285
|
integrity: { unreadableRows: a.unreadable + b.unreadable, foreignRows: foreign, sources: [a.source, b.source] },
|
|
@@ -212,7 +293,7 @@ export function readEvents(home, { limit = 200, since = null } = {}) {
|
|
|
212
293
|
"a claim older than the incarnation's latest kernel launched, restarted or stopped row belongs to an ended session and is not counted",
|
|
213
294
|
"integrity counts torn and foreign rows and names each source's status; truncated is true when the window cut rows or a source was read as a tail",
|
|
214
295
|
],
|
|
215
|
-
};
|
|
296
|
+
});
|
|
216
297
|
}
|
|
217
298
|
|
|
218
299
|
/** The instance's live waiting state, `{since, producer, reason, message}` or
|
|
@@ -221,11 +302,11 @@ export function readEvents(home, { limit = 200, since = null } = {}) {
|
|
|
221
302
|
* (the workspace log only when the home log is absent): every row a producer
|
|
222
303
|
* or a session boundary writes goes to both. */
|
|
223
304
|
export function liveWaiting(home) {
|
|
224
|
-
const
|
|
225
|
-
if (incarnation === null) return null;
|
|
226
|
-
let log = readLog(
|
|
227
|
-
if (log.source.status === "absent") log = readLog(
|
|
228
|
-
return waitingShape(positiveOf(claimsOf(admit(home, [log]).rows, incarnation)));
|
|
305
|
+
const a = addressOf(home);
|
|
306
|
+
if (a.incarnation === null) return null;
|
|
307
|
+
let log = readLog(a.homeLog, "home");
|
|
308
|
+
if (log.source.status === "absent") log = readLog(a.workspaceLog, "workspace");
|
|
309
|
+
return waitingShape(positiveOf(claimsOf(admit(a.home, [log]).rows, a.incarnation)));
|
|
229
310
|
}
|
|
230
311
|
|
|
231
312
|
const waitingError = (code, message) => Object.assign(new Error(message), { code });
|
|
@@ -238,8 +319,9 @@ const waitingError = (code, message) => Object.assign(new Error(message), { code
|
|
|
238
319
|
* other log; success means both logs took the row. Validates its input
|
|
239
320
|
* (E_BAD_ARGS); a home without a readable instance.json is
|
|
240
321
|
* E_SESSION_UNKNOWN; a write either log refused is E_EVENTS_FAILED (a retry
|
|
241
|
-
* repairs it).
|
|
242
|
-
|
|
322
|
+
* repairs it). The answer's `home` is the home as the caller spelled it; the row
|
|
323
|
+
* records the real path. Evidence, not authority: nothing else is touched. */
|
|
324
|
+
export function setWaiting(given, { producer, waiting, reason, message } = {}) {
|
|
243
325
|
if (typeof producer !== "string" || !WAITING_PRODUCER_RE.test(producer)) throw waitingError("E_BAD_ARGS", `--producer must match ${WAITING_PRODUCER_RE.source}`);
|
|
244
326
|
if (producer === "kernel") throw waitingError("E_BAD_ARGS", "--producer kernel is reserved for the kernel's own events");
|
|
245
327
|
if (typeof waiting !== "boolean") throw waitingError("E_BAD_ARGS", "waiting must be set or clear");
|
|
@@ -250,15 +332,16 @@ export function setWaiting(home, { producer, waiting, reason, message } = {}) {
|
|
|
250
332
|
if (reason !== undefined) throw waitingError("E_BAD_ARGS", "--reason is for set, not clear");
|
|
251
333
|
if (message !== undefined) throw waitingError("E_BAD_ARGS", "--message is for set, not clear");
|
|
252
334
|
}
|
|
253
|
-
const
|
|
254
|
-
|
|
255
|
-
|
|
335
|
+
const a = addressOf(given);
|
|
336
|
+
const { home, incarnation } = a;
|
|
337
|
+
if (incarnation === null) throw waitingError("E_SESSION_UNKNOWN", `${given} is not an instance home (no readable instance.json)`);
|
|
338
|
+
const live = [readLog(a.homeLog, "home"), readLog(a.workspaceLog, "workspace")]
|
|
256
339
|
.map((log) => { const c = claimsOf(admit(home, [log]).rows, incarnation).get(producer); return c?.waiting ? c : null; });
|
|
257
340
|
const agrees = (c) => (waiting ? !!c && c.reason === reason && c.message === (message ?? null) : !c);
|
|
258
|
-
const answer = (changed, claim) => ({ eventsApi: EVENTS_API, instance:
|
|
341
|
+
const answer = (changed, claim) => answered(a, { eventsApi: EVENTS_API, instance: a.instance, home, producer, changed, waitingOnYou: waitingShape(claim) });
|
|
259
342
|
if (live.every(agrees)) return answer(false, waiting ? live[0] : null);
|
|
260
343
|
const data = waiting ? { waitingOnYou: true, reason, ...(message !== undefined ? { message } : {}) } : { waitingOnYou: false };
|
|
261
|
-
const res =
|
|
344
|
+
const res = append(a, { producer, kind: "waiting", data }, { incarnation });
|
|
262
345
|
const failed = (res.results || []).filter((r) => !r.ok);
|
|
263
346
|
if (!res.ok || failed.length) throw waitingError("E_EVENTS_FAILED", `could not record the waiting claim in ${failed.map((r) => `${r.path} (${r.reason})`).join(", ") || res.reason}; retry to complete it`);
|
|
264
347
|
return answer(true, waiting ? { producer, since: res.row.at, reason, message: message ?? null } : null);
|
|
@@ -272,22 +355,23 @@ export function setWaiting(home, { producer, waiting, reason, message } = {}) {
|
|
|
272
355
|
* same data), and only when neither has it is the row made, still at the
|
|
273
356
|
* launch time, so a claim the new session made since is kept. `ok` is true
|
|
274
357
|
* only when every log holds the row; evidence, never authority. */
|
|
275
|
-
export function recordStartBoundary(
|
|
276
|
-
const
|
|
277
|
-
const instance =
|
|
358
|
+
export function recordStartBoundary(given, { startId, startedAt, ...data }) {
|
|
359
|
+
const a = addressOf(given);
|
|
360
|
+
const { home, instance } = a;
|
|
361
|
+
const paths = [a.homeLog, a.workspaceLog];
|
|
278
362
|
const isIt = (r) => r.instance === instance && r.home === home && (r.producer ?? "kernel") === "kernel" && r.kind === "launched" && r.data?.startId === startId;
|
|
279
363
|
const found = paths.map((path) => readLog(path, "log").rows.find(isIt) ?? null);
|
|
280
364
|
if (found.every(Boolean)) return { ok: true, existed: true };
|
|
281
365
|
const existing = found.find(Boolean);
|
|
282
366
|
if (!existing) {
|
|
283
|
-
const res =
|
|
367
|
+
const res = append(a, { kind: "launched", data: { ...data, startId } }, { at: startedAt });
|
|
284
368
|
return { ...res, ok: res.ok && (res.results || []).every((r) => r.ok) };
|
|
285
369
|
}
|
|
286
370
|
// The row exactly as the other log holds it: same time, incarnation and data.
|
|
287
371
|
const line = JSON.stringify(existing) + "\n";
|
|
288
372
|
const results = paths.filter((_, i) => !found[i]).map((path) => {
|
|
289
373
|
try {
|
|
290
|
-
if (path.
|
|
374
|
+
if (path === a.homeLog && !existsSync(home)) return { path, ok: false, reason: "home absent" };
|
|
291
375
|
mkdirSync(dirname(path), { recursive: true });
|
|
292
376
|
appendFileSync(path, line);
|
|
293
377
|
return { path, ok: true };
|
|
@@ -7,9 +7,10 @@ import { signalGroup } from "./process-group.mjs";
|
|
|
7
7
|
|
|
8
8
|
const { file, args, timeout, graceMs, maxBuffer } = JSON.parse(readFileSync(0, "utf8"));
|
|
9
9
|
const result = await new Promise((resolve) => {
|
|
10
|
-
let child, status = null, signal = null, error;
|
|
11
|
-
let exited = false, closed = false, stopping = false, hardEnd = false, settled = false;
|
|
10
|
+
let child, status = null, signal = null, error, interrupted = false;
|
|
11
|
+
let exited = false, closed = false, stopping = false, hardEnd = false, settled = false, termSent = false;
|
|
12
12
|
let deadline, escalation, probe, groupEmpty = false;
|
|
13
|
+
const signalHandlers = new Map();
|
|
13
14
|
const chunks = { stdout: [], stderr: [] }, sizes = { stdout: 0, stderr: 0 };
|
|
14
15
|
const groupAlive = () => {
|
|
15
16
|
if (groupEmpty || !child?.pid) return false;
|
|
@@ -21,6 +22,11 @@ const result = await new Promise((resolve) => {
|
|
|
21
22
|
if (process.platform === "win32") { if (!exited) child.kill(sig); }
|
|
22
23
|
else if (groupAlive()) signalGroup(child, sig);
|
|
23
24
|
};
|
|
25
|
+
const requestTerm = () => {
|
|
26
|
+
if (termSent || !child?.pid) return;
|
|
27
|
+
termSent = true;
|
|
28
|
+
send("SIGTERM");
|
|
29
|
+
};
|
|
24
30
|
const finish = () => {
|
|
25
31
|
if (settled || (!closed && !hardEnd) || (!exited && child?.pid)) return;
|
|
26
32
|
// A leader can close while a pipe-free descendant ignores TERM. Keep the
|
|
@@ -28,31 +34,55 @@ const result = await new Promise((resolve) => {
|
|
|
28
34
|
if (stopping && !hardEnd && groupAlive()) return;
|
|
29
35
|
settled = true;
|
|
30
36
|
clearTimeout(deadline); clearTimeout(escalation); clearInterval(probe);
|
|
31
|
-
|
|
37
|
+
for (const [sig, handler] of signalHandlers) process.off(sig, handler);
|
|
38
|
+
resolve({ status, signal, stdout: Buffer.concat(chunks.stdout).toString("utf8"), stderr: Buffer.concat(chunks.stderr).toString("utf8"), ...(error ? { error } : {}), ...(interrupted ? { interrupted: true } : {}) });
|
|
39
|
+
};
|
|
40
|
+
const watchGroup = () => {
|
|
41
|
+
probe ??= setInterval(() => { groupAlive(); finish(); }, Math.min(50, graceMs));
|
|
32
42
|
};
|
|
33
43
|
const stop = (cause) => {
|
|
34
44
|
if (stopping || settled) return;
|
|
35
|
-
stopping = true; error
|
|
36
|
-
|
|
45
|
+
stopping = true; error ??= cause;
|
|
46
|
+
requestTerm();
|
|
37
47
|
escalation = setTimeout(() => {
|
|
38
48
|
send("SIGKILL"); hardEnd = true;
|
|
39
49
|
// A detached descendant can escape the group yet retain inherited pipes.
|
|
40
50
|
// Its pipes must not prevent returning after the direct child's exit.
|
|
41
|
-
child
|
|
51
|
+
child?.stdout?.destroy(); child?.stderr?.destroy();
|
|
42
52
|
finish();
|
|
43
53
|
}, graceMs);
|
|
44
|
-
|
|
54
|
+
watchGroup();
|
|
45
55
|
};
|
|
56
|
+
// Catchable supervisor shutdown owns the same cleanup as timeout/overflow.
|
|
57
|
+
// Keep handlers throughout cleanup: repeated signals cannot bypass it.
|
|
58
|
+
// SIGKILL/OOM cannot run this path; a missing receipt stays unconfirmed.
|
|
59
|
+
for (const sig of ["SIGINT", "SIGTERM", "SIGHUP"]) {
|
|
60
|
+
const handler = () => {
|
|
61
|
+
// Preserve interruption even if overflow/timeout already began cleanup;
|
|
62
|
+
// the first error alone cannot describe both independent observations.
|
|
63
|
+
interrupted = true;
|
|
64
|
+
stop({ code: "EINTR", message: `schedule command supervisor interrupted by ${sig}` });
|
|
65
|
+
};
|
|
66
|
+
signalHandlers.set(sig, handler);
|
|
67
|
+
process.on(sig, handler);
|
|
68
|
+
}
|
|
46
69
|
try { child = spawn(file, args, { detached: process.platform !== "win32", stdio: ["ignore", "pipe", "pipe"] }); }
|
|
47
|
-
catch (e) { error
|
|
70
|
+
catch (e) { error ??= { code: e.code, message: e.message }; closed = true; finish(); return; }
|
|
48
71
|
for (const stream of ["stdout", "stderr"]) child[stream].on("data", (chunk) => {
|
|
49
72
|
const room = maxBuffer - sizes[stream];
|
|
50
73
|
if (room > 0) { const kept = chunk.subarray(0, room); chunks[stream].push(kept); sizes[stream] += kept.length; }
|
|
51
74
|
if (chunk.length > room) stop({ code: "ENOBUFS", message: `schedule command ${stream} exceeded ${maxBuffer} bytes` });
|
|
52
75
|
});
|
|
53
76
|
child.once("error", (e) => { error ??= { code: e.code, message: e.message }; });
|
|
54
|
-
child.once("exit", (code, sig) => {
|
|
77
|
+
child.once("exit", (code, sig) => {
|
|
78
|
+
exited = true; status = code; signal = sig;
|
|
79
|
+
// Do not wait for pipes or the deadline: an escaped session may hold pipes
|
|
80
|
+
// after this group empties. Once observed empty it is never signalled again.
|
|
81
|
+
if (groupAlive()) watchGroup();
|
|
82
|
+
finish();
|
|
83
|
+
});
|
|
55
84
|
child.once("close", () => { closed = true; finish(); });
|
|
85
|
+
if (stopping) requestTerm(); // Shutdown may have begun while spawn returned.
|
|
56
86
|
deadline = setTimeout(() => stop({ code: "ETIMEDOUT", message: `schedule command exceeded ${timeout} ms` }), timeout);
|
|
57
87
|
});
|
|
58
88
|
process.stdout.write(JSON.stringify(result));
|