@staix/agent-hub 0.12.5 → 0.12.7
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +21 -0
- package/docs/cooperbench.md +5 -3
- package/docs/operations.md +13 -8
- package/docs/specs/2026-09-19-agent-hub-design.md +1 -1
- package/docs/verification/2026-10-03-0.12.6.md +171 -0
- package/docs/verification/2026-10-03-0.12.7.md +58 -0
- package/package.json +1 -1
- package/plugins/agent-hub/.claude-plugin/plugin.json +1 -1
- package/plugins/agent-hub/server.js +632 -956
- package/src/adapters/acp.ts +7 -5
- package/src/adapters/codex-appserver.ts +8 -6
- package/src/adapters/pi.ts +9 -4
- package/src/hub/child-process.ts +26 -8
- package/src/hub/daemon.ts +7 -1
- package/src/hub/project.ts +5 -4
- package/src/hub/tasks.ts +23 -4
package/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,27 @@
|
|
|
2
2
|
|
|
3
3
|
Issue and pull request numbers in the entries for 0.7.7 and earlier refer to the previous repository, archived on 2026-09-30 when this repository's history was rewritten; the one exception is the open smoke-check issue, formerly #12, which moved here as #1. Numbers in newer entries refer to this repository.
|
|
4
4
|
|
|
5
|
+
## 0.12.7
|
|
6
|
+
|
|
7
|
+
- CooperBench runner and recovery (#120): from the moment the runner may change anything, before it locks its first input, `restoration.json` says `restored: false` and names the runner, until the runner writes its outcome. A runner killed after that point sends the recovery in, and the recovery waits while the runner it names still runs. A run directory whose earlier run is not restored is refused first, before the runner reads its inputs or any mode (the refusal names `restore.ts`), and checked again right before the runner locks anything, so a runner does not record locked modes as the originals; two runners started on one directory at the same moment are not guarded against. A directory whose earlier invocation ended restored, with no records, can still be used. `restore.ts` trusts `restored: true` only when the restoration ledger agrees (a ledger from before 0.12.5, with no runner identity, still has its locks restored; when `restoration.json` says `restored: true` or names a runner, its trust entry is kept as it was then), so a stale file over a re-run that died no longer reports a run restored while its inputs stay locked; it checks the runners again right before it restores anything. With no ledger, nothing was locked and nothing is restored.
|
|
8
|
+
- CooperBench manifests v2 and the #106 ablation pin hub 0.12.7.
|
|
9
|
+
- Bun 1.4.2 in CI, the release workflow and the plugin bundle build. Bun 1.4.2 still throws for a real path with a backslash, so `realPath` stays, and still runs the next test inside an outer spawnSync's event loop after a timeout, so `check.sh` keeps its test timeout and watchdog. On 1.4.2, a model relay that is closed while it streams a response gets an error printed by Bun when it aborts that stream ("model relay closed"): log noise, not a failure (#121).
|
|
10
|
+
|
|
11
|
+
## 0.12.6
|
|
12
|
+
|
|
13
|
+
- Agents the hub spawns through ACP (Kimi) and Pi run in their own process group and are stopped as one, as the Codex app-server is since 0.12.5. A stop with or without a group drops the child's pipes, so a process it left cannot keep the hub alive. Once a launcher has exited, its group is followed until it is gone, and members still finishing their exit get a bound before the stop fails. A Codex start that fails reports its own error. The group stop finishes as soon as the leader has exited and a read shows nothing left. As for Codex since 0.12.5, a Kimi or Pi that crashed and left processes in its group (an MCP server, a tool command) is not restarted until they are gone: the error names their pids (#115).
|
|
14
|
+
- Split predictions (shadow only): busy counts as taking the task in question only while the peer is still in the turn that task started (delivered at once to the idle peer) or in which it claimed the task. A task queued, held or steered into a turn about something else is not taken, a turn that ended takes its hand-overs with it, and a routed peer busy when the record is taken is not available (#109, #115).
|
|
15
|
+
- Claude Code's version is read from the transcript again only when the transcript changed (#115).
|
|
16
|
+
- Tests: `scripts/check.sh` runs `bun test` with a 20 s timeout and under `scripts/hang-watch.sh`, which samples and stops a run past its bound. In Bun 1.3.14 a test timeout that fires while `Bun.spawnSync` runs can start the next test inside spawnSync's event loop, where another spawnSync then spins for good: the stacks of two local hangs under load show it, and the macOS CI hang of 0.12.5 matches them (minimal reproductions did not hang). The permission test no longer depends on runner speed (#115).
|
|
17
|
+
- CooperBench runner and ledger (#115):
|
|
18
|
+
- The runner never adds or removes an Orca registration, which needs the user's explicit authorization: every fixture of the selection is looked up read-only before anything is changed, a missing or ambiguous one refuses the run, and each arm checks again, before touching its fixture, that its identity is the one preflighted (#117, #118).
|
|
19
|
+
- A trust write that never landed is `not_written`, not `changed_concurrently`, also after a failed restore and in the recovery; its temp files are removed in the runner and the recovery (also one a runner left when it died in its own restore), one that cannot be removed keeps the trust entry open for the next recovery. A write the runner knows never landed is `not_written` on every path, and neither the runner nor the recovery then touches an entry; for a runner that died mid write, the recovery takes back only an entry exactly as the runner would have written it; one the user changed meanwhile is never taken back.
|
|
20
|
+
- The ledger shows normal shutdown errors, and reports records kept in `recovery/` as withheld attempts, not missing; a record caught in the middle of the recovery's move counts once, as withheld, and attempts a still-locked `runs/` hides are `unreadable`, not missing, in the per-arm summary too, where an arm with no record at all is listed with what it owes.
|
|
21
|
+
- Records carry the platform, and the runner refuses non-macOS.
|
|
22
|
+
- An unresolved process's program name is its executable's name as the kernel recorded it (`ps -o ucomm`), checked against its start time, never its arguments; it is omitted when it cannot be read.
|
|
23
|
+
- The end reason is one function, and shared helpers are not duplicated.
|
|
24
|
+
- CooperBench manifests v2 and the #106 ablation pin hub 0.12.6.
|
|
25
|
+
|
|
5
26
|
## 0.12.5
|
|
6
27
|
|
|
7
28
|
- The Codex app-server runs in its own process group and is stopped as one, with what it started in groups of its own (MCP servers, tool commands): after SIGTERM, to its group and to each recorded process that leads a group of its own, and a grace period in which what it starts is recorded, the tree is frozen, read again and killed, and the stop is done only when the table shows none of it (or, when no table can be read at all, when its own group is gone); a launcher that exited before the stop while its group still has members fails the stop, its group unsignalled. `codex` is a node launcher, and mid-turn the native app-server does not exit on SIGTERM: the SIGKILL that followed reached the launcher alone, leaving the app-server at work under init and the hub process alive after `ahub kill` reported it stopped. The CooperBench runner's teardown is verified (#113): every process an arm starts that the reads see is recorded with its pid, start time and the evidence that it is the arm's (a fixture name in an argv is never proof on its own), read again every 5 s while the agents work and while Claude ends its turn (one that detaches between two reads, outside the fixture and without it in its argv, is not seen); a process with a fixture in its argv or as its working directory that is not proved the arm's is never signalled and keeps the cleanup open; teardown pauses the agents, lets a completed arm's Claude end its turn (up to 30 s, by the transcript's `turn_duration` row proved on the probe turn), asks for the normal shutdown, reads the process table back in the C locale and in UTC (a start time is part of a process's identity), signals only re-read identities, and records `clean`, `clean_with_fallback` or `incomplete_or_unknown` apart from the end reason. Evidence is taken after it, and the read locks and protected inputs come off only after a complete cleanup (otherwise `scripts/benchmarks/restore.ts` does it later, once the runner, every recorded process and anything in the fixture are gone). Run records carry the completion, cleanup, restoration, stage times and a summary of the Codex skills `skills/list` reports (its answer itself is not kept); a record from before 0.12.5 is judged by its own cleanup and trust flags, as then, and its teardown is shown as not verified (0.12.3 and 0.12.4 recorded only that the shutdown steps reported success, with no process readback); `teardown.ts` is pinned with the runner sources; runner commands run in their own process groups.
|
package/docs/cooperbench.md
CHANGED
|
@@ -37,6 +37,8 @@ python3 scripts/benchmarks/runner.py prepare \
|
|
|
37
37
|
--upstream-root /tmp/agent-hub-cooperbench-upstream \
|
|
38
38
|
--output /private/tmp/ahub-0123-case0
|
|
39
39
|
chmod 700 /private/tmp/ahub-0123-case0
|
|
40
|
+
# Orca registrations need the user's explicit authorization (see below): the runner never adds or removes one
|
|
41
|
+
# (#117) and refuses the run while any selected fixture is not registered.
|
|
40
42
|
bun scripts/benchmarks/native.ts \
|
|
41
43
|
--run /private/tmp/ahub-0123-case0 \
|
|
42
44
|
--private-inputs /private/tmp/ahub-0123-private-inputs \
|
|
@@ -49,9 +51,9 @@ bun scripts/benchmarks/native.ts \
|
|
|
49
51
|
--cases 0
|
|
50
52
|
```
|
|
51
53
|
|
|
52
|
-
The runner registers each exact fixture root with
|
|
54
|
+
The native runner runs on macOS only and refuses anything else (each run record says `platform`): the arms run in Orca terminals, and on Linux a clock step moves the start times its teardown proves processes by. Before native execution, obtain explicit user authorization before changing Orca registrations; a benchmark request alone is not authorization. Only after that authorization, an operator manually registers each exact fixture root with `orca repo add --path '<fixture>'` and confirms `orca worktree list --repo id:<repo-id>` shows that exact path. The runner performs read-only exact-path repo/worktree lookup; it preflights every selected fixture before changing fixture files, input modes, cohort or attempt records. If either identity is missing or mismatched, it exits nonzero with setup guidance and leaves the prepared run available for retry after explicitly authorized registration. Each arm rechecks its identity before changing fixture files. Benchmark and agent workflows never add or remove Orca registrations. The runner uses the project CLI to start and stop each daemon, and requires exact worktree/cwd readback for native sessions. Claude starts through the canonical `ahub claude` guard and loads this checkout's candidate bundle through an exact session-only `--mcp-config` server (`server:agent-hub`); it does not promote or mutate the globally installed plugin. Codex uses the native app-server adapter and `workspace-write` sandbox. Each agent must execute a setup-only `head -c 1` probe against the exact protected file and produce only the denied marker before scored tasks begin. There is one native sandbox layer per agent.
|
|
53
55
|
|
|
54
|
-
Every run record, diff, PTY/provider trace and evaluation artifact stays in the mode-0700 private run directory. Setup errors, structured provider quota errors, hub budget pauses, unsettled deliveries and interruptions stay separate and unscored. A quota snapshot or numeric `429` alone is not a provider error. On exit the runner restores the exact input and read-lock modes, the sibling artifact modes and the scoped Claude trust flag (removing its own fresh
|
|
56
|
+
Every run record, diff, PTY/provider trace and evaluation artifact stays in the mode-0700 private run directory. Setup errors, structured provider quota errors, hub budget pauses, unsettled deliveries and interruptions stay separate and unscored. A quota snapshot or numeric `429` alone is not a provider error. On exit the runner restores the exact input and read-lock modes, the sibling artifact modes and the scoped Claude trust flag (removing its own fresh Claude trust entry), under the condition below. Teardown (issue #113) proves what it stops. Each process an arm starts is recorded with its pid and start time and the evidence that it is the arm's: the daemon by the pid in its state directory and an argv that serves this fixture (a later daemon there is a replacement, never adopted), Claude's launch chain by the arm's own session id as its `--session-id` (and the Orca terminal shell that runs it, working in the fixture), the Codex app-server as the daemon's child. What runs below them, and what is in a group a recorded process leads while that group is known to be the same one (its leader alive, or a recorded member still in it), is the arm's too; the table is read every 5 s while the agents work and during the completion wait, so a tool command's background job is recorded while its parent runs (one that detaches between two reads, outside the fixture and without it in its argv, is not seen). A fixture name in an argv is never proof on its own (Codex's need not name the fixture), and every recorded actor counts even when a read of the table fails. Teardown pauses every agent, reads the deliveries still owed, lets a completed arm's Claude end the turn it is in (up to 30 s and never past the 300 s limit (the bound is taken when the wait starts; one 250 ms poll can pass it), judged by the `turn_duration` row Claude Code writes after a turn, proved on the arm's probe turn; a stopped or timed-out arm does not wait; a completed arm's tree is hashed at the end of its active time and again after the teardown, and a write in between flags the attempt invalid), asks for the normal shutdown (Claude's terminal closed, `ahub kill`, the hub's project registration removed (`ahub projects remove`); the runner leaves Orca registrations untouched; a step of it that fails is recorded in `cleanup.normal.errors` and does not by itself fail the attempt, since the process readback decides, so a hub project registration left behind is listed there), and reads the process table back, in the C locale and in UTC (start times are read as identities, the same in every reader). What is still running after a settle period gets SIGTERM; what is left then is frozen (SIGSTOP), the table read again, and it is killed (SIGKILL), what was frozen and not killed being continued: a frozen process starts nothing, so that read sees all of it. Every signal goes only to an identity read again just before, by the group the table shows it in now; a group leader is signalled with its group. What the read after the freeze shows for the first time is frozen too and read again before anything is killed, and when such a read fails only what a STOP reached is killed, by the last read that showed it (a stopped process keeps its pid). Anything with the fixture in its argv or as its working directory that is not proved the arm's (a background job that escaped the reads, someone's shell) is never signalled and keeps the cleanup open; the record keeps its pid, start time, program name (the name of its executable as the kernel recorded it at exec, `ps -o ucomm`, and only while the pid has the start time recorded; never `comm`, which on macOS is the process's own argv[0] and can carry its arguments; omitted when it cannot be read) and working directory, never its arguments. The record says `clean`, `clean_with_fallback` or `incomplete_or_unknown` (`cleanup`, with the normal shutdown's errors, every signal sent, what is still running and what is unresolved), and `cleanup_complete` is true only for the first two. Evidence follows the teardown: the transcript prefix (`transcriptBytes`, `transcriptSha256`, with the session and time it was taken) and the patch. The sibling read locks come off and the protected inputs become readable again only when the cleanup is complete; otherwise both stay, the record goes to `recovery/` beside the run's restoration ledger (which lists every arm's recorded processes and the runner's own identity), the cohort stops, and `bun scripts/benchmarks/restore.ts --run <run>` (also `runner.py restore`, which calls it) puts the modes back and takes back the trust entry once the runner, every recorded process and anything with a fixture in its argv or as its working directory are gone, then moves the kept records into `runs/`, where grading reads them as unavailable. A ledger written before runner identities needs `--runner-exited`, the operator's statement that the runner is gone. Runner commands are bounded (a hung one is killed), and SIGHUP stops the run like SIGINT and SIGTERM. `restoration.json` says `restored: true` only when the protected inputs and every arm's sibling locks are back. From before the runner locks its first input until it writes that outcome, it says `restored: false` and names the runner (#120), so a runner that is killed sends the recovery in, and the recovery waits while that runner still runs. The runner refuses a run directory whose earlier run is not restored (its `restoration.json` or its ledger says so) first, before it reads its inputs or any mode, and again right before it locks anything, since locked modes would become the originals; `restore.ts` first. One runner per directory at a time: two started on one directory at the same moment are not guarded against. `restore.ts` trusts `restored: true` only when the restoration ledger agrees (a ledger from before 0.12.5 still has its locks restored; when `restoration.json` says `restored: true` or names a runner, its trust entry is kept as it was then), and checks the runners again right before it restores anything. The trust entry is taken back with them (by the recovery when the cleanup is incomplete), and kept when the user changed it meanwhile; when the runner's own write never landed (it stopped between recording the lease and the rename), the outcome is `not_written` and nothing is taken back; for a runner that died there, the recovery takes back only an entry exactly as the runner would have written it, and none while the runner's temp file is still there. Orca registrations are not removed during benchmark teardown. Legacy benchmark-created entries may be removed only after provenance confirms the exact fixture path, using `orca project setup-delete --setup <id>` after verifying the exact repo and fixture path; unrelated registrations and fixture/Git/evidence data remain untouched. The ledger reports a record kept in `recovery/` as an attempt with `withheld: true` (unavailable: its cleanup was incomplete, or its sibling read locks could not be put back) until `restore.ts` moves it, never as missing (a record in both places, a move that did not finish, counts once and withheld); while `runs/` is still locked (a cleanup is incomplete), the attempts it cannot read there are listed as `unreadable`, never as missing, and its teardown row carries the normal shutdown's errors (`normal_errors`, for example a hub project registration left behind). The recovery also removes the temp files a runner left (one that died mid trust write or mid restore, or could not remove its own), each a copy of `~/.claude.json`; a write the runner knew never landed stays `not_written` there too, and no entry is touched. Run records carry `completion`, `cleanup`, `restoration` and `stages` (completion wait, shutdown, settle, fallback with the final readback, evidence and restoration times), apart from the end reason: `end_reason_detail` keeps the original classification, and `end_flags` (an unverified model, modified metadata, a tree changed or unverified after the active time) makes any end other than an interruption, a provider quota error or a budget pause an infrastructure error beside it, which grading and the ledger report as `end_story`; a record with `teardown_errors` (a patch, transcript prefix, last capture, fixture metadata or event log that could not be taken, or sibling read locks or the trust entry that could not be put back, a concurrent change of the trust entry included), an incomplete cleanup or a trust entry not taken back is unavailable to grading and to the ledger alike (a record from before 0.12.5 is judged as it was then, by its own `cleanup_complete` and `trust_restored`, and the ledger shows its teardown with `verified: false`: 0.12.3 and 0.12.4 set `cleanup_complete` when the shutdown steps reported success (commands, terminal close, lock and trust restores), with no process readback, which is no proof that the processes were gone); runner commands, the `ps` reads included, run in process groups of their own, so a Ctrl-C reaches the runner, which stops in this order, and not the command it is running; the completion wait's tree hash never writes the agent's index. `teardown.ts` and the process table it reads (`src/hub/child-process.ts`) are pinned in `prepared.json`, and `teardown.ts` in `cohort.json` and `grade.json` too, with the other runner sources.
|
|
55
57
|
|
|
56
58
|
## Manifest v2: the turn-free arm
|
|
57
59
|
|
|
@@ -71,7 +73,7 @@ Pass `--repeat <n>` for the n-th repeat of a case (0 for the first): the arm ord
|
|
|
71
73
|
python3 scripts/benchmarks/ledger.py --run /private/tmp/ahub-0124-r1 [--run /private/tmp/ahub-0124-r2 ...]
|
|
72
74
|
```
|
|
73
75
|
|
|
74
|
-
It reads each run record, the Claude transcript it names and the fixture's git history, and writes `ledger.json` into the first run directory with a `units` table that names the unit and coverage of every measure. Give `--run` once per directory to pool the repeats of one plan; an attempt a directory's `cohort.json` planned that wrote no record is listed as missing. Per attempt: completion (the runner's end reason and its detail, such as `wall-timeout`, and whether every task has a done); setup and active time; first candidate, completion intents, integration and check times, and each agent's settlement read from its own record (the end of its last native turn after the last done, turns started by late messages included; the last done itself when it did not work after it), and how long the agents could still write after the active time ended (`stopped_s`); usage per agent in task (to its last done on the board) and over the whole attempt: Codex turns, assistant messages, token-usage updates whose running total grew past the total before the window, and the token growth; Claude assistant messages (unique message ids), turns and the tokens their usage records; Claude's main-loop requests by request id (side requests and retries are not in its transcript), and Codex's provider requests unknown (app-server 0.159 does not send its response ids); Codex turns after its done and what started them; late replies, steered or at the next turn; the hub's fact offers, acknowledgements, bytes offered and acknowledged, build times, hook start-up times and steer round trips by path, in the task window; capability readbacks; validity by the grader's gates, and for turn-free the treatment received; held-back messages (`quiet` events) apart from [FYI] messages (which include the final [FYI] the instructions ask for); stale notices; shadow split predictions with their traces; the hooks each agent's records show, by a label that keeps paths and arguments out, with Claude's hook durations and the hub's own timing of every facts hook call; and contributions. Summaries give medians over valid completed attempts and over the (case, repeat) pairs every arm completed validly, totals over the attempts whose tasks were handed out and to which the measure applies (a Codex measure in a Claude-only arm is not counted as unknown), with the number of attempts each was unknown for, and the reasons for the rest. A repeat given twice is refused, and `--plan pilot` (or `study`) lists every planned attempt that wrote no record, whole repeats included. `ledger.json` holds code fragments from the agents' writes and local paths: keep it with the private run data and never commit it; the summary is what a verification record quotes. In-task windows end at different points by arm (a turn-free integration step comes before the done, an advisory completed-change notice turn after it), so arms are compared on whole-attempt usage.
|
|
76
|
+
It reads each run record, the Claude transcript it names and the fixture's git history, and writes `ledger.json` into the first run directory with a `units` table that names the unit and coverage of every measure. Give `--run` once per directory to pool the repeats of one plan; an attempt a directory's `cohort.json` planned that wrote no record is listed as missing, and one a still-locked `runs/` hides as unreadable, in the summary per arm too (an arm with no record at all is listed with what it owes). Per attempt: completion (the runner's end reason and its detail, such as `wall-timeout`, and whether every task has a done); setup and active time; first candidate, completion intents, integration and check times, and each agent's settlement read from its own record (the end of its last native turn after the last done, turns started by late messages included; the last done itself when it did not work after it), and how long the agents could still write after the active time ended (`stopped_s`); usage per agent in task (to its last done on the board) and over the whole attempt: Codex turns, assistant messages, token-usage updates whose running total grew past the total before the window, and the token growth; Claude assistant messages (unique message ids), turns and the tokens their usage records; Claude's main-loop requests by request id (side requests and retries are not in its transcript), and Codex's provider requests unknown (app-server 0.159 does not send its response ids); Codex turns after its done and what started them; late replies, steered or at the next turn; the hub's fact offers, acknowledgements, bytes offered and acknowledged, build times, hook start-up times and steer round trips by path, in the task window; capability readbacks; validity by the grader's gates, and for turn-free the treatment received; held-back messages (`quiet` events) apart from [FYI] messages (which include the final [FYI] the instructions ask for); stale notices; shadow split predictions with their traces; the hooks each agent's records show, by a label that keeps paths and arguments out, with Claude's hook durations and the hub's own timing of every facts hook call; and contributions. Summaries give medians over valid completed attempts and over the (case, repeat) pairs every arm completed validly, totals over the attempts whose tasks were handed out and to which the measure applies (a Codex measure in a Claude-only arm is not counted as unknown), with the number of attempts each was unknown for, and the reasons for the rest. A repeat given twice is refused, and `--plan pilot` (or `study`) lists every planned attempt that wrote no record, whole repeats included. `ledger.json` holds code fragments from the agents' writes and local paths: keep it with the private run data and never commit it; the summary is what a verification record quotes. In-task windows end at different points by arm (a turn-free integration step comes before the done, an advisory completed-change notice turn after it), so arms are compared on whole-attempt usage.
|
|
75
77
|
|
|
76
78
|
Contributions are a heuristic for possible loss, never a certificate: per agent and file, the identifiers and changed fragments its applied writes introduced (only a Claude tool call with a successful result, or a completed Codex patch, counts; a Write replaces the agent's earlier contribution and is not credited with what other agents wrote; a delete removes it; when a file is moved, every agent's contributions to it are checked at its new path) that the final tree lacks, a fragment counting as present anywhere in the file. A same-name overwrite shows as a lost fragment. Shell commands run during the task are not attributed and are counted under `coverage`, with a missing transcript or an unreadable file. Correctness comes from the official grader, for every arm.
|
|
77
79
|
|
package/docs/operations.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Operations guide
|
|
2
2
|
|
|
3
|
-
This guide describes ahub 0.12.
|
|
3
|
+
This guide describes ahub 0.12.7 and control protocol 13. Live verification
|
|
4
4
|
results and remaining prerequisites are recorded separately in [the smoke ledger](smoke.md).
|
|
5
5
|
|
|
6
6
|
## Install and start
|
|
@@ -502,9 +502,14 @@ stays the default until an evaluation says otherwise (`docs/cooperbench.md`).
|
|
|
502
502
|
the peer that failed, never the next owner), and a peer whose version is
|
|
503
503
|
unknown has no profile. The user's and
|
|
504
504
|
plugins' hooks are not part of it. It is unknown unless the units are equal
|
|
505
|
-
and known, both peers are available (
|
|
506
|
-
|
|
507
|
-
|
|
505
|
+
and known, both peers are available (idle, or busy taking the task in
|
|
506
|
+
question: in this hub run, it is still in the turn that task started (the
|
|
507
|
+
task delivered at once to the idle peer) or in which it claimed the task; a
|
|
508
|
+
task queued, held, or steered into a turn about something else is not taken,
|
|
509
|
+
and once that turn ends, busy is another turn; for the
|
|
510
|
+
other owner, while the overlapping task is not started; the routing record,
|
|
511
|
+
and a cohort record formed as a task is assigned, are taken before the task is
|
|
512
|
+
sent, so a busy routed peer is not available then) with no other open work (an overlapping task its owner
|
|
508
513
|
has started counts), and each has five measured tasks with no more than 30%
|
|
509
514
|
failures and comparable work times. The work stage of a task ends at its first
|
|
510
515
|
`hub_task_done`, and the task itself is never one of its own observations.
|
|
@@ -642,21 +647,21 @@ Rows without a live process are stale registrations; forget them with
|
|
|
642
647
|
|
|
643
648
|
Upgrade running projects with the target release's own coordinator. It accepts
|
|
644
649
|
a running source on control protocol 9 (0.6.x), 10 (0.7.0 through 0.12.0),
|
|
645
|
-
11 (0.12.1 and 0.12.2), 12 (0.12.3) or 13 (0.12.4) and only
|
|
650
|
+
11 (0.12.1 and 0.12.2), 12 (0.12.3) or 13 (0.12.4 through 0.12.6) and only
|
|
646
651
|
a target on its own protocol, so the target's coordinator fits every supported
|
|
647
652
|
source and carries every recovery fix released up to it. Protocol 8 and older
|
|
648
653
|
(0.5.x and earlier) are refused as `manual-bootstrap-required`. Run from the
|
|
649
654
|
project directory, without replacing the global CLI first:
|
|
650
655
|
|
|
651
656
|
```bash
|
|
652
|
-
bunx --package @staix/agent-hub@0.12.
|
|
653
|
-
bunx --package @staix/agent-hub@0.12.
|
|
657
|
+
bunx --package @staix/agent-hub@0.12.7 ahub upgrade --to 0.12.7 --dry-run
|
|
658
|
+
bunx --package @staix/agent-hub@0.12.7 ahub upgrade --to 0.12.7 --yes
|
|
654
659
|
```
|
|
655
660
|
|
|
656
661
|
| Running now | Coordinator to use |
|
|
657
662
|
| --- | --- |
|
|
658
663
|
| 0.6.x (protocol 9) | the target's, through `bunx` as above |
|
|
659
|
-
| 0.7.0 through 0.12.0 (protocol 10), 0.12.1 and 0.12.2 (protocol 11), 0.12.3 (protocol 12), 0.12.4 (protocol 13) | the target's, through `bunx` as above |
|
|
664
|
+
| 0.7.0 through 0.12.0 (protocol 10), 0.12.1 and 0.12.2 (protocol 11), 0.12.3 (protocol 12), 0.12.4 through 0.12.6 (protocol 13) | the target's, through `bunx` as above |
|
|
660
665
|
| any supported source, with the installed CLI already at the target | `ahub upgrade` below, which is the same coordinator |
|
|
661
666
|
| 0.5.x or earlier (protocol 8 and older) | not supported: bootstrap by hand with the matching CLI |
|
|
662
667
|
|
|
@@ -1214,7 +1214,7 @@ upgrade sources; ordinary clients must use protocol 12 (13 from 0.12.4, with 12
|
|
|
1214
1214
|
|
|
1215
1215
|
## Amendment: shadow split prediction (issue #109)
|
|
1216
1216
|
|
|
1217
|
-
- `predictSplit(input)` is a pure function next to `assign()`, which no longer takes a split option: assignment never changes. For the pair formed by the peer routing picks and the owner of an overlapping open task, it compares a split (the later of each peer's orientation plus one unit) with the best peer alone (orientation plus two units). It is unknown unless the units are equal and known (one task of the same class is one unit), both peers are available (the
|
|
1217
|
+
- `predictSplit(input)` is a pure function next to `assign()`, which no longer takes a split option: assignment never changes. For the pair formed by the peer routing picks and the owner of an overlapping open task, it compares a split (the later of each peer's orientation plus one unit) with the best peer alone (orientation plus two units). It is unknown unless the units are equal and known (one task of the same class is one unit), both peers are available (idle, or busy taking the task in question: in this hub run, it is still in the turn that task started (the task delivered at once to the idle peer) or in which it claimed the task; a task queued, held, or steered into a turn about something else is not taken, and once that turn ends, busy is another turn; for the other owner while the overlapping task is not started; the routing record, and a cohort record formed as a task is assigned, are taken before the task is sent, so a busy routed peer is not available then; #115) with no other open work (an overlapping task its owner has started counts), and each has five measured tasks with its current profile and at most 30% failures and work times whose IQR does not exceed their median; a difference under a tenth of the single time is inconclusive.
|
|
1218
1218
|
- `Tasks.splitObservations` reads one observation per hand-over to the peer with the profile it has now, across hub runs of the project (what happened from that hand-over up to and with the next one is that peer's: a decline or escalation away is its failure, never the next owner's), never the routed task itself, leaves out claims, types failures (failed check, changes requested, escalation, release, decline, unresolved integration), ends the work stage at the first done call (checks and integration are not work) and leaves the stages unknown for an accept recorded by the done itself. The `split` event carries the trace.
|
|
1219
1219
|
- An overlap that forms or changes a cohort records a `split` event for the task, routed or named (a benchmark names every owner); `ahub route explain <id>` appends the trace as it would be now. Rerouting needs held-out evidence through #110 first.
|
|
1220
1220
|
- 2026-10-03 (T3 calibration data, issue #109 amendment): every hand-over records the new owner's split profile on its history entry: the hub's version, the agent's version (Codex's from app-server's `initialize` answer, Claude Code's from its transcript rows) and the hook profile, which is the hub's own (the coordination mode: a turn-free project runs the facts hooks in Claude). Observations count for a prediction only with the peer's current profile; while a version is unknown there is no profile, and the prediction is unknown. The prediction calibration reads is the one recorded when routing chose the first owner (no single named candidate, no claim; not an escalation, budget relay or reassignment of work already begun) of a task that overlaps another owner's task not started yet (`where: "routing"`), and `route explain` prefers the same unstarted pair; cohort-time predictions stay in the log with `where: "cohort"` and are not calibration data. The user's and plugins' hooks are not seen by the hub (ponytail: read them from the native records as the benchmark ledger does), and only Claude and Codex report a version so far.
|
|
@@ -0,0 +1,171 @@
|
|
|
1
|
+
# Agent-hub 0.12.6 verification
|
|
2
|
+
|
|
3
|
+
Implementation and candidate verification for issue #115 (follow-ups of the 0.12.5 review), performed on 2026-10-03. Native CLIs: Codex 0.160.0, Claude Code 2.1.288, Kimi Code 2.1.1. Production application is verified separately after publication.
|
|
4
|
+
|
|
5
|
+
## The macOS CI hang of 0.12.5
|
|
6
|
+
|
|
7
|
+
On 0.12.5's CI, `test/daemon.test.ts` once hung until the job limit; no test reached its own timeout, so the event loop was blocked. It reproduced locally twice under heavy machine load (load average about 12), once on the 0.12.5 checkout and once on this branch. Each time `bun test` sat at 100% CPU for minutes, with nothing but the banner in its output in one case.
|
|
8
|
+
|
|
9
|
+
The stacks were sampled with macOS `sample` and symbolized with Bun's own profile build of 1.3.14, whose UUID matches the installed binary (`C7E7A979-F99B-3466-9AD6-E56A63373A35`). The hung main thread shows this chain:
|
|
10
|
+
|
|
11
|
+
1. An async continuation, resumed from a timer, calls `Bun.spawnSync` (`BunObject_callback_spawnSync`, `spawnMaybeSync`).
|
|
12
|
+
2. Inside spawnSync's own event loop, the test runner's timeout callback fires (`BunTest.bunTestTimeoutCallback`). Through `BunTest.run` and `runTestCallback`, it runs the next test's JavaScript.
|
|
13
|
+
3. That JavaScript calls `Bun.spawnSync` again, and the inner call spins in `SpawnSyncEventLoop.tickWithTimeout` for good.
|
|
14
|
+
|
|
15
|
+
So a test timeout that fires while a `Bun.spawnSync` runs can hang the whole run. Minimal attempts to reproduce it outside the suite did not hang: when a test times out, Bun kills the child that test's own spawnSync started. The precise trigger therefore also needs a spawnSync that Bun does not attribute to the timed-out test.
|
|
16
|
+
|
|
17
|
+
Mitigation, since the fix belongs in Bun:
|
|
18
|
+
|
|
19
|
+
- `scripts/check.sh` runs `bun test --timeout 20000`, so a timeout, which is the trigger, is much rarer under load.
|
|
20
|
+
- It runs the suite under `scripts/hang-watch.sh`. If the run is still going after 600 s, the watchdog prints a stack sample (macOS) or the thread states (Linux) and stops it, so a hang fails fast with evidence.
|
|
21
|
+
- The process-heavy stop tests carry timeouts of their own.
|
|
22
|
+
- Upgrading Bun (1.4.2 is current) is left for a separate change.
|
|
23
|
+
|
|
24
|
+
The permission-test flake of 0.12.5 had a different cause: the test rig's 200 ms permission timeout is shorter than a loaded runner takes to connect a peer and make two tool calls. That test now has its own 30 s window.
|
|
25
|
+
|
|
26
|
+
## Automated gate
|
|
27
|
+
|
|
28
|
+
`scripts/check.sh` at the head of the pull request, after the eleventh review round and the merge of main (#118): 754 tests passed, 0 failed, 3,807 expectations across 68 files in 135 s (macOS CI counts one expectation fewer) (the 68th file is #118's); `check: OK` (with the fourth round's fixes: 739 passed, 3,723 expectations; at `b03b035`: 739 passed, 3,710 expectations; at `fdc884e`: 738; at `4b98d2d`: 737 passed in 116 s at a load average of about 5). Mutation checks (each test fails with its fix removed; the last four are 0.12.5 tests checked again against this release's stop change): split availability under the 0.12.5 rule, the ACP launcher test without `detached`, pipes in the non-group stop, the fallback's newcomer freeze, the kill-only-stopped rule, the reused-pid stop, and the grace-period SIGTERM to own-group descendants.
|
|
29
|
+
|
|
30
|
+
CI on `4b98d2d`: macOS passed; Ubuntu failed the new ACP launcher test, because dash points a background job's stdin at `/dev/null` before the explicit `0<&0`, so the fake agent never got its handshake. The test now passes stdin through fd 3 (`bef35ab`); CI passed on both platforms there, at `fdc884e`, `5bfaf53`, `70b1daf` and `b03b035`.
|
|
31
|
+
|
|
32
|
+
Cost of the group stops: each ACP or Pi stop now reads the process table twice (a `ps` of about 1,300 processes takes about 60 ms here). `test/acp.test.ts` and `test/pi-daemon.test.ts` together take 8.1 s against 3.1 s on 0.12.5. The group stop now finishes as soon as the leader has exited and a read shows nothing left.
|
|
33
|
+
|
|
34
|
+
## Live checks
|
|
35
|
+
|
|
36
|
+
- Upgrade, on `4b98d2d`: `scripts/smoke-recovery-09-10.ts` recovered the published 0.12.5 package (protocol 13, integrity from the registry) into the candidate (protocol 13): operation `31f206a6-6dbc-4f42-adfb-40eac841417a`, queue and task identity and state digest preserved.
|
|
37
|
+
- Kimi through ACP, on `4b98d2d`, real `kimi acp` 2.1.1, in a disposable project with its own `AGENTHUB_HOME`: started by the hub as the daemon's child, it led its own process group; `ahub kill` said "hub stopped", nothing of that group was left, and the daemon was gone.
|
|
38
|
+
- Interrupted benchmark arms with real agents, on `4b98d2d` (A, B) and `bef35ab` (C, E, V, S2; it differs only in the ACP test):
|
|
39
|
+
|
|
40
|
+
| Run | Arm | Stopped at | Signal | Cleanup | Left running | Restoration |
|
|
41
|
+
|---|---|---|---|---|---|---|
|
|
42
|
+
| A | solo Codex | 31.1 s into the work | SIGTERM to the runner | `clean` (298 ms) | nothing | inputs, sibling locks |
|
|
43
|
+
| B | advisory | 40.3 s into the work | SIGINT to the group | `clean` (463 ms) | nothing | inputs, locks, trust |
|
|
44
|
+
| C1 | solo Claude | completed (57.7 s); turn end after 0.75 s, tree unchanged | none | `clean` (340 ms) | nothing | inputs, locks, trust |
|
|
45
|
+
| C2 | advisory | setup | SIGTERM to the runner | `clean` (424 ms) | nothing | inputs, locks, trust |
|
|
46
|
+
| E | turn-free | 31.2 s into the work, and again during teardown | SIGINT to the group, twice | `clean` (422 ms) | nothing | inputs, locks, trust |
|
|
47
|
+
| V | solo Codex with a visitor in the fixture | 31.0 s into the work | SIGTERM to the runner | `incomplete_or_unknown`, the visitor left alone and recorded as `program: sleep` | the visitor | withheld; `restore.ts` refused while it ran, restored once it left |
|
|
48
|
+
| S2 | advisory | native readiness | SIGINT to the group | `clean` (452 ms) | nothing | inputs, locks, trust |
|
|
49
|
+
|
|
50
|
+
Runs A, V and S2 were repeated on `fdc884e` with the same outcomes. Every one of these runs registered its fixtures with Orca, since the runner still added them then. The user has since directed that no fixture is added to Orca without an explicit registration (#117), and the registrations these runs left are removed under #117. Since `70b1daf` the runner only looks registrations up, and since `b03b035` it refuses an unregistered selection up front. A live check of the latter registered nothing. A run prepared afresh (`/private/tmp/ahub-0126-refusal`, 40 fixtures, none registered) was started with the runner of `b03b035`. It refused the whole run before anything was locked or recorded, naming every unregistered fixture. That lookup was later replaced by #118's read-only helper (`scripts/benchmarks/orca-workspace.ts`, merged into this branch from main), which refuses at the first fixture missing and checks each arm's identity against the preflight; #118 records its own checks, and this refusal check was not repeated on it. Orca's repository list was the same before and after (15 repositories), the upstream checkout kept its mode, and no attempt record was written. Before the release, no arm had run on the lookup-only runner, which needs fixtures registered explicitly first: the live arms above ran on the runner that still added them. Every record carries `platform: darwin`. The Codex skills condition was 177 user and 5 system skills in every arm above with Codex. Arms on the released runner follow.
|
|
51
|
+
|
|
52
|
+
After the release, with the user's authorization for the Orca registrations, three of these arms ran on the released runner (`44e04e5`, Bun 1.3.14, which 0.12.6's CI used). Each run's four case-0 fixtures were registered by hand before the run and removed after it, and Orca's repository list ended as it began (16 repositories, the same ids; the one more than at the refusal check is an unrelated project added meanwhile). Every record carries `platform: darwin`; the Codex skills condition was 177 user and 4 system skills in each.
|
|
53
|
+
|
|
54
|
+
| Run | Arm | Stopped | Cleanup | Left running | Restoration |
|
|
55
|
+
|---|---|---|---|---|---|
|
|
56
|
+
| A | solo Codex | SIGTERM to the runner, 30.8 s into the work | `clean` (313 ms) | nothing | inputs, sibling locks |
|
|
57
|
+
| B | advisory | SIGINT to the group, 40.7 s into the work | `clean` (447 ms) | nothing | inputs, locks, trust |
|
|
58
|
+
| V | solo Codex with a visitor in the fixture | SIGTERM to the runner, 30.6 s into the work | `incomplete_or_unknown`, the visitor left alone and recorded as `program: sleep` | the visitor | withheld; `restore.ts` refused while it ran, restored once it left |
|
|
59
|
+
|
|
60
|
+
## Review
|
|
61
|
+
|
|
62
|
+
Independent read-only reviews (Claude Code OCR delegation, with the OCR rule groups, `REVIEW.md`, the `AGENTS.md` invariants and issue #115), each on the head of the time. The first round, on `4b98d2d`, found one Important gap: this record and the spec amendment for the Bun cause were missing. The spec was amended (#115, Decisions of 2026-10-03), and this record was written. Its Minor findings were fixed:
|
|
63
|
+
|
|
64
|
+
- Hand-overs in the split prediction are keyed by owner generation: a task handed back to a peer that had it is a new hand-over.
|
|
65
|
+
- A failed Pi start reports its own error, and Pi's stop finishes its teardown when the group stop fails.
|
|
66
|
+
- Pi has a launcher test, as ACP does.
|
|
67
|
+
- The spawn-heavy stop tests' own timeouts are at least the suite's 20 s.
|
|
68
|
+
- The runner removes the trust write's temp file on every path, the recovery's included.
|
|
69
|
+
- `hang-watch.sh` signals only the process it began watching (pid and start time), escalates to SIGKILL, and is stopped once the tests end.
|
|
70
|
+
- Run records and the restoration ledger are written atomically.
|
|
71
|
+
- `programOf` looks at a bounded number of argv words.
|
|
72
|
+
- Several doc and comment fixes.
|
|
73
|
+
|
|
74
|
+
Deferred: the Codex start-failure test does not exercise a stop that fails, since no fake can make the group stop fail there without mocking. The second round, on `5bfaf53`, found no Critical or Important defect. Its Minor findings were fixed:
|
|
75
|
+
|
|
76
|
+
- The program name comes from `ps`'s name for the executable, not from stats that can hang on a dead mount.
|
|
77
|
+
- A failed restore keeps the trust write `pending` for the recovery.
|
|
78
|
+
- The trust temp files are named in one place, and removed when a write or a restore fails.
|
|
79
|
+
- The recovery's pending-trust path is tested.
|
|
80
|
+
- The Bun mechanism is worded as what the stacks show, not as a general rule.
|
|
81
|
+
|
|
82
|
+
The runtime review of that round also found Minor gaps. Fixed: a claimed task now counts as handed over for the split prediction; Claude's version is cached when the transcript has none yet; the watchdog test waits for its trap and always cleans up; `sample` writes its report to a temp file it removes. Deferred and listed in a comment on the pull request: a task steered into a turn about other work counts as taken (shadow only); the start-failure tests do not exercise a failing stop. The same commit makes the runner lookup-only in Orca (#117).
|
|
83
|
+
|
|
84
|
+
The third round, on `70b1daf`, found two Important defects; both are fixed.
|
|
85
|
+
|
|
86
|
+
- **Arguments recorded as a program name.** The program name read from `ps -o comm=` is, on macOS, a process's own argv[0], so a process that sets its title (Node's `process.title`, perl's `$0`) put its arguments into the record. It now reads `ps -o ucomm=`, the executable's name as the kernel recorded it, only while the pid has the start time recorded, and records nothing when it cannot be read. A test with a self-titled perl fails with `comm`.
|
|
87
|
+
- **A steered task counted as taken.** A task steered into a busy Codex turn counted as taken by it. Only a task that starts the owner's turn counts now (idle with nothing queued when it is sent, or claimed), and a test with a steering peer covers it.
|
|
88
|
+
|
|
89
|
+
Its Minor findings were fixed:
|
|
90
|
+
|
|
91
|
+
- The runner, whose own trust write threw, no longer touches the trust file.
|
|
92
|
+
- An unregistered fixture refuses the run up front instead of failing every arm, and the documented procedure has the registration step.
|
|
93
|
+
- The operator note on refused Kimi and Pi restarts was added.
|
|
94
|
+
- Several wording fixes.
|
|
95
|
+
|
|
96
|
+
Deferred, as before: the start-failure tests do not exercise a failing stop.
|
|
97
|
+
|
|
98
|
+
The fourth round, on `b03b035`, found two Important defects in the split prediction; both are fixed.
|
|
99
|
+
|
|
100
|
+
- **A hand-over outlived its turn.** A task that started its owner's turn stayed marked as taken after that turn ended, so a later turn about something else counted as taking it. The mark now lasts as long as that turn: the peer leaving busy drops it. It is set only once the task envelope was delivered at once (idle before, busy after, not queued), so a held delivery is not taken; a claim counts only in a turn the hub sees.
|
|
101
|
+
- **The queue was checked again at prediction time.** A peer still in the turn its task started, with a status message queued behind it, read as not available, and so did a claimant with one queued. Only the turn decides now.
|
|
102
|
+
|
|
103
|
+
Tests cover a turn that ended, a message queued behind the turn, a claim with one queued, a held delivery and the daemon's journal, and each fails with its part of the fix removed. Its Minor findings were fixed:
|
|
104
|
+
|
|
105
|
+
- The ledger counts a record caught in the middle of the recovery's move once, as withheld (also raised by the Codex review on the pull request), and refuses a run whose `runs/` is still locked instead of reporting its attempts missing.
|
|
106
|
+
- The Orca refusal names why each lookup failed, and each arm looks its fixture up before touching it.
|
|
107
|
+
- The recovery removes the temp file a runner left when it died in its own trust restore, and a temp file it cannot remove keeps the recovery open instead of aborting it.
|
|
108
|
+
- The tail permission test has the 30 s permission window too.
|
|
109
|
+
- The docs name the hub's project registration where they meant it, and the availability rule names the claim and the turn.
|
|
110
|
+
- Several nits: a nested ternary, the version cache stamped only after a read that worked, a failed `sample` said so.
|
|
111
|
+
|
|
112
|
+
Deferred, as before: the start-failure tests do not exercise a failing stop (a fake cannot end the launcher between the health check and the failure that follows it).
|
|
113
|
+
|
|
114
|
+
The fifth round, on `56585ea`, found no Critical or Important defect in the change. It found that main had moved: #118 (#117's Orca change) had been merged, so the pull request conflicted and no CI had run on its head. Main was merged in. The Orca lookup is #118's: its helper, its preflight and its per-arm identity check replace this branch's own preflight and lookup move. This branch keeps the guard that asks for a hub shutdown only once an arm got past its locks, which the merged code needs (its `orcaProject` is now always set). The docs keep #118's text on explicit authorization for Orca registrations. Its Minor findings were fixed:
|
|
115
|
+
|
|
116
|
+
- A pause no longer drops a hand-over: the turn's end is read from the adapter's own state.
|
|
117
|
+
- A temp file the recovery cannot remove keeps the trust entry unrestored, so the next recovery tries again; it had been marked restored and never retried. A test covers it.
|
|
118
|
+
- The ledger no longer refuses a run whose `runs/` is locked: the attempts it cannot read there are listed as `unreadable`, never as missing, and the withheld ones are rows as before.
|
|
119
|
+
- #115's Design and #109's Decisions were amended to the turn rule; test assertions that could pass without a routing record now require one.
|
|
120
|
+
|
|
121
|
+
CI on the merge (`46e22e0`): macOS passed; Ubuntu failed two of #118's tests. They run the runner's entrypoint, which on Linux stops at this release's macOS-only refusal before the Orca checks. The helper-pin test now expects that refusal off macOS, with the same checks that nothing was written or run, and the preflight test runs on macOS only, where the runner does. CI passed on both platforms at `e2d66ea`.
|
|
122
|
+
|
|
123
|
+
The sixth round, on `46e22e0`, found the Ubuntu failure above (fixed in `e2d66ea`) and checked the merge: #118's helper, its pin, its preflight before any change and its per-arm identity check are intact, and nothing of either side was lost or duplicated. Its Minor findings were fixed:
|
|
124
|
+
|
|
125
|
+
- The runner's own pending trust write whose temp file cannot be removed keeps the trust entry open for the recovery, as the recovery already did for its own; it had been marked settled.
|
|
126
|
+
- The ledger's per-arm summary lists the attempts a locked `runs/` hides (`unreadable`), so the summary a record quotes accounts for them.
|
|
127
|
+
- An interrupt right after the locks no longer asks a hub that never started to stop.
|
|
128
|
+
- A claim reads the claimant's own adapter state, as the turn's end does.
|
|
129
|
+
|
|
130
|
+
CI passed on both platforms at `e9137dd` and `5440086`. The seventh round, on `e9137dd`, found no Critical or Important defect. Its Minor findings were fixed:
|
|
131
|
+
|
|
132
|
+
- When the runner's own temp file could not be removed, the recovery treated the write as possibly landed and could take back an entry someone else had set meanwhile. The runner now records the write as `not_written`, and the recovery removes the copy without touching any entry. A test with someone else's entry covers the recovery's side and fails without its fix.
|
|
133
|
+
- An arm with no record at all appears in the ledger's summary with what it owes (`missing`, `unreadable`).
|
|
134
|
+
|
|
135
|
+
The claim's change was left untested in that round, on the premise that only the cohort record at the claim reads its mark; `route explain` reads it too (round 8).
|
|
136
|
+
|
|
137
|
+
The eighth round, on `5440086`, found no Critical or Important defect. It found the same class a fourth time: with an incomplete cleanup, a trust write the runner knew never landed was left `pending`, so the recovery could take back an entry someone else set meanwhile. Following `REVIEW.md`, the rule went into `AGENTS.md`. A write still `pending` in the runner's own process is `not_written` on every path, and neither the runner nor the recovery then touches an entry. Its other Minor findings were fixed:
|
|
138
|
+
|
|
139
|
+
- The recovery re-settled a `changed_concurrently` entry (the user's) when the run was unrestored for another reason, and could take it back. It now leaves it. This was older than this release (0.12.5). A test covers it and fails without the fix.
|
|
140
|
+
- The summary seeds an arm with no record into its full shape, so its common medians count it.
|
|
141
|
+
- A claim made while paused, then resumed in the same turn, is read as taken by `route explain`. A test covers it and fails with the bus's `paused` state.
|
|
142
|
+
- `restoration.json` names a temp file left behind, rather than a trust entry, when that is what is left; the summary's units and the docs describe `unreadable`.
|
|
143
|
+
|
|
144
|
+
CI passed on both platforms at `e5940e0`. The ninth round, on `e5940e0`, enumerated every state the trust entry can be left in by the runner, a runner that dies at any point, and the recovery. It found no Critical or Important defect, and every path the runner settles itself correct. Its Minor findings were fixed:
|
|
145
|
+
|
|
146
|
+
- A runner that died between its trust write and the rename leaves `pending`, and the recovery took back any entry set meanwhile. It now takes back only an entry exactly as the runner would have written it, and none while the runner's temp file is still there (the rename never happened). This was older than this release. A test covers the three cases and fails without either check.
|
|
147
|
+
- The recovery's own restore named its temp file by its own pid, which no later recovery looks for. It now uses the runner's.
|
|
148
|
+
- The runner's settlement is a function (`settleTrust`), tested as a table of stage, containment and temp removal; the `AGENTS.md` rule points to it. The table test fails with the earlier order (containment before `pending`).
|
|
149
|
+
- The common medians' change has a test: an arm with no record leaves no common pair. It fails without the change.
|
|
150
|
+
|
|
151
|
+
CI passed on both platforms at `bbf59d9`. The tenth round, on `bbf59d9`, found no Critical or Important defect. It compared `settleTrust` with the inline code it replaced, field by field, and found them equal. It walked every point at which a runner can die, every stage on disk, and the 0.12.3 to 0.12.5 ledgers, and found no path where the recovery changes an entry the runner never wrote or the user changed. Fixed:
|
|
152
|
+
|
|
153
|
+
- A test now checks that the recovery names its restore's temp file by the runner's pid. It fails when the recovery's own pid is used.
|
|
154
|
+
- The `TrustLease` comment now says that every ledger since 0.12.3 records `written`.
|
|
155
|
+
- The `AGENTS.md` rule no longer gives a count of rounds.
|
|
156
|
+
|
|
157
|
+
Deferred, older than this release and listed in the pull request:
|
|
158
|
+
|
|
159
|
+
- A recovery of a 0.12.3 or 0.12.4 ledger (no runner identity) killed mid restore leaves a temp file that no later recovery looks for.
|
|
160
|
+
- A 0.12.5 runner whose rename threw left its temp file behind a `restored: true` run.
|
|
161
|
+
- A run directory whose earlier invocation stopped before any arm keeps its `restoration.json` (`restored: true`). A recovery after a crashed re-run there reports the run restored and exits 0 while the re-run's locks and trust entry stay; it changes no entry.
|
|
162
|
+
|
|
163
|
+
CI passed on both platforms at `e0b71a0`. The eleventh round, on `e0b71a0`, found no Critical or Important defect. The new test now runs the recovery as this process (`self` is its pid), so a temp file named by either the recovery's pid or `self` fails it; the record's third deferred item says what the recovery reports.
|
|
164
|
+
|
|
165
|
+
Deferred: the start-failure tests, as before.
|
|
166
|
+
|
|
167
|
+
## Known limits
|
|
168
|
+
|
|
169
|
+
- The Bun 1.3.14 hang is mitigated, not fixed. A test that times out during a spawnSync can still hang a run, which the watchdog then stops after 600 s.
|
|
170
|
+
- Group stops cost about two process-table reads per stop.
|
|
171
|
+
- The limits recorded for 0.12.5 stand.
|
|
@@ -0,0 +1,58 @@
|
|
|
1
|
+
# Agent-hub 0.12.7 verification
|
|
2
|
+
|
|
3
|
+
Implementation and candidate verification for issue #120 (the CooperBench run's restoration state), performed on 2026-10-03, with Bun 1.4.2 (#121). Native CLIs: Codex 0.160.0, Claude Code 2.1.288. The hub runtime is unchanged from 0.12.6 apart from the plugin bundle, which is built with Bun 1.4.2. Production application is verified separately after publication.
|
|
4
|
+
|
|
5
|
+
## What changed
|
|
6
|
+
|
|
7
|
+
- The runner writes `restoration.json` as `{restored: false, runner}` right before it locks anything, and its outcome replaces it at the end. A runner killed in between leaves a file that sends the recovery in.
|
|
8
|
+
- The runner refuses a run directory whose earlier run is not restored (`reuseProblem`), before it reads any mode.
|
|
9
|
+
- `restore.ts` trusts `restored: true` only when the restoration ledger agrees (`unrestored`). It waits while a runner named by either file still runs. With no ledger, nothing was locked; with neither file, there is nothing to recover.
|
|
10
|
+
|
|
11
|
+
## Automated gate
|
|
12
|
+
|
|
13
|
+
`scripts/check.sh` on Bun 1.4.2, at the head of the pull request, with the fixes of three review rounds: 759 tests passed, 0 failed, 3,848 expectations across 68 files in 132 s; `check: OK` (757 and 3,833 before them). Mutation checks (each test fails with its part removed): the old early return on `restored: true`, the runner named by `restoration.json` ignored, and a ledger required.
|
|
14
|
+
|
|
15
|
+
## Live checks
|
|
16
|
+
|
|
17
|
+
Orca lookups in these runs were answered by a read-only stand-in for the `orca` CLI. It answered `repo list` and `worktree list` for the run's own fixtures, passed `worktree current` to the real Orca, and refused everything else. Nothing was registered in Orca. The arms were real: Codex 0.160.0 through the hub, with the inputs and sibling artifacts really locked.
|
|
18
|
+
|
|
19
|
+
- **SIGKILL mid-arm** (`/private/tmp/ahub-0127-K2`, solo Codex, killed 30 s into the work):
|
|
20
|
+
- Before the kill, `restoration.json` was `restored: false` and named the runner. It stayed so after the kill. The protected root and a sibling fixture were mode 000, and the arm's daemon and Codex app-server were still running.
|
|
21
|
+
- A re-run in the same directory was refused (exit 1), but by its first read of the locked private inputs (a bare EACCES), before the reuse check. The review moved the check first; see K4.
|
|
22
|
+
- `restore.ts` refused while the daemon, the app-server and what ran below them were running (exit 1, each named). After `ahub kill` it restored the run (exit 0): the modes were back, and `restoration.json` said `recovered`.
|
|
23
|
+
- **Reuse refused** (`/private/tmp/ahub-0127-K3`, a fresh preparation whose `restoration.json` said `restored: false`): the real runner refused it with "an earlier run here is not restored (restoration.json)". It refused before it read any mode or wrote a ledger or `cohort.json`; the protected root's mode was unchanged. `restore.ts` then found no ledger and a runner that was gone, and marked the directory restored.
|
|
24
|
+
- **SIGKILL mid-arm again, after the review** (`/private/tmp/ahub-0127-K4`, the same run as K2; its log names `58c8db0`, since the fixes were not committed yet, but its `prepared.json` and `cohort.json` pin the sources of `bdc5581`): the re-run in the same directory was now refused by the reuse check, naming `restore.ts` ("an earlier run here is not restored (restoration.json): run bun scripts/benchmarks/restore.ts --run ... first"). The recovery was held while the arm's processes ran, and restored the run after `ahub kill`.
|
|
25
|
+
- **SIGKILL mid-arm on the final runner** (`/private/tmp/ahub-0127-K5`, after the third review round; its log names `ff2480e`, since the fixes were not committed yet, but its `prepared.json` pins the final `native.ts` and `teardown.ts`; and so does its `cohort.json`; `restore.ts` is not pinned, and this run's ledger names its runner, so `owed` returns it unchanged and the run does not exercise the pre-0.12.5 rule): the same outcome as K4. The marker stayed, the re-run was refused by the reuse check naming `restore.ts`, the recovery was held while the arm's processes ran, and it restored the run after `ahub kill`.
|
|
26
|
+
- **A first attempt** (`/private/tmp/ahub-0127-K`) stopped before the marker, because the stand-in did not yet answer `worktree current`. Nothing was locked or written. The recovery of that directory, which had neither file, blocked on "the ledger does not name the runner"; it now reports nothing to recover (a test covers it).
|
|
27
|
+
- **Upgrade:** `scripts/smoke-recovery-09-10.ts` recovered the published 0.12.6 package into the candidate. Both run protocol 13, and the source's integrity came from the registry. Operation `b7643429-7f0c-4f70-8128-857bb965a5d2`: the queue was preserved, and the task digest is the same as in 0.12.6's check.
|
|
28
|
+
|
|
29
|
+
## Known limits
|
|
30
|
+
|
|
31
|
+
- One runner per run directory at a time: the reuse check runs again right before the marker, but two runners started on one directory at the same moment can both pass it (a `ponytail:` in `native.ts` names the exclusive claim that would close it).
|
|
32
|
+
- Grading still trusts `restoration.json` alone. With the marker, a run whose runner died says `restored: false`, so grading refuses it; a stale `restored: true` can only come from a run before 0.12.7.
|
|
33
|
+
- Deferred, as listed in #116: a recovery of a 0.12.3 or 0.12.4 ledger killed mid restore leaves a temp file no later recovery looks for, and a 0.12.5 runner whose rename threw left its trust temp file behind a `restored: true` run.
|
|
34
|
+
|
|
35
|
+
## Review
|
|
36
|
+
|
|
37
|
+
Independent read-only reviews (Claude Code OCR delegation), each on the head of the time. The first, on `58c8db0`, went through every state a run directory can be in, and through a runner and a recovery started together. It found no Critical or Important defect. Its Minor findings were fixed:
|
|
38
|
+
|
|
39
|
+
- The reuse check runs first, before the runner resolves its inputs, so a directory whose inputs a dead runner locked is refused with a message that names `restore.ts`, not a bare EACCES (K4). It runs again right before the marker. The limit for two runners started together is recorded above.
|
|
40
|
+
- `restore.ts` reads the process table before the files, so the runners the files name are checked again right before anything is restored. A test with a runner missing from the first table covers it.
|
|
41
|
+
- A ledger from before 0.12.5 (no runner identity) under `restored: true` keeps its trust entry as it was then. Those versions left a concurrently changed trust entry at `written`, and taking it back would remove the user's entry. A test covers it.
|
|
42
|
+
- `unrestored` says a `not_written` entry whose temp file is left is still unrestored, which is what the code did; the spec and the comment now say so, with table rows.
|
|
43
|
+
- The runner's outcome uses `unrestored`, so it and the recovery's agreement check cannot disagree. Smaller wording fixes.
|
|
44
|
+
|
|
45
|
+
The second round, on `bdc5581`, found no Critical or Important defect. Its Minor finding was fixed. The exception for ledgers from before 0.12.5 covered the whole ledger, so a re-run that died over such a directory kept its inputs locked. Meanwhile the reuse check sent the operator to a recovery that then did nothing. Now only the trust entry is kept as it was then (`owed` in `teardown.ts`); the runner's reuse check and the recovery use the same rule, and the locks are restored. A test covers it and fails with the earlier rule. Its nits were fixed too:
|
|
46
|
+
|
|
47
|
+
- Nothing is written to the run directory before the marker, so a run that died between the two checks keeps its `cohort.json`.
|
|
48
|
+
- The unused table parameter is gone.
|
|
49
|
+
- The reasons read "still unrestored: ...".
|
|
50
|
+
|
|
51
|
+
The third round, on `ff2480e`, found no Critical or Important defect. Its Minor finding was fixed: a 0.12.7 runner that accepted a directory with a pre-0.12.5 ledger, wrote its marker and died before a ledger of its own left that old ledger under a marker. The recovery then took the old trust entry back, because the rule keyed on `restored: true`. `owed` now also applies under a marker that names a runner; a test covers it and fails without it. Its nits were fixed too:
|
|
52
|
+
|
|
53
|
+
- The rule lives in `owed` alone, and the recovery settles the trust entry `owed` returns.
|
|
54
|
+
- The run directory's writes are inside the `try`, so a failure there still ends with an outcome.
|
|
55
|
+
- `AGENTS.md` names `owed`, and the spec wording is fixed.
|
|
56
|
+
- K5 ran the final runner.
|
|
57
|
+
|
|
58
|
+
The fourth round, on `68a0f9b`, found no Critical or Important defect. It ran the reuse check and the recovery over 35 directory shapes, and in none did the reuse check accept a directory where the recovery then restored anything. Its nits, all wording, were fixed: the `AGENTS.md` rule, the CHANGELOG and the docs now state the condition under which the old trust entry is kept, and K5's provenance is noted.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "agent-hub",
|
|
3
|
-
"version": "0.12.
|
|
3
|
+
"version": "0.12.7",
|
|
4
4
|
"description": "Channel between Claude Code and the agent-hub daemon: peer messages from Codex, Kimi and the local worker arrive as channel events; hub_send replies.",
|
|
5
5
|
"author": {
|
|
6
6
|
"name": "Young Joon Lee",
|