@staix/agent-hub 0.12.6 → 0.12.7
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +6 -0
- package/docs/cooperbench.md +1 -1
- package/docs/operations.md +5 -5
- package/docs/verification/2026-10-03-0.12.6.md +9 -1
- package/docs/verification/2026-10-03-0.12.7.md +58 -0
- package/package.json +1 -1
- package/plugins/agent-hub/.claude-plugin/plugin.json +1 -1
- package/plugins/agent-hub/server.js +632 -956
- package/src/hub/project.ts +5 -4
package/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,12 @@
|
|
|
2
2
|
|
|
3
3
|
Issue and pull request numbers in the entries for 0.7.7 and earlier refer to the previous repository, archived on 2026-09-30 when this repository's history was rewritten; the one exception is the open smoke-check issue, formerly #12, which moved here as #1. Numbers in newer entries refer to this repository.
|
|
4
4
|
|
|
5
|
+
## 0.12.7
|
|
6
|
+
|
|
7
|
+
- CooperBench runner and recovery (#120): from the moment the runner may change anything, before it locks its first input, `restoration.json` says `restored: false` and names the runner, until the runner writes its outcome. A runner killed after that point sends the recovery in, and the recovery waits while the runner it names still runs. A run directory whose earlier run is not restored is refused first, before the runner reads its inputs or any mode (the refusal names `restore.ts`), and checked again right before the runner locks anything, so a runner does not record locked modes as the originals; two runners started on one directory at the same moment are not guarded against. A directory whose earlier invocation ended restored, with no records, can still be used. `restore.ts` trusts `restored: true` only when the restoration ledger agrees (a ledger from before 0.12.5, with no runner identity, still has its locks restored; when `restoration.json` says `restored: true` or names a runner, its trust entry is kept as it was then), so a stale file over a re-run that died no longer reports a run restored while its inputs stay locked; it checks the runners again right before it restores anything. With no ledger, nothing was locked and nothing is restored.
|
|
8
|
+
- CooperBench manifests v2 and the #106 ablation pin hub 0.12.7.
|
|
9
|
+
- Bun 1.4.2 in CI, the release workflow and the plugin bundle build. Bun 1.4.2 still throws for a real path with a backslash, so `realPath` stays, and still runs the next test inside an outer spawnSync's event loop after a timeout, so `check.sh` keeps its test timeout and watchdog. On 1.4.2, a model relay that is closed while it streams a response gets an error printed by Bun when it aborts that stream ("model relay closed"): log noise, not a failure (#121).
|
|
10
|
+
|
|
5
11
|
## 0.12.6
|
|
6
12
|
|
|
7
13
|
- Agents the hub spawns through ACP (Kimi) and Pi run in their own process group and are stopped as one, as the Codex app-server is since 0.12.5. A stop with or without a group drops the child's pipes, so a process it left cannot keep the hub alive. Once a launcher has exited, its group is followed until it is gone, and members still finishing their exit get a bound before the stop fails. A Codex start that fails reports its own error. The group stop finishes as soon as the leader has exited and a read shows nothing left. As for Codex since 0.12.5, a Kimi or Pi that crashed and left processes in its group (an MCP server, a tool command) is not restarted until they are gone: the error names their pids (#115).
|
package/docs/cooperbench.md
CHANGED
|
@@ -53,7 +53,7 @@ bun scripts/benchmarks/native.ts \
|
|
|
53
53
|
|
|
54
54
|
The native runner runs on macOS only and refuses anything else (each run record says `platform`): the arms run in Orca terminals, and on Linux a clock step moves the start times its teardown proves processes by. Before native execution, obtain explicit user authorization before changing Orca registrations; a benchmark request alone is not authorization. Only after that authorization, an operator manually registers each exact fixture root with `orca repo add --path '<fixture>'` and confirms `orca worktree list --repo id:<repo-id>` shows that exact path. The runner performs read-only exact-path repo/worktree lookup; it preflights every selected fixture before changing fixture files, input modes, cohort or attempt records. If either identity is missing or mismatched, it exits nonzero with setup guidance and leaves the prepared run available for retry after explicitly authorized registration. Each arm rechecks its identity before changing fixture files. Benchmark and agent workflows never add or remove Orca registrations. The runner uses the project CLI to start and stop each daemon, and requires exact worktree/cwd readback for native sessions. Claude starts through the canonical `ahub claude` guard and loads this checkout's candidate bundle through an exact session-only `--mcp-config` server (`server:agent-hub`); it does not promote or mutate the globally installed plugin. Codex uses the native app-server adapter and `workspace-write` sandbox. Each agent must execute a setup-only `head -c 1` probe against the exact protected file and produce only the denied marker before scored tasks begin. There is one native sandbox layer per agent.
|
|
55
55
|
|
|
56
|
-
Every run record, diff, PTY/provider trace and evaluation artifact stays in the mode-0700 private run directory. Setup errors, structured provider quota errors, hub budget pauses, unsettled deliveries and interruptions stay separate and unscored. A quota snapshot or numeric `429` alone is not a provider error. On exit the runner restores the exact input and read-lock modes, the sibling artifact modes and the scoped Claude trust flag (removing its own fresh Claude trust entry), under the condition below. Teardown (issue #113) proves what it stops. Each process an arm starts is recorded with its pid and start time and the evidence that it is the arm's: the daemon by the pid in its state directory and an argv that serves this fixture (a later daemon there is a replacement, never adopted), Claude's launch chain by the arm's own session id as its `--session-id` (and the Orca terminal shell that runs it, working in the fixture), the Codex app-server as the daemon's child. What runs below them, and what is in a group a recorded process leads while that group is known to be the same one (its leader alive, or a recorded member still in it), is the arm's too; the table is read every 5 s while the agents work and during the completion wait, so a tool command's background job is recorded while its parent runs (one that detaches between two reads, outside the fixture and without it in its argv, is not seen). A fixture name in an argv is never proof on its own (Codex's need not name the fixture), and every recorded actor counts even when a read of the table fails. Teardown pauses every agent, reads the deliveries still owed, lets a completed arm's Claude end the turn it is in (up to 30 s and never past the 300 s limit (the bound is taken when the wait starts; one 250 ms poll can pass it), judged by the `turn_duration` row Claude Code writes after a turn, proved on the arm's probe turn; a stopped or timed-out arm does not wait; a completed arm's tree is hashed at the end of its active time and again after the teardown, and a write in between flags the attempt invalid), asks for the normal shutdown (Claude's terminal closed, `ahub kill`, the hub's project registration removed (`ahub projects remove`); the runner leaves Orca registrations untouched; a step of it that fails is recorded in `cleanup.normal.errors` and does not by itself fail the attempt, since the process readback decides, so a hub project registration left behind is listed there), and reads the process table back, in the C locale and in UTC (start times are read as identities, the same in every reader). What is still running after a settle period gets SIGTERM; what is left then is frozen (SIGSTOP), the table read again, and it is killed (SIGKILL), what was frozen and not killed being continued: a frozen process starts nothing, so that read sees all of it. Every signal goes only to an identity read again just before, by the group the table shows it in now; a group leader is signalled with its group. What the read after the freeze shows for the first time is frozen too and read again before anything is killed, and when such a read fails only what a STOP reached is killed, by the last read that showed it (a stopped process keeps its pid). Anything with the fixture in its argv or as its working directory that is not proved the arm's (a background job that escaped the reads, someone's shell) is never signalled and keeps the cleanup open; the record keeps its pid, start time, program name (the name of its executable as the kernel recorded it at exec, `ps -o ucomm`, and only while the pid has the start time recorded; never `comm`, which on macOS is the process's own argv[0] and can carry its arguments; omitted when it cannot be read) and working directory, never its arguments. The record says `clean`, `clean_with_fallback` or `incomplete_or_unknown` (`cleanup`, with the normal shutdown's errors, every signal sent, what is still running and what is unresolved), and `cleanup_complete` is true only for the first two. Evidence follows the teardown: the transcript prefix (`transcriptBytes`, `transcriptSha256`, with the session and time it was taken) and the patch. The sibling read locks come off and the protected inputs become readable again only when the cleanup is complete; otherwise both stay, the record goes to `recovery/` beside the run's restoration ledger (which lists every arm's recorded processes and the runner's own identity), the cohort stops, and `bun scripts/benchmarks/restore.ts --run <run>` (also `runner.py restore`, which calls it) puts the modes back and takes back the trust entry once the runner, every recorded process and anything with a fixture in its argv or as its working directory are gone, then moves the kept records into `runs/`, where grading reads them as unavailable. A ledger written before runner identities needs `--runner-exited`, the operator's statement that the runner is gone. Runner commands are bounded (a hung one is killed), and SIGHUP stops the run like SIGINT and SIGTERM. `restoration.json` says `restored: true` only when the protected inputs and every arm's sibling locks are back. The trust entry is taken back with them (by the recovery when the cleanup is incomplete), and kept when the user changed it meanwhile; when the runner's own write never landed (it stopped between recording the lease and the rename), the outcome is `not_written` and nothing is taken back; for a runner that died there, the recovery takes back only an entry exactly as the runner would have written it, and none while the runner's temp file is still there. Orca registrations are not removed during benchmark teardown. Legacy benchmark-created entries may be removed only after provenance confirms the exact fixture path, using `orca project setup-delete --setup <id>` after verifying the exact repo and fixture path; unrelated registrations and fixture/Git/evidence data remain untouched. The ledger reports a record kept in `recovery/` as an attempt with `withheld: true` (unavailable: its cleanup was incomplete, or its sibling read locks could not be put back) until `restore.ts` moves it, never as missing (a record in both places, a move that did not finish, counts once and withheld); while `runs/` is still locked (a cleanup is incomplete), the attempts it cannot read there are listed as `unreadable`, never as missing, and its teardown row carries the normal shutdown's errors (`normal_errors`, for example a hub project registration left behind). The recovery also removes the temp files a runner left (one that died mid trust write or mid restore, or could not remove its own), each a copy of `~/.claude.json`; a write the runner knew never landed stays `not_written` there too, and no entry is touched. Run records carry `completion`, `cleanup`, `restoration` and `stages` (completion wait, shutdown, settle, fallback with the final readback, evidence and restoration times), apart from the end reason: `end_reason_detail` keeps the original classification, and `end_flags` (an unverified model, modified metadata, a tree changed or unverified after the active time) makes any end other than an interruption, a provider quota error or a budget pause an infrastructure error beside it, which grading and the ledger report as `end_story`; a record with `teardown_errors` (a patch, transcript prefix, last capture, fixture metadata or event log that could not be taken, or sibling read locks or the trust entry that could not be put back, a concurrent change of the trust entry included), an incomplete cleanup or a trust entry not taken back is unavailable to grading and to the ledger alike (a record from before 0.12.5 is judged as it was then, by its own `cleanup_complete` and `trust_restored`, and the ledger shows its teardown with `verified: false`: 0.12.3 and 0.12.4 set `cleanup_complete` when the shutdown steps reported success (commands, terminal close, lock and trust restores), with no process readback, which is no proof that the processes were gone); runner commands, the `ps` reads included, run in process groups of their own, so a Ctrl-C reaches the runner, which stops in this order, and not the command it is running; the completion wait's tree hash never writes the agent's index. `teardown.ts` and the process table it reads (`src/hub/child-process.ts`) are pinned in `prepared.json`, and `teardown.ts` in `cohort.json` and `grade.json` too, with the other runner sources.
|
|
56
|
+
Every run record, diff, PTY/provider trace and evaluation artifact stays in the mode-0700 private run directory. Setup errors, structured provider quota errors, hub budget pauses, unsettled deliveries and interruptions stay separate and unscored. A quota snapshot or numeric `429` alone is not a provider error. On exit the runner restores the exact input and read-lock modes, the sibling artifact modes and the scoped Claude trust flag (removing its own fresh Claude trust entry), under the condition below. Teardown (issue #113) proves what it stops. Each process an arm starts is recorded with its pid and start time and the evidence that it is the arm's: the daemon by the pid in its state directory and an argv that serves this fixture (a later daemon there is a replacement, never adopted), Claude's launch chain by the arm's own session id as its `--session-id` (and the Orca terminal shell that runs it, working in the fixture), the Codex app-server as the daemon's child. What runs below them, and what is in a group a recorded process leads while that group is known to be the same one (its leader alive, or a recorded member still in it), is the arm's too; the table is read every 5 s while the agents work and during the completion wait, so a tool command's background job is recorded while its parent runs (one that detaches between two reads, outside the fixture and without it in its argv, is not seen). A fixture name in an argv is never proof on its own (Codex's need not name the fixture), and every recorded actor counts even when a read of the table fails. Teardown pauses every agent, reads the deliveries still owed, lets a completed arm's Claude end the turn it is in (up to 30 s and never past the 300 s limit (the bound is taken when the wait starts; one 250 ms poll can pass it), judged by the `turn_duration` row Claude Code writes after a turn, proved on the arm's probe turn; a stopped or timed-out arm does not wait; a completed arm's tree is hashed at the end of its active time and again after the teardown, and a write in between flags the attempt invalid), asks for the normal shutdown (Claude's terminal closed, `ahub kill`, the hub's project registration removed (`ahub projects remove`); the runner leaves Orca registrations untouched; a step of it that fails is recorded in `cleanup.normal.errors` and does not by itself fail the attempt, since the process readback decides, so a hub project registration left behind is listed there), and reads the process table back, in the C locale and in UTC (start times are read as identities, the same in every reader). What is still running after a settle period gets SIGTERM; what is left then is frozen (SIGSTOP), the table read again, and it is killed (SIGKILL), what was frozen and not killed being continued: a frozen process starts nothing, so that read sees all of it. Every signal goes only to an identity read again just before, by the group the table shows it in now; a group leader is signalled with its group. What the read after the freeze shows for the first time is frozen too and read again before anything is killed, and when such a read fails only what a STOP reached is killed, by the last read that showed it (a stopped process keeps its pid). Anything with the fixture in its argv or as its working directory that is not proved the arm's (a background job that escaped the reads, someone's shell) is never signalled and keeps the cleanup open; the record keeps its pid, start time, program name (the name of its executable as the kernel recorded it at exec, `ps -o ucomm`, and only while the pid has the start time recorded; never `comm`, which on macOS is the process's own argv[0] and can carry its arguments; omitted when it cannot be read) and working directory, never its arguments. The record says `clean`, `clean_with_fallback` or `incomplete_or_unknown` (`cleanup`, with the normal shutdown's errors, every signal sent, what is still running and what is unresolved), and `cleanup_complete` is true only for the first two. Evidence follows the teardown: the transcript prefix (`transcriptBytes`, `transcriptSha256`, with the session and time it was taken) and the patch. The sibling read locks come off and the protected inputs become readable again only when the cleanup is complete; otherwise both stay, the record goes to `recovery/` beside the run's restoration ledger (which lists every arm's recorded processes and the runner's own identity), the cohort stops, and `bun scripts/benchmarks/restore.ts --run <run>` (also `runner.py restore`, which calls it) puts the modes back and takes back the trust entry once the runner, every recorded process and anything with a fixture in its argv or as its working directory are gone, then moves the kept records into `runs/`, where grading reads them as unavailable. A ledger written before runner identities needs `--runner-exited`, the operator's statement that the runner is gone. Runner commands are bounded (a hung one is killed), and SIGHUP stops the run like SIGINT and SIGTERM. `restoration.json` says `restored: true` only when the protected inputs and every arm's sibling locks are back. From before the runner locks its first input until it writes that outcome, it says `restored: false` and names the runner (#120), so a runner that is killed sends the recovery in, and the recovery waits while that runner still runs. The runner refuses a run directory whose earlier run is not restored (its `restoration.json` or its ledger says so) first, before it reads its inputs or any mode, and again right before it locks anything, since locked modes would become the originals; `restore.ts` first. One runner per directory at a time: two started on one directory at the same moment are not guarded against. `restore.ts` trusts `restored: true` only when the restoration ledger agrees (a ledger from before 0.12.5 still has its locks restored; when `restoration.json` says `restored: true` or names a runner, its trust entry is kept as it was then), and checks the runners again right before it restores anything. The trust entry is taken back with them (by the recovery when the cleanup is incomplete), and kept when the user changed it meanwhile; when the runner's own write never landed (it stopped between recording the lease and the rename), the outcome is `not_written` and nothing is taken back; for a runner that died there, the recovery takes back only an entry exactly as the runner would have written it, and none while the runner's temp file is still there. Orca registrations are not removed during benchmark teardown. Legacy benchmark-created entries may be removed only after provenance confirms the exact fixture path, using `orca project setup-delete --setup <id>` after verifying the exact repo and fixture path; unrelated registrations and fixture/Git/evidence data remain untouched. The ledger reports a record kept in `recovery/` as an attempt with `withheld: true` (unavailable: its cleanup was incomplete, or its sibling read locks could not be put back) until `restore.ts` moves it, never as missing (a record in both places, a move that did not finish, counts once and withheld); while `runs/` is still locked (a cleanup is incomplete), the attempts it cannot read there are listed as `unreadable`, never as missing, and its teardown row carries the normal shutdown's errors (`normal_errors`, for example a hub project registration left behind). The recovery also removes the temp files a runner left (one that died mid trust write or mid restore, or could not remove its own), each a copy of `~/.claude.json`; a write the runner knew never landed stays `not_written` there too, and no entry is touched. Run records carry `completion`, `cleanup`, `restoration` and `stages` (completion wait, shutdown, settle, fallback with the final readback, evidence and restoration times), apart from the end reason: `end_reason_detail` keeps the original classification, and `end_flags` (an unverified model, modified metadata, a tree changed or unverified after the active time) makes any end other than an interruption, a provider quota error or a budget pause an infrastructure error beside it, which grading and the ledger report as `end_story`; a record with `teardown_errors` (a patch, transcript prefix, last capture, fixture metadata or event log that could not be taken, or sibling read locks or the trust entry that could not be put back, a concurrent change of the trust entry included), an incomplete cleanup or a trust entry not taken back is unavailable to grading and to the ledger alike (a record from before 0.12.5 is judged as it was then, by its own `cleanup_complete` and `trust_restored`, and the ledger shows its teardown with `verified: false`: 0.12.3 and 0.12.4 set `cleanup_complete` when the shutdown steps reported success (commands, terminal close, lock and trust restores), with no process readback, which is no proof that the processes were gone); runner commands, the `ps` reads included, run in process groups of their own, so a Ctrl-C reaches the runner, which stops in this order, and not the command it is running; the completion wait's tree hash never writes the agent's index. `teardown.ts` and the process table it reads (`src/hub/child-process.ts`) are pinned in `prepared.json`, and `teardown.ts` in `cohort.json` and `grade.json` too, with the other runner sources.
|
|
57
57
|
|
|
58
58
|
## Manifest v2: the turn-free arm
|
|
59
59
|
|
package/docs/operations.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Operations guide
|
|
2
2
|
|
|
3
|
-
This guide describes ahub 0.12.
|
|
3
|
+
This guide describes ahub 0.12.7 and control protocol 13. Live verification
|
|
4
4
|
results and remaining prerequisites are recorded separately in [the smoke ledger](smoke.md).
|
|
5
5
|
|
|
6
6
|
## Install and start
|
|
@@ -647,21 +647,21 @@ Rows without a live process are stale registrations; forget them with
|
|
|
647
647
|
|
|
648
648
|
Upgrade running projects with the target release's own coordinator. It accepts
|
|
649
649
|
a running source on control protocol 9 (0.6.x), 10 (0.7.0 through 0.12.0),
|
|
650
|
-
11 (0.12.1 and 0.12.2), 12 (0.12.3) or 13 (0.12.4
|
|
650
|
+
11 (0.12.1 and 0.12.2), 12 (0.12.3) or 13 (0.12.4 through 0.12.6) and only
|
|
651
651
|
a target on its own protocol, so the target's coordinator fits every supported
|
|
652
652
|
source and carries every recovery fix released up to it. Protocol 8 and older
|
|
653
653
|
(0.5.x and earlier) are refused as `manual-bootstrap-required`. Run from the
|
|
654
654
|
project directory, without replacing the global CLI first:
|
|
655
655
|
|
|
656
656
|
```bash
|
|
657
|
-
bunx --package @staix/agent-hub@0.12.
|
|
658
|
-
bunx --package @staix/agent-hub@0.12.
|
|
657
|
+
bunx --package @staix/agent-hub@0.12.7 ahub upgrade --to 0.12.7 --dry-run
|
|
658
|
+
bunx --package @staix/agent-hub@0.12.7 ahub upgrade --to 0.12.7 --yes
|
|
659
659
|
```
|
|
660
660
|
|
|
661
661
|
| Running now | Coordinator to use |
|
|
662
662
|
| --- | --- |
|
|
663
663
|
| 0.6.x (protocol 9) | the target's, through `bunx` as above |
|
|
664
|
-
| 0.7.0 through 0.12.0 (protocol 10), 0.12.1 and 0.12.2 (protocol 11), 0.12.3 (protocol 12), 0.12.4
|
|
664
|
+
| 0.7.0 through 0.12.0 (protocol 10), 0.12.1 and 0.12.2 (protocol 11), 0.12.3 (protocol 12), 0.12.4 through 0.12.6 (protocol 13) | the target's, through `bunx` as above |
|
|
665
665
|
| any supported source, with the installed CLI already at the target | `ahub upgrade` below, which is the same coordinator |
|
|
666
666
|
| 0.5.x or earlier (protocol 8 and older) | not supported: bootstrap by hand with the matching CLI |
|
|
667
667
|
|
|
@@ -47,7 +47,15 @@ Cost of the group stops: each ACP or Pi stop now reads the process table twice (
|
|
|
47
47
|
| V | solo Codex with a visitor in the fixture | 31.0 s into the work | SIGTERM to the runner | `incomplete_or_unknown`, the visitor left alone and recorded as `program: sleep` | the visitor | withheld; `restore.ts` refused while it ran, restored once it left |
|
|
48
48
|
| S2 | advisory | native readiness | SIGINT to the group | `clean` (452 ms) | nothing | inputs, locks, trust |
|
|
49
49
|
|
|
50
|
-
Runs A, V and S2 were repeated on `fdc884e` with the same outcomes. Every one of these runs registered its fixtures with Orca, since the runner still added them then. The user has since directed that no fixture is added to Orca without an explicit registration (#117), and the registrations these runs left are removed under #117. Since `70b1daf` the runner only looks registrations up, and since `b03b035` it refuses an unregistered selection up front. A live check of the latter registered nothing. A run prepared afresh (`/private/tmp/ahub-0126-refusal`, 40 fixtures, none registered) was started with the runner of `b03b035`. It refused the whole run before anything was locked or recorded, naming every unregistered fixture. That lookup was later replaced by #118's read-only helper (`scripts/benchmarks/orca-workspace.ts`, merged into this branch from main), which refuses at the first fixture missing and checks each arm's identity against the preflight; #118 records its own checks, and this refusal check was not repeated on it. Orca's repository list was the same before and after (15 repositories), the upstream checkout kept its mode, and no attempt record was written.
|
|
50
|
+
Runs A, V and S2 were repeated on `fdc884e` with the same outcomes. Every one of these runs registered its fixtures with Orca, since the runner still added them then. The user has since directed that no fixture is added to Orca without an explicit registration (#117), and the registrations these runs left are removed under #117. Since `70b1daf` the runner only looks registrations up, and since `b03b035` it refuses an unregistered selection up front. A live check of the latter registered nothing. A run prepared afresh (`/private/tmp/ahub-0126-refusal`, 40 fixtures, none registered) was started with the runner of `b03b035`. It refused the whole run before anything was locked or recorded, naming every unregistered fixture. That lookup was later replaced by #118's read-only helper (`scripts/benchmarks/orca-workspace.ts`, merged into this branch from main), which refuses at the first fixture missing and checks each arm's identity against the preflight; #118 records its own checks, and this refusal check was not repeated on it. Orca's repository list was the same before and after (15 repositories), the upstream checkout kept its mode, and no attempt record was written. Before the release, no arm had run on the lookup-only runner, which needs fixtures registered explicitly first: the live arms above ran on the runner that still added them. Every record carries `platform: darwin`. The Codex skills condition was 177 user and 5 system skills in every arm above with Codex. Arms on the released runner follow.
|
|
51
|
+
|
|
52
|
+
After the release, with the user's authorization for the Orca registrations, three of these arms ran on the released runner (`44e04e5`, Bun 1.3.14, which 0.12.6's CI used). Each run's four case-0 fixtures were registered by hand before the run and removed after it, and Orca's repository list ended as it began (16 repositories, the same ids; the one more than at the refusal check is an unrelated project added meanwhile). Every record carries `platform: darwin`; the Codex skills condition was 177 user and 4 system skills in each.
|
|
53
|
+
|
|
54
|
+
| Run | Arm | Stopped | Cleanup | Left running | Restoration |
|
|
55
|
+
|---|---|---|---|---|---|
|
|
56
|
+
| A | solo Codex | SIGTERM to the runner, 30.8 s into the work | `clean` (313 ms) | nothing | inputs, sibling locks |
|
|
57
|
+
| B | advisory | SIGINT to the group, 40.7 s into the work | `clean` (447 ms) | nothing | inputs, locks, trust |
|
|
58
|
+
| V | solo Codex with a visitor in the fixture | SIGTERM to the runner, 30.6 s into the work | `incomplete_or_unknown`, the visitor left alone and recorded as `program: sleep` | the visitor | withheld; `restore.ts` refused while it ran, restored once it left |
|
|
51
59
|
|
|
52
60
|
## Review
|
|
53
61
|
|
|
@@ -0,0 +1,58 @@
|
|
|
1
|
+
# Agent-hub 0.12.7 verification
|
|
2
|
+
|
|
3
|
+
Implementation and candidate verification for issue #120 (the CooperBench run's restoration state), performed on 2026-10-03, with Bun 1.4.2 (#121). Native CLIs: Codex 0.160.0, Claude Code 2.1.288. The hub runtime is unchanged from 0.12.6 apart from the plugin bundle, which is built with Bun 1.4.2. Production application is verified separately after publication.
|
|
4
|
+
|
|
5
|
+
## What changed
|
|
6
|
+
|
|
7
|
+
- The runner writes `restoration.json` as `{restored: false, runner}` right before it locks anything, and its outcome replaces it at the end. A runner killed in between leaves a file that sends the recovery in.
|
|
8
|
+
- The runner refuses a run directory whose earlier run is not restored (`reuseProblem`), before it reads any mode.
|
|
9
|
+
- `restore.ts` trusts `restored: true` only when the restoration ledger agrees (`unrestored`). It waits while a runner named by either file still runs. With no ledger, nothing was locked; with neither file, there is nothing to recover.
|
|
10
|
+
|
|
11
|
+
## Automated gate
|
|
12
|
+
|
|
13
|
+
`scripts/check.sh` on Bun 1.4.2, at the head of the pull request, with the fixes of three review rounds: 759 tests passed, 0 failed, 3,848 expectations across 68 files in 132 s; `check: OK` (757 and 3,833 before them). Mutation checks (each test fails with its part removed): the old early return on `restored: true`, the runner named by `restoration.json` ignored, and a ledger required.
|
|
14
|
+
|
|
15
|
+
## Live checks
|
|
16
|
+
|
|
17
|
+
Orca lookups in these runs were answered by a read-only stand-in for the `orca` CLI. It answered `repo list` and `worktree list` for the run's own fixtures, passed `worktree current` to the real Orca, and refused everything else. Nothing was registered in Orca. The arms were real: Codex 0.160.0 through the hub, with the inputs and sibling artifacts really locked.
|
|
18
|
+
|
|
19
|
+
- **SIGKILL mid-arm** (`/private/tmp/ahub-0127-K2`, solo Codex, killed 30 s into the work):
|
|
20
|
+
- Before the kill, `restoration.json` was `restored: false` and named the runner. It stayed so after the kill. The protected root and a sibling fixture were mode 000, and the arm's daemon and Codex app-server were still running.
|
|
21
|
+
- A re-run in the same directory was refused (exit 1), but by its first read of the locked private inputs (a bare EACCES), before the reuse check. The review moved the check first; see K4.
|
|
22
|
+
- `restore.ts` refused while the daemon, the app-server and what ran below them were running (exit 1, each named). After `ahub kill` it restored the run (exit 0): the modes were back, and `restoration.json` said `recovered`.
|
|
23
|
+
- **Reuse refused** (`/private/tmp/ahub-0127-K3`, a fresh preparation whose `restoration.json` said `restored: false`): the real runner refused it with "an earlier run here is not restored (restoration.json)". It refused before it read any mode or wrote a ledger or `cohort.json`; the protected root's mode was unchanged. `restore.ts` then found no ledger and a runner that was gone, and marked the directory restored.
|
|
24
|
+
- **SIGKILL mid-arm again, after the review** (`/private/tmp/ahub-0127-K4`, the same run as K2; its log names `58c8db0`, since the fixes were not committed yet, but its `prepared.json` and `cohort.json` pin the sources of `bdc5581`): the re-run in the same directory was now refused by the reuse check, naming `restore.ts` ("an earlier run here is not restored (restoration.json): run bun scripts/benchmarks/restore.ts --run ... first"). The recovery was held while the arm's processes ran, and restored the run after `ahub kill`.
|
|
25
|
+
- **SIGKILL mid-arm on the final runner** (`/private/tmp/ahub-0127-K5`, after the third review round; its log names `ff2480e`, since the fixes were not committed yet, but its `prepared.json` pins the final `native.ts` and `teardown.ts`; and so does its `cohort.json`; `restore.ts` is not pinned, and this run's ledger names its runner, so `owed` returns it unchanged and the run does not exercise the pre-0.12.5 rule): the same outcome as K4. The marker stayed, the re-run was refused by the reuse check naming `restore.ts`, the recovery was held while the arm's processes ran, and it restored the run after `ahub kill`.
|
|
26
|
+
- **A first attempt** (`/private/tmp/ahub-0127-K`) stopped before the marker, because the stand-in did not yet answer `worktree current`. Nothing was locked or written. The recovery of that directory, which had neither file, blocked on "the ledger does not name the runner"; it now reports nothing to recover (a test covers it).
|
|
27
|
+
- **Upgrade:** `scripts/smoke-recovery-09-10.ts` recovered the published 0.12.6 package into the candidate. Both run protocol 13, and the source's integrity came from the registry. Operation `b7643429-7f0c-4f70-8128-857bb965a5d2`: the queue was preserved, and the task digest is the same as in 0.12.6's check.
|
|
28
|
+
|
|
29
|
+
## Known limits
|
|
30
|
+
|
|
31
|
+
- One runner per run directory at a time: the reuse check runs again right before the marker, but two runners started on one directory at the same moment can both pass it (a `ponytail:` in `native.ts` names the exclusive claim that would close it).
|
|
32
|
+
- Grading still trusts `restoration.json` alone. With the marker, a run whose runner died says `restored: false`, so grading refuses it; a stale `restored: true` can only come from a run before 0.12.7.
|
|
33
|
+
- Deferred, as listed in #116: a recovery of a 0.12.3 or 0.12.4 ledger killed mid restore leaves a temp file no later recovery looks for, and a 0.12.5 runner whose rename threw left its trust temp file behind a `restored: true` run.
|
|
34
|
+
|
|
35
|
+
## Review
|
|
36
|
+
|
|
37
|
+
Independent read-only reviews (Claude Code OCR delegation), each on the head of the time. The first, on `58c8db0`, went through every state a run directory can be in, and through a runner and a recovery started together. It found no Critical or Important defect. Its Minor findings were fixed:
|
|
38
|
+
|
|
39
|
+
- The reuse check runs first, before the runner resolves its inputs, so a directory whose inputs a dead runner locked is refused with a message that names `restore.ts`, not a bare EACCES (K4). It runs again right before the marker. The limit for two runners started together is recorded above.
|
|
40
|
+
- `restore.ts` reads the process table before the files, so the runners the files name are checked again right before anything is restored. A test with a runner missing from the first table covers it.
|
|
41
|
+
- A ledger from before 0.12.5 (no runner identity) under `restored: true` keeps its trust entry as it was then. Those versions left a concurrently changed trust entry at `written`, and taking it back would remove the user's entry. A test covers it.
|
|
42
|
+
- `unrestored` says a `not_written` entry whose temp file is left is still unrestored, which is what the code did; the spec and the comment now say so, with table rows.
|
|
43
|
+
- The runner's outcome uses `unrestored`, so it and the recovery's agreement check cannot disagree. Smaller wording fixes.
|
|
44
|
+
|
|
45
|
+
The second round, on `bdc5581`, found no Critical or Important defect. Its Minor finding was fixed. The exception for ledgers from before 0.12.5 covered the whole ledger, so a re-run that died over such a directory kept its inputs locked. Meanwhile the reuse check sent the operator to a recovery that then did nothing. Now only the trust entry is kept as it was then (`owed` in `teardown.ts`); the runner's reuse check and the recovery use the same rule, and the locks are restored. A test covers it and fails with the earlier rule. Its nits were fixed too:
|
|
46
|
+
|
|
47
|
+
- Nothing is written to the run directory before the marker, so a run that died between the two checks keeps its `cohort.json`.
|
|
48
|
+
- The unused table parameter is gone.
|
|
49
|
+
- The reasons read "still unrestored: ...".
|
|
50
|
+
|
|
51
|
+
The third round, on `ff2480e`, found no Critical or Important defect. Its Minor finding was fixed: a 0.12.7 runner that accepted a directory with a pre-0.12.5 ledger, wrote its marker and died before a ledger of its own left that old ledger under a marker. The recovery then took the old trust entry back, because the rule keyed on `restored: true`. `owed` now also applies under a marker that names a runner; a test covers it and fails without it. Its nits were fixed too:
|
|
52
|
+
|
|
53
|
+
- The rule lives in `owed` alone, and the recovery settles the trust entry `owed` returns.
|
|
54
|
+
- The run directory's writes are inside the `try`, so a failure there still ends with an outcome.
|
|
55
|
+
- `AGENTS.md` names `owed`, and the spec wording is fixed.
|
|
56
|
+
- K5 ran the final runner.
|
|
57
|
+
|
|
58
|
+
The fourth round, on `68a0f9b`, found no Critical or Important defect. It ran the reuse check and the recovery over 35 directory shapes, and in none did the reuse check accept a directory where the recovery then restored anything. Its nits, all wording, were fixed: the `AGENTS.md` rule, the CHANGELOG and the docs now state the condition under which the old trust entry is kept, and K5's provenance is noted.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "agent-hub",
|
|
3
|
-
"version": "0.12.
|
|
3
|
+
"version": "0.12.7",
|
|
4
4
|
"description": "Channel between Claude Code and the agent-hub daemon: peer messages from Codex, Kimi and the local worker arrive as channel events; hub_send replies.",
|
|
5
5
|
"author": {
|
|
6
6
|
"name": "Young Joon Lee",
|