@staix/agent-hub 0.12.4 → 0.12.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -2,6 +2,28 @@
2
2
 
3
3
  Issue and pull request numbers in the entries for 0.7.7 and earlier refer to the previous repository, archived on 2026-09-30 when this repository's history was rewritten; the one exception is the open smoke-check issue, formerly #12, which moved here as #1. Numbers in newer entries refer to this repository.
4
4
 
5
+ ## 0.12.6
6
+
7
+ - Agents the hub spawns through ACP (Kimi) and Pi run in their own process group and are stopped as one, as the Codex app-server is since 0.12.5. A stop with or without a group drops the child's pipes, so a process it left cannot keep the hub alive. Once a launcher has exited, its group is followed until it is gone, and members still finishing their exit get a bound before the stop fails. A Codex start that fails reports its own error. The group stop finishes as soon as the leader has exited and a read shows nothing left. As for Codex since 0.12.5, a Kimi or Pi that crashed and left processes in its group (an MCP server, a tool command) is not restarted until they are gone: the error names their pids (#115).
8
+ - Split predictions (shadow only): busy counts as taking the task in question only while the peer is still in the turn that task started (delivered at once to the idle peer) or in which it claimed the task. A task queued, held or steered into a turn about something else is not taken, a turn that ended takes its hand-overs with it, and a routed peer busy when the record is taken is not available (#109, #115).
9
+ - Claude Code's version is read from the transcript again only when the transcript changed (#115).
10
+ - Tests: `scripts/check.sh` runs `bun test` with a 20 s timeout and under `scripts/hang-watch.sh`, which samples and stops a run past its bound. In Bun 1.3.14 a test timeout that fires while `Bun.spawnSync` runs can start the next test inside spawnSync's event loop, where another spawnSync then spins for good: the stacks of two local hangs under load show it, and the macOS CI hang of 0.12.5 matches them (minimal reproductions did not hang). The permission test no longer depends on runner speed (#115).
11
+ - CooperBench runner and ledger (#115):
12
+ - The runner never adds or removes an Orca registration, which needs the user's explicit authorization: every fixture of the selection is looked up read-only before anything is changed, a missing or ambiguous one refuses the run, and each arm checks again, before touching its fixture, that its identity is the one preflighted (#117, #118).
13
+ - A trust write that never landed is `not_written`, not `changed_concurrently`, also after a failed restore and in the recovery; its temp files are removed in the runner and the recovery (also one a runner left when it died in its own restore), one that cannot be removed keeps the trust entry open for the next recovery. A write the runner knows never landed is `not_written` on every path, and neither the runner nor the recovery then touches an entry; for a runner that died mid write, the recovery takes back only an entry exactly as the runner would have written it; one the user changed meanwhile is never taken back.
14
+ - The ledger shows normal shutdown errors, and reports records kept in `recovery/` as withheld attempts, not missing; a record caught in the middle of the recovery's move counts once, as withheld, and attempts a still-locked `runs/` hides are `unreadable`, not missing, in the per-arm summary too, where an arm with no record at all is listed with what it owes.
15
+ - Records carry the platform, and the runner refuses non-macOS.
16
+ - An unresolved process's program name is its executable's name as the kernel recorded it (`ps -o ucomm`), checked against its start time, never its arguments; it is omitted when it cannot be read.
17
+ - The end reason is one function, and shared helpers are not duplicated.
18
+ - CooperBench manifests v2 and the #106 ablation pin hub 0.12.6.
19
+
20
+ ## 0.12.5
21
+
22
+ - The Codex app-server runs in its own process group and is stopped as one, with what it started in groups of its own (MCP servers, tool commands): after SIGTERM, to its group and to each recorded process that leads a group of its own, and a grace period in which what it starts is recorded, the tree is frozen, read again and killed, and the stop is done only when the table shows none of it (or, when no table can be read at all, when its own group is gone); a launcher that exited before the stop while its group still has members fails the stop, its group unsignalled. `codex` is a node launcher, and mid-turn the native app-server does not exit on SIGTERM: the SIGKILL that followed reached the launcher alone, leaving the app-server at work under init and the hub process alive after `ahub kill` reported it stopped. The CooperBench runner's teardown is verified (#113): every process an arm starts that the reads see is recorded with its pid, start time and the evidence that it is the arm's (a fixture name in an argv is never proof on its own), read again every 5 s while the agents work and while Claude ends its turn (one that detaches between two reads, outside the fixture and without it in its argv, is not seen); a process with a fixture in its argv or as its working directory that is not proved the arm's is never signalled and keeps the cleanup open; teardown pauses the agents, lets a completed arm's Claude end its turn (up to 30 s, by the transcript's `turn_duration` row proved on the probe turn), asks for the normal shutdown, reads the process table back in the C locale and in UTC (a start time is part of a process's identity), signals only re-read identities, and records `clean`, `clean_with_fallback` or `incomplete_or_unknown` apart from the end reason. Evidence is taken after it, and the read locks and protected inputs come off only after a complete cleanup (otherwise `scripts/benchmarks/restore.ts` does it later, once the runner, every recorded process and anything in the fixture are gone). Run records carry the completion, cleanup, restoration, stage times and a summary of the Codex skills `skills/list` reports (its answer itself is not kept); a record from before 0.12.5 is judged by its own cleanup and trust flags, as then, and its teardown is shown as not verified (0.12.3 and 0.12.4 recorded only that the shutdown steps reported success, with no process readback); `teardown.ts` is pinned with the runner sources; runner commands run in their own process groups.
23
+ - Turn-free facts: files seen under a named directory beyond the 200 followed have a notice of their own; it and the touched-limit notice are spent only when an offer carrying them is read back (a name a PII pattern matches is counted, never named; an offer read back after its file was covered again and dropped again does not spend the later notice), and a new session hears only the drops of its own; it also starts its touched list afresh. A directory file rewritten with HEAD's bytes, which git lists until it refreshes its index, is neither named nor counted against `current()`; a file too large to read is compared by git's blob id of its raw bytes while all such files in one comparison total 64 MiB or less, and named otherwise (git's filters are not applied, so an end-of-line conversion or LFS makes it read as changed). The `fact` event carries `named` (directory files named without a diff), and the ledger counts them (#112).
24
+ - Split predictions (shadow only, routing unchanged): every hand-over records the new owner's profile (hub version, agent version from app-server or Claude Code's transcript, coordination mode); the other owner of an overlapping task not started yet counts as available while it is busy taking it (an owner goes busy as its task is delivered); observations accumulate across hub runs, one per hand-over (a decline or an escalation away counts against the peer that failed), and count only with the peer's current profile. A prediction is recorded when routing chooses the first owner of a task overlapping another owner's task not started yet (`where: "routing"`, the calibration record; `route explain` shows the same pair); cohort-time records carry `where: "cohort"` (#109).
25
+ - CooperBench manifest v2 and the #106 ablation pin hub 0.12.5, Codex 0.160.0 and Claude Code 2.1.288.
26
+
5
27
  ## 0.12.4
6
28
 
7
29
  - A completed-change or edit-conflict notice is checked again for each recipient right before it is handed over, after condensation: a copy whose task has closed, changed owner or is gone is dropped instead of starting a turn, recorded as discarded and as a `stale` event. Other recipients and unrecorded envelopes are delivered as before (#106).
@@ -37,6 +37,8 @@ python3 scripts/benchmarks/runner.py prepare \
37
37
  --upstream-root /tmp/agent-hub-cooperbench-upstream \
38
38
  --output /private/tmp/ahub-0123-case0
39
39
  chmod 700 /private/tmp/ahub-0123-case0
40
+ # Orca registrations need the user's explicit authorization (see below): the runner never adds or removes one
41
+ # (#117) and refuses the run while any selected fixture is not registered.
40
42
  bun scripts/benchmarks/native.ts \
41
43
  --run /private/tmp/ahub-0123-case0 \
42
44
  --private-inputs /private/tmp/ahub-0123-private-inputs \
@@ -49,17 +51,17 @@ bun scripts/benchmarks/native.ts \
49
51
  --cases 0
50
52
  ```
51
53
 
52
- The runner registers each exact fixture root with Orca, uses the project CLI to start and stop each daemon, and requires exact worktree/cwd readback for native sessions. Claude starts through the canonical `ahub claude` guard and loads this checkout's candidate bundle through an exact session-only `--mcp-config` server (`server:agent-hub`); it does not promote or mutate the globally installed plugin. Codex uses the native app-server adapter and `workspace-write` sandbox. Each agent must execute a setup-only `head -c 1` probe against the exact protected file and produce only the denied marker before scored tasks begin. There is one native sandbox layer per agent.
54
+ The native runner runs on macOS only and refuses anything else (each run record says `platform`): the arms run in Orca terminals, and on Linux a clock step moves the start times its teardown proves processes by. Before native execution, obtain explicit user authorization before changing Orca registrations; a benchmark request alone is not authorization. Only after that authorization, an operator manually registers each exact fixture root with `orca repo add --path '<fixture>'` and confirms `orca worktree list --repo id:<repo-id>` shows that exact path. The runner performs read-only exact-path repo/worktree lookup; it preflights every selected fixture before changing fixture files, input modes, cohort or attempt records. If either identity is missing or mismatched, it exits nonzero with setup guidance and leaves the prepared run available for retry after explicitly authorized registration. Each arm rechecks its identity before changing fixture files. Benchmark and agent workflows never add or remove Orca registrations. The runner uses the project CLI to start and stop each daemon, and requires exact worktree/cwd readback for native sessions. Claude starts through the canonical `ahub claude` guard and loads this checkout's candidate bundle through an exact session-only `--mcp-config` server (`server:agent-hub`); it does not promote or mutate the globally installed plugin. Codex uses the native app-server adapter and `workspace-write` sandbox. Each agent must execute a setup-only `head -c 1` probe against the exact protected file and produce only the denied marker before scored tasks begin. There is one native sandbox layer per agent.
53
55
 
54
- Every run record, diff, PTY/provider trace and evaluation artifact stays in the mode-0700 private run directory. Setup errors, structured provider quota errors, hub budget pauses, unsettled deliveries and interruptions stay separate and unscored. A quota snapshot or numeric `429` alone is not a provider error. The runner restores exact input/read-lock modes and the scoped Claude trust flag (removing its own fresh project entry) and sibling artifact modes on exit. It stops only its recorded Claude terminal, hub project and hub terminal handles. Orca currently has no repo removal command, so inactive exact-path fixture repos remain registered after a run; one case leaves one record per arm.
56
+ Every run record, diff, PTY/provider trace and evaluation artifact stays in the mode-0700 private run directory. Setup errors, structured provider quota errors, hub budget pauses, unsettled deliveries and interruptions stay separate and unscored. A quota snapshot or numeric `429` alone is not a provider error. On exit the runner restores the exact input and read-lock modes, the sibling artifact modes and the scoped Claude trust flag (removing its own fresh Claude trust entry), under the condition below. Teardown (issue #113) proves what it stops. Each process an arm starts is recorded with its pid and start time and the evidence that it is the arm's: the daemon by the pid in its state directory and an argv that serves this fixture (a later daemon there is a replacement, never adopted), Claude's launch chain by the arm's own session id as its `--session-id` (and the Orca terminal shell that runs it, working in the fixture), the Codex app-server as the daemon's child. What runs below them, and what is in a group a recorded process leads while that group is known to be the same one (its leader alive, or a recorded member still in it), is the arm's too; the table is read every 5 s while the agents work and during the completion wait, so a tool command's background job is recorded while its parent runs (one that detaches between two reads, outside the fixture and without it in its argv, is not seen). A fixture name in an argv is never proof on its own (Codex's need not name the fixture), and every recorded actor counts even when a read of the table fails. Teardown pauses every agent, reads the deliveries still owed, lets a completed arm's Claude end the turn it is in (up to 30 s and never past the 300 s limit (the bound is taken when the wait starts; one 250 ms poll can pass it), judged by the `turn_duration` row Claude Code writes after a turn, proved on the arm's probe turn; a stopped or timed-out arm does not wait; a completed arm's tree is hashed at the end of its active time and again after the teardown, and a write in between flags the attempt invalid), asks for the normal shutdown (Claude's terminal closed, `ahub kill`, the hub's project registration removed (`ahub projects remove`); the runner leaves Orca registrations untouched; a step of it that fails is recorded in `cleanup.normal.errors` and does not by itself fail the attempt, since the process readback decides, so a hub project registration left behind is listed there), and reads the process table back, in the C locale and in UTC (start times are read as identities, the same in every reader). What is still running after a settle period gets SIGTERM; what is left then is frozen (SIGSTOP), the table read again, and it is killed (SIGKILL), what was frozen and not killed being continued: a frozen process starts nothing, so that read sees all of it. Every signal goes only to an identity read again just before, by the group the table shows it in now; a group leader is signalled with its group. What the read after the freeze shows for the first time is frozen too and read again before anything is killed, and when such a read fails only what a STOP reached is killed, by the last read that showed it (a stopped process keeps its pid). Anything with the fixture in its argv or as its working directory that is not proved the arm's (a background job that escaped the reads, someone's shell) is never signalled and keeps the cleanup open; the record keeps its pid, start time, program name (the name of its executable as the kernel recorded it at exec, `ps -o ucomm`, and only while the pid has the start time recorded; never `comm`, which on macOS is the process's own argv[0] and can carry its arguments; omitted when it cannot be read) and working directory, never its arguments. The record says `clean`, `clean_with_fallback` or `incomplete_or_unknown` (`cleanup`, with the normal shutdown's errors, every signal sent, what is still running and what is unresolved), and `cleanup_complete` is true only for the first two. Evidence follows the teardown: the transcript prefix (`transcriptBytes`, `transcriptSha256`, with the session and time it was taken) and the patch. The sibling read locks come off and the protected inputs become readable again only when the cleanup is complete; otherwise both stay, the record goes to `recovery/` beside the run's restoration ledger (which lists every arm's recorded processes and the runner's own identity), the cohort stops, and `bun scripts/benchmarks/restore.ts --run <run>` (also `runner.py restore`, which calls it) puts the modes back and takes back the trust entry once the runner, every recorded process and anything with a fixture in its argv or as its working directory are gone, then moves the kept records into `runs/`, where grading reads them as unavailable. A ledger written before runner identities needs `--runner-exited`, the operator's statement that the runner is gone. Runner commands are bounded (a hung one is killed), and SIGHUP stops the run like SIGINT and SIGTERM. `restoration.json` says `restored: true` only when the protected inputs and every arm's sibling locks are back. The trust entry is taken back with them (by the recovery when the cleanup is incomplete), and kept when the user changed it meanwhile; when the runner's own write never landed (it stopped between recording the lease and the rename), the outcome is `not_written` and nothing is taken back; for a runner that died there, the recovery takes back only an entry exactly as the runner would have written it, and none while the runner's temp file is still there. Orca registrations are not removed during benchmark teardown. Legacy benchmark-created entries may be removed only after provenance confirms the exact fixture path, using `orca project setup-delete --setup <id>` after verifying the exact repo and fixture path; unrelated registrations and fixture/Git/evidence data remain untouched. The ledger reports a record kept in `recovery/` as an attempt with `withheld: true` (unavailable: its cleanup was incomplete, or its sibling read locks could not be put back) until `restore.ts` moves it, never as missing (a record in both places, a move that did not finish, counts once and withheld); while `runs/` is still locked (a cleanup is incomplete), the attempts it cannot read there are listed as `unreadable`, never as missing, and its teardown row carries the normal shutdown's errors (`normal_errors`, for example a hub project registration left behind). The recovery also removes the temp files a runner left (one that died mid trust write or mid restore, or could not remove its own), each a copy of `~/.claude.json`; a write the runner knew never landed stays `not_written` there too, and no entry is touched. Run records carry `completion`, `cleanup`, `restoration` and `stages` (completion wait, shutdown, settle, fallback with the final readback, evidence and restoration times), apart from the end reason: `end_reason_detail` keeps the original classification, and `end_flags` (an unverified model, modified metadata, a tree changed or unverified after the active time) makes any end other than an interruption, a provider quota error or a budget pause an infrastructure error beside it, which grading and the ledger report as `end_story`; a record with `teardown_errors` (a patch, transcript prefix, last capture, fixture metadata or event log that could not be taken, or sibling read locks or the trust entry that could not be put back, a concurrent change of the trust entry included), an incomplete cleanup or a trust entry not taken back is unavailable to grading and to the ledger alike (a record from before 0.12.5 is judged as it was then, by its own `cleanup_complete` and `trust_restored`, and the ledger shows its teardown with `verified: false`: 0.12.3 and 0.12.4 set `cleanup_complete` when the shutdown steps reported success (commands, terminal close, lock and trust restores), with no process readback, which is no proof that the processes were gone); runner commands, the `ps` reads included, run in process groups of their own, so a Ctrl-C reaches the runner, which stops in this order, and not the command it is running; the completion wait's tree hash never writes the agent's index. `teardown.ts` and the process table it reads (`src/hub/child-process.ts`) are pinned in `prepared.json`, and `teardown.ts` in `cohort.json` and `grade.json` too, with the other runner sources.
55
57
 
56
58
  ## Manifest v2: the turn-free arm
57
59
 
58
60
  `scripts/benchmarks/manifest-v2.json` keeps v1's cases, models and limits and adds a fourth arm, `hub-turnfree-codex-claude` (issue #110): the same two agents and assignment rotation as `hub-codex-claude`, with `coordination: "turn-free"` in the fixture's hub config and fixture instructions that tell the owners not to message each other. A v1 manifest still validates and prepares. Its cohorts are run, graded and reported with the release and runner sources they were prepared with: `native.ts` refuses a manifest whose hub version is not the checkout's (v1 pins 0.12.3), and `runner.py` refuses to grade a cohort whose runner sources differ from its own.
59
61
 
60
- Hooks are equal across arms: no arm runs the user's or a plugin's hooks, and no arm runs a status line (`disableAllHooks` turns it off in the other arms, so the turn-free arm leaves it out; Claude's quota reaches the hub in no arm). Claude starts with `--setting-sources project` and `--strict-mcp-config`; the solo and advisory arms set `disableAllHooks`, and the turn-free arm's session settings carry the hub's own hooks (before and after every tool call, and at Stop) and nothing else, because they are the treatment. Every Codex thread starts with `features.hooks` off: the turn-free arm's Codex boundary is the adapter's steer into the running turn, whose readback is the steered input coming back as a user message item. MCP servers are isolated too: Claude has only the hub's (`--strict-mcp-config`), and every Codex runs through a wrapper written into each arm's fixture that turns off the user's plugins, apps, sub-agents and turn-end notifier and disables each MCP server the user's config defines, by name, so it starts only the hub's; the user's config is not changed. Instructions: Claude reads the fixture's `AGENTS.md` through `--append-system-prompt-file` (Claude Code reads `CLAUDE.md`, not `AGENTS.md`), Codex as the project's `AGENTS.md`; Codex also reads the user's global `AGENTS.md`, which only a separate Codex home with its own login would leave out (the run records say so), and the user's Codex skills stay available to it (the wrapper leaves them on); both are the same in every arm. Each run record carries these `conditions`; its `events` carry the hub's `capability` readbacks.
62
+ Hooks are equal across arms: no arm runs the user's or a plugin's hooks, and no arm runs a status line (`disableAllHooks` turns it off in the other arms, so the turn-free arm leaves it out; Claude's quota reaches the hub in no arm). Claude starts with `--setting-sources project` and `--strict-mcp-config`; the solo and advisory arms set `disableAllHooks`, and the turn-free arm's session settings carry the hub's own hooks (before and after every tool call, and at Stop) and nothing else, because they are the treatment. Every Codex thread starts with `features.hooks` off: the turn-free arm's Codex boundary is the adapter's steer into the running turn, whose readback is the steered input coming back as a user message item. MCP servers are isolated too: Claude has only the hub's (`--strict-mcp-config`), and every Codex runs through a wrapper written into each arm's fixture that turns off the user's plugins, apps, sub-agents and turn-end notifier and disables each MCP server the user's config defines, by name, so it starts only the hub's; the user's config is not changed. Instructions: Claude reads the fixture's `AGENTS.md` through `--append-system-prompt-file` (Claude Code reads `CLAUDE.md`, not `AGENTS.md`), Codex as the project's `AGENTS.md`; Codex also reads the user's global `AGENTS.md`, which only a separate Codex home with its own login would leave out (the run records say so), and the user's Codex skills stay available to it (the wrapper leaves them on; `conditions.codex.skills` records what app-server's `skills/list` reports for the fixture: counts by scope, how many are enabled, and a hash of their names, never their bodies or paths, and the `skills/list` answer itself is not kept in `codexMessages`; Claude's Skill tool is denied); both are the same in every arm. Each run record carries these `conditions`; its `events` carry the hub's `capability` readbacks.
61
63
 
62
- Validity, decided by the grader and applied by the ledger alike: an attempt whose records show a hook or an MCP server that is not the hub's, or whose Claude transcript cannot be read, is unavailable. A turn-free attempt is valid only with its context paths working: both verified before its tasks, and none lost, no cohort lifted and none formed open while the agents worked; teardown comes after that and does not count. Whether the agents' plans overlapped, so that a cohort formed at all, is their doing after assignment and is not a condition: every turn-free attempt without a capability failure counts for the arm, and the ledger reports the treatment received and a median over treated attempts beside it. A capability failure while the agents work is the one exclusion after assignment, because #110 forbids reporting it as a turn-free run (AC3); such attempts are listed with their reasons, never dropped silently. Each attempt records its transcript's length and hash when it ends; validity and every Claude measure are read from that prefix (Claude Code may append rows after it exits: a response still being written at teardown is not counted, and that agent's settlement is then unknown), and a prefix that changed counts as unreadable.
64
+ Validity, decided by the grader and applied by the ledger alike: an attempt whose records show a hook or an MCP server that is not the hub's, or whose Claude transcript cannot be read, is unavailable. A turn-free attempt is valid only with its context paths working: both verified before its tasks, and none lost, no cohort lifted and none formed open while the agents worked; teardown comes after that and does not count. Whether the agents' plans overlapped, so that a cohort formed at all, is their doing after assignment and is not a condition: every turn-free attempt without a capability failure counts for the arm, and the ledger reports the treatment received and a median over treated attempts beside it. A capability failure while the agents work is the one exclusion after assignment, because #110 forbids reporting it as a turn-free run (AC3); such attempts are listed with their reasons, never dropped silently. Each attempt records its transcript's length and hash when it ends; validity and every Claude measure are read from that prefix (Claude Code may append rows after it exits: what it writes after the prefix was taken is never counted, the ledger reports its size as `late_append_bytes`, and a turn still open in the prefix leaves that agent's settlement unknown), and a prefix that changed counts as unreadable.
63
65
 
64
66
  `scripts/benchmarks/manifest-v2-ablation-106.json` is the #106 ablation: the advisory arm against the same arm with `experiments.stale_notices: "deliver"`, which turns stale-notice dropping off and changes nothing else, on the same release and conditions. Its attempts are counted apart from the four-arm plan.
65
67
 
@@ -71,7 +73,7 @@ Pass `--repeat <n>` for the n-th repeat of a case (0 for the first): the arm ord
71
73
  python3 scripts/benchmarks/ledger.py --run /private/tmp/ahub-0124-r1 [--run /private/tmp/ahub-0124-r2 ...]
72
74
  ```
73
75
 
74
- It reads each run record, the Claude transcript it names and the fixture's git history, and writes `ledger.json` into the first run directory with a `units` table that names the unit and coverage of every measure. Give `--run` once per directory to pool the repeats of one plan; an attempt a directory's `cohort.json` planned that wrote no record is listed as missing. Per attempt: completion (the runner's end reason and its detail, such as `wall-timeout`, and whether every task has a done); setup and active time; first candidate, completion intents, integration and check times, and each agent's settlement read from its own record (the end of its last native turn after the last done, turns started by late messages included; the last done itself when it did not work after it), and how long the agents could still write after the active time ended (`stopped_s`); usage per agent in task (to its last done on the board) and over the whole attempt: Codex turns, assistant messages, token-usage updates whose running total grew past the total before the window, and the token growth; Claude assistant messages (unique message ids), turns and the tokens their usage records; Claude's main-loop requests by request id (side requests and retries are not in its transcript), and Codex's provider requests unknown (app-server 0.159 does not send its response ids); Codex turns after its done and what started them; late replies, steered or at the next turn; the hub's fact offers, acknowledgements, bytes offered and acknowledged, build times, hook start-up times and steer round trips by path, in the task window; capability readbacks; validity by the grader's gates, and for turn-free the treatment received; held-back messages (`quiet` events) apart from [FYI] messages (which include the final [FYI] the instructions ask for); stale notices; shadow split predictions with their traces; the hooks each agent's records show, by a label that keeps paths and arguments out, with Claude's hook durations and the hub's own timing of every facts hook call; and contributions. Summaries give medians over valid completed attempts and over the (case, repeat) pairs every arm completed validly, totals over the attempts whose tasks were handed out and to which the measure applies (a Codex measure in a Claude-only arm is not counted as unknown), with the number of attempts each was unknown for, and the reasons for the rest. A repeat given twice is refused, and `--plan pilot` (or `study`) lists every planned attempt that wrote no record, whole repeats included. `ledger.json` holds code fragments from the agents' writes and local paths: keep it with the private run data and never commit it; the summary is what a verification record quotes. In-task windows end at different points by arm (a turn-free integration step comes before the done, an advisory completed-change notice turn after it), so arms are compared on whole-attempt usage.
76
+ It reads each run record, the Claude transcript it names and the fixture's git history, and writes `ledger.json` into the first run directory with a `units` table that names the unit and coverage of every measure. Give `--run` once per directory to pool the repeats of one plan; an attempt a directory's `cohort.json` planned that wrote no record is listed as missing, and one a still-locked `runs/` hides as unreadable, in the summary per arm too (an arm with no record at all is listed with what it owes). Per attempt: completion (the runner's end reason and its detail, such as `wall-timeout`, and whether every task has a done); setup and active time; first candidate, completion intents, integration and check times, and each agent's settlement read from its own record (the end of its last native turn after the last done, turns started by late messages included; the last done itself when it did not work after it), and how long the agents could still write after the active time ended (`stopped_s`); usage per agent in task (to its last done on the board) and over the whole attempt: Codex turns, assistant messages, token-usage updates whose running total grew past the total before the window, and the token growth; Claude assistant messages (unique message ids), turns and the tokens their usage records; Claude's main-loop requests by request id (side requests and retries are not in its transcript), and Codex's provider requests unknown (app-server 0.159 does not send its response ids); Codex turns after its done and what started them; late replies, steered or at the next turn; the hub's fact offers, acknowledgements, bytes offered and acknowledged, build times, hook start-up times and steer round trips by path, in the task window; capability readbacks; validity by the grader's gates, and for turn-free the treatment received; held-back messages (`quiet` events) apart from [FYI] messages (which include the final [FYI] the instructions ask for); stale notices; shadow split predictions with their traces; the hooks each agent's records show, by a label that keeps paths and arguments out, with Claude's hook durations and the hub's own timing of every facts hook call; and contributions. Summaries give medians over valid completed attempts and over the (case, repeat) pairs every arm completed validly, totals over the attempts whose tasks were handed out and to which the measure applies (a Codex measure in a Claude-only arm is not counted as unknown), with the number of attempts each was unknown for, and the reasons for the rest. A repeat given twice is refused, and `--plan pilot` (or `study`) lists every planned attempt that wrote no record, whole repeats included. `ledger.json` holds code fragments from the agents' writes and local paths: keep it with the private run data and never commit it; the summary is what a verification record quotes. In-task windows end at different points by arm (a turn-free integration step comes before the done, an advisory completed-change notice turn after it), so arms are compared on whole-attempt usage.
75
77
 
76
78
  Contributions are a heuristic for possible loss, never a certificate: per agent and file, the identifiers and changed fragments its applied writes introduced (only a Claude tool call with a successful result, or a completed Codex patch, counts; a Write replaces the agent's earlier contribution and is not credited with what other agents wrote; a delete removes it; when a file is moved, every agent's contributions to it are checked at its new path) that the final tree lacks, a fragment counting as present anywhere in the file. A same-name overwrite shows as a lost fragment. Shell commands run during the task are not attributed and are counted under `coverage`, with a missing transcript or an unreadable file. Correctness comes from the official grader, for every arm.
77
79
 
package/docs/events.md CHANGED
@@ -16,13 +16,13 @@ marked `private: true`, and PII tasks `pii: true`.
16
16
  | `overflow`, `undeliverable` | `id`, `from`, `peer` |
17
17
  | `stale` | `id`, `from`, `peer`, `task`: a notice about the recipient's open task, dropped unsent because, right before it would have been handed over, that task was closed, had another owner or was gone (issue #106) |
18
18
  | `quiet` | `id`, `from`, `peers`: an agent message held back from these members of a silent turn-free cohort; its other recipients got it (issue #107) |
19
- | `fact` | `peer`, `id` (the offer), `files` (files whose diff it carried), `plans`, `unknown` (files whose change it showed with attribution unknown), `bytes` (the injected text), `via` (`hook` for Claude, `steer` for Codex, `done` with an integration request), `ms` (the hub's time to build it), `hookMs` (the hook process's own start-up and connect time), `accepted` (whether app-server took the steer), `unanswered` (app-server did not answer it within 10 s: it may have gone in), `rttMs` (an accepted steer: from sending it to app-server's answer), `probe` (a context check), `coverage` (it named files earlier changes are not covered for): one fact offer (issue #108) |
19
+ | `fact` | `peer`, `id` (the offer), `files` (files whose diff it carried), `plans`, `unknown` (files whose change it showed with attribution unknown), `named` (files under a named directory it named without a diff, those counted as "N more" or held back by the PII filter included), `bytes` (the injected text), `via` (`hook` for Claude, `steer` for Codex, `done` with an integration request), `ms` (the hub's time to build it), `hookMs` (the hook process's own start-up and connect time), `accepted` (whether app-server took the steer), `unanswered` (app-server did not answer it within 10 s: it may have gone in), `rttMs` (an accepted steer: from sending it to app-server's answer), `probe` (a context check), `coverage` (it named files earlier changes are not covered for): one fact offer (issue #108) |
20
20
  | `fact_ack` | `peer`, `id`, `via` (`hook` and `steer`: a readback found the offer in the native session; `done`: the next `hub_task_done`), `ms` (from the offer): an acknowledged offer, the only thing that moves a peer's view |
21
21
  | `capability` | `peer`, `state` (`verified` or `lost`), `via`: a peer's context path for facts |
22
22
  | `native_turn_end` | `peer`: Claude's Stop hook, the end of its turn (Codex's is its `turn_end`); quiescence evidence for an integration |
23
23
  | `hook_stats` | `peer`, `n` (facts hook calls in the turn, its Stop included), `startupMs` and `maxStartupMs` (the hook processes' start-up and connect time, summed and the largest), `hubMs` (the hub's own time for them): at Claude's Stop (issue #108) |
24
24
  | `cohort` | `id`, `event` (`formed`, `joined`, `lifted`), `silent`, `tasks`, `owners`: owners of overlapping tasks formed a cohort, it changed membership, or it stopped being silent (issue #107). Recorded in every regime; only a turn-free project's cohorts can be silent |
25
- | `split` | `task`, `verdict` (`split`, `single`, `unknown`), `single` (the peer that would finish both units alone soonest), `splitS`, `singleS`, `reason` (for `unknown`), `trace` (the inputs and steps: peer names and numbers only): a shadow split prediction where an overlap forms or changes a cohort, routed or named; it never changes the assignment (issue #109) |
25
+ | `split` | `task`, `where` (`routing`: routing chose the first owner of a task overlapping another owner's task not started yet, not an escalation, relay or reassignment, the record calibration reads; `cohort`: an overlap formed or changed a cohort), `verdict` (`split`, `single`, `unknown`), `single` (the peer that would finish both units alone soonest), `splitS`, `singleS`, `reason` (for `unknown`), `trace` (the inputs and steps: peer names, their profiles of versions and coordination, and numbers only): a shadow split prediction; it never changes the assignment (issue #109) |
26
26
  | `state` | `peer`, `state` |
27
27
  | `turn_start` | `peer`, `turn` (`<peer>#<hub run>.<n>`, unique across restarts). A turn follows the adapter: pausing a busy peer does not end it |
28
28
  | `turn_end` | `peer`, `turn`, `ms`, `tokens` (when the adapter reported any during the turn), `files` and `snapshotMs` (when snapshots are on: how many files the turn changed, and the time both snapshots took) |
@@ -1,6 +1,6 @@
1
1
  # Operations guide
2
2
 
3
- This guide describes ahub 0.12.4 and control protocol 13. Live verification
3
+ This guide describes ahub 0.12.6 and control protocol 13. Live verification
4
4
  results and remaining prerequisites are recorded separately in [the smoke ledger](smoke.md).
5
5
 
6
6
  ## Install and start
@@ -408,14 +408,18 @@ stays the default until an evaluation says otherwise (`docs/cooperbench.md`).
408
408
  touched, with the other members' new plans; the last member still at work
409
409
  keeps the others' files after they finish. A directory a task names stands for
410
410
  git's changed (staged or not) files against HEAD, new and deleted files under
411
- it (200 at most; the fact says when more were cut, until it is read back and
412
- again in a new session). A file the peer has seen there stays covered after git
411
+ it (200 at most; the fact says when more were cut, until it is read back, and a
412
+ new session hears it again). A file the peer has seen there stays covered after git
413
413
  stops listing it (put back to HEAD's bytes, or the directory moved away), so
414
414
  the way back is shown (200 at most, the newest versions first; the rest are
415
- named once as no longer tracked). A file there that the peer has neither seen
416
- nor touched appears once someone changed or created it: it is named without a
417
- diff (what happened before is never shown) until the peer reads the fact back,
418
- and until then it counts as a change the peer has not been shown. `.git` directories
415
+ named as no longer followed, until the peer reads that back; a new session
416
+ hears only the drops of its own). A file there that the peer has neither seen nor touched appears once
417
+ someone changed it so that its bytes differ from HEAD's, or created it: it is
418
+ named without a diff (what happened before is never shown) until the peer reads
419
+ the fact back, and until then it counts as a change the peer has not been
420
+ shown. A rewrite with HEAD's bytes, which git lists until it refreshes its
421
+ index, is no change (a file too large to read is compared by git's own hash).
422
+ `.git` directories
419
423
  at any depth and what the denylist keeps from every agent
420
424
  (`src/local/deny.ts` and `local.deny`) are never read or shown. A file a peer touched before it had a view of it (a
421
425
  partial read, say) is compared with what it was then, so a change landing in
@@ -436,7 +440,8 @@ stays the default until an evaluation says otherwise (`docs/cooperbench.md`).
436
440
  acknowledgement says the context reached the native session, not that the
437
441
  model read it. 60 changed lines are shown at most, the cut files named; a
438
442
  history longer than the hub keeps is shown with its attribution unknown, and a
439
- file that falls out of the 64 a peer touched is named once. Only regular files
443
+ file that falls out of the 64 a peer touched is named until the peer reads
444
+ that back. Only regular files
440
445
  of 256 KB or less inside the project are read, re-resolved at every read and
441
446
  opened without following links; larger ones are named without a diff. Facts never go through the bus or the delivery journal;
442
447
  `events.jsonl` records `fact`, `fact_ack` and `capability` events with bytes
@@ -482,14 +487,29 @@ stays the default until an evaluation says otherwise (`docs/cooperbench.md`).
482
487
  silent cohort with the caller, and only then leaves out the request to settle
483
488
  by message. `templates/claude-hooks.json` holds only the check-path hook; the
484
489
  facts hooks need a hub-launched session.
485
- - Routing does not change. When a task's overlap with another owner's open task
486
- forms or changes a cohort, routed or named, the hub records a shadow split
487
- prediction (`split` event; `ahub route explain <id>` shows its trace as it would
488
- be now): whether splitting two equal units between the two peers
489
- (`o_s + u_s < o_f + 2u_f`) would finish sooner than the faster one alone, from
490
- this hub run's recorded task stages. It is unknown unless the units are equal
491
- and known, both peers are available (the other owner idle; the task's own peer
492
- idle or busy taking it) with no other open work (an overlapping task its owner
490
+ - Routing does not change. When routing chooses the first owner of a task that
491
+ overlaps another owner's task not started yet (`where: "routing"`, the record
492
+ calibration reads; an escalation, relay or reassignment is not one), and when an overlap forms or changes a cohort, routed or
493
+ named (`where: "cohort"`), the hub records a shadow split prediction (`split`
494
+ event; `ahub route explain <id>` shows its trace as it would be now): whether
495
+ splitting two equal units between the two peers (`o_s + u_s < o_f + 2u_f`)
496
+ would finish sooner than the faster one alone, from the recorded task stages of
497
+ each peer under the profile it has now. Every hand-over is tagged with the new
498
+ owner's profile: the hub's version, the agent's (Codex's from app-server,
499
+ Claude Code's from its transcript; other agents report none yet) and the
500
+ coordination mode, which decides the hub's own hooks; records accumulate across
501
+ hub runs, one per hand-over (a decline or an escalation away counts against
502
+ the peer that failed, never the next owner), and a peer whose version is
503
+ unknown has no profile. The user's and
504
+ plugins' hooks are not part of it. It is unknown unless the units are equal
505
+ and known, both peers are available (idle, or busy taking the task in
506
+ question: in this hub run, it is still in the turn that task started (the
507
+ task delivered at once to the idle peer) or in which it claimed the task; a
508
+ task queued, held, or steered into a turn about something else is not taken,
509
+ and once that turn ends, busy is another turn; for the
510
+ other owner, while the overlapping task is not started; the routing record,
511
+ and a cohort record formed as a task is assigned, are taken before the task is
512
+ sent, so a busy routed peer is not available then) with no other open work (an overlapping task its owner
493
513
  has started counts), and each has five measured tasks with no more than 30%
494
514
  failures and comparable work times. The work stage of a task ends at its first
495
515
  `hub_task_done`, and the task itself is never one of its own observations.
@@ -627,21 +647,21 @@ Rows without a live process are stale registrations; forget them with
627
647
 
628
648
  Upgrade running projects with the target release's own coordinator. It accepts
629
649
  a running source on control protocol 9 (0.6.x), 10 (0.7.0 through 0.12.0),
630
- 11 (0.12.1 and 0.12.2) or 12 (0.12.3) and only
650
+ 11 (0.12.1 and 0.12.2), 12 (0.12.3) or 13 (0.12.4 and 0.12.5) and only
631
651
  a target on its own protocol, so the target's coordinator fits every supported
632
652
  source and carries every recovery fix released up to it. Protocol 8 and older
633
653
  (0.5.x and earlier) are refused as `manual-bootstrap-required`. Run from the
634
654
  project directory, without replacing the global CLI first:
635
655
 
636
656
  ```bash
637
- bunx --package @staix/agent-hub@0.12.4 ahub upgrade --to 0.12.4 --dry-run
638
- bunx --package @staix/agent-hub@0.12.4 ahub upgrade --to 0.12.4 --yes
657
+ bunx --package @staix/agent-hub@0.12.6 ahub upgrade --to 0.12.6 --dry-run
658
+ bunx --package @staix/agent-hub@0.12.6 ahub upgrade --to 0.12.6 --yes
639
659
  ```
640
660
 
641
661
  | Running now | Coordinator to use |
642
662
  | --- | --- |
643
663
  | 0.6.x (protocol 9) | the target's, through `bunx` as above |
644
- | 0.7.0 through 0.12.0 (protocol 10), 0.12.1 and 0.12.2 (protocol 11), 0.12.3 (protocol 12) | the target's, through `bunx` as above |
664
+ | 0.7.0 through 0.12.0 (protocol 10), 0.12.1 and 0.12.2 (protocol 11), 0.12.3 (protocol 12), 0.12.4 and 0.12.5 (protocol 13) | the target's, through `bunx` as above |
645
665
  | any supported source, with the installed CLI already at the target | `ahub upgrade` below, which is the same coordinator |
646
666
  | 0.5.x or earlier (protocol 8 and older) | not supported: bootstrap by hand with the matching CLI |
647
667
 
@@ -1205,7 +1205,7 @@ upgrade sources; ordinary clients must use protocol 12 (13 from 0.12.4, with 12
1205
1205
 
1206
1206
  ## Amendment: turn-free facts at tool boundaries (issue #108)
1207
1207
 
1208
- - `src/hub/facts.ts` keeps three things apart per peer: what the hub observed in the tree (latest version and transitions per file), what it offered (bounded offers with their file versions and plans), and what the peer acknowledged (its view). Only an acknowledgement moves a view; a later boundary offers everything since the view again. Files are the paths every task of each live cohort in which the peer has an open task names (a directory expands to git's changed files against HEAD, staged or not, and its new and deleted files, 200 at most, with a note when more were cut, told until it is read back and again in a new session), contained in the project as real paths and never `.git` or anything the denylist (`src/local/deny.ts`, `local.deny`) keeps from agents, plus the last 64 it touched; a file that falls out of those is named once. A file the peer touched before it had a view of it is compared with its version at that touch, so a change in between is shown, not absorbed into a first look; a file it saw under a named directory stays covered after git stops listing it (put back to HEAD's bytes, or the directory moved away; 200 at most, the rest named once as no longer tracked). A directory file it has neither seen nor touched is named without a diff when it appears (diffing it against HEAD would show what changed while tracking was off, a PII window, and credit unseen changes by elimination), and counts as not yet shown until that fact is read back. Only regular files of 256 KB or less are read, re-resolved at every read and opened without following a final link or blocking. A history longer than the 64 kept transitions makes the change's attribution unknown.
1208
+ - `src/hub/facts.ts` keeps three things apart per peer: what the hub observed in the tree (latest version and transitions per file), what it offered (bounded offers with their file versions and plans), and what the peer acknowledged (its view). Only an acknowledgement moves a view; a later boundary offers everything since the view again. Files are the paths every task of each live cohort in which the peer has an open task names (a directory expands to git's changed files against HEAD, staged or not, and its new and deleted files, 200 at most, with a note when more were cut, told until it is read back and again in a new session), contained in the project as real paths and never `.git` or anything the denylist (`src/local/deny.ts`, `local.deny`) keeps from agents, plus the last 64 it touched in its current session; a file that falls out of those is named in every offer until one naming it is read back (#112). A file the peer touched before it had a view of it is compared with its version at that touch, so a change in between is shown, not absorbed into a first look; a file it saw under a named directory stays covered after git stops listing it (put back to HEAD's bytes, or the directory moved away; 200 at most, the rest named as no longer followed until that is read back; a new session hears only the drops of its own). A directory file it has neither seen nor touched is named without a diff when it appears with bytes that differ from HEAD's (a rewrite with HEAD's bytes, listed until git refreshes its index, is no change, for the fact and for whether the peer is current; a file too large to read is compared by git's blob id of its bytes, hashed in-process, while all such files in one comparison total 64 MiB or less, and is not compared otherwise (so the answer never depends on what an earlier comparison hashed); the bytes are the raw working-tree bytes, without git's filters, so with an end-of-line conversion or LFS such a file reads as changed and is named; #112) (diffing it against HEAD would show what changed while tracking was off, a PII window, and credit unseen changes by elimination), and counts as not yet shown until that fact is read back. Only regular files of 256 KB or less are read, re-resolved at every read and opened without following a final link or blocking. A history longer than the 64 kept transitions makes the change's attribution unknown.
1209
1209
  - Attribution: a transition is credited only with effect evidence: a Claude Edit, MultiEdit or Write whose result equals its input applied to the file as observed at PreToolUse with no observation in between, or a Codex `fileChange` whose changed lines equal the observed diff. Everything else is shown with its attribution unknown; an agent's own verified writes advance its view and are not shown back. A Claude Read advances its view only when it returned the whole file (no offset or limit, at most 2000 lines, none over 2000 characters); a Codex read action never does (it may be partial). A diff that matches a PII pattern is replaced by a note naming the file, and a changed file whose name matches one is counted without its name.
1210
1210
  - Acknowledgement and capability: Claude's hook runs before and after every tool and at Stop (`ahub claude` injects it in a turn-free project); the pre answer's text becomes `additionalContext`, and the row Claude Code writes in its transcript for it (matched by tool use id and offer id) is the readback. Codex's offer goes in with `steerText` (bound to the running turn), and the steered input coming back as a `userMessage` item is the readback. A readback marks the peer's path verified; a new session or thread, the peer going offline, or three offers past a minute without one (checked at boundaries and every 30 seconds; facts sent with an integration request wait for the next done and never count, a refused steer is dropped, not unread, and an unanswered one stays pending, since it may have gone in) unverifies it. Until verified, a peer gets a one-line probe at most three times per session, and no offers at all once three are unread; a Claude session whose transcript the hub cannot find gets none at all and loses a verified path.
1211
1211
  - The control contract is protocol 13: `facts` carries the phase, tool, input, tool use id, session id and transcript path; `silenced` serves check-path; `send` results name held-back recipients. Protocol 12 stays an authenticated recovery source.
@@ -1214,6 +1214,7 @@ upgrade sources; ordinary clients must use protocol 12 (13 from 0.12.4, with 12
1214
1214
 
1215
1215
  ## Amendment: shadow split prediction (issue #109)
1216
1216
 
1217
- - `predictSplit(input)` is a pure function next to `assign()`, which no longer takes a split option: assignment never changes. For the pair formed by the peer routing picks and the owner of an overlapping open task, it compares a split (the later of each peer's orientation plus one unit) with the best peer alone (orientation plus two units). It is unknown unless the units are equal and known (one task of the same class is one unit), both peers are available (the other owner idle; the task's own peer idle or busy taking it) with no other open work (an overlapping task its owner has started counts), and each has five measured tasks in this hub run with at most 30% failures and work times whose IQR does not exceed their median; a difference under a tenth of the single time is inconclusive.
1218
- - `Tasks.splitObservations` reads only tasks handed out in this hub run (one version and hook profile), never the routed task itself, leaves out claims, types failures (failed check, changes requested, escalation, release, decline, unresolved integration), ends the work stage at the first done call (checks and integration are not work) and leaves the stages unknown for an accept recorded by the done itself. The `split` event carries the trace.
1217
+ - `predictSplit(input)` is a pure function next to `assign()`, which no longer takes a split option: assignment never changes. For the pair formed by the peer routing picks and the owner of an overlapping open task, it compares a split (the later of each peer's orientation plus one unit) with the best peer alone (orientation plus two units). It is unknown unless the units are equal and known (one task of the same class is one unit), both peers are available (idle, or busy taking the task in question: in this hub run, it is still in the turn that task started (the task delivered at once to the idle peer) or in which it claimed the task; a task queued, held, or steered into a turn about something else is not taken, and once that turn ends, busy is another turn; for the other owner while the overlapping task is not started; the routing record, and a cohort record formed as a task is assigned, are taken before the task is sent, so a busy routed peer is not available then; #115) with no other open work (an overlapping task its owner has started counts), and each has five measured tasks with its current profile and at most 30% failures and work times whose IQR does not exceed their median; a difference under a tenth of the single time is inconclusive.
1218
+ - `Tasks.splitObservations` reads one observation per hand-over to the peer with the profile it has now, across hub runs of the project (what happened from that hand-over up to and with the next one is that peer's: a decline or escalation away is its failure, never the next owner's), never the routed task itself, leaves out claims, types failures (failed check, changes requested, escalation, release, decline, unresolved integration), ends the work stage at the first done call (checks and integration are not work) and leaves the stages unknown for an accept recorded by the done itself. The `split` event carries the trace.
1219
1219
  - An overlap that forms or changes a cohort records a `split` event for the task, routed or named (a benchmark names every owner); `ahub route explain <id>` appends the trace as it would be now. Rerouting needs held-out evidence through #110 first.
1220
+ - 2026-10-03 (T3 calibration data, issue #109 amendment): every hand-over records the new owner's split profile on its history entry: the hub's version, the agent's version (Codex's from app-server's `initialize` answer, Claude Code's from its transcript rows) and the hook profile, which is the hub's own (the coordination mode: a turn-free project runs the facts hooks in Claude). Observations count for a prediction only with the peer's current profile; while a version is unknown there is no profile, and the prediction is unknown. The prediction calibration reads is the one recorded when routing chose the first owner (no single named candidate, no claim; not an escalation, budget relay or reassignment of work already begun) of a task that overlaps another owner's task not started yet (`where: "routing"`), and `route explain` prefers the same unstarted pair; cohort-time predictions stay in the log with `where: "cohort"` and are not calibration data. The user's and plugins' hooks are not seen by the hub (ponytail: read them from the native records as the benchmark ledger does), and only Claude and Codex report a version so far.
@@ -0,0 +1,46 @@
1
+ # Agent-hub 0.12.5 verification
2
+
3
+ Implementation and candidate verification for issues #113, #112 and #109 (T3 groundwork), performed on 2026-10-03. Native CLIs: Codex 0.160.0 (it replaced 0.159.3 on this machine during the work), Claude Code 2.1.288; the v2 manifests pin both. Production application is verified separately after publication.
4
+
5
+ ## #113: the leftover after an interrupted arm
6
+
7
+ Cause, reproduced before any fix: a fixture hub with the real Codex, a Codex turn running `sleep 40`, torn down as the runner did (pause, close, `ahub kill`, `projects remove`). `codex` is a node launcher (`codex.js`) that forwards SIGTERM to the native app-server and waits; mid-turn the app-server does not exit, so after 1 s the hub SIGKILLed the launcher alone. The app-server went on under init (PPID 1) with its MCP servers, holding the launcher's stderr pipe, and the daemon, its state removed and its claim released, stayed alive on that pipe: `kill` answered "hub is not running". With the app-server in its own process group, stopped as one, nothing of the fixture was left and `kill` took 173 ms instead of 1,087 ms.
8
+
9
+ What Codex starts (MCP servers, tool commands such as `sleep 40`) leads process groups of its own (read from the live table). A group stop alone does not reach them if the native is SIGKILLed before it cleans up, so the hub's stop records them by parent links, before the stop and every 100 ms of the grace period, then freezes what is left (SIGSTOP) before reading the table again and killing it, and is done only when the table shows none of it. Tests (`test/shutdown.test.ts`): a launcher that waits for a child ignoring SIGTERM, a member that outlives its leader, a descendant in a group of its own, a process started and orphaned during the grace period (the test fails without the recording during the grace period), no process table at all (the group's own answer decides), and a table that keeps showing a member (the stop gives up at its deadline, never claims done).
10
+
11
+ Live interrupted-arm smokes, on case 0 (`pallets_click_task:2068`) with real agents, Codex 0.160.0 and Claude Code 2.1.288. Each prepared its own run directory and built a protect list of every earlier benchmark artifact and native session file. After each run exited, the process table was searched for anything with a fixture in its argv and `lsof` for anything with a fixture as its working directory.
12
+
13
+ | Run | Source | Arm | Stopped at | Signal | Cleanup | Left running (argv, cwd) | Restoration |
14
+ |---|---|---|---|---|---|---|---|
15
+ | A | `9d9f552` | solo Codex | 30.5 s into the work, Codex mid-turn | SIGTERM to the runner | `clean` (normal shutdown 533 ms) | nothing, nothing | inputs, sibling locks |
16
+ | B | `9d9f552` | advisory | 41.0 s into the work | SIGINT to the runner's process group | `clean` (918 ms) | nothing, nothing | inputs, locks, trust |
17
+ | C1 | `9d9f552` | solo Claude | completed normally (44.3 s) | none | `clean` (460 ms) | nothing, nothing | inputs, locks, trust |
18
+ | C2 | `9d9f552` | advisory (the next arm of C1's run) | 8.9 s into the work | SIGTERM to the runner | `clean` (707 ms) | nothing, nothing | inputs, locks, trust |
19
+ | D | `9d9f552` | advisory | setup, before the hub started (`ahub kill` said "hub is not running", recorded) | SIGTERM to the runner | `clean` (170 ms) | nothing, nothing | inputs, locks |
20
+ | E | `9d9f552` | turn-free | 31.1 s into the work, and again during teardown | SIGINT to the group, twice | `clean` (916 ms) | nothing, nothing | inputs, locks, trust |
21
+ | S1 | `9d9f552` | advisory | native startup: the hub up, Claude's channel prompt confirmed, its sandbox probe not run yet | SIGTERM to the runner | `clean` (380 ms) | nothing, nothing | inputs, locks, trust |
22
+ | S2 | `9d9f552` | advisory | native readiness: Claude's probe done, Codex's thread started and its MCP ready, the Codex probe not finished | SIGINT to the group | `clean` (561 ms) | nothing, nothing | inputs, locks, trust |
23
+ | V | `9d9f552` | solo Codex, with a visitor: a `sleep` started in the arm's fixture by the test, not by the arm | 30.4 s into the work | SIGTERM to the runner | `incomplete_or_unknown`: the visitor, working in the fixture, unresolved and left alone | the visitor only | withheld; `restore.ts` refused while the visitor ran, restored once it left |
24
+ | F | `c9d7ff0` | advisory | 40.2 s into the work | SIGINT to the group | `incomplete_or_unknown`: the Orca terminal shell working in the fixture was not recorded | nothing, nothing | withheld; `restore.ts` restored after its checks |
25
+
26
+ Runs A, B, E, V and S2 were repeated on `5a30876`, the last commit to change code, with the same outcomes (every cleanup `clean`, V `incomplete_or_unknown` and recovered once its visitor left, nothing of a fixture left running). V and F are the fail-closed path. In V, at the head's source, a process the arm did not start sits in the fixture: the runner leaves it alone, keeps the protected inputs and sibling locks in place, writes the record to `recovery/` and stops the cohort; `bun scripts/benchmarks/restore.ts --run <run>` refuses while the visitor runs (it names it twice: unresolved in the record, and working in the fixture), and once it is gone restores the modes, moves the record into `runs/` and leaves the upstream checkout readable again. In F, on an earlier source, the same path was taken for real: the Orca terminal's login shell, which runs Claude's launcher and works in the fixture, outlives the terminal close by a moment and was not recorded; since `d692534` it is recorded as part of Claude's launch chain (the parent of the launcher, working in the fixture). Setup-stage stops are recorded as `interrupted` with their error kept. C1's completed arm waited for Claude's turn end (bound 30 s): the session's `turn_duration` row (proved on the probe turn in every Claude arm that got that far) arrived after 1.01 s (on `1b33179` it had already been written when the wait began), and its tree was the same after the verified teardown as at the end of its active time. Every arm that reached Codex recorded `conditions.codex.skills` from `skills/list`, a summary only (counts by scope and a hash of the names): 177 user skills, all enabled, in every arm, and a number of system skills that Codex reported differently from arm to arm on the same machine and Codex configuration: 4 or 5 within the batch on `9d9f552`, 4 on `1b33179`, none on `65d6acc` and 5 on `24390de`, while direct `skills/list` calls in between saw 5. The cause is in Codex and was not established. Each record therefore says what `skills/list` answered for its arm at thread start, which a comparison of arms can read; that this was the session's effective skill availability is not shown. No record keeps the `skills/list` answer itself (no skill path in any record since `24390de`, against 183 paths in a record from the source before). D stopped before Codex and says so. Process start times in the records are read in UTC since `65d6acc`. Runs on the sources in between (`ab11387`, `488a2df`, `d692534`, `557c40c`, `28475c3`, `52972ff`, `24390de`, `65d6acc`, `3d017ed`, `1b33179`) gave the same outcomes; those before `c9d7ff0` did not yet scan working directories.
27
+
28
+ The turn-end marker: in the 0.12.4 release runs, all 126 Claude turns that ended with an `end_turn` answer were followed by a `turn_duration` row of the session before the next activity; the 5 transcripts whose last activity has no such row are cut-off attempts.
29
+
30
+ Fault fixtures (`test/benchmarks/teardown.test.ts`, in-memory process tables and one real process tree): a normal shutdown is `clean` with nothing signalled; a lost acknowledgement gets a fallback on proven identities only, group leaders by their group: `clean_with_fallback`; a fallback that cannot stop an actor, an unreadable table and unreadable working directories are `incomplete_or_unknown`; a first read that fails keeps every recorded actor for the fallback; a reused pid, a replacement hub on the same fixture and a foreign process naming the fixture are never signalled, and the last two leave the cleanup incomplete; Codex's argv without the fixture is owned by lineage; a background job in a recorded group is the arm's while that group is known, and one that left it, or a shell merely working in the fixture, is left alone and keeps the cleanup open. Turn end: tool-use-only rows, a final answer without the marker, another session's marker and a new turn after the marker do not pass; the bounded wait reports ended, timeout, interrupted and unreadable. Restoration: modes back parents first with failures named, a trust entry the user changed kept, and the recovery tool refusing while the runner, a recorded process or anything working in the fixture runs, or while working directories cannot be read, then restoring and moving the kept record into `runs/`. The ledger keeps the frozen transcript prefix when Claude Code appends after it, reports `late_append_bytes`, and leaves an open turn's settlement unknown.
31
+
32
+ ## #112 and #109
33
+
34
+ Unit tests: the seen-cap and touched-limit notices survive a dropped offer, are spent by an acknowledgement, and a new session starts both and its touched list afresh (the session test fails with the reset removed); a rewrite with HEAD's bytes under a named directory is neither named nor counted against `current()`, a real change is; a names-only offer counts its names in `named`, and the ledger reports them per path (unknown for events written before 0.12.5). Split records: a hand-over is tagged with the new owner's profile, observations count only with the current profile, an unknown version makes the prediction unknown, and a `routing` record is made only for a first assignment routing chose while the overlapped task has not started (the reassignment case fails with that rule removed); Codex's version comes from app-server's `userAgent`, Claude Code's from its transcript (daemon test). Calibration and the held-out assessment (T3) and opt-in rerouting (T4) stay open: the data source and sample size are preregistered in #110 first.
35
+
36
+ ## Upgrade
37
+
38
+ `scripts/smoke-recovery-09-10.ts` recovered the real published 0.12.4 package (protocol 13) into the candidate (protocol 13): operation `46160c3c-0f87-4be3-8cfa-3739f36bcab7`, queued envelope ids and the task identity and state digest preserved.
39
+
40
+ ## Automated gate and review
41
+
42
+ `scripts/check.sh` at `5a30876`: 730 tests passed, 0 failed, 3,677 expectations across 66 files; typecheck, bundle freshness, npm package contents and process ownership passed; `check: OK`. CI on `5a30876`, `241fbe7`, `9d9f552`, `c988054`, `1b33179`, `3d017ed` and `fcc27e9`: both platforms passed at the first attempt. CI on `65d6acc`: Ubuntu passed; macOS passed on its third attempt. The first hung in `test/daemon.test.ts` after `conflicts: with experiments.stale_notices deliver ...` until the job's 15-minute limit, with no test reaching its own timeout (the event loop was blocked, as in one local run earlier in this work that showed the test process at 99.6% CPU with no child process). The second failed `a peer can never answer a permission request` (`permission=cancelled`), a test that failed the same way on this branch before and on the 0.12.4 branch. Neither reproduced locally: 8 runs of `daemon.test.ts` alone and 8 of it with the facts, benchmark and events tests all passed. Both are recorded as open, not as fixed.
43
+
44
+ Independent host Claude Code OCR delegation reviews (read-only reviewers with the OCR rule groups, `REVIEW.md`, the `AGENTS.md` invariants and the issue contracts) on the exact head. The first round, on `ce1901d`, found that #113 had been amended ("Review correction") before the first commit and that the runner met only its earlier contract; the runner's teardown was rebuilt to the amended contract. Its other findings, all fixed: a localized `ps` that parsed to nothing and still read as clean; restoration while a survivor could read the inputs; the completion wait running before the agents were paused; a daemon still finishing its exit counted as a leftover; fallback by pid where a group was needed; a transcript read error able to skip the teardown; process groups of Codex's own children; group stops confirmed by the leader's exit alone; a new session told about the old one's touched files; the touched-limit notice without a name cap; the HEAD check reading a second version of a file; Claude's version lost behind a big transcript row; escalations and relays counted as routing records; `route explain` and the record describing different pairs; the skills condition a constant; `teardown.ts` unpinned; ledger gaps; the 0.12.4 upgrade claim without proof. The second round, on `615dce9`, found a macOS EPERM from a group whose members were all zombies (CI), recorded actors dropped when the first table read failed, a background job that could escape the reads, `restore.ts` restoring without a recovery record or while the runner ran, a replacement daemon adoptable at teardown, kept records invisible to grading, records with teardown errors graded, an end reason overwritten by later flags, `restoration.json` silent about sibling locks, the completion wait writing the agent's index, failures of an escalation or decline charged to the new owner's split observations, and smaller items; all were fixed, and the live runs above were repeated with the working-directory scan. The third round, on `049b2cd`, found that a process the tree started during the stop's grace period escaped it (the third gap of the same kind, so the stop was redesigned: freeze what is left, read the table again, then kill; the rule is in `AGENTS.md`), split observations that charged nobody for a decline or escalation away, large files outside the HEAD comparison, a run-level restoration that an early exception could let through, unbounded runner commands and an unhandled SIGHUP, a completion-wait hash that followed links, a registration left by an interrupted setup, `restore.ts` not taking back the trust entry and `runner.py restore` as a second, unchecked path, a doc overclaim and inconsistencies in this record; all were fixed and the live runs above repeated. The fourth round, on `a92b475`, found the benchmark fallback still reading the table while the actors ran (it now records what they start while it waits and freezes before SIGKILL), the ledger counting attempts grading had made unavailable (one gate now serves both), a completion wait able to run past the active limit or let a write into the patch unflagged, trust restored while containment was uncertain and never retried after a failure, a recovery that followed links, a ledger written in place, an unbounded `lsof`, an unrecorded member of a group whose id may have been reused (left alone, and the stop no longer called done), descendants without a grace period, a synchronous table read blocking the hub, and PII-named files dropped from the counts; all were fixed and the live runs above repeated, the fail-closed path (V) included. The fifth round, on `a390aec`, found a startup race in a new stop test that turned Linux CI red (its children now report readiness only after their handler is set), a HEAD check that could follow a link or block on a fifo swapped in for a large file (now hashed in-process, opened as facts open files), a narrow window that could leave a process frozen and not killed (what is frozen is now killed or continued, in the hub's stop and the benchmark fallback alike), group membership judged from a recorded instead of a current pgid, end flags that grading and the ledger did not report (now `end_story`), a tree check that ended before the kill (now from the end of the active time to after the verified teardown), the recovery's ledger written in place, the regime in force missing from split profiles, and smaller items; all were fixed and the live runs above repeated. The sixth round, on `768b60e`, found the `skills/list` answer, with absolute skill paths, kept in every run record beside its hashed summary (its response is no longer recorded), records from before 0.12.3 graded invalid by the ledger instead of unknown, the capture of an arm's processes and the new pins without tests (the capture is now a function tested on fixture tables, the pins are tested in `prepare` and `grade`), a stop that read as done when a launcher had exited and left members in its group (it now fails, its group unsignalled), pipes not dropped after a successful stop, an offer read back late spending the notice of a later drop of the same file, large files hashed again at every boundary (now once per version), and smaller items; all were fixed. The seventh round, on `24390de`, found no Critical or Important defect; its Minor findings were fixed: process start times read in local time, so a reader in another time zone (or a zone change during a run) would see every identity as changed (now read in UTC); descendants leading groups of their own given no SIGTERM before the freeze (a tool command such as `git` killed without a chance to remove its lock; now each gets its own SIGTERM and the grace period); a pre-0.12.3 record with a definite isolation failure reported as unknown; the restoration ledger rewritten every 5 s in full (now only when the recorded actors change, without restored sibling maps); a recovery that could reach through a replaced parent directory; and test gaps (the adapter's launcher-and-native group stop, a reused pid in the hub's stop). The eighth round, on `fcc27e9`, again found no Critical or Important defect; its Minor findings were fixed: a process that a read after the freeze showed for the first time was killed without being frozen (it is now frozen and the table read again, up to five times, in the hub's stop and the runner's fallback alike); after a failed read following the freeze, a pid whose STOP had failed could still be killed from the earlier read (now only what the STOP reached is killed); a frozen tree could wait on a slow `ps` for its SIGKILL (the read after the freeze is given up after 1 s); a launcher that had exited, whose pid now led someone else's group, failed every later stop and so blocked a Codex restart (a live process holding the leader's pid now proves the group is not the hub's, and zombies are not counted); large-file hashing in one call is capped at 64 MiB; the completion wait's bound is taken when the wait starts; a `skills/list` error is recorded as its class, not its message; a process not the arm's is recorded by program name, never its arguments; the tree hash and the metadata hash open files without blocking; and `restoration.json` and the recovered records are written atomically. The ninth round, on `3d017ed`, found no Critical or Important defect; its Minor findings were fixed: a per-call budget for large-file hashing that let an integration target move without a byte changing (whether large files are compared now depends on their total size alone, and a file is hashed only as the version observed), a calibration record for work routing gave back to its proposer, whose observation is left out (no record now either), a process frozen in a later round that was continued instead of killed when the next read failed (now killed by the last read that showed it), 0.12.3 and 0.12.4 records' `cleanup_complete` read as a verified teardown by the ledger (unknown now), and `setupMs` that counted the teardown of an arm stopped in setup. The tenth round, on `c988054`, found one Important defect, fixed: routing-time calibration records would almost always read `not available`, because the owner of the overlapped task is busy taking it when the next task is routed (it now counts as available while that task is not started, as the routed peer does; #109 amended). Its Minor findings were fixed too: an exited launcher's group id that could be reused before a later stop (the group is now followed after the launcher exits until it is gone), a group SIGTERM sent after the leader may have been reaped, no capture during the completion wait (now every 5 s there too), the tree hashed after the pause rather than first, pre-0.12.5 records re-judged by the ledger (now judged by their own flags, their teardown shown as not verified), and documentation of what `teardown_errors` covers. Live runs S1 and S2 now cover a stop during native startup and readiness. The eleventh round, on `241fbe7`, found no Critical or Important defect; its Minor findings were fixed: a process that led a group of its own only after it was recorded had that group's other members left out of what the runner owns, which could read as a complete cleanup (groups are now followed by the table as it is now), what the runner froze and did not kill was continued only when one more read succeeded (now by the last read that showed it), the failed stop of an exited launcher did not name the members it left (now it does), a routing record written after the board write without a guard (now guarded), and documentation of the Linux clock-step limit, the exited launcher's limit and what 0.12.3 and 0.12.4 recorded. The twelfth round, on `f63c5ea`, found no Critical or Important defect. Two of its Minor findings were documentation and were fixed: the 64 MiB rule described as a limit per file, where it is the total of the large files in one comparison, and the skills condition described as what each arm had, where it is what `skills/list` answered at thread start. Its other Minor findings were deferred. From the seventh round on, the round summaries above name the findings that were fixed; the Minor findings they do not name were recorded as known limits (below) or deferred. The pull request lists the deferred ones, and later rounds.
45
+
46
+ Known limits: the hub sees its own hooks only, so a profile does not tell the user's or plugins' hooks apart, and only Claude and Codex report a version; a process that leaves the recorded parent links and groups between two reads (every 5 s while the agents work) is not recorded, and keeps the cleanup open only if it names a fixture or works in one; the hub's own stop records what the tree starts during the grace period every 100 ms, so a process started and orphaned within one such interval escapes it (an unrecorded process left in a recorded group is never signalled, and the stop is then not called done); identity is the pid with a start time of one-second resolution, and on Linux `ps` derives that time from the boot time, which moves when the clock is stepped, so a stop across a clock step sees its processes as replaced and leaves them alone: the hub's own stop then reads incomplete, never done, while the runner's teardown counts such an owned process as gone and keeps the cleanup open only if it names or works in the fixture (the benchmark runs on macOS); the hub's own stop does not sweep the group of a Codex launcher that was killed from outside before the stop (while that group has members the stop fails instead of reading as done, naming them), nor processes that left that group for groups of their own before the stop (an MCP server or a tool command orphaned with the launcher): its pipes are dropped so they cannot keep the hub alive; a file too large to read is compared by its raw bytes, so with git's end-of-line conversion or LFS it reads as changed and is named; the trust entry is written by read, change and rename, so a write another Claude Code session makes to `~/.claude.json` in the milliseconds between is lost (as before this release).
@@ -0,0 +1,163 @@
1
+ # Agent-hub 0.12.6 verification
2
+
3
+ Implementation and candidate verification for issue #115 (follow-ups of the 0.12.5 review), performed on 2026-10-03. Native CLIs: Codex 0.160.0, Claude Code 2.1.288, Kimi Code 2.1.1. Production application is verified separately after publication.
4
+
5
+ ## The macOS CI hang of 0.12.5
6
+
7
+ On 0.12.5's CI, `test/daemon.test.ts` once hung until the job limit; no test reached its own timeout, so the event loop was blocked. It reproduced locally twice under heavy machine load (load average about 12), once on the 0.12.5 checkout and once on this branch. Each time `bun test` sat at 100% CPU for minutes, with nothing but the banner in its output in one case.
8
+
9
+ The stacks were sampled with macOS `sample` and symbolized with Bun's own profile build of 1.3.14, whose UUID matches the installed binary (`C7E7A979-F99B-3466-9AD6-E56A63373A35`). The hung main thread shows this chain:
10
+
11
+ 1. An async continuation, resumed from a timer, calls `Bun.spawnSync` (`BunObject_callback_spawnSync`, `spawnMaybeSync`).
12
+ 2. Inside spawnSync's own event loop, the test runner's timeout callback fires (`BunTest.bunTestTimeoutCallback`). Through `BunTest.run` and `runTestCallback`, it runs the next test's JavaScript.
13
+ 3. That JavaScript calls `Bun.spawnSync` again, and the inner call spins in `SpawnSyncEventLoop.tickWithTimeout` for good.
14
+
15
+ So a test timeout that fires while a `Bun.spawnSync` runs can hang the whole run. Minimal attempts to reproduce it outside the suite did not hang: when a test times out, Bun kills the child that test's own spawnSync started. The precise trigger therefore also needs a spawnSync that Bun does not attribute to the timed-out test.
16
+
17
+ Mitigation, since the fix belongs in Bun:
18
+
19
+ - `scripts/check.sh` runs `bun test --timeout 20000`, so a timeout, which is the trigger, is much rarer under load.
20
+ - It runs the suite under `scripts/hang-watch.sh`. If the run is still going after 600 s, the watchdog prints a stack sample (macOS) or the thread states (Linux) and stops it, so a hang fails fast with evidence.
21
+ - The process-heavy stop tests carry timeouts of their own.
22
+ - Upgrading Bun (1.4.2 is current) is left for a separate change.
23
+
24
+ The permission-test flake of 0.12.5 had a different cause: the test rig's 200 ms permission timeout is shorter than a loaded runner takes to connect a peer and make two tool calls. That test now has its own 30 s window.
25
+
26
+ ## Automated gate
27
+
28
+ `scripts/check.sh` at the head of the pull request, after the eleventh review round and the merge of main (#118): 754 tests passed, 0 failed, 3,807 expectations across 68 files in 135 s (macOS CI counts one expectation fewer) (the 68th file is #118's); `check: OK` (with the fourth round's fixes: 739 passed, 3,723 expectations; at `b03b035`: 739 passed, 3,710 expectations; at `fdc884e`: 738; at `4b98d2d`: 737 passed in 116 s at a load average of about 5). Mutation checks (each test fails with its fix removed; the last four are 0.12.5 tests checked again against this release's stop change): split availability under the 0.12.5 rule, the ACP launcher test without `detached`, pipes in the non-group stop, the fallback's newcomer freeze, the kill-only-stopped rule, the reused-pid stop, and the grace-period SIGTERM to own-group descendants.
29
+
30
+ CI on `4b98d2d`: macOS passed; Ubuntu failed the new ACP launcher test, because dash points a background job's stdin at `/dev/null` before the explicit `0<&0`, so the fake agent never got its handshake. The test now passes stdin through fd 3 (`bef35ab`); CI passed on both platforms there, at `fdc884e`, `5bfaf53`, `70b1daf` and `b03b035`.
31
+
32
+ Cost of the group stops: each ACP or Pi stop now reads the process table twice (a `ps` of about 1,300 processes takes about 60 ms here). `test/acp.test.ts` and `test/pi-daemon.test.ts` together take 8.1 s against 3.1 s on 0.12.5. The group stop now finishes as soon as the leader has exited and a read shows nothing left.
33
+
34
+ ## Live checks
35
+
36
+ - Upgrade, on `4b98d2d`: `scripts/smoke-recovery-09-10.ts` recovered the published 0.12.5 package (protocol 13, integrity from the registry) into the candidate (protocol 13): operation `31f206a6-6dbc-4f42-adfb-40eac841417a`, queue and task identity and state digest preserved.
37
+ - Kimi through ACP, on `4b98d2d`, real `kimi acp` 2.1.1, in a disposable project with its own `AGENTHUB_HOME`: started by the hub as the daemon's child, it led its own process group; `ahub kill` said "hub stopped", nothing of that group was left, and the daemon was gone.
38
+ - Interrupted benchmark arms with real agents, on `4b98d2d` (A, B) and `bef35ab` (C, E, V, S2; it differs only in the ACP test):
39
+
40
+ | Run | Arm | Stopped at | Signal | Cleanup | Left running | Restoration |
41
+ |---|---|---|---|---|---|---|
42
+ | A | solo Codex | 31.1 s into the work | SIGTERM to the runner | `clean` (298 ms) | nothing | inputs, sibling locks |
43
+ | B | advisory | 40.3 s into the work | SIGINT to the group | `clean` (463 ms) | nothing | inputs, locks, trust |
44
+ | C1 | solo Claude | completed (57.7 s); turn end after 0.75 s, tree unchanged | none | `clean` (340 ms) | nothing | inputs, locks, trust |
45
+ | C2 | advisory | setup | SIGTERM to the runner | `clean` (424 ms) | nothing | inputs, locks, trust |
46
+ | E | turn-free | 31.2 s into the work, and again during teardown | SIGINT to the group, twice | `clean` (422 ms) | nothing | inputs, locks, trust |
47
+ | V | solo Codex with a visitor in the fixture | 31.0 s into the work | SIGTERM to the runner | `incomplete_or_unknown`, the visitor left alone and recorded as `program: sleep` | the visitor | withheld; `restore.ts` refused while it ran, restored once it left |
48
+ | S2 | advisory | native readiness | SIGINT to the group | `clean` (452 ms) | nothing | inputs, locks, trust |
49
+
50
+ Runs A, V and S2 were repeated on `fdc884e` with the same outcomes. Every one of these runs registered its fixtures with Orca, since the runner still added them then. The user has since directed that no fixture is added to Orca without an explicit registration (#117), and the registrations these runs left are removed under #117. Since `70b1daf` the runner only looks registrations up, and since `b03b035` it refuses an unregistered selection up front. A live check of the latter registered nothing. A run prepared afresh (`/private/tmp/ahub-0126-refusal`, 40 fixtures, none registered) was started with the runner of `b03b035`. It refused the whole run before anything was locked or recorded, naming every unregistered fixture. That lookup was later replaced by #118's read-only helper (`scripts/benchmarks/orca-workspace.ts`, merged into this branch from main), which refuses at the first fixture missing and checks each arm's identity against the preflight; #118 records its own checks, and this refusal check was not repeated on it. Orca's repository list was the same before and after (15 repositories), the upstream checkout kept its mode, and no attempt record was written. No arm has run on the lookup-only runner, which needs fixtures registered explicitly first: the live arms of this record ran on the runner that still added them, and an arm on the released runner waits for the user to register fixtures. Every record carries `platform: darwin`. The Codex skills condition was 177 user and 5 system skills in every arm with Codex.
51
+
52
+ ## Review
53
+
54
+ Independent read-only reviews (Claude Code OCR delegation, with the OCR rule groups, `REVIEW.md`, the `AGENTS.md` invariants and issue #115), each on the head of the time. The first round, on `4b98d2d`, found one Important gap: this record and the spec amendment for the Bun cause were missing. The spec was amended (#115, Decisions of 2026-10-03), and this record was written. Its Minor findings were fixed:
55
+
56
+ - Hand-overs in the split prediction are keyed by owner generation: a task handed back to a peer that had it is a new hand-over.
57
+ - A failed Pi start reports its own error, and Pi's stop finishes its teardown when the group stop fails.
58
+ - Pi has a launcher test, as ACP does.
59
+ - The spawn-heavy stop tests' own timeouts are at least the suite's 20 s.
60
+ - The runner removes the trust write's temp file on every path, the recovery's included.
61
+ - `hang-watch.sh` signals only the process it began watching (pid and start time), escalates to SIGKILL, and is stopped once the tests end.
62
+ - Run records and the restoration ledger are written atomically.
63
+ - `programOf` looks at a bounded number of argv words.
64
+ - Several doc and comment fixes.
65
+
66
+ Deferred: the Codex start-failure test does not exercise a stop that fails, since no fake can make the group stop fail there without mocking. The second round, on `5bfaf53`, found no Critical or Important defect. Its Minor findings were fixed:
67
+
68
+ - The program name comes from `ps`'s name for the executable, not from stats that can hang on a dead mount.
69
+ - A failed restore keeps the trust write `pending` for the recovery.
70
+ - The trust temp files are named in one place, and removed when a write or a restore fails.
71
+ - The recovery's pending-trust path is tested.
72
+ - The Bun mechanism is worded as what the stacks show, not as a general rule.
73
+
74
+ The runtime review of that round also found Minor gaps. Fixed: a claimed task now counts as handed over for the split prediction; Claude's version is cached when the transcript has none yet; the watchdog test waits for its trap and always cleans up; `sample` writes its report to a temp file it removes. Deferred and listed in a comment on the pull request: a task steered into a turn about other work counts as taken (shadow only); the start-failure tests do not exercise a failing stop. The same commit makes the runner lookup-only in Orca (#117).
75
+
76
+ The third round, on `70b1daf`, found two Important defects; both are fixed.
77
+
78
+ - **Arguments recorded as a program name.** The program name read from `ps -o comm=` is, on macOS, a process's own argv[0], so a process that sets its title (Node's `process.title`, perl's `$0`) put its arguments into the record. It now reads `ps -o ucomm=`, the executable's name as the kernel recorded it, only while the pid has the start time recorded, and records nothing when it cannot be read. A test with a self-titled perl fails with `comm`.
79
+ - **A steered task counted as taken.** A task steered into a busy Codex turn counted as taken by it. Only a task that starts the owner's turn counts now (idle with nothing queued when it is sent, or claimed), and a test with a steering peer covers it.
80
+
81
+ Its Minor findings were fixed:
82
+
83
+ - The runner, whose own trust write threw, no longer touches the trust file.
84
+ - An unregistered fixture refuses the run up front instead of failing every arm, and the documented procedure has the registration step.
85
+ - The operator note on refused Kimi and Pi restarts was added.
86
+ - Several wording fixes.
87
+
88
+ Deferred, as before: the start-failure tests do not exercise a failing stop.
89
+
90
+ The fourth round, on `b03b035`, found two Important defects in the split prediction; both are fixed.
91
+
92
+ - **A hand-over outlived its turn.** A task that started its owner's turn stayed marked as taken after that turn ended, so a later turn about something else counted as taking it. The mark now lasts as long as that turn: the peer leaving busy drops it. It is set only once the task envelope was delivered at once (idle before, busy after, not queued), so a held delivery is not taken; a claim counts only in a turn the hub sees.
93
+ - **The queue was checked again at prediction time.** A peer still in the turn its task started, with a status message queued behind it, read as not available, and so did a claimant with one queued. Only the turn decides now.
94
+
95
+ Tests cover a turn that ended, a message queued behind the turn, a claim with one queued, a held delivery and the daemon's journal, and each fails with its part of the fix removed. Its Minor findings were fixed:
96
+
97
+ - The ledger counts a record caught in the middle of the recovery's move once, as withheld (also raised by the Codex review on the pull request), and refuses a run whose `runs/` is still locked instead of reporting its attempts missing.
98
+ - The Orca refusal names why each lookup failed, and each arm looks its fixture up before touching it.
99
+ - The recovery removes the temp file a runner left when it died in its own trust restore, and a temp file it cannot remove keeps the recovery open instead of aborting it.
100
+ - The tail permission test has the 30 s permission window too.
101
+ - The docs name the hub's project registration where they meant it, and the availability rule names the claim and the turn.
102
+ - Several nits: a nested ternary, the version cache stamped only after a read that worked, a failed `sample` said so.
103
+
104
+ Deferred, as before: the start-failure tests do not exercise a failing stop (a fake cannot end the launcher between the health check and the failure that follows it).
105
+
106
+ The fifth round, on `56585ea`, found no Critical or Important defect in the change. It found that main had moved: #118 (#117's Orca change) had been merged, so the pull request conflicted and no CI had run on its head. Main was merged in. The Orca lookup is #118's: its helper, its preflight and its per-arm identity check replace this branch's own preflight and lookup move. This branch keeps the guard that asks for a hub shutdown only once an arm got past its locks, which the merged code needs (its `orcaProject` is now always set). The docs keep #118's text on explicit authorization for Orca registrations. Its Minor findings were fixed:
107
+
108
+ - A pause no longer drops a hand-over: the turn's end is read from the adapter's own state.
109
+ - A temp file the recovery cannot remove keeps the trust entry unrestored, so the next recovery tries again; it had been marked restored and never retried. A test covers it.
110
+ - The ledger no longer refuses a run whose `runs/` is locked: the attempts it cannot read there are listed as `unreadable`, never as missing, and the withheld ones are rows as before.
111
+ - #115's Design and #109's Decisions were amended to the turn rule; test assertions that could pass without a routing record now require one.
112
+
113
+ CI on the merge (`46e22e0`): macOS passed; Ubuntu failed two of #118's tests. They run the runner's entrypoint, which on Linux stops at this release's macOS-only refusal before the Orca checks. The helper-pin test now expects that refusal off macOS, with the same checks that nothing was written or run, and the preflight test runs on macOS only, where the runner does. CI passed on both platforms at `e2d66ea`.
114
+
115
+ The sixth round, on `46e22e0`, found the Ubuntu failure above (fixed in `e2d66ea`) and checked the merge: #118's helper, its pin, its preflight before any change and its per-arm identity check are intact, and nothing of either side was lost or duplicated. Its Minor findings were fixed:
116
+
117
+ - The runner's own pending trust write whose temp file cannot be removed keeps the trust entry open for the recovery, as the recovery already did for its own; it had been marked settled.
118
+ - The ledger's per-arm summary lists the attempts a locked `runs/` hides (`unreadable`), so the summary a record quotes accounts for them.
119
+ - An interrupt right after the locks no longer asks a hub that never started to stop.
120
+ - A claim reads the claimant's own adapter state, as the turn's end does.
121
+
122
+ CI passed on both platforms at `e9137dd` and `5440086`. The seventh round, on `e9137dd`, found no Critical or Important defect. Its Minor findings were fixed:
123
+
124
+ - When the runner's own temp file could not be removed, the recovery treated the write as possibly landed and could take back an entry someone else had set meanwhile. The runner now records the write as `not_written`, and the recovery removes the copy without touching any entry. A test with someone else's entry covers the recovery's side and fails without its fix.
125
+ - An arm with no record at all appears in the ledger's summary with what it owes (`missing`, `unreadable`).
126
+
127
+ The claim's change was left untested in that round, on the premise that only the cohort record at the claim reads its mark; `route explain` reads it too (round 8).
128
+
129
+ The eighth round, on `5440086`, found no Critical or Important defect. It found the same class a fourth time: with an incomplete cleanup, a trust write the runner knew never landed was left `pending`, so the recovery could take back an entry someone else set meanwhile. Following `REVIEW.md`, the rule went into `AGENTS.md`. A write still `pending` in the runner's own process is `not_written` on every path, and neither the runner nor the recovery then touches an entry. Its other Minor findings were fixed:
130
+
131
+ - The recovery re-settled a `changed_concurrently` entry (the user's) when the run was unrestored for another reason, and could take it back. It now leaves it. This was older than this release (0.12.5). A test covers it and fails without the fix.
132
+ - The summary seeds an arm with no record into its full shape, so its common medians count it.
133
+ - A claim made while paused, then resumed in the same turn, is read as taken by `route explain`. A test covers it and fails with the bus's `paused` state.
134
+ - `restoration.json` names a temp file left behind, rather than a trust entry, when that is what is left; the summary's units and the docs describe `unreadable`.
135
+
136
+ CI passed on both platforms at `e5940e0`. The ninth round, on `e5940e0`, enumerated every state the trust entry can be left in by the runner, a runner that dies at any point, and the recovery. It found no Critical or Important defect, and every path the runner settles itself correct. Its Minor findings were fixed:
137
+
138
+ - A runner that died between its trust write and the rename leaves `pending`, and the recovery took back any entry set meanwhile. It now takes back only an entry exactly as the runner would have written it, and none while the runner's temp file is still there (the rename never happened). This was older than this release. A test covers the three cases and fails without either check.
139
+ - The recovery's own restore named its temp file by its own pid, which no later recovery looks for. It now uses the runner's.
140
+ - The runner's settlement is a function (`settleTrust`), tested as a table of stage, containment and temp removal; the `AGENTS.md` rule points to it. The table test fails with the earlier order (containment before `pending`).
141
+ - The common medians' change has a test: an arm with no record leaves no common pair. It fails without the change.
142
+
143
+ CI passed on both platforms at `bbf59d9`. The tenth round, on `bbf59d9`, found no Critical or Important defect. It compared `settleTrust` with the inline code it replaced, field by field, and found them equal. It walked every point at which a runner can die, every stage on disk, and the 0.12.3 to 0.12.5 ledgers, and found no path where the recovery changes an entry the runner never wrote or the user changed. Fixed:
144
+
145
+ - A test now checks that the recovery names its restore's temp file by the runner's pid. It fails when the recovery's own pid is used.
146
+ - The `TrustLease` comment now says that every ledger since 0.12.3 records `written`.
147
+ - The `AGENTS.md` rule no longer gives a count of rounds.
148
+
149
+ Deferred, older than this release and listed in the pull request:
150
+
151
+ - A recovery of a 0.12.3 or 0.12.4 ledger (no runner identity) killed mid restore leaves a temp file that no later recovery looks for.
152
+ - A 0.12.5 runner whose rename threw left its temp file behind a `restored: true` run.
153
+ - A run directory whose earlier invocation stopped before any arm keeps its `restoration.json` (`restored: true`). A recovery after a crashed re-run there reports the run restored and exits 0 while the re-run's locks and trust entry stay; it changes no entry.
154
+
155
+ CI passed on both platforms at `e0b71a0`. The eleventh round, on `e0b71a0`, found no Critical or Important defect. The new test now runs the recovery as this process (`self` is its pid), so a temp file named by either the recovery's pid or `self` fails it; the record's third deferred item says what the recovery reports.
156
+
157
+ Deferred: the start-failure tests, as before.
158
+
159
+ ## Known limits
160
+
161
+ - The Bun 1.3.14 hang is mitigated, not fixed. A test that times out during a spawnSync can still hang a run, which the watchdog then stops after 600 s.
162
+ - Group stops cost about two process-table reads per stop.
163
+ - The limits recorded for 0.12.5 stand.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@staix/agent-hub",
3
- "version": "0.12.4",
3
+ "version": "0.12.6",
4
4
  "description": "Native multi-agent hub: Claude Code, Codex, Kimi Code, Pi and local inference as peers in one project",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "agent-hub",
3
- "version": "0.12.4",
3
+ "version": "0.12.6",
4
4
  "description": "Channel between Claude Code and the agent-hub daemon: peer messages from Codex, Kimi and the local worker arrive as channel events; hub_send replies.",
5
5
  "author": {
6
6
  "name": "Young Joon Lee",