@staix/agent-hub 0.12.8 → 0.12.9

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -2,6 +2,16 @@
2
2
 
3
3
  Issue and pull request numbers in the entries for 0.7.7 and earlier refer to the previous repository, archived on 2026-09-30 when this repository's history was rewritten; the one exception is the open smoke-check issue, formerly #12, which moved here as #1. Numbers in newer entries refer to this repository.
4
4
 
5
+ ## Unreleased
6
+
7
+ ## 0.12.9
8
+
9
+ - The CooperBench native runner proves each prepared fixture root is still the directory preparation left before anything is locked, written or launched, and each arm checks its own again first (#119): a root replaced after preparation by a symlink to an equivalent outside tree passed the lexical `resolve()` comparison and the baseline content checks and would have redirected setup and agent writes there. The check is read-only (`lstat` and the real path, never a follow), so a substitution is refused without touching its target; a sibling fixture root that is itself a symlink is refused before the sibling-artifact walk reads it.
10
+ - The model relay no longer lets a gateway heartbeat name the served model (#137): an SSE event carries model identity only with generation activity (a delta with any field, a role-only first chunk included, a finish reason or a non-streaming message), so a keepalive with `choices: [{index: 0, delta: {}}]` cannot pin `actualModel` to the synthetic `keepalive` label and refuse the real model that follows. A heartbeat-only or cancelled-before-identification stream leaves the served model unknown; no requested alias is substituted.
11
+ - ACP tool identity is bound to the announced call, not the permission request title (#138): Qwen 0.24.7 announces an MCP call as `hub_send (agent-hub MCP Server)` and then titles the permission request with the serialized arguments, so the exact-name auto-approval never matched and the manual prompt showed the argument JSON twice. The announced title is cached per call id (cleared on reuse, evicted on completion, bound at the initial `tool_call` only — a later update's mutable display title never rewrites it), and a request resolves to the canonical `mcp__<server>__<tool>` name only when the server half is a server the session was configured with. Argument text never becomes an identity candidate, approval still picks only `allow_once`, and an unresolved or cut payload keeps its conservative manual path, displayed under the announced title.
12
+ - The model relay journals request-bound identity and cancellation provenance (#139): every upstream dispatch attempt (a fallback is its own record) closes exactly one sanitized `RelayRequestRecord` — resolved alias, upstream-configured model, observed provider and served model with their source (gateway header, generation SSE event, or local MLX configuration), outcome, duration and a confirmed-mismatch flag — exposed through `relay.requests()` (last 1000) and an `onRequest` hook that can neither break the proxied stream nor reject unobserved. A backend's mutable last-served label is never a request's evidence, HTTP 200 plus the requested alias identifies nothing, a request cancelled before identification stays explicitly unidentified, and a primary/auxiliary role stays unknown without native evidence. Records carry no prompts, tools, keys or Access headers.
13
+ - The Pi and Qwen native CooperBench study driver is versioned as manifest v3 and `scripts/benchmarks/native-pi-qwen.ts` (#140), pinning hub 0.12.9: the 2026-10-04 private calibration study (arms solo-pi, solo-qwen, joint-pi-qwen; Pi 1.0.1, Qwen 0.24.7; 60 preregistered attempts) becomes a regression-tested headless path that never touches Orca. The pinned build is the effective one — each native's `--version` runs under the final isolation environment, because Qwen's PATH bootstrap reported the managed 0.24.7 and fell back to base 0.24.1 under `QWEN_HOME` isolation; the protected-file probe counts only with structured denial evidence, never a model-written marker; source guards follow each case's `source_dirs` (`src/` for Click/Jinja, `dirty_equals/` for dirty_equals); new source files enter the binary submission patch (`git add -N`); setup resources are disposed on every exit path without erasing the active-window record; existing attempt evidence is rejected before any record is written (`wx` claims); served-model evidence is each request's own journaled record, where only cancelled-before-identification is non-evidence and an identified mismatch fails the gate whatever the outcome; and Qwen's peer tool approval is the exact canonical name through the adapter's #138 binding, not a title workaround. Grading flows through the same official evaluator adapter with quality, model and request-linkage coverage reported separately. The live cohort, official Docker controls and native readback remain manual live legs.
14
+
5
15
  ## 0.12.8
6
16
 
7
17
  - Port Switchyard's Stage signals/scoring, Plan/Execute, advisor gate and escalation policies in-process with source-based golden tests and Apache-2.0 attribution (#124, #125).
@@ -22,7 +22,7 @@ python3 scripts/benchmarks/runner.py prepare \
22
22
  --output /private/path/to/new-run
23
23
  ```
24
24
 
25
- The manifest's archive and prompt hashes are checked before a fixture is accepted. The `prepared.json` ledger binds the fixture roots, baseline commits, complete baseline path counts and manifest hash.
25
+ The manifest's archive and prompt hashes are checked before a fixture is accepted. The `prepared.json` ledger binds the fixture roots, baseline commits, complete baseline path counts and manifest hash. At run time each fixture root must still be the directory preparation left: a root replaced by a symlink to an equivalent outside tree, a non-directory, or a name whose real path differs from the prepared one is refused read-only before anything is locked, written or launched (#119).
26
26
 
27
27
  ## Native execution
28
28
 
@@ -51,15 +51,23 @@ bun scripts/benchmarks/native.ts \
51
51
  --cases 0
52
52
  ```
53
53
 
54
+ During a measured calibration block, suspend unrelated Claude and Codex sessions so that quota growth between snapshots can be attributed to the attempt. `--setup-only` does not enter an active-work window and leaves the pre reading explicitly unknown. The first study block validates that quota readings are actually collected for Claude-participating arms (non-`unknown` `claudeUsage.pre` and `claudeUsage.post` in those run records) before subsequent blocks rely on them.
55
+
54
56
  The native runner runs on macOS only and refuses anything else (each run record says `platform`): the arms run in Orca terminals, and on Linux a clock step moves the start times its teardown proves processes by. Before native execution, obtain explicit user authorization before changing Orca registrations; a benchmark request alone is not authorization. Only after that authorization, an operator manually registers each exact fixture root with `orca repo add --path '<fixture>'` and confirms `orca worktree list --repo id:<repo-id>` shows that exact path. The runner performs read-only exact-path repo/worktree lookup; it preflights every selected fixture before changing fixture files, input modes, cohort or attempt records. If either identity is missing or mismatched, it exits nonzero with setup guidance and leaves the prepared run available for retry after explicitly authorized registration. Each arm rechecks its identity before changing fixture files. Benchmark and agent workflows never add or remove Orca registrations. The runner uses the project CLI to start and stop each daemon, and requires exact worktree/cwd readback for native sessions. Claude starts through the canonical `ahub claude` guard and loads this checkout's candidate bundle through an exact session-only `--mcp-config` server (`server:agent-hub`); it does not promote or mutate the globally installed plugin. Codex uses the native app-server adapter and `workspace-write` sandbox. Each agent must execute a setup-only `head -c 1` probe against the exact protected file and produce only the denied marker before scored tasks begin. There is one native sandbox layer per agent.
55
57
 
58
+ Fixture roots are checked without following a root symlink, and their device/inode identity is captured before asynchronous setup (#119). After the baseline and sibling lookups, the synchronous setup mutation batch rechecks that identity before writing. Git setup, daemon startup, Claude terminal creation (after its terminal-list lookup), and Codex startup also recheck at their launch boundaries. A symlink or ordinary directory substituted during a preceding await is rejected before that boundary uses it. These checks do not pin a directory descriptor across filesystem syscalls or an external launcher: an operator must keep the prepared root and its ancestors unchanged throughout the run. Concurrent same-user renames between a check and a syscall or inside an external launcher are outside this guarantee; this is not a claim of atomic directory pinning.
59
+
56
60
  Every run record, diff, PTY/provider trace and evaluation artifact stays in the mode-0700 private run directory. Setup errors, structured provider quota errors, hub budget pauses, unsettled deliveries and interruptions stay separate and unscored. A quota snapshot or numeric `429` alone is not a provider error. On exit the runner restores the exact input and read-lock modes, the sibling artifact modes and the scoped Claude trust flag (removing its own fresh Claude trust entry), under the condition below. Teardown (issue #113) proves what it stops. Each process an arm starts is recorded with its pid and start time and the evidence that it is the arm's: the daemon by the pid in its state directory and an argv that serves this fixture (a later daemon there is a replacement, never adopted), Claude's launch chain by the arm's own session id as its `--session-id` (and the Orca terminal shell that runs it, working in the fixture), the Codex app-server as the daemon's child. What runs below them, and what is in a group a recorded process leads while that group is known to be the same one (its leader alive, or a recorded member still in it), is the arm's too; the table is read every 5 s while the agents work and during the completion wait, so a tool command's background job is recorded while its parent runs (one that detaches between two reads, outside the fixture and without it in its argv, is not seen). A fixture name in an argv is never proof on its own (Codex's need not name the fixture), and every recorded actor counts even when a read of the table fails. Teardown pauses every agent, reads the deliveries still owed, lets a completed arm's Claude end the turn it is in (up to 30 s and never past the 300 s limit (the bound is taken when the wait starts; one 250 ms poll can pass it), judged by the `turn_duration` row Claude Code writes after a turn, proved on the arm's probe turn; a stopped or timed-out arm does not wait; a completed arm's tree is hashed at the end of its active time and again after the teardown, and a write in between flags the attempt invalid), asks for the normal shutdown (Claude's terminal closed, `ahub kill`, the hub's project registration removed (`ahub projects remove`); the runner leaves Orca registrations untouched; a step of it that fails is recorded in `cleanup.normal.errors` and does not by itself fail the attempt, since the process readback decides, so a hub project registration left behind is listed there), and reads the process table back, in the C locale and in UTC (start times are read as identities, the same in every reader). What is still running after a settle period gets SIGTERM; what is left then is frozen (SIGSTOP), the table read again, and it is killed (SIGKILL), what was frozen and not killed being continued: a frozen process starts nothing, so that read sees all of it. Every signal goes only to an identity read again just before, by the group the table shows it in now; a group leader is signalled with its group. What the read after the freeze shows for the first time is frozen too and read again before anything is killed, and when such a read fails only what a STOP reached is killed, by the last read that showed it (a stopped process keeps its pid). Anything with the fixture in its argv or as its working directory that is not proved the arm's (a background job that escaped the reads, someone's shell) is never signalled and keeps the cleanup open; the record keeps its pid, start time, program name (the name of its executable as the kernel recorded it at exec, `ps -o ucomm`, and only while the pid has the start time recorded; never `comm`, which on macOS is the process's own argv[0] and can carry its arguments; omitted when it cannot be read) and working directory, never its arguments. The record says `clean`, `clean_with_fallback` or `incomplete_or_unknown` (`cleanup`, with the normal shutdown's errors, every signal sent, what is still running and what is unresolved), and `cleanup_complete` is true only for the first two. Evidence follows the teardown: the transcript prefix (`transcriptBytes`, `transcriptSha256`, with the session and time it was taken) and the patch. The sibling read locks come off and the protected inputs become readable again only when the cleanup is complete; otherwise both stay, the record goes to `recovery/` beside the run's restoration ledger (which lists every arm's recorded processes and the runner's own identity), the cohort stops, and `bun scripts/benchmarks/restore.ts --run <run>` (also `runner.py restore`, which calls it) puts the modes back and takes back the trust entry once the runner, every recorded process and anything with a fixture in its argv or as its working directory are gone, then moves the kept records into `runs/`, where grading reads them as unavailable. A ledger written before runner identities needs `--runner-exited`, the operator's statement that the runner is gone. Runner commands are bounded (a hung one is killed), and SIGHUP stops the run like SIGINT and SIGTERM. `restoration.json` says `restored: true` only when the protected inputs and every arm's sibling locks are back. From before the runner locks its first input until it writes that outcome, it says `restored: false` and names the runner (#120), so a runner that is killed sends the recovery in, and the recovery waits while that runner still runs. The runner refuses a run directory whose earlier run is not restored (its `restoration.json` or its ledger says so) first, before it reads its inputs or any mode, and again right before it locks anything, since locked modes would become the originals; `restore.ts` first. One runner per directory at a time: two started on one directory at the same moment are not guarded against. `restore.ts` trusts `restored: true` only when the restoration ledger agrees (a ledger from before 0.12.5 still has its locks restored; when `restoration.json` says `restored: true` or names a runner, its trust entry is kept as it was then), and checks the runners again right before it restores anything. The trust entry is taken back with them (by the recovery when the cleanup is incomplete), and kept when the user changed it meanwhile; when the runner's own write never landed (it stopped between recording the lease and the rename), the outcome is `not_written` and nothing is taken back; for a runner that died there, the recovery takes back only an entry exactly as the runner would have written it, and none while the runner's temp file is still there. Orca registrations are not removed during benchmark teardown. Legacy benchmark-created entries may be removed only after provenance confirms the exact fixture path, using `orca project setup-delete --setup <id>` after verifying the exact repo and fixture path; unrelated registrations and fixture/Git/evidence data remain untouched. The ledger reports a record kept in `recovery/` as an attempt with `withheld: true` (unavailable: its cleanup was incomplete, or its sibling read locks could not be put back) until `restore.ts` moves it, never as missing (a record in both places, a move that did not finish, counts once and withheld); while `runs/` is still locked (a cleanup is incomplete), the attempts it cannot read there are listed as `unreadable`, never as missing, and its teardown row carries the normal shutdown's errors (`normal_errors`, for example a hub project registration left behind). The recovery also removes the temp files a runner left (one that died mid trust write or mid restore, or could not remove its own), each a copy of `~/.claude.json`; a write the runner knew never landed stays `not_written` there too, and no entry is touched. Run records carry `completion`, `cleanup`, `restoration` and `stages` (completion wait, shutdown, settle, fallback with the final readback, evidence and restoration times), apart from the end reason: `end_reason_detail` keeps the original classification, and `end_flags` (an unverified model, modified metadata, a tree changed or unverified after the active time) makes any end other than an interruption, a provider quota error or a budget pause an infrastructure error beside it, which grading and the ledger report as `end_story`; a record with `teardown_errors` (a patch, transcript prefix, last capture, fixture metadata or event log that could not be taken, or sibling read locks or the trust entry that could not be put back, a concurrent change of the trust entry included), an incomplete cleanup or a trust entry not taken back is unavailable to grading and to the ledger alike (a record from before 0.12.5 is judged as it was then, by its own `cleanup_complete` and `trust_restored`, and the ledger shows its teardown with `verified: false`: 0.12.3 and 0.12.4 set `cleanup_complete` when the shutdown steps reported success (commands, terminal close, lock and trust restores), with no process readback, which is no proof that the processes were gone); runner commands, the `ps` reads included, run in process groups of their own, so a Ctrl-C reaches the runner, which stops in this order, and not the command it is running; the completion wait's tree hash never writes the agent's index. `teardown.ts` and the process table it reads (`src/hub/child-process.ts`) are pinned in `prepared.json`, and `teardown.ts` in `cohort.json` and `grade.json` too, with the other runner sources.
57
61
 
58
62
  ## Manifest v2: the turn-free arm
59
63
 
60
64
  `scripts/benchmarks/manifest-v2.json` keeps v1's cases, models and limits and adds a fourth arm, `hub-turnfree-codex-claude` (issue #110): the same two agents and assignment rotation as `hub-codex-claude`, with `coordination: "turn-free"` in the fixture's hub config and fixture instructions that tell the owners not to message each other. A v1 manifest still validates and prepares. Its cohorts are run, graded and reported with the release and runner sources they were prepared with: `native.ts` refuses a manifest whose hub version is not the checkout's (v1 pins 0.12.3), and `runner.py` refuses to grade a cohort whose runner sources differ from its own.
61
65
 
62
- Hooks are equal across arms: no arm runs the user's or a plugin's hooks, and no arm runs a status line (`disableAllHooks` turns it off in the other arms, so the turn-free arm leaves it out; Claude's quota reaches the hub in no arm). Claude starts with `--setting-sources project` and `--strict-mcp-config`; the solo and advisory arms set `disableAllHooks`, and the turn-free arm's session settings carry the hub's own hooks (before and after every tool call, and at Stop) and nothing else, because they are the treatment. Every Codex thread starts with `features.hooks` off: the turn-free arm's Codex boundary is the adapter's steer into the running turn, whose readback is the steered input coming back as a user message item. MCP servers are isolated too: Claude has only the hub's (`--strict-mcp-config`), and every Codex runs through a wrapper written into each arm's fixture that turns off the user's plugins, apps, sub-agents and turn-end notifier and disables each MCP server the user's config defines, by name, so it starts only the hub's; the user's config is not changed. Instructions: Claude reads the fixture's `AGENTS.md` through `--append-system-prompt-file` (Claude Code reads `CLAUDE.md`, not `AGENTS.md`), Codex as the project's `AGENTS.md`; Codex also reads the user's global `AGENTS.md`, which only a separate Codex home with its own login would leave out (the run records say so), and the user's Codex skills stay available to it (the wrapper leaves them on; `conditions.codex.skills` records what app-server's `skills/list` reports for the fixture: counts by scope, how many are enabled, and a hash of their names, never their bodies or paths, and the `skills/list` answer itself is not kept in `codexMessages`; Claude's Skill tool is denied); both are the same in every arm. Each run record carries these `conditions`; its `events` carry the hub's `capability` readbacks.
66
+ Every Claude-participating arm runs the same isolated status-line tee (`src/cli/statusline-tee.ts`) to produce `benchmark-usage/claude-usage.json` under the arm's hub state directory for #134. The daemon consumes only the parent `claude-usage.json` for budget and routing; the observer's separate directory leaves those inputs unchanged. Generated session settings set `disableAllHooks: false` because Claude Code gates status-line execution with that flag; the solo and advisory sessions have an empty hook map, and the turn-free session carries only the hub's facts hooks. The existing `--restricted` launch ignores user, project and local settings files; the explicit `--settings` file and managed settings still apply. User/plugin hook settings are not copied into the session. The tee has no original status-line command to run. This instrumentation profile is recorded as `conditions.claude.statusLine: true` and `disableAllHooks: false`, and source/settings hashes prevent mixing it with historical hookless runs. See [Claude Code status-line execution and troubleshooting](https://code.claude.com/docs/en/statusline#troubleshooting).
67
+
68
+ Every emitted arm record carries `claudeUsage.pre` and `claudeUsage.post`. Claude arms wait up to 1.5 seconds for an observer write after probe completion plus the 300 ms status-line debounce before starting active work. The isolated producer refreshes once per second. After active work, a post reading is trusted only if the final turn ended and an observer write follows that settled boundary plus the debounce. Timeout, interruption, unsupported/unsettled turn completion or a missing fresh write records explicit unknown; a cached pre value is never treated as post consumption. These waits stay outside active-work time and do not make collection failure fail an attempt. Solo-Codex launches no Claude just to measure quota, and records explicit unknown with the reason that there is no Claude actor. Readings retain the file's own `at`; expired or invalid-reset windows are marked stale, all-stale/missing/unreadable/malformed data is unknown, and collection failure does not fail an attempt. The file reader refuses symlinks, special files and oversized data, so a FIFO cannot hold teardown. Grade `native_usage.claude` keeps transcript-derived token counts and adds pre/post. Rate-limit fields are optional and appear only after an API response on supported accounts: producer wiring does not establish live provider collection. The first study block must validate actual non-unknown Claude-arm readings before later blocks rely on quota deltas.
69
+
70
+ Claude starts with `--setting-sources project` and `--strict-mcp-config`; the turn-free arm's before/after-tool and Stop hooks remain its treatment. Codex threads keep `features.hooks` off. MCP servers are isolated too: Claude has only the hub's (`--strict-mcp-config`), and every Codex runs through a wrapper written into each arm's fixture that turns off the user's plugins, apps, sub-agents and turn-end notifier and disables each MCP server the user's config defines, by name, so it starts only the hub's; the user's config is not changed. Instructions: Claude reads the fixture's `AGENTS.md` through `--append-system-prompt-file` (Claude Code reads `CLAUDE.md`, not `AGENTS.md`), Codex as the project's `AGENTS.md`; Codex also reads the user's global `AGENTS.md`, which only a separate Codex home with its own login would leave out (the run records say so), and the user's Codex skills stay available to it (the wrapper leaves them on; `conditions.codex.skills` records what app-server's `skills/list` reports for the fixture: counts by scope, how many are enabled, and a hash of their names, never their bodies or paths, and the `skills/list` answer itself is not kept in `codexMessages`; Claude's Skill tool is denied); both are the same in every arm. Each run record carries these `conditions`; its `events` carry the hub's `capability` readbacks.
63
71
 
64
72
  Validity, decided by the grader and applied by the ledger alike: an attempt whose records show a hook or an MCP server that is not the hub's, or whose Claude transcript cannot be read, is unavailable. A turn-free attempt is valid only with its context paths working: both verified before its tasks, and none lost, no cohort lifted and none formed open while the agents worked; teardown comes after that and does not count. Whether the agents' plans overlapped, so that a cohort formed at all, is their doing after assignment and is not a condition: every turn-free attempt without a capability failure counts for the arm, and the ledger reports the treatment received and a median over treated attempts beside it. A capability failure while the agents work is the one exclusion after assignment, because #110 forbids reporting it as a turn-free run (AC3); such attempts are listed with their reasons, never dropped silently. Each attempt records its transcript's length and hash when it ends; validity and every Claude measure are read from that prefix (Claude Code may append rows after it exits: what it writes after the prefix was taken is never counted, the ledger reports its size as `late_append_bytes`, and a turn still open in the prefix leaves that agent's settlement unknown), and a prefix that changed counts as unreadable.
65
73
 
@@ -67,6 +75,26 @@ Validity, decided by the grader and applied by the ledger alike: an attempt whos
67
75
 
68
76
  Pass `--repeat <n>` for the n-th repeat of a case (0 for the first): the arm order is row (case index + repeat) of a Williams design (0, 1, n-1, 2, n-2, ... shifted by the row), so over n consecutive rows every arm runs right before every other one once and repeats of one pair change which arm runs last. The manifest's `plan` names the release pilot (case 0, three repeats, 12 attempts, an active-time ceiling of one hour) and the study (ten cases, two repeats, 80 attempts, 6 hours 40 minutes at 300 s each, setup, grading and teardown excluded); `runner.py` refuses a plan whose attempts or ceiling do not follow from its arms, cases and repeats.
69
77
 
78
+ ## Manifest v3: the headless Pi and Qwen arms
79
+
80
+ `scripts/benchmarks/manifest-v3-pi-qwen.json` (issue #140) keeps the upstream commit, the ten cases and the 300 s active limit, and versions the 2026-10-04 private Pi/Qwen calibration study as a regression-tested execution path: arms `solo-pi`, `solo-qwen` and `joint-pi-qwen`, native builds Pi 1.0.1 and Qwen 0.24.7, one fixed backend (`fixed_backend`) with its expected served model and provider, and a preregistered plan of 60 attempts (ten cases, two repeats, three arms; the six-row odd-n Williams layout, joint ownership `(caseIndex+repeat)%2`, fixed before any model call). Prompt bodies, hidden tests, gold solutions and vendor state stay external, hashes only, exactly as in v1/v2.
81
+
82
+ The driver is `scripts/benchmarks/native-pi-qwen.ts`, headless: it spawns Pi and Qwen (ACP) itself over the hub's model relay and never touches Orca — no terminal, no registration, and the canonical Orca UI-path rules of the Claude/Codex runner above are unchanged. It keeps the runner's protections: private mode-0700 run root outside the repository, umask 0077, macOS-only, SIGHUP/SIGINT/SIGTERM stop the cohort in order, prepared fixture root identity checks (#119), the reuse refusal with `restoration.json`/`restoration-ledger.json` (#120), process ownership proven by `teardown.ts` (each peer process recorded with pid and start time, the table followed while the agents work, signals only to proven identities), and no attempt record ever written over another (the attempt directory must be new, `started.json` and `native-owner.json` are exclusive `wx` claims).
83
+
84
+ Its calibration corrections, relative to what an ad-hoc harness gets wrong:
85
+
86
+ 1. The pinned build is the EFFECTIVE build: each native's `--version` runs under the final isolation environment (Qwen inside its seatbelt profile with `QWEN_HOME`/`TMPDIR` set — the PATH bootstrap reported the managed 0.24.7 in the normal home and fell back to base 0.24.1 under isolation), and the binary path and version are recorded in every attempt record, so launch and recovery emit the same build.
87
+ 2. The protected-file probe passes only on a native read attempt with structured denial evidence — a guard denial for Pi, a failed read tool call with EPERM/EACCES for Qwen, plus the seatbelt kernel probe — never on a model-written `AHUB_PROBE_DENIED` marker alone.
88
+ 3. Source guards follow each case's `source_dirs` (`src/` for Click/Jinja, `dirty_equals/` for dirty_equals).
89
+ 4. New source files enter the binary submission patch (`git add -N` on the source dirs before `git diff --binary`); grading re-collects the same scoped patch and separately refuses any change outside the guarded dirs.
90
+ 5. Relay, tool server and peers are disposed on every exit path; a disposal failure is recorded in `teardown_errors` and never erases the active-window tree record captured before disposal, and the relay is never left open.
91
+ 6. Existing attempt evidence is rejected before any fallback record is written.
92
+ 7. The driver is preflighted by strict typechecking (`scripts/check.sh`) and executable lifecycle checks (every spawnable binary answers a trivial command before a fixture is touched), not transpilation alone.
93
+
94
+ Served-model evidence is per request, from the relay's journaled `RelayRequestRecord` (#139), not from a response wrapper: a completed request that was never identified fails the attempt's model gate (a heartbeat-only stream identifies nothing), an observed mismatch flags it, and a request cancelled before identification stays explicitly unidentified and never certifies another request. Qwen's peer tool is approved through the adapter's shipped tool-identity binding (#138): the announced `hub_send (pilot-peer-bus MCP Server)` title resolves to the canonical `mcp__pilot-peer-bus__hub_send`, the only name on the exact whitelist. The joint arm's MCP peer bus is the versioned `scripts/benchmarks/peer-bus-mcp.py`, pinned with the driver in `prepared.json` and `cohort.json` next to the candidate source pins (`src/adapters/pi.ts`, `src/adapters/acp.ts`, `src/models/relay.ts`).
95
+
96
+ Grading flows through the same official evaluator adapter and controls; a quality failure is never a retry selector and no evaluator feedback reaches the candidate agents during generation. Reports keep quality, model-identity and request-linkage coverage separate per arm, with unavailable attempts (failed, missing or unavailable) retained in the planned denominator, and name the usage units (Pi's incremental `onTokens` counter, Qwen's session `usage_update` running total — never added together) and the tool-surface difference (Pi's hub-moderated tools against Qwen's own seatbelted auto-edit tools). A live cohort, the official Docker controls and a native readback of an attempt's records remain manual live legs requiring accounts and the pinned archives; they are not part of the checked-in tests.
97
+
70
98
  ## Coordination ledger
71
99
 
72
100
  ```sh
@@ -647,21 +647,21 @@ Rows without a live process are stale registrations; forget them with
647
647
 
648
648
  Upgrade running projects with the target release's own coordinator. It accepts
649
649
  a running source on control protocol 9 (0.6.x), 10 (0.7.0 through 0.12.0),
650
- 11 (0.12.1 and 0.12.2), 12 (0.12.3) or 13 (0.12.4 through 0.12.8) and only
650
+ 11 (0.12.1 and 0.12.2), 12 (0.12.3) or 13 (0.12.4 through 0.12.9) and only
651
651
  a target on its own protocol, so the target's coordinator fits every supported
652
652
  source and carries every recovery fix released up to it. Protocol 8 and older
653
653
  (0.5.x and earlier) are refused as `manual-bootstrap-required`. Run from the
654
654
  project directory, without replacing the global CLI first:
655
655
 
656
656
  ```bash
657
- bunx --package @staix/agent-hub@0.12.8 ahub upgrade --to 0.12.8 --dry-run
658
- bunx --package @staix/agent-hub@0.12.8 ahub upgrade --to 0.12.8 --yes
657
+ bunx --package @staix/agent-hub@0.12.9 ahub upgrade --to 0.12.9 --dry-run
658
+ bunx --package @staix/agent-hub@0.12.9 ahub upgrade --to 0.12.9 --yes
659
659
  ```
660
660
 
661
661
  | Running now | Coordinator to use |
662
662
  | --- | --- |
663
663
  | 0.6.x (protocol 9) | the target's, through `bunx` as above |
664
- | 0.7.0 through 0.12.0 (protocol 10), 0.12.1 and 0.12.2 (protocol 11), 0.12.3 (protocol 12), 0.12.4 through 0.12.8 (protocol 13) | the target's, through `bunx` as above |
664
+ | 0.7.0 through 0.12.0 (protocol 10), 0.12.1 and 0.12.2 (protocol 11), 0.12.3 (protocol 12), 0.12.4 through 0.12.9 (protocol 13) | the target's, through `bunx` as above |
665
665
  | any supported source, with the installed CLI already at the target | `ahub upgrade` below, which is the same coordinator |
666
666
  | 0.5.x or earlier (protocol 8 and older) | not supported: bootstrap by hand with the matching CLI |
667
667
 
@@ -17,7 +17,7 @@ Issue: #25. Pi is a managed project peer for implementation, edits, tests, summa
17
17
  - `mlx/fast` resolves to the machine-shared loopback MLX server.
18
18
  - Apple Silicon setup pins Python 3.12, `mlx-lm==0.31.3`, and `Qwen/Qwen3-8B-MLX-4bit` revision `383413e909f3bc5303ce195ebbdf0339c5a1a2a3`.
19
19
  - The MLX input estimate is capped at 16,000 tokens; decode and prompt concurrency are one. DGX's default input estimate cap is 262,144 tokens.
20
- - The relay authenticates every call, rejects browser origins and unknown model aliases, and preserves streamed bytes. It records the requested route and model reported by the backend.
20
+ - The relay authenticates every call, rejects browser origins and unknown model aliases, and preserves streamed bytes. It records the requested route and model reported by the backend. Per request it also journals sanitized identity and lifecycle evidence (#139): resolved alias, upstream-configured model, observed provider and served model with their source (gateway header, generation SSE event, or local MLX configuration), outcome (completed, cancelled, failed), duration, and a confirmed-mismatch flag. A request cancelled before any identification stays explicitly unidentified; a backend's mutable last-served label is never a request's own evidence, and a primary/auxiliary role stays unknown without native evidence. Records carry no prompts, tools, keys or Access headers.
21
21
  - MLX may fall back to DGX before response streaming begins. Failed non-PII tasks can escalate to configured cloud peers with an explicit warning to reconcile partial effects.
22
22
  - Runtime ownership includes process start identity. Project shutdown releases its relay, while explicit `ahub models stop` controls the shared MLX process.
23
23
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@staix/agent-hub",
3
- "version": "0.12.8",
3
+ "version": "0.12.9",
4
4
  "description": "Native multi-agent hub: Claude Code, Codex, Kimi Code, Pi and local inference as peers in one project",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "agent-hub",
3
- "version": "0.12.8",
3
+ "version": "0.12.9",
4
4
  "description": "Channel between Claude Code and the agent-hub daemon: peer messages from Codex, Kimi and the local worker arrive as channel events; hub_send replies.",
5
5
  "author": {
6
6
  "name": "Young Joon Lee",
@@ -15403,7 +15403,7 @@ class ControlClient {
15403
15403
  // package.json
15404
15404
  var package_default = {
15405
15405
  name: "@staix/agent-hub",
15406
- version: "0.12.8",
15406
+ version: "0.12.9",
15407
15407
  description: "Native multi-agent hub: Claude Code, Codex, Kimi Code, Pi and local inference as peers in one project",
15408
15408
  license: "MIT",
15409
15409
  type: "module",
@@ -67,6 +67,7 @@ export class AcpPeer extends BasePeer {
67
67
  private chunks: string[] = [];
68
68
  private readonly toolInputs = new Map<string, unknown>();
69
69
  private readonly toolText = new Map<string, string>(); // streamed argument text, per call, until it finishes
70
+ private readonly toolTitles = new Map<string, string>(); // the title a call was announced with, per call id (#138)
70
71
  private primed = false;
71
72
  private turn = 0; // generation: a prompt cancelled by the watchdog must not touch the turn that followed it
72
73
  private activeDeliveryId: string | undefined;
@@ -210,13 +211,16 @@ export class AcpPeer extends BasePeer {
210
211
  // console would be asked to approve a bare tool name. Keep what the call said it would run (issue #31).
211
212
  // Kimi 2.1.1 sends no rawInput before the answer either: the argument JSON streams as content text (issue #72).
212
213
  else if ((u?.sessionUpdate === "tool_call" || u?.sessionUpdate === "tool_call_update") && typeof u.toolCallId === "string") {
213
- // A new call starts clean, so a reused id can never show the arguments of the call before it.
214
- if (u.sessionUpdate === "tool_call") (this.toolInputs.delete(u.toolCallId), this.toolText.delete(u.toolCallId));
214
+ // A new call starts clean, so a reused id can never show the arguments or identity of the call before it.
215
+ if (u.sessionUpdate === "tool_call") (this.toolInputs.delete(u.toolCallId), this.toolText.delete(u.toolCallId), this.toolTitles.delete(u.toolCallId));
215
216
  if (u.rawInput !== undefined) this.toolInputs.set(u.toolCallId, u.rawInput);
217
+ // Identity is bound at the announcement only: a later update's display title is mutable and must
218
+ // never rewrite what the call was announced as (#138 review).
219
+ if (u.sessionUpdate === "tool_call" && typeof u.title === "string") this.toolTitles.set(u.toolCallId, u.title);
216
220
  const text = Array.isArray(u.content) ? u.content.map((c: any) => (c?.type === "content" && c.content?.type === "text" ? String(c.content.text) : "")).join("") : "";
217
221
  if (text) this.toolText.set(u.toolCallId, text);
218
- if (u.status === "completed" || u.status === "failed") (this.toolInputs.delete(u.toolCallId), this.toolText.delete(u.toolCallId));
219
- for (const map of [this.toolInputs, this.toolText]) while (map.size > TOOL_INPUT_CAP) map.delete(map.keys().next().value as string);
222
+ if (u.status === "completed" || u.status === "failed") (this.toolInputs.delete(u.toolCallId), this.toolText.delete(u.toolCallId), this.toolTitles.delete(u.toolCallId));
223
+ for (const map of [this.toolInputs, this.toolText, this.toolTitles]) while (map.size > TOOL_INPUT_CAP) map.delete(map.keys().next().value as string);
220
224
  }
221
225
  else if (u?.sessionUpdate === "usage_update" && this.opts.onTokens) {
222
226
  const f = { ...u, ...(typeof u.usage === "object" ? u.usage : {}) } as Record<string, unknown>;
@@ -239,24 +243,37 @@ export class AcpPeer extends BasePeer {
239
243
  private async answerPermission(msg: any): Promise<void> {
240
244
  const call = msg.params?.toolCall ?? {};
241
245
  const once = (msg.params?.options ?? []).find((o: PermissionOption) => o.kind === "allow_once");
242
- if (once && typeof call.title === "string" && this.opts.autoApprove?.(call.title)) {
243
- this.opts.log?.(`permission auto-approved for ${this.id}: ${call.title}`); // the name only: arguments may quote a PII turn
246
+ const id = typeof call.toolCallId === "string" ? call.toolCallId : undefined;
247
+ const announced = id === undefined ? undefined : this.toolTitles.get(id);
248
+ // Qwen 0.24.7 announces an MCP call as `<tool> (<server> MCP Server)` and then titles the permission
249
+ // request with the serialized arguments (#138). Identity comes back from the announced title, bound to
250
+ // the call id, and only when the server half is a server this session was configured with: the canonical
251
+ // `mcp__<server>__<tool>` name is the only derived candidate the exact-match whitelist ever sees, and
252
+ // argument text is never one.
253
+ const canonical = (() => {
254
+ const m = announced?.match(/^([\w-]{1,80}) \((.+) MCP Server\)$/);
255
+ return m && this.opts.mcpServers?.some((s) => s.name === m[2]) ? `mcp__${m[2]}__${m[1]}` : undefined;
256
+ })();
257
+ const identity = [typeof call.title === "string" ? call.title : undefined, canonical].find((candidate) => candidate !== undefined && this.opts.autoApprove?.(candidate));
258
+ if (once && identity) {
259
+ this.opts.log?.(`permission auto-approved for ${this.id}: ${identity}`); // the name only: arguments may quote a PII turn
244
260
  return this.send({ jsonrpc: "2.0", id: msg.id, result: { outcome: { outcome: "selected", optionId: once.optionId } } });
245
261
  }
246
262
  // The approver has to see what runs, not only the tool's name ("Bash"). The request itself carries it
247
263
  // for some agents; for the rest it was on the `tool_call` update that announced the call, as rawInput or,
248
264
  // failing that, as streamed argument text that only counts once it is a complete JSON object.
249
- const id = typeof call.toolCallId === "string" ? call.toolCallId : undefined;
250
265
  const raw = call.rawInput ?? (id === undefined ? undefined : this.toolInputs.get(id) ?? jsonObject(this.toolText.get(id)));
251
266
  const full = raw === undefined ? "" : typeof raw === "string" ? raw : JSON.stringify(raw);
252
267
  // A cut payload hides its tail, and a tail can change what runs: it is marked and buys no session-wide grant.
253
268
  const cut = full.length > 600;
254
269
  const input = raw === undefined ? "" : `: ${cut ? `${full.slice(0, 600)} [cut, ${full.length} chars]` : full}`;
270
+ // A request titled with the argument JSON is no title at all: display what the call was announced as.
271
+ const named = typeof call.title === "string" && jsonObject(call.title) === undefined ? call.title : announced;
255
272
  // An unknown payload is never dressed up as a description, and it must not buy a blanket grant.
256
- const title: string = raw === undefined ? `${call.title ?? "tool call"} (payload not reported by the agent)` : `${call.title ?? "tool call"}${input}`;
273
+ const title: string = raw === undefined ? `${named ?? "tool call"} (payload not reported by the agent)` : `${named ?? "tool call"}${input}`;
257
274
  const options: PermissionOption[] = (msg.params?.options ?? []).filter((o: PermissionOption) => (raw !== undefined && !cut) || o.kind !== "allow_always");
258
275
  // Only a bare tool name travels on its own (a desktop notice shows it); anything prose-like stays in the title.
259
- const tool = typeof call.title === "string" && /^[\w.:@-]{1,80}$/.test(call.title) ? call.title : undefined;
276
+ const tool = typeof named === "string" && /^[\w.:@-]{1,80}$/.test(named) ? named : undefined;
260
277
  // Waiting for a person is not the agent going silent: keep the watchdog from cancelling the turn meanwhile.
261
278
  const turn = this.turn;
262
279
  const alive = setInterval(() => this.state === "busy" && turn === this.turn && this.touch(), Math.max(10, Math.min(30_000, Math.floor(this.watchdogMs / 3))));
package/src/hub/daemon.ts CHANGED
@@ -1567,9 +1567,10 @@ export async function startDaemon(opts: DaemonOptions) {
1567
1567
  cwd: opts.cwd,
1568
1568
  watchdogMs: config.watchdog_ms,
1569
1569
  onPermission,
1570
- // The hub's own tools pass without a console prompt, as Codex's do (approval_mode below; issue #72). This
1571
- // relies on the agent putting the tool name in `title`, as Kimi does; an ACP agent that titles calls with
1572
- // model-written text must not be configured as kimi_cmd. A stopping hub approves nothing.
1570
+ // The hub's own tools pass without a console prompt, as Codex's do (approval_mode below; issue #72).
1571
+ // Identity is the request title (Kimi names the canonical tool there) or, for an agent like Qwen that
1572
+ // titles the request with the argument JSON, the announced tool_call title resolved against this
1573
+ // session's configured MCP servers (issue #138). Exact names only; a stopping hub approves nothing.
1573
1574
  autoApprove: (title) => !stopping && HUB_TOOL_TITLES.has(title),
1574
1575
  log,
1575
1576
  onTokens: onKimiTokens,
@@ -31,6 +31,36 @@ export interface ModelRelayStatus {
31
31
  backends: RelayBackendStatus[];
32
32
  }
33
33
 
34
+ /** Sanitized per-request identity and lifecycle evidence. One record per upstream dispatch attempt
35
+ * (a fallback dispatch is its own record). Records carry no messages, tools, keys or Access headers.
36
+ * `identified: false` with `outcome: "cancelled"` is the cancelled-before-identification state; an
37
+ * observed `actualModel` that differs from the upstream-configured `requestedModel` sets `mismatch`.
38
+ * `identitySource` says where the served-model label came from: the gateway response header, a
39
+ * generation SSE event (#137 classification: heartbeats never identify), or the locally validated
40
+ * MLX configuration. HTTP 200, the requested alias and a previous request's label never identify. */
41
+ export interface RelayRequestRecord {
42
+ id: string;
43
+ /** Admission timestamp (start of the upstream dispatch attempt), ISO. */
44
+ at: string;
45
+ /** Resolved backend alias (the requested route). */
46
+ alias: string;
47
+ /** Physical model the relay asked the upstream for. */
48
+ requestedModel?: string;
49
+ /** Sanitized `x-omniroute-provider` header; absent stays unknown. */
50
+ provider?: string;
51
+ /** Observed served model; never read back from the backend's mutable last label. */
52
+ actualModel?: string;
53
+ identitySource: "header" | "stream" | "configured" | "none";
54
+ // ponytail: the relay cannot see native turn structure, so role stays "unknown"; a native surface
55
+ // that knows primary vs auxiliary work (benchmark wiring, issue #140) is the upgrade path.
56
+ role: "primary" | "auxiliary" | "unknown";
57
+ outcome: "completed" | "cancelled" | "failed";
58
+ identified: boolean;
59
+ /** Set only when the observed served model differs from the upstream-configured model. */
60
+ mismatch?: boolean;
61
+ durationMs: number;
62
+ }
63
+
34
64
  export interface ModelRelayOptions {
35
65
  omni: OmniRoute;
36
66
  /** Authoritative admission immediately before each upstream request, including fallbacks. */
@@ -51,6 +81,8 @@ export interface ModelRelayOptions {
51
81
  mlxAlias?: string;
52
82
  mlxModel?: string;
53
83
  fallbackDGXAlias?: string;
84
+ /** Called exactly once per journaled request, at its terminal close, with a sanitized copy. */
85
+ onRequest?: (record: RelayRequestRecord) => void;
54
86
  }
55
87
 
56
88
  export interface ModelRelay {
@@ -58,13 +90,22 @@ export interface ModelRelay {
58
90
  readonly token: string;
59
91
  readonly models: string[];
60
92
  readonly status: () => ModelRelayStatus;
93
+ /** Closed request records, oldest first, bounded to the last 1000. */
94
+ readonly requests: () => RelayRequestRecord[];
61
95
  readonly close: () => Promise<void>;
62
96
  }
63
97
 
98
+ interface RequestJournalEntry {
99
+ readonly record: RelayRequestRecord;
100
+ identify(model: string, source: "header" | "stream" | "configured"): void;
101
+ close(outcome: RelayRequestRecord["outcome"]): void;
102
+ }
103
+
64
104
  interface ActiveRequest {
65
105
  controller: AbortController;
66
106
  release?: () => void;
67
107
  cancel?: (reason?: unknown) => Promise<void>;
108
+ closeRecord?: (outcome: RelayRequestRecord["outcome"]) => void;
68
109
  cleanup: () => void;
69
110
  }
70
111
 
@@ -100,9 +141,23 @@ function bodyForUpstream(body: RelayRequest, model: string): Record<string, unkn
100
141
  };
101
142
  }
102
143
 
103
- function sseResponse(response: Response, release: () => void, onModel?: (model: string) => void, registerCancel?: (cancel: (reason?: unknown) => Promise<void>) => void): Response {
144
+ /** An SSE event carries model identity only with generation activity: a delta with any field (a role-only
145
+ * first chunk counts), a finish reason, or a non-streaming message. Empty choices and empty-delta events
146
+ * are transport heartbeats and say nothing about the served model. */
147
+ function isGenerationEvent(choices: unknown): boolean {
148
+ if (!Array.isArray(choices)) return false;
149
+ return choices.some((choice: any) => {
150
+ if (choice?.finish_reason) return true;
151
+ if (choice?.message && typeof choice.message === "object") return true;
152
+ const delta = choice?.delta;
153
+ return delta !== null && typeof delta === "object" && Object.keys(delta).length > 0;
154
+ });
155
+ }
156
+
157
+ function sseResponse(response: Response, release: () => void, onModel?: (model: string) => void, registerCancel?: (cancel: (reason?: unknown) => Promise<void>) => void, onClose?: (outcome: RelayRequestRecord["outcome"]) => void): Response {
104
158
  if (!response.body) {
105
159
  release();
160
+ onClose?.("failed");
106
161
  return new Response("upstream returned no stream", { status: 502 });
107
162
  }
108
163
  const reader = response.body.getReader();
@@ -122,8 +177,9 @@ function sseResponse(response: Response, release: () => void, onModel?: (model:
122
177
  if (!line.startsWith("data:") || line.slice(5).trim() === "[DONE]") continue;
123
178
  try {
124
179
  const value = JSON.parse(line.slice(5).trim()) as { model?: unknown; choices?: unknown[] };
125
- // Gateway heartbeat events can name a synthetic "keepalive" model with no choices.
126
- if (Array.isArray(value.choices) && value.choices.length && typeof value.model === "string" && value.model.length < 256) {
180
+ // Transport heartbeats are not model identity: a gateway keepalive can name a synthetic model on
181
+ // an event with no generation activity (no choices, or only empty deltas without a finish reason).
182
+ if (typeof value.model === "string" && value.model.length < 256 && isGenerationEvent(value.choices)) {
127
183
  inspectedModel = true;
128
184
  onModel(value.model);
129
185
  return;
@@ -137,15 +193,18 @@ function sseResponse(response: Response, release: () => void, onModel?: (model:
137
193
  const next = await reader.read();
138
194
  if (next.done) {
139
195
  release();
196
+ onClose?.("completed");
140
197
  controller.close();
141
198
  } else { inspect(next.value); controller.enqueue(next.value); }
142
199
  } catch (error) {
143
200
  release();
201
+ onClose?.("failed");
144
202
  controller.error(error);
145
203
  }
146
204
  },
147
205
  async cancel(reason) {
148
206
  release();
207
+ onClose?.("cancelled");
149
208
  await cancel(reason);
150
209
  },
151
210
  });
@@ -169,6 +228,43 @@ export async function startModelRelay(options: ModelRelayOptions): Promise<Model
169
228
  const states = new Map<string, RelayBackendStatus>();
170
229
  const activeRequests = new Set<ActiveRequest>();
171
230
  const activeByAlias = new Map<string, number>();
231
+ const journal: RelayRequestRecord[] = [];
232
+ const JOURNAL_LIMIT = 1000;
233
+
234
+ // The record object is the generation fence: every update goes through this entry's own closure,
235
+ // so interleaved requests for the same alias never write into each other's evidence.
236
+ const openRequestRecord = (alias: string): RequestJournalEntry => {
237
+ const start = Date.now();
238
+ const record: RelayRequestRecord = {
239
+ id: randomUUID(), at: new Date(start).toISOString(), alias,
240
+ identitySource: "none", role: "unknown", outcome: "completed", identified: false, durationMs: 0,
241
+ };
242
+ let closed = false;
243
+ const identify: RequestJournalEntry["identify"] = (model, source) => {
244
+ if (closed) return;
245
+ // First observation wins; a stream observation may still replace a configured label (observed
246
+ // beats configured), and a configured label never replaces an observation.
247
+ if (record.identified && (record.identitySource !== "configured" || source === "configured")) return;
248
+ record.actualModel = model;
249
+ record.identitySource = source;
250
+ record.identified = true;
251
+ };
252
+ const close: RequestJournalEntry["close"] = (outcome) => {
253
+ if (closed) return;
254
+ closed = true;
255
+ record.outcome = outcome;
256
+ record.durationMs = Date.now() - start;
257
+ if (record.identified && record.requestedModel !== undefined && record.actualModel !== record.requestedModel) record.mismatch = true;
258
+ journal.push({ ...record });
259
+ if (journal.length > JOURNAL_LIMIT) journal.shift();
260
+ try {
261
+ // An async hook fits the void signature: its rejection is handled too, never unobserved.
262
+ const notified = options.onRequest?.({ ...record }) as unknown;
263
+ if (notified instanceof Promise) notified.catch(() => { /* a persistence hook must never break the relay */ });
264
+ } catch { /* a persistence hook must never break the proxied stream it observes */ }
265
+ };
266
+ return { record, identify, close };
267
+ };
172
268
 
173
269
  const ensureMlxHandle = async (): Promise<MlxHandle> => (mlx ??= await (mlxStarting ??= ensureMlx(options.mlx).finally(() => { mlxStarting = undefined; })));
174
270
  const state = (backend: ModelBackend): RelayBackendStatus => states.get(aliasOf(backend, mlxAlias)) ?? {
@@ -192,7 +288,7 @@ export async function startModelRelay(options: ModelRelayOptions): Promise<Model
192
288
  return selected;
193
289
  };
194
290
 
195
- const upstream = async (request: RelayRequest, backend: ModelBackend, signal: AbortSignal): Promise<{ response: Response; release: () => void; onModel?: (model: string) => void }> => {
291
+ const upstream = async (request: RelayRequest, backend: ModelBackend, signal: AbortSignal, journalEntry: RequestJournalEntry): Promise<{ response: Response; release: () => void; onModel?: (model: string) => void }> => {
196
292
  const alias = aliasOf(backend, mlxAlias);
197
293
  let base: string;
198
294
  let model: string;
@@ -219,6 +315,8 @@ export async function startModelRelay(options: ModelRelayOptions): Promise<Model
219
315
  }
220
316
  const key = backend.kind === "dgx" ? options.omni.apiKey() : "";
221
317
  if (backend.kind === "dgx" && !key) throw new Error("DGX gateway key is unavailable");
318
+ // What the relay will ask the upstream for is known before the call: a failed dispatch keeps it too.
319
+ journalEntry.record.requestedModel = model;
222
320
  count(1);
223
321
  const releaseOnce = () => {
224
322
  if (released) return;
@@ -251,13 +349,16 @@ export async function startModelRelay(options: ModelRelayOptions): Promise<Model
251
349
  }
252
350
  const provider = safeHeader(response.headers.get("x-omniroute-provider"));
253
351
  const actualModel = safeHeader(response.headers.get("x-model-router-selected-model"));
352
+ if (provider) journalEntry.record.provider = provider;
353
+ if (actualModel) journalEntry.identify(actualModel, "header");
354
+ else if (backend.kind === "mlx") journalEntry.identify(model, "configured");
254
355
  setState(backend, { state: "ready", requestedModel: request.model, active: activeByAlias.get(alias) ?? 0,
255
356
  provider: provider ?? undefined, actualModel: actualModel ?? (backend.kind === "mlx" ? model : undefined) });
256
357
  const releaseWithStatus = () => {
257
358
  releaseOnce();
258
359
  setState(backend, { active: activeByAlias.get(alias) ?? 0 });
259
360
  };
260
- return { response, release: releaseWithStatus, ...(actualModel ? {} : { onModel: (value: string) => setState(backend, { actualModel: value }) }) };
361
+ return { response, release: releaseWithStatus, ...(actualModel ? {} : { onModel: (value: string) => { setState(backend, { actualModel: value }); journalEntry.identify(value, "stream"); } }) };
261
362
  };
262
363
 
263
364
  const server = Bun.serve({
@@ -311,8 +412,10 @@ export async function startModelRelay(options: ModelRelayOptions): Promise<Model
311
412
  return Response.json({ error: "input exceeds the model context budget" }, { status: 400 });
312
413
  }
313
414
  const fallback = backend.kind === "mlx" && options.fallbackDGXAlias ? { kind: "dgx", alias: options.fallbackDGXAlias } as ModelBackend : undefined;
314
- try {
315
- const result = await upstream(body, backend, controller.signal);
415
+ const dispatch = async (selected: ModelBackend, body: RelayRequest) => {
416
+ const journalEntry = openRequestRecord(aliasOf(selected, mlxAlias));
417
+ record.closeRecord = journalEntry.close;
418
+ const result = await upstream(body, selected, controller.signal, journalEntry);
316
419
  let released = false;
317
420
  const release = () => {
318
421
  if (released) return;
@@ -322,26 +425,23 @@ export async function startModelRelay(options: ModelRelayOptions): Promise<Model
322
425
  activeRequests.delete(record);
323
426
  };
324
427
  record.release = release;
325
- return sseResponse(result.response, release, result.onModel, (cancel) => { record.cancel = cancel; });
428
+ return { result, release, journalEntry };
429
+ };
430
+ try {
431
+ const { result, release, journalEntry } = await dispatch(backend, body);
432
+ return sseResponse(result.response, release, result.onModel, (cancel) => { record.cancel = cancel; }, journalEntry.close);
326
433
  } catch (error) {
434
+ record.closeRecord?.(controller.signal.aborted ? "cancelled" : "failed");
327
435
  if (!fallback || controller.signal.aborted || error instanceof ExecutionAdmissionError) {
328
436
  record.cleanup();
329
437
  activeRequests.delete(record);
330
438
  return Response.json({ error: error instanceof Error ? error.message : "backend unavailable" }, { status: 502 });
331
439
  }
332
440
  try {
333
- const result = await upstream({ ...body, model: fallback.alias }, fallback, controller.signal);
334
- let released = false;
335
- const release = () => {
336
- if (released) return;
337
- released = true;
338
- result.release();
339
- record.cleanup();
340
- activeRequests.delete(record);
341
- };
342
- record.release = release;
343
- return sseResponse(result.response, release, result.onModel, (cancel) => { record.cancel = cancel; });
441
+ const { result, release, journalEntry } = await dispatch(fallback, { ...body, model: fallback.alias });
442
+ return sseResponse(result.response, release, result.onModel, (cancel) => { record.cancel = cancel; }, journalEntry.close);
344
443
  } catch (fallbackError) {
444
+ record.closeRecord?.(controller.signal.aborted ? "cancelled" : "failed");
345
445
  record.cleanup();
346
446
  activeRequests.delete(record);
347
447
  return Response.json({ error: fallbackError instanceof Error ? fallbackError.message : "fallback unavailable" }, { status: 502 });
@@ -351,10 +451,15 @@ export async function startModelRelay(options: ModelRelayOptions): Promise<Model
351
451
  });
352
452
  const url = `http://${host}:${server.port}/v1`;
353
453
  const status = (): ModelRelayStatus => ({ url, models, backends: [...states.values()].map((value) => ({ ...value })) });
354
- return { url, token, models, status, close: async () => {
454
+ const requests = (): RelayRequestRecord[] => journal.map((record) => ({ ...record }));
455
+ return { url, token, models, status, requests, close: async () => {
355
456
  const closing = [...activeRequests].map(async (request) => {
356
- request.controller.abort(new Error("model relay closed"));
457
+ // The relay-initiated cancellation closes the record first: the abort below settles the stream
458
+ // as a completed read, and the first terminal transition is the one that counts. Cancelling the
459
+ // upstream reader before the abort keeps the aborted fetch body from rejecting unobserved.
460
+ request.closeRecord?.("cancelled");
357
461
  await request.cancel?.(new Error("model relay closed"));
462
+ request.controller.abort(new Error("model relay closed"));
358
463
  request.release?.();
359
464
  request.cleanup();
360
465
  });