@staix/agent-hub 0.12.10 → 0.12.12

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,21 @@ Issue and pull request numbers in the entries for 0.7.7 and earlier refer to the
4
4
 
5
5
  ## Unreleased
6
6
 
7
+ ## 0.12.12
8
+
9
+ - Restore owned native cohorts when a study supervisor is interrupted: bounded orderly cancellation with signal-time verified child identity, restoration settled only when the marker and ledger agree, a bounded identity-checked fallback with explicit restoration-failure records, and the interrupted phase, cause and restoration status persisted before the incomplete exit. Grading never starts after interrupted generation, and no resume, root reuse or retry is added (#168).
10
+ - Diagnose native sandbox probe failures with a bounded structured outcome: a fixed-enum, count-only probe trace bound to peer, session and setup window explains why the protected-file evidence predicate was not satisfied (agent behavior, tool event coverage or normalization, with unresolved causes explicitly unknown), exposed as the unavailable setup reason in serialized readiness. The #138 and #150 evidence predicates are unchanged and the safe view cannot carry paths, arguments, answers or arbitrary strings (#169).
11
+ - Persist scalar progress for long native studies separately from the verified milestone phase: command kind, repeat ordinal, retained/planned committed-record counts, updated timestamp and verified child identity, with a bounded `status` command that reports stale, crashed, pid-reused and unreadable observations as explicit unknown/stalled and never advances a milestone from progress (#170).
12
+ - Export an audited, versioned per-arm study summary (JSON, optional CSV) after the final audit: planned/scored/passed/unavailable/missing counts reconciled with the audited cells, setup and active medians with valid-completed and common-pair denominators, and native usage, relay usage and provider availability kept distinct. The summary is hash-bound to the manifest, audit, grades and ledger, refuses inconsistent bindings or unsealed studies, and is built allowlist-only so private row content cannot enter it (#171).
13
+
14
+ ## 0.12.11
15
+
16
+ - End native benchmark attempts on terminal Pi or Qwen failures, retain the first failure cause at the final deadline tick, and gate peer-failure patches on active-window tree preservation (#160).
17
+ - Validate ACP native counters without treating context occupancy as consumed tokens; record independent request usage and provider coverage with request-bound primary/fallback linkage (#161, #162).
18
+ - Add `mlx.enabled=false` to omit local inference while preserving remote `hub/auto` fast/coding selection, with explicit launch/recovery conflict diagnostics (#163).
19
+ - Separate exact native-response, primary-route and fallback verdicts in Pi smoke; add optional `--require-primary` (#164).
20
+ - Add a durable, version-bound native study supervisor and scalar-only evidence audit; keep generation ahead of evaluation, refuse reused roots and preserve private artifacts (#165).
21
+
7
22
  ## 0.12.10
8
23
 
9
24
  - Preserve each Pi/Qwen peer's observed sandbox-probe result in v3 readiness; failed or missing native probes never synthesize a denial (#150).
package/README.md CHANGED
@@ -182,3 +182,22 @@ Before upgrading an older hub, inspect the plan with the matching CLI.
182
182
  Incompatible or unverified processes are shown explicitly and are never killed by
183
183
  PID-name matching. See the [operations guide](docs/operations.md) and the
184
184
  [multi-project specification](docs/specs/2026-09-19-multi-project-design.md).
185
+
186
+ ## Agent instructions
187
+
188
+ `AGENTS.md` holds what every agent session loads: commands, conventions,
189
+ architecture and the rules that hold anywhere in the code. Rules that hold in one
190
+ area (benchmarks, adapters, the bus, the local worker, models, the task board,
191
+ budget, the daemon and recovery, tests) live in [`docs/agent-notes/`](docs/agent-notes/),
192
+ and `AGENTS.md` names the paths that call for each note, so a session loads only
193
+ the notes for what it edits (#158). Background behind some of those rules:
194
+
195
+ - Benchmark trust entry: the review rounds of #115 kept finding paths that broke
196
+ the runner's settlement of the Claude trust entry.
197
+ - Reply addressing: leaving `to` empty fanned a reply out to every peer, which is
198
+ what made one directed question cost every agent a turn.
199
+ - Bun test hang: the stacks of two local hangs show a test timeout starting the next
200
+ test inside `Bun.spawnSync`'s own event loop (#115,
201
+ `docs/verification/2026-10-03-0.12.6.md`); minimal reproductions did not hang, so
202
+ the full trigger is not pinned down. On Bun 1.4.2 the next test still starts there
203
+ (#121, checked directly); the spin was not reproduced on demand on either version.
@@ -0,0 +1,15 @@
1
+ # Native adapters and spawned agents
2
+
3
+ Scope: the Codex, ACP and Pi adapters and the Pi extension (turn `generation` and `activity` events), peer turn state, approvals, and the agent processes the hub spawns and stops. Read before editing `src/adapters/`, `src/pi/`, `src/hub/peers.ts`, `src/hub/child-process.ts`, `src/hub/lifecycle.ts` (shutdown deadline), `src/memory/capture.ts` (session end) or the process-tree stop in `scripts/benchmarks/teardown.ts`, or before adding a native peer.
4
+
5
+ - Codex: the adapter never sends its own `initialize`. It is a proxy; hub requests use negative ids and their responses must not reach the TUI.
6
+ - Codex 0.154.0 `agentMessage` items carry `text` and `phase` (not `content[]`); only the last non-`commentary` message of a turn is shared.
7
+ - Approval titles are agent-written text shown to a person: the daemon escapes control characters, and write/edit/bash show what will be written or run.
8
+ - An ACP permission request may carry no `rawInput` (Kimi 2.0.1): the arguments were on the earlier `tool_call` update, which `acp.ts` remembers by `toolCallId`. Kimi 2.1.1 sends no `rawInput` before the answer at all: the argument JSON streams as `content` text on `tool_call_update`, and only complete JSON counts. A payload that still cannot be resolved, or is too long to show whole (marked `[cut, N chars]`), withholds its `allow_always` option - never render a bare tool name as if it described the call.
9
+ - Auto-approval covers the hub's own tools by exact `mcp__agent-hub__<name>` title and only ever picks `allow_once`. Identity comes from the permission request title or, when an agent titles the request with the argument JSON (Qwen), from the announced `tool_call` title bound to the call id and resolved against the session's configured MCP servers, never from payload text, a prefix or an unrelated display title. Never widen it to a prefix, and never set Kimi's session mode to `auto` or `yolo` instead: those approve everything.
10
+ - Every agent the hub spawns (the Codex app-server, ACP agents such as Kimi, Pi) runs in its own process group and is stopped as one (`stopOwnedProcess(proc, { group: true })`, `trackGroup` after spawn): `codex` is a node launcher, and mid-turn its native app-server does not exit on SIGTERM within the grace period, so a SIGKILL to the launcher alone left the app-server running under init and holding the daemon alive (#113, #115). A new adapter that spawns an agent does the same.
11
+ - To stop a process tree, freeze it (SIGSTOP) before reading what it contains, then SIGKILL, and call it done only when the table shows none of it (with no table at all, the stop can only ask the group itself, which cannot see groups below it). A process a read after the freeze shows for the first time is frozen and read again before anything is killed. A snapshot taken while the tree runs misses what it starts next: three review rounds of #113 found such a gap (groups of its own, an unreadable first read, a child started during the grace period).
12
+ - Work that must survive `ahub kill` (claude-mem `session-end`) is awaited in `stop()`; fire-and-forget dies with the process.
13
+ - A peer must set `busy` synchronously inside `deliver()`, otherwise the bus drains the next envelope into a running turn (Kimi answers `turn.agent_busy`).
14
+ - A watchdog-cancelled turn still reports later. Anything a turn does on completion must check it is still the current turn (`turn` generation in `acp.ts`, `activeTurns` in `codex-appserver.ts`). A rejected prompt or abnormal prompt end reports through the adapter's `onTurnFailure` (same shape in `acp.ts` and `pi.ts`), even when ACP streamed partial text; a stale cancelled turn never reports it.
15
+ - Native owner teardown must handle a lost shutdown acknowledgement using verified process identity. Pending launches must be revocable, and active tools/streaming output must count as watchdog activity.
@@ -0,0 +1,8 @@
1
+ # Benchmark runners
2
+
3
+ Scope: the CooperBench runners, their fixtures, teardown and recovery. Read before editing `scripts/benchmarks/` or `test/benchmarks/`, or before running a benchmark.
4
+
5
+ - Never auto-register Orca benchmark fixtures, including ad-hoc or copied runners; require an explicit pre-existing repo and exact worktree registration, and preserve fixture directories, Git state and benchmark evidence during cleanup.
6
+ - The manifest-v3 Pi/Qwen CooperBench driver (`scripts/benchmarks/native-pi-qwen.ts`, #140) is headless and never touches Orca. Its pinned build is the effective one: each native's `--version` runs under the final isolation environment (Qwen inside its seatbelt profile with `QWEN_HOME` set; the PATH bootstrap lies about the managed build under isolation). Its protected-file probe counts only with structured denial evidence (a guard denial or a failed read tool call with EPERM/EACCES), never a model-written marker alone, and each peer's serialized `sandboxProbe` readiness reports only that peer's observed probe outcome (`denied` with its evidence, else explicit `failed`/`unknown`; #150), never a synthesized denial. Served-model evidence is each request's own journaled `RelayRequestRecord` (#139): a heartbeat-only or cancelled-before-identification record certifies nothing, and every identified record is evidence whatever its outcome; a confirmed mismatch on a request cancelled after identification fails the model gate. A backend's last-served label is never a request's evidence. Qwen's peer tool approval is the exact canonical `mcp__pilot-peer-bus__hub_send` through the adapter's #138 binding, never a title-prefix workaround.
7
+ - The benchmark run's restoration state: `restoration.json` says `restored: false` from before the runner locks anything until it writes its outcome, and the recovery trusts `restored: true` only when `restoration-ledger.json` agrees (`unrestored` in `scripts/benchmarks/teardown.ts`). A run directory that is not restored is never reused: its locked modes would become the originals. A ledger from before 0.12.5 (no runner identity) owes its locks but, when `restoration.json` says `restored: true` or names a runner, keeps its trust entry as it was then (`owed`; with no `restoration.json`, its runner did not finish and the entry is owed too): the reuse check and the recovery read both through these two functions, never on their own (#120).
8
+ - The benchmark runner's Claude trust entry: a write still `pending` in the runner's own process never landed, so it is `not_written` on every path (contained or not, its temp file removed or left), and neither the runner nor the recovery touches an entry then. Only a dead runner's `pending` leaves that question to the recovery, which takes back only an entry exactly as the runner would have written it, and none while the runner's temp file is still there (the rename never happened). `changed_concurrently` is the user's entry and is never taken back. The runner's settlement is `settleTrust` in `scripts/benchmarks/teardown.ts`, tested as a table: change it there, never inline.
@@ -0,0 +1,10 @@
1
+ # Budget, pause and handoff
2
+
3
+ Scope: quota readings, pause, handoff and resume, and the status line tee. Read before editing `src/hub/budget.ts`, `src/cli/statusline-tee.ts`, `src/hub/bus.ts`, `src/hub/delivery-journal.ts` or `src/hub/restart.ts` (pause persistence).
4
+
5
+ - Budget: checkpoint first, pause second (a paused peer receives nothing). A handoff that fails is left unmarked so the next tick or hub run retries it; never record it as done.
6
+ - A handoff needs somebody to hand over to: `canHandOff` is false right after a restart, when no peer is attached yet, and the handoff waits for a later tick instead of stripping tasks of their owner.
7
+ - A reading has its own timestamp. Numbers that arrive through a file (`claude-usage.json`) carry the file's `at`; a window whose `resetsAt` has passed says nothing any more.
8
+ - On resume the notice is published before the peer is released, so it leads the first delivery.
9
+ - The coordinator only lifts its own pauses: `manualPaused` in the daemon keeps a `ahub pause` in place, and `ahub resume` refuses while a budget record is open.
10
+ - The status line tee must never fail or slow the render: no throw, original command run with the same stdin, 5 s cap.
@@ -0,0 +1,13 @@
1
+ # Bus, digests and replies
2
+
3
+ Scope: delivery, digests and condensation, reply addressing, priority and limits. Read before editing `src/hub/bus.ts`, `src/hub/envelope.ts`, `src/hub/limits.ts`, `src/hub/delivery-journal.ts`, `src/hub/inference.ts` (digest condensation), or how an adapter addresses a reply or sets its priority.
4
+
5
+ - A failed digest is retried one envelope at a time, so a poison envelope cannot take its neighbours down with it.
6
+ - An `important` envelope being steered is not in the queue while the steer is in flight; queue it first and an idle transition delivers it twice.
7
+ - `replyParent()` decides what a reply answers (highest hop, never the `hub` preface). Use it for deliveries and steers alike, or the hop cap can be reset.
8
+ - What a peer is handed (`out`) and what it stands for (`originals`) differ once a delivery is condensed: the bus keeps both, registers `out` so `reply_to` resolves, puts the originals back on any failure (a thrown `deliver` or a later `onFailed`), and counts no attempt when the peer merely got busy while the delivery was prepared. A delivery with an `important` envelope is never condensed.
9
+ - A condensed digest is sent by `digest`, not `hub`: `replyParent()` skips `hub` items, and a reply to a digest must keep the highest hop of what it replaced. On a failed delivery the bus puts back the originals, never the digest.
10
+ - Limits admit what is sent: the envelope `newEnvelope` builds (a reply inherits its parent's sender, `digest` resolves to the originals, `capPriority` applies), never the raw `to` and priority. Admit first and build later, and implicit replies all count as broadcasts.
11
+ - A reply is addressed, never broadcast. An adapter that answers a delivery passes `to: replyAudience(envs)`; anything else with an `inReplyTo` inherits that envelope's sender. Leaving `to` empty fans the message out to every peer.
12
+ - `digest` is not a peer: addressing a reply at the envelopes the peer was handed sends it nowhere once a delivery was condensed. The bus resolves `digest` back through `lastDelivery.originals`; anything else that reads a reply's `to` has to do the same.
13
+ - A hub-native peer (Pi, the local worker) does not get to call its own message `important`: `capPriority` caps it unless the delivery it answers held an `important` envelope addressed to it. Check the delivery, not `replyParent`, which ties on hop and takes the later item. Do not bypass it by setting `priority` in the adapter.
@@ -0,0 +1,13 @@
1
+ # Daemon, control protocol and recovery
2
+
3
+ Scope: the control WS and its protocol, state files, the Claude channel's reconnect, `ahub setup`, git snapshots, upgrade and recovery. Read before editing `src/hub/daemon.ts`, `src/hub/control-client.ts`, `src/adapters/claude-channel.ts`, `src/hub/restart.ts`, `src/hub/recovery-store.ts`, `src/hub/snapshots.ts`, `src/hub/manager.ts`, `src/hub/lifecycle.ts`, `src/hub/crash.ts`, `src/cli/setup.ts`, `src/cli/upgrade*.ts`, `src/cli/terminal-recovery.ts`, `src/cli/recovery-package.ts` or `src/cli/facts-hook.ts`, or before adding a native peer.
4
+
5
+ - The plugin bundle is installed apart from the daemon. Any change to a control WS message shape bumps `PROTOCOL` in `control-client.ts`.
6
+ - `ahub setup` reads Claude Code's state from `--json` listings and takes one step at a time, re-reading after each; never match paths or names by substring.
7
+ - A copy of the git index keeps the original's mtime (`snapshot()`): git's racy-entry check compares entries with the index file's time in whole seconds, and a fresh copy makes a same-size edit look clean.
8
+ - State files that clients read (`status.json`, `control-token`) are written after the port is bound, and `status.json` via temp file + rename.
9
+ - The channel's reconnect loop stops on closes a retry cannot fix (`TERMINAL_CLOSES` in `claude-channel.ts`); a new daemon close code that means "do not come back" belongs there, or two clients fight over it forever.
10
+ - Close 4000 is the one close that ends on its own, so it is not in `TERMINAL_CLOSES`: the session stands by and reconnects only once `status.json` reports the peer offline and not `claiming`. Reconnecting blind would evict whoever holds the id now, and the two would trade it forever. `claim()` happens at hello and `attach()` only after the preface, so the peer reads as offline in between: the daemon writes the status file at the claim and the flag covers that window.
11
+ - Adding a native peer requires testing the command emitted by the actual recovery driver, not only its inspection or launch helpers. Never let a generic non-Codex branch treat a new peer as Claude.
12
+ - Cross-version recovery must use the authenticated source protocol for prepare/commit/abort and the target protocol after startup. Prove the transition against a real prior release before claiming compatibility.
13
+ - A zero-turn Claude session has no transcript, so it can never be resumed: every identity gate that compares session ids (restore, coordinator verify, daemon readiness) must tolerate a fresh session while the transcript is absent and stay strict the moment one exists. The commit snapshot records `sessionPersisted` for the restored daemon, and snapshots too old to have it are re-derived from disk. Never treat "session id changed" as proof of a lost conversation without checking the transcript first.
@@ -0,0 +1,12 @@
1
+ # Local worker and sandbox
2
+
3
+ Scope: the hub-native local worker, its tools, path guard, denylist and seatbelt sandbox. Read before editing `src/adapters/local-worker.ts`, `src/local/`, `src/memory/capture.ts` or `src/hub/facts.ts` (both read the denylist).
4
+
5
+ - Nothing the local worker executes may run outside `sandboxedExec`; a new tool that spawns a process goes through it, and a tool that touches a path goes through `guardPath`.
6
+ - The sandbox denies home reads by default (toolchain dirs, the project and its real git dir excepted) and all network, loopback included: claude-mem and the Codex app-server listen on loopback without auth. Tests that bind a local port therefore fail when the worker runs them, also with `local.bash_network: true` (the egress proxy keeps loopback closed, and the proxy variables send a test's plain-HTTP `fetch` to the proxy, which refuses it); that is the intended trade-off, and `"direct"` is the switch until 0.13.0 removes it.
7
+ - SBPL strings go through `sbplString` (the plain `"..."` form). In the raw `#"..."` form a backslash escapes nothing, so a `"` in a path ends the literal and the whole profile fails to parse.
8
+ - One denylist, `src/local/deny.ts`: the path guard, the seatbelt profile and the memory filter all read it. Seatbelt sees absolute paths, so `local.deny` entries are anchored under the project root (a bare `private/` once denied all of `/private/var`).
9
+ - `guardPath` walks with `lstat`: `existsSync` follows symlinks, so a dangling link looked like a new file and the write landed at its target.
10
+ - git arguments never pass through `guardPath`; `gitArgsProblem` refuses absolute paths, `..`, `--no-index` and denylisted `rev:path` forms.
11
+ - A worker turn builds its messages in a local array and joins the history only as a whole. Never push to `history` mid-turn: one tool call without its result poisons every later request.
12
+ - After a tool with side effects ran, a failed turn is reported, never redelivered.
@@ -0,0 +1,9 @@
1
+ # Model relay, inference and gateway
2
+
3
+ Scope: the model relay and its journal, the hub's own inference and its slots, OmniRoute and Switchyard. Read before editing `src/models/`, `src/hub/inference.ts`, `src/omniroute/` or `src/switchyard/`.
4
+
5
+ - The model relay journals per-request identity (`RelayRequestRecord`): a request's served model comes only from its own gateway header, its own generation SSE event (#137 heartbeat classification), or the locally validated MLX configuration. A backend's mutable last-served label is never a request's evidence, HTTP 200 plus the requested alias identifies nothing, and a request cancelled before identification stays explicitly unidentified (`identified: false`, `outcome: "cancelled"`). Journal records carry no prompts, tools, keys or Access headers.
6
+ - Inference is optional and fail-open: it returns its input or `undefined` on any failure and backs off, so a delivery never waits on a dead model twice. Its output is only ever capped text or a value checked against a closed list.
7
+ - Secrets stay inside `OmniRoute`: never put the key or Access values in a log line, an error message, a return value or the generated Switchyard file (the key goes by env var name).
8
+ - Switchyard's docs drift from the released binary. Any change to `switchyardToml` is checked with the real `switchyard-server --dry-run`, not only the stand-in in `test/fakes/`.
9
+ - Shared inference slots must recover after a hub process dies, without evicting a live owner. Cancellation has to reach slot acquisition from the relay caller, not just exist in the helper signature.
@@ -0,0 +1,7 @@
1
+ # Task board and hub tools
2
+
3
+ Scope: the task flow, assignment, the state machine and the hub's MCP tools. Read before editing `src/hub/tasks.ts`, `src/hub/board.ts`, `src/hub/routing.ts` or `src/hub/hub-tools.ts`.
4
+
5
+ - Tool callers are models: MCP `inputSchema` is not enforced on the way in. Normalize at the boundary (`cleanRefs`) before anything reaches the board, and never throw after a board write.
6
+ - `assign()` never defaults to the task's current owner, or a decline can only come back to the decliner.
7
+ - The state machine allows `in_progress -> approved` only for classes without a reviewer; `review()` checks `in_review` itself.
@@ -0,0 +1,7 @@
1
+ # Tests
2
+
3
+ Scope: tests, fakes and the test gate. Read before adding or changing a test or fake under `test/`, or editing `scripts/check.sh` or `scripts/hang-watch.sh`.
4
+
5
+ - Tests that are not about batching build the bus with `batchMs: 0`; with the default 15 s window a lone status envelope looks like a lost message.
6
+ - In Bun 1.3.14 a test timeout that fires while `Bun.spawnSync` runs can start the next test inside spawnSync's own event loop, where another spawnSync then spins at full CPU for good (trigger not pinned down; see README). Bun 1.4.2 still runs the next test inside the outer spawnSync's event loop. `scripts/check.sh` runs tests with a 20 s timeout and under `scripts/hang-watch.sh`; a test that spawns and reads the process table gets a timeout of its own.
7
+ - A child process gets a scrubbed environment, so test knobs for fakes travel in a wrapper script, not in `process.env`.
@@ -95,6 +95,8 @@ The v3 driver supplies a trusted `expectedServedModels` map to the relay: `fixed
95
95
 
96
96
  Served-model evidence is per request, from the relay's journaled `RelayRequestRecord` (#139), not from a response wrapper: a completed request that was never identified fails the attempt's model gate (a heartbeat-only stream identifies nothing), an observed mismatch flags it, and a request cancelled before identification stays explicitly unidentified and never certifies another request. Qwen's peer tool is approved through the adapter's shipped tool-identity binding (#138): the announced `hub_send (pilot-peer-bus MCP Server)` title resolves to the canonical `mcp__pilot-peer-bus__hub_send`, the only name on the exact whitelist. The joint arm's MCP peer bus is the versioned `scripts/benchmarks/peer-bus-mcp.py`, pinned with the driver in `prepared.json` and `cohort.json` next to the candidate source pins (`src/adapters/pi.ts`, `src/adapters/acp.ts`, `src/models/relay.ts`).
97
97
 
98
+ A required peer's terminal active-turn failure ends the whole attempt at once (#160), instead of waiting out the wall limit for an answer the failed peer cannot produce: in the fixed study two attempts idled 153 s and 193 s after Pi's 100-step ceiling before recording a wall-timeout. The driver latches the first failure either native's active-phase turn callback reports (Pi's and Qwen's `onTurnFailure`: a rejected prompt turn, or one that ended without normal completion, including partial ACP text), with the peer, the original failure class and the failure time preserved in the record's `peer_failure` field and beside the `active_end` event. The latch is fenced by phase and by generation: nothing before the active start latches (the setup probe's expected denied tool read is a tool error, never a terminal turn failure, and even a turn failure there does not end the attempt), and the end cause freezes at `active_end`, so a later stop or watchdog callback during teardown is recorded as a teardown event only and can never replace it. A joint attempt cancels both owned peers through the same owned-process teardown — a failed actor never leaves the other actor performing an undefined partial treatment — and no failure is ever silently reported as completed. A genuine wall limit with no latched failure stays `wall-timeout`/`timeout`. The preserved partial submission (the source tree at the end of the active window and its scoped patch) remains independently gradeable exactly when the existing setup, identity, metadata, source and teardown gates pass: `peer-failure` joins `completed` and `timeout` in the runner's graded end classes, and a flag beside it (modified metadata, an unverified tree) still makes the record an infrastructure error. Records from before #160 carry no `peer_failure` field and no peer-failure end class; the ledger and grader read them unchanged.
99
+
98
100
  Grading flows through the same official evaluator adapter and controls; a quality failure is never a retry selector and no evaluator feedback reaches the candidate agents during generation. Reports keep quality, model-identity and request-linkage coverage separate per arm, with unavailable attempts (failed, missing or unavailable) retained in the planned denominator, and name the usage units (Pi's incremental `onTokens` counter, Qwen's session `usage_update` running total — never added together) and the tool-surface difference (Pi's hub-moderated tools against Qwen's own seatbelted auto-edit tools). A native actor's usage is counted only in attempts where that actor participated (#152): every v3 grade row carries `native_participants`, the actors the attempt record shows actually started — the arm bounds the candidates (a solo-qwen attempt has no Pi peer), and among them a peer participated only once its readiness entry carries the sessionId its adapter reported after start. A peer that was constructed but never started (a setup failure: `elapsedMs` 0, no active start) is absent; its token counter's initial 0 is never read as a measurement. An absent actor reports zero known observations and a null total (`*_tokens_known: 0`, `*_tokens: null`), a participant whose native reading is missing stays unknown rather than zero, and a participant's genuinely observed 0 stays counted. The summary's `pi_participating`/`qwen_participating` fields give each coverage figure's participation denominator. Grade rows from before #152 carry no `native_participants`; the report takes their participants from the arm, which is exact for them: their `native_usage` was attached only to scored attempts, whose readiness gate had proved every required actor's session. A live cohort, the official Docker controls and a native readback of an attempt's records remain manual live legs requiring accounts and the pinned archives; they are not part of the checked-in tests.
99
101
 
100
102
  ## Coordination ledger
@@ -141,3 +143,80 @@ The fixed ten-pair sample is a convenience sample, not the full 652-pair suite.
141
143
  The optional `--codex-bin` selects an absolute native executable when PATH contains several Codex installations. Its reported version must match the manifest. Codex automatic memories and external agent memory import are disabled for the benchmark thread. Native usage totals include the unscored sandbox probe and are labelled as whole-session counters.
142
144
 
143
145
  `--setup-only` runs every selected arm through native readiness and read-denial probes without assigning feature work. Its cohort is marked as calibration and the grader refuses it. A zero-turn Claude session is bound through its verified Orca launch and then checked against native transcript session IDs after the probe. The instance-fenced metadata file enables the daemon's optional usage reader.
146
+
147
+ ## Independent usage and provider coverage (#161, #162)
148
+
149
+ ACP's native cumulative counter accepts checked nonnegative finite totals (including measured zero) and input/output pairs. The pinned Qwen 0.24.7 ACP source emits `usage_update {used,size}` from context occupancy (`collectContextData`), which is recorded as an unsupported cumulative-counter shape with context counters only, never consumed tokens. `usage_update` and prompt-result observations expose only a bounded counter projection with a supported, unsupported or invalid shape verdict. Cumulative readings replace the native counter; they are never summed. A started Qwen session with no compatible reading remains `no-reading`.
150
+
151
+ Relay `requestUsage` is an independent per-dispatch observation of OpenAI-shaped stream `usage` fields, including a final usage-only event after model identification. Missing fields make it partial; malformed readings remain invalid or unknown. This measurement never substitutes for, or adds to, a native session total. The relay records counters only when upstream supplies them, and does not claim billing or force an unsupported stream option.
152
+
153
+ Provider provenance is currently `header` (`x-omniroute-provider`) or `none`; no generation event provider contract has been verified. Provider absence leaves model qualification unchanged. `request_observability` in grading and ledger exports independently reconciles provider known/missing denominators for completed, cancelled and failed dispatches, including cancelled-after-identification. Older records remain readable: their own provider value supplies legacy header coverage, while absent request counters stay unknown. Heartbeats and another request never fill a missing provider.
154
+
155
+ The bounded live capability probe is `bun scripts/benchmarks/probe-qwen-usage.ts --qwen-package PINNED_PACKAGE --run NEW_PRIVATE_DIR --config-dir PROJECT`. It verifies Qwen 0.24.7 under the final seatbelt environment, uses host-owned OmniRoute authentication through a loopback relay, and saves only counter/source/availability projections. Use a fresh disposable directory; sealed historical unknowns remain unchanged.
156
+
157
+ ## Durable native v3 study supervision (#165)
158
+
159
+ Use `scripts/benchmarks/study_supervisor.py` for a manifest-defined study. Choose a
160
+ new directory under an existing private, durable parent outside this repository
161
+ and the OS temporary directory. The supervisor atomically claims that directory;
162
+ existing roots, including incomplete ones, are refused. It never retries a cell or
163
+ resumes a cohort. Prior archives remain read-only.
164
+
165
+ ```sh
166
+ python3 -B scripts/benchmarks/study_supervisor.py \
167
+ --manifest scripts/benchmarks/manifest-v3-pi-qwen.json \
168
+ --output "$STUDY_ROOT" --archives "$ARCHIVE_ROOT" \
169
+ --private-inputs "$PRIVATE_CASE_ROOT" --upstream-root "$UPSTREAM_ROOT" \
170
+ --probe-target "$PROTECTED_PROBE" --qwen-package "$QWEN_PACKAGE" \
171
+ --protect "$PRIOR_ARTIFACT_ROOT" --preflight-only
172
+ ```
173
+
174
+ Repeat the command without `--preflight-only` to execute. Every required archive,
175
+ upstream commit, prompt pin, native build, protected root, probe and cached
176
+ Docker image must be available before fixture preparation. Use `--python` to
177
+ select the evaluator Python environment with the upstream dependencies installed. Explicitly optional
178
+ prior roots use `--optional-protect`; their presence or absence is recorded.
179
+ For a bounded native lifecycle check, use a separate manifest with a declared
180
+ small plan and `--plan` to select it; preserve all case/model/order/budget pins.
181
+ `--generation-only` stops at restored generation, records outcome generation-only,
182
+ and never grades or seals that root. The fixed study plan remains unchanged.
183
+ Version mismatch fails unless `--bind-current-hub` explicitly changes only
184
+ `hub_version` and `versions.hub`. Both original and runtime hashes, source
185
+ identity, invocation, and preparation's archive-path additions are sealed in the
186
+ private provenance file. Native versions are checked in isolation at preflight
187
+ and rechecked by the existing driver in each arm's final environment.
188
+
189
+ The phase order is claimed, prepared, generated, restored, graded, sealed. All
190
+ repeat roots are freshly prepared, each generation runs sequentially, and all
191
+ planned generations finish before any official controls or grading. Restoration
192
+ markers, their ledgers, and live PID/start-time identities must agree before
193
+ grading. Every retained cell also passes a grade-independent readback of its
194
+ actual fixture metadata, guarded source diff, sealed baseline and native patch
195
+ bytes/hash. This same gate runs before reporting restored generation with
196
+ `--generation-only`; stored `metadata_clean` flags alone never qualify it.
197
+ Scored cells add the grade/evaluation bindings to that native evidence chain. A failure preserves an incomplete private root and logs for inspection.
198
+ The procedure uses the existing native driver, official grade/report commands,
199
+ and pooled ledger, without Orca registration or toolchain tree copying.
200
+
201
+ This entrypoint requires the evaluator service and images to be available. It
202
+ never starts or stops a shared VM or service. An operator who explicitly starts a
203
+ service must retain original-state and ownership evidence separately and restore
204
+ only that owned service after proving no unrelated usage; otherwise restoration
205
+ is deferred. Service ownership in this supervisor is recorded as require-existing.
206
+
207
+ Audit historical repeat roots without changing them:
208
+
209
+ ```sh
210
+ python3 -B scripts/benchmarks/study_audit.py \
211
+ --run "$STUDY_ROOT/r0" --run "$STUDY_ROOT/r1" \
212
+ --export "$NEW_SAFE_EXPORT"
213
+ ```
214
+
215
+ The auditor reconstructs the exact matrix and checks restoration, request-based
216
+ model gates, official controls, evaluation/patch bindings and owned live processes.
217
+ It constructs its output from fixed scalar fields and hashes. Nested official
218
+ test output, answers, native events, errors, provider identifiers and machine paths
219
+ are never copied to stdout or the safe aggregate. Unknown usage remains null.
220
+ Export creation is exclusive, so it cannot overwrite a historical artifact. The
221
+ complete private logs, evidence hash index, and safe aggregate stay in the durable
222
+ study root. A failed audit exits nonzero with a fixed error code and no raw error.
@@ -896,3 +896,31 @@ Pi exposes `hub/auto` for stage routing when available. Fixed `dgx/coding`, `dgx
896
896
  and `mlx/fast` aliases still pin the backend. Automatic MLX selection admits the complete
897
897
  input, tool schemas and requested output within the configured context window.
898
898
  Progress judgements suggest reassignment; they never change task ownership.
899
+
900
+ ## Disable local MLX while keeping remote auto routing
901
+
902
+ `mlx.enabled` is an optional boolean and defaults to `true`, including legacy
903
+ configurations. Operators without an available local service can set:
904
+
905
+ ```json
906
+ {
907
+ "pi": { "enabled": true, "backend": "auto" },
908
+ "mlx": { "enabled": false }
909
+ }
910
+ ```
911
+
912
+ This removes `mlx/fast` from the relay inventory. `hub/auto` continues selecting
913
+ `dgx/fast` for efficient work and `dgx/coding` for capable work. Neither a local
914
+ probe nor a local startup occurs. `ahub models status` reports `disabled` and
915
+ `ahub doctor` reports a successful disabled row without probing the endpoint.
916
+ `models setup`, `start`, and `stop` are refused while disabled; shared Ollama
917
+ continues under its existing owner.
918
+
919
+ Migrate an earlier `pi.backend=dgx` workaround to `auto` explicitly. In a project
920
+ routing file, remove `pi_backend="mlx"` pins to use `hub/auto`, or set `dgx` to
921
+ keep a remote class pin. The hub refuses conflicting operator-written routing
922
+ before startup, while inherited shipped MLX class defaults use `hub/auto` when
923
+ the capability is absent. It never rewrites these settings. Explicit
924
+ `--backend mlx`, `--model mlx/fast`, and recorded MLX recovery launches are
925
+ refused before any local startup. Re-enable MLX or explicitly migrate the
926
+ recorded launch before recovery.
package/docs/smoke.md CHANGED
@@ -1130,3 +1130,32 @@ Compare Claude usage events against its explicit native session transcript, dedu
1130
1130
  by assistant message ID. Compare local usage events with actual successful provider
1131
1131
  response counters; absent counters remain unknown and model aliases remain separate
1132
1132
  from reported served-model provenance. Do not infer dollar spend from token counts.
1133
+
1134
+ ## Native Pi auto-route dispatch verdicts (#164)
1135
+
1136
+ Run `bun scripts/smoke-pi-route.ts` for permissive connectivity, or add
1137
+ `--require-primary` to require a successful primary dispatch without fallback.
1138
+ `AHUB_SMOKE_PROJECT` selects the project configuration; the smoke creates a
1139
+ separate temporary Pi workspace, permits no tools and leaves existing hubs alone.
1140
+ The project `mlx.enabled = false` capability excludes local routing from this
1141
+ smoke as well as the daemon. The response must equal `PI_HUB_AUTO_OK` exactly.
1142
+
1143
+ The JSON verdict separates `nativeResponse`, `primaryRoute`, `fallbackRoute` and
1144
+ `fallbackOccurred`. `choices` records routing intent. `dispatches` records actual
1145
+ upstream attempts, with dispatch IDs and `fallbackOfId` linking a fallback to its
1146
+ failed primary; model identification comes only from that dispatch's journal.
1147
+ Backend readiness, a selected tier and a previous model label cannot certify
1148
+ that a generation completed. Failure diagnostics contain only bounded categories
1149
+ and HTTP status; raw errors, URLs, headers and native answer text are omitted.
1150
+
1151
+ For an unavailable MLX primary followed by a completed DGX fallback, expect
1152
+ `nativeResponse: passed`, `primaryRoute: failed`, `fallbackRoute: passed` and
1153
+ `fallbackOccurred: true`. Connectivity passes, but the local route failed.
1154
+ The same evidence fails with `--require-primary`. Primary-only completion reports
1155
+ `fallbackRoute: unknown`; a cancelled primary reports `primaryRoute: cancelled`.
1156
+ If both attempts fail, connectivity fails even if unrelated answer text contains
1157
+ the sentinel. Missing served-model evidence remains `identified: false`.
1158
+
1159
+ Before recording live acceptance, run a bounded unavailable-local/healthy-remote
1160
+ leg and retain the sanitized JSON plus process cleanup confirmation. Fixture
1161
+ verdicts do not certify provider availability or a real local generation.
@@ -383,8 +383,9 @@ inside a peer.
383
383
  rolling 5 h against `budget.kimi_tokens_5h` (off by default). Verified live on kimi
384
384
  2.0.1: one update per turn, payload `{"sessionUpdate":"usage_update","used":<tokens>,
385
385
  "size":<context window>}`; `used` is the session's context occupancy against `size`
386
- (1M), not billed quota; the parser matches it through the `used` fallback and it grows
387
- monotonically within a session (a compaction reads as a new session). `ahub budget set`
386
+ (1M), not billed quota. ACP records this shape as context occupancy diagnostics,
387
+ never as cumulative consumed tokens. Only validated cumulative totals or input/output
388
+ pairs feed native token readings; unsupported shapes remain unknown. `ahub budget set`
388
389
  feeds a reading by hand. `local` has no quota and is never paused.
389
390
  - Gate at `budget.gate` (default 0.9) on any fresh window; readings older than
390
391
  `budget.stale_min` are ignored. Checkpoint first, pause second: the peer gets one
@@ -1228,3 +1229,16 @@ upgrade sources; ordinary clients must use protocol 12 (13 from 0.12.4, with 12
1228
1229
  - `Tasks.splitObservations` reads one observation per hand-over to the peer with the profile it has now, across hub runs of the project (what happened from that hand-over up to and with the next one is that peer's: a decline or escalation away is its failure, never the next owner's), never the routed task itself, leaves out claims, types failures (failed check, changes requested, escalation, release, decline, unresolved integration), ends the work stage at the first done call (checks and integration are not work) and leaves the stages unknown for an accept recorded by the done itself. The `split` event carries the trace.
1229
1230
  - An overlap that forms or changes a cohort records a `split` event for the task, routed or named (a benchmark names every owner); `ahub route explain <id>` appends the trace as it would be now. Rerouting needs held-out evidence through #110 first.
1230
1231
  - 2026-10-03 (T3 calibration data, issue #109 amendment): every hand-over records the new owner's split profile on its history entry: the hub's version, the agent's version (Codex's from app-server's `initialize` answer, Claude Code's from its transcript rows) and the hook profile, which is the hub's own (the coordination mode: a turn-free project runs the facts hooks in Claude). Observations count for a prediction only with the peer's current profile; while a version is unknown there is no profile, and the prediction is unknown. The prediction calibration reads is the one recorded when routing chose the first owner (no single named candidate, no claim; not an escalation, budget relay or reassignment of work already begun) of a task that overlaps another owner's task not started yet (`where: "routing"`), and `route explain` prefers the same unstarted pair; cohort-time predictions stay in the log with `where: "cohort"` and are not calibration data. The user's and plugins' hooks are not seen by the hub (ponytail: read them from the native records as the benchmark ledger does), and only Claude and Codex report a version so far.
1232
+
1233
+ ### Optional local MLX capability (0.12.11)
1234
+
1235
+ `mlx.enabled` is an optional boolean, defaulting to true. When false, the daemon
1236
+ omits local MLX relay capability while remote `hub/auto` retains fast/coding
1237
+ selection. Explicit MLX backend/model, project policy and recovery conflicts are
1238
+ refused before local startup. Model diagnostics report disabled without probing
1239
+ that endpoint. The switch never starts or stops shared Ollama.
1240
+
1241
+ Relay native-session counters and per-dispatch transport usage are separate
1242
+ measurements. Request usage and provider availability are optional metadata,
1243
+ bound to their own dispatch IDs; provider absence does not change model
1244
+ qualification. Primary and fallback dispatches have independent outcomes.
@@ -293,7 +293,7 @@ related: 261004-review-agent-hub-switchyard-comparison.md, 261004-plan-agent-hub
293
293
  - `crates/libsy/src/prompts/*`, `crates/libsy-llm-client/src/{client,backend,run}.rs`(호출 정책).
294
294
  - 소스 분석 보고(조사 에이전트, 2026-10-04)와 주요 상수·공식의 직접 확인.
295
295
  - agent-hub(main `ab16c7b`)
296
- - `src/adapters/local-worker.ts`(`call()`, `commit()`), `src/local/tools.ts`(도구 이름), `src/models/relay.ts`(`selectBackend`, 폴백), `src/hub/daemon.ts`(relay 시작, Pi 설정), `src/hub/inference.ts`, `src/hub/facts.ts`, `src/adapters/codex-appserver.ts`, `test/fakes/model-server.ts`, `docs/events.md`, `AGENTS.md`(inference·history·PII 규칙).
296
+ - `src/adapters/local-worker.ts`(`call()`, `commit()`), `src/local/tools.ts`(도구 이름), `src/models/relay.ts`(`selectBackend`, 폴백), `src/hub/daemon.ts`(relay 시작, Pi 설정), `src/hub/inference.ts`, `src/hub/facts.ts`, `src/adapters/codex-appserver.ts`, `test/fakes/model-server.ts`, `docs/events.md`, `AGENTS.md`(PII·PII turn 규칙), `docs/agent-notes/models.md`(inference 규칙), `docs/agent-notes/local-worker.md`(history 규칙).
297
297
  - 기본 설정: Pi `dgx_coding = "coding"`, `dgx_fast = "fast"`, MLX `qwen3.5:4b-mlx`(8k).
298
298
 
299
299
  ## Implementation contract (approved 2026-10-04)
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@staix/agent-hub",
3
- "version": "0.12.10",
3
+ "version": "0.12.12",
4
4
  "description": "Native multi-agent hub: Claude Code, Codex, Kimi Code, Pi and local inference as peers in one project",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "agent-hub",
3
- "version": "0.12.10",
3
+ "version": "0.12.12",
4
4
  "description": "Channel between Claude Code and the agent-hub daemon: peer messages from Codex, Kimi and the local worker arrive as channel events; hub_send replies.",
5
5
  "author": {
6
6
  "name": "Young Joon Lee",
@@ -15403,7 +15403,7 @@ class ControlClient {
15403
15403
  // package.json
15404
15404
  var package_default = {
15405
15405
  name: "@staix/agent-hub",
15406
- version: "0.12.10",
15406
+ version: "0.12.12",
15407
15407
  description: "Native multi-agent hub: Claude Code, Codex, Kimi Code, Pi and local inference as peers in one project",
15408
15408
  license: "MIT",
15409
15409
  type: "module",
@@ -17,6 +17,40 @@ export interface PermissionRequest {
17
17
  tool?: string;
18
18
  }
19
19
 
20
+ export interface ACPUsageDiagnostic {
21
+ source: "usage_update" | "prompt_result";
22
+ availability: "known" | "unsupported" | "invalid";
23
+ shape: "total" | "input-output" | "context-used" | "none";
24
+ contextUsed?: number;
25
+ contextCapacity?: number;
26
+ total?: number;
27
+ }
28
+
29
+ /** Counter-only projection: no source payload, text, tools or metadata escapes. */
30
+ export function normalizeACPUsage(value: unknown, source: ACPUsageDiagnostic["source"] = "usage_update"): ACPUsageDiagnostic {
31
+ const unknown: ACPUsageDiagnostic = { source, availability: "unsupported", shape: "none" };
32
+ if (!value || typeof value !== "object" || Array.isArray(value)) return unknown;
33
+ const raw = value as Record<string, unknown>;
34
+ const f = raw.usage && typeof raw.usage === "object" && !Array.isArray(raw.usage) ? raw.usage as Record<string, unknown> : raw;
35
+ const valid = (v: unknown): v is number => typeof v === "number" && Number.isFinite(v) && v >= 0;
36
+ for (const key of ["totalTokens", "total_tokens"]) {
37
+ if (f[key] === undefined) continue;
38
+ const shape = "total";
39
+ return valid(f[key]) ? { source, availability: "known", shape, total: f[key] } : { source, availability: "invalid", shape };
40
+ }
41
+ for (const [input, output] of [["inputTokens", "outputTokens"], ["input_tokens", "output_tokens"]]) {
42
+ if (f[input!] === undefined && f[output!] === undefined) continue;
43
+ const a = f[input!], b = f[output!];
44
+ return valid(a) && valid(b) && Number.isFinite(a + b) ? { source, availability: "known", shape: "input-output", total: a + b } : { source, availability: "invalid", shape: "input-output" };
45
+ }
46
+ if (f.used !== undefined) {
47
+ if (!valid(f.used) || !valid(f.size) || !Number.isSafeInteger(f.used) || !Number.isSafeInteger(f.size) || f.size <= 0) return { source, availability: "invalid", shape: "context-used" };
48
+ // Qwen 0.24.7 derives this from collectContextData: occupancy is not cumulative consumption.
49
+ return { source, availability: "unsupported", shape: "context-used", contextUsed: f.used, contextCapacity: f.size };
50
+ }
51
+ return unknown;
52
+ }
53
+
20
54
  export interface AcpOptions {
21
55
  /** e.g. ["kimi", "acp"]. `opencode acp` fits the same adapter. */
22
56
  cmd: string[];
@@ -32,8 +66,12 @@ export interface AcpOptions {
32
66
  mcpServers?: { name: string; command: string; args: string[]; env: { name: string; value: string }[] }[];
33
67
  /** Appended to the standing instruction of the first delivery (role contract). */
34
68
  preamble?: string;
35
- /** The session's running token total from `usage_update` (inferred shape: totalTokens, else input + output, else `used`). Cumulative, not a delta. */
69
+ /** The session's running token total from `usage_update` (checked totalTokens or input/output pair; `used` context occupancy is diagnostic only). Cumulative, not a delta. */
36
70
  onTokens?: (sessionTotal: number, sessionId: string) => void;
71
+ /** The prompt rejected or ended without normal completion, even when it streamed partial text.
72
+ * Same shape as Pi's onTurnFailure; a stale cancelled turn never reports. */
73
+ onTurnFailure?: (envs: Envelope[], reason: string) => Promise<void> | void;
74
+ onUsageDiagnostic?: (observation: ACPUsageDiagnostic) => void;
37
75
  /** Resolve with an optionId, or undefined to cancel. Absent = every request is cancelled.
38
76
  * A request whose payload could not be resolved is titled as such and carries no session-wide allow option. */
39
77
  onPermission?: (req: PermissionRequest) => Promise<string | undefined>;
@@ -146,9 +184,12 @@ export class AcpPeer extends BasePeer {
146
184
  prompt
147
185
  .then((result) => {
148
186
  if (turn !== this.turn) return; // superseded: these chunks belong to a later turn
187
+ this.observeUsage(result, "prompt_result");
149
188
  this.primed = true;
150
189
  const body = this.chunks.join("").trim();
151
190
  if (body) this.onMessage?.(body, { inReplyTo: replyParent(envs), to: replyAudience(envs) });
191
+ // Partial output cannot certify normal completion (#160).
192
+ if (result?.stopReason !== "end_turn") this.reportTurnFailure(envs, `ACP prompt ended without normal completion (${String(result?.stopReason ?? "unknown")})`);
152
193
  if (deliveryId && this.activeDeliveryId === deliveryId) {
153
194
  this.acceptDelivery();
154
195
  this.delivery({ id: deliveryId, state: result?.stopReason === "end_turn" ? "completed" : "needs_review", ...(result?.stopReason === "end_turn" ? {} : { reason: "ACP prompt ended without normal completion" }) });
@@ -156,6 +197,7 @@ export class AcpPeer extends BasePeer {
156
197
  })
157
198
  .catch((e: Error) => {
158
199
  this.opts.log?.(`[${this.id}] prompt failed: ${e.message}`);
200
+ if (turn === this.turn) this.reportTurnFailure(envs, e.message); // a watchdog-cancelled turn reports stale, never as this turn's failure
159
201
  if (deliveryId && this.activeDeliveryId === deliveryId) this.delivery({ id: deliveryId, state: "needs_review", reason: e.message });
160
202
  else if (!deliveryId) this.onFailed?.(envs);
161
203
  })
@@ -165,6 +207,15 @@ export class AcpPeer extends BasePeer {
165
207
  });
166
208
  }
167
209
 
210
+ private reportTurnFailure(envs: Envelope[], reason: string): void {
211
+ try {
212
+ const result = this.opts.onTurnFailure?.(envs, reason);
213
+ if (result) void result.catch(() => this.opts.log?.("ACP failure handoff could not be completed"));
214
+ } catch {
215
+ this.opts.log?.("ACP failure handoff could not be completed");
216
+ }
217
+ }
218
+
168
219
  protected override onWatchdog(): void {
169
220
  const durable = !!this.activeDeliveryId;
170
221
  if (this.activeDeliveryId) this.delivery({ id: this.activeDeliveryId, state: "needs_review", reason: "ACP turn watchdog timeout" });
@@ -177,6 +228,12 @@ export class AcpPeer extends BasePeer {
177
228
  else super.onWatchdog();
178
229
  }
179
230
 
231
+ private observeUsage(value: unknown, source: ACPUsageDiagnostic["source"]): void {
232
+ const observation = normalizeACPUsage(value, source);
233
+ this.opts.onUsageDiagnostic?.(observation);
234
+ if (observation.availability === "known") this.opts.onTokens?.(observation.total!, this.sessionId);
235
+ }
236
+
180
237
  private acceptDelivery(): void {
181
238
  if (!this.activeDeliveryId || this.deliveryAccepted) return;
182
239
  this.deliveryAccepted = true;
@@ -222,12 +279,7 @@ export class AcpPeer extends BasePeer {
222
279
  if (u.status === "completed" || u.status === "failed") (this.toolInputs.delete(u.toolCallId), this.toolText.delete(u.toolCallId), this.toolTitles.delete(u.toolCallId));
223
280
  for (const map of [this.toolInputs, this.toolText, this.toolTitles]) while (map.size > TOOL_INPUT_CAP) map.delete(map.keys().next().value as string);
224
281
  }
225
- else if (u?.sessionUpdate === "usage_update" && this.opts.onTokens) {
226
- const f = { ...u, ...(typeof u.usage === "object" ? u.usage : {}) } as Record<string, unknown>;
227
- const num = (k: string) => (typeof f[k] === "number" ? (f[k] as number) : 0);
228
- const total = num("totalTokens") || num("total_tokens") || num("inputTokens") + num("outputTokens") || num("input_tokens") + num("output_tokens") || num("used");
229
- if (total > 0) this.opts.onTokens(total, this.sessionId);
230
- }
282
+ else if (u?.sessionUpdate === "usage_update") this.observeUsage(u, "usage_update");
231
283
  } else if (msg.method === "session/request_permission") {
232
284
  void this.answerPermission(msg);
233
285
  } else if (msg.method && msg.id !== undefined) {
package/src/cli/main.ts CHANGED
@@ -571,6 +571,11 @@ const commands: Record<string, () => Promise<void> | void> = {
571
571
  models: async () => {
572
572
  const action = args[0] ?? "status";
573
573
  const configured = projectConfig().mlx;
574
+ if (configured.enabled === false) {
575
+ if (action === "status") return console.log(JSON.stringify({ state: "disabled", enabled: false }, null, 2));
576
+ if (["setup", "start", "stop"].includes(action)) fail("MLX is disabled by mlx.enabled=false; models commands do not manage shared Ollama");
577
+ fail("usage: ahub models setup|status|start|stop");
578
+ }
574
579
  const runtimeDir = configured.runtimeDir ? resolve(cwd, configured.runtimeDir) : join(homedir(), ".agenthub", "runtimes", "mlx");
575
580
  const modelPath = configured.modelPath ? resolve(cwd, configured.modelPath) : join(homedir(), ".agenthub", "models", "qwen3-8b-mlx");
576
581
  const mlxOptions = configured.provider === "ollama" ? configured : { ...configured, runtimeDir, modelPath };
@@ -953,8 +958,11 @@ const commands: Record<string, () => Promise<void> | void> = {
953
958
  const sy = spawnSync(process.env.AGENTHUB_SWITCHYARD_BIN ?? "switchyard-server", ["--version"], { encoding: "utf8" });
954
959
  row(sy.status === 0 ? true : undefined, "switchyard", sy.status === 0 ? sy.stdout.trim() : "not installed: ahub local uses fixed_model on OmniRoute (cargo install --locked switchyard-server)");
955
960
  const mlxConfig = config.mlx;
956
- const mlx = await inspectMlx({ ...mlxConfig, runtimeDir: mlxConfig.runtimeDir ? resolve(cwd, mlxConfig.runtimeDir) : undefined, modelPath: mlxConfig.modelPath ? resolve(cwd, mlxConfig.modelPath) : undefined });
957
- row(mlx.state === "ready" || mlx.state === "stopped", "pi mlx", `${mlx.state}${mlx.model ? ` (${mlx.model})` : ""}${mlx.lastError ? `: ${mlx.lastError}` : ""}`);
961
+ if (mlxConfig.enabled === false) row(true, "pi mlx", "disabled (mlx.enabled=false; local endpoint not probed)");
962
+ else {
963
+ const mlx = await inspectMlx({ ...mlxConfig, runtimeDir: mlxConfig.runtimeDir ? resolve(cwd, mlxConfig.runtimeDir) : undefined, modelPath: mlxConfig.modelPath ? resolve(cwd, mlxConfig.modelPath) : undefined });
964
+ row(mlx.state === "ready" || mlx.state === "stopped", "pi mlx", `${mlx.state}${mlx.model ? ` (${mlx.model})` : ""}${mlx.lastError ? `: ${mlx.lastError}` : ""}`);
965
+ }
958
966
 
959
967
  const memory = new MemoryClient();
960
968
  const mem = await memory.health();
package/src/hub/daemon.ts CHANGED
@@ -69,7 +69,7 @@ export interface HubConfig {
69
69
  inference: InferenceConfig;
70
70
  omniroute: OmniRouteConfig;
71
71
  pi: { enabled: boolean; auto_start: boolean; cmd: string[]; backend: "auto" | "dgx" | "mlx"; dgx_coding: string; dgx_fast: string; max_steps: number };
72
- mlx: Pick<MlxOptions, "provider" | "host" | "runtimeDir" | "modelPath" | "port" | "model" | "sourceModel" | "contextWindow" | "maxInputTokens" | "maxTokens" | "maxConcurrency">;
72
+ mlx: { enabled: boolean } & Pick<MlxOptions, "provider" | "host" | "runtimeDir" | "modelPath" | "port" | "model" | "sourceModel" | "contextWindow" | "maxInputTokens" | "maxTokens" | "maxConcurrency">;
73
73
  /** `bash_network`: true is network through the egress proxy to `network_allow` (#65); "direct" is everything, until 0.13.0 (#83). */
74
74
  local: { deny: string[]; bash_network: boolean | "direct"; network_allow: string[]; max_steps: number; read_allow: string[] };
75
75
  /** Pending permission requests: how long they wait, and whether the desktop is told (issue #5). */
@@ -113,7 +113,7 @@ export const DEFAULT_CONFIG: HubConfig = {
113
113
  inference: DEFAULT_INFERENCE,
114
114
  omniroute: DEFAULT_OMNIROUTE,
115
115
  pi: { enabled: false, auto_start: false, cmd: ["pi"], backend: "auto", dgx_coding: "coding", dgx_fast: "fast", max_steps: 30 },
116
- mlx: { provider: "ollama", model: "agenthub-fast-mlx:4b-8k", sourceModel: "qwen3.5:4b-mlx", contextWindow: 8192, maxInputTokens: 6000, maxTokens: 2048, maxConcurrency: 1 },
116
+ mlx: { enabled: true, provider: "ollama", model: "agenthub-fast-mlx:4b-8k", sourceModel: "qwen3.5:4b-mlx", contextWindow: 8192, maxInputTokens: 6000, maxTokens: 2048, maxConcurrency: 1 },
117
117
  local: { deny: [], bash_network: false, network_allow: DEFAULT_NETWORK_ALLOW, max_steps: 30, read_allow: [] },
118
118
  // Off here, so tests and a hub without a config file stay silent; a project's config defaults it on for macOS.
119
119
  approvals: { timeout_s: 120, notify: false },
@@ -166,11 +166,13 @@ export function loadConfig(cwd: string): HubConfig {
166
166
  }
167
167
  if (file.local?.bash_network === "direct") retired.push('local.bash_network "direct" (the open network) goes in 0.13.0: set it to true and list the hosts in local.network_allow');
168
168
  if (file.mlx != null && (typeof file.mlx !== "object" || Array.isArray(file.mlx))) throw new Error("mlx configuration must be an object");
169
+ if (file.mlx?.enabled !== undefined && typeof file.mlx.enabled !== "boolean") throw new Error("mlx.enabled must be a boolean");
169
170
  if (file.mlx?.provider !== undefined && !["ollama", "legacy"].includes(file.mlx.provider)) throw new Error("mlx.provider must be ollama or legacy");
170
171
  if (file.mlx?.provider === undefined && (file.mlx?.modelPath || file.mlx?.runtimeDir || file.mlx?.bin || file.mlx?.port)) {
171
172
  throw new Error("legacy MLX configuration requires explicit mlx.provider=legacy; migrate to provider=ollama to avoid Python serving");
172
173
  }
173
- const mlx = { ...(file.mlx?.provider === "legacy" ? { provider: "legacy" as const, maxInputTokens: 16_000, maxTokens: 2048 } : DEFAULT_CONFIG.mlx), ...file.mlx };
174
+ const mlx = { ...(file.mlx?.provider === "legacy" ? { enabled: true, provider: "legacy" as const, maxInputTokens: 16_000, maxTokens: 2048 } : DEFAULT_CONFIG.mlx), ...file.mlx };
175
+ if (!mlx.enabled && file.pi?.backend === "mlx") throw new Error("Pi backend mlx conflicts with mlx.enabled=false");
174
176
  if (typeof mlx.runtimeDir === "string" && mlx.runtimeDir) mlx.runtimeDir = resolve(cwd, mlx.runtimeDir);
175
177
  if (typeof mlx.modelPath === "string" && mlx.modelPath) mlx.modelPath = resolve(cwd, mlx.modelPath);
176
178
  return {
@@ -200,6 +202,20 @@ export function loadConfig(cwd: string): HubConfig {
200
202
  };
201
203
  }
202
204
 
205
+ /** Explicit launch choices are never replaced when the local capability is disabled. */
206
+ export function mlxLaunchProblem(config: HubConfig, args: { backend?: unknown; model?: unknown }): string | undefined {
207
+ if (config.mlx.enabled !== false) return undefined;
208
+ if (args.backend === "mlx" || args.model === "mlx/fast") return "Pi MLX launch conflicts with mlx.enabled=false; select auto or a DGX alias";
209
+ return undefined;
210
+ }
211
+
212
+ /** Shipped defaults remain capability-aware; operator-written MLX pins require an explicit migration. */
213
+ function disabledMlxPolicyProblem(config: HubConfig, cwd: string): string | undefined {
214
+ if (config.mlx.enabled !== false || !existsSync(join(cwd, ".agenthub", "routing.toml"))) return undefined;
215
+ const conflict = Object.entries(currentRouting(cwd).classes).find(([, policy]) => policy?.pi_backend === "mlx");
216
+ return conflict ? `routing.toml: [classes.${conflict[0]}] pi_backend=mlx conflicts with mlx.enabled=false; remove the pin for hub/auto or select dgx` : undefined;
217
+ }
218
+
203
219
  export interface DaemonOptions {
204
220
  projectId?: string;
205
221
  instanceId?: string;
@@ -284,6 +300,8 @@ export async function startDaemon(opts: DaemonOptions) {
284
300
  let ready = false;
285
301
  try {
286
302
  const config = opts.config ?? loadConfig(opts.cwd);
303
+ const mlxConfigProblem = mlxLaunchProblem(config, config.pi) ?? disabledMlxPolicyProblem(config, opts.cwd);
304
+ if (mlxConfigProblem) throw new Error(mlxConfigProblem);
287
305
  mkdirSync(opts.stateDir, { recursive: true });
288
306
  const logFile = join(opts.stateDir, "hub.log");
289
307
  // The state dir can vanish under a running hub (issue #56): a log line must never take a handler down with it.
@@ -333,6 +351,13 @@ export async function startDaemon(opts: DaemonOptions) {
333
351
 
334
352
  // A session record left by a run that never shut down means it crashed (issue #37). A controlled restart has its own.
335
353
  const crashed = !recoveryOperation && !restartFilePresent ? readSessions(opts.stateDir) : undefined;
354
+ if (config.mlx.enabled === false) {
355
+ const launches = [...(restored?.peers ?? []).filter(peer => peer.id === "pi").map(peer => peer.launch), ...(crashed?.peers ?? []).filter(peer => peer.peer === "pi").map(peer => peer.meta.launch)];
356
+ for (const launch of launches) {
357
+ const problem = mlxLaunchProblem(config, (launch ?? {}) as { backend?: unknown; model?: unknown });
358
+ if (problem) throw new Error(`recorded Pi recovery: ${problem}`);
359
+ }
360
+ }
336
361
  // A controlled restart's source may have been cut short before it removed its record: this run is not a crash, and
337
362
  // a record left now would make the next ordinary start look like one.
338
363
  if (recoveryOperation || restartFilePresent) try { removeSessions(opts.stateDir); } catch { /* nothing to remove */ }
@@ -1516,6 +1541,8 @@ export async function startDaemon(opts: DaemonOptions) {
1516
1541
 
1517
1542
  async function startPeerBody(peer: string, args: { model?: string; route?: string; mode?: "headless" | "tui"; backend?: "auto" | "dgx" | "mlx"; sessionId?: string; sessionFile?: string; fresh?: boolean }, mute: (p: PeerAdapter) => void): Promise<Record<string, unknown>> {
1518
1543
  if (peer === "pi") {
1544
+ const problem = mlxLaunchProblem(config, { backend: args.backend ?? config.pi.backend, model: args.model }) ?? disabledMlxPolicyProblem(config, opts.cwd);
1545
+ if (problem) return { ok: false, error: problem };
1519
1546
  if (args.mode !== undefined && !["headless", "tui"].includes(args.mode)) return { ok: false, error: "invalid Pi mode" };
1520
1547
  if (args.backend !== undefined && !["auto", "dgx", "mlx"].includes(args.backend)) return { ok: false, error: "invalid Pi backend" };
1521
1548
  if (args.model !== undefined && !["dgx/coding", "dgx/fast", "mlx/fast"].includes(args.model)) return { ok: false, error: "unknown Pi model alias" };
@@ -1650,7 +1677,9 @@ export async function startDaemon(opts: DaemonOptions) {
1650
1677
  const mode = args.mode ?? "headless";
1651
1678
  const backend = args.backend ?? config.pi.backend;
1652
1679
  if (!["headless", "tui"].includes(mode) || !["auto", "dgx", "mlx"].includes(backend)) return { ok: false, error: "invalid Pi mode/backend" };
1653
- modelRelay ??= await startModelRelay({ omni, admitRequest: admitPiRequest, dgxMaxInputTokens: currentRouting(opts.cwd, log).pi.dgx_max_context_tokens, allowedDGXmodels: { "dgx/coding": config.pi.dgx_coding, "dgx/fast": config.pi.dgx_fast }, mlx: config.mlx, mlxAlias: "mlx/fast", fallbackDGXAlias: "dgx/fast", enableHubAuto: true,
1680
+ const problem = mlxLaunchProblem(config, { backend, model: args.model });
1681
+ if (problem) return { ok: false, error: problem };
1682
+ modelRelay ??= await startModelRelay({ omni, admitRequest: admitPiRequest, dgxMaxInputTokens: currentRouting(opts.cwd, log).pi.dgx_max_context_tokens, allowedDGXmodels: { "dgx/coding": config.pi.dgx_coding, "dgx/fast": config.pi.dgx_fast }, mlx: config.mlx.enabled === false ? undefined : config.mlx, mlxAlias: "mlx/fast", fallbackDGXAlias: "dgx/fast", enableHubAuto: true,
1654
1683
  routeSessionKey: () => { const session = bus.peers.get("pi")?.recoveryMetadata?.().sessionId; return typeof session === "string" ? session : undefined; },
1655
1684
  onRoute: record => event({ type: "route", peer: "pi", ...record }),
1656
1685
  });
@@ -1673,7 +1702,7 @@ export async function startDaemon(opts: DaemonOptions) {
1673
1702
  model: args.model,
1674
1703
  sessionId: args.sessionId, sessionFile: args.sessionFile,
1675
1704
  admitBudget: async (envs, unit) => unit === "model_calls" ? [] : tasks.admitExecutionEnvelopes(envs, "pi", unit),
1676
- relay: { url: modelRelay.url, token: modelRelay.token, models: modelRelay.models.map((id) => ({ id, contextWindow: id.startsWith("mlx/") ? Math.min(routing.pi.mlx_max_context_tokens, config.mlx.provider === "ollama" ? (config.mlx.contextWindow ?? 8192) : routing.pi.mlx_max_context_tokens) : routing.pi.dgx_max_context_tokens, maxTokens: id === "hub/auto" ? Math.min(config.mlx.maxTokens ?? 2048, 8192) : id.startsWith("mlx/") ? (config.mlx.maxTokens ?? 2048) : 8192 })) },
1705
+ relay: { url: modelRelay.url, token: modelRelay.token, models: modelRelay.models.map((id) => ({ id, contextWindow: id.startsWith("mlx/") ? Math.min(routing.pi.mlx_max_context_tokens, config.mlx.provider === "ollama" ? (config.mlx.contextWindow ?? 8192) : routing.pi.mlx_max_context_tokens) : routing.pi.dgx_max_context_tokens, maxTokens: id === "hub/auto" ? (config.mlx.enabled === false ? 8192 : Math.min(config.mlx.maxTokens ?? 2048, 8192)) : id.startsWith("mlx/") ? (config.mlx.maxTokens ?? 2048) : 8192 })) },
1677
1706
  tools: [...TOOL_SCHEMAS.map((t) => t.function), ...TASK_TOOLS.map((t) => ({ name: t.name, description: t.description, parameters: t.inputSchema }))],
1678
1707
  executeTool: async (name, raw, callId, sessionId, signal) => {
1679
1708
  if (signal?.aborted) return "error: turn cancelled before tool effects";
@@ -1709,7 +1738,12 @@ export async function startDaemon(opts: DaemonOptions) {
1709
1738
  const taskId = envs.find(env => env.refs?.task)?.refs?.task;
1710
1739
  const task = taskId ? board.get(Number(taskId)) : undefined;
1711
1740
  const policyBackend = task ? currentRouting(opts.cwd, log).classes[task.class]?.pi_backend : undefined;
1712
- if (policyBackend === "mlx") return "mlx/fast";
1741
+ if (policyBackend === "mlx") {
1742
+ if (config.mlx.enabled !== false) return "mlx/fast";
1743
+ const problem = disabledMlxPolicyProblem(config, opts.cwd);
1744
+ if (problem) throw new Error(problem);
1745
+ return "hub/auto";
1746
+ }
1713
1747
  if (policyBackend === "dgx") return task && ["bulk_edit", "test"].includes(task.class) ? "dgx/fast" : "dgx/coding";
1714
1748
  return "hub/auto";
1715
1749
  },
@@ -39,8 +39,20 @@ export interface ModelRelayStatus {
39
39
  * `identitySource` says where the served-model label came from: the gateway response header, a
40
40
  * generation SSE event (#137 classification: heartbeats never identify), or the locally validated
41
41
  * MLX configuration. HTTP 200, the requested alias and a previous request's label never identify. */
42
+ export interface RelayUsageObservation {
43
+ source: "openai-stream-usage";
44
+ completeness: "complete" | "partial";
45
+ promptTokens?: number;
46
+ completionTokens?: number;
47
+ totalTokens?: number;
48
+ }
49
+
42
50
  export interface RelayRequestRecord {
43
51
  id: string;
52
+ dispatchGroupId?: string;
53
+ fallbackOfId?: string;
54
+ failureClass?: "http" | "transport" | "startup" | "admission" | "cancelled";
55
+ httpStatus?: number;
44
56
  /** Admission timestamp (start of the upstream dispatch attempt), ISO. */
45
57
  at: string;
46
58
  /** Resolved backend alias (the requested route). */
@@ -49,6 +61,11 @@ export interface RelayRequestRecord {
49
61
  requestedModel?: string;
50
62
  /** Sanitized `x-omniroute-provider` header; absent stays unknown. */
51
63
  provider?: string;
64
+ providerSource?: "header" | "none";
65
+ providerAvailability?: "known" | "missing";
66
+ /** Independent transport counter, never a native session counter. */
67
+ requestUsage?: RelayUsageObservation;
68
+ usageAvailability?: "known" | "partial" | "missing" | "invalid";
52
69
  /** Observed served model; never read back from the backend's mutable last label. */
53
70
  actualModel?: string;
54
71
  identitySource: "header" | "stream" | "configured" | "none";
@@ -86,6 +103,9 @@ export interface ModelRelayOptions {
86
103
  mlxAlias?: string;
87
104
  mlxModel?: string;
88
105
  fallbackDGXAlias?: string;
106
+ /** Include explicit missing usage/provider metadata and dispatch groups, independently of observers.
107
+ * Omission preserves the legacy absent-metadata journal schema. */
108
+ observeRequestMetadata?: boolean;
89
109
  /** Called exactly once per journaled request, at its terminal close, with a sanitized copy. */
90
110
  onRequest?: (record: RelayRequestRecord) => void;
91
111
  }
@@ -103,6 +123,7 @@ export interface ModelRelay {
103
123
  interface RequestJournalEntry {
104
124
  readonly record: RelayRequestRecord;
105
125
  identify(model: string, source: "header" | "stream" | "configured"): void;
126
+ usage(value: unknown): void;
106
127
  close(outcome: RelayRequestRecord["outcome"]): void;
107
128
  }
108
129
 
@@ -116,7 +137,24 @@ interface ActiveRequest {
116
137
 
117
138
  class ExecutionAdmissionError extends Error {}
118
139
 
119
- const safeHeader = (value: string | null): string | undefined => value && value.length < 256 ? value : undefined;
140
+ const safeHeader = (value: string | null): string | undefined => value && value.length < 256 && !/\p{C}/u.test(value) ? value : undefined;
141
+
142
+ export function normalizeRelayUsage(value: unknown): RelayUsageObservation | undefined {
143
+ if (!value || typeof value !== "object" || Array.isArray(value)) return undefined;
144
+ const usage = value as Record<string, unknown>;
145
+ const fields = [["prompt_tokens", "promptTokens"], ["completion_tokens", "completionTokens"], ["total_tokens", "totalTokens"]] as const;
146
+ const result: RelayUsageObservation = { source: "openai-stream-usage", completeness: "partial" };
147
+ let count = 0;
148
+ for (const [input, output] of fields) {
149
+ if (usage[input] === undefined) continue;
150
+ const v = usage[input];
151
+ if (typeof v !== "number" || !Number.isFinite(v) || v < 0) return undefined;
152
+ result[output] = v; count++;
153
+ }
154
+ if (!count) return undefined;
155
+ if (count === 3) result.completeness = "complete";
156
+ return result;
157
+ }
120
158
 
121
159
  function assertLoopback(host: string): void {
122
160
  const value = host.toLowerCase();
@@ -159,7 +197,7 @@ function isGenerationEvent(choices: unknown): boolean {
159
197
  });
160
198
  }
161
199
 
162
- function sseResponse(response: Response, release: () => void, onModel?: (model: string) => void, registerCancel?: (cancel: (reason?: unknown) => Promise<void>) => void, onClose?: (outcome: RelayRequestRecord["outcome"]) => void): Response {
200
+ function sseResponse(response: Response, release: () => void, onModel?: (model: string) => void, registerCancel?: (cancel: (reason?: unknown) => Promise<void>) => void, onClose?: (outcome: RelayRequestRecord["outcome"]) => void, onUsage?: (value: unknown) => void): Response {
163
201
  if (!response.body) {
164
202
  release();
165
203
  onClose?.("failed");
@@ -172,22 +210,23 @@ function sseResponse(response: Response, release: () => void, onModel?: (model:
172
210
  registerCancel?.(cancel);
173
211
  let inspectBuffer = "";
174
212
  let inspectedModel = false;
213
+ const decoder = new TextDecoder();
175
214
  const inspect = (chunk: Uint8Array) => {
176
- if (!onModel || inspectedModel) return;
177
- inspectBuffer += new TextDecoder().decode(chunk);
215
+ if (!onUsage && (!onModel || inspectedModel)) return;
216
+ inspectBuffer += decoder.decode(chunk, { stream: true });
178
217
  if (inspectBuffer.length > 64_000) inspectBuffer = inspectBuffer.slice(-64_000);
179
218
  const lines = inspectBuffer.split(/\r?\n/);
180
219
  inspectBuffer = lines.pop() ?? "";
181
220
  for (const line of lines) {
182
221
  if (!line.startsWith("data:") || line.slice(5).trim() === "[DONE]") continue;
183
222
  try {
184
- const value = JSON.parse(line.slice(5).trim()) as { model?: unknown; choices?: unknown[] };
223
+ const value = JSON.parse(line.slice(5).trim()) as { model?: unknown; choices?: unknown[]; usage?: unknown };
224
+ if (value.usage !== undefined && value.usage !== null) onUsage?.(value.usage);
185
225
  // Transport heartbeats are not model identity: a gateway keepalive can name a synthetic model on
186
226
  // an event with no generation activity (no choices, or only empty deltas without a finish reason).
187
- if (typeof value.model === "string" && value.model.length < 256 && isGenerationEvent(value.choices)) {
227
+ if (!inspectedModel && typeof value.model === "string" && safeHeader(value.model) !== undefined && isGenerationEvent(value.choices)) {
188
228
  inspectedModel = true;
189
- onModel(value.model);
190
- return;
229
+ onModel?.(value.model);
191
230
  }
192
231
  } catch { /* incomplete or non-JSON SSE data */ }
193
232
  }
@@ -219,6 +258,8 @@ function sseResponse(response: Response, release: () => void, onModel?: (model:
219
258
  });
220
259
  }
221
260
 
261
+ const copyRecord = (record: RelayRequestRecord): RelayRequestRecord => ({ ...record, ...(record.requestUsage ? { requestUsage: { ...record.requestUsage } } : {}) });
262
+
222
263
  export async function startModelRelay(options: ModelRelayOptions): Promise<ModelRelay> {
223
264
  const host = options.host ?? "127.0.0.1";
224
265
  assertLoopback(host);
@@ -238,12 +279,13 @@ export async function startModelRelay(options: ModelRelayOptions): Promise<Model
238
279
 
239
280
  // The record object is the generation fence: every update goes through this entry's own closure,
240
281
  // so interleaved requests for the same alias never write into each other's evidence.
241
- const openRequestRecord = (alias: string): RequestJournalEntry => {
282
+ const openRequestRecord = (alias: string, dispatchGroupId: string): RequestJournalEntry => {
242
283
  const start = Date.now();
243
284
  const expectedServedModel = options.expectedServedModels?.[alias];
244
285
  const record: RelayRequestRecord = {
245
- id: randomUUID(), at: new Date(start).toISOString(), alias,
286
+ id: randomUUID(), ...(options.observeRequestMetadata ? { dispatchGroupId } : {}), at: new Date(start).toISOString(), alias,
246
287
  identitySource: "none", role: "unknown", outcome: "completed", identified: false, durationMs: 0,
288
+ ...(options.observeRequestMetadata ? { providerSource: "none" as const, providerAvailability: "missing" as const, usageAvailability: "missing" as const } : {}),
247
289
  };
248
290
  let closed = false;
249
291
  const identify: RequestJournalEntry["identify"] = (model, source) => {
@@ -262,15 +304,25 @@ export async function startModelRelay(options: ModelRelayOptions): Promise<Model
262
304
  record.durationMs = Date.now() - start;
263
305
  const expectedModel = expectedServedModel ?? record.requestedModel;
264
306
  if (record.identified && expectedModel !== undefined && record.actualModel !== expectedModel) record.mismatch = true;
265
- journal.push({ ...record });
307
+ journal.push(copyRecord(record));
266
308
  if (journal.length > JOURNAL_LIMIT) journal.shift();
267
309
  try {
268
310
  // An async hook fits the void signature: its rejection is handled too, never unobserved.
269
- const notified = options.onRequest?.({ ...record }) as unknown;
311
+ const notified = options.onRequest?.(copyRecord(record)) as unknown;
270
312
  if (notified instanceof Promise) notified.catch(() => { /* a persistence hook must never break the relay */ });
271
313
  } catch { /* a persistence hook must never break the proxied stream it observes */ }
272
314
  };
273
- return { record, identify, close };
315
+ const usage = (value: unknown) => {
316
+ if (closed) return;
317
+ record.providerSource ??= "none";
318
+ record.providerAvailability ??= "missing";
319
+ const observation = normalizeRelayUsage(value);
320
+ if (observation) {
321
+ record.requestUsage = observation;
322
+ record.usageAvailability = observation.completeness === "complete" ? "known" : "partial";
323
+ } else if (!record.requestUsage) record.usageAvailability = "invalid";
324
+ };
325
+ return { record, identify, usage, close };
274
326
  };
275
327
 
276
328
  const ensureMlxHandle = async (): Promise<MlxHandle> => (mlx ??= await (mlxStarting ??= ensureMlx(options.mlx).finally(() => { mlxStarting = undefined; })));
@@ -345,18 +397,25 @@ export async function startModelRelay(options: ModelRelayOptions): Promise<Model
345
397
  response = await fetch(`${base.replace(/\/$/, "")}/chat/completions`, { method: "POST", ...(isOllama ? { redirect: "error" as const } : {}), headers, body: JSON.stringify(bodyForUpstream(boundedRequest, model)), signal: AbortSignal.any([signal, AbortSignal.timeout(deadline)]) });
346
398
  } catch (error) {
347
399
  releaseOnce();
400
+ journalEntry.record.failureClass = "transport";
348
401
  setState(backend, { state: "error", lastError: error instanceof Error ? error.message.slice(0, 160) : "upstream request failed" });
349
402
  throw error;
350
403
  }
404
+ const provider = safeHeader(response.headers.get("x-omniroute-provider"));
405
+ if (provider) {
406
+ journalEntry.record.provider = provider;
407
+ journalEntry.record.providerSource = "header";
408
+ journalEntry.record.providerAvailability = "known";
409
+ }
351
410
  if (!response.ok) {
352
411
  releaseOnce();
412
+ journalEntry.record.failureClass = "http";
413
+ journalEntry.record.httpStatus = response.status;
353
414
  const message = `backend returned HTTP ${response.status}`;
354
415
  setState(backend, { state: "error", lastError: message });
355
416
  throw new Error(message);
356
417
  }
357
- const provider = safeHeader(response.headers.get("x-omniroute-provider"));
358
418
  const actualModel = safeHeader(response.headers.get("x-model-router-selected-model"));
359
- if (provider) journalEntry.record.provider = provider;
360
419
  if (actualModel) journalEntry.identify(actualModel, "header");
361
420
  else if (backend.kind === "mlx") journalEntry.identify(model, "configured");
362
421
  setState(backend, { state: "ready", requestedModel: request.model, active: activeByAlias.get(alias) ?? 0,
@@ -419,10 +478,23 @@ export async function startModelRelay(options: ModelRelayOptions): Promise<Model
419
478
  return Response.json({ error: "input exceeds the model context budget" }, { status: 400 });
420
479
  }
421
480
  const fallback = backend.kind === "mlx" && options.fallbackDGXAlias ? { kind: "dgx", alias: options.fallbackDGXAlias } as ModelBackend : undefined;
481
+ const dispatchGroupId = randomUUID();
482
+ let primaryDispatchId: string | undefined;
422
483
  const dispatch = async (selected: ModelBackend, body: RelayRequest) => {
423
- const journalEntry = openRequestRecord(aliasOf(selected, mlxAlias));
484
+ const journalEntry = openRequestRecord(aliasOf(selected, mlxAlias), dispatchGroupId);
485
+ if (primaryDispatchId) {
486
+ journalEntry.record.fallbackOfId = primaryDispatchId;
487
+ journalEntry.record.dispatchGroupId = dispatchGroupId;
488
+ }
489
+ else primaryDispatchId = journalEntry.record.id;
424
490
  record.closeRecord = journalEntry.close;
425
- const result = await upstream(body, selected, controller.signal, journalEntry);
491
+ let result: Awaited<ReturnType<typeof upstream>>;
492
+ try {
493
+ result = await upstream(body, selected, controller.signal, journalEntry);
494
+ } catch (error) {
495
+ journalEntry.record.failureClass ??= controller.signal.aborted ? "cancelled" : error instanceof ExecutionAdmissionError ? "admission" : "startup";
496
+ throw error;
497
+ }
426
498
  let released = false;
427
499
  const release = () => {
428
500
  if (released) return;
@@ -432,11 +504,12 @@ export async function startModelRelay(options: ModelRelayOptions): Promise<Model
432
504
  activeRequests.delete(record);
433
505
  };
434
506
  record.release = release;
435
- return { result, release, journalEntry };
507
+ const onUsage = journalEntry.usage;
508
+ return { result, release, journalEntry, onUsage };
436
509
  };
437
510
  try {
438
- const { result, release, journalEntry } = await dispatch(backend, body);
439
- return sseResponse(result.response, release, result.onModel, (cancel) => { record.cancel = cancel; }, journalEntry.close);
511
+ const { result, release, journalEntry, onUsage } = await dispatch(backend, body);
512
+ return sseResponse(result.response, release, result.onModel, (cancel) => { record.cancel = cancel; }, journalEntry.close, onUsage);
440
513
  } catch (error) {
441
514
  record.closeRecord?.(controller.signal.aborted ? "cancelled" : "failed");
442
515
  if (!fallback || controller.signal.aborted || error instanceof ExecutionAdmissionError) {
@@ -445,8 +518,8 @@ export async function startModelRelay(options: ModelRelayOptions): Promise<Model
445
518
  return Response.json({ error: error instanceof Error ? error.message : "backend unavailable" }, { status: 502 });
446
519
  }
447
520
  try {
448
- const { result, release, journalEntry } = await dispatch(fallback, { ...body, model: fallback.alias });
449
- return sseResponse(result.response, release, result.onModel, (cancel) => { record.cancel = cancel; }, journalEntry.close);
521
+ const { result, release, journalEntry, onUsage } = await dispatch(fallback, { ...body, model: fallback.alias });
522
+ return sseResponse(result.response, release, result.onModel, (cancel) => { record.cancel = cancel; }, journalEntry.close, onUsage);
450
523
  } catch (fallbackError) {
451
524
  record.closeRecord?.(controller.signal.aborted ? "cancelled" : "failed");
452
525
  record.cleanup();
@@ -458,7 +531,7 @@ export async function startModelRelay(options: ModelRelayOptions): Promise<Model
458
531
  });
459
532
  const url = `http://${host}:${server.port}/v1`;
460
533
  const status = (): ModelRelayStatus => ({ url, models, backends: [...states.values()].map((value) => ({ ...value })) });
461
- const requests = (): RelayRequestRecord[] => journal.map((record) => ({ ...record }));
534
+ const requests = (): RelayRequestRecord[] => journal.map(copyRecord);
462
535
  return { url, token, models, status, requests, close: async () => {
463
536
  const closing = [...activeRequests].map(async (request) => {
464
537
  // The relay-initiated cancellation closes the record first: the abort below settles the stream