@staix/agent-hub 0.12.9 → 0.12.10

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,13 @@ Issue and pull request numbers in the entries for 0.7.7 and earlier refer to the
4
4
 
5
5
  ## Unreleased
6
6
 
7
+ ## 0.12.10
8
+
9
+ - Preserve each Pi/Qwen peer's observed sandbox-probe result in v3 readiness; failed or missing native probes never synthesize a denial (#150).
10
+ - Dispatch native v3 runs through their own coordination ledger, preserving completion, timeout and setup-error classifications independently from official quality; derive model identity and linkage from the request journal and bound timing medians to valid completed attempts (#151).
11
+ - Count official native usage only for actors with started sessions, keeping Pi incremental counters and Qwen session totals separate (#152). Apply the same participation gate to v3 ledger aggregates, with known, unknown and absent measurement coverage (#156).
12
+ - Allow trusted served-model expectations in relay mismatch checks when a provider-prefixed upstream identifier differs from the physical generation model (#139, #140).
13
+
7
14
  ## 0.12.9
8
15
 
9
16
  - The CooperBench native runner proves each prepared fixture root is still the directory preparation left before anything is locked, written or launched, and each arm checks its own again first (#119): a root replaced after preparation by a symlink to an equivalent outside tree passed the lexical `resolve()` comparison and the baseline content checks and would have redirected setup and agent writes there. The check is read-only (`lstat` and the real path, never a follow), so a substitution is refused without touching its target; a sibling fixture root that is itself a symlink is refused before the sibling-artifact walk reads it.
@@ -91,9 +91,11 @@ Its calibration corrections, relative to what an ad-hoc harness gets wrong:
91
91
  6. Existing attempt evidence is rejected before any fallback record is written.
92
92
  7. The driver is preflighted by strict typechecking (`scripts/check.sh`) and executable lifecycle checks (every spawnable binary answers a trivial command before a fixture is touched), not transpilation alone.
93
93
 
94
+ The v3 driver supplies a trusted `expectedServedModels` map to the relay: `fixed_backend` is the provider/route identifier sent upstream, while `expected_served_model` is the physical name expected in generation. The expectation is captured when each request starts. These names may differ (for example `flashnext/qwen3.8-flash-next` and `qwen3.8-flash-next`); a route prefix alone is not a model mismatch. An observed model that differs from the physical expectation still fails qualification, including after cancellation. The Qwen read-denial probe binds the announced read, exact target path and denial result to the same tool call ID; failures from separate calls cannot be combined. Each peer's serialized `readiness[].sandboxProbe` reports only what that peer's own probe observed (#150): `checked: true, result: "denied"` solely with the structured denial evidence for that peer's exact call, otherwise an explicit `checked: false`/`result: "unknown"` (the probe never settled or never ran) or `result: "failed"` (it settled without the evidence, or reported the file accessible), with Qwen's seatbelt layer recorded separately in `kernelProbe`. Grading's readiness gate accepts only the verified denial, so a failed setup stays unavailable however the other fields read.
95
+
94
96
  Served-model evidence is per request, from the relay's journaled `RelayRequestRecord` (#139), not from a response wrapper: a completed request that was never identified fails the attempt's model gate (a heartbeat-only stream identifies nothing), an observed mismatch flags it, and a request cancelled before identification stays explicitly unidentified and never certifies another request. Qwen's peer tool is approved through the adapter's shipped tool-identity binding (#138): the announced `hub_send (pilot-peer-bus MCP Server)` title resolves to the canonical `mcp__pilot-peer-bus__hub_send`, the only name on the exact whitelist. The joint arm's MCP peer bus is the versioned `scripts/benchmarks/peer-bus-mcp.py`, pinned with the driver in `prepared.json` and `cohort.json` next to the candidate source pins (`src/adapters/pi.ts`, `src/adapters/acp.ts`, `src/models/relay.ts`).
95
97
 
96
- Grading flows through the same official evaluator adapter and controls; a quality failure is never a retry selector and no evaluator feedback reaches the candidate agents during generation. Reports keep quality, model-identity and request-linkage coverage separate per arm, with unavailable attempts (failed, missing or unavailable) retained in the planned denominator, and name the usage units (Pi's incremental `onTokens` counter, Qwen's session `usage_update` running total — never added together) and the tool-surface difference (Pi's hub-moderated tools against Qwen's own seatbelted auto-edit tools). A live cohort, the official Docker controls and a native readback of an attempt's records remain manual live legs requiring accounts and the pinned archives; they are not part of the checked-in tests.
98
+ Grading flows through the same official evaluator adapter and controls; a quality failure is never a retry selector and no evaluator feedback reaches the candidate agents during generation. Reports keep quality, model-identity and request-linkage coverage separate per arm, with unavailable attempts (failed, missing or unavailable) retained in the planned denominator, and name the usage units (Pi's incremental `onTokens` counter, Qwen's session `usage_update` running total — never added together) and the tool-surface difference (Pi's hub-moderated tools against Qwen's own seatbelted auto-edit tools). A native actor's usage is counted only in attempts where that actor participated (#152): every v3 grade row carries `native_participants`, the actors the attempt record shows actually started — the arm bounds the candidates (a solo-qwen attempt has no Pi peer), and among them a peer participated only once its readiness entry carries the sessionId its adapter reported after start. A peer that was constructed but never started (a setup failure: `elapsedMs` 0, no active start) is absent; its token counter's initial 0 is never read as a measurement. An absent actor reports zero known observations and a null total (`*_tokens_known: 0`, `*_tokens: null`), a participant whose native reading is missing stays unknown rather than zero, and a participant's genuinely observed 0 stays counted. The summary's `pi_participating`/`qwen_participating` fields give each coverage figure's participation denominator. Grade rows from before #152 carry no `native_participants`; the report takes their participants from the arm, which is exact for them: their `native_usage` was attached only to scored attempts, whose readiness gate had proved every required actor's session. A live cohort, the official Docker controls and a native readback of an attempt's records remain manual live legs requiring accounts and the pinned archives; they are not part of the checked-in tests.
97
99
 
98
100
  ## Coordination ledger
99
101
 
@@ -103,6 +105,8 @@ python3 scripts/benchmarks/ledger.py --run /private/tmp/ahub-0124-r1 [--run /pri
103
105
 
104
106
  It reads each run record, the Claude transcript it names and the fixture's git history, and writes `ledger.json` into the first run directory with a `units` table that names the unit and coverage of every measure. Give `--run` once per directory to pool the repeats of one plan; an attempt a directory's `cohort.json` planned that wrote no record is listed as missing, and one a still-locked `runs/` hides as unreadable, in the summary per arm too (an arm with no record at all is listed with what it owes). Per attempt: completion (the runner's end reason and its detail, such as `wall-timeout`, and whether every task has a done); setup and active time; first candidate, completion intents, integration and check times, and each agent's settlement read from its own record (the end of its last native turn after the last done, turns started by late messages included; the last done itself when it did not work after it), and how long the agents could still write after the active time ended (`stopped_s`); usage per agent in task (to its last done on the board) and over the whole attempt: Codex turns, assistant messages, token-usage updates whose running total grew past the total before the window, and the token growth; Claude assistant messages (unique message ids), turns and the tokens their usage records; Claude's main-loop requests by request id (side requests and retries are not in its transcript), and Codex's provider requests unknown (app-server 0.159 does not send its response ids); Codex turns after its done and what started them; late replies, steered or at the next turn; the hub's fact offers, acknowledgements, bytes offered and acknowledged, build times, hook start-up times and steer round trips by path, in the task window; capability readbacks; validity by the grader's gates, and for turn-free the treatment received; held-back messages (`quiet` events) apart from [FYI] messages (which include the final [FYI] the instructions ask for); stale notices; shadow split predictions with their traces; the hooks each agent's records show, by a label that keeps paths and arguments out, with Claude's hook durations and the hub's own timing of every facts hook call; and contributions. Summaries give medians over valid completed attempts and over the (case, repeat) pairs every arm completed validly, totals over the attempts whose tasks were handed out and to which the measure applies (a Codex measure in a Claude-only arm is not counted as unknown), with the number of attempts each was unknown for, and the reasons for the rest. A repeat given twice is refused, and `--plan pilot` (or `study`) lists every planned attempt that wrote no record, whole repeats included. `ledger.json` holds code fragments from the agents' writes and local paths: keep it with the private run data and never commit it; the summary is what a verification record quotes. In-task windows end at different points by arm (a turn-free integration step comes before the done, an advisory completed-change notice turn after it), so arms are compared on whole-attempt usage.
105
107
 
108
+ Manifest v3 run directories (headless Pi/Qwen, `protocol: native-pq-v3` in every record, #140) pool the same way: `--plan study` lists the unwritten cells of the fixed ten cases x three arms x two repeats matrix, and a repeated repeat is refused. A v3 record has no board `taskStates` (the driver coordinates over bus prompts), so the ledger dispatches on the record's protocol instead of requiring tasks: completion is the runner's own end classification (`completed`, `wall-timeout`, `infrastructure-error`, `interrupted`, with the flags beside them as `end_story`), independent of official quality, which stays the grader's. Per attempt the row carries setup and elapsed times, the request-linkage coverage and model-identity verification recomputed from the journaled `RelayRequestRecord`s by the grader's own per-request gate and coverage functions (#139 — never the record's cached `generationVerified`/`requestLinkage` summary: an identified request is evidence whatever its outcome, a cancelled-after-identification mismatch fails the verdict, and a record without its journal reports unknown), the teardown gates' validity, and the native usage exactly as recorded — Pi's incremental `onTokens` counter and Qwen's session `usage_update` running total, the whole attempt including the setup probes, two units that are never added together; an attempt that wrote no usage is unknown, never zero. Rows preserve those raw counters and expose `native_participants`, using the grader's started-session predicate. Aggregate usage counts only participating actors, with `<actor>_participating`, `<actor>_not_participating`, `<actor>_tokens_known` and `<actor>_tokens_unknown` coverage: an unstarted actor is absent, a participating actor without a reading is unknown, and a measured zero remains known (#156). The task-based measures of the Claude/Codex records do not exist for v3 and are absent from its rows. A record whose protocol is neither `native-pq-v3` nor `native-cc-v1` is refused before any output is written.
109
+
106
110
  Contributions are a heuristic for possible loss, never a certificate: per agent and file, the identifiers and changed fragments its applied writes introduced (only a Claude tool call with a successful result, or a completed Codex patch, counts; a Write replaces the agent's earlier contribution and is not credited with what other agents wrote; a delete removes it; when a file is moved, every agent's contributions to it are checked at its new path) that the final tree lacks, a fragment counting as present anywhere in the file. A same-name overwrite shows as a lost fragment. Shell commands run during the task are not attributed and are counted under `coverage`, with a missing transcript or an unreadable file. Correctness comes from the official grader, for every arm.
107
111
 
108
112
  ## Preregistration (v2)
@@ -647,21 +647,21 @@ Rows without a live process are stale registrations; forget them with
647
647
 
648
648
  Upgrade running projects with the target release's own coordinator. It accepts
649
649
  a running source on control protocol 9 (0.6.x), 10 (0.7.0 through 0.12.0),
650
- 11 (0.12.1 and 0.12.2), 12 (0.12.3) or 13 (0.12.4 through 0.12.9) and only
650
+ 11 (0.12.1 and 0.12.2), 12 (0.12.3) or 13 (0.12.4 through 0.12.10) and only
651
651
  a target on its own protocol, so the target's coordinator fits every supported
652
652
  source and carries every recovery fix released up to it. Protocol 8 and older
653
653
  (0.5.x and earlier) are refused as `manual-bootstrap-required`. Run from the
654
654
  project directory, without replacing the global CLI first:
655
655
 
656
656
  ```bash
657
- bunx --package @staix/agent-hub@0.12.9 ahub upgrade --to 0.12.9 --dry-run
658
- bunx --package @staix/agent-hub@0.12.9 ahub upgrade --to 0.12.9 --yes
657
+ bunx --package @staix/agent-hub@0.12.10 ahub upgrade --to 0.12.10 --dry-run
658
+ bunx --package @staix/agent-hub@0.12.10 ahub upgrade --to 0.12.10 --yes
659
659
  ```
660
660
 
661
661
  | Running now | Coordinator to use |
662
662
  | --- | --- |
663
663
  | 0.6.x (protocol 9) | the target's, through `bunx` as above |
664
- | 0.7.0 through 0.12.0 (protocol 10), 0.12.1 and 0.12.2 (protocol 11), 0.12.3 (protocol 12), 0.12.4 through 0.12.9 (protocol 13) | the target's, through `bunx` as above |
664
+ | 0.7.0 through 0.12.0 (protocol 10), 0.12.1 and 0.12.2 (protocol 11), 0.12.3 (protocol 12), 0.12.4 through 0.12.10 (protocol 13) | the target's, through `bunx` as above |
665
665
  | any supported source, with the installed CLI already at the target | `ahub upgrade` below, which is the same coordinator |
666
666
  | 0.5.x or earlier (protocol 8 and older) | not supported: bootstrap by hand with the matching CLI |
667
667
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@staix/agent-hub",
3
- "version": "0.12.9",
3
+ "version": "0.12.10",
4
4
  "description": "Native multi-agent hub: Claude Code, Codex, Kimi Code, Pi and local inference as peers in one project",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "agent-hub",
3
- "version": "0.12.9",
3
+ "version": "0.12.10",
4
4
  "description": "Channel between Claude Code and the agent-hub daemon: peer messages from Codex, Kimi and the local worker arrive as channel events; hub_send replies.",
5
5
  "author": {
6
6
  "name": "Young Joon Lee",
@@ -15403,7 +15403,7 @@ class ControlClient {
15403
15403
  // package.json
15404
15404
  var package_default = {
15405
15405
  name: "@staix/agent-hub",
15406
- version: "0.12.9",
15406
+ version: "0.12.10",
15407
15407
  description: "Native multi-agent hub: Claude Code, Codex, Kimi Code, Pi and local inference as peers in one project",
15408
15408
  license: "MIT",
15409
15409
  type: "module",
@@ -34,7 +34,8 @@ export interface ModelRelayStatus {
34
34
  /** Sanitized per-request identity and lifecycle evidence. One record per upstream dispatch attempt
35
35
  * (a fallback dispatch is its own record). Records carry no messages, tools, keys or Access headers.
36
36
  * `identified: false` with `outcome: "cancelled"` is the cancelled-before-identification state; an
37
- * observed `actualModel` that differs from the upstream-configured `requestedModel` sets `mismatch`.
37
+ * observed `actualModel` that differs from the trusted expected served model sets `mismatch`.
38
+ * Without an explicit expectation, the upstream-configured identifier remains the comparison default.
38
39
  * `identitySource` says where the served-model label came from: the gateway response header, a
39
40
  * generation SSE event (#137 classification: heartbeats never identify), or the locally validated
40
41
  * MLX configuration. HTTP 200, the requested alias and a previous request's label never identify. */
@@ -56,7 +57,7 @@ export interface RelayRequestRecord {
56
57
  role: "primary" | "auxiliary" | "unknown";
57
58
  outcome: "completed" | "cancelled" | "failed";
58
59
  identified: boolean;
59
- /** Set only when the observed served model differs from the upstream-configured model. */
60
+ /** Set only when the observed served model differs from the expected model. */
60
61
  mismatch?: boolean;
61
62
  durationMs: number;
62
63
  }
@@ -76,6 +77,10 @@ export interface ModelRelayOptions {
76
77
  routeSessionKey?: (request: RelayRequest) => string | undefined;
77
78
  onRoute?: (event: { route: "hub/auto"; tier: string; source: "override" | "dimensions" | "hold" | "classifier" | "default"; score: number; ms: number }) => void;
78
79
  allowedDGXmodels: Record<string, string>;
80
+ /** Trusted physical model expectations by backend alias. Gateway identifiers can include a provider
81
+ * prefix or route name that differs from the model reported by generation. Omission preserves the
82
+ * existing literal upstream-identifier comparison; never derive this map from a response. */
83
+ expectedServedModels?: Record<string, string>;
79
84
  dgxMaxInputTokens?: number;
80
85
  mlx?: MlxOptions;
81
86
  mlxAlias?: string;
@@ -235,6 +240,7 @@ export async function startModelRelay(options: ModelRelayOptions): Promise<Model
235
240
  // so interleaved requests for the same alias never write into each other's evidence.
236
241
  const openRequestRecord = (alias: string): RequestJournalEntry => {
237
242
  const start = Date.now();
243
+ const expectedServedModel = options.expectedServedModels?.[alias];
238
244
  const record: RelayRequestRecord = {
239
245
  id: randomUUID(), at: new Date(start).toISOString(), alias,
240
246
  identitySource: "none", role: "unknown", outcome: "completed", identified: false, durationMs: 0,
@@ -254,7 +260,8 @@ export async function startModelRelay(options: ModelRelayOptions): Promise<Model
254
260
  closed = true;
255
261
  record.outcome = outcome;
256
262
  record.durationMs = Date.now() - start;
257
- if (record.identified && record.requestedModel !== undefined && record.actualModel !== record.requestedModel) record.mismatch = true;
263
+ const expectedModel = expectedServedModel ?? record.requestedModel;
264
+ if (record.identified && expectedModel !== undefined && record.actualModel !== expectedModel) record.mismatch = true;
258
265
  journal.push({ ...record });
259
266
  if (journal.length > JOURNAL_LIMIT) journal.shift();
260
267
  try {