@staix/agent-hub 0.12.7 → 0.12.9

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (39) hide show
  1. package/CHANGELOG.md +19 -0
  2. package/LICENSES/Apache-2.0.txt +204 -0
  3. package/THIRD_PARTY_NOTICES.md +19 -0
  4. package/docs/cooperbench.md +30 -2
  5. package/docs/events.md +32 -0
  6. package/docs/operations.md +26 -4
  7. package/docs/specs/2026-09-19-agent-hub-design.md +10 -0
  8. package/docs/specs/2026-09-20-pi-local-models-design.md +1 -1
  9. package/docs/specs/2026-10-04-switchyard-source-port-design.md +309 -0
  10. package/package.json +4 -2
  11. package/plugins/agent-hub/.claude-plugin/plugin.json +1 -1
  12. package/plugins/agent-hub/server.js +4 -2
  13. package/src/adapters/acp.ts +26 -9
  14. package/src/adapters/codex-appserver.ts +2 -2
  15. package/src/adapters/local-worker.ts +138 -26
  16. package/src/adapters/pi.ts +8 -0
  17. package/src/hub/daemon.ts +59 -17
  18. package/src/hub/events.ts +5 -0
  19. package/src/hub/inference.ts +20 -0
  20. package/src/hub/progress.ts +251 -0
  21. package/src/hub/routing.ts +4 -0
  22. package/src/local/tools.ts +9 -0
  23. package/src/models/relay.ts +136 -22
  24. package/src/models/route/advisor.ts +102 -0
  25. package/src/models/route/config.ts +41 -0
  26. package/src/models/route/escalation.ts +82 -0
  27. package/src/models/route/judge.ts +23 -0
  28. package/src/models/route/labels.ts +167 -0
  29. package/src/models/route/normalize.ts +103 -0
  30. package/src/models/route/plan-execute.ts +52 -0
  31. package/src/models/route/prompts.ts +8 -0
  32. package/src/models/route/relay-selector.ts +102 -0
  33. package/src/models/route/runtime.ts +156 -0
  34. package/src/models/route/signals.ts +234 -0
  35. package/src/models/route/stage.ts +93 -0
  36. package/src/models/route/state.ts +54 -0
  37. package/src/models/route/text.ts +61 -0
  38. package/src/omniroute/client.ts +8 -1
  39. package/templates/routing.toml +28 -0
package/CHANGELOG.md CHANGED
@@ -2,6 +2,25 @@
2
2
 
3
3
  Issue and pull request numbers in the entries for 0.7.7 and earlier refer to the previous repository, archived on 2026-09-30 when this repository's history was rewritten; the one exception is the open smoke-check issue, formerly #12, which moved here as #1. Numbers in newer entries refer to this repository.
4
4
 
5
+ ## Unreleased
6
+
7
+ ## 0.12.9
8
+
9
+ - The CooperBench native runner proves each prepared fixture root is still the directory preparation left before anything is locked, written or launched, and each arm checks its own again first (#119): a root replaced after preparation by a symlink to an equivalent outside tree passed the lexical `resolve()` comparison and the baseline content checks and would have redirected setup and agent writes there. The check is read-only (`lstat` and the real path, never a follow), so a substitution is refused without touching its target; a sibling fixture root that is itself a symlink is refused before the sibling-artifact walk reads it.
10
+ - The model relay no longer lets a gateway heartbeat name the served model (#137): an SSE event carries model identity only with generation activity (a delta with any field, a role-only first chunk included, a finish reason or a non-streaming message), so a keepalive with `choices: [{index: 0, delta: {}}]` cannot pin `actualModel` to the synthetic `keepalive` label and refuse the real model that follows. A heartbeat-only or cancelled-before-identification stream leaves the served model unknown; no requested alias is substituted.
11
+ - ACP tool identity is bound to the announced call, not the permission request title (#138): Qwen 0.24.7 announces an MCP call as `hub_send (agent-hub MCP Server)` and then titles the permission request with the serialized arguments, so the exact-name auto-approval never matched and the manual prompt showed the argument JSON twice. The announced title is cached per call id (cleared on reuse, evicted on completion, bound at the initial `tool_call` only — a later update's mutable display title never rewrites it), and a request resolves to the canonical `mcp__<server>__<tool>` name only when the server half is a server the session was configured with. Argument text never becomes an identity candidate, approval still picks only `allow_once`, and an unresolved or cut payload keeps its conservative manual path, displayed under the announced title.
12
+ - The model relay journals request-bound identity and cancellation provenance (#139): every upstream dispatch attempt (a fallback is its own record) closes exactly one sanitized `RelayRequestRecord` — resolved alias, upstream-configured model, observed provider and served model with their source (gateway header, generation SSE event, or local MLX configuration), outcome, duration and a confirmed-mismatch flag — exposed through `relay.requests()` (last 1000) and an `onRequest` hook that can neither break the proxied stream nor reject unobserved. A backend's mutable last-served label is never a request's evidence, HTTP 200 plus the requested alias identifies nothing, a request cancelled before identification stays explicitly unidentified, and a primary/auxiliary role stays unknown without native evidence. Records carry no prompts, tools, keys or Access headers.
13
+ - The Pi and Qwen native CooperBench study driver is versioned as manifest v3 and `scripts/benchmarks/native-pi-qwen.ts` (#140), pinning hub 0.12.9: the 2026-10-04 private calibration study (arms solo-pi, solo-qwen, joint-pi-qwen; Pi 1.0.1, Qwen 0.24.7; 60 preregistered attempts) becomes a regression-tested headless path that never touches Orca. The pinned build is the effective one — each native's `--version` runs under the final isolation environment, because Qwen's PATH bootstrap reported the managed 0.24.7 and fell back to base 0.24.1 under `QWEN_HOME` isolation; the protected-file probe counts only with structured denial evidence, never a model-written marker; source guards follow each case's `source_dirs` (`src/` for Click/Jinja, `dirty_equals/` for dirty_equals); new source files enter the binary submission patch (`git add -N`); setup resources are disposed on every exit path without erasing the active-window record; existing attempt evidence is rejected before any record is written (`wx` claims); served-model evidence is each request's own journaled record, where only cancelled-before-identification is non-evidence and an identified mismatch fails the gate whatever the outcome; and Qwen's peer tool approval is the exact canonical name through the adapter's #138 binding, not a title workaround. Grading flows through the same official evaluator adapter with quality, model and request-linkage coverage reported separately. The live cohort, official Docker controls and native readback remain manual live legs.
14
+
15
+ ## 0.12.8
16
+
17
+ - Port Switchyard's Stage signals/scoring, Plan/Execute, advisor gate and escalation policies in-process with source-based golden tests and Apache-2.0 attribution (#124, #125).
18
+ - Add opt-in `hub/` local-worker routes with fixed-model fallback, bounded fail-open judges, completed-turn REDO history and campus-only PII transport. Each local routing decision has an id and exactly one joined outcome label, including failure, cancellation and later escalation; labels contain identifiers and closed values, never text (#126).
19
+ - Add Pi's `hub/auto` alias with session-scoped Stage hold, MLX input/output admission and DGX fallback. Preserve explicit backend pins and the class-specific DGX fast/coding defaults (#127).
20
+ - Record aggregate peer progress and suggest reassignment only from repeated failures across known native turns or spinning. Exclude PII, group native attempts, ignore replayed Pi receipts and stop judging after latch; retain observation series and coverage limits in benchmark ledgers (#128).
21
+ - Preserve retrieval semantics for shell reads, attached output redirections and Unicode fingerprints. Budget escalation anchors to retain the newest trajectory; this intentional upstream correction is documented in the source-port spec.
22
+ - The control protocol remains unchanged. Optional typed judges, Pi advisor stream replay, Rust differential verification and sidecar retirement are deferred.
23
+
5
24
  ## 0.12.7
6
25
 
7
26
  - CooperBench runner and recovery (#120): from the moment the runner may change anything, before it locks its first input, `restoration.json` says `restored: false` and names the runner, until the runner writes its outcome. A runner killed after that point sends the recovery in, and the recovery waits while the runner it names still runs. A run directory whose earlier run is not restored is refused first, before the runner reads its inputs or any mode (the refusal names `restore.ts`), and checked again right before the runner locks anything, so a runner does not record locked modes as the originals; two runners started on one directory at the same moment are not guarded against. A directory whose earlier invocation ended restored, with no records, can still be used. `restore.ts` trusts `restored: true` only when the restoration ledger agrees (a ledger from before 0.12.5, with no runner identity, still has its locks restored; when `restoration.json` says `restored: true` or names a runner, its trust entry is kept as it was then), so a stale file over a re-run that died no longer reports a run restored while its inputs stay locked; it checks the runners again right before it restores anything. With no ledger, nothing was locked and nothing is restored.
@@ -0,0 +1,204 @@
1
+ Copyright (c) 2024-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
2
+
3
+
4
+ Apache License
5
+ Version 2.0, January 2004
6
+ http://www.apache.org/licenses/
7
+
8
+ TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
9
+
10
+ 1. Definitions.
11
+
12
+ "License" shall mean the terms and conditions for use, reproduction,
13
+ and distribution as defined by Sections 1 through 9 of this document.
14
+
15
+ "Licensor" shall mean the copyright owner or entity authorized by
16
+ the copyright owner that is granting the License.
17
+
18
+ "Legal Entity" shall mean the union of the acting entity and all
19
+ other entities that control, are controlled by, or are under common
20
+ control with that entity. For the purposes of this definition,
21
+ "control" means (i) the power, direct or indirect, to cause the
22
+ direction or management of such entity, whether by contract or
23
+ otherwise, or (ii) ownership of fifty percent (50%) or more of the
24
+ outstanding shares, or (iii) beneficial ownership of such entity.
25
+
26
+ "You" (or "Your") shall mean an individual or Legal Entity
27
+ exercising permissions granted by this License.
28
+
29
+ "Source" form shall mean the preferred form for making modifications,
30
+ including but not limited to software source code, documentation
31
+ source, and configuration files.
32
+
33
+ "Object" form shall mean any form resulting from mechanical
34
+ transformation or translation of a Source form, including but
35
+ not limited to compiled object code, generated documentation,
36
+ and conversions to other media types.
37
+
38
+ "Work" shall mean the work of authorship, whether in Source or
39
+ Object form, made available under the License, as indicated by a
40
+ copyright notice that is included in or attached to the work
41
+ (an example is provided in the Appendix below).
42
+
43
+ "Derivative Works" shall mean any work, whether in Source or Object
44
+ form, that is based on (or derived from) the Work and for which the
45
+ editorial revisions, annotations, elaborations, or other modifications
46
+ represent, as a whole, an original work of authorship. For the purposes
47
+ of this License, Derivative Works shall not include works that remain
48
+ separable from, or merely link (or bind by name) to the interfaces of,
49
+ the Work and Derivative Works thereof.
50
+
51
+ "Contribution" shall mean any work of authorship, including
52
+ the original version of the Work and any modifications or additions
53
+ to that Work or Derivative Works thereof, that is intentionally
54
+ submitted to Licensor for inclusion in the Work by the copyright owner
55
+ or by an individual or Legal Entity authorized to submit on behalf of
56
+ the copyright owner. For the purposes of this definition, "submitted"
57
+ means any form of electronic, verbal, or written communication sent
58
+ to the Licensor or its representatives, including but not limited to
59
+ communication on electronic mailing lists, source code control systems,
60
+ and issue tracking systems that are managed by, or on behalf of, the
61
+ Licensor for the purpose of discussing and improving the Work, but
62
+ excluding communication that is conspicuously marked or otherwise
63
+ designated in writing by the copyright owner as "Not a Contribution."
64
+
65
+ "Contributor" shall mean Licensor and any individual or Legal Entity
66
+ on behalf of whom a Contribution has been received by Licensor and
67
+ subsequently incorporated within the Work.
68
+
69
+ 2. Grant of Copyright License. Subject to the terms and conditions of
70
+ this License, each Contributor hereby grants to You a perpetual,
71
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
72
+ copyright license to reproduce, prepare Derivative Works of,
73
+ publicly display, publicly perform, sublicense, and distribute the
74
+ Work and such Derivative Works in Source or Object form.
75
+
76
+ 3. Grant of Patent License. Subject to the terms and conditions of
77
+ this License, each Contributor hereby grants to You a perpetual,
78
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
79
+ (except as stated in this section) patent license to make, have made,
80
+ use, offer to sell, sell, import, and otherwise transfer the Work,
81
+ where such license applies only to those patent claims licensable
82
+ by such Contributor that are necessarily infringed by their
83
+ Contribution(s) alone or by combination of their Contribution(s)
84
+ with the Work to which such Contribution(s) was submitted. If You
85
+ institute patent litigation against any entity (including a
86
+ cross-claim or counterclaim in a lawsuit) alleging that the Work
87
+ or a Contribution incorporated within the Work constitutes direct
88
+ or contributory patent infringement, then any patent licenses
89
+ granted to You under this License for that Work shall terminate
90
+ as of the date such litigation is filed.
91
+
92
+ 4. Redistribution. You may reproduce and distribute copies of the
93
+ Work or Derivative Works thereof in any medium, with or without
94
+ modifications, and in Source or Object form, provided that You
95
+ meet the following conditions:
96
+
97
+ (a) You must give any other recipients of the Work or
98
+ Derivative Works a copy of this License; and
99
+
100
+ (b) You must cause any modified files to carry prominent notices
101
+ stating that You changed the files; and
102
+
103
+ (c) You must retain, in the Source form of any Derivative Works
104
+ that You distribute, all copyright, patent, trademark, and
105
+ attribution notices from the Source form of the Work,
106
+ excluding those notices that do not pertain to any part of
107
+ the Derivative Works; and
108
+
109
+ (d) If the Work includes a "NOTICE" text file as part of its
110
+ distribution, then any Derivative Works that You distribute must
111
+ include a readable copy of the attribution notices contained
112
+ within such NOTICE file, excluding those notices that do not
113
+ pertain to any part of the Derivative Works, in at least one
114
+ of the following places: within a NOTICE text file distributed
115
+ as part of the Derivative Works; within the Source form or
116
+ documentation, if provided along with the Derivative Works; or,
117
+ within a display generated by the Derivative Works, if and
118
+ wherever such third-party notices normally appear. The contents
119
+ of the NOTICE file are for informational purposes only and
120
+ do not modify the License. You may add Your own attribution
121
+ notices within Derivative Works that You distribute, alongside
122
+ or as an addendum to the NOTICE text from the Work, provided
123
+ that such additional attribution notices cannot be construed
124
+ as modifying the License.
125
+
126
+ You may add Your own copyright statement to Your modifications and
127
+ may provide additional or different license terms and conditions
128
+ for use, reproduction, or distribution of Your modifications, or
129
+ for any such Derivative Works as a whole, provided Your use,
130
+ reproduction, and distribution of the Work otherwise complies with
131
+ the conditions stated in this License.
132
+
133
+ 5. Submission of Contributions. Unless You explicitly state otherwise,
134
+ any Contribution intentionally submitted for inclusion in the Work
135
+ by You to the Licensor shall be under the terms and conditions of
136
+ this License, without any additional terms or conditions.
137
+ Notwithstanding the above, nothing herein shall supersede or modify
138
+ the terms of any separate license agreement you may have executed
139
+ with Licensor regarding such Contributions.
140
+
141
+ 6. Trademarks. This License does not grant permission to use the trade
142
+ names, trademarks, service marks, or product names of the Licensor,
143
+ except as required for reasonable and customary use in describing the
144
+ origin of the Work and reproducing the content of the NOTICE file.
145
+
146
+ 7. Disclaimer of Warranty. Unless required by applicable law or
147
+ agreed to in writing, Licensor provides the Work (and each
148
+ Contributor provides its Contributions) on an "AS IS" BASIS,
149
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
150
+ implied, including, without limitation, any warranties or conditions
151
+ of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
152
+ PARTICULAR PURPOSE. You are solely responsible for determining the
153
+ appropriateness of using or redistributing the Work and assume any
154
+ risks associated with Your exercise of permissions under this License.
155
+
156
+ 8. Limitation of Liability. In no event and under no legal theory,
157
+ whether in tort (including negligence), contract, or otherwise,
158
+ unless required by applicable law (such as deliberate and grossly
159
+ negligent acts) or agreed to in writing, shall any Contributor be
160
+ liable to You for damages, including any direct, indirect, special,
161
+ incidental, or consequential damages of any character arising as a
162
+ result of this License or out of the use or inability to use the
163
+ Work (including but not limited to damages for loss of goodwill,
164
+ work stoppage, computer failure or malfunction, or any and all
165
+ other commercial damages or losses), even if such Contributor
166
+ has been advised of the possibility of such damages.
167
+
168
+ 9. Accepting Warranty or Additional Liability. While redistributing
169
+ the Work or Derivative Works thereof, You may choose to offer,
170
+ and charge a fee for, acceptance of support, warranty, indemnity,
171
+ or other liability obligations and/or rights consistent with this
172
+ License. However, in accepting such obligations, You may act only
173
+ on Your own behalf and on Your sole responsibility, not on behalf
174
+ of any other Contributor, and only if You agree to indemnify,
175
+ defend, and hold each Contributor harmless for any liability
176
+ incurred by, or claims asserted against, such Contributor by reason
177
+ of your accepting any such warranty or additional liability.
178
+
179
+ END OF TERMS AND CONDITIONS
180
+
181
+ APPENDIX: How to apply the Apache License to your work.
182
+
183
+ To apply the Apache License to your work, attach the following
184
+ boilerplate notice, with the fields enclosed by brackets "[]"
185
+ replaced with your own identifying information. (Don't include
186
+ the brackets!) The text should be enclosed in the appropriate
187
+ comment syntax for the file format. We also recommend that a
188
+ file or class name and description of purpose be included on the
189
+ same "printed page" as the copyright notice for easier
190
+ identification within third-party archives.
191
+
192
+ Copyright (c) 2024-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
193
+
194
+ Licensed under the Apache License, Version 2.0 (the "License");
195
+ you may not use this file except in compliance with the License.
196
+ You may obtain a copy of the License at
197
+
198
+ http://www.apache.org/licenses/LICENSE-2.0
199
+
200
+ Unless required by applicable law or agreed to in writing, software
201
+ distributed under the License is distributed on an "AS IS" BASIS,
202
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
203
+ See the License for the specific language governing permissions and
204
+ limitations under the License.
@@ -0,0 +1,19 @@
1
+ # Third-party notices
2
+
3
+ ## NVIDIA NeMo Switchyard
4
+
5
+ The TypeScript routing modules in `src/models/route/` contain modified ports from NVIDIA NeMo Switchyard, commit `c8848511a7e2e1d605070c7a68905bdc24c6481a`. These portions are licensed under Apache-2.0; see [the license](LICENSES/Apache-2.0.txt). The remainder of agent-hub retains its existing MIT license.
6
+
7
+ Switchyard
8
+ Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
9
+
10
+ This product is licensed under the Apache License, Version 2.0 (the "License").
11
+ You may obtain a copy of the License in the LICENSE file at the root of this
12
+ repository, or at:
13
+
14
+ http://www.apache.org/licenses/LICENSE-2.0
15
+
16
+ Unless required by applicable law or agreed to in writing, software distributed
17
+ under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR
18
+ CONDITIONS OF ANY KIND, either express or implied. See the License for the
19
+ specific language governing permissions and limitations under the License.
@@ -22,7 +22,7 @@ python3 scripts/benchmarks/runner.py prepare \
22
22
  --output /private/path/to/new-run
23
23
  ```
24
24
 
25
- The manifest's archive and prompt hashes are checked before a fixture is accepted. The `prepared.json` ledger binds the fixture roots, baseline commits, complete baseline path counts and manifest hash.
25
+ The manifest's archive and prompt hashes are checked before a fixture is accepted. The `prepared.json` ledger binds the fixture roots, baseline commits, complete baseline path counts and manifest hash. At run time each fixture root must still be the directory preparation left: a root replaced by a symlink to an equivalent outside tree, a non-directory, or a name whose real path differs from the prepared one is refused read-only before anything is locked, written or launched (#119).
26
26
 
27
27
  ## Native execution
28
28
 
@@ -51,15 +51,23 @@ bun scripts/benchmarks/native.ts \
51
51
  --cases 0
52
52
  ```
53
53
 
54
+ During a measured calibration block, suspend unrelated Claude and Codex sessions so that quota growth between snapshots can be attributed to the attempt. `--setup-only` does not enter an active-work window and leaves the pre reading explicitly unknown. The first study block validates that quota readings are actually collected for Claude-participating arms (non-`unknown` `claudeUsage.pre` and `claudeUsage.post` in those run records) before subsequent blocks rely on them.
55
+
54
56
  The native runner runs on macOS only and refuses anything else (each run record says `platform`): the arms run in Orca terminals, and on Linux a clock step moves the start times its teardown proves processes by. Before native execution, obtain explicit user authorization before changing Orca registrations; a benchmark request alone is not authorization. Only after that authorization, an operator manually registers each exact fixture root with `orca repo add --path '<fixture>'` and confirms `orca worktree list --repo id:<repo-id>` shows that exact path. The runner performs read-only exact-path repo/worktree lookup; it preflights every selected fixture before changing fixture files, input modes, cohort or attempt records. If either identity is missing or mismatched, it exits nonzero with setup guidance and leaves the prepared run available for retry after explicitly authorized registration. Each arm rechecks its identity before changing fixture files. Benchmark and agent workflows never add or remove Orca registrations. The runner uses the project CLI to start and stop each daemon, and requires exact worktree/cwd readback for native sessions. Claude starts through the canonical `ahub claude` guard and loads this checkout's candidate bundle through an exact session-only `--mcp-config` server (`server:agent-hub`); it does not promote or mutate the globally installed plugin. Codex uses the native app-server adapter and `workspace-write` sandbox. Each agent must execute a setup-only `head -c 1` probe against the exact protected file and produce only the denied marker before scored tasks begin. There is one native sandbox layer per agent.
55
57
 
58
+ Fixture roots are checked without following a root symlink, and their device/inode identity is captured before asynchronous setup (#119). After the baseline and sibling lookups, the synchronous setup mutation batch rechecks that identity before writing. Git setup, daemon startup, Claude terminal creation (after its terminal-list lookup), and Codex startup also recheck at their launch boundaries. A symlink or ordinary directory substituted during a preceding await is rejected before that boundary uses it. These checks do not pin a directory descriptor across filesystem syscalls or an external launcher: an operator must keep the prepared root and its ancestors unchanged throughout the run. Concurrent same-user renames between a check and a syscall or inside an external launcher are outside this guarantee; this is not a claim of atomic directory pinning.
59
+
56
60
  Every run record, diff, PTY/provider trace and evaluation artifact stays in the mode-0700 private run directory. Setup errors, structured provider quota errors, hub budget pauses, unsettled deliveries and interruptions stay separate and unscored. A quota snapshot or numeric `429` alone is not a provider error. On exit the runner restores the exact input and read-lock modes, the sibling artifact modes and the scoped Claude trust flag (removing its own fresh Claude trust entry), under the condition below. Teardown (issue #113) proves what it stops. Each process an arm starts is recorded with its pid and start time and the evidence that it is the arm's: the daemon by the pid in its state directory and an argv that serves this fixture (a later daemon there is a replacement, never adopted), Claude's launch chain by the arm's own session id as its `--session-id` (and the Orca terminal shell that runs it, working in the fixture), the Codex app-server as the daemon's child. What runs below them, and what is in a group a recorded process leads while that group is known to be the same one (its leader alive, or a recorded member still in it), is the arm's too; the table is read every 5 s while the agents work and during the completion wait, so a tool command's background job is recorded while its parent runs (one that detaches between two reads, outside the fixture and without it in its argv, is not seen). A fixture name in an argv is never proof on its own (Codex's need not name the fixture), and every recorded actor counts even when a read of the table fails. Teardown pauses every agent, reads the deliveries still owed, lets a completed arm's Claude end the turn it is in (up to 30 s and never past the 300 s limit (the bound is taken when the wait starts; one 250 ms poll can pass it), judged by the `turn_duration` row Claude Code writes after a turn, proved on the arm's probe turn; a stopped or timed-out arm does not wait; a completed arm's tree is hashed at the end of its active time and again after the teardown, and a write in between flags the attempt invalid), asks for the normal shutdown (Claude's terminal closed, `ahub kill`, the hub's project registration removed (`ahub projects remove`); the runner leaves Orca registrations untouched; a step of it that fails is recorded in `cleanup.normal.errors` and does not by itself fail the attempt, since the process readback decides, so a hub project registration left behind is listed there), and reads the process table back, in the C locale and in UTC (start times are read as identities, the same in every reader). What is still running after a settle period gets SIGTERM; what is left then is frozen (SIGSTOP), the table read again, and it is killed (SIGKILL), what was frozen and not killed being continued: a frozen process starts nothing, so that read sees all of it. Every signal goes only to an identity read again just before, by the group the table shows it in now; a group leader is signalled with its group. What the read after the freeze shows for the first time is frozen too and read again before anything is killed, and when such a read fails only what a STOP reached is killed, by the last read that showed it (a stopped process keeps its pid). Anything with the fixture in its argv or as its working directory that is not proved the arm's (a background job that escaped the reads, someone's shell) is never signalled and keeps the cleanup open; the record keeps its pid, start time, program name (the name of its executable as the kernel recorded it at exec, `ps -o ucomm`, and only while the pid has the start time recorded; never `comm`, which on macOS is the process's own argv[0] and can carry its arguments; omitted when it cannot be read) and working directory, never its arguments. The record says `clean`, `clean_with_fallback` or `incomplete_or_unknown` (`cleanup`, with the normal shutdown's errors, every signal sent, what is still running and what is unresolved), and `cleanup_complete` is true only for the first two. Evidence follows the teardown: the transcript prefix (`transcriptBytes`, `transcriptSha256`, with the session and time it was taken) and the patch. The sibling read locks come off and the protected inputs become readable again only when the cleanup is complete; otherwise both stay, the record goes to `recovery/` beside the run's restoration ledger (which lists every arm's recorded processes and the runner's own identity), the cohort stops, and `bun scripts/benchmarks/restore.ts --run <run>` (also `runner.py restore`, which calls it) puts the modes back and takes back the trust entry once the runner, every recorded process and anything with a fixture in its argv or as its working directory are gone, then moves the kept records into `runs/`, where grading reads them as unavailable. A ledger written before runner identities needs `--runner-exited`, the operator's statement that the runner is gone. Runner commands are bounded (a hung one is killed), and SIGHUP stops the run like SIGINT and SIGTERM. `restoration.json` says `restored: true` only when the protected inputs and every arm's sibling locks are back. From before the runner locks its first input until it writes that outcome, it says `restored: false` and names the runner (#120), so a runner that is killed sends the recovery in, and the recovery waits while that runner still runs. The runner refuses a run directory whose earlier run is not restored (its `restoration.json` or its ledger says so) first, before it reads its inputs or any mode, and again right before it locks anything, since locked modes would become the originals; `restore.ts` first. One runner per directory at a time: two started on one directory at the same moment are not guarded against. `restore.ts` trusts `restored: true` only when the restoration ledger agrees (a ledger from before 0.12.5 still has its locks restored; when `restoration.json` says `restored: true` or names a runner, its trust entry is kept as it was then), and checks the runners again right before it restores anything. The trust entry is taken back with them (by the recovery when the cleanup is incomplete), and kept when the user changed it meanwhile; when the runner's own write never landed (it stopped between recording the lease and the rename), the outcome is `not_written` and nothing is taken back; for a runner that died there, the recovery takes back only an entry exactly as the runner would have written it, and none while the runner's temp file is still there. Orca registrations are not removed during benchmark teardown. Legacy benchmark-created entries may be removed only after provenance confirms the exact fixture path, using `orca project setup-delete --setup <id>` after verifying the exact repo and fixture path; unrelated registrations and fixture/Git/evidence data remain untouched. The ledger reports a record kept in `recovery/` as an attempt with `withheld: true` (unavailable: its cleanup was incomplete, or its sibling read locks could not be put back) until `restore.ts` moves it, never as missing (a record in both places, a move that did not finish, counts once and withheld); while `runs/` is still locked (a cleanup is incomplete), the attempts it cannot read there are listed as `unreadable`, never as missing, and its teardown row carries the normal shutdown's errors (`normal_errors`, for example a hub project registration left behind). The recovery also removes the temp files a runner left (one that died mid trust write or mid restore, or could not remove its own), each a copy of `~/.claude.json`; a write the runner knew never landed stays `not_written` there too, and no entry is touched. Run records carry `completion`, `cleanup`, `restoration` and `stages` (completion wait, shutdown, settle, fallback with the final readback, evidence and restoration times), apart from the end reason: `end_reason_detail` keeps the original classification, and `end_flags` (an unverified model, modified metadata, a tree changed or unverified after the active time) makes any end other than an interruption, a provider quota error or a budget pause an infrastructure error beside it, which grading and the ledger report as `end_story`; a record with `teardown_errors` (a patch, transcript prefix, last capture, fixture metadata or event log that could not be taken, or sibling read locks or the trust entry that could not be put back, a concurrent change of the trust entry included), an incomplete cleanup or a trust entry not taken back is unavailable to grading and to the ledger alike (a record from before 0.12.5 is judged as it was then, by its own `cleanup_complete` and `trust_restored`, and the ledger shows its teardown with `verified: false`: 0.12.3 and 0.12.4 set `cleanup_complete` when the shutdown steps reported success (commands, terminal close, lock and trust restores), with no process readback, which is no proof that the processes were gone); runner commands, the `ps` reads included, run in process groups of their own, so a Ctrl-C reaches the runner, which stops in this order, and not the command it is running; the completion wait's tree hash never writes the agent's index. `teardown.ts` and the process table it reads (`src/hub/child-process.ts`) are pinned in `prepared.json`, and `teardown.ts` in `cohort.json` and `grade.json` too, with the other runner sources.
57
61
 
58
62
  ## Manifest v2: the turn-free arm
59
63
 
60
64
  `scripts/benchmarks/manifest-v2.json` keeps v1's cases, models and limits and adds a fourth arm, `hub-turnfree-codex-claude` (issue #110): the same two agents and assignment rotation as `hub-codex-claude`, with `coordination: "turn-free"` in the fixture's hub config and fixture instructions that tell the owners not to message each other. A v1 manifest still validates and prepares. Its cohorts are run, graded and reported with the release and runner sources they were prepared with: `native.ts` refuses a manifest whose hub version is not the checkout's (v1 pins 0.12.3), and `runner.py` refuses to grade a cohort whose runner sources differ from its own.
61
65
 
62
- Hooks are equal across arms: no arm runs the user's or a plugin's hooks, and no arm runs a status line (`disableAllHooks` turns it off in the other arms, so the turn-free arm leaves it out; Claude's quota reaches the hub in no arm). Claude starts with `--setting-sources project` and `--strict-mcp-config`; the solo and advisory arms set `disableAllHooks`, and the turn-free arm's session settings carry the hub's own hooks (before and after every tool call, and at Stop) and nothing else, because they are the treatment. Every Codex thread starts with `features.hooks` off: the turn-free arm's Codex boundary is the adapter's steer into the running turn, whose readback is the steered input coming back as a user message item. MCP servers are isolated too: Claude has only the hub's (`--strict-mcp-config`), and every Codex runs through a wrapper written into each arm's fixture that turns off the user's plugins, apps, sub-agents and turn-end notifier and disables each MCP server the user's config defines, by name, so it starts only the hub's; the user's config is not changed. Instructions: Claude reads the fixture's `AGENTS.md` through `--append-system-prompt-file` (Claude Code reads `CLAUDE.md`, not `AGENTS.md`), Codex as the project's `AGENTS.md`; Codex also reads the user's global `AGENTS.md`, which only a separate Codex home with its own login would leave out (the run records say so), and the user's Codex skills stay available to it (the wrapper leaves them on; `conditions.codex.skills` records what app-server's `skills/list` reports for the fixture: counts by scope, how many are enabled, and a hash of their names, never their bodies or paths, and the `skills/list` answer itself is not kept in `codexMessages`; Claude's Skill tool is denied); both are the same in every arm. Each run record carries these `conditions`; its `events` carry the hub's `capability` readbacks.
66
+ Every Claude-participating arm runs the same isolated status-line tee (`src/cli/statusline-tee.ts`) to produce `benchmark-usage/claude-usage.json` under the arm's hub state directory for #134. The daemon consumes only the parent `claude-usage.json` for budget and routing; the observer's separate directory leaves those inputs unchanged. Generated session settings set `disableAllHooks: false` because Claude Code gates status-line execution with that flag; the solo and advisory sessions have an empty hook map, and the turn-free session carries only the hub's facts hooks. The existing `--restricted` launch ignores user, project and local settings files; the explicit `--settings` file and managed settings still apply. User/plugin hook settings are not copied into the session. The tee has no original status-line command to run. This instrumentation profile is recorded as `conditions.claude.statusLine: true` and `disableAllHooks: false`, and source/settings hashes prevent mixing it with historical hookless runs. See [Claude Code status-line execution and troubleshooting](https://code.claude.com/docs/en/statusline#troubleshooting).
67
+
68
+ Every emitted arm record carries `claudeUsage.pre` and `claudeUsage.post`. Claude arms wait up to 1.5 seconds for an observer write after probe completion plus the 300 ms status-line debounce before starting active work. The isolated producer refreshes once per second. After active work, a post reading is trusted only if the final turn ended and an observer write follows that settled boundary plus the debounce. Timeout, interruption, unsupported/unsettled turn completion or a missing fresh write records explicit unknown; a cached pre value is never treated as post consumption. These waits stay outside active-work time and do not make collection failure fail an attempt. Solo-Codex launches no Claude just to measure quota, and records explicit unknown with the reason that there is no Claude actor. Readings retain the file's own `at`; expired or invalid-reset windows are marked stale, all-stale/missing/unreadable/malformed data is unknown, and collection failure does not fail an attempt. The file reader refuses symlinks, special files and oversized data, so a FIFO cannot hold teardown. Grade `native_usage.claude` keeps transcript-derived token counts and adds pre/post. Rate-limit fields are optional and appear only after an API response on supported accounts: producer wiring does not establish live provider collection. The first study block must validate actual non-unknown Claude-arm readings before later blocks rely on quota deltas.
69
+
70
+ Claude starts with `--setting-sources project` and `--strict-mcp-config`; the turn-free arm's before/after-tool and Stop hooks remain its treatment. Codex threads keep `features.hooks` off. MCP servers are isolated too: Claude has only the hub's (`--strict-mcp-config`), and every Codex runs through a wrapper written into each arm's fixture that turns off the user's plugins, apps, sub-agents and turn-end notifier and disables each MCP server the user's config defines, by name, so it starts only the hub's; the user's config is not changed. Instructions: Claude reads the fixture's `AGENTS.md` through `--append-system-prompt-file` (Claude Code reads `CLAUDE.md`, not `AGENTS.md`), Codex as the project's `AGENTS.md`; Codex also reads the user's global `AGENTS.md`, which only a separate Codex home with its own login would leave out (the run records say so), and the user's Codex skills stay available to it (the wrapper leaves them on; `conditions.codex.skills` records what app-server's `skills/list` reports for the fixture: counts by scope, how many are enabled, and a hash of their names, never their bodies or paths, and the `skills/list` answer itself is not kept in `codexMessages`; Claude's Skill tool is denied); both are the same in every arm. Each run record carries these `conditions`; its `events` carry the hub's `capability` readbacks.
63
71
 
64
72
  Validity, decided by the grader and applied by the ledger alike: an attempt whose records show a hook or an MCP server that is not the hub's, or whose Claude transcript cannot be read, is unavailable. A turn-free attempt is valid only with its context paths working: both verified before its tasks, and none lost, no cohort lifted and none formed open while the agents worked; teardown comes after that and does not count. Whether the agents' plans overlapped, so that a cohort formed at all, is their doing after assignment and is not a condition: every turn-free attempt without a capability failure counts for the arm, and the ledger reports the treatment received and a median over treated attempts beside it. A capability failure while the agents work is the one exclusion after assignment, because #110 forbids reporting it as a turn-free run (AC3); such attempts are listed with their reasons, never dropped silently. Each attempt records its transcript's length and hash when it ends; validity and every Claude measure are read from that prefix (Claude Code may append rows after it exits: what it writes after the prefix was taken is never counted, the ledger reports its size as `late_append_bytes`, and a turn still open in the prefix leaves that agent's settlement unknown), and a prefix that changed counts as unreadable.
65
73
 
@@ -67,6 +75,26 @@ Validity, decided by the grader and applied by the ledger alike: an attempt whos
67
75
 
68
76
  Pass `--repeat <n>` for the n-th repeat of a case (0 for the first): the arm order is row (case index + repeat) of a Williams design (0, 1, n-1, 2, n-2, ... shifted by the row), so over n consecutive rows every arm runs right before every other one once and repeats of one pair change which arm runs last. The manifest's `plan` names the release pilot (case 0, three repeats, 12 attempts, an active-time ceiling of one hour) and the study (ten cases, two repeats, 80 attempts, 6 hours 40 minutes at 300 s each, setup, grading and teardown excluded); `runner.py` refuses a plan whose attempts or ceiling do not follow from its arms, cases and repeats.
69
77
 
78
+ ## Manifest v3: the headless Pi and Qwen arms
79
+
80
+ `scripts/benchmarks/manifest-v3-pi-qwen.json` (issue #140) keeps the upstream commit, the ten cases and the 300 s active limit, and versions the 2026-10-04 private Pi/Qwen calibration study as a regression-tested execution path: arms `solo-pi`, `solo-qwen` and `joint-pi-qwen`, native builds Pi 1.0.1 and Qwen 0.24.7, one fixed backend (`fixed_backend`) with its expected served model and provider, and a preregistered plan of 60 attempts (ten cases, two repeats, three arms; the six-row odd-n Williams layout, joint ownership `(caseIndex+repeat)%2`, fixed before any model call). Prompt bodies, hidden tests, gold solutions and vendor state stay external, hashes only, exactly as in v1/v2.
81
+
82
+ The driver is `scripts/benchmarks/native-pi-qwen.ts`, headless: it spawns Pi and Qwen (ACP) itself over the hub's model relay and never touches Orca — no terminal, no registration, and the canonical Orca UI-path rules of the Claude/Codex runner above are unchanged. It keeps the runner's protections: private mode-0700 run root outside the repository, umask 0077, macOS-only, SIGHUP/SIGINT/SIGTERM stop the cohort in order, prepared fixture root identity checks (#119), the reuse refusal with `restoration.json`/`restoration-ledger.json` (#120), process ownership proven by `teardown.ts` (each peer process recorded with pid and start time, the table followed while the agents work, signals only to proven identities), and no attempt record ever written over another (the attempt directory must be new, `started.json` and `native-owner.json` are exclusive `wx` claims).
83
+
84
+ Its calibration corrections, relative to what an ad-hoc harness gets wrong:
85
+
86
+ 1. The pinned build is the EFFECTIVE build: each native's `--version` runs under the final isolation environment (Qwen inside its seatbelt profile with `QWEN_HOME`/`TMPDIR` set — the PATH bootstrap reported the managed 0.24.7 in the normal home and fell back to base 0.24.1 under isolation), and the binary path and version are recorded in every attempt record, so launch and recovery emit the same build.
87
+ 2. The protected-file probe passes only on a native read attempt with structured denial evidence — a guard denial for Pi, a failed read tool call with EPERM/EACCES for Qwen, plus the seatbelt kernel probe — never on a model-written `AHUB_PROBE_DENIED` marker alone.
88
+ 3. Source guards follow each case's `source_dirs` (`src/` for Click/Jinja, `dirty_equals/` for dirty_equals).
89
+ 4. New source files enter the binary submission patch (`git add -N` on the source dirs before `git diff --binary`); grading re-collects the same scoped patch and separately refuses any change outside the guarded dirs.
90
+ 5. Relay, tool server and peers are disposed on every exit path; a disposal failure is recorded in `teardown_errors` and never erases the active-window tree record captured before disposal, and the relay is never left open.
91
+ 6. Existing attempt evidence is rejected before any fallback record is written.
92
+ 7. The driver is preflighted by strict typechecking (`scripts/check.sh`) and executable lifecycle checks (every spawnable binary answers a trivial command before a fixture is touched), not transpilation alone.
93
+
94
+ Served-model evidence is per request, from the relay's journaled `RelayRequestRecord` (#139), not from a response wrapper: a completed request that was never identified fails the attempt's model gate (a heartbeat-only stream identifies nothing), an observed mismatch flags it, and a request cancelled before identification stays explicitly unidentified and never certifies another request. Qwen's peer tool is approved through the adapter's shipped tool-identity binding (#138): the announced `hub_send (pilot-peer-bus MCP Server)` title resolves to the canonical `mcp__pilot-peer-bus__hub_send`, the only name on the exact whitelist. The joint arm's MCP peer bus is the versioned `scripts/benchmarks/peer-bus-mcp.py`, pinned with the driver in `prepared.json` and `cohort.json` next to the candidate source pins (`src/adapters/pi.ts`, `src/adapters/acp.ts`, `src/models/relay.ts`).
95
+
96
+ Grading flows through the same official evaluator adapter and controls; a quality failure is never a retry selector and no evaluator feedback reaches the candidate agents during generation. Reports keep quality, model-identity and request-linkage coverage separate per arm, with unavailable attempts (failed, missing or unavailable) retained in the planned denominator, and name the usage units (Pi's incremental `onTokens` counter, Qwen's session `usage_update` running total — never added together) and the tool-surface difference (Pi's hub-moderated tools against Qwen's own seatbelted auto-edit tools). A live cohort, the official Docker controls and a native readback of an attempt's records remain manual live legs requiring accounts and the pinned archives; they are not part of the checked-in tests.
97
+
70
98
  ## Coordination ledger
71
99
 
72
100
  ```sh
package/docs/events.md CHANGED
@@ -18,6 +18,11 @@ marked `private: true`, and PII tasks `pii: true`.
18
18
  | `quiet` | `id`, `from`, `peers`: an agent message held back from these members of a silent turn-free cohort; its other recipients got it (issue #107) |
19
19
  | `fact` | `peer`, `id` (the offer), `files` (files whose diff it carried), `plans`, `unknown` (files whose change it showed with attribution unknown), `named` (files under a named directory it named without a diff, those counted as "N more" or held back by the PII filter included), `bytes` (the injected text), `via` (`hook` for Claude, `steer` for Codex, `done` with an integration request), `ms` (the hub's time to build it), `hookMs` (the hook process's own start-up and connect time), `accepted` (whether app-server took the steer), `unanswered` (app-server did not answer it within 10 s: it may have gone in), `rttMs` (an accepted steer: from sending it to app-server's answer), `probe` (a context check), `coverage` (it named files earlier changes are not covered for): one fact offer (issue #108) |
20
20
  | `fact_ack` | `peer`, `id`, `via` (`hook` and `steer`: a readback found the offer in the native session; `done`: the next `hub_task_done`), `ms` (from the offer): an acknowledged offer, the only thing that moves a peer's view |
21
+ | `route` | `peer`, `route`, `tier`, `source` (`override`, `dimensions`, `hold`, `classifier`, `default`), `score`, `ms`. Local decisions also carry `decision`, `turn`, optional numeric `task`, `pii`, and decision-time `severity`, `spinning`, `exploring`, `production`; no prompt text |
22
+ | `route_outcome` | `peer`, `decision`, `turnId`, `turn` (`completed` or `failed`), optional `task`, `pii`, `latched`, optional `next` (`severity`, `tests`: `pass`/`fail`/`none`, `repeat`) and `advisor` (`approve`/`redo`/`failed`). Exactly once per local decision at turn settlement, including failure/cancellation |
23
+ | `advisor` | `peer`, `route`, `trigger`, `verdict` (`approve`, `redo`, `failed`), `discardedChars`: an advisor check result; no transcript or feedback text |
24
+ | `progress` | `peer`, `task` (task id), `severity`, `spinning`, `exploring`, `production`: normalized tool activity dimensions for an open task; coverage differs by peer and no commands or task text are recorded |
25
+ | `stuck` | `peer`, `task` (task id), `category` (`repetition`, `false_progress`, `drift`, `desperation`, `capability_gap`), `streak`, `latched`: an escalation verdict; it recommends considering reassignment and never reassigns automatically |
21
26
  | `capability` | `peer`, `state` (`verified` or `lost`), `via`: a peer's context path for facts |
22
27
  | `native_turn_end` | `peer`: Claude's Stop hook, the end of its turn (Codex's is its `turn_end`); quiescence evidence for an integration |
23
28
  | `hook_stats` | `peer`, `n` (facts hook calls in the turn, its Stop included), `startupMs` and `maxStartupMs` (the hook processes' start-up and connect time, summed and the largest), `hubMs` (the hub's own time for them): at Claude's Stop (issue #108) |
@@ -47,3 +52,30 @@ Token usage by adapter:
47
52
 
48
53
  The file is local and never uploaded. It grows without rotation; delete it to start
49
54
  over (the hub recreates it).
55
+
56
+ ## Offline route outcome joins
57
+
58
+ For local-worker decisions, join `route_outcome.decision` to `route.decision` (one-to-one).
59
+ The route row records the chosen tier and decision-time dimensions; `next` describes only
60
+ the tool results answering that response, not an aggregate over later unrelated calls.
61
+ A final no-tool response has no `next`. Advisor checks add their closed verdict.
62
+ `latched` includes escalation that occurred later in the same turn.
63
+
64
+ Join `route.turn` or `route_outcome.turnId` to `turn_start.turn` / `turn_end.turn`.
65
+ The outcome's `turn` field is the completion status, not the turn id. To add offline
66
+ review supervision, join `route.task` to `task.id` and inspect subsequent task events
67
+ whose `state` is `approved` or `changes_requested`. Keep their timestamps so a review
68
+ from a different task revision is not treated as immediate per-call evidence.
69
+
70
+ Exclude both decisions and outcomes with `pii: true` before exporting a training set.
71
+ Labels contain ids, numbers, booleans and closed values only; they never retain tool
72
+ output, task text, advisor feedback, or raw failure fingerprints. A failed telemetry
73
+ sink does not affect model selection, budget admission or delivery. Pi outcome labels
74
+ are deferred: its existing route events retain their original shape. No control WebSocket
75
+ message changed, so the protocol number is unchanged.
76
+
77
+ Progress coverage is asymmetric: Codex contributes completed command and file-change items, local and Pi contribute
78
+ per-model-step tool observations, and Claude contributes tool inputs from turn-free pre-hooks without tool results or
79
+ a native turn id. Missing peers or intervals therefore mean unobserved, not zero activity. Escalation uses only known
80
+ native turn ids; observations without one still contribute progress dimensions but cannot establish separate attempts
81
+ or the eight-turn spinning threshold.
@@ -647,21 +647,21 @@ Rows without a live process are stale registrations; forget them with
647
647
 
648
648
  Upgrade running projects with the target release's own coordinator. It accepts
649
649
  a running source on control protocol 9 (0.6.x), 10 (0.7.0 through 0.12.0),
650
- 11 (0.12.1 and 0.12.2), 12 (0.12.3) or 13 (0.12.4 through 0.12.6) and only
650
+ 11 (0.12.1 and 0.12.2), 12 (0.12.3) or 13 (0.12.4 through 0.12.9) and only
651
651
  a target on its own protocol, so the target's coordinator fits every supported
652
652
  source and carries every recovery fix released up to it. Protocol 8 and older
653
653
  (0.5.x and earlier) are refused as `manual-bootstrap-required`. Run from the
654
654
  project directory, without replacing the global CLI first:
655
655
 
656
656
  ```bash
657
- bunx --package @staix/agent-hub@0.12.7 ahub upgrade --to 0.12.7 --dry-run
658
- bunx --package @staix/agent-hub@0.12.7 ahub upgrade --to 0.12.7 --yes
657
+ bunx --package @staix/agent-hub@0.12.9 ahub upgrade --to 0.12.9 --dry-run
658
+ bunx --package @staix/agent-hub@0.12.9 ahub upgrade --to 0.12.9 --yes
659
659
  ```
660
660
 
661
661
  | Running now | Coordinator to use |
662
662
  | --- | --- |
663
663
  | 0.6.x (protocol 9) | the target's, through `bunx` as above |
664
- | 0.7.0 through 0.12.0 (protocol 10), 0.12.1 and 0.12.2 (protocol 11), 0.12.3 (protocol 12), 0.12.4 through 0.12.6 (protocol 13) | the target's, through `bunx` as above |
664
+ | 0.7.0 through 0.12.0 (protocol 10), 0.12.1 and 0.12.2 (protocol 11), 0.12.3 (protocol 12), 0.12.4 through 0.12.9 (protocol 13) | the target's, through `bunx` as above |
665
665
  | any supported source, with the installed CLI already at the target | `ahub upgrade` below, which is the same coordinator |
666
666
  | 0.5.x or earlier (protocol 8 and older) | not supported: bootstrap by hand with the matching CLI |
667
667
 
@@ -874,3 +874,25 @@ headroom for templates/tool metadata and choose a larger dedicated recipe
874
874
  only after measuring memory. This path does not change local-worker PII
875
875
  routing or the DGX backend. Rollback requires explicitly restoring the old
876
876
  config with `provider: "legacy"`; Ollama errors never launch Python.
877
+
878
+ ### In-process model routes
879
+
880
+ `routing.toml` supports `[hub_routes."hub/<id>"]`, separate from sidecar `[routes]`.
881
+ Set `[local] route = "hub/stage"` or a task class's `route` to enable a hub route.
882
+ The shipped examples use OmniRoute `fast` and `coding`; `fixed_model` remains the fallback.
883
+ `--model` pins the worker and bypasses both routing engines. Sidecar `sy/` routes remain optional.
884
+
885
+ Route types: `stage` scores recent tools and holds capable recovery for two calls;
886
+ `plan_execute` plans on capable until the first mutation and then stays efficient;
887
+ `advisor` holds no-tool answers for APPROVE/REDO review; `escalation` starts efficient
888
+ and latches capable after two same-category judgements with new evidence.
889
+ Optional judges have an eight-second deadline and five-minute failure backoff.
890
+ REDO feedback is kept in the completed local turn history and counts toward `max_steps`.
891
+ PII calls require a positively confirmed campus gateway immediately before transport;
892
+ PII turns never enter shared history, memory capture, or progress observation.
893
+ The `route`, `advisor`, `progress` and `stuck` events contain identifiers and aggregates only.
894
+
895
+ Pi exposes `hub/auto` for stage routing when available. Fixed `dgx/coding`, `dgx/fast`
896
+ and `mlx/fast` aliases still pin the backend. Automatic MLX selection admits the complete
897
+ input, tool schemas and requested output within the configured context window.
898
+ Progress judgements suggest reassignment; they never change task ownership.
@@ -317,6 +317,16 @@ pii = "local_only" # hard: on-prem models only
317
317
  budget_paused = "skip_peer"
318
318
  ```
319
319
 
320
+ L2 in-process routing (2026-10-04 source port): `[hub_routes."hub/<id>"]` selects
321
+ stage, plan_execute, advisor or escalation policies inside the hub. The host retains
322
+ transport, budget admission and complete-turn history ownership; algorithms make decisions
323
+ from normalized chat input and bounded session state. See
324
+ [the source-port contract](2026-10-04-switchyard-source-port-design.md) for upstream
325
+ provenance, intentional fail-open escalation and REDO history differences. These tables
326
+ are never copied to sidecar configuration. Existing `sy/` routes and fixed-model fallback
327
+ remain supported. The Pi `hub/auto` alias uses the same stage signals with session-scoped
328
+ hold and context admission; peer progress observations suggest but never execute reassignment.
329
+
320
330
  L2, Switchyard sidecar (amended in M3, verified against `switchyard-server` 0.2.0 built with
321
331
  `cargo install --locked switchyard-server`; there is no prebuilt macOS binary).
322
332
  `routing.toml` carries `[targets.*]` and `[routes."sy/..."]` in Switchyard's own table
@@ -17,7 +17,7 @@ Issue: #25. Pi is a managed project peer for implementation, edits, tests, summa
17
17
  - `mlx/fast` resolves to the machine-shared loopback MLX server.
18
18
  - Apple Silicon setup pins Python 3.12, `mlx-lm==0.31.3`, and `Qwen/Qwen3-8B-MLX-4bit` revision `383413e909f3bc5303ce195ebbdf0339c5a1a2a3`.
19
19
  - The MLX input estimate is capped at 16,000 tokens; decode and prompt concurrency are one. DGX's default input estimate cap is 262,144 tokens.
20
- - The relay authenticates every call, rejects browser origins and unknown model aliases, and preserves streamed bytes. It records the requested route and model reported by the backend.
20
+ - The relay authenticates every call, rejects browser origins and unknown model aliases, and preserves streamed bytes. It records the requested route and model reported by the backend. Per request it also journals sanitized identity and lifecycle evidence (#139): resolved alias, upstream-configured model, observed provider and served model with their source (gateway header, generation SSE event, or local MLX configuration), outcome (completed, cancelled, failed), duration, and a confirmed-mismatch flag. A request cancelled before any identification stays explicitly unidentified; a backend's mutable last-served label is never a request's own evidence, and a primary/auxiliary role stays unknown without native evidence. Records carry no prompts, tools, keys or Access headers.
21
21
  - MLX may fall back to DGX before response streaming begins. Failed non-PII tasks can escalate to configured cloud peers with an explicit warning to reconcile partial effects.
22
22
  - Runtime ownership includes process start identity. Project shutdown releases its relay, while explicit `ahub models stop` controls the shared MLX process.
23
23