evalroute 0.6.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (51) hide show
  1. evalroute-0.6.0/.gitignore +9 -0
  2. evalroute-0.6.0/CHANGELOG.md +53 -0
  3. evalroute-0.6.0/PKG-INFO +351 -0
  4. evalroute-0.6.0/README.md +337 -0
  5. evalroute-0.6.0/evalroute/__init__.py +18 -0
  6. evalroute-0.6.0/evalroute/adjudicate.py +109 -0
  7. evalroute-0.6.0/evalroute/cli.py +140 -0
  8. evalroute-0.6.0/evalroute/contract.py +32 -0
  9. evalroute-0.6.0/evalroute/data/facets.yaml +56 -0
  10. evalroute-0.6.0/evalroute/data/routes.yaml +224 -0
  11. evalroute-0.6.0/evalroute/dataset.py +151 -0
  12. evalroute-0.6.0/evalroute/dispatch.py +163 -0
  13. evalroute-0.6.0/evalroute/flywheel.py +305 -0
  14. evalroute-0.6.0/evalroute/harness/__init__.py +1 -0
  15. evalroute-0.6.0/evalroute/harness/tier_a.py +501 -0
  16. evalroute-0.6.0/evalroute/paths.py +19 -0
  17. evalroute-0.6.0/evalroute/routes_from_labels.py +211 -0
  18. evalroute-0.6.0/evalroute/routes_from_report.py +316 -0
  19. evalroute-0.6.0/evalroute/routing.py +607 -0
  20. evalroute-0.6.0/evalroute/runners/__init__.py +1 -0
  21. evalroute-0.6.0/evalroute/runners/hermes_shim.py +157 -0
  22. evalroute-0.6.0/evalroute/schemas.py +36 -0
  23. evalroute-0.6.0/examples/artifacts/tier-a-alignment/report.csv +6 -0
  24. evalroute-0.6.0/examples/artifacts/tier-a-alignment/runs.jsonl +149 -0
  25. evalroute-0.6.0/examples/artifacts/tier-a-alignment/tasks.jsonl +10 -0
  26. evalroute-0.6.0/examples/artifacts/tier-a-dl-ml/report.csv +6 -0
  27. evalroute-0.6.0/examples/artifacts/tier-a-dl-ml/runs.jsonl +150 -0
  28. evalroute-0.6.0/examples/artifacts/tier-a-dl-ml/tasks.jsonl +10 -0
  29. evalroute-0.6.0/examples/artifacts/tier-a-models.json +7 -0
  30. evalroute-0.6.0/examples/artifacts/tier-a-rc-models.json +6 -0
  31. evalroute-0.6.0/examples/artifacts/tier-a-routine-coding/report.csv +5 -0
  32. evalroute-0.6.0/examples/artifacts/tier-a-routine-coding/runs.jsonl +120 -0
  33. evalroute-0.6.0/examples/artifacts/tier-a-routine-coding/tasks.jsonl +10 -0
  34. evalroute-0.6.0/pyproject.toml +22 -0
  35. evalroute-0.6.0/scripts/publish_dataset.py +166 -0
  36. evalroute-0.6.0/tests/conftest.py +28 -0
  37. evalroute-0.6.0/tests/test_adjudicate.py +79 -0
  38. evalroute-0.6.0/tests/test_artifacts.py +66 -0
  39. evalroute-0.6.0/tests/test_cli_flags.py +112 -0
  40. evalroute-0.6.0/tests/test_contract.py +28 -0
  41. evalroute-0.6.0/tests/test_dataset.py +298 -0
  42. evalroute-0.6.0/tests/test_dispatch.py +211 -0
  43. evalroute-0.6.0/tests/test_facets.py +122 -0
  44. evalroute-0.6.0/tests/test_flywheel.py +308 -0
  45. evalroute-0.6.0/tests/test_harness.py +209 -0
  46. evalroute-0.6.0/tests/test_isolation.py +56 -0
  47. evalroute-0.6.0/tests/test_rate_confirm.py +40 -0
  48. evalroute-0.6.0/tests/test_routes_from_report.py +228 -0
  49. evalroute-0.6.0/tests/test_shim.py +124 -0
  50. evalroute-0.6.0/tests/test_tools.py +300 -0
  51. evalroute-0.6.0/tests/test_workflow.py +113 -0
@@ -0,0 +1,9 @@
1
+ .venv/
2
+ __pycache__/
3
+ *.pyc
4
+ .DS_Store
5
+ .pytest_cache/
6
+ dist/
7
+ build/
8
+ *.egg-info/
9
+ data/flywheel/labels.jsonl
@@ -0,0 +1,53 @@
1
+ # Changelog
2
+
3
+ ## [0.6.0] - 2026-10-03
4
+
5
+ The routing core moved out of `hermes-plugin-evalroute` v0.5.1 into this
6
+ library. A move, not a rewrite: no behaviour change. The Hermes plugin
7
+ remains at `keppy/hermes-plugin-evalroute` and becomes a thin adapter over
8
+ the nine contract names in `evalroute/contract.py`.
9
+
10
+ Moved to the library:
11
+
12
+ - `tools.py` → split mechanically along existing function boundaries into
13
+ `evalroute/routing.py` (table load/validate, classifier, facets, route
14
+ card, `_tool_result`, `_note_route`, `install_routes`) and `evalroute/cli.py`
15
+ (`setup_cli`, `evalroute_cli`, the workflow epilog).
16
+ - `flywheel.py`, `dispatch.py`, `dataset.py`, `schemas.py`, `adjudicate.py`,
17
+ `routes_from_report.py`, `routes_from_labels.py` (imports made relative;
18
+ the path-load + sibling-injection hack in `evalroute_cli` and its
19
+ `if "fw" not in globals()` guard deleted — `dispatch` is now
20
+ `from . import dispatch`).
21
+ - `harness/evalroute.py` → `evalroute/harness/tier_a.py`;
22
+ `runners/hermes-shim.py` → `evalroute/runners/hermes_shim.py`.
23
+ - `data/routes.yaml`, `data/facets.yaml` → `evalroute/data/` as package data
24
+ (resolved via `importlib.resources`).
25
+ - `examples/artifacts/` (repo root, not package data) and
26
+ `scripts/publish_dataset.py`.
27
+ - The tests, minus `test_registration.py` and `test_sniff.py`, which stay in
28
+ the plugin (registration wiring and the Hermes sniff hook are plugin
29
+ concepts). `sniff.py` and `__init__.py`'s `register(ctx)` stay in the
30
+ plugin too.
31
+
32
+ Adapted:
33
+
34
+ - The `hermes_constants.get_hermes_home` try/except fallback at four sites
35
+ is now one function, `evalroute.paths.hermes_home()`.
36
+ - `PLUGIN_DIR` anchoring is gone: package data comes from
37
+ `importlib.resources`, and `_plugin_version()` (which read `plugin.yaml`)
38
+ becomes the package version — `pyproject.toml` via
39
+ `importlib.metadata.version("evalroute")`, literal fallback when not
40
+ installed. The card's provenance line reads `table: bundled
41
+ (evalroute 0.6.0)`.
42
+ - `evalroute.contract` is new: `CONTRACT_VERSION = 1` and the nine names the
43
+ plugin relies on (`routing.set_llm_facade`, `routing.evalroute_route`,
44
+ `routing.handle_route_command`, `cli.setup_cli`, `cli.evalroute_cli`,
45
+ `flywheel.handle_rate`, `flywheel.on_pre_command`,
46
+ `flywheel.on_post_llm_call`, `schemas.EVALROUTE_ROUTE`).
47
+ - A standalone `evalroute` console script wraps the same argparse tree, so
48
+ `evalroute route|dispatch|sync|rate ...` matches `hermes evalroute ...`
49
+ byte-for-byte.
50
+
51
+ Kept in the plugin: `__init__.py` (`register(ctx)`), `sniff.py`, `skills/`,
52
+ `plugin.yaml`, `catalog/`, `tests/test_registration.py`,
53
+ `tests/test_sniff.py`.
@@ -0,0 +1,351 @@
1
+ Metadata-Version: 2.5
2
+ Name: evalroute
3
+ Version: 0.6.0
4
+ Summary: Model routing with a verified-success flywheel: classify tasks, measure arms, feed results back
5
+ License: MIT
6
+ Requires-Python: >=3.11
7
+ Requires-Dist: pyyaml<7,>=6
8
+ Provides-Extra: dev
9
+ Requires-Dist: gonogo-eval<0.4,>=0.3; extra == 'dev'
10
+ Requires-Dist: pytest>=7; extra == 'dev'
11
+ Provides-Extra: hub
12
+ Requires-Dist: huggingface-hub>=0.20; extra == 'hub'
13
+ Description-Content-Type: text/markdown
14
+
15
+ # evalroute
16
+
17
+ Route a task to the right **(model, reasoning effort) arm** before you start.
18
+ evalroute is a Python package built around the evalroute procedure — a harness
19
+ that measures **cost per verified success** per task lane (`evalroute/harness/
20
+ tier_a.py`: measure, then serve the results) — and turns that output into a
21
+ route table with provenance on every row: pick the lane, pick the model, pick
22
+ the effort, as data.
23
+
24
+ The Hermes plugin is a thin adapter over this package, at
25
+ `keppy/hermes-plugin-evalroute` (`/route`, the `evalroute_route` tool, the
26
+ first-turn sniff, and the bundled skill live there; the routing core lives
27
+ here).
28
+
29
+ ## Install
30
+
31
+ ```bash
32
+ pip install evalroute # runtime: pyyaml only
33
+ pip install "evalroute[hub]" # + huggingface_hub, for `sync`
34
+ ```
35
+
36
+ ## Standalone CLI
37
+
38
+ The same argparse tree the plugin registers, standalone and byte-identical in
39
+ output to `hermes evalroute ...`:
40
+
41
+ ```bash
42
+ evalroute route --lane routine-coding "fix the failing test" # route card
43
+ evalroute route --json "read this 80-page spec and summarize" # JSON envelope
44
+ evalroute dispatch brief.md # route a brief, spawn hermes chat on that arm,
45
+ # print the rate line (--dry-run prints the argv)
46
+ evalroute sync --status # which route table is active
47
+ evalroute rate pass --note "why" # label the last routed task
48
+ evalroute install-routes --dry-run # route table -> agent.reasoning_overrides
49
+ ```
50
+
51
+ Works with no Hermes installed: the Hermes home falls back to `HERMES_HOME`
52
+ or `~/.hermes`, and the ledger, dataset pins, and effort-override reads all
53
+ honor it.
54
+
55
+ ## The nine-name contract
56
+
57
+ The plugin adapter relies on exactly these names; everything else in the
58
+ package is private to the library (see `evalroute/contract.py`,
59
+ `CONTRACT_VERSION = 1`):
60
+
61
+ | contract name | module | what it is |
62
+ | --- | --- | --- |
63
+ | `routing.set_llm_facade` | `evalroute/routing.py` | stash the host LLM facade; `None` disables the fallback |
64
+ | `routing.evalroute_route` | `evalroute/routing.py` | tool handler (`evalroute_route`) |
65
+ | `routing.handle_route_command` | `evalroute/routing.py` | `/route` slash command |
66
+ | `cli.setup_cli` | `evalroute/cli.py` | argparse wiring (`register_cli_command` setup_fn) |
67
+ | `cli.evalroute_cli` | `evalroute/cli.py` | CLI handler |
68
+ | `flywheel.handle_rate` | `evalroute/flywheel.py` | `/rate` pass\|fail\|skip |
69
+ | `flywheel.on_pre_command` | `evalroute/flywheel.py` | `/model` + `/reasoning` observer |
70
+ | `flywheel.on_post_llm_call` | `evalroute/flywheel.py` | last-seen-model diagnostic |
71
+ | `schemas.EVALROUTE_ROUTE` | `evalroute/schemas.py` | tool schema |
72
+
73
+ ## The route table
74
+
75
+ `evalroute/data/routes.yaml` — one row per lane: `id`, `keywords` (the rule
76
+ layer), `model`, `effort`, `escalation`, `provenance`, `notes`. Lane taxonomy
77
+ is the union of the two source tables (9 lanes); where they disagreed
78
+ (long-doc merged into web-research in one, orchestration only in the other)
79
+ both are kept as distinct lanes. Model ids must match `/model` spelling
80
+ exactly.
81
+
82
+ The table began with a **2026-09-26 priors snapshot** (public benchmarks,
83
+ many vendor-run). Three lanes now have small measured batches; the others
84
+ remain marked `priors`. Replace those rows only after your own controlled
85
+ data, and keep each row's `provenance` visible.
86
+
87
+ ### Route table: bundled or synced
88
+
89
+ The table you route against is either the bundled one or a pinned dataset
90
+ revision; the card says which. The dataset is fetched only when you run
91
+ `sync` (never on install, never while routing), it is pinned to a resolved
92
+ revision, and the library routes fully offline without it:
93
+
94
+ ```bash
95
+ evalroute sync --revision <sha> # pin the published table (default: main)
96
+ evalroute sync --status # bundled, or dataset @ <sha>
97
+ evalroute sync --clear # back to the bundled table
98
+ ```
99
+
100
+ `sync` downloads only the `routes/` config of
101
+ [keppy/evalroute-flywheel](https://huggingface.co/datasets/keppy/evalroute-flywheel)
102
+ into `<hermes home>/evalroute/dataset/<sha>/` — never the measured evidence
103
+ (grows over time; leave it on the Hub). It needs `pip install
104
+ huggingface_hub` (the `hub` extra). Routing data lands only under the Hermes
105
+ home, like the ledger.
106
+
107
+ ### Regenerating from measured data
108
+
109
+ ```bash
110
+ # in the harness venv (openai + anthropic; it stays out of the runtime venv)
111
+ python -m evalroute.harness.tier_a run -m models.json -t tasks.jsonl -k 3
112
+ python -m evalroute.harness.tier_a report -o runs.jsonl -t tasks.jsonl -k 3 --csv report.csv
113
+ python -m evalroute.routes_from_report --csv report.csv --runs runs.jsonl \
114
+ --models models.json --k 3 --out routes.generated.yaml
115
+ ```
116
+
117
+ `report` needs the matching `--tasks` file: it checks each run's prompt/checker
118
+ contract and treats missing declared tasks as incomplete. Without that file it
119
+ prints diagnostics but selects no route. The `report.csv` files under
120
+ `examples/artifacts/` were regenerated from their `runs.jsonl` and
121
+ `tasks.jsonl` with this harness and carry the `complete` column. A legacy CSV
122
+ (no `complete` column) is refused even with `--runs`, since the old winner
123
+ selection may have ignored pending cells; regenerate it the same way.
124
+
125
+ The harness is packaged so the loop is complete inside one repo: write
126
+ tasksets (deterministic `python` checkers where possible — validate every
127
+ checker against a reference solution before paid runs), run k samples per
128
+ arm, report, then flip the lane's row. `routes_from_report` applies the
129
+ report's own routing rule (coverage-gated lowest all-in $/success), stamps
130
+ `provenance: measured ...` with a gonogo McNemar stamp when the winner and
131
+ runner-up shared cases, preserves each lane's `keywords`/`match_hint`/
132
+ `escalation`/`notes`, and carries unmeasured lanes over verbatim —
133
+ regeneration never silently deletes a route. The tier-a tasksets and runs
134
+ that produced the current measured lanes are under `examples/artifacts/`
135
+ (sets: `tasks.jsonl`; raw run records: `runs.jsonl`).
136
+
137
+ Inspect the generated YAML before replacing the bundled table. `--models`
138
+ maps harness arm names such as `glm-5.3@high` to `/model` IDs; omission is
139
+ only safe if the CSV already contains routable IDs. `--k` defaults to 3;
140
+ graded sample counts must divide evenly by k, or the generator refuses to
141
+ invent a task count. The harness v2 resume key includes the full task and
142
+ checker spec, model price/config and effective max tokens, plus the judge
143
+ configuration when used. It keeps legacy JSONL reportable, but reporting
144
+ mixed legacy/v2 or multiple prompt/checker versions together fails explicitly.
145
+
146
+ ### The Hermes shim (`api: "cmd"` arms)
147
+
148
+ To run Hermes itself as an evalroute arm (agentic cells):
149
+
150
+ ```json
151
+ {"name": "hermes-glm@high", "api": "cmd", "effort": "high",
152
+ "cmd": "python -m evalroute.runners.herbes_shim --model {model} --effort {effort} --prompt {prompt_file}",
153
+ "model": "z-ai/glm-5.3", "in": 0.91, "out": 2.86, "timeout": 3600}
154
+ ```
155
+
156
+ The shim wraps `hermes -z` (one-shot; tools, memory, AGENTS.md loaded as
157
+ normal; approvals auto-bypassed), reads the usage report (`--usage-file`),
158
+ and emits the contract evalroute expects: the answer on stdout, then one
159
+ JSON last line `{"text": ..., "usage": {"inp", "out", "cache_read"}}`. A
160
+ non-zero hermes exit becomes an `error` field in that line (the harness
161
+ records an error row and retries on the next `run`); the shim itself always
162
+ exits 0. `--system` is prepended to the prompt. `--hermes PATH` overrides
163
+ the executable (tests use this to point at a fake — no real runs).
164
+
165
+ ## Classification: rules first, LLM when weak
166
+
167
+ The classifier's first layer is deterministic keyword rules over
168
+ `evalroute/data/routes.yaml` — free, no API calls — but rules alone misroute
169
+ paraphrase ("manage life, writing, and researchy tasks" has zero keyword
170
+ signal) and negation ("not usually hard math though" used to count as a
171
+ math hit; the rules now guard negated keywords). So routing is two-layer:
172
+
173
+ 1. **Strong rules** (2+ distinct keyword hits on the winning lane) — trusted
174
+ outright, no LLM call.
175
+ 2. **Weak signal** (0-1 hits) — one structured call via the facade set with
176
+ `routing.set_llm_facade` (the host's own model and auth; the plugin sets
177
+ it at register time, but **it consumes tokens and may incur provider
178
+ charges**). `None` (the default, and what a bare `evalroute` CLI sees)
179
+ disables the fallback. The LLM judges what the work IS — a description of
180
+ an assistant's duties routes to `orchestration`, not to whatever nouns
181
+ appear.
182
+
183
+ The card always prints which layer decided: `rules match`, `LLM fallback`,
184
+ or `no keyword hit - defaulted`. If the LLM call fails (offline, no facade),
185
+ the weak rules result stands and the card says so. Pin manually with
186
+ `--lane <id>` when you know better.
187
+
188
+ ## Facets: labels with dimensions
189
+
190
+ A lane is the routing decision; facets are the label. Every route also
191
+ captures the task's shape along three axes, defined in
192
+ `evalroute/data/facets.yaml`:
193
+
194
+ - **input-shape**: `long-doc` | `interactive`
195
+ - **domain**: `domain-dlml` | `domain-alignment` | `domain-math` | `domain-prose` | `domain-research`
196
+ - **demand-tier**: `tier-routine` | `tier-hard` | `tier-orchestration`
197
+
198
+ So "audit my RL training plan files" is recorded as `long-doc +
199
+ domain-dlml + tier-hard` — three facts about one task — instead of one
200
+ collapsed lane. The rules layer derives facets from keyword evidence
201
+ (conservative: only lanes that drew hits claim facets); the LLM fallback
202
+ names them semantically in the same structured call.
203
+
204
+ When a task claims facets on multiple axes, the card describes the
205
+ conjunction. **Facets do not alter the chosen arm**: the lane classifier
206
+ selects the arm; no domain/tier precedence is implemented:
207
+
208
+ ```
209
+ facets: long-doc + domain-dlml (descriptive conjunction; lane chooses arm)
210
+ ```
211
+
212
+ Facet conjunctions aggregate in `routes_from_labels`, so the
213
+ high-dimensional nodes — "how do long-doc x dl-ml tasks fare on arm X?" —
214
+ fill in from daily use without controlled-batch spend.
215
+
216
+ ## Flywheel: labels from daily workflow
217
+
218
+ The controlled harness is not the only source of data. As you route in daily
219
+ sessions, the library quietly builds an observational dataset:
220
+
221
+ - **`route`** logs the assignment (lane, recommended arm, method,
222
+ confidence, facets) to `<hermes home>/evalroute/labels.jsonl` — the
223
+ task text you typed is the label.
224
+ - **`/model` or `/reasoning` after a route** logs a process-global switch
225
+ observation. Without a session join it is not a verified route rejection
226
+ and cannot supply the actual arm for a flip.
227
+ - **`rate pass|fail [--route-id <id>] [--lane <id>] [--model <id> --effort <level>] [--note ...]`**
228
+ labels the outcome when you finish. Use both arm flags to self-report what
229
+ actually ran; otherwise the arm stays unknown. `--lane` corrects a lane;
230
+ `skip` discards that route. Prefer `--route-id` in overlapping sessions.
231
+ - Nothing else is recorded: no response bodies or turn telemetry. Task text,
232
+ switch arguments and optional notes are recorded. The last-seen model is
233
+ process-global memory only and is **not** assigned to a route as fact.
234
+
235
+ The ledger therefore holds two row qualities: an **observed arm** (a
236
+ `/model` or `/reasoning` switch after a route — a process-global candidate,
237
+ not proof) and a **caller-stated arm** (`evalroute dispatch <brief.md>`
238
+ records the spawn arguments as `arm_attribution: explicit_user` on the
239
+ outcome row). `dispatch` is the one-line form of the flywheel: route the
240
+ brief, spawn `hermes chat` on exactly that arm, print the `rate it:` line —
241
+ it still never auto-rates `pass`.
242
+
243
+ The ledger is **profile-wide**, not session-scoped: command hooks do not
244
+ supply a reliable session ID for route and rate. The card prints a route ID;
245
+ when tasks overlap, select it with `--route-id`. Without it, `rate` consumes
246
+ the latest pending route in that profile, which may be another session's.
247
+ `/model` and `/reasoning` observations are process-global candidates, not
248
+ proof of which model served a given route; only an explicit `rate --model
249
+ ... --effort ...` confirms an observational arm. To replace a pinned card,
250
+ use `route --lane <lane> --replace-route-id <old-id> <same task>`; identical
251
+ task text alone never consumes another pending route. Inspect the route ID,
252
+ task and lane before trusting a label.
253
+
254
+ **Continuing across sessions (turn caps).** A task that outlives its session —
255
+ the turn limit hits, the terminal closes mid-task — is a continuation, not a
256
+ new route. The row being labeled is (task, arm), not (task, session):
257
+
258
+ - Prefer staying in the session: `continue: <what remains>` gets a fresh
259
+ iteration budget, and the arm is session-scoped and persists.
260
+ - Otherwise `hermes -c` continues the same conversation, or paste the capped
261
+ session's final turn as the new session's opener.
262
+ - Never re-`route` the continuation. If the new session is a different task,
263
+ `rate skip --route-id <id>` clears that specific pending row first if it
264
+ belongs to you; don't consume another user's pending row.
265
+ - Steer as much as you like. The observed layer is defined as
266
+ daily-workflow-with-a-human-in-the-loop; a directive continuation is normal
267
+ operation, and it matches a detailed original prompt better than a bare
268
+ "continue" — which quietly tests prompt-luck instead of the arm. Measured
269
+ rows are untouched: they come from fixed-prompt, fresh-context harness cells.
270
+ - Put the methodology in the note: `--note "completed across two sessions
271
+ (turn cap), directive continuation"`. The label records neither cost nor
272
+ session boundaries, and a session-spanning pass re-reads the accumulated
273
+ context at full input price — the note is where that lives.
274
+
275
+ **Turning labels into route data:** `python -m evalroute.routes_from_labels`
276
+ prints per-lane, per-arm pass rates, lane corrections, and facet conjunction
277
+ outcomes; `--apply` writes `routes.observed.yaml`. Observed rows carry
278
+ honest, weaker provenance:
279
+
280
+ ```
281
+ observed 23 outcomes, same-maintainer observational single-arm, pass 78%, 2026-10-30;
282
+ not independent trials or a controlled comparison
283
+ ```
284
+
285
+ and only where the lane has no `measured` row — observational data can contest
286
+ a priors row, never overwrite a measured one. When gonogo is installed, each
287
+ observed row also carries its decide() verdict, so a row with 3 outcomes reads
288
+ as INSUFFICIENT_EVIDENCE rather than a pass rate someone will trust. The one
289
+ provisional flip threshold is 3 user-confirmed arm failures on the recommended
290
+ arm and 2 user-confirmed wins on another observed arm. This is still
291
+ same-maintainer, non-randomized evidence — review it and verify with paired
292
+ controlled cases before treating it as a quality comparison. The two
293
+ same-arm outcomes in the old snapshot never justify a route flip.
294
+
295
+ **Publishing the labels.** The live ledger stays local and append-only; do
296
+ not check raw labels into a public repo. Task text and notes can expose
297
+ paths and private project details even without response bodies. Do not claim
298
+ retroactive erasure for any label data once shared.
299
+
300
+ ## What the harness measures
301
+
302
+ Three lanes measured with the evalroute harness via the Nous inference API —
303
+ 10 tasks x 3 samples per arm (four routine-coding, five DL/ML, five
304
+ alignment arms, of which only four alignment arms have graded samples: 119
305
+ graded, 30 judge-pending, and one missing cell). Raw API and judge cost in
306
+ the vendored rows totals **$3.0590051**, excluding verification time across
307
+ routine coding, DL/ML research engineering, and alignment reasoning. Every
308
+ winner was statistically indistinguishable from its runner-up at n=10
309
+ (McNemar, via gonogo) — the empirical paired gap is zero, but the
310
+ conservative interval spans [-33.4%, +33.4%]; $p=1$ is not a population
311
+ equivalence test. All-coverage lanes have a ceiling on this taskset; the
312
+ routes choose cost among observed ties, not quality parity. The alignment
313
+ judge has no blind checker audit; its 30 pending Qwen outputs are excluded
314
+ from the route comparison. No quality claim spans that arm. The historical
315
+ v1 run records omit model IDs; the arm-name-to-ID mapping is the vendored
316
+ `models.json`, not an ID echoed by those records. Routine coding overturned
317
+ the priors' vendor pick: glm-5.3-flash at medium effort covered every task
318
+ at $0.00005/success, 2.4–3.9x cheaper than the V4.1 Flash arms at equal
319
+ coverage.
320
+
321
+ Every row is `priors`, `observed`, or `measured`. Nothing hypothesis-shaped
322
+ masquerades as a result — the card prints the row's provenance verbatim,
323
+ statistical stamp included.
324
+
325
+ ## Known limitations
326
+
327
+ - **The rule layer is keywords.** Deterministic, free, and misses
328
+ paraphrase — which is what the LLM fallback is for. The fallback's
329
+ quality tracks whatever model the host is on; it costs one small
330
+ structured call (temp 0, 256 tokens) only when rules are weak.
331
+ - **The library cannot switch the model for you.** `route` prints the card;
332
+ you run `/model <id>`. Run it before turn 1 — mid-session switches re-read
333
+ the whole context at full input price.
334
+ - **One effort slot per model id** (`agent.reasoning_overrides`); when a
335
+ model serves two lanes, `install-routes` keeps the higher effort. A
336
+ lower-effort lane's card explicitly prints `/reasoning <lane-effort>`
337
+ after `/model`, because the installed override alone would run the wrong
338
+ arm.
339
+ - **Provenance is priors until you measure.** Several rows are explicitly
340
+ untested/unmeasured/contested; the card prints the provenance verbatim so
341
+ nobody mistakes a hypothesis for a result.
342
+ - **The harness stays out of the runtime venv.** Run it in its own venv with
343
+ `openai`/`anthropic`; `EVALROUTE_PYTHON` points the runners at that
344
+ interpreter. The package itself requires nothing beyond `pyyaml`
345
+ (`huggingface_hub` only for `sync`).
346
+ - **The sniff hook is advisory only** (plugin side): it speaks when the
347
+ classifier is confident and the session's model disagrees with the lane's
348
+ route; it never rewrites, blocks, or switches.
349
+ - **`install-routes` needs the Hermes config module to write.** Standalone,
350
+ writing `agent.reasoning_overrides` works inside the `hermes` process;
351
+ `--dry-run` works anywhere.