evalroute 0.6.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- evalroute-0.6.0/.gitignore +9 -0
- evalroute-0.6.0/CHANGELOG.md +53 -0
- evalroute-0.6.0/PKG-INFO +351 -0
- evalroute-0.6.0/README.md +337 -0
- evalroute-0.6.0/evalroute/__init__.py +18 -0
- evalroute-0.6.0/evalroute/adjudicate.py +109 -0
- evalroute-0.6.0/evalroute/cli.py +140 -0
- evalroute-0.6.0/evalroute/contract.py +32 -0
- evalroute-0.6.0/evalroute/data/facets.yaml +56 -0
- evalroute-0.6.0/evalroute/data/routes.yaml +224 -0
- evalroute-0.6.0/evalroute/dataset.py +151 -0
- evalroute-0.6.0/evalroute/dispatch.py +163 -0
- evalroute-0.6.0/evalroute/flywheel.py +305 -0
- evalroute-0.6.0/evalroute/harness/__init__.py +1 -0
- evalroute-0.6.0/evalroute/harness/tier_a.py +501 -0
- evalroute-0.6.0/evalroute/paths.py +19 -0
- evalroute-0.6.0/evalroute/routes_from_labels.py +211 -0
- evalroute-0.6.0/evalroute/routes_from_report.py +316 -0
- evalroute-0.6.0/evalroute/routing.py +607 -0
- evalroute-0.6.0/evalroute/runners/__init__.py +1 -0
- evalroute-0.6.0/evalroute/runners/hermes_shim.py +157 -0
- evalroute-0.6.0/evalroute/schemas.py +36 -0
- evalroute-0.6.0/examples/artifacts/tier-a-alignment/report.csv +6 -0
- evalroute-0.6.0/examples/artifacts/tier-a-alignment/runs.jsonl +149 -0
- evalroute-0.6.0/examples/artifacts/tier-a-alignment/tasks.jsonl +10 -0
- evalroute-0.6.0/examples/artifacts/tier-a-dl-ml/report.csv +6 -0
- evalroute-0.6.0/examples/artifacts/tier-a-dl-ml/runs.jsonl +150 -0
- evalroute-0.6.0/examples/artifacts/tier-a-dl-ml/tasks.jsonl +10 -0
- evalroute-0.6.0/examples/artifacts/tier-a-models.json +7 -0
- evalroute-0.6.0/examples/artifacts/tier-a-rc-models.json +6 -0
- evalroute-0.6.0/examples/artifacts/tier-a-routine-coding/report.csv +5 -0
- evalroute-0.6.0/examples/artifacts/tier-a-routine-coding/runs.jsonl +120 -0
- evalroute-0.6.0/examples/artifacts/tier-a-routine-coding/tasks.jsonl +10 -0
- evalroute-0.6.0/pyproject.toml +22 -0
- evalroute-0.6.0/scripts/publish_dataset.py +166 -0
- evalroute-0.6.0/tests/conftest.py +28 -0
- evalroute-0.6.0/tests/test_adjudicate.py +79 -0
- evalroute-0.6.0/tests/test_artifacts.py +66 -0
- evalroute-0.6.0/tests/test_cli_flags.py +112 -0
- evalroute-0.6.0/tests/test_contract.py +28 -0
- evalroute-0.6.0/tests/test_dataset.py +298 -0
- evalroute-0.6.0/tests/test_dispatch.py +211 -0
- evalroute-0.6.0/tests/test_facets.py +122 -0
- evalroute-0.6.0/tests/test_flywheel.py +308 -0
- evalroute-0.6.0/tests/test_harness.py +209 -0
- evalroute-0.6.0/tests/test_isolation.py +56 -0
- evalroute-0.6.0/tests/test_rate_confirm.py +40 -0
- evalroute-0.6.0/tests/test_routes_from_report.py +228 -0
- evalroute-0.6.0/tests/test_shim.py +124 -0
- evalroute-0.6.0/tests/test_tools.py +300 -0
- evalroute-0.6.0/tests/test_workflow.py +113 -0
|
@@ -0,0 +1,53 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
## [0.6.0] - 2026-10-03
|
|
4
|
+
|
|
5
|
+
The routing core moved out of `hermes-plugin-evalroute` v0.5.1 into this
|
|
6
|
+
library. A move, not a rewrite: no behaviour change. The Hermes plugin
|
|
7
|
+
remains at `keppy/hermes-plugin-evalroute` and becomes a thin adapter over
|
|
8
|
+
the nine contract names in `evalroute/contract.py`.
|
|
9
|
+
|
|
10
|
+
Moved to the library:
|
|
11
|
+
|
|
12
|
+
- `tools.py` → split mechanically along existing function boundaries into
|
|
13
|
+
`evalroute/routing.py` (table load/validate, classifier, facets, route
|
|
14
|
+
card, `_tool_result`, `_note_route`, `install_routes`) and `evalroute/cli.py`
|
|
15
|
+
(`setup_cli`, `evalroute_cli`, the workflow epilog).
|
|
16
|
+
- `flywheel.py`, `dispatch.py`, `dataset.py`, `schemas.py`, `adjudicate.py`,
|
|
17
|
+
`routes_from_report.py`, `routes_from_labels.py` (imports made relative;
|
|
18
|
+
the path-load + sibling-injection hack in `evalroute_cli` and its
|
|
19
|
+
`if "fw" not in globals()` guard deleted — `dispatch` is now
|
|
20
|
+
`from . import dispatch`).
|
|
21
|
+
- `harness/evalroute.py` → `evalroute/harness/tier_a.py`;
|
|
22
|
+
`runners/hermes-shim.py` → `evalroute/runners/hermes_shim.py`.
|
|
23
|
+
- `data/routes.yaml`, `data/facets.yaml` → `evalroute/data/` as package data
|
|
24
|
+
(resolved via `importlib.resources`).
|
|
25
|
+
- `examples/artifacts/` (repo root, not package data) and
|
|
26
|
+
`scripts/publish_dataset.py`.
|
|
27
|
+
- The tests, minus `test_registration.py` and `test_sniff.py`, which stay in
|
|
28
|
+
the plugin (registration wiring and the Hermes sniff hook are plugin
|
|
29
|
+
concepts). `sniff.py` and `__init__.py`'s `register(ctx)` stay in the
|
|
30
|
+
plugin too.
|
|
31
|
+
|
|
32
|
+
Adapted:
|
|
33
|
+
|
|
34
|
+
- The `hermes_constants.get_hermes_home` try/except fallback at four sites
|
|
35
|
+
is now one function, `evalroute.paths.hermes_home()`.
|
|
36
|
+
- `PLUGIN_DIR` anchoring is gone: package data comes from
|
|
37
|
+
`importlib.resources`, and `_plugin_version()` (which read `plugin.yaml`)
|
|
38
|
+
becomes the package version — `pyproject.toml` via
|
|
39
|
+
`importlib.metadata.version("evalroute")`, literal fallback when not
|
|
40
|
+
installed. The card's provenance line reads `table: bundled
|
|
41
|
+
(evalroute 0.6.0)`.
|
|
42
|
+
- `evalroute.contract` is new: `CONTRACT_VERSION = 1` and the nine names the
|
|
43
|
+
plugin relies on (`routing.set_llm_facade`, `routing.evalroute_route`,
|
|
44
|
+
`routing.handle_route_command`, `cli.setup_cli`, `cli.evalroute_cli`,
|
|
45
|
+
`flywheel.handle_rate`, `flywheel.on_pre_command`,
|
|
46
|
+
`flywheel.on_post_llm_call`, `schemas.EVALROUTE_ROUTE`).
|
|
47
|
+
- A standalone `evalroute` console script wraps the same argparse tree, so
|
|
48
|
+
`evalroute route|dispatch|sync|rate ...` matches `hermes evalroute ...`
|
|
49
|
+
byte-for-byte.
|
|
50
|
+
|
|
51
|
+
Kept in the plugin: `__init__.py` (`register(ctx)`), `sniff.py`, `skills/`,
|
|
52
|
+
`plugin.yaml`, `catalog/`, `tests/test_registration.py`,
|
|
53
|
+
`tests/test_sniff.py`.
|
evalroute-0.6.0/PKG-INFO
ADDED
|
@@ -0,0 +1,351 @@
|
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
|
+
Name: evalroute
|
|
3
|
+
Version: 0.6.0
|
|
4
|
+
Summary: Model routing with a verified-success flywheel: classify tasks, measure arms, feed results back
|
|
5
|
+
License: MIT
|
|
6
|
+
Requires-Python: >=3.11
|
|
7
|
+
Requires-Dist: pyyaml<7,>=6
|
|
8
|
+
Provides-Extra: dev
|
|
9
|
+
Requires-Dist: gonogo-eval<0.4,>=0.3; extra == 'dev'
|
|
10
|
+
Requires-Dist: pytest>=7; extra == 'dev'
|
|
11
|
+
Provides-Extra: hub
|
|
12
|
+
Requires-Dist: huggingface-hub>=0.20; extra == 'hub'
|
|
13
|
+
Description-Content-Type: text/markdown
|
|
14
|
+
|
|
15
|
+
# evalroute
|
|
16
|
+
|
|
17
|
+
Route a task to the right **(model, reasoning effort) arm** before you start.
|
|
18
|
+
evalroute is a Python package built around the evalroute procedure — a harness
|
|
19
|
+
that measures **cost per verified success** per task lane (`evalroute/harness/
|
|
20
|
+
tier_a.py`: measure, then serve the results) — and turns that output into a
|
|
21
|
+
route table with provenance on every row: pick the lane, pick the model, pick
|
|
22
|
+
the effort, as data.
|
|
23
|
+
|
|
24
|
+
The Hermes plugin is a thin adapter over this package, at
|
|
25
|
+
`keppy/hermes-plugin-evalroute` (`/route`, the `evalroute_route` tool, the
|
|
26
|
+
first-turn sniff, and the bundled skill live there; the routing core lives
|
|
27
|
+
here).
|
|
28
|
+
|
|
29
|
+
## Install
|
|
30
|
+
|
|
31
|
+
```bash
|
|
32
|
+
pip install evalroute # runtime: pyyaml only
|
|
33
|
+
pip install "evalroute[hub]" # + huggingface_hub, for `sync`
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
## Standalone CLI
|
|
37
|
+
|
|
38
|
+
The same argparse tree the plugin registers, standalone and byte-identical in
|
|
39
|
+
output to `hermes evalroute ...`:
|
|
40
|
+
|
|
41
|
+
```bash
|
|
42
|
+
evalroute route --lane routine-coding "fix the failing test" # route card
|
|
43
|
+
evalroute route --json "read this 80-page spec and summarize" # JSON envelope
|
|
44
|
+
evalroute dispatch brief.md # route a brief, spawn hermes chat on that arm,
|
|
45
|
+
# print the rate line (--dry-run prints the argv)
|
|
46
|
+
evalroute sync --status # which route table is active
|
|
47
|
+
evalroute rate pass --note "why" # label the last routed task
|
|
48
|
+
evalroute install-routes --dry-run # route table -> agent.reasoning_overrides
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
Works with no Hermes installed: the Hermes home falls back to `HERMES_HOME`
|
|
52
|
+
or `~/.hermes`, and the ledger, dataset pins, and effort-override reads all
|
|
53
|
+
honor it.
|
|
54
|
+
|
|
55
|
+
## The nine-name contract
|
|
56
|
+
|
|
57
|
+
The plugin adapter relies on exactly these names; everything else in the
|
|
58
|
+
package is private to the library (see `evalroute/contract.py`,
|
|
59
|
+
`CONTRACT_VERSION = 1`):
|
|
60
|
+
|
|
61
|
+
| contract name | module | what it is |
|
|
62
|
+
| --- | --- | --- |
|
|
63
|
+
| `routing.set_llm_facade` | `evalroute/routing.py` | stash the host LLM facade; `None` disables the fallback |
|
|
64
|
+
| `routing.evalroute_route` | `evalroute/routing.py` | tool handler (`evalroute_route`) |
|
|
65
|
+
| `routing.handle_route_command` | `evalroute/routing.py` | `/route` slash command |
|
|
66
|
+
| `cli.setup_cli` | `evalroute/cli.py` | argparse wiring (`register_cli_command` setup_fn) |
|
|
67
|
+
| `cli.evalroute_cli` | `evalroute/cli.py` | CLI handler |
|
|
68
|
+
| `flywheel.handle_rate` | `evalroute/flywheel.py` | `/rate` pass\|fail\|skip |
|
|
69
|
+
| `flywheel.on_pre_command` | `evalroute/flywheel.py` | `/model` + `/reasoning` observer |
|
|
70
|
+
| `flywheel.on_post_llm_call` | `evalroute/flywheel.py` | last-seen-model diagnostic |
|
|
71
|
+
| `schemas.EVALROUTE_ROUTE` | `evalroute/schemas.py` | tool schema |
|
|
72
|
+
|
|
73
|
+
## The route table
|
|
74
|
+
|
|
75
|
+
`evalroute/data/routes.yaml` — one row per lane: `id`, `keywords` (the rule
|
|
76
|
+
layer), `model`, `effort`, `escalation`, `provenance`, `notes`. Lane taxonomy
|
|
77
|
+
is the union of the two source tables (9 lanes); where they disagreed
|
|
78
|
+
(long-doc merged into web-research in one, orchestration only in the other)
|
|
79
|
+
both are kept as distinct lanes. Model ids must match `/model` spelling
|
|
80
|
+
exactly.
|
|
81
|
+
|
|
82
|
+
The table began with a **2026-09-26 priors snapshot** (public benchmarks,
|
|
83
|
+
many vendor-run). Three lanes now have small measured batches; the others
|
|
84
|
+
remain marked `priors`. Replace those rows only after your own controlled
|
|
85
|
+
data, and keep each row's `provenance` visible.
|
|
86
|
+
|
|
87
|
+
### Route table: bundled or synced
|
|
88
|
+
|
|
89
|
+
The table you route against is either the bundled one or a pinned dataset
|
|
90
|
+
revision; the card says which. The dataset is fetched only when you run
|
|
91
|
+
`sync` (never on install, never while routing), it is pinned to a resolved
|
|
92
|
+
revision, and the library routes fully offline without it:
|
|
93
|
+
|
|
94
|
+
```bash
|
|
95
|
+
evalroute sync --revision <sha> # pin the published table (default: main)
|
|
96
|
+
evalroute sync --status # bundled, or dataset @ <sha>
|
|
97
|
+
evalroute sync --clear # back to the bundled table
|
|
98
|
+
```
|
|
99
|
+
|
|
100
|
+
`sync` downloads only the `routes/` config of
|
|
101
|
+
[keppy/evalroute-flywheel](https://huggingface.co/datasets/keppy/evalroute-flywheel)
|
|
102
|
+
into `<hermes home>/evalroute/dataset/<sha>/` — never the measured evidence
|
|
103
|
+
(grows over time; leave it on the Hub). It needs `pip install
|
|
104
|
+
huggingface_hub` (the `hub` extra). Routing data lands only under the Hermes
|
|
105
|
+
home, like the ledger.
|
|
106
|
+
|
|
107
|
+
### Regenerating from measured data
|
|
108
|
+
|
|
109
|
+
```bash
|
|
110
|
+
# in the harness venv (openai + anthropic; it stays out of the runtime venv)
|
|
111
|
+
python -m evalroute.harness.tier_a run -m models.json -t tasks.jsonl -k 3
|
|
112
|
+
python -m evalroute.harness.tier_a report -o runs.jsonl -t tasks.jsonl -k 3 --csv report.csv
|
|
113
|
+
python -m evalroute.routes_from_report --csv report.csv --runs runs.jsonl \
|
|
114
|
+
--models models.json --k 3 --out routes.generated.yaml
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
`report` needs the matching `--tasks` file: it checks each run's prompt/checker
|
|
118
|
+
contract and treats missing declared tasks as incomplete. Without that file it
|
|
119
|
+
prints diagnostics but selects no route. The `report.csv` files under
|
|
120
|
+
`examples/artifacts/` were regenerated from their `runs.jsonl` and
|
|
121
|
+
`tasks.jsonl` with this harness and carry the `complete` column. A legacy CSV
|
|
122
|
+
(no `complete` column) is refused even with `--runs`, since the old winner
|
|
123
|
+
selection may have ignored pending cells; regenerate it the same way.
|
|
124
|
+
|
|
125
|
+
The harness is packaged so the loop is complete inside one repo: write
|
|
126
|
+
tasksets (deterministic `python` checkers where possible — validate every
|
|
127
|
+
checker against a reference solution before paid runs), run k samples per
|
|
128
|
+
arm, report, then flip the lane's row. `routes_from_report` applies the
|
|
129
|
+
report's own routing rule (coverage-gated lowest all-in $/success), stamps
|
|
130
|
+
`provenance: measured ...` with a gonogo McNemar stamp when the winner and
|
|
131
|
+
runner-up shared cases, preserves each lane's `keywords`/`match_hint`/
|
|
132
|
+
`escalation`/`notes`, and carries unmeasured lanes over verbatim —
|
|
133
|
+
regeneration never silently deletes a route. The tier-a tasksets and runs
|
|
134
|
+
that produced the current measured lanes are under `examples/artifacts/`
|
|
135
|
+
(sets: `tasks.jsonl`; raw run records: `runs.jsonl`).
|
|
136
|
+
|
|
137
|
+
Inspect the generated YAML before replacing the bundled table. `--models`
|
|
138
|
+
maps harness arm names such as `glm-5.3@high` to `/model` IDs; omission is
|
|
139
|
+
only safe if the CSV already contains routable IDs. `--k` defaults to 3;
|
|
140
|
+
graded sample counts must divide evenly by k, or the generator refuses to
|
|
141
|
+
invent a task count. The harness v2 resume key includes the full task and
|
|
142
|
+
checker spec, model price/config and effective max tokens, plus the judge
|
|
143
|
+
configuration when used. It keeps legacy JSONL reportable, but reporting
|
|
144
|
+
mixed legacy/v2 or multiple prompt/checker versions together fails explicitly.
|
|
145
|
+
|
|
146
|
+
### The Hermes shim (`api: "cmd"` arms)
|
|
147
|
+
|
|
148
|
+
To run Hermes itself as an evalroute arm (agentic cells):
|
|
149
|
+
|
|
150
|
+
```json
|
|
151
|
+
{"name": "hermes-glm@high", "api": "cmd", "effort": "high",
|
|
152
|
+
"cmd": "python -m evalroute.runners.herbes_shim --model {model} --effort {effort} --prompt {prompt_file}",
|
|
153
|
+
"model": "z-ai/glm-5.3", "in": 0.91, "out": 2.86, "timeout": 3600}
|
|
154
|
+
```
|
|
155
|
+
|
|
156
|
+
The shim wraps `hermes -z` (one-shot; tools, memory, AGENTS.md loaded as
|
|
157
|
+
normal; approvals auto-bypassed), reads the usage report (`--usage-file`),
|
|
158
|
+
and emits the contract evalroute expects: the answer on stdout, then one
|
|
159
|
+
JSON last line `{"text": ..., "usage": {"inp", "out", "cache_read"}}`. A
|
|
160
|
+
non-zero hermes exit becomes an `error` field in that line (the harness
|
|
161
|
+
records an error row and retries on the next `run`); the shim itself always
|
|
162
|
+
exits 0. `--system` is prepended to the prompt. `--hermes PATH` overrides
|
|
163
|
+
the executable (tests use this to point at a fake — no real runs).
|
|
164
|
+
|
|
165
|
+
## Classification: rules first, LLM when weak
|
|
166
|
+
|
|
167
|
+
The classifier's first layer is deterministic keyword rules over
|
|
168
|
+
`evalroute/data/routes.yaml` — free, no API calls — but rules alone misroute
|
|
169
|
+
paraphrase ("manage life, writing, and researchy tasks" has zero keyword
|
|
170
|
+
signal) and negation ("not usually hard math though" used to count as a
|
|
171
|
+
math hit; the rules now guard negated keywords). So routing is two-layer:
|
|
172
|
+
|
|
173
|
+
1. **Strong rules** (2+ distinct keyword hits on the winning lane) — trusted
|
|
174
|
+
outright, no LLM call.
|
|
175
|
+
2. **Weak signal** (0-1 hits) — one structured call via the facade set with
|
|
176
|
+
`routing.set_llm_facade` (the host's own model and auth; the plugin sets
|
|
177
|
+
it at register time, but **it consumes tokens and may incur provider
|
|
178
|
+
charges**). `None` (the default, and what a bare `evalroute` CLI sees)
|
|
179
|
+
disables the fallback. The LLM judges what the work IS — a description of
|
|
180
|
+
an assistant's duties routes to `orchestration`, not to whatever nouns
|
|
181
|
+
appear.
|
|
182
|
+
|
|
183
|
+
The card always prints which layer decided: `rules match`, `LLM fallback`,
|
|
184
|
+
or `no keyword hit - defaulted`. If the LLM call fails (offline, no facade),
|
|
185
|
+
the weak rules result stands and the card says so. Pin manually with
|
|
186
|
+
`--lane <id>` when you know better.
|
|
187
|
+
|
|
188
|
+
## Facets: labels with dimensions
|
|
189
|
+
|
|
190
|
+
A lane is the routing decision; facets are the label. Every route also
|
|
191
|
+
captures the task's shape along three axes, defined in
|
|
192
|
+
`evalroute/data/facets.yaml`:
|
|
193
|
+
|
|
194
|
+
- **input-shape**: `long-doc` | `interactive`
|
|
195
|
+
- **domain**: `domain-dlml` | `domain-alignment` | `domain-math` | `domain-prose` | `domain-research`
|
|
196
|
+
- **demand-tier**: `tier-routine` | `tier-hard` | `tier-orchestration`
|
|
197
|
+
|
|
198
|
+
So "audit my RL training plan files" is recorded as `long-doc +
|
|
199
|
+
domain-dlml + tier-hard` — three facts about one task — instead of one
|
|
200
|
+
collapsed lane. The rules layer derives facets from keyword evidence
|
|
201
|
+
(conservative: only lanes that drew hits claim facets); the LLM fallback
|
|
202
|
+
names them semantically in the same structured call.
|
|
203
|
+
|
|
204
|
+
When a task claims facets on multiple axes, the card describes the
|
|
205
|
+
conjunction. **Facets do not alter the chosen arm**: the lane classifier
|
|
206
|
+
selects the arm; no domain/tier precedence is implemented:
|
|
207
|
+
|
|
208
|
+
```
|
|
209
|
+
facets: long-doc + domain-dlml (descriptive conjunction; lane chooses arm)
|
|
210
|
+
```
|
|
211
|
+
|
|
212
|
+
Facet conjunctions aggregate in `routes_from_labels`, so the
|
|
213
|
+
high-dimensional nodes — "how do long-doc x dl-ml tasks fare on arm X?" —
|
|
214
|
+
fill in from daily use without controlled-batch spend.
|
|
215
|
+
|
|
216
|
+
## Flywheel: labels from daily workflow
|
|
217
|
+
|
|
218
|
+
The controlled harness is not the only source of data. As you route in daily
|
|
219
|
+
sessions, the library quietly builds an observational dataset:
|
|
220
|
+
|
|
221
|
+
- **`route`** logs the assignment (lane, recommended arm, method,
|
|
222
|
+
confidence, facets) to `<hermes home>/evalroute/labels.jsonl` — the
|
|
223
|
+
task text you typed is the label.
|
|
224
|
+
- **`/model` or `/reasoning` after a route** logs a process-global switch
|
|
225
|
+
observation. Without a session join it is not a verified route rejection
|
|
226
|
+
and cannot supply the actual arm for a flip.
|
|
227
|
+
- **`rate pass|fail [--route-id <id>] [--lane <id>] [--model <id> --effort <level>] [--note ...]`**
|
|
228
|
+
labels the outcome when you finish. Use both arm flags to self-report what
|
|
229
|
+
actually ran; otherwise the arm stays unknown. `--lane` corrects a lane;
|
|
230
|
+
`skip` discards that route. Prefer `--route-id` in overlapping sessions.
|
|
231
|
+
- Nothing else is recorded: no response bodies or turn telemetry. Task text,
|
|
232
|
+
switch arguments and optional notes are recorded. The last-seen model is
|
|
233
|
+
process-global memory only and is **not** assigned to a route as fact.
|
|
234
|
+
|
|
235
|
+
The ledger therefore holds two row qualities: an **observed arm** (a
|
|
236
|
+
`/model` or `/reasoning` switch after a route — a process-global candidate,
|
|
237
|
+
not proof) and a **caller-stated arm** (`evalroute dispatch <brief.md>`
|
|
238
|
+
records the spawn arguments as `arm_attribution: explicit_user` on the
|
|
239
|
+
outcome row). `dispatch` is the one-line form of the flywheel: route the
|
|
240
|
+
brief, spawn `hermes chat` on exactly that arm, print the `rate it:` line —
|
|
241
|
+
it still never auto-rates `pass`.
|
|
242
|
+
|
|
243
|
+
The ledger is **profile-wide**, not session-scoped: command hooks do not
|
|
244
|
+
supply a reliable session ID for route and rate. The card prints a route ID;
|
|
245
|
+
when tasks overlap, select it with `--route-id`. Without it, `rate` consumes
|
|
246
|
+
the latest pending route in that profile, which may be another session's.
|
|
247
|
+
`/model` and `/reasoning` observations are process-global candidates, not
|
|
248
|
+
proof of which model served a given route; only an explicit `rate --model
|
|
249
|
+
... --effort ...` confirms an observational arm. To replace a pinned card,
|
|
250
|
+
use `route --lane <lane> --replace-route-id <old-id> <same task>`; identical
|
|
251
|
+
task text alone never consumes another pending route. Inspect the route ID,
|
|
252
|
+
task and lane before trusting a label.
|
|
253
|
+
|
|
254
|
+
**Continuing across sessions (turn caps).** A task that outlives its session —
|
|
255
|
+
the turn limit hits, the terminal closes mid-task — is a continuation, not a
|
|
256
|
+
new route. The row being labeled is (task, arm), not (task, session):
|
|
257
|
+
|
|
258
|
+
- Prefer staying in the session: `continue: <what remains>` gets a fresh
|
|
259
|
+
iteration budget, and the arm is session-scoped and persists.
|
|
260
|
+
- Otherwise `hermes -c` continues the same conversation, or paste the capped
|
|
261
|
+
session's final turn as the new session's opener.
|
|
262
|
+
- Never re-`route` the continuation. If the new session is a different task,
|
|
263
|
+
`rate skip --route-id <id>` clears that specific pending row first if it
|
|
264
|
+
belongs to you; don't consume another user's pending row.
|
|
265
|
+
- Steer as much as you like. The observed layer is defined as
|
|
266
|
+
daily-workflow-with-a-human-in-the-loop; a directive continuation is normal
|
|
267
|
+
operation, and it matches a detailed original prompt better than a bare
|
|
268
|
+
"continue" — which quietly tests prompt-luck instead of the arm. Measured
|
|
269
|
+
rows are untouched: they come from fixed-prompt, fresh-context harness cells.
|
|
270
|
+
- Put the methodology in the note: `--note "completed across two sessions
|
|
271
|
+
(turn cap), directive continuation"`. The label records neither cost nor
|
|
272
|
+
session boundaries, and a session-spanning pass re-reads the accumulated
|
|
273
|
+
context at full input price — the note is where that lives.
|
|
274
|
+
|
|
275
|
+
**Turning labels into route data:** `python -m evalroute.routes_from_labels`
|
|
276
|
+
prints per-lane, per-arm pass rates, lane corrections, and facet conjunction
|
|
277
|
+
outcomes; `--apply` writes `routes.observed.yaml`. Observed rows carry
|
|
278
|
+
honest, weaker provenance:
|
|
279
|
+
|
|
280
|
+
```
|
|
281
|
+
observed 23 outcomes, same-maintainer observational single-arm, pass 78%, 2026-10-30;
|
|
282
|
+
not independent trials or a controlled comparison
|
|
283
|
+
```
|
|
284
|
+
|
|
285
|
+
and only where the lane has no `measured` row — observational data can contest
|
|
286
|
+
a priors row, never overwrite a measured one. When gonogo is installed, each
|
|
287
|
+
observed row also carries its decide() verdict, so a row with 3 outcomes reads
|
|
288
|
+
as INSUFFICIENT_EVIDENCE rather than a pass rate someone will trust. The one
|
|
289
|
+
provisional flip threshold is 3 user-confirmed arm failures on the recommended
|
|
290
|
+
arm and 2 user-confirmed wins on another observed arm. This is still
|
|
291
|
+
same-maintainer, non-randomized evidence — review it and verify with paired
|
|
292
|
+
controlled cases before treating it as a quality comparison. The two
|
|
293
|
+
same-arm outcomes in the old snapshot never justify a route flip.
|
|
294
|
+
|
|
295
|
+
**Publishing the labels.** The live ledger stays local and append-only; do
|
|
296
|
+
not check raw labels into a public repo. Task text and notes can expose
|
|
297
|
+
paths and private project details even without response bodies. Do not claim
|
|
298
|
+
retroactive erasure for any label data once shared.
|
|
299
|
+
|
|
300
|
+
## What the harness measures
|
|
301
|
+
|
|
302
|
+
Three lanes measured with the evalroute harness via the Nous inference API —
|
|
303
|
+
10 tasks x 3 samples per arm (four routine-coding, five DL/ML, five
|
|
304
|
+
alignment arms, of which only four alignment arms have graded samples: 119
|
|
305
|
+
graded, 30 judge-pending, and one missing cell). Raw API and judge cost in
|
|
306
|
+
the vendored rows totals **$3.0590051**, excluding verification time across
|
|
307
|
+
routine coding, DL/ML research engineering, and alignment reasoning. Every
|
|
308
|
+
winner was statistically indistinguishable from its runner-up at n=10
|
|
309
|
+
(McNemar, via gonogo) — the empirical paired gap is zero, but the
|
|
310
|
+
conservative interval spans [-33.4%, +33.4%]; $p=1$ is not a population
|
|
311
|
+
equivalence test. All-coverage lanes have a ceiling on this taskset; the
|
|
312
|
+
routes choose cost among observed ties, not quality parity. The alignment
|
|
313
|
+
judge has no blind checker audit; its 30 pending Qwen outputs are excluded
|
|
314
|
+
from the route comparison. No quality claim spans that arm. The historical
|
|
315
|
+
v1 run records omit model IDs; the arm-name-to-ID mapping is the vendored
|
|
316
|
+
`models.json`, not an ID echoed by those records. Routine coding overturned
|
|
317
|
+
the priors' vendor pick: glm-5.3-flash at medium effort covered every task
|
|
318
|
+
at $0.00005/success, 2.4–3.9x cheaper than the V4.1 Flash arms at equal
|
|
319
|
+
coverage.
|
|
320
|
+
|
|
321
|
+
Every row is `priors`, `observed`, or `measured`. Nothing hypothesis-shaped
|
|
322
|
+
masquerades as a result — the card prints the row's provenance verbatim,
|
|
323
|
+
statistical stamp included.
|
|
324
|
+
|
|
325
|
+
## Known limitations
|
|
326
|
+
|
|
327
|
+
- **The rule layer is keywords.** Deterministic, free, and misses
|
|
328
|
+
paraphrase — which is what the LLM fallback is for. The fallback's
|
|
329
|
+
quality tracks whatever model the host is on; it costs one small
|
|
330
|
+
structured call (temp 0, 256 tokens) only when rules are weak.
|
|
331
|
+
- **The library cannot switch the model for you.** `route` prints the card;
|
|
332
|
+
you run `/model <id>`. Run it before turn 1 — mid-session switches re-read
|
|
333
|
+
the whole context at full input price.
|
|
334
|
+
- **One effort slot per model id** (`agent.reasoning_overrides`); when a
|
|
335
|
+
model serves two lanes, `install-routes` keeps the higher effort. A
|
|
336
|
+
lower-effort lane's card explicitly prints `/reasoning <lane-effort>`
|
|
337
|
+
after `/model`, because the installed override alone would run the wrong
|
|
338
|
+
arm.
|
|
339
|
+
- **Provenance is priors until you measure.** Several rows are explicitly
|
|
340
|
+
untested/unmeasured/contested; the card prints the provenance verbatim so
|
|
341
|
+
nobody mistakes a hypothesis for a result.
|
|
342
|
+
- **The harness stays out of the runtime venv.** Run it in its own venv with
|
|
343
|
+
`openai`/`anthropic`; `EVALROUTE_PYTHON` points the runners at that
|
|
344
|
+
interpreter. The package itself requires nothing beyond `pyyaml`
|
|
345
|
+
(`huggingface_hub` only for `sync`).
|
|
346
|
+
- **The sniff hook is advisory only** (plugin side): it speaks when the
|
|
347
|
+
classifier is confident and the session's model disagrees with the lane's
|
|
348
|
+
route; it never rewrites, blocks, or switches.
|
|
349
|
+
- **`install-routes` needs the Hermes config module to write.** Standalone,
|
|
350
|
+
writing `agent.reasoning_overrides` works inside the `hermes` process;
|
|
351
|
+
`--dry-run` works anywhere.
|