openmerit 0.1.0 → 0.1.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -32,10 +32,10 @@ pi task A → saved pi trace → extension trial job → pi trials B, C → scor
32
32
  ([`src/frontier.ts`](src/frontier.ts), 3 objectives: quality ↑, price ↓,
33
33
  latency ↓). The aggregate across all of an agent's tasks is the union of
34
34
  its task frontiers (`openmerit frontier`).
35
- 4. **Model discovery** — the OpenRouter catalog is snapshotted and diffed;
36
- newly appeared models are queued for future release experiments. The
37
- per-task loop shortlists currently available pi/OpenRouter models with
38
- OpenRouter benchmark signals and the existing strategist
35
+ 4. **Model discovery** — the active Pi session's scoped models (or Pi's
36
+ authenticated available-model registry when unscoped) are the candidate
37
+ pool. OpenRouter is an optional route and enrichment source: its catalog can
38
+ be snapshotted and its benchmark signals can improve the shortlist
39
39
  ([`src/catalog.ts`](src/catalog.ts) + [`src/daemon.ts`](src/daemon.ts)).
40
40
  5. **Informing the main harness** — recommendations land in
41
41
  `~/.openmerit/recommendations.jsonl`. The pi extension
@@ -47,8 +47,9 @@ pi task A → saved pi trace → extension trial job → pi trials B, C → scor
47
47
 
48
48
  ## Install from npm
49
49
 
50
- You need Node 22.18+, pi 0.85.1+, and an OpenRouter API key with available
51
- credits. Install the package into pi, then initialize its policy and state:
50
+ You need Node 22.18+, pi 0.85.1+, and at least two eligible model routes
51
+ authenticated in Pi. OpenRouter is supported but not required. Install the
52
+ package into pi, then initialize its policy and state:
52
53
 
53
54
  ```bash
54
55
  pi install npm:openmerit
@@ -59,7 +60,7 @@ pi list
59
60
  `pi install` makes the extension and its trial engine available to pi. The
60
61
  one-off `npx` command creates `~/.openmerit/policy.json`; install OpenMerit
61
62
  globally with `npm install --global openmerit` only if you also want persistent
62
- shell access to `openmerit status`, `frontier`, or the optional watcher. Install
63
+ shell access to `openmerit status`, `doctor`, `frontier`, or the optional watcher. Install
63
64
  the extension from only one source—remove any older copied `openmerit.ts` first
64
65
  so pi does not load it twice.
65
66
 
@@ -85,47 +86,64 @@ Keep that checkout available because pi loads a local package from its path.
85
86
  policy**. The shipped policy uses `"mode": "recommend"` and does not change
86
87
  the active model without approval.
87
88
 
88
- 2. Give **Pi and the extension trial job** OpenRouter access. The job reads
89
- `OPENROUTER_API_KEY` from its environment or `~/.openmerit/.env`. If using
90
- the file, edit it to contain `OPENROUTER_API_KEY=your-key` and protect it:
89
+ 2. Authenticate the providers you want to compare in Pi, using `/login` or
90
+ Pi's normal environment/model configuration. OpenMerit takes its eligible
91
+ model routes, capabilities, prices, and credentials from Pi; judge and
92
+ strategist calls use the same routes and do not require duplicate keys.
93
+
94
+ For example, verify one provider without printing its credential:
91
95
 
92
96
  ```bash
93
- nano ~/.openmerit/.env
94
- chmod 600 ~/.openmerit/.env
97
+ pi auth check --provider openai --json
95
98
  ```
96
99
 
97
- Pi does **not** read OpenMerit's `.env` file. Log in to OpenRouter with
98
- `/login openrouter` inside pi, or export `OPENROUTER_API_KEY` in the shell
99
- that starts pi. Verify pi's side without printing the key:
100
+ Add an OpenRouter route the same way if you want its routed catalog:
100
101
 
101
102
  ```bash
102
103
  pi auth check --provider openrouter --json
103
104
  ```
104
105
 
105
- A ready configuration prints `{"status":"ready","provider":"openrouter",...}`.
106
- Older `model_search/.env` files are not read by the trial engine.
106
+ An `OPENROUTER_API_KEY` in the process environment or
107
+ `~/.openmerit/.env` additionally enables OpenRouter catalog and public-
108
+ benchmark enrichment. It is optional for native-provider comparisons. If
109
+ using the file, protect it:
110
+
111
+ ```bash
112
+ nano ~/.openmerit/.env
113
+ chmod 600 ~/.openmerit/.env
114
+ ```
115
+
116
+ Pi does not read OpenMerit's `.env` file for its ordinary sessions, so an
117
+ OpenRouter route still needs Pi authentication. Older `model_search/.env`
118
+ files are not read.
107
119
 
108
120
  3. Verify the extension appears in `pi list`. The instruction file
109
121
  [`instructions/OPENMERIT.md`](instructions/OPENMERIT.md) can be added to a
110
122
  pi project's AGENTS.md for agent context, but the extension does
111
- not require it.
123
+ not require it. Run `npx openmerit doctor` after starting Pi once to inspect
124
+ the eligible route snapshot and configuration without printing secrets.
112
125
 
113
- 4. Start pi in one terminal with an OpenRouter model and tools disabled:
126
+ 4. Start pi in one terminal with any configured model. For example:
114
127
 
115
128
  ```bash
116
- pi --provider openrouter --model openai/gpt-4o-mini --no-tools
129
+ pi --provider openrouter --model openai/gpt-4o-mini
117
130
  ```
118
131
 
119
132
  For a first text task, ask: “Give the shortest valid word ladder from cat
120
133
  to dog. Each step changes one letter and must be a common English word.
121
134
  Return only the path.” Wait for A to finish and leave the pi session open.
122
- The extension automatically starts B and C **sequentially through pi**,
135
+ The extension automatically starts B and C **sequentially through their
136
+ selected Pi routes**,
123
137
  reports each score in Pi, and writes a recommendation for this exact
124
138
  session. With the shipped supervised policy, use
125
139
  `/openmerit` inside pi to inspect the evidence and `/openmerit apply` to
126
140
  switch. Then send a second message to see which model actually handles it.
127
141
  You do not need a watcher terminal or the standalone `trial` command.
128
142
 
143
+ Candidate comparisons have no tools by default, even when the observed task
144
+ used tools. See **Alpha boundaries** before explicitly enabling candidate
145
+ tools.
146
+
129
147
  To opt in to automatic swaps, edit `~/.openmerit/policy.json`, set
130
148
  `"mode": "auto"` and `"auto_apply.enabled": true`, review the score-gain
131
149
  and price-ratio thresholds, and run the next task.
@@ -133,7 +151,7 @@ Keep that checkout available because pi loads a local package from its path.
133
151
  ### Try an invoice-to-JSON task
134
152
 
135
153
  Use the same one-terminal setup. Start pi with a vision-capable model, for
136
- example `openai/gpt-4o-mini`, and `--no-tools`. In pi, type `@` to select
154
+ example `openai/gpt-4o-mini`. In pi, type `@` to select
137
155
  [`benchmark/invoice_ocr/data/invoice_01_row_2.jpg`](benchmark/invoice_ocr/data/invoice_01_row_2.jpg)
138
156
  and paste the **entire, unchanged** text from
139
157
  [`examples/invoice-prompt.txt`](examples/invoice-prompt.txt) into the same
@@ -161,7 +179,8 @@ schema, so its published scores are not directly comparable to these pi runs.
161
179
 
162
180
  ### Inspect or troubleshoot a run
163
181
 
164
- `npx openmerit status` shows the latest pi model and pending recommendations;
182
+ `npx openmerit doctor` checks Pi, policy, eligible routes, and optional
183
+ OpenRouter enrichment. `npx openmerit status` shows the latest pi model and pending recommendations;
165
184
  `npx openmerit frontier` shows measured quality, blended price, latency, and
166
185
  the chosen frontier per task. `/openmerit` inside pi shows the current model,
167
186
  fallback, the model currently being compared, completed models with quality,
@@ -176,43 +195,74 @@ for local experiments while keeping `budgets.max_usd_per_day` as a conservative
176
195
  candidate-spend threshold.
177
196
 
178
197
  If no comparison starts after Pi settles, confirm that pi loaded the extension
179
- (`pi list`), the task finished, the pi session is saved (do not use
180
- `--no-session`), and the task used no tools. Image candidates must advertise
181
- image input in OpenRouter's current model catalog. The extension queues
198
+ (`pi list`), the task finished, and the pi session is saved (do not use
199
+ `--no-session`). Image candidates must advertise image input in Pi's model
200
+ registry. The extension queues
182
201
  completed tasks from its current session and runs one comparison at a time.
183
202
  Closing or switching the session cancels the active job.
184
203
  Candidate runs use the same text and uploaded image bytes, but they do not
185
204
  replay earlier answers or file changes. An exact task in another session
186
205
  (including identical image bytes) appears as **advice**, not a pending swap;
187
206
  the new session still gets its own comparison. The optional `trial` command
188
- remains for controlled text-task runs from a task JSON file.
207
+ remains for controlled text-task runs from a task JSON file; that legacy
208
+ standalone command is still OpenRouter-specific in 0.1.2. The automatic Pi
209
+ session path is provider-neutral.
189
210
 
190
211
  ## Alpha boundaries
191
212
 
192
- - With the extension installed and an OpenRouter key configured, each supported
193
- settled task can start comparison calls automatically. Candidate, judge, and
194
- strategist requests send the task text, attached images, and candidate output
195
- to OpenRouter and can incur charges. Use non-sensitive test tasks and
196
- conservative account limits while evaluating this alpha.
197
- - The extension queues tasks from its current pi session, compares one at a
198
- time, and skips tasks that use tools. Candidate sessions replay the latest
199
- text and images, not earlier conversation context or workspace changes.
200
- - `ledger.json` counts reported candidate-trial spend and the daily number of
201
- trials. It does not yet include rubric, judge, or strategist request cost,
202
- and `max_usd_per_trial` is not yet a hard provider-side cap.
213
+ - With the extension installed and at least two eligible Pi routes, each
214
+ supported settled task can start comparison calls automatically. Candidate,
215
+ judge, and strategist requests send the task text, attached files or images,
216
+ and candidate output to the configured providers and can incur charges. Use
217
+ non-sensitive test tasks and conservative account limits while evaluating
218
+ this alpha.
219
+ - The extension queues tasks from its current pi session and compares one at a
220
+ time. Candidate runs use isolated temporary copies of the original working
221
+ directory and have no Pi tools by default. Their temporary changes are
222
+ discarded and recorded as counts, and raw Pi JSON event streams are saved
223
+ under `~/.openmerit/traces/trials/`.
224
+ Candidate sessions replay the task text and images, not earlier conversation
225
+ context or workspace changes.
226
+ - Set `OPENMERIT_PI_TRIAL_TOOLS` to an explicit comma-separated allowlist such
227
+ as `read,grep,find,ls` to let candidate models use tools; `none` keeps them
228
+ disabled. Any tool access is an advanced opt-in: the copied working directory
229
+ prevents ordinary project writes from touching the original, but it is not an
230
+ OS sandbox. Tools may accept absolute paths, and shell tools may access the
231
+ network. Keep the no-tools default for untrusted or sensitive projects.
232
+ - File attachments that Pi records as `<file name="…">` are copied into every
233
+ candidate sandbox and passed back to Pi as `@` file inputs. This covers PDFs,
234
+ CSVs, spreadsheets, and other files that Pi can open; OpenMerit does not
235
+ implement a separate parser for them.
236
+ - `ledger.json` counts reported candidate, rubric, judge, and strategist spend
237
+ for automatic session comparisons plus the daily candidate count.
238
+ `max_usd_per_trial` is a conservative admission estimate based on known Pi
239
+ prices and a 4K answer; it is not a provider-side hard cap. Routes without
240
+ known pricing are excluded from automatic comparisons and cannot auto-apply.
203
241
  - Each comparison has one observed baseline plus a small candidate slate and
204
242
  one quality score per answer. Treat recommendations as experimental evidence,
205
243
  not a universal model ranking.
244
+ - Candidate execution, judging, and strategy use the provider/model routes
245
+ exposed by Pi. OpenRouter remains an optional route plus catalog/benchmark
246
+ enrichment source. `HarnessAdapter`, `ModelProviderAdapter`,
247
+ `ObservationSource`, and `EventSink` remain separate integration boundaries;
248
+ Pi and local JSONL are the implementations shipped in this release.
249
+ - Candidate subprocesses can use Pi built-ins and custom/local routes available
250
+ without loading extensions (for example routes from Pi's model
251
+ configuration). A provider registered only at runtime by another extension
252
+ is visible in the route snapshot but cannot yet be executed by the isolated
253
+ subprocess.
206
254
 
207
255
  ## State layout (`~/.openmerit/`)
208
256
 
209
257
  | file | contents |
210
258
  |---|---|
211
259
  | `policy.json` | the policy file (gate thresholds, budgets, intervals) |
212
- | `harness-state.json` | written by the pi extension: current model, fallback, latest session and settled task |
213
- | `recommendations.jsonl` | shared append-only channel: session-bound swap recommendations + status updates |
214
- | `trials.jsonl` | every model trial point (source, score, price, latency) |
215
- | `observations.jsonl` | task observations extracted from session traces |
260
+ | `harness-state.json` | current/fallback routes, Pi's eligible route snapshot, latest session and settled task |
261
+ | `recommendations.jsonl` | append-only session-bound recommendations, routes, evidence, gate reasons, and status updates |
262
+ | `trials.jsonl` | every model trial point, including its provider route when known |
263
+ | `traces/observations.jsonl` | task observations extracted from session traces |
264
+ | `events.jsonl` | versioned provider-neutral observation, trial, and recommendation events for future sinks |
265
+ | `traces/trials/*.jsonl` | raw Pi JSON event streams for candidate trials |
216
266
  | `catalog/snapshot.json` + `candidates.json` | catalog snapshot (including input modalities) + new-model queue |
217
267
  | `benchmarks/digest.json` | public-benchmark scores per model (seed + refresh) |
218
268
  | `ledger.json` | daily trial spend (budget enforcement) |
package/dist/catalog.js CHANGED
@@ -1,7 +1,17 @@
1
1
  /** OpenRouter catalog fetch + snapshot diffing: the "new model release" watch. */
2
2
  import { readJson, paths, writeJson } from "./store.js";
3
3
  const OR = "https://openrouter.ai/api/v1";
4
- export async function fetchCatalog(key) {
4
+ export class OpenRouterCatalogProvider {
5
+ key;
6
+ id = "openrouter";
7
+ constructor(key) {
8
+ this.key = key;
9
+ }
10
+ fetchCatalog() { return fetchCatalog(this.key); }
11
+ }
12
+ export async function fetchCatalog(key, provider = "openrouter") {
13
+ if (provider !== "openrouter")
14
+ return new Map();
5
15
  const res = await fetch(`${OR}/models`, {
6
16
  headers: { Authorization: `Bearer ${key}` },
7
17
  });
package/dist/cli.js CHANGED
@@ -1,5 +1,5 @@
1
1
  #!/usr/bin/env node
2
- /** openmerit CLI: init / watch / trial / frontier / recommend / status. */
2
+ /** openmerit CLI: init / watch / trial / frontier / recommend / status / doctor. */
3
3
  import { copyFileSync, existsSync, mkdirSync, readFileSync } from "node:fs";
4
4
  import { dirname, join } from "node:path";
5
5
  import { fileURLToPath } from "node:url";
@@ -11,10 +11,12 @@ import { buildRecommendation } from "./recommend.js";
11
11
  import { appendJsonl, paths, readJson, readJsonl, taskKey } from "./store.js";
12
12
  import { autoTaskTick, recommendTick, runDaemon, tickOnce } from "./daemon.js";
13
13
  import { budgetOk, recordTrialSpend } from "./trials.js";
14
- import { availablePiModels, recordedActiveTask, runPiTrial, scorePiRun, settledActiveTask } from "./pi-trials.js";
14
+ import { availablePiModels, availablePiRoutes, piExecutable, recordedActiveTask, runPiTrial, scorePiRun, settledActiveTask } from "./pi-trials.js";
15
15
  import { pickNext, STRAT_PREFS } from "./strategist.js";
16
16
  import { JUDGE_PREFS } from "./judge.js";
17
17
  import { openRouterBenchmarkCandidates, relevantBenchmarks } from "./benchmarks.js";
18
+ import { modelRoute, routeLabel } from "./routes.js";
19
+ import { spawnSync } from "node:child_process";
18
20
  const ROOT = dirname(dirname(fileURLToPath(import.meta.url)));
19
21
  function resolvePref(cat, prefs, label) {
20
22
  for (const p of prefs)
@@ -43,7 +45,48 @@ function cmdInit() {
43
45
  console.log(`ext. -> ${join(ROOT, "extension", "openmerit.ts")}`);
44
46
  console.log(`pi npm -> pi install npm:openmerit`);
45
47
  console.log(`pi local-> pi install "${ROOT}"`);
46
- console.log("\nNext: set OPENROUTER_API_KEY (env or ~/.openmerit/.env), install one extension source into pi, then start pi; the extension launches comparisons automatically.");
48
+ console.log("\nNext: authenticate at least two models in pi, install one extension source, then start pi. OpenRouter is optional enrichment; candidate tools are disabled by default.");
49
+ }
50
+ function optionalOpenRouterKey() {
51
+ try {
52
+ return loadKey();
53
+ }
54
+ catch {
55
+ return null;
56
+ }
57
+ }
58
+ function cmdDoctor() {
59
+ const executable = piExecutable();
60
+ const version = spawnSync(executable, ["--version"], { encoding: "utf8" });
61
+ const state = readJson(paths.harnessState(), {});
62
+ console.log(`pi executable : ${executable}`);
63
+ console.log(`pi detected : ${version.status === 0 ? `yes (${version.stdout.trim()})` : "no"}`);
64
+ try {
65
+ const policy = loadPolicy(paths.policy());
66
+ console.log(`policy : valid (v${policy.version}, ${policy.mode})`);
67
+ }
68
+ catch (error) {
69
+ console.log(`policy : invalid (${error.message})`);
70
+ }
71
+ let routes = state.routes ?? [];
72
+ const hasLiveSnapshot = routes.length > 0;
73
+ if (routes.length === 0 && version.status === 0) {
74
+ try {
75
+ routes = availablePiRoutes();
76
+ }
77
+ catch { /* reported below */ }
78
+ }
79
+ console.log(`eligible routes: ${routes.length}`);
80
+ for (const route of routes.slice(0, 8))
81
+ console.log(` - ${routeLabel(route)}${route.cost ? "" : " (price unknown)"}`);
82
+ if (routes.length > 8)
83
+ console.log(` … ${routes.length - 8} more`);
84
+ console.log(`active route : ${state.currentRoute ? routeLabel(state.currentRoute) : "not reported; start pi once"}`);
85
+ const pricedRoutes = routes.filter((route) => !!route.cost).length;
86
+ console.log(`alternates : ${hasLiveSnapshot
87
+ ? pricedRoutes >= 2 ? "ready" : "need at least two priced eligible pi routes"
88
+ : "start Pi once to capture route prices and capabilities"}`);
89
+ console.log(`OpenRouter enrichment: ${optionalOpenRouterKey() ? "available" : "not configured (optional)"}`);
47
90
  }
48
91
  async function cmdTrial(taskFile, rounds) {
49
92
  const cfg = JSON.parse(readFileSync(taskFile, "utf8"));
@@ -189,7 +232,8 @@ function cmdStatus() {
189
232
  latestById.set(r.id, r);
190
233
  const recs = [...latestById.values()].filter((r) => r.status === "pending");
191
234
  const frontiers = loadFrontiers();
192
- console.log(`harness model : ${harness.currentModel ?? "unknown"} (as of ${harness.updatedAt ?? "n/a"})`);
235
+ console.log(`harness model : ${harness.currentRoute ? routeLabel(harness.currentRoute) : harness.currentModel ?? "unknown"} ` +
236
+ `(as of ${harness.updatedAt ?? "n/a"})`);
193
237
  console.log(`tasks tracked : ${frontiers.length}`);
194
238
  console.log(`new models queued for trial: ${queue.newModels.length}${queue.newModels.length ? " — " + queue.newModels.slice(0, 5).map((m) => m.id).join(", ") + (queue.newModels.length > 5 ? "…" : "") : ""}`);
195
239
  console.log(`pending recommendations: ${recs.length}`);
@@ -205,16 +249,22 @@ async function main() {
205
249
  cmdInit();
206
250
  break;
207
251
  case "session-trial": {
208
- const [sessionFile, settledAt, currentModel, settledTaskKey, cwd, bytesText] = args;
252
+ const [sessionFile, settledAt, currentModel, settledTaskKey, cwd, bytesText, provider, providerModelId] = args;
209
253
  if (!sessionFile || !settledAt || !currentModel || !settledTaskKey || !cwd || !bytesText)
210
254
  throw new Error("usage: openmerit session-trial <session-file> <settled-at> <model> <task-key> <cwd> <session-bytes>");
211
255
  const sessionBytes = Number(bytesText);
212
256
  if (!Number.isSafeInteger(sessionBytes) || sessionBytes <= 0)
213
257
  throw new Error("invalid session byte limit");
214
- const task = settledActiveTask({ sessionFile, settledAt, currentModel, settledTaskKey, cwd, sessionBytes });
258
+ const state = readJson(paths.harnessState(), {});
259
+ const currentRoute = provider && providerModelId
260
+ ? state.routes?.find((route) => route.provider === provider && route.modelId === providerModelId) ??
261
+ modelRoute(provider, providerModelId)
262
+ : state.currentRoute;
263
+ const task = settledActiveTask({ sessionFile, settledAt, currentModel, currentRoute,
264
+ routes: state.routes, settledTaskKey, cwd, sessionBytes });
215
265
  if (!task)
216
266
  throw new Error("the specified pi session has no completed, supported task matching this marker");
217
- await autoTaskTick(loadKey(), policy, task);
267
+ await autoTaskTick(optionalOpenRouterKey(), policy, task);
218
268
  break;
219
269
  }
220
270
  case "watch":
@@ -240,6 +290,9 @@ async function main() {
240
290
  case "status":
241
291
  cmdStatus();
242
292
  break;
293
+ case "doctor":
294
+ cmdDoctor();
295
+ break;
243
296
  default:
244
297
  console.log(`openmerit — external model-merit harness
245
298
 
@@ -250,7 +303,8 @@ usage: openmerit <command>
250
303
  trial <task.json> iterative model search on one task spec [--rounds N]
251
304
  frontier [taskKey] print pareto frontier(s)
252
305
  recommend emit recommendations now
253
- status harness model, queued models, pending recommendations`);
306
+ status harness model, queued models, pending recommendations
307
+ doctor verify pi, policy, routes, and optional enrichment`);
254
308
  if (cmd && cmd !== "help" && cmd !== "--help")
255
309
  process.exitCode = 1;
256
310
  }