openmerit 0.1.3 → 0.1.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -142,6 +142,12 @@ Keep that checkout available because pi loads a local package from its path.
142
142
  switch. Then send a second message to see which model actually handles it.
143
143
  You do not need a watcher terminal or the standalone `trial` command.
144
144
 
145
+ Use `/openmerit pause` to stop the active comparison and suppress automatic
146
+ comparisons, `/openmerit resume` to enable them for future completed tasks,
147
+ and `/openmerit compare` to explicitly compare the latest completed task
148
+ even while automatic comparisons are paused. `/openmerit doctor` runs the
149
+ sanitized setup checks without leaving Pi.
150
+
145
151
  Candidate comparisons have no tools by default, even when the observed task
146
152
  used tools. See **Alpha boundaries** before explicitly enabling candidate
147
153
  tools.
@@ -150,6 +156,57 @@ Keep that checkout available because pi loads a local package from its path.
150
156
  `"mode": "auto"` and `"auto_apply.enabled": true`, review the score-gain
151
157
  and price-ratio thresholds, and run the next task.
152
158
 
159
+ ### Configure custom or local routes
160
+
161
+ Pi normally supplies each route's price, context window, output limit, and
162
+ modalities. Some custom and local providers omit that metadata. OpenMerit does
163
+ not guess that a local model is free: add an override keyed by the exact
164
+ `provider:modelId` route in `~/.openmerit/policy.json` instead:
165
+
166
+ ```json
167
+ {
168
+ "route_overrides": {
169
+ "ollama:qwen3:8b": {
170
+ "cost": { "input": 0, "output": 0 },
171
+ "context_window": 32768,
172
+ "max_tokens": 4096,
173
+ "input": ["text"]
174
+ }
175
+ }
176
+ }
177
+ ```
178
+
179
+ All fields are optional, but both `cost.input` and `cost.output` are required
180
+ when declaring cost. Prices use Pi's dollars-per-million-token units. The
181
+ override applies only to that exact route; it does not change the stable
182
+ `vendor/model` identity or another provider's route to the same model.
183
+
184
+ If the provider itself is registered at runtime by a Pi extension, explicitly
185
+ allow that provider-registration file in the same policy. Paths must be
186
+ absolute, existing files:
187
+
188
+ ```json
189
+ {
190
+ "pi": {
191
+ "provider_extensions": [
192
+ "/absolute/path/to/ollama-provider.ts"
193
+ ]
194
+ }
195
+ }
196
+ ```
197
+
198
+ OpenMerit still launches subprocesses with extension discovery disabled, then
199
+ loads only these explicit files. Do not add OpenMerit's own extension. An
200
+ allowlisted extension executes code in every candidate, judge, and strategist
201
+ Pi subprocess, so list only provider extensions you trust. Run
202
+ `npx openmerit doctor` or `/openmerit doctor` to validate the files and confirm
203
+ that route overrides match Pi's visible routes.
204
+
205
+ Existing `0.1.x` policy files remain valid: missing `route_overrides` and
206
+ `pi.provider_extensions` fields normalize to empty safe defaults. `openmerit init`
207
+ continues to preserve an existing policy, so add these fields manually
208
+ only when you need them.
209
+
153
210
  ### Try an invoice-to-JSON task
154
211
 
155
212
  Use the same one-terminal setup. Start pi with a vision-capable model, for
@@ -193,7 +250,9 @@ the chosen frontier per task. `/openmerit` inside pi shows the current model,
193
250
  fallback, the model currently being compared, completed models with quality,
194
251
  cost, and latency, trial budget, pending recommendations, and any exact-task result
195
252
  measured in another session. When the gate declines an automatic swap, it
196
- prints reasons and `/openmerit apply` remains available.
253
+ prints reasons and `/openmerit apply` remains available. It also lists every
254
+ route skipped during the current comparison with the relevant policy, pricing,
255
+ modality, or per-trial budget reason.
197
256
 
198
257
  The daily trial-count and dollar limits come from `~/.openmerit/policy.json`.
199
258
  If a limit is reached, the extension reports why it skipped the comparison;
@@ -255,7 +314,9 @@ Pi. An OpenRouter key only adds optional catalog and public-benchmark metadata.
255
314
  for automatic session comparisons plus the daily candidate count.
256
315
  `max_usd_per_trial` is a conservative admission estimate based on known Pi
257
316
  prices and a 4K answer; it is not a provider-side hard cap. Routes without
258
- known pricing are excluded from automatic comparisons and cannot auto-apply.
317
+ known pricing are excluded as candidates and cannot auto-apply. Use an exact
318
+ `route_overrides` entry for a custom/local route whose price is known; zero
319
+ cost must be stated explicitly.
259
320
  - Each comparison has one observed baseline plus a small candidate slate and
260
321
  one quality score per answer. Treat recommendations as experimental evidence,
261
322
  not a universal model ranking.
@@ -264,11 +325,10 @@ Pi. An OpenRouter key only adds optional catalog and public-benchmark metadata.
264
325
  enrichment source. `HarnessAdapter`, `ModelProviderAdapter`,
265
326
  `ObservationSource`, and `EventSink` remain separate integration boundaries;
266
327
  Pi and local JSONL are the implementations shipped in this release.
267
- - Candidate subprocesses can use Pi built-ins and custom/local routes available
268
- without loading extensions (for example routes from Pi's model
269
- configuration). A provider registered only at runtime by another extension
270
- is visible in the route snapshot but cannot yet be executed by the isolated
271
- subprocess.
328
+ - Candidate subprocesses can use Pi built-ins, configured custom/local routes,
329
+ and providers registered by explicitly allowlisted extension files. Normal
330
+ extension discovery remains disabled, and OpenMerit refuses to load its own
331
+ extension recursively.
272
332
  - JSON state snapshots are replaced atomically. Append-only readers skip and
273
333
  report malformed or interrupted lines while retaining later valid records.
274
334
  A job owned by a crashed process is reclaimable instead of remaining stuck
@@ -278,8 +338,8 @@ Pi. An OpenRouter key only adds optional catalog and public-benchmark metadata.
278
338
 
279
339
  | file | contents |
280
340
  |---|---|
281
- | `policy.json` | the policy file (gate thresholds, budgets, intervals) |
282
- | `harness-state.json` | current/fallback routes, Pi's eligible route snapshot, latest session and settled task |
341
+ | `policy.json` | gate thresholds, budgets, intervals, exact-route metadata overrides, and allowlisted Pi provider extensions |
342
+ | `harness-state.json` | current/fallback routes, Pi's eligible route snapshot, pause state, latest session and settled task |
283
343
  | `recommendations.jsonl` | append-only session-bound recommendations, routes, evidence, gate reasons, and status updates |
284
344
  | `trials.jsonl` | every model trial point, including its provider route when known |
285
345
  | `traces/observations.jsonl` | task observations extracted from session traces |
package/dist/daemon.js CHANGED
@@ -11,10 +11,10 @@ import { buildRecommendation } from "./recommend.js";
11
11
  import { appendJsonl, paths, readJson, readJsonl, writeJson, } from "./store.js";
12
12
  import { budgetOk, ensureRubric, recordMeritSpend, recordTrialSpend, runTrial, trialBudgetOk } from "./trials.js";
13
13
  import { JUDGE_PREFS } from "./judge.js";
14
- import { runPiTrial, scorePiRun, settledActiveTask } from "./pi-trials.js";
14
+ import { PiHarnessAdapter, runPiTrial, scorePiRun, settledActiveTask } from "./pi-trials.js";
15
15
  import { taskInputKey } from "./task-input.js";
16
16
  import { pickNext, STRAT_PREFS } from "./strategist.js";
17
- import { enrichRouteEntry, routeCatalog, routeKey, routeLabel } from "./routes.js";
17
+ import { applyRouteOverrides, enrichRouteEntry, routeCatalog, routeKey, routeLabel } from "./routes.js";
18
18
  import { PiCliChatClient } from "./llm.js";
19
19
  import { LocalJsonlEventSink, PiTraceObservationSource, meritEvent } from "./integrations.js";
20
20
  function resolveRoutePref(cat, prefs, label, requireImages = false) {
@@ -172,9 +172,13 @@ export async function trialTick(key, policy) {
172
172
  /** Compare the one completed active pi task, sequentially, then target its session. */
173
173
  export async function autoTaskTick(key, policy, settledTask) {
174
174
  const sink = new LocalJsonlEventSink();
175
- const task = settledTask === undefined ? settledActiveTask() : settledTask;
175
+ let task = settledTask === undefined ? settledActiveTask() : settledTask;
176
176
  if (!task)
177
177
  return null;
178
+ const routes = applyRouteOverrides(task.routes, policy.route_overrides);
179
+ const currentRoute = applyRouteOverrides([task.route], policy.route_overrides)[0];
180
+ task = { ...task, routes, route: currentRoute };
181
+ const tKey = taskInputKey(task.task, task.images, task.files);
178
182
  const marker = settledTask === undefined
179
183
  ? `${task.sessionFile}:${task.settledAt}`
180
184
  : `${task.sessionFile}:${task.sessionBytes ?? task.settledAt}`;
@@ -202,13 +206,28 @@ export async function autoTaskTick(key, policy, settledTask) {
202
206
  console.log(`[openmerit] OpenRouter enrichment unavailable: ${error.message}`);
203
207
  }
204
208
  }
209
+ const excluded = [];
205
210
  for (const [id, entry] of [...cat]) {
206
- if ((entry.priceKnown !== false && entry.price > policy.max_usd_per_m) ||
207
- !providerAllowed(policy, entry.id, entry.route?.provider) ||
208
- (task.images.length > 0 && !entry.inputModalities?.includes("image")) ||
209
- (id !== routeKey(task.route) && !trialBudgetOk(policy, entry, task.task.length).ok))
211
+ const reasons = [];
212
+ if (entry.priceKnown !== false && entry.price > policy.max_usd_per_m)
213
+ reasons.push(`price $${entry.price.toFixed(2)}/M exceeds $${policy.max_usd_per_m.toFixed(2)}/M policy limit`);
214
+ if (!providerAllowed(policy, entry.id, entry.route?.provider))
215
+ reasons.push("provider is denied by policy");
216
+ if (task.images.length > 0 && !entry.inputModalities?.includes("image"))
217
+ reasons.push("route does not advertise image input");
218
+ if (id !== routeKey(task.route)) {
219
+ const admission = trialBudgetOk(policy, entry, task.task.length);
220
+ if (!admission.ok && admission.reason)
221
+ reasons.push(admission.reason);
222
+ }
223
+ if (reasons.length) {
210
224
  cat.delete(id);
225
+ excluded.push({ model: entry.route ? routeLabel(entry.route) : entry.id, reasons: [...new Set(reasons)] });
226
+ }
211
227
  }
228
+ if (settledTask !== undefined)
229
+ emitTrialProgress({ phase: "selection", taskKey: tKey,
230
+ eligible: [...cat.values()].map((entry) => entry.route ? routeLabel(entry.route) : entry.id), excluded });
212
231
  const activeEntry = cat.get(routeKey(task.route));
213
232
  if (!activeEntry)
214
233
  throw new Error(`active route ${task.route.provider}/${task.route.modelId} is excluded by policy or task capabilities`);
@@ -216,7 +235,8 @@ export async function autoTaskTick(key, policy, settledTask) {
216
235
  throw new Error("Pi exposes no eligible alternate model route for this task");
217
236
  const judge = resolveRoutePref(cat, [policy.judge_model, ...JUDGE_PREFS, task.model], "judge", task.images.length > 0);
218
237
  const strategist = resolveRoutePref(cat, [policy.strategist_model, ...STRAT_PREFS, task.model], "strategist");
219
- const client = new PiCliChatClient(undefined, recordMeritSpend);
238
+ const client = new PiCliChatClient(undefined, recordMeritSpend, policy.pi.provider_extensions);
239
+ const harness = new PiHarnessAdapter(policy.pi.provider_extensions);
220
240
  const { category, benchmarks } = relevantBenchmarks(task.task);
221
241
  let benchmarkCandidates = [];
222
242
  try {
@@ -227,7 +247,6 @@ export async function autoTaskTick(key, policy, settledTask) {
227
247
  console.log(`[openmerit] benchmark shortlist unavailable: ${e.message}`);
228
248
  }
229
249
  const rubric = await ensureRubric(key ?? "", judge.route, task.task, new Map(), client);
230
- const tKey = taskInputKey(task.task, task.images, task.files);
231
250
  const points = [];
232
251
  const tried = new Set();
233
252
  const failedVendors = new Set();
@@ -256,7 +275,7 @@ export async function autoTaskTick(key, policy, settledTask) {
256
275
  emitTrialProgress({ phase: "start", taskKey: tKey,
257
276
  model: pick.route ? routeLabel(pick.route) : pick.model, index: i + 1, total });
258
277
  const entry = pick.route ? cat.get(routeKey(pick.route)) : cat.get(pick.model);
259
- const { point, costUsd, error } = await runPiTrial(key ?? "", judge.route, task.task, rubric, pick.model, entry, task.cwd, task.images, task.files, undefined, client);
278
+ const { point, costUsd, error } = await runPiTrial(key ?? "", judge.route, task.task, rubric, pick.model, entry, task.cwd, task.images, task.files, harness, client);
260
279
  tried.add(pick.route ? routeKey(pick.route) : pick.model);
261
280
  points.push(point);
262
281
  recordTrialSpend(costUsd);
@@ -1,13 +1,14 @@
1
1
  /** Read-only diagnostics plus a provider-free self-test for first-run support. */
2
2
  import { spawnSync } from "node:child_process";
3
- import { appendFileSync, existsSync, mkdtempSync, readFileSync, rmSync } from "node:fs";
3
+ import { appendFileSync, existsSync, mkdtempSync, readFileSync, rmSync, writeFileSync } from "node:fs";
4
4
  import { tmpdir } from "node:os";
5
5
  import { basename, join } from "node:path";
6
6
  import { LocalJsonlEventSink, meritEvent } from "./integrations.js";
7
7
  import { piExecutable, availablePiRoutes } from "./pi-trials.js";
8
8
  import { DEFAULT_POLICY, loadPolicy, providerAllowed } from "./policy.js";
9
9
  import { buildRecommendation } from "./recommend.js";
10
- import { modelRoute, routeCatalog, routeKey, routeLabel } from "./routes.js";
10
+ import { applyRouteOverrides, modelRoute, routeCatalog, routeKey, routeLabel } from "./routes.js";
11
+ import { validatePiProviderExtensions } from "./pi-config.js";
11
12
  import { appendJsonl, paths, readJsonReport, readJsonlReport, writeJson, } from "./store.js";
12
13
  function check(id, status, message, details) {
13
14
  return { id, status, message, ...(details?.length ? { details } : {}) };
@@ -32,13 +33,34 @@ export function collectDoctorReport() {
32
33
  checks.push(piVersion
33
34
  ? check("pi", "pass", `Pi detected (${piVersion})`)
34
35
  : check("pi", "fail", "Pi is not executable"));
36
+ let policy = DEFAULT_POLICY;
35
37
  try {
36
- const policy = loadPolicy(paths.policy());
38
+ policy = loadPolicy(paths.policy());
37
39
  checks.push(check("policy", "pass", `Policy is valid (v${policy.version}, ${policy.mode})`));
38
40
  }
39
41
  catch (error) {
40
42
  checks.push(check("policy", "fail", `Policy is invalid: ${error.message}`));
41
43
  }
44
+ const extensions = validatePiProviderExtensions(policy.pi.provider_extensions);
45
+ const safeExtensionErrors = extensions.errors.map((message) => {
46
+ const configured = policy.pi.provider_extensions.find((file) => message.startsWith(`${file}:`));
47
+ return configured ? message.replace(configured, `local:${basename(configured)}`) : message;
48
+ });
49
+ let isolatedExtensionRoutes = null;
50
+ let extensionLoadError = null;
51
+ if (!safeExtensionErrors.length && policy.pi.provider_extensions.length && piVersion) {
52
+ try {
53
+ isolatedExtensionRoutes = availablePiRoutes(undefined, extensions.valid);
54
+ }
55
+ catch {
56
+ extensionLoadError = "Pi could not load the allowlisted provider extensions in isolated mode";
57
+ }
58
+ }
59
+ checks.push(safeExtensionErrors.length || extensionLoadError
60
+ ? check("provider-extensions", "fail", "One or more allowlisted Pi provider extensions are unavailable", [...safeExtensionErrors, ...(extensionLoadError ? [extensionLoadError] : [])])
61
+ : check("provider-extensions", "pass", policy.pi.provider_extensions.length
62
+ ? `${policy.pi.provider_extensions.length} explicit Pi provider extension${policy.pi.provider_extensions.length === 1 ? " loads" : "s load"} in isolated mode`
63
+ : "No Pi provider extensions are allowlisted"));
42
64
  const stateRead = readJsonReport(paths.harnessState(), {});
43
65
  checks.push(!stateRead.valid
44
66
  ? check("harness-state", "fail", "harness-state.json is malformed; start Pi to replace it atomically")
@@ -49,16 +71,17 @@ export function collectDoctorReport() {
49
71
  const liveSnapshot = routes.length > 0;
50
72
  if (!routes.length && piVersion) {
51
73
  try {
52
- routes = availablePiRoutes();
74
+ routes = isolatedExtensionRoutes ?? availablePiRoutes(undefined, extensions.valid);
53
75
  }
54
76
  catch { /* the check below explains it */ }
55
77
  }
56
- let policy = DEFAULT_POLICY;
57
- try {
58
- policy = loadPolicy(paths.policy());
59
- }
60
- catch { /* already reported */ }
78
+ routes = applyRouteOverrides(routes, policy.route_overrides);
79
+ const visibleRouteKeys = new Set(routes.map(routeKey));
61
80
  routes = routes.filter((route) => providerAllowed(policy, route.meritId, route.provider));
81
+ const unusedOverrides = Object.keys(policy.route_overrides).filter((key) => !visibleRouteKeys.has(key));
82
+ checks.push(unusedOverrides.length
83
+ ? check("route-overrides", "warn", `${unusedOverrides.length} route metadata override${unusedOverrides.length === 1 ? " does" : "s do"} not match a visible route`, unusedOverrides)
84
+ : check("route-overrides", "pass", `${Object.keys(policy.route_overrides).length} route metadata override${Object.keys(policy.route_overrides).length === 1 ? "" : "s"} applied or ready`));
62
85
  const priced = routes.filter((route) => !!route.cost);
63
86
  checks.push(routes.length >= 2
64
87
  ? check("routes", "pass", `${routes.length} Pi model routes are visible`, routes.slice(0, 8).map((route) => `${routeLabel(route)}${route.cost ? "" : " (price unknown)"}`))
@@ -156,6 +179,16 @@ export function runOfflineVerify() {
156
179
  recommendation.policy.reasons.length > 0
157
180
  ? check("provider-neutral", "pass", "Exact provider routes and policy reasons survive recommendation")
158
181
  : check("provider-neutral", "fail", "Provider route or policy evidence was lost"));
182
+ const configured = applyRouteOverrides([modelRoute("custom", "model-c")], {
183
+ "custom:model-c": { cost: { input: 0, output: 0 }, context_window: 4096,
184
+ max_tokens: 1024, input: ["text"] },
185
+ })[0];
186
+ const providerExtension = join(root, "custom-provider.ts");
187
+ writeFileSync(providerExtension, "export default function provider() {}\n");
188
+ const extension = validatePiProviderExtensions([providerExtension]);
189
+ checks.push(configured.cost?.input === 0 && configured.contextWindow === 4096 && extension.valid.length === 1
190
+ ? check("custom-routes", "pass", "Exact route metadata and allowlisted provider extensions validate")
191
+ : check("custom-routes", "fail", "Custom route or provider-extension configuration failed"));
159
192
  if (recommendation)
160
193
  appendJsonl(paths.recommendations(), recommendation);
161
194
  appendJsonl(paths.recommendations(), { schemaVersion: 0, id: "legacy", status: "dismissed" });
package/dist/llm.js CHANGED
@@ -5,6 +5,7 @@ import { tmpdir } from "node:os";
5
5
  import { join } from "node:path";
6
6
  import { paths } from "./store.js";
7
7
  import { legacyOpenRouterRoute } from "./routes.js";
8
+ import { piProviderExtensionArgs } from "./pi-config.js";
8
9
  const OR = "https://openrouter.ai/api/v1";
9
10
  /** Load OPENROUTER_API_KEY from env or ~/.openmerit/.env. */
10
11
  export function loadKey() {
@@ -123,10 +124,12 @@ function textContent(content) {
123
124
  export class PiCliChatClient {
124
125
  executable;
125
126
  onSpend;
127
+ providerExtensions;
126
128
  id = "pi-cli";
127
- constructor(executable = process.env.OPENMERIT_PI_BIN?.trim() || "pi", onSpend) {
129
+ constructor(executable = process.env.OPENMERIT_PI_BIN?.trim() || "pi", onSpend, providerExtensions = []) {
128
130
  this.executable = executable;
129
131
  this.onSpend = onSpend;
132
+ this.providerExtensions = providerExtensions;
130
133
  }
131
134
  chat(model, prompt, _maxTokens, _temperature) {
132
135
  return this.run(model, prompt, []);
@@ -148,7 +151,8 @@ export class PiCliChatClient {
148
151
  });
149
152
  const child = spawn(this.executable, [
150
153
  "--provider", route.provider, "--model", route.modelId, "--mode", "json", "--offline",
151
- "--no-extensions", "--no-tools", "--print", ...imageArgs, prompt,
154
+ "--no-extensions", ...piProviderExtensionArgs(this.providerExtensions),
155
+ "--no-tools", "--print", ...imageArgs, prompt,
152
156
  ], { cwd: dir, stdio: ["ignore", "pipe", "pipe"] });
153
157
  let stdout = "";
154
158
  let stderr = "";
@@ -0,0 +1,46 @@
1
+ /** Safe, explicit Pi subprocess configuration. */
2
+ import { existsSync, realpathSync, statSync } from "node:fs";
3
+ import { basename, isAbsolute } from "node:path";
4
+ /**
5
+ * Provider extensions execute as code. Require exact, existing, absolute files
6
+ * and refuse OpenMerit itself so an isolated trial cannot recursively launch
7
+ * another merit loop.
8
+ */
9
+ export function validatePiProviderExtensions(paths) {
10
+ const valid = [];
11
+ const errors = [];
12
+ for (const configured of [...new Set(paths)]) {
13
+ if (!isAbsolute(configured)) {
14
+ errors.push(`${configured}: path is not absolute`);
15
+ continue;
16
+ }
17
+ if (!existsSync(configured)) {
18
+ errors.push(`${configured}: file does not exist`);
19
+ continue;
20
+ }
21
+ let actual = configured;
22
+ try {
23
+ actual = realpathSync(configured);
24
+ if (!statSync(actual).isFile()) {
25
+ errors.push(`${configured}: path is not a file`);
26
+ continue;
27
+ }
28
+ }
29
+ catch (error) {
30
+ errors.push(`${configured}: ${error.message}`);
31
+ continue;
32
+ }
33
+ if (/^openmerit\.[cm]?[jt]s$/i.test(basename(actual))) {
34
+ errors.push(`${configured}: refusing to load OpenMerit recursively`);
35
+ continue;
36
+ }
37
+ valid.push(actual);
38
+ }
39
+ return { valid, errors };
40
+ }
41
+ export function piProviderExtensionArgs(paths) {
42
+ const checked = validatePiProviderExtensions(paths);
43
+ if (checked.errors.length)
44
+ throw new Error(`invalid pi.provider_extensions: ${checked.errors.join("; ")}`);
45
+ return checked.valid.flatMap((file) => ["--extension", file]);
46
+ }
package/dist/pi-trials.js CHANGED
@@ -10,6 +10,7 @@ import { paths, readJson } from "./store.js";
10
10
  import { sessionFiles, sessionImages, taskInputKey } from "./task-input.js";
11
11
  import { legacyOpenRouterRoute, meritModelId, modelRoute, routeKey } from "./routes.js";
12
12
  import { directChatClient } from "./llm.js";
13
+ import { piProviderExtensionArgs } from "./pi-config.js";
13
14
  function answerText(content) {
14
15
  if (typeof content === "string")
15
16
  return content;
@@ -20,8 +21,12 @@ function answerText(content) {
20
21
  }
21
22
  /** Pi-backed harness adapter. Future harnesses can implement the same seam. */
22
23
  export class PiHarnessAdapter {
24
+ providerExtensions;
25
+ constructor(providerExtensions = []) {
26
+ this.providerExtensions = providerExtensions;
27
+ }
23
28
  runTask(model, input) {
24
- return executePiTask(model, input.task, input.cwd, input.images, input.files);
29
+ return executePiTask(model, input.task, input.cwd, input.images, input.files, this.providerExtensions);
25
30
  }
26
31
  }
27
32
  const defaultHarness = new PiHarnessAdapter();
@@ -192,10 +197,11 @@ function compactNumber(value) {
192
197
  return Math.round(Number(match[1]) * scale);
193
198
  }
194
199
  /** Read provider/model routes from pi when no live extension snapshot is available. */
195
- export function availablePiRoutes(snapshot) {
200
+ export function availablePiRoutes(snapshot, providerExtensions = []) {
196
201
  if (snapshot?.length)
197
202
  return snapshot;
198
- const result = spawnSync(piExecutable(), ["--offline", "--list-models"], { encoding: "utf8" });
203
+ const result = spawnSync(piExecutable(), ["--offline", "--no-extensions",
204
+ ...piProviderExtensionArgs(providerExtensions), "--list-models"], { encoding: "utf8" });
199
205
  if (result.error || result.status !== 0)
200
206
  throw new Error("could not list models from pi");
201
207
  const routes = [];
@@ -216,7 +222,7 @@ export function availablePiModels() {
216
222
  return new Set(availablePiRoutes().map((route) => route.meritId));
217
223
  }
218
224
  /** Each invocation is a fresh pi run, so model B does not inherit model A's answer. */
219
- export async function executePiTask(model, task, cwd = process.cwd(), images = [], files = []) {
225
+ export async function executePiTask(model, task, cwd = process.cwd(), images = [], files = [], providerExtensions = []) {
220
226
  const route = typeof model === "string" ? legacyOpenRouterRoute(model) : model;
221
227
  const imageDir = images.length ? mkdtempSync(join(tmpdir(), "openmerit-pi-image-")) : null;
222
228
  const suffix = {
@@ -247,7 +253,8 @@ export async function executePiTask(model, task, cwd = process.cwd(), images = [
247
253
  const traceLines = [];
248
254
  const child = spawn(piExecutable(), [
249
255
  "--provider", route.provider, "--model", route.modelId, "--mode", "json",
250
- "--offline", "--no-extensions", "--approve", ...piTrialToolArgs(),
256
+ "--offline", "--no-extensions", ...piProviderExtensionArgs(providerExtensions),
257
+ "--approve", ...piTrialToolArgs(),
251
258
  "--print", ...imageArgs, ...fileArgs, candidateTask,
252
259
  ], { cwd: workspace, stdio: ["ignore", "pipe", "pipe"] });
253
260
  let buffer = "";
package/dist/policy.js CHANGED
@@ -1,7 +1,7 @@
1
1
  /** Policy loading, validation, and the swap gate. */
2
2
  import { existsSync, readFileSync } from "node:fs";
3
3
  export const DEFAULT_POLICY = {
4
- version: 1,
4
+ version: 2,
5
5
  mode: "recommend",
6
6
  auto_apply: {
7
7
  enabled: false,
@@ -16,10 +16,78 @@ export const DEFAULT_POLICY = {
16
16
  judge_model: null,
17
17
  strategist_model: null,
18
18
  max_usd_per_m: 20.0,
19
+ route_overrides: {},
20
+ pi: { provider_extensions: [] },
19
21
  };
20
22
  function num(v, fallback) {
21
23
  return typeof v === "number" && Number.isFinite(v) ? v : fallback;
22
24
  }
25
+ function finiteAtLeast(value, min, label) {
26
+ if (typeof value !== "number" || !Number.isFinite(value) || value < min)
27
+ throw new Error(`${label} must be a finite number >= ${min}`);
28
+ return value;
29
+ }
30
+ function positiveInteger(value, label) {
31
+ const result = finiteAtLeast(value, 1, label);
32
+ if (!Number.isInteger(result))
33
+ throw new Error(`${label} must be an integer`);
34
+ return result;
35
+ }
36
+ function routeOverrides(value) {
37
+ if (value === undefined)
38
+ return {};
39
+ if (!value || typeof value !== "object" || Array.isArray(value))
40
+ throw new Error("route_overrides must be an object keyed by provider:modelId");
41
+ const result = {};
42
+ for (const [key, raw] of Object.entries(value)) {
43
+ const separator = key.indexOf(":");
44
+ if (separator <= 0 || separator === key.length - 1)
45
+ throw new Error(`route override ${key} must use provider:modelId`);
46
+ if (!raw || typeof raw !== "object" || Array.isArray(raw))
47
+ throw new Error(`route override ${key} must be an object`);
48
+ const input = raw.input;
49
+ if (input !== undefined && (!Array.isArray(input) || input.length === 0 ||
50
+ input.some((item) => item !== "text" && item !== "image")))
51
+ throw new Error(`route override ${key}.input must contain only text or image`);
52
+ const cost = raw.cost;
53
+ let normalizedCost;
54
+ if (cost !== undefined) {
55
+ if (!cost || typeof cost !== "object" || Array.isArray(cost))
56
+ throw new Error(`route override ${key}.cost must be an object`);
57
+ const c = cost;
58
+ normalizedCost = {
59
+ input: finiteAtLeast(c.input, 0, `route override ${key}.cost.input`),
60
+ output: finiteAtLeast(c.output, 0, `route override ${key}.cost.output`),
61
+ ...(c.cacheRead === undefined ? {} : {
62
+ cacheRead: finiteAtLeast(c.cacheRead, 0, `route override ${key}.cost.cacheRead`),
63
+ }),
64
+ ...(c.cacheWrite === undefined ? {} : {
65
+ cacheWrite: finiteAtLeast(c.cacheWrite, 0, `route override ${key}.cost.cacheWrite`),
66
+ }),
67
+ };
68
+ }
69
+ const contextWindow = raw.context_window;
70
+ const maxTokens = raw.max_tokens;
71
+ result[key] = {
72
+ ...(normalizedCost ? { cost: normalizedCost } : {}),
73
+ ...(contextWindow === undefined ? {} : {
74
+ context_window: positiveInteger(contextWindow, `route override ${key}.context_window`),
75
+ }),
76
+ ...(maxTokens === undefined ? {} : {
77
+ max_tokens: positiveInteger(maxTokens, `route override ${key}.max_tokens`),
78
+ }),
79
+ ...(input === undefined ? {} : { input: [...new Set(input)] }),
80
+ };
81
+ }
82
+ return result;
83
+ }
84
+ function providerExtensions(value) {
85
+ if (value === undefined)
86
+ return [];
87
+ if (!Array.isArray(value) || value.some((item) => typeof item !== "string" || !item.trim()))
88
+ throw new Error("pi.provider_extensions must be an array of non-empty file paths");
89
+ return [...new Set(value.map((item) => String(item).trim()))];
90
+ }
23
91
  /** Load a policy file, filling defaults for missing keys. Throws on invalid JSON. */
24
92
  export function loadPolicy(path) {
25
93
  if (!existsSync(path))
@@ -58,6 +126,8 @@ export function loadPolicy(path) {
58
126
  judge_model: raw.judge_model ?? null,
59
127
  strategist_model: raw.strategist_model ?? null,
60
128
  max_usd_per_m: num(raw.max_usd_per_m, d.max_usd_per_m),
129
+ route_overrides: routeOverrides(raw.route_overrides),
130
+ pi: { provider_extensions: providerExtensions(raw.pi?.provider_extensions) },
61
131
  };
62
132
  }
63
133
  /** Match either a logical vendor/model id or its concrete route provider. */
package/dist/routes.js CHANGED
@@ -19,6 +19,21 @@ export function legacyOpenRouterRoute(meritId) {
19
19
  export function routeKey(route) {
20
20
  return `${route.provider}:${route.modelId}`;
21
21
  }
22
+ /** Apply user metadata only to the exact provider route it names. */
23
+ export function applyRouteOverrides(routes, overrides) {
24
+ return routes.map((route) => {
25
+ const override = overrides[routeKey(route)];
26
+ if (!override)
27
+ return route;
28
+ return {
29
+ ...route,
30
+ ...(override.cost ? { cost: { ...override.cost } } : {}),
31
+ ...(override.context_window === undefined ? {} : { contextWindow: override.context_window }),
32
+ ...(override.max_tokens === undefined ? {} : { maxTokens: override.max_tokens }),
33
+ ...(override.input ? { input: [...override.input] } : {}),
34
+ };
35
+ });
36
+ }
22
37
  export function pointKey(point) {
23
38
  return point.route ? routeKey(point.route) : `openrouter:${point.model}`;
24
39
  }
@@ -5,10 +5,10 @@ import { fetchCatalog, saveSnapshot } from "./catalog.js";
5
5
  import { paretoFrontier, pickBest, pickFallback } from "./frontier.js";
6
6
  import { JUDGE_PREFS } from "./judge.js";
7
7
  import { loadKey, PiCliChatClient } from "./llm.js";
8
- import { availablePiRoutes, recordedActiveTask, runPiTrial, scorePiRun } from "./pi-trials.js";
8
+ import { availablePiRoutes, PiHarnessAdapter, recordedActiveTask, runPiTrial, scorePiRun } from "./pi-trials.js";
9
9
  import { loadPolicy, providerAllowed } from "./policy.js";
10
10
  import { buildRecommendation } from "./recommend.js";
11
- import { enrichRouteEntry, routeCatalog, routeKey, routeLabel } from "./routes.js";
11
+ import { applyRouteOverrides, enrichRouteEntry, routeCatalog, routeKey, routeLabel } from "./routes.js";
12
12
  import { appendJsonl, paths, readJson, taskKey } from "./store.js";
13
13
  import { pickNext, STRAT_PREFS } from "./strategist.js";
14
14
  import { budgetOk, recordMeritSpend, recordTrialSpend, trialBudgetOk } from "./trials.js";
@@ -20,13 +20,16 @@ function optionalOpenRouterKey() {
20
20
  return null;
21
21
  }
22
22
  }
23
- function configuredRoutes() {
23
+ function configuredRoutes(policy) {
24
24
  const state = readJson(paths.harnessState(), {});
25
- let routes = state.routes?.length ? state.routes : availablePiRoutes();
25
+ let routes = state.routes?.length ? state.routes : availablePiRoutes(undefined, policy.pi.provider_extensions);
26
26
  if (state.currentRoute && !routes.some((route) => routeKey(route) === routeKey(state.currentRoute)))
27
27
  routes = [state.currentRoute, ...routes];
28
+ routes = applyRouteOverrides(routes, policy.route_overrides);
28
29
  const unique = new Map(routes.map((route) => [routeKey(route), route]));
29
- return { routes: [...unique.values()], currentRoute: state.currentRoute ?? null };
30
+ const currentRoute = state.currentRoute
31
+ ? applyRouteOverrides([state.currentRoute], policy.route_overrides)[0] : null;
32
+ return { routes: [...unique.values()], currentRoute };
30
33
  }
31
34
  function selectorKey(selector) {
32
35
  if (!selector)
@@ -111,7 +114,7 @@ export async function runStandaloneTrial(taskFile, rounds) {
111
114
  throw new Error("--rounds must be a positive integer");
112
115
  const policy = loadPolicy(paths.policy());
113
116
  const key = optionalOpenRouterKey();
114
- const { routes, currentRoute } = configuredRoutes();
117
+ const { routes, currentRoute } = configuredRoutes(policy);
115
118
  let cat = await buildCatalog(routes, key);
116
119
  for (const [id, entry] of [...cat]) {
117
120
  if (!entry.route || !providerAllowed(policy, entry.id, entry.route.provider) ||
@@ -123,7 +126,8 @@ export async function runStandaloneTrial(taskFile, rounds) {
123
126
  const initial = resolveInitial(cat, cfg, currentRoute);
124
127
  const judge = resolveMeritRoute(cat, [cfg.judge_model, policy.judge_model, ...JUDGE_PREFS, initial.id], "judge");
125
128
  const strategist = resolveMeritRoute(cat, [cfg.strategist_model, policy.strategist_model, ...STRAT_PREFS, initial.id], "strategist");
126
- const client = new PiCliChatClient(undefined, recordMeritSpend);
129
+ const client = new PiCliChatClient(undefined, recordMeritSpend, policy.pi.provider_extensions);
130
+ const harness = new PiHarnessAdapter(policy.pi.provider_extensions);
127
131
  const { category, benchmarks } = relevantBenchmarks(cfg.task);
128
132
  let benchmarkCandidates = [];
129
133
  if (key) {
@@ -180,7 +184,7 @@ export async function runStandaloneTrial(taskFile, rounds) {
180
184
  : " active trace unavailable; running the initial route in a fresh Pi session");
181
185
  const outcome = observed
182
186
  ? await scorePiRun(key ?? "", judge.route, cfg.task, cfg.eval, entry.id, entry, observed, "trace", [], client)
183
- : await runPiTrial(key ?? "", judge.route, cfg.task, cfg.eval, entry.id, entry, cfg.cwd ?? process.cwd(), [], [], undefined, client);
187
+ : await runPiTrial(key ?? "", judge.route, cfg.task, cfg.eval, entry.id, entry, cfg.cwd ?? process.cwd(), [], [], harness, client);
184
188
  tried.add(keyForRoute);
185
189
  results.push(outcome.point);
186
190
  // Reusing the active Pi answer is observation, not a new candidate call.
@@ -37,7 +37,7 @@ const JOB_DIR = join(HOME, "watch", "jobs");
37
37
  const ENGINE_FILE = join(fileURLToPath(new URL("..", import.meta.url)), "dist", "cli.js");
38
38
  const PROGRESS_PREFIX = "[openmerit/progress] ";
39
39
 
40
- interface TrialProgress {
40
+ interface TrialRunProgress {
41
41
  phase: "start" | "complete";
42
42
  taskKey: string;
43
43
  model: string;
@@ -53,19 +53,35 @@ interface TrialProgress {
53
53
  changedFiles?: number;
54
54
  }
55
55
 
56
+ interface TrialSelectionProgress {
57
+ phase: "selection";
58
+ taskKey: string;
59
+ eligible: string[];
60
+ excluded: { model: string; reasons: string[] }[];
61
+ }
62
+
63
+ type TrialProgress = TrialRunProgress | TrialSelectionProgress;
64
+
56
65
  function parseTrialProgress(line: string): TrialProgress | null {
57
66
  if (!line.startsWith(PROGRESS_PREFIX)) return null;
58
67
  try {
59
68
  const value = JSON.parse(line.slice(PROGRESS_PREFIX.length)) as TrialProgress;
60
- if ((value.phase !== "start" && value.phase !== "complete") ||
61
- typeof value.taskKey !== "string" || typeof value.model !== "string" ||
69
+ if (typeof value.taskKey !== "string") return null;
70
+ if (value.phase === "selection") {
71
+ if (!Array.isArray(value.eligible) || !Array.isArray(value.excluded) ||
72
+ value.eligible.some((item) => typeof item !== "string") ||
73
+ value.excluded.some((item) => typeof item?.model !== "string" || !Array.isArray(item.reasons) ||
74
+ item.reasons.some((reason) => typeof reason !== "string"))) return null;
75
+ return value;
76
+ }
77
+ if ((value.phase !== "start" && value.phase !== "complete") || typeof value.model !== "string" ||
62
78
  !Number.isInteger(value.index) || !Number.isInteger(value.total) ||
63
79
  value.index < 1 || value.total < value.index) return null;
64
80
  return value;
65
81
  } catch { return null; }
66
82
  }
67
83
 
68
- function completedTrialText(trial: TrialProgress): string {
84
+ function completedTrialText(trial: TrialRunProgress): string {
69
85
  return `${trial.model}: quality ${(trial.score ?? 0).toFixed(2)}, ` +
70
86
  `$${(trial.price ?? 0).toFixed(2)}/M, ${Math.round(trial.latencyMs ?? 0)}ms, ` +
71
87
  `run $${(trial.costUsd ?? 0).toFixed(4)}, tools ${trial.toolCalls ?? 0}` +
@@ -77,6 +93,12 @@ interface ExtensionPolicy {
77
93
  mode?: string;
78
94
  fallback?: { apply_on_error?: boolean };
79
95
  budgets?: { max_trials_per_day?: number; max_usd_per_day?: number };
96
+ route_overrides?: Record<string, {
97
+ cost?: { input?: number; output?: number; cacheRead?: number; cacheWrite?: number };
98
+ context_window?: number;
99
+ max_tokens?: number;
100
+ input?: ("text" | "image")[];
101
+ }>;
80
102
  }
81
103
 
82
104
  function trialBudget(): { summary: string; reason: string | null } {
@@ -237,6 +259,7 @@ function latestSessionTaskKey(ctx: ExtensionContext): string | null {
237
259
  }
238
260
 
239
261
  let settledTaskKey: string | null = null;
262
+ let comparisonsPaused = false;
240
263
 
241
264
  function pendingRecommendations(ctx?: ExtensionContext): Recommendation[] {
242
265
  const pending = latestRecommendations().filter((r) => r.status === "pending");
@@ -295,8 +318,8 @@ let fallbackRoute: ModelRoute | null = null;
295
318
 
296
319
  function routeFor(model: { provider: string; id: string; input?: ("text" | "image")[];
297
320
  cost?: { input?: number; output?: number; cacheRead?: number; cacheWrite?: number };
298
- contextWindow?: number; maxTokens?: number }): ModelRoute {
299
- return {
321
+ contextWindow?: number; maxTokens?: number }, overrides = loadPolicy().route_overrides ?? {}): ModelRoute {
322
+ const route: ModelRoute = {
300
323
  provider: model.provider,
301
324
  modelId: model.id,
302
325
  meritId: canonicalModel(model.provider, model.id),
@@ -308,6 +331,33 @@ function routeFor(model: { provider: string; id: string; input?: ("text" | "imag
308
331
  contextWindow: model.contextWindow,
309
332
  maxTokens: model.maxTokens,
310
333
  };
334
+ const override = overrides[`${route.provider}:${route.modelId}`];
335
+ if (!override) return route;
336
+ const finite = (value: unknown, min: number) =>
337
+ typeof value === "number" && Number.isFinite(value) && value >= min ? value : undefined;
338
+ const integer = (value: unknown) => {
339
+ const normalized = finite(value, 1);
340
+ return normalized !== undefined && Number.isInteger(normalized) ? normalized : undefined;
341
+ };
342
+ const input = Array.isArray(override.input) && override.input.length > 0 &&
343
+ override.input.every((item) => item === "text" || item === "image")
344
+ ? [...new Set(override.input)] : undefined;
345
+ const costInput = finite(override.cost?.input, 0);
346
+ const costOutput = finite(override.cost?.output, 0);
347
+ return {
348
+ ...route,
349
+ ...(costInput === undefined || costOutput === undefined ? {} : { cost: {
350
+ input: costInput,
351
+ output: costOutput,
352
+ cacheRead: finite(override.cost?.cacheRead, 0),
353
+ cacheWrite: finite(override.cost?.cacheWrite, 0),
354
+ }}),
355
+ ...(integer(override.context_window) === undefined ? {} : {
356
+ contextWindow: integer(override.context_window),
357
+ }),
358
+ ...(integer(override.max_tokens) === undefined ? {} : { maxTokens: integer(override.max_tokens) }),
359
+ ...(input ? { input } : {}),
360
+ };
311
361
  }
312
362
 
313
363
  function eligibleRoutes(ctx: ExtensionContext): ModelRoute[] {
@@ -317,8 +367,9 @@ function eligibleRoutes(ctx: ExtensionContext): ModelRoute[] {
317
367
  if (ctx.model && !models.some((model) => model.provider === ctx.model!.provider && model.id === ctx.model!.id))
318
368
  models.unshift(ctx.model);
319
369
  const byRoute = new Map<string, ModelRoute>();
370
+ const overrides = loadPolicy().route_overrides ?? {};
320
371
  for (const model of models) {
321
- const route = routeFor(model);
372
+ const route = routeFor(model, overrides);
322
373
  byRoute.set(`${route.provider}:${route.modelId}`, route);
323
374
  }
324
375
  return [...byRoute.values()];
@@ -347,7 +398,7 @@ function reportHarnessState(ctx: ExtensionContext, settled = false): void {
347
398
  const settledAt = settled ? new Date().toISOString()
348
399
  : prior.sessionFile === sessionFile ? prior.settledAt ?? null : null;
349
400
  writeState({
350
- schemaVersion: 1,
401
+ schemaVersion: 2,
351
402
  currentModel: model,
352
403
  currentRoute,
353
404
  routes: eligibleRoutes(ctx),
@@ -356,6 +407,7 @@ function reportHarnessState(ctx: ExtensionContext, settled = false): void {
356
407
  sessionFile,
357
408
  settledTaskKey,
358
409
  settledAt,
410
+ comparisonsPaused,
359
411
  cwd: ctx.cwd,
360
412
  updatedAt: new Date().toISOString(),
361
413
  });
@@ -465,6 +517,34 @@ function ensureModel(pi: ExtensionAPI, ctx: ExtensionContext, openrouterId: stri
465
517
  return registry.find?.(INJECT_PROVIDER, openrouterId);
466
518
  }
467
519
 
520
+ async function doctorSummary(): Promise<{ text: string; ok: boolean }> {
521
+ if (!existsSync(ENGINE_FILE)) return {
522
+ text: "trial engine missing; reinstall OpenMerit or run npm ci in the checkout",
523
+ ok: false,
524
+ };
525
+ const child = spawn(process.execPath, [ENGINE_FILE, "doctor", "--json"], {
526
+ env: process.env, stdio: ["ignore", "pipe", "pipe"],
527
+ });
528
+ let stdout = "";
529
+ let stderr = "";
530
+ child.stdout.on("data", (chunk: Buffer) => { stdout += chunk.toString(); });
531
+ child.stderr.on("data", (chunk: Buffer) => { stderr += chunk.toString(); });
532
+ const code = await new Promise<number>((resolve, reject) => {
533
+ child.on("error", reject);
534
+ child.on("close", (value) => resolve(value ?? 1));
535
+ });
536
+ try {
537
+ const report = JSON.parse(stdout) as { ok?: boolean; checks?: {
538
+ id?: string; status?: string; message?: string;
539
+ }[] };
540
+ const lines = (report.checks ?? []).map((item) =>
541
+ `${(item.status ?? "?").toUpperCase().padEnd(4)} ${item.id ?? "check"}: ${item.message ?? ""}`);
542
+ return { text: lines.join("\n") || "doctor returned no checks", ok: report.ok === true };
543
+ } catch {
544
+ return { text: stderr.trim().slice(0, 500) || stdout.trim().slice(0, 500) || `doctor exited ${code}`, ok: false };
545
+ }
546
+ }
547
+
468
548
  export default function openmerit(pi: ExtensionAPI) {
469
549
  let pollTimer: ReturnType<typeof setInterval> | null = null;
470
550
  let trialJob: ChildProcessByStdio<null, Readable, Readable> | null = null;
@@ -472,8 +552,8 @@ export default function openmerit(pi: ExtensionAPI) {
472
552
  key: string; cwd: string; bytes: number }[] = [];
473
553
  const startedJobs = new Set<string>();
474
554
  let activeSession: string | null = null;
475
- let progress: { sessionFile: string; taskKey: string; current: TrialProgress | null;
476
- completed: TrialProgress[] } | null = null;
555
+ let progress: { sessionFile: string; taskKey: string; current: TrialRunProgress | null;
556
+ completed: TrialRunProgress[]; eligible: string[]; excluded: TrialSelectionProgress["excluded"] } | null = null;
477
557
 
478
558
  function stopTrialJob(): void {
479
559
  if (!trialJob) return;
@@ -501,14 +581,20 @@ export default function openmerit(pi: ExtensionAPI) {
501
581
  detached: process.platform !== "win32",
502
582
  });
503
583
  trialJob = child;
504
- progress = { sessionFile: job.sessionFile, taskKey: job.key, current: null, completed: [] };
584
+ progress = { sessionFile: job.sessionFile, taskKey: job.key, current: null, completed: [],
585
+ eligible: [], excluded: [] };
505
586
  if (ctx.hasUI) ctx.ui.notify(`openmerit: comparing models for task ${job.key} (A=${job.model})`, "info");
506
587
  let stdout = "";
507
588
  let stderr = "";
508
589
  function onLine(line: string): void {
509
590
  const trial = parseTrialProgress(line);
510
591
  if (trial && job.sessionFile === activeSession && progress?.taskKey === trial.taskKey) {
511
- if (trial.phase === "start") {
592
+ if (trial.phase === "selection") {
593
+ progress.eligible = trial.eligible;
594
+ progress.excluded = trial.excluded;
595
+ if (ctx.hasUI) ctx.ui.notify(`openmerit: ${trial.eligible.length} routes eligible; ` +
596
+ `${trial.excluded.length} skipped (/openmerit for reasons)`, "info");
597
+ } else if (trial.phase === "start") {
512
598
  progress.current = trial;
513
599
  if (ctx.hasUI) ctx.ui.notify(
514
600
  `openmerit: trial ${trial.index}/${trial.total} comparing ${trial.model}` +
@@ -551,11 +637,13 @@ export default function openmerit(pi: ExtensionAPI) {
551
637
  });
552
638
  }
553
639
 
554
- function queueSettledTask(ctx: ExtensionContext): void {
640
+ function queueSettledTask(ctx: ExtensionContext, force = false): string {
641
+ if (comparisonsPaused && !force) return "comparisons are paused";
555
642
  const sessionFile = ctx.sessionManager.getSessionFile();
556
643
  const model = ctx.model ? canonicalModel(ctx.model.provider, ctx.model.id) : null;
557
644
  const route = ctx.model ? routeFor(ctx.model) : null;
558
- if (!sessionFile || !model || !route || !settledTaskKey || !existsSync(sessionFile)) return;
645
+ if (!sessionFile || !model || !route || !settledTaskKey || !existsSync(sessionFile))
646
+ return "no saved completed task is available";
559
647
  // The alpha compares completed response tasks. The engine validates images,
560
648
  // tool use and the baseline answer before any provider calls.
561
649
  let completed = false;
@@ -568,19 +656,20 @@ export default function openmerit(pi: ExtensionAPI) {
568
656
  userTaskKeyFromAnswer(entry.message.content)) completed = true;
569
657
  } catch { /* partial session line */ }
570
658
  }
571
- if (!completed) return;
659
+ if (!completed) return "the latest task has not completed successfully";
572
660
  const settledAt = new Date().toISOString();
573
661
  const bytes = readFileSync(sessionFile).length;
574
662
  const marker = `${sessionFile}:${settledTaskKey}:${model}:${bytes}`;
575
- if (startedJobs.has(marker)) return;
663
+ if (startedJobs.has(marker)) return "this task is already queued or was compared in this Pi session";
576
664
  startedJobs.add(marker);
577
665
  const budget = trialBudget();
578
666
  if (budget.reason) {
579
667
  if (ctx.hasUI) ctx.ui.notify(`openmerit: comparison skipped: ${budget.reason}. ${budget.summary}`, "warning");
580
- return;
668
+ return budget.reason;
581
669
  }
582
670
  queuedJobs.push({ sessionFile, settledAt, model, route, key: settledTaskKey, cwd: ctx.cwd, bytes });
583
671
  startNextJob(ctx);
672
+ return "comparison queued";
584
673
  }
585
674
 
586
675
  function userTaskKeyFromAnswer(content: unknown): boolean {
@@ -667,11 +756,12 @@ export default function openmerit(pi: ExtensionAPI) {
667
756
  restoreFallback();
668
757
  try {
669
758
  const st = JSON.parse(readFileSync(STATE_FILE, "utf8")) as {
670
- sessionFile?: string | null; settledTaskKey?: string | null;
759
+ sessionFile?: string | null; settledTaskKey?: string | null; comparisonsPaused?: boolean;
671
760
  };
761
+ comparisonsPaused = st.comparisonsPaused === true;
672
762
  settledTaskKey = st.sessionFile === ctx.sessionManager.getSessionFile()
673
763
  ? st.settledTaskKey ?? null : null;
674
- } catch { settledTaskKey = null; }
764
+ } catch { settledTaskKey = null; comparisonsPaused = false; }
675
765
  reportHarnessState(ctx);
676
766
  await surface(ctx);
677
767
  if (pollTimer) clearInterval(pollTimer);
@@ -748,7 +838,7 @@ export default function openmerit(pi: ExtensionAPI) {
748
838
  });
749
839
 
750
840
  pi.registerCommand("openmerit", {
751
- description: "OpenMerit: status, pending recommendations, apply/dismiss",
841
+ description: "OpenMerit: status, compare, pause/resume, doctor, apply/dismiss",
752
842
  handler: async (args, ctx) => {
753
843
  const pending = pendingRecommendations(ctx);
754
844
  const prior = priorExactTaskSuggestion(ctx);
@@ -782,6 +872,37 @@ export default function openmerit(pi: ExtensionAPI) {
782
872
  }
783
873
  return;
784
874
  }
875
+ if (sub === "pause") {
876
+ comparisonsPaused = true;
877
+ queuedJobs.length = 0;
878
+ stopTrialJob();
879
+ startedJobs.clear();
880
+ progress = null;
881
+ reportHarnessState(ctx);
882
+ ctx.ui.notify("openmerit: automatic comparisons paused; `/openmerit compare` remains available", "info");
883
+ return;
884
+ }
885
+ if (sub === "resume") {
886
+ comparisonsPaused = false;
887
+ reportHarnessState(ctx);
888
+ ctx.ui.notify("openmerit: automatic comparisons resumed for future completed tasks", "info");
889
+ return;
890
+ }
891
+ if (sub === "compare") {
892
+ settledTaskKey = latestSessionTaskKey(ctx);
893
+ const result = queueSettledTask(ctx, true);
894
+ if (result !== "comparison queued") ctx.ui.notify(`openmerit: ${result}`, "warning");
895
+ return;
896
+ }
897
+ if (sub === "doctor") {
898
+ try {
899
+ const result = await doctorSummary();
900
+ ctx.ui.notify(`openmerit doctor:\n${result.text}`, result.ok ? "info" : "warning");
901
+ } catch (error) {
902
+ ctx.ui.notify(`openmerit doctor failed: ${(error as Error).message}`, "error");
903
+ }
904
+ return;
905
+ }
785
906
 
786
907
  const currentRoute = ctx.model ? routeFor(ctx.model) : undefined;
787
908
  const current = currentRoute ? routeDisplay(currentRoute, currentRoute.meritId) : "unknown";
@@ -791,6 +912,7 @@ export default function openmerit(pi: ExtensionAPI) {
791
912
  ? `running — ${liveProgress.current.model} (trial ${liveProgress.current.index}/${liveProgress.current.total})`
792
913
  : "running — selecting next model"
793
914
  : queuedJobs.length ? `queued (${queuedJobs.length})`
915
+ : comparisonsPaused ? "paused"
794
916
  : latestSessionJob(ctx.sessionManager.getSessionFile() ?? undefined) ?? "not started";
795
917
  const lines = [
796
918
  `current model : ${current}`,
@@ -803,6 +925,11 @@ export default function openmerit(pi: ExtensionAPI) {
803
925
  lines.push("", "models tested:");
804
926
  for (const trial of liveProgress.completed) lines.push(` ${completedTrialText(trial)}`);
805
927
  }
928
+ if (liveProgress?.excluded.length) {
929
+ lines.push("", "routes skipped:");
930
+ for (const route of liveProgress.excluded)
931
+ lines.push(` ${route.model}: ${route.reasons.join("; ")}`);
932
+ }
806
933
  if (prior) lines.push("", suggestionText(prior.recommendation, prior.trials));
807
934
  for (const r of pending.slice(-5)) {
808
935
  lines.push(
@@ -39,6 +39,12 @@ the best model for the task at hand, at the best price, with a vetted fallback.
39
39
  recommendations.
40
40
  - `/openmerit apply` / `/openmerit dismiss` — act on a pending recommendation
41
41
  when automatic application is declined by the policy gate.
42
+ - `/openmerit pause` / `/openmerit resume` — stop or restart automatic task
43
+ comparisons without changing the saved policy.
44
+ - `/openmerit compare` — explicitly compare the latest completed task, even
45
+ while automatic comparisons are paused.
46
+ - `/openmerit doctor` — inspect sanitized installation, route, pricing, state,
47
+ and provider-extension diagnostics inside Pi.
42
48
  - Policy changes (auto-apply thresholds, budgets, provider allow-lists) are
43
49
  made by the user in `~/.openmerit/policy.json`, not by you.
44
50
 
@@ -52,3 +58,6 @@ the best model for the task at hand, at the best price, with a vetted fallback.
52
58
  - The shipped policy is supervised. Swaps happen automatically only after the
53
59
  user opts in and the configured quality/cost guardrails pass; otherwise they
54
60
  remain recommendations for a human to approve.
61
+ - Isolated subprocesses disable extension discovery. Provider-registration
62
+ extensions run only when the user lists their absolute file paths in
63
+ `policy.json`; OpenMerit refuses to load itself recursively.
@@ -1,5 +1,5 @@
1
1
  {
2
- "version": 1,
2
+ "version": 2,
3
3
  "mode": "recommend",
4
4
  "auto_apply": {
5
5
  "enabled": false,
@@ -29,5 +29,9 @@
29
29
  },
30
30
  "judge_model": null,
31
31
  "strategist_model": null,
32
- "max_usd_per_m": 20.0
32
+ "max_usd_per_m": 20.0,
33
+ "route_overrides": {},
34
+ "pi": {
35
+ "provider_extensions": []
36
+ }
33
37
  }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "openmerit",
3
- "version": "0.1.3",
3
+ "version": "0.1.4",
4
4
  "description": "Find better models for each pi task by comparing quality, cost, and latency.",
5
5
  "type": "module",
6
6
  "keywords": [
package/rules.md CHANGED
@@ -16,6 +16,8 @@
16
16
 
17
17
  - **Respect the harness's eligible model pool.** Prefer models already available and configured in the active harness; external catalogs may enrich or expand discovery but must not silently override harness scope or credentials.
18
18
 
19
+ - **Require explicit metadata instead of price guesses.** Missing price, context, output, or modality data for a custom route may be supplied by an exact-route override; never assume a local route is free or copy metadata across providers.
20
+
19
21
  - **Treat public benchmarks as priors, not proof.** Benchmarks help shortlist candidates, but merit comes from trials on the user's actual task.
20
22
 
21
23
  - **Keep four integration boundaries distinct.** Model execution (`ModelProviderAdapter`), agent execution (`HarnessAdapter`), incoming traces (`ObservationSource`), and outgoing telemetry (`EventSink`) solve different problems and must not be coupled.
@@ -30,6 +32,8 @@
30
32
 
31
33
  - **Default to safe, explicit trials.** Candidate tool access stays off unless deliberately allowed, trials remain isolated, and temporary workspaces are cleaned up because trying a model must not expose or damage a user's project by surprise.
32
34
 
35
+ - **Allowlist subprocess extensions.** Keep Pi extension discovery disabled in isolated trials and load only explicitly named provider-registration files; never recursively load OpenMerit or inherit unrelated harness extensions.
36
+
33
37
  - **Keep the shadow track isolated, not necessarily concurrent.** Candidate trials may run sequentially to respect cost, rate, and safety limits while remaining separate from the user's live task.
34
38
 
35
39
  - **Treat local JSONL as the default integration, not a lock-in.** Local traces and events should work without an external service; systems such as Langfuse can later plug in as observation sources or event sinks.