openmerit 0.1.0 → 0.1.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -105,15 +105,23 @@ Keep that checkout available because pi loads a local package from its path.
105
105
  A ready configuration prints `{"status":"ready","provider":"openrouter",...}`.
106
106
  Older `model_search/.env` files are not read by the trial engine.
107
107
 
108
+ OpenRouter is currently required for candidate runs and catalog discovery.
109
+ Judge and strategist calls can use OpenAI or Gemini directly, but require
110
+ `OPENMERIT_DIRECT_PROVIDER=openai` or
111
+ `OPENMERIT_DIRECT_PROVIDER=google` plus the corresponding provider key;
112
+ OpenRouter model IDs remain on OpenRouter by default. The harness and
113
+ provider interfaces are extension seams for future integrations, not a
114
+ claim that other harnesses already work.
115
+
108
116
  3. Verify the extension appears in `pi list`. The instruction file
109
117
  [`instructions/OPENMERIT.md`](instructions/OPENMERIT.md) can be added to a
110
118
  pi project's AGENTS.md for agent context, but the extension does
111
119
  not require it.
112
120
 
113
- 4. Start pi in one terminal with an OpenRouter model and tools disabled:
121
+ 4. Start pi in one terminal with an OpenRouter model:
114
122
 
115
123
  ```bash
116
- pi --provider openrouter --model openai/gpt-4o-mini --no-tools
124
+ pi --provider openrouter --model openai/gpt-4o-mini
117
125
  ```
118
126
 
119
127
  For a first text task, ask: “Give the shortest valid word ladder from cat
@@ -126,6 +134,10 @@ Keep that checkout available because pi loads a local package from its path.
126
134
  switch. Then send a second message to see which model actually handles it.
127
135
  You do not need a watcher terminal or the standalone `trial` command.
128
136
 
137
+ Candidate comparisons have no tools by default, even when the observed task
138
+ used tools. See **Alpha boundaries** before explicitly enabling candidate
139
+ tools.
140
+
129
141
  To opt in to automatic swaps, edit `~/.openmerit/policy.json`, set
130
142
  `"mode": "auto"` and `"auto_apply.enabled": true`, review the score-gain
131
143
  and price-ratio thresholds, and run the next task.
@@ -133,7 +145,7 @@ Keep that checkout available because pi loads a local package from its path.
133
145
  ### Try an invoice-to-JSON task
134
146
 
135
147
  Use the same one-terminal setup. Start pi with a vision-capable model, for
136
- example `openai/gpt-4o-mini`, and `--no-tools`. In pi, type `@` to select
148
+ example `openai/gpt-4o-mini`. In pi, type `@` to select
137
149
  [`benchmark/invoice_ocr/data/invoice_01_row_2.jpg`](benchmark/invoice_ocr/data/invoice_01_row_2.jpg)
138
150
  and paste the **entire, unchanged** text from
139
151
  [`examples/invoice-prompt.txt`](examples/invoice-prompt.txt) into the same
@@ -176,8 +188,8 @@ for local experiments while keeping `budgets.max_usd_per_day` as a conservative
176
188
  candidate-spend threshold.
177
189
 
178
190
  If no comparison starts after Pi settles, confirm that pi loaded the extension
179
- (`pi list`), the task finished, the pi session is saved (do not use
180
- `--no-session`), and the task used no tools. Image candidates must advertise
191
+ (`pi list`), the task finished, and the pi session is saved (do not use
192
+ `--no-session`). Image candidates must advertise
181
193
  image input in OpenRouter's current model catalog. The extension queues
182
194
  completed tasks from its current session and runs one comparison at a time.
183
195
  Closing or switching the session cancels the active job.
@@ -191,18 +203,35 @@ remains for controlled text-task runs from a task JSON file.
191
203
 
192
204
  - With the extension installed and an OpenRouter key configured, each supported
193
205
  settled task can start comparison calls automatically. Candidate, judge, and
194
- strategist requests send the task text, attached images, and candidate output
195
- to OpenRouter and can incur charges. Use non-sensitive test tasks and
196
- conservative account limits while evaluating this alpha.
197
- - The extension queues tasks from its current pi session, compares one at a
198
- time, and skips tasks that use tools. Candidate sessions replay the latest
199
- text and images, not earlier conversation context or workspace changes.
206
+ strategist requests send the task text, attached files or images, and
207
+ candidate output to OpenRouter and can incur charges. Use non-sensitive test
208
+ tasks and conservative account limits while evaluating this alpha.
209
+ - The extension queues tasks from its current pi session and compares one at a
210
+ time. Candidate runs use isolated temporary copies of the original working
211
+ directory and have no Pi tools by default. Their temporary changes are
212
+ discarded and recorded as counts, and raw Pi JSON event streams are saved
213
+ under `~/.openmerit/traces/trials/`.
214
+ Candidate sessions replay the task text and images, not earlier conversation
215
+ context or workspace changes.
216
+ - Set `OPENMERIT_PI_TRIAL_TOOLS` to an explicit comma-separated allowlist such
217
+ as `read,grep,find,ls` to let candidate models use tools; `none` keeps them
218
+ disabled. Any tool access is an advanced opt-in: the copied working directory
219
+ prevents ordinary project writes from touching the original, but it is not an
220
+ OS sandbox. Tools may accept absolute paths, and shell tools may access the
221
+ network. Keep the no-tools default for untrusted or sensitive projects.
222
+ - File attachments that Pi records as `<file name="…">` are copied into every
223
+ candidate sandbox and passed back to Pi as `@` file inputs. This covers PDFs,
224
+ CSVs, spreadsheets, and other files that Pi can open; OpenMerit does not
225
+ implement a separate parser for them.
200
226
  - `ledger.json` counts reported candidate-trial spend and the daily number of
201
227
  trials. It does not yet include rubric, judge, or strategist request cost,
202
228
  and `max_usd_per_trial` is not yet a hard provider-side cap.
203
229
  - Each comparison has one observed baseline plus a small candidate slate and
204
230
  one quality score per answer. Treat recommendations as experimental evidence,
205
231
  not a universal model ranking.
232
+ - Candidate execution and catalog discovery currently use Pi through
233
+ OpenRouter. `HarnessAdapter`, `ModelProvider`, and `CatalogProvider` are
234
+ implementation boundaries for future harness/provider integrations.
206
235
 
207
236
  ## State layout (`~/.openmerit/`)
208
237
 
@@ -213,6 +242,7 @@ remains for controlled text-task runs from a task JSON file.
213
242
  | `recommendations.jsonl` | shared append-only channel: session-bound swap recommendations + status updates |
214
243
  | `trials.jsonl` | every model trial point (source, score, price, latency) |
215
244
  | `observations.jsonl` | task observations extracted from session traces |
245
+ | `traces/trials/*.jsonl` | raw Pi JSON event streams for candidate trials |
216
246
  | `catalog/snapshot.json` + `candidates.json` | catalog snapshot (including input modalities) + new-model queue |
217
247
  | `benchmarks/digest.json` | public-benchmark scores per model (seed + refresh) |
218
248
  | `ledger.json` | daily trial spend (budget enforcement) |
package/dist/catalog.js CHANGED
@@ -1,7 +1,17 @@
1
1
  /** OpenRouter catalog fetch + snapshot diffing: the "new model release" watch. */
2
2
  import { readJson, paths, writeJson } from "./store.js";
3
3
  const OR = "https://openrouter.ai/api/v1";
4
- export async function fetchCatalog(key) {
4
+ export class OpenRouterCatalogProvider {
5
+ key;
6
+ id = "openrouter";
7
+ constructor(key) {
8
+ this.key = key;
9
+ }
10
+ fetchCatalog() { return fetchCatalog(this.key); }
11
+ }
12
+ export async function fetchCatalog(key, provider = "openrouter") {
13
+ if (provider !== "openrouter")
14
+ return new Map();
5
15
  const res = await fetch(`${OR}/models`, {
6
16
  headers: { Authorization: `Bearer ${key}` },
7
17
  });
package/dist/cli.js CHANGED
@@ -43,7 +43,7 @@ function cmdInit() {
43
43
  console.log(`ext. -> ${join(ROOT, "extension", "openmerit.ts")}`);
44
44
  console.log(`pi npm -> pi install npm:openmerit`);
45
45
  console.log(`pi local-> pi install "${ROOT}"`);
46
- console.log("\nNext: set OPENROUTER_API_KEY (env or ~/.openmerit/.env), install one extension source into pi, then start pi; the extension launches comparisons automatically.");
46
+ console.log("\nNext: set OPENROUTER_API_KEY (env or ~/.openmerit/.env), install one extension source into pi, then start pi; the extension launches comparisons automatically with candidate tools disabled by default.");
47
47
  }
48
48
  async function cmdTrial(taskFile, rounds) {
49
49
  const cfg = JSON.parse(readFileSync(taskFile, "utf8"));
package/dist/daemon.js CHANGED
@@ -168,7 +168,8 @@ export async function autoTaskTick(key, policy, settledTask) {
168
168
  emitTrialProgress({ phase: "complete", taskKey: tKey,
169
169
  model: task.model, index: 1, total, score: baseline.point.score,
170
170
  price: baseline.point.price, latencyMs: baseline.point.latencyMs,
171
- costUsd: baseline.costUsd });
171
+ costUsd: baseline.costUsd, toolCalls: baseline.point.toolCalls,
172
+ toolErrors: baseline.point.toolErrors, changedFiles: baseline.point.changedFiles });
172
173
  for (let i = 1; i < total; i++) {
173
174
  if (!budgetOk(policy).ok)
174
175
  break;
@@ -178,18 +179,21 @@ export async function autoTaskTick(key, policy, settledTask) {
178
179
  if (settledTask !== undefined)
179
180
  emitTrialProgress({ phase: "start", taskKey: tKey,
180
181
  model: pick.model, index: i + 1, total });
181
- const { point, costUsd, error } = await runPiTrial(key, judge, task.task, rubric, pick.model, cat.get(pick.model), task.cwd, task.images);
182
+ const { point, costUsd, error } = await runPiTrial(key, judge, task.task, rubric, pick.model, cat.get(pick.model), task.cwd, task.images, task.files);
182
183
  tried.add(pick.model);
183
184
  points.push(point);
184
185
  recordTrialSpend(costUsd);
185
186
  appendJsonl(paths.trials(), { ...point, taskKey: tKey, sessionFile: task.sessionFile });
186
187
  if (error)
187
188
  failedVendors.add(pick.model.split("/")[0]);
188
- console.log(`[openmerit] task ${tKey}: ${pick.model} score=${point.score.toFixed(2)} cost=$${costUsd.toFixed(4)}${error ? ` error=${error.slice(0, 80)}` : ""}`);
189
+ console.log(`[openmerit] task ${tKey}: ${pick.model} score=${point.score.toFixed(2)} cost=$${costUsd.toFixed(4)} ` +
190
+ `tools=${point.toolCalls ?? 0} toolErrors=${point.toolErrors ?? 0} changedFiles=${point.changedFiles ?? 0}` +
191
+ `${error ? ` error=${error.slice(0, 80)}` : ""}`);
189
192
  if (settledTask !== undefined)
190
193
  emitTrialProgress({ phase: "complete", taskKey: tKey,
191
194
  model: pick.model, index: i + 1, total, score: point.score, price: point.price,
192
- latencyMs: point.latencyMs, costUsd, error: error?.slice(0, 120) });
195
+ latencyMs: point.latencyMs, costUsd, error: error?.slice(0, 120),
196
+ toolCalls: point.toolCalls, toolErrors: point.toolErrors, changedFiles: point.changedFiles });
193
197
  }
194
198
  const rec = buildRecommendation(tKey, task.task.slice(0, 120), task.model, points, policy, task.sessionFile);
195
199
  if (rec) {
@@ -0,0 +1 @@
1
+ export {};
package/dist/llm.js CHANGED
@@ -1,4 +1,4 @@
1
- /** Minimal OpenRouter chat client for catalog and trial requests. */
1
+ /** Minimal provider clients for judge and strategist requests. */
2
2
  import { existsSync, readFileSync } from "node:fs";
3
3
  import { paths } from "./store.js";
4
4
  const OR = "https://openrouter.ai/api/v1";
@@ -20,9 +20,9 @@ export function loadKey() {
20
20
  throw new Error("missing OPENROUTER_API_KEY (env or ~/.openmerit/.env)");
21
21
  return key;
22
22
  }
23
- async function request(key, path, body, tries = 3) {
23
+ async function request(key, path, body, baseUrl = OR, tries = 3) {
24
24
  for (let attempt = 0; attempt < tries; attempt++) {
25
- const res = await fetch(OR + path, {
25
+ const res = await fetch(baseUrl + path, {
26
26
  method: body === undefined ? "GET" : "POST",
27
27
  headers: {
28
28
  Authorization: `Bearer ${key}`,
@@ -41,26 +41,62 @@ async function request(key, path, body, tries = 3) {
41
41
  }
42
42
  throw new Error("unreachable");
43
43
  }
44
+ class OpenAICompatibleProvider {
45
+ id;
46
+ baseUrl;
47
+ key;
48
+ constructor(id, baseUrl, key) {
49
+ this.id = id;
50
+ this.baseUrl = baseUrl;
51
+ this.key = key;
52
+ }
53
+ model(model) { return this.id === "openrouter" ? model : model.split("/", 2).at(-1) ?? model; }
54
+ async chat(model, prompt, maxTokens, temperature) {
55
+ const r = await request(this.key, "/chat/completions", { model: this.model(model), messages: [{ role: "user", content: prompt }], max_tokens: maxTokens, temperature }, this.baseUrl);
56
+ return { content: r.choices[0]?.message?.content ?? "", usage: r.usage ?? {} };
57
+ }
58
+ async chatWithImages(model, prompt, images, maxTokens) {
59
+ const content = [{ type: "text", text: prompt }, ...images.map((image) => ({ type: "image_url", image_url: { url: `data:${image.mimeType};base64,${image.data}` } }))];
60
+ const r = await request(this.key, "/chat/completions", { model: this.model(model), messages: [{ role: "user", content }], max_tokens: maxTokens, temperature: 0 }, this.baseUrl);
61
+ return { content: r.choices[0]?.message?.content ?? "", usage: r.usage ?? {} };
62
+ }
63
+ }
64
+ class GeminiProvider {
65
+ key;
66
+ id = "google";
67
+ constructor(key) {
68
+ this.key = key;
69
+ }
70
+ async call(model, prompt, images = [], maxTokens) {
71
+ const name = model.split("/", 2).at(-1) ?? model;
72
+ const parts = [{ text: prompt }, ...images.map((image) => ({ inline_data: { mime_type: image.mimeType, data: image.data } }))];
73
+ const res = await fetch(`https://generativelanguage.googleapis.com/v1beta/models/${encodeURIComponent(name)}:generateContent?key=${encodeURIComponent(this.key)}`, {
74
+ method: "POST", headers: { "Content-Type": "application/json" },
75
+ body: JSON.stringify({ contents: [{ role: "user", parts }], generationConfig: { maxOutputTokens: maxTokens, temperature: images.length ? 0 : undefined } }),
76
+ });
77
+ if (!res.ok)
78
+ throw new Error(`Google Gemini HTTP ${res.status}: ${(await res.text()).slice(0, 300)}`);
79
+ const raw = await res.json();
80
+ return { content: raw.candidates?.[0]?.content?.parts?.map((p) => p.text ?? "").join("") ?? "", usage: { prompt_tokens: raw.usageMetadata?.promptTokenCount, completion_tokens: raw.usageMetadata?.candidatesTokenCount, total_tokens: raw.usageMetadata?.totalTokenCount } };
81
+ }
82
+ chat(model, prompt, maxTokens) { return this.call(model, prompt, [], maxTokens); }
83
+ chatWithImages(model, prompt, images, maxTokens) { return this.call(model, prompt, images, maxTokens); }
84
+ }
85
+ export function providerFor(model, fallbackKey) {
86
+ const id = model.split("/", 1)[0] || "openrouter";
87
+ // OpenRouter model IDs are vendor/model (for example openai/gpt-4o-mini).
88
+ // Direct providers require an explicit opt-in to avoid routing those IDs
89
+ // away from OpenRouter when users have multiple credentials configured.
90
+ if (id === "openai" && process.env.OPENMERIT_DIRECT_PROVIDER === "openai" && process.env.OPENAI_API_KEY)
91
+ return new OpenAICompatibleProvider(id, "https://api.openai.com/v1", process.env.OPENAI_API_KEY);
92
+ if (id === "google" && process.env.OPENMERIT_DIRECT_PROVIDER === "google" && process.env.GEMINI_API_KEY)
93
+ return new GeminiProvider(process.env.GEMINI_API_KEY);
94
+ return new OpenAICompatibleProvider("openrouter", OR, fallbackKey);
95
+ }
44
96
  export async function chat(key, model, prompt, maxTokens, temperature) {
45
- const r = await request(key, "/chat/completions", {
46
- model,
47
- messages: [{ role: "user", content: prompt }],
48
- max_tokens: maxTokens,
49
- temperature,
50
- });
51
- return { content: r.choices[0]?.message?.content ?? "", usage: r.usage ?? {} };
97
+ return providerFor(model, key).chat(model, prompt, maxTokens, temperature);
52
98
  }
53
99
  /** Send a visual grading request through the same OpenRouter account. */
54
100
  export async function chatWithImages(key, model, prompt, images, maxTokens) {
55
- const r = await request(key, "/chat/completions", {
56
- model,
57
- messages: [{ role: "user", content: [
58
- { type: "text", text: prompt },
59
- ...images.map((image) => ({ type: "image_url",
60
- image_url: { url: `data:${image.mimeType};base64,${image.data}` } })),
61
- ] }],
62
- max_tokens: maxTokens,
63
- temperature: 0,
64
- });
65
- return { content: r.choices[0]?.message?.content ?? "", usage: r.usage ?? {} };
101
+ return providerFor(model, key).chatWithImages(model, prompt, images, maxTokens);
66
102
  }
package/dist/pi-trials.js CHANGED
@@ -1,12 +1,13 @@
1
1
  /** Run a task through pi for each model, using pi's JSON event stream as the measurement source. */
2
2
  import { spawn, spawnSync } from "node:child_process";
3
- import { existsSync, mkdtempSync, readFileSync, rmSync, writeFileSync } from "node:fs";
3
+ import { cpSync, existsSync, mkdirSync, mkdtempSync, readdirSync, readFileSync, rmSync, statSync, writeFileSync } from "node:fs";
4
4
  import { tmpdir } from "node:os";
5
- import { join } from "node:path";
5
+ import { dirname, join, relative } from "node:path";
6
+ import { createHash } from "node:crypto";
6
7
  import { judge, judgeWithImages } from "./judge.js";
7
8
  import { knownInvoiceScore } from "./invoice-eval.js";
8
9
  import { paths, readJson } from "./store.js";
9
- import { sessionImages, taskInputKey } from "./task-input.js";
10
+ import { sessionFiles, sessionImages, taskInputKey } from "./task-input.js";
10
11
  function answerText(content) {
11
12
  if (typeof content === "string")
12
13
  return content;
@@ -15,6 +16,59 @@ function answerText(content) {
15
16
  return content.filter((x) => !!x && typeof x === "object" && x.type === "text")
16
17
  .map((x) => x.text ?? "").join("\n");
17
18
  }
19
+ /** Pi-backed harness adapter. Future harnesses can implement the same seam. */
20
+ export class PiHarnessAdapter {
21
+ runTask(model, input) {
22
+ return executePiTask(model, input.task, input.cwd, input.images, input.files);
23
+ }
24
+ }
25
+ const defaultHarness = new PiHarnessAdapter();
26
+ const KNOWN_TRIAL_TOOLS = new Set([
27
+ "read", "grep", "find", "ls", "bash", "powershell", "edit", "write",
28
+ ]);
29
+ /**
30
+ * Candidate runs have no tools by default. Users may explicitly opt into a
31
+ * list, but a copied cwd is not an OS sandbox for absolute paths or network.
32
+ */
33
+ export function piTrialToolArgs(setting = process.env.OPENMERIT_PI_TRIAL_TOOLS) {
34
+ if (!setting?.trim())
35
+ return ["--no-tools"];
36
+ const requested = [...new Set(setting.split(",").map((tool) => tool.trim()).filter(Boolean))];
37
+ if (requested.length === 1 && requested[0] === "none")
38
+ return ["--no-tools"];
39
+ const unknown = requested.filter((tool) => !KNOWN_TRIAL_TOOLS.has(tool));
40
+ if (requested.length === 0 || unknown.length > 0) {
41
+ throw new Error(`invalid OPENMERIT_PI_TRIAL_TOOLS${unknown.length ? `: ${unknown.join(", ")}` : ""}`);
42
+ }
43
+ return ["--tools", requested.join(",")];
44
+ }
45
+ function fileSnapshot(root) {
46
+ const out = new Map();
47
+ function walk(dir) {
48
+ for (const name of readdirSync(dir)) {
49
+ if (name === ".git" || name === "node_modules" || name === ".openmerit")
50
+ continue;
51
+ const file = join(dir, name);
52
+ const rel = relative(root, file);
53
+ const st = statSync(file);
54
+ if (st.isDirectory())
55
+ walk(file);
56
+ else if (st.isFile() && st.size < 20_000_000) {
57
+ out.set(rel, createHash("sha1").update(readFileSync(file)).digest("hex"));
58
+ }
59
+ }
60
+ }
61
+ walk(root);
62
+ return out;
63
+ }
64
+ function isolatedWorkspace(cwd) {
65
+ const root = mkdtempSync(join(tmpdir(), "openmerit-pi-workspace-"));
66
+ cpSync(cwd, root, { recursive: true, filter: (src) => {
67
+ const rel = relative(cwd, src);
68
+ return !rel.split("/").some((part) => part === ".git" || part === "node_modules" || part === ".openmerit");
69
+ } });
70
+ return root;
71
+ }
18
72
  /** Use model A's completed result from the active pi session when it matches this task. */
19
73
  export function recordedActiveTask(task, model, images = [], sessionFile, sessionBytes) {
20
74
  const state = sessionFile ? null : readJson(paths.harnessState(), {});
@@ -23,6 +77,8 @@ export function recordedActiveTask(task, model, images = [], sessionFile, sessio
23
77
  return null;
24
78
  let user = null;
25
79
  const assistants = [];
80
+ let toolCalls = 0;
81
+ let toolErrors = 0;
26
82
  for (const line of readFileSync(file).subarray(0, sessionBytes).toString("utf8").split("\n")) {
27
83
  if (!line.trim())
28
84
  continue;
@@ -38,9 +94,17 @@ export function recordedActiveTask(task, model, images = [], sessionFile, sessio
38
94
  if (entry.message.role === "user") {
39
95
  user = entry.message;
40
96
  assistants.length = 0;
97
+ toolCalls = 0;
98
+ toolErrors = 0;
41
99
  }
42
- else if (user && entry.message.role === "assistant")
100
+ else if (user && entry.message.role === "assistant") {
43
101
  assistants.push(entry.message);
102
+ if (Array.isArray(entry.message.content))
103
+ toolCalls += entry.message.content.filter((part) => part?.type === "toolCall").length;
104
+ }
105
+ else if (user && entry.message.role === "toolResult" &&
106
+ entry.message.isError)
107
+ toolErrors++;
44
108
  }
45
109
  const userImages = user ? sessionImages(user.content) : null;
46
110
  if (!user || !userImages || taskInputKey(answerText(user.content), userImages) !==
@@ -59,6 +123,9 @@ export function recordedActiveTask(task, model, images = [], sessionFile, sessio
59
123
  costUsd: matching.reduce((n, m) => n + (m.usage?.cost?.total ?? 0), 0),
60
124
  latencyMs: Math.max(0, last - first),
61
125
  errors: 0,
126
+ toolCalls,
127
+ toolErrors,
128
+ changedFiles: 0,
62
129
  };
63
130
  }
64
131
  /** Read the exact latest completed text task from the pi extension's session marker. */
@@ -91,19 +158,20 @@ export function settledActiveTask(snapshot) {
91
158
  Array.isArray(entry.message.content) && entry.message.content.some((b) => b?.type === "toolCall"))
92
159
  usedTools = true;
93
160
  }
94
- if (!userContent || usedTools)
161
+ if (!userContent)
95
162
  return null;
96
163
  const images = sessionImages(userContent);
97
164
  if (!images)
98
165
  return null;
99
166
  const task = answerText(userContent);
100
- if (!task.trim() || taskInputKey(task, images) !== st.settledTaskKey)
167
+ const files = sessionFiles(task);
168
+ if (!task.trim() || taskInputKey(task, images, files) !== st.settledTaskKey)
101
169
  return null;
102
170
  const run = recordedActiveTask(task, st.currentModel, images, st.sessionFile, st.sessionBytes);
103
171
  if (!run)
104
172
  return null;
105
173
  return { task, images, model: st.currentModel, run, sessionFile: st.sessionFile,
106
- settledAt: st.settledAt, sessionBytes: st.sessionBytes, cwd: st.cwd ?? process.cwd() };
174
+ settledAt: st.settledAt, sessionBytes: st.sessionBytes, cwd: st.cwd ?? process.cwd(), usedTools, files };
107
175
  }
108
176
  /** Limit candidates to models this installed pi can actually select. */
109
177
  export function availablePiModels() {
@@ -115,7 +183,7 @@ export function availablePiModels() {
115
183
  .filter((id) => !!id && !id.startsWith("~")));
116
184
  }
117
185
  /** Each invocation is a fresh pi run, so model B does not inherit model A's answer. */
118
- export async function executePiTask(model, task, cwd = process.cwd(), images = []) {
186
+ export async function executePiTask(model, task, cwd = process.cwd(), images = [], files = []) {
119
187
  const imageDir = images.length ? mkdtempSync(join(tmpdir(), "openmerit-pi-image-")) : null;
120
188
  const suffix = {
121
189
  "image/jpeg": "jpg", "image/png": "png", "image/webp": "webp", "image/gif": "gif",
@@ -127,73 +195,109 @@ export async function executePiTask(model, task, cwd = process.cwd(), images = [
127
195
  writeFileSync(file, Buffer.from(image.data, "base64"));
128
196
  return `@${file}`;
129
197
  });
130
- const child = spawn("pi", [
131
- "--provider", "openrouter", "--model", model, "--mode", "json",
132
- "--offline", "--no-extensions", "--no-tools", "--print", ...imageArgs, task,
133
- ], { cwd, stdio: ["ignore", "pipe", "pipe"] });
134
- let buffer = "";
135
- let stderr = "";
136
- let answer = "";
137
- let costUsd = 0;
138
- let errors = 0;
139
- let sessionId;
140
- const started = Date.now();
141
- function consume(line) {
142
- if (!line.trim())
143
- return;
144
- let event;
145
- try {
146
- event = JSON.parse(line);
147
- }
148
- catch {
149
- return;
150
- }
151
- if (event.type === "session" && event.id)
152
- sessionId = event.id;
153
- if (event.type !== "message_end" || event.message?.role !== "assistant")
154
- return;
155
- answer = answerText(event.message.content) || answer;
156
- costUsd += event.message.usage?.cost?.total ?? 0;
157
- if (event.message.stopReason === "error")
158
- errors++;
159
- }
160
- child.stdout.on("data", (chunk) => {
161
- buffer += chunk.toString("utf8");
162
- let n;
163
- while ((n = buffer.indexOf("\n")) >= 0) {
164
- consume(buffer.slice(0, n));
165
- buffer = buffer.slice(n + 1);
166
- }
167
- });
168
- child.stderr.on("data", (chunk) => { stderr += chunk.toString("utf8"); });
169
- let exitCode;
198
+ const workspace = isolatedWorkspace(cwd);
170
199
  try {
171
- exitCode = await new Promise((resolve, reject) => {
200
+ const attachmentDir = join(workspace, ".openmerit-attachments");
201
+ mkdirSync(attachmentDir, { recursive: true });
202
+ const attachmentPaths = new Map();
203
+ const fileArgs = files.map((file, i) => {
204
+ const safeName = `${String(i).padStart(3, "0")}-${file.name.replace(/[^A-Za-z0-9._-]/g, "_")}`;
205
+ const target = join(attachmentDir, safeName);
206
+ cpSync(file.path, target);
207
+ attachmentPaths.set(file.path, target);
208
+ return `@${target}`;
209
+ });
210
+ const candidateTask = [...attachmentPaths.entries()].reduce((text, [source, target]) => text.split(source).join(target), task);
211
+ const before = fileSnapshot(workspace);
212
+ const traceFile = paths.trialTrace(`${model}:${task}:${Date.now()}`);
213
+ const traceLines = [];
214
+ const child = spawn("pi", [
215
+ "--provider", "openrouter", "--model", model, "--mode", "json",
216
+ "--offline", "--no-extensions", "--approve", ...piTrialToolArgs(),
217
+ "--print", ...imageArgs, ...fileArgs, candidateTask,
218
+ ], { cwd: workspace, stdio: ["ignore", "pipe", "pipe"] });
219
+ let buffer = "";
220
+ let stderr = "";
221
+ let answer = "";
222
+ let costUsd = 0;
223
+ let errors = 0;
224
+ let sessionId;
225
+ let toolCalls = 0;
226
+ let toolErrors = 0;
227
+ const started = Date.now();
228
+ function consume(line) {
229
+ if (!line.trim())
230
+ return;
231
+ traceLines.push(line);
232
+ let event;
233
+ try {
234
+ event = JSON.parse(line);
235
+ }
236
+ catch {
237
+ return;
238
+ }
239
+ if (event.type === "session" && event.id)
240
+ sessionId = event.id;
241
+ const content = event.message?.content;
242
+ if (event.message?.role === "assistant" && Array.isArray(content))
243
+ toolCalls += content.filter((part) => part?.type === "toolCall").length;
244
+ if (event.message?.role === "toolResult" && event.message.isError)
245
+ toolErrors++;
246
+ if (event.type !== "message_end" || event.message?.role !== "assistant")
247
+ return;
248
+ answer = answerText(event.message.content) || answer;
249
+ costUsd += event.message.usage?.cost?.total ?? 0;
250
+ if (event.message.stopReason === "error")
251
+ errors++;
252
+ }
253
+ child.stdout.on("data", (chunk) => {
254
+ buffer += chunk.toString("utf8");
255
+ let n;
256
+ while ((n = buffer.indexOf("\n")) >= 0) {
257
+ consume(buffer.slice(0, n));
258
+ buffer = buffer.slice(n + 1);
259
+ }
260
+ });
261
+ child.stderr.on("data", (chunk) => { stderr += chunk.toString("utf8"); });
262
+ const exitCode = await new Promise((resolve, reject) => {
172
263
  child.on("error", reject);
173
264
  child.on("close", (code) => resolve(code ?? 1));
174
265
  });
266
+ consume(buffer);
267
+ mkdirSync(dirname(traceFile), { recursive: true });
268
+ writeFileSync(traceFile, traceLines.join("\n") + (traceLines.length ? "\n" : ""));
269
+ const after = fileSnapshot(workspace);
270
+ let changedFiles = 0;
271
+ for (const [file, hash] of after)
272
+ if (before.get(file) !== hash)
273
+ changedFiles++;
274
+ for (const file of before.keys())
275
+ if (!after.has(file))
276
+ changedFiles++;
277
+ if (exitCode !== 0 || !answer.trim()) {
278
+ throw new Error(`pi trial ${model} failed: ${stderr.trim().slice(0, 180) || `exit ${exitCode}, empty answer`}`);
279
+ }
280
+ return { answer, costUsd, latencyMs: Date.now() - started, errors, sessionId,
281
+ toolCalls, toolErrors, changedFiles, traceFile };
175
282
  }
176
283
  finally {
177
284
  if (imageDir)
178
285
  rmSync(imageDir, { recursive: true, force: true });
286
+ rmSync(workspace, { recursive: true, force: true });
179
287
  }
180
- consume(buffer);
181
- if (exitCode !== 0 || !answer.trim()) {
182
- throw new Error(`pi trial ${model} failed: ${stderr.trim().slice(0, 180) || `exit ${exitCode}, empty answer`}`);
183
- }
184
- return { answer, costUsd, latencyMs: Date.now() - started, errors, sessionId };
185
288
  }
186
289
  /** Quality uses the same OpenMerit judge; cost and latency come from pi itself. */
187
- export async function runPiTrial(key, judgeModel, task, rubric, model, entry, cwd = process.cwd(), images = []) {
290
+ export async function runPiTrial(key, judgeModel, task, rubric, model, entry, cwd = process.cwd(), images = [], files = [], harness = defaultHarness) {
188
291
  try {
189
- const run = await executePiTask(model, task, cwd, images);
292
+ const run = await harness.runTask(model, { task, cwd, images, files });
190
293
  return await scorePiRun(key, judgeModel, task, rubric, model, entry, run, "pi_trial", images);
191
294
  }
192
295
  catch (e) {
193
296
  const error = String(e);
194
297
  return {
195
298
  point: { model, score: 0, price: entry?.price ?? 0,
196
- ts: new Date().toISOString(), source: "pi_trial", why: error.slice(0, 160) },
299
+ ts: new Date().toISOString(), source: "pi_trial", toolCalls: 0, toolErrors: 1,
300
+ changedFiles: 0, why: error.slice(0, 160) },
197
301
  costUsd: 0, error,
198
302
  };
199
303
  }
@@ -216,7 +320,9 @@ export async function scorePiRun(key, judgeModel, task, rubric, model, entry, ru
216
320
  point: {
217
321
  model, score, price: entry?.price ?? 0,
218
322
  latencyMs: run.latencyMs, ts: new Date().toISOString(), source,
219
- why: `${why}${run.errors ? `; pi errors: ${run.errors}` : ""}`,
323
+ why: `${why}${run.errors ? `; pi errors: ${run.errors}` : ""}; ` +
324
+ `tools=${run.toolCalls}, toolErrors=${run.toolErrors}, changedFiles=${run.changedFiles}`,
325
+ toolCalls: run.toolCalls, toolErrors: run.toolErrors, changedFiles: run.changedFiles,
220
326
  },
221
327
  costUsd: run.costUsd,
222
328
  sessionId: run.sessionId,
@@ -0,0 +1 @@
1
+ export {};
package/dist/store.js CHANGED
@@ -22,6 +22,7 @@ export const paths = {
22
22
  ledger: () => join(stateDir(), "ledger.json"),
23
23
  watchProcessed: () => join(stateDir(), "watch", "processed.json"),
24
24
  sessionJob: (marker) => join(stateDir(), "watch", "jobs", sha1(marker) + ".json"),
25
+ trialTrace: (id) => join(stateDir(), "traces", "trials", `${sha1(id)}.jsonl`),
25
26
  envFile: () => join(stateDir(), ".env"),
26
27
  };
27
28
  export function sha1(text) {
@@ -1,6 +1,29 @@
1
1
  /** A pi text task may carry images embedded in its saved session message. */
2
2
  import { createHash } from "node:crypto";
3
+ import { existsSync, readFileSync, statSync } from "node:fs";
4
+ import { basename } from "node:path";
3
5
  import { taskKey } from "./store.js";
6
+ /** Recover files Pi expanded into a task message (for example @invoice.pdf). */
7
+ export function sessionFiles(text) {
8
+ const files = [];
9
+ const seen = new Set();
10
+ const re = /<file\s+name=["']([^"']+)["'][^>]*>/g;
11
+ for (const match of text.matchAll(re)) {
12
+ const file = match[1];
13
+ if (!file || seen.has(file) || !existsSync(file))
14
+ continue;
15
+ try {
16
+ const stat = statSync(file);
17
+ if (!stat.isFile() || stat.size > 100_000_000)
18
+ continue;
19
+ files.push({ path: file, name: basename(file), size: stat.size,
20
+ sha256: createHash("sha256").update(readFileSync(file)).digest("hex") });
21
+ seen.add(file);
22
+ }
23
+ catch { /* inaccessible attachment */ }
24
+ }
25
+ return files;
26
+ }
4
27
  const SUPPORTED = new Set(["image/jpeg", "image/png", "image/webp", "image/gif"]);
5
28
  export function sessionImages(content) {
6
29
  if (!Array.isArray(content))
@@ -20,11 +43,12 @@ export function sessionImages(content) {
20
43
  return images;
21
44
  }
22
45
  /** Text-only keys remain compatible; image bytes distinguish same-prompt documents. */
23
- export function taskInputKey(text, images) {
24
- if (images.length === 0)
46
+ export function taskInputKey(text, images, files = []) {
47
+ if (images.length === 0 && files.length === 0)
25
48
  return taskKey(text);
26
49
  const normalized = text.toLowerCase().replace(/\s+/g, " ").trim();
27
50
  const hashes = images.map((image) => `${image.mimeType}:` +
28
51
  createHash("sha256").update(Buffer.from(image.data, "base64")).digest("hex"));
29
- return taskKey(`${normalized}\nimages:${hashes.join(",")}`);
52
+ const fileHashes = files.map((file) => `${file.name}:${file.sha256}`).join(",");
53
+ return taskKey(`${normalized}\nimages:${hashes.join(",")}\nfiles:${fileHashes}`);
30
54
  }
@@ -19,7 +19,7 @@
19
19
  import type { ExtensionAPI, ExtensionContext } from "@earendil-works/pi-coding-agent";
20
20
  import { appendFileSync, existsSync, mkdirSync, readFileSync, readdirSync, writeFileSync } from "node:fs";
21
21
  import { homedir } from "node:os";
22
- import { join } from "node:path";
22
+ import { basename, join } from "node:path";
23
23
  import { createHash } from "node:crypto";
24
24
  import { spawn, type ChildProcessByStdio } from "node:child_process";
25
25
  import type { Readable } from "node:stream";
@@ -46,6 +46,9 @@ interface TrialProgress {
46
46
  latencyMs?: number;
47
47
  costUsd?: number;
48
48
  error?: string;
49
+ toolCalls?: number;
50
+ toolErrors?: number;
51
+ changedFiles?: number;
49
52
  }
50
53
 
51
54
  function parseTrialProgress(line: string): TrialProgress | null {
@@ -63,7 +66,9 @@ function parseTrialProgress(line: string): TrialProgress | null {
63
66
  function completedTrialText(trial: TrialProgress): string {
64
67
  return `${trial.model}: quality ${(trial.score ?? 0).toFixed(2)}, ` +
65
68
  `$${(trial.price ?? 0).toFixed(2)}/M, ${Math.round(trial.latencyMs ?? 0)}ms, ` +
66
- `run $${(trial.costUsd ?? 0).toFixed(4)}` + (trial.error ? ` (failed: ${trial.error})` : "");
69
+ `run $${(trial.costUsd ?? 0).toFixed(4)}, tools ${trial.toolCalls ?? 0}` +
70
+ (trial.toolErrors || trial.changedFiles ? `, tool errors ${trial.toolErrors ?? 0}, changed files ${trial.changedFiles ?? 0}` : "") +
71
+ (trial.error ? ` (failed: ${trial.error})` : "");
67
72
  }
68
73
 
69
74
  interface ExtensionPolicy {
@@ -171,10 +176,18 @@ function userTaskKey(content: unknown): string | null {
171
176
  const images = parts.filter((b) => b?.type === "image");
172
177
  if (parts.length !== parts.filter((b) => b?.type === "text" || b?.type === "image").length ||
173
178
  images.some((b) => typeof b.data !== "string" || typeof b.mimeType !== "string")) return null;
174
- if (images.length === 0) return taskKey(text);
179
+ const files: string[] = [];
180
+ const fileRe = /<file\s+name=["']([^"']+)["'][^>]*>/g;
181
+ for (const match of text.matchAll(fileRe)) {
182
+ const file = match[1];
183
+ if (!file || !existsSync(file) || files.includes(file)) continue;
184
+ try { files.push(`${basename(file)}:` + createHash("sha256").update(readFileSync(file)).digest("hex")); }
185
+ catch { /* inaccessible attachment */ }
186
+ }
187
+ if (images.length === 0 && files.length === 0) return taskKey(text);
175
188
  const hashes = images.map((b) => `${b.mimeType}:` +
176
189
  createHash("sha256").update(Buffer.from(b.data, "base64")).digest("hex"));
177
- return taskKey(`${text.toLowerCase().replace(/\s+/g, " ").trim()}\nimages:${hashes.join(",")}`);
190
+ return taskKey(`${text.toLowerCase().replace(/\s+/g, " ").trim()}\nimages:${hashes.join(",")}\nfiles:${files.join(",")}`);
178
191
  }
179
192
 
180
193
  function latestSessionTaskKey(ctx: ExtensionContext): string | null {
@@ -8,8 +8,9 @@ the best model for the task at hand, at the best price, with a vetted fallback.
8
8
 
9
9
  1. **Observes** this harness's session traces (models used, tokens, cost,
10
10
  latency, errors) without intercepting or slowing down your work.
11
- 2. **Evaluates** candidate models for each completed text or image task sequentially
12
- through pi and scores their outputs with a judge model.
11
+ 2. **Evaluates** candidate models for each completed task sequentially through
12
+ pi and scores their outputs with a judge model. Candidate tools are disabled
13
+ by default.
13
14
  3. **Maintains a pareto frontier** per task (quality vs. cost vs. latency) and
14
15
  an aggregate frontier across all of your tasks.
15
16
  4. **Watches for new model releases** (provider catalogs + public benchmarks)
@@ -43,9 +44,11 @@ the best model for the task at hand, at the best price, with a vetted fallback.
43
44
 
44
45
  ## Guarantees
45
46
 
46
- - Per-task comparisons send the task text, uploaded images (when present), and
47
- candidate answers through pi/OpenRouter for model runs and judging. Use
48
- non-sensitive examples while evaluating this alpha.
47
+ - Per-task comparisons send the task text, uploaded files or images (when
48
+ present), and candidate answers through pi/OpenRouter for model runs and
49
+ judging. Candidate Pi runs default to no tools; any tool access is an
50
+ explicit user opt-in. Use non-sensitive examples while
51
+ evaluating this alpha.
49
52
  - The shipped policy is supervised. Swaps happen automatically only after the
50
53
  user opts in and the configured quality/cost guardrails pass; otherwise they
51
54
  remain recommendations for a human to approve.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "openmerit",
3
- "version": "0.1.0",
3
+ "version": "0.1.1",
4
4
  "description": "Find better models for each pi task by comparing quality, cost, and latency.",
5
5
  "type": "module",
6
6
  "keywords": [