openmerit 0.1.0 → 0.1.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +41 -11
- package/dist/catalog.js +11 -1
- package/dist/cli.js +1 -1
- package/dist/daemon.js +8 -4
- package/dist/harness.js +1 -0
- package/dist/llm.js +57 -21
- package/dist/pi-trials.js +164 -58
- package/dist/providers.js +1 -0
- package/dist/store.js +1 -0
- package/dist/task-input.js +27 -3
- package/extension/openmerit.ts +17 -4
- package/instructions/OPENMERIT.md +8 -5
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -105,15 +105,23 @@ Keep that checkout available because pi loads a local package from its path.
|
|
|
105
105
|
A ready configuration prints `{"status":"ready","provider":"openrouter",...}`.
|
|
106
106
|
Older `model_search/.env` files are not read by the trial engine.
|
|
107
107
|
|
|
108
|
+
OpenRouter is currently required for candidate runs and catalog discovery.
|
|
109
|
+
Judge and strategist calls can use OpenAI or Gemini directly, but require
|
|
110
|
+
`OPENMERIT_DIRECT_PROVIDER=openai` or
|
|
111
|
+
`OPENMERIT_DIRECT_PROVIDER=google` plus the corresponding provider key;
|
|
112
|
+
OpenRouter model IDs remain on OpenRouter by default. The harness and
|
|
113
|
+
provider interfaces are extension seams for future integrations, not a
|
|
114
|
+
claim that other harnesses already work.
|
|
115
|
+
|
|
108
116
|
3. Verify the extension appears in `pi list`. The instruction file
|
|
109
117
|
[`instructions/OPENMERIT.md`](instructions/OPENMERIT.md) can be added to a
|
|
110
118
|
pi project's AGENTS.md for agent context, but the extension does
|
|
111
119
|
not require it.
|
|
112
120
|
|
|
113
|
-
4. Start pi in one terminal with an OpenRouter model
|
|
121
|
+
4. Start pi in one terminal with an OpenRouter model:
|
|
114
122
|
|
|
115
123
|
```bash
|
|
116
|
-
pi --provider openrouter --model openai/gpt-4o-mini
|
|
124
|
+
pi --provider openrouter --model openai/gpt-4o-mini
|
|
117
125
|
```
|
|
118
126
|
|
|
119
127
|
For a first text task, ask: “Give the shortest valid word ladder from cat
|
|
@@ -126,6 +134,10 @@ Keep that checkout available because pi loads a local package from its path.
|
|
|
126
134
|
switch. Then send a second message to see which model actually handles it.
|
|
127
135
|
You do not need a watcher terminal or the standalone `trial` command.
|
|
128
136
|
|
|
137
|
+
Candidate comparisons have no tools by default, even when the observed task
|
|
138
|
+
used tools. See **Alpha boundaries** before explicitly enabling candidate
|
|
139
|
+
tools.
|
|
140
|
+
|
|
129
141
|
To opt in to automatic swaps, edit `~/.openmerit/policy.json`, set
|
|
130
142
|
`"mode": "auto"` and `"auto_apply.enabled": true`, review the score-gain
|
|
131
143
|
and price-ratio thresholds, and run the next task.
|
|
@@ -133,7 +145,7 @@ Keep that checkout available because pi loads a local package from its path.
|
|
|
133
145
|
### Try an invoice-to-JSON task
|
|
134
146
|
|
|
135
147
|
Use the same one-terminal setup. Start pi with a vision-capable model, for
|
|
136
|
-
example `openai/gpt-4o-mini
|
|
148
|
+
example `openai/gpt-4o-mini`. In pi, type `@` to select
|
|
137
149
|
[`benchmark/invoice_ocr/data/invoice_01_row_2.jpg`](benchmark/invoice_ocr/data/invoice_01_row_2.jpg)
|
|
138
150
|
and paste the **entire, unchanged** text from
|
|
139
151
|
[`examples/invoice-prompt.txt`](examples/invoice-prompt.txt) into the same
|
|
@@ -176,8 +188,8 @@ for local experiments while keeping `budgets.max_usd_per_day` as a conservative
|
|
|
176
188
|
candidate-spend threshold.
|
|
177
189
|
|
|
178
190
|
If no comparison starts after Pi settles, confirm that pi loaded the extension
|
|
179
|
-
(`pi list`), the task finished, the pi session is saved (do not use
|
|
180
|
-
`--no-session`)
|
|
191
|
+
(`pi list`), the task finished, and the pi session is saved (do not use
|
|
192
|
+
`--no-session`). Image candidates must advertise
|
|
181
193
|
image input in OpenRouter's current model catalog. The extension queues
|
|
182
194
|
completed tasks from its current session and runs one comparison at a time.
|
|
183
195
|
Closing or switching the session cancels the active job.
|
|
@@ -191,18 +203,35 @@ remains for controlled text-task runs from a task JSON file.
|
|
|
191
203
|
|
|
192
204
|
- With the extension installed and an OpenRouter key configured, each supported
|
|
193
205
|
settled task can start comparison calls automatically. Candidate, judge, and
|
|
194
|
-
strategist requests send the task text, attached images, and
|
|
195
|
-
to OpenRouter and can incur charges. Use non-sensitive test
|
|
196
|
-
conservative account limits while evaluating this alpha.
|
|
197
|
-
- The extension queues tasks from its current pi session
|
|
198
|
-
time
|
|
199
|
-
|
|
206
|
+
strategist requests send the task text, attached files or images, and
|
|
207
|
+
candidate output to OpenRouter and can incur charges. Use non-sensitive test
|
|
208
|
+
tasks and conservative account limits while evaluating this alpha.
|
|
209
|
+
- The extension queues tasks from its current pi session and compares one at a
|
|
210
|
+
time. Candidate runs use isolated temporary copies of the original working
|
|
211
|
+
directory and have no Pi tools by default. Their temporary changes are
|
|
212
|
+
discarded and recorded as counts, and raw Pi JSON event streams are saved
|
|
213
|
+
under `~/.openmerit/traces/trials/`.
|
|
214
|
+
Candidate sessions replay the task text and images, not earlier conversation
|
|
215
|
+
context or workspace changes.
|
|
216
|
+
- Set `OPENMERIT_PI_TRIAL_TOOLS` to an explicit comma-separated allowlist such
|
|
217
|
+
as `read,grep,find,ls` to let candidate models use tools; `none` keeps them
|
|
218
|
+
disabled. Any tool access is an advanced opt-in: the copied working directory
|
|
219
|
+
prevents ordinary project writes from touching the original, but it is not an
|
|
220
|
+
OS sandbox. Tools may accept absolute paths, and shell tools may access the
|
|
221
|
+
network. Keep the no-tools default for untrusted or sensitive projects.
|
|
222
|
+
- File attachments that Pi records as `<file name="…">` are copied into every
|
|
223
|
+
candidate sandbox and passed back to Pi as `@` file inputs. This covers PDFs,
|
|
224
|
+
CSVs, spreadsheets, and other files that Pi can open; OpenMerit does not
|
|
225
|
+
implement a separate parser for them.
|
|
200
226
|
- `ledger.json` counts reported candidate-trial spend and the daily number of
|
|
201
227
|
trials. It does not yet include rubric, judge, or strategist request cost,
|
|
202
228
|
and `max_usd_per_trial` is not yet a hard provider-side cap.
|
|
203
229
|
- Each comparison has one observed baseline plus a small candidate slate and
|
|
204
230
|
one quality score per answer. Treat recommendations as experimental evidence,
|
|
205
231
|
not a universal model ranking.
|
|
232
|
+
- Candidate execution and catalog discovery currently use Pi through
|
|
233
|
+
OpenRouter. `HarnessAdapter`, `ModelProvider`, and `CatalogProvider` are
|
|
234
|
+
implementation boundaries for future harness/provider integrations.
|
|
206
235
|
|
|
207
236
|
## State layout (`~/.openmerit/`)
|
|
208
237
|
|
|
@@ -213,6 +242,7 @@ remains for controlled text-task runs from a task JSON file.
|
|
|
213
242
|
| `recommendations.jsonl` | shared append-only channel: session-bound swap recommendations + status updates |
|
|
214
243
|
| `trials.jsonl` | every model trial point (source, score, price, latency) |
|
|
215
244
|
| `observations.jsonl` | task observations extracted from session traces |
|
|
245
|
+
| `traces/trials/*.jsonl` | raw Pi JSON event streams for candidate trials |
|
|
216
246
|
| `catalog/snapshot.json` + `candidates.json` | catalog snapshot (including input modalities) + new-model queue |
|
|
217
247
|
| `benchmarks/digest.json` | public-benchmark scores per model (seed + refresh) |
|
|
218
248
|
| `ledger.json` | daily trial spend (budget enforcement) |
|
package/dist/catalog.js
CHANGED
|
@@ -1,7 +1,17 @@
|
|
|
1
1
|
/** OpenRouter catalog fetch + snapshot diffing: the "new model release" watch. */
|
|
2
2
|
import { readJson, paths, writeJson } from "./store.js";
|
|
3
3
|
const OR = "https://openrouter.ai/api/v1";
|
|
4
|
-
export
|
|
4
|
+
export class OpenRouterCatalogProvider {
|
|
5
|
+
key;
|
|
6
|
+
id = "openrouter";
|
|
7
|
+
constructor(key) {
|
|
8
|
+
this.key = key;
|
|
9
|
+
}
|
|
10
|
+
fetchCatalog() { return fetchCatalog(this.key); }
|
|
11
|
+
}
|
|
12
|
+
export async function fetchCatalog(key, provider = "openrouter") {
|
|
13
|
+
if (provider !== "openrouter")
|
|
14
|
+
return new Map();
|
|
5
15
|
const res = await fetch(`${OR}/models`, {
|
|
6
16
|
headers: { Authorization: `Bearer ${key}` },
|
|
7
17
|
});
|
package/dist/cli.js
CHANGED
|
@@ -43,7 +43,7 @@ function cmdInit() {
|
|
|
43
43
|
console.log(`ext. -> ${join(ROOT, "extension", "openmerit.ts")}`);
|
|
44
44
|
console.log(`pi npm -> pi install npm:openmerit`);
|
|
45
45
|
console.log(`pi local-> pi install "${ROOT}"`);
|
|
46
|
-
console.log("\nNext: set OPENROUTER_API_KEY (env or ~/.openmerit/.env), install one extension source into pi, then start pi; the extension launches comparisons automatically.");
|
|
46
|
+
console.log("\nNext: set OPENROUTER_API_KEY (env or ~/.openmerit/.env), install one extension source into pi, then start pi; the extension launches comparisons automatically with candidate tools disabled by default.");
|
|
47
47
|
}
|
|
48
48
|
async function cmdTrial(taskFile, rounds) {
|
|
49
49
|
const cfg = JSON.parse(readFileSync(taskFile, "utf8"));
|
package/dist/daemon.js
CHANGED
|
@@ -168,7 +168,8 @@ export async function autoTaskTick(key, policy, settledTask) {
|
|
|
168
168
|
emitTrialProgress({ phase: "complete", taskKey: tKey,
|
|
169
169
|
model: task.model, index: 1, total, score: baseline.point.score,
|
|
170
170
|
price: baseline.point.price, latencyMs: baseline.point.latencyMs,
|
|
171
|
-
costUsd: baseline.costUsd
|
|
171
|
+
costUsd: baseline.costUsd, toolCalls: baseline.point.toolCalls,
|
|
172
|
+
toolErrors: baseline.point.toolErrors, changedFiles: baseline.point.changedFiles });
|
|
172
173
|
for (let i = 1; i < total; i++) {
|
|
173
174
|
if (!budgetOk(policy).ok)
|
|
174
175
|
break;
|
|
@@ -178,18 +179,21 @@ export async function autoTaskTick(key, policy, settledTask) {
|
|
|
178
179
|
if (settledTask !== undefined)
|
|
179
180
|
emitTrialProgress({ phase: "start", taskKey: tKey,
|
|
180
181
|
model: pick.model, index: i + 1, total });
|
|
181
|
-
const { point, costUsd, error } = await runPiTrial(key, judge, task.task, rubric, pick.model, cat.get(pick.model), task.cwd, task.images);
|
|
182
|
+
const { point, costUsd, error } = await runPiTrial(key, judge, task.task, rubric, pick.model, cat.get(pick.model), task.cwd, task.images, task.files);
|
|
182
183
|
tried.add(pick.model);
|
|
183
184
|
points.push(point);
|
|
184
185
|
recordTrialSpend(costUsd);
|
|
185
186
|
appendJsonl(paths.trials(), { ...point, taskKey: tKey, sessionFile: task.sessionFile });
|
|
186
187
|
if (error)
|
|
187
188
|
failedVendors.add(pick.model.split("/")[0]);
|
|
188
|
-
console.log(`[openmerit] task ${tKey}: ${pick.model} score=${point.score.toFixed(2)} cost=$${costUsd.toFixed(4)}
|
|
189
|
+
console.log(`[openmerit] task ${tKey}: ${pick.model} score=${point.score.toFixed(2)} cost=$${costUsd.toFixed(4)} ` +
|
|
190
|
+
`tools=${point.toolCalls ?? 0} toolErrors=${point.toolErrors ?? 0} changedFiles=${point.changedFiles ?? 0}` +
|
|
191
|
+
`${error ? ` error=${error.slice(0, 80)}` : ""}`);
|
|
189
192
|
if (settledTask !== undefined)
|
|
190
193
|
emitTrialProgress({ phase: "complete", taskKey: tKey,
|
|
191
194
|
model: pick.model, index: i + 1, total, score: point.score, price: point.price,
|
|
192
|
-
latencyMs: point.latencyMs, costUsd, error: error?.slice(0, 120)
|
|
195
|
+
latencyMs: point.latencyMs, costUsd, error: error?.slice(0, 120),
|
|
196
|
+
toolCalls: point.toolCalls, toolErrors: point.toolErrors, changedFiles: point.changedFiles });
|
|
193
197
|
}
|
|
194
198
|
const rec = buildRecommendation(tKey, task.task.slice(0, 120), task.model, points, policy, task.sessionFile);
|
|
195
199
|
if (rec) {
|
package/dist/harness.js
ADDED
|
@@ -0,0 +1 @@
|
|
|
1
|
+
export {};
|
package/dist/llm.js
CHANGED
|
@@ -1,4 +1,4 @@
|
|
|
1
|
-
/** Minimal
|
|
1
|
+
/** Minimal provider clients for judge and strategist requests. */
|
|
2
2
|
import { existsSync, readFileSync } from "node:fs";
|
|
3
3
|
import { paths } from "./store.js";
|
|
4
4
|
const OR = "https://openrouter.ai/api/v1";
|
|
@@ -20,9 +20,9 @@ export function loadKey() {
|
|
|
20
20
|
throw new Error("missing OPENROUTER_API_KEY (env or ~/.openmerit/.env)");
|
|
21
21
|
return key;
|
|
22
22
|
}
|
|
23
|
-
async function request(key, path, body, tries = 3) {
|
|
23
|
+
async function request(key, path, body, baseUrl = OR, tries = 3) {
|
|
24
24
|
for (let attempt = 0; attempt < tries; attempt++) {
|
|
25
|
-
const res = await fetch(
|
|
25
|
+
const res = await fetch(baseUrl + path, {
|
|
26
26
|
method: body === undefined ? "GET" : "POST",
|
|
27
27
|
headers: {
|
|
28
28
|
Authorization: `Bearer ${key}`,
|
|
@@ -41,26 +41,62 @@ async function request(key, path, body, tries = 3) {
|
|
|
41
41
|
}
|
|
42
42
|
throw new Error("unreachable");
|
|
43
43
|
}
|
|
44
|
+
class OpenAICompatibleProvider {
|
|
45
|
+
id;
|
|
46
|
+
baseUrl;
|
|
47
|
+
key;
|
|
48
|
+
constructor(id, baseUrl, key) {
|
|
49
|
+
this.id = id;
|
|
50
|
+
this.baseUrl = baseUrl;
|
|
51
|
+
this.key = key;
|
|
52
|
+
}
|
|
53
|
+
model(model) { return this.id === "openrouter" ? model : model.split("/", 2).at(-1) ?? model; }
|
|
54
|
+
async chat(model, prompt, maxTokens, temperature) {
|
|
55
|
+
const r = await request(this.key, "/chat/completions", { model: this.model(model), messages: [{ role: "user", content: prompt }], max_tokens: maxTokens, temperature }, this.baseUrl);
|
|
56
|
+
return { content: r.choices[0]?.message?.content ?? "", usage: r.usage ?? {} };
|
|
57
|
+
}
|
|
58
|
+
async chatWithImages(model, prompt, images, maxTokens) {
|
|
59
|
+
const content = [{ type: "text", text: prompt }, ...images.map((image) => ({ type: "image_url", image_url: { url: `data:${image.mimeType};base64,${image.data}` } }))];
|
|
60
|
+
const r = await request(this.key, "/chat/completions", { model: this.model(model), messages: [{ role: "user", content }], max_tokens: maxTokens, temperature: 0 }, this.baseUrl);
|
|
61
|
+
return { content: r.choices[0]?.message?.content ?? "", usage: r.usage ?? {} };
|
|
62
|
+
}
|
|
63
|
+
}
|
|
64
|
+
class GeminiProvider {
|
|
65
|
+
key;
|
|
66
|
+
id = "google";
|
|
67
|
+
constructor(key) {
|
|
68
|
+
this.key = key;
|
|
69
|
+
}
|
|
70
|
+
async call(model, prompt, images = [], maxTokens) {
|
|
71
|
+
const name = model.split("/", 2).at(-1) ?? model;
|
|
72
|
+
const parts = [{ text: prompt }, ...images.map((image) => ({ inline_data: { mime_type: image.mimeType, data: image.data } }))];
|
|
73
|
+
const res = await fetch(`https://generativelanguage.googleapis.com/v1beta/models/${encodeURIComponent(name)}:generateContent?key=${encodeURIComponent(this.key)}`, {
|
|
74
|
+
method: "POST", headers: { "Content-Type": "application/json" },
|
|
75
|
+
body: JSON.stringify({ contents: [{ role: "user", parts }], generationConfig: { maxOutputTokens: maxTokens, temperature: images.length ? 0 : undefined } }),
|
|
76
|
+
});
|
|
77
|
+
if (!res.ok)
|
|
78
|
+
throw new Error(`Google Gemini HTTP ${res.status}: ${(await res.text()).slice(0, 300)}`);
|
|
79
|
+
const raw = await res.json();
|
|
80
|
+
return { content: raw.candidates?.[0]?.content?.parts?.map((p) => p.text ?? "").join("") ?? "", usage: { prompt_tokens: raw.usageMetadata?.promptTokenCount, completion_tokens: raw.usageMetadata?.candidatesTokenCount, total_tokens: raw.usageMetadata?.totalTokenCount } };
|
|
81
|
+
}
|
|
82
|
+
chat(model, prompt, maxTokens) { return this.call(model, prompt, [], maxTokens); }
|
|
83
|
+
chatWithImages(model, prompt, images, maxTokens) { return this.call(model, prompt, images, maxTokens); }
|
|
84
|
+
}
|
|
85
|
+
export function providerFor(model, fallbackKey) {
|
|
86
|
+
const id = model.split("/", 1)[0] || "openrouter";
|
|
87
|
+
// OpenRouter model IDs are vendor/model (for example openai/gpt-4o-mini).
|
|
88
|
+
// Direct providers require an explicit opt-in to avoid routing those IDs
|
|
89
|
+
// away from OpenRouter when users have multiple credentials configured.
|
|
90
|
+
if (id === "openai" && process.env.OPENMERIT_DIRECT_PROVIDER === "openai" && process.env.OPENAI_API_KEY)
|
|
91
|
+
return new OpenAICompatibleProvider(id, "https://api.openai.com/v1", process.env.OPENAI_API_KEY);
|
|
92
|
+
if (id === "google" && process.env.OPENMERIT_DIRECT_PROVIDER === "google" && process.env.GEMINI_API_KEY)
|
|
93
|
+
return new GeminiProvider(process.env.GEMINI_API_KEY);
|
|
94
|
+
return new OpenAICompatibleProvider("openrouter", OR, fallbackKey);
|
|
95
|
+
}
|
|
44
96
|
export async function chat(key, model, prompt, maxTokens, temperature) {
|
|
45
|
-
|
|
46
|
-
model,
|
|
47
|
-
messages: [{ role: "user", content: prompt }],
|
|
48
|
-
max_tokens: maxTokens,
|
|
49
|
-
temperature,
|
|
50
|
-
});
|
|
51
|
-
return { content: r.choices[0]?.message?.content ?? "", usage: r.usage ?? {} };
|
|
97
|
+
return providerFor(model, key).chat(model, prompt, maxTokens, temperature);
|
|
52
98
|
}
|
|
53
99
|
/** Send a visual grading request through the same OpenRouter account. */
|
|
54
100
|
export async function chatWithImages(key, model, prompt, images, maxTokens) {
|
|
55
|
-
|
|
56
|
-
model,
|
|
57
|
-
messages: [{ role: "user", content: [
|
|
58
|
-
{ type: "text", text: prompt },
|
|
59
|
-
...images.map((image) => ({ type: "image_url",
|
|
60
|
-
image_url: { url: `data:${image.mimeType};base64,${image.data}` } })),
|
|
61
|
-
] }],
|
|
62
|
-
max_tokens: maxTokens,
|
|
63
|
-
temperature: 0,
|
|
64
|
-
});
|
|
65
|
-
return { content: r.choices[0]?.message?.content ?? "", usage: r.usage ?? {} };
|
|
101
|
+
return providerFor(model, key).chatWithImages(model, prompt, images, maxTokens);
|
|
66
102
|
}
|
package/dist/pi-trials.js
CHANGED
|
@@ -1,12 +1,13 @@
|
|
|
1
1
|
/** Run a task through pi for each model, using pi's JSON event stream as the measurement source. */
|
|
2
2
|
import { spawn, spawnSync } from "node:child_process";
|
|
3
|
-
import { existsSync, mkdtempSync, readFileSync, rmSync, writeFileSync } from "node:fs";
|
|
3
|
+
import { cpSync, existsSync, mkdirSync, mkdtempSync, readdirSync, readFileSync, rmSync, statSync, writeFileSync } from "node:fs";
|
|
4
4
|
import { tmpdir } from "node:os";
|
|
5
|
-
import { join } from "node:path";
|
|
5
|
+
import { dirname, join, relative } from "node:path";
|
|
6
|
+
import { createHash } from "node:crypto";
|
|
6
7
|
import { judge, judgeWithImages } from "./judge.js";
|
|
7
8
|
import { knownInvoiceScore } from "./invoice-eval.js";
|
|
8
9
|
import { paths, readJson } from "./store.js";
|
|
9
|
-
import { sessionImages, taskInputKey } from "./task-input.js";
|
|
10
|
+
import { sessionFiles, sessionImages, taskInputKey } from "./task-input.js";
|
|
10
11
|
function answerText(content) {
|
|
11
12
|
if (typeof content === "string")
|
|
12
13
|
return content;
|
|
@@ -15,6 +16,59 @@ function answerText(content) {
|
|
|
15
16
|
return content.filter((x) => !!x && typeof x === "object" && x.type === "text")
|
|
16
17
|
.map((x) => x.text ?? "").join("\n");
|
|
17
18
|
}
|
|
19
|
+
/** Pi-backed harness adapter. Future harnesses can implement the same seam. */
|
|
20
|
+
export class PiHarnessAdapter {
|
|
21
|
+
runTask(model, input) {
|
|
22
|
+
return executePiTask(model, input.task, input.cwd, input.images, input.files);
|
|
23
|
+
}
|
|
24
|
+
}
|
|
25
|
+
const defaultHarness = new PiHarnessAdapter();
|
|
26
|
+
const KNOWN_TRIAL_TOOLS = new Set([
|
|
27
|
+
"read", "grep", "find", "ls", "bash", "powershell", "edit", "write",
|
|
28
|
+
]);
|
|
29
|
+
/**
|
|
30
|
+
* Candidate runs have no tools by default. Users may explicitly opt into a
|
|
31
|
+
* list, but a copied cwd is not an OS sandbox for absolute paths or network.
|
|
32
|
+
*/
|
|
33
|
+
export function piTrialToolArgs(setting = process.env.OPENMERIT_PI_TRIAL_TOOLS) {
|
|
34
|
+
if (!setting?.trim())
|
|
35
|
+
return ["--no-tools"];
|
|
36
|
+
const requested = [...new Set(setting.split(",").map((tool) => tool.trim()).filter(Boolean))];
|
|
37
|
+
if (requested.length === 1 && requested[0] === "none")
|
|
38
|
+
return ["--no-tools"];
|
|
39
|
+
const unknown = requested.filter((tool) => !KNOWN_TRIAL_TOOLS.has(tool));
|
|
40
|
+
if (requested.length === 0 || unknown.length > 0) {
|
|
41
|
+
throw new Error(`invalid OPENMERIT_PI_TRIAL_TOOLS${unknown.length ? `: ${unknown.join(", ")}` : ""}`);
|
|
42
|
+
}
|
|
43
|
+
return ["--tools", requested.join(",")];
|
|
44
|
+
}
|
|
45
|
+
function fileSnapshot(root) {
|
|
46
|
+
const out = new Map();
|
|
47
|
+
function walk(dir) {
|
|
48
|
+
for (const name of readdirSync(dir)) {
|
|
49
|
+
if (name === ".git" || name === "node_modules" || name === ".openmerit")
|
|
50
|
+
continue;
|
|
51
|
+
const file = join(dir, name);
|
|
52
|
+
const rel = relative(root, file);
|
|
53
|
+
const st = statSync(file);
|
|
54
|
+
if (st.isDirectory())
|
|
55
|
+
walk(file);
|
|
56
|
+
else if (st.isFile() && st.size < 20_000_000) {
|
|
57
|
+
out.set(rel, createHash("sha1").update(readFileSync(file)).digest("hex"));
|
|
58
|
+
}
|
|
59
|
+
}
|
|
60
|
+
}
|
|
61
|
+
walk(root);
|
|
62
|
+
return out;
|
|
63
|
+
}
|
|
64
|
+
function isolatedWorkspace(cwd) {
|
|
65
|
+
const root = mkdtempSync(join(tmpdir(), "openmerit-pi-workspace-"));
|
|
66
|
+
cpSync(cwd, root, { recursive: true, filter: (src) => {
|
|
67
|
+
const rel = relative(cwd, src);
|
|
68
|
+
return !rel.split("/").some((part) => part === ".git" || part === "node_modules" || part === ".openmerit");
|
|
69
|
+
} });
|
|
70
|
+
return root;
|
|
71
|
+
}
|
|
18
72
|
/** Use model A's completed result from the active pi session when it matches this task. */
|
|
19
73
|
export function recordedActiveTask(task, model, images = [], sessionFile, sessionBytes) {
|
|
20
74
|
const state = sessionFile ? null : readJson(paths.harnessState(), {});
|
|
@@ -23,6 +77,8 @@ export function recordedActiveTask(task, model, images = [], sessionFile, sessio
|
|
|
23
77
|
return null;
|
|
24
78
|
let user = null;
|
|
25
79
|
const assistants = [];
|
|
80
|
+
let toolCalls = 0;
|
|
81
|
+
let toolErrors = 0;
|
|
26
82
|
for (const line of readFileSync(file).subarray(0, sessionBytes).toString("utf8").split("\n")) {
|
|
27
83
|
if (!line.trim())
|
|
28
84
|
continue;
|
|
@@ -38,9 +94,17 @@ export function recordedActiveTask(task, model, images = [], sessionFile, sessio
|
|
|
38
94
|
if (entry.message.role === "user") {
|
|
39
95
|
user = entry.message;
|
|
40
96
|
assistants.length = 0;
|
|
97
|
+
toolCalls = 0;
|
|
98
|
+
toolErrors = 0;
|
|
41
99
|
}
|
|
42
|
-
else if (user && entry.message.role === "assistant")
|
|
100
|
+
else if (user && entry.message.role === "assistant") {
|
|
43
101
|
assistants.push(entry.message);
|
|
102
|
+
if (Array.isArray(entry.message.content))
|
|
103
|
+
toolCalls += entry.message.content.filter((part) => part?.type === "toolCall").length;
|
|
104
|
+
}
|
|
105
|
+
else if (user && entry.message.role === "toolResult" &&
|
|
106
|
+
entry.message.isError)
|
|
107
|
+
toolErrors++;
|
|
44
108
|
}
|
|
45
109
|
const userImages = user ? sessionImages(user.content) : null;
|
|
46
110
|
if (!user || !userImages || taskInputKey(answerText(user.content), userImages) !==
|
|
@@ -59,6 +123,9 @@ export function recordedActiveTask(task, model, images = [], sessionFile, sessio
|
|
|
59
123
|
costUsd: matching.reduce((n, m) => n + (m.usage?.cost?.total ?? 0), 0),
|
|
60
124
|
latencyMs: Math.max(0, last - first),
|
|
61
125
|
errors: 0,
|
|
126
|
+
toolCalls,
|
|
127
|
+
toolErrors,
|
|
128
|
+
changedFiles: 0,
|
|
62
129
|
};
|
|
63
130
|
}
|
|
64
131
|
/** Read the exact latest completed text task from the pi extension's session marker. */
|
|
@@ -91,19 +158,20 @@ export function settledActiveTask(snapshot) {
|
|
|
91
158
|
Array.isArray(entry.message.content) && entry.message.content.some((b) => b?.type === "toolCall"))
|
|
92
159
|
usedTools = true;
|
|
93
160
|
}
|
|
94
|
-
if (!userContent
|
|
161
|
+
if (!userContent)
|
|
95
162
|
return null;
|
|
96
163
|
const images = sessionImages(userContent);
|
|
97
164
|
if (!images)
|
|
98
165
|
return null;
|
|
99
166
|
const task = answerText(userContent);
|
|
100
|
-
|
|
167
|
+
const files = sessionFiles(task);
|
|
168
|
+
if (!task.trim() || taskInputKey(task, images, files) !== st.settledTaskKey)
|
|
101
169
|
return null;
|
|
102
170
|
const run = recordedActiveTask(task, st.currentModel, images, st.sessionFile, st.sessionBytes);
|
|
103
171
|
if (!run)
|
|
104
172
|
return null;
|
|
105
173
|
return { task, images, model: st.currentModel, run, sessionFile: st.sessionFile,
|
|
106
|
-
settledAt: st.settledAt, sessionBytes: st.sessionBytes, cwd: st.cwd ?? process.cwd() };
|
|
174
|
+
settledAt: st.settledAt, sessionBytes: st.sessionBytes, cwd: st.cwd ?? process.cwd(), usedTools, files };
|
|
107
175
|
}
|
|
108
176
|
/** Limit candidates to models this installed pi can actually select. */
|
|
109
177
|
export function availablePiModels() {
|
|
@@ -115,7 +183,7 @@ export function availablePiModels() {
|
|
|
115
183
|
.filter((id) => !!id && !id.startsWith("~")));
|
|
116
184
|
}
|
|
117
185
|
/** Each invocation is a fresh pi run, so model B does not inherit model A's answer. */
|
|
118
|
-
export async function executePiTask(model, task, cwd = process.cwd(), images = []) {
|
|
186
|
+
export async function executePiTask(model, task, cwd = process.cwd(), images = [], files = []) {
|
|
119
187
|
const imageDir = images.length ? mkdtempSync(join(tmpdir(), "openmerit-pi-image-")) : null;
|
|
120
188
|
const suffix = {
|
|
121
189
|
"image/jpeg": "jpg", "image/png": "png", "image/webp": "webp", "image/gif": "gif",
|
|
@@ -127,73 +195,109 @@ export async function executePiTask(model, task, cwd = process.cwd(), images = [
|
|
|
127
195
|
writeFileSync(file, Buffer.from(image.data, "base64"));
|
|
128
196
|
return `@${file}`;
|
|
129
197
|
});
|
|
130
|
-
const
|
|
131
|
-
"--provider", "openrouter", "--model", model, "--mode", "json",
|
|
132
|
-
"--offline", "--no-extensions", "--no-tools", "--print", ...imageArgs, task,
|
|
133
|
-
], { cwd, stdio: ["ignore", "pipe", "pipe"] });
|
|
134
|
-
let buffer = "";
|
|
135
|
-
let stderr = "";
|
|
136
|
-
let answer = "";
|
|
137
|
-
let costUsd = 0;
|
|
138
|
-
let errors = 0;
|
|
139
|
-
let sessionId;
|
|
140
|
-
const started = Date.now();
|
|
141
|
-
function consume(line) {
|
|
142
|
-
if (!line.trim())
|
|
143
|
-
return;
|
|
144
|
-
let event;
|
|
145
|
-
try {
|
|
146
|
-
event = JSON.parse(line);
|
|
147
|
-
}
|
|
148
|
-
catch {
|
|
149
|
-
return;
|
|
150
|
-
}
|
|
151
|
-
if (event.type === "session" && event.id)
|
|
152
|
-
sessionId = event.id;
|
|
153
|
-
if (event.type !== "message_end" || event.message?.role !== "assistant")
|
|
154
|
-
return;
|
|
155
|
-
answer = answerText(event.message.content) || answer;
|
|
156
|
-
costUsd += event.message.usage?.cost?.total ?? 0;
|
|
157
|
-
if (event.message.stopReason === "error")
|
|
158
|
-
errors++;
|
|
159
|
-
}
|
|
160
|
-
child.stdout.on("data", (chunk) => {
|
|
161
|
-
buffer += chunk.toString("utf8");
|
|
162
|
-
let n;
|
|
163
|
-
while ((n = buffer.indexOf("\n")) >= 0) {
|
|
164
|
-
consume(buffer.slice(0, n));
|
|
165
|
-
buffer = buffer.slice(n + 1);
|
|
166
|
-
}
|
|
167
|
-
});
|
|
168
|
-
child.stderr.on("data", (chunk) => { stderr += chunk.toString("utf8"); });
|
|
169
|
-
let exitCode;
|
|
198
|
+
const workspace = isolatedWorkspace(cwd);
|
|
170
199
|
try {
|
|
171
|
-
|
|
200
|
+
const attachmentDir = join(workspace, ".openmerit-attachments");
|
|
201
|
+
mkdirSync(attachmentDir, { recursive: true });
|
|
202
|
+
const attachmentPaths = new Map();
|
|
203
|
+
const fileArgs = files.map((file, i) => {
|
|
204
|
+
const safeName = `${String(i).padStart(3, "0")}-${file.name.replace(/[^A-Za-z0-9._-]/g, "_")}`;
|
|
205
|
+
const target = join(attachmentDir, safeName);
|
|
206
|
+
cpSync(file.path, target);
|
|
207
|
+
attachmentPaths.set(file.path, target);
|
|
208
|
+
return `@${target}`;
|
|
209
|
+
});
|
|
210
|
+
const candidateTask = [...attachmentPaths.entries()].reduce((text, [source, target]) => text.split(source).join(target), task);
|
|
211
|
+
const before = fileSnapshot(workspace);
|
|
212
|
+
const traceFile = paths.trialTrace(`${model}:${task}:${Date.now()}`);
|
|
213
|
+
const traceLines = [];
|
|
214
|
+
const child = spawn("pi", [
|
|
215
|
+
"--provider", "openrouter", "--model", model, "--mode", "json",
|
|
216
|
+
"--offline", "--no-extensions", "--approve", ...piTrialToolArgs(),
|
|
217
|
+
"--print", ...imageArgs, ...fileArgs, candidateTask,
|
|
218
|
+
], { cwd: workspace, stdio: ["ignore", "pipe", "pipe"] });
|
|
219
|
+
let buffer = "";
|
|
220
|
+
let stderr = "";
|
|
221
|
+
let answer = "";
|
|
222
|
+
let costUsd = 0;
|
|
223
|
+
let errors = 0;
|
|
224
|
+
let sessionId;
|
|
225
|
+
let toolCalls = 0;
|
|
226
|
+
let toolErrors = 0;
|
|
227
|
+
const started = Date.now();
|
|
228
|
+
function consume(line) {
|
|
229
|
+
if (!line.trim())
|
|
230
|
+
return;
|
|
231
|
+
traceLines.push(line);
|
|
232
|
+
let event;
|
|
233
|
+
try {
|
|
234
|
+
event = JSON.parse(line);
|
|
235
|
+
}
|
|
236
|
+
catch {
|
|
237
|
+
return;
|
|
238
|
+
}
|
|
239
|
+
if (event.type === "session" && event.id)
|
|
240
|
+
sessionId = event.id;
|
|
241
|
+
const content = event.message?.content;
|
|
242
|
+
if (event.message?.role === "assistant" && Array.isArray(content))
|
|
243
|
+
toolCalls += content.filter((part) => part?.type === "toolCall").length;
|
|
244
|
+
if (event.message?.role === "toolResult" && event.message.isError)
|
|
245
|
+
toolErrors++;
|
|
246
|
+
if (event.type !== "message_end" || event.message?.role !== "assistant")
|
|
247
|
+
return;
|
|
248
|
+
answer = answerText(event.message.content) || answer;
|
|
249
|
+
costUsd += event.message.usage?.cost?.total ?? 0;
|
|
250
|
+
if (event.message.stopReason === "error")
|
|
251
|
+
errors++;
|
|
252
|
+
}
|
|
253
|
+
child.stdout.on("data", (chunk) => {
|
|
254
|
+
buffer += chunk.toString("utf8");
|
|
255
|
+
let n;
|
|
256
|
+
while ((n = buffer.indexOf("\n")) >= 0) {
|
|
257
|
+
consume(buffer.slice(0, n));
|
|
258
|
+
buffer = buffer.slice(n + 1);
|
|
259
|
+
}
|
|
260
|
+
});
|
|
261
|
+
child.stderr.on("data", (chunk) => { stderr += chunk.toString("utf8"); });
|
|
262
|
+
const exitCode = await new Promise((resolve, reject) => {
|
|
172
263
|
child.on("error", reject);
|
|
173
264
|
child.on("close", (code) => resolve(code ?? 1));
|
|
174
265
|
});
|
|
266
|
+
consume(buffer);
|
|
267
|
+
mkdirSync(dirname(traceFile), { recursive: true });
|
|
268
|
+
writeFileSync(traceFile, traceLines.join("\n") + (traceLines.length ? "\n" : ""));
|
|
269
|
+
const after = fileSnapshot(workspace);
|
|
270
|
+
let changedFiles = 0;
|
|
271
|
+
for (const [file, hash] of after)
|
|
272
|
+
if (before.get(file) !== hash)
|
|
273
|
+
changedFiles++;
|
|
274
|
+
for (const file of before.keys())
|
|
275
|
+
if (!after.has(file))
|
|
276
|
+
changedFiles++;
|
|
277
|
+
if (exitCode !== 0 || !answer.trim()) {
|
|
278
|
+
throw new Error(`pi trial ${model} failed: ${stderr.trim().slice(0, 180) || `exit ${exitCode}, empty answer`}`);
|
|
279
|
+
}
|
|
280
|
+
return { answer, costUsd, latencyMs: Date.now() - started, errors, sessionId,
|
|
281
|
+
toolCalls, toolErrors, changedFiles, traceFile };
|
|
175
282
|
}
|
|
176
283
|
finally {
|
|
177
284
|
if (imageDir)
|
|
178
285
|
rmSync(imageDir, { recursive: true, force: true });
|
|
286
|
+
rmSync(workspace, { recursive: true, force: true });
|
|
179
287
|
}
|
|
180
|
-
consume(buffer);
|
|
181
|
-
if (exitCode !== 0 || !answer.trim()) {
|
|
182
|
-
throw new Error(`pi trial ${model} failed: ${stderr.trim().slice(0, 180) || `exit ${exitCode}, empty answer`}`);
|
|
183
|
-
}
|
|
184
|
-
return { answer, costUsd, latencyMs: Date.now() - started, errors, sessionId };
|
|
185
288
|
}
|
|
186
289
|
/** Quality uses the same OpenMerit judge; cost and latency come from pi itself. */
|
|
187
|
-
export async function runPiTrial(key, judgeModel, task, rubric, model, entry, cwd = process.cwd(), images = []) {
|
|
290
|
+
export async function runPiTrial(key, judgeModel, task, rubric, model, entry, cwd = process.cwd(), images = [], files = [], harness = defaultHarness) {
|
|
188
291
|
try {
|
|
189
|
-
const run = await
|
|
292
|
+
const run = await harness.runTask(model, { task, cwd, images, files });
|
|
190
293
|
return await scorePiRun(key, judgeModel, task, rubric, model, entry, run, "pi_trial", images);
|
|
191
294
|
}
|
|
192
295
|
catch (e) {
|
|
193
296
|
const error = String(e);
|
|
194
297
|
return {
|
|
195
298
|
point: { model, score: 0, price: entry?.price ?? 0,
|
|
196
|
-
ts: new Date().toISOString(), source: "pi_trial",
|
|
299
|
+
ts: new Date().toISOString(), source: "pi_trial", toolCalls: 0, toolErrors: 1,
|
|
300
|
+
changedFiles: 0, why: error.slice(0, 160) },
|
|
197
301
|
costUsd: 0, error,
|
|
198
302
|
};
|
|
199
303
|
}
|
|
@@ -216,7 +320,9 @@ export async function scorePiRun(key, judgeModel, task, rubric, model, entry, ru
|
|
|
216
320
|
point: {
|
|
217
321
|
model, score, price: entry?.price ?? 0,
|
|
218
322
|
latencyMs: run.latencyMs, ts: new Date().toISOString(), source,
|
|
219
|
-
why: `${why}${run.errors ? `; pi errors: ${run.errors}` : ""}
|
|
323
|
+
why: `${why}${run.errors ? `; pi errors: ${run.errors}` : ""}; ` +
|
|
324
|
+
`tools=${run.toolCalls}, toolErrors=${run.toolErrors}, changedFiles=${run.changedFiles}`,
|
|
325
|
+
toolCalls: run.toolCalls, toolErrors: run.toolErrors, changedFiles: run.changedFiles,
|
|
220
326
|
},
|
|
221
327
|
costUsd: run.costUsd,
|
|
222
328
|
sessionId: run.sessionId,
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
export {};
|
package/dist/store.js
CHANGED
|
@@ -22,6 +22,7 @@ export const paths = {
|
|
|
22
22
|
ledger: () => join(stateDir(), "ledger.json"),
|
|
23
23
|
watchProcessed: () => join(stateDir(), "watch", "processed.json"),
|
|
24
24
|
sessionJob: (marker) => join(stateDir(), "watch", "jobs", sha1(marker) + ".json"),
|
|
25
|
+
trialTrace: (id) => join(stateDir(), "traces", "trials", `${sha1(id)}.jsonl`),
|
|
25
26
|
envFile: () => join(stateDir(), ".env"),
|
|
26
27
|
};
|
|
27
28
|
export function sha1(text) {
|
package/dist/task-input.js
CHANGED
|
@@ -1,6 +1,29 @@
|
|
|
1
1
|
/** A pi text task may carry images embedded in its saved session message. */
|
|
2
2
|
import { createHash } from "node:crypto";
|
|
3
|
+
import { existsSync, readFileSync, statSync } from "node:fs";
|
|
4
|
+
import { basename } from "node:path";
|
|
3
5
|
import { taskKey } from "./store.js";
|
|
6
|
+
/** Recover files Pi expanded into a task message (for example @invoice.pdf). */
|
|
7
|
+
export function sessionFiles(text) {
|
|
8
|
+
const files = [];
|
|
9
|
+
const seen = new Set();
|
|
10
|
+
const re = /<file\s+name=["']([^"']+)["'][^>]*>/g;
|
|
11
|
+
for (const match of text.matchAll(re)) {
|
|
12
|
+
const file = match[1];
|
|
13
|
+
if (!file || seen.has(file) || !existsSync(file))
|
|
14
|
+
continue;
|
|
15
|
+
try {
|
|
16
|
+
const stat = statSync(file);
|
|
17
|
+
if (!stat.isFile() || stat.size > 100_000_000)
|
|
18
|
+
continue;
|
|
19
|
+
files.push({ path: file, name: basename(file), size: stat.size,
|
|
20
|
+
sha256: createHash("sha256").update(readFileSync(file)).digest("hex") });
|
|
21
|
+
seen.add(file);
|
|
22
|
+
}
|
|
23
|
+
catch { /* inaccessible attachment */ }
|
|
24
|
+
}
|
|
25
|
+
return files;
|
|
26
|
+
}
|
|
4
27
|
const SUPPORTED = new Set(["image/jpeg", "image/png", "image/webp", "image/gif"]);
|
|
5
28
|
export function sessionImages(content) {
|
|
6
29
|
if (!Array.isArray(content))
|
|
@@ -20,11 +43,12 @@ export function sessionImages(content) {
|
|
|
20
43
|
return images;
|
|
21
44
|
}
|
|
22
45
|
/** Text-only keys remain compatible; image bytes distinguish same-prompt documents. */
|
|
23
|
-
export function taskInputKey(text, images) {
|
|
24
|
-
if (images.length === 0)
|
|
46
|
+
export function taskInputKey(text, images, files = []) {
|
|
47
|
+
if (images.length === 0 && files.length === 0)
|
|
25
48
|
return taskKey(text);
|
|
26
49
|
const normalized = text.toLowerCase().replace(/\s+/g, " ").trim();
|
|
27
50
|
const hashes = images.map((image) => `${image.mimeType}:` +
|
|
28
51
|
createHash("sha256").update(Buffer.from(image.data, "base64")).digest("hex"));
|
|
29
|
-
|
|
52
|
+
const fileHashes = files.map((file) => `${file.name}:${file.sha256}`).join(",");
|
|
53
|
+
return taskKey(`${normalized}\nimages:${hashes.join(",")}\nfiles:${fileHashes}`);
|
|
30
54
|
}
|
package/extension/openmerit.ts
CHANGED
|
@@ -19,7 +19,7 @@
|
|
|
19
19
|
import type { ExtensionAPI, ExtensionContext } from "@earendil-works/pi-coding-agent";
|
|
20
20
|
import { appendFileSync, existsSync, mkdirSync, readFileSync, readdirSync, writeFileSync } from "node:fs";
|
|
21
21
|
import { homedir } from "node:os";
|
|
22
|
-
import { join } from "node:path";
|
|
22
|
+
import { basename, join } from "node:path";
|
|
23
23
|
import { createHash } from "node:crypto";
|
|
24
24
|
import { spawn, type ChildProcessByStdio } from "node:child_process";
|
|
25
25
|
import type { Readable } from "node:stream";
|
|
@@ -46,6 +46,9 @@ interface TrialProgress {
|
|
|
46
46
|
latencyMs?: number;
|
|
47
47
|
costUsd?: number;
|
|
48
48
|
error?: string;
|
|
49
|
+
toolCalls?: number;
|
|
50
|
+
toolErrors?: number;
|
|
51
|
+
changedFiles?: number;
|
|
49
52
|
}
|
|
50
53
|
|
|
51
54
|
function parseTrialProgress(line: string): TrialProgress | null {
|
|
@@ -63,7 +66,9 @@ function parseTrialProgress(line: string): TrialProgress | null {
|
|
|
63
66
|
function completedTrialText(trial: TrialProgress): string {
|
|
64
67
|
return `${trial.model}: quality ${(trial.score ?? 0).toFixed(2)}, ` +
|
|
65
68
|
`$${(trial.price ?? 0).toFixed(2)}/M, ${Math.round(trial.latencyMs ?? 0)}ms, ` +
|
|
66
|
-
`run $${(trial.costUsd ?? 0).toFixed(4)}
|
|
69
|
+
`run $${(trial.costUsd ?? 0).toFixed(4)}, tools ${trial.toolCalls ?? 0}` +
|
|
70
|
+
(trial.toolErrors || trial.changedFiles ? `, tool errors ${trial.toolErrors ?? 0}, changed files ${trial.changedFiles ?? 0}` : "") +
|
|
71
|
+
(trial.error ? ` (failed: ${trial.error})` : "");
|
|
67
72
|
}
|
|
68
73
|
|
|
69
74
|
interface ExtensionPolicy {
|
|
@@ -171,10 +176,18 @@ function userTaskKey(content: unknown): string | null {
|
|
|
171
176
|
const images = parts.filter((b) => b?.type === "image");
|
|
172
177
|
if (parts.length !== parts.filter((b) => b?.type === "text" || b?.type === "image").length ||
|
|
173
178
|
images.some((b) => typeof b.data !== "string" || typeof b.mimeType !== "string")) return null;
|
|
174
|
-
|
|
179
|
+
const files: string[] = [];
|
|
180
|
+
const fileRe = /<file\s+name=["']([^"']+)["'][^>]*>/g;
|
|
181
|
+
for (const match of text.matchAll(fileRe)) {
|
|
182
|
+
const file = match[1];
|
|
183
|
+
if (!file || !existsSync(file) || files.includes(file)) continue;
|
|
184
|
+
try { files.push(`${basename(file)}:` + createHash("sha256").update(readFileSync(file)).digest("hex")); }
|
|
185
|
+
catch { /* inaccessible attachment */ }
|
|
186
|
+
}
|
|
187
|
+
if (images.length === 0 && files.length === 0) return taskKey(text);
|
|
175
188
|
const hashes = images.map((b) => `${b.mimeType}:` +
|
|
176
189
|
createHash("sha256").update(Buffer.from(b.data, "base64")).digest("hex"));
|
|
177
|
-
return taskKey(`${text.toLowerCase().replace(/\s+/g, " ").trim()}\nimages:${hashes.join(",")}`);
|
|
190
|
+
return taskKey(`${text.toLowerCase().replace(/\s+/g, " ").trim()}\nimages:${hashes.join(",")}\nfiles:${files.join(",")}`);
|
|
178
191
|
}
|
|
179
192
|
|
|
180
193
|
function latestSessionTaskKey(ctx: ExtensionContext): string | null {
|
|
@@ -8,8 +8,9 @@ the best model for the task at hand, at the best price, with a vetted fallback.
|
|
|
8
8
|
|
|
9
9
|
1. **Observes** this harness's session traces (models used, tokens, cost,
|
|
10
10
|
latency, errors) without intercepting or slowing down your work.
|
|
11
|
-
2. **Evaluates** candidate models for each completed
|
|
12
|
-
|
|
11
|
+
2. **Evaluates** candidate models for each completed task sequentially through
|
|
12
|
+
pi and scores their outputs with a judge model. Candidate tools are disabled
|
|
13
|
+
by default.
|
|
13
14
|
3. **Maintains a pareto frontier** per task (quality vs. cost vs. latency) and
|
|
14
15
|
an aggregate frontier across all of your tasks.
|
|
15
16
|
4. **Watches for new model releases** (provider catalogs + public benchmarks)
|
|
@@ -43,9 +44,11 @@ the best model for the task at hand, at the best price, with a vetted fallback.
|
|
|
43
44
|
|
|
44
45
|
## Guarantees
|
|
45
46
|
|
|
46
|
-
- Per-task comparisons send the task text, uploaded images (when
|
|
47
|
-
candidate answers through pi/OpenRouter for model runs and
|
|
48
|
-
|
|
47
|
+
- Per-task comparisons send the task text, uploaded files or images (when
|
|
48
|
+
present), and candidate answers through pi/OpenRouter for model runs and
|
|
49
|
+
judging. Candidate Pi runs default to no tools; any tool access is an
|
|
50
|
+
explicit user opt-in. Use non-sensitive examples while
|
|
51
|
+
evaluating this alpha.
|
|
49
52
|
- The shipped policy is supervised. Swaps happen automatically only after the
|
|
50
53
|
user opts in and the configured quality/cost guardrails pass; otherwise they
|
|
51
54
|
remain recommendations for a human to approve.
|