champollion-mcp-server 0.1.0 → 0.1.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +34 -0
- package/README.md +2 -2
- package/instructions.md +5 -5
- package/package.json +2 -1
- package/src/index.js +50 -17
- package/src/tools/forge.js +1 -1
- package/src/tools/harness.js +67 -49
- package/src/tools/queue.js +392 -82
package/CHANGELOG.md
ADDED
|
@@ -0,0 +1,34 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
## 0.1.1 (2026-08-27)
|
|
4
|
+
|
|
5
|
+
Fixes the queue-tool timeouts found by live testing of the 0.1.0 npm release:
|
|
6
|
+
with the queue at 211k+ open items, `list_queue`, `estimate_cost`,
|
|
7
|
+
`get_project_info`, and `run_benchmark` all exceeded MCP clients' default 60s
|
|
8
|
+
request timeout, because the DB fetch path drained the entire `queue_top`
|
|
9
|
+
ranking (~423 sequential pages ≈ 3 minutes) before answering anything.
|
|
10
|
+
|
|
11
|
+
- **Bounded, purpose-fit queue fetching.** Metadata comes from
|
|
12
|
+
queue-preview.json plus a live open-item count from the unpaged
|
|
13
|
+
`queue_pairs` RPC; ranked items are paged from `queue_top` only as deep as
|
|
14
|
+
the caller's selection needs (fetch-until-satisfied, bounded by
|
|
15
|
+
`CHAMPOLLION_QUEUE_MAX_PAGES`, default 20 pages / 10,000 rows); single
|
|
16
|
+
items are read by primary key over PostgREST with a verified-coverage
|
|
17
|
+
probe. When a bound truncates a search, the tool says how deep it looked —
|
|
18
|
+
no silent caps.
|
|
19
|
+
- **`get_queue_item` / `run_benchmark(item_id)`** now do a direct by-id (or
|
|
20
|
+
mode+priority) lookup instead of scanning a full drain, and refuse items
|
|
21
|
+
already covered by a VERIFIED run instead of re-spending on them.
|
|
22
|
+
- **Queue-mode `dry_run` is backgrounded** like a real run (the installed
|
|
23
|
+
harness loads the full queue before printing its plan): it returns a job id
|
|
24
|
+
immediately; the plan arrives via `get_run_status`.
|
|
25
|
+
- **Failure ladder hardened.** A slow-but-alive DB degrades to a
|
|
26
|
+
truncated-but-honest prefix; a dead DB falls back to the static queue.json
|
|
27
|
+
blob; when both are down the error names both causes.
|
|
28
|
+
- Harness (mt-eval, monorepo): `--top N` runs now page the DB queue only as
|
|
29
|
+
deep as selection needs, and `CHAMPOLLION_QUEUE_SOURCE=blob` is accepted as
|
|
30
|
+
a sentinel (previously read as a literal file path).
|
|
31
|
+
|
|
32
|
+
## 0.1.0 (2026-08-27)
|
|
33
|
+
|
|
34
|
+
Initial npm release.
|
package/README.md
CHANGED
|
@@ -27,7 +27,7 @@ When connected to an agent (Claude Code, Antigravity, Cursor, etc.), the server
|
|
|
27
27
|
|
|
28
28
|
#### Training tools (nmt-forge)
|
|
29
29
|
|
|
30
|
-
These wrap the [nmt-forge](https://github.com/gamedaysuits/Champollion) training suite. forge is part of the Champollion monorepo (not on PyPI) — clone the repo and set `CHAMPOLLION_FORGE_DIR` to its `forge/` directory; scoring additionally needs the eval harness (`pip install mt-eval`). Without forge present these tools return an actionable error rather than crashing.
|
|
30
|
+
These wrap the [nmt-forge](https://github.com/gamedaysuits/Champollion) training suite. forge is part of the Champollion monorepo (not on PyPI) — clone the repo and set `CHAMPOLLION_FORGE_DIR` to its `forge/` directory; scoring additionally needs the eval harness (`pip install mt-eval-harness`). Without forge present these tools return an actionable error rather than crashing.
|
|
31
31
|
|
|
32
32
|
| Tool | Type | Description |
|
|
33
33
|
|---|---|---|
|
|
@@ -132,7 +132,7 @@ Once connected, you can talk to your agent naturally:
|
|
|
132
132
|
>
|
|
133
133
|
> **Agent** uses `list_queue` with `budget: 10` → sees what's available
|
|
134
134
|
>
|
|
135
|
-
> **Agent:** "
|
|
135
|
+
> **Agent:** "The queue has over 200,000 open benchmark items. Your $10 could fund dozens of runs. Any preference on languages?"
|
|
136
136
|
>
|
|
137
137
|
> **You:** "West African languages"
|
|
138
138
|
>
|
package/instructions.md
CHANGED
|
@@ -19,13 +19,13 @@ Start with `get_project_info` to understand what Champollion is and how contribu
|
|
|
19
19
|
|
|
20
20
|
#### Running is asynchronous (important)
|
|
21
21
|
|
|
22
|
-
A real benchmark runs a corpus through a live model and takes **minutes**, which is longer than the default 60-second request timeout most MCP clients (Claude Code, Cursor) enforce. So `run_benchmark` does **not** wait for the run to finish — it launches the run in the background and returns a `job id` right away. Treat that `STARTED` response as success, **not** completion.
|
|
22
|
+
A real benchmark runs a corpus through a live model and takes **minutes**, which is longer than the default 60-second request timeout most MCP clients (Claude Code, Cursor) enforce. So `run_benchmark` does **not** wait for the run to finish — it launches the run in the background and returns a `job id` right away. Treat that `STARTED` response as success, **not** completion. (A queue-mode `dry_run` is backgrounded the same way — the harness loads the full ranked queue before printing its plan — so poll `get_run_status` for the plan; an `item_id` dry run answers inline.)
|
|
23
23
|
|
|
24
24
|
- After `run_benchmark` returns, call `get_run_status { "job_id": "run-N" }`. Each poll returns instantly: `RUNNING` (keep polling), `COMPLETED` (output is in the response), `FAILED`, or `ERROR`.
|
|
25
25
|
- Do **not** re-call `run_benchmark` because nothing "came back" — that would start a **second** run and spend tokens twice. The first call already started it; poll `get_run_status` instead.
|
|
26
26
|
- Jobs live in the server process's memory, so a job id is only pollable from the same session. Call `get_run_status` with no `job_id` to list every job started this session.
|
|
27
27
|
|
|
28
|
-
The estimate you show in step 4 is what executes: `run_benchmark` runs budget/top items in **deterministic top-of-queue order** (it passes `--no-spread`), so the selection matches the `estimate_cost` / `list_queue` preview item-for-item. (One caveat for honesty: `estimate_cost` samples up to ~500 matching items
|
|
28
|
+
The estimate you show in step 4 is what executes: `run_benchmark` runs budget/top items in **deterministic top-of-queue order** (it passes `--no-spread`), so the selection matches the `estimate_cost` / `list_queue` preview item-for-item. (One caveat for honesty: `estimate_cost` samples up to ~500 matching items from a bounded scan of the top of the ranking — its reply says how deep it looked when the bound bites — so for a very large budget or a very narrow filter, treat its count/total as a lower bound.) A live run additionally skips any (corpus, model, condition) combo already on the leaderboard, so the executed set can be a subset of the preview — never a different, unseen set.
|
|
29
29
|
|
|
30
30
|
To spend tokens for **scoring/validation without writing to the leaderboard**, pass `publish: false` to `run_benchmark` (budget/top mode). A single `item_id` run is always scored locally and is never auto-published — publish it afterward with `mt-eval publish`, or use budget/top mode to auto-publish.
|
|
31
31
|
|
|
@@ -111,7 +111,7 @@ per-call report of what was cached, what was validated, and what it cost.
|
|
|
111
111
|
|
|
112
112
|
### "What languages need help?"
|
|
113
113
|
|
|
114
|
-
1. Call `list_queue`
|
|
114
|
+
1. Call `list_queue` to see what's available (it fetches only as deep as your `limit` needs, so a generous limit is fine)
|
|
115
115
|
2. Look for languages with the highest ECV (Expected Chain Value) — these have the most impact per dollar
|
|
116
116
|
3. Use `search_languages` to find context: family, speakers, region, endonym
|
|
117
117
|
|
|
@@ -214,7 +214,7 @@ https://champollion.dev/docs/network/getting-started/contributing-compute
|
|
|
214
214
|
- **Trust the queue ranking.** Items are ordered by ECV — the expected improvement in translation quality per dollar. Don't re-sort or second-guess the ranking.
|
|
215
215
|
- **Budget mode skips, it doesn't stop.** If an item exceeds the remaining budget, the system skips it and continues to cheaper items further down the queue. This is by design — it maximizes what gets done within a budget.
|
|
216
216
|
- **Items without cost estimates are skipped in budget mode.** Unknown cost ≠ free.
|
|
217
|
-
- **The preview is what runs.** `run_benchmark` executes in deterministic top-of-queue order (`--no-spread`), so what `estimate_cost`/`list_queue` showed is what spends. Don't assume a different set ran.
|
|
217
|
+
- **The preview is what runs.** `run_benchmark` executes in deterministic top-of-queue order (`--no-spread`), so what `estimate_cost`/`list_queue` showed is what spends. Don't assume a different set ran. (Both previews scan a bounded top slice of the ranking and say so when the bound bites; a very large budget can execute deeper than the preview sampled — the executed order is still the same ranking, top first.)
|
|
218
218
|
- **Publishing writes to a public, production leaderboard.** Budget/top runs auto-publish each result by default. For a scoring/validation run with no leaderboard write, pass `publish: false`.
|
|
219
219
|
- **Use `translate` for translation; don't improvise.** The tool's Translation Memory makes repeats free and its quality gate rejects garbage deterministically — a hand-rolled prompt has neither. Never present a gate-FAILED text as a translation.
|
|
220
220
|
- **Translation ≠ evidence.** `translate` output is production translation; quality claims about methods and models come only from benchmark runs and the leaderboard.
|
|
@@ -225,7 +225,7 @@ The champollion.dev homepage map is an idealization — read the data, not
|
|
|
225
225
|
the picture. The full endpoint table lives in the
|
|
226
226
|
`champollion://network-data` resource.
|
|
227
227
|
|
|
228
|
-
- **Queue**: served LIVE from the public database
|
|
228
|
+
- **Queue**: served LIVE from the public database by default — stats from queue-preview.json + the unpaged `queue_pairs` RPC, ranked items paged from the `queue_top` RPC only as deep as each question needs, single items by primary key — with https://champollion.dev/queue.json as the fallback when the DB is unreachable (cached 5 minutes); small preview at https://champollion.dev/queue-preview.json
|
|
229
229
|
- **Mesh**: https://champollion.dev/mesh.json — the measured/registered pair network behind the map
|
|
230
230
|
- **Corpus registry**: https://champollion.dev/registry.json — every registered eval corpus with license lane + attribution
|
|
231
231
|
- **Provider coverage**: `shared/catalogue/method-coverage.json` (repo) — each provider's published language list, cited + as-of + `tier`. The map's green is two tiers by exact ISO-639-3 code: bright = a deployed service lists it (Google/Microsoft/DeepL/LibreTranslate); dim = only an open research model lists it (NLLB/OPUS/M2M-100/MADLAD-400 — a model-card code, not a usable service). "Covered" is a published-list claim, never a quality claim.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "champollion-mcp-server",
|
|
3
|
-
"version": "0.1.
|
|
3
|
+
"version": "0.1.1",
|
|
4
4
|
"description": "MCP server for Champollion \u2014 exposes the public benchmark queue, language metadata, and mt-eval harness controls to AI agents.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"bin": {
|
|
@@ -12,6 +12,7 @@
|
|
|
12
12
|
"src",
|
|
13
13
|
"instructions.md",
|
|
14
14
|
"README.md",
|
|
15
|
+
"CHANGELOG.md",
|
|
15
16
|
"LICENSE"
|
|
16
17
|
],
|
|
17
18
|
"scripts": {
|
package/src/index.js
CHANGED
|
@@ -42,7 +42,9 @@ import { readFile } from 'node:fs/promises';
|
|
|
42
42
|
import { resolve, dirname } from 'node:path';
|
|
43
43
|
import { fileURLToPath } from 'node:url';
|
|
44
44
|
|
|
45
|
-
import {
|
|
45
|
+
import {
|
|
46
|
+
fetchQueueMeta, selectFromQueue, lookupQueueItem, estimateCost,
|
|
47
|
+
} from './tools/queue.js';
|
|
46
48
|
import { searchLanguages, loadLanguageIndex } from './tools/languages.js';
|
|
47
49
|
import { runBenchmark, getRunStatus } from './tools/harness.js';
|
|
48
50
|
import { fetchResults, formatResults, fetchRunCard } from './tools/results.js';
|
|
@@ -84,8 +86,14 @@ export async function createServer() {
|
|
|
84
86
|
// no instructions if the file is missing rather than failing startup.
|
|
85
87
|
const instructions = await loadInstructions(resolve(__dirname, '..', 'instructions.md'));
|
|
86
88
|
|
|
89
|
+
// Version from package.json — the SSOT — so the handshake can never claim
|
|
90
|
+
// a stale release (0.1.1 shipped while the handshake still said 0.1.0).
|
|
91
|
+
const { version } = JSON.parse(
|
|
92
|
+
await readFile(resolve(__dirname, '..', 'package.json'), 'utf-8'),
|
|
93
|
+
);
|
|
94
|
+
|
|
87
95
|
const server = new McpServer(
|
|
88
|
-
{ name: 'champollion', version
|
|
96
|
+
{ name: 'champollion', version },
|
|
89
97
|
instructions ? { instructions } : undefined,
|
|
90
98
|
);
|
|
91
99
|
|
|
@@ -307,10 +315,10 @@ export async function createServer() {
|
|
|
307
315
|
},
|
|
308
316
|
async ({ budget, language, source_language, model, condition, limit }) => {
|
|
309
317
|
try {
|
|
310
|
-
const
|
|
311
|
-
const items = filterQueue(queue.items, {
|
|
318
|
+
const sel = await selectFromQueue({
|
|
312
319
|
budget, language, source_language, model, condition, limit,
|
|
313
320
|
});
|
|
321
|
+
const items = sel.selected;
|
|
314
322
|
|
|
315
323
|
// Compute summary statistics for the filtered set
|
|
316
324
|
const totalCost = items.reduce((s, it) => s + (it.est_cost_usd || 0), 0);
|
|
@@ -324,11 +332,19 @@ export async function createServer() {
|
|
|
324
332
|
+ `[${it.condition}]`
|
|
325
333
|
);
|
|
326
334
|
|
|
335
|
+
// Honest truncation: the DB path only pages as deep as the selection
|
|
336
|
+
// needs (bounded); if the bound hit before `limit` matches were found,
|
|
337
|
+
// say how far the search went instead of implying the queue ran dry.
|
|
338
|
+
const depthNote = (!sel.complete && items.length < limit)
|
|
339
|
+
? `Searched the top ${sel.scannedRows.toLocaleString()} ranked items — deeper matches may exist. Narrow the filters, or raise CHAMPOLLION_QUEUE_MAX_PAGES to scan deeper.`
|
|
340
|
+
: '';
|
|
341
|
+
|
|
327
342
|
const summary = [
|
|
328
|
-
`Found ${items.length} items (of ${
|
|
343
|
+
`Found ${items.length} items (of ${sel.metadata.open_items} total open).`,
|
|
329
344
|
`Estimated cost: $${totalCost.toFixed(2)}`,
|
|
330
345
|
`Languages: ${languages.join(', ')}`,
|
|
331
|
-
`Models in queue: ${
|
|
346
|
+
`Models in queue: ${sel.metadata.models.map(m => m.split('/').pop()).join(', ')}`,
|
|
347
|
+
...(depthNote ? [depthNote] : []),
|
|
332
348
|
'',
|
|
333
349
|
...lines,
|
|
334
350
|
].join('\n');
|
|
@@ -363,16 +379,21 @@ export async function createServer() {
|
|
|
363
379
|
};
|
|
364
380
|
}
|
|
365
381
|
try {
|
|
366
|
-
|
|
367
|
-
|
|
382
|
+
// Single-item lookups go straight to the queue_items primary key (or
|
|
383
|
+
// the mode+priority index) — never the ranked paging path.
|
|
384
|
+
const { item, covered, truncatedNote } = await lookupQueueItem({ id, priority });
|
|
368
385
|
if (!item) {
|
|
386
|
+
const note = truncatedNote ? ` ${truncatedNote}` : '';
|
|
369
387
|
return {
|
|
370
|
-
content: [{ type: 'text', text: `No queue item found for ${id ? `id="${id}"` : `priority=${priority}`}
|
|
388
|
+
content: [{ type: 'text', text: `No queue item found for ${id ? `id="${id}"` : `priority=${priority}`}.${note}` }],
|
|
371
389
|
isError: true,
|
|
372
390
|
};
|
|
373
391
|
}
|
|
392
|
+
const coveredNote = covered === true
|
|
393
|
+
? '\n\nNote: this item is already covered by a VERIFIED run — it is no longer an open work item and will not appear in list_queue.'
|
|
394
|
+
: '';
|
|
374
395
|
return {
|
|
375
|
-
content: [{ type: 'text', text: JSON.stringify(item, null, 2) }],
|
|
396
|
+
content: [{ type: 'text', text: JSON.stringify(item, null, 2) + coveredNote }],
|
|
376
397
|
};
|
|
377
398
|
} catch (err) {
|
|
378
399
|
return {
|
|
@@ -405,10 +426,17 @@ export async function createServer() {
|
|
|
405
426
|
},
|
|
406
427
|
async ({ budget, language, source_language, model, condition }) => {
|
|
407
428
|
try {
|
|
408
|
-
|
|
409
|
-
|
|
429
|
+
// Deepen the ranked prefix toward the estimate cap (500), then run
|
|
430
|
+
// the same pure aggregation as before over what was scanned.
|
|
431
|
+
const sel = await selectFromQueue({
|
|
432
|
+
budget, language, source_language, model, condition, limit: 500,
|
|
433
|
+
});
|
|
434
|
+
const result = estimateCost(sel.scanned, {
|
|
410
435
|
budget, language, source_language, model, condition,
|
|
411
436
|
});
|
|
437
|
+
const depthNote = (!sel.complete && !result.capped)
|
|
438
|
+
? `Scanned the top ${sel.scannedRows.toLocaleString()} ranked items — treat count/total as a lower bound; deeper matches may exist.`
|
|
439
|
+
: '';
|
|
412
440
|
return {
|
|
413
441
|
content: [{
|
|
414
442
|
type: 'text',
|
|
@@ -419,6 +447,7 @@ export async function createServer() {
|
|
|
419
447
|
`Most expensive item: $${result.mostExpensive.toFixed(4)}`,
|
|
420
448
|
`Languages covered: ${result.languages.join(', ')}`,
|
|
421
449
|
budget ? `Budget remaining: $${(budget - result.totalCost).toFixed(2)}` : '',
|
|
450
|
+
depthNote,
|
|
422
451
|
].filter(Boolean).join('\n'),
|
|
423
452
|
}],
|
|
424
453
|
};
|
|
@@ -498,8 +527,9 @@ export async function createServer() {
|
|
|
498
527
|
{},
|
|
499
528
|
async () => {
|
|
500
529
|
try {
|
|
501
|
-
|
|
502
|
-
|
|
530
|
+
// Stats only — metadata comes from the preview + one aggregate RPC;
|
|
531
|
+
// this tool never needs a single ranked item.
|
|
532
|
+
const { metadata: meta } = await fetchQueueMeta();
|
|
503
533
|
return {
|
|
504
534
|
content: [{
|
|
505
535
|
type: 'text',
|
|
@@ -520,7 +550,7 @@ export async function createServer() {
|
|
|
520
550
|
'The easiest way to help: donate some API tokens to run benchmarks',
|
|
521
551
|
'from the public queue. Anyone with an API key can contribute:',
|
|
522
552
|
'',
|
|
523
|
-
'1. Install the harness: `pipx install mt-eval`',
|
|
553
|
+
'1. Install the harness: `pipx install mt-eval-harness`',
|
|
524
554
|
'2. Set your API key: `export OPENROUTER_API_KEY=sk-or-...`',
|
|
525
555
|
'3. Run from the queue: `mt-eval queue --budget 5`',
|
|
526
556
|
' (runs top items up to $5 estimated cost)',
|
|
@@ -537,7 +567,7 @@ export async function createServer() {
|
|
|
537
567
|
'',
|
|
538
568
|
'## Current Queue Stats',
|
|
539
569
|
'',
|
|
540
|
-
`- Open items: ${meta.open_items.toLocaleString()}`,
|
|
570
|
+
`- Open items: ${meta.open_items.toLocaleString()}${meta.open_items_basis === 'generation' ? ' (as of last queue generation)' : ''}`,
|
|
541
571
|
`- Corpora: ${meta.corpora}`,
|
|
542
572
|
`- Models: ${meta.models.map(m => m.split('/').pop()).join(', ')}`,
|
|
543
573
|
`- Conditions: ${meta.conditions.join(', ')}`,
|
|
@@ -1033,7 +1063,10 @@ export async function createServer() {
|
|
|
1033
1063
|
.describe('Run a specific queue item by ID. One of '
|
|
1034
1064
|
+ 'budget/top/item_id is REQUIRED for a real run.'),
|
|
1035
1065
|
dry_run: z.boolean().default(false)
|
|
1036
|
-
.describe('If true, show what would be run without executing anything.'
|
|
1066
|
+
.describe('If true, show what would be run without executing anything. '
|
|
1067
|
+
+ 'A queue-mode (budget/top) dry run is backgrounded like a real run '
|
|
1068
|
+
+ '— it returns a job id; poll get_run_status for the plan. An '
|
|
1069
|
+
+ 'item_id dry run returns its plan inline.'),
|
|
1037
1070
|
provider: z.enum(['openrouter', 'openai', 'anthropic', 'gemini']).optional()
|
|
1038
1071
|
.describe('API provider. Auto-detected from environment if not specified.'),
|
|
1039
1072
|
confirm: z.boolean().default(false)
|
package/src/tools/forge.js
CHANGED
|
@@ -91,7 +91,7 @@ export function runForge(subArgs, { workspace, timeout = 120_000 } = {}) {
|
|
|
91
91
|
+ 'Champollion monorepo (not on PyPI): clone '
|
|
92
92
|
+ 'https://github.com/gamedaysuits/Champollion and point '
|
|
93
93
|
+ 'CHAMPOLLION_FORGE_DIR at its forge/ directory. Scoring also needs '
|
|
94
|
-
+ 'the eval harness: pip install mt-eval';
|
|
94
|
+
+ 'the eval harness: pip install mt-eval-harness';
|
|
95
95
|
}
|
|
96
96
|
resolvePromise({ code: code ?? 1, stdout, stderr });
|
|
97
97
|
});
|
package/src/tools/harness.js
CHANGED
|
@@ -146,22 +146,6 @@ export function stripAbsolutePaths(text) {
|
|
|
146
146
|
.replace(/(?<![\w:/])\/(?:[^\s/'")\],]+\/)+([^\s/'")\],]+)/g, '$1');
|
|
147
147
|
}
|
|
148
148
|
|
|
149
|
-
/**
|
|
150
|
-
* Reduce raw subprocess output to a single, path-stripped final error line.
|
|
151
|
-
* A Python traceback ends with the `ExceptionType: message` line — the only
|
|
152
|
-
* line worth relaying; the frames above it are local paths and stack noise.
|
|
153
|
-
*
|
|
154
|
-
* @param {string} stderr
|
|
155
|
-
* @param {string} stdout
|
|
156
|
-
* @returns {string}
|
|
157
|
-
*/
|
|
158
|
-
function finalErrorLine(stderr, stdout) {
|
|
159
|
-
const source = (stderr || '').trim() || (stdout || '').trim();
|
|
160
|
-
if (!source) return '(no error output captured)';
|
|
161
|
-
const lines = source.split('\n').map((l) => l.trim()).filter(Boolean);
|
|
162
|
-
return stripAbsolutePaths(lines[lines.length - 1]);
|
|
163
|
-
}
|
|
164
|
-
|
|
165
149
|
// ---------------------------------------------------------------------------
|
|
166
150
|
// Command construction — reconstruct argv locally, NEVER shell the queue.
|
|
167
151
|
// ---------------------------------------------------------------------------
|
|
@@ -295,6 +279,7 @@ function tail(text) {
|
|
|
295
279
|
* @param {object} spec
|
|
296
280
|
* @param {string[]} spec.argv argv AFTER the `mt-eval` program name
|
|
297
281
|
* @param {'item'|'queue'} spec.mode
|
|
282
|
+
* @param {boolean} [spec.dryRun] true = plan-only run, spends nothing
|
|
298
283
|
* @param {string} spec.label human description of what's running
|
|
299
284
|
* @param {string} [spec.estLabel] estimated-cost label for the start message
|
|
300
285
|
* @param {boolean} [spec.publish] queue mode: will results publish?
|
|
@@ -302,11 +287,12 @@ function tail(text) {
|
|
|
302
287
|
* @param {function} spec.exec execCapture (injectable for tests)
|
|
303
288
|
* @returns {object} the job record
|
|
304
289
|
*/
|
|
305
|
-
function launchJob({ argv, mode, label, estLabel, publish, timeout, exec }) {
|
|
290
|
+
function launchJob({ argv, mode, dryRun, label, estLabel, publish, timeout, exec }) {
|
|
306
291
|
const id = newJobId();
|
|
307
292
|
const job = {
|
|
308
293
|
id,
|
|
309
294
|
mode,
|
|
295
|
+
dryRun: dryRun === true,
|
|
310
296
|
label,
|
|
311
297
|
estLabel: estLabel ?? null,
|
|
312
298
|
publish: publish !== false,
|
|
@@ -357,6 +343,22 @@ function launchJob({ argv, mode, label, estLabel, publish, timeout, exec }) {
|
|
|
357
343
|
|
|
358
344
|
/** Build the "STARTED — poll get_run_status" message for a freshly launched job. */
|
|
359
345
|
function formatJobStarted(job) {
|
|
346
|
+
if (job.dryRun) {
|
|
347
|
+
return [
|
|
348
|
+
'DRY-RUN STARTED — computing the queue plan in the background. No tokens',
|
|
349
|
+
'will be spent.',
|
|
350
|
+
'',
|
|
351
|
+
`Job id: ${job.id}`,
|
|
352
|
+
`Running: ${job.label}`,
|
|
353
|
+
'',
|
|
354
|
+
'The harness loads the full ranked queue before printing its plan, which',
|
|
355
|
+
'can take a few minutes — longer than a default 60-second MCP client',
|
|
356
|
+
'request timeout — so it runs detached from this tool call.',
|
|
357
|
+
'',
|
|
358
|
+
`Next: poll get_run_status with { "job_id": "${job.id}" } every ~15-30s until`,
|
|
359
|
+
'it reports COMPLETED. The plan is in the job output.',
|
|
360
|
+
].join('\n');
|
|
361
|
+
}
|
|
360
362
|
const publishLine = job.mode === 'queue'
|
|
361
363
|
? (job.publish
|
|
362
364
|
? 'Each result auto-publishes to the public leaderboard as it finishes.'
|
|
@@ -427,12 +429,15 @@ export function getRunStatus(jobId) {
|
|
|
427
429
|
}
|
|
428
430
|
|
|
429
431
|
if (job.status === 'completed') {
|
|
430
|
-
const closing = job.
|
|
431
|
-
?
|
|
432
|
-
|
|
433
|
-
|
|
434
|
-
|
|
435
|
-
|
|
432
|
+
const closing = job.dryRun
|
|
433
|
+
? 'Dry-run plan above — no tokens were spent. Run again with confirm: true '
|
|
434
|
+
+ 'to execute it.'
|
|
435
|
+
: job.mode === 'queue'
|
|
436
|
+
? (job.publish
|
|
437
|
+
? 'Each result was published to the public leaderboard — call get_results '
|
|
438
|
+
+ '(filtered to this pair/model) to see it.'
|
|
439
|
+
: 'Results were scored but NOT published (validation run).')
|
|
440
|
+
: 'Scored locally (single-item runs are not auto-published).';
|
|
436
441
|
return [
|
|
437
442
|
`COMPLETED — job ${job.id} (took ${secs}s)`,
|
|
438
443
|
job.label,
|
|
@@ -495,8 +500,11 @@ export function resetJobs() {
|
|
|
495
500
|
* subprocess is launched in the BACKGROUND and a job handle is returned
|
|
496
501
|
* immediately, so the call returns well under any MCP client request timeout
|
|
497
502
|
* (the SDK default is 60s, far shorter than a real run). The agent then polls
|
|
498
|
-
* get_run_status with the returned job id until the job settles.
|
|
499
|
-
*
|
|
503
|
+
* get_run_status with the returned job id until the job settles. A queue-mode
|
|
504
|
+
* dry_run is ALSO backgrounded — the harness loads the full ranked queue
|
|
505
|
+
* before printing its plan, which can take minutes. Item-mode dry_run and the
|
|
506
|
+
* confirmation prompts stay synchronous (local argv construction, no
|
|
507
|
+
* subprocess).
|
|
500
508
|
*
|
|
501
509
|
* @param {object} params
|
|
502
510
|
* @param {number} [params.budget] Budget cap in USD
|
|
@@ -521,7 +529,7 @@ export async function runBenchmark(
|
|
|
521
529
|
) {
|
|
522
530
|
const {
|
|
523
531
|
isMtEvalInstalled: checkInstalled = isMtEvalInstalled,
|
|
524
|
-
|
|
532
|
+
lookupQueueItem = null,
|
|
525
533
|
execCapture: exec = execCapture,
|
|
526
534
|
} = deps;
|
|
527
535
|
|
|
@@ -532,7 +540,7 @@ export async function runBenchmark(
|
|
|
532
540
|
'mt-eval is not installed on this machine.',
|
|
533
541
|
'',
|
|
534
542
|
'To install it, run:',
|
|
535
|
-
' pipx install mt-eval',
|
|
543
|
+
' pipx install mt-eval-harness',
|
|
536
544
|
'',
|
|
537
545
|
'After installation, set your API key:',
|
|
538
546
|
' export OPENROUTER_API_KEY=sk-or-...',
|
|
@@ -543,13 +551,23 @@ export async function runBenchmark(
|
|
|
543
551
|
|
|
544
552
|
// ----- Specific item -----------------------------------------------------
|
|
545
553
|
if (item_id) {
|
|
546
|
-
|
|
547
|
-
|
|
548
|
-
const
|
|
549
|
-
|
|
554
|
+
// Direct primary-key lookup — the 0.1.0 code drained the entire ranked
|
|
555
|
+
// queue (211k+ items, minutes of paging) to run one .find().
|
|
556
|
+
const lookup = lookupQueueItem
|
|
557
|
+
?? (await import('./queue.js')).lookupQueueItem;
|
|
558
|
+
const { item, covered } = await lookup({ id: item_id });
|
|
550
559
|
if (!item) {
|
|
551
560
|
return `Queue item "${item_id}" not found. Use list_queue to see available items.`;
|
|
552
561
|
}
|
|
562
|
+
if (covered === true) {
|
|
563
|
+
return [
|
|
564
|
+
`REFUSED — queue item "${item_id}" is already covered by a VERIFIED run.`,
|
|
565
|
+
'',
|
|
566
|
+
'Running it again would spend tokens re-measuring a combination the',
|
|
567
|
+
'leaderboard already has a refereed result for. Use list_queue to pick',
|
|
568
|
+
'an open item instead (open items are coverage-filtered automatically).',
|
|
569
|
+
].join('\n');
|
|
570
|
+
}
|
|
553
571
|
|
|
554
572
|
// Reconstruct a shell-free argv from STRUCTURED fields. The item's
|
|
555
573
|
// network-supplied run_command is NEVER executed.
|
|
@@ -638,26 +656,26 @@ export async function runBenchmark(
|
|
|
638
656
|
|
|
639
657
|
if (dry_run) {
|
|
640
658
|
args.push('--dry-run');
|
|
641
|
-
|
|
659
|
+
// Background even the dry-run: `mt-eval queue` loads the FULL ranked
|
|
660
|
+
// queue before printing its plan (the DB queue is 211k+ items — minutes
|
|
661
|
+
// of paging on harness versions that drain it), so awaiting it inline
|
|
662
|
+
// blew the client's 60s request timeout every time. The plan lands in
|
|
663
|
+
// the job's output; the agent polls get_run_status like any run.
|
|
664
|
+
const scope = budget != null
|
|
665
|
+
? `dry-run plan for up to $${Number(budget).toFixed(2)}`
|
|
666
|
+
: top != null
|
|
667
|
+
? `dry-run plan for the top ${top} item(s)`
|
|
668
|
+
: 'dry-run plan for the full queue';
|
|
669
|
+
const job = launchJob({
|
|
670
|
+
argv: args,
|
|
671
|
+
mode: 'queue',
|
|
672
|
+
dryRun: true,
|
|
673
|
+
label: scope,
|
|
674
|
+
publish,
|
|
642
675
|
timeout: 1800_000,
|
|
676
|
+
exec,
|
|
643
677
|
});
|
|
644
|
-
|
|
645
|
-
// Never relay the raw harness output here — a Python traceback carries
|
|
646
|
-
// local absolute paths. Surface the final error line (path-stripped)
|
|
647
|
-
// and keep the full detail in the debug log only.
|
|
648
|
-
await writeDebugLog(
|
|
649
|
-
`queue dry-run failed (exit ${code}) — argv: mt-eval ${args.join(' ')}`,
|
|
650
|
-
debugDetail(stdout, stderr),
|
|
651
|
-
);
|
|
652
|
-
return [
|
|
653
|
-
`Queue dry-run failed (exit code ${code}): ${finalErrorLine(stderr, stdout)}`,
|
|
654
|
-
'',
|
|
655
|
-
'Hint: this error came from the local mt-eval harness, not the queue. '
|
|
656
|
-
+ 'Run the same mt-eval command in a terminal to reproduce; the full '
|
|
657
|
-
+ `unedited output was saved to the debug log (${DEBUG_LOG_HINT}).`,
|
|
658
|
-
].join('\n');
|
|
659
|
-
}
|
|
660
|
-
return stdout || 'Queue dry-run completed (no output captured).';
|
|
678
|
+
return formatJobStarted(job);
|
|
661
679
|
}
|
|
662
680
|
|
|
663
681
|
// ENFORCED bound on scope. Without a selector, argv is
|
package/src/tools/queue.js
CHANGED
|
@@ -3,11 +3,20 @@
|
|
|
3
3
|
*
|
|
4
4
|
* Data source (default): the live queue served from Postgres via the queue_top
|
|
5
5
|
* RPC — a ranked list of (corpus, model, condition) items, coverage-filtered
|
|
6
|
-
* against VERIFIED runs, so nothing stale or already-done is shown.
|
|
7
|
-
*
|
|
8
|
-
*
|
|
9
|
-
*
|
|
10
|
-
*
|
|
6
|
+
* against VERIFIED runs, so nothing stale or already-done is shown.
|
|
7
|
+
*
|
|
8
|
+
* The queue is six figures deep (211k+ open items as of 2026-08), so the DB
|
|
9
|
+
* path NEVER drains it: metadata comes from the small queue-preview.json (with
|
|
10
|
+
* a live open-item count from the unpaged queue_pairs RPC), ranked items are
|
|
11
|
+
* paged from queue_top only as deep as the caller's selection needs (bounded
|
|
12
|
+
* by MAX_DB_PAGES), and single-item lookups go straight to the queue_items
|
|
13
|
+
* primary key over PostgREST. Draining was the 0.1.0 failure mode: 423
|
|
14
|
+
* sequential pages ≈ 3 minutes, past every MCP client's 60s request timeout.
|
|
15
|
+
*
|
|
16
|
+
* If the DB is unreachable, every entry point falls back to the static
|
|
17
|
+
* queue.json blob, so the tools never break. Set CHAMPOLLION_QUEUE_SOURCE=blob
|
|
18
|
+
* to force the blob. An in-memory CACHE_TTL_MS cache (a growing ranked prefix,
|
|
19
|
+
* never re-fetched from page 0) avoids re-fetching on every tool call.
|
|
11
20
|
*/
|
|
12
21
|
|
|
13
22
|
const QUEUE_URL = 'https://champollion.dev/queue.json';
|
|
@@ -15,7 +24,9 @@ const QUEUE_URL = 'https://champollion.dev/queue.json';
|
|
|
15
24
|
// coverage-filtered against VERIFIED runs, so items are never stale. Metadata
|
|
16
25
|
// (open_items, models, priority_model, cost_basis, how_to_run) comes from the
|
|
17
26
|
// small queue-preview.json companion. The full static blob remains the FALLBACK
|
|
18
|
-
// so this tool never breaks if the DB is unreachable.
|
|
27
|
+
// so this tool never breaks if the DB is unreachable. (The blob itself is a
|
|
28
|
+
// self-describing top slice when the ranking outgrows its size cap — see
|
|
29
|
+
// metadata.blob_truncated — so "fallback" never silently means "everything".)
|
|
19
30
|
const PREVIEW_URL = 'https://champollion.dev/queue-preview.json';
|
|
20
31
|
const SUPABASE_URL = process.env.MT_EVAL_SUPABASE_URL
|
|
21
32
|
|| 'https://sjdomynysdljkbemupqa.supabase.co';
|
|
@@ -26,6 +37,21 @@ const QUEUE_TOP_PAGE = 500; // matches the RPC's hard page cap
|
|
|
26
37
|
// the legacy static file (used by the existing fetch tests).
|
|
27
38
|
const QUEUE_SOURCE = process.env.CHAMPOLLION_QUEUE_SOURCE || 'db';
|
|
28
39
|
const CACHE_TTL_MS = 5 * 60 * 1000; // 5 minutes
|
|
40
|
+
// Selection-depth bound for the DB path: at most this many queue_top pages per
|
|
41
|
+
// cache generation (default 20 → 10,000 ranked rows ≈ 8-10s of paging). Every
|
|
42
|
+
// consumer needs ≤500 post-filter items, so the bound only bites on narrow
|
|
43
|
+
// filters over a huge queue — and then the tools SAY how deep they looked
|
|
44
|
+
// (no silent caps). Env-tunable for agents that want to scan deeper.
|
|
45
|
+
const MAX_DB_PAGES = (() => {
|
|
46
|
+
const n = Number.parseInt(process.env.CHAMPOLLION_QUEUE_MAX_PAGES ?? '', 10);
|
|
47
|
+
return Number.isInteger(n) && n > 0 ? n : 20;
|
|
48
|
+
})();
|
|
49
|
+
// Wall-clock budget for one exported queue operation's DB round trips.
|
|
50
|
+
// Chosen so that even the worst ladder — DB hangs to the full deadline, THEN
|
|
51
|
+
// the blob fallback takes its whole 30s request window — still lands under
|
|
52
|
+
// the 60s MCP client default (25s + 30s + parse < 60s). A slow network
|
|
53
|
+
// degrades to a truncated-but-honest answer, not a dead tool.
|
|
54
|
+
const DB_DEADLINE_MS = 25_000;
|
|
29
55
|
|
|
30
56
|
// Served item fields (what queue.json publishes / consumers rely on). queue_top
|
|
31
57
|
// rows also carry rank_mode/map_value/diagnostics/generation_id/generated_at —
|
|
@@ -37,103 +63,387 @@ const SERVED_FIELDS = [
|
|
|
37
63
|
'run_command',
|
|
38
64
|
];
|
|
39
65
|
|
|
40
|
-
|
|
66
|
+
// The queue cache: ONE generation at a time. On the DB path `items` is a
|
|
67
|
+
// ranked PREFIX that grows in place as callers ask deeper (never re-fetching
|
|
68
|
+
// page 0); on the blob path it is whatever the blob shipped. `complete` means
|
|
69
|
+
// "we have seen the end of the served ranking". A generation lives CACHE_TTL_MS
|
|
70
|
+
// from its first fetch, then the whole thing resets.
|
|
71
|
+
let _cache = null; // { source, metadata, items, complete, pages }
|
|
41
72
|
let _cacheTime = 0;
|
|
42
73
|
|
|
74
|
+
/** Standard headers for Supabase REST/RPC calls (anon key is publishable). */
|
|
75
|
+
function dbHeaders() {
|
|
76
|
+
return {
|
|
77
|
+
'apikey': SUPABASE_ANON_KEY,
|
|
78
|
+
'Authorization': `Bearer ${SUPABASE_ANON_KEY}`,
|
|
79
|
+
'Content-Type': 'application/json',
|
|
80
|
+
'Accept': 'application/json',
|
|
81
|
+
};
|
|
82
|
+
}
|
|
83
|
+
|
|
84
|
+
/** Per-request abort signal that also respects the operation deadline. */
|
|
85
|
+
function signalFor(deadline) {
|
|
86
|
+
const remaining = deadline - Date.now();
|
|
87
|
+
if (remaining <= 0) throw new Error('queue operation deadline exceeded');
|
|
88
|
+
return AbortSignal.timeout(Math.min(30_000, remaining));
|
|
89
|
+
}
|
|
90
|
+
|
|
91
|
+
/** Project a queue_top/queue_items row down to the served item shape. */
|
|
92
|
+
function projectRow(row) {
|
|
93
|
+
const item = {};
|
|
94
|
+
for (const f of SERVED_FIELDS) if (row[f] !== undefined) item[f] = row[f];
|
|
95
|
+
// The restricted-corpus `transmission` stamp is a SERVED extra on
|
|
96
|
+
// queue.json but not a queue_items COLUMN — the ranker writes it into
|
|
97
|
+
// the diagnostics JSONB. Projecting columns alone dropped it from every
|
|
98
|
+
// DB-served item, so agents pulling work through MCP lost the no-train
|
|
99
|
+
// channel requirement the blob discloses. Lift it back.
|
|
100
|
+
const stamp = row.diagnostics?.transmission;
|
|
101
|
+
if (stamp && typeof stamp === 'object' && stamp.policy) item.transmission = stamp;
|
|
102
|
+
return item;
|
|
103
|
+
}
|
|
104
|
+
|
|
105
|
+
/** Return the fresh cache generation for `source`, or null. */
|
|
106
|
+
function freshCache(source) {
|
|
107
|
+
if (_cache && _cache.source === source
|
|
108
|
+
&& (Date.now() - _cacheTime) < CACHE_TTL_MS) {
|
|
109
|
+
return _cache;
|
|
110
|
+
}
|
|
111
|
+
return null;
|
|
112
|
+
}
|
|
113
|
+
|
|
43
114
|
/**
|
|
44
|
-
*
|
|
115
|
+
* Ensure a cache generation exists: metadata loaded, items array started.
|
|
45
116
|
*
|
|
46
|
-
*
|
|
47
|
-
*
|
|
48
|
-
*
|
|
49
|
-
*
|
|
50
|
-
*
|
|
51
|
-
* and parsed here, so a non-JSON response becomes ONE clean, user-facing
|
|
52
|
-
* error at this shared choke point. Failures are never cached — the next
|
|
53
|
-
* call re-fetches.
|
|
117
|
+
* DB path: metadata from queue-preview.json, then the LIVE open-item count
|
|
118
|
+
* from one unpaged queue_pairs call (SUM of per-pair counts — same verified-
|
|
119
|
+
* coverage filter as queue_top, so it is the true served total without
|
|
120
|
+
* draining anything). If queue_pairs fails, the preview's generation-time
|
|
121
|
+
* open_items stands, stamped `open_items_basis: 'generation'`.
|
|
54
122
|
*
|
|
55
|
-
*
|
|
56
|
-
*
|
|
57
|
-
*
|
|
123
|
+
* Blob path: the whole blob (metadata + items, complete by definition —
|
|
124
|
+
* "complete" meaning the end of what the blob SERVES; a size-capped blob
|
|
125
|
+
* says so itself via metadata.blob_truncated).
|
|
126
|
+
*
|
|
127
|
+
* Throws on failure so callers can fall back — never caches a failure.
|
|
58
128
|
*/
|
|
59
|
-
|
|
60
|
-
const
|
|
61
|
-
if (
|
|
62
|
-
return _cache;
|
|
63
|
-
}
|
|
129
|
+
async function ensureCache(fetchImpl, source, deadline) {
|
|
130
|
+
const hit = freshCache(source);
|
|
131
|
+
if (hit) return hit;
|
|
64
132
|
|
|
65
|
-
let
|
|
133
|
+
let gen;
|
|
66
134
|
if (source === 'blob') {
|
|
67
|
-
data = await fetchQueueFromBlob(fetchImpl);
|
|
135
|
+
const data = await fetchQueueFromBlob(fetchImpl);
|
|
136
|
+
gen = {
|
|
137
|
+
source, metadata: data.metadata, items: data.items,
|
|
138
|
+
complete: true, pages: 0,
|
|
139
|
+
};
|
|
68
140
|
} else {
|
|
69
|
-
//
|
|
70
|
-
//
|
|
141
|
+
// Metadata + preview from the companion file (a few hundred KB — ~292 KB
|
|
142
|
+
// as of 2026-08; it grows with the queue, so never assume it is tiny).
|
|
143
|
+
const pResp = await fetchImpl(PREVIEW_URL, {
|
|
144
|
+
headers: { 'Accept': 'application/json' },
|
|
145
|
+
signal: signalFor(deadline),
|
|
146
|
+
});
|
|
147
|
+
if (!pResp.ok) throw new Error(`preview HTTP ${pResp.status}`);
|
|
148
|
+
const preview = JSON.parse(await pResp.text());
|
|
149
|
+
const metadata = preview?.metadata;
|
|
150
|
+
if (metadata === null || typeof metadata !== 'object') {
|
|
151
|
+
throw new Error('queue-preview.json missing metadata');
|
|
152
|
+
}
|
|
153
|
+
gen = {
|
|
154
|
+
source,
|
|
155
|
+
metadata: { ...metadata, open_items_basis: 'generation' },
|
|
156
|
+
items: [],
|
|
157
|
+
complete: false,
|
|
158
|
+
pages: 0,
|
|
159
|
+
};
|
|
160
|
+
// Live open count — best-effort; the generation-time count stands if the
|
|
161
|
+
// aggregate RPC is down (the ranked items themselves are unaffected).
|
|
71
162
|
try {
|
|
72
|
-
|
|
163
|
+
const resp = await fetchImpl(`${SUPABASE_URL}/rest/v1/rpc/queue_pairs`, {
|
|
164
|
+
method: 'POST',
|
|
165
|
+
headers: dbHeaders(),
|
|
166
|
+
body: JSON.stringify({ p_rank_mode: metadata.rank_mode || 'map' }),
|
|
167
|
+
signal: signalFor(deadline),
|
|
168
|
+
});
|
|
169
|
+
if (resp.ok) {
|
|
170
|
+
const pairs = JSON.parse(await resp.text());
|
|
171
|
+
if (Array.isArray(pairs)) {
|
|
172
|
+
// Live count = SUM over per-pair counts. Never used to short-circuit
|
|
173
|
+
// paging — a misreported 0 must not suppress the ranked fetch.
|
|
174
|
+
const live = pairs.reduce((s, p) => s + (Number(p.item_count) || 0), 0);
|
|
175
|
+
gen.metadata.open_items = live;
|
|
176
|
+
gen.metadata.open_items_basis = 'live';
|
|
177
|
+
}
|
|
178
|
+
}
|
|
73
179
|
} catch {
|
|
74
|
-
|
|
180
|
+
// keep generation-time open_items
|
|
75
181
|
}
|
|
76
182
|
}
|
|
77
183
|
|
|
78
|
-
_cache =
|
|
79
|
-
_cacheTime = now;
|
|
80
|
-
return
|
|
184
|
+
_cache = gen;
|
|
185
|
+
_cacheTime = Date.now();
|
|
186
|
+
return gen;
|
|
81
187
|
}
|
|
82
188
|
|
|
83
|
-
/**
|
|
84
|
-
*
|
|
85
|
-
* Throws on
|
|
86
|
-
async function
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
189
|
+
/** Fetch ONE queue_top page at the generation's current watermark and append
|
|
190
|
+
* it. Returns the number of rows appended. Marks the generation complete on
|
|
191
|
+
* a short page. Throws on failure (callers decide whether that is fatal). */
|
|
192
|
+
async function fetchNextPage(fetchImpl, gen, deadline) {
|
|
193
|
+
const resp = await fetchImpl(`${SUPABASE_URL}/rest/v1/rpc/queue_top`, {
|
|
194
|
+
method: 'POST',
|
|
195
|
+
headers: dbHeaders(),
|
|
196
|
+
body: JSON.stringify({
|
|
197
|
+
p_rank_mode: gen.metadata.rank_mode || 'map',
|
|
198
|
+
p_limit: QUEUE_TOP_PAGE,
|
|
199
|
+
p_offset: gen.items.length,
|
|
200
|
+
}),
|
|
201
|
+
signal: signalFor(deadline),
|
|
92
202
|
});
|
|
93
|
-
if (!
|
|
94
|
-
const
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
203
|
+
if (!resp.ok) throw new Error(`queue_top HTTP ${resp.status}`);
|
|
204
|
+
const page = JSON.parse(await resp.text());
|
|
205
|
+
if (!Array.isArray(page)) throw new Error('queue_top did not return an array');
|
|
206
|
+
for (const row of page) gen.items.push(projectRow(row));
|
|
207
|
+
gen.pages += 1;
|
|
208
|
+
if (page.length < QUEUE_TOP_PAGE) gen.complete = true;
|
|
209
|
+
return page.length;
|
|
210
|
+
}
|
|
211
|
+
|
|
212
|
+
/**
|
|
213
|
+
* Deepen the generation's ranked prefix until it satisfies `isSatisfied(items)`
|
|
214
|
+
* or a stop condition: ranking complete, MAX_DB_PAGES spent, deadline reached,
|
|
215
|
+
* or a mid-flight page error AFTER some rows already arrived (page-1 errors
|
|
216
|
+
* rethrow so the caller can fall back to the blob; later errors degrade to a
|
|
217
|
+
* truncated-but-honest prefix rather than throwing away everything fetched).
|
|
218
|
+
*/
|
|
219
|
+
async function deepenUntil(fetchImpl, gen, deadline, isSatisfied) {
|
|
220
|
+
while (!gen.complete
|
|
221
|
+
&& !isSatisfied(gen.items)
|
|
222
|
+
&& gen.pages < MAX_DB_PAGES
|
|
223
|
+
&& Date.now() < deadline) {
|
|
224
|
+
try {
|
|
225
|
+
await fetchNextPage(fetchImpl, gen, deadline);
|
|
226
|
+
} catch (err) {
|
|
227
|
+
if (gen.items.length === 0) throw err;
|
|
228
|
+
break; // keep the partial prefix; tools report the depth honestly
|
|
229
|
+
}
|
|
98
230
|
}
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
|
|
122
|
-
|
|
123
|
-
|
|
124
|
-
|
|
125
|
-
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
|
|
231
|
+
}
|
|
232
|
+
|
|
233
|
+
/**
|
|
234
|
+
* Fetch queue metadata only — no ranked items. This is what get_project_info
|
|
235
|
+
* and any stats display should use: on the DB path it costs one preview GET
|
|
236
|
+
* plus one unpaged aggregate RPC, never a queue_top page.
|
|
237
|
+
*
|
|
238
|
+
* Falls back to the blob's metadata if the DB path fails outright.
|
|
239
|
+
*
|
|
240
|
+
* @returns {Promise<{ metadata: object }>}
|
|
241
|
+
*/
|
|
242
|
+
export async function fetchQueueMeta({ fetchImpl = fetch, source = QUEUE_SOURCE } = {}) {
|
|
243
|
+
const deadline = Date.now() + DB_DEADLINE_MS;
|
|
244
|
+
if (source === 'blob') {
|
|
245
|
+
const gen = await ensureCache(fetchImpl, 'blob', deadline);
|
|
246
|
+
return { metadata: gen.metadata };
|
|
247
|
+
}
|
|
248
|
+
try {
|
|
249
|
+
const gen = await ensureCache(fetchImpl, source, deadline);
|
|
250
|
+
return { metadata: gen.metadata };
|
|
251
|
+
} catch (dbErr) {
|
|
252
|
+
return { metadata: (await blobFallback(fetchImpl, dbErr, source)).metadata };
|
|
253
|
+
}
|
|
254
|
+
}
|
|
255
|
+
|
|
256
|
+
/**
|
|
257
|
+
* Fetch enough of the ranked queue to satisfy a filterQueue selection, then
|
|
258
|
+
* run the (untouched, SSOT-locked) filterQueue over it.
|
|
259
|
+
*
|
|
260
|
+
* The deepening loop exists because filter/budget semantics make "rows
|
|
261
|
+
* scanned" ≠ "items selected": budget mode SKIPS over-budget items and keeps
|
|
262
|
+
* walking, and narrow language/model filters may match sparsely. So we page,
|
|
263
|
+
* re-filter, and page again until the selection is satisfied or a bound hits.
|
|
264
|
+
*
|
|
265
|
+
* @param {object} [opts] filterQueue filters + { limit, fetchImpl, source }
|
|
266
|
+
* @returns {Promise<{
|
|
267
|
+
* selected: object[], // filterQueue's picks, in ranking order
|
|
268
|
+
* scanned: object[], // the ranked prefix examined (for estimateCost)
|
|
269
|
+
* metadata: object,
|
|
270
|
+
* complete: boolean, // true = the WHOLE served ranking was examined
|
|
271
|
+
* scannedRows: number, // how deep the examination went
|
|
272
|
+
* }>}
|
|
273
|
+
*/
|
|
274
|
+
export async function selectFromQueue({
|
|
275
|
+
budget = null, language = null, source_language = null, model = null,
|
|
276
|
+
condition = null, limit = 20,
|
|
277
|
+
fetchImpl = fetch, source = QUEUE_SOURCE,
|
|
278
|
+
} = {}) {
|
|
279
|
+
const deadline = Date.now() + DB_DEADLINE_MS;
|
|
280
|
+
const filters = { budget, language, source_language, model, condition, limit };
|
|
281
|
+
|
|
282
|
+
let gen;
|
|
283
|
+
if (source === 'blob') {
|
|
284
|
+
gen = await ensureCache(fetchImpl, 'blob', deadline);
|
|
285
|
+
} else {
|
|
286
|
+
try {
|
|
287
|
+
gen = await ensureCache(fetchImpl, source, deadline);
|
|
288
|
+
await deepenUntil(fetchImpl, gen, deadline,
|
|
289
|
+
(items) => filterQueue(items, filters).length >= limit);
|
|
290
|
+
} catch (dbErr) {
|
|
291
|
+
gen = await blobFallback(fetchImpl, dbErr, source);
|
|
292
|
+
}
|
|
293
|
+
}
|
|
294
|
+
|
|
295
|
+
return {
|
|
296
|
+
selected: filterQueue(gen.items, filters),
|
|
297
|
+
scanned: gen.items,
|
|
298
|
+
metadata: gen.metadata,
|
|
299
|
+
complete: gen.complete === true,
|
|
300
|
+
scannedRows: gen.items.length,
|
|
301
|
+
};
|
|
302
|
+
}
|
|
303
|
+
|
|
304
|
+
/**
|
|
305
|
+
* Look up ONE queue item by id or priority rank — without touching the ranked
|
|
306
|
+
* paging at all. On the DB path this is a primary-key (or mode+priority index)
|
|
307
|
+
* read on queue_items over PostgREST, plus a one-row run_cards probe that
|
|
308
|
+
* replicates queue_top's verified-coverage filter (queue_items itself is the
|
|
309
|
+
* raw registered set, so a row can exist yet already be verified-covered).
|
|
310
|
+
*
|
|
311
|
+
* NOTE: never offset arithmetic — queue_top's coverage filter means row N of
|
|
312
|
+
* the served ranking is NOT the row with priority N, so priority lookups
|
|
313
|
+
* match the stored priority field, exactly like the in-memory getQueueItem.
|
|
314
|
+
*
|
|
315
|
+
* @param {{ id?: string, priority?: number, fetchImpl?, source? }} opts
|
|
316
|
+
* @returns {Promise<{
|
|
317
|
+
* item: object|null,
|
|
318
|
+
* covered: boolean|null, // true = exists but a VERIFIED run already covers
|
|
319
|
+
* // it (not an open work item); null = probe failed
|
|
320
|
+
* truncatedNote: string|null, // set when a not-found came from a truncated
|
|
321
|
+
* // blob fallback and the item might exist deeper
|
|
322
|
+
* }>}
|
|
323
|
+
*/
|
|
324
|
+
export async function lookupQueueItem({
|
|
325
|
+
id, priority, fetchImpl = fetch, source = QUEUE_SOURCE,
|
|
326
|
+
} = {}) {
|
|
327
|
+
const deadline = Date.now() + DB_DEADLINE_MS;
|
|
328
|
+
|
|
329
|
+
const fromItems = (gen) => {
|
|
330
|
+
const item = getQueueItem(gen.items, { id, priority });
|
|
331
|
+
const truncated = !item && gen.metadata?.blob_truncated
|
|
332
|
+
? `Note: the fallback queue snapshot is a top slice (${gen.metadata.blob_truncated.kept} of ${gen.metadata.blob_truncated.total} items) — the item may exist deeper in the live queue.`
|
|
333
|
+
: null;
|
|
334
|
+
// Blob/queue.json items are coverage-filtered at generation time, so a hit
|
|
335
|
+
// there is an open item by construction.
|
|
336
|
+
return { item, covered: item ? false : null, truncatedNote: truncated };
|
|
337
|
+
};
|
|
338
|
+
|
|
339
|
+
if (source === 'blob') {
|
|
340
|
+
return fromItems(await ensureCache(fetchImpl, 'blob', deadline));
|
|
341
|
+
}
|
|
342
|
+
|
|
343
|
+
try {
|
|
344
|
+
const gen = await ensureCache(fetchImpl, source, deadline);
|
|
345
|
+
const rankMode = gen.metadata.rank_mode || 'map';
|
|
346
|
+
const query = id
|
|
347
|
+
? `id=eq.${encodeURIComponent(id)}`
|
|
348
|
+
: `rank_mode=eq.${encodeURIComponent(rankMode)}&priority=eq.${Number(priority)}`;
|
|
349
|
+
const resp = await fetchImpl(
|
|
350
|
+
`${SUPABASE_URL}/rest/v1/queue_items?${query}&limit=1`,
|
|
351
|
+
{ headers: dbHeaders(), signal: signalFor(deadline) },
|
|
352
|
+
);
|
|
353
|
+
if (!resp.ok) throw new Error(`queue_items lookup HTTP ${resp.status}`);
|
|
354
|
+
const rows = JSON.parse(await resp.text());
|
|
355
|
+
if (!Array.isArray(rows)) throw new Error('queue_items lookup did not return an array');
|
|
356
|
+
if (rows.length === 0) return { item: null, covered: null, truncatedNote: null };
|
|
357
|
+
|
|
358
|
+
const row = rows[0];
|
|
359
|
+
const item = projectRow(row);
|
|
360
|
+
|
|
361
|
+
// Coverage probe — the same NOT EXISTS queue_top applies (migration 059):
|
|
362
|
+
// a verified run_card for (corpus, model, condition) closes the item.
|
|
363
|
+
let covered = null;
|
|
364
|
+
try {
|
|
365
|
+
const probeQ = `dataset_id=eq.${encodeURIComponent(row.corpus_id)}`
|
|
366
|
+
+ `&model_slug=eq.${encodeURIComponent(row.model)}`
|
|
367
|
+
+ `&condition=eq.${encodeURIComponent(row.condition)}`
|
|
368
|
+
+ '&trust=eq.verified&select=id&limit=1';
|
|
369
|
+
const probe = await fetchImpl(
|
|
370
|
+
`${SUPABASE_URL}/rest/v1/run_cards?${probeQ}`,
|
|
371
|
+
{ headers: dbHeaders(), signal: signalFor(deadline) },
|
|
372
|
+
);
|
|
373
|
+
if (probe.ok) {
|
|
374
|
+
const hits = JSON.parse(await probe.text());
|
|
375
|
+
if (Array.isArray(hits)) covered = hits.length > 0;
|
|
376
|
+
}
|
|
377
|
+
} catch {
|
|
378
|
+
// covered stays null (unknown) — the item itself is still served
|
|
131
379
|
}
|
|
132
|
-
|
|
380
|
+
return { item, covered, truncatedNote: null };
|
|
381
|
+
} catch (dbErr) {
|
|
382
|
+
return fromItems(await blobFallback(fetchImpl, dbErr, source));
|
|
383
|
+
}
|
|
384
|
+
}
|
|
385
|
+
|
|
386
|
+
/** Fall back to the blob after a DB-path failure, chaining both causes if the
|
|
387
|
+
* blob is down too — the agent should see WHY both lanes failed, not just
|
|
388
|
+
* the second one. Caches the blob generation so retries stay cheap. */
|
|
389
|
+
async function blobFallback(fetchImpl, dbErr, source = QUEUE_SOURCE) {
|
|
390
|
+
try {
|
|
391
|
+
const data = await fetchQueueFromBlob(fetchImpl);
|
|
392
|
+
// Cached under the CALLER's source key so retries within the TTL reuse
|
|
393
|
+
// the blob instead of hammering a DB that just failed.
|
|
394
|
+
const gen = {
|
|
395
|
+
source, metadata: data.metadata, items: data.items,
|
|
396
|
+
complete: true, pages: 0,
|
|
397
|
+
};
|
|
398
|
+
_cache = gen;
|
|
399
|
+
_cacheTime = Date.now();
|
|
400
|
+
return gen;
|
|
401
|
+
} catch (blobErr) {
|
|
402
|
+
throw new Error(
|
|
403
|
+
`live queue unavailable (${dbErr.message}) and the static fallback also `
|
|
404
|
+
+ `failed (${blobErr.message})`,
|
|
405
|
+
);
|
|
133
406
|
}
|
|
407
|
+
}
|
|
134
408
|
|
|
135
|
-
|
|
136
|
-
|
|
409
|
+
/**
|
|
410
|
+
* Fetch the queue, returning the cached version if still fresh.
|
|
411
|
+
*
|
|
412
|
+
* LEGACY entry point: returns a `{ metadata, items }` queue object. On the
|
|
413
|
+
* blob path this is the whole blob, exactly as before. On the DB path it is
|
|
414
|
+
* now a BOUNDED ranked prefix (up to MAX_DB_PAGES pages), never the 0.1.0
|
|
415
|
+
* full drain — prefer fetchQueueMeta / selectFromQueue / lookupQueueItem,
|
|
416
|
+
* which fetch only what the caller's question needs.
|
|
417
|
+
*
|
|
418
|
+
* The static host can serve an HTML holding page with HTTP 200 (site gated,
|
|
419
|
+
* maintenance, CDN error page) — `resp.json()` would then surface a raw
|
|
420
|
+
* `SyntaxError: Unexpected token '<'` to every queue-backed tool. The body is
|
|
421
|
+
* therefore read as text and parsed at one choke point, so a non-JSON
|
|
422
|
+
* response becomes ONE clean, user-facing error. Failures are never cached —
|
|
423
|
+
* the next call re-fetches.
|
|
424
|
+
*
|
|
425
|
+
* @param {object} [opts]
|
|
426
|
+
* @param {typeof fetch} [opts.fetchImpl] Injectable fetch (for tests).
|
|
427
|
+
* @returns {Promise<{ metadata: object, items: object[] }>}
|
|
428
|
+
*/
|
|
429
|
+
export async function fetchQueue({ fetchImpl = fetch, source = QUEUE_SOURCE } = {}) {
|
|
430
|
+
const deadline = Date.now() + DB_DEADLINE_MS;
|
|
431
|
+
let gen;
|
|
432
|
+
if (source === 'blob') {
|
|
433
|
+
gen = await ensureCache(fetchImpl, 'blob', deadline);
|
|
434
|
+
} else {
|
|
435
|
+
try {
|
|
436
|
+
gen = await ensureCache(fetchImpl, source, deadline);
|
|
437
|
+
await deepenUntil(fetchImpl, gen, deadline, () => false);
|
|
438
|
+
} catch (dbErr) {
|
|
439
|
+
gen = await blobFallback(fetchImpl, dbErr, source);
|
|
440
|
+
}
|
|
441
|
+
}
|
|
442
|
+
// One stable view per generation — callers within a TTL get the SAME object
|
|
443
|
+
// (metadata/items are the live cache references, so a deepened prefix shows
|
|
444
|
+
// through), preserving the original fetchQueue cache contract.
|
|
445
|
+
if (!gen.legacyView) gen.legacyView = { metadata: gen.metadata, items: gen.items };
|
|
446
|
+
return gen.legacyView;
|
|
137
447
|
}
|
|
138
448
|
|
|
139
449
|
/** The legacy path: the full static queue.json blob. Also the DB-path fallback. */
|