champollion-mcp-server 0.1.0 → 0.1.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md ADDED
@@ -0,0 +1,34 @@
1
+ # Changelog
2
+
3
+ ## 0.1.1 (2026-08-27)
4
+
5
+ Fixes the queue-tool timeouts found by live testing of the 0.1.0 npm release:
6
+ with the queue at 211k+ open items, `list_queue`, `estimate_cost`,
7
+ `get_project_info`, and `run_benchmark` all exceeded MCP clients' default 60s
8
+ request timeout, because the DB fetch path drained the entire `queue_top`
9
+ ranking (~423 sequential pages ≈ 3 minutes) before answering anything.
10
+
11
+ - **Bounded, purpose-fit queue fetching.** Metadata comes from
12
+ queue-preview.json plus a live open-item count from the unpaged
13
+ `queue_pairs` RPC; ranked items are paged from `queue_top` only as deep as
14
+ the caller's selection needs (fetch-until-satisfied, bounded by
15
+ `CHAMPOLLION_QUEUE_MAX_PAGES`, default 20 pages / 10,000 rows); single
16
+ items are read by primary key over PostgREST with a verified-coverage
17
+ probe. When a bound truncates a search, the tool says how deep it looked —
18
+ no silent caps.
19
+ - **`get_queue_item` / `run_benchmark(item_id)`** now do a direct by-id (or
20
+ mode+priority) lookup instead of scanning a full drain, and refuse items
21
+ already covered by a VERIFIED run instead of re-spending on them.
22
+ - **Queue-mode `dry_run` is backgrounded** like a real run (the installed
23
+ harness loads the full queue before printing its plan): it returns a job id
24
+ immediately; the plan arrives via `get_run_status`.
25
+ - **Failure ladder hardened.** A slow-but-alive DB degrades to a
26
+ truncated-but-honest prefix; a dead DB falls back to the static queue.json
27
+ blob; when both are down the error names both causes.
28
+ - Harness (mt-eval, monorepo): `--top N` runs now page the DB queue only as
29
+ deep as selection needs, and `CHAMPOLLION_QUEUE_SOURCE=blob` is accepted as
30
+ a sentinel (previously read as a literal file path).
31
+
32
+ ## 0.1.0 (2026-08-27)
33
+
34
+ Initial npm release.
package/README.md CHANGED
@@ -27,7 +27,7 @@ When connected to an agent (Claude Code, Antigravity, Cursor, etc.), the server
27
27
 
28
28
  #### Training tools (nmt-forge)
29
29
 
30
- These wrap the [nmt-forge](https://github.com/gamedaysuits/Champollion) training suite. forge is part of the Champollion monorepo (not on PyPI) — clone the repo and set `CHAMPOLLION_FORGE_DIR` to its `forge/` directory; scoring additionally needs the eval harness (`pip install mt-eval`). Without forge present these tools return an actionable error rather than crashing.
30
+ These wrap the [nmt-forge](https://github.com/gamedaysuits/Champollion) training suite. forge is part of the Champollion monorepo (not on PyPI) — clone the repo and set `CHAMPOLLION_FORGE_DIR` to its `forge/` directory; scoring additionally needs the eval harness (`pip install mt-eval-harness`). Without forge present these tools return an actionable error rather than crashing.
31
31
 
32
32
  | Tool | Type | Description |
33
33
  |---|---|---|
@@ -132,7 +132,7 @@ Once connected, you can talk to your agent naturally:
132
132
  >
133
133
  > **Agent** uses `list_queue` with `budget: 10` → sees what's available
134
134
  >
135
- > **Agent:** "I found a few thousand open benchmark items. Your $10 could fund dozens of runs. Any preference on languages?"
135
+ > **Agent:** "The queue has over 200,000 open benchmark items. Your $10 could fund dozens of runs. Any preference on languages?"
136
136
  >
137
137
  > **You:** "West African languages"
138
138
  >
package/instructions.md CHANGED
@@ -19,13 +19,13 @@ Start with `get_project_info` to understand what Champollion is and how contribu
19
19
 
20
20
  #### Running is asynchronous (important)
21
21
 
22
- A real benchmark runs a corpus through a live model and takes **minutes**, which is longer than the default 60-second request timeout most MCP clients (Claude Code, Cursor) enforce. So `run_benchmark` does **not** wait for the run to finish — it launches the run in the background and returns a `job id` right away. Treat that `STARTED` response as success, **not** completion.
22
+ A real benchmark runs a corpus through a live model and takes **minutes**, which is longer than the default 60-second request timeout most MCP clients (Claude Code, Cursor) enforce. So `run_benchmark` does **not** wait for the run to finish — it launches the run in the background and returns a `job id` right away. Treat that `STARTED` response as success, **not** completion. (A queue-mode `dry_run` is backgrounded the same way — the harness loads the full ranked queue before printing its plan — so poll `get_run_status` for the plan; an `item_id` dry run answers inline.)
23
23
 
24
24
  - After `run_benchmark` returns, call `get_run_status { "job_id": "run-N" }`. Each poll returns instantly: `RUNNING` (keep polling), `COMPLETED` (output is in the response), `FAILED`, or `ERROR`.
25
25
  - Do **not** re-call `run_benchmark` because nothing "came back" — that would start a **second** run and spend tokens twice. The first call already started it; poll `get_run_status` instead.
26
26
  - Jobs live in the server process's memory, so a job id is only pollable from the same session. Call `get_run_status` with no `job_id` to list every job started this session.
27
27
 
28
- The estimate you show in step 4 is what executes: `run_benchmark` runs budget/top items in **deterministic top-of-queue order** (it passes `--no-spread`), so the selection matches the `estimate_cost` / `list_queue` preview item-for-item. (One caveat for honesty: `estimate_cost` samples up to ~500 matching items, so for a very large budget that funds more than that, treat its count/total as a lower bound.) A live run additionally skips any (corpus, model, condition) combo already on the leaderboard, so the executed set can be a subset of the preview — never a different, unseen set.
28
+ The estimate you show in step 4 is what executes: `run_benchmark` runs budget/top items in **deterministic top-of-queue order** (it passes `--no-spread`), so the selection matches the `estimate_cost` / `list_queue` preview item-for-item. (One caveat for honesty: `estimate_cost` samples up to ~500 matching items from a bounded scan of the top of the ranking — its reply says how deep it looked when the bound bites — so for a very large budget or a very narrow filter, treat its count/total as a lower bound.) A live run additionally skips any (corpus, model, condition) combo already on the leaderboard, so the executed set can be a subset of the preview — never a different, unseen set.
29
29
 
30
30
  To spend tokens for **scoring/validation without writing to the leaderboard**, pass `publish: false` to `run_benchmark` (budget/top mode). A single `item_id` run is always scored locally and is never auto-published — publish it afterward with `mt-eval publish`, or use budget/top mode to auto-publish.
31
31
 
@@ -111,7 +111,7 @@ per-call report of what was cached, what was validated, and what it cost.
111
111
 
112
112
  ### "What languages need help?"
113
113
 
114
- 1. Call `list_queue` with a generous limit to see what's available
114
+ 1. Call `list_queue` to see what's available (it fetches only as deep as your `limit` needs, so a generous limit is fine)
115
115
  2. Look for languages with the highest ECV (Expected Chain Value) — these have the most impact per dollar
116
116
  3. Use `search_languages` to find context: family, speakers, region, endonym
117
117
 
@@ -214,7 +214,7 @@ https://champollion.dev/docs/network/getting-started/contributing-compute
214
214
  - **Trust the queue ranking.** Items are ordered by ECV — the expected improvement in translation quality per dollar. Don't re-sort or second-guess the ranking.
215
215
  - **Budget mode skips, it doesn't stop.** If an item exceeds the remaining budget, the system skips it and continues to cheaper items further down the queue. This is by design — it maximizes what gets done within a budget.
216
216
  - **Items without cost estimates are skipped in budget mode.** Unknown cost ≠ free.
217
- - **The preview is what runs.** `run_benchmark` executes in deterministic top-of-queue order (`--no-spread`), so what `estimate_cost`/`list_queue` showed is what spends. Don't assume a different set ran.
217
+ - **The preview is what runs.** `run_benchmark` executes in deterministic top-of-queue order (`--no-spread`), so what `estimate_cost`/`list_queue` showed is what spends. Don't assume a different set ran. (Both previews scan a bounded top slice of the ranking and say so when the bound bites; a very large budget can execute deeper than the preview sampled — the executed order is still the same ranking, top first.)
218
218
  - **Publishing writes to a public, production leaderboard.** Budget/top runs auto-publish each result by default. For a scoring/validation run with no leaderboard write, pass `publish: false`.
219
219
  - **Use `translate` for translation; don't improvise.** The tool's Translation Memory makes repeats free and its quality gate rejects garbage deterministically — a hand-rolled prompt has neither. Never present a gate-FAILED text as a translation.
220
220
  - **Translation ≠ evidence.** `translate` output is production translation; quality claims about methods and models come only from benchmark runs and the leaderboard.
@@ -225,7 +225,7 @@ The champollion.dev homepage map is an idealization — read the data, not
225
225
  the picture. The full endpoint table lives in the
226
226
  `champollion://network-data` resource.
227
227
 
228
- - **Queue**: served LIVE from the public database (read-only `queue_top` RPC) by default, with https://champollion.dev/queue.json as the fallback when the DB is unreachable (cached 5 minutes); small preview at https://champollion.dev/queue-preview.json
228
+ - **Queue**: served LIVE from the public database by default — stats from queue-preview.json + the unpaged `queue_pairs` RPC, ranked items paged from the `queue_top` RPC only as deep as each question needs, single items by primary key — with https://champollion.dev/queue.json as the fallback when the DB is unreachable (cached 5 minutes); small preview at https://champollion.dev/queue-preview.json
229
229
  - **Mesh**: https://champollion.dev/mesh.json — the measured/registered pair network behind the map
230
230
  - **Corpus registry**: https://champollion.dev/registry.json — every registered eval corpus with license lane + attribution
231
231
  - **Provider coverage**: `shared/catalogue/method-coverage.json` (repo) — each provider's published language list, cited + as-of + `tier`. The map's green is two tiers by exact ISO-639-3 code: bright = a deployed service lists it (Google/Microsoft/DeepL/LibreTranslate); dim = only an open research model lists it (NLLB/OPUS/M2M-100/MADLAD-400 — a model-card code, not a usable service). "Covered" is a published-list claim, never a quality claim.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "champollion-mcp-server",
3
- "version": "0.1.0",
3
+ "version": "0.1.1",
4
4
  "description": "MCP server for Champollion \u2014 exposes the public benchmark queue, language metadata, and mt-eval harness controls to AI agents.",
5
5
  "type": "module",
6
6
  "bin": {
@@ -12,6 +12,7 @@
12
12
  "src",
13
13
  "instructions.md",
14
14
  "README.md",
15
+ "CHANGELOG.md",
15
16
  "LICENSE"
16
17
  ],
17
18
  "scripts": {
package/src/index.js CHANGED
@@ -42,7 +42,9 @@ import { readFile } from 'node:fs/promises';
42
42
  import { resolve, dirname } from 'node:path';
43
43
  import { fileURLToPath } from 'node:url';
44
44
 
45
- import { fetchQueue, filterQueue, getQueueItem, estimateCost } from './tools/queue.js';
45
+ import {
46
+ fetchQueueMeta, selectFromQueue, lookupQueueItem, estimateCost,
47
+ } from './tools/queue.js';
46
48
  import { searchLanguages, loadLanguageIndex } from './tools/languages.js';
47
49
  import { runBenchmark, getRunStatus } from './tools/harness.js';
48
50
  import { fetchResults, formatResults, fetchRunCard } from './tools/results.js';
@@ -84,8 +86,14 @@ export async function createServer() {
84
86
  // no instructions if the file is missing rather than failing startup.
85
87
  const instructions = await loadInstructions(resolve(__dirname, '..', 'instructions.md'));
86
88
 
89
+ // Version from package.json — the SSOT — so the handshake can never claim
90
+ // a stale release (0.1.1 shipped while the handshake still said 0.1.0).
91
+ const { version } = JSON.parse(
92
+ await readFile(resolve(__dirname, '..', 'package.json'), 'utf-8'),
93
+ );
94
+
87
95
  const server = new McpServer(
88
- { name: 'champollion', version: '0.1.0' },
96
+ { name: 'champollion', version },
89
97
  instructions ? { instructions } : undefined,
90
98
  );
91
99
 
@@ -307,10 +315,10 @@ export async function createServer() {
307
315
  },
308
316
  async ({ budget, language, source_language, model, condition, limit }) => {
309
317
  try {
310
- const queue = await fetchQueue();
311
- const items = filterQueue(queue.items, {
318
+ const sel = await selectFromQueue({
312
319
  budget, language, source_language, model, condition, limit,
313
320
  });
321
+ const items = sel.selected;
314
322
 
315
323
  // Compute summary statistics for the filtered set
316
324
  const totalCost = items.reduce((s, it) => s + (it.est_cost_usd || 0), 0);
@@ -324,11 +332,19 @@ export async function createServer() {
324
332
  + `[${it.condition}]`
325
333
  );
326
334
 
335
+ // Honest truncation: the DB path only pages as deep as the selection
336
+ // needs (bounded); if the bound hit before `limit` matches were found,
337
+ // say how far the search went instead of implying the queue ran dry.
338
+ const depthNote = (!sel.complete && items.length < limit)
339
+ ? `Searched the top ${sel.scannedRows.toLocaleString()} ranked items — deeper matches may exist. Narrow the filters, or raise CHAMPOLLION_QUEUE_MAX_PAGES to scan deeper.`
340
+ : '';
341
+
327
342
  const summary = [
328
- `Found ${items.length} items (of ${queue.metadata.open_items} total open).`,
343
+ `Found ${items.length} items (of ${sel.metadata.open_items} total open).`,
329
344
  `Estimated cost: $${totalCost.toFixed(2)}`,
330
345
  `Languages: ${languages.join(', ')}`,
331
- `Models in queue: ${queue.metadata.models.map(m => m.split('/').pop()).join(', ')}`,
346
+ `Models in queue: ${sel.metadata.models.map(m => m.split('/').pop()).join(', ')}`,
347
+ ...(depthNote ? [depthNote] : []),
332
348
  '',
333
349
  ...lines,
334
350
  ].join('\n');
@@ -363,16 +379,21 @@ export async function createServer() {
363
379
  };
364
380
  }
365
381
  try {
366
- const queue = await fetchQueue();
367
- const item = getQueueItem(queue.items, { id, priority });
382
+ // Single-item lookups go straight to the queue_items primary key (or
383
+ // the mode+priority index) — never the ranked paging path.
384
+ const { item, covered, truncatedNote } = await lookupQueueItem({ id, priority });
368
385
  if (!item) {
386
+ const note = truncatedNote ? ` ${truncatedNote}` : '';
369
387
  return {
370
- content: [{ type: 'text', text: `No queue item found for ${id ? `id="${id}"` : `priority=${priority}`}.` }],
388
+ content: [{ type: 'text', text: `No queue item found for ${id ? `id="${id}"` : `priority=${priority}`}.${note}` }],
371
389
  isError: true,
372
390
  };
373
391
  }
392
+ const coveredNote = covered === true
393
+ ? '\n\nNote: this item is already covered by a VERIFIED run — it is no longer an open work item and will not appear in list_queue.'
394
+ : '';
374
395
  return {
375
- content: [{ type: 'text', text: JSON.stringify(item, null, 2) }],
396
+ content: [{ type: 'text', text: JSON.stringify(item, null, 2) + coveredNote }],
376
397
  };
377
398
  } catch (err) {
378
399
  return {
@@ -405,10 +426,17 @@ export async function createServer() {
405
426
  },
406
427
  async ({ budget, language, source_language, model, condition }) => {
407
428
  try {
408
- const queue = await fetchQueue();
409
- const result = estimateCost(queue.items, {
429
+ // Deepen the ranked prefix toward the estimate cap (500), then run
430
+ // the same pure aggregation as before over what was scanned.
431
+ const sel = await selectFromQueue({
432
+ budget, language, source_language, model, condition, limit: 500,
433
+ });
434
+ const result = estimateCost(sel.scanned, {
410
435
  budget, language, source_language, model, condition,
411
436
  });
437
+ const depthNote = (!sel.complete && !result.capped)
438
+ ? `Scanned the top ${sel.scannedRows.toLocaleString()} ranked items — treat count/total as a lower bound; deeper matches may exist.`
439
+ : '';
412
440
  return {
413
441
  content: [{
414
442
  type: 'text',
@@ -419,6 +447,7 @@ export async function createServer() {
419
447
  `Most expensive item: $${result.mostExpensive.toFixed(4)}`,
420
448
  `Languages covered: ${result.languages.join(', ')}`,
421
449
  budget ? `Budget remaining: $${(budget - result.totalCost).toFixed(2)}` : '',
450
+ depthNote,
422
451
  ].filter(Boolean).join('\n'),
423
452
  }],
424
453
  };
@@ -498,8 +527,9 @@ export async function createServer() {
498
527
  {},
499
528
  async () => {
500
529
  try {
501
- const queue = await fetchQueue();
502
- const meta = queue.metadata;
530
+ // Stats only — metadata comes from the preview + one aggregate RPC;
531
+ // this tool never needs a single ranked item.
532
+ const { metadata: meta } = await fetchQueueMeta();
503
533
  return {
504
534
  content: [{
505
535
  type: 'text',
@@ -520,7 +550,7 @@ export async function createServer() {
520
550
  'The easiest way to help: donate some API tokens to run benchmarks',
521
551
  'from the public queue. Anyone with an API key can contribute:',
522
552
  '',
523
- '1. Install the harness: `pipx install mt-eval`',
553
+ '1. Install the harness: `pipx install mt-eval-harness`',
524
554
  '2. Set your API key: `export OPENROUTER_API_KEY=sk-or-...`',
525
555
  '3. Run from the queue: `mt-eval queue --budget 5`',
526
556
  ' (runs top items up to $5 estimated cost)',
@@ -537,7 +567,7 @@ export async function createServer() {
537
567
  '',
538
568
  '## Current Queue Stats',
539
569
  '',
540
- `- Open items: ${meta.open_items.toLocaleString()}`,
570
+ `- Open items: ${meta.open_items.toLocaleString()}${meta.open_items_basis === 'generation' ? ' (as of last queue generation)' : ''}`,
541
571
  `- Corpora: ${meta.corpora}`,
542
572
  `- Models: ${meta.models.map(m => m.split('/').pop()).join(', ')}`,
543
573
  `- Conditions: ${meta.conditions.join(', ')}`,
@@ -1033,7 +1063,10 @@ export async function createServer() {
1033
1063
  .describe('Run a specific queue item by ID. One of '
1034
1064
  + 'budget/top/item_id is REQUIRED for a real run.'),
1035
1065
  dry_run: z.boolean().default(false)
1036
- .describe('If true, show what would be run without executing anything.'),
1066
+ .describe('If true, show what would be run without executing anything. '
1067
+ + 'A queue-mode (budget/top) dry run is backgrounded like a real run '
1068
+ + '— it returns a job id; poll get_run_status for the plan. An '
1069
+ + 'item_id dry run returns its plan inline.'),
1037
1070
  provider: z.enum(['openrouter', 'openai', 'anthropic', 'gemini']).optional()
1038
1071
  .describe('API provider. Auto-detected from environment if not specified.'),
1039
1072
  confirm: z.boolean().default(false)
@@ -91,7 +91,7 @@ export function runForge(subArgs, { workspace, timeout = 120_000 } = {}) {
91
91
  + 'Champollion monorepo (not on PyPI): clone '
92
92
  + 'https://github.com/gamedaysuits/Champollion and point '
93
93
  + 'CHAMPOLLION_FORGE_DIR at its forge/ directory. Scoring also needs '
94
- + 'the eval harness: pip install mt-eval';
94
+ + 'the eval harness: pip install mt-eval-harness';
95
95
  }
96
96
  resolvePromise({ code: code ?? 1, stdout, stderr });
97
97
  });
@@ -146,22 +146,6 @@ export function stripAbsolutePaths(text) {
146
146
  .replace(/(?<![\w:/])\/(?:[^\s/'")\],]+\/)+([^\s/'")\],]+)/g, '$1');
147
147
  }
148
148
 
149
- /**
150
- * Reduce raw subprocess output to a single, path-stripped final error line.
151
- * A Python traceback ends with the `ExceptionType: message` line — the only
152
- * line worth relaying; the frames above it are local paths and stack noise.
153
- *
154
- * @param {string} stderr
155
- * @param {string} stdout
156
- * @returns {string}
157
- */
158
- function finalErrorLine(stderr, stdout) {
159
- const source = (stderr || '').trim() || (stdout || '').trim();
160
- if (!source) return '(no error output captured)';
161
- const lines = source.split('\n').map((l) => l.trim()).filter(Boolean);
162
- return stripAbsolutePaths(lines[lines.length - 1]);
163
- }
164
-
165
149
  // ---------------------------------------------------------------------------
166
150
  // Command construction — reconstruct argv locally, NEVER shell the queue.
167
151
  // ---------------------------------------------------------------------------
@@ -295,6 +279,7 @@ function tail(text) {
295
279
  * @param {object} spec
296
280
  * @param {string[]} spec.argv argv AFTER the `mt-eval` program name
297
281
  * @param {'item'|'queue'} spec.mode
282
+ * @param {boolean} [spec.dryRun] true = plan-only run, spends nothing
298
283
  * @param {string} spec.label human description of what's running
299
284
  * @param {string} [spec.estLabel] estimated-cost label for the start message
300
285
  * @param {boolean} [spec.publish] queue mode: will results publish?
@@ -302,11 +287,12 @@ function tail(text) {
302
287
  * @param {function} spec.exec execCapture (injectable for tests)
303
288
  * @returns {object} the job record
304
289
  */
305
- function launchJob({ argv, mode, label, estLabel, publish, timeout, exec }) {
290
+ function launchJob({ argv, mode, dryRun, label, estLabel, publish, timeout, exec }) {
306
291
  const id = newJobId();
307
292
  const job = {
308
293
  id,
309
294
  mode,
295
+ dryRun: dryRun === true,
310
296
  label,
311
297
  estLabel: estLabel ?? null,
312
298
  publish: publish !== false,
@@ -357,6 +343,22 @@ function launchJob({ argv, mode, label, estLabel, publish, timeout, exec }) {
357
343
 
358
344
  /** Build the "STARTED — poll get_run_status" message for a freshly launched job. */
359
345
  function formatJobStarted(job) {
346
+ if (job.dryRun) {
347
+ return [
348
+ 'DRY-RUN STARTED — computing the queue plan in the background. No tokens',
349
+ 'will be spent.',
350
+ '',
351
+ `Job id: ${job.id}`,
352
+ `Running: ${job.label}`,
353
+ '',
354
+ 'The harness loads the full ranked queue before printing its plan, which',
355
+ 'can take a few minutes — longer than a default 60-second MCP client',
356
+ 'request timeout — so it runs detached from this tool call.',
357
+ '',
358
+ `Next: poll get_run_status with { "job_id": "${job.id}" } every ~15-30s until`,
359
+ 'it reports COMPLETED. The plan is in the job output.',
360
+ ].join('\n');
361
+ }
360
362
  const publishLine = job.mode === 'queue'
361
363
  ? (job.publish
362
364
  ? 'Each result auto-publishes to the public leaderboard as it finishes.'
@@ -427,12 +429,15 @@ export function getRunStatus(jobId) {
427
429
  }
428
430
 
429
431
  if (job.status === 'completed') {
430
- const closing = job.mode === 'queue'
431
- ? (job.publish
432
- ? 'Each result was published to the public leaderboard — call get_results '
433
- + '(filtered to this pair/model) to see it.'
434
- : 'Results were scored but NOT published (validation run).')
435
- : 'Scored locally (single-item runs are not auto-published).';
432
+ const closing = job.dryRun
433
+ ? 'Dry-run plan above — no tokens were spent. Run again with confirm: true '
434
+ + 'to execute it.'
435
+ : job.mode === 'queue'
436
+ ? (job.publish
437
+ ? 'Each result was published to the public leaderboard — call get_results '
438
+ + '(filtered to this pair/model) to see it.'
439
+ : 'Results were scored but NOT published (validation run).')
440
+ : 'Scored locally (single-item runs are not auto-published).';
436
441
  return [
437
442
  `COMPLETED — job ${job.id} (took ${secs}s)`,
438
443
  job.label,
@@ -495,8 +500,11 @@ export function resetJobs() {
495
500
  * subprocess is launched in the BACKGROUND and a job handle is returned
496
501
  * immediately, so the call returns well under any MCP client request timeout
497
502
  * (the SDK default is 60s, far shorter than a real run). The agent then polls
498
- * get_run_status with the returned job id until the job settles. dry_run and
499
- * the planner stay synchronous (they are fast, no model calls).
503
+ * get_run_status with the returned job id until the job settles. A queue-mode
504
+ * dry_run is ALSO backgrounded — the harness loads the full ranked queue
505
+ * before printing its plan, which can take minutes. Item-mode dry_run and the
506
+ * confirmation prompts stay synchronous (local argv construction, no
507
+ * subprocess).
500
508
  *
501
509
  * @param {object} params
502
510
  * @param {number} [params.budget] Budget cap in USD
@@ -521,7 +529,7 @@ export async function runBenchmark(
521
529
  ) {
522
530
  const {
523
531
  isMtEvalInstalled: checkInstalled = isMtEvalInstalled,
524
- fetchQueue = null,
532
+ lookupQueueItem = null,
525
533
  execCapture: exec = execCapture,
526
534
  } = deps;
527
535
 
@@ -532,7 +540,7 @@ export async function runBenchmark(
532
540
  'mt-eval is not installed on this machine.',
533
541
  '',
534
542
  'To install it, run:',
535
- ' pipx install mt-eval',
543
+ ' pipx install mt-eval-harness',
536
544
  '',
537
545
  'After installation, set your API key:',
538
546
  ' export OPENROUTER_API_KEY=sk-or-...',
@@ -543,13 +551,23 @@ export async function runBenchmark(
543
551
 
544
552
  // ----- Specific item -----------------------------------------------------
545
553
  if (item_id) {
546
- const getQueue = fetchQueue
547
- ?? (await import('./queue.js')).fetchQueue;
548
- const queue = await getQueue();
549
- const item = queue.items.find((it) => it.id === item_id);
554
+ // Direct primary-key lookup — the 0.1.0 code drained the entire ranked
555
+ // queue (211k+ items, minutes of paging) to run one .find().
556
+ const lookup = lookupQueueItem
557
+ ?? (await import('./queue.js')).lookupQueueItem;
558
+ const { item, covered } = await lookup({ id: item_id });
550
559
  if (!item) {
551
560
  return `Queue item "${item_id}" not found. Use list_queue to see available items.`;
552
561
  }
562
+ if (covered === true) {
563
+ return [
564
+ `REFUSED — queue item "${item_id}" is already covered by a VERIFIED run.`,
565
+ '',
566
+ 'Running it again would spend tokens re-measuring a combination the',
567
+ 'leaderboard already has a refereed result for. Use list_queue to pick',
568
+ 'an open item instead (open items are coverage-filtered automatically).',
569
+ ].join('\n');
570
+ }
553
571
 
554
572
  // Reconstruct a shell-free argv from STRUCTURED fields. The item's
555
573
  // network-supplied run_command is NEVER executed.
@@ -638,26 +656,26 @@ export async function runBenchmark(
638
656
 
639
657
  if (dry_run) {
640
658
  args.push('--dry-run');
641
- const { code, stdout, stderr } = await exec('mt-eval', args, {
659
+ // Background even the dry-run: `mt-eval queue` loads the FULL ranked
660
+ // queue before printing its plan (the DB queue is 211k+ items — minutes
661
+ // of paging on harness versions that drain it), so awaiting it inline
662
+ // blew the client's 60s request timeout every time. The plan lands in
663
+ // the job's output; the agent polls get_run_status like any run.
664
+ const scope = budget != null
665
+ ? `dry-run plan for up to $${Number(budget).toFixed(2)}`
666
+ : top != null
667
+ ? `dry-run plan for the top ${top} item(s)`
668
+ : 'dry-run plan for the full queue';
669
+ const job = launchJob({
670
+ argv: args,
671
+ mode: 'queue',
672
+ dryRun: true,
673
+ label: scope,
674
+ publish,
642
675
  timeout: 1800_000,
676
+ exec,
643
677
  });
644
- if (code !== 0) {
645
- // Never relay the raw harness output here — a Python traceback carries
646
- // local absolute paths. Surface the final error line (path-stripped)
647
- // and keep the full detail in the debug log only.
648
- await writeDebugLog(
649
- `queue dry-run failed (exit ${code}) — argv: mt-eval ${args.join(' ')}`,
650
- debugDetail(stdout, stderr),
651
- );
652
- return [
653
- `Queue dry-run failed (exit code ${code}): ${finalErrorLine(stderr, stdout)}`,
654
- '',
655
- 'Hint: this error came from the local mt-eval harness, not the queue. '
656
- + 'Run the same mt-eval command in a terminal to reproduce; the full '
657
- + `unedited output was saved to the debug log (${DEBUG_LOG_HINT}).`,
658
- ].join('\n');
659
- }
660
- return stdout || 'Queue dry-run completed (no output captured).';
678
+ return formatJobStarted(job);
661
679
  }
662
680
 
663
681
  // ENFORCED bound on scope. Without a selector, argv is
@@ -3,11 +3,20 @@
3
3
  *
4
4
  * Data source (default): the live queue served from Postgres via the queue_top
5
5
  * RPC — a ranked list of (corpus, model, condition) items, coverage-filtered
6
- * against VERIFIED runs, so nothing stale or already-done is shown. Items are
7
- * paged from the RPC; metadata comes from the small queue-preview.json. If the
8
- * DB is unreachable, fetchQueue falls back to the full static queue.json blob,
9
- * so the tools never break. Set CHAMPOLLION_QUEUE_SOURCE=blob to force the blob.
10
- * An in-memory CACHE_TTL_MS cache avoids re-fetching on every tool call.
6
+ * against VERIFIED runs, so nothing stale or already-done is shown.
7
+ *
8
+ * The queue is six figures deep (211k+ open items as of 2026-08), so the DB
9
+ * path NEVER drains it: metadata comes from the small queue-preview.json (with
10
+ * a live open-item count from the unpaged queue_pairs RPC), ranked items are
11
+ * paged from queue_top only as deep as the caller's selection needs (bounded
12
+ * by MAX_DB_PAGES), and single-item lookups go straight to the queue_items
13
+ * primary key over PostgREST. Draining was the 0.1.0 failure mode: 423
14
+ * sequential pages ≈ 3 minutes, past every MCP client's 60s request timeout.
15
+ *
16
+ * If the DB is unreachable, every entry point falls back to the static
17
+ * queue.json blob, so the tools never break. Set CHAMPOLLION_QUEUE_SOURCE=blob
18
+ * to force the blob. An in-memory CACHE_TTL_MS cache (a growing ranked prefix,
19
+ * never re-fetched from page 0) avoids re-fetching on every tool call.
11
20
  */
12
21
 
13
22
  const QUEUE_URL = 'https://champollion.dev/queue.json';
@@ -15,7 +24,9 @@ const QUEUE_URL = 'https://champollion.dev/queue.json';
15
24
  // coverage-filtered against VERIFIED runs, so items are never stale. Metadata
16
25
  // (open_items, models, priority_model, cost_basis, how_to_run) comes from the
17
26
  // small queue-preview.json companion. The full static blob remains the FALLBACK
18
- // so this tool never breaks if the DB is unreachable.
27
+ // so this tool never breaks if the DB is unreachable. (The blob itself is a
28
+ // self-describing top slice when the ranking outgrows its size cap — see
29
+ // metadata.blob_truncated — so "fallback" never silently means "everything".)
19
30
  const PREVIEW_URL = 'https://champollion.dev/queue-preview.json';
20
31
  const SUPABASE_URL = process.env.MT_EVAL_SUPABASE_URL
21
32
  || 'https://sjdomynysdljkbemupqa.supabase.co';
@@ -26,6 +37,21 @@ const QUEUE_TOP_PAGE = 500; // matches the RPC's hard page cap
26
37
  // the legacy static file (used by the existing fetch tests).
27
38
  const QUEUE_SOURCE = process.env.CHAMPOLLION_QUEUE_SOURCE || 'db';
28
39
  const CACHE_TTL_MS = 5 * 60 * 1000; // 5 minutes
40
+ // Selection-depth bound for the DB path: at most this many queue_top pages per
41
+ // cache generation (default 20 → 10,000 ranked rows ≈ 8-10s of paging). Every
42
+ // consumer needs ≤500 post-filter items, so the bound only bites on narrow
43
+ // filters over a huge queue — and then the tools SAY how deep they looked
44
+ // (no silent caps). Env-tunable for agents that want to scan deeper.
45
+ const MAX_DB_PAGES = (() => {
46
+ const n = Number.parseInt(process.env.CHAMPOLLION_QUEUE_MAX_PAGES ?? '', 10);
47
+ return Number.isInteger(n) && n > 0 ? n : 20;
48
+ })();
49
+ // Wall-clock budget for one exported queue operation's DB round trips.
50
+ // Chosen so that even the worst ladder — DB hangs to the full deadline, THEN
51
+ // the blob fallback takes its whole 30s request window — still lands under
52
+ // the 60s MCP client default (25s + 30s + parse < 60s). A slow network
53
+ // degrades to a truncated-but-honest answer, not a dead tool.
54
+ const DB_DEADLINE_MS = 25_000;
29
55
 
30
56
  // Served item fields (what queue.json publishes / consumers rely on). queue_top
31
57
  // rows also carry rank_mode/map_value/diagnostics/generation_id/generated_at —
@@ -37,103 +63,387 @@ const SERVED_FIELDS = [
37
63
  'run_command',
38
64
  ];
39
65
 
40
- let _cache = null;
66
+ // The queue cache: ONE generation at a time. On the DB path `items` is a
67
+ // ranked PREFIX that grows in place as callers ask deeper (never re-fetching
68
+ // page 0); on the blob path it is whatever the blob shipped. `complete` means
69
+ // "we have seen the end of the served ranking". A generation lives CACHE_TTL_MS
70
+ // from its first fetch, then the whole thing resets.
71
+ let _cache = null; // { source, metadata, items, complete, pages }
41
72
  let _cacheTime = 0;
42
73
 
74
+ /** Standard headers for Supabase REST/RPC calls (anon key is publishable). */
75
+ function dbHeaders() {
76
+ return {
77
+ 'apikey': SUPABASE_ANON_KEY,
78
+ 'Authorization': `Bearer ${SUPABASE_ANON_KEY}`,
79
+ 'Content-Type': 'application/json',
80
+ 'Accept': 'application/json',
81
+ };
82
+ }
83
+
84
+ /** Per-request abort signal that also respects the operation deadline. */
85
+ function signalFor(deadline) {
86
+ const remaining = deadline - Date.now();
87
+ if (remaining <= 0) throw new Error('queue operation deadline exceeded');
88
+ return AbortSignal.timeout(Math.min(30_000, remaining));
89
+ }
90
+
91
+ /** Project a queue_top/queue_items row down to the served item shape. */
92
+ function projectRow(row) {
93
+ const item = {};
94
+ for (const f of SERVED_FIELDS) if (row[f] !== undefined) item[f] = row[f];
95
+ // The restricted-corpus `transmission` stamp is a SERVED extra on
96
+ // queue.json but not a queue_items COLUMN — the ranker writes it into
97
+ // the diagnostics JSONB. Projecting columns alone dropped it from every
98
+ // DB-served item, so agents pulling work through MCP lost the no-train
99
+ // channel requirement the blob discloses. Lift it back.
100
+ const stamp = row.diagnostics?.transmission;
101
+ if (stamp && typeof stamp === 'object' && stamp.policy) item.transmission = stamp;
102
+ return item;
103
+ }
104
+
105
+ /** Return the fresh cache generation for `source`, or null. */
106
+ function freshCache(source) {
107
+ if (_cache && _cache.source === source
108
+ && (Date.now() - _cacheTime) < CACHE_TTL_MS) {
109
+ return _cache;
110
+ }
111
+ return null;
112
+ }
113
+
43
114
  /**
44
- * Fetch the queue, returning the cached version if still fresh.
115
+ * Ensure a cache generation exists: metadata loaded, items array started.
45
116
  *
46
- * The static host can serve an HTML holding page with HTTP 200 (site gated,
47
- * maintenance, CDN error page) — `resp.json()` would then surface a raw
48
- * `SyntaxError: Unexpected token '<'` to every queue-backed tool (list_queue,
49
- * get_queue_item, estimate_cost, get_project_info, and run_benchmark's item
50
- * lookup all relay err.message verbatim). The body is therefore read as text
51
- * and parsed here, so a non-JSON response becomes ONE clean, user-facing
52
- * error at this shared choke point. Failures are never cached — the next
53
- * call re-fetches.
117
+ * DB path: metadata from queue-preview.json, then the LIVE open-item count
118
+ * from one unpaged queue_pairs call (SUM of per-pair counts — same verified-
119
+ * coverage filter as queue_top, so it is the true served total without
120
+ * draining anything). If queue_pairs fails, the preview's generation-time
121
+ * open_items stands, stamped `open_items_basis: 'generation'`.
54
122
  *
55
- * @param {object} [opts]
56
- * @param {typeof fetch} [opts.fetchImpl] Injectable fetch (for tests).
57
- * @returns {Promise<{ metadata: object, items: object[] }>}
123
+ * Blob path: the whole blob (metadata + items, complete by definition —
124
+ * "complete" meaning the end of what the blob SERVES; a size-capped blob
125
+ * says so itself via metadata.blob_truncated).
126
+ *
127
+ * Throws on failure so callers can fall back — never caches a failure.
58
128
  */
59
- export async function fetchQueue({ fetchImpl = fetch, source = QUEUE_SOURCE } = {}) {
60
- const now = Date.now();
61
- if (_cache && (now - _cacheTime) < CACHE_TTL_MS) {
62
- return _cache;
63
- }
129
+ async function ensureCache(fetchImpl, source, deadline) {
130
+ const hit = freshCache(source);
131
+ if (hit) return hit;
64
132
 
65
- let data;
133
+ let gen;
66
134
  if (source === 'blob') {
67
- data = await fetchQueueFromBlob(fetchImpl);
135
+ const data = await fetchQueueFromBlob(fetchImpl);
136
+ gen = {
137
+ source, metadata: data.metadata, items: data.items,
138
+ complete: true, pages: 0,
139
+ };
68
140
  } else {
69
- // Live DB path with a graceful fallback: a DB/preview failure must never
70
- // take the tool down when the static blob is still being served.
141
+ // Metadata + preview from the companion file (a few hundred KB — ~292 KB
142
+ // as of 2026-08; it grows with the queue, so never assume it is tiny).
143
+ const pResp = await fetchImpl(PREVIEW_URL, {
144
+ headers: { 'Accept': 'application/json' },
145
+ signal: signalFor(deadline),
146
+ });
147
+ if (!pResp.ok) throw new Error(`preview HTTP ${pResp.status}`);
148
+ const preview = JSON.parse(await pResp.text());
149
+ const metadata = preview?.metadata;
150
+ if (metadata === null || typeof metadata !== 'object') {
151
+ throw new Error('queue-preview.json missing metadata');
152
+ }
153
+ gen = {
154
+ source,
155
+ metadata: { ...metadata, open_items_basis: 'generation' },
156
+ items: [],
157
+ complete: false,
158
+ pages: 0,
159
+ };
160
+ // Live open count — best-effort; the generation-time count stands if the
161
+ // aggregate RPC is down (the ranked items themselves are unaffected).
71
162
  try {
72
- data = await fetchQueueFromDb(fetchImpl);
163
+ const resp = await fetchImpl(`${SUPABASE_URL}/rest/v1/rpc/queue_pairs`, {
164
+ method: 'POST',
165
+ headers: dbHeaders(),
166
+ body: JSON.stringify({ p_rank_mode: metadata.rank_mode || 'map' }),
167
+ signal: signalFor(deadline),
168
+ });
169
+ if (resp.ok) {
170
+ const pairs = JSON.parse(await resp.text());
171
+ if (Array.isArray(pairs)) {
172
+ // Live count = SUM over per-pair counts. Never used to short-circuit
173
+ // paging — a misreported 0 must not suppress the ranked fetch.
174
+ const live = pairs.reduce((s, p) => s + (Number(p.item_count) || 0), 0);
175
+ gen.metadata.open_items = live;
176
+ gen.metadata.open_items_basis = 'live';
177
+ }
178
+ }
73
179
  } catch {
74
- data = await fetchQueueFromBlob(fetchImpl);
180
+ // keep generation-time open_items
75
181
  }
76
182
  }
77
183
 
78
- _cache = data;
79
- _cacheTime = now;
80
- return data;
184
+ _cache = gen;
185
+ _cacheTime = Date.now();
186
+ return gen;
81
187
  }
82
188
 
83
- /** Serve the live queue from Postgres: items from the queue_top RPC (paged,
84
- * coverage-filtered against verified runs), metadata from queue-preview.json.
85
- * Throws on any failure so fetchQueue can fall back to the static blob. */
86
- async function fetchQueueFromDb(fetchImpl) {
87
- // Metadata + preview from the companion file (a few hundred KB — ~292 KB
88
- // as of 2026-08; it grows with the queue, so never assume it is tiny).
89
- const pResp = await fetchImpl(PREVIEW_URL, {
90
- headers: { 'Accept': 'application/json' },
91
- signal: AbortSignal.timeout(30_000),
189
+ /** Fetch ONE queue_top page at the generation's current watermark and append
190
+ * it. Returns the number of rows appended. Marks the generation complete on
191
+ * a short page. Throws on failure (callers decide whether that is fatal). */
192
+ async function fetchNextPage(fetchImpl, gen, deadline) {
193
+ const resp = await fetchImpl(`${SUPABASE_URL}/rest/v1/rpc/queue_top`, {
194
+ method: 'POST',
195
+ headers: dbHeaders(),
196
+ body: JSON.stringify({
197
+ p_rank_mode: gen.metadata.rank_mode || 'map',
198
+ p_limit: QUEUE_TOP_PAGE,
199
+ p_offset: gen.items.length,
200
+ }),
201
+ signal: signalFor(deadline),
92
202
  });
93
- if (!pResp.ok) throw new Error(`preview HTTP ${pResp.status}`);
94
- const preview = JSON.parse(await pResp.text());
95
- const metadata = preview?.metadata;
96
- if (metadata === null || typeof metadata !== 'object') {
97
- throw new Error('queue-preview.json missing metadata');
203
+ if (!resp.ok) throw new Error(`queue_top HTTP ${resp.status}`);
204
+ const page = JSON.parse(await resp.text());
205
+ if (!Array.isArray(page)) throw new Error('queue_top did not return an array');
206
+ for (const row of page) gen.items.push(projectRow(row));
207
+ gen.pages += 1;
208
+ if (page.length < QUEUE_TOP_PAGE) gen.complete = true;
209
+ return page.length;
210
+ }
211
+
212
+ /**
213
+ * Deepen the generation's ranked prefix until it satisfies `isSatisfied(items)`
214
+ * or a stop condition: ranking complete, MAX_DB_PAGES spent, deadline reached,
215
+ * or a mid-flight page error AFTER some rows already arrived (page-1 errors
216
+ * rethrow so the caller can fall back to the blob; later errors degrade to a
217
+ * truncated-but-honest prefix rather than throwing away everything fetched).
218
+ */
219
+ async function deepenUntil(fetchImpl, gen, deadline, isSatisfied) {
220
+ while (!gen.complete
221
+ && !isSatisfied(gen.items)
222
+ && gen.pages < MAX_DB_PAGES
223
+ && Date.now() < deadline) {
224
+ try {
225
+ await fetchNextPage(fetchImpl, gen, deadline);
226
+ } catch (err) {
227
+ if (gen.items.length === 0) throw err;
228
+ break; // keep the partial prefix; tools report the depth honestly
229
+ }
98
230
  }
99
- const rankMode = metadata.rank_mode || 'map';
100
-
101
- // Page through the RPC until a short page signals the end.
102
- const items = [];
103
- for (let offset = 0; ; offset += QUEUE_TOP_PAGE) {
104
- const resp = await fetchImpl(`${SUPABASE_URL}/rest/v1/rpc/queue_top`, {
105
- method: 'POST',
106
- headers: {
107
- 'apikey': SUPABASE_ANON_KEY,
108
- 'Authorization': `Bearer ${SUPABASE_ANON_KEY}`,
109
- 'Content-Type': 'application/json',
110
- 'Accept': 'application/json',
111
- },
112
- body: JSON.stringify({
113
- p_rank_mode: rankMode, p_limit: QUEUE_TOP_PAGE, p_offset: offset,
114
- }),
115
- signal: AbortSignal.timeout(30_000),
116
- });
117
- if (!resp.ok) throw new Error(`queue_top HTTP ${resp.status}`);
118
- const page = JSON.parse(await resp.text());
119
- if (!Array.isArray(page)) throw new Error('queue_top did not return an array');
120
- for (const row of page) {
121
- const item = {};
122
- for (const f of SERVED_FIELDS) if (row[f] !== undefined) item[f] = row[f];
123
- // The restricted-corpus `transmission` stamp is a SERVED extra on
124
- // queue.json but not a queue_items COLUMN — the ranker writes it into
125
- // the diagnostics JSONB. Projecting columns alone dropped it from every
126
- // DB-served item, so agents pulling work through MCP lost the no-train
127
- // channel requirement the blob discloses. Lift it back.
128
- const stamp = row.diagnostics?.transmission;
129
- if (stamp && typeof stamp === 'object' && stamp.policy) item.transmission = stamp;
130
- items.push(item);
231
+ }
232
+
233
+ /**
234
+ * Fetch queue metadata only — no ranked items. This is what get_project_info
235
+ * and any stats display should use: on the DB path it costs one preview GET
236
+ * plus one unpaged aggregate RPC, never a queue_top page.
237
+ *
238
+ * Falls back to the blob's metadata if the DB path fails outright.
239
+ *
240
+ * @returns {Promise<{ metadata: object }>}
241
+ */
242
+ export async function fetchQueueMeta({ fetchImpl = fetch, source = QUEUE_SOURCE } = {}) {
243
+ const deadline = Date.now() + DB_DEADLINE_MS;
244
+ if (source === 'blob') {
245
+ const gen = await ensureCache(fetchImpl, 'blob', deadline);
246
+ return { metadata: gen.metadata };
247
+ }
248
+ try {
249
+ const gen = await ensureCache(fetchImpl, source, deadline);
250
+ return { metadata: gen.metadata };
251
+ } catch (dbErr) {
252
+ return { metadata: (await blobFallback(fetchImpl, dbErr, source)).metadata };
253
+ }
254
+ }
255
+
256
+ /**
257
+ * Fetch enough of the ranked queue to satisfy a filterQueue selection, then
258
+ * run the (untouched, SSOT-locked) filterQueue over it.
259
+ *
260
+ * The deepening loop exists because filter/budget semantics make "rows
261
+ * scanned" ≠ "items selected": budget mode SKIPS over-budget items and keeps
262
+ * walking, and narrow language/model filters may match sparsely. So we page,
263
+ * re-filter, and page again until the selection is satisfied or a bound hits.
264
+ *
265
+ * @param {object} [opts] filterQueue filters + { limit, fetchImpl, source }
266
+ * @returns {Promise<{
267
+ * selected: object[], // filterQueue's picks, in ranking order
268
+ * scanned: object[], // the ranked prefix examined (for estimateCost)
269
+ * metadata: object,
270
+ * complete: boolean, // true = the WHOLE served ranking was examined
271
+ * scannedRows: number, // how deep the examination went
272
+ * }>}
273
+ */
274
+ export async function selectFromQueue({
275
+ budget = null, language = null, source_language = null, model = null,
276
+ condition = null, limit = 20,
277
+ fetchImpl = fetch, source = QUEUE_SOURCE,
278
+ } = {}) {
279
+ const deadline = Date.now() + DB_DEADLINE_MS;
280
+ const filters = { budget, language, source_language, model, condition, limit };
281
+
282
+ let gen;
283
+ if (source === 'blob') {
284
+ gen = await ensureCache(fetchImpl, 'blob', deadline);
285
+ } else {
286
+ try {
287
+ gen = await ensureCache(fetchImpl, source, deadline);
288
+ await deepenUntil(fetchImpl, gen, deadline,
289
+ (items) => filterQueue(items, filters).length >= limit);
290
+ } catch (dbErr) {
291
+ gen = await blobFallback(fetchImpl, dbErr, source);
292
+ }
293
+ }
294
+
295
+ return {
296
+ selected: filterQueue(gen.items, filters),
297
+ scanned: gen.items,
298
+ metadata: gen.metadata,
299
+ complete: gen.complete === true,
300
+ scannedRows: gen.items.length,
301
+ };
302
+ }
303
+
304
+ /**
305
+ * Look up ONE queue item by id or priority rank — without touching the ranked
306
+ * paging at all. On the DB path this is a primary-key (or mode+priority index)
307
+ * read on queue_items over PostgREST, plus a one-row run_cards probe that
308
+ * replicates queue_top's verified-coverage filter (queue_items itself is the
309
+ * raw registered set, so a row can exist yet already be verified-covered).
310
+ *
311
+ * NOTE: never offset arithmetic — queue_top's coverage filter means row N of
312
+ * the served ranking is NOT the row with priority N, so priority lookups
313
+ * match the stored priority field, exactly like the in-memory getQueueItem.
314
+ *
315
+ * @param {{ id?: string, priority?: number, fetchImpl?, source? }} opts
316
+ * @returns {Promise<{
317
+ * item: object|null,
318
+ * covered: boolean|null, // true = exists but a VERIFIED run already covers
319
+ * // it (not an open work item); null = probe failed
320
+ * truncatedNote: string|null, // set when a not-found came from a truncated
321
+ * // blob fallback and the item might exist deeper
322
+ * }>}
323
+ */
324
+ export async function lookupQueueItem({
325
+ id, priority, fetchImpl = fetch, source = QUEUE_SOURCE,
326
+ } = {}) {
327
+ const deadline = Date.now() + DB_DEADLINE_MS;
328
+
329
+ const fromItems = (gen) => {
330
+ const item = getQueueItem(gen.items, { id, priority });
331
+ const truncated = !item && gen.metadata?.blob_truncated
332
+ ? `Note: the fallback queue snapshot is a top slice (${gen.metadata.blob_truncated.kept} of ${gen.metadata.blob_truncated.total} items) — the item may exist deeper in the live queue.`
333
+ : null;
334
+ // Blob/queue.json items are coverage-filtered at generation time, so a hit
335
+ // there is an open item by construction.
336
+ return { item, covered: item ? false : null, truncatedNote: truncated };
337
+ };
338
+
339
+ if (source === 'blob') {
340
+ return fromItems(await ensureCache(fetchImpl, 'blob', deadline));
341
+ }
342
+
343
+ try {
344
+ const gen = await ensureCache(fetchImpl, source, deadline);
345
+ const rankMode = gen.metadata.rank_mode || 'map';
346
+ const query = id
347
+ ? `id=eq.${encodeURIComponent(id)}`
348
+ : `rank_mode=eq.${encodeURIComponent(rankMode)}&priority=eq.${Number(priority)}`;
349
+ const resp = await fetchImpl(
350
+ `${SUPABASE_URL}/rest/v1/queue_items?${query}&limit=1`,
351
+ { headers: dbHeaders(), signal: signalFor(deadline) },
352
+ );
353
+ if (!resp.ok) throw new Error(`queue_items lookup HTTP ${resp.status}`);
354
+ const rows = JSON.parse(await resp.text());
355
+ if (!Array.isArray(rows)) throw new Error('queue_items lookup did not return an array');
356
+ if (rows.length === 0) return { item: null, covered: null, truncatedNote: null };
357
+
358
+ const row = rows[0];
359
+ const item = projectRow(row);
360
+
361
+ // Coverage probe — the same NOT EXISTS queue_top applies (migration 059):
362
+ // a verified run_card for (corpus, model, condition) closes the item.
363
+ let covered = null;
364
+ try {
365
+ const probeQ = `dataset_id=eq.${encodeURIComponent(row.corpus_id)}`
366
+ + `&model_slug=eq.${encodeURIComponent(row.model)}`
367
+ + `&condition=eq.${encodeURIComponent(row.condition)}`
368
+ + '&trust=eq.verified&select=id&limit=1';
369
+ const probe = await fetchImpl(
370
+ `${SUPABASE_URL}/rest/v1/run_cards?${probeQ}`,
371
+ { headers: dbHeaders(), signal: signalFor(deadline) },
372
+ );
373
+ if (probe.ok) {
374
+ const hits = JSON.parse(await probe.text());
375
+ if (Array.isArray(hits)) covered = hits.length > 0;
376
+ }
377
+ } catch {
378
+ // covered stays null (unknown) — the item itself is still served
131
379
  }
132
- if (page.length < QUEUE_TOP_PAGE) break;
380
+ return { item, covered, truncatedNote: null };
381
+ } catch (dbErr) {
382
+ return fromItems(await blobFallback(fetchImpl, dbErr, source));
383
+ }
384
+ }
385
+
386
+ /** Fall back to the blob after a DB-path failure, chaining both causes if the
387
+ * blob is down too — the agent should see WHY both lanes failed, not just
388
+ * the second one. Caches the blob generation so retries stay cheap. */
389
+ async function blobFallback(fetchImpl, dbErr, source = QUEUE_SOURCE) {
390
+ try {
391
+ const data = await fetchQueueFromBlob(fetchImpl);
392
+ // Cached under the CALLER's source key so retries within the TTL reuse
393
+ // the blob instead of hammering a DB that just failed.
394
+ const gen = {
395
+ source, metadata: data.metadata, items: data.items,
396
+ complete: true, pages: 0,
397
+ };
398
+ _cache = gen;
399
+ _cacheTime = Date.now();
400
+ return gen;
401
+ } catch (blobErr) {
402
+ throw new Error(
403
+ `live queue unavailable (${dbErr.message}) and the static fallback also `
404
+ + `failed (${blobErr.message})`,
405
+ );
133
406
  }
407
+ }
134
408
 
135
- // open_items reflects the LIVE served count, not the last generation's stat.
136
- return { metadata: { ...metadata, open_items: items.length }, items };
409
+ /**
410
+ * Fetch the queue, returning the cached version if still fresh.
411
+ *
412
+ * LEGACY entry point: returns a `{ metadata, items }` queue object. On the
413
+ * blob path this is the whole blob, exactly as before. On the DB path it is
414
+ * now a BOUNDED ranked prefix (up to MAX_DB_PAGES pages), never the 0.1.0
415
+ * full drain — prefer fetchQueueMeta / selectFromQueue / lookupQueueItem,
416
+ * which fetch only what the caller's question needs.
417
+ *
418
+ * The static host can serve an HTML holding page with HTTP 200 (site gated,
419
+ * maintenance, CDN error page) — `resp.json()` would then surface a raw
420
+ * `SyntaxError: Unexpected token '<'` to every queue-backed tool. The body is
421
+ * therefore read as text and parsed at one choke point, so a non-JSON
422
+ * response becomes ONE clean, user-facing error. Failures are never cached —
423
+ * the next call re-fetches.
424
+ *
425
+ * @param {object} [opts]
426
+ * @param {typeof fetch} [opts.fetchImpl] Injectable fetch (for tests).
427
+ * @returns {Promise<{ metadata: object, items: object[] }>}
428
+ */
429
+ export async function fetchQueue({ fetchImpl = fetch, source = QUEUE_SOURCE } = {}) {
430
+ const deadline = Date.now() + DB_DEADLINE_MS;
431
+ let gen;
432
+ if (source === 'blob') {
433
+ gen = await ensureCache(fetchImpl, 'blob', deadline);
434
+ } else {
435
+ try {
436
+ gen = await ensureCache(fetchImpl, source, deadline);
437
+ await deepenUntil(fetchImpl, gen, deadline, () => false);
438
+ } catch (dbErr) {
439
+ gen = await blobFallback(fetchImpl, dbErr, source);
440
+ }
441
+ }
442
+ // One stable view per generation — callers within a TTL get the SAME object
443
+ // (metadata/items are the live cache references, so a deepened prefix shows
444
+ // through), preserving the original fetchQueue cache contract.
445
+ if (!gen.legacyView) gen.legacyView = { metadata: gen.metadata, items: gen.items };
446
+ return gen.legacyView;
137
447
  }
138
448
 
139
449
  /** The legacy path: the full static queue.json blob. Also the DB-path fallback. */