champollion-mcp-server 0.1.0 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md ADDED
@@ -0,0 +1,468 @@
1
+ # Changelog
2
+
3
+ ## 0.2.0 (2026-10-03)
4
+
5
+ Built around one flow: someone tells their agent "let's build a Cree model
6
+ for our school" (or "an Atya phrasebook for the clinic") and the agent can
7
+ take them from discovery to a measured, deployable model. Synthetic-user
8
+ testing against the 0.1.x npm install found the gaps this release closes.
9
+
10
+ - **Round 13 (synthetic-user findings, 2026-10-04).**
11
+ - New `preview_publish`: the READ-ONLY half of `publish_report` — the
12
+ harness's own `mt-eval publish <report> --dry-run`, WHAT GETS PUBLISHED,
13
+ the exact `publish_ack` and the exact `publish_report` call that would
14
+ publish it. It has no `confirm` or `publish_ack` (refused by name) and
15
+ carries the MCP annotation `readOnlyHint: true`, so an agent host that
16
+ gates writes can allow it alone (a host had blocked publish_report's
17
+ preview as a production deploy). `publish_report` is annotated
18
+ `destructiveHint` / `openWorldHint`; without `confirm` it still returns
19
+ the preview. 34 tools.
20
+ - `run_benchmark`: `metricx` (+ `metricx_model`) and `fuse` pass the
21
+ harness's opt-in `--metricx` / `--metricx-model` / `--fuse`; `comet`
22
+ requires COMET (the harness computes it whenever unbabel-comet is
23
+ installed — there is no run flag). The plan says, from the harness's
24
+ Python, whether each is installed, what to install and what it
25
+ downloads; a confirmed run that asks for a metric the harness cannot
26
+ compute is REFUSED, never run without it. Item and corpus runs only.
27
+ - `run_benchmark`: a corpus inside a folder `mt-eval contest prepare`
28
+ marks releasable (`.champollion-releasable.json` in it or above it, or a
29
+ `public/` beside `local/manifest.json`) runs into `<contest>/runs/` —
30
+ results and cache — never into the released folder; the plan's
31
+ `Results:` line says where and why. No output folder inside a releasable
32
+ folder is ever passed to the harness.
33
+ - The coached plan gives ONE verdict on the coaching file, the harness's
34
+ sentence: ✓ it names the language (as "…"), ⚠ it names neither, or not
35
+ checked — the 'Not sent: …' line is gone.
36
+ - A method plugin's cost says why, in the harness's words
37
+ (`method_loader.plugin_cost_basis`, mirrored and compared), in the plan
38
+ and the start message.
39
+ - `language_overview`: speaker counts keep their scope notes (ELCat's
40
+ "10-99" for Plains Cree is British Columbia only) and say when two
41
+ records of one source have nothing on the card telling them apart. Its
42
+ FST line and the run plan read the analyzer and the pyhfst runtime from
43
+ the harness's `config.fst_state` (forge's reading), and name the Python
44
+ they asked.
45
+ - The forge tools relay mt-eval's `score_caveats` as forge gives them
46
+ (never reworded or recomputed): `forge_export`'s summary (and each
47
+ twin-free sibling's), `forge_status`'s exports
48
+ (`summary.export_caveats`), `forge_compare` (`summary.score_caveats`)
49
+ and `forge_lint`'s `R9-harness-score-caveat`; a MAJOR one leads the next
50
+ step — a score is never "the number to quote" without it. No caveat
51
+ list → nothing said. `forge_export` also returns `hypotheses` (the file
52
+ `forge_compare` takes) and its compare command; `forge_status` in
53
+ `no-dev-set` says to call `get_training_guardrails` once, then
54
+ `forge_split`.
55
+ - **Round 12 (synthetic-user findings, 2026-10-04).**
56
+ - `run_benchmark`: a method plugin whose method.json declares dependency
57
+ class S or O with no `gateway` / `external-api` dependency is "$0 API
58
+ cost (runs on this machine)" in the plan, by the harness's rule (it was
59
+ "unknown (the plugin prices its own calls)"); it still needs the user's
60
+ attestation for a local-only corpus. A coached run's plan on a
61
+ local-only corpus says why the coaching file's first line is withheld.
62
+ - `get_run_status`: the Results block lists every score caveat the report
63
+ records (the harness's new near-constant-output caveat among them).
64
+ - `publish_report`: for a local-only corpus, WHAT GETS PUBLISHED relays
65
+ which of its facts go public with the score, which stay on this
66
+ machine, and how others read a score on a set they cannot open.
67
+ - `forge_export` returns a summary like the other forge tools;
68
+ `forge_status` names the real next step (export, or the re-audit when
69
+ the twin-free corpus predates the dev split) and a `serving` state;
70
+ `forge_register_eval` refuses a missing file with the path it looked
71
+ for and the fix.
72
+ - The `register-corpus` command the tools suggest for a test set carries
73
+ `--role test`; `search_languages` docs say a location is shown only with
74
+ its source, otherwise a link to its Glottolog record.
75
+ - **Round 11 (synthetic-user findings, 2026-10-04).**
76
+ - `run_benchmark` names a language given as a code ("sme" → "Northern
77
+ Sami", through the card) instead of prompting with the code; the plan
78
+ shows the prompt the model gets (a coaching file by its first line and
79
+ hash, and that it REPLACES the built-in prompt); for a card with several
80
+ scripts it reports the references' script share (counts only) and the
81
+ script it prompts for; it names the translation cache folder; and it
82
+ prints the exact `compare` command for runs on a registered corpus id.
83
+ - `search_languages` matches the start of a name word ("North Sami" finds
84
+ Northern Sami), ranked below exact matches; a candidate whose location
85
+ cannot be cited links its Glottolog record instead.
86
+ - `language_overview` takes its runnable open models from the CLI's
87
+ `recommend` (one rule, `-ct2` and ONNX repos excluded).
88
+ - `translate`: each row is marked `cache: pair|fallback` and every count
89
+ is derived from the rows; a discarded cached answer names what it
90
+ collided with; an unpriced or unknown model reports "cost unknown —
91
+ <why>".
92
+ - **Round 10 (synthetic-user findings, 2026-10-04).** Next steps, register
93
+ hints and the run plan put the forge steps (register, leak-audit,
94
+ predictions) before any baseline; `local-model` requires a model;
95
+ `allow_model_pair_mismatch`; `publish_report` repeats the harness's method
96
+ lines; `get_results` says publishing is explicit; locations are shown only
97
+ with a source.
98
+ - **Round 9 (synthetic-user findings, 2026-10-04).**
99
+ - New `forge_compare`: compares two systems on a registered eval set and
100
+ prints the near-twin caveat when either was trained on near-copies of
101
+ the test rows ("never the winner alone"). instructions.md now lists
102
+ which forge commands have a tool and which run in a terminal.
103
+ - `forge_prereg` takes `config_hash`; a test checks that every flag
104
+ forge's advice names is an argument of the tool that runs that command.
105
+ - `forge_leak_audit` takes `companion_config` and relays the twin-free
106
+ model's config and train command; `forge_status`,
107
+ `forge_register_eval` and `forge_leak_audit` name the preregistration
108
+ as the next step, before any benchmark, once a test set is registered.
109
+ - `get_training_guardrails` points to the public diagnosing-training
110
+ page, never a monorepo path.
111
+ - `run_benchmark` passes the resolved `--target-lang-code` (a
112
+ private-use `qaa` included), takes `script` for two-script targets
113
+ (`--target-script`), and shell-quotes every printed argument.
114
+ - `language_overview` on a private-use code (qaa–qtz) explains what the
115
+ code is and what still applies, instead of a dead end.
116
+ - `translate` reports `estimated_api_cost_label` ("$0 API cost (runs on
117
+ this machine)" for a local model) instead of "0 USD".
118
+ - `get_metric_reliability` no longer prints internal review wording.
119
+ - **Round 8 (synthetic-user findings, 2026-10-04).**
120
+ - `get_run_status`: the output trim keeps warning and notice lines (⚠,
121
+ ✗, ❌, `[WARN`, `Warning:`, `Note:`) and their indented continuation
122
+ lines from the trimmed middle, in place, up to 2,000 characters. Every
123
+ other line is still counted in a marker. A run's "COMET not computed"
124
+ notice used to sit exactly where the trim cut.
125
+ - `publish_report` passes `--anonymous` to the harness's dry run as well,
126
+ and relays the preview's trust tier (self-benchmarked, `unverified`)
127
+ and score lane.
128
+ - `run_benchmark` plans: a missing FST alone reads "the run PROCEEDS
129
+ without it" (harness 2026-10-04+), with the setup command and the
130
+ `mt-eval test <run log>` re-score. The overview's FST line uses the
131
+ harness's own sentence and never says the FST "downloads on the first
132
+ evaluation".
133
+ - `search_languages` matches the Roman-letter form of a card's
134
+ non-Latin endonym, produced by the CLI's script converter and labelled
135
+ as derived (`nêhiyawêwin` finds crk).
136
+ - `run_benchmark` no longer asks for `target_language` when one was
137
+ passed for a code with no card (private-use `qaa`).
138
+ - `translate` says when a cached answer was discarded by the
139
+ shared-output check, and does not count it as served.
140
+ - **Round 7 (synthetic-user findings, 2026-10-04).**
141
+ - `translate` with `project_dir` resolves language NAMES to the project's
142
+ own locale codes ("English"/"French" → en/fr, through the CLI's
143
+ `resolveLanguageInput`) before matching the pair and before any cache
144
+ read or write; an ambiguous name refuses with the choices. Names used to
145
+ miss the pair, run the engine default model and write cache entries
146
+ under a locale literally named "French".
147
+ - `publish_report`: publish a finished run's `*_report.json`, scores-only
148
+ if the user wants (`scores_only`, `redact_coaching`, `anonymous`), behind
149
+ the same gate as `run_benchmark` — every call shows the harness's own
150
+ `mt-eval publish --dry-run` preview and the exact `publish_ack`; only
151
+ `confirm` with those words publishes (`--prod` for production).
152
+ - `run_benchmark` plans open with the target language's `EVAL PACK:`
153
+ status (missing — with the setup command — / ready / none needed) and
154
+ name the corpus licence and `do_not_train` term the run accepts by
155
+ passing `--yes` ("unknown" when nobody states them), asked of the
156
+ harness's registry and its own eval-pack gate. New `skip_fst` /
157
+ `skip_eval_standard` (`--skip-fst` / `--skip-eval-standard`) score
158
+ without them, marked not computed; a run the harness stops for a missing
159
+ pack says so and names both.
160
+ - `forge_split` takes `near_dupe` (`--near-dupe`) and `max_group`
161
+ (`--max-group`, with `near_dupe`) and relays forge's near-twin advice
162
+ instead of promising `--near-dupe 0.6`.
163
+ - `get_run_status` keeps a long log's first and last lines and trims the
164
+ middle with `… [N lines trimmed] …` (it began mid-sentence); the
165
+ harness's `EVAL PACK` lines lead the answer.
166
+ - The register-corpus command the guidance prints now runs as printed
167
+ (`--yes`, `--domain`; it writes the local-only sidecar itself).
168
+ - Significance wording: a small Δ can be significant yet not meaningful
169
+ (check the CI on Δ and metric reliability; p-values are per metric,
170
+ uncorrected) — no more "probably noise" beside a significant result.
171
+ - Counts agree with their nouns: "1 is marked do_not_train", "1 report".
172
+ - **Round 6 (synthetic hospital, school and researcher personas, 2026-10-04).**
173
+ - `run_benchmark { publish: true }` shows WHAT GETS PUBLISHED before
174
+ anything runs — every row with its sentence text, or scores only; the
175
+ prompt published (a coaching file in full), redacted, or none; the
176
+ target — read from the harness's own publish gates. A real publish needs
177
+ `publish_ack` in the exact words the plan prints; without them (or when
178
+ the harness cannot be asked) it is REFUSED and nothing runs.
179
+ - One attestation rule everywhere: the harness's own `method:
180
+ "local-model"` runs in-process and needs no attestation (an
181
+ `attest_local_transport` for it is refused as meaningless); MT engines
182
+ and `method_dir` plugins still need the user's. The server instructions,
183
+ the refusal message and the tool descriptions used to give two answers.
184
+ - `model` with `method_dir` is passed to the plugin (`-m`, its own naming),
185
+ as the methods spec and the CLI say; `provider` with a method stays
186
+ refused. A local model directory is accepted as `model` for
187
+ `local-model` too.
188
+ - `source_language` for corpus mode (`--source-lang`); when the steward's
189
+ sidecar or the corpus card it names states the pair, its source code is
190
+ passed (`--source-code`) and the harness names the language — the plan
191
+ says which. The run card's source language used to be blank.
192
+ - One cost rule: a loopback or in-process run is "$0 API cost (runs on
193
+ this machine)" in the plan as in the start message and the run card (the
194
+ plan said "unknown, never $0").
195
+ - `forge_prereg_verdict`: record the USER's verdict on a free-text
196
+ prediction (ledgered with who, when and a note; shown as a human verdict).
197
+ - `forge_status` after `forge_init` reports `initialized` and points at
198
+ `forge_split` (it said "discover, then init" again); `summary.runs` lists
199
+ every run with its checkpoint, dev score and exports, and
200
+ `summary.warnings` carries a saturated dev set. `forge_report` on an
201
+ exported run includes the test result.
202
+ - `translate` with `project_dir` runs the project pair's `fallback` exactly
203
+ as `champollion sync` does (the CLI's own `translateWithFallback`) and
204
+ marks each text the fallback produced, with its method.
205
+ - `search_languages` matches a multi-word query part by part
206
+ ("Plains Cree nêhiyawêwin" → crk), ranked by how much of the query each
207
+ language's recorded names cover.
208
+ - `language_overview` separates "an FST exists (the card)" from "the
209
+ installed harness has a pin for it and can use it here".
210
+
211
+ - **`translate` with `project_dir` shares `champollion sync`'s cache —
212
+ exactly.** The tool keyed entries itself: the resolved code (`fra`) for the
213
+ project's `fr`, a hash of the register text for the card preset sync keys
214
+ by name (`formal-vous`), and an endpoint suffix sync never adds. A persona
215
+ got 0 of 2 hits on strings sync had just cached. With `project_dir` the
216
+ pair now comes from the CLI's own config and pair code — what
217
+ `champollion sync --method <method> [--model <model>]` runs there — so
218
+ model, register, script, cache key and locale code are sync's, in both
219
+ directions. The answer names the project pair and the key. Without
220
+ `project_dir` the server's own cache keeps its keys. A register given as a
221
+ card preset name is now sent as that preset's instructions (the name itself
222
+ used to reach the model).
223
+ - **`translate` never picks an orthography.** A target with two real writing
224
+ systems (Plains Cree: SRO or Syllabics) is refused until one is chosen, as
225
+ `champollion sync` refuses: pass the new `script` argument (`"Latn"`,
226
+ `"Cans"`), or let the `project_dir` pair's `script` decide. A chosen display
227
+ script is produced from the cached working-script translation, exactly as
228
+ sync writes it.
229
+ - **`translate` reports `isError: true` when no text was translated** (an
230
+ unreachable engine, say: it answered "0 of 2" as a success). A partial
231
+ result stays a normal answer that lists each failure.
232
+ - **`search_languages` tells same-named languages apart.** Every result says
233
+ where the language is spoken (countries, Glottolog's point, macroarea) and
234
+ its other names, each with its source ("Atya" → six Ayta languages, all in
235
+ the Philippines: only their points differ). Ties are called out. In an npm
236
+ install, name-only results are filled in from their published cards (first
237
+ 10, within 6 s, cached).
238
+ - **One argument name for one language.** Every tool that takes one language
239
+ (`search_languages`, `get_language`, `language_overview`,
240
+ `get_metric_reliability`, `forge_discover`, `forge_init`) also accepts
241
+ `language`; `get_metric_reliability {"language": "crk"}` no longer fails
242
+ with -32602. The original names still work; a missing or conflicting
243
+ language is a tool error naming both arguments. README, instructions.md and
244
+ the champollion.dev MCP page list every tool's arguments, and tests hold
245
+ them equal to the schemas.
246
+ - **`forge_status` knows the `training` state.** While a run holds the
247
+ workspace run lock, nmt-forge reports `training`; the tool's description
248
+ lists it and its hint says wait — never start another run — instead of the
249
+ generic `nmt-forge run config.json` example.
250
+ - **`language_overview` shows every source on a disputed fact.** The summary
251
+ line kept three unattributed endangerment values (Plains Cree has five,
252
+ from three sources) and joined disputed families without their sources; it
253
+ now lists each value with its source and says the sources differ.
254
+
255
+ - **New `language_overview`** — the stage-1 answer. One honest page per
256
+ language: what the index knows, which benchmarks exist, published results,
257
+ which methods can run here and with what evidence (the CLI's own
258
+ `champollion recommend` logic), tooling (FSTs, dictionaries, keyboards),
259
+ licence/consent constraints read from the corpora, contests, and numbered
260
+ next steps that each name the exact tool or command (protect your data →
261
+ baseline → build → prove → deploy). Every section that cannot be reached
262
+ says "unavailable: why" instead of sinking the answer.
263
+ - **New `get_language`** — the full cited card for one language, resolved
264
+ exactly as the `champollion` CLI resolves it (bundled card → per-user cache
265
+ → champollion.dev's published card tables, via the package's own async
266
+ prefetch). Every value carries its source; disagreements list every claim;
267
+ absent fields are named, with what absence means for the tier the card came
268
+ from. From an npm install, most low-resource languages used to answer
269
+ "family: unknown, speakers: unknown".
270
+ - **Fuzzy `search_languages`.** No exact or whole-word hit → the closest
271
+ names by restricted Damerau-Levenshtein distance (a swapped letter pair
272
+ costs ½; budget 1 edit for 4–5 letters, 2 beyond; names only, never codes).
273
+ "Atya" now surfaces the six Ayta languages first. Matching is accent- and
274
+ punctuation-insensitive, and a whole-word hit ("Cree" in "Plains Cree")
275
+ outranks a prefix ("Creek").
276
+ - **`run_benchmark` no longer publishes by default.** 0.1.x auto-published
277
+ every budget/top queue run to the PRODUCTION leaderboard unless the agent
278
+ passed `publish:false`. Now nothing is published unless `publish: true`;
279
+ every plan, start and completion message names the target (production, or
280
+ the non-production project in `MT_EVAL_SUPABASE_URL`), and a production
281
+ publish passes the harness's separate `--prod` opt-in.
282
+ - **`run_benchmark` corpus mode + local models.** Run ANY corpus — a registry
283
+ id or a test file the user holds — on any model: `provider: "local"` (an
284
+ OpenAI-compatible server on this machine; Ollama by default, `base_url` /
285
+ `LOCAL_API_BASE` otherwise), `method: "local-model"` (NLLB / OPUS-MT /
286
+ MADLAD weights), a hosted provider, or an MT API. Optional `max_cost`,
287
+ `attest_no_training`, `accept_nc_terms`, `anonymous`, field names.
288
+ - **Refusals stay refusals.** A data steward's `<file>.champollion.json`
289
+ `{"transmission":"local-only"}` mark refuses every remote model up front
290
+ (no job, no spawn). The harness's transmission-policy, NC-terms, prod-
291
+ publish and cost-cap stops come back from `get_run_status` as REFUSED with
292
+ what IS allowed — never as a generic FAILED, and never with advice to try
293
+ another provider.
294
+ - **Cost "unknown" is never $0.** `get_results` shows each row's cost and
295
+ "unknown" for unpriced/local runs (sorted last by cost); `get_run_card`
296
+ says so when `total_cost_usd` is null.
297
+ - **New read-only contest tools `list_contests` / `get_contest`.** Phases
298
+ (with the active window), the organizer's declared terms (the frozen
299
+ PROMISE keys, mirrored from `contest_policy.py` and parity-tested) with a
300
+ computed SHA-256 digest, results visibility, and the public ranking when one
301
+ is visible (the frozen final ranking with CIs and tie groups after close;
302
+ interim published scores when results are immediate). Bounded anon reads;
303
+ an entrant's sign-in email (`submitted_by`, `created_by`, `closed_by`) is
304
+ never read or shown; what an anonymous reader cannot see is said plainly.
305
+ Entering a contest remains a human-authorized CLI flow.
306
+ - **forge tools work from a pip install.** forge resolves as
307
+ `NMT_FORGE_BIN` → `CHAMPOLLION_FORGE_DIR` (a clone; set-but-wrong is an
308
+ error) → the monorepo sibling → `nmt-forge` on `PATH` → `python -m
309
+ nmt_forge.cli` from the active Python. The misleading "forge is not on
310
+ PyPI" message is gone; a missing forge now says `pip install nmt-forge`
311
+ (and `pip install 'nmt-forge[hf]'` to train), and common cold-start
312
+ tracebacks (missing harness, missing torch, no card directory) map to fixes.
313
+ - **forge tools speak nmt-forge 0.2.0's `--json` contract.** Every forge
314
+ tool passes `--json` and returns `{result, summary?, next?}` (forge 0.2.0
315
+ prints human text by default for `init`, `split`, `leak-audit` and
316
+ `registry add`, which the old text fallback relayed without structure). A
317
+ refusal — `{"error": {type, guard, message, why, fix, …}}`, exit 2 — comes
318
+ back as a tool error with what / why / fix plus the envelope; a pre-0.2.0
319
+ forge, a crash or a missing dependency comes back as the fix, never a
320
+ traceback. Each call is bounded (2 min; 10 min for evaluate/export): past
321
+ the bound forge gets SIGINT, so its own cleanup runs, then SIGKILL, and the
322
+ tool names the terminal command.
323
+ - **New `forge_export`**: score the test battery once (prereg-gated, CIs)
324
+ and package the model, an mt-eval RunLog + TestReport, a champollion
325
+ `method.json` and DEPLOY.md. **New `forge_prereg_template`**: the one
326
+ valid predictions format, written to edit.
327
+ - `forge_preflight` takes the `evaluate` / `export` / `serve` targets and a
328
+ `config`; a failing gate (forge exits 2) is an answer — `summary.passed`,
329
+ `summary.failing` — not a tool error.
330
+ - `forge_init` takes `model` (`cpu-tiny` default, `cpu-finetune` + `base`,
331
+ `nllb-600m`), `no_card` + `name` and `cards_dir`, and drops `workspace`
332
+ (forge's `init` never read it: the workspace is `<dir>/.forge`); `forge_split` takes
333
+ `test: 0`; `forge_evaluate` takes `harness_out`; corpora and eval files
334
+ may be `.tsv`.
335
+ - `forge_status` knows the `exported` state and maps its next command onto
336
+ tools (`summary.tools`); training and `nmt-forge serve` are terminal
337
+ steps, never tools.
338
+ - Every forge tool but `forge_init` (which creates the project at `dir`)
339
+ takes `project_dir`: forge runs from there, as in its
340
+ own `cd <project> && nmt-forge …` advice, so config.json's relative paths
341
+ and the `.forge` workspace resolve as `forge_init` wrote them.
342
+ - `forge_discover` no longer says it needs a card directory — forge 0.2.0
343
+ resolves cards from a pip install.
344
+ - **`get_metric_reliability` works from an npm install** — it now reads the
345
+ index the `champollion` package ships (`shared/catalogue/`), not only the
346
+ monorepo copy.
347
+ - **`translate` with the local engine works.** Keylessness is derived from
348
+ the method registry's loopback `default_base_url`, so method `local` no
349
+ longer demands an endpoint variable as if it were a key, and engines other
350
+ than the OpenRouter lane no longer receive the OpenRouter default model slug.
351
+ - **`translate` can target the model you just deployed.** A synthetic user
352
+ served their model with `nmt-forge serve`, passed `endpoint` to `translate`
353
+ (which had no such argument) — zod stripped it silently and the call ran on
354
+ a different local model. Now `base_url` points method `local` (or `openai`)
355
+ at an OpenAI-compatible server (serve's `/v1` URL), and `endpoint` drives
356
+ the new method `api` (the champollion API contract, serve's `/translate`;
357
+ key `CHAMPOLLION_API_KEY`, none needed on loopback). The engine is checked
358
+ to have taken the target before anything is sent. The schema is strict: an
359
+ unknown argument is refused by name (so is `run_benchmark`'s), and a
360
+ `model` sent to a machine-translation API, a `base_url` for an engine that
361
+ cannot use one, or an `endpoint` without method `api` is refused too. Every
362
+ answer has an `Engine:` line naming the engine that ran and, where it has
363
+ them, its model (or "engine default"), endpoint and key source — and names
364
+ the Translation Memory file it used. The local lanes' TM entries are keyed on the endpoint, so your model
365
+ is never served Ollama's cached output.
366
+ - **`translate` TM location is explicit; `project_dir` uses a project's.**
367
+ The tool's own TM is `~/.champollion-mcp/.champollion/tm.json`, separate
368
+ from every project's (`CHAMPOLLION_MCP_HOME` moves it). `project_dir` reads
369
+ and writes `<project>/.champollion/tm.json` instead (the file `champollion
370
+ sync` uses) and runs the engine there, with that project's coaching and
371
+ `.env`, as the CLI does.
372
+ - **`translate` failures say why, and engine output never reaches stdout.**
373
+ The engines report failures on the console and return nothing; the tool
374
+ used to answer "translation failed". Their output is now captured per call
375
+ (the reason becomes the failure, the last lines are shown) and echoed to
376
+ stderr. Progress lines that DeepL / Google / Microsoft / LibreTranslate and
377
+ the api engine print to stdout no longer corrupt the JSON-RPC stream.
378
+ - **`run_benchmark` jobs survive a server restart.** Hosts restart MCP
379
+ servers; a job then "never existed" although the harness had written its
380
+ results. Every job is now recorded in `~/.champollion-mcp/jobs.json` (the
381
+ newest 50; running jobs are kept) with its command, working directory,
382
+ where results land, status, start time and pid — never its environment.
383
+ mt-eval runs under a small detached supervisor that writes the run's
384
+ output and exit status to `jobs/<id>/`, so the run outlives the server and
385
+ its outcome is recorded with nobody watching. `get_run_status` on any
386
+ server answers from that evidence: RUNNING (supervisor alive, checked
387
+ against its command line so a recycled pid does not count), COMPLETED /
388
+ FAILED from the exit record, or INTERRUPTED with the log tail — with the
389
+ harness's report path and headline numbers read from the results folder.
390
+ Item and corpus runs get their own `--output-dir` inside the job folder;
391
+ queue runs report under the harness's `eval/logs/harness/queue/`. Job ids
392
+ are now random (`run-3f9c2a7b1e04`), unique across restarts. Corpus-mode
393
+ plans name the endpoint a local run will call.
394
+ - **New prompt `start_language_project`**; `explore_language` now points at
395
+ `get_language` / `language_overview`. `get_project_info` reports the
396
+ language count from the loaded index instead of a hand-typed "7,900+".
397
+ - **Docs match the server.** README tool tables, `instructions.md` (now led
398
+ by the north-star flow and which tool serves each stage) and the tool
399
+ descriptions are checked against the live tool list by
400
+ `test/server-surface.test.js` (in-memory MCP client).
401
+ - **Contract suite** (`npm run test:contract`, opt-in): packs this server and
402
+ the local CLI, installs both into a temp prefix, spawns the INSTALLED server
403
+ over stdio with the MCP SDK client, and calls every tool with realistic
404
+ arguments under a scrubbed environment (no API keys, temp HOME) — each must
405
+ return a non-error result or an error naming an actionable prerequisite; no
406
+ stack traces, no "[object Object]", nothing over 60 s.
407
+ - Requires `champollion` ^0.4.0.
408
+
409
+ ## 0.1.2 (2026-09-06)
410
+
411
+ Closes the gap found by the 2026-09-06 benchmark-hosting beta: an agent could
412
+ list the *queue* but had no way to ask "what benchmarks exist for eng→yor?".
413
+
414
+ - **New tool `list_corpora`** (read-only, 24th tool). Lists the registered
415
+ evaluation corpora for a language pair and/or benchmark family from the
416
+ corpus registry — metadata cards only (size, license, contamination grade,
417
+ domain, provider) plus what the harness can actually do with each entry
418
+ (`fetch` on demand from its pinned upstream / `gated` behind an access
419
+ token with the exact accept-terms instructions / `quarantined`). Corpus
420
+ content is never hosted or returned. Quarantined entries are hidden by
421
+ default but always counted, so a catalogued-but-held pair (eng→crk) says
422
+ so instead of looking unsupported. Source ladder: the in-repo
423
+ `arena/datasets/registry.json` when running inside a checkout →
424
+ `champollion.dev/registry.json` (HTML-holding-page guarded) → the prod
425
+ `datasets` mirror over PostgREST (last, and labelled as lagging).
426
+ `CHAMPOLLION_CORPORA_SOURCE=registry|remote|db` forces a rung. Requires
427
+ at least one filter — the registry is thousands of rows and an unfiltered
428
+ dump is never a useful answer. Its normalizer is the JS twin of the
429
+ harness's `corpora_browse.normalize_entry`.
430
+ - **Live open-item count was under-reported 3×.** `queue_pairs` returns
431
+ one row per open pair (3,620 on prod) and PostgREST serves at most 1,000
432
+ per response; the live count summed only that first page (69,695 against
433
+ a true 211,082). The fetch now pages with `limit`/`offset` until a short
434
+ page. (A `Range` header is not honoured on this RPC by the deployed
435
+ PostgREST.) Ledger entry D15 in `arena/DATABASE_SCHEMA.md`.
436
+
437
+ ## 0.1.1 (2026-08-27)
438
+
439
+ Fixes the queue-tool timeouts found by live testing of the 0.1.0 npm release:
440
+ with the queue at 211k+ open items, `list_queue`, `estimate_cost`,
441
+ `get_project_info`, and `run_benchmark` all exceeded MCP clients' default 60s
442
+ request timeout, because the DB fetch path drained the entire `queue_top`
443
+ ranking (~423 sequential pages ≈ 3 minutes) before answering anything.
444
+
445
+ - **Bounded, purpose-fit queue fetching.** Metadata comes from
446
+ queue-preview.json plus a live open-item count from the unpaged
447
+ `queue_pairs` RPC; ranked items are paged from `queue_top` only as deep as
448
+ the caller's selection needs (fetch-until-satisfied, bounded by
449
+ `CHAMPOLLION_QUEUE_MAX_PAGES`, default 20 pages / 10,000 rows); single
450
+ items are read by primary key over PostgREST with a verified-coverage
451
+ probe. When a bound truncates a search, the tool says how deep it looked —
452
+ no silent caps.
453
+ - **`get_queue_item` / `run_benchmark(item_id)`** now do a direct by-id (or
454
+ mode+priority) lookup instead of scanning a full drain, and refuse items
455
+ already covered by a VERIFIED run instead of re-spending on them.
456
+ - **Queue-mode `dry_run` is backgrounded** like a real run (the installed
457
+ harness loads the full queue before printing its plan): it returns a job id
458
+ immediately; the plan arrives via `get_run_status`.
459
+ - **Failure ladder hardened.** A slow-but-alive DB degrades to a
460
+ truncated-but-honest prefix; a dead DB falls back to the static queue.json
461
+ blob; when both are down the error names both causes.
462
+ - Harness (mt-eval, monorepo): `--top N` runs now page the DB queue only as
463
+ deep as selection needs, and `CHAMPOLLION_QUEUE_SOURCE=blob` is accepted as
464
+ a sentinel (previously read as a literal file path).
465
+
466
+ ## 0.1.0 (2026-08-27)
467
+
468
+ Initial npm release.