champollion-mcp-server 0.1.1 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,451 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.3.0 (2026-10-05)
4
+
5
+ Requires champollion 0.5.0.
6
+
7
+ - **Exact model slugs only.** The `model` arguments of `translate` and
8
+ `run_benchmark` take an exact slug (`google/gemini-3.5-flash`); a retired
9
+ alias (`gemini-flash`) or a floating id (`…-latest`, `~vendor/…`) is
10
+ refused with the slug to write, as the CLI and the harness refuse it.
11
+ - Through champollion 0.5.0, `translate` never writes an `[EN] ` marker: a
12
+ part the quality gate refuses twice keeps its source text, and the result
13
+ says so.
14
+
15
+ ## 0.2.0 (2026-10-03)
16
+
17
+ Built around one flow: someone tells their agent "let's build a Cree model
18
+ for our school" (or "an Atya phrasebook for the clinic") and the agent can
19
+ take them from discovery to a measured, deployable model. Synthetic-user
20
+ testing against the 0.1.x npm install found the gaps this release closes.
21
+
22
+ - **Round 13 (synthetic-user findings, 2026-10-04).**
23
+ - New `preview_publish`: the READ-ONLY half of `publish_report` — the
24
+ harness's own `mt-eval publish <report> --dry-run`, WHAT GETS PUBLISHED,
25
+ the exact `publish_ack` and the exact `publish_report` call that would
26
+ publish it. It has no `confirm` or `publish_ack` (refused by name) and
27
+ carries the MCP annotation `readOnlyHint: true`, so an agent host that
28
+ gates writes can allow it alone (a host had blocked publish_report's
29
+ preview as a production deploy). `publish_report` is annotated
30
+ `destructiveHint` / `openWorldHint`; without `confirm` it still returns
31
+ the preview. 34 tools.
32
+ - `run_benchmark`: `metricx` (+ `metricx_model`) and `fuse` pass the
33
+ harness's opt-in `--metricx` / `--metricx-model` / `--fuse`; `comet`
34
+ requires COMET (the harness computes it whenever unbabel-comet is
35
+ installed — there is no run flag). The plan says, from the harness's
36
+ Python, whether each is installed, what to install and what it
37
+ downloads; a confirmed run that asks for a metric the harness cannot
38
+ compute is REFUSED, never run without it. Item and corpus runs only.
39
+ - `run_benchmark`: a corpus inside a folder `mt-eval contest prepare`
40
+ marks releasable (`.champollion-releasable.json` in it or above it, or a
41
+ `public/` beside `local/manifest.json`) runs into `<contest>/runs/` —
42
+ results and cache — never into the released folder; the plan's
43
+ `Results:` line says where and why. No output folder inside a releasable
44
+ folder is ever passed to the harness.
45
+ - The coached plan gives ONE verdict on the coaching file, the harness's
46
+ sentence: ✓ it names the language (as "…"), ⚠ it names neither, or not
47
+ checked — the 'Not sent: …' line is gone.
48
+ - A method plugin's cost says why, in the harness's words
49
+ (`method_loader.plugin_cost_basis`, mirrored and compared), in the plan
50
+ and the start message.
51
+ - `language_overview`: speaker counts keep their scope notes (ELCat's
52
+ "10-99" for Plains Cree is British Columbia only) and say when two
53
+ records of one source have nothing on the card telling them apart. Its
54
+ FST line and the run plan read the analyzer and the pyhfst runtime from
55
+ the harness's `config.fst_state` (forge's reading), and name the Python
56
+ they asked.
57
+ - The forge tools relay mt-eval's `score_caveats` as forge gives them
58
+ (never reworded or recomputed): `forge_export`'s summary (and each
59
+ twin-free sibling's), `forge_status`'s exports
60
+ (`summary.export_caveats`), `forge_compare` (`summary.score_caveats`)
61
+ and `forge_lint`'s `R9-harness-score-caveat`; a MAJOR one leads the next
62
+ step — a score is never "the number to quote" without it. No caveat
63
+ list → nothing said. `forge_export` also returns `hypotheses` (the file
64
+ `forge_compare` takes) and its compare command; `forge_status` in
65
+ `no-dev-set` says to call `get_training_guardrails` once, then
66
+ `forge_split`.
67
+ - **Round 12 (synthetic-user findings, 2026-10-04).**
68
+ - `run_benchmark`: a method plugin whose method.json declares dependency
69
+ class S or O with no `gateway` / `external-api` dependency is "$0 API
70
+ cost (runs on this machine)" in the plan, by the harness's rule (it was
71
+ "unknown (the plugin prices its own calls)"); it still needs the user's
72
+ attestation for a local-only corpus. A coached run's plan on a
73
+ local-only corpus says why the coaching file's first line is withheld.
74
+ - `get_run_status`: the Results block lists every score caveat the report
75
+ records (the harness's new near-constant-output caveat among them).
76
+ - `publish_report`: for a local-only corpus, WHAT GETS PUBLISHED relays
77
+ which of its facts go public with the score, which stay on this
78
+ machine, and how others read a score on a set they cannot open.
79
+ - `forge_export` returns a summary like the other forge tools;
80
+ `forge_status` names the real next step (export, or the re-audit when
81
+ the twin-free corpus predates the dev split) and a `serving` state;
82
+ `forge_register_eval` refuses a missing file with the path it looked
83
+ for and the fix.
84
+ - The `register-corpus` command the tools suggest for a test set carries
85
+ `--role test`; `search_languages` docs say a location is shown only with
86
+ its source, otherwise a link to its Glottolog record.
87
+ - **Round 11 (synthetic-user findings, 2026-10-04).**
88
+ - `run_benchmark` names a language given as a code ("sme" → "Northern
89
+ Sami", through the card) instead of prompting with the code; the plan
90
+ shows the prompt the model gets (a coaching file by its first line and
91
+ hash, and that it REPLACES the built-in prompt); for a card with several
92
+ scripts it reports the references' script share (counts only) and the
93
+ script it prompts for; it names the translation cache folder; and it
94
+ prints the exact `compare` command for runs on a registered corpus id.
95
+ - `search_languages` matches the start of a name word ("North Sami" finds
96
+ Northern Sami), ranked below exact matches; a candidate whose location
97
+ cannot be cited links its Glottolog record instead.
98
+ - `language_overview` takes its runnable open models from the CLI's
99
+ `recommend` (one rule, `-ct2` and ONNX repos excluded).
100
+ - `translate`: each row is marked `cache: pair|fallback` and every count
101
+ is derived from the rows; a discarded cached answer names what it
102
+ collided with; an unpriced or unknown model reports "cost unknown —
103
+ <why>".
104
+ - **Round 10 (synthetic-user findings, 2026-10-04).** Next steps, register
105
+ hints and the run plan put the forge steps (register, leak-audit,
106
+ predictions) before any baseline; `local-model` requires a model;
107
+ `allow_model_pair_mismatch`; `publish_report` repeats the harness's method
108
+ lines; `get_results` says publishing is explicit; locations are shown only
109
+ with a source.
110
+ - **Round 9 (synthetic-user findings, 2026-10-04).**
111
+ - New `forge_compare`: compares two systems on a registered eval set and
112
+ prints the near-twin caveat when either was trained on near-copies of
113
+ the test rows ("never the winner alone"). instructions.md now lists
114
+ which forge commands have a tool and which run in a terminal.
115
+ - `forge_prereg` takes `config_hash`; a test checks that every flag
116
+ forge's advice names is an argument of the tool that runs that command.
117
+ - `forge_leak_audit` takes `companion_config` and relays the twin-free
118
+ model's config and train command; `forge_status`,
119
+ `forge_register_eval` and `forge_leak_audit` name the preregistration
120
+ as the next step, before any benchmark, once a test set is registered.
121
+ - `get_training_guardrails` points to the public diagnosing-training
122
+ page, never a monorepo path.
123
+ - `run_benchmark` passes the resolved `--target-lang-code` (a
124
+ private-use `qaa` included), takes `script` for two-script targets
125
+ (`--target-script`), and shell-quotes every printed argument.
126
+ - `language_overview` on a private-use code (qaa–qtz) explains what the
127
+ code is and what still applies, instead of a dead end.
128
+ - `translate` reports `estimated_api_cost_label` ("$0 API cost (runs on
129
+ this machine)" for a local model) instead of "0 USD".
130
+ - `get_metric_reliability` no longer prints internal review wording.
131
+ - **Round 8 (synthetic-user findings, 2026-10-04).**
132
+ - `get_run_status`: the output trim keeps warning and notice lines (⚠,
133
+ ✗, ❌, `[WARN`, `Warning:`, `Note:`) and their indented continuation
134
+ lines from the trimmed middle, in place, up to 2,000 characters. Every
135
+ other line is still counted in a marker. A run's "COMET not computed"
136
+ notice used to sit exactly where the trim cut.
137
+ - `publish_report` passes `--anonymous` to the harness's dry run as well,
138
+ and relays the preview's trust tier (self-benchmarked, `unverified`)
139
+ and score lane.
140
+ - `run_benchmark` plans: a missing FST alone reads "the run PROCEEDS
141
+ without it" (harness 2026-10-04+), with the setup command and the
142
+ `mt-eval test <run log>` re-score. The overview's FST line uses the
143
+ harness's own sentence and never says the FST "downloads on the first
144
+ evaluation".
145
+ - `search_languages` matches the Roman-letter form of a card's
146
+ non-Latin endonym, produced by the CLI's script converter and labelled
147
+ as derived (`nêhiyawêwin` finds crk).
148
+ - `run_benchmark` no longer asks for `target_language` when one was
149
+ passed for a code with no card (private-use `qaa`).
150
+ - `translate` says when a cached answer was discarded by the
151
+ shared-output check, and does not count it as served.
152
+ - **Round 7 (synthetic-user findings, 2026-10-04).**
153
+ - `translate` with `project_dir` resolves language NAMES to the project's
154
+ own locale codes ("English"/"French" → en/fr, through the CLI's
155
+ `resolveLanguageInput`) before matching the pair and before any cache
156
+ read or write; an ambiguous name refuses with the choices. Names used to
157
+ miss the pair, run the engine default model and write cache entries
158
+ under a locale literally named "French".
159
+ - `publish_report`: publish a finished run's `*_report.json`, scores-only
160
+ if the user wants (`scores_only`, `redact_coaching`, `anonymous`), behind
161
+ the same gate as `run_benchmark` — every call shows the harness's own
162
+ `mt-eval publish --dry-run` preview and the exact `publish_ack`; only
163
+ `confirm` with those words publishes (`--prod` for production).
164
+ - `run_benchmark` plans open with the target language's `EVAL PACK:`
165
+ status (missing — with the setup command — / ready / none needed) and
166
+ name the corpus licence and `do_not_train` term the run accepts by
167
+ passing `--yes` ("unknown" when nobody states them), asked of the
168
+ harness's registry and its own eval-pack gate. New `skip_fst` /
169
+ `skip_eval_standard` (`--skip-fst` / `--skip-eval-standard`) score
170
+ without them, marked not computed; a run the harness stops for a missing
171
+ pack says so and names both.
172
+ - `forge_split` takes `near_dupe` (`--near-dupe`) and `max_group`
173
+ (`--max-group`, with `near_dupe`) and relays forge's near-twin advice
174
+ instead of promising `--near-dupe 0.6`.
175
+ - `get_run_status` keeps a long log's first and last lines and trims the
176
+ middle with `… [N lines trimmed] …` (it began mid-sentence); the
177
+ harness's `EVAL PACK` lines lead the answer.
178
+ - The register-corpus command the guidance prints now runs as printed
179
+ (`--yes`, `--domain`; it writes the local-only sidecar itself).
180
+ - Significance wording: a small Δ can be significant yet not meaningful
181
+ (check the CI on Δ and metric reliability; p-values are per metric,
182
+ uncorrected) — no more "probably noise" beside a significant result.
183
+ - Counts agree with their nouns: "1 is marked do_not_train", "1 report".
184
+ - **Round 6 (synthetic hospital, school and researcher personas, 2026-10-04).**
185
+ - `run_benchmark { publish: true }` shows WHAT GETS PUBLISHED before
186
+ anything runs — every row with its sentence text, or scores only; the
187
+ prompt published (a coaching file in full), redacted, or none; the
188
+ target — read from the harness's own publish gates. A real publish needs
189
+ `publish_ack` in the exact words the plan prints; without them (or when
190
+ the harness cannot be asked) it is REFUSED and nothing runs.
191
+ - One attestation rule everywhere: the harness's own `method:
192
+ "local-model"` runs in-process and needs no attestation (an
193
+ `attest_local_transport` for it is refused as meaningless); MT engines
194
+ and `method_dir` plugins still need the user's. The server instructions,
195
+ the refusal message and the tool descriptions used to give two answers.
196
+ - `model` with `method_dir` is passed to the plugin (`-m`, its own naming),
197
+ as the methods spec and the CLI say; `provider` with a method stays
198
+ refused. A local model directory is accepted as `model` for
199
+ `local-model` too.
200
+ - `source_language` for corpus mode (`--source-lang`); when the steward's
201
+ sidecar or the corpus card it names states the pair, its source code is
202
+ passed (`--source-code`) and the harness names the language — the plan
203
+ says which. The run card's source language used to be blank.
204
+ - One cost rule: a loopback or in-process run is "$0 API cost (runs on
205
+ this machine)" in the plan as in the start message and the run card (the
206
+ plan said "unknown, never $0").
207
+ - `forge_prereg_verdict`: record the USER's verdict on a free-text
208
+ prediction (ledgered with who, when and a note; shown as a human verdict).
209
+ - `forge_status` after `forge_init` reports `initialized` and points at
210
+ `forge_split` (it said "discover, then init" again); `summary.runs` lists
211
+ every run with its checkpoint, dev score and exports, and
212
+ `summary.warnings` carries a saturated dev set. `forge_report` on an
213
+ exported run includes the test result.
214
+ - `translate` with `project_dir` runs the project pair's `fallback` exactly
215
+ as `champollion sync` does (the CLI's own `translateWithFallback`) and
216
+ marks each text the fallback produced, with its method.
217
+ - `search_languages` matches a multi-word query part by part
218
+ ("Plains Cree nêhiyawêwin" → crk), ranked by how much of the query each
219
+ language's recorded names cover.
220
+ - `language_overview` separates "an FST exists (the card)" from "the
221
+ installed harness has a pin for it and can use it here".
222
+
223
+ - **`translate` with `project_dir` shares `champollion sync`'s cache —
224
+ exactly.** The tool keyed entries itself: the resolved code (`fra`) for the
225
+ project's `fr`, a hash of the register text for the card preset sync keys
226
+ by name (`formal-vous`), and an endpoint suffix sync never adds. A persona
227
+ got 0 of 2 hits on strings sync had just cached. With `project_dir` the
228
+ pair now comes from the CLI's own config and pair code — what
229
+ `champollion sync --method <method> [--model <model>]` runs there — so
230
+ model, register, script, cache key and locale code are sync's, in both
231
+ directions. The answer names the project pair and the key. Without
232
+ `project_dir` the server's own cache keeps its keys. A register given as a
233
+ card preset name is now sent as that preset's instructions (the name itself
234
+ used to reach the model).
235
+ - **`translate` never picks an orthography.** A target with two real writing
236
+ systems (Plains Cree: SRO or Syllabics) is refused until one is chosen, as
237
+ `champollion sync` refuses: pass the new `script` argument (`"Latn"`,
238
+ `"Cans"`), or let the `project_dir` pair's `script` decide. A chosen display
239
+ script is produced from the cached working-script translation, exactly as
240
+ sync writes it.
241
+ - **`translate` reports `isError: true` when no text was translated** (an
242
+ unreachable engine, say: it answered "0 of 2" as a success). A partial
243
+ result stays a normal answer that lists each failure.
244
+ - **`search_languages` tells same-named languages apart.** Every result says
245
+ where the language is spoken (countries, Glottolog's point, macroarea) and
246
+ its other names, each with its source ("Atya" → six Ayta languages, all in
247
+ the Philippines: only their points differ). Ties are called out. In an npm
248
+ install, name-only results are filled in from their published cards (first
249
+ 10, within 6 s, cached).
250
+ - **One argument name for one language.** Every tool that takes one language
251
+ (`search_languages`, `get_language`, `language_overview`,
252
+ `get_metric_reliability`, `forge_discover`, `forge_init`) also accepts
253
+ `language`; `get_metric_reliability {"language": "crk"}` no longer fails
254
+ with -32602. The original names still work; a missing or conflicting
255
+ language is a tool error naming both arguments. README, instructions.md and
256
+ the champollion.dev MCP page list every tool's arguments, and tests hold
257
+ them equal to the schemas.
258
+ - **`forge_status` knows the `training` state.** While a run holds the
259
+ workspace run lock, nmt-forge reports `training`; the tool's description
260
+ lists it and its hint says wait — never start another run — instead of the
261
+ generic `nmt-forge run config.json` example.
262
+ - **`language_overview` shows every source on a disputed fact.** The summary
263
+ line kept three unattributed endangerment values (Plains Cree has five,
264
+ from three sources) and joined disputed families without their sources; it
265
+ now lists each value with its source and says the sources differ.
266
+
267
+ - **New `language_overview`** — the stage-1 answer. One honest page per
268
+ language: what the index knows, which benchmarks exist, published results,
269
+ which methods can run here and with what evidence (the CLI's own
270
+ `champollion recommend` logic), tooling (FSTs, dictionaries, keyboards),
271
+ licence/consent constraints read from the corpora, contests, and numbered
272
+ next steps that each name the exact tool or command (protect your data →
273
+ baseline → build → prove → deploy). Every section that cannot be reached
274
+ says "unavailable: why" instead of sinking the answer.
275
+ - **New `get_language`** — the full cited card for one language, resolved
276
+ exactly as the `champollion` CLI resolves it (bundled card → per-user cache
277
+ → champollion.dev's published card tables, via the package's own async
278
+ prefetch). Every value carries its source; disagreements list every claim;
279
+ absent fields are named, with what absence means for the tier the card came
280
+ from. From an npm install, most low-resource languages used to answer
281
+ "family: unknown, speakers: unknown".
282
+ - **Fuzzy `search_languages`.** No exact or whole-word hit → the closest
283
+ names by restricted Damerau-Levenshtein distance (a swapped letter pair
284
+ costs ½; budget 1 edit for 4–5 letters, 2 beyond; names only, never codes).
285
+ "Atya" now surfaces the six Ayta languages first. Matching is accent- and
286
+ punctuation-insensitive, and a whole-word hit ("Cree" in "Plains Cree")
287
+ outranks a prefix ("Creek").
288
+ - **`run_benchmark` no longer publishes by default.** 0.1.x auto-published
289
+ every budget/top queue run to the PRODUCTION leaderboard unless the agent
290
+ passed `publish:false`. Now nothing is published unless `publish: true`;
291
+ every plan, start and completion message names the target (production, or
292
+ the non-production project in `MT_EVAL_SUPABASE_URL`), and a production
293
+ publish passes the harness's separate `--prod` opt-in.
294
+ - **`run_benchmark` corpus mode + local models.** Run ANY corpus — a registry
295
+ id or a test file the user holds — on any model: `provider: "local"` (an
296
+ OpenAI-compatible server on this machine; Ollama by default, `base_url` /
297
+ `LOCAL_API_BASE` otherwise), `method: "local-model"` (NLLB / OPUS-MT /
298
+ MADLAD weights), a hosted provider, or an MT API. Optional `max_cost`,
299
+ `attest_no_training`, `accept_nc_terms`, `anonymous`, field names.
300
+ - **Refusals stay refusals.** A data steward's `<file>.champollion.json`
301
+ `{"transmission":"local-only"}` mark refuses every remote model up front
302
+ (no job, no spawn). The harness's transmission-policy, NC-terms, prod-
303
+ publish and cost-cap stops come back from `get_run_status` as REFUSED with
304
+ what IS allowed — never as a generic FAILED, and never with advice to try
305
+ another provider.
306
+ - **Cost "unknown" is never $0.** `get_results` shows each row's cost and
307
+ "unknown" for unpriced/local runs (sorted last by cost); `get_run_card`
308
+ says so when `total_cost_usd` is null.
309
+ - **New read-only contest tools `list_contests` / `get_contest`.** Phases
310
+ (with the active window), the organizer's declared terms (the frozen
311
+ PROMISE keys, mirrored from `contest_policy.py` and parity-tested) with a
312
+ computed SHA-256 digest, results visibility, and the public ranking when one
313
+ is visible (the frozen final ranking with CIs and tie groups after close;
314
+ interim published scores when results are immediate). Bounded anon reads;
315
+ an entrant's sign-in email (`submitted_by`, `created_by`, `closed_by`) is
316
+ never read or shown; what an anonymous reader cannot see is said plainly.
317
+ Entering a contest remains a human-authorized CLI flow.
318
+ - **forge tools work from a pip install.** forge resolves as
319
+ `NMT_FORGE_BIN` → `CHAMPOLLION_FORGE_DIR` (a clone; set-but-wrong is an
320
+ error) → the monorepo sibling → `nmt-forge` on `PATH` → `python -m
321
+ nmt_forge.cli` from the active Python. The misleading "forge is not on
322
+ PyPI" message is gone; a missing forge now says `pip install nmt-forge`
323
+ (and `pip install 'nmt-forge[hf]'` to train), and common cold-start
324
+ tracebacks (missing harness, missing torch, no card directory) map to fixes.
325
+ - **forge tools speak nmt-forge 0.2.0's `--json` contract.** Every forge
326
+ tool passes `--json` and returns `{result, summary?, next?}` (forge 0.2.0
327
+ prints human text by default for `init`, `split`, `leak-audit` and
328
+ `registry add`, which the old text fallback relayed without structure). A
329
+ refusal — `{"error": {type, guard, message, why, fix, …}}`, exit 2 — comes
330
+ back as a tool error with what / why / fix plus the envelope; a pre-0.2.0
331
+ forge, a crash or a missing dependency comes back as the fix, never a
332
+ traceback. Each call is bounded (2 min; 10 min for evaluate/export): past
333
+ the bound forge gets SIGINT, so its own cleanup runs, then SIGKILL, and the
334
+ tool names the terminal command.
335
+ - **New `forge_export`**: score the test battery once (prereg-gated, CIs)
336
+ and package the model, an mt-eval RunLog + TestReport, a champollion
337
+ `method.json` and DEPLOY.md. **New `forge_prereg_template`**: the one
338
+ valid predictions format, written to edit.
339
+ - `forge_preflight` takes the `evaluate` / `export` / `serve` targets and a
340
+ `config`; a failing gate (forge exits 2) is an answer — `summary.passed`,
341
+ `summary.failing` — not a tool error.
342
+ - `forge_init` takes `model` (`cpu-tiny` default, `cpu-finetune` + `base`,
343
+ `nllb-600m`), `no_card` + `name` and `cards_dir`, and drops `workspace`
344
+ (forge's `init` never read it: the workspace is `<dir>/.forge`); `forge_split` takes
345
+ `test: 0`; `forge_evaluate` takes `harness_out`; corpora and eval files
346
+ may be `.tsv`.
347
+ - `forge_status` knows the `exported` state and maps its next command onto
348
+ tools (`summary.tools`); training and `nmt-forge serve` are terminal
349
+ steps, never tools.
350
+ - Every forge tool but `forge_init` (which creates the project at `dir`)
351
+ takes `project_dir`: forge runs from there, as in its
352
+ own `cd <project> && nmt-forge …` advice, so config.json's relative paths
353
+ and the `.forge` workspace resolve as `forge_init` wrote them.
354
+ - `forge_discover` no longer says it needs a card directory — forge 0.2.0
355
+ resolves cards from a pip install.
356
+ - **`get_metric_reliability` works from an npm install** — it now reads the
357
+ index the `champollion` package ships (`shared/catalogue/`), not only the
358
+ monorepo copy.
359
+ - **`translate` with the local engine works.** Keylessness is derived from
360
+ the method registry's loopback `default_base_url`, so method `local` no
361
+ longer demands an endpoint variable as if it were a key, and engines other
362
+ than the OpenRouter lane no longer receive the OpenRouter default model slug.
363
+ - **`translate` can target the model you just deployed.** A synthetic user
364
+ served their model with `nmt-forge serve`, passed `endpoint` to `translate`
365
+ (which had no such argument) — zod stripped it silently and the call ran on
366
+ a different local model. Now `base_url` points method `local` (or `openai`)
367
+ at an OpenAI-compatible server (serve's `/v1` URL), and `endpoint` drives
368
+ the new method `api` (the champollion API contract, serve's `/translate`;
369
+ key `CHAMPOLLION_API_KEY`, none needed on loopback). The engine is checked
370
+ to have taken the target before anything is sent. The schema is strict: an
371
+ unknown argument is refused by name (so is `run_benchmark`'s), and a
372
+ `model` sent to a machine-translation API, a `base_url` for an engine that
373
+ cannot use one, or an `endpoint` without method `api` is refused too. Every
374
+ answer has an `Engine:` line naming the engine that ran and, where it has
375
+ them, its model (or "engine default"), endpoint and key source — and names
376
+ the Translation Memory file it used. The local lanes' TM entries are keyed on the endpoint, so your model
377
+ is never served Ollama's cached output.
378
+ - **`translate` TM location is explicit; `project_dir` uses a project's.**
379
+ The tool's own TM is `~/.champollion-mcp/.champollion/tm.json`, separate
380
+ from every project's (`CHAMPOLLION_MCP_HOME` moves it). `project_dir` reads
381
+ and writes `<project>/.champollion/tm.json` instead (the file `champollion
382
+ sync` uses) and runs the engine there, with that project's coaching and
383
+ `.env`, as the CLI does.
384
+ - **`translate` failures say why, and engine output never reaches stdout.**
385
+ The engines report failures on the console and return nothing; the tool
386
+ used to answer "translation failed". Their output is now captured per call
387
+ (the reason becomes the failure, the last lines are shown) and echoed to
388
+ stderr. Progress lines that DeepL / Google / Microsoft / LibreTranslate and
389
+ the api engine print to stdout no longer corrupt the JSON-RPC stream.
390
+ - **`run_benchmark` jobs survive a server restart.** Hosts restart MCP
391
+ servers; a job then "never existed" although the harness had written its
392
+ results. Every job is now recorded in `~/.champollion-mcp/jobs.json` (the
393
+ newest 50; running jobs are kept) with its command, working directory,
394
+ where results land, status, start time and pid — never its environment.
395
+ mt-eval runs under a small detached supervisor that writes the run's
396
+ output and exit status to `jobs/<id>/`, so the run outlives the server and
397
+ its outcome is recorded with nobody watching. `get_run_status` on any
398
+ server answers from that evidence: RUNNING (supervisor alive, checked
399
+ against its command line so a recycled pid does not count), COMPLETED /
400
+ FAILED from the exit record, or INTERRUPTED with the log tail — with the
401
+ harness's report path and headline numbers read from the results folder.
402
+ Item and corpus runs get their own `--output-dir` inside the job folder;
403
+ queue runs report under the harness's `eval/logs/harness/queue/`. Job ids
404
+ are now random (`run-3f9c2a7b1e04`), unique across restarts. Corpus-mode
405
+ plans name the endpoint a local run will call.
406
+ - **New prompt `start_language_project`**; `explore_language` now points at
407
+ `get_language` / `language_overview`. `get_project_info` reports the
408
+ language count from the loaded index instead of a hand-typed "7,900+".
409
+ - **Docs match the server.** README tool tables, `instructions.md` (now led
410
+ by the north-star flow and which tool serves each stage) and the tool
411
+ descriptions are checked against the live tool list by
412
+ `test/server-surface.test.js` (in-memory MCP client).
413
+ - **Contract suite** (`npm run test:contract`, opt-in): packs this server and
414
+ the local CLI, installs both into a temp prefix, spawns the INSTALLED server
415
+ over stdio with the MCP SDK client, and calls every tool with realistic
416
+ arguments under a scrubbed environment (no API keys, temp HOME) — each must
417
+ return a non-error result or an error naming an actionable prerequisite; no
418
+ stack traces, no "[object Object]", nothing over 60 s.
419
+ - Requires `champollion` ^0.4.0.
420
+
421
+ ## 0.1.2 (2026-09-06)
422
+
423
+ Closes the gap found by the 2026-09-06 benchmark-hosting beta: an agent could
424
+ list the *queue* but had no way to ask "what benchmarks exist for eng→yor?".
425
+
426
+ - **New tool `list_corpora`** (read-only, 24th tool). Lists the registered
427
+ evaluation corpora for a language pair and/or benchmark family from the
428
+ corpus registry — metadata cards only (size, license, contamination grade,
429
+ domain, provider) plus what the harness can actually do with each entry
430
+ (`fetch` on demand from its pinned upstream / `gated` behind an access
431
+ token with the exact accept-terms instructions / `quarantined`). Corpus
432
+ content is never hosted or returned. Quarantined entries are hidden by
433
+ default but always counted, so a catalogued-but-held pair (eng→crk) says
434
+ so instead of looking unsupported. Source ladder: the in-repo
435
+ `arena/datasets/registry.json` when running inside a checkout →
436
+ `champollion.dev/registry.json` (HTML-holding-page guarded) → the prod
437
+ `datasets` mirror over PostgREST (last, and labelled as lagging).
438
+ `CHAMPOLLION_CORPORA_SOURCE=registry|remote|db` forces a rung. Requires
439
+ at least one filter — the registry is thousands of rows and an unfiltered
440
+ dump is never a useful answer. Its normalizer is the JS twin of the
441
+ harness's `corpora_browse.normalize_entry`.
442
+ - **Live open-item count was under-reported 3×.** `queue_pairs` returns
443
+ one row per open pair (3,620 on prod) and PostgREST serves at most 1,000
444
+ per response; the live count summed only that first page (69,695 against
445
+ a true 211,082). The fetch now pages with `limit`/`offset` until a short
446
+ page. (A `Range` header is not honoured on this RPC by the deployed
447
+ PostgREST.) Ledger entry D15 in `arena/DATABASE_SCHEMA.md`.
448
+
3
449
  ## 0.1.1 (2026-08-27)
4
450
 
5
451
  Fixes the queue-tool timeouts found by live testing of the 0.1.0 npm release: