champollion-mcp-server 0.1.0 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,47 +1,74 @@
1
1
  # champollion-mcp-server
2
2
 
3
- MCP (Model Context Protocol) server for Champollion. Lets AI agents browse the public benchmark queue, search language metadata, and run `mt-eval` benchmarks — all through natural conversation.
3
+ MCP (Model Context Protocol) server for Champollion. Lets an AI agent take a person from "we want machine translation for our language" to a measured, deployable model — find the language and what exists for it, keep the community's data private, baseline existing models (including ones on your own machine), train with guardrails, prove the result, and translate — plus browse and run the public benchmark queue.
4
4
 
5
- > **Champollion** is infrastructure for trustworthy machine translation across every language — source-available and free for noncommercial use (the evaluation harness and shared registries are open source) — the test sets and the map that show who can translate what, how good each method is, and where the gaps are. Public benchmarks on open data rank every method (human and machine); sovereign benchmarks are secret community-owned test sets we never see. The infrastructure is source-available and singly stewarded; the test sets and the methods for a community's language belong to that community — built with communities, never scraped from them. This server is the agent-facing door into that network ([champollion.dev/docs/network](https://champollion.dev/docs/network/)). This server itself is PolyForm Noncommercial 1.0.0 (see [LICENSE](LICENSE)).
5
+ > **Champollion** is infrastructure for trustworthy machine-translation evaluation across every language — source-available and free for noncommercial use (the evaluation harness and shared registries are open source) — the test sets and the map that show who can translate what, how good each method is, and where the gaps are. Public benchmarks on open data rank every method (human and machine); sovereign benchmarks are secret community-owned test sets we never see. The infrastructure is source-available and singly stewarded; the test sets and the methods for a community's language belong to that community — designed to work with communities, never hosting their corpora. This server is the agent-facing door into that network ([champollion.dev/docs/network](https://champollion.dev/docs/network/)). This server itself is PolyForm Noncommercial 1.0.0 (see [LICENSE](LICENSE)): free for noncommercial use; using it for a commercial purpose is not covered by this license. Who is covered, in plain words with examples: [Who may use this](https://champollion.dev/docs/getting-started/who-may-use-this).
6
6
 
7
7
  ## What it does
8
8
 
9
- When connected to an agent (Claude Code, Antigravity, Cursor, etc.), the server exposes tools, resources, and prompts:
9
+ When connected to an agent (Claude Code, Cursor, …), the server exposes tools, resources, and prompts built around one flow: **"let's build a model for our language — how do we start?"** — discover what exists → protect your own data → baseline → build → prove → deploy. (Contributing compute to the public benchmark is the second flow.)
10
10
 
11
11
  ### Tools
12
12
 
13
- | Tool | Type | Description |
14
- |---|---|---|
15
- | `list_queue` | Read-only | Browse open benchmark items, filter by language/model/budget |
16
- | `get_queue_item` | Read-only | Get full details for a specific queue item |
17
- | `estimate_cost` | Read-only | Estimate cost for a set of benchmark runs |
18
- | `search_languages` | Read-only | Search language cards by name, code, family, or region |
19
- | `get_project_info` | Read-only | Get a Champollion project overview |
20
- | `get_results` | Read-only | Read scored runs from the public leaderboard (closes the run → see-impact loop) |
21
- | `get_run_card` | Read-only | Get one run card (scores + method/config metadata) by id |
22
- | `get_metric_reliability` | Read-only | Which metric to TRUST for a target language — correlations with WMT human judgments, per language family ([methodology](https://champollion.dev/docs/network/specifications/metric-reliability)) |
23
- | `get_training_guardrails` | Read-only | How to train an NMT model without fooling yourself — the guardrail rules (group-disjoint splits, dev-fence, leak audits, preregistration, …) extracted from real measured failures, each naming its enforcing tool in `forge/` (nmt-forge) |
24
- | `translate` | Action | Translate texts through champollion's tested pipeline — engine choice, register conditioning, persistent Translation Memory (repeats are free), deterministic quality gate. Spends API tokens only on cache misses |
25
- | `run_benchmark` | Action | Start benchmarks via the mt-eval harness — launches in the background and returns a job id immediately |
26
- | `get_run_status` | Read-only | Poll a benchmark job by id until it completes (the run continues past the client's 60s timeout) |
27
-
28
- #### Training tools (nmt-forge)
29
-
30
- These wrap the [nmt-forge](https://github.com/gamedaysuits/Champollion) training suite. forge is part of the Champollion monorepo (not on PyPI) — clone the repo and set `CHAMPOLLION_FORGE_DIR` to its `forge/` directory; scoring additionally needs the eval harness (`pip install mt-eval`). Without forge present these tools return an actionable error rather than crashing.
31
-
32
- | Tool | Type | Description |
33
- |---|---|---|
34
- | `forge_status` | Read-only | Where am I in an nmt-forge project and what do I run next — call first and after every step |
35
- | `forge_preflight` | Read-only | Will this command refuse? Renders every gate it will hit (✓/✗ with the fix for each ✗) |
36
- | `forge_discover` | Read-only | What a language HAS — reads the SSOT language card (scripts, analyzers, dictionaries, corpora) |
37
- | `forge_init` | Action | Scaffold a forge project from a language card: workspace + starter config + NEXT_STEPS brief |
38
- | `forge_split` | Action | Carve a parallel corpus into GROUP-DISJOINT train/dev/test (shared-source/target pairs stay together) |
39
- | `forge_leak_audit` | Read-only | Screen a corpus against every registered eval set BEFORE training (exact/near-dupe detection) |
40
- | `forge_register_eval` | Action | Register an eval file in the workspace with a role: dev (fenced selection) or test (prereg-gated) |
41
- | `forge_prereg` | Action | Preregister falsifiable predictions for a test/sealed set BEFORE scoring it |
42
- | `forge_evaluate` | Action | Close the loop: decode the config's battery with the selected checkpoint, score via the mt-eval harness (forge implements zero metrics itself) |
43
- | `forge_lint` | Read-only | Diagnose a battery manifest: weak registers and the likeliest cause given co-occurring signals |
44
- | `forge_report` | Read-only | Re-render the plain-language training report (with the Diagnosis section) from a manifest |
13
+ Arguments: `name` is required, `name?` is optional. Every tool that takes ONE language also accepts it as `language` — `code` / `language` means either name works (pass one). The original names keep working.
14
+
15
+ #### Discover — what exists for a language
16
+
17
+ | Tool | Arguments | Type | Description |
18
+ |---|---|---|---|
19
+ | `search_languages` | `query` / `language`, `limit?` | Read-only | Find a language by name, endonym, code, family or region — misspellings included (no exact hit → the closest names by edit distance, e.g. "Atya" → the Ayta languages). Each result says where the language is spoken (countries, Glottolog's point, macroarea) and its other names — only the facts its card cites, each with its source, so same-named languages can be told apart; a location without a source is never shown, and the line links the language's Glottolog record instead. From an npm install, name-only results are filled in from their published cards first; those rows carry no per-field sources until champollion.dev's card tables are next uploaded, so their lines carry the Glottolog link rather than a location |
20
+ | `language_overview` | `code` / `language`, `source?` | Read-only | **Start here.** One page per language: what the index knows, benchmarks, published results, runnable methods + evidence, tooling (FSTs, dictionaries), licence/consent constraints, and numbered next steps naming the exact tool/command |
21
+ | `get_language` | `code` / `language`, `format?` | Read-only | The full cited language card — every value with its source, disagreements shown, absences stated. Resolved exactly like the `champollion` CLI (bundled card → per-user cache → champollion.dev's published card tables) |
22
+ | `list_corpora` | `source_language?`, `target_language?`, `family?` (at least one), `include_quarantined?`, `limit?` | Read-only | Registered eval corpora (benchmarks) for a pair or family — size, licence, contamination, fetch/gated/quarantined. Metadata only; corpus content is never returned |
23
+ | `get_results` | `source_language?`, `target_language?`, `model?`, `sort?`, `limit?` | Read-only | Scored runs from the public leaderboard (cost "unknown" when a run could not be priced — never $0) |
24
+ | `get_run_card` | `id` | Read-only | One run card (scores + method/config metadata) by id |
25
+ | `get_metric_reliability` | `target` / `language` | Read-only | Which metric to TRUST for a target language — correlations with WMT human judgments ([methodology](https://champollion.dev/docs/network/specifications/metric-reliability)) |
26
+
27
+ #### Baseline — measure before you build
28
+
29
+ | Tool | Arguments | Type | Description |
30
+ |---|---|---|---|
31
+ | `run_benchmark` | one mode: `budget?` or `top?` (queue), `item_id?`, or `corpus?` with `model?` (with `method_dir`: the model the plugin loads), `method?` or `method_dir?` (a method plugin directory; `local-model` needs `model` — it has no default), `allow_model_pair_mismatch?` (`local-model`: an OPUS-MT pair model naming another pair, as a related-language baseline), `attest_local_transport?` (an MT engine or plugin; never for `local-model`), `provider?`, `base_url?`, `target_language?`, `script?` (LLM runs: the ISO 15924 script the output must be in, e.g. `Cans` or `Latn` — ask the user when the target's card lists several), `source_language?`, `source_field?`, `target_field?`, `max_cost?`, `coaching_file?`, `glossary?`, `attest_no_training?`, `accept_nc_terms?`, `skip_fst?`, `skip_eval_standard?` (item and corpus runs: score without the FST / the eval-standard metrics, marked not computed), `comet?` (require COMET: refused at confirm while unbabel-comet is not installed — the harness computes it whenever it is), `metricx?` + `metricx_model?` and `fuse?` (item and corpus runs: the harness's opt-in MetricX-24 and FUSE-style comparator — the plan says whether each is installed and what it downloads; refused at confirm while not installed); then `dry_run?`, `confirm?`, `publish?`, `publish_ack?`, `anonymous?` | Action | Run the mt-eval harness: queue items (`budget`/`top`/`item_id`) **or any corpus** — a registry id or a test file you hold — on any model, including one on this machine (`provider: "local"`, `method: "local-model"`, run in-process with no attestation). Plans without `confirm: true` — the plan names the corpus licence and `do_not_train` terms the run accepts (it passes `--yes`) and the target language's `EVAL PACK:` status (a missing FST or runtime never stops the run: FST acceptance is marked not computed; any other missing piece stops it before translating), whether COMET will be computed, and where results and the cache land (a corpus in a folder `mt-eval contest prepare` marks releasable runs into the contest's `runs/`, never into the released folder); **publishes nothing unless `publish: true`** — the plan then lists what goes public (sentence text or scores only, the prompt, the target) and a real publish needs `publish_ack` in the exact words it gives. Returns a job id immediately |
32
+ | `get_run_status` | `job_id?` | Read-only | Poll a benchmark job until it completes; refusals (local-only data, licence, NC terms, a missing eval pack) come back as refusals with what is allowed. A long log keeps its first and last lines, with `… [N lines trimmed] …` between. Jobs survive a server restart — see [Local state](#local-state) |
33
+ | `preview_publish` | `report`, `scores_only?`, `redact_coaching?`, `anonymous?` | Read-only | What publishing a finished run's `*_report.json` would put on the board — the harness's own `mt-eval publish <report> --dry-run` preview: sentence text or scores only, the prompt or its sha256, the target (normally PRODUCTION) — the exact `publish_ack` words and the exact `publish_report` call that would publish it. It has no `confirm` and cannot publish (MCP annotation `readOnlyHint: true`), so a host that gates writes can allow it on its own |
34
+ | `publish_report` | `report`, `scores_only?`, `redact_coaching?`, `anonymous?`, `confirm?`, `publish_ack?` | Action (writes) | Publish a finished run's `*_report.json` (the MCP twin of `mt-eval publish`) — annotated `destructiveHint` / `openWorldHint`, so a host asks first. Every call first runs the same dry run as `preview_publish`; only `confirm: true` with that exact `publish_ack` publishes. Without `confirm` it returns the preview and writes nothing (kept for older callers; use `preview_publish` to preview) |
35
+
36
+ #### Build and prove — nmt-forge
37
+
38
+ The forge tools drive [nmt-forge](https://github.com/gamedaysuits/Champollion/tree/main/forge), a guarded NMT training suite. Install it with `python3 -m pip install nmt-forge` (0.2.0 or later; training and serving backend: `python3 -m pip install 'nmt-forge[hf]'`); the server finds it on `PATH` or in the active Python, or in a clone via `CHAMPOLLION_FORGE_DIR`. No clone is needed otherwise: forge finds language cards on its own (a card directory, a checkout, else the public card index). Without forge, every forge tool returns that install instruction instead of a traceback.
39
+
40
+ Every forge tool runs `nmt-forge … --json` and returns `{result, summary?, next?}`; a guard's refusal comes back as a tool error with forge's what / why / fix. Every forge tool except `forge_init` takes an optional `project_dir` — the directory forge runs from, as in forge's own `cd <project> && nmt-forge …`; `forge_init` creates it (at `dir`) and returns it as `project`. Training (`nmt-forge run`) and serving (`nmt-forge serve`) outlive any tool call, so they are terminal steps; `forge_status` hands back their exact commands.
41
+
42
+ | Tool | Arguments | Type | Description |
43
+ |---|---|---|---|
44
+ | `get_training_guardrails` | `topic?` | Read-only | The training rules extracted from real measured failures (group-disjoint splits, dev-fence, leak audits, preregistration, …) — call before building a pipeline |
45
+ | `forge_status` | `workspace?`, `project_dir?` | Read-only | Where am I in a forge project (initialized → … → ready-to-train → training (wait: a run is in progress) → ready-to-score → exported → serving (the chosen model answers; the app's sync is next)) and what do I run next, with the tool for each step; every run listed with its dev score, a saturated dev set flagged — call first and after every step |
46
+ | `forge_preflight` | `target`, `config?`, `workspace?`, `project_dir?` | Read-only | Will this command refuse? Every gate `run`, `evaluate`, `export`, `serve`, `score`, `split`, `prereg` or `leak-audit` will hit, with the fix (optional `config`); for `run` the same checks `run` makes (dev set, leak audit, decode length), and `next` names the command once every gate passes |
47
+ | `forge_discover` | `code` / `language`, `cards_dir?`, `workspace?`, `project_dir?` | Read-only | What a language has, as forge sees it — works from a pip install (card directory → checkout → the public card index) |
48
+ | `forge_init` | `code` / `language`, `dir?`, `pair?`, `model?`, `base?`, `no_card?`, `name?`, `cards_dir?` | Action | Scaffold a forge project from a language card; model preset `cpu-tiny` (default — a laptop CPU, minutes), `cpu-finetune` or `nllb-600m`; `no_card` + `name` for a language the index lacks |
49
+ | `forge_split` | `corpus`, `test`, `seed`, `out?` (default `data/split`), `dev?`, `register?` (a prefix, or `true` for `project`), `allow_rotate?`, `near_dupe?` (a Jaccard threshold, e.g. 0.6: near-duplicates stay on one side), `max_group?` (with `near_dupe`: cap near-duplicate groups at N rows), `workspace?`, `project_dir?` | Action | Carve a parallel corpus (`.jsonl` or `.tsv`) into GROUP-DISJOINT train/dev/test; `test: 0` keeps your own test file separate; relays forge's near-twin advice (when it recommends `--near-dupe 0.6`, or names another route) |
50
+ | `forge_leak_audit` | `corpus`, `strict?`, `clean_to?`, `drop_test_twins?`, `companion_config?`, `overwrite?`, `full_indices?`, `workspace?`, `project_dir?` | Read-only | Screen a corpus against every registered eval set before training — leaks dropped, template siblings kept and reported (`clean_to` writes the survivors); `drop_test_twins` writes the twin-free corpus to its own `clean_to` (e.g. `corpus.notwins.jsonl`) and the twin-free model's config (`config-notwins.json`, or `companion_config`; never overwritten), and names the command that trains it. A `clean_to` that a config, a run, a split or another audit uses is refused unless `overwrite`; long row-number lists come back as `{count, first}` unless `full_indices` |
51
+ | `forge_register_eval` | `name`, `path`, `role`, `source_field?`, `target_field?`, `allow_rotate?`, `workspace?`, `project_dir?` | Action | Register an eval file (`.jsonl` or `.tsv`) with a role: dev, test, or sealed. `path` is absolute or relative to `project_dir` (a file beside the project folder is `../data/test.tsv`); one found only from the server's directory is refused with the path to pass |
52
+ | `forge_prereg_template` | `out?`, `force?`, `project_dir?` | Action | The one valid predictions-file format, and a template to edit (written to `out`) |
53
+ | `forge_prereg` | `id`, `eval_set`, `predictions`, `author?`, `config_hash?`, `allow_after_reads?`, `workspace?`, `project_dir?` | Action | Preregister predictions for a test/sealed set BEFORE scoring it — and before any benchmark on it (a scoring read blocks a later prereg); `config_hash` pins it to one run |
54
+ | `forge_prereg_verdict` | `id`, `prediction`, `verdict`, `by`, `note?`, `revise?`, `workspace?`, `project_dir?` | Action | Record the USER's verdict (held / missed) on a prediction forge cannot judge — a free-text range — ledgered with who, when and a note; shown as a human verdict everywhere, never as computed |
55
+ | `forge_export` | `run_manifest`, `out`, `config?`, `no_eval?`, `no_model?`, `glossary?`, `endpoint?`, `port?`, `name?`, `force?`, `prereg?`, `workspace?`, `project_dir?` | Action | Score the test battery once (prereg-gated, CIs) and package the model, an mt-eval report, a champollion `method.json` and DEPLOY.md — then `nmt-forge serve` it in a terminal. `summary` carries the scores with CIs, the near-twin reading, the prereg's verdict counts, the twin-free model to quote (or, planned but not exported yet, its next step), mt-eval's `score_caveats` on the score and on each twin-free model's, verbatim (a MAJOR one — e.g. a near-constant output — leads the next step: never quote the score without it), and `hypotheses`, the file `forge_compare` takes |
56
+ | `forge_evaluate` | `run_manifest`, `config?`, `out_hyps?`, `harness_out?`, `glossary?`, `prereg?`, `workspace?`, `project_dir?` | Action | Score only: decode + score the selected checkpoint through the harness, with a diagnosis |
57
+ | `forge_lint` | `manifest`, `run_manifest?`, `workspace?`, `project_dir?` | Read-only | Diagnose a battery manifest: weak registers, likely cause, the lever to pull; mt-eval's caveats on the scores come through as `R9-harness-score-caveat` (`summary.harness_score_caveats`), relayed first |
58
+ | `forge_report` | `manifest`, `workspace?`, `project_dir?` | Read-only | Re-render the plain-language training report — for an exported run, with its test score, caveats and prereg verdicts |
59
+ | `forge_compare` | `eval_set`, `hyps_a`, `hyps_b`, `label_a?`, `label_b?`, `run_a?`, `run_b?`, `metric?`, `target_lang?`, `config_hash?`, `prereg?`, `override_respend?`, `workspace?`, `project_dir?` | Action | A/B two systems on a registered eval set (paired approximate randomization, prereg-gated, ledgered); the result carries each system's near-twin caveat, so a win on recall of training phrases is said to be one, and mt-eval's own score caveats on each system's outputs (`summary.score_caveats`) — no score is quotable without them |
60
+ | `list_contests` | `status?`, `language?`, `limit?` | Read-only | Contests and shared-task editions visible to an anonymous reader |
61
+ | `get_contest` | `id` | Read-only | One contest: phases, the organizer's declared terms (+ a digest), results visibility, and the public ranking when one is visible. Entering stays a human-authorized CLI flow |
62
+
63
+ #### Deploy and contribute
64
+
65
+ | Tool | Arguments | Type | Description |
66
+ |---|---|---|---|
67
+ | `translate` | `texts`, `source_language`, `target_language`, `method?`, `model?`, `base_url?`, `endpoint?`, `register?`, `project_dir?`, `context?`, `script?`, `use_tm?`, `validate?` | Action | Translate through champollion's tested pipeline — engine choice, including a model you deployed (`method: "local"` + `base_url`, or `method: "api"` + `endpoint`, e.g. `nmt-forge serve`), register conditioning, Translation Memory (repeats are free; `project_dir` uses a project's), deterministic quality gate. Every answer names the engine that ran, with its model and endpoint. With `project_dir` the repeat check is the project's: a sentence its sync caught a model repeating, or one its files hold for another text, is refused and goes to the pair's fallback. `context` is a gettext msgctxt (one for all texts, or one per text): the cache is keyed with it as sync keys that catalog entry, and the model is told it. An unknown argument is refused, not ignored |
68
+ | `get_project_info` | none | Read-only | Project overview + live queue statistics |
69
+ | `list_queue` | `language?`, `source_language?`, `model?`, `budget?`, `condition?`, `limit?` | Read-only | Open public-benchmark items, ranked; filter by language/model/budget |
70
+ | `get_queue_item` | `id?` or `priority?` (one of them) | Read-only | One queue item in full |
71
+ | `estimate_cost` | `budget?`, `language?`, `source_language?`, `model?`, `condition?` | Read-only | What a budget or filter would cost to run |
45
72
 
46
73
  ### Resources (read-only data)
47
74
 
@@ -55,6 +82,7 @@ These wrap the [nmt-forge](https://github.com/gamedaysuits/Champollion) training
55
82
 
56
83
  | Prompt | Arguments | Description |
57
84
  |---|---|---|
85
+ | `start_language_project` | `language`, `purpose?` | "We want MT for our language (for our school / clinic) — how do we start?" |
58
86
  | `contribute_compute` | `budget?`, `language?` | "I want to help — what would $X buy?" |
59
87
  | `compete_for_prize` | `language?` | "I want to build a competitive method — any prizes?" |
60
88
  | `explore_language` | `language` | "Tell me about [language] in Champollion" |
@@ -124,37 +152,30 @@ Add to `.cursor/mcp.json` (same two options):
124
152
 
125
153
  ## What a conversation looks like
126
154
 
127
- Once connected, you can talk to your agent naturally:
128
-
129
- > **You:** "I want to help with Champollion — can you devote $10 in API credits to it?"
130
- >
131
- > **Agent** uses `get_project_info` → learns about the project
132
- >
133
- > **Agent** uses `list_queue` with `budget: 10` → sees what's available
134
- >
135
- > **Agent:** "I found a few thousand open benchmark items. Your $10 could fund dozens of runs. Any preference on languages?"
136
- >
137
- > **You:** "West African languages"
138
- >
139
- > **Agent** uses `list_queue` with `language: "african"` → filters results
155
+ > **You:** "We want to build an Atya phrasebook model for our clinic. How do we start?"
140
156
  >
141
- > **Agent** uses `estimate_cost` → calculates the plan
157
+ > **Agent** uses `search_languages { query: "Atya" }` → *no exact match — six Ayta languages, equally close: Ambala Ayta (abc, 14.82°N 120.28°E), Abellen Ayta (abp, 15.41°N 120.20°E), Sorsogon Ayta (ays, 13.04°N 124.17°E) …, each located with its source* → asks which community
142
158
  >
143
- > **Agent:** "I found 18 items for Yoruba, Hausa, Igbo, Zulu, Xhosa, and Luganda. Total: ~$1.64. Ready to run?"
159
+ > **Agent** uses `language_overview { code: "abp" }` → what the index knows (speakers, endangerment — every claim cited), no registered benchmark, no MT service lists it, and the next steps
144
160
  >
145
- > **You:** "Go for it"
161
+ > **Agent:** "First, your sentences stay on your machine — we'll mark the file local-only. Then let's measure what exists before building anything."
146
162
  >
147
- > **Agent** uses `run_benchmark` with `budget: 10` → gets a **job id** back immediately (the run continues in the background)
163
+ > **Agent** uses `run_benchmark` `{ corpus: "~/clinic/test.jsonl", provider: "local", model: "llama3.1", dry_run: true }` → shows the plan → you agree → `confirm: true` → **job id**; polls `get_run_status`
148
164
  >
149
- > **Agent** polls `get_run_status` with that job id until it reports `COMPLETED`, then uses `get_results` to show what was scored
165
+ > **Agent** uses `get_training_guardrails`, then `forge_status { project_dir }` / `forge_split { corpus, test, seed, out }` / `forge_prereg { id, eval_set, predictions }` … to train your own model without fooling yourselves, `forge_export { run_manifest, out }` to prove and package it, and `nmt-forge serve` (in a terminal) to put it behind a local endpoint for `translate { texts, source_language, target_language, method: "local", base_url }`
166
+
167
+ Contributing compute is the other path: `get_project_info` → `list_queue { language?, budget? }` / `estimate_cost { budget }` → `run_benchmark { budget: 5, confirm: true }` (add `publish: true` only if you want your results on the public leaderboard) → `get_run_status { job_id }` → `get_results { target_language }`.
150
168
 
151
169
  ## Testing
152
170
 
153
171
  ```bash
154
- npm test
172
+ npm test # unit + surface tests (mocked network; parity with arena where present)
173
+ npm run test:contract # opt-in: packs this server + the local CLI, installs both into a
174
+ # temp prefix, spawns the INSTALLED server over stdio and calls
175
+ # every tool (needs network for the anon reads)
155
176
  ```
156
177
 
157
- Tests use mock data and don't make network calls. To test the server interactively:
178
+ To test the server interactively:
158
179
 
159
180
  ```bash
160
181
  npx @modelcontextprotocol/inspector node bin/server.js
@@ -165,27 +186,32 @@ npx @modelcontextprotocol/inspector node bin/server.js
165
186
  ```
166
187
  mcp-server/
167
188
  ├── bin/server.js Entry point (stdio transport)
168
- ├── instructions.md Agent behavioral guide (loaded at connect time)
189
+ ├── instructions.md Agent guide (sent in the MCP initialize result)
169
190
  ├── src/
170
- │ ├── index.js Server setup + tool/resource/prompt registration
191
+ │ ├── index.js Tool/resource/prompt registration
171
192
  │ └── tools/
193
+ │ ├── args.js The shared `language` argument (alias of code/target/query)
194
+ │ ├── languages.js Language index + search (exact, then fuzzy; where each is spoken)
195
+ │ ├── language-card.js get_language — cards via the champollion package
196
+ │ ├── overview.js language_overview — composes the other tools
197
+ │ ├── corpora.js Corpus registry (list_corpora)
198
+ │ ├── results.js Public leaderboard reads
199
+ │ ├── contests.js Read-only contest views
200
+ │ ├── reliability.js Metric-reliability lookups
172
201
  │ ├── queue.js Queue fetch, filter, cost estimation
173
- │ ├── languages.js Language card index + search
174
- │ ├── results.js Public leaderboard reads (scored run_cards)
175
- │ ├── reliability.js Metric-reliability lookups (which metric to trust)
176
- │ ├── training.js Training guardrails (get_training_guardrails)
177
- │ ├── translate.js Champollion translate pipeline wrapper
178
- │ └── harness.js mt-eval CLI wrapper
179
- ├── test/
180
- │ ├── tools.test.js Unit tests (node --test) + SSOT shared vectors
181
- │ ├── harness.test.js Unit tests for run_benchmark + the async job model
182
- │ ├── results.test.js Unit tests for the leaderboard read tools
183
- │ ├── queue-fetch.test.js Unit tests for queue fetching/caching
184
- │ ├── reliability.test.js Unit tests for metric-reliability lookups
185
- │ ├── training.test.js Unit tests for the training-guardrails tool
186
- │ └── translate.test.js Unit tests for the translate tool
187
- ├── package.json
188
- └── README.md
202
+ │ ├── harness.js mt-eval wrapper (run_benchmark, get_run_status)
203
+ │ ├── run-plan.js A plan's licence / do_not_train / EVAL PACK lines, asked of the harness
204
+ │ ├── publish-report.js preview_publish (read-only) + publish_report — the harness's own publish preview + exact acknowledgement
205
+ │ ├── output-trim.js Long job output: head + tail, the middle trimmed with a marker
206
+ │ ├── jobs.js Durable job history (jobs.json) + the job runner
207
+ │ ├── job-supervisor.js Detached runner: logs + exit record per job
208
+ │ ├── state.js The state directory (~/.champollion-mcp)
209
+ │ ├── forge.js nmt-forge wrapper (forge_*)
210
+ │ ├── training.js Training guardrails
211
+ │ └── translate.js Champollion translate pipeline wrapper
212
+ └── test/
213
+ ├── *.test.js Unit tests (node --test)
214
+ └── contract.e2e.js Installed-package contract suite (npm run test:contract)
189
215
  ```
190
216
 
191
217
  ## Protocol version — and the 2026-07-28 stateless spec
@@ -216,12 +242,13 @@ transport sessions. `run_benchmark` already works exactly that way: it returns a
216
242
  job id, and the agent passes that id to `get_run_status`. No transport-level
217
243
  session is ever involved.
218
244
 
219
- **One assumption to know about.** The job registry in `src/tools/harness.js` is
220
- in-memory and assumes a single server process for the agent's lifetime. That
221
- holds for stdio. It would **not** hold behind a stateless HTTP deployment with
222
- more than one process, where a poll could land on a process that never started
223
- the job. Anyone adding an HTTP transport must move that registry to shared
224
- storage first.
245
+ **One assumption to know about.** The job history is a file on the machine
246
+ running the server (`~/.champollion-mcp/jobs.json`, see [Local
247
+ state](#local-state)), and the runs it describes are processes on that machine.
248
+ Server processes on the same machine share it — a restarted server, or a second
249
+ one, finds every job — but servers on different machines do not. Anyone adding
250
+ an HTTP transport served from more than one machine must move the history (and
251
+ the runs) to shared storage first.
225
252
 
226
253
  **Why we have not upgraded.** The published TypeScript SDK does not speak the
227
254
  new revision yet: `@modelcontextprotocol/sdk@1.30.0` is the only dist-tag on
@@ -230,6 +257,27 @@ was raised to `^1.30.0` (from a stale `^1.12.1`, eighteen releases behind) so
230
257
  installs resolve current. Re-check when a `2026-07-28`-capable SDK publishes;
231
258
  the migration should be small given the table above.
232
259
 
260
+ ## Local state
261
+
262
+ The server keeps state in one directory, `~/.champollion-mcp/`
263
+ (`CHAMPOLLION_MCP_HOME` moves it):
264
+
265
+ | Path | What it is |
266
+ |---|---|
267
+ | `.champollion/tm.json` | The `translate` tool's own Translation Memory. It is **not** any project's `.champollion/tm.json`: pass `project_dir` to read and write a project's TM instead (the file `champollion sync` uses there — the engine then also uses that project's coaching, `.env` and its own pair for the languages, so the cache key is exactly sync's) |
268
+ | `jobs.json` | The `run_benchmark` job history — the newest 50 jobs (running jobs are never dropped): id, command, working directory, where results land, status, start time, pid while running. No environment or keys |
269
+ | `jobs/<job-id>/` | One job's `stdout.log`, `stderr.log`, `exit.json` and — for an item run or a registered corpus — `results/`, the harness's `--output-dir` (run log + `*_report.json`). A run on a file the user holds writes its results (`results/mcp-<job-id>/`) and cache (`results/cache/`) beside that file instead — or, for a file inside a folder `mt-eval contest prepare` marks releasable, to the contest's `runs/` (never into the released folder). Queue runs write their reports where the harness always does: `eval/logs/harness/queue/` under the server's working directory. When a job leaves the history its logs are deleted; its results never are |
270
+
271
+ **Jobs outlive the server.** mt-eval runs under a small detached supervisor
272
+ (`src/tools/job-supervisor.js`) that writes the run's output and exit status to
273
+ the job folder itself. If the host restarts the server mid-run, the run keeps
274
+ going, and `get_run_status` on the new server reports it from that evidence:
275
+ RUNNING (the supervisor is alive — checked against its command line, so a
276
+ recycled pid is not mistaken for it), COMPLETED / FAILED from the exit record
277
+ with the results read from the output folder, or INTERRUPTED with the log tail
278
+ when the process vanished without one (a reboot, a kill). To stop a run by hand,
279
+ `kill` the supervisor's pid (shown while RUNNING); the stop is recorded.
280
+
233
281
  ## Data sources
234
282
 
235
283
  The champollion.dev homepage map is an idealization of this data — agents
@@ -238,8 +286,13 @@ resource carries the full endpoint table).
238
286
 
239
287
  - **Queue**: Fetched from `champollion.dev/queue.json` (tens of MB — it grows with coverage; cached 5 min in memory). Small slice: `champollion.dev/queue-preview.json`. For the live open-item count, call `get_project_info`
240
288
  - **Mesh**: `champollion.dev/mesh.json` — the measured/registered pair network behind the homepage map
241
- - **Corpus registry**: `champollion.dev/registry.json` — every registered eval corpus with license lane, attribution, checksum
289
+ - **Corpus registry**: `champollion.dev/registry.json` — every registered eval corpus with license lane, attribution, checksum. The `list_corpora` tool reads the in-repo copy when run inside a checkout, else this file, else the prod `datasets` mirror (labelled as lagging); `CHAMPOLLION_CORPORA_SOURCE=registry|remote|db` forces one
242
290
  - **Provider coverage**: `shared/catalogue/method-coverage.json` — each provider's published language list, cited + as-of + `tier` (the data behind the map's covered/uncovered split). The map's green has two tiers by exact ISO-639-3 code: bright = a deployed service lists it (Google/Microsoft/DeepL/LibreTranslate); dim = only an open research model lists it (NLLB/OPUS/M2M-100/MADLAD-400 — a model-card code, not a usable service)
243
- - **Languages**: Loaded from `cli/shared/language-cards/` on startup (falls back to built-in index of 40 languages if the directory isn't accessible)
291
+ - **Languages**: the monorepo's `cli/shared/language-cards/` in a checkout; from an npm install, the `champollion` package's bundled cards (core languages in full, every catalogued language by name), with `get_language` materializing any other card through the CLI's own resolver — per-user cache (`~/.champollion/cards`), then champollion.dev's published card tables (read-only). `search_languages` fills in its name-only results the same way (the first 10 per search, within 6 s; a slower card keeps downloading into the cache and the line says so), so each can say where it is spoken once its published row carries per-field sources; until the tables' next upload, a filled-in line withholds the uncited location and links the language's Glottolog record. `CHAMPOLLION_OFFLINE=1` disables the fetch
292
+ - **Contests**: the public `contests`, `contest_phases`, `contest_submissions` (never the email column) and `shared_tasks` tables, anon read; withheld results (`contest_deferred_results`) are not anon-readable and the tools say so
244
293
  - **Results**: Read from the public Supabase leaderboard (`run_cards`) — the same anon read path the champollion.dev leaderboard uses. Scored aggregates and run-card metadata only; per-entry test sentences are never read. Override the project with `CHAMPOLLION_SUPABASE_URL` / `CHAMPOLLION_SUPABASE_ANON_KEY`.
245
- - **Harness**: Shells out to `mt-eval` CLI (must be installed separately)
294
+ - **Harness**: Shells out to the `mt-eval` CLI (`pipx install mt-eval-harness`). Publishing goes to `MT_EVAL_SUPABASE_URL` (default: the production project) and only when `publish: true`
295
+
296
+ ## License
297
+
298
+ PolyForm Noncommercial 1.0.0 ([LICENSE](LICENSE)): free to use, change and share for noncommercial purposes; using it for a commercial purpose is not covered by this license. A school, a public hospital or clinic, a charity, or a personal or research project is covered; a for-profit business's product is not. In full, with the harness (AGPL-3.0-or-later) and the other packages: [Who may use this](https://champollion.dev/docs/getting-started/who-may-use-this) — a summary, not legal advice; the license text governs.
package/bin/server.js CHANGED
@@ -8,9 +8,10 @@
8
8
  *
9
9
  * node bin/server.js
10
10
  *
11
- * The server exposes read-only tools for exploring the Champollion
12
- * benchmark queue and language metadata, plus an action tool for
13
- * running benchmarks (which requires user confirmation in the agent).
11
+ * The server exposes read-only tools for discovering what exists for a
12
+ * language (cited cards, benchmarks, results, contests), plus action tools
13
+ * that run benchmarks, drive nmt-forge training and translate — every one
14
+ * that spends or publishes requires the user's explicit confirmation.
14
15
  */
15
16
 
16
17
  import { createServer } from '../src/index.js';