champollion-mcp-server 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/LICENSE ADDED
@@ -0,0 +1,133 @@
1
+ # PolyForm Noncommercial License 1.0.0
2
+
3
+ <https://polyformproject.org/licenses/noncommercial/1.0.0>
4
+
5
+ ## Acceptance
6
+
7
+ In order to get any license under these terms, you must agree
8
+ to them as both strict obligations and conditions to all
9
+ your licenses.
10
+
11
+ ## Copyright License
12
+
13
+ The licensor grants you a copyright license for the
14
+ software to do everything you might do with the software
15
+ that would otherwise infringe the licensor's copyright
16
+ in it for any permitted purpose. However, you may
17
+ only distribute the software according to [Distribution
18
+ License](#distribution-license) and make changes or new works
19
+ based on the software according to [Changes and New Works
20
+ License](#changes-and-new-works-license).
21
+
22
+ ## Distribution License
23
+
24
+ The licensor grants you an additional copyright license
25
+ to distribute copies of the software. Your license
26
+ to distribute covers distributing the software with
27
+ changes and new works permitted by [Changes and New Works
28
+ License](#changes-and-new-works-license).
29
+
30
+ ## Notices
31
+
32
+ You must ensure that anyone who gets a copy of any part of
33
+ the software from you also gets a copy of these terms or the
34
+ URL for them above, as well as copies of any plain-text lines
35
+ beginning with `Required Notice:` that the licensor provided
36
+ with the software. For example:
37
+
38
+ > Required Notice: Copyright Yoyodyne, Inc. (http://example.com)
39
+
40
+ ## Changes and New Works License
41
+
42
+ The licensor grants you an additional copyright license to
43
+ make changes and new works based on the software for any
44
+ permitted purpose.
45
+
46
+ ## Patent License
47
+
48
+ The licensor grants you a patent license for the software that
49
+ covers patent claims the licensor can license, or becomes able
50
+ to license, that you would infringe by using the software.
51
+
52
+ ## Noncommercial Purposes
53
+
54
+ Any noncommercial purpose is a permitted purpose.
55
+
56
+ ## Personal Uses
57
+
58
+ Personal use for research, experiment, and testing for
59
+ the benefit of public knowledge, personal study, private
60
+ entertainment, hobby projects, amateur pursuits, or religious
61
+ observance, without any anticipated commercial application,
62
+ is use for a permitted purpose.
63
+
64
+ ## Noncommercial Organizations
65
+
66
+ Use by any charitable organization, educational institution,
67
+ public research organization, public safety or health
68
+ organization, environmental protection organization,
69
+ or government institution is use for a permitted purpose
70
+ regardless of the source of funding or obligations resulting
71
+ from the funding.
72
+
73
+ ## Fair Use
74
+
75
+ You may have "fair use" rights for the software under the
76
+ law. These terms do not limit them.
77
+
78
+ ## No Other Rights
79
+
80
+ These terms do not allow you to sublicense or transfer any of
81
+ your licenses to anyone else, or prevent the licensor from
82
+ granting licenses to anyone else. These terms do not imply
83
+ any other licenses.
84
+
85
+ ## Patent Defense
86
+
87
+ If you make any written claim that the software infringes or
88
+ contributes to infringement of any patent, your patent license
89
+ for the software granted under these terms ends immediately. If
90
+ your company makes such a claim, your patent license ends
91
+ immediately for work on behalf of your company.
92
+
93
+ ## Violations
94
+
95
+ The first time you are notified in writing that you have
96
+ violated any of these terms, or done anything with the software
97
+ not covered by your licenses, your licenses can nonetheless
98
+ continue if you come into full compliance with these terms,
99
+ and take practical steps to correct past violations, within
100
+ 32 days of receiving notice. Otherwise, all your licenses
101
+ end immediately.
102
+
103
+ ## No Liability
104
+
105
+ ***As far as the law allows, the software comes as is, without
106
+ any warranty or condition, and the licensor will not be liable
107
+ to you for any damages arising out of these terms or the use
108
+ or nature of the software, under any kind of legal claim.***
109
+
110
+ ## Definitions
111
+
112
+ The **licensor** is the individual or entity offering these
113
+ terms, and the **software** is the software the licensor makes
114
+ available under these terms.
115
+
116
+ **You** refers to the individual or entity agreeing to these
117
+ terms.
118
+
119
+ **Your company** is any legal entity, sole proprietorship,
120
+ or other kind of organization that you work for, plus all
121
+ organizations that have control over, are under the control of,
122
+ or are under common control with that organization. **Control**
123
+ means ownership of substantially all the assets of an entity,
124
+ or the power to direct its management and policies by vote,
125
+ contract, or otherwise. Control can be direct or indirect.
126
+
127
+ **Your licenses** are all the licenses granted to you for the
128
+ software under these terms.
129
+
130
+ **Use** means anything you do with the software requiring one
131
+ of your licenses.
132
+
133
+ Required Notice: Copyright Curtis Forbes — Champollion (https://champollion.dev)
package/README.md ADDED
@@ -0,0 +1,245 @@
1
+ # champollion-mcp-server
2
+
3
+ MCP (Model Context Protocol) server for Champollion. Lets AI agents browse the public benchmark queue, search language metadata, and run `mt-eval` benchmarks — all through natural conversation.
4
+
5
+ > **Champollion** is infrastructure for trustworthy machine translation across every language — source-available and free for noncommercial use (the evaluation harness and shared registries are open source) — the test sets and the map that show who can translate what, how good each method is, and where the gaps are. Public benchmarks on open data rank every method (human and machine); sovereign benchmarks are secret community-owned test sets we never see. The infrastructure is source-available and singly stewarded; the test sets and the methods for a community's language belong to that community — built with communities, never scraped from them. This server is the agent-facing door into that network ([champollion.dev/docs/network](https://champollion.dev/docs/network/)). This server itself is PolyForm Noncommercial 1.0.0 (see [LICENSE](LICENSE)).
6
+
7
+ ## What it does
8
+
9
+ When connected to an agent (Claude Code, Antigravity, Cursor, etc.), the server exposes tools, resources, and prompts:
10
+
11
+ ### Tools
12
+
13
+ | Tool | Type | Description |
14
+ |---|---|---|
15
+ | `list_queue` | Read-only | Browse open benchmark items, filter by language/model/budget |
16
+ | `get_queue_item` | Read-only | Get full details for a specific queue item |
17
+ | `estimate_cost` | Read-only | Estimate cost for a set of benchmark runs |
18
+ | `search_languages` | Read-only | Search language cards by name, code, family, or region |
19
+ | `get_project_info` | Read-only | Get a Champollion project overview |
20
+ | `get_results` | Read-only | Read scored runs from the public leaderboard (closes the run → see-impact loop) |
21
+ | `get_run_card` | Read-only | Get one run card (scores + method/config metadata) by id |
22
+ | `get_metric_reliability` | Read-only | Which metric to TRUST for a target language — correlations with WMT human judgments, per language family ([methodology](https://champollion.dev/docs/network/specifications/metric-reliability)) |
23
+ | `get_training_guardrails` | Read-only | How to train an NMT model without fooling yourself — the guardrail rules (group-disjoint splits, dev-fence, leak audits, preregistration, …) extracted from real measured failures, each naming its enforcing tool in `forge/` (nmt-forge) |
24
+ | `translate` | Action | Translate texts through champollion's tested pipeline — engine choice, register conditioning, persistent Translation Memory (repeats are free), deterministic quality gate. Spends API tokens only on cache misses |
25
+ | `run_benchmark` | Action | Start benchmarks via the mt-eval harness — launches in the background and returns a job id immediately |
26
+ | `get_run_status` | Read-only | Poll a benchmark job by id until it completes (the run continues past the client's 60s timeout) |
27
+
28
+ #### Training tools (nmt-forge)
29
+
30
+ These wrap the [nmt-forge](https://github.com/gamedaysuits/Champollion) training suite. forge is part of the Champollion monorepo (not on PyPI) — clone the repo and set `CHAMPOLLION_FORGE_DIR` to its `forge/` directory; scoring additionally needs the eval harness (`pip install mt-eval`). Without forge present these tools return an actionable error rather than crashing.
31
+
32
+ | Tool | Type | Description |
33
+ |---|---|---|
34
+ | `forge_status` | Read-only | Where am I in an nmt-forge project and what do I run next — call first and after every step |
35
+ | `forge_preflight` | Read-only | Will this command refuse? Renders every gate it will hit (✓/✗ with the fix for each ✗) |
36
+ | `forge_discover` | Read-only | What a language HAS — reads the SSOT language card (scripts, analyzers, dictionaries, corpora) |
37
+ | `forge_init` | Action | Scaffold a forge project from a language card: workspace + starter config + NEXT_STEPS brief |
38
+ | `forge_split` | Action | Carve a parallel corpus into GROUP-DISJOINT train/dev/test (shared-source/target pairs stay together) |
39
+ | `forge_leak_audit` | Read-only | Screen a corpus against every registered eval set BEFORE training (exact/near-dupe detection) |
40
+ | `forge_register_eval` | Action | Register an eval file in the workspace with a role: dev (fenced selection) or test (prereg-gated) |
41
+ | `forge_prereg` | Action | Preregister falsifiable predictions for a test/sealed set BEFORE scoring it |
42
+ | `forge_evaluate` | Action | Close the loop: decode the config's battery with the selected checkpoint, score via the mt-eval harness (forge implements zero metrics itself) |
43
+ | `forge_lint` | Read-only | Diagnose a battery manifest: weak registers and the likeliest cause given co-occurring signals |
44
+ | `forge_report` | Read-only | Re-render the plain-language training report (with the Diagnosis section) from a manifest |
45
+
46
+ ### Resources (read-only data)
47
+
48
+ | Resource | URI | Description |
49
+ |---|---|---|
50
+ | Contributing guide | `champollion://contributing-guide` | CONTRIBUTING.md — how to help with the project |
51
+ | Queue schema | `champollion://queue-schema` | Field definitions for every queue.json item |
52
+ | Network data map | `champollion://network-data` | Which machine artifact answers which question (queue, mesh, registry, coverage) |
53
+
54
+ ### Prompts (conversation starters)
55
+
56
+ | Prompt | Arguments | Description |
57
+ |---|---|---|
58
+ | `contribute_compute` | `budget?`, `language?` | "I want to help — what would $X buy?" |
59
+ | `compete_for_prize` | `language?` | "I want to build a competitive method — any prizes?" |
60
+ | `explore_language` | `language` | "Tell me about [language] in Champollion" |
61
+
62
+ ## Quick start
63
+
64
+ **From the published package** (no clone needed):
65
+
66
+ ```bash
67
+ npx champollion-mcp-server # start on stdio (for agent connection)
68
+
69
+ ```
70
+
71
+ **From source:**
72
+
73
+ ```bash
74
+ cd mcp-server
75
+ npm install
76
+ npm test # run unit tests
77
+ npm start # start on stdio (for agent connection)
78
+ ```
79
+
80
+ ## Connect to your agent
81
+
82
+ ### Claude Code / Antigravity
83
+
84
+ Add to your MCP configuration — published package:
85
+
86
+ ```json
87
+ {
88
+ "mcpServers": {
89
+ "champollion": {
90
+ "command": "npx",
91
+ "args": ["-y", "champollion-mcp-server"]
92
+ }
93
+ }
94
+ }
95
+ ```
96
+
97
+ Or from a local checkout:
98
+
99
+ ```json
100
+ {
101
+ "mcpServers": {
102
+ "champollion": {
103
+ "command": "node",
104
+ "args": ["/path/to/Champollion/mcp-server/bin/server.js"]
105
+ }
106
+ }
107
+ }
108
+ ```
109
+
110
+ ### Cursor
111
+
112
+ Add to `.cursor/mcp.json` (same two options):
113
+
114
+ ```json
115
+ {
116
+ "mcpServers": {
117
+ "champollion": {
118
+ "command": "npx",
119
+ "args": ["-y", "champollion-mcp-server"]
120
+ }
121
+ }
122
+ }
123
+ ```
124
+
125
+ ## What a conversation looks like
126
+
127
+ Once connected, you can talk to your agent naturally:
128
+
129
+ > **You:** "I want to help with Champollion — can you devote $10 in API credits to it?"
130
+ >
131
+ > **Agent** uses `get_project_info` → learns about the project
132
+ >
133
+ > **Agent** uses `list_queue` with `budget: 10` → sees what's available
134
+ >
135
+ > **Agent:** "I found a few thousand open benchmark items. Your $10 could fund dozens of runs. Any preference on languages?"
136
+ >
137
+ > **You:** "West African languages"
138
+ >
139
+ > **Agent** uses `list_queue` with `language: "african"` → filters results
140
+ >
141
+ > **Agent** uses `estimate_cost` → calculates the plan
142
+ >
143
+ > **Agent:** "I found 18 items for Yoruba, Hausa, Igbo, Zulu, Xhosa, and Luganda. Total: ~$1.64. Ready to run?"
144
+ >
145
+ > **You:** "Go for it"
146
+ >
147
+ > **Agent** uses `run_benchmark` with `budget: 10` → gets a **job id** back immediately (the run continues in the background)
148
+ >
149
+ > **Agent** polls `get_run_status` with that job id until it reports `COMPLETED`, then uses `get_results` to show what was scored
150
+
151
+ ## Testing
152
+
153
+ ```bash
154
+ npm test
155
+ ```
156
+
157
+ Tests use mock data and don't make network calls. To test the server interactively:
158
+
159
+ ```bash
160
+ npx @modelcontextprotocol/inspector node bin/server.js
161
+ ```
162
+
163
+ ## Architecture
164
+
165
+ ```
166
+ mcp-server/
167
+ ├── bin/server.js Entry point (stdio transport)
168
+ ├── instructions.md Agent behavioral guide (loaded at connect time)
169
+ ├── src/
170
+ │ ├── index.js Server setup + tool/resource/prompt registration
171
+ │ └── tools/
172
+ │ ├── queue.js Queue fetch, filter, cost estimation
173
+ │ ├── languages.js Language card index + search
174
+ │ ├── results.js Public leaderboard reads (scored run_cards)
175
+ │ ├── reliability.js Metric-reliability lookups (which metric to trust)
176
+ │ ├── training.js Training guardrails (get_training_guardrails)
177
+ │ ├── translate.js Champollion translate pipeline wrapper
178
+ │ └── harness.js mt-eval CLI wrapper
179
+ ├── test/
180
+ │ ├── tools.test.js Unit tests (node --test) + SSOT shared vectors
181
+ │ ├── harness.test.js Unit tests for run_benchmark + the async job model
182
+ │ ├── results.test.js Unit tests for the leaderboard read tools
183
+ │ ├── queue-fetch.test.js Unit tests for queue fetching/caching
184
+ │ ├── reliability.test.js Unit tests for metric-reliability lookups
185
+ │ ├── training.test.js Unit tests for the training-guardrails tool
186
+ │ └── translate.test.js Unit tests for the translate tool
187
+ ├── package.json
188
+ └── README.md
189
+ ```
190
+
191
+ ## Protocol version — and the 2026-07-28 stateless spec
192
+
193
+ **Transport: stdio only.** One server process per agent, launched by the client.
194
+
195
+ MCP's [2026-07-28 revision](https://blog.modelcontextprotocol.io/posts/2026-07-28/)
196
+ made the protocol **stateless by default** — the largest change since
197
+ authorization. It retires the `initialize`/`initialized` handshake and the
198
+ `Mcp-Session-Id` header, requires `Mcp-Method`/`Mcp-Name` HTTP headers so
199
+ gateways can route without parsing bodies, replaces held-open bidirectional
200
+ streams with Multi Round-Trip Requests, and deprecates Roots, Sampling, Logging
201
+ and the legacy HTTP+SSE transport (twelve-month support window).
202
+
203
+ **Where this server stands (checked 2026-08-01):**
204
+
205
+ | | Status |
206
+ |---|---|
207
+ | Deprecated capabilities (Roots / Sampling / Logging) | **None used.** |
208
+ | Legacy HTTP+SSE transport | **Not used** — stdio only. |
209
+ | `Mcp-Session-Id`, header routing, MRTR | **Not applicable** to stdio. |
210
+ | Application state across calls | **Already uses the prescribed pattern** — see below. |
211
+ | SDK support for `2026-07-28` | **Not yet available.** |
212
+
213
+ The new spec's guidance for cross-call state is to "mint an explicit handle from
214
+ a tool and have the model pass it back as an argument" rather than lean on
215
+ transport sessions. `run_benchmark` already works exactly that way: it returns a
216
+ job id, and the agent passes that id to `get_run_status`. No transport-level
217
+ session is ever involved.
218
+
219
+ **One assumption to know about.** The job registry in `src/tools/harness.js` is
220
+ in-memory and assumes a single server process for the agent's lifetime. That
221
+ holds for stdio. It would **not** hold behind a stateless HTTP deployment with
222
+ more than one process, where a poll could land on a process that never started
223
+ the job. Anyone adding an HTTP transport must move that registry to shared
224
+ storage first.
225
+
226
+ **Why we have not upgraded.** The published TypeScript SDK does not speak the
227
+ new revision yet: `@modelcontextprotocol/sdk@1.30.0` is the only dist-tag on
228
+ npm and its `LATEST_PROTOCOL_VERSION` is `2025-11-25`. The dependency floor here
229
+ was raised to `^1.30.0` (from a stale `^1.12.1`, eighteen releases behind) so
230
+ installs resolve current. Re-check when a `2026-07-28`-capable SDK publishes;
231
+ the migration should be small given the table above.
232
+
233
+ ## Data sources
234
+
235
+ The champollion.dev homepage map is an idealization of this data — agents
236
+ should read the sources, not the picture (the `champollion://network-data`
237
+ resource carries the full endpoint table).
238
+
239
+ - **Queue**: Fetched from `champollion.dev/queue.json` (tens of MB — it grows with coverage; cached 5 min in memory). Small slice: `champollion.dev/queue-preview.json`. For the live open-item count, call `get_project_info`
240
+ - **Mesh**: `champollion.dev/mesh.json` — the measured/registered pair network behind the homepage map
241
+ - **Corpus registry**: `champollion.dev/registry.json` — every registered eval corpus with license lane, attribution, checksum
242
+ - **Provider coverage**: `shared/catalogue/method-coverage.json` — each provider's published language list, cited + as-of + `tier` (the data behind the map's covered/uncovered split). The map's green has two tiers by exact ISO-639-3 code: bright = a deployed service lists it (Google/Microsoft/DeepL/LibreTranslate); dim = only an open research model lists it (NLLB/OPUS/M2M-100/MADLAD-400 — a model-card code, not a usable service)
243
+ - **Languages**: Loaded from `cli/shared/language-cards/` on startup (falls back to built-in index of 40 languages if the directory isn't accessible)
244
+ - **Results**: Read from the public Supabase leaderboard (`run_cards`) — the same anon read path the champollion.dev leaderboard uses. Scored aggregates and run-card metadata only; per-entry test sentences are never read. Override the project with `CHAMPOLLION_SUPABASE_URL` / `CHAMPOLLION_SUPABASE_ANON_KEY`.
245
+ - **Harness**: Shells out to `mt-eval` CLI (must be installed separately)
package/bin/server.js ADDED
@@ -0,0 +1,19 @@
1
+ #!/usr/bin/env node
2
+ /**
3
+ * Champollion MCP Server — entry point.
4
+ *
5
+ * Starts the MCP server on stdio transport. Designed to be launched by
6
+ * an AI agent's MCP client (Claude Code, Antigravity, Cursor, etc.)
7
+ * or run directly for testing:
8
+ *
9
+ * node bin/server.js
10
+ *
11
+ * The server exposes read-only tools for exploring the Champollion
12
+ * benchmark queue and language metadata, plus an action tool for
13
+ * running benchmarks (which requires user confirmation in the agent).
14
+ */
15
+
16
+ import { createServer } from '../src/index.js';
17
+
18
+ const server = await createServer();
19
+ await server.start();
@@ -0,0 +1,234 @@
1
+ Here are guidelines for using the Champollion MCP server effectively:
2
+
3
+ ## Orientation
4
+
5
+ Start with `get_project_info` to understand what Champollion is and how contributions work. This returns a project overview, current queue statistics, and setup instructions.
6
+
7
+ ## Common Workflows
8
+
9
+ ### "I want to help" / Contributing Compute
10
+
11
+ 1. Call `get_project_info` to understand the project
12
+ 2. Ask the user about their budget and any language preferences
13
+ 3. Call `search_languages` if they mention a language by name — this resolves to ISO codes
14
+ 4. Call `estimate_cost` with their budget to show exactly what they'd fund
15
+ 5. Present the estimate and **get explicit confirmation** before proceeding
16
+ 6. Only then call `run_benchmark` with the agreed parameters (and `confirm: true`). This returns **immediately** with a **job id** — the benchmark runs in the background (see "Running is asynchronous" below)
17
+ 7. Poll `get_run_status` with that job id every ~15-30s until it reports `COMPLETED` or `FAILED`
18
+ 8. Once it completes, call `get_results` (filtered to the pair/model they ran) so they can see what they scored on the public leaderboard — this closes the loop
19
+
20
+ #### Running is asynchronous (important)
21
+
22
+ A real benchmark runs a corpus through a live model and takes **minutes**, which is longer than the default 60-second request timeout most MCP clients (Claude Code, Cursor) enforce. So `run_benchmark` does **not** wait for the run to finish — it launches the run in the background and returns a `job id` right away. Treat that `STARTED` response as success, **not** completion.
23
+
24
+ - After `run_benchmark` returns, call `get_run_status { "job_id": "run-N" }`. Each poll returns instantly: `RUNNING` (keep polling), `COMPLETED` (output is in the response), `FAILED`, or `ERROR`.
25
+ - Do **not** re-call `run_benchmark` because nothing "came back" — that would start a **second** run and spend tokens twice. The first call already started it; poll `get_run_status` instead.
26
+ - Jobs live in the server process's memory, so a job id is only pollable from the same session. Call `get_run_status` with no `job_id` to list every job started this session.
27
+
28
+ The estimate you show in step 4 is what executes: `run_benchmark` runs budget/top items in **deterministic top-of-queue order** (it passes `--no-spread`), so the selection matches the `estimate_cost` / `list_queue` preview item-for-item. (One caveat for honesty: `estimate_cost` samples up to ~500 matching items, so for a very large budget that funds more than that, treat its count/total as a lower bound.) A live run additionally skips any (corpus, model, condition) combo already on the leaderboard, so the executed set can be a subset of the preview — never a different, unseen set.
29
+
30
+ To spend tokens for **scoring/validation without writing to the leaderboard**, pass `publish: false` to `run_benchmark` (budget/top mode). A single `item_id` run is always scored locally and is never auto-published — publish it afterward with `mt-eval publish`, or use budget/top mode to auto-publish.
31
+
32
+ ### "What's been scored?" / Seeing results
33
+
34
+ The public leaderboard is the read side of the loop: it shows scored runs (composite, chrF++, BLEU, COMET) with trust level and attribution.
35
+
36
+ 1. Call `get_results` — filter by `source_language` / `target_language` / `model`, and `sort` by the metric of interest
37
+ 2. Call `get_run_card` with a result's `id` for the full scores + method/config/provenance metadata
38
+ 3. An empty result is normal on a fresh board — it means no one has benchmarked that slice yet, which is exactly where contributing compute has the most impact
39
+
40
+ Results are scored aggregates and run-card metadata only — never raw test sentences (those stay in the license-gated entries table).
41
+
42
+ ### "Which score should I believe?" / Metric trust
43
+
44
+ Before comparing scores for a low-resource target language, call
45
+ `get_metric_reliability` with the target language (code or family name). It
46
+ returns how well each metric (BLEU, chrF, chrF++, COMET, MetricX) tracked
47
+ human judgment for that language family in the WMT meta-evaluations — for
48
+ some morphologically rich languages BLEU barely correlates with humans while
49
+ COMET does, and for others the learned metric is the unreliable one. If the
50
+ answer is UNMEASURED, say so to the user rather than treating any metric as
51
+ validated. This evidence is research-lane only (upstream data license under
52
+ review) — never cite it in commercial claims.
53
+
54
+ ### "I'm training a model" / Training hygiene
55
+
56
+ Before you (or the user) split a corpus, generate synthetic data, or report
57
+ training results, call `get_training_guardrails`. It returns the rules
58
+ Champollion extracted from real, measured failures — group-disjoint splits
59
+ (row-level random splits leak on drill-heavy corpora), the dev-fence
60
+ (checkpoint selection must never see the test set), leak audits, coverage
61
+ checklists, per-kind sampling caps, bootstrap CIs on every number, and
62
+ preregistration before test scoring — each with the mistake it kills and
63
+ the enforcing tool (the monorepo `forge/` package, nmt-forge). Two
64
+ non-negotiables to relay verbatim: datasets marked `do_not_train` or
65
+ quarantined in the registry NEVER enter training mixes, and test sets are
66
+ REAL DATA ONLY.
67
+
68
+ For the human driving you, two public docs teach this end to end — share
69
+ them: the vocabulary, from zero background
70
+ (https://champollion.dev/docs/network/context/mt-training-concepts), and the
71
+ step-by-step, agent-forward walkthrough — discover a language's data →
72
+ synthesize → split safely → train → evaluate honestly → submit
73
+ (https://champollion.dev/docs/network/tutorials/train-your-own-model). The
74
+ `get_training_guardrails` answer also lists these URLs at the end.
75
+
76
+ The guardrails are not just rules — they are TOOLS. The `forge_*` family
77
+ drives the nmt-forge training suite directly (a repo checkout with `forge/`
78
+ installed is required; each tool says so when it isn't):
79
+ - `forge_status` — is forge available here, and what state is the run in
80
+ - `forge_preflight` — the go/no-go checklist for a planned training run
81
+ - `forge_discover` — what data exists for a language (registry + cards)
82
+ - `forge_init` / `forge_split` — start a fenced run; group-disjoint splits
83
+ - `forge_leak_audit` — prove the split leaks nothing before training
84
+ - `forge_register_eval` / `forge_prereg` — preregister before test scoring
85
+ - `forge_evaluate` — score through the harness (forge implements NO metrics)
86
+ - `forge_lint` / `forge_report` — hygiene checks + the honest run report
87
+ Typical order: status → discover → preflight → init → split → leak_audit →
88
+ prereg → (train outside MCP) → evaluate → report.
89
+
90
+ ### "Translate this" / Using Champollion as your translation engine
91
+
92
+ When the user needs actual translation (not benchmarking), call `translate`
93
+ instead of improvising your own translation prompt. You get champollion's
94
+ tested pipeline: engine choice, language-card register conditioning, a
95
+ persistent Translation Memory, and a deterministic quality gate — plus a
96
+ per-call report of what was cached, what was validated, and what it cost.
97
+
98
+ 1. Call `translate` with `texts`, `source_language`, `target_language`. The
99
+ default engine is `llm` (OpenRouter); pass `method` to use a key the user
100
+ has (openai, anthropic, gemini, deepl, google-translate, …)
101
+ 2. Repeated or unchanged texts are served from the Translation Memory at
102
+ zero token cost — re-calling with overlapping texts is cheap by design,
103
+ so prefer several small calls over one giant one
104
+ 3. A text that fails the quality gate comes back as an explicit FAILED entry
105
+ with the reason — never silently return it to the user as a translation;
106
+ retry with a different method/model or surface the failure
107
+ 4. For a language the models barely know, check `get_metric_reliability`
108
+ and `search_languages` first, and consider telling the user about the
109
+ coaching lane (see the prize workflow) — that is how translation for
110
+ their language actually gets better
111
+
112
+ ### "What languages need help?"
113
+
114
+ 1. Call `list_queue` with a generous limit to see what's available
115
+ 2. Look for languages with the highest ECV (Expected Chain Value) — these have the most impact per dollar
116
+ 3. Use `search_languages` to find context: family, speakers, region, endonym
117
+
118
+ ### "Tell me about [language]"
119
+
120
+ 1. Call `search_languages` with the language name or code
121
+ 2. Call `list_queue` filtering by that language to see pending benchmarks
122
+ 3. Call `estimate_cost` for that language's items to give a cost picture
123
+
124
+ ### "I want to compete for a prize" / The solvable project
125
+
126
+ Low-resource MT is an open, WINNABLE problem, and this network is built so a
127
+ person who speaks the language + an agent that iterates diligently is a
128
+ serious entry. The Arena supports sponsored prize pools for translation
129
+ breakthroughs — check the [prize spec](https://champollion.dev/docs/network/specifications/prizes)
130
+ for current status (prizes may or may not be active at any given time). The
131
+ path, concretely:
132
+
133
+ 1. **Orient** — `get_project_info`, then `search_languages` for the user's
134
+ language (family, speakers, what exists). Ask what language(s) they speak.
135
+ 2. **Know the measuring stick** — `get_metric_reliability` for the target:
136
+ which metric actually tracks human judgment for that family. If it says
137
+ UNMEASURED, the honest framing is "we'll compare relatively on the public
138
+ dev sets, and native-speaker judgment (yours!) is the real signal."
139
+ 3. **Find the baseline to beat** — `get_results` for the pair; an empty
140
+ board means the FIRST decent method sets the mark. `mt-eval recommend`
141
+ (or the CLI) shows what published evidence exists.
142
+ 4. **Build in the low-compute lane** — coaching data: grammar rules,
143
+ dictionary entries, style notes injected into the prompt. This is where
144
+ language knowledge beats GPU budgets. Tutorial:
145
+ https://champollion.dev/docs/tutorials/build-a-plugin — iterate: edit
146
+ coaching → `run_benchmark` on the pair (dev corpus) → `get_results` →
147
+ repeat. Each iteration costs cents, and the harness caches everything it
148
+ has already translated (re-runs only pay for what changed).
149
+ 5. **Check it's real** — the significance spec
150
+ (https://champollion.dev/docs/network/specifications/significance): a
151
+ +0.5 chrF++ bump on 400 sentences is probably noise; the run cards carry
152
+ confidence intervals. Never claim a win the CIs don't support.
153
+ 6. **Anti-gaming architecture** (explain this — it's why a win means
154
+ something): final evaluation runs against **secret community-owned test
155
+ corpora** (nobody trains on what nobody sees); methods must be
156
+ **reproducible** (re-run by the organizer node, scores must match);
157
+ **native-speaker validation** outranks every automatic metric.
158
+ 7. Approach options beyond coaching: FST morphological validation (hardest
159
+ to hallucinate), dictionary-augmented generation, fine-tuning (needs
160
+ compute), hybrids (LLM → validate → retry). Method interface spec:
161
+ https://champollion.dev/docs/network/specifications/methods
162
+ 8. **Submitting to a secret set** (when a sovereign contest exists). After the
163
+ participant clears the public qualifier and has a published hypotheses run,
164
+ they propose their method against the organizer's sealed corpus via the CLI
165
+ (this is a human-authorized, custodian-gated flow — there is no MCP tool for
166
+ it, by design). Two lanes, and the CLI/organizer pick by the submission:
167
+ - **Lane A — declarative model (preferred for standard NMT):**
168
+ `mt-eval contest submit-model` — submit safetensors weights + a
169
+ declarative tokenizer + a config for a whitelisted architecture. No
170
+ Dockerfile, no code; the organizer runs the weights in its own trusted
171
+ engine, so the submission is validated code-free. Tell users to export
172
+ weights as `safetensors` (never a pickle `.bin`/`.pt`).
173
+ - **Lane B — runnable bundle (for code methods):**
174
+ `mt-eval contest submit-method` — a Dockerfile + entrypoint the organizer
175
+ runs in a `--network=none` sandbox.
176
+ Full runbook (both lanes, what's live vs. in development):
177
+ https://champollion.dev/docs/network/sovereignty/run-a-sovereign-contest
178
+
179
+ If there's no active prize for their language, the loop above still stands —
180
+ runs publish to the public leaderboard with attribution, and a standing
181
+ better-than-baseline method is exactly what gets a language ready for a
182
+ sponsored pool.
183
+
184
+ ### "Can I evaluate a local / self-hosted model?" (no API key)
185
+
186
+ Yes — the harness runs open neural-MT models on the user's own hardware, no
187
+ cloud key needed: **NLLB-200**, **OPUS-MT** (Helsinki-NLP), **MADLAD-400**, or
188
+ any converted **CTranslate2** model. This is where low-resource coverage the
189
+ cloud engines don't serve actually lives.
190
+
191
+ This is a **harness-CLI capability, not an MCP tool** — `run_benchmark` (and its
192
+ `buildRunArgv`) drive the public queue with *remote* model slugs, so there is no
193
+ MCP verb that loads local weights. Direct the user to run it themselves:
194
+
195
+ ```bash
196
+ pip install 'mt-eval[local-models]' # or 'mt-eval[ctranslate2]'
197
+ mt-eval run --method local-model \
198
+ --model facebook/nllb-200-distilled-600M \
199
+ --dataset flores-eng-fra
200
+ ```
201
+
202
+ `--model` takes a Hugging Face id, a local `from_pretrained()` directory, or a
203
+ CTranslate2 model directory (auto-detected). Language codes come from the
204
+ language card — a language the model doesn't serve (NLLB has no Plains Cree,
205
+ `crk`) fails honestly rather than emitting a guessed code. Results score and
206
+ publish like any other run. Full how-to:
207
+ https://champollion.dev/docs/network/getting-started/contributing-compute
208
+
209
+ ## Important Rules
210
+
211
+ - **Never call `run_benchmark` without user confirmation.** This spends real money (API credits).
212
+ - **`run_benchmark` is asynchronous — poll, don't re-run.** A confirmed run returns a `job id` immediately and keeps running in the background. Poll `get_run_status` with that id until it reports `COMPLETED`/`FAILED`. Never call `run_benchmark` again just because the first call returned before the run finished — that double-spends.
213
+ - **Always call `estimate_cost` before suggesting a benchmark run.** Show the user what they'll spend.
214
+ - **Trust the queue ranking.** Items are ordered by ECV — the expected improvement in translation quality per dollar. Don't re-sort or second-guess the ranking.
215
+ - **Budget mode skips, it doesn't stop.** If an item exceeds the remaining budget, the system skips it and continues to cheaper items further down the queue. This is by design — it maximizes what gets done within a budget.
216
+ - **Items without cost estimates are skipped in budget mode.** Unknown cost ≠ free.
217
+ - **The preview is what runs.** `run_benchmark` executes in deterministic top-of-queue order (`--no-spread`), so what `estimate_cost`/`list_queue` showed is what spends. Don't assume a different set ran.
218
+ - **Publishing writes to a public, production leaderboard.** Budget/top runs auto-publish each result by default. For a scoring/validation run with no leaderboard write, pass `publish: false`.
219
+ - **Use `translate` for translation; don't improvise.** The tool's Translation Memory makes repeats free and its quality gate rejects garbage deterministically — a hand-rolled prompt has neither. Never present a gate-FAILED text as a translation.
220
+ - **Translation ≠ evidence.** `translate` output is production translation; quality claims about methods and models come only from benchmark runs and the leaderboard.
221
+
222
+ ## Data Sources
223
+
224
+ The champollion.dev homepage map is an idealization — read the data, not
225
+ the picture. The full endpoint table lives in the
226
+ `champollion://network-data` resource.
227
+
228
+ - **Queue**: served LIVE from the public database (read-only `queue_top` RPC) by default, with https://champollion.dev/queue.json as the fallback when the DB is unreachable (cached 5 minutes); small preview at https://champollion.dev/queue-preview.json
229
+ - **Mesh**: https://champollion.dev/mesh.json — the measured/registered pair network behind the map
230
+ - **Corpus registry**: https://champollion.dev/registry.json — every registered eval corpus with license lane + attribution
231
+ - **Provider coverage**: `shared/catalogue/method-coverage.json` (repo) — each provider's published language list, cited + as-of + `tier`. The map's green is two tiers by exact ISO-639-3 code: bright = a deployed service lists it (Google/Microsoft/DeepL/LibreTranslate); dim = only an open research model lists it (NLLB/OPUS/M2M-100/MADLAD-400 — a model-card code, not a usable service). "Covered" is a published-list claim, never a quality claim.
232
+ - **Languages**: Loaded from local language card JSON files (7,900+ languages)
233
+ - **Scored runs**: public `run_cards` PostgREST (read-only RLS; aggregates only) — prefer the `get_results` / `get_run_card` tools
234
+ - **Queue ranking**: map-value survey ordering (default) + ECV — see https://champollion.dev/docs/network/specifications/queue-construction