champollion-mcp-server 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +133 -0
- package/README.md +245 -0
- package/bin/server.js +19 -0
- package/instructions.md +234 -0
- package/package.json +50 -0
- package/src/index.js +1106 -0
- package/src/tools/forge.js +140 -0
- package/src/tools/harness.js +727 -0
- package/src/tools/languages.js +329 -0
- package/src/tools/queue.js +313 -0
- package/src/tools/reliability.js +190 -0
- package/src/tools/results.js +346 -0
- package/src/tools/training.js +349 -0
- package/src/tools/translate.js +385 -0
package/LICENSE
ADDED
|
@@ -0,0 +1,133 @@
|
|
|
1
|
+
# PolyForm Noncommercial License 1.0.0
|
|
2
|
+
|
|
3
|
+
<https://polyformproject.org/licenses/noncommercial/1.0.0>
|
|
4
|
+
|
|
5
|
+
## Acceptance
|
|
6
|
+
|
|
7
|
+
In order to get any license under these terms, you must agree
|
|
8
|
+
to them as both strict obligations and conditions to all
|
|
9
|
+
your licenses.
|
|
10
|
+
|
|
11
|
+
## Copyright License
|
|
12
|
+
|
|
13
|
+
The licensor grants you a copyright license for the
|
|
14
|
+
software to do everything you might do with the software
|
|
15
|
+
that would otherwise infringe the licensor's copyright
|
|
16
|
+
in it for any permitted purpose. However, you may
|
|
17
|
+
only distribute the software according to [Distribution
|
|
18
|
+
License](#distribution-license) and make changes or new works
|
|
19
|
+
based on the software according to [Changes and New Works
|
|
20
|
+
License](#changes-and-new-works-license).
|
|
21
|
+
|
|
22
|
+
## Distribution License
|
|
23
|
+
|
|
24
|
+
The licensor grants you an additional copyright license
|
|
25
|
+
to distribute copies of the software. Your license
|
|
26
|
+
to distribute covers distributing the software with
|
|
27
|
+
changes and new works permitted by [Changes and New Works
|
|
28
|
+
License](#changes-and-new-works-license).
|
|
29
|
+
|
|
30
|
+
## Notices
|
|
31
|
+
|
|
32
|
+
You must ensure that anyone who gets a copy of any part of
|
|
33
|
+
the software from you also gets a copy of these terms or the
|
|
34
|
+
URL for them above, as well as copies of any plain-text lines
|
|
35
|
+
beginning with `Required Notice:` that the licensor provided
|
|
36
|
+
with the software. For example:
|
|
37
|
+
|
|
38
|
+
> Required Notice: Copyright Yoyodyne, Inc. (http://example.com)
|
|
39
|
+
|
|
40
|
+
## Changes and New Works License
|
|
41
|
+
|
|
42
|
+
The licensor grants you an additional copyright license to
|
|
43
|
+
make changes and new works based on the software for any
|
|
44
|
+
permitted purpose.
|
|
45
|
+
|
|
46
|
+
## Patent License
|
|
47
|
+
|
|
48
|
+
The licensor grants you a patent license for the software that
|
|
49
|
+
covers patent claims the licensor can license, or becomes able
|
|
50
|
+
to license, that you would infringe by using the software.
|
|
51
|
+
|
|
52
|
+
## Noncommercial Purposes
|
|
53
|
+
|
|
54
|
+
Any noncommercial purpose is a permitted purpose.
|
|
55
|
+
|
|
56
|
+
## Personal Uses
|
|
57
|
+
|
|
58
|
+
Personal use for research, experiment, and testing for
|
|
59
|
+
the benefit of public knowledge, personal study, private
|
|
60
|
+
entertainment, hobby projects, amateur pursuits, or religious
|
|
61
|
+
observance, without any anticipated commercial application,
|
|
62
|
+
is use for a permitted purpose.
|
|
63
|
+
|
|
64
|
+
## Noncommercial Organizations
|
|
65
|
+
|
|
66
|
+
Use by any charitable organization, educational institution,
|
|
67
|
+
public research organization, public safety or health
|
|
68
|
+
organization, environmental protection organization,
|
|
69
|
+
or government institution is use for a permitted purpose
|
|
70
|
+
regardless of the source of funding or obligations resulting
|
|
71
|
+
from the funding.
|
|
72
|
+
|
|
73
|
+
## Fair Use
|
|
74
|
+
|
|
75
|
+
You may have "fair use" rights for the software under the
|
|
76
|
+
law. These terms do not limit them.
|
|
77
|
+
|
|
78
|
+
## No Other Rights
|
|
79
|
+
|
|
80
|
+
These terms do not allow you to sublicense or transfer any of
|
|
81
|
+
your licenses to anyone else, or prevent the licensor from
|
|
82
|
+
granting licenses to anyone else. These terms do not imply
|
|
83
|
+
any other licenses.
|
|
84
|
+
|
|
85
|
+
## Patent Defense
|
|
86
|
+
|
|
87
|
+
If you make any written claim that the software infringes or
|
|
88
|
+
contributes to infringement of any patent, your patent license
|
|
89
|
+
for the software granted under these terms ends immediately. If
|
|
90
|
+
your company makes such a claim, your patent license ends
|
|
91
|
+
immediately for work on behalf of your company.
|
|
92
|
+
|
|
93
|
+
## Violations
|
|
94
|
+
|
|
95
|
+
The first time you are notified in writing that you have
|
|
96
|
+
violated any of these terms, or done anything with the software
|
|
97
|
+
not covered by your licenses, your licenses can nonetheless
|
|
98
|
+
continue if you come into full compliance with these terms,
|
|
99
|
+
and take practical steps to correct past violations, within
|
|
100
|
+
32 days of receiving notice. Otherwise, all your licenses
|
|
101
|
+
end immediately.
|
|
102
|
+
|
|
103
|
+
## No Liability
|
|
104
|
+
|
|
105
|
+
***As far as the law allows, the software comes as is, without
|
|
106
|
+
any warranty or condition, and the licensor will not be liable
|
|
107
|
+
to you for any damages arising out of these terms or the use
|
|
108
|
+
or nature of the software, under any kind of legal claim.***
|
|
109
|
+
|
|
110
|
+
## Definitions
|
|
111
|
+
|
|
112
|
+
The **licensor** is the individual or entity offering these
|
|
113
|
+
terms, and the **software** is the software the licensor makes
|
|
114
|
+
available under these terms.
|
|
115
|
+
|
|
116
|
+
**You** refers to the individual or entity agreeing to these
|
|
117
|
+
terms.
|
|
118
|
+
|
|
119
|
+
**Your company** is any legal entity, sole proprietorship,
|
|
120
|
+
or other kind of organization that you work for, plus all
|
|
121
|
+
organizations that have control over, are under the control of,
|
|
122
|
+
or are under common control with that organization. **Control**
|
|
123
|
+
means ownership of substantially all the assets of an entity,
|
|
124
|
+
or the power to direct its management and policies by vote,
|
|
125
|
+
contract, or otherwise. Control can be direct or indirect.
|
|
126
|
+
|
|
127
|
+
**Your licenses** are all the licenses granted to you for the
|
|
128
|
+
software under these terms.
|
|
129
|
+
|
|
130
|
+
**Use** means anything you do with the software requiring one
|
|
131
|
+
of your licenses.
|
|
132
|
+
|
|
133
|
+
Required Notice: Copyright Curtis Forbes — Champollion (https://champollion.dev)
|
package/README.md
ADDED
|
@@ -0,0 +1,245 @@
|
|
|
1
|
+
# champollion-mcp-server
|
|
2
|
+
|
|
3
|
+
MCP (Model Context Protocol) server for Champollion. Lets AI agents browse the public benchmark queue, search language metadata, and run `mt-eval` benchmarks — all through natural conversation.
|
|
4
|
+
|
|
5
|
+
> **Champollion** is infrastructure for trustworthy machine translation across every language — source-available and free for noncommercial use (the evaluation harness and shared registries are open source) — the test sets and the map that show who can translate what, how good each method is, and where the gaps are. Public benchmarks on open data rank every method (human and machine); sovereign benchmarks are secret community-owned test sets we never see. The infrastructure is source-available and singly stewarded; the test sets and the methods for a community's language belong to that community — built with communities, never scraped from them. This server is the agent-facing door into that network ([champollion.dev/docs/network](https://champollion.dev/docs/network/)). This server itself is PolyForm Noncommercial 1.0.0 (see [LICENSE](LICENSE)).
|
|
6
|
+
|
|
7
|
+
## What it does
|
|
8
|
+
|
|
9
|
+
When connected to an agent (Claude Code, Antigravity, Cursor, etc.), the server exposes tools, resources, and prompts:
|
|
10
|
+
|
|
11
|
+
### Tools
|
|
12
|
+
|
|
13
|
+
| Tool | Type | Description |
|
|
14
|
+
|---|---|---|
|
|
15
|
+
| `list_queue` | Read-only | Browse open benchmark items, filter by language/model/budget |
|
|
16
|
+
| `get_queue_item` | Read-only | Get full details for a specific queue item |
|
|
17
|
+
| `estimate_cost` | Read-only | Estimate cost for a set of benchmark runs |
|
|
18
|
+
| `search_languages` | Read-only | Search language cards by name, code, family, or region |
|
|
19
|
+
| `get_project_info` | Read-only | Get a Champollion project overview |
|
|
20
|
+
| `get_results` | Read-only | Read scored runs from the public leaderboard (closes the run → see-impact loop) |
|
|
21
|
+
| `get_run_card` | Read-only | Get one run card (scores + method/config metadata) by id |
|
|
22
|
+
| `get_metric_reliability` | Read-only | Which metric to TRUST for a target language — correlations with WMT human judgments, per language family ([methodology](https://champollion.dev/docs/network/specifications/metric-reliability)) |
|
|
23
|
+
| `get_training_guardrails` | Read-only | How to train an NMT model without fooling yourself — the guardrail rules (group-disjoint splits, dev-fence, leak audits, preregistration, …) extracted from real measured failures, each naming its enforcing tool in `forge/` (nmt-forge) |
|
|
24
|
+
| `translate` | Action | Translate texts through champollion's tested pipeline — engine choice, register conditioning, persistent Translation Memory (repeats are free), deterministic quality gate. Spends API tokens only on cache misses |
|
|
25
|
+
| `run_benchmark` | Action | Start benchmarks via the mt-eval harness — launches in the background and returns a job id immediately |
|
|
26
|
+
| `get_run_status` | Read-only | Poll a benchmark job by id until it completes (the run continues past the client's 60s timeout) |
|
|
27
|
+
|
|
28
|
+
#### Training tools (nmt-forge)
|
|
29
|
+
|
|
30
|
+
These wrap the [nmt-forge](https://github.com/gamedaysuits/Champollion) training suite. forge is part of the Champollion monorepo (not on PyPI) — clone the repo and set `CHAMPOLLION_FORGE_DIR` to its `forge/` directory; scoring additionally needs the eval harness (`pip install mt-eval`). Without forge present these tools return an actionable error rather than crashing.
|
|
31
|
+
|
|
32
|
+
| Tool | Type | Description |
|
|
33
|
+
|---|---|---|
|
|
34
|
+
| `forge_status` | Read-only | Where am I in an nmt-forge project and what do I run next — call first and after every step |
|
|
35
|
+
| `forge_preflight` | Read-only | Will this command refuse? Renders every gate it will hit (✓/✗ with the fix for each ✗) |
|
|
36
|
+
| `forge_discover` | Read-only | What a language HAS — reads the SSOT language card (scripts, analyzers, dictionaries, corpora) |
|
|
37
|
+
| `forge_init` | Action | Scaffold a forge project from a language card: workspace + starter config + NEXT_STEPS brief |
|
|
38
|
+
| `forge_split` | Action | Carve a parallel corpus into GROUP-DISJOINT train/dev/test (shared-source/target pairs stay together) |
|
|
39
|
+
| `forge_leak_audit` | Read-only | Screen a corpus against every registered eval set BEFORE training (exact/near-dupe detection) |
|
|
40
|
+
| `forge_register_eval` | Action | Register an eval file in the workspace with a role: dev (fenced selection) or test (prereg-gated) |
|
|
41
|
+
| `forge_prereg` | Action | Preregister falsifiable predictions for a test/sealed set BEFORE scoring it |
|
|
42
|
+
| `forge_evaluate` | Action | Close the loop: decode the config's battery with the selected checkpoint, score via the mt-eval harness (forge implements zero metrics itself) |
|
|
43
|
+
| `forge_lint` | Read-only | Diagnose a battery manifest: weak registers and the likeliest cause given co-occurring signals |
|
|
44
|
+
| `forge_report` | Read-only | Re-render the plain-language training report (with the Diagnosis section) from a manifest |
|
|
45
|
+
|
|
46
|
+
### Resources (read-only data)
|
|
47
|
+
|
|
48
|
+
| Resource | URI | Description |
|
|
49
|
+
|---|---|---|
|
|
50
|
+
| Contributing guide | `champollion://contributing-guide` | CONTRIBUTING.md — how to help with the project |
|
|
51
|
+
| Queue schema | `champollion://queue-schema` | Field definitions for every queue.json item |
|
|
52
|
+
| Network data map | `champollion://network-data` | Which machine artifact answers which question (queue, mesh, registry, coverage) |
|
|
53
|
+
|
|
54
|
+
### Prompts (conversation starters)
|
|
55
|
+
|
|
56
|
+
| Prompt | Arguments | Description |
|
|
57
|
+
|---|---|---|
|
|
58
|
+
| `contribute_compute` | `budget?`, `language?` | "I want to help — what would $X buy?" |
|
|
59
|
+
| `compete_for_prize` | `language?` | "I want to build a competitive method — any prizes?" |
|
|
60
|
+
| `explore_language` | `language` | "Tell me about [language] in Champollion" |
|
|
61
|
+
|
|
62
|
+
## Quick start
|
|
63
|
+
|
|
64
|
+
**From the published package** (no clone needed):
|
|
65
|
+
|
|
66
|
+
```bash
|
|
67
|
+
npx champollion-mcp-server # start on stdio (for agent connection)
|
|
68
|
+
|
|
69
|
+
```
|
|
70
|
+
|
|
71
|
+
**From source:**
|
|
72
|
+
|
|
73
|
+
```bash
|
|
74
|
+
cd mcp-server
|
|
75
|
+
npm install
|
|
76
|
+
npm test # run unit tests
|
|
77
|
+
npm start # start on stdio (for agent connection)
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
## Connect to your agent
|
|
81
|
+
|
|
82
|
+
### Claude Code / Antigravity
|
|
83
|
+
|
|
84
|
+
Add to your MCP configuration — published package:
|
|
85
|
+
|
|
86
|
+
```json
|
|
87
|
+
{
|
|
88
|
+
"mcpServers": {
|
|
89
|
+
"champollion": {
|
|
90
|
+
"command": "npx",
|
|
91
|
+
"args": ["-y", "champollion-mcp-server"]
|
|
92
|
+
}
|
|
93
|
+
}
|
|
94
|
+
}
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
Or from a local checkout:
|
|
98
|
+
|
|
99
|
+
```json
|
|
100
|
+
{
|
|
101
|
+
"mcpServers": {
|
|
102
|
+
"champollion": {
|
|
103
|
+
"command": "node",
|
|
104
|
+
"args": ["/path/to/Champollion/mcp-server/bin/server.js"]
|
|
105
|
+
}
|
|
106
|
+
}
|
|
107
|
+
}
|
|
108
|
+
```
|
|
109
|
+
|
|
110
|
+
### Cursor
|
|
111
|
+
|
|
112
|
+
Add to `.cursor/mcp.json` (same two options):
|
|
113
|
+
|
|
114
|
+
```json
|
|
115
|
+
{
|
|
116
|
+
"mcpServers": {
|
|
117
|
+
"champollion": {
|
|
118
|
+
"command": "npx",
|
|
119
|
+
"args": ["-y", "champollion-mcp-server"]
|
|
120
|
+
}
|
|
121
|
+
}
|
|
122
|
+
}
|
|
123
|
+
```
|
|
124
|
+
|
|
125
|
+
## What a conversation looks like
|
|
126
|
+
|
|
127
|
+
Once connected, you can talk to your agent naturally:
|
|
128
|
+
|
|
129
|
+
> **You:** "I want to help with Champollion — can you devote $10 in API credits to it?"
|
|
130
|
+
>
|
|
131
|
+
> **Agent** uses `get_project_info` → learns about the project
|
|
132
|
+
>
|
|
133
|
+
> **Agent** uses `list_queue` with `budget: 10` → sees what's available
|
|
134
|
+
>
|
|
135
|
+
> **Agent:** "I found a few thousand open benchmark items. Your $10 could fund dozens of runs. Any preference on languages?"
|
|
136
|
+
>
|
|
137
|
+
> **You:** "West African languages"
|
|
138
|
+
>
|
|
139
|
+
> **Agent** uses `list_queue` with `language: "african"` → filters results
|
|
140
|
+
>
|
|
141
|
+
> **Agent** uses `estimate_cost` → calculates the plan
|
|
142
|
+
>
|
|
143
|
+
> **Agent:** "I found 18 items for Yoruba, Hausa, Igbo, Zulu, Xhosa, and Luganda. Total: ~$1.64. Ready to run?"
|
|
144
|
+
>
|
|
145
|
+
> **You:** "Go for it"
|
|
146
|
+
>
|
|
147
|
+
> **Agent** uses `run_benchmark` with `budget: 10` → gets a **job id** back immediately (the run continues in the background)
|
|
148
|
+
>
|
|
149
|
+
> **Agent** polls `get_run_status` with that job id until it reports `COMPLETED`, then uses `get_results` to show what was scored
|
|
150
|
+
|
|
151
|
+
## Testing
|
|
152
|
+
|
|
153
|
+
```bash
|
|
154
|
+
npm test
|
|
155
|
+
```
|
|
156
|
+
|
|
157
|
+
Tests use mock data and don't make network calls. To test the server interactively:
|
|
158
|
+
|
|
159
|
+
```bash
|
|
160
|
+
npx @modelcontextprotocol/inspector node bin/server.js
|
|
161
|
+
```
|
|
162
|
+
|
|
163
|
+
## Architecture
|
|
164
|
+
|
|
165
|
+
```
|
|
166
|
+
mcp-server/
|
|
167
|
+
├── bin/server.js Entry point (stdio transport)
|
|
168
|
+
├── instructions.md Agent behavioral guide (loaded at connect time)
|
|
169
|
+
├── src/
|
|
170
|
+
│ ├── index.js Server setup + tool/resource/prompt registration
|
|
171
|
+
│ └── tools/
|
|
172
|
+
│ ├── queue.js Queue fetch, filter, cost estimation
|
|
173
|
+
│ ├── languages.js Language card index + search
|
|
174
|
+
│ ├── results.js Public leaderboard reads (scored run_cards)
|
|
175
|
+
│ ├── reliability.js Metric-reliability lookups (which metric to trust)
|
|
176
|
+
│ ├── training.js Training guardrails (get_training_guardrails)
|
|
177
|
+
│ ├── translate.js Champollion translate pipeline wrapper
|
|
178
|
+
│ └── harness.js mt-eval CLI wrapper
|
|
179
|
+
├── test/
|
|
180
|
+
│ ├── tools.test.js Unit tests (node --test) + SSOT shared vectors
|
|
181
|
+
│ ├── harness.test.js Unit tests for run_benchmark + the async job model
|
|
182
|
+
│ ├── results.test.js Unit tests for the leaderboard read tools
|
|
183
|
+
│ ├── queue-fetch.test.js Unit tests for queue fetching/caching
|
|
184
|
+
│ ├── reliability.test.js Unit tests for metric-reliability lookups
|
|
185
|
+
│ ├── training.test.js Unit tests for the training-guardrails tool
|
|
186
|
+
│ └── translate.test.js Unit tests for the translate tool
|
|
187
|
+
├── package.json
|
|
188
|
+
└── README.md
|
|
189
|
+
```
|
|
190
|
+
|
|
191
|
+
## Protocol version — and the 2026-07-28 stateless spec
|
|
192
|
+
|
|
193
|
+
**Transport: stdio only.** One server process per agent, launched by the client.
|
|
194
|
+
|
|
195
|
+
MCP's [2026-07-28 revision](https://blog.modelcontextprotocol.io/posts/2026-07-28/)
|
|
196
|
+
made the protocol **stateless by default** — the largest change since
|
|
197
|
+
authorization. It retires the `initialize`/`initialized` handshake and the
|
|
198
|
+
`Mcp-Session-Id` header, requires `Mcp-Method`/`Mcp-Name` HTTP headers so
|
|
199
|
+
gateways can route without parsing bodies, replaces held-open bidirectional
|
|
200
|
+
streams with Multi Round-Trip Requests, and deprecates Roots, Sampling, Logging
|
|
201
|
+
and the legacy HTTP+SSE transport (twelve-month support window).
|
|
202
|
+
|
|
203
|
+
**Where this server stands (checked 2026-08-01):**
|
|
204
|
+
|
|
205
|
+
| | Status |
|
|
206
|
+
|---|---|
|
|
207
|
+
| Deprecated capabilities (Roots / Sampling / Logging) | **None used.** |
|
|
208
|
+
| Legacy HTTP+SSE transport | **Not used** — stdio only. |
|
|
209
|
+
| `Mcp-Session-Id`, header routing, MRTR | **Not applicable** to stdio. |
|
|
210
|
+
| Application state across calls | **Already uses the prescribed pattern** — see below. |
|
|
211
|
+
| SDK support for `2026-07-28` | **Not yet available.** |
|
|
212
|
+
|
|
213
|
+
The new spec's guidance for cross-call state is to "mint an explicit handle from
|
|
214
|
+
a tool and have the model pass it back as an argument" rather than lean on
|
|
215
|
+
transport sessions. `run_benchmark` already works exactly that way: it returns a
|
|
216
|
+
job id, and the agent passes that id to `get_run_status`. No transport-level
|
|
217
|
+
session is ever involved.
|
|
218
|
+
|
|
219
|
+
**One assumption to know about.** The job registry in `src/tools/harness.js` is
|
|
220
|
+
in-memory and assumes a single server process for the agent's lifetime. That
|
|
221
|
+
holds for stdio. It would **not** hold behind a stateless HTTP deployment with
|
|
222
|
+
more than one process, where a poll could land on a process that never started
|
|
223
|
+
the job. Anyone adding an HTTP transport must move that registry to shared
|
|
224
|
+
storage first.
|
|
225
|
+
|
|
226
|
+
**Why we have not upgraded.** The published TypeScript SDK does not speak the
|
|
227
|
+
new revision yet: `@modelcontextprotocol/sdk@1.30.0` is the only dist-tag on
|
|
228
|
+
npm and its `LATEST_PROTOCOL_VERSION` is `2025-11-25`. The dependency floor here
|
|
229
|
+
was raised to `^1.30.0` (from a stale `^1.12.1`, eighteen releases behind) so
|
|
230
|
+
installs resolve current. Re-check when a `2026-07-28`-capable SDK publishes;
|
|
231
|
+
the migration should be small given the table above.
|
|
232
|
+
|
|
233
|
+
## Data sources
|
|
234
|
+
|
|
235
|
+
The champollion.dev homepage map is an idealization of this data — agents
|
|
236
|
+
should read the sources, not the picture (the `champollion://network-data`
|
|
237
|
+
resource carries the full endpoint table).
|
|
238
|
+
|
|
239
|
+
- **Queue**: Fetched from `champollion.dev/queue.json` (tens of MB — it grows with coverage; cached 5 min in memory). Small slice: `champollion.dev/queue-preview.json`. For the live open-item count, call `get_project_info`
|
|
240
|
+
- **Mesh**: `champollion.dev/mesh.json` — the measured/registered pair network behind the homepage map
|
|
241
|
+
- **Corpus registry**: `champollion.dev/registry.json` — every registered eval corpus with license lane, attribution, checksum
|
|
242
|
+
- **Provider coverage**: `shared/catalogue/method-coverage.json` — each provider's published language list, cited + as-of + `tier` (the data behind the map's covered/uncovered split). The map's green has two tiers by exact ISO-639-3 code: bright = a deployed service lists it (Google/Microsoft/DeepL/LibreTranslate); dim = only an open research model lists it (NLLB/OPUS/M2M-100/MADLAD-400 — a model-card code, not a usable service)
|
|
243
|
+
- **Languages**: Loaded from `cli/shared/language-cards/` on startup (falls back to built-in index of 40 languages if the directory isn't accessible)
|
|
244
|
+
- **Results**: Read from the public Supabase leaderboard (`run_cards`) — the same anon read path the champollion.dev leaderboard uses. Scored aggregates and run-card metadata only; per-entry test sentences are never read. Override the project with `CHAMPOLLION_SUPABASE_URL` / `CHAMPOLLION_SUPABASE_ANON_KEY`.
|
|
245
|
+
- **Harness**: Shells out to `mt-eval` CLI (must be installed separately)
|
package/bin/server.js
ADDED
|
@@ -0,0 +1,19 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
/**
|
|
3
|
+
* Champollion MCP Server — entry point.
|
|
4
|
+
*
|
|
5
|
+
* Starts the MCP server on stdio transport. Designed to be launched by
|
|
6
|
+
* an AI agent's MCP client (Claude Code, Antigravity, Cursor, etc.)
|
|
7
|
+
* or run directly for testing:
|
|
8
|
+
*
|
|
9
|
+
* node bin/server.js
|
|
10
|
+
*
|
|
11
|
+
* The server exposes read-only tools for exploring the Champollion
|
|
12
|
+
* benchmark queue and language metadata, plus an action tool for
|
|
13
|
+
* running benchmarks (which requires user confirmation in the agent).
|
|
14
|
+
*/
|
|
15
|
+
|
|
16
|
+
import { createServer } from '../src/index.js';
|
|
17
|
+
|
|
18
|
+
const server = await createServer();
|
|
19
|
+
await server.start();
|
package/instructions.md
ADDED
|
@@ -0,0 +1,234 @@
|
|
|
1
|
+
Here are guidelines for using the Champollion MCP server effectively:
|
|
2
|
+
|
|
3
|
+
## Orientation
|
|
4
|
+
|
|
5
|
+
Start with `get_project_info` to understand what Champollion is and how contributions work. This returns a project overview, current queue statistics, and setup instructions.
|
|
6
|
+
|
|
7
|
+
## Common Workflows
|
|
8
|
+
|
|
9
|
+
### "I want to help" / Contributing Compute
|
|
10
|
+
|
|
11
|
+
1. Call `get_project_info` to understand the project
|
|
12
|
+
2. Ask the user about their budget and any language preferences
|
|
13
|
+
3. Call `search_languages` if they mention a language by name — this resolves to ISO codes
|
|
14
|
+
4. Call `estimate_cost` with their budget to show exactly what they'd fund
|
|
15
|
+
5. Present the estimate and **get explicit confirmation** before proceeding
|
|
16
|
+
6. Only then call `run_benchmark` with the agreed parameters (and `confirm: true`). This returns **immediately** with a **job id** — the benchmark runs in the background (see "Running is asynchronous" below)
|
|
17
|
+
7. Poll `get_run_status` with that job id every ~15-30s until it reports `COMPLETED` or `FAILED`
|
|
18
|
+
8. Once it completes, call `get_results` (filtered to the pair/model they ran) so they can see what they scored on the public leaderboard — this closes the loop
|
|
19
|
+
|
|
20
|
+
#### Running is asynchronous (important)
|
|
21
|
+
|
|
22
|
+
A real benchmark runs a corpus through a live model and takes **minutes**, which is longer than the default 60-second request timeout most MCP clients (Claude Code, Cursor) enforce. So `run_benchmark` does **not** wait for the run to finish — it launches the run in the background and returns a `job id` right away. Treat that `STARTED` response as success, **not** completion.
|
|
23
|
+
|
|
24
|
+
- After `run_benchmark` returns, call `get_run_status { "job_id": "run-N" }`. Each poll returns instantly: `RUNNING` (keep polling), `COMPLETED` (output is in the response), `FAILED`, or `ERROR`.
|
|
25
|
+
- Do **not** re-call `run_benchmark` because nothing "came back" — that would start a **second** run and spend tokens twice. The first call already started it; poll `get_run_status` instead.
|
|
26
|
+
- Jobs live in the server process's memory, so a job id is only pollable from the same session. Call `get_run_status` with no `job_id` to list every job started this session.
|
|
27
|
+
|
|
28
|
+
The estimate you show in step 4 is what executes: `run_benchmark` runs budget/top items in **deterministic top-of-queue order** (it passes `--no-spread`), so the selection matches the `estimate_cost` / `list_queue` preview item-for-item. (One caveat for honesty: `estimate_cost` samples up to ~500 matching items, so for a very large budget that funds more than that, treat its count/total as a lower bound.) A live run additionally skips any (corpus, model, condition) combo already on the leaderboard, so the executed set can be a subset of the preview — never a different, unseen set.
|
|
29
|
+
|
|
30
|
+
To spend tokens for **scoring/validation without writing to the leaderboard**, pass `publish: false` to `run_benchmark` (budget/top mode). A single `item_id` run is always scored locally and is never auto-published — publish it afterward with `mt-eval publish`, or use budget/top mode to auto-publish.
|
|
31
|
+
|
|
32
|
+
### "What's been scored?" / Seeing results
|
|
33
|
+
|
|
34
|
+
The public leaderboard is the read side of the loop: it shows scored runs (composite, chrF++, BLEU, COMET) with trust level and attribution.
|
|
35
|
+
|
|
36
|
+
1. Call `get_results` — filter by `source_language` / `target_language` / `model`, and `sort` by the metric of interest
|
|
37
|
+
2. Call `get_run_card` with a result's `id` for the full scores + method/config/provenance metadata
|
|
38
|
+
3. An empty result is normal on a fresh board — it means no one has benchmarked that slice yet, which is exactly where contributing compute has the most impact
|
|
39
|
+
|
|
40
|
+
Results are scored aggregates and run-card metadata only — never raw test sentences (those stay in the license-gated entries table).
|
|
41
|
+
|
|
42
|
+
### "Which score should I believe?" / Metric trust
|
|
43
|
+
|
|
44
|
+
Before comparing scores for a low-resource target language, call
|
|
45
|
+
`get_metric_reliability` with the target language (code or family name). It
|
|
46
|
+
returns how well each metric (BLEU, chrF, chrF++, COMET, MetricX) tracked
|
|
47
|
+
human judgment for that language family in the WMT meta-evaluations — for
|
|
48
|
+
some morphologically rich languages BLEU barely correlates with humans while
|
|
49
|
+
COMET does, and for others the learned metric is the unreliable one. If the
|
|
50
|
+
answer is UNMEASURED, say so to the user rather than treating any metric as
|
|
51
|
+
validated. This evidence is research-lane only (upstream data license under
|
|
52
|
+
review) — never cite it in commercial claims.
|
|
53
|
+
|
|
54
|
+
### "I'm training a model" / Training hygiene
|
|
55
|
+
|
|
56
|
+
Before you (or the user) split a corpus, generate synthetic data, or report
|
|
57
|
+
training results, call `get_training_guardrails`. It returns the rules
|
|
58
|
+
Champollion extracted from real, measured failures — group-disjoint splits
|
|
59
|
+
(row-level random splits leak on drill-heavy corpora), the dev-fence
|
|
60
|
+
(checkpoint selection must never see the test set), leak audits, coverage
|
|
61
|
+
checklists, per-kind sampling caps, bootstrap CIs on every number, and
|
|
62
|
+
preregistration before test scoring — each with the mistake it kills and
|
|
63
|
+
the enforcing tool (the monorepo `forge/` package, nmt-forge). Two
|
|
64
|
+
non-negotiables to relay verbatim: datasets marked `do_not_train` or
|
|
65
|
+
quarantined in the registry NEVER enter training mixes, and test sets are
|
|
66
|
+
REAL DATA ONLY.
|
|
67
|
+
|
|
68
|
+
For the human driving you, two public docs teach this end to end — share
|
|
69
|
+
them: the vocabulary, from zero background
|
|
70
|
+
(https://champollion.dev/docs/network/context/mt-training-concepts), and the
|
|
71
|
+
step-by-step, agent-forward walkthrough — discover a language's data →
|
|
72
|
+
synthesize → split safely → train → evaluate honestly → submit
|
|
73
|
+
(https://champollion.dev/docs/network/tutorials/train-your-own-model). The
|
|
74
|
+
`get_training_guardrails` answer also lists these URLs at the end.
|
|
75
|
+
|
|
76
|
+
The guardrails are not just rules — they are TOOLS. The `forge_*` family
|
|
77
|
+
drives the nmt-forge training suite directly (a repo checkout with `forge/`
|
|
78
|
+
installed is required; each tool says so when it isn't):
|
|
79
|
+
- `forge_status` — is forge available here, and what state is the run in
|
|
80
|
+
- `forge_preflight` — the go/no-go checklist for a planned training run
|
|
81
|
+
- `forge_discover` — what data exists for a language (registry + cards)
|
|
82
|
+
- `forge_init` / `forge_split` — start a fenced run; group-disjoint splits
|
|
83
|
+
- `forge_leak_audit` — prove the split leaks nothing before training
|
|
84
|
+
- `forge_register_eval` / `forge_prereg` — preregister before test scoring
|
|
85
|
+
- `forge_evaluate` — score through the harness (forge implements NO metrics)
|
|
86
|
+
- `forge_lint` / `forge_report` — hygiene checks + the honest run report
|
|
87
|
+
Typical order: status → discover → preflight → init → split → leak_audit →
|
|
88
|
+
prereg → (train outside MCP) → evaluate → report.
|
|
89
|
+
|
|
90
|
+
### "Translate this" / Using Champollion as your translation engine
|
|
91
|
+
|
|
92
|
+
When the user needs actual translation (not benchmarking), call `translate`
|
|
93
|
+
instead of improvising your own translation prompt. You get champollion's
|
|
94
|
+
tested pipeline: engine choice, language-card register conditioning, a
|
|
95
|
+
persistent Translation Memory, and a deterministic quality gate — plus a
|
|
96
|
+
per-call report of what was cached, what was validated, and what it cost.
|
|
97
|
+
|
|
98
|
+
1. Call `translate` with `texts`, `source_language`, `target_language`. The
|
|
99
|
+
default engine is `llm` (OpenRouter); pass `method` to use a key the user
|
|
100
|
+
has (openai, anthropic, gemini, deepl, google-translate, …)
|
|
101
|
+
2. Repeated or unchanged texts are served from the Translation Memory at
|
|
102
|
+
zero token cost — re-calling with overlapping texts is cheap by design,
|
|
103
|
+
so prefer several small calls over one giant one
|
|
104
|
+
3. A text that fails the quality gate comes back as an explicit FAILED entry
|
|
105
|
+
with the reason — never silently return it to the user as a translation;
|
|
106
|
+
retry with a different method/model or surface the failure
|
|
107
|
+
4. For a language the models barely know, check `get_metric_reliability`
|
|
108
|
+
and `search_languages` first, and consider telling the user about the
|
|
109
|
+
coaching lane (see the prize workflow) — that is how translation for
|
|
110
|
+
their language actually gets better
|
|
111
|
+
|
|
112
|
+
### "What languages need help?"
|
|
113
|
+
|
|
114
|
+
1. Call `list_queue` with a generous limit to see what's available
|
|
115
|
+
2. Look for languages with the highest ECV (Expected Chain Value) — these have the most impact per dollar
|
|
116
|
+
3. Use `search_languages` to find context: family, speakers, region, endonym
|
|
117
|
+
|
|
118
|
+
### "Tell me about [language]"
|
|
119
|
+
|
|
120
|
+
1. Call `search_languages` with the language name or code
|
|
121
|
+
2. Call `list_queue` filtering by that language to see pending benchmarks
|
|
122
|
+
3. Call `estimate_cost` for that language's items to give a cost picture
|
|
123
|
+
|
|
124
|
+
### "I want to compete for a prize" / The solvable project
|
|
125
|
+
|
|
126
|
+
Low-resource MT is an open, WINNABLE problem, and this network is built so a
|
|
127
|
+
person who speaks the language + an agent that iterates diligently is a
|
|
128
|
+
serious entry. The Arena supports sponsored prize pools for translation
|
|
129
|
+
breakthroughs — check the [prize spec](https://champollion.dev/docs/network/specifications/prizes)
|
|
130
|
+
for current status (prizes may or may not be active at any given time). The
|
|
131
|
+
path, concretely:
|
|
132
|
+
|
|
133
|
+
1. **Orient** — `get_project_info`, then `search_languages` for the user's
|
|
134
|
+
language (family, speakers, what exists). Ask what language(s) they speak.
|
|
135
|
+
2. **Know the measuring stick** — `get_metric_reliability` for the target:
|
|
136
|
+
which metric actually tracks human judgment for that family. If it says
|
|
137
|
+
UNMEASURED, the honest framing is "we'll compare relatively on the public
|
|
138
|
+
dev sets, and native-speaker judgment (yours!) is the real signal."
|
|
139
|
+
3. **Find the baseline to beat** — `get_results` for the pair; an empty
|
|
140
|
+
board means the FIRST decent method sets the mark. `mt-eval recommend`
|
|
141
|
+
(or the CLI) shows what published evidence exists.
|
|
142
|
+
4. **Build in the low-compute lane** — coaching data: grammar rules,
|
|
143
|
+
dictionary entries, style notes injected into the prompt. This is where
|
|
144
|
+
language knowledge beats GPU budgets. Tutorial:
|
|
145
|
+
https://champollion.dev/docs/tutorials/build-a-plugin — iterate: edit
|
|
146
|
+
coaching → `run_benchmark` on the pair (dev corpus) → `get_results` →
|
|
147
|
+
repeat. Each iteration costs cents, and the harness caches everything it
|
|
148
|
+
has already translated (re-runs only pay for what changed).
|
|
149
|
+
5. **Check it's real** — the significance spec
|
|
150
|
+
(https://champollion.dev/docs/network/specifications/significance): a
|
|
151
|
+
+0.5 chrF++ bump on 400 sentences is probably noise; the run cards carry
|
|
152
|
+
confidence intervals. Never claim a win the CIs don't support.
|
|
153
|
+
6. **Anti-gaming architecture** (explain this — it's why a win means
|
|
154
|
+
something): final evaluation runs against **secret community-owned test
|
|
155
|
+
corpora** (nobody trains on what nobody sees); methods must be
|
|
156
|
+
**reproducible** (re-run by the organizer node, scores must match);
|
|
157
|
+
**native-speaker validation** outranks every automatic metric.
|
|
158
|
+
7. Approach options beyond coaching: FST morphological validation (hardest
|
|
159
|
+
to hallucinate), dictionary-augmented generation, fine-tuning (needs
|
|
160
|
+
compute), hybrids (LLM → validate → retry). Method interface spec:
|
|
161
|
+
https://champollion.dev/docs/network/specifications/methods
|
|
162
|
+
8. **Submitting to a secret set** (when a sovereign contest exists). After the
|
|
163
|
+
participant clears the public qualifier and has a published hypotheses run,
|
|
164
|
+
they propose their method against the organizer's sealed corpus via the CLI
|
|
165
|
+
(this is a human-authorized, custodian-gated flow — there is no MCP tool for
|
|
166
|
+
it, by design). Two lanes, and the CLI/organizer pick by the submission:
|
|
167
|
+
- **Lane A — declarative model (preferred for standard NMT):**
|
|
168
|
+
`mt-eval contest submit-model` — submit safetensors weights + a
|
|
169
|
+
declarative tokenizer + a config for a whitelisted architecture. No
|
|
170
|
+
Dockerfile, no code; the organizer runs the weights in its own trusted
|
|
171
|
+
engine, so the submission is validated code-free. Tell users to export
|
|
172
|
+
weights as `safetensors` (never a pickle `.bin`/`.pt`).
|
|
173
|
+
- **Lane B — runnable bundle (for code methods):**
|
|
174
|
+
`mt-eval contest submit-method` — a Dockerfile + entrypoint the organizer
|
|
175
|
+
runs in a `--network=none` sandbox.
|
|
176
|
+
Full runbook (both lanes, what's live vs. in development):
|
|
177
|
+
https://champollion.dev/docs/network/sovereignty/run-a-sovereign-contest
|
|
178
|
+
|
|
179
|
+
If there's no active prize for their language, the loop above still stands —
|
|
180
|
+
runs publish to the public leaderboard with attribution, and a standing
|
|
181
|
+
better-than-baseline method is exactly what gets a language ready for a
|
|
182
|
+
sponsored pool.
|
|
183
|
+
|
|
184
|
+
### "Can I evaluate a local / self-hosted model?" (no API key)
|
|
185
|
+
|
|
186
|
+
Yes — the harness runs open neural-MT models on the user's own hardware, no
|
|
187
|
+
cloud key needed: **NLLB-200**, **OPUS-MT** (Helsinki-NLP), **MADLAD-400**, or
|
|
188
|
+
any converted **CTranslate2** model. This is where low-resource coverage the
|
|
189
|
+
cloud engines don't serve actually lives.
|
|
190
|
+
|
|
191
|
+
This is a **harness-CLI capability, not an MCP tool** — `run_benchmark` (and its
|
|
192
|
+
`buildRunArgv`) drive the public queue with *remote* model slugs, so there is no
|
|
193
|
+
MCP verb that loads local weights. Direct the user to run it themselves:
|
|
194
|
+
|
|
195
|
+
```bash
|
|
196
|
+
pip install 'mt-eval[local-models]' # or 'mt-eval[ctranslate2]'
|
|
197
|
+
mt-eval run --method local-model \
|
|
198
|
+
--model facebook/nllb-200-distilled-600M \
|
|
199
|
+
--dataset flores-eng-fra
|
|
200
|
+
```
|
|
201
|
+
|
|
202
|
+
`--model` takes a Hugging Face id, a local `from_pretrained()` directory, or a
|
|
203
|
+
CTranslate2 model directory (auto-detected). Language codes come from the
|
|
204
|
+
language card — a language the model doesn't serve (NLLB has no Plains Cree,
|
|
205
|
+
`crk`) fails honestly rather than emitting a guessed code. Results score and
|
|
206
|
+
publish like any other run. Full how-to:
|
|
207
|
+
https://champollion.dev/docs/network/getting-started/contributing-compute
|
|
208
|
+
|
|
209
|
+
## Important Rules
|
|
210
|
+
|
|
211
|
+
- **Never call `run_benchmark` without user confirmation.** This spends real money (API credits).
|
|
212
|
+
- **`run_benchmark` is asynchronous — poll, don't re-run.** A confirmed run returns a `job id` immediately and keeps running in the background. Poll `get_run_status` with that id until it reports `COMPLETED`/`FAILED`. Never call `run_benchmark` again just because the first call returned before the run finished — that double-spends.
|
|
213
|
+
- **Always call `estimate_cost` before suggesting a benchmark run.** Show the user what they'll spend.
|
|
214
|
+
- **Trust the queue ranking.** Items are ordered by ECV — the expected improvement in translation quality per dollar. Don't re-sort or second-guess the ranking.
|
|
215
|
+
- **Budget mode skips, it doesn't stop.** If an item exceeds the remaining budget, the system skips it and continues to cheaper items further down the queue. This is by design — it maximizes what gets done within a budget.
|
|
216
|
+
- **Items without cost estimates are skipped in budget mode.** Unknown cost ≠ free.
|
|
217
|
+
- **The preview is what runs.** `run_benchmark` executes in deterministic top-of-queue order (`--no-spread`), so what `estimate_cost`/`list_queue` showed is what spends. Don't assume a different set ran.
|
|
218
|
+
- **Publishing writes to a public, production leaderboard.** Budget/top runs auto-publish each result by default. For a scoring/validation run with no leaderboard write, pass `publish: false`.
|
|
219
|
+
- **Use `translate` for translation; don't improvise.** The tool's Translation Memory makes repeats free and its quality gate rejects garbage deterministically — a hand-rolled prompt has neither. Never present a gate-FAILED text as a translation.
|
|
220
|
+
- **Translation ≠ evidence.** `translate` output is production translation; quality claims about methods and models come only from benchmark runs and the leaderboard.
|
|
221
|
+
|
|
222
|
+
## Data Sources
|
|
223
|
+
|
|
224
|
+
The champollion.dev homepage map is an idealization — read the data, not
|
|
225
|
+
the picture. The full endpoint table lives in the
|
|
226
|
+
`champollion://network-data` resource.
|
|
227
|
+
|
|
228
|
+
- **Queue**: served LIVE from the public database (read-only `queue_top` RPC) by default, with https://champollion.dev/queue.json as the fallback when the DB is unreachable (cached 5 minutes); small preview at https://champollion.dev/queue-preview.json
|
|
229
|
+
- **Mesh**: https://champollion.dev/mesh.json — the measured/registered pair network behind the map
|
|
230
|
+
- **Corpus registry**: https://champollion.dev/registry.json — every registered eval corpus with license lane + attribution
|
|
231
|
+
- **Provider coverage**: `shared/catalogue/method-coverage.json` (repo) — each provider's published language list, cited + as-of + `tier`. The map's green is two tiers by exact ISO-639-3 code: bright = a deployed service lists it (Google/Microsoft/DeepL/LibreTranslate); dim = only an open research model lists it (NLLB/OPUS/M2M-100/MADLAD-400 — a model-card code, not a usable service). "Covered" is a published-list claim, never a quality claim.
|
|
232
|
+
- **Languages**: Loaded from local language card JSON files (7,900+ languages)
|
|
233
|
+
- **Scored runs**: public `run_cards` PostgREST (read-only RLS; aggregates only) — prefer the `get_results` / `get_run_card` tools
|
|
234
|
+
- **Queue ranking**: map-value survey ordering (default) + ECV — see https://champollion.dev/docs/network/specifications/queue-construction
|